[Docs](/docs) / People

# People

Beta

`people_search` finds the public profile pages of people: a person's own website and the platform profiles it points to. Every result carries only what the page publishes about its subject, the identity Wikidata knows for them, and the id of the rule that decided the page is a profile.

## What `people_search` is

It is a keyword search over profile pages, by name, role, employer or topic.

**Query.** The query is free text of 1 to 500 characters. Relevance is computed over the profile's name, which is weighted highest, the page title, the job title, the employer, the affiliation and the page text.

**Results.** `num_results` is 1 to 50, and 10 by default. Each result is one profile page with its fields, a snippet and a score.

**Domains.** `include_domains` keeps only profiles hosted on those domains, and `exclude_domains` drops them. Subdomains match, and each list takes up to 50 domains. Use `["github.com"]` for GitHub profiles only, or exclude it to see personal sites.

A result is a pointer, not the page. When the agent needs the text behind a profile, it calls `web_fetch` with the result's `url`; the same endpoint serves both tools.

## Which pages are indexed

The index holds a seed list from Wikidata and the pages those seeds declare as theirs. Nothing a crawler happens upon becomes a profile.

**Wikidata seeds.** The seeds are the humans (`P31 = Q5`) that Wikidata lists an official website (`P856`) and an English Wikipedia article for, with their ORCID iD (`P496`) and GitHub username (`P2037`) when it has them. Each seed is the official website's URL plus that identity.

**The person's site.** A weekly job fetches each seed's URL and honours `robots.txt`; a blocked or failed fetch stores nothing. The page is read by the `people.v1` rules below and, when one matches, it is indexed as a profile with the fields it publishes.

**Platform profiles the site declares.** The GitHub, GitLab, ORCID, Mastodon and other profiles a page links as `sameAs` or `rel=me` are returned in `same_as`, pooled with the links Wikidata knows. When the official website Wikidata lists is itself a platform profile, rule P5 reads it as one.

In this release only seeds are indexed: the declared profiles are listed, not fetched, and a page is a profile only if a rule says so. Pages that ask not to be indexed (`noindex`) and pages about several people, such as team and listing pages, are never profiles.

## The rules

The `people.v1` rules apply in order. The first match wins, and its id is stored with the page as `rule_id`, so every result says why it counts as a profile.

| Rule                               | Matches when                                                                                                                                                                                                                                                                                                                                    |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| people.P1Profile page              | The page is a schema.org `ProfilePage` whose `mainEntity` is a named `Person`, inline or by `@id`.                                                                                                                                                                                                                                              |
| people.P2Single Person             | The page has exactly one top-level schema.org `Person` whose `url`, `@id` or `mainEntityOfPage` is the canonical URL, or that sits on the root of a personal site. A `Person` that is only the author of an article or blog post never matches.                                                                                                 |
| people.P3Representative h-card     | The page has exactly one top-level `h-card` whose `u-url` or `u-uid` is the canonical URL, with a `p-name`.                                                                                                                                                                                                                                     |
| people.P4Wikidata official website | The page is a seed that none of the page rules matched, and Wikidata lists it as the person's official website (`P856`). Its fields come from the seed alone: the name, the QID, the ORCID iD and `sameAs`.                                                                                                                                     |
| people.P5Platform profile          | The canonical URL matches a known platform pattern and the page carries that platform's marker: `github.com/{login}` with a `Person` plus `og:type profile`, `gitlab.com/{username}` with a `Person`, `orcid.org/{id}` with a `Person` at that URL, and `/@{user}` on any Fediverse host with an ActivityPub actor link plus `og:type profile`. |

P1, P2, P3 and P5 are read from the page's markup: JSON-LD, microdata, RDFa, microformats and OpenGraph. P4 applies only to a seed that no page rule matched.

**Excluded.** A page with a `noindex` robots directive, as a meta tag or an `X-Robots-Tag` header, is never a profile. Neither is a page with several named `Person` nodes and no `mainEntity`, nor a URL or person on the removal list.

## Fields and identity

Only self-published fields are kept: what the page says about its subject. A result never includes an email address or a phone number, even when the page shows one.

| Field       | Where it comes from                                                                                                                                                                              |
| ----------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| url         | It is the profile page.                                                                                                                                                                          |
| name        | It comes from `name`, from `givenName` and `familyName` when a page has only those, or from the h-card `p-name`.                                                                                 |
| job_title   | It comes from `jobTitle` or the h-card `p-job-title`.                                                                                                                                            |
| works_for   | It comes from `worksFor`, as a string or an `Organization`'s name, or from the h-card `p-org`.                                                                                                   |
| affiliation | It comes from `affiliation`, as a string or an `Organization`'s name.                                                                                                                            |
| description | It comes from `description` or the h-card `p-note`.                                                                                                                                              |
| image       | It comes from `image`, as a URL or an `ImageObject`, from the h-card `u-photo`, or from `og:image`.                                                                                              |
| locality    | It comes from `address.addressLocality`, else `homeLocation`, or from the h-card `p-locality`.                                                                                                   |
| same_as     | It lists other profiles of the same person: the page's `sameAs`, `url` and `rel=me` links, pooled with the seed's, which are the Wikipedia article, the orcid.org record and the GitHub profile. |
| wikidata    | It is the person's Wikidata QID, from the seed list.                                                                                                                                             |
| orcid       | It is the person's ORCID iD, from the seed list (Wikidata `P496`).                                                                                                                               |
| rule_id     | It is the rule that labelled the page, `people.P1` to `people.P5`.                                                                                                                               |
| snippet     | It is a plain-text fragment of the page around your query terms, and it is empty when the page sets `nosnippet`.                                                                                 |
| score       | It is the relevance. Higher is better, and scores compare within one response only.                                                                                                              |

**Identity** is the Wikidata QID when the seed list has one, else the ORCID iD, else the page URL. Profiles are tied together only through the links the page and Wikidata declare: `same_as` is the union of the page's own `sameAs` and `rel=me` links and the seed's. Nothing is inferred from a matching name. A field the page does not publish is absent from the result.

## Call it

`people_search` and `web_fetch` are served on `/people` and, like every tool you have switched on, on `/mcp`.

```
https://mcp.deep.navy/people
```

```
https://mcp.deep.navy/mcp
```

| Tool          | What it does                                                                                                        |
| ------------- | ------------------------------------------------------------------------------------------------------------------- |
| people_search | It runs a keyword search over the indexed profile pages and returns the profile fields, the identity and a snippet. |
| web_fetch     | It returns the profile page itself, as deep.navy stored it, with its earlier versions.                              |

Tool arguments and results are protojson, so field names are `camelCase`. This call was made on production and asks for computer scientists. The first two pages were labelled by P4, so their fields come from the seed: the name, the QID, the ORCID iD and `sameAs`. The third was labelled by P2, a single `Person` on the page, so it also carries the affiliation and image the page publishes. Fields the page does not publish are left out of the JSON.

`people_search`

### Find people by role

Profile pages of computer scientists on university websites, each with the rule that decided the page is a profile and the person's Wikipedia, ORCID and GitHub links.

**Request**

```
{
  "query": "computer scientist",
  "numResults": 3
}
```

**Response**

```
{
  "results": [
    {
      "url": "http://hcil.umd.edu/catherine-plaisant/",
      "name": "Catherine Plaisant",
      "sameAs": [
        "https://en.wikipedia.org/wiki/Catherine_Plaisant",
        "https://orcid.org/0000-0003-4049-5848"
      ],
      "wikidata": "Q23008424",
      "orcid": "0000-0003-4049-5848",
      "ruleId": "people.P4",
      "snippet": "Catherine Plaisant is a Research Scientist Emerita at the [University of Maryland Institute for Advanced Computer Studies](http://www.umiacs.umd.edu/) and a member of the [Human-Computer Interaction Lab",
      "score": 10.120917
    },
    {
      "url": "https://www.uidaho.edu/engr/departments/ece/our-people/faculty/dennis-sullivan",
      "name": "Dennis Michael Sullivan",
      "sameAs": [
        "https://en.wikipedia.org/wiki/Dennis_Michael_Sullivan",
        "https://orcid.org/0000-0001-6536-2641"
      ],
      "wikidata": "Q29387632",
      "orcid": "0000-0001-6536-2641",
      "ruleId": "people.P4",
      "snippet": "Apply\nJoin the Department of Electrical and Computer Engineering to make a difference in our world.",
      "score": 9.015392
    },
    {
      "url": "https://samueli.ucla.edu/people/paul-eggert/",
      "name": "Paul Eggert",
      "affiliation": "Computer Science",
      "image": "https://samueli.ucla.edu/wp-content/uploads/samueli/Paul_Eggert_header.jpg",
      "sameAs": [
        "https://en.wikipedia.org/wiki/Paul_Eggert",
        "https://github.com/eggert"
      ],
      "wikidata": "Q66732288",
      "ruleId": "people.P2",
      "snippet": "Scientist Keeps Digital Clocks Ticking | UC IT Blog](https://cio.ucop.edu/time-zone-king-how-one-ucla-computer-scientist-keeps-digital-clocks-ticking/), March 2022\n- [A New Bill Could Do Away With Daylight",
      "score": 8.760338
    }
  ]
}
```

Captured 2026-10-06 00:47 UTC from production

## Data policy

deep.navy reads public, self-published pages, as they ask to be read.

**Sources.** The sources are a person's own official website, as listed on Wikidata, and the public platform profiles that page itself links as `sameAs` or `rel=me`. There are no social graphs, no data brokers and no pages about a person written by someone else.

**Fields.** Only what the page publishes about its subject is kept: the name, job title, employer, affiliation, description, image, locality and the profiles it links. Email addresses and phone numbers are never kept.

**Robots.** Fetching honours `robots.txt`. A page with `noindex` is never a profile, and a page with `nosnippet` is listed without a snippet.

**Identity.** Identity is the Wikidata QID, then the ORCID iD, then the URL. Both identifiers are public, and the person's Wikidata item already carries them.

**Auditability.** Every result names the rule that labelled it in `rule_id`. `web_fetch` returns the page itself, so an agent can check any field against its source.

## Removal

In this release removals are handled by the deep.navy team; there is no self-service form yet.

To have a profile removed, send a request to `privacy@deep.navy` with the page's URL or the person's Wikidata item. We remove the page from the index and add it to the seed suppression list, so the weekly sync never fetches it again.

A removal covers the URL or the Wikidata item you name. If the person has several listed websites, name each one, or name the item, which covers them all.

## Beta

**Coverage.** Coverage is the Wikidata seed list: people with an official website and an English Wikipedia article. A person without a Wikidata item is not in the index.

**Not in this release.** Fetching the profiles a page declares is not in this release; they are listed in `same_as`, not indexed. Neither are profiles found by crawling, nor a public removal form.

**May change.** The fields, the ranking and the platform patterns P5 knows may change. The rule ids are stable.

---

This is the Markdown copy of https://deep.navy/docs/people.
