deep.navy

Docs / People

People

Beta

people_search finds the public profile pages of people: a person's own website and the platform profiles it points to. Every result carries only what the page publishes about its subject, the identity Wikidata knows for them, and the id of the rule that decided the page is a profile.

What people_search is

It is a keyword search over profile pages, by name, role, employer or topic.

Query. The query is free text of 1 to 500 characters. Relevance is computed over the profile's name, which is weighted highest, the page title, the job title, the employer, the affiliation and the page text.

Results. num_results is 1 to 50, and 10 by default. Each result is one profile page with its fields, a snippet and a score.

Domains. include_domains keeps only profiles hosted on those domains, and exclude_domains drops them. Subdomains match, and each list takes up to 50 domains. Use ["github.com"] for GitHub profiles only, or exclude it to see personal sites.

A result is a pointer, not the page. When the agent needs the text behind a profile, it calls web_fetch with the result's url; the same endpoint serves both tools.

Which pages are indexed

The index holds a seed list from Wikidata and the pages those seeds declare as theirs. Nothing a crawler happens upon becomes a profile.

Wikidata seeds. The seeds are the humans (P31 = Q5) that Wikidata lists an official website (P856) and an English Wikipedia article for, with their ORCID iD (P496) and GitHub username (P2037) when it has them. Each seed is the official website's URL plus that identity.

The person's site. A weekly job fetches each seed's URL and honours robots.txt; a blocked or failed fetch stores nothing. The page is read by the people.v1 rules below and, when one matches, it is indexed as a profile with the fields it publishes.

Platform profiles the site declares. The GitHub, GitLab, ORCID, Mastodon and other profiles a page links as sameAs or rel=me are returned in same_as, pooled with the links Wikidata knows. When the official website Wikidata lists is itself a platform profile, rule P5 reads it as one.

In this release only seeds are indexed: the declared profiles are listed, not fetched, and a page is a profile only if a rule says so. Pages that ask not to be indexed (noindex) and pages about several people, such as team and listing pages, are never profiles.

The rules

The people.v1 rules apply in order. The first match wins, and its id is stored with the page as rule_id, so every result says why it counts as a profile.

Rule Matches when
people.P1
Profile page
The page is a schema.org ProfilePage whose mainEntity is a named Person, inline or by @id.
people.P2
Single Person
The page has exactly one top-level schema.org Person whose url, @id or mainEntityOfPage is the canonical URL, or that sits on the root of a personal site. A Person that is only the author of an article or blog post never matches.
people.P3
Representative h-card
The page has exactly one top-level h-card whose u-url or u-uid is the canonical URL, with a p-name.
people.P4
Wikidata official website
The page is a seed that none of the page rules matched, and Wikidata lists it as the person's official website (P856). Its fields come from the seed alone: the name, the QID, the ORCID iD and sameAs.
people.P5
Platform profile
The canonical URL matches a known platform pattern and the page carries that platform's marker: github.com/{login} with a Person plus og:type profile, gitlab.com/{username} with a Person, orcid.org/{id} with a Person at that URL, and /@{user} on any Fediverse host with an ActivityPub actor link plus og:type profile.

P1, P2, P3 and P5 are read from the page's markup: JSON-LD, microdata, RDFa, microformats and OpenGraph. P4 applies only to a seed that no page rule matched.

Excluded. A page with a noindex robots directive, as a meta tag or an X-Robots-Tag header, is never a profile. Neither is a page with several named Person nodes and no mainEntity, nor a URL or person on the removal list.

Fields and identity

Only self-published fields are kept: what the page says about its subject. A result never includes an email address or a phone number, even when the page shows one.

Field Where it comes from
url It is the profile page.
name It comes from name, from givenName and familyName when a page has only those, or from the h-card p-name.
job_title It comes from jobTitle or the h-card p-job-title.
works_for It comes from worksFor, as a string or an Organization's name, or from the h-card p-org.
affiliation It comes from affiliation, as a string or an Organization's name.
description It comes from description or the h-card p-note.
image It comes from image, as a URL or an ImageObject, from the h-card u-photo, or from og:image.
locality It comes from address.addressLocality, else homeLocation, or from the h-card p-locality.
same_as It lists other profiles of the same person: the page's sameAs, url and rel=me links, pooled with the seed's, which are the Wikipedia article, the orcid.org record and the GitHub profile.
wikidata It is the person's Wikidata QID, from the seed list.
orcid It is the person's ORCID iD, from the seed list (Wikidata P496).
rule_id It is the rule that labelled the page, people.P1 to people.P5.
snippet It is a plain-text fragment of the page around your query terms, and it is empty when the page sets nosnippet.
score It is the relevance. Higher is better, and scores compare within one response only.

Identity is the Wikidata QID when the seed list has one, else the ORCID iD, else the page URL. Profiles are tied together only through the links the page and Wikidata declare: same_as is the union of the page's own sameAs and rel=me links and the seed's. Nothing is inferred from a matching name. A field the page does not publish is absent from the result.

Call it

people_search and web_fetch are served on /people and, like every tool you have switched on, on /mcp.

			https://mcp.deep.navy/people
		
			https://mcp.deep.navy/mcp
		
Tool What it does
people_search It runs a keyword search over the indexed profile pages and returns the profile fields, the identity and a snippet.
web_fetch It returns the profile page itself, as deep.navy stored it, with its earlier versions.

Tool arguments and results are protojson, so field names are camelCase. This call was made on production and asks for computer scientists. The first two pages were labelled by P4, so their fields come from the seed: the name, the QID, the ORCID iD and sameAs. The third was labelled by P2, a single Person on the page, so it also carries the affiliation and image the page publishes. Fields the page does not publish are left out of the JSON.

people_search

Find people by role

Profile pages of computer scientists on university websites, each with the rule that decided the page is a profile and the person's Wikipedia, ORCID and GitHub links.

{
  "results": [
    {
      "url": "http://hcil.umd.edu/catherine-plaisant/",
      "name": "Catherine Plaisant",
      "sameAs": [
        "https://en.wikipedia.org/wiki/Catherine_Plaisant",
        "https://orcid.org/0000-0003-4049-5848"
      ],
      "wikidata": "Q23008424",
      "orcid": "0000-0003-4049-5848",
      "ruleId": "people.P4",
      "snippet": "Catherine Plaisant is a Research Scientist Emerita at the [University of Maryland Institute for Advanced Computer Studies](http://www.umiacs.umd.edu/) and a member of the [Human-Computer Interaction Lab",
      "score": 10.120917
    },
    {
      "url": "https://www.uidaho.edu/engr/departments/ece/our-people/faculty/dennis-sullivan",
      "name": "Dennis Michael Sullivan",
      "sameAs": [
        "https://en.wikipedia.org/wiki/Dennis_Michael_Sullivan",
        "https://orcid.org/0000-0001-6536-2641"
      ],
      "wikidata": "Q29387632",
      "orcid": "0000-0001-6536-2641",
      "ruleId": "people.P4",
      "snippet": "Apply\nJoin the Department of Electrical and Computer Engineering to make a difference in our world.",
      "score": 9.015392
    },
    {
      "url": "https://samueli.ucla.edu/people/paul-eggert/",
      "name": "Paul Eggert",
      "affiliation": "Computer Science",
      "image": "https://samueli.ucla.edu/wp-content/uploads/samueli/Paul_Eggert_header.jpg",
      "sameAs": [
        "https://en.wikipedia.org/wiki/Paul_Eggert",
        "https://github.com/eggert"
      ],
      "wikidata": "Q66732288",
      "ruleId": "people.P2",
      "snippet": "Scientist Keeps Digital Clocks Ticking | UC IT Blog](https://cio.ucop.edu/time-zone-king-how-one-ucla-computer-scientist-keeps-digital-clocks-ticking/), March 2022\n- [A New Bill Could Do Away With Daylight",
      "score": 8.760338
    }
  ]
}

Captured 2026-10-06 00:47 UTC from production

Data policy

deep.navy reads public, self-published pages, as they ask to be read.

Sources. The sources are a person's own official website, as listed on Wikidata, and the public platform profiles that page itself links as sameAs or rel=me. There are no social graphs, no data brokers and no pages about a person written by someone else.

Fields. Only what the page publishes about its subject is kept: the name, job title, employer, affiliation, description, image, locality and the profiles it links. Email addresses and phone numbers are never kept.

Robots. Fetching honours robots.txt. A page with noindex is never a profile, and a page with nosnippet is listed without a snippet.

Identity. Identity is the Wikidata QID, then the ORCID iD, then the URL. Both identifiers are public, and the person's Wikidata item already carries them.

Auditability. Every result names the rule that labelled it in rule_id. web_fetch returns the page itself, so an agent can check any field against its source.

Removal

In this release removals are handled by the deep.navy team; there is no self-service form yet.

To have a profile removed, send a request to privacy@deep.navy with the page's URL or the person's Wikidata item. We remove the page from the index and add it to the seed suppression list, so the weekly sync never fetches it again.

A removal covers the URL or the Wikidata item you name. If the person has several listed websites, name each one, or name the item, which covers them all.

Beta

Coverage. Coverage is the Wikidata seed list: people with an official website and an English Wikipedia article. A person without a Wikidata item is not in the index.

Not in this release. Fetching the profiles a page declares is not in this release; they are listed in same_as, not indexed. Neither are profiles found by crawling, nor a public removal form.

May change. The fields, the ranking and the platform patterns P5 knows may change. The rule ids are stable.