[Docs](/docs) / Fetch

# Fetch

GA

`web_fetch` reads HTML and text URLs as markdown, plain text or the original response body. Every fetch is stored, so `web_versions` lists what a page looked like over time and `web_diff` shows what changed between two versions.

## Call it

The three tools are served on `/fetch` and, like every tool you have switched on, on `/mcp`. `/news` and `/people` also serve `web_fetch`, so an agent can open a result on the same endpoint.

```
https://mcp.deep.navy/fetch
```

```
https://mcp.deep.navy/mcp
```

| Tool                          | What it does                                                                                                 |
| ----------------------------- | ------------------------------------------------------------------------------------------------------------ |
| [web_fetch](#web_fetch)       | It fetches HTML and text URLs as markdown, text or the original response body, latest or any stored version. |
| [web_versions](#web_versions) | It lists the stored versions of a URL, newest first.                                                         |
| [web_diff](#web_diff)         | It diffs two stored versions of a URL.                                                                       |

## Cache, live and versions

A page is served from the index when a stored copy is fresh enough, and fetched live otherwise.

**Freshness.** `web_fetch` takes up to 20 URLs in one call. A stored copy is served when it is younger than `maxAgeSeconds`; with 0, the default, any stored copy that may be served is used, and a URL with no stored copy is fetched live. Each result says which it got in `source`, `cache` or `live`.

**Formats.** `formats` takes `FORMAT_MARKDOWN`, the default, `FORMAT_TEXT` and `FORMAT_HTML`, the original response body. When `maxCharacters` is omitted or 0, each requested format is capped at 50,000 characters per URL. A positive value overrides that cap. A capped result has `truncated` set; the response cap does not shorten stored versions.

**Content types.** HTML pages are extracted into main-content markdown and plain text. Other `text/*` and JSON responses are returned as decoded text in `FORMAT_TEXT` and `FORMAT_MARKDOWN`. PDFs and other binary responses return a per-URL error.

**Versions.** Every fetch is stored, and a new version is written only when the main content changes, so page chrome and ad churn do not create versions. Pass `version.versionNo` for a version by number or `version.asOf` for the version that was current at a time, and `web_fetch` returns it from storage.

Live fetches honour `robots.txt`, and a stored page is served only if its site allowed the fetch and did not set `noarchive`.

## web_fetch

GA

It returns one result per URL with the content in each format you asked for, what the page declares about itself, where it came from, and the version you got.

### What comes back

| Field                                         | Type              | Meaning                                                                                                                                      |
| --------------------------------------------- | ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
| results                                       | ContentsResult\[] | They are one result per URL you passed, in the same order.                                                                                   |
| results.​url                                  | string            | It is the URL as you passed it.                                                                                                              |
| results.​markdown                             | string            | It is the page's main content as markdown, returned for `FORMAT_MARKDOWN`, the default.                                                      |
| results.​html                                 | string            | It is the full original page exactly as fetched, returned for `FORMAT_HTML`.                                                                 |
| results.​text                                 | string            | It is the main content as plain text, returned for `FORMAT_TEXT`.                                                                            |
| results.​truncated                            | bool              | It is true when a format was cut at `maxCharacters`.                                                                                         |
| results.​version.​no                          | int32             | It is the version number, starting at 1.                                                                                                     |
| results.​version.​firstSeenAt                 | timestamp         | It is when this content was first fetched.                                                                                                   |
| results.​version.​lastSeenAt                  | timestamp         | It is when a fetch last found this same content.                                                                                             |
| results.​version.​isLatest                    | bool              | It is true for the newest version and left out for older ones.                                                                               |
| results.​version.​totalVersions               | int32             | It is how many versions of the URL are stored.                                                                                               |
| results.​version.​contentHash                 | string            | It is the hash of the main content. A new version is stored only when this hash changes.                                                     |
| results.​metadata.​title · author · publisher | string            | They are what the page declares about itself.                                                                                                |
| results.​metadata.​publishedAt                | timestamp         | It is the publication time the page declares.                                                                                                |
| results.​metadata.​categories                 | string\[]         | They are the labels the index gave the page, such as `web`, `news` or `fetched`.                                                             |
| results.​metadata.​discoveredVia              | string            | It is how the page entered the index, such as `ccnews` for the CC-NEWS crawl or `fetch` for a `web_fetch`.                                   |
| results.​metadata.​person                     | Person            | It is the profile a people page publishes: name, job title, employer, affiliation, `sameAs` links and the `people.v1` rule that labelled it. |
| results.​source                               | string            | It is `cache` when the stored copy was served and `live` when the page was fetched now.                                                      |
| results.​error                                | string            | It says why the URL could not be fetched or served.                                                                                          |

Unset optional fields are omitted; explicitly set zero and false values are preserved. Empty lists are explicit arrays, and 64-bit integers are strings, as protojson writes them.

`web_fetch`

### Read a news article as markdown

A press release found by news_search, served from the cache with its author, publisher, publication time and the crawl that discovered it.

**Request**

```
{
  "urls": [
    "https://www.biospace.com/press-releases/exthymic-secures-up-to-25m-arpa-h-contract-to-enable-hospital-based-cell-therapy-production-prototype"
  ],
  "maxCharacters": 1500
}
```

**Response**

```
{
  "results": [
    {
      "url": "https://www.biospace.com/press-releases/exthymic-secures-up-to-25m-arpa-h-contract-to-enable-hospital-based-cell-therapy-production-prototype",
      "markdown": "SAN DIEGO--(BUSINESS WIRE)--Exthymic Corporation (Exthymic) today announced that the Advanced Research Projects Agency for Health (ARPA-H), an agency within the U.S. Department of Health and Human Services, has awarded Exthymic a contract to develop a cell therapy device to address key issues underlying patient access and affordability.\n\nThe ARPA-H Scalable Cell Therapies through Optimized Phenotyping and Enrichment (SCOPE) project will provide Exthymic with up to $25 million of funding over two years to develop and test LITTLESTAR, a next generation device that enables the safe, reliable production of engineered cell therapies like CAR-T cell therapy by managing patient variability during the production process.\n\n*Using a patient's own cells is what makes cell therapies powerful. It's also what makes them so hard to produce,\"* says Rich Stoner, Exthymic CEO*. \"The industry has spent a decade optimizing production, but it still can't tell you what makes one patient's cells work and another's fail. So it settles for a process it can get approved, and then decides which patients that process will accept. We won't solve access or scale commercially that way. Simply put, we're failing patients.\"*\n\nThe LITTLESTAR device functions by directly measuring and reacting to differences in cell phenotype over a full production process. This adaptive control ensures each final dose is produced safely and reliably without the overhead of a dedicated clean room or quality control laboratory,",
      "truncated": true,
      "version": {
        "no": 1,
        "firstSeenAt": "2026-10-05T15:50:53.368246060Z",
        "lastSeenAt": "2026-10-05T15:50:53.368246060Z",
        "isLatest": true,
        "totalVersions": 1,
        "contentHash": "45ae47750ef11fff18e07732d689e1663c057dae3542c2c7b780998920dd1304"
      },
      "metadata": {
        "title": "Exthymic Secures up to $25M ARPA-H Contract to Enable Hospital-Based Cell Therapy Production Prototype",
        "author": "Gabrielle Masson",
        "publishedAt": "2026-10-05T13:48:04.298Z",
        "publisher": "BioSpace",
        "categories": [
          "web",
          "news"
        ],
        "discoveredVia": "ccnews"
      },
      "source": "cache"
    }
  ]
}
```

Captured 2026-10-05 17:28 UTC from production

`web_fetch`

### Read an earlier version of a page

Version 1 of the Hacker News front page, served from storage exactly as it was first seen, while a newer version exists.

**Request**

```
{
  "urls": [
    "https://news.ycombinator.com/"
  ],
  "version": {
    "versionNo": 1
  },
  "maxCharacters": 1200
}
```

**Response**

```
{
  "results": [
    {
      "url": "https://news.ycombinator.com/",
      "markdown": "Hacker News\nnew\n|\npast\n|\ncomments\n|\nask\n|\nshow\n|\njobs\n|\nsubmit\nlogin\n1.\nBorland Turbo Basic\n(\ndosdays.co.uk\n)\n68 points\nby\nsssilver\n1 hour ago\n|\nhide\n|\n35 comments\n2.\nWeb Search API\n(\ncloudflare.com\n)\n255 points\nby\ntosh\n6 hours ago\n|\nhide\n|\n125 comments\n3.\nMaking a GTK application in Haskell, part 1\n(\nfloreal.tech\n)\n45 points\nby\nVosporos\n2 hours ago\n|\nhide\n|\n2 comments\n4.\nDenmark Data Breach Exposes 8.8M People's Personal Data\n(\ncpr.dk\n)\n355 points\nby\nclan\n8 hours ago\n|\nhide\n|\n261 comments\n5.\nMold Linker Version 3.0.0 Release – Rewritten in Rust\n(\ngithub.com/rui314\n)\n98 points\nby\nroflcopter69\n5 hours ago\n|\nhide\n|\n37 comments\n6.\nThe technology to eradicate mosquito-borne disease exists\n(\nworksinprogress.co\n)\n129 points\nby\nbenbreen\n2 hours ago\n|\nhide\n|\n104 comments\n7.\nThe future of independence is interdependence\n(\nonlys.ky\n)\n16 points\nby\neustoria\n1 hour ago\n|\nhide\n|\n8 comments\n8.\nPixel 11 doesn't yet meet the GrapheneOS security standards and may be skipped\n(\ngrapheneos.org\n)\n277 points\nby\nfinnlab\n3 hours ago\n|\nhide\n|\n152 comments\n9.\nHuawei and Qualcomm Announce Broad Patent License Agreement\n(\nhuawei.com\n)\n129 points\nby\n0xedb\n9 hours ago\n|\nhide\n|\n79 comments\n10.\nType Safe Generic D",
      "truncated": true,
      "version": {
        "no": 1,
        "firstSeenAt": "2026-10-05T17:01:32.444631267Z",
        "lastSeenAt": "2026-10-05T17:01:32.444631267Z",
        "totalVersions": 2,
        "contentHash": "f6e5a2f1e943d0bfc483d87895fefa73514a951089a969110704163c7bda37de"
      },
      "metadata": {
        "title": "Hacker News",
        "publishedAt": "2026-10-05T00:00:00Z",
        "publisher": "news.ycombinator.com",
        "categories": [
          "web",
          "fetched"
        ],
        "discoveredVia": "fetch"
      },
      "source": "cache"
    }
  ]
}
```

Captured 2026-10-05 17:28 UTC from production

## web_versions

GA

It lists a URL's stored versions, newest first, with when each was first and last seen, its size, and how many lines changed from the version before. `limit` takes up to 100.

### What comes back

| Field                               | Type              | Meaning                                                                                                              |
| ----------------------------------- | ----------------- | -------------------------------------------------------------------------------------------------------------------- |
| versions                            | VersionSummary\[] | They are the stored versions of the URL, newest first.                                                               |
| versions.​info.​no                  | int32             | It is the version number, starting at 1.                                                                             |
| versions.​info.​firstSeenAt         | timestamp         | It is when this content was first fetched.                                                                           |
| versions.​info.​lastSeenAt          | timestamp         | It is when a fetch last found this same content.                                                                     |
| versions.​info.​isLatest            | bool              | It is true for the newest version and left out for older ones.                                                       |
| versions.​info.​totalVersions       | int32             | It is how many versions of the URL are stored.                                                                       |
| versions.​info.​contentHash         | string            | It is the hash of the main content. A new version is stored only when this hash changes.                             |
| versions.​title                     | string            | It is the page's title in that version.                                                                              |
| versions.​bytes                     | int64             | It is the size of the stored version in bytes. As an int64 it is a string in the JSON.                               |
| versions.​linesAdded · linesRemoved | int32             | They are the lines of main content added and removed since the version before; version 1 counts every line as added. |

Unset optional fields are omitted; explicitly set zero and false values are preserved. Empty lists are explicit arrays, and 64-bit integers are strings, as protojson writes them.

`web_versions`

### List the versions of a page

Every stored version of the Hacker News front page, newest first, with when each was seen and how many lines changed from the one before.

**Request**

```
{
  "url": "https://news.ycombinator.com/",
  "limit": 10
}
```

**Response**

```
{
  "versions": [
    {
      "info": {
        "no": 2,
        "firstSeenAt": "2026-10-05T17:05:51.357118150Z",
        "lastSeenAt": "2026-10-05T17:05:51.357118150Z",
        "isLatest": true,
        "totalVersions": 2,
        "contentHash": "638d25ca9026f28f5dbcfb08f9f87c65a9501fd57a5497d4470cfcdabd50d373"
      },
      "title": "Hacker News",
      "bytes": "34685",
      "linesAdded": 97,
      "linesRemoved": 97
    },
    {
      "info": {
        "no": 1,
        "firstSeenAt": "2026-10-05T17:01:32.444631267Z",
        "lastSeenAt": "2026-10-05T17:01:32.444631267Z",
        "totalVersions": 2,
        "contentHash": "f6e5a2f1e943d0bfc483d87895fefa73514a951089a969110704163c7bda37de"
      },
      "title": "Hacker News",
      "bytes": "34683",
      "linesAdded": 425
    }
  ]
}
```

Captured 2026-10-05 17:28 UTC from production

## web_diff

GA

It returns the unified diff of the main-content markdown between `fromVersion` and `toVersion`, with the counts of lines added and removed.

### What comes back

| Field        | Type   | Meaning                                                                                 |
| ------------ | ------ | --------------------------------------------------------------------------------------- |
| unifiedDiff  | string | It is the unified diff of the main-content markdown, from `--- v<from>` to `+++ v<to>`. |
| linesAdded   | int32  | It is the number of lines added.                                                        |
| linesRemoved | int32  | It is the number of lines removed.                                                      |

Unset optional fields are omitted; explicitly set zero and false values are preserved. Empty lists are explicit arrays, and 64-bit integers are strings, as protojson writes them.

`web_diff`

### Compare two versions of a page

A unified diff of the main content between versions 1 and 2 of the Hacker News front page, where points and comment counts moved.

**Request**

```
{
  "url": "https://news.ycombinator.com/",
  "fromVersion": 1,
  "toVersion": 2
}
```

**Response**

```
{
  "unifiedDiff": "--- v1\n+++ v2\n@@ -18,33 +18,33 @@\n (\n dosdays.co.uk\n )\n-68 points\n+70 points\n by\n sssilver\n 1 hour ago\n |\n hide\n |\n-35 comments\n+36 comments\n 2.\n Web Search API\n (\n cloudflare.com\n )\n-255 points\n+260 points\n by\n tosh\n 6 hours ago\n |\n hide\n |\n-125 comments\n+127 comments\n 3.\n Making a GTK application in Haskell, part 1\n (\n floreal.tech\n )\n-45 points\n+46 points\n by\n Vosporos\n 2 hours ago\n@@ -57,36 +35,36 @@\n (\n cpr.dk\n )\n-355 points\n+358 points\n by\n clan\n 8 hours ago\n |\n hide\n |\n-261 comments\n+262 comments\n 5.\n Mold Linker Version 3.0.0 Release – Rewritten in Rust\n (\n github.com/rui314\n )\n-98 points\n+100 points\n by\n roflcopter69\n 5 hours ago\n |\n hide\n |\n-37 comments\n+39 comments\n 6.\n The technology to eradicate mosquito-borne disease exists\n (\n worksinprogress.co\n )\n-129 points\n+133 points\n by\n benbreen\n-2 hours ago\n+3 hours ago\n |\n hide\n |\n@@ -96,33 +50,33 @@\n (\n onlys.ky\n )\n-16 points\n+17 points\n by\n eustoria\n 1 hour ago\n |\n hide\n |\n-8 comments\n+9 comments\n 8.\n Pixel 11 doesn't yet meet the GrapheneOS security standards and may be skipped\n (\n grapheneos.org\n )\n-277 points\n+279 points\n by\n finnlab\n-3 hours ago\n+4 hours ago\n |\n hide\n |\n-152 comments\n+153 comments\n 9.\n Huawei and Qualcomm Announce Broad Patent License Agreement\n (\n huawei.com\n )\n-129 points\n+131 points\n by\n 0xedb\n 9 hours ago\n@@ -135,20 +68,20 @@\n (\n danielchasehooper.com\n )\n-102 points\n+103 points\n by\n AlexeyBrin\n 8 hours ago\n |\n hide\n |\n-57 comments\n+58 comments\n 11.\n Differences Between \\`Foldl\\` and \\`Foldr\\`\n… [trimmed: 2,844 more characters]",
  "linesAdded": 97,
  "linesRemoved": 97
}
```

Captured 2026-10-05 17:28 UTC from production

Read [provenance and trust](/docs/provenance) for source timestamps, historical limitations and handling untrusted text. See [limits and errors](/docs/limits) for request receipts and retry behavior.

---

This is the Markdown copy of https://deep.navy/docs/fetch.
