# AI and MCP Every archive Archilyzer builds is machine-navigable by design. That is not a bolt-on: the same paginated JSON the site's own search reads is public, described by an index that tells a client how to navigate it, and served with the CORS headers that let anything fetch it. ## The published contract Every archive serves two files that exist for machines: - **`/corpus.json`** — a machine-readable index. It does not contain transcripts; a corpus of tens of thousands of videos cannot be enumerated per-video without running into a host's file-count limits. Instead it *documents how to navigate the shards*: where each channel's manifest is, how to turn a video's id into a page number, and what a record looks like when you get there. It also names the software that built the archive. - **`/llms.txt`** — the same thing as prose, following the llmstxt.org convention, for a client that reads a page rather than an API. Between them, an LLM with a fetch tool can navigate a whole archive without scraping a single HTML page, and without any code of ours running on its side. ## The MCP server Included in the source is an MCP server that exposes an archive — a single site, or several at once — to Claude Code, Claude Desktop, Cursor, or any other MCP client. It is **a local tool you run yourself**. It changes nothing: it only reads already-published static JSON, either from a directory on disk or over HTTP. The one exception is `fetch_clip`, which asks a local Archilyzer editor for a clip's media: the editor writes it, and the server itself never writes. ### Ten-minute setup To run Claude Code against a published archive, with nothing of your own hosted — the Jeralyzer here, and any archive's URL works the same: ```sh git clone https://archilyzer.pages.dev/source/archilyzer.git archilyzer # or the tarball on /downloads/ cd archilyzer && pnpm install claude mcp add archilyzer \ --env TRANSCRIPT_SITE_URL=https://jeralyzer.pages.dev \ -- pnpm --silent -C "$PWD" archilyzer mcp claude # then: /ask what has he said about … ``` - **What you need:** Node.js 20.9 or newer, pnpm 9 or newer, and Claude Code. No corpus, no yt-dlp, no GPU, nothing hosted. - **The source** is the project's read-only [git mirror](/source/); without git, the same tree is a [tarball](/downloads/): unpack it and carry on from `cd archilyzer`. - **Register it as `archilyzer`.** The shipped `/ask` and `/sweep` commands call `mcp__archilyzer__ask_plan` / `mcp__archilyzer__sweep_plan`, and that tool name embeds the server name as you registered it. - **Keep `--silent`.** Claude Code reads the server's replies on its standard output, and without it some pnpm versions print a line of their own there first. `archilyzer mcp` is the source's own command for starting the server. - **Clips:** two optional `--env` lines, `ARCHILYZER_EDITOR_URL` and `WORKER_TOKEN` (the editor's own), let `fetch_clip` ask a local editor for clip media; leave them out for research alone. - **One archive or several:** `TRANSCRIPT_SITE_URL` reads one archive; `TRANSCRIPT_HUB_URL`, given a hub's URL, federates every archive on the hub. - **Another client** (Claude Desktop, Cursor): the same server as an `mcp.json` entry, `"command": "pnpm"` with `"args": ["--silent", "-C", "/ABS/PATH/TO/archilyzer", "archilyzer", "mcp"]`, is in [mcp/README.md](https://archilyzer.pages.dev/source/tree/mcp/README.md). - **On Windows,** run all of this inside WSL2, Claude Code included: see “Claude Code on Windows” in the [README](https://archilyzer.pages.dev/source/tree/README.md). ### What it can do - **Search transcripts** for a term, a phrase or a regular expression, with timestamped snippets. Every timestamp is a link to that exact second of the recording, in the archive's own viewer. - **Filter before scanning** — by availability state, upload date range, media type, or scope. This is not just convenience: a filtered query computes exactly which pages it needs before reading a single transcript byte, which is the difference between reading a whole corpus and reading a fraction of it. - **Enumerate a query's complete match set** as a worklist in one pass, so an agent can cover everything rather than reporting a number it took from the first page. - **Batch-read** many videos at once, as bounded excerpt windows around a query rather than whole transcripts. - **Report everything known about one video** — metadata, engagement counts, whether other copies of the same recording are archived, and how much of its runtime the transcript actually covers. ## The honesty features, which are the point Anything that lets a model summarise an archive can also let it summarise the archive *wrongly*, confidently. Several behaviours exist specifically to make that harder: - **A partial result says so first, not in a footnote.** If a scan hits a cap, the first line of the response says the result is a sample and names which channels went unread. - **Counts are of recordings, not uploads.** When the same recording exists on two platforms, it collapses to one row — and says how many it collapsed, and which. A count that double-counts mirrors looks exactly like a correct one. - **A truncated transcript is flagged as a warning**, not reported as a percentage, because a transcript covering 40% of a video will otherwise support a confident conclusion that something was never said. - **A timestamp is never translated across two copies** of a recording unless the two were actually measured as aligned. A mirror with a different intro carries the same words at different times, and a translated citation would look perfectly plausible while pointing at the wrong moment. ## In the browser, without any of this The published site also carries a chat interface that searches the archive and answers with citations, using an API key the visitor supplies themselves. The key stays in their browser; nothing is proxied through the archive. Next: [Questions](/docs/faq/).