AI and MCP
Every archive is machine-navigable. Point Claude Code, Cursor, or a browser chat at it.
Every archive Archilyzer builds is machine-navigable by design. That is not a bolt-on: the same paginated JSON the site's own search reads is public, described by an index that tells a client how to navigate it, and served with the CORS headers that let anything fetch it.
The published contract
Every archive serves two files that exist for machines:
/corpus.json— a machine-readable index. It does not contain transcripts; a corpus of tens of thousands of videos cannot be enumerated per-video without running into a host's file-count limits. Instead it documents how to navigate the shards: where each channel's manifest is, how to turn a video's id into a page number, and what a record looks like when you get there. It also names the software that built the archive./llms.txt— the same thing as prose, following the llmstxt.org convention, for a client that reads a page rather than an API.
Between them, an LLM with a fetch tool can navigate a whole archive without scraping a single HTML page, and without any code of ours running on its side.
The MCP server
Included in the source is an MCP server that exposes an archive — a single site, or several at once — to Claude Code, Claude Desktop, Cursor, or any other MCP client.
It is a local tool you run yourself. It changes nothing: it only reads
already-published static JSON, either from a directory on disk or over HTTP.
The one exception is fetch_clip, which asks a local Archilyzer editor for a
clip's media: the editor writes it, and the server itself never writes.
Ten-minute setup
To run Claude Code against a published archive, with nothing of your own hosted — the Jeralyzer here, and any archive's URL works the same:
git clone https://archilyzer.pages.dev/source/archilyzer.git archilyzer # or the tarball on /downloads/
cd archilyzer && pnpm install
claude mcp add archilyzer \
--env TRANSCRIPT_SITE_URL=https://jeralyzer.pages.dev \
-- pnpm --silent -C "$PWD" archilyzer mcp
claude # then: /ask what has he said about …
- What you need: Node.js 20.9 or newer, pnpm 9 or newer, and Claude Code. No corpus, no yt-dlp, no GPU, nothing hosted.
- The source is the project's read-only git mirror; without
git, the same tree is a tarball: unpack it and carry on from
cd archilyzer. - Register it as
archilyzer. The shipped/askand/sweepcommands callmcp__archilyzer__ask_plan/mcp__archilyzer__sweep_plan, and that tool name embeds the server name as you registered it. - Keep
--silent. Claude Code reads the server's replies on its standard output, and without it some pnpm versions print a line of their own there first.archilyzer mcpis the source's own command for starting the server. - Clips: two optional
--envlines,ARCHILYZER_EDITOR_URLandWORKER_TOKEN(the editor's own), letfetch_clipask a local editor for clip media; leave them out for research alone. - One archive or several:
TRANSCRIPT_SITE_URLreads one archive;TRANSCRIPT_HUB_URL, given a hub's URL, federates every archive on the hub. - Another client (Claude Desktop, Cursor): the same server as an
mcp.jsonentry,"command": "pnpm"with"args": ["--silent", "-C", "/ABS/PATH/TO/archilyzer", "archilyzer", "mcp"], is in mcp/README.md. - On Windows, run all of this inside WSL2, Claude Code included: see “Claude Code on Windows” in the README.
What it can do
- Search transcripts for a term, a phrase or a regular expression, with timestamped snippets. Every timestamp is a link to that exact second of the recording, in the archive's own viewer.
- Filter before scanning — by availability state, upload date range, media type, or scope. This is not just convenience: a filtered query computes exactly which pages it needs before reading a single transcript byte, which is the difference between reading a whole corpus and reading a fraction of it.
- Enumerate a query's complete match set as a worklist in one pass, so an agent can cover everything rather than reporting a number it took from the first page.
- Batch-read many videos at once, as bounded excerpt windows around a query rather than whole transcripts.
- Report everything known about one video — metadata, engagement counts, whether other copies of the same recording are archived, and how much of its runtime the transcript actually covers.
The honesty features, which are the point
Anything that lets a model summarise an archive can also let it summarise the archive wrongly, confidently. Several behaviours exist specifically to make that harder:
- A partial result says so first, not in a footnote. If a scan hits a cap, the first line of the response says the result is a sample and names which channels went unread.
- Counts are of recordings, not uploads. When the same recording exists on two platforms, it collapses to one row — and says how many it collapsed, and which. A count that double-counts mirrors looks exactly like a correct one.
- A truncated transcript is flagged as a warning, not reported as a percentage, because a transcript covering 40% of a video will otherwise support a confident conclusion that something was never said.
- A timestamp is never translated across two copies of a recording unless the two were actually measured as aligned. A mirror with a different intro carries the same words at different times, and a translated citation would look perfectly plausible while pointing at the wrong moment.
In the browser, without any of this
The published site also carries a chat interface that searches the archive and answers with citations, using an API key the visitor supplies themselves. The key stays in their browser; nothing is proxied through the archive.
Next: Questions.