# yt-dlp-transcript-mcp An [MCP](https://modelcontextprotocol.io) server that exposes a transcript archive — a single site, or a federated hub of sites — to Claude Code, Claude Desktop, Cursor, and any other MCP client. It is a **local tool you run yourself**. It does not change the archive: it only reads the site's already-published static JSON shards (`corpus.json` + `transcripts//…`), either from disk or over HTTP. Nothing is hosted for you. The exceptions all ask a local Archilyzer editor (`ARCHILYZER_EDITOR_URL` + `WORKER_TOKEN`, the editor's own) to do the work — `fetch_clip` fetches a clip window, `enqueue` queues archival work, `get_job` / `channel_coverage` read the editor's jobs and disk, and `get_transcript` falls back to the editor's disk for a video not yet published; `notes` reads umtool (`UMTOOL_URL`). The MCP itself still writes nothing. Settings, storage and deletes are not reachable from it (`pnpm ops` only). ## Tools | Tool | What it does | |------|--------------| | `list_channels` | List channels **organized under their channel groups** (name, slug, video count; site in hub mode), with a compact group cheat-sheet (`id · name · N channels`) for scoping. | | `list_tags` | The **curated tags** the archive publishes (the operator's cross-channel vocabulary, what `tags` on `search_transcripts` / `enumerate_matches` filters by): each tag's id, label, group, how many videos carry it on this source, and the per-channel breakdown. Not the yt-dlp keywords `scopes:["tags"]` searches. A source that publishes none says so. | | `list_reports` / `get_report` | The cited reports a site publishes (corpus spec 5): the index (title, kind, claim and citation counts, a fact-check's verdict tally), then one report's sections and claims with every citation's **verbatim quote, original URL and the site's moment page**. See [Cited reports](#cited-reports). | | `search_transcripts` | Search captions for a term/phrase (or regex); returns matching videos with timestamped snippets — **each `[mm:ss]` is a clickable link to that exact moment** (or a compact `[mm:ss\|sec]` with `link_style:"base"`). Alias-aware, pageable, and **filterable** (`states`, `date_from`/`date_to`, `media_type`, `age`, `exclude`, `scopes`). A page that isn't the whole match set is flagged **above** the hits. | | `enumerate_matches` | A query's **complete** match set as a worklist (id/title/channel/date + batch count) in **one scan**. Takes the **same filters** as `search_transcripts`, so the two can never disagree about coverage. The tool to use whenever you need to count or cover everything. | | `get_transcripts` | Batch-read up to 20 videos in one call — bounded, timestamped **excerpt windows** around one query or up to 8 (`queries`), with per-query counts; or full transcripts without a query. Reads posts too. | | `get_transcript` | One video's full transcript as clean markdown (metadata + **linked** timestamped captions). A video the archive does not hold yet (imported or transcribed since the last build) is read off the **local editor's disk** when one is configured (`GET /api/ops/transcript`: a fresh `transcript.cues.json`, else the raw transcript normalized in memory, else the English VTT), marked as such, with no moment links. | | `get_post` / `get_thread` | One archived social post, or its whole thread. Posts have no timeline — cite them with no `@ mm:ss`. | | `get_video_metadata` | Everything known about one video without the transcript body: metadata, plus **view/like counts, cue count and transcript coverage** (`stats/`), **other archived copies of the same recording** with an explicit timings-aligned verdict (`duplicates.json`), and **AI chapters/tags** where they exist (`digests/`). | | `fetch_clip` | The media behind a cited moment, **fetched by the local editor** (`POST /api/media/fetch-window`) through its paced, cookie-aware, provenanced job — never a yt-dlp run by hand. Needs `ARCHILYZER_EDITOR_URL` (default `http://localhost:3001`) and `WORKER_TOKEN` (the editor's own) in this server's env; without them it says so and fetches nothing. The editor must already archive the cited channel (a channel dir under its `transcripts/`), else it answers 404 `Channel "" not found`: an MCP pointed at a public site with a fresh editor gets that on every clip. A window is the cited span ± `pad` (default 3 s), at most 15 min, and lands at `channels//data//clips/`; `full: true` fetches the whole recording into the saved-video store (needs a video the editor already knows). `maxHeight` (144–2160) caps the source height: a window is fetched at or under it (default 720); a whole recording at 720 or less is saved as the editor's 720p H.264 preset and above 720 at the original quality (omitted, the channel's source-video quality applies). A file already on disk is returned as it is, never re-fetched for a different cap, and the answer gives its height and says when it is taller than asked. Waits up to `wait_seconds` (default 90, max 300), then returns the job id to resume with `job`; a client with a 60 s default request timeout must raise it or pass `wait_seconds` ≤ 50 — the fetch continues on the editor either way; resume it with `job`, and once it has finished the same request finds it cached. While it waits it sends one progress notification per poll to a client that asked for progress (a `progressToken`), which keeps a reset-on-progress timeout alive. A Rumble embed id is mapped to the editor's slug id through the record's `webpageUrl`, so pass the citing corpus as `source`; a video not in `source` is passed through as cited (known limitation). The file is a read-only corpus artifact. | | `get_job` | One editor job's state — kind, channel, status, times, exit code, where it waits in its queue — and the last `tail` (default 40) log lines (`GET /api/ops/job/`). For the job id `enqueue` or `fetch_clip` returned. Needs the editor. | | `enqueue` | Queue archival work on a channel the editor archives, as its page's buttons do (`POST /api/ops/`): `sync` (`full`), `download-missing`, `retry-bucket` (`bucket`, `ids` — the way to download chosen videos), `transcribe-bucket` (`ids`), `fetch-posts` (`full` / `older`), `import-video` (`url`). Answers the job id; the editor runs it on its own queues at each platform's pace. Only when the operator asks. Needs the editor. | | `channel_coverage` | What a channel **holds** on the editor's disk by date (`GET /api/ops/coverage`): count, first and last day, per year, and every gap longer than `gap_days` (30); `date_from` / `date_to` narrow it, `list` lists the videos. Each video is dated by its **recorded date** when the channel has a title rule (a VOD mirror, release 20), else its upload date. Needs the editor. | | `notes` | The operator's notes in umtool on articles and report-video projects: `list` (every open note, by target), `read` (`target` as `list` names it — the digest with anchors resolved and the source file to edit). `reply` answers with the `umtool notes reply` command, because umtool records a reply sent over HTTP as the operator's. Needs `UMTOOL_URL`. | | `open_link` | Paste an archilyzer viewer **share link** to re-run that exact search here (query tree + every filter, at full fidelity) — plan, results and corpus handle in **one** call. `dry_run:true` for the plan alone. | | `list_sources` | Show the **default** corpus and, with a hub, its member sites as ready-to-paste handles. A site with reports says how many; a cited-only site says it is one. | | `resolve_source` | Turn a URL or site name into the canonical `source` handle and check it can be read. Changes nothing. | | `sweep_plan` / `ask_plan` | Turn a plain-English request (plus an optional pasted link) into a resolved, step-by-step plan. What `/sweep` and `/ask` call. | **Clickable moment links.** Every timestamp the read tools emit is a Markdown link to the exact second — an **archilyzer viewer** deep link (`…/?v=&t=`, opening the transcript modal at the moment) when the source has a public origin, otherwise the video's platform watch page with a per-platform time param (YouTube `&t=s`, Odysee/Twitch equivalents). A `--local` source now **cites the site it was composed for**, not the platform: a composed public dir is not an anonymous pile of JSON, it names its own deployed origin in `corpus.json` (`site.url`), and that is read the first time the channel list is loaded. Following a citation therefore lands in the archive — at the cited second, with the transcript around it and the neighbouring videos one click away — instead of on the platform page, where the archive's whole point (that a copy still exists here) is invisible. Set `TRANSCRIPT_PLATFORM_LINKS=1` for the old behaviour, which is the right choice when a local build's declared site URL is not actually deployed. A dir with no `corpus.json` still falls back to platform links. **Compact base links (`link_style:"base"`).** Inline links are ~70–90 chars *per line* — bulk an agent pipeline shouldn't pay for. `search_transcripts` and `get_transcripts` accept `link_style:"base"`: each video gets **one** `- moment_base:` header line (a URL ending in `t=`) and every stamp becomes the compact `[mm:ss|]`. The expansion rule: **full moment link = ``** — append the integer after the `|`, e.g. `[title @ 2:36](156)`. The seconds are floored exactly like the inline links', so both styles cite the identical second. A video whose base can't be built (no viewer origin and a platform whose time param doesn't take raw seconds — Twitch — or doesn't exist — Rumble/Kick/archive.org/BitChute) omits the line: cite its `- source:` URL plain instead. `open_link` results and `get_transcript` stay inline-linked (future work). ### `search_transcripts` Beyond `query`, `regex`, and `limit`: - **Scope** — restrict the scan to a subset of the corpus. Both scope fields are optional and **additive** (the search runs over the union); with none, it covers everything. - **`channels`** — a list of channel slugs/names. - **`groups`** — a list of channel groups, matched by **group id or display name** (case-insensitive), expanded to the channels in that group. e.g. `groups:["other"]` and `groups:["Extended Universe"]` resolve the same set. (In hub mode groups are deferred — a group token reports "unknown" while channel scoping still works.) - The singular `channel`/`group` are no longer advertised (the engine always unioned them anyway) but are still parsed, so an old habit doesn't break. The footer names the resolved scope (e.g. *"scope: group Extended Universe (7 channels)"* or *"scope: 6 channels"*) and flags any channel/group token that matched nothing — so a typo is surfaced, not silently a whole-corpus scan. - **`offset`** (default 0) — skip this many matches. A page that is not the whole match set is headed with an unmissable **`⚠ INCOMPLETE PAGE — N total, showing a–b. Do NOT report a count from this page.`** (the footer's `total`/`has_more` are still there for existing consumers). **Do not page this to cover a query — use `enumerate_matches`**: the engine materialises the whole match set and *then* slices, so paging costs one full corpus scan per page while enumerating costs one, total. Paging was the documented advice and it produced a report claiming 319 videos swept after seeing 200. - **`include_snippets`** (default true) — set `false` for a cheap worklist (id / title / channel / date / match count, no cue text — the per-video `- source:` line is dropped too). Ideal for the planning pass of a sweep. - **`link_style`** (default `"inline"`) — `"base"` switches to the compact agent-pipeline form: a `- moment_base:` line per hit and `[mm:ss|]` snippet stamps (see *Compact base links* above). - **`use_aliases`** (default true) — expand the query through the site's curated search aliases. A plain query that matches a curated trigger also searches the alias's regex, so mis-transcribed spellings are caught (e.g. `k cups` also matches `cake cup`). The footer reports which aliases fired, e.g. *"expanded via alias K-Cups → `(k|cake)[ -]?cup`"*. Explicit `regex` queries are used verbatim (no expansion). - **`max_pages`** (default 400) — scan cap. If reached (or a very common term passes the 2000-video cap), the footer flags coverage as **PARTIAL** rather than silently truncating — and **names which channels were fully scanned and which were never reached**. The cap cuts in channel iteration order, never at random, so "partial" without that list hides the shape of the bias: the sample is simply whatever sorted first. With the list, it is actionable — re-run scoped to the remainder. #### Filters (also on `enumerate_matches`) These are the share-link filters, reachable without a pasted link. Both search tools take **all of them**, from one shared schema constant, because a filter reachable from one and not the other would make the two disagree about coverage — exactly the failure the stateless rebuild set out to make impossible. - **`states`** — keep only videos in these presence states: `available`, `maybe_missing`, `deleted`, `private`, `members_only`, `unlisted`. This is how you ask the corpus's signature question — *what did the videos that have since been removed say about X* — which previously had no path at all. - **`date_from`** / **`date_to`** — inclusive upload-date bounds, `YYYYMMDD`. A differently-shaped date is **reported and ignored** rather than compared lexicographically into a wrong answer. - **`media_type`** — `video` or `livestream`. **`age`** — `all_ages` or `restricted`. - **`exclude`** — video-level NOT: drop any video that also contains one of these terms. This is `"cup"` but not `"world cup"`. Compiled the same way the query is (regex too) but **never** alias-expanded — an exclusion stays exactly as narrow as you wrote it. Arbitrary boolean trees remain `open_link`'s job; that is what a share link is for. - **`scopes`** — which layers to match: `transcripts`, `chat`, `description`, `tags`, `metadata` (title + channel), `posts`. Omitted means captions plus the title, exactly as before. Naming scopes **replaces** that default, so `scopes:["description"]` searches descriptions and *not* captions. Non-timed layers emit `[description]`-tagged snippets with **no timestamp link**, since citing a description line as `@ 0:00` would assert someone said it. `chat` reads the separate ~1.2 GB live-chat shards lazily — scope it. - **`collapse_duplicates`** (default **true**) — see *Counts are of recordings*. An unrecognised token in any of these is **named in the footer** rather than dropped silently: a typo'd state would otherwise widen the search back to the whole corpus and return a perfectly legitimate-looking answer. ### `get_transcripts` Reads a whole batch of videos with one round-trip. `video_ids` (max 20; extras dropped and noted). With a **`query`** (alias-aware, like `search_transcripts`; optional `regex`/`use_aliases`), each transcript is reduced to bounded windows of timestamped lines around the matches (`before`/`after` seconds, default 30) — high-signal context for folding a batch into a report. Without a query, each video's full transcript comes back as markdown. Missing ids are reported inline. - **`queries`** (max 8, unioned with `query`) — window around several terms in **one** pass. The windows merge per video, and each video's header reports a **per-query count**, so a term that matched nothing in that video is visible rather than absorbed into the merged excerpt. Use this instead of re-reading the same videos once per term. - **`content_types`** — `"video"` and/or `"post"` (default both). An id that isn't a video falls through to the post corpus, so a mixed batch works. `channel`/`channels` are optional owning-channel hints that speed the per-id lookup when a batch spans several channels (the ids are already scoped by the search that produced them). - **`link_style`** (default `"inline"`) — `"base"` emits one `- moment_base:` line per video and compact `[mm:ss|]` stamps instead of a full Markdown link per line (see *Compact base links* above). - **`max_lines`** (default 200) — cap on merged excerpt lines per video when a query is given (the earliest lines are kept). Ignored without a query. ### `open_link` — paste a viewer share link, at full fidelity The archilyzer viewer's **Share** button produces a URL that encodes the whole search: the origin, a `qt=` composite **query tree** (any mix of scopes — transcripts, live-chat, title/channel, description, tags — combined with AND/OR/NOT), and the filter block (`fc` channels, `ft` type, `fa` audience, `fav` availability, `fdf`/`fdt` upload-date range). `open_link` re-runs that exact search here — no manual source-switching or query reconstruction, nothing lost. **One call does the whole job.** `open_link` decodes the link, resolves its origin to a corpus handle (hub vs single-site, auto-probed from `corpus.json`), validates the channel scope against that live corpus, and returns the plan, the first page of **linked** results, and the handle to pass as `source` from then on. It switches nothing — the old `apply:true` mutated a global active source, which is exactly the footgun this rebuild removed. - **The plan** names the resolved corpus handle, the query tree rendered readably, every active filter, the validated channel scope, and anything ignored or warned (the `fk` subtitle-track token is vestigial in the composite share model — decoded and reported as ignored). - **`dry_run: true`** returns the plan *without* searching, for confirming a link's scope first. - **Adjust** — re-call with structured `overrides` to honor a natural-language edit: `clear_availability` (drop the availability filter), `clear_type`, `clear_age`, `clear_dates`, `clear_filters`, `clear_channels`, `channels:[…]` (re-scope), `date_from`/`date_to`, or `query`/`regex`/ `query_scope` (replace the search; `query_scope:"posts"` targets the social-post corpus — declared but silently dropped until now). - Pageable via `limit`/`offset`. Read-only throughout. ``` open_link link="https://rekietalyzer.pages.dev/?qt=…&fv=1&fc=Rekieta%20Law&ft=v&fav=a" open_link link="…" dry_run=true # plan only open_link link="…" overrides={ "clear_availability": true } # "drop the availability filter" ``` ## `/sweep` and `/ask` Two entry points that turn **Claude Code itself** into the corpus sweep engine — the same "batch matching transcripts into a running report" the browser does with a BYO AI key, but driven by your Claude **plan usage** (no API key), with the report written to a file. ``` /sweep https://hasanalyzer.pages.dev/?qt=…&fav=deleted This search finds deleted videos. Why might have Rekieta privated these? batch_size=12 /ask what did they actually say about the settlement? channels=rekieta-law ``` Type the whole thing on one line, in your own words, punctuation intact. Paste a share link anywhere in it. **Why a tool and not a prompt argument.** Claude Code parses an MCP *prompt*'s arguments as `text.trim().split(/\s+/)` zipped against the declared argument names. It is not quote-aware, the last argument does **not** absorb the remainder, and tokens past the declared count are **dropped silently**. With the nine arguments `sweep` declares, the request above became `link="This"`, `channel="search"`, `group="deleted"`, `directive="videos."`, `batch_size="Why"` — rendered literally into `` ceil(N / Why) `` — and everything from *"Rekieta"* onward vanished with no warning. So `.claude/commands/sweep.md` and `ask.md` are thin shims that pass `$ARGUMENTS` — the **entire raw argument string, untokenised** — to the `sweep_plan` / `ask_plan` tools, whose arguments are structured JSON and arrive intact. The MCP `sweep` prompt still exists for form-based clients (Claude Desktop, Cursor), where each argument gets its own field; it routes through the same parser and validators. **It cannot be used as the shredding trap it used to be.** In Claude Code that prompt still appears in the slash list as `/mcp____sweep`, so it now **refuses** rather than sweeping the wrong thing: a typed argument holding prose (`link` = `"search"`, `batch_size` = `"might"`) is the tokenizer's fingerprint, and the prompt answers with what it detected, the words that survived in the order they were typed, and the `/sweep` line to use instead. A form client only trips it by genuinely typing a bad value — which deserves the error too. Note the tool path stays forgiving (default + `⚠`); only the prompt path is strict. **Settings.** Append `key=value` for any of `channels=`, `groups=`, `batch_size=`, `parse_model=`, `report=`, `directive=`, `source=`, `content_types=`, `regex=` (quote a multi-word value). The whitelist is closed: an unrecognised `x=y` **stays in the question** and warns rather than being eaten, with a typo hint if it is one edit away from a real key. Values are validated — `batch_size` an integer 1–20, `parse_model` a single token, `report` a single `.md` path with no `..` — and a bad one falls back to the default with a `⚠` line at the top of the plan, never into the arithmetic. **What the plan does.** The tools pre-resolve what they can — canonicalising the corpus handle, validating your channel/group tokens against the live corpus, inlining the group roster when you gave no scope — so a three-call preamble becomes none. Then: **pick the scope** if none was given (never a silent whole-corpus sweep) → **`enumerate_matches` once** for the complete worklist → **state N and `ceil(N / batch_size)` batches** → **per batch, map-reduce** → finish with a summary that says how many of the N were actually read. An unresolvable channel halts the plan instead of quietly widening it. The MCP stays read-only; only the report file is written. **Dumb extractors on the cheap model.** Each batch is handed to a **subagent** (Claude Code's Task tool) spawned as a *verbatim extractor* on the cheapest model — the plan requests `parse_model` (default `haiku`) via the Task tool's model override, and gracefully spawns on the default when no override exists. The extractor calls `get_transcripts` with `link_style:"base"` and returns *only* the directive-relevant lines, **verbatim** — grouped per video under its `moment_base:` line, in their compact `[mm:ss|sec]` form, under a hard budget of **≤40 lines (~600 words) per batch**. No analysis, no summarizing: the heavy transcript bulk lives and dies inside the cheap subagent. The **orchestrator does all the synthesis** — claims, contradictions, cross-referencing — expanding each kept citation to a full link by appending the seconds to the video's `moment_base`, then discards the fragment. Every subagent is given the corpus handle explicitly, since one that omitted it would read the server default and quote the wrong archive. Batches are independent, so several can run in parallel. **Linked citations.** Every source is cited as a clickable `[title @ mm:ss]()` link — built by appending the cited integer seconds to the video's `moment_base` from the tool output, so a click seeks to the exact second (see *Compact base links* above; a video with no base is cited by its `source:` URL). A post has no timeline, so it is cited `[post by , ]()` and never with `@ mm:ss`. The sweep discipline lives in `src/instructions.ts` — one builder shared by both tools and the prompt — rather than duplicated into the markdown command files, so it cannot drift out of sync with the tools it names. A test asserts every tool the instructions mention actually exists. ## Counts are of recordings, not uploads Some videos exist in the archive more than once — the same recording mirrored to another platform, or re-uploaded on another channel. The site already ships the detector's output as `duplicates.json`; the MCP never opened it, so a sweep counted the same recording twice and said "N videos". `search_transcripts` and `enumerate_matches` now **collapse cluster members to one row by default**, and report it: *"12 mirror(s) collapsed across 9 cluster(s) — the 43 above are distinct RECORDINGS, not uploads"*. Two rules keep that safe: - **Nothing is hidden.** The collapsed copies are **named on the row they fold into** (`⧉ same recording also archived as: …`). Pass `collapse_duplicates:false` for one row per upload. - **The surviving copy is never dropped.** The kept row is the cluster's canonical member *only when that member is itself among the matches*; otherwise it is simply the first match. A mirror is frequently the only surviving copy of a deleted upload, and preferring an absent canonical would delete precisely the evidence a `states:["deleted"]` question is asking for. **This changes reported totals**, so a report generated before and after will disagree. That is a fix, not a regression: the earlier number was double-counting mirrors. `get_video_metadata` lists a video's other copies, and states for each whether the two were **measured as aligned**. If they were not — including when alignment was simply never measured — it says so and tells you not to map a timestamp across. A mirror with a different intro carries the same words at shifted times, so a translated citation would look perfectly plausible and point at the wrong moment of a different upload. Absent means *not measured*, and not measured means *no*. ## Cited reports A site may publish cited reports (corpus spec 5): `/reports/index.json` lists them and `/reports//page.json` is one report with every citation resolved — the same files the site's Reports pages read. `list_reports` reads the index; `get_report` (`report`, optionally one `section`) renders a report as text: its verdict tally, then each claim with its verdict, its findings and, per citation, the quote, the original (the platform at the cited second, the post, the document) and the site's moment page (`/m///-/`). A **cited-only** site (`corpus.json` `site.scope: "cited"`; built from a site with search off, `site.json` `search: false`) publishes its reports and the moments they cite and nothing else: no channels, no transcripts, nothing to search. `list_channels`, `list_sources` and `resolve_source` say "cited-only site: N report(s)" there. Reports are per site: a hub has none of its own, so name a member with `source:"remote:"`. ## How it reads the corpus Three changes, in increasing order of how much they buy: **1. Nothing is read twice.** Channel manifests (~551 KB for the whole corpus) are cached outright; shard pages go in a **byte-budgeted LRU**, default 48 MB of raw page bytes, configurable with `TRANSCRIPT_MCP_PAGE_CACHE_MB`. The budget is denominated in bytes rather than entries because page sizes differ by an order of magnitude across corpora, and 48 MB is chosen from measurement: a parsed page retains about **2.7× its file bytes**, so the ceiling is ~130 MB resident. Caches hold **promises**, so concurrent callers for the same page coalesce onto one read. `list_channels({refresh:true})` drops all of them — the explicit escape hatch for a corpus rebuilt under a long-lived server, deliberately not a TTL. What this fixes: a 20-id `get_transcripts` batch used to re-read the same 7.4 MB page once per id and, without a channel hint, do ~300 manifest reads. It is now one read per distinct page. **2. Pages are read concurrently**, in windows clamped to the remaining `max_pages` budget and folded back in page order — so hit ordering is unchanged and `max_pages:1` still reads exactly one page. Local sources use a window of 4 (reads overlap parses; parse dominates locally), remote and hub 8 (latency dominates). **3. Filter-first scanning — the big one.** `summaries/` is a **global index of every video** (slug, channel, title, upload date, livestream, age-restricted, presence state) that parses in well under a second, against ~1.3 GB and tens of seconds for the transcripts. The MCP already read it and threw away everything except the availability state. It now keeps the whole record, and because each channel manifest carries `slugToPage`, **a filtered query's exact page set is computable before a single transcript byte is read**. The invariant that makes this safe: **the index only prunes pages; the record predicate still decides every hit.** Both call the same `passesFilters`, so they cannot drift, and a video the index has never heard of gets its page read unconditionally. A stale, partial or missing summaries set therefore costs time, never correctness. It pays in proportion to how selective the filter is, and it says which path ran (*"filter-pruned: planned 8 of 170 page(s)"*), so a slow query is explicable. An unfiltered query never reads the index at all — a filter that excludes nothing is treated as no filter, so planning can never make a whole-corpus scan slower. ### Benchmark `mcp/bench/` drives the **real server over stdio**, through `pnpm --filter … exec tsx src/index.ts` — the command `archilyzer mcp`, which a client is registered with, runs — and times a fixed query set: ```bash pnpm --filter yt-dlp-transcript-mcp bench pnpm --filter yt-dlp-transcript-mcp bench -- --repeat 3 --json out.json ``` It reports two kinds of number and treats them differently. **Pages read and bytes parsed are structural** — properties of the query plan, identical on an idle box and a hammered one, so a before/after comparison of them is always valid. **Wall time is contingent**: this box is shared, so the bench checks the load average first and, above `--max-load` (default 0.7/core), **refuses to run** unless given `--force`, in which case every wall figure is stamped `UNRELIABLE` in both the table and the JSON. A number taken under load cannot later be quoted as if it weren't. Every run prints a **corpus fingerprint** first, because the composed dir gets rebuilt and a before/after that silently spans two corpora is worse than no measurement at all. **The run this work was built against** — corpus `local:export/public`, 5 channels, 170 transcript pages, 1,288 MB, 3,330 videos in `summaries/`; 3 repetitions, median, taken at 0.64 load/core (i.e. trusted — no `UNRELIABLE` stamp): | query | wall ms | page reads | MB parsed | |---|---|---|---| | `list_channels` (cold) | 1 | 0 | 0 | | rare term, whole corpus | 16,967 | 170 | 1,261 | | common term, whole corpus | 16,446 | 170 | 1,288 | | common term, one channel | 42 | 4 | 27 | | common term + one upload year | 2,647 | **32** | **202** | | common term + missing states only | **823** | **8** | **63** | | `enumerate_matches`, whole corpus | 23,422 | 170 | 1,288 | | `get_transcripts` × 20 ids, one channel | 24 | 4 | 27 | Read it as three facts. **Filter-first is worth ~21× on a selective filter** (8 pages against 170) and ~5× on a one-year date range — and the comparison to make is within this same run: the removed-videos question costs 0.82 s where the only previously-available way to ask it, an unfiltered scan, costs 16.4 s. **The unfiltered rows are unchanged at 170 reads**, which is the point — planning must never make a whole-corpus scan slower. **The 20-id batch is 4 reads**, cold; it was up to 20 re-reads of the same page. *On the prebuilt-index question* (deferred until the speed win was a number): these numbers say **not yet**. Scoped and filtered questions — the ones people actually ask — now land between 0.04 s and 2.6 s. Only the unfiltered whole-corpus scan is still slow, and that is a sweep's one-off first step before it goes on to read hundreds of transcripts. If it ever does become the bottleneck, the right artifact is a narrow token → `(channel, page)` postings list emitted at export time: it would prune pages exactly the way the filter planner already does, which makes that machinery its prerequisite rather than its competitor. ## Data source (pick one) Resolved from flags or env — precedence hub > remote > local: | Mode | Flag | Env | Notes | |------|------|-----|-------| | Hub | `--hub ` | `TRANSCRIPT_HUB_URL` | Federate across every member of a hub's `corpus.json`. | | Remote | `--remote ` | `TRANSCRIPT_SITE_URL` | One deployed site origin. | | Local | `--local ` | `TRANSCRIPT_LOCAL_DIR` | A composed public dir on disk (default `./export/public`). | ## Choosing a corpus per call The server is added to your MCP client **once**, pointed at a startup source (above) — its **default** corpus. To read something else, pass a `source` handle on the call. There is no "active source": nothing is switched, and nothing is remembered between calls. A handle **is** the serialised spec, in a canonical round-trippable form: | Handle | Means | |---|---| | `default` | the startup source — the same as omitting `source` | | `local:/dir` | a composed public dir on disk | | `remote:https://site` | one deployed site origin | | `hub:https://hub` | federate every member of a hub | | `hub:https://hub#alpha,beta` | a hub **subset** (`#`, so it can never collide with a query string) | Shorthands are accepted and normalised: a **bare site or hub URL** (classified by probing its `corpus.json`), a hub member's **siteId or title**, `site:`/`sites:`, and `./dir`. The canonical form is echoed back, so you learn it as you go. **Every result names the corpus it read**, as a trailing `(corpus: )` — errors included. It is `corpus:` rather than `source:` because `- source:` already means "this video's URL" in the output. ``` list_channels # the default corpus list_channels source="remote:https://rekietalyzer.pages.dev" search_transcripts query="k cups" source="hub:https://archilyzer-hub.pages.dev#jeralyzer,rekietalyzer" resolve_source source="jeralyzer" # → remote:https://jeralyzer.pages.dev ``` **Why it works this way.** The server used to hold a mutable active source and persist it to a state file under `$XDG_STATE_HOME/yt-dlp-transcript-mcp`. That was silently wrong: a server registered with `--local …/export/public` had `{"activeSpec":{"kind":"remote","url":"https://hasanalyzer.pages.dev"}}` on disk from some earlier session, so **every call read the wrong archive and no result said so**. Protocol revision 2026-07-28 removes protocol-level sessions and directs servers needing cross-call state to use "explicit, server-minted handles passed as ordinary tool arguments" — which is exactly the fix. Any leftover state files are now inert and can be deleted. `use_source` survives one release as an unadvertised alias that resolves its target and tells you the handle to pass; `reset_source` is gone (there is nothing to reset). **Still read-only.** A `source` handle only changes *which* already-published static shards are read — the same capability the startup flags already grant this locally-run tool. This server never writes to any corpus; `fetch_clip` and `enqueue` ask the editor, and the editor writes. ## Protocol The server speaks MCP revision **2026-07-28** through `@modelcontextprotocol/server@2`'s `serveStdio`, which owns the era decision: the opening exchange selects `modern` (2026-07-28, negotiated via `server/discover`) or `legacy` (the 2025 `initialize` handshake), and one server instance is pinned for the connection. Both eras serve the identical tools — `src/protocol.test.ts` spawns the real process and asserts it on each. `resultType` is stamped and stripped by the SDK. Cache hints (`ttlMs` / `cacheScope`) are a construction-time policy: `tools/list`, `prompts/list` and `server/discover` are literal constants carrying no corpus data, so they are `public` for an hour, and every read tool's result stays uncached. An invalid hint throws at startup rather than on the wire. The low-level `Server` is used deliberately (it is marked "advanced use" rather than removed): literal JSON Schema tool definitions keep the zod dependency out of the package entirely. ## Run it From the monorepo root, through the repo's CLI. `archilyzer mcp` (`common/bin/mcp.ts`) starts the server exactly as `pnpm --filter yt-dlp-transcript-mcp exec tsx src/index.ts` does — mcp's own `tsx`, cwd this package's dir (so the TypeScript path alias resolves), the environment and every argument after `mcp` passed through: ```sh # a deployed site pnpm --silent archilyzer mcp --remote https://rekietalyzer.pages.dev # a hub, federating every member site pnpm --silent archilyzer mcp --hub https://archilyzer-hub.pages.dev # local shards on disk (a relative path resolves from mcp/, so give an absolute one) pnpm --silent archilyzer mcp --local "$PWD/export/public" ``` Logs go to stderr; stdout is the MCP JSON-RPC channel. **Keep `--silent`** (spelled out — `claude mcp add` has its own `-s`): `pnpm archilyzer` is a `pnpm run`, and without it pnpm 9 and 10 print their `> …` script banner to stdout before the server starts, which a client reads as broken protocol (pnpm 11 prints `$ …` to stderr). With it, nothing precedes the protocol on 9, 10 or 11 — on an installed checkout: after a pull that changed the lockfile, pnpm 11 may install first and print that to stdout whatever the flags, so run `pnpm install` first. ## Add to Claude Code ```sh claude mcp add archilyzer \ --env TRANSCRIPT_SITE_URL=https://rekietalyzer.pages.dev \ --env ARCHILYZER_EDITOR_URL=http://localhost:3001 \ --env WORKER_TOKEN=… \ -- pnpm --silent -C /ABS/PATH/TO/archilyzer archilyzer mcp ``` The two editor lines are optional: they let `fetch_clip` ask a local editor for clip media (`WORKER_TOKEN` is the editor's own, from `editor/.env`; the URL defaults to `http://localhost:3001`). Leave them out and every other tool works the same; `fetch_clip` then says it has no editor and fetches nothing. **The `/sweep` and `/ask` commands.** `.claude/commands/{sweep,ask}.md` in this repo call `mcp__archilyzer__sweep_plan` / `mcp__archilyzer__ask_plan` — the tool name embeds the MCP server name **as you registered it**, which is why every example registers it as `archilyzer`. If you use another name (a site's own, say `rekietalyzer`), change the `mcp____` prefix in those two files to match. `foo:bar` namespacing is plugin-only, so what you type stays `/sweep`, not `/archilyzer:sweep`. ## Add to any MCP client (mcp.json) ```json { "mcpServers": { "archilyzer": { "command": "pnpm", "args": [ "--silent", "-C", "/ABS/PATH/TO/archilyzer", "archilyzer", "mcp" ], "env": { "TRANSCRIPT_SITE_URL": "https://rekietalyzer.pages.dev", "ARCHILYZER_EDITOR_URL": "http://localhost:3001", "WORKER_TOKEN": "…" } } } } ``` Swap `TRANSCRIPT_SITE_URL` for `TRANSCRIPT_HUB_URL` to federate a whole hub, or `TRANSCRIPT_LOCAL_DIR` to read a local build. `ARCHILYZER_EDITOR_URL` and `WORKER_TOKEN` are optional, for `fetch_clip` only, as above.