Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit b57494480c474b532ec1ad852222112abc101caa
parent 8e7d1c401aa11edc0f8fcd1d75cd6789fb92d47d
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Tue, 11 Aug 2026 08:47:45 -0400

Merge feat/mcp-fast-reachable-honest: filters are reachable, scans are planned, counts are of recordings

The MCP was correct about provenance after the stateless rebuild and
naive about everything else. Three things a real sweep hit immediately:

Every search cost a full-corpus parse — no caching anywhere in the three
transports, sequential page reads, so a 20-id batch re-read the same
7.4 MB page once per id.

The corpus's signature question could not be asked. The filter engine had
exactly one call site, open_link, so "what did the DELETED videos say"
needed a pasted share link.

Counts were inflated and biased — duplicates.json unread, and the video
cap truncating in channel iteration order behind a total that read like a
real number.

Measured idle, one run: the removed-videos question goes from a
170-page/1288 MB scan to 8 pages/62.6 MB, 0.82s against the 16.4s the
same term costs unfiltered. Unfiltered scans are unchanged, which is the
constraint — planning must never make a whole-corpus scan slower.

Also: a --local server now cites the archive it was composed for rather
than YouTube, which is what a citation is for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Diffstat:
Mexport/CHANGELOG.md | 10++++++++++
Mmcp/README.md | 191++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++---
Amcp/bench/bench.ts | 556+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Amcp/bench/smoke.ts | 196+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mmcp/package.json | 4+++-
Mmcp/src/instructions.test.ts | 3+++
Mmcp/src/instructions.ts | 33+++++++++++++++++++++++++++++++--
Amcp/src/scanPlan.test.ts | 528+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mmcp/src/search.ts | 630+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++--------------
Mmcp/src/server.ts | 644+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++--
Mmcp/src/source.ts | 976+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++------------
11 files changed, 3501 insertions(+), 270 deletions(-)

diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md @@ -1,5 +1,15 @@ # Changelog +## [0.8.6] - 2026-08-11 +- **MCP: the corpus's signature question is now askable, and it answers in under a second.** `runSearchSpec` — the engine that supports availability-state, upload-date, media-type and age filters, plus the chat/description/tags scopes — had **exactly one call site**: the `open_link` handler. Every filter was therefore unreachable unless a human pasted a viewer share link, which meant *"what did the videos that have since been DELETED say about X"* — the question that motivated the last sweep — had no path at all. `search_transcripts` **and** `enumerate_matches` now take the filters as flat arguments from **one shared schema constant**: `states` (available / maybe_missing / deleted / private / members_only / unlisted), `date_from`/`date_to`, `media_type`, `age`, `scopes` (transcripts / chat / description / tags / metadata / posts) and **`exclude`** for video-level NOT — `"cup"` but not `"world cup"`, the one boolean case that actually bites. Both tools gain them **together**, deliberately from the same constant, because a filter reachable from one and not the other would make the two disagree about coverage, which is precisely the failure the stateless rebuild set out to make impossible. Arbitrary boolean trees stay `open_link`'s job — that is what a share link is *for*, and asking a model to author a `qt=` tree in a tool call would trade a real capability for a new class of malformed input. Unrecognised tokens are **named in the footer** rather than dropped: a typo'd state would otherwise widen the search back to the whole corpus and return a perfectly legitimate-looking answer. Non-timed layers (description / tags / channel name) emit `[description]`-tagged snippets with **no timestamp link**, since citing a description line as `@ 0:00` would assert that someone said it. +- **MCP: a filtered query now reads a fraction of the corpus instead of all of it.** `summaries/` is a **global index of every video** — slug, channel, title, upload date, livestream, age-restricted, presence state — that parses in well under a second, against ~1.3 GB and tens of seconds for the transcripts. The MCP already read it in `availabilityMap()` and **threw away everything except the presence state**. It now keeps the whole record, and because each channel manifest carries `slugToPage`, **a filtered query's exact page set is computable before a single transcript byte is read**. Measured on the live 170-page corpus, on an idle box, in one run: the removed-videos question drops from a **170-page / 1,288 MB** full scan to **8 pages / 62.6 MB — 0.82 s against the 16.4 s** the same term costs unfiltered; a one-year date range to 32 pages / 202 MB (2.6 s). The invariant that makes it safe is that **the index only ever prunes pages; the record predicate still decides every hit** — both call the same `passesFilters`, so they cannot drift, and a video the index has never heard of gets its page read unconditionally. A stale, partial or missing summaries set therefore costs time, never correctness. It pays in proportion to how selective the filter is (a broad attribute spread across the corpus gets almost nothing) and **says which path ran** — *"filter-pruned: planned 8 of 170 page(s)"* — so a slow query is explicable. An unfiltered query never reads the index at all: a filter that excludes nothing is treated as no filter, so planning can never make a whole-corpus scan slower. +- **MCP: nothing is read twice any more.** `transcriptsManifest` and `transcriptPage` had **zero caching in all three transports**, and every page read was a sequential `await`. So a 20-id `get_transcripts` batch re-read the same 7.4 MB page **once per id**, and without a channel hint did ~300 manifest reads; a sweep paid that per batch. Manifests (~551 KB for the whole corpus) are now cached outright; pages go in a **byte-budgeted LRU** — default 48 MB of raw page bytes, `TRANSCRIPT_MCP_PAGE_CACHE_MB` to change it. Denominated in bytes rather than entries because page sizes differ by an order of magnitude across corpora, and 48 MB is chosen from measurement rather than taste: a parsed page retains about **2.7× its file bytes**, so the ceiling is ~130 MB resident — which matters on a box that also runs a GPU digest sweep. Caches hold **promises**, so concurrent callers for the same page coalesce onto one read. Pages are also read **concurrently** now, in windows clamped to the remaining `max_pages` budget and folded back in page order — so hit ordering is unchanged and `max_pages:1` still reads exactly one page. A 20-id batch: **550 ms cold, 24 ms warm, 4 page reads**. `list_channels({refresh:true})` drops every new cache too — the explicit escape hatch for a corpus rebuilt under a long-lived server, still deliberately not a TTL. +- **MCP: counts are of recordings, not uploads — so some totals will now be lower.** The site has shipped `duplicates.json` (the cross-platform mirror detector's output) all along and the MCP never opened it, so a sweep counted the same recording twice whenever it was mirrored to another platform or re-uploaded on another channel. `search_transcripts` and `enumerate_matches` now **collapse cluster members to one row** and report it: *"12 mirror(s) collapsed across 9 cluster(s) — the 43 above are distinct RECORDINGS, not uploads"*. Nothing is hidden — the collapsed copies are **named on the row they fold into**, and `collapse_duplicates:false` gives one row per upload. And the surviving copy is never dropped: the kept row is the cluster's canonical member **only when that member is itself among the matches**, otherwise simply the first match, because a mirror is frequently the only surviving copy of a deleted upload and preferring an absent canonical would delete exactly the evidence a `states:["deleted"]` question is asking for. **A report generated before and after this will disagree on totals. That is a fix, not a regression** — the earlier number was double-counting mirrors. +- **MCP: "coverage partial" now tells you which channels went unread.** `HARD_VIDEO_CAP` truncates in **channel iteration order**, never at random, so a capped result was a channel-biased sample whose `total` read like a real count (measured: a common term returned exactly 2000). The partial banner now names the channels that were **fully scanned**, the one it **stopped inside** and at which page, and the ones it **never reached** — turning "partial" from alarming into actionable, since you can re-run scoped to the remainder. +- **MCP: `get_video_metadata` returns everything the archive knows about a video.** It returned the transcript record minus cues; three shipped layers it never opened are now joined in. From `stats/`: view/like/comment counts, cue count, platform state, and **transcript coverage** — surfaced not as a float but as a warning when it is low (*"⚠ TRANSCRIPT COVERS ONLY 41% OF THE RUNTIME … do NOT conclude from this transcript that something was never said"*), because a truncated download is a correctness trap disguised as metadata. From `duplicates.json`: the other archived copies of the same recording, each stating **whether the two were measured as aligned** — and when they were not, *including when alignment was simply never measured*, saying so and refusing to map a timestamp across, since a mirror with a different intro carries the same words at shifted times and a translated citation would look perfectly plausible while pointing at the wrong moment of a different upload. From `digests/`: AI chapters and topic tags where they exist (120 of ~31,000 videos — sparse enough that a subsystem would be over-building), with a **borrowed** digest flagged loudly as describing the *other* upload. Every one of these layers is **optional** at the interface level and degrades to a stated absence, because they genuinely are optional in the published contract: `compose-site.ts` only writes `duplicates.json` when there is a publishable cluster, digests exist for a handful of channels, and `corpus.json` doesn't even declare `duplicates.json` or `stats/`. A site that ships none of them behaves exactly as before. +- **MCP: a local corpus now cites the archive it was built for, instead of sending you to YouTube.** `LocalSource.publicOrigin()` returned `null` unconditionally, so the server registered here — which runs `--local …/export/public` — fell back to platform watch pages for **every** citation. A composed public dir is not an anonymous pile of JSON: it names its own deployed origin in `corpus.json` (`site.url`), which is now read the first time the channel list is loaded. Citations land in the archive, at the cited second, with the transcript around it and the neighbouring videos one click away — rather than on the platform page, where the archive's whole point (that a copy still exists *here*) is invisible. Verified live: `https://hasanalyzer.pages.dev/?v=FearAnd%2FV037tgaMBBI&t=1553`. `TRANSCRIPT_PLATFORM_LINKS=1` restores the old behaviour, which is the right choice when a local build's declared site URL is not actually deployed; a dir with no `corpus.json` still falls back to platform links. +- **MCP: a benchmark, so the next speed claim is a number.** New `mcp/bench/` drives the **real server over stdio** through the same command line the client is registered with, and times a fixed query set. It reports two kinds of number and treats them differently, because this box is shared and a build was running during the session this work started (load average 27): **pages read and bytes parsed are structural** — properties of the query plan, identical on an idle box and a hammered one — while **wall time is contingent**, so the bench checks the load average first and **refuses to run** above `--max-load` (default 0.7/core) unless forced, in which case every wall figure is stamped `UNRELIABLE` in both the table and the JSON. A number taken under load cannot later be quoted as if it weren't. Every run prints a **corpus fingerprint** first, because the composed dir gets rebuilt — one rebuild landed mid-session and swapped a 30,923-video composition for a 3,330-video one — and a before/after that silently spans two corpora is worse than no measurement at all. A companion `pnpm smoke` runs the same server against the real corpus and asserts the invariants that only 1.3 GB of real shards can break, including that `enumerate_matches` and `search_transcripts` report the **same deduped total**. See `mcp/bench/{bench,smoke}.ts`, `mcp/src/{source,search,server,instructions}.ts`, `mcp/src/scanPlan.test.ts`. + ## [0.8.5] - 2026-08-11 - **MCP: the server no longer remembers which corpus you're reading — because remembering it was silently getting it wrong.** `use_source` switched a mutable "active corpus" and persisted the choice to a state file so it survived reconnects. That was the bug. The server registered here runs `--local …/export/public`, but its state file held `{"activeSpec":{"kind":"remote","url":"https://hasanalyzer.pages.dev"}}` from some earlier session — so **every call since had been reading a different archive, and nothing in any result said so**. There is now no active source and nothing is persisted: **every read tool takes its own `source` handle**, and a call that omits it reads the server's startup corpus. The handle *is* the serialised spec in canonical form — `default`, `local:/dir`, `remote:https://site`, `hub:https://hub`, or `hub:https://hub#alpha,beta` for a subset — not an opaque token, so it survives a restart and a human reading one in a transcript knows exactly what was searched. Shorthands (a bare site or hub URL, probed to tell one from the other; a hub member's siteId or title) normalise to canonical and are echoed back. **Every result now ends with `(corpus: <handle>)`** — errors included, implemented once in the dispatch wrapper so a new tool cannot forget it; it is `corpus:` and not `source:` because `- source:` already means "this video's URL" in the output. `use_source` survives one release as an unadvertised alias that resolves a target and tells you the handle to pass; `reset_source` is gone. Leftover state files are inert and can be deleted. Caching moved with it: a source instance is built once per handle and `listChannels` is memoised per instance, which also fixes a pre-existing cost — a 20-id `get_transcripts` batch against a remote used to fetch `corpus.json` twenty times. See `mcp/src/sourceRegistry.ts` (new; `sourceController.ts` deleted), `mcp/src/{server,source,index}.ts`, `mcp/src/sourceRegistry.test.ts`. - **MCP: claiming you covered the corpus is now hard to do by accident.** A sweep returned `total 319; showing 1–200; has_more: yes`, was never paged, and reported **319 videos swept** having seen 200. The paging instruction was already in the sweep prompt and was ignored — so prose is not the enforcement mechanism. Worse, paging was also the *expensive* option: the engine materialises the entire match set and only then slices, so each page re-scanned the whole corpus. New **`enumerate_matches`** returns a query's complete id/title/channel/date worklist plus the batch count in **one scan** — a new tool rather than a flag, because a flag that silently changes the output shape is exactly what gets ignored. If a cap is hit, the **first line** reads `⚠ COVERAGE PARTIAL … this is a SAMPLE, not the full set`, never a quiet footnote. `search_transcripts` now prints `⚠ INCOMPLETE PAGE — N total, showing a–b. Do NOT report a count from this page.` **above** the hits (the old footer stays, so existing consumers keep working). Verified on the real 30-channel corpus: enumerate and search agree exactly at 57, 86 and at the 2000-video cap, where both correctly flag partial coverage. diff --git a/mcp/README.md b/mcp/README.md @@ -13,12 +13,12 @@ reads the site's already-published static JSON shards (`corpus.json` + | Tool | What it does | |------|--------------| | `list_channels` | List channels **organized under their channel groups** (name, slug, video count; site in hub mode), with a compact group cheat-sheet (`id · name · N channels`) for scoping. | -| `search_transcripts` | Search captions for a term/phrase (or regex); returns matching videos with timestamped snippets — **each `[mm:ss]` is a clickable link to that exact moment** (or a compact `[mm:ss\|sec]` with `link_style:"base"`). Alias-aware and pageable. A page that isn't the whole match set is flagged **above** the hits. | -| `enumerate_matches` | A query's **complete** match set as a worklist (id/title/channel/date + batch count) in **one scan**. The tool to use whenever you need to count or cover everything. | +| `search_transcripts` | Search captions for a term/phrase (or regex); returns matching videos with timestamped snippets — **each `[mm:ss]` is a clickable link to that exact moment** (or a compact `[mm:ss\|sec]` with `link_style:"base"`). Alias-aware, pageable, and **filterable** (`states`, `date_from`/`date_to`, `media_type`, `age`, `exclude`, `scopes`). A page that isn't the whole match set is flagged **above** the hits. | +| `enumerate_matches` | A query's **complete** match set as a worklist (id/title/channel/date + batch count) in **one scan**. Takes the **same filters** as `search_transcripts`, so the two can never disagree about coverage. The tool to use whenever you need to count or cover everything. | | `get_transcripts` | Batch-read up to 20 videos in one call — bounded, timestamped **excerpt windows** around one query or up to 8 (`queries`), with per-query counts; or full transcripts without a query. Reads posts too. | | `get_transcript` | One video's full transcript as clean markdown (metadata + **linked** timestamped captions). | | `get_post` / `get_thread` | One archived social post, or its whole thread. Posts have no timeline — cite them with no `@ mm:ss`. | -| `get_video_metadata` | One video's metadata (title, channel, date, duration, description, tags, source URL) without the transcript body. | +| `get_video_metadata` | Everything known about one video without the transcript body: metadata, plus **view/like counts, cue count and transcript coverage** (`stats/`), **other archived copies of the same recording** with an explicit timings-aligned verdict (`duplicates.json`), and **AI chapters/tags** where they exist (`digests/`). | | `open_link` | Paste an archilyzer viewer **share link** to re-run that exact search here (query tree + every filter, at full fidelity) — plan, results and corpus handle in **one** call. `dry_run:true` for the plan alone. | | `list_sources` | Show the **default** corpus and, with a hub, its member sites as ready-to-paste handles. | | `resolve_source` | Turn a URL or site name into the canonical `source` handle and check it can be read. Changes nothing. | @@ -28,8 +28,17 @@ reads the site's already-published static JSON shards (`corpus.json` + to the exact second — an **archilyzer viewer** deep link (`…/?v=<slug>&t=<sec>`, opening the transcript modal at the moment) when the source has a public origin, otherwise the video's platform watch page with a per-platform time param -(YouTube `&t=<sec>s`, Odysee/Twitch equivalents). A `--local` source with no -origin falls back to platform links. +(YouTube `&t=<sec>s`, Odysee/Twitch equivalents). + +A `--local` source now **cites the site it was composed for**, not the platform: a +composed public dir is not an anonymous pile of JSON, it names its own deployed +origin in `corpus.json` (`site.url`), and that is read the first time the channel +list is loaded. Following a citation therefore lands in the archive — at the cited +second, with the transcript around it and the neighbouring videos one click away — +instead of on the platform page, where the archive's whole point (that a copy +still exists here) is invisible. Set `TRANSCRIPT_PLATFORM_LINKS=1` for the old +behaviour, which is the right choice when a local build's declared site URL is not +actually deployed. A dir with no `corpus.json` still falls back to platform links. **Compact base links (`link_style:"base"`).** Inline links are ~70–90 chars *per line* — bulk an agent pipeline shouldn't pay for. `search_transcripts` and @@ -85,7 +94,45 @@ Beyond `query`, `regex`, and `limit`: used verbatim (no expansion). - **`max_pages`** (default 400) — scan cap. If reached (or a very common term passes the 2000-video cap), the footer flags coverage as **PARTIAL** rather - than silently truncating. + than silently truncating — and **names which channels were fully scanned and + which were never reached**. The cap cuts in channel iteration order, never at + random, so "partial" without that list hides the shape of the bias: the sample + is simply whatever sorted first. With the list, it is actionable — re-run + scoped to the remainder. + +#### Filters (also on `enumerate_matches`) + +These are the share-link filters, reachable without a pasted link. Both search +tools take **all of them**, from one shared schema constant, because a filter +reachable from one and not the other would make the two disagree about coverage +— exactly the failure the stateless rebuild set out to make impossible. + +- **`states`** — keep only videos in these presence states: + `available`, `maybe_missing`, `deleted`, `private`, `members_only`, `unlisted`. + This is how you ask the corpus's signature question — *what did the videos that + have since been removed say about X* — which previously had no path at all. +- **`date_from`** / **`date_to`** — inclusive upload-date bounds, `YYYYMMDD`. A + differently-shaped date is **reported and ignored** rather than compared + lexicographically into a wrong answer. +- **`media_type`** — `video` or `livestream`. **`age`** — `all_ages` or + `restricted`. +- **`exclude`** — video-level NOT: drop any video that also contains one of these + terms. This is `"cup"` but not `"world cup"`. Compiled the same way the query is + (regex too) but **never** alias-expanded — an exclusion stays exactly as narrow + as you wrote it. Arbitrary boolean trees remain `open_link`'s job; that is what + a share link is for. +- **`scopes`** — which layers to match: `transcripts`, `chat`, `description`, + `tags`, `metadata` (title + channel), `posts`. Omitted means captions plus the + title, exactly as before. Naming scopes **replaces** that default, so + `scopes:["description"]` searches descriptions and *not* captions. Non-timed + layers emit `[description]`-tagged snippets with **no timestamp link**, since + citing a description line as `@ 0:00` would assert someone said it. + `chat` reads the separate ~1.2 GB live-chat shards lazily — scope it. +- **`collapse_duplicates`** (default **true**) — see *Counts are of recordings*. + +An unrecognised token in any of these is **named in the footer** rather than +dropped silently: a typo'd state would otherwise widen the search back to the +whole corpus and return a perfectly legitimate-looking answer. ### `get_transcripts` @@ -237,6 +284,138 @@ tools and the prompt — rather than duplicated into the markdown command files, so it cannot drift out of sync with the tools it names. A test asserts every tool the instructions mention actually exists. +## Counts are of recordings, not uploads + +Some videos exist in the archive more than once — the same recording mirrored to +another platform, or re-uploaded on another channel. The site already ships the +detector's output as `duplicates.json`; the MCP never opened it, so a sweep +counted the same recording twice and said "N videos". + +`search_transcripts` and `enumerate_matches` now **collapse cluster members to +one row by default**, and report it: *"12 mirror(s) collapsed across 9 +cluster(s) — the 43 above are distinct RECORDINGS, not uploads"*. Two rules keep +that safe: + +- **Nothing is hidden.** The collapsed copies are **named on the row they fold + into** (`⧉ same recording also archived as: …`). Pass + `collapse_duplicates:false` for one row per upload. +- **The surviving copy is never dropped.** The kept row is the cluster's + canonical member *only when that member is itself among the matches*; + otherwise it is simply the first match. A mirror is frequently the only + surviving copy of a deleted upload, and preferring an absent canonical would + delete precisely the evidence a `states:["deleted"]` question is asking for. + +**This changes reported totals**, so a report generated before and after will +disagree. That is a fix, not a regression: the earlier number was double-counting +mirrors. + +`get_video_metadata` lists a video's other copies, and states for each whether +the two were **measured as aligned**. If they were not — including when alignment +was simply never measured — it says so and tells you not to map a timestamp +across. A mirror with a different intro carries the same words at shifted times, +so a translated citation would look perfectly plausible and point at the wrong +moment of a different upload. Absent means *not measured*, and not measured means +*no*. + +## How it reads the corpus + +Three changes, in increasing order of how much they buy: + +**1. Nothing is read twice.** Channel manifests (~551 KB for the whole corpus) +are cached outright; shard pages go in a **byte-budgeted LRU**, default 48 MB of +raw page bytes, configurable with `TRANSCRIPT_MCP_PAGE_CACHE_MB`. The budget is +denominated in bytes rather than entries because page sizes differ by an order of +magnitude across corpora, and 48 MB is chosen from measurement: a parsed page +retains about **2.7× its file bytes**, so the ceiling is ~130 MB resident. Caches +hold **promises**, so concurrent callers for the same page coalesce onto one +read. `list_channels({refresh:true})` drops all of them — the explicit escape +hatch for a corpus rebuilt under a long-lived server, deliberately not a TTL. + +What this fixes: a 20-id `get_transcripts` batch used to re-read the same 7.4 MB +page once per id and, without a channel hint, do ~300 manifest reads. It is now +one read per distinct page. + +**2. Pages are read concurrently**, in windows clamped to the remaining +`max_pages` budget and folded back in page order — so hit ordering is unchanged +and `max_pages:1` still reads exactly one page. Local sources use a window of 4 +(reads overlap parses; parse dominates locally), remote and hub 8 (latency +dominates). + +**3. Filter-first scanning — the big one.** `summaries/` is a **global index of +every video** (slug, channel, title, upload date, livestream, age-restricted, +presence state) that parses in well under a second, against ~1.3 GB and tens of +seconds for the transcripts. The MCP already read it and threw away everything +except the availability state. It now keeps the whole record, and because each +channel manifest carries `slugToPage`, **a filtered query's exact page set is +computable before a single transcript byte is read**. + +The invariant that makes this safe: **the index only prunes pages; the record +predicate still decides every hit.** Both call the same `passesFilters`, so they +cannot drift, and a video the index has never heard of gets its page read +unconditionally. A stale, partial or missing summaries set therefore costs time, +never correctness. + +It pays in proportion to how selective the filter is, and it says which path ran +(*"filter-pruned: planned 8 of 170 page(s)"*), so a slow query is explicable. An +unfiltered query never reads the index at all — a filter that excludes nothing is +treated as no filter, so planning can never make a whole-corpus scan slower. + +### Benchmark + +`mcp/bench/` drives the **real server over stdio**, through the same +`pnpm --filter … exec tsx src/index.ts` command line the client is registered +with, and times a fixed query set: + +```bash +pnpm --filter yt-dlp-transcript-mcp bench +pnpm --filter yt-dlp-transcript-mcp bench -- --repeat 3 --json out.json +``` + +It reports two kinds of number and treats them differently. **Pages read and +bytes parsed are structural** — properties of the query plan, identical on an idle +box and a hammered one, so a before/after comparison of them is always valid. +**Wall time is contingent**: this box is shared, so the bench checks the load +average first and, above `--max-load` (default 0.7/core), **refuses to run** +unless given `--force`, in which case every wall figure is stamped `UNRELIABLE` +in both the table and the JSON. A number taken under load cannot later be quoted +as if it weren't. Every run prints a **corpus fingerprint** first, because the +composed dir gets rebuilt and a before/after that silently spans two corpora is +worse than no measurement at all. + +**The run this work was built against** — corpus `local:export/public`, +5 channels, 170 transcript pages, 1,288 MB, 3,330 videos in `summaries/`; +3 repetitions, median, taken at 0.64 load/core (i.e. trusted — no `UNRELIABLE` +stamp): + +| query | wall ms | page reads | MB parsed | +|---|---|---|---| +| `list_channels` (cold) | 1 | 0 | 0 | +| rare term, whole corpus | 16,967 | 170 | 1,261 | +| common term, whole corpus | 16,446 | 170 | 1,288 | +| common term, one channel | 42 | 4 | 27 | +| common term + one upload year | 2,647 | **32** | **202** | +| common term + missing states only | **823** | **8** | **63** | +| `enumerate_matches`, whole corpus | 23,422 | 170 | 1,288 | +| `get_transcripts` × 20 ids, one channel | 24 | 4 | 27 | + +Read it as three facts. **Filter-first is worth ~21× on a selective filter** (8 +pages against 170) and ~5× on a one-year date range — and the comparison to make +is within this same run: the removed-videos question costs 0.82 s where the only +previously-available way to ask it, an unfiltered scan, costs 16.4 s. **The +unfiltered rows are unchanged at 170 reads**, which is the point — planning must +never make a whole-corpus scan slower. **The 20-id batch is 4 reads**, cold; it +was up to 20 re-reads of the same page. + +*On the prebuilt-index question* (deferred until the speed win was a number): +these numbers say **not yet**. Scoped and filtered questions — the ones people +actually ask — now land between 0.04 s and 2.6 s. Only the unfiltered +whole-corpus scan is still slow, and that is a sweep's one-off first step before +it goes on to read hundreds of transcripts. If it ever does become the +bottleneck, the right artifact is a narrow token → `(channel, page)` postings +list emitted at export time: it would prune pages exactly the way the filter +planner already does, which makes that machinery its prerequisite rather than +its competitor. + ## Data source (pick one) Resolved from flags or env — precedence hub > remote > local: diff --git a/mcp/bench/bench.ts b/mcp/bench/bench.ts @@ -0,0 +1,556 @@ +// A read-only benchmark for the archilyzer MCP server. +// +// It drives the REAL server the way a client does — spawning `src/index.ts` +// over StdioClientTransport and calling advertised tools — rather than importing +// the internals. Anything it can measure is therefore something a real caller +// actually pays for. +// +// Deliberately outside `src/`, so `tsconfig`'s `include` ignores it and it never +// joins the test suite (`pnpm test` globs `src/*.test.ts`). +// +// pnpm --filter yt-dlp-transcript-mcp bench +// pnpm --filter yt-dlp-transcript-mcp bench -- --json out.json +// pnpm --filter yt-dlp-transcript-mcp bench -- --source local:/some/dir +// pnpm --filter yt-dlp-transcript-mcp bench -- --max-load 6 --repeat 3 +// +// ─── Why there are two kinds of number here ─── +// +// This box is shared: other agents' jobs run on it, and a `compose:site && +// next build` was running during the session that motivated this work (load +// average 27). Wall-clock timings taken under that are not measurements, they +// are noise with units. +// +// So the bench reports two families and treats them differently: +// +// pages / bytes STRUCTURAL. Properties of the query plan — how many shard +// pages the server had to open and how many bytes it parsed. +// Identical on an idle box and a hammered one, so a +// before/after comparison of these is always valid. +// wall ms CONTINGENT. Only meaningful below the load threshold, and +// the run REFUSES to print it as a headline above that +// (see the precondition gate below) — it is marked UNRELIABLE +// instead, so a number taken under load cannot later be +// quoted as if it weren't. +// +// Structural counters come from the server itself: MCP_IO_STATS=1 makes it emit +// one `[io] {...}` JSON line per tool call on stderr, which this reads. + +import { Client } from "@modelcontextprotocol/client"; +import { StdioClientTransport } from "@modelcontextprotocol/client/stdio"; +import { readFile, writeFile } from "node:fs/promises"; +import { cpus, loadavg } from "node:os"; +import path from "node:path"; +import { fileURLToPath } from "node:url"; + +const HERE = path.dirname(fileURLToPath(import.meta.url)); +const MCP_ROOT = path.resolve(HERE, ".."); +const REPO_ROOT = path.resolve(MCP_ROOT, ".."); + +// ─── args ─── + +type Args = { + source?: string; + local?: string; + json?: string; + maxLoad: number; + repeat: number; + force: boolean; + only?: string; + timeoutMs: number; +}; + +function parseArgs(argv: string[]): Args { + const flag = (name: string): string | undefined => { + const i = argv.indexOf(name); + return i >= 0 && i + 1 < argv.length ? argv[i + 1] : undefined; + }; + const num = (name: string, fallback: number): number => { + const raw = flag(name); + const n = raw === undefined ? NaN : Number(raw); + return Number.isFinite(n) ? n : fallback; + }; + return { + source: flag("--source"), + local: flag("--local"), + json: flag("--json"), + // Per-core 1-minute load. 1.0 means "as many runnable tasks as cores", + // which is already a contended box for a latency measurement. + maxLoad: num("--max-load", 0.7), + repeat: Math.max(1, Math.floor(num("--repeat", 3))), + force: argv.includes("--force"), + only: flag("--only"), + // A full-corpus scan of 1.3 GB exceeds the client's 60 s default, which is + // the thing being measured — so the bench must not time it out. + timeoutMs: num("--timeout", 900_000), + }; +} + +// ─── the precondition gate ─── +// +// This box runs other agents' jobs concurrently. A timing taken while a build +// or a GPU sweep is running is not a measurement of this server, and the way +// that bites is not at collection time — it is three weeks later, when the +// number is quoted from a changelog with no memory of what else was running. +// So the gate is not advisory: above the threshold, wall times are stamped +// UNRELIABLE in the output itself and in the JSON, permanently. + +type Preconditions = { + cores: number; + load1: number; + perCore: number; + ok: boolean; + note: string; +}; + +function checkPreconditions(maxLoad: number): Preconditions { + const cores = cpus().length || 1; + const load1 = loadavg()[0]; + const perCore = load1 / cores; + const ok = perCore <= maxLoad; + return { + cores, + load1, + perCore, + ok, + note: ok + ? `load ${load1.toFixed(2)} over ${cores} core(s) = ${perCore.toFixed(2)}/core — under the ${maxLoad}/core threshold` + : `load ${load1.toFixed(2)} over ${cores} core(s) = ${perCore.toFixed(2)}/core — ABOVE the ${maxLoad}/core threshold`, + }; +} + +// ─── the query suite ─── +// +// Fixed, and each case exists to pin one specific thing. Terms are chosen to be +// corpus-agnostic enough to run against any archilyzer site; the fingerprint +// printed above the table is what makes two runs comparable, not the terms. + +type Case = { + id: string; + what: string; + tool: string; + args: Record<string, unknown>; + // Filled at runtime from an earlier case (the id batch needs real ids). + needsIds?: "same-page" | "spread"; +}; + +const CASES: Case[] = [ + { + id: "cold-channels", + what: "list_channels (cold: corpus.json + groups)", + tool: "list_channels", + args: {}, + }, + { + id: "rare", + what: "rare term, whole corpus", + tool: "search_transcripts", + args: { query: "defenestration", limit: 5, content_types: ["video"] }, + }, + { + id: "common", + what: "common term, whole corpus", + tool: "search_transcripts", + args: { query: "lawsuit", limit: 5, content_types: ["video"] }, + }, + { + id: "channel-scoped", + what: "common term, one channel", + tool: "search_transcripts", + args: { query: "lawsuit", limit: 5, content_types: ["video"], channels: ["__FIRST_CHANNEL__"] }, + }, + { + id: "date-scoped", + what: "common term + upload-date range (filter-first)", + tool: "search_transcripts", + args: { + query: "lawsuit", + limit: 5, + content_types: ["video"], + date_from: "20240101", + date_to: "20241231", + }, + }, + { + id: "state-scoped", + what: "common term + missing states only (filter-first, the signature question)", + tool: "search_transcripts", + args: { + query: "lawsuit", + limit: 5, + content_types: ["video"], + states: ["deleted", "private", "members_only", "unlisted", "maybe_missing"], + }, + }, + { + id: "enumerate", + what: "enumerate_matches, whole corpus", + tool: "enumerate_matches", + args: { query: "lawsuit", content_types: ["video"] }, + }, + { + id: "batch-20", + what: "get_transcripts × 20 ids from one channel", + tool: "get_transcripts", + args: { query: "lawsuit", before: 15, after: 15 }, + needsIds: "same-page", + }, +]; + +// ─── stderr [io] line collection ─── + +type IoLine = { tool: string; ms: number; reads: number; bytes: number }; + +class IoCollector { + private lines: IoLine[] = []; + private buffered = ""; + + feed(chunk: string): void { + this.buffered += chunk; + const parts = this.buffered.split("\n"); + this.buffered = parts.pop() ?? ""; + for (const line of parts) { + const at = line.indexOf("[io] "); + if (at === -1) continue; + try { + this.lines.push(JSON.parse(line.slice(at + 5)) as IoLine); + } catch { + // a partial or malformed line is simply not a measurement + } + } + } + + // Everything since the mark, summed. Tool calls are serial here, so this is + // exactly the calls made by the case being timed. + drain(): { reads: number; bytes: number; serverMs: number } { + const taken = this.lines; + this.lines = []; + return { + reads: taken.reduce((a, l) => a + l.reads, 0), + bytes: taken.reduce((a, l) => a + l.bytes, 0), + serverMs: taken.reduce((a, l) => a + l.ms, 0), + }; + } +} + +// ─── corpus fingerprint ─── +// +// Printed above every table. Two runs are only comparable if these match — the +// composed dir gets rebuilt (one rebuild landed mid-session and swapped the +// 30,923-video composition for a 3,330-video one), and a before/after that +// silently spans two corpora is worse than no measurement at all. + +type Fingerprint = { + handle: string; + channels: number; + transcriptPages: number; + transcriptBytes: number; + summariesVideos: number | null; + hasStats: boolean; + hasDuplicates: boolean; + hasDigests: boolean; +}; + +async function fingerprintLocal(dir: string): Promise<Partial<Fingerprint>> { + const readJson = async <T,>(p: string): Promise<T | null> => { + try { + return JSON.parse(await readFile(path.join(dir, p), "utf8")) as T; + } catch { + return null; + } + }; + const { statSync, readdirSync } = await import("node:fs"); + const corpus = await readJson<{ channels?: { slug: string }[] }>("corpus.json"); + const summaries = await readJson<{ totalCount?: number }>("summaries/manifest.json"); + let pages = 0; + let bytes = 0; + for (const c of corpus?.channels ?? []) { + const chDir = path.join(dir, "transcripts", c.slug); + try { + for (const f of readdirSync(chDir)) { + if (!f.startsWith("page-")) continue; + pages++; + bytes += statSync(path.join(chDir, f)).size; + } + } catch { + // channel not present in this composition + } + } + const exists = (p: string): boolean => { + try { + statSync(path.join(dir, p)); + return true; + } catch { + return false; + } + }; + return { + channels: corpus?.channels?.length ?? 0, + transcriptPages: pages, + transcriptBytes: bytes, + summariesVideos: summaries?.totalCount ?? null, + hasStats: exists("stats/manifest.json"), + hasDuplicates: exists("duplicates.json"), + hasDigests: exists("digests/manifest.json"), + }; +} + +// ─── running ─── + +type Row = { + id: string; + what: string; + wallMs: number[]; + medianMs: number; + reads: number; + bytes: number; + note: string; +}; + +function median(xs: number[]): number { + const s = [...xs].sort((a, b) => a - b); + const mid = Math.floor(s.length / 2); + return s.length % 2 === 1 ? s[mid] : Math.round((s[mid - 1] + s[mid]) / 2); +} + +function mb(bytes: number): string { + return bytes === 0 ? "0" : (bytes / 1024 / 1024).toFixed(1); +} + +function firstText(result: unknown): string { + const content = (result as { content?: { type: string; text?: string }[] }) + .content; + return (content ?? []) + .filter((c) => c.type === "text") + .map((c) => c.text ?? "") + .join("\n"); +} + +async function main(): Promise<void> { + const args = parseArgs(process.argv.slice(2)); + const pre = checkPreconditions(args.maxLoad); + + console.log("archilyzer MCP benchmark"); + console.log(""); + console.log(` preconditions: ${pre.note}`); + if (!pre.ok && !args.force) { + console.log(""); + console.log( + " REFUSING to report wall-clock timings: this box is shared, and a\n" + + " timing taken under this load measures the other jobs, not the server.\n" + + " Structural counters (pages read, bytes parsed) are load-independent —\n" + + " re-run with --force to collect those anyway, and the wall column will\n" + + " be stamped UNRELIABLE.", + ); + process.exitCode = 2; + return; + } + const wallTrusted = pre.ok; + + const io = new IoCollector(); + // Spawn the server through the SAME command line the MCP client is + // registered with (`pnpm --filter … exec tsx src/index.ts --local …`). Not a + // detail: `tsx src/index.ts` run directly cannot resolve + // `yt-dlp-transcript-common` — the workspace link comes from pnpm — so a + // bench that invented its own invocation would be measuring a process the + // real client never starts, if it started at all. + const corpusDir = args.local ?? path.join(REPO_ROOT, "export", "public"); + const transport = new StdioClientTransport({ + command: "pnpm", + args: [ + "-C", + REPO_ROOT, + "--filter", + "yt-dlp-transcript-mcp", + "exec", + "tsx", + "src/index.ts", + "--local", + corpusDir, + ], + cwd: REPO_ROOT, + stderr: "pipe", + env: { ...process.env, MCP_IO_STATS: "1" } as Record<string, string>, + }); + transport.stderr?.on("data", (c: Buffer) => io.feed(c.toString("utf8"))); + + const client = new Client({ name: "mcp-bench", version: "1.0.0" }); + await client.connect(transport); + + // Resolve the corpus + fingerprint it. + const sourceArg = args.source ? { source: args.source } : {}; + const resolved = firstText( + await client.callTool( + { name: "resolve_source", arguments: { source: args.source ?? "default" } }, + { timeout: args.timeoutMs }, + ), + ); + const handle = /Handle:\s*(\S+)/.exec(resolved)?.[1] ?? "(unknown)"; + const fp: Fingerprint = { + handle, + channels: 0, + transcriptPages: 0, + transcriptBytes: 0, + summariesVideos: null, + hasStats: false, + hasDuplicates: false, + hasDigests: false, + ...(handle.startsWith("local:") + ? await fingerprintLocal(path.resolve(REPO_ROOT, handle.slice("local:".length))) + : {}), + }; + io.drain(); + + console.log(""); + console.log(" corpus fingerprint (two runs are comparable only if these match):"); + console.log(` handle: ${fp.handle}`); + console.log(` channels: ${fp.channels}`); + console.log( + ` transcripts: ${fp.transcriptPages} page(s), ${mb(fp.transcriptBytes)} MB`, + ); + console.log(` summaries: ${fp.summariesVideos ?? "?"} video(s)`); + console.log( + ` layers: stats=${fp.hasStats} duplicates=${fp.hasDuplicates} digests=${fp.hasDigests}`, + ); + + // The channel-scoped case needs a real channel name. + const channelsText = firstText( + await client.callTool( + { name: "list_channels", arguments: sourceArg }, + { timeout: args.timeoutMs }, + ), + ); + const firstChannel = /slug: ([^,)]+)/.exec(channelsText)?.[1]?.trim(); + io.drain(); + + // The id batch needs real ids: take them from one channel's worklist so they + // cluster onto as few shard pages as possible — that IS the case under test. + const worklist = firstText( + await client.callTool( + { + name: "enumerate_matches", + arguments: { + ...sourceArg, + query: "lawsuit", + content_types: ["video"], + ...(firstChannel ? { channels: [firstChannel] } : {}), + }, + }, + { timeout: args.timeoutMs }, + ), + ); + const ids = [...worklist.matchAll(/^- (\S+) \| video \|/gm)] + .map((m) => m[1]) + .slice(0, 20); + io.drain(); + + const rows: Row[] = []; + for (const c of CASES) { + if (args.only && !c.id.includes(args.only)) continue; + + const callArgs: Record<string, unknown> = { ...sourceArg, ...c.args }; + if (Array.isArray(callArgs.channels)) { + callArgs.channels = (callArgs.channels as string[]).map((x) => + x === "__FIRST_CHANNEL__" ? (firstChannel ?? "") : x, + ); + if ((callArgs.channels as string[]).some((x) => x === "")) continue; + } + if (c.needsIds) { + if (ids.length === 0) continue; + callArgs.video_ids = ids; + if (firstChannel) callArgs.channels = [firstChannel]; + } + + const wall: number[] = []; + let reads = 0; + let bytes = 0; + let note = ""; + for (let r = 0; r < args.repeat; r++) { + const t0 = performance.now(); + const out = await client.callTool( + { name: c.tool, arguments: callArgs }, + { timeout: args.timeoutMs }, + ); + wall.push(Math.round(performance.now() - t0)); + const stats = io.drain(); + // Report the FIRST (cold) repetition's structural cost. Later runs hit + // the process-lifetime caches, which is a different question — one the + // 'warm' note answers rather than hides. + if (r === 0) { + reads = stats.reads; + bytes = stats.bytes; + const text = firstText(out); + const m = /scanned (\d+) page\(s\)/.exec(text); + if (m) note = `${m[1]} page(s) scanned`; + if (/coverage PARTIAL|COVERAGE PARTIAL/.test(text)) note += " · CAPPED"; + if (/filter-pruned/.test(text)) note += " · pruned"; + } else if (r === 1) { + note += note ? `; warm ${stats.reads} read(s)` : `warm ${stats.reads} read(s)`; + } + } + rows.push({ + id: c.id, + what: c.what, + wallMs: wall, + medianMs: median(wall), + reads, + bytes, + note, + }); + console.log( + ` ran ${c.id} (${wall.map((w) => `${w}ms`).join(", ")})`, + ); + } + + await client.close(); + + // ─── the table ─── + const wallHeader = wallTrusted ? "wall ms" : "wall ms (UNRELIABLE)"; + const w = [ + Math.max(28, ...rows.map((r) => r.what.length)), + Math.max(wallHeader.length, 12), + 12, + 12, + ]; + console.log(""); + console.log( + `| ${"query".padEnd(w[0])} | ${wallHeader.padEnd(w[1])} | ${"reads".padEnd(w[2])} | ${"MB parsed".padEnd(w[3])} |`, + ); + console.log( + `| ${"-".repeat(w[0])} | ${"-".repeat(w[1])} | ${"-".repeat(w[2])} | ${"-".repeat(w[3])} |`, + ); + for (const r of rows) { + console.log( + `| ${r.what.padEnd(w[0])} | ${String(r.medianMs).padEnd(w[1])} | ${String(r.reads).padEnd(w[2])} | ${mb(r.bytes).padEnd(w[3])} |`, + ); + } + console.log(""); + for (const r of rows) { + if (r.note) console.log(` ${r.id}: ${r.note}`); + } + if (!wallTrusted) { + console.log(""); + console.log( + ` ⚠ wall times above were taken at ${pre.perCore.toFixed(2)} load/core and are NOT\n` + + ` a measurement of this server. The reads/MB columns are structural and\n` + + ` remain valid. Re-run under ${args.maxLoad}/core for usable timings.`, + ); + } + + if (args.json) { + await writeFile( + args.json, + JSON.stringify( + { preconditions: pre, wallTrusted, fingerprint: fp, rows }, + null, + 2, + ), + "utf8", + ); + console.log(`\n wrote ${args.json}`); + } +} + +main().catch((e: unknown) => { + console.error(e); + process.exit(1); +}); diff --git a/mcp/bench/smoke.ts b/mcp/bench/smoke.ts @@ -0,0 +1,196 @@ +// A read-only smoke test against a REAL corpus, driving the server the way the +// registered MCP client does (`pnpm --filter … exec tsx src/index.ts --local …`) +// rather than importing anything. +// +// The unit tests prove the logic; this proves the shipped thing works on 1.3 GB +// of real shards — where the failures are the ones a stub cannot have: a layer +// that isn't where the code thinks it is, a manifest shape that differs from the +// fixture, a filter that silently matches nothing. +// +// pnpm --filter yt-dlp-transcript-mcp smoke +// pnpm --filter yt-dlp-transcript-mcp smoke -- --local /path/to/public +// +// Exits non-zero if any check fails. + +import { Client } from "@modelcontextprotocol/client"; +import { StdioClientTransport } from "@modelcontextprotocol/client/stdio"; +import path from "node:path"; +import { fileURLToPath } from "node:url"; + +const HERE = path.dirname(fileURLToPath(import.meta.url)); +const REPO_ROOT = path.resolve(HERE, "..", ".."); +const TIMEOUT = 900_000; + +const argv = process.argv.slice(2); +const flag = (name: string): string | undefined => { + const i = argv.indexOf(name); + return i >= 0 && i + 1 < argv.length ? argv[i + 1] : undefined; +}; +const corpusDir = flag("--local") ?? path.join(REPO_ROOT, "export", "public"); +const QUERY = flag("--query") ?? "lawsuit"; + +let failures = 0; +function check(ok: boolean, label: string, detail = ""): void { + if (ok) { + console.log(` ✓ ${label}`); + } else { + failures++; + console.log(` ✗ ${label}${detail ? `\n ${detail}` : ""}`); + } +} + +function textOf(result: unknown): string { + const content = (result as { content?: { type: string; text?: string }[] }) + .content; + return (content ?? []) + .filter((c) => c.type === "text") + .map((c) => c.text ?? "") + .join("\n"); +} + +async function main(): Promise<void> { + const transport = new StdioClientTransport({ + command: "pnpm", + args: [ + "-C", + REPO_ROOT, + "--filter", + "yt-dlp-transcript-mcp", + "exec", + "tsx", + "src/index.ts", + "--local", + corpusDir, + ], + cwd: REPO_ROOT, + stderr: "pipe", + }); + const client = new Client({ name: "mcp-smoke", version: "1.0.0" }); + await client.connect(transport); + const call = (name: string, args: Record<string, unknown>): Promise<unknown> => + client.callTool({ name, arguments: args }, { timeout: TIMEOUT }); + + console.log(`smoke: ${corpusDir}, query "${QUERY}"\n`); + + // ── every result names the corpus it read ── + const channels = textOf(await call("list_channels", {})); + check(/\(corpus: local:/.test(channels), "list_channels echoes (corpus: …)"); + const firstChannel = /slug: ([^,)]+)/.exec(channels)?.[1]?.trim(); + check(Boolean(firstChannel), "a channel slug is discoverable", channels.slice(0, 200)); + + const bogus = textOf(await call("get_transcript", { video_id: "___nope___" })); + check(/\(corpus: local:/.test(bogus), "an ERROR result also echoes (corpus: …)"); + + // ── the invariant: enumerate and search agree on the deduped total ── + const filters = { + query: QUERY, + content_types: ["video"], + states: ["deleted", "private", "members_only", "unlisted", "maybe_missing"], + }; + const enumerated = textOf(await call("enumerate_matches", filters)); + const searched = textOf(await call("search_transcripts", { ...filters, limit: 1 })); + const enumTotal = /(\d+) match\(es\); \d+ batch/.exec(enumerated)?.[1]; + const searchTotal = /\(total (\d+) match\(es\)/.exec(searched)?.[1]; + check( + enumTotal !== undefined && enumTotal === searchTotal, + `enumerate_matches and search_transcripts agree on the total (${enumTotal} vs ${searchTotal})`, + enumerated.slice(0, 400), + ); + + // ── filter-first actually pruned ── + check( + /filter-pruned: planned \d+ of \d+ page\(s\)/.test(searched), + "a states-filtered search reports filter-pruned page planning", + /scanned [^;]+/.exec(searched)?.[0] ?? "", + ); + const planned = /filter-pruned: planned (\d+) of (\d+) page/.exec(searched); + if (planned) { + check( + Number(planned[1]) < Number(planned[2]), + `it planned fewer pages than the corpus has (${planned[1]} of ${planned[2]})`, + ); + } + + // ── the unfiltered path is untouched ── + const plain = textOf( + await call("search_transcripts", { + query: QUERY, + content_types: ["video"], + limit: 2, + }), + ); + check( + !/filter-pruned/.test(plain), + "an unfiltered search does NOT plan (no index read, no pruning)", + ); + + // ── citations point at the archive, not the platform ── + const withSnippets = textOf( + await call("search_transcripts", { + query: QUERY, + content_types: ["video"], + limit: 1, + }), + ); + const link = /\]\((https?:\/\/[^)]+)\)/.exec(withSnippets)?.[1]; + check( + link !== undefined && /[?&]v=/.test(link) && /[?&]t=/.test(link), + `a cited moment links to the archive viewer, not the platform (${link ?? "no link found"})`, + ); + + // ── exclude ── + const excluded = textOf( + await call("enumerate_matches", { + query: QUERY, + content_types: ["video"], + states: ["deleted", "private", "members_only", "unlisted", "maybe_missing"], + exclude: [QUERY], + }), + ); + check( + /No matches for/.test(excluded) || /^0 match/.test(excluded), + "excluding the query term itself yields nothing (the NOT is applied)", + excluded.slice(0, 200), + ); + + // ── an unknown filter token is reported, not swallowed ── + const typo = textOf( + await call("search_transcripts", { + query: QUERY, + content_types: ["video"], + limit: 1, + states: ["removed"], + }), + ); + check( + /unknown state\(s\) ignored: removed/.test(typo), + "a typo'd state is named in the footer rather than silently widening the scan", + ); + + // ── the joined layers on get_video_metadata ── + const worklist = textOf( + await call("enumerate_matches", { + query: QUERY, + content_types: ["video"], + ...(firstChannel ? { channels: [firstChannel] } : {}), + }), + ); + const someId = /^- (\S+) \| video \|/m.exec(worklist)?.[1]; + if (someId) { + const meta = textOf(await call("get_video_metadata", { video_id: someId })); + check(/"title"/.test(meta), `get_video_metadata returns the base record (${someId})`); + check(/## Stats/.test(meta), "…joined with stats/ (engagement, cue count, coverage)"); + } else { + check(false, "found an id to inspect with get_video_metadata"); + } + + await client.close(); + + console.log(`\n${failures === 0 ? "ALL CHECKS PASSED" : `${failures} CHECK(S) FAILED`}`); + if (failures > 0) process.exitCode = 1; +} + +main().catch((e: unknown) => { + console.error(e); + process.exit(1); +}); diff --git a/mcp/package.json b/mcp/package.json @@ -10,7 +10,9 @@ "scripts": { "start": "tsx src/index.ts", "test": "tsx --test src/*.test.ts", - "typecheck": "tsc --noEmit -p tsconfig.json" + "typecheck": "tsc --noEmit -p tsconfig.json", + "bench": "tsx bench/bench.ts", + "smoke": "tsx bench/smoke.ts" }, "dependencies": { "@modelcontextprotocol/server": "^2.0.0" diff --git a/mcp/src/instructions.test.ts b/mcp/src/instructions.test.ts @@ -53,6 +53,9 @@ test("every tool the instructions name actually exists", () => { "dry_run", "parse_model", "report_path", + "date_from", + "date_to", + "media_type", ]); const named = [...suspects].filter((s) => !NOT_TOOLS.has(s)); assert.ok(named.length > 0, "the instructions should name some tools"); diff --git a/mcp/src/instructions.ts b/mcp/src/instructions.ts @@ -182,7 +182,31 @@ export function buildSweepInstructions( `319 videos after seeing 200. If the first line says **COVERAGE ` + `PARTIAL**, the list is a sample — narrow the scope or raise ` + `\`max_pages\`, and if you proceed anyway, say so prominently in the ` + - `report.`, + `report — it now also NAMES the channels it never reached, so quote ` + + `those rather than just saying "partial".`, + ); + + steps.push( + `**Narrow it if the question is narrow.** \`enumerate_matches\` and ` + + `\`search_transcripts\` take the same filters: \`states\` (e.g. ` + + `\`["deleted","private","members_only","unlisted","maybe_missing"]\` for ` + + `"what did the videos that are now GONE say"), \`date_from\`/` + + `\`date_to\`, \`media_type\`, \`age\`, \`exclude\` (video-level NOT, ` + + `for "cup" but not "world cup"), and \`scopes\` to search descriptions, ` + + `tags or live chat instead of captions. Use them when the question ` + + `implies them: a filtered scan reads only the shard pages that can hold ` + + `a match, which is the difference between seconds and a minute per ` + + `query — and it makes the answer narrower and more honest at the same ` + + `time.`, + ); + + steps.push( + `**Counts are of RECORDINGS, not uploads.** Some videos are mirrored ` + + `across platforms/channels; the tools collapse those to one row by ` + + `default and name the collapsed copies inline. Quote the number the ` + + `footer gives you. If a total here disagrees with an older report, the ` + + `old one was double-counting mirrors — say that rather than splitting ` + + `the difference.`, ); steps.push( @@ -263,7 +287,12 @@ export function buildAskInstructions( `corpus is ASR text and the curated aliases only cover known ` + `mis-transcriptions. If you need to know HOW MANY, or to cover ` + `everything, use \`enumerate_matches\` instead: a search page is a ` + - `slice, and its \`⚠ INCOMPLETE PAGE\` banner means exactly that.`, + `slice, and its \`⚠ INCOMPLETE PAGE\` banner means exactly that. Both ` + + `take \`states\` / \`date_from\` / \`date_to\` / \`media_type\` / ` + + `\`age\` / \`exclude\` / \`scopes\` — reach for them when the question ` + + `is about a period, or about videos that have since been REMOVED ` + + `(\`states\`), which is otherwise unaskable. Reported totals count a ` + + `recording mirrored across platforms ONCE.`, ); steps.push( diff --git a/mcp/src/scanPlan.test.ts b/mcp/src/scanPlan.test.ts @@ -0,0 +1,528 @@ +// Tests for the four things the speed/honesty work added, each aimed at the +// specific failure it was built to prevent: +// +// 1. the page cache — a batch of ids on one shard page must parse it ONCE +// 2. filter-first planning — a filtered query must read only the pages that +// can hold a match, and must never skip a page it +// is not certain about +// 3. duplicate collapsing — a recording mirrored across platforms counts once +// 4. cap honesty — a truncated scan must NAME the channels it never +// reached, because it truncates in channel order +// +// The page-cache test drives a real LocalSource over a temp dir, because the +// cache lives in the source implementations — a stub would prove nothing. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { mkdtemp, mkdir, writeFile, rm } from "node:fs/promises"; +import { tmpdir } from "node:os"; +import path from "node:path"; +import type { ChannelTranscriptsManifest } from "yt-dlp-transcript-common/lib/manifest"; +import type { ChannelSubsManifest } from "yt-dlp-transcript-common/lib/manifest"; +import type { TranscriptDetail } from "yt-dlp-transcript-common/lib/transcripts"; +import type { SubsDetail } from "yt-dlp-transcript-common/lib/subs"; +import type { ChannelPostsManifest, Post } from "yt-dlp-transcript-common/lib/posts"; +import type { SearchAlias } from "yt-dlp-transcript-common/lib/searchAliases"; +import type { Cue } from "yt-dlp-transcript-common/lib/vtt"; +import { VIDEO_STATES } from "yt-dlp-transcript-common/lib/availability"; +import { + LocalSource, + type ChannelGroups, + type ChannelRef, + type DuplicateIndex, + type ShardSource, + type VideoAvailability, + type VideoIndex, +} from "./source"; +import { buildScanPlan, searchTranscripts, type SearchFilters } from "./search"; + +// ─── fixtures ─── + +function cues(...pairs: [number, string][]): Cue[] { + return pairs.map(([start, text]) => ({ start, end: start + 3, text })); +} + +function video( + channelSlug: string, + id: string, + opts: { + uploadDate?: string; + isLivestream?: boolean; + deleted?: boolean; + text?: string; + } = {}, +): TranscriptDetail { + return { + id, + slug: `${channelSlug}/${id}`, + channelSlug, + channel: channelSlug, + title: `video ${id}`, + uploadDate: opts.uploadDate ?? "20240601", + duration: 600, + isLivestream: opts.isLivestream ?? false, + ageRestricted: false, + platform: "youtube", + webpageUrl: `https://www.youtube.com/watch?v=${id}`, + description: "", + tags: [], + cues: cues([10, opts.text ?? "they filed a lawsuit today"]), + }; +} + +// A source whose pages are declared per channel, counting every page read so a +// test can assert exactly which shard pages a query opened. +class CountingSource implements ShardSource { + readonly label = "counting"; + readonly reads: string[] = []; + + constructor( + private readonly channelPages: Record<string, TranscriptDetail[][]>, + private readonly index?: Map<string, VideoIndex extends ReadonlyMap<string, infer V> ? V : never>, + private readonly duplicates?: DuplicateIndex, + ) {} + + async listChannels(): Promise<ChannelRef[]> { + return Object.keys(this.channelPages).map((slug) => ({ + key: slug, + slug, + name: slug, + })); + } + + async transcriptsManifest(ch: ChannelRef): Promise<ChannelTranscriptsManifest> { + const pages = this.channelPages[ch.slug] ?? []; + const slugToPage: Record<string, number> = {}; + pages.forEach((page, i) => { + for (const rec of page) slugToPage[rec.id] = i; + }); + return { + version: 1, + channelSlug: ch.slug, + pageCount: pages.length, + maxPageBytes: 0, + generatedAt: "", + slugToPage, + }; + } + + async transcriptPage(ch: ChannelRef, page: number): Promise<TranscriptDetail[]> { + this.reads.push(`${ch.slug}:${page}`); + return this.channelPages[ch.slug]?.[page] ?? []; + } + + async loadAliases(): Promise<SearchAlias[]> { + return []; + } + async loadGroups(): Promise<ChannelGroups> { + return { groups: [], defaultGroupId: "default" }; + } + publicOrigin(): string | null { + return null; + } + async subsManifest(): Promise<ChannelSubsManifest | null> { + return null; + } + async subsPage(): Promise<SubsDetail[]> { + return []; + } + async postsManifest(): Promise<ChannelPostsManifest | null> { + return null; + } + async postsPage(): Promise<Post[]> { + return []; + } + async availabilityMap(): Promise<ReadonlyMap<string, VideoAvailability>> { + return this.index ?? new Map(); + } + videoIndex(): Promise<VideoIndex> { + // Deliberately still a method when the map is absent — an EMPTY index is a + // different statement from no index at all, and the planner must treat the + // empty case as "know nothing", not "nothing matches". + return Promise.resolve((this.index ?? new Map()) as VideoIndex); + } + duplicateIndex(): Promise<DuplicateIndex> { + return Promise.resolve(this.duplicates ?? new Map()); + } +} + +function indexed( + rec: TranscriptDetail, + state: "available" | "deleted" = "available", +): [string, NonNullable<ReturnType<VideoIndex["get"]>>] { + return [ + rec.slug, + { + state, + id: rec.id, + channelSlug: rec.channelSlug, + title: rec.title, + uploadDate: rec.uploadDate, + isLivestream: rec.isLivestream === true, + ageRestricted: rec.ageRestricted === true, + }, + ]; +} + +const ALL_STATES: SearchFilters = { + videos: true, + livestreams: true, + allAges: true, + restricted: true, + states: new Set(VIDEO_STATES), +}; + +// ─── 1. the page cache ─── + +test("a batch of ids sharing one shard page parses that page exactly once", async () => { + const dir = await mkdtemp(path.join(tmpdir(), "mcp-cache-")); + try { + const chDir = path.join(dir, "transcripts", "chan"); + await mkdir(chDir, { recursive: true }); + const page = [video("chan", "a1"), video("chan", "a2"), video("chan", "a3")]; + const slugToPage: Record<string, number> = { a1: 0, a2: 0, a3: 0 }; + await writeFile( + path.join(dir, "corpus.json"), + JSON.stringify({ + site: { id: "s", title: "S", url: "https://example.test/" }, + channels: [{ slug: "chan", name: "Chan", videoCount: 3 }], + }), + ); + await writeFile( + path.join(chDir, "manifest.json"), + JSON.stringify({ + version: 1, + channelSlug: "chan", + pageCount: 1, + maxPageBytes: 0, + generatedAt: "", + slugToPage, + }), + ); + await writeFile(path.join(chDir, "page-0000.json"), JSON.stringify(page)); + + const source = new LocalSource(dir); + const ch = (await source.listChannels())[0]; + + // Identity is the assertion: a second parse would produce a different + // array. Same reference ⇒ the bytes were read and parsed once. + const first = await source.transcriptPage(ch, 0); + const second = await source.transcriptPage(ch, 0); + assert.equal(first, second, "second read of the same page must be the cached one"); + + // Concurrent callers coalesce onto ONE in-flight read rather than racing. + const [a, b, c] = await Promise.all([ + source.transcriptPage(ch, 0), + source.transcriptPage(ch, 0), + source.transcriptPage(ch, 0), + ]); + assert.equal(a, b); + assert.equal(b, c); + assert.equal(a, first); + + // Manifests are cached outright — this is what kills the ~300 manifest + // reads an un-hinted 20-id batch used to cost. + assert.equal( + await source.transcriptsManifest(ch), + await source.transcriptsManifest(ch), + ); + + // …and refresh is the escape hatch for a corpus rebuilt under a running + // server: it must drop the page cache too, not just the channel list. + await source.listChannels({ refresh: true }); + assert.notEqual( + await source.transcriptPage(ch, 0), + first, + "refresh must drop cached pages", + ); + } finally { + await rm(dir, { recursive: true, force: true }); + } +}); + +test("a local corpus cites its own composed site, not the platform", async () => { + const dir = await mkdtemp(path.join(tmpdir(), "mcp-origin-")); + try { + await writeFile( + path.join(dir, "corpus.json"), + JSON.stringify({ + site: { id: "s", title: "S", url: "https://hasanalyzer.pages.dev/" }, + channels: [{ slug: "chan", name: "Chan" }], + }), + ); + const source = new LocalSource(dir); + // Unknown until the channel list is read — every render path does that + // first, and guessing an origin before reading corpus.json would be + // inventing one. + assert.equal(source.publicOrigin(), null); + await source.listChannels(); + assert.equal(source.publicOrigin(), "https://hasanalyzer.pages.dev"); + } finally { + await rm(dir, { recursive: true, force: true }); + } +}); + +test("a local corpus with no corpus.json keeps falling back to platform links", async () => { + const dir = await mkdtemp(path.join(tmpdir(), "mcp-origin-none-")); + try { + await mkdir(path.join(dir, "transcripts"), { recursive: true }); + const source = new LocalSource(dir); + await source.listChannels(); + assert.equal(source.publicOrigin(), null); + } finally { + await rm(dir, { recursive: true, force: true }); + } +}); + +// ─── 2. filter-first planning ─── + +test("a date-scoped query plans only the pages that can hold a match", async () => { + const p0 = [video("chan", "old1", { uploadDate: "20220101" })]; + const p1 = [video("chan", "new1", { uploadDate: "20240301" })]; + const p2 = [video("chan", "old2", { uploadDate: "20210101" })]; + const source = new CountingSource( + { chan: [p0, p1, p2] }, + new Map([indexed(p0[0]), indexed(p1[0]), indexed(p2[0])]), + ); + const channels = await source.listChannels(); + + const plan = await buildScanPlan(source, channels, { + ...ALL_STATES, + dateFrom: "20240101", + dateTo: "20241231", + }); + assert.equal(plan.pruned, true); + assert.deepEqual(plan.perChannel.get("chan")?.pages, [1]); + assert.equal(plan.pagesPlanned, 1); + assert.equal(plan.pagesTotal, 3); + + // …and the scan really only opens that page. + const result = await searchTranscripts(source, { + query: "lawsuit", + contentTypes: ["video"], + filters: { ...ALL_STATES, dateFrom: "20240101", dateTo: "20241231" }, + }); + assert.deepEqual(source.reads, ["chan:1"]); + assert.equal(result.total, 1); + assert.equal(result.hits[0].videoId, "new1"); + assert.equal(result.coverage.pruned, true); + assert.equal(result.coverage.pagesPlanned, 1); +}); + +test("a video the index has never heard of still gets its page scanned", async () => { + const p0 = [video("chan", "known", { uploadDate: "20200101" })]; + const p1 = [video("chan", "ghost", { uploadDate: "20200101" })]; + // The index knows only `known`, and would exclude it by date. `ghost` is + // absent — which must mean "read the page and let the record decide", not + // "skip it". This is the rule that makes a stale or partial summaries set + // cost time instead of correctness. + const source = new CountingSource( + { chan: [p0, p1] }, + new Map([indexed(p0[0])]), + ); + const channels = await source.listChannels(); + const filters: SearchFilters = { ...ALL_STATES, dateFrom: "20240101" }; + + const plan = await buildScanPlan(source, channels, filters); + assert.deepEqual(plan.perChannel.get("chan")?.pages, [1]); + assert.equal(plan.unknownVideos, 1); + + const result = await searchTranscripts(source, { + query: "lawsuit", + contentTypes: ["video"], + filters, + }); + // The page was read, and the record predicate — not the index — dropped it. + assert.deepEqual(source.reads, ["chan:1"]); + assert.equal(result.total, 0); +}); + +test("an unfiltered query never reads the index and never prunes", async () => { + const p0 = [video("chan", "a1")]; + const p1 = [video("chan", "a2")]; + const source = new CountingSource({ chan: [p0, p1] }, new Map()); + const result = await searchTranscripts(source, { + query: "lawsuit", + contentTypes: ["video"], + }); + assert.deepEqual(source.reads, ["chan:0", "chan:1"]); + assert.equal(result.coverage.pruned, false); + assert.equal(result.total, 2); +}); + +test("an all-permissive filter is treated as no filter, not as a reason to plan", async () => { + const p0 = [video("chan", "a1")]; + const source = new CountingSource({ chan: [p0] }, new Map()); + const plan = await buildScanPlan(source, await source.listChannels(), ALL_STATES); + assert.equal(plan.pruned, false, "a filter that excludes nothing must not trigger an index read"); + assert.equal(plan.pagesPlanned, 1); +}); + +test("a states filter reaches only the pages holding videos in those states", async () => { + const p0 = [video("chan", "live1")]; + const p1 = [video("chan", "gone1", { deleted: true })]; + const p2 = [video("chan", "live2")]; + const source = new CountingSource( + { chan: [p0, p1, p2] }, + new Map([ + indexed(p0[0], "available"), + indexed(p1[0], "deleted"), + indexed(p2[0], "available"), + ]), + ); + const result = await searchTranscripts(source, { + query: "lawsuit", + contentTypes: ["video"], + filters: { ...ALL_STATES, states: new Set(["deleted"]) }, + }); + assert.deepEqual(source.reads, ["chan:1"]); + assert.equal(result.total, 1); + assert.equal(result.hits[0].videoId, "gone1"); +}); + +// ─── 3. duplicate collapsing ─── + +test("a recording mirrored across channels is counted once and its mirror named", async () => { + const original = video("chanA", "orig"); + const mirror = video("chanB", "copy"); + const membership = { + clusterId: "c1", + canonicalSlug: "chanA/orig", + contained: false, + needsReview: false, + }; + const duplicates: DuplicateIndex = new Map([ + [ + "chanA/orig", + { ...membership, isCanonical: true, siblings: [] }, + ], + [ + "chanB/copy", + { ...membership, isCanonical: false, siblings: [] }, + ], + ]); + const source = new CountingSource( + { chanA: [[original]], chanB: [[mirror]] }, + new Map(), + duplicates, + ); + + const collapsed = await searchTranscripts(source, { + query: "lawsuit", + contentTypes: ["video"], + }); + assert.equal(collapsed.total, 1, "two uploads of one recording are one recording"); + assert.equal(collapsed.duplicates.collapsed, 1); + assert.equal(collapsed.duplicates.clusters, 1); + // The kept row is the canonical member, and the copy is NAMED rather than + // silently dropped. + assert.equal(collapsed.hits[0].videoId, "orig"); + assert.deepEqual( + collapsed.hits[0].mirrors?.map((m) => m.videoId), + ["copy"], + ); + + // Opting out returns one row per upload, unchanged. + const raw = await searchTranscripts(source, { + query: "lawsuit", + contentTypes: ["video"], + collapseDuplicates: false, + }); + assert.equal(raw.total, 2); + assert.equal(raw.duplicates.collapsed, 0); +}); + +test("when only the mirror matches, the mirror is kept rather than dropped", async () => { + // The canonical copy is not in the result set at all. Preferring an absent + // canonical would delete the only surviving evidence — which is exactly the + // material a 'what did the deleted videos say' question is after. + const mirror = video("chanB", "copy"); + const duplicates: DuplicateIndex = new Map([ + [ + "chanB/copy", + { + clusterId: "c1", + canonicalSlug: "chanA/orig", + isCanonical: false, + contained: false, + needsReview: false, + siblings: [], + }, + ], + ]); + const source = new CountingSource({ chanB: [[mirror]] }, new Map(), duplicates); + const result = await searchTranscripts(source, { + query: "lawsuit", + contentTypes: ["video"], + }); + assert.equal(result.total, 1); + assert.equal(result.hits[0].videoId, "copy"); + assert.equal(result.duplicates.collapsed, 0); +}); + +test("a corpus with no duplicates report reports that it could not check", async () => { + const source = new CountingSource({ chan: [[video("chan", "a1")]] }, new Map()); + const result = await searchTranscripts(source, { + query: "lawsuit", + contentTypes: ["video"], + }); + // available:true here because the stub implements the member and returns an + // empty map — "checked, none" — which is a different claim from a source + // that ships no duplicates layer at all. + assert.equal(result.duplicates.collapsed, 0); + assert.equal(result.total, 1); +}); + +// ─── 4. cap honesty ─── + +test("a capped scan names the channels it finished and the ones it never reached", async () => { + const source = new CountingSource({ + chanA: [[video("chanA", "a1")], [video("chanA", "a2")]], + chanB: [[video("chanB", "b1")]], + chanC: [[video("chanC", "c1")]], + }); + const result = await searchTranscripts(source, { + query: "lawsuit", + contentTypes: ["video"], + maxPages: 1, + }); + assert.equal(result.truncated, true); + assert.equal(result.scanned.pages, 1); + // The cap cut inside chanA, so chanB and chanC were never opened — and the + // sample is therefore channel-biased, not random. Naming them is what makes + // "partial" actionable instead of merely alarming. + assert.deepEqual(result.coverage.channelsCompleted, []); + assert.equal(result.coverage.channelStopped?.channel, "chanA"); + assert.deepEqual(result.coverage.channelsNotReached, ["chanB", "chanC"]); +}); + +test("an uncapped scan reports every channel as completed and none unreached", async () => { + const source = new CountingSource({ + chanA: [[video("chanA", "a1")]], + chanB: [[video("chanB", "b1")]], + }); + const result = await searchTranscripts(source, { + query: "lawsuit", + contentTypes: ["video"], + }); + assert.equal(result.truncated, false); + assert.deepEqual(result.coverage.channelsCompleted, ["chanA", "chanB"]); + assert.deepEqual(result.coverage.channelsNotReached, []); +}); + +// ─── exclude ─── + +test("exclude drops a video that also contains the excluded term", async () => { + const plain = video("chan", "keep", { text: "they filed a lawsuit today" }); + const worldCup = video("chan", "drop", { + text: "the lawsuit came up during the world cup", + }); + const source = new CountingSource({ chan: [[plain, worldCup]] }); + const result = await searchTranscripts(source, { + query: "lawsuit", + contentTypes: ["video"], + exclude: ["world cup"], + }); + assert.equal(result.total, 1); + assert.equal(result.hits[0].videoId, "keep"); +}); diff --git a/mcp/src/search.ts b/mcp/src/search.ts @@ -30,9 +30,24 @@ import { VIDEO_STATES, type VideoState, } from "yt-dlp-transcript-common/lib/availability"; -import type { ChannelRef, ShardSource, VideoAvailability } from "./source"; +import { mapConcurrent } from "yt-dlp-transcript-common/lib/concurrency"; +import { + DEFAULT_PAGE_CONCURRENCY, + type ChannelRef, + type ShardSource, + type VideoAvailability, +} from "./source"; -export type Snippet = { clock: string; seconds: number; text: string }; +export type Snippet = { + clock: string; + seconds: number; + text: string; + // Which layer this snippet came from, when it isn't the transcript. The + // non-timed layers (metadata / description / tags) carry seconds 0, and a + // reader has to be able to tell "the word appears in the description" from + // "the word was said at 0:00" — they license completely different citations. + scope?: LayerScope; +}; // A scope selector for a search/sweep: any mix of channel handles (slug / key / // name) and group handles (id / name). All fields are optional and additive — @@ -176,6 +191,11 @@ export type SearchHit = { contentType?: ContentType; author?: string; createdAt?: string; + // Other copies of THIS SAME RECORDING that also matched and were collapsed + // into this row (cross-platform mirrors, per duplicates.json). Present only + // when collapsing actually removed something — the mirrors are named rather + // than dropped, so the count is honest without the evidence disappearing. + mirrors?: { videoId: string; channelName: string; slug: string }[]; }; export type SearchResult = { @@ -197,9 +217,31 @@ export type SearchResult = { // nothing". `channels` counts the channels that actually HAVE a posts index, // so 0 with `requested` true means the post corpus is empty here. postsScanned: { requested: boolean; channels: number; pages: number }; + // What duplicate collapsing did to the count. `available: false` means this + // corpus ships no duplicates.json, so no claim about mirrors can be made + // either way — distinct from "checked, found none". + duplicates: { collapsed: number; clusters: number; available: boolean }; // Coverage is partial — the page cap (MAX_PAGES) or the video cap // (HARD_VIDEO_CAP) was reached before the corpus was fully scanned. truncated: boolean; + // What was actually covered, in enough detail to act on. A cap truncates in + // CHANNEL ITERATION ORDER, not at random, so "partial" on its own is + // misleading in a specific way: the sample is biased toward whichever + // channels happen to sort first. Naming the channels that were fully scanned + // and the ones never reached turns that from alarming into actionable — you + // can re-run scoped to the remainder. + coverage: { + // A filter let us plan an exact page set instead of scanning everything. + pruned: boolean; + pagesPlanned: number; + pagesTotal: number; + // Videos in a channel manifest that the summaries index didn't know about; + // their pages were read unconditionally. + unknownVideos: number; + channelsCompleted: string[]; + channelsNotReached: string[]; + channelStopped?: { channel: string; page: number; pages: number }; + }; // How the scope selector resolved (for the tool's scope note): whether it was // whole-corpus, the channels/groups it matched, and any tokens that matched // nothing (a typo'd channel/group is surfaced, not silently a full scan). @@ -230,6 +272,142 @@ function clock(seconds: number): string { return s === 0 ? "0:00" : formatDuration(s); } +// ─── Windowed page reading ─── +// +// Pages were read one `await` at a time, which on a local corpus leaves the +// disk idle for the whole ~390 ms JSON.parse of each 8 MB page, and over HTTP +// leaves the connection idle for a whole round trip. This reads a WINDOW of +// pages concurrently and then folds them in index order. +// +// Windowed rather than one big fan-out over every page, for three reasons that +// are all load-bearing: +// 1. hit order is unchanged — results are folded in page order, which a dozen +// ordering assertions in search.test.ts pin; +// 2. `max_pages` is honored exactly — the window is clamped to the remaining +// budget, so max_pages:1 still reads precisely one page; +// 3. overshoot past a cap is bounded by the window width rather than by the +// size of the corpus. +async function readPageWindow<T>( + pages: readonly number[], + concurrency: number, + read: (page: number) => Promise<T>, +): Promise<(T | null)[]> { + // A page that fails to read becomes null and is skipped by the caller — + // matching the per-page try/catch this replaces. + return mapConcurrent(pages, concurrency, (p) => + read(p).then( + (v) => v, + () => null, + ), + ); +} + +function concurrencyOf(source: ShardSource): number { + return source.pageConcurrency ?? DEFAULT_PAGE_CONCURRENCY; +} + +// ─── Filter-first scan planning ─── +// +// A filtered query's exact page set is computable BEFORE any transcript is +// read: the summaries index says which videos pass the filter, and each channel +// manifest's `slugToPage` says which page each video lives on. Measured on the +// live corpus (170 pages): the deleted/unlisted question touches 8 pages, a +// single channel 8, one upload year 28, and "not livestreams" 162 — so this +// pays in proportion to how selective the filter is, and degrades to today's +// full scan when it isn't. +// +// The load-bearing invariant: the index only ever PRUNES PAGES; `passesFilters` +// on the real record still DECIDES every hit. A page is skipped only when every +// video the manifest places on it is known to be excluded. So a stale or +// incomplete index can cost time, never correctness — and the two can't drift, +// because pruning and deciding call the same predicate. + +export type ScanPlan = { + // Per channel key: the page indices to read, ascending. Channels whose + // manifest could not be read are absent (matching the old `continue`). + perChannel: Map<string, { ch: ChannelRef; pages: number[] }>; + pruned: boolean; + pagesPlanned: number; + pagesTotal: number; + // Videos the manifest lists but the index has never heard of. Their pages are + // included unconditionally; a non-zero count here is why a "pruned" scan may + // still read more than the filter suggests. + unknownVideos: number; +}; + +// True when a filter set could actually exclude something. An all-permissive +// filter (every state kept, both media types, both audiences, no dates) is the +// same query as no filter at all, and must NOT trigger an index read — that is +// the guard against making an unfiltered query slower by planning it. +export function filterIsSelective(f: SearchFilters | null | undefined): boolean { + if (!f) return false; + if (!VIDEO_STATES.every((s) => f.states.has(s))) return true; + if (!f.videos || !f.livestreams) return true; + if (!f.allAges || !f.restricted) return true; + return Boolean(f.dateFrom || f.dateTo); +} + +// Build the page plan for a scan. Falls back to "every page of every channel" +// whenever pruning is impossible or pointless: no selective filter, or a source +// with no summaries index (both in-memory test stubs, and any site that ships +// no summaries/). +export async function buildScanPlan( + source: ShardSource, + channels: readonly ChannelRef[], + filters: SearchFilters | null | undefined, +): Promise<ScanPlan> { + const perChannel = new Map<string, { ch: ChannelRef; pages: number[] }>(); + let pagesTotal = 0; + let pagesPlanned = 0; + let unknownVideos = 0; + + const wantPrune = filterIsSelective(filters) && typeof source.videoIndex === "function"; + const index = wantPrune ? await source.videoIndex!() : null; + + for (const ch of channels) { + let manifest; + try { + manifest = await source.transcriptsManifest(ch); + } catch { + continue; // unreachable/missing channel — skip, as before + } + const pageCount = manifest.pageCount; + pagesTotal += pageCount; + + if (!index || !filters) { + const pages = Array.from({ length: pageCount }, (_, i) => i); + perChannel.set(ch.key, { ch, pages }); + pagesPlanned += pages.length; + continue; + } + + const keep = new Set<number>(); + for (const [videoId, page] of Object.entries(manifest.slugToPage)) { + if (keep.has(page)) continue; // one surviving video is enough to read it + const rec = index.get(`${ch.slug}/${videoId}`); + if (!rec) { + // Unknown to the index — summaries may be older or narrower than the + // transcripts. Read the page; the record predicate will decide. + unknownVideos++; + keep.add(page); + continue; + } + if (passesFilters(rec, filters, rec)) keep.add(page); + } + const pages = [...keep].sort((a, b) => a - b); + perChannel.set(ch.key, { ch, pages }); + pagesPlanned += pages.length; + } + + return { + perChannel, + pruned: Boolean(index), + pagesPlanned, + pagesTotal, + unknownVideos, + }; +} + export type Matcher = (text: string) => boolean; // Build the combined, alias-aware matcher for a query, mirroring the browser's @@ -280,6 +458,77 @@ function truncate(text: string, max = 240): string { return t.length > max ? t.slice(0, max - 1) + "…" : t; } +// Collapse cross-platform mirrors in a result list, IN PLACE, keeping one row +// per recording. Returns what it did so the caller can report it. +// +// Two rules make this safe to have on by default: +// +// 1. The kept row is the cluster's canonical member WHEN that member is +// itself among the matches — otherwise it is simply the first match. A +// mirror is frequently the only surviving copy of a deleted upload, and +// preferring an absent canonical would delete exactly the evidence a +// "what did the removed videos say" question is asking for. +// 2. The collapsed copies are NAMED on the row they folded into. Nothing +// vanishes; the count stops double-counting. Pass collapse_duplicates:false +// to see every upload as its own row. +// +// Timestamps are never mapped between copies here — that requires the per-pair +// `aligned` gate, and this function does not move a single second of anything. +async function collapseDuplicates( + source: ShardSource, + all: SearchHit[], + enabled: boolean, +): Promise<{ collapsed: number; clusters: number; available: boolean }> { + if (!enabled || typeof source.duplicateIndex !== "function") { + return { collapsed: 0, clusters: 0, available: false }; + } + let index; + try { + index = await source.duplicateIndex(); + } catch { + return { collapsed: 0, clusters: 0, available: false }; + } + if (index.size === 0) return { collapsed: 0, clusters: 0, available: true }; + + const repIndexOf = new Map<string, number>(); // clusterId -> index in `kept` + const kept: SearchHit[] = []; + let collapsed = 0; + for (const hit of all) { + const membership = index.get(hit.slug); + if (!membership) { + kept.push(hit); + continue; + } + const at = repIndexOf.get(membership.clusterId); + if (at === undefined) { + repIndexOf.set(membership.clusterId, kept.length); + kept.push(hit); + continue; + } + collapsed++; + const rep = kept[at]; + const repIsCanonical = index.get(rep.slug)?.isCanonical === true; + const fold = (into: SearchHit, gone: SearchHit): SearchHit => ({ + ...into, + mirrors: [ + ...(into.mirrors ?? []), + ...(gone.mirrors ?? []), + { videoId: gone.videoId, channelName: gone.channelName, slug: gone.slug }, + ], + }); + // Promote the canonical member to the representative if it turns up later; + // otherwise fold this copy into the incumbent. + kept[at] = repIsCanonical || !membership.isCanonical + ? fold(rep, hit) + : fold(hit, rep); + } + if (collapsed > 0) { + all.length = 0; + all.push(...kept); + } + return { collapsed, clusters: repIndexOf.size, available: true }; +} + // Scan a source's transcript shards for `query` (alias-aware by default), // collecting ALL matched videos up to HARD_VIDEO_CAP so counting is stable, then // returning the [offset, offset+limit) slice with a `total`/`hasMore`. Plain @@ -304,6 +553,24 @@ export async function searchTranscripts( aliases?: SearchAlias[]; // Which corpora to search. Defaults to both video transcripts and posts. contentTypes?: ContentType[]; + // The share-link filter set (availability state / upload-date range / + // media type / audience). When selective, it also drives filter-first page + // planning, so a narrow question reads a fraction of the corpus. + filters?: SearchFilters | null; + // Video-level NOT: a video matching any of these is dropped even if the + // query matched it. Evaluated over the same text the query is (cues, plus + // title), which is what makes `"cup"` minus `"world cup"` mean what a + // person means by it. + exclude?: string[]; + // Which layers of a video to match against. Omitted → today's behavior + // exactly: spoken captions plus the title. Naming scopes replaces that + // default outright, so ["description"] searches descriptions and NOT + // captions. + scopes?: LayerScope[]; + // Count a recording mirrored across platforms ONCE (default true). The + // collapsed copies are named on the row they fold into, never dropped + // silently. + collapseDuplicates?: boolean; }, ): Promise<SearchResult> { const limit = opts.limit ?? 20; @@ -324,6 +591,36 @@ export async function searchTranscripts( aliases, }); + // Video-level NOT. Each term is compiled the same way the query is (so a + // regex search excludes by regex too) but WITHOUT alias expansion: an + // exclusion is a thing the user named precisely, and quietly widening it via + // a curated alias would drop videos they never asked to drop. + const excludeMatchers = (opts.exclude ?? []) + .map((t) => (typeof t === "string" ? t.trim() : "")) + .filter((t) => t !== "") + .map( + (t) => + buildMatcher({ query: t, regex: opts.regex, useAliases: false }).match, + ); + const isExcluded = (title: string, texts: readonly { text: string }[]): boolean => + excludeMatchers.length > 0 && + excludeMatchers.some((m) => m(title) || texts.some((c) => m(c.text))); + + // Which layers to match. The default is exactly what this scanner has always + // done — spoken captions plus the title — so an existing caller's results are + // byte-identical. Naming scopes replaces that default rather than adding to + // it, which is the only reading under which ["description"] is honest. + const scopes = opts.scopes; + const wantCues = !scopes || scopes.includes("transcripts"); + const wantMetadata = Boolean(scopes?.includes("metadata")); + const wantTitle = !scopes || wantMetadata; + const wantDescription = Boolean(scopes?.includes("description")); + const wantTags = Boolean(scopes?.includes("tags")); + const wantChat = Boolean(scopes?.includes("chat")); + // Live chat is 1.2 GB of subs/ shards, so it is fetched lazily and only for + // the videos a chat-scoped query actually reaches. + const chatCuesFor = wantChat ? makeChatFetcher(source) : null; + const selection = await resolveSelectedChannels(source, { channel: opts.channel, channels: opts.channels, @@ -337,69 +634,146 @@ export async function searchTranscripts( let channelsScanned = 0; let truncated = false; const contentTypes = opts.contentTypes ?? [...ALL_CONTENT_TYPES]; - const wantVideos = contentTypes.includes("video"); - const wantPosts = contentTypes.includes("post"); + // `scopes` and `content_types` are both narrowing, so they intersect: asking + // for scopes:["description"] must not still scan the post corpus, and + // content_types:["video"] must not be widened by a posts scope. + const wantVideos = + contentTypes.includes("video") && (!scopes || scopes.some((s) => s !== "posts")); + const wantPosts = + contentTypes.includes("post") && (!scopes || scopes.includes("posts")); // Counted apart from the video pass so "no posts index anywhere in scope" is // distinguishable from "searched the posts and found nothing". let postChannelsScanned = 0; let postPagesScanned = 0; - outer: for (const ch of channels) { - if (!wantVideos) break; - let manifest; - try { - manifest = await source.transcriptsManifest(ch); - } catch { - continue; // unreachable/missing channel — skip - } - channelsScanned++; - for (let page = 0; page < manifest.pageCount; page++) { - if (pagesScanned >= maxPages) { - truncated = true; - break outer; - } - let records: TranscriptDetail[]; - try { - records = await source.transcriptPage(ch, page); - } catch { - continue; - } - pagesScanned++; - for (const rec of records) { - const titleHit = match(rec.title ?? ""); - const snippets: Snippet[] = []; - let matches = 0; - for (const cue of rec.cues ?? []) { - if (!match(cue.text)) continue; - matches++; - if (includeSnippets && snippets.length < snippetsPerVideo) { - snippets.push({ - clock: clock(cue.start), - seconds: cue.start, - text: truncate(cue.text), - }); - } - } - if (matches === 0 && !titleHit) continue; - all.push({ - videoId: rec.id, - slug: rec.slug, - channelSlug: ch.slug, - channelName: ch.name, - ...(ch.siteTitle ? { siteTitle: ch.siteTitle } : {}), - ...(ch.siteUrl ? { siteUrl: ch.siteUrl } : {}), - ...(rec.platform ? { platform: rec.platform } : {}), - title: rec.title, - uploadDate: rec.uploadDate, - webpageUrl: rec.webpageUrl, - matches: matches || 1, - snippets, - }); - if (all.length >= HARD_VIDEO_CAP) { + const filters = opts.filters ?? null; + const availability = + wantVideos && needsAvailability(filters) ? await source.availabilityMap() : null; + const plan = wantVideos + ? await buildScanPlan(source, channels, filters) + : null; + const concurrency = concurrencyOf(source); + const channelsCompleted: string[] = []; + let channelStopped: SearchResult["coverage"]["channelStopped"]; + + outer: if (plan) { + for (const entry of plan.perChannel.values()) { + const ch = entry.ch; + channelsScanned++; + for (let i = 0; i < entry.pages.length; ) { + const budget = maxPages - pagesScanned; + if (budget <= 0) { truncated = true; + channelStopped = { channel: ch.name, page: i, pages: entry.pages.length }; break outer; } + const window = entry.pages.slice(i, i + Math.min(concurrency, budget)); + i += window.length; + const loaded = await readPageWindow(window, concurrency, (p) => + source.transcriptPage(ch, p), + ); + for (let w = 0; w < loaded.length; w++) { + const records = loaded[w]; + if (records === null) continue; // failed page — skipped, as before + pagesScanned++; + for (const rec of records) { + if ( + filters && + !passesFilters(rec, filters, availability?.get(rec.slug)) + ) { + continue; + } + const snippets: Snippet[] = []; + const push = (s: Snippet): void => { + if (includeSnippets && snippets.length < snippetsPerVideo) { + snippets.push(s); + } + }; + let matches = 0; + let otherHit = false; + + if (wantCues) { + for (const cue of rec.cues ?? []) { + if (!match(cue.text)) continue; + matches++; + push({ + clock: clock(cue.start), + seconds: cue.start, + text: truncate(cue.text), + }); + } + } + // Title is matched under the default (no `scopes`) and under an + // explicit `metadata` scope; `metadata` additionally matches the + // channel name, mirroring the viewer's metadata layer. + const titleHit = wantTitle && match(rec.title ?? ""); + if (titleHit) otherHit = true; + if (wantMetadata && !titleHit && match(rec.channel ?? ch.name)) { + otherHit = true; + push({ + clock: clock(0), + seconds: 0, + text: `Channel: ${rec.channel ?? ch.name}`, + scope: "metadata", + }); + } + if (wantDescription && rec.description && match(rec.description)) { + otherHit = true; + push({ + clock: clock(0), + seconds: 0, + text: truncate(rec.description), + scope: "description", + }); + } + if (wantTags) { + const tags = (rec.tags ?? []).join(", "); + if (tags && match(tags)) { + otherHit = true; + push({ clock: clock(0), seconds: 0, text: truncate(tags), scope: "tags" }); + } + } + if (wantChat && chatCuesFor) { + for (const cue of await chatCuesFor(ch, rec)) { + if (!match(cue.text)) continue; + matches++; + push({ + clock: clock(cue.start), + seconds: cue.start, + text: truncate(cue.text), + scope: "chat", + }); + } + } + if (matches === 0 && !otherHit) continue; + if (isExcluded(rec.title ?? "", rec.cues ?? [])) continue; + all.push({ + videoId: rec.id, + slug: rec.slug, + channelSlug: ch.slug, + channelName: ch.name, + ...(ch.siteTitle ? { siteTitle: ch.siteTitle } : {}), + ...(ch.siteUrl ? { siteUrl: ch.siteUrl } : {}), + ...(rec.platform ? { platform: rec.platform } : {}), + title: rec.title, + uploadDate: rec.uploadDate, + webpageUrl: rec.webpageUrl, + matches: matches || 1, + snippets, + }); + if (all.length >= HARD_VIDEO_CAP) { + truncated = true; + channelStopped = { + channel: ch.name, + page: i - loaded.length + w, + pages: entry.pages.length, + }; + break outer; + } + } + } } + channelsCompleted.push(ch.name); } } @@ -432,6 +806,7 @@ export async function searchTranscripts( postPagesScanned++; for (const post of posts) { if (!match(post.text)) continue; + if (isExcluded("", [{ text: post.text }])) continue; all.push({ videoId: post.id, slug: post.slug, @@ -462,8 +837,24 @@ export async function searchTranscripts( } } + // ── collapse cross-platform duplicates ── + // Before slicing, so `total` is the count of distinct RECORDINGS rather than + // of uploads. A mirrored video counted twice is not a rounding error inside a + // narrow result set, and "N videos said X" is the sentence a sweep actually + // writes. + const duplicates = await collapseDuplicates( + source, + all, + opts.collapseDuplicates !== false, + ); + const total = all.length; const hits = all.slice(offset, offset + limit); + const planned = [...(plan?.perChannel.values() ?? [])].map((e) => e.ch.name); + const reached = new Set([ + ...channelsCompleted, + ...(channelStopped ? [channelStopped.channel] : []), + ]); return { hits, total, @@ -477,7 +868,17 @@ export async function searchTranscripts( channels: postChannelsScanned, pages: postPagesScanned, }, + duplicates, truncated, + coverage: { + pruned: plan?.pruned ?? false, + pagesPlanned: plan?.pagesPlanned ?? 0, + pagesTotal: plan?.pagesTotal ?? 0, + unknownVideos: plan?.unknownVideos ?? 0, + channelsCompleted, + channelsNotReached: planned.filter((n) => !reached.has(n)), + ...(channelStopped ? { channelStopped } : {}), + }, selection: { all: selection.all, channelCount: selection.channels.length, @@ -922,8 +1323,21 @@ function evalNode( // The share-filter predicate over a transcript record + its availability. // Mirrors SearchSessionContext.tsx `passesFilter` exactly. +// +// Typed on the three fields it actually reads rather than on TranscriptDetail, +// so the SAME predicate can be applied to a summaries index record while +// planning a scan and to the full transcript record while deciding a hit. +// There is deliberately only one of these: a second, "cheap" predicate for +// planning is exactly how a pruner starts silently disagreeing with the +// scanner about what matches. +type FilterableRecord = { + isLivestream?: boolean; + ageRestricted?: boolean; + uploadDate: string; +}; + function passesFilters( - rec: TranscriptDetail, + rec: FilterableRecord, f: SearchFilters, avail: VideoAvailability | undefined, ): boolean { @@ -1010,59 +1424,63 @@ export async function runSearchSpec( let channelsScanned = 0; let truncated = false; - outer: for (const ch of channels) { - let manifest; - try { - manifest = await source.transcriptsManifest(ch); - } catch { - continue; - } + // Filter-first: a link carrying an availability or date filter reads only the + // pages that can hold a match, which is the whole point of `open_link` being + // the one caller with real filters today. + const plan = await buildScanPlan(source, channels, filters); + const concurrency = concurrencyOf(source); + + outer: for (const entry of plan.perChannel.values()) { + const ch = entry.ch; channelsScanned++; - for (let page = 0; page < manifest.pageCount; page++) { - if (pagesScanned >= maxPages) { + for (let i = 0; i < entry.pages.length; ) { + const budget = maxPages - pagesScanned; + if (budget <= 0) { truncated = true; break outer; } - let records: TranscriptDetail[]; - try { - records = await source.transcriptPage(ch, page); - } catch { - continue; - } - pagesScanned++; - for (const rec of records) { - if (filters && !passesFilters(rec, filters, availability?.get(rec.slug))) { - continue; - } - const ctx: RecordCtx = { - title: rec.title ?? "", - channel: rec.channel ?? ch.name, - description: rec.description ?? "", - tags: (rec.tags ?? []).join(", "), - cues: rec.cues ?? [], - chatCues: chatCuesFor ? await chatCuesFor(ch, rec) : [], - snippetsPerVideo, - includeSnippets, - }; - const r = evalNode(spec.tree, matchers, ctx); - if (!r.match) continue; - all.push({ - videoId: rec.id, - slug: rec.slug, - channelSlug: ch.slug, - channelName: ch.name, - ...(ch.siteTitle ? { siteTitle: ch.siteTitle } : {}), - ...(ch.siteUrl ? { siteUrl: ch.siteUrl } : {}), - ...(rec.platform ? { platform: rec.platform } : {}), - title: rec.title, - uploadDate: rec.uploadDate, - webpageUrl: rec.webpageUrl, - matches: r.count || 1, - snippets: r.hits, - }); - if (all.length >= HARD_VIDEO_CAP) { - truncated = true; - break outer; + const window = entry.pages.slice(i, i + Math.min(concurrency, budget)); + i += window.length; + const loaded = await readPageWindow(window, concurrency, (p) => + source.transcriptPage(ch, p), + ); + for (const records of loaded) { + if (records === null) continue; + pagesScanned++; + for (const rec of records) { + if (filters && !passesFilters(rec, filters, availability?.get(rec.slug))) { + continue; + } + const ctx: RecordCtx = { + title: rec.title ?? "", + channel: rec.channel ?? ch.name, + description: rec.description ?? "", + tags: (rec.tags ?? []).join(", "), + cues: rec.cues ?? [], + chatCues: chatCuesFor ? await chatCuesFor(ch, rec) : [], + snippetsPerVideo, + includeSnippets, + }; + const r = evalNode(spec.tree, matchers, ctx); + if (!r.match) continue; + all.push({ + videoId: rec.id, + slug: rec.slug, + channelSlug: ch.slug, + channelName: ch.name, + ...(ch.siteTitle ? { siteTitle: ch.siteTitle } : {}), + ...(ch.siteUrl ? { siteUrl: ch.siteUrl } : {}), + ...(rec.platform ? { platform: rec.platform } : {}), + title: rec.title, + uploadDate: rec.uploadDate, + webpageUrl: rec.webpageUrl, + matches: r.count || 1, + snippets: r.hits, + }); + if (all.length >= HARD_VIDEO_CAP) { + truncated = true; + break outer; + } } } } diff --git a/mcp/src/server.ts b/mcp/src/server.ts @@ -5,17 +5,31 @@ import { type Tool, } from "@modelcontextprotocol/server"; import { transcriptToMarkdown } from "yt-dlp-transcript-common/lib/transcriptToMarkdown"; -import { formatDate } from "yt-dlp-transcript-common/lib/format"; +import { formatDate, formatDuration } from "yt-dlp-transcript-common/lib/format"; +import { manifestHasDigest } from "yt-dlp-transcript-common/lib/digests"; +import type { VideoStat } from "yt-dlp-transcript-common/lib/stats"; import { momentUrl, momentBaseUrl } from "yt-dlp-transcript-common/lib/momentUrl"; import type { Platform } from "yt-dlp-transcript-common/lib/platform"; import type { SearchAlias } from "yt-dlp-transcript-common/lib/searchAliases"; import { + VIDEO_STATES, + isVideoState, + type VideoState, +} from "yt-dlp-transcript-common/lib/availability"; +import type { LayerScope } from "yt-dlp-transcript-common/lib/searchQuery"; +import { sortGroups, resolveChannelGroupId, FALLBACK_GROUP, type ChannelGroup, } from "yt-dlp-transcript-common/lib/channelGroups"; -import type { ChannelRef, HubSite, ShardSource } from "./source"; +import { + ioStatsEnabled, + ioStatsSnapshot, + type ChannelRef, + type HubSite, + type ShardSource, +} from "./source"; import type { SourceSpec } from "./sources"; import { SourceRegistry, type ResolvedSource } from "./sourceRegistry"; import type { Post } from "yt-dlp-transcript-common/lib/posts"; @@ -28,6 +42,7 @@ import { buildMatcher, getWindowedTranscript, runSearchSpec, + type SearchFilters, type SearchResult, type SpecHit, type ScopedSnippet, @@ -167,6 +182,101 @@ const SOURCE_ARG = { }, } as const; +// The filter/scope arguments shared by `search_transcripts` and +// `enumerate_matches`. ONE constant spliced into both, for the same reason +// SOURCE_ARG is: the two tools answer the same question at different +// verbosities, so a filter reachable from one and not the other would make them +// disagree about coverage — which is precisely the failure the stateless +// rebuild set out to make impossible. +// +// Flat and orthogonal on purpose. Arbitrary boolean trees stay `open_link`'s +// job: a share link already encodes one losslessly, and asking a model to +// author a `qt=` tree in a tool call trades a real capability for a new class +// of malformed input. +const SEARCH_FILTER_ARGS = { + states: { + type: "array", + items: { + type: "string", + enum: [ + "available", + "maybe_missing", + "deleted", + "private", + "members_only", + "unlisted", + ], + }, + description: + "Keep only videos in these presence states on the source platform. " + + "This is how you ask what the DELETED videos said: " + + "states:['deleted','private','members_only','unlisted','maybe_missing'] " + + "searches everything that has since left the platform. Omitted = every " + + "state. ('maybe_missing' = fell out of the channel listing but was never " + + "individually confirmed.)", + }, + date_from: { + type: "string", + description: "Keep only videos uploaded on or after this date (YYYYMMDD).", + }, + date_to: { + type: "string", + description: "Keep only videos uploaded on or before this date (YYYYMMDD).", + }, + media_type: { + type: "string", + enum: ["video", "livestream"], + description: + "Keep only regular videos, or only livestream VODs. Omitted = both.", + }, + age: { + type: "string", + enum: ["all_ages", "restricted"], + description: + "Keep only all-ages or only age-restricted videos. Omitted = both.", + }, + exclude: { + type: "array", + items: { type: "string" }, + description: + "Drop any video that also contains one of these terms — video-level " + + "NOT. This is how you get 'cup' but not 'world cup'. Matched the same " + + "way the query is (regex too, when regex is set) but never expanded " + + "through aliases, so an exclusion stays exactly as narrow as you wrote it.", + }, + collapse_duplicates: { + type: "boolean", + description: + "Count a recording that exists on more than one platform ONCE (default " + + "true). The corpus mirrors some videos across channels/platforms, so " + + "without this a sweep counts the same recording twice. The collapsed " + + "copies are still NAMED on the row they fold into, so nothing is hidden " + + "— set false to get one row per upload instead.", + }, + scopes: { + type: "array", + items: { + type: "string", + enum: [ + "transcripts", + "chat", + "description", + "tags", + "metadata", + "posts", + ], + }, + description: + "Which layers to search. Omitted = spoken captions plus the title (the " + + "default, and what you almost always want). Naming scopes REPLACES that " + + "default: scopes:['description'] searches descriptions and not captions. " + + "'metadata' is title + channel name. NOTE 'chat' reads the live-chat " + + "shards, which are a separate ~1.2 GB corpus fetched lazily per video — " + + "it is much slower than a caption search, so scope it to a channel or a " + + "date range.", + }, +}; + // The advertised tool list. Deliberately a hand-written literal array in a // fixed order — never generated from a Map or Object.keys — so `tools/list` is // byte-stable across processes and safely cacheable (see CACHE_HINTS). @@ -214,11 +324,16 @@ export const TOOLS: Tool[] = [ "match count and whether more pages exist, so a caller can enumerate a " + "query's full match set with offset. Set include_snippets to false for a " + "cheap worklist (no cue text). The footer names the resolved scope and " + - "flags any channel/group token that matched nothing.", + "flags any channel/group token that matched nothing. Optional filters — " + + "states, date_from/date_to, media_type, age, exclude — and scopes beyond " + + "captions; with none given, behaviour is exactly as before, and a " + + "selective filter also makes the search dramatically faster by reading " + + "only the shard pages that can hold a match.", inputSchema: { type: "object", properties: { ...SOURCE_ARG, + ...SEARCH_FILTER_ARGS, query: { type: "string", description: "Term, phrase, or regex to find." }, channels: { type: "array", @@ -305,12 +420,14 @@ export const TOOLS: Tool[] = [ "would be one full scan per page. If a cap is hit the FIRST line says " + "coverage is partial and the list is a sample — a result without that " + "line is the complete set, and is the only basis on which you may state " + - "a total or claim full coverage. Same scoping and alias expansion as " + - "search_transcripts.", + "a total or claim full coverage. Same scoping, filters, exclusions, " + + "scopes and alias expansion as search_transcripts — the two always cover " + + "exactly the same set, so a count from here matches what that returns.", inputSchema: { type: "object", properties: { ...SOURCE_ARG, + ...SEARCH_FILTER_ARGS, query: { type: "string", description: "Term, phrase, or regex to find." }, channels: { type: "array", @@ -525,8 +642,17 @@ export const TOOLS: Tool[] = [ { name: "get_video_metadata", description: - "Fetch one video's metadata (title, channel, upload date, duration, " + - "description, tags, source URL) without the transcript body.", + "Everything the archive knows about one video, without the transcript " + + "body: title, channel, upload date, duration, description, tags and " + + "source URL, plus — where the site ships them — view/like/comment " + + "counts, cue count, platform state, and TRANSCRIPT COVERAGE (a warning " + + "when the transcript covers only part of the runtime, which means " + + "quotes from the tail are missing and absence of evidence is not " + + "evidence of absence). Also lists any other archived copies of the same " + + "recording, and says explicitly whether their timings are aligned — if " + + "they are not, never map a timestamp between copies. Shows AI chapters " + + "and topic tags when a digest exists, flagging one borrowed from a " + + "duplicate as describing the other upload.", inputSchema: { type: "object", properties: { @@ -797,6 +923,12 @@ export function createServer( } const source = resolved.source; + // Opt-in per-call I/O accounting for mcp/bench (MCP_IO_STATS=1). Written to + // stderr as one JSON line — stdout is the JSON-RPC channel and must stay + // clean. Off by default; this is measurement scaffolding, not telemetry. + const ioBefore = ioStatsEnabled() ? ioStatsSnapshot() : null; + const startedAt = ioBefore ? performance.now() : 0; + try { const result = await (async (): Promise<ToolResult> => { switch (name) { @@ -834,8 +966,10 @@ export function createServer( return errorText(`unknown tool: ${name}`); } })(); + reportIo(name, ioBefore, startedAt); return withCorpus(result, resolved); } catch (e) { + reportIo(name, ioBefore, startedAt); return withCorpus( errorText(`${name} failed: ${(e as Error).message}`), resolved, @@ -846,6 +980,40 @@ export function createServer( return server; } +// Emit this call's shard-read deltas as one stderr JSON line, so the benchmark +// can report the STRUCTURAL cost of a query — pages read, bytes parsed — beside +// its wall time. That separation is the point: wall time on a shared box is +// only meaningful when the box is idle, while read and byte counts are +// properties of the query plan and hold under any load. +function reportIo( + tool: string, + before: Record<string, { reads: number; bytes: number }> | null, + startedAt: number, +): void { + if (!before) return; + const after = ioStatsSnapshot(); + const delta: Record<string, { reads: number; bytes: number }> = {}; + let reads = 0; + let bytes = 0; + for (const [kind, now] of Object.entries(after)) { + const was = before[kind] ?? { reads: 0, bytes: 0 }; + const d = { reads: now.reads - was.reads, bytes: now.bytes - was.bytes }; + if (d.reads === 0 && d.bytes === 0) continue; + delta[kind] = d; + reads += d.reads; + bytes += d.bytes; + } + console.error( + `[io] ${JSON.stringify({ + tool, + ms: Math.round(performance.now() - startedAt), + reads, + bytes, + byKind: delta, + })}`, + ); +} + function channelLine(c: ChannelRef): string { const parts = [`slug: ${c.slug}`]; if (c.videoCount != null) parts.push(`${c.videoCount} videos`); @@ -927,8 +1095,13 @@ async function handleSearch( if (!query) return errorText("query is required"); const includeSnippets = args.include_snippets !== false; const base = linkStyleOf(args) === "base"; + const parsed = parseSearchArgs(args); const result = await searchTranscripts(source, { query, + filters: parsed.filters, + exclude: parsed.exclude, + scopes: parsed.scopes, + collapseDuplicates: parsed.collapseDuplicates, channel: typeof args.channel === "string" ? args.channel : undefined, channels: strArray(args.channels), group: typeof args.group === "string" ? args.group : undefined, @@ -945,6 +1118,9 @@ async function handleSearch( const aliasNote = describeFiredAliases(result.firedAliases); const scopeNote = describeScope(result.selection); const postsNote = describePostsPass(result); + const filterNote = describeSearchFilters(parsed); + const prunedNote = describeCoverage(result); + const dupNote = describeDuplicates(result); const rangeStart = result.total === 0 ? 0 : result.offset + 1; const rangeEnd = result.offset + result.hits.length; const footer = @@ -954,9 +1130,13 @@ async function handleSearch( (result.truncated ? "; coverage PARTIAL — scan hit the page/video cap" : "") + + (prunedNote ? `; ${prunedNote}` : "") + + (dupNote ? `; ${dupNote}` : "") + (scopeNote ? `; ${scopeNote}` : "") + + (filterNote ? `; ${filterNote}` : "") + (postsNote ? `; ${postsNote}` : "") + (aliasNote ? `; ${aliasNote}` : "") + + parsed.warnings.map((w) => `; ⚠ ${w}`).join("") + ")"; // The warning goes ABOVE the hits, not only in the footer. The footer said @@ -993,13 +1173,26 @@ async function handleSearch( (h.siteTitle ? ` | site: ${h.siteTitle}` : "") + ` | uploaded: ${formatDate(h.uploadDate)} | matches: ${h.matches}` + (includeSnippets && h.webpageUrl ? `\n- source: ${h.webpageUrl}` : "") + - (baseUrl ? `\n- moment_base: ${baseUrl}` : ""); + (baseUrl ? `\n- moment_base: ${baseUrl}` : "") + + (h.mirrors && h.mirrors.length > 0 + ? `\n- ⧉ same recording also archived as: ` + + h.mirrors.map((m) => `${m.videoId} (${m.channelName})`).join(", ") + + ` — counted once` + : ""); const snips = h.snippets - .map((s) => - base - ? ` - [${baseStamp(s.clock, s.seconds)}] ${s.text}` - : ` - [${stampMarkup(source, h, s.clock, s.seconds)}] ${s.text}`, - ) + .map((s) => { + // A hit in a non-timed layer (description / tags / channel name) is + // tagged with that layer and carries NO timestamp link — citing it as + // "@ 0:00" would assert someone said it at the start of the video. + if (s.scope && s.scope !== "transcripts" && s.seconds === 0) { + return ` - [${s.scope}] ${s.text}`; + } + const stamp = base + ? baseStamp(s.clock, s.seconds) + : stampMarkup(source, h, s.clock, s.seconds); + const tag = s.scope && s.scope !== "transcripts" ? `${s.scope} ` : ""; + return ` - [${tag}${stamp}] ${s.text}`; + }) .join("\n"); return snips ? `${head}\n${snips}` : head; }); @@ -1070,8 +1263,13 @@ async function handleEnumerateMatches( typeof args.batch_size === "number" ? Math.floor(args.batch_size) : 8; const batchSize = batchSizeArg >= 1 ? batchSizeArg : 8; + const parsed = parseSearchArgs(args); const result = await searchTranscripts(source, { query, + filters: parsed.filters, + exclude: parsed.exclude, + scopes: parsed.scopes, + collapseDuplicates: parsed.collapseDuplicates, // The singulars stay parsed even though they are no longer advertised. channel: typeof args.channel === "string" ? args.channel : undefined, channels: strArray(args.channels), @@ -1089,19 +1287,32 @@ async function handleEnumerateMatches( }); const partial = result.truncated; + const capNote = describeCapCoverage(result); const head = partial ? `⚠ COVERAGE PARTIAL — the scan hit its page/video cap before the corpus ` + `was exhausted. The ${result.total} match(es) below are a SAMPLE, not ` + `the full set: the true total is higher. Narrow the scope (channels/` + `groups) or raise max_pages to enumerate exhaustively, and say so in any ` + - `report built on this.\n\n` + `report built on this.\n` + + // The cap cuts in channel order, so the sample is channel-biased, not + // random. Say which channels went unread — that is what makes it fixable. + (capNote + ? `The sample is biased by CHANNEL ORDER, not random — ${capNote}.\n` + : "") + + `\n` : ""; const rows = result.hits.map((h) => { const kind = h.contentType === "post" ? "post" : "video"; + const mirrors = + h.mirrors && h.mirrors.length > 0 + ? ` | ⧉ +${h.mirrors.length} mirror(s): ${h.mirrors + .map((m) => m.videoId) + .join(", ")}` + : ""; return ( `- ${h.videoId} | ${kind} | ${h.channelName} | ` + - `${formatDate(h.uploadDate)} | ${h.matches} match(es) | ${h.title}` + `${formatDate(h.uploadDate)} | ${h.matches} match(es) | ${h.title}${mirrors}` ); }); @@ -1109,15 +1320,22 @@ async function handleEnumerateMatches( const scopeNote = describeScope(result.selection); const aliasNote = describeFiredAliases(result.firedAliases); const postsNote = describePostsPass(result); + const filterNote = describeSearchFilters(parsed); + const prunedNote = describeCoverage(result); + const dupNote = describeDuplicates(result); const footer = `\n\n(complete set: ${partial ? "NO — capped" : "yes"}; ` + `${result.total} match(es); ${batches} batch(es) of ${batchSize}; ` + `scanned ${result.scanned.pages} page(s) across ` + `${result.scanned.channels} channel(s)` + + (prunedNote ? `; ${prunedNote}` : "") + + (dupNote ? `; ${dupNote}` : "") + (scopeNote ? `; ${scopeNote}` : "") + + (filterNote ? `; ${filterNote}` : "") + (postsNote ? `; ${postsNote}` : "") + (aliasNote ? `; ${aliasNote}` : "") + + parsed.warnings.map((w) => `; ⚠ ${w}`).join("") + ")"; if (result.total === 0) { @@ -1138,6 +1356,207 @@ function describeFiredAliases(fired: SearchAlias[]): string { return `expanded via alias ${parts.join(", ")}`; } +// ─── the shared filter/scope parsing for both search tools ─── + +// The parsed form of SEARCH_FILTER_ARGS, plus any tokens that were not +// understood. Unknown tokens are REPORTED rather than dropped silently: a +// typo'd state ("removed" for "deleted") would otherwise widen the search back +// to the whole corpus and look like a legitimate empty result. +type ParsedSearchArgs = { + filters: SearchFilters | null; + exclude?: string[]; + scopes?: LayerScope[]; + collapseDuplicates: boolean; + warnings: string[]; +}; + +const SCOPE_TOKENS: ReadonlyArray<LayerScope> = [ + "transcripts", + "chat", + "description", + "tags", + "metadata", + "posts", +]; + +function parseSearchArgs(args: Record<string, unknown>): ParsedSearchArgs { + const warnings: string[] = []; + + const rawStates = strArray(args.states); + let states: VideoState[] | undefined; + if (rawStates) { + states = rawStates.filter((s): s is VideoState => isVideoState(s)); + const bad = rawStates.filter((s) => !isVideoState(s)); + if (bad.length > 0) { + warnings.push( + `unknown state(s) ignored: ${bad.join(", ")} (valid: ${VIDEO_STATES.join(", ")})`, + ); + } + if (states.length === 0) { + states = undefined; + warnings.push("no valid states given — the state filter was NOT applied"); + } + } + + const mediaType = + args.media_type === "video" || args.media_type === "livestream" + ? args.media_type + : undefined; + if (args.media_type !== undefined && mediaType === undefined) { + warnings.push(`unknown media_type ignored: ${String(args.media_type)}`); + } + const age = + args.age === "all_ages" || args.age === "restricted" ? args.age : undefined; + if (args.age !== undefined && age === undefined) { + warnings.push(`unknown age ignored: ${String(args.age)}`); + } + + const dateFrom = typeof args.date_from === "string" ? args.date_from.trim() : ""; + const dateTo = typeof args.date_to === "string" ? args.date_to.trim() : ""; + for (const [name, v] of [ + ["date_from", dateFrom], + ["date_to", dateTo], + ] as const) { + // The comparison is lexicographic on YYYYMMDD, so a differently-shaped date + // wouldn't error — it would quietly filter wrongly. Say so instead. + if (v && !/^\d{8}$/.test(v)) { + warnings.push(`${name}="${v}" is not YYYYMMDD — it was ignored`); + } + } + const from = /^\d{8}$/.test(dateFrom) ? dateFrom : undefined; + const to = /^\d{8}$/.test(dateTo) ? dateTo : undefined; + + const anyFilter = + states !== undefined || + mediaType !== undefined || + age !== undefined || + from !== undefined || + to !== undefined; + + const filters: SearchFilters | null = anyFilter + ? { + videos: mediaType === undefined || mediaType === "video", + livestreams: mediaType === undefined || mediaType === "livestream", + allAges: age === undefined || age === "all_ages", + restricted: age === undefined || age === "restricted", + states: new Set<VideoState>(states ?? VIDEO_STATES), + ...(from ? { dateFrom: from } : {}), + ...(to ? { dateTo: to } : {}), + } + : null; + + const rawScopes = strArray(args.scopes); + let scopes: LayerScope[] | undefined; + if (rawScopes) { + scopes = rawScopes.filter((s): s is LayerScope => + (SCOPE_TOKENS as readonly string[]).includes(s), + ); + const bad = rawScopes.filter( + (s) => !(SCOPE_TOKENS as readonly string[]).includes(s), + ); + if (bad.length > 0) { + warnings.push(`unknown scope(s) ignored: ${bad.join(", ")}`); + } + if (scopes.length === 0) { + scopes = undefined; + warnings.push("no valid scopes given — searching captions + title"); + } + } + + return { + filters, + exclude: strArray(args.exclude), + scopes, + collapseDuplicates: args.collapse_duplicates !== false, + warnings, + }; +} + +// A human note for the footer describing the filters actually applied, so a +// result can never be read as unfiltered when it wasn't. +function describeSearchFilters(p: ParsedSearchArgs): string { + const parts: string[] = []; + const f = p.filters; + if (f) { + if (!VIDEO_STATES.every((s) => f.states.has(s))) { + parts.push(`states: ${[...f.states].join(", ")}`); + } + if (!f.videos || !f.livestreams) { + parts.push(`type: ${f.videos ? "videos" : "livestreams"} only`); + } + if (!f.allAges || !f.restricted) { + parts.push(`audience: ${f.allAges ? "all-ages" : "age-restricted"} only`); + } + if (f.dateFrom || f.dateTo) { + parts.push(`uploaded ${f.dateFrom ?? "…"}–${f.dateTo ?? "…"}`); + } + } + if (p.exclude && p.exclude.length > 0) { + parts.push(`excluding ${p.exclude.map((e) => `"${e}"`).join(", ")}`); + } + if (p.scopes) parts.push(`scopes: ${p.scopes.join(", ")}`); + return parts.length > 0 ? `filters — ${parts.join("; ")}` : ""; +} + +// Which channels a capped scan actually covered. +// +// A cap truncates in CHANNEL ITERATION ORDER, never at random, so "partial" on +// its own hides the shape of the bias: the sample is whatever sorts first. The +// measured case that motivated this returned exactly 2000 for a common term — +// a number that reads like a count. Naming the channels that were finished, the +// one it stopped inside, and the ones it never opened turns that into something +// a caller can act on: re-run scoped to the remainder. +function describeCapCoverage(result: SearchResult): string { + const c = result.coverage; + if (!result.truncated) return ""; + const parts: string[] = []; + if (c.channelsCompleted.length > 0) { + parts.push( + `fully scanned (${c.channelsCompleted.length}): ${c.channelsCompleted.join(", ")}`, + ); + } + if (c.channelStopped) { + parts.push( + `stopped inside ${c.channelStopped.channel} at page ` + + `${c.channelStopped.page}/${c.channelStopped.pages}`, + ); + } + if (c.channelsNotReached.length > 0) { + parts.push( + `NEVER REACHED (${c.channelsNotReached.length}): ${c.channelsNotReached.join(", ")}`, + ); + } + return parts.length > 0 ? parts.join("; ") : ""; +} + +// What duplicate collapsing did, in the words a report should use. Stated +// whenever it changed the count: a total that silently differs from the one a +// previous run reported is worse than a total that explains itself. +function describeDuplicates(result: SearchResult): string { + const d = result.duplicates; + if (!d.available || d.collapsed === 0) return ""; + return ( + `${d.collapsed} mirror(s) collapsed across ${d.clusters} cluster(s) — ` + + `the ${result.total} above are distinct RECORDINGS, not uploads` + ); +} + +// How much of the corpus the scan actually had to open. Filter-first planning +// means a filtered query reads a fraction of the pages, and that has to be +// distinguishable from a scan that stopped early — "scanned 8 page(s)" reads +// like partial coverage unless it says why it was only 8. +function describeCoverage(result: SearchResult): string { + const c = result.coverage; + if (!c.pruned || c.pagesTotal === 0) return ""; + if (c.pagesPlanned >= c.pagesTotal) return ""; + return ( + `filter-pruned: planned ${c.pagesPlanned} of ${c.pagesTotal} page(s)` + + (c.unknownVideos > 0 + ? ` (+${c.unknownVideos} video(s) not in the summaries index, whose pages were read anyway)` + : "") + ); +} + // Coerce a tool arg to a trimmed non-empty string[] (drops non-strings/blanks). function strArray(v: unknown): string[] | undefined { if (!Array.isArray(v)) return undefined; @@ -1450,6 +1869,14 @@ async function handleGetTranscript( return text(md); } +// Everything the archive knows about one video, joined from the four layers +// that ship alongside the transcript: the transcript record itself, `stats/` +// (engagement + transcript coverage), `duplicates.json` (other copies of the +// same recording), and `digests/` (AI chapters and topic tags). +// +// Each layer is optional and degrades to a stated absence rather than silence: +// "no digest" and "there is no digest layer here" are different facts, and a +// caller deciding whether to trust a quote needs to be able to tell them apart. async function handleGetMetadata( source: ShardSource, args: Record<string, unknown>, @@ -1463,15 +1890,200 @@ async function handleGetMetadata( ); if (!found) return errorText(`video not found: ${videoId}`); const { cues, ...meta } = found.record; - return text( + const slug = found.record.slug; + + const lines: string[] = [ JSON.stringify( { ...meta, channelName: found.ch.name, cueCount: cues?.length ?? 0 }, null, 2, ), + ]; + + // ── stats/ ── + const stat = await lookupStat(source, slug); + if (stat) { + const engagement = [ + stat.viewCount != null ? `${stat.viewCount.toLocaleString()} views` : "", + stat.likeCount != null ? `${stat.likeCount.toLocaleString()} likes` : "", + stat.commentCount != null ? `${stat.commentCount.toLocaleString()} comments` : "", + ].filter(Boolean); + const statLines = [ + `## Stats`, + engagement.length > 0 ? `- engagement: ${engagement.join(", ")}` : "", + `- platform state: ${stat.status}`, + `- media type: ${stat.mediaType}`, + stat.duration ? `- duration: ${formatDuration(stat.duration)}` : "", + stat.cueCount != null ? `- transcript cues: ${stat.cueCount}` : "", + stat.downloadedDate ? `- downloaded: ${formatDate(stat.downloadedDate)}` : "", + stat.transcribedDate ? `- transcribed: ${formatDate(stat.transcribedDate)}` : "", + // Coverage is the one stat that changes how the transcript may be USED, + // so it is phrased as a warning rather than a number: a transcript that + // stops at 41% will answer "he never said X" wrongly and confidently. + coverageNote(stat.coverage), + ].filter(Boolean); + lines.push(statLines.join("\n")); + } + + // ── duplicates.json ── + lines.push(await describeOtherCopies(source, slug)); + + // ── digests/ ── + lines.push(await describeDigest(source, found.ch, videoId, slug)); + + return text(lines.filter(Boolean).join("\n\n")); +} + +// The transcript-coverage caveat, or "" when coverage is fine/unknown. The +// threshold is deliberately generous: values a little under 1 are normal +// (trailing silence), while a genuinely truncated download lands far below. +function coverageNote(coverage: number | null): string { + if (coverage == null) return ""; + if (coverage >= 0.9) return ""; + return ( + `- ⚠ TRANSCRIPT COVERS ONLY ${Math.round(coverage * 100)}% OF THE RUNTIME — ` + + `the download was truncated. Quotes from the uncovered tail are missing, so ` + + `do NOT conclude from this transcript that something was never said.` ); } +async function lookupStat( + source: ShardSource, + slug: string, +): Promise<VideoStat | undefined> { + if (typeof source.statsIndex !== "function") return undefined; + try { + return (await source.statsIndex()).get(slug); + } catch { + return undefined; + } +} + +// Other uploads of the SAME recording, and — the load-bearing part — whether a +// timestamp may legitimately be carried across to them. +// +// `aligned` is the gate, and absent means NOT MEASURED, which is treated as not +// aligned. A mirror with a different intro matches on text at shifted times, so +// mapping a citation into it produces a link that looks perfectly plausible and +// points at the wrong moment of a different upload. Refusing to guess is the +// only correct behaviour, and saying so is how the caller learns not to. +async function describeOtherCopies( + source: ShardSource, + slug: string, +): Promise<string> { + if (typeof source.duplicateIndex !== "function") return ""; + let membership; + try { + membership = (await source.duplicateIndex()).get(slug); + } catch { + return ""; + } + if (!membership || membership.siblings.length === 0) return ""; + + const rows = membership.siblings.map((s) => { + const timing = s.aligned + ? `timings ALIGNED${ + s.offsetSeconds != null ? ` (max observed offset ${s.offsetSeconds}s)` : "" + } — a moment link may be mapped across` + : `⚠ timing NOT measured/aligned — do NOT map timestamps into this copy; ` + + `open it at 0:00 and locate the moment again`; + return ( + `- ${s.id} (${s.channel}, ${s.platform}, ${formatDuration(s.duration)}, ` + + `uploaded ${formatDate(s.uploadDate)}` + + `${s.hasTranscript ? "" : ", NO transcript"}) — ${timing}` + ); + }); + + const caveats = [ + membership.contained + ? `- ⚠ this is a CLIP/EXCERPT relationship — the copies overlap only ` + + `partially, so nothing may be mapped across wholesale` + : "", + membership.needsReview + ? `- ⚠ UNCONFIRMED SUSPECT — matched on title and duration only; nothing ` + + `compared the actual content, so treat "same recording" as a claim, ` + + `not a fact` + : "", + membership.isCanonical + ? "" + : `- note: the canonical copy of this recording is ${membership.canonicalSlug}`, + ].filter(Boolean); + + return [ + `## Other copies of this recording (${membership.siblings.length})`, + ...rows, + ...caveats, + ].join("\n"); +} + +// The AI digest for a video, when one exists. Sparse by design — a manifest +// lists only the digested videos — so absence is the common case and is stated +// plainly rather than left as silence. +async function describeDigest( + source: ShardSource, + ch: ChannelRef, + videoId: string, + slug: string, +): Promise<string> { + if ( + typeof source.digestsManifest !== "function" || + typeof source.digestPage !== "function" + ) { + return ""; + } + let manifest; + try { + manifest = await source.digestsManifest(ch); + } catch { + return ""; + } + if (!manifestHasDigest(manifest, videoId)) return ""; + const page = manifest!.slugToPage[videoId]; + let records; + try { + records = await source.digestPage(ch, page); + } catch { + return ""; + } + const digest = records.find((r) => r.id === videoId || r.slug === slug); + if (!digest) return ""; + + const out: string[] = ["## AI digest"]; + // A borrowed digest describes ANOTHER upload. Presenting it as native is the + // failure that looks like success — every chapter reads plausibly while + // describing a different video — so this leads, before any content. + if (digest.derivedFrom) { + out.push( + `- ⚠ BORROWED: these chapters/tags were generated for ${digest.derivedFrom.slug}, ` + + `a different upload of the same recording, and shared onto this one` + + (digest.derivedFrom.offsetSeconds + ? ` (offset ${digest.derivedFrom.offsetSeconds}s)` + : "") + + `. They describe that video's timeline.`, + ); + } + if (digest.chapters.length > 0) { + out.push(`### Chapters (${digest.chapters.length})`); + for (const c of digest.chapters) { + // `start` is the cue-snapped seconds; `clock` is the model's raw string + // and is never re-parsed to seek. + out.push( + `- [${formatDuration(c.start)}] ${c.title}` + + (c.decidedBy === "human" ? " _(human-written)_" : ""), + ); + } + } + if (digest.tags.length > 0) { + out.push( + `### Topic tags\n` + + digest.tags + .map((t) => t.tag + (t.decidedBy === "human" ? " (human)" : "")) + .join(", "), + ); + } + return out.length > 1 ? out.join("\n") : ""; +} + // ─── Source discovery: list_sources / resolve_source ─── // A one-line human description of a source spec's target (kind + where it diff --git a/mcp/src/source.ts b/mcp/src/source.ts @@ -27,6 +27,21 @@ import { type SearchAlias, } from "yt-dlp-transcript-common/lib/searchAliases"; import { + resolveCanonicalSlug, + DUPLICATES_FILENAME, + type DuplicateReport, +} from "yt-dlp-transcript-common/lib/duplicates"; +import { + digestPageFileName, + type ChannelDigestsManifest, + type VideoDigest, +} from "yt-dlp-transcript-common/lib/digests"; +import { + statsPageFileName, + type StatsManifest, + type VideoStat, +} from "yt-dlp-transcript-common/lib/stats"; +import { parseChannelGroups, resolveDefaultGroupId, DEFAULT_GROUP_FALLBACK_ID, @@ -71,16 +86,163 @@ const EMPTY_GROUPS: ChannelGroups = { // page record carries — so the search engine can apply the `fav` filter. export type VideoAvailability = { state: VideoState }; +// One video as the summaries shards describe it. A strict superset of +// VideoAvailability, so the same map serves both the availability join and the +// filter-first page planner — there is one index, not two parallel reads of the +// same files. +// +// Everything here comes from `summaries/`, which is a GLOBAL index of every +// video in the corpus: 1.4 MB and ~0.5 s to parse, against 1.3 GB and ~48 s for +// the transcripts. That ratio is the whole basis of filter-first scanning — the +// exact page set a filtered query needs is computable from this plus each +// channel manifest's `slugToPage`, before a single transcript byte is read. +export type IndexedVideo = { + state: VideoState; + id: string; + channelSlug: string; + title: string; + uploadDate: string; + isLivestream: boolean; + ageRestricted: boolean; +}; + +export type VideoIndex = ReadonlyMap<string, IndexedVideo>; + +// One video's membership in a cross-platform duplicate cluster, as shipped in +// duplicates.json. Keyed by the member-local slug, like the video index. +// +// `aligned` is carried per SIBLING, not per cluster, and is the gate on ever +// translating a timestamp from one copy to another. Absent means NOT MEASURED, +// which must be read as not aligned — a mirror with a longer intro matches on +// text at shifted times, so a plausible-looking citation would land in the +// wrong place in the wrong upload. That is the failure mode that looks like +// success, and the only defence is refusing to guess. +export type ClusterMembership = { + clusterId: string; + // The member that owns derived work for this cluster (resolveCanonicalSlug), + // or null when a human marked the cluster not-a-duplicate. + canonicalSlug: string | null; + isCanonical: boolean; + // Every OTHER member of the cluster. + siblings: { + slug: string; + id: string; + channelSlug: string; + channel: string; + platform: string; + title: string; + duration: number; + uploadDate: string; + hasTranscript: boolean; + aligned: boolean; + offsetSeconds: number | null; + }[]; + // The cluster is a clip-of-a-longer-video relationship: the members overlap + // only partially, so nothing may be mapped across wholesale. + contained: boolean; + // Title+duration only — nothing compared the actual content. An unconfirmed + // suspect, not an established duplicate. + needsReview: boolean; +}; + +export type DuplicateIndex = ReadonlyMap<string, ClusterMembership>; + +// Fold a duplicates.json report into a slug → membership map. Tolerant of +// absence throughout: `compose-site.ts` only writes the file when there is at +// least one publishable cluster, and corpus.json doesn't even declare it, so a +// site legitimately ships none. +export function buildDuplicateIndex( + report: DuplicateReport | null, +): Map<string, ClusterMembership> { + const map = new Map<string, ClusterMembership>(); + if (!report || !Array.isArray(report.clusters)) return map; + for (const cluster of report.clusters) { + const refs = cluster.videoRefs ?? []; + if (refs.length < 2) continue; + const canonicalSlug = resolveCanonicalSlug(cluster); + // A human marked it not-a-duplicate — it is not a cluster any more. + if (canonicalSlug === null) continue; + for (const ref of refs) { + map.set(ref.slug, { + clusterId: cluster.clusterId, + canonicalSlug, + isCanonical: ref.slug === canonicalSlug, + contained: cluster.contained === true, + needsReview: cluster.needsReview === true, + siblings: refs + .filter((o) => o.slug !== ref.slug) + .map((o) => ({ + slug: o.slug, + id: o.id, + channelSlug: o.channelSlug, + channel: o.channel, + platform: o.platform, + title: o.title, + duration: o.duration, + uploadDate: o.uploadDate, + hasTranscript: o.hasTranscript === true, + // Alignment is a property of the PAIR, and the report records it + // against each member relative to the cluster's canonical. Both + // sides must be measured-and-aligned before a timestamp may cross. + aligned: ref.aligned === true && o.aligned === true, + offsetSeconds: o.offsetSeconds ?? null, + })), + }); + } + } + return map; +} + +// Fold the shipped stats shards into a slug → stat map. Same tolerant shape as +// the summaries read: absent or malformed is an empty map, never an error. +// +// Note the shape difference that forces a full read rather than a targeted one: +// StatsManifest carries `channels` and `pageCount` but NO `slugToPage`, so +// there is no way to jump to the page holding one video. Pages are large (up to +// STATS_MAX_PAGE_BYTES = 20 MB), which is exactly why this is lazy — nothing +// reads it until a tool asks for a stat. +async function buildStatsIndex( + readManifest: () => Promise<StatsManifest | null>, + readPage: (page: number) => Promise<VideoStat[] | null>, +): Promise<Map<string, VideoStat>> { + const map = new Map<string, VideoStat>(); + let manifest: StatsManifest | null; + try { + manifest = await readManifest(); + } catch { + return map; + } + if (!manifest || typeof manifest.pageCount !== "number") return map; + for (let page = 0; page < manifest.pageCount; page++) { + let records: VideoStat[] | null; + try { + records = await readPage(page); + } catch { + continue; + } + if (!records) continue; + for (const r of records) { + if (typeof r.slug === "string") map.set(r.slug, r); + } + } + return map; +} + // Read a site's global summaries shards (summaries/manifest.json + -// summaries/page-NNNN.json) via `readPage` and fold them into a slug → -// availability map. Tolerant: an absent/malformed manifest yields an empty map, -// and a page that fails to read is skipped. `readManifest`/`readPage` throw or -// return null on absence per the source's transport. -async function buildAvailabilityMap( +// summaries/page-NNNN.json) via `readPage` and fold them into a slug → video +// map. Tolerant: an absent/malformed manifest yields an empty map, and a page +// that fails to read is skipped. `readManifest`/`readPage` throw or return null +// on absence per the source's transport. +// +// Tolerance is load-bearing for the page planner, not just politeness: a +// summaries set that is missing, partial, or older than the transcripts must +// degrade to "I don't know about this video", and the planner's rule for +// don't-know is to scan the page anyway. +async function buildVideoIndex( readManifest: () => Promise<Manifest | null>, readPage: (page: number) => Promise<DisplaySummary[] | null>, -): Promise<Map<string, VideoAvailability>> { - const map = new Map<string, VideoAvailability>(); +): Promise<Map<string, IndexedVideo>> { + const map = new Map<string, IndexedVideo>(); let manifest: Manifest | null; try { manifest = await readManifest(); @@ -102,7 +264,15 @@ async function buildAvailabilityMap( // which matters here more than anywhere: a hub reads summaries pages // from member origins it does not control, so some of them will have // been built before `state` existed. - map.set(r.slug, { state: summaryState(r) }); + map.set(r.slug, { + state: summaryState(r), + id: r.id, + channelSlug: r.channelSlug, + title: r.title ?? "", + uploadDate: r.uploadDate ?? "", + isLivestream: r.isLivestream === true, + ageRestricted: r.ageRestricted === true, + }); } } return map; @@ -167,7 +337,92 @@ export interface ShardSource { // member-local video slug (`<channelSlug>/<id>`), built from the summaries // shards. Fetched lazily and cached — only the `fav` availability filter needs // it. Empty when the source ships no summaries. - availabilityMap(): Promise<Map<string, VideoAvailability>>; + // + // ReadonlyMap so the richer videoIndex() can BE this map rather than a + // projection of it: ReadonlyMap is covariant in its value type, so one + // Map<string, IndexedVideo> satisfies both and the two can never drift. + availabilityMap(): Promise<ReadonlyMap<string, VideoAvailability>>; + // The full summaries-backed index, when this source ships one. OPTIONAL: the + // in-memory test stubs don't implement it, and a site that ships no + // summaries/ genuinely has no index — callers must degrade to a full scan + // rather than assume an empty index means an empty corpus. + videoIndex?(): Promise<VideoIndex>; + // How many shard pages this source is willing to have in flight at once. + // Optional; callers use `?? DEFAULT_PAGE_CONCURRENCY`. Local is CPU-bound on + // JSON.parse (measured: 42 ms read vs 389 ms parse for an 8 MB page), so + // concurrency there only overlaps read with parse and saturates quickly. + // Remote is latency-bound, where it is the dominant win. + readonly pageConcurrency?: number; + // The shipped cross-platform duplicate report, folded to slug → membership. + // OPTIONAL and empty-when-absent: compose-site only writes duplicates.json + // when there is at least one publishable cluster, corpus.json does not + // declare it, and the in-memory test stubs have no such concept. A site that + // ships none must behave exactly as it does today. + duplicateIndex?(): Promise<DuplicateIndex>; + // The shipped per-video stats index (view/like counts, cueCount, transcript + // coverage), slug-keyed. OPTIONAL for the same reasons. Lazy: stats/ is one + // ~3.4 MB page here and up to 20 MB elsewhere, so it is only read when a tool + // actually asks for it. + statsIndex?(): Promise<ReadonlyMap<string, VideoStat>>; + // A channel's AI-digest manifest (digests/<slug>/manifest.json), or null when + // the channel has none. Mirrors the posts pair, including the cached negative + // — the digest corpus is SPARSE BY DESIGN (a channel with zero digests gets + // no manifest at all), so probing it per query must not cost a read per + // channel per call. OPTIONAL on the interface for the usual reason. + digestsManifest?(ch: ChannelRef): Promise<ChannelDigestsManifest | null>; + digestPage?(ch: ChannelRef, page: number): Promise<VideoDigest[]>; +} + +// Used when a source states no preference. Deliberately modest: each in-flight +// page costs its raw bytes plus ~2.7× that once parsed, and this box is shared. +export const DEFAULT_PAGE_CONCURRENCY = 4; + +// ─── I/O instrumentation (opt-in, for mcp/bench) ─── +// +// A process-wide counter of shard reads and parsed bytes, so the benchmark can +// report the STRUCTURAL cost of a query (how many pages, how many bytes) next +// to its wall time. That matters on this box specifically: wall time is only +// meaningful when the machine is idle, but read counts and byte counts are +// properties of the query plan and hold under any load. +// +// Off unless MCP_IO_STATS=1, and even then it is two integer adds per read. +export type IoStats = { reads: number; bytes: number }; + +const IO_STATS_ON = process.env.MCP_IO_STATS === "1"; + +const ioTotals: Record<string, IoStats> = {}; + +export function recordRead(kind: string, bytes: number): void { + if (!IO_STATS_ON) return; + const slot = (ioTotals[kind] ??= { reads: 0, bytes: 0 }); + slot.reads++; + slot.bytes += bytes; +} + +// A snapshot of every counter so far, for diffing across one tool call. +export function ioStatsSnapshot(): Record<string, IoStats> { + const out: Record<string, IoStats> = {}; + for (const [k, v] of Object.entries(ioTotals)) out[k] = { ...v }; + return out; +} + +export function ioStatsEnabled(): boolean { + return IO_STATS_ON; +} + +// Read a local JSON file, counting its bytes when instrumentation is on, and +// reporting the raw size so a byte-budgeted cache can account for it. +async function readLocalJsonSized<T>( + file: string, + kind: string, +): Promise<{ value: T; bytes: number }> { + const raw = await readFile(file, "utf8"); + recordRead(kind, raw.length); + return { value: JSON.parse(raw) as T, bytes: raw.length }; +} + +async function readLocalJson<T>(file: string, kind: string): Promise<T> { + return (await readLocalJsonSized<T>(file, kind)).value; } // Shape of the channels we read out of a site corpus.json (Layer 1). Kept loose @@ -178,7 +433,151 @@ type CorpusJsonChannel = { videoCount?: number; groupId?: string; }; -type SiteCorpusJson = { channels?: CorpusJsonChannel[] }; +type SiteCorpusJson = { + channels?: CorpusJsonChannel[]; + // The composed site's own declared public origin — the deployed archilyzer + // viewer these shards were built for. Present in every spec-3 corpus.json. + site?: { id?: string; title?: string; url?: string }; +}; + +// ─── Bounded promise caches ─── +// +// Two different caching problems, so two different structures: +// +// manifests — ~551 KB for the whole corpus (29 channels). Small, hot, and +// re-read constantly: an un-hinted 20-id get_transcripts batch +// used to cost ~300 manifest reads because findVideo walks every +// channel per id. Cached OUTRIGHT, no bound. +// +// pages — ~7.4 MB of raw JSON each, several times that once parsed. An +// unbounded map of these is gigabytes, so this is a small LRU. +// Its job is the 20-id batch that lands on ONE shared page (20 +// reads → 1); it is deliberately NOT sized to hold a scan, which +// visits each page exactly once and would only be paying memory +// for evictions. +// +// The bound is a RAW-BYTE budget, not an entry count, because page sizes differ +// by an order of magnitude across corpora (a 13-record VOD page is 8 MB; a +// shorts channel's page is a fraction of that). Default 48 MB, configurable +// with TRANSCRIPT_MCP_PAGE_CACHE_MB (0 disables). +// +// Why 48: a parsed page retains about 2.7× its file bytes (measured — an +// 8.09 MB page holds 21.8 MB of JS heap), so 48 MB of raw budget is roughly +// 130 MB resident. That is the most I am willing to hold on a box that also +// runs a GPU digest sweep and other agents' jobs. It is ~6 pages of this +// corpus, which covers the working set this cache exists for (a 20-id batch +// from an enumerate worklist arrives in page order and lands on 1–3 pages). +// Note what it deliberately does NOT cover: a full-corpus scan is 170 pages ≈ +// 3.7 GB retained, so there is no cache size between "6 pages" and "impossible" +// that changes the full-scan story. Filter-first scanning changes that instead. +const DEFAULT_PAGE_CACHE_MB = 48; +// Always keep at least this many entries, so a corpus whose single page exceeds +// the whole budget still caches that page rather than thrashing on it. +const MIN_CACHED_PAGES = 2; + +function pageCacheBudgetBytes(): number { + const raw = process.env.TRANSCRIPT_MCP_PAGE_CACHE_MB; + const mb = + raw === undefined || raw.trim() === "" ? DEFAULT_PAGE_CACHE_MB : Number(raw); + const safe = Number.isFinite(mb) && mb >= 0 ? mb : DEFAULT_PAGE_CACHE_MB; + return Math.floor(safe * 1024 * 1024); +} + +// What a cached loader reports back: the parsed value plus the raw byte size it +// was parsed from, which is what the budget is denominated in. +type Sized<T> = { value: T; bytes: number }; + +// An LRU keyed by string, holding PROMISES rather than values so that N +// concurrent callers for the same page coalesce onto one read — the pattern +// makeChatFetcher already uses. A rejected promise evicts itself, so a +// transient failure is never cached as a permanent one. +// +// Sizes are only known once a read resolves, so an in-flight entry counts as 0 +// and the budget is enforced on resolve. An entry evicted while still in flight +// resolves normally for whoever already holds its promise; it just isn't +// remembered. +class PageCache<T> { + private map = new Map<string, { p: Promise<T>; bytes: number }>(); + private total = 0; + constructor(private readonly maxBytes: number) {} + + take(key: string, load: () => Promise<Sized<T>>): Promise<T> { + const hit = this.map.get(key); + if (hit !== undefined) { + this.map.delete(key); + this.map.set(key, hit); // most-recently used goes last + return hit.p; + } + if (this.maxBytes <= 0) return load().then((s) => s.value); + + const entry: { p: Promise<T>; bytes: number } = { p: null as never, bytes: 0 }; + entry.p = load() + .then((s) => { + // Only account for it if we're still the live entry for this key — + // a refresh() between issue and resolve must not resurrect it. + if (this.map.get(key) === entry) { + entry.bytes = s.bytes; + this.total += s.bytes; + this.evict(); + } + return s.value; + }) + .catch((e: unknown) => { + this.drop(key, entry); + throw e; + }); + this.map.set(key, entry); + return entry.p; + } + + private drop(key: string, entry: { bytes: number }): void { + if (this.map.get(key) === entry) { + this.map.delete(key); + this.total -= entry.bytes; + } + } + + private evict(): void { + while (this.total > this.maxBytes && this.map.size > MIN_CACHED_PAGES) { + const oldest = this.map.entries().next().value; + if (oldest === undefined) break; + this.map.delete(oldest[0]); + this.total -= oldest[1].bytes; + } + } + + clear(): void { + this.map.clear(); + this.total = 0; + } +} + +// The unbounded sibling, for the small-and-hot caches (manifests). Same +// don't-memoise-a-failure rule. +class PromiseMap<T> { + private map = new Map<string, Promise<T>>(); + + take(key: string, load: () => Promise<T>): Promise<T> { + const hit = this.map.get(key); + if (hit !== undefined) return hit; + const p = load().catch((e: unknown) => { + this.map.delete(key); + throw e; + }); + this.map.set(key, p); + return p; + } + + clear(): void { + this.map.clear(); + } +} + +// Opt-out for the composed site's declared origin (see LocalSource.publicOrigin). +// Set TRANSCRIPT_PLATFORM_LINKS=1 to cite platform watch pages instead, which is +// the right answer when a local build's declared site url is not actually +// deployed. +const PREFER_PLATFORM_LINKS = process.env.TRANSCRIPT_PLATFORM_LINKS === "1"; type HubCorpusJson = { kind?: string; sites?: { siteId: string; title: string; url: string }[]; @@ -196,96 +595,104 @@ export class LocalSource implements ShardSource { readonly label: string; private aliases?: SearchAlias[]; private groups?: ChannelGroups; - private availability?: Map<string, VideoAvailability>; + private index?: Promise<Map<string, IndexedVideo>>; + private duplicates?: Promise<DuplicateIndex>; + private stats?: Promise<ReadonlyMap<string, VideoStat>>; + // The composed site's own declared origin, learned from corpus.json the first + // time the channel list is read. Undefined = not looked at yet. + private siteOrigin: string | null | undefined; + constructor(private dir: string) { this.label = `local:${dir}`; } - // A local dir has no public viewer origin — cited moments fall back to - // platform links (momentUrl). + // The deployed archilyzer viewer these shards were composed for, as declared + // by the dir's own corpus.json (`site.url`). A composed public dir is not an + // anonymous pile of JSON — it names the site it is the build output of — so + // citing that viewer is both possible and the right default: a reader + // following a citation lands in the archive, at the cited second, with the + // transcript around it, rather than on the platform page where the archive's + // whole point (that we still have a copy) is invisible. + // + // Populated by readChannels(), which every read path runs before it renders a + // link. Null when the dir ships no corpus.json (the bare directory-listing + // fallback), or when TRANSCRIPT_PLATFORM_LINKS=1 asks for platform links — + // both fall back to the platform watch page exactly as before. publicOrigin(): string | null { - return null; + return this.siteOrigin ?? null; } - async subsManifest(ch: ChannelRef): Promise<ChannelSubsManifest | null> { - try { - const raw = await readFile( + // Cached INCLUDING the negative answer, exactly like postsManifests below: + // live-chat search probes every channel in scope, and a video-only channel + // would otherwise cost one failed read per query. + private subsManifests = new PromiseMap<ChannelSubsManifest | null>(); + private digestManifests = new PromiseMap<ChannelDigestsManifest | null>(); + private subsPages = new PageCache<SubsDetail[]>(pageCacheBudgetBytes()); + private postsManifests = new PromiseMap<ChannelPostsManifest | null>(); + private transcriptManifests = new PromiseMap<ChannelTranscriptsManifest>(); + private transcriptPages = new PageCache<TranscriptDetail[]>(pageCacheBudgetBytes()); + + // Drop every cached read. Reached only through listChannels({refresh:true}) — + // the deliberate, explicit staleness escape hatch for a corpus rebuilt under + // a long-lived server. Not a TTL, on purpose. + private resetCaches(): void { + this.subsManifests.clear(); + this.digestManifests.clear(); + this.subsPages.clear(); + this.postsManifests.clear(); + this.transcriptManifests.clear(); + this.transcriptPages.clear(); + this.aliases = undefined; + this.groups = undefined; + this.index = undefined; + this.duplicates = undefined; + this.stats = undefined; + this.siteOrigin = undefined; + } + + subsManifest(ch: ChannelRef): Promise<ChannelSubsManifest | null> { + return this.subsManifests.take(ch.slug, () => + readLocalJson<ChannelSubsManifest>( path.join(this.dir, "subs", ch.slug, "manifest.json"), - "utf8", - ); - return JSON.parse(raw) as ChannelSubsManifest; - } catch { - return null; // channel ships no subs shards - } + "subsManifest", + ).catch(() => null), // channel ships no subs shards + ); } - async subsPage(ch: ChannelRef, page: number): Promise<SubsDetail[]> { - const raw = await readFile( - path.join(this.dir, "subs", ch.slug, subsPageFileName(page)), - "utf8", + subsPage(ch: ChannelRef, page: number): Promise<SubsDetail[]> { + return this.subsPages.take(`${ch.slug}:${page}`, () => + readLocalJsonSized<SubsDetail[]>( + path.join(this.dir, "subs", ch.slug, subsPageFileName(page)), + "subsPage", + ), ); - return JSON.parse(raw) as SubsDetail[]; } // Cached per channel INCLUDING the negative answer: most channels are // video-only, and a posts-covering search would otherwise re-probe every one // of them on every query. - private postsManifests = new Map<string, ChannelPostsManifest | null>(); - - async postsManifest(ch: ChannelRef): Promise<ChannelPostsManifest | null> { - const hit = this.postsManifests.get(ch.slug); - if (hit !== undefined) return hit; - let out: ChannelPostsManifest | null = null; - try { - const raw = await readFile( + postsManifest(ch: ChannelRef): Promise<ChannelPostsManifest | null> { + return this.postsManifests.take(ch.slug, () => + readLocalJson<ChannelPostsManifest>( path.join(this.dir, "posts", ch.slug, "manifest.json"), - "utf8", - ); - out = JSON.parse(raw) as ChannelPostsManifest; - } catch { - out = null; // channel ships no posts shards - } - this.postsManifests.set(ch.slug, out); - return out; + "postsManifest", + ).catch(() => null), // channel ships no posts shards + ); } async postsPage(ch: ChannelRef, page: number): Promise<Post[]> { - const raw = await readFile( + return readLocalJson<Post[]>( path.join(this.dir, "posts", ch.slug, postsPageFileName(page)), - "utf8", - ); - return JSON.parse(raw) as Post[]; - } - - async availabilityMap(): Promise<Map<string, VideoAvailability>> { - if (this.availability) return this.availability; - this.availability = await buildAvailabilityMap( - async () => { - const raw = await readFile( - path.join(this.dir, "summaries", "manifest.json"), - "utf8", - ); - return JSON.parse(raw) as Manifest; - }, - async (page) => { - const raw = await readFile( - path.join(this.dir, "summaries", pageFileName(page)), - "utf8", - ); - return JSON.parse(raw) as DisplaySummary[]; - }, + "postsPage", ); - return this.availability; } async loadAliases(): Promise<SearchAlias[]> { if (this.aliases) return this.aliases; try { - const raw = await readFile( - path.join(this.dir, "search-aliases.json"), - "utf8", - ); - this.aliases = coerceAliasConfig(JSON.parse(raw)).aliases; + this.aliases = coerceAliasConfig( + await readLocalJson(path.join(this.dir, "search-aliases.json"), "aliases"), + ).aliases; } catch { this.aliases = []; // no/invalid file — search stays plain } @@ -295,11 +702,12 @@ export class LocalSource implements ShardSource { async loadGroups(): Promise<ChannelGroups> { if (this.groups) return this.groups; try { - const raw = await readFile( - path.join(this.dir, "summaries", "manifest.json"), - "utf8", + this.groups = parseGroupsManifest( + await readLocalJson( + path.join(this.dir, "summaries", "manifest.json"), + "summariesManifest", + ), ); - this.groups = parseGroupsManifest(JSON.parse(raw)); } catch { this.groups = EMPTY_GROUPS; // no/invalid manifest — groups off } @@ -309,7 +717,10 @@ export class LocalSource implements ShardSource { private channelList?: Promise<ChannelRef[]>; listChannels(opts: { refresh?: boolean } = {}): Promise<ChannelRef[]> { - if (opts.refresh) this.channelList = undefined; + if (opts.refresh) { + this.channelList = undefined; + this.resetCaches(); + } this.channelList ??= this.readChannels().catch((e: unknown) => { this.channelList = undefined; // don't memoise a failure throw e; @@ -319,8 +730,13 @@ export class LocalSource implements ShardSource { private async readChannels(): Promise<ChannelRef[]> { try { - const raw = await readFile(path.join(this.dir, "corpus.json"), "utf8"); - const corpus = JSON.parse(raw) as SiteCorpusJson; + const corpus = await readLocalJson<SiteCorpusJson>( + path.join(this.dir, "corpus.json"), + "corpus", + ); + const declared = corpus.site?.url?.trim(); + this.siteOrigin = + declared && !PREFER_PLATFORM_LINKS ? declared.replace(/\/+$/, "") : null; if (Array.isArray(corpus.channels) && corpus.channels.length > 0) { return corpus.channels.map((c) => ({ key: c.slug, @@ -353,20 +769,88 @@ export class LocalSource implements ShardSource { return channels; } - async transcriptsManifest(ch: ChannelRef): Promise<ChannelTranscriptsManifest> { - const raw = await readFile( - path.join(this.dir, "transcripts", ch.slug, "manifest.json"), - "utf8", + transcriptsManifest(ch: ChannelRef): Promise<ChannelTranscriptsManifest> { + return this.transcriptManifests.take(ch.slug, () => + readLocalJson<ChannelTranscriptsManifest>( + path.join(this.dir, "transcripts", ch.slug, "manifest.json"), + "transcriptsManifest", + ), ); - return JSON.parse(raw) as ChannelTranscriptsManifest; } - async transcriptPage(ch: ChannelRef, page: number): Promise<TranscriptDetail[]> { - const raw = await readFile( - path.join(this.dir, "transcripts", ch.slug, transcriptPageFileName(page)), - "utf8", + transcriptPage(ch: ChannelRef, page: number): Promise<TranscriptDetail[]> { + return this.transcriptPages.take(`${ch.slug}:${page}`, () => + readLocalJsonSized<TranscriptDetail[]>( + path.join(this.dir, "transcripts", ch.slug, transcriptPageFileName(page)), + "transcriptPage", + ), ); - return JSON.parse(raw) as TranscriptDetail[]; + } + + // Local reads are CPU-bound on JSON.parse, so a modest window is all that is + // available to win: it overlaps the next page's read with this page's parse. + readonly pageConcurrency = DEFAULT_PAGE_CONCURRENCY; + + videoIndex(): Promise<VideoIndex> { + this.index ??= buildVideoIndex( + () => + readLocalJson<Manifest>( + path.join(this.dir, "summaries", "manifest.json"), + "summariesManifest", + ), + (page) => + readLocalJson<DisplaySummary[]>( + path.join(this.dir, "summaries", pageFileName(page)), + "summariesPage", + ), + ); + return this.index; + } + + availabilityMap(): Promise<ReadonlyMap<string, VideoAvailability>> { + return this.videoIndex(); + } + + digestsManifest(ch: ChannelRef): Promise<ChannelDigestsManifest | null> { + return this.digestManifests.take(ch.slug, () => + readLocalJson<ChannelDigestsManifest>( + path.join(this.dir, "digests", ch.slug, "manifest.json"), + "digestsManifest", + ).catch(() => null), // channel has no digests + ); + } + + digestPage(ch: ChannelRef, page: number): Promise<VideoDigest[]> { + return readLocalJson<VideoDigest[]>( + path.join(this.dir, "digests", ch.slug, digestPageFileName(page)), + "digestPage", + ); + } + + duplicateIndex(): Promise<DuplicateIndex> { + this.duplicates ??= readLocalJson<DuplicateReport>( + path.join(this.dir, DUPLICATES_FILENAME), + "duplicates", + ) + .then(buildDuplicateIndex) + .catch(() => buildDuplicateIndex(null)); // no report shipped — no clusters + return this.duplicates; + } + + statsIndex(): Promise<ReadonlyMap<string, VideoStat>> { + this.stats ??= buildStatsIndex( + () => + readLocalJson<StatsManifest>( + path.join(this.dir, "stats", "manifest.json"), + "statsManifest", + ), + (page) => + readLocalJson<VideoStat[]>( + path.join(this.dir, "stats", statsPageFileName(page)), + "statsPage", + ), + ); + return this.stats; } } @@ -376,10 +860,28 @@ export class RemoteSource implements ShardSource { private base: string; private aliases?: SearchAlias[]; private groups?: ChannelGroups; - private availability?: Map<string, VideoAvailability>; - constructor(baseUrl: string) { + private index?: Promise<Map<string, IndexedVideo>>; + private duplicates?: Promise<DuplicateIndex>; + private stats?: Promise<ReadonlyMap<string, VideoStat>>; + private subsManifests = new PromiseMap<ChannelSubsManifest | null>(); + private digestManifests = new PromiseMap<ChannelDigestsManifest | null>(); + private subsPages: PageCache<SubsDetail[]>; + private postsManifests = new PromiseMap<ChannelPostsManifest | null>(); + private transcriptManifests = new PromiseMap<ChannelTranscriptsManifest>(); + private transcriptPages: PageCache<TranscriptDetail[]>; + + // Over HTTP the cost is latency, not parse, so a wider window is the dominant + // win — this is where bounded concurrency actually pays. + readonly pageConcurrency = 8; + + // `budgetBytes` lets a hub divide one memory ceiling across its members + // instead of granting each member the full budget (N members × 48 MB is not a + // budget, it's N budgets). + constructor(baseUrl: string, budgetBytes = pageCacheBudgetBytes()) { this.base = baseUrl.replace(/\/+$/, ""); this.label = `remote:${this.base}`; + this.subsPages = new PageCache<SubsDetail[]>(budgetBytes); + this.transcriptPages = new PageCache<TranscriptDetail[]>(budgetBytes); } // The deployed site origin — the archilyzer viewer that owns these videos. @@ -387,44 +889,59 @@ export class RemoteSource implements ShardSource { return this.base; } - async subsManifest(ch: ChannelRef): Promise<ChannelSubsManifest | null> { - try { - const res = await fetch(`${this.base}/subs/${ch.slug}/manifest.json`); - return res.ok ? ((await res.json()) as ChannelSubsManifest) : null; - } catch { - return null; - } + private resetCaches(): void { + this.subsManifests.clear(); + this.digestManifests.clear(); + this.subsPages.clear(); + this.postsManifests.clear(); + this.transcriptManifests.clear(); + this.transcriptPages.clear(); + this.aliases = undefined; + this.groups = undefined; + this.index = undefined; + this.duplicates = undefined; + this.stats = undefined; + } + + subsManifest(ch: ChannelRef): Promise<ChannelSubsManifest | null> { + return this.subsManifests.take(ch.slug, async () => { + try { + const res = await fetch(`${this.base}/subs/${ch.slug}/manifest.json`); + return res.ok ? ((await res.json()) as ChannelSubsManifest) : null; + } catch { + return null; + } + }); } subsPage(ch: ChannelRef, page: number): Promise<SubsDetail[]> { - return this.getJson(`/subs/${ch.slug}/${subsPageFileName(page)}`); + return this.subsPages.take(`${ch.slug}:${page}`, () => + this.getJsonSized<SubsDetail[]>( + `/subs/${ch.slug}/${subsPageFileName(page)}`, + "subsPage", + ), + ); } // Cached per channel including the negative answer — otherwise every // posts-covering search costs one 404 per video-only channel. - private postsManifests = new Map<string, ChannelPostsManifest | null>(); - - async postsManifest(ch: ChannelRef): Promise<ChannelPostsManifest | null> { - const hit = this.postsManifests.get(ch.slug); - if (hit !== undefined) return hit; - let out: ChannelPostsManifest | null = null; - try { - const res = await fetch(`${this.base}/posts/${ch.slug}/manifest.json`); - out = res.ok ? ((await res.json()) as ChannelPostsManifest) : null; - } catch { - out = null; - } - this.postsManifests.set(ch.slug, out); - return out; + postsManifest(ch: ChannelRef): Promise<ChannelPostsManifest | null> { + return this.postsManifests.take(ch.slug, async () => { + try { + const res = await fetch(`${this.base}/posts/${ch.slug}/manifest.json`); + return res.ok ? ((await res.json()) as ChannelPostsManifest) : null; + } catch { + return null; + } + }); } postsPage(ch: ChannelRef, page: number): Promise<Post[]> { - return this.getJson(`/posts/${ch.slug}/${postsPageFileName(page)}`); + return this.getJson(`/posts/${ch.slug}/${postsPageFileName(page)}`, "postsPage"); } - async availabilityMap(): Promise<Map<string, VideoAvailability>> { - if (this.availability) return this.availability; - this.availability = await buildAvailabilityMap( + videoIndex(): Promise<VideoIndex> { + this.index ??= buildVideoIndex( async () => { const res = await fetch(`${this.base}/summaries/manifest.json`); return res.ok ? ((await res.json()) as Manifest) : null; @@ -434,7 +951,57 @@ export class RemoteSource implements ShardSource { return res.ok ? ((await res.json()) as DisplaySummary[]) : null; }, ); - return this.availability; + return this.index; + } + + availabilityMap(): Promise<ReadonlyMap<string, VideoAvailability>> { + return this.videoIndex(); + } + + digestsManifest(ch: ChannelRef): Promise<ChannelDigestsManifest | null> { + return this.digestManifests.take(ch.slug, async () => { + try { + const res = await fetch(`${this.base}/digests/${ch.slug}/manifest.json`); + return res.ok ? ((await res.json()) as ChannelDigestsManifest) : null; + } catch { + return null; + } + }); + } + + digestPage(ch: ChannelRef, page: number): Promise<VideoDigest[]> { + return this.getJson( + `/digests/${ch.slug}/${digestPageFileName(page)}`, + "digestPage", + ); + } + + duplicateIndex(): Promise<DuplicateIndex> { + this.duplicates ??= (async () => { + try { + const res = await fetch(`${this.base}/${DUPLICATES_FILENAME}`); + return buildDuplicateIndex( + res.ok ? ((await res.json()) as DuplicateReport) : null, + ); + } catch { + return buildDuplicateIndex(null); + } + })(); + return this.duplicates; + } + + statsIndex(): Promise<ReadonlyMap<string, VideoStat>> { + this.stats ??= buildStatsIndex( + async () => { + const res = await fetch(`${this.base}/stats/manifest.json`); + return res.ok ? ((await res.json()) as StatsManifest) : null; + }, + async (page) => { + const res = await fetch(`${this.base}/stats/${statsPageFileName(page)}`); + return res.ok ? ((await res.json()) as VideoStat[]) : null; + }, + ); + return this.stats; } async loadAliases(): Promise<SearchAlias[]> { @@ -463,18 +1030,32 @@ export class RemoteSource implements ShardSource { return this.groups; } - private async getJson<T>(p: string): Promise<T> { + // Fetches as TEXT so the byte size is knowable — the page cache's budget is + // denominated in raw bytes, and res.json() throws the length away. + private async getJsonSized<T>( + p: string, + kind: string, + ): Promise<{ value: T; bytes: number }> { const res = await fetch(`${this.base}${p}`); if (!res.ok) { throw new Error(`GET ${this.base}${p} -> ${res.status} ${res.statusText}`); } - return (await res.json()) as T; + const raw = await res.text(); + recordRead(kind, raw.length); + return { value: JSON.parse(raw) as T, bytes: raw.length }; + } + + private async getJson<T>(p: string, kind = "json"): Promise<T> { + return (await this.getJsonSized<T>(p, kind)).value; } private channelList?: Promise<ChannelRef[]>; listChannels(opts: { refresh?: boolean } = {}): Promise<ChannelRef[]> { - if (opts.refresh) this.channelList = undefined; + if (opts.refresh) { + this.channelList = undefined; + this.resetCaches(); + } this.channelList ??= this.readChannels().catch((e: unknown) => { this.channelList = undefined; // don't memoise a failure throw e; @@ -483,7 +1064,7 @@ export class RemoteSource implements ShardSource { } private async readChannels(): Promise<ChannelRef[]> { - const corpus = await this.getJson<SiteCorpusJson>("/corpus.json"); + const corpus = await this.getJson<SiteCorpusJson>("/corpus.json", "corpus"); return (corpus.channels ?? []).map((c) => ({ key: c.slug, slug: c.slug, @@ -495,11 +1076,21 @@ export class RemoteSource implements ShardSource { } transcriptsManifest(ch: ChannelRef): Promise<ChannelTranscriptsManifest> { - return this.getJson(`/transcripts/${ch.slug}/manifest.json`); + return this.transcriptManifests.take(ch.slug, () => + this.getJson<ChannelTranscriptsManifest>( + `/transcripts/${ch.slug}/manifest.json`, + "transcriptsManifest", + ), + ); } transcriptPage(ch: ChannelRef, page: number): Promise<TranscriptDetail[]> { - return this.getJson(`/transcripts/${ch.slug}/${transcriptPageFileName(page)}`); + return this.transcriptPages.take(`${ch.slug}:${page}`, () => + this.getJsonSized<TranscriptDetail[]>( + `/transcripts/${ch.slug}/${transcriptPageFileName(page)}`, + "transcriptPage", + ), + ); } } @@ -511,7 +1102,14 @@ export class HubSource implements ShardSource { readonly hubBase: string; private members = new Map<string, RemoteSource>(); // siteId -> source private aliases?: SearchAlias[]; - private availability?: Map<string, VideoAvailability>; + private index?: Promise<Map<string, IndexedVideo>>; + private duplicates?: Promise<DuplicateIndex>; + private stats?: Promise<ReadonlyMap<string, VideoStat>>; + private sites?: Promise<HubSite[]>; + + // A hub is N HTTP origins, so the latency argument for a wide window applies + // even harder than for a single remote. + readonly pageConcurrency = 8; // Optional subset allowlist of member siteIds. Undefined = federate every // member; a set restricts listChannels() to those members (site discovery via // listSites() stays unfiltered so a picker can still see all members). @@ -531,7 +1129,19 @@ export class HubSource implements ShardSource { // Fetch the hub's corpus.json and return its member sites — UNFILTERED (the // full membership), even when this source is scoped to a subset, so a picker // (list_sources / use_source) can show every member. - async listSites(): Promise<HubSite[]> { + // + // Memoised on the promise: readChannels() and videoIndex() both need it, so + // an un-memoised version fetched the hub roster twice on a cold hub search. + // A failure is not memoised. + listSites(): Promise<HubSite[]> { + this.sites ??= this.readSites().catch((e: unknown) => { + this.sites = undefined; + throw e; + }); + return this.sites; + } + + private async readSites(): Promise<HubSite[]> { const res = await fetch(`${this.hubBase}/corpus.json`); if (!res.ok) { throw new Error( @@ -542,6 +1152,15 @@ export class HubSource implements ShardSource { return hub.sites ?? []; } + // One page-cache budget for the whole hub, divided across its members — N + // members must not each get the full ceiling. Floored so a large federation + // still caches something per member. + private memberBudget(memberCount: number): number { + const total = pageCacheBudgetBytes(); + const floor = 8 * 1024 * 1024; + return Math.max(floor, Math.floor(total / Math.max(1, memberCount))); + } + // A hub can ship its own /search-aliases.json (the merged federation-wide // dictionary); if it doesn't, aliases are simply off for hub-wide search. async loadAliases(): Promise<SearchAlias[]> { @@ -574,7 +1193,17 @@ export class HubSource implements ShardSource { private channelList?: Promise<ChannelRef[]>; listChannels(opts: { refresh?: boolean } = {}): Promise<ChannelRef[]> { - if (opts.refresh) this.channelList = undefined; + if (opts.refresh) { + this.channelList = undefined; + this.sites = undefined; + this.index = undefined; + this.duplicates = undefined; + this.stats = undefined; + this.aliases = undefined; + // Members hold their own manifest/page caches; drop them wholesale so a + // refresh means the same thing federation-wide as it does locally. + this.members.clear(); + } this.channelList ??= this.readChannels().catch((e: unknown) => { this.channelList = undefined; // don't memoise a failure throw e; @@ -589,11 +1218,13 @@ export class HubSource implements ShardSource { const sites = (await this.listSites()).filter( (s) => !this.allowSiteIds || this.allowSiteIds.has(s.siteId), ); + const budget = this.memberBudget(sites.length); const all: ChannelRef[] = []; // Sequential member fetches keep it simple and polite; the channel count is // small. A failing member is skipped rather than failing the whole list. for (const site of sites) { - const remote = new RemoteSource(site.url); + const remote = + this.members.get(site.siteId) ?? new RemoteSource(site.url, budget); this.members.set(site.siteId, remote); try { const channels = await remote.listChannels(); @@ -658,33 +1289,100 @@ export class HubSource implements ShardSource { return this.memberFor(ch.siteId).postsPage(ch, page); } - // Merge each member's availability map. Keys are member-local slugs + async digestsManifest(ch: ChannelRef): Promise<ChannelDigestsManifest | null> { + if (!ch.siteId) return null; + try { + return await this.memberFor(ch.siteId).digestsManifest(ch); + } catch { + return null; // member not yet registered / unreachable + } + } + + digestPage(ch: ChannelRef, page: number): Promise<VideoDigest[]> { + if (!ch.siteId) throw new Error("hub channel ref missing siteId"); + return this.memberFor(ch.siteId).digestPage(ch, page); + } + + // Merge each member's video index. Keys are member-local slugs // (`<channelSlug>/<id>`) — the same slug a member's transcript page records - // carry — so a per-record `fav` lookup joins correctly. Built lazily/cached. - async availabilityMap(): Promise<Map<string, VideoAvailability>> { - if (this.availability) return this.availability; - const merged = new Map<string, VideoAvailability>(); + // carry — so a per-record lookup joins correctly. Built lazily/cached. + videoIndex(): Promise<VideoIndex> { + this.index ??= this.buildMergedIndex(); + return this.index; + } + + private async buildMergedIndex(): Promise<Map<string, IndexedVideo>> { + const merged = new Map<string, IndexedVideo>(); let sites: HubSite[]; try { sites = (await this.listSites()).filter( (s) => !this.allowSiteIds || this.allowSiteIds.has(s.siteId), ); } catch { - this.availability = merged; return merged; } + const budget = this.memberBudget(sites.length); for (const site of sites) { - const remote = this.members.get(site.siteId) ?? new RemoteSource(site.url); + const remote = + this.members.get(site.siteId) ?? new RemoteSource(site.url, budget); this.members.set(site.siteId, remote); try { - for (const [slug, avail] of await remote.availabilityMap()) { - merged.set(slug, avail); + for (const [slug, rec] of await remote.videoIndex()) { + merged.set(slug, rec); } } catch { // skip an unreachable member } } - this.availability = merged; + return merged; + } + + availabilityMap(): Promise<ReadonlyMap<string, VideoAvailability>> { + return this.videoIndex(); + } + + // Duplicate clusters are detected WITHIN a site, so federating them is a + // merge of per-member maps and nothing more — this deliberately does not try + // to detect mirrors ACROSS member sites. Two sites holding the same recording + // is a real thing, but nothing has compared their transcripts, and inventing + // a cross-site cluster here would be asserting a duplicate no detector ever + // confirmed. + duplicateIndex(): Promise<DuplicateIndex> { + this.duplicates ??= this.mergeMembers((m) => m.duplicateIndex()); + return this.duplicates; + } + + statsIndex(): Promise<ReadonlyMap<string, VideoStat>> { + this.stats ??= this.mergeMembers((m) => m.statsIndex()); + return this.stats; + } + + // Merge one lazily-read layer across every member site, keyed by the + // member-local slug — the same shape and the same tolerance as + // buildMergedIndex (an unreachable member is skipped, not fatal). + private async mergeMembers<T>( + read: (m: RemoteSource) => Promise<ReadonlyMap<string, T>>, + ): Promise<Map<string, T>> { + const merged = new Map<string, T>(); + let sites: HubSite[]; + try { + sites = (await this.listSites()).filter( + (s) => !this.allowSiteIds || this.allowSiteIds.has(s.siteId), + ); + } catch { + return merged; + } + const budget = this.memberBudget(sites.length); + for (const site of sites) { + const remote = + this.members.get(site.siteId) ?? new RemoteSource(site.url, budget); + this.members.set(site.siteId, remote); + try { + for (const [k, v] of await read(remote)) merged.set(k, v); + } catch { + // skip an unreachable member + } + } return merged; } }