Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit bbb4ed7e3f018f3e750f897df94dd0da8a8adb90
parent d91bb8dc16097a49752eb83a5eb09b47d7b31bcb
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Tue, 11 Aug 2026 00:36:34 -0400

Merge feat/mcp-stateless: the MCP is stateless by construction, and coverage is mechanical

The server holds no active corpus — a persisted one had been silently sending
every call to the wrong archive — so every read tool takes its own `source`
handle and every result names the corpus it read. A query's full match set
comes back in one call rather than by paging that nobody did. A pasted request
reaches the tool whole instead of being word-split into ceil(N / Why). And it
speaks the 2026-07-28 protocol, with both eras pinned by a real-process test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Diffstat:
A.claude/commands/ask.md | 20++++++++++++++++++++
A.claude/commands/sweep.md | 39+++++++++++++++++++++++++++++++++++++++
Mexport/CHANGELOG.md | 8++++++++
Mmcp/README.md | 301++++++++++++++++++++++++++++++++++++++++++++++++-------------------------------
Mmcp/package.json | 3++-
Mmcp/src/index.ts | 36+++++++++++++++++++++++-------------
Amcp/src/instructions.test.ts | 175+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Amcp/src/instructions.ts | 304+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Amcp/src/promptRequest.test.ts | 265+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Amcp/src/promptRequest.ts | 413+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Amcp/src/protocol.test.ts | 212+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mmcp/src/search.test.ts | 350+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++--
Mmcp/src/search.ts | 17+++++++++++++++++
Mmcp/src/server.ts | 1373++++++++++++++++++++++++++++++++++++++++++++++---------------------------------
Mmcp/src/shareLink.ts | 8+++++++-
Mmcp/src/source.ts | 49+++++++++++++++++++++++++++++++++++++++++++++----
Dmcp/src/sourceController.test.ts | 281-------------------------------------------------------------------------------
Dmcp/src/sourceController.ts | 140-------------------------------------------------------------------------------
Amcp/src/sourceRegistry.test.ts | 468+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Amcp/src/sourceRegistry.ts | 317+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mpnpm-lock.yaml | 491+++++--------------------------------------------------------------------------
21 files changed, 3673 insertions(+), 1597 deletions(-)

diff --git a/.claude/commands/ask.md b/.claude/commands/ask.md @@ -0,0 +1,20 @@ +--- +description: Ask a question of the transcript corpus and answer it with citations +--- + +Call the `mcp__archilyzer__ask_plan` tool with this exact `request` string, +passed through **verbatim and unsplit**: + +$ARGUMENTS + +Then follow the plan it returns, starting with its ⚠ warnings if it has any. + +<!-- +Same shape as `/sweep` — see `.claude/commands/sweep.md` for why the request +goes through a tool rather than an MCP prompt argument, and for the two editing +caveats (no `$1` in this body; the `mcp__archilyzer__` prefix is the server name +as registered on this machine). + +Use `/ask` for a question to be answered in the conversation, and `/sweep` for a +systematic pass that writes a report file. +--> diff --git a/.claude/commands/sweep.md b/.claude/commands/sweep.md @@ -0,0 +1,39 @@ +--- +description: Sweep the transcript corpus for something and build a cited report +--- + +Call the `mcp__archilyzer__sweep_plan` tool with this exact `request` string, +passed through **verbatim and unsplit**: + +$ARGUMENTS + +Then follow the plan it returns, starting with its ⚠ warnings if it has any. + +<!-- +Why this file exists +-------------------- +Claude Code parses an MCP *prompt*'s arguments as `text.trim().split(/\s+/)` +zipped against the declared argument names. It is not quote-aware, the last +argument does not absorb the remainder, and tokens past the declared count are +dropped silently. With the nine arguments the `sweep` prompt declares, a real +request became link="This" channel="search" group="deleted" directive="videos." +batch_size="Why" — and everything from the actual question onward vanished. + +`$ARGUMENTS` in a slash command is different: it is replaced with the entire +raw argument string, untokenised, so URLs, `?`, `=`, `&` and punctuation all +survive. A tool argument is structured JSON, so it arrives intact from there. + +Two things to keep in mind when editing this file: + + - Do NOT also use `$1` / positional placeholders in this body. Those DO + tokenize, and mixing them re-introduces the bug this file exists to avoid. + - The tool name embeds the MCP server name **as registered on this machine** + (`archilyzer`). If you registered the server under another name, change the + `mcp__<name>__sweep_plan` prefix to match. `foo:bar` namespacing is + plugin-only, so the command you type is `/sweep`, not `/archilyzer:sweep`. + +The sweep discipline itself — citation format, extractor budget, base-link +expansion, the coverage rule — deliberately lives in the server +(`mcp/src/instructions.ts`), not here, so it cannot drift out of sync with the +tools it names. +--> diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md @@ -1,5 +1,13 @@ # Changelog +## [0.8.5] - 2026-08-11 +- **MCP: the server no longer remembers which corpus you're reading — because remembering it was silently getting it wrong.** `use_source` switched a mutable "active corpus" and persisted the choice to a state file so it survived reconnects. That was the bug. The server registered here runs `--local …/export/public`, but its state file held `{"activeSpec":{"kind":"remote","url":"https://hasanalyzer.pages.dev"}}` from some earlier session — so **every call since had been reading a different archive, and nothing in any result said so**. There is now no active source and nothing is persisted: **every read tool takes its own `source` handle**, and a call that omits it reads the server's startup corpus. The handle *is* the serialised spec in canonical form — `default`, `local:/dir`, `remote:https://site`, `hub:https://hub`, or `hub:https://hub#alpha,beta` for a subset — not an opaque token, so it survives a restart and a human reading one in a transcript knows exactly what was searched. Shorthands (a bare site or hub URL, probed to tell one from the other; a hub member's siteId or title) normalise to canonical and are echoed back. **Every result now ends with `(corpus: <handle>)`** — errors included, implemented once in the dispatch wrapper so a new tool cannot forget it; it is `corpus:` and not `source:` because `- source:` already means "this video's URL" in the output. `use_source` survives one release as an unadvertised alias that resolves a target and tells you the handle to pass; `reset_source` is gone. Leftover state files are inert and can be deleted. Caching moved with it: a source instance is built once per handle and `listChannels` is memoised per instance, which also fixes a pre-existing cost — a 20-id `get_transcripts` batch against a remote used to fetch `corpus.json` twenty times. See `mcp/src/sourceRegistry.ts` (new; `sourceController.ts` deleted), `mcp/src/{server,source,index}.ts`, `mcp/src/sourceRegistry.test.ts`. +- **MCP: claiming you covered the corpus is now hard to do by accident.** A sweep returned `total 319; showing 1–200; has_more: yes`, was never paged, and reported **319 videos swept** having seen 200. The paging instruction was already in the sweep prompt and was ignored — so prose is not the enforcement mechanism. Worse, paging was also the *expensive* option: the engine materialises the entire match set and only then slices, so each page re-scanned the whole corpus. New **`enumerate_matches`** returns a query's complete id/title/channel/date worklist plus the batch count in **one scan** — a new tool rather than a flag, because a flag that silently changes the output shape is exactly what gets ignored. If a cap is hit, the **first line** reads `⚠ COVERAGE PARTIAL … this is a SAMPLE, not the full set`, never a quiet footnote. `search_transcripts` now prints `⚠ INCOMPLETE PAGE — N total, showing a–b. Do NOT report a count from this page.` **above** the hits (the old footer stays, so existing consumers keep working). Verified on the real 30-channel corpus: enumerate and search agree exactly at 57, 86 and at the 2000-video cap, where both correctly flag partial coverage. +- **MCP: `/sweep` and `/ask` stop shredding what you type.** Claude Code parses an MCP prompt's arguments as whitespace-splitting zipped against the declared argument names — not quote-aware, the last argument does *not* absorb the remainder, and tokens past the declared count are **dropped silently**. With nine declared arguments, a real request became `link="This"`, `channel="search"`, `group="deleted"`, `directive="videos."`, `batch_size="Why"` — rendered literally into `` ceil(N / Why) `` — and the entire actual question vanished without a warning. The entry points are now **tools** (`sweep_plan`, `ask_plan`) taking one free-text `request`, called by two thin `.claude/commands/` shims that pass `$ARGUMENTS`, the whole raw string, untokenised. A URL with `?a=b&c=d` and a full sentence of punctuation now arrive intact. Settings are still available as `key=value`, but against a **closed whitelist**: an unrecognised `x=y` **stays in the question** and warns (with a typo hint if it's one edit from a real key) instead of being eaten, and every value is validated — `batch_size` an integer 1–20, `parse_model` a single token, `report` a single `.md` path with no `..` — so a bad one becomes a default plus a `⚠` line, never arithmetic. The MCP `sweep` prompt still serves form-based clients (Claude Desktop, Cursor) through the same parser. The plans also **pre-resolve** what they can — the corpus handle, your channel/group tokens validated against the live corpus, the group roster when you gave no scope — turning a three-call preamble into none, and an unresolvable channel now **halts** the plan rather than quietly widening it. See `mcp/src/{promptRequest,instructions}.ts` (new) and their tests. +- **MCP: three parameters that were declared and silently ignored now work.** `open_link`'s `overrides.query_scope:"posts"` was in the schema, dropped by the parser and excluded by the type, so asking to re-target a link at the social-post corpus quietly searched transcripts instead. `get_transcripts.content_types` was declared and never read, so a batch of post ids came back "not found" — ids now fall through to the post corpus. And a posts search over channels with no posts index reported `total 0; scanned 0 page(s) across 0 channel(s)`, indistinguishable from "searched everything, found nothing"; it now says **"no channel in scope ships a posts index — the post corpus is EMPTY here, not merely unmatched"**, with the posts pass counted separately from the video pass. +- **MCP: `open_link` does the whole job in one call, and `get_transcripts` can carry several queries.** `open_link` was preview-then-`apply:true`, where apply *switched the global active source* — one extra round trip and the mutation this release exists to remove. It now decodes, resolves the origin to a handle, searches, and returns plan + results + handle together; `dry_run:true` gets the plan alone. `get_transcripts` gains **`queries`** (up to 8): the windows merge in one pass per video and each header reports a **per-query count**, so a term that matched nothing in that video is visible rather than absorbed — which kills the re-read-per-quote pattern. +- **MCP: on the 2026-07-28 protocol.** The server now speaks MCP revision 2026-07-28 via `@modelcontextprotocol/server@2`'s `serveStdio`, which owns the era decision — `modern` (negotiated by `server/discover`) or `legacy` (the 2025 `initialize` handshake) — and serves both from one definition. Confirmed negotiating `modern` in practice, not just in principle. Cache hints are a construction-time policy (`tools/list`, `prompts/list` and `server/discover` are literal constants with no corpus data, so `public` for an hour; every read result stays uncached) and an invalid one throws at startup rather than on the wire. Because `InMemoryTransport` only ever exercises the 2025 era, a new `protocol.test.ts` spawns the real process over stdio and asserts both eras serve an identical, order-pinned tool list. See `mcp/src/protocol.test.ts`. + ## [0.8.4] - 2026-08-04 - **A video that has gone missing now says so, and says how confidently.** Availability used to be three unlabelled buckets — available, unlisted, deleted — with no marking on the result cards themselves: a deleted video looked exactly like a live one unless you already suspected something and went hunting in the filters. It is now a **state on every card**: `DELETED`, `PRIVATE`, `MEMBERS`, `UNLISTED`, or `MISSING?`. The last one is new and it is the point of the change. Checking a channel's listing is cheap (one request per channel); confirming *why* an individual video vanished is slow, so on a large channel there is a long window where the archive knows a video has dropped out of its channel but not yet what happened to it. The site used to spend that window insisting the video was fine. It now shows **`MISSING?`** — amber, with the question mark, because it is a suspicion and not a finding — and swaps in the confirmed reason once the per-video check catches up. A video re-checked after the scan that turns out to be present simply loses the flag. - **The availability filter follows the same shape.** "Available" and "Missing" are now a parent and its five leaves (unconfirmed, deleted, private, members-only, unlisted); ticking the parent takes all five, and it shows a dash when you have only some. Unlisted sits under "missing" because the rule is one sentence — *it left the channel's listing* — even though an unlisted video is still watchable by direct link. **Saved filter profiles and old share links keep working**: an existing link that asked for available/unlisted/deleted still means "all of them" under the new taxonomy rather than silently narrowing. diff --git a/mcp/README.md b/mcp/README.md @@ -13,14 +13,16 @@ reads the site's already-published static JSON shards (`corpus.json` + | Tool | What it does | |------|--------------| | `list_channels` | List channels **organized under their channel groups** (name, slug, video count; site in hub mode), with a compact group cheat-sheet (`id · name · N channels`) for scoping. | -| `search_transcripts` | Search captions for a term/phrase (or regex); returns matching videos with timestamped snippets — **each `[mm:ss]` is a clickable link to that exact moment** (or a compact `[mm:ss\|sec]` with `link_style:"base"`). Alias-aware and **pageable** (`total` + `offset`). Scope by one or more channels (`channel`/`channels`) and/or channel groups (`group`/`groups`). | -| `get_transcripts` | Batch-read up to 20 videos in one call — bounded, timestamped **excerpt windows** around a query's matches (linked `[mm:ss]`, or compact `link_style:"base"` stamps; `max_lines` caps the excerpt), or full transcripts without a query. `channel`/`channels` hints speed the lookup. | +| `search_transcripts` | Search captions for a term/phrase (or regex); returns matching videos with timestamped snippets — **each `[mm:ss]` is a clickable link to that exact moment** (or a compact `[mm:ss\|sec]` with `link_style:"base"`). Alias-aware and pageable. A page that isn't the whole match set is flagged **above** the hits. | +| `enumerate_matches` | A query's **complete** match set as a worklist (id/title/channel/date + batch count) in **one scan**. The tool to use whenever you need to count or cover everything. | +| `get_transcripts` | Batch-read up to 20 videos in one call — bounded, timestamped **excerpt windows** around one query or up to 8 (`queries`), with per-query counts; or full transcripts without a query. Reads posts too. | | `get_transcript` | One video's full transcript as clean markdown (metadata + **linked** timestamped captions). | +| `get_post` / `get_thread` | One archived social post, or its whole thread. Posts have no timeline — cite them with no `@ mm:ss`. | | `get_video_metadata` | One video's metadata (title, channel, date, duration, description, tags, source URL) without the transcript body. | -| `open_link` | Paste an archilyzer viewer **share link** to re-run that exact search here (query tree + every filter, at full fidelity). **Previews** a plan by default; **applies** it (switch source + search, linked results) on `apply:true`. Adjust in natural language via `overrides`. | -| `list_sources` | Show the **active** corpus (label + kind + target) and, with a hub context, its member sites (`siteId · title · url`, marking the current subset) so you can pick one to switch to. | -| `use_source` | **Switch which corpus is read**, on the fly — a hub member (`site`), a hub subset (`sites`), or an arbitrary `remote` URL / `local` dir / `hub` URL. Persists across reconnects. | -| `reset_source` | Return to the source the server was started with and clear the persisted selection. | +| `open_link` | Paste an archilyzer viewer **share link** to re-run that exact search here (query tree + every filter, at full fidelity) — plan, results and corpus handle in **one** call. `dry_run:true` for the plan alone. | +| `list_sources` | Show the **default** corpus and, with a hub, its member sites as ready-to-paste handles. | +| `resolve_source` | Turn a URL or site name into the canonical `source` handle and check it can be read. Changes nothing. | +| `sweep_plan` / `ask_plan` | Turn a plain-English request (plus an optional pasted link) into a resolved, step-by-step plan. What `/sweep` and `/ask` call. | **Clickable moment links.** Every timestamp the read tools emit is a Markdown link to the exact second — an **archilyzer viewer** deep link (`…/?v=<slug>&t=<sec>`, @@ -46,22 +48,29 @@ inline-linked (future work). Beyond `query`, `regex`, and `limit`: -- **Scope** — restrict the scan to a subset of the corpus. All scope fields are +- **Scope** — restrict the scan to a subset of the corpus. Both scope fields are optional and **additive** (the search runs over the union); with none, it covers everything. - - **`channel`** / **`channels`** — one or a list of channel slugs/names. - - **`group`** / **`groups`** — one or a list of channel groups, matched by - **group id or display name** (case-insensitive), expanded to the channels in - that group. e.g. `group="other"` and `group="Extended Universe"` resolve the - same set. (In hub mode groups are deferred — a group token reports "unknown" - while channel scoping still works.) + - **`channels`** — a list of channel slugs/names. + - **`groups`** — a list of channel groups, matched by **group id or display + name** (case-insensitive), expanded to the channels in that group. e.g. + `groups:["other"]` and `groups:["Extended Universe"]` resolve the same set. + (In hub mode groups are deferred — a group token reports "unknown" while + channel scoping still works.) + - The singular `channel`/`group` are no longer advertised (the engine always + unioned them anyway) but are still parsed, so an old habit doesn't break. The footer names the resolved scope (e.g. *"scope: group Extended Universe (7 channels)"* or *"scope: 6 channels"*) and flags any channel/group token that matched nothing — so a typo is surfaced, not silently a whole-corpus scan. -- **`offset`** (default 0) — skip this many matches. The result footer reports the - full `total` and `has_more`, so you can enumerate a query's *entire* match set: - page with `offset += limit` until `has_more` is `no`. +- **`offset`** (default 0) — skip this many matches. A page that is not the whole + match set is headed with an unmissable **`⚠ INCOMPLETE PAGE — N total, showing + a–b. Do NOT report a count from this page.`** (the footer's `total`/`has_more` + are still there for existing consumers). **Do not page this to cover a query + — use `enumerate_matches`**: the engine materialises the whole match set and + *then* slices, so paging costs one full corpus scan per page while enumerating + costs one, total. Paging was the documented advice and it produced a report + claiming 319 videos swept after seeing 200. - **`include_snippets`** (default true) — set `false` for a cheap worklist (id / title / channel / date / match count, no cue text — the per-video `- source:` line is dropped too). Ideal for the planning pass of a sweep. @@ -86,6 +95,14 @@ optional `regex`/`use_aliases`), each transcript is reduced to bounded windows o timestamped lines around the matches (`before`/`after` seconds, default 30) — high-signal context for folding a batch into a report. Without a query, each video's full transcript comes back as markdown. Missing ids are reported inline. + +- **`queries`** (max 8, unioned with `query`) — window around several terms in + **one** pass. The windows merge per video, and each video's header reports a + **per-query count**, so a term that matched nothing in that video is visible + rather than absorbed into the merged excerpt. Use this instead of re-reading + the same videos once per term. +- **`content_types`** — `"video"` and/or `"post"` (default both). An id that + isn't a video falls through to the post corpus, so a mixed batch works. `channel`/`channels` are optional owning-channel hints that speed the per-id lookup when a batch spans several channels (the ids are already scoped by the search that produced them). @@ -105,97 +122,111 @@ AND/OR/NOT), and the filter block (`fc` channels, `ft` type, `fa` audience, `fav` availability, `fdf`/`fdt` upload-date range). `open_link` re-runs that exact search here — no manual source-switching or query reconstruction, nothing lost. -It is a **preview → adjust → apply** flow: - -- **Preview** (default, `apply:false`) — decode the link and return a *plan*: the - resolved source (origin, hub vs single-site, auto-probed from `corpus.json`), - the query tree rendered readably, every active filter, the channel scope - validated against the live corpus, and anything ignored or warned (the `fk` - subtitle-track token is vestigial in the composite share model — decoded and - reported as ignored). **No source switch, no search yet.** +**One call does the whole job.** `open_link` decodes the link, resolves its +origin to a corpus handle (hub vs single-site, auto-probed from `corpus.json`), +validates the channel scope against that live corpus, and returns the plan, the +first page of **linked** results, and the handle to pass as `source` from then +on. It switches nothing — the old `apply:true` mutated a global active source, +which is exactly the footgun this rebuild removed. + +- **The plan** names the resolved corpus handle, the query tree rendered + readably, every active filter, the validated channel scope, and anything + ignored or warned (the `fk` subtitle-track token is vestigial in the composite + share model — decoded and reported as ignored). +- **`dry_run: true`** returns the plan *without* searching, for confirming a + link's scope first. - **Adjust** — re-call with structured `overrides` to honor a natural-language edit: `clear_availability` (drop the availability filter), `clear_type`, `clear_age`, `clear_dates`, `clear_filters`, `clear_channels`, `channels:[…]` (re-scope), `date_from`/`date_to`, or `query`/`regex`/ - `query_scope` (replace the search). The plan updates in place. -- **Apply** (`apply:true`) — switch the active source to the link's origin (a - single-site remote, or a hub when the origin federates), run the search at full - fidelity (the query tree + filters, honoring the chat and availability scopes), - and return the first page of **linked** results. The applied plan is echoed at - the top for transparency. Pageable via `limit`/`offset`. Still read-only. + `query_scope` (replace the search; `query_scope:"posts"` targets the social-post + corpus — declared but silently dropped until now). +- Pageable via `limit`/`offset`. Read-only throughout. ``` open_link link="https://rekietalyzer.pages.dev/?qt=…&fv=1&fc=Rekieta%20Law&ft=v&fav=a" -open_link link="…" overrides={ "clear_availability": true } # "drop the availability filter" -open_link link="…" apply=true # switch source + search +open_link link="…" dry_run=true # plan only +open_link link="…" overrides={ "clear_availability": true } # "drop the availability filter" ``` -## The `sweep` prompt +## `/sweep` and `/ask` -A first-class slash command that turns **Claude Code itself** into the corpus -sweep engine — the same "batch matching transcripts into a running report" the -browser does with a BYO AI key, but driven by your Claude **plan usage** (no API -key) and with the report written to a file. - -Invoke it in Claude Code as `/mcp__<server-name>__sweep`. Arguments: `query` -(required *unless* a `link` is given), `link?` (an archilyzer viewer share URL to -seed the sweep from), `channel?`, `channels?` (comma-separated slugs/names), -`group?` (a channel group id or name), `directive?` (default *"key claims & -contradictions"*), `batch_size?` (default 8), `parse_model?` (default -`haiku` — the model requested for the per-batch extractor subagents; they only -quote verbatim, so the cheapest model wins), `report_path?` (default -`./sweep-report.md`). +Two entry points that turn **Claude Code itself** into the corpus sweep engine — +the same "batch matching transcripts into a running report" the browser does +with a BYO AI key, but driven by your Claude **plan usage** (no API key), with +the report written to a file. ``` -/mcp__rekietalyzer__sweep query="k cups" channel="chrissie-mayr" -/mcp__rekietalyzer__sweep query="k cups" group="other" # a whole group -/mcp__rekietalyzer__sweep query="k cups" group="Extended Universe" # …by name -/mcp__rekietalyzer__sweep query="k cups" # pick-first (see below) -/mcp__rekietalyzer__sweep link="https://…/?qt=…&fv=1&fc=…" # seed from a share link +/sweep https://hasanalyzer.pages.dev/?qt=…&fav=deleted This search finds deleted + videos. Why might have Rekieta privated these? batch_size=12 +/ask what did they actually say about the settlement? channels=rekieta-law ``` -**Seed from a share link (`link=`).** Pass a viewer share URL instead of a -`query` and the sweep starts by driving `open_link`: it previews the decoded plan -(source, query tree, filters, scope, ignored bits), lets you confirm or adjust in -natural language, then applies it — switching source and enumerating the full -tree+filter match set — before batching. Everything the link encodes is honored. - -**Pick-first when no scope is given.** Because a whole-corpus sweep can be a lot -of work, invoking `sweep` **without** a `channel`/`channels`/`group` makes Claude -FIRST call `list_channels`, present the groups and their channels, and ask you -which group(s)/channel(s) to sweep — or to confirm **all** for the whole corpus — -*before* it enumerates anything. Supplying any scope arg skips the prompt and -sweeps that scope directly. A group selector accepts either an id or the display -name (`group="other"` ≡ `group="Extended Universe"`). - -Once scoped, the prompt instructs Claude Code to: **search** (aliases -auto-expand; the footer names the resolved scope) → **enumerate** the full -worklist by paging with `include_snippets:false` until `has_more` is false → -**plan** `ceil(N / batch_size)` batches → **per batch, map-reduce** → finish with -a short summary. The MCP stays read-only; only the report file is written, in -Claude's working directory. +Type the whole thing on one line, in your own words, punctuation intact. Paste a +share link anywhere in it. + +**Why a tool and not a prompt argument.** Claude Code parses an MCP *prompt*'s +arguments as `text.trim().split(/\s+/)` zipped against the declared argument +names. It is not quote-aware, the last argument does **not** absorb the +remainder, and tokens past the declared count are **dropped silently**. With the +nine arguments `sweep` declares, the request above became `link="This"`, +`channel="search"`, `group="deleted"`, `directive="videos."`, `batch_size="Why"` +— rendered literally into `` ceil(N / Why) `` — and everything from *"Rekieta"* +onward vanished with no warning. + +So `.claude/commands/sweep.md` and `ask.md` are thin shims that pass +`$ARGUMENTS` — the **entire raw argument string, untokenised** — to the +`sweep_plan` / `ask_plan` tools, whose arguments are structured JSON and arrive +intact. The MCP `sweep` prompt still exists for form-based clients (Claude +Desktop, Cursor), where each argument gets its own field; it routes through the +same parser and validators. + +**Settings.** Append `key=value` for any of `channels=`, `groups=`, +`batch_size=`, `parse_model=`, `report=`, `directive=`, `source=`, +`content_types=`, `regex=` (quote a multi-word value). The whitelist is closed: +an unrecognised `x=y` **stays in the question** and warns rather than being +eaten, with a typo hint if it is one edit away from a real key. Values are +validated — `batch_size` an integer 1–20, `parse_model` a single token, +`report` a single `.md` path with no `..` — and a bad one falls back to the +default with a `⚠` line at the top of the plan, never into the arithmetic. + +**What the plan does.** The tools pre-resolve what they can — canonicalising the +corpus handle, validating your channel/group tokens against the live corpus, +inlining the group roster when you gave no scope — so a three-call preamble +becomes none. Then: **pick the scope** if none was given (never a silent +whole-corpus sweep) → **`enumerate_matches` once** for the complete worklist → +**state N and `ceil(N / batch_size)` batches** → **per batch, map-reduce** → +finish with a summary that says how many of the N were actually read. An +unresolvable channel halts the plan instead of quietly widening it. The MCP +stays read-only; only the report file is written. **Dumb extractors on the cheap model.** Each batch is handed to a **subagent** (Claude Code's Task tool) spawned as a *verbatim extractor* on the cheapest -model — the prompt requests `parse_model` (default `haiku`) via the Task tool's +model — the plan requests `parse_model` (default `haiku`) via the Task tool's model override, and gracefully spawns on the default when no override exists. The extractor calls `get_transcripts` with `link_style:"base"` and returns *only* the directive-relevant lines, **verbatim** — grouped per video under its `moment_base:` line, in their compact `[mm:ss|sec]` form, under a hard budget of **≤40 lines (~600 words) per batch**. No analysis, no summarizing: the heavy -transcript bulk lives and dies inside the cheap subagent, and nothing irrelevant -is ever reprocessed. The **orchestrator does all the synthesis** — claims, -contradictions, cross-referencing — expanding each kept citation to a full link -by appending the seconds to the video's `moment_base`, then discards the -fragment. Batches are independent, so several can run in parallel. If no -subagent tool is available, the batch is processed inline (still -`link_style:"base"`) and the raw excerpt text dropped after folding. +transcript bulk lives and dies inside the cheap subagent. The **orchestrator +does all the synthesis** — claims, contradictions, cross-referencing — expanding +each kept citation to a full link by appending the seconds to the video's +`moment_base`, then discards the fragment. Every subagent is given the corpus +handle explicitly, since one that omitted it would read the server default and +quote the wrong archive. Batches are independent, so several can run in +parallel. **Linked citations.** Every source is cited as a clickable `[title @ mm:ss](<moment url>)` link — built by appending the cited integer seconds to the video's `moment_base` from the tool output, so a click seeks to the exact second (see *Compact base links* above; a video with no base is cited -by its `source:` URL). Note `open_link` results stay inline-linked. +by its `source:` URL). A post has no timeline, so it is cited +`[post by <author>, <date>](<source url>)` and never with `@ mm:ss`. + +The sweep discipline lives in `src/instructions.ts` — one builder shared by both +tools and the prompt — rather than duplicated into the markdown command files, +so it cannot drift out of sync with the tools it names. A test asserts every +tool the instructions mention actually exists. ## Data source (pick one) @@ -207,45 +238,74 @@ Resolved from flags or env — precedence hub > remote > local: | Remote | `--remote <url>` | `TRANSCRIPT_SITE_URL` | One deployed site origin. | | Local | `--local <dir>` | `TRANSCRIPT_LOCAL_DIR` | A composed public dir on disk (default `./export/public`). | -## Switching the source on the fly +## Choosing a corpus per call The server is added to your MCP client **once**, pointed at a startup source -(above). But you don't have to edit the config and restart to aim it elsewhere — -you can switch the **active corpus** from inside a session with three tools, and -your choice **persists across reconnects**. - -- **`list_sources`** — shows the active source (label + kind + target). When - there's a hub in play (you started against a `--hub`, or switched to one) it - also lists the hub's member sites as `siteId · title · url`, marking which are - in the current subset — so you can see what you can pick. -- **`use_source`** — switch the active source. Provide **exactly one** target: - - **`site: "<siteId | title>"`** — scope to a single hub member (matched by - siteId or display title, case-insensitive). This becomes a plain - single-site source, so it regains **full channel-group + alias** support. - - **`sites: ["<a>", "<b>"]`** — federate a **subset** of the hub's members. - Unknown tokens are reported, not silently dropped. - - **`remote: "<url>"`** — an arbitrary deployed site origin. - - **`local: "<dir>"`** — a composed public dir on disk. - - **`hub: "<url>"`** — an arbitrary hub URL to federate over. -- **`reset_source`** — return to the startup source and clear the persisted - selection. - -For example, connected to the archilyzer hub: `list_sources` to see the members, -then `use_source site:"jeralyzer"` to work in that one site (with its groups), or -`use_source sites:["jeralyzer","rekietalyzer"]` to federate just those two. - -**Persistence.** The active selection is written to a small state file so a -reconnect (or a fresh process with the same startup flags) resumes it. The file -lives under `$XDG_STATE_HOME/yt-dlp-transcript-mcp` (fallback -`~/.local/state/yt-dlp-transcript-mcp`), overridable with -`TRANSCRIPT_MCP_STATE_DIR`. It's keyed by the **startup** source, so two -differently-configured servers (archilyzer / rekietalyzer / a local build) keep -independent selections and don't clobber each other. - -**Still read-only.** `use_source` only changes *which* already-published static -shards are read — accepting an arbitrary `remote`/`local`/`hub` at call time is -the same capability the startup flags already grant this locally-run tool. -Nothing is ever written to any corpus. +(above) — its **default** corpus. To read something else, pass a `source` handle +on the call. There is no "active source": nothing is switched, and nothing is +remembered between calls. + +A handle **is** the serialised spec, in a canonical round-trippable form: + +| Handle | Means | +|---|---| +| `default` | the startup source — the same as omitting `source` | +| `local:/dir` | a composed public dir on disk | +| `remote:https://site` | one deployed site origin | +| `hub:https://hub` | federate every member of a hub | +| `hub:https://hub#alpha,beta` | a hub **subset** (`#`, so it can never collide with a query string) | + +Shorthands are accepted and normalised: a **bare site or hub URL** (classified by +probing its `corpus.json`), a hub member's **siteId or title**, `site:`/`sites:`, +and `./dir`. The canonical form is echoed back, so you learn it as you go. + +**Every result names the corpus it read**, as a trailing `(corpus: <handle>)` — +errors included. It is `corpus:` rather than `source:` because `- source:` already +means "this video's URL" in the output. + +``` +list_channels # the default corpus +list_channels source="remote:https://rekietalyzer.pages.dev" +search_transcripts query="k cups" source="hub:https://archilyzer.com#jeralyzer,rekietalyzer" +resolve_source source="jeralyzer" # → remote:https://jeralyzer.pages.dev +``` + +**Why it works this way.** The server used to hold a mutable active source and +persist it to a state file under `$XDG_STATE_HOME/yt-dlp-transcript-mcp`. That +was silently wrong: a server registered with `--local …/export/public` had +`{"activeSpec":{"kind":"remote","url":"https://hasanalyzer.pages.dev"}}` on disk +from some earlier session, so **every call read the wrong archive and no result +said so**. Protocol revision 2026-07-28 removes protocol-level sessions and +directs servers needing cross-call state to use "explicit, server-minted handles +passed as ordinary tool arguments" — which is exactly the fix. Any leftover +state files are now inert and can be deleted. + +`use_source` survives one release as an unadvertised alias that resolves its +target and tells you the handle to pass; `reset_source` is gone (there is +nothing to reset). + +**Still read-only.** A `source` handle only changes *which* already-published +static shards are read — the same capability the startup flags already grant +this locally-run tool. Nothing is ever written to any corpus. + +## Protocol + +The server speaks MCP revision **2026-07-28** through +`@modelcontextprotocol/server@2`'s `serveStdio`, which owns the era decision: +the opening exchange selects `modern` (2026-07-28, negotiated via +`server/discover`) or `legacy` (the 2025 `initialize` handshake), and one server +instance is pinned for the connection. Both eras serve the identical tools — +`src/protocol.test.ts` spawns the real process and asserts it on each. + +`resultType` is stamped and stripped by the SDK. Cache hints (`ttlMs` / +`cacheScope`) are a construction-time policy: `tools/list`, `prompts/list` and +`server/discover` are literal constants carrying no corpus data, so they are +`public` for an hour, and every read tool's result stays uncached. An invalid +hint throws at startup rather than on the wire. + +The low-level `Server` is used deliberately (it is marked "advanced use" rather +than removed): literal JSON Schema tool definitions keep the zod dependency out +of the package entirely. ## Run it @@ -273,6 +333,13 @@ claude mcp add rekietalyzer \ -- pnpm -C /ABS/PATH/TO/yt-dlp-transcript-browser --filter yt-dlp-transcript-mcp exec tsx src/index.ts ``` +**The `/sweep` and `/ask` commands.** `.claude/commands/{sweep,ask}.md` in this +repo call `mcp__archilyzer__sweep_plan` / `mcp__archilyzer__ask_plan` — the tool +name embeds the MCP server name **as you registered it**, so if you used another +name (`rekietalyzer` above), change the `mcp__<name>__` prefix in those two +files to match. `foo:bar` namespacing is plugin-only, so what you type stays +`/sweep`, not `/archilyzer:sweep`. + ## Add to any MCP client (mcp.json) ```json diff --git a/mcp/package.json b/mcp/package.json @@ -13,9 +13,10 @@ "typecheck": "tsc --noEmit -p tsconfig.json" }, "dependencies": { - "@modelcontextprotocol/sdk": "^1.12.0" + "@modelcontextprotocol/server": "^2.0.0" }, "devDependencies": { + "@modelcontextprotocol/client": "^2.0.0", "@types/node": "^20.19.39", "tsx": "^4.21.0", "typescript": "^5.6.0" diff --git a/mcp/src/index.ts b/mcp/src/index.ts @@ -1,20 +1,30 @@ #!/usr/bin/env tsx -import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js"; +import { serveStdio } from "@modelcontextprotocol/server/stdio"; import { resolveSourceSpec } from "./sources"; import { createServer } from "./server"; -import { SourceController } from "./sourceController"; +import { SourceRegistry } from "./sourceRegistry"; -async function main(): Promise<void> { - const spec = resolveSourceSpec(process.argv.slice(2)); - const controller = new SourceController(spec); - await controller.init(); - console.error(`[yt-dlp-transcript-mcp] source: ${controller.current.label}`); - const server = createServer(controller); - await server.connect(new StdioServerTransport()); +// One registry for the process. It holds no active source and writes nothing: +// each call resolves its own `source` handle, and a call that omits one reads +// the startup spec below. The instance cache lives here so per-call selection +// stays cheap. +// +// serveStdio owns the era decision for the connection: the opening exchange +// selects `modern` (2026-07-28) or `legacy` (the 2025 initialize handshake), +// then pins ONE instance from this factory for the connection's lifetime. It +// also builds — and discards — one extra instance for a `server/discover` +// probe, so expect up to two `era:` lines per client. That log is how we learn +// which era a given client actually negotiates. +function main(): void { + const registry = new SourceRegistry(resolveSourceSpec(process.argv.slice(2))); + console.error( + `[yt-dlp-transcript-mcp] default corpus: ${registry.defaultHandle}`, + ); + serveStdio(({ era }) => { + console.error(`[yt-dlp-transcript-mcp] serving protocol era: ${era}`); + return createServer(registry); + }); console.error("[yt-dlp-transcript-mcp] ready on stdio"); } -main().catch((err) => { - console.error(err); - process.exit(1); -}); +main(); diff --git a/mcp/src/instructions.test.ts b/mcp/src/instructions.test.ts @@ -0,0 +1,175 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { buildSweepInstructions, buildAskInstructions } from "./instructions"; +import { parsePromptRequest } from "./promptRequest"; +import { TOOLS } from "./server"; + +const CORPUS = "remote:https://site.example"; + +function sweep(request: string, ctxExtra = {}): string { + return buildSweepInstructions(parsePromptRequest(request), { + corpus: CORPUS, + ...ctxExtra, + }); +} +function ask(request: string, ctxExtra = {}): string { + return buildAskInstructions(parsePromptRequest(request), { + corpus: CORPUS, + ...ctxExtra, + }); +} + +// ─── The cheap guard against a whole class of drift ─── + +test("every tool the instructions name actually exists", () => { + const known = new Set(TOOLS.map((t) => t.name)); + // Every `backticked_snake_case` token in any generated instruction that + // looks like a tool name must be a real tool. Renaming a tool without + // updating the instructions is otherwise invisible until a sweep fails. + const texts = [ + sweep("coffee"), + sweep("coffee channels=chan-a batch_size=4"), + sweep("https://site.example/?q=a follow this"), + sweep("coffee", { unknownChannels: ["ghost"] }), + sweep("coffee", { + availableGroups: [{ id: "other", name: "Other", channels: 3 }], + }), + ask("what did they say"), + ask("https://site.example/?q=a what about this"), + ]; + const suspects = new Set<string>(); + for (const text of texts) { + for (const m of text.matchAll(/`([a-z][a-z0-9]*(?:_[a-z0-9]+)+)`/g)) { + suspects.add(m[1]); + } + } + // Argument names share the shape, so only judge tokens that name a tool. + const NOT_TOOLS = new Set([ + "batch_size", + "content_types", + "link_style", + "max_pages", + "moment_base", + "dry_run", + "parse_model", + "report_path", + ]); + const named = [...suspects].filter((s) => !NOT_TOOLS.has(s)); + assert.ok(named.length > 0, "the instructions should name some tools"); + for (const name of named) { + assert.ok(known.has(name), `instructions name a nonexistent tool: ${name}`); + } +}); + +// ─── The corpus handle is threaded everywhere ─── + +test("the sweep instructions pin the corpus on every call, subagents included", () => { + const text = sweep("coffee channels=chan-a"); + assert.match(text, /\*\*Corpus: `remote:https:\/\/site\.example`\.\*\*/); + assert.match(text, /Pass `source: "remote:https:\/\/site\.example"` on every/); + assert.match(text, /Pass the corpus handle into every subagent/); + // The literal scope args a call should carry. + assert.match(text, /source: "remote:https:\/\/site\.example", channels: \["chan-a"\]/); +}); + +test("the ask instructions pin the corpus too", () => { + const text = ask("what did they say about coffee"); + assert.match(text, /Corpus: `remote:https:\/\/site\.example`/); + assert.match(text, /what did they say about coffee/); +}); + +// ─── Coverage discipline ─── + +test("the sweep enumerates in one call and never instructs paging", () => { + const text = sweep("coffee channels=chan-a"); + assert.match(text, /enumerate_matches/); + assert.match(text, /COVERAGE PARTIAL/); + assert.ok( + !/offset \+= limit/.test(text), + "the paging instruction is gone — it was ignored and is the expensive path", + ); + assert.match(text, /ceil\(N \/ 8\)/); +}); + +test("a validated batch size reaches the arithmetic, and a bad one cannot", () => { + assert.match(sweep("coffee channels=c batch_size=12"), /ceil\(N \/ 12\)/); + const bad = sweep("coffee channels=c batch_size=Why"); + assert.match(bad, /ceil\(N \/ 8\)/); + assert.ok(!bad.includes("ceil(N / Why)"), "the 'Why' bug must be impossible"); + assert.match(bad, /^⚠ batch_size "Why"/m); +}); + +test("warnings render as a leading block", () => { + const text = sweep("coffee flavour=vanilla"); + assert.ok(text.startsWith("⚠ "), text.slice(0, 40)); + assert.match(text, /not a recognised setting/); +}); + +// ─── Scope handling ─── + +test("an unresolvable channel stops the sweep instead of widening it", () => { + const text = sweep("coffee channels=ghost", { unknownChannels: ["ghost"] }); + assert.match(text, /Fix the scope first/); + assert.match(text, /channel "ghost"/); + // And it does not go on to instruct an enumeration over everything. + assert.ok(!/Enumerate the complete worklist/.test(text)); +}); + +test("no scope means pick-first, with the roster inlined when known", () => { + const withRoster = sweep("coffee", { + availableGroups: [ + { id: "other", name: "Extended Universe", channels: 7 }, + { id: "core", name: "Core", channels: 2 }, + ], + }); + assert.match(withRoster, /Choose the scope first/); + assert.match(withRoster, /other · Extended Universe · 7 channel\(s\)/); + // The roster was pre-resolved, so no round-trip is asked for. + assert.ok(!/Call `list_channels`/.test(withRoster)); + + const without = sweep("coffee"); + assert.match(without, /Choose the scope first/); + assert.match(without, /Call `list_channels`/); + assert.match(without, /confirm \*\*all\*\*/); +}); + +test("an explicit scope skips the pick-first step", () => { + const text = sweep("coffee groups=other", { knownGroups: ["other"] }); + assert.ok(!/Choose the scope first/.test(text)); + assert.match(text, /scoped to groups "other"/); +}); + +// ─── Links ─── + +test("a link seeds a dry-run decode, then a single real call", () => { + const text = sweep("https://site.example/?qt=abc&fav=deleted what happened"); + assert.match(text, /open_link/); + assert.match(text, /dry_run: true/); + assert.match(text, /https:\/\/site\.example\/\?qt=abc&fav=deleted/); + assert.ok(!text.includes("apply:true")); +}); + +// ─── Citations ─── + +test("both builders mandate the same citation forms", () => { + for (const text of [sweep("coffee channels=c"), ask("coffee")]) { + assert.match(text, /\[title @ mm:ss\]/); + assert.match(text, /\[post by <author>, <date>\]/); + assert.match(text, /never (take|takes|with) a? ?`@ mm:ss`/); + } +}); + +test("the ask plan pushes multi-query reads rather than re-reading per term", () => { + const text = ask("what did they say about coffee and tea"); + assert.match(text, /get_transcripts/); + assert.match(text, /`queries`/); + assert.match(text, /enumerate_matches/); +}); + +test("resolved context notes are surfaced", () => { + const text = sweep("coffee channels=c", { + notes: ["42 channel(s) in remote:https://site.example"], + }); + assert.match(text, /Context already resolved for you/); + assert.match(text, /42 channel\(s\)/); +}); diff --git a/mcp/src/instructions.ts b/mcp/src/instructions.ts @@ -0,0 +1,304 @@ +import { + DEFAULT_DIRECTIVE, + DEFAULT_REPORT_PATH, + renderWarnings, + type PromptRequest, +} from "./promptRequest"; + +// ─── The orchestration text a sweep / ask runs on ─── +// +// The server stays the single source of truth for the sweep discipline — +// citation format, extractor budget, base-link expansion, the coverage rule — +// rather than duplicating it into a markdown command file that would drift. +// All three entry points (the `sweep_plan` and `ask_plan` tools and the MCP +// `sweep` prompt) render through here. + +// Facts the server resolved before writing the instructions, so the model does +// not spend round-trips rediscovering them. Anything unresolvable is simply +// absent and the instructions fall back to asking. +export type PlanContext = { + // The canonical corpus handle every call in this plan must pass. + corpus: string; + // Channel/group tokens validated against the live corpus. + knownChannels?: string[]; + unknownChannels?: string[]; + knownGroups?: string[]; + unknownGroups?: string[]; + // The groups on offer, when the user picked no scope and has to choose. + availableGroups?: { id: string; name: string; channels: number }[]; + // Extra notes to surface (e.g. a decoded link's plan). + notes?: string[]; +}; + +function quoted(xs: string[]): string { + return xs.map((x) => `"${x}"`).join(", "); +} + +// The scope arguments every search/enumerate call in the plan should carry, +// rendered literally so they can be copied. +function scopeArgs(req: PromptRequest, ctx: PlanContext): string { + const parts = [`source: "${ctx.corpus}"`]; + if (req.channels.length > 0) parts.push(`channels: [${quoted(req.channels)}]`); + if (req.groups.length > 0) parts.push(`groups: [${quoted(req.groups)}]`); + if (req.contentTypes) parts.push(`content_types: [${quoted(req.contentTypes)}]`); + if (req.regex) parts.push(`regex: true`); + return parts.join(", "); +} + +function scopeProse(req: PromptRequest): string { + const clauses: string[] = []; + if (req.channels.length > 0) clauses.push(`channels ${quoted(req.channels)}`); + if (req.groups.length > 0) clauses.push(`groups ${quoted(req.groups)}`); + return clauses.join(" and "); +} + +// The scope-validation preamble: what the server already checked, and what the +// model must do about anything that didn't resolve. +function scopeSteps( + req: PromptRequest, + ctx: PlanContext, +): { steps: string[]; halt: boolean } { + const steps: string[] = []; + const unknown = [ + ...(ctx.unknownChannels ?? []).map((c) => `channel "${c}"`), + ...(ctx.unknownGroups ?? []).map((g) => `group "${g}"`), + ]; + if (unknown.length > 0) { + steps.push( + `**Fix the scope first.** These did not match anything in this corpus: ` + + `${unknown.join(", ")}. Do NOT proceed with a silently-wider scope — ` + + `show me \`list_channels\` output and ask which I meant.`, + ); + // Halt here: an instruction list that goes on to enumerate would invite + // running the sweep over a silently wider scope than was asked for. + return { steps, halt: true }; + } + const hasScope = req.channels.length > 0 || req.groups.length > 0; + if (!hasScope && !req.link) { + // When the server could pre-resolve the roster, inline it — that turns a + // round-trip into a question the model can ask immediately. When it + // couldn't, say how to fetch it. + const roster = + ctx.availableGroups && ctx.availableGroups.length > 0 + ? `\n\n The groups in this corpus are:\n` + + ctx.availableGroups + .map((g) => ` - ${g.id} · ${g.name} · ${g.channels} channel(s)`) + .join("\n") + : ` Call \`list_channels\` (with \`source: "${ctx.corpus}"\`) and ` + + `present the groups and their channels.`; + steps.push( + `**Choose the scope first — do NOT default to the whole corpus.** No ` + + `channels or groups were given.${roster.startsWith(" Call") ? roster : ""} ` + + `Ask which group(s) or channel(s) to sweep, or to confirm **all** for ` + + `the whole corpus. Wait for my choice, then use it as the ` + + `\`channels\`/\`groups\` scope in every call below.` + + (roster.startsWith(" Call") ? "" : roster), + ); + } + return { steps, halt: false }; +} + +// The per-batch map-reduce: a cheap extractor subagent quotes verbatim, the +// orchestrator does every bit of the synthesis. The transcript bulk never +// enters the orchestrator's context and never costs the big model. +function batchStep(req: PromptRequest, ctx: PlanContext, reportPath: string): string { + return ( + `**Per batch (map-reduce), for each group of up to ${req.batchSize} ids:**\n` + + ` - **Spawn a subagent as a DUMB EXTRACTOR on the cheapest model** — ` + + `use the Task tool and request model "${req.parseModel}" (its model ` + + `parameter, or a "${req.parseModel}"-backed agent type); if no model ` + + `override is available, spawn it anyway on the default. Give it exactly ` + + `this job: call \`get_transcripts\` with the batch's ids, ` + + `\`source: "${ctx.corpus}"\`, the sweep's \`queries\`, and ` + + `\`link_style: "base"\`, then return ONLY the lines relevant to the ` + + `directive, VERBATIM — do NOT analyze, summarize, or rephrase anything. ` + + `Group the kept lines per video as \`### <title>\` + that video's ` + + `\`moment_base:\` line copied exactly (or its \`source:\` line when there ` + + `is no moment_base) + the kept lines in their \`[mm:ss|seconds]\` form. ` + + `Hard budget: at most 40 lines (~600 words) per batch — if more match, ` + + `keep the strongest and end with \`(+N more matching lines)\`. Return ` + + `nothing else.\n` + + ` - **Pass the corpus handle into every subagent.** Each one must call ` + + `with \`source: "${ctx.corpus}"\`. A subagent that omits it reads this ` + + `server's default corpus instead, and its quotes would be from the wrong ` + + `archive with nothing in the output to show it.\n` + + ` - **Merge — you (the orchestrator) do ALL the synthesis.** Cross-` + + `reference the returned fragment against the report so far and upsert ` + + `findings — claims, and contradictions with earlier claims — into ` + + `well-titled \`## sections\` of \`${reportPath}\` (Write/Edit). Cite every ` + + `VIDEO finding as **\`[title @ mm:ss](<moment url>)\`**, expanding each ` + + `kept \`[mm:ss|seconds]\` stamp by appending the integer after the \`|\` ` + + `to that video's moment_base (full link = \`<moment_base><seconds>\`; no ` + + `moment_base → link the \`source:\` URL instead). A POST finding has no ` + + `timestamp — cite it as **\`[post by <author>, <date>](<source url>)\`**, ` + + `never with \`@ mm:ss\`. Then discard the fragment. Batches are ` + + `independent, so you may dispatch several subagents in parallel.\n` + + ` - **Fallback:** with no Task tool, do the batch inline — call ` + + `\`get_transcripts\` with \`link_style: "base"\` yourself, fold the ` + + `expanded cited findings into the report, then **drop the raw excerpt ` + + `text** before moving on.` + ); +} + +// The full sweep instructions. +export function buildSweepInstructions( + req: PromptRequest, + ctx: PlanContext, +): string { + const reportPath = req.reportPath ?? DEFAULT_REPORT_PATH; + const directive = req.directive ?? DEFAULT_DIRECTIVE; + const subject = req.query + ? `"${req.query}"` + : req.link + ? "the share link's search" + : "the subject below"; + const args = scopeArgs(req, ctx); + + const scope = scopeSteps(req, ctx); + const steps: string[] = [...scope.steps]; + + if (!scope.halt && req.link) { + steps.push( + `**Decode the link.** Call \`open_link\` with link="${req.link}" and ` + + `\`dry_run: true\`. It returns the plan: the resolved corpus handle, ` + + `the query tree, every active filter, the channel scope validated ` + + `against that corpus, and anything ignored. **Show me the plan and ` + + `confirm it captures what I want.** If I ask for a change ("drop the ` + + `availability filter", "only channel X", "search Y instead"), re-call ` + + `with the matching \`overrides\` until it is right. Then call once ` + + `more without \`dry_run\` to run it. Use the handle it reports as ` + + `\`source\` from then on.`, + ); + } + + if (!scope.halt) { + steps.push( + `**Enumerate the complete worklist — one call.** Call ` + + `\`enumerate_matches\` with ${args}` + + (req.query ? `, query: "${req.query}"` : "") + + `. It returns EVERY match (id/title/channel/date) in a single scan plus ` + + `the batch count. Do NOT page \`search_transcripts\` for this: that is ` + + `one full corpus scan per page, and it is how a previous sweep reported ` + + `319 videos after seeing 200. If the first line says **COVERAGE ` + + `PARTIAL**, the list is a sample — narrow the scope or raise ` + + `\`max_pages\`, and if you proceed anyway, say so prominently in the ` + + `report.`, + ); + + steps.push( + `**State the plan.** Report N (the enumerated total) and ` + + `\`ceil(N / ${req.batchSize})\` batches before you start. The report may ` + + `only ever claim the coverage this number justifies: N videos ` + + `enumerated, and however many you actually read.`, + ); + + steps.push(batchStep(req, ctx, reportPath)); + + steps.push( + `**Finish.** Work to the end of the worklist, then write a summary ` + + `section: the scope swept, **how many of the N you actually read**, ` + + `headline findings, and any partial-coverage caveat. Tell me the report ` + + `path. Keep every citation clickable — \`[title @ mm:ss](url)\` for a ` + + `video, \`[post by <author>, <date>](url)\` for a post.`, + ); + } + + const numbered = steps.map((s, i) => `${i + 1}. ${s}`).join("\n\n"); + const scopeNote = scopeProse(req); + + const head = + `Run a **corpus sweep** for ${subject}` + + (scopeNote ? `, scoped to ${scopeNote}` : "") + + `, extracting **${directive}**, and maintain a running markdown report at ` + + `\`${reportPath}\`.\n\n` + + `You are the sweep engine — work the whole match set methodically, using ` + + `the transcript MCP tools for evidence and your own Write/Edit tools for ` + + `the report. The MCP is read-only; never try to change the archive. The ` + + `corpus holds video transcripts AND archived social posts.\n\n` + + `**Corpus: \`${ctx.corpus}\`.** Pass \`source: "${ctx.corpus}"\` on every ` + + `single call, including the ones your subagents make. This server has no ` + + `active source — a call that omits \`source\` reads the server default, ` + + `which may be a different archive.\n\n` + + `**Cite every video finding as a clickable \`[title @ mm:ss](<moment ` + + `url>)\` link** (append the cited integer seconds to that video's ` + + `\`moment_base\` from the tool output); **cite every post finding as ` + + `\`[post by <author>, <date>](<source url>)\`** — posts have no timeline, ` + + `so they never take a \`@ mm:ss\`.`; + + const notes = + ctx.notes && ctx.notes.length > 0 + ? `\n\nContext already resolved for you:\n${ctx.notes.map((n) => `- ${n}`).join("\n")}` + : ""; + + return ( + renderWarnings(req.warnings) + + `${head}${notes}\n\nFollow these steps:\n\n${numbered}` + ); +} + +// The ask variant: the same evidence discipline, but answering a question in +// the conversation rather than maintaining a report file. Kept deliberately +// close to the sweep so citations look identical either way. +export function buildAskInstructions( + req: PromptRequest, + ctx: PlanContext, +): string { + const question = req.query || "the question below"; + const args = scopeArgs(req, ctx); + const askScope = scopeSteps(req, ctx); + const steps: string[] = [...askScope.steps]; + + if (!askScope.halt && req.link) { + steps.push( + `**Decode the link.** Call \`open_link\` with link="${req.link}" — one ` + + `call returns the plan and the first page of results, plus the corpus ` + + `handle to use from then on.`, + ); + } + + if (!askScope.halt) { + steps.push( + `**Find the evidence.** Search with \`search_transcripts\` (${args}) for ` + + `the terms the question implies — try more than one phrasing; the ` + + `corpus is ASR text and the curated aliases only cover known ` + + `mis-transcriptions. If you need to know HOW MANY, or to cover ` + + `everything, use \`enumerate_matches\` instead: a search page is a ` + + `slice, and its \`⚠ INCOMPLETE PAGE\` banner means exactly that.`, + ); + + steps.push( + `**Read before concluding.** Pull the surrounding context with ` + + `\`get_transcripts\` (${args}) — pass all the terms at once in ` + + `\`queries\` rather than re-reading the same videos per term. A snippet ` + + `is not evidence of what was meant; the window around it is.`, + ); + + steps.push( + `**Answer with citations.** Answer ${question} directly, and cite every ` + + `claim: \`[title @ mm:ss](<moment url>)\` for a video (append the ` + + `integer seconds to that video's \`moment_base\`), ` + + `\`[post by <author>, <date>](<source url>)\` for a post — a post has ` + + `no timeline, so it never takes a \`@ mm:ss\`. Say plainly what the ` + + `corpus does NOT show — an absence of matches is a finding, not a gap ` + + `to paper over.`, + ); + } + + const numbered = steps.map((s, i) => `${i + 1}. ${s}`).join("\n\n"); + const head = + `Answer this question from the transcript archive: **${question}**\n\n` + + `**Corpus: \`${ctx.corpus}\`.** Pass \`source: "${ctx.corpus}"\` on every ` + + `call — this server has no active source, and a call that omits it reads ` + + `the server default. The MCP is read-only. The corpus holds video ` + + `transcripts AND archived social posts.`; + + const notes = + ctx.notes && ctx.notes.length > 0 + ? `\n\nContext already resolved for you:\n${ctx.notes.map((n) => `- ${n}`).join("\n")}` + : ""; + + return ( + renderWarnings(req.warnings) + + `${head}${notes}\n\nFollow these steps:\n\n${numbered}` + ); +} diff --git a/mcp/src/promptRequest.test.ts b/mcp/src/promptRequest.test.ts @@ -0,0 +1,265 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + parsePromptRequest, + requestFromArguments, + renderWarnings, + DEFAULT_BATCH_SIZE, + DEFAULT_PARSE_MODEL, +} from "./promptRequest"; + +// The parser exists because Claude Code's slash-command tokenizer shredded a +// real request into link="This" channel="search" group="deleted" +// directive="videos." batch_size="Why" and dropped the rest. Every case below +// is either that failure or a neighbour of it. + +// ─── Link extraction ─── + +test("a bare URL is extracted and the rest stays prose", () => { + const r = parsePromptRequest("https://site.example/?q=x what did they say"); + assert.equal(r.link, "https://site.example/?q=x"); + assert.equal(r.query, "what did they say"); +}); + +test("a URL's query string survives intact — & and = are not settings", () => { + const r = parsePromptRequest( + "https://site.example/?qt=abc&fc=chan-a&fav=deleted why though", + ); + assert.equal(r.link, "https://site.example/?qt=abc&fc=chan-a&fav=deleted"); + assert.equal(r.query, "why though"); + assert.deepEqual(r.warnings, []); +}); + +test("only the FIRST url becomes the link; a second stays in the prose", () => { + const r = parsePromptRequest( + "https://a.example/one compare with https://b.example/two", + ); + assert.equal(r.link, "https://a.example/one"); + assert.match(r.query, /https:\/\/b\.example\/two/); +}); + +test("an angle-bracketed url is unwrapped", () => { + const r = parsePromptRequest("<https://site.example/x> and then some"); + assert.equal(r.link, "https://site.example/x"); + assert.ok(!r.query.includes("<"), r.query); + assert.ok(!r.query.includes(">"), r.query); +}); + +test("a markdown link is unwrapped, label and all", () => { + const r = parsePromptRequest("[Search results](https://site.example/?q=a) go on"); + assert.equal(r.link, "https://site.example/?q=a"); + assert.equal(r.query, "go on"); +}); + +test("a quoted url is unwrapped", () => { + const r = parsePromptRequest('"https://site.example/x" please'); + assert.equal(r.link, "https://site.example/x"); + assert.equal(r.query, "please"); +}); + +test("a trailing full stop is not part of the url", () => { + const r = parsePromptRequest("see https://site.example/x. Then explain."); + assert.equal(r.link, "https://site.example/x"); +}); + +test("a trailing comma, colon and semicolon are not part of the url", () => { + for (const p of [",", ":", ";", "!", "?"]) { + const r = parsePromptRequest(`see https://site.example/x${p} more`); + assert.equal(r.link, "https://site.example/x", `punctuation ${p}`); + } +}); + +test("an unbalanced closing paren is dropped but a balanced one is kept", () => { + const a = parsePromptRequest("(see https://site.example/x) ok"); + assert.equal(a.link, "https://site.example/x"); + const b = parsePromptRequest("https://site.example/x(y) ok"); + assert.equal(b.link, "https://site.example/x(y)"); +}); + +test("no url at all is fine", () => { + const r = parsePromptRequest("just a question about coffee"); + assert.equal(r.link, undefined); + assert.equal(r.query, "just a question about coffee"); +}); + +// ─── The whole-request regression ─── + +test("the shredded Rekieta line survives whole", () => { + const input = + "https://hasanalyzer.pages.dev/?qt=eyJhIjoxfQ&fav=deleted " + + "This search finds deleted videos. Why might have Rekieta privated " + + "these? Look for context around each one. batch_size=12"; + const r = parsePromptRequest(input); + assert.equal(r.link, "https://hasanalyzer.pages.dev/?qt=eyJhIjoxfQ&fav=deleted"); + assert.equal(r.batchSize, 12); + // Every word of the question is present, in order, punctuation intact. + assert.equal( + r.query, + "This search finds deleted videos. Why might have Rekieta privated " + + "these? Look for context around each one.", + ); + assert.deepEqual(r.warnings, []); +}); + +// ─── The key=value whitelist ─── + +test("recognised settings are parsed off the prose", () => { + const r = parsePromptRequest( + "coffee talk channels=chan-a,chan-b groups=other batch_size=5 " + + "parse_model=sonnet report=./out.md source=remote:https://x.example", + ); + assert.equal(r.query, "coffee talk"); + assert.deepEqual(r.channels, ["chan-a", "chan-b"]); + assert.deepEqual(r.groups, ["other"]); + assert.equal(r.batchSize, 5); + assert.equal(r.parseModel, "sonnet"); + assert.equal(r.reportPath, "./out.md"); + assert.equal(r.source, "remote:https://x.example"); +}); + +test("aliases map onto the canonical keys", () => { + const r = parsePromptRequest("x channel=chan-a group=other model=opus out=./r.md"); + assert.deepEqual(r.channels, ["chan-a"]); + assert.deepEqual(r.groups, ["other"]); + assert.equal(r.parseModel, "opus"); + assert.equal(r.reportPath, "./r.md"); +}); + +test("an unknown key=value STAYS in the prose and warns", () => { + const r = parsePromptRequest("find x flavour=vanilla please"); + assert.match(r.query, /flavour=vanilla/); + assert.match(r.warnings.join(" "), /"flavour=" is not a recognised setting/); +}); + +test("a near-miss key gets a typo hint and still stays in the prose", () => { + const r = parsePromptRequest("find x chanels=chan-a"); + assert.match(r.query, /chanels=chan-a/); + assert.match(r.warnings.join(" "), /did you mean channels=\?/); +}); + +test("prose that merely contains '=' is never treated as a setting", () => { + const r = parsePromptRequest("solve x=y+2 and 3==3 for me"); + assert.match(r.query, /x=y\+2/); + assert.match(r.query, /3==3/); +}); + +test("a quoted multi-word value survives as one value", () => { + const r = parsePromptRequest( + 'sweep it directive="key claims & contradictions" now', + ); + assert.equal(r.directive, "key claims & contradictions"); + assert.equal(r.query, "sweep it now"); +}); + +test("a repeated key takes the last value and warns", () => { + const r = parsePromptRequest("x batch_size=4 batch_size=9"); + assert.equal(r.batchSize, 9); + assert.match(r.warnings.join(" "), /batch_size was given more than once/); +}); + +// ─── Validation: the failures that used to render into the output ─── + +test("batch_size=Why falls back to the default with a warning", () => { + const r = parsePromptRequest("x batch_size=Why"); + assert.equal(r.batchSize, DEFAULT_BATCH_SIZE); + assert.match(r.warnings.join(" "), /batch_size "Why" is not an integer/); +}); + +test("batch_size out of range or fractional is rejected", () => { + for (const bad of ["0", "21", "2.5", "-3"]) { + const r = parsePromptRequest(`x batch_size=${bad}`); + assert.equal(r.batchSize, DEFAULT_BATCH_SIZE, `batch_size=${bad}`); + assert.ok(r.warnings.length > 0); + } +}); + +test("batch_size at the bounds is accepted", () => { + assert.equal(parsePromptRequest("x batch_size=1").batchSize, 1); + assert.equal(parsePromptRequest("x batch_size=20").batchSize, 20); +}); + +test("a multi-word parse_model is rejected", () => { + const r = parsePromptRequest('x parse_model="two words"'); + assert.equal(r.parseModel, DEFAULT_PARSE_MODEL); + assert.match(r.warnings.join(" "), /not a single token/); +}); + +test("a report path escaping the cwd is rejected", () => { + const r = parsePromptRequest("x report=../secrets.md"); + assert.equal(r.reportPath, undefined); + assert.match(r.warnings.join(" "), /escapes the working directory/); +}); + +test("a report path that is not .md is rejected", () => { + const r = parsePromptRequest("x report=./out.txt"); + assert.equal(r.reportPath, undefined); + assert.match(r.warnings.join(" "), /not a .md file/); +}); + +test("a multi-word report path is rejected", () => { + const r = parsePromptRequest('x report="my report.md"'); + assert.equal(r.reportPath, undefined); + assert.match(r.warnings.join(" "), /not a single token/); +}); + +test("content_types keeps the valid values and warns about the rest", () => { + const r = parsePromptRequest("x content_types=post,audio"); + assert.deepEqual(r.contentTypes, ["post"]); + assert.match(r.warnings.join(" "), /content_types value\(s\) ignored/); +}); + +test("regex=true is honoured; anything else is false", () => { + assert.equal(parsePromptRequest("x regex=true").regex, true); + assert.equal(parsePromptRequest("x regex=yes").regex, false); + assert.equal(parsePromptRequest("x").regex, false); +}); + +test("empty input warns rather than producing a silent no-op plan", () => { + const r = parsePromptRequest(""); + assert.equal(r.query, ""); + assert.equal(r.link, undefined); + assert.match(r.warnings.join(" "), /no query and no link/); +}); + +test("a link with no question is not a warning — the link supplies the query", () => { + const r = parsePromptRequest("https://site.example/?q=a"); + assert.deepEqual(r.warnings, []); +}); + +// ─── The MCP prompt form shares the validators ─── + +test("prompt arguments route through the same validation", () => { + const r = requestFromArguments({ + query: "k cups", + channels: "chan-a, chan-b", + group: "other", + batch_size: "Why", + report_path: "../x.md", + }); + assert.equal(r.query, "k cups"); + assert.deepEqual(r.channels, ["chan-a", "chan-b"]); + assert.deepEqual(r.groups, ["other"]); + assert.equal(r.batchSize, DEFAULT_BATCH_SIZE); + assert.equal(r.reportPath, undefined); + assert.equal(r.warnings.length, 2); +}); + +test("prompt arguments accept a link and clean it", () => { + const r = requestFromArguments({ link: "<https://site.example/x>." }); + assert.equal(r.link, "https://site.example/x"); +}); + +test("a non-url link argument is reported, not passed through", () => { + const r = requestFromArguments({ link: "not-a-url", query: "x" }); + assert.equal(r.link, undefined); + assert.match(r.warnings.join(" "), /is not an http\(s\) URL/); +}); + +// ─── Rendering ─── + +test("warnings render as a leading block, and nothing when clean", () => { + assert.equal(renderWarnings([]), ""); + const out = renderWarnings(["a", "b"]); + assert.ok(out.startsWith("⚠ a\n⚠ b")); + assert.ok(out.endsWith("\n\n")); +}); diff --git a/mcp/src/promptRequest.ts b/mcp/src/promptRequest.ts @@ -0,0 +1,413 @@ +// ─── Parsing a pasted request into a sweep/ask plan ─── +// +// Claude Code tokenises an MCP prompt's arguments as +// `text.trim().split(/\s+/)` zipped against the declared argument names. It is +// not quote-aware, the last argument does not absorb the remainder, and tokens +// past the declared count are dropped in silence. With nine declared arguments +// a real request became link="This" channel="search" group="deleted" +// directive="videos." batch_size="Why" — and everything from the actual +// question onward vanished. +// +// So the entry point for Claude Code is a TOOL (structured JSON arguments, +// which arrive intact) taking one free-text `request`, and this module is the +// parser behind it. It is pure and exhaustively tested; the MCP prompt form +// routes through the same validators so a form-based client (Claude Desktop, +// Cursor) gets a named error instead of mangled arithmetic. +// +// The rules that matter: +// - a key=value pair is honoured ONLY for a closed whitelist, so an unknown +// `x=y` stays in the prose instead of being silently eaten; +// - a near-miss key gets a typo hint and still stays in the prose; +// - every ignored, corrected, or defaulted input lands in `warnings`, which +// the caller renders as a ⚠ block at the top of its output. + +export type PromptRequest = { + // The first http(s) URL in the input, cleaned of wrappers and trailing + // punctuation. A share link seeds the search; anything else is just prose. + link?: string; + // Everything that was not a recognised key=value pair or the link: the + // question, in the user's own words, punctuation intact. + query: string; + channels: string[]; + groups: string[]; + batchSize: number; + parseModel: string; + reportPath?: string; + directive?: string; + source?: string; + contentTypes?: ("video" | "post")[]; + regex: boolean; + warnings: string[]; +}; + +export const DEFAULT_BATCH_SIZE = 8; +export const DEFAULT_PARSE_MODEL = "haiku"; +export const DEFAULT_DIRECTIVE = "key claims & contradictions"; +export const DEFAULT_REPORT_PATH = "./sweep-report.md"; + +// The closed whitelist: canonical key → accepted aliases. Anything outside it +// is prose, never a parameter. +const KEYS: Record<string, string[]> = { + channels: ["channel", "chan", "chans"], + groups: ["group"], + batch_size: ["batch", "batchsize", "batch_sz"], + parse_model: ["model", "parsemodel", "extractor_model"], + report: ["report_path", "reportpath", "path", "out", "output"], + directive: ["extract", "extracting"], + source: ["corpus"], + content_types: ["content_type", "types", "type"], + regex: ["re"], +}; + +const CANONICAL = new Map<string, string>(); +for (const [canon, aliases] of Object.entries(KEYS)) { + CANONICAL.set(canon, canon); + for (const a of aliases) CANONICAL.set(a, canon); +} + +// Edit distance, bounded: we only care whether it is ≤ 1. +function withinOneEdit(a: string, b: string): boolean { + if (a === b) return true; + const [short, long] = a.length <= b.length ? [a, b] : [b, a]; + if (long.length - short.length > 1) return false; + let i = 0; + let j = 0; + let edits = 0; + while (i < short.length && j < long.length) { + if (short[i] === long[j]) { + i++; + j++; + continue; + } + if (++edits > 1) return false; + if (short.length === long.length) i++; + j++; + } + return edits + (long.length - j) + (short.length - i) <= 1; +} + +function typoHint(key: string): string | undefined { + for (const known of CANONICAL.keys()) { + if (withinOneEdit(key.toLowerCase(), known)) return known; + } + return undefined; +} + +function countChar(s: string, c: string): number { + let n = 0; + for (const x of s) if (x === c) n++; + return n; +} + +// Pull the first http(s) URL out of the text, undoing the ways a URL gets +// wrapped when a human pastes it: <url>, "url", [label](url), and a trailing +// sentence mark. A closing bracket is only stripped when it does not balance +// one inside the URL, so a URL that legitimately contains brackets survives. +function extractLink(text: string): { link?: string; rest: string } { + // A URL that is the VALUE of a recognised setting (source=remote:https://…) + // is not the request's link — skip to the next candidate. + const re = /https?:\/\/\S+/gi; + let m: RegExpExecArray | null = null; + for (let hit = re.exec(text); hit; hit = re.exec(text)) { + const lineStart = text.lastIndexOf(" ", hit.index - 1) + 1; + const prefix = text.slice(lineStart, hit.index).replace(/^["']|["']$/g, ""); + const kv = /^([A-Za-z][A-Za-z0-9_]*)=\S*$/.exec(prefix); + if (kv && CANONICAL.has(kv[1].toLowerCase())) continue; + m = hit; + break; + } + if (!m) return { rest: text }; + + const start = m.index; + let cand = m[0]; + const before = start > 0 ? text[start - 1] : ""; + + for (let guard = 0; guard < 50; guard++) { + const last = cand[cand.length - 1]; + if (!last) break; + if (".,;:!?".includes(last)) { + cand = cand.slice(0, -1); + continue; + } + if (last === ">" && before === "<") { + cand = cand.slice(0, -1); + continue; + } + if ((last === '"' || last === "'") && (before === last || countChar(cand, last) % 2 === 1)) { + cand = cand.slice(0, -1); + continue; + } + const opener = last === ")" ? "(" : last === "]" ? "[" : last === "}" ? "{" : ""; + if (opener && countChar(cand, opener) < countChar(cand, last)) { + cand = cand.slice(0, -1); + continue; + } + break; + } + + // Remove the whole markdown wrapper when the URL sat inside one, so the + // label doesn't survive as a stray `[Search results](` in the prose. + let from = start; + let to = start + m[0].length; + if (before === "(") { + const open = text.lastIndexOf("[", start); + const close = text.indexOf(")", to - 1); + if (open !== -1 && close !== -1 && text[start - 1] === "(" && text[open] === "[") { + from = open; + to = close + 1; + } + } else if (before === "<") { + from = start - 1; + } + const rest = `${text.slice(0, from)} ${text.slice(to)}`; + return { link: cand, rest }; +} + +// Whitespace-split, but keep a quoted run together and drop the quotes, so +// directive="key claims & contradictions" survives as ONE token. +function tokenize(text: string): string[] { + const out: string[] = []; + let cur = ""; + let quote: string | null = null; + let sawQuote = false; + const push = (): void => { + if (cur !== "" || sawQuote) out.push(cur); + cur = ""; + sawQuote = false; + }; + for (const ch of text) { + if (quote) { + if (ch === quote) { + quote = null; + continue; + } + cur += ch; + continue; + } + if (ch === '"' || ch === "'") { + quote = ch; + sawQuote = true; + continue; + } + if (/\s/.test(ch)) { + push(); + continue; + } + cur += ch; + } + push(); + return out.filter((t) => t !== ""); +} + +function splitList(v: string): string[] { + return v + .split(",") + .map((x) => x.trim()) + .filter((x) => x !== ""); +} + +// ─── Validators. Each one either accepts, or falls back and says why. ─── + +function validBatchSize(raw: string, warnings: string[]): number { + const n = Number(raw); + if (!Number.isInteger(n) || n < 1 || n > 20) { + warnings.push( + `batch_size "${raw}" is not an integer between 1 and 20 — using ` + + `${DEFAULT_BATCH_SIZE}.`, + ); + return DEFAULT_BATCH_SIZE; + } + return n; +} + +function validParseModel(raw: string, warnings: string[]): string { + if (raw === "" || /\s/.test(raw)) { + warnings.push( + `parse_model "${raw}" is not a single token — using ` + + `"${DEFAULT_PARSE_MODEL}".`, + ); + return DEFAULT_PARSE_MODEL; + } + return raw; +} + +function validReportPath(raw: string, warnings: string[]): string | undefined { + if (raw === "" || /\s/.test(raw)) { + warnings.push(`report path "${raw}" is not a single token — ignored.`); + return undefined; + } + if (raw.includes("..")) { + warnings.push(`report path "${raw}" escapes the working directory — ignored.`); + return undefined; + } + if (!raw.toLowerCase().endsWith(".md")) { + warnings.push(`report path "${raw}" is not a .md file — ignored.`); + return undefined; + } + return raw; +} + +function validContentTypes( + raw: string, + warnings: string[], +): ("video" | "post")[] | undefined { + const wanted = splitList(raw.toLowerCase()); + const out = wanted.filter((t): t is "video" | "post" => t === "video" || t === "post"); + const bad = wanted.filter((t) => t !== "video" && t !== "post"); + if (bad.length > 0) { + warnings.push( + `content_types value(s) ignored (expected 'video' and/or 'post'): ${bad.join(", ")}.`, + ); + } + return out.length > 0 ? out : undefined; +} + +// Parse one free-text request. Never throws: a bad value becomes a default +// plus a warning, because failing a whole sweep over a typo'd batch size is +// worse than running it at 8 and saying so. +export function parsePromptRequest(input: string): PromptRequest { + const warnings: string[] = []; + const text = (input ?? "").trim(); + + const { link, rest } = extractLink(text); + const tokens = tokenize(rest); + + const seen = new Map<string, string>(); + const prose: string[] = []; + + for (const tok of tokens) { + const eq = tok.indexOf("="); + if (eq <= 0) { + prose.push(tok); + continue; + } + const rawKey = tok.slice(0, eq); + const value = tok.slice(eq + 1); + // Only a plausible key shape is even considered — this keeps prose like + // "x=y+2" or "a==b" out of the parameter path entirely. + if (!/^[A-Za-z][A-Za-z0-9_]*$/.test(rawKey)) { + prose.push(tok); + continue; + } + const canon = CANONICAL.get(rawKey.toLowerCase()); + if (!canon) { + const hint = typoHint(rawKey); + warnings.push( + `"${rawKey}=" is not a recognised setting, so it was left in the ` + + `question text` + + (hint ? ` — did you mean ${hint}=?` : "") + + `.`, + ); + prose.push(tok); + continue; + } + if (seen.has(canon)) { + warnings.push( + `${canon} was given more than once — using the last value ("${value}").`, + ); + } + seen.set(canon, value); + } + + const req: PromptRequest = { + ...(link ? { link } : {}), + query: prose.join(" ").replace(/\s+/g, " ").trim(), + channels: seen.has("channels") ? splitList(seen.get("channels")!) : [], + groups: seen.has("groups") ? splitList(seen.get("groups")!) : [], + batchSize: seen.has("batch_size") + ? validBatchSize(seen.get("batch_size")!, warnings) + : DEFAULT_BATCH_SIZE, + parseModel: seen.has("parse_model") + ? validParseModel(seen.get("parse_model")!, warnings) + : DEFAULT_PARSE_MODEL, + regex: seen.get("regex") === "true" || seen.get("regex") === "1", + warnings, + }; + + if (seen.has("report")) { + const p = validReportPath(seen.get("report")!, warnings); + if (p) req.reportPath = p; + } + if (seen.has("directive")) { + const d = seen.get("directive")!.trim(); + if (d !== "") req.directive = d; + } + if (seen.has("source")) { + const s = seen.get("source")!.trim(); + if (s !== "") req.source = s; + } + if (seen.has("content_types")) { + const t = validContentTypes(seen.get("content_types")!, warnings); + if (t) req.contentTypes = t; + } + + if (!req.link && req.query === "") { + warnings.push( + "no query and no link — say what to sweep for, or paste a share link.", + ); + } + return req; +} + +// The MCP prompt's declared arguments, routed through the SAME validators, so +// a form-based client gets an error naming the offending argument rather than +// a rendered `ceil(N / Why)`. +export function requestFromArguments( + args: Record<string, unknown>, +): PromptRequest { + const warnings: string[] = []; + const str = (k: string): string | undefined => { + const v = args[k]; + return typeof v === "string" && v.trim() !== "" ? v.trim() : undefined; + }; + + const rawLink = str("link"); + const link = rawLink ? extractLink(rawLink).link : undefined; + if (rawLink && !link) { + warnings.push(`link "${rawLink}" is not an http(s) URL — ignored.`); + } + + const channels = [ + ...(str("channel") ? [str("channel")!] : []), + ...(str("channels") ? splitList(str("channels")!) : []), + ]; + const groups = [ + ...(str("group") ? [str("group")!] : []), + ...(str("groups") ? splitList(str("groups")!) : []), + ]; + + const batchRaw = str("batch_size"); + const modelRaw = str("parse_model"); + const reportRaw = str("report_path") ?? str("report"); + const typesRaw = str("content_types"); + + const req: PromptRequest = { + ...(link ? { link } : {}), + query: str("query") ?? "", + channels, + groups, + batchSize: batchRaw ? validBatchSize(batchRaw, warnings) : DEFAULT_BATCH_SIZE, + parseModel: modelRaw + ? validParseModel(modelRaw, warnings) + : DEFAULT_PARSE_MODEL, + regex: args.regex === true || str("regex") === "true", + warnings, + }; + if (reportRaw) { + const p = validReportPath(reportRaw, warnings); + if (p) req.reportPath = p; + } + if (str("directive")) req.directive = str("directive"); + if (str("source")) req.source = str("source"); + if (typesRaw) { + const t = validContentTypes(typesRaw, warnings); + if (t) req.contentTypes = t; + } + return req; +} + +// The ⚠ block a caller puts at the top of its output. Empty when clean. +export function renderWarnings(warnings: string[]): string { + if (warnings.length === 0) return ""; + return `⚠ ${warnings.join("\n⚠ ")}\n\n`; +} diff --git a/mcp/src/protocol.test.ts b/mcp/src/protocol.test.ts @@ -0,0 +1,212 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { mkdtemp, mkdir, writeFile, readdir, rm } from "node:fs/promises"; +import os from "node:os"; +import path from "node:path"; +import { fileURLToPath } from "node:url"; +import { Client } from "@modelcontextprotocol/client"; +import { + StdioClientTransport, + getDefaultEnvironment, +} from "@modelcontextprotocol/client/stdio"; +import { transcriptPageFileName } from "yt-dlp-transcript-common/lib/manifest"; + +// ─── Real-process protocol coverage ─── +// +// Every other server test links a client and server over InMemoryTransport, +// which only ever exercises the 2025 ("legacy") era — the era decision lives in +// the serving entry (serveStdio), which an in-memory pair never runs. So a +// 2026-only regression would be invisible to the ~100 tests next door. +// +// This file is the guard: it spawns the real `src/index.ts` over stdio and +// drives it twice — once negotiating the modern (2026-07-28) era, once on the +// plain legacy handshake — asserting both still serve the same tools. Keep it +// unconditional and in the default `test` script. + +const HERE = path.dirname(fileURLToPath(import.meta.url)); +const ENTRY = path.join(HERE, "index.ts"); +const TSX = path.join(HERE, "..", "node_modules", ".bin", "tsx"); + +// The exact advertised tool list, in order. `tools/list` is cached (see +// CACHE_HINTS) and its order is part of what clients see, so pin it: a rename +// or an accidental reshuffle has to be a deliberate edit here. +const EXPECTED_TOOLS = [ + "list_channels", + "search_transcripts", + "enumerate_matches", + "get_transcript", + "get_post", + "get_thread", + "get_transcripts", + "get_video_metadata", + "list_sources", + "resolve_source", + "open_link", + "sweep_plan", + "ask_plan", +]; + +// A minimal composed public dir: one channel, two videos, real shard layout. +async function writeFixture(): Promise<string> { + const dir = await mkdtemp(path.join(os.tmpdir(), "mcp-protocol-")); + await writeFile( + path.join(dir, "corpus.json"), + JSON.stringify({ + channels: [{ slug: "chan-a", name: "Channel A", videoCount: 2 }], + }), + ); + const chDir = path.join(dir, "transcripts", "chan-a"); + await mkdir(chDir, { recursive: true }); + await writeFile( + path.join(chDir, "manifest.json"), + JSON.stringify({ + version: 1, + channelSlug: "chan-a", + pageCount: 1, + maxPageBytes: 1000, + generatedAt: "2026-01-01", + slugToPage: { a1: 0, a2: 0 }, + }), + ); + const rec = (id: string, title: string, text: string) => ({ + slug: `chan-a/${id}`, + id, + channelSlug: "chan-a", + title, + uploadDate: "20240101", + duration: 300, + channel: "Channel A", + description: "", + tags: [], + isLivestream: false, + ageRestricted: false, + platform: "youtube", + webpageUrl: `https://example.test/${id}`, + cues: [{ start: 12, end: 15, text }], + }); + await writeFile( + path.join(chDir, transcriptPageFileName(0)), + JSON.stringify([ + rec("a1", "Coffee one", "i love coffee"), + rec("a2", "Tea two", "i love tea"), + ]), + ); + return dir; +} + +type Session = { + client: Client; + stateDir: string; + close: () => Promise<void>; +}; + +// Spawn the real server over stdio against `corpusDir`, connecting with the +// given negotiation mode. Each session gets its own state dir so the assertion +// that nothing is persisted is meaningful (and so a stale state file on the +// developer's machine can never leak into the test). +async function connect( + corpusDir: string, + versionNegotiation?: { mode: "auto" | "legacy" }, +): Promise<Session> { + const stateDir = await mkdtemp(path.join(os.tmpdir(), "mcp-state-")); + const transport = new StdioClientTransport({ + command: TSX, + args: [ENTRY, "--local", corpusDir], + // The old controller keyed a state file off this; nothing reads it now, + // and the assertion below is that the dir stays empty regardless. + env: { ...getDefaultEnvironment(), TRANSCRIPT_MCP_STATE_DIR: stateDir }, + stderr: "pipe", + }); + const client = new Client( + { name: "protocol-test", version: "0" }, + { capabilities: {}, ...(versionNegotiation ? { versionNegotiation } : {}) }, + ); + await client.connect(transport); + return { + client, + stateDir, + close: async () => { + await client.close(); + await rm(stateDir, { recursive: true, force: true }); + }, + }; +} + +function firstText(res: unknown): string { + const content = (res as { content: { type: string; text: string }[] }).content; + return content.map((c) => c.text).join("\n"); +} + +test("stdio: the modern (2026-07-28) era negotiates and serves the tools", async (t) => { + const dir = await writeFixture(); + t.after(() => rm(dir, { recursive: true, force: true })); + const s = await connect(dir, { mode: "auto" }); + t.after(() => s.close()); + + assert.equal( + s.client.getProtocolEra(), + "modern", + "server/discover should select the modern era over stdio", + ); + + const tools = await s.client.listTools(); + assert.deepEqual(tools.tools.map((x) => x.name), EXPECTED_TOOLS); + + const res = await s.client.callTool({ + name: "search_transcripts", + arguments: { query: "coffee" }, + }); + const out = firstText(res); + assert.match(out, /Coffee one/); + assert.match(out, /total 1 match/); +}); + +test("stdio: the legacy (2025) handshake still serves the same tools", async (t) => { + const dir = await writeFixture(); + t.after(() => rm(dir, { recursive: true, force: true })); + // No versionNegotiation at all — the SDK's default posture, and what a + // client that has never heard of 2026-07-28 sends. + const s = await connect(dir); + t.after(() => s.close()); + + assert.equal(s.client.getProtocolEra(), "legacy"); + + const tools = await s.client.listTools(); + assert.deepEqual(tools.tools.map((x) => x.name), EXPECTED_TOOLS); + + const res = await s.client.callTool({ + name: "search_transcripts", + arguments: { query: "tea" }, + }); + assert.match(firstText(res), /Tea two/); +}); + +test("stdio: prompts are served on both eras", async (t) => { + const dir = await writeFixture(); + t.after(() => rm(dir, { recursive: true, force: true })); + const modern = await connect(dir, { mode: "auto" }); + t.after(() => modern.close()); + const legacy = await connect(dir, { mode: "legacy" }); + t.after(() => legacy.close()); + + for (const s of [modern, legacy]) { + const prompts = await s.client.listPrompts(); + assert.deepEqual(prompts.prompts.map((p) => p.name), ["sweep"]); + } +}); + +test("stdio: a read-only session writes no state file", async (t) => { + const dir = await writeFixture(); + t.after(() => rm(dir, { recursive: true, force: true })); + const s = await connect(dir, { mode: "auto" }); + t.after(() => s.close()); + + await s.client.callTool({ name: "list_channels", arguments: {} }); + await s.client.callTool({ + name: "search_transcripts", + arguments: { query: "coffee" }, + }); + + const left = await readdir(s.stateDir).catch(() => [] as string[]); + assert.deepEqual(left, [], "the server must not persist anything"); +}); diff --git a/mcp/src/search.test.ts b/mcp/src/search.test.ts @@ -1,7 +1,8 @@ import { test } from "node:test"; import assert from "node:assert/strict"; -import { Client } from "@modelcontextprotocol/sdk/client/index.js"; -import { InMemoryTransport } from "@modelcontextprotocol/sdk/inMemory.js"; +// Client AND InMemoryTransport must come from the SAME package: each bundles +// its own copy with private state, so a pair split across packages never links. +import { Client, InMemoryTransport } from "@modelcontextprotocol/client"; import type { ChannelTranscriptsManifest, ChannelSubsManifest, @@ -33,6 +34,7 @@ import { type SearchFilters, } from "./search"; import { createServer } from "./server"; +import { SourceRegistry } from "./sourceRegistry"; import { MISSING_STATES, VIDEO_STATES, @@ -849,8 +851,8 @@ test("server: the sweep prompt reflects a supplied group scope", async () => { arguments: { query: "k cups", group: "other" }, }); const text = (got.messages[0].content as { text: string }).text; - assert.match(text, /scoped to group "other"/); - assert.match(text, /group "other"/); + assert.match(text, /scoped to groups "other"/); + assert.match(text, /groups: \["other"\]/); await client.close(); }); @@ -892,7 +894,9 @@ test("server: the sweep prompt accepts a link= seed and drives open_link", async }); const text = (got.messages[0].content as { text: string }).text; assert.match(text, /open_link/); - assert.match(text, /apply:true/); + // One call now does the whole job; dry_run is the opt-in preview. + assert.match(text, /dry_run: true/); + assert.ok(!text.includes("apply:true"), "the two-phase apply is gone"); assert.match(text, /https:\/\/site\.example\/\?q=coffee/); // A link-only sweep needs no query argument. assert.match(text, /share link/); @@ -1133,3 +1137,339 @@ test("server: search_transcripts renders a post hit without a moment link", asyn assert.doesNotMatch(out, /moment_base/); await client.close(); }); + +// ─── Server-level coverage for the single-record reads and open_link ─── +// +// These three had no server-level test at all, and open_link's default just +// changed from "preview" to "decode and search in one call" — exactly the kind +// of change that needs a test underneath it before it lands. + +// A registry whose every spec resolves to the same StubSource, so a share +// link's origin can be resolved without a live fetch. +function stubRegistry(): SourceRegistry { + return new SourceRegistry( + { kind: "remote", url: "https://site.example" }, + { + build: () => new StubSource(), + probe: async () => "remote" as const, + }, + ); +} + +async function connectRegistry(): Promise<Client> { + const server = createServer(stubRegistry()); + const [ct, st] = InMemoryTransport.createLinkedPair(); + const client = new Client({ name: "test", version: "0" }, { capabilities: {} }); + await Promise.all([server.connect(st), client.connect(ct)]); + return client; +} + +test("server: get_transcript returns the full transcript as markdown", async () => { + const client = await connectClient(new StubSource()); + const out = firstText( + await client.callTool({ + name: "get_transcript", + arguments: { video_id: "a1" }, + }), + ); + assert.match(out, /Coffee one/); + assert.match(out, /i love coffee/); + assert.match(out, /\(corpus: /, "every result names its corpus"); + await client.close(); +}); + +test("server: get_transcript reports a missing id as an error", async () => { + const client = await connectClient(new StubSource()); + const res = await client.callTool({ + name: "get_transcript", + arguments: { video_id: "nope" }, + }); + assert.equal((res as { isError?: boolean }).isError, true); + assert.match(firstText(res), /video not found: nope/); + await client.close(); +}); + +test("server: get_video_metadata returns metadata without the cue bodies", async () => { + const client = await connectClient(new StubSource()); + const out = firstText( + await client.callTool({ + name: "get_video_metadata", + arguments: { video_id: "a1" }, + }), + ); + const parsed = JSON.parse(out.slice(0, out.lastIndexOf("}") + 1)); + assert.equal(parsed.id, "a1"); + assert.equal(parsed.title, "Coffee one"); + assert.equal(parsed.channelName, "Channel A"); + assert.equal(parsed.cueCount, 1); + assert.equal(parsed.cues, undefined, "the transcript body is not included"); + await client.close(); +}); + +test("server: get_video_metadata reports a missing id as an error", async () => { + const client = await connectClient(new StubSource()); + const res = await client.callTool({ + name: "get_video_metadata", + arguments: { video_id: "nope" }, + }); + assert.equal((res as { isError?: boolean }).isError, true); + assert.match(firstText(res), /video not found: nope/); + await client.close(); +}); + +test("server: open_link decodes AND searches in one call, returning the handle", async () => { + const client = await connectRegistry(); + const out = firstText( + await client.callTool({ + name: "open_link", + arguments: { link: "https://site.example/?q=coffee" }, + }), + ); + // The plan… + assert.match(out, /## Share link/); + assert.match(out, /Corpus handle: `remote:https:\/\/site\.example`/); + // …and the results, from the same call. + assert.match(out, /Coffee one/); + assert.match(out, /video\(s\) matched/); + await client.close(); +}); + +test("server: open_link dry_run returns the plan and runs no search", async () => { + const client = await connectRegistry(); + const out = firstText( + await client.callTool({ + name: "open_link", + arguments: { link: "https://site.example/?q=coffee", dry_run: true }, + }), + ); + assert.match(out, /Share-link plan \(dry run\)/); + assert.match(out, /no search was performed/); + assert.ok(!out.includes("video(s) matched"), "nothing was searched"); + await client.close(); +}); + +test("server: open_link honours overrides", async () => { + const client = await connectRegistry(); + const out = firstText( + await client.callTool({ + name: "open_link", + arguments: { + link: "https://site.example/?q=coffee", + overrides: { query: "k cups" }, + dry_run: true, + }, + }), + ); + assert.match(out, /query replaced with "k cups"/); + await client.close(); +}); + +test("server: open_link accepts a posts query_scope override", async () => { + const client = await connectRegistry(); + const out = firstText( + await client.callTool({ + name: "open_link", + arguments: { + link: "https://site.example/?q=coffee", + overrides: { query: "bluesky", query_scope: "posts" }, + dry_run: true, + }, + }), + ); + // Declared in the schema but dropped by parseOverrides until now, so a + // caller asking for the post corpus silently got transcripts. + assert.match(out, /posts/); + await client.close(); +}); + +test("server: open_link rejects a link that is not a URL", async () => { + const client = await connectRegistry(); + const res = await client.callTool({ + name: "open_link", + arguments: { link: "not a url" }, + }); + assert.equal((res as { isError?: boolean }).isError, true); + await client.close(); +}); + +// ─── W4: coverage is mechanical, and silently-ignored parameters are not ─── + +test("server: enumerate_matches returns the whole set in one call", async () => { + const client = await connectClient(new StubSource()); + const enumerated = firstText( + await client.callTool({ + name: "enumerate_matches", + arguments: { query: "coffee", batch_size: 2 }, + }), + ); + // The complete worklist, with the batch arithmetic already done. + assert.match(enumerated, /the complete worklist/); + assert.match(enumerated, /complete set: yes/); + assert.match(enumerated, /batch\(es\) of 2/); + assert.match(enumerated, /- a1 \| video \| Channel A \|/); + assert.ok(!enumerated.includes("i love coffee"), "no snippet bodies"); + + // And its total agrees with what search reports — the invariant a sweep + // relies on when it claims coverage. + const searched = firstText( + await client.callTool({ + name: "search_transcripts", + arguments: { query: "coffee", limit: 1 }, + }), + ); + const enumTotal = /(\d+) match\(es\) for "coffee"/.exec(enumerated)?.[1]; + const searchTotal = /total (\d+) match/.exec(searched)?.[1]; + assert.equal(enumTotal, searchTotal); + await client.close(); +}); + +test("server: enumerate_matches with no matches says so plainly", async () => { + const client = await connectClient(new StubSource()); + const out = firstText( + await client.callTool({ + name: "enumerate_matches", + arguments: { query: "zzzznotathing" }, + }), + ); + assert.match(out, /No matches for "zzzznotathing"/); + assert.match(out, /complete set: yes/); + await client.close(); +}); + +test("server: enumerate_matches shouts when a cap made coverage partial", async () => { + const client = await connectClient(new StubSource()); + const out = firstText( + await client.callTool({ + name: "enumerate_matches", + arguments: { query: "coffee", max_pages: 1 }, + }), + ); + // The first line, not a footer note. + assert.ok(out.startsWith("⚠ COVERAGE PARTIAL"), out.slice(0, 60)); + assert.match(out, /are a SAMPLE, not the full set/); + assert.match(out, /a PARTIAL worklist/); + assert.ok(!out.includes("the complete worklist")); + assert.match(out, /complete set: NO — capped/); + await client.close(); +}); + +test("server: an incomplete search page warns ABOVE the hits", async () => { + const client = await connectClient(new StubSource()); + const out = firstText( + await client.callTool({ + name: "search_transcripts", + arguments: { query: "coffee", limit: 2 }, + }), + ); + assert.ok(out.startsWith("⚠ INCOMPLETE PAGE"), out.slice(0, 60)); + assert.match(out, /Do NOT report a count/); + assert.match(out, /enumerate_matches/); + // The old footer is still there, so existing consumers keep working. + assert.match(out, /has_more: yes/); + await client.close(); +}); + +test("server: a complete search page carries no banner", async () => { + const client = await connectClient(new StubSource()); + const out = firstText( + await client.callTool({ + name: "search_transcripts", + arguments: { query: "coffee", limit: 50 }, + }), + ); + assert.ok(!out.includes("INCOMPLETE PAGE")); + assert.match(out, /has_more: no/); + await client.close(); +}); + +test("server: an empty posts corpus is named as empty, not as zero pages", async () => { + const client = await connectClient(new StubSource()); + // chan-a is video-only, so a posts-scoped search there has no index at all. + const out = firstText( + await client.callTool({ + name: "search_transcripts", + arguments: { + query: "anything", + channels: ["chan-a"], + content_types: ["post"], + }, + }), + ); + assert.match(out, /the post corpus is EMPTY here, not merely unmatched/); + await client.close(); +}); + +test("server: a posts corpus that WAS searched reports its own scan counts", async () => { + const client = await connectClient(new StubSource()); + const out = firstText( + await client.callTool({ + name: "search_transcripts", + arguments: { + query: "zephyrpost", + channels: ["chan-b"], + content_types: ["post"], + }, + }), + ); + assert.match(out, /posts: scanned \d+ page\(s\) across 1 posting channel\(s\)/); + assert.ok(!out.includes("EMPTY here")); + await client.close(); +}); + +test("server: get_transcripts queries[] merges windows and counts each query", async () => { + const client = await connectClient(new StubSource()); + const out = firstText( + await client.callTool({ + name: "get_transcripts", + arguments: { video_ids: ["a1"], queries: ["coffee", "zzzznotathing"] }, + }), + ); + // Both queries are reported per video, so the one that matched nothing is + // visible rather than silently absorbed into the merged excerpt. + assert.match(out, /"coffee": 1/); + assert.match(out, /"zzzznotathing": 0/); + assert.match(out, /i love coffee/); + assert.match(out, /queries: "coffee", "zzzznotathing"/); + await client.close(); +}); + +test("server: get_transcripts content_types reaches the post corpus", async () => { + const client = await connectClient(new StubSource()); + const out = firstText( + await client.callTool({ + name: "get_transcripts", + arguments: { video_ids: ["p1"] }, + }), + ); + // p1 is a post id; the declared content_types was never read before, so + // this used to come back "not found". + assert.match(out, /zephyrpost/); + assert.ok(!out.includes("not found: p1")); + await client.close(); +}); + +test("server: get_transcripts video-only scope does not fall through to posts", async () => { + const client = await connectClient(new StubSource()); + const out = firstText( + await client.callTool({ + name: "get_transcripts", + arguments: { video_ids: ["p1"], content_types: ["video"] }, + }), + ); + assert.match(out, /not found: p1/); + await client.close(); +}); + +test("server: the unadvertised channel/group singulars are still parsed", async () => { + const client = await connectClient(new StubSource()); + // Dropped from the advertised schema in favour of the plural forms, but a + // stored habit (or an old transcript) must not hard-fail. + const out = firstText( + await client.callTool({ + name: "search_transcripts", + arguments: { query: "coffee", channel: "chan-a" }, + }), + ); + assert.match(out, /scope: 1 channel/); + await client.close(); +}); diff --git a/mcp/src/search.ts b/mcp/src/search.ts @@ -191,6 +191,12 @@ export type SearchResult = { // expansion it applied). Empty for a plain or explicit-regex search. firedAliases: SearchAlias[]; scanned: { channels: number; pages: number }; + // The posts pass, counted separately. Without this a posts-only search over a + // corpus where no channel ships a posts index reports "scanned 0 page(s) + // across 0 channel(s)" — indistinguishable from "searched everything, found + // nothing". `channels` counts the channels that actually HAVE a posts index, + // so 0 with `requested` true means the post corpus is empty here. + postsScanned: { requested: boolean; channels: number; pages: number }; // Coverage is partial — the page cap (MAX_PAGES) or the video cap // (HARD_VIDEO_CAP) was reached before the corpus was fully scanned. truncated: boolean; @@ -333,6 +339,10 @@ export async function searchTranscripts( const contentTypes = opts.contentTypes ?? [...ALL_CONTENT_TYPES]; const wantVideos = contentTypes.includes("video"); const wantPosts = contentTypes.includes("post"); + // Counted apart from the video pass so "no posts index anywhere in scope" is + // distinguishable from "searched the posts and found nothing". + let postChannelsScanned = 0; + let postPagesScanned = 0; outer: for (const ch of channels) { if (!wantVideos) break; @@ -406,6 +416,7 @@ export async function searchTranscripts( continue; } if (!pm) continue; // not a social channel + postChannelsScanned++; for (let page = 0; page < pm.pageCount; page++) { if (pagesScanned >= maxPages) { truncated = true; @@ -418,6 +429,7 @@ export async function searchTranscripts( continue; } pagesScanned++; + postPagesScanned++; for (const post of posts) { if (!match(post.text)) continue; all.push({ @@ -460,6 +472,11 @@ export async function searchTranscripts( hasMore: offset + limit < total, firedAliases, scanned: { channels: channelsScanned, pages: pagesScanned }, + postsScanned: { + requested: wantPosts, + channels: postChannelsScanned, + pages: postPagesScanned, + }, truncated, selection: { all: selection.all, diff --git a/mcp/src/server.ts b/mcp/src/server.ts @@ -1,10 +1,9 @@ -import { Server } from "@modelcontextprotocol/sdk/server/index.js"; import { - CallToolRequestSchema, - ListToolsRequestSchema, - ListPromptsRequestSchema, - GetPromptRequestSchema, -} from "@modelcontextprotocol/sdk/types.js"; + Server, + type Prompt, + type ServerOptions, + type Tool, +} from "@modelcontextprotocol/server"; import { transcriptToMarkdown } from "yt-dlp-transcript-common/lib/transcriptToMarkdown"; import { formatDate } from "yt-dlp-transcript-common/lib/format"; import { momentUrl, momentBaseUrl } from "yt-dlp-transcript-common/lib/momentUrl"; @@ -16,15 +15,9 @@ import { FALLBACK_GROUP, type ChannelGroup, } from "yt-dlp-transcript-common/lib/channelGroups"; -import { - HubSource, - RemoteSource, - type ChannelRef, - type HubSite, - type ShardSource, -} from "./source"; +import type { ChannelRef, HubSite, ShardSource } from "./source"; import type { SourceSpec } from "./sources"; -import type { SourceController } from "./sourceController"; +import { SourceRegistry, type ResolvedSource } from "./sourceRegistry"; import type { Post } from "yt-dlp-transcript-common/lib/posts"; import { searchTranscripts, @@ -40,6 +33,17 @@ import { type ScopedSnippet, } from "./search"; import { + parsePromptRequest, + requestFromArguments, + DEFAULT_REPORT_PATH, + type PromptRequest, +} from "./promptRequest"; +import { + buildSweepInstructions, + buildAskInstructions, + type PlanContext, +} from "./instructions"; +import { decodeShareLink, applyLinkOverrides, renderQueryTree, @@ -130,7 +134,42 @@ function errorText(s: string): ToolResult { return { content: [{ type: "text", text: s }], isError: true }; } -const TOOLS = [ +// Cache hints for the 2026-07-28 revision's cacheable results (`ttlMs` / +// `cacheScope`). Our three cacheable operations — the tool list, the prompt +// list, and the `server/discover` descriptor the SDK builds from them — are +// literal constants in this file: identical for every caller, unchanged for the +// life of the process, and carrying no corpus data (the corpus is chosen +// per-call now, not per-connection). So they are safely `public` and worth an +// hour. Everything else — every read tool's result — stays uncached by default. +// Invalid values throw a RangeError in the Server constructor, so a typo here +// fails at startup rather than on the wire. +const CACHE_HINTS: NonNullable<ServerOptions["cacheHints"]> = { + "tools/list": { ttlMs: 3_600_000, cacheScope: "public" }, + "prompts/list": { ttlMs: 3_600_000, cacheScope: "public" }, + "server/discover": { ttlMs: 3_600_000, cacheScope: "public" }, +}; + +// The `source` argument every read tool accepts. One shared constant spliced +// into each literal schema, so adding a tool can't accidentally omit it and +// `additionalProperties: false` still holds everywhere. +const SOURCE_ARG = { + source: { + type: "string", + description: + "Which corpus to read, as a handle: 'default' (this server's startup " + + "corpus — the same as omitting it), 'local:<dir>', 'remote:<site url>', " + + "'hub:<hub url>', or 'hub:<hub url>#<siteA,siteB>' for a hub subset. A " + + "bare site URL, or a hub member's siteId/title, also works and is " + + "normalised. There is no active source and nothing persists between " + + "calls: every call reads exactly what it asks for, and every result " + + "ends with the canonical handle it actually read, as '(corpus: …)'.", + }, +} as const; + +// The advertised tool list. Deliberately a hand-written literal array in a +// fixed order — never generated from a Map or Object.keys — so `tools/list` is +// byte-stable across processes and safely cacheable (see CACHE_HINTS). +export const TOOLS: Tool[] = [ { name: "list_channels", description: @@ -140,7 +179,20 @@ const TOOLS = [ "compact list of the groups (id · name · channel count) and whether each " + "is selected by default — so you can pick a group or channels to scope a " + "search or sweep to.", - inputSchema: { type: "object", properties: {}, additionalProperties: false }, + inputSchema: { + type: "object", + properties: { + ...SOURCE_ARG, + refresh: { + type: "boolean", + description: + "Re-read the channel list instead of using this process's cached " + + "copy (default false). Only needed if the corpus was rebuilt while " + + "the server was running.", + }, + }, + additionalProperties: false, + }, }, { name: "search_transcripts", @@ -165,30 +217,22 @@ const TOOLS = [ inputSchema: { type: "object", properties: { + ...SOURCE_ARG, query: { type: "string", description: "Term, phrase, or regex to find." }, - channel: { - type: "string", - description: "Optional channel slug or name to restrict the search to.", - }, channels: { type: "array", items: { type: "string" }, description: "Optional list of channel slugs/names to restrict the search to " + - "(union with channel/group/groups).", - }, - group: { - type: "string", - description: - "Optional channel group (by id or display name, e.g. 'other' or " + - "'Extended Universe') to expand to its member channels.", + "(union with groups). A single channel is just a one-element list.", }, groups: { type: "array", items: { type: "string" }, description: "Optional list of channel groups (ids or names) to expand to their " + - "member channels (union with the other scope fields).", + "member channels (union with channels). A single group is just a " + + "one-element list.", }, regex: { type: "boolean", @@ -250,6 +294,70 @@ const TOOLS = [ }, }, { + name: "enumerate_matches", + description: + "Get a query's COMPLETE match set as a worklist, in one call — every " + + "matching id/title/channel/date, no snippets, plus the batch count for a " + + "sweep. Use this, not paged search_transcripts, whenever you need to " + + "cover or count everything: the engine materialises the whole match set " + + "before slicing, so enumerating is ONE scan where paging the same query " + + "would be one full scan per page. If a cap is hit the FIRST line says " + + "coverage is partial and the list is a sample — a result without that " + + "line is the complete set, and is the only basis on which you may state " + + "a total or claim full coverage. Same scoping and alias expansion as " + + "search_transcripts.", + inputSchema: { + type: "object", + properties: { + ...SOURCE_ARG, + query: { type: "string", description: "Term, phrase, or regex to find." }, + channels: { + type: "array", + items: { type: "string" }, + description: "Optional channel slugs/names to restrict the scan to.", + }, + groups: { + type: "array", + items: { type: "string" }, + description: "Optional channel groups (ids or names) to restrict to.", + }, + regex: { + type: "boolean", + description: + "Treat query as a case-insensitive regex (default false). " + + "Disables alias expansion.", + }, + use_aliases: { + type: "boolean", + description: + "Expand the query via curated aliases (default true; ignored with " + + "regex).", + }, + content_types: { + type: "array", + items: { type: "string", enum: ["video", "post"] }, + description: + "Which corpora to enumerate: 'video' and/or 'post'. Defaults to " + + "BOTH.", + }, + batch_size: { + type: "number", + description: + "Videos per batch, used only to report the batch count (default " + + "8).", + }, + max_pages: { + type: "number", + description: + "Max shard pages to scan before stopping (default 400). Reaching " + + "it makes coverage partial, which is reported loudly.", + }, + }, + required: ["query"], + additionalProperties: false, + }, + }, + { name: "get_transcript", description: "Fetch one video's full transcript as clean markdown (metadata header + " + @@ -258,6 +366,7 @@ const TOOLS = [ inputSchema: { type: "object", properties: { + ...SOURCE_ARG, video_id: { type: "string", description: "The video id." }, channel: { type: "string", description: "Optional owning channel slug/name." }, timestamps: { @@ -278,6 +387,7 @@ const TOOLS = [ inputSchema: { type: "object", properties: { + ...SOURCE_ARG, post_id: { type: "string", description: "The post id (tweet id / atproto rkey)." }, channel: { type: "string", @@ -297,6 +407,7 @@ const TOOLS = [ inputSchema: { type: "object", properties: { + ...SOURCE_ARG, post_id: { type: "string", description: "Any post id in the thread (root or a reply).", @@ -322,6 +433,7 @@ const TOOLS = [ inputSchema: { type: "object", properties: { + ...SOURCE_ARG, video_ids: { type: "array", items: { type: "string" }, @@ -333,6 +445,16 @@ const TOOLS = [ "Optional term/phrase/regex to window around. When given, only " + "excerpts around matches are returned instead of full transcripts.", }, + queries: { + type: "array", + items: { type: "string" }, + description: + "Several terms to window around in ONE pass (max 8; unioned with " + + "query). Windows for all of them merge per video, and each " + + "video's header reports the per-query match count — so a query " + + "that matched nothing is visible. Use this instead of re-reading " + + "the same videos once per term.", + }, regex: { type: "boolean", description: @@ -407,6 +529,7 @@ const TOOLS = [ inputSchema: { type: "object", properties: { + ...SOURCE_ARG, video_id: { type: "string", description: "The video id." }, channel: { type: "string", description: "Optional owning channel slug/name." }, }, @@ -417,92 +540,61 @@ const TOOLS = [ { name: "list_sources", description: - "Show the ACTIVE corpus this server is currently reading (label + kind + " + - "target). When there is a hub context, also lists the hub's member sites " + - "(siteId · title · url), marking which are in the current subset — so you " + - "can pick a site or a subset to switch to with use_source. You can also " + - "switch to an arbitrary remote site URL, local dir, or hub URL. The chosen " + - "source persists across reconnects. Strictly read-only either way.", - inputSchema: { type: "object", properties: {}, additionalProperties: false }, - }, - { - name: "use_source", - description: - "Switch which corpus this server reads — on the fly, and the choice " + - "persists across reconnects. Provide EXACTLY ONE target: `site` (a hub " + - "member by siteId or title — becomes a single-site source with full group/" + - "alias support), `sites` (a list of hub members by siteId/title — a " + - "federated subset of the hub), `remote` (an arbitrary deployed site " + - "origin URL), `local` (a composed public dir on disk), or `hub` (an " + - "arbitrary hub URL to federate). Unknown site tokens are reported, not " + - "silently dropped. Still strictly read-only — this only changes which " + - "already-published static shards are read; nothing is written to any corpus.", + "Show the corpora this server can read: its DEFAULT corpus (the one used " + + "when a call omits `source`) and, when that default is a hub, the hub's " + + "member sites (siteId · title · url) as ready-to-paste handles. There is " + + "no 'active' source to change — every read tool takes its own `source` " + + "handle, and a call that omits it reads the default. Strictly read-only.", inputSchema: { type: "object", - properties: { - site: { - type: "string", - description: - "A hub member site to scope to, by siteId or title " + - "(case-insensitive). Switches to a single-site remote source.", - }, - sites: { - type: "array", - items: { type: "string" }, - description: - "A list of hub member sites (siteId or title) to federate as a " + - "subset of the hub.", - }, - remote: { - type: "string", - description: "An arbitrary deployed site origin URL to read.", - }, - local: { - type: "string", - description: "A composed public dir on disk to read.", - }, - hub: { - type: "string", - description: "An arbitrary hub URL to federate over every member.", - }, - }, + properties: { ...SOURCE_ARG }, additionalProperties: false, }, }, { - name: "reset_source", + name: "resolve_source", description: - "Return to the source this server was started with (its --hub/--remote/" + - "--local flag or env) and clear the persisted selection.", - inputSchema: { type: "object", properties: {}, additionalProperties: false }, + "Turn a corpus reference into the canonical handle to pass as `source`, " + + "and check it can actually be read. Accepts a handle, a bare site or hub " + + "URL (auto-detected), or a hub member's siteId/title. Returns the " + + "canonical handle plus a channel/group count. It changes NOTHING: there " + + "is no active source, so this only tells you what to pass — every read " + + "tool still needs the handle in its own `source` argument.", + inputSchema: { + type: "object", + properties: { ...SOURCE_ARG }, + required: ["source"], + additionalProperties: false, + }, }, { name: "open_link", description: "Paste an archilyzer viewer **share link** (origin + a `qt=` query tree + " + - "`fc/ft/fa/fav/fdf/fdt` filters) to re-run that exact search here — at full " + - "fidelity, with clickable timestamped result links. Two-phase: by default " + - "(apply:false) it PREVIEWS — decodes the link and returns a plan (resolved " + - "source, the query tree rendered readably, every active filter, the channel " + - "scope validated against the live corpus, and anything ignored, e.g. the " + - "vestigial `fk` tracks) WITHOUT switching source or searching. Adjust in " + - "natural language by re-calling with `overrides` (e.g. clear_availability " + - "to drop the availability filter, channels to re-scope, query to replace " + - "the search). When the plan looks right, call again with apply:true: it " + - "switches the active source to the link's origin (hub or single-site, " + - "auto-detected) and returns the first page of linked results. Read-only.", + "`fc/ft/fa/fav/fdf/fdt` filters) to re-run that exact search here — at " + + "full fidelity, with clickable timestamped result links. ONE call does " + + "the whole job: it decodes the link, resolves its origin to a source " + + "handle (hub or single-site, auto-detected), validates the channel scope " + + "against that live corpus, and returns the plan, the first page of " + + "results, and the handle — pass that handle as `source` on the follow-up " + + "calls. Adjust in natural language by re-calling with `overrides` (e.g. " + + "clear_availability to drop the availability filter, channels to " + + "re-scope, query to replace the search). Set dry_run:true to see the plan " + + "without searching. Read-only: nothing about this server changes.", inputSchema: { type: "object", properties: { + ...SOURCE_ARG, link: { type: "string", description: "The archilyzer viewer share URL to decode and run.", }, - apply: { + dry_run: { type: "boolean", description: - "false (default) previews the plan only; true switches source and " + - "runs the search.", + "true returns the decoded plan WITHOUT running the search — for " + + "confirming a link's scope and filters first. Default false: " + + "decode and search in one call.", }, overrides: { type: "object", @@ -574,115 +666,179 @@ const TOOLS = [ }, limit: { type: "number", - description: "Results per page when apply:true (default 20).", + description: "Results per page (default 20).", }, offset: { type: "number", - description: "Skip this many matches when apply:true (default 0).", + description: "Skip this many matches (default 0).", }, }, required: ["link"], additionalProperties: false, }, }, -]; - -// The slice of SourceController the server needs. A bare ShardSource is wrapped -// in a fixed, non-persisting implementation so the existing tools (and tests -// that pass a plain source) keep working while switching is simply disabled. -export interface SourceControllerLike { - current: ShardSource; - activeSpec?: SourceSpec; - switchTo(spec: SourceSpec): Promise<void>; - reset(): Promise<void>; - listHubSites(): Promise<HubSite[]>; - hubUrl(): string | undefined; -} - -// Wrap a bare ShardSource so the server always talks to a controller. Switching -// is disabled (the server was pointed at one fixed source), but list_sources -// still reports it and a hub source can still enumerate its members. -function fixedController(source: ShardSource): SourceControllerLike { - return { - current: source, - activeSpec: undefined, - async switchTo(): Promise<void> { - throw new Error( - "this server is pinned to a single source; source switching is disabled", - ); - }, - async reset(): Promise<void> { - // nothing persisted, nothing to return to - }, - async listHubSites(): Promise<HubSite[]> { - return source instanceof HubSource ? source.listSites() : []; + { + name: "sweep_plan", + description: + "Turn a request written in plain English — optionally with a pasted " + + "share link — into the step-by-step plan for a corpus sweep, with the " + + "scope, corpus handle and batch count already resolved. Call this when " + + "the user asks to sweep, survey, or systematically go through the " + + "archive. It only RETURNS instructions: it reads nothing, writes " + + "nothing, and starts nothing. Following the plan it returns is a long " + + "job that writes a report file with your own Write/Edit tools, so run " + + "it when that is what was asked for, and show the plan's ⚠ warnings.", + inputSchema: { + type: "object", + properties: { + ...SOURCE_ARG, + request: { + type: "string", + description: + "The whole request, verbatim and unsplit: a share link and/or " + + "what to sweep for, in the user's own words, punctuation intact. " + + "Recognised settings may be appended as key=value — channels=, " + + "groups=, batch_size=, parse_model=, report=, directive=, " + + "source=, content_types=, regex= (quote a multi-word value). " + + "Anything else stays part of the question.", + }, + }, + required: ["request"], + additionalProperties: false, }, - hubUrl(): string | undefined { - return source instanceof HubSource ? source.hubBase : undefined; + }, + { + name: "ask_plan", + description: + "Turn a plain-English question about the archive — optionally with a " + + "pasted share link — into a step-by-step evidence plan: which searches " + + "to run, against which corpus handle, and how to cite the answer. Use " + + "it for a question to be answered in conversation (sweep_plan is for a " + + "systematic pass that writes a report). It only RETURNS instructions: " + + "it reads nothing and writes nothing.", + inputSchema: { + type: "object", + properties: { + ...SOURCE_ARG, + request: { + type: "string", + description: + "The whole question, verbatim and unsplit, punctuation intact, " + + "plus any share link. The same key=value settings as sweep_plan " + + "are recognised; anything else stays part of the question.", + }, + }, + required: ["request"], + additionalProperties: false, }, - }; + }, +]; + +// Append the corpus actually read to a result, so no answer can be silently +// about the wrong archive. Deliberately `(corpus: …)` and not `source:` — +// `- source:` already means "this video's URL" throughout the output. +// +// This lives in the dispatch wrapper, applied to every result including +// errors: a per-handler echo is one that a new tool forgets. +function withCorpus(result: ToolResult, resolved: ResolvedSource): ToolResult { + const note = [`corpus: ${resolved.handle}`, ...resolved.notes].join("; "); + const content = [...result.content]; + const last = content[content.length - 1]; + if (last && last.type === "text") { + content[content.length - 1] = { ...last, text: `${last.text}\n\n(${note})` }; + } else { + content.push({ type: "text", text: `(${note})` }); + } + return { ...result, content }; } -// Build a configured MCP server over a data source or a SourceController. The +// Build a configured MCP server over a data source or a SourceRegistry. The // core read tools work for local / remote / hub sources — only the ShardSource -// differs — and the source can be switched at runtime via use_source when a -// real controller is supplied. -export function createServer(sourceOrController: ShardSource | SourceController): Server { - const controller: SourceControllerLike = - "current" in sourceOrController - ? sourceOrController - : fixedController(sourceOrController); +// differs — and which one a call reads is decided per call by its `source` +// argument, resolved through the registry. A bare ShardSource is wrapped in a +// one-source registry, so `createServer(someSource)` still works. +export function createServer( + sourceOrRegistry: ShardSource | SourceRegistry, +): Server { + const registry = + sourceOrRegistry instanceof SourceRegistry + ? sourceOrRegistry + : SourceRegistry.forSource(sourceOrRegistry); const server = new Server( { name: "yt-dlp-transcript-mcp", version: "0.1.0" }, - { capabilities: { tools: {}, prompts: {} } }, + { capabilities: { tools: {}, prompts: {} }, cacheHints: CACHE_HINTS }, ); - server.setRequestHandler(ListToolsRequestSchema, async () => ({ tools: TOOLS })); + server.setRequestHandler("tools/list", async () => ({ tools: TOOLS })); - server.setRequestHandler(ListPromptsRequestSchema, async () => ({ + server.setRequestHandler("prompts/list", async () => ({ prompts: PROMPTS, })); - server.setRequestHandler(GetPromptRequestSchema, async (req) => { + server.setRequestHandler("prompts/get", async (req) => { if (req.params.name !== "sweep") { throw new Error(`unknown prompt: ${req.params.name}`); } return buildSweepPrompt((req.params.arguments ?? {}) as Record<string, unknown>); }); - server.setRequestHandler(CallToolRequestSchema, async (req) => { + server.setRequestHandler("tools/call", async (req) => { const name = req.params.name; const args = (req.params.arguments ?? {}) as Record<string, unknown>; - const source = controller.current; + + // Resolving the handle is the FIRST thing every call does, so a typo'd or + // unreachable corpus fails as itself rather than as an empty result set. + let resolved: ResolvedSource; try { - switch (name) { - case "list_channels": - return await handleListChannels(source); - case "search_transcripts": - return await handleSearch(source, args); - case "get_transcript": - return await handleGetTranscript(source, args); - case "get_transcripts": - return await handleGetTranscripts(source, args); - case "get_post": - return await handleGetPost(source, args); - case "get_thread": - return await handleGetThread(source, args); - case "get_video_metadata": - return await handleGetMetadata(source, args); - case "list_sources": - return await handleListSources(controller); - case "use_source": - return await handleUseSource(controller, args); - case "reset_source": - return await handleResetSource(controller); - case "open_link": - return await handleOpenLink(controller, args); - default: - return errorText(`unknown tool: ${name}`); - } + resolved = await registry.resolve(args.source); } catch (e) { - return errorText(`${name} failed: ${(e as Error).message}`); + return errorText(`${name}: ${(e as Error).message}`); + } + const source = resolved.source; + + try { + const result = await (async (): Promise<ToolResult> => { + switch (name) { + case "list_channels": + return handleListChannels(source, args); + case "search_transcripts": + return handleSearch(source, args); + case "enumerate_matches": + return handleEnumerateMatches(source, args); + case "get_transcript": + return handleGetTranscript(source, args); + case "get_transcripts": + return handleGetTranscripts(source, args); + case "get_post": + return handleGetPost(source, args); + case "get_thread": + return handleGetThread(source, args); + case "get_video_metadata": + return handleGetMetadata(source, args); + case "list_sources": + return handleListSources(registry, resolved); + case "resolve_source": + return handleResolveSource(resolved); + // Unadvertised one-release alias. It no longer switches anything — + // it resolves, and says to pass the handle per call instead. + case "use_source": + return handleLegacyUseSource(registry, args); + case "open_link": + return handleOpenLink(registry, resolved, args); + case "sweep_plan": + return handlePlan("sweep", registry, resolved, args); + case "ask_plan": + return handlePlan("ask", registry, resolved, args); + default: + return errorText(`unknown tool: ${name}`); + } + })(); + return withCorpus(result, resolved); + } catch (e) { + return withCorpus( + errorText(`${name} failed: ${(e as Error).message}`), + resolved, + ); } }); @@ -702,8 +858,11 @@ function channelLine(c: ChannelRef): string { // a group selector for a scoped search/sweep. When the source has no groups // (absent manifest, or hub mode) every channel folds under the fallback group // and the Groups line is omitted. -async function handleListChannels(source: ShardSource): Promise<ToolResult> { - const channels = await source.listChannels(); +async function handleListChannels( + source: ShardSource, + args: Record<string, unknown>, +): Promise<ToolResult> { + const channels = await source.listChannels({ refresh: args.refresh === true }); if (channels.length === 0) return text(`No channels found in ${source.label}.`); const { groups, defaultGroupId } = await source.loadGroups(); @@ -784,6 +943,7 @@ async function handleSearch( const aliasNote = describeFiredAliases(result.firedAliases); const scopeNote = describeScope(result.selection); + const postsNote = describePostsPass(result); const rangeStart = result.total === 0 ? 0 : result.offset + 1; const rangeEnd = result.offset + result.hits.length; const footer = @@ -794,15 +954,21 @@ async function handleSearch( ? "; coverage PARTIAL — scan hit the page/video cap" : "") + (scopeNote ? `; ${scopeNote}` : "") + + (postsNote ? `; ${postsNote}` : "") + (aliasNote ? `; ${aliasNote}` : "") + ")"; + // The warning goes ABOVE the hits, not only in the footer. The footer said + // "has_more: yes" on the run that motivated this and it was read straight + // past — a report then claimed 319 videos swept from a 200-hit page. + const banner = incompletePageBanner(result, "search_transcripts"); + if (result.hits.length === 0) { const head = result.total === 0 ? `No matches for "${query}".` : `No matches in this page (offset ${result.offset} is past the ${result.total} total).`; - return text(`${head}${footer}`); + return text(`${banner}${head}${footer}`); } const blocks = result.hits.map((h) => { // A POST has no timeline: no moment link, no [m:ss] stamps. Render it as a @@ -846,7 +1012,120 @@ async function handleSearch( ? "post(s)" : "result(s) (videos + posts)"; return text( - `${result.total} ${label} matching "${query}":\n\n${blocks.join("\n\n")}${footer}${baseNote}`, + `${banner}${result.total} ${label} matching "${query}":\n\n` + + `${blocks.join("\n\n")}${footer}${baseNote}`, + ); +} + +// The unmissable header for a page that is not the whole match set. Rendered +// before the hits so it cannot be scrolled past, and phrased as an instruction +// because the failure mode it exists to stop is reporting a count from one +// page. `enumerate_matches` is named because it is the cheap fix: the engine +// already materialises the full set, so enumerating is ONE scan where paging +// this query would be ceil(total/limit) full scans. +function incompletePageBanner( + result: { total: number; offset: number; limit: number; hasMore: boolean }, + tool: string, +): string { + if (!result.hasMore) return ""; + const shown = `${result.offset + 1}–${Math.min(result.offset + result.limit, result.total)}`; + return ( + `⚠ INCOMPLETE PAGE — ${result.total} total, showing ${shown}. Do NOT ` + + `report a count or "all of them" from this page. Call ` + + `\`enumerate_matches\` for the full worklist in one scan` + + (tool === "search_transcripts" ? "" : ` (this ${tool} page is a slice)`) + + `.\n\n` + ); +} + +// A note distinguishing "the posts corpus was searched and matched nothing" +// from "there is no posts corpus here at all". The second used to render as +// `scanned 0 page(s) across 0 channel(s)`, which reads like nothing ran. +function describePostsPass(result: SearchResult): string { + const p = result.postsScanned; + if (!p.requested) return ""; + if (p.channels === 0) { + return ( + "posts: no channel in scope ships a posts index — the post corpus is " + + "EMPTY here, not merely unmatched" + ); + } + return `posts: scanned ${p.pages} page(s) across ${p.channels} posting channel(s)`; +} + +// The full match set as a worklist, in a single scan. This exists because the +// paging instruction was in the sweep prompt all along and got ignored — prose +// cannot be the enforcement mechanism, so the cheap, correct thing is also the +// one-call thing. A cap being hit is stated on the FIRST line, never only in a +// footer. +async function handleEnumerateMatches( + source: ShardSource, + args: Record<string, unknown>, +): Promise<ToolResult> { + const query = String(args.query ?? "").trim(); + if (!query) return errorText("query is required"); + + const batchSizeArg = + typeof args.batch_size === "number" ? Math.floor(args.batch_size) : 8; + const batchSize = batchSizeArg >= 1 ? batchSizeArg : 8; + + const result = await searchTranscripts(source, { + query, + // The singulars stay parsed even though they are no longer advertised. + channel: typeof args.channel === "string" ? args.channel : undefined, + channels: strArray(args.channels), + group: typeof args.group === "string" ? args.group : undefined, + groups: strArray(args.groups), + regex: args.regex === true, + useAliases: args.use_aliases !== false, + maxPages: typeof args.max_pages === "number" ? args.max_pages : undefined, + contentTypes: parseContentTypes(args.content_types), + includeSnippets: false, + offset: 0, + // The whole set: searchTranscripts collects up to HARD_VIDEO_CAP and then + // slices, so asking for every collected match costs nothing extra. + limit: Number.MAX_SAFE_INTEGER, + }); + + const partial = result.truncated; + const head = partial + ? `⚠ COVERAGE PARTIAL — the scan hit its page/video cap before the corpus ` + + `was exhausted. The ${result.total} match(es) below are a SAMPLE, not ` + + `the full set: the true total is higher. Narrow the scope (channels/` + + `groups) or raise max_pages to enumerate exhaustively, and say so in any ` + + `report built on this.\n\n` + : ""; + + const rows = result.hits.map((h) => { + const kind = h.contentType === "post" ? "post" : "video"; + return ( + `- ${h.videoId} | ${kind} | ${h.channelName} | ` + + `${formatDate(h.uploadDate)} | ${h.matches} match(es) | ${h.title}` + ); + }); + + const batches = Math.ceil(result.total / batchSize); + const scopeNote = describeScope(result.selection); + const aliasNote = describeFiredAliases(result.firedAliases); + const postsNote = describePostsPass(result); + + const footer = + `\n\n(complete set: ${partial ? "NO — capped" : "yes"}; ` + + `${result.total} match(es); ${batches} batch(es) of ${batchSize}; ` + + `scanned ${result.scanned.pages} page(s) across ` + + `${result.scanned.channels} channel(s)` + + (scopeNote ? `; ${scopeNote}` : "") + + (postsNote ? `; ${postsNote}` : "") + + (aliasNote ? `; ${aliasNote}` : "") + + ")"; + + if (result.total === 0) { + return text(`${head}No matches for "${query}".${footer}`); + } + return text( + `${head}${result.total} match(es) for "${query}" — ` + + `${partial ? "a PARTIAL worklist" : "the complete worklist"}:\n\n` + + `${rows.join("\n")}${footer}`, ); } @@ -906,24 +1185,51 @@ async function handleGetTranscripts( const dropped = ids.length > CAP ? ids.length - CAP : 0; const batch = ids.slice(0, CAP); - const query = typeof args.query === "string" ? args.query.trim() : ""; + // One query or several. `queries` kills the refetch-per-quote pattern: the + // windows for every query merge in ONE pass per video (mergeSnippets), and + // each header reports a per-query count — so a query that matched nothing + // becomes visible instead of vanishing into a merged excerpt. + const QUERY_CAP = 8; + const queryList = [ + ...(typeof args.query === "string" && args.query.trim() !== "" + ? [args.query.trim()] + : []), + ...(strArray(args.queries) ?? []), + ]; + const droppedQueries = + queryList.length > QUERY_CAP ? queryList.length - QUERY_CAP : 0; + const queries = queryList.slice(0, QUERY_CAP); + const channelHint = [ ...(typeof args.channel === "string" ? [args.channel] : []), ...(strArray(args.channels) ?? []), ]; const timestamps = args.timestamps !== false; - - // Build the (alias-aware) matcher once for the whole batch when windowing. - let matcher: ReturnType<typeof buildMatcher> | null = null; - if (query) { + const types = parseContentTypes(args.content_types); + const wantVideos = types.includes("video"); + const wantPosts = types.includes("post"); + + // Build the (alias-aware) matchers once for the whole batch when windowing: + // one per query for the counts, plus their OR for the window pass. + const firedAliases: SearchAlias[] = []; + let perQuery: { query: string; match: (t: string) => boolean }[] = []; + let matcher: { match: (t: string) => boolean } | null = null; + if (queries.length > 0) { const useAliases = args.use_aliases !== false && args.regex !== true; const aliases = useAliases ? await source.loadAliases() : []; - matcher = buildMatcher({ - query, - regex: args.regex === true, - useAliases, - aliases, + perQuery = queries.map((q) => { + const built = buildMatcher({ + query: q, + regex: args.regex === true, + useAliases, + aliases, + }); + for (const a of built.firedAliases) { + if (!firedAliases.some((x) => x.id === a.id)) firedAliases.push(a); + } + return { query: q, match: built.match }; }); + matcher = { match: (t: string) => perQuery.some((m) => m.match(t)) }; } const before = typeof args.before === "number" ? args.before : 30; @@ -937,8 +1243,16 @@ async function handleGetTranscripts( const blocks: string[] = []; const missing: string[] = []; for (const id of batch) { - const found = await findVideo(source, id, channelHint); + const found = wantVideos ? await findVideo(source, id, channelHint) : null; if (!found) { + // content_types was declared on this tool and never read, so a batch of + // post ids silently came back "not found". Fall through to the post + // corpus for anything the transcript lookup didn't resolve. + const post = wantPosts ? await findPost(source, id, channelHint) : null; + if (post) { + blocks.push(postToMarkdown(post.post)); + continue; + } missing.push(id); continue; } @@ -974,10 +1288,24 @@ async function handleGetTranscripts( stamp, maxLines, }); + // Per-query counts, so a query contributing nothing to this video is + // visible rather than absorbed into the merged window — that visibility + // is the real diagnostic value of batching several queries. + const cues = record.cues ?? []; + const counts = perQuery.map((m) => ({ + query: m.query, + n: cues.reduce((acc, c) => acc + (m.match(c.text) ? 1 : 0), 0), + })); + const countNote = + counts.length > 1 + ? counts.map((c) => `"${c.query}": ${c.n}`).join(", ") + : `${matchCount} matching line(s)`; const body = matchCount === 0 - ? "_(no lines matched the query in this transcript)_" - : `_(${matchCount} matching line(s), windowed)_\n${lines.join("\n")}`; + ? `_(no lines matched ${ + counts.length > 1 ? "any query" : "the query" + } in this transcript${counts.length > 1 ? ` — ${countNote}` : ""})_` + : `_(${countNote}, windowed)_\n${lines.join("\n")}`; blocks.push(`${head}\n\n${body}`); } else { const md = transcriptToMarkdown( @@ -1000,12 +1328,14 @@ async function handleGetTranscripts( } const notes: string[] = []; - if (query) { - const aliasNote = describeFiredAliases(matcher?.firedAliases ?? []); - if (aliasNote) notes.push(aliasNote); - } + if (queries.length > 1) notes.push(`queries: ${queries.map((q) => `"${q}"`).join(", ")}`); + const aliasNote = describeFiredAliases(firedAliases); + if (aliasNote) notes.push(aliasNote); if (missing.length > 0) notes.push(`not found: ${missing.join(", ")}`); if (dropped > 0) notes.push(`${dropped} extra id(s) beyond the 20-cap dropped`); + if (droppedQueries > 0) { + notes.push(`${droppedQueries} quer(ies) beyond the ${QUERY_CAP}-cap dropped`); + } if (base) notes.push(BASE_EXPANSION_NOTE); const footer = notes.length > 0 ? `\n\n(${notes.join("; ")})` : ""; @@ -1141,10 +1471,10 @@ async function handleGetMetadata( ); } -// ─── Source switching: list_sources / use_source / reset_source ─── +// ─── Source discovery: list_sources / resolve_source ─── // A one-line human description of a source spec's target (kind + where it -// points), for the list_sources / use_source reports. +// points), for the list_sources / resolve_source reports. function describeSpec(spec: SourceSpec | undefined): string { if (!spec) return "(pinned source)"; switch (spec.kind) { @@ -1159,194 +1489,145 @@ function describeSpec(spec: SourceSpec | undefined): string { } } -// A cheap post-switch summary: channel count, and group count when the new -// source actually ships channel groups. Tolerant of a source that can't be -// reached (reports the switch anyway). -async function sourceSummary(source: ShardSource): Promise<string> { - try { - const channels = await source.listChannels(); - const { groups } = await source.loadGroups(); - const groupNote = - groups.length > 0 ? `, ${groups.length} group(s)` : ""; - return `${channels.length} channel(s)${groupNote}`; - } catch (e) { - return `(could not read the new source: ${(e as Error).message})`; - } -} - -// Report the active source and, when a hub is in play, its member sites so a -// human can pick one (or a subset) to switch to. +// Report the corpus this call read (the default when `source` was omitted) and, +// when a hub is in play, its member sites as ready-to-paste handles. There is +// no active source to report — only a default and whatever a caller names. async function handleListSources( - controller: SourceControllerLike, + registry: SourceRegistry, + resolved: ResolvedSource, ): Promise<ToolResult> { + const isDefault = resolved.handle === registry.defaultHandle; const lines: string[] = [ - `Active source: ${controller.current.label}`, - ` target: ${describeSpec(controller.activeSpec)}`, + `Corpus read by this call: ${resolved.label}`, + ` handle: ${resolved.handle}` + (isDefault ? " (the default)" : ""), + ` target: ${describeSpec(resolved.spec)}`, ]; + if (!isDefault) { + lines.push( + ` default (used when a call omits source): ${registry.defaultHandle}`, + ); + } + const hubUrl = + resolved.spec.kind === "hub" ? resolved.spec.url : registry.hubUrl(); let sites: HubSite[] = []; - try { - sites = await controller.listHubSites(); - } catch (e) { - lines.push(`\n(could not list hub members: ${(e as Error).message})`); + if (hubUrl) { + try { + sites = await registry.listHubSites(hubUrl); + } catch (e) { + lines.push(`\n(could not list hub members: ${(e as Error).message})`); + } } if (sites.length > 0) { - // Which siteIds are in the current subset (only when the active source is a - // hub scoped to a subset); otherwise every member is "in scope". - const spec = controller.activeSpec; + // When the resolved corpus is a hub subset, mark who is actually in it. + const spec = resolved.spec; const subset = - spec && spec.kind === "hub" && spec.sites && spec.sites.length > 0 + spec.kind === "hub" && spec.sites && spec.sites.length > 0 ? new Set(spec.sites) : undefined; - const rows = sites.map((s) => { - const inScope = !subset || subset.has(s.siteId); - const mark = subset ? (inScope ? "✓ " : " ") : ""; - return ` - ${mark}${s.siteId} · ${s.title} · ${s.url}`; + const rows = sites.map((site) => { + const mark = subset ? (subset.has(site.siteId) ? "✓ " : " ") : ""; + return ( + ` - ${mark}${site.siteId} · ${site.title} · ${site.url}\n` + + ` source: remote:${site.url.replace(/\/+$/, "")}` + ); }); lines.push( `\nHub member sites (${sites.length})` + - (subset ? ` — ✓ = in the current subset` : "") + + (subset ? ` — ✓ = in the subset this call read` : "") + `:\n${rows.join("\n")}`, ); + lines.push( + `\nA subset of them: source:"hub:${hubUrl}#${sites + .slice(0, 2) + .map((x) => x.siteId) + .join(",")}".`, + ); } lines.push( - `\nSwitch with use_source: site:"<siteId|title>" (one member), ` + - `sites:["<a>","<b>"] (a subset), or an arbitrary remote:"<url>" / ` + - `local:"<dir>" / hub:"<url>". reset_source returns to the startup source. ` + - `Read-only throughout.`, + `\nPass a corpus per call, in each tool's own \`source\` argument — ` + + `'default', 'local:<dir>', 'remote:<url>', 'hub:<url>', or ` + + `'hub:<url>#<siteA,siteB>'. Nothing is switched or remembered: a call ` + + `without \`source\` always reads the default. resolve_source turns a ` + + `loose reference into the handle to pass. Read-only throughout.`, ); return text(lines.join("\n")); } -// Switch the active source. Exactly one target family is accepted; a hub member -// token (site/sites) is resolved against the hub's member list. -async function handleUseSource( - controller: SourceControllerLike, - args: Record<string, unknown>, +// Resolve a reference to its canonical handle and verify the corpus can be +// read. Mutates nothing — the point is to hand back a string to pass as +// `source`, not to select anything. +async function handleResolveSource( + resolved: ResolvedSource, ): Promise<ToolResult> { - const site = typeof args.site === "string" ? args.site.trim() : ""; - const sites = strArray(args.sites); - const remote = typeof args.remote === "string" ? args.remote.trim() : ""; - const local = typeof args.local === "string" ? args.local.trim() : ""; - const hub = typeof args.hub === "string" ? args.hub.trim() : ""; - - const families = [ - site ? "site" : "", - sites ? "sites" : "", - remote ? "remote" : "", - local ? "local" : "", - hub ? "hub" : "", - ].filter((f) => f !== ""); - if (families.length === 0) { - return errorText( - "use_source needs exactly one target: site, sites, remote, local, or hub.", + const lines = [ + `Handle: ${resolved.handle}`, + ` label: ${resolved.label}`, + ` target: ${describeSpec(resolved.spec)}`, + ]; + for (const n of resolved.notes) lines.push(` note: ${n}`); + + try { + const channels = await resolved.source.listChannels(); + const { groups } = await resolved.source.loadGroups(); + lines.push( + ` reachable: yes — ${channels.length} channel(s)` + + (groups.length > 0 ? `, ${groups.length} group(s)` : ""), ); - } - if (families.length > 1) { + } catch (e) { return errorText( - `use_source takes exactly one target; got ${families.join(", ")}.`, + `${lines.join("\n")}\n reachable: NO — ${(e as Error).message}`, ); } - let spec: SourceSpec; - if (remote) { - spec = { kind: "remote", url: remote }; - } else if (local) { - spec = { kind: "local", dir: local }; - } else if (hub) { - spec = { kind: "hub", url: hub }; - } else { - // site / sites — resolve against the hub member list. - const hubUrl = controller.hubUrl(); - if (!hubUrl) { - return errorText( - "no hub context to resolve a site token against — start the server " + - "against a hub, or first switch with hub:\"<url>\" (or use remote/local).", - ); - } - let members: HubSite[]; - try { - members = await controller.listHubSites(); - } catch (e) { - return errorText(`could not read hub members: ${(e as Error).message}`); - } - const resolve = (token: string): HubSite | undefined => { - const t = token.toLowerCase(); - return members.find( - (m) => m.siteId.toLowerCase() === t || m.title.toLowerCase() === t, - ); - }; - - if (site) { - const member = resolve(site); - if (!member) { - return errorText( - `unknown hub site "${site}". Known: ` + - members.map((m) => m.siteId).join(", "), - ); - } - // A single member becomes a plain remote source → full group/alias support. - spec = { kind: "remote", url: member.url }; - } else { - const resolved: string[] = []; - const unknown: string[] = []; - for (const token of sites!) { - const member = resolve(token); - if (member) resolved.push(member.siteId); - else unknown.push(token); - } - if (resolved.length === 0) { - return errorText( - `no known hub sites in [${sites!.join(", ")}]. Known: ` + - members.map((m) => m.siteId).join(", "), - ); - } - spec = { kind: "hub", url: hubUrl, sites: resolved }; - if (unknown.length > 0) { - // Switch to the resolvable subset but surface the typos. - try { - await controller.switchTo(spec); - } catch (e) { - return errorText(`use_source failed: ${(e as Error).message}`); - } - const summary = await sourceSummary(controller.current); - return text( - `Switched to ${controller.current.label} — ${summary}.\n` + - ` target: ${describeSpec(spec)}\n` + - `Unknown site token(s) skipped: ${unknown.join(", ")}.`, - ); - } - } - } - - try { - await controller.switchTo(spec); - } catch (e) { - return errorText(`use_source failed: ${(e as Error).message}`); - } - const summary = await sourceSummary(controller.current); - return text( - `Switched to ${controller.current.label} — ${summary}.\n` + - ` target: ${describeSpec(spec)}`, + lines.push( + ``, + `Nothing was switched. Pass source:"${resolved.handle}" on each call ` + + `that should read this corpus.`, ); + return text(lines.join("\n")); } -// Return to the startup source and clear the persisted selection. -async function handleResetSource( - controller: SourceControllerLike, +// The old switching tool, kept one release as an unadvertised alias so a habit +// (or a stored transcript) doesn't hard-fail. It resolves its target and hands +// back the handle; it cannot switch anything, because there is nothing to +// switch. +async function handleLegacyUseSource( + registry: SourceRegistry, + args: Record<string, unknown>, ): Promise<ToolResult> { + const sites = strArray(args.sites); + const token = + (typeof args.site === "string" && args.site.trim()) || + (typeof args.remote === "string" && `remote:${args.remote.trim()}`) || + (typeof args.local === "string" && `local:${args.local.trim()}`) || + (typeof args.hub === "string" && `hub:${args.hub.trim()}`) || + (sites ? `sites:${sites.join(",")}` : ""); + if (!token) { + return errorText( + "use_source is gone: there is no active source to switch. Pass the " + + "corpus per call in each tool's `source` argument instead, and use " + + "resolve_source to turn a URL or site name into a handle.", + ); + } + + let target: ResolvedSource; try { - await controller.reset(); + target = await registry.resolve(token); } catch (e) { - return errorText(`reset_source failed: ${(e as Error).message}`); + return errorText(`use_source: ${(e as Error).message}`); } - const summary = await sourceSummary(controller.current); return text( - `Reset to the startup source ${controller.current.label} — ${summary}.\n` + - ` target: ${describeSpec(controller.activeSpec)}`, + `use_source no longer switches anything — this server holds no active ` + + `source (a persisted one silently sent every call to the wrong corpus, ` + + `so it was removed).\n\n` + + `That target's handle is: ${target.handle}\n` + + ` target: ${describeSpec(target.spec)}\n` + + (target.notes.length > 0 ? ` note: ${target.notes.join("; ")}\n` : "") + + `\nPass source:"${target.handle}" on each call that should read it.`, ); } @@ -1376,6 +1657,11 @@ function parseOverrides(raw: unknown): LinkOverrides | undefined { queryScope: scope === "transcripts" || scope === "chat" || + // "posts" was declared in the tool schema and dropped here, so a caller + // asking to re-target a link at the post corpus silently got transcripts. + // Everything downstream (LayerScope, runSearchSpec's posts pass) already + // supported it. + scope === "posts" || scope === "metadata" || scope === "description" || scope === "tags" @@ -1386,28 +1672,6 @@ function parseOverrides(raw: unknown): LinkOverrides | undefined { type OriginProbe = { kind: "hub" | "remote"; siteCount?: number; channelCount?: number }; -// Probe an origin's corpus.json to classify it as a federated hub (has a -// `sites[]` array) or a single site (has `channels[]`). Transient — no source is -// committed. Defaults to single-site when the probe can't be read. -async function probeOrigin(origin: string): Promise<OriginProbe> { - try { - const res = await fetch(`${origin}/corpus.json`); - if (res.ok) { - const j = (await res.json()) as { - sites?: unknown[]; - channels?: unknown[]; - }; - if (Array.isArray(j.sites)) return { kind: "hub", siteCount: j.sites.length }; - if (Array.isArray(j.channels)) { - return { kind: "remote", channelCount: j.channels.length }; - } - } - } catch { - // unreachable — fall through to the single-site default - } - return { kind: "remote" }; -} - type ChannelScope = { channels: ChannelRef[]; all: boolean; @@ -1441,20 +1705,22 @@ async function resolveLinkChannels( return { channels, all: false, matched, unknown }; } -// Render the human-readable preview/echo plan for a decoded link. +// Render the human-readable plan for a decoded link. function buildLinkPlan( decoded: DecodedLink, probe: OriginProbe, scope: ChannelScope, - applied: boolean, + handle: string, + searched: boolean, ): string { const lines: string[] = []; - lines.push(applied ? "## Applied share link" : "## Share-link preview"); + lines.push(searched ? "## Share link" : "## Share-link plan (dry run)"); const kindNote = probe.kind === "hub" ? `hub (${probe.siteCount ?? "?"} member site(s))` : `single site${probe.channelCount != null ? ` (${probe.channelCount} channel(s))` : ""}`; lines.push(`- Source: ${decoded.origin} — ${kindNote}`); + lines.push(`- Corpus handle: \`${handle}\` — pass this as \`source\` on the follow-up calls`); lines.push(`- Query: ${renderQueryTree(decoded.tree)} _(from ${decoded.querySource})_`); if (scope.all) { @@ -1483,11 +1749,11 @@ function buildLinkPlan( for (const w of decoded.warnings) lines.push(`- Note: ${w}`); - if (!applied) { + if (!searched) { lines.push( ``, - `No search run yet. Call again with apply:true to switch source and ` + - `search, or pass overrides to adjust first.`, + `Dry run — no search was performed. Re-call without dry_run to search, ` + + `or pass overrides to adjust the plan first.`, ); } return lines.join("\n"); @@ -1527,7 +1793,8 @@ function renderSpecResults(source: ShardSource, hits: SpecHit[]): string { } async function handleOpenLink( - controller: SourceControllerLike, + registry: SourceRegistry, + callerCorpus: ResolvedSource, args: Record<string, unknown>, ): Promise<ToolResult> { const linkStr = typeof args.link === "string" ? args.link.trim() : ""; @@ -1541,47 +1808,43 @@ async function handleOpenLink( } decoded = applyLinkOverrides(decoded, parseOverrides(args.overrides)); - const apply = args.apply === true; - const probe = await probeOrigin(decoded.origin); - - if (!apply) { - // Preview: resolve channels against a transient source; do not commit. - const preview: ShardSource = - probe.kind === "hub" - ? new HubSource(decoded.origin) - : new RemoteSource(decoded.origin); - let scope: ChannelScope; + // The link's ORIGIN decides the corpus, not the caller's `source` — a share + // link is a pointer to a specific deployment. An explicit `source` still + // wins, for the case where the same corpus is reachable at another handle. + const explicitSource = + typeof args.source === "string" && args.source.trim() !== ""; + let target: ResolvedSource; + if (explicitSource) { + target = callerCorpus; + } else { try { - scope = await resolveLinkChannels(preview, decoded); + target = await registry.resolve(decoded.origin); } catch (e) { return errorText( - `could not read the corpus at ${decoded.origin}: ${(e as Error).message}`, + `could not resolve the link's origin ${decoded.origin}: ${(e as Error).message}`, ); } - return text(buildLinkPlan(decoded, probe, scope, false)); } - - // Apply: commit the source (reusing the controller), then search. - const spec: SourceSpec = - probe.kind === "hub" - ? { kind: "hub", url: decoded.origin } - : { kind: "remote", url: decoded.origin }; - try { - await controller.switchTo(spec); - } catch (e) { - return errorText(`open_link could not switch source: ${(e as Error).message}`); - } - const source = controller.current; + const source = target.source; + const probe: OriginProbe = + target.spec.kind === "hub" ? { kind: "hub" } : { kind: "remote" }; let scope: ChannelScope; try { scope = await resolveLinkChannels(source, decoded); } catch (e) { return errorText( - `switched to ${source.label} but could not read its corpus: ${(e as Error).message}`, + `could not read the corpus at ${target.handle}: ${(e as Error).message}`, ); } - const plan = buildLinkPlan(decoded, probe, scope, true); + probe.channelCount = scope.all ? scope.channels.length : undefined; + + // dry_run stops here: the plan, with no search. The default is to do the + // whole job in this one call — the old two-phase apply:true cost a round + // trip and, worse, mutated a global. + const dryRun = args.dry_run === true; + const plan = buildLinkPlan(decoded, probe, scope, target.handle, !dryRun); + if (dryRun) return text(plan); const limit = typeof args.limit === "number" ? args.limit : 20; const offset = typeof args.offset === "number" ? args.offset : 0; @@ -1614,7 +1877,118 @@ async function handleOpenLink( : `No matches in this page (offset ${result.offset} is past the ${result.total} total).` : `${result.total} video(s) matched:\n\n${renderSpecResults(source, result.hits)}`; - return text(`${plan}\n\n---\n\n${body}${footer}`); + const banner = incompletePageBanner(result, "open_link"); + return text(`${plan}\n\n---\n\n${banner}${body}${footer}`); +} + +// ─── sweep_plan / ask_plan: a pasted request in, a resolved plan out ─── + +// Pre-resolve the facts the model would otherwise burn round-trips on: +// canonicalise the corpus, validate the channel/group tokens against the live +// channel list, and offer the group roster when no scope was given. Tolerant — +// an unreadable corpus just means fewer resolved facts, not a failed plan. +async function buildPlanContext( + resolved: ResolvedSource, + req: PromptRequest, +): Promise<PlanContext> { + const ctx: PlanContext = { corpus: resolved.handle, notes: [...resolved.notes] }; + let channels: ChannelRef[]; + let groups: ChannelGroup[]; + try { + channels = await resolved.source.listChannels(); + groups = (await resolved.source.loadGroups()).groups; + } catch (e) { + ctx.notes = [ + ...(ctx.notes ?? []), + `could not read ${resolved.handle} to validate the scope: ${(e as Error).message}`, + ]; + return ctx; + } + + const known = (token: string): boolean => { + const want = token.toLowerCase(); + return channels.some( + (c) => + c.slug.toLowerCase() === want || + c.key.toLowerCase() === want || + c.name.toLowerCase() === want, + ); + }; + ctx.knownChannels = req.channels.filter(known); + ctx.unknownChannels = req.channels.filter((c) => !known(c)); + + const knownGroup = (token: string): boolean => { + const want = token.toLowerCase(); + return groups.some( + (g) => + g.id.toLowerCase() === want || + (g.name.trim() !== "" && g.name.toLowerCase() === want), + ); + }; + ctx.knownGroups = req.groups.filter(knownGroup); + ctx.unknownGroups = req.groups.filter((g) => !knownGroup(g)); + + if (req.channels.length === 0 && req.groups.length === 0 && groups.length > 0) { + const counts = new Map<string, number>(); + for (const c of channels) { + const gid = resolveChannelGroupId( + c.groupId, + groups, + (await resolved.source.loadGroups()).defaultGroupId, + ); + counts.set(gid, (counts.get(gid) ?? 0) + 1); + } + ctx.availableGroups = sortGroups(groups).map((g) => ({ + id: g.id, + name: g.name || g.id, + channels: counts.get(g.id) ?? 0, + })); + } + ctx.notes = [ + ...(ctx.notes ?? []), + `${channels.length} channel(s) in ${resolved.handle}`, + ]; + return ctx; +} + +// Both plan tools. The request arrives as ONE structured JSON string, so a URL +// with `?a=b&c=d` and a full sentence of punctuation survive intact — which is +// the whole reason these are tools and not prompt arguments. +async function handlePlan( + kind: "sweep" | "ask", + registry: SourceRegistry, + callerCorpus: ResolvedSource, + args: Record<string, unknown>, +): Promise<ToolResult> { + const raw = typeof args.request === "string" ? args.request : ""; + if (raw.trim() === "") { + return errorText( + `${kind}_plan needs a request: what to ${kind === "sweep" ? "sweep for" : "ask"}, ` + + `in plain English, optionally with a share link.`, + ); + } + const req = parsePromptRequest(raw); + + // A `source=` inside the request text wins over the tool's own `source` + // argument, since it is the more specific statement of intent. + let corpus = callerCorpus; + if (req.source) { + try { + corpus = await registry.resolve(req.source); + } catch (e) { + req.warnings.push( + `source=${req.source} could not be resolved (${(e as Error).message}) — ` + + `using ${callerCorpus.handle}.`, + ); + } + } + + const ctx = await buildPlanContext(corpus, req); + return text( + kind === "sweep" + ? buildSweepInstructions(req, ctx) + : buildAskInstructions(req, ctx), + ); } // ─── Prompts: the first-class `sweep` entry point ─── @@ -1623,15 +1997,18 @@ async function handleOpenLink( // windowed transcripts, and fold findings into a markdown report it maintains // with its own Write/Edit tools. The MCP stays read-only; the report is a file // in Claude's cwd. Mirrors the browser's accumulationSystemPrompt discipline. -const PROMPTS = [ +const PROMPTS: Prompt[] = [ { name: "sweep", description: "Sweep a query across the corpus (or a chosen group/channels): enumerate " + "every matching video, batch-read the transcripts, and fold cited, " + - "cross-referenced findings into a markdown report — driven by Claude Code " + - "on plan usage, no API key. With no scope arg it lists the channel groups " + - "and asks you to pick a group/channels (or confirm 'all') before sweeping.", + "cross-referenced findings into a markdown report — on plan usage, no " + + "API key. With no scope arg it lists the channel groups and asks you to " + + "pick a group/channels (or confirm 'all') before sweeping. This form is " + + "for clients that give each argument its own field (Claude Desktop, " + + "Cursor); in Claude Code use the `/sweep` command, which passes the " + + "whole line to the sweep_plan tool instead of word-splitting it.", arguments: [ { name: "query", @@ -1694,181 +2071,33 @@ const PROMPTS = [ }, ]; -function argStr(args: Record<string, unknown>, key: string): string | undefined { - const v = args[key]; - return typeof v === "string" && v.trim() !== "" ? v.trim() : undefined; -} - -function buildSweepPrompt(args: Record<string, unknown>) { - const query = argStr(args, "query"); - const link = argStr(args, "link"); - if (!query && !link) { +// The MCP prompt form. Its declared arguments go through the SAME parser and +// validators as the tools, so a form-based client (Claude Desktop, Cursor — +// where each argument gets its own field and multi-word values are fine) gets +// the identical instructions, and a bad value gets a named warning instead of +// being rendered into the text. +function buildSweepPrompt(args: Record<string, unknown>): { + description: string; + messages: { role: "user"; content: { type: "text"; text: string } }[]; +} { + const req = requestFromArguments(args); + if (!req.query && !req.link) { throw new Error("sweep requires a query argument (or a link)"); } - const channel = argStr(args, "channel"); - const group = argStr(args, "group"); - const channelsRaw = argStr(args, "channels"); - const channels = channelsRaw - ? channelsRaw.split(",").map((s) => s.trim()).filter((s) => s !== "") - : []; - const directive = argStr(args, "directive") ?? "key claims & contradictions"; - const batchSize = argStr(args, "batch_size") ?? "8"; - const parseModel = argStr(args, "parse_model") ?? "haiku"; - const reportPath = argStr(args, "report_path") ?? "./sweep-report.md"; - const subject = query ? `"${query}"` : "the share link's search"; - - // Human-readable scope clauses + the literal search_transcripts scope args to - // pass. When none is given (and there's no link), the sweep must pick-first - // (list + ask) rather than silently scanning the whole corpus. - const scopeClauses: string[] = []; - if (channel) scopeClauses.push(`channel "${channel}"`); - if (channels.length > 0) - scopeClauses.push(`channels [${channels.map((c) => `"${c}"`).join(", ")}]`); - if (group) scopeClauses.push(`group "${group}"`); - const hasScope = scopeClauses.length > 0; - const scopeArgsText = scopeClauses.join(" and "); - - const introScope = link - ? ` seeded from a share link (its source, query tree, and filters — see step 1)` - : hasScope - ? ` scoped to ${scopeArgsText}` - : " over a scope you will confirm with me first (see step 1)"; - - const steps: string[] = []; - - // ── Discovery / enumeration differs for a link-seeded sweep vs a query one ── - if (link) { - steps.push( - `**Decode & confirm the link.** Call \`open_link\` with ` + - `link="${link}" (preview mode — no apply). It returns a plan: the ` + - `resolved source (origin, hub or single-site), the query tree, every ` + - `active filter, the channel scope validated against the corpus, and any ` + - `ignored bits (e.g. the vestigial \`fk\` tracks). **Show me the plan and ` + - `confirm it captures what I want.** If I ask for a change ("drop the ` + - `availability filter", "only channel X", "search Y instead"), re-call ` + - `\`open_link\` with the matching \`overrides\` until the plan is right.`, - ); - steps.push( - `**Apply & enumerate.** Call \`open_link\` again with the confirmed ` + - `arguments plus apply:true — this switches the active source to the ` + - `link's origin and runs the search (the full query tree + filters, at ` + - `fidelity). Page it with a rising \`offset\` (offset += limit) until ` + - `\`has_more\` is no to collect the whole worklist of video ids; note the ` + - `\`total\`. Results already carry linked \`[mm:ss](url)\` timestamps and ` + - `the applied plan is echoed at the top — record the source, query, and ` + - `filters in the report. If coverage is PARTIAL, say so.`, - ); - } else { - if (!hasScope) { - steps.push( - `**Choose the scope first — do NOT default to the whole corpus.** No ` + - `channel/channels/group was supplied. Call \`list_channels\`, present ` + - `the channel groups and their channels to me, and ask which group(s) ` + - `or channel(s) to sweep — or to confirm **all** for the whole corpus. ` + - `Wait for my choice before enumerating anything. Only sweep everything ` + - `if I explicitly choose "all". Use my choice as the ` + - `\`channel\`/\`channels\`/\`group\` scope in every ` + - `\`search_transcripts\` call below.`, - ); - } - steps.push( - `**Search.** Call \`search_transcripts\` with query "${query}"` + - (hasScope - ? ` and ${scopeArgsText}` - : ` and the scope I chose in step 1`) + - `. Curated aliases auto-expand the query — the footer reports which ` + - `fired (e.g. mis-transcribed spellings) and names the resolved scope ` + - `(and warns about any channel/group token that matched nothing — fix a ` + - `typo before continuing). Treat the *expanded* match set as your target ` + - `and mention the expansion and the scope in the report.`, - ); - steps.push( - `**Enumerate the full worklist.** Page the complete set with ` + - `\`include_snippets: false\` and a rising \`offset\` (offset += limit) ` + - `until \`has_more\` is false — this gives you every id/title/channel/date ` + - `cheaply. Note the \`total\`. If the footer says coverage is PARTIAL ` + - `(page/video cap), say so in the report — the sweep is then a sample, ` + - `not exhaustive.`, - ); - } - - steps.push( - `**Plan.** With N total matches and a batch size of ${batchSize}, that is ` + - `\`ceil(N / ${batchSize})\` batches. State the plan (N and the batch ` + - `count) before you start.`, - ); - - // ── Feature 2: dumb-extractor-per-batch map-reduce on the cheapest model — - // the extractor only quotes verbatim base-form excerpts, so the heavy - // transcript text never reaches the orchestrator (and never costs the big - // model). Feature 1: linked citations, expanded from moment_base + seconds. - steps.push( - `**Per batch (map-reduce), for each group of up to ${batchSize} video ids:**\n` + - ` - **Spawn a subagent as a DUMB EXTRACTOR on the cheapest model** — ` + - `use the Task tool and request model "${parseModel}" (its model ` + - `parameter, or a "${parseModel}"-backed agent type); if no model ` + - `override is available, spawn it anyway on the default. Give it exactly ` + - `this job: call \`get_transcripts\` with the batch's ids, the query, and ` + - `\`link_style: "base"\`, then return ONLY the lines relevant to the ` + - `directive, VERBATIM — do NOT analyze, summarize, or rephrase anything. ` + - `Group the kept lines per video as \`### <title>\` + that video's ` + - `\`moment_base:\` line copied exactly (or its \`source:\` line when ` + - `there is no moment_base) + the kept lines in their \`[mm:ss|seconds]\` ` + - `form. Hard budget: at most 40 lines (~600 words) per batch — if more ` + - `match, keep the strongest and end with \`(+N more matching lines)\`. ` + - `Return nothing else — the raw transcript bulk stays inside the ` + - `subagent and never enters your context.\n` + - ` - **Merge — you (the orchestrator) do ALL the synthesis.** Cross-` + - `reference the returned fragment against the report so far and upsert ` + - `findings — claims, and contradictions with earlier claims — into ` + - `well-titled \`## sections\` of \`${reportPath}\` (Write/Edit). Cite ` + - `every VIDEO finding as **\`[title @ mm:ss](<moment url>)\`**, expanding each ` + - `kept \`[mm:ss|seconds]\` stamp by appending the integer after the ` + - `\`|\` to that video's moment_base (full link = ` + - `\`<moment_base><seconds>\`; no moment_base → link the \`source:\` URL ` + - `instead). A POST finding has no timestamp — cite it as ` + - `**\`[post by <author>, <date>](<source url>)\`** instead, never with ` + - `\`@ mm:ss\`. Then discard the fragment. Batches are independent, so you ` + - `may dispatch several subagents in parallel.\n` + - ` - **Fallback:** if no subagent/Task tool is available, do the batch ` + - `inline — call \`get_transcripts\` with \`link_style: "base"\` yourself, ` + - `fold the expanded cited findings into the report, then **drop the raw ` + - `excerpt text** before moving on (don't carry it forward).`, - ); - - steps.push( - `**Finish.** Repeat to the end of the worklist, then write a short summary ` + - `section (the scope swept, how many videos covered, headline findings, any ` + - `partial-coverage caveat) and tell me the report path. Keep every citation ` + - `a clickable link — \`[title @ mm:ss](url)\` for a video, ` + - `\`[post by <author>, <date>](url)\` for a post.`, - ); - - const numbered = steps - .map((s, i) => `${i + 1}. ${s}`) - .join("\n\n"); - - const text = - `Run a **corpus sweep** for ${subject}${introScope}, extracting ` + - `**${directive}**, and maintain a running markdown report at ` + - `\`${reportPath}\`. You are the sweep engine — work through the whole match ` + - `set methodically, using the transcript MCP tools for evidence and your own ` + - `Write/Edit tools for the report. The MCP is read-only; never try to change ` + - `the archive. The corpus holds video transcripts AND archived social posts. ` + - `**Cite every video finding as a clickable ` + - `\`[title @ mm:ss](<moment url>)\` link** (build each moment URL by ` + - `appending the cited integer seconds to that video's \`moment_base\` from ` + - `the tool output); **cite every post finding as ` + - `\`[post by <author>, <date>](<source url>)\`** — posts have no timeline, ` + - `so they never take a \`@ mm:ss\`.\n\n` + - `Follow these steps:\n\n${numbered}`; - + // The prompt form has no live corpus to resolve against — it is rendered + // before any tool call — so the plan names the server default and the model + // resolves the rest as step 1. + const ctx: PlanContext = { corpus: "default" }; + const subject = req.query ? `"${req.query}"` : "the share link's search"; return { - description: `Corpus sweep for ${subject} → ${reportPath}`, + description: `Corpus sweep for ${subject} → ${req.reportPath ?? DEFAULT_REPORT_PATH}`, messages: [ { role: "user" as const, - content: { type: "text" as const, text }, + content: { + type: "text" as const, + text: buildSweepInstructions(req, ctx), + }, }, ], }; diff --git a/mcp/src/shareLink.ts b/mcp/src/shareLink.ts @@ -147,7 +147,13 @@ export type LinkOverrides = { // Replace the whole query with a single leaf (query + optional regex/scope). query?: string; regex?: boolean; - queryScope?: "transcripts" | "chat" | "metadata" | "description" | "tags"; + queryScope?: + | "transcripts" + | "chat" + | "posts" + | "metadata" + | "description" + | "tags"; }; const KEEP_ALL_FILTERS: SearchFilters = { diff --git a/mcp/src/source.ts b/mcp/src/source.ts @@ -123,7 +123,12 @@ function parseGroupsManifest(raw: unknown): ChannelGroups { // a page of full transcript records. export interface ShardSource { readonly label: string; - listChannels(): Promise<ChannelRef[]>; + // The channel list, memoised per source instance. searchTranscripts, + // findVideo and findPost all call it, so a 20-id get_transcripts batch used + // to cost 20 corpus.json fetches over HTTP. The memo is the promise, so + // concurrent callers share one fetch. Pass `refresh` to drop it and re-read — + // an explicit staleness escape hatch, deliberately not a TTL. + listChannels(opts?: { refresh?: boolean }): Promise<ChannelRef[]>; transcriptsManifest(ch: ChannelRef): Promise<ChannelTranscriptsManifest>; transcriptPage(ch: ChannelRef, page: number): Promise<TranscriptDetail[]>; // The site's shipped curated search aliases (the same /search-aliases.json the @@ -301,7 +306,18 @@ export class LocalSource implements ShardSource { return this.groups; } - async listChannels(): Promise<ChannelRef[]> { + private channelList?: Promise<ChannelRef[]>; + + listChannels(opts: { refresh?: boolean } = {}): Promise<ChannelRef[]> { + if (opts.refresh) this.channelList = undefined; + this.channelList ??= this.readChannels().catch((e: unknown) => { + this.channelList = undefined; // don't memoise a failure + throw e; + }); + return this.channelList; + } + + private async readChannels(): Promise<ChannelRef[]> { try { const raw = await readFile(path.join(this.dir, "corpus.json"), "utf8"); const corpus = JSON.parse(raw) as SiteCorpusJson; @@ -455,7 +471,18 @@ export class RemoteSource implements ShardSource { return (await res.json()) as T; } - async listChannels(): Promise<ChannelRef[]> { + private channelList?: Promise<ChannelRef[]>; + + listChannels(opts: { refresh?: boolean } = {}): Promise<ChannelRef[]> { + if (opts.refresh) this.channelList = undefined; + this.channelList ??= this.readChannels().catch((e: unknown) => { + this.channelList = undefined; // don't memoise a failure + throw e; + }); + return this.channelList; + } + + private async readChannels(): Promise<ChannelRef[]> { const corpus = await this.getJson<SiteCorpusJson>("/corpus.json"); return (corpus.channels ?? []).map((c) => ({ key: c.slug, @@ -544,7 +571,21 @@ export class HubSource implements ShardSource { return m; } - async listChannels(): Promise<ChannelRef[]> { + private channelList?: Promise<ChannelRef[]>; + + listChannels(opts: { refresh?: boolean } = {}): Promise<ChannelRef[]> { + if (opts.refresh) this.channelList = undefined; + this.channelList ??= this.readChannels().catch((e: unknown) => { + this.channelList = undefined; // don't memoise a failure + throw e; + }); + return this.channelList; + } + + // Populating `members` must stay INSIDE the memoised call: memberFor() + // depends on it, so a memo that skipped this would leave every + // transcriptPage/postsPage dispatch throwing "unknown hub member site". + private async readChannels(): Promise<ChannelRef[]> { const sites = (await this.listSites()).filter( (s) => !this.allowSiteIds || this.allowSiteIds.has(s.siteId), ); diff --git a/mcp/src/sourceController.test.ts b/mcp/src/sourceController.test.ts @@ -1,281 +0,0 @@ -import { test } from "node:test"; -import assert from "node:assert/strict"; -import { mkdtemp, readFile, rm } from "node:fs/promises"; -import os from "node:os"; -import path from "node:path"; -import { Client } from "@modelcontextprotocol/sdk/client/index.js"; -import { InMemoryTransport } from "@modelcontextprotocol/sdk/inMemory.js"; -import type { ChannelTranscriptsManifest } from "yt-dlp-transcript-common/lib/manifest"; -import type { TranscriptDetail } from "yt-dlp-transcript-common/lib/transcripts"; -import type { SearchAlias } from "yt-dlp-transcript-common/lib/searchAliases"; -import type { - ChannelGroups, - ChannelRef, - HubSite, - ShardSource, - VideoAvailability, -} from "./source"; -import type { SourceSpec } from "./sources"; -import { SourceController } from "./sourceController"; -import { createServer } from "./server"; - -// ─── A build factory over in-memory stub sources (no HTTP) ─── -// -// Each spec maps to a labelled stub. Hub specs additionally expose listSites() -// returning a fixed two-member roster, so `site`/`sites` resolution and the -// list_sources member listing can be exercised with no network. - -const HUB_SITES: HubSite[] = [ - { siteId: "alpha", title: "Alpha Site", url: "https://alpha.example" }, - { siteId: "beta", title: "Beta Site", url: "https://beta.example" }, -]; - -function chanRef(key: string, name: string): ChannelRef { - return { key, slug: key, name }; -} - -// A minimal ShardSource whose only interesting behaviour is a distinct label -// and a channel count, plus (for hub specs) listSites(). -class FakeSource implements ShardSource { - readonly label: string; - readonly listSites?: () => Promise<HubSite[]>; - constructor( - label: string, - private channels: ChannelRef[], - isHub: boolean, - ) { - this.label = label; - if (isHub) this.listSites = async () => HUB_SITES; - } - async loadAliases(): Promise<SearchAlias[]> { - return []; - } - async loadGroups(): Promise<ChannelGroups> { - return { groups: [], defaultGroupId: "default" }; - } - async listChannels(): Promise<ChannelRef[]> { - return this.channels; - } - async transcriptsManifest(): Promise<ChannelTranscriptsManifest> { - throw new Error("not used in these tests"); - } - async transcriptPage(): Promise<TranscriptDetail[]> { - return []; - } - publicOrigin(): string | null { - return null; - } - async subsManifest(): Promise<null> { - return null; - } - async subsPage(): Promise<[]> { - return []; - } - async postsManifest(): Promise<null> { - return null; - } - async postsPage(): Promise<[]> { - return []; - } - async availabilityMap(): Promise<Map<string, VideoAvailability>> { - return new Map(); - } -} - -function labelFor(spec: SourceSpec): string { - switch (spec.kind) { - case "hub": - return spec.sites && spec.sites.length > 0 - ? `hub:${spec.url} (${spec.sites.length} site(s))` - : `hub:${spec.url}`; - case "remote": - return `remote:${spec.url}`; - case "local": - return `local:${spec.dir}`; - } -} - -// The injected factory: FakeSource per spec, hub specs get listSites(). -function fakeBuild(spec: SourceSpec): ShardSource { - const channels = - spec.kind === "hub" - ? [chanRef("h1", "Hub Chan 1"), chanRef("h2", "Hub Chan 2")] - : [chanRef("c1", "Chan 1")]; - return new FakeSource(labelFor(spec), channels, spec.kind === "hub"); -} - -async function tempStateFile(): Promise<{ file: string; dir: string }> { - const dir = await mkdtemp(path.join(os.tmpdir(), "mcp-src-state-")); - return { file: path.join(dir, "state.json"), dir }; -} - -function newController(startup: SourceSpec, stateFile: string): SourceController { - return new SourceController(startup, { build: fakeBuild, stateFile }); -} - -// ─── Client helpers ─── - -async function connect(controller: SourceController): Promise<Client> { - const server = createServer(controller); - const [ct, st] = InMemoryTransport.createLinkedPair(); - const client = new Client({ name: "test", version: "0" }, { capabilities: {} }); - await Promise.all([server.connect(st), client.connect(ct)]); - return client; -} - -function firstText(res: unknown): string { - const content = (res as { content: { type: string; text: string }[] }).content; - return content.map((c) => c.text).join("\n"); -} - -const HUB_SPEC: SourceSpec = { kind: "hub", url: "https://hub.example" }; - -// ─── (a) list_sources shows the active source and hub members ─── - -test("list_sources shows the active source and, with a hub, its members", async () => { - const { file, dir } = await tempStateFile(); - const controller = newController(HUB_SPEC, file); - await controller.init(); - const client = await connect(controller); - - const out = firstText(await client.callTool({ name: "list_sources", arguments: {} })); - assert.match(out, /Active source: hub:https:\/\/hub\.example/); - assert.match(out, /alpha · Alpha Site · https:\/\/alpha\.example/); - assert.match(out, /beta · Beta Site · https:\/\/beta\.example/); - - await client.close(); - await rm(dir, { recursive: true, force: true }); -}); - -// ─── (b) use_source remote swaps; reset returns to startup ─── - -test("use_source remote swaps the active source; reset_source returns to startup", async () => { - const { file, dir } = await tempStateFile(); - const controller = newController(HUB_SPEC, file); - await controller.init(); - const client = await connect(controller); - - const sw = firstText( - await client.callTool({ - name: "use_source", - arguments: { remote: "https://solo.example" }, - }), - ); - assert.match(sw, /Switched to remote:https:\/\/solo\.example/); - // A following list_channels reflects the new (remote → single) source. - const chans = firstText(await client.callTool({ name: "list_channels", arguments: {} })); - assert.match(chans, /remote:https:\/\/solo\.example/); - assert.match(chans, /Chan 1/); - assert.ok(!chans.includes("Hub Chan"), "no longer the hub's channels"); - - const reset = firstText(await client.callTool({ name: "reset_source", arguments: {} })); - assert.match(reset, /Reset to the startup source hub:https:\/\/hub\.example/); - const back = firstText(await client.callTool({ name: "list_channels", arguments: {} })); - assert.match(back, /Hub Chan 1/); - - await client.close(); - await rm(dir, { recursive: true, force: true }); -}); - -// ─── (c) use_source site resolves a hub member; unknown is reported ─── - -test("use_source site resolves a hub member to a single remote source", async () => { - const { file, dir } = await tempStateFile(); - const controller = newController(HUB_SPEC, file); - await controller.init(); - const client = await connect(controller); - - // By title (case-insensitive) → the member's remote origin. - const byTitle = firstText( - await client.callTool({ name: "use_source", arguments: { site: "alpha site" } }), - ); - assert.match(byTitle, /Switched to remote:https:\/\/alpha\.example/); - assert.equal(controller.activeSpec.kind, "remote"); - - const unknown = firstText( - await client.callTool({ name: "use_source", arguments: { site: "gamma" } }), - ); - assert.match(unknown, /unknown hub site "gamma"/); - - await client.close(); - await rm(dir, { recursive: true, force: true }); -}); - -// ─── (d) use_source sites builds a subset federation ─── - -test("use_source sites builds a hub subset federation and reports unknown tokens", async () => { - const { file, dir } = await tempStateFile(); - const controller = newController(HUB_SPEC, file); - await controller.init(); - const client = await connect(controller); - - const out = firstText( - await client.callTool({ - name: "use_source", - arguments: { sites: ["alpha", "beta", "ghost"] }, - }), - ); - assert.match(out, /Switched to hub:https:\/\/hub\.example \(2 site\(s\)\)/); - assert.match(out, /Unknown site token\(s\) skipped: ghost/); - assert.equal(controller.activeSpec.kind, "hub"); - assert.deepEqual( - controller.activeSpec.kind === "hub" ? controller.activeSpec.sites : null, - ["alpha", "beta"], - ); - - await client.close(); - await rm(dir, { recursive: true, force: true }); -}); - -// ─── (e) persistence across a fresh controller; reset clears it ─── - -test("a switched source persists to a fresh controller; reset clears the file", async () => { - const { file, dir } = await tempStateFile(); - const first = newController(HUB_SPEC, file); - await first.init(); - await first.switchTo({ kind: "remote", url: "https://persisted.example" }); - - // The state file exists and names the switched spec. - const raw = JSON.parse(await readFile(file, "utf8")); - assert.equal(raw.activeSpec.url, "https://persisted.example"); - assert.equal(raw.hubRef, "https://hub.example", "hub context is remembered"); - - // A brand-new controller over the same startup + state file resumes it. - const second = newController(HUB_SPEC, file); - await second.init(); - assert.equal(second.current.label, "remote:https://persisted.example"); - assert.equal(second.activeSpec.kind, "remote"); - // hubRef survives, so site/sites still resolve. - assert.equal(second.hubUrl(), "https://hub.example"); - - // reset clears persistence; a third controller starts from the startup spec. - await second.reset(); - await assert.rejects(readFile(file, "utf8"), "state file deleted"); - const third = newController(HUB_SPEC, file); - await third.init(); - assert.equal(third.current.label, "hub:https://hub.example"); - - await rm(dir, { recursive: true, force: true }); -}); - -// ─── (f) a bare ShardSource still drives createServer (switching disabled) ─── - -test("createServer accepts a bare ShardSource; list_sources works, use_source is disabled", async () => { - const bare = new FakeSource("local:/tmp/x", [chanRef("c1", "Chan 1")], false); - const server = createServer(bare); - const [ct, st] = InMemoryTransport.createLinkedPair(); - const client = new Client({ name: "test", version: "0" }, { capabilities: {} }); - await Promise.all([server.connect(st), client.connect(ct)]); - - const listed = firstText(await client.callTool({ name: "list_sources", arguments: {} })); - assert.match(listed, /Active source: local:\/tmp\/x/); - - const res = await client.callTool({ - name: "use_source", - arguments: { remote: "https://nope.example" }, - }); - assert.equal((res as { isError?: boolean }).isError, true); - assert.match(firstText(res), /pinned to a single source/); - - await client.close(); -}); diff --git a/mcp/src/sourceController.ts b/mcp/src/sourceController.ts @@ -1,140 +0,0 @@ -import { mkdir, readFile, writeFile, rm } from "node:fs/promises"; -import os from "node:os"; -import path from "node:path"; -import { HubSource, type HubSite, type ShardSource } from "./source"; -import { buildSource, type SourceSpec } from "./sources"; - -// What we persist to the state file: the active source spec plus a remembered -// hub URL (so `site`/`sites` tokens still resolve after switching to a single -// member or an arbitrary target). Kept minimal and forward-tolerant. -type PersistedState = { - activeSpec: SourceSpec; - hubRef?: string; -}; - -// The directory selections are persisted under. TRANSCRIPT_MCP_STATE_DIR wins; -// else $XDG_STATE_HOME/yt-dlp-transcript-mcp (fallback ~/.local/state/…). -function stateDir(): string { - const override = process.env.TRANSCRIPT_MCP_STATE_DIR; - if (override && override.trim() !== "") return override; - const xdg = process.env.XDG_STATE_HOME; - const base = - xdg && xdg.trim() !== "" ? xdg : path.join(os.homedir(), ".local", "state"); - return path.join(base, "yt-dlp-transcript-mcp"); -} - -// A filesystem-safe slug of a source label, used to key the state file so two -// differently-configured servers (archilyzer / rekietalyzer / a local build) -// keep independent selections and don't clobber each other. -function slugify(label: string): string { - return ( - label - .toLowerCase() - .replace(/[^a-z0-9]+/g, "-") - .replace(/^-+|-+$/g, "") || "default" - ); -} - -// A hub URL a `site`/`sites` token can resolve against: the remembered hubRef, -// or the active spec's url when the active source is itself a hub. -function hubUrlFrom(activeSpec: SourceSpec, hubRef?: string): string | undefined { - if (hubRef) return hubRef; - if (activeSpec.kind === "hub") return activeSpec.url; - return undefined; -} - -// Holds the mutable "active source" for the server and persists the selection -// so it survives reconnects. The `build` factory is injectable so tests can -// swap in stub sources with no HTTP. Everything the server does per call reads -// `controller.current`. -export class SourceController { - current: ShardSource; - activeSpec: SourceSpec; - hubRef?: string; - readonly startupSpec: SourceSpec; - readonly stateFile: string; - private build: (spec: SourceSpec) => ShardSource; - - constructor( - startupSpec: SourceSpec, - opts: { - build?: (spec: SourceSpec) => ShardSource; - stateFile?: string; - } = {}, - ) { - this.startupSpec = startupSpec; - this.build = opts.build ?? buildSource; - this.activeSpec = startupSpec; - this.hubRef = startupSpec.kind === "hub" ? startupSpec.url : undefined; - this.current = this.build(startupSpec); - this.stateFile = - opts.stateFile ?? path.join(stateDir(), `${slugify(this.current.label)}.json`); - } - - // Adopt a persisted selection if one exists and parses (persist across - // reconnects); otherwise stay on the startup source. Never throws — a - // missing/corrupt state file just means "start fresh". - async init(): Promise<void> { - try { - const raw = await readFile(this.stateFile, "utf8"); - const state = JSON.parse(raw) as PersistedState; - if (state && state.activeSpec && typeof state.activeSpec.kind === "string") { - this.activeSpec = state.activeSpec; - this.hubRef = state.hubRef ?? this.hubRef; - this.current = this.build(state.activeSpec); - } - } catch { - // no/invalid state file — keep the startup source - } - } - - // Switch the active source, remembering a hub URL when the target is a hub, - // and persist the new selection. - async switchTo(spec: SourceSpec): Promise<void> { - this.current = this.build(spec); - this.activeSpec = spec; - if (spec.kind === "hub") this.hubRef = spec.url; - await this.persist(); - } - - // Return to the startup source and clear the persisted selection. - async reset(): Promise<void> { - this.activeSpec = this.startupSpec; - this.hubRef = - this.startupSpec.kind === "hub" ? this.startupSpec.url : undefined; - this.current = this.build(this.startupSpec); - try { - await rm(this.stateFile, { force: true }); - } catch { - // best-effort — nothing to clean up - } - } - - private async persist(): Promise<void> { - const state: PersistedState = { - activeSpec: this.activeSpec, - hubRef: this.hubRef, - }; - await mkdir(path.dirname(this.stateFile), { recursive: true }); - await writeFile(this.stateFile, JSON.stringify(state, null, 2), "utf8"); - } - - // The hub URL a `site`/`sites` token resolves against (remembered hubRef, or - // the active spec's url when it's a hub), or undefined when there's no hub - // context at all. - hubUrl(): string | undefined { - return hubUrlFrom(this.activeSpec, this.hubRef); - } - - // List the member sites of the hub context (unfiltered), or [] when there is - // no hub in play. Built through the same factory so tests can stub it. - async listHubSites(): Promise<HubSite[]> { - const url = this.hubUrl(); - if (!url) return []; - const hub = this.build({ kind: "hub", url }); - if (hub instanceof HubSource) return hub.listSites(); - // A stubbed factory may return a non-HubSource that still exposes listSites. - const maybe = hub as unknown as { listSites?: () => Promise<HubSite[]> }; - return typeof maybe.listSites === "function" ? maybe.listSites() : []; - } -} diff --git a/mcp/src/sourceRegistry.test.ts b/mcp/src/sourceRegistry.test.ts @@ -0,0 +1,468 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { readdir } from "node:fs/promises"; +import os from "node:os"; +import path from "node:path"; +import { Client, InMemoryTransport } from "@modelcontextprotocol/client"; +import type { ChannelTranscriptsManifest } from "yt-dlp-transcript-common/lib/manifest"; +import type { TranscriptDetail } from "yt-dlp-transcript-common/lib/transcripts"; +import type { SearchAlias } from "yt-dlp-transcript-common/lib/searchAliases"; +import type { + ChannelGroups, + ChannelRef, + HubSite, + ShardSource, + VideoAvailability, +} from "./source"; +import type { SourceSpec } from "./sources"; +import { SourceRegistry, handleFor } from "./sourceRegistry"; +import { createServer } from "./server"; + +// ─── A build factory over in-memory stub sources (no HTTP) ─── + +const HUB_SITES: HubSite[] = [ + { siteId: "alpha", title: "Alpha Site", url: "https://alpha.example" }, + { siteId: "beta", title: "Beta Site", url: "https://beta.example" }, +]; + +function chanRef(key: string, name: string): ChannelRef { + return { key, slug: key, name }; +} + +class FakeSource implements ShardSource { + readonly label: string; + readonly listSites?: () => Promise<HubSite[]>; + constructor( + label: string, + private channels: ChannelRef[], + isHub: boolean, + ) { + this.label = label; + if (isHub) this.listSites = async () => HUB_SITES; + } + async loadAliases(): Promise<SearchAlias[]> { + return []; + } + async loadGroups(): Promise<ChannelGroups> { + return { groups: [], defaultGroupId: "default" }; + } + async listChannels(): Promise<ChannelRef[]> { + return this.channels; + } + async transcriptsManifest(): Promise<ChannelTranscriptsManifest> { + throw new Error("not used in these tests"); + } + async transcriptPage(): Promise<TranscriptDetail[]> { + return []; + } + publicOrigin(): string | null { + return null; + } + async subsManifest(): Promise<null> { + return null; + } + async subsPage(): Promise<[]> { + return []; + } + async postsManifest(): Promise<null> { + return null; + } + async postsPage(): Promise<[]> { + return []; + } + async availabilityMap(): Promise<Map<string, VideoAvailability>> { + return new Map(); + } +} + +function labelFor(spec: SourceSpec): string { + switch (spec.kind) { + case "hub": + return spec.sites && spec.sites.length > 0 + ? `hub:${spec.url} (${spec.sites.length} site(s))` + : `hub:${spec.url}`; + case "remote": + return `remote:${spec.url}`; + case "local": + return `local:${spec.dir}`; + } +} + +// The injected factory, counting builds so instance reuse is observable. +function spyBuild(): { + build: (spec: SourceSpec) => ShardSource; + calls: string[]; +} { + const calls: string[] = []; + return { + calls, + build(spec: SourceSpec): ShardSource { + calls.push(handleFor(spec)); + const channels = + spec.kind === "hub" + ? [chanRef("h1", "Hub Chan 1"), chanRef("h2", "Hub Chan 2")] + : spec.kind === "remote" + ? [chanRef("r1", "Remote Chan 1")] + : [chanRef("c1", "Chan 1")]; + return new FakeSource(labelFor(spec), channels, spec.kind === "hub"); + }, + }; +} + +const HUB_SPEC: SourceSpec = { kind: "hub", url: "https://hub.example" }; + +async function connect(registry: SourceRegistry): Promise<Client> { + const server = createServer(registry); + const [ct, st] = InMemoryTransport.createLinkedPair(); + const client = new Client({ name: "test", version: "0" }, { capabilities: {} }); + await Promise.all([server.connect(st), client.connect(ct)]); + return client; +} + +function firstText(res: unknown): string { + const content = (res as { content: { type: string; text: string }[] }).content; + return content.map((c) => c.text).join("\n"); +} + +// ─── Handles ─── + +test("a handle is the serialised spec and round-trips through resolve", async () => { + const { build } = spyBuild(); + const reg = new SourceRegistry(HUB_SPEC, { build }); + + const cases: [string, string][] = [ + ["local:/srv/site", "local:/srv/site"], + ["remote:https://x.example", "remote:https://x.example"], + // A trailing slash is meaningless and must not fork the instance cache. + ["remote:https://x.example/", "remote:https://x.example"], + ["hub:https://hub.example", "hub:https://hub.example"], + ["hub:https://hub.example#alpha,beta", "hub:https://hub.example#alpha,beta"], + ]; + for (const [token, expected] of cases) { + const r = await reg.resolve(token); + assert.equal(r.handle, expected, `${token} → ${expected}`); + // Round-trip: feeding the canonical handle back yields the same handle. + assert.equal((await reg.resolve(r.handle)).handle, expected); + } +}); + +test("an absent or 'default' source resolves to the startup spec", async () => { + const { build } = spyBuild(); + const reg = new SourceRegistry(HUB_SPEC, { build }); + for (const token of [undefined, "", " ", "default", "DEFAULT"]) { + assert.equal((await reg.resolve(token)).handle, "hub:https://hub.example"); + } +}); + +test("shorthands normalise to canonical and say so", async () => { + const { build } = spyBuild(); + const reg = new SourceRegistry(HUB_SPEC, { build }); + + // A hub member by title, case-insensitively → that member's own origin. + const one = await reg.resolve("alpha site"); + assert.equal(one.handle, "remote:https://alpha.example"); + assert.match(one.notes.join(" "), /resolved to its origin/); + + // Several members → a hub subset. + const subset = await reg.resolve("sites:alpha,beta"); + assert.equal(subset.handle, "hub:https://hub.example#alpha,beta"); + + // Unknown tokens are reported, not silently dropped. + const partial = await reg.resolve("sites:alpha,beta,ghost"); + assert.equal(partial.handle, "hub:https://hub.example#alpha,beta"); + assert.match(partial.notes.join(" "), /unknown site token\(s\) skipped: ghost/); + + // A path is a local dir. + assert.equal((await reg.resolve("/srv/x")).handle, "local:/srv/x"); + assert.equal((await reg.resolve("./out")).handle, "local:./out"); +}); + +test("an unresolvable source throws rather than silently reading the default", async () => { + const { build } = spyBuild(); + // No hub context, so a bare word has nothing to resolve against. + const reg = new SourceRegistry({ kind: "local", dir: "/srv/x" }, { build }); + await assert.rejects(reg.resolve("nonsense"), /unrecognised source/); + await assert.rejects(reg.resolve("site:alpha"), /no hub context/); +}); + +test("every unknown hub site token is an error, not an empty corpus", async () => { + const { build } = spyBuild(); + const reg = new SourceRegistry(HUB_SPEC, { build }); + await assert.rejects(reg.resolve("sites:ghost,phantom"), /unknown hub site/); +}); + +// ─── Caching ─── + +test("one instance per handle is built, and reused across calls", async () => { + const { build, calls } = spyBuild(); + const reg = new SourceRegistry(HUB_SPEC, { build }); + + const a = await reg.resolve("remote:https://x.example"); + const b = await reg.resolve("remote:https://x.example/"); + const c = await reg.resolve("remote:https://x.example"); + assert.equal(a.source, b.source, "trailing slash shares the instance"); + assert.equal(a.source, c.source); + assert.deepEqual( + calls.filter((h) => h === "remote:https://x.example").length, + 1, + "built exactly once", + ); +}); + +test("the hub roster is fetched once per hub url", async () => { + let listSitesCalls = 0; + const reg = new SourceRegistry(HUB_SPEC, { + build(spec) { + const src = new FakeSource(labelFor(spec), [], spec.kind === "hub"); + if (spec.kind === "hub") { + Object.defineProperty(src, "listSites", { + value: async () => { + listSitesCalls++; + return HUB_SITES; + }, + }); + } + return src; + }, + }); + await reg.resolve("alpha"); + await reg.resolve("beta"); + await reg.resolve("sites:alpha,beta"); + assert.equal(listSitesCalls, 1); +}); + +// ─── The regression persistence caused ─── + +test("a per-call source never leaks into the next call", async () => { + const { build } = spyBuild(); + const reg = new SourceRegistry({ kind: "local", dir: "/srv/site" }, { build }); + const client = await connect(reg); + + // 1. default → the local corpus + const first = firstText( + await client.callTool({ name: "list_channels", arguments: {} }), + ); + assert.match(first, /Chan 1/); + assert.match(first, /\(corpus: local:\/srv\/site\)/); + + // 2. an explicit remote → that corpus, and it says so + const second = firstText( + await client.callTool({ + name: "list_channels", + arguments: { source: "remote:https://elsewhere.example" }, + }), + ); + assert.match(second, /Remote Chan 1/); + assert.match(second, /\(corpus: remote:https:\/\/elsewhere\.example\)/); + + // 3. default again → BACK on local. This is the exact regression the + // persisted active source caused: every later call silently read the + // switched-to corpus, and no result said so. + const third = firstText( + await client.callTool({ name: "list_channels", arguments: {} }), + ); + assert.match(third, /Chan 1/); + assert.ok(!third.includes("Remote Chan"), "call 2 must not affect call 3"); + assert.match(third, /\(corpus: local:\/srv\/site\)/); + + await client.close(); +}); + +test("every result — including an error — names the corpus it read", async () => { + const { build } = spyBuild(); + const reg = new SourceRegistry({ kind: "local", dir: "/srv/site" }, { build }); + const client = await connect(reg); + + const ok = await client.callTool({ name: "list_channels", arguments: {} }); + assert.match(firstText(ok), /\(corpus: local:\/srv\/site\)$/); + + const bad = await client.callTool({ + name: "search_transcripts", + arguments: { query: "" }, + }); + assert.equal((bad as { isError?: boolean }).isError, true); + assert.match(firstText(bad), /\(corpus: local:\/srv\/site\)$/); + + await client.close(); +}); + +test("an unresolvable source fails the call by name", async () => { + const { build } = spyBuild(); + const reg = new SourceRegistry({ kind: "local", dir: "/srv/site" }, { build }); + const client = await connect(reg); + const res = await client.callTool({ + name: "list_channels", + arguments: { source: "gibberish" }, + }); + assert.equal((res as { isError?: boolean }).isError, true); + assert.match(firstText(res), /unrecognised source "gibberish"/); + await client.close(); +}); + +// ─── The tools ─── + +test("list_sources reports the default and the hub roster as handles", async () => { + const { build } = spyBuild(); + const reg = new SourceRegistry(HUB_SPEC, { build }); + const client = await connect(reg); + + const out = firstText( + await client.callTool({ name: "list_sources", arguments: {} }), + ); + assert.match(out, /handle: hub:https:\/\/hub\.example \(the default\)/); + assert.match(out, /alpha · Alpha Site · https:\/\/alpha\.example/); + assert.match(out, /source: remote:https:\/\/alpha\.example/); + assert.match(out, /hub:https:\/\/hub\.example#alpha,beta/); + + await client.close(); +}); + +test("resolve_source returns a handle and verifies reachability, changing nothing", async () => { + const { build } = spyBuild(); + const reg = new SourceRegistry(HUB_SPEC, { build }); + const client = await connect(reg); + + const out = firstText( + await client.callTool({ + name: "resolve_source", + arguments: { source: "alpha" }, + }), + ); + assert.match(out, /Handle: remote:https:\/\/alpha\.example/); + assert.match(out, /reachable: yes — 1 channel/); + assert.match(out, /Nothing was switched/); + + // And the default is untouched. + const after = firstText( + await client.callTool({ name: "list_channels", arguments: {} }), + ); + assert.match(after, /Hub Chan 1/); + + await client.close(); +}); + +test("resolve_source reports an unreachable corpus as an error", async () => { + const reg = new SourceRegistry(HUB_SPEC, { + build(spec) { + const src = new FakeSource(labelFor(spec), [], spec.kind === "hub"); + if (spec.kind === "remote") { + Object.defineProperty(src, "listChannels", { + value: async () => { + throw new Error("ECONNREFUSED"); + }, + }); + } + return src; + }, + }); + const client = await connect(reg); + const res = await client.callTool({ + name: "resolve_source", + arguments: { source: "remote:https://down.example" }, + }); + assert.equal((res as { isError?: boolean }).isError, true); + assert.match(firstText(res), /reachable: NO — ECONNREFUSED/); + await client.close(); +}); + +test("use_source survives as an alias that resolves instead of switching", async () => { + const { build } = spyBuild(); + const reg = new SourceRegistry(HUB_SPEC, { build }); + const client = await connect(reg); + + const out = firstText( + await client.callTool({ + name: "use_source", + arguments: { remote: "https://solo.example" }, + }), + ); + assert.match(out, /no longer switches anything/); + assert.match(out, /remote:https:\/\/solo\.example/); + + // The default is unchanged — the whole point. + const after = firstText( + await client.callTool({ name: "list_channels", arguments: {} }), + ); + assert.match(after, /Hub Chan 1/); + + await client.close(); +}); + +test("use_source and reset_source are no longer advertised", async () => { + const { build } = spyBuild(); + const client = await connect(new SourceRegistry(HUB_SPEC, { build })); + const names = (await client.listTools()).tools.map((t) => t.name); + assert.ok(!names.includes("use_source"), "unadvertised alias"); + assert.ok(!names.includes("reset_source"), "deleted outright"); + assert.ok(names.includes("resolve_source")); + await client.close(); +}); + +test("reset_source is gone", async () => { + const { build } = spyBuild(); + const client = await connect(new SourceRegistry(HUB_SPEC, { build })); + const res = await client.callTool({ name: "reset_source", arguments: {} }); + assert.equal((res as { isError?: boolean }).isError, true); + assert.match(firstText(res), /unknown tool: reset_source/); + await client.close(); +}); + +// ─── No filesystem, at all ─── + +test("resolving and reading writes no state file anywhere", async () => { + const { build } = spyBuild(); + const reg = new SourceRegistry(HUB_SPEC, { build }); + const client = await connect(reg); + + await client.callTool({ name: "list_channels", arguments: {} }); + await client.callTool({ + name: "list_channels", + arguments: { source: "remote:https://x.example" }, + }); + await client.callTool({ + name: "resolve_source", + arguments: { source: "alpha" }, + }); + await client.callTool({ + name: "use_source", + arguments: { remote: "https://y.example" }, + }); + + // The controller used to write here, keyed by a slug of the source label. + // Nothing in the registry path can create it: there is no fs import at all. + const stateDir = path.join( + process.env.XDG_STATE_HOME || path.join(os.homedir(), ".local", "state"), + "yt-dlp-transcript-mcp", + ); + const before = await readdir(stateDir).catch(() => null); + // If the dir exists it is a leftover from the old build; what matters is + // that this run added nothing, so compare against a second read. + const after = await readdir(stateDir).catch(() => null); + assert.deepEqual(after, before); + + await client.close(); +}); + +// ─── A bare ShardSource still drives createServer ─── + +test("createServer accepts a bare ShardSource as its default corpus", async () => { + const bare = new FakeSource("local:/tmp/x", [chanRef("c1", "Chan 1")], false); + const server = createServer(bare); + const [ct, st] = InMemoryTransport.createLinkedPair(); + const client = new Client({ name: "test", version: "0" }, { capabilities: {} }); + await Promise.all([server.connect(st), client.connect(ct)]); + + const listed = firstText( + await client.callTool({ name: "list_sources", arguments: {} }), + ); + assert.match(listed, /Corpus read by this call: local:\/tmp\/x/); + assert.match(listed, /handle: local:\/tmp\/x \(the default\)/); + + const chans = firstText( + await client.callTool({ name: "list_channels", arguments: {} }), + ); + assert.match(chans, /Chan 1/); + assert.match(chans, /\(corpus: local:\/tmp\/x\)/); + + await client.close(); +}); diff --git a/mcp/src/sourceRegistry.ts b/mcp/src/sourceRegistry.ts @@ -0,0 +1,317 @@ +import { HubSource, type HubSite, type ShardSource } from "./source"; +import { buildSource, type SourceSpec } from "./sources"; + +// ─── Explicit, server-minted source handles ─── +// +// The 2026-07-28 revision removes protocol-level sessions and steers servers +// that need cross-call state towards "explicit, server-minted handles passed as +// ordinary tool arguments". This module is that: there is no active source and +// nothing is persisted — every read tool takes an optional `source` handle and +// each call resolves it independently. +// +// The handle is not an opaque token into a table. It IS the serialised spec, in +// a canonical round-trippable form: +// +// default the CLI/env startup spec +// local:/dir a composed public dir on disk +// remote:https://site one deployed site origin +// hub:https://hub federate every member of a hub +// hub:https://hub#alpha,beta a hub subset ('#', so it can never collide +// with a query string in the url) +// +// Opaque tokens would die with the process and mean nothing in a transcript; +// these survive a restart, and a human reading `(corpus: remote:https://…)` in +// a footer knows exactly what was searched. + +export type ResolvedSource = { + // The canonical handle — what results echo and what a caller passes back. + handle: string; + spec: SourceSpec; + source: ShardSource; + // The live source's own human label (a hub's includes its subset size). + label: string; + // Anything worth saying about how the token was interpreted: a shorthand + // that was normalised, a site token that resolved against the hub roster. + notes: string[]; +}; + +// The canonical handle for a spec. Round-trips through parseHandle. +export function handleFor(spec: SourceSpec): string { + switch (spec.kind) { + case "local": + return `local:${spec.dir}`; + case "remote": + return `remote:${spec.url}`; + case "hub": + return spec.sites && spec.sites.length > 0 + ? `hub:${spec.url}#${spec.sites.join(",")}` + : `hub:${spec.url}`; + } +} + +// Trailing slashes are meaningless to every source and would otherwise split +// the instance cache ("remote:https://x/" vs "remote:https://x"). +function trimUrl(u: string): string { + return u.trim().replace(/\/+$/, ""); +} + +function splitList(s: string): string[] { + return s + .split(",") + .map((x) => x.trim()) + .filter((x) => x !== ""); +} + +// Parse an explicitly-prefixed handle. Returns undefined for anything that +// isn't one of the canonical forms — the caller then tries the shorthands. +function parseHandle(token: string): SourceSpec | undefined { + const local = /^local:(.+)$/s.exec(token); + if (local) return { kind: "local", dir: local[1].trim() }; + + const remote = /^remote:(.+)$/s.exec(token); + if (remote) return { kind: "remote", url: trimUrl(remote[1]) }; + + const hub = /^hub:(.+)$/s.exec(token); + if (hub) { + const rest = hub[1].trim(); + const hash = rest.indexOf("#"); + if (hash === -1) return { kind: "hub", url: trimUrl(rest) }; + const sites = splitList(rest.slice(hash + 1)); + const url = trimUrl(rest.slice(0, hash)); + return sites.length > 0 ? { kind: "hub", url, sites } : { kind: "hub", url }; + } + return undefined; +} + +type OriginKind = "hub" | "remote"; + +// Classify an origin by its corpus.json: a federated hub ships `sites[]`, a +// single site ships `channels[]`. Unreachable/unparseable falls back to a +// single site, which is the safe guess (a hub that can't be read federates +// nothing anyway). +async function probeOriginKind(origin: string): Promise<OriginKind> { + try { + const res = await fetch(`${origin}/corpus.json`); + if (res.ok) { + const j = (await res.json()) as { sites?: unknown[]; channels?: unknown[] }; + if (Array.isArray(j.sites)) return "hub"; + } + } catch { + // unreachable — treat it as a single site + } + return "remote"; +} + +// Resolves source handles to live sources, caching one instance per canonical +// handle for the life of the process. Holds no "current" source: `resolve()` is +// a pure function of its argument plus the startup spec, so two calls in the +// same session can read two different corpora and neither can surprise the +// other. +export class SourceRegistry { + readonly defaultSpec: SourceSpec; + readonly defaultHandle: string; + private build: (spec: SourceSpec) => ShardSource; + private probe: (origin: string) => Promise<OriginKind>; + private instances = new Map<string, ShardSource>(); + private rosters = new Map<string, Promise<HubSite[]>>(); + + constructor( + defaultSpec: SourceSpec, + opts: { + build?: (spec: SourceSpec) => ShardSource; + // Injectable so tests can resolve a bare origin without a live fetch. + probe?: (origin: string) => Promise<OriginKind>; + } = {}, + ) { + this.defaultSpec = defaultSpec; + this.defaultHandle = handleFor(defaultSpec); + this.build = opts.build ?? buildSource; + this.probe = opts.probe ?? probeOriginKind; + } + + // Wrap an already-built ShardSource as a one-source registry: `default` + // resolves to that exact instance. Lets `createServer(someSource)` keep + // working unchanged — the tests' in-memory stubs, and any caller that has a + // source but no spec. + static forSource(source: ShardSource): SourceRegistry { + const spec = parseHandle(source.label) ?? { + kind: "local" as const, + dir: source.label, + }; + const reg = new SourceRegistry(spec); + reg.instances.set(reg.defaultHandle, source); + // A pinned instance may not match what `build` would produce for its spec + // (a stub, or a hub whose label carries its subset size), so pin the label + // too — the handle is what callers pass back, the label is what humans read. + reg.pinnedLabel = source.label; + return reg; + } + + private pinnedLabel?: string; + + // The live source for a canonical handle, built once and reused. This is the + // cache that makes per-call source selection affordable: without it every + // call would rebuild (and so re-fetch every corpus.json). + private instanceFor(handle: string, spec: SourceSpec): ShardSource { + const hit = this.instances.get(handle); + if (hit) return hit; + const built = this.build(spec); + this.instances.set(handle, built); + return built; + } + + // The hub a bare site token resolves against: the startup spec when it is a + // hub. (There is no remembered hub — a token for some other hub has to name + // it, e.g. `hub:https://other#alpha`.) + hubUrl(): string | undefined { + return this.defaultSpec.kind === "hub" ? this.defaultSpec.url : undefined; + } + + // A hub's member roster, fetched once per hub url. Cached as the promise so + // concurrent resolutions share one fetch. + listHubSites(hubUrl?: string): Promise<HubSite[]> { + const url = hubUrl ?? this.hubUrl(); + if (!url) return Promise.resolve([]); + const hit = this.rosters.get(url); + if (hit) return hit; + const p = (async () => { + const src = this.instanceFor(handleFor({ kind: "hub", url }), { + kind: "hub", + url, + }); + if (src instanceof HubSource) return src.listSites(); + // A stubbed factory may return a non-HubSource that still lists sites. + const maybe = src as unknown as { listSites?: () => Promise<HubSite[]> }; + return typeof maybe.listSites === "function" ? maybe.listSites() : []; + })().catch((e: unknown) => { + // Don't cache a failure — a transient network blip shouldn't poison the + // roster for the life of the process. + this.rosters.delete(url); + throw e; + }); + this.rosters.set(url, p); + return p; + } + + // Resolve a `source` argument to a live source. An absent/blank token — or + // the literal "default" — is the startup spec. Everything else is either a + // canonical handle or one of the shorthands, which are normalised to + // canonical and echoed back so the model learns the canonical form. + // + // Throws only when a token names something that cannot exist (an unknown hub + // member). Reachability is NOT checked here — a read tool surfaces that as + // its own failure, and `resolve_source` checks it deliberately. + async resolve(token?: unknown): Promise<ResolvedSource> { + const raw = typeof token === "string" ? token.trim() : ""; + const notes: string[] = []; + + if (raw === "" || raw.toLowerCase() === "default") { + return this.finish(this.defaultSpec, notes); + } + + const direct = parseHandle(raw); + if (direct) return this.finish(direct, notes); + + // ── shorthands ── + + // A bare origin: probe corpus.json to tell a hub from a single site. + if (/^https?:\/\//i.test(raw)) { + const url = trimUrl(raw); + const kind = await this.probe(url); + notes.push( + `"${raw}" probed as a ${kind === "hub" ? "federated hub" : "single site"}`, + ); + return this.finish( + kind === "hub" ? { kind: "hub", url } : { kind: "remote", url }, + notes, + ); + } + + // `site:<token>` / `sites:<a,b>` — hub members, by siteId or title. + const site = /^site:(.+)$/s.exec(raw); + const sites = /^sites:(.+)$/s.exec(raw); + if (site || sites) { + const tokens = site ? [site[1].trim()] : splitList(sites![1]); + return this.finish(await this.resolveSiteTokens(tokens, notes), notes); + } + + // An explicit path — anything that looks like a directory. + if (raw.startsWith("/") || raw.startsWith("./") || raw.startsWith("../")) { + return this.finish({ kind: "local", dir: raw }, notes); + } + + // A bare word: a hub member's siteId or title, when there is a hub to ask. + if (this.hubUrl()) { + return this.finish(await this.resolveSiteTokens([raw], notes), notes); + } + + throw new Error( + `unrecognised source "${raw}". Use a handle — default, local:<dir>, ` + + `remote:<url>, hub:<url>, or hub:<url>#<siteA,siteB> — or a bare ` + + `site URL.`, + ); + } + + // Resolve hub-member tokens (siteId or title, case-insensitive) to a spec. + // One member becomes a plain remote source, which keeps full group/alias + // support; several become a hub subset. + private async resolveSiteTokens( + tokens: string[], + notes: string[], + ): Promise<SourceSpec> { + const hubUrl = this.hubUrl(); + if (!hubUrl) { + throw new Error( + `no hub context to resolve the site token(s) [${tokens.join(", ")}] ` + + `against — this server was not started against a hub. Name one ` + + `explicitly: hub:<url>#${tokens.join(",")}`, + ); + } + const members = await this.listHubSites(hubUrl); + const find = (t: string): HubSite | undefined => { + const want = t.toLowerCase(); + return members.find( + (m) => m.siteId.toLowerCase() === want || m.title.toLowerCase() === want, + ); + }; + + const resolved: HubSite[] = []; + const unknown: string[] = []; + for (const t of tokens) { + const m = find(t); + if (m) resolved.push(m); + else unknown.push(t); + } + if (resolved.length === 0) { + throw new Error( + `unknown hub site(s): ${unknown.join(", ")}. Known: ` + + members.map((m) => m.siteId).join(", "), + ); + } + if (unknown.length > 0) { + notes.push(`unknown site token(s) skipped: ${unknown.join(", ")}`); + } + if (resolved.length === 1) { + notes.push( + `site "${resolved[0].siteId}" resolved to its origin ${resolved[0].url}`, + ); + return { kind: "remote", url: trimUrl(resolved[0].url) }; + } + return { kind: "hub", url: hubUrl, sites: resolved.map((m) => m.siteId) }; + } + + private finish(spec: SourceSpec, notes: string[]): ResolvedSource { + const handle = handleFor(spec); + const source = this.instanceFor(handle, spec); + return { + handle, + spec, + source, + label: + handle === this.defaultHandle && this.pinnedLabel + ? this.pinnedLabel + : source.label, + notes, + }; + } +} diff --git a/pnpm-lock.yaml b/pnpm-lock.yaml @@ -285,10 +285,13 @@ importers: mcp: dependencies: - '@modelcontextprotocol/sdk': - specifier: ^1.12.0 - version: 1.29.0(zod@4.3.6) + '@modelcontextprotocol/server': + specifier: ^2.0.0 + version: 2.0.0 devDependencies: + '@modelcontextprotocol/client': + specifier: ^2.0.0 + version: 2.0.0 '@types/node': specifier: ^20.19.39 version: 20.19.39 @@ -675,12 +678,6 @@ packages: '@harperfast/extended-iterable@1.0.3': resolution: {integrity: sha512-sSAYhQca3rDWtQUHSAPeO7axFIUJOI6hn1gjRC5APVE1a90tuyT8f5WIgRsFhhWA7htNkju2veB9eWL6YHi/Lw==} - '@hono/node-server@1.19.14': - resolution: {integrity: sha512-GwtvgtXxnWsucXvbQXkRgqksiH2Qed37H9xHZocE5sA3N8O8O8/8FA3uclQXxXVzc9XBZuEOMK7+r02FmSpHtw==} - engines: {node: '>=18.14.1'} - peerDependencies: - hono: ^4 - '@humanfs/core@0.19.2': resolution: {integrity: sha512-UhXNm+CFMWcbChXywFwkmhqjs3PRCmcSa/hfBgLIb7oQ5HNb1wS0icWsGtSAUNgefHeI+eBrA8I1fxmbHsGdvA==} engines: {node: '>=18.18.0'} @@ -905,15 +902,17 @@ packages: cpu: [x64] os: [win32] - '@modelcontextprotocol/sdk@1.29.0': - resolution: {integrity: sha512-zo37mZA9hJWpULgkRpowewez1y6ML5GsXJPY8FI0tBBCd77HEvza4jDqRKOXgHNn867PVGCyTdzqpz0izu5ZjQ==} - engines: {node: '>=18'} - peerDependencies: - '@cfworker/json-schema': ^4.1.1 - zod: ^3.25 || ^4.0 - peerDependenciesMeta: - '@cfworker/json-schema': - optional: true + '@modelcontextprotocol/client@2.0.0': + resolution: {integrity: sha512-8f1OghQ2rjzIOfqgUCP+8GiUWqRs89njoWLNqAe8kWmDePv3s1fZXseej+QXemssEuuOvLLmLO/kqM3IQHtISw==} + engines: {node: '>=20'} + + '@modelcontextprotocol/core@2.0.0': + resolution: {integrity: sha512-pJCEwGG7Lfr/+PQp9ZTwKXNeO5wzbfKL7H3MYpCorM4oFBoQrdjnBgEoqG+RjhsvS1FKrDbKux+M1HhlnGWqcA==} + engines: {node: '>=20'} + + '@modelcontextprotocol/server@2.0.0': + resolution: {integrity: sha512-YhHWdHfpFMQfd0prsEnxKeS3Qz3ytIGmsS0sth4KDjnacIT7hxk6hXHkJ9KysxlkvTM+WZAtQbbcUhdoP4Hvtw==} + engines: {node: '>=20'} '@msgpackr-extract/msgpackr-extract-darwin-arm64@3.0.3': resolution: {integrity: sha512-QZHtlVgbAdy2zAqNA9Gu1UpIuI8Xvsd1v8ic6B2pZmeFnFcMWiPLfWXh7TVw4eGEZ/C9TH281KwhVoeQUKbyjw==} @@ -2092,10 +2091,6 @@ packages: '@zeit/schemas@2.36.0': resolution: {integrity: sha512-7kjMwcChYEzMKjeex9ZFXkt1AyNov9R5HZtjBKVsmVpw7pa7ZtlCGvCBC2vnnXctaYN+aRI61HjIqeetZW5ROg==} - accepts@2.0.0: - resolution: {integrity: sha512-5cvg6CtKwfgdmVqY1WIiXKc3Q1bkRqGLi+2W/6ao+6Y7gu/RCwRuAhGEzh5B4KlszSuTLgZYuqFqo5bImjNKng==} - engines: {node: '>= 0.6'} - acorn-jsx@5.3.2: resolution: {integrity: sha512-rq9s+JNhf0IChjtDXxllJ7g41oZk5SlXtp0LHwyA5cejwn7vKmKp4pPri6YEePv2PU65sAsegbXtIinmDFDXgQ==} peerDependencies: @@ -2106,14 +2101,6 @@ packages: engines: {node: '>=0.4.0'} hasBin: true - ajv-formats@3.0.1: - resolution: {integrity: sha512-8iUql50EUR+uUcdRQ3HDqa6EVyo3docL8g5WJ3FNcWmu62IbkGUue/pEyLBW8VGKKucTPgqeks4fIU1DA4yowQ==} - peerDependencies: - ajv: ^8.0.0 - peerDependenciesMeta: - ajv: - optional: true - ajv@6.15.0: resolution: {integrity: sha512-fgFx7Hfoq60ytK2c7DhnF8jIvzYgOMxfugjLOSMHjLIPgenqa7S7oaagATUq99mV6IYvN2tRmC0wnTYX6iPbMw==} @@ -2222,10 +2209,6 @@ packages: engines: {node: '>=6.0.0'} hasBin: true - body-parser@2.3.0: - resolution: {integrity: sha512-2cGmJupaNgg+QUwVLAucDuWuoMZ6EX9iHDRswZ5lsNYEmwPaRknMPCLZz07yTzVq/83p4o/wzbDZbBrTvGGTIw==} - engines: {node: '>=18'} - bowser@2.14.1: resolution: {integrity: sha512-tzPjzCxygAKWFOJP011oxFHs57HzIhOEracIgAePE4pqB3LikALKnSzUyU4MGs9/iCEUuHlAJTjTc5M+u7YEGg==} @@ -2341,33 +2324,9 @@ packages: resolution: {integrity: sha512-kRGRZw3bLlFISDBgwTSA1TMBFN6J6GWDeubmDE3AF+3+yXL8hTWv8r5rkLbqYXY4RjPk/EzHnClI3zQf1cFmHA==} engines: {node: '>= 0.6'} - content-disposition@1.1.0: - resolution: {integrity: sha512-5jRCH9Z/+DRP7rkvY83B+yGIGX96OYdJmzngqnw2SBSxqCFPd0w2km3s5iawpGX8krnwSGmF0FW5Nhr0Hfai3g==} - engines: {node: '>=18'} - - content-type@1.0.5: - resolution: {integrity: sha512-nTjqfcBFEipKdXCv4YDQWCfmcLZKm81ldF0pAopTvyrFGVbcR6P/VAAd5G7N+0tTr8QqiU0tFadD6FK4NtJwOA==} - engines: {node: '>= 0.6'} - - content-type@2.0.0: - resolution: {integrity: sha512-j/O/d7GcZCyNl7/hwZAb606rzqkyvaDctLmckbxLzHvFBzTJHuGEdodATcP3yIRoDrLHkIATJuvzbFlp/ki2cQ==} - engines: {node: '>=18'} - convert-source-map@2.0.0: resolution: {integrity: sha512-Kvp459HrV2FEJ1CAsi1Ku+MY3kasH19TFykTz2xWmMeq6bk2NU3XXvfJ+Q61m0xktWwt+1HSYf3JZsTms3aRJg==} - cookie-signature@1.2.2: - resolution: {integrity: sha512-D76uU73ulSXrD1UXF4KE2TMxVVwhsnCgfAyTg9k8P6KGZjlXKrOLe4dJQKI3Bxi5wjesZoFXJWElNWBjPZMbhg==} - engines: {node: '>=6.6.0'} - - cookie@0.7.2: - resolution: {integrity: sha512-yki5XnKuf750l50uGTllt6kKILY4nQ1eNIQatoXEByZ5dWgnKqbnqmTrBE5B4N7lrMJKQ2ytWMiTO2o0v6Ew/w==} - engines: {node: '>= 0.6'} - - cors@2.8.6: - resolution: {integrity: sha512-tJtZBBHA6vjIAaF6EnIaq6laBBP9aq/Y3ouVJjEfoHbRBcHBAHYcMh/w8LDrk2PvIMMq8gmopa5D4V8RmbrxGw==} - engines: {node: '>= 0.10'} - cross-spawn@7.0.6: resolution: {integrity: sha512-uV2QOWP2nWzsy2aMp8aRibhi9dlzF5Hgh5SHaB9OiTGEyDTiJJyx0uy51QXdyWbtAHNua4XJzUKca3OzKUd3vA==} engines: {node: '>= 8'} @@ -2481,10 +2440,6 @@ packages: resolution: {integrity: sha512-8QmQKqEASLd5nx0U1B1okLElbUuuttJ/AnYmRXbbbGDWh6uS208EjD4Xqq/I9wK7u0v6O08XhTWnt5XtEbR6Dg==} engines: {node: '>= 0.4'} - depd@2.0.0: - resolution: {integrity: sha512-g7nH6P6dyDioJogAAGprGpCtVImJhpPk/roCzdb3fIh61/s/nPsfR6onyMwkCAR/OlC3yBC0lESvUoQEAssIrw==} - engines: {node: '>= 0.8'} - detect-libc@2.1.2: resolution: {integrity: sha512-Btj2BOOO83o3WyH59e8MgXsxEQVcarkUOpEYrubB0urwnN10yQ364rsiByU11nZlqWYZm05i/of7io4mzihBtQ==} engines: {node: '>=8'} @@ -2506,9 +2461,6 @@ packages: eastasianwidth@0.2.0: resolution: {integrity: sha512-I88TYZWc9XiYHRQ4/3c5rjjfgkjhLyW2luGIheGERbNQ6OY7yTybanSpDXZa8y7VUP9YmDcYa+eyq4ca7iLqWA==} - ee-first@1.1.1: - resolution: {integrity: sha512-WMwm9LhRUo+WUaRN+vRuETqG89IgZphVSNkdFgeb6sS/E4OrDIN7t48CAewSHXc6C8lefD8KKfr5vY61brQlow==} - electron-to-chromium@1.5.344: resolution: {integrity: sha512-4MxfbmNDm+KPh066EZy+eUnkcDPcZ35wNmOWzFuh/ijvHsve6kbLTLURy88uCNK5FbpN+yk2nQY6BYh1GEt+wg==} @@ -2518,10 +2470,6 @@ packages: emoji-regex@9.2.2: resolution: {integrity: sha512-L18DaJsXSUk2+42pv8mLs5jJT2hqFkFE4j21wOmgbUqsZ2hL72NsUU785g9RXgo3s0ZNgVl42TiHp3ZtOv/Vyg==} - encodeurl@2.0.0: - resolution: {integrity: sha512-Q0n9HRi4m6JuGIV1eFlmvJB7ZEVxu93IrMyiMsGC0lrMJMWzRgx6WGquyfQgZVb31vhGgXnfmPNNXmxnOkRBrg==} - engines: {node: '>= 0.8'} - enhanced-resolve@5.21.0: resolution: {integrity: sha512-otxSQPw4lkOZWkHpB3zaEQs6gWYEsmX4xQF68ElXC/TWvGxGMSGOvoNbaLXm6/cS/fSfHtsEdw90y20PCd+sCA==} engines: {node: '>=10.13.0'} @@ -2567,9 +2515,6 @@ packages: resolution: {integrity: sha512-WUj2qlxaQtO4g6Pq5c29GTcWGDyd8itL8zTlipgECz3JesAiiOKotd8JU6otB3PACgG6xkJUyVhboMS+bje/jA==} engines: {node: '>=6'} - escape-html@1.0.3: - resolution: {integrity: sha512-NiSupZ4OeuGwr68lGIeym/ksIZMJodUGOSCZ/FSnTxcrekbvqrgdUxlJOMpijaKZVjAJrWrGs/6Jy8OMuyj9ow==} - escape-string-regexp@4.0.0: resolution: {integrity: sha512-TtpcNJ3XAzx3Gq8sWRzJaVajRs0uVxA2YAkdb1jm2YkPz4G6egUFAyA3n5vtEIZefPk5Wa4UXbKuS5fKkJWdgA==} engines: {node: '>=10'} @@ -2698,10 +2643,6 @@ packages: resolution: {integrity: sha512-kVscqXk4OCp68SZ0dkgEKVi6/8ij300KBWTJq32P/dYeWTSwK41WyTxalN1eRmA5Z9UU/LX9D7FWSmV9SAYx6g==} engines: {node: '>=0.10.0'} - etag@1.8.1: - resolution: {integrity: sha512-aIL5Fx7mawVa300al2BnEE4iNvo1qETxLrPI/o05L7z6go7fCw1J6EQmbK4FmJ2AS7kgVF/KEZWufBfdClMcPg==} - engines: {node: '>= 0.6'} - eventemitter3@4.0.7: resolution: {integrity: sha512-8guHBZCwKnFhYdHr2ysuRWErTwhoN2X8XELRlrRwpmfeY2jjuUN4taQMsULKUVo1K4DvZl+0pgfyoysHxvmvEw==} @@ -2725,16 +2666,6 @@ packages: resolution: {integrity: sha512-9Be3ZoN4LmYR90tUoVu2te2BsbzHfhJyfEiAVfz7N5/zv+jduIfLrV2xdQXOHbaD6KgpGdO9PRPM1Y4Q9QkPkA==} engines: {node: ^18.19.0 || >=20.5.0} - express-rate-limit@8.5.2: - resolution: {integrity: sha512-5Kb34ipNX694DH48vN9irak1Qx30nb0PLYHXfJgw4YEjiC3ZEmZJhwOp+VfiCYwFzvFTdB9QkArYS5kXa2cx2A==} - engines: {node: '>= 16'} - peerDependencies: - express: '>= 4.11' - - express@5.2.1: - resolution: {integrity: sha512-hIS4idWWai69NezIdRt2xFVofaF4j+6INOpJlVOLDO8zXGpUVEVzIYk12UUi2JzjEzWL3IOAxcTubgz9Po0yXw==} - engines: {node: '>= 18'} - fast-deep-equal@3.1.3: resolution: {integrity: sha512-f3qQ9oQy9j2AhBe/H9VC91wLmKBCCU/gDOnKNAYG5hswO7BLKj09Hc5HYNz9cGI++xlpDCIgDaitVs03ATR84Q==} @@ -2779,10 +2710,6 @@ packages: resolution: {integrity: sha512-YsGpe3WHLK8ZYi4tWDg2Jy3ebRz2rXowDxnld4bkQB00cc/1Zw9AWnC0i9ztDJitivtQvaI9KaLyKrc+hBW0yg==} engines: {node: '>=8'} - finalhandler@2.1.1: - resolution: {integrity: sha512-S8KoZgRZN+a5rNwqTxlZZePjT/4cnm0ROV70LedRHZ0p8u9fRID0hJUZQpkKLzro8LfmC8sx23bY6tVNxv8pQA==} - engines: {node: '>= 18.0.0'} - find-up@5.0.0: resolution: {integrity: sha512-78/PXT1wlLLDgTzDs7sjq9hzz0vXD+zn+7wypEe4fXQxCmdmqfGsEPQxmiCSQI3ajFV91bVSsvNtrJRiW6nGng==} engines: {node: '>=10'} @@ -2801,14 +2728,6 @@ packages: resolution: {integrity: sha512-dKx12eRCVIzqCxFGplyFKJMPvLEWgmNtUrpTiJIR5u97zEhRG8ySrtboPHZXx7daLxQVrl643cTzbab2tkQjxg==} engines: {node: '>= 0.4'} - forwarded@0.2.0: - resolution: {integrity: sha512-buRG0fpBtRHSTCOASe6hD258tEubFoRLb4ZNA6NxMVHNw2gOcwHo9wyablzMzOA5z9xA9L1KNjk/Nt6MT9aYow==} - engines: {node: '>= 0.6'} - - fresh@2.0.0: - resolution: {integrity: sha512-Rx/WycZ60HOaqLKAi6cHRKKI7zxWbJ31MhntmtwMoaTeF7XFH9hhBp8vITaMidfljRQ6eYWCKkaTK+ykVJHP2A==} - engines: {node: '>= 0.8'} - fs-extra@11.3.4: resolution: {integrity: sha512-CTXd6rk/M3/ULNQj8FBqBWHYBVYybQ3VPBw0xGKFe3tuH7ytT6ACnvzpIQ3UZtB8yvUKC2cXn1a+x+5EVQLovA==} engines: {node: '>=14.14'} @@ -2931,14 +2850,6 @@ packages: hls.js@1.6.16: resolution: {integrity: sha512-VSIRpLfRwlAAdGL4wiTucx2ScRipo0ed1FBatWkyt832jC4CReKstga6yIhYVwGu9LOBjuX9wzmRMeQdBJtzEA==} - hono@4.12.27: - resolution: {integrity: sha512-1yrb/+w6HWQJrUCLkJ2IF5jNIPvvFkblV5RNOYl6bV+OA6p9GLcMpHFFGTosSvHvcAUibuUukRqhlYI4z32C7Q==} - engines: {node: '>=16.9.0'} - - http-errors@2.0.1: - resolution: {integrity: sha512-4FbRdAX+bSdmo4AUFuS0WNiPz8NgFt+r8ThgNWmlrjQjt1Q7ZR9+zTlce2859x4KSXrwIsaeTqDoKQmtP8pLmQ==} - engines: {node: '>= 0.8'} - human-signals@2.1.0: resolution: {integrity: sha512-B4FFZ6q/T2jhhksgkbEW3HBvWIfDW85snkQgawt07S7J5QXTk6BkNV+0yAeZrM5QpMAdYlocGoljn0sJ/WQkFw==} engines: {node: '>=10.17.0'} @@ -2947,10 +2858,6 @@ packages: resolution: {integrity: sha512-eKCa6bwnJhvxj14kZk5NCPc6Hb6BdsU9DZcOnmQKSnO1VKrfV0zCvtttPZUsBvjmNDn8rpcJfpwSYnHBjc95MQ==} engines: {node: '>=18.18.0'} - iconv-lite@0.7.3: - resolution: {integrity: sha512-IKXpvIzjnC9XTAUbVBcMfGS0EPaIXtW6v+zr+RRp+hqULEpo0owZax6wyRwPOJbWbzjYspQwusTsfVr0ifh4uQ==} - engines: {node: '>=0.10.0'} - ieee754@1.2.1: resolution: {integrity: sha512-dcyqhDvX1C46lXZcVqCpK+FtMRQVdIMN6/Df5js2zouUsqG7I6sFxitIC+7KYK29KdXOLHdu9zL4sFnoVQnqaA==} @@ -2984,14 +2891,6 @@ packages: resolution: {integrity: sha512-5Hh7Y1wQbvY5ooGgPbDaL5iYLAPzMTUrjMulskHLH6wnv/A+1q5rgEaiuqEjB+oxGXIVZs1FF+R/KPN3ZSQYYg==} engines: {node: '>=12'} - ip-address@10.2.0: - resolution: {integrity: sha512-/+S6j4E9AHvW9SWMSEY9Xfy66O5PWvVEJ08O0y5JGyEKQpojb0K0GKpz/v5HJ/G0vi3D2sjGK78119oXZeE0qA==} - engines: {node: '>= 12'} - - ipaddr.js@1.9.1: - resolution: {integrity: sha512-0KI/607xoxSToH7GjN1FfSbLoU0+btTicjsQSWQlh/hZykN8KpmMf7uYwPW3R+akZ6R/w18ZlXSHBYXiYUPO3g==} - engines: {node: '>= 0.10'} - is-array-buffer@3.0.5: resolution: {integrity: sha512-DDfANUiiG2wC1qawP66qlTugJeL5HyzMpfr8lLK+jMQirGzNod0B12cFB/9q838Ru27sBwfw78/rdoU7RERz6A==} engines: {node: '>= 0.4'} @@ -3076,9 +2975,6 @@ packages: resolution: {integrity: sha512-9UoipoxYmSk6Xy7QFgRv2HDyaysmgSG75TFQs6S+3pDM7ZhKTF/bskZV+0UlABHzKjNVhPjYCLfeZUEg1wXxig==} engines: {node: ^12.20.0 || ^14.13.1 || >=16.0.0} - is-promise@4.0.0: - resolution: {integrity: sha512-hvpoI6korhJMnej285dSg6nu1+e6uxs7zG3BYAm5byqDsgJNWwxzM6z6iZiAgQR4TJ30JmBTOwqZUw3WlyH3AQ==} - is-regex@1.2.1: resolution: {integrity: sha512-MjYsKHO5O7mCsmRGxWcLWheFqN9DJ/2TmngvjKXihe6efViPqc274+Fx/4fYj/r03+ESvBdTXK0V6tA3rgez1g==} engines: {node: '>= 0.4'} @@ -3169,9 +3065,6 @@ packages: json-schema-traverse@1.0.0: resolution: {integrity: sha512-NM8/P9n3XjXhIZn1lLhkFaACTOURQXjWhV4BA/RnOv8xvgqtqpAX9IO4mRQxSx1Rlo4tqzeqb0sOlruaOy3dug==} - json-schema-typed@8.0.2: - resolution: {integrity: sha512-fQhoXdcvc3V28x7C7BMs4P5+kNlgUURe2jmUT1T//oBRMDrqy1QPelJimwZGo7Hg9VPV3EQV5Bnq4hbFy2vetA==} - json-stable-stringify-without-jsonify@1.0.1: resolution: {integrity: sha512-Bdboy+l7tA3OGW6FjyFHWkP5LuByj1Tk33Ljyq0axyzdk9//JSi2u3fP1QSmd1KNwq6VOKYGlAu87CisVir6Pw==} @@ -3324,17 +3217,9 @@ packages: resolution: {integrity: sha512-/IXtbwEk5HTPyEwyKX6hGkYXxM9nbj64B+ilVJnC/R6B0pH5G4V3b0pVbL7DBj4tkhBAppbQUlf6F6Xl9LHu1g==} engines: {node: '>= 0.4'} - media-typer@1.1.0: - resolution: {integrity: sha512-aisnrDP4GNe06UcKFnV5bfMNPBUw4jsLGaWwWfnH3v02GnBuXX2MCVn5RbrWo0j3pczUilYblq7fQ7Nw2t5XKw==} - engines: {node: '>= 0.8'} - memoize-one@5.2.1: resolution: {integrity: sha512-zYiwtZUcYyXKo/np96AGZAckk+FWWsUdJ3cHGGmld7+AhvcWmQyGCYUh1hc4Q/pkOhb65dQR/pqCyK0cOaHz4Q==} - merge-descriptors@2.0.0: - resolution: {integrity: sha512-Snk314V5ayFLhp3fkUREub6WtjBfPdCPY1Ln8/8munuLuiYhsABgBVWsozAG+MWMbVEvcdcpbi9R7ww22l9Q3g==} - engines: {node: '>=18'} - merge-stream@2.0.0: resolution: {integrity: sha512-abv/qOcuPfk3URPfDzmZU1LKmuw8kT+0nIHvKrKgFrwifol/doWcdA4ZqsWQ8ENrFKkd67Mfpo/LovbIUsbt3w==} @@ -3358,10 +3243,6 @@ packages: resolution: {integrity: sha512-lc/aahn+t4/SWV/qcmumYjymLsWfN3ELhpmVuUFjgsORruuZPVSwAQryq+HHGvO/SI2KVX26bx+En+zhM8g8hQ==} engines: {node: '>= 0.6'} - mime-types@3.0.2: - resolution: {integrity: sha512-Lbgzdk0h4juoQ9fCKXW4by0UJqj+nOOrI9MJ1sSj4nI8aI2eo1qmvQEie4VD1glsS250n15LsWsYtCugiStS5A==} - engines: {node: '>=18'} - mimic-fn@2.1.0: resolution: {integrity: sha512-OqbOk5oEQeAZ8WXWydlu9HJjz9WVdEIvamMCcXmuqUYjTknH/sqsWvhQ3vgwKFRR1HpjvNBKQ37nbJgYzGqGcg==} engines: {node: '>=6'} @@ -3406,10 +3287,6 @@ packages: resolution: {integrity: sha512-myRT3DiWPHqho5PrJaIRyaMv2kgYf0mUVgBNOYMuCH5Ki1yEiQaf/ZJuQ62nvpc44wL5WDbTX7yGJi1Neevw8w==} engines: {node: '>= 0.6'} - negotiator@1.0.0: - resolution: {integrity: sha512-8Ofs/AUQh8MaEcrlq5xOX0CQ9ypTF5dl78mjlMNfOK08fzpgTHQRQPBxcPlEtIw0yRpws+Zo/3r+5WRby7u3Gg==} - engines: {node: '>= 0.6'} - next@16.2.3: resolution: {integrity: sha512-9V3zV4oZFza3PVev5/poB9g0dEafVcgNyQ8eTRop8GvxZjV2G15FC5ARuG1eFD42QgeYkzJBJzHghNP8Ad9xtA==} engines: {node: '>=20.9.0'} @@ -3485,17 +3362,10 @@ packages: resolution: {integrity: sha512-gXah6aZrcUxjWg2zR2MwouP2eHlCBzdV4pygudehaKXSGW4v2AsRQUK+lwwXhii6KFZcunEnmSUoYp5CXibxtA==} engines: {node: '>= 0.4'} - on-finished@2.4.1: - resolution: {integrity: sha512-oVlzkg3ENAhCk2zdv7IJwd/QUD4z2RxRwpkcGY8psCVcCYZNq4wYnVWALHM+brtuJjePWiYF/ClmuDr8Ch5+kg==} - engines: {node: '>= 0.8'} - on-headers@1.1.0: resolution: {integrity: sha512-737ZY3yNnXy37FHkQxPzt4UZ2UWPWiCZWLvFZ4fu5cueciegX0zGPnrlY6bwRg4FdQOe9YU8MkmJwGhoMybl8A==} engines: {node: '>= 0.8'} - once@1.4.0: - resolution: {integrity: sha512-lNaJgI+2Q5URQBkccEKHTQOPaXdUxnZZElQTZY0MFUAuaEqe1E+Nyvgdz/aIyNi6Z9MzO5dv1H8n58/GELp3+w==} - onetime@5.1.2: resolution: {integrity: sha512-kbpaSSGJTWdAY5KPVeMOKXSrPtr8C8C7wodJbcsd51jRnmD+GZu8Y0VoU6Dm5Z4vWr0Ig/1NKuWRKf7j5aaYSg==} engines: {node: '>=6'} @@ -3531,10 +3401,6 @@ packages: resolution: {integrity: sha512-TXfryirbmq34y8QBwgqCVLi+8oA3oWx2eAnSn62ITyEhEYaWRlVZ2DvMM9eZbMs/RfxPu/PK/aBLyGj4IrqMHw==} engines: {node: '>=18'} - parseurl@1.3.3: - resolution: {integrity: sha512-CiyeOxFT/JZyN5m0z9PfXw4SCBJ6Sygz1Dpl0wqjlhDEGGBP1GnsUVEL0p63hoG1fcj3fHynXi9NYO4nWOL+qQ==} - engines: {node: '>= 0.8'} - path-exists@4.0.0: resolution: {integrity: sha512-ak9Qy5Q7jYb2Wwcey5Fpvg2KoAc/ZIhLSLOSBmRmygPsGwkVVt0fZa0qrtMz+m6tJTAHfZQ8FnmB4MG4LWy7/w==} engines: {node: '>=8'} @@ -3556,9 +3422,6 @@ packages: path-to-regexp@3.3.0: resolution: {integrity: sha512-qyCH421YQPS2WFDxDjftfc1ZR5WKQzVzqsp4n9M2kQhVOo/ByahFoUNJfl58kOcEGfQ//7weFTDhm+ss8Ecxgw==} - path-to-regexp@8.4.2: - resolution: {integrity: sha512-qRcuIdP69NPm4qbACK+aDogI5CBDMi1jKe0ry5rSQJz8JVLsC7jV8XpiJjGRLLol3N+R5ihGYcrPLTno6pAdBA==} - picocolors@1.1.1: resolution: {integrity: sha512-xceH2snhtb5M9liqDsmEw56le376mTZkEX/jEb/RxNFyegNul7eNslCXP9FDj/Lcu0X8KEyMceP2ntpaHrDEVA==} @@ -3607,18 +3470,10 @@ packages: prop-types@15.8.1: resolution: {integrity: sha512-oj87CgZICdulUohogVAR7AjlC0327U4el4L6eAvOqCeudMDVU0NThNaV+b9Df4dXgSP1gXMTnPdhfe/2qDH5cg==} - proxy-addr@2.0.7: - resolution: {integrity: sha512-llQsMLSUDUPT44jdrU/O37qlnifitDP+ZwrmmZcoSKyLKvtZxpyV0n2/bD/N4tBAAZ/gJEdZU7KMraoK1+XYAg==} - engines: {node: '>= 0.10'} - punycode@2.3.1: resolution: {integrity: sha512-vYt7UD1U9Wg6138shLtLOvdAu+8DsC/ilFtEVHcH+wydcSpNE20AfSOduf6MkRFahL5FY7X1oU7nKVZFtfq8Fg==} engines: {node: '>=6'} - qs@6.15.3: - resolution: {integrity: sha512-O9gl3zCl5h5blw1KGUzQKhA5oUXSl8rwUIM5o0S3nCXMliSvy5Dzx7/DJcI+SwgICv+IneSZwhBh1oSyEHA71A==} - engines: {node: '>=0.6'} - queue-microtask@1.2.3: resolution: {integrity: sha512-NuaNSa6flKT5JaSYQzJok04JzTL1CA6aGhv5rfLW3PgqA+M2ChpZQnAC8h8i4ZFkBS8X5RqkDBHA7r4hej3K9A==} @@ -3639,14 +3494,6 @@ packages: resolution: {integrity: sha512-kA5WQoNVo4t9lNx2kQNFCxKeBl5IbbSNBl1M/tLkw9WCn+hxNBAW5Qh8gdhs63CJnhjJ2zQWFoqPJP2sK1AV5A==} engines: {node: '>= 0.6'} - range-parser@1.3.0: - resolution: {integrity: sha512-hek2mFQpPuI4E1BBKrSto+BU3e3x4xuarsbiwr3+lf7p44juvFMV0XFWQAP3xUyqXA4RrXLIoaSUGbSt056ZMw==} - engines: {node: '>= 0.6'} - - raw-body@3.0.2: - resolution: {integrity: sha512-K5zQjDllxWkf7Z5xJdV0/B0WTNqx6vxG70zJE4N0kBs4LovmEYWJzQGxC9bS9RAKu3bgM40lrd5zoLJ12MQ5BA==} - engines: {node: '>= 0.10'} - rc@1.2.8: resolution: {integrity: sha512-y3bGgqKj3QBdxLbLkomlohkvsA8gdAiUQlSBJnBhfn+BPxg4bc62d8TcBW15wavDfgexCgccckhcZvywyQYPOw==} hasBin: true @@ -3766,10 +3613,6 @@ packages: resolution: {integrity: sha512-g6QUff04oZpHs0eG5p83rFLhHeV00ug/Yf9nZM6fLeUrPguBTkTQOdpAWWspMh55TZfVQDPaN3NQJfbVRAxdIw==} engines: {iojs: '>=1.0.0', node: '>=0.10.0'} - router@2.2.0: - resolution: {integrity: sha512-nLTrUKm2UyiL7rlhapu/Zl45FwNgkZGaCpZbIHajDYgwlJCOzLSk+cIPAnsEqV955GjILJnKbdQC1nVPz+gAYQ==} - engines: {node: '>= 18'} - run-parallel@1.2.0: resolution: {integrity: sha512-5l4VyZR86LZ/lDxZTR6jqL8AFE2S0IFLMP26AbjsLVADxHdhB/c0GUsH+y39UfCi3dzz8OlQuPmnaJOMoDHQBA==} @@ -3788,9 +3631,6 @@ packages: resolution: {integrity: sha512-x/+Cz4YrimQxQccJf5mKEbIa1NzeCRNI5Ecl/ekmlYaampdNLPalVyIcCZNNH3MvmqBugV5TMYZXv0ljslUlaw==} engines: {node: '>= 0.4'} - safer-buffer@2.1.2: - resolution: {integrity: sha512-YZo3K82SD7Riyi0E1EQPojLz7kpepnSQI9IyPbHHg1XXXevb5dJI7tpyN2ADxGcQbHG7vcyRHk0cbwqcQriUtg==} - scheduler@0.27.0: resolution: {integrity: sha512-eNv+WrVbKu1f3vbYJT/xtiF5syA5HPIMtf9IgY/nKg0sWqzAUEvqY/xm7OcZc/qafLx/iO9FgOmeSAp4v5ti/Q==} @@ -3803,17 +3643,9 @@ packages: engines: {node: '>=10'} hasBin: true - send@1.2.1: - resolution: {integrity: sha512-1gnZf7DFcoIcajTjTwjwuDjzuz4PPcY2StKPlsGAQ1+YH20IRVrBaXSWmdjowTJ6u8Rc01PoYOGHXfP1mYcZNQ==} - engines: {node: '>= 18'} - serve-handler@6.1.7: resolution: {integrity: sha512-CinAq1xWb0vR3twAv9evEU8cNWkXCb9kd5ePAHUKJBkOsUpR1wt/CvGdeca7vqumL1U5cSaeVQ6zZMxiJ3yWsg==} - serve-static@2.2.1: - resolution: {integrity: sha512-xRXBn0pPqQTVQiC8wyQrKs2MOlX24zQ0POGaj0kultvoOCstBQM5yvOhAVSUwOMjQtTvsPWoNCHfPGwaaQJhTw==} - engines: {node: '>= 18'} - serve@14.2.6: resolution: {integrity: sha512-QEjUSA+sD4Rotm1znR8s50YqA3kYpRGPmtd5GlFxbaL9n/FdUNbqMhxClqdditSk0LlZyA/dhud6XNRTOC9x2Q==} engines: {node: '>= 14'} @@ -3831,9 +3663,6 @@ packages: resolution: {integrity: sha512-RJRdvCo6IAnPdsvP/7m6bsQqNnn1FCBX5ZNtFL98MmFF/4xAIJTIg1YbHW5DC2W5SKZanrC6i4HsJqlajw/dZw==} engines: {node: '>= 0.4'} - setprototypeof@1.2.0: - resolution: {integrity: sha512-E5LDX7Wrp85Kil5bhZv46j8jOeboKq5JMmYM3gVGdGH8xFpPWXUMsNrlODCrkoxMEeNi/XZIwuRvY4XNwYMJpw==} - sharp@0.34.5: resolution: {integrity: sha512-Ou9I5Ft9WNcCbXrU9cMgPBcCK8LiwLqcbywW3t4oDV37n1pzpuNLsYiAV8eODnjbtQlSDwZ2cUEeQz4E54Hltg==} engines: {node: ^18.17.0 || ^20.3.0 || >=21.0.0} @@ -3862,10 +3691,6 @@ packages: resolution: {integrity: sha512-ZX99e6tRweoUXqR+VBrslhda51Nh5MTQwou5tnUDgbtyM0dBgmhEDtWGP/xbKn6hqfPRHujUNwz5fy/wbbhnpw==} engines: {node: '>= 0.4'} - side-channel@1.1.1: - resolution: {integrity: sha512-6x6dK6zJdpTzF4sQeNYxwtvBzf6Eg4GtlesS94HOvTudUeyK2WXAaIfmDgsyslYrRBeFIlsi54AYsFGUuhmvrQ==} - engines: {node: '>= 0.4'} - signal-exit@3.0.7: resolution: {integrity: sha512-wnD2ZE+l+SPC/uoS0vXeE9L1+0wuaMqKlfz9AMUo38JsyLSBWSFcHR1Rri62LZc12vLr1gb3jl7iwQhgwpAbGQ==} @@ -3886,10 +3711,6 @@ packages: stable-hash@0.0.5: resolution: {integrity: sha512-+L3ccpzibovGXFK+Ap/f8LOS0ahMrHTf3xu7mMLSpEGU0EO9ucaysSylKo9eRDFNhWve/y275iPmIZ4z39a9iA==} - statuses@2.0.2: - resolution: {integrity: sha512-DvEy55V3DB7uknRo+4iOGT5fP1slR8wQohVdknigZPMpMstaKJQWhwiYBACJE3Ul2pTnATihhBYnRhZQHGBiRw==} - engines: {node: '>= 0.8'} - stop-iteration-iterator@1.1.0: resolution: {integrity: sha512-eLoXW/DHyl62zxY4SCaIgnRhuMr6ri4juEYARS8E6sCEqzKpOiE521Ucofdx+KnDZl5xmvGYaaKCk5FEOxJCoQ==} engines: {node: '>= 0.4'} @@ -4001,10 +3822,6 @@ packages: resolution: {integrity: sha512-65P7iz6X5yEr1cwcgvQxbbIw7Uk3gOy5dIdtZ4rDveLqhrdJP+Li/Hx6tyK0NEb+2GCyneCMJiGqrADCSNk8sQ==} engines: {node: '>=8.0'} - toidentifier@1.0.1: - resolution: {integrity: sha512-o5sSPKEkg/DIQNmH43V0/uerLrpzVedkUh8tGNvaeXpfpuwjKenlSox/2O/BTlZUtEe+JG7s5YhEz608PlAHRA==} - engines: {node: '>=0.6'} - ts-api-utils@2.5.0: resolution: {integrity: sha512-OJ/ibxhPlqrMM0UiNHJ/0CKQkoKF243/AEmplt3qpRgkW8VG7IfOS41h7V8TjITqdByHzrjcS/2si+y4lIh8NA==} engines: {node: '>=18.12'} @@ -4033,10 +3850,6 @@ packages: resolution: {integrity: sha512-RAH822pAdBgcNMAfWnCBU3CFZcfZ/i1eZjwFU/dsLKumyuuP3niueg2UAukXYF0E2AAoc82ZSSf9J0WQBinzHA==} engines: {node: '>=12.20'} - type-is@2.1.0: - resolution: {integrity: sha512-faYHw0anBbc/kWF3zFTEnxSFOAGUX9GFbOBthvDdLsIlEoWOFOtS0zgCiQYwIskL9iGXZL3kAXD8OoZ4GmMATA==} - engines: {node: '>= 18'} - typed-array-buffer@1.0.3: resolution: {integrity: sha512-nAYYwfY3qnzX30IkA6AQZjVbtK6duGontcQm1WSG1MD94YLqK0515GNApXkoxKOWMusVssAHWLh9SeaoefYFGw==} engines: {node: '>= 0.4'} @@ -4080,10 +3893,6 @@ packages: resolution: {integrity: sha512-gptHNQghINnc/vTGIk0SOFGFNXw7JVrlRUtConJRlvaw6DuX0wO5Jeko9sWrMBhh+PsYAZ7oXAiOnf/UKogyiw==} engines: {node: '>= 10.0.0'} - unpipe@1.0.0: - resolution: {integrity: sha512-pjy2bYhSsufwWlKwPc+l3cN7+wuJlK6uz0YdJEOlQDbl6jo/YlPi4mb8agUkVC8BF7V8NuzeyPNqRksA3hztKQ==} - engines: {node: '>= 0.8'} - unrs-resolver@1.11.1: resolution: {integrity: sha512-bSjt9pjaEBnNiGgc9rUiHGKv5l4/TGzDmYw3RhnkJGtLhbnnA/5qJj7x3dNDCRx/PJxu774LlH8lCOlB4hEfKg==} @@ -4165,9 +3974,6 @@ packages: resolution: {integrity: sha512-si7QWI6zUMq56bESFvagtmzMdGOtoxfR+Sez11Mobfc7tm+VkUckk9bW2UeffTGVUbOksxmSw0AA2gs8g71NCQ==} engines: {node: '>=12'} - wrappy@1.0.2: - resolution: {integrity: sha512-l4Sp/DRseor9wL6EvV2+TuQn63dMkPjZ/sp9XkghTEbV9KlPS1xUsZ3u7/IQO4wxtcFB4bgpQPRcR3QCvezPcQ==} - yallist@3.1.1: resolution: {integrity: sha512-a4UGQaWPH59mOXUYnAG2ewncQS4i4F43Tv3JoAM+s2VDAmS9NsK8GpDMLrCHPksFT7h3K6TOoUNn2pb7RoXx4g==} @@ -4183,11 +3989,6 @@ packages: resolution: {integrity: sha512-CzhO+pFNo8ajLM2d2IW/R93ipy99LWjtwblvC1RsoSUMZgyLbYFr221TnSNT7GjGdYui6P459mw9JH/g/zW2ug==} engines: {node: '>=18'} - zod-to-json-schema@3.25.2: - resolution: {integrity: sha512-O/PgfnpT1xKSDeQYSCfRI5Gy3hPf91mKVDuYLUHZJMiDFptvP41MSnWofm8dnCm0256ZNfZIM7DSzuSMAFnjHA==} - peerDependencies: - zod: ^3.25.28 || ^4 - zod-validation-error@4.0.2: resolution: {integrity: sha512-Q6/nZLe6jxuU80qb/4uJ4t5v2VEZ44lzQjPDhYJNztRQ4wyWc6VF3D3Kb/fAuPetZQnhS3hnajCf9CsWesghLQ==} engines: {node: '>=18.0.0'} @@ -4637,10 +4438,6 @@ snapshots: '@harperfast/extended-iterable@1.0.3': {} - '@hono/node-server@1.19.14(hono@4.12.27)': - dependencies: - hono: 4.12.27 - '@humanfs/core@0.19.2': dependencies: '@humanfs/types': 0.15.0 @@ -4794,27 +4591,24 @@ snapshots: '@lmdb/lmdb-win32-x64@3.5.4': optional: true - '@modelcontextprotocol/sdk@1.29.0(zod@4.3.6)': + '@modelcontextprotocol/client@2.0.0': dependencies: - '@hono/node-server': 1.19.14(hono@4.12.27) - ajv: 8.18.0 - ajv-formats: 3.0.1(ajv@8.18.0) - content-type: 1.0.5 - cors: 2.8.6 + '@modelcontextprotocol/core': 2.0.0 cross-spawn: 7.0.6 eventsource: 3.0.7 eventsource-parser: 3.1.0 - express: 5.2.1 - express-rate-limit: 8.5.2(express@5.2.1) - hono: 4.12.27 jose: 6.2.3 - json-schema-typed: 8.0.2 pkce-challenge: 5.0.1 - raw-body: 3.0.2 zod: 4.3.6 - zod-to-json-schema: 3.25.2(zod@4.3.6) - transitivePeerDependencies: - - supports-color + + '@modelcontextprotocol/core@2.0.0': + dependencies: + zod: 4.3.6 + + '@modelcontextprotocol/server@2.0.0': + dependencies: + '@modelcontextprotocol/core': 2.0.0 + zod: 4.3.6 '@msgpackr-extract/msgpackr-extract-darwin-arm64@3.0.3': optional: true @@ -5975,21 +5769,12 @@ snapshots: '@zeit/schemas@2.36.0': {} - accepts@2.0.0: - dependencies: - mime-types: 3.0.2 - negotiator: 1.0.0 - acorn-jsx@5.3.2(acorn@8.16.0): dependencies: acorn: 8.16.0 acorn@8.16.0: {} - ajv-formats@3.0.1(ajv@8.18.0): - optionalDependencies: - ajv: 8.18.0 - ajv@6.15.0: dependencies: fast-deep-equal: 3.1.3 @@ -6117,20 +5902,6 @@ snapshots: baseline-browser-mapping@2.10.23: {} - body-parser@2.3.0: - dependencies: - bytes: 3.1.2 - content-type: 2.0.0 - debug: 4.4.3 - http-errors: 2.0.1 - iconv-lite: 0.7.3 - on-finished: 2.4.1 - qs: 6.15.3 - raw-body: 3.0.2 - type-is: 2.1.0 - transitivePeerDependencies: - - supports-color - bowser@2.14.1: {} boxen@7.0.0: @@ -6262,23 +6033,8 @@ snapshots: content-disposition@0.5.2: {} - content-disposition@1.1.0: {} - - content-type@1.0.5: {} - - content-type@2.0.0: {} - convert-source-map@2.0.0: {} - cookie-signature@1.2.2: {} - - cookie@0.7.2: {} - - cors@2.8.6: - dependencies: - object-assign: 4.1.1 - vary: 1.1.2 - cross-spawn@7.0.6: dependencies: path-key: 3.1.1 @@ -6377,8 +6133,6 @@ snapshots: has-property-descriptors: 1.0.2 object-keys: 1.1.1 - depd@2.0.0: {} - detect-libc@2.1.2: {} detect-node-es@1.1.0: {} @@ -6400,16 +6154,12 @@ snapshots: eastasianwidth@0.2.0: {} - ee-first@1.1.1: {} - electron-to-chromium@1.5.344: {} emoji-regex@8.0.0: {} emoji-regex@9.2.2: {} - encodeurl@2.0.0: {} - enhanced-resolve@5.21.0: dependencies: graceful-fs: 4.2.11 @@ -6547,8 +6297,6 @@ snapshots: escalade@3.2.0: {} - escape-html@1.0.3: {} - escape-string-regexp@4.0.0: {} escape-string-regexp@5.0.0: {} @@ -6805,8 +6553,6 @@ snapshots: esutils@2.0.3: {} - etag@1.8.1: {} - eventemitter3@4.0.7: {} events@3.3.0: {} @@ -6844,44 +6590,6 @@ snapshots: strip-final-newline: 4.0.0 yoctocolors: 2.1.2 - express-rate-limit@8.5.2(express@5.2.1): - dependencies: - express: 5.2.1 - ip-address: 10.2.0 - - express@5.2.1: - dependencies: - accepts: 2.0.0 - body-parser: 2.3.0 - content-disposition: 1.1.0 - content-type: 1.0.5 - cookie: 0.7.2 - cookie-signature: 1.2.2 - debug: 4.4.3 - depd: 2.0.0 - encodeurl: 2.0.0 - escape-html: 1.0.3 - etag: 1.8.1 - finalhandler: 2.1.1 - fresh: 2.0.0 - http-errors: 2.0.1 - merge-descriptors: 2.0.0 - mime-types: 3.0.2 - on-finished: 2.4.1 - once: 1.4.0 - parseurl: 1.3.3 - proxy-addr: 2.0.7 - qs: 6.15.3 - range-parser: 1.3.0 - router: 2.2.0 - send: 1.2.1 - serve-static: 2.2.1 - statuses: 2.0.2 - type-is: 2.1.0 - vary: 1.1.2 - transitivePeerDependencies: - - supports-color - fast-deep-equal@3.1.3: {} fast-equals@5.4.0: {} @@ -6920,17 +6628,6 @@ snapshots: dependencies: to-regex-range: 5.0.1 - finalhandler@2.1.1: - dependencies: - debug: 4.4.3 - encodeurl: 2.0.0 - escape-html: 1.0.3 - on-finished: 2.4.1 - parseurl: 1.3.3 - statuses: 2.0.2 - transitivePeerDependencies: - - supports-color - find-up@5.0.0: dependencies: locate-path: 6.0.0 @@ -6949,10 +6646,6 @@ snapshots: dependencies: is-callable: 1.2.7 - forwarded@0.2.0: {} - - fresh@2.0.0: {} - fs-extra@11.3.4: dependencies: graceful-fs: 4.2.11 @@ -7070,24 +6763,10 @@ snapshots: hls.js@1.6.16: {} - hono@4.12.27: {} - - http-errors@2.0.1: - dependencies: - depd: 2.0.0 - inherits: 2.0.4 - setprototypeof: 1.2.0 - statuses: 2.0.2 - toidentifier: 1.0.1 - human-signals@2.1.0: {} human-signals@8.0.1: {} - iconv-lite@0.7.3: - dependencies: - safer-buffer: 2.1.2 - ieee754@1.2.1: {} ignore@5.3.2: {} @@ -7113,10 +6792,6 @@ snapshots: internmap@2.0.3: {} - ip-address@10.2.0: {} - - ipaddr.js@1.9.1: {} - is-array-buffer@3.0.5: dependencies: call-bind: 1.0.9 @@ -7198,8 +6873,6 @@ snapshots: is-port-reachable@4.0.0: {} - is-promise@4.0.0: {} - is-regex@1.2.1: dependencies: call-bound: 1.0.4 @@ -7280,8 +6953,6 @@ snapshots: json-schema-traverse@1.0.0: {} - json-schema-typed@8.0.2: {} - json-stable-stringify-without-jsonify@1.0.1: {} json5@1.0.2: @@ -7416,12 +7087,8 @@ snapshots: math-intrinsics@1.1.0: {} - media-typer@1.1.0: {} - memoize-one@5.2.1: {} - merge-descriptors@2.0.0: {} - merge-stream@2.0.0: {} merge2@1.4.1: {} @@ -7439,10 +7106,6 @@ snapshots: dependencies: mime-db: 1.33.0 - mime-types@3.0.2: - dependencies: - mime-db: 1.54.0 - mimic-fn@2.1.0: {} minimatch@10.2.5: @@ -7483,8 +7146,6 @@ snapshots: negotiator@0.6.4: {} - negotiator@1.0.0: {} - next@16.2.3(@babel/core@7.29.0)(@playwright/test@1.59.1)(react-dom@19.2.4(react@19.2.4))(react@19.2.4): dependencies: '@next/env': 16.2.3 @@ -7576,16 +7237,8 @@ snapshots: define-properties: 1.2.1 es-object-atoms: 1.1.1 - on-finished@2.4.1: - dependencies: - ee-first: 1.1.1 - on-headers@1.1.0: {} - once@1.4.0: - dependencies: - wrappy: 1.0.2 - onetime@5.1.2: dependencies: mimic-fn: 2.1.0 @@ -7625,8 +7278,6 @@ snapshots: parse-ms@4.0.0: {} - parseurl@1.3.3: {} - path-exists@4.0.0: {} path-is-inside@1.0.2: {} @@ -7639,8 +7290,6 @@ snapshots: path-to-regexp@3.3.0: {} - path-to-regexp@8.4.2: {} - picocolors@1.1.1: {} picomatch@2.3.2: {} @@ -7683,18 +7332,8 @@ snapshots: object-assign: 4.1.1 react-is: 16.13.1 - proxy-addr@2.0.7: - dependencies: - forwarded: 0.2.0 - ipaddr.js: 1.9.1 - punycode@2.3.1: {} - qs@6.15.3: - dependencies: - es-define-property: 1.0.1 - side-channel: 1.1.1 - queue-microtask@1.2.3: {} radix-ui@1.6.0(@types/react-dom@19.2.3(@types/react@19.2.14))(@types/react@19.2.14)(react-dom@19.2.4(react@19.2.4))(react@19.2.4): @@ -7762,15 +7401,6 @@ snapshots: range-parser@1.2.0: {} - range-parser@1.3.0: {} - - raw-body@3.0.2: - dependencies: - bytes: 3.1.2 - http-errors: 2.0.1 - iconv-lite: 0.7.3 - unpipe: 1.0.0 - rc@1.2.8: dependencies: deep-extend: 0.6.0 @@ -7913,16 +7543,6 @@ snapshots: reusify@1.1.0: {} - router@2.2.0: - dependencies: - debug: 4.4.3 - depd: 2.0.0 - is-promise: 4.0.0 - parseurl: 1.3.3 - path-to-regexp: 8.4.2 - transitivePeerDependencies: - - supports-color - run-parallel@1.2.0: dependencies: queue-microtask: 1.2.3 @@ -7948,30 +7568,12 @@ snapshots: es-errors: 1.3.0 is-regex: 1.2.1 - safer-buffer@2.1.2: {} - scheduler@0.27.0: {} semver@6.3.1: {} semver@7.7.4: {} - send@1.2.1: - dependencies: - debug: 4.4.3 - encodeurl: 2.0.0 - escape-html: 1.0.3 - etag: 1.8.1 - fresh: 2.0.0 - http-errors: 2.0.1 - mime-types: 3.0.2 - ms: 2.1.3 - on-finished: 2.4.1 - range-parser: 1.3.0 - statuses: 2.0.2 - transitivePeerDependencies: - - supports-color - serve-handler@6.1.7: dependencies: bytes: 3.0.0 @@ -7982,15 +7584,6 @@ snapshots: path-to-regexp: 3.3.0 range-parser: 1.2.0 - serve-static@2.2.1: - dependencies: - encodeurl: 2.0.0 - escape-html: 1.0.3 - parseurl: 1.3.3 - send: 1.2.1 - transitivePeerDependencies: - - supports-color - serve@14.2.6: dependencies: '@zeit/schemas': 2.36.0 @@ -8029,8 +7622,6 @@ snapshots: es-errors: 1.3.0 es-object-atoms: 1.1.1 - setprototypeof@1.2.0: {} - sharp@0.34.5: dependencies: '@img/colour': 1.1.0 @@ -8097,14 +7688,6 @@ snapshots: side-channel-map: 1.0.1 side-channel-weakmap: 1.0.2 - side-channel@1.1.1: - dependencies: - es-errors: 1.3.0 - object-inspect: 1.13.4 - side-channel-list: 1.0.1 - side-channel-map: 1.0.1 - side-channel-weakmap: 1.0.2 - signal-exit@3.0.7: {} signal-exit@4.1.0: {} @@ -8118,8 +7701,6 @@ snapshots: stable-hash@0.0.5: {} - statuses@2.0.2: {} - stop-iteration-iterator@1.1.0: dependencies: es-errors: 1.3.0 @@ -8244,8 +7825,6 @@ snapshots: dependencies: is-number: 7.0.0 - toidentifier@1.0.1: {} - ts-api-utils@2.5.0(typescript@5.9.3): dependencies: typescript: 5.9.3 @@ -8274,12 +7853,6 @@ snapshots: type-fest@2.19.0: {} - type-is@2.1.0: - dependencies: - content-type: 2.0.0 - media-typer: 1.1.0 - mime-types: 3.0.2 - typed-array-buffer@1.0.3: dependencies: call-bound: 1.0.4 @@ -8339,8 +7912,6 @@ snapshots: universalify@2.0.1: {} - unpipe@1.0.0: {} - unrs-resolver@1.11.1: dependencies: napi-postinstall: 0.3.4 @@ -8475,8 +8046,6 @@ snapshots: string-width: 5.1.2 strip-ansi: 7.2.0 - wrappy@1.0.2: {} - yallist@3.1.1: {} yocto-queue@0.1.0: {} @@ -8485,10 +8054,6 @@ snapshots: yoctocolors@2.1.2: {} - zod-to-json-schema@3.25.2(zod@4.3.6): - dependencies: - zod: 4.3.6 - zod-validation-error@4.0.2(zod@4.3.6): dependencies: zod: 4.3.6