commit a2de17b53c2a809018e00d2366dad3b71f364727
parent 58569da848f69b3b3ba75c0c6b4a0239d687f3b8
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Sat, 26 Sep 2026 14:55:20 -0400
Merge mcp/fetch-clip — release 10 slice M: fetch_clip, the MCP asks the editor for clip media (window or full: true) through POST /api/media/fetch-window; a Rumble embed id maps to the slug dir via the record's webpageUrl; the ask/sweep plans send media through the tool and never a hand-run yt-dlp; 15 s per-request timeout, progress notifications per poll, the job id survives a poll failure
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Diffstat:
13 files changed, 2353 insertions(+), 21 deletions(-)
diff --git a/AGENTS.md b/AGENTS.md
@@ -52,11 +52,17 @@ Register the MCP server against a public instance and the corpus is readable ove
```sh
claude mcp add archilyzer \
--env TRANSCRIPT_SITE_URL=https://jeralyzer.pages.dev \
+ --env ARCHILYZER_EDITOR_URL=http://localhost:3001 \
+ --env WORKER_TOKEN=… \
-- pnpm -C "$PWD" --filter yt-dlp-transcript-mcp exec tsx src/index.ts
```
+The two editor lines are optional: they let `fetch_clip` ask a local editor for clip
+media (`WORKER_TOKEN` is the editor's own, from `editor/.env`).
+
`TRANSCRIPT_HUB_URL` federates several sites; `TRANSCRIPT_LOCAL_DIR` reads a local
-build off disk. The server never writes to an archive.
+build off disk. The server never writes to an archive — `fetch_clip` asks the editor,
+and the editor writes.
**Register it as `archilyzer`.** `.claude/commands/{ask,sweep}.md` are tracked in git
and call `mcp__archilyzer__ask_plan` / `mcp__archilyzer__sweep_plan`. That tool name
@@ -74,9 +80,14 @@ written to a file.
## Clips and report-to-video
-The high-value loop for a repo with no corpus: point the MCP at a public instance, ask
-about a subject, then use `yt-dlp --download-sections` to pull **just the cited
-seconds** rather than whole videos. Searching text first is what makes fetching cheap.
+The high-value loop: point the MCP at a public instance, ask about a subject, then pull
+**just the cited seconds** rather than whole videos. Searching text first is what makes
+fetching cheap. Clip media for a cited moment goes through the `fetch_clip` MCP tool
+(→ the editor's `POST /api/media/fetch-window`, paced, cookie-aware and provenanced),
+never `yt-dlp --download-sections` by hand; that command is only the fallback for a
+machine with no editor. The editor fetches only for a channel it already archives (a
+`transcripts/channels/<slug>/`), so an MCP pointed at a public site with a fresh editor
+gets a 404 `Channel "<slug>" not found` on every clip.
`umtool/report-to-video/` renders a cited sweep report to an mp4. What it needs:
diff --git a/README.md b/README.md
@@ -68,10 +68,13 @@ Claude Code:
```bash
claude mcp add archilyzer \
--env TRANSCRIPT_SITE_URL=https://jeralyzer.pages.dev \
+ --env ARCHILYZER_EDITOR_URL=http://localhost:3001 \
+ --env WORKER_TOKEN=… \
-- pnpm -C /ABS/PATH/TO/this/repo --filter yt-dlp-transcript-mcp exec tsx src/index.ts
```
-Use `TRANSCRIPT_HUB_URL` instead to federate a whole hub of sites, or
+The two editor lines are optional: they let `fetch_clip` ask a local editor for clip
+media. Use `TRANSCRIPT_HUB_URL` instead to federate a whole hub of sites, or
`TRANSCRIPT_LOCAL_DIR` to read a local build off disk.
It gives a client full-text search with clickable second-level citations, complete
@@ -318,10 +321,15 @@ quietly sampling.
pnpm install
claude mcp add archilyzer \
--env TRANSCRIPT_SITE_URL=https://jeralyzer.pages.dev \
+ --env ARCHILYZER_EDITOR_URL=http://localhost:3001 \
+ --env WORKER_TOKEN=… \
-- pnpm -C "$PWD" --filter yt-dlp-transcript-mcp exec tsx src/index.ts
claude # then try: /ask what has he said about magic tournaments?
```
+The two editor lines are optional: they let `fetch_clip` ask a local editor for clip
+media (`WORKER_TOKEN` is the editor's own). Leave them out for research alone.
+
> **Register the server as `archilyzer`.** The shipped commands call
> `mcp__archilyzer__ask_plan` / `mcp__archilyzer__sweep_plan`, and that tool name
> embeds the server name **as you registered it**. Under any other name the commands
@@ -338,24 +346,34 @@ no GPU, nothing hosted.** Point `TRANSCRIPT_SITE_URL` at any instance, or use
### Clips and video: download the moment, not the movie
This is the part that makes an archive worth more than a search box. A transcript gives
-you the **exact second** something was said, and `yt-dlp --download-sections` can fetch
-just those seconds. So the loop is:
+you the **exact second** something was said, and the editor can fetch just those
+seconds. So the loop is:
1. Point the MCP server at an instance and `/ask` or `/sweep` about a subject.
2. Get back citations that resolve to precise moments in real recordings.
-3. Pull **just those clips** with yt-dlp — seconds of media, not hours.
+3. Ask for **just those clips** with the `fetch_clip` MCP tool — the editor fetches the
+ window through its paced, cookie-aware, provenanced job; the file lands in the
+ corpus beside the video (`channels/<slug>/data/<id>/clips/`). Seconds of media, not
+ hours. `full: true` fetches the whole recording into the saved-video store instead
+ (it needs a video the editor already knows).
4. Optionally, render them into a finished video.
You are never downloading a back catalogue to find a quote. You search text, then fetch
the few seconds you actually want. Steps 1–2 need no corpus and no media at all; step 3
-needs `yt-dlp`, and step 4 adds `ffmpeg`/`ffprobe` and **ImageMagick with Pango** for
-the chrome.
-
-The clips **are** kept on disk — they land under the report's own `out/` directory
-(`clips-raw/`, `segments/`, `cards/`, and the finished `<slug>.mp4`) and are reused on
-a rebuild. "No corpus" means you are not mirroring a channel's back catalogue, not that
-nothing is stored: your disk use scales with the clips you actually pull, which for a
-report is minutes of video rather than years of it.
+needs a local Archilyzer editor (the MCP registered with `ARCHILYZER_EDITOR_URL` and
+`WORKER_TOKEN`) that already archives the cited channel: the editor fetches only for a
+channel it has under its `transcripts/`, so an MCP pointed at a public site with a fresh
+editor gets a 404 (`Channel "<slug>" not found`) on every clip. Step 4 adds
+`ffmpeg`/`ffprobe` and **ImageMagick with Pango** for the chrome. With no editor (a
+public-only setup), or none that archives the channel, the fallback is running yt-dlp
+yourself — `yt-dlp --download-sections` fetches just the cited seconds.
+
+A window fetched through the editor is kept in the corpus and reused by every later ask
+and render; the editor prunes clips by age. A render keeps its own `segments/`,
+`cards/` and the finished `<slug>.mp4` under the report's `out/`. "No corpus" means you
+are not mirroring a channel's back catalogue, not that nothing is stored: your disk use
+scales with the clips you actually pull, which for a report is minutes of video rather
+than years of it.
### Rendering a report to video
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -6,6 +6,7 @@
- **Jobs a restart left queued are settled even when a drive hangs.** At boot the editor settles those jobs after it has checked where its storage locations are. A hung network mount could stall that check forever, and the jobs then stayed "queued" on `/jobs`. The settling now waits at most 60 seconds, logs `[boot] storage pass still running after 60 s …` and carries on. The storage check keeps running and logs when it ends. A job re-queued for a channel on the hung drive itself still waits for that drive, and any re-queued after it wait too.
- **`/jobs` says why a job was cancelled at boot.** A job the boot settled shows its reason under its status on `/jobs` and as *Cancelled because* on its own page: for example "server restarted; the scheduler re-derives syncs" or "superseded by a newer queued job (…)". The reason used to be only in the job's log.
- **The server log says how often a queued job skips its page refresh.** When a queued job finishes outside any request, the editor skips its page refresh and notes it in the log. The note used to appear once and never again. Now the first one after a quiet spell is logged at once, any more in the next 10 minutes are counted, and one line at the end gives the count, with a running total.
+- **An agent working through the MCP server asks the editor for a clip instead of running yt-dlp.** The MCP server has a new tool, `fetch_clip`. Given a citation's channel, video id, start and end and a one-line reason, it asks the local editor for that window through `POST /api/media/fetch-window`: the same paced, cookie-aware job umtool uses, which records who asked and why beside the file. It answers with the file's path in the corpus (`channels/<slug>/data/<id>/clips/`). The window is the cited span with 3 seconds either side, at most 15 minutes. `full: true` asks for the whole recording instead, which lands in the saved-video store and needs a video the editor already knows. A Rumble citation's id (the embed id the archive publishes) is mapped to the id the editor names the video's folder by, through the archive record's link. The editor must already archive the channel: pointed at a public site with a fresh editor, every clip gets a 404 `Channel "<slug>" not found`. The tool waits up to 90 seconds by default (at most 300) and otherwise returns the job's id, to wait on with `job`; the fetch carries on in the editor either way. While it waits it sends a progress notification per poll to a client that asks for progress. A client whose requests time out at 60 seconds (the MCP SDK's default) must raise that or pass `wait_seconds` of 50 or less. No request to the editor waits more than 15 seconds. If the editor stops answering mid-fetch, the answer gives the job's id and says not to ask again from scratch. The `/ask` and `/sweep` plans now tell the agent to use it and never to run yt-dlp itself. The MCP needs `ARCHILYZER_EDITOR_URL` and `WORKER_TOKEN` (the editor's own) in its environment, so re-register it with the two `--env` lines in the README; without them the tool says so and fetches nothing. The MCP server itself still writes nothing. The README's `yt-dlp --download-sections` command is now only the fallback for a machine with no editor.
## [0.9.0] - 2026-09-26
- **Every page now has a ground and an accent to choose, and the five theme families are gone.** The theme menu (the palette button beside the quick toggle, in the editor's sidebar and in the header of every published site, the hub and the homepage) has two groups. **Base** is System, Light, Sepia or Dark; Sepia is new, a warm paper ground for long reading. **Accent** is Signal, Brass, Vermilion, Violet, Sakura, Blue or Green, with the site's own tagged *default*; a site with a custom hex offers it first as *Site colour*. The quick toggle cycles System → Light → Sepia → Dark. A published site opens on the reader's system setting, in the accent its site form sets. The hub and the homepage open on Dark, in Signal, even with JavaScript off, and the editor follows the system, in Signal. Each accent has a value for each ground that reads at 4.5:1, and a custom hex is darkened or lightened per ground to match. A reader's accent is remembered only while it differs from the site's: picking the site's own again forgets it, so the reader follows the site if its accent changes later. Base, Archive, Selenized, Swiss and Archilyzer are gone. A choice made before this update carries over once: light stays light (Archive light becomes Sepia), dark stays dark and system stays system; the family itself is dropped. Headings are Archivo, text is IBM Plex Sans and figures are IBM Plex Mono everywhere, with one corner radius. Success, warning and other status text reads at 4.5:1 on its own tinted fill on every ground; on Light, success and warning are a shade deeper than before for it. Chart colours are fixed per ground and never follow the accent; the third is a violet, well clear of the red that marks a recording as gone. The phone's browser bar takes the page's ground, not the accent. Needs a rebuild and deploy of every site, the hub and the homepage.
diff --git a/mcp/README.md b/mcp/README.md
@@ -7,6 +7,8 @@ Desktop, Cursor, and any other MCP client.
It is a **local tool you run yourself**. It does not change the archive: it only
reads the site's already-published static JSON shards (`corpus.json` +
`transcripts/<slug>/…`), either from disk or over HTTP. Nothing is hosted for you.
+The one exception is `fetch_clip`, which asks a local Archilyzer editor to fetch a
+clip window; the MCP itself still writes nothing.
## Tools
@@ -19,6 +21,7 @@ reads the site's already-published static JSON shards (`corpus.json` +
| `get_transcript` | One video's full transcript as clean markdown (metadata + **linked** timestamped captions). |
| `get_post` / `get_thread` | One archived social post, or its whole thread. Posts have no timeline — cite them with no `@ mm:ss`. |
| `get_video_metadata` | Everything known about one video without the transcript body: metadata, plus **view/like counts, cue count and transcript coverage** (`stats/`), **other archived copies of the same recording** with an explicit timings-aligned verdict (`duplicates.json`), and **AI chapters/tags** where they exist (`digests/`). |
+| `fetch_clip` | The media behind a cited moment, **fetched by the local editor** (`POST /api/media/fetch-window`) through its paced, cookie-aware, provenanced job — never a yt-dlp run by hand. Needs `ARCHILYZER_EDITOR_URL` (default `http://localhost:3001`) and `WORKER_TOKEN` (the editor's own) in this server's env; without them it says so and fetches nothing. The editor must already archive the cited channel (a channel dir under its `transcripts/`), else it answers 404 `Channel "<slug>" not found`: an MCP pointed at a public site with a fresh editor gets that on every clip. A window is the cited span ± `pad` (default 3 s), at most 15 min, and lands at `channels/<slug>/data/<id>/clips/`; `full: true` fetches the whole recording into the saved-video store (needs a video the editor already knows). Waits up to `wait_seconds` (default 90, max 300), then returns the job id to resume with `job`; a client with a 60 s default request timeout must raise it or pass `wait_seconds` ≤ 50 — the fetch continues on the editor either way; resume it with `job`, and once it has finished the same request finds it cached. While it waits it sends one progress notification per poll to a client that asked for progress (a `progressToken`), which keeps a reset-on-progress timeout alive. A Rumble embed id is mapped to the editor's slug id through the record's `webpageUrl`, so pass the citing corpus as `source`; a video not in `source` is passed through as cited (known limitation). The file is a read-only corpus artifact. |
| `open_link` | Paste an archilyzer viewer **share link** to re-run that exact search here (query tree + every filter, at full fidelity) — plan, results and corpus handle in **one** call. `dry_run:true` for the plan alone. |
| `list_sources` | Show the **default** corpus and, with a hub, its member sites as ready-to-paste handles. |
| `resolve_source` | Turn a URL or site name into the canonical `source` handle and check it can be read. Changes nothing. |
@@ -474,7 +477,8 @@ nothing to reset).
**Still read-only.** A `source` handle only changes *which* already-published
static shards are read — the same capability the startup flags already grant
-this locally-run tool. Nothing is ever written to any corpus.
+this locally-run tool. This server never writes to any corpus; `fetch_clip` asks
+the editor, and the editor writes.
## Protocol
diff --git a/mcp/src/fetchClip.test.ts b/mcp/src/fetchClip.test.ts
@@ -0,0 +1,746 @@
+import { test } from "node:test";
+import assert from "node:assert/strict";
+import http from "node:http";
+import type { AddressInfo } from "node:net";
+import {
+ fetchClip,
+ parseSeconds,
+ planWindow,
+ renderFetchClip,
+ validateFetchClipArgs,
+ type FetchClipDeps,
+ type FetchClipRequest,
+ type HttpInit,
+} from "./fetchClip";
+
+// ─── A scripted editor: no network, no timers ───
+//
+// `respond` answers each request in turn; `sleep` records and advances the
+// fake clock by exactly what it was asked to wait, so a 90 s wait costs no real
+// time and the poll count is exact.
+
+type Answer = { status: number; body: unknown } | Error;
+type Call = {
+ url: string;
+ method: string;
+ headers: Record<string, string>;
+ body?: unknown;
+ signal?: AbortSignal;
+};
+
+function editor(
+ respond: (call: Call, n: number) => Answer,
+ env: Record<string, string | undefined> = {
+ WORKER_TOKEN: "tok",
+ ARCHILYZER_EDITOR_URL: "http://editor.test/",
+ },
+) {
+ const calls: Call[] = [];
+ const sleeps: number[] = [];
+ let clock = 1_000_000;
+ const deps: FetchClipDeps = {
+ env,
+ fetch: async (url: string, init?: HttpInit) => {
+ const call: Call = {
+ url,
+ method: init?.method ?? "GET",
+ headers: init?.headers ?? {},
+ body: init?.body !== undefined ? JSON.parse(init.body) : undefined,
+ signal: init?.signal,
+ };
+ calls.push(call);
+ const a = respond(call, calls.length - 1);
+ if (a instanceof Error) throw a;
+ return { status: a.status, json: async () => a.body };
+ },
+ sleep: async (ms: number) => {
+ sleeps.push(ms);
+ clock += ms;
+ },
+ now: () => clock,
+ };
+ return { deps, calls, sleeps };
+}
+
+// Answers in order; running off the end is a test bug, said loudly.
+function seq(...answers: Answer[]) {
+ return (call: Call, n: number): Answer => {
+ const a = answers[n];
+ if (!a) throw new Error(`unscripted request #${n}: ${call.method} ${call.url}`);
+ return a;
+ };
+}
+
+const CLIP_FILE =
+ "/corpus/channels/chan/data/vid/clips/7.00-23.00.mp4";
+
+function windowRequest(extra: Partial<Record<string, unknown>> = {}): FetchClipRequest {
+ const v = validateFetchClipArgs({
+ channel: "chan",
+ video: "vid",
+ start: 10,
+ end: 20,
+ reason: "the quote in the report",
+ report: "rep-1",
+ ...extra,
+ });
+ assert.ok(v.ok, v.ok ? "" : v.error);
+ return v.request;
+}
+
+function fullRequest(extra: Partial<Record<string, unknown>> = {}): FetchClipRequest {
+ const v = validateFetchClipArgs({
+ channel: "chan",
+ video: "vid",
+ full: true,
+ reason: "summarise the whole stream",
+ report: "rep-1",
+ ...extra,
+ });
+ assert.ok(v.ok, v.ok ? "" : v.error);
+ return v.request;
+}
+
+const CTX = { channel: "chan", video: "vid" };
+
+// ─── Times and windows ───
+
+test("parseSeconds reads seconds, mm:ss and h:mm:ss", () => {
+ assert.equal(parseSeconds("1:02:03"), 3723);
+ assert.equal(parseSeconds("1:02"), 62);
+ assert.equal(parseSeconds("75:30"), 4530);
+ assert.equal(parseSeconds("0:07.5"), 7.5);
+ assert.equal(parseSeconds("12"), 12);
+ assert.equal(parseSeconds(" 12.25 "), 12.25);
+ assert.equal(parseSeconds(12.5), 12.5);
+ assert.equal(parseSeconds(0), 0);
+ for (const bad of ["abc", "1:60", "1:2:3:4", "1:60:00", "-3", "", -3, Number.NaN, null, undefined, {}]) {
+ assert.equal(parseSeconds(bad), null, `${String(bad)} is not a time`);
+ }
+});
+
+test("planWindow pads, rounds to the editor's two decimals, and clamps at 0", () => {
+ assert.deepEqual(planWindow({ video: "vid", start: 10, end: 20, pad: 3 }), { from: 7, to: 23 });
+ // 7.004 → "7.00", 23.006 → "23.01": the name IS the window.
+ assert.deepEqual(
+ planWindow({ video: "vid", start: 10.004, end: 20.006, pad: 3 }),
+ { from: 7, to: 23.01 },
+ );
+ assert.deepEqual(planWindow({ video: "vid", start: 1, end: 5, pad: 3 }), { from: 0, to: 8 });
+ assert.deepEqual(planWindow({ video: "vid", start: 1, end: 5, pad: 0 }), { from: 1, to: 5 });
+});
+
+test("planWindow caps the window at 900 s AFTER padding", () => {
+ // 3 → 897 padded is exactly 0 → 900: allowed.
+ assert.deepEqual(planWindow({ video: "v", start: 3, end: 897, pad: 3 }), { from: 0, to: 900 });
+ // 895 s of citation is 901 s once padded: refused, with the padded numbers.
+ assert.deepEqual(planWindow({ video: "v", start: 10, end: 905, pad: 3 }), {
+ error:
+ "fetch_clip: the window 7.00–908.00 is 901s; the editor fetches at most " +
+ "900s per window — cite a narrower span",
+ });
+});
+
+test("planWindow refuses a bad id and an inverted span, in the plan's words", () => {
+ for (const video of ["..", ".", "a/b", "a b", ""]) {
+ assert.deepEqual(planWindow({ video, start: 1, end: 2, pad: 3 }), {
+ error: `fetch_clip: video "${video}" must match /^[\\w.-]+$/`,
+ });
+ }
+ assert.deepEqual(planWindow({ video: "v", start: 20, end: 10, pad: 3 }), {
+ error: "fetch_clip: start (20) must be less than end (10)",
+ });
+ assert.deepEqual(planWindow({ video: "v", start: 10, end: 10, pad: 3 }), {
+ error: "fetch_clip: start (10) must be less than end (10)",
+ });
+});
+
+// ─── Arguments ───
+
+test("argument errors say what to fix, before any HTTP", () => {
+ const err = (args: Record<string, unknown>) => {
+ const v = validateFetchClipArgs(args);
+ assert.equal(v.ok, false);
+ return v.ok ? "" : v.error;
+ };
+ const base = { channel: "chan", video: "vid", start: 10, end: 20, reason: "why" };
+ assert.equal(
+ err({ ...base, channel: undefined }),
+ "fetch_clip: channel is required (the channel slug)",
+ );
+ assert.equal(
+ err({ ...base, video: " " }),
+ "fetch_clip: video is required (the archive's video id)",
+ );
+ assert.equal(err({ ...base, video: "../x" }), 'fetch_clip: video "../x" must match /^[\\w.-]+$/');
+ assert.equal(
+ err({ ...base, start: "soon" }),
+ 'fetch_clip: start "soon" is not a time (use seconds, mm:ss or h:mm:ss)',
+ );
+ assert.equal(
+ err({ ...base, end: "1:99" }),
+ 'fetch_clip: end "1:99" is not a time (use seconds, mm:ss or h:mm:ss)',
+ );
+ assert.equal(err({ ...base, start: "0:20", end: "0:10" }), "fetch_clip: start (20) must be less than end (10)");
+ assert.equal(
+ err({ ...base, reason: " " }),
+ "fetch_clip: reason is required — one line saying why these seconds are " +
+ "needed (it is stored beside the file)",
+ );
+ assert.match(err({ ...base, pad: -1 }), /^fetch_clip: pad "-1" must be a finite number/);
+ assert.match(err({ ...base, pad: "3" }), /^fetch_clip: pad "3" must be a finite number/);
+ assert.match(err({ ...base, start: undefined }), /^fetch_clip: start is required/);
+});
+
+test("wait_seconds is clamped to [0, 300] and defaults to 90", () => {
+ const wait = (w: unknown) => {
+ const v = validateFetchClipArgs({ job: "j1", wait_seconds: w });
+ assert.ok(v.ok);
+ return v.request.waitSeconds;
+ };
+ assert.equal(wait(undefined), 90);
+ assert.equal(wait(500), 300);
+ assert.equal(wait(-5), 0);
+ assert.equal(wait(12), 12);
+ assert.equal(wait("30"), 90);
+});
+
+test("a job alone is a whole request; everything else is then ignored", () => {
+ const v = validateFetchClipArgs({ job: " j7 ", video: "../bad", start: "nope" });
+ assert.deepEqual(v, { ok: true, request: { job: "j7", waitSeconds: 90 } });
+});
+
+test("full mode ignores start/end/pad entirely — not required, not validated", () => {
+ const v = validateFetchClipArgs({
+ channel: "chan",
+ video: "vid",
+ full: true,
+ start: "not a time",
+ end: -4,
+ pad: -1,
+ reason: "why",
+ });
+ assert.ok(v.ok, v.ok ? "" : v.error);
+ assert.deepEqual(v.request, {
+ target: { kind: "full", channel: "chan", video: "vid" },
+ reason: "why",
+ report: undefined,
+ waitSeconds: 90,
+ });
+ // …but a reason is still required.
+ const noReason = validateFetchClipArgs({ channel: "chan", video: "vid", full: true });
+ assert.equal(noReason.ok, false);
+});
+
+// ─── The POST ───
+
+test("the POST carries requestedBy mcp, the report as manifest, and a reason cut to 400", async () => {
+ const { deps, calls } = editor(
+ seq({ status: 200, body: { cached: true, file: CLIP_FILE, from: 7, to: 23, bytes: 10, provenance: null } }),
+ );
+ const long = "x".repeat(500);
+ const req = windowRequest({ reason: long });
+ if ("target" in req && req.target.kind === "window") {
+ req.target.webpageUrl = "https://www.youtube.com/watch?v=vid";
+ }
+ await fetchClip(req, deps);
+ assert.equal(calls.length, 1);
+ assert.equal(calls[0].method, "POST");
+ // The trailing slash on ARCHILYZER_EDITOR_URL does not double up.
+ assert.equal(calls[0].url, "http://editor.test/api/media/fetch-window");
+ assert.equal(calls[0].headers.authorization, "Bearer tok");
+ assert.deepEqual(calls[0].body, {
+ channelSlug: "chan",
+ videoId: "vid",
+ webpageUrl: "https://www.youtube.com/watch?v=vid",
+ from: 7,
+ to: 23,
+ pad: 3,
+ requestedBy: "mcp",
+ manifest: "rep-1",
+ reason: "x".repeat(400),
+ });
+});
+
+test("no editor configured is an outcome, and nothing is sent", async () => {
+ for (const env of [{}, { ARCHILYZER_EDITOR_URL: "http://localhost:3001" }, { WORKER_TOKEN: " " }]) {
+ const { deps, calls } = editor(seq(), env);
+ const outcome = await fetchClip(windowRequest(), deps);
+ assert.deepEqual(outcome, { kind: "no_editor" });
+ assert.equal(calls.length, 0);
+ const r = renderFetchClip(outcome);
+ assert.equal(r.isError, true);
+ assert.match(r.text, /^fetch_clip: no editor configured\./);
+ assert.match(r.text, /ARCHILYZER_EDITOR_URL/);
+ assert.match(r.text, /WORKER_TOKEN/);
+ assert.match(r.text, /no-editor fallback/);
+ }
+});
+
+test("the editor URL defaults to localhost:3001", async () => {
+ const { deps, calls } = editor(
+ seq({ status: 200, body: { cached: true, file: CLIP_FILE, from: 7, to: 23, bytes: 1 } }),
+ { WORKER_TOKEN: "tok" },
+ );
+ await fetchClip(windowRequest(), deps);
+ assert.equal(calls[0].url, "http://localhost:3001/api/media/fetch-window");
+});
+
+// ─── Answers ───
+
+test("200 cached: a WIDER window is named as such, with its own span", async () => {
+ const wide = "/corpus/channels/chan/data/vid/clips/0.00-60.00.mp4";
+ const { deps } = editor(
+ seq({
+ status: 200,
+ body: { cached: true, file: wide, from: 0, to: 60, bytes: 123456, provenance: { requestedBy: "umtool" } },
+ }),
+ );
+ const outcome = await fetchClip(windowRequest(), deps);
+ const r = renderFetchClip(outcome, CTX);
+ assert.equal(r.isError, false);
+ assert.equal(
+ r.text,
+ [
+ "Already on disk — a WIDER cached window that contains 7.00–23.00: 0.00–60.00.",
+ "",
+ `file: ${wide}`,
+ "window: 0.00–60.00 (60s)",
+ "bytes: 123456",
+ "requested by umtool",
+ "This path is a read-only corpus artifact: play or copy it, never move, " +
+ "edit or delete it. The editor prunes clips by age (evict-clips); " +
+ "provenance sits beside it as 0.00-60.00.json.",
+ ].join("\n"),
+ );
+});
+
+test("200 cached: the exact window (within 0.02 s) says so", async () => {
+ const { deps } = editor(
+ seq({ status: 200, body: { cached: true, file: CLIP_FILE, from: 7.01, to: 23, bytes: 5, provenance: null } }),
+ );
+ const r = renderFetchClip(await fetchClip(windowRequest(), deps), CTX);
+ assert.match(r.text, /^Already on disk — the exact window: 7\.01–23\.00\./);
+ assert.ok(!/requested by/.test(r.text));
+});
+
+test("202 → running → done: the poll's file, span and bytes", async () => {
+ const { deps, calls, sleeps } = editor(
+ seq(
+ { status: 202, body: { cached: false, jobId: "j1", file: CLIP_FILE, from: 7, to: 23 } },
+ { status: 200, body: { status: "running", jobId: "j1" } },
+ { status: 200, body: { status: "done", jobId: "j1", file: CLIP_FILE, from: 7, to: 23, bytes: 4096 } },
+ ),
+ );
+ const outcome = await fetchClip(windowRequest(), deps);
+ assert.equal(outcome.kind, "fetched");
+ assert.deepEqual(calls.map((c) => c.method), ["POST", "GET", "GET"]);
+ assert.equal(calls[1].url, "http://editor.test/api/media/fetch-window/j1");
+ assert.equal(calls[1].headers.authorization, "Bearer tok");
+ assert.deepEqual(sleeps, [1000, 1000]);
+ const r = renderFetchClip(outcome, CTX);
+ assert.equal(r.isError, false);
+ assert.match(r.text, /^Fetched 7\.00–23\.00 of chan\/vid \(job j1, 2s waited\)\.\n\n/);
+ assert.match(r.text, new RegExp(`\\nfile: ${CLIP_FILE.replace(/\./g, "\\.")}\\n`));
+ assert.match(r.text, /\nwindow: 7\.00–23\.00 \(16s\)\n/);
+ assert.match(r.text, /\nbytes: 4096\n/);
+ assert.match(r.text, /provenance sits beside it as 7\.00-23\.00\.json\.$/);
+});
+
+test("202 → failed: the job's log tail comes back", async () => {
+ const { deps } = editor(
+ seq(
+ { status: 202, body: { cached: false, jobId: "j1", file: CLIP_FILE, from: 7, to: 23 } },
+ { status: 200, body: { status: "failed", jobId: "j1", error: "ERROR: [youtube] vid: Video unavailable" } },
+ ),
+ );
+ const r = renderFetchClip(await fetchClip(windowRequest(), deps), CTX);
+ assert.equal(r.isError, true);
+ assert.equal(r.text, "Editor job j1 failed.\n\nlog tail:\nERROR: [youtube] vid: Video unavailable");
+});
+
+test("the wait runs out → queued with the job id, without real timers", async () => {
+ const { deps, calls, sleeps } = editor((call, n) =>
+ n === 0
+ ? { status: 202, body: { cached: false, jobId: "j1", file: CLIP_FILE, from: 7, to: 23 } }
+ : { status: 200, body: { status: "running", jobId: "j1" } },
+ );
+ const outcome = await fetchClip(windowRequest({ wait_seconds: 3 }), deps);
+ assert.deepEqual(outcome, { kind: "queued", jobId: "j1", status: "running", waited: 3 });
+ assert.deepEqual(sleeps, [1000, 1000, 1000]);
+ assert.equal(calls.length, 4);
+ const r = renderFetchClip(outcome, CTX);
+ assert.equal(r.isError, false, "a wait that runs out is not a failure");
+ assert.equal(
+ r.text,
+ 'Still running on the editor (job j1, waited 3s). Call fetch_clip again with job: "j1" ' +
+ "to keep waiting — never repeat the original request while it runs (that would queue " +
+ "a second fetch). Nothing is lost: the fetch continues on the editor, and once it has " +
+ "finished the same request finds it cached.",
+ );
+});
+
+test("wait_seconds 0 returns the job at once, still queued", async () => {
+ const { deps, calls, sleeps } = editor(
+ seq({ status: 202, body: { cached: false, jobId: "j1", file: CLIP_FILE, from: 7, to: 23 } }),
+ );
+ const outcome = await fetchClip(windowRequest({ wait_seconds: 0 }), deps);
+ assert.deepEqual(outcome, { kind: "queued", jobId: "j1", status: "queued", waited: 0 });
+ assert.equal(calls.length, 1);
+ assert.deepEqual(sleeps, []);
+});
+
+test("resume by job issues no POST, and polls before it sleeps", async () => {
+ const { deps, calls, sleeps } = editor(
+ seq({ status: 200, body: { status: "done", jobId: "j9", file: CLIP_FILE, from: 7, to: 23, bytes: 77 } }),
+ );
+ const outcome = await fetchClip({ job: "j9", waitSeconds: 90 }, deps);
+ assert.deepEqual(calls.map((c) => `${c.method} ${c.url}`), [
+ "GET http://editor.test/api/media/fetch-window/j9",
+ ]);
+ assert.deepEqual(sleeps, []);
+ // A resume knows no channel/video; the window comes from the poll.
+ const r = renderFetchClip(outcome);
+ assert.match(r.text, /^Fetched 7\.00–23\.00 \(job j9, 0s waited\)\.\n\n/);
+ assert.match(r.text, /\nwindow: 7\.00–23\.00 \(16s\)\n/);
+});
+
+test("409: the platform and the seconds left", async () => {
+ const { deps } = editor(
+ seq({
+ status: 409,
+ body: { error: "youtube is in a rate-limit cooldown (13s remaining).", cooldownMs: 12_500, platform: "youtube" },
+ }),
+ );
+ const r = renderFetchClip(await fetchClip(windowRequest(), deps), CTX);
+ assert.equal(r.isError, true);
+ assert.equal(
+ r.text,
+ "The editor is in a youtube rate-limit cooldown — 13s remaining. Wait, then call " +
+ "fetch_clip again. (youtube is in a rate-limit cooldown (13s remaining).)",
+ );
+});
+
+test("401 and a disabled 503 are token problems; 503 disabled says the endpoint is off", async () => {
+ const bad = editor(seq({ status: 401, body: { error: "invalid worker token" } }));
+ assert.deepEqual(renderFetchClip(await fetchClip(windowRequest(), bad.deps), CTX), {
+ text:
+ "The editor refused the token (HTTP 401): invalid worker token. WORKER_TOKEN must " +
+ "equal the value the editor runs with (editor/.env).",
+ isError: true,
+ });
+ const off = editor(
+ seq({ status: 503, body: { error: "worker endpoint disabled (set WORKER_TOKEN to enable)" } }),
+ );
+ assert.deepEqual(renderFetchClip(await fetchClip(windowRequest(), off.deps), CTX), {
+ text:
+ "The editor refused the token (HTTP 503): worker endpoint disabled (set WORKER_TOKEN " +
+ "to enable). WORKER_TOKEN must equal the value the editor runs with (editor/.env). " +
+ "The editor has no WORKER_TOKEN set, so its fetch endpoint is off.",
+ isError: true,
+ });
+});
+
+test("a 503 that is not the token (an unmounted drive) passes through verbatim", async () => {
+ const drive = "Channel media for chan is unreachable: /mnt/platter is not mounted";
+ const { deps } = editor(seq({ status: 503, body: { error: drive } }));
+ assert.deepEqual(renderFetchClip(await fetchClip(windowRequest(), deps), CTX), {
+ text: `The editor refused (HTTP 503): ${drive}`,
+ isError: true,
+ });
+});
+
+test("any other refusal is passed through verbatim", async () => {
+ const { deps } = editor(seq({ status: 507, body: { error: "Low disk space: 2.1 GB free" } }));
+ assert.deepEqual(renderFetchClip(await fetchClip(windowRequest(), deps), CTX), {
+ text: "The editor refused (HTTP 507): Low disk space: 2.1 GB free",
+ isError: true,
+ });
+});
+
+test("an unreachable editor is named, with its URL", async () => {
+ const { deps } = editor(seq(new Error("connect ECONNREFUSED 127.0.0.1:3001")));
+ assert.deepEqual(renderFetchClip(await fetchClip(windowRequest(), deps), CTX), {
+ text: "Could not reach the editor at http://editor.test: connect ECONNREFUSED 127.0.0.1:3001. Is it running?",
+ isError: true,
+ });
+});
+
+test("a poll the editor does not know says it may have restarted", async () => {
+ const { deps } = editor(seq({ status: 404, body: { error: "no such job" } }));
+ assert.deepEqual(renderFetchClip(await fetchClip({ job: "j1", waitSeconds: 5 }, deps)), {
+ text:
+ "Polling editor job j1 failed (HTTP 404): no such job; the job is unknown to this " +
+ "editor (restarted? wrong ARCHILYZER_EDITOR_URL?)",
+ isError: true,
+ });
+});
+
+// ─── full: true — the whole recording ───
+
+test("full: the POST is exactly {channelSlug, videoId, full, requestedBy, manifest, reason}", async () => {
+ // start/end/pad given anyway: they are ignored, not sent.
+ for (const req of [fullRequest(), fullRequest({ start: 10, end: 20, pad: 5 })]) {
+ const { deps, calls } = editor(
+ seq({ status: 202, body: { cached: false, jobId: "j2", file: null, from: 0, to: 0 } }),
+ );
+ await fetchClip({ ...req, waitSeconds: 0 } as FetchClipRequest, deps);
+ assert.deepEqual(calls[0].body, {
+ channelSlug: "chan",
+ videoId: "vid",
+ full: true,
+ requestedBy: "mcp",
+ manifest: "rep-1",
+ reason: "summarise the whole stream",
+ });
+ }
+});
+
+test("full: 200 cached names the saved-video store, with no window line", async () => {
+ const saved = "/corpus/saved-videos/chan/vid/source.mkv";
+ const { deps } = editor(
+ seq({
+ status: 200,
+ body: { cached: true, file: saved, from: 0, to: 0, bytes: 9_000_000, provenance: { requestedBy: "umtool" } },
+ }),
+ );
+ const r = renderFetchClip(await fetchClip(fullRequest(), deps), CTX);
+ assert.equal(r.isError, false);
+ assert.equal(
+ r.text,
+ [
+ "Already on disk — the whole recording of chan/vid.",
+ "",
+ `file: ${saved}`,
+ "bytes: 9000000",
+ "requested by umtool",
+ "This path is a read-only corpus artifact: play or copy it, never move, edit " +
+ "or delete it. It lives in the editor's saved-video store, not in clips/; " +
+ "the editor's keep-videos rule decides how long it stays.",
+ ].join("\n"),
+ );
+});
+
+test("full: 202 → done carries the file and bytes, and no window line", async () => {
+ const saved = "/corpus/saved-videos/chan/vid/source.mkv";
+ const { deps } = editor(
+ seq(
+ { status: 202, body: { cached: false, jobId: "j2", file: null, from: 0, to: 0 } },
+ { status: 200, body: { status: "done", jobId: "j2", file: saved, bytes: 42 } },
+ ),
+ );
+ const r = renderFetchClip(await fetchClip(fullRequest(), deps), CTX);
+ assert.equal(r.isError, false);
+ assert.match(r.text, /^Fetched the whole recording of chan\/vid \(job j2, 1s waited\)\.\n\n/);
+ assert.match(r.text, /\nbytes: 42\n/);
+ assert.ok(!/window:/.test(r.text), r.text);
+ assert.match(r.text, /saved-video store, not in clips\//);
+});
+
+test("full: a done job that names no file is an error", async () => {
+ const { deps } = editor(
+ seq(
+ { status: 202, body: { cached: false, jobId: "j2", file: null, from: 0, to: 0 } },
+ { status: 200, body: { status: "done", jobId: "j2" } },
+ ),
+ );
+ assert.deepEqual(renderFetchClip(await fetchClip(fullRequest(), deps), CTX), {
+ text: "Editor job j2 finished but named no file; check the video's page in the editor.",
+ isError: true,
+ });
+});
+
+test("full: a 404 says a whole recording needs a video the editor knows; a window's does not", async () => {
+ const unknown =
+ "Could not determine the video URL: no metadata.info.json and the playlist does " +
+ "not contain a matching entry.";
+ const full = editor(seq({ status: 404, body: { error: unknown } }));
+ assert.deepEqual(renderFetchClip(await fetchClip(fullRequest(), full.deps), CTX), {
+ text:
+ `The editor refused (HTTP 404): ${unknown} A whole-recording fetch needs a video ` +
+ "the editor already knows (a metadata.info.json or a playlist entry); fetch a " +
+ "window instead, or add the video to the channel first.",
+ isError: true,
+ });
+ const win = editor(seq({ status: 404, body: { error: "Channel not found" } }));
+ assert.deepEqual(renderFetchClip(await fetchClip(windowRequest(), win.deps), CTX), {
+ text: "The editor refused (HTTP 404): Channel not found",
+ isError: true,
+ });
+});
+
+test("a resumed whole-recording job renders as one (no from/to in its answer)", async () => {
+ const saved = "/corpus/saved-videos/chan/vid/source.mkv";
+ const { deps } = editor(
+ seq({ status: 200, body: { status: "done", jobId: "j2", file: saved, bytes: 42 } }),
+ );
+ const r = renderFetchClip(await fetchClip({ job: "j2", waitSeconds: 90 }, deps));
+ assert.match(r.text, /^Fetched the whole recording \(job j2, 0s waited\)\./);
+ assert.ok(!/window:/.test(r.text));
+});
+
+test("the editor dropping out WHILE POLLING hands back the job, and says not to repeat the request", async () => {
+ // Full mode is where it matters: a repeated full: true would queue a second
+ // whole-recording download, because the saved-video pointer is only checked
+ // when a request arrives.
+ const { deps, calls } = editor(
+ seq(
+ { status: 202, body: { cached: false, jobId: "j3", file: null, from: 0, to: 0 } },
+ new Error("socket hang up"),
+ ),
+ );
+ const outcome = await fetchClip(fullRequest(), deps);
+ assert.deepEqual(calls.map((c) => c.method), ["POST", "GET"]);
+ const r = renderFetchClip(outcome, CTX);
+ assert.deepEqual(r, {
+ text:
+ "Could not reach the editor at http://editor.test while polling job j3: socket hang " +
+ 'up. Is it running? Call fetch_clip again with job: "j3" — do not repeat the original ' +
+ "request, the fetch may still be running.",
+ isError: true,
+ });
+ // …and the call it names is a resume: one poll, no second POST.
+ const again = editor(
+ seq({ status: 200, body: { status: "done", jobId: "j3", file: "/corpus/saved-videos/x.mkv", bytes: 9 } }),
+ );
+ const v = validateFetchClipArgs({ job: "j3" });
+ assert.ok(v.ok);
+ const resumed = renderFetchClip(await fetchClip(v.request, again.deps));
+ assert.deepEqual(again.calls.map((c) => c.method), ["GET"]);
+ assert.match(resumed.text, /^Fetched the whole recording \(job j3, 0s waited\)\./);
+});
+
+// ─── Every request is bounded, and a failure says why ───
+
+test("every request — the POST and each poll — carries its own timeout signal", async () => {
+ const { deps, calls } = editor(
+ seq(
+ { status: 202, body: { cached: false, jobId: "j1", file: CLIP_FILE, from: 7, to: 23 } },
+ { status: 200, body: { status: "running", jobId: "j1" } },
+ { status: 200, body: { status: "done", jobId: "j1", file: CLIP_FILE, from: 7, to: 23, bytes: 1 } },
+ ),
+ );
+ await fetchClip(windowRequest(), deps);
+ assert.equal(calls.length, 3);
+ for (const c of calls) assert.ok(c.signal instanceof AbortSignal, `${c.method} has a signal`);
+ // One signal per request, not one shared deadline for the whole call.
+ assert.equal(new Set(calls.map((c) => c.signal)).size, 3);
+});
+
+test("Node's bare 'fetch failed' gets its cause appended", async () => {
+ const { deps } = editor(
+ seq(new TypeError("fetch failed", { cause: new Error("connect ECONNREFUSED 127.0.0.1:3001") })),
+ );
+ assert.equal(
+ renderFetchClip(await fetchClip(windowRequest(), deps), CTX).text,
+ "Could not reach the editor at http://editor.test: fetch failed (connect ECONNREFUSED " +
+ "127.0.0.1:3001). Is it running?",
+ );
+});
+
+test("a timed-out POST says so, and that the fetch may have been queued anyway", async () => {
+ const { deps } = editor(
+ seq(new DOMException("The operation was aborted due to timeout", "TimeoutError")),
+ );
+ assert.deepEqual(renderFetchClip(await fetchClip(fullRequest(), deps), CTX), {
+ text:
+ "Could not reach the editor at http://editor.test: no answer within 15 s (request " +
+ "timed out). Is it running? It may have queued the fetch anyway: check the editor's " +
+ "/jobs page before asking again.",
+ isError: true,
+ });
+});
+
+// A real socket, a real fetch: an editor that accepts the connection and never
+// answers — a POST stuck on a hung mount — must not hold the call.
+async function silentEditor(answerPost?: unknown) {
+ const server = http.createServer((req, res) => {
+ if (answerPost && req.method === "POST") {
+ res.writeHead(202, { "content-type": "application/json" });
+ res.end(JSON.stringify(answerPost));
+ }
+ // Anything else: never answer.
+ });
+ await new Promise<void>((r) => server.listen(0, "127.0.0.1", r));
+ const { port } = server.address() as AddressInfo;
+ const close = () => {
+ server.closeAllConnections();
+ return new Promise<void>((r) => server.close(() => r()));
+ };
+ return { url: `http://127.0.0.1:${port}`, port, close };
+}
+
+function realDeps(url: string, requestTimeoutMs: number): FetchClipDeps {
+ let clock = 0;
+ return {
+ env: { WORKER_TOKEN: "tok", ARCHILYZER_EDITOR_URL: url },
+ fetch: (u, init) => globalThis.fetch(u, init),
+ sleep: async (ms) => {
+ clock += ms;
+ },
+ now: () => clock,
+ requestTimeoutMs,
+ };
+}
+
+test("a real editor that never answers the POST is given up on at the request timeout", { timeout: 10_000 }, async () => {
+ const ed = await silentEditor();
+ try {
+ const t0 = Date.now();
+ const r = renderFetchClip(await fetchClip(windowRequest(), realDeps(ed.url, 150)), CTX);
+ const took = Date.now() - t0;
+ assert.ok(took < 5000, `gave up after ${took} ms`);
+ assert.equal(r.isError, true);
+ assert.equal(
+ r.text,
+ `Could not reach the editor at ${ed.url}: no answer within 0.15 s (request timed out). ` +
+ "Is it running? It may have queued the fetch anyway: check the editor's /jobs page " +
+ "before asking again.",
+ );
+ } finally {
+ await ed.close();
+ }
+});
+
+test("a real editor that stops answering a poll hands back the job", { timeout: 10_000 }, async () => {
+ const ed = await silentEditor({ cached: false, jobId: "j4", file: null, from: 0, to: 0 });
+ try {
+ const r = renderFetchClip(await fetchClip(fullRequest(), realDeps(ed.url, 150)), CTX);
+ assert.equal(
+ r.text,
+ `Could not reach the editor at ${ed.url} while polling job j4: no answer within 0.15 s ` +
+ '(request timed out). Is it running? Call fetch_clip again with job: "j4" — do not ' +
+ "repeat the original request, the fetch may still be running.",
+ );
+ } finally {
+ await ed.close();
+ }
+});
+
+test("a real refused connection names its cause", { timeout: 10_000 }, async () => {
+ const ed = await silentEditor();
+ const url = ed.url;
+ await ed.close(); // nothing listens there now
+ const r = renderFetchClip(await fetchClip(windowRequest(), realDeps(url, 5000)), CTX);
+ assert.match(r.text, /: fetch failed \(connect ECONNREFUSED 127\.0\.0\.1:\d+\)\. Is it running\?$/);
+});
+
+test("onPoll is told after every poll that finds the job still waiting, and cannot break the fetch", async () => {
+ const { deps } = editor(
+ seq(
+ { status: 202, body: { cached: false, jobId: "j1", file: CLIP_FILE, from: 7, to: 23 } },
+ { status: 200, body: { status: "queued", jobId: "j1" } },
+ { status: 200, body: { status: "running", jobId: "j1" } },
+ { status: 200, body: { status: "done", jobId: "j1", file: CLIP_FILE, from: 7, to: 23, bytes: 1 } },
+ ),
+ );
+ const seen: string[] = [];
+ const outcome = await fetchClip(windowRequest(), deps, {
+ onPoll: (p) => {
+ seen.push(`${p.polls} ${p.jobId} ${p.status} ${p.waited}s`);
+ throw new Error("the client went away");
+ },
+ });
+ assert.deepEqual(seen, ["1 j1 queued 1s", "2 j1 running 2s"]);
+ assert.equal(outcome.kind, "fetched");
+});
diff --git a/mcp/src/fetchClip.tool.test.ts b/mcp/src/fetchClip.tool.test.ts
@@ -0,0 +1,299 @@
+import { test } from "node:test";
+import assert from "node:assert/strict";
+import { Client, InMemoryTransport } from "@modelcontextprotocol/client";
+import type { ChannelTranscriptsManifest } from "yt-dlp-transcript-common/lib/manifest";
+import type { TranscriptDetail } from "yt-dlp-transcript-common/lib/transcripts";
+import type { SearchAlias } from "yt-dlp-transcript-common/lib/searchAliases";
+import type {
+ ChannelGroups,
+ ChannelRef,
+ ShardSource,
+ VideoAvailability,
+} from "./source";
+import { SourceRegistry } from "./sourceRegistry";
+import { createServer } from "./server";
+import type { FetchClipDeps, HttpInit } from "./fetchClip";
+
+// ─── fetch_clip through the real tools/call path ───
+//
+// The unit tests pin the HTTP exchange; these pin what only the server does:
+// `source` resolves first, the cited id is mapped to the editor's directory id
+// through the corpus record, and the corpus trailer lands on every answer.
+
+const RUMBLE = "the-quartering-rumble";
+const YT = "chan-yt";
+
+// A published Rumble record is keyed by yt-dlp's native id — the EMBED id —
+// while its webpageUrl is the slug URL the editor names the directory for.
+const RECORDS: Record<string, Record<string, unknown>> = {
+ [RUMBLE]: {
+ id: "vxe1ae",
+ slug: `${RUMBLE}/vxe1ae`,
+ title: "a rumble stream",
+ webpageUrl: "https://rumble.com/v1007ay-x.html?e9s=1",
+ },
+ [YT]: {
+ id: "dQw4w9WgXcQ",
+ slug: `${YT}/dQw4w9WgXcQ`,
+ title: "a youtube video",
+ webpageUrl: "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
+ },
+};
+
+// How many times any source was asked for its channels — the first step of
+// every corpus lookup.
+let channelListings = 0;
+
+class RecordSource implements ShardSource {
+ readonly label = "local:/srv/fixture";
+ async loadAliases(): Promise<SearchAlias[]> {
+ return [];
+ }
+ async loadGroups(): Promise<ChannelGroups> {
+ return { groups: [], defaultGroupId: "default" };
+ }
+ async listChannels(): Promise<ChannelRef[]> {
+ channelListings++;
+ return Object.keys(RECORDS).map((slug) => ({ key: slug, slug, name: slug }));
+ }
+ async transcriptsManifest(ch: ChannelRef): Promise<ChannelTranscriptsManifest> {
+ const rec = RECORDS[ch.slug];
+ return { slugToPage: { [String(rec.id)]: 0 } } as unknown as ChannelTranscriptsManifest;
+ }
+ async transcriptPage(ch: ChannelRef): Promise<TranscriptDetail[]> {
+ return [RECORDS[ch.slug] as unknown as TranscriptDetail];
+ }
+ publicOrigin(): string | null {
+ return null;
+ }
+ async subsManifest(): Promise<null> {
+ return null;
+ }
+ async subsPage(): Promise<[]> {
+ return [];
+ }
+ async postsManifest(): Promise<null> {
+ return null;
+ }
+ async postsPage(): Promise<[]> {
+ return [];
+ }
+ async availabilityMap(): Promise<Map<string, VideoAvailability>> {
+ return new Map();
+ }
+}
+
+type Posted = { url: string; method: string; body?: Record<string, unknown> };
+
+// A fake editor that answers every POST "cached" (or with `answer`), and a
+// deps bundle that records what was sent.
+function fakeEditor(
+ env: Record<string, string | undefined> = { WORKER_TOKEN: "tok" },
+ answer: (call: Posted) => { status: number; body: unknown } = () => ({
+ status: 200,
+ body: {
+ cached: true,
+ file: "/corpus/channels/x/data/y/clips/7.00-23.00.mp4",
+ from: 7,
+ to: 23,
+ bytes: 10,
+ provenance: null,
+ },
+ }),
+): { deps: Partial<FetchClipDeps>; calls: Posted[] } {
+ const calls: Posted[] = [];
+ let clock = 0;
+ return {
+ calls,
+ deps: {
+ env,
+ fetch: async (url: string, init?: HttpInit) => {
+ const call: Posted = {
+ url,
+ method: init?.method ?? "GET",
+ body: init?.body ? JSON.parse(init.body) : undefined,
+ };
+ calls.push(call);
+ const a = answer(call);
+ return { status: a.status, json: async () => a.body };
+ },
+ sleep: async (ms: number) => {
+ clock += ms;
+ },
+ now: () => clock,
+ },
+ };
+}
+
+async function connect(deps: Partial<FetchClipDeps>): Promise<Client> {
+ const registry = SourceRegistry.forSource(new RecordSource());
+ const server = createServer(registry, { fetchClipDeps: deps });
+ const [ct, st] = InMemoryTransport.createLinkedPair();
+ const client = new Client({ name: "test", version: "0" }, { capabilities: {} });
+ await Promise.all([server.connect(st), client.connect(ct)]);
+ return client;
+}
+
+function textOf(res: unknown): string {
+ const content = (res as { content: { type: string; text: string }[] }).content;
+ return content.map((c) => c.text).join("\n");
+}
+const isError = (res: unknown) => (res as { isError?: boolean }).isError === true;
+
+test("a Rumble embed id is fetched under the editor's slug id, with the record's URL", async () => {
+ const { deps, calls } = fakeEditor();
+ const client = await connect(deps);
+ const res = await client.callTool({
+ name: "fetch_clip",
+ arguments: { channel: RUMBLE, video: "vxe1ae", start: 10, end: 20, reason: "proof" },
+ });
+ assert.equal(isError(res), false, textOf(res));
+ assert.equal(calls.length, 1);
+ assert.equal(calls[0].body?.videoId, "v1007ay");
+ assert.equal(calls[0].body?.webpageUrl, "https://rumble.com/v1007ay-x.html?e9s=1");
+ assert.equal(calls[0].body?.channelSlug, RUMBLE);
+ assert.match(textOf(res), /\n\n\(corpus: local:\/srv\/fixture\)$/);
+});
+
+test("the channel is the record's own slug, however the citation spelled it", async () => {
+ const { deps, calls } = fakeEditor();
+ const client = await connect(deps);
+ await client.callTool({
+ name: "fetch_clip",
+ arguments: { channel: "The-Quartering-Rumble", video: "vxe1ae", start: 10, end: 20, reason: "proof" },
+ });
+ assert.equal(calls[0].body?.channelSlug, RUMBLE);
+});
+
+test("a YouTube id maps to itself", async () => {
+ const { deps, calls } = fakeEditor();
+ const client = await connect(deps);
+ await client.callTool({
+ name: "fetch_clip",
+ arguments: { channel: YT, video: "dQw4w9WgXcQ", start: "1:00", end: "1:10", reason: "proof" },
+ });
+ assert.equal(calls[0].body?.videoId, "dQw4w9WgXcQ");
+ assert.equal(calls[0].body?.from, 57);
+ assert.equal(calls[0].body?.to, 73);
+});
+
+test("full: true still maps a Rumble id to the slug, and sends no URL or span", async () => {
+ const { deps, calls } = fakeEditor();
+ const client = await connect(deps);
+ await client.callTool({
+ name: "fetch_clip",
+ arguments: { channel: RUMBLE, video: "vxe1ae", full: true, start: 1, end: 2, reason: "summary" },
+ });
+ assert.deepEqual(calls[0].body, {
+ channelSlug: RUMBLE,
+ videoId: "v1007ay",
+ full: true,
+ requestedBy: "mcp",
+ reason: "summary",
+ });
+});
+
+test("no editor configured: an error naming both variables, and nothing sent", async () => {
+ const { deps, calls } = fakeEditor({});
+ const client = await connect(deps);
+ const res = await client.callTool({
+ name: "fetch_clip",
+ arguments: { channel: YT, video: "dQw4w9WgXcQ", start: 1, end: 2, reason: "proof" },
+ });
+ assert.equal(isError(res), true);
+ const t = textOf(res);
+ assert.match(t, /ARCHILYZER_EDITOR_URL/);
+ assert.match(t, /WORKER_TOKEN/);
+ assert.match(t, /\(corpus: local:\/srv\/fixture\)$/);
+ assert.equal(calls.length, 0);
+});
+
+test("a video the corpus does not hold goes through as cited, with a note", async () => {
+ const { deps, calls } = fakeEditor();
+ const client = await connect(deps);
+ const res = await client.callTool({
+ name: "fetch_clip",
+ arguments: { channel: YT, video: "zzzz", start: 1, end: 2, reason: "proof" },
+ });
+ assert.equal(isError(res), false);
+ assert.equal(calls[0].body?.videoId, "zzzz");
+ assert.equal(calls[0].body?.channelSlug, YT);
+ assert.ok(!("webpageUrl" in (calls[0].body ?? {})));
+ assert.ok(
+ textOf(res).startsWith(
+ 'note: "zzzz" was not found in corpus local:/srv/fixture; the id was passed to ' +
+ "the editor as-is (for a Rumble citation this may be the embed id — pass the " +
+ "corpus the citation came from as source).\n\nAlready on disk",
+ ),
+ textOf(res),
+ );
+});
+
+test("a job alone is a valid call: one poll, no POST, no corpus lookup needed", async () => {
+ const { deps, calls } = fakeEditor({ WORKER_TOKEN: "tok" }, () => ({
+ status: 200,
+ body: { status: "running", jobId: "j5" },
+ }));
+ const client = await connect(deps);
+ const before = channelListings;
+ const res = await client.callTool({
+ name: "fetch_clip",
+ arguments: { job: "j5", wait_seconds: 0 },
+ });
+ assert.equal(isError(res), false, textOf(res));
+ assert.deepEqual(calls.map((c) => c.method), ["GET"]);
+ assert.equal(channelListings, before, "a resume reads nothing from the corpus");
+ assert.match(textOf(res), /^Still running on the editor \(job j5, waited 0s\)/);
+});
+
+test("a bad argument is refused before any HTTP", async () => {
+ const { deps, calls } = fakeEditor();
+ const client = await connect(deps);
+ const res = await client.callTool({
+ name: "fetch_clip",
+ arguments: { channel: YT, video: "dQw4w9WgXcQ", start: 10, end: 5, reason: "proof" },
+ });
+ assert.equal(isError(res), true);
+ assert.match(textOf(res), /^fetch_clip: start \(10\) must be less than end \(5\)/);
+ assert.equal(calls.length, 0);
+});
+
+test("the tool is advertised with job-only calls allowed", async () => {
+ const client = await connect(fakeEditor().deps);
+ const { tools } = await client.listTools();
+ const tool = tools.find((t) => t.name === "fetch_clip");
+ assert.ok(tool, "fetch_clip is listed");
+ assert.deepEqual(tool.inputSchema.required ?? [], []);
+ const props = Object.keys(tool.inputSchema.properties ?? {});
+ for (const p of ["source", "channel", "video", "start", "end", "pad", "full", "reason", "report", "wait_seconds", "job"]) {
+ assert.ok(props.includes(p), `has ${p}`);
+ }
+ assert.match(tool.description ?? "", /NEVER run yt-dlp/);
+});
+
+test("a client that asks for progress gets one notification per poll", async () => {
+ let polls = 0;
+ const { deps } = fakeEditor({ WORKER_TOKEN: "tok" }, (call) => {
+ if (call.method === "POST") {
+ return { status: 202, body: { cached: false, jobId: "j6", file: "/f.mp4", from: 7, to: 23 } };
+ }
+ polls++;
+ return polls < 3
+ ? { status: 200, body: { status: "running", jobId: "j6" } }
+ : { status: 200, body: { status: "done", jobId: "j6", file: "/f.mp4", from: 7, to: 23, bytes: 3 } };
+ });
+ const client = await connect(deps);
+ const got: { progress: number; message?: string }[] = [];
+ const res = await client.callTool(
+ {
+ name: "fetch_clip",
+ arguments: { channel: YT, video: "dQw4w9WgXcQ", start: 10, end: 20, reason: "proof" },
+ },
+ { onprogress: (p) => got.push({ progress: p.progress, message: p.message }), resetTimeoutOnProgress: true },
+ );
+ assert.equal(isError(res), false, textOf(res));
+ assert.deepEqual(got, [
+ { progress: 1, message: "editor job j6: running, 1s waited" },
+ { progress: 2, message: "editor job j6: running, 2s waited" },
+ ]);
+});
diff --git a/mcp/src/fetchClip.ts b/mcp/src/fetchClip.ts
@@ -0,0 +1,721 @@
+import { MAX_CLIP_WINDOW_SECONDS } from "yt-dlp-transcript-common/lib/clipWindow";
+
+// ─── fetch_clip: ask the local Archilyzer editor for a clip's media ───
+//
+// The operator's rule is that no fetch happens by hand: every byte of media
+// goes through the editor, which has the cookie policy, the per-platform
+// sleeps, the 429 cooldown and the provenance note. umtool already obeys it
+// (umtool/report-to-video/fetch-via-editor.mjs); this is the same client for
+// the MCP, so an agent that needs the seconds behind a citation asks the
+// editor instead of shelling out to yt-dlp.
+//
+// THE MCP PROCESS STILL WRITES NOTHING. Everything here is HTTP through
+// `deps.fetch`; the editor decides where the bytes go and says so. The file it
+// names is a corpus artifact the agent may read, never one it may move.
+//
+// Pure apart from the injected deps (env, fetch, sleep, now), so the tests need
+// no network, no timers and no editor. No MCP imports: server.ts adapts.
+
+export type HttpInit = {
+ method?: string;
+ headers?: Record<string, string>;
+ body?: string;
+ signal?: AbortSignal;
+};
+export type HttpResponse = { status: number; json(): Promise<unknown> };
+export type FetchLike = (url: string, init?: HttpInit) => Promise<HttpResponse>;
+
+export type FetchClipDeps = {
+ env: Record<string, string | undefined>;
+ fetch: FetchLike;
+ sleep: (ms: number) => Promise<void>;
+ now: () => number;
+ // Per-request bound (default REQUEST_TIMEOUT_MS). Only tests shorten it.
+ requestTimeoutMs?: number;
+};
+
+// The client family's names and defaults (fetch-via-editor.mjs,
+// scripts/archilyzer-ops.mjs): one editor URL, one shared token.
+export const DEFAULT_EDITOR_URL = "http://localhost:3001";
+export const POLL_MS = 1000;
+// EVERY request — the POST and each poll — gives up after this. Without it a
+// POST that stats a hung mount, or an editor whose event loop has stalled,
+// holds the call until undici's own 300 s headers timeout, and wait_seconds
+// bounds nothing. A healthy editor answers both in milliseconds: the POST only
+// checks the cache and queues, the poll only reads the job's state.
+export const REQUEST_TIMEOUT_MS = 15_000;
+export const DEFAULT_PAD_SECONDS = 3;
+export const DEFAULT_WAIT_SECONDS = 90;
+export const MAX_WAIT_SECONDS = 300;
+export const MAX_REASON_CHARS = 400;
+// Who asked, as the editor records it beside the file.
+export const REQUESTED_BY = "mcp";
+// A read-back window can sit a hair off the request (2 dp names); the same
+// tolerance as common/lib/clipWindow.ts WIN_EPS.
+const SAME_WINDOW_EPS = 0.02;
+
+export const NO_EDITOR_TEXT =
+ "fetch_clip: no editor configured. Set ARCHILYZER_EDITOR_URL (e.g. " +
+ "http://localhost:3001) and WORKER_TOKEN (the editor's own WORKER_TOKEN) " +
+ "when registering the MCP server. A public-only setup — no local Archilyzer " +
+ "editor — cannot fetch media through Archilyzer; see README \"Clips and " +
+ "video\" for the no-editor fallback.";
+
+// The editor's own id grammar (editor/app/api/media/fetch-window/route.ts):
+// anchored, and `.` / `..` refused separately because the class allows a dot.
+const ID_RE = /^[\w.-]+$/;
+export function isVideoId(v: string): boolean {
+ return ID_RE.test(v) && v !== "." && v !== "..";
+}
+
+export type EditorConfig = { url: string; token: string };
+
+// The editor to ask, or null when this MCP was registered without a token —
+// which is every public-only setup, and is not an error until a fetch is asked.
+export function editorFromEnv(
+ env: Record<string, string | undefined>,
+): EditorConfig | null {
+ const token = (env.WORKER_TOKEN ?? "").trim();
+ if (!token) return null;
+ const raw = (env.ARCHILYZER_EDITOR_URL ?? "").trim() || DEFAULT_EDITOR_URL;
+ return { url: raw.replace(/\/+$/, ""), token };
+}
+
+// A citation time: a number of seconds, or "ss", "mm:ss", "h:mm:ss" (a
+// fraction allowed on the seconds). Minutes may run past 59 in the two-part
+// form, because "75:30" is how a long video's moment is often written.
+export function parseSeconds(v: unknown): number | null {
+ if (typeof v === "number") return Number.isFinite(v) && v >= 0 ? v : null;
+ if (typeof v !== "string") return null;
+ const s = v.trim();
+ let m = /^(\d+(?:\.\d+)?)$/.exec(s);
+ if (m) return Number(m[1]);
+ m = /^(\d+):([0-5]?\d(?:\.\d+)?)$/.exec(s);
+ if (m) return Number(m[1]) * 60 + Number(m[2]);
+ m = /^(\d+):([0-5]?\d):([0-5]?\d(?:\.\d+)?)$/.exec(s);
+ if (m) return Number(m[1]) * 3600 + Number(m[2]) * 60 + Number(m[3]);
+ return null;
+}
+
+// Seconds as the editor names a window: two decimals, always.
+export function fmtSeconds(n: number): string {
+ return n.toFixed(2);
+}
+function fmtSpan(n: number): string {
+ return `${Number(n.toFixed(2))}s`;
+}
+
+// The padded window, exactly as umtool's client computes it: the name IS the
+// window, so a request that rounds differently addresses a different file and
+// the cache misses forever. The 900 s cap applies AFTER padding — it is the
+// span the editor is asked for.
+export function planWindow(a: {
+ video: string;
+ start: number;
+ end: number;
+ pad: number;
+}): { from: number; to: number } | { error: string } {
+ if (!isVideoId(a.video)) {
+ return { error: `fetch_clip: video "${a.video}" must match /^[\\w.-]+$/` };
+ }
+ if (!(a.start < a.end)) {
+ return {
+ error: `fetch_clip: start (${a.start}) must be less than end (${a.end})`,
+ };
+ }
+ const from = Number(Math.max(0, a.start - a.pad).toFixed(2));
+ const to = Number((a.end + a.pad).toFixed(2));
+ if (!(from < to)) {
+ return { error: `fetch_clip: start (${a.start}) must be less than end (${a.end})` };
+ }
+ const span = Number((to - from).toFixed(2));
+ if (span > MAX_CLIP_WINDOW_SECONDS) {
+ return {
+ error:
+ `fetch_clip: the window ${fmtSeconds(from)}–${fmtSeconds(to)} is ` +
+ `${span}s; the editor fetches at most ${MAX_CLIP_WINDOW_SECONDS}s per ` +
+ `window — cite a narrower span`,
+ };
+ }
+ return { from, to };
+}
+
+// ─── Arguments ───
+
+export type ClipTarget =
+ | {
+ kind: "window";
+ channel: string;
+ video: string;
+ webpageUrl?: string;
+ from: number;
+ to: number;
+ pad: number;
+ }
+ | { kind: "full"; channel: string; video: string };
+
+export type FetchClipRequest =
+ | { job: string; waitSeconds: number }
+ | {
+ target: ClipTarget;
+ reason: string;
+ report?: string;
+ waitSeconds: number;
+ };
+
+const trimmed = (v: unknown): string =>
+ typeof v === "string" ? v.trim() : "";
+
+function waitSecondsOf(v: unknown): number {
+ const n = typeof v === "number" ? v : Number.NaN;
+ if (!Number.isFinite(n)) return DEFAULT_WAIT_SECONDS;
+ return Math.min(MAX_WAIT_SECONDS, Math.max(0, n));
+}
+
+// Validate the tool's raw arguments into a request, before any HTTP. A `job`
+// alone is a whole request (resume): every other argument is then ignored. In
+// full mode start/end/pad are ignored — not required, not validated. The
+// returned video id is the one AS CITED; server.ts maps it to the editor's
+// directory id (Rumble) before fetching.
+export function validateFetchClipArgs(
+ args: Record<string, unknown>,
+): { ok: true; request: FetchClipRequest } | { ok: false; error: string } {
+ const waitSeconds = waitSecondsOf(args.wait_seconds);
+ const job = trimmed(args.job);
+ if (job) return { ok: true, request: { job, waitSeconds } };
+
+ const channel = trimmed(args.channel);
+ if (!channel) {
+ return { ok: false, error: "fetch_clip: channel is required (the channel slug)" };
+ }
+ const video = trimmed(args.video);
+ if (!video) {
+ return {
+ ok: false,
+ error: "fetch_clip: video is required (the archive's video id)",
+ };
+ }
+ if (!isVideoId(video)) {
+ return { ok: false, error: `fetch_clip: video "${video}" must match /^[\\w.-]+$/` };
+ }
+
+ let target: ClipTarget;
+ if (args.full === true) {
+ target = { kind: "full", channel, video };
+ } else {
+ const times: Record<"start" | "end", number> = { start: 0, end: 0 };
+ for (const key of ["start", "end"] as const) {
+ const raw = args[key];
+ if (raw === undefined || raw === null || raw === "") {
+ return {
+ ok: false,
+ error:
+ `fetch_clip: ${key} is required (seconds, mm:ss or h:mm:ss) — or ` +
+ `pass full: true for the whole recording`,
+ };
+ }
+ const secs = parseSeconds(raw);
+ if (secs === null) {
+ return {
+ ok: false,
+ error:
+ `fetch_clip: ${key} "${String(raw)}" is not a time (use seconds, ` +
+ `mm:ss or h:mm:ss)`,
+ };
+ }
+ times[key] = secs;
+ }
+ const pad = args.pad === undefined ? DEFAULT_PAD_SECONDS : args.pad;
+ if (typeof pad !== "number" || !Number.isFinite(pad) || pad < 0) {
+ return {
+ ok: false,
+ error: `fetch_clip: pad "${String(pad)}" must be a finite number of seconds, 0 or more`,
+ };
+ }
+ const w = planWindow({ video, start: times.start, end: times.end, pad });
+ if ("error" in w) return { ok: false, error: w.error };
+ target = { kind: "window", channel, video, from: w.from, to: w.to, pad };
+ }
+
+ const reason = trimmed(args.reason);
+ if (!reason) {
+ return {
+ ok: false,
+ error:
+ "fetch_clip: reason is required — one line saying why these seconds " +
+ "are needed (it is stored beside the file)",
+ };
+ }
+ const report = trimmed(args.report) || undefined;
+ return { ok: true, request: { target, reason, report, waitSeconds } };
+}
+
+// ─── The HTTP exchange ───
+
+export type FetchClipOutcome =
+ | { kind: "no_editor" }
+ | {
+ kind: "cached";
+ mode: "window" | "full";
+ // The span that was asked for (window mode).
+ reqFrom: number;
+ reqTo: number;
+ file: string;
+ from: number;
+ to: number;
+ bytes: number;
+ requestedBy?: string;
+ }
+ | {
+ kind: "fetched";
+ // For a resume, read off the poll's answer: a window job's carries
+ // from/to, a whole-recording job's does not.
+ mode: "window" | "full";
+ jobId: string;
+ file?: string;
+ from?: number;
+ to?: number;
+ bytes?: number;
+ waited: number;
+ }
+ | { kind: "queued"; jobId: string; status: string; waited: number }
+ | { kind: "cooldown"; platform: string; cooldownMs: number; error: string }
+ | {
+ kind: "refused";
+ phase: "post" | "poll";
+ status: number;
+ error: string;
+ jobId?: string;
+ full?: boolean;
+ }
+ | { kind: "failed"; jobId: string; status: string; log: string }
+ // `jobId` when the editor stopped answering WHILE POLLING: the job was
+ // queued and may still be running, so the answer must hand back the id to
+ // resume with. Repeating the original request instead would queue a second
+ // fetch — for full: true, a second whole-recording download.
+ | {
+ kind: "unreachable";
+ url: string;
+ message: string;
+ jobId?: string;
+ // The POST timed out: the editor may have queued the fetch anyway.
+ postTimedOut?: boolean;
+ };
+
+type Json = Record<string, unknown>;
+
+function isTimeout(e: unknown): boolean {
+ return typeof e === "object" && e !== null && (e as { name?: unknown }).name === "TimeoutError";
+}
+
+// What went wrong, in words the operator can act on. Node's fetch throws a
+// bare `TypeError: fetch failed` and keeps the reason (ECONNREFUSED, ENOTFOUND,
+// a reset) in `cause`, so the cause is appended; a timeout says how long.
+export function describeFetchError(e: unknown, timeoutMs: number): string {
+ if (isTimeout(e)) return `no answer within ${timeoutMs / 1000} s (request timed out)`;
+ const obj = typeof e === "object" && e !== null ? (e as { message?: unknown; cause?: unknown }) : null;
+ const message = obj && typeof obj.message === "string" ? obj.message : String(e);
+ const cause = obj?.cause;
+ const causeText =
+ typeof cause === "string"
+ ? cause
+ : typeof cause === "object" && cause !== null &&
+ typeof (cause as { message?: unknown }).message === "string"
+ ? (cause as { message: string }).message
+ : "";
+ return causeText && causeText !== message ? `${message} (${causeText})` : message;
+}
+
+async function readJson(res: HttpResponse): Promise<Json> {
+ try {
+ const body = await res.json();
+ return body && typeof body === "object" ? (body as Json) : {};
+ } catch {
+ return {};
+ }
+}
+
+const num = (v: unknown): number | undefined =>
+ typeof v === "number" && Number.isFinite(v) ? v : undefined;
+const str = (v: unknown): string | undefined =>
+ typeof v === "string" && v !== "" ? v : undefined;
+
+// Told after every poll that finds the job still waiting — server.ts turns it
+// into an MCP progress notification, which keeps a client's reset-on-progress
+// request timeout alive through a long wait.
+export type PollProgress = {
+ jobId: string;
+ status: string;
+ polls: number;
+ waited: number;
+};
+
+// POST the request (unless it is a resume), then poll every POLL_MS until the
+// job is terminal or the wait runs out. A wait that runs out is not a failure:
+// the job keeps running on the editor, and the answer says how to resume.
+export async function fetchClip(
+ request: FetchClipRequest,
+ deps: FetchClipDeps,
+ hooks: { onPoll?: (p: PollProgress) => void | Promise<void> } = {},
+): Promise<FetchClipOutcome> {
+ const editor = editorFromEnv(deps.env);
+ if (!editor) return { kind: "no_editor" };
+ const headers = {
+ authorization: `Bearer ${editor.token}`,
+ "content-type": "application/json",
+ };
+ const started = deps.now();
+ const deadline = started + request.waitSeconds * 1000;
+ const waited = () => Math.round((deps.now() - started) / 1000);
+
+ const timeoutMs = deps.requestTimeoutMs ?? REQUEST_TIMEOUT_MS;
+ async function call(
+ url: string,
+ init: HttpInit,
+ ): Promise<HttpResponse | { failed: string; timedOut: boolean }> {
+ try {
+ return await deps.fetch(url, { ...init, signal: AbortSignal.timeout(timeoutMs) });
+ } catch (e) {
+ return { failed: describeFetchError(e, timeoutMs), timedOut: isTimeout(e) };
+ }
+ }
+
+ // Poll until terminal or out of time. `pollFirst` is a resume: the job may
+ // well have finished since it was queued, so ask before sleeping.
+ async function waitFor(
+ jobId: string,
+ mode: "window" | "full" | "unknown",
+ pollFirst: boolean,
+ ): Promise<FetchClipOutcome> {
+ let status = "queued";
+ let skipSleep = pollFirst;
+ let polls = 0;
+ for (;;) {
+ if (!skipSleep) {
+ if (deps.now() >= deadline) {
+ return { kind: "queued", jobId, status, waited: waited() };
+ }
+ await deps.sleep(POLL_MS);
+ }
+ skipSleep = false;
+ const res = await call(
+ `${editor!.url}/api/media/fetch-window/${encodeURIComponent(jobId)}`,
+ { method: "GET", headers },
+ );
+ if ("failed" in res) {
+ return { kind: "unreachable", url: editor!.url, message: res.failed, jobId };
+ }
+ const body = await readJson(res);
+ if (res.status < 200 || res.status >= 300) {
+ return {
+ kind: "refused",
+ phase: "poll",
+ status: res.status,
+ error: str(body.error) ?? "no reason given",
+ jobId,
+ };
+ }
+ status = str(body.status) ?? status;
+ if (status === "done") {
+ const from = num(body.from);
+ const to = num(body.to);
+ return {
+ kind: "fetched",
+ mode:
+ mode === "unknown"
+ ? from !== undefined && to !== undefined
+ ? "window"
+ : "full"
+ : mode,
+ jobId,
+ file: str(body.file),
+ from,
+ to,
+ bytes: num(body.bytes),
+ waited: waited(),
+ };
+ }
+ if (status === "failed" || status === "cancelled") {
+ return { kind: "failed", jobId, status, log: str(body.error) ?? "" };
+ }
+ polls++;
+ if (hooks.onPoll) {
+ // A progress report is a courtesy: a client that went away must not
+ // turn a running fetch into an error.
+ try {
+ await hooks.onPoll({ jobId, status, polls, waited: waited() });
+ } catch {
+ /* ignored */
+ }
+ }
+ }
+ }
+
+ if ("job" in request) return waitFor(request.job, "unknown", true);
+
+ const { target } = request;
+ const provenance = {
+ requestedBy: REQUESTED_BY,
+ manifest: request.report,
+ reason: request.reason.slice(0, MAX_REASON_CHARS),
+ };
+ // `full` REPLACES the span (as in fetch-via-editor.mjs): the route reads
+ // `full === true` before it validates from/to, and the saved-video path
+ // resolves the URL itself, so a webpageUrl would describe a request the
+ // editor does not have.
+ const body =
+ target.kind === "full"
+ ? {
+ channelSlug: target.channel,
+ videoId: target.video,
+ full: true,
+ ...provenance,
+ }
+ : {
+ channelSlug: target.channel,
+ videoId: target.video,
+ webpageUrl: target.webpageUrl,
+ from: target.from,
+ to: target.to,
+ pad: target.pad,
+ ...provenance,
+ };
+ const res = await call(`${editor.url}/api/media/fetch-window`, {
+ method: "POST",
+ headers,
+ body: JSON.stringify(body),
+ });
+ if ("failed" in res) {
+ return {
+ kind: "unreachable",
+ url: editor.url,
+ message: res.failed,
+ postTimedOut: res.timedOut,
+ };
+ }
+ const answer = await readJson(res);
+ const reqFrom = target.kind === "window" ? target.from : 0;
+ const reqTo = target.kind === "window" ? target.to : 0;
+
+ if (res.status === 200 && str(answer.file)) {
+ const prov = answer.provenance as Json | null | undefined;
+ return {
+ kind: "cached",
+ mode: target.kind,
+ reqFrom,
+ reqTo,
+ file: str(answer.file)!,
+ from: num(answer.from) ?? reqFrom,
+ to: num(answer.to) ?? reqTo,
+ bytes: num(answer.bytes) ?? 0,
+ requestedBy: prov && typeof prov === "object" ? str(prov.requestedBy) : undefined,
+ };
+ }
+ if (res.status === 202 && str(answer.jobId)) {
+ return waitFor(str(answer.jobId)!, target.kind, false);
+ }
+ if (res.status === 409) {
+ return {
+ kind: "cooldown",
+ platform: str(answer.platform) ?? "platform",
+ cooldownMs: num(answer.cooldownMs) ?? 0,
+ error: str(answer.error) ?? "",
+ };
+ }
+ return {
+ kind: "refused",
+ phase: "post",
+ status: res.status,
+ error: str(answer.error) ?? "no reason given",
+ full: target.kind === "full",
+ };
+}
+
+// ─── The answer an agent reads ───
+
+const READ_ONLY_NOTE =
+ "This path is a read-only corpus artifact: play or copy it, never move, " +
+ "edit or delete it.";
+
+function footer(
+ mode: "window" | "full",
+ f: { file: string; from?: number; to?: number; bytes?: number; requestedBy?: string },
+): string {
+ const lines = [`file: ${f.file}`];
+ if (mode === "window" && f.from !== undefined && f.to !== undefined) {
+ lines.push(
+ `window: ${fmtSeconds(f.from)}–${fmtSeconds(f.to)} (${fmtSpan(f.to - f.from)})`,
+ );
+ }
+ if (f.bytes !== undefined) lines.push(`bytes: ${f.bytes}`);
+ if (f.requestedBy) lines.push(`requested by ${f.requestedBy}`);
+ if (mode === "window") {
+ const name =
+ f.from !== undefined && f.to !== undefined
+ ? `${fmtSeconds(f.from)}-${fmtSeconds(f.to)}.json`
+ : "<from>-<to>.json";
+ lines.push(
+ `${READ_ONLY_NOTE} The editor prunes clips by age (evict-clips); ` +
+ `provenance sits beside it as ${name}.`,
+ );
+ } else {
+ lines.push(
+ `${READ_ONLY_NOTE} It lives in the editor's saved-video store, not in ` +
+ `clips/; the editor's keep-videos rule decides how long it stays.`,
+ );
+ }
+ return lines.join("\n");
+}
+
+const TOKEN_HINT =
+ "WORKER_TOKEN must equal the value the editor runs with (editor/.env).";
+
+// Only a 503 that SAYS the endpoint is disabled is a token problem: the same
+// status also carries an unreachable-media refusal (a relocated channel whose
+// drive is not mounted), which is the editor operator's to fix, verbatim.
+function isTokenRefusal(status: number, error: string): boolean {
+ return status === 401 || (status === 503 && /disabled/i.test(error));
+}
+
+export function renderFetchClip(
+ outcome: FetchClipOutcome,
+ ctx: { channel?: string; video?: string } = {},
+): { text: string; isError: boolean } {
+ const of = ctx.channel && ctx.video ? ` of ${ctx.channel}/${ctx.video}` : "";
+ switch (outcome.kind) {
+ case "no_editor":
+ return { text: NO_EDITOR_TEXT, isError: true };
+ case "cached": {
+ let head: string;
+ if (outcome.mode === "full") {
+ head = `Already on disk — the whole recording${of}.`;
+ } else {
+ const same =
+ Math.abs(outcome.from - outcome.reqFrom) <= SAME_WINDOW_EPS &&
+ Math.abs(outcome.to - outcome.reqTo) <= SAME_WINDOW_EPS;
+ head = same
+ ? `Already on disk — the exact window: ${fmtSeconds(outcome.from)}–${fmtSeconds(outcome.to)}.`
+ : `Already on disk — a WIDER cached window that contains ` +
+ `${fmtSeconds(outcome.reqFrom)}–${fmtSeconds(outcome.reqTo)}: ` +
+ `${fmtSeconds(outcome.from)}–${fmtSeconds(outcome.to)}.`;
+ }
+ return { text: `${head}\n\n${footer(outcome.mode, outcome)}`, isError: false };
+ }
+ case "fetched": {
+ if (!outcome.file) {
+ return {
+ text:
+ `Editor job ${outcome.jobId} finished but named no file; check ` +
+ `the video's page in the editor.`,
+ isError: true,
+ };
+ }
+ const head =
+ outcome.mode === "full"
+ ? `Fetched the whole recording${of} (job ${outcome.jobId}, ${outcome.waited}s waited).`
+ : `Fetched ${
+ outcome.from !== undefined && outcome.to !== undefined
+ ? `${fmtSeconds(outcome.from)}–${fmtSeconds(outcome.to)}`
+ : "the window"
+ }${of} (job ${outcome.jobId}, ${outcome.waited}s waited).`;
+ return {
+ text: `${head}\n\n${footer(outcome.mode, { ...outcome, file: outcome.file })}`,
+ isError: false,
+ };
+ }
+ case "queued":
+ return {
+ text:
+ `Still ${outcome.status} on the editor (job ${outcome.jobId}, waited ` +
+ `${outcome.waited}s). Call fetch_clip again with job: ` +
+ `"${outcome.jobId}" to keep waiting — never repeat the original ` +
+ `request while it runs (that would queue a second fetch). Nothing is ` +
+ `lost: the fetch continues on the editor, and once it has finished ` +
+ `the same request finds it cached.`,
+ isError: false,
+ };
+ case "cooldown":
+ return {
+ text:
+ `The editor is in a ${outcome.platform} rate-limit cooldown — ` +
+ `${Math.ceil(outcome.cooldownMs / 1000)}s remaining. Wait, then call ` +
+ `fetch_clip again. (${outcome.error})`,
+ isError: true,
+ };
+ case "refused": {
+ if (outcome.phase === "poll") {
+ const extra =
+ outcome.status === 404
+ ? "; the job is unknown to this editor (restarted? wrong ARCHILYZER_EDITOR_URL?)"
+ : isTokenRefusal(outcome.status, outcome.error)
+ ? `. ${TOKEN_HINT}`
+ : "";
+ return {
+ text:
+ `Polling editor job ${outcome.jobId} failed (HTTP ${outcome.status}): ` +
+ `${outcome.error}${extra}`,
+ isError: true,
+ };
+ }
+ if (isTokenRefusal(outcome.status, outcome.error)) {
+ const off =
+ outcome.status === 503
+ ? " The editor has no WORKER_TOKEN set, so its fetch endpoint is off."
+ : "";
+ return {
+ text:
+ `The editor refused the token (HTTP ${outcome.status}): ` +
+ `${outcome.error}. ${TOKEN_HINT}${off}`,
+ isError: true,
+ };
+ }
+ const known =
+ outcome.full && outcome.status === 404
+ ? " A whole-recording fetch needs a video the editor already knows " +
+ "(a metadata.info.json or a playlist entry); fetch a window " +
+ "instead, or add the video to the channel first."
+ : "";
+ return {
+ text: `The editor refused (HTTP ${outcome.status}): ${outcome.error}${known}`,
+ isError: true,
+ };
+ }
+ case "failed":
+ return {
+ text:
+ `Editor job ${outcome.jobId} ${outcome.status}.\n\nlog tail:\n` +
+ (outcome.log || "(the job left no log)"),
+ isError: true,
+ };
+ case "unreachable":
+ if (outcome.jobId) {
+ return {
+ text:
+ `Could not reach the editor at ${outcome.url} while polling job ` +
+ `${outcome.jobId}: ${outcome.message}. Is it running? Call ` +
+ `fetch_clip again with job: "${outcome.jobId}" — do not repeat the ` +
+ `original request, the fetch may still be running.`,
+ isError: true,
+ };
+ }
+ return {
+ text:
+ `Could not reach the editor at ${outcome.url}: ${outcome.message}. Is it running?` +
+ (outcome.postTimedOut
+ ? " It may have queued the fetch anyway: check the editor's /jobs page " +
+ "before asking again."
+ : ""),
+ isError: true,
+ };
+ }
+}
+
+// The prefix for a cited id the corpus does not hold: the id goes to the editor
+// as-is, which for a Rumble EMBED id is the wrong directory.
+export function notFoundNote(video: string, handle: string): string {
+ return (
+ `note: "${video}" was not found in corpus ${handle}; the id was passed to ` +
+ `the editor as-is (for a Rumble citation this may be the embed id — pass ` +
+ `the corpus the citation came from as source).`
+ );
+}
diff --git a/mcp/src/instructions.test.ts b/mcp/src/instructions.test.ts
@@ -188,3 +188,29 @@ test("the sweep reads operator corrections from a sibling manifest before the fi
assert.ok(at > 0 && batch > at, `corrections step at ${at}, batch step at ${batch}`);
assert.match(text, /never re-asserted/);
});
+
+// ─── Media for a citation ───
+
+test("both builders send clip media through fetch_clip, never a hand-run yt-dlp", () => {
+ for (const text of [sweep("coffee channels=c"), ask("coffee")]) {
+ assert.match(text, /`fetch_clip`/);
+ assert.match(text, /NEVER run yt-dlp yourself/);
+ // The corpus rides along, so a Rumble embed id maps to the editor's dir.
+ assert.match(text, /pass source: "remote:https:\/\/site\.example" so a Rumble id/);
+ // The whole recording is on offer, but the window is the default.
+ assert.match(text, /full: true when the ask genuinely needs the whole recording/);
+ // Only the operator falls back to yt-dlp.
+ assert.match(text, /the operator's fallback, not yours/);
+ }
+ // In the sweep, before Finish: a report that needs a clip asks while the
+ // citation is in hand, not after the summary is written.
+ const s = sweep("coffee channels=c");
+ const clip = s.indexOf("Media for a cited moment");
+ const finish = s.indexOf("**Finish.**");
+ assert.ok(clip > 0 && finish > clip, `clip step at ${clip}, Finish at ${finish}`);
+ // In the ask, after the answer's citation rule.
+ const a = ask("coffee");
+ const cite = a.indexOf("Answer with citations");
+ const askClip = a.indexOf("Media for a cited moment");
+ assert.ok(cite > 0 && askClip > cite, `citations at ${cite}, clip step at ${askClip}`);
+});
diff --git a/mcp/src/instructions.ts b/mcp/src/instructions.ts
@@ -140,6 +140,25 @@ function batchStep(req: PromptRequest, ctx: PlanContext, reportPath: string): st
);
}
+// Media for a citation goes through the editor's fetch job, never a yt-dlp the
+// agent runs itself (the operator's rule: every fetch through Archilyzer). One
+// step, shared by both builders so the two cannot drift. `wait_seconds` and
+// `full` stay un-backticked: they are arguments, not tools.
+function clipStep(ctx: PlanContext): string {
+ return (
+ `**Media for a cited moment — through the editor only.** When I ask for ` +
+ `the clip behind a citation (or a report needs one), call \`fetch_clip\` ` +
+ `with that citation's channel slug, video id, start/end in seconds (or ` +
+ `mm:ss) and a one-line reason (or full: true when the ask genuinely needs ` +
+ `the whole recording — a window is the default), and pass ` +
+ `source: "${ctx.corpus}" so a Rumble id resolves to the right directory. ` +
+ `NEVER run yt-dlp yourself, in any form. If the tool reports no editor is ` +
+ `configured, say so and stop — the README's yt-dlp command is the ` +
+ `operator's fallback, not yours. If it returns queued, call it again with ` +
+ `the job it names.`
+ );
+}
+
// The full sweep instructions.
export function buildSweepInstructions(
req: PromptRequest,
@@ -236,6 +255,8 @@ export function buildSweepInstructions(
steps.push(batchStep(req, ctx, reportPath));
+ steps.push(clipStep(ctx));
+
steps.push(
`**Finish.** Work to the end of the worklist, then write a summary ` +
`section: the scope swept, **how many of the N you actually read**, ` +
@@ -329,6 +350,8 @@ export function buildAskInstructions(
`corpus does NOT show — an absence of matches is a finding, not a gap ` +
`to paper over.`,
);
+
+ steps.push(clipStep(ctx));
}
const numbered = steps.map((s, i) => `${i + 1}. ${s}`).join("\n\n");
diff --git a/mcp/src/protocol.test.ts b/mcp/src/protocol.test.ts
@@ -2,6 +2,8 @@ import { test } from "node:test";
import assert from "node:assert/strict";
import { mkdtemp, mkdir, writeFile, readdir, rm } from "node:fs/promises";
import os from "node:os";
+import http from "node:http";
+import type { AddressInfo } from "node:net";
import path from "node:path";
import { fileURLToPath } from "node:url";
import { Client } from "@modelcontextprotocol/client";
@@ -40,6 +42,7 @@ const EXPECTED_TOOLS = [
"get_thread",
"get_transcripts",
"get_video_metadata",
+ "fetch_clip",
"list_sources",
"resolve_source",
"open_link",
@@ -147,6 +150,7 @@ type Session = {
async function connect(
corpusDir: string,
versionNegotiation?: { mode: "auto" | "legacy" },
+ extraEnv: Record<string, string> = {},
): Promise<Session> {
const stateDir = await mkdtemp(path.join(os.tmpdir(), "mcp-state-"));
const transport = new StdioClientTransport({
@@ -154,7 +158,7 @@ async function connect(
args: [ENTRY, "--local", corpusDir],
// The old controller keyed a state file off this; nothing reads it now,
// and the assertion below is that the dir stays empty regardless.
- env: { ...getDefaultEnvironment(), TRANSCRIPT_MCP_STATE_DIR: stateDir },
+ env: { ...getDefaultEnvironment(), TRANSCRIPT_MCP_STATE_DIR: stateDir, ...extraEnv },
stderr: "pipe",
});
const client = new Client(
@@ -349,3 +353,48 @@ test("stdio: a read-only session writes no state file", async (t) => {
const left = await readdir(s.stateDir).catch(() => [] as string[]);
assert.deepEqual(left, [], "the server must not persist anything");
});
+
+// fetch_clip over the real process, on the modern era: the env it was
+// registered with reaches the editor as the bearer token, and a client that
+// asks for progress gets one notification per poll — the keep-alive for a
+// client whose request timeout resets on progress. The editor is a stub on an
+// ephemeral port; a resume by `job` needs no corpus read.
+test("stdio (modern): fetch_clip asks the editor with its registered token and reports progress", async (t) => {
+ const seen: string[] = [];
+ let polls = 0;
+ const editor = http.createServer((req, res) => {
+ seen.push(`${req.method} ${req.url} ${req.headers.authorization}`);
+ polls++;
+ const body =
+ polls < 3
+ ? { status: "running", jobId: "j1" }
+ : { status: "done", jobId: "j1", file: "/corpus/clips/7.00-23.00.mp4", from: 7, to: 23, bytes: 5 };
+ res.writeHead(200, { "content-type": "application/json" });
+ res.end(JSON.stringify(body));
+ });
+ await new Promise<void>((r) => editor.listen(0, "127.0.0.1", r));
+ t.after(() => new Promise<void>((r) => editor.close(() => r())));
+ const { port } = editor.address() as AddressInfo;
+
+ const dir = await writeFixture();
+ t.after(() => rm(dir, { recursive: true, force: true }));
+ const s = await connect(dir, { mode: "auto" }, {
+ ARCHILYZER_EDITOR_URL: `http://127.0.0.1:${port}`,
+ WORKER_TOKEN: "tok-proto",
+ });
+ t.after(() => s.close());
+ assert.equal(s.client.getProtocolEra(), "modern");
+
+ const progress: number[] = [];
+ const res = await s.client.callTool(
+ { name: "fetch_clip", arguments: { job: "j1" } },
+ { onprogress: (p) => progress.push(p.progress), resetTimeoutOnProgress: true },
+ );
+ assert.match(firstText(res), /^Fetched 7\.00–23\.00 \(job j1, \d+s waited\)\./);
+ assert.deepEqual(seen, [
+ "GET /api/media/fetch-window/j1 Bearer tok-proto",
+ "GET /api/media/fetch-window/j1 Bearer tok-proto",
+ "GET /api/media/fetch-window/j1 Bearer tok-proto",
+ ]);
+ assert.deepEqual(progress, [1, 2]);
+});
diff --git a/mcp/src/server.ts b/mcp/src/server.ts
@@ -62,6 +62,18 @@ import {
type PlanContext,
} from "./instructions";
import {
+ fetchClip,
+ renderFetchClip,
+ validateFetchClipArgs,
+ editorFromEnv,
+ notFoundNote,
+ isVideoId,
+ NO_EDITOR_TEXT,
+ type FetchClipDeps,
+ type PollProgress,
+} from "./fetchClip";
+import { extractVideoId } from "yt-dlp-transcript-common/lib/videoId";
+import {
decodeShareLink,
applyLinkOverrides,
renderQueryTree,
@@ -702,6 +714,94 @@ export const TOOLS: Tool[] = [
},
},
{
+ name: "fetch_clip",
+ description:
+ "Get the media behind a cited moment — by asking the local Archilyzer " +
+ "editor, which fetches it through its own paced, cookie-aware, " +
+ "provenanced job (the per-platform sleeps, the rate-limit cooldown, a " +
+ "note beside the file saying who asked and why). NEVER run yt-dlp " +
+ "yourself instead. Give the citation's channel slug, video id, start " +
+ "and end, and a one-line reason; the window is padded (pad, default 3 " +
+ "s) and may be at most 15 min. full: true fetches the whole recording " +
+ "instead, into the editor's saved-video store. The answer names the " +
+ "file on disk: a read-only corpus artifact to play or copy, never to " +
+ "move, edit or delete. The call waits up to wait_seconds; if the fetch " +
+ "is still running it returns the job id — call again with job to keep " +
+ "waiting. It sends a progress notification per poll when the client asks " +
+ "for progress. A client with a 60 s default request timeout must raise " +
+ "it or pass wait_seconds ≤ 50 — the fetch continues on the editor either " +
+ "way; resume it with job, and once it has finished the same request " +
+ "finds it cached. The editor must already archive " +
+ "the cited channel. Needs ARCHILYZER_EDITOR_URL and WORKER_TOKEN in this server's " +
+ "environment; without them it says so and fetches nothing.",
+ inputSchema: {
+ type: "object",
+ properties: {
+ ...SOURCE_ARG,
+ channel: {
+ type: "string",
+ description: "The channel slug, as the citation names it.",
+ },
+ video: {
+ type: "string",
+ description:
+ "The archive's video id, as cited. Pass the corpus the citation " +
+ "came from as `source`, so a Rumble embed id resolves to the " +
+ "editor's directory.",
+ },
+ start: {
+ oneOf: [{ type: "number" }, { type: "string" }],
+ description: "Start of the cited span: seconds, or mm:ss / h:mm:ss.",
+ },
+ end: {
+ oneOf: [{ type: "number" }, { type: "string" }],
+ description: "End of the cited span: seconds, or mm:ss / h:mm:ss.",
+ },
+ pad: {
+ type: "number",
+ description:
+ "Seconds added before start and after end (default 3). The padded " +
+ "window is what is fetched, and it may be at most 900 s.",
+ },
+ full: {
+ type: "boolean",
+ description:
+ "The whole recording instead of a window (default false) — for a " +
+ "video that must be re-cut freely or watched end to end. It lands " +
+ "in the editor's saved-video store, is much larger than a window, " +
+ "and needs a video the editor already knows (its metadata or " +
+ "playlist entry). start, end and pad are ignored.",
+ },
+ reason: {
+ type: "string",
+ description:
+ "Required: one line saying why these seconds are needed. Stored " +
+ "beside the file (at most 400 characters are kept).",
+ },
+ report: {
+ type: "string",
+ description:
+ "Optional: the report or manifest this clip is for, recorded with " +
+ "the file.",
+ },
+ wait_seconds: {
+ type: "number",
+ description:
+ "How long to wait for the fetch before returning its job id " +
+ "(default 90, max 300). Nothing is lost when it runs out.",
+ },
+ job: {
+ type: "string",
+ description:
+ "Resume: the job id an earlier call returned. When set, every " +
+ "other argument is ignored and the call just waits on that job.",
+ },
+ },
+ required: [],
+ additionalProperties: false,
+ },
+ },
+ {
name: "list_sources",
description:
"Show the corpora this server can read: its DEFAULT corpus (the one used " +
@@ -916,6 +1016,93 @@ function withCorpus(result: ToolResult, resolved: ResolvedSource): ToolResult {
return { ...result, content };
}
+// The real world for fetch_clip. `env` is process.env itself, so a token set
+// after startup is still seen.
+const DEFAULT_FETCH_CLIP_DEPS: FetchClipDeps = {
+ env: process.env,
+ fetch: (url, init) => globalThis.fetch(url, init),
+ sleep: (ms) => new Promise((resolve) => setTimeout(resolve, ms)),
+ now: () => Date.now(),
+};
+
+// fetch_clip: validate, map the cited id to the editor's directory id, ask the
+// editor, and render its answer. The one tool that causes a write — and the
+// editor makes it, not this process.
+//
+// THE RUMBLE TRAP. A published record's `id` is yt-dlp's native id, which on
+// Rumble is the EMBED id (`vxe1ae`); the editor's directory is the canonical id
+// from the URL slug (`v1007ay`). Passing the embed id would make the editor
+// create a NEW data/vxe1ae/ and fetch rumble.com/vxe1ae. The record's
+// `webpageUrl` is the slug URL, so the canonical id comes from it — the same
+// `extractVideoId` the editor names its directories with. On YouTube the two
+// are equal. A video the corpus does not hold goes through as cited, with a
+// note saying what that risks.
+async function handleFetchClip(
+ source: ShardSource,
+ resolved: ResolvedSource,
+ args: Record<string, unknown>,
+ deps: FetchClipDeps,
+ onPoll?: (p: PollProgress) => Promise<void>,
+): Promise<ToolResult> {
+ // No editor, no fetch — said first, before a corpus read that could only be
+ // wasted.
+ if (!editorFromEnv(deps.env)) return errorText(NO_EDITOR_TEXT);
+ const v = validateFetchClipArgs(args);
+ if (!v.ok) return errorText(v.error);
+ const request = v.request;
+
+ let note = "";
+ let ctx: { channel?: string; video?: string } = {};
+ if ("target" in request) {
+ const target = request.target;
+ const found = await findVideo(source, target.video, target.channel);
+ if (found) {
+ const webpageUrl = found.record.webpageUrl;
+ const mapped = webpageUrl ? extractVideoId(webpageUrl) : null;
+ target.video = mapped && isVideoId(mapped) ? mapped : target.video;
+ // The channel the record lives in, which is the editor's directory even
+ // when the citation spelled it by name or in another case.
+ target.channel = found.ch.slug;
+ if (target.kind === "window" && webpageUrl) target.webpageUrl = webpageUrl;
+ } else {
+ note = `${notFoundNote(target.video, resolved.handle)}\n\n`;
+ }
+ ctx = { channel: target.channel, video: target.video };
+ }
+
+ const outcome = await fetchClip(request, deps, { onPoll });
+ const rendered = renderFetchClip(outcome, ctx);
+ const body = `${note}${rendered.text}`;
+ return rendered.isError ? errorText(body) : text(body);
+}
+
+// One MCP progress notification per poll while fetch_clip waits, when the
+// client asked for progress (a `progressToken` in the request's _meta). A
+// client whose request timeout resets on progress then keeps waiting for as
+// long as the editor keeps answering. `progress` is the poll count, so it only
+// ever increases; there is no total, because a queue's length is not knowable
+// from here. No token, no notifications.
+function progressNotifier(
+ progressToken: unknown,
+ notify: (n: {
+ method: "notifications/progress";
+ params: { progressToken: string | number; progress: number; message: string };
+ }) => Promise<void>,
+): ((p: PollProgress) => Promise<void>) | undefined {
+ if (typeof progressToken !== "string" && typeof progressToken !== "number") {
+ return undefined;
+ }
+ return (p) =>
+ notify({
+ method: "notifications/progress",
+ params: {
+ progressToken,
+ progress: p.polls,
+ message: `editor job ${p.jobId}: ${p.status}, ${p.waited}s waited`,
+ },
+ });
+}
+
// Build a configured MCP server over a data source or a SourceRegistry. The
// core read tools work for local / remote / hub sources — only the ShardSource
// differs — and which one a call reads is decided per call by its `source`
@@ -923,7 +1110,15 @@ function withCorpus(result: ToolResult, resolved: ResolvedSource): ToolResult {
// one-source registry, so `createServer(someSource)` still works.
export function createServer(
sourceOrRegistry: ShardSource | SourceRegistry,
+ opts: { fetchClipDeps?: Partial<FetchClipDeps> } = {},
): Server {
+ // fetch_clip's world — env, HTTP, the clock — injected so a test drives it
+ // with no network and no timers. Read per call, never cached: the env is the
+ // live process.env unless a test says otherwise.
+ const fetchClipDeps: FetchClipDeps = {
+ ...DEFAULT_FETCH_CLIP_DEPS,
+ ...opts.fetchClipDeps,
+ };
const registry =
sourceOrRegistry instanceof SourceRegistry
? sourceOrRegistry
@@ -946,7 +1141,7 @@ export function createServer(
return buildSweepPrompt((req.params.arguments ?? {}) as Record<string, unknown>);
});
- server.setRequestHandler("tools/call", async (req) => {
+ server.setRequestHandler("tools/call", async (req, ctx) => {
const name = req.params.name;
const args = (req.params.arguments ?? {}) as Record<string, unknown>;
@@ -987,6 +1182,14 @@ export function createServer(
return handleGetThread(source, args);
case "get_video_metadata":
return handleGetMetadata(source, args);
+ case "fetch_clip":
+ return handleFetchClip(
+ source,
+ resolved,
+ args,
+ fetchClipDeps,
+ progressNotifier(ctx.mcpReq._meta?.progressToken, ctx.mcpReq.notify),
+ );
case "list_sources":
return handleListSources(registry, resolved);
case "resolve_source":
diff --git a/plans/mcp-fetch-clip.md b/plans/mcp-fetch-clip.md
@@ -106,8 +106,13 @@ it as <from>-<to>.json.`
- 202 → done: `Fetched <from>–<to> of <channel>/<canonical> (job <id>, <n>s waited).` + footer from
the poll's `file/from/to/bytes`.
- wait expired, still queued/running (NOT `isError`): `Still <status> on the editor (job <jobId>,
- waited <n>s). Call fetch_clip again with job: "<jobId>" to keep waiting — nothing is lost, the fetch
- continues on the editor and the next ask finds it cached.`
+ waited <n>s). Call fetch_clip again with job: "<jobId>" to keep waiting — never repeat the original
+ request while it runs (that would queue a second fetch). Nothing is lost: the fetch continues on
+ the editor, and once it has finished the same request finds it cached.` (Amended 2026-09-26 after
+ the review: the editor does not dedupe a running whole-recording job — `archiveSourceVideo` →
+ `runManagedFunction` has no running-job check — so a repeated `full: true` request while it runs
+ queues a second download, and "the next ask finds it cached" was true only once the job had
+ finished.)
- 409: `The editor is in a <platform> rate-limit cooldown — <ceil(cooldownMs/1000)>s remaining. Wait,
then call fetch_clip again. (<error>)`
- 401 / 503: `The editor refused the token (HTTP <s>): <error>. WORKER_TOKEN must equal the value the
diff --git a/plans/release-10.md b/plans/release-10.md
@@ -940,6 +940,232 @@ unchanged by all three.
| `9e3047e3` | `plans:` FACTS: the two worktree workarounds superseded (L1), and the release 10 facts for S4, L1 and L2 |
| _this_ | `plans:` this record, STATE ("Merged to main (2026-09-26), NOT rolled out"), the Rollout steps |
+### Slice M, as shipped — fetch_clip (2026-09-26)
+
+Branch `mcp/fetch-clip` off `main` `4ac32a2e`, worktree `/home/user/Projects/mcp-fetch-clip`, one
+Opus implementer. The plan is [`mcp-fetch-clip.md`](mcp-fetch-clip.md), §1–§5 and its Amendment
+(`full: true` exposed). The MCP server gains ONE tool, `fetch_clip`, which asks the local editor for
+the media behind a cited moment through its existing `POST /api/media/fetch-window`. So an agent
+following `/ask` or `/sweep` no longer has to shell out to yt-dlp (unpaced, no cookies, bytes
+outside the corpus) or stop. The guidance now makes the tool the way, and `yt-dlp
+--download-sections` only the no-editor fallback. Nothing in `common/`, `editor/`, `export/`,
+`scripts/` or `umtool/` changed; no settings, site or channel key; nothing on disk.
+
+**The client** (`mcp/src/fetchClip.ts`, pure; `env`, `fetch`, `sleep` and `now` injected, no MCP
+imports).
+- `parseSeconds`: a number, `ss`, `mm:ss` (minutes may pass 59) or `h:mm:ss`.
+- `planWindow`: umtool's arithmetic, `from = max(0, start-pad)` and `to = end+pad`, each
+ `.toFixed(2)`; the 900 s cap (`MAX_CLIP_WINDOW_SECONDS`, imported) applies AFTER padding.
+- `validateFetchClipArgs`: a `job` alone is a whole request and every other argument is ignored;
+ `full: true` ignores `start`/`end`/`pad` (not required, not validated); `reason` is required in
+ both modes; `wait_seconds` is clamped to [0, 300], default 90.
+- `fetchClip`: POST (unless resuming), then poll every 1 s via `deps.sleep` until the job is
+ terminal or `deps.now()` passes the deadline. A resume polls before it sleeps. The body is
+ `{channelSlug, videoId, webpageUrl?, from, to, pad, requestedBy: "mcp", manifest, reason}`, the
+ reason cut to 400; in full mode exactly `{channelSlug, videoId, full: true, requestedBy, manifest,
+ reason}`.
+- `renderFetchClip`: the plan's texts, verbatim where the plan gave them.
+- HTTP only: the MCP process writes nothing.
+
+**The tool** (`mcp/src/server.ts`): the schema after `get_video_metadata`, `required: []`,
+`additionalProperties: false`. `createServer(source, {fetchClipDeps})` defaults to `process.env`,
+`globalThis.fetch`, a `setTimeout` sleep and `Date.now`. `source` resolves first and `withCorpus`
+adds the trailer, as for every tool. `handleFetchClip`:
+- **no editor configured** is said before argument validation and before any corpus read;
+- `findVideo(source, video, channel)`;
+- the id sent is `extractVideoId(record.webpageUrl) ?? video` (the editor's own directory naming —
+ a Rumble embed id `vxe1ae` becomes `v1007ay`), and the record's `webpageUrl` rides along in
+ window mode;
+- a video `source` does not hold goes through as cited, with the plan's `note:` prefix.
+
+**The plans** (`mcp/src/instructions.ts`): one shared `clipStep`, after "Answer with citations" in
+ask and before **Finish** in the sweep. It carries the plan's sentence with the amendment's
+clause. `wait_seconds` and `full` are not backticked, and `NOT_TOOLS` is unchanged.
+
+**Docs.**
+- `README.md`: both `claude mcp add` blocks carry the two optional `--env` lines. In "Clips and
+ video", step 3 is `fetch_clip` (window → `clips/`, `full: true` → the saved-video store), the
+ editor path no longer claims `out/clips-raw`, and yt-dlp is the no-editor fallback.
+- `mcp/README.md`: the one exception to "writes nothing", the tool row, and "Still read-only"
+ now says the editor writes.
+- `AGENTS.md`: the env lines, and the clips loop through `fetch_clip`.
+- `grep -n 'download-sections' README.md AGENTS.md mcp/README.md` gives two lines, both in
+ no-editor fallback sentences (`AGENTS.md:87`, `README.md:366`).
+
+**Where the texts depart from, or fill in, the plan** (for the reviewer):
+- **A 503 is a token problem only when its error says "disabled".** The plan put every 401/503 on
+ the token text. But `fetchWindowAction` answers 503 for an unreachable-media refusal
+ (`videoActions.ts`, the `!res.ok` after `runManagedFunction`), and `fetchFullSourceAction` for
+ any `archiveSourceVideo` error that is not low disk. Under the token text, "plug the drive in"
+ would have read as "fix your token". Such a 503 goes through the generic `The editor refused
+ (HTTP 503): <error verbatim>`.
+- **The channel sent is the record's own `ch.slug` when the video is found** (the plan: the
+ `channel` argument). A citation that spells the channel by name or in another case still reaches
+ the right directory. When the video is not found, the argument goes as given.
+- Texts the plan left open:
+ - `Already on disk — the exact window: <from>–<to>.`
+ - `requested by <x>` is a footer line, shown for a cached whole recording too (`SavedVideoOrigin`
+ carries it).
+ - The footer names the window's real sidecar (`7.00-23.00.json`).
+ - A missing start/end reads `fetch_clip: start is required (seconds, mm:ss or h:mm:ss) — or pass
+ full: true for the whole recording`.
+ - A bad pad reads `fetch_clip: pad "<p>" must be a finite number of seconds, 0 or more`.
+ - A done window job with no file uses the amendment's full-mode "finished but named no file"
+ text.
+ - A poll 401, or a disabled 503, adds the token sentence.
+ - A resume knows no channel or video, so its "Fetched …" line omits "of <channel>/<video>", and
+ window or full is read off the poll's answer (a window's carries `from`/`to`).
+ - A 409 is `isError`; only the expired wait (`Still <status> …`) is not.
+
+| sha | what |
+|---|---|
+| `55f81128` | `mcp: fetchClip` — the pure client and its texts; `fetchClip.test.ts` (30) |
+| `2cd35fbf` | `mcp: fetch_clip` tool — schema, `createServer(…, {fetchClipDeps})`, `handleFetchClip` (Rumble mapping, not-found note); `EXPECTED_TOOLS` +1; `fetchClip.tool.test.ts` (9) |
+| `345bf26d` | `mcp:` the ask and sweep plans' `clipStep`; `instructions.test.ts` +1 |
+| `a0326dfd` | `docs:` README (both blocks, "Clips and video"), `mcp/README.md`, `AGENTS.md` |
+| _this_ | `plans:` this record; the editor `[Unreleased]` bullet |
+
+**Gates**, all from the worktree root.
+- **tsc** (`pnpm -r --no-bail --workspace-concurrency=1 exec tsc --noEmit`) clean on the full tree
+ before the first commit (`m-gate1.log`). The mcp package's `tsc --noEmit` is also clean on each of
+ the three code commits checked out alone (`m-tsc-per-commit.log`). The docs commit has no code.
+- **mcp 259/259** (219 + 40: `fetchClip.test.ts` 30, `fetchClip.tool.test.ts` 9,
+ `instructions.test.ts` +1; `protocol.test.ts` still 8 tests, its list now 15 names), 23 s.
+- **`test:scripts` 173 + 1 skip of 174**; **common 1,954/1,954** (`m-gate2.log`). Both are `main`'s
+ counts: the prompt's 162 + 1 and 1,845 predate L1, L2 and S4. This slice changes neither
+ (`git diff --stat 4ac32a2e a0326dfd -- common scripts umtool editor export homepage` is empty;
+ the `plans:` commit adds only the `editor/CHANGELOG.md` line).
+- **Builds and e2e: none, by design.** No editor, export, homepage or umtool code changed, and no
+ e2e spec names an MCP tool: `grep -rln 'fetch_clip\|get_video_metadata\|ask_plan' editor/e2e
+ export/e2e homepage/e2e umtool/e2e` is empty.
+- **The tests bite.** With the mapping line reverted to pass the cited id, 2 of the 9 tool tests
+ fail: the Rumble window and the full-mode Rumble. The timeout test runs its 3 s wait (three
+ 1 s polls) on a fake clock that advances by each `sleep`, so no real timer is involved.
+- **Numbers tools: none.** Nothing here was run against the live :3001 editor. The live proofs are
+ rollout step 1.
+
+**Found and left.**
+- **`mcp/bench/smoke.ts --clip` (optional) was not done.** The rollout's live proof is an in-memory
+ client script, as the plan says. A `--clip` smoke would need the stdio transport to pass
+ `WORKER_TOKEN` through: `StdioClientTransport` gives the child only `getDefaultEnvironment()`
+ unless `env` is set.
+- **Both plan heads still say "The MCP is read-only".** That is true of the process. The new step
+ sits beside it and says the editor fetches.
+- **`mcp/README.md`'s own "Add to Claude Code" and `mcp.json` examples** do not carry the two env
+ lines. The plan named only README.md's two blocks and AGENTS.md's; the tool row names both
+ variables.
+- **The plan's known limitations stand**, plus one the review found (S2, now documented):
+ - The editor fetches only for a channel it already archives: a fresh editor behind an MCP
+ pointed at a public site answers 404 `Channel "<slug>" not found` on every clip.
+ - A video absent from `source` is passed through as cited, with the note.
+ - Full mode sends no `webpageUrl`, so a video the editor has never seen gets the editor's 404
+ (with the added sentence), where a window can still fetch it by URL.
+ - A public-only setup gets the no-editor error.
+- **FACTS and STATE are not edited** (shared files): the plan's rollout step 4 adds "`fetch_clip`
+ is the only MCP tool that causes a write, and the editor does it".
+- **Registration is owed by the operator** (rollout step 2): this machine's `archilyzer` entry has
+ `"env": {}`, so until it is re-registered `fetch_clip` answers "no editor configured".
+- ~~N2: the `queued` text ends "the next ask finds it cached"~~: taken by the plan owner with the
+ re-read's R2, below.
+- **Review nits not taken** (`m-review.md`):
+ - N3: a poll-phase disabled 503 gets the token sentence but not "the endpoint is off". Cosmetic.
+ - N4: the sweep's step says "(or a report needs one)", so extractors could fetch per finding.
+ Plan text, left for the operator.
+ - N5 (pre-existing): `AGENTS.md`'s "What it needs: **yt-dlp**" line for umtool's report-to-video
+ is stale, because umtool asks the editor by default.
+- **A request whose body stalls after its headers** reads as `{}`. `readJson` swallows the
+ aborted body, so a POST reads as `The editor refused (HTTP 200): no reason given` and a poll
+ keeps polling until the wait ends. Rare; the timeout still bounds it.
+
+**Review fixes** (review SHIP AFTER FIXES, `m-review.md`: no must-fix, four should-fix, all fixed;
+nit N1 taken).
+
+1. **S1: the job id survives an editor that stops answering mid-poll** (`868d4ab7`).
+ - **The failure.** A network error while polling answered "Could not reach the editor" with no
+ job id. The natural retry was the same call. With `full: true` that queues a second
+ whole-recording download, because `fetchFullSourceAction` checks the saved-video pointer
+ only when a request arrives.
+ - **The fix.** The poll-phase outcome carries the `jobId`, and the text is `Could not reach the
+ editor at <url> while polling job <jobId>: <message>. Is it running? Call fetch_clip again
+ with job: "<jobId>" — do not repeat the original request, the fetch may still be running.`
+2. **S2: the docs say the editor must already archive the channel** (`d117e533`). The route answers
+ 404 `Channel "<slug>" not found` for a channel with no dir under the editor's `transcripts/`, so
+ pointing the MCP at a public site with a fresh editor gets a 404 on every clip. One sentence
+ each in README "Clips and video" (whose yt-dlp fallback now also covers an editor that does not
+ archive the channel), AGENTS.md's clips loop and the `mcp/README.md` row. The tool description
+ says it too. The `download-sections` grep still gives two lines, both fallback sentences
+ (`AGENTS.md:87`, `README.md:369`).
+3. **S3 + N1: every request is bounded, and a failure says why** (`0777bfd1`).
+ - Each request, the POST and every poll, carries its own `AbortSignal.timeout(REQUEST_TIMEOUT_MS
+ = 15 s)`. Before, a POST that stats a hung mount, or an editor whose event loop had stalled,
+ held the call until undici's 300 s headers timeout. Now `wait_seconds` is the bound, give or
+ take one request (at most POST 15 s + the wait + one poll of 15 s).
+ - A timeout reads `no answer within 15 s (request timed out)`. A timed-out POST adds `It may
+ have queued the fetch anyway: check the editor's /jobs page before asking again.`
+ - Node's bare `fetch failed` carries its cause, e.g. `fetch failed (connect ECONNREFUSED
+ 127.0.0.1:3001)`.
+ - `requestTimeoutMs` is injectable in `FetchClipDeps` for the tests only.
+4. **S4: the 60 s client timeout** (`8cf50cf6`).
+ - (a) The tool description and the `mcp/README.md` row say that a client with a 60 s default
+ request timeout must raise it or pass `wait_seconds` ≤ 50. The fetch continues on the editor
+ either way; resume it with `job` (wording corrected in the re-read, R2 below).
+ - (b) **Progress notifications: done, because it was cheap.** The `tools/call` handler now takes
+ `ctx`. When `ctx.mcpReq._meta.progressToken` is set, `progressNotifier` turns `fetchClip`'s new
+ `onPoll` hook into one `notifications/progress` per poll that finds the job still waiting
+ (`progress` = the poll count, `message` `editor job <id>: <status>, <n>s waited`, no `total`),
+ sent with `ctx.mcpReq.notify`. A client's `resetTimeoutOnProgress` then keeps the call alive.
+ A failing notify never fails the fetch, and with no token none is sent.
+ - Claude Code sets its own, longer tool timeout. **The rollout's in-memory proof script must
+ still pass `{ timeout: 330_000 }`** (or `onprogress` with `resetTimeoutOnProgress`) to
+ `callTool`.
+
+| sha | what |
+|---|---|
+| `868d4ab7` | S1: the poll-phase `unreachable` keeps `jobId`; the resume text; `fetchClip.test.ts` +1 (POST 202, poll throws, then the named resume polls once and posts nothing) |
+| `d117e533` | S2: README, AGENTS.md, mcp/README: the editor must already archive the channel |
+| `0777bfd1` | S3 + N1: `AbortSignal.timeout(15 s)` per request, `describeFetchError` (cause, timeout), the POST-timeout `/jobs` sentence; `fetchClip.test.ts` +6 (a signal per request, the cause, a fake `TimeoutError`, and over a real socket: a POST never answered, a poll never answered → the job id, a refused connection) |
+| `8cf50cf6` | S4: `onPoll` → `notifications/progress`; the 60 s clause in the tool description and the README row; `fetchClip.test.ts` +1, `fetchClip.tool.test.ts` +1 (in-memory progress), `protocol.test.ts` +1 (the real process on the modern era with a stub editor) |
+| _this_ | `plans:` these review fixes; the changelog bullet names the channel requirement, the progress and the 60 s clause |
+
+**Gates on `8cf50cf6`** (`m-gate3.log`):
+- **tsc:**
+ - The whole workspace is clean.
+ - The mcp `tsc --noEmit` is clean on `868d4ab7` and `0777bfd1`, each checked out alone.
+ - `d117e533` is docs only.
+- **mcp 269/269** (259 + 10), 17 s.
+- common and `test:scripts` were not re-run: nothing outside `mcp/` and the three docs changed.
+- **The new tests bite:**
+ - Without the timeout signal, the signal test fails, and the real never-answering POST hangs
+ past an 8 s test timeout.
+ - With the progress wiring disabled, the in-memory progress test fails.
+- **The modern era carries progress.** The protocol test drives the real `src/index.ts` over stdio,
+ negotiated `modern`, with `ARCHILYZER_EDITOR_URL` and `WORKER_TOKEN` passed through the
+ transport's `env`. The stub editor sees three polls, each with `Bearer tok-proto`, and the client
+ gets progress `[1, 2]`.
+
+**Re-read of the fixes** (`m-review.md`: SHIP). The plan owner took R1 and R2 in one commit.
+- **R2: the queued text no longer invites a repeat.** The editor does not dedupe a running
+ whole-recording job (`archiveSourceVideo` → `runManagedFunction` has no running-job check), so
+ "the next ask finds it cached" held only once the job had finished. A repeated `full: true`
+ while it ran queued a second download.
+ - The queued text now reads `Still <status> on the editor (job <jobId>, waited <n>s). Call
+ fetch_clip again with job: "<jobId>" to keep waiting — never repeat the original request while
+ it runs (that would queue a second fetch). Nothing is lost: the fetch continues on the editor,
+ and once it has finished the same request finds it cached.`
+ - The two 60 s clauses (tool description, `mcp/README.md` row) now end: "the fetch continues on
+ the editor either way; resume it with job, and once it has finished the same request finds it
+ cached".
+ - The instructions step already said "call it again with the job it names" and is unchanged.
+ - The plan's queued text in `mcp-fetch-clip.md` is amended to match, with the reason.
+- **R1:** the three real-socket tests carry `{ timeout: 10_000 }`, so a timeout regression fails in
+ 10 s instead of hanging.
+
+| sha | what |
+|---|---|
+| _this_ | R2's queued text and the 60 s clauses; the test that asserts the queued text; R1's test timeouts; the plan's queued text; this paragraph and N2 struck above |
+
+**Gates** (`m-gate4.log`): tsc clean (whole workspace); **mcp 269/269**, 16 s.
+
## Rollout
Nothing is rolled out, except that **Jeralyzer is already on the brand, in Signal** (a build-deploy