# `fetch_clip` — the MCP asks the editor for clip media (no yt-dlp by hand) Written 2026-09-26 in plan mode; self-contained for a fresh session. Anchors verified on `main` `c9223978` by two Explore passes and one Plan pass. Rules: `plans/tools/implementer-rules.md` (one Opus implementer in a worktree, one read-only Opus review, the parent merges). Memory to read first: `no-yt-dlp-by-hand`, `release-9-in-flight`. ## Context The operator's rule since 2026-09-20: every fetch goes through Archilyzer (editor lanes, umtool actions), never yt-dlp by hand — the editor path has the cookie policy, the per-platform sleeps, the 429 cooldown, and provenance. umtool already obeys it (`umtool/report-to-video/fetch-via-editor.mjs` → `POST /api/media/fetch-window`). The MCP server (`mcp/`) does not: it has no media tool, and the docs it points agents at (`README.md` "Clips and video", `AGENTS.md` "Clips and report-to-video") still say "use `yt-dlp --download-sections`". An agent following `/ask` or `/sweep` that needs the clip behind a citation therefore either shells out to yt-dlp (unpaced, no cookies, bytes outside the corpus) or stops. This adds ONE tool, `fetch_clip`, that calls the editor's existing fetch-window job, and rewrites the guidance so the tool is the way and yt-dlp is only the no-editor fallback. ## What exists (verified) - **Editor endpoint** `POST /api/media/fetch-window` (`editor/app/api/media/fetch-window/route.ts`), auth `Authorization: Bearer $WORKER_TOKEN` (`authorizeWorkerRequest`, `common/lib/workerToken.ts:23`; 503 "worker endpoint disabled" when the editor has no token, 401 on a bad one). Body `{channelSlug, videoId, webpageUrl?, from, to, pad?, requestedBy (required), manifest?, clipId?, reason?}`. Window ≤ `MAX_CLIP_WINDOW_SECONDS` = 900 (`common/lib/clipWindow.ts:53`). Responses: 200 `{cached:true, file, from, to, bytes, provenance}` (a NARROWER request inside an existing window returns the wider file and ITS span), 202 `{cached:false, jobId, file, from, to}`, 409 `{error, cooldownMs, platform}`, 400/404/503/507 `{error}`. Poll `GET /api/media/fetch-window/{jobId}` → `{status: queued|running|done|failed|cancelled, jobId, file?, from?, to?, bytes?, error?}` (`error` = last 2000 bytes of the job log). File: `channels//data//clips/-.mp4` (`.toFixed(2)`, 720p H.264/AAC) + `-.json` provenance; never writes `metadata.info.json`; pruned only by `evict-clips` by age (no TTL). Job kind `fetch-window` (`common/jobs/jobKinds.ts:255`) is platform-queued, paced with `--sleep-requests 1`, records the platform backoff on a 429. - **Reference client** `umtool/report-to-video/fetch-via-editor.mjs`: env `ARCHILYZER_EDITOR_URL` (default `http://localhost:3001`) + `WORKER_TOKEN`; `POLL_MS` 1000; pad default 3 s; `from = max(0, start-pad).toFixed(2)`, `to = (end+pad).toFixed(2)`; `reason` ≤ 400 chars. The ops script `scripts/archilyzer-ops.mjs:311-318` uses the same env names. No shared HTTP client in `common/` — the tool gets its own small one. - **MCP shape** `mcp/src/server.ts`: `TOOLS: Tool[]` literal JSON-schema array (~:299, `SOURCE_ARG` spread ~:173); `createServer(...)` :924; the `tools/call` handler :949 resolves `source` FIRST (`registry.resolve(args.source)`), dispatches through one `switch`, and passes every result through `withCorpus()` (~:907, appends "(corpus: …)"). Handlers return `text()`/`errorText()` (`{content:[{type:"text",text}], isError?}`). No tool does authenticated HTTP or polling today. `findVideo(source, videoId, channelHint?)` (`mcp/src/search.ts:924`) returns the record incl. `webpageUrl`. Tests: `tsx --test src/*.test.ts`, in-memory client harness (`sourceRegistry.test.ts:112-127` `connect()` + `firstText()`); `protocol.test.ts:33-48` pins the EXACT tool list `EXPECTED_TOOLS`. `instructions.test.ts:24-65` requires every backticked snake_case token in plan text to be a real tool (or listed in `NOT_TOOLS` :47). `mcp/src/instructions.ts` builds the `ask_plan`/`sweep_plan` text and never mentions yt-dlp today. `.claude/commands/{ask,sweep}.md` need no change. - **Rumble ids (the one real trap).** A published record's `id` is yt-dlp's native id = the EMBED id (`common/lib/transcripts-server.ts:100-103`; live proof: `the-quartering-rumble/data/v1007ay/metadata.info.json` has `"id": "vxe1ae"`, the export manifest keys by `vxe1ae`). Citations carry the embed id. The editor's dir is the canonical id from the URL SLUG (`videoDirOf`, `videoActions.ts:887`; `extractVideoId`, `common/lib/videoId.ts:11`, Rumble branch :21-29). Passing the embed id would create a NEW `data/vxe1ae/` and fetch `rumble.com/vxe1ae` — wrong dir. Nothing in the editor maps embed → slug. But the record's `webpageUrl` (`transcripts-server.ts:118`, `meta.webpage_url`) IS the slug URL, and the MCP already reads it. YouTube: canonical === id. ## Decisions (made; do not re-open) Tool name `fetch_clip`. Env `ARCHILYZER_EDITOR_URL` (default `http://localhost:3001`) + `WORKER_TOKEN` (the client family's names). ONE tool that waits up to `wait_seconds` (default 90, max 300) and, if the job is still running, returns `queued` + `jobId` and says to call again with `job` — no second tool. `requestedBy: "mcp"`; `reason` required; `report` optional → `manifest`. Window-only (`full` not exposed). `source` stays optional for symmetry (the corpus trailer is fine) and is what the Rumble mapping reads. No `webpage_url` override: a video not found in `source` is passed through as-is with a note (documented limitation). The MCP process writes nothing. No editor configured → `isError` naming both variables; the README keeps the raw yt-dlp command as the NO-EDITOR fallback only. ## Implementation (one Opus slice, branch `mcp/fetch-clip`, worktree off `main`) ### 1. `mcp/src/fetchClip.ts` (new, pure; no MCP imports) - `FetchClipDeps = { env, fetch, sleep, now }` — all injected so tests need no timers or network. - `parseSeconds(v)`: number, `ss`, `mm:ss`, `h:mm:ss` → seconds or null. - `planWindow({video, start, end, pad})` → `{from, to}` or `{error}`: id regex `/^[\w.-]+$/` and not `.`/`..` (mirror `route.ts:34-36`); `from = Number(Math.max(0, start-pad).toFixed(2))`, `to = Number((end+pad).toFixed(2))`; `from < to`; `to-from ≤ 900` AFTER padding. - `fetchClip(args, deps)` → typed `FetchClipOutcome` (`cached | fetched | queued | cooldown | refused | failed | unreachable | no_editor`): POST (unless `job` is set → poll only), then poll every 1 s via `deps.sleep` until terminal or `deps.now()` passes the deadline. Body: `{channelSlug, videoId, webpageUrl?, from, to, pad, requestedBy: "mcp", manifest: report, reason: reason.slice(0,400)}`. - `renderFetchClip(outcome, ctx)` → `{text, isError}` with the exact texts below. **Client-side errors (before any HTTP), `isError`:** `fetch_clip: channel is required (the channel slug)` · `fetch_clip: video is required (the archive's video id)` · `fetch_clip: video "" must match /^[\w.-]+$/` · `fetch_clip: start "" is not a time (use seconds, mm:ss or h:mm:ss)` (same for end) · `fetch_clip: start () must be less than end ()` · `fetch_clip: the window – is s; the editor fetches at most 900s per window — cite a narrower span` · `fetch_clip: reason is required — one line saying why these seconds are needed (it is stored beside the file)` · `pad` finite ≥ 0 · `wait_seconds` clamped to [0, 300] · no `WORKER_TOKEN`: `fetch_clip: no editor configured. Set ARCHILYZER_EDITOR_URL (e.g. http://localhost:3001) and WORKER_TOKEN (the editor's own WORKER_TOKEN) when registering the MCP server. A public-only setup — no local Archilyzer editor — cannot fetch media through Archilyzer; see README "Clips and video" for the no-editor fallback.` **Result texts.** Footer for any file result, one line each: `file: ` · `window: – (s)` · `bytes: ` · `This path is a read-only corpus artifact: play or copy it, never move, edit or delete it. The editor prunes clips by age (evict-clips); provenance sits beside it as -.json.` - 200 cached: `Already on disk — a WIDER cached window that contains –: –.` (or "the exact window" when equal within 0.02 s) + footer (+ `requested by ` when present). - 202 → done: `Fetched – of / (job , s waited).` + footer from the poll's `file/from/to/bytes`. - wait expired, still queued/running (NOT `isError`): `Still on the editor (job , waited s). Call fetch_clip again with job: "" to keep waiting — never repeat the original request while it runs (that would queue a second fetch). Nothing is lost: the fetch continues on the editor, and once it has finished the same request finds it cached.` (Amended 2026-09-26 after the review: the editor does not dedupe a running whole-recording job — `archiveSourceVideo` → `runManagedFunction` has no running-job check — so a repeated `full: true` request while it runs queues a second download, and "the next ask finds it cached" was true only once the job had finished.) - 409: `The editor is in a rate-limit cooldown — s remaining. Wait, then call fetch_clip again. ()` - 401 / 503: `The editor refused the token (HTTP ): . WORKER_TOKEN must equal the value the editor runs with (editor/.env).` — 503 whose error says "disabled" adds `The editor has no WORKER_TOKEN set, so its fetch endpoint is off.` - other 4xx/5xx: `The editor refused (HTTP ): `. - poll failed/cancelled: `Editor job .\n\nlog tail:\n`. - poll HTTP error: `Polling editor job failed (HTTP ): `; on 404 add `the job is unknown to this editor (restarted? wrong ARCHILYZER_EDITOR_URL?)`. - network: `Could not reach the editor at : . Is it running?` - not found in `source` (prefix, not an error): `note: "