Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 1e7f5c9847a7c190f28b7f2897d8fa0a01a9fbbc
parent 65172404180bc9553b7b544652cc9800b282ae01
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sat, 26 Sep 2026 14:14:09 -0400

plans: fetch_clip — the MCP asks the editor for clip media (no yt-dlp by hand); amended: full: true exposed (saved-video store, editor's existing branch)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
Aplans/mcp-fetch-clip.md | 302++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
1 file changed, 302 insertions(+), 0 deletions(-)

diff --git a/plans/mcp-fetch-clip.md b/plans/mcp-fetch-clip.md @@ -0,0 +1,302 @@ +# `fetch_clip` — the MCP asks the editor for clip media (no yt-dlp by hand) + +Written 2026-09-26 in plan mode; self-contained for a fresh session. Anchors verified on `main` +`c9223978` by two Explore passes and one Plan pass. Rules: `plans/tools/implementer-rules.md` (one +Opus implementer in a worktree, one read-only Opus review, the parent merges). Memory to read first: +`no-yt-dlp-by-hand`, `release-9-in-flight`. + +## Context + +The operator's rule since 2026-09-20: every fetch goes through Archilyzer (editor lanes, umtool +actions), never yt-dlp by hand — the editor path has the cookie policy, the per-platform sleeps, +the 429 cooldown, and provenance. umtool already obeys it (`umtool/report-to-video/fetch-via-editor.mjs` +→ `POST /api/media/fetch-window`). The MCP server (`mcp/`) does not: it has no media tool, and the +docs it points agents at (`README.md` "Clips and video", `AGENTS.md` "Clips and report-to-video") +still say "use `yt-dlp --download-sections`". An agent following `/ask` or `/sweep` that needs the +clip behind a citation therefore either shells out to yt-dlp (unpaced, no cookies, bytes outside the +corpus) or stops. This adds ONE tool, `fetch_clip`, that calls the editor's existing fetch-window +job, and rewrites the guidance so the tool is the way and yt-dlp is only the no-editor fallback. + +## What exists (verified) + +- **Editor endpoint** `POST /api/media/fetch-window` (`editor/app/api/media/fetch-window/route.ts`), + auth `Authorization: Bearer $WORKER_TOKEN` (`authorizeWorkerRequest`, `common/lib/workerToken.ts:23`; + 503 "worker endpoint disabled" when the editor has no token, 401 on a bad one). Body + `{channelSlug, videoId, webpageUrl?, from, to, pad?, requestedBy (required), manifest?, clipId?, reason?}`. + Window ≤ `MAX_CLIP_WINDOW_SECONDS` = 900 (`common/lib/clipWindow.ts:53`). Responses: 200 + `{cached:true, file, from, to, bytes, provenance}` (a NARROWER request inside an existing window + returns the wider file and ITS span), 202 `{cached:false, jobId, file, from, to}`, 409 + `{error, cooldownMs, platform}`, 400/404/503/507 `{error}`. Poll `GET /api/media/fetch-window/{jobId}` + → `{status: queued|running|done|failed|cancelled, jobId, file?, from?, to?, bytes?, error?}` (`error` + = last 2000 bytes of the job log). File: `channels/<slug>/data/<id>/clips/<from>-<to>.mp4` + (`.toFixed(2)`, 720p H.264/AAC) + `<from>-<to>.json` provenance; never writes `metadata.info.json`; + pruned only by `evict-clips` by age (no TTL). Job kind `fetch-window` (`common/jobs/jobKinds.ts:255`) + is platform-queued, paced with `--sleep-requests 1`, records the platform backoff on a 429. +- **Reference client** `umtool/report-to-video/fetch-via-editor.mjs`: env `ARCHILYZER_EDITOR_URL` + (default `http://localhost:3001`) + `WORKER_TOKEN`; `POLL_MS` 1000; pad default 3 s; + `from = max(0, start-pad).toFixed(2)`, `to = (end+pad).toFixed(2)`; `reason` ≤ 400 chars. The ops + script `scripts/archilyzer-ops.mjs:311-318` uses the same env names. No shared HTTP client in + `common/` — the tool gets its own small one. +- **MCP shape** `mcp/src/server.ts`: `TOOLS: Tool[]` literal JSON-schema array (~:299, `SOURCE_ARG` + spread ~:173); `createServer(...)` :924; the `tools/call` handler :949 resolves `source` FIRST + (`registry.resolve(args.source)`), dispatches through one `switch`, and passes every result + through `withCorpus()` (~:907, appends "(corpus: …)"). Handlers return `text()`/`errorText()` + (`{content:[{type:"text",text}], isError?}`). No tool does authenticated HTTP or polling today. + `findVideo(source, videoId, channelHint?)` (`mcp/src/search.ts:924`) returns the record incl. + `webpageUrl`. Tests: `tsx --test src/*.test.ts`, in-memory client harness (`sourceRegistry.test.ts:112-127` + `connect()` + `firstText()`); `protocol.test.ts:33-48` pins the EXACT tool list `EXPECTED_TOOLS`. + `instructions.test.ts:24-65` requires every backticked snake_case token in plan text to be a real + tool (or listed in `NOT_TOOLS` :47). `mcp/src/instructions.ts` builds the `ask_plan`/`sweep_plan` + text and never mentions yt-dlp today. `.claude/commands/{ask,sweep}.md` need no change. +- **Rumble ids (the one real trap).** A published record's `id` is yt-dlp's native id = the EMBED id + (`common/lib/transcripts-server.ts:100-103`; live proof: `the-quartering-rumble/data/v1007ay/metadata.info.json` + has `"id": "vxe1ae"`, the export manifest keys by `vxe1ae`). Citations carry the embed id. The + editor's dir is the canonical id from the URL SLUG (`videoDirOf`, `videoActions.ts:887`; + `extractVideoId`, `common/lib/videoId.ts:11`, Rumble branch :21-29). Passing the embed id would + create a NEW `data/vxe1ae/` and fetch `rumble.com/vxe1ae` — wrong dir. Nothing in the editor maps + embed → slug. But the record's `webpageUrl` (`transcripts-server.ts:118`, `meta.webpage_url`) IS + the slug URL, and the MCP already reads it. YouTube: canonical === id. + +## Decisions (made; do not re-open) + +Tool name `fetch_clip`. Env `ARCHILYZER_EDITOR_URL` (default `http://localhost:3001`) + `WORKER_TOKEN` +(the client family's names). ONE tool that waits up to `wait_seconds` (default 90, max 300) and, if +the job is still running, returns `queued` + `jobId` and says to call again with `job` — no second +tool. `requestedBy: "mcp"`; `reason` required; `report` optional → `manifest`. Window-only (`full` +not exposed). `source` stays optional for symmetry (the corpus trailer is fine) and is what the +Rumble mapping reads. No `webpage_url` override: a video not found in `source` is passed through +as-is with a note (documented limitation). The MCP process writes nothing. No editor configured → +`isError` naming both variables; the README keeps the raw yt-dlp command as the NO-EDITOR fallback +only. + +## Implementation (one Opus slice, branch `mcp/fetch-clip`, worktree off `main`) + +### 1. `mcp/src/fetchClip.ts` (new, pure; no MCP imports) + +- `FetchClipDeps = { env, fetch, sleep, now }` — all injected so tests need no timers or network. +- `parseSeconds(v)`: number, `ss`, `mm:ss`, `h:mm:ss` → seconds or null. +- `planWindow({video, start, end, pad})` → `{from, to}` or `{error}`: id regex `/^[\w.-]+$/` and not + `.`/`..` (mirror `route.ts:34-36`); `from = Number(Math.max(0, start-pad).toFixed(2))`, + `to = Number((end+pad).toFixed(2))`; `from < to`; `to-from ≤ 900` AFTER padding. +- `fetchClip(args, deps)` → typed `FetchClipOutcome` (`cached | fetched | queued | cooldown | refused + | failed | unreachable | no_editor`): POST (unless `job` is set → poll only), then poll every 1 s via + `deps.sleep` until terminal or `deps.now()` passes the deadline. Body: `{channelSlug, videoId, + webpageUrl?, from, to, pad, requestedBy: "mcp", manifest: report, reason: reason.slice(0,400)}`. +- `renderFetchClip(outcome, ctx)` → `{text, isError}` with the exact texts below. + +**Client-side errors (before any HTTP), `isError`:** +`fetch_clip: channel is required (the channel slug)` · `fetch_clip: video is required (the archive's +video id)` · `fetch_clip: video "<v>" must match /^[\w.-]+$/` · `fetch_clip: start "<x>" is not a +time (use seconds, mm:ss or h:mm:ss)` (same for end) · `fetch_clip: start (<s>) must be less than +end (<e>)` · `fetch_clip: the window <from>–<to> is <n>s; the editor fetches at most 900s per window +— cite a narrower span` · `fetch_clip: reason is required — one line saying why these seconds are +needed (it is stored beside the file)` · `pad` finite ≥ 0 · `wait_seconds` clamped to [0, 300] · +no `WORKER_TOKEN`: `fetch_clip: no editor configured. Set ARCHILYZER_EDITOR_URL (e.g. +http://localhost:3001) and WORKER_TOKEN (the editor's own WORKER_TOKEN) when registering the MCP +server. A public-only setup — no local Archilyzer editor — cannot fetch media through Archilyzer; +see README "Clips and video" for the no-editor fallback.` + +**Result texts.** Footer for any file result, one line each: `file: <abs path>` · `window: +<from>–<to> (<n>s)` · `bytes: <n>` · `This path is a read-only corpus artifact: play or copy it, +never move, edit or delete it. The editor prunes clips by age (evict-clips); provenance sits beside +it as <from>-<to>.json.` +- 200 cached: `Already on disk — a WIDER cached window that contains <req from>–<req to>: + <body.from>–<body.to>.` (or "the exact window" when equal within 0.02 s) + footer (+ `requested + by <provenance.requestedBy>` when present). +- 202 → done: `Fetched <from>–<to> of <channel>/<canonical> (job <id>, <n>s waited).` + footer from + the poll's `file/from/to/bytes`. +- wait expired, still queued/running (NOT `isError`): `Still <status> on the editor (job <jobId>, + waited <n>s). Call fetch_clip again with job: "<jobId>" to keep waiting — nothing is lost, the fetch + continues on the editor and the next ask finds it cached.` +- 409: `The editor is in a <platform> rate-limit cooldown — <ceil(cooldownMs/1000)>s remaining. Wait, + then call fetch_clip again. (<error>)` +- 401 / 503: `The editor refused the token (HTTP <s>): <error>. WORKER_TOKEN must equal the value the + editor runs with (editor/.env).` — 503 whose error says "disabled" adds `The editor has no + WORKER_TOKEN set, so its fetch endpoint is off.` +- other 4xx/5xx: `The editor refused (HTTP <s>): <error verbatim>`. +- poll failed/cancelled: `Editor job <id> <status>.\n\nlog tail:\n<error>`. +- poll HTTP error: `Polling editor job <id> failed (HTTP <s>): <error>`; on 404 add `the job is + unknown to this editor (restarted? wrong ARCHILYZER_EDITOR_URL?)`. +- network: `Could not reach the editor at <url>: <message>. Is it running?` +- not found in `source` (prefix, not an error): `note: "<video>" was not found in corpus <handle>; + the id was passed to the editor as-is (for a Rumble citation this may be the embed id — pass the + corpus the citation came from as source).` + +### 2. `mcp/src/server.ts` + +- `TOOLS` entry after `get_video_metadata` (~:704): `channel` (string, the channel slug), `video` + (string, the archive's video id as cited), `start`, `end` (`oneOf [number, string]`; seconds or + mm:ss / h:mm:ss), `pad` (number, default 3), `reason` (string, one line, ≤ 400), `report` + (string, optional), `wait_seconds` (number, default 90, max 300), `job` (string; resume — when set + the other args are ignored), `...SOURCE_ARG`; `additionalProperties: false`; `required: []` with + validation in code (so `job` alone is valid). Description says: the editor fetches through its + paced, cookie-aware, provenanced job; the file is a read-only corpus artifact; ≤ 15 min per + window; never run yt-dlp yourself. +- `createServer(sourceOrRegistry, opts?: { fetchClipDeps?: Partial<FetchClipDeps> })`, default deps + `{env: process.env, fetch: globalThis.fetch.bind(globalThis), sleep: setTimeout-promise, now: Date.now}`. +- `case "fetch_clip": return handleFetchClip(source, resolved, args, deps)`: validate → (unless + `job`) `findVideo(source, video, channel)` → `canonical = extractVideoId(record.webpageUrl) ?? video` + (import from `yt-dlp-transcript-common/lib/videoId`, a leaf module) and `webpageUrl = + record.webpageUrl` → `fetchClip` → `renderFetchClip` → `text()`/`errorText()`. `withCorpus` adds + the trailer as for every tool. +- `protocol.test.ts` `EXPECTED_TOOLS`: insert `"fetch_clip"` after `"get_video_metadata"` (the + test pins the exact ordered list and fails otherwise). + +### 3. `mcp/src/instructions.ts` + +One step in BOTH builders, voice-matched, after the citation step (`buildAskInstructions` after +~:331; `buildSweepInstructions` before **Finish**, after ~:237): +`**Media for a cited moment — through the editor only.** When I ask for the clip behind a citation +(or a report needs one), call \`fetch_clip\` with that citation's channel slug, video id, start/end +in seconds (or mm:ss) and a one-line reason, and pass source: "<ctx.corpus>" so a Rumble id +resolves to the right directory. NEVER run yt-dlp yourself, in any form. If the tool reports no +editor is configured, say so and stop — the README's yt-dlp command is the operator's fallback, +not yours. If it returns queued, call it again with the job it names.` +Do not backtick `wait_seconds` in the prose (the token test) or add it to `NOT_TOOLS`. + +### 4. Docs (short edits) + +- `README.md` ~:61-75 and ~:319-333 (both `claude mcp add` blocks): add + `--env ARCHILYZER_EDITOR_URL=http://localhost:3001 \` and `--env WORKER_TOKEN=… \` with one line + "optional: lets `fetch_clip` ask a local editor for clip media". ~:338-356 "Clips and video": + step 3 becomes "ask for just those clips with the `fetch_clip` MCP tool — the editor fetches the + window through its paced, cookie-aware, provenanced job; the file lands in the corpus beside the + video (`channels/<slug>/data/<id>/clips/`)"; keep `yt-dlp --download-sections` as "with no editor + (a public-only setup) the fallback is running yt-dlp yourself"; drop the `out/clips-raw` claim + for the editor path. +- `mcp/README.md` ~:7-9: "The one exception is `fetch_clip`, which asks a local Archilyzer editor to + fetch a clip window; the MCP itself still writes nothing." Tool table ~:14-25: one row (env vars, + ≤ 15 min, waits then returns a job to resume, Rumble mapping via the record's `webpageUrl`, a + video not in `source` passed through as-is = known limitation). +- `AGENTS.md` ~:53-58: the two `--env` lines with the "optional" note; ~:75-79: clip media for a + cited moment goes through the `fetch_clip` MCP tool (→ editor `POST /api/media/fetch-window`), + never `yt-dlp --download-sections` by hand; that command is only the fallback for a machine with + no editor. + +### 5. Tests + +- `mcp/src/fetchClip.test.ts` (scripted fake `fetch`, recording `sleep`, `now` advancing per sleep): + 200 cached wider window; 202 → running → done (file/bytes); 202 → failed with log tail; 409 + (seconds + platform); 401 and 503 texts; timeout → `queued` with jobId without real timers; resume + by `job` issues no POST (assert the fetch call list); `planWindow` every error text, padding + rounding, the 900 cap after pad; `parseSeconds("1:02:03") === 3723`; POST body has + `requestedBy: "mcp"`, `manifest`, truncated `reason`; no-env outcome. +- Tool-level (new `fetchClip.tool.test.ts`, via `connect()`): `createServer(registry, {fetchClipDeps})` + with a fake source record `{id: "vxe1ae", webpageUrl: "https://rumble.com/v1007ay-x.html?e9s=1"}` → + the POST body has `videoId: "v1007ay"` and the `webpageUrl`; the text ends with `(corpus: …)`; + no-env is `isError` naming both variables; a YouTube id maps to itself. +- `protocol.test.ts` +1 name; `instructions.test.ts` +1: both builders contain `fetch_clip` and + `NEVER run yt-dlp`, and the sweep's step comes before "Finish". +- Optional: `mcp/bench/smoke.ts` gains an opt-in `--clip <channel>/<video>@<start>-<end>` (the default + smoke stays write-free). + +### 6. Gates and rollout + +Gates (worktree root): tsc clean; `pnpm --filter yt-dlp-transcript-mcp test` (219 baseline → record +the new count); `pnpm run test:scripts` (162 + 1 skip); `pnpm --filter yt-dlp-transcript-common test` +unchanged (1,845). No e2e: no editor/export/hub spec lists MCP tool names. Record: `### Slice M, +as shipped — fetch_clip` in `plans/release-10.md` (create with the release-9 shape) + CHANGELOG +bullet; report to `$T/m-report.md`. + +Rollout (parent, after merge; no editor restart — Claude Code relaunches the MCP per session): +1. Live proof from the primary with the merged code, against the real editor on :3001, through an + in-memory client script (not by hand-running yt-dlp): `fetch_clip` on a tiny YouTube window from + a channel already archived (e.g. one teamrcn video, `start` 10 → `end` 20, reason "release 10 + live proof") → expect 202 → done, a file under `data/<id>/clips/10.00-23.00.mp4` (pad 3), a + provenance sidecar with `requestedBy: "mcp"`, no `metadata.info.json` written; then the same + call again → cached. Then one Rumble citation from `the-quartering-rumble` (embed id from the + published archive as `video`, `source: remote:https://jeralyzer.pages.dev`) → the POST carries the + slug id and the clip lands in the slug-named dir. +2. Re-register the MCP for this machine (the operator, or via `! …`): + `claude mcp remove archilyzer && claude mcp add archilyzer --env ARCHILYZER_EDITOR_URL=http://localhost:3001 --env WORKER_TOKEN="$(grep '^WORKER_TOKEN=' editor/.env | cut -d= -f2-)" -- pnpm -C /home/user/Projects/yt-dlp-transcript-browser --filter yt-dlp-transcript-mcp exec tsx src/index.ts --local /home/user/Projects/yt-dlp-transcript-browser/export/public` + (the current entry has `"env": {}`; `umtool/.env.local` already carries both values). +3. `/ask` proof in a fresh session: ask for the clip behind one citation; the plan step names + `fetch_clip`; the tool answers with a corpus path. +4. STATE, FACTS ("`fetch_clip` is the only MCP tool that causes a write, and the editor does it"), + memory (`no-yt-dlp-by-hand` gains the tool), runbook. + +## Verification + +- Unit: the test list above, all green; `EXPECTED_TOOLS` updated; the instructions token test green. +- Live: the two proofs in rollout step 1 (YouTube: fetched then cached; Rumble: slug dir), checked + on disk under `transcripts/channels/…/clips/` and in the editor's `/jobs` (kind `fetch-window`, + `done`). +- Docs: `grep -n 'download-sections' README.md AGENTS.md mcp/README.md` shows it only in the + no-editor fallback sentences. + +## Known limitations (documented, not fixed here) + +- A video absent from the `source` corpus is passed through by its cited id; for Rumble that can be + the embed id and the editor would fetch into the wrong dir — the tool says so in its note; pass + the corpus the citation came from as `source`. +- `full` (whole source) stays umtool/editor-only. +- The MCP cannot start an editor; a public-only setup gets the no-editor error and the README's + yt-dlp fallback. + +## Amendment (operator, 2026-09-26, before implementation) — `full: true` IS exposed + +The operator re-opened one decision: "the ability to download a full video may be useful, e.g. if +asking to summarize one long video." So the "Window-only (`full` not exposed)" decision above is +REPLACED by this section; everything else stands. The editor already does the work — verified: + +- `route.ts:143` reads `body.full === true` BEFORE validating `from`/`to` and calls + `fetchFullSourceAction` (`editor/app/channels/[slug]/videos/[id]/videoActions.ts`), which uses + the **saved-video store** (`transcripts/saved-videos/`, pointer under `data/<id>/`), not `clips/`. + It does NOT take `webpageUrl`: the URL comes from `findVideoSourceUrl` (the video's + `metadata.info.json` or the channel `playlist`), so a video the editor has never seen returns 404 + `Could not determine the video URL: …`. 200 cached → `{cached:true, file, from:0, to:0, bytes, + provenance}` (provenance = the pointer's `origin`, a `SavedVideoOrigin` with `requestedBy`, + `manifest`, `clipId`, `reason`, `requestedAt`). 202 → `{cached:false, jobId, file: null, from:0, + to:0}` — the path is unknowable until yt-dlp picks the container. The job kind is + `redownload-archive`; the poll route (`[jobId]/route.ts:50-52`) accepts that kind and on `done` + reads `file`/`bytes` from the saved-video pointer (no `from`/`to` keys). +- umtool's client sends `{full: true}` INSTEAD of `from`/`to`/`pad` (`fetch-via-editor.mjs:153-167`, + "full REPLACES the span"). Do the same. + +**Tool surface.** Add `full` (boolean, default false) to the `fetch_clip` schema: "the whole +recording instead of a window — for a video that must be re-cut freely or watched end to end; it +lands in the editor's saved-video store, is much larger than a window, and needs a video the editor +already knows (its metadata or playlist entry)." When `full` is true, `start`/`end`/`pad` are +ignored (not required, not validated); `reason` is still required; `channel`/`video` as before; the +Rumble id mapping still applies (the dir is the slug id). `planWindow` is not called. POST body: +`{channelSlug, videoId, full: true, requestedBy: "mcp", manifest: report, reason}` — no +`webpageUrl`, no `from`/`to`/`pad`. + +**Texts (full mode).** +- 200 cached: `Already on disk — the whole recording of <channel>/<canonical>.` + footer. +- 202 → done: `Fetched the whole recording of <channel>/<canonical> (job <id>, <n>s waited).` + + footer. If the poll's `done` carries no `file` (pointer missing), say + `Editor job <id> finished but named no file; check the video's page in the editor.` (`isError`). +- Footer for full: `file: <abs path>` · `bytes: <n>` · `This path is a read-only corpus artifact: + play or copy it, never move, edit or delete it. It lives in the editor's saved-video store, not + in clips/; the editor's keep-videos rule decides how long it stays.` (no `window:` line). +- 404 in full mode comes back through the generic `The editor refused (HTTP 404): <error verbatim>` + path; ADD one sentence after it when `full` was set: `A whole-recording fetch needs a video the + editor already knows (a metadata.info.json or a playlist entry); fetch a window instead, or add + the video to the channel first.` +- `queued` text is unchanged (resume by `job` works for both kinds: the poll route answers for + `redownload-archive` too). +- A `wait_seconds` default of 90 will usually expire on a full download; that is fine — the text + already says the fetch continues and to call again with `job`. + +**Instructions step** (both builders): extend the sentence — "… and a one-line reason (or +full: true when the ask genuinely needs the whole recording — a window is the default), and pass +source: …". Do not backtick `full`. + +**Docs.** README "Clips and video" + `mcp/README.md` row: one clause each — "`full: true` fetches +the whole recording into the saved-video store (needs a video the editor already knows)". + +**Tests** (add to §5): `fetchClip` full → POST body is exactly `{channelSlug, videoId, full: true, +requestedBy, manifest, reason}` with no `from`/`to`/`pad`/`webpageUrl`; 200 cached full text; 202 → +done with `file` + `bytes` and no `window:` line; done without `file` → the isError text; 404 in +full mode carries the extra sentence; `full` with `start`/`end` also given → they are ignored +(assert the body). Tool-level: `full: true` + a Rumble record still maps `videoId` to the slug id. + +**Rollout step 1 gains 1c**: after the window proofs, one `full: true` on a SHORT teamrcn video +(under ~5 min; one download, at the editor's own pace — never a second one in parallel) → expect +202 → `queued` at 90 s or `done`, then a resume by `job` → `file` under `transcripts/saved-videos/`, +the pointer's `origin.requestedBy === "mcp"`; the same call again → cached. + +**Known limitations** gains: full mode cannot take a `webpageUrl`, so a video absent from the +editor's channel (not in its playlist, no metadata) gets the editor's 404; the window path can +still fetch it by URL.