commit 05056a8f2eab10acdc8a9792b43cee5af565d627
parent a2de17b53c2a809018e00d2366dad3b71f364727
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Sat, 26 Sep 2026 15:02:38 -0400
plans: slice M merged (a2de17b5) and proved live — the three fetch_clip proofs against :3001 (YouTube window then cached, Rumble embed id → slug dir, whole recording then cached), the archilyzer re-registration, two editor-side full-mode lows (silent no-op on a captions-only video with a transcript; metadata.info.json rewritten); STATE's merged block, a FACTS section
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Diffstat:
3 files changed, 80 insertions(+), 0 deletions(-)
diff --git a/plans/FACTS.md b/plans/FACTS.md
@@ -6737,3 +6737,31 @@ S4, as shipped"; `release-10.md` "Slice L2 / L1, as shipped". Every anchor below
mount the guard's own `stat` (`common/lib/channelMedia.ts`) hangs on the same syscall, and since
re-queues run one at a time, every re-queue after it stays `queued` until the mount answers.
Cancels all run first. Left, not fixed (`bootQueuedJobs.ts:113-125`).
+
+## `fetch_clip` — the one MCP tool that causes a write, and the editor does it (verified 2026-09-26)
+
+- `mcp/src/fetchClip.ts` is HTTP only (`env`, `fetch`, `sleep`, `now` injected); the MCP process
+ writes nothing. It POSTs `/api/media/fetch-window` with `Authorization: Bearer $WORKER_TOKEN`,
+ `requestedBy: "mcp"`, `reason` (≤ 400, required), `manifest` = the optional `report`; then polls
+ `GET /api/media/fetch-window/<jobId>` every 1 s until `wait_seconds` (default 90, max 300) and
+ returns `queued` + the job id (not an error) to resume with `job`. Every request has a 15 s
+ `AbortSignal.timeout`; a poll network error keeps the job id in its text. When the client sends a
+ `progressToken`, one `notifications/progress` per poll keeps a reset-on-progress client alive
+ (the client SDK's default request timeout is 60 s, `DEFAULT_REQUEST_TIMEOUT_MSEC`).
+- **Rumble:** the cited id is the EMBED id (the record's `id`); the editor's dir is the URL slug.
+ `server.ts` maps through `findVideo(source, video, channel)` → `extractVideoId(record.webpageUrl)`
+ and sends that as `videoId` (+ `webpageUrl` in window mode). Not found in `source` → passed as-is
+ with a `note:` prefix. Live: `vxe1ae` → `data/v1007ay/clips/…`, and 80 of 8,051 sampled
+ `the-quartering-rumble` dirs follow the slug rule.
+- **`full: true`** sends `{channelSlug, videoId, full: true, …}` with no from/to/pad/webpageUrl;
+ the editor's `fetchFullSourceAction` uses the saved-video store (`transcripts/saved-videos/<slug>/
+ <id>/source-media.<ext>`, pointer `data/<id>/saved-video.json` with `keepReason: "override"` and
+ `origin`), job kind `redownload-archive`, and takes NO URL (404 when the video has neither
+ metadata nor a playlist entry). It is a silent no-op on a youtube-handling video that already has
+ a transcript (`downloadOneManaged.ts:1233-1235`: the keep-source override reaches only transcribe
+ handling or the no-captions fallback) and it rewrites `metadata.info.json`; the window path never
+ writes metadata. Neither path dedupes a running job: a repeated request while one runs queues a
+ second fetch, which is why every text says "resume with job".
+- The editor must already HAVE the channel (404 `Channel "<slug>" not found` otherwise); a
+ public-only setup gets the "no editor configured" error naming both env vars, and the README's
+ `yt-dlp --download-sections` is the no-editor fallback only.
diff --git a/plans/STATE.md b/plans/STATE.md
@@ -12,6 +12,16 @@ and L1 records, then "Integration (S4 + L1 + L2), as merged"); the brand's plan
records, S4's included: [`brand-and-themes.md`](brand-and-themes.md).
**Merged to main (2026-09-26), NOT rolled out:**
+- **Slice M, `fetch_clip`, merged after the integration pass and ALREADY IN EFFECT:** `mcp/fetch-clip`
+ @ `fa44984a` → `main` `64f26a71`. The MCP gains ONE tool that asks the live editor's
+ `POST /api/media/fetch-window` for a cited moment's media (a window ≤ 15 min, or `full: true` for
+ the whole recording into the saved-video store); the ask/sweep plans send media through it and
+ never a hand-run yt-dlp. Live-proved 2026-09-26 evening against :3001 (YouTube window fetched then
+ cached; Rumble embed id → slug dir; full recording fetched then cached); `archilyzer` re-registered
+ with `ARCHILYZER_EDITOR_URL` + `WORKER_TOKEN`. Nothing to restart. Owed: the `/ask` proof in a
+ fresh session. Two editor-side lows found by the proof (record, "Slice M — fetch_clip, merged and
+ proved live"): full mode is a silent no-op on a captions-only channel whose video already has a
+ transcript; full mode rewrites `metadata.info.json`. Plan: [`mcp-fetch-clip.md`](mcp-fetch-clip.md).
- **The merges**, in the planned order: S4 `dcb04f61` (`brand/media` @ `333d2826`) → L2 `41ddc382`
(`r10/runner-lows` @ `14e96739`) → L1 `5c0a6ef9` (`r10/hub-lows` @ `b6b47ec1`). The code merged
clean; the conflicts were `plans/release-10.md` (the record sections, all kept) and one
diff --git a/plans/release-10.md b/plans/release-10.md
@@ -1240,3 +1240,45 @@ archive off, is disabled with one line saying why; an archive whose live chat fa
pauses the platform like a 429 (the batch stops, the runner defers the video 6 h, a scan stops,
the availability check stops) instead of excluding the video as deleted; cleanup keeps the audio of
every video its confirmation check did not reach; `/jobs` says why a job was cancelled at boot.
+
+### Slice M — fetch_clip, merged and proved live (2026-09-26 evening)
+
+Merged AFTER the integration pass: `mcp/fetch-clip` @ `fa44984a` → `main` `64f26a71` (the one
+conflict was this file's `## Rollout` insertion point; both sections kept). The runbook's `FINAL`
+predates it — the primary's `main` is the tip to check, not the integration tip. Nothing to
+restart: the MCP is relaunched per Claude Code session, and the live :3001 editor already serves
+`POST /api/media/fetch-window` (every proof below ran against it, from the primary with the merged
+code, through an in-memory MCP client — no yt-dlp by hand).
+
+- **Merged-tree gates:** mcp 269/269 (219 + 50), mcp tsc clean; nothing outside `mcp/` and the docs
+ changed after the integration gates.
+- **YouTube window** (`teamrcn/2Xu514rMOAI`, start 10 → end 20, reason "release 10 live proof"):
+ 202 → done in 16 s; `data/2Xu514rMOAI/clips/7.00-23.00.mp4` (5,746,423 bytes; pad 3 both sides),
+ sidecar `7.00-23.00.json` with `requestedBy: "mcp"` and the reason; `metadata.info.json` kept its
+ 07:01 mtime; the job (`01M3FH6T4NKQT5H12CQBREPE1M`, kind fetch-window) polls `done`. The same
+ call again: "Already on disk — the exact window" in 0.1 s; a narrower 12 → 15 inside it: "a WIDER
+ cached window that contains 9.00–18.00: 7.00–23.00".
+- **Rumble mapping** (`the-quartering-rumble`, cited embed id `vxe1ae`, 83.72 → 86.76): the tool
+ posted the slug id — the clip landed in `data/v1007ay/clips/80.72-89.76.mp4` (1,614,518 bytes,
+ 13 s), no `data/vxe1ae/` was created. The Rumble window itself downloaded (the 2026-09-24 embedJS
+ 403 did not bite this path).
+- **Whole recording** (`full: true`): FIRST attempt on the captions-only `teamrcn/dbnS-cBgStY`
+ (`handling: "youtube"`, transcript on disk) was a 7 s no-op — both yt-dlp passes ran with
+ `--skip-download` and the job named no file; the tool answered with the planned text ("finished
+ but named no file; check the video's page in the editor"). Cause, editor-side and pre-existing:
+ the persistence plan with `keepSourceVideoOverride` applies only to transcribe handling or the
+ no-captions fallback (`downloadOneManaged.ts:1233-1235`), so `fetchFullSourceAction` (umtool's
+ `--full` too) cannot archive the source of a youtube-handling video that already has a
+ transcript. It DID rewrite `metadata.info.json` (the full path writes metadata; the window path
+ never does). **Low, owed:** full mode on a captions-only channel should either force the media
+ step or refuse with one line. SECOND attempt on transcribe-handling `midwestly/
+ 79934a667be267dc6b2126cad3fdde79266739da` (Odysee, 49 s): 202 → done in 32 s, file
+ `transcripts/saved-videos/midwestly/<id>/source-media.mp4` (6,561,079 bytes), pointer
+ `saved-video.json` with `keepReason: "override"` and `origin.requestedBy: "mcp"`; the same call
+ again: "Already on disk — the whole recording" in 0.1 s. Side effect of the editor's path, not
+ the tool's: the transcribe download also left `audio.mp3` beside the transcript. The `note:`
+ prefix fired correctly (midwestly is not in the local Jeralyzer export).
+- **Re-registered** `archilyzer` for this machine with `ARCHILYZER_EDITOR_URL=http://localhost:3001`
+ and `WORKER_TOKEN` (from `editor/.env`); `claude mcp get archilyzer` shows both and Connected.
+- **Owed (operator):** the `/ask` proof in a fresh session — ask for the clip behind one citation;
+ the plan step names `fetch_clip` and the tool answers with a corpus path.