commit 8920be42b8489d4eda66688cb3ad4a5962da61c2
parent 990e1cd3ad3d97fedeebb8b6c5c69f3479cb15ec
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Thu, 27 Aug 2026 10:23:43 -0400
plans: backfill re-acquire on subtitle channels planned
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Diffstat:
1 file changed, 306 insertions(+), 0 deletions(-)
diff --git a/plans/backfill-reacquire-subtitle-channels.md b/plans/backfill-reacquire-subtitle-channels.md
@@ -0,0 +1,306 @@
+# Backfill re-acquire on subtitle channels — fetch audio, and keep the cues fresh
+
+## Context
+
+**Every fact here was verified read-only against the tree at `9ffa437` (clean) and the real
+corpus on 2026-08-26/27.** Corpus-wide digest `deferred` is 16,156. A census with the real
+`isCuesJsonFresh` over all 78,128 transcribed video dirs: **16,081 `stale`** (cues.json older
+than a rewritten `metadata.info.json`; in 7,849 of those a re-fetched VTT too), 75 `missing`,
+1 `no-meta`. The stale ones sit on eight `handling: "youtube"` channels, written in contiguous
+serial blocks 08-22 → 08-26 (the-quartering 3,762, nux-taku 1,802, quartering-live 805,
+destiny 3,869, leaflit 1,612, kirsche 1,197, HasanAbiVODs3 348, chibi-reviews 2,699), each
+dir's `download-outcome.json` recording a `metadata-prefetch` + `primary` attempt with
+`handling: "youtube"`.
+
+**The mechanism** (`settings.json`: `backfill.reach: corpus`, `order: newest`,
+`allowRedownload: true`, diarization enabled): the sweep walked those channels; every video on
+a subtitle-downloading channel is diarization **`missing-input`** (no audio, by design of that
+handling — `operations.ts:679`); `allowRedownload` turns that into `dispatch`
+(`backfillBatch.ts:194-196`); `reacquireMediaFor` then calls `downloadOneManaged` with the
+channel's **unmodified** config (`backfillReacquire.ts:116-133`) — there is no `handling`
+check anywhere in that file — so yt-dlp ran `--skip-download --write-subs --write-auto-subs`,
+rewrote metadata (+ VTT when YouTube served one), landed no audio, "nothing usable landed",
+next video, ~12/min. ~16k wasted fetches (with firefox cookies), zero diarizations, and every
+touched video now reads `deferred` to the digest lane and `no-transcript` → `skipped` to
+attribution (`attributeOne.ts:139`, `operations.ts:983`), because both gate on cues freshness.
+
+**Two defects, and the fix needs both:**
+
+1. **Re-acquire on a subtitle channel can never yield audio with the channel's own config.**
+ The precedent for the fix is already in the tree: `autoRunner.ts:996-1010` forces
+ `{ ...rawConfig, handling: "transcribe" }` for one video so a youtube-handling channel
+ downloads audio, "the channel's stored config untouched"; `downloadOneManaged`'s own
+ no-subs fallback (`:932-950`) does the same thing under the same name. This is also the
+ operator's stated model: no channel is truly subs-only — when subs are not what's needed,
+ fall back to audio.
+2. **Every re-acquire rewrites `metadata.info.json`** — the metadata prefetch
+ (`downloadOneManaged.ts:517-548`) runs before any attempt, and the transcribe download
+ needs that info json (`reuseInfoJson`). So even a *correct* re-acquire leaves
+ `transcript.cues.json` stale by mtime (`normalizeTranscript.ts:243`), which makes the
+ very next lane in the same sweep (`attribution-diarized`) skip the video with
+ `no-transcript`, and the digest lane defer it, until an operator runs Normalize. The
+ precedent for the fix is `transcribeOne.ts:236-243`: normalize right after the thing that
+ changed the inputs. It is cheap (returns `fresh` when nothing moved), and it invalidates no
+ digest — `isSectionFresh` (`digest.ts:513-534`) is provenance-keyed (app/model/prompt/
+ contextHash), not cues-mtime.
+
+**Content is still correct:** 230 of 240 sampled re-fetched VTTs parse to byte-identical
+cues (the 10 differ by a few cues of YouTube ASR drift); the 8,231 metadata-only cases never
+touched the transcript. So the 16,081 are an mtime false positive that one Normalize pass
+clears — an operator action, not this slice's.
+
+**Not running now:** `sweepEnabled: false`, no editor process; the newest outcome files are
+ordinary sync. **If re-armed unfixed:** 66,540 youtube-handling videos are diarization
+`missingInput` corpus-wide (chrissie-mayr 3,514, nuxanor 3,211, rev-says-desu 2,770 next in
+`newest` order, plus ~7,500 unvisited on the-quartering).
+
+## Step 0 — the plan on disk
+
+Write this file verbatim to `plans/backfill-reacquire-subtitle-channels.md` and commit it
+alone: `plans: backfill re-acquire on subtitle channels planned`. (Tree is clean at
+`9ffa437`; nothing else goes in.)
+
+## Order: three commits
+
+1. **The fix** — `reacquireMediaFor` forces transcribe handling and re-normalizes after the
+ fetch; diarization's `state()` reads the duration cap before dispatching a download.
+ Units for each.
+2. **The e2e and the wording** — a backfill spec on a youtube-handling channel; the four
+ places whose prose says "re-acquire" without saying "audio".
+3. **Docs** — FACTS, STATE (correct the wrong framing in place), CHANGELOG, memory.
+
+tsc in all six packages + `pnpm -C common test` + editor units after each. e2e once after
+commit 2, **detached** (memory `e2e-run-detached`; the queue lock is serial and a Bash call
+caps at 10 min): `setsid nohup … pnpm e2e -- backfill.spec.ts auto-subs-replace.spec.ts`.
+Edit nothing while it runs. If port 3011 is held by another session's server, run on an
+offset block (`PORT=3111 EXPORT_PORT=3110 OLLAMA_STUB_PORT=11535`) — do not kill it.
+
+## Commit 1 — the fix
+
+### `common/controller/backfillReacquire.ts`
+
+- Add a **pure, exported** helper, and use it for the config passed to *both*
+ `findVideoSourceUrl` (`:96`) and `downloadOneManaged` (`:118`), the way autoRunner passes
+ the overridden config to both:
+
+ ```ts
+ // A backfill wants AUDIO. A `handling: "youtube"` channel's own download is
+ // --skip-download --write-subs --write-auto-subs: it would re-fetch the captions
+ // the video already has and land nothing a diarizer can read — which is exactly
+ // what happened to ~16,000 videos on eight channels, 2026-08-22 → 08-26. Same
+ // override, same reason, as autoRunner's replaceAutoSubs unit and
+ // downloadOneManaged's own no-subs fallback: transcribe-handling for this one
+ // video, the channel's stored config untouched.
+ export function reacquireConfigFor(config: ChannelConfig): ChannelConfig {
+ return config.handling === "transcribe"
+ ? config
+ : { ...config, handling: "transcribe" };
+ }
+ ```
+
+ Log the override once per video when it applies (`Re-acquiring media for <id> — subtitle
+ channel, downloading audio (handling override: transcribe)`), so a run log says what it
+ did. `resolveCookiePolicy(settings, config)` is unaffected either way (it reads cookie
+ fields only) — pass it the same overridden config for consistency.
+- After `downloadOneManaged` has run — **whether it returned or threw** (the prefetch wrote
+ metadata before any failure; that is the 16,081) — re-normalize:
+
+ ```ts
+ // The fetch rewrote metadata.info.json, which the cues embed and which
+ // isCuesJsonFresh compares against. Without this the very next lane in the
+ // same sweep (attribution) skips the video as no-transcript and the digest
+ // lane defers it, until an operator runs Normalize. Same step transcribeOne
+ // takes after a transcription, for the same reason. Cheap: `fresh` when
+ // nothing moved. Digest freshness is provenance-keyed, so this invalidates
+ // nothing.
+ ```
+
+ Implement as a small non-exported `refreshCues(videoDir, channelSlug, config, videoId,
+ log)` that calls `normalizeTranscript({ videoDir, channelSlug, configName: config.name,
+ log })` (the shape `archiveTranscripts.ts:191` uses), logs `wrote` at info and any throw
+ as a warning (never rethrow — this runs on the failure path too). Call it right after the
+ `try/catch` around the download, before the "nothing usable landed" check. It is not part of
+ `cleanup()`: cues.json is a record, not media, and `buildCleanup` deliberately never touches
+ non-media files (`:158-165`).
+- Header comment (`:1-29`): add the fifth guard — "AUDIO, NOT THE CHANNEL'S DEFAULT" — with
+ the 2026-08-22→26 measurement, and a line for the re-normalize. Reword the `:129-130`
+ comment ("Audio now, container discarded") to note the handling override is what makes
+ there be audio to extract on a subtitle channel.
+
+### `common/lib/operations.ts` — the cap before the download
+
+`state()` for diarization (`:679-689`) returns `missing-input` *before* it reads the duration
+cap, deliberately (the cap read parses metadata and `countBackfillWork` runs `state()` per
+video). But with `allowRedownload` armed, `missing-input` dispatches a *download*, and
+`diarizeOneVideo` does not enforce the cap itself — so a 10-hour VOD over the cap would be
+fetched and diarized. One metadata read is nothing next to a download:
+
+```ts
+if (!(await hasDiarizableInput(videoDir, files))) {
+ // Read the cap here ONLY when this branch can dispatch a download: with
+ // re-download armed, missing-input is a fetch, and diarizeOneVideo does not
+ // enforce the cap. Otherwise stay cheap — this runs per video per job start.
+ if (
+ settings.backfill.allowRedownload &&
+ (await isOverDiarizationCap(videoDir, settings.diarization))
+ ) {
+ return "deferred";
+ }
+ return "missing-input";
+}
+```
+
+On this corpus `maxAudioHours` is 0 (cap off), so no count moves today; say so in the commit
+body. Update the `:680-689` comment, which currently says the cap is read "ONLY here" in the
+would-be-`missing` branch.
+
+### Tests (commit 1)
+
+- New `common/controller/backfillReacquire.test.ts` (the file has none; `backfillBatch.test.ts`
+ tests only pure functions and the repo does not module-mock — keep to the pure surface):
+ `reacquireConfigFor` returns the same object for transcribe handling; returns a copy with
+ `handling: "transcribe"` and every other field intact (`audioFormat`, `name`, `platform`,
+ `cookies…`) for youtube handling; never mutates its input.
+- For the re-normalize, a tmp-dir test in the same file **if** `normalizeTranscript` can be
+ driven on a synthetic dir the way `normalizeAll.test.ts` does (check its fixture shape
+ first): write `metadata.info.json` + `transcript.en.vtt` + an older `transcript.cues.json`,
+ bump the metadata mtime, call `refreshCues` (export it for the test, or test through
+ `normalizeTranscript` directly if exporting is churn), assert `isCuesJsonFresh` is true
+ after. If `normalizeAll.test.ts` shows this needs more scaffolding than a dozen lines,
+ drop it and rely on the e2e in commit 2, and say so in the report.
+- `common/lib/operations.test.ts`: diarization `state()` on a transcribed video with no
+ audio → `missing-input` with the cap set but `allowRedownload` off; → `deferred` with both
+ set and `duration` over the cap; → `missing-input` with both set and duration under the cap
+ or absent. Find the existing diarization `state()` cases in that file and follow their
+ fixture shape (`files`, `settings`, a tmp `videoDir` with `metadata.info.json`).
+
+Commit message: `backfill: re-acquire fetches audio on a subtitle channel and keeps the cues
+fresh`. Body: the mechanism, the census numbers (16,081 stale / eight channels / dates), the
+two precedents, the cap guard and that it moves nothing at `maxAudioHours: 0`.
+
+## Commit 2 — the e2e and the wording
+
+### `editor/e2e/backfill.spec.ts` — a new numbered case after (10)
+
+"(11) THE RE-DOWNLOAD ON A SUBTITLE CHANNEL: audio, not captions." Seed a
+`handling: "youtube"` channel the way `auto-subs-replace.spec.ts` `seedChannel` (`:83-130`)
+does — `config.json` with `handling: "youtube"`, `audioFormat: "mp3"`, a `playlist` line per
+id, per video `metadata.info.json` (with `automatic_captions`) + `transcript.en.vtt` and **no
+audio** — copy the helper locally rather than importing across specs unless `helpers.ts`
+already has one. Also write a `transcript.cues.json` older than the metadata, or simply omit
+it (either way the assertion below holds). Then, with
+`backfillSettings({ backfill: { allowRedownload: true } })`, open the channel's speakers stage
+and click "Run speaker work" as cases (4)/(5) do. Assert:
+
+- `diarization.json` lands (the fake yt-dlp writes `audio.mp3` for a transcribe-handling
+ invocation — `fixtures/bin/fake-ytdlp.mjs:167-198`, the same path auto-subs-replace relies
+ on at its `:203`);
+- the audio does not survive (`audioFiles(id)` → `[]`), as in case (4);
+- `transcript.en.vtt` is byte-identical to what was seeded (no subtitle re-fetch);
+- `download-outcome.json`'s last attempt has `handling: "transcribe"`;
+- `transcript.cues.json` exists and its mtime ≥ `metadata.info.json`'s — the property the
+ 16,081 lost.
+
+If the fake yt-dlp's prefetch branch (`:621+`) needs the youtube URL shape to write metadata
+for a transcribe-handling invocation, check auto-subs-replace's fixture ids first; it already
+exercises exactly this override end to end.
+
+### Wording — four places that say "re-acquire" without saying "audio"
+
+- `common/lib/settings.ts:325-331` (`allowRedownload` doc comment): add that on a
+ `handling: "youtube"` channel this downloads **audio** with a per-video transcribe
+ override — the channel's config is not changed.
+- `editor/app/operations/components/settings/LaneSettingsForm.tsx:96-112` (the checkbox
+ hint): one sentence, same content, operator-facing.
+- `editor/app/channels/[slug]/components/stages/SpeakersStage.tsx:165-167`: "re-download is
+ on, so this run will fetch **audio** and then delete it …".
+- `common/lib/operations.ts:178-186` ("backfillReacquire answers it by fetching AUDIO") is
+ now true; leave it, but the `:786-790` sentence ("What it needs re-acquiring media for is
+ already controller/backfillReacquire.ts") can stay as is.
+
+Commit message: `backfill: the subtitle-channel re-acquire is pinned end to end`.
+
+## Commit 3 — docs and memory
+
+- `plans/FACTS.md` — new dated section "Verified 2026-08-27 — why `deferred` was 16,156":
+ the census (16,081 stale / 75 missing / 1 no-meta; 8,231 meta-only vs 7,849 meta+VTT;
+ dates; eight channels with counts; per-channel redownloads ≈ deferred ≈ a slice of
+ `missingInput`), the mechanism with the file:line trail, the content check (230/240),
+ `isSectionFresh` being provenance-keyed (why normalize is safe), the 66,540 exposure, and
+ the two precedents for the override.
+- `plans/STATE.md` — **correct the finding in place, dated**: the "Recommended next" #2
+ (`:152-160`) and the "THE FINDING TO ACT ON" paragraph (`:211-218`) both frame it as
+ accumulation-vs-regression; it is neither. Replace with: cause found, fix shipped (shas),
+ and the operator runbook below. Keep "GPU yield on a quiet box" as #1.
+- `editor/CHANGELOG.md` [Unreleased]: re-acquiring media for the speaker lane on a
+ subtitle-downloading channel now downloads audio (per-video transcribe override, channel
+ config unchanged) instead of re-fetching captions; a re-acquire re-normalizes the transcript
+ so the digest and attribution lanes do not defer the video; a video over the diarization
+ duration cap is deferred before any download is spent on it.
+- Memory: update `slice-3-chosen-next.md` ("CAUSE FOUND" paragraph → fixed, shas); index
+ line to match.
+
+Commit message: `plans: the deferred cause is fixed and recorded`.
+
+## Operator runbook (for STATE.md — not the agent's work, and NOT run from this slice)
+
+1. Keep `sweepEnabled: false` until commit 1 is on the running editor. (It is off now.)
+2. **Clear the 16,081**: Digest stage card → *Normalize transcripts* on the-quartering,
+ destiny, chibi-reviews, nux-taku, leaflit, kirsche, quartering-live, HasanAbiVODs3 (and
+ shondo-vods, 22). Pure re-parse, no network; identical cues in 230/240 sampled cases;
+ digests are not invalidated. Expect corpus `deferred` → ~76 (the `missing`/`no-meta`
+ tail — the same button on their channels clears those; the census script lists them).
+3. **Re-affirm `allowRedownload` knowing what it now does.** With the fix, the sweep will
+ *really* fetch audio for the 66,540 youtube-handling `missingInput` videos in `newest`
+ order across the corpus (sleep 30 s between downloads → ≥ 23 days of network before any
+ diarization time), and delete each file after use. Scope it with `sweepChannels`, or turn
+ the flag off, before arming — that is the policy decision, and it is yours.
+4. **A known interaction, flagged not fixed:** `autoQueue.transcription` is enabled with
+ `replaceAutoSubs: true`, and it picks from snapshots that regenerate ~1 s after a download
+ finishes. A re-acquired `audio.mp3` on an ASR-only video is exactly what it looks for, so
+ during a long diarization it may start whisper over the auto-captions; the backfill's
+ `finally` then deletes the audio under it (a failed unit, nothing lost) — or whisper wins
+ first and the video gets a better transcript. Worth deciding whether that is wanted before
+ a corpus-wide run.
+
+## Verification
+
+1. After each code commit: `pnpm -C <pkg> exec tsc --noEmit` for `common editor export
+ homepage umtool mcp`; `pnpm -C common test` (827 → +N new); editor units
+ `pnpm -C editor exec tsx --test "app/**/*.test.ts"` (85 → ±0).
+2. Grep gates after commit 2: `grep -n "handling" common/controller/backfillReacquire.ts`
+ shows the override; `grep -rn "reacquireConfigFor" common editor/app` → definition, two
+ call sites, tests.
+3. e2e as above, detached, once: `backfill.spec.ts` (all cases, including the new (11)) and
+ `auto-subs-replace.spec.ts` (the override's precedent must still pass).
+4. **Read-only corpus check, no editor boot** (AGENTS.md): nothing in this slice writes under
+ `transcripts/`; the only corpus figure to re-state is the one already measured. Do not
+ re-run the census; do not run Normalize.
+5. Manual, on the e2e fixture (`PORT=3021 pnpm dev:test`): the speakers stage on a
+ youtube-handling channel with re-download on says it will fetch audio; kill the server,
+ remove `editor/test-transcripts` and `editor/test-settings.json`, `git status` clean.
+
+## Out of scope
+
+- Running Normalize over the eight channels (operator button, runbook step 2).
+- Any change to `isCuesJsonFresh` — a metadata rewrite *is* staleness (the cues embed the
+ summary); the fix is to re-normalize at the one place that rewrites metadata under a
+ transcript, not to loosen the gate.
+- The auto-transcribe interaction (runbook step 4).
+- Scoping/throttling the corpus-wide re-download (runbook step 3). A per-channel
+ "re-acquirable" flag or a third handling value — the binary `handling` is a deliberate
+ invariant (`channelConfig.ts:11-13`).
+- The `deferred` count for `attribution-*` lanes on those videos: it resolves with the
+ normalize pass; nothing to build.
+
+## Handoff — the cadence
+
+On approval, Fable does not implement (memory `plan-then-opus-implements`). It spawns one
+`general-purpose` agent, `model: "opus"`, with: the plan path (after step 0, which the agent
+commits), the fish-shell caveat (commit via `git commit -F <file under $CLAUDE_JOB_DIR/tmp>`;
+quote `[slug]` paths; `cd` persists), never boot against `transcripts/`, e2e detached, tmp
+files under `$CLAUDE_JOB_DIR/tmp`, and the report contract: commit shas with one line each;
+exact tsc/test outputs; e2e pass/fail per spec with any retry; every divergence from the plan
+and why; anything undone. Fable reviews on return (`git log --oneline 9ffa437..`, the hunks in
+`backfillReacquire.ts`, `operations.ts` and the new spec; re-runs grep gates + `pnpm -C common
+test` + editor units, not e2e), sends fixes to the same agent via SendMessage, and reports.