# Backfill re-acquire on subtitle channels — fetch audio, and keep the cues fresh ## Context **Every fact here was verified read-only against the tree at `9ffa437` (clean) and the real corpus on 2026-08-26/27.** Corpus-wide digest `deferred` is 16,156. A census with the real `isCuesJsonFresh` over all 78,128 transcribed video dirs: **16,081 `stale`** (cues.json older than a rewritten `metadata.info.json`; in 7,849 of those a re-fetched VTT too), 75 `missing`, 1 `no-meta`. The stale ones sit on eight `handling: "youtube"` channels, written in contiguous serial blocks 08-22 → 08-26 (the-quartering 3,762, nux-taku 1,802, quartering-live 805, destiny 3,869, leaflit 1,612, kirsche 1,197, HasanAbiVODs3 348, chibi-reviews 2,699), each dir's `download-outcome.json` recording a `metadata-prefetch` + `primary` attempt with `handling: "youtube"`. **The mechanism** (`settings.json`: `backfill.reach: corpus`, `order: newest`, `allowRedownload: true`, diarization enabled): the sweep walked those channels; every video on a subtitle-downloading channel is diarization **`missing-input`** (no audio, by design of that handling — `operations.ts:679`); `allowRedownload` turns that into `dispatch` (`backfillBatch.ts:194-196`); `reacquireMediaFor` then calls `downloadOneManaged` with the channel's **unmodified** config (`backfillReacquire.ts:116-133`) — there is no `handling` check anywhere in that file — so yt-dlp ran `--skip-download --write-subs --write-auto-subs`, rewrote metadata (+ VTT when YouTube served one), landed no audio, "nothing usable landed", next video, ~12/min. ~16k wasted fetches (with firefox cookies), zero diarizations, and every touched video now reads `deferred` to the digest lane and `no-transcript` → `skipped` to attribution (`attributeOne.ts:139`, `operations.ts:983`), because both gate on cues freshness. **Two defects, and the fix needs both:** 1. **Re-acquire on a subtitle channel can never yield audio with the channel's own config.** The precedent for the fix is already in the tree: `autoRunner.ts:996-1010` forces `{ ...rawConfig, handling: "transcribe" }` for one video so a youtube-handling channel downloads audio, "the channel's stored config untouched"; `downloadOneManaged`'s own no-subs fallback (`:932-950`) does the same thing under the same name. This is also the operator's stated model: no channel is truly subs-only — when subs are not what's needed, fall back to audio. 2. **Every re-acquire rewrites `metadata.info.json`** — the metadata prefetch (`downloadOneManaged.ts:517-548`) runs before any attempt, and the transcribe download needs that info json (`reuseInfoJson`). So even a *correct* re-acquire leaves `transcript.cues.json` stale by mtime (`normalizeTranscript.ts:243`), which makes the very next lane in the same sweep (`attribution-diarized`) skip the video with `no-transcript`, and the digest lane defer it, until an operator runs Normalize. The precedent for the fix is `transcribeOne.ts:236-243`: normalize right after the thing that changed the inputs. It is cheap (returns `fresh` when nothing moved), and it invalidates no digest — `isSectionFresh` (`digest.ts:513-534`) is provenance-keyed (app/model/prompt/ contextHash), not cues-mtime. **Content is still correct:** 230 of 240 sampled re-fetched VTTs parse to byte-identical cues (the 10 differ by a few cues of YouTube ASR drift); the 8,231 metadata-only cases never touched the transcript. So the 16,081 are an mtime false positive that one Normalize pass clears — an operator action, not this slice's. **Not running now:** `sweepEnabled: false`, no editor process; the newest outcome files are ordinary sync. **If re-armed unfixed:** 66,540 youtube-handling videos are diarization `missingInput` corpus-wide (chrissie-mayr 3,514, nuxanor 3,211, rev-says-desu 2,770 next in `newest` order, plus ~7,500 unvisited on the-quartering). ## Step 0 — the plan on disk Write this file verbatim to `plans/backfill-reacquire-subtitle-channels.md` and commit it alone: `plans: backfill re-acquire on subtitle channels planned`. (Tree is clean at `9ffa437`; nothing else goes in.) ## Order: three commits 1. **The fix** — `reacquireMediaFor` forces transcribe handling and re-normalizes after the fetch; diarization's `state()` reads the duration cap before dispatching a download. Units for each. 2. **The e2e and the wording** — a backfill spec on a youtube-handling channel; the four places whose prose says "re-acquire" without saying "audio". 3. **Docs** — FACTS, STATE (correct the wrong framing in place), CHANGELOG, memory. tsc in all six packages + `pnpm -C common test` + editor units after each. e2e once after commit 2, **detached** (memory `e2e-run-detached`; the queue lock is serial and a Bash call caps at 10 min): `setsid nohup … pnpm e2e -- backfill.spec.ts auto-subs-replace.spec.ts`. Edit nothing while it runs. If port 3011 is held by another session's server, run on an offset block (`PORT=3111 EXPORT_PORT=3110 OLLAMA_STUB_PORT=11535`) — do not kill it. ## Commit 1 — the fix ### `common/controller/backfillReacquire.ts` - Add a **pure, exported** helper, and use it for the config passed to *both* `findVideoSourceUrl` (`:96`) and `downloadOneManaged` (`:118`), the way autoRunner passes the overridden config to both: ```ts // A backfill wants AUDIO. A `handling: "youtube"` channel's own download is // --skip-download --write-subs --write-auto-subs: it would re-fetch the captions // the video already has and land nothing a diarizer can read — which is exactly // what happened to ~16,000 videos on eight channels, 2026-08-22 → 08-26. Same // override, same reason, as autoRunner's replaceAutoSubs unit and // downloadOneManaged's own no-subs fallback: transcribe-handling for this one // video, the channel's stored config untouched. export function reacquireConfigFor(config: ChannelConfig): ChannelConfig { return config.handling === "transcribe" ? config : { ...config, handling: "transcribe" }; } ``` Log the override once per video when it applies (`Re-acquiring media for — subtitle channel, downloading audio (handling override: transcribe)`), so a run log says what it did. `resolveCookiePolicy(settings, config)` is unaffected either way (it reads cookie fields only) — pass it the same overridden config for consistency. - After `downloadOneManaged` has run — **whether it returned or threw** (the prefetch wrote metadata before any failure; that is the 16,081) — re-normalize: ```ts // The fetch rewrote metadata.info.json, which the cues embed and which // isCuesJsonFresh compares against. Without this the very next lane in the // same sweep (attribution) skips the video as no-transcript and the digest // lane defers it, until an operator runs Normalize. Same step transcribeOne // takes after a transcription, for the same reason. Cheap: `fresh` when // nothing moved. Digest freshness is provenance-keyed, so this invalidates // nothing. ``` Implement as a small non-exported `refreshCues(videoDir, channelSlug, config, videoId, log)` that calls `normalizeTranscript({ videoDir, channelSlug, configName: config.name, log })` (the shape `archiveTranscripts.ts:191` uses), logs `wrote` at info and any throw as a warning (never rethrow — this runs on the failure path too). Call it right after the `try/catch` around the download, before the "nothing usable landed" check. It is not part of `cleanup()`: cues.json is a record, not media, and `buildCleanup` deliberately never touches non-media files (`:158-165`). - Header comment (`:1-29`): add the fifth guard — "AUDIO, NOT THE CHANNEL'S DEFAULT" — with the 2026-08-22→26 measurement, and a line for the re-normalize. Reword the `:129-130` comment ("Audio now, container discarded") to note the handling override is what makes there be audio to extract on a subtitle channel. ### `common/lib/operations.ts` — the cap before the download `state()` for diarization (`:679-689`) returns `missing-input` *before* it reads the duration cap, deliberately (the cap read parses metadata and `countBackfillWork` runs `state()` per video). But with `allowRedownload` armed, `missing-input` dispatches a *download*, and `diarizeOneVideo` does not enforce the cap itself — so a 10-hour VOD over the cap would be fetched and diarized. One metadata read is nothing next to a download: ```ts if (!(await hasDiarizableInput(videoDir, files))) { // Read the cap here ONLY when this branch can dispatch a download: with // re-download armed, missing-input is a fetch, and diarizeOneVideo does not // enforce the cap. Otherwise stay cheap — this runs per video per job start. if ( settings.backfill.allowRedownload && (await isOverDiarizationCap(videoDir, settings.diarization)) ) { return "deferred"; } return "missing-input"; } ``` On this corpus `maxAudioHours` is 0 (cap off), so no count moves today; say so in the commit body. Update the `:680-689` comment, which currently says the cap is read "ONLY here" in the would-be-`missing` branch. ### Tests (commit 1) - New `common/controller/backfillReacquire.test.ts` (the file has none; `backfillBatch.test.ts` tests only pure functions and the repo does not module-mock — keep to the pure surface): `reacquireConfigFor` returns the same object for transcribe handling; returns a copy with `handling: "transcribe"` and every other field intact (`audioFormat`, `name`, `platform`, `cookies…`) for youtube handling; never mutates its input. - For the re-normalize, a tmp-dir test in the same file **if** `normalizeTranscript` can be driven on a synthetic dir the way `normalizeAll.test.ts` does (check its fixture shape first): write `metadata.info.json` + `transcript.en.vtt` + an older `transcript.cues.json`, bump the metadata mtime, call `refreshCues` (export it for the test, or test through `normalizeTranscript` directly if exporting is churn), assert `isCuesJsonFresh` is true after. If `normalizeAll.test.ts` shows this needs more scaffolding than a dozen lines, drop it and rely on the e2e in commit 2, and say so in the report. - `common/lib/operations.test.ts`: diarization `state()` on a transcribed video with no audio → `missing-input` with the cap set but `allowRedownload` off; → `deferred` with both set and `duration` over the cap; → `missing-input` with both set and duration under the cap or absent. Find the existing diarization `state()` cases in that file and follow their fixture shape (`files`, `settings`, a tmp `videoDir` with `metadata.info.json`). Commit message: `backfill: re-acquire fetches audio on a subtitle channel and keeps the cues fresh`. Body: the mechanism, the census numbers (16,081 stale / eight channels / dates), the two precedents, the cap guard and that it moves nothing at `maxAudioHours: 0`. ## Commit 2 — the e2e and the wording ### `editor/e2e/backfill.spec.ts` — a new numbered case after (10) "(11) THE RE-DOWNLOAD ON A SUBTITLE CHANNEL: audio, not captions." Seed a `handling: "youtube"` channel the way `auto-subs-replace.spec.ts` `seedChannel` (`:83-130`) does — `config.json` with `handling: "youtube"`, `audioFormat: "mp3"`, a `playlist` line per id, per video `metadata.info.json` (with `automatic_captions`) + `transcript.en.vtt` and **no audio** — copy the helper locally rather than importing across specs unless `helpers.ts` already has one. Also write a `transcript.cues.json` older than the metadata, or simply omit it (either way the assertion below holds). Then, with `backfillSettings({ backfill: { allowRedownload: true } })`, open the channel's speakers stage and click "Run speaker work" as cases (4)/(5) do. Assert: - `diarization.json` lands (the fake yt-dlp writes `audio.mp3` for a transcribe-handling invocation — `fixtures/bin/fake-ytdlp.mjs:167-198`, the same path auto-subs-replace relies on at its `:203`); - the audio does not survive (`audioFiles(id)` → `[]`), as in case (4); - `transcript.en.vtt` is byte-identical to what was seeded (no subtitle re-fetch); - `download-outcome.json`'s last attempt has `handling: "transcribe"`; - `transcript.cues.json` exists and its mtime ≥ `metadata.info.json`'s — the property the 16,081 lost. If the fake yt-dlp's prefetch branch (`:621+`) needs the youtube URL shape to write metadata for a transcribe-handling invocation, check auto-subs-replace's fixture ids first; it already exercises exactly this override end to end. ### Wording — four places that say "re-acquire" without saying "audio" - `common/lib/settings.ts:325-331` (`allowRedownload` doc comment): add that on a `handling: "youtube"` channel this downloads **audio** with a per-video transcribe override — the channel's config is not changed. - `editor/app/operations/components/settings/LaneSettingsForm.tsx:96-112` (the checkbox hint): one sentence, same content, operator-facing. - `editor/app/channels/[slug]/components/stages/SpeakersStage.tsx:165-167`: "re-download is on, so this run will fetch **audio** and then delete it …". - `common/lib/operations.ts:178-186` ("backfillReacquire answers it by fetching AUDIO") is now true; leave it, but the `:786-790` sentence ("What it needs re-acquiring media for is already controller/backfillReacquire.ts") can stay as is. Commit message: `backfill: the subtitle-channel re-acquire is pinned end to end`. ## Commit 3 — docs and memory - `plans/FACTS.md` — new dated section "Verified 2026-08-27 — why `deferred` was 16,156": the census (16,081 stale / 75 missing / 1 no-meta; 8,231 meta-only vs 7,849 meta+VTT; dates; eight channels with counts; per-channel redownloads ≈ deferred ≈ a slice of `missingInput`), the mechanism with the file:line trail, the content check (230/240), `isSectionFresh` being provenance-keyed (why normalize is safe), the 66,540 exposure, and the two precedents for the override. - `plans/STATE.md` — **correct the finding in place, dated**: the "Recommended next" #2 (`:152-160`) and the "THE FINDING TO ACT ON" paragraph (`:211-218`) both frame it as accumulation-vs-regression; it is neither. Replace with: cause found, fix shipped (shas), and the operator runbook below. Keep "GPU yield on a quiet box" as #1. - `editor/CHANGELOG.md` [Unreleased]: re-acquiring media for the speaker lane on a subtitle-downloading channel now downloads audio (per-video transcribe override, channel config unchanged) instead of re-fetching captions; a re-acquire re-normalizes the transcript so the digest and attribution lanes do not defer the video; a video over the diarization duration cap is deferred before any download is spent on it. - Memory: update `slice-3-chosen-next.md` ("CAUSE FOUND" paragraph → fixed, shas); index line to match. Commit message: `plans: the deferred cause is fixed and recorded`. ## Operator runbook (for STATE.md — not the agent's work, and NOT run from this slice) 1. Keep `sweepEnabled: false` until commit 1 is on the running editor. (It is off now.) 2. **Clear the 16,081**: Digest stage card → *Normalize transcripts* on the-quartering, destiny, chibi-reviews, nux-taku, leaflit, kirsche, quartering-live, HasanAbiVODs3 (and shondo-vods, 22). Pure re-parse, no network; identical cues in 230/240 sampled cases; digests are not invalidated. Expect corpus `deferred` → ~76 (the `missing`/`no-meta` tail — the same button on their channels clears those; the census script lists them). 3. **Re-affirm `allowRedownload` knowing what it now does.** With the fix, the sweep will *really* fetch audio for the 66,540 youtube-handling `missingInput` videos in `newest` order across the corpus (sleep 30 s between downloads → ≥ 23 days of network before any diarization time), and delete each file after use. Scope it with `sweepChannels`, or turn the flag off, before arming — that is the policy decision, and it is yours. 4. **A known interaction, flagged not fixed:** `autoQueue.transcription` is enabled with `replaceAutoSubs: true`, and it picks from snapshots that regenerate ~1 s after a download finishes. A re-acquired `audio.mp3` on an ASR-only video is exactly what it looks for, so during a long diarization it may start whisper over the auto-captions; the backfill's `finally` then deletes the audio under it (a failed unit, nothing lost) — or whisper wins first and the video gets a better transcript. Worth deciding whether that is wanted before a corpus-wide run. ## Verification 1. After each code commit: `pnpm -C exec tsc --noEmit` for `common editor export homepage umtool mcp`; `pnpm -C common test` (827 → +N new); editor units `pnpm -C editor exec tsx --test "app/**/*.test.ts"` (85 → ±0). 2. Grep gates after commit 2: `grep -n "handling" common/controller/backfillReacquire.ts` shows the override; `grep -rn "reacquireConfigFor" common editor/app` → definition, two call sites, tests. 3. e2e as above, detached, once: `backfill.spec.ts` (all cases, including the new (11)) and `auto-subs-replace.spec.ts` (the override's precedent must still pass). 4. **Read-only corpus check, no editor boot** (AGENTS.md): nothing in this slice writes under `transcripts/`; the only corpus figure to re-state is the one already measured. Do not re-run the census; do not run Normalize. 5. Manual, on the e2e fixture (`PORT=3021 pnpm dev:test`): the speakers stage on a youtube-handling channel with re-download on says it will fetch audio; kill the server, remove `editor/test-transcripts` and `editor/test-settings.json`, `git status` clean. ## Out of scope - Running Normalize over the eight channels (operator button, runbook step 2). - Any change to `isCuesJsonFresh` — a metadata rewrite *is* staleness (the cues embed the summary); the fix is to re-normalize at the one place that rewrites metadata under a transcript, not to loosen the gate. - The auto-transcribe interaction (runbook step 4). - Scoping/throttling the corpus-wide re-download (runbook step 3). A per-channel "re-acquirable" flag or a third handling value — the binary `handling` is a deliberate invariant (`channelConfig.ts:11-13`). - The `deferred` count for `attribution-*` lanes on those videos: it resolves with the normalize pass; nothing to build. ## Handoff — the cadence On approval, Fable does not implement (memory `plan-then-opus-implements`). It spawns one `general-purpose` agent, `model: "opus"`, with: the plan path (after step 0, which the agent commits), the fish-shell caveat (commit via `git commit -F `; quote `[slug]` paths; `cd` persists), never boot against `transcripts/`, e2e detached, tmp files under `$CLAUDE_JOB_DIR/tmp`, and the report contract: commit shas with one line each; exact tsc/test outputs; e2e pass/fail per spec with any retry; every divergence from the plan and why; anything undone. Fable reviews on return (`git log --oneline 9ffa437..`, the hunks in `backfillReacquire.ts`, `operations.ts` and the new spec; re-runs grep gates + `pnpm -C common test` + editor units, not e2e), sends fixes to the same agent via SendMessage, and reports.