Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 8920be42b8489d4eda66688cb3ad4a5962da61c2
parent 990e1cd3ad3d97fedeebb8b6c5c69f3479cb15ec
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Thu, 27 Aug 2026 10:23:43 -0400

plans: backfill re-acquire on subtitle channels planned

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

Diffstat:
Aplans/backfill-reacquire-subtitle-channels.md | 306+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
1 file changed, 306 insertions(+), 0 deletions(-)

diff --git a/plans/backfill-reacquire-subtitle-channels.md b/plans/backfill-reacquire-subtitle-channels.md @@ -0,0 +1,306 @@ +# Backfill re-acquire on subtitle channels — fetch audio, and keep the cues fresh + +## Context + +**Every fact here was verified read-only against the tree at `9ffa437` (clean) and the real +corpus on 2026-08-26/27.** Corpus-wide digest `deferred` is 16,156. A census with the real +`isCuesJsonFresh` over all 78,128 transcribed video dirs: **16,081 `stale`** (cues.json older +than a rewritten `metadata.info.json`; in 7,849 of those a re-fetched VTT too), 75 `missing`, +1 `no-meta`. The stale ones sit on eight `handling: "youtube"` channels, written in contiguous +serial blocks 08-22 → 08-26 (the-quartering 3,762, nux-taku 1,802, quartering-live 805, +destiny 3,869, leaflit 1,612, kirsche 1,197, HasanAbiVODs3 348, chibi-reviews 2,699), each +dir's `download-outcome.json` recording a `metadata-prefetch` + `primary` attempt with +`handling: "youtube"`. + +**The mechanism** (`settings.json`: `backfill.reach: corpus`, `order: newest`, +`allowRedownload: true`, diarization enabled): the sweep walked those channels; every video on +a subtitle-downloading channel is diarization **`missing-input`** (no audio, by design of that +handling — `operations.ts:679`); `allowRedownload` turns that into `dispatch` +(`backfillBatch.ts:194-196`); `reacquireMediaFor` then calls `downloadOneManaged` with the +channel's **unmodified** config (`backfillReacquire.ts:116-133`) — there is no `handling` +check anywhere in that file — so yt-dlp ran `--skip-download --write-subs --write-auto-subs`, +rewrote metadata (+ VTT when YouTube served one), landed no audio, "nothing usable landed", +next video, ~12/min. ~16k wasted fetches (with firefox cookies), zero diarizations, and every +touched video now reads `deferred` to the digest lane and `no-transcript` → `skipped` to +attribution (`attributeOne.ts:139`, `operations.ts:983`), because both gate on cues freshness. + +**Two defects, and the fix needs both:** + +1. **Re-acquire on a subtitle channel can never yield audio with the channel's own config.** + The precedent for the fix is already in the tree: `autoRunner.ts:996-1010` forces + `{ ...rawConfig, handling: "transcribe" }` for one video so a youtube-handling channel + downloads audio, "the channel's stored config untouched"; `downloadOneManaged`'s own + no-subs fallback (`:932-950`) does the same thing under the same name. This is also the + operator's stated model: no channel is truly subs-only — when subs are not what's needed, + fall back to audio. +2. **Every re-acquire rewrites `metadata.info.json`** — the metadata prefetch + (`downloadOneManaged.ts:517-548`) runs before any attempt, and the transcribe download + needs that info json (`reuseInfoJson`). So even a *correct* re-acquire leaves + `transcript.cues.json` stale by mtime (`normalizeTranscript.ts:243`), which makes the + very next lane in the same sweep (`attribution-diarized`) skip the video with + `no-transcript`, and the digest lane defer it, until an operator runs Normalize. The + precedent for the fix is `transcribeOne.ts:236-243`: normalize right after the thing that + changed the inputs. It is cheap (returns `fresh` when nothing moved), and it invalidates no + digest — `isSectionFresh` (`digest.ts:513-534`) is provenance-keyed (app/model/prompt/ + contextHash), not cues-mtime. + +**Content is still correct:** 230 of 240 sampled re-fetched VTTs parse to byte-identical +cues (the 10 differ by a few cues of YouTube ASR drift); the 8,231 metadata-only cases never +touched the transcript. So the 16,081 are an mtime false positive that one Normalize pass +clears — an operator action, not this slice's. + +**Not running now:** `sweepEnabled: false`, no editor process; the newest outcome files are +ordinary sync. **If re-armed unfixed:** 66,540 youtube-handling videos are diarization +`missingInput` corpus-wide (chrissie-mayr 3,514, nuxanor 3,211, rev-says-desu 2,770 next in +`newest` order, plus ~7,500 unvisited on the-quartering). + +## Step 0 — the plan on disk + +Write this file verbatim to `plans/backfill-reacquire-subtitle-channels.md` and commit it +alone: `plans: backfill re-acquire on subtitle channels planned`. (Tree is clean at +`9ffa437`; nothing else goes in.) + +## Order: three commits + +1. **The fix** — `reacquireMediaFor` forces transcribe handling and re-normalizes after the + fetch; diarization's `state()` reads the duration cap before dispatching a download. + Units for each. +2. **The e2e and the wording** — a backfill spec on a youtube-handling channel; the four + places whose prose says "re-acquire" without saying "audio". +3. **Docs** — FACTS, STATE (correct the wrong framing in place), CHANGELOG, memory. + +tsc in all six packages + `pnpm -C common test` + editor units after each. e2e once after +commit 2, **detached** (memory `e2e-run-detached`; the queue lock is serial and a Bash call +caps at 10 min): `setsid nohup … pnpm e2e -- backfill.spec.ts auto-subs-replace.spec.ts`. +Edit nothing while it runs. If port 3011 is held by another session's server, run on an +offset block (`PORT=3111 EXPORT_PORT=3110 OLLAMA_STUB_PORT=11535`) — do not kill it. + +## Commit 1 — the fix + +### `common/controller/backfillReacquire.ts` + +- Add a **pure, exported** helper, and use it for the config passed to *both* + `findVideoSourceUrl` (`:96`) and `downloadOneManaged` (`:118`), the way autoRunner passes + the overridden config to both: + + ```ts + // A backfill wants AUDIO. A `handling: "youtube"` channel's own download is + // --skip-download --write-subs --write-auto-subs: it would re-fetch the captions + // the video already has and land nothing a diarizer can read — which is exactly + // what happened to ~16,000 videos on eight channels, 2026-08-22 → 08-26. Same + // override, same reason, as autoRunner's replaceAutoSubs unit and + // downloadOneManaged's own no-subs fallback: transcribe-handling for this one + // video, the channel's stored config untouched. + export function reacquireConfigFor(config: ChannelConfig): ChannelConfig { + return config.handling === "transcribe" + ? config + : { ...config, handling: "transcribe" }; + } + ``` + + Log the override once per video when it applies (`Re-acquiring media for <id> — subtitle + channel, downloading audio (handling override: transcribe)`), so a run log says what it + did. `resolveCookiePolicy(settings, config)` is unaffected either way (it reads cookie + fields only) — pass it the same overridden config for consistency. +- After `downloadOneManaged` has run — **whether it returned or threw** (the prefetch wrote + metadata before any failure; that is the 16,081) — re-normalize: + + ```ts + // The fetch rewrote metadata.info.json, which the cues embed and which + // isCuesJsonFresh compares against. Without this the very next lane in the + // same sweep (attribution) skips the video as no-transcript and the digest + // lane defers it, until an operator runs Normalize. Same step transcribeOne + // takes after a transcription, for the same reason. Cheap: `fresh` when + // nothing moved. Digest freshness is provenance-keyed, so this invalidates + // nothing. + ``` + + Implement as a small non-exported `refreshCues(videoDir, channelSlug, config, videoId, + log)` that calls `normalizeTranscript({ videoDir, channelSlug, configName: config.name, + log })` (the shape `archiveTranscripts.ts:191` uses), logs `wrote` at info and any throw + as a warning (never rethrow — this runs on the failure path too). Call it right after the + `try/catch` around the download, before the "nothing usable landed" check. It is not part of + `cleanup()`: cues.json is a record, not media, and `buildCleanup` deliberately never touches + non-media files (`:158-165`). +- Header comment (`:1-29`): add the fifth guard — "AUDIO, NOT THE CHANNEL'S DEFAULT" — with + the 2026-08-22→26 measurement, and a line for the re-normalize. Reword the `:129-130` + comment ("Audio now, container discarded") to note the handling override is what makes + there be audio to extract on a subtitle channel. + +### `common/lib/operations.ts` — the cap before the download + +`state()` for diarization (`:679-689`) returns `missing-input` *before* it reads the duration +cap, deliberately (the cap read parses metadata and `countBackfillWork` runs `state()` per +video). But with `allowRedownload` armed, `missing-input` dispatches a *download*, and +`diarizeOneVideo` does not enforce the cap itself — so a 10-hour VOD over the cap would be +fetched and diarized. One metadata read is nothing next to a download: + +```ts +if (!(await hasDiarizableInput(videoDir, files))) { + // Read the cap here ONLY when this branch can dispatch a download: with + // re-download armed, missing-input is a fetch, and diarizeOneVideo does not + // enforce the cap. Otherwise stay cheap — this runs per video per job start. + if ( + settings.backfill.allowRedownload && + (await isOverDiarizationCap(videoDir, settings.diarization)) + ) { + return "deferred"; + } + return "missing-input"; +} +``` + +On this corpus `maxAudioHours` is 0 (cap off), so no count moves today; say so in the commit +body. Update the `:680-689` comment, which currently says the cap is read "ONLY here" in the +would-be-`missing` branch. + +### Tests (commit 1) + +- New `common/controller/backfillReacquire.test.ts` (the file has none; `backfillBatch.test.ts` + tests only pure functions and the repo does not module-mock — keep to the pure surface): + `reacquireConfigFor` returns the same object for transcribe handling; returns a copy with + `handling: "transcribe"` and every other field intact (`audioFormat`, `name`, `platform`, + `cookies…`) for youtube handling; never mutates its input. +- For the re-normalize, a tmp-dir test in the same file **if** `normalizeTranscript` can be + driven on a synthetic dir the way `normalizeAll.test.ts` does (check its fixture shape + first): write `metadata.info.json` + `transcript.en.vtt` + an older `transcript.cues.json`, + bump the metadata mtime, call `refreshCues` (export it for the test, or test through + `normalizeTranscript` directly if exporting is churn), assert `isCuesJsonFresh` is true + after. If `normalizeAll.test.ts` shows this needs more scaffolding than a dozen lines, + drop it and rely on the e2e in commit 2, and say so in the report. +- `common/lib/operations.test.ts`: diarization `state()` on a transcribed video with no + audio → `missing-input` with the cap set but `allowRedownload` off; → `deferred` with both + set and `duration` over the cap; → `missing-input` with both set and duration under the cap + or absent. Find the existing diarization `state()` cases in that file and follow their + fixture shape (`files`, `settings`, a tmp `videoDir` with `metadata.info.json`). + +Commit message: `backfill: re-acquire fetches audio on a subtitle channel and keeps the cues +fresh`. Body: the mechanism, the census numbers (16,081 stale / eight channels / dates), the +two precedents, the cap guard and that it moves nothing at `maxAudioHours: 0`. + +## Commit 2 — the e2e and the wording + +### `editor/e2e/backfill.spec.ts` — a new numbered case after (10) + +"(11) THE RE-DOWNLOAD ON A SUBTITLE CHANNEL: audio, not captions." Seed a +`handling: "youtube"` channel the way `auto-subs-replace.spec.ts` `seedChannel` (`:83-130`) +does — `config.json` with `handling: "youtube"`, `audioFormat: "mp3"`, a `playlist` line per +id, per video `metadata.info.json` (with `automatic_captions`) + `transcript.en.vtt` and **no +audio** — copy the helper locally rather than importing across specs unless `helpers.ts` +already has one. Also write a `transcript.cues.json` older than the metadata, or simply omit +it (either way the assertion below holds). Then, with +`backfillSettings({ backfill: { allowRedownload: true } })`, open the channel's speakers stage +and click "Run speaker work" as cases (4)/(5) do. Assert: + +- `diarization.json` lands (the fake yt-dlp writes `audio.mp3` for a transcribe-handling + invocation — `fixtures/bin/fake-ytdlp.mjs:167-198`, the same path auto-subs-replace relies + on at its `:203`); +- the audio does not survive (`audioFiles(id)` → `[]`), as in case (4); +- `transcript.en.vtt` is byte-identical to what was seeded (no subtitle re-fetch); +- `download-outcome.json`'s last attempt has `handling: "transcribe"`; +- `transcript.cues.json` exists and its mtime ≥ `metadata.info.json`'s — the property the + 16,081 lost. + +If the fake yt-dlp's prefetch branch (`:621+`) needs the youtube URL shape to write metadata +for a transcribe-handling invocation, check auto-subs-replace's fixture ids first; it already +exercises exactly this override end to end. + +### Wording — four places that say "re-acquire" without saying "audio" + +- `common/lib/settings.ts:325-331` (`allowRedownload` doc comment): add that on a + `handling: "youtube"` channel this downloads **audio** with a per-video transcribe + override — the channel's config is not changed. +- `editor/app/operations/components/settings/LaneSettingsForm.tsx:96-112` (the checkbox + hint): one sentence, same content, operator-facing. +- `editor/app/channels/[slug]/components/stages/SpeakersStage.tsx:165-167`: "re-download is + on, so this run will fetch **audio** and then delete it …". +- `common/lib/operations.ts:178-186` ("backfillReacquire answers it by fetching AUDIO") is + now true; leave it, but the `:786-790` sentence ("What it needs re-acquiring media for is + already controller/backfillReacquire.ts") can stay as is. + +Commit message: `backfill: the subtitle-channel re-acquire is pinned end to end`. + +## Commit 3 — docs and memory + +- `plans/FACTS.md` — new dated section "Verified 2026-08-27 — why `deferred` was 16,156": + the census (16,081 stale / 75 missing / 1 no-meta; 8,231 meta-only vs 7,849 meta+VTT; + dates; eight channels with counts; per-channel redownloads ≈ deferred ≈ a slice of + `missingInput`), the mechanism with the file:line trail, the content check (230/240), + `isSectionFresh` being provenance-keyed (why normalize is safe), the 66,540 exposure, and + the two precedents for the override. +- `plans/STATE.md` — **correct the finding in place, dated**: the "Recommended next" #2 + (`:152-160`) and the "THE FINDING TO ACT ON" paragraph (`:211-218`) both frame it as + accumulation-vs-regression; it is neither. Replace with: cause found, fix shipped (shas), + and the operator runbook below. Keep "GPU yield on a quiet box" as #1. +- `editor/CHANGELOG.md` [Unreleased]: re-acquiring media for the speaker lane on a + subtitle-downloading channel now downloads audio (per-video transcribe override, channel + config unchanged) instead of re-fetching captions; a re-acquire re-normalizes the transcript + so the digest and attribution lanes do not defer the video; a video over the diarization + duration cap is deferred before any download is spent on it. +- Memory: update `slice-3-chosen-next.md` ("CAUSE FOUND" paragraph → fixed, shas); index + line to match. + +Commit message: `plans: the deferred cause is fixed and recorded`. + +## Operator runbook (for STATE.md — not the agent's work, and NOT run from this slice) + +1. Keep `sweepEnabled: false` until commit 1 is on the running editor. (It is off now.) +2. **Clear the 16,081**: Digest stage card → *Normalize transcripts* on the-quartering, + destiny, chibi-reviews, nux-taku, leaflit, kirsche, quartering-live, HasanAbiVODs3 (and + shondo-vods, 22). Pure re-parse, no network; identical cues in 230/240 sampled cases; + digests are not invalidated. Expect corpus `deferred` → ~76 (the `missing`/`no-meta` + tail — the same button on their channels clears those; the census script lists them). +3. **Re-affirm `allowRedownload` knowing what it now does.** With the fix, the sweep will + *really* fetch audio for the 66,540 youtube-handling `missingInput` videos in `newest` + order across the corpus (sleep 30 s between downloads → ≥ 23 days of network before any + diarization time), and delete each file after use. Scope it with `sweepChannels`, or turn + the flag off, before arming — that is the policy decision, and it is yours. +4. **A known interaction, flagged not fixed:** `autoQueue.transcription` is enabled with + `replaceAutoSubs: true`, and it picks from snapshots that regenerate ~1 s after a download + finishes. A re-acquired `audio.mp3` on an ASR-only video is exactly what it looks for, so + during a long diarization it may start whisper over the auto-captions; the backfill's + `finally` then deletes the audio under it (a failed unit, nothing lost) — or whisper wins + first and the video gets a better transcript. Worth deciding whether that is wanted before + a corpus-wide run. + +## Verification + +1. After each code commit: `pnpm -C <pkg> exec tsc --noEmit` for `common editor export + homepage umtool mcp`; `pnpm -C common test` (827 → +N new); editor units + `pnpm -C editor exec tsx --test "app/**/*.test.ts"` (85 → ±0). +2. Grep gates after commit 2: `grep -n "handling" common/controller/backfillReacquire.ts` + shows the override; `grep -rn "reacquireConfigFor" common editor/app` → definition, two + call sites, tests. +3. e2e as above, detached, once: `backfill.spec.ts` (all cases, including the new (11)) and + `auto-subs-replace.spec.ts` (the override's precedent must still pass). +4. **Read-only corpus check, no editor boot** (AGENTS.md): nothing in this slice writes under + `transcripts/`; the only corpus figure to re-state is the one already measured. Do not + re-run the census; do not run Normalize. +5. Manual, on the e2e fixture (`PORT=3021 pnpm dev:test`): the speakers stage on a + youtube-handling channel with re-download on says it will fetch audio; kill the server, + remove `editor/test-transcripts` and `editor/test-settings.json`, `git status` clean. + +## Out of scope + +- Running Normalize over the eight channels (operator button, runbook step 2). +- Any change to `isCuesJsonFresh` — a metadata rewrite *is* staleness (the cues embed the + summary); the fix is to re-normalize at the one place that rewrites metadata under a + transcript, not to loosen the gate. +- The auto-transcribe interaction (runbook step 4). +- Scoping/throttling the corpus-wide re-download (runbook step 3). A per-channel + "re-acquirable" flag or a third handling value — the binary `handling` is a deliberate + invariant (`channelConfig.ts:11-13`). +- The `deferred` count for `attribution-*` lanes on those videos: it resolves with the + normalize pass; nothing to build. + +## Handoff — the cadence + +On approval, Fable does not implement (memory `plan-then-opus-implements`). It spawns one +`general-purpose` agent, `model: "opus"`, with: the plan path (after step 0, which the agent +commits), the fish-shell caveat (commit via `git commit -F <file under $CLAUDE_JOB_DIR/tmp>`; +quote `[slug]` paths; `cd` persists), never boot against `transcripts/`, e2e detached, tmp +files under `$CLAUDE_JOB_DIR/tmp`, and the report contract: commit shas with one line each; +exact tsc/test outputs; e2e pass/fail per spec with any retry; every divergence from the plan +and why; anything undone. Fable reviews on return (`git log --oneline 9ffa437..`, the hunks in +`backfillReacquire.ts`, `operations.ts` and the new spec; re-runs grep gates + `pnpm -C common +test` + editor units, not e2e), sends fixes to the same agent via SendMessage, and reports.