Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit cea8c991af9aac51f5a4b57b4043c42d9ea1c1be
parent 88f045ef11a3edfa85e46538f5d50d8a22150b77
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Fri, 28 Aug 2026 00:21:24 -0400

backfill: the hand-off is pinned end to end, and the docs say so

Two e2e cases in backfill.spec.ts, on the same subtitle channel case (11) seeds,
walking the two sides of the keep decision through the real code path:

  (12) with autoQueue.transcription enabled, replaceAutoSubs on and one
       catch-all leaf: diarization.json still lands, the per-video log says
       "handed to auto-transcribe", the run's summary line reads "0 cleaned up,
       1 handed to auto-transcribe", the audio SURVIVES the item that fetched it
       — and the channel's snapshot, regenerated at job end, lists the video
       under buckets.downloadedAutoSubsOnly. That last assertion is the hand-off
       actually completing rather than a file left lying around.

  (12b) with no autoQueue policy: "Removed re-acquired media", no "handed to",
        no audio. Today's behaviour, pinned through the new code.

The seed had to change to make (12) mean anything. seedSubtitleChannel takes the
VTT now, because case (11)'s SEEDED_VTT carries no ASR fingerprints — provenance
falls back to metadata.info.json, which the fake yt-dlp's re-acquire rewrites
WITHOUT automatic_captions, so the video would not be ASR-only and the hand-off
would never fire for the wrong reason. Both new cases seed ASR_VTT (the copy
from auto-subs-replace.spec.ts) so the 4 KB sniff is what decides.

The auto runner is deliberately NOT started: startAutoRunnersIfEnabled runs only
at boot and the invalidate-cache route starts nothing, so the policy is read and
nothing races the assertions. Driving the runner end to end here would be a
timing-bound test of the lane auto-subs-replace.spec.ts already walks.

Three prose sites that promised unconditional deletion now name the exception:
settings.ts's allowRedownload doc, the lane settings checkbox, and the speakers
stage's re-download sentence.

FACTS.md gains "Verified 2026-08-27 — the re-acquire / auto-transcribe hand-off":
the no-per-video-lock anchor, the seven-step collision trail, the two facts that
made a hand-off the right shape rather than a lock, the resume-mark choice, the
remote ENOENT wrapping, the dead "failed" cleanup — and two corrections. The
sweep has NO inter-download sleep and is channel-major
(backfillSweep.ts:55, 335-368), so STATE.md's "30 s between downloads → ≥ 23
days" was wrong; and the digest lane is equally blind to a transcript
replacement, recorded as a follow-up with its reason for not being fixed.

STATE.md's runbook item 2 step 4 is struck through and RESOLVED with the sha,
including the part the original note got wrong: the losing race is not "a failed
unit, nothing lost" but a video appended to failed-transcriptions, which the
manual per-channel batch honours permanently. The operator no longer needs
replaceAutoSubs off for the corpus run; the trade is that handed-off audio waits
for a Clean-audio sweep.

CHANGELOG [Unreleased] gets the three operator-facing lines. Memory: new
reacquire-handoff.md, its MEMORY.md pointer, and slice-3-chosen-next amended
where it still said step 4 was open and quoted the 23-day figure.

Nothing under transcripts/ was read or written for this commit.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

Diffstat:
Mcommon/lib/settings.ts | 4+++-
Meditor/CHANGELOG.md | 3+++
Meditor/app/channels/[slug]/components/stages/SpeakersStage.tsx | 2+-
Meditor/app/operations/components/settings/LaneSettingsForm.tsx | 7+++++--
Meditor/e2e/backfill.spec.ts | 171+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++------
Mplans/FACTS.md | 82+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mplans/STATE.md | 37+++++++++++++++++++++++++++----------
7 files changed, 279 insertions(+), 27 deletions(-)

diff --git a/common/lib/settings.ts b/common/lib/settings.ts @@ -327,7 +327,9 @@ export type BackfillSettings = { // 836 videos still have media and ~76,270 would need a re-download — 91x the // reachable work, against 45 GB free at 97% full. When on, each re-fetched // file is removed in a `finally` as soon as the backfill has used it, unless - // the video is marked do-not-clean. + // the video is marked do-not-clean, or unless the auto-transcribe policy would + // replace its auto-captions (`replaceAutoSubs`, or a leaf on + // `downloadedAutoSubsOnly`), in which case the audio is kept for that runner. // // WHAT IT DOWNLOADS IS AUDIO, on every channel. On a `handling: "youtube"` // channel — which normally only fetches subtitles — the re-acquire applies a diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,9 @@ # Changelog ## [Unreleased] +- **Re-acquired audio is handed to auto-transcribe instead of being deleted under it.** On a channel that only downloads subtitles, the audio the speaker lane fetches sits next to YouTube's own auto-captions — which is precisely the shape the *replace auto-captions* transcription runner looks for, and the channel's report is rebuilt about a second after any work on it finishes. Nothing coordinates the two, so the runner could start on a file the speaker lane was about to delete: the transcription then fails, and the video is written to the channel's permanently-honoured failed list. The lane now checks, at the moment it would delete, whether the transcription policy would take this video — and if it would, leaves the audio for it, saying so per video and in the run's summary line. That audio then behaves like any other download: it stays until you run *Clean audio from transcribed*. Audio is still deleted immediately in every other case, still kept for a video marked *do not clean*, and the hand-off is refused when free disk is below the mark the download runner itself would need — the lane never keeps a file the runner would have refused to fetch. +- **A video whose audio disappears mid-transcription is skipped, not marked failed.** Losing the input file is a media problem that fixes itself on the next attempt; recording it as a failed transcription blacklisted the video for the channel's manual whisper run for good. Both the local engine and a delegated remote one now report it as "no audio", which the queue already treats as "try again later". +- **Speaker names go stale when the transcript they were read from is replaced.** A name is a claim about a particular text; replacing a video's auto-captions with our own transcript changes the wording, the timings and sometimes who is in it, and nothing recorded which text the names came from — so the old names kept reading as current. New speaker records note it and are re-made when it changes. Records already on disk are untouched: they never carried the field, and inventing an answer for them would re-run speaker attribution across the whole archive. - **Re-acquiring media for the speaker lane now downloads audio, on every channel.** On a channel set to *download subtitles* rather than transcribe, the re-acquire ran that channel's own download — `--skip-download --write-subs --write-auto-subs` — so it re-fetched the captions the video already had, landed nothing the diarizer could read, and moved on. On this archive that spent about **16,000 fetches across eight channels for zero speaker records**. It now applies a per-video *transcribe* override, exactly as the "replace auto-captions" download already did: the channel's stored setting is not touched, and the fetched audio is still deleted the moment the lane has finished with it. - **A re-acquire no longer leaves the video looking un-transcribed.** Every fetch rewrites the video's metadata before it does anything else — even one that then fails — and that made the derived cue file read as out of date. The next operation in the very same run therefore skipped the video as having no transcript, and the digest lane deferred it, until someone pressed *Normalize transcripts* by hand. A re-acquire now re-normalizes the transcript straight after the fetch, on both the success and the failure path. It costs nothing when nothing moved, and it invalidates no digest. - **A video longer than the diarization length limit is deferred before a download is spent on it.** With re-acquire switched on, "the media is gone" meant "fetch it" — and the length limit was only consulted afterwards, by which point the file was already on disk. The limit is read first now, and only in the case that can actually spend a download. On this install the limit is switched off, so no count changes. diff --git a/editor/app/channels/[slug]/components/stages/SpeakersStage.tsx b/editor/app/channels/[slug]/components/stages/SpeakersStage.tsx @@ -163,7 +163,7 @@ export function SpeakersStage({ {missingInput === 1 ? "video needs" : "videos need"} their media re-acquired first {allowRedownload - ? " — re-download is on, so this run will fetch audio (even on a subtitle-only channel) and then delete it, bounded by the free-disk floor." + ? " — re-download is on, so this run will fetch audio (even on a subtitle-only channel) and then delete it — or keep it for auto-transcribe when that policy would replace the auto-captions — bounded by the free-disk floor." : " — re-download is off, so this run skips them."} </> ), diff --git a/editor/app/operations/components/settings/LaneSettingsForm.tsx b/editor/app/operations/components/settings/LaneSettingsForm.tsx @@ -108,8 +108,11 @@ export function LaneSettingsForm({ have media on disk and ~76,270 would need re-downloading — 91× the reachable work, against 45 GB free. With this on, each file is fetched, used, and <strong>deleted again immediately</strong>{" "} - (unless the video is marked &quot;do not clean&quot;), and nothing - starts at all when free space is under the disk floor. What it + (unless the video is marked &quot;do not clean&quot;, or the + auto-transcribe policy would replace that video&apos;s + auto-captions — then the audio is kept for that runner instead), + and nothing starts at all when free space is under the disk floor. + What it fetches is <strong>audio</strong>, including on channels that normally only download subtitles — those use a per-video transcribe override, and their stored config is not changed. diff --git a/editor/e2e/backfill.spec.ts b/editor/e2e/backfill.spec.ts @@ -841,7 +841,41 @@ const SEEDED_VTT = "WEBVTT\n\n00:00:00.000 --> 00:00:05.000\nSeeded caption line one.\n\n" + "00:00:05.000 --> 00:00:10.000\nSeeded caption line two.\n"; -async function seedSubtitleChannel(videoId: string): Promise<void> { +// Shaped after real yt-dlp --write-auto-subs output (a copy of +// auto-subs-replace.spec.ts's). It matters WHICH vtt a case seeds: the hand-off +// asks isAutoSubsOnly, whose 4 KB sniff is what decides ASR vs manual, and +// SEEDED_VTT above carries no ASR fingerprints at all — provenance then falls +// back to metadata.info.json, which the fake yt-dlp's re-acquire rewrites +// WITHOUT automatic_captions. +const ASR_VTT = `WEBVTT +Kind: captions +Language: en + +00:00:00.030 --> 00:00:03.919 align:start position:0% +so<00:00:00.719> today<00:00:01.199> we're<00:00:01.439> going<00:00:01.680> to + +00:00:03.919 --> 00:00:03.929 align:start position:0% +so today we're going to + +00:00:03.929 --> 00:00:07.070 align:start position:0% +so today we're going to +talk<00:00:04.320> about<00:00:04.639> the<00:00:04.879> whole<00:00:05.199> thing +`; + +// Audio this channel's video actually has on disk right now. +async function ytAudioFiles(videoId: string): Promise<string[]> { + const entries = await readdir( + resolvePath(`${YT_ROOT}/data/${videoId}`), + ).catch(() => [] as string[]); + return entries.filter( + (e) => e.startsWith("audio.") && !e.endsWith(".info.json"), + ); +} + +async function seedSubtitleChannel( + videoId: string, + vtt: string = SEEDED_VTT, +): Promise<void> { const dir = resolvePath(`${YT_ROOT}/data/${videoId}`); await mkdir(dir, { recursive: true }); await writeFile( @@ -858,7 +892,7 @@ async function seedSubtitleChannel(videoId: string): Promise<void> { resolvePath(`${YT_ROOT}/playlist`), `https://www.youtube.com/watch?v=${videoId}\n`, ); - await writeFile(`${dir}/transcript.en.vtt`, SEEDED_VTT); + await writeFile(`${dir}/transcript.en.vtt`, vtt); await writeFile( `${dir}/metadata.info.json`, JSON.stringify({ @@ -905,17 +939,7 @@ test("re-acquiring on a subtitle channel downloads audio and re-normalizes the c // …and the audio still does not survive the item that fetched it (case (4)'s // property, on a channel that never had audio to begin with). await expect - .poll( - async () => { - const entries = await readdir( - resolvePath(`${YT_ROOT}/data/${VID}`), - ).catch(() => [] as string[]); - return entries.filter( - (e) => e.startsWith("audio.") && !e.endsWith(".info.json"), - ); - }, - { timeout: 30_000 }, - ) + .poll(async () => ytAudioFiles(VID), { timeout: 30_000 }) .toEqual([]); // NO SUBTITLE RE-FETCH. The captions on disk are byte-identical to what was @@ -944,6 +968,127 @@ test("re-acquiring on a subtitle channel downloads audio and re-normalizes the c expect(cuesStat.mtimeMs).toBeGreaterThanOrEqual(metaStat.mtimeMs); }); +// (12) THE HAND-OFF. The audio case (11) deletes is exactly what +// autoQueue.transcription draws from `downloadedAutoSubsOnly` (ASR VTT AND audio +// present), and the snapshot regenerates ~1 s after any unit on the channel +// finishes — with no per-video lock anywhere. Deleting it under a running +// whisper costs that video's run and blacklists it in failed-transcriptions. So +// when the policy WOULD take it, the backfill hands it over instead. +// +// The runner is deliberately NOT started here: startAutoRunnersIfEnabled runs +// only at boot (editor/instrumentation.ts), and the invalidate-cache route +// starts nothing, so the policy is READ but nothing races the assertions. Once +// the audio is in that bucket the rest of the lane is what +// auto-subs-replace.spec.ts already walks end to end. +test("re-acquired audio is handed to auto-transcribe when the policy would replace the auto-captions", async ({ + page, +}) => { + test.setTimeout(SLOW); + const VID = "ytsubs000002"; + await resetData(); + await writeSettings({ + ...backfillSettings({ backfill: { allowRedownload: true } }), + autoQueue: { + transcription: { + enabled: true, + maxWorkers: 1, + replaceAutoSubs: true, + root: { + id: "root", + mode: "strict", + children: [{ id: "leaf-all", match: { type: "all" } }], + }, + }, + download: {}, + }, + }); + // ASR-shaped captions: the 4 KB sniff is what makes this video ASR-only, and + // therefore a candidate for the bucket. + await seedSubtitleChannel(VID, ASR_VTT); + + await generateReport(page, YT_SLUG); + await page.goto(channelStage(YT_SLUG, "speakers")); + await page + .getByRole("button", { name: "Run speaker work", exact: true }) + .click(); + + // The diarization still happens — the hand-off changes what happens to the + // audio AFTERWARDS, not whether the backfill does its work. + await expect + .poll(async () => pathExists(ytRel(VID, "diarization.json")), { + timeout: 60_000, + }) + .toBe(true); + + // Said so, per video and in the run's one-line summary. An unexplained file on + // a full disk is how a leak gets discovered the hard way, so "kept" must never + // be silent. + await expect(page.getByLabel("Run speaker work output")).toContainText( + "handed to auto-transcribe", + { timeout: 30_000 }, + ); + await expect(page.getByLabel("Run speaker work output")).toContainText( + "0 cleaned up, 1 handed to auto-transcribe", + { timeout: 30_000 }, + ); + + // The file itself survives the item that fetched it — the ONE case where that + // is correct. + expect((await ytAudioFiles(VID)).length).toBeGreaterThan(0); + + // …and it lands where the runner will find it. The snapshot regenerates at + // job end (operationJobs.ts), so this is the hand-off actually completing + // rather than a file left lying around. + await expect + .poll( + async () => { + const snap = await readJson<{ + buckets?: { downloadedAutoSubsOnly?: string[] }; + }>(`${YT_ROOT}/snapshot.json`).catch(() => null); + return snap?.buckets?.downloadedAutoSubsOnly ?? []; + }, + { timeout: 60_000 }, + ) + .toContain(VID); +}); + +// (12b) …and with no such policy, today's behaviour, through the same new code +// path. The keep decision has one "yes" and six "no"s; this is the "no" that +// every non-opted-in corpus gets. +test("re-acquired audio is still removed when no auto-transcribe policy would take it", async ({ + page, +}) => { + test.setTimeout(SLOW); + const VID = "ytsubs000003"; + await resetData(); + await writeSettings( + backfillSettings({ backfill: { allowRedownload: true } }), + ); + await seedSubtitleChannel(VID, ASR_VTT); + + await generateReport(page, YT_SLUG); + await page.goto(channelStage(YT_SLUG, "speakers")); + await page + .getByRole("button", { name: "Run speaker work", exact: true }) + .click(); + + await expect + .poll(async () => pathExists(ytRel(VID, "diarization.json")), { + timeout: 60_000, + }) + .toBe(true); + await expect(page.getByLabel("Run speaker work output")).toContainText( + "Removed re-acquired media", + { timeout: 30_000 }, + ); + expect( + await page.getByLabel("Run speaker work output").textContent(), + ).not.toContain("handed to"); + await expect + .poll(async () => ytAudioFiles(VID), { timeout: 30_000 }) + .toEqual([]); +}); + // (N) THE REGRESSION THIS STEP'S DESIGN EXISTS TO PREVENT. // // The channel snapshot now carries a work-list entry for EVERY catalog diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -2530,3 +2530,85 @@ override already existed twice: And the precedent for re-normalizing right after the thing that changed the inputs is `transcribeOne.ts:236-243`. + +--- + +## Verified 2026-08-27 — the re-acquire / auto-transcribe hand-off + +Read-only against `da339a4`. This section is the trail behind the hand-off shipped in +`e450c2c`; it is also the answer to STATE.md runbook step 4. + +### There is no per-video lock anywhere + +`common/jobs/registry.ts:78-80` — the registry serializes on `queueKey` alone. Two lanes +holding different queue keys can touch the same video dir at the same moment, and the +backfill (`BACKFILL_QUEUE`) and the transcription runner (`TRANSCRIPTION_QUEUE`) do exactly +that, by design. + +### The collision, with the trail + +1. `backfillBatch.ts:684-693` — the item's `finally` calls `reacquired.cleanup()`, which + unlinks the audio the re-acquire just fetched. +2. `channelSnapshot.ts:997` — `downloadedAutoSubsOnly` is `asrVtt && files.audioFiles.length + > 0`. A re-acquired `audio.mp3` next to YouTube auto-captions puts the video in it. +3. `autoRunner.ts:871` — a snapshot regen is requested on every unit completion, 1 s + debounce. (The BACKFILL regenerates once, at job end, `operationJobs.ts:138`.) +4. So `autoQueue.transcription` can start on the file between (2) and (1). +5. `scripts/parakeet-stitch.mjs:350` — parakeet re-opens the audio once per 480 s window, so + the unlink surfaces mid-run as a `sliceWav` failure. whisper.cpp opens once, so a short + video may survive; parakeet on a long one will not. +6. `transcribeOneFromQueue.ts:164-166` — anything that is not a `TranscribeError` with + `failureClass === "no-audio"` (`:156-163`) appends the id to `failed-transcriptions`. +7. `whisperBatch.ts:101-107` — the manual per-channel batch honours that list PERMANENTLY. + The auto runner passes no `failedSet` and would retry. + +### Two facts that shaped the fix rather than a lock + +- **Nothing automatic deletes audio after a transcription.** `cleanAudioFromTranscribed` is + a manual channel button (`whisperActions.ts:415`). So audio kept for the runner behaves + like any other download: `transcribedWithAudio` after transcription, waiting for the + operator's Clean-audio sweep. The steady-state cost is bounded by the existing UI and the + disk floor. +- **`replaceAutoSubs: false` is not a refusal.** A leaf whose `match.bucket` is + `downloadedAutoSubsOnly` draws it regardless (`autoQueuePolicy.ts:147-149, :180-184`). The + correct question is "does any leaf covering this channel draw that bucket", where a + bucket-less leaf draws `defaultBucketsForPolicy` — the only place the flag enters. That is + `policyDrawsBucket`. + +### The disk bar is the resume mark, not the floor + +`autoRunner.ts:663-668` — the download runner resumes only at `resumeBytes` (floor + +margin). A hand-off is a download the backfill was about to give back, so keeping audio the +runner would have refused to fetch would be incoherent. `decideKeep` calls +`evaluateDiskGate({ ...disk, latched: true }).ok` on the PURE core; `diskGate()` in enforce +mode mutates a shared module latch (`diskSpace.ts:231`) and must never be called from a keep +decision. + +### The remote branch wraps ENOENT as "transport" + +`remoteTranscribe.ts:111, :126-133`. A vanished-audio check at the local rethrow alone would +still let the remote path blacklist the video, which is why `throwIfAudioVanished` also +wraps `transcribeViaRemote`. + +### The `"failed"` cleanup was dead code + +`backfillReacquire.ts:196-208` has always returned a real partial-file cleanup for the +`"failed"` outcome, and the batch's `finally` only ever called `cleanup()` for `"fetched"`. +It runs now (`fetched || failed`). + +### The sweep has no inter-download sleep, and it is channel-major + +`backfillSweep.ts:55, 335-368`. STATE.md's runbook step 3 previously said "sleep 30 s between +downloads → ≥ 23 days"; `sleepBetweenDownloadsSeconds` is the CHANNEL sync path's setting and +the backfill sweep does not consult it. The sweep walks channel by channel, each channel's +candidates in the configured order. + +### Follow-up, deliberately not fixed: the digest lane is equally blind + +`DigestProvenance` records no transcript source, and `contextHash` is +`hashDigestContext(note)` over the channel's `digest-context.md` +(`digestContext-server.ts:47-60`) — nothing about the transcript. So an ASR→whisper +replacement leaves every digest section looking fresh over rewritten text, exactly as +attribution did before `f661677`. Same field, same record-carries-it rule, ~10 lines. NOT +done here because it would re-queue a local LLM digest for every replaced transcript, which +is a cost decision for the operator, not a correctness one. diff --git a/plans/STATE.md b/plans/STATE.md @@ -3,7 +3,15 @@ The working memory for the local-AI derived-corpus work. Rewritten at the end of every session, before context is cleared. See [`README.md`](README.md) for the protocol. -**Last updated:** 2026-08-27 — **the `deferred` cause is found and fixed** (`fed4b01`, +**Last updated:** 2026-08-28 — **the re-acquire hand-off shipped** (`e450c2c`, `f661677`, +and the e2e/docs commit after them): re-acquired audio on a subtitle channel is no longer +deleted under a transcription that the auto-queue policy would have started on it. It is +handed over when `autoQueue.transcription` would draw the video from +`downloadedAutoSubsOnly`, kept while a transcription task is running on it, and removed +otherwise; audio that vanishes mid-transcription is now a skip rather than a permanent +`failed-transcriptions` entry; and speaker attribution records which transcript the names +were made from, so an ASR→whisper replacement goes stale (for records written from now on). +Runbook item 2 **step 4 is resolved** — see below. Previously: 2026-08-27 — **the `deferred` cause is found and fixed** (`fed4b01`, `ca42f7b`): a backfill re-acquire on a `handling: "youtube"` channel was running the channel's own subtitle download, so it re-fetched captions, landed no audio, and left ~16,000 videos' `transcript.cues.json` stale by mtime. It now forces a per-video transcribe override and @@ -173,17 +181,26 @@ nothing renders. digests are not invalidated (`isSectionFresh` is provenance-keyed). Expect corpus `deferred` → ~76, the `missing`/`no-meta` tail. 3. **Re-affirm `allowRedownload` knowing what it now does.** With the fix the sweep will - *really* fetch audio for the **66,540** youtube-handling `missingInput` videos in - `newest` order corpus-wide (sleep 30 s between downloads → ≥ 23 days of network before - any diarization time), deleting each file after use. Scope it with `sweepChannels`, or + *really* fetch audio for the **66,540** youtube-handling `missingInput` videos + corpus-wide, deleting each file after use (or handing it to auto-transcribe — step 4). + The walk is **channel-major, in the configured order, with no inter-download sleep**: + `sleepBetweenDownloadsSeconds` is the channel SYNC path's setting and the backfill + sweep does not consult it (`backfillSweep.ts:55, 335-368`), so the earlier "30 s + between downloads → ≥ 23 days" figure here was wrong. Scope it with `sweepChannels`, or turn the flag off, before arming. That is a policy decision, not a code one. - 4. **A known interaction, flagged not fixed:** `autoQueue.transcription` is enabled with + 4. ~~**A known interaction, flagged not fixed**~~ — **RESOLVED 2026-08-28 by `e450c2c`.** + The interaction was real: `autoQueue.transcription` is enabled with `replaceAutoSubs: true` and picks from snapshots that regenerate ~1 s after a download - finishes. A re-acquired `audio.mp3` on an ASR-only video is exactly what it looks for, - so during a long diarization it may start whisper over the auto-captions; the - backfill's `finally` then deletes the audio under it (a failed unit, nothing lost) — or - whisper wins first and the video gets a better transcript. Worth deciding before a - corpus-wide run. + finishes, so a re-acquired `audio.mp3` on an ASR-only video is exactly what it looks + for — and the backfill's `finally` would delete it mid-whisper, which is not "a failed + unit, nothing lost" but a video appended to `failed-transcriptions`, honoured + permanently by the manual per-channel batch. Now the backfill hands the file over + instead (`decideKeep`), vetoes cleanup while a transcribe task is running on the video, + and `transcribeOne` treats vanished audio as a `no-audio` SKIP on both the local and the + remote path. **The operator no longer needs to turn `replaceAutoSubs` off for the corpus + run**; the trade is that handed-off audio persists like any other download until a + Clean-audio sweep. Mechanics and the file:line trail: FACTS.md "Verified 2026-08-27 — + the re-acquire / auto-transcribe hand-off". 3. **Editor IA slice 4** (`/actionable` dissolves) — **UNBLOCKED**: it wanted the per-state split (`blocked`, `deferred`, `partial`) and the coverage pair (`eligible`, `present`), which is what `snapshot.backfill.digest` carries and the deleted bucket did not.