Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 11d1a0ec63e054c852cb6ea0a2d2829309b15e4b
parent 8a34d340a63f625381acdcc8c58a118bbaa81667
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Thu, 27 Aug 2026 10:33:38 -0400

plans: the deferred cause is fixed and recorded

FACTS.md gains "Verified 2026-08-27 — why `deferred` was 16,156": the census
(16,081 stale / 75 missing / 1 no-meta over 78,128 transcribed dirs; 8,231
metadata-only vs 7,849 metadata+VTT; the eight channels with counts and the
08-22 → 08-26 dates), the mechanism with its file:line trail, the content check
(230 of 240 sampled VTTs byte-identical), why re-normalizing invalidates no
digest, the 66,540-video exposure had the sweep been re-armed unfixed, and the
two precedents for the transcribe override.

STATE.md is corrected in place rather than appended to. Both places that framed
this as accumulation-vs-regression said so on the strength of the same numbers
the census then explained, so both now say what it actually was, dated, with the
shas: "Recommended next" #2 becomes the four-step operator runbook (hold the
sweep, run Normalize over the eight channels, re-decide allowRedownload's scope
knowing it will now really fetch 66,540 videos' audio, and the flagged
auto-transcribe interaction), and the "THE FINDING TO ACT ON" paragraph keeps
its numbers under a dated correction block. "GPU yield on a quiet box" stays #1.

CHANGELOG [Unreleased] gets the three operator-facing lines.

Nothing under transcripts/ was read or written for this commit; the only corpus
figures restated are the ones already measured on 2026-08-26/27.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 3+++
Mplans/FACTS.md | 86+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mplans/STATE.md | 64+++++++++++++++++++++++++++++++++++++++++++++++-----------------
3 files changed, 136 insertions(+), 17 deletions(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,9 @@ # Changelog ## [Unreleased] +- **Re-acquiring media for the speaker lane now downloads audio, on every channel.** On a channel set to *download subtitles* rather than transcribe, the re-acquire ran that channel's own download — `--skip-download --write-subs --write-auto-subs` — so it re-fetched the captions the video already had, landed nothing the diarizer could read, and moved on. On this archive that spent about **16,000 fetches across eight channels for zero speaker records**. It now applies a per-video *transcribe* override, exactly as the "replace auto-captions" download already did: the channel's stored setting is not touched, and the fetched audio is still deleted the moment the lane has finished with it. +- **A re-acquire no longer leaves the video looking un-transcribed.** Every fetch rewrites the video's metadata before it does anything else — even one that then fails — and that made the derived cue file read as out of date. The next operation in the very same run therefore skipped the video as having no transcript, and the digest lane deferred it, until someone pressed *Normalize transcripts* by hand. A re-acquire now re-normalizes the transcript straight after the fetch, on both the success and the failure path. It costs nothing when nothing moved, and it invalidates no digest. +- **A video longer than the diarization length limit is deferred before a download is spent on it.** With re-acquire switched on, "the media is gone" meant "fetch it" — and the length limit was only consulted afterwards, by which point the file was already on disk. The limit is read first now, and only in the case that can actually spend a download. On this install the limit is switched off, so no count changes. - **"How many videos still need a digest" now has exactly one answer.** A channel's report used to carry two: a plain list of videos with no current digest, and the operation registry's own entry — the one the digest run itself dispatches from. They were built from different rules and disagreed by **11,777 videos** across this archive, because only the registry asks whether a video has a transcript at all and whether its normalized cue file is current. The plain list is gone from the report; every screen that shows a digest figure — the dashboard's *needs work* card, the `/channels` Digest column, a channel's Digest stage and its transit line, the pipelines band and the monitor widget — reads the registry's entry, which is what all of them were already doing. **No count on any page changes**, and nothing published changes: an exported site never reads a channel report. Reports written before this keep an unused key until the next time they are regenerated; nothing needs to be re-run. - **An operation's settings are on its own page now.** Digest, Diarization and Speaker attribution each moved off **Settings** and onto that operation's page under **Operations**, next to its backlog, its sweep and its pause — so switching an operation on and seeing what it would do are no longer two different screens. The **Speaker work lane** settings (whether the lane runs, its resource share, its concurrency, and re-acquiring deleted media) moved too, and appear on *all three* speaker operations' pages, because those operations genuinely share one queue: change it in one place and you have changed it for the others. Each block has its own **Save** button naming what it saves — *Save digest settings*, *Save lane settings*, and so on — so saving one can no longer touch another. Nothing on disk changed and no setting was renamed; the values you had are where you left them. **Settings** keeps the machine and the site: transcription workers, the sync scheduler, the build pipeline, social links and system paths, plus a line pointing at the four pages. - **The word "backfill" now means one thing: the shared queue.** It had been doing double duty — naming the queue that diarization, speaker attribution and the digest share, *and* standing in for each of those operations wherever a screen had no better word. The `/channels` group button that read **Backfill** is named after the operations it runs: **Speakers** once a speaker operation is switched on, **Derived data** while none is — the same derived label the channel page's stage already carried, so the button never claims work its lane is not doing. The `/actionable` section and every row of the channel page's Speakers card are labelled by the operations, not the queue, and the count column there says *reachable* rather than *to backfill*. *Pause backfill*, *Resume backfill* and *Start backfill sweep* keep their names, because those act on the queue itself. Under the hood the operation registry is `operations.ts` and its types say *Operation*, not *BackfillKind*; nothing stored on disk changed. diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -2444,3 +2444,89 @@ downloads have accumulated with no `cues.json` (nothing automatic runs `normaliz for a subtitle-downloading channel) or the cues gate regressed. It is exactly what the Digest card's **Normalize transcripts** button exists for, and it was invisible to a reader of the old bucket. + +--- + +## Verified 2026-08-27 — why `deferred` was 16,156 + +The 08-26 finding ("either accumulation or a regression") was **neither**. It was a single +sweep, in code, on a defect that is now fixed (`fed4b01`, `ca42f7b`). + +### The census + +Run read-only with the real `isCuesJsonFresh` over all **78,128** transcribed video dirs: + +| reason | count | +|---|---| +| `stale` (cues.json older than its inputs) | **16,081** | +| `missing` (no cues.json at all) | 75 | +| `no-meta` | 1 | + +16,081 + 75 = 16,156, which is the corpus-wide digest `deferred`. Of the stale ones, **7,849** +also had a re-fetched VTT and **8,231** were metadata-only rewrites. They sit on eight +`handling: "youtube"` channels, written in contiguous serial blocks **2026-08-22 → 08-26**: + +| channel | stale | +|---|---| +| destiny | 3,869 | +| the-quartering | 3,762 | +| chibi-reviews | 2,699 | +| nux-taku | 1,802 | +| leaflit | 1,612 | +| kirsche | 1,197 | +| quartering-live | 805 | +| HasanAbiVODs3 | 348 | + +Per channel, redownloads ≈ deferred ≈ a slice of that channel's `missingInput` population — +the signature of one ordered walk, not of accumulation. Each dir's `download-outcome.json` +records a `metadata-prefetch` + `primary` attempt with `handling: "youtube"`. + +### The mechanism, with the trail + +`settings.json` had `backfill.reach: corpus`, `order: newest`, `allowRedownload: true` and +diarization enabled. + +1. Every video on a subtitle-downloading channel is diarization **`missing-input`** — no + audio, by design of that handling (`common/lib/operations.ts`, the `hasDiarizableInput` + branch of the diarization `state()`). +2. `allowRedownload` turns `missing-input` into a **dispatch** (`backfillBatch.ts:194-196`). +3. `reacquireMediaFor` called `downloadOneManaged` with the channel's **unmodified** config + (`backfillReacquire.ts:116-133` before the fix — there was no `handling` check anywhere in + the file), so yt-dlp ran `--skip-download --write-subs --write-auto-subs`. +4. That rewrote `metadata.info.json` (and a VTT where YouTube served one), landed no audio, + logged "nothing usable landed", and moved to the next video at ~12/min. +5. `isCuesJsonFresh` compares `transcript.cues.json`'s mtime against metadata and the raw + transcript (`normalizeTranscript.ts:243`), so every touched video went stale — which reads + `deferred` to the digest lane and `no-transcript` → `skipped` to attribution + (`attributeOne.ts:139`, `operations.ts:983`). + +~16,000 wasted fetches (with firefox cookies), zero diarizations. + +### The content is fine + +230 of 240 sampled re-fetched VTTs parse to **byte-identical** cues; the 10 that differ do so +by a few cues of YouTube ASR drift. The 8,231 metadata-only cases never touched the transcript +at all. So the 16,081 are an **mtime false positive**, and one Normalize pass clears them. + +### Why re-normalizing is safe + +`isSectionFresh` (`digest.ts:513-534`) is **provenance-keyed** — app / model / prompt / +contextHash — not cues-mtime. Re-normalizing invalidates no digest. + +### The exposure, had it been re-armed unfixed + +**66,540** `handling: "youtube"` videos are diarization `missingInput` corpus-wide. Next in +`newest` order: chrissie-mayr 3,514, nuxanor 3,211, rev-says-desu 2,770, plus ~7,500 +unvisited on the-quartering. + +### The two precedents for the override + +The fix (`reacquireConfigFor`) is not novel; the same one-video `handling: "transcribe"` +override already existed twice: + +- `autoRunner.ts:996-1010` — the `replaceAutoSubs` unit, "the channel's stored config + untouched". +- `downloadOneManaged.ts:932-950` — its own no-subs fallback, same name, same shape. + +And the precedent for re-normalizing right after the thing that changed the inputs is +`transcribeOne.ts:236-243`. diff --git a/plans/STATE.md b/plans/STATE.md @@ -3,7 +3,12 @@ The working memory for the local-AI derived-corpus work. Rewritten at the end of every session, before context is cleared. See [`README.md`](README.md) for the protocol. -**Last updated:** 2026-08-26 (late) — **unified-ops step 1 landed**: `buckets.noDigest` is +**Last updated:** 2026-08-27 — **the `deferred` cause is found and fixed** (`fed4b01`, +`ca42f7b`): a backfill re-acquire on a `handling: "youtube"` channel was running the channel's +own subtitle download, so it re-fetched captions, landed no audio, and left ~16,000 videos' +`transcript.cues.json` stale by mtime. It now forces a per-video transcribe override and +re-normalizes after every fetch. The census is in FACTS.md; the operator's four remaining +steps are item 2 of "Recommended next". Previously: 2026-08-26 (late) — **unified-ops step 1 landed**: `buckets.noDigest` is deleted and `snapshot.backfill.digest` is the one digest work list. No rendered number moved (the fallback's migration was already complete on disk, 68/68), and **IA slice 4 is now unblocked**. One finding for the operator: corpus-wide `deferred` is **16,156**, against 1 on @@ -149,16 +154,36 @@ nothing renders. **Recommended next**, with unified-ops step 1 (the old #2) now shipped: 1. **GPU yield on a quiet box** — still the gate on arming the sweep, still unmeasured. -2. **Run Normalize transcripts over the 16,156 `deferred` videos** — NEW, and the one thing - this session found rather than built. Corpus `deferred` was **1** on 2026-08-10 after the - normalize pass and is **16,156** today (`destiny` 3,869, `the-quartering` 3,762, - `chibi-reviews` 2,699, `nux-taku` 1,802, `leaflit` 1,612, `kirsche` 1,197). Those videos - have a transcript and no current `cues.json`, so the digest lane will not touch them and - nothing automatic will fix it: `transcribeOne.ts` is the only automatic caller of - `normalizeTranscript` and a `handling: "youtube"` channel downloads subtitles with - `--skip-download`. **Decide whether this is accumulation or a regression before running - anything** — if it is accumulation, the missing piece is a normalize step on the - subtitle-download path, and re-running normalize by hand only defers the question. +2. **The operator runbook for the 16,156 `deferred`** — **CAUSE FOUND AND FIXED 2026-08-27** + (`fed4b01` the fix + units, `ca42f7b` the e2e and the wording). It was **neither + accumulation nor a regression in the cues gate**: it was one armed backfill sweep. With + `allowRedownload: true`, every video on a `handling: "youtube"` channel is diarization + `missing-input`, which dispatched a re-acquire that ran the channel's OWN + `--skip-download --write-subs --write-auto-subs` download — rewriting + `metadata.info.json`, landing no audio, and leaving `transcript.cues.json` stale by mtime. + ~16,000 videos on eight channels, 2026-08-22 → 08-26, zero diarizations. Census, trail and + numbers: FACTS.md "Verified 2026-08-27 — why `deferred` was 16,156". The fix forces a + per-video transcribe override and re-normalizes after every fetch. **Four operator steps + remain, and none was run from that slice:** + + 1. Keep `sweepEnabled: false` until `fed4b01` is on the running editor. (It is off now.) + 2. **Clear the 16,081**: Digest stage card → *Normalize transcripts* on the-quartering, + destiny, chibi-reviews, nux-taku, leaflit, kirsche, quartering-live, HasanAbiVODs3 (and + shondo-vods, 22). Pure re-parse, no network; identical cues in 230/240 sampled cases; + digests are not invalidated (`isSectionFresh` is provenance-keyed). Expect corpus + `deferred` → ~76, the `missing`/`no-meta` tail. + 3. **Re-affirm `allowRedownload` knowing what it now does.** With the fix the sweep will + *really* fetch audio for the **66,540** youtube-handling `missingInput` videos in + `newest` order corpus-wide (sleep 30 s between downloads → ≥ 23 days of network before + any diarization time), deleting each file after use. Scope it with `sweepChannels`, or + turn the flag off, before arming. That is a policy decision, not a code one. + 4. **A known interaction, flagged not fixed:** `autoQueue.transcription` is enabled with + `replaceAutoSubs: true` and picks from snapshots that regenerate ~1 s after a download + finishes. A re-acquired `audio.mp3` on an ASR-only video is exactly what it looks for, + so during a long diarization it may start whisper over the auto-captions; the + backfill's `finally` then deletes the audio under it (a failed unit, nothing lost) — or + whisper wins first and the video gets a better transcript. Worth deciding before a + corpus-wide run. 3. **Editor IA slice 4** (`/actionable` dissolves) — **UNBLOCKED**: it wanted the per-state split (`blocked`, `deferred`, `partial`) and the coverage pair (`eligible`, `present`), which is what `snapshot.backfill.digest` carries and the deleted bucket did not. @@ -210,12 +235,17 @@ measurements in [`FACTS.md`](FACTS.md) under "Verified 2026-08-26 — unified-op **THE FINDING TO ACT ON: corpus-wide `deferred` is 16,156.** It was **1** on 2026-08-10, right after the normalize pass. `destiny` 3,869, `the-quartering` 3,762, `chibi-reviews` 2,699, -`nux-taku` 1,802, `leaflit` 1,612, `kirsche` 1,197. Either `handling: "youtube"` downloads -have accumulated with no `cues.json` (nothing automatic runs `normalizeTranscript` for a -subtitle-downloading channel) or the cues gate regressed. That is 16,156 videos the digest -lane will not touch, and it was invisible to a reader of the old bucket. The fix is the Digest -card's **Normalize transcripts** button, run deliberately — deliberately NOT run from this -slice. +`nux-taku` 1,802, `leaflit` 1,612, `kirsche` 1,197. That is 16,156 videos the digest lane will +not touch, and it was invisible to a reader of the old bucket. + +> **Corrected 2026-08-27.** The "either accumulation or a regression" framing above was wrong +> on both counts, and it is left here only because the census that settled it started from +> these numbers. It was ONE ARMED BACKFILL SWEEP re-acquiring media on subtitle channels with +> the channel's own subtitle-download config: metadata rewritten, no audio landed, cues stale +> by mtime. Fixed in `fed4b01` / `ca42f7b`; the census and the file:line trail are in FACTS.md +> under "Verified 2026-08-27 — why `deferred` was 16,156", and the operator runbook is item 2 +> of "Recommended next" above. The Digest card's **Normalize transcripts** button is still +> what clears the 16,081 already on disk — deliberately NOT run from either slice. ---