commit 347ef052527ce17f7dba6c4785307c7047d37c66
parent e9ee8bd7c214999483e735383e7d882ed94f086c
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Mon, 10 Aug 2026 09:50:22 -0400
Record what the deferred bucket actually was, and that it is cleared
The measured split (1,942 missing vs 47 stale) contradicts what both docs
recorded, and the post-normalize numbers are measured with the real
digest.state() rather than predicted: deferred 1,987 -> 1, reachable
75,406 -> 77,392, piratesoftware ~114 -> 1,797.
Also records the omnimirror answer: its 712 are downloaded-but-never-
transcribed, so they are correctly blocked and normalize does nothing for
them. The earlier per-channel list had conflated the two populations.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Diffstat:
2 files changed, 55 insertions(+), 3 deletions(-)
diff --git a/plans/FACTS.md b/plans/FACTS.md
@@ -1978,13 +1978,51 @@ binary, and why the scope comparison at `backfillSweep.ts:382` was repeatedly mi
| --- | --- |
| video dirs / `eligible` for digest | 78,963 / 78,885 (so exactly **78** untranscribable) |
| `blocked` (no transcript) | **1,631** |
-| `deferred` (stale `cues.json`) | **1,933**, of which `piratesoftware` alone is **1,683** |
+| `deferred` (no current `cues.json`) | **1,933**, of which `piratesoftware` alone is **1,683** — see the correction below: these are overwhelmingly MISSING, not stale |
| new reachable | **75,199** |
| old `noDigest` (54 channels that have the bucket) | 75,613 |
| `present` (digested at the current identity) | 122 |
| snapshots with a `backfill` block | 44 of 66 |
| channels predating the `noDigest` bucket | 11 (~1,516 videos) |
+**CORRECTED 2026-08-10 — `deferred` was NEVER mostly "stale cues.json".** A per-video walk of
+all 79,219 dirs split the bucket by reason for the first time:
+
+| reason | videos |
+| --- | --- |
+| **no `transcript.cues.json` at all** | **1,942** (`piratesoftware` 1,683 · `jfg-tonight` 151 · `angryjoeshow` 44 · `destiny` 30 · `omnibased` 14 · `jeremy-hambly` 8 · `the-quartering` 6 · `elissa-clips` 4 · 2 more) |
+| genuinely superseded (stale) | **47**, all `shondo-vods` |
+
+So the superseded case the branch was written for is **2.4%** of what it caught, and the claim
+that it "resolves itself" was false: `transcribeOne.ts:238` is the ONLY automatic caller of
+`normalizeTranscript`, and a `handling: "youtube"` channel downloads subtitles with
+`--skip-download` (`downloadOneManaged.ts:133-146`) so it never runs. It went unseen because
+`buildIndex.ts:636-679` treats `cues.json` as a CACHE and silently re-parses the raw VTT — the
+published site is correct, and only the digest lane, which has no fallback, can see it.
+
+**RESOLVED the same day.** `normalizeChannelTranscripts` was added and run over the 11 affected
+channels: **1,991 sidecars written, 0 failures**, ~1.9 min total (57 s of it `piratesoftware`).
+Re-measured with the REAL `digest.state()` over the whole corpus afterwards:
+
+| measure | before | after |
+| --- | --- | --- |
+| digest `reachable` | 75,406 (live post-regen) | **77,392** |
+| `deferred` | 1,987 | **1** |
+| `blocked` | 1,626 | **1,626** (unchanged — correctly) |
+| `present` / `not-applicable` | — | 122 / 78 |
+| `piratesoftware` reachable | ~114 | **1,797** (of 1,801; 4 blocked) |
+
+The single remaining `deferred` is a live-corpus artifact (a transcript rewritten mid-run).
+A related one in `nerdrotic-live` has no `metadata.info.json` at all and so reports `no-meta` —
+normalize cannot fix that one, and reporting it honestly is why `isCuesJsonFresh` now returns a
+`reason` (`missing` / `stale` / `no-raw` / `no-meta`) alongside `fresh`.
+
+**NOT a separate bug: `omnimirror`'s 712.** Checked because it is `handling: "transcribe"`, so
+`--skip-download` could not explain it. All 712 have `audio.mp3`, `metadata.info.json` and NO
+transcript of any kind — downloaded, never transcribed. They classify `blocked` correctly and
+normalizing would do nothing. Same for `RagingGoldenEagle` (659). The earlier per-channel list
+had conflated untranscribed dirs with unnormalized ones.
+
**The plan predicted the count would RISE by ~1,530; it FALLS by 414.** `deferred` was
underestimated. `digest.state()` costs **0.34 ms/video** (125-video channel) to **0.38
ms/video** (773-video channel), i.e. 0.7x–1.2x the `readVideoFiles` the snapshot already pays.
diff --git a/plans/STATE.md b/plans/STATE.md
@@ -29,12 +29,26 @@ count to RISE by ~1,530. It FALLS by 414, and the decomposition closes exactly:
| --- | --- |
| old `noDigest` total (54 channels that have the bucket) | 75,613 |
| + the 11 pre-bucket channels, now counted | +1,516 |
-| − newly `deferred` (stale `cues.json`) | −1,933 |
+| − newly `deferred` (no current `cues.json`) | −1,933 |
| = new reachable | **75,199** |
`deferred` is far larger than the plan assumed and it is CONCENTRATED: **`piratesoftware`
alone has 1,683**, so 94% of that channel's 1,797-video apparent digest backlog was work
-`digestVideo` would have taken and immediately put back. Everything else the plan stated held:
+`digestVideo` would have taken and immediately put back.
+
+**FOLLOW-UP, 2026-08-10 — the `deferred` bucket is now CLEARED, and the reason for it was
+wrong.** Splitting it by cause showed 1,942 videos with NO `cues.json` at all against 47
+genuinely superseded ones, so "the raw transcript changed underneath, and the normalize pass
+clears it" described 2.4% of the bucket and self-resolution described none of it: nothing
+automatic ever runs `normalizeTranscript` for a channel that downloads its subtitles. Fixed in
+`426ae51` — `isCuesJsonFresh` returns a `reason`, `normalizeAllTranscripts` takes
+`channelSlugs`, `normalize-transcripts` is a registered job kind, and the Digest card carries a
+**Normalize transcripts** button next to the count. Run over the 11 affected channels: **1,991
+written, 0 failed**; measured with the real `digest.state()` afterwards, corpus `deferred`
+**1,987 → 1** and reachable **75,406 → 77,392**, with `piratesoftware` going **~114 → 1,797**.
+`blocked` stayed at 1,626, unchanged, which is the point — those are waiting on transcription
+and normalize was never going to touch them. Full split and the `omnimirror` finding in
+FACTS.md. Everything else the plan stated held:
`blocked` = **1,631** exactly (the untranscribed count), `eligible` = 78,885 of 78,963 videos,
i.e. exactly **78** untranscribable. Coverage is **122** digested at the current identity.