Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 347ef052527ce17f7dba6c4785307c7047d37c66
parent e9ee8bd7c214999483e735383e7d882ed94f086c
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon, 10 Aug 2026 09:50:22 -0400

Record what the deferred bucket actually was, and that it is cleared

The measured split (1,942 missing vs 47 stale) contradicts what both docs
recorded, and the post-normalize numbers are measured with the real
digest.state() rather than predicted: deferred 1,987 -> 1, reachable
75,406 -> 77,392, piratesoftware ~114 -> 1,797.

Also records the omnimirror answer: its 712 are downloaded-but-never-
transcribed, so they are correctly blocked and normalize does nothing for
them. The earlier per-channel list had conflated the two populations.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Diffstat:
Mplans/FACTS.md | 40+++++++++++++++++++++++++++++++++++++++-
Mplans/STATE.md | 18++++++++++++++++--
2 files changed, 55 insertions(+), 3 deletions(-)

diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -1978,13 +1978,51 @@ binary, and why the scope comparison at `backfillSweep.ts:382` was repeatedly mi | --- | --- | | video dirs / `eligible` for digest | 78,963 / 78,885 (so exactly **78** untranscribable) | | `blocked` (no transcript) | **1,631** | -| `deferred` (stale `cues.json`) | **1,933**, of which `piratesoftware` alone is **1,683** | +| `deferred` (no current `cues.json`) | **1,933**, of which `piratesoftware` alone is **1,683** — see the correction below: these are overwhelmingly MISSING, not stale | | new reachable | **75,199** | | old `noDigest` (54 channels that have the bucket) | 75,613 | | `present` (digested at the current identity) | 122 | | snapshots with a `backfill` block | 44 of 66 | | channels predating the `noDigest` bucket | 11 (~1,516 videos) | +**CORRECTED 2026-08-10 — `deferred` was NEVER mostly "stale cues.json".** A per-video walk of +all 79,219 dirs split the bucket by reason for the first time: + +| reason | videos | +| --- | --- | +| **no `transcript.cues.json` at all** | **1,942** (`piratesoftware` 1,683 · `jfg-tonight` 151 · `angryjoeshow` 44 · `destiny` 30 · `omnibased` 14 · `jeremy-hambly` 8 · `the-quartering` 6 · `elissa-clips` 4 · 2 more) | +| genuinely superseded (stale) | **47**, all `shondo-vods` | + +So the superseded case the branch was written for is **2.4%** of what it caught, and the claim +that it "resolves itself" was false: `transcribeOne.ts:238` is the ONLY automatic caller of +`normalizeTranscript`, and a `handling: "youtube"` channel downloads subtitles with +`--skip-download` (`downloadOneManaged.ts:133-146`) so it never runs. It went unseen because +`buildIndex.ts:636-679` treats `cues.json` as a CACHE and silently re-parses the raw VTT — the +published site is correct, and only the digest lane, which has no fallback, can see it. + +**RESOLVED the same day.** `normalizeChannelTranscripts` was added and run over the 11 affected +channels: **1,991 sidecars written, 0 failures**, ~1.9 min total (57 s of it `piratesoftware`). +Re-measured with the REAL `digest.state()` over the whole corpus afterwards: + +| measure | before | after | +| --- | --- | --- | +| digest `reachable` | 75,406 (live post-regen) | **77,392** | +| `deferred` | 1,987 | **1** | +| `blocked` | 1,626 | **1,626** (unchanged — correctly) | +| `present` / `not-applicable` | — | 122 / 78 | +| `piratesoftware` reachable | ~114 | **1,797** (of 1,801; 4 blocked) | + +The single remaining `deferred` is a live-corpus artifact (a transcript rewritten mid-run). +A related one in `nerdrotic-live` has no `metadata.info.json` at all and so reports `no-meta` — +normalize cannot fix that one, and reporting it honestly is why `isCuesJsonFresh` now returns a +`reason` (`missing` / `stale` / `no-raw` / `no-meta`) alongside `fresh`. + +**NOT a separate bug: `omnimirror`'s 712.** Checked because it is `handling: "transcribe"`, so +`--skip-download` could not explain it. All 712 have `audio.mp3`, `metadata.info.json` and NO +transcript of any kind — downloaded, never transcribed. They classify `blocked` correctly and +normalizing would do nothing. Same for `RagingGoldenEagle` (659). The earlier per-channel list +had conflated untranscribed dirs with unnormalized ones. + **The plan predicted the count would RISE by ~1,530; it FALLS by 414.** `deferred` was underestimated. `digest.state()` costs **0.34 ms/video** (125-video channel) to **0.38 ms/video** (773-video channel), i.e. 0.7x–1.2x the `readVideoFiles` the snapshot already pays. diff --git a/plans/STATE.md b/plans/STATE.md @@ -29,12 +29,26 @@ count to RISE by ~1,530. It FALLS by 414, and the decomposition closes exactly: | --- | --- | | old `noDigest` total (54 channels that have the bucket) | 75,613 | | + the 11 pre-bucket channels, now counted | +1,516 | -| − newly `deferred` (stale `cues.json`) | −1,933 | +| − newly `deferred` (no current `cues.json`) | −1,933 | | = new reachable | **75,199** | `deferred` is far larger than the plan assumed and it is CONCENTRATED: **`piratesoftware` alone has 1,683**, so 94% of that channel's 1,797-video apparent digest backlog was work -`digestVideo` would have taken and immediately put back. Everything else the plan stated held: +`digestVideo` would have taken and immediately put back. + +**FOLLOW-UP, 2026-08-10 — the `deferred` bucket is now CLEARED, and the reason for it was +wrong.** Splitting it by cause showed 1,942 videos with NO `cues.json` at all against 47 +genuinely superseded ones, so "the raw transcript changed underneath, and the normalize pass +clears it" described 2.4% of the bucket and self-resolution described none of it: nothing +automatic ever runs `normalizeTranscript` for a channel that downloads its subtitles. Fixed in +`426ae51` — `isCuesJsonFresh` returns a `reason`, `normalizeAllTranscripts` takes +`channelSlugs`, `normalize-transcripts` is a registered job kind, and the Digest card carries a +**Normalize transcripts** button next to the count. Run over the 11 affected channels: **1,991 +written, 0 failed**; measured with the real `digest.state()` afterwards, corpus `deferred` +**1,987 → 1** and reachable **75,406 → 77,392**, with `piratesoftware` going **~114 → 1,797**. +`blocked` stayed at 1,626, unchanged, which is the point — those are waiting on transcription +and normalize was never going to touch them. Full split and the `omnimirror` finding in +FACTS.md. Everything else the plan stated held: `blocked` = **1,631** exactly (the untranscribed count), `eligible` = 78,885 of 78,963 videos, i.e. exactly **78** untranscribable. Coverage is **122** digested at the current identity.