# The stats cache key (fix, 2026-09-28) Branch `fix/stats-cache-key` from `main` `ac438bbc`. **Merged `10cefd15` (2026-09-28). Rolled out with release 15** (the recount; the homepage at 77,547 transcripts on 2026-09-30, `release-15.md` "Rollout"). Verified in the tree 2026-10-09: the cache keys on `indexSignature` (`common/controller/buildStats.ts:119`), and `STATS_DOWNGRADE_ENV` (`:134`) guards a downgrade. ## What was wrong - **The stats cache was keyed on the wrong file.** The cache (`statsByPath`, in the index LMDB) was keyed on `metadata.info.json`'s mtime alone. But `hasTranscript`, `cueCount`, `coverage` and `transcribedDate` come from the index and the transcript files. - **So a late transcript never reached its stat.** That covers a Whisper run days after the download, a Normalize run weeks later, and a stats run made before `build:index` had the video. The hub and homepage composers make that last kind of run. - **Caption videos could not be dated at all.** A caption-only video (a YouTube VTT, no `transcript.json`, no outcome sidecar) got no `transcribedDate` even when computed fresh. - **The homepage then dropped every undated transcript.** Its fold required `hasTranscript && transcribedDate`. Jasolyzer served 1,889 videos and showed 0 transcripts, 0 channels, 0 hours. Every site's counts and "Transcribed over time" charts were low. ## Numbers, before and after (ESTIMATES) **How they were measured.** The whole-pool stats pages of 2026-09-28T20:40Z (`homepage/public/stats`) were classified record by record against the files on disk now. This was read-only: listings, `stat` and small JSON reads. No LMDB was opened and nothing was rebuilt. - **Shown** is the published homepage summary. - **After** counts the records a rebuild with this fix would count. These are: - `hasTranscript` with a date source on disk; - stale `hasTranscript: false` records whose `transcript.cues.json` has cues; - caption-only records, which the new date fallback reaches. The real numbers come from the first rebuild. | Site | Records | Transcripts shown | Transcripts after (est.) | | --- | ---: | ---: | ---: | | Anilyzer | 29,836 | 8,370 | ~29,665 | | Jeralyzer | 32,994 | 29,983 | ~31,906 (+ up to ~290, see below) | | Bonnellyzer | 8,085 | 5,540 | ~7,244 | | Hasanalyzer | 3,432 | 3,048 | ~3,335 | | Rekietalyzer | 2,931 | 2,855 | ~2,923 | | Jasolyzer | 1,889 | 0 | ~1,751 (~3,808 h) | | Whole pool | 79,385 | 49,798 | ~76,990 | - **Why so many were missing:** - 24,710 records had `hasTranscript` with no date. Of those, 23,198 have a date source on disk now, and 1,512 are caption-only. - 2,484 carried a stale `hasTranscript: false`. - **About 290 more are probably stale but are not counted above.** These are records with large caption files and `cueCount: null` (Jeralyzer's paramount-tactical, and the pool-only candace-owens). The live Jeralyzer shards serve cues for the three that were checked. - **The whole-pool total includes pool-only channels,** so it is more than the sum of the sites. ## The fix | Commit | What | | --- | --- | | `ad152529` | `common:` the key becomes the metadata mtime AND buildIndex's `mtimes.transcriptMs`. `transcribedDate` falls back through the outcome sidecar, then the transcript the index read (`transcript.json`, else the caption VTT), then `transcript.cues.json`, then `downloadedDate`, so a transcript always has one. `STATS_SCHEMA_VERSION` 5 → 6. New `buildStats.test.ts`. | | `b7a733ad` | `common:` test (b) asserts the heal before the new count. | | `9a8bded7` | `common:` the homepage fold counts a transcript with no date (totals, channels, hours, the card's total, upload-month placement). Only the transcribed series, "this month" and the recent rail need the date. | | `e765b168` | `mcp:` `get_video_metadata`'s stats block on a recomputed stat: the real cue count and no false "truncated". | | `6c7ff664` | `plans:` the first record, FACTS, changelogs. | | `9bc3c635` | `common:` buildIndex records `meta.scannedAt`, the time its last completed scan began. | | `e4c61c77` | `common:` the review's fixes (see "Review" below): the downgrade guard, the whole index record as the key with the cues read under its `indexKey`, "not indexed yet" and "not indexable" counted apart, an unmounted drive's stats kept (and a clear refused), and test hygiene. | | `9e61b119` | `common(test):` `source.test.ts` expands the `~` the kept-scratch log prints (release 12's test). | | `000273c0` | `common:` test (i) asserts the kept stats before the new result field. | | `22ec383e` | `plans:` the record's review, gates and rollout; FACTS; STATE; the three changelogs. | | `b30b52c1` | `common:` re-review R3 and R4: the "not indexable" line names an unreachable drive, and the held-channel refusal names every way out, without paths. Test (z) asserts the LMDB file is under the temp root. | | this commit | `plans:` re-review R1, R2, R4 (record), R5 and the nits: the rollout's commands as typed, the index's `Diff:` check, the hub's summary check, numbered preconditions, and the index follow-up in the record and STATE. | - **Caption videos are dated by when their captions arrived** (the VTT's mtime), not by a later Normalize run. Whisper videos resolve as before. Their "Transcribed over time" curves move: on Jasolyzer, 1,683 videos would otherwise all have landed on the Normalize day. - **The published stats page format did not change.** `STATS_MANIFEST_VERSION` stays 1, and `HOMEPAGE_SUMMARY_VERSION` stays 5. - **The compose paths warn and do not refuse.** `compose-hub` and `compose-homepage` still run `buildStats` against the index as it stands, and they still do not build the index. - A video they meet before the index has it is keyed `NOT_INDEXED`. It is recomputed on the first run after the next index build. - The log counts separately the videos downloaded since the last index build and the ones older than it that it did not index: no `upload_date`, a failure, or their channel's media unreachable during that build. - A refusal would stop these builds whenever the editor had downloaded since the last index build, which is almost always. - **Concurrent stats builds: an operator rule, not a lock** (review L3, brief item 5). - There is no cross-process lock primitive in `common/`. `scripts/queue-lock.mjs` is the e2e queue's flock wrapper, run as a separate holder process. Building one here would be a new lock, with its own stale-lock story. - The rule is: **one stats build at a time.** The editor's build jobs share the queue `"build"` by default. **The CLI is outside every queue.** - Two concurrent runs are harmless unless one clears the cache, which only a schema change does. ## Review **Verdict: SHIP AFTER FIXES** (the review is `j-review.md` in the job's scratch). The reviewer found no code defect; every fix was made on this branch. | Finding | Where | | --- | --- | | L1: the record overstated which paths run old code | `22ec383e`: the rollout says only the in-process **Build stats dataset** runs old code, and step 1 is "rebuild the editor bundle, then restart"; the editor changelog says the same. | | L2: wrong rollout commands | `22ec383e`: every command checked against `pnpm archilyzer --help` and `pnpm ops --help`. The homepage and the hub are each built and deployed ONE way, the hub before the sites, and `build-deploy` takes `{"all":true}`. | | L3: concurrent runs; the removable drive | `e4c61c77`: an unmounted drive's channel is held and its stats kept, and a cache clear with one refuses. Rollout step 2 is the precondition. The lock was **left**: see "Concurrent stats builds" above. | | L4: test hygiene | `e4c61c77`: every `getPaths()` path is pinned under the temp root, the root is removed in `after`, and case (z) spies on node:fs and node:fs/promises (async and sync) and fails on any write outside it. | | L5: the "not in the index yet" line was wrong for skipped videos | `9bc3c635` + `e4c61c77`: `notIndexedYet` and `notIndexable`, logged apart, with case (h). | | L6: the time estimate left out the USB drive | `22ec383e`: 10–30 minutes, with the reason; interruptible and resumes. | | L7: no homepage changelog bullet | `22ec383e`: `homepage/CHANGELOG.md`, and the export bullet says "once the site is rebuilt". | | The guard (ruled: add it) | `e4c61c77`: a build refuses to clear a cache a newer schema wrote, and names both versions and `ARCHILYZER_STATS_ALLOW_DOWNGRADE`. The variable is declared in `envVars.ts`; `ENVIRONMENT.md` is regenerated and `--check` is clean. Case (j) covers older → cleared, newer → refused (also the CLI, exit non-zero, cache untouched), and the override → cleared. | | O1 (ruled: do it) | `e4c61c77`: the cues are read under the `mtimes` record's `indexKey` (case (f)), and the key is the whole record (case (g)). About 40 lines with comments, and no new I/O on the unchanged path: the same one LMDB get per video. | | O2: date from the index's mtime | **Left.** The fresh readdir happens only on the recompute path, and it keeps the Whisper rule byte-identical to before. | | O3: the spy saw only async fs | `e4c61c77`: the spy now covers node:fs sync and callback APIs too. | | O4: the reviewer's extra cases | **Partly.** Case (e) now includes an index rebuild with no churn. The removed-transcript and two-channel cases stay in the reviewer's scratch, where they pass. | | O5: manual captions parse to 0 cues | **Left**, noted in FACTS. | | `source.test.ts` under a home `TMPDIR` | `9e61b119`: 13 of 13 with `TMPDIR` unset and with it under `~`. | **The new cases fail on the pre-review code** (`6c7ff664`'s `buildStats.ts`, `buildIndex.ts` and `stats.ts` swapped in once): | Case | Result | | --- | --- | | (b) | `undefined` for `notIndexedYet` (the field did not exist) | | (e) | `NaN` counts, for the same reason | | (f) | `hasTranscript` false: the cues missed under the metadata's upload date | | (g) | changed 0, not 1: the cue-count drift | | (h) | the fields did not exist | | (i) | removed 2, not 0: the drive's stats were dropped | | (j) | "Missing expected rejection": the newer cache was cleared | (a), (c), (d) and (z) pass there, as they should: (a) to (d) were fixed before the review. ### Re-review **Verdict: SHIP.** No code defect; five Low touch-ups and two nits, all made: | Finding | Where | | --- | --- | | R1: `archilyzer` is not on PATH | this commit: every rollout command is written as typed from the primary checkout's root, `pnpm archilyzer …`, as release 12's records spell it. Nothing else on the branch spells a bare command: the code's messages name none, and FACTS and the changelogs name scripts or pages. | | R2: step 3 had no post-check | this commit: the index step now reads its `Diff:` line. Thousands removed means a drive was missing; mount it and re-run the index before step 4. | | R3: "not indexable" blamed the video for a missing drive | `b30b52c1`: the line adds "or its channel's media was unreachable during that build (run an index build with every drive mounted)", and case (h) expects it. | | R4: the refusal gave one way out | `b30b52c1` (message) and this commit (record): mount its media first; else repair or re-point its location on /storage, finish or clear its move, or delete the channel or set `excludeFromBuild`. The message carries no path: a held channel is named with its location's label. Case (i) asserts both. | | R5: a hub build with no figures still deploys | this commit: step 6 builds the hub, checks the log for `hub-summary.json covers N official instance(s)`, and deploys only then. `hub-summary.json skipped: …` means stop. | | Nits | this commit numbers the preconditions; `b30b52c1` asserts `paths.lmdbPath` is under the temp root in case (z). | | Follow-up (recommended) | recorded under "Left" and in STATE: **the index build still treats an unmounted drive as an empty channel.** It needs its own slice before routine builds resume. Until then, R2's check is the safeguard. | ## Gates (worktree, 2026-09-28, at the review fixes) - **tsc:** clean (43 s). - **Common tests, with `TMPDIR` unset:** **2,161, all pass**. That is 2,149 at the branch point, plus 11 buildStats cases and 1 homepage fold case. - **Other unit suites:** - editor unit: 85 of 85; - `test:scripts`: 185 pass, 1 skipped; - mcp: 271 of 271; - homepage unit: 7 of 7. - **Docs:** `pnpm archilyzer docs env --check` is clean. - **Builds:** `next build` succeeded for export (36 s), editor (55 s) and homepage (23 s). - **e2e:** - Before the review (at `6c7ff664`): - editor `duplicate-shorts`, `build`, `site-scope`, `sites-homepage` and `deploy-page`: 21 passed, 0 failed, 2.3 min; - export `charts.spec.ts`: 8 passed, 28 s; - homepage, full suite: 36 passed, 1.0 min. - After the review fixes: the same editor list again, because `duplicate-shorts` drives Build index and Build stats dataset: 21 passed, 0 failed, 1.2 min. - Export and homepage were not rerun. The review fixes change no export or homepage code; the homepage fold is unchanged since `9a8bded7`. - **At the re-review touch-ups (`b30b52c1`):** - tsc is clean; - `buildStats.test.ts` and `envVars.test.ts` pass 20 of 20; - `pnpm archilyzer docs env --check` is clean; - no e2e was run, since only wording and one assertion changed. ## Rollout (operator) — follow it literally, in this order Every command below is typed **from the primary checkout's root**. There is no `archilyzer` on PATH, so it is `pnpm archilyzer …`. **What runs which code.** - **The live :3001 editor runs its BUILT bundle** until it is rebuilt and restarted. The only stats path that runs inside that bundle is the **Build stats dataset** button (`buildStatsAction`, in-process). On the old code it has no guard: against the new cache it would clear it and refill it the old way. - **Everything else spawns the checkout's code from disk**, so it runs the new code the moment `main` has this merge: - a site build's data phase (`pnpm run build:data`); - the hub (`compose:hub`) and the homepage (`compose`); - every CLI command. - So **the first of those after the merge is the first schema-6 stats run.** It clears the cache and does the whole pass inside that job. **Preconditions for steps 3 and 4.** 1. **The removable media drive is mounted.** `/storage` shows every location **Available**. - The stats build now refuses a cache clear while any channel's media is unreachable. - **The index build has no such guard.** Run with the drive absent, it drops those channels from the index, and the next site build publishes them as gone. Step 3's `Diff:` check is what catches it. 2. **No other index, stats or site build is running.** - `/jobs` shows no `build-index`, `build-stats`, `build-site`, `build-deploy`, `build-hub` or `build-homepage` job running or queued, on any queue. - No CLI or spawned build is running: `pgrep -af 'archilyzer\.ts (index|build|compose)'` prints nothing. Every CLI build and every spawned data phase or compose goes through `archilyzer.ts`; the in-process editor jobs do not show here, and `/jobs` covers them. **The steps.** 1. **Rebuild the editor bundle, then restart :3001 onto it:** `pnpm --filter editor build` in the primary checkout, then restart the editor the way it is normally run. Between the merge and this restart: - **never press Build stats dataset**; - **start no site, hub or homepage build**. It would do step 4's full pass itself, inside that job, unannounced. 2. **Check preconditions 1 and 2 above.** 3. **Index.** Use `pnpm archilyzer index`, or `/sites` → **Build index** (`pnpm ops build-index --wait`). It writes the same LMDB as the editor, so run it only with no build job running (precondition 2). **Then check its `Diff:` line**, `Diff: +A added, ~C changed, -R removed, N total.`: - with `pnpm archilyzer index` or `pnpm ops build-index --wait`, it is in the terminal output; - from the button, it is in the `build-index` job's log on `/jobs`. **R should be 0, or a handful.** Thousands removed means a drive was missing during the build, and those channels just left the index. Stop, mount the drive (precondition 1), and run step 3 again before step 4. 4. **Stats.** Use `pnpm archilyzer build stats`: the CLI, with the editor idle. The in-process button stalls the editor for the length of the pass. - The first run logs `Stats schema change (5 -> 6); clearing stats cache.` and re-extracts every video: **about 10–30 minutes, longer with a cold cache.** About a quarter of the video dirs are on the USB drive, at 4–5 random reads each. - It can be interrupted (Ctrl-C, or cancelling the job) and **resumes**: the schema is written at the clear, so the next run only finishes the rest. - **If it refuses because a channel cannot be read,** the message names the channel and its location. The ways out, in order: 1. mount its media and run step 4 again; 2. repair or re-point its location on `/storage`; 3. finish or clear its move, from the channel's Storage panel; 4. if the channel is gone for good, delete it or set `excludeFromBuild` in its config. - **Let it finish before step 5.** 5. **Homepage.** Use exactly ONE of: - `pnpm archilyzer build homepage && pnpm archilyzer deploy homepage`; - `pnpm ops build-homepage --json '{"deploy":true}' --wait`. Building the homepage **also runs release 12's source publish**. So the homepage waits on release 12's rollout step 0: the denylist is complete and `source publish --check` is clean. The homepage and source in `main` at that moment must be the ones the operator has judged. 6. **Hub, before the sites** (in basic mode the hub and the sites share `export/out`). Build it, check its log, and only then deploy it. Use exactly ONE pair: - `pnpm archilyzer build hub`, then `pnpm archilyzer deploy hub`; - `pnpm ops build-hub --wait`, then `pnpm ops deploy-hub --wait`. **Between the two,** the build's output (the compose line, `compose-hub: …`) must end with `hub-summary.json covers N official instance(s)`, where N is the number of public sites (6 today). **`hub-summary.json skipped: …` means the hub would deploy with no figures on its cards.** Stop and fix the cause it names before deploying. 7. **The six sites:** `pnpm ops build-deploy --json '{"all":true}' --wait`. **Live check.** - The homepage's Jasolyzer card shows about 1,751 transcripts, 1 channel and about 3,808 hours. - `https://jasolyzer.pages.dev/stats/page-0000.json` has no record with `hasTranscript: true` and `transcribedDate: null`. - Step 4's log has no `Channel …: … cached stat(s) are kept` line, which would mean a held channel. ## Left - **MCP "truncated" for no transcript at all.** A video with no transcript has coverage 0 (`transcriptCoverage(undefined, d > 0)`), so MCP `get_video_metadata` tells it "covers only 0% — truncated". This is older than the cache bug and out of scope. The coverage should be null when there are no cues. - **FOLLOW-UP, its own slice: the index build still treats an unmounted drive as an empty channel.** - What it does now: it drops that channel's index records, and the next site build publishes the channel as gone. - What it needs: `buildIndex` gets the same hold the stats build now has (keep the channel's `mtimes`, cues and pages), or at least a refusal with an override. - When: schedule it before routine builds resume after this rollout. - Until then, the rollout's step 3 `Diff:` check is the safeguard. - **Closed by release 15 slice IG** ([`release-15.md`](release-15.md), "Slice IG, as shipped"): the hold, and a refusal with an override for a full rebuild. - **A cross-process lock for builds** (see "Concurrent stats builds" above). The rule stands in for it. - **O5:** a manual English caption (no inline timing tags) indexes as 0 cues (FACTS).