commit b365aef422eeced8dc3947f80ef087691877b135
parent c2849a59576096830f6b52b49781e82258a6cbaf
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Mon, 28 Sep 2026 18:54:50 -0400
plans: the stats cache key — the review, the rewritten rollout, FACTS, STATE, changelogs
stats-cache-key.md: the commit table with the review's fixes, a "Review"
section (every finding to a sha or "left, why"), the proof the new cases fail
on the pre-review code, the gates, and the rollout rewritten to be followed
literally: which paths run old code until the restart (only the in-process
Build stats dataset), the preconditions (the drive mounted, no build running,
and how to check each), commands checked against the CLI's own usage, one way
each for the homepage and the hub, the hub before the sites, 10-30 minutes,
interruptible and resumable, and the homepage waiting on release 12's step 0.
FACTS: the section rewritten for the whole-record key, the guard, the two
counts, the drive rule, one stats build at a time. STATE: a "Now" entry.
Changelogs: the editor bullet corrected and a second one for the guards, the
export bullet "once the site is rebuilt", and a homepage bullet.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Diffstat:
6 files changed, 260 insertions(+), 123 deletions(-)
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -1,7 +1,8 @@
# Changelog
## [Unreleased]
-- **Transcripts that arrived after a video was first seen are counted.** The stats behind the homepage, the hub and every site's charts were cached per video and refreshed only when the video's metadata changed, so a transcript that came later — a Whisper run days after the download, or a video downloaded after the last index build — never reached them, and a video with YouTube captions alone had no transcription date. Counts and charts were low; the homepage could show a site with 0 transcripts, 0 channels and 0 hours while it served its videos. A stat is now also redone when the index's transcript for the video changes, every transcript has a date, and a captioned video is dated by when its captions arrived rather than by a later Normalize run, so its place on "Transcribed over time" can move. **Restart the editor on this version before the next stats build:** the first one re-reads every video once (6–12 minutes on a large archive), and an editor still on the old version would clear the new stats and build them the old way. Then build the index and the stats, the homepage and the hub, and the sites. A stats build that meets videos the index does not have yet says how many.
+- **Transcripts that arrived after a video was first seen are counted.** The stats behind the homepage, the hub and every site's charts were cached per video and refreshed only when the video's metadata changed, so a transcript that came later — a Whisper run days after the download, or a video downloaded after the last index build — never reached them, and a video with YouTube captions alone had no transcription date. Counts and charts were low; the homepage could show a site with 0 transcripts, 0 channels and 0 hours while it served its videos. A stat is now also redone whenever the index re-reads the video, every transcript has a date, and a captioned video is dated by when its captions arrived rather than by a later Normalize run, so its place on "Transcribed over time" can move. **After updating, rebuild and restart the editor before anything else:** until then, **Build stats dataset** runs the old code and would undo the new stats, while a site, hub or homepage build already runs the new code — and the first stats build of any kind re-reads every video once (about 10–30 minutes on a large archive; it can be stopped and picks up where it stopped). Then build the index, the stats, the homepage, the hub, and the sites.
+- **A stats build keeps the stats of a channel whose drive is not mounted, and will not undo a newer version's stats.** A channel whose media is on a drive that is not mounted (or is being moved) is left as it was instead of being read as a channel with no videos; a stats rebuild that has to start over refuses until the drive is back. A stats build refuses to clear stats written by a newer version of the editor; set `ARCHILYZER_STATS_ALLOW_DOWNGRADE=1` to roll back on purpose. Its log also says apart how many videos were downloaded since the last index build (they catch up after the next one) and how many the index skipped (no upload date, or it failed on them).
- **Building the homepage now publishes the source: a read-only git mirror, its raw tree and a fresh tarball, behind a gate.** `archilyzer build homepage`, the `/sites` Homepage jobs and `pnpm ops build-homepage` run `archilyzer source publish` between compose and `next build`. It makes a fresh clone of the private `main` (the repository itself is never rewritten), rewrites that copy with git-filter-repo using your scrub rules (file contents and commit messages; your home directory becomes `/home/user` without a rule), and publishes it under `homepage/public` for `git clone https://archilyzer.pages.dev/source/archilyzer.git`, beside `/source/tree/` and the Downloads tarball. Before anything is written, every object of the rewritten history and every file about to be published is searched for every string you have denied; **one hit refuses the build**, and its log names the string only by where you wrote it (`denylist line 3 (len 5)`) and each hit by its object, field and byte offset — never a byte of the object. **A refusal withdraws the source**: the last publish is removed from `homepage/public` and the last build's copy from `homepage/out`, and **Deploy homepage refuses** a build whose source was not audited under today's rules and today's `main` ("run `archilyzer build homepage`, then deploy"). The rules live outside the repo, in `~/.config/archilyzer/source-scrub.txt` and `source-denylist.txt` (`ARCHILYZER_CONFIG_DIR`, `SOURCE_SCRUB_FILE`, `SOURCE_DENYLIST_FILE`); **without them the build refuses**, naming the missing file. **Put everything private in the denylist before any deploy, a preview included**: previews are public, and every deployment stays reachable at its own address until you delete it. Install git-filter-repo once (`pipx install git-filter-repo`; the editor's process needs `~/.local/bin` on its `PATH` to find it) — without it the build fetches it through `pipx run`, which needs the network — and gitleaks if you want its secret scan too. An unchanged `main` with unchanged rules is skipped, so a rebuild costs about 20 seconds only when something moved. A checkout with no git repository (the docker image, a tarball install) builds with the /source page's empty state. `archilyzer source publish --check` audits without writing, `archilyzer source audit <clone>/.git` checks any clone, `archilyzer build homepage --no-source` removes the published source instead, and `archilyzer doctor` reports the tools, the two files (rule counts and permissions, never their contents) and the last publish. `create-archives.sh` is gone. See PUBLISH.md, "The source mirror (homepage)".
- **umtool reads the corpus from its checkout (or `TRANSCRIPTS_DIR`), and the song project's data defaults to `~/.local/share/archilyzer/song`.** If yours is elsewhere, link it there before restarting umtool: `mkdir -p ~/.local/share/archilyzer && ln -s <where the data is> ~/.local/share/archilyzer/song` (the data stays where it is). With no `CHANNELS_DIR`, umtool reads the corpus at `$TRANSCRIPTS_DIR/channels`, else the checkout's own `transcripts/channels`; it used to fall back to an absolute path that existed on one machine only. The song project's videos default to `~/reports/quartering-uh-song/videos`; `SONG_DIR` and `VIDEO_ROOT` still win. The song project's tracked manifests record their paths relative to the song folders, and the twenty one-off `umtool/song/*.sh` run logs, which only ever ran on the machine that wrote them, are gone.
diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md
@@ -1,7 +1,7 @@
# Changelog
## [Unreleased]
-- **The charts count every transcript.** A transcript that arrived after its video was first indexed was missing from the charts' transcript and cue counts and from "Transcribed over time", and a video with YouTube captions alone had no transcription date. Both are counted now, and a captioned video is dated by when its captions arrived.
+- **The charts count every transcript, once the site is rebuilt.** A transcript that arrived after its video was first indexed was missing from the charts' transcript and cue counts and from "Transcribed over time", and a video with YouTube captions alone had no transcription date. Both are counted now, and a captioned video is dated by when its captions arrived.
## [0.10.0] - 2026-09-28
- **A video whose recheck failed shows as possibly missing rather than available.** When a video drops out of its channel's listing it is marked "Missing?" until a recheck says why. A recheck that could not reach the video — a blocked request or a network error — used to clear the mark as if the video had been found. It now leaves "Missing?" in place until a recheck actually reaches the video. Needs a rebuild and deploy of every export site.
diff --git a/homepage/CHANGELOG.md b/homepage/CHANGELOG.md
@@ -2,6 +2,7 @@
## [Unreleased]
+- **A site's card counts every transcript, and never shows 0 channels while it serves recordings.** A transcript that arrived after its video was first indexed, or a video with YouTube captions alone, could be left out of the family's numbers: one site served 1,889 recordings and its card said 0 transcripts, 0 channels and 0 hours. Such transcripts are counted now — in the card, the family totals and the archive-growth chart — and one with no transcription date is left off only what is placed by that date: the charts by transcription date, "this month" and the recent list. The official-instance figures on the hub move with them.
- **The source is on the site, with its history: `/source/`.** A new **Source** page (and nav entry) gives `git clone https://archilyzer.pages.dev/source/archilyzer.git`, a read-only mirror of the main branch regenerated with every deploy, with its head, the private commit it reflects, a link to browse every file raw at `/source/tree/`, and the tarball with its size and sha256. Commit ids differ from the private repository's, because machine paths are scrubbed on the way out, and the page says so. A build without a published source says "No source published in this build." instead of offering a clone. The Downloads tarball is now regenerated by every build (its commit is the mirror's), and the page points at the mirror for history. The docs that said there is no public repository (*Install*, the FAQ, *What is Archilyzer*) now say how to clone. Below `md` the header's nav drops to its own row, as it did below `sm`, because five labels no longer fit beside the wordmark. `_headers` serves the raw tree as plain text.
- **The docs' *Building several sites at once* page says what Build all does.** It called the container pipeline opt-in, turned on in the settings. Build all sites builds every site in parallel in containers whenever a container engine is available, and one after another when none is; there is nothing to switch on.
- **A single-colour social icon shows on every ground.** The footer's social icons are the operator's (`homepage.json`'s, else `settings.socialLinks`), normalized when they are saved (`normalizeSocialSvg`, release 11 slice O1). An icon drawn in one colour now takes the footer's colour throughout; before, a part that carried its own colour kept it, so X's official logo, which is white, was invisible on the Light ground. An icon of two or more colours, such as YouTube's red mark with its white triangle, keeps its colours as pasted. "No fill", gradients, masks, clip paths and animation timing are never changed, and a clip path's own colour does not count, so a one-colour icon exported from Figma follows the footer too. It applies when the settings are next saved, then needs a rebuild and deploy of the homepage.
diff --git a/plans/FACTS.md b/plans/FACTS.md
@@ -5705,7 +5705,7 @@ Every `file:line` below was grepped at `7dfd7508`, whose code is byte-identical
- Export pages (buildIndex, buildStats, compose-homepage): compact, no newline.
- Chart, alias and tag stores: indented, no newline.
- Everything else: indented, with a newline.
- The six local `function writeJsonAtomic` left in `buildIndex.ts:404`, `buildStats.ts:282`,
+ The six local `function writeJsonAtomic` left in `buildIndex.ts:404`, `buildStats.ts:354`,
`bin/compose-homepage.ts:40`, `aliasesStore.ts:23`, `chartsStore.ts:50` and
`curatedTagsStore.ts:61` are one-line wrappers that pin those bytes over the shared writer.
They are not copies.
@@ -5810,7 +5810,7 @@ Every `file:line` below was grepped at `7dfd7508`, whose code is byte-identical
- **`readChannelConfigFile(file)`** `:165-170`. It never throws, and answers null for absent,
unreadable, not JSON, or not a channel. `readChannelConfig(paths, slug)` `:172` wraps it,
- and `buildIndex.ts:296` and `buildStats.ts:202` call it directly. The header `:152-157` names
+ and `buildIndex.ts:296` and `buildStats.ts:267` call it directly. The header `:152-157` names
the raw readers that bypass it: `channelMedia.ts:179` (the `dataDir` guard) and the legacy
migrations (`migrateToSites.ts:101` for `group`, `bin/migrate-channel-priority.ts` for
`excludeFromSync`, which also WRITES raw at `:206`).
@@ -5894,7 +5894,7 @@ which is the same race class 4b fixed for `config.json`.
Not in the 14:
-- The export build's streamed page writers `buildIndex.ts:971` and `buildStats.ts:264`. These
+- The export build's streamed page writers `buildIndex.ts:971` and `buildStats.ts:336`. These
are JSON writers still on the per-pid name, not in the list above only because there is one
writer per build. They are owed with the rest (16 + 2).
- `metadataScanStore.ts:188-206` and `autoQueueState.ts:141-156`. They carry a module-level
@@ -6142,7 +6142,7 @@ complete with this release.
- **The two write counters are deleted.** `git grep -n 'writeSeq\|nextWriteSeq' -- common
editor` is empty.
- **What `git grep -n 'tmp-${process.pid}' -- common editor` still finds:**
- - `buildIndex.ts:971` and `buildStats.ts:264`, the export page writers, out of scope;
+ - `buildIndex.ts:971` and `buildStats.ts:336`, the export page writers, out of scope;
- `transcode.ts:31` (ffmpeg's output, renamed at `:55`);
- `transcribeOne.ts:142` (the transcription app's `outputBase`);
- the shared writer's own comment `:9`, code `:136`, and test `:68`.
@@ -7322,66 +7322,101 @@ Slices Q (`4855f70b`) and R (`ffdeb2cd`): [`release-12.md`](release-12.md), the
## The stats cache key (verified 2026-09-28, branch `fix/stats-cache-key`)
-The record is [`stats-cache-key.md`](stats-cache-key.md). Anchors are at the branch tip.
-
-- **The stats cache is keyed on metadata AND the index's transcript record.**
- - `statsByPath` (`common/controller/buildStats.ts`) holds `{metaMs, idxMs, stat}` per
- `[channelSlug, videoDir]` (`:74`). A stat is recomputed when `metaMs` or `idxMs` moved (`:357`).
- - `idxMs` is buildIndex's `mtimes.transcriptMs` for the same key: the mtime of the transcript
- the index took the cues from (`pickIndexTranscript`, `videoStatus.ts:212`, picked at
- `buildIndex.ts:332`, written at `:835`). It is `null` when the video was indexed with no
- transcript, and `NOT_INDEXED` (−1, `:78`) when the index has no record.
- - It is read at `:320-327`: one LMDB get per video, and no file I/O on the unchanged path.
- - Why exactly this: buildIndex re-processes a video, and rewrites its `cues` (`:714`), when
- `transcriptMs` moves (`:579`). `hasTranscript`, `cueCount` and `coverage` are read from those cues.
+The record is [`stats-cache-key.md`](stats-cache-key.md). Anchors are at the branch tip. Two notes
+on anchors elsewhere in this file:
+- The branch added lines to `buildIndex.ts`: anchors above this section that point past `:66`
+ moved by +1, and past `:559` by +6. They are not rewritten in place.
+- The three `buildStats.ts` anchors in the slice-W sections were refreshed here.
+
+- **The stats cache is keyed on the metadata AND the index's own record for the video.**
+ - `statsByPath` (`common/controller/buildStats.ts`) holds `{metaMs, idx, stat}` per
+ `[channelSlug, videoDir]` (`:87`). A stat is recomputed when `metaMs` or `idx` moved (`:458`).
+ - `idx` is `indexSignature` (`:109`) of buildIndex's whole `mtimes` record: `metaMs`,
+ `transcriptMs`, `subsMs`, `availabilityMs`, `digestMs` (every input that makes buildIndex
+ re-process the video, `buildIndex.ts:585`) and `indexKey`. It is `NOT_INDEXED` (`"-"`,
+ `:102`) when the index has no record.
+ - The cues are read under the record's own `indexKey` (`:519`). The key buildIndex used when
+ `transcript.cues.json` is fresh comes from that file's `uploadDate`, which a later metadata
+ rewrite can differ from; the computed key is only the fallback for a video the index lacks.
+ - Cost on the unchanged path: one LMDB get per video, and no file I/O.
- **Until schema 6 the key was `metaMs` alone.** A transcript that arrived after a video was
first seen never reached its stat: Whisper days later, a Normalize run, or a stats run made
before `build:index` had the video. The pool composers make exactly that last kind of run:
`poolSummary.ts:76` runs `buildStats` with no index build. A later
`build:index && build:stats` then reported "added 0, changed 0" over those videos.
- - **Race, benign:** buildStats reads `idxMs` before the cues, and buildIndex writes the cues
- (`:714`) before `mtimes` (`:835`). A concurrent index build can only pair an older `idxMs`
- with newer cues, which the next run recomputes. It can never freeze a stale stat.
- - **Residual:** buildIndex does not re-index when only `transcript.cues.json` changes; its key has
- no cues.json mtime. When it does re-index for another reason (subs, availability, digest) it
- may read the fresher cues.json, so `cueCount` can drift without `idxMs` moving. `hasTranscript`
- and the date cannot drift that way.
- - Videos the index lacks are counted in `BuildStatsResult.unindexed` and logged as "N video(s)
- are not in the index yet …" (`:366`). The run does not refuse them: the editor downloads
- between index builds, so a refusal would stop the hub and homepage builds routinely.
+ - **Why the whole record, not just the transcript mtime:** buildIndex does not re-index when only
+ `transcript.cues.json` changes, but a re-index for another reason (subs, availability, digest)
+ reads the fresher cues.json and changes the cue count. Keying on the whole record redoes the
+ stat then. The cost is a recompute on those rarer changes.
+ - **Race, benign:** buildStats reads the record before the cues, and buildIndex writes the cues
+ (`buildIndex.ts:720`) before `mtimes` (`:841`). A concurrent index build can only pair an
+ older record with newer cues, which the next run recomputes.
+- **Two kinds of video have no index record, and they are counted apart.**
+ - buildIndex writes `meta.scannedAt` (`INDEX_SCANNED_AT_KEY`, `common/lib/stats.ts:19`) at
+ `buildIndex.ts:2006`, when a build completes. The value is the time its scan began (`:564`).
+ - **`notIndexedYet`:** metadata newer than `scannedAt`, or no build has completed yet. The video
+ was downloaded since, and it heals on the first stats run after the next index build.
+ - **`notIndexable`:** the last scan saw the video and did not index it. That means no
+ `upload_date` (`buildIndex.ts:701`) or a processing failure. It stays until fixed, and is
+ logged as such rather than as pending (`:471`).
- **A transcript always has a date, and a caption video takes its captions' arrival.**
- `resolveAcquisitionDates` (`:163`) tries, in order:
+ `resolveAcquisitionDates` (`:221`) tries, in order:
1. transcribe-outcome's `transcribedAt`;
- 2. the mtime of the picked index transcript, which is `transcript.json`, else the caption VTT (`:177`);
+ 2. the mtime of the picked index transcript, which is `transcript.json`, else the caption VTT (`:235`);
3. `transcript.cues.json`;
- 4. `downloadedDate` (`:181`).
+ 4. `downloadedDate` (`:239`).
Whisper videos resolve as before. The one difference is an outcome sidecar whose date will not
- parse: it now falls through to step 2 instead of giving null. A caption video used to get the
- Normalize run's date (cues.json's mtime) or none at all.
-- **`STATS_SCHEMA_VERSION` is 6** (`common/lib/stats.ts:11`). It versions the CACHE. The pages have
- their own version, `STATS_MANIFEST_VERSION`, which stays 1: the page shape did not change.
- - `buildStats` clears the cache on a mismatch.
- - `digestPlan.ts:395,447` and `duplicateShorts.ts:228` only warn on a mismatch. They read
- `value.stat` alone, so `idxMs` is invisible to them.
+ parse: it now falls through instead of giving null.
+- **The downgrade guard.** A build never clears a cache that a NEWER schema wrote (`:402`). It
+ throws, naming both versions and `ARCHILYZER_STATS_ALLOW_DOWNGRADE` (`:124`). That variable is
+ declared in `envVars.ts` for a deliberate rollback.
+ - The guard helps FUTURE bumps only. Schema-5 code has no guard, and would clear a schema-6 cache.
+ - `STATS_SCHEMA_VERSION` is 6 (`stats.ts:11`). It versions the cache. The pages' version,
+ `STATS_MANIFEST_VERSION`, stays 1.
+ - `digestPlan.ts:395,447` and `duplicateShorts.ts:228` only warn on a mismatch, and read
+ `value.stat` alone.
+- **An unmounted media drive is not an empty channel, for stats either.**
+ - `scanSource` asks `inspectChannelMedia` per channel (`:276`). A channel that is not `ok` or
+ `in-place` is "held": not rescanned, its cached stats kept and published (`:485`), and logged.
+ - A schema clear with any channel held REFUSES (`:421`), because the clear would drop that
+ channel's stats for good.
+ - This is the build's own guard. The job registry's `needsMedia` check
+ (`streamCommand.ts refuseForUnreachableMedia`) is per channel and needs a `channelSlug`, so it
+ never covered this pool-wide build, from the editor or from the CLI.
+ - **buildIndex has no such guard.** An index build with a drive unmounted drops those channels'
+ index records, and the site pages built from it lose them.
+- **One stats build at a time: an operator rule, not a lock.**
+ - Two concurrent runs are harmless unless one clears the cache (a schema change) after the other
+ has scanned. The other then collects a partly refilled `statsByPath` and publishes truncated
+ pages.
+ - There is no cross-process lock primitive in `common/`: `scripts/queue-lock.mjs` is the e2e
+ queue's flock wrapper.
+ - The editor's build jobs share the queue `"build"` by default. The per-button queue fields can
+ split them.
+ - **The CLI is outside every queue.**
- **The homepage fold counts a transcript that has no date** (`common/lib/homepageSummary.ts:407-408`).
- - It counts toward transcripts, channels and hours, the card's `transcribed.total`
- (`withUndated`, `:330`), and upload-month placement.
- - Only the transcribed series, `transcribedThisMonth` and `recent` need the date.
- - `HOMEPAGE_SUMMARY_VERSION` stays 5.
-- **MCP `get_video_metadata`** prints its "## Stats" block straight from the archive's stats pages
- (`mcp/src/server.ts:2273`), so its "covers only N% — truncated" note (`coverageNote`, `:2309`) is
- exactly as good as the stat. **Still owed:** a video with no transcript at all has coverage 0
- (`transcriptCoverage(undefined, d > 0)`), and gets that note too.
-- **A caption test fixture must carry YouTube's inline timing tags.** `parseVtt`
- (`common/lib/vtt.ts:30`) keeps only cue lines containing `<hh:mm:ss.mmm>` (`TIMING_TAG_RE`, `:3`).
- A plain `WEBVTT` cue parses to zero cues, so a hand-written VTT silently reads as "no
- transcript". `maybeMissingBuild.test.ts`'s VTT is such a file, which is harmless there.
+ It counts toward totals, channels, hours, the card's `transcribed.total` (`withUndated`) and
+ upload-month placement. Only the transcribed series, "this month" and the recent rail need the
+ date.
+- **MCP `get_video_metadata`** prints its "## Stats" block from the stats pages
+ (`mcp/src/server.ts:2273`). **Still owed:** a video with no transcript has coverage 0 and gets the
+ "covers only 0% — truncated" note (`coverageNote`, `:2309`).
+- **A caption test fixture must carry YouTube's inline timing tags.** `parseVtt` (`vtt.ts:30`) keeps
+ only cue lines containing `<hh:mm:ss.mmm>` (`TIMING_TAG_RE`, `:3`). A plain `WEBVTT` cue parses to
+ zero cues. `maybeMissingBuild.test.ts`'s VTT is such a file, which is harmless there.
+ - The same rule makes a video whose English track is a *manual* caption (no inline tags) index
+ as 0 cues. That is rare: 0 in 3,000 sampled of the-quartering, 6 of chibi-reviews.
- **Measured before the fix,** on the whole-pool stats of 2026-09-28T20:40Z:
- - 49,798 transcripts were shown, of about 77,000 on disk.
- - 24,710 records had `hasTranscript` but no `transcribedDate`.
- - 2,484 carried a stale `hasTranscript: false`.
- - Jasolyzer showed 0 of its 1,889.
-- **Cost of the schema bump on the real corpus:** one full re-extraction. That is 79,500 videos and
- 39.3 GB of `metadata.info.json` to read and parse (measured at 149 MB/s, so about 4.5 min), plus
- about 77,000 cue decodes and the sidecars: 6–12 minutes in all.
+ - 49,798 transcripts shown, of about 77,000 on disk;
+ - 24,710 records with `hasTranscript` but no `transcribedDate`;
+ - 2,484 with a stale `hasTranscript: false`;
+ - Jasolyzer 0 of 1,889.
+- **The schema bump's cost:** one full re-extraction. That is 79,500 videos and 39.3 GB of
+ metadata, measured at 149 MB/s on NVMe. About a quarter of the video dirs are on a USB drive, at
+ 4–5 random reads per recompute. Expect **about 10–30 minutes, longer with a cold cache**.
+ - The pass commits per batch of 200.
+ - It is interruptible, and it resumes: the schema is written at the clear, so the next run
+ finishes the rest.
+ - Run in the editor, it stalls the editor's event loop for the length of the pass. Prefer the CLI
+ with the editor idle.
diff --git a/plans/STATE.md b/plans/STATE.md
@@ -3,6 +3,24 @@
The working memory for the local-AI derived-corpus work. Rewritten at the end of every
session, before context is cleared. See [`README.md`](README.md) for the protocol.
+**Now (2026-09-28, night): the stats cache key fix — built, reviewed (SHIP AFTER FIXES, fixes
+done), not merged.** The branch is `fix/stats-cache-key`, and [`stats-cache-key.md`](stats-cache-key.md)
+holds the record, the review and the rollout. FACTS has "The stats cache key".
+- **What it fixes:** the homepage showed Jasolyzer as 0 transcripts, 0 channels, 0 hours while it
+ served 1,889 videos. Instance-wide it showed 49,798 transcripts of about 77,000.
+- **The cause:** `statsByPath` was keyed on the metadata mtime alone. It is now keyed on the index's
+ own record as well, and a transcript always has a date.
+- **Added by the review:**
+ - a guard against clearing a newer cache (`ARCHILYZER_STATS_ALLOW_DOWNGRADE`);
+ - an unmounted drive's stats are kept, and a cache clear with one refuses;
+ - "not indexed yet" and "not indexable" are counted apart.
+- **Owed after the merge:** the rollout in the record, in its order. First, rebuild and restart
+ :3001 (until then, never press "Build stats dataset"). Then index, then one full stats pass of
+ 10–30 min, then the homepage, the hub and the sites. The homepage deploy waits on release 12's
+ step 0: it runs the source publish.
+- **Merge note:** `homepage/social-visible` conflicts only in the changelogs' `[Unreleased]`. Keep
+ both sides.
+
**Now (2026-09-28, evening): release 12 — the source mirror — is merged to `main` and NOT rolled
out.** [`release-12.md`](release-12.md) holds Q's and R's records, their reviews, "Merged" and
"Rollout". The operator's runbook is `~/reports/release-12/RUNBOOK.html`, with its scripts in
diff --git a/plans/stats-cache-key.md b/plans/stats-cache-key.md
@@ -53,11 +53,16 @@ The real numbers come from the first rebuild.
| Commit | What |
| --- | --- |
-| `ad152529` | `common:` the key is the metadata mtime AND buildIndex's `mtimes.transcriptMs` (−1 when the video is not indexed yet), one LMDB get per video and no file I/O on the unchanged path. `transcribedDate` falls back through the outcome sidecar, then the transcript the index read (`transcript.json`, else the caption VTT), then `transcript.cues.json`, then `downloadedDate`, so a transcript always has one. `BuildStatsResult.unindexed` and a log line cover videos the index lacks. `STATS_SCHEMA_VERSION` 5 → 6. New `buildStats.test.ts`. |
-| `b7a733ad` | `common:` test (b) asserts the heal before the new `unindexed` count. |
+| `ad152529` | `common:` the key becomes the metadata mtime AND buildIndex's `mtimes.transcriptMs`. `transcribedDate` falls back through the outcome sidecar, then the transcript the index read (`transcript.json`, else the caption VTT), then `transcript.cues.json`, then `downloadedDate`, so a transcript always has one. `STATS_SCHEMA_VERSION` 5 → 6. New `buildStats.test.ts`. |
+| `b7a733ad` | `common:` test (b) asserts the heal before the new count. |
| `9a8bded7` | `common:` the homepage fold counts a transcript with no date (totals, channels, hours, the card's total, upload-month placement). Only the transcribed series, "this month" and the recent rail need the date. |
| `e765b168` | `mcp:` `get_video_metadata`'s stats block on a recomputed stat: the real cue count and no false "truncated". |
-| this commit | `plans:` this record, FACTS, changelogs. |
+| `6c7ff664` | `plans:` the first record, FACTS, changelogs. |
+| `9bc3c635` | `common:` buildIndex records `meta.scannedAt`, the time its last completed scan began. |
+| `e4c61c77` | `common:` the review's fixes (see "Review" below): the downgrade guard, the whole index record as the key with the cues read under its `indexKey`, "not indexed yet" and "not indexable" counted apart, an unmounted drive's stats kept (and a clear refused), and test hygiene. |
+| `9e61b119` | `common(test):` `source.test.ts` expands the `~` the kept-scratch log prints (release 12's test). |
+| `000273c0` | `common:` test (i) asserts the kept stats before the new result field. |
+| this commit | `plans:` the record's review, gates and rollout; FACTS; STATE; the three changelogs. |
- **Caption videos are dated by when their captions arrived** (the VTT's mtime), not by a later
Normalize run. Whisper videos resolve as before. Their "Transcribed over time" curves move: on
@@ -65,78 +70,155 @@ The real numbers come from the first rebuild.
- **The published stats page format did not change.** `STATS_MANIFEST_VERSION` stays 1, and
`HOMEPAGE_SUMMARY_VERSION` stays 5.
- **The compose paths warn and do not refuse.** `compose-hub` and `compose-homepage` still run
- `buildStats` against the index as it stands, and they still do not build the index. A video
- they meet before the index has it is keyed `NOT_INDEXED`. It is recomputed on the first run
- after the next index build, and the run's log says how many there are. A refusal would stop
- these builds whenever the editor had downloaded since the last index build, which is almost
- always.
-
-## Gates (worktree, 2026-09-28)
-
-- **tsc:** `pnpm -r --no-bail --workspace-concurrency=1 exec tsc --noEmit` is clean (36 s).
-- **Common tests:** 2,155 (2,149 at the branch point, plus 5 buildStats cases and 1 homepage fold
- case). 2,154 pass and 1 fails. The failure is the same test at the branch point:
- `source.test.ts`, "a denied literal no rule removes". It fails only because `TMPDIR` was set
- under the home directory, and the step prints that path with `~`.
+ `buildStats` against the index as it stands, and they still do not build the index.
+ - A video they meet before the index has it is keyed `NOT_INDEXED`. It is recomputed on the first
+ run after the next index build.
+ - The log counts separately the videos downloaded since the last index build and the ones that
+ build skipped (no `upload_date`, or a failure).
+ - A refusal would stop these builds whenever the editor had downloaded since the last index build,
+ which is almost always.
+- **Concurrent stats builds: an operator rule, not a lock** (review L3, brief item 5).
+ - There is no cross-process lock primitive in `common/`. `scripts/queue-lock.mjs` is the e2e
+ queue's flock wrapper, run as a separate holder process. Building one here would be a new
+ lock, with its own stale-lock story.
+ - The rule is: **one stats build at a time.** The editor's build jobs share the queue `"build"`
+ by default. **The CLI is outside every queue.**
+ - Two concurrent runs are harmless unless one clears the cache, which only a schema change does.
+
+## Review
+
+**Verdict: SHIP AFTER FIXES** (the review is `j-review.md` in the job's scratch). The reviewer found
+no code defect; every fix was made on this branch.
+
+| Finding | Where |
+| --- | --- |
+| L1: the record overstated which paths run old code | this commit: the rollout says only the in-process **Build stats dataset** runs old code, and step 1 is "rebuild the editor bundle, then restart"; the editor changelog says the same. |
+| L2: wrong rollout commands | this commit: every command checked against `archilyzer --help` and `pnpm ops --help`. The homepage and the hub are each built and deployed ONE way, the hub before the sites, and `build-deploy` takes `{"all":true}`. |
+| L3: concurrent runs; the removable drive | `e4c61c77`: an unmounted drive's channel is held and its stats kept, and a cache clear with one refuses. Rollout step 2 is the precondition. The lock was **left**: see "Concurrent stats builds" above. |
+| L4: test hygiene | `e4c61c77`: every `getPaths()` path is pinned under the temp root, the root is removed in `after`, and case (z) spies on node:fs and node:fs/promises (async and sync) and fails on any write outside it. |
+| L5: the "not in the index yet" line was wrong for skipped videos | `9bc3c635` + `e4c61c77`: `notIndexedYet` and `notIndexable`, logged apart, with case (h). |
+| L6: the time estimate left out the USB drive | this commit: 10–30 minutes, with the reason; interruptible and resumes. |
+| L7: no homepage changelog bullet | this commit: `homepage/CHANGELOG.md`, and the export bullet says "once the site is rebuilt". |
+| The guard (ruled: add it) | `e4c61c77`: a build refuses to clear a cache a newer schema wrote, and names both versions and `ARCHILYZER_STATS_ALLOW_DOWNGRADE`. The variable is declared in `envVars.ts`; `ENVIRONMENT.md` is regenerated and `--check` is clean. Case (j) covers older → cleared, newer → refused (also the CLI, exit non-zero, cache untouched), and the override → cleared. |
+| O1 (ruled: do it) | `e4c61c77`: the cues are read under the `mtimes` record's `indexKey` (case (f)), and the key is the whole record (case (g)). About 40 lines with comments, and no new I/O on the unchanged path: the same one LMDB get per video. |
+| O2: date from the index's mtime | **Left.** The fresh readdir happens only on the recompute path, and it keeps the Whisper rule byte-identical to before. |
+| O3: the spy saw only async fs | `e4c61c77`: the spy now covers node:fs sync and callback APIs too. |
+| O4: the reviewer's extra cases | **Partly.** Case (e) now includes an index rebuild with no churn. The removed-transcript and two-channel cases stay in the reviewer's scratch, where they pass. |
+| O5: manual captions parse to 0 cues | **Left**, noted in FACTS. |
+| `source.test.ts` under a home `TMPDIR` | `9e61b119`: 13 of 13 with `TMPDIR` unset and with it under `~`. |
+
+**The new cases fail on the pre-review code** (`6c7ff664`'s `buildStats.ts`, `buildIndex.ts` and
+`stats.ts` swapped in once):
+
+| Case | Result |
+| --- | --- |
+| (b) | `undefined` for `notIndexedYet` (the field did not exist) |
+| (e) | `NaN` counts, for the same reason |
+| (f) | `hasTranscript` false: the cues missed under the metadata's upload date |
+| (g) | changed 0, not 1: the cue-count drift |
+| (h) | the fields did not exist |
+| (i) | removed 2, not 0: the drive's stats were dropped |
+| (j) | "Missing expected rejection": the newer cache was cleared |
+
+(a), (c), (d) and (z) pass there, as they should: (a) to (d) were fixed before the review.
+
+## Gates (worktree, 2026-09-28, at the review fixes)
+
+- **tsc:** clean (43 s).
+- **Common tests, with `TMPDIR` unset:** **2,161, all pass**. That is 2,149 at the branch point,
+ plus 11 buildStats cases and 1 homepage fold case.
- **Other unit suites:**
- editor unit: 85 of 85;
- `test:scripts`: 185 pass, 1 skipped;
- - mcp: 271 of 271 (269 plus 2);
+ - mcp: 271 of 271;
- homepage unit: 7 of 7.
-- **Builds:** `next build` succeeded for export (24 s), editor (43 s) and homepage (16 s).
-- **e2e (queued, from the worktree root):**
- - editor `duplicate-shorts`, `build`, `site-scope`, `sites-homepage` and `deploy-page`: 21 passed,
- 0 failed, 2.3 min. `duplicate-shorts` drives Build index and Build stats dataset.
- - export `charts.spec.ts`: 8 passed, 0 failed, 28 s.
- - homepage, full suite: 36 passed, 0 failed, 1.0 min.
-- **The new tests fail against `main`'s code** (`main`'s `buildStats.ts`, `stats.ts` and
- `homepageSummary.ts` swapped in once):
- - (a): changed 0 !== 1. The late transcript never reaches the stat.
- - (b): changed 0 !== 1. No heal after the index build.
- - (c): the `vtt-only` date is null, not `20260711`.
- - (d): `after` is null. `with-outcome`, `no-outcome` and `hybrid` resolve to the same dates as on
- `main`, which proves the Whisper rule is unchanged.
- - (e): the heal redoes 0 stats, not 1. Its steady-state no-I/O assertions pin that the fix adds
- no reads; they would hold on `main` too.
- - Homepage fold: the undated site shows `[0, 0, 0]`, not `[1, 2, 2]`, for channels,
- transcripts and hours.
-
-## Rollout (operator), in this order
-
-1. **Restart the live editor onto the new build first.** An editor still on schema 5 that runs a
- stats build (Build stats dataset, a site build's data phase, a hub or homepage build) sees
- 6 ≠ 5. It clears the cache and recomputes everything with the old code. The two versions would
- then clear each other's cache, at a full pass each time.
-2. **Index.** Use `archilyzer index`, or `/sites` → **Build index** (`pnpm ops build-index`).
- Run it through the editor, or with the editor's build queue idle: `build:index` writes the same
- LMDB.
-3. **Stats.** Use `archilyzer build stats`, or `/sites` → **Build stats dataset**. The first run
- logs `Stats schema change (5 -> 6); clearing stats cache.` and re-extracts every video.
-4. **Homepage and hub.** Build and deploy both: `archilyzer build homepage` then the homepage
- deploy (`pnpm ops build-homepage` with `{"deploy": true}`), and `archilyzer build hub` then
- `pnpm ops deploy-hub`.
-5. **The six sites.** Build and deploy them so their `/stats` bundles and charts carry the new
- stats (`pnpm ops build-deploy`, or `archilyzer build all` and the deploy).
-
-**Cost:** one full pass in step 3. On the real corpus that is about 79,500 videos and 39.3 GB of
-`metadata.info.json`. Metadata was measured to read and parse at 149 MB/s, which is about 4.5 min.
-With about 77,000 cue decodes and the sidecars, expect **6–12 minutes**. Later runs are incremental
-again.
-
-**Live check:**
+- **Docs:** `archilyzer docs env --check` is clean.
+- **Builds:** `next build` succeeded for export (36 s), editor (55 s) and homepage (23 s).
+- **e2e:**
+ - Before the review (at `6c7ff664`):
+ - editor `duplicate-shorts`, `build`, `site-scope`, `sites-homepage` and `deploy-page`: 21
+ passed, 0 failed, 2.3 min;
+ - export `charts.spec.ts`: 8 passed, 28 s;
+ - homepage, full suite: 36 passed, 1.0 min.
+ - After the review fixes: the same editor list again, because `duplicate-shorts` drives Build
+ index and Build stats dataset: 21 passed, 0 failed, 1.2 min.
+ - Export and homepage were not rerun. The review fixes change no export or homepage code; the
+ homepage fold is unchanged since `9a8bded7`.
+
+## Rollout (operator) — follow it literally, in this order
+
+**What runs which code.**
+- **The live :3001 editor runs its BUILT bundle** until it is rebuilt and restarted. The only
+ stats path that runs inside that bundle is the **Build stats dataset** button (`buildStatsAction`,
+ in-process). On the old code it has no guard: against the new cache it would clear it and refill
+ it the old way.
+- **Everything else spawns the checkout's code from disk**, so it runs the new code the moment
+ `main` has this merge:
+ - a site build's data phase (`pnpm run build:data`);
+ - the hub (`compose:hub`) and the homepage (`compose`);
+ - every CLI command.
+- So **the first of those after the merge is the first schema-6 stats run.** It clears the cache
+ and does the whole pass inside that job.
+
+**Preconditions for steps 3 and 4.**
+- **The removable media drive is mounted.** `/storage` shows every location **Available**.
+ - The stats build now refuses a cache clear while any channel's media is unreachable.
+ - **The index build has no such guard.** Run with the drive absent, it drops those channels from
+ the index, and the next site build publishes them as gone.
+- **No other index, stats or site build is running.**
+ - `/jobs` shows no `build-index`, `build-stats`, `build-site`, `build-deploy`, `build-hub` or
+ `build-homepage` job running or queued, on any queue.
+ - No CLI or spawned build is running:
+ `pgrep -af 'archilyzer\.ts (index|build|compose)'` prints nothing. Every CLI build and every
+ spawned data phase or compose goes through `archilyzer.ts`; the in-process editor jobs do not
+ show here, and `/jobs` covers them.
+
+**The steps.**
+
+1. **Rebuild the editor bundle, then restart :3001 onto it:** `pnpm --filter editor build` in the
+ primary checkout, then restart the editor the way it is normally run. Between the merge and this
+ restart:
+ - **never press Build stats dataset**;
+ - **start no site, hub or homepage build**. It would do step 4's full pass itself, inside that
+ job, unannounced.
+2. **Check the preconditions above.**
+3. **Index.** Use `archilyzer index`, or `/sites` → **Build index** (`pnpm ops build-index --wait`).
+ It writes the same LMDB as the editor, so run it only with no build job running (precondition 2).
+4. **Stats.** Use `archilyzer build stats`: the CLI, with the editor idle. The in-process button
+ stalls the editor for the length of the pass.
+ - The first run logs `Stats schema change (5 -> 6); clearing stats cache.` and re-extracts every
+ video: **about 10–30 minutes, longer with a cold cache.** About a quarter of the video dirs
+ are on the USB drive, at 4–5 random reads each.
+ - It can be interrupted (Ctrl-C, or cancelling the job) and **resumes**: the schema is written at
+ the clear, so the next run only finishes the rest.
+ - **Let it finish before step 5.**
+5. **Homepage.** Use exactly ONE of:
+ - `archilyzer build homepage && archilyzer deploy homepage`;
+ - `pnpm ops build-homepage --json '{"deploy":true}' --wait`.
+
+ Building the homepage **also runs release 12's source publish**. So the homepage waits on
+ release 12's rollout step 0: the denylist is complete and `source publish --check` is clean.
+ The homepage and source in `main` at that moment must be the ones the operator has judged.
+6. **Hub, before the sites** (in basic mode the hub and the sites share `export/out`). Use exactly
+ ONE of:
+ - `archilyzer build hub && archilyzer deploy hub`;
+ - `pnpm ops build-hub --json '{"deploy":true}' --wait`.
+7. **The six sites:** `pnpm ops build-deploy --json '{"all":true}' --wait`.
+
+**Live check.**
- The homepage's Jasolyzer card shows about 1,751 transcripts, 1 channel and about 3,808 hours.
- `https://jasolyzer.pages.dev/stats/page-0000.json` has no record with `hasTranscript: true`
and `transcribedDate: null`.
+- Step 4's log names no held channel.
## Left
- **MCP "truncated" for no transcript at all.** A video with no transcript has coverage 0
(`transcriptCoverage(undefined, d > 0)`), so MCP `get_video_metadata` tells it "covers only
- 0% — truncated". This is older than the cache bug and out of scope. The coverage should be
- null when there are no cues.
-- **`cueCount` can drift.** It can move without the key moving when buildIndex re-indexes for an
- unrelated reason and reads a newer `transcript.cues.json` (FACTS, "The stats cache key").
- `hasTranscript` and the date are not affected.
-- **One metadata-key split is unmeasured.** When `transcript.cues.json` is fresh, buildIndex keys
- the cues by that file's `uploadDate`, while stats key them by the metadata's. They agree for
- all 1,865 of Jasolyzer's cues files. It was not measured corpus-wide.
+ 0% — truncated". This is older than the cache bug and out of scope. The coverage should be null
+ when there are no cues.
+- **buildIndex treats an unmounted drive as an empty channel** and drops its index records. The
+ stats build no longer does; the index build still needs the precondition.
+- **A cross-process lock for builds** (see "Concurrent stats builds" above). The rule stands in
+ for it.
+- **O5:** a manual English caption (no inline timing tags) indexes as 0 cues (FACTS).