Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 539d3a878e2825ea58587c2d2c8ca5e99e5b2a3d
parent a4f7145efbd053cfc8fc353eafc9b4674da516df
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon, 28 Sep 2026 18:09:56 -0400

plans: the stats cache key — record, FACTS, changelogs

plans/stats-cache-key.md: what was wrong, the per-site numbers before and
after (estimates, and how they were measured), the commits, the gates, the
proof each new test fails on main, and the operator's rollout in order
(restart the editor onto the new build first; one full stats pass, 6-12 min).
FACTS: a new "The stats cache key" section, the chunk census's "all" corrected,
and three buildStats.ts anchors moved by the fix. [Unreleased] bullets in
editor/CHANGELOG.md and export/CHANGELOG.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 1+
Mexport/CHANGELOG.md | 3+++
Mplans/FACTS.md | 78+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-----
Aplans/stats-cache-key.md | 142+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
4 files changed, 219 insertions(+), 5 deletions(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,7 @@ # Changelog ## [Unreleased] +- **Transcripts that arrived after a video was first seen are counted.** The stats behind the homepage, the hub and every site's charts were cached per video and refreshed only when the video's metadata changed, so a transcript that came later — a Whisper run days after the download, or a video downloaded after the last index build — never reached them, and a video with YouTube captions alone had no transcription date. Counts and charts were low; the homepage could show a site with 0 transcripts, 0 channels and 0 hours while it served its videos. A stat is now also redone when the index's transcript for the video changes, every transcript has a date, and a captioned video is dated by when its captions arrived rather than by a later Normalize run, so its place on "Transcribed over time" can move. **Restart the editor on this version before the next stats build:** the first one re-reads every video once (6–12 minutes on a large archive), and an editor still on the old version would clear the new stats and build them the old way. Then build the index and the stats, the homepage and the hub, and the sites. A stats build that meets videos the index does not have yet says how many. - **Building the homepage now publishes the source: a read-only git mirror, its raw tree and a fresh tarball, behind a gate.** `archilyzer build homepage`, the `/sites` Homepage jobs and `pnpm ops build-homepage` run `archilyzer source publish` between compose and `next build`. It makes a fresh clone of the private `main` (the repository itself is never rewritten), rewrites that copy with git-filter-repo using your scrub rules (file contents and commit messages; your home directory becomes `/home/user` without a rule), and publishes it under `homepage/public` for `git clone https://archilyzer.pages.dev/source/archilyzer.git`, beside `/source/tree/` and the Downloads tarball. Before anything is written, every object of the rewritten history and every file about to be published is searched for every string you have denied; **one hit refuses the build**, and its log names the string only by where you wrote it (`denylist line 3 (len 5)`) and each hit by its object, field and byte offset — never a byte of the object. **A refusal withdraws the source**: the last publish is removed from `homepage/public` and the last build's copy from `homepage/out`, and **Deploy homepage refuses** a build whose source was not audited under today's rules and today's `main` ("run `archilyzer build homepage`, then deploy"). The rules live outside the repo, in `~/.config/archilyzer/source-scrub.txt` and `source-denylist.txt` (`ARCHILYZER_CONFIG_DIR`, `SOURCE_SCRUB_FILE`, `SOURCE_DENYLIST_FILE`); **without them the build refuses**, naming the missing file. **Put everything private in the denylist before any deploy, a preview included**: previews are public, and every deployment stays reachable at its own address until you delete it. Install git-filter-repo once (`pipx install git-filter-repo`; the editor's process needs `~/.local/bin` on its `PATH` to find it) — without it the build fetches it through `pipx run`, which needs the network — and gitleaks if you want its secret scan too. An unchanged `main` with unchanged rules is skipped, so a rebuild costs about 20 seconds only when something moved. A checkout with no git repository (the docker image, a tarball install) builds with the /source page's empty state. `archilyzer source publish --check` audits without writing, `archilyzer source audit <clone>/.git` checks any clone, `archilyzer build homepage --no-source` removes the published source instead, and `archilyzer doctor` reports the tools, the two files (rule counts and permissions, never their contents) and the last publish. `create-archives.sh` is gone. See PUBLISH.md, "The source mirror (homepage)". - **umtool reads the corpus from its checkout (or `TRANSCRIPTS_DIR`), and the song project's data defaults to `~/.local/share/archilyzer/song`.** If yours is elsewhere, link it there before restarting umtool: `mkdir -p ~/.local/share/archilyzer && ln -s <where the data is> ~/.local/share/archilyzer/song` (the data stays where it is). With no `CHANNELS_DIR`, umtool reads the corpus at `$TRANSCRIPTS_DIR/channels`, else the checkout's own `transcripts/channels`; it used to fall back to an absolute path that existed on one machine only. The song project's videos default to `~/reports/quartering-uh-song/videos`; `SONG_DIR` and `VIDEO_ROOT` still win. The song project's tracked manifests record their paths relative to the song folders, and the twenty one-off `umtool/song/*.sh` run logs, which only ever ran on the machine that wrote them, are gone. diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md @@ -1,5 +1,8 @@ # Changelog +## [Unreleased] +- **The charts count every transcript.** A transcript that arrived after its video was first indexed was missing from the charts' transcript and cue counts and from "Transcribed over time", and a video with YouTube captions alone had no transcription date. Both are counted now, and a captioned video is dated by when its captions arrived. + ## [0.10.0] - 2026-09-28 - **A video whose recheck failed shows as possibly missing rather than available.** When a video drops out of its channel's listing it is marked "Missing?" until a recheck says why. A recheck that could not reach the video — a blocked request or a network error — used to clear the mark as if the video had been found. It now leaves "Missing?" in place until a recheck actually reaches the video. Needs a rebuild and deploy of every export site. - **The hub's Ask AI says when the hub has no archives.** On a hub whose list of archives is empty or could not be read, with none added in this browser, the question box read "Loading transcripts…" forever. It is now disabled, and one line under it says the hub has no archives yet, with a link to the front page, where one can be added. diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -655,7 +655,10 @@ chunk they agree to within ~2%. #### The chunk census Computed free from `statsByPath.cueCount` — present on **all 73,367** transcribed -videos, so it needed no GPU time and no transcript reads. At the shipped config +videos, so it needed no GPU time and no transcript reads. (Corrected 2026-09-28: "all" was +all the CACHE knew of. Until stats schema 6 a video transcribed after its stat was first +cached kept `cueCount: null` — about 2,500 of them on 2026-09-28; see "The stats cache key" +at the end of this file.) At the shipped config (`maxCues` 600, overlap 40, step 560): ``` @@ -5702,7 +5705,7 @@ Every `file:line` below was grepped at `7dfd7508`, whose code is byte-identical - Export pages (buildIndex, buildStats, compose-homepage): compact, no newline. - Chart, alias and tag stores: indented, no newline. - Everything else: indented, with a newline. - The six local `function writeJsonAtomic` left in `buildIndex.ts:404`, `buildStats.ts:238`, + The six local `function writeJsonAtomic` left in `buildIndex.ts:404`, `buildStats.ts:282`, `bin/compose-homepage.ts:40`, `aliasesStore.ts:23`, `chartsStore.ts:50` and `curatedTagsStore.ts:61` are one-line wrappers that pin those bytes over the shared writer. They are not copies. @@ -5807,7 +5810,7 @@ Every `file:line` below was grepped at `7dfd7508`, whose code is byte-identical - **`readChannelConfigFile(file)`** `:165-170`. It never throws, and answers null for absent, unreadable, not JSON, or not a channel. `readChannelConfig(paths, slug)` `:172` wraps it, - and `buildIndex.ts:296` and `buildStats.ts:158` call it directly. The header `:152-157` names + and `buildIndex.ts:296` and `buildStats.ts:202` call it directly. The header `:152-157` names the raw readers that bypass it: `channelMedia.ts:179` (the `dataDir` guard) and the legacy migrations (`migrateToSites.ts:101` for `group`, `bin/migrate-channel-priority.ts` for `excludeFromSync`, which also WRITES raw at `:206`). @@ -5891,7 +5894,7 @@ which is the same race class 4b fixed for `config.json`. Not in the 14: -- The export build's streamed page writers `buildIndex.ts:971` and `buildStats.ts:220`. These +- The export build's streamed page writers `buildIndex.ts:971` and `buildStats.ts:264`. These are JSON writers still on the per-pid name, not in the list above only because there is one writer per build. They are owed with the rest (16 + 2). - `metadataScanStore.ts:188-206` and `autoQueueState.ts:141-156`. They carry a module-level @@ -6139,7 +6142,7 @@ complete with this release. - **The two write counters are deleted.** `git grep -n 'writeSeq\|nextWriteSeq' -- common editor` is empty. - **What `git grep -n 'tmp-${process.pid}' -- common editor` still finds:** - - `buildIndex.ts:971` and `buildStats.ts:220`, the export page writers, out of scope; + - `buildIndex.ts:971` and `buildStats.ts:264`, the export page writers, out of scope; - `transcode.ts:31` (ffmpeg's output, renamed at `:55`); - `transcribeOne.ts:142` (the transcription app's `outputBase`); - the shared writer's own comment `:9`, code `:136`, and test `:68`. @@ -7317,3 +7320,68 @@ Slices Q (`4855f70b`) and R (`ffdeb2cd`): [`release-12.md`](release-12.md), the - **The live :3001 editor runs its BUILT bundle.** Until it is rebuilt on a tree with release 12, its `/sites` Homepage jobs have no source step, no withdrawal and no deploy check. +## The stats cache key (verified 2026-09-28, branch `fix/stats-cache-key`) + +The record is [`stats-cache-key.md`](stats-cache-key.md). Anchors are at the branch tip. + +- **The stats cache is keyed on metadata AND the index's transcript record.** + - `statsByPath` (`common/controller/buildStats.ts`) holds `{metaMs, idxMs, stat}` per + `[channelSlug, videoDir]` (`:74`). A stat is recomputed when `metaMs` or `idxMs` moved (`:357`). + - `idxMs` is buildIndex's `mtimes.transcriptMs` for the same key: the mtime of the transcript + the index took the cues from (`pickIndexTranscript`, `videoStatus.ts:212`, picked at + `buildIndex.ts:332`, written at `:835`). It is `null` when the video was indexed with no + transcript, and `NOT_INDEXED` (−1, `:78`) when the index has no record. + - It is read at `:320-327`: one LMDB get per video, and no file I/O on the unchanged path. + - Why exactly this: buildIndex re-processes a video, and rewrites its `cues` (`:714`), when + `transcriptMs` moves (`:579`). `hasTranscript`, `cueCount` and `coverage` are read from those cues. + - **Until schema 6 the key was `metaMs` alone.** A transcript that arrived after a video was + first seen never reached its stat: Whisper days later, a Normalize run, or a stats run made + before `build:index` had the video. The pool composers make exactly that last kind of run: + `poolSummary.ts:76` runs `buildStats` with no index build. A later + `build:index && build:stats` then reported "added 0, changed 0" over those videos. + - **Race, benign:** buildStats reads `idxMs` before the cues, and buildIndex writes the cues + (`:714`) before `mtimes` (`:835`). A concurrent index build can only pair an older `idxMs` + with newer cues, which the next run recomputes. It can never freeze a stale stat. + - **Residual:** buildIndex does not re-index when only `transcript.cues.json` changes; its key has + no cues.json mtime. When it does re-index for another reason (subs, availability, digest) it + may read the fresher cues.json, so `cueCount` can drift without `idxMs` moving. `hasTranscript` + and the date cannot drift that way. + - Videos the index lacks are counted in `BuildStatsResult.unindexed` and logged as "N video(s) + are not in the index yet …" (`:366`). The run does not refuse them: the editor downloads + between index builds, so a refusal would stop the hub and homepage builds routinely. +- **A transcript always has a date, and a caption video takes its captions' arrival.** + `resolveAcquisitionDates` (`:163`) tries, in order: + 1. transcribe-outcome's `transcribedAt`; + 2. the mtime of the picked index transcript, which is `transcript.json`, else the caption VTT (`:177`); + 3. `transcript.cues.json`; + 4. `downloadedDate` (`:181`). + + Whisper videos resolve as before. The one difference is an outcome sidecar whose date will not + parse: it now falls through to step 2 instead of giving null. A caption video used to get the + Normalize run's date (cues.json's mtime) or none at all. +- **`STATS_SCHEMA_VERSION` is 6** (`common/lib/stats.ts:11`). It versions the CACHE. The pages have + their own version, `STATS_MANIFEST_VERSION`, which stays 1: the page shape did not change. + - `buildStats` clears the cache on a mismatch. + - `digestPlan.ts:395,447` and `duplicateShorts.ts:228` only warn on a mismatch. They read + `value.stat` alone, so `idxMs` is invisible to them. +- **The homepage fold counts a transcript that has no date** (`common/lib/homepageSummary.ts:407-408`). + - It counts toward transcripts, channels and hours, the card's `transcribed.total` + (`withUndated`, `:330`), and upload-month placement. + - Only the transcribed series, `transcribedThisMonth` and `recent` need the date. + - `HOMEPAGE_SUMMARY_VERSION` stays 5. +- **MCP `get_video_metadata`** prints its "## Stats" block straight from the archive's stats pages + (`mcp/src/server.ts:2273`), so its "covers only N% — truncated" note (`coverageNote`, `:2309`) is + exactly as good as the stat. **Still owed:** a video with no transcript at all has coverage 0 + (`transcriptCoverage(undefined, d > 0)`), and gets that note too. +- **A caption test fixture must carry YouTube's inline timing tags.** `parseVtt` + (`common/lib/vtt.ts:30`) keeps only cue lines containing `<hh:mm:ss.mmm>` (`TIMING_TAG_RE`, `:3`). + A plain `WEBVTT` cue parses to zero cues, so a hand-written VTT silently reads as "no + transcript". `maybeMissingBuild.test.ts`'s VTT is such a file, which is harmless there. +- **Measured before the fix,** on the whole-pool stats of 2026-09-28T20:40Z: + - 49,798 transcripts were shown, of about 77,000 on disk. + - 24,710 records had `hasTranscript` but no `transcribedDate`. + - 2,484 carried a stale `hasTranscript: false`. + - Jasolyzer showed 0 of its 1,889. +- **Cost of the schema bump on the real corpus:** one full re-extraction. That is 79,500 videos and + 39.3 GB of `metadata.info.json` to read and parse (measured at 149 MB/s, so about 4.5 min), plus + about 77,000 cue decodes and the sidecars: 6–12 minutes in all. diff --git a/plans/stats-cache-key.md b/plans/stats-cache-key.md @@ -0,0 +1,142 @@ +# The stats cache key (fix, 2026-09-28) + +Branch `fix/stats-cache-key` from `main` `ac438bbc`. Merged: not yet. Rolled out: not yet. + +## What was wrong + +- **The stats cache was keyed on the wrong file.** The cache (`statsByPath`, in the index LMDB) + was keyed on `metadata.info.json`'s mtime alone. But `hasTranscript`, `cueCount`, `coverage` and + `transcribedDate` come from the index and the transcript files. +- **So a late transcript never reached its stat.** That covers a Whisper run days after the + download, a Normalize run weeks later, and a stats run made before `build:index` had the video. + The hub and homepage composers make that last kind of run. +- **Caption videos could not be dated at all.** A caption-only video (a YouTube VTT, no + `transcript.json`, no outcome sidecar) got no `transcribedDate` even when computed fresh. +- **The homepage then dropped every undated transcript.** Its fold required `hasTranscript && + transcribedDate`. Jasolyzer served 1,889 videos and showed 0 transcripts, 0 channels, 0 hours. + Every site's counts and "Transcribed over time" charts were low. + +## Numbers, before and after (ESTIMATES) + +**How they were measured.** The whole-pool stats pages of 2026-09-28T20:40Z (`homepage/public/stats`) +were classified record by record against the files on disk now. This was read-only: listings, +`stat` and small JSON reads. No LMDB was opened and nothing was rebuilt. + +- **Shown** is the published homepage summary. +- **After** counts the records a rebuild with this fix would count. These are: + - `hasTranscript` with a date source on disk; + - stale `hasTranscript: false` records whose `transcript.cues.json` has cues; + - caption-only records, which the new date fallback reaches. + +The real numbers come from the first rebuild. + +| Site | Records | Transcripts shown | Transcripts after (est.) | +| --- | ---: | ---: | ---: | +| Anilyzer | 29,836 | 8,370 | ~29,665 | +| Jeralyzer | 32,994 | 29,983 | ~31,906 (+ up to ~290, see below) | +| Bonnellyzer | 8,085 | 5,540 | ~7,244 | +| Hasanalyzer | 3,432 | 3,048 | ~3,335 | +| Rekietalyzer | 2,931 | 2,855 | ~2,923 | +| Jasolyzer | 1,889 | 0 | ~1,751 (~3,808 h) | +| Whole pool | 79,385 | 49,798 | ~76,990 | + +- **Why so many were missing:** + - 24,710 records had `hasTranscript` with no date. Of those, 23,198 have a date source on disk + now, and 1,512 are caption-only. + - 2,484 carried a stale `hasTranscript: false`. +- **About 290 more are probably stale but are not counted above.** These are records with large + caption files and `cueCount: null` (Jeralyzer's paramount-tactical, and the pool-only + candace-owens). The live Jeralyzer shards serve cues for the three that were checked. +- **The whole-pool total includes pool-only channels,** so it is more than the sum of the sites. + +## The fix + +| Commit | What | +| --- | --- | +| `ad152529` | `common:` the key is the metadata mtime AND buildIndex's `mtimes.transcriptMs` (−1 when the video is not indexed yet), one LMDB get per video and no file I/O on the unchanged path. `transcribedDate` falls back through the outcome sidecar, then the transcript the index read (`transcript.json`, else the caption VTT), then `transcript.cues.json`, then `downloadedDate`, so a transcript always has one. `BuildStatsResult.unindexed` and a log line cover videos the index lacks. `STATS_SCHEMA_VERSION` 5 → 6. New `buildStats.test.ts`. | +| `b7a733ad` | `common:` test (b) asserts the heal before the new `unindexed` count. | +| `9a8bded7` | `common:` the homepage fold counts a transcript with no date (totals, channels, hours, the card's total, upload-month placement). Only the transcribed series, "this month" and the recent rail need the date. | +| `e765b168` | `mcp:` `get_video_metadata`'s stats block on a recomputed stat: the real cue count and no false "truncated". | +| this commit | `plans:` this record, FACTS, changelogs. | + +- **Caption videos are dated by when their captions arrived** (the VTT's mtime), not by a later + Normalize run. Whisper videos resolve as before. Their "Transcribed over time" curves move: on + Jasolyzer, 1,683 videos would otherwise all have landed on the Normalize day. +- **The published stats page format did not change.** `STATS_MANIFEST_VERSION` stays 1, and + `HOMEPAGE_SUMMARY_VERSION` stays 5. +- **The compose paths warn and do not refuse.** `compose-hub` and `compose-homepage` still run + `buildStats` against the index as it stands, and they still do not build the index. A video + they meet before the index has it is keyed `NOT_INDEXED`. It is recomputed on the first run + after the next index build, and the run's log says how many there are. A refusal would stop + these builds whenever the editor had downloaded since the last index build, which is almost + always. + +## Gates (worktree, 2026-09-28) + +- **tsc:** `pnpm -r --no-bail --workspace-concurrency=1 exec tsc --noEmit` is clean (36 s). +- **Common tests:** 2,155 (2,149 at the branch point, plus 5 buildStats cases and 1 homepage fold + case). 2,154 pass and 1 fails. The failure is the same test at the branch point: + `source.test.ts`, "a denied literal no rule removes". It fails only because `TMPDIR` was set + under the home directory, and the step prints that path with `~`. +- **Other unit suites:** + - editor unit: 85 of 85; + - `test:scripts`: 185 pass, 1 skipped; + - mcp: 271 of 271 (269 plus 2); + - homepage unit: 7 of 7. +- **Builds:** `next build` succeeded for export (24 s), editor (43 s) and homepage (16 s). +- **e2e (queued, from the worktree root):** + - editor `duplicate-shorts`, `build`, `site-scope`, `sites-homepage` and `deploy-page`: 21 passed, + 0 failed, 2.3 min. `duplicate-shorts` drives Build index and Build stats dataset. + - export `charts.spec.ts`: 8 passed, 0 failed, 28 s. + - homepage, full suite: 36 passed, 0 failed, 1.0 min. +- **The new tests fail against `main`'s code** (`main`'s `buildStats.ts`, `stats.ts` and + `homepageSummary.ts` swapped in once): + - (a): changed 0 !== 1. The late transcript never reaches the stat. + - (b): changed 0 !== 1. No heal after the index build. + - (c): the `vtt-only` date is null, not `20260711`. + - (d): `after` is null. `with-outcome`, `no-outcome` and `hybrid` resolve to the same dates as on + `main`, which proves the Whisper rule is unchanged. + - (e): the heal redoes 0 stats, not 1. Its steady-state no-I/O assertions pin that the fix adds + no reads; they would hold on `main` too. + - Homepage fold: the undated site shows `[0, 0, 0]`, not `[1, 2, 2]`, for channels, + transcripts and hours. + +## Rollout (operator), in this order + +1. **Restart the live editor onto the new build first.** An editor still on schema 5 that runs a + stats build (Build stats dataset, a site build's data phase, a hub or homepage build) sees + 6 ≠ 5. It clears the cache and recomputes everything with the old code. The two versions would + then clear each other's cache, at a full pass each time. +2. **Index.** Use `archilyzer index`, or `/sites` → **Build index** (`pnpm ops build-index`). + Run it through the editor, or with the editor's build queue idle: `build:index` writes the same + LMDB. +3. **Stats.** Use `archilyzer build stats`, or `/sites` → **Build stats dataset**. The first run + logs `Stats schema change (5 -> 6); clearing stats cache.` and re-extracts every video. +4. **Homepage and hub.** Build and deploy both: `archilyzer build homepage` then the homepage + deploy (`pnpm ops build-homepage` with `{"deploy": true}`), and `archilyzer build hub` then + `pnpm ops deploy-hub`. +5. **The six sites.** Build and deploy them so their `/stats` bundles and charts carry the new + stats (`pnpm ops build-deploy`, or `archilyzer build all` and the deploy). + +**Cost:** one full pass in step 3. On the real corpus that is about 79,500 videos and 39.3 GB of +`metadata.info.json`. Metadata was measured to read and parse at 149 MB/s, which is about 4.5 min. +With about 77,000 cue decodes and the sidecars, expect **6–12 minutes**. Later runs are incremental +again. + +**Live check:** +- The homepage's Jasolyzer card shows about 1,751 transcripts, 1 channel and about 3,808 hours. +- `https://jasolyzer.pages.dev/stats/page-0000.json` has no record with `hasTranscript: true` + and `transcribedDate: null`. + +## Left + +- **MCP "truncated" for no transcript at all.** A video with no transcript has coverage 0 + (`transcriptCoverage(undefined, d > 0)`), so MCP `get_video_metadata` tells it "covers only + 0% — truncated". This is older than the cache bug and out of scope. The coverage should be + null when there are no cues. +- **`cueCount` can drift.** It can move without the key moving when buildIndex re-indexes for an + unrelated reason and reads a newer `transcript.cues.json` (FACTS, "The stats cache key"). + `hasTranscript` and the date are not affected. +- **One metadata-key split is unmeasured.** When `transcript.cues.json` is fresh, buildIndex keys + the cues by that file's `uploadDate`, while stats key them by the metadata's. They agree for + all 1,865 of Jasolyzer's cues files. It was not measured corpus-wide.