# Verified codebase facts Facts established by reading the tree, each with a `file:line` anchor. **Trust these over re-deriving them** — that is the entire point of this file. If one turns out to be stale, correct it in place and note the correction in `STATE.md`. Append as you verify new things. Do not add anything here you have not actually checked. Last full verification pass: **2026-07-26**, against `main` @ `18c5a7a`. --- ## Corrections to the original draft plan These were confidently stated and **wrong**. They are the most valuable entries here. ### `windowCues()` is not a chunker `common/lib/transcriptWindow.ts:21` — `windowCues(cues, centerSeconds, { before = 45, after = 45, maxCues = 60 })`. It is a **center-based** window around a search hit, with `before`/`after` in **seconds**. It cannot split a transcript into sequential context-sized chunks. Callers: `export/app/ask/useAskChat.ts:1333`, `export/app/lib/searchAgent.ts:509,742`. → A sequential `chunkCuesForContext()` must be written. Put it in the same file; it is pure and `transcriptWindow.test.ts` already exists. ### No LMDB `maxDbs` bump is needed `maxDbs` caps the named DBs opened *by one handle*, not the DBs present in the environment. - `common/controller/buildIndex.ts` — `maxDbs: 17`, now opens **15**: the original 12 (`sums`, `cues`, `subs`, `mtimes`, `byChannel`, `pageHashes`, `subPageHashes`, `channelStats`, `posts`, `postPageHashes`, `channelPostsStats`, `meta`) plus the three Phase 2 added (`digests`, `digestPageHashes`, `channelDigestStats`). Still under. - `common/lib/channelSignature.ts:48` — `maxDbs: 14`, opens **2** (`mtimes`, `meta`). - `common/controller/duplicateShorts.ts:111` — `maxDbs: 12`, opens **3**. None of these break. The "17 vs 14 inconsistency" is a red herring. → What **was** required and easy to miss: the schema-invalidation block enumerates every sub-DB by hand with `clearAsync()`. **Done (2026-07-29)** — the three digest sub-DBs are listed there and `SCHEMA_VERSION` is **13**. The next person to add a sub-DB still has to add its `clearAsync()` line by hand; nothing enforces it. ### `JobProgressMetric` is not a closed union in practice — **RESOLVED, table was stale** **This entry is kept only to stop the next session re-doing it. Verified 2026-07-29: all of it is already done.** The union is extended and every site is correct: - `common/jobs/registry.ts:17` — `JobProgressMetric` is `"downloads" | "transcripts" | "digests"`, and `:25` `JobTaskKind` is `"download" | "transcribe" | "digest"`. - `editor/app/jobs/components/RunningJobsList.tsx:6-9` **imports** `JobProgressMetric` and `JobTaskKind` rather than re-spelling them, so TypeScript now does flag the next member. A comment on the union at `registry.ts:14` records that this is why. - `editor/app/widget/components/MonitorWidget.tsx:628` is a `METRIC_PREFIX` lookup table (`Record`), not the two copies of a `metric === "downloads" ? …` binary the old table described — so it is exhaustiveness-checked and there is nothing left to duplicate. PLAN.md's Phase 2.5 opens by demanding this fix "first". It is done; start 2.5 at its second item. ### Queue concurrency is not declared per key `common/lib/queueKeys.ts` has **no** concurrency map. `common/jobs/registry.ts:158` hardcodes `concurrency: 1` for every non-empty `queueKey`. A `queueKey` of `""` means "run immediately, unserialized". `TRANSCRIPTION_QUEUE` is defined at `common/lib/platform.ts:112` and merely re-exported from `queueKeys.ts:2-8`. `resolveQueueKey(defaultKey, override)` is at `queueKeys.ts:27-32`: `undefined` → default, `""` → immediate, else trimmed. ### `resolveCitationSeconds` snaps to snippets, not cues `export/app/ask/citations.tsx:76` scans `src.snippets` — only the ~5 matched lines per video (`export/app/lib/askRetrieval.ts:132`). It never sees a cue list, so it cannot be reused for cue snapping. → The right reference for a cue-boundary snap is `findActiveIndex` (`common/components/TranscriptModal.tsx:473-491`), a hand-rolled binary search over cue starts. ### There is no vitest or jest Unit tests are `node:test` run via `tsx --test`. Only `mcp/package.json:12` has a `test` script (`tsx --test src/*.test.ts`). `common/package.json` has **no `scripts` key at all**; neither `editor` nor `export` has a `test` script. ~46 `*.test.ts` files across the repo are unrunnable as a suite today. *(Amended 2026-09-28, release 13 slice W2: no longer true. Every package with unit tests has a `test` script — `common` (`pnpm --filter yt-dlp-transcript-common test`), `mcp`, `homepage`, `export` (added by W2: `tsx --test "app/**/*.test.ts"`, 11 files / 98 tests) and `editor` (added by W1 in the same release). The list is in `CONTRIBUTING.md`, "Tests", and the gate list in `plans/tools/implementer-rules.md`. Still no vitest or jest.)* ### `subsCache.ts` has no IndexedDB - `common/components/subsCache.ts` — memory + react-query only. - `common/components/summariesCache.ts` — pure react-query, eager over all pages, no IDB, no `slugToPage`. - `common/components/transcriptCache.ts` — the lazy fetch/memo layer (manifest → `slugToPage` → page), and it opportunistically warms every entry in a fetched page. - `common/components/transcriptStore.ts` — the IDB layer. `DB_NAME = "yt-dlp-transcript-browser"`, **`DB_VERSION = 4`** (`:15`), `STORE = "transcripts"`. Its `onupgradeneeded` (`:43-47`) **deletes and recreates** the store, so a version bump wipes the cache. → Copy the `transcriptCache` + `transcriptStore` **pair**. → `editor/e2e/export-player-platform-cache.spec.ts:115` hardcodes DB version **3** and is already stale relative to `DB_VERSION = 4`. ### Social-posts incrementality cannot be copied for a per-video artifact `buildIndex.ts:233` — `scanSource()` does `if (isSocialChannel(cfg)) continue;`. Social channels never enter the mtime scan at all, which is *why* posts use a per-channel shard-signature (`buildIndex.ts:1088-1118`) instead of `MtimeRecord`. → For a per-video artifact use the mtime mechanism (`digestMs` in `MtimeRecord`). Follow the posts precedent (`buildIndex.ts:1065-1240`) only for the **page-tree emission shape**. ### Archive gating already has a seam `common/controller/archiveTranscripts.ts:177` computes `const requiresRewrite = !build.includeMetadata || build.prettyPrint;` and branches at `:205-209` to `writeTransformedCues(...)` vs `link(...)` (a hard link straight from the source tree). → Phase 8 extends that predicate and that transform, rather than building a parallel gate. --- ## Naming hazards **`summaries` is taken.** `exportSummariesDir` (`common/lib/paths.ts:58,151`) → `/summaries/{manifest,page-NNNN}.json` is the cross-channel browse index of video *metadata* (`DisplaySummary` records), documented in `common/lib/corpus.ts:35-36` and served by `common/components/summariesCache.ts`. Nothing to do with AI summaries. Use **`digest`**. **Never name a sidecar `transcript..`.** `common/lib/videoStatus.ts:88` — `export const SUB_FILE_RE = /^transcript\.([^.]+)\.([^.]+)$/;` treats any such file as a subtitle track. The exclusion list in `readSubTracks` is hardcoded. Use `ai-digest.json`, `diarization.json`, `attribution.json`. Since one-core phase 3 slice 4b (2026-09-24) a sidecar is declared with `sidecar(filename, schema)` (`common/lib/sidecar-server.ts`), which THROWS at declaration — module load — on a name `SUB_FILE_RE` matches; `sidecar-server.test.ts` enumerates `SIDECAR_FILENAMES`. **`report` already means three things**: `reportDebouncePreset` in settings, the `refresh-report` job kind, and `/ask`'s `ReportPanel`. Use `feedback` / `review` instead. **`tags` means three unrelated things, and the record field is `curatedTags`.** | Which | Where | What it is | | --- | --- | --- | | yt-dlp keywords | `common/lib/transcripts.ts:21` — `TranscriptSummary.tags` | The platform's own keywords from `metadata.info.json`, searchable via the `"tags"` scope. The UI label for this scope is being changed to **"Keywords"** (`QueryLeafView.tsx:46,156-159`); the code token, the URL and the MCP `scopes` enum stay `"tags"`. | | AI digest topic tags | `common/lib/digests.ts` — `VideoDigest.tags` | Model-generated topics on a digested video. | | **Curated tags** | `common/lib/curatedTags.ts`; record field **`curatedTags`** | The operator's cross-channel vocabulary ("every stream where X collabs"), published at `/tags.json`. | Never name the curated field `tags`. Never assume a `tags.json` is the keyword list: `transcripts/tags.json` and `sites//tags.json` are curated tags. **`listed` / `unlisted` mean three unrelated things** (release 14, slice HS): | Which | Where | What it is | | --- | --- | --- | | A video's visibility | `common/lib/availability.ts` (the `"unlisted"` state, `isUnlisted`), `common/lib/transcripts{,-server}.ts`, `common/components/shareUrl.ts`, `common/controller/buildIndex.ts` | The platform's own "unlisted" (reachable by link, not listed on the channel). | | The hub's list has loaded | `export/app/components/hub/useHubSites.ts` — `listed` | `/hub-sites.json` has been answered and `/hub-summary.json` has settled. | | **A site the family lists** | `site.json` `listed` (`common/lib/siteSchema.ts` — `isListedSite`, `channelsOnlyOnUnlistedSites`) | Absent = listed; `false` keeps the site off the homepage, the hub and the other sites' footers, and out of the public totals. A PRIVATE site (`audience: "private"`, release 17 XP) is never listed, whatever `listed` says. | A grep for either word finds all three; read the file before assuming which. **Never publish a path segment named `.git`.** wrangler's Pages upload drops `**/.git` (and `**/node_modules`) SILENTLY, and Cloudflare's managed rules block `/.git/` requests. The source mirror is `/source/archilyzer.git/` for that reason (release 12, below). A mirror named `.git` would deploy "successfully" and 404. --- ## Reusable helpers (do not rewrite these) | What | Where | Notes | | --- | --- | --- | | Transcript → text for an AI | `common/lib/transcriptToMarkdown.ts:60` | Declared single source of truth. `stampForCue?: (clock, seconds) => string` at `:43` takes precedence over `linkForCue`. | | Zero-padded `HH:MM:SS` | `common/lib/aiHandoff.ts:54` — `hms(s)` | So the stamper is `(_clock, s) => hms(s)`. No new formatter needed. | | Read a normalized transcript | `common/controller/normalizeTranscript.ts:161` | `readNormalizedTranscript(cuesPath)`. | | Freshness check | `normalizeTranscript.ts:193` | `isCuesJsonFresh(videoDir)`. | | tmp+rename write idiom | `normalizeTranscript.ts:139-141` | `${path}.tmp-${process.pid}` then `rename`. | | Engine registry + client-safe listing | `common/lib/transcriptionApps.ts:62-81, 248-278` | The pattern for `digestApps.ts`. | | Inheritance resolve (global → channel) | `common/lib/cookiePolicy.ts:53-66` | Two rules: truthiness-after-trim for free text, type-guard for enums. | | Child process into a job log | `common/jobs/runChild.ts:22` | `runChildIntoLog(onLog, signal, opts)`. | | Bounded concurrent pool | `common/jobs/concurrentRunner.ts:29` | `runPool()`; honors `signal` + `drainSignal`. | | Managed server action | `editor/app/channels/[slug]/whisperActions.ts:45-103` | The `runManagedFunction` template. | | Paginated page writer | `common/controller/buildIndex.ts:694` | `createPageWriter()` — size budget + sha1 skip-if-unchanged. | | Health probe | `common/controller/remoteTranscribe.ts:51` | `pingRemoteHealth(worker, timeoutMs)`. | | ENOENT → friendly message | `common/social/xGalleryDlFetcher.ts:207-213` | `/ENOENT/.test(message) ? "X not found (set X_BIN)" : message`. | | ETA estimator | `editor/app/jobs/active/buildActiveJobs.ts:38` | `computeEtaSeconds`. | | Binary search over cue starts | `common/components/TranscriptModal.tsx:473-491` | `findActiveIndex`. | | Curated-tag model (pure) | `common/lib/curatedTags.ts` | `sanitizeTagsConfig` / `mergeTagDefs` / `effectiveTagsFor` / `compileTagRules` + `evaluateCompiledRules`. Compile ONCE per build, evaluate per video. | | Curated-tag persistence | `common/lib/curatedTagsStore.ts:198` | `applyTagAssignments` — the ONE write path for every tag writer (editor, ops, umtool); one atomic write per call, and it refuses to pin a tag the vocabulary does not define. | | A chat cue's author | `common/lib/curatedTags.ts:495` — `chatAuthorOf` | The `: ` prefix; there is no `author` field on `Cue`. | --- ## Transcription engines (Phase 0) All three are **already selectable** — `common/lib/transcriptionApps.ts:248-252`: | id | Label | Line | Fields | Notes | | --- | --- | --- | --- | --- | | `whisper-cpp` | whisper.cpp (whisper-cli) | 157 | model, customArgs | `DEFAULT_TRANSCRIPTION_APP_ID` (`:254`) | | `chough` | chough | 179 | model, remoteUrl, chunkSize | writes exactly the `-o` path | | `parakeet` | parakeet.cpp (overlapping segments) | 216 | model, chunkSize, device | **`supportsPartialStop: true`** (`:222`) — stops after the current window and stitches a partial transcript on SIGTERM, so a long run is interruptible | Exposed to the settings form via `listTranscriptionApps()` (`:272`), consumed at `editor/app/settings/page.tsx:45`. Per-worker `appId` means a mixed fleet already works (`Worker.appId`, `common/lib/workers.ts:102`; this anchor was `settings.ts:837` before 2026-09-23). **Nothing measures relative speed.** Phase 0 fills that gap — record the numbers below. > _Benchmark results: not yet measured._ --- ## Build pipeline constants | Constant | Where | Value | | --- | --- | --- | | `SCHEMA_VERSION` | `buildIndex.ts:165` | **13** — and curated tags deliberately did NOT bump it (2026-09-21); see below | | `CUES_FILE_VERSION` | `normalizeTranscript.ts:33` | 2 | | `CORPUS_SPEC_VERSION` | `common/lib/archive/contract.ts:35` (re-exported `corpus.ts:59`) | **4** (3 → 4 for `/tags.json`, 2026-09-21) | | `MANIFEST_VERSION` (site summaries) | `common/lib/manifest.ts:36` | 3 | | `TRANSCRIPTS_MANIFEST_VERSION` | `manifest.ts:43` | 1 | | `SUBS_MANIFEST_VERSION` | `manifest.ts:59` | 4 | | `SUMMARIES_PAGE_SIZE` | `manifest.ts:37` | 1000 | | `DEFAULT_ARCHIVE_MAX_BYTES` | `common/lib/archiveOptions.ts:72` | 25 MB — over it, **all** archives are dropped | `MtimeRecord` (`buildIndex.ts:146-154`): `metaMs`, `transcriptMs`, `subsMs`, `availabilityMs`, `isDeleted`, `isUnlisted`, `indexKey`. Changed-detection comparison at `:455-468`; the single `mtimes.put` at `:631-639`. `channelSignature.ts` uses a **narrower structural subset** (`:23-27`): only `metaMs`, `transcriptMs`, `subsMs`, and its hash (`:65-83`) covers only those three. **A new per-video artifact that does not move one of those three will not invalidate archive/compose caches.** `compose-site.ts`: `ComposeCache` at `:486-495`; `reconcileChannelTree()` at `:579`, called three times (`:667` transcripts, `:675` subs, `:686` posts); it passes `ignoreBasename: "manifest.json"` internally at `:597` because `generatedAt` churns. Transcript shards ship the **full `cues` array inline** — `buildIndex.ts:830-832`. ### Curated tags invalidate through `meta`, not through mtimes (2026-09-21) `curatedTags` is DERIVED at build time — `(rule hits ∪ pins) − suppressions` — and **rule hits are never persisted**. So nothing in the per-video mtime diff can see that a rule was edited or a video pinned, and the mechanism is four keys in the existing `meta` sub-DB (`common/controller/curatedTagsIndex.ts:44-52`), checked **unconditionally on every build**: | Key | Covers | Deliberately excludes | | --- | --- | --- | | `curatedRulesHash` | every def's id + `order`, and every **enabled** rule's kind, pattern, channel scope and date range | label / colour / group / hidden — relabelling a tag must not re-page 30,000 videos | | `curatedAssignHash` | every pin and suppression | — | | `curatedAssignSigs` | per-key `manual|suppressed` signature | not a hash: it is what makes an assignment-only edit re-derive just the videos that moved | | `curatedPagesPending` | "records were re-derived, the shared pages have not been written yet" | set when the hashes are recorded WITH changes, cleared only after the shared page build — without it a Ctrl-C between the two would store the hashes and leave the shards stale for ever, since the next build would compare equal and dirty nothing | `reapplyCuratedTags` (`:279`) runs after the mtime diff: rules changed → every video is a candidate (cue reads gated per video by channel/date scope); assignments changed → only the symmetric difference of the signatures, located by a key-only `byChannel` range scan per affected channel. **It is `async` and yields to the event loop every `CURATED_REAPPLY_YIELD_EVERY` (200) records walked**, in both loops — a rule edit makes this - **Key-only LMDB range scans use `getRange({start,end})` WITHOUT `values: false`** (verified 2026-09-22): with that option lmdb-js yields bare keys, not `{key, value}`, so `for (const { key } of …)` reads `undefined`. The first assignment-only re-apply on the real corpus crashed on exactly this (`8db8a3be`); `recencyIndex.ts:111` documents the same convention, and the test fake now yields bare keys under `values: false`. the one whole-corpus pass in the build, and at ~500 cue-decodes/second an unscoped caption rule is 2–3 minutes; run synchronously inside the editor that blocked every page render for the duration (observed live 2026-09-21, `/channels` `/jobs` `/tags` all >25 s). The whole-corpus scan therefore passes `snapshot: false` to `sums.getRange`, lmdb-js's long-duration-iterator idiom, so the read transaction is not held open across minutes of writes. `ReapplyResult.yields` reports the count; it is 0 below the interval. It re-derives out of LMDB — **no video directory is read twice** — and whatever it changes flips `sharedNeedsBuild`, because the page writer's sha1 skip is what then leaves the untouched pages untouched — and `curatedPagesPending` is what covers the build that was interrupted between the two. The per-site fingerprint (`buildIndex.ts:1765`) includes both hashes, or a tag-only edit would leave a site reporting "up to date" with yesterday's counts. **No new sub-DB**, so the `clearAsync()` enumeration is unchanged. **And no `SCHEMA_VERSION` bump — on purpose.** Every prior bump existed because something would otherwise never be derived at all (v13 populates a new sub-DB; nothing re-derives it lazily). Curated tags re-derive lazily *by design*: `curatedRulesHash` is absent on an index built before them, so it can never equal the current hash and the first build re-derives every record on its own — ~80 s across 30k videos, out of LMDB, and **zero page writes on an untagged corpus** because an empty derivation leaves each summary byte-identical (tested: "an UNTAGGED corpus cold-starts to zero changes"). A bump would instead wipe the cache and re-read 76k video directories to reach exactly the same state. **The site layer is presentation-only.** `sites//tags.json` relabels, recolours, reorders, hides and may add site-only tags; it can never delete a corpus rule or an assignment. A site-layer RULE is accepted by `mergeTagDefs` (harmless on a published def) but **never tags a record**: rules are evaluated once at index time over records SHARED by every site carrying the channel, so there is no per-site record for a per-site hit to live in. Rules bind only from `transcripts/tags.json` — promote one there to make it fire. The editor therefore offers no site-side rule editor. **`CuratedTagDef.label` is OPTIONAL, and that is what makes the overlay work.** `sanitizeTagsConfig(raw, {layer})` defaults a missing label to the id at the `global` layer and leaves it ABSENT at the `site` layer (`readSiteTags`/`writeSiteTags` pass `site`, in both directions, because the file is sanitized on write AND on every read). Without that, a site row that set only a colour would arrive at `mergeTagDefs` carrying `label: ` and rename the tag on that site — publishing "eva-collab" where the corpus says "Collab". Every reader therefore renders `label ?? id`; the one place the fallback is baked in is `publishedTagsFrom`, because `/tags.json` is read, not layered. **A chat cue's author is a string prefix, not a field.** `common/lib/liveChat.ts:100-105` emits every live-chat cue as ``text: author ? `${author}: ${text}` : text`` (`:104`) — `Cue` is `{start, end, text}` (`vtt.ts:1`) and has no `author`. So a `chat-author` rule matches `chatAuthorOf(text)` (`curatedTags.ts:495`): everything before the FIRST `": "`, and `null` when there is none (a cue with no author can never match). Anything that wants the author of a chat line must use this, not a second copy of the split. ### Measured: what a curated rule costs (2026-09-21, real corpus) `legal-mindset`, 518 videos / 1.37 M caption cues / 1.47 M live-chat cues, the three seed rules (metadata + chat-author + caption), run offline against the production LMDB (warm): | Rules | Wall | LMDB decode | Regex | | --- | --- | --- | --- | | metadata only | 18 ms | 0 | 4 ms | | chat-author only | 848 ms | 727 ms | 110 ms | | caption only | 421 ms | 302 ms | 109 ms | | all three | 1.28 s | 1.03 s | 232 ms | **The cost is the cue decode, not the regex** (4:1). A metadata-only vocabulary is free; a cue-backed rule costs about 2.5 ms per video with cues, so a whole-corpus re-derive at Jeralyzer's scale (30,886 videos) projects to ~80 s — a rule edit, not a rebuild. Scoping a rule with `channels`/`dateFrom` skips the decode entirely for everything out of scope. Hits on that channel: 7 `eva-collab` (metadata), 96 `eva-in-chat` (chat author), 1 `eva-topic` (caption `\belf ?pire\b|elfpyre|legal loli`) — the caption number is a pattern problem, not a plumbing one: a spoken name rarely appears spelled that way in ASR output. --- ## Digest bake-off (measured 2026-07-26) Harness: `plans/tools/digest-bakeoff.ts`. Fixed stratified sample in `plans/bakeoff/sample.json` (8 videos, 7 channels, 13 min - 8 h), never written to the corpus. Full tables in `plans/bakeoff/round{1,2}.{json,md}`. Corpus totals captured by the same scan that picked the sample: **76,354 videos, 73,367 with transcripts, 77,298 audio-hours**; 6,038 videos over 4 h holding 40,153 h (8.2% of videos, 52% of audio). Sweep days below = measured seconds-per-audio-hour x 77,298 h, one lane, no parallelism. ### Round 1 — screening (short+medium, absolute) | Candidate | Zero-yield | Chapters/h | Rejection | Max gap | Generic | Sweep days | | --- | --- | --- | --- | --- | --- | --- | | `gemma2:9b@8192` | 0/10 | 15.78 | 0% | 14:59 | 2.3% | 288.1 | | `qwen3:8b@16384` | 1/7 | 15.42 | 23.6% | 39:57 | 11.9% | 204.9 | | `mistral-nemo:12b@16384` | 1/7 | 10.28 | 28.2% | 31:20 | 7.1% | 447.3 | | `qwen3:14b@8192` | 2/10 | 16.52 | 4.3% | 47:51 | 6.7% | 541.3 | | `qwen2.5:7b@16384` (incumbent) | 2/7 | 6.61 | 55% | 41:37 | 11.1% | 59.0 | qwen3 candidates were run with `think: false`; with thinking on they spend most of their output budget on a reasoning block the pinned schema then discards. Two `qwen3:14b` chunks failed with `fetch failed` (ollama dropped the connection under memory pressure — a 9.3 GB model on an 8 GB card) and are counted as zero-yield, which inflates that row. It does not change its exclusion at 541 days. **The incumbent is the fastest by 3.5x and the worst on quality**: it discards 55% of what it generates and lands at 6.6 chapters/h against a ~15 target. ### Round 2 — the long tail (2 x ~3.5 h videos, both timestamp modes) | Candidate | Zero-yield | Chapters/h | Rejection | Max gap | Generic | Sweep days | | --- | --- | --- | --- | --- | --- | --- | | `gemma2:9b@8192` **chunk-local** | 0/17 (0%) | 12.93 | 14.3% | 24:04 | 10% | 132.4 | | `qwen2.5:7b@8192` **chunk-local** | 2/17 (11.8%) | 13.37 | 19.1% | 24:13 | 31.2% | **24.2** | | `qwen2.5:7b@16384` chunk-local | 1/9 (11.1%) | 9.34 | 41.4% | 1:27:48 | 20% | 25.1 | | `gemma2:9b@8192` absolute | 3/17 (17.7%) | 11.07 | 28% | 1:04:07 | 5.2% | 127.9 | | `qwen2.5:7b@8192` absolute | 5/17 (29.4%) | 10.35 | 47.5% | 1:05:16 | 18.1% | 27.7 | | `qwen2.5:7b@16384` absolute (today's default) | 3/9 (33.3%) | 8.19 | 50.4% | 1:27:48 | 17.5% | 26.8 | **Three findings, all measured, none assumed:** 1. **chunk-local beats absolute for every model tested**, on the metric this stage exists to fix. Zero-yield chunks: 33.3% -> 11.1% (qwen2.5@16k), 29.4% -> 11.8% (qwen2.5@8k), 17.7% -> 0% (gemma2). Rejection rate falls with it in every case. The hypothesis — that the model cannot hold a large absolute offset across a long chunk and reverts to counting from zero — is confirmed. 2. **Smaller chunks are a second, independent win, and they are nearly free.** Halving the context (16k -> 8k, 1200 -> 600 cues) took qwen2.5 from 9.34 to 13.37 chapters/h and its max coverage gap from **1:27:48 to 24:13**. The expected cost did not materialise: 24.2 vs 25.1 projected days. Twice the calls at half the prompt each is the same seconds-per-audio-hour. The plan's working assumption that a smaller context roughly doubles sweep cost is **wrong** — it is flat. 3. **Sweep days are far lower than the short-video estimate suggested.** Round 1 measured 59 days on short+medium; Round 2 measures ~25 on the long tail, because long videos amortise the fixed per-call overhead. The 77,298-hour corpus is dominated by long videos, so ~25 days is the number to plan against. **Recommendation: `qwen2.5:7b` at `num_ctx` 8192 with `timestampMode: "chunk-local"`.** It is a strict improvement over today's default on every axis measured, including throughput (24.2 vs 26.8 days). Its one weakness is title quality — 31.2% generic, the worst in the table; by hand its titles read "Introduction and Context" / "Conclusion and Final Thoughts" where gemma2 writes "Andrew Wilson and His Comparison to a Serial Killer". **`gemma2:9b@8192` + chunk-local is the quality option**: 0% zero-yield, 10% generic, comparable density — at 5.5x the wall-clock (132 days). It is the right engine for a targeted re-run of high-value channels, not for the first full sweep. Note gemma2 is hard-capped at an 8192 context by the model itself, so it has no 16k row to compare. Not yet measured: the >4 h `verylong` stratum (Round 2 used the `long` bucket), and the ~100-video validation run. --- ## Pilot: community-notes (measured 2026-07-26) 39 videos, local lane, `qwen2.5:7b` / `num_ctx` 8192 / `chunk-local`, launched from the channel page's Digest stage through the real managed-job path. | | | | --- | --- | | Result | **39 generated, 0 failed**, 90 model calls, 140 warnings | | Chunks | **90 / 90 usable (100%)** | | Chapters | 435 across 39 videos | | Warning codes | `out-of-range` 130, `non-monotonic` 9, `seam-duplicate` 1 | | Generic titles | 103 / 435 (23.7%) | | Re-run | `countMissingDigests` = 0; batch reports **0 generated, 39 fresh, 0 engine calls, 0.1 s** | **The defect video is fixed.** `community-notes/v2chrch` (3505 cues) is the 2.3 h video whose third chunk previously produced NOTHING — nine of eleven starts rejected as out-of-range because the model had reverted to counting from zero. At 8 k / chunk-local it is 7 chunks, **7/7 usable, 42 chapters, one out-of-range warning**, covering 00:16:59 → 02:12:55 with no gap larger than the leading one. The other two 6-chunk videos behave the same: `v2on7b3` 36 chapters 6/6, `v2apmfn` 20 chapters 6/6. **Two honest caveats:** - 130 `out-of-range` rejections remain across the channel. The guard is catching them and the chunks still yield, so this is degraded output rather than lost output — but it is not zero, and it is the metric to watch in the validation run. - **Coverage of the opening minutes is weak on long videos.** The first chapter lands at 00:16:59 (`v2chrch`), 00:15:30 (`v2on7b3`) and **00:41:47** (`v2apmfn`) — the early chunks yielded little or nothing, which is the mirror image of the original defect. Not investigated in this stage. Worth checking whether these channels open with long pre-shows, or whether the first chunk is systematically weaker. Generic-title rate on real data (23.7%) sits between the bake-off's long-tail sample figure for this config (31.2%) and gemma2's (10%), which is consistent rather than surprising. --- ## Duplicate detection at corpus scale (measured 2026-07-26) ### Why the first attempt failed (the counts here are correct and are the reason blocking exists) `duplicate-shorts.ts --all-durations` originally died two different ways: | Heap | Outcome | Elapsed | | --- | --- | --- | | default (~4 GB) | `FATAL ERROR: Ineffective mark-compacts near heap limit` | 87 s | | `--max-old-space-size=24576` | `RangeError: Invalid array length` at the containment `candidates.push` | 34 s | Counting the pairs the old generator would build, without building them: | | Count | | --- | --- | | Videos scanned | 76,354 | | Participants (`duration > 0`) | 76,318 | | Shorts (≤ 180 s) | 6,970 | | Duration buckets (2 s) | 11,008 | | Duration-bucket pairs (same + adjacent) | 9,131,916 | | **Containment pairs** (short × every video ≥ 1.5× its length) | **497,860,972** | | Total candidate pairs | 506,992,888 | Two independent defects, not one: 1. **The containment pass was an unblocked cartesian product** — 98.2% of all candidates. V8 throws `RangeError` because a fast-mode object array caps well below 500 M elements. 2. **Fingerprints were held for the whole corpus at once.** `fingerprintFrom` builds a `Set` of 5-word shingles over the entire transcript (~40 k strings for a 3.5 h video), and the duration pass alone made `needed` ≈ the corpus. That is what killed the 4 GB run before the array was even reached. ### The conclusion that was wrong The earlier note here said corpus-wide detection "does not complete on this corpus" and recorded that "as a cost, not worked around". **That was wrong.** Neither failure is inherent to corpus-wide detection: blocking fixes (1) and block-at-a-time streaming fixes (2). Both are now implemented, and corpus-wide detection completes comfortably in minutes on a default heap. ### Corpus-wide runs, measured (same corpus, 76,354 videos) `detectDuplicateShorts` now iterates blocks — fingerprint a block, evaluate its pairs, keep the confirmed matches, release — so peak memory is O(largest block), and no multi-million-element pair array is ever materialised. `--blocking` selects the nomination strategy. | | `--blocking title` | `--blocking duration` | | --- | --- | --- | | Wall clock | **206 s** | **2,889 s** (48 min) | | Peak RSS | **2,001 MB** | **2,450 MB** | | Blocks | 7,551 title groups (of 67,064 distinct titles) | 11,008 buckets of 2 s | | Oversized blocks skipped | **0** (cap 40) | **0** (cap 2,000) | | Nominated pairs | 8,352 | 9,131,476 | | Confirmed by content | 2,717 | **5,708** | | Rejected by content | 5,250 | 9,125,768 | | Untestable suspects | 385 | **0** (by design — see below) | | Videos fingerprinted | 15,479 | 74,943 | | Clusters written | 2,994 / 6,044 videos (335 `needsReview`) | 2,917 / 5,955 videos (0 `needsReview`) | | Cluster kinds | 79 exact, 2,580 near, 335 suspect | 108 exact, 2,809 near | | Alignment | 1,599 of 2,689 within 5 s | 1,818 of 3,028 within 5 s | The duration run's 9,131,476 nominated pairs land within 0.005% of the 9,131,916 predicted above — the old counting was right, only the conclusion drawn from it was wrong. **Duration blocking yields no suspects, and that is deliberate.** A shared runtime alone was never evidence (it produced enormous false clusters of unrelated same-length videos), so when a duration-nominated pair cannot be content-tested it is dropped, not queued for review. Only title blocking produces suspects, because only *title + near-identical runtime* is a claim worth a human's attention. Title-run cluster sizes are overwhelmingly pairs: 2,950×2, 39×3, 3×4, 1×6, 1×9. 2,728 cross-platform, 2,891 cross-channel. ### Measured recall: what title blocking actually misses Comparing the two runs' clusters directly (title's suspects excluded, since they are not confirmed): | | Count | | --- | --- | | Clusters with identical membership in both | 2,613 | | Title-only clusters | 46 | | Duration-only clusters | 304 | | Videos in title clusters / in duration clusters | 5,351 / 5,955 | | **Videos found only by duration** (re-titled mirrors) | **672** | | Videos found only by title (runtime drift past the bucket) | 68 | So **title blocking has ~89% of duration blocking's video-level recall at 1/14th the wall clock**, and the 672 it misses are exactly the re-titled mirrors it structurally cannot see. `--blocking both` exists for when that 11% matters; the title default is the right routine choice. The containment sweep is now budget-capped (`MAX_CONTAINMENT_PAIRS`, 5 M) and reports itself skipped corpus-wide: `6,970 shorts × 76,318 videos ≈ 531,936,460 pairs`. It still runs, unchanged, at shorts scale. Making it scale is the one piece genuinely deferred — see STATE.md. ### 63% of title matches were rejected at 0.6 — and most of those rejections were wrong Of 8,352 pairs nominated by an exact normalized title *and* a compatible runtime, **5,250 were rejected once their transcripts were compared** at the default 0.6 5-gram Jaccard. Only 2,717 held up. That was read as "the title-level count was right; the inference that they are duplicates was not." **Measurement has since shown the opposite: the threshold was wrong, not the inference.** See the next section. ### MEASURED: the 0.6 near threshold was rejecting real cross-platform mirrors Re-run corpus-wide, title blocking, everything else identical, only `--near` changed (2026-07-27): | | `--near 0.6` (default) | `--near 0.35` | | --- | --- | --- | | Nominated pairs | 8,352 | 8,352 | | **Confirmed by content** | 2,717 | **7,329** | | Rejected by content | 5,250 | **638** | | Untestable suspects | 385 | 385 | | Clusters written | 2,994 / 6,044 videos | **7,450 / 15,030 videos** | | `needsReview` | 335 | 334 | | **Timing-aligned mirrors** | 1,599 of 2,689 | **3,952 of 7,221** | | Wall clock | 206 s | 249 s | The confirmed set grows **+170%** and is a strict superset: **0 videos** present in the 0.6 clusters are absent from the 0.35 clusters. **The 4,456 entirely-new clusters were inspected, not just counted**, and they are real mirrors. Evidence, in descending order of strength: - **4,423 of 4,456 (99.3%) are cross-platform AND cross-channel.** 3,954 pair a channel with its own platform-suffixed sibling (`the-quartering` ↔ `the-quartering-rumble`); of the 502 that don't, the sampled ones pair a creator's alternate channels (`jeremy-hambly` / `quartering-live` / `the-quartering`, `HasanAbiVODs` / `HasanAbiVODsbackup`). - **4,290 of 4,456 have a byte-identical title** across every copy; 4,023 agree on runtime within 2 s; 3,572 share an upload date. - The score histogram is a tight band at **0.45–0.60** (1,995 at 0.55–0.60, 1,481 at 0.50–0.55, 665 at 0.45–0.50) — i.e. they pile up *just under* the old cutoff, which is the signature of a systematic offset, not of noise. - The **138 riskiest** (score < 0.42) and the **502 non-sibling** clusters were sampled across the whole band. Every one examined is the same recording: same title, same runtime to the second, same upload date, mirrored platform. The mechanism is the one the caveat predicted: the two sides of a cross-platform mirror are transcribed by **different ASR engines**, and a 5-gram Jaccard is unforgiving of word-level disagreement. Two transcripts of the same audio from different engines land at **~0.35–0.60**, not ≥ 0.6. The old default was set for same-engine text and silently failed the cross-platform case, which is precisely the case duplicate detection exists to catch. **Consequence for the digest backfill.** The ceiling on shared-digest savings is not 2,717 pairs. At 0.35 it is **7,329 confirmed pairs, of which 3,952 are timing-aligned** enough for a shared digest to be placed correctly — **2.5× the 1,599 previously banked**. Against a corpus of 76,318 videos that is ~5.2% of the sweep avoidable by sharing rather than ~2.1%. **CHANGED 2026-07-29: `DEFAULT_NEAR_THRESHOLD` now ships at 0.35**, in its own commit, after bracketing it at four values — see the next section. Note that the "consequence for the digest backfill" above was **overstated**: the saving is real in pair count but nearly absent in audio-hours, which is the unit the sweep is actually priced in. ~5.2% of pairs is ~1.7 sweep days of 80. Reproduce with: `pnpm exec tsx bin/duplicate-shorts.ts --all-durations --blocking title --near 0.35` (back up `transcripts/duplicates.json` first — the run overwrites it). ### Bracketing the near threshold — 4 runs, identical inputs (measured 2026-07-29) Corpus-wide, `--all-durations --blocking title`, only `--near` varied. 76,354 videos scanned, 8,352 pairs nominated in every run. | `--near` | confirmed | rejected | clusters | videos | aligned | needsReview | wall | | --- | --- | --- | --- | --- | --- | --- | --- | | 0.6 | 2,736 | 5,404 | 2,846 | 5,746 | 4,273 | 168 | 124 s | | 0.45 | 7,147 | 993 | 7,110 | 14,344 | 10,747 | 166 | 199 s | | **0.35** | **7,483** | **657** | **7,434** | **14,997** | **11,217** | 166 | 225 s | | 0.25 | 7,542 | 598 | 7,493 | 15,115 | 11,297 | 166 | ~230 s | Every step down is a **strict superset** — 0 videos lost at any step. Returns collapse: 0.6→0.45 admits **4,264** entirely-new clusters, 0.45→0.35 admits **324**, 0.35→0.25 admits **59**. The marginal bands were read, not counted: | band | new clusters | byte-identical title | runtime ±2 s | cross-platform | SAME-channel | | --- | --- | --- | --- | --- | --- | | 0.6 → 0.45 | 4,264 | 96.2% | 91.1% | 99.5% | **3** | | 0.45 → 0.35 | 324 | 97.2% | 83.3% | 96.0% | **2** | | 0.35 → 0.25 | 59 | 94.9% | 86.4% | 93.2% | **0** | The riskiest cases in each band were printed individually. In the 0.6→0.45 band they are all YouTube↔Rumble pairs agreeing on title (modulo whitespace), runtime to the second, and upload date. The 0.45→0.35 band contains the only two arguable cases in the whole sweep: a `HasanAbiVODs` pair with the same title but runtimes 9 minutes apart, and an `omnibased` pair with identical title/duration/date. Neither is a clear false positive; the first would be refused by the alignment gate regardless. **Adopted: 0.35.** The recall knee is at **0.45** — the value to take if a more conservative assertion is ever wanted. ### The digest sweep, priced in CHUNKS (measured 2026-07-29) **A chunk — one model call — is the unit of cost. Seconds-per-audio-hour is NOT a stable unit here and must not be used for any projection.** Chunk density varies **fourfold** across the corpus, so an s/audio-hour figure measured on one length mix does not transfer to another. Three separate cost measurements looked irreconcilable (27 / 90 / 151 s/audio-hour) purely because of this; re-priced per chunk they agree to within ~2%. #### The chunk census Computed free from `statsByPath.cueCount` — present on **all 73,367** transcribed videos, so it needed no GPU time and no transcript reads. (Corrected 2026-09-28: "all" was all the CACHE knew of. Until stats schema 6 a video transcribed after its stat was first cached kept `cueCount: null` — about 2,500 of them on 2026-09-28; see "The stats cache key" at the end of this file.) At the shipped config (`maxCues` 600, overlap 40, step 560): ``` chunks(c) = c === 0 ? 0 : c <= 600 ? 1 : 1 + ceil((c - 600) / 560) ``` | band | videos | audio-h | % of audio | chunks | **chunks/audio-h** | % of chunks | | --- | --- | --- | --- | --- | --- | --- | | < 15 min | 41,960 | 6,203 | 8.0% | 41,960 | **6.76** | 22.0% | | 15–60 min | 14,036 | 6,631 | 8.6% | 22,446 | 3.39 | 11.7% | | 1–2 h | 5,847 | 8,480 | 11.0% | 22,353 | 2.64 | 11.7% | | 2–4 h | 5,486 | 15,832 | 20.5% | 33,106 | 2.09 | 17.3% | | 4–8 h | 4,504 | 25,304 | 32.7% | 46,363 | 1.83 | 24.3% | | > 8 h | 1,534 | 14,850 | 19.2% | 24,888 | **1.68** | 13.0% | | **total** | **73,367** | **77,298** | | **191,116** | **2.47** | | **4 h+ videos are 52% of the AUDIO but only 37% of the WORK; sub-hour videos are 16% of audio and 34% of chunks.** Any reasoning that prices this sweep in audio-hours gets that backwards. Pinned as a regression test in `controller/digestPlan.test.ts` (it re-derives 191,116 from LMDB and fails if the plan and the chunker ever drift). #### Re-pricing every measurement taken so far | per-chunk cost | source | × 191,116 | previously claimed | | --- | --- | --- | --- | | 11.2 s | round 2, idle box, **engine** time | **24.8 d** | 24.2 ✓ | | 24.7 s | 102-video run, contended, **wall** time | **54.6 d** | 81 ✗ | | 60.6 s | round 2 gemma2, idle, engine | **134 d** | 132 ✓ | **The honest range is ~25 days idle to ~55 contended — not 81.** The audio-hour model reproduces none of the three; the chunk model reproduces all three. #### The plan output itself `common/bin/digest-plan.ts --no-freshness`. Total digestable corpus: **77,298 audio-hours over 73,367 videos** (2,987 have no transcript) = **191,116 chunks**, of which **185,483 chunks / 75,638 audio-hours** remain to generate at `--near` 0.35 → **53.5 sweep days**. Sharing avoids 5,633 chunks (1.6 days). Per-channel density is now reported, because it is what explains a changing rate mid-sweep — the queue runs heaviest-first and the heaviest channels are the CHEAPEST per audio-hour: `HasanAbiVODs3` 1.71 c/h against `chibi-reviews` 6.38. The audio-hour table below is kept as measured, but its `days` column was computed at 90 s/audio-hour and is **retracted**. | `--near` | canonical | mirror-aligned (free) | mirror-unaligned | unclustered | **to generate** | **days** | | --- | --- | --- | --- | --- | --- | --- | | 0.6 | 853 h | 494 h | 344 h | 75,607 h | 76,804 h | 80.0 | | 0.45 | 3,472 h | 1,448 h | 1,866 h | 70,513 h | 75,851 h | 79.0 | | 0.35 | 4,070 h | 1,661 h | 2,206 h | 69,361 h | 75,638 h | 78.8 | | 0.25 | 4,137 h | 1,683 h | 2,242 h | 69,236 h | 75,616 h | 78.8 | **Duplicate sharing is not a meaningful cost lever, and `digestSharing.ts`'s "~11% of the sweep" header is wrong.** Cluster members are 19% of the corpus by video count but **4.3% of its audio-hours**: mirrors skew SHORT while the sweep is dominated by unclustered long-form VODs. Loosening the threshold as far as it goes moves the sweep from 80.0 to 78.8 days. Note the `mirror-unaligned` column: only ~55% of mirrors pass the alignment gate, so counting all mirrors as free would overstate the saving by nearly half. The plan splits them. Heaviest channels (the sweep's work-queue order), audio-hours to generate at 0.35: | channel | h | days | | --- | --- | --- | | HasanAbiVODs3 | 8,331 | 8.7 | | omnibased | 7,783 | 8.1 | | HasanAbiVODs | 5,380 | 5.6 | | rekietalaw | 4,993 | 5.2 | | destiny | 4,020 | 4.2 | | shondo-vods | 3,332 | 3.5 | `HasanAbiVODs3` alone outweighs every duplicate mirror in the corpus combined. ### The id-vs-directory mismatch (measured 2026-07-29) A slug is `${channelSlug}/${id}` and an id is **not** the on-disk directory name for **11,175 of 77,106 indexed videos (14.5%)**, concentrated almost entirely in Rumble re-uploads: | channel | dir !== id | | --- | --- | | the-quartering-rumble | 7,870 (of 7,870 — every video) | | leaflit-rumble | 838 | | midwestly | 758 | | rekietalaw-rumble | 641 | | omnimirror | 522 | | cornbreadman | 337 | Verified directly: for all 7,870 `the-quartering-rumble` entries, `stat.slug === channel/id` and `stat.slug !== channel/dir`. `digestBatch` looked the cluster plan up by DIRECTORY, so every one of those lookups missed silently and the mirror regenerated. See the `videoDir` field on `DuplicateVideoRef`. ### What this unblocks Corpus-wide detection is no longer off. `buildDigestClusterPlan` reads `duplicates.json` and returns an empty plan when it is absent, so the digest batch stays correct either way — it just re-generates mirrors instead of sharing them. --- ## Digest validation run — 102 videos, real corpus (measured 2026-07-27) The first digest run at corpus scale rather than on a hand-picked bake-off sample. Shipped defaults (`qwen2.5:7b` @ 8192, chunk-local, 600 cues, `PROMPT_VERSION` 2), production build, driven through the channel page's **Digest channel** button. `community-notes` (39 videos, 33.5 audio-h — the pilot channel, every digest invalidated by the version bump) plus `friendofrc` (63 videos, 8.7 audio-h, never digested). Score with: ``` cd common && pnpm exec tsx bin/digest-validate.ts community-notes friendofrc ``` **102 digests, 42.2 audio-hours, 153 chunks, 622 chapters, 0 failures.** ### Quality: better than projected on everything except coverage gap | Metric | round 2 projection | **validation (102 videos)** | | | --- | --- | --- | --- | | Zero-yield chunks | 11.8% (2 of 17) | **0.0% (0 of 153)** | far better | | Chapters/hour | 13.37 | **14.74** | better | | Rejection rate | 19.1% | **18.4%** | matches | | Generic titles | 31.2% | **27.0%** | better | | Duplicate titles | 3.2% | **0.3%** | far better | | Worst coverage gap | 00:24:13 | **01:09:48** | **2.9× worse** | **Zero-yield went to actually zero.** Not one of 153 chunks came back empty. Round 2's 11.8% came from 2 chunks in a 17-chunk sample; at 9× the chunks the rate is 0. The chunk-local + 8192 pairing does not produce dead chunks. **Rejection rate landed within 0.7 points of a 2-video projection**, which is the strongest evidence available that the bake-off sample was representative on quality. **Rejections are now almost entirely one guard:** | Guard | Count | round 2 | | --- | --- | --- | | `out-of-range` | **130** (93%) | 13 | | `non-monotonic` | 9 | 2 | | `seam-duplicate` | 1 | 0 | | `language-drift` | **0** | 7 | Language drift disappeared across 42 audio-hours (round 2 saw 7 in 7). The range clamp is doing effectively all of the work, which means it is the guard whose removal would silently corrupt the corpus. **The worst coverage gap is the one metric that got worse at scale, and the per-video panel explains it.** `community-notes/v2fkbw7` (1:58:42, 20 chapters, gap 1:09:48) recorded **13 `out-of-range` rejections** — the model proposed chapters for that hour and every one was thrown away by the clamp. So the gap is not "the model ignored an hour of video", it is "the model's output for that hour was unusable". Those are different problems with different fixes, and the distinction is only visible because the artifact records rejections instead of dropping them. **The median gap is 00:05:19**, so this is a tail, not the norm — report the median alongside the max. ### Throughput: 3.3× worse than projected — **and ~a third of that was the unit** **CORRECTED 2026-07-29.** The `~81 days` this section claimed is wrong on this run's own data. Re-priced against the 153 chunks the run's 102 sidecars actually record in `provenance.chunks`, it says **54.6 days**: | | audio-h | wall | chunks | chunks/audio-h | s per audio-h | **s per chunk** | | --- | --- | --- | --- | --- | --- | --- | | round 2 (long videos, idle box) | 6.96 | — | — | 2.44 | 27 | **11.2** | | `community-notes` (avg 51 min/video) | 33.5 | 42.1 min | 90 | 2.69 | 75 | **28.07** | | `friendofrc` (avg 8 min/video) | 8.7 | 20.8 min | 63 | 7.24 | 143 | **19.81** | | **combined** | **42.2** | **62.9 min** | **153** | **3.63** | **89.4** | **24.67** | The 3.33× factors almost exactly: **1.49× sample length mix (3.63 vs 2.44 chunks/audio-h) × 2.20× per-chunk cost = 3.27×**, against 3.33× observed. **The first cause as previously written was an artifact of the unit, and inverts.** It said "short videos cost ~2× per audio-hour (143 vs 75 s)". True, and irrelevant: **per chunk the short-video channel is 29% CHEAPER** (19.81 vs 28.07 s), because a short video's single chunk is a *partial* chunk — fewer cues to prefill and fewer chapters to decode. Short videos are not expensive work; they are expensive *per hour of audio*, which is not a thing the GPU charges for. So only ONE compounding cause survives, plus a third of the gap being sampling: 1. ~~Short videos cost ~2× per audio-hour~~ — **retracted as a unit artifact.** What is true: round 2 sampled only the `long` bucket, whose 2.44 chunks/audio-h is well below the corpus's 2.47… and *far* below the validation sample's 3.63. The sample mix, not the videos, contributed 1.49× of the 3.33×. 2. **The box was not idle.** `auto-download` and `auto-transcribe` ran throughout (load average 10–14, the transcription engine contending for the same GPU). This is representative of production but not comparable to the bake-off's conditions. This remains the one real cause — but its size is a **hypothesis, not a measurement**, because the 2.20× per-chunk residual compares this run's **wall** clock against round 2's **engine** time. Those are different quantities, and the residual therefore also contains model-load, prefill, and the yield gate's own 3 s-poll idling. `digestApps.ts` now records load/prefill/decode per call precisely so the next measurement can separate them. **Do not plan against ~80–90 days.** Plan against **~55 days on a shared box and ~25 idle**, and treat closing that gap as the open question. Two things follow that were not visible before: the yield shipped with a CPU-worker bug (it stopped the digest lane for `device: "cpu"` transcription competing for zero shaders — fixed, `digest.yieldToCpuWorkers`), and `gemma2:9b` **offloads 1,029 MB of its 6.43 GB footprint to CPU** at 8192 ctx on this 8 GB card while `qwen2.5:7b` (5.12 GB) is fully resident — so gemma2's 134 re-priced days are partly a property of the card, not the model. Verified via `/api/ps` (`size_vram` < `size`). ### Built, specified, and never run once — a recurring defect class Four cases found on 2026-07-29, all the same shape: the code exists, reads correctly, has a settings surface or a caller signature, and **had never executed**. Worth naming as a class, because a code review cannot see it and every test passed. | Case | Evidence it never ran | | --- | --- | | `updateDuplicateOverride` | zero callers | | `shareClusterFromCanonical` | zero callers | | `setProgress` (one path) | referenced nowhere | | **Digest TAGS** | **0 of 109 sidecars had a tags section** | The tags case was the largest: `parseTags`, a tags JSON schema, a `sections` branch in `digestVideo` and a per-section checkbox in `SettingsForm.tsx` had all shipped, `digest-validate.ts` had no tag metrics, `editor/e2e/digest.spec.ts` had no tags coverage, and neither bake-off round scored them. It works — 11 videos, 0 warnings, `parseTags` accepted every live output — and it is now generated, scored and covered by e2e. See the tags section below. The pattern to take from it: **a capability with no metric and no e2e is indistinguishable from a broken one.** The fix is not more review; it is that anything with a settings surface gets one live run and one assertion. ### Notes for any future bake-off round (not gates now) - **`BUCKET_BOUNDS` has holes** (`plans/tools/digest-bakeoff.ts:70-75`). `bucketFor()` returns `null` for any duration in a gap, so those videos are silently unsamplable: **30–45 min**, **90 min – 3 h**, and **4.5 – 6 h** (`long` caps at 4.5 h, `verylong` starts at 6 h). The gaps look deliberate — they keep the strata well separated — but nothing says so and nothing reports how many videos fall in them. The 4.5–6 h hole matters most, since that band is dense in this corpus. - **`sample.json` does not record the strata that produced it.** So a round cannot prove what mix it measured, which is exactly the confound that made round 2's 27 s/audio-hour non-transferable. Any future round should record chunks/audio-hour for its sample beside the per-bucket counts. ### Two things the run confirmed about the plumbing - **The counters agree now.** Before the sweep, `noDigest` reported **39** and `countMissingDigests` independently reported **39** for `community-notes` — the version bump correctly invalidating every pilot digest. Under the old bucket the Digest stage read "All digested" on the same corpus. Snapshot regeneration with the added freshness target takes **0.1 s** for a 39-video channel, so the hot path did not get meaningfully slower. - **Progress bars cannot move during a REGENERATION.** `setProgress` uses `initial: stat.digestCount` and `target: initial + missing`, but `digestCount` counts digests that *exist* — and a regeneration rewrites in place, so the count never grows and `pct` sits at 0 for the whole job (observed: `{initial:39,current:39,target:72,pct:0}`). Cosmetic, but it makes a long re-sweep look wedged. Fixing it means counting digests *at the current identity* rather than digests-on-disk. ## Editor surfaces | Surface | File | Note | | --- | --- | --- | | System paths table | `editor/app/settings/page.tsx:17-28` | Hand-curated tuple list, **not** `Object.entries(paths)`. The blurb at `:57-64` is already stale (missing parakeet/gallery-dl). | | Job kinds | `common/jobs/jobKinds.ts:21-40, 44-255` | ~30 entries. **No entry sets `defaultTier`** — declared but unused so far. | | Replay handlers | `editor/app/jobs/jobReplayRegistry.ts:75-236` | Bucket kinds re-derive ids from the live snapshot (`:50-57`), never a frozen list. | | Snapshot buckets | `common/controller/channelSnapshot.ts:57-143` | 21 buckets, all `string[]` of video ids. Readers default `?.length ?? 0`. `undownloadedIds` is deliberately **outside** `buckets` (`:144`). | | Channel-work sections | `editor/app/components/channelWork/sections.tsx` | `SectionConfig` + `channelWorkSections()` + `sectionsFor(op)`; rendered by `ChannelWorkTable.tsx` (a SERVER component — `primaryAction` is a function). Eight sections: three on `/operations/download`, two on `/operations/transcription`, one on `/operations/digest`, two (`operation: null`) on `/cleanup`. Counters live in `editor/app/lib/actionable/loadActionable.ts`; the review half is `editor/app/review/lib/loadReview.ts`. Unit-tested in `sections.test.ts`. | | Widget sync payload | `common/views/widgetSync.ts:22` (`WidgetSyncPayload`) | Comment at `:18` states it deliberately stays "a handful of scalars". `buildWidgetSyncPayload(inputs)` (`:174`) is also used for SSR seeding. Served as view `widgetSync`; `/api/widget/sync` is a rewrite (2026-09-23; the route file is gone). | | Pipeline band | `editor/app/components/dashboard/PipelineBand.tsx:63-114` | Pure props, no fetching. `` encodes state color. | | Operations board | `editor/app/operations/page.tsx` | `/operations`. Rail + sync row, SSR-seeded, polls `/api/view/autoQueueStatus` every 3 s (`useOperationsStatus.ts:26`; the old path is a rewrite since 2026-09-23). `data-board="operations"` carries `data-hydrated`. (The arbiter bar was here until slice 1.3.) | | One operation | `editor/app/operations/[id]/page.tsx` | `/operations/`, **routed off `operationCatalog()`** — an unknown id is `notFound()`, a new registry entry needs no route work. Every operation with a lane (`pauseLaneFor`) renders `RunnerOperationView` inside `
`; since slice 1.3 that is the only lane section on the page. | | Retired editor routes | `editor/next.config.ts` `redirects()` | `/auto-queue` → `/operations` and `/actionable` → `/operations`, both **temporary (307)**, not permanent — a 308 on a self-hosted admin surface is a support call with no remedy. Query strings pass through. The API paths `/api/auto-queue/{status,control}` and `/api/widget/actionable` did **not** move — since 2026-09-23 `/api/auto-queue/status` and `/api/widget/actionable` are `rewrites()` onto `/api/view/` (`:139-148`). **API paths are rewritten, never redirected.** | --- ## Viewer surfaces `ModalMode` — `common/components/urlState.ts:13`, currently `"transcript" | "chat" | "post"`. Parsed at `:63-65` (unrecognized → `"transcript"`), written at `:126-130`. Referenced in **only three files**: `urlState.ts`, `PlayerProvider.tsx`, `TranscriptModal.tsx`. Lazy chat fetch effect — `PlayerProvider.tsx:590-640`, deps `[urlSlug, modalMode]`, with a ref-based in-flight guard and a comment at `:585-589` explaining why reducer state in the dep array drops results. It snaps back via `writeUrlParams({ vm: "transcript" })` on missing. `seekTo` at `:522-530`. **Styling:** `TranscriptModal.tsx` and `PlayerProvider.tsx` are **not** on semantic theme tokens. Verified hard-coded: `bg-black/70` (`:199`), `bg-zinc-900/80 ring-white/10` (`:210`), `text-white` (`:331`), `text-zinc-300` (`:336`), `border-zinc-700 bg-zinc-900` (`:357`), `bg-blue-950/50` (`:419`). By contrast `export/app/ask/citations.tsx` uses semantic tokens throughout — the two surfaces are on different systems. Do not mix them. `export/app/lib/askProvider.ts`: `Provider` union at `:19`, `PROVIDERS` at `:34-59`, `askStream` switch at `:128-137`, `askOpenAI` at `:336-397`. Provider settings live at `export/app/ask/ProviderSettings.tsx` (**not** `app/components/`); its `Props` is a 21-field flat prop bag (`:13-36`). Two independent search index paths: `common/components/searchIndex.worker.ts` (FlexSearch, client) and `mcp/src/search.ts` (server). MCP has 11 tools (`mcp/src/server.ts:132`) and a 14-member `ShardSource` interface (`mcp/src/source.ts:119`) implemented three times — `LocalSource` (`:185`), `RemoteSource` (`:353`), `HubSource` (`:477`). --- ## E2E environment **`build:hub` clobbers the export e2e fixture in the same worktree (hit 2026-09-12).** `export`'s `build:hub` runs `compose:hub`, which overwrites `export/public/corpus.json` and `sw.js` and drops the hub files beside them. An export e2e run after it in the same worktree then runs against a hub, not a site. Run the export suites first, or restore the site fixture (`corpus.json` from the duplicates-page worktree, remove the hub files, recopy `site-sw.js`) before them. All of those paths are gitignored, so nothing shows in `git status`. **Since release 10 L1 (2026-09-26)** the compose writes land in the worktree's own files, never through its `export/public` links (`common/bin/_publicFile.ts`), so the restore is the per-path seed again (`plans/tools/implementer-rules.md`); the export config re-copies `sw.js` itself. - `editor/playwright.config.ts` — editor on `PORT ?? 3011`, export on `EXPORT_PORT ?? 3010`. `E2E_MODE=start` → `pnpm start:test`, otherwise `pnpm dev:test` (`:10-11`). `fullyParallel: false`, `workers: 1`, `timeout: 30_000`. **AMENDED 2026-10-10 (e2e speed S1):** `start` is the DEFAULT (`E2E_MODE=dev` for `next dev`). Before the servers start, `scripts/e2e-stamp.mjs ensure editor` checks the build's stamp (`editor/.next/e2e/e2e-stamp.json`) and rebuilds through the heavy slot when the tree moved; the test server serves `.next/e2e` (`E2E_NEXT_DIST_DIR`), never `.next`. `CI` keeps `.next` and skips the stamp (the sharded image builds its own). umtool likewise (`.next-e2e-start`); export and homepage stay `next dev`. - `editor/e2e/helpers.ts` — `resetData` (`:32`), `writeSettings` (`:49`), `readJson` (`:90`). - Fake binaries live in `editor/e2e/fixtures/bin/`. The **`SLOWOP` sentinel** (checked case-insensitively against cwd or video id) makes an instant fake emit paced progress — three independent copies at `fake-whisper.mjs:29-40`, `fake-chough.mjs:25`, `fake-ytdlp.mjs:173-183`. - Binary env vars are wired in `editor/package.json`'s `dev:test` **and** `start:test` scripts — a single line duplicated verbatim, differing only in `next dev` vs `next start`. **Adding one env var means editing both.** **AMENDED 2026-09-28 (release 11 O6-B):** that holds for the paths and binaries. A TEST-ONLY variable no longer goes there: it is `E2E_`-prefixed and set in `editor/playwright.config.ts`' `E2E_SERVER_ENV` (`:37-44`, the webServer's `env`) — see "Release 11 — slices O1–O6". - Route-stubbing scaffold: `editor/e2e/export-player-platform-cache.spec.ts:39-111`. ### Known-failing on base — not regressions **Re-verified 2026-07-26** by stashing all working changes and re-running the suspect specs against `8c041fd`. Full suite in default dev mode: **21 failed / 363 passed of 384**, and every one of the 21 is accounted for below. Do not chase these. | Spec | Count | Evidence | | --- | --- | --- | | `actionable.spec` — 2 "queues a job" + the "Needs attention" dashboard card | 3 | fails on base | | `bulk-actions.spec` raw-job-kind tests | 3 | fails on base (known) | | `channels-actions.spec:42` | 1 | fails on base (known) | | `cleanup-actionable.spec` inline "Clean audio" | 1 | fails on base | | `deploy-page.spec` (all three) | 3 | fails on base — `getByRole('button', {name:'Build & deploy'})` also matches "Build & deploy all sites"; `getByLabel('Deploy after build')` matches 2 checkboxes | | `new-channel-onboarding.spec` (all five) | 5 | fails on base — `getByLabel('URL')` resolves to 4 elements (a `