# Better digest inputs — round 1 notes Written 2026-08-12. Companion to the harness-generated reports in this directory. `chapters-baseline.{json,md}` is the Stage 2a result; `speakers-round1.{json,md}` is reserved for the paired A/B, which has **not** run (see "What did not happen"). The plan this implements asked for a number, not a feature. Here is what was bought. --- ## 1. The corpus facts, re-measured rather than assumed Every load-bearing figure in the plan was re-derived from disk. They hold: | | plan said | measured | | | --- | --- | --- | --- | | videos with a non-empty uploader `chapters` array | ~19.5% sampled | **14,923** | ✓ | | + a transcript + >=4 chapters (the oracle pool) | ~11,366 | **12,139**, across **43 channels** | ✓ | | + retained audio (the Stage 2b pool) | 34 | **34** | ✓ exact | | Stage 2b pool by channel | shondo-vods 20, shondo 7, HasanAbiVODs3 3, kirsche 2, omnibased 2 | identical | ✓ exact | | already-digested videos with >=4 marks | 14, all omnibased | **15**, all omnibased | ✓ | Median 7 uploader marks per video in the oracle pool, max 136. ## 2. What was built **`common/lib/boundaryScore.ts`** — the metric the harness did not have. Every score `digest-bakeoff.ts` produced before this one (`zeroYieldRate`, `chaptersPerHour`, `rejectionRate`, `maxGapSeconds`, `genericTitleRate`, `duplicateTitleRate`) is a *defect counter*: it says how malformed the output is, never whether a boundary is in the right place. So the harness could rank candidates by which was least broken and still not say which segmented better. Design decisions that matter, each pinned by a unit test (12 of them): - **The oracle is independent of the experiment.** Scoring "chapters align with speaker changes" against a variant *given* the speaker changes proves nothing. Uploader chapters are authored by a human who never saw our prompt. - **The t=0 boundary is dropped.** Nearly every uploader list opens at 00:00 and every segmentation trivially has a boundary there. Counting it hands both arms of any A/B a free true positive worth ~1/n. A metric that cannot be lost is not a measurement. - **Matching is one-to-one, closest pair first.** Ten boundaries crammed around one uploader mark score *one* match, not ten — that over-segmentation failure is exactly what precision exists to punish, and a nearest-neighbour count would reward it. - **Boilerplate marks are dropped** ("Intro", "Sponsor", "Outro"). A boundary in the ad read is a boundary in the furniture, not the subject. - **Uploader *titles* are never scored against.** Boundaries are facts; titles are the uploader's expression. This is what keeps the metric free of the copyright objection that rules out admitting their chapter text. **`readUploaderChapters`** in `common/lib/videoStatus.ts`, beside `readVideoDurationSec` and using the existing `META_FILENAME` — the location the plan specified, so no second path to this data can grow later. Nothing in the codebase read `chapters` before this. **Bake-off integration** — pooled precision/recall/F1 at 30 s and 60 s (pooled, not a mean of per-video rates, so a 3-mark video cannot outweigh a 29-mark one), median offset, a `--pick-chapters` sampler, a `--score-existing` zero-GPU mode, a `--speakers on,off` axis, and `--max-cues` for A/B headroom. ### Two production files were touched, both additively and both pinned The plan scoped code to the bake-off. Two exceptions were unavoidable to run the experiment honestly, and each renders **byte-identically when unused**: - `transcriptToMarkdown.prefixForCue` — a per-line prefix hook. The label cannot go inside the timestamp bracket: the prompt tells the model to copy `[HH:MM:SS]` markers verbatim and `HMS_PATTERN` rejects the echo. Returning null for every cue reproduces the current line exactly (test: *"prefixForCue returning null for every cue is byte-identical to omitting it"*). - `ChapterPromptInput.speakerRoster` — a per-chunk cast list. Deliberately *not* `contextNote`: that is per-channel, hand-authored and hashed into `contextHash`, and folding a derived per-video roster into it would mislabel it to the model and make a channel note and a roster mutually exclusive (test: *"a prompt with no speakerRoster renders byte-identically to today"*). **No `PROMPT_VERSION` bump, and none is needed** — the shipped lane sets neither field, so every digest the sweep produces is byte-for-byte what it produced before. That is the same compatibility argument that let `timestampMode` in at version 1. ## 3. Stage 2a — the baseline, and it caught a real bug on day one **Free result first.** `--score-existing` scores digests already on disk. Over the 15 scorable ones, with no model calls and no writes: > **P@30 10.8%, R@30 26.1%, R@60 35.9%, F1@30 15.2%** across 153 uploader marks. Per-video F1@30 ranged 0% to 50% — the metric discriminates rather than saturating. **Precision is partly a density artifact and must not be read naively.** The digest targets one chapter per 4 minutes and is deliberately denser than uploader chaptering (372 generated vs 153 uploader marks here). Recall is the half that answers "did it find the human's boundaries"; precision should be read together with `chaptersPerHour`. **The metric immediately found a defect.** `omnibased/3fl_uRFXxLk`: the video runs 0–5119 s and the uploader marked chapters across all of it, but the digest emitted chapters only between 3184 s and 4865 s — it summarized the last third and skipped the rest. Median offset 1916 s, F1 0%. `maxGapSeconds` corroborates, but boundary scoring quantified it independently. That is the metric earning its keep before any A/B. The generated baseline runs over a fresh 12-video, 8-channel, 4-bucket sample (`chapter-sample.json`, 170 uploader marks) at the **production configuration** (`chunk-local`, 8192, 600 cues) — the bake-off's own default is `absolute`, which is *not* what the sweep runs. Re-run it with: ```sh cd common && pnpm exec tsx ../plans/tools/digest-bakeoff.ts --label chapters-baseline \ --sample ../plans/bakeoff/chapter-sample.json --modes chunk-local \ --candidates 'qwen2.5:7b@8192' ``` **Final result, all 12 videos, 8 channels, 163 scorable uploader marks, 87 chunks:** | | full sample (12 videos, 8 channels) | short/medium only (5 videos) | `--score-existing` (15 videos, omnibased) | | --- | ---: | ---: | ---: | | **R@30** | **21.5%** | 12.2% | **26.1%** | | R@60 / F1@60 | 14.1% (F1) | 19.5% (R) | 35.9% (R) | | P@30 | 5.6% | 20.8% | 10.8% | | F1@30 | 8.8% | 15.4% | 15.2% | | median offset | 109 s | — | — | | within 30 s of a mark | 22.7% | — | — | Guard rails on the same run: 8.1% zero-yield (7/87 chunks), 15.96 chapters/h, 14.8% rejection, 16.5% generic titles, 47 s/audio-hour → **43.1 projected sweep days** (against `digest-plan`'s 53.8; this run was partly contended, so read it as the same order, not as a refinement). Name-in-title is 0%, correctly — this arm has no roster. **READ RECALL. PRECISION IS MOSTLY A DENSITY ARTIFACT, AND THIS TABLE PROVES IT.** Recall@30 lands at 21.5% / 12.2% / 26.1% across three samples that share no videos — a narrow band. Precision swings 5.6% → 20.8% over the *same* metric on the *same* engine, purely with the duration mix: on short videos the digest emits fewer boundaries than the uploader (24 generated vs 41 marks), on 8-hour VODs far more (149 vs 9). F1 inherits that swing, which is why the full-sample F1 (8.8%) reads worse than either subsample without the digest having gotten worse at anything. Practical rule for any future candidate comparison: **compare on recall, and read `chaptersPerHour` beside it.** A variant that "improves F1" by matching the sample's density has improved nothing. ### Checkpointing was added because a long run cannot be trusted to reach its own end The reports are only written when the whole run finishes. A first attempt spent ~26 minutes of engine time and produced nothing readable, and the postmortem was itself instructive: it had **not** crashed. Three bake-off runs were alive at once — a relaunch issued on the mistaken belief the first had died, which truncated the first's log and made it look restarted. The three starved each other. With exactly one run, per-chunk time fell from **20-42 s to 3 s**. Two lessons, both now encoded: `scoreVideo` results are checkpointed to `