commit 348813a3afa0d7df4a1a2b75c8417d16939ed5e9
parent c31ee48452e94eb68af9053b39da0a17488c3cde
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Tue, 18 Aug 2026 19:59:26 -0400
digest-bakeoff: measure whether a chapter boundary is in the RIGHT PLACE
Every metric this harness had was a defect counter -- zeroYieldRate,
chaptersPerHour, rejectionRate, maxGapSeconds, genericTitleRate,
duplicateTitleRate all say how malformed the output is, never whether the
segmentation is correct. So it could rank candidates by which was least broken
and still not say which produced better chapters.
lib/boundaryScore.ts scores generated chapter starts against the UPLOADER's own
marks from metadata.info.json -- an oracle the model never sees, on 12,139
videos that also carry a transcript and >=4 marks, across 43 channels. Nothing
in the codebase read that field before. Boundaries are facts, so scoring
against them reproduces nothing; the uploader's chapter TITLES are expression
and are never scored against, only checked for boilerplate.
Four decisions, each pinned by a test:
- the t=0 boundary is dropped. Every segmentation trivially has one, so
counting it hands every candidate a free true positive worth ~1/n.
- matching is one-to-one, closest pair first. Ten boundaries crammed around
one uploader mark score ONE match; nearest-neighbour counting would reward
the exact over-segmentation precision exists to punish.
- boilerplate marks ("Intro", "Sponsor") are dropped -- a boundary in the ad
read is a boundary in the furniture, not the subject.
- aggregation is POOLED, not a mean of per-video rates, so a 3-mark video
cannot outweigh a 29-mark one.
READ RECALL, NOT PRECISION. Across three samples sharing no videos, recall@30
holds a narrow band (21.5% / 12.2% / 26.1%) while precision swings 5.6% ->
20.8% on the same engine purely with the duration mix: the digest targets one
chapter per 4 minutes, so on short videos it emits fewer boundaries than the
uploader (24 vs 41) and on 8-hour VODs far more (149 vs 9). F1 inherits the
swing.
The metric earned its keep before any experiment: omnibased/3fl_uRFXxLk runs
0-5119s with uploader marks throughout, but was digested only between
3184-4865s -- the head two-thirds unsummarized.
Two production files gain one optional field each, both byte-identical when
unused and both tested, so NO PROMPT_VERSION BUMP and the running sweep's
output is unchanged:
- transcriptToMarkdown.prefixForCue. The speaker label has to sit OUTSIDE the
[HH:MM:SS] bracket: the prompt tells the model to copy those markers
verbatim and HMS_PATTERN rejects an echo.
- ChapterPromptInput.speakerRoster. Deliberately NOT contextNote, which is
per-channel, hand-authored and hashed into contextHash; a derived per-video
roster folded in there would mislabel it to the model and make a channel
note and a roster mutually exclusive.
Also: --score-existing (scores digests already on disk, zero GPU, zero writes),
--pick-chapters, --speakers off,on, --max-cues, and a per-video checkpoint to
<label>.partial.json because the reports are only written at the end and a long
run does not reliably get there.
The speaker A/B is NOT run. Sample and pipeline are ready
(plans/bakeoff/speakers-round1-notes.md has the resume command), but the box
was loaded, which is the plan's own stated precondition for not starting. What
the one diarized video shows is not encouraging: labels cost only 4.4% prompt
inflation, but they come out as ROLES ("Host", "Caller") rather than names, and
they invent speaker changes -- anti-signal for boundary detection.
Note the 11 digest-context.suggested.md channel notes this work produced live
under transcripts/ and are gitignored; plans/digest-context-review.md documents
them but the note text is not in this commit.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Diffstat:
18 files changed, 3742 insertions(+), 27 deletions(-)
diff --git a/common/bin/digest-bakeoff.ts b/common/bin/digest-bakeoff.ts
@@ -17,23 +17,47 @@
// projected sweep days alone, however good its chapters look, which is why every
// run prints days alongside quality.
//
-// Two modes:
+// BOUNDARY ACCURACY, AND WHY IT WAS ADDED. Every metric this harness scored
+// originally — zeroYieldRate, chaptersPerHour, rejectionRate, maxGapSeconds,
+// genericTitleRate, duplicateTitleRate — is a DEFECT COUNTER: it says how
+// malformed the output is, never whether a boundary landed in the right place.
+// So the harness could rank candidates by which was least broken and still not
+// say which segmented better. lib/boundaryScore.ts closes that with an oracle
+// the model never sees: the UPLOADER's own chapter marks from
+// metadata.info.json, present on ~12,139 videos that also have a transcript and
+// >=4 chapters. Boundaries are facts, so scoring against them reproduces
+// nothing; the uploader's chapter TITLES are expression and are never scored
+// against, only used to spot boilerplate ("Intro", "Sponsor").
//
-// --pick scan the stats cache and write the fixed stratified sample
-// (plans/bakeoff/sample.json). Run ONCE. A moving sample makes the
-// comparison between rounds meaningless.
-// (default) run the candidates over that sample and write a JSON + Markdown
-// report under plans/bakeoff/.
+// Modes:
+//
+// --pick scan the stats cache and write the fixed stratified sample
+// (plans/bakeoff/sample.json). Run ONCE. A moving sample
+// makes the comparison between rounds meaningless.
+// --pick-chapters write a sample restricted to videos that carry >=4
+// uploader chapters, so boundary accuracy is scorable. Kept
+// as a SEPARATE file (speaker-sample.json /
+// chapter-sample.json) so rounds 1-2 stay comparable.
+// --score-existing score the ai-digest.json files ALREADY on disk against
+// their uploader chapters. Zero GPU, zero writes — the way
+// to prove the scorer works before spending a generation
+// run on it.
+// (default) run the candidates over that sample and write a JSON +
+// Markdown report under plans/bakeoff/.
//
// Examples:
// tsx bin/digest-bakeoff.ts --pick
+// tsx bin/digest-bakeoff.ts --pick-chapters --sample ../plans/bakeoff/chapter-sample.json
+// tsx bin/digest-bakeoff.ts --score-existing
// tsx bin/digest-bakeoff.ts --label round1 --buckets short,medium \
// --candidates 'qwen2.5:7b@16384,qwen3:8b@16384,gemma2:9b@16384'
// tsx bin/digest-bakeoff.ts --label round2 --buckets long,verylong \
// --candidates 'qwen2.5:7b@16384' --modes absolute,chunk-local
+// tsx bin/digest-bakeoff.ts --label speakers-round1 --speakers off,on \
+// --sample ../plans/bakeoff/speaker-sample.json --candidates 'qwen2.5:7b@8192'
import path from "node:path";
-import { mkdir, writeFile } from "node:fs/promises";
+import { mkdir, readdir, writeFile } from "node:fs/promises";
import { readFile } from "node:fs/promises";
import { open } from "lmdb";
import { getPaths } from "../lib/paths";
@@ -43,6 +67,7 @@ import { getDigestApp } from "../lib/digestApps";
import type { DigestAppConfig, DigestTimestampMode } from "../lib/digest";
import {
CHAPTER_SYSTEM_PROMPT,
+ DIGEST_MINUTES_PER_CHAPTER,
DIGEST_OVERLAP_CUES,
buildChapterPrompt,
chapterSchema,
@@ -53,7 +78,25 @@ import { parseChapters, type DigestChunkOutput } from "../lib/digestParse";
import { chunkCuesForContext } from "../lib/transcriptWindow";
import { transcriptToMarkdown } from "../lib/transcriptToMarkdown";
import { readNormalizedTranscript } from "../controller/normalizeTranscript";
-import { CUES_JSON_FILENAME } from "../lib/videoStatus";
+import {
+ CUES_JSON_FILENAME,
+ isRealAudioFile,
+ readUploaderChapters,
+ type UploaderChapter,
+} from "../lib/videoStatus";
+import {
+ BOUNDARY_TOLERANCES_SECONDS,
+ isBoilerplateChapterTitle,
+ scoreBoundaries,
+ type BoundaryReport,
+} from "../lib/boundaryScore";
+import { loadAttribution } from "../lib/attribution-server";
+import type { AttributionRecord } from "../lib/attribution";
+import {
+ ATTRIBUTION_MAX_SPEAKERS,
+ ATTRIBUTION_MIN_CLUSTER_SHARE,
+} from "../lib/attributionPrompt";
+import { loadDigest } from "../lib/digest-server";
import type { Cue } from "../lib/vtt";
// ---------------------------------------------------------------------------
@@ -83,6 +126,10 @@ type SampleVideo = {
bucket: Bucket;
durationSeconds: number;
cueCount: number;
+ // Non-boilerplate uploader marks counted at pick time. Recorded for the
+ // record only — the scorer re-reads them from disk, because metadata can be
+ // refetched and a frozen copy would let the sample and the corpus disagree.
+ uploaderChapters?: number;
};
type Sample = {
@@ -251,6 +298,12 @@ type Candidate = {
maxCues: number;
timestampMode: DigestTimestampMode;
think?: boolean;
+ // Render speaker labels into the transcript from attribution.json.
+ //
+ // FALSE MUST BE BYTE-IDENTICAL TO PRODUCTION. The speakers-off arm is not a
+ // control unless it renders exactly what the sweep renders, which is why the
+ // prefix hook is a no-op that returns null rather than a second renderer.
+ speakers: boolean;
};
type VideoScore = {
@@ -271,6 +324,18 @@ type VideoScore = {
// Kept so a table can be sanity-checked against real output by hand, which is
// the only way to catch a model that scores well and reads badly.
sampleTitles: string[];
+ // Boundary accuracy against the uploader's own chapter marks. Null when the
+ // video carries none, which is the normal case for ~85% of the corpus.
+ boundary: BoundaryReport | null;
+ // How many named speakers the render actually carried, and how many kept
+ // titles mention one of them.
+ //
+ // NAME UPTAKE IS THE SPEAKER HYPOTHESIS' OWN METRIC. speakers-off cannot
+ // produce it by construction (the roster is never in the prompt), so a
+ // non-zero value on the off arm means a name leaked in some other way — a
+ // useful tripwire, not a score.
+ speakersRendered: number;
+ titlesWithSpeakerName: number;
};
type CandidateScore = {
@@ -295,13 +360,153 @@ type CandidateScore = {
tokensPerSecond: number;
secondsPerAudioHour: number;
projectedSweepDays: number;
+ // Boundary accuracy, pooled over the videos that had an oracle.
+ //
+ // POOLED, NOT AVERAGED PER VIDEO. A per-video mean lets a 4-chapter video
+ // weigh as much as a 90-chapter one, so a candidate could win by doing well
+ // on the shortest lists. Pooling counts matched/generated/reference across
+ // the whole sample and derives precision and recall from the totals.
+ boundary: {
+ videosScored: number;
+ referenceBoundaries: number;
+ medianOffsetSeconds: number | null;
+ withinThirtySecondsRate: number;
+ byTolerance: {
+ toleranceSeconds: number;
+ matched: number;
+ precision: number;
+ recall: number;
+ f1: number;
+ }[];
+ };
+ speakerNameTitleRate: number;
};
};
+// ---------------------------------------------------------------------------
+// Speaker rendering (the speakers-on arm)
+// ---------------------------------------------------------------------------
+
+// A label per cue, or null where no NAMED speaker covers it.
+//
+// CONSUMES attribution.json, NEVER diarization.json. A bare cluster index
+// rendered as "Speaker 7:" is prompt cost with no semantic content, and it
+// invites chapter titles like "Speaker 7 responds". lib/diarization.ts already
+// argues a cluster index is not an identity; this honours that.
+type SpeakerContext = {
+ // Indexed by position in the FULL cue array, so a chunk can slice it.
+ labelByCueIndex: (string | null)[];
+ roster: string[];
+ // Cues that fell inside a named turn. Reported so a report can say how much
+ // of the transcript the labels actually reached.
+ labelledCues: number;
+};
+
+// Speakers worth rendering: the heaviest by attributed speech, capped and
+// floored exactly as attributionPrompt caps the naming call itself.
+//
+// The cap is doing real work here. Diarization over-splits (median 35 clusters
+// per video on this corpus, max 325), and 8 of the 9 attribution records that
+// existed when this was written were the REJECTED text-only pilot, one of them
+// carrying 346 "speakers" that were raw transcript fragments. Rendering that
+// unfiltered would bury the transcript in noise and measure the noise.
+function namedSpeakers(record: AttributionRecord): Map<number, string> {
+ const seconds = new Map<number, number>();
+ for (const seg of record.segments) {
+ const d = Math.max(0, seg.end - seg.start);
+ seconds.set(seg.speaker, (seconds.get(seg.speaker) ?? 0) + d);
+ }
+ const total = Array.from(seconds.values()).reduce((a, b) => a + b, 0);
+ if (total <= 0) return new Map();
+
+ const ranked = record.speakers
+ .map((s) => ({ index: s.index, label: s.label?.trim() ?? "", share: (seconds.get(s.index) ?? 0) / total }))
+ .filter((s) => s.label.length > 0 && s.share >= ATTRIBUTION_MIN_CLUSTER_SHARE)
+ .sort((a, b) => b.share - a.share)
+ .slice(0, ATTRIBUTION_MAX_SPEAKERS);
+
+ return new Map(ranked.map((s) => [s.index, s.label]));
+}
+
+// Map cues onto turns by MAXIMUM OVERLAP, not by containment.
+//
+// Cue boundaries do not align to speaker turns — measured on the diarized set,
+// only 64.4% of cues fall wholly inside a named turn. Containment would leave a
+// third of the transcript unlabelled for a reason that has nothing to do with
+// speaker identity; max-overlap assigns each cue to whoever does most of the
+// talking during it, and still yields null when no named turn touches it.
+function buildSpeakerContext(
+ cues: Cue[],
+ record: AttributionRecord,
+): SpeakerContext {
+ const named = namedSpeakers(record);
+ const segments = record.segments
+ .filter((s) => named.has(s.speaker))
+ .sort((a, b) => a.start - b.start);
+
+ const labelByCueIndex: (string | null)[] = new Array(cues.length).fill(null);
+ const rosterSeen = new Set<string>();
+ const roster: string[] = [];
+ let labelledCues = 0;
+
+ let cursor = 0;
+ for (let i = 0; i < cues.length; i++) {
+ const cue = cues[i];
+ const cueEnd = cue.end > cue.start ? cue.end : cue.start + 1;
+ // Segments are sorted, so the scan only ever moves forward.
+ while (cursor < segments.length && segments[cursor].end <= cue.start) cursor++;
+ let best: { label: string; overlap: number } | null = null;
+ for (let j = cursor; j < segments.length; j++) {
+ const seg = segments[j];
+ if (seg.start >= cueEnd) break;
+ const overlap = Math.min(cueEnd, seg.end) - Math.max(cue.start, seg.start);
+ if (overlap > 0 && (!best || overlap > best.overlap)) {
+ best = { label: named.get(seg.speaker)!, overlap };
+ }
+ }
+ if (best) {
+ labelByCueIndex[i] = best.label;
+ labelledCues++;
+ if (!rosterSeen.has(best.label)) {
+ rosterSeen.add(best.label);
+ roster.push(best.label);
+ }
+ }
+ }
+
+ return { labelByCueIndex, roster, labelledCues };
+}
+
+// Label on speaker CHANGE only, never per line.
+//
+// MEASURED COST. Per-line labels inflate a rendered chunk by ~12% of characters
+// at the median and up to ~30%, which on the ~10-tokens-per-cue budget that
+// sizes the 600-cue chunk to an 8k window is enough to start truncating — and
+// ollama truncates SILENTLY. Change-only lands at ~3.4% median. Since the
+// speaker changes on only ~26% of cues, the two carry the same information.
+//
+// A cue with no named speaker gets NO prefix and does not count as a change, so
+// a coverage gap reads as "the previous speaker continues" rather than as a
+// fake new person.
+function speakerPrefixer(
+ context: SpeakerContext,
+ chunkStartIndex: number,
+): (cue: Cue, index: number) => string | null {
+ let previous: string | null = null;
+ return (_cue, index) => {
+ const label = context.labelByCueIndex[chunkStartIndex + index] ?? null;
+ if (!label) return null;
+ if (label === previous) return null;
+ previous = label;
+ return `${label}: `;
+ };
+}
+
function renderChunk(
meta: { id: string; title: string; channel?: string; duration?: number },
cues: Cue[],
offsetSeconds: number,
+ prefixForCue?: (cue: Cue, index: number) => string | null,
): string {
return transcriptToMarkdown(
{ ...meta, cues },
@@ -311,23 +516,61 @@ function renderChunk(
includeTags: false,
stampForCue: (_clock, seconds) =>
toHms(Math.max(0, seconds - offsetSeconds)),
+ // Omitted entirely on the speakers-off arm, so that arm's bytes are the
+ // sweep's bytes.
+ ...(prefixForCue ? { prefixForCue } : {}),
},
);
}
+// The cast list for ONE chunk, not for the video.
+//
+// A 12-name roster on a chunk where two people speak is misleading and wastes
+// context. The preamble also has to explain the change-only convention, or the
+// model reads an unlabelled line as an unknown speaker.
+function speakerPreamble(names: string[]): string | undefined {
+ if (names.length === 0) return undefined;
+ return [
+ "Speakers in this section (a name before a line means that speaker begins",
+ "there; unlabelled lines continue the previous speaker):",
+ ...names.map((n) => `- ${n}`),
+ ].join("\n");
+}
+
+// Where a sample video's sidecars live. One definition, because three call
+// sites now need it (scoring, the oracle read, attribution).
+function videoDirFor(video: SampleVideo): string {
+ const paths = getPaths();
+ return path.join(paths.channelsDir, video.channelSlug, "data", video.videoDir);
+}
+
+// The oracle, read FRESH at score time rather than frozen into the sample.
+//
+// A video's uploader chapters can change when metadata is refetched. Freezing
+// them would let the sample and the disk disagree silently; reading them here
+// means a report always scored against what the uploader currently says.
+//
+// Boilerplate marks are dropped. "Intro"/"Sponsor"/"Outro" are boundaries in the
+// video's FURNITURE, not in its subject, and crediting a model for finding the
+// sponsor read measures the wrong thing — the same class of junk the
+// `boilerplate` context field exists to remove.
+async function uploaderBoundaries(
+ video: SampleVideo,
+): Promise<{ starts: number[]; chapters: UploaderChapter[] } | null> {
+ const chapters = await readUploaderChapters(videoDirFor(video));
+ if (!chapters) return null;
+ const kept = chapters.filter((c) => !isBoilerplateChapterTitle(c.title));
+ if (kept.length === 0) return null;
+ return { starts: kept.map((c) => c.start), chapters: kept };
+}
+
async function scoreVideo(
video: SampleVideo,
candidate: Candidate,
log: (m: string) => void,
): Promise<VideoScore | null> {
- const paths = getPaths();
- const cuesPath = path.join(
- paths.channelsDir,
- video.channelSlug,
- "data",
- video.videoDir,
- CUES_JSON_FILENAME,
- );
+ const videoDir = videoDirFor(video);
+ const cuesPath = path.join(videoDir, CUES_JSON_FILENAME);
const transcript = await readNormalizedTranscript(cuesPath);
if (!transcript || !transcript.cues?.length) {
log(` ${video.slug}: no transcript on disk, skipped`);
@@ -339,6 +582,42 @@ async function scoreVideo(
overlapCues: DIGEST_OVERLAP_CUES,
});
+ // Speakers are loaded per VIDEO, once, even though they are rendered per
+ // chunk: the roster has to be sliced to the chunk but the cue->turn mapping
+ // is a whole-video computation.
+ let speakerContext: SpeakerContext | null = null;
+ if (candidate.speakers) {
+ const record = await loadAttribution(videoDir);
+ if (record) {
+ speakerContext = buildSpeakerContext(cues, record);
+ log(
+ ` ${video.slug}: ${speakerContext.roster.length} named speaker(s), ` +
+ `${pct(cues.length > 0 ? speakerContext.labelledCues / cues.length : 0)} of cues labelled`,
+ );
+ } else {
+ // NOT an error and NOT a skip. A video with no attribution renders
+ // exactly as the off arm renders it, which is the behaviour any shipped
+ // version would need for the ~98% of the corpus that can never be
+ // diarized. Silently degrading is the feature.
+ log(` ${video.slug}: no attribution on disk, rendering without speakers`);
+ }
+ }
+
+ // Chunk i starts at this index in the full cue array. chunkCuesForContext
+ // overlaps by DIGEST_OVERLAP_CUES, so this is not i * maxCues.
+ const chunkStartIndices: number[] = [];
+ {
+ let cursor = 0;
+ for (const chunk of chunks) {
+ const first = chunk[0];
+ // Cues are unique by identity here, so indexOf from the last position is
+ // both correct and linear overall.
+ const found = cues.indexOf(first, Math.max(0, cursor - chunk.length));
+ chunkStartIndices.push(found >= 0 ? found : cursor);
+ cursor = (found >= 0 ? found : cursor) + chunk.length;
+ }
+ }
+
const app = getDigestApp("ollama-direct");
const config: DigestAppConfig = {
model: candidate.model,
@@ -365,6 +644,9 @@ async function scoreVideo(
genericTitles: 0,
duplicateTitles: 0,
sampleTitles: [],
+ boundary: null,
+ speakersRendered: speakerContext?.roster.length ?? 0,
+ titlesWithSpeakerName: 0,
};
for (let i = 0; i < chunks.length; i++) {
@@ -375,6 +657,29 @@ async function scoreVideo(
Math.ceil(chunk[chunk.length - 1].end || chunk[chunk.length - 1].start),
);
const offset = candidate.timestampMode === "chunk-local" ? startSeconds : 0;
+
+ // The roster is per-CHUNK: the names that actually appear in THIS slice.
+ // A 12-name cast list over a chunk where two people speak is misleading and
+ // spends context for nothing.
+ let prefixForCue: ((cue: Cue, index: number) => string | null) | undefined;
+ let speakerRoster: string | undefined;
+ if (speakerContext) {
+ const startIndex = chunkStartIndices[i];
+ const namesHere: string[] = [];
+ const seenHere = new Set<string>();
+ for (let k = 0; k < chunk.length; k++) {
+ const label = speakerContext.labelByCueIndex[startIndex + k];
+ if (label && !seenHere.has(label)) {
+ seenHere.add(label);
+ namesHere.push(label);
+ }
+ }
+ if (namesHere.length > 0) {
+ prefixForCue = speakerPrefixer(speakerContext, startIndex);
+ speakerRoster = speakerPreamble(namesHere);
+ }
+ }
+
const promptInput = {
title: transcript.title || video.videoId,
channel: transcript.channel || video.channelSlug,
@@ -389,8 +694,10 @@ async function scoreVideo(
},
chunk,
offset,
+ prefixForCue,
),
timestampMode: candidate.timestampMode,
+ ...(speakerRoster ? { speakerRoster } : {}),
};
try {
const result = await app.run({
@@ -453,9 +760,42 @@ async function scoreVideo(
}
score.sampleTitles = parsed.chapters.slice(0, 8).map((c) => `${c.clock} ${c.title}`);
+ // BOUNDARY ACCURACY against the uploader. The one metric here that measures
+ // whether the segmentation is RIGHT rather than well-formed.
+ const oracle = await uploaderBoundaries(video);
+ if (oracle) {
+ score.boundary = scoreBoundaries(
+ oracle.starts,
+ parsed.chapters.map((c) => c.start),
+ );
+ }
+
+ // SPEAKER NAME UPTAKE. Counted against the roster the render actually used,
+ // so the off arm can be checked for leakage: a non-zero value there means a
+ // name reached the titles by some route other than the labels.
+ if (speakerContext && speakerContext.roster.length > 0) {
+ // WORD-BOUNDARY MATCHING, not substring. normalizeTitle collapses to
+ // space-separated words, so a substring test would count "ghost stories" as
+ // containing the speaker "Host" — and role labels like "Host" and "Caller"
+ // are exactly what the diarized lane produces when the transcript does not
+ // support a real name, so the false positives would not be rare.
+ const needles = speakerContext.roster
+ .map((n) => normalizeTitle(n))
+ .filter((n) => n.length >= 3)
+ .map((n) => ` ${n} `);
+ for (const c of parsed.chapters) {
+ const t = ` ${normalizeTitle(c.title)} `;
+ if (needles.some((n) => t.includes(n))) score.titlesWithSpeakerName++;
+ }
+ }
+
log(
` ${video.slug} [${video.bucket}] ${chunks.length} chunk(s) → ${score.kept} chapter(s), ` +
- `${score.zeroYieldChunks} zero-yield, ${Math.round(score.engineSeconds)}s engine`,
+ `${score.zeroYieldChunks} zero-yield, ${Math.round(score.engineSeconds)}s engine` +
+ (score.boundary
+ ? `, boundary F1@30 ${pct(score.boundary.scores[0]?.f1 ?? 0)} ` +
+ `(${score.boundary.referenceCount} uploader mark(s))`
+ : ", no uploader chapters"),
);
return score;
}
@@ -481,6 +821,38 @@ function aggregate(candidate: Candidate, videos: VideoScore[]): CandidateScore {
const secondsPerAudioHour = audioHours > 0 ? engineSeconds / audioHours : 0;
+ // Pooled boundary accuracy. See the comment on CandidateScore.totals.boundary
+ // for why this is not a mean of per-video F1s.
+ const scored = videos.filter((v) => v.boundary);
+ const referenceBoundaries = scored.reduce((a, v) => a + v.boundary!.referenceCount, 0);
+ const generatedBoundaries = scored.reduce((a, v) => a + v.boundary!.generatedCount, 0);
+ const within30 = scored.reduce((a, v) => a + v.boundary!.withinThirtySeconds, 0);
+ const allOffsets: number[] = [];
+ for (const v of scored) {
+ // medianOffsetSeconds is per video; pooling the medians is not a median, so
+ // the report quotes the median OF the per-video medians and says so.
+ if (v.boundary!.medianOffsetSeconds !== null) {
+ allOffsets.push(v.boundary!.medianOffsetSeconds);
+ }
+ }
+ allOffsets.sort((a, b) => a - b);
+
+ const byTolerance = BOUNDARY_TOLERANCES_SECONDS.map((tol) => {
+ const matched = scored.reduce(
+ (a, v) => a + (v.boundary!.scores.find((s) => s.toleranceSeconds === tol)?.matched ?? 0),
+ 0,
+ );
+ const precision = generatedBoundaries > 0 ? matched / generatedBoundaries : 0;
+ const recall = referenceBoundaries > 0 ? matched / referenceBoundaries : 0;
+ return {
+ toleranceSeconds: tol,
+ matched,
+ precision: round(precision, 4),
+ recall: round(recall, 4),
+ f1: round(precision + recall > 0 ? (2 * precision * recall) / (precision + recall) : 0, 4),
+ };
+ });
+
return {
candidate,
videos,
@@ -506,6 +878,16 @@ function aggregate(candidate: Candidate, videos: VideoScore[]): CandidateScore {
tokensPerSecond: engineSeconds > 0 ? round(tokens / engineSeconds, 1) : 0,
secondsPerAudioHour: Math.round(secondsPerAudioHour),
projectedSweepDays: 0, // filled in once corpus hours are known
+ boundary: {
+ videosScored: scored.length,
+ referenceBoundaries,
+ medianOffsetSeconds:
+ allOffsets.length > 0 ? allOffsets[Math.floor(allOffsets.length / 2)] : null,
+ withinThirtySecondsRate:
+ referenceBoundaries > 0 ? round(within30 / referenceBoundaries, 4) : 0,
+ byTolerance,
+ },
+ speakerNameTitleRate: kept > 0 ? round(sum((v) => v.titlesWithSpeakerName) / kept, 4) : 0,
},
};
}
@@ -553,6 +935,33 @@ function markdownReport(
);
}
lines.push("");
+ lines.push("## Boundary accuracy vs the uploader's own chapters");
+ lines.push("");
+ lines.push(
+ "The oracle is `metadata.info.json.chapters` — marks a human authored while",
+ "watching, who never saw our prompt. Boilerplate marks (\"Intro\", \"Sponsor\")",
+ "and the boundary at 00:00 are dropped before scoring: neither carries",
+ "segmentation information, and the origin would be a free hit for every",
+ "candidate. Matching is one-to-one and closest-pair-first, so a cluster of",
+ "boundaries around one uploader mark scores one match, not many.",
+ );
+ lines.push("");
+ lines.push(
+ "| Candidate | Videos scored | Uploader marks | Median offset | Within 30s | P@30 | R@30 | **F1@30** | F1@60 | Name-in-title |",
+ );
+ lines.push("| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |");
+ for (const s of scores) {
+ const b = s.totals.boundary;
+ const t30 = b.byTolerance.find((x) => x.toleranceSeconds === 30);
+ const t60 = b.byTolerance.find((x) => x.toleranceSeconds === 60);
+ lines.push(
+ `| \`${s.candidate.key}\` | ${b.videosScored} | ${b.referenceBoundaries} | ` +
+ `${b.medianOffsetSeconds === null ? "—" : `${b.medianOffsetSeconds}s`} | ` +
+ `${pct(b.withinThirtySecondsRate)} | ${pct(t30?.precision ?? 0)} | ${pct(t30?.recall ?? 0)} | ` +
+ `**${pct(t30?.f1 ?? 0)}** | ${pct(t60?.f1 ?? 0)} | ${pct(s.totals.speakerNameTitleRate)} |`,
+ );
+ }
+ lines.push("");
lines.push("## Rejections by guard");
lines.push("");
const codes = Array.from(
@@ -570,13 +979,18 @@ function markdownReport(
lines.push("");
lines.push("## Per-video");
lines.push("");
- lines.push("| Candidate | Video | Bucket | Chunks | Zero-yield | Chapters | Max gap | Engine s |");
- lines.push("| --- | --- | --- | --- | --- | --- | --- | --- |");
+ lines.push(
+ "| Candidate | Video | Bucket | Chunks | Zero-yield | Chapters | Max gap | Engine s | Marks | F1@30 | Speakers |",
+ );
+ lines.push("| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |");
for (const s of scores) {
for (const v of s.videos) {
+ const f1 = v.boundary?.scores.find((x) => x.toleranceSeconds === 30)?.f1;
lines.push(
`| \`${s.candidate.key}\` | \`${v.slug}\` | ${v.bucket} | ${v.chunks} | ` +
- `${v.zeroYieldChunks} | ${v.kept} | ${toHms(v.maxGapSeconds)} | ${Math.round(v.engineSeconds)} |`,
+ `${v.zeroYieldChunks} | ${v.kept} | ${toHms(v.maxGapSeconds)} | ${Math.round(v.engineSeconds)} | ` +
+ `${v.boundary?.referenceCount ?? "—"} | ${f1 === undefined ? "—" : pct(f1)} | ` +
+ `${v.speakersRendered || "—"} |`,
);
}
}
@@ -612,19 +1026,37 @@ function pct(v: number): string {
// plus room for prompt and response), so halving the context must halve the
// slice or every call silently truncates — the exact failure that made the first
// smoke test summarize a fragment.
-function parseCandidate(spec: string, mode: DigestTimestampMode): Candidate {
+function parseCandidate(
+ spec: string,
+ mode: DigestTimestampMode,
+ speakers: boolean,
+ maxCuesOverride: number | null,
+): Candidate {
const [modelPart, rest] = spec.split("@");
const [ctxPart, thinkPart] = (rest ?? "").split(":");
const numCtx = ctxPart ? Number(ctxPart) : undefined;
// The SAME derivation production uses, not a parallel copy — otherwise the
// bake-off scores a chunk size the sweep would never actually run.
- const maxCues = maxCuesForContext(numCtx);
+ //
+ // --max-cues breaks that tie deliberately, for one reason: a paired A/B needs
+ // HEADROOM. Speaker labels add ~10% to the prompt, and on a pool of long VODs
+ // the un-labelled chunk is already near the window — so at the derived size
+ // the speakers-on arm would truncate where speakers-off did not, and the
+ // measurement would be of truncation. Raising numCtx while pinning maxCues
+ // gives both arms the same cue count with room to spare. BOTH ARMS ALWAYS GET
+ // THE SAME VALUE; the report records it.
+ const maxCues = maxCuesOverride ?? maxCuesForContext(numCtx);
return {
- key: `${modelPart}@${numCtx ?? "default"}/${mode}`,
+ // The speaker axis and a cue override only show in the key when set, so
+ // existing round labels keep reading the way rounds 1-2 wrote them.
+ key:
+ `${modelPart}@${numCtx ?? "default"}/${mode}` +
+ `${maxCuesOverride ? `/${maxCuesOverride}cues` : ""}${speakers ? "/speakers" : ""}`,
model: modelPart,
...(numCtx ? { numCtx } : {}),
maxCues,
timestampMode: mode,
+ speakers,
...(thinkPart === "think"
? { think: true }
: thinkPart === "nothink"
@@ -633,16 +1065,267 @@ function parseCandidate(spec: string, mode: DigestTimestampMode): Candidate {
};
}
+// ---------------------------------------------------------------------------
+// The chapter-oracle sample
+// ---------------------------------------------------------------------------
+
+// Pick a sample restricted to videos that can actually be SCORED — i.e. that
+// carry >=`minChapters` non-boilerplate uploader marks.
+//
+// WHY IT DOES NOT READ EVERY METADATA FILE. metadata.info.json runs to ~100 KB
+// and only ~19% of videos carry chapters, so parsing all 77k to find them would
+// read several GB to answer a question a stable-hash walk answers in a few
+// hundred reads: candidates are visited in the same deterministic order
+// pickSample uses, and the walk stops as soon as every stratum is full.
+//
+// KEPT IN A SEPARATE FILE from sample.json, deliberately. Rounds 1 and 2 are
+// scored against that sample; repointing it would silently invalidate the
+// comparison this harness exists to protect.
+async function pickChapterSample(
+ outPath: string,
+ opts: { minChapters: number; perBucket: number | null; requireAudio: boolean },
+): Promise<void> {
+ const paths = getPaths();
+ const root = open({ path: paths.lmdbPath, maxDbs: 12, compression: true });
+ const statsByPath = root.openDB<
+ { metaMs: number; stat: VideoStat },
+ [string, string]
+ >({ name: "statsByPath", encoding: "msgpack" });
+
+ const byBucket = new Map<Bucket, SampleVideo[]>();
+ for (const b of BUCKETS) byBucket.set(b, []);
+
+ let videosScanned = 0;
+ let videosWithTranscript = 0;
+ let totalSeconds = 0;
+ let longTailVideos = 0;
+ let longTailSeconds = 0;
+
+ for (const { key, value } of statsByPath.getRange()) {
+ const stat = value.stat;
+ videosScanned++;
+ if (!stat.hasTranscript || !(stat.duration > 0)) continue;
+ videosWithTranscript++;
+ totalSeconds += stat.duration;
+ if (stat.duration > 4 * 3600) {
+ longTailVideos++;
+ longTailSeconds += stat.duration;
+ }
+ const bucket = bucketFor(stat.duration);
+ if (!bucket) continue;
+ if (!stat.cueCount || stat.cueCount < 30) continue;
+ byBucket.get(bucket)!.push({
+ slug: stat.slug,
+ channelSlug: stat.channelSlug,
+ videoId: stat.id,
+ videoDir: (key as [string, string])[1],
+ title: stat.title,
+ bucket,
+ durationSeconds: Math.round(stat.duration),
+ cueCount: stat.cueCount,
+ });
+ }
+ await root.close();
+
+ const videos: SampleVideo[] = [];
+ for (const b of BUCKETS) {
+ const pool = byBucket.get(b)!;
+ pool.sort((a, c) => stableHash(a.slug) - stableHash(c.slug));
+ const want = opts.perBucket ?? BUCKET_BOUNDS[b].want;
+ const seenChannels = new Set<string>();
+ const taken: SampleVideo[] = [];
+ let probed = 0;
+ for (const v of pool) {
+ if (taken.length >= want) break;
+ // One per channel first, for the same reason pickSample does it: a
+ // stratum drawn from one creator measures a house style.
+ if (seenChannels.has(v.channelSlug)) continue;
+ probed++;
+ const oracle = await uploaderBoundaries(v);
+ if (!oracle || oracle.starts.length < opts.minChapters) continue;
+ if (opts.requireAudio) {
+ // isRealAudioFile, not an extension test of my own: it is the same
+ // predicate the cleanup and diarization lanes use, so "has audio" means
+ // here exactly what it means to the job that would diarize it.
+ const files = await readdir(videoDirFor(v)).catch(() => [] as string[]);
+ if (!files.some((f) => isRealAudioFile(f))) continue;
+ }
+ seenChannels.add(v.channelSlug);
+ taken.push({ ...v, uploaderChapters: oracle.starts.length });
+ }
+ if (taken.length < want) {
+ console.warn(
+ `Warning: bucket ${b} wanted ${want} scorable videos but only ${taken.length} qualify ` +
+ `(probed ${probed} candidates).`,
+ );
+ }
+ videos.push(...taken);
+ }
+
+ const sample: Sample = {
+ version: 1,
+ pickedAt: new Date().toISOString(),
+ corpus: {
+ videosScanned,
+ videosWithTranscript,
+ audioHours: Math.round(totalSeconds / 3600),
+ longTailVideos,
+ longTailAudioHours: Math.round(longTailSeconds / 3600),
+ },
+ videos,
+ };
+ await mkdir(path.dirname(outPath), { recursive: true });
+ await writeFile(outPath, `${JSON.stringify(sample, null, 2)}\n`);
+ for (const v of videos) {
+ console.log(
+ ` ${v.bucket.padEnd(8)} ${toHms(v.durationSeconds)} ${String(v.uploaderChapters).padStart(3)} marks ` +
+ `${v.slug} ${v.title.slice(0, 55)}`,
+ );
+ }
+ console.log(`Wrote ${outPath} (${videos.length} scorable video(s))`);
+}
+
+// ---------------------------------------------------------------------------
+// Free scorer smoke test
+// ---------------------------------------------------------------------------
+
+// Score the digests ALREADY on disk against their uploader chapters.
+//
+// This is the cheapest honest check available: no model call, no write, and it
+// answers "does the metric move, and is the plumbing right" before a generation
+// run is spent finding out. If this prints nonsense — every F1 at 0 or 1 — the
+// scorer is wrong and no A/B built on it would mean anything.
+async function scoreExisting(minChapters: number, limit: number): Promise<void> {
+ const paths = getPaths();
+ const root = open({ path: paths.lmdbPath, maxDbs: 12, compression: true });
+ const statsByPath = root.openDB<
+ { metaMs: number; stat: VideoStat },
+ [string, string]
+ >({ name: "statsByPath", encoding: "msgpack" });
+
+ const candidates: SampleVideo[] = [];
+ for (const { key, value } of statsByPath.getRange()) {
+ const stat = value.stat;
+ if (!stat.hasTranscript || !(stat.duration > 0)) continue;
+ candidates.push({
+ slug: stat.slug,
+ channelSlug: stat.channelSlug,
+ videoId: stat.id,
+ videoDir: (key as [string, string])[1],
+ title: stat.title,
+ bucket: bucketFor(stat.duration) ?? "short",
+ durationSeconds: Math.round(stat.duration),
+ cueCount: stat.cueCount ?? 0,
+ });
+ }
+ await root.close();
+
+ const rows: {
+ slug: string;
+ marks: number;
+ generated: number;
+ medianOffset: number | null;
+ precision30: number;
+ recall30: number;
+ recall60: number;
+ f1at30: number;
+ f1at60: number;
+ }[] = [];
+
+ for (const v of candidates) {
+ if (rows.length >= limit) break;
+ const dir = videoDirFor(v);
+ const digest = await loadDigest(dir);
+ const items = digest?.sections?.chapters?.items ?? [];
+ if (items.length === 0) continue;
+ const oracle = await uploaderBoundaries(v);
+ if (!oracle || oracle.starts.length < minChapters) continue;
+ const report = scoreBoundaries(oracle.starts, items.map((c) => c.start));
+ if (report.referenceCount === 0) continue;
+ const s30 = report.scores.find((s) => s.toleranceSeconds === 30);
+ const s60 = report.scores.find((s) => s.toleranceSeconds === 60);
+ rows.push({
+ slug: v.slug,
+ marks: report.referenceCount,
+ generated: report.generatedCount,
+ medianOffset: report.medianOffsetSeconds,
+ precision30: s30?.precision ?? 0,
+ recall30: s30?.recall ?? 0,
+ recall60: s60?.recall ?? 0,
+ f1at30: s30?.f1 ?? 0,
+ f1at60: s60?.f1 ?? 0,
+ });
+ }
+
+ if (rows.length === 0) {
+ console.log(
+ `No video on disk has both a digest and >=${minChapters} non-boilerplate uploader chapters.`,
+ );
+ return;
+ }
+
+ console.log(
+ `Scored ${rows.length} existing digest(s) against uploader chapters (no model calls, no writes).\n`,
+ );
+ console.log(
+ " slug marks gen medOff P@30 R@30 R@60 F1@30 F1@60",
+ );
+ for (const r of rows) {
+ console.log(
+ ` ${r.slug.padEnd(36).slice(0, 36)} ${String(r.marks).padStart(5)} ` +
+ `${String(r.generated).padStart(3)} ${String(r.medianOffset ?? "—").padStart(6)} ` +
+ `${pct(r.precision30).padStart(5)} ${pct(r.recall30).padStart(5)} ` +
+ `${pct(r.recall60).padStart(5)} ${pct(r.f1at30).padStart(5)} ${pct(r.f1at60).padStart(5)}`,
+ );
+ }
+ const sumOf = (f: (r: (typeof rows)[number]) => number): number =>
+ rows.reduce((a, r) => a + f(r), 0);
+ // POOLED, not a mean of per-video rates: a 3-mark video must not weigh the
+ // same as a 29-mark one.
+ const marks = sumOf((r) => r.marks);
+ const generated = sumOf((r) => r.generated);
+ const matched30 = sumOf((r) => r.recall30 * r.marks);
+ const matched60 = sumOf((r) => r.recall60 * r.marks);
+ const p30 = generated > 0 ? matched30 / generated : 0;
+ const r30 = marks > 0 ? matched30 / marks : 0;
+ console.log(
+ `\n POOLED: ${marks} uploader mark(s), ${generated} generated. ` +
+ `P@30 ${pct(p30)}, R@30 ${pct(r30)}, R@60 ${pct(marks > 0 ? matched60 / marks : 0)}, ` +
+ `F1@30 ${pct(p30 + r30 > 0 ? (2 * p30 * r30) / (p30 + r30) : 0)}`,
+ );
+ console.log(
+ ` NOTE: the digest targets one chapter per ${DIGEST_MINUTES_PER_CHAPTER} minutes and so is\n` +
+ ` deliberately DENSER than uploader chaptering (${generated} vs ${marks} here). Precision is\n` +
+ ` therefore partly a density artifact; recall is the half that answers "did it find the\n` +
+ ` human's boundaries". Compare candidates on recall AND on chapters/h together.`,
+ );
+}
+
async function main(): Promise<void> {
const flags = parseFlags(process.argv.slice(2));
const outDir = flags.outDir ?? path.join(process.cwd(), "..", "plans", "bakeoff");
const samplePath = flags.sample ?? path.join(outDir, "sample.json");
+ const minChapters = flags["min-chapters"] ? Number(flags["min-chapters"]) : 4;
if (flags.pick === "true") {
await pickSample(samplePath);
return;
}
+ if (flags["pick-chapters"] === "true") {
+ await pickChapterSample(samplePath, {
+ minChapters,
+ perBucket: flags["per-bucket"] ? Number(flags["per-bucket"]) : null,
+ requireAudio: flags["require-audio"] === "true",
+ });
+ return;
+ }
+
+ if (flags["score-existing"] === "true") {
+ await scoreExisting(minChapters, flags.limit ? Number(flags.limit) : 200);
+ return;
+ }
+
const sample = JSON.parse(await readFile(samplePath, "utf8")) as Sample;
const label = flags.label ?? "round";
const buckets = (flags.buckets ?? BUCKETS.join(","))
@@ -657,13 +1340,25 @@ async function main(): Promise<void> {
.split(",")
.map((s) => s.trim())
.filter(Boolean);
+ // The paired A/B axis. Default "off" — the same rendering every earlier round
+ // used, so omitting the flag reproduces them.
+ const speakerArms = (flags.speakers ?? "off")
+ .split(",")
+ .map((s) => s.trim())
+ .filter((s) => s === "on" || s === "off")
+ .map((s) => s === "on");
const videos = sample.videos.filter((v) => buckets.includes(v.bucket));
if (videos.length === 0) throw new Error(`No sample videos in buckets ${buckets.join(",")}`);
+ const maxCuesOverride = flags["max-cues"] ? Number(flags["max-cues"]) : null;
const candidates: Candidate[] = [];
for (const spec of specs) {
- for (const mode of modes) candidates.push(parseCandidate(spec, mode));
+ for (const mode of modes) {
+ for (const speakers of speakerArms) {
+ candidates.push(parseCandidate(spec, mode, speakers, maxCuesOverride));
+ }
+ }
}
console.log(
@@ -671,6 +1366,10 @@ async function main(): Promise<void> {
`(${Math.round(videos.reduce((a, v) => a + v.durationSeconds, 0) / 3600)} audio-hours each).`,
);
+ // Up front, not just before the final write, so the per-video checkpoint below
+ // has somewhere to land from the very first video.
+ await mkdir(outDir, { recursive: true });
+
const scores: CandidateScore[] = [];
for (const candidate of candidates) {
console.log(`\n=== ${candidate.key} (maxCues ${candidate.maxCues}) ===`);
@@ -679,6 +1378,16 @@ async function main(): Promise<void> {
for (const video of videos) {
const s = await scoreVideo(video, candidate, (m) => console.log(m));
if (s) perVideo.push(s);
+ // CHECKPOINT AFTER EVERY VIDEO, because the reports are only written when
+ // the whole run finishes and a long run does not reliably get there. A
+ // 12-video run was OOM-killed on its last video after 26 minutes of engine
+ // time and left NOTHING behind — no error, no partial, just a dead process
+ // (the kill is a SIGKILL, so no handler can save it either). On a box that
+ // swaps, "it completed 11 of 12" has to survive.
+ await writeFile(
+ path.join(outDir, `${label}.partial.json`),
+ `${JSON.stringify({ label, candidate: candidate.key, videos: perVideo }, null, 2)}\n`,
+ ).catch(() => {});
}
const agg = aggregate(candidate, perVideo);
// Days, from measured seconds-per-audio-hour against the corpus total
diff --git a/common/lib/boundaryScore.test.ts b/common/lib/boundaryScore.test.ts
@@ -0,0 +1,111 @@
+import { test } from "node:test";
+import assert from "node:assert/strict";
+import {
+ BOUNDARY_ORIGIN_EPSILON_SECONDS,
+ isBoilerplateChapterTitle,
+ matchBoundaries,
+ normalizeBoundaries,
+ scoreBoundaries,
+ scoreBoundariesAt,
+} from "./boundaryScore";
+
+test("the origin boundary is dropped, because everything has one", () => {
+ // A boundary at t=0 is free for any segmentation and carries no information
+ // about segmentation skill. If it survived, both arms of an A/B would collect
+ // an unearned true positive.
+ assert.deepEqual(normalizeBoundaries([0, 120, 300]), [120, 300]);
+ assert.deepEqual(normalizeBoundaries([BOUNDARY_ORIGIN_EPSILON_SECONDS, 120]), [120]);
+ assert.deepEqual(normalizeBoundaries([BOUNDARY_ORIGIN_EPSILON_SECONDS + 1, 120]), [6, 120]);
+});
+
+test("normalization sorts, rounds and de-dups", () => {
+ assert.deepEqual(normalizeBoundaries([300.4, 120.6, 300.2, NaN, 121]), [121, 300]);
+});
+
+test("a perfect segmentation scores 1.0 at every tolerance", () => {
+ const truth = [120, 300, 600];
+ const report = scoreBoundaries(truth, truth);
+ for (const s of report.scores) {
+ assert.equal(s.f1, 1);
+ assert.equal(s.precision, 1);
+ assert.equal(s.recall, 1);
+ }
+ assert.equal(report.medianOffsetSeconds, 0);
+ assert.equal(report.withinThirtySeconds, 3);
+});
+
+test("a disjoint segmentation scores 0", () => {
+ const report = scoreBoundaries([120, 300], [5000, 6000]);
+ for (const s of report.scores) {
+ assert.equal(s.matched, 0);
+ assert.equal(s.f1, 0);
+ }
+ assert.equal(report.withinThirtySeconds, 0);
+});
+
+test("matching is ONE-TO-ONE, so clustering is not rewarded", () => {
+ // The over-segmentation failure precision exists to punish: ten boundaries
+ // crammed into one 30 s window around a single uploader boundary. A
+ // nearest-neighbour count would score this 10 matches and look excellent.
+ const generated = [115, 116, 117, 118, 119, 120, 121, 122, 123, 124];
+ const score = scoreBoundariesAt([120], generated, 30);
+ assert.equal(score.matched, 1, "one reference boundary can absorb only one match");
+ assert.equal(score.recall, 1);
+ assert.equal(score.precision, 1 / 10);
+ assert.ok(score.f1 < 0.2);
+});
+
+test("closest pair wins when two references contend for one boundary", () => {
+ // 300 is 5 s away, 280 is 15 s away. Greedy-by-distance must give the
+ // generated boundary to 300 and leave 280 unmatched, not the reverse.
+ const matches = matchBoundaries([280, 300], [295], 30);
+ assert.equal(matches.length, 1);
+ assert.deepEqual(matches[0], { reference: 300, generated: 295, distance: 5 });
+});
+
+test("tolerance is a hard edge, and the two tolerances can disagree", () => {
+ // 45 s off: missed at 30 s, found at 60 s. A variant that trades tightness
+ // for recall has to be visible as exactly that.
+ const report = scoreBoundaries([300], [345]);
+ const at30 = report.scores.find((s) => s.toleranceSeconds === 30)!;
+ const at60 = report.scores.find((s) => s.toleranceSeconds === 60)!;
+ assert.equal(at30.matched, 0);
+ assert.equal(at60.matched, 1);
+ assert.equal(report.medianOffsetSeconds, 45);
+});
+
+test("offset is nearest-neighbour and F1 is not, so they fail differently", () => {
+ // One generated boundary sitting near one of three references: a good median
+ // offset would be misleading on its own, and precision/recall is what catches
+ // it. Both are reported for this reason.
+ const report = scoreBoundaries([300, 900, 1500], [305]);
+ assert.equal(report.medianOffsetSeconds, 595);
+ const at30 = report.scores.find((s) => s.toleranceSeconds === 30)!;
+ assert.equal(at30.precision, 1);
+ assert.equal(at30.recall, 1 / 3);
+});
+
+test("an empty generated set scores zero rather than dividing by zero", () => {
+ const report = scoreBoundaries([120, 300], []);
+ for (const s of report.scores) {
+ assert.equal(s.precision, 0);
+ assert.equal(s.recall, 0);
+ assert.equal(s.f1, 0);
+ }
+ assert.equal(report.medianOffsetSeconds, null);
+});
+
+test("a video whose only uploader chapter is the origin yields no reference", () => {
+ const report = scoreBoundaries([0], [120, 300]);
+ assert.equal(report.referenceCount, 0);
+ assert.equal(report.medianOffsetSeconds, null);
+});
+
+test("boilerplate chapter titles are recognized", () => {
+ for (const t of ["Intro", "intro", "Outro", "Sponsor", "Ad read", "Merch", "BRB"]) {
+ assert.ok(isBoilerplateChapterTitle(t), `${t} should be boilerplate`);
+ }
+ for (const t of ["Court filing deadlines", "Introducing the new tariff rules"]) {
+ assert.ok(!isBoilerplateChapterTitle(t), `${t} should NOT be boilerplate`);
+ }
+});
diff --git a/common/lib/boundaryScore.ts b/common/lib/boundaryScore.ts
@@ -0,0 +1,208 @@
+// Does a generated chapter boundary land where a human put one?
+//
+// WHY THIS EXISTS. Every metric digest-bakeoff.ts scored before this one —
+// zeroYieldRate, chaptersPerHour, rejectionRate, maxGapSeconds,
+// genericTitleRate, duplicateTitleRate — is a DEFECT COUNTER. Each answers "how
+// broken is the output", none answers "is the segmentation right". So the
+// bake-off could rank candidates by which was least malformed and still could
+// not say which produced better chapters. That is the gap this closes.
+//
+// THE ORACLE IS THE UPLOADER'S OWN CHAPTERS, and the choice is deliberate on two
+// grounds. First, independence: scoring "chapter starts align with speaker
+// changes" against a variant that was GIVEN the speaker changes proves nothing,
+// so the ground truth has to come from outside the experiment. Uploader chapters
+// are authored by a human who watched the video and never saw our prompt.
+// Second, posture: a boundary is a FACT about where a subject changes, not
+// expression. Scoring against timestamps reproduces nothing, which is why this
+// is safe where admitting the uploader's chapter TITLES into the corpus would
+// not be (see PLAN.md on why the archive publishes new expression about works).
+//
+// Pure — no I/O, no clock — so it is unit-testable and can be reused by any
+// future digest change, which is worth more than the one experiment that
+// prompted it.
+
+// The reference boundary at t=0 is dropped before scoring, always.
+//
+// Nearly every uploader chapter list opens with a chapter at 00:00 ("Intro"),
+// and every segmentation trivially has a boundary at the start of the video.
+// Counting it would hand both arms of any A/B a free true-positive that carries
+// no information about segmentation skill, inflating precision, recall and F1 by
+// roughly 1/n on short chapter lists. A metric that cannot be lost is not a
+// measurement.
+export const BOUNDARY_ORIGIN_EPSILON_SECONDS = 5;
+
+// The tolerances every report quotes. 30 s is "the model found the same moment";
+// 60 s is "the model found the same transition, a little late". Both are
+// reported because a variant that trades tightness for recall should be visible
+// as exactly that, not averaged into one number.
+export const BOUNDARY_TOLERANCES_SECONDS = [30, 60] as const;
+
+export type BoundaryMatch = {
+ reference: number;
+ generated: number;
+ distance: number;
+};
+
+export type BoundaryScore = {
+ toleranceSeconds: number;
+ referenceCount: number;
+ generatedCount: number;
+ matched: number;
+ // matched / generatedCount — of the boundaries we emitted, how many a human
+ // agrees with. Low precision means over-segmentation.
+ precision: number;
+ // matched / referenceCount — of the boundaries a human drew, how many we
+ // found. Low recall means we missed real transitions.
+ recall: number;
+ f1: number;
+ matches: BoundaryMatch[];
+};
+
+export type BoundaryReport = {
+ referenceCount: number;
+ generatedCount: number;
+ // Per reference boundary, the distance to the NEAREST generated boundary,
+ // independent of any matching. This is the "how far off were we" view, and it
+ // is reported alongside F1 because the two fail differently: a model that
+ // emits one boundary 5 s from every reference scores a perfect median offset
+ // and a terrible precision.
+ medianOffsetSeconds: number | null;
+ meanOffsetSeconds: number | null;
+ withinThirtySeconds: number;
+ scores: BoundaryScore[];
+};
+
+function median(values: number[]): number | null {
+ if (values.length === 0) return null;
+ const sorted = [...values].sort((a, b) => a - b);
+ const mid = Math.floor(sorted.length / 2);
+ return sorted.length % 2 === 0
+ ? (sorted[mid - 1] + sorted[mid]) / 2
+ : sorted[mid];
+}
+
+// Drop the origin boundary and any duplicate/unsorted noise, so callers can pass
+// raw starts from either side without pre-cleaning them.
+export function normalizeBoundaries(
+ starts: readonly number[],
+ originEpsilon = BOUNDARY_ORIGIN_EPSILON_SECONDS,
+): number[] {
+ const seen = new Set<number>();
+ const out: number[] = [];
+ for (const raw of starts) {
+ if (!Number.isFinite(raw)) continue;
+ const s = Math.round(raw);
+ if (s <= originEpsilon) continue;
+ if (seen.has(s)) continue;
+ seen.add(s);
+ out.push(s);
+ }
+ return out.sort((a, b) => a - b);
+}
+
+// Greedy one-to-one matching, closest pair first.
+//
+// ONE-TO-ONE IS THE POINT. A model that emits ten boundaries clustered inside
+// one 30 s window must not be credited with ten matches against a single
+// uploader boundary — that is the exact over-segmentation failure precision is
+// supposed to punish, and a nearest-neighbour count would reward it instead.
+// Closest-first (rather than left-to-right) keeps the pairing stable when two
+// references sit within one tolerance of the same generated boundary.
+export function matchBoundaries(
+ reference: readonly number[],
+ generated: readonly number[],
+ toleranceSeconds: number,
+): BoundaryMatch[] {
+ const pairs: BoundaryMatch[] = [];
+ for (const r of reference) {
+ for (const g of generated) {
+ const distance = Math.abs(r - g);
+ if (distance <= toleranceSeconds) pairs.push({ reference: r, generated: g, distance });
+ }
+ }
+ // Ties broken by position so the result is deterministic across runs.
+ pairs.sort(
+ (a, b) =>
+ a.distance - b.distance || a.reference - b.reference || a.generated - b.generated,
+ );
+ const usedReference = new Set<number>();
+ const usedGenerated = new Set<number>();
+ const matches: BoundaryMatch[] = [];
+ for (const p of pairs) {
+ if (usedReference.has(p.reference) || usedGenerated.has(p.generated)) continue;
+ usedReference.add(p.reference);
+ usedGenerated.add(p.generated);
+ matches.push(p);
+ }
+ return matches.sort((a, b) => a.reference - b.reference);
+}
+
+export function scoreBoundariesAt(
+ reference: readonly number[],
+ generated: readonly number[],
+ toleranceSeconds: number,
+): BoundaryScore {
+ const matches = matchBoundaries(reference, generated, toleranceSeconds);
+ const matched = matches.length;
+ const precision = generated.length > 0 ? matched / generated.length : 0;
+ const recall = reference.length > 0 ? matched / reference.length : 0;
+ const f1 = precision + recall > 0 ? (2 * precision * recall) / (precision + recall) : 0;
+ return {
+ toleranceSeconds,
+ referenceCount: reference.length,
+ generatedCount: generated.length,
+ matched,
+ precision,
+ recall,
+ f1,
+ matches,
+ };
+}
+
+// The whole picture for one video. `reference` and `generated` are raw starts in
+// seconds; normalization (origin drop, de-dup, sort) happens here so every
+// caller gets the same treatment.
+export function scoreBoundaries(
+ referenceRaw: readonly number[],
+ generatedRaw: readonly number[],
+ tolerances: readonly number[] = BOUNDARY_TOLERANCES_SECONDS,
+): BoundaryReport {
+ const reference = normalizeBoundaries(referenceRaw);
+ const generated = normalizeBoundaries(generatedRaw);
+
+ const offsets: number[] = [];
+ for (const r of reference) {
+ let best = Infinity;
+ for (const g of generated) best = Math.min(best, Math.abs(r - g));
+ if (Number.isFinite(best)) offsets.push(best);
+ }
+
+ return {
+ referenceCount: reference.length,
+ generatedCount: generated.length,
+ medianOffsetSeconds: median(offsets),
+ meanOffsetSeconds:
+ offsets.length > 0
+ ? offsets.reduce((a, b) => a + b, 0) / offsets.length
+ : null,
+ withinThirtySeconds: offsets.filter((d) => d <= 30).length,
+ scores: tolerances.map((t) => scoreBoundariesAt(reference, generated, t)),
+ };
+}
+
+// ---------------------------------------------------------------------------
+// Uploader chapter titles: usable as boilerplate, never as a quality reference
+// ---------------------------------------------------------------------------
+
+// Uploader chapter lists are dominated by structural marks — the modal first
+// chapter across this corpus is literally "Intro". They are therefore NOT a
+// reference for title quality (scoring our titles against them would score us
+// against boilerplate), but the marks themselves are worth recognizing because a
+// boundary at "Sponsor" is a boundary in the AD READ, not in the subject, and
+// crediting a model for finding it measures the wrong thing.
+const BOILERPLATE_TITLE_RE =
+ /^(intro(duction)?|outro|start|beginning|end(ing)?|sponsor(ed)?( segment| message)?|ad( read| break)?|advert(isement)?|thanks?( for watching)?|subscribe|like and subscribe|patreon|merch|announcements?|housekeeping|stream starts?( soon)?|waiting( screen)?|brb|be right back|credits|outro music|music)\b/i;
+
+export function isBoilerplateChapterTitle(title: string): boolean {
+ return BOILERPLATE_TITLE_RE.test(title.trim());
+}
diff --git a/common/lib/digestPrompt.test.ts b/common/lib/digestPrompt.test.ts
@@ -117,6 +117,48 @@ test("absolute mode renders byte-identically to the pre-timestampMode prompt", (
);
});
+test("a prompt with no speakerRoster renders byte-identically to today", () => {
+ // THE SAME COMPATIBILITY ARGUMENT AS absolute-mode ABOVE, and the reason
+ // adding the roster field needed no PROMPT_VERSION bump: the shipped lane
+ // never sets it, so every digest the sweep produces is byte-for-byte what it
+ // produced before the field existed. It is also what makes the bake-off's
+ // speakers-OFF arm a real control rather than a third variant.
+ const base = {
+ title: "T",
+ channel: "C",
+ startSeconds: 100,
+ endSeconds: 700,
+ transcript: "[00:01:40] hello",
+ };
+ const shipped = buildChapterPrompt(base);
+ assert.equal(buildChapterPrompt({ ...base, speakerRoster: undefined }), shipped);
+ // Whitespace-only is treated as absent, so a caller that builds an empty
+ // roster string cannot accidentally change the prompt.
+ assert.equal(buildChapterPrompt({ ...base, speakerRoster: " \n" }), shipped);
+ assert.equal(buildChapterPrompt({ ...base, speakerRoster: "" }), shipped);
+});
+
+test("a speakerRoster is stated separately from the channel context note", () => {
+ const base = {
+ title: "T",
+ channel: "C",
+ startSeconds: 100,
+ endSeconds: 700,
+ transcript: "[00:01:40] hello",
+ contextNote: "The host is Alice.",
+ speakerRoster: "Speakers in this section:\n- Alice\n- Bob",
+ };
+ const prompt = buildChapterPrompt(base);
+ // Both present, and the roster is NOT filed under "context for this channel"
+ // — it is per-chunk and derived, not hand-authored per channel.
+ assert.match(prompt, /Context for this channel[^\n]*\nThe host is Alice\./);
+ assert.match(prompt, /Speakers in this section:\n- Alice\n- Bob/);
+ // The guard against titling chapters after whoever is speaking.
+ assert.match(prompt, /A title still names the SUBJECT under discussion/);
+ // The transcript still comes last, so the body is what the model ends on.
+ assert.ok(prompt.indexOf("Speakers in this section") < prompt.indexOf("Transcript section:"));
+});
+
test("promptOffsetSeconds is zero in absolute mode by construction", () => {
assert.equal(promptOffsetSeconds({ startSeconds: 5503 }), 0);
assert.equal(
diff --git a/common/lib/digestPrompt.ts b/common/lib/digestPrompt.ts
@@ -141,6 +141,21 @@ export type ChapterPromptInput = {
// Optional per-channel context note (Phase 1.5). Plumbed from the start so
// adding notes later doesn't invalidate the corpus — see contextHash.
contextNote?: string;
+ // Optional cast list for THIS chunk, when the caller rendered speaker labels
+ // into `transcript`. Experimental: set by the bake-off's speakers-on arm only,
+ // never by the shipped lane.
+ //
+ // SEPARATE FROM contextNote ON PURPOSE, though both carry names. contextNote
+ // is per-CHANNEL, hand-authored, and hashed into contextHash for freshness; a
+ // roster is per-CHUNK and derived from that video's attribution.json. Folding
+ // the roster into contextNote would mislabel it to the model ("context for
+ // this channel"), and would make a channel note and a roster mutually
+ // exclusive.
+ //
+ // ABSENT MUST RENDER BYTE-IDENTICALLY, which is what keeps this addition free
+ // of a PROMPT_VERSION bump — the same argument that let timestampMode in. See
+ // digestPrompt.test.ts.
+ speakerRoster?: string;
// Which numbering the caller rendered `transcript` with. Defaults to
// "absolute". Under "chunk-local" the caller has re-based every marker to
// 00:00:00, so the range this prompt states must be re-based to match — a
@@ -229,6 +244,17 @@ export function buildChapterPrompt(input: ChapterPromptInput): string {
lines.push("Context for this channel (use it for names and recurring topics):");
lines.push(input.contextNote.trim());
}
+ if (input.speakerRoster?.trim()) {
+ lines.push("");
+ lines.push(input.speakerRoster.trim());
+ // Without this the model titles sections after whoever is talking
+ // ("Erica Kirk responds"), which is a worse table of contents than the
+ // generic titles it replaces. A title names a SUBJECT.
+ lines.push(
+ "Use the speakers to tell apart who is talking and to spell names correctly.",
+ "A title still names the SUBJECT under discussion, never the speaker.",
+ );
+ }
lines.push("");
lines.push("Transcript section:");
lines.push("");
diff --git a/common/lib/transcriptToMarkdown.test.ts b/common/lib/transcriptToMarkdown.test.ts
@@ -78,3 +78,53 @@ test("extraMeta lines render as `- <line>` in the metadata header", () => {
"extra meta sits in the header, before the body",
);
});
+
+// ---------------------------------------------------------------------------
+// prefixForCue — the speaker-label hook
+// ---------------------------------------------------------------------------
+
+test("prefixForCue returning null for every cue is byte-identical to omitting it", () => {
+ // THE GUARANTEE THAT LETS A SPEAKER-AWARE CALLER SHARE THIS RENDERER. If a
+ // no-op prefix changed a single byte, the speakers-off arm of an A/B would
+ // stop being comparable to production output and the measurement would be
+ // of the renderer, not of the speakers.
+ const plain = transcriptToMarkdown(base);
+ assert.equal(transcriptToMarkdown(base, { prefixForCue: () => null }), plain);
+ assert.equal(transcriptToMarkdown(base, { prefixForCue: () => "" }), plain);
+});
+
+test("prefixForCue sits OUTSIDE the timestamp bracket", () => {
+ // A speaker name inside the bracket would break the digest prompt's contract
+ // that every [HH:MM:SS] marker is copyable verbatim, and HMS_PATTERN would
+ // reject the echo.
+ const md = transcriptToMarkdown(base, {
+ prefixForCue: (cue) => (cue.start === 0 ? "Alice: " : null),
+ });
+ assert.match(md, /\[0:00\] Alice: Hello there/);
+ assert.match(md, /\[1:01:01\] General Kenobi/);
+});
+
+test("prefixForCue composes with stampForCue and with timestamps off", () => {
+ const stamped = transcriptToMarkdown(base, {
+ stampForCue: (clock) => `${clock}|x`,
+ prefixForCue: () => "Bob: ",
+ });
+ assert.match(stamped, /\[0:00\|x\] Bob: Hello there/);
+
+ const bare = transcriptToMarkdown(base, {
+ timestamps: false,
+ prefixForCue: () => "Bob: ",
+ });
+ assert.match(bare, /^Bob: Hello there$/m);
+});
+
+test("prefixForCue receives the cue index of emitted lines", () => {
+ const seen: number[] = [];
+ transcriptToMarkdown(base, {
+ prefixForCue: (_cue, i) => {
+ seen.push(i);
+ return null;
+ },
+ });
+ assert.deepEqual(seen, [0, 1]);
+});
diff --git a/common/lib/transcriptToMarkdown.ts b/common/lib/transcriptToMarkdown.ts
@@ -44,6 +44,19 @@ export type TranscriptMarkdownOptions = {
// Extra metadata lines rendered as `- <line>` after the tags line (e.g.
// "moment_base: <url>").
extraMeta?: string[];
+ // Optional per-line prefix inserted between the timestamp and the cue text
+ // (`[0:12] Mr. Obvious: and that is ...`). Return null for "no prefix on this
+ // line" — returning null for EVERY cue reproduces the un-prefixed line byte
+ // for byte, which is what lets a speaker-aware caller share this renderer
+ // instead of forking a second one. Pinned by a test in
+ // transcriptToMarkdown.test.ts.
+ //
+ // DELIBERATELY NOT FOLDED INTO stampForCue. That formats the BRACKET's
+ // contents, and the digest prompt instructs the model to copy the [HH:MM:SS]
+ // markers verbatim while HMS_PATTERN rejects anything else — a speaker name
+ // smuggled inside the bracket would make every start the model echoed
+ // unparseable. The prefix has to live outside the bracket.
+ prefixForCue?: (cue: Cue, index: number) => string | null;
};
// [h:mm:ss] / [m:ss] label for a cue start. formatDuration returns "" for 0, so
@@ -68,6 +81,7 @@ export function transcriptToMarkdown(
maxCues,
linkForCue,
stampForCue,
+ prefixForCue,
extraMeta,
} = options;
@@ -112,8 +126,12 @@ export function transcriptToMarkdown(
: cues.length;
for (let i = 0; i < limit; i++) {
const cue = cues[i];
- const text = cue.text.trim();
- if (!text) continue;
+ const rawText = cue.text.trim();
+ if (!rawText) continue;
+ // Empty string and null both mean "nothing here", so a caller can return
+ // either without changing the output.
+ const prefix = prefixForCue ? prefixForCue(cue, i) : null;
+ const text = prefix ? `${prefix}${rawText}` : rawText;
if (!timestamps) {
lines.push(text);
continue;
diff --git a/common/lib/videoStatus.ts b/common/lib/videoStatus.ts
@@ -252,3 +252,61 @@ export async function readVideoDurationSec(
return null;
}
}
+
+// One chapter as the UPLOADER authored it. Facts (`start`, `end`) and expression
+// (`title`) in one record, and callers are expected to treat them differently —
+// see readUploaderChapters.
+export type UploaderChapter = {
+ start: number;
+ end: number | null;
+ title: string;
+};
+
+// The uploader's own chapter marks from metadata.info.json, sorted by start, or
+// null when there is no metadata, it will not parse, or it carries no chapters.
+//
+// ALREADY ON DISK FOR THE WHOLE ARCHIVE, AND READ BY NOTHING ELSE. yt-dlp writes
+// `chapters` alongside `duration`, so this is a free signal for the ~19% of
+// videos that have it (measured: 14,923 video dirs carry a non-empty array,
+// 12,139 of those with >=4 chapters and a transcript). Nothing in the codebase
+// consumed it before boundary scoring did.
+//
+// THE TIMESTAMPS AND THE TITLES HAVE DIFFERENT STANDING. A chapter start is a
+// fact about where a subject changes; a chapter title is the uploader's own
+// expression. Boundary scoring uses the starts as ground truth, which reproduces
+// nothing. Admitting the titles into the corpus as our own chapter text would
+// republish the uploader's copy and is rejected on the grounds PLAN.md sets out.
+// Callers that touch `title` should be doing something like boilerplate
+// detection, not authorship.
+//
+// Same cost warning as readVideoDurationSec, more so: metadata.info.json runs to
+// ~100 KB and this parses all of it. Narrow the population with a cheap check
+// first — never map this over the whole archive in a render path.
+export async function readUploaderChapters(
+ videoDir: string,
+): Promise<UploaderChapter[] | null> {
+ try {
+ const raw = await readFile(path.join(videoDir, META_FILENAME), "utf8");
+ const meta = JSON.parse(raw) as { chapters?: unknown };
+ if (!Array.isArray(meta?.chapters) || meta.chapters.length === 0) return null;
+ const chapters: UploaderChapter[] = [];
+ for (const entry of meta.chapters) {
+ const c = entry as { start_time?: unknown; end_time?: unknown; title?: unknown };
+ // start_time is the only required field: a chapter with no start is not a
+ // boundary and cannot be scored against.
+ if (typeof c?.start_time !== "number" || !Number.isFinite(c.start_time)) continue;
+ chapters.push({
+ start: c.start_time,
+ end:
+ typeof c.end_time === "number" && Number.isFinite(c.end_time)
+ ? c.end_time
+ : null,
+ title: typeof c.title === "string" ? c.title : "",
+ });
+ }
+ if (chapters.length === 0) return null;
+ return chapters.sort((a, b) => a.start - b.start);
+ } catch {
+ return null;
+ }
+}
diff --git a/plans/STATE.md b/plans/STATE.md
@@ -3,7 +3,93 @@
The working memory for the local-AI derived-corpus work. Rewritten at the end of every
session, before context is cleared. See [`README.md`](README.md) for the protocol.
-**Last updated:** 2026-08-09 (later) — **Phase D step 1: the digest counters collapsed into the
+**Last updated:** 2026-08-12 — **the digest bake-off can now measure QUALITY, not just
+defects; and the sweep was running 8× slower than its own projection.** Uncommitted on
+`main`. See [`bakeoff/speakers-round1-notes.md`](bakeoff/speakers-round1-notes.md) and
+[`digest-context-review.md`](digest-context-review.md).
+
+**EVERY METRIC THE BAKE-OFF HAD WAS A DEFECT COUNTER.** `zeroYieldRate`,
+`chaptersPerHour`, `rejectionRate`, `maxGapSeconds`, `genericTitleRate`,
+`duplicateTitleRate` — all say how malformed the output is, none says whether a boundary
+is in the right place. `lib/boundaryScore.ts` (new, 12 unit tests) scores generated
+chapter starts against the **uploader's own chapter marks**, an oracle the model never
+sees. Boundaries are facts, so scoring against them reproduces nothing; the uploader's
+chapter *titles* are expression and are never scored against. The t=0 mark and
+boilerplate marks are dropped, and matching is one-to-one closest-pair-first so
+over-segmentation is punished rather than rewarded.
+
+**THE ORACLE IS FAR BIGGER THAN THE SPEAKER DATA IT WAS BUILT TO JUDGE.** Measured, not
+sampled: **14,923** video dirs carry a non-empty `chapters` array; **12,139** of those
+also have a transcript and >=4 marks, across **43 channels**. Nothing in the codebase read
+`metadata.info.json.chapters` before this. Against that, only **34** videos in the whole
+archive have audio + transcript + >=4 marks — the plan's figure, confirmed exactly,
+including its channel split.
+
+**FIRST BASELINE, FOR FREE, AND IT FOUND A BUG.** `--score-existing` scores digests
+already on disk with no model calls and no writes. Over the 15 scorable ones:
+**P@30 10.8%, R@30 26.1%, R@60 35.9%, F1@30 15.2%** across 153 uploader marks; per-video
+F1 spans 0–50%, so the metric discriminates. It immediately caught
+`omnibased/3fl_uRFXxLk`: video runs 0–5119 s, uploader marked throughout, digest emitted
+chapters only between 3184–4865 s — the head two-thirds unsummarized.
+
+**READ RECALL, NOT PRECISION — and there are now three samples proving it.** A generated
+run over a fresh 12-video / 8-channel sample at the production config
+(`chunk-local`/8192/600) gives **R@30 21.5%, P@30 5.6%, F1@30 8.8%**, median offset 109 s,
+163 marks, 87 chunks. Against the free `--score-existing` run (R@30 26.1%, P@30 10.8%) and
+a short/medium subsample (R@30 12.2%, P@30 20.8%): **recall stays in a narrow band across
+three samples sharing no videos, while precision swings 5.6%→20.8% on the same engine
+purely with the duration mix.** The digest targets 1 chapter/4 min, so on short videos it
+emits fewer boundaries than the uploader (24 vs 41) and on 8-hour VODs far more (149 vs
+9). F1 inherits the swing. **Compare candidates on recall with `chaptersPerHour` beside
+it**; a variant that "improves F1" by matching the sample's density has improved nothing.
+
+Same run's guard rails: 8.1% zero-yield, 14.8% rejection, 16.5% generic titles,
+47 s/audio-hour → 43.1 projected sweep days (same order as `digest-plan`'s 53.8).
+
+**THE SWEEP WAS PAYING AN 8× CONCURRENCY TAX.** The live `digest-channel-local` job was
+running at **~200 s/chunk against `MEASURED_SECONDS_PER_CHUNK` of 24.7** — `ollama ps`
+read `85%/15% CPU/GPU`, i.e. memory pressure had pushed the model mostly onto CPU
+(prefill ~43 tok/s). With the sweep and backfill paused and memory freed, the identical
+model loads **100% GPU**. The 53.8-day projection was on the order of 400+ days at the
+contended rate. This dwarfs every prompt-level improvement currently under discussion.
+
+**TWO PRODUCTION FILES TOUCHED, BOTH ADDITIVE, BOTH BYTE-IDENTICAL WHEN UNUSED, NO
+`PROMPT_VERSION` BUMP.** `transcriptToMarkdown.prefixForCue` (the speaker label must sit
+OUTSIDE the `[HH:MM:SS]` bracket — the prompt tells the model to copy those markers
+verbatim and `HMS_PATTERN` rejects an echo) and `ChapterPromptInput.speakerRoster`
+(deliberately NOT `contextNote`, which is per-channel and hashed into `contextHash`).
+The shipped lane sets neither, so sweep output is unchanged — the same argument that let
+`timestampMode` in at version 1. Both pinned by tests; 701/701 common tests pass, `tsc`
+clean in common/editor/export.
+
+**THE SPEAKER A/B DID NOT RUN, DELIBERATELY.** Sample built (`bakeoff/speaker-sample.json`,
+12 videos / 27.3 audio-h / 178 marks) and the pipeline proven end to end on one video, but
+the plan's own precondition — confirm the box is not loaded — failed: a second sherpa
+process alongside the operator's already-running `diarize-channel` took the box to
+**262 MB free, 14.1 GB swapped, memory-pressure full-stall 2.6%**. Stopped rather than
+corrupt both the operator's job and the bake-off beside it. Exact resume command is in
+the notes file.
+
+**WHAT THE ONE DIARIZED VIDEO SHOWS, AND IT IS NOT ENCOURAGING.** `shondo/OWuZuOrbQ10`:
+8 clusters → 3 speakers, 98.8% of cues labelled, speaker changes on 16.9% of cues,
+**prompt inflation only 4.4%** (well under the ~10% budgeted, because labels are emitted
+on speaker CHANGE only). But (a) the labels are ROLES — "Host", "Caller", "Interviewer" —
+so the title-quality half of the hypothesis is untestable on this material, there is no
+name to put in a title; and (b) they **invent speaker changes**, splitting one continuous
+narration across two labels. False changes are anti-signal for boundary detection. n=1,
+but both facts were invisible before the render existed, and both say the A/B must be
+read for harm as carefully as for benefit.
+
+**ELEVEN CHANNEL CONTEXT NOTES DRAFTED, NONE PROMOTED.** `digest-context.suggested.md` for
+the 11 heaviest channels by remaining sweep work (~113k of 188,269 chunks, ~60%). A note
+only reaches the prompt as `digest-context.md`, and only a human renames it — auto-applying
+model-authored context stays forbidden. `omnibased`'s note is the strongest: 11 of 150
+sampled descriptions carry the channel's OWN segment taxonomy (`AFK:`, `Chatting:`,
+`Gaming:`, `Clicking Links:`, `Watching:`, `Debate:` …), so that note transcribes the
+channel's vocabulary rather than inferring one. Provenance and a review checklist are in
+`digest-context-review.md`.
+
+**2026-08-09 (later)** — **Phase D step 1: the digest counters collapsed into the
operation registry.** Branch `feat/digest-in-registry`, off a fast-forward merge of
`feat/diarization-oom-wall` into `main` (8 commits, no conflicts).
diff --git a/plans/attribution-pilot/speakers-smoke-quality.md b/plans/attribution-pilot/speakers-smoke-quality.md
@@ -0,0 +1,81 @@
+# Attribution pilot — speakers-smoke (diarized): what the labels actually say
+
+Read this file, not the metrics, to answer *are the labels right*. Each speaker
+is shown with its talk-time share and the verbatim transcript at its longest
+segments — the same text, in the same `[HH:MM:SS] line` form, the model saw.
+
+## shondo/OWuZuOrbQ10
+
+**[ASMR] Asking You Lots Of Questions | THE SURVEY [Keyboard Typing] [Softly Spoken]**
+
+00:42:26 · 1 chunk(s) · outcome `attributed`
+
+### Host — 67% of attributed time, confidence 0.90, first seen in chunk 1, present in 1/1 chunk(s)
+
+_00:00:41 – 00:00:49_
+
+```
+[00:00:41] You may hear your results being typed out on the computer so that they can be recorded in our online portal.
+```
+
+_00:03:47 – 00:03:56_
+
+```
+[00:03:47] It is really important to us to record your sincere answers.
+[00:03:53] Your answers are really important to us.
+```
+
+_00:21:16 – 00:21:24_
+
+```
+[00:21:16] Question twenty eight
+[00:21:19] standardized testing an accurate measure of a child's intellect
+```
+
+### Caller — 18% of attributed time, confidence 0.50, first seen in chunk 1, present in 1/1 chunk(s)
+
+_00:04:44 – 00:04:48_
+
+```
+[00:04:44] The option that you choose will be recorded in our online portal.
+```
+
+_00:17:28 – 00:17:33_
+
+```
+[00:17:32] live inside of it
+```
+
+_00:29:26 – 00:29:31_
+
+```
+[00:29:26] you may hear your chosen option being typed on the computer
+```
+
+### Interviewer — 15% of attributed time, confidence 0.80, first seen in chunk 1, present in 1/1 chunk(s)
+
+_00:36:40 – 00:36:45_
+
+```
+[00:36:40] Question three one
+[00:36:43] intelligence
+[00:36:44] or beauty
+```
+
+_00:41:07 – 00:41:12_
+
+```
+[00:41:07] there are no more questions in Tony's survey.
+```
+
+_00:41:29 – 00:41:34_
+
+```
+[00:41:29] thank you so much for volunteering for today's survey. You have answered all of the
+```
+
+- [ ] labels are right
+- [ ] a label is wrong (which: ____)
+- [ ] a name appears that the transcript never states
+- [ ] one person is split across two labels
+- [ ] two people are merged into one label
diff --git a/plans/attribution-pilot/speakers-smoke.json b/plans/attribution-pilot/speakers-smoke.json
@@ -0,0 +1,177 @@
+{
+ "version": 1,
+ "label": "speakers-smoke",
+ "method": "diarized",
+ "sample": "video:shondo/OWuZuOrbQ10",
+ "startedAt": "2026-08-12T18:18:55.516Z",
+ "finishedAt": "2026-08-12T18:19:09.922Z",
+ "engine": {
+ "appId": "ollama-direct",
+ "modelRequested": "qwen2.5:7b",
+ "numCtx": 8192,
+ "maxCues": 600,
+ "overlapCues": 40,
+ "promptVersion": 1
+ },
+ "settingsOverride": {
+ "enabled": true,
+ "appId": "ollama-direct",
+ "model": "",
+ "diarizedEnabled": true,
+ "textOnlyEnabled": false,
+ "promptVersion": 1
+ },
+ "skipped": [],
+ "box": {
+ "start": {
+ "loadavg": [
+ 13.83,
+ 16.17,
+ 15.04
+ ],
+ "freeMemGb": 2.61,
+ "totalMemGb": 15.53
+ },
+ "end": {
+ "loadavg": [
+ 12.17,
+ 15.69,
+ 14.9
+ ],
+ "freeMemGb": 2.36,
+ "totalMemGb": 15.53
+ }
+ },
+ "census": {
+ "bands": [
+ {
+ "label": "< 15 min",
+ "videos": 42454,
+ "audioSeconds": 22469452,
+ "chunks": 42454
+ },
+ {
+ "label": "15–60 min",
+ "videos": 14278,
+ "audioSeconds": 24266683,
+ "chunks": 22830
+ },
+ {
+ "label": "1–2 h",
+ "videos": 6068,
+ "audioSeconds": 31401966,
+ "chunks": 23129
+ },
+ {
+ "label": "2–4 h",
+ "videos": 5598,
+ "audioSeconds": 58120282,
+ "chunks": 33669
+ },
+ {
+ "label": "4–8 h",
+ "videos": 4622,
+ "audioSeconds": 93628258,
+ "chunks": 47375
+ },
+ {
+ "label": "> 8 h",
+ "videos": 1581,
+ "audioSeconds": 55055934,
+ "chunks": 25454
+ }
+ ],
+ "videos": 74601,
+ "audioSeconds": 284942575,
+ "chunks": 194911,
+ "estimated": 0,
+ "chunksPerAudioHour": 2.462529862376656,
+ "statsSchemaStale": false
+ },
+ "cost": {
+ "videosMeasured": 1,
+ "chunks": 1,
+ "chunksOk": 1,
+ "calls": 1,
+ "callsWithEngineTiming": 1,
+ "audioSeconds": 2546,
+ "chunksPerAudioHour": 1.4139827179890023,
+ "secondsPerChunkWall": 14.383,
+ "secondsPerChunkCallWall": 14.4,
+ "secondsPerChunkEngine": 14.3,
+ "secondsPerChunkEngineExLoad": 4.7,
+ "meanLoadSeconds": 0.4,
+ "decodeTokensPerSecond": 42.72727272727273,
+ "unparsedLogLines": 0
+ },
+ "baseline": {
+ "digestSecondsPerChunkEngine": 11.2
+ },
+ "videos": [
+ {
+ "slug": "shondo/OWuZuOrbQ10",
+ "title": "[ASMR] Asking You Lots Of Questions | THE SURVEY [Keyboard Typing] [Softly Spoken]",
+ "durationSeconds": 2546,
+ "expectedChunks": 1,
+ "outcome": "attributed",
+ "videoWallSeconds": 14.383,
+ "calls": [
+ {
+ "wallSeconds": 14.4,
+ "inputTokens": 777,
+ "outputTokens": 94,
+ "loadSeconds": 0.4,
+ "prefillSeconds": 2.5,
+ "decodeSeconds": 2.2,
+ "engineSeconds": 14.3
+ }
+ ],
+ "unparsedLogLines": 0,
+ "warningsByCode": {},
+ "metrics": {
+ "labels": [
+ "Host",
+ "Caller",
+ "Interviewer"
+ ],
+ "labelsPerVideo": 3,
+ "newLabelsPerChunk": null,
+ "singletonLabelRate": null,
+ "dominantSpeakerShare": 0.6657032755298651,
+ "talkShares": [
+ 0.6657032755298651,
+ 0.18304431599229287,
+ 0.151252408477842
+ ],
+ "nearDuplicateLabelRate": 0,
+ "nearDuplicatePairs": [],
+ "genericLabelRate": 1,
+ "genericSecondsShare": 1,
+ "nonAsciiLabelRate": 0,
+ "coverageRate": 0.4075439905734475,
+ "maxGapSeconds": 52.903999999999996,
+ "medianGapSeconds": 1.3329999999998563,
+ "markFreeChunks": 0,
+ "segments": 487,
+ "chunkReconstructionOk": true,
+ "labelFirstChunk": [
+ 0,
+ 0,
+ 0
+ ],
+ "labelChunkCount": [
+ 1,
+ 1,
+ 1
+ ]
+ },
+ "clustersOffered": 3,
+ "clustersNamed": 3,
+ "confidences": [
+ 0.9,
+ 0.5,
+ 0.8
+ ]
+ }
+ ]
+}
diff --git a/plans/attribution-pilot/speakers-smoke.md b/plans/attribution-pilot/speakers-smoke.md
@@ -0,0 +1,48 @@
+# Attribution pilot — speakers-smoke
+
+Lane **diarized** · sample `video:shondo/OWuZuOrbQ10` · 1 video(s) · model `qwen2.5:7b` @ numCtx 8192 (maxCues 600) · 2026-08-12T18:18:55.516Z
+
+Outcomes: attributed 1
+
+## Diarized lane — the claim under test is ONE call per video
+
+| video | calls | clusters offered | clusters named | confidences |
+| --- | ---: | ---: | ---: | --- |
+| shondo/OWuZuOrbQ10 | 1 | 3 | 3 | 0.90, 0.50, 0.80 |
+
+**Calls per video: 1.00** (1 call(s) over 1 video(s)). The claim holds iff this is 1.00.
+
+Warnings by code: none
+
+## Cost
+
+Measured over the 1 video(s) this run actually attributed (1/1 chunk(s) ok, 1 call(s), 1 with engine timing). Sample density **1.41 chunks/audio-hour** — the corpus census reads 2.46.
+
+| quantity | value |
+| --- | ---: |
+| s/chunk, wall (video wall ÷ chunks) | 14.4 |
+| s/chunk, per-call wall | 14.4 |
+| s/chunk, engine | 14.3 |
+| s/chunk, engine excl. model load | 4.7 |
+| mean model-load s/call | 0.40 |
+| decode tokens/s | 42.7 |
+| digest chapters baseline, s/chunk engine | 11.2 |
+
+## Box
+
+Start: `{"loadavg":[13.83,16.17,15.04],"freeMemGb":2.61,"totalMemGb":15.53}`
+
+End: `{"loadavg":[12.17,15.69,14.9],"freeMemGb":2.36,"totalMemGb":15.53}`
+
+## Units
+
+- The unit of text-only work is the **chunk** (one model call). Seconds-per-
+ audio-hour is not a unit here: chunk density varies 4× across this corpus, and
+ pricing in it is the error behind a retracted throughput headline. This report
+ does not print one.
+- The unit of diarized work is **calls per video**.
+- Wall and engine are different quantities and are never substituted for each
+ other. Wall is the realistic ceiling on this shared box; engine excl. load is
+ the floor.
+- Every rate carries its denominator. At n ≈ 8 videos a bare percentage would be
+ a lie of precision.
diff --git a/plans/bakeoff/chapter-sample.json b/plans/bakeoff/chapter-sample.json
@@ -0,0 +1,145 @@
+{
+ "version": 1,
+ "pickedAt": "2026-08-12T18:05:55.714Z",
+ "corpus": {
+ "videosScanned": 78525,
+ "videosWithTranscript": 74601,
+ "audioHours": 79151,
+ "longTailVideos": 6200,
+ "longTailAudioHours": 41289
+ },
+ "videos": [
+ {
+ "slug": "nuxanor/9bgVoJY0jEg",
+ "channelSlug": "nuxanor",
+ "videoId": "9bgVoJY0jEg",
+ "videoDir": "9bgVoJY0jEg",
+ "title": "New Vivziepop allegations are unsettling...",
+ "bucket": "short",
+ "durationSeconds": 817,
+ "cueCount": 407,
+ "uploaderChapters": 9
+ },
+ {
+ "slug": "chibi-reviews/XYyejb64QpU",
+ "channelSlug": "chibi-reviews",
+ "videoId": "XYyejb64QpU",
+ "videoDir": "XYyejb64QpU",
+ "title": "Boku no Hero Academia Chapter 57 Manga Review - Wonderful New Designs 僕のヒーローアカデミア",
+ "bucket": "short",
+ "durationSeconds": 619,
+ "cueCount": 275,
+ "uploaderChapters": 8
+ },
+ {
+ "slug": "the-quartering/br9koSf6D0Q",
+ "channelSlug": "the-quartering",
+ "videoId": "br9koSf6D0Q",
+ "videoDir": "br9koSf6D0Q",
+ "title": "Father Of 4 Doxxed By Journos For Trolling Joe Biden & Twitter Trended His Information!",
+ "bucket": "short",
+ "durationSeconds": 606,
+ "cueCount": 244,
+ "uploaderChapters": 13
+ },
+ {
+ "slug": "destiny/5nmDzKB23OU",
+ "channelSlug": "destiny",
+ "videoId": "5nmDzKB23OU",
+ "videoDir": "5nmDzKB23OU",
+ "title": "Can Science Answer All Questions?",
+ "bucket": "medium",
+ "durationSeconds": 4419,
+ "cueCount": 2072,
+ "uploaderChapters": 7
+ },
+ {
+ "slug": "angryjoeshow/8AJsKyh0x7w",
+ "channelSlug": "angryjoeshow",
+ "videoId": "8AJsKyh0x7w",
+ "videoDir": "8AJsKyh0x7w",
+ "title": "Anthem Angry Review",
+ "bucket": "medium",
+ "durationSeconds": 3041,
+ "cueCount": 867,
+ "uploaderChapters": 6
+ },
+ {
+ "slug": "kirsche/GchuqKQsnHk",
+ "channelSlug": "kirsche",
+ "videoId": "GchuqKQsnHk",
+ "videoDir": "GchuqKQsnHk",
+ "title": "Journalist Targeted By Karmelo Anthony Supporters?",
+ "bucket": "medium",
+ "durationSeconds": 4155,
+ "cueCount": 1821,
+ "uploaderChapters": 17
+ },
+ {
+ "slug": "HasanAbiVODs/p7-GXvM3eXg",
+ "channelSlug": "HasanAbiVODs",
+ "videoId": "p7-GXvM3eXg",
+ "videoDir": "p7-GXvM3eXg",
+ "title": "2/2 HasanAbi June 18, 2021 – Criminal Psychology, LSF, Money Laundering REACT, 🎮Ratchet & Clank🎮",
+ "bucket": "long",
+ "durationSeconds": 15523,
+ "cueCount": 3951,
+ "uploaderChapters": 5
+ },
+ {
+ "slug": "omnibased/S-9F_Jr3jmA",
+ "channelSlug": "omnibased",
+ "videoId": "S-9F_Jr3jmA",
+ "videoDir": "S-9F_Jr3jmA",
+ "title": "2026-01-24 Another ICE Shooting | Tectone's Lost Court Case?? | Chillin' with mooty 7BTSlPSUwWE",
+ "bucket": "long",
+ "durationSeconds": 12705,
+ "cueCount": 3746,
+ "uploaderChapters": 36
+ },
+ {
+ "slug": "destiny/WL9htojKbGY",
+ "channelSlug": "destiny",
+ "videoId": "WL9htojKbGY",
+ "videoDir": "WL9htojKbGY",
+ "title": "Destiny Debates 7 Pro-MAGA Asmongold Fans On Venezuela",
+ "bucket": "long",
+ "durationSeconds": 14259,
+ "cueCount": 7329,
+ "uploaderChapters": 9
+ },
+ {
+ "slug": "HasanAbiVODs/GjPX_ueTdfc",
+ "channelSlug": "HasanAbiVODs",
+ "videoId": "GjPX_ueTdfc",
+ "videoDir": "GjPX_ueTdfc",
+ "title": "2/2 HasanAbi April 7, 2021 - Sunburn OMEGALUL, Basketball Memes, OKBUDDY, 🎮GTA NoPixel🎮 FULL VOD",
+ "bucket": "verylong",
+ "durationSeconds": 28965,
+ "cueCount": 7223,
+ "uploaderChapters": 10
+ },
+ {
+ "slug": "HasanAbiVODs3/bW2lKpHlmMQ",
+ "channelSlug": "HasanAbiVODs3",
+ "videoId": "bW2lKpHlmMQ",
+ "videoDir": "bW2lKpHlmMQ",
+ "title": "HasanAbi November 29, 2024 – Cenk Drama, Syria Opposition Offensive, End of England, Ordinary Gamers",
+ "bucket": "verylong",
+ "durationSeconds": 28740,
+ "cueCount": 8764,
+ "uploaderChapters": 16
+ },
+ {
+ "slug": "HasanAbiVODsbackup/5IufSgJSI0E",
+ "channelSlug": "HasanAbiVODsbackup",
+ "videoId": "5IufSgJSI0E",
+ "videoDir": "5IufSgJSI0E",
+ "title": "HasanAbi February 25, 2022 – RUSSIA INVADES UKRAINE Day",
+ "bucket": "verylong",
+ "durationSeconds": 28747,
+ "cueCount": 0,
+ "uploaderChapters": 34
+ }
+ ]
+}
diff --git a/plans/bakeoff/chapters-baseline.json b/plans/bakeoff/chapters-baseline.json
@@ -0,0 +1,1241 @@
+{
+ "label": "chapters-baseline",
+ "sample": "../plans/bakeoff/chapter-sample.json",
+ "corpus": {
+ "videosScanned": 78525,
+ "videosWithTranscript": 74601,
+ "audioHours": 79151,
+ "longTailVideos": 6200,
+ "longTailAudioHours": 41289
+ },
+ "buckets": [
+ "short",
+ "medium",
+ "long",
+ "verylong"
+ ],
+ "scores": [
+ {
+ "candidate": {
+ "key": "qwen2.5:7b@8192/chunk-local",
+ "model": "qwen2.5:7b",
+ "numCtx": 8192,
+ "maxCues": 600,
+ "timestampMode": "chunk-local",
+ "speakers": false
+ },
+ "videos": [
+ {
+ "slug": "nuxanor/9bgVoJY0jEg",
+ "bucket": "short",
+ "durationSeconds": 817,
+ "chunks": 1,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 4,
+ "maxGapSeconds": 339,
+ "engineSeconds": 16.49,
+ "inputTokens": 8126,
+ "outputTokens": 133,
+ "warningsByCode": {},
+ "genericTitles": 2,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:00:04 Introduction to the controversy",
+ "00:05:02 Discussion on Alistair's design and Voodoo symbols",
+ "00:10:40 Harassment and death threats",
+ "00:12:51 Personal reflections on the issue"
+ ],
+ "boundary": {
+ "referenceCount": 9,
+ "generatedCount": 3,
+ "medianOffsetSeconds": 109,
+ "meanOffsetSeconds": 103.77777777777777,
+ "withinThirtySeconds": 2,
+ "scores": [
+ {
+ "toleranceSeconds": 30,
+ "referenceCount": 9,
+ "generatedCount": 3,
+ "matched": 1,
+ "precision": 0.3333333333333333,
+ "recall": 0.1111111111111111,
+ "f1": 0.16666666666666666,
+ "matches": [
+ {
+ "reference": 649,
+ "generated": 640,
+ "distance": 9
+ }
+ ]
+ },
+ {
+ "toleranceSeconds": 60,
+ "referenceCount": 9,
+ "generatedCount": 3,
+ "matched": 2,
+ "precision": 0.6666666666666666,
+ "recall": 0.2222222222222222,
+ "f1": 0.3333333333333333,
+ "matches": [
+ {
+ "reference": 355,
+ "generated": 301,
+ "distance": 54
+ },
+ {
+ "reference": 649,
+ "generated": 640,
+ "distance": 9
+ }
+ ]
+ }
+ ]
+ },
+ "speakersRendered": 0,
+ "titlesWithSpeakerName": 0
+ },
+ {
+ "slug": "chibi-reviews/XYyejb64QpU",
+ "bucket": "short",
+ "durationSeconds": 619,
+ "chunks": 1,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 2,
+ "maxGapSeconds": 448,
+ "engineSeconds": 2.924,
+ "inputTokens": 5578,
+ "outputTokens": 71,
+ "warningsByCode": {},
+ "genericTitles": 0,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:00:07 Villain Character Designs",
+ "00:02:52 Media and News Impact on Villains"
+ ],
+ "boundary": {
+ "referenceCount": 8,
+ "generatedCount": 2,
+ "medianOffsetSeconds": 153,
+ "meanOffsetSeconds": 177.625,
+ "withinThirtySeconds": 1,
+ "scores": [
+ {
+ "toleranceSeconds": 30,
+ "referenceCount": 8,
+ "generatedCount": 2,
+ "matched": 1,
+ "precision": 0.5,
+ "recall": 0.125,
+ "f1": 0.2,
+ "matches": [
+ {
+ "reference": 163,
+ "generated": 171,
+ "distance": 8
+ }
+ ]
+ },
+ {
+ "toleranceSeconds": 60,
+ "referenceCount": 8,
+ "generatedCount": 2,
+ "matched": 2,
+ "precision": 1,
+ "recall": 0.25,
+ "f1": 0.4,
+ "matches": [
+ {
+ "reference": 62,
+ "generated": 7,
+ "distance": 55
+ },
+ {
+ "reference": 163,
+ "generated": 171,
+ "distance": 8
+ }
+ ]
+ }
+ ]
+ },
+ "speakersRendered": 0,
+ "titlesWithSpeakerName": 0
+ },
+ {
+ "slug": "the-quartering/br9koSf6D0Q",
+ "bucket": "short",
+ "durationSeconds": 606,
+ "chunks": 1,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 2,
+ "maxGapSeconds": 337,
+ "engineSeconds": 3.453,
+ "inputTokens": 4879,
+ "outputTokens": 71,
+ "warningsByCode": {},
+ "genericTitles": 1,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:00:01 Introduction and Media Criticism",
+ "00:04:29 Prank Call Details and Backstory"
+ ],
+ "boundary": {
+ "referenceCount": 13,
+ "generatedCount": 1,
+ "medianOffsetSeconds": 126,
+ "meanOffsetSeconds": 133.84615384615384,
+ "withinThirtySeconds": 1,
+ "scores": [
+ {
+ "toleranceSeconds": 30,
+ "referenceCount": 13,
+ "generatedCount": 1,
+ "matched": 1,
+ "precision": 1,
+ "recall": 0.07692307692307693,
+ "f1": 0.14285714285714288,
+ "matches": [
+ {
+ "reference": 270,
+ "generated": 269,
+ "distance": 1
+ }
+ ]
+ },
+ {
+ "toleranceSeconds": 60,
+ "referenceCount": 13,
+ "generatedCount": 1,
+ "matched": 1,
+ "precision": 1,
+ "recall": 0.07692307692307693,
+ "f1": 0.14285714285714288,
+ "matches": [
+ {
+ "reference": 270,
+ "generated": 269,
+ "distance": 1
+ }
+ ]
+ }
+ ]
+ },
+ "speakersRendered": 0,
+ "titlesWithSpeakerName": 0
+ },
+ {
+ "slug": "destiny/5nmDzKB23OU",
+ "bucket": "medium",
+ "durationSeconds": 4419,
+ "chunks": 4,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 1,
+ "kept": 15,
+ "maxGapSeconds": 1201,
+ "engineSeconds": 81.166,
+ "inputTokens": 19982,
+ "outputTokens": 731,
+ "warningsByCode": {
+ "out-of-range": 10
+ },
+ "genericTitles": 3,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:00:06 Introduction and Initial Agreement",
+ "00:01:42 Defining Religion and Atheism",
+ "00:04:31 The Role of Reason in Belief",
+ "00:07:44 Atheist vs. Agnostic Perspective",
+ "00:10:35 Reasoning Processes and Their Differences",
+ "00:13:51 Discussion on God of the Gaps Argument",
+ "00:16:14 Impact of Organized Religion on Society",
+ "00:19:29 Introduction and Skepticism"
+ ],
+ "boundary": {
+ "referenceCount": 6,
+ "generatedCount": 15,
+ "medianOffsetSeconds": 111.5,
+ "meanOffsetSeconds": 246.66666666666666,
+ "withinThirtySeconds": 0,
+ "scores": [
+ {
+ "toleranceSeconds": 30,
+ "referenceCount": 6,
+ "generatedCount": 15,
+ "matched": 0,
+ "precision": 0,
+ "recall": 0,
+ "f1": 0,
+ "matches": []
+ },
+ {
+ "toleranceSeconds": 60,
+ "referenceCount": 6,
+ "generatedCount": 15,
+ "matched": 1,
+ "precision": 0.06666666666666667,
+ "recall": 0.16666666666666666,
+ "f1": 0.09523809523809522,
+ "matches": [
+ {
+ "reference": 3607,
+ "generated": 3566,
+ "distance": 41
+ }
+ ]
+ }
+ ]
+ },
+ "speakersRendered": 0,
+ "titlesWithSpeakerName": 0
+ },
+ {
+ "slug": "angryjoeshow/8AJsKyh0x7w",
+ "bucket": "medium",
+ "durationSeconds": 3041,
+ "chunks": 2,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 1,
+ "kept": 3,
+ "maxGapSeconds": 2074,
+ "engineSeconds": 48.357,
+ "inputTokens": 10332,
+ "outputTokens": 350,
+ "warningsByCode": {
+ "language-drift": 9
+ },
+ "genericTitles": 0,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:34:34 Gameplay Issues",
+ "00:43:21 Endgame and Crafting Issues",
+ "00:47:11 Final Verdict and Future Prospects"
+ ],
+ "boundary": {
+ "referenceCount": 5,
+ "generatedCount": 3,
+ "medianOffsetSeconds": 57,
+ "meanOffsetSeconds": 752.2,
+ "withinThirtySeconds": 2,
+ "scores": [
+ {
+ "toleranceSeconds": 30,
+ "referenceCount": 5,
+ "generatedCount": 3,
+ "matched": 2,
+ "precision": 0.6666666666666666,
+ "recall": 0.4,
+ "f1": 0.5,
+ "matches": [
+ {
+ "reference": 2057,
+ "generated": 2074,
+ "distance": 17
+ },
+ {
+ "reference": 2831,
+ "generated": 2830,
+ "distance": 1
+ }
+ ]
+ },
+ {
+ "toleranceSeconds": 60,
+ "referenceCount": 5,
+ "generatedCount": 3,
+ "matched": 2,
+ "precision": 0.6666666666666666,
+ "recall": 0.4,
+ "f1": 0.5,
+ "matches": [
+ {
+ "reference": 2057,
+ "generated": 2074,
+ "distance": 17
+ },
+ {
+ "reference": 2831,
+ "generated": 2830,
+ "distance": 1
+ }
+ ]
+ }
+ ]
+ },
+ "speakersRendered": 0,
+ "titlesWithSpeakerName": 0
+ },
+ {
+ "slug": "kirsche/GchuqKQsnHk",
+ "bucket": "medium",
+ "durationSeconds": 4155,
+ "chunks": 4,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 28,
+ "maxGapSeconds": 925,
+ "engineSeconds": 78.359,
+ "inputTokens": 15303,
+ "outputTokens": 785,
+ "warningsByCode": {},
+ "genericTitles": 4,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:15:16 Shooting at Wazada High School Graduation Ceremony",
+ "00:15:35 University of Minnesota Police Response",
+ "00:16:30 Surveillance Video and Suspect Description",
+ "00:17:46 Suspect's Arrest and Charges",
+ "00:19:46 University of Minnesota Decision to Stop Hosting Graduations",
+ "00:21:06 Impact on Local School Districts",
+ "00:21:52 Anoka-Henipin School District's Response",
+ "00:37:17 Introduction of the incident"
+ ],
+ "boundary": {
+ "referenceCount": 17,
+ "generatedCount": 28,
+ "medianOffsetSeconds": 178,
+ "meanOffsetSeconds": 275.05882352941177,
+ "withinThirtySeconds": 4,
+ "scores": [
+ {
+ "toleranceSeconds": 30,
+ "referenceCount": 17,
+ "generatedCount": 28,
+ "matched": 4,
+ "precision": 0.14285714285714285,
+ "recall": 0.23529411764705882,
+ "f1": 0.17777777777777778,
+ "matches": [
+ {
+ "reference": 2249,
+ "generated": 2249,
+ "distance": 0
+ },
+ {
+ "reference": 2615,
+ "generated": 2617,
+ "distance": 2
+ },
+ {
+ "reference": 2661,
+ "generated": 2661,
+ "distance": 0
+ },
+ {
+ "reference": 2699,
+ "generated": 2709,
+ "distance": 10
+ }
+ ]
+ },
+ {
+ "toleranceSeconds": 60,
+ "referenceCount": 17,
+ "generatedCount": 28,
+ "matched": 6,
+ "precision": 0.21428571428571427,
+ "recall": 0.35294117647058826,
+ "f1": 0.26666666666666666,
+ "matches": [
+ {
+ "reference": 1366,
+ "generated": 1312,
+ "distance": 54
+ },
+ {
+ "reference": 2249,
+ "generated": 2249,
+ "distance": 0
+ },
+ {
+ "reference": 2615,
+ "generated": 2617,
+ "distance": 2
+ },
+ {
+ "reference": 2661,
+ "generated": 2661,
+ "distance": 0
+ },
+ {
+ "reference": 2699,
+ "generated": 2709,
+ "distance": 10
+ },
+ {
+ "reference": 3787,
+ "generated": 3831,
+ "distance": 44
+ }
+ ]
+ }
+ ]
+ },
+ "speakersRendered": 0,
+ "titlesWithSpeakerName": 0
+ },
+ {
+ "slug": "HasanAbiVODs/p7-GXvM3eXg",
+ "bucket": "long",
+ "durationSeconds": 15523,
+ "chunks": 7,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 51,
+ "maxGapSeconds": 2366,
+ "engineSeconds": 155.98499999999999,
+ "inputTokens": 28686,
+ "outputTokens": 1732,
+ "warningsByCode": {
+ "out-of-range": 4,
+ "non-monotonic": 13
+ },
+ "genericTitles": 10,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:23:08 Initial frustration and request to stop",
+ "00:23:11 Discussion of RP frogs",
+ "00:23:14 Request for an ad break",
+ "00:23:18 Realistic portrayal of money laundering in movies",
+ "00:24:01 Money laundering basics",
+ "00:25:13 Realistic portrayal of money laundering in 'Gotcha'",
+ "00:27:02 Drug money laundering through gems and emeralds",
+ "00:28:02 Realistic portrayal of drug dealer's cash stashes"
+ ],
+ "boundary": {
+ "referenceCount": 4,
+ "generatedCount": 51,
+ "medianOffsetSeconds": 49.5,
+ "meanOffsetSeconds": 59,
+ "withinThirtySeconds": 2,
+ "scores": [
+ {
+ "toleranceSeconds": 30,
+ "referenceCount": 4,
+ "generatedCount": 51,
+ "matched": 2,
+ "precision": 0.0392156862745098,
+ "recall": 0.5,
+ "f1": 0.07272727272727272,
+ "matches": [
+ {
+ "reference": 1397,
+ "generated": 1398,
+ "distance": 1
+ },
+ {
+ "reference": 9709,
+ "generated": 9721,
+ "distance": 12
+ }
+ ]
+ },
+ {
+ "toleranceSeconds": 60,
+ "referenceCount": 4,
+ "generatedCount": 51,
+ "matched": 2,
+ "precision": 0.0392156862745098,
+ "recall": 0.5,
+ "f1": 0.07272727272727272,
+ "matches": [
+ {
+ "reference": 1397,
+ "generated": 1398,
+ "distance": 1
+ },
+ {
+ "reference": 9709,
+ "generated": 9721,
+ "distance": 12
+ }
+ ]
+ }
+ ]
+ },
+ "speakersRendered": 0,
+ "titlesWithSpeakerName": 0
+ },
+ {
+ "slug": "omnibased/S-9F_Jr3jmA",
+ "bucket": "long",
+ "durationSeconds": 12705,
+ "chunks": 7,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 66,
+ "maxGapSeconds": 1351,
+ "engineSeconds": 169.464,
+ "inputTokens": 32383,
+ "outputTokens": 1951,
+ "warningsByCode": {
+ "out-of-range": 4
+ },
+ "genericTitles": 23,
+ "duplicateTitles": 15,
+ "sampleTitles": [
+ "00:00:48 Introduction and Opening Remarks",
+ "00:02:23 Discussion on the Situation in Minneapolis",
+ "00:05:16 Response to Questions about ICE Operations",
+ "00:08:17 Explanation of the Shooting Incident",
+ "00:14:44 Discussion on the Storm and Emergency Preparedness",
+ "00:19:33 Response to Questions about the Shooting Incident",
+ "00:25:45 Conclusion and Closing Remarks",
+ "00:29:02 Introduction and Setup"
+ ],
+ "boundary": {
+ "referenceCount": 35,
+ "generatedCount": 66,
+ "medianOffsetSeconds": 73,
+ "meanOffsetSeconds": 101.37142857142857,
+ "withinThirtySeconds": 10,
+ "scores": [
+ {
+ "toleranceSeconds": 30,
+ "referenceCount": 35,
+ "generatedCount": 66,
+ "matched": 9,
+ "precision": 0.13636363636363635,
+ "recall": 0.2571428571428571,
+ "f1": 0.17821782178217818,
+ "matches": [
+ {
+ "reference": 1154,
+ "generated": 1172,
+ "distance": 18
+ },
+ {
+ "reference": 6672,
+ "generated": 6673,
+ "distance": 1
+ },
+ {
+ "reference": 7001,
+ "generated": 6982,
+ "distance": 19
+ },
+ {
+ "reference": 7220,
+ "generated": 7214,
+ "distance": 6
+ },
+ {
+ "reference": 7474,
+ "generated": 7471,
+ "distance": 3
+ },
+ {
+ "reference": 9944,
+ "generated": 9943,
+ "distance": 1
+ },
+ {
+ "reference": 10151,
+ "generated": 10176,
+ "distance": 25
+ },
+ {
+ "reference": 11659,
+ "generated": 11662,
+ "distance": 3
+ },
+ {
+ "reference": 12411,
+ "generated": 12427,
+ "distance": 16
+ }
+ ]
+ },
+ {
+ "toleranceSeconds": 60,
+ "referenceCount": 35,
+ "generatedCount": 66,
+ "matched": 14,
+ "precision": 0.21212121212121213,
+ "recall": 0.4,
+ "f1": 0.2772277227722772,
+ "matches": [
+ {
+ "reference": 1154,
+ "generated": 1172,
+ "distance": 18
+ },
+ {
+ "reference": 4867,
+ "generated": 4923,
+ "distance": 56
+ },
+ {
+ "reference": 5087,
+ "generated": 5042,
+ "distance": 45
+ },
+ {
+ "reference": 5198,
+ "generated": 5149,
+ "distance": 49
+ },
+ {
+ "reference": 6672,
+ "generated": 6673,
+ "distance": 1
+ },
+ {
+ "reference": 7001,
+ "generated": 6982,
+ "distance": 19
+ },
+ {
+ "reference": 7220,
+ "generated": 7214,
+ "distance": 6
+ },
+ {
+ "reference": 7474,
+ "generated": 7471,
+ "distance": 3
+ },
+ {
+ "reference": 9413,
+ "generated": 9446,
+ "distance": 33
+ },
+ {
+ "reference": 9944,
+ "generated": 9943,
+ "distance": 1
+ },
+ {
+ "reference": 10014,
+ "generated": 10065,
+ "distance": 51
+ },
+ {
+ "reference": 10151,
+ "generated": 10176,
+ "distance": 25
+ },
+ {
+ "reference": 11659,
+ "generated": 11662,
+ "distance": 3
+ },
+ {
+ "reference": 12411,
+ "generated": 12427,
+ "distance": 16
+ }
+ ]
+ }
+ ]
+ },
+ "speakersRendered": 0,
+ "titlesWithSpeakerName": 0
+ },
+ {
+ "slug": "destiny/WL9htojKbGY",
+ "bucket": "long",
+ "durationSeconds": 14259,
+ "chunks": 14,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 1,
+ "kept": 67,
+ "maxGapSeconds": 1168,
+ "engineSeconds": 262.87899999999996,
+ "inputTokens": 54563,
+ "outputTokens": 2127,
+ "warningsByCode": {
+ "language-drift": 6,
+ "out-of-range": 1
+ },
+ "genericTitles": 19,
+ "duplicateTitles": 6,
+ "sampleTitles": [
+ "00:19:29 Introduction and Context",
+ "00:22:03 Discussion on Asmin Gold's Views",
+ "00:24:46 Criticism of the Video Content",
+ "00:27:58 Investigations and Prosecutions under Biden Administration",
+ "00:32:52 Fraud Allegations in Daycares and Charities",
+ "00:36:14 Investigation Outcome Uncertainty",
+ "00:37:21 Introduction and Context",
+ "00:37:58 Entertainment vs. Journalism Debate"
+ ],
+ "boundary": {
+ "referenceCount": 9,
+ "generatedCount": 67,
+ "medianOffsetSeconds": 90,
+ "meanOffsetSeconds": 221.11111111111111,
+ "withinThirtySeconds": 1,
+ "scores": [
+ {
+ "toleranceSeconds": 30,
+ "referenceCount": 9,
+ "generatedCount": 67,
+ "matched": 1,
+ "precision": 0.014925373134328358,
+ "recall": 0.1111111111111111,
+ "f1": 0.026315789473684213,
+ "matches": [
+ {
+ "reference": 1331,
+ "generated": 1325,
+ "distance": 6
+ }
+ ]
+ },
+ {
+ "toleranceSeconds": 60,
+ "referenceCount": 9,
+ "generatedCount": 67,
+ "matched": 3,
+ "precision": 0.04477611940298507,
+ "recall": 0.3333333333333333,
+ "f1": 0.07894736842105263,
+ "matches": [
+ {
+ "reference": 1331,
+ "generated": 1325,
+ "distance": 6
+ },
+ {
+ "reference": 3205,
+ "generated": 3145,
+ "distance": 60
+ },
+ {
+ "reference": 4034,
+ "generated": 4077,
+ "distance": 43
+ }
+ ]
+ }
+ ]
+ },
+ "speakersRendered": 0,
+ "titlesWithSpeakerName": 0
+ },
+ {
+ "slug": "HasanAbiVODs/GjPX_ueTdfc",
+ "bucket": "verylong",
+ "durationSeconds": 28965,
+ "chunks": 13,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 149,
+ "maxGapSeconds": 2123,
+ "engineSeconds": 304.272,
+ "inputTokens": 53274,
+ "outputTokens": 4054,
+ "warningsByCode": {
+ "non-monotonic": 7
+ },
+ "genericTitles": 14,
+ "duplicateTitles": 1,
+ "sampleTitles": [
+ "00:35:24 Stream Snipers and Stream Saviors",
+ "00:36:13 Conservative Floridian Mom Lobster",
+ "00:37:50 Bondo Simulator 2021",
+ "00:39:05 Boomer of the Year",
+ "00:41:15 Compliments and Reactions",
+ "00:45:28 Chat Union Demands",
+ "00:49:12 Food Cravings",
+ "00:53:25 Sweat Dripping Down"
+ ],
+ "boundary": {
+ "referenceCount": 9,
+ "generatedCount": 149,
+ "medianOffsetSeconds": 263,
+ "meanOffsetSeconds": 453.1111111111111,
+ "withinThirtySeconds": 1,
+ "scores": [
+ {
+ "toleranceSeconds": 30,
+ "referenceCount": 9,
+ "generatedCount": 149,
+ "matched": 1,
+ "precision": 0.006711409395973154,
+ "recall": 0.1111111111111111,
+ "f1": 0.012658227848101266,
+ "matches": [
+ {
+ "reference": 3334,
+ "generated": 3328,
+ "distance": 6
+ }
+ ]
+ },
+ {
+ "toleranceSeconds": 60,
+ "referenceCount": 9,
+ "generatedCount": 149,
+ "matched": 1,
+ "precision": 0.006711409395973154,
+ "recall": 0.1111111111111111,
+ "f1": 0.012658227848101266,
+ "matches": [
+ {
+ "reference": 3334,
+ "generated": 3328,
+ "distance": 6
+ }
+ ]
+ }
+ ]
+ },
+ "speakersRendered": 0,
+ "titlesWithSpeakerName": 0
+ },
+ {
+ "slug": "HasanAbiVODs3/bW2lKpHlmMQ",
+ "bucket": "verylong",
+ "durationSeconds": 28740,
+ "chunks": 16,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 1,
+ "kept": 122,
+ "maxGapSeconds": 3112,
+ "engineSeconds": 351.075,
+ "inputTokens": 68571,
+ "outputTokens": 3755,
+ "warningsByCode": {
+ "out-of-range": 14
+ },
+ "genericTitles": 17,
+ "duplicateTitles": 3,
+ "sampleTitles": [
+ "00:07:55 Introduction and Podcast Update",
+ "00:09:20 Kyrie Irving's Performance in Game 7",
+ "00:12:31 LeBron James' Jab Step and Its Impact on the Game",
+ "00:16:24 Discussion on Kyrie Irving's Shot Selection",
+ "00:23:22 Podcast Episode Rankings and Pod Save America",
+ "00:31:41 Support for Raes Texas",
+ "00:35:56 TikTok Video with Twin Peaks Music",
+ "00:39:44 Thanksgiving and Business Continuation"
+ ],
+ "boundary": {
+ "referenceCount": 15,
+ "generatedCount": 122,
+ "medianOffsetSeconds": 60,
+ "meanOffsetSeconds": 191.86666666666667,
+ "withinThirtySeconds": 3,
+ "scores": [
+ {
+ "toleranceSeconds": 30,
+ "referenceCount": 15,
+ "generatedCount": 122,
+ "matched": 3,
+ "precision": 0.02459016393442623,
+ "recall": 0.2,
+ "f1": 0.043795620437956206,
+ "matches": [
+ {
+ "reference": 5587,
+ "generated": 5579,
+ "distance": 8
+ },
+ {
+ "reference": 6000,
+ "generated": 6001,
+ "distance": 1
+ },
+ {
+ "reference": 17388,
+ "generated": 17382,
+ "distance": 6
+ }
+ ]
+ },
+ {
+ "toleranceSeconds": 60,
+ "referenceCount": 15,
+ "generatedCount": 122,
+ "matched": 8,
+ "precision": 0.06557377049180328,
+ "recall": 0.5333333333333333,
+ "f1": 0.11678832116788322,
+ "matches": [
+ {
+ "reference": 605,
+ "generated": 561,
+ "distance": 44
+ },
+ {
+ "reference": 2199,
+ "generated": 2156,
+ "distance": 43
+ },
+ {
+ "reference": 5587,
+ "generated": 5579,
+ "distance": 8
+ },
+ {
+ "reference": 6000,
+ "generated": 6001,
+ "distance": 1
+ },
+ {
+ "reference": 6953,
+ "generated": 7013,
+ "distance": 60
+ },
+ {
+ "reference": 10782,
+ "generated": 10824,
+ "distance": 42
+ },
+ {
+ "reference": 17388,
+ "generated": 17382,
+ "distance": 6
+ },
+ {
+ "reference": 28439,
+ "generated": 28382,
+ "distance": 57
+ }
+ ]
+ }
+ ]
+ },
+ "speakersRendered": 0,
+ "titlesWithSpeakerName": 0
+ },
+ {
+ "slug": "HasanAbiVODsbackup/5IufSgJSI0E",
+ "bucket": "verylong",
+ "durationSeconds": 28747,
+ "chunks": 17,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 3,
+ "kept": 123,
+ "maxGapSeconds": 2974,
+ "engineSeconds": 368.595,
+ "inputTokens": 68039,
+ "outputTokens": 4508,
+ "warningsByCode": {
+ "non-monotonic": 3,
+ "language-drift": 7,
+ "out-of-range": 32,
+ "seam-duplicate": 1
+ },
+ "genericTitles": 11,
+ "duplicateTitles": 3,
+ "sampleTitles": [
+ "00:30:09 Fear of Russia and NATO's Protection Racket",
+ "00:30:28 Mafia Analogy for NATO Membership",
+ "00:31:16 Finland's Consideration to Join NATO",
+ "00:32:21 Anti-NATO Stance and Protection Racket Argument",
+ "00:34:02 Putin's Actions in Ukraine",
+ "00:35:05 Criticism of Putin's Military Strategy",
+ "00:36:17 Humanitarian Impact and Social Media Response",
+ "00:37:32 NATO Expansion Concerns"
+ ],
+ "boundary": {
+ "referenceCount": 33,
+ "generatedCount": 123,
+ "medianOffsetSeconds": 96,
+ "meanOffsetSeconds": 310.969696969697,
+ "withinThirtySeconds": 10,
+ "scores": [
+ {
+ "toleranceSeconds": 30,
+ "referenceCount": 33,
+ "generatedCount": 123,
+ "matched": 10,
+ "precision": 0.08130081300813008,
+ "recall": 0.30303030303030304,
+ "f1": 0.1282051282051282,
+ "matches": [
+ {
+ "reference": 3532,
+ "generated": 3528,
+ "distance": 4
+ },
+ {
+ "reference": 4003,
+ "generated": 4013,
+ "distance": 10
+ },
+ {
+ "reference": 4491,
+ "generated": 4472,
+ "distance": 19
+ },
+ {
+ "reference": 6355,
+ "generated": 6344,
+ "distance": 11
+ },
+ {
+ "reference": 7871,
+ "generated": 7874,
+ "distance": 3
+ },
+ {
+ "reference": 11312,
+ "generated": 11294,
+ "distance": 18
+ },
+ {
+ "reference": 14647,
+ "generated": 14630,
+ "distance": 17
+ },
+ {
+ "reference": 16752,
+ "generated": 16727,
+ "distance": 25
+ },
+ {
+ "reference": 20920,
+ "generated": 20940,
+ "distance": 20
+ },
+ {
+ "reference": 24232,
+ "generated": 24233,
+ "distance": 1
+ }
+ ]
+ },
+ {
+ "toleranceSeconds": 60,
+ "referenceCount": 33,
+ "generatedCount": 123,
+ "matched": 14,
+ "precision": 0.11382113821138211,
+ "recall": 0.42424242424242425,
+ "f1": 0.17948717948717946,
+ "matches": [
+ {
+ "reference": 3532,
+ "generated": 3528,
+ "distance": 4
+ },
+ {
+ "reference": 4003,
+ "generated": 4013,
+ "distance": 10
+ },
+ {
+ "reference": 4491,
+ "generated": 4472,
+ "distance": 19
+ },
+ {
+ "reference": 4824,
+ "generated": 4784,
+ "distance": 40
+ },
+ {
+ "reference": 5299,
+ "generated": 5333,
+ "distance": 34
+ },
+ {
+ "reference": 6355,
+ "generated": 6344,
+ "distance": 11
+ },
+ {
+ "reference": 7871,
+ "generated": 7874,
+ "distance": 3
+ },
+ {
+ "reference": 11216,
+ "generated": 11174,
+ "distance": 42
+ },
+ {
+ "reference": 11312,
+ "generated": 11294,
+ "distance": 18
+ },
+ {
+ "reference": 14647,
+ "generated": 14630,
+ "distance": 17
+ },
+ {
+ "reference": 16752,
+ "generated": 16727,
+ "distance": 25
+ },
+ {
+ "reference": 18566,
+ "generated": 18603,
+ "distance": 37
+ },
+ {
+ "reference": 20920,
+ "generated": 20940,
+ "distance": 20
+ },
+ {
+ "reference": 24232,
+ "generated": 24233,
+ "distance": 1
+ }
+ ]
+ }
+ ]
+ },
+ "speakersRendered": 0,
+ "titlesWithSpeakerName": 0
+ }
+ ],
+ "totals": {
+ "videos": 12,
+ "audioHours": 39.61,
+ "chunks": 87,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 7,
+ "zeroYieldRate": 0.0805,
+ "kept": 632,
+ "chaptersPerHour": 15.96,
+ "maxGapSeconds": 3112,
+ "meanGapSeconds": 1535,
+ "genericTitleRate": 0.1646,
+ "duplicateTitleRate": 0.0443,
+ "warningsByCode": {
+ "out-of-range": 65,
+ "language-drift": 22,
+ "non-monotonic": 23,
+ "seam-duplicate": 1
+ },
+ "rejectionRate": 0.1482,
+ "engineSeconds": 1843,
+ "tokensPerSecond": 211.6,
+ "secondsPerAudioHour": 47,
+ "projectedSweepDays": 43.1,
+ "boundary": {
+ "videosScored": 12,
+ "referenceBoundaries": 163,
+ "medianOffsetSeconds": 109,
+ "withinThirtySecondsRate": 0.227,
+ "byTolerance": [
+ {
+ "toleranceSeconds": 30,
+ "matched": 35,
+ "precision": 0.0556,
+ "recall": 0.2147,
+ "f1": 0.0883
+ },
+ {
+ "toleranceSeconds": 60,
+ "matched": 56,
+ "precision": 0.0889,
+ "recall": 0.3436,
+ "f1": 0.1412
+ }
+ ]
+ },
+ "speakerNameTitleRate": 0
+ }
+ }
+ ]
+}
diff --git a/plans/bakeoff/chapters-baseline.md b/plans/bakeoff/chapters-baseline.md
@@ -0,0 +1,160 @@
+# Digest bake-off — chapters-baseline
+
+Sample: 12 video(s) from `plans/bakeoff/sample.json` (buckets: short, medium, long, verylong), 39.61 audio-hours.
+Sweep days are projected as measured seconds-per-audio-hour x 79151 corpus audio-hours, one lane, no parallelism.
+
+| Candidate | Zero-yield chunks | Chapters/h | Rejection rate | Max gap | Generic | Dup | tok/s | s per audio-h | **Sweep days** |
+| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
+| `qwen2.5:7b@8192/chunk-local` | 7/87 (8.1%) | 15.96 | 14.8% | 00:51:52 | 16.5% | 4.4% | 211.6 | 47 | **43.1** |
+
+## Boundary accuracy vs the uploader's own chapters
+
+The oracle is `metadata.info.json.chapters` — marks a human authored while
+watching, who never saw our prompt. Boilerplate marks ("Intro", "Sponsor")
+and the boundary at 00:00 are dropped before scoring: neither carries
+segmentation information, and the origin would be a free hit for every
+candidate. Matching is one-to-one and closest-pair-first, so a cluster of
+boundaries around one uploader mark scores one match, not many.
+
+| Candidate | Videos scored | Uploader marks | Median offset | Within 30s | P@30 | R@30 | **F1@30** | F1@60 | Name-in-title |
+| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
+| `qwen2.5:7b@8192/chunk-local` | 12 | 163 | 109s | 22.7% | 5.6% | 21.5% | **8.8%** | 14.1% | 0% |
+
+## Rejections by guard
+
+| Candidate | language-drift | non-monotonic | out-of-range | seam-duplicate |
+| --- | --- | --- | --- | --- |
+| `qwen2.5:7b@8192/chunk-local` | 22 | 23 | 65 | 1 |
+
+## Per-video
+
+| Candidate | Video | Bucket | Chunks | Zero-yield | Chapters | Max gap | Engine s | Marks | F1@30 | Speakers |
+| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
+| `qwen2.5:7b@8192/chunk-local` | `nuxanor/9bgVoJY0jEg` | short | 1 | 0 | 4 | 00:05:39 | 16 | 9 | 16.7% | — |
+| `qwen2.5:7b@8192/chunk-local` | `chibi-reviews/XYyejb64QpU` | short | 1 | 0 | 2 | 00:07:28 | 3 | 8 | 20% | — |
+| `qwen2.5:7b@8192/chunk-local` | `the-quartering/br9koSf6D0Q` | short | 1 | 0 | 2 | 00:05:37 | 3 | 13 | 14.3% | — |
+| `qwen2.5:7b@8192/chunk-local` | `destiny/5nmDzKB23OU` | medium | 4 | 1 | 15 | 00:20:01 | 81 | 6 | 0% | — |
+| `qwen2.5:7b@8192/chunk-local` | `angryjoeshow/8AJsKyh0x7w` | medium | 2 | 1 | 3 | 00:34:34 | 48 | 5 | 50% | — |
+| `qwen2.5:7b@8192/chunk-local` | `kirsche/GchuqKQsnHk` | medium | 4 | 0 | 28 | 00:15:25 | 78 | 17 | 17.8% | — |
+| `qwen2.5:7b@8192/chunk-local` | `HasanAbiVODs/p7-GXvM3eXg` | long | 7 | 0 | 51 | 00:39:26 | 156 | 4 | 7.3% | — |
+| `qwen2.5:7b@8192/chunk-local` | `omnibased/S-9F_Jr3jmA` | long | 7 | 0 | 66 | 00:22:31 | 169 | 35 | 17.8% | — |
+| `qwen2.5:7b@8192/chunk-local` | `destiny/WL9htojKbGY` | long | 14 | 1 | 67 | 00:19:28 | 263 | 9 | 2.6% | — |
+| `qwen2.5:7b@8192/chunk-local` | `HasanAbiVODs/GjPX_ueTdfc` | verylong | 13 | 0 | 149 | 00:35:23 | 304 | 9 | 1.3% | — |
+| `qwen2.5:7b@8192/chunk-local` | `HasanAbiVODs3/bW2lKpHlmMQ` | verylong | 16 | 1 | 122 | 00:51:52 | 351 | 15 | 4.4% | — |
+| `qwen2.5:7b@8192/chunk-local` | `HasanAbiVODsbackup/5IufSgJSI0E` | verylong | 17 | 3 | 123 | 00:49:34 | 369 | 33 | 12.8% | — |
+
+## Sample output (first chapters per video)
+
+### `qwen2.5:7b@8192/chunk-local`
+
+**nuxanor/9bgVoJY0jEg** (short, 00:13:37)
+
+- 00:00:04 Introduction to the controversy
+- 00:05:02 Discussion on Alistair's design and Voodoo symbols
+- 00:10:40 Harassment and death threats
+- 00:12:51 Personal reflections on the issue
+
+**chibi-reviews/XYyejb64QpU** (short, 00:10:19)
+
+- 00:00:07 Villain Character Designs
+- 00:02:52 Media and News Impact on Villains
+
+**the-quartering/br9koSf6D0Q** (short, 00:10:06)
+
+- 00:00:01 Introduction and Media Criticism
+- 00:04:29 Prank Call Details and Backstory
+
+**destiny/5nmDzKB23OU** (medium, 01:13:39)
+
+- 00:00:06 Introduction and Initial Agreement
+- 00:01:42 Defining Religion and Atheism
+- 00:04:31 The Role of Reason in Belief
+- 00:07:44 Atheist vs. Agnostic Perspective
+- 00:10:35 Reasoning Processes and Their Differences
+- 00:13:51 Discussion on God of the Gaps Argument
+- 00:16:14 Impact of Organized Religion on Society
+- 00:19:29 Introduction and Skepticism
+
+**angryjoeshow/8AJsKyh0x7w** (medium, 00:50:41)
+
+- 00:34:34 Gameplay Issues
+- 00:43:21 Endgame and Crafting Issues
+- 00:47:11 Final Verdict and Future Prospects
+
+**kirsche/GchuqKQsnHk** (medium, 01:09:15)
+
+- 00:15:16 Shooting at Wazada High School Graduation Ceremony
+- 00:15:35 University of Minnesota Police Response
+- 00:16:30 Surveillance Video and Suspect Description
+- 00:17:46 Suspect's Arrest and Charges
+- 00:19:46 University of Minnesota Decision to Stop Hosting Graduations
+- 00:21:06 Impact on Local School Districts
+- 00:21:52 Anoka-Henipin School District's Response
+- 00:37:17 Introduction of the incident
+
+**HasanAbiVODs/p7-GXvM3eXg** (long, 04:18:43)
+
+- 00:23:08 Initial frustration and request to stop
+- 00:23:11 Discussion of RP frogs
+- 00:23:14 Request for an ad break
+- 00:23:18 Realistic portrayal of money laundering in movies
+- 00:24:01 Money laundering basics
+- 00:25:13 Realistic portrayal of money laundering in 'Gotcha'
+- 00:27:02 Drug money laundering through gems and emeralds
+- 00:28:02 Realistic portrayal of drug dealer's cash stashes
+
+**omnibased/S-9F_Jr3jmA** (long, 03:31:45)
+
+- 00:00:48 Introduction and Opening Remarks
+- 00:02:23 Discussion on the Situation in Minneapolis
+- 00:05:16 Response to Questions about ICE Operations
+- 00:08:17 Explanation of the Shooting Incident
+- 00:14:44 Discussion on the Storm and Emergency Preparedness
+- 00:19:33 Response to Questions about the Shooting Incident
+- 00:25:45 Conclusion and Closing Remarks
+- 00:29:02 Introduction and Setup
+
+**destiny/WL9htojKbGY** (long, 03:57:39)
+
+- 00:19:29 Introduction and Context
+- 00:22:03 Discussion on Asmin Gold's Views
+- 00:24:46 Criticism of the Video Content
+- 00:27:58 Investigations and Prosecutions under Biden Administration
+- 00:32:52 Fraud Allegations in Daycares and Charities
+- 00:36:14 Investigation Outcome Uncertainty
+- 00:37:21 Introduction and Context
+- 00:37:58 Entertainment vs. Journalism Debate
+
+**HasanAbiVODs/GjPX_ueTdfc** (verylong, 08:02:45)
+
+- 00:35:24 Stream Snipers and Stream Saviors
+- 00:36:13 Conservative Floridian Mom Lobster
+- 00:37:50 Bondo Simulator 2021
+- 00:39:05 Boomer of the Year
+- 00:41:15 Compliments and Reactions
+- 00:45:28 Chat Union Demands
+- 00:49:12 Food Cravings
+- 00:53:25 Sweat Dripping Down
+
+**HasanAbiVODs3/bW2lKpHlmMQ** (verylong, 07:59:00)
+
+- 00:07:55 Introduction and Podcast Update
+- 00:09:20 Kyrie Irving's Performance in Game 7
+- 00:12:31 LeBron James' Jab Step and Its Impact on the Game
+- 00:16:24 Discussion on Kyrie Irving's Shot Selection
+- 00:23:22 Podcast Episode Rankings and Pod Save America
+- 00:31:41 Support for Raes Texas
+- 00:35:56 TikTok Video with Twin Peaks Music
+- 00:39:44 Thanksgiving and Business Continuation
+
+**HasanAbiVODsbackup/5IufSgJSI0E** (verylong, 07:59:07)
+
+- 00:30:09 Fear of Russia and NATO's Protection Racket
+- 00:30:28 Mafia Analogy for NATO Membership
+- 00:31:16 Finland's Consideration to Join NATO
+- 00:32:21 Anti-NATO Stance and Protection Racket Argument
+- 00:34:02 Putin's Actions in Ukraine
+- 00:35:05 Criticism of Putin's Military Strategy
+- 00:36:17 Humanitarian Impact and Social Media Response
+- 00:37:32 NATO Expansion Concerns
+
diff --git a/plans/bakeoff/speaker-sample.json b/plans/bakeoff/speaker-sample.json
@@ -0,0 +1,145 @@
+{
+ "version": 1,
+ "pickedAt": "2026-08-12T18:10:34.039Z",
+ "corpus": {
+ "videosScanned": 76354,
+ "videosWithTranscript": 73367,
+ "audioHours": 77298,
+ "longTailVideos": 6038,
+ "longTailAudioHours": 40153
+ },
+ "videos": [
+ {
+ "slug": "shondo/uVYjZKWEWmk",
+ "channelSlug": "shondo",
+ "videoId": "uVYjZKWEWmk",
+ "videoDir": "uVYjZKWEWmk",
+ "title": "1000 Mwahs To Help You Sleep ASMR! ❤️ Personal Attention, Silliness & Tingles!",
+ "bucket": "medium",
+ "durationSeconds": 3622,
+ "cueCount": 0,
+ "uploaderChapters": 13
+ },
+ {
+ "slug": "shondo/KJCmnasraVg",
+ "channelSlug": "shondo",
+ "videoId": "KJCmnasraVg",
+ "videoDir": "KJCmnasraVg",
+ "title": "ASMR Sus Doctor Pokes Your Brain & Examines You! 💉Weird Personal Attention, Gloves, Squishy Noises!?",
+ "bucket": "medium",
+ "durationSeconds": 4646,
+ "cueCount": 0,
+ "uploaderChapters": 15
+ },
+ {
+ "slug": "kirsche/pPt39Tf-vNo",
+ "channelSlug": "kirsche",
+ "videoId": "pPt39Tf-vNo",
+ "videoDir": "pPt39Tf-vNo",
+ "title": "DEIagnosis: Incompetence and Inclusion",
+ "bucket": "medium",
+ "durationSeconds": 3632,
+ "cueCount": 0,
+ "uploaderChapters": 11
+ },
+ {
+ "slug": "shondo/pl0Wt0P3J4w",
+ "channelSlug": "shondo",
+ "videoId": "pl0Wt0P3J4w",
+ "videoDir": "pl0Wt0P3J4w",
+ "title": "ASMR Strange Doctor Treats You Like A Big Dummy! 🥴💉 Personal Attention, Gloves, Hearing & Eye Exam!",
+ "bucket": "medium",
+ "durationSeconds": 3511,
+ "cueCount": 0,
+ "uploaderChapters": 10
+ },
+ {
+ "slug": "kirsche/Y9hdSH8Ky10",
+ "channelSlug": "kirsche",
+ "videoId": "Y9hdSH8Ky10",
+ "videoDir": "Y9hdSH8Ky10",
+ "title": "Do NOT The Cousin",
+ "bucket": "medium",
+ "durationSeconds": 3319,
+ "cueCount": 0,
+ "uploaderChapters": 8
+ },
+ {
+ "slug": "shondo/ZUh5_uvFjEs",
+ "channelSlug": "shondo",
+ "videoId": "ZUh5_uvFjEs",
+ "videoDir": "ZUh5_uvFjEs",
+ "title": "[ASMR] Yandere Twins Keep You For Themselves 💕🔪 Headpats, Ear Cleaning, Personal Attention! ft Silvi",
+ "bucket": "medium",
+ "durationSeconds": 3640,
+ "cueCount": 0,
+ "uploaderChapters": 8
+ },
+ {
+ "slug": "shondo/CIU_khXnH5I",
+ "channelSlug": "shondo",
+ "videoId": "CIU_khXnH5I",
+ "videoDir": "CIU_khXnH5I",
+ "title": "Welcome to the ASMR Maid Cafe!! ✨ ✨ 1 Hour of Twin Ear Cleaning, Ear Massage & Personal Attention!!",
+ "bucket": "medium",
+ "durationSeconds": 4863,
+ "cueCount": 0,
+ "uploaderChapters": 8
+ },
+ {
+ "slug": "shondo/OWuZuOrbQ10",
+ "channelSlug": "shondo",
+ "videoId": "OWuZuOrbQ10",
+ "videoDir": "OWuZuOrbQ10",
+ "title": "[ASMR] Asking You Lots Of Questions | THE SURVEY [Keyboard Typing] [Softly Spoken]",
+ "bucket": "medium",
+ "durationSeconds": 2546,
+ "cueCount": 0,
+ "uploaderChapters": 5
+ },
+ {
+ "slug": "shondo-vods/Iq3svKzK-yQ",
+ "channelSlug": "shondo-vods",
+ "videoId": "Iq3svKzK-yQ",
+ "videoDir": "Iq3svKzK-yQ",
+ "title": "SHONDO COMEBACK!! NEW OUTFIT NEW EVERYTHING NEW LORE HUGE ANNOUNCEMENTS VERY EXCITING AAAAH HI HI HI",
+ "bucket": "long",
+ "durationSeconds": 14515,
+ "cueCount": 0,
+ "uploaderChapters": 19
+ },
+ {
+ "slug": "shondo-vods/F7_IyckDc3M",
+ "channelSlug": "shondo-vods",
+ "videoId": "F7_IyckDc3M",
+ "videoDir": "F7_IyckDc3M",
+ "title": "IMOUTO WIFE KARAOKE DATE!!! LET'S SING SONGS, CATCH UP AND HAVE FUN TODAY!!! 🎀 vtuber!",
+ "bucket": "long",
+ "durationSeconds": 17638,
+ "cueCount": 0,
+ "uploaderChapters": 45
+ },
+ {
+ "slug": "shondo-vods/1d1BEZMo2F0",
+ "channelSlug": "shondo-vods",
+ "videoId": "1d1BEZMo2F0",
+ "videoDir": "1d1BEZMo2F0",
+ "title": "SDFIOGHJDROIGHEDRFOIUGHDUFHFGFSDUIHFGDSUIFGHDUIFHGDIFGJDPOFGUJJID-R9090T8ER09TUFGKBDNFHDSFGOIDSFHJ ♥",
+ "bucket": "verylong",
+ "durationSeconds": 21875,
+ "cueCount": 0,
+ "uploaderChapters": 27
+ },
+ {
+ "slug": "shondo-vods/CcSsT1cSn8E",
+ "channelSlug": "shondo-vods",
+ "videoId": "CcSsT1cSn8E",
+ "videoDir": "CcSsT1cSn8E",
+ "title": "CHRISTMAS COUNTDOWN WITH YOUR IMOUTOWIFE!!! WE'RE WAITING FOR SANTA!!! 🎀 vtuber!",
+ "bucket": "long",
+ "durationSeconds": 14308,
+ "cueCount": 0,
+ "uploaderChapters": 9
+ }
+ ]
+}
diff --git a/plans/bakeoff/speakers-round1-notes.md b/plans/bakeoff/speakers-round1-notes.md
@@ -0,0 +1,266 @@
+# Better digest inputs — round 1 notes
+
+Written 2026-08-12. Companion to the harness-generated reports in this directory.
+`chapters-baseline.{json,md}` is the Stage 2a result; `speakers-round1.{json,md}` is
+reserved for the paired A/B, which has **not** run (see "What did not happen").
+
+The plan this implements asked for a number, not a feature. Here is what was bought.
+
+---
+
+## 1. The corpus facts, re-measured rather than assumed
+
+Every load-bearing figure in the plan was re-derived from disk. They hold:
+
+| | plan said | measured | |
+| --- | --- | --- | --- |
+| videos with a non-empty uploader `chapters` array | ~19.5% sampled | **14,923** | ✓ |
+| + a transcript + >=4 chapters (the oracle pool) | ~11,366 | **12,139**, across **43 channels** | ✓ |
+| + retained audio (the Stage 2b pool) | 34 | **34** | ✓ exact |
+| Stage 2b pool by channel | shondo-vods 20, shondo 7, HasanAbiVODs3 3, kirsche 2, omnibased 2 | identical | ✓ exact |
+| already-digested videos with >=4 marks | 14, all omnibased | **15**, all omnibased | ✓ |
+
+Median 7 uploader marks per video in the oracle pool, max 136.
+
+## 2. What was built
+
+**`common/lib/boundaryScore.ts`** — the metric the harness did not have. Every score
+`digest-bakeoff.ts` produced before this one (`zeroYieldRate`, `chaptersPerHour`,
+`rejectionRate`, `maxGapSeconds`, `genericTitleRate`, `duplicateTitleRate`) is a
+*defect counter*: it says how malformed the output is, never whether a boundary is in
+the right place. So the harness could rank candidates by which was least broken and
+still not say which segmented better.
+
+Design decisions that matter, each pinned by a unit test (12 of them):
+
+- **The oracle is independent of the experiment.** Scoring "chapters align with speaker
+ changes" against a variant *given* the speaker changes proves nothing. Uploader
+ chapters are authored by a human who never saw our prompt.
+- **The t=0 boundary is dropped.** Nearly every uploader list opens at 00:00 and every
+ segmentation trivially has a boundary there. Counting it hands both arms of any A/B a
+ free true positive worth ~1/n. A metric that cannot be lost is not a measurement.
+- **Matching is one-to-one, closest pair first.** Ten boundaries crammed around one
+ uploader mark score *one* match, not ten — that over-segmentation failure is exactly
+ what precision exists to punish, and a nearest-neighbour count would reward it.
+- **Boilerplate marks are dropped** ("Intro", "Sponsor", "Outro"). A boundary in the ad
+ read is a boundary in the furniture, not the subject.
+- **Uploader *titles* are never scored against.** Boundaries are facts; titles are the
+ uploader's expression. This is what keeps the metric free of the copyright objection
+ that rules out admitting their chapter text.
+
+**`readUploaderChapters`** in `common/lib/videoStatus.ts`, beside `readVideoDurationSec`
+and using the existing `META_FILENAME` — the location the plan specified, so no second
+path to this data can grow later. Nothing in the codebase read `chapters` before this.
+
+**Bake-off integration** — pooled precision/recall/F1 at 30 s and 60 s (pooled, not a
+mean of per-video rates, so a 3-mark video cannot outweigh a 29-mark one), median
+offset, a `--pick-chapters` sampler, a `--score-existing` zero-GPU mode, a
+`--speakers on,off` axis, and `--max-cues` for A/B headroom.
+
+### Two production files were touched, both additively and both pinned
+
+The plan scoped code to the bake-off. Two exceptions were unavoidable to run the
+experiment honestly, and each renders **byte-identically when unused**:
+
+- `transcriptToMarkdown.prefixForCue` — a per-line prefix hook. The label cannot go
+ inside the timestamp bracket: the prompt tells the model to copy `[HH:MM:SS]` markers
+ verbatim and `HMS_PATTERN` rejects the echo. Returning null for every cue reproduces
+ the current line exactly (test: *"prefixForCue returning null for every cue is
+ byte-identical to omitting it"*).
+- `ChapterPromptInput.speakerRoster` — a per-chunk cast list. Deliberately *not*
+ `contextNote`: that is per-channel, hand-authored and hashed into `contextHash`, and
+ folding a derived per-video roster into it would mislabel it to the model and make a
+ channel note and a roster mutually exclusive (test: *"a prompt with no speakerRoster
+ renders byte-identically to today"*).
+
+**No `PROMPT_VERSION` bump, and none is needed** — the shipped lane sets neither field,
+so every digest the sweep produces is byte-for-byte what it produced before. That is the
+same compatibility argument that let `timestampMode` in at version 1.
+
+## 3. Stage 2a — the baseline, and it caught a real bug on day one
+
+**Free result first.** `--score-existing` scores digests already on disk. Over the 15
+scorable ones, with no model calls and no writes:
+
+> **P@30 10.8%, R@30 26.1%, R@60 35.9%, F1@30 15.2%** across 153 uploader marks.
+
+Per-video F1@30 ranged 0% to 50% — the metric discriminates rather than saturating.
+
+**Precision is partly a density artifact and must not be read naively.** The digest
+targets one chapter per 4 minutes and is deliberately denser than uploader chaptering
+(372 generated vs 153 uploader marks here). Recall is the half that answers "did it find
+the human's boundaries"; precision should be read together with `chaptersPerHour`.
+
+**The metric immediately found a defect.** `omnibased/3fl_uRFXxLk`: the video runs
+0–5119 s and the uploader marked chapters across all of it, but the digest emitted
+chapters only between 3184 s and 4865 s — it summarized the last third and skipped the
+rest. Median offset 1916 s, F1 0%. `maxGapSeconds` corroborates, but boundary scoring
+quantified it independently. That is the metric earning its keep before any A/B.
+
+The generated baseline runs over a fresh 12-video, 8-channel, 4-bucket sample
+(`chapter-sample.json`, 170 uploader marks) at the **production configuration**
+(`chunk-local`, 8192, 600 cues) — the bake-off's own default is `absolute`, which is
+*not* what the sweep runs. Re-run it with:
+
+```sh
+cd common && pnpm exec tsx bin/digest-bakeoff.ts --label chapters-baseline \
+ --sample ../plans/bakeoff/chapter-sample.json --modes chunk-local \
+ --candidates 'qwen2.5:7b@8192'
+```
+
+**Final result, all 12 videos, 8 channels, 163 scorable uploader marks, 87 chunks:**
+
+| | full sample (12 videos, 8 channels) | short/medium only (5 videos) | `--score-existing` (15 videos, omnibased) |
+| --- | ---: | ---: | ---: |
+| **R@30** | **21.5%** | 12.2% | **26.1%** |
+| R@60 / F1@60 | 14.1% (F1) | 19.5% (R) | 35.9% (R) |
+| P@30 | 5.6% | 20.8% | 10.8% |
+| F1@30 | 8.8% | 15.4% | 15.2% |
+| median offset | 109 s | — | — |
+| within 30 s of a mark | 22.7% | — | — |
+
+Guard rails on the same run: 8.1% zero-yield (7/87 chunks), 15.96 chapters/h, 14.8%
+rejection, 16.5% generic titles, 47 s/audio-hour → **43.1 projected sweep days**
+(against `digest-plan`'s 53.8; this run was partly contended, so read it as the same
+order, not as a refinement). Name-in-title is 0%, correctly — this arm has no roster.
+
+**READ RECALL. PRECISION IS MOSTLY A DENSITY ARTIFACT, AND THIS TABLE PROVES IT.**
+Recall@30 lands at 21.5% / 12.2% / 26.1% across three samples that share no videos — a
+narrow band. Precision swings 5.6% → 20.8% over the *same* metric on the *same* engine,
+purely with the duration mix: on short videos the digest emits fewer boundaries than the
+uploader (24 generated vs 41 marks), on 8-hour VODs far more (149 vs 9). F1 inherits that
+swing, which is why the full-sample F1 (8.8%) reads worse than either subsample without
+the digest having gotten worse at anything.
+
+Practical rule for any future candidate comparison: **compare on recall, and read
+`chaptersPerHour` beside it.** A variant that "improves F1" by matching the sample's
+density has improved nothing.
+
+### Checkpointing was added because a long run cannot be trusted to reach its own end
+
+The reports are only written when the whole run finishes. A first attempt spent ~26
+minutes of engine time and produced nothing readable, and the postmortem was itself
+instructive: it had **not** crashed. Three bake-off runs were alive at once — a relaunch
+issued on the mistaken belief the first had died, which truncated the first's log and
+made it look restarted. The three starved each other. With exactly one run, per-chunk
+time fell from **20-42 s to 3 s**.
+
+Two lessons, both now encoded: `scoreVideo` results are checkpointed to
+`<label>.partial.json` after **every video**, so a kill leaves usable data; and
+`pkill -f <pattern>` on this box will match *its own invoking shell* when the pattern
+appears in the command line — it killed the shell that issued it earlier in the same
+session. Kill by PID.
+
+## 4. What did not happen, and why
+
+**The paired speakers-on/off A/B did not run.** The sample is built
+(`speaker-sample.json`: 12 videos, 27.3 audio-hours, 178 uploader marks, stratified by
+duration and weighted toward marks-per-CPU-hour) and the pipeline is proven end to end,
+but the diarization it depends on was stopped deliberately.
+
+**Why stopped.** The plan says to confirm the box is not already loaded before
+committing diarization, because that precondition has invalidated measurements in this
+project before. It was loaded: with a pre-existing `diarize-channel` job already running,
+adding a second sherpa process took the box to **262 MB free, 14.1 GB swapped, and
+memory-pressure full-stall at 2.6%** — the same pressure that had earlier forced ollama
+off the GPU. Continuing would have degraded the operator's own job and corrupted the
+bake-off running beside it. One video was diarized before stopping; the rest are queued.
+
+**A related finding worth more than the A/B would have been.** The live sweep was
+running at **~200 s/chunk against its own 24.7 s/chunk projection** — `ollama ps` showed
+`85%/15% CPU/GPU`, i.e. memory pressure had pushed the model mostly onto CPU (prefill
+~43 tok/s). With the sweep and backfill paused and memory freed, the same model loads
+**100% GPU**. The 53.8-day projection was really on the order of **400+ days at the
+contended rate**. Concurrency on this box is not free; it is roughly an 8× tax on the
+sweep.
+
+## 5. What the one diarized video already shows
+
+`shondo/OWuZuOrbQ10` (42 min) was diarized (445 s) and attributed via the shipped
+`--method diarized` lane. Rendering both arms and diffing, with no model call:
+
+- 8 diarization clusters collapsed to **3 named speakers**
+- **98.8% of cues** carried a label; the speaker changed on **16.9%** of cues
+- **prompt inflation 4.4%** — comfortably below the ~10% the plan budgeted, because
+ labels are emitted on speaker *change* only, not per line
+
+Two observations that shape whether the A/B is even worth running:
+
+1. **The labels are roles, not names**: "Host", "Caller", "Interviewer". The diarized
+ lane names a role when the transcript does not support a proper name. So on this kind
+ of material the *boundary-detection* half of the hypothesis is testable, but the
+ *title-quality* half ("Guest X on Y" beats "Discussion") is **not** — there is no name
+ to put in a title. Name-uptake will read ~0 here for a reason that is not a failure of
+ the prompt.
+2. **The labels invent speaker changes.** The rendered opening reads
+ `Caller: Hello` / `Host: and thank you for volunteering...` and later
+ `Caller: The first answer that occurs to you...` — one continuous narration split
+ across two labels. False speaker changes are *anti-signal* for boundary detection:
+ they would push the model to segment where nothing changed. This is the over-split
+ problem the top-12/>=1% filter is supposed to contain, surviving the filter.
+
+Neither observation settles the question at n=1. Both say the A/B should be read for
+*harm* as carefully as for benefit, and both were invisible before the render existed.
+
+## 6. Running the A/B when the box is quiet
+
+```sh
+cd common
+# 1. Data prep (writes sidecars; ~27 audio-hours, ~4 CPU-hours on a CLEAN box).
+# diarize-backfill skips anything already fresh, so this is safe to re-run.
+for v in shondo/uVYjZKWEWmk kirsche/Y9hdSH8Ky10 shondo/pl0Wt0P3J4w shondo/ZUh5_uvFjEs \
+ kirsche/pPt39Tf-vNo shondo/FC57QxT8jUU shondo/KJCmnasraVg shondo/CIU_khXnH5I \
+ shondo-vods/CcSsT1cSn8E shondo-vods/Iq3svKzK-yQ shondo-vods/F7_IyckDc3M \
+ shondo-vods/1d1BEZMo2F0 ; do
+ pnpm exec tsx bin/diarize-backfill.ts --scope "video:$v"
+ pnpm exec tsx bin/attribution-pilot.ts --sample "video:$v" --method diarized --label speakers-prep
+done
+
+# 2. The paired A/B. Both arms get the SAME cue count with context headroom, so the
+# speakers-on arm cannot be the only one that truncates.
+pnpm exec tsx bin/digest-bakeoff.ts --label speakers-round1 \
+ --sample ../plans/bakeoff/speaker-sample.json \
+ --modes chunk-local --candidates 'qwen2.5:7b@16384' --max-cues 600 --speakers off,on
+```
+
+**Check the box first** (`free -m`, `/proc/pressure/memory`, `ollama ps` showing
+`100% GPU`). If `ollama ps` says anything other than 100% GPU, stop — the timing numbers
+will be meaningless and the run will take ~8× longer.
+
+### Headroom is not optional
+
+Measured across the 34-video Stage 2b pool, a 600-cue chunk at the ~4 chars/token ratio
+the live sweep log implies (4,098 tokens for a 600-cue chunk) puts **24 of 34 over
+~6,000 transcript tokens** — already at or past what an 8192 window holds once prompt and
+response are allowed for, *before* any speaker labels. **ollama truncates silently.** Run
+the A/B at 8192 and it measures truncation, not speakers. Hence `@16384 --max-cues 600`
+for both arms, recorded in the candidate key.
+
+### How to read the result
+
+- **Recall@30 and @60** are the headline. Precision moves with density.
+- **The single-speaker controls are the honesty check.** If speakers-on changes output
+ on a video with one speaker at all, the render is leaking noise, not signal.
+- **`genericTitleRate` and `projectedSweepDays` are guard rails.** A variant that
+ improves boundaries but costs 20% more chunks still loses.
+- **The sample is 10 of 12 videos from two creators** (shondo/shondo-vods, kirsche).
+ A win is a signal to widen the sample, **never** a corpus-wide result. That constraint
+ is structural: only 34 videos in the entire archive have audio, a transcript and >=4
+ uploader chapters.
+
+## 7. Standing recommendation
+
+Ordered by value per GPU-second, on the evidence now in hand:
+
+1. **Promote channel context notes.** Eleven candidates are drafted and cover ~60% of
+ remaining sweep work. Zero GPU, zero code, invalidates only the channel it touches,
+ and reaches 100% of the corpus rather than the 1.6% that can ever have speaker data.
+ See `plans/digest-context-review.md`.
+2. **Keep using the boundary metric.** It is now the only quality measure the harness
+ has, it costs nothing on already-generated digests, and it found a real coverage bug
+ on its first run.
+3. **Fix the concurrency tax before spending more GPU.** An 8× throughput swing from
+ memory pressure dwarfs every prompt-level improvement under discussion.
+4. **The speaker A/B is last**, and its ceiling is low: 1.6% of the archive retains
+ audio, the labels observed so far are roles rather than names, and they already
+ introduce false speaker changes. Run it to bound the harm as much as to find a win.
diff --git a/plans/digest-context-review.md b/plans/digest-context-review.md
@@ -0,0 +1,144 @@
+# Channel context notes — candidates awaiting review
+
+Eleven candidate `digest-context.suggested.md` notes, drafted 2026-08-12 from corpus
+evidence. **None is active.** A note only reaches the prompt under the name
+`digest-context.md`, and only a human puts it there:
+
+```sh
+cd transcripts/channels/<slug>
+mv digest-context.suggested.md digest-context.md
+```
+
+Auto-applying model-authored context is forbidden (PLAN.md:245-248), and the reason is
+specific: a wrong inference here degrades every future summary for the channel, and the
+output still looks plausible. So each note below lists what it asserts and what that
+assertion rests on. **Check the identity claims first** — they are the ones that would
+do damage if wrong.
+
+## Why these channels, in this order
+
+Ordered by remaining sweep work from `tsx bin/digest-plan.ts` (the sweep's own queue
+order). The eleven cover **~113,000 of the 188,269 chunks still to generate — about 60%
+of the remaining sweep.**
+
+| # | Channel | Chunks remaining | Note |
+| ---: | --- | ---: | --- |
+| 1 | `omnibased` | 16,747 | drafted |
+| 2 | `HasanAbiVODs3` | 14,496 | drafted |
+| 3 | `destiny` | 13,775 | drafted |
+| 4 | `rekietalaw` | 12,518 | drafted |
+| 5 | `the-quartering` | 11,604 | drafted |
+| 6 | `chibi-reviews` | 10,129 | drafted |
+| 7 | `chrissie-mayr` | 9,297 | drafted |
+| 8 | `HasanAbiVODs` | 9,146 | drafted |
+| 9 | `kirsche` | 8,063 | drafted |
+| 10 | `nuxanor` | 7,436 | drafted |
+| 15 | `HasanAbiVODsbackup` | 4,451 | drafted (same show as #2) |
+
+## Evidence behind each note
+
+Drafted from a deterministic 150-video spread per channel (stable-hash ordering, so it
+is not the first 150 on disk), counting proper-noun phrases once per video across
+titles and descriptions, plus a mid-video transcript slice from 25 of them.
+
+### `omnibased` — **check this one first**
+- **Claims it is a VOD archive of Destiny (Steven Bonnell II).** Rests on: titles
+ "Destiny Playing Diablo III", "Destiny Duo FPP", "@ destiny gg embed chat"; the
+ sibling channel `destiny` credits "TheOmniLiberal" (Bonnell's handle) in 23 of 150
+ descriptions, which is also what the "omni" in `omnibased`/`omnimirror`/
+ `omnivods-odysee` points at. Strong, but it is an inference — confirm it.
+- **The segment taxonomy is not invented.** 11 of 150 descriptions carry "Chapters
+ (auto-generated from community timestamps)" using exactly these labels: `AFK:`,
+ `Chatting:`, `Gaming:`, `Clicking Links:`, `Watching:`, `Research:`, `Reading:`,
+ `Tech Support:`, `Debate:`, `Other:`, `/r/LSF:`. This is the channel's own
+ vocabulary, which is why the note is worth more here than anywhere else.
+- Games and topics listed are all observed in sampled titles.
+
+### `HasanAbiVODs3`, `HasanAbiVODs`, `HasanAbiVODsbackup`
+- Same show (HasanAbi / Hasan Piker), three archive channels; the note is shared, with
+ `HasanAbiVODs` additionally describing its `1/2` / `2/2` split uploads, "FULL VOD"
+ marking, and NoPixel GTA segments — all observed only on that channel.
+- "HasanAbi" appears in 150/150 descriptions on VODs3. Will Neff and AustinShow are the
+ only recurring non-Hasan people in sampled titles.
+- Political figures listed are the ones actually recurring in the transcript sample.
+
+### `destiny`
+- Distinguished from `omnibased` deliberately: this channel's uploads are **edited
+ excerpts**, evidenced by editor credits (Voddity, SL5000), "Date streamed", and short
+ single-subject titles. The note tells the model not to over-split, which is the
+ opposite of the advice `omnibased` gets. **If the archive/main-channel split is wrong,
+ both notes are wrong together.**
+
+### `rekietalaw`
+- Nick Rekieta, lawyer: "Rekieta Law" in 22/150 titles, "RekietaMedia" in descriptions.
+- Named recurring people (Ty Beard 4, Andrew Branca, Dick Masterson, Maddox, Drex) and
+ the case list are all from sampled titles.
+- The "boundaries follow the court's schedule" advice is a judgement call, not evidence.
+
+### `the-quartering`
+- Jeremy Hambly. Note the channel slug and the separate `jeremy-hambly`,
+ `quartering-live`, `quarteringvlogs` channels — those have **no note yet**.
+- Sampled titles yielded almost no recurring proper nouns, which is itself the finding:
+ the value here is the format advice (one story per video, do not manufacture topic
+ changes) rather than a cast list.
+
+### `chibi-reviews`
+- "Anime Review" in 56/150 titles, "Manga Review" in 22. The title-pattern claim is the
+ strongest evidence of any note here.
+- The advice to prefer the title's spelling of Japanese names over the transcript's is a
+ judgement call aimed at a known ASR weakness; it is untested.
+
+### `chrissie-mayr`
+- "Chrissie Mayr" 58/150 titles, "CMP" 26, "SimpCast" 22.
+- Co-hosts listed (Brittany Venti 9, Lila Hart 9, Keanu Thompson 7, Anna TSWG, Melonie
+ Mac, Riss Flex) are all from sampled titles. This is the best-evidenced cast list.
+
+### `kirsche`
+- Kirsche Verstahl, VTuber. "Foxu News" in 4+3 sampled titles; Gamersupps, tip jar and
+ artist credits in ~140/150 descriptions.
+- The observation that titles are often contentless is evidence-based and is the most
+ useful line in the note.
+
+### `nuxanor`
+- Nux / Nux Taku. "Nuxanor" and "NuxTaku" in 145/150 descriptions; Live2D credits
+ (AkatsukiEnma, Ironvertex) in 141.
+- "Stevie Blunder" appears in 92/150 descriptions as a ghost writer credit. The note
+ does **not** claim he speaks in the videos — worth keeping that way unless you know
+ otherwise.
+
+## What to check before promoting
+
+1. **Identity claims** — `omnibased` = Destiny's archive, and `destiny` = his main
+ edited channel. Everything else in those two notes hangs off it.
+2. **Cast lists** — a name that is actually a one-off guest, listed as a regular, tells
+ the model to expect someone who is not there.
+3. **Anything phrased as advice** ("do not over-split", "prefer the title's spelling")
+ is my judgement, not measured. Cut anything you disagree with; a shorter true note
+ beats a longer speculative one.
+
+## What promoting one costs, and how to tell it worked
+
+Promoting a note changes that channel's `contextHash`, so **every digest already
+generated for that channel becomes stale and regenerates.** That is the intended
+behaviour (`digestContext-server.ts:45-55`), and it is cheap now and expensive later:
+only 390 videos corpus-wide are digested today, 15 of them on `omnibased`.
+
+Nothing else is touched — `contextHash` is per channel, so no other channel goes stale.
+
+To measure the effect, per the plan:
+
+```sh
+cd common
+pnpm exec tsx bin/digest-validate.ts <slug> # before promoting, and after regenerating
+```
+
+Compare generic-title rate and coverage gap. Boundary accuracy is now also measurable
+on the ~12,139 videos that carry uploader chapters:
+
+```sh
+cd common
+pnpm exec tsx bin/digest-bakeoff.ts --score-existing # no model calls, no writes
+```
+
+Baseline for that command as of 2026-08-12, over the 15 scorable digests on disk:
+**R@30 26.1%, R@60 35.9%, F1@30 15.2%** across 153 uploader marks.