Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 01c3073598b4a553981ac11ad8c6daf0c6986789
parent 3598bf4dc79149beabca5a701dd115adeaf9c823
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon, 27 Jul 2026 00:08:59 -0400

Digest layer Stage B1: engine registry, batch runner, guards, e2e

Completes the per-video AI digest layer so a sweep is startable from a
browser and provable end to end.

What lands:

- Engine registry + per-app config (digestApps.ts, settings.digest.apps),
  with numCtx reaching ollama explicitly. Its 4096 default truncates
  silently and the model then summarizes whatever fragment survived —
  measured, and the easiest way to get quietly-wrong output at scale.
- Chunk size derived FROM the configured context (maxCuesForContext), for
  the same reason.
- Two timestamp numberings, both shipped: "absolute" and "chunk-local".
  Which is better was a measured question, not a hunch, so a bake-off
  harness (common/bin/digest-bakeoff.ts) settles it rather than one being
  pre-applied as a fix. Reports in plans/bakeoff/.
- The prompt shape is folded into the recorded identity via
  digestPromptVariant(), so switching modes or chunk sizes invalidates the
  corpus instead of skipping it as fresh.
- Parser guards (range clamp, monotonicity, seam de-dup, language drift,
  empty title), each rejection recorded as a warning on the artifact
  rather than dropped — a multi-week sweep is only tunable if its failures
  are inspectable.
- runDigestBatch + a Digest stage on the channel page (count, lane picker,
  order, limit, queue control), on the two lanes' separate queue keys.

Two wiring bugs, both of which only e2e could have caught:

- taskHooks.ts fed log lines through `downloadParser!.feed(part)`. Both
  parsers are null for a digest task, so it threw on the FIRST log line of
  every digest job; the batch caught it per-video, so a whole sweep would
  have reported "0 generated, N failed" and looked like an engine fault.
- saveSettingsAction rebuilt settings.digest without carrying
  timestampMode/promptVariant through, so an unrelated Settings save would
  have silently reset the prompt shape — invalidating every digest
  generated under a non-default one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Diffstat:
Acommon/bin/digest-bakeoff.ts | 716+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/controller/digestBatch.ts | 21++++++++++++++++++++-
Mcommon/controller/digestVideo.ts | 43+++++++++++++++++++++++++++++++++++++++----
Mcommon/jobs/taskHooks.ts | 7++++++-
Mcommon/lib/digest.test.ts | 93+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/lib/digest.ts | 105++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-
Mcommon/lib/digestApps.ts | 13++++++++++---
Mcommon/lib/digestParse.test.ts | 134+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/lib/digestParse.ts | 47++++++++++++++++++++++++++++++++++++++---------
Acommon/lib/digestPrompt.test.ts | 101+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/lib/digestPrompt.ts | 67+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++------
Mcommon/lib/settings.ts | 27+++++++++++++++++++++++++++
Meditor/CHANGELOG.md | 1+
Aeditor/app/channels/[slug]/components/stages/DigestStage.tsx | 165+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/app/channels/[slug]/lib/stageStatus.ts | 34++++++++++++++++++++++++++++++++++
Meditor/app/channels/[slug]/page.tsx | 19++++++++++++++++++-
Meditor/app/settings/actions.ts | 8++++++++
Aeditor/e2e/digest.spec.ts | 550+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aeditor/e2e/fixtures/bin/fake-claude.mjs | 86+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aeditor/e2e/fixtures/ollama-stub.mjs | 152+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/e2e/helpers.ts | 155+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/package.json | 4++--
Meditor/playwright.config.ts | 16++++++++++++++++
Aplans/bakeoff/round1.json | 771+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aplans/bakeoff/round1.md | 262+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aplans/bakeoff/round2.json | 562+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aplans/bakeoff/round2.md | 188+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aplans/bakeoff/sample.json | 93+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
28 files changed, 4412 insertions(+), 28 deletions(-)

diff --git a/common/bin/digest-bakeoff.ts b/common/bin/digest-bakeoff.ts @@ -0,0 +1,716 @@ +#!/usr/bin/env tsx +// Score digest engine candidates against a FIXED sample, without writing a +// single digest to disk. +// +// WHY THIS IS A SCRIPT AND NOT INFRASTRUCTURE. It answers one question — which +// (model x context x timestamp mode) should carry a multi-week sweep — and the +// answer is a number in a table, not a feature. It therefore drives digestApps + +// digestPrompt + digestParse DIRECTLY and scores in memory. It NEVER calls +// digestVideo or writeDigestSection, so a losing candidate cannot leave anything +// behind in the corpus, and no freshness record has to be invalidated afterwards +// to undo a round. +// +// THROUGHPUT IS A FIRST-CLASS METRIC, not a footnote. Measured: 46 s per +// 8,194-token chunk on qwen2.5:7b, which over the corpus is ~64 days on one +// lane. A 14B at half the context roughly doubles the chunk count and halves the +// token rate — order 250 days. A candidate can therefore be ruled out on +// projected sweep days alone, however good its chapters look, which is why every +// run prints days alongside quality. +// +// Two modes: +// +// --pick scan the stats cache and write the fixed stratified sample +// (plans/bakeoff/sample.json). Run ONCE. A moving sample makes the +// comparison between rounds meaningless. +// (default) run the candidates over that sample and write a JSON + Markdown +// report under plans/bakeoff/. +// +// Examples: +// tsx bin/digest-bakeoff.ts --pick +// tsx bin/digest-bakeoff.ts --label round1 --buckets short,medium \ +// --candidates 'qwen2.5:7b@16384,qwen3:8b@16384,gemma2:9b@16384' +// tsx bin/digest-bakeoff.ts --label round2 --buckets long,verylong \ +// --candidates 'qwen2.5:7b@16384' --modes absolute,chunk-local + +import path from "node:path"; +import { mkdir, writeFile } from "node:fs/promises"; +import { readFile } from "node:fs/promises"; +import { open } from "lmdb"; +import { getPaths } from "../lib/paths"; +import { parseFlags } from "./_parseFlags"; +import type { VideoStat } from "../lib/stats"; +import { getDigestApp } from "../lib/digestApps"; +import type { DigestAppConfig, DigestTimestampMode } from "../lib/digest"; +import { + CHAPTER_SYSTEM_PROMPT, + DIGEST_OVERLAP_CUES, + buildChapterPrompt, + chapterSchema, + maxCuesForContext, + toHms, +} from "../lib/digestPrompt"; +import { parseChapters, type DigestChunkOutput } from "../lib/digestParse"; +import { chunkCuesForContext } from "../lib/transcriptWindow"; +import { transcriptToMarkdown } from "../lib/transcriptToMarkdown"; +import { readNormalizedTranscript } from "../controller/normalizeTranscript"; +import { CUES_JSON_FILENAME } from "../lib/videoStatus"; +import type { Cue } from "../lib/vtt"; + +// --------------------------------------------------------------------------- +// The sample +// --------------------------------------------------------------------------- + +// Duration strata. Chosen to match the corpus shape recorded in FACTS.md rather +// than to be round numbers: the >4 h bucket is 8.2% of videos but 46% of all +// transcript tokens, so a sample that under-represents it measures the cheap +// half of the sweep and misses the half where chunk-seam bugs live. +const BUCKETS = ["short", "medium", "long", "verylong"] as const; +type Bucket = (typeof BUCKETS)[number]; + +const BUCKET_BOUNDS: Record<Bucket, { min: number; max: number; want: number }> = { + short: { min: 5 * 60, max: 30 * 60, want: 3 }, + medium: { min: 45 * 60, max: 90 * 60, want: 2 }, + long: { min: 3 * 3600, max: 4.5 * 3600, want: 2 }, + verylong: { min: 6 * 3600, max: 14 * 3600, want: 1 }, +}; + +type SampleVideo = { + slug: string; + channelSlug: string; + videoId: string; + videoDir: string; + title: string; + bucket: Bucket; + durationSeconds: number; + cueCount: number; +}; + +type Sample = { + version: 1; + pickedAt: string; + // Corpus-wide totals, captured at pick time. The sweep-days projection is + // computed from measured seconds-per-audio-hour times THIS number, so the + // projection and the sample come from one scan and can't drift apart. + corpus: { + videosScanned: number; + videosWithTranscript: number; + audioHours: number; + longTailVideos: number; + longTailAudioHours: number; + }; + videos: SampleVideo[]; +}; + +function bucketFor(seconds: number): Bucket | null { + for (const b of BUCKETS) { + const { min, max } = BUCKET_BOUNDS[b]; + if (seconds >= min && seconds <= max) return b; + } + return null; +} + +// Deterministic pick, so re-running --pick on an unchanged corpus reproduces the +// same sample. No Math.random: a sample that moves between rounds is not a +// sample, it is noise. Videos are ordered by a stable hash of the slug and the +// first N per bucket are taken, spreading the pick across channels instead of +// clustering on whichever channel sorts first. +function stableHash(s: string): number { + let h = 2166136261; + for (let i = 0; i < s.length; i++) { + h ^= s.charCodeAt(i); + h = Math.imul(h, 16777619); + } + return h >>> 0; +} + +async function pickSample(outPath: string): Promise<void> { + const paths = getPaths(); + const root = open({ path: paths.lmdbPath, maxDbs: 12, compression: true }); + const statsByPath = root.openDB< + { metaMs: number; stat: VideoStat }, + [string, string] + >({ name: "statsByPath", encoding: "msgpack" }); + + const byBucket = new Map<Bucket, SampleVideo[]>(); + for (const b of BUCKETS) byBucket.set(b, []); + + let videosScanned = 0; + let videosWithTranscript = 0; + let totalSeconds = 0; + let longTailVideos = 0; + let longTailSeconds = 0; + + for (const { key, value } of statsByPath.getRange()) { + const stat = value.stat; + videosScanned++; + if (!stat.hasTranscript || !(stat.duration > 0)) continue; + videosWithTranscript++; + totalSeconds += stat.duration; + if (stat.duration > 4 * 3600) { + longTailVideos++; + longTailSeconds += stat.duration; + } + const bucket = bucketFor(stat.duration); + if (!bucket) continue; + // A digest needs cues; a transcript flagged present but empty is useless + // here and would silently shrink a stratum. + if (!stat.cueCount || stat.cueCount < 30) continue; + byBucket.get(bucket)!.push({ + slug: stat.slug, + channelSlug: stat.channelSlug, + videoId: stat.id, + videoDir: (key as [string, string])[1], + title: stat.title, + bucket, + durationSeconds: Math.round(stat.duration), + cueCount: stat.cueCount, + }); + } + await root.close(); + + const videos: SampleVideo[] = []; + for (const b of BUCKETS) { + const pool = byBucket.get(b)!; + pool.sort((a, c) => stableHash(a.slug) - stableHash(c.slug)); + // One per channel first, so a stratum can't come entirely from one + // creator's house style — a model that happens to suit one show would + // otherwise look like a model that suits the corpus. + const seenChannels = new Set<string>(); + const spread: SampleVideo[] = []; + for (const v of pool) { + if (seenChannels.has(v.channelSlug)) continue; + seenChannels.add(v.channelSlug); + spread.push(v); + } + const want = BUCKET_BOUNDS[b].want; + const taken = (spread.length >= want ? spread : pool).slice(0, want); + if (taken.length < want) { + console.warn( + `Warning: bucket ${b} wanted ${want} videos but only ${taken.length} qualify.`, + ); + } + videos.push(...taken); + } + + const sample: Sample = { + version: 1, + pickedAt: new Date().toISOString(), + corpus: { + videosScanned, + videosWithTranscript, + audioHours: Math.round(totalSeconds / 3600), + longTailVideos, + longTailAudioHours: Math.round(longTailSeconds / 3600), + }, + videos, + }; + await mkdir(path.dirname(outPath), { recursive: true }); + await writeFile(outPath, `${JSON.stringify(sample, null, 2)}\n`); + console.log( + `Scanned ${videosScanned} videos (${videosWithTranscript} with transcripts, ` + + `${sample.corpus.audioHours} audio-hours; ${longTailVideos} over 4 h holding ` + + `${sample.corpus.longTailAudioHours} h).`, + ); + for (const v of videos) { + console.log( + ` ${v.bucket.padEnd(8)} ${toHms(v.durationSeconds)} ${v.cueCount + .toString() + .padStart(5)} cues ${v.slug} ${v.title.slice(0, 60)}`, + ); + } + console.log(`Wrote ${outPath}`); +} + +// --------------------------------------------------------------------------- +// Scoring +// --------------------------------------------------------------------------- + +// Titles that carry no information about what was actually said. A cheap proxy +// for title quality: a model that segments correctly but names every section +// "Discussion" has produced a table of contents nobody can navigate. +const GENERIC_TITLE_RE = + /^(the\s+)?(intro(duction)?|outro|conclusion|discussion|continued|continuation|overview|summary|recap|closing( remarks)?|opening( remarks)?|final thoughts|misc(ellaneous)?|other|general|topics?|segment|section|chapter|part)\b/i; +const GENERIC_TITLE_TAIL_RE = /\b(part|section|segment|chapter)\s+(\d+|one|two|three|four|five|six|seven|eight|nine|ten)$/i; + +function isGenericTitle(title: string): boolean { + const t = title.trim(); + return GENERIC_TITLE_RE.test(t) || GENERIC_TITLE_TAIL_RE.test(t); +} + +function normalizeTitle(title: string): string { + return title.toLowerCase().replace(/[^\p{L}\p{N}]+/gu, " ").trim(); +} + +type Candidate = { + key: string; + model: string; + // Absent means "the engine's default" — the same thing an unset + // settings.digest.apps[id].numCtx means, so the default row measures exactly + // what a default-configured sweep would do. + numCtx?: number; + maxCues: number; + timestampMode: DigestTimestampMode; + think?: boolean; +}; + +type VideoScore = { + slug: string; + bucket: Bucket; + durationSeconds: number; + chunks: number; + chunksFailed: number; + zeroYieldChunks: number; + kept: number; + maxGapSeconds: number; + engineSeconds: number; + inputTokens: number; + outputTokens: number; + warningsByCode: Record<string, number>; + genericTitles: number; + duplicateTitles: number; + // Kept so a table can be sanity-checked against real output by hand, which is + // the only way to catch a model that scores well and reads badly. + sampleTitles: string[]; +}; + +type CandidateScore = { + candidate: Candidate; + videos: VideoScore[]; + totals: { + videos: number; + audioHours: number; + chunks: number; + chunksFailed: number; + zeroYieldChunks: number; + zeroYieldRate: number; + kept: number; + chaptersPerHour: number; + maxGapSeconds: number; + meanGapSeconds: number; + genericTitleRate: number; + duplicateTitleRate: number; + warningsByCode: Record<string, number>; + rejectionRate: number; + engineSeconds: number; + tokensPerSecond: number; + secondsPerAudioHour: number; + projectedSweepDays: number; + }; +}; + +function renderChunk( + meta: { id: string; title: string; channel?: string; duration?: number }, + cues: Cue[], + offsetSeconds: number, +): string { + return transcriptToMarkdown( + { ...meta, cues }, + { + timestamps: true, + includeDescription: false, + includeTags: false, + stampForCue: (_clock, seconds) => + toHms(Math.max(0, seconds - offsetSeconds)), + }, + ); +} + +async function scoreVideo( + video: SampleVideo, + candidate: Candidate, + log: (m: string) => void, +): Promise<VideoScore | null> { + const paths = getPaths(); + const cuesPath = path.join( + paths.channelsDir, + video.channelSlug, + "data", + video.videoDir, + CUES_JSON_FILENAME, + ); + const transcript = await readNormalizedTranscript(cuesPath); + if (!transcript || !transcript.cues?.length) { + log(` ${video.slug}: no transcript on disk, skipped`); + return null; + } + const cues = transcript.cues; + const chunks = chunkCuesForContext(cues, { + maxCues: candidate.maxCues, + overlapCues: DIGEST_OVERLAP_CUES, + }); + + const app = getDigestApp("ollama-direct"); + const config: DigestAppConfig = { + model: candidate.model, + numCtx: candidate.numCtx, + temperature: 0, + timeoutMs: 20 * 60_000, + ...(candidate.think !== undefined ? { think: candidate.think } : {}), + }; + + const outputs: DigestChunkOutput[] = []; + const score: VideoScore = { + slug: video.slug, + bucket: video.bucket, + durationSeconds: video.durationSeconds, + chunks: chunks.length, + chunksFailed: 0, + zeroYieldChunks: 0, + kept: 0, + maxGapSeconds: 0, + engineSeconds: 0, + inputTokens: 0, + outputTokens: 0, + warningsByCode: {}, + genericTitles: 0, + duplicateTitles: 0, + sampleTitles: [], + }; + + for (let i = 0; i < chunks.length; i++) { + const chunk = chunks[i]; + const startSeconds = Math.max(0, Math.floor(chunk[0].start)); + const endSeconds = Math.max( + startSeconds, + Math.ceil(chunk[chunk.length - 1].end || chunk[chunk.length - 1].start), + ); + const offset = candidate.timestampMode === "chunk-local" ? startSeconds : 0; + const promptInput = { + title: transcript.title || video.videoId, + channel: transcript.channel || video.channelSlug, + startSeconds, + endSeconds, + transcript: renderChunk( + { + id: transcript.id, + title: transcript.title, + channel: transcript.channel, + duration: transcript.duration, + }, + chunk, + offset, + ), + timestampMode: candidate.timestampMode, + }; + try { + const result = await app.run({ + system: CHAPTER_SYSTEM_PROMPT, + prompt: buildChapterPrompt(promptInput), + schema: chapterSchema(endSeconds - startSeconds), + config, + }); + score.engineSeconds += result.durationMs / 1000; + score.inputTokens += result.inputTokens ?? 0; + score.outputTokens += result.outputTokens ?? 0; + outputs.push({ + index: i, + startSeconds, + endSeconds, + data: result.data, + timestampMode: candidate.timestampMode, + }); + } catch (err) { + score.chunksFailed++; + score.warningsByCode["chunk-failed"] = + (score.warningsByCode["chunk-failed"] ?? 0) + 1; + log(` ${video.slug} chunk ${i + 1}/${chunks.length} failed: ${(err as Error).message}`); + } + } + + // ZERO-YIELD CHUNKS — the headline defect metric. Computed by parsing each + // chunk ALONE, because the merged parse cannot attribute a kept chapter back + // to the chunk that produced it, and "chunk 3 produced nothing" is precisely + // the failure this whole stage exists to fix. A chunk the engine never + // answered counts too: from the corpus's point of view the outcome is the + // same, an interval of the video with no chapters in it. + for (const out of outputs) { + if (parseChapters([out], cues).chapters.length === 0) score.zeroYieldChunks++; + } + score.zeroYieldChunks += score.chunksFailed; + + const parsed = parseChapters(outputs, cues); + score.kept = parsed.chapters.length; + for (const w of parsed.warnings) { + score.warningsByCode[w.code] = (score.warningsByCode[w.code] ?? 0) + 1; + } + + // MAX COVERAGE GAP — catches "summarised the tail, skipped the head". The + // leading gap (0 -> first chapter) and the trailing one (last chapter -> end) + // are included deliberately: a video whose chapters all sit in the last 20 + // minutes has a coverage failure that consecutive-gap-only scoring hides. + const starts = parsed.chapters.map((c) => c.start); + const bounds = [0, ...starts, video.durationSeconds]; + for (let i = 1; i < bounds.length; i++) { + score.maxGapSeconds = Math.max(score.maxGapSeconds, bounds[i] - bounds[i - 1]); + } + + const seen = new Set<string>(); + for (const c of parsed.chapters) { + if (isGenericTitle(c.title)) score.genericTitles++; + const n = normalizeTitle(c.title); + if (seen.has(n)) score.duplicateTitles++; + seen.add(n); + } + score.sampleTitles = parsed.chapters.slice(0, 8).map((c) => `${c.clock} ${c.title}`); + + log( + ` ${video.slug} [${video.bucket}] ${chunks.length} chunk(s) → ${score.kept} chapter(s), ` + + `${score.zeroYieldChunks} zero-yield, ${Math.round(score.engineSeconds)}s engine`, + ); + return score; +} + +function aggregate(candidate: Candidate, videos: VideoScore[]): CandidateScore { + const sum = (f: (v: VideoScore) => number): number => + videos.reduce((a, v) => a + f(v), 0); + const audioHours = sum((v) => v.durationSeconds) / 3600; + const chunks = sum((v) => v.chunks); + const kept = sum((v) => v.kept); + const engineSeconds = sum((v) => v.engineSeconds); + const tokens = sum((v) => v.inputTokens + v.outputTokens); + + const warningsByCode: Record<string, number> = {}; + for (const v of videos) { + for (const [code, n] of Object.entries(v.warningsByCode)) { + warningsByCode[code] = (warningsByCode[code] ?? 0) + n; + } + } + const rejections = Object.entries(warningsByCode) + .filter(([code]) => code !== "seam-duplicate") + .reduce((a, [, n]) => a + n, 0); + + const secondsPerAudioHour = audioHours > 0 ? engineSeconds / audioHours : 0; + + return { + candidate, + videos, + totals: { + videos: videos.length, + audioHours: round(audioHours, 2), + chunks, + chunksFailed: sum((v) => v.chunksFailed), + zeroYieldChunks: sum((v) => v.zeroYieldChunks), + zeroYieldRate: chunks > 0 ? round(sum((v) => v.zeroYieldChunks) / chunks, 4) : 0, + kept, + chaptersPerHour: audioHours > 0 ? round(kept / audioHours, 2) : 0, + maxGapSeconds: videos.reduce((a, v) => Math.max(a, v.maxGapSeconds), 0), + meanGapSeconds: + videos.length > 0 ? Math.round(sum((v) => v.maxGapSeconds) / videos.length) : 0, + genericTitleRate: kept > 0 ? round(sum((v) => v.genericTitles) / kept, 4) : 0, + duplicateTitleRate: kept > 0 ? round(sum((v) => v.duplicateTitles) / kept, 4) : 0, + warningsByCode, + // Rejections per kept chapter — the ratio that says how much of what the + // model produced the guards had to throw away. + rejectionRate: kept + rejections > 0 ? round(rejections / (kept + rejections), 4) : 0, + engineSeconds: Math.round(engineSeconds), + tokensPerSecond: engineSeconds > 0 ? round(tokens / engineSeconds, 1) : 0, + secondsPerAudioHour: Math.round(secondsPerAudioHour), + projectedSweepDays: 0, // filled in once corpus hours are known + }, + }; +} + +function round(n: number, places: number): number { + const f = 10 ** places; + return Math.round(n * f) / f; +} + +// --------------------------------------------------------------------------- +// Report +// --------------------------------------------------------------------------- + +function markdownReport( + label: string, + sample: Sample, + buckets: Bucket[], + scores: CandidateScore[], +): string { + const lines: string[] = []; + lines.push(`# Digest bake-off — ${label}`); + lines.push(""); + lines.push( + `Sample: ${scores[0]?.totals.videos ?? 0} video(s) from \`plans/bakeoff/sample.json\`` + + ` (buckets: ${buckets.join(", ")}), ${scores[0]?.totals.audioHours ?? 0} audio-hours.`, + ); + lines.push( + `Sweep days are projected as measured seconds-per-audio-hour x ` + + `${sample.corpus.audioHours} corpus audio-hours, one lane, no parallelism.`, + ); + lines.push(""); + lines.push( + "| Candidate | Zero-yield chunks | Chapters/h | Rejection rate | Max gap | Generic | Dup | tok/s | s per audio-h | **Sweep days** |", + ); + lines.push( + "| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |", + ); + for (const s of scores) { + const t = s.totals; + lines.push( + `| \`${s.candidate.key}\` | ${t.zeroYieldChunks}/${t.chunks} (${pct(t.zeroYieldRate)}) | ` + + `${t.chaptersPerHour} | ${pct(t.rejectionRate)} | ${toHms(t.maxGapSeconds)} | ` + + `${pct(t.genericTitleRate)} | ${pct(t.duplicateTitleRate)} | ${t.tokensPerSecond} | ` + + `${t.secondsPerAudioHour} | **${t.projectedSweepDays}** |`, + ); + } + lines.push(""); + lines.push("## Rejections by guard"); + lines.push(""); + const codes = Array.from( + new Set(scores.flatMap((s) => Object.keys(s.totals.warningsByCode))), + ).sort(); + lines.push(`| Candidate | ${codes.join(" | ")} |`); + lines.push(`| --- | ${codes.map(() => "---").join(" | ")} |`); + for (const s of scores) { + lines.push( + `| \`${s.candidate.key}\` | ${codes + .map((c) => s.totals.warningsByCode[c] ?? 0) + .join(" | ")} |`, + ); + } + lines.push(""); + lines.push("## Per-video"); + lines.push(""); + lines.push("| Candidate | Video | Bucket | Chunks | Zero-yield | Chapters | Max gap | Engine s |"); + lines.push("| --- | --- | --- | --- | --- | --- | --- | --- |"); + for (const s of scores) { + for (const v of s.videos) { + lines.push( + `| \`${s.candidate.key}\` | \`${v.slug}\` | ${v.bucket} | ${v.chunks} | ` + + `${v.zeroYieldChunks} | ${v.kept} | ${toHms(v.maxGapSeconds)} | ${Math.round(v.engineSeconds)} |`, + ); + } + } + lines.push(""); + lines.push("## Sample output (first chapters per video)"); + lines.push(""); + for (const s of scores) { + lines.push(`### \`${s.candidate.key}\``); + lines.push(""); + for (const v of s.videos) { + lines.push(`**${v.slug}** (${v.bucket}, ${toHms(v.durationSeconds)})`); + lines.push(""); + for (const t of v.sampleTitles) lines.push(`- ${t}`); + if (v.sampleTitles.length === 0) lines.push("- _(nothing survived the guards)_"); + lines.push(""); + } + } + return `${lines.join("\n")}\n`; +} + +function pct(v: number): string { + return `${Math.round(v * 1000) / 10}%`; +} + +// --------------------------------------------------------------------------- +// Main +// --------------------------------------------------------------------------- + +// "model@ctx" or "model@ctx:think" / "model@ctx:nothink". +// +// maxCues is DERIVED from the context rather than configured: the 1200-cue +// default is sized for a 16k window (~10 tokens/cue -> ~12k tokens of transcript +// plus room for prompt and response), so halving the context must halve the +// slice or every call silently truncates — the exact failure that made the first +// smoke test summarize a fragment. +function parseCandidate(spec: string, mode: DigestTimestampMode): Candidate { + const [modelPart, rest] = spec.split("@"); + const [ctxPart, thinkPart] = (rest ?? "").split(":"); + const numCtx = ctxPart ? Number(ctxPart) : undefined; + // The SAME derivation production uses, not a parallel copy — otherwise the + // bake-off scores a chunk size the sweep would never actually run. + const maxCues = maxCuesForContext(numCtx); + return { + key: `${modelPart}@${numCtx ?? "default"}/${mode}`, + model: modelPart, + ...(numCtx ? { numCtx } : {}), + maxCues, + timestampMode: mode, + ...(thinkPart === "think" + ? { think: true } + : thinkPart === "nothink" + ? { think: false } + : {}), + }; +} + +async function main(): Promise<void> { + const flags = parseFlags(process.argv.slice(2)); + const outDir = flags.outDir ?? path.join(process.cwd(), "..", "plans", "bakeoff"); + const samplePath = flags.sample ?? path.join(outDir, "sample.json"); + + if (flags.pick === "true") { + await pickSample(samplePath); + return; + } + + const sample = JSON.parse(await readFile(samplePath, "utf8")) as Sample; + const label = flags.label ?? "round"; + const buckets = (flags.buckets ?? BUCKETS.join(",")) + .split(",") + .map((b) => b.trim()) + .filter((b): b is Bucket => (BUCKETS as readonly string[]).includes(b)); + const modes = (flags.modes ?? "absolute") + .split(",") + .map((m) => m.trim()) + .filter((m): m is DigestTimestampMode => m === "absolute" || m === "chunk-local"); + const specs = (flags.candidates ?? "qwen2.5:7b@16384") + .split(",") + .map((s) => s.trim()) + .filter(Boolean); + + const videos = sample.videos.filter((v) => buckets.includes(v.bucket)); + if (videos.length === 0) throw new Error(`No sample videos in buckets ${buckets.join(",")}`); + + const candidates: Candidate[] = []; + for (const spec of specs) { + for (const mode of modes) candidates.push(parseCandidate(spec, mode)); + } + + console.log( + `Bake-off ${label}: ${candidates.length} candidate(s) x ${videos.length} video(s) ` + + `(${Math.round(videos.reduce((a, v) => a + v.durationSeconds, 0) / 3600)} audio-hours each).`, + ); + + const scores: CandidateScore[] = []; + for (const candidate of candidates) { + console.log(`\n=== ${candidate.key} (maxCues ${candidate.maxCues}) ===`); + const startedAt = Date.now(); + const perVideo: VideoScore[] = []; + for (const video of videos) { + const s = await scoreVideo(video, candidate, (m) => console.log(m)); + if (s) perVideo.push(s); + } + const agg = aggregate(candidate, perVideo); + // Days, from measured seconds-per-audio-hour against the corpus total + // captured in the same scan that picked the sample. + agg.totals.projectedSweepDays = round( + (agg.totals.secondsPerAudioHour * sample.corpus.audioHours) / 86400, + 1, + ); + scores.push(agg); + console.log( + `--- ${candidate.key}: ${agg.totals.kept} chapters, ` + + `${agg.totals.zeroYieldChunks}/${agg.totals.chunks} zero-yield, ` + + `${agg.totals.tokensPerSecond} tok/s, ` + + `projected ${agg.totals.projectedSweepDays} sweep days ` + + `(wall ${Math.round((Date.now() - startedAt) / 60000)} min)`, + ); + } + + scores.sort((a, b) => a.totals.zeroYieldRate - b.totals.zeroYieldRate); + + await mkdir(outDir, { recursive: true }); + const jsonPath = path.join(outDir, `${label}.json`); + const mdPath = path.join(outDir, `${label}.md`); + await writeFile( + jsonPath, + `${JSON.stringify({ label, sample: samplePath, corpus: sample.corpus, buckets, scores }, null, 2)}\n`, + ); + await writeFile(mdPath, markdownReport(label, sample, buckets, scores)); + console.log(`\nWrote ${jsonPath}\nWrote ${mdPath}`); +} + +main().catch((err) => { + console.error(err); + process.exit(1); +}); diff --git a/common/controller/digestBatch.ts b/common/controller/digestBatch.ts @@ -29,11 +29,12 @@ import { type DigestAppConfig, } from "../lib/digestApps"; import { + digestPromptVariant, isSectionFresh, type DigestSectionKind, } from "../lib/digest"; import { loadDigest } from "../lib/digest-server"; -import { PROMPT_VERSION } from "../lib/digestPrompt"; +import { PROMPT_VERSION, maxCuesForContext } from "../lib/digestPrompt"; import { readDigestContext } from "../lib/digestContext-server"; import { CUES_JSON_FILENAME } from "../lib/videoStatus"; import { digestVideo } from "./digestVideo"; @@ -147,6 +148,14 @@ export async function runDigestBatch( const config: DigestAppConfig = digestSettings.apps[app.id] ?? {}; const sections = opts.sections ?? digestSettings.sections; const modelRequested = config.model?.trim() || app.defaultModel(); + const timestampMode = digestSettings.timestampMode; + // Must match what digestVideo will actually chunk with, or the batch's + // freshness check and the writer would disagree on the identity and every + // video would look stale forever. + const promptVariant = digestPromptVariant({ + ...digestSettings, + maxCues: maxCuesForContext(config.numCtx), + }); const dataDir = path.join(opts.paths.channelsDir, opts.channelSlug, "data"); const context = await readDigestContext(opts.paths, opts.channelSlug); @@ -239,6 +248,7 @@ export async function runDigestBatch( model: modelRequested, promptVersion: PROMPT_VERSION, contextHash: context.hash, + ...(promptVariant ? { promptVariant } : {}), }; // Ids already handed out this run. Combined with the cursor below, this is what @@ -314,6 +324,10 @@ export async function runDigestBatch( appId: app.id, config, context, + timestampMode, + ...(digestSettings.promptVariant + ? { promptVariant: digestSettings.promptVariant } + : {}), force: opts.force, onLog: task ? task.onLog : opts.onLog, signal: runSignal, @@ -424,11 +438,16 @@ export async function countMissingDigests( const app = getDigestApp(settings.localAppId); const config = settings.apps[app.id] ?? {}; const context = await readDigestContext(paths, channelSlug); + const promptVariant = digestPromptVariant({ + ...settings, + maxCues: maxCuesForContext(config.numCtx), + }); const target = { appId: app.id, model: config.model?.trim() || app.defaultModel(), promptVersion: PROMPT_VERSION, contextHash: context.hash, + ...(promptVariant ? { promptVariant } : {}), }; const dataDir = path.join(paths.channelsDir, channelSlug, "data"); const dirs = ids diff --git a/common/controller/digestVideo.ts b/common/controller/digestVideo.ts @@ -20,7 +20,6 @@ import path from "node:path"; import type { Paths } from "../lib/paths"; import { - DIGEST_MAX_CUES_PER_CHUNK, DIGEST_OVERLAP_CUES, PROMPT_VERSION, CHAPTER_SYSTEM_PROMPT, @@ -28,13 +27,16 @@ import { buildChapterPrompt, buildTagPrompt, chapterSchema, + maxCuesForContext, tagSchema, toHms, } from "../lib/digestPrompt"; import { getDigestApp, type DigestAppConfig } from "../lib/digestApps"; import { + digestPromptVariant, isSectionFresh, type DigestItem, + type DigestTimestampMode, type DigestProvenance, type DigestRecord, type DigestSectionKind, @@ -60,6 +62,10 @@ export type DigestVideoOptions = { sections?: DigestSectionKind[]; appId?: string; config?: DigestAppConfig; + // Prompt SHAPE. Both default, so an existing caller keeps producing records + // with no promptVariant and every digest already on disk stays fresh. + timestampMode?: DigestTimestampMode; + promptVariant?: string; // Pre-read channel context, so a batch reads it once per channel instead of // once per video. Omitted → read here. context?: DigestContext; @@ -125,12 +131,25 @@ export async function digestVideo( const context = opts.context ?? (await readDigestContext(opts.paths, opts.channelSlug)); + const timestampMode = opts.timestampMode ?? "absolute"; + // Sized to the CONFIGURED context, not to a constant. A 8192-token window with + // 1200-cue chunks overflows and ollama truncates without saying so. + const maxCues = maxCuesForContext(config.numCtx); + // One derivation, shared with the batch and the bake-off, so the same config + // never produces two different identities. + const promptVariant = digestPromptVariant({ + promptVariant: opts.promptVariant, + timestampMode, + maxCues, + }); + const existing = await loadDigest(videoDir); const target = { appId: app.id, model: modelRequested, promptVersion: PROMPT_VERSION, contextHash: context.hash, + ...(promptVariant ? { promptVariant } : {}), }; const stale = sections.filter( (section) => opts.force || !isSectionFresh(existing, section, target), @@ -138,7 +157,7 @@ export async function digestVideo( if (stale.length === 0) return { status: "fresh" }; const chunks = chunkCuesForContext(cues, { - maxCues: DIGEST_MAX_CUES_PER_CHUNK, + maxCues, overlapCues: DIGEST_OVERLAP_CUES, }); log( @@ -170,7 +189,14 @@ export async function digestVideo( channel: transcript.channel || opts.channelSlug, startSeconds, endSeconds, - transcript: renderChunk(transcript, chunk), + // The renderer and the prompt MUST agree on the numbering, so both are + // driven from the same value rather than each deciding for itself. + transcript: renderChunk( + transcript, + chunk, + timestampMode === "chunk-local" ? startSeconds : 0, + ), + timestampMode, ...(context.note ? { contextNote: context.note } : {}), }; const span = endSeconds - startSeconds; @@ -201,6 +227,7 @@ export async function digestVideo( startSeconds, endSeconds, data: result.data, + timestampMode, }); } catch (err) { // A cancel is not a chunk failure — let it propagate so the batch stops. @@ -251,6 +278,7 @@ export async function digestVideo( lane: app.lane, generatedAt: new Date().toISOString(), promptVersion: PROMPT_VERSION, + ...(promptVariant ? { promptVariant } : {}), contextHash: context.hash, chunks: chunks.length, chunksOk: outputs.length, @@ -299,9 +327,15 @@ export async function digestVideo( // // The description is omitted: it is the uploader's own promotional copy and // biases titles toward it. +// +// `offsetSeconds` is subtracted from every marker. It is 0 in absolute mode and +// the chunk's own start in chunk-local mode, which is the whole of what +// "re-basing" means — the cue list itself is never modified, only how it is +// stamped, so the parser's cue snap still works against real cue times. function renderChunk( transcript: { id: string; title: string; channel?: string; duration?: number }, cues: Cue[], + offsetSeconds = 0, ): string { return transcriptToMarkdown( { @@ -315,7 +349,8 @@ function renderChunk( timestamps: true, includeDescription: false, includeTags: false, - stampForCue: (_clock, seconds) => toHms(seconds), + stampForCue: (_clock, seconds) => + toHms(Math.max(0, seconds - offsetSeconds)), }, ); } diff --git a/common/jobs/taskHooks.ts b/common/jobs/taskHooks.ts @@ -80,9 +80,14 @@ export function makeTaskTracker( // split on both \r and \n to see each discrete update. for (const part of line.split(/[\r\n]+/)) { if (!part) continue; + // Both parsers are null for a digest task (see above), so this must + // tolerate having neither. The `!` this replaces was a lie that threw + // on the FIRST log line of every digest job — the batch caught it as a + // per-video failure, so a whole sweep would have reported "0 + // generated, N failed" while looking like an engine problem. const update = transcribeParser ? transcribeParser.feed(part) - : downloadParser!.feed(part); + : downloadParser?.feed(part); if (update) ctx.updateTask(id, update); } }; diff --git a/common/lib/digest.test.ts b/common/lib/digest.test.ts @@ -3,6 +3,7 @@ import assert from "node:assert/strict"; import { DIGEST_SCHEMA_VERSION, chapterId, + digestPromptVariant, effectiveDigest, isSectionFresh, isSharedFrom, @@ -253,3 +254,95 @@ test("isSharedFrom recognizes a digest copied from a cluster's canonical member" assert.equal(isSharedFrom(shared, "Quartering/other"), false); assert.equal(isSharedFrom(machineRecord(), "Quartering/abc"), false); }); + +// --------------------------------------------------------------------------- +// promptVariant — the identity field that lets several prompt shapes coexist +// --------------------------------------------------------------------------- + +test("digestPromptVariant: the default configuration has no variant", () => { + assert.equal(digestPromptVariant({}), undefined); + assert.equal(digestPromptVariant({ timestampMode: "absolute" }), undefined); + assert.equal(digestPromptVariant({ promptVariant: " " }), undefined); +}); + +test("digestPromptVariant: a non-default timestampMode is folded in", () => { + // Folded in rather than left to the operator to remember: a knob that changes + // the output but not the identity would let a re-run under a different mode + // skip every video as "fresh". + assert.equal( + digestPromptVariant({ timestampMode: "chunk-local" }), + "chunk-local", + ); + assert.equal( + digestPromptVariant({ promptVariant: "dense", timestampMode: "chunk-local" }), + "dense+chunk-local", + ); + assert.equal(digestPromptVariant({ promptVariant: "dense" }), "dense"); +}); + +test("isSectionFresh: a record with no promptVariant stays fresh by default", () => { + // The compatibility guarantee. Every digest written before the field existed + // must keep matching, or adding it would have invalidated the whole corpus. + assert.equal(isSectionFresh(machineRecord(), "chapters", TARGET), true); +}); + +test("isSectionFresh: a variant change invalidates the section", () => { + assert.equal( + isSectionFresh(machineRecord(), "chapters", { + ...TARGET, + promptVariant: "chunk-local", + }), + false, + "a chunk-local re-run must not skip absolute-mode records as fresh", + ); +}); + +test("isSectionFresh: a variant record is stale against the default target", () => { + const record = machineRecord({ + sections: { + chapters: { + provenance: { + appId: "ollama-direct", + model: "qwen2.5:7b", + modelRequested: "qwen2.5:7b", + lane: "local-gpu", + generatedAt: "2026-07-26T00:00:00.000Z", + promptVersion: 1, + promptVariant: "chunk-local", + contextHash: "abc123", + }, + items: [], + }, + }, + }); + assert.equal(isSectionFresh(record, "chapters", TARGET), false); + assert.equal( + isSectionFresh(record, "chapters", { + ...TARGET, + promptVariant: "chunk-local", + }), + true, + ); +}); + +test("digestPromptVariant: the default chunk size contributes nothing", () => { + assert.equal(digestPromptVariant({ maxCues: 1200 }), undefined); + assert.equal(digestPromptVariant({ maxCues: 600 }), "c600"); + assert.equal( + digestPromptVariant({ timestampMode: "chunk-local", maxCues: 600 }), + "chunk-local+c600", + ); +}); + +test("isSectionFresh: a chunk-size change invalidates the section", () => { + // Measured: halving the chunk took qwen2.5:7b from 9.34 to 13.37 chapters per + // hour and its worst coverage gap from 1:27:48 to 24:13. A change that large + // must not be able to hide behind an unchanged identity. + assert.equal( + isSectionFresh(machineRecord(), "chapters", { + ...TARGET, + promptVariant: digestPromptVariant({ maxCues: 600 }), + }), + false, + ); +}); diff --git a/common/lib/digest.ts b/common/lib/digest.ts @@ -35,6 +35,80 @@ export const DIGEST_OVERRIDES_VERSION = 1; // digestApps.ts re-exports them so the registry still reads as one unit. export type DigestLane = "local-gpu" | "remote-api"; +// Default ollama context, and the chunk size sized to it. Both live HERE rather +// than in digestApps.ts / digestPrompt.ts so the pure prompt module can size a +// chunk against them without reaching a module that imports execa, and so the +// identity helper below can compare against the default without a cycle. +// +// 16k is comfortable on an 8 GB card (qwen2.5:7b KV cache ~56 KB/token -> +// ~0.9 GB at 16k, atop 4.7 GB of weights). digestPrompt.ts re-exports the cue +// count as DIGEST_MAX_CUES_PER_CHUNK and derives other sizes from it. +export const DEFAULT_DIGEST_NUM_CTX = 16384; +export const DEFAULT_DIGEST_MAX_CUES_PER_CHUNK = 1200; + +// How the transcript markers inside ONE chunk are numbered, and therefore what +// the model is asked to copy. +// +// absolute the chunk's cues carry their real video times, and the prompt +// states the chunk's real range ("01:31:43 to 02:16:09"). +// chunk-local the chunk is re-based to 00:00:00 and the prompt states +// "00:00:00 to 00:44:26". The parser adds the offset back before +// any guard runs, so the range clamp still checks the chunk's +// REAL range and warnings still report real video times. +// +// This exists because of a measured failure, not a hunch. On a 2.3 h video +// (community-notes/v2chrch, 3505 cues → 3 chunks) the third chunk — range +// 01:31:43–02:16:09 — came back with nine starts of 00:00:00, 00:03:54, +// 00:12:26 …: the model had reverted to counting from zero. The per-chunk clamp +// caught all nine, which is exactly its job, but a caught error is still a lost +// chunk, and >4 h videos are 8.2% of the corpus by count and 46% of its tokens. +// chunk-local removes the large offset the model has to hold. Which mode is +// actually better is a BAKE-OFF QUESTION (common/bin/digest-bakeoff.ts), which +// is why both are shipped rather than one being pre-applied as a fix. +export const DIGEST_TIMESTAMP_MODES = ["absolute", "chunk-local"] as const; +export type DigestTimestampMode = (typeof DIGEST_TIMESTAMP_MODES)[number]; +export const DEFAULT_DIGEST_TIMESTAMP_MODE: DigestTimestampMode = "absolute"; + +export function isDigestTimestampMode(v: unknown): v is DigestTimestampMode { + return ( + typeof v === "string" && + (DIGEST_TIMESTAMP_MODES as readonly string[]).includes(v) + ); +} + +// The ONE place the recorded promptVariant string is derived, so every writer +// (the controller, the bake-off harness) produces the same identity for the same +// configuration. +// +// timestampMode is folded in rather than left to the operator to remember: a +// knob that changes the output but not the identity would let a re-run under a +// different mode skip every video as "fresh", which is precisely the failure +// this identity exists to prevent. The default configuration maps to `undefined` +// so pre-existing records stay fresh. +export function digestPromptVariant(input: { + promptVariant?: string; + timestampMode?: DigestTimestampMode; + // Cues per chunk. Folded in for the same reason as timestampMode, and on the + // same evidence: halving it took qwen2.5:7b from 9.34 to 13.37 chapters/hour + // and its worst coverage gap from 1:27:48 to 24:13 on a 3.5 h video. A knob + // that changes the output that much cannot sit outside the identity, or a + // re-run at a new chunk size would skip the whole corpus as "fresh". + maxCues?: number; +}): string | undefined { + const parts: string[] = []; + const named = input.promptVariant?.trim(); + if (named) parts.push(named); + if (input.timestampMode && input.timestampMode !== DEFAULT_DIGEST_TIMESTAMP_MODE) { + parts.push(input.timestampMode); + } + // The default chunk size contributes NOTHING, so every record written before + // this field existed still compares equal. + if (input.maxCues && input.maxCues !== DEFAULT_DIGEST_MAX_CUES_PER_CHUNK) { + parts.push(`c${Math.floor(input.maxCues)}`); + } + return parts.length > 0 ? parts.join("+") : undefined; +} + export const OLLAMA_DIGEST_APP_ID = "ollama-direct"; export const CLAUDE_DIGEST_APP_ID = "claude-code"; export const DEFAULT_DIGEST_APP_ID = OLLAMA_DIGEST_APP_ID; @@ -55,6 +129,15 @@ export type DigestAppConfig = { numCtx?: number; // Sampling temperature. 0 for a structured extraction task. temperature?: number; + // Reasoning-model toggle (ollama's top-level `think`). Only sent when set, so + // a model that does not support thinking is never handed a field it rejects. + // + // It matters for throughput, not correctness: measured on this box, qwen3:8b + // with thinking on spends most of its output budget on a `thinking` block + // before the JSON body the schema constrains. For an extraction task with a + // pinned schema that reasoning buys little and costs a multiple of the tokens, + // and tokens are what a multi-week sweep is priced in. + think?: boolean; // Per-request wall-clock ceiling (ms). A wedged engine must not stall a sweep. timeoutMs?: number; }; @@ -146,6 +229,15 @@ export type DigestProvenance = { generatedAt: string; // Bumped when the prompt/schema changes. Invalidates this section only. promptVersion: number; + // Which prompt SHAPE produced this section, when it was not the default one. + // promptVersion answers "has the prompt changed since?"; this answers "which + // of several concurrently-supported shapes was used?" — the two are different + // questions and a bake-off needs both. Absent means the default shape + // (timestampMode "absolute", no named variant), so every record written before + // this field existed stays valid: isSectionFresh treats absent-on-both as + // equal. Same "one identity field per thing that can change the output" rule + // the rest of this record already follows. + promptVariant?: string; // Hash of the channel-context inputs. Plumbed from the start even while // context files are empty, so adding them later doesn't invalidate the corpus. contextHash: string; @@ -319,9 +411,19 @@ export type DigestFreshnessTarget = { model: string; promptVersion: number; contextHash: string; + // Absent (or empty) means the default prompt shape — see DigestProvenance. + promptVariant?: string; schemaVersion?: number; }; +// Absent and "" are the same thing (the default shape), so that a record written +// before promptVariant existed compares equal to one written today with the +// default settings. Without this normalization, adding the field would have +// invalidated every digest on disk. +function sameVariant(a: string | undefined, b: string | undefined): boolean { + return (a ?? "") === (b ?? ""); +} + export function isSectionFresh( record: DigestRecord | null, section: DigestSectionKind, @@ -340,7 +442,8 @@ export function isSectionFresh( p.appId === target.appId && (p.modelRequested ?? p.model) === target.model && p.promptVersion === target.promptVersion && - p.contextHash === target.contextHash + p.contextHash === target.contextHash && + sameVariant(p.promptVariant, target.promptVariant) ); } diff --git a/common/lib/digestApps.ts b/common/lib/digestApps.ts @@ -22,6 +22,7 @@ import { extractJsonObject, parseStdoutJson } from "./parseStdoutJson"; import { CLAUDE_DIGEST_APP_ID, DEFAULT_DIGEST_APP_ID, + DEFAULT_DIGEST_NUM_CTX, OLLAMA_DIGEST_APP_ID, type DigestAppConfig, type DigestLane, @@ -86,9 +87,11 @@ export type DigestApp = { run: (input: DigestRunInput) => Promise<DigestRunResult>; }; -// Fallbacks shared by both apps. 16k context is comfortable on an 8 GB card -// (qwen2.5:7b KV cache ≈ 56 KB/token → ~0.9 GB at 16k, atop 4.7 GB of weights). -const DEFAULT_NUM_CTX = 16384; +// Fallbacks shared by both apps. The context default lives in digest.ts, so the +// pure prompt module can size a chunk against the same number this resolves to — +// two copies would let the chunker and the engine disagree, and ollama truncates +// the excess SILENTLY. +const DEFAULT_NUM_CTX = DEFAULT_DIGEST_NUM_CTX; const DEFAULT_TEMPERATURE = 0; const DEFAULT_TIMEOUT_MS = 10 * 60_000; @@ -161,6 +164,10 @@ const ollamaDirect: DigestApp = { body: JSON.stringify({ model, stream: false, + // Omitted entirely unless configured: ollama rejects `think` for + // models that have no reasoning mode, so sending a default would + // break every non-reasoning engine. + ...(typeof config.think === "boolean" ? { think: config.think } : {}), // Schema-constrained decoding. This — not prompt wording — is what // made a 7B model emit well-formed timestamps. format: schema, diff --git a/common/lib/digestParse.test.ts b/common/lib/digestParse.test.ts @@ -286,3 +286,137 @@ test("parseTags flags a drifted tag", () => { ); assert.equal(warnings[0].code, "language-drift"); }); + +// --------------------------------------------------------------------------- +// chunk-local timestamps +// +// These pin the EXACT failure that motivated the mode. Digesting +// community-notes/v2chrch (2.3 h, 3505 cues -> 3 chunks), the third chunk — +// range 01:31:43-02:16:09 — came back with nine starts of 00:00:00, 00:03:54, +// 00:12:26 ...: the model had reverted to counting from zero. The per-chunk +// clamp rejected all nine, so the chunk yielded nothing. +// --------------------------------------------------------------------------- + +// The real third chunk's range, to the second. +const CHUNK3_START = 91 * 60 + 43; // 01:31:43 = 5503 +const CHUNK3_END = 2 * 3600 + 16 * 60 + 9; // 02:16:09 = 8169 + +// Cues every 10s across the whole 2.3h video, so the snap has real boundaries. +const LONG_CUES: Cue[] = Array.from({ length: 830 }, (_, i) => cue(i * 10)); + +function chunk3(chapters: unknown[], mode?: "absolute" | "chunk-local"): DigestChunkOutput { + return { + index: 2, + startSeconds: CHUNK3_START, + endSeconds: CHUNK3_END, + data: { chapters }, + ...(mode ? { timestampMode: mode } : {}), + }; +} + +test("chunk-local: a 00:03:00 start in the third chunk resolves to 01:34:43", () => { + const { chapters, warnings } = parseChapters( + [chunk3([{ start: "00:03:00", title: "Filing deadlines" }], "chunk-local")], + LONG_CUES, + ); + assert.equal(chapters.length, 1, "the offset is added back, so it is in range"); + // 5503 + 180 = 5683, snapped to the nearest cue boundary (5680). + assert.equal(chapters[0].start, 5680); + // `clock` is stored in REAL video time, never in the numbering the model used. + assert.equal(chapters[0].clock, "01:34:43"); + assert.equal( + warnings.length, + 0, + "nothing is rejected: this is exactly the output the mode exists to accept", + ); +}); + +test("absolute: the same 00:03:00 is rejected as out-of-range", () => { + const { chapters, warnings } = parseChapters( + [chunk3([{ start: "00:03:00", title: "Filing deadlines" }])], + LONG_CUES, + ); + assert.equal(chapters.length, 0); + const w = warnings.find((x) => x.code === "out-of-range"); + assert.ok(w, "the per-chunk clamp is what caught the measured failure"); + assert.equal(w?.value, "00:03:00"); + assert.match(String(w?.detail), /01:31:43/); +}); + +test("chunk-local: the measured nine-start chunk yields chapters instead of nothing", () => { + // The first three of the nine starts the model actually emitted. + const emitted = ["00:00:00", "00:03:54", "00:12:26"]; + const absolute = parseChapters( + [chunk3(emitted.map((start, i) => ({ start, title: `Topic ${i}` })))], + LONG_CUES, + ); + assert.equal(absolute.chapters.length, 0, "the measured outcome: a lost chunk"); + assert.equal( + absolute.warnings.filter((w) => w.code === "out-of-range").length, + 3, + ); + + const local = parseChapters( + [ + chunk3( + emitted.map((start, i) => ({ start, title: `Topic ${i}` })), + "chunk-local", + ), + ], + LONG_CUES, + ); + assert.equal(local.chapters.length, 3); + assert.deepEqual( + local.chapters.map((c) => c.clock), + ["01:31:43", "01:35:37", "01:44:09"], + ); +}); + +test("chunk-local does not weaken the range clamp", () => { + // 00:50:00 local is 02:21:43 absolute — past the chunk's real end, so it is + // still rejected. The guard checks the chunk's REAL range either way. + const { chapters, warnings } = parseChapters( + [chunk3([{ start: "00:50:00", title: "Past the end" }], "chunk-local")], + LONG_CUES, + ); + assert.equal(chapters.length, 0); + const w = warnings.find((x) => x.code === "out-of-range"); + assert.ok(w); + // The warning reports REAL video time, with the model's own value recorded + // alongside it so a prompt regression stays diagnosable from the artifact. + assert.equal(w?.value, "02:21:43"); + assert.match(String(w?.detail), /model emitted 00:50:00/); +}); + +test("chunk-local: monotonicity is still checked, in real video time", () => { + const { chapters, warnings } = parseChapters( + [ + chunk3( + [ + { start: "00:10:00", title: "Second thing" }, + { start: "00:02:00", title: "Backwards" }, + ], + "chunk-local", + ), + ], + LONG_CUES, + ); + assert.deepEqual( + chapters.map((c) => c.title), + ["Second thing"], + ); + const w = warnings.find((x) => x.code === "non-monotonic"); + assert.ok(w); + assert.equal(w?.value, "01:33:43"); +}); + +test("a chunk starting at 0 is identical under both modes", () => { + const entries = [{ start: "00:01:00", title: "Opening" }]; + const abs = parseChapters([chunk(entries)], CUES); + const local = parseChapters( + [{ ...chunk(entries), timestampMode: "chunk-local" as const }], + CUES, + ); + assert.deepEqual(abs.chapters, local.chapters); + assert.deepEqual(abs.warnings, local.warnings); +}); diff --git a/common/lib/digestParse.ts b/common/lib/digestParse.ts @@ -19,9 +19,15 @@ import { tagId, type DigestChapter, type DigestTag, + type DigestTimestampMode, type DigestWarning, } from "./digest"; -import { HMS_RE, hmsToSeconds, toHms } from "./digestPrompt"; +import { + HMS_RE, + hmsToSeconds, + promptOffsetSeconds, + toHms, +} from "./digestPrompt"; import type { Cue } from "./vtt"; // One engine call's output, with the range that call was RESPONSIBLE for. The @@ -33,6 +39,17 @@ export type DigestChunkOutput = { endSeconds: number; // Whatever the engine returned. Unvalidated by construction. data: unknown; + // Which numbering the engine was PROMPTED in. Must match what the caller + // rendered the chunk with; digestVideo passes the same value to both. Defaults + // to "absolute", so every existing caller is unaffected. + // + // Under "chunk-local" the model counted from 00:00:00 and the parser adds + // startSeconds back BEFORE any guard runs. That ordering is deliberate: the + // range clamp then still checks the chunk's real range (unchanged in + // strength), monotonicity is still compared in real video time, the cue snap + // still lands on a real cue, and every warning still names a real video time + // rather than an offset the reader would have to undo by hand. + timestampMode?: DigestTimestampMode; }; // How far outside its own range a chunk's start may fall before it is rejected. @@ -151,6 +168,9 @@ function parseChapterChunk( return { chapters: [], warnings }; } + // Added back to every emitted stamp before the guards run. Zero unless the + // chunk was prompted in chunk-local numbering. + const offset = promptOffsetSeconds(chunk); const lo = chunk.startSeconds - RANGE_TOLERANCE_SECONDS; const hi = chunk.endSeconds + RANGE_TOLERANCE_SECONDS; const kept: DigestChapter[] = []; @@ -169,14 +189,21 @@ function parseChapterChunk( warn("malformed-timestamp", rawStart || String(entry?.start)); continue; } - const seconds = hmsToSeconds(rawStart); - if (seconds === null) { + const emitted = hmsToSeconds(rawStart); + if (emitted === null) { warn("malformed-timestamp", rawStart); continue; } + // From here on everything is in REAL VIDEO TIME. `clock` follows suit, so a + // stored chapter never carries a stamp in one numbering next to a `start` in + // the other. In absolute mode both are identical to what the model emitted. + const seconds = emitted + offset; + const clock = toHms(seconds); + // Only worth saying when the two differ, i.e. chunk-local. + const emittedNote = clock === rawStart ? "" : ` (model emitted ${rawStart})`; if (!rawTitle) { - warn("empty-title", rawStart); + warn("empty-title", clock); continue; } @@ -186,19 +213,21 @@ function parseChapterChunk( continue; } - // GUARD 3 — per-chunk range clamp. + // GUARD 3 — per-chunk range clamp, always against the chunk's REAL range. + // Under chunk-local this is the guard the offset is designed to let the + // model pass; it is not weakened to do so. if (seconds < lo || seconds > hi) { warn( "out-of-range", - rawStart, - `outside ${toHms(chunk.startSeconds)}–${toHms(chunk.endSeconds)}`, + clock, + `outside ${toHms(chunk.startSeconds)}–${toHms(chunk.endSeconds)}${emittedNote}`, ); continue; } // GUARD 4a — monotonic starts within the chunk. if (seconds <= lastStart) { - warn("non-monotonic", rawStart, `after ${toHms(lastStart)}`); + warn("non-monotonic", clock, `after ${toHms(lastStart)}${emittedNote}`); continue; } lastStart = seconds; @@ -208,7 +237,7 @@ function parseChapterChunk( kept.push({ id: chapterId(snapped), start: snapped, - clock: rawStart, + clock, title: rawTitle, decidedBy: "ai", }); diff --git a/common/lib/digestPrompt.test.ts b/common/lib/digestPrompt.test.ts @@ -0,0 +1,101 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + DIGEST_MAX_CUES_PER_CHUNK, + buildChapterPrompt, + buildTagPrompt, + maxCuesForContext, + promptOffsetSeconds, +} from "./digestPrompt"; +import { DEFAULT_DIGEST_MAX_CUES_PER_CHUNK } from "./digest"; + +// The chunk size MUST track the configured context. Lowering num_ctx without +// lowering it feeds ollama more transcript than its window holds, and ollama +// truncates SILENTLY — the model then summarizes a fragment and the result reads +// as a bad model rather than as a misconfiguration. + +test("the two copies of the default chunk size agree", () => { + // digestPrompt.ts re-exports digest.ts's constant; this pins them equal so the + // identity helper (which compares against digest.ts's copy) and the chunker + // (which uses this one) can never drift apart. + assert.equal(DIGEST_MAX_CUES_PER_CHUNK, DEFAULT_DIGEST_MAX_CUES_PER_CHUNK); +}); + +test("maxCuesForContext scales the slice with the window", () => { + assert.equal(maxCuesForContext(16384), 1200); + assert.equal(maxCuesForContext(8192), 600); + assert.equal(maxCuesForContext(32768), 2400); + // Unset or nonsense falls back to the default window, never to zero cues. + assert.equal(maxCuesForContext(undefined), 1200); + assert.equal(maxCuesForContext(0), 1200); + // Floored, so a tiny window still yields a usable slice rather than one cue. + assert.equal(maxCuesForContext(256), 100); +}); + +test("chunk-local re-bases the range the chapter prompt states", () => { + const base = { + title: "T", + channel: "C", + startSeconds: 5503, // 01:31:43 — the measured third chunk + endSeconds: 8169, // 02:16:09 + transcript: "[00:00:00] hello", + }; + const absolute = buildChapterPrompt(base); + assert.match(absolute, /This section covers 01:31:43 to 02:16:09/); + assert.match(absolute, /between 01:31:43 and 02:16:09 inclusive/); + + const local = buildChapterPrompt({ ...base, timestampMode: "chunk-local" }); + assert.match(local, /This section covers 00:00:00 to 00:44:26/); + assert.match(local, /between 00:00:00 and 00:44:26 inclusive/); + + // The SPAN is unchanged, so the density target and the schema's minItems floor + // are identical in both modes — only the numbering moves. + assert.match(absolute, /44 minutes of material/); + assert.match(local, /44 minutes of material/); +}); + +test("the tag prompt is re-based the same way", () => { + // It sees the same rendered transcript, so a range stated in the other + // numbering would contradict the markers in front of it. + const base = { + title: "T", + channel: "C", + startSeconds: 5503, + endSeconds: 8169, + transcript: "[00:00:00] hello", + }; + assert.match(buildTagPrompt(base), /covering 01:31:43 to 02:16:09/); + assert.match( + buildTagPrompt({ ...base, timestampMode: "chunk-local" }), + /covering 00:00:00 to 00:44:26/, + ); +}); + +test("absolute mode renders byte-identically to the pre-timestampMode prompt", () => { + // The compatibility guarantee that let PROMPT_VERSION stay at 1: adding the + // mode must not change what the default configuration sends, or every digest + // on disk would have been invalidated. + const base = { + title: "T", + channel: "C", + startSeconds: 100, + endSeconds: 700, + transcript: "[00:01:40] hello", + }; + assert.equal( + buildChapterPrompt(base), + buildChapterPrompt({ ...base, timestampMode: "absolute" }), + ); +}); + +test("promptOffsetSeconds is zero in absolute mode by construction", () => { + assert.equal(promptOffsetSeconds({ startSeconds: 5503 }), 0); + assert.equal( + promptOffsetSeconds({ startSeconds: 5503, timestampMode: "absolute" }), + 0, + ); + assert.equal( + promptOffsetSeconds({ startSeconds: 5503, timestampMode: "chunk-local" }), + 5503, + ); +}); diff --git a/common/lib/digestPrompt.ts b/common/lib/digestPrompt.ts @@ -17,14 +17,47 @@ // chunking constants below — a section whose recorded promptVersion differs is // regenerated, and a section whose version matches is skipped. That is what // keeps a re-run minutes long instead of weeks. +import { + DEFAULT_DIGEST_MAX_CUES_PER_CHUNK, + DEFAULT_DIGEST_NUM_CTX, + type DigestTimestampMode, +} from "./digest"; + export const PROMPT_VERSION = 1; +// NOT bumped by the addition of timestampMode below. In "absolute" mode — the +// default — the rendered prompt is byte-identical to what version 1 always +// produced, so every digest on disk stays fresh. The non-default shape is +// distinguished by provenance.promptVariant instead (see digestPromptVariant), +// which is the field that exists precisely so several shapes can be compared +// without a version bump invalidating the corpus between rounds. + // Chunking. 12k usable tokens of transcript per call inside a 16k context leaves // room for the prompt and the response. Expressed in CUES because that is what // the chunker slices; ~10 tokens/cue is the corpus average, so 1200 cues ≈ 12k // tokens. The overlap exists so a topic straddling a seam is visible whole to at // least one call; the parser de-dups the resulting near-identical chapters. -export const DIGEST_MAX_CUES_PER_CHUNK = 1200; +export const DIGEST_MAX_CUES_PER_CHUNK = DEFAULT_DIGEST_MAX_CUES_PER_CHUNK; + +// Cues per chunk, SIZED TO THE CONFIGURED CONTEXT. +// +// The 1200-cue default is sized for a 16k window (~10 tokens/cue -> ~12k tokens +// of transcript, leaving room for the prompt and the response). Lowering +// `numCtx` without lowering this feeds the engine more transcript than its +// window holds, and ollama TRUNCATES SILENTLY — the model then summarizes +// whatever fragment survived and the result looks like a bad model rather than a +// misconfiguration. That failure is measured (see FACTS.md, smoke test 1) and it +// is why this derivation exists rather than a bare constant. +// +// Halving the context is close to free on throughput — measured 24.2 vs 25.1 +// projected sweep days — because twice as many calls each carry half the prompt. +export function maxCuesForContext(numCtx?: number): number { + const ctx = numCtx && numCtx > 0 ? numCtx : DEFAULT_DIGEST_NUM_CTX; + return Math.max( + 100, + Math.round((DIGEST_MAX_CUES_PER_CHUNK * ctx) / DEFAULT_DIGEST_NUM_CTX), + ); +} export const DIGEST_OVERLAP_CUES = 40; // Segmentation density target. 3 chapters for 22 minutes (the measured @@ -82,8 +115,8 @@ export function maxChaptersForSpan(spanSeconds: number): number { export type ChapterPromptInput = { title: string; channel: string; - // The chunk's OWN range, in seconds. Stating it in the prompt is one of the - // three changes that fixed out-of-range output. + // The chunk's OWN range, in seconds, in REAL video time. Stating it in the + // prompt is one of the three changes that fixed out-of-range output. startSeconds: number; endSeconds: number; // The transcript slice, already rendered as `[HH:MM:SS] text` lines by @@ -92,8 +125,25 @@ export type ChapterPromptInput = { // Optional per-channel context note (Phase 1.5). Plumbed from the start so // adding notes later doesn't invalidate the corpus — see contextHash. contextNote?: string; + // Which numbering the caller rendered `transcript` with. Defaults to + // "absolute". Under "chunk-local" the caller has re-based every marker to + // 00:00:00, so the range this prompt states must be re-based to match — a + // prompt that says "01:31:43 to 02:16:09" over markers that start at 00:00:00 + // would contradict itself and is worse than either mode alone. + timestampMode?: DigestTimestampMode; }; +// The offset the caller subtracted from every marker, and that the parser must +// add back. Zero in absolute mode by construction. +export function promptOffsetSeconds(input: { + startSeconds: number; + timestampMode?: DigestTimestampMode; +}): number { + return input.timestampMode === "chunk-local" + ? Math.max(0, Math.floor(input.startSeconds)) + : 0; +} + export function chapterSchema(spanSeconds: number): Record<string, unknown> { return { type: "object", @@ -126,8 +176,9 @@ export const CHAPTER_SYSTEM_PROMPT = [ ].join(" "); export function buildChapterPrompt(input: ChapterPromptInput): string { - const from = toHms(input.startSeconds); - const to = toHms(input.endSeconds); + const offset = promptOffsetSeconds(input); + const from = toHms(input.startSeconds - offset); + const to = toHms(input.endSeconds - offset); const span = Math.max(0, input.endSeconds - input.startSeconds); const minItems = minChaptersForSpan(span); const lines: string[] = []; @@ -199,10 +250,14 @@ export const TAG_SYSTEM_PROMPT = [ ].join(" "); export function buildTagPrompt(input: ChapterPromptInput): string { + // Re-based the same way as the chapter prompt: the tag prompt sees the SAME + // rendered transcript, so a range stated in the other numbering would + // contradict the markers in front of it. + const offset = promptOffsetSeconds(input); const lines: string[] = []; lines.push( `Below is one section of the transcript of "${input.title}" (${input.channel}),`, - `covering ${toHms(input.startSeconds)} to ${toHms(input.endSeconds)}.`, + `covering ${toHms(input.startSeconds - offset)} to ${toHms(input.endSeconds - offset)}.`, ); lines.push(""); lines.push( diff --git a/common/lib/settings.ts b/common/lib/settings.ts @@ -34,10 +34,14 @@ import { import { CLAUDE_DIGEST_APP_ID, DEFAULT_DIGEST_APP_ID, + DEFAULT_DIGEST_TIMESTAMP_MODE, DIGEST_SECTION_KINDS, + DIGEST_TIMESTAMP_MODES, isDigestSectionKind, + isDigestTimestampMode, type DigestAppConfig, type DigestSectionKind, + type DigestTimestampMode, } from "./digest"; export type { Worker } from "./workers"; @@ -218,6 +222,15 @@ export type DigestSettings = { // Which sections a sweep generates. Chapters alone is the default: tags double // the call count for a smaller payoff. sections: DigestSectionKind[]; + // How each chunk's transcript markers are numbered — see DigestTimestampMode. + // A scored variable in the bake-off, not a pre-applied fix, so its effect on + // long videos is measured against the alternatives rather than assumed. + timestampMode: DigestTimestampMode; + // A free-text label for a non-default prompt shape, folded into the recorded + // provenance by digestPromptVariant(). Setting it invalidates every digest + // generated under a different label, which is exactly what makes a bake-off + // round re-run its sample instead of skipping it as fresh. Empty = default. + promptVariant: string; }; // "basic" — `pnpm run build` in export/, serialized on the build queue (shared @@ -487,6 +500,8 @@ export function defaultDigest(): DigestSettings { digestsPaused: false, spendCapUsd: 0, sections: ["chapters"], + timestampMode: DEFAULT_DIGEST_TIMESTAMP_MODE, + promptVariant: "", }; } @@ -517,6 +532,8 @@ export function sanitizeDigestApps( if (typeof r.timeoutMs === "number" && r.timeoutMs > 0) { cfg.timeoutMs = Math.floor(r.timeoutMs); } + // Only carried when explicitly set — see DigestAppConfig.think. + if (typeof r.think === "boolean") cfg.think = r.think; out[id] = cfg; } return out; @@ -555,11 +572,21 @@ export function sanitizeDigest(value: unknown): DigestSettings { // An empty/garbage list would silently generate nothing, so fall back to the // default rather than honoring it. sections: sections.length > 0 ? sections : d.sections, + timestampMode: isDigestTimestampMode(r.timestampMode) + ? r.timestampMode + : d.timestampMode, + // Trimmed and length-capped: it goes into provenance on every record, and a + // runaway value would bloat 119k sidecars. + promptVariant: + typeof r.promptVariant === "string" + ? r.promptVariant.trim().slice(0, 40) + : d.promptVariant, }; } // Every known section kind, for the settings UI's checkbox list. export const DIGEST_SECTION_OPTIONS = DIGEST_SECTION_KINDS; +export const DIGEST_TIMESTAMP_MODE_OPTIONS = DIGEST_TIMESTAMP_MODES; export const BUILD_MAX_PARALLEL_DEFAULT = 2; export const BUILD_MAX_PARALLEL_MAX = 16; diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,7 @@ # Changelog ## [Unreleased] +- **Every transcript can now be given AI-generated chapters and topic tags, from a model running on your own hardware.** A new **Digest** stage on each channel page sweeps its transcripts through a local model (ollama by default) and writes an `ai-digest.json` next to each one — a list of titled moments with timestamps, plus a short set of topic tags. Nothing leaves the machine unless you opt in: a second, **metered** lane (Claude Code) exists for the long tail of very long videos and is **off by default**, gated behind its own spend cap. The two lanes run on separate job queues on purpose — the local lane is GPU-bound and the metered one is network-bound, so sharing a queue would have halved the throughput of a sweep measured in weeks. **A re-run is cheap by construction.** Each generated section records the exact identity that produced it — engine, requested model, prompt version, prompt shape, and a hash of the channel-context inputs — and regeneration skips any section whose identity already matches. That is what makes "digest this channel again" take minutes instead of restarting a multi-week job, and it is why an alias like `qwen2.5` resolving to `qwen2.5:7b` is deliberately *not* treated as a model change. **Hand corrections are kept in a separate file** (`ai-digest.overrides.json`) that generation never opens, so no merge bug in the generator can destroy work a human did; readers compose the two, an override replaces a generated item by id, and `enabled: false` suppresses one without deleting it, so a regeneration can't resurrect something you rejected. **The output is guarded, not trusted.** Schema-constrained decoding pins every timestamp to a full `HH:MM:SS` and demands English titles, and the parser re-checks each item against the chunk's real time range, monotonicity, seam duplication and empty titles — every rejection recorded as a `warning` on the artifact rather than silently dropped, because a sweep this long is only tunable if its failures are inspectable. Chunk size is **sized to the configured context window**: ollama's default 4096 silently truncates over-long input and the model then summarizes whatever fragment survived, which looks like a bad model and is actually a misconfiguration. Two timestamp numberings are shipped and both are selectable (`absolute`, and `chunk-local`, which re-bases each chunk to `00:00:00` and adds the offset back before any guard runs) because which one is better was a measured question, not a guess — a bake-off harness (`common/bin/digest-bakeoff.ts`, reports under `plans/bakeoff/`) exists to settle it, and the choice is folded into the recorded identity so switching modes correctly invalidates the corpus instead of silently skipping it. Digests are also **shared across duplicate clusters**: a confirmed mirror of an already-digested video borrows its digest rather than paying for it twice, but only when detection *measured* the two as aligned, and the borrowed copy records where it came from and by what offset. Also fixes two wiring bugs found only end-to-end: the job log parser threw on the first line of every digest job (both of its parsers are null for this task, and the code asserted one was not — a whole sweep would have reported "0 generated, N failed" and looked like an engine fault), and saving Settings rebuilt the digest block without carrying `timestampMode`/`promptVariant` through, which would have silently reset the prompt shape on an unrelated save and invalidated every digest generated under it. See `common/lib/{digest,digestPrompt,digestParse,digestApps}.ts`, `common/controller/{digestVideo,digestBatch,digestSharing}.ts`, `editor/app/channels/[slug]/{digestActions.ts,components/stages/DigestStage.tsx}`, and `editor/e2e/digest.spec.ts`. - **Replace YouTube's auto-captions with transcripts of our own.** Most of the corpus rides on YouTube ASR captions, which are noticeably worse than what the transcription workers produce — no punctuation, rolling duplicate cues, `[Music]` filler — and they were *sticky*: `isVideoTranscribed()` counts any English VTT, so a video with only auto-captions was permanently invisible to every transcribe bucket and every transcribe job. There is now an opt-in, strictly-lowest-priority lane that finds those videos, downloads their audio, and transcribes them properly; whisper's `transcript.json` then wins the index pick automatically. Provenance is decided by a 4 KB sniff of the VTT itself (YouTube ASR marks ~96–100% of cues with `align:start position:N%` plus inline word timings; manual tracks mark 0%), with the 490 KB `metadata.info.json` parse kept only as a tie-breaker — so the per-regen cost is one small read per English-VTT-having video. Three snapshot buckets carry it: `autoSubsOnly` (needs audio) → `downloadedAutoSubsOnly` (needs whisper) → `supersededAutoSubs` (done, old VTT kept as a backup). Nothing is automatic by default: the auto-queue gains a per-runner **Replace YouTube auto-captions** switch that appends the bucket to the *tail* of the default union (real work always drains first), and a leaf can target the bucket by name for per-channel opt-in — the default unions are byte-identical to before, so existing setups are untouched. Manually, the Transcribe stage gains a two-step "YouTube auto-captions only" section and a single-video *Replace auto-captions* action, and each subtitle track is now labelled *YouTube auto-captions* / *manual captions*. A caption track whose provenance can't be proven machine-generated is **never** a candidate, and the original VTT is never deleted automatically — the Cleanup stage's purge button is the only thing that removes it (English ASR tracks only, honouring `do-not-clean`), so an AI-vs-YouTube comparison stays possible. See `common/lib/subtitleProvenance.ts`, `common/controller/purgeSupersededAutoSubs.ts`, `common/jobs/autoQueuePolicy.ts`, `editor/e2e/auto-subs-replace.spec.ts`. - **Social posts are archived as a parallel corpus to video transcripts.** The archive can now ingest X/Twitter and Bluesky accounts from the same commentators and search them *together* with video transcripts — one corpus, one set of searches. A channel gains `sourceKind: "social"` (a separate axis from `handling`, so every existing `handling === "transcribe" ? … : …` branch stays binary and can never misroute), plus `postFetcher` and `socialHandle`. Posts are modelled on the live-chat layer, not the video layer: their own month-sharded JSONL on disk (`channels/<slug>/posts/YYYY-MM.jsonl` + a `posts-archive` of seen ids, so a re-run is a no-op), their own LMDB sub-DB keyed `[createdAt, channelSlug, id]` (ISO-8601 sorts chronologically, fixing the intra-day ordering the `YYYYMMDD` video key has), and their own `/posts/<slug>/{manifest,page-NNNN}.json` page tree. Ingest is a pluggable `SocialFetcher` registry (`common/social/fetchers.ts`) modelled on the transcription-app registry: **`bluesky-atproto`** (pure `fetch` against the public AT Protocol — no auth, no binary, verified end-to-end against a live account), **`x-gallery-dl`** (the primary X path: a light headless subprocess, cookies via the existing `cookiePolicy.ts`), and **`x-playwright`** (the fallback, immune to the GraphQL query-id rotations that periodically break gallery-dl). One `fetch-posts` job kind carries it, drainable and bookmarkable, routed to `platform:x.com` / `platform:bsky.app` by the existing queue keys. A social channel gets a minimal two-stage rail (Fetch → Index) instead of the six video stages, and the channel form hides every video-only control (audio format, download format, keep-source-video, extraction mode, saved-video dir). Auto-sync works unchanged — but the scheduler's *dispatch* now routes social channels to a post fetch rather than a yt-dlp video sync. See `common/lib/posts.ts`, `common/social/*`, `common/controller/fetchPosts.ts`, `editor/e2e/social-channel.spec.ts`. - **Connect an X account once, instead of re-supplying cookies every few days.** X session cookies expire within days, which is gallery-dl's worst flaw as an archiving path. Settings gains an X-session broker: a headed "Connect X account" flow opens a browser on the editor host so the operator logs in by hand (2FA and captcha included — the login is deliberately never automated), storing a persistent browser profile. The fetcher then re-exports a fresh `cookies.txt` from that profile on demand and prefers it over `--cookies-from-browser`, so the session stops being the thing that breaks. Only X's own cookies are exported. See `common/social/xSessionBroker.ts`, `editor/app/settings/components/XSessionSection.tsx`. diff --git a/editor/app/channels/[slug]/components/stages/DigestStage.tsx b/editor/app/channels/[slug]/components/stages/DigestStage.tsx @@ -0,0 +1,165 @@ +"use client"; + +// The minimum surface needed to RUN a digest sweep on a channel: a count, a lane +// picker, and a button. Deliberately not a review UI — the per-video review +// panel, the /actionable cluster section and the settings form are a separate +// piece of work. What this exists for is that the sweep has to be startable from +// a browser at all, both for an operator and for the e2e suite that proves the +// job path end to end. +// +// Modelled on TranscribeStage: same StreamActionLog + QueueControl shape, same +// bucket-count heading, so the stage rail reads as one system. + +import { useState } from "react"; +import { StreamActionLog } from "yt-dlp-transcript-common/components/StreamActionLog"; +import { QueueControl } from "../../../../components/QueueControl"; +import { cancelJobAction } from "../../../../jobs/actions"; +import { + digestChannelAction, + type DigestLaneChoice, +} from "../../digestActions"; +import { VideoIdList } from "../VideoIdList"; + +type Props = { + slug: string; + existingQueues: string[]; + // Videos with a transcript but no digest at the current identity. + noDigestIds: string[]; + // The two lanes' default queue keys. They are DIFFERENT on purpose: the local + // lane is GPU-bound and the metered lane is network-bound, and the registry + // hardcodes concurrency 1 per key, so a shared key would serialize them. + localQueueKey: string; + remoteQueueKey: string; + // settings.digest.remoteEnabled. False disables the metered option outright + // rather than letting the action reject it after the fact. + remoteEnabled: boolean; + // Longest-first is worth offering because the >4h tail is 8.2% of videos but + // 46% of the tokens — the half where a prompt problem is most expensive to + // discover late. + defaultOrder?: DigestOrder; +}; + +type DigestOrder = "shortest-first" | "longest-first"; + +export function DigestStage({ + slug, + existingQueues, + noDigestIds, + localQueueKey, + remoteQueueKey, + remoteEnabled, + defaultOrder = "shortest-first", +}: Props) { + const [lane, setLane] = useState<DigestLaneChoice>("local"); + const [queue, setQueue] = useState(localQueueKey); + const [order, setOrder] = useState<DigestOrder>(defaultOrder); + const [limit, setLimit] = useState(""); + + const onLaneChange = (next: DigestLaneChoice): void => { + setLane(next); + // Follow the lane's own default key unless the operator has typed a custom + // one — picking "metered" and silently keeping the local key would put both + // lanes on one key and serialize them, which is the exact failure the two + // keys exist to avoid. + setQueue((current) => + current === localQueueKey || current === remoteQueueKey + ? next === "remote" + ? remoteQueueKey + : localQueueKey + : current, + ); + }; + + const parsedLimit = Number.parseInt(limit, 10); + const limitCount = + Number.isFinite(parsedLimit) && parsedLimit > 0 ? parsedLimit : undefined; + + return ( + <div + aria-label="digest section" + className="flex flex-col gap-3" + > + <div> + <h3 className="text-base font-semibold"> + Generate digests ({noDigestIds.length}) + </h3> + <p className="text-sm text-muted-foreground"> + Chapters (and optionally topic tags) derived from each transcript by a + local model, written to <code>ai-digest.json</code> next to it. A + re-run regenerates only what the current model and prompt version have + not already produced, so running this twice costs nothing the second + time. Hand corrections live in a separate{" "} + <code>ai-digest.overrides.json</code> and are never overwritten. + </p> + </div> + + <VideoIdList + slug={slug} + ids={noDigestIds} + ariaLabel="videos without a digest list" + emptyAriaLabel="videos without a digest empty" + emptyMessage="Every transcript has a current digest." + itemAriaLabel={(id) => `video without a digest ${id}`} + /> + + <StreamActionLog + trigger={() => + digestChannelAction(slug, lane, queue, order, limitCount) + } + cancelAction={cancelJobAction} + buttonLabel="Digest channel" + runningLabel="Digesting…" + label="Digest channel" + extraControls={ + <> + <label className="flex items-center gap-1 text-xs text-muted-foreground"> + Lane + <select + value={lane} + onChange={(e) => onLaneChange(e.target.value as DigestLaneChoice)} + aria-label="lane for Digest channel" + className="rounded border border-border bg-card px-1 py-0.5" + > + <option value="local">local (ollama)</option> + <option value="remote" disabled={!remoteEnabled}> + metered{remoteEnabled ? "" : " — disabled in Settings"} + </option> + </select> + </label> + <label className="flex items-center gap-1 text-xs text-muted-foreground"> + Order + <select + value={order} + onChange={(e) => setOrder(e.target.value as DigestOrder)} + aria-label="order for Digest channel" + className="rounded border border-border bg-card px-1 py-0.5" + > + <option value="shortest-first">shortest first</option> + <option value="longest-first">longest first</option> + </select> + </label> + <label className="flex items-center gap-1 text-xs text-muted-foreground"> + Limit + <input + type="text" + inputMode="numeric" + value={limit} + onChange={(e) => setLimit(e.target.value)} + placeholder="all" + aria-label="limit for Digest channel" + className="w-16 rounded border border-border bg-card px-1 py-0.5" + /> + </label> + <QueueControl + value={queue} + onChange={setQueue} + defaultQueueKey={lane === "remote" ? remoteQueueKey : localQueueKey} + existingQueues={existingQueues} + actionLabel="Digest channel" + /> + </> + } + /> + </div> + ); +} diff --git a/editor/app/channels/[slug]/lib/stageStatus.ts b/editor/app/channels/[slug]/lib/stageStatus.ts @@ -44,6 +44,7 @@ export type StageId = | "download" | "transcode" | "transcribe" + | "digest" | "cleanup" | "diagnostics" | "danger"; @@ -76,6 +77,9 @@ const JOB_KIND_TO_STAGE: Record<string, StageId> = { "transcribe-one": "transcribe", "whisper-video": "transcribe", "whisper-bucket-auto-subs": "transcribe", + "digest-channel-local": "digest", + "digest-channel-remote": "digest", + "digest-share-cluster": "digest", "clean-audio-transcribed": "cleanup", "purge-superseded-auto-subs": "cleanup", "clean-extra-audio-formats": "cleanup", @@ -328,6 +332,35 @@ export function computeStageStatuses( }), }; + // Videos with a transcript but no digest at the current prompt/model identity. + // Counted as pending work rather than merely informational: unlike the + // auto-captions lane, every transcribed video is eventually meant to have one. + const digestRunning = runningByStage.has("digest"); + const digestPending = buckets.noDigest.length; + const digest: StageStatus = { + id: "digest", + title: "Digest", + pending: digestPending, + failed: 0, + running: digestRunning, + defaultOpen: true, + summary: digestRunning + ? "Running…" + : digestPending > 0 + ? pluralize( + digestPending, + "transcript without a digest", + "transcripts without a digest", + ) + : "All digested.", + tone: pickTone({ + running: digestRunning, + pending: digestPending, + failed: 0, + fallback: "ok", + }), + }; + const cleanupRunning = runningByStage.has("cleanup"); const cleanupParts: string[] = []; if (cleanupPending > 0) { @@ -408,6 +441,7 @@ export function computeStageStatuses( download, transcode, transcribe, + digest, cleanup, diagnostics, danger, diff --git a/editor/app/channels/[slug]/page.tsx b/editor/app/channels/[slug]/page.tsx @@ -45,6 +45,10 @@ import { queueKeyForUrl, TRANSCRIPTION_QUEUE, } from "yt-dlp-transcript-common/lib/platform"; +import { + DIGEST_LOCAL_QUEUE, + DIGEST_REMOTE_QUEUE, +} from "yt-dlp-transcript-common/lib/queueKeys"; import { getRegistry } from "yt-dlp-transcript-common/jobs/registry"; import { listSites } from "yt-dlp-transcript-common/lib/site"; import { sortGroups } from "yt-dlp-transcript-common/lib/channelGroups"; @@ -64,6 +68,7 @@ import { DownloadStage } from "./components/stages/DownloadStage"; import { PlaylistStage } from "./components/stages/PlaylistStage"; import { TranscodeStage } from "./components/stages/TranscodeStage"; import { TranscribeStage } from "./components/stages/TranscribeStage"; +import { DigestStage } from "./components/stages/DigestStage"; import { VideoListPane } from "./components/VideoListPane"; import { VideoListPaneSection } from "./components/VideoListPaneSection"; import { VideoPanel, type VideoFile } from "./videos/[id]/components/VideoPanel"; @@ -223,7 +228,8 @@ export default async function ChannelDetailPage({ // Saved-video store summary for this channel + whether backups are configured, // for the Cleanup stage's Retention & persistence section (Phase 5). const savedTotals = await savedVideoTotals({ paths, channelSlug: slug }); - const backupConfigured = getSettings().savedVideoBackup.dest.trim() !== ""; + const settings = getSettings(); + const backupConfigured = settings.savedVideoBackup.dest.trim() !== ""; // The retry-failures panel is interactive (a click mutates the file), so // read it fresh on every render — the snapshot bucket only reflects state // at refresh time. @@ -282,6 +288,7 @@ export default async function ChannelDetailPage({ "download", ...(transcodeApplies ? (["transcode"] as const) : []), "transcribe", + "digest", "cleanup", "diagnostics", "danger", @@ -356,6 +363,16 @@ export default async function ChannelDetailPage({ missingShard={transcribeMissingShard} /> ), + digest: ( + <DigestStage + slug={slug} + existingQueues={existingQueues} + noDigestIds={buckets.noDigest} + localQueueKey={DIGEST_LOCAL_QUEUE} + remoteQueueKey={DIGEST_REMOTE_QUEUE} + remoteEnabled={settings.digest.remoteEnabled} + /> + ), cleanup: ( <CleanupStage slug={slug} diff --git a/editor/app/settings/actions.ts b/editor/app/settings/actions.ts @@ -241,6 +241,14 @@ export async function saveSettingsAction( String(formData.get("digestSpendCapUsd") ?? "").trim(), ), sections: digestSectionsRaw.length > 0 ? digestSectionsRaw : dD.sections, + // Prompt SHAPE. Carried through from the current value rather than + // read from the form, because the form has no fields for them yet — + // and a field this block rebuilds but does not carry is a field that + // silently resets on the next save. Both are freshness-affecting + // (see digestPromptVariant), so a silent reset would invalidate every + // digest generated under the non-default shape. + timestampMode: dD.timestampMode, + promptVariant: dD.promptVariant, } : dD ) as SiteSettings["digest"]; diff --git a/editor/e2e/digest.spec.ts b/editor/e2e/digest.spec.ts @@ -0,0 +1,550 @@ +import { test, expect } from "@playwright/test"; +import { writeFile, readFile, rm } from "node:fs/promises"; +import { join } from "node:path"; +// Imported by RELATIVE path, not by package name. The other specs only ever +// import types from `yt-dlp-transcript-common/...`, which are erased at compile +// time; `common` publishes no `exports` map, so a runtime import of the package +// specifier does not resolve under playwright's loader. The relative path does, +// the same way ./helpers does — and pulling in the REAL effectiveDigest matters +// here: asserting override survival against a reimplementation of the merge +// would prove nothing about the merge that ships. +import { + effectiveDigest, + type DigestOverrides, + type DigestRecord, +} from "../../common/lib/digest"; +import { + pathExists, + readJson, + resetData, + resolvePath, + writeChannelConfig, + writeDigestVideo, + writeSettings, +} from "./helpers"; + +// The digest lane, end to end through the real job path. +// +// The local engine is an HTTP stub (e2e/fixtures/ollama-stub.mjs, a third +// playwright webServer) rather than a fake binary, because ollama-direct POSTs +// to ${ollamaUrl}/api/chat and has no binary to replace. The metered lane's +// engine IS a subprocess and so uses the ordinary fake-binary idiom +// (e2e/fixtures/bin/fake-claude.mjs, wired in via CLAUDE_BIN). +// +// Post-mutation assertions poll with reload where they read a rendered page: +// the channel page serves a persisted snapshot regenerated on a ~1s debounce. +// Assertions that read the SIDECARS go straight to disk and need no polling — +// the batch"s closing summary line is the happens-before edge. + +const CHANNEL = "digest-channel"; +const VIDEO = "digestvid0001"; + +function digestPath(channel: string, video: string): string { + return join("test-transcripts", "channels", channel, "data", video, "ai-digest.json"); +} + +function overridesPath(channel: string, video: string): string { + return join( + "test-transcripts", + "channels", + channel, + "data", + video, + "ai-digest.overrides.json", + ); +} + +// Settings with the digest block spelled out. sanitizeDigest fills the rest. +function digestSettings(over: Record<string, unknown> = {}) { + return { + adminTitle: "Test Admin", + maxTranscriptPageBytes: 8388608, + sleepBetweenDownloadsSeconds: 0, + minFreeDiskGB: 0, + digest: { + localAppId: "ollama-direct", + remoteAppId: "claude-code", + sections: ["chapters"], + ...over, + }, + }; +} + +// The batch's closing summary line. Waiting on THAT rather than on a generic +// "succeeded" is what makes these assertions a happens-before edge for the +// sidecar reads below: the line is emitted after the last write. +const BATCH_DONE = "Digest batch:"; + +async function runDigest(page: import("@playwright/test").Page, slug: string) { + await page.goto(`/channels/${slug}`); + await page.getByRole("button", { name: "Digest channel" }).click(); + await expect(page.getByLabel("Digest channel output")).toContainText( + BATCH_DONE, + { timeout: 60_000 }, + ); +} + +// --------------------------------------------------------------------------- +// 1. Override survival — the single most important test in this file. +// +// A full sweep is weeks of wall-clock, so a hand correction that regeneration +// destroys is work that can never be affordably redone. The two sidecars exist +// for exactly this, and nothing else in the suite would notice if the generator +// started writing through to the human file. +// --------------------------------------------------------------------------- + +test("a human override survives a regeneration and reads back as human", async ({ + page, +}) => { + await resetData(null); + await writeSettings(digestSettings()); + await writeChannelConfig(CHANNEL); + await writeDigestVideo({ channelSlug: CHANNEL, videoId: VIDEO }); + + await runDigest(page, CHANNEL); + + const first = await readJson<DigestRecord>(digestPath(CHANNEL, VIDEO)); + const generated = first.sections.chapters?.items ?? []; + expect(generated.length).toBeGreaterThan(0); + expect(generated[0].decidedBy).toBe("ai"); + + // A human retitles the first chapter and suppresses the second, using the + // ids the generator produced. + const overrides: DigestOverrides = { + version: 1, + chapters: [ + { + ...generated[0], + title: "Hand-written chapter title", + decidedBy: "human", + }, + ...(generated[1] + ? [{ ...generated[1], decidedBy: "human" as const, enabled: false }] + : []), + ], + note: "corrected by hand in e2e", + }; + await writeFile( + resolvePath(overridesPath(CHANNEL, VIDEO)), + JSON.stringify(overrides, null, 2), + ); + const overridesBefore = await readFile( + resolvePath(overridesPath(CHANNEL, VIDEO)), + "utf8", + ); + + // Force a genuine regeneration by changing the requested model — that is a + // real identity change, not a test-only escape hatch, so this exercises the + // same path a prompt or model change would take in production. + await writeSettings( + digestSettings({ apps: { "ollama-direct": { model: "qwen2.5:7b-v2" } } }), + ); + await runDigest(page, CHANNEL); + + const second = await readJson<DigestRecord>(digestPath(CHANNEL, VIDEO)); + expect(second.sections.chapters?.provenance.modelRequested).toBe( + "qwen2.5:7b-v2", + ); + + // The human file is untouched, byte for byte. The generator must not be able + // to reach it at all. + const overridesAfter = await readFile( + resolvePath(overridesPath(CHANNEL, VIDEO)), + "utf8", + ); + expect(overridesAfter).toBe(overridesBefore); + + // And the composed view a reader would see honors it: the retitle wins and + // reports decidedBy "human"; the suppressed item is gone without being + // deleted from either file. + const composed = effectiveDigest(second, overrides); + const kept = composed.chapters.find( + (c) => c.title === "Hand-written chapter title", + ); + expect(kept).toBeTruthy(); + expect(kept?.decidedBy).toBe("human"); + expect(composed.hasOverrides).toBe(true); + if (generated[1]) { + expect(composed.chapters.some((c) => c.id === generated[1].id)).toBe(false); + } +}); + +// --------------------------------------------------------------------------- +// 2. Dedup sharing — the correctness crux. +// +// Content similarity says nothing about TIMING. A mirror with a longer intro +// matches on text at shifted times, so sharing a digest onto it would place +// every chapter wrong while the artifact looked perfectly healthy. +// --------------------------------------------------------------------------- + +type ClusterFixture = { + clusterId: string; + members: string[]; + contained?: boolean; + // Pinned explicitly in these fixtures rather than left to pickCanonicalSlug. + // The rule's tie-break for equal-duration, equally-transcribed members is + // LEXICOGRAPHIC on the slug, which silently made the intended mirror the + // canonical and inverted the assertion. Naming it here keeps each test about + // the thing it is testing — the alignment gate — instead of about the + // tie-break. + canonicalSlug?: string; +}; + +async function writeDuplicateReport(clusters: ClusterFixture[]) { + const report = { + version: 1, + generatedAt: new Date().toISOString(), + runConfig: { + thresholdSeconds: null, + durationToleranceSeconds: 2, + nearThreshold: 0.6, + containmentThreshold: 0.8, + shingleSize: 5, + }, + totals: { + videosScanned: clusters.reduce((a, c) => a + c.members.length, 0), + clusters: clusters.length, + videosInClusters: clusters.reduce((a, c) => a + c.members.length, 0), + }, + clusters: clusters.map((c) => ({ + clusterId: c.clusterId, + ...(c.canonicalSlug ? { canonicalSlug: c.canonicalSlug } : {}), + matchKind: "transcript-exact", + score: 1, + contained: c.contained ?? false, + durationBucket: 300, + crossPlatform: true, + crossChannel: true, + videoRefs: c.members.map((slug) => { + const [channelSlug, id] = slug.split("/"); + return { + slug, + channelSlug, + channel: channelSlug, + platform: "youtube", + id, + title: `Mirror of ${id}`, + duration: 600, + uploadDate: "20240101", + hasTranscript: true, + }; + }), + })), + }; + await writeFile( + resolvePath(join("test-transcripts", "duplicates.json")), + JSON.stringify(report, null, 2), + ); +} + +test("a digest is shared to an aligned mirror and refused to a shifted one", async ({ + page, +}) => { + await resetData(null); + await writeSettings(digestSettings()); + await writeChannelConfig(CHANNEL); + // canonical + a byte-aligned mirror + a mirror whose cues are shifted 40s by + // a longer intro. All three carry identical TEXT — only the timings differ, + // which is precisely the case a text-similarity check cannot distinguish. + await writeDigestVideo({ channelSlug: CHANNEL, videoId: "canon00000001" }); + await writeDigestVideo({ channelSlug: CHANNEL, videoId: "aligned000001" }); + await writeDigestVideo({ channelSlug: CHANNEL, videoId: "canon00000002" }); + await writeDigestVideo({ + channelSlug: CHANNEL, + videoId: "shifted000001", + startOffsetSeconds: 40, + }); + + // Each cluster gets its OWN canonical member. Clusters partition the corpus, + // and DigestClusterPlan.bySlug is keyed by slug — so a video listed in two + // clusters would have its role silently overwritten by whichever cluster is + // read last, and only one of the two outcomes would ever be observed. + await writeDuplicateReport([ + { + clusterId: "cluster-aligned", + members: [`${CHANNEL}/canon00000001`, `${CHANNEL}/aligned000001`], + canonicalSlug: `${CHANNEL}/canon00000001`, + }, + { + clusterId: "cluster-shifted", + members: [`${CHANNEL}/canon00000002`, `${CHANNEL}/shifted000001`], + canonicalSlug: `${CHANNEL}/canon00000002`, + }, + ]); + + await runDigest(page, CHANNEL); + + const aligned = await readJson<DigestRecord>( + digestPath(CHANNEL, "aligned000001"), + ); + expect(aligned.derivedFrom?.slug).toBe(`${CHANNEL}/canon00000001`); + // Sharing only happens at near-zero offset, and the record says why it was + // considered safe. + expect(Math.abs(aligned.derivedFrom?.offsetSeconds ?? 99)).toBeLessThanOrEqual(5); + expect(aligned.sections.chapters?.items.length).toBeGreaterThan(0); + + // The shifted mirror was refused the share. It may still have been generated + // for in its own right — what must never happen is it CARRYING the canonical + // member's digest. + const shifted = await pathExists(digestPath(CHANNEL, "shifted000001")); + if (shifted) { + const rec = await readJson<DigestRecord>(digestPath(CHANNEL, "shifted000001")); + expect(rec.derivedFrom).toBeUndefined(); + } +}); + +test("a contained cluster never shares, however well it aligns", async ({ + page, +}) => { + await resetData(null); + await writeSettings(digestSettings()); + await writeChannelConfig(CHANNEL); + await writeDigestVideo({ channelSlug: CHANNEL, videoId: "canon00000001" }); + await writeDigestVideo({ channelSlug: CHANNEL, videoId: "clipped000001" }); + + // Identical timings — this pair would sail through the alignment gate. It is + // refused earlier, because containment means one member is a CLIP: the long + // video's chapters describe material the clip does not contain. + await writeDuplicateReport([ + { + clusterId: "cluster-contained", + members: [`${CHANNEL}/canon00000001`, `${CHANNEL}/clipped000001`], + canonicalSlug: `${CHANNEL}/canon00000001`, + contained: true, + }, + ]); + + await runDigest(page, CHANNEL); + + const clip = await readJson<DigestRecord>(digestPath(CHANNEL, "clipped000001")); + expect(clip.derivedFrom).toBeUndefined(); + // It got its own digest instead of being skipped — a clip is a different + // artifact and deserves one. + expect(clip.sections.chapters?.items.length).toBeGreaterThan(0); +}); + +test("a human canonical override beats the rule", async ({ page }) => { + await resetData(null); + await writeSettings(digestSettings()); + await writeChannelConfig(CHANNEL); + // Equal duration and both transcribed, so the rule falls through to the + // lexicographic tie-break and would pick "aaa…". The human names the other. + await writeDigestVideo({ channelSlug: CHANNEL, videoId: "aaa000000001" }); + await writeDigestVideo({ channelSlug: CHANNEL, videoId: "zzz000000001" }); + + await writeDuplicateReport([ + { + clusterId: "cluster-human", + members: [`${CHANNEL}/aaa000000001`, `${CHANNEL}/zzz000000001`], + }, + ]); + await writeFile( + resolvePath(join("test-transcripts", "duplicates.overrides.json")), + JSON.stringify({ + version: 1, + clusters: { + "cluster-human": { + canonicalSlug: `${CHANNEL}/zzz000000001`, + decidedAt: new Date().toISOString(), + note: "e2e: operator picked the later id", + }, + }, + }), + ); + + await runDigest(page, CHANNEL); + + // The human's choice generated; the other received the share. + const shared = await readJson<DigestRecord>(digestPath(CHANNEL, "aaa000000001")); + expect(shared.derivedFrom?.slug).toBe(`${CHANNEL}/zzz000000001`); + const canonical = await readJson<DigestRecord>( + digestPath(CHANNEL, "zzz000000001"), + ); + expect(canonical.derivedFrom).toBeUndefined(); +}); + +// --------------------------------------------------------------------------- +// 3. The job path +// --------------------------------------------------------------------------- + +test("the Digest stage queues a job and writes the sidecar", async ({ page }) => { + await resetData(null); + await writeSettings(digestSettings()); + await writeChannelConfig(CHANNEL); + await writeDigestVideo({ channelSlug: CHANNEL, videoId: VIDEO }); + + await runDigest(page, CHANNEL); + + await page.goto("/jobs"); + const firstRow = page.getByRole("row").nth(1); + await expect(firstRow).toContainText("digest:local"); + await expect(firstRow).toContainText("Digest channel (local)"); + + const record = await readJson<DigestRecord>(digestPath(CHANNEL, VIDEO)); + expect(record.sections.chapters?.provenance.appId).toBe("ollama-direct"); + expect(record.sections.chapters?.provenance.lane).toBe("local-gpu"); + expect(record.digestSchemaVersion).toBe(1); + // A default-configuration record carries NO promptVariant, which is what + // keeps every digest written before that field existed fresh. + expect(record.sections.chapters?.provenance.promptVariant).toBeUndefined(); +}); + +test("a second run with unchanged versions is a no-op", async ({ page }) => { + await resetData(null); + await writeSettings(digestSettings()); + await writeChannelConfig(CHANNEL); + await writeDigestVideo({ channelSlug: CHANNEL, videoId: VIDEO }); + + await runDigest(page, CHANNEL); + const first = await readJson<DigestRecord>(digestPath(CHANNEL, VIDEO)); + const generatedAt = first.sections.chapters?.provenance.generatedAt; + + await runDigest(page, CHANNEL); + await expect(page.getByLabel("Digest channel output")).toContainText( + "1 already current", + ); + + // The freshness skip is what makes a re-run minutes instead of a second + // multi-week sweep, so "nothing was rewritten" is the assertion. + const second = await readJson<DigestRecord>(digestPath(CHANNEL, VIDEO)); + expect(second.sections.chapters?.provenance.generatedAt).toBe(generatedAt); +}); + +// --------------------------------------------------------------------------- +// 4. Two lanes, two queue keys +// --------------------------------------------------------------------------- + +test("the two lanes land on different queue keys", async ({ page }) => { + await resetData(null); + // longTailSeconds is lowered so the fixture video falls INSIDE the metered + // lane's window. digestChannelAction routes the remote lane with + // minDurationSeconds: settings.digest.longTailSeconds, because the metered + // lane exists for the >4h tail — at the default the 10-minute fixture would be + // out of scope and the run would legitimately find nothing to do. + await writeSettings(digestSettings({ remoteEnabled: true, longTailSeconds: 60 })); + await writeChannelConfig(CHANNEL); + await writeDigestVideo({ channelSlug: CHANNEL, videoId: VIDEO }); + + await page.goto(`/channels/${CHANNEL}`); + await page.getByRole("button", { name: "Digest channel" }).click(); + await expect(page.getByLabel("Digest channel output")).toContainText( + BATCH_DONE, + { timeout: 60_000 }, + ); + + // Switch to the metered lane. The control follows the lane's own default key — + // if it did not, both lanes would share one key and the registry's + // concurrency-1-per-key rule would serialize a GPU lane behind a network one. + await page.goto(`/channels/${CHANNEL}`); + await page.getByLabel("lane for Digest channel").selectOption("remote"); + await expect(page.getByLabel("queue for Digest channel")).toHaveValue( + "digest:remote", + ); + // The metered lane must actually have work to do, so invalidate the identity. + await writeSettings( + digestSettings({ + remoteEnabled: true, + longTailSeconds: 60, + apps: { "claude-code": { model: "haiku" } }, + }), + ); + await page.reload(); + await page.getByLabel("lane for Digest channel").selectOption("remote"); + await page.getByRole("button", { name: "Digest channel" }).click(); + await expect(page.getByLabel("Digest channel output")).toContainText( + BATCH_DONE, + { timeout: 60_000 }, + ); + + await page.goto("/jobs"); + const rows = page.getByRole("row"); + await expect(rows.nth(1)).toContainText("digest:remote"); + await expect(rows.nth(2)).toContainText("digest:local"); + + // And the metered lane's own engine really ran: the record carries its lane + // and the cost the CLI wrapper reported. + const record = await readJson<DigestRecord>(digestPath(CHANNEL, VIDEO)); + expect(record.sections.chapters?.provenance.lane).toBe("remote-api"); + expect(record.sections.chapters?.provenance.appId).toBe("claude-code"); + expect(record.sections.chapters?.provenance.costUsd).toBeGreaterThan(0); +}); + +test("the metered lane is refused while it is disabled in settings", async ({ + page, +}) => { + await resetData(null); + await writeSettings(digestSettings({ remoteEnabled: false })); + await writeChannelConfig(CHANNEL); + await writeDigestVideo({ channelSlug: CHANNEL, videoId: VIDEO }); + + await page.goto(`/channels/${CHANNEL}`); + const lane = page.getByLabel("lane for Digest channel"); + // Nothing can spend money until it is explicitly turned on, so the option is + // disabled at the control rather than rejected after the click. + // + // Asserted via the ATTRIBUTE: playwright's toBeDisabled() reports an <option> + // as enabled regardless of its disabled attribute (it models form-control + // disabled state, and an <option> is not a form control). + await expect(lane.locator("option[value='remote']")).toHaveAttribute( + "disabled", + "", + ); +}); + +// --------------------------------------------------------------------------- +// 5. The guards, through the real pipeline +// +// The unit tests pin each guard in isolation; this pins that the guards are +// actually WIRED — that a bad engine response reaches warnings[] on disk and +// that nothing bogus reaches items[]. +// --------------------------------------------------------------------------- + +test("bad engine output is recorded in warnings and never reaches items", async ({ + page, +}) => { + await resetData(null); + await writeSettings(digestSettings()); + await writeChannelConfig(CHANNEL); + // The stub keys its bad-output mode off this sentinel in the request body — + // the SLOWOP convention the other fakes use. + await writeDigestVideo({ channelSlug: CHANNEL, videoId: "badout000001" }); + + await runDigest(page, CHANNEL); + + const record = await readJson<DigestRecord>( + digestPath(CHANNEL, "badout000001"), + ); + const codes = record.warnings.map((w) => w.code); + expect(codes).toContain("malformed-timestamp"); + expect(codes).toContain("language-drift"); + expect(codes).toContain("out-of-range"); + + // Every rejection names the offending value verbatim, because a multi-week + // sweep is only tunable if its failures are inspectable from the artifact. + const malformed = record.warnings.find( + (w) => w.code === "malformed-timestamp", + ); + expect(malformed?.value).toBe(":00:27"); + expect(malformed?.section).toBe("chapters"); + + const items = record.sections.chapters?.items ?? []; + expect(items.length).toBeGreaterThan(0); + for (const item of items) { + expect(item.clock).toMatch(/^\d\d:\d\d:\d\d$/); + // Nothing that failed a guard survived. + expect(item.title).not.toMatch(/Malformed stamp|Out of range topic/); + expect(item.title).toMatch(/^[\p{L}\p{N}\p{P}\p{Zs}]+$/u); + expect(item.start).toBeLessThanOrEqual(600); + } +}); + +test.afterAll(async () => { + // duplicates.json / duplicates.overrides.json live at the transcripts ROOT, so + // they survive a channel-scoped reset and would leak into unrelated specs. + await rm(resolvePath(join("test-transcripts", "duplicates.json")), { + force: true, + }); + await rm(resolvePath(join("test-transcripts", "duplicates.overrides.json")), { + force: true, + }); +}); diff --git a/editor/e2e/fixtures/bin/fake-claude.mjs b/editor/e2e/fixtures/bin/fake-claude.mjs @@ -0,0 +1,86 @@ +#!/usr/bin/env node +// E2E fake `claude` CLI — the metered digest lane's engine. +// +// Unlike the ollama stub (an HTTP server, because ollama-direct talks HTTP), +// claude-code IS a subprocess, so this follows the existing fake-binary idiom: +// a .mjs in e2e/fixtures/bin/ wired in through an env var (CLAUDE_BIN). +// +// Mimics what digestApps.ts's claudeCode.run() actually invokes and parses: +// claude -p --output-format json [--model <m>] with the prompt on STDIN +// and expects on stdout the CLI's own JSON WRAPPER, whose `result` field is a +// STRING containing the model's JSON body. Reproducing both layers matters — +// the double parse is exactly where a real integration breaks. +// +// There is no constrained decoding over the CLI, so this deliberately emits the +// contract by hand, the same way the real lane depends on the parser's guards +// rather than on a schema. + +import { readFileSync } from "node:fs"; + +const argv = process.argv.slice(2); +function flag(name) { + const i = argv.indexOf(name); + return i < 0 ? undefined : argv[i + 1]; +} + +// The reachability probe digestBatch runs before touching 74k videos — it +// shells out to `claude --version` and treats a non-zero exit as "engine down". +// Answering it is not optional: without this the metered lane fails at the probe +// and never reaches the code the rest of this fake exists to exercise. +if (argv.includes("--version")) { + process.stdout.write("1.0.0 (fake-claude)\n"); + process.exit(0); +} + +// Fail loudly if the invocation drifts: a fake that accepts anything stops +// testing the thing it exists to test. +if (!argv.includes("-p") || flag("--output-format") !== "json") { + process.stderr.write( + `[fake-claude] unexpected argv: ${JSON.stringify(argv)}\n`, + ); + process.exit(2); +} + +let stdin = ""; +try { + stdin = readFileSync(0, "utf8"); +} catch { + stdin = ""; +} + +function hms(total) { + const n = Math.max(0, Math.floor(total)); + return [Math.floor(n / 3600), Math.floor((n % 3600) / 60), n % 60] + .map((v) => String(v).padStart(2, "0")) + .join(":"); +} + +function toSeconds(clock) { + const m = /^(\d\d):(\d\d):(\d\d)$/.exec(clock); + return m ? Number(m[1]) * 3600 + Number(m[2]) * 60 + Number(m[3]) : null; +} + +const m = /This section covers (\d\d:\d\d:\d\d) to (\d\d:\d\d:\d\d)/.exec(stdin); +const start = m ? (toSeconds(m[1]) ?? 0) : 0; +const end = m ? (toSeconds(m[2]) ?? 600) : 600; +const span = Math.max(1, end - start); + +const body = /"tags"/.test(stdin) + ? { tags: ["metered lane", "court filings", "podcast"] } + : { + chapters: [ + { start: hms(start), title: "Metered lane opening remarks" }, + { start: hms(start + Math.floor(span / 3)), title: "Ruling analysis" }, + ], + }; + +process.stdout.write( + JSON.stringify({ + type: "result", + is_error: false, + // The inner body is a STRING, exactly as the real CLI reports it. + result: JSON.stringify(body), + total_cost_usd: 0.0123, + usage: { input_tokens: 1000, output_tokens: 120 }, + }), +); diff --git a/editor/e2e/fixtures/ollama-stub.mjs b/editor/e2e/fixtures/ollama-stub.mjs @@ -0,0 +1,152 @@ +#!/usr/bin/env node +// A stand-in for the ollama HTTP API, for e2e. +// +// WHY THIS IS A SERVER AND NOT A FAKE BINARY. Every other engine in this repo is +// a subprocess, so the e2e idiom is a fake executable in e2e/fixtures/bin/ wired +// in through an env var. `ollama-direct` is different: it POSTs to +// `${ollamaUrl}/api/chat` (common/lib/digestApps.ts), so there is no binary to +// replace. It therefore runs as a third playwright `webServer` and the editor is +// pointed at it with OLLAMA_URL. +// +// It answers from the PROMPT, not from a canned script, so it exercises the real +// contract: the chunk range the model is told to stay inside is parsed back out +// of the prompt text and every emitted start is placed within it. That means the +// stub works unchanged for both timestamp modes — in chunk-local mode the prompt +// states 00:00:00–00:44:26 and the stub answers in that numbering, which is +// exactly what a real model does and what the parser must offset back. +// +// DETERMINISTIC BAD-OUTPUT MODE. When the request carries the BADOUT sentinel +// (the SLOWOP convention from fake-whisper.mjs et al — matched case-insensitively +// anywhere in the request body, so it can be carried by a video id or a title), +// the stub emits one valid chapter plus one of each measured failure: a malformed +// stamp, a non-Latin title, and an out-of-range start. One VALID chapter is +// included on purpose: digestVideo deliberately leaves the sidecar untouched when +// a section yields nothing (so the video retries), so an all-bad response would +// write no file and there would be no warnings[] on disk to assert against. + +import { createServer } from "node:http"; + +const PORT = Number(process.env.OLLAMA_STUB_PORT ?? 11435); +const MODEL = process.env.OLLAMA_STUB_MODEL ?? "qwen2.5:7b"; +const SENTINEL = /badout/i; + +function hms(total) { + const n = Math.max(0, Math.floor(total)); + return [Math.floor(n / 3600), Math.floor((n % 3600) / 60), n % 60] + .map((v) => String(v).padStart(2, "0")) + .join(":"); +} + +function toSeconds(clock) { + const m = /^(\d\d):(\d\d):(\d\d)$/.exec(clock); + if (!m) return null; + return Number(m[1]) * 3600 + Number(m[2]) * 60 + Number(m[3]); +} + +// The prompt states its own range ("This section covers HH:MM:SS to HH:MM:SS"), +// which is the one piece of state the stub needs to answer plausibly. +function rangeFromPrompt(prompt) { + const m = /This section covers (\d\d:\d\d:\d\d) to (\d\d:\d\d:\d\d)/.exec( + prompt ?? "", + ); + if (!m) return { start: 0, end: 600 }; + return { start: toSeconds(m[1]) ?? 0, end: toSeconds(m[2]) ?? 600 }; +} + +// Titles are concrete noun phrases, not "Discussion" — the stub should not be +// the thing that fails a generic-title assertion. +const TITLES = [ + "Court filing deadlines", + "Sponsor read and housekeeping", + "Audience questions on the ruling", + "Closing arguments recap", +]; + +function goodChapters({ start, end }) { + const span = Math.max(1, end - start); + const out = []; + // Three evenly spaced starts, the last comfortably inside the range so a + // rounding difference can never push it past the clamp. + for (let i = 0; i < 3; i++) { + const at = start + Math.floor((span * i) / 4); + out.push({ start: hms(at), title: TITLES[i % TITLES.length] }); + } + return out; +} + +function badChapters(range) { + const [first] = goodChapters(range); + return [ + // Survives every guard, so the section is written and warnings[] lands on + // disk where a spec can read it. + first, + // GUARD 1 — the exact malformed shape the naive prompt produced. + { start: ":00:27", title: "Malformed stamp" }, + // GUARD 2 — language drift. + { start: hms(range.start + 5), title: "法廷の締め切り" }, + // GUARD 3 — an hour past the end of this chunk's range. + { start: hms(range.end + 3600), title: "Out of range topic" }, + ]; +} + +function readBody(req) { + return new Promise((resolve, reject) => { + let raw = ""; + req.on("data", (c) => { + raw += c; + }); + req.on("end", () => resolve(raw)); + req.on("error", reject); + }); +} + +const server = createServer(async (req, res) => { + const url = req.url ?? "/"; + + // The reachability probe digestBatch runs before touching 74k videos. + if (req.method === "GET" && url.startsWith("/api/tags")) { + res.writeHead(200, { "content-type": "application/json" }); + res.end(JSON.stringify({ models: [{ name: MODEL, model: MODEL }] })); + return; + } + + if (req.method === "POST" && url.startsWith("/api/chat")) { + const raw = await readBody(req); + let body = {}; + try { + body = JSON.parse(raw); + } catch { + /* fall through to the default range */ + } + const messages = Array.isArray(body.messages) ? body.messages : []; + const prompt = messages.map((m) => m?.content ?? "").join("\n"); + const range = rangeFromPrompt(prompt); + const bad = SENTINEL.test(raw); + + // Tags ask for a `tags` array; chapters ask for `chapters`. Answer whichever + // the caller's schema names, so the stub covers both sections. + const wantsTags = Boolean(body.format?.properties?.tags); + const data = wantsTags + ? { tags: ["court filings", "podcast", "legal news"] } + : { chapters: bad ? badChapters(range) : goodChapters(range) }; + + res.writeHead(200, { "content-type": "application/json" }); + res.end( + JSON.stringify({ + model: body.model || MODEL, + message: { role: "assistant", content: JSON.stringify(data) }, + done: true, + prompt_eval_count: 100, + eval_count: 50, + }), + ); + return; + } + + res.writeHead(404, { "content-type": "application/json" }); + res.end(JSON.stringify({ error: `no stub route for ${req.method} ${url}` })); +}); + +server.listen(PORT, "127.0.0.1", () => { + console.log(`ollama stub listening on http://127.0.0.1:${PORT}`); +}); diff --git a/editor/e2e/helpers.ts b/editor/e2e/helpers.ts @@ -6,6 +6,7 @@ import { readFile, rm, stat, + utimes, writeFile, } from "node:fs/promises"; import { baseUrl } from "./baseUrl"; @@ -107,3 +108,157 @@ export async function copyFixture(name: string) { await mkdir(dst, { recursive: true }); await cp(testTranscriptsDir, dst, { recursive: true }); } + +// --------------------------------------------------------------------------- +// Digest fixtures +// --------------------------------------------------------------------------- + +// Write a video that the digest lane will actually accept. +// +// digestVideo refuses to run unless transcript.cues.json is FRESH — at least as +// new as both metadata.info.json and the raw transcript (isCuesJsonFresh) — so a +// digest never describes text that is about to be rewritten. Copying a fixture +// tree can't guarantee that ordering, because `cp` stamps every file with the +// time of the copy and the resulting order is whatever the walk produced. So the +// mtimes are set EXPLICITLY here: the sources are backdated and cues.json is +// left at "now". Without this the whole digest suite fails intermittently with +// "stale-cues", which reads like a product bug and isn't one. +export async function writeDigestVideo(opts: { + channelSlug: string; + videoId: string; + title?: string; + // Total video length in seconds. Cues are laid down every 5s across it. + durationSeconds?: number; + channelName?: string; + // Shift every cue by this many seconds while keeping the TEXT identical — a + // mirror with a longer intro. The alignment gate must refuse to share a digest + // onto one of these: the content matches, so a text-only check would pass it, + // and every shared chapter would then land at the wrong moment while the + // artifact looked perfectly healthy. + startOffsetSeconds?: number; +}) { + const { + channelSlug, + videoId, + title = "Synthetic Digest Video", + durationSeconds = 600, + channelName = channelSlug, + startOffsetSeconds = 0, + } = opts; + const dir = join(testTranscriptsDir, "channels", channelSlug, "data", videoId); + await mkdir(dir, { recursive: true }); + + // Every cue's words are GLOBALLY UNIQUE. measureAlignment anchors on 8-word + // n-grams that occur exactly once on each side, so the obvious fixture — the + // same sentence in every cue with only the number changed — yields no unique + // anchors at all and reports "too-few-anchors" instead of the offset result a + // sharing test is actually trying to observe. + const cues = []; + let i = 0; + for (let t = 0; t + 5 <= durationSeconds; t += 5, i++) { + const words = [ + "segment", + "filing", + "deadline", + "schedule", + "ruling", + "hearing", + "motion", + "brief", + "docket", + "counsel", + "exhibit", + "transcript", + ].map((w) => `${w}${i}`); + cues.push({ + start: t + startOffsetSeconds, + end: t + 5 + startOffsetSeconds, + text: words.join(" "), + }); + } + + const meta = { + id: videoId, + title, + channel: channelName, + channel_id: `UC${channelSlug}`, + uploader: channelName, + upload_date: "20240101", + duration: durationSeconds + startOffsetSeconds, + description: "", + is_live: false, + was_live: false, + live_status: "not_live", + age_limit: 0, + extractor_key: "Youtube", + webpage_url: `https://www.youtube.com/watch?v=${videoId}`, + }; + const vtt = [ + "WEBVTT", + "Kind: captions", + "Language: en", + "", + ...cues.flatMap((c) => [ + `${vttStamp(c.start)} --> ${vttStamp(c.end)}`, + c.text, + "", + ]), + ].join("\n"); + + const detail = { + version: 2, + source: "vtt", + transcriptFormat: "vtt", + id: videoId, + slug: `${channelSlug}/${videoId}`, + channelSlug, + channel: channelName, + title, + uploadDate: "20240101", + duration: durationSeconds + startOffsetSeconds, + isLivestream: false, + cues, + }; + + const metaPath = join(dir, "metadata.info.json"); + const vttPath = join(dir, "transcript.en.vtt"); + const cuesPath = join(dir, "transcript.cues.json"); + await writeFile(metaPath, JSON.stringify(meta, null, 2)); + await writeFile(vttPath, vtt); + await writeFile(cuesPath, JSON.stringify(detail)); + + const older = new Date(Date.now() - 60_000); + await utimes(metaPath, older, older); + await utimes(vttPath, older, older); + return { dir, cuesPath, cues }; +} + +function vttStamp(seconds: number): string { + const n = Math.max(0, Math.floor(seconds)); + const h = String(Math.floor(n / 3600)).padStart(2, "0"); + const m = String(Math.floor((n % 3600) / 60)).padStart(2, "0"); + const s = String(n % 60).padStart(2, "0"); + return `${h}:${m}:${s}.000`; +} + +// Minimal channel config so the channel page renders and the batch can run. +export async function writeChannelConfig( + channelSlug: string, + config: Record<string, unknown> = {}, +) { + const dir = join(testTranscriptsDir, "channels", channelSlug); + await mkdir(dir, { recursive: true }); + await writeFile( + join(dir, "config.json"), + JSON.stringify( + { + handling: config.handling ?? "youtube", + name: config.name ?? channelSlug, + url: config.url ?? "https://www.youtube.com/@example/videos", + ...config, + }, + null, + 2, + ), + ); +} diff --git a/editor/package.json b/editor/package.json @@ -5,8 +5,8 @@ "type": "module", "scripts": { "dev": "next dev --port ${EDITOR_PORT:-3001}", - "dev:test": "WORKER_TOKEN=test-worker-token TRANSCRIPTS_DIR=$(pwd)/test-transcripts EXPORT_PUBLIC_DIR=$(pwd)/test-transcripts/.export-public SETTINGS_FILE=$(pwd)/test-settings.json YTDLP_BIN=$(pwd)/e2e/fixtures/bin/fake-ytdlp.mjs GALLERY_DL_BIN=$(pwd)/e2e/fixtures/bin/fake-gallery-dl.mjs WHISPER_BIN=$(pwd)/e2e/fixtures/bin/fake-whisper.mjs WHISPER_MODEL=/dev/null CHOUGH_BIN=$(pwd)/e2e/fixtures/bin/fake-chough.mjs CHOUGH_MODEL=/dev/null PARAKEET_STITCH_BIN=$(pwd)/e2e/fixtures/bin/fake-parakeet-stitch.mjs PARAKEET_CLI=/dev/null PARAKEET_MODEL=/dev/null FFMPEG_BIN=$(pwd)/e2e/fixtures/bin/fake-ffmpeg.mjs FFPROBE_BIN=$(pwd)/e2e/fixtures/bin/fake-ffprobe.mjs AUDIO_CHECK_INTERVAL_MS_OVERRIDE=300 AUDIO_CHECK_SIZE_GATE_OVERRIDE=4096 AUDIO_CHECK_INTERVAL_FLOOR_MS_OVERRIDE=50 AUDIO_CHECK_RECOVER_STEP_MS_OVERRIDE=100 AUDIO_CHECK_RECOVER_AFTER_OVERRIDE=2 next dev --port ${PORT:-3011}", - "start:test": "WORKER_TOKEN=test-worker-token TRANSCRIPTS_DIR=$(pwd)/test-transcripts EXPORT_PUBLIC_DIR=$(pwd)/test-transcripts/.export-public SETTINGS_FILE=$(pwd)/test-settings.json YTDLP_BIN=$(pwd)/e2e/fixtures/bin/fake-ytdlp.mjs GALLERY_DL_BIN=$(pwd)/e2e/fixtures/bin/fake-gallery-dl.mjs WHISPER_BIN=$(pwd)/e2e/fixtures/bin/fake-whisper.mjs WHISPER_MODEL=/dev/null CHOUGH_BIN=$(pwd)/e2e/fixtures/bin/fake-chough.mjs CHOUGH_MODEL=/dev/null PARAKEET_STITCH_BIN=$(pwd)/e2e/fixtures/bin/fake-parakeet-stitch.mjs PARAKEET_CLI=/dev/null PARAKEET_MODEL=/dev/null FFMPEG_BIN=$(pwd)/e2e/fixtures/bin/fake-ffmpeg.mjs FFPROBE_BIN=$(pwd)/e2e/fixtures/bin/fake-ffprobe.mjs AUDIO_CHECK_INTERVAL_MS_OVERRIDE=300 AUDIO_CHECK_SIZE_GATE_OVERRIDE=4096 AUDIO_CHECK_INTERVAL_FLOOR_MS_OVERRIDE=50 AUDIO_CHECK_RECOVER_STEP_MS_OVERRIDE=100 AUDIO_CHECK_RECOVER_AFTER_OVERRIDE=2 next start --port ${PORT:-3011}", + "dev:test": "WORKER_TOKEN=test-worker-token TRANSCRIPTS_DIR=$(pwd)/test-transcripts EXPORT_PUBLIC_DIR=$(pwd)/test-transcripts/.export-public SETTINGS_FILE=$(pwd)/test-settings.json YTDLP_BIN=$(pwd)/e2e/fixtures/bin/fake-ytdlp.mjs GALLERY_DL_BIN=$(pwd)/e2e/fixtures/bin/fake-gallery-dl.mjs WHISPER_BIN=$(pwd)/e2e/fixtures/bin/fake-whisper.mjs WHISPER_MODEL=/dev/null CHOUGH_BIN=$(pwd)/e2e/fixtures/bin/fake-chough.mjs CHOUGH_MODEL=/dev/null PARAKEET_STITCH_BIN=$(pwd)/e2e/fixtures/bin/fake-parakeet-stitch.mjs PARAKEET_CLI=/dev/null PARAKEET_MODEL=/dev/null FFMPEG_BIN=$(pwd)/e2e/fixtures/bin/fake-ffmpeg.mjs FFPROBE_BIN=$(pwd)/e2e/fixtures/bin/fake-ffprobe.mjs OLLAMA_URL=http://127.0.0.1:${OLLAMA_STUB_PORT:-11435} CLAUDE_BIN=$(pwd)/e2e/fixtures/bin/fake-claude.mjs AUDIO_CHECK_INTERVAL_MS_OVERRIDE=300 AUDIO_CHECK_SIZE_GATE_OVERRIDE=4096 AUDIO_CHECK_INTERVAL_FLOOR_MS_OVERRIDE=50 AUDIO_CHECK_RECOVER_STEP_MS_OVERRIDE=100 AUDIO_CHECK_RECOVER_AFTER_OVERRIDE=2 next dev --port ${PORT:-3011}", + "start:test": "WORKER_TOKEN=test-worker-token TRANSCRIPTS_DIR=$(pwd)/test-transcripts EXPORT_PUBLIC_DIR=$(pwd)/test-transcripts/.export-public SETTINGS_FILE=$(pwd)/test-settings.json YTDLP_BIN=$(pwd)/e2e/fixtures/bin/fake-ytdlp.mjs GALLERY_DL_BIN=$(pwd)/e2e/fixtures/bin/fake-gallery-dl.mjs WHISPER_BIN=$(pwd)/e2e/fixtures/bin/fake-whisper.mjs WHISPER_MODEL=/dev/null CHOUGH_BIN=$(pwd)/e2e/fixtures/bin/fake-chough.mjs CHOUGH_MODEL=/dev/null PARAKEET_STITCH_BIN=$(pwd)/e2e/fixtures/bin/fake-parakeet-stitch.mjs PARAKEET_CLI=/dev/null PARAKEET_MODEL=/dev/null FFMPEG_BIN=$(pwd)/e2e/fixtures/bin/fake-ffmpeg.mjs FFPROBE_BIN=$(pwd)/e2e/fixtures/bin/fake-ffprobe.mjs OLLAMA_URL=http://127.0.0.1:${OLLAMA_STUB_PORT:-11435} CLAUDE_BIN=$(pwd)/e2e/fixtures/bin/fake-claude.mjs AUDIO_CHECK_INTERVAL_MS_OVERRIDE=300 AUDIO_CHECK_SIZE_GATE_OVERRIDE=4096 AUDIO_CHECK_INTERVAL_FLOOR_MS_OVERRIDE=50 AUDIO_CHECK_RECOVER_STEP_MS_OVERRIDE=100 AUDIO_CHECK_RECOVER_AFTER_OVERRIDE=2 next start --port ${PORT:-3011}", "build": "next build", "start": "next start --port ${EDITOR_PORT:-3001}", "lint": "eslint", diff --git a/editor/playwright.config.ts b/editor/playwright.config.ts @@ -10,6 +10,13 @@ process.env.PLAYWRIGHT_BASE_URL = baseURL; const webServerCommand = process.env.E2E_MODE === "start" ? "pnpm start:test" : "pnpm dev:test"; +// The digest lane's local engine is reached over HTTP, not spawned, so it gets a +// webServer entry instead of a fake binary in e2e/fixtures/bin/. Its port is +// exported so the editor's dev:test / start:test scripts point OLLAMA_URL at the +// same place, and so an offset-port worktree does not collide. +const OLLAMA_STUB_PORT = Number(process.env.OLLAMA_STUB_PORT ?? 11435); +process.env.OLLAMA_STUB_PORT = String(OLLAMA_STUB_PORT); + const EXPORT_PORT = Number(process.env.EXPORT_PORT ?? 3010); const exportBaseURL = `http://localhost:${EXPORT_PORT}`; // Share the editor's test-settings.json with the export server so the @@ -54,6 +61,15 @@ export default defineConfig({ SITE_ID: "testsite", }, }, + { + command: `node ${path.resolve(process.cwd(), "e2e", "fixtures", "ollama-stub.mjs")}`, + // /api/tags is the same endpoint digestApps.ts probes, so playwright's + // readiness check and the app's own reachability check agree. + url: `http://127.0.0.1:${OLLAMA_STUB_PORT}/api/tags`, + timeout: 30_000, + reuseExistingServer: !process.env.CI, + env: { OLLAMA_STUB_PORT: String(OLLAMA_STUB_PORT) }, + }, ], use: { baseURL, diff --git a/plans/bakeoff/round1.json b/plans/bakeoff/round1.json @@ -0,0 +1,771 @@ +{ + "label": "round1", + "sample": "/home/user/Projects/yt-dlp-transcript-browser/plans/bakeoff/sample.json", + "corpus": { + "videosScanned": 76354, + "videosWithTranscript": 73367, + "audioHours": 77298, + "longTailVideos": 6038, + "longTailAudioHours": 40153 + }, + "buckets": [ + "short", + "medium" + ], + "scores": [ + { + "candidate": { + "key": "gemma2:9b@8192/absolute", + "model": "gemma2:9b", + "numCtx": 8192, + "maxCues": 600, + "timestampMode": "absolute" + }, + "videos": [ + { + "slug": "the-quartering-rumble/v6ve4n0", + "bucket": "short", + "durationSeconds": 791, + "chunks": 1, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 3, + "maxGapSeconds": 360, + "engineSeconds": 78.767, + "inputTokens": 4644, + "outputTokens": 131, + "warningsByCode": {}, + "genericTitles": 0, + "duplicateTitles": 0, + "sampleTitles": [ + "00:00:00 Trump Trolls Colbert While Setting Up a Trap for Democrats", + "00:04:00 Colbert's Cancellation and the Changing Landscape of Late Night Television", + "00:10:00 Trump's DC Move Forces Democrats to Address Crime" + ] + }, + { + "slug": "chibi-reviews/VQykVuHd9xQ", + "bucket": "short", + "durationSeconds": 433, + "chunks": 1, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 2, + "maxGapSeconds": 248, + "engineSeconds": 51.266, + "inputTokens": 3630, + "outputTokens": 81, + "warningsByCode": {}, + "genericTitles": 0, + "duplicateTitles": 0, + "sampleTitles": [ + "00:00:02 Crunchyroll's Translation Issues", + "00:04:10 Impact on Viewership Experience" + ] + }, + { + "slug": "the-quartering/J0ySGwzP4Nw", + "bucket": "short", + "durationSeconds": 785, + "chunks": 1, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 3, + "maxGapSeconds": 331, + "engineSeconds": 104.735, + "inputTokens": 6119, + "outputTokens": 122, + "warningsByCode": {}, + "genericTitles": 0, + "duplicateTitles": 0, + "sampleTitles": [ + "00:00:04 Tim Pool Discusses Censorship on Social Media", + "00:04:01 Twitter's Handling of the 'G Word'", + "00:09:32 Concerns About LGBTQ+ Content in Schools" + ] + }, + { + "slug": "destiny/5nmDzKB23OU", + "bucket": "medium", + "durationSeconds": 4419, + "chunks": 4, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 20, + "maxGapSeconds": 804, + "engineSeconds": 408.8299999999999, + "inputTokens": 20127, + "outputTokens": 724, + "warningsByCode": {}, + "genericTitles": 0, + "duplicateTitles": 0, + "sampleTitles": [ + "00:13:24 The Pure Concept of Religion", + "00:16:53 Assumptions and Definitions", + "00:20:02 Beyond Christianity", + "00:20:41 Reconciling Perspectives", + "00:20:56 The Importance of Openness", + "00:32:58 Methodological Approach to Incorporating Evidence for Moral Facts", + "00:34:02 Determining Which Brain States are 'Moral'", + "00:37:01 Skepticism and the Burden of Proof" + ] + }, + { + "slug": "angryjoeshow/RPJqkewZP5I", + "bucket": "medium", + "durationSeconds": 3380, + "chunks": 3, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 15, + "maxGapSeconds": 899, + "engineSeconds": 233.37699999999998, + "inputTokens": 12297, + "outputTokens": 552, + "warningsByCode": {}, + "genericTitles": 1, + "duplicateTitles": 0, + "sampleTitles": [ + "00:14:59 Game's Defense and Blame Shifting", + "00:15:02 Content Creators and Gamers Blamed", + "00:15:38 One Peg Response - Game's Issues Highlighted", + "00:16:21 Mixed Opinions on the Game", + "00:17:34 Studio's Responsibility and No Man's Sky Comparison", + "00:18:25 Microtransactions and Lack of Appeal", + "00:20:39 Jeff Keley's Role and Personal Taste", + "00:34:25 Twitter and Patch Notes Struggles" + ] + } + ], + "totals": { + "videos": 5, + "audioHours": 2.72, + "chunks": 10, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "zeroYieldRate": 0, + "kept": 43, + "chaptersPerHour": 15.78, + "maxGapSeconds": 899, + "meanGapSeconds": 528, + "genericTitleRate": 0.0233, + "duplicateTitleRate": 0, + "warningsByCode": {}, + "rejectionRate": 0, + "engineSeconds": 877, + "tokensPerSecond": 55.2, + "secondsPerAudioHour": 322, + "projectedSweepDays": 288.1 + } + }, + { + "candidate": { + "key": "qwen3:8b@16384/absolute", + "model": "qwen3:8b", + "numCtx": 16384, + "maxCues": 1200, + "timestampMode": "absolute", + "think": false + }, + "videos": [ + { + "slug": "the-quartering-rumble/v6ve4n0", + "bucket": "short", + "durationSeconds": 791, + "chunks": 1, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 6, + "maxGapSeconds": 246, + "engineSeconds": 62.781, + "inputTokens": 4523, + "outputTokens": 215, + "warningsByCode": {}, + "genericTitles": 0, + "duplicateTitles": 0, + "sampleTitles": [ + "00:00:00 Trump Trolls Colbert by Taking His Prime Time Job", + "00:02:03 Speculation About Secret Side Deals and Late Night Television", + "00:06:09 Paramount's Merger and the Late Night Show Format Crisis", + "00:08:31 Democrats Grapple with Crime and Trump's Strategy", + "00:10:06 Trump's Impact on Crime Perception and Democratic Vulnerabilities", + "00:12:01 Woke DAs and the Changing Crime Statistics" + ] + }, + { + "slug": "chibi-reviews/VQykVuHd9xQ", + "bucket": "short", + "durationSeconds": 433, + "chunks": 1, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 4, + "maxGapSeconds": 230, + "engineSeconds": 34.821, + "inputTokens": 3576, + "outputTokens": 91, + "warningsByCode": {}, + "genericTitles": 0, + "duplicateTitles": 0, + "sampleTitles": [ + "00:00:02 Crunchyroll's Translation Issues", + "00:01:00 Translation Errors in Gundam Episode", + "00:02:33 Unwatchable Episode Due to Subtitle Delays", + "00:03:23 Translation Problems in Overlord Series},{" + ] + }, + { + "slug": "the-quartering/J0ySGwzP4Nw", + "bucket": "short", + "durationSeconds": 785, + "chunks": 1, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 7, + "maxGapSeconds": 168, + "engineSeconds": 76.327, + "inputTokens": 6067, + "outputTokens": 223, + "warningsByCode": {}, + "genericTitles": 0, + "duplicateTitles": 0, + "sampleTitles": [ + "00:00:04 Tim Pool's War With Twitter", + "00:01:13 The G Word and Banning", + "00:04:01 Free Speech and Platform Rules", + "00:06:00 Twitter's Alleged Protection of Creeps", + "00:08:01 Twitter's Uneven Enforcement", + "00:10:00 Concerns About Teachers and Young Ones", + "00:12:05 Criticizing the Platform's Policies" + ] + }, + { + "slug": "destiny/5nmDzKB23OU", + "bucket": "medium", + "durationSeconds": 4419, + "chunks": 2, + "chunksFailed": 0, + "zeroYieldChunks": 1, + "kept": 8, + "maxGapSeconds": 1900, + "engineSeconds": 231.563, + "inputTokens": 16388, + "outputTokens": 527, + "warningsByCode": { + "out-of-range": 12 + }, + "genericTitles": 3, + "duplicateTitles": 0, + "sampleTitles": [ + "00:02:00 Exploring Moral Systems and Preferences", + "00:06:00 Meta Ethics and the Nature of Moral Truth", + "00:12:00 The Role of Neuroscience in Understanding Morality", + "00:18:00 Skepticism and the Burden of Proof", + "00:24:00 Normativity and Its Interpretations", + "00:30:00 Conclusion and Reflection on Moral Frameworks", + "00:36:00 Final Thoughts and Open Questions", + "00:42:00 Closing Remarks and Further Exploration" + ] + }, + { + "slug": "angryjoeshow/RPJqkewZP5I", + "bucket": "medium", + "durationSeconds": 3380, + "chunks": 2, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 17, + "maxGapSeconds": 2397, + "engineSeconds": 218.47199999999998, + "inputTokens": 15792, + "outputTokens": 499, + "warningsByCode": { + "out-of-range": 1 + }, + "genericTitles": 2, + "duplicateTitles": 0, + "sampleTitles": [ + "00:00:28 Silent Hill and New Game", + "00:00:30 John Wick and God of War", + "00:00:40 Marathon and High Guard", + "00:00:50 Crimson Desert and AI", + "00:01:00 State of Play and AI Controversy", + "00:01:10 Conclusion and Final Thoughts", + "00:01:20 Wrap-Up and Farewell", + "00:01:30 End of Episode" + ] + } + ], + "totals": { + "videos": 5, + "audioHours": 2.72, + "chunks": 7, + "chunksFailed": 0, + "zeroYieldChunks": 1, + "zeroYieldRate": 0.1429, + "kept": 42, + "chaptersPerHour": 15.42, + "maxGapSeconds": 2397, + "meanGapSeconds": 988, + "genericTitleRate": 0.119, + "duplicateTitleRate": 0, + "warningsByCode": { + "out-of-range": 13 + }, + "rejectionRate": 0.2364, + "engineSeconds": 624, + "tokensPerSecond": 76.8, + "secondsPerAudioHour": 229, + "projectedSweepDays": 204.9 + } + }, + { + "candidate": { + "key": "mistral-nemo:12b@16384/absolute", + "model": "mistral-nemo:12b", + "numCtx": 16384, + "maxCues": 1200, + "timestampMode": "absolute" + }, + "videos": [ + { + "slug": "the-quartering-rumble/v6ve4n0", + "bucket": "short", + "durationSeconds": 791, + "chunks": 1, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 3, + "maxGapSeconds": 297, + "engineSeconds": 106.735, + "inputTokens": 4552, + "outputTokens": 110, + "warningsByCode": {}, + "genericTitles": 0, + "duplicateTitles": 0, + "sampleTitles": [ + "00:00:00 Trump's Ruthless Troll Against Colbert", + "00:04:09 Colbert's Late Show Cancellation and Trump's Involvement", + "00:08:14 Economic Challenges of Late Night Shows" + ] + }, + { + "slug": "chibi-reviews/VQykVuHd9xQ", + "bucket": "short", + "durationSeconds": 433, + "chunks": 1, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 2, + "maxGapSeconds": 364, + "engineSeconds": 68.864, + "inputTokens": 3590, + "outputTokens": 73, + "warningsByCode": {}, + "genericTitles": 0, + "duplicateTitles": 0, + "sampleTitles": [ + "00:00:02 Crunchyroll's translations issues", + "00:01:09 Gundam episode with subtitle delays" + ] + }, + { + "slug": "the-quartering/J0ySGwzP4Nw", + "bucket": "short", + "durationSeconds": 785, + "chunks": 1, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 3, + "maxGapSeconds": 609, + "engineSeconds": 138.169, + "inputTokens": 6093, + "outputTokens": 107, + "warningsByCode": {}, + "genericTitles": 0, + "duplicateTitles": 0, + "sampleTitles": [ + "00:00:04 Word meanings and political aisles", + "00:01:13 Tim Pool's Twitter ban for using the 'G word'", + "00:02:56 Twitter's advertising restrictions on Tim Pool" + ] + }, + { + "slug": "destiny/5nmDzKB23OU", + "bucket": "medium", + "durationSeconds": 4419, + "chunks": 2, + "chunksFailed": 0, + "zeroYieldChunks": 1, + "kept": 8, + "maxGapSeconds": 1880, + "engineSeconds": 550.203, + "inputTokens": 16390, + "outputTokens": 504, + "warningsByCode": { + "out-of-range": 10 + }, + "genericTitles": 2, + "duplicateTitles": 0, + "sampleTitles": [ + "00:26:10 Segment Transcript", + "00:27:58 Example of Moral System", + "00:34:30 General Claim about Psychology", + "00:36:17 Claim about Matrix Facts", + "00:40:22 Default Position on Normativity", + "00:41:04 Definition of Normativity", + "00:41:36 Christine Korsgaard's Theory of Normativity", + "00:42:19 Step Made by Christine Korsgaard" + ] + }, + { + "slug": "angryjoeshow/RPJqkewZP5I", + "bucket": "medium", + "durationSeconds": 3380, + "chunks": 2, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 12, + "maxGapSeconds": 1707, + "engineSeconds": 497.004, + "inputTokens": 15847, + "outputTokens": 355, + "warningsByCode": { + "out-of-range": 1 + }, + "genericTitles": 0, + "duplicateTitles": 0, + "sampleTitles": [ + "00:28:27 Silent Hill", + "00:29:50 High Guard", + "00:34:41 Crimson Desert", + "00:36:02 Chicken Man", + "00:37:19 God of War Remake", + "00:38:40 AI Deepfakes", + "00:41:37 Copyright Infringement in AI-Generated Content", + "00:42:23 Cance 2.0" + ] + } + ], + "totals": { + "videos": 5, + "audioHours": 2.72, + "chunks": 7, + "chunksFailed": 0, + "zeroYieldChunks": 1, + "zeroYieldRate": 0.1429, + "kept": 28, + "chaptersPerHour": 10.28, + "maxGapSeconds": 1880, + "meanGapSeconds": 971, + "genericTitleRate": 0.0714, + "duplicateTitleRate": 0, + "warningsByCode": { + "out-of-range": 11 + }, + "rejectionRate": 0.2821, + "engineSeconds": 1361, + "tokensPerSecond": 35, + "secondsPerAudioHour": 500, + "projectedSweepDays": 447.3 + } + }, + { + "candidate": { + "key": "qwen3:14b@8192/absolute", + "model": "qwen3:14b", + "numCtx": 8192, + "maxCues": 600, + "timestampMode": "absolute", + "think": false + }, + "videos": [ + { + "slug": "the-quartering-rumble/v6ve4n0", + "bucket": "short", + "durationSeconds": 791, + "chunks": 1, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 3, + "maxGapSeconds": 344, + "engineSeconds": 168.879, + "inputTokens": 4523, + "outputTokens": 119, + "warningsByCode": {}, + "genericTitles": 0, + "duplicateTitles": 0, + "sampleTitles": [ + "00:00:00 Trump's Troll on Colbert and the Kennedy Center Honors", + "00:03:19 Late Night Television's Decline and the Impact of Trump Derangement Syndrome", + "00:09:03 Democrats' Struggle with Crime and Trump's Political Strategy" + ] + }, + { + "slug": "chibi-reviews/VQykVuHd9xQ", + "bucket": "short", + "durationSeconds": 433, + "chunks": 1, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 1, + "maxGapSeconds": 431, + "engineSeconds": 92.794, + "inputTokens": 3576, + "outputTokens": 41, + "warningsByCode": {}, + "genericTitles": 0, + "duplicateTitles": 0, + "sampleTitles": [ + "00:00:02 Translation Errors on Crunchyroll" + ] + }, + { + "slug": "the-quartering/J0ySGwzP4Nw", + "bucket": "short", + "durationSeconds": 785, + "chunks": 1, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 3, + "maxGapSeconds": 301, + "engineSeconds": 207.878, + "inputTokens": 6067, + "outputTokens": 110, + "warningsByCode": {}, + "genericTitles": 2, + "duplicateTitles": 0, + "sampleTitles": [ + "00:00:04 Introduction to the Controversy Over Word Usage", + "00:03:06 Tim Pool's Twitter Ban and His Response", + "00:08:07 Discussion on Twitter's Enforcement and LGBTQ+ Terminology" + ] + }, + { + "slug": "destiny/5nmDzKB23OU", + "bucket": "medium", + "durationSeconds": 4419, + "chunks": 4, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 30, + "maxGapSeconds": 816, + "engineSeconds": 977.384, + "inputTokens": 19990, + "outputTokens": 954, + "warningsByCode": { + "seam-duplicate": 1 + }, + "genericTitles": 0, + "duplicateTitles": 0, + "sampleTitles": [ + "00:13:15 Different Reasoning Processes and the Concept of a Higher Power", + "00:13:45 Faith vs. Science in Practical Applications", + "00:15:05 The Role of Faith in Metaphysical vs. Material Claims", + "00:16:04 Philosophical vs. Societal Definitions of Religion", + "00:17:56 Assumptions About Religion and Their Implications", + "00:19:17 The Frustration with Narrow Interpretations of Religion", + "00:20:32 Conceding the Common Vernacular of Religion", + "00:32:50 The Challenge of Moral Facts and Neuroscience" + ] + }, + { + "slug": "angryjoeshow/RPJqkewZP5I", + "bucket": "medium", + "durationSeconds": 3380, + "chunks": 3, + "chunksFailed": 2, + "zeroYieldChunks": 2, + "kept": 8, + "maxGapSeconds": 2871, + "engineSeconds": 201.428, + "inputTokens": 4098, + "outputTokens": 196, + "warningsByCode": { + "chunk-failed": 2 + }, + "genericTitles": 1, + "duplicateTitles": 0, + "sampleTitles": [ + "00:47:51 AI in Healthcare and Accuracy Concerns", + "00:48:10 AI's Role in Managing Health Records", + "00:49:00 Escaping AI Algorithms and Content", + "00:50:00 AI Memes and Cultural Impact", + "00:51:00 AI and the Future of Entertainment", + "00:52:00 AI's Environmental Impact", + "00:54:00 Water Usage and AI", + "00:55:00 Conclusion and Advertisements" + ] + } + ], + "totals": { + "videos": 5, + "audioHours": 2.72, + "chunks": 10, + "chunksFailed": 2, + "zeroYieldChunks": 2, + "zeroYieldRate": 0.2, + "kept": 45, + "chaptersPerHour": 16.52, + "maxGapSeconds": 2871, + "meanGapSeconds": 953, + "genericTitleRate": 0.0667, + "duplicateTitleRate": 0, + "warningsByCode": { + "seam-duplicate": 1, + "chunk-failed": 2 + }, + "rejectionRate": 0.0426, + "engineSeconds": 1648, + "tokensPerSecond": 24.1, + "secondsPerAudioHour": 605, + "projectedSweepDays": 541.3 + } + }, + { + "candidate": { + "key": "qwen2.5:7b@16384/absolute", + "model": "qwen2.5:7b", + "numCtx": 16384, + "maxCues": 1200, + "timestampMode": "absolute" + }, + "videos": [ + { + "slug": "the-quartering-rumble/v6ve4n0", + "bucket": "short", + "durationSeconds": 791, + "chunks": 1, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 4, + "maxGapSeconds": 297, + "engineSeconds": 6.85, + "inputTokens": 4515, + "outputTokens": 137, + "warningsByCode": {}, + "genericTitles": 1, + "duplicateTitles": 0, + "sampleTitles": [ + "00:00:00 Introduction and Discord Promotion", + "00:01:49 Trump's Trolling of Stephen Colbert", + "00:03:56 Announcement of Trump Hosting Kennedy Center Honors", + "00:08:14 Late Night Television Challenges and Viewership Shifts" + ] + }, + { + "slug": "chibi-reviews/VQykVuHd9xQ", + "bucket": "short", + "durationSeconds": 433, + "chunks": 1, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 2, + "maxGapSeconds": 224, + "engineSeconds": 3.725, + "inputTokens": 3568, + "outputTokens": 71, + "warningsByCode": {}, + "genericTitles": 0, + "duplicateTitles": 0, + "sampleTitles": [ + "00:00:05 Crunchyroll's Translation Issues", + "00:03:49 Multiple Shows with Translation Errors" + ] + }, + { + "slug": "the-quartering/J0ySGwzP4Nw", + "bucket": "short", + "durationSeconds": 785, + "chunks": 1, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 3, + "maxGapSeconds": 516, + "engineSeconds": 5.496, + "inputTokens": 6059, + "outputTokens": 109, + "warningsByCode": {}, + "genericTitles": 1, + "duplicateTitles": 0, + "sampleTitles": [ + "00:00:48 Discussion on the redefinition of words", + "00:01:35 Tim Pool's ban from Twitter and his response", + "00:04:29 Journalist Tim Pool's criticism of Twitter's policies" + ] + }, + { + "slug": "destiny/5nmDzKB23OU", + "bucket": "medium", + "durationSeconds": 4419, + "chunks": 2, + "chunksFailed": 0, + "zeroYieldChunks": 1, + "kept": 6, + "maxGapSeconds": 1882, + "engineSeconds": 85.707, + "inputTokens": 16388, + "outputTokens": 507, + "warningsByCode": { + "out-of-range": 12 + }, + "genericTitles": 0, + "duplicateTitles": 0, + "sampleTitles": [ + "00:26:11 The easiest way to prove that ever exists", + "00:27:05 Does conceding the possibility of metaphysical truth change anything?", + "00:30:02 Is it possible to find moral facts by analyzing brain states?", + "00:40:00 Arguments against the existence of moral facts", + "00:41:03 Definition and explanation of normativity", + "00:42:17 Normativity in practical problems" + ] + }, + { + "slug": "angryjoeshow/RPJqkewZP5I", + "bucket": "medium", + "durationSeconds": 3380, + "chunks": 2, + "chunksFailed": 0, + "zeroYieldChunks": 1, + "kept": 3, + "maxGapSeconds": 2497, + "engineSeconds": 77.36699999999999, + "inputTokens": 15784, + "outputTokens": 363, + "warningsByCode": { + "out-of-range": 10 + }, + "genericTitles": 0, + "duplicateTitles": 0, + "sampleTitles": [ + "00:41:37 Copyright Issues with Unauthorized AI Creations", + "00:45:24 Sony Considering Delaying PlayStation Console Launch Due to AI Demand", + "00:48:10 AI Accuracy in Healthcare Records Management" + ] + } + ], + "totals": { + "videos": 5, + "audioHours": 2.72, + "chunks": 7, + "chunksFailed": 0, + "zeroYieldChunks": 2, + "zeroYieldRate": 0.2857, + "kept": 18, + "chaptersPerHour": 6.61, + "maxGapSeconds": 2497, + "meanGapSeconds": 1083, + "genericTitleRate": 0.1111, + "duplicateTitleRate": 0, + "warningsByCode": { + "out-of-range": 22 + }, + "rejectionRate": 0.55, + "engineSeconds": 179, + "tokensPerSecond": 265.2, + "secondsPerAudioHour": 66, + "projectedSweepDays": 59 + } + } + ] +} diff --git a/plans/bakeoff/round1.md b/plans/bakeoff/round1.md @@ -0,0 +1,262 @@ +# Digest bake-off — round1 + +Sample: 5 video(s) from `plans/bakeoff/sample.json` (buckets: short, medium), 2.72 audio-hours. +Sweep days are projected as measured seconds-per-audio-hour x 77298 corpus audio-hours, one lane, no parallelism. + +| Candidate | Zero-yield chunks | Chapters/h | Rejection rate | Max gap | Generic | Dup | tok/s | s per audio-h | **Sweep days** | +| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | +| `gemma2:9b@8192/absolute` | 0/10 (0%) | 15.78 | 0% | 00:14:59 | 2.3% | 0% | 55.2 | 322 | **288.1** | +| `qwen3:8b@16384/absolute` | 1/7 (14.3%) | 15.42 | 23.6% | 00:39:57 | 11.9% | 0% | 76.8 | 229 | **204.9** | +| `mistral-nemo:12b@16384/absolute` | 1/7 (14.3%) | 10.28 | 28.2% | 00:31:20 | 7.1% | 0% | 35 | 500 | **447.3** | +| `qwen3:14b@8192/absolute` | 2/10 (20%) | 16.52 | 4.3% | 00:47:51 | 6.7% | 0% | 24.1 | 605 | **541.3** | +| `qwen2.5:7b@16384/absolute` | 2/7 (28.6%) | 6.61 | 55% | 00:41:37 | 11.1% | 0% | 265.2 | 66 | **59** | + +## Rejections by guard + +| Candidate | chunk-failed | out-of-range | seam-duplicate | +| --- | --- | --- | --- | +| `gemma2:9b@8192/absolute` | 0 | 0 | 0 | +| `qwen3:8b@16384/absolute` | 0 | 13 | 0 | +| `mistral-nemo:12b@16384/absolute` | 0 | 11 | 0 | +| `qwen3:14b@8192/absolute` | 2 | 0 | 1 | +| `qwen2.5:7b@16384/absolute` | 0 | 22 | 0 | + +## Per-video + +| Candidate | Video | Bucket | Chunks | Zero-yield | Chapters | Max gap | Engine s | +| --- | --- | --- | --- | --- | --- | --- | --- | +| `gemma2:9b@8192/absolute` | `the-quartering-rumble/v6ve4n0` | short | 1 | 0 | 3 | 00:06:00 | 79 | +| `gemma2:9b@8192/absolute` | `chibi-reviews/VQykVuHd9xQ` | short | 1 | 0 | 2 | 00:04:08 | 51 | +| `gemma2:9b@8192/absolute` | `the-quartering/J0ySGwzP4Nw` | short | 1 | 0 | 3 | 00:05:31 | 105 | +| `gemma2:9b@8192/absolute` | `destiny/5nmDzKB23OU` | medium | 4 | 0 | 20 | 00:13:24 | 409 | +| `gemma2:9b@8192/absolute` | `angryjoeshow/RPJqkewZP5I` | medium | 3 | 0 | 15 | 00:14:59 | 233 | +| `qwen3:8b@16384/absolute` | `the-quartering-rumble/v6ve4n0` | short | 1 | 0 | 6 | 00:04:06 | 63 | +| `qwen3:8b@16384/absolute` | `chibi-reviews/VQykVuHd9xQ` | short | 1 | 0 | 4 | 00:03:50 | 35 | +| `qwen3:8b@16384/absolute` | `the-quartering/J0ySGwzP4Nw` | short | 1 | 0 | 7 | 00:02:48 | 76 | +| `qwen3:8b@16384/absolute` | `destiny/5nmDzKB23OU` | medium | 2 | 1 | 8 | 00:31:40 | 232 | +| `qwen3:8b@16384/absolute` | `angryjoeshow/RPJqkewZP5I` | medium | 2 | 0 | 17 | 00:39:57 | 218 | +| `mistral-nemo:12b@16384/absolute` | `the-quartering-rumble/v6ve4n0` | short | 1 | 0 | 3 | 00:04:57 | 107 | +| `mistral-nemo:12b@16384/absolute` | `chibi-reviews/VQykVuHd9xQ` | short | 1 | 0 | 2 | 00:06:04 | 69 | +| `mistral-nemo:12b@16384/absolute` | `the-quartering/J0ySGwzP4Nw` | short | 1 | 0 | 3 | 00:10:09 | 138 | +| `mistral-nemo:12b@16384/absolute` | `destiny/5nmDzKB23OU` | medium | 2 | 1 | 8 | 00:31:20 | 550 | +| `mistral-nemo:12b@16384/absolute` | `angryjoeshow/RPJqkewZP5I` | medium | 2 | 0 | 12 | 00:28:27 | 497 | +| `qwen3:14b@8192/absolute` | `the-quartering-rumble/v6ve4n0` | short | 1 | 0 | 3 | 00:05:44 | 169 | +| `qwen3:14b@8192/absolute` | `chibi-reviews/VQykVuHd9xQ` | short | 1 | 0 | 1 | 00:07:11 | 93 | +| `qwen3:14b@8192/absolute` | `the-quartering/J0ySGwzP4Nw` | short | 1 | 0 | 3 | 00:05:01 | 208 | +| `qwen3:14b@8192/absolute` | `destiny/5nmDzKB23OU` | medium | 4 | 0 | 30 | 00:13:36 | 977 | +| `qwen3:14b@8192/absolute` | `angryjoeshow/RPJqkewZP5I` | medium | 3 | 2 | 8 | 00:47:51 | 201 | +| `qwen2.5:7b@16384/absolute` | `the-quartering-rumble/v6ve4n0` | short | 1 | 0 | 4 | 00:04:57 | 7 | +| `qwen2.5:7b@16384/absolute` | `chibi-reviews/VQykVuHd9xQ` | short | 1 | 0 | 2 | 00:03:44 | 4 | +| `qwen2.5:7b@16384/absolute` | `the-quartering/J0ySGwzP4Nw` | short | 1 | 0 | 3 | 00:08:36 | 5 | +| `qwen2.5:7b@16384/absolute` | `destiny/5nmDzKB23OU` | medium | 2 | 1 | 6 | 00:31:22 | 86 | +| `qwen2.5:7b@16384/absolute` | `angryjoeshow/RPJqkewZP5I` | medium | 2 | 1 | 3 | 00:41:37 | 77 | + +## Sample output (first chapters per video) + +### `gemma2:9b@8192/absolute` + +**the-quartering-rumble/v6ve4n0** (short, 00:13:11) + +- 00:00:00 Trump Trolls Colbert While Setting Up a Trap for Democrats +- 00:04:00 Colbert's Cancellation and the Changing Landscape of Late Night Television +- 00:10:00 Trump's DC Move Forces Democrats to Address Crime + +**chibi-reviews/VQykVuHd9xQ** (short, 00:07:13) + +- 00:00:02 Crunchyroll's Translation Issues +- 00:04:10 Impact on Viewership Experience + +**the-quartering/J0ySGwzP4Nw** (short, 00:13:05) + +- 00:00:04 Tim Pool Discusses Censorship on Social Media +- 00:04:01 Twitter's Handling of the 'G Word' +- 00:09:32 Concerns About LGBTQ+ Content in Schools + +**destiny/5nmDzKB23OU** (medium, 01:13:39) + +- 00:13:24 The Pure Concept of Religion +- 00:16:53 Assumptions and Definitions +- 00:20:02 Beyond Christianity +- 00:20:41 Reconciling Perspectives +- 00:20:56 The Importance of Openness +- 00:32:58 Methodological Approach to Incorporating Evidence for Moral Facts +- 00:34:02 Determining Which Brain States are 'Moral' +- 00:37:01 Skepticism and the Burden of Proof + +**angryjoeshow/RPJqkewZP5I** (medium, 00:56:20) + +- 00:14:59 Game's Defense and Blame Shifting +- 00:15:02 Content Creators and Gamers Blamed +- 00:15:38 One Peg Response - Game's Issues Highlighted +- 00:16:21 Mixed Opinions on the Game +- 00:17:34 Studio's Responsibility and No Man's Sky Comparison +- 00:18:25 Microtransactions and Lack of Appeal +- 00:20:39 Jeff Keley's Role and Personal Taste +- 00:34:25 Twitter and Patch Notes Struggles + +### `qwen3:8b@16384/absolute` + +**the-quartering-rumble/v6ve4n0** (short, 00:13:11) + +- 00:00:00 Trump Trolls Colbert by Taking His Prime Time Job +- 00:02:03 Speculation About Secret Side Deals and Late Night Television +- 00:06:09 Paramount's Merger and the Late Night Show Format Crisis +- 00:08:31 Democrats Grapple with Crime and Trump's Strategy +- 00:10:06 Trump's Impact on Crime Perception and Democratic Vulnerabilities +- 00:12:01 Woke DAs and the Changing Crime Statistics + +**chibi-reviews/VQykVuHd9xQ** (short, 00:07:13) + +- 00:00:02 Crunchyroll's Translation Issues +- 00:01:00 Translation Errors in Gundam Episode +- 00:02:33 Unwatchable Episode Due to Subtitle Delays +- 00:03:23 Translation Problems in Overlord Series},{ + +**the-quartering/J0ySGwzP4Nw** (short, 00:13:05) + +- 00:00:04 Tim Pool's War With Twitter +- 00:01:13 The G Word and Banning +- 00:04:01 Free Speech and Platform Rules +- 00:06:00 Twitter's Alleged Protection of Creeps +- 00:08:01 Twitter's Uneven Enforcement +- 00:10:00 Concerns About Teachers and Young Ones +- 00:12:05 Criticizing the Platform's Policies + +**destiny/5nmDzKB23OU** (medium, 01:13:39) + +- 00:02:00 Exploring Moral Systems and Preferences +- 00:06:00 Meta Ethics and the Nature of Moral Truth +- 00:12:00 The Role of Neuroscience in Understanding Morality +- 00:18:00 Skepticism and the Burden of Proof +- 00:24:00 Normativity and Its Interpretations +- 00:30:00 Conclusion and Reflection on Moral Frameworks +- 00:36:00 Final Thoughts and Open Questions +- 00:42:00 Closing Remarks and Further Exploration + +**angryjoeshow/RPJqkewZP5I** (medium, 00:56:20) + +- 00:00:28 Silent Hill and New Game +- 00:00:30 John Wick and God of War +- 00:00:40 Marathon and High Guard +- 00:00:50 Crimson Desert and AI +- 00:01:00 State of Play and AI Controversy +- 00:01:10 Conclusion and Final Thoughts +- 00:01:20 Wrap-Up and Farewell +- 00:01:30 End of Episode + +### `mistral-nemo:12b@16384/absolute` + +**the-quartering-rumble/v6ve4n0** (short, 00:13:11) + +- 00:00:00 Trump's Ruthless Troll Against Colbert +- 00:04:09 Colbert's Late Show Cancellation and Trump's Involvement +- 00:08:14 Economic Challenges of Late Night Shows + +**chibi-reviews/VQykVuHd9xQ** (short, 00:07:13) + +- 00:00:02 Crunchyroll's translations issues +- 00:01:09 Gundam episode with subtitle delays + +**the-quartering/J0ySGwzP4Nw** (short, 00:13:05) + +- 00:00:04 Word meanings and political aisles +- 00:01:13 Tim Pool's Twitter ban for using the 'G word' +- 00:02:56 Twitter's advertising restrictions on Tim Pool + +**destiny/5nmDzKB23OU** (medium, 01:13:39) + +- 00:26:10 Segment Transcript +- 00:27:58 Example of Moral System +- 00:34:30 General Claim about Psychology +- 00:36:17 Claim about Matrix Facts +- 00:40:22 Default Position on Normativity +- 00:41:04 Definition of Normativity +- 00:41:36 Christine Korsgaard's Theory of Normativity +- 00:42:19 Step Made by Christine Korsgaard + +**angryjoeshow/RPJqkewZP5I** (medium, 00:56:20) + +- 00:28:27 Silent Hill +- 00:29:50 High Guard +- 00:34:41 Crimson Desert +- 00:36:02 Chicken Man +- 00:37:19 God of War Remake +- 00:38:40 AI Deepfakes +- 00:41:37 Copyright Infringement in AI-Generated Content +- 00:42:23 Cance 2.0 + +### `qwen3:14b@8192/absolute` + +**the-quartering-rumble/v6ve4n0** (short, 00:13:11) + +- 00:00:00 Trump's Troll on Colbert and the Kennedy Center Honors +- 00:03:19 Late Night Television's Decline and the Impact of Trump Derangement Syndrome +- 00:09:03 Democrats' Struggle with Crime and Trump's Political Strategy + +**chibi-reviews/VQykVuHd9xQ** (short, 00:07:13) + +- 00:00:02 Translation Errors on Crunchyroll + +**the-quartering/J0ySGwzP4Nw** (short, 00:13:05) + +- 00:00:04 Introduction to the Controversy Over Word Usage +- 00:03:06 Tim Pool's Twitter Ban and His Response +- 00:08:07 Discussion on Twitter's Enforcement and LGBTQ+ Terminology + +**destiny/5nmDzKB23OU** (medium, 01:13:39) + +- 00:13:15 Different Reasoning Processes and the Concept of a Higher Power +- 00:13:45 Faith vs. Science in Practical Applications +- 00:15:05 The Role of Faith in Metaphysical vs. Material Claims +- 00:16:04 Philosophical vs. Societal Definitions of Religion +- 00:17:56 Assumptions About Religion and Their Implications +- 00:19:17 The Frustration with Narrow Interpretations of Religion +- 00:20:32 Conceding the Common Vernacular of Religion +- 00:32:50 The Challenge of Moral Facts and Neuroscience + +**angryjoeshow/RPJqkewZP5I** (medium, 00:56:20) + +- 00:47:51 AI in Healthcare and Accuracy Concerns +- 00:48:10 AI's Role in Managing Health Records +- 00:49:00 Escaping AI Algorithms and Content +- 00:50:00 AI Memes and Cultural Impact +- 00:51:00 AI and the Future of Entertainment +- 00:52:00 AI's Environmental Impact +- 00:54:00 Water Usage and AI +- 00:55:00 Conclusion and Advertisements + +### `qwen2.5:7b@16384/absolute` + +**the-quartering-rumble/v6ve4n0** (short, 00:13:11) + +- 00:00:00 Introduction and Discord Promotion +- 00:01:49 Trump's Trolling of Stephen Colbert +- 00:03:56 Announcement of Trump Hosting Kennedy Center Honors +- 00:08:14 Late Night Television Challenges and Viewership Shifts + +**chibi-reviews/VQykVuHd9xQ** (short, 00:07:13) + +- 00:00:05 Crunchyroll's Translation Issues +- 00:03:49 Multiple Shows with Translation Errors + +**the-quartering/J0ySGwzP4Nw** (short, 00:13:05) + +- 00:00:48 Discussion on the redefinition of words +- 00:01:35 Tim Pool's ban from Twitter and his response +- 00:04:29 Journalist Tim Pool's criticism of Twitter's policies + +**destiny/5nmDzKB23OU** (medium, 01:13:39) + +- 00:26:11 The easiest way to prove that ever exists +- 00:27:05 Does conceding the possibility of metaphysical truth change anything? +- 00:30:02 Is it possible to find moral facts by analyzing brain states? +- 00:40:00 Arguments against the existence of moral facts +- 00:41:03 Definition and explanation of normativity +- 00:42:17 Normativity in practical problems + +**angryjoeshow/RPJqkewZP5I** (medium, 00:56:20) + +- 00:41:37 Copyright Issues with Unauthorized AI Creations +- 00:45:24 Sony Considering Delaying PlayStation Console Launch Due to AI Demand +- 00:48:10 AI Accuracy in Healthcare Records Management + diff --git a/plans/bakeoff/round2.json b/plans/bakeoff/round2.json @@ -0,0 +1,562 @@ +{ + "label": "round2", + "sample": "/home/user/Projects/yt-dlp-transcript-browser/plans/bakeoff/sample.json", + "corpus": { + "videosScanned": 76354, + "videosWithTranscript": 73367, + "audioHours": 77298, + "longTailVideos": 6038, + "longTailAudioHours": 40153 + }, + "buckets": [ + "long" + ], + "scores": [ + { + "candidate": { + "key": "gemma2:9b@8192/chunk-local", + "model": "gemma2:9b", + "numCtx": 8192, + "maxCues": 600, + "timestampMode": "chunk-local" + }, + "videos": [ + { + "slug": "rekietalaw/EsZhaCfc8HQ", + "bucket": "long", + "durationSeconds": 12497, + "chunks": 8, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 44, + "maxGapSeconds": 1228, + "engineSeconds": 502.5470000000001, + "inputTokens": 30597, + "outputTokens": 1816, + "warningsByCode": { + "out-of-range": 4, + "non-monotonic": 3 + }, + "genericTitles": 7, + "duplicateTitles": 1, + "sampleTitles": [ + "00:19:17 Introduction and Bronca's Take", + "00:20:36 Discussing the Shooting Incident", + "00:24:18 Andrew Wilson and His Comparison to a Serial Killer", + "00:24:54 Analyzing the Lawyer's Argument", + "00:26:32 The Fourth Amendment and Reasonable Belief", + "00:27:37 Minnesota Law and Police Shootings", + "00:42:50 ICE Jurisdiction and State Laws", + "00:48:03 The Incident: Vehicle Approach and Officer Orders" + ] + }, + { + "slug": "chrissie-mayr/cfLF2o2-0BA", + "bucket": "long", + "durationSeconds": 12553, + "chunks": 9, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 46, + "maxGapSeconds": 1444, + "engineSeconds": 527.155, + "inputTokens": 37239, + "outputTokens": 1958, + "warningsByCode": { + "non-monotonic": 7, + "out-of-range": 1 + }, + "genericTitles": 2, + "duplicateTitles": 0, + "sampleTitles": [ + "00:20:33 Initial Reaction and Expectations", + "00:20:57 Humor and Critique of Masculinity", + "00:22:04 Addressing the 'Woke' Label", + "00:24:51 Comparison to Past Chick Flicks", + "00:26:37 The Impact of Online Movie Reviews", + "00:27:33 Ryan Gosling's Performance and Humor", + "00:44:41 Barbie's Throwaway Boyfriend and Teresa's Absence", + "00:51:40 The Appeal of Barbie History and Gender Differences" + ] + } + ], + "totals": { + "videos": 2, + "audioHours": 6.96, + "chunks": 17, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "zeroYieldRate": 0, + "kept": 90, + "chaptersPerHour": 12.93, + "maxGapSeconds": 1444, + "meanGapSeconds": 1336, + "genericTitleRate": 0.1, + "duplicateTitleRate": 0.0111, + "warningsByCode": { + "out-of-range": 5, + "non-monotonic": 10 + }, + "rejectionRate": 0.1429, + "engineSeconds": 1030, + "tokensPerSecond": 69.5, + "secondsPerAudioHour": 148, + "projectedSweepDays": 132.4 + } + }, + { + "candidate": { + "key": "qwen2.5:7b@16384/chunk-local", + "model": "qwen2.5:7b", + "numCtx": 16384, + "maxCues": 1200, + "timestampMode": "chunk-local" + }, + "videos": [ + { + "slug": "rekietalaw/EsZhaCfc8HQ", + "bucket": "long", + "durationSeconds": 12497, + "chunks": 4, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 36, + "maxGapSeconds": 2981, + "engineSeconds": 95.23400000000001, + "inputTokens": 34715, + "outputTokens": 1360, + "warningsByCode": { + "out-of-range": 17, + "seam-duplicate": 1 + }, + "genericTitles": 4, + "duplicateTitles": 1, + "sampleTitles": [ + "00:00:19 Introduction and Context", + "00:01:42 The Incident and Initial Analysis", + "00:05:15 Legal Justification for Use of Force", + "00:10:46 Use of Deadly Force and Self-Defense", + "00:14:18 Tactical Considerations and Officer Safety", + "00:19:05 The Role of Bystanders and Video Evidence", + "00:23:16 Public Perception and Media Coverage", + "00:30:48 Legal Analysis and Case Law" + ] + }, + { + "slug": "chrissie-mayr/cfLF2o2-0BA", + "bucket": "long", + "durationSeconds": 12553, + "chunks": 5, + "chunksFailed": 0, + "zeroYieldChunks": 1, + "kept": 29, + "maxGapSeconds": 5268, + "engineSeconds": 100.58200000000001, + "inputTokens": 34157, + "outputTokens": 1592, + "warningsByCode": { + "out-of-range": 29 + }, + "genericTitles": 9, + "duplicateTitles": 1, + "sampleTitles": [ + "01:27:49 Opinions on the Barbie movie", + "01:31:24 Trust in opinions based on experience", + "01:33:28 Comparison to other movies and remakes", + "01:35:38 Amy Schumer's involvement with the Barbie project", + "01:40:01 Discussion on feminism in the Barbie movie", + "01:43:27 Dreams of childhood and their impact", + "01:45:52 Barbie movies for a different age group", + "01:46:35 Introduction and Opening Remarks" + ] + } + ], + "totals": { + "videos": 2, + "audioHours": 6.96, + "chunks": 9, + "chunksFailed": 0, + "zeroYieldChunks": 1, + "zeroYieldRate": 0.1111, + "kept": 65, + "chaptersPerHour": 9.34, + "maxGapSeconds": 5268, + "meanGapSeconds": 4125, + "genericTitleRate": 0.2, + "duplicateTitleRate": 0.0308, + "warningsByCode": { + "out-of-range": 46, + "seam-duplicate": 1 + }, + "rejectionRate": 0.4144, + "engineSeconds": 196, + "tokensPerSecond": 366.8, + "secondsPerAudioHour": 28, + "projectedSweepDays": 25.1 + } + }, + { + "candidate": { + "key": "qwen2.5:7b@8192/chunk-local", + "model": "qwen2.5:7b", + "numCtx": 8192, + "maxCues": 600, + "timestampMode": "chunk-local" + }, + "videos": [ + { + "slug": "rekietalaw/EsZhaCfc8HQ", + "bucket": "long", + "durationSeconds": 12497, + "chunks": 8, + "chunksFailed": 0, + "zeroYieldChunks": 0, + "kept": 47, + "maxGapSeconds": 1382, + "engineSeconds": 85.91499999999999, + "inputTokens": 30565, + "outputTokens": 1441, + "warningsByCode": { + "out-of-range": 4, + "non-monotonic": 2 + }, + "genericTitles": 16, + "duplicateTitles": 3, + "sampleTitles": [ + "00:00:19 Introduction and Context", + "00:02:15 Overview of the Incident", + "00:04:42 Legal Analysis of Use of Force", + "00:07:57 Discussion on Objective Reasonable Fear", + "00:11:15 Subjective Fear and Its Role", + "00:14:48 Minnesota's Specific Law on Use of Force", + "00:17:57 Impact of the Lawsuit on Objective Reasonable Fear Standard", + "00:21:15 Conclusion and Final Thoughts" + ] + }, + { + "slug": "chrissie-mayr/cfLF2o2-0BA", + "bucket": "long", + "durationSeconds": 12553, + "chunks": 9, + "chunksFailed": 0, + "zeroYieldChunks": 2, + "kept": 46, + "maxGapSeconds": 1453, + "engineSeconds": 104.55800000000002, + "inputTokens": 37164, + "outputTokens": 1785, + "warningsByCode": { + "out-of-range": 9, + "language-drift": 7 + }, + "genericTitles": 13, + "duplicateTitles": 0, + "sampleTitles": [ + "00:20:21 Personal enjoyment and expectations", + "00:20:57 Anti-male themes in the movie", + "00:22:04 Balancing empowerment with criticism", + "00:24:17 Target audience and gender dynamics", + "00:25:49 Impact of political correctness on entertainment", + "00:27:33 Ryan Gosling's humor in the film", + "00:51:45 Introduction and Personal Background", + "00:58:20 Discussion on the impact of moving to Alabama" + ] + } + ], + "totals": { + "videos": 2, + "audioHours": 6.96, + "chunks": 17, + "chunksFailed": 0, + "zeroYieldChunks": 2, + "zeroYieldRate": 0.1176, + "kept": 93, + "chaptersPerHour": 13.37, + "maxGapSeconds": 1453, + "meanGapSeconds": 1418, + "genericTitleRate": 0.3118, + "duplicateTitleRate": 0.0323, + "warningsByCode": { + "out-of-range": 13, + "non-monotonic": 2, + "language-drift": 7 + }, + "rejectionRate": 0.1913, + "engineSeconds": 190, + "tokensPerSecond": 372.5, + "secondsPerAudioHour": 27, + "projectedSweepDays": 24.2 + } + }, + { + "candidate": { + "key": "gemma2:9b@8192/absolute", + "model": "gemma2:9b", + "numCtx": 8192, + "maxCues": 600, + "timestampMode": "absolute" + }, + "videos": [ + { + "slug": "rekietalaw/EsZhaCfc8HQ", + "bucket": "long", + "durationSeconds": 12497, + "chunks": 8, + "chunksFailed": 0, + "zeroYieldChunks": 1, + "kept": 37, + "maxGapSeconds": 1634, + "engineSeconds": 456.12999999999994, + "inputTokens": 30597, + "outputTokens": 1810, + "warningsByCode": { + "non-monotonic": 7, + "out-of-range": 7 + }, + "genericTitles": 3, + "duplicateTitles": 0, + "sampleTitles": [ + "00:19:17 Introduction and Bronca's Take", + "00:20:36 Discussing the Shooting Video", + "00:24:18 Andrew Wilson and Scarlett Hampton", + "00:25:00 Legal Analysis of the Shooting", + "00:27:32 Objective vs. Subjective Fear", + "00:27:58 Continuing the Legal Analysis", + "00:42:50 ICE Jurisdiction and State Laws", + "00:43:12 State Charges Against Federal Agents" + ] + }, + { + "slug": "chrissie-mayr/cfLF2o2-0BA", + "bucket": "long", + "durationSeconds": 12553, + "chunks": 9, + "chunksFailed": 0, + "zeroYieldChunks": 2, + "kept": 40, + "maxGapSeconds": 3847, + "engineSeconds": 535.993, + "inputTokens": 37239, + "outputTokens": 2021, + "warningsByCode": { + "non-monotonic": 1, + "out-of-range": 15 + }, + "genericTitles": 1, + "duplicateTitles": 0, + "sampleTitles": [ + "00:20:23 Barbie Movie Review - A Divided Opinion", + "00:21:48 The Pressure of Manhood and Womanhood", + "00:23:57 Barbie vs. Fight Club - A Comparison", + "00:26:19 The Political Brain Rot of Entertainment", + "00:27:48 Ryan Gosling's Ken - A Hilarious Performance", + "00:28:19 The Making of a Ken Doll - Ryan Gosling's Daughter", + "00:44:41 Barbie's Throwaway Boyfriend and Teresa's Absence", + "00:45:21 Greta Gerwig and the Brunettes of the World" + ] + } + ], + "totals": { + "videos": 2, + "audioHours": 6.96, + "chunks": 17, + "chunksFailed": 0, + "zeroYieldChunks": 3, + "zeroYieldRate": 0.1765, + "kept": 77, + "chaptersPerHour": 11.07, + "maxGapSeconds": 3847, + "meanGapSeconds": 2741, + "genericTitleRate": 0.0519, + "duplicateTitleRate": 0, + "warningsByCode": { + "non-monotonic": 8, + "out-of-range": 22 + }, + "rejectionRate": 0.2804, + "engineSeconds": 992, + "tokensPerSecond": 72.2, + "secondsPerAudioHour": 143, + "projectedSweepDays": 127.9 + } + }, + { + "candidate": { + "key": "qwen2.5:7b@8192/absolute", + "model": "qwen2.5:7b", + "numCtx": 8192, + "maxCues": 600, + "timestampMode": "absolute" + }, + "videos": [ + { + "slug": "rekietalaw/EsZhaCfc8HQ", + "bucket": "long", + "durationSeconds": 12497, + "chunks": 8, + "chunksFailed": 0, + "zeroYieldChunks": 2, + "kept": 32, + "maxGapSeconds": 3291, + "engineSeconds": 104.74499999999999, + "inputTokens": 30565, + "outputTokens": 1955, + "warningsByCode": { + "out-of-range": 37 + }, + "genericTitles": 8, + "duplicateTitles": 0, + "sampleTitles": [ + "00:01:36 Overview of the Incident", + "00:04:25 Use of Force Analysis", + "00:07:38 Legal Framework and Reasonable Fear", + "00:11:29 Subjective vs. Objective Reasonable Fear", + "00:14:56 Minnesota's Use of Force Law", + "00:18:37 Discussion with Bronca", + "00:22:41 Critique of Andrew Wilson's Analysis", + "00:25:34 Conclusion and Legal Implications" + ] + }, + { + "slug": "chrissie-mayr/cfLF2o2-0BA", + "bucket": "long", + "durationSeconds": 12553, + "chunks": 9, + "chunksFailed": 0, + "zeroYieldChunks": 3, + "kept": 40, + "maxGapSeconds": 3916, + "engineSeconds": 108.733, + "inputTokens": 37164, + "outputTokens": 1969, + "warningsByCode": { + "out-of-range": 28 + }, + "genericTitles": 5, + "duplicateTitles": 0, + "sampleTitles": [ + "00:19:56 Review of the movie", + "00:20:07 Enjoyment and target audience", + "00:20:34 Criticism of the review culture", + "00:21:55 Discussion on gender roles and pressures", + "00:24:07 Comparison with other movies", + "00:26:58 Sensationalism in movie reviews", + "01:08:05 Personal Background and Travel Plans", + "01:08:34 East Coast vs. Southern East Coast" + ] + } + ], + "totals": { + "videos": 2, + "audioHours": 6.96, + "chunks": 17, + "chunksFailed": 0, + "zeroYieldChunks": 5, + "zeroYieldRate": 0.2941, + "kept": 72, + "chaptersPerHour": 10.35, + "maxGapSeconds": 3916, + "meanGapSeconds": 3604, + "genericTitleRate": 0.1806, + "duplicateTitleRate": 0, + "warningsByCode": { + "out-of-range": 65 + }, + "rejectionRate": 0.4745, + "engineSeconds": 213, + "tokensPerSecond": 335.6, + "secondsPerAudioHour": 31, + "projectedSweepDays": 27.7 + } + }, + { + "candidate": { + "key": "qwen2.5:7b@16384/absolute", + "model": "qwen2.5:7b", + "numCtx": 16384, + "maxCues": 1200, + "timestampMode": "absolute" + }, + "videos": [ + { + "slug": "rekietalaw/EsZhaCfc8HQ", + "bucket": "long", + "durationSeconds": 12497, + "chunks": 4, + "chunksFailed": 0, + "zeroYieldChunks": 1, + "kept": 22, + "maxGapSeconds": 3527, + "engineSeconds": 102.381, + "inputTokens": 34715, + "outputTokens": 1322, + "warningsByCode": { + "out-of-range": 31 + }, + "genericTitles": 3, + "duplicateTitles": 0, + "sampleTitles": [ + "00:01:23 The Incident and Initial Analysis", + "00:05:46 Legal Justification for Use of Force", + "00:17:28 Analysis of the Video Evidence", + "00:35:15 Police Authority and Self-Defense", + "00:44:12 Discussion on Legal Standards and Analysis", + "00:48:05 Video Evidence and Analysis", + "00:53:52 Imminence of Threat", + "00:56:02 Conclusion and Final Thoughts" + ] + }, + { + "slug": "chrissie-mayr/cfLF2o2-0BA", + "bucket": "long", + "durationSeconds": 12553, + "chunks": 5, + "chunksFailed": 0, + "zeroYieldChunks": 2, + "kept": 35, + "maxGapSeconds": 5268, + "engineSeconds": 104.44, + "inputTokens": 34157, + "outputTokens": 1783, + "warningsByCode": { + "out-of-range": 27 + }, + "genericTitles": 7, + "duplicateTitles": 0, + "sampleTitles": [ + "01:27:49 Disagreement about the impact of your opinion on others", + "01:28:06 Discussion on a potential conservative Barbie movie", + "01:28:54 Criticism of the original Barbie movie's content and tone", + "01:30:09 Comparison between different opinions on the same topic", + "01:31:00 Defense of a male reviewer's opinion on Barbie", + "01:32:05 Criticism of Ben Shapiro's review as overly negative and unbalanced", + "01:34:06 Discussion about potential Ken movies with Ryan Gosling", + "01:37:03 Exploration of different interpretations of feminism" + ] + } + ], + "totals": { + "videos": 2, + "audioHours": 6.96, + "chunks": 9, + "chunksFailed": 0, + "zeroYieldChunks": 3, + "zeroYieldRate": 0.3333, + "kept": 57, + "chaptersPerHour": 8.19, + "maxGapSeconds": 5268, + "meanGapSeconds": 4398, + "genericTitleRate": 0.1754, + "duplicateTitleRate": 0, + "warningsByCode": { + "out-of-range": 58 + }, + "rejectionRate": 0.5043, + "engineSeconds": 207, + "tokensPerSecond": 348, + "secondsPerAudioHour": 30, + "projectedSweepDays": 26.8 + } + } + ] +} diff --git a/plans/bakeoff/round2.md b/plans/bakeoff/round2.md @@ -0,0 +1,188 @@ +# Digest bake-off — round2 + +Sample: 2 video(s) from `plans/bakeoff/sample.json` (buckets: long), 6.96 audio-hours. +Sweep days are projected as measured seconds-per-audio-hour x 77298 corpus audio-hours, one lane, no parallelism. + +| Candidate | Zero-yield chunks | Chapters/h | Rejection rate | Max gap | Generic | Dup | tok/s | s per audio-h | **Sweep days** | +| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | +| `gemma2:9b@8192/chunk-local` | 0/17 (0%) | 12.93 | 14.3% | 00:24:04 | 10% | 1.1% | 69.5 | 148 | **132.4** | +| `qwen2.5:7b@16384/chunk-local` | 1/9 (11.1%) | 9.34 | 41.4% | 01:27:48 | 20% | 3.1% | 366.8 | 28 | **25.1** | +| `qwen2.5:7b@8192/chunk-local` | 2/17 (11.8%) | 13.37 | 19.1% | 00:24:13 | 31.2% | 3.2% | 372.5 | 27 | **24.2** | +| `gemma2:9b@8192/absolute` | 3/17 (17.7%) | 11.07 | 28% | 01:04:07 | 5.2% | 0% | 72.2 | 143 | **127.9** | +| `qwen2.5:7b@8192/absolute` | 5/17 (29.4%) | 10.35 | 47.5% | 01:05:16 | 18.1% | 0% | 335.6 | 31 | **27.7** | +| `qwen2.5:7b@16384/absolute` | 3/9 (33.3%) | 8.19 | 50.4% | 01:27:48 | 17.5% | 0% | 348 | 30 | **26.8** | + +## Rejections by guard + +| Candidate | language-drift | non-monotonic | out-of-range | seam-duplicate | +| --- | --- | --- | --- | --- | +| `gemma2:9b@8192/chunk-local` | 0 | 10 | 5 | 0 | +| `qwen2.5:7b@16384/chunk-local` | 0 | 0 | 46 | 1 | +| `qwen2.5:7b@8192/chunk-local` | 7 | 2 | 13 | 0 | +| `gemma2:9b@8192/absolute` | 0 | 8 | 22 | 0 | +| `qwen2.5:7b@8192/absolute` | 0 | 0 | 65 | 0 | +| `qwen2.5:7b@16384/absolute` | 0 | 0 | 58 | 0 | + +## Per-video + +| Candidate | Video | Bucket | Chunks | Zero-yield | Chapters | Max gap | Engine s | +| --- | --- | --- | --- | --- | --- | --- | --- | +| `gemma2:9b@8192/chunk-local` | `rekietalaw/EsZhaCfc8HQ` | long | 8 | 0 | 44 | 00:20:28 | 503 | +| `gemma2:9b@8192/chunk-local` | `chrissie-mayr/cfLF2o2-0BA` | long | 9 | 0 | 46 | 00:24:04 | 527 | +| `qwen2.5:7b@16384/chunk-local` | `rekietalaw/EsZhaCfc8HQ` | long | 4 | 0 | 36 | 00:49:41 | 95 | +| `qwen2.5:7b@16384/chunk-local` | `chrissie-mayr/cfLF2o2-0BA` | long | 5 | 1 | 29 | 01:27:48 | 101 | +| `qwen2.5:7b@8192/chunk-local` | `rekietalaw/EsZhaCfc8HQ` | long | 8 | 0 | 47 | 00:23:02 | 86 | +| `qwen2.5:7b@8192/chunk-local` | `chrissie-mayr/cfLF2o2-0BA` | long | 9 | 2 | 46 | 00:24:13 | 105 | +| `gemma2:9b@8192/absolute` | `rekietalaw/EsZhaCfc8HQ` | long | 8 | 1 | 37 | 00:27:14 | 456 | +| `gemma2:9b@8192/absolute` | `chrissie-mayr/cfLF2o2-0BA` | long | 9 | 2 | 40 | 01:04:07 | 536 | +| `qwen2.5:7b@8192/absolute` | `rekietalaw/EsZhaCfc8HQ` | long | 8 | 2 | 32 | 00:54:51 | 105 | +| `qwen2.5:7b@8192/absolute` | `chrissie-mayr/cfLF2o2-0BA` | long | 9 | 3 | 40 | 01:05:16 | 109 | +| `qwen2.5:7b@16384/absolute` | `rekietalaw/EsZhaCfc8HQ` | long | 4 | 1 | 22 | 00:58:47 | 102 | +| `qwen2.5:7b@16384/absolute` | `chrissie-mayr/cfLF2o2-0BA` | long | 5 | 2 | 35 | 01:27:48 | 104 | + +## Sample output (first chapters per video) + +### `gemma2:9b@8192/chunk-local` + +**rekietalaw/EsZhaCfc8HQ** (long, 03:28:17) + +- 00:19:17 Introduction and Bronca's Take +- 00:20:36 Discussing the Shooting Incident +- 00:24:18 Andrew Wilson and His Comparison to a Serial Killer +- 00:24:54 Analyzing the Lawyer's Argument +- 00:26:32 The Fourth Amendment and Reasonable Belief +- 00:27:37 Minnesota Law and Police Shootings +- 00:42:50 ICE Jurisdiction and State Laws +- 00:48:03 The Incident: Vehicle Approach and Officer Orders + +**chrissie-mayr/cfLF2o2-0BA** (long, 03:29:13) + +- 00:20:33 Initial Reaction and Expectations +- 00:20:57 Humor and Critique of Masculinity +- 00:22:04 Addressing the 'Woke' Label +- 00:24:51 Comparison to Past Chick Flicks +- 00:26:37 The Impact of Online Movie Reviews +- 00:27:33 Ryan Gosling's Performance and Humor +- 00:44:41 Barbie's Throwaway Boyfriend and Teresa's Absence +- 00:51:40 The Appeal of Barbie History and Gender Differences + +### `qwen2.5:7b@16384/chunk-local` + +**rekietalaw/EsZhaCfc8HQ** (long, 03:28:17) + +- 00:00:19 Introduction and Context +- 00:01:42 The Incident and Initial Analysis +- 00:05:15 Legal Justification for Use of Force +- 00:10:46 Use of Deadly Force and Self-Defense +- 00:14:18 Tactical Considerations and Officer Safety +- 00:19:05 The Role of Bystanders and Video Evidence +- 00:23:16 Public Perception and Media Coverage +- 00:30:48 Legal Analysis and Case Law + +**chrissie-mayr/cfLF2o2-0BA** (long, 03:29:13) + +- 01:27:49 Opinions on the Barbie movie +- 01:31:24 Trust in opinions based on experience +- 01:33:28 Comparison to other movies and remakes +- 01:35:38 Amy Schumer's involvement with the Barbie project +- 01:40:01 Discussion on feminism in the Barbie movie +- 01:43:27 Dreams of childhood and their impact +- 01:45:52 Barbie movies for a different age group +- 01:46:35 Introduction and Opening Remarks + +### `qwen2.5:7b@8192/chunk-local` + +**rekietalaw/EsZhaCfc8HQ** (long, 03:28:17) + +- 00:00:19 Introduction and Context +- 00:02:15 Overview of the Incident +- 00:04:42 Legal Analysis of Use of Force +- 00:07:57 Discussion on Objective Reasonable Fear +- 00:11:15 Subjective Fear and Its Role +- 00:14:48 Minnesota's Specific Law on Use of Force +- 00:17:57 Impact of the Lawsuit on Objective Reasonable Fear Standard +- 00:21:15 Conclusion and Final Thoughts + +**chrissie-mayr/cfLF2o2-0BA** (long, 03:29:13) + +- 00:20:21 Personal enjoyment and expectations +- 00:20:57 Anti-male themes in the movie +- 00:22:04 Balancing empowerment with criticism +- 00:24:17 Target audience and gender dynamics +- 00:25:49 Impact of political correctness on entertainment +- 00:27:33 Ryan Gosling's humor in the film +- 00:51:45 Introduction and Personal Background +- 00:58:20 Discussion on the impact of moving to Alabama + +### `gemma2:9b@8192/absolute` + +**rekietalaw/EsZhaCfc8HQ** (long, 03:28:17) + +- 00:19:17 Introduction and Bronca's Take +- 00:20:36 Discussing the Shooting Video +- 00:24:18 Andrew Wilson and Scarlett Hampton +- 00:25:00 Legal Analysis of the Shooting +- 00:27:32 Objective vs. Subjective Fear +- 00:27:58 Continuing the Legal Analysis +- 00:42:50 ICE Jurisdiction and State Laws +- 00:43:12 State Charges Against Federal Agents + +**chrissie-mayr/cfLF2o2-0BA** (long, 03:29:13) + +- 00:20:23 Barbie Movie Review - A Divided Opinion +- 00:21:48 The Pressure of Manhood and Womanhood +- 00:23:57 Barbie vs. Fight Club - A Comparison +- 00:26:19 The Political Brain Rot of Entertainment +- 00:27:48 Ryan Gosling's Ken - A Hilarious Performance +- 00:28:19 The Making of a Ken Doll - Ryan Gosling's Daughter +- 00:44:41 Barbie's Throwaway Boyfriend and Teresa's Absence +- 00:45:21 Greta Gerwig and the Brunettes of the World + +### `qwen2.5:7b@8192/absolute` + +**rekietalaw/EsZhaCfc8HQ** (long, 03:28:17) + +- 00:01:36 Overview of the Incident +- 00:04:25 Use of Force Analysis +- 00:07:38 Legal Framework and Reasonable Fear +- 00:11:29 Subjective vs. Objective Reasonable Fear +- 00:14:56 Minnesota's Use of Force Law +- 00:18:37 Discussion with Bronca +- 00:22:41 Critique of Andrew Wilson's Analysis +- 00:25:34 Conclusion and Legal Implications + +**chrissie-mayr/cfLF2o2-0BA** (long, 03:29:13) + +- 00:19:56 Review of the movie +- 00:20:07 Enjoyment and target audience +- 00:20:34 Criticism of the review culture +- 00:21:55 Discussion on gender roles and pressures +- 00:24:07 Comparison with other movies +- 00:26:58 Sensationalism in movie reviews +- 01:08:05 Personal Background and Travel Plans +- 01:08:34 East Coast vs. Southern East Coast + +### `qwen2.5:7b@16384/absolute` + +**rekietalaw/EsZhaCfc8HQ** (long, 03:28:17) + +- 00:01:23 The Incident and Initial Analysis +- 00:05:46 Legal Justification for Use of Force +- 00:17:28 Analysis of the Video Evidence +- 00:35:15 Police Authority and Self-Defense +- 00:44:12 Discussion on Legal Standards and Analysis +- 00:48:05 Video Evidence and Analysis +- 00:53:52 Imminence of Threat +- 00:56:02 Conclusion and Final Thoughts + +**chrissie-mayr/cfLF2o2-0BA** (long, 03:29:13) + +- 01:27:49 Disagreement about the impact of your opinion on others +- 01:28:06 Discussion on a potential conservative Barbie movie +- 01:28:54 Criticism of the original Barbie movie's content and tone +- 01:30:09 Comparison between different opinions on the same topic +- 01:31:00 Defense of a male reviewer's opinion on Barbie +- 01:32:05 Criticism of Ben Shapiro's review as overly negative and unbalanced +- 01:34:06 Discussion about potential Ken movies with Ryan Gosling +- 01:37:03 Exploration of different interpretations of feminism + diff --git a/plans/bakeoff/sample.json b/plans/bakeoff/sample.json @@ -0,0 +1,93 @@ +{ + "version": 1, + "pickedAt": "2026-07-26T17:58:35.059Z", + "corpus": { + "videosScanned": 76354, + "videosWithTranscript": 73367, + "audioHours": 77298, + "longTailVideos": 6038, + "longTailAudioHours": 40153 + }, + "videos": [ + { + "slug": "the-quartering-rumble/v6ve4n0", + "channelSlug": "the-quartering-rumble", + "videoId": "v6ve4n0", + "videoDir": "v6xl0vu", + "title": "Donald Trump TROLLS Stephen Colbert While Setting Up GENIUS Trap That Woke Leftists Immediately Fall", + "bucket": "short", + "durationSeconds": 791, + "cueCount": 134 + }, + { + "slug": "chibi-reviews/VQykVuHd9xQ", + "channelSlug": "chibi-reviews", + "videoId": "VQykVuHd9xQ", + "videoDir": "VQykVuHd9xQ", + "title": "Crunchyroll This is Unacceptable! Your Service is LITERALLY Unwatchable", + "bucket": "short", + "durationSeconds": 433, + "cueCount": 173 + }, + { + "slug": "the-quartering/J0ySGwzP4Nw", + "channelSlug": "the-quartering", + "videoId": "J0ySGwzP4Nw", + "videoDir": "J0ySGwzP4Nw", + "title": "Tim Pool Goes To WAR With Twitter After Being Banned For Saying The FORBIDDEN Word On Timcast IRL", + "bucket": "short", + "durationSeconds": 785, + "cueCount": 315 + }, + { + "slug": "destiny/5nmDzKB23OU", + "channelSlug": "destiny", + "videoId": "5nmDzKB23OU", + "videoDir": "5nmDzKB23OU", + "title": "Can Science Answer All Questions?", + "bucket": "medium", + "durationSeconds": 4419, + "cueCount": 2072 + }, + { + "slug": "angryjoeshow/RPJqkewZP5I", + "channelSlug": "angryjoeshow", + "videoId": "RPJqkewZP5I", + "videoDir": "RPJqkewZP5I", + "title": "AJSN WK8A- AngryJoe Eye Surgery, HIGHGUARD Laysoff 80%, Riot Laysoff 2XKO Devs, Sony State of Play!", + "bucket": "medium", + "durationSeconds": 3380, + "cueCount": 1526 + }, + { + "slug": "rekietalaw/EsZhaCfc8HQ", + "channelSlug": "rekietalaw", + "videoId": "EsZhaCfc8HQ", + "videoDir": "EsZhaCfc8HQ", + "title": "ICE Agent Kills Woman In Minneapolis: Twitter Most Affected", + "bucket": "long", + "durationSeconds": 12497, + "cueCount": 3997 + }, + { + "slug": "chrissie-mayr/cfLF2o2-0BA", + "channelSlug": "chrissie-mayr", + "videoId": "cfLF2o2-0BA", + "videoDir": "cfLF2o2-0BA", + "title": "SimpCast 85! Chrissie Mayr, Wicked Virtue, Monica Paige, Brittany Venti, Lauren, - Barbie, Whatever", + "bucket": "long", + "durationSeconds": 12553, + "cueCount": 4692 + }, + { + "slug": "HasanAbiVODs/GjPX_ueTdfc", + "channelSlug": "HasanAbiVODs", + "videoId": "GjPX_ueTdfc", + "videoDir": "GjPX_ueTdfc", + "title": "2/2 HasanAbi April 7, 2021 - Sunburn OMEGALUL, Basketball Memes, OKBUDDY, 🎮GTA NoPixel🎮 FULL VOD", + "bucket": "verylong", + "durationSeconds": 28965, + "cueCount": 7223 + } + ] +}