Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 3598bf4dc79149beabca5a701dd115adeaf9c823
parent 483631920159b2cd8c3725ef8f07eb799f2b9afe
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sun, 26 Jul 2026 04:47:43 -0400

Digest layer: engine registry, artifact + override model, parser guards, batch

The generation spine for the derived-corpus backfill (PLAN.md Stage A) plus the
duplicate-cluster groundwork it depends on (Stage 0). No UI yet; no viewer
exposure at all.

Engines (common/lib/digestApps.ts) mirror the transcriptionApps registry:
ollama-direct on the local GPU with a regex-pinned JSON schema, and claude-code
as a metered lane that is OFF by default (settings.digest.remoteEnabled). The
schema pin is load-bearing, not polish: measured on qwen2.5:7b, pinning `start`
to ^\d\d:\d\d:\d\d$, stating each chunk's own time range, and demanding English
titles took malformed stamps, out-of-range starts and language drift from
"every request" to zero.

The artifact is two sidecars. ai-digest.json is machine-owned and freely
overwritten; ai-digest.overrides.json is human-owned and the generator can never
reach it. A full sweep is weeks of wall-clock, so a correction that regeneration
destroys is work that cannot be affordably redone — hence separate files rather
than a merge, with the id-keyed shadow idiom from searchAliases (later source
wins, enabled:false suppresses without deleting). Regeneration is skipped unless
(schemaVersion, appId, model, promptVersion, contextHash) differs, which is what
makes a prompt change cost minutes over a sample instead of a second sweep.
contextHash is plumbed now, while channel notes are still empty, so adding notes
later invalidates one channel rather than the corpus.

digestParse absorbs model sloppiness with five guards, each traceable to a
measured failure, and records every rejection in warnings[] instead of dropping
it. Verified on real data: a 2.3h video's third chunk re-emitted nine
whole-video-relative timestamps and the per-chunk clamp caught all nine — a
whole-video range check would have accepted them.

Duplicates get a canonical member (rule + human override in a sibling overrides
file, since the report is rewritten on every run) and a timestamp-alignment gate.
The gate is the correctness crux: content similarity says nothing about timing,
so a mirror with a longer intro matches on text at shifted times and a shared
digest would place every chapter wrong while looking fine. Clusters flagged
contained (a clip of a longer video) never share at all.

Two hazards a 119k-item sweep would otherwise hit are closed: the digest kinds
are in NO_REGEN_KINDS (recordTaskDone would otherwise arm a full per-channel
snapshot fan-out after every video) and the batch is one job per channel (logs
keep 500, the registry keeps 100). The lanes get separate queue keys because
registry.ts hardcodes concurrency 1 per key, so sharing one would serialize a
GPU-bound lane behind a network-bound one.

Also fixes the re-spelled JobProgressMetric/JobTaskKind literals that would not
have failed the build when the unions grew, and adds the `common` test script:
36 test files and 306 tests were previously unrunnable as a suite. 64 of those
tests are new here — the parser guards, the override merge, and the alignment
gate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Diffstat:
Mcommon/controller/channelSnapshot.ts | 43+++++++++++++++++++++++++++++++++++++++++++
Mcommon/controller/channels.ts | 29+++++++++++++++++++++++------
Mcommon/controller/checkAvailability.ts | 18++++--------------
Acommon/controller/digestBatch.ts | 446+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/controller/digestSharing.ts | 236+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/controller/digestVideo.ts | 321+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/controller/duplicateShorts.ts | 76++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++--
Mcommon/jobs/jobKinds.ts | 31+++++++++++++++++++++++++++++++
Mcommon/jobs/registry.ts | 7+++++--
Mcommon/jobs/snapshotScheduler.ts | 15++++++++++++++-
Mcommon/jobs/taskHooks.ts | 11++++++++++-
Acommon/lib/digest-server.ts | 298+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/digest.test.ts | 255+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/digest.ts | 355+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/digestApps.ts | 375+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/digestContext-server.ts | 63+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/digestParse.test.ts | 288+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/digestParse.ts | 325+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/digestPrompt.ts | 223+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/duplicates.test.ts | 273+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/lib/duplicates.ts | 337+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/parseStdoutJson.ts | 58++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/lib/paths.ts | 15+++++++++++++++
Mcommon/lib/queueKeys.ts | 11+++++++++++
Mcommon/lib/settings.ts | 139+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/lib/transcriptWindow.test.ts | 71++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-
Mcommon/lib/transcriptWindow.ts | 41+++++++++++++++++++++++++++++++++++++++++
Mcommon/package.json | 3+++
Aeditor/app/channels/[slug]/digestActions.ts | 249+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/app/channels/[slug]/lib/stageStatus.ts | 1+
Meditor/app/jobs/active/buildActiveJobs.ts | 9++++++++-
Meditor/app/jobs/components/RunningJobsList.tsx | 59+++++++++++++++++++++++++++++++++++++++++------------------
Meditor/app/jobs/jobReplayRegistry.ts | 29+++++++++++++++++++++++++++++
Meditor/app/settings/actions.ts | 47+++++++++++++++++++++++++++++++++++++++++++++++
Meditor/app/widget/components/MonitorWidget.tsx | 43+++++++++++++++++++++++++++++++------------
35 files changed, 4742 insertions(+), 58 deletions(-)

diff --git a/common/controller/channelSnapshot.ts b/common/controller/channelSnapshot.ts @@ -23,6 +23,7 @@ import { resolveEffectiveAvailability, } from "../lib/availability-server"; import { isDoNotClean } from "../lib/doNotClean-server"; +import { loadDigest } from "../lib/digest-server"; import { isExcludedFromTruncatedCheck } from "../lib/excludeTruncatedCheck-server"; import { loadDownloadOutcome } from "../lib/downloadOutcome-server"; import type { Paths } from "../lib/paths"; @@ -54,6 +55,13 @@ export type ChannelSnapshot = { transcribed: number; downloaded: number; }; + // Digest coverage by engine: appId -> count of videos whose ai-digest.json was + // produced by it. Beside `totals`, NOT in `buckets` — buckets are a closed + // literal of `string[]` id lists and a Record<string, number> does not belong + // there. During a multi-week sweep this split is what tells you whether the + // local lane is actually carrying the corpus. Optional: older snapshots lack + // it; readers default to {}. + digestEngines?: Record<string, number>; buckets: { noTranscript: string[]; downloadedNoTranscript: string[]; @@ -140,6 +148,12 @@ export type ChannelSnapshot = { // (retry-bucket with forceCookies). Optional: older snapshots lack it; // readers must default to []. needsCookies: string[]; + // Transcribed videos with no (or an outdated-shape) ai-digest.json — the AI + // digest layer's work list, and the denominator for corpus coverage during + // the backfill. Only TRANSCRIBED videos are listed: a video without a + // transcript is a transcription problem, not a digest one. Optional: older + // snapshots lack it; readers must default to []. + noDigest: string[]; }; undownloadedIds: string[]; excludedFromDownload?: ExcludedFromDownload; @@ -350,6 +364,13 @@ export async function generateChannelSnapshot( const vttProvenance = files.ytVttFile ? await resolveVttProvenance(dir, files.ytVttFile) : null; + // Only transcribed videos can carry a digest, so everything else skips + // the sidecar read entirely — the same conditional per-video + // sidecar-read pattern as the two reads above. + const digest = + isVideoTranscribed(files) && !files.isUntranscribable + ? await loadDigest(dir) + : null; return { id, files, @@ -362,6 +383,7 @@ export async function generateChannelSnapshot( outcome, coverage, vttProvenance, + digest, }; }), ), @@ -434,6 +456,8 @@ export async function generateChannelSnapshot( const autoSubsOnly: string[] = []; const downloadedAutoSubsOnly: string[] = []; const supersededAutoSubs: string[] = []; + const noDigest: string[] = []; + const digestEngines: Record<string, number> = {}; let transcribedWithAudioBytes = 0; let multipleAudioFormatsBytes = 0; let foreignAudioBytes = 0; @@ -446,6 +470,7 @@ export async function generateChannelSnapshot( outcome, coverage, vttProvenance, + digest, } of perVideo) { if (isVideoTranscribed(files)) transcribed++; if (isVideoDownloaded(files)) downloaded++; @@ -575,6 +600,22 @@ export async function generateChannelSnapshot( if (files.audioFiles.length > 0) downloadedAutoSubsOnly.push(id); else autoSubsOnly.push(id); } + // --- AI digest coverage ------------------------------------------------- + // Counted before the untranscribable/no-transcript `continue`s below so the + // accounting is unambiguous: a video is either a digest candidate or not a + // transcript at all. + if (isVideoTranscribed(files) && !files.isUntranscribable) { + const chapters = digest?.sections.chapters; + const tags = digest?.sections.tags; + const engine = chapters?.provenance.appId ?? tags?.provenance.appId; + const hasItems = + (chapters?.items.length ?? 0) > 0 || (tags?.items.length ?? 0) > 0; + if (engine && hasItems) { + digestEngines[engine] = (digestEngines[engine] ?? 0) + 1; + } else { + noDigest.push(id); + } + } if (files.isUntranscribable) { untranscribable.push(id); continue; @@ -681,7 +722,9 @@ export async function generateChannelSnapshot( downloadedAutoSubsOnly: downloadedAutoSubsOnly.sort(), supersededAutoSubs: supersededAutoSubs.sort(), needsCookies: needsCookies.sort(), + noDigest: noDigest.sort(), }, + digestEngines, undownloadedIds, excludedFromDownload, keptCount: keptIds.size, diff --git a/common/controller/channels.ts b/common/controller/channels.ts @@ -11,6 +11,7 @@ import { isVideoTranscribed, readVideoFiles, } from "../lib/videoStatus"; +import { hasDigest } from "../lib/digest-server"; export type ChannelStat = { slug: string; @@ -19,6 +20,11 @@ export type ChannelStat = { videoCount: number; transcriptCount: number; downloadCount: number; + // Videos carrying a non-empty ai-digest.json. The Active Jobs progress bar + // re-counts `current` from disk rather than trusting the runner, so a digest + // job needs this counter to have a bar at all. Optional so a caller reading an + // older serialized stat still type-checks. + digestCount?: number; }; // A channel slug is also its directory name under transcripts/channels/, so it @@ -43,32 +49,42 @@ async function exists(p: string): Promise<boolean> { } } -async function countDataFiles( - dataDir: string, -): Promise<{ videos: number; transcripts: number; downloads: number }> { +async function countDataFiles(dataDir: string): Promise<{ + videos: number; + transcripts: number; + downloads: number; + digests: number; +}> { let dirs: Dirent[]; try { dirs = await readdir(dataDir, { withFileTypes: true }); } catch { - return { videos: 0, transcripts: 0, downloads: 0 }; + return { videos: 0, transcripts: 0, downloads: 0, digests: 0 }; } const videoDirs = dirs.filter((d) => d.isDirectory()); const flags = await Promise.all( videoDirs.map(async (d) => { - const files = await readVideoFiles(path.join(dataDir, d.name)); + const dir = path.join(dataDir, d.name); + const files = await readVideoFiles(dir); return { transcript: isVideoTranscribed(files), download: isVideoDownloaded(files), + // Only transcribed videos can carry a digest, so the sidecar read is + // skipped for the rest — the same conditional per-video sidecar-read + // pattern channelSnapshot.ts uses for coverage and VTT provenance. + digest: isVideoTranscribed(files) ? await hasDigest(dir) : false, }; }), ); let transcripts = 0; let downloads = 0; + let digests = 0; for (const f of flags) { if (f.transcript) transcripts++; if (f.download) downloads++; + if (f.digest) digests++; } - return { videos: videoDirs.length, transcripts, downloads }; + return { videos: videoDirs.length, transcripts, downloads, digests }; } async function countPlaylist(p: string): Promise<number | null> { @@ -122,6 +138,7 @@ export async function readChannelStat( videoCount: counts.videos, transcriptCount: counts.transcripts, downloadCount: counts.downloads, + digestCount: counts.digests, }; } diff --git a/common/controller/checkAvailability.ts b/common/controller/checkAvailability.ts @@ -20,6 +20,10 @@ import { resolveCookiePolicy, } from "../lib/cookiePolicy"; import { getSettings } from "../lib/settings"; +// yt-dlp --dump-json emits one JSON object per video, but may print warnings to +// stdout first — so the payload is the LAST parseable line. Shared with the +// digest registry's claude-code lane, which has the same problem. +import { parseStdoutJson } from "../lib/parseStdoutJson"; import { readChannelConfig } from "./channels"; import { resolveShardItems } from "./shard"; import type { Paths } from "../lib/paths"; @@ -88,20 +92,6 @@ async function readWebpageUrl(videoDir: string): Promise<string | null> { } } -function parseStdoutJson(stdout: string): unknown | null { - // yt-dlp --dump-json emits one JSON object per video; for a single URL we - // expect one line. If something else printed warnings to stdout, the JSON - // is the last non-empty line. - const lines = stdout.split("\n").map((l) => l.trim()).filter(Boolean); - for (let i = lines.length - 1; i >= 0; i--) { - try { - return JSON.parse(lines[i]); - } catch { - continue; - } - } - return null; -} export async function runAvailabilityCheck({ channelSlug, diff --git a/common/controller/digestBatch.ts b/common/controller/digestBatch.ts @@ -0,0 +1,446 @@ +// Channel-scoped digest sweep: the thing that actually runs for weeks. +// +// Three mechanics here are copied from existing code for specific reasons, and +// changing any of them breaks a property a multi-week sweep depends on: +// +// 1. RESUME BY RE-DERIVING FROM DISK. Eligibility is re-checked against the +// sidecar on every next() pull (the auto-runner's shape), NOT frozen into an +// array with an index cursor (whisperBatch's shape). The cursor form does not +// survive a restart, and it cannot see a video that became eligible mid-run +// (a transcript that finished, a mirror that got shared to). +// 2. PAUSE BY RETURNING limit() === 0. runPool idle-WAITS at a zero limit +// rather than finishing (concurrentRunner.ts's documented invariant), so a +// pause holds the job open instead of ending it. The flag is re-read from +// settings on every pull — the downloadsPaused pattern, which needs no boot +// hook, unlike transcriptionsPaused. +// 3. ONE JOB PER CHANNEL, never one per video. Job logs keep the newest 500 / +// 30 days and the in-memory registry keeps 100 records; 119,600 jobs would +// evict everything, including the running ones' history. + +import path from "node:path"; +import { open, readdir } from "node:fs/promises"; +import type { Paths } from "../lib/paths"; +import { getSettings } from "../lib/settings"; +import { runPool } from "../jobs/concurrentRunner"; +import type { TaskTracker } from "../jobs/taskHooks"; +import type { JobProgress } from "../jobs/registry"; +import { + getDigestApp, + type DigestAppConfig, +} from "../lib/digestApps"; +import { + isSectionFresh, + type DigestSectionKind, +} from "../lib/digest"; +import { loadDigest } from "../lib/digest-server"; +import { PROMPT_VERSION } from "../lib/digestPrompt"; +import { readDigestContext } from "../lib/digestContext-server"; +import { CUES_JSON_FILENAME } from "../lib/videoStatus"; +import { digestVideo } from "./digestVideo"; +import { + buildDigestClusterPlan, + shareDigestToCluster, + type DigestClusterPlan, +} from "./digestSharing"; + +export type DigestOrder = "shortest-first" | "longest-first"; + +export type DigestBatchOptions = { + channelSlug: string; + paths: Paths; + // Which lane to run. "local" is the default and the only one enabled unless + // settings.digest.remoteEnabled is true; asking for "remote" while it is off is + // an error the caller surfaces, not a silent downgrade. + lane?: "local" | "remote"; + // When set, only consider these video ids (intersected with what's on disk). + ids?: string[]; + sections?: DigestSectionKind[]; + // The repo only had a `reverse` flag before this; digest ordering is explicit + // because lane routing and the pilot both depend on it. Shortest-first by + // default: it converts the backlog into visible coverage fastest, and the + // long-tail 8% is where a prompt bug is most expensive to discover late. + order?: DigestOrder; + // Duration window, for splitting the corpus between lanes (e.g. local takes + // ≤ longTailSeconds, the metered lane takes the tail). + minDurationSeconds?: number; + maxDurationSeconds?: number; + // Stop after this many successful videos. For the pilot and Stage B's sample. + limitCount?: number; + concurrency?: number; + force?: boolean; + // Skip mirrors and share the canonical member's digest to aligned ones. On by + // default: worth ~11% of the sweep. Pass a prebuilt plan to avoid re-reading + // the duplicates report. + useClusters?: boolean; + clusterPlan?: DigestClusterPlan; + setProgress?: (snap: JobProgress) => void; + progressBaseline?: number; + onLog?: (msg: string) => void; + signal?: AbortSignal; + drainSignal?: AbortSignal; + tracker?: TaskTracker; +}; + +export type DigestBatchResult = { + attempted: number; + succeeded: number; + skipped: number; + failed: number; + fresh: number; + // Cluster mirrors that received the canonical member's digest, and mirrors the + // alignment gate refused. A refusal is a normal outcome, not an error. + shared: number; + misaligned: number; + engineCalls: number; + costUsd: number; + warnings: number; + // True when the run stopped early because the spend cap was reached. + spendCapped: boolean; +}; + +type Candidate = { id: string; duration: number }; + +// Read a video's duration cheaply. transcript.cues.json is +// {version, source, transcriptFormat, ...summary, cues} — the summary (and so +// `duration`) is serialized BEFORE the multi-megabyte cues array, so the head of +// the file is enough and a full parse is only the fallback. At 74k videos this is +// the difference between a few seconds and reading 6.9 GB to sort a list. +const DURATION_HEAD_BYTES = 8192; + +async function readDurationFast(cuesPath: string): Promise<number | null> { + let handle: Awaited<ReturnType<typeof open>> | null = null; + try { + handle = await open(cuesPath, "r"); + const buf = Buffer.alloc(DURATION_HEAD_BYTES); + const { bytesRead } = await handle.read(buf, 0, DURATION_HEAD_BYTES, 0); + const head = buf.subarray(0, bytesRead).toString("utf8"); + const m = head.match(/"duration"\s*:\s*([0-9]+(?:\.[0-9]+)?)/); + if (m) return Number(m[1]); + // Short file: the whole thing is in `head`, so parse it properly. + if (bytesRead < DURATION_HEAD_BYTES) { + const parsed = JSON.parse(head) as { duration?: unknown }; + return typeof parsed.duration === "number" ? parsed.duration : null; + } + return null; + } catch { + return null; + } finally { + await handle?.close().catch(() => {}); + } +} + +export async function runDigestBatch( + opts: DigestBatchOptions, +): Promise<DigestBatchResult> { + const log = opts.onLog ?? ((m: string) => console.log(m)); + const settings = getSettings(); + const digestSettings = settings.digest; + const lane = opts.lane ?? "local"; + if (lane === "remote" && !digestSettings.remoteEnabled) { + throw new Error( + "The metered digest lane is disabled (settings.digest.remoteEnabled). Enable it in Settings before running it.", + ); + } + const appId = + lane === "remote" ? digestSettings.remoteAppId : digestSettings.localAppId; + const app = getDigestApp(appId); + const config: DigestAppConfig = digestSettings.apps[app.id] ?? {}; + const sections = opts.sections ?? digestSettings.sections; + const modelRequested = config.model?.trim() || app.defaultModel(); + + const dataDir = path.join(opts.paths.channelsDir, opts.channelSlug, "data"); + const context = await readDigestContext(opts.paths, opts.channelSlug); + + // Fail fast and loudly rather than 1,100 times in a row: an unreachable engine + // is a configuration problem, and discovering it per-item wastes the log. + if (!(await app.probe(config))) { + throw new Error( + `Digest engine ${app.id} is not reachable. ` + + (app.lane === "local-gpu" + ? "Is the ollama service running (systemctl status ollama)?" + : "Is the claude CLI installed and on PATH (set CLAUDE_BIN)?"), + ); + } + + const allDirs = await readdir(dataDir).catch(() => [] as string[]); + const onDisk = new Set(allDirs); + const wanted = opts.ids + ? opts.ids.filter((id) => onDisk.has(id)) + : allDirs; + + const clusterPlan = + opts.useClusters === false + ? null + : (opts.clusterPlan ?? (await buildDigestClusterPlan(opts.paths))); + + // Durations, read once. Videos with no normalized transcript have no duration + // and are dropped here — they are a transcription problem, not a digest one. + const candidates: Candidate[] = []; + let noTranscript = 0; + let outOfWindow = 0; + for (const id of wanted) { + opts.signal?.throwIfAborted(); + const duration = await readDurationFast( + path.join(dataDir, id, CUES_JSON_FILENAME), + ); + if (duration === null || duration <= 0) { + noTranscript++; + continue; + } + if ( + opts.minDurationSeconds !== undefined && + duration < opts.minDurationSeconds + ) { + outOfWindow++; + continue; + } + if ( + opts.maxDurationSeconds !== undefined && + duration > opts.maxDurationSeconds + ) { + outOfWindow++; + continue; + } + candidates.push({ id, duration }); + } + + const order = opts.order ?? "shortest-first"; + candidates.sort((a, b) => + order === "longest-first" + ? b.duration - a.duration || a.id.localeCompare(b.id) + : a.duration - b.duration || a.id.localeCompare(b.id), + ); + + log( + `Digest ${opts.channelSlug} (${lane} lane, ${app.id}/${modelRequested}, sections: ${sections.join(", ")}): ` + + `${candidates.length} candidate(s) of ${wanted.length} on disk ` + + `(${noTranscript} without a transcript, ${outOfWindow} outside the duration window), ${order}.` + + (clusterPlan + ? ` Duplicate plan: ${clusterPlan.clusters} cluster(s), ${clusterPlan.bySlug.size} member(s) mapped.` + : " Duplicate sharing off."), + ); + + const result: DigestBatchResult = { + attempted: 0, + succeeded: 0, + skipped: 0, + failed: 0, + fresh: 0, + shared: 0, + misaligned: 0, + engineCalls: 0, + costUsd: 0, + warnings: 0, + spendCapped: false, + }; + + const freshnessTarget = { + appId: app.id, + model: modelRequested, + promptVersion: PROMPT_VERSION, + contextHash: context.hash, + }; + + // Ids already handed out this run. Combined with the cursor below, this is what + // makes the disk re-derivation O(n) overall rather than O(n²): the cursor only + // ever moves forward past ids that have been attempted, while eligibility for + // the id it stops on is always re-read from disk. + const attempted = new Set<string>(); + let cursor = 0; + + const slugOf = (id: string): string => `${opts.channelSlug}/${id}`; + + const next = async (): Promise<Candidate | null> => { + while (cursor < candidates.length) { + if (opts.limitCount !== undefined && result.succeeded >= opts.limitCount) { + return null; + } + const candidate = candidates[cursor]; + if (attempted.has(candidate.id)) { + cursor++; + continue; + } + const videoDir = path.join(dataDir, candidate.id); + + // A cluster mirror is not this lane's work: its canonical member owns the + // generation and shares the result here. + const role = clusterPlan?.bySlug.get(slugOf(candidate.id)); + if (role?.kind === "mirror") { + attempted.add(candidate.id); + cursor++; + result.skipped++; + log( + `Skipping ${candidate.id}: duplicate of ${role.canonicalSlug}, which owns the digest for this cluster.`, + ); + continue; + } + + // RE-DERIVED FROM DISK, every pull. A restart, a concurrent lane, or a + // share that landed while this job ran are all visible here. + if (!opts.force) { + const record = await loadDigest(videoDir); + const allFresh = sections.every((section) => + isSectionFresh(record, section, freshnessTarget), + ); + if (allFresh) { + attempted.add(candidate.id); + cursor++; + result.fresh++; + continue; + } + } + attempted.add(candidate.id); + cursor++; + return candidate; + } + return null; + }; + + const runOne = async ( + candidate: Candidate, + runSignal: AbortSignal, + ): Promise<void> => { + const task = opts.tracker?.start({ + id: candidate.id, + label: `digest ${opts.channelSlug}/${candidate.id}`, + kind: "digest", + }); + try { + const outcome = await digestVideo({ + paths: opts.paths, + channelSlug: opts.channelSlug, + videoId: candidate.id, + sections, + appId: app.id, + config, + context, + force: opts.force, + onLog: task ? task.onLog : opts.onLog, + signal: runSignal, + }); + if (outcome.status === "fresh") { + result.fresh++; + return; + } + if (outcome.status === "skipped") { + result.skipped++; + log(`Skipped ${candidate.id}: ${outcome.reason}.`); + return; + } + result.attempted++; + result.succeeded++; + result.engineCalls += outcome.engineCalls; + result.costUsd += outcome.costUsd; + result.warnings += outcome.warningCount; + + // Share to this cluster's aligned mirrors, right after the canonical + // member's digest lands — so a mirror never sits un-digested waiting for a + // second pass, and a crash mid-sweep leaves a consistent cluster. + const role = clusterPlan?.bySlug.get(slugOf(candidate.id)); + if (role?.kind === "canonical") { + const outcomes = await shareDigestToCluster({ + paths: opts.paths, + clusterId: role.clusterId, + canonicalSlug: slugOf(candidate.id), + mirrors: role.mirrors, + onLog: log, + }); + for (const o of outcomes) { + if (o.status === "shared") result.shared++; + else if (o.status === "misaligned") result.misaligned++; + } + } + } catch (err) { + if (runSignal.aborted || opts.signal?.aborted) throw err; + result.attempted++; + result.failed++; + log( + `Failed ${candidate.id}: ${(err as Error)?.message ?? String(err)}`, + ); + } finally { + task?.end(); + } + }; + + const NEVER = new AbortController().signal; + // Local lane: 1. The GPU is the bottleneck and a second concurrent generation + // just thrashes the same 8 GB of VRAM. The metered lane is network-bound, so it + // can overlap — but modestly, since it is paying per call. + const defaultConcurrency = app.lane === "local-gpu" ? 1 : 2; + const concurrency = Math.max(1, opts.concurrency ?? defaultConcurrency); + + await runPool<Candidate>({ + next, + run: runOne, + limit: () => { + // Re-read at DISPATCH time so a pause takes effect within one poll and + // survives a restart with no boot hook. Returning 0 makes runPool + // idle-wait, which is a pause; returning null from next() would END the + // batch, which is not. + if (getSettings().digest.digestsPaused) return 0; + if ( + app.metered && + digestSettings.spendCapUsd > 0 && + result.costUsd >= digestSettings.spendCapUsd + ) { + if (!result.spendCapped) { + result.spendCapped = true; + log( + `Spend cap reached ($${result.costUsd.toFixed(2)} of $${digestSettings.spendCapUsd.toFixed(2)}) — parking the metered lane.`, + ); + } + return 0; + } + return concurrency; + }, + signal: opts.signal ?? NEVER, + drainSignal: opts.drainSignal ?? NEVER, + finite: true, + idlePollMs: 3000, + }); + + // Metered accounting is logged unconditionally when the lane is metered, even + // at zero calls: "this run cost nothing" is information too. + if (app.metered) { + log( + `Metered lane: ${result.engineCalls} model call(s), $${result.costUsd.toFixed(4)} total` + + (digestSettings.spendCapUsd > 0 + ? ` (cap $${digestSettings.spendCapUsd.toFixed(2)})` + : " (no cap set)"), + ); + } + return result; +} + +// How many of a channel's videos still need a digest at the CURRENT prompt/model +// identity. Used to scope a job's progress bar, and cheap enough to call before +// starting one (a small sidecar read per video, no transcript parsing). +export async function countMissingDigests( + paths: Paths, + channelSlug: string, + ids?: ReadonlyArray<string>, +): Promise<number> { + const settings = getSettings().digest; + const app = getDigestApp(settings.localAppId); + const config = settings.apps[app.id] ?? {}; + const context = await readDigestContext(paths, channelSlug); + const target = { + appId: app.id, + model: config.model?.trim() || app.defaultModel(), + promptVersion: PROMPT_VERSION, + contextHash: context.hash, + }; + const dataDir = path.join(paths.channelsDir, channelSlug, "data"); + const dirs = ids + ? [...ids] + : await readdir(dataDir).catch(() => [] as string[]); + let missing = 0; + for (const id of dirs) { + const record = await loadDigest(path.join(dataDir, id)); + const fresh = settings.sections.every((section) => + isSectionFresh(record, section, target), + ); + if (!fresh) missing++; + } + return missing; +} diff --git a/common/controller/digestSharing.ts b/common/controller/digestSharing.ts @@ -0,0 +1,236 @@ +// Where duplicate detection meets the digest layer: generate ONCE per cluster, +// share to the mirrors that are safe to share to. +// +// This is worth ~11% of the sweep on its own — 8,204 exact-title redundancies +// across the corpus, 5,863 of them Quartering YouTube↔Rumble mirror pairs — and +// at weeks of wall-clock per pass, not re-generating those is real time. +// +// Two rules keep the sharing honest, and both are the difference between a +// correct optimisation and a plausible-looking wrong one: +// +// 1. A `contained` cluster NEVER shares. Containment means one member is a CLIP +// of a longer video; the longer video's chapters describe material the clip +// does not contain. +// 2. A mirror only receives a digest when its cue TIMINGS align with the +// canonical member's. Content similarity says nothing about timing — a +// mirror with a longer intro matches on text at shifted times — so a shared +// digest would place every chapter wrong while looking perfectly fine. The +// gate measures the offset at several anchors and requires near-zero. + +import path from "node:path"; +import type { Paths } from "../lib/paths"; +import { + DEFAULT_ALIGNMENT_TOLERANCE_SECONDS, + clusterMaySharePartial, + measureAlignment, + resolveCanonicalSlug, + type AlignmentResult, + type DuplicateCluster, + type DuplicateOverrides, +} from "../lib/duplicates"; +import { CUES_JSON_FILENAME } from "../lib/videoStatus"; +import type { Cue } from "../lib/vtt"; +import { loadDigest, writeSharedDigest } from "../lib/digest-server"; +import type { DigestRecord } from "../lib/digest"; +import { readNormalizedTranscript } from "./normalizeTranscript"; +import { readDuplicateOverrides, readDuplicateReport } from "./duplicateShorts"; + +// What a video's cluster membership means for the batch. +export type DigestClusterRole = + // Generate for this video: it owns its cluster's derived work. + | { kind: "canonical"; clusterId: string; mirrors: string[] } + // Do NOT generate: the canonical member owns it and will share it here. + | { kind: "mirror"; clusterId: string; canonicalSlug: string }; + +export type DigestClusterPlan = { + // `${channelSlug}/${id}` → role. Videos absent from the map are in no cluster + // and are generated normally, which is the overwhelming majority. + bySlug: Map<string, DigestClusterRole>; + // Clusters that exist but share nothing (contained, or marked not-a-duplicate). + // Their members are absent from bySlug and are each generated independently — + // a clip is a different artifact and deserves its own digest. + independentClusters: number; + clusters: number; +}; + +export function emptyDigestClusterPlan(): DigestClusterPlan { + return { bySlug: new Map(), independentClusters: 0, clusters: 0 }; +} + +// Build the plan from the last detection run. Absent report → an empty plan, so +// the digest batch works fine before duplicates have ever been detected (it just +// generates for mirrors twice, which is correct, only slower). +export async function buildDigestClusterPlan( + paths: Paths, + opts: { overrides?: DuplicateOverrides | null } = {}, +): Promise<DigestClusterPlan> { + const report = await readDuplicateReport(paths); + if (!report) return emptyDigestClusterPlan(); + const overrides = opts.overrides ?? (await readDuplicateOverrides(paths)); + const plan = emptyDigestClusterPlan(); + plan.clusters = report.clusters.length; + + for (const cluster of report.clusters) { + const canonicalSlug = resolveCanonicalSlug(cluster, overrides); + // null → a human said "not a duplicate": every member stands alone. + if (!canonicalSlug || !clusterMaySharePartial(cluster)) { + plan.independentClusters++; + continue; + } + const mirrors = cluster.videoRefs + .map((r) => r.slug) + .filter((slug) => slug !== canonicalSlug); + if (mirrors.length === 0) { + plan.independentClusters++; + continue; + } + plan.bySlug.set(canonicalSlug, { + kind: "canonical", + clusterId: cluster.clusterId, + mirrors, + }); + for (const slug of mirrors) { + plan.bySlug.set(slug, { + kind: "mirror", + clusterId: cluster.clusterId, + canonicalSlug, + }); + } + } + return plan; +} + +function videoDirForSlug(paths: Paths, slug: string): string | null { + const at = slug.indexOf("/"); + if (at <= 0) return null; + return path.join( + paths.channelsDir, + slug.slice(0, at), + "data", + slug.slice(at + 1), + ); +} + +async function readCues(paths: Paths, slug: string): Promise<Cue[] | null> { + const dir = videoDirForSlug(paths, slug); + if (!dir) return null; + const t = await readNormalizedTranscript(path.join(dir, CUES_JSON_FILENAME)); + return t?.cues ?? null; +} + +export type ShareOutcome = { + slug: string; + status: "shared" | "misaligned" | "no-transcript" | "already-shared" | "failed"; + alignment?: AlignmentResult; + error?: string; +}; + +export type ShareDigestOptions = { + paths: Paths; + clusterId: string; + canonicalSlug: string; + mirrors: string[]; + toleranceSeconds?: number; + onLog?: (msg: string) => void; +}; + +// Copy the canonical member's digest onto each mirror that passes the alignment +// gate. Returns one outcome per mirror — a misaligned mirror is a normal, +// expected result, not an error, and it simply stays on the generation worklist. +export async function shareDigestToCluster( + opts: ShareDigestOptions, +): Promise<ShareOutcome[]> { + const log = opts.onLog ?? (() => {}); + const tolerance = + opts.toleranceSeconds ?? DEFAULT_ALIGNMENT_TOLERANCE_SECONDS; + const canonicalDir = videoDirForSlug(opts.paths, opts.canonicalSlug); + if (!canonicalDir) return []; + const source: DigestRecord | null = await loadDigest(canonicalDir); + if (!source) return []; + const canonicalCues = await readCues(opts.paths, opts.canonicalSlug); + if (!canonicalCues || canonicalCues.length === 0) return []; + + const sharedAt = new Date().toISOString(); + const outcomes: ShareOutcome[] = []; + for (const slug of opts.mirrors) { + const dir = videoDirForSlug(opts.paths, slug); + if (!dir) { + outcomes.push({ slug, status: "failed", error: "unparseable slug" }); + continue; + } + try { + const existing = await loadDigest(dir); + // Already carrying this canonical member's digest at the same provenance: + // nothing to do. Keeps a re-run a genuine no-op. + if ( + existing?.derivedFrom?.slug === opts.canonicalSlug && + existing.promptVersion === source.promptVersion && + existing.contextHash === source.contextHash + ) { + outcomes.push({ slug, status: "already-shared" }); + continue; + } + const cues = await readCues(opts.paths, slug); + if (!cues || cues.length === 0) { + outcomes.push({ slug, status: "no-transcript" }); + continue; + } + const alignment = measureAlignment(canonicalCues, cues, { + toleranceSeconds: tolerance, + }); + if (!alignment.aligned) { + log( + `Not sharing to ${slug}: ${alignment.reason} (max offset ${ + Number.isFinite(alignment.maxOffsetSeconds) + ? `${alignment.maxOffsetSeconds.toFixed(1)}s` + : "n/a" + }, ${alignment.matchedAnchors}/${alignment.totalAnchors} anchors matched).`, + ); + outcomes.push({ slug, status: "misaligned", alignment }); + continue; + } + await writeSharedDigest(dir, source, { + slug: opts.canonicalSlug, + clusterId: opts.clusterId, + sharedAt, + offsetSeconds: Math.round(alignment.maxOffsetSeconds * 100) / 100, + }); + log( + `Shared digest ${opts.canonicalSlug} → ${slug} (max offset ${alignment.maxOffsetSeconds.toFixed(2)}s).`, + ); + outcomes.push({ slug, status: "shared", alignment }); + } catch (err) { + outcomes.push({ + slug, + status: "failed", + error: (err as Error)?.message ?? String(err), + }); + } + } + return outcomes; +} + +// Convenience for the /actionable per-cluster action: share from whatever the +// effective canonical member currently is. +export async function shareClusterFromCanonical( + paths: Paths, + cluster: DuplicateCluster, + opts: { + overrides?: DuplicateOverrides | null; + onLog?: (msg: string) => void; + } = {}, +): Promise<ShareOutcome[]> { + if (!clusterMaySharePartial(cluster)) return []; + const overrides = opts.overrides ?? (await readDuplicateOverrides(paths)); + const canonicalSlug = resolveCanonicalSlug(cluster, overrides); + if (!canonicalSlug) return []; + return shareDigestToCluster({ + paths, + clusterId: cluster.clusterId, + canonicalSlug, + mirrors: cluster.videoRefs + .map((r) => r.slug) + .filter((slug) => slug !== canonicalSlug), + onLog: opts.onLog, + }); +} diff --git a/common/controller/digestVideo.ts b/common/controller/digestVideo.ts @@ -0,0 +1,321 @@ +// Generate the AI digest for ONE video: chapters and/or topic tags, written to +// the ai-digest.json sidecar next to the transcript. +// +// Pipeline (each step reusing the declared single source of truth for its job): +// isCuesJsonFresh — don't digest a transcript that's about to change +// readNormalizedTranscript +// chunkCuesForContext — sequential overlapping slices +// transcriptToMarkdown — cues -> text for an AI, with hms() stamps +// digest app — ollama (local) or claude CLI (metered, opt-in) +// parseChapters/parseTags — the guards; every rejection recorded +// writeDigestSection — read-modify-write ONE section, atomic rename +// +// THE FRESHNESS SKIP IS THE POINT. A full sweep of the 74k-transcript corpus is +// weeks of wall-clock on one local lane, so a redo is unaffordable. A section is +// regenerated only when its recorded (schemaVersion, appId, model, promptVersion, +// contextHash) differs from what we would produce now — which makes a re-run +// after a prompt change minutes long over a sample instead of weeks over the +// corpus. + +import path from "node:path"; +import type { Paths } from "../lib/paths"; +import { + DIGEST_MAX_CUES_PER_CHUNK, + DIGEST_OVERLAP_CUES, + PROMPT_VERSION, + CHAPTER_SYSTEM_PROMPT, + TAG_SYSTEM_PROMPT, + buildChapterPrompt, + buildTagPrompt, + chapterSchema, + tagSchema, + toHms, +} from "../lib/digestPrompt"; +import { getDigestApp, type DigestAppConfig } from "../lib/digestApps"; +import { + isSectionFresh, + type DigestItem, + type DigestProvenance, + type DigestRecord, + type DigestSectionKind, + type DigestWarning, +} from "../lib/digest"; +import { loadDigest, writeDigestSection } from "../lib/digest-server"; +import { readDigestContext, type DigestContext } from "../lib/digestContext-server"; +import { parseChapters, parseTags, type DigestChunkOutput } from "../lib/digestParse"; +import { chunkCuesForContext } from "../lib/transcriptWindow"; +import { transcriptToMarkdown } from "../lib/transcriptToMarkdown"; +import type { Cue } from "../lib/vtt"; +import { + isCuesJsonFresh, + readNormalizedTranscript, +} from "./normalizeTranscript"; + +export type DigestVideoOptions = { + paths: Paths; + channelSlug: string; + videoId: string; + // Which sections to generate. Defaults to chapters only — tags are cheap but + // double the call count, so the sweep opts into them explicitly. + sections?: DigestSectionKind[]; + appId?: string; + config?: DigestAppConfig; + // Pre-read channel context, so a batch reads it once per channel instead of + // once per video. Omitted → read here. + context?: DigestContext; + // Regenerate even when the recorded provenance matches. For Stage B iteration + // on a fixed sample; never set for a sweep. + force?: boolean; + onLog?: (msg: string) => void; + signal?: AbortSignal; +}; + +export type DigestVideoOutcome = + | { + status: "wrote"; + record: DigestRecord; + sections: DigestSectionKind[]; + itemCount: number; + warningCount: number; + costUsd: number; + engineCalls: number; + } + | { status: "fresh" } + | { + status: "skipped"; + reason: "no-transcript" | "stale-cues" | "no-cues" | "no-metadata"; + }; + +// A per-chunk engine failure must not lose the chunks that DID work: a 12-hour +// video is 30+ calls and one 500 from ollama should cost that chunk, not the +// video. The failure is recorded as a warning and the section is written from +// what survived, with chunks/chunksOk in provenance so a partially-generated +// section is identifiable later. +export async function digestVideo( + opts: DigestVideoOptions, +): Promise<DigestVideoOutcome> { + const log = opts.onLog ?? (() => {}); + const videoDir = path.join( + opts.paths.channelsDir, + opts.channelSlug, + "data", + opts.videoId, + ); + const sections = opts.sections ?? ["chapters"]; + + // Don't digest a transcript that is about to be rewritten: a stale cues.json + // means the raw transcript changed under it, so the digest would describe + // superseded text and then look "fresh" forever. + const { fresh: cuesFresh, cuesPath } = await isCuesJsonFresh(videoDir); + if (!cuesFresh) { + const existing = await readNormalizedTranscript(cuesPath); + return { + status: "skipped", + reason: existing ? "stale-cues" : "no-transcript", + }; + } + const transcript = await readNormalizedTranscript(cuesPath); + if (!transcript) return { status: "skipped", reason: "no-transcript" }; + const cues = transcript.cues ?? []; + if (cues.length === 0) return { status: "skipped", reason: "no-cues" }; + + const app = getDigestApp(opts.appId); + const config = opts.config ?? {}; + const modelRequested = config.model?.trim() || app.defaultModel(); + const context = + opts.context ?? (await readDigestContext(opts.paths, opts.channelSlug)); + + const existing = await loadDigest(videoDir); + const target = { + appId: app.id, + model: modelRequested, + promptVersion: PROMPT_VERSION, + contextHash: context.hash, + }; + const stale = sections.filter( + (section) => opts.force || !isSectionFresh(existing, section, target), + ); + if (stale.length === 0) return { status: "fresh" }; + + const chunks = chunkCuesForContext(cues, { + maxCues: DIGEST_MAX_CUES_PER_CHUNK, + overlapCues: DIGEST_OVERLAP_CUES, + }); + log( + `${opts.channelSlug}/${opts.videoId}: ${cues.length} cues → ${chunks.length} chunk(s), sections: ${stale.join(", ")}`, + ); + + let record: DigestRecord | null = existing; + let itemCount = 0; + let warningCount = 0; + let costUsd = 0; + let engineCalls = 0; + + for (const section of stale) { + const outputs: DigestChunkOutput[] = []; + const runWarnings: DigestWarning[] = []; + let reportedModel = modelRequested; + let sectionCost = 0; + + for (let i = 0; i < chunks.length; i++) { + opts.signal?.throwIfAborted(); + const chunk = chunks[i]; + const startSeconds = Math.max(0, Math.floor(chunk[0].start)); + const endSeconds = Math.max( + startSeconds, + Math.ceil(chunk[chunk.length - 1].end || chunk[chunk.length - 1].start), + ); + const promptInput = { + title: transcript.title || opts.videoId, + channel: transcript.channel || opts.channelSlug, + startSeconds, + endSeconds, + transcript: renderChunk(transcript, chunk), + ...(context.note ? { contextNote: context.note } : {}), + }; + const span = endSeconds - startSeconds; + const request = + section === "chapters" + ? { + system: CHAPTER_SYSTEM_PROMPT, + prompt: buildChapterPrompt(promptInput), + schema: chapterSchema(span), + } + : { + system: TAG_SYSTEM_PROMPT, + prompt: buildTagPrompt(promptInput), + schema: tagSchema(), + }; + try { + const result = await app.run({ + ...request, + config, + signal: opts.signal, + onLog: opts.onLog, + }); + engineCalls++; + reportedModel = result.model || reportedModel; + if (typeof result.costUsd === "number") sectionCost += result.costUsd; + outputs.push({ + index: i, + startSeconds, + endSeconds, + data: result.data, + }); + } catch (err) { + // A cancel is not a chunk failure — let it propagate so the batch stops. + if (opts.signal?.aborted) throw err; + const message = (err as Error)?.message ?? String(err); + runWarnings.push({ + code: "chunk-failed", + section, + chunk: i, + detail: message.slice(0, 300), + }); + log( + `${opts.channelSlug}/${opts.videoId}: chunk ${i + 1}/${chunks.length} failed: ${message}`, + ); + } + } + + if (outputs.length === 0) { + // Nothing usable. Do NOT write a section — an empty section with current + // provenance would read as "fresh" and the video would never be retried. + log( + `${opts.channelSlug}/${opts.videoId}: ${section} produced no usable output; leaving the sidecar untouched so it retries.`, + ); + warningCount += runWarnings.length; + continue; + } + + const parsed = + section === "chapters" + ? parseChapters(outputs, cues) + : parseTags(outputs); + const items: DigestItem[] = + "chapters" in parsed ? parsed.chapters : parsed.tags; + const warnings = [...runWarnings, ...parsed.warnings]; + + if (items.length === 0) { + log( + `${opts.channelSlug}/${opts.videoId}: every ${section} entry was rejected (${warnings.length} warning(s)); leaving the sidecar untouched so it retries.`, + ); + warningCount += warnings.length; + continue; + } + + const provenance: DigestProvenance = { + appId: app.id, + model: reportedModel, + modelRequested, + lane: app.lane, + generatedAt: new Date().toISOString(), + promptVersion: PROMPT_VERSION, + contextHash: context.hash, + chunks: chunks.length, + chunksOk: outputs.length, + ...(app.metered && sectionCost > 0 ? { costUsd: sectionCost } : {}), + }; + record = await writeDigestSection(videoDir, { + section, + items, + provenance, + warnings, + }); + itemCount += items.length; + warningCount += warnings.length; + costUsd += sectionCost; + log( + `${opts.channelSlug}/${opts.videoId}: ${section} → ${items.length} item(s), ${warnings.length} warning(s)` + + (sectionCost > 0 ? `, $${sectionCost.toFixed(4)}` : ""), + ); + } + + if (!record || itemCount === 0) { + // Every requested section failed. Report it as a skip rather than a write so + // the batch's counters stay honest. + return { status: "skipped", reason: "no-cues" }; + } + return { + status: "wrote", + record, + sections: stale, + itemCount, + warningCount, + costUsd, + engineCalls, + }; +} + +// Render one chunk as the text the engine sees. transcriptToMarkdown is the +// declared single source of truth for "transcript -> text for an AI"; the +// stampForCue override makes every line carry a HH:MM:SS marker, which is what +// the prompt tells the model to copy its `start` values from. +// +// toHms — NOT aiHandoff's hms() — because hms drops the hour field below an hour +// ("2:36") and abbreviates it above one ("1:00:00"), while the schema pattern +// requires two digits in all three fields. Feeding the model markers it cannot +// legally echo would reintroduce the malformed-timestamp failure the pin fixed. +// +// The description is omitted: it is the uploader's own promotional copy and +// biases titles toward it. +function renderChunk( + transcript: { id: string; title: string; channel?: string; duration?: number }, + cues: Cue[], +): string { + return transcriptToMarkdown( + { + id: transcript.id, + title: transcript.title, + channel: transcript.channel, + duration: transcript.duration, + cues, + }, + { + timestamps: true, + includeDescription: false, + includeTags: false, + stampForCue: (_clock, seconds) => toHms(seconds), + }, + ); +} diff --git a/common/controller/duplicateShorts.ts b/common/controller/duplicateShorts.ts @@ -37,12 +37,16 @@ import { DEFAULT_SHORT_THRESHOLD_SECONDS, DUPLICATES_FILENAME, DUPLICATE_REPORT_VERSION, + DUPLICATE_OVERRIDES_FILENAME, UnionFind, comparisonText, containment, jaccard, + pickCanonicalSlug, + sanitizeDuplicateOverrides, shingles, strongerMatch, + type DuplicateOverrides, type DuplicateCluster, type DuplicateMatchKind, type DuplicateReport, @@ -89,6 +93,10 @@ export async function detectDuplicateShorts( opts: DetectDuplicateShortsOptions, ): Promise<DuplicateReport> { const log = opts.onLog ?? ((m: string) => console.log(m)); + // Corpus-wide (thresholdSeconds: null) is untested at 74k videos across the + // long tail, so the run reports its own wall-clock: the cost of the pass has to + // be a measured number before anything is wired to it on a schedule. + const startedAt = Date.now(); const thresholdSeconds = opts.thresholdSeconds === undefined ? DEFAULT_SHORT_THRESHOLD_SECONDS @@ -297,7 +305,7 @@ export async function detectDuplicateShorts( const platforms = new Set(refs.map((r) => r.platform)); const channels = new Set(refs.map((r) => r.channelSlug)); const minDuration = Math.min(...refs.map((r) => r.duration)); - return { + const cluster: DuplicateCluster = { clusterId: sha1(refs.map((r) => r.slug).join("\n")), matchKind: agg.matchKind, score: agg.score, @@ -307,6 +315,11 @@ export async function detectDuplicateShorts( crossChannel: channels.size > 1, videoRefs: refs, }; + // Record the rule's choice of canonical member: the one derived work (an AI + // digest today, attribution later) is generated for and shared FROM. A human + // can override it in duplicates.overrides.json — which is a separate file + // precisely because this report is rewritten wholesale on every run. + return { ...cluster, canonicalSlug: pickCanonicalSlug(cluster) }; }); const rank: Record<DuplicateMatchKind, number> = { @@ -336,7 +349,9 @@ export async function detectDuplicateShorts( await writeFile(tmp, JSON.stringify(report)); await rename(tmp, outPath); log( - `Done: ${clusters.length} duplicate cluster(s) over ${videosInClusters} video(s) → ${DUPLICATES_FILENAME}.`, + `Done: ${clusters.length} duplicate cluster(s) over ${videosInClusters} video(s) → ${DUPLICATES_FILENAME} ` + + `in ${Math.round((Date.now() - startedAt) / 1000)}s ` + + `(${videosScanned} scanned, ${candidates.length} candidate pairs, ${matches.length} confirmed).`, ); return report; } @@ -357,6 +372,63 @@ export async function readDuplicateReport( } } +// --------------------------------------------------------------------------- +// Human review decisions (canonical choice / not-a-duplicate) +// --------------------------------------------------------------------------- + +export function duplicateOverridesPath(paths: Paths): string { + return path.join(paths.transcriptsDir, DUPLICATE_OVERRIDES_FILENAME); +} + +// Never throws: an unreadable or malformed overrides file reads as "no decisions +// recorded", so the duplicates page still renders. +export async function readDuplicateOverrides( + paths: Paths, +): Promise<DuplicateOverrides> { + try { + const raw = await readFile(duplicateOverridesPath(paths), "utf8"); + return sanitizeDuplicateOverrides(JSON.parse(raw)); + } catch { + return sanitizeDuplicateOverrides(null); + } +} + +// Record one cluster's decision, preserving every other cluster's. Read-modify- +// write with the same atomic tmp+rename the report itself uses. +export async function updateDuplicateOverride( + paths: Paths, + clusterId: string, + patch: { canonicalSlug?: string; notDuplicate?: boolean; note?: string }, +): Promise<DuplicateOverrides> { + const current = await readDuplicateOverrides(paths); + const existing = current.clusters[clusterId] ?? {}; + const next = { + ...existing, + ...(patch.canonicalSlug !== undefined + ? { canonicalSlug: patch.canonicalSlug } + : {}), + ...(patch.notDuplicate !== undefined + ? { notDuplicate: patch.notDuplicate } + : {}), + ...(patch.note !== undefined ? { note: patch.note } : {}), + decidedAt: new Date().toISOString(), + }; + // An empty patch clears the decision (back to "awaiting review") rather than + // leaving a decidedAt-only stub that would read as reviewed. + if (!next.canonicalSlug && next.notDuplicate !== true) { + delete current.clusters[clusterId]; + } else { + current.clusters[clusterId] = next; + } + const out = sanitizeDuplicateOverrides(current); + const file = duplicateOverridesPath(paths); + await mkdir(path.dirname(file), { recursive: true }); + const tmp = `${file}.tmp-${process.pid}`; + await writeFile(tmp, JSON.stringify(out, null, 2) + "\n"); + await rename(tmp, file); + return out; +} + // --- helpers --------------------------------------------------------------- function durationsClose( diff --git a/common/jobs/jobKinds.ts b/common/jobs/jobKinds.ts @@ -99,6 +99,37 @@ const JOB_KINDS: Record<string, JobKindMeta> = { bookmarkable: true, queueKeyStrategy: "custom", }, + // AI digest sweep, local (ollama) lane — the one that carries the corpus. Both + // digest kinds are drainable (the batch honors the drain signal: it stops + // pulling new videos and lets the in-flight one finish) and bookmarkable, since + // a channel-scoped sweep is exactly the kind of thing an operator re-launches. + "digest-channel-local": { + kind: "digest-channel-local", + label: "Digest channel (local)", + drainable: true, + bookmarkable: true, + queueKeyStrategy: "custom", + }, + // Same batch, metered lane. Off unless settings.digest.remoteEnabled is true, + // and it lands on its own queue key so it runs CONCURRENTLY with the local lane + // rather than behind it. + "digest-channel-remote": { + kind: "digest-channel-remote", + label: "Digest channel (metered)", + drainable: true, + bookmarkable: true, + queueKeyStrategy: "custom", + }, + // Copy a duplicate cluster's canonical digest onto its aligned mirrors. A fast + // file operation gated by the timestamp-alignment check, so it is not drainable + // but is bookmarkable. + "digest-share-cluster": { + kind: "digest-share-cluster", + label: "Share cluster digest", + drainable: false, + bookmarkable: false, + queueKeyStrategy: "parallel", + }, "redownload-incomplete-bucket": { kind: "redownload-incomplete-bucket", label: "Re-download truncated transcripts", diff --git a/common/jobs/registry.ts b/common/jobs/registry.ts @@ -11,7 +11,10 @@ export type JobStatus = | "failed" | "cancelled"; -export type JobProgressMetric = "downloads" | "transcripts"; +// Extended for the digest sweep. NOTE: this union has one re-spelled copy in +// RunningJobsList.tsx's RunningJobsListItem — kept as an IMPORT there now, so +// TypeScript actually flags the next member added here. +export type JobProgressMetric = "downloads" | "transcripts" | "digests"; export type JobProgress = { metric: JobProgressMetric; @@ -19,7 +22,7 @@ export type JobProgress = { target: number; }; -export type JobTaskKind = "download" | "transcribe"; +export type JobTaskKind = "download" | "transcribe" | "digest"; // A single in-flight sub-operation within a job (one video download or one // transcription). Only currently-running tasks are kept on the record — they diff --git a/common/jobs/snapshotScheduler.ts b/common/jobs/snapshotScheduler.ts @@ -27,7 +27,20 @@ import { drainStream } from "./drainStream"; // — including it would make the central runManagedFunction hook re-arm the timer // from inside the regen job, an infinite loop. "detect-duplicates" is a global // read-only scan with no per-channel snapshot impact. -const NO_REGEN_KINDS = new Set<string>(["refresh-report", "detect-duplicates"]); +// +// The digest kinds are here for a different, load-bearing reason: SCALE. +// ctx.recordTaskDone arms a channel-snapshot regen after EVERY completed +// sub-operation, and each regen is a full 16-way per-video fan-out over the whole +// channel. A digest sweep is ~119,600 sub-operations corpus-wide, so leaving them +// in would spend more machine time regenerating snapshots than generating +// digests. The digest actions instead regenerate ONCE per channel, at job end. +const NO_REGEN_KINDS = new Set<string>([ + "refresh-report", + "detect-duplicates", + "digest-channel-local", + "digest-channel-remote", + "digest-share-cluster", +]); export function shouldRequestSnapshot(kind: string): boolean { return !NO_REGEN_KINDS.has(kind); diff --git a/common/jobs/taskHooks.ts b/common/jobs/taskHooks.ts @@ -63,7 +63,16 @@ export function makeTaskTracker( // (DLOM_PROBE) drive the probe phase and probe-aware ETA. Transcribe // tasks without an appId (remote workers stream pre-parsed progress) fall // through to it too — harmless for non-matching lines. - const downloadParser = transcribeParser ? null : createDownloadProgressParser(); + // + // Digest tasks get NO parser: a digest's per-chunk progress is discrete + // (chunk 3 of 7), not a byte/second rate, and its log lines carry model + // token counts that the download parser would happily misread as a + // percentage. The task row shows an indeterminate bar instead, which is + // honest. + const downloadParser = + transcribeParser || kind === "digest" + ? null + : createDownloadProgressParser(); let ended = false; const onLog = (line: string) => { forwardLog(line); diff --git a/common/lib/digest-server.ts b/common/lib/digest-server.ts @@ -0,0 +1,298 @@ +// Server-only disk I/O for the two digest sidecars. Modelled on +// availability-server.ts, which is the only per-video sidecar writer in the repo +// with all four properties this needs: atomic tmp+rename, a tolerant validated +// read that degrades to null, append-on-change history so re-runs are auditable, +// and a PARTIAL-OWNERSHIP merge that preserves fields another writer owns. +// +// The partial-ownership property is the load-bearing one here: writeDigestSection +// replaces exactly one section and leaves the other alone, so a metered tags run +// never clobbers local chapters (and vice versa). + +import path from "node:path"; +import { readFile, rename, rm, writeFile } from "node:fs/promises"; +import { + DIGEST_FILENAME, + DIGEST_OVERRIDES_FILENAME, + DIGEST_OVERRIDES_VERSION, + DIGEST_SCHEMA_VERSION, + isDigestSectionKind, + type DigestChapter, + type DigestDerivedFrom, + type DigestHistoryEntry, + type DigestItem, + type DigestOverrides, + type DigestProvenance, + type DigestRecord, + type DigestSectionKind, + type DigestTag, + type DigestWarning, +} from "./digest"; + +// Cap the audit trail so a video that is regenerated across many prompt +// iterations during Stage B tuning cannot grow an unbounded sidecar. +const MAX_HISTORY_ENTRIES = 40; + +export function digestPath(videoDir: string): string { + return path.join(videoDir, DIGEST_FILENAME); +} + +export function digestOverridesPath(videoDir: string): string { + return path.join(videoDir, DIGEST_OVERRIDES_FILENAME); +} + +// --------------------------------------------------------------------------- +// Reads (tolerant: anything unparseable or structurally wrong reads as absent, +// so one corrupt sidecar can never fail a channel-wide sweep) +// --------------------------------------------------------------------------- + +export async function loadDigest( + videoDir: string, +): Promise<DigestRecord | null> { + try { + const raw = await readFile(digestPath(videoDir), "utf8"); + const parsed = JSON.parse(raw) as Partial<DigestRecord>; + if (typeof parsed?.digestSchemaVersion !== "number") return null; + if (!parsed.sections || typeof parsed.sections !== "object") return null; + return { + digestSchemaVersion: parsed.digestSchemaVersion, + promptVersion: + typeof parsed.promptVersion === "number" ? parsed.promptVersion : 0, + contextHash: + typeof parsed.contextHash === "string" ? parsed.contextHash : "", + warnings: Array.isArray(parsed.warnings) + ? (parsed.warnings as DigestWarning[]) + : [], + sections: parsed.sections, + ...(Array.isArray(parsed.history) + ? { history: parsed.history as DigestHistoryEntry[] } + : {}), + ...(parsed.derivedFrom + ? { derivedFrom: parsed.derivedFrom as DigestDerivedFrom } + : {}), + }; + } catch { + return null; + } +} + +// Cheap existence/coverage check for the snapshot's noDigest bucket and the +// channel digestCount, which run over every video dir in a channel. Reads the +// file (a few KB) rather than statting, because a digest whose sections are all +// empty is not coverage. +export async function hasDigest(videoDir: string): Promise<boolean> { + const record = await loadDigest(videoDir); + if (!record) return false; + return ( + (record.sections.chapters?.items.length ?? 0) > 0 || + (record.sections.tags?.items.length ?? 0) > 0 + ); +} + +export async function loadDigestOverrides( + videoDir: string, +): Promise<DigestOverrides | null> { + try { + const raw = await readFile(digestOverridesPath(videoDir), "utf8"); + const parsed = JSON.parse(raw) as Partial<DigestOverrides>; + // A hand-authored file may omit `version`; treat it as current rather than + // discarding human work over a missing scalar. + const chapters = sanitizeChapters(parsed.chapters); + const tags = sanitizeTags(parsed.tags); + if (chapters.length === 0 && tags.length === 0 && !parsed.note) return null; + return { + version: + typeof parsed.version === "number" + ? parsed.version + : DIGEST_OVERRIDES_VERSION, + ...(chapters.length > 0 ? { chapters } : {}), + ...(tags.length > 0 ? { tags } : {}), + ...(typeof parsed.note === "string" ? { note: parsed.note } : {}), + ...(typeof parsed.updatedAt === "string" + ? { updatedAt: parsed.updatedAt } + : {}), + }; + } catch { + return null; + } +} + +// Hand-authored overrides are coerced, not trusted: an entry missing an id or a +// body is dropped rather than poisoning the merge. +function sanitizeChapters(value: unknown): DigestChapter[] { + if (!Array.isArray(value)) return []; + const out: DigestChapter[] = []; + for (const raw of value) { + if (!raw || typeof raw !== "object") continue; + const r = raw as Record<string, unknown>; + if (typeof r.id !== "string" || !r.id.trim()) continue; + const title = typeof r.title === "string" ? r.title.trim() : ""; + const start = + typeof r.start === "number" && Number.isFinite(r.start) + ? Math.max(0, Math.floor(r.start)) + : undefined; + // An override may be a patch (id + enabled:false) with no title at all. + if (!title && r.enabled !== false) continue; + out.push({ + id: r.id, + ...(start !== undefined ? { start } : {}), + ...(typeof r.clock === "string" ? { clock: r.clock } : {}), + title, + decidedBy: "human", + ...(r.enabled === false ? { enabled: false } : {}), + } as DigestChapter); + } + return out; +} + +function sanitizeTags(value: unknown): DigestTag[] { + if (!Array.isArray(value)) return []; + const out: DigestTag[] = []; + for (const raw of value) { + if (!raw || typeof raw !== "object") continue; + const r = raw as Record<string, unknown>; + if (typeof r.id !== "string" || !r.id.trim()) continue; + const tag = typeof r.tag === "string" ? r.tag.trim() : ""; + if (!tag && r.enabled !== false) continue; + out.push({ + id: r.id, + tag, + decidedBy: "human", + ...(r.enabled === false ? { enabled: false } : {}), + }); + } + return out; +} + +// --------------------------------------------------------------------------- +// Writes +// --------------------------------------------------------------------------- + +async function writeJsonAtomic(file: string, value: unknown): Promise<void> { + const tmp = `${file}.tmp-${process.pid}`; + await writeFile(tmp, JSON.stringify(value, null, 2) + "\n"); + await rename(tmp, file); +} + +export async function writeDigest( + videoDir: string, + record: DigestRecord, +): Promise<void> { + await writeJsonAtomic(digestPath(videoDir), record); +} + +export type WriteDigestSectionInput = { + section: DigestSectionKind; + items: DigestItem[]; + provenance: DigestProvenance; + // Warnings from THIS pass. Warnings belonging to the section being rewritten + // are replaced; another section's warnings are preserved. + warnings: DigestWarning[]; +}; + +// Read-modify-write ONE section, preserving everything another writer owns. +// This is the only path the generator uses, and it never touches +// ai-digest.overrides.json — that file is human-owned, full stop. +export async function writeDigestSection( + videoDir: string, + input: WriteDigestSectionInput, +): Promise<DigestRecord> { + const existing = await loadDigest(videoDir); + const sections = { ...(existing?.sections ?? {}) }; + // The cast is contained here: DigestSectionKind and the item type are paired + // by construction at every call site (digestVideo builds both together). + (sections as Record<string, unknown>)[input.section] = { + provenance: input.provenance, + items: input.items, + }; + + // Drop the prior pass's warnings for THIS section only. + const keptWarnings = (existing?.warnings ?? []).filter( + (w) => w.section !== input.section, + ); + + const entry: DigestHistoryEntry = { + section: input.section, + generatedAt: input.provenance.generatedAt, + appId: input.provenance.appId, + model: input.provenance.model, + promptVersion: input.provenance.promptVersion, + contextHash: input.provenance.contextHash, + itemCount: input.items.length, + warningCount: input.warnings.length, + }; + const history = [...(existing?.history ?? []), entry].slice( + -MAX_HISTORY_ENTRIES, + ); + + const record: DigestRecord = { + digestSchemaVersion: DIGEST_SCHEMA_VERSION, + promptVersion: input.provenance.promptVersion, + contextHash: input.provenance.contextHash, + warnings: [...keptWarnings, ...input.warnings], + sections, + history, + }; + // A freshly generated section makes the record this video's own again: it is + // no longer a copy of a cluster's canonical member. + await writeDigest(videoDir, record); + return record; +} + +// Copy a canonical member's digest onto an aligned duplicate, stamped with the +// provenance of where it came from and the measured timing offset that made +// sharing safe. The receiving video's own overrides are untouched, so a human +// correction on a mirror still wins over the shared machine content. +export async function writeSharedDigest( + videoDir: string, + source: DigestRecord, + derivedFrom: DigestDerivedFrom, +): Promise<DigestRecord> { + const record: DigestRecord = { + ...source, + // The share is an event in the receiving video's history too. + history: [ + ...(source.history ?? []), + ...Object.entries(source.sections) + .filter(([kind]) => isDigestSectionKind(kind)) + .map(([kind, section]) => ({ + section: kind as DigestSectionKind, + generatedAt: derivedFrom.sharedAt, + appId: section!.provenance.appId, + model: section!.provenance.model, + promptVersion: section!.provenance.promptVersion, + contextHash: section!.provenance.contextHash, + itemCount: section!.items.length, + warningCount: 0, + })), + ].slice(-MAX_HISTORY_ENTRIES), + derivedFrom, + }; + await writeDigest(videoDir, record); + return record; +} + +// Human-authored write. Separate function, separate file, so nothing in the +// generation path can reach it. +export async function writeDigestOverrides( + videoDir: string, + overrides: DigestOverrides, +): Promise<void> { + const hasContent = + (overrides.chapters?.length ?? 0) > 0 || + (overrides.tags?.length ?? 0) > 0 || + Boolean(overrides.note); + const file = digestOverridesPath(videoDir); + if (!hasContent) { + // Emptying the override list means "revert to machine output" — remove the + // file rather than leaving an empty shadow behind. + await rm(file, { force: true }); + return; + } + await writeJsonAtomic(file, { + version: DIGEST_OVERRIDES_VERSION, + ...(overrides.chapters?.length ? { chapters: overrides.chapters } : {}), + ...(overrides.tags?.length ? { tags: overrides.tags } : {}), + ...(overrides.note ? { note: overrides.note } : {}), + updatedAt: overrides.updatedAt ?? new Date().toISOString(), + }); +} diff --git a/common/lib/digest.test.ts b/common/lib/digest.test.ts @@ -0,0 +1,255 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + DIGEST_SCHEMA_VERSION, + chapterId, + effectiveDigest, + isSectionFresh, + isSharedFrom, + tagId, + type DigestOverrides, + type DigestRecord, +} from "./digest"; + +// The override model is the single most important thing in the digest layer: a +// full sweep is weeks of wall-clock, so a human correction that a regeneration +// destroys is work that can never be affordably redone. + +function machineRecord(over: Partial<DigestRecord> = {}): DigestRecord { + return { + digestSchemaVersion: DIGEST_SCHEMA_VERSION, + promptVersion: 1, + contextHash: "abc123", + warnings: [], + sections: { + chapters: { + provenance: { + appId: "ollama-direct", + model: "qwen2.5:7b", + modelRequested: "qwen2.5:7b", + lane: "local-gpu", + generatedAt: "2026-07-26T00:00:00.000Z", + promptVersion: 1, + contextHash: "abc123", + }, + items: [ + { id: "c0", start: 0, clock: "00:00:00", title: "Intro", decidedBy: "ai" }, + { + id: "c300", + start: 300, + clock: "00:05:00", + title: "Court filing deadlnes", + decidedBy: "ai", + }, + ], + }, + }, + ...over, + }; +} + +test("effectiveDigest returns machine content when there are no overrides", () => { + const e = effectiveDigest(machineRecord(), null); + assert.deepEqual( + e.chapters.map((c) => c.title), + ["Intro", "Court filing deadlnes"], + ); + assert.equal(e.hasOverrides, false); + assert.ok(e.chapters.every((c) => c.decidedBy === "ai")); +}); + +test("a human override REPLACES the machine item and reports decidedBy: human", () => { + const overrides: DigestOverrides = { + version: 1, + chapters: [ + { id: "c300", start: 300, clock: "00:05:00", title: "Court filing deadlines", decidedBy: "human" }, + ], + }; + const e = effectiveDigest(machineRecord(), overrides); + const fixed = e.chapters.find((c) => c.id === "c300"); + assert.equal(fixed?.title, "Court filing deadlines", "the typo is corrected"); + assert.equal(fixed?.decidedBy, "human"); + assert.equal(e.chapters.length, 2, "the untouched machine chapter survives"); + assert.equal(e.hasOverrides, true); +}); + +test("an override authored as a patch does not blank the machine's start", () => { + const overrides: DigestOverrides = { + version: 1, + // Title-only patch: no `start`. + chapters: [{ id: "c300", title: "Retitled", decidedBy: "human" }] as never, + }; + const e = effectiveDigest(machineRecord(), overrides); + const patched = e.chapters.find((c) => c.id === "c300"); + assert.equal(patched?.start, 300, "the machine's start is preserved"); + assert.equal(patched?.title, "Retitled"); +}); + +test("enabled: false SUPPRESSES a generated item without deleting it", () => { + const overrides: DigestOverrides = { + version: 1, + chapters: [{ id: "c0", title: "", decidedBy: "human", enabled: false }] as never, + }; + const e = effectiveDigest(machineRecord(), overrides); + assert.deepEqual( + e.chapters.map((c) => c.id), + ["c300"], + ); + // The suppression is a recorded decision, so re-generating the same item does + // not resurrect something a human rejected. + assert.equal(e.hasOverrides, true); +}); + +test("an override with a new id APPENDS a human-authored chapter", () => { + const overrides: DigestOverrides = { + version: 1, + chapters: [ + { id: "c600", start: 600, clock: "00:10:00", title: "Missed topic", decidedBy: "human" }, + ], + }; + const e = effectiveDigest(machineRecord(), overrides); + assert.deepEqual( + e.chapters.map((c) => c.start), + [0, 300, 600], + "composed list stays sorted by start", + ); +}); + +test("overrides survive a regeneration that rewrites the machine file", () => { + const overrides: DigestOverrides = { + version: 1, + chapters: [{ id: "c300", title: "Court filing deadlines", decidedBy: "human" }] as never, + }; + // A later run at a new promptVersion re-emits a DIFFERENT title at the same + // snapped start — the id is derived from the start, so the human's correction + // still shadows it. + const regenerated = machineRecord({ + promptVersion: 2, + sections: { + chapters: { + provenance: { + appId: "ollama-direct", + model: "qwen2.5:7b", + modelRequested: "qwen2.5:7b", + lane: "local-gpu", + generatedAt: "2026-08-01T00:00:00.000Z", + promptVersion: 2, + contextHash: "abc123", + }, + items: [ + { + id: "c300", + start: 300, + clock: "00:05:00", + title: "Some new machine title", + decidedBy: "ai", + }, + ], + }, + }, + }); + const e = effectiveDigest(regenerated, overrides); + assert.equal(e.chapters[0].title, "Court filing deadlines"); + assert.equal(e.chapters[0].decidedBy, "human"); +}); + +test("chapterId is derived from the snapped start, so it is stable", () => { + assert.equal(chapterId(300), "c300"); + assert.equal(chapterId(300.9), "c300"); +}); + +test("tagId normalizes so a re-generated tag lands on the same id", () => { + assert.equal(tagId("Court Filings"), "tcourt-filings"); + assert.equal(tagId("court filings!"), "tcourt-filings"); +}); + +// --------------------------------------------------------------------------- +// Freshness — what keeps a re-run from becoming a second multi-week sweep +// --------------------------------------------------------------------------- + +const TARGET = { + appId: "ollama-direct", + model: "qwen2.5:7b", + promptVersion: 1, + contextHash: "abc123", +}; + +test("an unchanged identity is fresh (a re-run is a no-op)", () => { + assert.equal(isSectionFresh(machineRecord(), "chapters", TARGET), true); +}); + +test("a missing record or section is stale", () => { + assert.equal(isSectionFresh(null, "chapters", TARGET), false); + assert.equal(isSectionFresh(machineRecord(), "tags", TARGET), false); +}); + +test("a bumped promptVersion makes the section stale", () => { + assert.equal( + isSectionFresh(machineRecord(), "chapters", { ...TARGET, promptVersion: 2 }), + false, + ); +}); + +test("a changed contextHash makes the section stale", () => { + // This is why contextHash is plumbed BEFORE the sweep even with empty notes: + // adding a channel note later must invalidate that channel, not the corpus. + assert.equal( + isSectionFresh(machineRecord(), "chapters", { ...TARGET, contextHash: "zzz" }), + false, + ); +}); + +test("a different engine or model makes the section stale", () => { + assert.equal( + isSectionFresh(machineRecord(), "chapters", { ...TARGET, appId: "claude-code" }), + false, + ); + assert.equal( + isSectionFresh(machineRecord(), "chapters", { ...TARGET, model: "llama3" }), + false, + ); +}); + +test("freshness compares the REQUESTED model, not the resolved one", () => { + // An alias resolving to a full tag is not a model change; treating it as one + // would re-run the entire corpus. + const record = machineRecord({ + sections: { + chapters: { + provenance: { + appId: "ollama-direct", + model: "qwen2.5:7b", + modelRequested: "qwen2.5", + lane: "local-gpu", + generatedAt: "2026-07-26T00:00:00.000Z", + promptVersion: 1, + contextHash: "abc123", + }, + items: [], + }, + }, + }); + assert.equal( + isSectionFresh(record, "chapters", { ...TARGET, model: "qwen2.5" }), + true, + ); +}); + +test("a schema-version bump invalidates every digest on disk", () => { + const record = machineRecord({ digestSchemaVersion: 0 }); + assert.equal(isSectionFresh(record, "chapters", TARGET), false); +}); + +test("isSharedFrom recognizes a digest copied from a cluster's canonical member", () => { + const shared = machineRecord({ + derivedFrom: { + slug: "Quartering/abc", + clusterId: "cluster1", + sharedAt: "2026-07-26T00:00:00.000Z", + offsetSeconds: 0.4, + }, + }); + assert.equal(isSharedFrom(shared, "Quartering/abc"), true); + assert.equal(isSharedFrom(shared, "Quartering/other"), false); + assert.equal(isSharedFrom(machineRecord(), "Quartering/abc"), false); +}); diff --git a/common/lib/digest.ts b/common/lib/digest.ts @@ -0,0 +1,355 @@ +// Client-safe types, constants and PURE merge helpers for the per-video AI +// digest (chapters + topic tags). No node-only imports — the editor's video +// panel is a "use client" file that pulls this in for the review UI. Server-only +// disk I/O lives in digest-server.ts. Same split, for the same reason, as +// doNotClean.ts / doNotClean-server.ts. +// +// TWO SIDECARS PER VIDEO DIR, and the split is the whole point: +// +// ai-digest.json machine-generated; freely overwritten by a re-run +// ai-digest.overrides.json human-authored; regeneration NEVER writes it +// +// A sweep over 74k transcripts takes weeks, so a redo is unaffordable and human +// corrections must survive one. Keeping the two in separate files means no merge +// bug in the generator can destroy hand-written work: the generator only ever +// opens the machine file. Readers compose the two with effectiveDigest(). +// +// NEVER rename these to `transcript.<x>.<y>` — SUB_FILE_RE in videoStatus.ts +// would claim such a file as a subtitle track. + +export const DIGEST_FILENAME = "ai-digest.json"; +export const DIGEST_OVERRIDES_FILENAME = "ai-digest.overrides.json"; + +// Shape of ai-digest.json itself. Bumping this invalidates every digest on +// disk, so it changes only when the FILE LAYOUT changes — not when a prompt +// changes (that is promptVersion, which invalidates per section). +export const DIGEST_SCHEMA_VERSION = 1; +export const DIGEST_OVERRIDES_VERSION = 1; + +// Which engine lane produced a section. "local-gpu" is the default and carries +// the corpus; "remote-api" is the opt-in metered overflow. +// +// The lane type, the app ids and the per-app config live HERE rather than in +// digestApps.ts because settings.ts needs them and digestApps.ts imports execa — +// a node-only dependency that must never be reachable from a client bundle. +// digestApps.ts re-exports them so the registry still reads as one unit. +export type DigestLane = "local-gpu" | "remote-api"; + +export const OLLAMA_DIGEST_APP_ID = "ollama-direct"; +export const CLAUDE_DIGEST_APP_ID = "claude-code"; +export const DEFAULT_DIGEST_APP_ID = OLLAMA_DIGEST_APP_ID; + +// Per-app configuration persisted under settings.digest.apps[id]. Every field is +// optional; an app falls back to its own defaults. +export type DigestAppConfig = { + // Binary path/name override (process-based apps only). + bin?: string; + // Base URL override (HTTP apps only). + baseUrl?: string; + // Model id, e.g. "qwen2.5:7b" or "haiku". + model?: string; + // Context window in tokens. MUST reach the engine explicitly for ollama: its + // 4096 default silently truncates the input and the model then summarizes + // whatever fragment survived — measured, and the single easiest way to get + // quietly-wrong output at scale. + numCtx?: number; + // Sampling temperature. 0 for a structured extraction task. + temperature?: number; + // Per-request wall-clock ceiling (ms). A wedged engine must not stall a sweep. + timeoutMs?: number; +}; + +// The two generated sections. Per-SECTION provenance (not per-file) because the +// controller does a read-modify-write merge, so one video can legitimately hold +// local chapters and metered tags. +export const DIGEST_SECTION_KINDS = ["chapters", "tags"] as const; +export type DigestSectionKind = (typeof DIGEST_SECTION_KINDS)[number]; + +export function isDigestSectionKind(v: unknown): v is DigestSectionKind { + return ( + typeof v === "string" && + (DIGEST_SECTION_KINDS as readonly string[]).includes(v) + ); +} + +// Who decided this item. Every AI decision is marked as such and is +// human-overridable — the same mechanism Phase 9 attribution will reuse rather +// than inventing a second one. +export type DecidedBy = "ai" | "human"; + +// A chapter: a titled moment. `start` is SECONDS (the cue unit — see vtt.ts), +// snapped to a real cue boundary by the parser. `clock` is the HH:MM:SS form the +// model emitted, kept for auditing what the model actually said. +export type DigestChapter = { + // Stable id so an override can shadow exactly one generated item. Derived from + // the snapped start (see chapterId) — deterministic, no Date.now/random. + id: string; + start: number; + clock: string; + title: string; + decidedBy: DecidedBy; + // Defaults true. An override sets false to SUPPRESS a generated item without + // deleting it (the searchAliases.ts `enabled` idiom), so a regeneration that + // re-emits the same item does not resurrect something a human rejected. + enabled?: boolean; +}; + +export type DigestTag = { + id: string; + tag: string; + decidedBy: DecidedBy; + enabled?: boolean; +}; + +export type DigestItem = DigestChapter | DigestTag; + +// Why an item or a whole chunk was rejected. Recorded, never silently dropped — +// a multi-week sweep is only tunable if its failures are inspectable, and the +// Phase 11a review queue is built entirely on this array. +export type DigestWarningCode = + | "malformed-timestamp" + | "out-of-range" + | "language-drift" + | "non-monotonic" + | "seam-duplicate" + | "empty-title" + | "empty-output" + | "parse-failed" + | "chunk-failed"; + +export type DigestWarning = { + code: DigestWarningCode; + // Which section's generation produced it, so warnings stay attributable after + // a read-modify-write merge of two separately-generated sections. + section: DigestSectionKind; + // Zero-based index of the transcript chunk that produced it, when applicable. + chunk?: number; + // The offending value, verbatim, so a prompt regression is diagnosable from + // the artifact alone. + value?: string; + detail?: string; +}; + +// What a regeneration compares against to decide "skip". This is what makes a +// re-run targeted instead of a second multi-week sweep. +export type DigestProvenance = { + appId: string; + // The model that ACTUALLY ran, as the engine reported it (e.g. "qwen2.5:7b"). + // The audit record. + model: string; + // What the config ASKED for (e.g. "qwen2.5"). Freshness compares this, not + // `model`: an alias resolving to a full tag is not a model change, and treating + // it as one would re-run the entire corpus. Older records lack it, so readers + // fall back to `model`. + modelRequested?: string; + lane: DigestLane; + generatedAt: string; + // Bumped when the prompt/schema changes. Invalidates this section only. + promptVersion: number; + // Hash of the channel-context inputs. Plumbed from the start even while + // context files are empty, so adding them later doesn't invalidate the corpus. + contextHash: string; + // Number of transcript chunks the engine was asked to process, and how many + // came back usable. A section generated from 3 of 5 chunks is suspect. + chunks?: number; + chunksOk?: number; + // Metered lanes only: what this section cost. + costUsd?: number; +}; + +export type DigestSection<T extends DigestItem> = { + provenance: DigestProvenance; + items: T[]; +}; + +export type DigestSections = { + chapters?: DigestSection<DigestChapter>; + tags?: DigestSection<DigestTag>; +}; + +// One entry per generation pass, appended on change so a re-run is auditable +// (the availability-server.ts history idiom). +export type DigestHistoryEntry = { + section: DigestSectionKind; + generatedAt: string; + appId: string; + model: string; + promptVersion: number; + contextHash: string; + itemCount: number; + warningCount: number; +}; + +export type DigestRecord = { + digestSchemaVersion: number; + // The most recent generation pass's prompt/context identity. Per-section + // provenance is authoritative for freshness (a video can hold sections made + // at different versions); these mirror whichever pass wrote last, so a + // corpus-wide sweep can be surveyed without opening every section. + promptVersion: number; + contextHash: string; + warnings: DigestWarning[]; + sections: DigestSections; + history?: DigestHistoryEntry[]; + // Set when this digest was SHARED from another video (a duplicate cluster's + // canonical member) rather than generated for this one. Keeps the sharing + // honest and lets a later correction propagate to the whole cluster. + derivedFrom?: DigestDerivedFrom; +}; + +export type DigestDerivedFrom = { + // `${channelSlug}/${id}` of the canonical member the digest was generated for. + slug: string; + clusterId: string; + sharedAt: string; + // Measured max cue-timing offset (seconds) between the two videos at the + // sampled anchors. Sharing only happens at near-zero offset, so this records + // WHY it was considered safe. + offsetSeconds: number; +}; + +// --------------------------------------------------------------------------- +// Overrides +// --------------------------------------------------------------------------- + +// Human-authored shadow of the machine file, id-keyed exactly like a per-site +// search-alias list shadows the global one (mergeAliases: later source wins). +// An override entry with an id that matches a generated item REPLACES it; an id +// with no match APPENDS. `enabled: false` suppresses without deleting. +export type DigestOverrides = { + version: number; + chapters?: DigestChapter[]; + tags?: DigestTag[]; + // Free-text operator note (why this was corrected). Never read by code. + note?: string; + updatedAt?: string; +}; + +// --------------------------------------------------------------------------- +// Pure helpers +// --------------------------------------------------------------------------- + +// Deterministic per-item ids. A chapter's identity is its (snapped) start +// second: that is what a human is correcting when they retitle a chapter, and it +// stays stable across a regeneration that produces the same segmentation. +export function chapterId(startSeconds: number): string { + return `c${Math.max(0, Math.floor(startSeconds))}`; +} + +// A tag's identity is its normalized text, so re-generating the same tag lands +// on the same id (and thus keeps honoring a human's `enabled: false`). +export function tagId(tag: string): string { + const slug = tag + .toLowerCase() + .replace(/[^\p{L}\p{N}]+/gu, "-") + .replace(/^-+|-+$/g, ""); + return `t${slug || "tag"}`; +} + +// Merge a machine-generated item list with its human shadow. Later source wins +// per id (mergeAliases), an overridden item reports decidedBy: "human", and +// suppressed items (enabled === false) are dropped from the effective list. +// `sort` keeps the composed list in the caller's canonical order. +function mergeItems<T extends DigestItem>( + machine: T[], + overrides: T[] | undefined, + sort: (a: T, b: T) => number, +): T[] { + const byId = new Map<string, T>(); + for (const item of machine) byId.set(item.id, item); + for (const item of overrides ?? []) { + const existing = byId.get(item.id); + // A human edit is authoritative for the fields it names, but an override + // authored as a patch (id + title only) must not blank the machine's start. + byId.set(item.id, { ...(existing ?? {}), ...item, decidedBy: "human" } as T); + } + return Array.from(byId.values()) + .filter((item) => item.enabled !== false) + .sort(sort); +} + +const byStart = (a: DigestChapter, b: DigestChapter): number => + a.start - b.start || a.title.localeCompare(b.title); +const byTag = (a: DigestTag, b: DigestTag): number => a.tag.localeCompare(b.tag); + +export type EffectiveDigest = { + chapters: DigestChapter[]; + tags: DigestTag[]; + // True when any effective item came from the override file — the signal the + // editor uses to show "edited by hand". + hasOverrides: boolean; + warnings: DigestWarning[]; + sections: DigestSections; + derivedFrom?: DigestDerivedFrom; +}; + +// Compose the machine file and the human shadow into what a reader should see. +// Either side may be null (no digest yet / no corrections yet). +export function effectiveDigest( + machine: DigestRecord | null, + overrides: DigestOverrides | null, +): EffectiveDigest { + const chapters = mergeItems( + machine?.sections.chapters?.items ?? [], + overrides?.chapters, + byStart, + ); + const tags = mergeItems( + machine?.sections.tags?.items ?? [], + overrides?.tags, + byTag, + ); + return { + chapters, + tags, + hasOverrides: + (overrides?.chapters?.length ?? 0) > 0 || + (overrides?.tags?.length ?? 0) > 0, + warnings: machine?.warnings ?? [], + sections: machine?.sections ?? {}, + ...(machine?.derivedFrom ? { derivedFrom: machine.derivedFrom } : {}), + }; +} + +// The freshness test that keeps a re-run from becoming a second sweep: a section +// is regenerated ONLY when its recorded identity differs from what we would +// produce now. Missing section → stale (never generated). +export type DigestFreshnessTarget = { + appId: string; + model: string; + promptVersion: number; + contextHash: string; + schemaVersion?: number; +}; + +export function isSectionFresh( + record: DigestRecord | null, + section: DigestSectionKind, + target: DigestFreshnessTarget, +): boolean { + if (!record) return false; + if ( + record.digestSchemaVersion !== + (target.schemaVersion ?? DIGEST_SCHEMA_VERSION) + ) { + return false; + } + const p = record.sections[section]?.provenance; + if (!p) return false; + return ( + p.appId === target.appId && + (p.modelRequested ?? p.model) === target.model && + p.promptVersion === target.promptVersion && + p.contextHash === target.contextHash + ); +} + +// A shared digest is fresh for the receiving video as long as it still points at +// the same canonical member; the canonical member's own freshness is what drives +// regeneration, and the share is re-applied from it. +export function isSharedFrom( + record: DigestRecord | null, + canonicalSlug: string, +): boolean { + return record?.derivedFrom?.slug === canonicalSlug; +} diff --git a/common/lib/digestApps.ts b/common/lib/digestApps.ts @@ -0,0 +1,375 @@ +// First-class registry of DIGEST apps — the engines that turn a transcript into +// derived content (chapters, topic tags). Deliberately mirrors +// transcriptionApps.ts: each app owns how it is invoked, how its config maps onto +// that invocation, and how its raw output becomes a parsed JSON body. The editor +// selects one app per section, storing a small per-app config block +// (settings.digest.apps[id]); digestVideo resolves the app and runs it. +// +// LOCAL-FIRST, BY DECISION. Hosted AI is not a dependency this project can rely +// on given the nature of the archived content, so `ollama-direct` carries the +// corpus and making it good enough IS the goal. `claude-code` is built but OFF by +// default (settings.digest.remoteEnabled === false) — an opt-in overflow for the +// >4h tail or a channel where local quality is poor, never the default. +// +// Pure definitions only (no I/O beyond reading process.env / getPaths inside +// builders) so both the `common` controllers and the editor UI can import this. +// Dependency direction is one-way: settings.ts -> digestApps.ts -> { paths, +// parseStdoutJson }. Keep it that way to avoid import cycles. + +import { execa } from "execa"; +import { getPaths } from "./paths"; +import { extractJsonObject, parseStdoutJson } from "./parseStdoutJson"; +import { + CLAUDE_DIGEST_APP_ID, + DEFAULT_DIGEST_APP_ID, + OLLAMA_DIGEST_APP_ID, + type DigestAppConfig, + type DigestLane, +} from "./digest"; + +// The lane type, the app ids and DigestAppConfig are defined in the client-safe +// digest.ts and re-exported here, so settings.ts can consume them without +// reaching this module (which imports execa). `lane` is also WHY the two engines +// get SEPARATE queue keys: registry.ts hardcodes concurrency 1 per key, so one +// shared key would serialize a GPU-bound lane behind a network-bound one and +// waste half the throughput of a multi-week sweep. +export { CLAUDE_DIGEST_APP_ID, DEFAULT_DIGEST_APP_ID, OLLAMA_DIGEST_APP_ID }; +export type { DigestAppConfig, DigestLane }; + +export type DigestRunInput = { + system: string; + prompt: string; + // JSON schema for the expected body. Engines with constrained decoding + // (ollama's `format`) enforce it; CLI lanes get it embedded in the prompt. + schema: Record<string, unknown>; + config: DigestAppConfig; + signal?: AbortSignal; + onLog?: (msg: string) => void; +}; + +export type DigestRunResult = { + // The parsed JSON body. Shape validation is the parser's job (digestParse.ts), + // not the engine's — an engine only guarantees "this is JSON". + data: unknown; + // The model that ACTUALLY ran, as reported by the engine when it says so. + // Recorded in provenance, so a config of "qwen2.5" resolving to "qwen2.5:7b" + // doesn't later look like a model change and trigger a needless regeneration. + model: string; + durationMs: number; + // Metered lanes only. + costUsd?: number; + inputTokens?: number; + outputTokens?: number; +}; + +export type DigestApp = { + id: string; + label: string; + lane: DigestLane; + // True when a run costs money. Metered apps are opt-in, never auto-queued, and + // their per-call count + cumulative cost is logged by the batch. + metered: boolean; + // Which config fields this app surfaces in the settings UI / consumes. + fields: { + bin?: boolean; + baseUrl?: boolean; + model?: boolean; + numCtx?: boolean; + temperature?: boolean; + }; + // Model used when the per-app `model` override is empty. + defaultModel: () => string; + // Cheap reachability probe, so the editor can say "ollama is down" instead of + // failing 74k items one at a time. Mirrors pingRemoteHealth's contract: true + // only on a positive response, never throws. + probe: (config: DigestAppConfig) => Promise<boolean>; + run: (input: DigestRunInput) => Promise<DigestRunResult>; +}; + +// Fallbacks shared by both apps. 16k context is comfortable on an 8 GB card +// (qwen2.5:7b KV cache ≈ 56 KB/token → ~0.9 GB at 16k, atop 4.7 GB of weights). +const DEFAULT_NUM_CTX = 16384; +const DEFAULT_TEMPERATURE = 0; +const DEFAULT_TIMEOUT_MS = 10 * 60_000; + +function resolveNumCtx(config: DigestAppConfig): number { + return typeof config.numCtx === "number" && config.numCtx > 0 + ? Math.floor(config.numCtx) + : DEFAULT_NUM_CTX; +} + +function resolveTemperature(config: DigestAppConfig): number { + return typeof config.temperature === "number" && config.temperature >= 0 + ? config.temperature + : DEFAULT_TEMPERATURE; +} + +function resolveTimeoutMs(config: DigestAppConfig): number { + return typeof config.timeoutMs === "number" && config.timeoutMs > 0 + ? Math.floor(config.timeoutMs) + : DEFAULT_TIMEOUT_MS; +} + +// --------------------------------------------------------------------------- +// ollama-direct — the local GPU lane that carries the corpus +// --------------------------------------------------------------------------- + +function ollamaBase(config: DigestAppConfig): string { + const override = config.baseUrl?.trim(); + // Strip trailing slashes the way remoteTranscribe.ts normalizes a worker base, + // so callers can always concatenate a leading-slash path. + return (override ? override.replace(/\/+$/, "") : getPaths().ollamaUrl) || ""; +} + +const ollamaDirect: DigestApp = { + id: OLLAMA_DIGEST_APP_ID, + label: "ollama (local, JSON schema)", + lane: "local-gpu", + metered: false, + fields: { baseUrl: true, model: true, numCtx: true, temperature: true }, + defaultModel: () => process.env.OLLAMA_DIGEST_MODEL ?? "qwen2.5:7b", + async probe(config) { + const base = ollamaBase(config); + if (!base) return false; + try { + const res = await fetch(`${base}/api/tags`, { + signal: AbortSignal.timeout(3000), + }); + return res.ok; + } catch { + return false; + } + }, + async run({ system, prompt, schema, config, signal, onLog }) { + const base = ollamaBase(config); + if (!base) throw new Error("ollama URL is not configured (set OLLAMA_URL)"); + const model = config.model?.trim() || ollamaDirect.defaultModel(); + const numCtx = resolveNumCtx(config); + const startedAt = Date.now(); + + // The timeout is combined with the caller's cancel signal so a drain/cancel + // stops a long generation promptly AND a wedged engine can't stall forever. + const timeout = AbortSignal.timeout(resolveTimeoutMs(config)); + const composed = signal ? AbortSignal.any([signal, timeout]) : timeout; + + let res: Response; + try { + res = await fetch(`${base}/api/chat`, { + method: "POST", + headers: { "content-type": "application/json" }, + signal: composed, + body: JSON.stringify({ + model, + stream: false, + // Schema-constrained decoding. This — not prompt wording — is what + // made a 7B model emit well-formed timestamps. + format: schema, + options: { + // Explicit, always. See DigestAppConfig.numCtx. + num_ctx: numCtx, + temperature: resolveTemperature(config), + }, + messages: [ + { role: "system", content: system }, + { role: "user", content: prompt }, + ], + }), + }); + } catch (err) { + const message = (err as Error)?.message ?? String(err); + if (signal?.aborted) throw err; + throw new Error( + `ollama request failed (${base}): ${message}. Is the ollama service running?`, + ); + } + if (!res.ok) { + const body = await res.text().catch(() => ""); + throw new Error( + `ollama returned ${res.status}: ${body.slice(0, 300) || res.statusText}`, + ); + } + const body = (await res.json()) as { + model?: string; + message?: { content?: string }; + prompt_eval_count?: number; + eval_count?: number; + }; + const content = body.message?.content ?? ""; + const data = extractJsonObject(content); + if (data === null) { + throw new Error( + `ollama returned unparseable content: ${content.slice(0, 300)}`, + ); + } + onLog?.( + `ollama ${body.model ?? model}: ${body.prompt_eval_count ?? "?"} in / ${ + body.eval_count ?? "?" + } out tokens in ${Math.round((Date.now() - startedAt) / 100) / 10}s (num_ctx ${numCtx})`, + ); + return { + data, + model: body.model ?? model, + durationMs: Date.now() - startedAt, + inputTokens: body.prompt_eval_count, + outputTokens: body.eval_count, + }; + }, +}; + +// --------------------------------------------------------------------------- +// claude-code — the metered overflow lane, OFF by default +// --------------------------------------------------------------------------- + +const claudeCode: DigestApp = { + id: CLAUDE_DIGEST_APP_ID, + label: "Claude Code CLI (metered)", + lane: "remote-api", + metered: true, + fields: { bin: true, model: true }, + defaultModel: () => process.env.CLAUDE_DIGEST_MODEL ?? "", + async probe(config) { + const bin = config.bin?.trim() || getPaths().claudeBin; + try { + const res = await execa(bin, ["--version"], { + buffer: true, + reject: false, + timeout: 10_000, + }); + return res.exitCode === 0; + } catch { + return false; + } + }, + async run({ system, prompt, schema, config, signal, onLog }) { + const bin = config.bin?.trim() || getPaths().claudeBin; + const model = config.model?.trim() || claudeCode.defaultModel(); + const startedAt = Date.now(); + const argv = ["-p", "--output-format", "json"]; + if (model) argv.push("--model", model); + + // No constrained decoding over the CLI, so the contract goes in the prompt + // and the parser's guards do the enforcing. The prompt is passed on stdin + // rather than argv: a 12k-token chunk is ~50 KB and does not belong in an + // argument list. + const stdin = [ + system, + "", + "Reply with a single JSON object and nothing else — no prose, no code fence.", + "It must validate against this JSON schema:", + JSON.stringify(schema), + "", + prompt, + ].join("\n"); + + let result: { exitCode: number | null; stdout: string; stderr: string }; + try { + result = (await execa(bin, argv, { + input: stdin, + buffer: true, + reject: false, + cancelSignal: signal, + timeout: resolveTimeoutMs(config), + })) as typeof result; + } catch (err) { + const message = (err as Error)?.message ?? String(err); + // Same ENOENT-to-friendly-message treatment as the gallery-dl fetcher: a + // missing optional binary is a configuration problem, not a crash. + throw new Error( + /ENOENT/.test(message) + ? `claude CLI not found (set CLAUDE_BIN)` + : message, + ); + } + if (result.exitCode !== 0) { + throw new Error( + `claude exited ${result.exitCode}: ${(result.stderr || result.stdout).slice(0, 300)}`, + ); + } + // Two layers: the CLI's own JSON wrapper, then the model's JSON body inside + // the wrapper's `result` string. + const wrapper = parseStdoutJson(result.stdout) as { + result?: unknown; + total_cost_usd?: unknown; + usage?: { input_tokens?: number; output_tokens?: number }; + is_error?: unknown; + } | null; + if (!wrapper || typeof wrapper !== "object") { + throw new Error( + `claude produced no parseable JSON wrapper: ${result.stdout.slice(0, 300)}`, + ); + } + if (wrapper.is_error === true) { + throw new Error( + `claude reported an error: ${String(wrapper.result).slice(0, 300)}`, + ); + } + const inner = typeof wrapper.result === "string" ? wrapper.result : ""; + const data = extractJsonObject(inner); + if (data === null) { + throw new Error( + `claude returned unparseable content: ${inner.slice(0, 300)}`, + ); + } + const costUsd = + typeof wrapper.total_cost_usd === "number" + ? wrapper.total_cost_usd + : undefined; + onLog?.( + `claude ${model || "(default model)"}: ${ + costUsd !== undefined ? `$${costUsd.toFixed(4)}` : "cost unknown" + } in ${Math.round((Date.now() - startedAt) / 100) / 10}s`, + ); + return { + data, + model: model || "claude-default", + durationMs: Date.now() - startedAt, + ...(costUsd !== undefined ? { costUsd } : {}), + inputTokens: wrapper.usage?.input_tokens, + outputTokens: wrapper.usage?.output_tokens, + }; + }, +}; + +// --------------------------------------------------------------------------- +// Registry +// --------------------------------------------------------------------------- + +export const DIGEST_APPS: Record<string, DigestApp> = { + [ollamaDirect.id]: ollamaDirect, + [claudeCode.id]: claudeCode, +}; + +// Total by construction — an unknown id falls back to the local default rather +// than throwing, so a hand-edited settings.json can never crash a sweep. +export function getDigestApp(id: string | undefined): DigestApp { + return ( + (id ? DIGEST_APPS[id] : undefined) ?? DIGEST_APPS[DEFAULT_DIGEST_APP_ID] + ); +} + +export function isDigestAppId(id: unknown): id is string { + return typeof id === "string" && Boolean(DIGEST_APPS[id]); +} + +// A client-safe view of an app (no functions). Build it on the server and pass it +// to the settings form so this module — which reaches getPaths()/process.env — +// never ends up in the client bundle. +export type DigestAppDescriptor = { + id: string; + label: string; + lane: DigestLane; + metered: boolean; + fields: DigestApp["fields"]; + defaultModel: string; +}; + +export function listDigestApps(): DigestAppDescriptor[] { + return Object.values(DIGEST_APPS).map((a) => ({ + id: a.id, + label: a.label, + lane: a.lane, + metered: a.metered, + fields: a.fields, + defaultModel: a.defaultModel(), + })); +} diff --git a/common/lib/digestContext-server.ts b/common/lib/digestContext-server.ts @@ -0,0 +1,63 @@ +// Per-channel context notes and the `contextHash` every generated section +// records. +// +// This is the MINIMUM of PLAN.md's Phase 1.5, deliberately: the full +// channel-context feature (structured frontmatter, per-channel entity lists) is +// not built here, but the KEY is plumbed now. PLAN.md's "do not backfill before +// 1.5" trap is exactly this — a digest generated without a contextHash cannot +// tell whether it predates a channel's context note, so adding notes later would +// invalidate the whole corpus. Hashing the (currently usually empty) note from +// day one means adding a note later invalidates only that channel. +// +// The note is a plain Markdown file in the channel dir, hand-authored. Stage B's +// compounding step is writing a correction here — "the co-host is Sam, not Sand" +// — rather than patching individual videos, because a note improves every future +// generation for that channel. + +import path from "node:path"; +import { readFile } from "node:fs/promises"; +import { createHash } from "node:crypto"; +import type { Paths } from "./paths"; + +export const DIGEST_CONTEXT_FILENAME = "digest-context.md"; + +// Cap what reaches the prompt: a note is guidance, not a second transcript, and +// an unbounded one would eat the context window the transcript needs. +const MAX_CONTEXT_CHARS = 4000; + +export type DigestContext = { + // The note text to embed in the prompt. "" when the channel has none. + note: string; + // Stable hash of every context input. Recorded in each section's provenance and + // compared on re-run. The empty-note hash is a real, stable value — NOT "" — + // so "no note" and "note removed" are the same state and neither is confused + // with "generated before contextHash existed" (which reads as "" and is stale). + hash: string; +}; + +export function digestContextPath(paths: Paths, channelSlug: string): string { + return path.join(paths.channelsDir, channelSlug, DIGEST_CONTEXT_FILENAME); +} + +// Hash the context inputs. Versioned by a literal prefix so the hashing scheme +// itself can change later without colliding with old values. +export function hashDigestContext(note: string): string { + return createHash("sha1") + .update(`digest-context-v1\n${note}`) + .digest("hex") + .slice(0, 16); +} + +export async function readDigestContext( + paths: Paths, + channelSlug: string, +): Promise<DigestContext> { + let note = ""; + try { + const raw = await readFile(digestContextPath(paths, channelSlug), "utf8"); + note = raw.trim().slice(0, MAX_CONTEXT_CHARS); + } catch { + note = ""; + } + return { note, hash: hashDigestContext(note) }; +} diff --git a/common/lib/digestParse.test.ts b/common/lib/digestParse.test.ts @@ -0,0 +1,288 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { parseChapters, parseTags, snapToCueStart } from "./digestParse"; +import type { DigestChunkOutput } from "./digestParse"; +import type { Cue } from "./vtt"; + +// Each guard here corresponds to a failure MEASURED on qwen2.5:7b before the +// prompt/schema were hardened, so these tests are regression pins for real +// output, not hypotheticals. + +function cue(start: number, text = `t${start}`): Cue { + return { start, end: start + 5, text }; +} + +// A 20-minute transcript with a cue every 10s. +const CUES: Cue[] = Array.from({ length: 120 }, (_, i) => cue(i * 10)); + +function chunk( + chapters: unknown[], + opts: { index?: number; start?: number; end?: number } = {}, +): DigestChunkOutput { + return { + index: opts.index ?? 0, + startSeconds: opts.start ?? 0, + endSeconds: opts.end ?? 1200, + data: { chapters }, + }; +} + +test("guard 1: rejects a malformed timestamp and records it verbatim", () => { + // ":00:27" is the exact shape the naive prompt produced. + const { chapters, warnings } = parseChapters( + [ + chunk([ + { start: ":00:27", title: "Bad stamp" }, + { start: "00:01:00", title: "Good stamp" }, + ]), + ], + CUES, + ); + assert.deepEqual( + chapters.map((c) => c.title), + ["Good stamp"], + ); + const w = warnings.find((x) => x.code === "malformed-timestamp"); + assert.ok(w, "the rejection is recorded, never silently dropped"); + assert.equal(w?.value, ":00:27"); + assert.equal(w?.section, "chapters"); + assert.equal(w?.chunk, 0); +}); + +test("guard 1: rejects single-digit fields the regex pin forbids", () => { + const { chapters, warnings } = parseChapters( + [chunk([{ start: "1:2:3", title: "Loose stamp" }])], + CUES, + ); + assert.equal(chapters.length, 0); + assert.equal(warnings[0].code, "malformed-timestamp"); +}); + +test("guard 1: rejects an in-shape stamp with an impossible field", () => { + const { chapters, warnings } = parseChapters( + [chunk([{ start: "00:99:00", title: "Ninety-nine minutes" }])], + CUES, + ); + assert.equal(chapters.length, 0); + assert.equal(warnings[0].code, "malformed-timestamp"); +}); + +test("guard 2: flags a title that drifted out of English", () => { + // The measured drift was into Chinese on an 8.2k-token chunk. + const { chapters, warnings } = parseChapters( + [ + chunk([ + { start: "00:00:10", title: "法庭文件截止日期" }, + { start: "00:02:00", title: "Court filing deadlines" }, + ]), + ], + CUES, + ); + assert.deepEqual( + chapters.map((c) => c.title), + ["Court filing deadlines"], + ); + const w = warnings.find((x) => x.code === "language-drift"); + assert.ok(w); + assert.equal(w?.value, "法庭文件截止日期"); +}); + +test("guard 3: clamps to the CHUNK's range, not the video's", () => { + // The measured failure: 01:10:29 emitted for an input spanning 00:04:45 to + // 00:15:36. The video was 176 minutes long, so a whole-video range check would + // have ACCEPTED it. This is why the clamp is per-chunk. + const { chapters, warnings } = parseChapters( + [ + chunk([{ start: "01:10:29", title: "Way past the end" }], { + start: 285, + end: 936, + }), + ], + CUES, + ); + assert.equal(chapters.length, 0); + const w = warnings.find((x) => x.code === "out-of-range"); + assert.ok(w); + assert.equal(w?.value, "01:10:29"); + assert.match(w?.detail ?? "", /00:04:45/); +}); + +test("guard 3: accepts a start inside the chunk's own range", () => { + const { chapters } = parseChapters( + [ + chunk([{ start: "00:05:00", title: "Inside the window" }], { + start: 285, + end: 936, + }), + ], + CUES, + ); + assert.equal(chapters.length, 1); + assert.equal(chapters[0].title, "Inside the window"); +}); + +test("guard 4a: drops a non-monotonic entry within a chunk", () => { + const { chapters, warnings } = parseChapters( + [ + chunk([ + { start: "00:01:00", title: "First" }, + { start: "00:00:30", title: "Backwards" }, + { start: "00:02:00", title: "Third" }, + ]), + ], + CUES, + ); + assert.deepEqual( + chapters.map((c) => c.title), + ["First", "Third"], + ); + assert.ok(warnings.some((w) => w.code === "non-monotonic")); +}); + +test("guard 4b: de-dups equivalent chapters across a chunk seam", () => { + const { chapters, warnings } = parseChapters( + [ + chunk([{ start: "00:05:00", title: "Court filing deadlines" }], { + index: 0, + start: 0, + end: 400, + }), + chunk([{ start: "00:05:10", title: "court filing deadlines." }], { + index: 1, + start: 300, + end: 700, + }), + ], + CUES, + ); + assert.equal(chapters.length, 1, "the overlap's duplicate view is collapsed"); + assert.ok(warnings.some((w) => w.code === "seam-duplicate")); +}); + +test("guard 4b: keeps a genuinely different topic near a seam", () => { + const { chapters } = parseChapters( + [ + chunk([{ start: "00:05:00", title: "Court filing deadlines" }], { + index: 0, + start: 0, + end: 400, + }), + chunk([{ start: "00:05:20", title: "Jury selection" }], { + index: 1, + start: 300, + end: 700, + }), + ], + CUES, + ); + assert.equal(chapters.length, 2); +}); + +test("guard 4b: collapses two chunks that snap onto the same cue", () => { + const { chapters, warnings } = parseChapters( + [ + chunk([{ start: "00:05:00", title: "Alpha" }], { index: 0, end: 700 }), + chunk([{ start: "00:05:02", title: "Completely unrelated beta" }], { + index: 1, + end: 700, + }), + ], + CUES, + ); + assert.equal(chapters.length, 1); + const w = warnings.find((w) => w.code === "seam-duplicate"); + assert.match(w?.detail ?? "", /same start/); +}); + +test("guard 5: snaps a start onto the nearest cue boundary", () => { + // 00:05:04 (304s) sits between cues at 300s and 310s; 300 is nearer. + const { chapters } = parseChapters( + [chunk([{ start: "00:05:04", title: "Between cues" }])], + CUES, + ); + assert.equal(chapters[0].start, 300); + assert.equal(chapters[0].clock, "00:05:04", "the model's own stamp is kept for audit"); + assert.equal(chapters[0].id, "c300", "the id derives from the SNAPPED start"); +}); + +test("guard 5: snaps forward when the later cue is nearer", () => { + const { chapters } = parseChapters( + [chunk([{ start: "00:05:08", title: "Nearer the next cue" }])], + CUES, + ); + assert.equal(chapters[0].start, 310); +}); + +test("snapToCueStart handles a time before the first cue", () => { + assert.equal(snapToCueStart([cue(40), cue(50)], 5), 40); +}); + +test("snapToCueStart handles an empty cue list", () => { + assert.equal(snapToCueStart([], 42.7), 42); +}); + +test("every kept chapter is marked decidedBy: ai", () => { + const { chapters } = parseChapters( + [chunk([{ start: "00:01:00", title: "Machine-decided" }])], + CUES, + ); + assert.equal(chapters[0].decidedBy, "ai"); +}); + +test("a chunk with no chapters array is recorded as parse-failed", () => { + const { chapters, warnings } = parseChapters( + [{ index: 0, startSeconds: 0, endSeconds: 600, data: { nope: true } }], + CUES, + ); + assert.equal(chapters.length, 0); + assert.equal(warnings[0].code, "parse-failed"); +}); + +test("an empty chapters array is recorded as empty-output", () => { + const { warnings } = parseChapters([chunk([])], CUES); + assert.equal(warnings[0].code, "empty-output"); +}); + +test("an entry with no title is recorded, not silently dropped", () => { + const { chapters, warnings } = parseChapters( + [chunk([{ start: "00:01:00", title: " " }])], + CUES, + ); + assert.equal(chapters.length, 0); + assert.equal(warnings[0].code, "empty-title"); +}); + +// --------------------------------------------------------------------------- +// Tags +// --------------------------------------------------------------------------- + +test("parseTags lowercases, dedups across chunks, and keeps ids stable", () => { + const { tags } = parseTags([ + { index: 0, startSeconds: 0, endSeconds: 600, data: { tags: ["Court Filings", "appeals"] } }, + { index: 1, startSeconds: 500, endSeconds: 1200, data: { tags: ["court filings", "sentencing"] } }, + ]); + assert.deepEqual( + tags.map((t) => t.tag), + ["appeals", "court filings", "sentencing"], + ); + assert.equal(tags[1].id, "tcourt-filings"); +}); + +test("parseTags does NOT warn about an expected overlap duplicate", () => { + const { warnings } = parseTags([ + { index: 0, startSeconds: 0, endSeconds: 600, data: { tags: ["appeals"] } }, + { index: 1, startSeconds: 500, endSeconds: 1200, data: { tags: ["appeals"] } }, + ]); + assert.equal(warnings.length, 0); +}); + +test("parseTags flags a drifted tag", () => { + const { tags, warnings } = parseTags([ + { index: 0, startSeconds: 0, endSeconds: 600, data: { tags: ["上訴", "appeals"] } }, + ]); + assert.deepEqual( + tags.map((t) => t.tag), + ["appeals"], + ); + assert.equal(warnings[0].code, "language-drift"); +}); diff --git a/common/lib/digestParse.ts b/common/lib/digestParse.ts @@ -0,0 +1,325 @@ +// The correctness layer between a local 7B model and the artifact on disk. +// +// Every guard here is traceable to a MEASURED failure, not to defensive +// instinct. On a real transcript, before the prompt/schema were hardened, +// qwen2.5:7b produced: a malformed stamp (":00:27"), 11 of 11 starts outside the +// input's range (one 55 minutes past the end), and titles that drifted into +// Chinese. The hardened schema eliminated all three at the decoder — but a parser +// that trusts the decoder is a parser that breaks the day an engine ignores the +// schema, so each rule is enforced here too. +// +// Nothing is ever silently dropped. Every rejection lands in `warnings[]` with +// the offending value, because a multi-week sweep is only tunable if its failures +// are inspectable, and the Phase 11a review queue is built entirely on that array. +// +// Pure (no I/O) — unit-tested in digestParse.test.ts. + +import { + chapterId, + tagId, + type DigestChapter, + type DigestTag, + type DigestWarning, +} from "./digest"; +import { HMS_RE, hmsToSeconds, toHms } from "./digestPrompt"; +import type { Cue } from "./vtt"; + +// One engine call's output, with the range that call was RESPONSIBLE for. The +// range is per-chunk on purpose: a whole-video check would have accepted the +// measured 01:10:29 for a 15-minute input, because the video was 176 minutes long. +export type DigestChunkOutput = { + index: number; + startSeconds: number; + endSeconds: number; + // Whatever the engine returned. Unvalidated by construction. + data: unknown; +}; + +// How far outside its own range a chunk's start may fall before it is rejected. +// Small and non-zero: a cue that begins a hair before the slice's first cue start +// is a rounding artifact, not a hallucination. +const RANGE_TOLERANCE_SECONDS = 2; + +// Two chapters closer than this, with equivalent titles, are the same chapter +// seen through the chunk overlap. Sized to the overlap (40 cues ≈ 2-4 minutes of +// speech) so a genuine topic change at a 3-minute gap survives. +const SEAM_DEDUP_SECONDS = 90; + +// Scripts that mean the model stopped writing English. The prompt and system +// message both demand English titles, so ANY character from one of these is +// drift, not a loanword — and drift is the signal that a chunk confused the +// model, which makes the rest of its output for that chunk suspect too. +const NON_LATIN_RE = + /[\p{Script=Han}\p{Script=Hiragana}\p{Script=Katakana}\p{Script=Hangul}\p{Script=Cyrillic}\p{Script=Arabic}\p{Script=Hebrew}\p{Script=Devanagari}\p{Script=Thai}\p{Script=Greek}]/u; + +export type ParsedChapters = { + chapters: DigestChapter[]; + warnings: DigestWarning[]; +}; + +export type ParsedTags = { + tags: DigestTag[]; + warnings: DigestWarning[]; +}; + +// GUARD 5 — snap a model-emitted second to the nearest real cue boundary. +// +// A chapter must start where someone actually starts speaking, or the viewer +// jumps into the middle of a sentence. Hand-rolled binary search over cue starts, +// the same shape as findActiveIndex in TranscriptModal.tsx. (resolveCitationSeconds +// in the viewer scans SNIPPETS — the ~5 matched lines per video — not cues, so it +// cannot be reused here.) +export function snapToCueStart(cues: Cue[], seconds: number): number { + if (cues.length === 0) return Math.max(0, Math.floor(seconds)); + let lo = 0; + let hi = cues.length - 1; + let found = -1; + while (lo <= hi) { + const mid = (lo + hi) >> 1; + if (cues[mid].start <= seconds) { + found = mid; + lo = mid + 1; + } else { + hi = mid - 1; + } + } + // Before the first cue: the first cue is the only sensible boundary. + if (found < 0) return Math.max(0, Math.floor(cues[0].start)); + const before = cues[found].start; + const after = found + 1 < cues.length ? cues[found + 1].start : null; + if (after !== null && after - seconds < seconds - before) { + return Math.max(0, Math.floor(after)); + } + return Math.max(0, Math.floor(before)); +} + +// Titles are compared for seam de-dup after this normalization, so "Court filing +// deadlines" and "court filing deadlines." collapse. +function normalizeTitle(title: string): string { + return title + .toLowerCase() + .replace(/[^\p{L}\p{N}]+/gu, " ") + .trim(); +} + +function titlesEquivalent(a: string, b: string): boolean { + const na = normalizeTitle(a); + const nb = normalizeTitle(b); + if (!na || !nb) return false; + if (na === nb) return true; + // The overlap frequently yields one call's fuller phrasing of the other's. + return na.includes(nb) || nb.includes(na); +} + +type RawChapter = { start: unknown; title: unknown }; + +function readRawChapters(data: unknown): RawChapter[] | null { + if (!data || typeof data !== "object") return null; + const arr = (data as { chapters?: unknown }).chapters; + if (!Array.isArray(arr)) return null; + return arr as RawChapter[]; +} + +// Parse ONE chunk's chapters, applying guards 1-4 in the order that keeps the +// warnings legible: shape, then range, then language, then monotonicity. +function parseChapterChunk( + chunk: DigestChunkOutput, + cues: Cue[], +): ParsedChapters { + const warnings: DigestWarning[] = []; + const warn = ( + code: DigestWarning["code"], + value?: string, + detail?: string, + ): void => { + warnings.push({ + code, + section: "chapters", + chunk: chunk.index, + ...(value !== undefined ? { value } : {}), + ...(detail !== undefined ? { detail } : {}), + }); + }; + + const raw = readRawChapters(chunk.data); + if (raw === null) { + warn("parse-failed", JSON.stringify(chunk.data ?? null).slice(0, 200)); + return { chapters: [], warnings }; + } + if (raw.length === 0) { + warn("empty-output"); + return { chapters: [], warnings }; + } + + const lo = chunk.startSeconds - RANGE_TOLERANCE_SECONDS; + const hi = chunk.endSeconds + RANGE_TOLERANCE_SECONDS; + const kept: DigestChapter[] = []; + // Monotonicity is checked against the model's EMITTED order, not sorted order: + // a chunk that jumps backwards has lost track of where it is, and that entry is + // the suspect one. + let lastStart = -1; + + for (const entry of raw) { + const rawStart = typeof entry?.start === "string" ? entry.start : ""; + const rawTitle = typeof entry?.title === "string" ? entry.title.trim() : ""; + + // GUARD 1 — timestamp shape. The schema pins this, so a hit here means an + // engine ignored the schema; recording it is how we'd find that out. + if (!HMS_RE.test(rawStart)) { + warn("malformed-timestamp", rawStart || String(entry?.start)); + continue; + } + const seconds = hmsToSeconds(rawStart); + if (seconds === null) { + warn("malformed-timestamp", rawStart); + continue; + } + + if (!rawTitle) { + warn("empty-title", rawStart); + continue; + } + + // GUARD 2 — language drift. + if (NON_LATIN_RE.test(rawTitle)) { + warn("language-drift", rawTitle); + continue; + } + + // GUARD 3 — per-chunk range clamp. + if (seconds < lo || seconds > hi) { + warn( + "out-of-range", + rawStart, + `outside ${toHms(chunk.startSeconds)}–${toHms(chunk.endSeconds)}`, + ); + continue; + } + + // GUARD 4a — monotonic starts within the chunk. + if (seconds <= lastStart) { + warn("non-monotonic", rawStart, `after ${toHms(lastStart)}`); + continue; + } + lastStart = seconds; + + // GUARD 5 — snap onto a real cue boundary. + const snapped = snapToCueStart(cues, seconds); + kept.push({ + id: chapterId(snapped), + start: snapped, + clock: rawStart, + title: rawTitle, + decidedBy: "ai", + }); + } + + return { chapters: kept, warnings }; +} + +// Parse every chunk and merge, applying guard 4b (seam de-dup) across chunk +// boundaries. Chunks are processed in `index` order so "first wins" is stable. +export function parseChapters( + chunks: DigestChunkOutput[], + cues: Cue[], +): ParsedChapters { + const warnings: DigestWarning[] = []; + const all: DigestChapter[] = []; + for (const chunk of [...chunks].sort((a, b) => a.index - b.index)) { + const parsed = parseChapterChunk(chunk, cues); + warnings.push(...parsed.warnings); + all.push(...parsed.chapters); + } + + all.sort((a, b) => a.start - b.start || a.title.localeCompare(b.title)); + const kept: DigestChapter[] = []; + for (const chapter of all) { + const prev = kept[kept.length - 1]; + if (prev) { + // Snapping can land two chunks' views of one moment on the same cue. + if (prev.start === chapter.start) { + warnings.push({ + code: "seam-duplicate", + section: "chapters", + value: chapter.title, + detail: `same start as "${prev.title}"`, + }); + continue; + } + if ( + chapter.start - prev.start <= SEAM_DEDUP_SECONDS && + titlesEquivalent(prev.title, chapter.title) + ) { + warnings.push({ + code: "seam-duplicate", + section: "chapters", + value: chapter.title, + detail: `equivalent to "${prev.title}" ${chapter.start - prev.start}s earlier`, + }); + continue; + } + } + kept.push(chapter); + } + + return { chapters: kept, warnings }; +} + +// --------------------------------------------------------------------------- +// Tags +// --------------------------------------------------------------------------- + +export function parseTags(chunks: DigestChunkOutput[]): ParsedTags { + const warnings: DigestWarning[] = []; + const byId = new Map<string, DigestTag>(); + for (const chunk of [...chunks].sort((a, b) => a.index - b.index)) { + const data = chunk.data; + const arr = + data && typeof data === "object" + ? (data as { tags?: unknown }).tags + : undefined; + if (!Array.isArray(arr)) { + warnings.push({ + code: "parse-failed", + section: "tags", + chunk: chunk.index, + value: JSON.stringify(data ?? null).slice(0, 200), + }); + continue; + } + if (arr.length === 0) { + warnings.push({ code: "empty-output", section: "tags", chunk: chunk.index }); + continue; + } + for (const raw of arr) { + const tag = typeof raw === "string" ? raw.trim().toLowerCase() : ""; + if (!tag) { + warnings.push({ + code: "empty-title", + section: "tags", + chunk: chunk.index, + value: String(raw), + }); + continue; + } + if (NON_LATIN_RE.test(tag)) { + warnings.push({ + code: "language-drift", + section: "tags", + chunk: chunk.index, + value: tag, + }); + continue; + } + const id = tagId(tag); + // Chunks overlap, so the same tag arrives repeatedly — that is expected, + // not a failure, so it is deduped WITHOUT a warning (unlike a chapter seam + // duplicate, which indicates a real segmentation ambiguity). + if (!byId.has(id)) byId.set(id, { id, tag, decidedBy: "ai" }); + } + } + return { + tags: Array.from(byId.values()).sort((a, b) => a.tag.localeCompare(b.tag)), + warnings, + }; +} diff --git a/common/lib/digestPrompt.ts b/common/lib/digestPrompt.ts @@ -0,0 +1,223 @@ +// The prompt + JSON schema the digest engines are driven with, and the ONE +// constant (PROMPT_VERSION) that invalidates generated sections when either +// changes. Pure — no I/O — so it can be unit-tested and imported anywhere. +// +// WHY THE SCHEMA IS THIS STRICT. Measured on a real transcript with qwen2.5:7b: +// +// naive prompt, loose schema → malformed stamps (":00:27"), 11 of 11 starts +// out of range, output drifted to Chinese +// hardened prompt + this → 0 malformed, 0 out of range, 0 drift +// +// Pinning `start` to a full HH:MM:SS regex, stating the chunk's own time range in +// the prompt, and demanding English titles eliminated EVERY correctness failure. +// A 7B model will not voluntarily honor a text contract; schema-constrained +// decoding is what makes the local lane usable. Do not loosen the pattern. +// +// Bump PROMPT_VERSION for any change to the prompt text, the schema, or the +// chunking constants below — a section whose recorded promptVersion differs is +// regenerated, and a section whose version matches is skipped. That is what +// keeps a re-run minutes long instead of weeks. +export const PROMPT_VERSION = 1; + +// Chunking. 12k usable tokens of transcript per call inside a 16k context leaves +// room for the prompt and the response. Expressed in CUES because that is what +// the chunker slices; ~10 tokens/cue is the corpus average, so 1200 cues ≈ 12k +// tokens. The overlap exists so a topic straddling a seam is visible whole to at +// least one call; the parser de-dups the resulting near-identical chapters. +export const DIGEST_MAX_CUES_PER_CHUNK = 1200; +export const DIGEST_OVERLAP_CUES = 40; + +// Segmentation density target. 3 chapters for 22 minutes (the measured +// hardened-prompt result) is too thin to be useful, so the prompt states an +// explicit rate and the schema carries a matching minItems floor. This is the +// open tuning question Stage B exists to settle — it is a knob, not a fact. +export const DIGEST_MINUTES_PER_CHAPTER = 4; +// Never demand more than this from one chunk, however long it is: an unreachable +// minItems floor makes a constrained decoder pad with junk. +export const DIGEST_MAX_CHAPTERS_PER_CHUNK = 24; + +// The regex that eliminated every malformed timestamp. Two digits per field, so +// ":00:27" and "1:2:3" are both rejected by the decoder itself. +export const HMS_PATTERN = "^[0-9][0-9]:[0-9][0-9]:[0-9][0-9]$"; +// The parser's own copy of the same rule (guard 1): the schema pins it, but the +// parser must still reject a bad stamp in case an engine ignores the schema. +export const HMS_RE = /^[0-9][0-9]:[0-9][0-9]:[0-9][0-9]$/; + +export function hmsToSeconds(clock: string): number | null { + if (!HMS_RE.test(clock)) return null; + const [h, m, s] = clock.split(":").map(Number); + if (m > 59 || s > 59) return null; + return h * 3600 + m * 60 + s; +} + +// Zero-padded HH:MM:SS. Deliberately NOT aiHandoff's hms(), which drops the hour +// field for short videos ("2:36") — the schema pattern requires all three fields. +export function toHms(totalSeconds: number): string { + const n = Math.max(0, Math.floor(totalSeconds)); + const h = Math.floor(n / 3600); + const m = Math.floor((n % 3600) / 60); + const s = n % 60; + return [h, m, s].map((v) => String(v).padStart(2, "0")).join(":"); +} + +export function minChaptersForSpan(spanSeconds: number): number { + const byRate = Math.floor(spanSeconds / 60 / DIGEST_MINUTES_PER_CHAPTER); + return Math.max(1, Math.min(DIGEST_MAX_CHAPTERS_PER_CHUNK, byRate)); +} + +export function maxChaptersForSpan(spanSeconds: number): number { + return Math.max( + minChaptersForSpan(spanSeconds) + 2, + Math.min( + DIGEST_MAX_CHAPTERS_PER_CHUNK, + Math.ceil(spanSeconds / 60 / Math.max(1, DIGEST_MINUTES_PER_CHAPTER - 2)), + ), + ); +} + +// --------------------------------------------------------------------------- +// Chapters +// --------------------------------------------------------------------------- + +export type ChapterPromptInput = { + title: string; + channel: string; + // The chunk's OWN range, in seconds. Stating it in the prompt is one of the + // three changes that fixed out-of-range output. + startSeconds: number; + endSeconds: number; + // The transcript slice, already rendered as `[HH:MM:SS] text` lines by + // transcriptToMarkdown with stampForCue. + transcript: string; + // Optional per-channel context note (Phase 1.5). Plumbed from the start so + // adding notes later doesn't invalidate the corpus — see contextHash. + contextNote?: string; +}; + +export function chapterSchema(spanSeconds: number): Record<string, unknown> { + return { + type: "object", + properties: { + chapters: { + type: "array", + minItems: minChaptersForSpan(spanSeconds), + maxItems: maxChaptersForSpan(spanSeconds), + items: { + type: "object", + properties: { + // The pin. Every correctness failure measured before this existed. + start: { type: "string", pattern: HMS_PATTERN }, + title: { type: "string", minLength: 3, maxLength: 90 }, + }, + required: ["start", "title"], + }, + }, + }, + required: ["chapters"], + }; +} + +export const CHAPTER_SYSTEM_PROMPT = [ + "You segment transcripts into chapters.", + "You reply with JSON only, matching the provided schema exactly.", + "Every title you write is in ENGLISH, regardless of the transcript's language.", + "You never invent a timestamp: every start you emit is copied from a", + "[HH:MM:SS] marker that appears in the transcript you were given.", +].join(" "); + +export function buildChapterPrompt(input: ChapterPromptInput): string { + const from = toHms(input.startSeconds); + const to = toHms(input.endSeconds); + const span = Math.max(0, input.endSeconds - input.startSeconds); + const minItems = minChaptersForSpan(span); + const lines: string[] = []; + + lines.push( + `Below is one section of the transcript of "${input.title}" (${input.channel}).`, + ); + lines.push(""); + lines.push( + `This section covers ${from} to ${to} — ${Math.round(span / 60)} minutes of material.`, + ); + lines.push( + `EVERY start you emit MUST be between ${from} and ${to} inclusive. A start outside`, + `that range is wrong even if the topic is real. Copy starts from the [HH:MM:SS]`, + "markers in the transcript; do not compute or estimate them.", + ); + lines.push(""); + lines.push( + `Aim for roughly one chapter per ${DIGEST_MINUTES_PER_CHAPTER} minutes of material —`, + `at least ${minItems} for this section. A chapter marks where the subject genuinely`, + "changes; do not split one continuous discussion into several chapters, and do not", + "merge unrelated subjects into one.", + ); + lines.push(""); + lines.push( + "Each title is a specific, concrete English noun phrase naming what is discussed", + '(e.g. "Court filing deadlines" — not "Discussion" or "Part two"). Do not use the', + "speaker's own words as a quote, and do not editorialize.", + ); + if (input.contextNote?.trim()) { + lines.push(""); + lines.push("Context for this channel (use it for names and recurring topics):"); + lines.push(input.contextNote.trim()); + } + lines.push(""); + lines.push("Transcript section:"); + lines.push(""); + lines.push(input.transcript); + return lines.join("\n"); +} + +// --------------------------------------------------------------------------- +// Tags +// --------------------------------------------------------------------------- + +export const TAG_MIN_ITEMS = 3; +export const TAG_MAX_ITEMS = 12; + +export function tagSchema(): Record<string, unknown> { + return { + type: "object", + properties: { + tags: { + type: "array", + minItems: TAG_MIN_ITEMS, + maxItems: TAG_MAX_ITEMS, + items: { type: "string", minLength: 2, maxLength: 40 }, + }, + }, + required: ["tags"], + }; +} + +export const TAG_SYSTEM_PROMPT = [ + "You extract topic tags from transcripts.", + "You reply with JSON only, matching the provided schema exactly.", + "Every tag is in ENGLISH, lowercase, and is a topic — not a sentence,", + "not a summary, and not a person's opinion of the topic.", +].join(" "); + +export function buildTagPrompt(input: ChapterPromptInput): string { + const lines: string[] = []; + lines.push( + `Below is one section of the transcript of "${input.title}" (${input.channel}),`, + `covering ${toHms(input.startSeconds)} to ${toHms(input.endSeconds)}.`, + ); + lines.push(""); + lines.push( + `List between ${TAG_MIN_ITEMS} and ${TAG_MAX_ITEMS} lowercase English topic tags for`, + "what this section is ABOUT. Prefer the specific over the generic: name the", + "subject, event, or field, not the format of the video.", + ); + if (input.contextNote?.trim()) { + lines.push(""); + lines.push("Context for this channel:"); + lines.push(input.contextNote.trim()); + } + lines.push(""); + lines.push("Transcript section:"); + lines.push(""); + lines.push(input.transcript); + return lines.join("\n"); +} diff --git a/common/lib/duplicates.test.ts b/common/lib/duplicates.test.ts @@ -0,0 +1,273 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + clusterMaySharePartial, + isClusterReviewed, + measureAlignment, + pickCanonicalSlug, + resolveCanonicalSlug, + sanitizeDuplicateOverrides, + type DuplicateCluster, + type DuplicateVideoRef, +} from "./duplicates"; +import type { Cue } from "./vtt"; + +function ref(over: Partial<DuplicateVideoRef> = {}): DuplicateVideoRef { + return { + slug: "chan/vid", + channelSlug: "chan", + channel: "Chan", + platform: "youtube", + id: "vid", + title: "A video", + duration: 600, + uploadDate: "20250101", + hasTranscript: true, + ...over, + }; +} + +function cluster( + refs: DuplicateVideoRef[], + over: Partial<DuplicateCluster> = {}, +): DuplicateCluster { + return { + clusterId: "cluster1", + matchKind: "transcript-exact", + score: 1, + contained: false, + durationBucket: 600, + crossPlatform: true, + crossChannel: false, + videoRefs: refs, + ...over, + }; +} + +// --------------------------------------------------------------------------- +// Canonical selection +// --------------------------------------------------------------------------- + +test("a member with a transcript beats one without", () => { + const c = cluster([ + ref({ slug: "a/1", hasTranscript: false, duration: 900 }), + ref({ slug: "b/2", hasTranscript: true, duration: 600 }), + ]); + assert.equal(pickCanonicalSlug(c), "b/2"); +}); + +test("the longest (most complete) member wins next", () => { + // Load-bearing: a mirror that cut the intro would place every shared chapter + // wrong, so the fullest artifact owns the digest. + const c = cluster([ + ref({ slug: "a/1", duration: 600 }), + ref({ slug: "b/2", duration: 640 }), + ]); + assert.equal(pickCanonicalSlug(c), "b/2"); +}); + +test("the preferred platform breaks a duration tie", () => { + const c = cluster([ + ref({ slug: "a/1", platform: "rumble" }), + ref({ slug: "b/2", platform: "youtube" }), + ]); + assert.equal(pickCanonicalSlug(c), "b/2"); +}); + +test("the earliest upload breaks a platform tie", () => { + const c = cluster([ + ref({ slug: "a/1", uploadDate: "20250301" }), + ref({ slug: "b/2", uploadDate: "20250101" }), + ]); + assert.equal(pickCanonicalSlug(c), "b/2"); +}); + +test("selection is deterministic when everything ties", () => { + const refs = [ref({ slug: "b/2" }), ref({ slug: "a/1" })]; + assert.equal(pickCanonicalSlug(cluster(refs)), "a/1"); + assert.equal(pickCanonicalSlug(cluster([...refs].reverse())), "a/1"); +}); + +test("a human canonical override WINS over the rule", () => { + const c = cluster([ + ref({ slug: "a/1", duration: 600 }), + ref({ slug: "b/2", duration: 900 }), + ]); + assert.equal(pickCanonicalSlug(c), "b/2", "the rule prefers the longer one"); + const overrides = sanitizeDuplicateOverrides({ + clusters: { cluster1: { canonicalSlug: "a/1" } }, + }); + assert.equal(resolveCanonicalSlug(c, overrides), "a/1"); +}); + +test("an override naming a non-member is ignored, not obeyed", () => { + const c = cluster([ref({ slug: "a/1" }), ref({ slug: "b/2" })]); + const overrides = sanitizeDuplicateOverrides({ + clusters: { cluster1: { canonicalSlug: "someone/else" } }, + }); + assert.equal(resolveCanonicalSlug(c, overrides), pickCanonicalSlug(c)); +}); + +test("notDuplicate suppresses the cluster entirely (nothing is shared)", () => { + const c = cluster([ref({ slug: "a/1" }), ref({ slug: "b/2" })]); + const overrides = sanitizeDuplicateOverrides({ + clusters: { cluster1: { notDuplicate: true } }, + }); + assert.equal(resolveCanonicalSlug(c, overrides), null); +}); + +test("a recorded canonicalSlug on the report is honored over the rule", () => { + const c = cluster( + [ref({ slug: "a/1", duration: 600 }), ref({ slug: "b/2", duration: 900 })], + { canonicalSlug: "a/1" }, + ); + assert.equal(resolveCanonicalSlug(c, null), "a/1"); +}); + +test("isClusterReviewed distinguishes a decision from a bare stub", () => { + const c = cluster([ref({ slug: "a/1" }), ref({ slug: "b/2" })]); + assert.equal(isClusterReviewed(c, null), false); + assert.equal( + isClusterReviewed( + c, + sanitizeDuplicateOverrides({ clusters: { cluster1: { canonicalSlug: "a/1" } } }), + ), + true, + ); + // A decidedAt-only entry is dropped by the sanitizer, so it never reads as + // reviewed. + assert.equal( + isClusterReviewed( + c, + sanitizeDuplicateOverrides({ clusters: { cluster1: { decidedAt: "now" } } }), + ), + false, + ); +}); + +test("sanitizeDuplicateOverrides drops malformed entries rather than throwing", () => { + const o = sanitizeDuplicateOverrides({ + clusters: { + good: { canonicalSlug: "a/1" }, + blank: { canonicalSlug: " " }, + junk: 42, + arrayish: [], + }, + }); + assert.deepEqual(Object.keys(o.clusters), ["good"]); +}); + +test("sanitizeDuplicateOverrides survives a non-object file", () => { + assert.deepEqual(sanitizeDuplicateOverrides(null).clusters, {}); + assert.deepEqual(sanitizeDuplicateOverrides([1, 2, 3]).clusters, {}); +}); + +// --------------------------------------------------------------------------- +// The alignment gate +// --------------------------------------------------------------------------- + +// A synthetic transcript whose sentences do NOT share long phrases, so the +// 8-word anchors are unambiguous — which is what the gate requires and what a +// real transcript mostly provides. (An earlier fixture repeated the same nine +// words in every cue; the gate correctly refused to measure it, which is how the +// ambiguous-anchor rule got written.) +const VOCAB = [ + "filing", "deadline", "jury", "selection", "witness", "testimony", "exhibit", + "objection", "sustained", "overruled", "docket", "motion", "dismissal", + "appeal", "verdict", "sentencing", "transcript", "counsel", "recess", + "subpoena", "affidavit", "discovery", "deposition", "settlement", "mediation", + "injunction", "damages", "liability", "negligence", "statute", "precedent", + "jurisdiction", "venue", "indictment", "arraignment", "plea", "bail", + "custody", "warrant", "evidence", +]; + +function makeCues(shiftSeconds = 0, count = 40): Cue[] { + return Array.from({ length: count }, (_, i) => { + // Each cue draws a distinct rotation of the vocabulary, so no 8-word window + // recurs anywhere in the transcript. + const words = Array.from( + { length: 12 }, + (_, j) => VOCAB[(i * 7 + j * 3) % VOCAB.length], + ); + return { + start: i * 15 + shiftSeconds, + end: i * 15 + 14 + shiftSeconds, + text: `${words.join(" ")} marker${i}`, + }; + }); +} + +test("identical timings are aligned at a zero offset", () => { + const a = makeCues(0); + const result = measureAlignment(a, makeCues(0)); + assert.equal(result.aligned, true); + assert.equal(result.maxOffsetSeconds, 0); + assert.equal(result.matchedAnchors, result.totalAnchors); +}); + +test("a mirror with a shifted intro is REFUSED", () => { + // The failure mode that looks like success: identical text, every chapter + // placed 40s wrong. Content similarity cannot see this — shingles are a set. + const result = measureAlignment(makeCues(0), makeCues(40)); + assert.equal(result.aligned, false); + assert.equal(result.reason, "offset-exceeded"); + assert.ok(result.maxOffsetSeconds >= 39 && result.maxOffsetSeconds <= 41); +}); + +test("a sub-tolerance shift is still aligned", () => { + const result = measureAlignment(makeCues(0), makeCues(2)); + assert.equal(result.aligned, true); +}); + +test("the tolerance is configurable", () => { + assert.equal( + measureAlignment(makeCues(0), makeCues(8), { toleranceSeconds: 10 }).aligned, + true, + ); + assert.equal( + measureAlignment(makeCues(0), makeCues(8), { toleranceSeconds: 3 }).aligned, + false, + ); +}); + +test("unrelated transcripts are refused for want of anchors", () => { + const other: Cue[] = Array.from({ length: 40 }, (_, i) => ({ + start: i * 15, + end: i * 15 + 14, + text: `zebra${i} quartz${i} lantern${i} beacon${i} pumice${i} sorrel${i} thicket${i} vellum${i} wicket${i}`, + })); + const result = measureAlignment(makeCues(0), other); + assert.equal(result.aligned, false); + assert.equal(result.reason, "too-few-anchors"); +}); + +test("a mirror aligned at the start but drifting mid-way is refused", () => { + // Sampling SEVERAL anchors is what catches an ad break inserted in the middle. + const canonical = makeCues(0, 40); + const drifting = canonical.map((c, i) => ({ + ...c, + start: i < 20 ? c.start : c.start + 45, + end: i < 20 ? c.end : c.end + 45, + })); + const result = measureAlignment(canonical, drifting); + assert.equal(result.aligned, false); + assert.equal(result.reason, "offset-exceeded"); +}); + +test("an empty transcript is refused, not treated as aligned", () => { + assert.equal(measureAlignment([], makeCues(0)).aligned, false); + assert.equal(measureAlignment(makeCues(0), []).reason, "empty-transcript"); +}); + +test("a contained cluster (clip of a longer video) never shares", () => { + // Correct even at a perfect zero offset: a clip is a different artifact, and + // the longer video's chapters describe material it does not contain. + assert.equal( + clusterMaySharePartial(cluster([ref()], { contained: true })), + false, + ); + assert.equal( + clusterMaySharePartial(cluster([ref()], { contained: false })), + true, + ); +}); diff --git a/common/lib/duplicates.ts b/common/lib/duplicates.ts @@ -50,6 +50,12 @@ export type DuplicateCluster = { crossPlatform: boolean; // members span more than one platform crossChannel: boolean; // members span more than one channelSlug videoRefs: DuplicateVideoRef[]; + // The member that OWNS derived work for this cluster: the one an AI digest is + // generated for, and the one aligned mirrors copy it from. Chosen by + // pickCanonicalSlug() at detection time and human-overridable afterwards + // (see DuplicateOverrides). Optional: reports written before this field + // existed lack it, so readers fall back to pickCanonicalSlug(). + canonicalSlug?: string; }; export type DuplicateRunConfig = { @@ -75,6 +81,337 @@ export type DuplicateReport = { export const DUPLICATES_FILENAME = "duplicates.json"; // --------------------------------------------------------------------------- +// Human review: canonical choice + not-a-duplicate, kept OUT of the report +// --------------------------------------------------------------------------- + +// duplicates.json is regenerated wholesale by every detection run, so a human +// decision recorded in it would be destroyed on the next run. It therefore lives +// in a sibling override file — the same separation, for the same reason, as +// ai-digest.overrides.json vs ai-digest.json. +export const DUPLICATE_OVERRIDES_FILENAME = "duplicates.overrides.json"; +export const DUPLICATE_OVERRIDES_VERSION = 1; + +export type DuplicateClusterOverride = { + // Operator's choice of canonical member (a `${channelSlug}/${id}` slug). Wins + // over the rule in pickCanonicalSlug. Ignored when the slug is not a member. + canonicalSlug?: string; + // "These are not the same video." Suppresses the cluster entirely: it stops + // being offered for review AND stops sharing derived work. The detector will + // keep finding it (content really is similar), which is exactly why the + // decision has to be recorded outside the report. + notDuplicate?: boolean; + decidedAt?: string; + note?: string; +}; + +export type DuplicateOverrides = { + version: number; + // Keyed by clusterId — a stable sha1 over the sorted member slugs, so the key + // survives re-detection as long as the membership does. A cluster that GAINS a + // member gets a new id and returns to review, which is the honest behavior: + // the canonical choice was made over a different set of videos. + clusters: Record<string, DuplicateClusterOverride>; +}; + +export function emptyDuplicateOverrides(): DuplicateOverrides { + return { version: DUPLICATE_OVERRIDES_VERSION, clusters: {} }; +} + +// Coerce a raw overrides file, dropping ill-typed entries rather than throwing — +// one hand-edit typo must not break the duplicates page. +export function sanitizeDuplicateOverrides(value: unknown): DuplicateOverrides { + const out = emptyDuplicateOverrides(); + if (!value || typeof value !== "object") return out; + const raw = (value as { clusters?: unknown }).clusters; + if (!raw || typeof raw !== "object" || Array.isArray(raw)) return out; + for (const [clusterId, entry] of Object.entries(raw as Record<string, unknown>)) { + if (!entry || typeof entry !== "object") continue; + const e = entry as Record<string, unknown>; + const override: DuplicateClusterOverride = {}; + if (typeof e.canonicalSlug === "string" && e.canonicalSlug.trim()) { + override.canonicalSlug = e.canonicalSlug.trim(); + } + if (e.notDuplicate === true) override.notDuplicate = true; + if (typeof e.decidedAt === "string") override.decidedAt = e.decidedAt; + if (typeof e.note === "string" && e.note.trim()) override.note = e.note.trim(); + if (Object.keys(override).length === 0) continue; + out.clusters[clusterId] = override; + } + return out; +} + +// --------------------------------------------------------------------------- +// Canonical member selection +// --------------------------------------------------------------------------- + +// Platform preference for the canonical member, most-preferred first. YouTube +// leads because its videos carry the richest metadata and the most reliable +// caption tracks, so a digest generated there is the best one to share. +const PLATFORM_PREFERENCE: ReadonlyArray<string> = [ + "youtube", + "rumble", + "odysee", + "kick", + "twitch", +]; + +function platformRank(platform: string): number { + const i = PLATFORM_PREFERENCE.indexOf(platform); + return i < 0 ? PLATFORM_PREFERENCE.length : i; +} + +// Pick the cluster member that should own derived work, by rule. Ordered by +// what actually makes a shared digest good: +// 1. has a transcript — you cannot digest a video without one +// 2. longest duration — the most complete artifact; a mirror that cuts +// the intro would place every shared chapter wrong +// 3. preferred platform — richest metadata +// 4. earliest upload — the original, where the same content appears twice +// 5. slug — a total order, so the choice is deterministic +// Deterministic and pure: no Date.now(), no Math.random(), so re-detection over +// unchanged inputs picks the same member. +export function pickCanonicalSlug(cluster: DuplicateCluster): string { + const refs = cluster.videoRefs; + if (refs.length === 0) return ""; + const best = refs.reduce((a, b) => (canonicalBetter(b, a) ? b : a)); + return best.slug; +} + +function canonicalBetter( + candidate: DuplicateVideoRef, + incumbent: DuplicateVideoRef, +): boolean { + if (candidate.hasTranscript !== incumbent.hasTranscript) { + return candidate.hasTranscript; + } + if (candidate.duration !== incumbent.duration) { + return candidate.duration > incumbent.duration; + } + const pc = platformRank(candidate.platform); + const pi = platformRank(incumbent.platform); + if (pc !== pi) return pc < pi; + if (candidate.uploadDate !== incumbent.uploadDate) { + // Empty upload dates sort last so a dated member wins over an undated one. + if (!candidate.uploadDate) return false; + if (!incumbent.uploadDate) return true; + return candidate.uploadDate < incumbent.uploadDate; + } + return candidate.slug.localeCompare(incumbent.slug) < 0; +} + +// The effective canonical member: the human choice when it names a real member, +// otherwise the recorded one, otherwise the rule. Returns null for a cluster a +// human marked not-a-duplicate — such a cluster shares nothing. +export function resolveCanonicalSlug( + cluster: DuplicateCluster, + overrides?: DuplicateOverrides | null, +): string | null { + const override = overrides?.clusters[cluster.clusterId]; + if (override?.notDuplicate) return null; + const members = new Set(cluster.videoRefs.map((r) => r.slug)); + if (override?.canonicalSlug && members.has(override.canonicalSlug)) { + return override.canonicalSlug; + } + if (cluster.canonicalSlug && members.has(cluster.canonicalSlug)) { + return cluster.canonicalSlug; + } + return pickCanonicalSlug(cluster) || null; +} + +// Whether a human has recorded a decision for this cluster. Drives the +// /actionable "awaiting review" list — a cluster with no decision is work. +export function isClusterReviewed( + cluster: DuplicateCluster, + overrides?: DuplicateOverrides | null, +): boolean { + const o = overrides?.clusters[cluster.clusterId]; + if (!o) return false; + return o.notDuplicate === true || Boolean(o.canonicalSlug); +} + +// --------------------------------------------------------------------------- +// The timestamp-alignment gate — the correctness crux of digest sharing +// --------------------------------------------------------------------------- + +// Content similarity does NOT imply timing alignment, and this is the failure +// mode that looks like success: a mirror with a 40-second-longer intro has +// matching text at shifted times, so a shared digest places EVERY chapter wrong +// while looking perfectly plausible. Nothing in the 5-gram Jaccard cascade +// notices, because shingles are a set — order and position are discarded. +// +// So before sharing we measure it directly: sample several anchor phrases spread +// through the canonical transcript, find where each occurs in the mirror, and +// require every offset to be near zero. + +// Max |offset| (seconds) at any anchor for two transcripts to count as aligned. +// Tight on purpose: a chapter placed 5s early still lands in the right sentence, +// 15s does not. +export const DEFAULT_ALIGNMENT_TOLERANCE_SECONDS = 5; +// How many anchors to sample. Several, spread out, because a mirror can share a +// start time and then diverge at an ad break in the middle. +export const DEFAULT_ALIGNMENT_ANCHORS = 5; +// An anchor must be this many words long to be a reliable locator; shorter +// phrases recur. +const ANCHOR_WORDS = 8; +// How far forward to look for a USABLE anchor when the phrase at the sampled +// position is ambiguous (occurs more than once on either side). Real transcripts +// repeat themselves — intros, catchphrases, ad reads — so without this the gate +// refuses perfectly aligned mirrors and the sharing optimisation never fires. +const ANCHOR_SEARCH_WORDS = 400; + +export type AlignmentResult = { + aligned: boolean; + // Largest |offset| observed across the matched anchors, in seconds. + maxOffsetSeconds: number; + // How many anchors were located in the other transcript. Too few and the + // measurement is not trustworthy, so `aligned` is false regardless of offset. + matchedAnchors: number; + totalAnchors: number; + reason?: + | "too-few-anchors" + | "offset-exceeded" + | "empty-transcript" + | "contained"; +}; + +function normalizeWords(cues: Cue[]): { word: string; start: number }[] { + const out: { word: string; start: number }[] = []; + for (const cue of cues) { + const words = cue.text + .toLowerCase() + .replace(/[^\p{L}\p{N}\s]/gu, " ") + .split(/\s+/) + .filter(Boolean); + for (const word of words) out.push({ word, start: cue.start }); + } + return out; +} + +// Word n-gram -> { first start time, occurrence count }. The count is what makes +// an ambiguous phrase skippable instead of silently mismatched. +function buildAnchorIndex( + words: { word: string; start: number }[], +): Map<string, { start: number; count: number }> { + const index = new Map<string, { start: number; count: number }>(); + for (let i = 0; i + ANCHOR_WORDS <= words.length; i++) { + const key = words + .slice(i, i + ANCHOR_WORDS) + .map((w) => w.word) + .join(" "); + const existing = index.get(key); + if (existing) existing.count++; + else index.set(key, { start: words[i].start, count: 1 }); + } + return index; +} + +// Measure timing alignment between two cue lists. Pure, so the gate is directly +// unit-testable (and it IS tested: a shifted-intro mirror must fail). +export function measureAlignment( + canonical: Cue[], + other: Cue[], + opts: { + toleranceSeconds?: number; + anchors?: number; + } = {}, +): AlignmentResult { + const tolerance = + opts.toleranceSeconds ?? DEFAULT_ALIGNMENT_TOLERANCE_SECONDS; + const anchorCount = Math.max(1, opts.anchors ?? DEFAULT_ALIGNMENT_ANCHORS); + + const a = normalizeWords(canonical); + const b = normalizeWords(other); + if (a.length < ANCHOR_WORDS || b.length < ANCHOR_WORDS) { + return { + aligned: false, + maxOffsetSeconds: Infinity, + matchedAnchors: 0, + totalAnchors: 0, + reason: "empty-transcript", + }; + } + + // Index BOTH sides' word n-grams, counting occurrences — not just recording the + // first. A phrase that occurs twice is not a locator: matching it to whichever + // copy came first would report a bogus offset, which fails in both directions + // (a false refusal wastes a generation; a false match shares a misplaced + // digest). So an ambiguous phrase is skipped rather than guessed at. + const indexB = buildAnchorIndex(b); + const indexA = buildAnchorIndex(a); + + // Anchors spread evenly through the canonical transcript, skipping the very + // start and end (intros/outros are exactly where mirrors differ). + const usable = a.length - ANCHOR_WORDS; + const positions: number[] = []; + for (let k = 1; k <= anchorCount; k++) { + positions.push(Math.floor((usable * k) / (anchorCount + 1))); + } + + let matched = 0; + let maxOffset = 0; + const usedTargets = new Set<number>(); + for (const pos of positions) { + // Walk forward from the sampled position until a phrase is unique on BOTH + // sides. Bounded, so a pathologically repetitive stretch just yields no + // anchor here rather than scanning the whole transcript. + const limit = Math.min(usable, pos + ANCHOR_SEARCH_WORDS); + for (let i = pos; i <= limit; i++) { + const key = a + .slice(i, i + ANCHOR_WORDS) + .map((w) => w.word) + .join(" "); + if ((indexA.get(key)?.count ?? 0) !== 1) continue; + const hit = indexB.get(key); + if (!hit || hit.count !== 1) continue; + // Don't let two sampled positions collapse onto the same anchor — that + // would report "2 anchors matched" from one measurement. + if (usedTargets.has(hit.start)) break; + usedTargets.add(hit.start); + matched++; + maxOffset = Math.max(maxOffset, Math.abs(hit.start - a[i].start)); + break; + } + } + + // Require a majority of anchors to be located: a mirror we can only match in + // one place is not a mirror we can trust timings from. + const enough = matched >= Math.ceil(positions.length / 2); + if (!enough) { + return { + aligned: false, + maxOffsetSeconds: matched === 0 ? Infinity : maxOffset, + matchedAnchors: matched, + totalAnchors: positions.length, + reason: "too-few-anchors", + }; + } + if (maxOffset > tolerance) { + return { + aligned: false, + maxOffsetSeconds: maxOffset, + matchedAnchors: matched, + totalAnchors: positions.length, + reason: "offset-exceeded", + }; + } + return { + aligned: true, + maxOffsetSeconds: maxOffset, + matchedAnchors: matched, + totalAnchors: positions.length, + }; +} + +// Whether a cluster may share derived work AT ALL, before any timing is measured. +// A `contained` cluster matched by containment — one member is a CLIP of a longer +// video, not a mirror of it. A clip is a different artifact: the longer video's +// chapters describe material the clip does not contain, so sharing wholesale +// would be wrong even at a perfect zero offset. +export function clusterMaySharePartial(cluster: DuplicateCluster): boolean { + return !cluster.contained; +} + +// --------------------------------------------------------------------------- // Pure helpers (no I/O) — exported for direct testing. // --------------------------------------------------------------------------- diff --git a/common/lib/parseStdoutJson.ts b/common/lib/parseStdoutJson.ts @@ -0,0 +1,58 @@ +// Parse the JSON a CLI printed on stdout, scanning lines LAST-TO-FIRST. +// +// Lifted from checkAvailability.ts (where it read yt-dlp --dump-json) because +// the digest registry needs exactly the same tolerance for `claude -p +// --output-format json`: both tools may print warnings, deprecation notices, or +// progress to stdout BEFORE the payload, so the JSON is the last parseable line, +// not the first. Reading forward finds the warning; reading backward finds the +// answer. +export function parseStdoutJson(stdout: string): unknown | null { + const lines = stdout + .split("\n") + .map((l) => l.trim()) + .filter(Boolean); + for (let i = lines.length - 1; i >= 0; i--) { + try { + return JSON.parse(lines[i]); + } catch { + continue; + } + } + // Fall back to the whole buffer: `--output-format json` pretty-prints across + // several lines, so no single line parses on its own. + try { + return JSON.parse(stdout); + } catch { + return null; + } +} + +// Pull a JSON object out of model prose. A CLI lane has no constrained decoding, +// so the body may arrive fenced (```json … ```) or with a sentence in front of +// it. Tries the whole string, then the fenced block, then the outermost braces. +export function extractJsonObject(text: string): unknown | null { + const trimmed = text.trim(); + try { + return JSON.parse(trimmed); + } catch { + // fall through + } + const fence = trimmed.match(/```(?:json)?\s*([\s\S]*?)```/); + if (fence) { + try { + return JSON.parse(fence[1].trim()); + } catch { + // fall through + } + } + const first = trimmed.indexOf("{"); + const last = trimmed.lastIndexOf("}"); + if (first >= 0 && last > first) { + try { + return JSON.parse(trimmed.slice(first, last + 1)); + } catch { + return null; + } + } + return null; +} diff --git a/common/lib/paths.ts b/common/lib/paths.ts @@ -108,6 +108,16 @@ export type Paths = { parakeetBin: string; parakeetCliBin: string; parakeetModel: string; + // Base URL of the local ollama server, the local-GPU digest lane + // (common/lib/digestApps.ts posts to `${ollamaUrl}/api/chat`). The FIRST + // URL-valued entry in Paths, so it is normalized here the way + // remoteTranscribe.ts normalizes a worker base: trailing slashes stripped, so + // callers can always concatenate a leading-slash path. + ollamaUrl: string; + // The `claude` CLI, driving the opt-in metered digest lane (off by default — + // settings.digest.remoteEnabled). Not bundled; install it separately and point + // CLAUDE_BIN at it if it isn't on PATH. + claudeBin: string; }; let cached: Paths | null = null; @@ -190,6 +200,11 @@ export function getPaths(): Paths { path.join(monorepoRoot, "scripts", "parakeet-stitch.mjs"), parakeetCliBin: process.env.PARAKEET_CLI ?? "parakeet-cli", parakeetModel: process.env.PARAKEET_MODEL ?? "", + ollamaUrl: (process.env.OLLAMA_URL ?? "http://127.0.0.1:11434").replace( + /\/+$/, + "", + ), + claudeBin: process.env.CLAUDE_BIN ?? "claude", }; return cached; } diff --git a/common/lib/queueKeys.ts b/common/lib/queueKeys.ts @@ -7,6 +7,17 @@ import { export { TRANSCRIPTION_QUEUE }; +// The two digest lanes get SEPARATE queue keys, and that separation is the whole +// point. registry.ts submits every non-empty queueKey with concurrency 1, so: +// - one shared digest key would serialize the lanes, wasting the network lane +// while the GPU works (and vice versa) across a multi-week sweep; +// - putting the local lane on TRANSCRIPTION_QUEUE would let ollama and the +// transcription engine thrash the same 8 GB of VRAM. +// (A queueKey of "" means "run immediately, untracked" — never right for either +// of these, both of which must be serialized against themselves.) +export const DIGEST_LOCAL_QUEUE = "digest:local"; +export const DIGEST_REMOTE_QUEUE = "digest:remote"; + // Per-channel queue for channel-local bookkeeping jobs (clean/clear/verify). export function channelQueueKey(slug: string): string { return `channel:${slug}`; diff --git a/common/lib/settings.ts b/common/lib/settings.ts @@ -29,6 +29,16 @@ import { isCookieMode, type CookieMode, } from "./cookiePolicy"; +// From the CLIENT-SAFE digest module, deliberately — digestApps.ts imports execa, +// and settings.ts must stay reachable from anywhere. +import { + CLAUDE_DIGEST_APP_ID, + DEFAULT_DIGEST_APP_ID, + DIGEST_SECTION_KINDS, + isDigestSectionKind, + type DigestAppConfig, + type DigestSectionKind, +} from "./digest"; export type { Worker } from "./workers"; export type { AutoQueueSettings } from "../jobs/autoQueuePolicy"; @@ -175,6 +185,39 @@ export type SiteSettings = { // pipeline itself is a follow-up; this block persists the chosen mode plus the // container/concurrency knobs the deploy page and the future orchestrator read. buildPipeline: BuildPipelineSettings; + // AI digest generation (chapters + topic tags over the existing transcripts). + // Local-first: the metered lane is off by default. See DigestSettings. + digest: DigestSettings; +}; + +// Configuration for the derived-corpus digest layer. Local-first by decision: +// `remoteEnabled` gates the metered lane and defaults to false, so nothing here +// can spend money until it is explicitly turned on. +export type DigestSettings = { + // Master switch for the metered (remote-api) lane. OFF by default — an opt-in + // overflow for the long tail or a channel where local quality is poor, never + // the default path. + remoteEnabled: boolean; + // Videos longer than this are "long tail": 8.2% of the corpus by count, 46% of + // all transcript tokens. The batch's duration-aware ordering and the optional + // remote overflow both key off it. + longTailSeconds: number; + // The engine each lane uses (ids from common/lib/digestApps.ts). + localAppId: string; + remoteAppId: string; + // Per-app config, keyed by app id — the same id-keyed sub-record shape as + // transcriptionApps. + apps: Record<string, DigestAppConfig>; + // Global pause. Read at DISPATCH time by the batch (the downloadsPaused + // pattern), so a pause survives a restart with no boot hook — unlike + // transcriptionsPaused, which needs editor/instrumentation.ts to re-apply it. + digestsPaused: boolean; + // Hard ceiling on cumulative metered spend per job, USD. 0 = no cap. Only ever + // consulted for a metered app. + spendCapUsd: number; + // Which sections a sweep generates. Chapters alone is the default: tags double + // the call count for a smaller payoff. + sections: DigestSectionKind[]; }; // "basic" — `pnpm run build` in export/, serialized on the build queue (shared @@ -427,6 +470,97 @@ export function sanitizeSavedVideoBackup( }; } +// 4 hours. Measured: videos over this are 8.2% of the corpus by count but hold +// 46% of all transcript tokens, so they are where a sweep's wall-clock actually +// goes and where chunk-seam bugs live. +export const DIGEST_LONG_TAIL_DEFAULT_SECONDS = 4 * 3600; +export const DIGEST_LONG_TAIL_MAX_SECONDS = 24 * 3600; + +export function defaultDigest(): DigestSettings { + return { + // OFF. The metered lane is built but never the default — see PLAN.md. + remoteEnabled: false, + longTailSeconds: DIGEST_LONG_TAIL_DEFAULT_SECONDS, + localAppId: DEFAULT_DIGEST_APP_ID, + remoteAppId: CLAUDE_DIGEST_APP_ID, + apps: {}, + digestsPaused: false, + spendCapUsd: 0, + sections: ["chapters"], + }; +} + +// Coerce a raw settings.digest.apps value into a clean keyed map of +// DigestAppConfig. Mirrors sanitizeTranscriptionApps — INCLUDING its +// Array.isArray guard, without which a JSON array would pass the typeof check and +// produce numeric-keyed garbage. +export function sanitizeDigestApps( + value: unknown, +): Record<string, DigestAppConfig> { + if (!value || typeof value !== "object" || Array.isArray(value)) return {}; + const out: Record<string, DigestAppConfig> = {}; + for (const [id, raw] of Object.entries(value as Record<string, unknown>)) { + if (!raw || typeof raw !== "object") continue; + const r = raw as Record<string, unknown>; + const cfg: DigestAppConfig = {}; + if (typeof r.bin === "string" && r.bin.trim()) cfg.bin = r.bin.trim(); + if (typeof r.baseUrl === "string" && r.baseUrl.trim()) { + cfg.baseUrl = r.baseUrl.trim(); + } + if (typeof r.model === "string" && r.model.trim()) cfg.model = r.model.trim(); + if (typeof r.numCtx === "number" && r.numCtx > 0) { + cfg.numCtx = Math.floor(r.numCtx); + } + if (typeof r.temperature === "number" && r.temperature >= 0) { + cfg.temperature = r.temperature; + } + if (typeof r.timeoutMs === "number" && r.timeoutMs > 0) { + cfg.timeoutMs = Math.floor(r.timeoutMs); + } + out[id] = cfg; + } + return out; +} + +export function sanitizeDigest(value: unknown): DigestSettings { + const d = defaultDigest(); + if (!value || typeof value !== "object") return d; + const r = value as Record<string, unknown>; + const sections = Array.isArray(r.sections) + ? (r.sections.filter(isDigestSectionKind) as DigestSectionKind[]) + : []; + return { + remoteEnabled: r.remoteEnabled === true, + longTailSeconds: clampPositiveInt( + r.longTailSeconds, + d.longTailSeconds, + DIGEST_LONG_TAIL_MAX_SECONDS, + ), + // Unknown app ids are not rejected here: getDigestApp() is total and falls + // back to the local default, so a stale id degrades rather than breaking. + localAppId: + typeof r.localAppId === "string" && r.localAppId.trim() + ? r.localAppId.trim() + : d.localAppId, + remoteAppId: + typeof r.remoteAppId === "string" && r.remoteAppId.trim() + ? r.remoteAppId.trim() + : d.remoteAppId, + apps: sanitizeDigestApps(r.apps), + digestsPaused: r.digestsPaused === true, + spendCapUsd: + typeof r.spendCapUsd === "number" && r.spendCapUsd > 0 + ? Math.round(r.spendCapUsd * 100) / 100 + : 0, + // An empty/garbage list would silently generate nothing, so fall back to the + // default rather than honoring it. + sections: sections.length > 0 ? sections : d.sections, + }; +} + +// Every known section kind, for the settings UI's checkbox list. +export const DIGEST_SECTION_OPTIONS = DIGEST_SECTION_KINDS; + export const BUILD_MAX_PARALLEL_DEFAULT = 2; export const BUILD_MAX_PARALLEL_MAX = 16; export const DEFAULT_BUILD_IMAGE = "yt-dlp-transcript-browser-build"; @@ -498,6 +632,9 @@ function defaults(): SiteSettings { homepageUrl: "", savedVideoBackup: defaultSavedVideoBackup(), buildPipeline: defaultBuildPipeline(), + // Must be listed here or the allowlist loop in getSettings() drops the key + // entirely and the whole section is never read from disk. + digest: defaultDigest(), }; } @@ -712,6 +849,7 @@ export function getSettings(): SiteSettings { merged.homepageUrl = normalizeHomepageUrl(merged.homepageUrl); merged.savedVideoBackup = sanitizeSavedVideoBackup(merged.savedVideoBackup); merged.buildPipeline = sanitizeBuildPipeline(merged.buildPipeline); + merged.digest = sanitizeDigest(merged.digest); // Workers. When the file predates the worker model (no `workers` key), // synthesize a default list from the (now-settled) active app + per-app // configs so existing installs behave identically. Otherwise sanitize the @@ -904,6 +1042,7 @@ export async function writeSettings(next: SiteSettings): Promise<void> { homepageUrl: normalizeHomepageUrl(next.homepageUrl), savedVideoBackup: sanitizeSavedVideoBackup(next.savedVideoBackup), buildPipeline: sanitizeBuildPipeline(next.buildPipeline), + digest: sanitizeDigest(next.digest), }; const tmp = `${file}.tmp-${process.pid}`; await fs.promises.writeFile(tmp, JSON.stringify(merged, null, 2) + "\n"); diff --git a/common/lib/transcriptWindow.test.ts b/common/lib/transcriptWindow.test.ts @@ -1,6 +1,11 @@ import { test } from "node:test"; import assert from "node:assert/strict"; -import { windowCues, cuesToSnippets, mergeSnippets } from "./transcriptWindow"; +import { + chunkCuesForContext, + cuesToSnippets, + mergeSnippets, + windowCues, +} from "./transcriptWindow"; import type { Cue } from "./vtt"; function cue(start: number, text = `t${start}`): Cue { @@ -82,3 +87,67 @@ test("mergeSnippets never drops existing hits and stops adding at cap", () => { [1, 2, 3], ); }); + +// --------------------------------------------------------------------------- +// chunkCuesForContext — the digest chunker. windowCues cannot do this job (it is +// center-based and measured in seconds), so these cases pin the contract the +// digest generator depends on: full coverage, in order, with a shared overlap. +// --------------------------------------------------------------------------- + +test("chunkCuesForContext returns one chunk when the transcript fits", () => { + const cues = [cue(0), cue(5), cue(10)]; + const chunks = chunkCuesForContext(cues, { maxCues: 10, overlapCues: 2 }); + assert.equal(chunks.length, 1); + assert.deepEqual(chunks[0], cues); +}); + +test("chunkCuesForContext covers every cue at least once", () => { + const cues = Array.from({ length: 25 }, (_, i) => cue(i * 10)); + const chunks = chunkCuesForContext(cues, { maxCues: 10, overlapCues: 3 }); + const seen = new Set(chunks.flat().map((c) => c.start)); + assert.equal(seen.size, cues.length, "every cue appears in some chunk"); +}); + +test("chunkCuesForContext overlaps consecutive chunks by overlapCues", () => { + const cues = Array.from({ length: 25 }, (_, i) => cue(i * 10)); + const chunks = chunkCuesForContext(cues, { maxCues: 10, overlapCues: 3 }); + for (let i = 1; i < chunks.length; i++) { + const prevTail = chunks[i - 1].slice(-3).map((c) => c.start); + const head = chunks[i].slice(0, 3).map((c) => c.start); + assert.deepEqual(head, prevTail, `chunk ${i} shares its head with the previous tail`); + } +}); + +test("chunkCuesForContext emits chunks in chronological order", () => { + const cues = Array.from({ length: 40 }, (_, i) => cue(i * 10)); + const chunks = chunkCuesForContext(cues, { maxCues: 12, overlapCues: 4 }); + for (let i = 1; i < chunks.length; i++) { + assert.ok( + chunks[i][0].start > chunks[i - 1][0].start, + "each chunk starts later than the previous one", + ); + } +}); + +test("chunkCuesForContext never emits a trailing chunk that is pure overlap", () => { + // 13 cues at maxCues 10 / overlap 3 steps by 7: [0..9], [7..12]. A third chunk + // starting at 14 would be empty, and one starting at 12 would be re-work. + const cues = Array.from({ length: 13 }, (_, i) => cue(i * 10)); + const chunks = chunkCuesForContext(cues, { maxCues: 10, overlapCues: 3 }); + assert.equal(chunks.length, 2); + assert.equal(chunks[1][chunks[1].length - 1].start, 120); +}); + +test("chunkCuesForContext clamps an overlap that would stall the loop", () => { + // overlapCues >= maxCues would make step 0 and loop forever; it is clamped so + // progress is guaranteed. + const cues = Array.from({ length: 30 }, (_, i) => cue(i * 10)); + const chunks = chunkCuesForContext(cues, { maxCues: 5, overlapCues: 99 }); + assert.ok(chunks.length > 1 && chunks.length < 40); + const seen = new Set(chunks.flat().map((c) => c.start)); + assert.equal(seen.size, cues.length); +}); + +test("chunkCuesForContext handles an empty transcript", () => { + assert.deepEqual(chunkCuesForContext([], { maxCues: 10 }), []); +}); diff --git a/common/lib/transcriptWindow.ts b/common/lib/transcriptWindow.ts @@ -40,6 +40,47 @@ export function windowCues( .sort((a, b) => a.start - b.start); } +// Split a whole transcript into SEQUENTIAL, overlapping, context-sized slices — +// the chunker the digest generator drives its engine with, one call per slice. +// +// This is a different job from windowCues() above and cannot be expressed with +// it: that one is CENTER-based (a window around a search hit, measured in +// seconds) and has no notion of covering a transcript exactly once. Here the +// contract is coverage: every cue appears in at least one slice, slices are in +// order, and consecutive slices share `overlapCues` cues. +// +// The overlap exists because a topic that straddles a seam is otherwise invisible +// to both calls — each sees half a discussion and titles it wrong. With the +// overlap, at least one call sees it whole; the parser's seam de-dup then drops +// the resulting near-identical chapters. +// +// `maxCues` is measured in cues rather than tokens deliberately: cues are what we +// slice, and the token estimate that sizes them belongs with the prompt (see +// DIGEST_MAX_CUES_PER_CHUNK), not here. +export function chunkCuesForContext( + cues: Cue[], + opts: { maxCues?: number; overlapCues?: number } = {}, +): Cue[][] { + const maxCues = Math.max(1, Math.floor(opts.maxCues ?? 1200)); + // Overlap must leave forward progress, or the loop never advances. + const overlapCues = Math.max( + 0, + Math.min(Math.floor(opts.overlapCues ?? 0), maxCues - 1), + ); + if (cues.length === 0) return []; + if (cues.length <= maxCues) return [cues.slice()]; + + const step = maxCues - overlapCues; + const out: Cue[][] = []; + for (let start = 0; start < cues.length; start += step) { + out.push(cues.slice(start, start + maxCues)); + // Stop once this slice reached the end, so we never emit a trailing slice + // that is pure overlap (it would be re-processed for nothing). + if (start + maxCues >= cues.length) break; + } + return out; +} + // Render cues as snippet objects (clock + seconds + collapsed, capped text). // Empty cues are dropped. export function cuesToSnippets(cues: Cue[]): WindowSnippet[] { diff --git a/common/package.json b/common/package.json @@ -3,6 +3,9 @@ "version": "0.1.0", "private": true, "type": "module", + "scripts": { + "test": "tsx --test \"{lib,controller,jobs,social,ytdlp,components}/*.test.ts\"" + }, "dependencies": { "@sindresorhus/slugify": "^3.0.0", "@tanstack/react-query": "^5.99.1", diff --git a/editor/app/channels/[slug]/digestActions.ts b/editor/app/channels/[slug]/digestActions.ts @@ -0,0 +1,249 @@ +"use server"; + +import { revalidatePath } from "next/cache"; +import { getPaths } from "yt-dlp-transcript-common/lib/paths"; +import { getSettings } from "yt-dlp-transcript-common/lib/settings"; +import { + DIGEST_LOCAL_QUEUE, + DIGEST_REMOTE_QUEUE, + resolveQueueKey, +} from "yt-dlp-transcript-common/lib/queueKeys"; +import { readChannelStat } from "yt-dlp-transcript-common/controller/channels"; +import { + countMissingDigests, + runDigestBatch, + type DigestOrder, +} from "yt-dlp-transcript-common/controller/digestBatch"; +import { + runManagedFunction, + type StreamActionResult, +} from "yt-dlp-transcript-common/jobs/streamCommand"; +import { makeTaskTracker } from "yt-dlp-transcript-common/jobs/taskHooks"; +import { requestChannelSnapshot } from "yt-dlp-transcript-common/jobs/snapshotScheduler"; +import { + loadDigest, + loadDigestOverrides, + writeDigestOverrides, +} from "yt-dlp-transcript-common/lib/digest-server"; +import { + effectiveDigest, + type DigestChapter, + type DigestOverrides, + type DigestTag, + type EffectiveDigest, +} from "yt-dlp-transcript-common/lib/digest"; +import path from "node:path"; + +function isDigestOrder(v: unknown): v is DigestOrder { + return v === "shortest-first" || v === "longest-first"; +} + +export type DigestLaneChoice = "local" | "remote"; + +// Run the digest sweep over one channel. ONE JOB PER CHANNEL, deliberately: +// job logs keep only the newest 500 (30 days) and the in-memory registry keeps +// 100 records, so a job per video would evict the whole history of a sweep — +// including the running jobs' own logs. +export async function digestChannelAction( + slug: string, + lane: DigestLaneChoice = "local", + queueKey?: string, + order?: string, + limitCount?: number, + force?: boolean, +): Promise<StreamActionResult> { + const paths = getPaths(); + const settings = getSettings(); + if (lane === "remote" && !settings.digest.remoteEnabled) { + return { + ok: false, + error: + "The metered digest lane is off. Enable it in Settings → Digest before running it.", + }; + } + const remote = lane === "remote"; + const kind = remote ? "digest-channel-remote" : "digest-channel-local"; + return runManagedFunction({ + kind, + // Separate keys per lane so the two run CONCURRENTLY — the local lane is + // GPU-bound and the metered lane is network-bound, so serializing them would + // waste half the throughput of a multi-week sweep. + queueKey: resolveQueueKey( + remote ? DIGEST_REMOTE_QUEUE : DIGEST_LOCAL_QUEUE, + queueKey, + ), + paths, + channelSlug: slug, + spec: { + kind, + slug, + params: { queueKey, lane, order, limitCount, force }, + }, + fn: async (onLog, signal, setProgress, ctx) => { + const stat = await readChannelStat(paths, slug); + if (stat) { + const missing = await countMissingDigests(paths, slug); + setProgress({ + metric: "digests", + initial: stat.digestCount ?? 0, + target: (stat.digestCount ?? 0) + missing, + }); + } + const result = await runDigestBatch({ + channelSlug: slug, + paths, + lane, + order: isDigestOrder(order) ? order : undefined, + // The local lane takes everything up to the long-tail cutoff; the metered + // lane exists for the tail above it. Passing no window (the default) runs + // the whole channel on one lane, which is what a single-lane sweep wants. + ...(remote + ? { minDurationSeconds: settings.digest.longTailSeconds } + : {}), + limitCount: + typeof limitCount === "number" && limitCount > 0 ? limitCount : undefined, + force: force === true, + setProgress, + progressBaseline: stat?.digestCount ?? 0, + onLog, + signal, + drainSignal: ctx.drainSignal, + tracker: makeTaskTracker(ctx, onLog), + }); + onLog( + `Digest batch: ${result.succeeded} generated, ${result.fresh} already current, ` + + `${result.shared} shared to mirrors, ${result.misaligned} mirror(s) refused by the alignment gate, ` + + `${result.skipped} skipped, ${result.failed} failed; ` + + `${result.engineCalls} model call(s), ${result.warnings} warning(s)` + + (result.costUsd > 0 ? `, $${result.costUsd.toFixed(4)}` : "") + + (result.spendCapped ? " (stopped at the spend cap)" : "") + + ".", + ); + // ONCE, at job end — the digest kinds are in NO_REGEN_KINDS precisely so + // that per-video regens don't fire. See snapshotScheduler.ts. + requestChannelSnapshot(paths, slug); + revalidatePath(`/channels/${slug}`); + }, + }); +} + +// Digest an explicit id set (a bucket selection, or the pilot's single channel +// slice). Shares the batch machinery; scoped ids are intersected with disk. +export async function digestBucketAction( + slug: string, + ids: string[], + lane: DigestLaneChoice = "local", + queueKey?: string, +): Promise<StreamActionResult> { + const paths = getPaths(); + const settings = getSettings(); + const cleaned = Array.from(new Set(ids.map((id) => id.trim()).filter(Boolean))); + if (cleaned.length === 0) { + return { ok: false, error: "No video ids supplied" }; + } + if (lane === "remote" && !settings.digest.remoteEnabled) { + return { ok: false, error: "The metered digest lane is off." }; + } + const remote = lane === "remote"; + const kind = remote ? "digest-channel-remote" : "digest-channel-local"; + return runManagedFunction({ + kind, + queueKey: resolveQueueKey( + remote ? DIGEST_REMOTE_QUEUE : DIGEST_LOCAL_QUEUE, + queueKey, + ), + paths, + channelSlug: slug, + fn: async (onLog, signal, setProgress, ctx) => { + const stat = await readChannelStat(paths, slug); + if (stat) { + const missing = await countMissingDigests(paths, slug, cleaned); + setProgress({ + metric: "digests", + initial: stat.digestCount ?? 0, + target: (stat.digestCount ?? 0) + missing, + }); + } + const result = await runDigestBatch({ + channelSlug: slug, + paths, + lane, + ids: cleaned, + onLog, + signal, + drainSignal: ctx.drainSignal, + tracker: makeTaskTracker(ctx, onLog), + }); + onLog( + `Digest bucket: ${result.succeeded} generated, ${result.fresh} already current, ${result.failed} failed.`, + ); + requestChannelSnapshot(paths, slug); + revalidatePath(`/channels/${slug}`); + }, + }); +} + +// --------------------------------------------------------------------------- +// Per-video review: read the composed digest, write human corrections +// --------------------------------------------------------------------------- + +export type VideoDigestView = { + digest: EffectiveDigest; + hasMachineDigest: boolean; + overrides: DigestOverrides | null; +}; + +function videoDir(slug: string, id: string): string { + return path.join(getPaths().channelsDir, slug, "data", id); +} + +export async function readVideoDigestAction( + slug: string, + id: string, +): Promise<VideoDigestView> { + const dir = videoDir(slug, id); + const [machine, overrides] = await Promise.all([ + loadDigest(dir), + loadDigestOverrides(dir), + ]); + return { + digest: effectiveDigest(machine, overrides), + hasMachineDigest: machine !== null, + overrides, + }; +} + +export type SaveDigestOverridesResult = + | { ok: true; digest: EffectiveDigest } + | { ok: false; error: string }; + +// Write human corrections. This ONLY ever touches ai-digest.overrides.json — +// never ai-digest.json — which is what makes a correction survive the next +// regeneration of a corpus too large to re-generate twice. +export async function saveDigestOverridesAction( + slug: string, + id: string, + input: { + chapters?: DigestChapter[]; + tags?: DigestTag[]; + note?: string; + }, +): Promise<SaveDigestOverridesResult> { + const dir = videoDir(slug, id); + try { + await writeDigestOverrides(dir, { + version: 1, + ...(input.chapters ? { chapters: input.chapters } : {}), + ...(input.tags ? { tags: input.tags } : {}), + ...(input.note ? { note: input.note } : {}), + }); + } catch (e) { + return { ok: false, error: (e as Error).message }; + } + const [machine, overrides] = await Promise.all([ + loadDigest(dir), + loadDigestOverrides(dir), + ]); + revalidatePath(`/channels/${slug}/videos/${id}`); + return { ok: true, digest: effectiveDigest(machine, overrides) }; +} diff --git a/editor/app/channels/[slug]/lib/stageStatus.ts b/editor/app/channels/[slug]/lib/stageStatus.ts @@ -34,6 +34,7 @@ export function normalizeBuckets( downloadedAutoSubsOnly: raw?.downloadedAutoSubsOnly ?? [], supersededAutoSubs: raw?.supersededAutoSubs ?? [], needsCookies: raw?.needsCookies ?? [], + noDigest: raw?.noDigest ?? [], }; } diff --git a/editor/app/jobs/active/buildActiveJobs.ts b/editor/app/jobs/active/buildActiveJobs.ts @@ -62,8 +62,15 @@ function computeJobProgressView( ): RunningJobsListItem["progress"] { const snap = job.progress; if (!snap || !stat) return undefined; + // `current` is RE-COUNTED from disk (readChannelStat), never reported by the + // runner — which is why the digest metric needed its own on-disk counter + // (digestCount) rather than a number the batch could have just told us. const current = - snap.metric === "downloads" ? stat.downloadCount : stat.transcriptCount; + snap.metric === "downloads" + ? stat.downloadCount + : snap.metric === "digests" + ? (stat.digestCount ?? 0) + : stat.transcriptCount; const range = Math.max(0, snap.target - snap.initial); const advance = Math.max(0, current - snap.initial); const pct = diff --git a/editor/app/jobs/components/RunningJobsList.tsx b/editor/app/jobs/components/RunningJobsList.tsx @@ -3,6 +3,10 @@ import Link from "next/link"; import { useEffect, useState } from "react"; import { formatDuration } from "yt-dlp-transcript-common/lib/format"; +import type { + JobProgressMetric, + JobTaskKind, +} from "yt-dlp-transcript-common/jobs/registry"; import { JobLogTail } from "../[id]/components/JobLogTail"; import { jobKindLabel } from "../jobKindLabels"; import { DrainJobButton } from "./DrainJobButton"; @@ -14,7 +18,9 @@ import { ReorderJobButtons } from "./ReorderJobButtons"; export type RunningJobsTask = { id: string; label: string; - kind: "download" | "transcribe"; + // Imported for the same reason as `metric` below: a re-spelled literal here + // would not fail the build when JobTaskKind grew a member. + kind: JobTaskKind; fraction?: number; detail?: string; // Epoch ms when this sub-operation started, for the live "running for" timer. @@ -38,7 +44,9 @@ export type RunningJobsListItem = { channelSlug?: string; videoId?: string; progress?: { - metric: "downloads" | "transcripts"; + // Imported, NOT re-spelled: a literal copy here silently drifted from + // JobProgressMetric and would not fail the build when the union grew. + metric: JobProgressMetric; initial: number; current: number; target: number; @@ -214,16 +222,25 @@ function useNow(): number | null { return now; } +// Records over JobTaskKind, so adding a kind is a compile error here. +const TASK_KIND_VERB: Record<JobTaskKind, string> = { + download: "Downloading", + transcribe: "Transcribing", + digest: "Digesting", +}; + +const METRIC_FILL_BY_TASK: Record<JobTaskKind, string> = { + download: "bg-success/60", + transcribe: "bg-success", + digest: "bg-info", +}; + function TaskProgressBar({ task }: { task: RunningJobsTask }) { // How long this task has been running. Null until mounted (see useNow); // formatDuration returns "" for 0, so the just-started case shows "0:00". const now = useNow(); const probing = task.phase === "probing"; - const verb = probing - ? "Probing audio" - : task.kind === "download" - ? "Downloading" - : "Transcribing"; + const verb = probing ? "Probing audio" : TASK_KIND_VERB[task.kind]; // While probing, fill against the estimated probe duration (a distinct violet // "scanning" bar) rather than the frozen download fraction. yt-dlp is paused, // so the download fraction wouldn't advance anyway. Falls back to an @@ -246,9 +263,7 @@ function TaskProgressBar({ task }: { task: RunningJobsTask }) { const pct = hasFraction ? Math.round((fraction as number) * 100) : 0; const fillClass = probing ? "bg-violet-400 dark:bg-violet-500" - : task.kind === "download" - ? "bg-success/60" - : "bg-success"; + : METRIC_FILL_BY_TASK[task.kind]; const pulseClass = probing ? "bg-violet-400 dark:bg-violet-500" : "bg-warning"; @@ -304,15 +319,26 @@ function TaskProgressBar({ task }: { task: RunningJobsTask }) { ); } +// One place per metric, so adding a metric to JobProgressMetric is a compile +// error here (Record over the union) rather than a silently-wrong label. +const METRIC_LABELS: Record<JobProgressMetric, string> = { + downloads: "Downloads", + transcripts: "Transcripts", + digests: "Digests", +}; + +const METRIC_FILL: Record<JobProgressMetric, string> = { + downloads: "bg-success/60", + transcripts: "bg-success", + digests: "bg-info", +}; + function JobProgressBar({ progress, }: { progress: NonNullable<RunningJobsListItem["progress"]>; }) { - const label = - progress.metric === "downloads" - ? `Downloads: ${progress.current} / ${progress.target}` - : `Transcripts: ${progress.current} / ${progress.target}`; + const label = `${METRIC_LABELS[progress.metric]}: ${progress.current} / ${progress.target}`; // Append an ETA once the batch has a measured average. formatDuration returns // "" for 0/falsy, so guard against printing a bare "·". const remaining = progress.target - progress.current; @@ -322,10 +348,7 @@ function JobProgressBar({ : typeof progress.etaSeconds === "number" ? `~${formatDuration(Math.max(1, Math.round(progress.etaSeconds)))} left` : "estimating…"; - const fillClass = - progress.metric === "downloads" - ? "bg-success/60" - : "bg-success"; + const fillClass = METRIC_FILL[progress.metric]; return ( <div className="flex flex-col gap-1"> <div diff --git a/editor/app/jobs/jobReplayRegistry.ts b/editor/app/jobs/jobReplayRegistry.ts @@ -35,6 +35,10 @@ import { transcribeBucketAction, transcribeMissingAction, } from "../channels/[slug]/whisperActions"; +import { + digestChannelAction, + type DigestLaneChoice, +} from "../channels/[slug]/digestActions"; import { persistKeptAction } from "../channels/[slug]/persistActions"; import { fetchPostsAction } from "../channels/[slug]/socialActions"; import { @@ -73,6 +77,31 @@ function params(spec: JobSpec): { } export const JOB_REPLAY_HANDLERS: Record<string, ReplayHandler> = { + // Both digest lanes replay through one action; the lane comes from params so a + // bookmarked metered run stays metered (and is refused if the lane has since + // been turned off, rather than quietly falling back to local). + "digest-channel-local": (spec) => { + const { p, queueKey } = params(spec); + return digestChannelAction( + spec.slug, + (str(p.lane) as DigestLaneChoice | undefined) ?? "local", + queueKey, + str(p.order), + num(p.limitCount), + bool(p.force), + ); + }, + "digest-channel-remote": (spec) => { + const { p, queueKey } = params(spec); + return digestChannelAction( + spec.slug, + (str(p.lane) as DigestLaneChoice | undefined) ?? "remote", + queueKey, + str(p.order), + num(p.limitCount), + bool(p.force), + ); + }, "whisper-all": (spec) => { const { p, queueKey } = params(spec); return transcribeMissingAction( diff --git a/editor/app/settings/actions.ts b/editor/app/settings/actions.ts @@ -201,6 +201,50 @@ export async function saveSettingsAction( // Build pipeline. Values are clamped/coerced by sanitizeBuildPipeline inside // writeSettings, so we only read the form here (NaN/blank → default). The // deploy-page toggle also writes `mode`; whichever saves last wins. + // Digest. The metered lane's switch is a checkbox like any other, but note the + // asymmetry: everything else here defaults to the CURRENT value on a partial + // save, while `remoteEnabled` is read straight from the form so it can never be + // turned on by an unrelated save. writeSettings re-sanitizes the whole block. + const dD = getSettings().digest; + const digestAppsRaw = String(formData.get("digestAppsJson") ?? "").trim(); + let digestApps: unknown = dD.apps; + if (digestAppsRaw) { + try { + digestApps = JSON.parse(digestAppsRaw); + } catch { + return { ok: false, error: "Digest app config payload is malformed" }; + } + } + const digestSectionsRaw = formData.getAll("digestSections").map(String); + // A hidden marker, because unchecked checkboxes are simply ABSENT from a + // FormData: without it, a submit from any form that lacks the digest fields + // would read remoteEnabled as false and silently reset the block. + const digestFormPresent = formData.get("digestFormPresent") === "1"; + const digestSettings = ( + digestFormPresent + ? { + remoteEnabled: formData.get("digestRemoteEnabled") === "on", + longTailSeconds: Number.parseInt( + String(formData.get("digestLongTailSeconds") ?? "").trim(), + 10, + ), + localAppId: + String(formData.get("digestLocalAppId") ?? "").trim() || + dD.localAppId, + remoteAppId: + String(formData.get("digestRemoteAppId") ?? "").trim() || + dD.remoteAppId, + apps: digestApps, + // Not edited by this form — the dashboard/channel controls own the pause. + digestsPaused: dD.digestsPaused, + spendCapUsd: Number.parseFloat( + String(formData.get("digestSpendCapUsd") ?? "").trim(), + ), + sections: digestSectionsRaw.length > 0 ? digestSectionsRaw : dD.sections, + } + : dD + ) as SiteSettings["digest"]; + const dB = defaultBuildPipeline(); const buildModeRaw = String(formData.get("buildMode") ?? "").trim(); const buildPipeline = { @@ -247,6 +291,9 @@ export async function saveSettingsAction( // Saved Videos page edits it). writeSettings re-sanitizes it regardless. savedVideoBackup: getSettings().savedVideoBackup, buildPipeline, + // Same: the Digest section of this form owns these fields, but an unrelated + // save must not reset them (and must never silently flip remoteEnabled on). + digest: digestSettings, }; try { await writeSettings(next); diff --git a/editor/app/widget/components/MonitorWidget.tsx b/editor/app/widget/components/MonitorWidget.tsx @@ -9,6 +9,10 @@ import type { DiskStatusView, } from "../../jobs/active/buildActiveJobs"; import type { RunningJobsListItem } from "../../jobs/components/RunningJobsList"; +import type { + JobProgressMetric, + JobTaskKind, +} from "yt-dlp-transcript-common/jobs/registry"; import { jobKindLabel } from "../../jobs/jobKindLabels"; import type { WorkersPayload, WorkerView } from "../../workers/components/WorkersView"; import { InlineActionButton } from "../../actionable/components/InlineActionButton"; @@ -604,6 +608,29 @@ function JobRow({ ); } +// Per-task-kind glyph/fill, as Records over JobTaskKind so a new kind is a +// compile error rather than silently rendering as a transcription. +const TASK_KIND_VERB: Record<JobTaskKind, string> = { + download: "\u2193", + transcribe: "\u270e", + digest: "\u00b6", +}; + +const TASK_KIND_FILL: Record<JobTaskKind, string> = { + download: "bg-success/60", + transcribe: "bg-success", + digest: "bg-info", +}; + +// Per-metric glyph for the compact widget line. A Record over JobProgressMetric +// so a new metric is a compile error, not a mislabelled bar (the two copies of +// this ternary previously had to be kept in sync by hand). +const METRIC_PREFIX: Record<JobProgressMetric, string> = { + downloads: "\u2193 ", + transcripts: "", + digests: "\u00b6 ", +}; + // One-line textual summary of a job's batch progress, e.g. "↓ 5/10 · ~2m left". // Shared by the job heading (headingProgress) and the job bar's caption so the // label/ETA formatting stays in one place. @@ -611,10 +638,7 @@ function jobProgressText( progress: NonNullable<RunningJobsListItem["progress"]>, showEta: boolean, ): string { - const label = - progress.metric === "downloads" - ? `↓ ${progress.current}/${progress.target}` - : `${progress.current}/${progress.target}`; + const label = `${METRIC_PREFIX[progress.metric]}${progress.current}/${progress.target}`; const remaining = progress.target - progress.current; const etaText = !showEta || remaining <= 0 @@ -632,10 +656,7 @@ function JobProgressBar({ progress: NonNullable<RunningJobsListItem["progress"]>; showEta: boolean; }) { - const label = - progress.metric === "downloads" - ? `↓ ${progress.current}/${progress.target}` - : `${progress.current}/${progress.target}`; + const label = `${METRIC_PREFIX[progress.metric]}${progress.current}/${progress.target}`; const remaining = progress.target - progress.current; const etaText = !showEta || remaining <= 0 @@ -673,7 +694,7 @@ function TaskBar({ }) { const now = useNow(); const probing = task.phase === "probing"; - const verb = probing ? "🔍" : task.kind === "download" ? "↓" : "✎"; + const verb = probing ? "🔍" : TASK_KIND_VERB[task.kind]; // While probing, fill against the estimated probe duration (violet "scanning" // bar) rather than the frozen download fraction. See TaskProgressBar. const probeFraction = @@ -694,9 +715,7 @@ function TaskBar({ const pct = hasFraction ? Math.round((fraction as number) * 100) : 0; const fillClass = probing ? "bg-violet-400 dark:bg-violet-500" - : task.kind === "download" - ? "bg-success/60" - : "bg-success"; + : TASK_KIND_FILL[task.kind]; const pulseClass = probing ? "bg-violet-400 dark:bg-violet-500" : "bg-warning";