commit 62b4cfc1753d1302b4518ba6deedfb497bc15b1b
parent dd7f18d2c5bde8350eb76df2a77f1cf5253a6c31
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Wed, 29 Jul 2026 12:23:40 -0400
Merge docs/local-ai-roadmap: get the digest backfill to the starting line
A corpus-wide digest sweep can now be started from one control, survives a
restart, pauses without being lost, yields the GPU to transcription, and reports
coverage plus an ETA in audio-hours — the unit the work is actually priced in.
The plan this delivers led with duplicate-sharing as the cost lever. Measuring it
in audio-hours showed it is worth ~1.7 sweep days of 80, not the ~11% the code
claimed: cluster members are 19% of the corpus by video count but 4.3% of its
audio-hours. GPU contention with whisper is worth ~56 days on the same
measurement, and that arbitration is now implemented but still unmeasured in
production — which is the single most valuable thing to do next.
Nine bugs fixed along the way that the plan did not anticipate, most found by
building or running the thing rather than by types or tests. The largest:
digest cluster sharing was silently dead for 14.5% of the corpus — including
100% of the channel it was written for.
Diffstat:
117 files changed, 18978 insertions(+), 397 deletions(-)
diff --git a/.gitignore b/.gitignore
@@ -55,6 +55,7 @@ yarn-error.log*
# them, leaving generated data showing up as untracked in every worktree.
/export/public/subs
/export/public/posts
+/export/public/digests
/export/public/summaries
/export/public/transcripts
/export/public/stats
diff --git a/AGENTS.md b/AGENTS.md
@@ -10,3 +10,10 @@ To run more than one checkout at once (parallel dev servers / e2e), use git work
per-worktree non-colliding ports. `pnpm wt add <branch>` creates one; `pnpm wt list` shows
each worktree's port block. `pnpm dev:editor` / `pnpm e2e` auto-assign ports per worktree.
See [WORKTREES.md](WORKTREES.md) for the port scheme and the shared-data caveat.
+
+# Roadmap
+
+Long-running work on the local-AI derived corpus is tracked in [PLAN.md](PLAN.md).
+Before doing any phase work, read `plans/STATE.md` (current status + decisions) and
+`plans/FACTS.md` (verified codebase facts — trust these over re-deriving them).
+The phase in flight has a detailed plan at `plans/phase-N-*.md`.
diff --git a/PLAN.md b/PLAN.md
@@ -0,0 +1,618 @@
+# Roadmap: the local-AI transformative layer
+
+Long-running work adding a derived, transformative layer over the verbatim transcript
+archive: timestamped chapters, topic tags, speaker attribution, a graduated visibility
+policy, and eventually a site that leads with the derived corpus rather than raw
+reproduction.
+
+This file is the **durable roadmap** and changes rarely. Three companions:
+
+- [`plans/FACTS.md`](plans/FACTS.md) — verified codebase facts. Trust it over re-deriving.
+- [`plans/STATE.md`](plans/STATE.md) — current status, decisions log, open questions.
+- `plans/phase-N-*.md` — file-level detail for the phase in flight, written when it starts.
+
+See [`plans/README.md`](plans/README.md) for the context-clear protocol.
+
+## Why
+
+Two goals drive the work.
+
+**Utility.** A timestamped summary per video lets a reader navigate a 3-hour livestream
+without scrubbing, and gives search something to match beyond raw speech.
+
+**Copyright posture.** The archive is more defensible when it publishes new expression
+*about* works alongside — and eventually instead of — full reproduction of them. Chapters,
+topic tags, and analysis are genuinely transformative; a graduated per-channel visibility
+policy replaces all-or-nothing takedown handling; and filtering out embedded third-party
+clips reduces reproduction of works the archive has the weakest claim to.
+
+## Non-goal: transcripts are never rewritten
+
+Verbatim accuracy is what makes search, `[n @ mm:ss]` citations, and the MCP corpus sweeps
+trustworthy. A full-length paraphrase would still be a derivative work while adding
+hallucination risk. Transformative value is layered **on top of** verbatim cues, and
+reduction of exposure happens by **withholding** cues, not by altering them.
+
+## Architecture
+
+One derived sidecar per video directory, mirroring the `transcript.cues.json` precedent:
+versioned, self-describing, written tmp+rename, with an mtime freshness check.
+
+```
+transcripts/channels/<slug>/data/<videoId>/
+ transcript.cues.json # verbatim source, never rewritten
+ ai-digest.json # chapters + tags + per-section provenance
+ diarization.json # phase 9
+ attribution.json # phase 9
+```
+
+**Never name a sidecar `transcript.<x>.<y>`** — `SUB_FILE_RE` claims any such file as a
+subtitle track. Likewise **never use the name `summaries`** for the new artifact; that is
+already the cross-channel metadata browse index. Both hazards are detailed in `FACTS.md`.
+
+```mermaid
+flowchart LR
+ cues["transcript.cues.json<br/>(verbatim, immutable)"] --> md["transcriptToMarkdown<br/>stampForCue = hms"]
+ ctx["AI-CONTEXT.md → channel → video<br/>(phase 1.5)"] --> md
+ md --> chunk["chunkCuesForContext<br/>(NEW — windowCues cannot do this)"]
+ chunk --> eng{"digestApps registry"}
+ eng -->|"local-gpu lane"| oll["ollama-direct<br/>/api/chat + JSON schema"]
+ eng -->|"remote-api lane"| cc["claude-code<br/>claude -p --output-format json"]
+ oll --> parse["digestParse<br/>range · monotonic · snap · warnings"]
+ cc --> parse
+ parse --> dig["ai-digest.json<br/>+ per-section provenance"]
+ dig --> bi["buildIndex: digestMs → digests page tree"]
+ bi --> pub["/digests/<slug>/page-NNNN.json"]
+ pub --> ui["viewer ?vm=summary · search · MCP · corpus.json"]
+```
+
+### Engine registries, and why there are two lanes
+
+The repo's established idiom for "swappable engine" is a **registry of descriptors with a
+client-safe listing function** — `TRANSCRIPTION_APPS` / `listTranscriptionApps()`. The
+digest layer copies that idiom exactly rather than inventing a second pattern. Transcription
+already has three engines; digests get two.
+
+| Lane | Engine | Queue key | Runs alongside | Cost |
+| --- | --- | --- | --- | --- |
+| `local-gpu` | `ollama-direct` (default) | `TRANSCRIPTION_QUEUE` | nothing — the 8 GB card is shared with whisper/parakeet | free |
+| `remote-api` | `claude-code` | `DIGEST_REMOTE_QUEUE` | the GPU lane | metered |
+
+Two lanes on **different queue keys** is the structural idea that makes Claude Code useful
+rather than merely available: a backfill can drive both at once, Ollama saturating the GPU
+while Claude Code works the same backlog over the network. Putting both on one key would
+serialize them and waste the network lane. Putting the local engine on its own key would let
+Ollama and the transcription engine thrash the same 8 GB of VRAM.
+
+Metered engines are **opt-in** (`digest.remoteEnabled`, default `false`) and are never the
+auto-queue default by inheritance.
+
+### Provenance — which AI was used
+
+Authoritative per **section**, because the digest controller does a read–modify–write merge
+and a video can legitimately carry Ollama chapters and Claude Code tags:
+
+```jsonc
+{
+ "version": 1,
+ "contextHash": "sha1…", // phase 1.5 — invalidates on channel-context edits
+ "chapters": [{ "start": 0, "title": "COLD OPEN" }],
+ "tags": ["game review"],
+ "warnings": [{ "section": "chapters", "reason": "timestamp-out-of-range", "raw": "01:10:29 …" }],
+ "sections": {
+ "chapters": { "appId": "ollama-direct", "model": "qwen2.5:7b", "lane": "local-gpu",
+ "generatedAt": "2026-07-25T…", "promptVersion": 1 },
+ "tags": { "appId": "claude-code", "model": "claude-opus-5", "lane": "remote-api",
+ "generatedAt": "2026-07-26T…", "promptVersion": 1 }
+ }
+}
+```
+
+`engines: string[]` (the distinct `appId`s) is derived on read for cheap display and
+filtering — do not persist it as a second source of truth.
+
+Provenance must reach every surface, not just disk: the per-video editor badge, a
+`digestEngines: Record<string, number>` stat on `ChannelSnapshot` (a sibling of `totals`,
+**not** a `buckets` entry — buckets are work-lanes of video ids), the digest record shipped
+to the export, and a viewer badge. The Phase 10 honesty requirement extends to *which*
+model, not merely *that* it was generated.
+
+## Phases
+
+```mermaid
+flowchart TD
+ P0["0 · Benchmark engines"] --> P1["1 · Generation harness"]
+ P1 --> P15["1.5 · Channel context"]
+ P15 --> P2["2 · Digest corpus in build"]
+ P1 --> P25["2.5 · Observability"]
+ P25 --> BF(["BACKFILL"])
+ P15 --> BF
+ P11a["11a · Review queue"] --> BF
+ P2 --> P3["3 · Viewer ?vm=summary"]
+ P2 --> P4["4 · Search indexing"]
+ P25 --> P5["5 · Auto-queue"]
+ P3 --> P10["10 · Lead with the corpus"]
+ P4 --> P10
+ P8["8 · Visibility policy"] --> P10
+ P2 --> P9["9 · Attribution + quote filter"]
+ P6["6 · Ollama /ask provider<br/>(no dependencies)"]
+ P1 --> P7["7 · Tags + chat highlights"]
+ P10 --> P11b["11b · Viewer feedback"]
+```
+
+### Phase 0 — Benchmark the transcription engines
+
+Small, and it belongs first because **transcription is the upstream bottleneck for the
+digest backfill**: a video cannot be digested until it has a transcript.
+
+`parakeet` is **already** a fully selectable engine — the third entry in
+`TRANSCRIPTION_APPS`, exposed in settings through `listTranscriptionApps()`, with a
+`device` field and `supportsPartialStop: true` (it stops after the current window and
+stitches a partial transcript on SIGTERM, so a long run is interruptible — directly
+relevant to a multi-day sweep). The default is `whisper-cpp`. Nothing is missing
+structurally.
+
+What *is* missing is evidence. Nothing in the repo measures relative throughput on this
+hardware, so "if parakeet is faster" is currently unanswerable.
+
+- Time `whisper-cpp`, `chough`, and `parakeet` over the same handful of real videos of
+ varying length, on this GPU. Record wall-clock per audio-minute, VRAM, and transcript
+ quality spot-checks.
+- Write the numbers into `plans/FACTS.md`. This is a fact that will be re-asked for months
+ and should never need re-measuring.
+- If parakeet wins, change `DEFAULT_TRANSCRIPTION_APP_ID` and note it in `STATE.md`.
+ Because workers already carry a per-worker `appId`, a mixed fleet is also possible
+ without new machinery.
+- Deliverable is measurements and possibly a one-line default change — **not** a new
+ benchmark harness. Do not build infrastructure for a question asked once.
+
+### Phase 1 — Generation harness
+
+Modelled on the transcription-app registry, **not** on `askProvider.ts` — that runs in the
+visitor's browser in the static export and cannot spawn a process.
+
+- `common/lib/paths.ts` — add `ollamaUrl` (`OLLAMA_URL`) and `claudeBin` (`CLAUDE_BIN`),
+ following the `process.env.X ?? default` idiom. Add both to the hand-curated `pathRows`
+ in the settings page's System paths table.
+- `common/lib/digestApps.ts` *(new)* — mirror `transcriptionApps.ts`: a `DigestApp` type, a
+ `DIGEST_APPS` record, a `getDigestApp(id)` that never throws, and a client-safe
+ `listDigestApps()` descriptor export so the registry never reaches the client bundle.
+ Each app carries `lane` and `metered`.
+ - **`ollama-direct`** — POST `${ollamaUrl}/api/chat` with `format: <JSON schema>`,
+ `options: { num_ctx: 16384, temperature: 0 }`, `stream: false`. **`num_ctx` must be
+ explicit**: the 4096 default silently truncates, the single easiest way to get
+ quietly-wrong output at scale.
+ - **`claude-code`** — shell out to `claudeBin -p <prompt> --output-format json`, parse the
+ wrapper's `result` string, then the JSON body. No context-window concern, but keep the
+ same chunking so output shape stays engine-independent.
+ - **`fabric`** — deferred. Measured to ignore the `HH:MM:SS TOPIC` contract at 7B, twice,
+ at two input sizes. Its 256-pattern library stays valuable as prompt source material;
+ it becomes a third registry entry when a prose-shaped output needs it, with no rework.
+- `common/lib/transcriptWindow.ts` — add `chunkCuesForContext(cues, { maxCues, overlapCues })`
+ returning sequential overlapping slices. **`windowCues()` cannot do this** — it is a
+ center-based helper for search hits (see `FACTS.md`).
+- `common/lib/digestParse.ts` *(new)* — parse engine output into `{ start, title }[]`. Must
+ reject timestamps beyond the video duration, **clamp each chunk's output to that chunk's
+ own range**, enforce monotonic starts, drop malformed entries, snap starts to the nearest
+ cue boundary, and de-duplicate across chunk seams. Every rejection is **recorded in a
+ `warnings` array, never silently dropped** — Phase 11's review queue is built entirely on
+ this. Unit-tested; it is the layer absorbing local-model sloppiness.
+- `common/controller/digestVideo.ts` *(new)* — freshness check → `readNormalizedTranscript`
+ → `transcriptToMarkdown` with `stampForCue: (_c, s) => hms(s)` → chunk → engine → parse →
+ read-modify-write merge → tmp+rename. Reachability probe for the Ollama lane; a friendly
+ ENOENT message for the Claude lane.
+- `common/controller/digestBatch.ts` *(new)* — `runPool()`, honoring `signal` and
+ `ctx.drainSignal`. Logs a running count of **metered** calls so a remote backfill's cost
+ is visible in the job log as it happens.
+- `common/lib/queueKeys.ts` — `DIGEST_REMOTE_QUEUE` plus `digestQueueKey(appId)` mapping
+ lane → key. No concurrency declaration; the registry hardcodes 1 per key.
+- One `jobKinds.ts` entry (`digest-channel`), one server action modelled on
+ `whisperActions.ts`, one replay handler.
+- Settings: `defaultDigest()` + `sanitizeDigest()` wired into `getSettings()`'s normalize
+ chain, with per-engine config keyed by `appId` and `remoteEnabled: false`.
+- UI: a channel stage component, an engine picker, a per-video button, and a **provenance
+ badge** (appId + model + generated-at) on the video page.
+- `common/package.json` — add a `test` script; there is currently none, and ~46 unit test
+ files are unrunnable as a suite.
+
+### Phase 1.5 — Channel context layers
+
+Cascading human-authored notes injected into every prompt, modelled on this repo's own
+`CLAUDE.md` → `AGENTS.md`. Lands early because every later phase improves from it, and
+Phase 9's attribution is close to unusable without it.
+
+```
+transcripts/AI-CONTEXT.md # corpus-wide
+transcripts/channels/<slug>/context.md # the main one
+transcripts/channels/<slug>/data/<id>/context.md # rare, for oddities
+```
+
+Markdown with optional frontmatter: `hosts`, `recurring_guests`,
+`plays_third_party_media`, `boilerplate`.
+
+- `common/lib/aiContext.ts` *(new)* — `resolveAiContext(paths, slug, videoId?)` follows the
+ `resolveCookiePolicy` inheritance shape. Caps total tokens (~1500) so context never crowds
+ out transcript. **There is no YAML dependency in the repo** — hand-roll a minimal
+ scalar/list frontmatter parser rather than adding one for three field types.
+- `boilerplate` earns its keep with no model involvement: a deterministic filter dropping
+ recurring sponsor/subscribe cues from digest input and chapter candidates. Cheap, exact,
+ and it removes the most common junk chapter.
+- `plays_third_party_media: false` lets Phase 9 skip attribution entirely for that channel —
+ a large accuracy and compute win.
+- Editing UI on the channel and settings pages (tmp+rename). Frontmatter errors surface
+ inline, never at job time.
+- A `propose-channel-context` job samples N transcripts and drafts `context.suggested.md`.
+ It **never** writes `context.md` — auto-applying model-authored context would let one bad
+ inference silently degrade every downstream summary for a channel, invisibly, because the
+ output would still look plausible.
+- Hash the resolved context into `ai-digest.json.contextHash` so editing notes correctly
+ marks digests stale. **Corrected 2026-07-29:** this is NOT why backfill must wait. The
+ empty note already hashes to a stable value, so a note added later invalidates only its own
+ channel. The corpus-wide risk is the deliberate one — bumping the `digest-context-v1` hash
+ prefix — so what must precede a sweep is that DECISION, not this phase.
+
+### Phase 2 — Digests as a first-class corpus — **DONE (2026-07-29)**
+
+> Shipped as specified, plus three things this spec omitted: `exportDigestsDir` /
+> `exportSharedDigestsDir` in `paths.ts` (compose cannot be written without them), a
+> per-SITE digests manifest so `corpus.json` can source per-channel counts the way it does
+> for posts, and `digests` added to the service worker's `SHARD_RE` (without it the tree has
+> no offline story — the gap already recorded for `duplicates.json`). Note `?vm=digest`, not
+> `?vm=summary`; see the naming decision in `plans/STATE.md`.
+
+An earlier draft inlined the digest into `TranscriptDetail`. That is wrong: under
+`summary-only` visibility a transcript page ships **no cues**, so a digest must never be
+bundled inside a payload that policy may withhold. Digests get their own page tree.
+
+- `buildIndex.ts` — three new LMDB sub-DBs (`digests`, `digestPageHashes`,
+ `channelDigestStats`); **no `maxDbs` change needed**, but three new `clearAsync()` lines in
+ the hand-enumerated schema-invalidation block; `SCHEMA_VERSION` 12 → 13; `digestMs` (later
+ `attributionMs`) added to `MtimeRecord`, `LiveEntry`, the `scanSource()` stat loop, the
+ changed-detection comparison, and the `mtimes.put` call; emit
+ `.export-index/shared/digests/<slug>/{manifest,page-NNNN}.json` via `createPageWriter()`.
+ Each record carries its `sections` provenance.
+- `common/lib/digests.ts` *(new)* — `VideoDigest` / `DigestPage` / `ChannelDigestsManifest`
+ with `slugToPage`, mirroring `manifest.ts`.
+- `channelSignature.ts` — add `digestMs` to its three-field structural-subset `MtimeRecord`
+ **and** to the hash input, or archives will not rebuild on a digest-only change.
+- `compose-site.ts` — a `digests?` key in `ComposeCache` and a fourth
+ `reconcileChannelTree()` call. It already passes `ignoreBasename: "manifest.json"`
+ internally; nothing to add there.
+- `common/components/digestCache.ts` + `digestStore.ts` *(new)* — copy the
+ `transcriptCache.ts` + `transcriptStore.ts` pair (**not** `subsCache.ts`, which has no
+ IndexedDB). Use a **separate IDB store** so the transcript store's `DB_VERSION` need not
+ bump and existing transcript caches survive.
+- `corpus.ts` — `CORPUS_SPEC_VERSION` 2 → 3, a `DIGEST_SCHEME` modelled on `POST_SCHEME`,
+ and a digest manifest pointer on `CorpusChannel.manifests`. Update `corpus.test.ts`.
+- **Carry `derivedFrom` through to the client.** *(Added by the duplicate-detection work.)*
+ Once digests are shared across a duplicate cluster, a video's digest may have been
+ generated for a *different* video. `DigestRecord.derivedFrom` already records that
+ (`{slug, clusterId, sharedAt, offsetSeconds}`, written by `writeSharedDigest`), but it
+ stops at the record — the page-tree schema above must include it, or the viewer cannot
+ tell a native digest from a borrowed one. Sharing that is invisible is sharing that is
+ indistinguishable from a claim, which Phase 3 then has no way to be honest about. This is
+ a dependency Phase 2 creates, not an optional extra.
+
+### Phase 2.5 — Observability for the backfill
+
+Build this **before** backfilling. A multi-day sweep you cannot observe is one you cannot
+tune or safely interrupt.
+
+- **First**, fix the duplicated literal union in `RunningJobsList.tsx` to import
+ `JobProgressMetric`. Until then the compiler-driven audit the rest of this phase relies on
+ has a silent hole. Six sites total — see `FACTS.md`.
+- Extend `JobProgressMetric` with `"digests"` and `JobTaskKind` with `"digest"`.
+- The two copies of the `metric === "downloads" ? … : …` binary in `MonitorWidget.tsx` become
+ **one lookup table keyed by metric** (label, glyph, color), rather than growing a third
+ branch — otherwise the next metric repeats the bug.
+- **Coverage, not just progress.** Per-job bars answer "how's this job"; during backfill the
+ question is "how much of the corpus is done" — `digested / total`, **split by engine**.
+ That is the single most useful number during a multi-day sweep, and with two lanes running
+ it is also how you see whether the remote lane is pulling its weight.
+- Widget sync payload gains coverage **scalars only** — honor the file's stated design
+ constraint that it stays a handful of scalars rather than shipping per-channel breakdowns
+ on every poll. Its builder is already reused for dashboard SSR seeding, so both surfaces
+ get it free.
+- A digest instrument in `PipelineBand`, `noDigest` counts in `NeedsWorkPanel` /
+ `ChannelsTable`, command-palette actions, and ETA via the existing `computeEtaSeconds`.
+- `channelSnapshot.ts` gains `noDigest` **here**, not in Phase 5; Phase 5 then consumes what
+ already exists.
+
+### Phase 3 — Viewer — **DONE (2026-07-29), as `?vm=digest`**
+
+> Every `"summary"` below reads `"digest"` in the shipped code: "summary" already means a
+> video listing card AND the `/summaries/` page tree, and this file's own naming-hazard rule
+> forbids a third meaning. Also note the control is HIDDEN where no digest exists (0.1%
+> coverage makes an always-present button a dead end), and the client cache is versioned by
+> page content hash rather than `generatedAt` — reasons for both in `plans/STATE.md`.
+
+- `urlState.ts` — add `"summary"` to `ModalMode`, the parse chain, and the `writeUrlParams`
+ `vm` branch. `ModalMode` appears in only three files.
+- `PlayerProvider.tsx` — lazy-fetch the digest when `modalMode === "summary"`, mirroring the
+ chat effect including its ref-based in-flight guard (there is a comment explaining why
+ reducer state in the dep array drops results) and its snap-back-to-transcript on missing.
+- `TranscriptModal.tsx` — a toolbar `ControlButton` beside the transcript/chat toggle;
+ chapters as buttons calling `seekTo(start)` with `scrollKindRef.current = "smooth"`;
+ active-chapter highlight off `currentTime`; a provenance line; an empty state.
+- **Styling:** this modal and `PlayerProvider` are **not** on semantic theme tokens — they
+ use hard-coded `bg-zinc-900/80`, `ring-white/10`, `text-white` throughout. Match that
+ local palette; do not introduce `bg-card`/`text-foreground` here. Migrating the modal
+ chrome is out of scope.
+- **Show a borrowed digest as borrowed.** *(Added by the duplicate-detection work.)* When
+ `derivedFrom` is set (Phase 2), the provenance line must name the member the digest was
+ generated *for* and the measured offset — not present it as native to the video being
+ watched. The chapters are placed by the canonical member's timeline; the detector's
+ `aligned` verdict is why sharing was allowed at all, and its `offsetSeconds` is the
+ residual error the viewer is looking at. A shared digest presented as native is the
+ failure mode that looks like success: every chapter plausible, all of them describing a
+ different upload.
+
+### Phase 4 — Search indexing
+
+Fold chapter titles and tags into the search corpus with a `source: "ai"` marker and the
+producing `appId`. **Two independent index paths exist** and both need it: the client-side
+FlexSearch worker and the server-side MCP search. Verbatim and generated hits must stay
+visually distinguishable — that honesty requirement is what makes indexing generated text
+acceptable at all.
+
+### Phase 5 — Auto-queue
+
+A third `AutoQueuePolicy` alongside `transcription`/`download`, a dispatch branch in
+`autoRunner.ts` re-reading settings per iteration so a pause flag takes effect live, exposure
+in the policy tree editor, a `digestsPaused` flag, and a pause button copied from the
+downloads one. **The policy carries an explicit `engineId` defaulting to the local lane** — a
+metered engine must never become the auto-queue default by inheritance.
+
+### Phase 6 — Ollama provider for `/ask`
+
+Zero dependencies on any other phase; a good early win. Widen the `Provider` union, add a
+`PROVIDERS` entry (no key required), a `switch` arm, and an `askOllama` using Ollama's
+OpenAI-compatible `/v1/chat/completions` SSE shape — the existing `askOpenAI` is a
+near-template. Make the base URL editable in the provider settings panel.
+
+**Caveat to surface in the UI:** the export is static and runs in the visitor's browser, so
+this only works when the visitor can reach an Ollama instance, and requires `OLLAMA_ORIGINS`
+for CORS. Local/LAN use only.
+
+Knock-on benefit: the MCP corpus sweep's extraction step can then run locally.
+
+### Phase 7 — Tags and live-chat highlights
+
+Tags come free with Phase 1 into the same `ai-digest.json`. Live-chat highlights reuse the
+whole harness against `live_chat.cues.json`, writing `ai-chat-digest.json` and rendering in
+the existing `?vm=chat` view.
+
+### Phase 8 — Transcript visibility policy
+
+Mandatory, not optional — leading with derived data means the published site's default
+posture is set here.
+
+- `common/lib/visibilityPolicy.ts` *(new)* — copy `cookiePolicy.ts` exactly:
+ `"full" | "excerpt" | "summary-only"`, a values array, a default, a type guard, and
+ `resolveVisibility(settings, channelConfig)`. Note cookiePolicy uses **two** inheritance
+ rules — truthiness-after-trim for free-text values, type-guard for enums; this is the enum
+ case.
+- Wire through settings, `parseChannelConfig`, both forms — and **list the key in
+ `CHANNEL_FORM_FIELDS`** or clearing back to inherit silently won't work.
+- **Enforced at build time** in the `buildIndex.ts` page writer, the single place that
+ decides what `cues` array ships. One enforcement point covers viewer, MCP, `/ask`, and
+ search, because all four read the same shards.
+- **Archives bypass the page writer** — `archiveTranscripts.ts` hard-links
+ `transcript.cues.json` straight from the source tree. Extend its existing `requiresRewrite`
+ predicate and `writeTransformedCues` path. This is the highest-exposure surface and
+ inherits nothing for free.
+
+### Phase 9 — Attribution (two lanes) + quote filtering
+
+Largest and least certain. Ship 1–8 first; keep it off by default.
+
+**Blocking constraint:** audio is deleted once a video is transcribed, except for dirs
+holding `do-not-clean.json` or entries in the saved-video store. **Most of the existing
+archive has no audio to diarize.**
+
+Two first-class lanes producing the same artifact at different quality tiers:
+
+| Lane | When | Input | Recorded as |
+| --- | --- | --- | --- |
+| Diarization-assisted | Going forward; on-demand upgrade | audio → speaker turns → LLM labels the turns | `method: "diarized"` |
+| Text-only | Legacy videos with no audio | cues alone → LLM segments and labels from content | `method: "text-only"` |
+
+Neither is a fallback for the other in code. Going forward, run diarization right after
+transcription while audio is still on disk and **before** the cleanup sweep. For legacy, run
+the text-only pass over the whole backlog immediately, then upgrade selectively.
+
+- `scripts/diarize.mjs` behind a `DIARIZE_BIN` path, exactly as the parakeet wrapper works.
+ Keeps the engine swappable: evaluate `sherpa-onnx` (CPU-friendly, no HF token) against
+ `pyannote` (better, needs token + GPU) **on real audio before committing** — quality here
+ cannot be judged from code.
+- `attributeCues.ts` writes `attribution.json` with ranges, confidence, method, and engine
+ provenance.
+- Filter in the page writer alongside Phase 8. **Bias toward dropping on uncertainty** — the
+ cost is asymmetric. Make the threshold policy-driven so a channel can filter on `diarized`
+ only and ignore `text-only` labels it doesn't trust.
+- Status everywhere: a client-safe `attributionStatus.ts`
+ (`"none" | "text-only" | "diarized" | "stale"`), new snapshot buckets, a per-video badge,
+ per-channel counts. Mark a video **ineligible** when availability says deleted/private and
+ no audio is retained — a distinct state, not an upgrade button that can only fail.
+- Re-processing: per-video re-attribute / upgrade, per-channel bulk over the text-only
+ bucket. The upgrade job re-downloads audio, diarizes, attributes, then removes the
+ re-fetched audio **in a `finally`** unless `do-not-clean.json` is present — this job can
+ pull gigabytes and a cancelled run must not silently fill the disk.
+- `attributionMs` threads through the build the same way `digestMs` does.
+- **Set expectations honestly:** this misfires on rapid back-and-forth, and auto-caption
+ channels have no speaker turns at all with cue boundaries that don't align to them. A
+ good-faith reduction in reproduction, not a guarantee; the UI must not claim otherwise.
+
+### Phase 10 — Leading with the derived corpus
+
+Depends on 2, 3, 4, 8.
+
+- **MCP** — a `get_digest` tool and digest-aware search, plus digest members on the
+ `ShardSource` interface implemented in **all three** classes (local, remote, hub). The hub
+ path needs it too, or federated sites silently lack digests.
+- **Archives** — a digests zip. Remember the 25 MB default cap drops *all* archives for large
+ sites.
+- **UI inversion** — search results lead with chapters and tags, verbatim cue matches
+ secondary and badged; listings surface chapter counts and topic tags; the modal defaults to
+ `?vm=summary` for `summary-only` channels (a default-selection change — the mode is already
+ URL-driven).
+- **Graceful fallback** where a digest is absent. A partially-digested corpus is the normal
+ state for a long time and must not look broken.
+- **Honesty:** generated content visually distinct from verbatim everywhere, labelled with
+ the producing model. This matters more once generated text is the primary thing users see.
+
+### Phase 11 — Human review queue & viewer feedback
+
+Without this, a corpus-wide backfill is a one-shot gamble on prompt quality.
+
+**Naming hazard:** `report` already means three different things in this repo. Use
+**`feedback`** for viewer-submitted items and **`review`** for triage state. Do not add a
+fourth meaning of "report".
+
+**11a — review queue (land before backfilling).** Extend the existing `/actionable` page
+rather than building a parallel one: its `SectionConfig` array with counters in
+`loadActionable.ts` is the designed extension point. New sections: digests needing review
+(driven by the `warnings` array Phase 1 persists), proposed channel context awaiting
+promotion, uncertain attribution, viewer feedback, and **duplicate clusters awaiting
+confirmation**.
+
+**Duplicate clusters awaiting confirmation** *(added by the duplicate-detection work — UI
+only, the data and the write path already exist).* A `title-duration` cluster is a suspect:
+two videos share a title and a near-identical runtime, and nothing compared their content
+because at least one side has no transcript. It ships nowhere (`clusterIsPublishable`) and
+shares no derived work (`clusterMaySharePartial`) until a human records `confirmed: true`.
+Corpus-wide that is currently **335 clusters**. Everything needed is in place:
+- the queue is `report.clusters.filter(c => c.needsReview)` minus `isClusterReviewed(...)`;
+- `updateDuplicateOverride(paths, clusterId, {confirmed})` already accepts and persists it,
+ and its clear-heuristic already refuses to delete a confirmation;
+- `/actionable` already renders cluster cards with a `needs review` badge.
+The missing piece is two buttons — *these are the same* / *these are not* — on that card.
+
+Actions per item: approve · edit inline · regenerate · dismiss · **"add note to channel
+context"**. The last is the one that compounds — a correction applied to one video fixes one
+video; the same correction written into `context.md` fixes every future generation for that
+channel. **Regenerate should offer the other engine**: "this looks wrong, redo it on Claude
+Code" is the most natural use of a second lane, and the provenance record makes it obvious
+when one engine is systematically weaker on a given channel.
+
+**11b — viewer feedback (can follow the corpus going public).** A "flag this" control per
+chapter, with categories including **"misrepresents what was said"**. Transport in preference
+order: an optional per-site `feedbackUrl` POST, else copy-to-clipboard / download-JSON.
+
+**Collector evaluated — the answer is no, not `r2-proxy/`.** *(Recorded by the
+duplicate-detection work so the question is not re-opened.)* The proxy is **read-only by
+explicit design**: `src/index.ts:35-41` rejects every non-GET/HEAD method, and `KEY_RE` at
+`:27` carries a comment stating the proxy must never become a general oracle over the bucket.
+Turning it into a write endpoint works directly against that intent, for a feature that does
+not need a server at all. Its one genuinely reusable piece is the `RATE_LIMITER` binding.
+**Use the zero-infrastructure fallback this phase already names** — copy-to-clipboard /
+download-JSON, pasted into an editor action. It needs no deployment and matches the existing
+idiom at `PlayerProvider.tsx:131-136`.
+
+**New trigger for this phase:** sharing one digest across a duplicate pair is precisely the
+case a viewer needs to be able to flag. A borrowed digest can be misplaced (the mirror drifts
+after the anchors that were measured) or simply wrong for that upload, and the viewer is the
+only party who will ever notice — the detector already believes the two are the same video.
+
+Ingestion is an editor action accepting pasted JSON. Treat submissions
+as **untrusted input rendered in an admin UI**: escape it, never feed it into a prompt
+unreviewed. Store under `transcripts/.feedback/`, a sibling of `.jobs` and `.bookmarks`,
+outside the build trees.
+
+**Why that category is not boilerplate:** the smoke test summarized allegation-heavy content
+about named individuals and restated those allegations as plain fact. A published AI summary
+that misstates what a real person said or did is the highest-risk output this system can
+produce, and the one failure mode no technical guard catches. Wire `dismiss` so it can
+**suppress a digest from the next build**, not merely flag it for later.
+
+## Verification strategy
+
+Per the repo's `verification-playwright-first` convention: verify with Playwright, open a
+browser only to diagnose failures.
+
+1. **Stub Ollama** — the default engine is an HTTP call from the Next server, so the
+ fake-binary trick does not cover it. Add an `ollama-stub.mjs` fixture launched from
+ `playwright.config.ts` alongside the existing web servers, with `OLLAMA_URL` pointed at
+ it. Give it a **deterministic bad-output mode** (out-of-range, non-monotonic) so the
+ parser guards are exercised by tests rather than by luck in production.
+2. **Fake `claude`** — a `fake-claude.mjs` echoing the `--output-format json` wrapper, plus
+ the `SLOWOP` paced-output trick so drain/progress specs can observe it. Wire `CLAUDE_BIN`
+ into **both** the `dev:test` and `start:test` lines; they are duplicated verbatim.
+3. **Unit** — `digestParse.test.ts` and the chunker cases, using `node:test`, runnable via
+ the new `common` test script.
+4. **Editor e2e** — queue a job, assert the row label, assert `ai-digest.json` lands with the
+ right `sections.chapters.appId`. Post-mutation assertions must **poll-with-reload**; the
+ channel page serves a snapshot regenerated on a ~1 s debounce.
+5. **Two-lane e2e** — queue both engines and assert they occupy **different queue keys** and
+ run concurrently. This is the behavior that makes the Claude Code lane worth having, so it
+ needs a test, not a comment.
+6. **Export e2e** — clone the route-stubbing scaffold in
+ `export-player-platform-cache.spec.ts`: stub manifest/pages with a digest-bearing fixture,
+ open `?v=<slug>&vm=summary`, assert chapters render, clicking one seeks, and the engine
+ badge shows.
+7. **Glance surfaces** — assert the sync payload carries coverage split by engine and the
+ pipeline band shows the digest chip and paused state. Run widget assertions under
+ `E2E_MODE=start` (see the Dev Tools caveat in `FACTS.md`).
+8. **Full suite** — `pnpm e2e` in **default dev mode**; `E2E_MODE=start` serves a stale build.
+ Kill stale dev servers by port between runs. The known-failing-on-base list is in
+ `FACTS.md` — do not chase those as regressions.
+9. **Real end-to-end** — one real channel against live Ollama, then `pnpm build:index &&
+ pnpm --filter export run build`; confirm the digest survives into
+ `export/public/digests/<slug>/page-*.json`. For Phase 9, confirm a `summary-only` channel
+ ships **zero** cues by inspecting the shard directly, not the UI.
+10. **Changelogs** — `editor/CHANGELOG.md` and `export/CHANGELOG.md` per repo convention.
+
+## Sequencing
+
+| Group | Why they group |
+| --- | --- |
+| 0 | Measurement only. Answers the transcription-speed question once, permanently. |
+| 1, 1.5, 2, 2.5, 3 | The shippable core: generate → context → build → observe → display. |
+| 4, 5, 6, 7 | Each independently shippable. **6 has no dependencies** — good early win. |
+| 8 | Self-contained; prerequisite for 10. |
+| 9 | Largest and least certain. Off by default. |
+| 10 | Depends on 2, 3, 4, 8. |
+| 11a / 11b | Review queue before backfill; viewer feedback after the corpus is public. |
+
+### Ordering traps
+
+- **CORRECTED (2026-07-29): 1.5 is not a backfill gate — a hash-prefix decision is.**
+ `digestContext-server.ts:31-35` hashes the EMPTY note to a stable value, so adding a
+ channel note later invalidates only that channel, not the corpus. What would invalidate
+ everything is bumping the `digest-context-v1` prefix (`:42`), which 1.5-as-specced does.
+ So the gate is a decision to take before sweeping, not work to do first.
+- **CORRECTED (2026-07-29): "before 2.5 and 11a" named the wrong things.** Both were largely
+ landed — the metric union, `METRIC_PREFIX`, per-job progress and the `noDigest` bucket all
+ existed. The things actually missing were never on this list: **no corpus-wide launcher,
+ no boot-time resume, no pause writer, and no GPU arbitration**. All four are now built
+ (`controller/digestSweep.ts`, `controller/digestYield.ts`, `instrumentation.ts`,
+ `DigestSweepControls`). The remaining honest gate is the review path, and its minimum —
+ persisted failure warnings plus a `digestWarnings` bucket — is in.
+- **Benchmark before committing to backfill.** The smoke test measured 15.5 s for a ~2.7k
+ token chunk; a 176-minute podcast needs roughly a dozen chunks. Budget minutes per long
+ video and multiply by corpus size. The remote lane changes this arithmetic — measure both,
+ then split the backlog between them.
+- **RESOLVED: the duplicated metric union is fixed.** `RunningJobsList` imports the union
+ and `MonitorWidget` uses a `METRIC_PREFIX` lookup table, so extending it is now
+ compiler-checked.
+- **Price a cost lever in AUDIO-HOURS before believing it** (`common/bin/digest-plan.ts`).
+ Measured 2026-07-29: duplicate-cluster sharing is worth **~1.7 sweep days of 80**, not the
+ "~11%" `digestSharing.ts`'s header claims — cluster members are 19% of the corpus by video
+ count but **4.3% of its audio-hours**, because mirrors skew SHORT and the sweep is
+ dominated by unclustered long-form VODs. GPU contention with whisper is worth ~56 days on
+ the same measurement, roughly fifty times more than every duplicate lever combined.
+
+### Hardware and prerequisites
+
+Radeon RX 6600, 8 GB VRAM, 15 GB system RAM. This is the binding constraint on model choice:
+target a 7–8B at Q4 (~5 GB), leaving ~3 GB for KV cache. A 14B at Q4 (~9 GB) spills to CPU
+and makes corpus-wide generation impractical. `ollama-vulkan` 0.32.4 verified with `100% GPU`
+offload. The systemd unit ships **inactive and disabled**, with no `~/.ollama` and no models:
+
+```bash
+sudo systemctl enable --now ollama
+ollama pull qwen2.5:7b # ~4.7 GB Q4; llama3.1:8b is the alternative
+```
+
+The Claude Code lane needs the `claude` CLI installed and authenticated on the host, plus
+`digest.remoteEnabled` turned on in settings.
diff --git a/common/bin/compose-site.ts b/common/bin/compose-site.ts
@@ -22,11 +22,15 @@ import { getSite, resolveSocialLinks, resolveHubUrl, type Site } from "../lib/si
import { getSettings } from "../lib/settings";
import {
DUPLICATES_FILENAME,
+ DUPLICATE_OVERRIDES_FILENAME,
+ clusterIsPublishable,
filterClusterToChannels,
+ sanitizeDuplicateOverrides,
type DuplicateReport,
} from "../lib/duplicates";
import type { Manifest, SubsManifest } from "../lib/manifest";
import type { PostsManifest } from "../lib/posts";
+import type { DigestsManifest } from "../lib/digests";
import { buildSiteDescriptor, type PublicSiteDescriptor } from "../lib/siteDescriptor";
import { effectiveSiteAliases } from "../lib/aliasesStore";
import {
@@ -147,7 +151,25 @@ async function emitAiFiles(paths: ReturnType<typeof getPaths>): Promise<void> {
/* no posts manifest for this site */
}
- const corpus = buildSiteCorpus(descriptor, { hasArchives, postCounts });
+ // Per-channel digest counts, from the site digests manifest composed above.
+ // Absent (no digested channels) leaves corpus.json without a digest scheme.
+ const digestCounts: Record<string, number> = {};
+ try {
+ const raw = await readFile(
+ path.join(paths.exportDigestsDir, "manifest.json"),
+ "utf8",
+ );
+ const dm = JSON.parse(raw) as DigestsManifest;
+ for (const ch of dm.channels ?? []) digestCounts[ch.slug] = ch.digestCount;
+ } catch {
+ /* no digests manifest for this site */
+ }
+
+ const corpus = buildSiteCorpus(descriptor, {
+ hasArchives,
+ postCounts,
+ digestCounts,
+ });
await writeFile(
path.join(paths.exportPublicDir, "corpus.json"),
JSON.stringify(corpus),
@@ -489,6 +511,9 @@ type ComposeCache = {
// Optional for backwards compat: a cache written before the posts corpus
// existed simply has no entry, so every social channel composes once.
posts?: Record<string, string>;
+ // Same, for the AI-digest corpus: an existing cache composes digests once
+ // rather than erroring on a missing key.
+ digests?: Record<string, string>;
summaries?: string;
stats?: string;
duplicates?: string;
@@ -512,6 +537,7 @@ async function readComposeCache(p: string): Promise<ComposeCache> {
transcripts: parsed?.transcripts ?? {},
subs: parsed?.subs ?? {},
posts: parsed?.posts ?? {},
+ digests: parsed?.digests ?? {},
summaries: parsed?.summaries,
stats: parsed?.stats,
duplicates: parsed?.duplicates,
@@ -691,6 +717,17 @@ async function main(): Promise<void> {
cache.posts ?? {},
console.log,
);
+ // The AI-digest corpus: same shared-tree shape again. Sparse — only channels
+ // with at least one digest have a source dir, and reconcileChannelTree treats
+ // a missing one as "nothing to copy", so passing every member slug is right.
+ cache.digests = await reconcileChannelTree(
+ "digests",
+ paths.exportSharedDigestsDir,
+ paths.exportDigestsDir,
+ memberSlugs,
+ cache.digests ?? {},
+ console.log,
+ );
// Subs also carries a per-site manifest.json (tiny — copied every build).
const subsManifestSrc = path.join(
paths.exportSitesIndexDir,
@@ -712,6 +749,20 @@ async function main(): Promise<void> {
await mkdir(paths.exportPostsDir, { recursive: true });
await cp(postsManifestSrc, path.join(paths.exportPostsDir, "manifest.json"));
}
+ // Same for the per-site digests manifest (which channels carry digests).
+ const digestsManifestSrc = path.join(
+ paths.exportSitesIndexDir,
+ siteId,
+ "digests",
+ "manifest.json",
+ );
+ if (await exists(digestsManifestSrc)) {
+ await mkdir(paths.exportDigestsDir, { recursive: true });
+ await cp(
+ digestsManifestSrc,
+ path.join(paths.exportDigestsDir, "manifest.json"),
+ );
+ }
// --- charts dashboard ---
const templatesSrc = path.join(
@@ -743,11 +794,24 @@ async function main(): Promise<void> {
// site (per-site `duplicates` opt-out) AND there's at least one in-scope
// cluster — so its mere presence is what hasDuplicates() keys off to show the
// nav link. No file → the page shows its empty state and the link self-hides.
+ //
+ // UNCONFIRMED SUSPECTS ARE NOT SHIPPED. A `needsReview` cluster is evidence
+ // that two videos share a title and a runtime — nothing compared their
+ // content — so it is an internal review queue, not something to assert to a
+ // viewer. It reaches the public site only once a human records `confirmed` in
+ // duplicates.overrides.json, which is why that file (never read at compose
+ // time before) is read here.
const dupSrc = path.join(paths.transcriptsDir, DUPLICATES_FILENAME);
const dupDest = path.join(paths.exportPublicDir, DUPLICATES_FILENAME);
- // Gate the re-filter/re-serialize on the source's mtime+size, the member set,
- // and the feature flag. The cache value is prefixed written|/empty| so a skip
- // can self-heal if public/ was wiped out of band (presence must match).
+ const dupOverridesSrc = path.join(
+ paths.transcriptsDir,
+ DUPLICATE_OVERRIDES_FILENAME,
+ );
+ // Gate the re-filter/re-serialize on the source's mtime+size, the OVERRIDES'
+ // mtime+size (a confirmation changes what ships without touching the report),
+ // the member set, and the feature flag. The cache value is prefixed
+ // written|/empty| so a skip can self-heal if public/ was wiped out of band
+ // (presence must match).
let dupStat: { mtimeMs: number; size: number } | null = null;
try {
const s = await stat(dupSrc);
@@ -755,8 +819,18 @@ async function main(): Promise<void> {
} catch {
dupStat = null;
}
+ let dupOvStat: { mtimeMs: number; size: number } | null = null;
+ try {
+ const s = await stat(dupOverridesSrc);
+ dupOvStat = { mtimeMs: s.mtimeMs, size: s.size };
+ } catch {
+ dupOvStat = null;
+ }
const dupEnabled = site.duplicates !== false && dupStat !== null;
- const dupKey = `${dupEnabled}|${dupStat?.mtimeMs ?? ""}|${dupStat?.size ?? ""}|${[...memberSlugs].sort().join(",")}`;
+ const dupKey =
+ `${dupEnabled}|${dupStat?.mtimeMs ?? ""}|${dupStat?.size ?? ""}` +
+ `|${dupOvStat?.mtimeMs ?? ""}|${dupOvStat?.size ?? ""}` +
+ `|${[...memberSlugs].sort().join(",")}`;
const dupDestPresent = await exists(dupDest);
const dupCachedWritten = cache.duplicates?.startsWith("written|") ?? false;
const dupCacheHit =
@@ -772,7 +846,18 @@ async function main(): Promise<void> {
await readFile(dupSrc, "utf8"),
) as DuplicateReport;
const memberSet = new Set(memberSlugs);
+ // Never throws: an unreadable or malformed overrides file reads as "no
+ // decisions recorded", which fails CLOSED — every suspect stays internal.
+ let dupOverrides = sanitizeDuplicateOverrides(null);
+ try {
+ dupOverrides = sanitizeDuplicateOverrides(
+ JSON.parse(await readFile(dupOverridesSrc, "utf8")),
+ );
+ } catch {
+ // no decisions recorded
+ }
const clusters = report.clusters
+ .filter((c) => clusterIsPublishable(c, dupOverrides))
.map((c) => filterClusterToChannels(c, memberSet))
.filter((c): c is NonNullable<typeof c> => c !== null);
if (clusters.length > 0) {
diff --git a/common/bin/digest-bakeoff.ts b/common/bin/digest-bakeoff.ts
@@ -0,0 +1,716 @@
+#!/usr/bin/env tsx
+// Score digest engine candidates against a FIXED sample, without writing a
+// single digest to disk.
+//
+// WHY THIS IS A SCRIPT AND NOT INFRASTRUCTURE. It answers one question — which
+// (model x context x timestamp mode) should carry a multi-week sweep — and the
+// answer is a number in a table, not a feature. It therefore drives digestApps +
+// digestPrompt + digestParse DIRECTLY and scores in memory. It NEVER calls
+// digestVideo or writeDigestSection, so a losing candidate cannot leave anything
+// behind in the corpus, and no freshness record has to be invalidated afterwards
+// to undo a round.
+//
+// THROUGHPUT IS A FIRST-CLASS METRIC, not a footnote. Measured: 46 s per
+// 8,194-token chunk on qwen2.5:7b, which over the corpus is ~64 days on one
+// lane. A 14B at half the context roughly doubles the chunk count and halves the
+// token rate — order 250 days. A candidate can therefore be ruled out on
+// projected sweep days alone, however good its chapters look, which is why every
+// run prints days alongside quality.
+//
+// Two modes:
+//
+// --pick scan the stats cache and write the fixed stratified sample
+// (plans/bakeoff/sample.json). Run ONCE. A moving sample makes the
+// comparison between rounds meaningless.
+// (default) run the candidates over that sample and write a JSON + Markdown
+// report under plans/bakeoff/.
+//
+// Examples:
+// tsx bin/digest-bakeoff.ts --pick
+// tsx bin/digest-bakeoff.ts --label round1 --buckets short,medium \
+// --candidates 'qwen2.5:7b@16384,qwen3:8b@16384,gemma2:9b@16384'
+// tsx bin/digest-bakeoff.ts --label round2 --buckets long,verylong \
+// --candidates 'qwen2.5:7b@16384' --modes absolute,chunk-local
+
+import path from "node:path";
+import { mkdir, writeFile } from "node:fs/promises";
+import { readFile } from "node:fs/promises";
+import { open } from "lmdb";
+import { getPaths } from "../lib/paths";
+import { parseFlags } from "./_parseFlags";
+import type { VideoStat } from "../lib/stats";
+import { getDigestApp } from "../lib/digestApps";
+import type { DigestAppConfig, DigestTimestampMode } from "../lib/digest";
+import {
+ CHAPTER_SYSTEM_PROMPT,
+ DIGEST_OVERLAP_CUES,
+ buildChapterPrompt,
+ chapterSchema,
+ maxCuesForContext,
+ toHms,
+} from "../lib/digestPrompt";
+import { parseChapters, type DigestChunkOutput } from "../lib/digestParse";
+import { chunkCuesForContext } from "../lib/transcriptWindow";
+import { transcriptToMarkdown } from "../lib/transcriptToMarkdown";
+import { readNormalizedTranscript } from "../controller/normalizeTranscript";
+import { CUES_JSON_FILENAME } from "../lib/videoStatus";
+import type { Cue } from "../lib/vtt";
+
+// ---------------------------------------------------------------------------
+// The sample
+// ---------------------------------------------------------------------------
+
+// Duration strata. Chosen to match the corpus shape recorded in FACTS.md rather
+// than to be round numbers: the >4 h bucket is 8.2% of videos but 46% of all
+// transcript tokens, so a sample that under-represents it measures the cheap
+// half of the sweep and misses the half where chunk-seam bugs live.
+const BUCKETS = ["short", "medium", "long", "verylong"] as const;
+type Bucket = (typeof BUCKETS)[number];
+
+const BUCKET_BOUNDS: Record<Bucket, { min: number; max: number; want: number }> = {
+ short: { min: 5 * 60, max: 30 * 60, want: 3 },
+ medium: { min: 45 * 60, max: 90 * 60, want: 2 },
+ long: { min: 3 * 3600, max: 4.5 * 3600, want: 2 },
+ verylong: { min: 6 * 3600, max: 14 * 3600, want: 1 },
+};
+
+type SampleVideo = {
+ slug: string;
+ channelSlug: string;
+ videoId: string;
+ videoDir: string;
+ title: string;
+ bucket: Bucket;
+ durationSeconds: number;
+ cueCount: number;
+};
+
+type Sample = {
+ version: 1;
+ pickedAt: string;
+ // Corpus-wide totals, captured at pick time. The sweep-days projection is
+ // computed from measured seconds-per-audio-hour times THIS number, so the
+ // projection and the sample come from one scan and can't drift apart.
+ corpus: {
+ videosScanned: number;
+ videosWithTranscript: number;
+ audioHours: number;
+ longTailVideos: number;
+ longTailAudioHours: number;
+ };
+ videos: SampleVideo[];
+};
+
+function bucketFor(seconds: number): Bucket | null {
+ for (const b of BUCKETS) {
+ const { min, max } = BUCKET_BOUNDS[b];
+ if (seconds >= min && seconds <= max) return b;
+ }
+ return null;
+}
+
+// Deterministic pick, so re-running --pick on an unchanged corpus reproduces the
+// same sample. No Math.random: a sample that moves between rounds is not a
+// sample, it is noise. Videos are ordered by a stable hash of the slug and the
+// first N per bucket are taken, spreading the pick across channels instead of
+// clustering on whichever channel sorts first.
+function stableHash(s: string): number {
+ let h = 2166136261;
+ for (let i = 0; i < s.length; i++) {
+ h ^= s.charCodeAt(i);
+ h = Math.imul(h, 16777619);
+ }
+ return h >>> 0;
+}
+
+async function pickSample(outPath: string): Promise<void> {
+ const paths = getPaths();
+ const root = open({ path: paths.lmdbPath, maxDbs: 12, compression: true });
+ const statsByPath = root.openDB<
+ { metaMs: number; stat: VideoStat },
+ [string, string]
+ >({ name: "statsByPath", encoding: "msgpack" });
+
+ const byBucket = new Map<Bucket, SampleVideo[]>();
+ for (const b of BUCKETS) byBucket.set(b, []);
+
+ let videosScanned = 0;
+ let videosWithTranscript = 0;
+ let totalSeconds = 0;
+ let longTailVideos = 0;
+ let longTailSeconds = 0;
+
+ for (const { key, value } of statsByPath.getRange()) {
+ const stat = value.stat;
+ videosScanned++;
+ if (!stat.hasTranscript || !(stat.duration > 0)) continue;
+ videosWithTranscript++;
+ totalSeconds += stat.duration;
+ if (stat.duration > 4 * 3600) {
+ longTailVideos++;
+ longTailSeconds += stat.duration;
+ }
+ const bucket = bucketFor(stat.duration);
+ if (!bucket) continue;
+ // A digest needs cues; a transcript flagged present but empty is useless
+ // here and would silently shrink a stratum.
+ if (!stat.cueCount || stat.cueCount < 30) continue;
+ byBucket.get(bucket)!.push({
+ slug: stat.slug,
+ channelSlug: stat.channelSlug,
+ videoId: stat.id,
+ videoDir: (key as [string, string])[1],
+ title: stat.title,
+ bucket,
+ durationSeconds: Math.round(stat.duration),
+ cueCount: stat.cueCount,
+ });
+ }
+ await root.close();
+
+ const videos: SampleVideo[] = [];
+ for (const b of BUCKETS) {
+ const pool = byBucket.get(b)!;
+ pool.sort((a, c) => stableHash(a.slug) - stableHash(c.slug));
+ // One per channel first, so a stratum can't come entirely from one
+ // creator's house style — a model that happens to suit one show would
+ // otherwise look like a model that suits the corpus.
+ const seenChannels = new Set<string>();
+ const spread: SampleVideo[] = [];
+ for (const v of pool) {
+ if (seenChannels.has(v.channelSlug)) continue;
+ seenChannels.add(v.channelSlug);
+ spread.push(v);
+ }
+ const want = BUCKET_BOUNDS[b].want;
+ const taken = (spread.length >= want ? spread : pool).slice(0, want);
+ if (taken.length < want) {
+ console.warn(
+ `Warning: bucket ${b} wanted ${want} videos but only ${taken.length} qualify.`,
+ );
+ }
+ videos.push(...taken);
+ }
+
+ const sample: Sample = {
+ version: 1,
+ pickedAt: new Date().toISOString(),
+ corpus: {
+ videosScanned,
+ videosWithTranscript,
+ audioHours: Math.round(totalSeconds / 3600),
+ longTailVideos,
+ longTailAudioHours: Math.round(longTailSeconds / 3600),
+ },
+ videos,
+ };
+ await mkdir(path.dirname(outPath), { recursive: true });
+ await writeFile(outPath, `${JSON.stringify(sample, null, 2)}\n`);
+ console.log(
+ `Scanned ${videosScanned} videos (${videosWithTranscript} with transcripts, ` +
+ `${sample.corpus.audioHours} audio-hours; ${longTailVideos} over 4 h holding ` +
+ `${sample.corpus.longTailAudioHours} h).`,
+ );
+ for (const v of videos) {
+ console.log(
+ ` ${v.bucket.padEnd(8)} ${toHms(v.durationSeconds)} ${v.cueCount
+ .toString()
+ .padStart(5)} cues ${v.slug} ${v.title.slice(0, 60)}`,
+ );
+ }
+ console.log(`Wrote ${outPath}`);
+}
+
+// ---------------------------------------------------------------------------
+// Scoring
+// ---------------------------------------------------------------------------
+
+// Titles that carry no information about what was actually said. A cheap proxy
+// for title quality: a model that segments correctly but names every section
+// "Discussion" has produced a table of contents nobody can navigate.
+const GENERIC_TITLE_RE =
+ /^(the\s+)?(intro(duction)?|outro|conclusion|discussion|continued|continuation|overview|summary|recap|closing( remarks)?|opening( remarks)?|final thoughts|misc(ellaneous)?|other|general|topics?|segment|section|chapter|part)\b/i;
+const GENERIC_TITLE_TAIL_RE = /\b(part|section|segment|chapter)\s+(\d+|one|two|three|four|five|six|seven|eight|nine|ten)$/i;
+
+function isGenericTitle(title: string): boolean {
+ const t = title.trim();
+ return GENERIC_TITLE_RE.test(t) || GENERIC_TITLE_TAIL_RE.test(t);
+}
+
+function normalizeTitle(title: string): string {
+ return title.toLowerCase().replace(/[^\p{L}\p{N}]+/gu, " ").trim();
+}
+
+type Candidate = {
+ key: string;
+ model: string;
+ // Absent means "the engine's default" — the same thing an unset
+ // settings.digest.apps[id].numCtx means, so the default row measures exactly
+ // what a default-configured sweep would do.
+ numCtx?: number;
+ maxCues: number;
+ timestampMode: DigestTimestampMode;
+ think?: boolean;
+};
+
+type VideoScore = {
+ slug: string;
+ bucket: Bucket;
+ durationSeconds: number;
+ chunks: number;
+ chunksFailed: number;
+ zeroYieldChunks: number;
+ kept: number;
+ maxGapSeconds: number;
+ engineSeconds: number;
+ inputTokens: number;
+ outputTokens: number;
+ warningsByCode: Record<string, number>;
+ genericTitles: number;
+ duplicateTitles: number;
+ // Kept so a table can be sanity-checked against real output by hand, which is
+ // the only way to catch a model that scores well and reads badly.
+ sampleTitles: string[];
+};
+
+type CandidateScore = {
+ candidate: Candidate;
+ videos: VideoScore[];
+ totals: {
+ videos: number;
+ audioHours: number;
+ chunks: number;
+ chunksFailed: number;
+ zeroYieldChunks: number;
+ zeroYieldRate: number;
+ kept: number;
+ chaptersPerHour: number;
+ maxGapSeconds: number;
+ meanGapSeconds: number;
+ genericTitleRate: number;
+ duplicateTitleRate: number;
+ warningsByCode: Record<string, number>;
+ rejectionRate: number;
+ engineSeconds: number;
+ tokensPerSecond: number;
+ secondsPerAudioHour: number;
+ projectedSweepDays: number;
+ };
+};
+
+function renderChunk(
+ meta: { id: string; title: string; channel?: string; duration?: number },
+ cues: Cue[],
+ offsetSeconds: number,
+): string {
+ return transcriptToMarkdown(
+ { ...meta, cues },
+ {
+ timestamps: true,
+ includeDescription: false,
+ includeTags: false,
+ stampForCue: (_clock, seconds) =>
+ toHms(Math.max(0, seconds - offsetSeconds)),
+ },
+ );
+}
+
+async function scoreVideo(
+ video: SampleVideo,
+ candidate: Candidate,
+ log: (m: string) => void,
+): Promise<VideoScore | null> {
+ const paths = getPaths();
+ const cuesPath = path.join(
+ paths.channelsDir,
+ video.channelSlug,
+ "data",
+ video.videoDir,
+ CUES_JSON_FILENAME,
+ );
+ const transcript = await readNormalizedTranscript(cuesPath);
+ if (!transcript || !transcript.cues?.length) {
+ log(` ${video.slug}: no transcript on disk, skipped`);
+ return null;
+ }
+ const cues = transcript.cues;
+ const chunks = chunkCuesForContext(cues, {
+ maxCues: candidate.maxCues,
+ overlapCues: DIGEST_OVERLAP_CUES,
+ });
+
+ const app = getDigestApp("ollama-direct");
+ const config: DigestAppConfig = {
+ model: candidate.model,
+ numCtx: candidate.numCtx,
+ temperature: 0,
+ timeoutMs: 20 * 60_000,
+ ...(candidate.think !== undefined ? { think: candidate.think } : {}),
+ };
+
+ const outputs: DigestChunkOutput[] = [];
+ const score: VideoScore = {
+ slug: video.slug,
+ bucket: video.bucket,
+ durationSeconds: video.durationSeconds,
+ chunks: chunks.length,
+ chunksFailed: 0,
+ zeroYieldChunks: 0,
+ kept: 0,
+ maxGapSeconds: 0,
+ engineSeconds: 0,
+ inputTokens: 0,
+ outputTokens: 0,
+ warningsByCode: {},
+ genericTitles: 0,
+ duplicateTitles: 0,
+ sampleTitles: [],
+ };
+
+ for (let i = 0; i < chunks.length; i++) {
+ const chunk = chunks[i];
+ const startSeconds = Math.max(0, Math.floor(chunk[0].start));
+ const endSeconds = Math.max(
+ startSeconds,
+ Math.ceil(chunk[chunk.length - 1].end || chunk[chunk.length - 1].start),
+ );
+ const offset = candidate.timestampMode === "chunk-local" ? startSeconds : 0;
+ const promptInput = {
+ title: transcript.title || video.videoId,
+ channel: transcript.channel || video.channelSlug,
+ startSeconds,
+ endSeconds,
+ transcript: renderChunk(
+ {
+ id: transcript.id,
+ title: transcript.title,
+ channel: transcript.channel,
+ duration: transcript.duration,
+ },
+ chunk,
+ offset,
+ ),
+ timestampMode: candidate.timestampMode,
+ };
+ try {
+ const result = await app.run({
+ system: CHAPTER_SYSTEM_PROMPT,
+ prompt: buildChapterPrompt(promptInput),
+ schema: chapterSchema(endSeconds - startSeconds),
+ config,
+ });
+ score.engineSeconds += result.durationMs / 1000;
+ score.inputTokens += result.inputTokens ?? 0;
+ score.outputTokens += result.outputTokens ?? 0;
+ outputs.push({
+ index: i,
+ startSeconds,
+ endSeconds,
+ data: result.data,
+ timestampMode: candidate.timestampMode,
+ });
+ } catch (err) {
+ score.chunksFailed++;
+ score.warningsByCode["chunk-failed"] =
+ (score.warningsByCode["chunk-failed"] ?? 0) + 1;
+ log(` ${video.slug} chunk ${i + 1}/${chunks.length} failed: ${(err as Error).message}`);
+ }
+ }
+
+ // ZERO-YIELD CHUNKS — the headline defect metric. Computed by parsing each
+ // chunk ALONE, because the merged parse cannot attribute a kept chapter back
+ // to the chunk that produced it, and "chunk 3 produced nothing" is precisely
+ // the failure this whole stage exists to fix. A chunk the engine never
+ // answered counts too: from the corpus's point of view the outcome is the
+ // same, an interval of the video with no chapters in it.
+ for (const out of outputs) {
+ if (parseChapters([out], cues).chapters.length === 0) score.zeroYieldChunks++;
+ }
+ score.zeroYieldChunks += score.chunksFailed;
+
+ const parsed = parseChapters(outputs, cues);
+ score.kept = parsed.chapters.length;
+ for (const w of parsed.warnings) {
+ score.warningsByCode[w.code] = (score.warningsByCode[w.code] ?? 0) + 1;
+ }
+
+ // MAX COVERAGE GAP — catches "summarised the tail, skipped the head". The
+ // leading gap (0 -> first chapter) and the trailing one (last chapter -> end)
+ // are included deliberately: a video whose chapters all sit in the last 20
+ // minutes has a coverage failure that consecutive-gap-only scoring hides.
+ const starts = parsed.chapters.map((c) => c.start);
+ const bounds = [0, ...starts, video.durationSeconds];
+ for (let i = 1; i < bounds.length; i++) {
+ score.maxGapSeconds = Math.max(score.maxGapSeconds, bounds[i] - bounds[i - 1]);
+ }
+
+ const seen = new Set<string>();
+ for (const c of parsed.chapters) {
+ if (isGenericTitle(c.title)) score.genericTitles++;
+ const n = normalizeTitle(c.title);
+ if (seen.has(n)) score.duplicateTitles++;
+ seen.add(n);
+ }
+ score.sampleTitles = parsed.chapters.slice(0, 8).map((c) => `${c.clock} ${c.title}`);
+
+ log(
+ ` ${video.slug} [${video.bucket}] ${chunks.length} chunk(s) → ${score.kept} chapter(s), ` +
+ `${score.zeroYieldChunks} zero-yield, ${Math.round(score.engineSeconds)}s engine`,
+ );
+ return score;
+}
+
+function aggregate(candidate: Candidate, videos: VideoScore[]): CandidateScore {
+ const sum = (f: (v: VideoScore) => number): number =>
+ videos.reduce((a, v) => a + f(v), 0);
+ const audioHours = sum((v) => v.durationSeconds) / 3600;
+ const chunks = sum((v) => v.chunks);
+ const kept = sum((v) => v.kept);
+ const engineSeconds = sum((v) => v.engineSeconds);
+ const tokens = sum((v) => v.inputTokens + v.outputTokens);
+
+ const warningsByCode: Record<string, number> = {};
+ for (const v of videos) {
+ for (const [code, n] of Object.entries(v.warningsByCode)) {
+ warningsByCode[code] = (warningsByCode[code] ?? 0) + n;
+ }
+ }
+ const rejections = Object.entries(warningsByCode)
+ .filter(([code]) => code !== "seam-duplicate")
+ .reduce((a, [, n]) => a + n, 0);
+
+ const secondsPerAudioHour = audioHours > 0 ? engineSeconds / audioHours : 0;
+
+ return {
+ candidate,
+ videos,
+ totals: {
+ videos: videos.length,
+ audioHours: round(audioHours, 2),
+ chunks,
+ chunksFailed: sum((v) => v.chunksFailed),
+ zeroYieldChunks: sum((v) => v.zeroYieldChunks),
+ zeroYieldRate: chunks > 0 ? round(sum((v) => v.zeroYieldChunks) / chunks, 4) : 0,
+ kept,
+ chaptersPerHour: audioHours > 0 ? round(kept / audioHours, 2) : 0,
+ maxGapSeconds: videos.reduce((a, v) => Math.max(a, v.maxGapSeconds), 0),
+ meanGapSeconds:
+ videos.length > 0 ? Math.round(sum((v) => v.maxGapSeconds) / videos.length) : 0,
+ genericTitleRate: kept > 0 ? round(sum((v) => v.genericTitles) / kept, 4) : 0,
+ duplicateTitleRate: kept > 0 ? round(sum((v) => v.duplicateTitles) / kept, 4) : 0,
+ warningsByCode,
+ // Rejections per kept chapter — the ratio that says how much of what the
+ // model produced the guards had to throw away.
+ rejectionRate: kept + rejections > 0 ? round(rejections / (kept + rejections), 4) : 0,
+ engineSeconds: Math.round(engineSeconds),
+ tokensPerSecond: engineSeconds > 0 ? round(tokens / engineSeconds, 1) : 0,
+ secondsPerAudioHour: Math.round(secondsPerAudioHour),
+ projectedSweepDays: 0, // filled in once corpus hours are known
+ },
+ };
+}
+
+function round(n: number, places: number): number {
+ const f = 10 ** places;
+ return Math.round(n * f) / f;
+}
+
+// ---------------------------------------------------------------------------
+// Report
+// ---------------------------------------------------------------------------
+
+function markdownReport(
+ label: string,
+ sample: Sample,
+ buckets: Bucket[],
+ scores: CandidateScore[],
+): string {
+ const lines: string[] = [];
+ lines.push(`# Digest bake-off — ${label}`);
+ lines.push("");
+ lines.push(
+ `Sample: ${scores[0]?.totals.videos ?? 0} video(s) from \`plans/bakeoff/sample.json\`` +
+ ` (buckets: ${buckets.join(", ")}), ${scores[0]?.totals.audioHours ?? 0} audio-hours.`,
+ );
+ lines.push(
+ `Sweep days are projected as measured seconds-per-audio-hour x ` +
+ `${sample.corpus.audioHours} corpus audio-hours, one lane, no parallelism.`,
+ );
+ lines.push("");
+ lines.push(
+ "| Candidate | Zero-yield chunks | Chapters/h | Rejection rate | Max gap | Generic | Dup | tok/s | s per audio-h | **Sweep days** |",
+ );
+ lines.push(
+ "| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |",
+ );
+ for (const s of scores) {
+ const t = s.totals;
+ lines.push(
+ `| \`${s.candidate.key}\` | ${t.zeroYieldChunks}/${t.chunks} (${pct(t.zeroYieldRate)}) | ` +
+ `${t.chaptersPerHour} | ${pct(t.rejectionRate)} | ${toHms(t.maxGapSeconds)} | ` +
+ `${pct(t.genericTitleRate)} | ${pct(t.duplicateTitleRate)} | ${t.tokensPerSecond} | ` +
+ `${t.secondsPerAudioHour} | **${t.projectedSweepDays}** |`,
+ );
+ }
+ lines.push("");
+ lines.push("## Rejections by guard");
+ lines.push("");
+ const codes = Array.from(
+ new Set(scores.flatMap((s) => Object.keys(s.totals.warningsByCode))),
+ ).sort();
+ lines.push(`| Candidate | ${codes.join(" | ")} |`);
+ lines.push(`| --- | ${codes.map(() => "---").join(" | ")} |`);
+ for (const s of scores) {
+ lines.push(
+ `| \`${s.candidate.key}\` | ${codes
+ .map((c) => s.totals.warningsByCode[c] ?? 0)
+ .join(" | ")} |`,
+ );
+ }
+ lines.push("");
+ lines.push("## Per-video");
+ lines.push("");
+ lines.push("| Candidate | Video | Bucket | Chunks | Zero-yield | Chapters | Max gap | Engine s |");
+ lines.push("| --- | --- | --- | --- | --- | --- | --- | --- |");
+ for (const s of scores) {
+ for (const v of s.videos) {
+ lines.push(
+ `| \`${s.candidate.key}\` | \`${v.slug}\` | ${v.bucket} | ${v.chunks} | ` +
+ `${v.zeroYieldChunks} | ${v.kept} | ${toHms(v.maxGapSeconds)} | ${Math.round(v.engineSeconds)} |`,
+ );
+ }
+ }
+ lines.push("");
+ lines.push("## Sample output (first chapters per video)");
+ lines.push("");
+ for (const s of scores) {
+ lines.push(`### \`${s.candidate.key}\``);
+ lines.push("");
+ for (const v of s.videos) {
+ lines.push(`**${v.slug}** (${v.bucket}, ${toHms(v.durationSeconds)})`);
+ lines.push("");
+ for (const t of v.sampleTitles) lines.push(`- ${t}`);
+ if (v.sampleTitles.length === 0) lines.push("- _(nothing survived the guards)_");
+ lines.push("");
+ }
+ }
+ return `${lines.join("\n")}\n`;
+}
+
+function pct(v: number): string {
+ return `${Math.round(v * 1000) / 10}%`;
+}
+
+// ---------------------------------------------------------------------------
+// Main
+// ---------------------------------------------------------------------------
+
+// "model@ctx" or "model@ctx:think" / "model@ctx:nothink".
+//
+// maxCues is DERIVED from the context rather than configured: the 1200-cue
+// default is sized for a 16k window (~10 tokens/cue -> ~12k tokens of transcript
+// plus room for prompt and response), so halving the context must halve the
+// slice or every call silently truncates — the exact failure that made the first
+// smoke test summarize a fragment.
+function parseCandidate(spec: string, mode: DigestTimestampMode): Candidate {
+ const [modelPart, rest] = spec.split("@");
+ const [ctxPart, thinkPart] = (rest ?? "").split(":");
+ const numCtx = ctxPart ? Number(ctxPart) : undefined;
+ // The SAME derivation production uses, not a parallel copy — otherwise the
+ // bake-off scores a chunk size the sweep would never actually run.
+ const maxCues = maxCuesForContext(numCtx);
+ return {
+ key: `${modelPart}@${numCtx ?? "default"}/${mode}`,
+ model: modelPart,
+ ...(numCtx ? { numCtx } : {}),
+ maxCues,
+ timestampMode: mode,
+ ...(thinkPart === "think"
+ ? { think: true }
+ : thinkPart === "nothink"
+ ? { think: false }
+ : {}),
+ };
+}
+
+async function main(): Promise<void> {
+ const flags = parseFlags(process.argv.slice(2));
+ const outDir = flags.outDir ?? path.join(process.cwd(), "..", "plans", "bakeoff");
+ const samplePath = flags.sample ?? path.join(outDir, "sample.json");
+
+ if (flags.pick === "true") {
+ await pickSample(samplePath);
+ return;
+ }
+
+ const sample = JSON.parse(await readFile(samplePath, "utf8")) as Sample;
+ const label = flags.label ?? "round";
+ const buckets = (flags.buckets ?? BUCKETS.join(","))
+ .split(",")
+ .map((b) => b.trim())
+ .filter((b): b is Bucket => (BUCKETS as readonly string[]).includes(b));
+ const modes = (flags.modes ?? "absolute")
+ .split(",")
+ .map((m) => m.trim())
+ .filter((m): m is DigestTimestampMode => m === "absolute" || m === "chunk-local");
+ const specs = (flags.candidates ?? "qwen2.5:7b@16384")
+ .split(",")
+ .map((s) => s.trim())
+ .filter(Boolean);
+
+ const videos = sample.videos.filter((v) => buckets.includes(v.bucket));
+ if (videos.length === 0) throw new Error(`No sample videos in buckets ${buckets.join(",")}`);
+
+ const candidates: Candidate[] = [];
+ for (const spec of specs) {
+ for (const mode of modes) candidates.push(parseCandidate(spec, mode));
+ }
+
+ console.log(
+ `Bake-off ${label}: ${candidates.length} candidate(s) x ${videos.length} video(s) ` +
+ `(${Math.round(videos.reduce((a, v) => a + v.durationSeconds, 0) / 3600)} audio-hours each).`,
+ );
+
+ const scores: CandidateScore[] = [];
+ for (const candidate of candidates) {
+ console.log(`\n=== ${candidate.key} (maxCues ${candidate.maxCues}) ===`);
+ const startedAt = Date.now();
+ const perVideo: VideoScore[] = [];
+ for (const video of videos) {
+ const s = await scoreVideo(video, candidate, (m) => console.log(m));
+ if (s) perVideo.push(s);
+ }
+ const agg = aggregate(candidate, perVideo);
+ // Days, from measured seconds-per-audio-hour against the corpus total
+ // captured in the same scan that picked the sample.
+ agg.totals.projectedSweepDays = round(
+ (agg.totals.secondsPerAudioHour * sample.corpus.audioHours) / 86400,
+ 1,
+ );
+ scores.push(agg);
+ console.log(
+ `--- ${candidate.key}: ${agg.totals.kept} chapters, ` +
+ `${agg.totals.zeroYieldChunks}/${agg.totals.chunks} zero-yield, ` +
+ `${agg.totals.tokensPerSecond} tok/s, ` +
+ `projected ${agg.totals.projectedSweepDays} sweep days ` +
+ `(wall ${Math.round((Date.now() - startedAt) / 60000)} min)`,
+ );
+ }
+
+ scores.sort((a, b) => a.totals.zeroYieldRate - b.totals.zeroYieldRate);
+
+ await mkdir(outDir, { recursive: true });
+ const jsonPath = path.join(outDir, `${label}.json`);
+ const mdPath = path.join(outDir, `${label}.md`);
+ await writeFile(
+ jsonPath,
+ `${JSON.stringify({ label, sample: samplePath, corpus: sample.corpus, buckets, scores }, null, 2)}\n`,
+ );
+ await writeFile(mdPath, markdownReport(label, sample, buckets, scores));
+ console.log(`\nWrote ${jsonPath}\nWrote ${mdPath}`);
+}
+
+main().catch((err) => {
+ console.error(err);
+ process.exit(1);
+});
diff --git a/common/bin/digest-plan.ts b/common/bin/digest-plan.ts
@@ -0,0 +1,136 @@
+#!/usr/bin/env tsx
+// Price the digest backfill, in audio-hours, before spending GPU-weeks on it.
+//
+// Reads the last duplicate-detection run and build:stats' cache; writes nothing.
+// Run it before and after a cheapness lever (a --near threshold change, a batch
+// of confirmed clusters) and diff `generate`: that number, not a video count, is
+// what the sweep's wall clock is made of.
+//
+// Flags:
+// --lane local|remote which engine identity to test freshness against
+// --rate N seconds per audio-hour (default 90, the measured local
+// rate from the 102-video validation run)
+// --no-freshness skip the per-video sidecar read; reports the
+// from-scratch cost instead of the remaining one
+// --channels a,b restrict to these channel slugs
+// --top N how many channels to list (default 15, 0 = all)
+// --json machine-readable, for diffing two runs
+import { getPaths } from "../lib/paths";
+import {
+ buildDigestSweepPlan,
+ audioHours,
+ sweepDays,
+ DIGEST_PLAN_ROLES,
+ MEASURED_SECONDS_PER_AUDIO_HOUR,
+ roleMustGenerate,
+} from "../controller/digestPlan";
+import { parseFlags } from "./_parseFlags";
+
+const flags = parseFlags(process.argv.slice(2));
+
+const lane = flags.lane === "remote" ? "remote" : "local";
+const rate =
+ flags.rate !== undefined
+ ? Number(flags.rate)
+ : MEASURED_SECONDS_PER_AUDIO_HOUR;
+const top = flags.top !== undefined ? Number(flags.top) : 15;
+const asJson = flags.json === "true";
+
+function hours(seconds: number): string {
+ return audioHours(seconds).toLocaleString("en-US", {
+ maximumFractionDigits: 0,
+ });
+}
+
+function pct(part: number, whole: number): string {
+ return whole > 0 ? `${((part / whole) * 100).toFixed(1)}%` : "—";
+}
+
+async function main(): Promise<void> {
+ const plan = await buildDigestSweepPlan({
+ paths: getPaths(),
+ lane,
+ checkFreshness: flags["no-freshness"] !== "true",
+ channelSlugs: flags.channels ? flags.channels.split(",") : undefined,
+ onLog: (m) => {
+ if (!asJson) console.log(m);
+ },
+ });
+
+ if (asJson) {
+ console.log(
+ JSON.stringify(
+ {
+ ...plan,
+ secondsPerAudioHour: rate,
+ generateAudioHours: audioHours(plan.generateSeconds),
+ sharedAudioHours: audioHours(plan.sharedSeconds),
+ sweepDays: sweepDays(plan.generateSeconds, rate),
+ },
+ null,
+ 2,
+ ),
+ );
+ return;
+ }
+
+ const eligible =
+ plan.generateSeconds + plan.sharedSeconds + plan.fresh.audioSeconds;
+
+ console.log("");
+ console.log(
+ `Digest sweep plan — ${lane} lane, ${rate}s per audio-hour` +
+ (plan.freshnessChecked ? "" : ", FROM SCRATCH (freshness not checked)"),
+ );
+ console.log(
+ `Duplicate report: ${plan.clusters.toLocaleString()} cluster(s), ${plan.clusterMembersMapped.toLocaleString()} member(s) mapped.`,
+ );
+ if (plan.statsSchemaStale)
+ console.log("!! stats cache is stale — run build:stats.");
+ console.log("");
+
+ console.log("Remaining work by role (audio-hours):");
+ for (const role of DIGEST_PLAN_ROLES) {
+ const t = plan.remaining[role];
+ console.log(
+ ` ${role.padEnd(17)} ${hours(t.audioSeconds).padStart(8)} h ` +
+ `${t.videos.toLocaleString().padStart(7)} videos ` +
+ `${roleMustGenerate(role) ? "GENERATE" : "shared free"}`,
+ );
+ }
+ console.log(
+ ` ${"already fresh".padEnd(17)} ${hours(plan.fresh.audioSeconds).padStart(8)} h ` +
+ `${plan.fresh.videos.toLocaleString().padStart(7)} videos skipped`,
+ );
+ console.log(
+ ` ${"ineligible".padEnd(17)} ${"—".padStart(8)} ${plan.ineligible.toLocaleString().padStart(7)} videos no transcript`,
+ );
+ console.log("");
+
+ console.log(
+ `TO GENERATE : ${hours(plan.generateSeconds)} audio-hours → ${sweepDays(plan.generateSeconds, rate).toFixed(1)} sweep days`,
+ );
+ console.log(
+ `SAVED by sharing: ${hours(plan.sharedSeconds)} audio-hours (${pct(plan.sharedSeconds, eligible)} of eligible) → ${sweepDays(plan.sharedSeconds, rate).toFixed(1)} days avoided`,
+ );
+ console.log("");
+
+ const listed = top > 0 ? plan.channels.slice(0, top) : plan.channels;
+ console.log(`Heaviest channels (the sweep's work queue order):`);
+ for (const c of listed) {
+ if (c.generateSeconds <= 0) continue;
+ console.log(
+ ` ${c.channelSlug.padEnd(28)} ${hours(c.generateSeconds).padStart(7)} h generate ` +
+ `${hours(c.sharedSeconds).padStart(6)} h shared ` +
+ `${sweepDays(c.generateSeconds, rate).toFixed(1).padStart(5)} d`,
+ );
+ }
+ const rest = plan.channels.length - listed.length;
+ if (rest > 0) console.log(` … and ${rest} more channel(s).`);
+ console.log("");
+}
+
+main().catch((err) => {
+ console.error(err);
+ process.exit(1);
+});
diff --git a/common/bin/digest-validate.ts b/common/bin/digest-validate.ts
@@ -0,0 +1,213 @@
+#!/usr/bin/env tsx
+// Score digests that ALREADY EXIST on disk, using the same metric definitions as
+// the bake-off harness (common/bin/digest-bakeoff.ts).
+//
+// The bake-off generates into a scratch directory and scores what it generated;
+// this reads what a real sweep wrote. That is the difference between "how does
+// this configuration behave on a hand-picked sample" and "how did it behave on
+// the corpus", and the second is what a validation run is for.
+//
+// Usage: pnpm exec tsx bin/digest-validate.ts <channel> [<channel> ...]
+
+import path from "node:path";
+import { readdir, readFile } from "node:fs/promises";
+import { getPaths } from "../lib/paths";
+import { loadDigest } from "../lib/digest-server";
+import { toHms } from "../lib/digestPrompt";
+import type { DigestRecord, DigestWarningCode } from "../lib/digest";
+
+// Copied verbatim from digest-bakeoff.ts so the two are comparable. A model that
+// segments correctly but names every section "Discussion" has produced a table
+// of contents nobody can navigate.
+const GENERIC_TITLE_RE =
+ /^(the\s+)?(intro(duction)?|outro|conclusion|discussion|continued|continuation|overview|summary|recap|closing( remarks)?|opening( remarks)?|final thoughts|misc(ellaneous)?|other|general|topics?|segment|section|chapter|part)\b/i;
+const GENERIC_TITLE_TAIL_RE =
+ /\b(part|section|segment|chapter)\s+(\d+|one|two|three|four|five|six|seven|eight|nine|ten)$/i;
+
+function isGenericTitle(title: string): boolean {
+ const t = title.trim();
+ return GENERIC_TITLE_RE.test(t) || GENERIC_TITLE_TAIL_RE.test(t);
+}
+
+function normalizeTitle(title: string): string {
+ return title.trim().toLowerCase().replace(/\s+/g, " ");
+}
+
+type VideoScore = {
+ slug: string;
+ durationSeconds: number;
+ chunks: number;
+ chunksOk: number;
+ kept: number;
+ maxGapSeconds: number;
+ generic: number;
+ duplicate: number;
+ warnings: number;
+ derived: boolean;
+};
+
+async function readDuration(videoDir: string): Promise<number> {
+ try {
+ const raw = await readFile(
+ path.join(videoDir, "transcript.cues.json"),
+ "utf8",
+ );
+ const j = JSON.parse(raw);
+ if (typeof j?.duration === "number" && j.duration > 0) return j.duration;
+ const cues = Array.isArray(j?.cues) ? j.cues : [];
+ return cues.length > 0 ? Number(cues[cues.length - 1]?.end ?? 0) : 0;
+ } catch {
+ return 0;
+ }
+}
+
+function scoreVideo(
+ slug: string,
+ record: DigestRecord,
+ durationSeconds: number,
+): VideoScore {
+ const section = record.sections.chapters;
+ const items = section?.items ?? [];
+ const p = section?.provenance;
+
+ // MAX COVERAGE GAP, including the leading and trailing gaps — a video whose
+ // chapters all sit in the last 20 minutes has a coverage failure that
+ // consecutive-gap-only scoring hides.
+ const starts = items.map((c) => c.start).sort((a, b) => a - b);
+ const bounds = [0, ...starts, durationSeconds];
+ let maxGap = 0;
+ for (let i = 1; i < bounds.length; i++) {
+ maxGap = Math.max(maxGap, bounds[i] - bounds[i - 1]);
+ }
+
+ let generic = 0;
+ let duplicate = 0;
+ const seen = new Set<string>();
+ for (const c of items) {
+ if (isGenericTitle(c.title)) generic++;
+ const n = normalizeTitle(c.title);
+ if (seen.has(n)) duplicate++;
+ seen.add(n);
+ }
+
+ return {
+ slug,
+ durationSeconds,
+ chunks: p?.chunks ?? 0,
+ chunksOk: p?.chunksOk ?? 0,
+ kept: items.length,
+ maxGapSeconds: durationSeconds > 0 ? maxGap : 0,
+ generic,
+ duplicate,
+ warnings: record.warnings.length,
+ derived: record.derivedFrom != null,
+ };
+}
+
+async function main(): Promise<void> {
+ const channels = process.argv.slice(2);
+ if (channels.length === 0) {
+ console.error("usage: digest-validate.ts <channel> [<channel> ...]");
+ process.exit(1);
+ }
+ const paths = getPaths();
+ const videos: VideoScore[] = [];
+ const byCode = new Map<DigestWarningCode, number>();
+ const provenanceSeen = new Map<string, number>();
+
+ for (const channelSlug of channels) {
+ const dataDir = path.join(paths.channelsDir, channelSlug, "data");
+ const ids = await readdir(dataDir).catch(() => [] as string[]);
+ for (const id of ids) {
+ const videoDir = path.join(dataDir, id);
+ const record = await loadDigest(videoDir);
+ if (!record) continue;
+ const duration = await readDuration(videoDir);
+ videos.push(scoreVideo(`${channelSlug}/${id}`, record, duration));
+ for (const w of record.warnings) {
+ byCode.set(w.code, (byCode.get(w.code) ?? 0) + 1);
+ }
+ const p = record.sections.chapters?.provenance;
+ if (p) {
+ const key = `${p.appId}/${p.modelRequested ?? p.model} v${p.promptVersion}${
+ p.promptVariant ? `+${p.promptVariant}` : ""
+ }`;
+ provenanceSeen.set(key, (provenanceSeen.get(key) ?? 0) + 1);
+ }
+ }
+ }
+
+ if (videos.length === 0) {
+ console.log("No digests found.");
+ return;
+ }
+
+ // Videos whose digest was SHARED from a cluster canonical are excluded from
+ // the quality metrics: they measure the canonical's generation, not this
+ // video's, and counting them twice would flatter whichever number they help.
+ const native = videos.filter((v) => !v.derived);
+ const sum = (f: (v: VideoScore) => number): number =>
+ native.reduce((a, v) => a + f(v), 0);
+ const chunks = sum((v) => v.chunks);
+ const chunksOk = sum((v) => v.chunksOk);
+ const zeroYield = chunks - chunksOk;
+ const kept = sum((v) => v.kept);
+ const audioHours = sum((v) => v.durationSeconds) / 3600;
+ const warningsTotal = sum((v) => v.warnings);
+ const pct = (n: number): string => `${(n * 100).toFixed(1)}%`;
+
+ console.log(`Digests scored: ${videos.length} (${native.length} natively generated, ${videos.length - native.length} shared from a duplicate)`);
+ console.log(`Audio hours (native): ${audioHours.toFixed(1)}`);
+ console.log("");
+ console.log(`Zero-yield chunks: ${zeroYield}/${chunks} (${chunks > 0 ? pct(zeroYield / chunks) : "n/a"})`);
+ console.log(`Chapters/hour: ${audioHours > 0 ? (kept / audioHours).toFixed(2) : "n/a"} (${kept} kept)`);
+ console.log(`Generic titles: ${sum((v) => v.generic)}/${kept} (${kept > 0 ? pct(sum((v) => v.generic) / kept) : "n/a"})`);
+ console.log(`Duplicate titles: ${sum((v) => v.duplicate)}/${kept} (${kept > 0 ? pct(sum((v) => v.duplicate) / kept) : "n/a"})`);
+ console.log(`Warnings recorded: ${warningsTotal}`);
+ const rejected = warningsTotal;
+ console.log(`Rejection rate: ${kept + rejected > 0 ? pct(rejected / (kept + rejected)) : "n/a"} (rejected / (kept + rejected))`);
+
+ const gaps = native
+ .filter((v) => v.durationSeconds > 0)
+ .sort((a, b) => b.maxGapSeconds - a.maxGapSeconds);
+ console.log(
+ `Worst coverage gap: ${gaps[0] ? `${toHms(gaps[0].maxGapSeconds)} (${gaps[0].slug}, ${toHms(gaps[0].durationSeconds)} long)` : "n/a"}`,
+ );
+ const medianGap = gaps.length > 0 ? gaps[Math.floor(gaps.length / 2)].maxGapSeconds : 0;
+ console.log(`Median coverage gap: ${toHms(medianGap)}`);
+
+ console.log("");
+ console.log("Rejections by guard:");
+ if (byCode.size === 0) console.log(" (none)");
+ for (const [code, n] of [...byCode.entries()].sort((a, b) => b[1] - a[1])) {
+ console.log(` ${code.padEnd(20)} ${n}`);
+ }
+
+ console.log("");
+ console.log("Provenance seen (this is what freshness compares):");
+ for (const [key, n] of [...provenanceSeen.entries()].sort((a, b) => b[1] - a[1])) {
+ console.log(` ${key} ×${n}`);
+ }
+
+ console.log("");
+ console.log("Worst 8 by coverage gap:");
+ for (const v of gaps.slice(0, 8)) {
+ console.log(
+ ` ${toHms(v.maxGapSeconds).padStart(9)} of ${toHms(v.durationSeconds).padStart(9)} ${String(v.kept).padStart(3)} ch ${v.slug}`,
+ );
+ }
+
+ const zeroChapter = native.filter((v) => v.kept === 0);
+ if (zeroChapter.length > 0) {
+ console.log("");
+ console.log(`Videos with NO chapters at all: ${zeroChapter.length}`);
+ for (const v of zeroChapter.slice(0, 10)) {
+ console.log(` ${v.slug} (${toHms(v.durationSeconds)})`);
+ }
+ }
+}
+
+main().catch((err) => {
+ console.error(err);
+ process.exit(1);
+});
diff --git a/common/bin/duplicate-shorts.ts b/common/bin/duplicate-shorts.ts
@@ -1,15 +1,21 @@
#!/usr/bin/env tsx
-// On-demand duplicate-shorts detection. Runs AFTER build:index + build:stats
-// (it reads the cues + statsByPath those populate), so it is deliberately not
-// chained into build:data.
+// On-demand duplicate detection. Runs AFTER build:index + build:stats (it reads
+// the cues + statsByPath those populate), so it is deliberately not chained into
+// build:data.
//
// Flags:
// --threshold N duration cutoff in seconds (default 180)
-// --all-durations no length filter; enables containment matching
+// --all-durations no length filter; enables the containment sweep
+// --blocking S how candidate pairs are nominated: title | duration |
+// both. Defaults to `title` corpus-wide and `duration` in
+// shorts mode. Nomination is never evidence — every pair is
+// then judged on transcript content — so this trades recall
+// against runtime, not correctness.
// --near F Jaccard near-duplicate threshold (default 0.6)
// --tolerance N duration bucketing window in seconds (default 2)
import { getPaths } from "../lib/paths";
import { detectDuplicateShorts } from "../controller/duplicateShorts";
+import type { DuplicateBlocking } from "../lib/duplicates";
import { parseFlags } from "./_parseFlags";
const flags = parseFlags(process.argv.slice(2));
@@ -21,9 +27,22 @@ const thresholdSeconds =
? Number(flags.threshold)
: undefined;
+const BLOCKING: ReadonlyArray<DuplicateBlocking> = ["title", "duration", "both"];
+const blockingFlag = flags.blocking;
+if (
+ blockingFlag !== undefined &&
+ !BLOCKING.includes(blockingFlag as DuplicateBlocking)
+) {
+ console.error(
+ `--blocking must be one of: ${BLOCKING.join(", ")} (got "${blockingFlag}")`,
+ );
+ process.exit(1);
+}
+
detectDuplicateShorts({
paths: getPaths(),
thresholdSeconds,
+ blocking: blockingFlag as DuplicateBlocking | undefined,
nearThreshold: flags.near !== undefined ? Number(flags.near) : undefined,
durationToleranceSeconds:
flags.tolerance !== undefined ? Number(flags.tolerance) : undefined,
diff --git a/common/components/PlayerProvider.tsx b/common/components/PlayerProvider.tsx
@@ -14,6 +14,8 @@ import {
import type ReactPlayerType from "react-player";
import { fetchTranscript } from "./transcriptCache";
import { fetchSubs } from "./subsCache";
+import { fetchDigest, hasDigest } from "./digestCache";
+import type { VideoDigest } from "../lib/digests";
import { useUrlParams, writeUrlParams } from "./urlState";
import type { ModalMode } from "./urlState";
import { formatDate, formatDuration } from "../lib/format";
@@ -78,6 +80,9 @@ export type TranscriptData = {
type Status = "idle" | "loading" | "ready";
type ChatStatus = "idle" | "loading" | "ready" | "missing";
+// Parallel to ChatStatus. "missing" is a first-class outcome here rather than
+// an edge case: only ~0.1% of the corpus is digested.
+export type DigestStatus = "idle" | "loading" | "ready" | "missing";
export type DisplayMode = "modal" | "mini" | "hidden";
type Detail = {
@@ -111,7 +116,18 @@ type PlayerState = {
modalMode: ModalMode;
chatCues: Cue[] | null;
chatStatus: ChatStatus;
- chatNotice: string | null;
+ // The modal's transient snap-back banner, shared by every mode that can fail
+ // to load and fall back to the transcript (live chat, digest). One channel,
+ // because only one such message can be relevant at a time.
+ modalNotice: string | null;
+ digest: VideoDigest | null;
+ digestStatus: DigestStatus;
+ // Whether the ACTIVE video has a digest at all, resolved from the channel
+ // manifest without fetching a page. Gates the modal's Digest control: with
+ // coverage at 0.1% of the corpus an always-present button would be a dead end
+ // on almost every video. Same idea as hasDuplicates() gating the Duplicates
+ // nav link (export/app/lib/duplicates.ts).
+ digestAvailable: boolean;
clipStart: number | null;
clipEnd: number | null;
openTranscript: (
@@ -190,6 +206,16 @@ type Core = {
cues: Cue[] | null;
notice: string | null;
};
+ // The derived layer, shaped exactly like `chat` and fetched the same lazy
+ // way. `available` is resolved separately (and eagerly, on slug change) from
+ // the channel manifest, because the toolbar needs it before anyone asks for
+ // the panel.
+ digest: {
+ slug: string | null;
+ status: DigestStatus;
+ data: VideoDigest | null;
+ available: boolean;
+ };
};
const INITIAL_CORE: Core = {
@@ -197,6 +223,7 @@ const INITIAL_CORE: Core = {
displayState: null,
clip: { slug: null, start: null, end: null },
chat: { slug: null, status: "idle", cues: null, notice: null },
+ digest: { slug: null, status: "idle", data: null, available: false },
};
type CoreAction =
@@ -211,7 +238,12 @@ type CoreAction =
| { type: "CHAT_LOADING"; slug: string }
| { type: "CHAT_READY"; slug: string; cues: Cue[] }
| { type: "CHAT_MISSING"; slug: string; notice: string }
- | { type: "CHAT_NOTICE_CLEAR" };
+ | { type: "MODAL_NOTICE_CLEAR" }
+ | { type: "DIGEST_RESET" }
+ | { type: "DIGEST_AVAILABLE"; slug: string; available: boolean }
+ | { type: "DIGEST_LOADING"; slug: string }
+ | { type: "DIGEST_READY"; slug: string; data: VideoDigest }
+ | { type: "DIGEST_MISSING"; slug: string; notice: string };
function coreReducer(state: Core, action: CoreAction): Core {
switch (action.type) {
@@ -273,8 +305,53 @@ function coreReducer(state: Core, action: CoreAction): Core {
notice: action.notice,
},
};
- case "CHAT_NOTICE_CLEAR":
+ case "MODAL_NOTICE_CLEAR":
return { ...state, chat: { ...state.chat, notice: null } };
+ case "DIGEST_RESET":
+ return { ...state, digest: INITIAL_CORE.digest };
+ case "DIGEST_AVAILABLE":
+ return {
+ ...state,
+ digest: {
+ ...state.digest,
+ // Availability is per-video and arrives asynchronously; keep whatever
+ // slug the fetch/status fields already refer to unless this is the
+ // first thing we know about this video.
+ slug: state.digest.slug ?? action.slug,
+ available: action.available,
+ },
+ };
+ case "DIGEST_LOADING":
+ return {
+ ...state,
+ digest: {
+ slug: action.slug,
+ status: "loading",
+ data: null,
+ available: state.digest.available,
+ },
+ };
+ case "DIGEST_READY":
+ return {
+ ...state,
+ digest: {
+ slug: action.slug,
+ status: "ready",
+ data: action.data,
+ available: true,
+ },
+ };
+ case "DIGEST_MISSING":
+ return {
+ ...state,
+ chat: { ...state.chat, notice: action.notice },
+ digest: {
+ slug: action.slug,
+ status: "missing",
+ data: null,
+ available: false,
+ },
+ };
default:
return state;
}
@@ -287,7 +364,7 @@ export function PlayerProvider({
}) {
const { v: urlSlug, t: urlTime, vm: urlVm } = useUrlParams();
const [core, dispatch] = useReducer(coreReducer, INITIAL_CORE);
- const { detail, displayState, clip, chat } = core;
+ const { detail, displayState, clip, chat, digest } = core;
const [playing, setPlaying] = useState(false);
const [currentTime, setCurrentTime] = useState(0);
// Set when the Kick HLS manifest fails to load (usually an expired VOD), so
@@ -304,6 +381,13 @@ export function PlayerProvider({
// chat-fetch effect needs to decide "re-toggle on already-missing slug".
const chatStatusRef = useRef<ChatStatus>("idle");
chatStatusRef.current = chat.status;
+ // The digest slice's equivalents, for exactly the same reason: including
+ // `digest.slug`/`digest.status` in the fetch effect's deps makes React run
+ // cleanup before the fetch resolves, and the closure's cancelled-flag then
+ // drops the result.
+ const digestInFlightForRef = useRef<string | null>(null);
+ const digestStatusRef = useRef<DigestStatus>("idle");
+ digestStatusRef.current = digest.status;
const activeSlug = urlSlug;
const modalMode: ModalMode = urlVm;
@@ -354,12 +438,7 @@ export function PlayerProvider({
writeUrlParams({
v: slug,
t,
- vm:
- opts?.mode === "chat"
- ? "chat"
- : opts?.mode === "post"
- ? "post"
- : "transcript",
+ vm: opts?.mode ?? "transcript",
});
},
[activeSlug],
@@ -422,7 +501,9 @@ export function PlayerProvider({
const params = new URLSearchParams();
params.set("v", data.slug);
if (secs > 0) params.set("t", String(secs));
- if (modalMode === "chat") params.set("vm", "chat");
+ // Carry the current panel so a shared link reopens what the sharer was
+ // looking at. "transcript" is the default and stays absent from the URL.
+ if (modalMode !== "transcript") params.set("vm", modalMode);
const url = `${window.location.origin}${window.location.pathname}?${params.toString()}`;
try {
await navigator.clipboard.writeText(url);
@@ -535,6 +616,8 @@ export function PlayerProvider({
useEffect(() => {
dispatch({ type: "CHAT_RESET" });
chatInFlightForRef.current = null;
+ dispatch({ type: "DIGEST_RESET" });
+ digestInFlightForRef.current = null;
setKickError(false);
if (!urlSlug) return;
// Posts are not transcripts — never run the video fetch for one.
@@ -640,13 +723,86 @@ export function PlayerProvider({
};
}, [urlSlug, modalMode]);
+ // Does this video have a digest? Resolved eagerly on every slug change (one
+ // small, cached manifest request per channel) because the modal's toolbar has
+ // to decide whether to render the Digest control before anyone clicks it.
+ // Never fetches a page — see digestCache.hasDigest.
+ useEffect(() => {
+ if (!urlSlug) return;
+ if (urlVm === "post") return;
+ let cancelled = false;
+ void hasDigest(urlSlug).then((available) => {
+ if (cancelled) return;
+ dispatch({ type: "DIGEST_AVAILABLE", slug: urlSlug, available });
+ });
+ return () => {
+ cancelled = true;
+ };
+ // urlVm intentionally omitted for the same reason as the transcript fetch
+ // above: a plain mode toggle must not re-run this.
+ // eslint-disable-next-line react-hooks/exhaustive-deps
+ }, [urlSlug]);
+
+ // Lazy-fetch the digest only when the modal enters digest mode for the
+ // current video, and snap back to the transcript on missing/error so a reader
+ // is never stranded on an empty panel. Structurally identical to the chat
+ // effect above, including why the two refs exist instead of reducer state in
+ // the dep array (cleanup would run before the fetch resolves and the result
+ // would be dropped).
+ useEffect(() => {
+ if (!urlSlug) return;
+ if (modalMode !== "digest") return;
+ if (
+ digestInFlightForRef.current === urlSlug &&
+ digestStatusRef.current === "missing"
+ ) {
+ dispatch({
+ type: "DIGEST_MISSING",
+ slug: urlSlug,
+ notice: "No AI digest for this video — showing transcript.",
+ });
+ writeUrlParams({ vm: "transcript" });
+ return;
+ }
+ if (digestInFlightForRef.current === urlSlug) return;
+ digestInFlightForRef.current = urlSlug;
+ dispatch({ type: "DIGEST_LOADING", slug: urlSlug });
+ let cancelled = false;
+ fetchDigest(urlSlug)
+ .then((found) => {
+ if (cancelled) return;
+ if (found && (found.chapters.length > 0 || found.tags.length > 0)) {
+ dispatch({ type: "DIGEST_READY", slug: urlSlug, data: found });
+ } else {
+ dispatch({
+ type: "DIGEST_MISSING",
+ slug: urlSlug,
+ notice: "No AI digest for this video — showing transcript.",
+ });
+ writeUrlParams({ vm: "transcript" });
+ }
+ })
+ .catch(() => {
+ if (cancelled) return;
+ dispatch({
+ type: "DIGEST_MISSING",
+ slug: urlSlug,
+ notice: "Couldn't load the AI digest — showing transcript.",
+ });
+ writeUrlParams({ vm: "transcript" });
+ });
+ return () => {
+ cancelled = true;
+ };
+ }, [urlSlug, modalMode]);
+
// Auto-clear the snap-back notice after a brief window. Replaces the
// hand-managed setTimeout/clearTimeout ref dance — the effect's cleanup is
// the cancellation.
useEffect(() => {
if (!chat.notice) return;
const id = window.setTimeout(() => {
- dispatch({ type: "CHAT_NOTICE_CLEAR" });
+ dispatch({ type: "MODAL_NOTICE_CLEAR" });
}, 4000);
return () => {
window.clearTimeout(id);
@@ -710,7 +866,10 @@ export function PlayerProvider({
modalMode,
chatCues: chat.slug === activeSlug ? chat.cues : null,
chatStatus: chat.slug === activeSlug ? chat.status : "idle",
- chatNotice: chat.notice,
+ modalNotice: chat.notice,
+ digest: digest.slug === activeSlug ? digest.data : null,
+ digestStatus: digest.slug === activeSlug ? digest.status : "idle",
+ digestAvailable: digest.slug === activeSlug && digest.available,
clipStart,
clipEnd,
openTranscript,
@@ -735,6 +894,7 @@ export function PlayerProvider({
modalOpen,
modalMode,
chat,
+ digest,
clipStart,
clipEnd,
openTranscript,
diff --git a/common/components/SearchResults.tsx b/common/components/SearchResults.tsx
@@ -14,9 +14,20 @@ import { Checkbox } from "./ui/checkbox";
import type { LayerHit } from "./searchPipeline";
import {
AgeRestrictedBadge,
+ DuplicateBadge,
LivestreamBadge,
VodExpiredBadge,
} from "./badges";
+import {
+ duplicateJumpSeconds,
+ duplicateSiblings,
+ useDuplicates,
+ type DuplicateLookup,
+} from "./duplicatesCache";
+import type { DuplicateVideoRef } from "../lib/duplicates";
+import { splitId } from "./originId";
+import { usePlayer } from "./PlayerProvider";
+import { formatTimestamp } from "../lib/vtt";
import { vodExpiry } from "../lib/vodExpiry";
import { Button } from "./ui/button";
import { LayerSwatch } from "./LayerSwatch";
@@ -298,6 +309,20 @@ function VirtualResultList({
}) {
const flatRows = useMemo(() => buildResultRows(resultGroups), [resultGroups]);
+ // "This exists elsewhere in the archive." One root-level JSON, fetched once;
+ // absent on most sites, which is why every use of it is additive.
+ const duplicates = useDuplicates();
+ const { openTranscript } = usePlayer();
+ // Opening a sibling is not opening a hit, so it goes straight to the player
+ // rather than through openWithMode: there is no LayerHit to carry a scope, and
+ // the target is always a video's transcript view.
+ const openDuplicate = useCallback(
+ (slug: string, seconds: number) => {
+ openTranscript(slug, seconds, { mode: "transcript" });
+ },
+ [openTranscript],
+ );
+
const listRef = useRef<HTMLDivElement | null>(null);
const [scrollMargin, setScrollMargin] = useState(0);
@@ -361,6 +386,8 @@ function VirtualResultList({
selected={selectedSlugs.has(row.group.slug)}
onToggleSelect={onToggleSelect}
onAsk={onAsk}
+ duplicates={duplicates}
+ openDuplicate={openDuplicate}
/>
)}
</VirtualRow>
@@ -380,6 +407,8 @@ const ResultCard = memo(function ResultCard({
selected,
onToggleSelect,
onAsk,
+ duplicates,
+ openDuplicate,
}: {
group: ResultGroup;
leavesById: ReadonlyMap<string, LeafInfo>;
@@ -390,7 +419,29 @@ const ResultCard = memo(function ResultCard({
selected: boolean;
onToggleSelect: (slug: string) => void;
onAsk: (slug: string) => void;
+ duplicates: DuplicateLookup;
+ openDuplicate: (slug: string, seconds: number) => void;
}) {
+ // Other copies of this video elsewhere in the archive.
+ //
+ // FEDERATION GUARD: in hub mode a group's slug may be an origin-qualified id,
+ // and /duplicates.json is same-origin only — a cross-origin row's bare slug
+ // could collide with a local one and point the viewer at the wrong video. A
+ // post has no duplicate story at all. Both are simply excluded.
+ const siblings = useMemo(() => {
+ if (group.post) return [];
+ if (splitId(group.slug).origin !== "") return [];
+ return duplicateSiblings(duplicates, group.slug);
+ }, [duplicates, group.slug, group.post]);
+
+ // Where the viewer currently is in THIS copy: the moment the card is showing
+ // them (the player's playhead when this card is what's open, otherwise the top
+ // hit). Only ever honoured for a sibling the detector measured as aligned.
+ const hereSeconds =
+ activeVideo === group.slug && activeTime !== null
+ ? activeTime
+ : group.hits[0]?.start;
+
// Bucket hits per contributing leaf so the user sees one section per
// layer rather than an interleaved mishmash.
const buckets = useMemo(() => {
@@ -451,6 +502,9 @@ const ResultCard = memo(function ResultCard({
</span>
{group.isLivestream && <LivestreamBadge />}
{group.ageRestricted && <AgeRestrictedBadge />}
+ {siblings.length > 0 && (
+ <DuplicateBadge count={siblings.length} />
+ )}
{(() => {
const e = vodExpiry(group.platform, group.uploadDate);
return e?.likelyExpired ? (
@@ -477,6 +531,13 @@ const ResultCard = memo(function ResultCard({
<MessageSquareIcon className="size-3.5" /> Ask
</button>
</div>
+ {siblings.length > 0 && (
+ <DuplicateStrip
+ siblings={siblings}
+ hereSeconds={hereSeconds}
+ openDuplicate={openDuplicate}
+ />
+ )}
<ul className="flex flex-col divide-y divide-border">
{Array.from(buckets.entries()).map(([leafId, hits]) => {
if (hits.length === 0) return null;
@@ -531,6 +592,59 @@ const ResultCard = memo(function ResultCard({
});
ResultCard.displayName = "ResultCard";
+// The jump controls for a card's other copies. Lives outside the card's
+// open-the-video <Button> so these can be real buttons, and handles more than
+// one sibling, which a single control in the header could not.
+//
+// Each button STATES where it will land, because the two cases are genuinely
+// different promises: an aligned sibling was measured to carry the same words at
+// the same times, so the same timestamp is meaningful there; an unmeasured or
+// misaligned one is opened from the start, because a mirror with a longer intro
+// would otherwise drop the viewer mid-sentence in the wrong place and look like
+// it worked. Never claim the moment we did not measure.
+function DuplicateStrip({
+ siblings,
+ hereSeconds,
+ openDuplicate,
+}: {
+ siblings: ReadonlyArray<DuplicateVideoRef>;
+ hereSeconds: number | undefined;
+ openDuplicate: (slug: string, seconds: number) => void;
+}) {
+ return (
+ <div
+ data-duplicate-strip=""
+ className="flex flex-wrap items-center gap-2 border-b border-border bg-muted/40 px-3 py-1.5 text-xs"
+ >
+ <span className="text-muted-foreground">Also in this archive:</span>
+ {siblings.map((ref) => {
+ const to = duplicateJumpSeconds(ref, hereSeconds);
+ const label = ref.channel || ref.channelSlug;
+ return (
+ <button
+ key={ref.slug}
+ type="button"
+ data-duplicate-slug={ref.slug}
+ data-duplicate-aligned={ref.aligned === true ? "true" : "false"}
+ onClick={() => openDuplicate(ref.slug, to)}
+ title={
+ to > 0
+ ? `Open the ${ref.platform} copy on ${label} at the same moment — its timings were measured to line up with this one`
+ : ref.aligned === true
+ ? `Open the ${ref.platform} copy on ${label}`
+ : `Open the ${ref.platform} copy on ${label} from the start — its timings were not verified to line up with this one`
+ }
+ className="rounded border border-border bg-background px-1.5 py-0.5 transition-colors hover:bg-accent hover:text-accent-foreground"
+ >
+ {label} · {ref.platform}
+ {to > 0 ? ` @ ${formatTimestamp(to)}` : ""}
+ </button>
+ );
+ })}
+ </div>
+ );
+}
+
// One hit row inside a card. Memoized so that (a) scroll-frame re-renders of
// the parent list don't re-run highlight() for every hit, and (b) when the
// modal target changes only the previously-active and newly-active rows
diff --git a/common/components/StreamActionLog.tsx b/common/components/StreamActionLog.tsx
@@ -45,6 +45,24 @@ export function StreamActionLog({
}: Props) {
const accessibleName = label ?? buttonLabel;
const router = useRouter();
+ // A click before hydration is SILENTLY DISCARDED — the server-rendered button
+ // has no handler yet, so nothing fires: no request, no job, no log, no error.
+ // On a heavy page that is a long window (the 39-video channel page takes
+ // seconds to hydrate even in production, minutes under a dev server), and it
+ // is the whole of STATE.md's "the pilot's un-created job": clicking "Digest
+ // channel" produced no job record and an empty log, with no dedupe guard in
+ // runManagedFunction to explain it — because the action was never invoked at
+ // all. Reproduced against the real corpus: the first click fires no POST, a
+ // second click a few seconds later starts the job normally.
+ //
+ // Disabling until mounted makes the swallowed click impossible rather than
+ // merely unlikely, and it turns the failure from "nothing happened" into a
+ // button that is visibly not ready yet. It also fixes the same latent race for
+ // every other action built on this component (sync, download, transcribe),
+ // and makes e2e clicks WAIT for enablement instead of losing the click — see
+ // the pre-hydration lost-click that made the charts specs flaky.
+ const [hydrated, setHydrated] = useState(false);
+ useEffect(() => setHydrated(true), []);
const [running, setRunning] = useState(false);
const [log, setLog] = useState("");
const [error, setError] = useState<string | null>(null);
@@ -163,7 +181,7 @@ export function StreamActionLog({
type="button"
size="sm"
onClick={handleClick}
- disabled={running || disabled}
+ disabled={!hydrated || running || disabled}
>
{running ? runningLabel : buttonLabel}
</Button>
diff --git a/common/components/TranscriptModal.tsx b/common/components/TranscriptModal.tsx
@@ -7,7 +7,9 @@ import {
toHMS,
usePlayer,
usePlayerTime,
+ type DigestStatus,
} from "./PlayerProvider";
+import type { VideoDigest } from "../lib/digests";
import { AgeRestrictedBadge, LivestreamBadge } from "./badges";
import { VirtualRow } from "./VirtualRow";
import { formatTimestamp } from "../lib/vtt";
@@ -31,7 +33,10 @@ export default function TranscriptModal() {
setModalMode,
chatCues,
chatStatus,
- chatNotice,
+ modalNotice,
+ digest,
+ digestStatus,
+ digestAvailable,
clipStart,
clipEnd,
setDisplayMode,
@@ -63,6 +68,7 @@ export default function TranscriptModal() {
const scrollKindRef = useRef<"smooth" | "auto">("smooth");
const isChat = modalMode === "chat";
+ const isDigest = modalMode === "digest";
const canDownloadFile = isChat
? (chatCues?.length ?? 0) > 0
: (data?.cues?.length ?? 0) > 0;
@@ -93,7 +99,10 @@ export default function TranscriptModal() {
// doesn't redo `indexOf`/`slice` on every progress tick. Memo key is the
// identity of the underlying cues array.
const displayCues: DisplayCue[] = useMemo(() => {
- const raw = isChat ? (chatCues ?? []) : (data?.cues ?? []);
+ // The digest panel renders chapters, not cues, and is NOT virtualized — so
+ // the cue list is empty in that mode and the virtualizer below measures
+ // nothing rather than a list that isn't on screen.
+ const raw = isDigest ? [] : isChat ? (chatCues ?? []) : (data?.cues ?? []);
return raw.map((c) => {
const sepIdx = isChat ? c.text.indexOf(": ") : -1;
const author = sepIdx > 0 ? c.text.slice(0, sepIdx) : null;
@@ -106,7 +115,14 @@ export default function TranscriptModal() {
body,
};
});
- }, [isChat, chatCues, data?.cues]);
+ }, [isDigest, isChat, chatCues, data?.cues]);
+
+ // Chapters are tens of rows, so they render as a plain list. Their active-row
+ // highlight uses the same binary search as the cue list — and reads the clock
+ // from usePlayerTime(), which is why that context is split out: a 4 Hz tick
+ // must not re-render every usePlayer() consumer.
+ const chapters = digest?.chapters ?? [];
+ const activeChapter = findActiveIndex(chapters, currentTime);
const cueStatus: "loading" | "ready" =
isChat
@@ -164,11 +180,30 @@ export default function TranscriptModal() {
[seekTo],
);
- const onToggleMode = useCallback(() => {
+ // Three-way selector, expressed as two toggles that each fall back to the
+ // transcript. A cycling single button would make "get me back to the
+ // transcript" take a variable number of clicks.
+ const onToggleChat = useCallback(() => {
scrollKindRef.current = "smooth";
setModalMode(isChat ? "transcript" : "chat");
}, [isChat, setModalMode]);
+ const onToggleDigest = useCallback(() => {
+ scrollKindRef.current = "smooth";
+ setModalMode(isDigest ? "transcript" : "digest");
+ }, [isDigest, setModalMode]);
+
+ // Seek from a chapter. ALWAYS on `start` — the parser already snapped it to a
+ // real cue boundary — and never by re-parsing `clock`, which is the raw model
+ // output kept for auditing and can be wrong.
+ const onSeekChapter = useCallback(
+ (start: number) => {
+ scrollKindRef.current = "smooth";
+ seekTo(start);
+ },
+ [seekTo],
+ );
+
if (!modalOpen || !activeSlug) return null;
const canDownload = clipStart !== null && clipEnd !== null && clipEnd > clipStart;
@@ -238,9 +273,20 @@ export default function TranscriptModal() {
disabled={!canDownload}
/>
<div className="flex-1" />
+ {/* Only offered where a digest exists. ~0.1% of the corpus is
+ digested, so an always-present control would be a dead end on
+ almost every video. */}
+ {digestAvailable && (
+ <ControlButton
+ title={isDigest ? "Show transcript" : "Show AI chapters"}
+ onClick={onToggleDigest}
+ char={isDigest ? "📜" : "✦"}
+ highlight={isDigest}
+ />
+ )}
<ControlButton
title={isChat ? "Show transcript" : "Show live chat"}
- onClick={onToggleMode}
+ onClick={onToggleChat}
char={isChat ? "📜" : "💬"}
highlight={isChat}
/>
@@ -344,18 +390,27 @@ export default function TranscriptModal() {
)}
</div>
- {chatNotice && (
+ {modalNotice && (
<p
role="status"
className="text-xs text-amber-200 bg-amber-500/15 ring-1 ring-amber-400/40 rounded px-3 py-1.5"
>
- {chatNotice}
+ {modalNotice}
</p>
)}
<div
ref={scrollRef}
className="flex-1 overflow-y-auto rounded-lg border border-zinc-700 bg-zinc-900"
>
+ {isDigest ? (
+ <DigestPanel
+ digest={digest}
+ status={digestStatus}
+ activeChapter={activeChapter}
+ onSeek={onSeekChapter}
+ />
+ ) : (
+ <>
{cueStatus === "loading" && (
<p className="text-sm text-zinc-400 p-4">
{isChat ? "Loading live chat…" : "Loading transcript…"}
@@ -388,6 +443,8 @@ export default function TranscriptModal() {
})}
</ol>
)}
+ </>
+ )}
</div>
</div>
</div>
@@ -395,6 +452,140 @@ export default function TranscriptModal() {
);
}
+// The digest panel: chapters as a plain (non-virtualized) list, the topic tags,
+// a provenance line, and — when the digest was borrowed from another video — a
+// prominent notice saying so.
+//
+// Styling deliberately matches this modal's local hard-coded zinc/white palette
+// rather than the semantic tokens used elsewhere in the app. This component and
+// PlayerProvider are the one part of the codebase still on literal colors
+// (bg-zinc-900/80, ring-white/10, active row bg-blue-950/50); migrating them is
+// its own change, and half-migrating one panel inside them would just look
+// broken.
+function DigestPanel({
+ digest,
+ status,
+ activeChapter,
+ onSeek,
+}: {
+ digest: VideoDigest | null;
+ status: DigestStatus;
+ activeChapter: number;
+ onSeek: (start: number) => void;
+}) {
+ if (status === "loading" || status === "idle") {
+ return <p className="text-sm text-zinc-400 p-4">Loading AI digest…</p>;
+ }
+ if (!digest) {
+ return (
+ <p className="text-sm text-zinc-400 p-4">
+ No AI digest for this video.
+ </p>
+ );
+ }
+
+ const prov = digest.provenance.chapters ?? digest.provenance.tags;
+ const borrowed = digest.derivedFrom;
+
+ return (
+ <div className="text-zinc-100">
+ {borrowed && (
+ <div className="m-3 rounded px-3 py-2 text-xs bg-amber-500/15 ring-1 ring-amber-400/40 text-amber-100">
+ <p className="font-medium">Borrowed from a duplicate upload</p>
+ <p className="mt-1 text-amber-200/90">
+ These chapters were generated for{" "}
+ <span className="font-mono">{borrowed.slug}</span>, a near-identical
+ copy of this video, and copied here. Timings were measured to differ
+ by {formatOffset(borrowed.offsetSeconds)}, so they should line up —
+ but the titles describe that upload, not this one.
+ </p>
+ </div>
+ )}
+
+ {digest.chapters.length > 0 ? (
+ <ol className="divide-y divide-zinc-800">
+ {digest.chapters.map((c, i) => (
+ <li key={c.id} className={i === activeChapter ? "bg-blue-950/50" : ""}>
+ <button
+ type="button"
+ onClick={() => onSeek(c.start)}
+ // Announces the chapter the playhead is currently inside — for
+ // a screen reader, and it is also the only DOM-observable proof
+ // that a chapter click actually moved the player.
+ aria-current={i === activeChapter ? "true" : undefined}
+ className="w-full text-left flex gap-3 px-3 py-2 hover:bg-zinc-800"
+ >
+ <span className="text-xs font-mono text-zinc-400 shrink-0 w-16 pt-0.5">
+ {formatTimestamp(c.start)}
+ </span>
+ <span className="text-sm min-w-0 flex-1">{c.title}</span>
+ {c.decidedBy === "human" && (
+ <span
+ title="Written or corrected by a person"
+ className="text-[10px] uppercase tracking-wide text-emerald-300/80 shrink-0 pt-1"
+ >
+ edited
+ </span>
+ )}
+ </button>
+ </li>
+ ))}
+ </ol>
+ ) : (
+ <p className="text-sm text-zinc-400 px-3 py-4">
+ No chapters in this digest.
+ </p>
+ )}
+
+ {digest.tags.length > 0 && (
+ <div className="border-t border-zinc-800 px-3 py-3">
+ <p className="text-[11px] uppercase tracking-wide text-zinc-500 mb-1.5">
+ Topics
+ </p>
+ <ul className="flex flex-wrap gap-1.5">
+ {digest.tags.map((t) => (
+ <li
+ key={t.id}
+ className="rounded bg-white/10 px-2 py-0.5 text-xs text-zinc-200"
+ >
+ {t.tag}
+ </li>
+ ))}
+ </ul>
+ </div>
+ )}
+
+ {/* Always say where this came from. A derived layer that doesn't
+ announce itself as machine-generated is the one that misleads. */}
+ <p className="border-t border-zinc-800 px-3 py-2 text-[11px] leading-relaxed text-zinc-500">
+ {prov ? (
+ <>
+ Chapters and topics generated by{" "}
+ <span className="font-mono text-zinc-400">{prov.model}</span>
+ {prov.generatedAt ? ` on ${prov.generatedAt.slice(0, 10)}` : ""}
+ . AI-generated and not reviewed unless marked{" "}
+ <span className="text-emerald-300/80">edited</span>.
+ </>
+ ) : (
+ // No machine provenance at all: this digest is entirely hand-written
+ // (an overrides file with no generated counterpart). Claiming
+ // "AI-generated" here would be the same dishonesty in reverse.
+ "Written by hand."
+ )}
+ </p>
+ </div>
+ );
+}
+
+// The measured cue-timing offset between a borrowed digest's source video and
+// this one. Sharing only happens at near-zero offset, so this is normally well
+// under a second — say so precisely rather than rounding it away to "0s".
+function formatOffset(seconds: number): string {
+ const s = Math.abs(seconds);
+ if (s < 1) return `under a second`;
+ return `${s.toFixed(1)}s`;
+}
+
const CueRow = memo(function CueRow({
cue,
isActive,
diff --git a/common/components/badges.tsx b/common/components/badges.tsx
@@ -23,6 +23,26 @@ export function AgeRestrictedBadge() {
);
}
+// Flags a search result that also exists elsewhere in the archive — a mirror on
+// another platform, or a re-upload on another channel. Non-interactive on
+// purpose: it sits inside the card's open-the-video button, and the jump
+// controls live in their own strip below the header where they can be real
+// buttons (and where there is room for more than one sibling).
+export function DuplicateBadge({ count }: { count: number }) {
+ return (
+ <span
+ title={
+ count === 1
+ ? "This video also exists elsewhere in the archive"
+ : `This video also exists in ${count} other places in the archive`
+ }
+ className={`${badgeBase} bg-sky-100 text-sky-800 dark:bg-sky-900/40 dark:text-sky-200`}
+ >
+ {count === 1 ? "Dupe" : `${count} dupes`}
+ </span>
+ );
+}
+
// Flags a stream VOD that is likely past its platform's retention window (Kick
// ~30d, Twitch 7–60d), so playback probably fails. `tooltip` explains the
// platform-specific retention on hover (see lib/vodExpiry).
diff --git a/common/components/digestCache.ts b/common/components/digestCache.ts
@@ -0,0 +1,167 @@
+"use client";
+
+import type { ChannelDigestsManifest, VideoDigest } from "../lib/digests";
+import { digestPageFileName, manifestHasDigest } from "../lib/digests";
+import { idbGet, idbPut } from "./digestStore";
+import { makeId, splitId, idBaseUrl } from "./originId";
+
+// Fetch layer for the derived corpus, mirroring transcriptCache.ts: a memory
+// map, in-flight dedupe, manifest -> page resolution and opportunistic warming
+// of every record in a fetched page.
+//
+// The one structural difference is that THE LAYER IS SPARSE. A transcript
+// exists for every indexed video, so transcriptCache treats a missing one as an
+// error. Digests cover 102 of ~76,000 videos, so "absent" is the normal answer
+// and must be a value, not a throw:
+//
+// hasDigest() -> boolean, answered from the channel manifest alone
+// fetchDigest() -> VideoDigest | null, null meaning "not digested"
+//
+// A channel with no digests has no manifest at all, so its 404 is also a normal
+// answer and is cached as `null` — otherwise every video opened in an
+// undigested channel would re-request the same missing file.
+
+const resolved = new Map<string, VideoDigest | null>();
+const inFlight = new Map<string, Promise<VideoDigest | null>>();
+const channelManifests = new Map<
+ string,
+ Promise<ChannelDigestsManifest | null>
+>();
+const pagePromises = new Map<string, Promise<VideoDigest[]>>();
+
+// `id` is an OriginId: a bare "channelSlug/videoId" for same-origin content, or
+// "origin\tchannelSlug/videoId" for a federated cross-origin video.
+export function fetchDigest(id: string): Promise<VideoDigest | null> {
+ const hit = resolved.get(id);
+ if (hit !== undefined) return Promise.resolve(hit);
+ const flying = inFlight.get(id);
+ if (flying) return flying;
+
+ const p = load(id).then((digest) => {
+ resolved.set(id, digest);
+ inFlight.delete(id);
+ return digest;
+ });
+ p.catch(() => {
+ inFlight.delete(id);
+ });
+ inFlight.set(id, p);
+ return p;
+}
+
+// Does this video have a digest? Answered from the channel manifest, which is
+// one small cached request per channel — never a per-video fetch. This is what
+// gates the viewer's Digest control: with coverage at 0.1% of the corpus, an
+// always-present button would be a dead end almost everywhere.
+export async function hasDigest(id: string): Promise<boolean> {
+ const parts = parseId(id);
+ if (!parts) return false;
+ try {
+ const manifest = await fetchChannelManifest(
+ parts.channelSlug,
+ parts.origin,
+ );
+ return manifestHasDigest(manifest, parts.videoId);
+ } catch {
+ return false;
+ }
+}
+
+function parseId(
+ id: string,
+): { origin: string; slug: string; channelSlug: string; videoId: string } | null {
+ const { origin, slug } = splitId(id);
+ const slashIdx = slug.indexOf("/");
+ if (slashIdx < 0) return null;
+ return {
+ origin,
+ slug,
+ channelSlug: slug.slice(0, slashIdx),
+ videoId: slug.slice(slashIdx + 1),
+ };
+}
+
+// Resolves to null (not a rejection) when the channel ships no digests, so the
+// absence is cached like any other answer.
+function fetchChannelManifest(
+ channelSlug: string,
+ origin: string,
+): Promise<ChannelDigestsManifest | null> {
+ const key = makeId(origin, channelSlug);
+ let p = channelManifests.get(key);
+ if (!p) {
+ p = fetch(`${idBaseUrl(origin)}/digests/${channelSlug}/manifest.json`)
+ .then((r) => {
+ if (r.status === 404) return null;
+ if (!r.ok) {
+ throw new Error(`Failed to fetch digests manifest for ${channelSlug}`);
+ }
+ return r.json() as Promise<ChannelDigestsManifest>;
+ })
+ .catch((err) => {
+ // A transport failure is not proof of absence, so it must not be
+ // cached as one — drop the entry and let the next caller retry.
+ channelManifests.delete(key);
+ throw err;
+ });
+ channelManifests.set(key, p);
+ }
+ return p;
+}
+
+function fetchPage(
+ channelSlug: string,
+ pageIndex: number,
+ origin: string,
+): Promise<VideoDigest[]> {
+ const key = `${makeId(origin, channelSlug)}:${pageIndex}`;
+ let p = pagePromises.get(key);
+ if (!p) {
+ p = fetch(
+ `${idBaseUrl(origin)}/digests/${channelSlug}/${digestPageFileName(pageIndex)}`,
+ ).then((r) => {
+ if (!r.ok) {
+ throw new Error(`Failed to fetch digest page ${channelSlug}/${pageIndex}`);
+ }
+ return r.json() as Promise<VideoDigest[]>;
+ });
+ p.catch(() => pagePromises.delete(key));
+ pagePromises.set(key, p);
+ }
+ return p;
+}
+
+async function load(id: string): Promise<VideoDigest | null> {
+ const parts = parseId(id);
+ if (!parts) return null;
+ const { origin, slug, channelSlug, videoId } = parts;
+
+ const manifest = await fetchChannelManifest(channelSlug, origin);
+ if (!manifest) return null;
+ const pageIndex = manifest.slugToPage[videoId];
+ // Not in slugToPage = not digested. The normal case, and not an error.
+ if (pageIndex === undefined) return null;
+
+ // The page's content hash is this record's cache version. Absent (an older
+ // manifest) means every cached entry misses, which is the safe direction.
+ const version = manifest.pageHashes?.[pageIndex] ?? "";
+ const stored = await idbGet(id, version);
+ if (stored) return stored;
+
+ const page = await fetchPage(channelSlug, pageIndex, origin);
+ let found: VideoDigest | null = null;
+ for (const entry of page) {
+ // Re-key warmed entries by OriginId so a cross-origin channelSlug/videoId
+ // can't shadow a same-origin one with the same slug.
+ const entryId = makeId(origin, entry.slug);
+ if (entry.slug === slug) found = entry;
+ resolved.set(entryId, entry);
+ idbPut(entryId, entry, version);
+ }
+ // In slugToPage but missing from the page means the manifest and the page
+ // tree disagree — a real build fault, not a coverage gap, so it throws.
+ if (!found) {
+ throw new Error(`Digest ${slug} missing from page ${pageIndex}`);
+ }
+ return found;
+}
diff --git a/common/components/digestStore.ts b/common/components/digestStore.ts
@@ -0,0 +1,204 @@
+"use client";
+
+import type { VideoDigest } from "../lib/digests";
+
+// A SEPARATE DATABASE, not a second store inside transcriptStore.ts's.
+//
+// `DB_VERSION` is a property of the DATABASE, not of an object store, and
+// transcriptStore.ts's `onupgradeneeded` does deleteObjectStore +
+// createObjectStore on every upgrade. Adding a `digests` store there would
+// force a version bump and WIPE every existing client's transcript cache — a
+// large, silent, entirely avoidable cost for shipping a new layer. Following
+// searchLayerCache.ts instead: own DB name, own version, independent evolution.
+const DB_NAME = "yt-dlp-transcript-browser:digests";
+const DB_VERSION = 1;
+const STORE = "digests";
+
+// PER-ENTRY VERSIONING, which transcripts deliberately do without.
+//
+// transcriptStore.ts gets away with an ad-hoc shape sniff (does `platform` look
+// valid?) because a transcript is near-immutable: once written it does not
+// change, so a structurally-valid cached copy is also a CURRENT one. Digests
+// break that assumption — they are regenerated in place whenever the prompt,
+// model or a human correction changes, and the regenerated record has exactly
+// the same shape. A shape sniff cannot see the difference, so a returning
+// reader would keep being served the superseded digest forever, which is the
+// one failure this whole layer exists to avoid (a corrected chapter that never
+// reaches anyone). We therefore store the record's own `generatedAt` alongside
+// it and treat any mismatch with the manifest-fresh copy as a miss.
+//
+// This is the same class of bug already recorded for the transcript cache when
+// TranscriptDetail's shape changed; here it is designed out rather than patched
+// with a version bump after the fact.
+//
+// The version we compare is the CONTENT HASH of the page the record came from
+// (ChannelDigestsManifest.pageHashes), not the record's own `generatedAt`. Both
+// detect a regeneration, but only the page hash is knowable from the manifest
+// alone — i.e. before paying for the fetch the cache exists to avoid — and it
+// additionally catches changes that leave `generatedAt` untouched, such as a
+// human retitling a chapter through the overrides file.
+type StoredEntry = {
+ digest: VideoDigest;
+ // The manifest pageHash this record was fetched under.
+ version: string;
+};
+
+type Mode = "pending" | "ok" | "unavailable";
+
+let mode: Mode = "pending";
+let dbPromise: Promise<IDBDatabase | null> | null = null;
+
+function openDb(): Promise<IDBDatabase | null> {
+ if (typeof indexedDB === "undefined") {
+ mode = "unavailable";
+ return Promise.resolve(null);
+ }
+ if (dbPromise) return dbPromise;
+ dbPromise = new Promise<IDBDatabase | null>((resolve) => {
+ let req: IDBOpenDBRequest;
+ try {
+ req = indexedDB.open(DB_NAME, DB_VERSION);
+ } catch {
+ downgrade("open threw");
+ resolve(null);
+ return;
+ }
+ req.onupgradeneeded = () => {
+ const db = req.result;
+ if (db.objectStoreNames.contains(STORE)) {
+ db.deleteObjectStore(STORE);
+ }
+ // Out-of-line keys: the caller supplies an OriginId (see idbPut).
+ db.createObjectStore(STORE);
+ };
+ req.onsuccess = () => {
+ mode = "ok";
+ resolve(req.result);
+ };
+ req.onerror = () => {
+ downgrade("open failed");
+ resolve(null);
+ };
+ req.onblocked = () => {
+ downgrade("open blocked");
+ resolve(null);
+ };
+ });
+ return dbPromise;
+}
+
+function downgrade(reason: string): void {
+ if (mode === "unavailable") return;
+ mode = "unavailable";
+ console.warn(
+ `[digestStore] IndexedDB unavailable (${reason}); falling back to network-only.`,
+ );
+}
+
+// Read a cached digest, but only if its recorded version matches `expected`
+// (the page hash the manifest says is current). A mismatch — or an entry
+// written before this field existed — resolves null, so the caller re-fetches.
+// An empty `expected` also misses: a record we cannot version is one we must
+// not serve from cache.
+export async function idbGet(
+ id: string,
+ expected: string,
+): Promise<VideoDigest | null> {
+ if (!expected) return null;
+ const db = await openDb();
+ if (!db) return null;
+ return new Promise<VideoDigest | null>((resolve) => {
+ let req: IDBRequest<StoredEntry | undefined>;
+ try {
+ const tx = db.transaction(STORE, "readonly");
+ req = tx.objectStore(STORE).get(id) as IDBRequest<StoredEntry | undefined>;
+ } catch {
+ downgrade("read tx threw");
+ resolve(null);
+ return;
+ }
+ req.onsuccess = () => {
+ const entry = req.result;
+ if (!entry || !entry.digest || entry.version !== expected) {
+ resolve(null);
+ return;
+ }
+ resolve(entry.digest);
+ };
+ req.onerror = () => resolve(null);
+ });
+}
+
+type PendingEntry = { id: string; entry: StoredEntry };
+let pending: PendingEntry[] = [];
+let flushScheduled = false;
+let flushInFlight: Promise<void> | null = null;
+
+// Queue a digest for persistence under the given OriginId, stamped with the
+// page hash it was fetched under. Batched via queueMicrotask exactly as
+// transcriptStore.ts does, so warming a whole page's worth of digests costs one
+// transaction. An unversioned write is dropped rather than stored — an entry
+// with no version could never be validated and would only ever waste quota.
+export function idbPut(id: string, digest: VideoDigest, version: string): void {
+ if (mode === "unavailable") return;
+ if (!version) return;
+ pending.push({ id, entry: { digest, version } });
+ if (flushScheduled) return;
+ flushScheduled = true;
+ queueMicrotask(() => {
+ flushScheduled = false;
+ void flush();
+ });
+}
+
+async function flush(): Promise<void> {
+ if (flushInFlight) {
+ await flushInFlight;
+ }
+ if (pending.length === 0) return;
+ const batch: PendingEntry[] = pending;
+ pending = [];
+ flushInFlight = writeBatch(batch).finally(() => {
+ flushInFlight = null;
+ if (pending.length > 0 && !flushScheduled) {
+ flushScheduled = true;
+ queueMicrotask(() => {
+ flushScheduled = false;
+ void flush();
+ });
+ }
+ });
+ await flushInFlight;
+}
+
+async function writeBatch(batch: PendingEntry[]): Promise<void> {
+ const db = await openDb();
+ if (!db) return;
+ return new Promise<void>((resolve) => {
+ let tx: IDBTransaction;
+ try {
+ tx = db.transaction(STORE, "readwrite");
+ } catch {
+ downgrade("write tx threw");
+ resolve();
+ return;
+ }
+ const store = tx.objectStore(STORE);
+ for (const item of batch) {
+ try {
+ store.put(item.entry, item.id);
+ } catch {
+ // Per-entry errors (e.g. unclonable values) shouldn't fail the batch.
+ }
+ }
+ tx.oncomplete = () => resolve();
+ tx.onerror = () => {
+ downgrade("write tx error");
+ resolve();
+ };
+ tx.onabort = () => {
+ downgrade("write tx abort");
+ resolve();
+ };
+ });
+}
diff --git a/common/components/duplicatesCache.ts b/common/components/duplicatesCache.ts
@@ -0,0 +1,83 @@
+"use client";
+
+// Client fetch for a site's shipped /duplicates.json, indexed by member slug so
+// a search result can answer "does this video exist elsewhere?" in O(1). Mirrors
+// aliasesCache / summariesCache: one root-level JSON, fetched once, cached
+// forever (it is static per export build).
+//
+// A missing file resolves to an empty map rather than an error. That is the
+// common case, not an edge one — compose-site writes the file only when the site
+// has at least one shippable cluster, and every duplicate affordance here is
+// purely additive, so its absence must never break search.
+//
+// WHAT SHIPS HERE IS ALREADY FILTERED. compose-site drops `needsReview` clusters
+// unless a human confirmed them, so anything in this map has had its CONTENT
+// compared, not just its title and runtime. The one thing still worth checking
+// per member is `aligned` — see duplicateSiblings below.
+
+import { useQuery } from "@tanstack/react-query";
+import { idBaseUrl } from "./originId";
+import type {
+ DuplicateCluster,
+ DuplicateReport,
+ DuplicateVideoRef,
+} from "../lib/duplicates";
+
+export type DuplicateLookup = ReadonlyMap<string, DuplicateCluster>;
+
+const EMPTY: DuplicateLookup = new Map();
+
+export async function fetchDuplicates(origin = ""): Promise<DuplicateLookup> {
+ let report: DuplicateReport | null = null;
+ try {
+ const r = await fetch(`${idBaseUrl(origin)}/duplicates.json`);
+ if (!r.ok) return EMPTY;
+ report = (await r.json()) as DuplicateReport;
+ } catch {
+ return EMPTY;
+ }
+ const out = new Map<string, DuplicateCluster>();
+ for (const cluster of report?.clusters ?? []) {
+ for (const ref of cluster?.videoRefs ?? []) {
+ if (ref?.slug) out.set(ref.slug, cluster);
+ }
+ }
+ return out;
+}
+
+export function useDuplicates(origin = ""): DuplicateLookup {
+ const { data } = useQuery<DuplicateLookup>({
+ queryKey: ["duplicates", origin],
+ queryFn: () => fetchDuplicates(origin),
+ staleTime: Infinity, // static per export build
+ });
+ return data ?? EMPTY;
+}
+
+// The other members of `slug`'s cluster, or [] when it is in none.
+export function duplicateSiblings(
+ lookup: DuplicateLookup,
+ slug: string,
+): DuplicateVideoRef[] {
+ const cluster = lookup.get(slug);
+ if (!cluster) return [];
+ return cluster.videoRefs.filter((r) => r.slug !== slug);
+}
+
+// Where to land when opening a sibling, given where the viewer is in THIS copy.
+//
+// The honesty rule. Matching content does NOT imply matching timings: a mirror
+// with a longer intro or an extra ad break carries the same words at shifted
+// times, so seeking to the same timestamp lands in the wrong place while looking
+// perfectly plausible. `aligned` is the detector's measured verdict on exactly
+// that, and it is only ever true when several anchors were located within the
+// tolerance. Absent (an older report, or no timed cues to measure) reads as NOT
+// aligned, so the fallback is the start of the video — a jump the viewer can
+// always make sense of.
+export function duplicateJumpSeconds(
+ ref: DuplicateVideoRef,
+ seconds: number | undefined,
+): number {
+ if (ref.aligned !== true) return 0;
+ return typeof seconds === "number" && seconds > 0 ? seconds : 0;
+}
diff --git a/common/components/urlState.ts b/common/components/urlState.ts
@@ -9,8 +9,14 @@ export type SearchMode = "transcripts" | "subs" | "posts";
// Per-video modal content mode. Independent from the search page's `mode` so
// the modal can be toggled without disturbing search state. Absence on the
-// URL means "transcript" — only `"chat"` is persisted.
-export type ModalMode = "transcript" | "chat" | "post";
+// URL means "transcript" — every other value is persisted verbatim.
+//
+// "digest" and NOT "summary", deliberately: `DisplaySummary` is already a video
+// LISTING CARD (common/lib/transcripts.ts) and `/summaries/` is already the
+// browse-index page tree the search index serves. A third meaning of "summary"
+// is exactly the naming hazard PLAN.md warns about. "digest" matches the
+// artifact, the generation stage, the settings section and the page tree.
+export type ModalMode = "transcript" | "chat" | "post" | "digest";
export type UrlParams = {
q: string;
@@ -62,7 +68,13 @@ function parse(search: string): UrlParams {
modeRaw === "subs" ? "subs" : modeRaw === "posts" ? "posts" : "transcripts";
const vmRaw = p.get("vm");
const vm: ModalMode =
- vmRaw === "chat" ? "chat" : vmRaw === "post" ? "post" : "transcript";
+ vmRaw === "chat"
+ ? "chat"
+ : vmRaw === "post"
+ ? "post"
+ : vmRaw === "digest"
+ ? "digest"
+ : "transcript";
return {
q: p.get("q") ?? "",
re: p.get("re") === "1",
@@ -126,6 +138,9 @@ export function writeUrlParams(patch: Patch) {
if (patch.vm !== undefined) {
if (patch.vm === "chat") params.set("vm", "chat");
else if (patch.vm === "post") params.set("vm", "post");
+ else if (patch.vm === "digest") params.set("vm", "digest");
+ // "transcript" is the fallthrough and DELETES the param rather than
+ // setting vm=transcript — the default must stay absent from the URL.
else params.delete("vm");
}
if (patch.ch !== undefined) {
diff --git a/common/controller/buildIndex.ts b/common/controller/buildIndex.ts
@@ -67,6 +67,7 @@ import {
import { getSettings } from "../lib/settings";
import {
listSites,
+ siteDigestsDir,
sitePostsDir,
siteSummariesDir,
siteSubsDir,
@@ -104,6 +105,18 @@ import {
readPostShard,
} from "../lib/posts-server";
import { isSocialChannel } from "../lib/channelConfig";
+import {
+ DIGESTS_MANIFEST_VERSION,
+ SITE_DIGESTS_MANIFEST_VERSION,
+ digestPageFileName,
+ newestGeneratedAt,
+ type ChannelDigestsManifest,
+ type DigestsChannelEntry,
+ type DigestsManifest,
+ type VideoDigest,
+} from "../lib/digests";
+import { DIGEST_FILENAME, DIGEST_OVERRIDES_FILENAME, effectiveDigest } from "../lib/digest";
+import { loadDigest, loadDigestOverrides } from "../lib/digest-server";
// v10: multi-site build. Shared per-channel transcript/subs pages are written
// once; per-site summaries + subs manifests are filtered selections. Bumped to
@@ -113,7 +126,11 @@ import { isSocialChannel } from "../lib/channelConfig";
// v12: the social-post corpus. A `posts` sub-DB keyed [createdAt, channelSlug,
// id] (ISO-8601 sorts correctly, unlike the [uploadDate, …] tuple videos use)
// plus a shared /posts/<slug>/ page tree.
-const SCHEMA_VERSION = 12;
+// v13: the AI-digest corpus. A `digests` sub-DB holding the COMPOSED digest
+// (effectiveDigest of the machine sidecar + the human override file) plus a
+// shared /digests/<slug>/ page tree. Bumped so existing indexes populate the
+// new sub-DB — nothing re-derives it lazily.
+const SCHEMA_VERSION = 13;
// Per-channel post stats, persisted so per-site aggregates survive a no-op
// rebuild that doesn't re-encode the post pages. Mirrors ChannelSubsStat.
@@ -126,6 +143,14 @@ type ChannelPostsStat = {
signature: string;
};
+// Per-channel digest stats, persisted so the per-site digests manifest survives
+// a no-op rebuild that doesn't re-encode the digest pages. Mirrors
+// ChannelSubsStat / ChannelPostsStat.
+type ChannelDigestStat = {
+ name: string;
+ digestCount: number;
+};
+
// Per-channel subtitle stats, collected while writing the shared subs pages and
// persisted to LMDB so per-site subs manifests can be assembled on a no-op
// rebuild without re-encoding every channel's pages.
@@ -148,6 +173,13 @@ type MtimeRecord = {
transcriptMs: number | null;
subsMs: number | null;
availabilityMs: number | null;
+ // Newest mtime across BOTH digest sidecars — ai-digest.json (machine) and
+ // ai-digest.overrides.json (human). Max-of-two, not a single stat, because a
+ // human correction only ever touches the overrides file: keying off the
+ // machine file alone would leave a corrected digest producing no index
+ // mutation, so the correction would never ship. That is the entire reason
+ // the overrides file exists. See the subsMs multi-file loop it follows.
+ digestMs: number | null;
isDeleted: boolean;
isUnlisted: boolean;
indexKey: IndexKey;
@@ -176,6 +208,7 @@ type LiveEntry = {
subTracks: SubTrack[];
subsMs: number | null;
availabilityMs: number | null;
+ digestMs: number | null;
};
async function exists(p: string): Promise<boolean> {
@@ -283,6 +316,19 @@ async function scanSource(
} catch {
availabilityMs = null;
}
+ // MAX of both digest sidecars, following the subsMs loop above rather
+ // than the single-stat availabilityMs below it. The overrides file is the
+ // one a human writes; if it did not move this number, a hand-corrected
+ // digest would produce no mutation and never reach a built site.
+ let digestMs: number | null = null;
+ for (const name of [DIGEST_FILENAME, DIGEST_OVERRIDES_FILENAME]) {
+ try {
+ const ms = (await stat(path.join(fullVideoDir, name))).mtimeMs;
+ if (digestMs === null || ms > digestMs) digestMs = ms;
+ } catch {
+ // Sidecar absent — the common case (102 of ~76,000 videos have one).
+ }
+ }
live.push({
channelSlug: ch.name,
handling: cfg.handling,
@@ -296,6 +342,7 @@ async function scanSource(
subTracks,
subsMs,
availabilityMs,
+ digestMs,
});
}
}
@@ -353,11 +400,13 @@ export async function buildIndex({
const transcriptsOutDir = paths.exportSharedTranscriptsDir;
const subsOutDir = paths.exportSharedSubsDir;
const postsOutDir = paths.exportSharedPostsDir;
+ const digestsOutDir = paths.exportSharedDigestsDir;
await mkdir(path.dirname(dbPath), { recursive: true });
await mkdir(transcriptsOutDir, { recursive: true });
await mkdir(subsOutDir, { recursive: true });
await mkdir(postsOutDir, { recursive: true });
+ await mkdir(digestsOutDir, { recursive: true });
await mkdir(paths.exportSitesIndexDir, { recursive: true });
const root = open({
@@ -414,6 +463,22 @@ export async function buildIndex({
name: "channelPostsStats",
encoding: "msgpack",
});
+ // The AI-digest corpus. Holds the COMPOSED digest (effectiveDigest of the
+ // machine sidecar + the human overrides), keyed like sums/cues/subs, so the
+ // page writer below can stream a channel's digests straight out of LMDB
+ // without re-reading 76k video dirs on a no-op rebuild.
+ const digests = root.openDB<VideoDigest, IndexKey>({
+ name: "digests",
+ encoding: "msgpack",
+ });
+ const digestPageHashes = root.openDB<PageHashRecord, PageHashKey>({
+ name: "digestPageHashes",
+ encoding: "msgpack",
+ });
+ const channelDigestStatsDb = root.openDB<ChannelDigestStat, string>({
+ name: "channelDigestStats",
+ encoding: "msgpack",
+ });
const meta = root.openDB<unknown, string>({
name: "meta",
encoding: "msgpack",
@@ -436,6 +501,9 @@ export async function buildIndex({
await posts.clearAsync();
await postPageHashes.clearAsync();
await channelPostsStatsDb.clearAsync();
+ await digests.clearAsync();
+ await digestPageHashes.clearAsync();
+ await channelDigestStatsDb.clearAsync();
await meta.put("schema", SCHEMA_VERSION);
}
@@ -461,7 +529,8 @@ export async function buildIndex({
prev.metaMs !== s.metaMs ||
prev.transcriptMs !== s.transcriptMs ||
(prev.subsMs ?? null) !== s.subsMs ||
- (prev.availabilityMs ?? null) !== s.availabilityMs
+ (prev.availabilityMs ?? null) !== s.availabilityMs ||
+ (prev.digestMs ?? null) !== s.digestMs
) {
changed.push(s);
}
@@ -580,6 +649,7 @@ export async function buildIndex({
sums.remove(prev.indexKey);
cues.remove(prev.indexKey);
subs.remove(prev.indexKey);
+ digests.remove(prev.indexKey);
byChannel.remove(indexToChannelKey(prev.indexKey));
}
@@ -619,6 +689,54 @@ export async function buildIndex({
if (parsedSubs.length > 0) subs.put(indexKey, parsedSubs);
else subs.remove(indexKey);
+ // The derived layer. Only opened when the scan saw a sidecar, so the
+ // ~76k videos without one cost zero extra reads. What is stored is
+ // effectiveDigest(machine, overrides) — human corrections applied,
+ // `enabled: false` items dropped, chapters sorted by start — so the
+ // corpus ships what a human approved rather than raw model output.
+ if (s.digestMs !== null) {
+ const [machine, overrides] = await Promise.all([
+ loadDigest(videoFullDir),
+ loadDigestOverrides(videoFullDir),
+ ]);
+ const eff = effectiveDigest(machine, overrides);
+ if (eff.chapters.length > 0 || eff.tags.length > 0) {
+ const provenance = {
+ ...(eff.sections.chapters
+ ? { chapters: eff.sections.chapters.provenance }
+ : {}),
+ ...(eff.sections.tags
+ ? { tags: eff.sections.tags.provenance }
+ : {}),
+ };
+ digests.put(indexKey, {
+ slug: `${s.channelSlug}/${summary.id}`,
+ id: summary.id,
+ chapters: eff.chapters.map((c) => ({
+ id: c.id,
+ start: c.start,
+ clock: c.clock,
+ title: c.title,
+ decidedBy: c.decidedBy,
+ })),
+ tags: eff.tags.map((t) => ({
+ id: t.id,
+ tag: t.tag,
+ decidedBy: t.decidedBy,
+ })),
+ generatedAt: newestGeneratedAt(provenance),
+ provenance,
+ // Carried straight through: a shared digest must stay
+ // identifiable as borrowed all the way to the viewer.
+ ...(eff.derivedFrom ? { derivedFrom: eff.derivedFrom } : {}),
+ });
+ } else {
+ digests.remove(indexKey);
+ }
+ } else {
+ digests.remove(indexKey);
+ }
+
let isDeleted = false;
let isUnlisted = false;
if (s.availabilityMs !== null) {
@@ -633,6 +751,7 @@ export async function buildIndex({
transcriptMs: s.transcriptMs,
subsMs: s.subsMs,
availabilityMs: s.availabilityMs,
+ digestMs: s.digestMs,
isDeleted,
isUnlisted,
indexKey,
@@ -652,6 +771,7 @@ export async function buildIndex({
sums.remove(indexKey);
cues.remove(indexKey);
subs.remove(indexKey);
+ digests.remove(indexKey);
byChannel.remove(indexToChannelKey(indexKey));
mtimes.remove(pathKey);
}
@@ -659,6 +779,7 @@ export async function buildIndex({
await sums.flushed;
await cues.flushed;
await subs.flushed;
+ await digests.flushed;
await byChannel.flushed;
await mtimes.flushed;
@@ -689,6 +810,12 @@ export async function buildIndex({
pagesWritten: number;
pagesSkipped: number;
slugToPage: Record<string, number>;
+ // Content hash of each emitted page, by page index. Recorded for BOTH the
+ // written and the skipped branch (a skipped page's hash is by definition
+ // the one already on disk), so this is the true current content hash
+ // regardless of whether the page was rewritten this build. The digests tree
+ // publishes these in its manifest as a client cache version.
+ pageHashes: string[];
};
const createPageWriter = (opts: PageWriterOpts) => {
@@ -705,6 +832,7 @@ export async function buildIndex({
let pagesSkippedLocal = 0;
let dirEnsured = !opts.ensureDir;
const slugToPage: Record<string, number> = {};
+ const pageHashes: string[] = [];
const writeChunk = async (chunk: string): Promise<void> => {
const s = stream;
@@ -743,6 +871,7 @@ export async function buildIndex({
});
});
const digest = hash.digest("hex");
+ pageHashes[pageIdx] = digest;
const prev = opts.getPrevHash(pageIdx);
if (prev === digest && (await exists(outPath))) {
await rm(tmpPath, { force: true });
@@ -794,6 +923,7 @@ export async function buildIndex({
pagesWritten: pagesWrittenLocal,
pagesSkipped: pagesSkippedLocal,
slugToPage,
+ pageHashes,
};
};
@@ -1240,6 +1370,156 @@ export async function buildIndex({
}
// ---------------------------------------------------------------------------
+ // The AI-digest corpus: a third shared per-channel page tree, holding the
+ // chapters + topic tags derived from each transcript. Unlike transcripts and
+ // subs it is SPARSE — a channel emits a manifest only if at least one of its
+ // videos has been digested, and a manifest's slugToPage lists only those
+ // videos. That sparsity is load-bearing downstream: it is what lets the
+ // viewer decide whether to offer a Digest control without a per-video fetch.
+ // ---------------------------------------------------------------------------
+ const channelDigestStats = new Map<string, ChannelDigestStat>();
+ let digestPagesWritten = 0;
+ let digestPagesSkipped = 0;
+ let digestTotalCount = 0;
+
+ if (sharedNeedsBuild) {
+ for (const channelSlug of Array.from(channelConfigs.keys()).sort()) {
+ const cfg = channelConfigs.get(channelSlug)!;
+ const digestChannelDir = path.join(digestsOutDir, channelSlug);
+ let digestCount = 0;
+
+ const digestWriter = createPageWriter({
+ outDir: digestChannelDir,
+ fileName: digestPageFileName,
+ maxPageBytes: maxTranscriptPageBytes,
+ ensureDir: true,
+ getPrevHash: (idx) => digestPageHashes.get([channelSlug, idx])?.hash,
+ setHash: (idx, record) => {
+ digestPageHashes.put([channelSlug, idx], record);
+ },
+ onLog: log,
+ });
+
+ // Oldest-first (the byChannel key order), matching the transcript tree so
+ // adding a newer digest only dirties the last page.
+ for (const { key } of byChannel.getRange({
+ start: [channelSlug],
+ end: [channelSlug, ""],
+ })) {
+ const ck = key as ChannelKey;
+ if (ck[0] !== channelSlug) continue;
+ const indexKey: IndexKey = [ck[1], ck[0], ck[2]];
+ const digest = digests.get(indexKey);
+ if (!digest) continue;
+ await digestWriter.push(JSON.stringify(digest), digest.id);
+ digestCount++;
+ }
+
+ const {
+ pageCount: digestPageCount,
+ pagesWritten: chDigestsWritten,
+ pagesSkipped: chDigestsSkipped,
+ slugToPage: digestSlugToPage,
+ pageHashes: digestHashes,
+ } = await digestWriter.finish();
+ digestPagesWritten += chDigestsWritten;
+ digestPagesSkipped += chDigestsSkipped;
+
+ if (digestCount === 0) {
+ // No digests in this channel — drop any stale tree + hashes so a
+ // channel whose digests were deleted stops advertising them.
+ await rm(digestChannelDir, { recursive: true, force: true });
+ for (const { key } of digestPageHashes.getRange({
+ start: [channelSlug],
+ end: [channelSlug, Number.MAX_SAFE_INTEGER],
+ })) {
+ digestPageHashes.remove(key as PageHashKey);
+ }
+ continue;
+ }
+
+ const digestKeep = new Set<string>(["manifest.json"]);
+ for (let i = 0; i < digestPageCount; i++) {
+ digestKeep.add(digestPageFileName(i));
+ }
+ for (const name of await readdir(digestChannelDir).catch(
+ () => [] as string[],
+ )) {
+ if (digestKeep.has(name)) continue;
+ await rm(path.join(digestChannelDir, name), { force: true });
+ }
+ for (const { key } of digestPageHashes.getRange({
+ start: [channelSlug, digestPageCount],
+ end: [channelSlug, Number.MAX_SAFE_INTEGER],
+ })) {
+ digestPageHashes.remove(key as PageHashKey);
+ }
+
+ const channelDigestsManifest: ChannelDigestsManifest = {
+ version: DIGESTS_MANIFEST_VERSION,
+ channelSlug,
+ pageCount: digestPageCount,
+ maxPageBytes: maxTranscriptPageBytes,
+ generatedAt,
+ slugToPage: digestSlugToPage,
+ pageHashes: digestHashes,
+ };
+ await writeJsonAtomic(
+ path.join(digestChannelDir, "manifest.json"),
+ channelDigestsManifest,
+ );
+ channelDigestStats.set(channelSlug, {
+ name: cfg.name ?? channelSlug,
+ digestCount,
+ });
+ digestTotalCount += digestCount;
+ }
+
+ await digestPageHashes.flushed;
+
+ // Drop shared digest dirs for channels that no longer have any.
+ const topDigestEntries = await readdir(digestsOutDir, {
+ withFileTypes: true,
+ }).catch(() => [] as Dirent[]);
+ for (const e of topDigestEntries) {
+ if (e.isDirectory()) {
+ if (!channelDigestStats.has(e.name)) {
+ await rm(path.join(digestsOutDir, e.name), {
+ recursive: true,
+ force: true,
+ });
+ }
+ } else if (e.isFile()) {
+ // The site-level digests manifest is per-site; the shared root holds
+ // only per-channel dirs.
+ await rm(path.join(digestsOutDir, e.name), { force: true });
+ }
+ }
+
+ await channelDigestStatsDb.clearAsync();
+ for (const [slug, statRec] of channelDigestStats) {
+ channelDigestStatsDb.put(slug, statRec);
+ }
+ await channelDigestStatsDb.flushed;
+
+ if (digestTotalCount > 0 || digestPagesWritten > 0) {
+ log(
+ `Digest pages: ${digestPagesWritten} written, ${digestPagesSkipped} unchanged ` +
+ `across ${channelDigestStats.size} channel(s) (${digestTotalCount} digests).`,
+ );
+ }
+ } else {
+ for (const { key, value } of channelDigestStatsDb.getRange()) {
+ channelDigestStats.set(key as string, value as ChannelDigestStat);
+ }
+ if (channelDigestStats.size > 0) {
+ log(
+ `Shared digest pages up to date; ${channelDigestStats.size} channel(s) with digests.`,
+ );
+ }
+ }
+
+ // ---------------------------------------------------------------------------
// Per-site aggregates: a filtered summaries index (pages + manifest) and a
// site-level subs manifest, one bundle per configured site. The heavy
// per-channel page trees above are shared; here we only select + regroup.
@@ -1277,9 +1557,11 @@ export async function buildIndex({
const summariesOut = siteSummariesDir(paths, site.siteId);
const subsOut = siteSubsDir(paths, site.siteId);
const postsOut = sitePostsDir(paths, site.siteId);
+ const digestsOut = siteDigestsDir(paths, site.siteId);
const summariesManifestPath = path.join(summariesOut, "manifest.json");
const subsManifestPath = path.join(subsOut, "manifest.json");
const postsManifestPath = path.join(postsOut, "manifest.json");
+ const digestsManifestPath = path.join(digestsOut, "manifest.json");
const fingerprint = JSON.stringify({
gen: generation,
@@ -1447,6 +1729,33 @@ export async function buildIndex({
await mkdir(postsOut, { recursive: true });
await writeJsonAtomic(postsManifestPath, sitePostsManifest);
+ // --- site-level digests manifest (filtered to member channels) ---
+ // Only channels that actually carry digests contribute, so a site with none
+ // ships a manifest with an empty channel list rather than no manifest —
+ // keeping the client's fetch unconditional, as with posts.
+ const digestEntries: DigestsChannelEntry[] = [];
+ let digestsTotalForSite = 0;
+ for (const slug of memberSlugs) {
+ const statRec = channelDigestStats.get(slug);
+ if (!statRec || statRec.digestCount === 0) continue;
+ digestEntries.push({
+ name: statRec.name,
+ slug,
+ digestCount: statRec.digestCount,
+ groupId: slugGroup.get(slug) ?? site.defaultGroupId,
+ });
+ digestsTotalForSite += statRec.digestCount;
+ }
+ const siteDigestsManifest: DigestsManifest = {
+ version: SITE_DIGESTS_MANIFEST_VERSION,
+ channels: digestEntries.sort((a, b) => a.name.localeCompare(b.name)),
+ totalCount: digestsTotalForSite,
+ generatedAt: new Date().toISOString(),
+ siteId: site.siteId,
+ };
+ await mkdir(digestsOut, { recursive: true });
+ await writeJsonAtomic(digestsManifestPath, siteDigestsManifest);
+
await meta.put(fpKey, fingerprint);
sitesBuilt++;
aggregateSummaryPages += pageIndex;
diff --git a/common/controller/channelSnapshot.ts b/common/controller/channelSnapshot.ts
@@ -23,6 +23,9 @@ import {
resolveEffectiveAvailability,
} from "../lib/availability-server";
import { isDoNotClean } from "../lib/doNotClean-server";
+import { loadDigest } from "../lib/digest-server";
+import { isSectionFresh } from "../lib/digest";
+import { resolveDigestTarget } from "./digestTarget";
import { isExcludedFromTruncatedCheck } from "../lib/excludeTruncatedCheck-server";
import { loadDownloadOutcome } from "../lib/downloadOutcome-server";
import type { Paths } from "../lib/paths";
@@ -54,6 +57,13 @@ export type ChannelSnapshot = {
transcribed: number;
downloaded: number;
};
+ // Digest coverage by engine: appId -> count of videos whose ai-digest.json was
+ // produced by it. Beside `totals`, NOT in `buckets` — buckets are a closed
+ // literal of `string[]` id lists and a Record<string, number> does not belong
+ // there. During a multi-week sweep this split is what tells you whether the
+ // local lane is actually carrying the corpus. Optional: older snapshots lack
+ // it; readers default to {}.
+ digestEngines?: Record<string, number>;
buckets: {
noTranscript: string[];
downloadedNoTranscript: string[];
@@ -140,6 +150,31 @@ export type ChannelSnapshot = {
// (retry-bucket with forceCookies). Optional: older snapshots lack it;
// readers must default to [].
needsCookies: string[];
+ // Transcribed videos whose digest is missing or STALE against the local
+ // lane's current freshness target (engine, requested model, prompt version,
+ // prompt shape, context hash) — the AI digest layer's work list, and the
+ // denominator for corpus coverage during the backfill. This is the same
+ // target countMissingDigests and runDigestBatch compute, deliberately: a
+ // bucket that silently meant something weaker than its name is how a sweep
+ // reports "nothing to do" on a corpus that needs redoing.
+ //
+ // Only TRANSCRIBED videos are listed: a video without a transcript is a
+ // transcription problem, not a digest one. Optional: older snapshots lack
+ // it; readers must default to [].
+ noDigest: string[];
+ // Videos whose digest pass recorded something a human should look at:
+ // either warnings alongside a section that WAS written, or a total failure
+ // that wrote no section at all (DigestRecord.failures).
+ //
+ // Both kinds matter and the second is the one that used to be invisible —
+ // a total failure deliberately writes no section so the video retries, and
+ // before `failures` existed its warnings survived only in a rotating job
+ // log. A review queue keyed on written warnings alone would have been blind
+ // to precisely the worst outputs.
+ //
+ // Costs no extra I/O: the sidecar is already loaded here for noDigest.
+ // Optional: older snapshots lack it; readers must default to [].
+ digestWarnings?: string[];
};
undownloadedIds: string[];
excludedFromDownload?: ExcludedFromDownload;
@@ -350,6 +385,13 @@ export async function generateChannelSnapshot(
const vttProvenance = files.ytVttFile
? await resolveVttProvenance(dir, files.ytVttFile)
: null;
+ // Only transcribed videos can carry a digest, so everything else skips
+ // the sidecar read entirely — the same conditional per-video
+ // sidecar-read pattern as the two reads above.
+ const digest =
+ isVideoTranscribed(files) && !files.isUntranscribable
+ ? await loadDigest(dir)
+ : null;
return {
id,
files,
@@ -362,6 +404,7 @@ export async function generateChannelSnapshot(
outcome,
coverage,
vttProvenance,
+ digest,
};
}),
),
@@ -434,6 +477,27 @@ export async function generateChannelSnapshot(
const autoSubsOnly: string[] = [];
const downloadedAutoSubsOnly: string[] = [];
const supersededAutoSubs: string[] = [];
+ const noDigest: string[] = [];
+ const digestWarnings: string[] = [];
+ const digestEngines: Record<string, number> = {};
+ // The SAME freshness target countMissingDigests and runDigestBatch use, so
+ // the Digest stage's count and the batch runner's progress target cannot
+ // disagree. This bucket used to ask only "is there an ai-digest.json with
+ // items?", which meant that after any prompt/model/context change the stage
+ // read "All digested." while the batch reported the whole channel as stale.
+ //
+ // Affordable in this hot path because the per-video cost is ZERO extra I/O:
+ // the digest sidecar is already loaded above, and isSectionFresh is a pure
+ // comparison. Only the target itself is new work, and it is resolved ONCE per
+ // channel (a settings read, a registry lookup, and one small context file).
+ //
+ // The LOCAL lane is the target on purpose: it is the lane that carries the
+ // corpus, and the metered lane exists only for the long tail.
+ const digestTarget = await resolveDigestTarget({
+ paths,
+ channelSlug: slug,
+ lane: "local",
+ });
let transcribedWithAudioBytes = 0;
let multipleAudioFormatsBytes = 0;
let foreignAudioBytes = 0;
@@ -446,6 +510,7 @@ export async function generateChannelSnapshot(
outcome,
coverage,
vttProvenance,
+ digest,
} of perVideo) {
if (isVideoTranscribed(files)) transcribed++;
if (isVideoDownloaded(files)) downloaded++;
@@ -575,6 +640,44 @@ export async function generateChannelSnapshot(
if (files.audioFiles.length > 0) downloadedAutoSubsOnly.push(id);
else autoSubsOnly.push(id);
}
+ // --- AI digest coverage -------------------------------------------------
+ // Counted before the untranscribable/no-transcript `continue`s below so the
+ // accounting is unambiguous: a video is either a digest candidate or not a
+ // transcript at all.
+ if (isVideoTranscribed(files) && !files.isUntranscribable) {
+ const chapters = digest?.sections.chapters;
+ const tags = digest?.sections.tags;
+ const engine = chapters?.provenance.appId ?? tags?.provenance.appId;
+ const hasItems =
+ (chapters?.items.length ?? 0) > 0 || (tags?.items.length ?? 0) > 0;
+ // Coverage-by-engine answers "what produced what is on disk?" and stays
+ // identity-BLIND — a section generated by an older prompt version was
+ // still generated by that engine, and hiding it would make the coverage
+ // split lie during exactly the config change it exists to survey.
+ if (engine && hasItems) {
+ digestEngines[engine] = (digestEngines[engine] ?? 0) + 1;
+ }
+ // The WORK LIST is identity-aware: a video whose digest predates the
+ // current identity is work, not coverage. A digest shared from a
+ // duplicate cluster's canonical member counts as done — the canonical
+ // member's own freshness is what drives regeneration, and the share is
+ // re-applied from it (isSharedFrom's contract).
+ const fresh =
+ digest?.derivedFrom != null ||
+ digestTarget.sections.every((section) =>
+ isSectionFresh(digest, section, digestTarget.target),
+ );
+ if (!fresh) noDigest.push(id);
+ // Reviewable regardless of freshness: a video that failed outright is
+ // ALSO in noDigest (it has no section), and a video whose section landed
+ // with warnings is fresh and would otherwise never be surfaced again.
+ if (
+ (digest?.warnings?.length ?? 0) > 0 ||
+ (digest?.failures?.length ?? 0) > 0
+ ) {
+ digestWarnings.push(id);
+ }
+ }
if (files.isUntranscribable) {
untranscribable.push(id);
continue;
@@ -681,7 +784,10 @@ export async function generateChannelSnapshot(
downloadedAutoSubsOnly: downloadedAutoSubsOnly.sort(),
supersededAutoSubs: supersededAutoSubs.sort(),
needsCookies: needsCookies.sort(),
+ noDigest: noDigest.sort(),
+ digestWarnings: digestWarnings.sort(),
},
+ digestEngines,
undownloadedIds,
excludedFromDownload,
keptCount: keptIds.size,
diff --git a/common/controller/channels.ts b/common/controller/channels.ts
@@ -11,6 +11,7 @@ import {
isVideoTranscribed,
readVideoFiles,
} from "../lib/videoStatus";
+import { hasDigest } from "../lib/digest-server";
export type ChannelStat = {
slug: string;
@@ -19,6 +20,11 @@ export type ChannelStat = {
videoCount: number;
transcriptCount: number;
downloadCount: number;
+ // Videos carrying a non-empty ai-digest.json. The Active Jobs progress bar
+ // re-counts `current` from disk rather than trusting the runner, so a digest
+ // job needs this counter to have a bar at all. Optional so a caller reading an
+ // older serialized stat still type-checks.
+ digestCount?: number;
};
// A channel slug is also its directory name under transcripts/channels/, so it
@@ -43,32 +49,42 @@ async function exists(p: string): Promise<boolean> {
}
}
-async function countDataFiles(
- dataDir: string,
-): Promise<{ videos: number; transcripts: number; downloads: number }> {
+async function countDataFiles(dataDir: string): Promise<{
+ videos: number;
+ transcripts: number;
+ downloads: number;
+ digests: number;
+}> {
let dirs: Dirent[];
try {
dirs = await readdir(dataDir, { withFileTypes: true });
} catch {
- return { videos: 0, transcripts: 0, downloads: 0 };
+ return { videos: 0, transcripts: 0, downloads: 0, digests: 0 };
}
const videoDirs = dirs.filter((d) => d.isDirectory());
const flags = await Promise.all(
videoDirs.map(async (d) => {
- const files = await readVideoFiles(path.join(dataDir, d.name));
+ const dir = path.join(dataDir, d.name);
+ const files = await readVideoFiles(dir);
return {
transcript: isVideoTranscribed(files),
download: isVideoDownloaded(files),
+ // Only transcribed videos can carry a digest, so the sidecar read is
+ // skipped for the rest — the same conditional per-video sidecar-read
+ // pattern channelSnapshot.ts uses for coverage and VTT provenance.
+ digest: isVideoTranscribed(files) ? await hasDigest(dir) : false,
};
}),
);
let transcripts = 0;
let downloads = 0;
+ let digests = 0;
for (const f of flags) {
if (f.transcript) transcripts++;
if (f.download) downloads++;
+ if (f.digest) digests++;
}
- return { videos: videoDirs.length, transcripts, downloads };
+ return { videos: videoDirs.length, transcripts, downloads, digests };
}
async function countPlaylist(p: string): Promise<number | null> {
@@ -122,6 +138,7 @@ export async function readChannelStat(
videoCount: counts.videos,
transcriptCount: counts.transcripts,
downloadCount: counts.downloads,
+ digestCount: counts.digests,
};
}
@@ -180,6 +197,11 @@ export async function listChannels(paths: Paths): Promise<ChannelStat[]> {
videoCount: counts.videos,
transcriptCount: counts.transcripts,
downloadCount: counts.downloads,
+ // `countDataFiles` has always computed this and `listChannels` has always
+ // discarded it — unlike readChannelStat, which emits it. One line, and
+ // every cross-channel surface (the dashboard, /actionable, the widget)
+ // gets a coverage counter it was already paying the I/O for.
+ digestCount: counts.digests,
});
}
return out.sort((a, b) => a.slug.localeCompare(b.slug));
diff --git a/common/controller/checkAvailability.ts b/common/controller/checkAvailability.ts
@@ -20,6 +20,10 @@ import {
resolveCookiePolicy,
} from "../lib/cookiePolicy";
import { getSettings } from "../lib/settings";
+// yt-dlp --dump-json emits one JSON object per video, but may print warnings to
+// stdout first — so the payload is the LAST parseable line. Shared with the
+// digest registry's claude-code lane, which has the same problem.
+import { parseStdoutJson } from "../lib/parseStdoutJson";
import { readChannelConfig } from "./channels";
import { resolveShardItems } from "./shard";
import type { Paths } from "../lib/paths";
@@ -88,20 +92,6 @@ async function readWebpageUrl(videoDir: string): Promise<string | null> {
}
}
-function parseStdoutJson(stdout: string): unknown | null {
- // yt-dlp --dump-json emits one JSON object per video; for a single URL we
- // expect one line. If something else printed warnings to stdout, the JSON
- // is the last non-empty line.
- const lines = stdout.split("\n").map((l) => l.trim()).filter(Boolean);
- for (let i = lines.length - 1; i >= 0; i--) {
- try {
- return JSON.parse(lines[i]);
- } catch {
- continue;
- }
- }
- return null;
-}
export async function runAvailabilityCheck({
channelSlug,
diff --git a/common/controller/digestBatch.ts b/common/controller/digestBatch.ts
@@ -0,0 +1,571 @@
+// Channel-scoped digest sweep: the thing that actually runs for weeks.
+//
+// Three mechanics here are copied from existing code for specific reasons, and
+// changing any of them breaks a property a multi-week sweep depends on:
+//
+// 1. RESUME BY RE-DERIVING FROM DISK. Eligibility is re-checked against the
+// sidecar on every next() pull (the auto-runner's shape), NOT frozen into an
+// array with an index cursor (whisperBatch's shape). The cursor form does not
+// survive a restart, and it cannot see a video that became eligible mid-run
+// (a transcript that finished, a mirror that got shared to).
+// 2. PAUSE BY RETURNING limit() === 0. runPool idle-WAITS at a zero limit
+// rather than finishing (concurrentRunner.ts's documented invariant), so a
+// pause holds the job open instead of ending it. The flag is re-read from
+// settings on every pull — the downloadsPaused pattern, which needs no boot
+// hook, unlike transcriptionsPaused.
+// 3. ONE JOB PER CHANNEL, never one per video. Job logs keep the newest 500 /
+// 30 days and the in-memory registry keeps 100 records; 119,600 jobs would
+// evict everything, including the running ones' history.
+
+import path from "node:path";
+import { open, readdir } from "node:fs/promises";
+import type { Paths } from "../lib/paths";
+import { getSettings } from "../lib/settings";
+import { runPool } from "../jobs/concurrentRunner";
+import type { TaskTracker } from "../jobs/taskHooks";
+import type { JobProgress } from "../jobs/registry";
+import { isSectionFresh, type DigestSectionKind } from "../lib/digest";
+import { loadDigest } from "../lib/digest-server";
+import {
+ resolveDigestTarget,
+ type DigestLaneChoice,
+} from "./digestTarget";
+import { CUES_JSON_FILENAME } from "../lib/videoStatus";
+import { digestVideo } from "./digestVideo";
+import { transcriptionActivity } from "./digestYield";
+import {
+ buildDigestClusterPlan,
+ planSlugForDir,
+ shareDigestToCluster,
+ type DigestClusterPlan,
+} from "./digestSharing";
+
+export type DigestOrder = "shortest-first" | "longest-first";
+
+export type DigestBatchOptions = {
+ channelSlug: string;
+ paths: Paths;
+ // Which lane to run. "local" is the default and the only one enabled unless
+ // settings.digest.remoteEnabled is true; asking for "remote" while it is off is
+ // an error the caller surfaces, not a silent downgrade.
+ lane?: "local" | "remote";
+ // When set, only consider these video ids (intersected with what's on disk).
+ ids?: string[];
+ sections?: DigestSectionKind[];
+ // The repo only had a `reverse` flag before this; digest ordering is explicit
+ // because lane routing and the pilot both depend on it. Shortest-first by
+ // default: it converts the backlog into visible coverage fastest, and the
+ // long-tail 8% is where a prompt bug is most expensive to discover late.
+ order?: DigestOrder;
+ // Duration window, for splitting the corpus between lanes (e.g. local takes
+ // ≤ longTailSeconds, the metered lane takes the tail).
+ minDurationSeconds?: number;
+ maxDurationSeconds?: number;
+ // Stop after this many successful videos. For the pilot and Stage B's sample.
+ limitCount?: number;
+ concurrency?: number;
+ force?: boolean;
+ // Skip mirrors and share the canonical member's digest to aligned ones. On by
+ // default: worth ~11% of the sweep. Pass a prebuilt plan to avoid re-reading
+ // the duplicates report.
+ useClusters?: boolean;
+ clusterPlan?: DigestClusterPlan;
+ setProgress?: (snap: JobProgress) => void;
+ // Where the bar starts: how many videos already counted as done before this
+ // run. The batch reports `initial + completed` against it rather than letting
+ // the UI re-count files from disk, which cannot see a regeneration.
+ progressBaseline?: number;
+ // Where the bar ends. Supplied by the caller because it already resolves it
+ // (countMissingDigests) to size the job; the batch raises it if sharing turns
+ // out to satisfy more videos than were counted missing.
+ progressTarget?: number;
+ onLog?: (msg: string) => void;
+ signal?: AbortSignal;
+ drainSignal?: AbortSignal;
+ tracker?: TaskTracker;
+};
+
+export type DigestBatchResult = {
+ attempted: number;
+ succeeded: number;
+ skipped: number;
+ failed: number;
+ fresh: number;
+ // Cluster mirrors that received the canonical member's digest, and mirrors the
+ // alignment gate refused. A refusal is a normal outcome, not an error.
+ shared: number;
+ misaligned: number;
+ engineCalls: number;
+ costUsd: number;
+ warnings: number;
+ // True when the run stopped early because the spend cap was reached.
+ spendCapped: boolean;
+};
+
+type Candidate = { id: string; duration: number };
+
+// Read a video's duration cheaply. transcript.cues.json is
+// {version, source, transcriptFormat, ...summary, cues} — the summary (and so
+// `duration`) is serialized BEFORE the multi-megabyte cues array, so the head of
+// the file is enough and a full parse is only the fallback. At 74k videos this is
+// the difference between a few seconds and reading 6.9 GB to sort a list.
+const DURATION_HEAD_BYTES = 8192;
+
+async function readDurationFast(cuesPath: string): Promise<number | null> {
+ let handle: Awaited<ReturnType<typeof open>> | null = null;
+ try {
+ handle = await open(cuesPath, "r");
+ const buf = Buffer.alloc(DURATION_HEAD_BYTES);
+ const { bytesRead } = await handle.read(buf, 0, DURATION_HEAD_BYTES, 0);
+ const head = buf.subarray(0, bytesRead).toString("utf8");
+ const m = head.match(/"duration"\s*:\s*([0-9]+(?:\.[0-9]+)?)/);
+ if (m) return Number(m[1]);
+ // Short file: the whole thing is in `head`, so parse it properly.
+ if (bytesRead < DURATION_HEAD_BYTES) {
+ const parsed = JSON.parse(head) as { duration?: unknown };
+ return typeof parsed.duration === "number" ? parsed.duration : null;
+ }
+ return null;
+ } catch {
+ return null;
+ } finally {
+ await handle?.close().catch(() => {});
+ }
+}
+
+export async function runDigestBatch(
+ opts: DigestBatchOptions,
+): Promise<DigestBatchResult> {
+ const log = opts.onLog ?? ((m: string) => console.log(m));
+ const settings = getSettings();
+ const digestSettings = settings.digest;
+ const lane = opts.lane ?? "local";
+ if (lane === "remote" && !digestSettings.remoteEnabled) {
+ throw new Error(
+ "The metered digest lane is disabled (settings.digest.remoteEnabled). Enable it in Settings before running it.",
+ );
+ }
+ // ONE derivation, shared with countMissingDigests and the channel snapshot's
+ // noDigest bucket — see digestTarget.ts. It must match what digestVideo will
+ // actually chunk with, or the batch's freshness check and the writer would
+ // disagree on the identity and every video would look stale forever.
+ const resolved = await resolveDigestTarget({
+ paths: opts.paths,
+ channelSlug: opts.channelSlug,
+ lane,
+ sections: opts.sections,
+ });
+ const { app, config, sections, modelRequested, timestampMode, context } =
+ resolved;
+
+ const dataDir = path.join(opts.paths.channelsDir, opts.channelSlug, "data");
+
+ // Fail fast and loudly rather than 1,100 times in a row: an unreachable engine
+ // is a configuration problem, and discovering it per-item wastes the log.
+ if (!(await app.probe(config))) {
+ throw new Error(
+ `Digest engine ${app.id} is not reachable. ` +
+ (app.lane === "local-gpu"
+ ? "Is the ollama service running (systemctl status ollama)?"
+ : "Is the claude CLI installed and on PATH (set CLAUDE_BIN)?"),
+ );
+ }
+
+ const allDirs = await readdir(dataDir).catch(() => [] as string[]);
+ const onDisk = new Set(allDirs);
+ const wanted = opts.ids
+ ? opts.ids.filter((id) => onDisk.has(id))
+ : allDirs;
+
+ const clusterPlan =
+ opts.useClusters === false
+ ? null
+ : (opts.clusterPlan ?? (await buildDigestClusterPlan(opts.paths)));
+
+ // Durations, read once. Videos with no normalized transcript have no duration
+ // and are dropped here — they are a transcription problem, not a digest one.
+ const candidates: Candidate[] = [];
+ let noTranscript = 0;
+ let outOfWindow = 0;
+ for (const id of wanted) {
+ opts.signal?.throwIfAborted();
+ const duration = await readDurationFast(
+ path.join(dataDir, id, CUES_JSON_FILENAME),
+ );
+ if (duration === null || duration <= 0) {
+ noTranscript++;
+ continue;
+ }
+ if (
+ opts.minDurationSeconds !== undefined &&
+ duration < opts.minDurationSeconds
+ ) {
+ outOfWindow++;
+ continue;
+ }
+ if (
+ opts.maxDurationSeconds !== undefined &&
+ duration > opts.maxDurationSeconds
+ ) {
+ outOfWindow++;
+ continue;
+ }
+ candidates.push({ id, duration });
+ }
+
+ const order = opts.order ?? "shortest-first";
+ candidates.sort((a, b) =>
+ order === "longest-first"
+ ? b.duration - a.duration || a.id.localeCompare(b.id)
+ : a.duration - b.duration || a.id.localeCompare(b.id),
+ );
+
+ log(
+ `Digest ${opts.channelSlug} (${lane} lane, ${app.id}/${modelRequested}, sections: ${sections.join(", ")}): ` +
+ `${candidates.length} candidate(s) of ${wanted.length} on disk ` +
+ `(${noTranscript} without a transcript, ${outOfWindow} outside the duration window), ${order}.` +
+ (clusterPlan
+ ? ` Duplicate plan: ${clusterPlan.clusters} cluster(s), ${clusterPlan.bySlug.size} member(s) mapped.`
+ : " Duplicate sharing off."),
+ );
+
+ const result: DigestBatchResult = {
+ attempted: 0,
+ succeeded: 0,
+ skipped: 0,
+ failed: 0,
+ fresh: 0,
+ shared: 0,
+ misaligned: 0,
+ engineCalls: 0,
+ costUsd: 0,
+ warnings: 0,
+ spendCapped: false,
+ };
+
+ const freshnessTarget = resolved.target;
+
+ // Ids already handed out this run. Combined with the cursor below, this is what
+ // makes the disk re-derivation O(n) overall rather than O(n²): the cursor only
+ // ever moves forward past ids that have been attempted, while eligibility for
+ // the id it stops on is always re-read from disk.
+ const attempted = new Set<string>();
+ let cursor = 0;
+
+ // `candidate.id` is a DIRECTORY name (it comes from readdir), and a slug is
+ // `${channelSlug}/${metadataId}`. Those disagree for 14.5% of the corpus —
+ // every Rumble re-upload — so resolving one to the other through the plan is
+ // what makes the cluster lookups below hit at all. See planSlugForDir.
+ const slugOf = (videoDir: string): string =>
+ planSlugForDir(clusterPlan, opts.channelSlug, videoDir);
+
+ const next = async (): Promise<Candidate | null> => {
+ while (cursor < candidates.length) {
+ if (opts.limitCount !== undefined && result.succeeded >= opts.limitCount) {
+ return null;
+ }
+ const candidate = candidates[cursor];
+ if (attempted.has(candidate.id)) {
+ cursor++;
+ continue;
+ }
+ const videoDir = path.join(dataDir, candidate.id);
+
+ // A cluster mirror is not this lane's work: its canonical member owns the
+ // generation and shares the result here.
+ const role = clusterPlan?.bySlug.get(slugOf(candidate.id));
+ if (role?.kind === "mirror") {
+ attempted.add(candidate.id);
+ cursor++;
+ resolvedAudioSeconds += candidate.duration;
+ result.skipped++;
+ log(
+ `Skipping ${candidate.id}: duplicate of ${role.canonicalSlug}, which owns the digest for this cluster.`,
+ );
+ continue;
+ }
+
+ // RE-DERIVED FROM DISK, every pull. A restart, a concurrent lane, or a
+ // share that landed while this job ran are all visible here.
+ if (!opts.force) {
+ const record = await loadDigest(videoDir);
+ const allFresh = sections.every((section) =>
+ isSectionFresh(record, section, freshnessTarget),
+ );
+ if (allFresh) {
+ attempted.add(candidate.id);
+ cursor++;
+ resolvedAudioSeconds += candidate.duration;
+ result.fresh++;
+ continue;
+ }
+ }
+ attempted.add(candidate.id);
+ cursor++;
+ return candidate;
+ }
+ return null;
+ };
+
+ // WIRING THE PLUMBING THAT WAS DECLARED AND NEVER USED.
+ //
+ // `setProgress` and `progressBaseline` were on the options type and passed by
+ // digestActions.ts, but nothing in this file referenced either. Progress
+ // appeared to work only because buildActiveJobs re-counts `current` from disk
+ // — and that re-count is exactly what a REGENERATION defeats: a regenerated
+ // digest is rewritten in place, the file count never moves, and the bar sits
+ // at 0% for the whole job.
+ //
+ // The batch is the only thing that knows the truth, for two reasons the disk
+ // cannot express: it knows a regenerate happened, and it knows sharing moved
+ // the denominator — a canonical member's digest can satisfy a dozen mirrors
+ // at once, so a sweep genuinely CHANGES its own target as it runs.
+ const baseline = opts.progressBaseline ?? 0;
+ const progressTarget = opts.progressTarget ?? null;
+ // Audio-seconds retired from the worklist, by ANY route — generated, shared
+ // to, found fresh, or skipped as a mirror. All four remove work, and an ETA
+ // that only counted generations would keep quoting time for videos that are
+ // already done.
+ const totalAudioSeconds = candidates.reduce((n, c) => n + c.duration, 0);
+ let resolvedAudioSeconds = 0;
+ const reportProgress = (): void => {
+ if (!opts.setProgress) return;
+ // Counted against the TARGET, which is the count of videos that were not
+ // fresh when the job was sized. So:
+ // succeeded + shared — work this run did, and did against the target;
+ // skipped — was in the target and is resolved anyway (a mirror
+ // its canonical owns, or a video with no transcript);
+ // without it the bar stalls short of done forever.
+ // `fresh` is deliberately EXCLUDED: a video already fresh was never in the
+ // target, and counting it would drive the bar past 100% on any channel that
+ // is mostly done — 90 fresh of 100 would report 100/10.
+ const done = result.succeeded + result.shared + result.skipped;
+ const remainingAudioSeconds = Math.max(
+ 0,
+ totalAudioSeconds - resolvedAudioSeconds,
+ );
+ // Never let the bar exceed its target: sharing can satisfy more videos than
+ // the target was sized for, and a bar past 100% reads as a bug rather than
+ // as good news.
+ const current = baseline + done;
+ opts.setProgress({
+ metric: "digests",
+ initial: baseline,
+ target: Math.max(progressTarget ?? current, current),
+ current,
+ remainingAudioSeconds,
+ });
+ };
+
+ const runOne = async (
+ candidate: Candidate,
+ runSignal: AbortSignal,
+ ): Promise<void> => {
+ const task = opts.tracker?.start({
+ id: candidate.id,
+ label: `digest ${opts.channelSlug}/${candidate.id}`,
+ kind: "digest",
+ // What makes seconds-per-audio-hour computable. A digest's cost is
+ // proportional to the transcript's LENGTH, not to it being one video, so
+ // a task-count average is the wrong denominator for a sweep ETA.
+ audioSeconds: candidate.duration,
+ });
+ try {
+ const outcome = await digestVideo({
+ paths: opts.paths,
+ channelSlug: opts.channelSlug,
+ videoId: candidate.id,
+ sections,
+ appId: app.id,
+ config,
+ context,
+ timestampMode,
+ ...(digestSettings.promptVariant
+ ? { promptVariant: digestSettings.promptVariant }
+ : {}),
+ force: opts.force,
+ onLog: task ? task.onLog : opts.onLog,
+ signal: runSignal,
+ });
+ if (outcome.status === "fresh") {
+ result.fresh++;
+ return;
+ }
+ if (outcome.status === "skipped") {
+ result.skipped++;
+ log(`Skipped ${candidate.id}: ${outcome.reason}.`);
+ return;
+ }
+ result.attempted++;
+ result.succeeded++;
+ result.engineCalls += outcome.engineCalls;
+ result.costUsd += outcome.costUsd;
+ result.warnings += outcome.warningCount;
+
+ // Share to this cluster's aligned mirrors, right after the canonical
+ // member's digest lands — so a mirror never sits un-digested waiting for a
+ // second pass, and a crash mid-sweep leaves a consistent cluster.
+ const role = clusterPlan?.bySlug.get(slugOf(candidate.id));
+ if (role?.kind === "canonical") {
+ const outcomes = await shareDigestToCluster({
+ paths: opts.paths,
+ clusterId: role.clusterId,
+ canonicalSlug: slugOf(candidate.id),
+ mirrors: role.mirrors,
+ dirBySlug: clusterPlan?.dirBySlug,
+ onLog: log,
+ });
+ for (const o of outcomes) {
+ if (o.status === "shared") result.shared++;
+ else if (o.status === "misaligned") result.misaligned++;
+ }
+ }
+ reportProgress();
+ } catch (err) {
+ if (runSignal.aborted || opts.signal?.aborted) throw err;
+ result.attempted++;
+ result.failed++;
+ log(
+ `Failed ${candidate.id}: ${(err as Error)?.message ?? String(err)}`,
+ );
+ } finally {
+ resolvedAudioSeconds += candidate.duration;
+ task?.end();
+ }
+ };
+
+ // Edge-triggered, so the yield logs twice per contention window rather than
+ // once per poll over a multi-week sweep.
+ let yielding = false;
+
+ const NEVER = new AbortController().signal;
+ // Local lane: 1. The GPU is the bottleneck and a second concurrent generation
+ // just thrashes the same 8 GB of VRAM. The metered lane is network-bound, so it
+ // can overlap — but modestly, since it is paying per call.
+ const defaultConcurrency = app.lane === "local-gpu" ? 1 : 2;
+ const concurrency = Math.max(1, opts.concurrency ?? defaultConcurrency);
+
+ // Seed the bar before the first video finishes, so a job that spends its
+ // first minutes on a 3-hour VOD does not look like it never started.
+ reportProgress();
+
+ await runPool<Candidate>({
+ next,
+ run: runOne,
+ limit: () => {
+ // Re-read at DISPATCH time so a pause takes effect within one poll and
+ // survives a restart with no boot hook. Returning 0 makes runPool
+ // idle-wait, which is a pause; returning null from next() would END the
+ // batch, which is not.
+ if (getSettings().digest.digestsPaused) return 0;
+ // Step aside for whisper. Same mechanism as the pause and for the same
+ // reason it works: a zero limit HOLDS the pool instead of ending the
+ // batch, so the sweep resumes the moment the card is free without
+ // re-deriving anything. Only the local (GPU) lane yields — the metered
+ // lane is network-bound and competes for nothing here.
+ if (
+ app.lane === "local-gpu" &&
+ getSettings().digest.yieldToTranscription
+ ) {
+ const activity = transcriptionActivity();
+ if (activity.busy) {
+ // Logged on the EDGE only. An operator watching a sweep sit at zero
+ // throughput has to be able to tell yielding from wedged, but a line
+ // per poll would bury the job log over a multi-week run.
+ if (!yielding) {
+ yielding = true;
+ log(
+ `Yielding the GPU to transcription (${activity.reason}); the digest lane will resume when it is free.`,
+ );
+ }
+ return 0;
+ }
+ if (yielding) {
+ yielding = false;
+ log("Transcription finished; resuming the digest lane.");
+ }
+ }
+ if (
+ app.metered &&
+ digestSettings.spendCapUsd > 0 &&
+ result.costUsd >= digestSettings.spendCapUsd
+ ) {
+ if (!result.spendCapped) {
+ result.spendCapped = true;
+ log(
+ `Spend cap reached ($${result.costUsd.toFixed(2)} of $${digestSettings.spendCapUsd.toFixed(2)}) — parking the metered lane.`,
+ );
+ }
+ return 0;
+ }
+ return concurrency;
+ },
+ signal: opts.signal ?? NEVER,
+ drainSignal: opts.drainSignal ?? NEVER,
+ finite: true,
+ idlePollMs: 3000,
+ });
+
+ // Metered accounting is logged unconditionally when the lane is metered, even
+ // at zero calls: "this run cost nothing" is information too.
+ if (app.metered) {
+ log(
+ `Metered lane: ${result.engineCalls} model call(s), $${result.costUsd.toFixed(4)} total` +
+ (digestSettings.spendCapUsd > 0
+ ? ` (cap $${digestSettings.spendCapUsd.toFixed(2)})`
+ : " (no cap set)"),
+ );
+ }
+ return result;
+}
+
+// How many of a channel's videos still need a digest at the CURRENT prompt/model
+// identity. Used to scope a job's progress bar, and cheap enough to call before
+// starting one (a small sidecar read per video, no transcript parsing).
+//
+// `lane` is REQUIRED-in-spirit: it used to be absent and the local app was
+// resolved unconditionally, so a metered-lane job's progress bar was sized
+// against the LOCAL engine's identity — and, once numCtx joined the identity,
+// against the local app's context size too. Both call sites already know their
+// lane. It still defaults to "local" so the meaning of an unqualified call is
+// the same as it always was, rather than silently changing under old callers.
+export async function countMissingDigests(
+ paths: Paths,
+ channelSlug: string,
+ ids?: ReadonlyArray<string>,
+ lane: DigestLaneChoice = "local",
+): Promise<number> {
+ const { target, sections } = await resolveDigestTarget({
+ paths,
+ channelSlug,
+ lane,
+ });
+ const dataDir = path.join(paths.channelsDir, channelSlug, "data");
+ const dirs = ids
+ ? [...ids]
+ : await readdir(dataDir).catch(() => [] as string[]);
+ let missing = 0;
+ for (const id of dirs) {
+ // SAME ELIGIBILITY AS THE BATCH, and it has to be. A video with no
+ // normalized transcript is a transcription problem, not a digest one:
+ // runDigestBatch drops it before it is ever a candidate, so counting it here
+ // sizes the progress bar against work the batch will never do and the bar
+ // stalls one short of complete — for the rest of the run.
+ //
+ // Measured: `teamrcn` reported target 8, current 7. The eighth directory is
+ // a channel-id-named folder with no transcript. Corpus-wide that is 2,987
+ // videos, i.e. most channels would never show as finished, which over an
+ // 81-day sweep is indistinguishable from wedged. This is exactly the
+ // counter-disagreement digestTarget.ts's header was written about.
+ const duration = await readDurationFast(
+ path.join(dataDir, id, CUES_JSON_FILENAME),
+ );
+ if (duration === null || duration <= 0) continue;
+ const record = await loadDigest(path.join(dataDir, id));
+ const fresh = sections.every((section) =>
+ isSectionFresh(record, section, target),
+ );
+ if (!fresh) missing++;
+ }
+ return missing;
+}
diff --git a/common/controller/digestPlan.test.ts b/common/controller/digestPlan.test.ts
@@ -0,0 +1,83 @@
+import { test } from "node:test";
+import assert from "node:assert/strict";
+import type { DigestClusterRole } from "./digestSharing";
+import {
+ audioHours,
+ classifyDigestRole,
+ DIGEST_PLAN_ROLES,
+ MEASURED_SECONDS_PER_AUDIO_HOUR,
+ roleMustGenerate,
+ sweepDays,
+} from "./digestPlan";
+
+// Run with:
+// pnpm --filter yt-dlp-transcript-common exec tsx --test common/controller/digestPlan.test.ts
+
+test("a video in no cluster is unclustered work", () => {
+ assert.equal(classifyDigestRole(undefined, undefined), "unclustered");
+ // Alignment is meaningless without a cluster and must not change the answer.
+ assert.equal(classifyDigestRole(undefined, true), "unclustered");
+});
+
+test("a canonical member is work regardless of alignment", () => {
+ const role: DigestClusterRole = {
+ kind: "canonical",
+ clusterId: "c1",
+ mirrors: ["a"],
+ };
+ assert.equal(classifyDigestRole(role, true), "canonical");
+ assert.equal(classifyDigestRole(role, false), "canonical");
+});
+
+// The distinction the whole plan turns on: only an ALIGNED mirror is free. The
+// sharing pass refuses to place a digest without measured alignment, so counting
+// every mirror as free overstates the saving by however many it will refuse —
+// nearly half, on the current corpus.
+test("only an explicitly-aligned mirror is free", () => {
+ const role: DigestClusterRole = {
+ kind: "mirror",
+ clusterId: "c1",
+ canonicalSlug: "x",
+ };
+ assert.equal(classifyDigestRole(role, true), "mirror-aligned");
+ assert.equal(classifyDigestRole(role, false), "mirror-unaligned");
+});
+
+// Absent means "never measured", which sharing treats as NOT aligned. The plan
+// must under-promise here, never over-promise.
+test("an unmeasured mirror counts as work, not as a saving", () => {
+ const role: DigestClusterRole = {
+ kind: "mirror",
+ clusterId: "c1",
+ canonicalSlug: "x",
+ };
+ assert.equal(classifyDigestRole(role, undefined), "mirror-unaligned");
+ assert.equal(roleMustGenerate(classifyDigestRole(role, undefined)), true);
+});
+
+test("exactly one role is free; every other role costs GPU time", () => {
+ const free = DIGEST_PLAN_ROLES.filter((r) => !roleMustGenerate(r));
+ assert.deepEqual(free, ["mirror-aligned"]);
+});
+
+test("audio-hours and the day projection agree with the measured rate", () => {
+ assert.equal(audioHours(3600), 1);
+ // One audio-hour at 90 s/audio-hour is 90 seconds of wall clock.
+ assert.equal(sweepDays(3600, 90), 90 / 86400);
+ // The corpus figure, as a regression pin on the headline: 77,298 audio-hours
+ // at the measured rate is ~80 days, which is what the whole plan is about.
+ const days = sweepDays(77_298 * 3600, MEASURED_SECONDS_PER_AUDIO_HOUR);
+ assert.ok(days > 79 && days < 82, `expected ~80 sweep days, got ${days}`);
+});
+
+// A saving counted in videos is the mistake this module exists to prevent, so
+// pin the arithmetic that makes it visible: many short mirrors are worth far
+// less than one long VOD.
+test("audio-hours, not video count, is what a saving is measured in", () => {
+ const fourThousandShorts = 4_000 * 120; // 4,000 two-minute mirrors
+ const oneVodChannel = 8_331 * 3600; // HasanAbiVODs3
+ assert.ok(
+ audioHours(oneVodChannel) > audioHours(fourThousandShorts) * 60,
+ "one VOD channel outweighs thousands of short mirrors",
+ );
+});
diff --git a/common/controller/digestPlan.ts b/common/controller/digestPlan.ts
@@ -0,0 +1,322 @@
+// What the digest backfill actually costs, measured in AUDIO-HOURS.
+//
+// A saving counted in videos is close to meaningless here: the corpus is 77k
+// videos but 77k audio-HOURS, and the two are not proportional per channel —
+// mirrors skew long (Hasan VODs, Quartering re-uploads), shorts skew numerous.
+// The sweep's wall clock is (audio-hours × seconds-per-audio-hour), so a plan
+// that moves 4,000 videos and 200 audio-hours has moved nothing. Everything
+// here is therefore denominated in seconds of audio.
+//
+// Two consumers, deliberately one implementation:
+// - `bin/digest-plan.ts`, to price the sweep before committing GPU-weeks to it
+// and to prove that a cheapness lever (a duplicate threshold change, a batch
+// of confirmed clusters) actually moved the number;
+// - the sweep orchestrator, which needs the same remaining-audio-hours figure
+// as an ETA denominator and the same per-channel ordering as a work queue.
+//
+// Source of truth is the `statsByPath` LMDB sub-DB — build:stats' output, keyed
+// [channelSlug, videoDir], which is the batch runner's own enumeration unit and
+// carries duration + hasTranscript in one scan. It can lag the corpus; the
+// schema version is checked and a mismatch is reported rather than swallowed.
+
+import { open } from "lmdb";
+import path from "node:path";
+import type { Paths } from "../lib/paths";
+import type { VideoStat } from "../lib/stats";
+import { STATS_SCHEMA_VERSION } from "../lib/stats";
+import { isSectionFresh } from "../lib/digest";
+import { loadDigest } from "../lib/digest-server";
+import {
+ buildDigestClusterPlan,
+ type DigestClusterPlan,
+ type DigestClusterRole,
+} from "./digestSharing";
+import { resolveDigestTarget, type DigestLaneChoice } from "./digestTarget";
+
+// The measured cost of the local lane, from the 102-video validation run on the
+// real corpus (NOT the bake-off's 27 s, which was projected from long videos on
+// an idle box — see plans/FACTS.md). Overridable, because it is the one number
+// here that is an estimate rather than a measurement of this corpus.
+export const MEASURED_SECONDS_PER_AUDIO_HOUR = 90;
+
+// Why a video is or is not this sweep's work. The mirror split is the point: a
+// mirror is only free if its cues actually ALIGN with its canonical member's.
+// Detection measures that and records it per ref, and the sharing pass refuses
+// to place a digest without it — so counting all mirrors as free overstates the
+// saving by however many the alignment gate will later reject. On the current
+// corpus that is nearly half of them.
+export type DigestPlanRole =
+ | "canonical" // owns its cluster's digest: must generate
+ | "mirror-aligned" // receives a share: free
+ | "mirror-unaligned" // share will be refused: must generate after all
+ | "unclustered"; // in no cluster: must generate
+
+export const DIGEST_PLAN_ROLES: readonly DigestPlanRole[] = [
+ "canonical",
+ "mirror-aligned",
+ "mirror-unaligned",
+ "unclustered",
+];
+
+// A role must generate unless the digest arrives by sharing.
+export function roleMustGenerate(role: DigestPlanRole): boolean {
+ return role !== "mirror-aligned";
+}
+
+export type DigestPlanTotals = {
+ videos: number;
+ audioSeconds: number;
+};
+
+function emptyTotals(): DigestPlanTotals {
+ return { videos: 0, audioSeconds: 0 };
+}
+
+function addTo(t: DigestPlanTotals, seconds: number): void {
+ t.videos++;
+ t.audioSeconds += seconds;
+}
+
+export type DigestChannelPlan = {
+ channelSlug: string;
+ // Videos with no fresh digest at the current identity, split by role.
+ remaining: Record<DigestPlanRole, DigestPlanTotals>;
+ // Already digested at the current identity — the sweep skips these.
+ fresh: DigestPlanTotals;
+ // Indexed but not digestable: no transcript, or no usable duration.
+ ineligible: number;
+ // Convenience rollups over `remaining`.
+ generateSeconds: number; // what this channel costs the GPU
+ sharedSeconds: number; // what cluster sharing takes off the bill
+};
+
+export type DigestSweepPlan = {
+ // Ordered by generateSeconds descending — the sweep's work queue. Channels
+ // with nothing to do are still present, with zeroed totals, so a caller can
+ // tell "done" apart from "not in the corpus".
+ channels: DigestChannelPlan[];
+ remaining: Record<DigestPlanRole, DigestPlanTotals>;
+ fresh: DigestPlanTotals;
+ ineligible: number;
+ generateSeconds: number;
+ sharedSeconds: number;
+ clusters: number;
+ clusterMembersMapped: number;
+ // True when build:stats' cache is not at the version this code expects, i.e.
+ // every number below may be stale. Reported, never silently tolerated.
+ statsSchemaStale: boolean;
+ // Set when freshness was NOT consulted, so `fresh` is 0 by construction and
+ // `remaining` is the cost of a sweep from scratch.
+ freshnessChecked: boolean;
+};
+
+export type BuildDigestSweepPlanOptions = {
+ paths: Paths;
+ lane?: DigestLaneChoice;
+ // Off makes the scan pure-LMDB and near-instant, at the cost of reporting the
+ // from-scratch cost rather than the remaining one. At 0.13% coverage the two
+ // are nearly the same number; mid-sweep they are not.
+ checkFreshness?: boolean;
+ // Restrict to these channels (the orchestrator uses it to re-price one).
+ channelSlugs?: string[];
+ clusterPlan?: DigestClusterPlan;
+ onLog?: (msg: string) => void;
+};
+
+function emptyRoleTotals(): Record<DigestPlanRole, DigestPlanTotals> {
+ return {
+ canonical: emptyTotals(),
+ "mirror-aligned": emptyTotals(),
+ "mirror-unaligned": emptyTotals(),
+ unclustered: emptyTotals(),
+ };
+}
+
+function mergeInto(
+ into: Record<DigestPlanRole, DigestPlanTotals>,
+ from: Record<DigestPlanRole, DigestPlanTotals>,
+): void {
+ for (const role of DIGEST_PLAN_ROLES) {
+ into[role].videos += from[role].videos;
+ into[role].audioSeconds += from[role].audioSeconds;
+ }
+}
+
+// The whole role decision, as one pure function so it can be asserted on
+// without an LMDB corpus behind it.
+//
+// `aligned` comes from the report's per-ref measurement and is deliberately
+// tri-state at the call site: absent means NOT MEASURED, which the sharing pass
+// treats as not aligned. So anything other than an explicit `true` lands in
+// mirror-unaligned and the plan under-promises rather than over-promises.
+export function classifyDigestRole(
+ clusterRole: DigestClusterRole | undefined,
+ aligned: boolean | undefined,
+): DigestPlanRole {
+ if (clusterRole?.kind === "canonical") return "canonical";
+ if (clusterRole?.kind === "mirror") {
+ return aligned === true ? "mirror-aligned" : "mirror-unaligned";
+ }
+ return "unclustered";
+}
+
+export async function buildDigestSweepPlan(
+ opts: BuildDigestSweepPlanOptions,
+): Promise<DigestSweepPlan> {
+ const log = opts.onLog ?? (() => {});
+ const checkFreshness = opts.checkFreshness !== false;
+ const clusterPlan =
+ opts.clusterPlan ?? (await buildDigestClusterPlan(opts.paths));
+
+ // `aligned` lives on the report's refs, not on the derived plan, so read it
+ // straight from the report rather than widening DigestClusterRole for one
+ // consumer.
+ const alignedBySlug = await readAlignmentBySlug(opts.paths);
+
+ const root = open({
+ path: opts.paths.lmdbPath,
+ maxDbs: 12,
+ compression: true,
+ });
+ const statsByPath = root.openDB<
+ { metaMs: number; stat: VideoStat },
+ [string, string]
+ >({ name: "statsByPath", encoding: "msgpack" });
+ const meta = root.openDB<unknown, string>({
+ name: "statsMeta",
+ encoding: "msgpack",
+ });
+ const statsSchemaStale =
+ (meta.get("schema") as number | undefined) !== STATS_SCHEMA_VERSION;
+ if (statsSchemaStale) {
+ log(
+ `Warning: stats cache schema ${meta.get("schema") ?? "<none>"} != ${STATS_SCHEMA_VERSION}; run build:stats for accurate audio-hours.`,
+ );
+ }
+
+ const wanted = opts.channelSlugs ? new Set(opts.channelSlugs) : null;
+
+ // Group by channel first: the freshness target is resolved ONCE per channel
+ // (it reads settings + the channel-context note), and resolving it per video
+ // would dominate the scan.
+ type Row = { videoDir: string; stat: VideoStat };
+ const byChannel = new Map<string, Row[]>();
+ for (const { key, value } of statsByPath.getRange()) {
+ const [channelSlug, videoDir] = key as [string, string];
+ if (wanted && !wanted.has(channelSlug)) continue;
+ let rows = byChannel.get(channelSlug);
+ if (!rows) byChannel.set(channelSlug, (rows = []));
+ rows.push({ videoDir, stat: value.stat });
+ }
+
+ const channels: DigestChannelPlan[] = [];
+ for (const [channelSlug, rows] of byChannel) {
+ const entry: DigestChannelPlan = {
+ channelSlug,
+ remaining: emptyRoleTotals(),
+ fresh: emptyTotals(),
+ ineligible: 0,
+ generateSeconds: 0,
+ sharedSeconds: 0,
+ };
+
+ const resolved = checkFreshness
+ ? await resolveDigestTarget({
+ paths: opts.paths,
+ channelSlug,
+ lane: opts.lane,
+ })
+ : null;
+ const dataDir = path.join(opts.paths.channelsDir, channelSlug, "data");
+
+ for (const { videoDir, stat } of rows) {
+ // Exactly the batch runner's eligibility: a video with no transcript is a
+ // transcription problem, not a digest one, and it is not this sweep's work.
+ if (!stat.hasTranscript || !(stat.duration > 0)) {
+ entry.ineligible++;
+ continue;
+ }
+
+ if (resolved) {
+ const record = await loadDigest(path.join(dataDir, videoDir));
+ const allFresh = resolved.sections.every((section) =>
+ isSectionFresh(record, section, resolved.target),
+ );
+ if (allFresh) {
+ addTo(entry.fresh, stat.duration);
+ continue;
+ }
+ }
+
+ const planRole = classifyDigestRole(
+ clusterPlan.bySlug.get(stat.slug),
+ alignedBySlug.get(stat.slug),
+ );
+ addTo(entry.remaining[planRole], stat.duration);
+ }
+
+ for (const role of DIGEST_PLAN_ROLES) {
+ if (roleMustGenerate(role))
+ entry.generateSeconds += entry.remaining[role].audioSeconds;
+ else entry.sharedSeconds += entry.remaining[role].audioSeconds;
+ }
+ channels.push(entry);
+ }
+
+ channels.sort(
+ (a, b) =>
+ b.generateSeconds - a.generateSeconds ||
+ a.channelSlug.localeCompare(b.channelSlug),
+ );
+
+ const totals: DigestSweepPlan = {
+ channels,
+ remaining: emptyRoleTotals(),
+ fresh: emptyTotals(),
+ ineligible: 0,
+ generateSeconds: 0,
+ sharedSeconds: 0,
+ clusters: clusterPlan.clusters,
+ clusterMembersMapped: clusterPlan.bySlug.size,
+ statsSchemaStale,
+ freshnessChecked: checkFreshness,
+ };
+ for (const c of channels) {
+ mergeInto(totals.remaining, c.remaining);
+ totals.fresh.videos += c.fresh.videos;
+ totals.fresh.audioSeconds += c.fresh.audioSeconds;
+ totals.ineligible += c.ineligible;
+ totals.generateSeconds += c.generateSeconds;
+ totals.sharedSeconds += c.sharedSeconds;
+ }
+ return totals;
+}
+
+// Per-ref `aligned` from the last detection run, keyed by slug. Kept separate
+// from DigestClusterPlan because only the pricing path needs it.
+async function readAlignmentBySlug(
+ paths: Paths,
+): Promise<Map<string, boolean>> {
+ const { readDuplicateReport } = await import("./duplicateShorts");
+ const report = await readDuplicateReport(paths);
+ const map = new Map<string, boolean>();
+ if (!report) return map;
+ for (const cluster of report.clusters) {
+ for (const ref of cluster.videoRefs) {
+ if (ref.aligned === true) map.set(ref.slug, true);
+ }
+ }
+ return map;
+}
+
+export function audioHours(seconds: number): number {
+ return seconds / 3600;
+}
+
+// The projection the whole plan exists to produce.
+export function sweepDays(
+ audioSeconds: number,
+ secondsPerAudioHour: number = MEASURED_SECONDS_PER_AUDIO_HOUR,
+): number {
+ return (audioHours(audioSeconds) * secondsPerAudioHour) / 86400;
+}
diff --git a/common/controller/digestSharing.test.ts b/common/controller/digestSharing.test.ts
@@ -0,0 +1,186 @@
+import { test } from "node:test";
+import assert from "node:assert/strict";
+import { mkdir, mkdtemp, rm, writeFile } from "node:fs/promises";
+import { tmpdir } from "node:os";
+import path from "node:path";
+import type { Paths } from "../lib/paths";
+import {
+ DUPLICATES_FILENAME,
+ DUPLICATE_REPORT_VERSION,
+ type DuplicateCluster,
+ type DuplicateVideoRef,
+} from "../lib/duplicates";
+import {
+ buildDigestClusterPlan,
+ dirBySlugForCluster,
+ planSlugForDir,
+} from "./digestSharing";
+
+// Run with:
+// pnpm --filter yt-dlp-transcript-common exec tsx --test common/controller/digestSharing.test.ts
+
+async function withPaths(fn: (paths: Paths) => Promise<void>): Promise<void> {
+ const dir = await mkdtemp(path.join(tmpdir(), "ttb-digest-sharing-"));
+ const paths = {
+ transcriptsDir: dir,
+ channelsDir: path.join(dir, "channels"),
+ } as Paths;
+ await mkdir(paths.transcriptsDir, { recursive: true });
+ try {
+ await fn(paths);
+ } finally {
+ await rm(dir, { recursive: true, force: true });
+ }
+}
+
+function ref(
+ channelSlug: string,
+ id: string,
+ extra: Partial<DuplicateVideoRef> = {},
+): DuplicateVideoRef {
+ return {
+ slug: `${channelSlug}/${id}`,
+ channelSlug,
+ channel: channelSlug,
+ platform: "youtube",
+ id,
+ title: "t",
+ duration: 600,
+ uploadDate: "20240101",
+ hasTranscript: true,
+ ...extra,
+ };
+}
+
+async function writeReport(
+ paths: Paths,
+ clusters: DuplicateCluster[],
+): Promise<void> {
+ await writeFile(
+ path.join(paths.transcriptsDir, DUPLICATES_FILENAME),
+ JSON.stringify({
+ version: DUPLICATE_REPORT_VERSION,
+ generatedAt: new Date(0).toISOString(),
+ runConfig: {
+ thresholdSeconds: null,
+ durationToleranceSeconds: 2,
+ nearThreshold: 0.35,
+ containmentThreshold: 0.8,
+ shingleSize: 5,
+ },
+ totals: { videosScanned: 2, clusters: clusters.length, videosInClusters: 2 },
+ clusters,
+ }),
+ );
+}
+
+function cluster(
+ clusterId: string,
+ videoRefs: DuplicateVideoRef[],
+ extra: Partial<DuplicateCluster> = {},
+): DuplicateCluster {
+ return {
+ clusterId,
+ matchKind: "transcript-near",
+ score: 0.9,
+ contained: false,
+ durationBucket: 600,
+ crossPlatform: true,
+ crossChannel: true,
+ videoRefs,
+ canonicalSlug: videoRefs[0].slug,
+ ...extra,
+ };
+}
+
+// The bug this whole mapping exists for: a slug is `${channelSlug}/${id}` while
+// the batch runner enumerates DIRECTORIES, and those disagree for 14.5% of the
+// real corpus (every Rumble re-upload). Before the mapping, the lookup below
+// missed and every such mirror was regenerated instead of shared.
+test("plan maps a directory name to its slug when the two differ", async () => {
+ await withPaths(async (paths) => {
+ await writeReport(paths, [
+ cluster("c1", [
+ ref("chan-yt", "canon1"),
+ // The Rumble shape: on-disk dir `v1007ay`, metadata id `vxe1ae`.
+ ref("chan-rumble", "vxe1ae", { videoDir: "v1007ay" }),
+ ]),
+ ]);
+ const plan = await buildDigestClusterPlan(paths, { overrides: null });
+
+ assert.equal(
+ planSlugForDir(plan, "chan-rumble", "v1007ay"),
+ "chan-rumble/vxe1ae",
+ "the directory name must resolve to the metadata-id slug",
+ );
+ assert.equal(
+ plan.bySlug.get(planSlugForDir(plan, "chan-rumble", "v1007ay"))?.kind,
+ "mirror",
+ "and that slug must then hit the plan as a mirror",
+ );
+ assert.equal(
+ plan.dirBySlug.get("chan-rumble/vxe1ae"),
+ "v1007ay",
+ "the reverse direction is what lets the sharing pass open the mirror's files",
+ );
+ });
+});
+
+test("a member whose dir equals its id is absent from the mapping", async () => {
+ await withPaths(async (paths) => {
+ await writeReport(paths, [
+ cluster("c1", [ref("chan-a", "same1"), ref("chan-b", "same2")]),
+ ]);
+ const plan = await buildDigestClusterPlan(paths, { overrides: null });
+ assert.equal(plan.dirBySlug.size, 0);
+ assert.equal(plan.slugByDir.size, 0);
+ // The fallback is the old behaviour, and it is the correct one here.
+ assert.equal(planSlugForDir(plan, "chan-a", "same1"), "chan-a/same1");
+ });
+});
+
+// A report written before `videoDir` existed must not regress: it keeps the
+// pre-mapping behaviour rather than resolving to nothing.
+test("an old report with no videoDir falls back to the id", async () => {
+ await withPaths(async (paths) => {
+ await writeReport(paths, [
+ cluster("c1", [ref("chan-a", "aaa"), ref("chan-b", "bbb")]),
+ ]);
+ const plan = await buildDigestClusterPlan(paths, { overrides: null });
+ assert.equal(planSlugForDir(plan, "chan-b", "bbb"), "chan-b/bbb");
+ assert.equal(planSlugForDir(null, "chan-b", "bbb"), "chan-b/bbb");
+ });
+});
+
+test("dirBySlugForCluster carries only the members that differ", () => {
+ const c = cluster("c1", [
+ ref("chan-yt", "canon1"),
+ ref("chan-rumble", "vxe1ae", { videoDir: "v1007ay" }),
+ ref("chan-odysee", "plain", { videoDir: "plain" }),
+ ]);
+ const map = dirBySlugForCluster(c);
+ assert.deepEqual([...map], [["chan-rumble/vxe1ae", "v1007ay"]]);
+});
+
+// The directory mapping is recorded for EVERY cluster, including ones that
+// share nothing — their members are still generated, and a caller holding the
+// plan still needs a correct path for them.
+test("an unshareable cluster still contributes its directory mapping", async () => {
+ await withPaths(async (paths) => {
+ await writeReport(paths, [
+ cluster(
+ "c1",
+ [
+ ref("chan-yt", "canon1"),
+ ref("chan-rumble", "vxe1ae", { videoDir: "v1007ay" }),
+ ],
+ // needsReview and unconfirmed → clusterMaySharePartial fails closed.
+ { needsReview: true },
+ ),
+ ]);
+ const plan = await buildDigestClusterPlan(paths, { overrides: null });
+ assert.equal(plan.bySlug.size, 0, "it shares nothing");
+ assert.equal(plan.independentClusters, 1);
+ assert.equal(plan.dirBySlug.get("chan-rumble/vxe1ae"), "v1007ay");
+ });
+});
diff --git a/common/controller/digestSharing.ts b/common/controller/digestSharing.ts
@@ -0,0 +1,322 @@
+// Where duplicate detection meets the digest layer: generate ONCE per cluster,
+// share to the mirrors that are safe to share to.
+//
+// This is worth ~11% of the sweep on its own — 8,204 exact-title redundancies
+// across the corpus, 5,863 of them Quartering YouTube↔Rumble mirror pairs — and
+// at weeks of wall-clock per pass, not re-generating those is real time.
+//
+// Two rules keep the sharing honest, and both are the difference between a
+// correct optimisation and a plausible-looking wrong one:
+//
+// 1. A `contained` cluster NEVER shares. Containment means one member is a CLIP
+// of a longer video; the longer video's chapters describe material the clip
+// does not contain.
+// 2. A mirror only receives a digest when its cue TIMINGS align with the
+// canonical member's. Content similarity says nothing about timing — a
+// mirror with a longer intro matches on text at shifted times — so a shared
+// digest would place every chapter wrong while looking perfectly fine. The
+// gate measures the offset at several anchors and requires near-zero.
+
+import path from "node:path";
+import type { Paths } from "../lib/paths";
+import {
+ DEFAULT_ALIGNMENT_TOLERANCE_SECONDS,
+ clusterMaySharePartial,
+ measureAlignment,
+ resolveCanonicalSlug,
+ type AlignmentResult,
+ type DuplicateCluster,
+ type DuplicateOverrides,
+} from "../lib/duplicates";
+import { CUES_JSON_FILENAME } from "../lib/videoStatus";
+import type { Cue } from "../lib/vtt";
+import { loadDigest, writeSharedDigest } from "../lib/digest-server";
+import type { DigestRecord } from "../lib/digest";
+import { readNormalizedTranscript } from "./normalizeTranscript";
+import { readDuplicateOverrides, readDuplicateReport } from "./duplicateShorts";
+
+// What a video's cluster membership means for the batch.
+export type DigestClusterRole =
+ // Generate for this video: it owns its cluster's derived work.
+ | { kind: "canonical"; clusterId: string; mirrors: string[] }
+ // Do NOT generate: the canonical member owns it and will share it here.
+ | { kind: "mirror"; clusterId: string; canonicalSlug: string };
+
+export type DigestClusterPlan = {
+ // `${channelSlug}/${id}` → role. Videos absent from the map are in no cluster
+ // and are generated normally, which is the overwhelming majority.
+ bySlug: Map<string, DigestClusterRole>;
+ // The two directions of the id ↔ on-disk-directory mapping, populated ONLY for
+ // cluster members whose directory name is not their metadata id.
+ //
+ // These exist because a slug is `${channelSlug}/${id}` and is not a path,
+ // while every caller that reaches this plan starts from one side or the other:
+ // the batch runner enumerates DIRECTORIES off disk, and the sharing pass has
+ // SLUGS from the report and has to open their files. 14.5% of the corpus has
+ // dir !== id, and it is not spread evenly — it is 100% of `the-quartering-
+ // rumble` (7,870 videos), i.e. exactly the mirror set cluster-sharing was
+ // written to exploit. Keying either lookup with the wrong half misses silently
+ // and simply regenerates the mirror, which is why this went unnoticed.
+ //
+ // Empty for a report written before `DuplicateVideoRef.videoDir` existed;
+ // lookups fall back to treating the id as the directory, the old behaviour.
+ slugByDir: Map<string, string>; // `${channelSlug}/${videoDir}` → slug
+ dirBySlug: Map<string, string>; // slug → videoDir (bare name, not a path)
+ // Clusters that exist but share nothing (contained, or marked not-a-duplicate).
+ // Their members are absent from bySlug and are each generated independently —
+ // a clip is a different artifact and deserves its own digest.
+ independentClusters: number;
+ clusters: number;
+};
+
+export function emptyDigestClusterPlan(): DigestClusterPlan {
+ return {
+ bySlug: new Map(),
+ slugByDir: new Map(),
+ dirBySlug: new Map(),
+ independentClusters: 0,
+ clusters: 0,
+ };
+}
+
+// The batch runner enumerates on-disk directories; the plan is keyed by slug.
+// Falls back to the directory name as the id, which is correct for the ~85.5%
+// where they agree and is what the code did before the mapping existed.
+export function planSlugForDir(
+ plan: DigestClusterPlan | null | undefined,
+ channelSlug: string,
+ videoDir: string,
+): string {
+ const key = `${channelSlug}/${videoDir}`;
+ return plan?.slugByDir.get(key) ?? key;
+}
+
+// Build the plan from the last detection run. Absent report → an empty plan, so
+// the digest batch works fine before duplicates have ever been detected (it just
+// generates for mirrors twice, which is correct, only slower).
+export async function buildDigestClusterPlan(
+ paths: Paths,
+ opts: { overrides?: DuplicateOverrides | null } = {},
+): Promise<DigestClusterPlan> {
+ const report = await readDuplicateReport(paths);
+ if (!report) return emptyDigestClusterPlan();
+ const overrides = opts.overrides ?? (await readDuplicateOverrides(paths));
+ const plan = emptyDigestClusterPlan();
+ plan.clusters = report.clusters.length;
+
+ for (const cluster of report.clusters) {
+ // Recorded for EVERY cluster, including the ones that share nothing: an
+ // independent cluster's members are still generated, and a caller that
+ // reaches the plan at all deserves a correct path for them.
+ for (const ref of cluster.videoRefs) {
+ if (!ref.videoDir || ref.videoDir === ref.id) continue;
+ plan.dirBySlug.set(ref.slug, ref.videoDir);
+ plan.slugByDir.set(`${ref.channelSlug}/${ref.videoDir}`, ref.slug);
+ }
+ const canonicalSlug = resolveCanonicalSlug(cluster, overrides);
+ // null → a human said "not a duplicate": every member stands alone.
+ // The overrides also carry the `confirmed` flag that is the ONLY thing
+ // letting a needsReview cluster share, so they must be passed here — without
+ // them every confirmed suspect would silently keep generating twice.
+ if (!canonicalSlug || !clusterMaySharePartial(cluster, overrides)) {
+ plan.independentClusters++;
+ continue;
+ }
+ const mirrors = cluster.videoRefs
+ .map((r) => r.slug)
+ .filter((slug) => slug !== canonicalSlug);
+ if (mirrors.length === 0) {
+ plan.independentClusters++;
+ continue;
+ }
+ plan.bySlug.set(canonicalSlug, {
+ kind: "canonical",
+ clusterId: cluster.clusterId,
+ mirrors,
+ });
+ for (const slug of mirrors) {
+ plan.bySlug.set(slug, {
+ kind: "mirror",
+ clusterId: cluster.clusterId,
+ canonicalSlug,
+ });
+ }
+ }
+ return plan;
+}
+
+// A slug is `${channelSlug}/${id}`, and an id is NOT always the directory name —
+// see DigestClusterPlan.slugByDir. `dirBySlug` carries the exceptions; without
+// it this resolves the 14.5% of the corpus with dir !== id to a path that does
+// not exist, which reads as "no transcript" and silently declines to share.
+function videoDirForSlug(
+ paths: Paths,
+ slug: string,
+ dirBySlug?: ReadonlyMap<string, string>,
+): string | null {
+ const at = slug.indexOf("/");
+ if (at <= 0) return null;
+ return path.join(
+ paths.channelsDir,
+ slug.slice(0, at),
+ "data",
+ dirBySlug?.get(slug) ?? slug.slice(at + 1),
+ );
+}
+
+async function readCues(
+ paths: Paths,
+ slug: string,
+ dirBySlug?: ReadonlyMap<string, string>,
+): Promise<Cue[] | null> {
+ const dir = videoDirForSlug(paths, slug, dirBySlug);
+ if (!dir) return null;
+ const t = await readNormalizedTranscript(path.join(dir, CUES_JSON_FILENAME));
+ return t?.cues ?? null;
+}
+
+export type ShareOutcome = {
+ slug: string;
+ status: "shared" | "misaligned" | "no-transcript" | "already-shared" | "failed";
+ alignment?: AlignmentResult;
+ error?: string;
+};
+
+export type ShareDigestOptions = {
+ paths: Paths;
+ clusterId: string;
+ canonicalSlug: string;
+ mirrors: string[];
+ // slug → on-disk directory, for members where the two differ. Take it from
+ // DigestClusterPlan.dirBySlug (or build it from the cluster's own refs).
+ // Omitting it is not an error, but every member with dir !== id will resolve
+ // to a nonexistent path and be reported "no-transcript".
+ dirBySlug?: ReadonlyMap<string, string>;
+ toleranceSeconds?: number;
+ onLog?: (msg: string) => void;
+};
+
+// Copy the canonical member's digest onto each mirror that passes the alignment
+// gate. Returns one outcome per mirror — a misaligned mirror is a normal,
+// expected result, not an error, and it simply stays on the generation worklist.
+export async function shareDigestToCluster(
+ opts: ShareDigestOptions,
+): Promise<ShareOutcome[]> {
+ const log = opts.onLog ?? (() => {});
+ const tolerance =
+ opts.toleranceSeconds ?? DEFAULT_ALIGNMENT_TOLERANCE_SECONDS;
+ const canonicalDir = videoDirForSlug(
+ opts.paths,
+ opts.canonicalSlug,
+ opts.dirBySlug,
+ );
+ if (!canonicalDir) return [];
+ const source: DigestRecord | null = await loadDigest(canonicalDir);
+ if (!source) return [];
+ const canonicalCues = await readCues(
+ opts.paths,
+ opts.canonicalSlug,
+ opts.dirBySlug,
+ );
+ if (!canonicalCues || canonicalCues.length === 0) return [];
+
+ const sharedAt = new Date().toISOString();
+ const outcomes: ShareOutcome[] = [];
+ for (const slug of opts.mirrors) {
+ const dir = videoDirForSlug(opts.paths, slug, opts.dirBySlug);
+ if (!dir) {
+ outcomes.push({ slug, status: "failed", error: "unparseable slug" });
+ continue;
+ }
+ try {
+ const existing = await loadDigest(dir);
+ // Already carrying this canonical member's digest at the same provenance:
+ // nothing to do. Keeps a re-run a genuine no-op.
+ if (
+ existing?.derivedFrom?.slug === opts.canonicalSlug &&
+ existing.promptVersion === source.promptVersion &&
+ existing.contextHash === source.contextHash
+ ) {
+ outcomes.push({ slug, status: "already-shared" });
+ continue;
+ }
+ const cues = await readCues(opts.paths, slug, opts.dirBySlug);
+ if (!cues || cues.length === 0) {
+ outcomes.push({ slug, status: "no-transcript" });
+ continue;
+ }
+ const alignment = measureAlignment(canonicalCues, cues, {
+ toleranceSeconds: tolerance,
+ });
+ if (!alignment.aligned) {
+ log(
+ `Not sharing to ${slug}: ${alignment.reason} (max offset ${
+ Number.isFinite(alignment.maxOffsetSeconds)
+ ? `${alignment.maxOffsetSeconds.toFixed(1)}s`
+ : "n/a"
+ }, ${alignment.matchedAnchors}/${alignment.totalAnchors} anchors matched).`,
+ );
+ outcomes.push({ slug, status: "misaligned", alignment });
+ continue;
+ }
+ await writeSharedDigest(dir, source, {
+ slug: opts.canonicalSlug,
+ clusterId: opts.clusterId,
+ sharedAt,
+ offsetSeconds: Math.round(alignment.maxOffsetSeconds * 100) / 100,
+ });
+ log(
+ `Shared digest ${opts.canonicalSlug} → ${slug} (max offset ${alignment.maxOffsetSeconds.toFixed(2)}s).`,
+ );
+ outcomes.push({ slug, status: "shared", alignment });
+ } catch (err) {
+ outcomes.push({
+ slug,
+ status: "failed",
+ error: (err as Error)?.message ?? String(err),
+ });
+ }
+ }
+ return outcomes;
+}
+
+// Convenience for the /actionable per-cluster action: share from whatever the
+// effective canonical member currently is.
+export async function shareClusterFromCanonical(
+ paths: Paths,
+ cluster: DuplicateCluster,
+ opts: {
+ overrides?: DuplicateOverrides | null;
+ onLog?: (msg: string) => void;
+ } = {},
+): Promise<ShareOutcome[]> {
+ // Load the overrides BEFORE the gate, not after: `confirmed` lives in them and
+ // is what unblocks a needsReview cluster, so testing the gate first would
+ // refuse to share from every cluster a human had just approved.
+ const overrides = opts.overrides ?? (await readDuplicateOverrides(paths));
+ if (!clusterMaySharePartial(cluster, overrides)) return [];
+ const canonicalSlug = resolveCanonicalSlug(cluster, overrides);
+ if (!canonicalSlug) return [];
+ return shareDigestToCluster({
+ paths,
+ clusterId: cluster.clusterId,
+ canonicalSlug,
+ mirrors: cluster.videoRefs
+ .map((r) => r.slug)
+ .filter((slug) => slug !== canonicalSlug),
+ dirBySlug: dirBySlugForCluster(cluster),
+ onLog: opts.onLog,
+ });
+}
+
+// The single-cluster equivalent of DigestClusterPlan.dirBySlug, for callers that
+// hold one cluster rather than a whole plan.
+export function dirBySlugForCluster(
+ cluster: DuplicateCluster,
+): Map<string, string> {
+ const map = new Map<string, string>();
+ for (const ref of cluster.videoRefs) {
+ if (ref.videoDir && ref.videoDir !== ref.id) map.set(ref.slug, ref.videoDir);
+ }
+ return map;
+}
diff --git a/common/controller/digestSweep.ts b/common/controller/digestSweep.ts
@@ -0,0 +1,411 @@
+// The corpus-wide digest sweep: one launcher for the thing that runs for weeks.
+//
+// Before this, "sweep the corpus" meant clicking Digest on 63 channel pages by
+// hand and losing the whole thing to a server restart. Everything underneath —
+// the per-channel batch, resume-by-re-deriving-from-disk, the pause, the GPU
+// yield — already existed and was already correct. What was missing was
+// something to walk the channels.
+//
+// Three shapes are deliberate:
+//
+// 1. IT LAUNCHES THE EXISTING PER-CHANNEL JOB, in sequence. It does not
+// introduce a batch over videos. digestBatch.ts is explicit that one job
+// per video would be 119,600 jobs against a 100-record registry and a
+// 500-record log, evicting the history of the very run it is recording.
+// 63 sequential channel jobs is the same contract, driven.
+//
+// 2. IT STORES NO CURSOR. Ordering is recomputed from
+// `buildDigestSweepPlan` each pass and eligibility is re-derived from disk
+// inside every batch, so a restart, a newly finished transcription, a
+// human confirming a duplicate cluster, and a settings change that
+// invalidates the corpus are all just visible on the next pass. A frozen
+// work-list would be wrong within hours of a multi-week run starting.
+//
+// 3. IT LOOPS UNTIL THE PLAN IS EMPTY, rather than passing over the channel
+// list once. A batch can legitimately leave work behind — the spend cap,
+// a drain, a video whose transcript landed mid-pass — and a one-pass sweep
+// would report itself finished with the corpus unfinished.
+//
+// Heaviest channel first, by remaining AUDIO-HOURS. Cost is audio, not videos:
+// HasanAbiVODs3 is 8,331 audio-hours and outweighs every duplicate mirror in
+// the corpus combined, so leaving it for last is how a sweep spends 70 days
+// looking nearly done.
+
+import type { Paths } from "../lib/paths";
+import { getPaths } from "../lib/paths";
+import { getSettings, writeSettings } from "../lib/settings";
+import { getRegistry } from "../jobs/registry";
+import { runManagedFunction } from "../jobs/streamCommand";
+import { drainStream } from "../jobs/drainStream";
+import { makeTaskTracker } from "../jobs/taskHooks";
+import { requestChannelSnapshot } from "../jobs/snapshotScheduler";
+import { DIGEST_LOCAL_QUEUE, DIGEST_REMOTE_QUEUE } from "../lib/queueKeys";
+import { countMissingDigests, runDigestBatch } from "./digestBatch";
+import {
+ audioHours,
+ buildDigestSweepPlan,
+ sweepDays,
+ type DigestSweepPlan,
+} from "./digestPlan";
+import type { DigestLaneChoice } from "./digestTarget";
+
+export const DIGEST_SWEEP_KIND = "digest-sweep";
+
+// How long the loop waits before recomputing the plan when a pass did no work.
+// Generous: nothing here is latency-sensitive and recomputing the plan reads
+// every channel's sidecars.
+const IDLE_POLL_MS = 60_000;
+
+// Give up after this many consecutive passes that move no work. Three, not one:
+// a single barren pass is legitimate (every remaining video failed a guard this
+// time round and may not next time), but three in a row is a broken engine or a
+// corpus the sweep cannot make progress on, and both want a human.
+const MAX_BARREN_PASSES = 3;
+
+type SweepLive = {
+ jobId: string;
+ startedAt: number;
+ // The per-channel job the sweep is currently waiting on. Tracked so a stop
+ // can drain THAT too — without it, "stop" means "after the current channel
+ // finishes", and the heaviest channel is 8.7 days long.
+ channelJobId: string | null;
+};
+type SweepSingleton = { live: SweepLive | null };
+
+declare global {
+ // eslint-disable-next-line no-var
+ var __yttDigestSweep__: SweepSingleton | undefined;
+}
+
+function getSingleton(): SweepSingleton {
+ if (!globalThis.__yttDigestSweep__) {
+ globalThis.__yttDigestSweep__ = { live: null };
+ }
+ return globalThis.__yttDigestSweep__;
+}
+
+export function getDigestSweepJobId(): string | null {
+ const live = getSingleton().live;
+ if (!live) return null;
+ return getRegistry().get(live.jobId)?.status === "running"
+ ? live.jobId
+ : null;
+}
+
+export type DigestSweepOptions = {
+ paths?: Paths;
+ lane?: DigestLaneChoice;
+ // Restrict the sweep to these channels. Absent = the whole corpus.
+ channelSlugs?: string[];
+};
+
+// Run ONE channel through the existing per-channel batch, as its own managed
+// job — same kind, same queue and same progress metric a hand-clicked run
+// produces, so the sweep is inspectable with the tools that already exist
+// rather than being an opaque mega-job.
+export async function runDigestChannelJob(opts: {
+ paths: Paths;
+ channelSlug: string;
+ lane: DigestLaneChoice;
+ background?: boolean;
+ // Called with the job id as soon as it exists, so a caller can drain this
+ // specific job rather than only the thing that launched it.
+ onStarted?: (jobId: string) => void;
+ onDone?: () => void;
+}): Promise<{ ok: boolean; error?: string }> {
+ const { paths, channelSlug, lane } = opts;
+ const settings = getSettings();
+ const remote = lane === "remote";
+ const kind = remote ? "digest-channel-remote" : "digest-channel-local";
+
+ const result = await runManagedFunction({
+ kind,
+ queueKey: remote ? DIGEST_REMOTE_QUEUE : DIGEST_LOCAL_QUEUE,
+ paths,
+ channelSlug,
+ background: opts.background,
+ spec: { kind, slug: channelSlug, params: { lane } },
+ fn: async (onLog, signal, setProgress, ctx) => {
+ const missing = await countMissingDigests(
+ paths,
+ channelSlug,
+ undefined,
+ lane,
+ );
+ const batch = await runDigestBatch({
+ channelSlug,
+ paths,
+ lane,
+ ...(remote
+ ? { minDurationSeconds: settings.digest.longTailSeconds }
+ : {}),
+ setProgress,
+ // The bar measures THIS run, from zero. Seeding it with the count of
+ // digests already on disk was the old shape and it could not represent
+ // a regeneration, where the file count never moves.
+ progressBaseline: 0,
+ progressTarget: missing,
+ onLog,
+ signal,
+ drainSignal: ctx.drainSignal,
+ tracker: makeTaskTracker(ctx, onLog),
+ });
+ onLog(
+ `Digest batch: ${batch.succeeded} generated, ${batch.fresh} already current, ` +
+ `${batch.shared} shared to mirrors, ${batch.misaligned} mirror(s) refused by the alignment gate, ` +
+ `${batch.skipped} skipped, ${batch.failed} failed; ` +
+ `${batch.engineCalls} model call(s), ${batch.warnings} warning(s)` +
+ (batch.costUsd > 0 ? `, $${batch.costUsd.toFixed(4)}` : "") +
+ (batch.spendCapped ? " (stopped at the spend cap)" : "") +
+ ".",
+ );
+ requestChannelSnapshot(paths, channelSlug);
+ opts.onDone?.();
+ },
+ });
+ if (!result.ok) return { ok: false, error: result.error };
+ opts.onStarted?.(result.jobId);
+ // Wait for the channel to finish before returning: the sweep is sequential by
+ // design (one GPU), and the queue would serialize these anyway — awaiting
+ // makes that explicit and lets the loop re-plan against real results.
+ await drainStream(result.stream);
+ return { ok: true };
+}
+
+async function runSweepLoop(
+ paths: Paths,
+ lane: DigestLaneChoice,
+ channelSlugs: string[] | undefined,
+ live: SweepLive,
+ onLog: (msg: string) => void,
+ signal: AbortSignal,
+ drainSignal: AbortSignal,
+): Promise<void> {
+ let pass = 0;
+ // Consecutive passes that left the remaining work UNCHANGED. This is the real
+ // termination guard, and it has to measure the plan rather than whether jobs
+ // started: an unreachable engine makes every channel job start normally and
+ // then throw inside its body, which looks exactly like progress from out here.
+ // Without this the loop would spin through all 63 channels as fast as they can
+ // fail, forever, evicting the job registry with its own wreckage.
+ let barrenPasses = 0;
+ let lastGenerateSeconds: number | null = null;
+ for (;;) {
+ if (signal.aborted || drainSignal.aborted) return;
+ // The operator turned the sweep off: stop cleanly rather than being
+ // cancelled, so the job ends "done" and the queue is released.
+ if (!getSettings().digest.sweepEnabled) {
+ onLog("Sweep disarmed in settings — stopping.");
+ return;
+ }
+
+ pass++;
+ const plan: DigestSweepPlan = await buildDigestSweepPlan({
+ paths,
+ lane,
+ channelSlugs,
+ onLog,
+ });
+ const work = plan.channels.filter((c) => c.generateSeconds > 0);
+ if (work.length === 0) {
+ onLog(
+ `Pass ${pass}: nothing left to generate — ${plan.fresh.videos.toLocaleString()} video(s) already digested at the current identity.`,
+ );
+ return;
+ }
+ onLog(
+ `Pass ${pass}: ${work.length} channel(s), ` +
+ `${audioHours(plan.generateSeconds).toFixed(0)} audio-hours to generate ` +
+ `(~${sweepDays(plan.generateSeconds).toFixed(1)} days at the measured rate), ` +
+ `${audioHours(plan.sharedSeconds).toFixed(0)} audio-hours covered by cluster sharing.`,
+ );
+
+ let didWork = false;
+ for (const channel of work) {
+ if (signal.aborted || drainSignal.aborted) return;
+ if (!getSettings().digest.sweepEnabled) {
+ onLog("Sweep disarmed in settings — stopping.");
+ return;
+ }
+ onLog(
+ `→ ${channel.channelSlug}: ${audioHours(channel.generateSeconds).toFixed(0)} audio-hours ` +
+ `(~${sweepDays(channel.generateSeconds).toFixed(1)} days).`,
+ );
+ const outcome = await runDigestChannelJob({
+ paths,
+ channelSlug: channel.channelSlug,
+ lane,
+ // Behind anything an operator clicks by hand: a sweep is weeks long and
+ // must never make a deliberate single-channel run wait for it.
+ background: true,
+ onStarted: (jobId) => {
+ live.channelJobId = jobId;
+ },
+ });
+ live.channelJobId = null;
+ if (!outcome.ok) {
+ // One channel failing to START is not the sweep failing.
+ onLog(`!! ${channel.channelSlug}: ${outcome.error ?? "failed to start"}`);
+ continue;
+ }
+ didWork = true;
+ }
+
+ // Did the pass actually MOVE anything? Compared against the previous pass's
+ // remaining work, because that is the only measure a failing engine cannot
+ // fake. Some tolerance: a pass that shifts less than a minute of audio is
+ // noise, not progress.
+ const moved =
+ lastGenerateSeconds === null ||
+ lastGenerateSeconds - plan.generateSeconds > 60;
+ barrenPasses = moved ? 0 : barrenPasses + 1;
+ lastGenerateSeconds = plan.generateSeconds;
+
+ if (barrenPasses >= MAX_BARREN_PASSES) {
+ onLog(
+ `Pass ${pass} made no progress for ${MAX_BARREN_PASSES} passes in a row ` +
+ `(${audioHours(plan.generateSeconds).toFixed(0)} audio-hours still outstanding). ` +
+ `Stopping rather than spinning — check the engine and the job logs, then restart the sweep.`,
+ );
+ return;
+ }
+
+ if (!didWork || !moved) {
+ // Nothing started, or nothing moved. Back off before re-planning: a tight
+ // retry against a down engine helps nobody and buries its own diagnosis.
+ onLog(
+ `Pass ${pass} made no progress; waiting ${IDLE_POLL_MS / 1000}s before re-planning.`,
+ );
+ await sleep(IDLE_POLL_MS, signal);
+ }
+ }
+}
+
+function sleep(ms: number, signal: AbortSignal): Promise<void> {
+ return new Promise((resolve) => {
+ const t = setTimeout(resolve, ms);
+ signal.addEventListener(
+ "abort",
+ () => {
+ clearTimeout(t);
+ resolve();
+ },
+ { once: true },
+ );
+ });
+}
+
+// Arm and start the corpus-wide sweep. Persists `sweepEnabled` so a restart
+// resumes it (see resumeDigestSweepIfEnabled).
+export async function startDigestSweep(
+ opts: DigestSweepOptions = {},
+): Promise<string | null> {
+ const paths = opts.paths ?? getPaths();
+ const running = getDigestSweepJobId();
+ if (running) return running;
+
+ const settings = getSettings();
+ // The SCOPE is persisted with the flag, not just the flag. The boot hook
+ // re-launches from settings alone, so arming without recording the scope
+ // would resurrect a deliberately-bounded run as a corpus-wide one.
+ const scope = opts.channelSlugs ?? settings.digest.sweepChannels;
+ if (
+ !settings.digest.sweepEnabled ||
+ settings.digest.sweepChannels.join("\u0000") !== scope.join("\u0000")
+ ) {
+ // AWAITED, and it matters: the loop reads `sweepEnabled` at the top of its
+ // very first pass, so an un-awaited write races it and the sweep quits
+ // immediately with "disarmed in settings" — refusing to start at all.
+ await writeSettings({
+ ...settings,
+ digest: {
+ ...settings.digest,
+ sweepEnabled: true,
+ sweepChannels: scope,
+ },
+ });
+ }
+
+ const live: SweepLive = {
+ jobId: "",
+ startedAt: Date.now(),
+ channelJobId: null,
+ };
+ const lane = opts.lane ?? "local";
+ const result = await runManagedFunction({
+ kind: DIGEST_SWEEP_KIND,
+ // Empty key: the orchestrator itself does no work, it waits on per-channel
+ // jobs that serialize on the digest queue. Putting it ON that queue would
+ // deadlock — it would hold the only slot while waiting for a job that needs
+ // the same slot.
+ queueKey: "",
+ paths,
+ fn: async (onLog, signal, _setProgress, ctx) => {
+ live.jobId = ctx.jobId;
+ try {
+ await runSweepLoop(
+ paths,
+ lane,
+ scope.length > 0 ? scope : undefined,
+ live,
+ onLog,
+ signal,
+ ctx.drainSignal,
+ );
+ } finally {
+ if (getSingleton().live === live) getSingleton().live = null;
+ }
+ },
+ });
+ if (!result.ok) return null;
+ live.jobId = result.jobId;
+ getSingleton().live = live;
+ return result.jobId;
+}
+
+// Boot hook. The auto-transcribe/auto-download runners are restarted at server
+// start the same way; digests were not, so a restart silently ended a sweep
+// that had been running for days and nothing said so.
+export async function resumeDigestSweepIfEnabled(
+ paths: Paths = getPaths(),
+): Promise<void> {
+ if (!getSettings().digest.sweepEnabled) return;
+ await startDigestSweep({ paths });
+}
+
+// Disarm and stop. Persisting the flag is the point: without it a restart would
+// resurrect a sweep the operator had just stopped.
+export async function stopDigestSweep(): Promise<boolean> {
+ const settings = getSettings();
+ // Note the second clause: guarding on `sweepEnabled` alone means a stop on an
+ // already-disarmed sweep skips the write entirely and leaves the SCOPE behind
+ // — which is how a stale scope survives to narrow the next sweep. Measured:
+ // after a stop, settings still read sweepChannels ["teamrcn"].
+ if (
+ settings.digest.sweepEnabled ||
+ settings.digest.sweepChannels.length > 0
+ ) {
+ // Awaited so a restart cannot resurrect a sweep the operator just stopped.
+ await writeSettings({
+ ...settings,
+ digest: {
+ ...settings.digest,
+ sweepEnabled: false,
+ // Cleared too. A leftover scope is worse than no scope: the next
+ // "Start Digest Sweep" would silently cover only the channels of a
+ // sweep somebody stopped weeks ago, and report itself finished.
+ sweepChannels: [],
+ },
+ });
+ }
+ const live = getSingleton().live;
+ if (!live) return false;
+ // Drain BOTH: the orchestrator, and the per-channel job it is waiting on.
+ // Draining only the orchestrator would mean "stop" waits for the current
+ // channel to finish — up to 8.7 days on the heaviest one — while the button
+ // sat there looking wedged. A drain is still graceful: the video in flight
+ // completes, the batch stops taking new ones, and nothing part-generated is
+ // thrown away.
+ if (live.channelJobId) getRegistry().requestDrain(live.channelJobId);
+ return getRegistry().requestDrain(live.jobId);
+}
diff --git a/common/controller/digestTarget.ts b/common/controller/digestTarget.ts
@@ -0,0 +1,95 @@
+// The ONE place a digest freshness target is derived.
+//
+// Three callers need to answer "would we regenerate this section?" — the batch
+// runner (to skip), countMissingDigests (to size a progress bar), and the
+// channel snapshot (to fill the noDigest bucket). They MUST agree, and before
+// this module they did not: the snapshot asked only "does an ai-digest.json with
+// items exist?", so after any config change the Digest stage read "All digested"
+// while the batch reported the whole channel as stale. Two counters that
+// disagree by an entire channel is how a sweep reports "nothing to do" on a
+// corpus that needs redoing.
+//
+// The lane is a PARAMETER, never assumed. countMissingDigests used to resolve
+// settings.localAppId unconditionally, so a metered-lane job's progress target
+// was computed against the local engine's identity — and, once numCtx moved into
+// the recorded identity, mis-derived promptVariant from the local app's context
+// size as well.
+
+import type { Paths } from "../lib/paths";
+import { getSettings } from "../lib/settings";
+import { getDigestApp, type DigestApp, type DigestAppConfig } from "../lib/digestApps";
+import {
+ digestPromptVariant,
+ type DigestFreshnessTarget,
+ type DigestSectionKind,
+ type DigestTimestampMode,
+} from "../lib/digest";
+import { PROMPT_VERSION, maxCuesForContext } from "../lib/digestPrompt";
+import {
+ readDigestContext,
+ type DigestContext,
+} from "../lib/digestContext-server";
+
+// Which lane to resolve against. The string form the editor actions already
+// speak, rather than DigestLane, so call sites pass what they already hold.
+export type DigestLaneChoice = "local" | "remote";
+
+export type ResolvedDigestTarget = {
+ target: DigestFreshnessTarget;
+ // Which sections must ALL be fresh for a video to count as digested.
+ sections: DigestSectionKind[];
+ app: DigestApp;
+ config: DigestAppConfig;
+ // What the config asked for. Distinct from what the engine reports having run
+ // — freshness compares this one (see DigestProvenance.modelRequested).
+ modelRequested: string;
+ timestampMode: DigestTimestampMode;
+ // Cues per chunk at this app's context size. The batch passes it on to the
+ // chunker, so the identity and the actual chunking cannot drift apart.
+ maxCues: number;
+ // The channel-context note + its hash. Returned rather than re-read because
+ // the hash is already IN the target above: reading it twice invites the two
+ // copies to disagree.
+ context: DigestContext;
+};
+
+export async function resolveDigestTarget(opts: {
+ paths: Paths;
+ channelSlug: string;
+ lane?: DigestLaneChoice;
+ // Overrides the lane's configured app. The bake-off harness drives this.
+ appId?: string;
+ sections?: DigestSectionKind[];
+}): Promise<ResolvedDigestTarget> {
+ const digestSettings = getSettings().digest;
+ const appId =
+ opts.appId ??
+ (opts.lane === "remote"
+ ? digestSettings.remoteAppId
+ : digestSettings.localAppId);
+ const app = getDigestApp(appId);
+ const config: DigestAppConfig = digestSettings.apps[app.id] ?? {};
+ const modelRequested = config.model?.trim() || app.defaultModel();
+ const maxCues = maxCuesForContext(config.numCtx);
+ const promptVariant = digestPromptVariant({
+ ...digestSettings,
+ maxCues,
+ });
+ const context = await readDigestContext(opts.paths, opts.channelSlug);
+ return {
+ target: {
+ appId: app.id,
+ model: modelRequested,
+ promptVersion: PROMPT_VERSION,
+ contextHash: context.hash,
+ ...(promptVariant ? { promptVariant } : {}),
+ },
+ sections: opts.sections ?? digestSettings.sections,
+ app,
+ config,
+ modelRequested,
+ timestampMode: digestSettings.timestampMode,
+ maxCues,
+ context,
+ };
+}
diff --git a/common/controller/digestVideo.ts b/common/controller/digestVideo.ts
@@ -0,0 +1,391 @@
+// Generate the AI digest for ONE video: chapters and/or topic tags, written to
+// the ai-digest.json sidecar next to the transcript.
+//
+// Pipeline (each step reusing the declared single source of truth for its job):
+// isCuesJsonFresh — don't digest a transcript that's about to change
+// readNormalizedTranscript
+// chunkCuesForContext — sequential overlapping slices
+// transcriptToMarkdown — cues -> text for an AI, with hms() stamps
+// digest app — ollama (local) or claude CLI (metered, opt-in)
+// parseChapters/parseTags — the guards; every rejection recorded
+// writeDigestSection — read-modify-write ONE section, atomic rename
+//
+// THE FRESHNESS SKIP IS THE POINT. A full sweep of the 74k-transcript corpus is
+// weeks of wall-clock on one local lane, so a redo is unaffordable. A section is
+// regenerated only when its recorded (schemaVersion, appId, model, promptVersion,
+// contextHash) differs from what we would produce now — which makes a re-run
+// after a prompt change minutes long over a sample instead of weeks over the
+// corpus.
+
+import path from "node:path";
+import type { Paths } from "../lib/paths";
+import {
+ DIGEST_OVERLAP_CUES,
+ PROMPT_VERSION,
+ CHAPTER_SYSTEM_PROMPT,
+ TAG_SYSTEM_PROMPT,
+ buildChapterPrompt,
+ buildTagPrompt,
+ chapterSchema,
+ maxCuesForContext,
+ tagSchema,
+ toHms,
+} from "../lib/digestPrompt";
+import { getDigestApp, type DigestAppConfig } from "../lib/digestApps";
+import {
+ digestPromptVariant,
+ isSectionFresh,
+ type DigestItem,
+ type DigestTimestampMode,
+ type DigestProvenance,
+ type DigestRecord,
+ type DigestSectionKind,
+ type DigestWarning,
+} from "../lib/digest";
+import {
+ loadDigest,
+ writeDigestFailure,
+ writeDigestSection,
+} from "../lib/digest-server";
+import { readDigestContext, type DigestContext } from "../lib/digestContext-server";
+import { parseChapters, parseTags, type DigestChunkOutput } from "../lib/digestParse";
+import { chunkCuesForContext } from "../lib/transcriptWindow";
+import { transcriptToMarkdown } from "../lib/transcriptToMarkdown";
+import type { Cue } from "../lib/vtt";
+import {
+ isCuesJsonFresh,
+ readNormalizedTranscript,
+} from "./normalizeTranscript";
+
+export type DigestVideoOptions = {
+ paths: Paths;
+ channelSlug: string;
+ videoId: string;
+ // Which sections to generate. Defaults to chapters only — tags are cheap but
+ // double the call count, so the sweep opts into them explicitly.
+ sections?: DigestSectionKind[];
+ appId?: string;
+ config?: DigestAppConfig;
+ // Prompt SHAPE. Both default, so an existing caller keeps producing records
+ // with no promptVariant and every digest already on disk stays fresh.
+ timestampMode?: DigestTimestampMode;
+ promptVariant?: string;
+ // Pre-read channel context, so a batch reads it once per channel instead of
+ // once per video. Omitted → read here.
+ context?: DigestContext;
+ // Regenerate even when the recorded provenance matches. For Stage B iteration
+ // on a fixed sample; never set for a sweep.
+ force?: boolean;
+ onLog?: (msg: string) => void;
+ signal?: AbortSignal;
+};
+
+export type DigestVideoOutcome =
+ | {
+ status: "wrote";
+ record: DigestRecord;
+ sections: DigestSectionKind[];
+ itemCount: number;
+ warningCount: number;
+ costUsd: number;
+ engineCalls: number;
+ }
+ | { status: "fresh" }
+ | {
+ status: "skipped";
+ reason: "no-transcript" | "stale-cues" | "no-cues" | "no-metadata";
+ };
+
+// A per-chunk engine failure must not lose the chunks that DID work: a 12-hour
+// video is 30+ calls and one 500 from ollama should cost that chunk, not the
+// video. The failure is recorded as a warning and the section is written from
+// what survived, with chunks/chunksOk in provenance so a partially-generated
+// section is identifiable later.
+export async function digestVideo(
+ opts: DigestVideoOptions,
+): Promise<DigestVideoOutcome> {
+ const log = opts.onLog ?? (() => {});
+ const videoDir = path.join(
+ opts.paths.channelsDir,
+ opts.channelSlug,
+ "data",
+ opts.videoId,
+ );
+ const sections = opts.sections ?? ["chapters"];
+
+ // Don't digest a transcript that is about to be rewritten: a stale cues.json
+ // means the raw transcript changed under it, so the digest would describe
+ // superseded text and then look "fresh" forever.
+ const { fresh: cuesFresh, cuesPath } = await isCuesJsonFresh(videoDir);
+ if (!cuesFresh) {
+ const existing = await readNormalizedTranscript(cuesPath);
+ return {
+ status: "skipped",
+ reason: existing ? "stale-cues" : "no-transcript",
+ };
+ }
+ const transcript = await readNormalizedTranscript(cuesPath);
+ if (!transcript) return { status: "skipped", reason: "no-transcript" };
+ const cues = transcript.cues ?? [];
+ if (cues.length === 0) return { status: "skipped", reason: "no-cues" };
+
+ const app = getDigestApp(opts.appId);
+ const config = opts.config ?? {};
+ const modelRequested = config.model?.trim() || app.defaultModel();
+ const context =
+ opts.context ?? (await readDigestContext(opts.paths, opts.channelSlug));
+
+ const timestampMode = opts.timestampMode ?? "absolute";
+ // Sized to the CONFIGURED context, not to a constant. A 8192-token window with
+ // 1200-cue chunks overflows and ollama truncates without saying so.
+ const maxCues = maxCuesForContext(config.numCtx);
+ // One derivation, shared with the batch and the bake-off, so the same config
+ // never produces two different identities.
+ const promptVariant = digestPromptVariant({
+ promptVariant: opts.promptVariant,
+ timestampMode,
+ maxCues,
+ });
+
+ const existing = await loadDigest(videoDir);
+ const target = {
+ appId: app.id,
+ model: modelRequested,
+ promptVersion: PROMPT_VERSION,
+ contextHash: context.hash,
+ ...(promptVariant ? { promptVariant } : {}),
+ };
+ const stale = sections.filter(
+ (section) => opts.force || !isSectionFresh(existing, section, target),
+ );
+ if (stale.length === 0) return { status: "fresh" };
+
+ const chunks = chunkCuesForContext(cues, {
+ maxCues,
+ overlapCues: DIGEST_OVERLAP_CUES,
+ });
+ log(
+ `${opts.channelSlug}/${opts.videoId}: ${cues.length} cues → ${chunks.length} chunk(s), sections: ${stale.join(", ")}`,
+ );
+
+ let record: DigestRecord | null = existing;
+ let itemCount = 0;
+ let warningCount = 0;
+ let costUsd = 0;
+ let engineCalls = 0;
+
+ for (const section of stale) {
+ const outputs: DigestChunkOutput[] = [];
+ const runWarnings: DigestWarning[] = [];
+ let reportedModel = modelRequested;
+ let sectionCost = 0;
+
+ for (let i = 0; i < chunks.length; i++) {
+ opts.signal?.throwIfAborted();
+ const chunk = chunks[i];
+ const startSeconds = Math.max(0, Math.floor(chunk[0].start));
+ const endSeconds = Math.max(
+ startSeconds,
+ Math.ceil(chunk[chunk.length - 1].end || chunk[chunk.length - 1].start),
+ );
+ const promptInput = {
+ title: transcript.title || opts.videoId,
+ channel: transcript.channel || opts.channelSlug,
+ startSeconds,
+ endSeconds,
+ // The renderer and the prompt MUST agree on the numbering, so both are
+ // driven from the same value rather than each deciding for itself.
+ transcript: renderChunk(
+ transcript,
+ chunk,
+ timestampMode === "chunk-local" ? startSeconds : 0,
+ ),
+ timestampMode,
+ ...(context.note ? { contextNote: context.note } : {}),
+ };
+ const span = endSeconds - startSeconds;
+ const request =
+ section === "chapters"
+ ? {
+ system: CHAPTER_SYSTEM_PROMPT,
+ prompt: buildChapterPrompt(promptInput),
+ schema: chapterSchema(span),
+ }
+ : {
+ system: TAG_SYSTEM_PROMPT,
+ prompt: buildTagPrompt(promptInput),
+ schema: tagSchema(),
+ };
+ try {
+ const result = await app.run({
+ ...request,
+ config,
+ signal: opts.signal,
+ onLog: opts.onLog,
+ });
+ engineCalls++;
+ reportedModel = result.model || reportedModel;
+ if (typeof result.costUsd === "number") sectionCost += result.costUsd;
+ outputs.push({
+ index: i,
+ startSeconds,
+ endSeconds,
+ data: result.data,
+ timestampMode,
+ });
+ } catch (err) {
+ // A cancel is not a chunk failure — let it propagate so the batch stops.
+ if (opts.signal?.aborted) throw err;
+ const message = (err as Error)?.message ?? String(err);
+ runWarnings.push({
+ code: "chunk-failed",
+ section,
+ chunk: i,
+ detail: message.slice(0, 300),
+ });
+ log(
+ `${opts.channelSlug}/${opts.videoId}: chunk ${i + 1}/${chunks.length} failed: ${message}`,
+ );
+ }
+ }
+
+ if (outputs.length === 0) {
+ // Nothing usable. Do NOT write a section — an empty section with current
+ // provenance would read as "fresh" and the video would never be retried.
+ // But DO persist the warnings, in the sibling `failures` field that
+ // freshness never reads: this is the worst outcome the generator has, and
+ // it used to leave no trace anywhere except a job log that rotates.
+ log(
+ `${opts.channelSlug}/${opts.videoId}: ${section} produced no usable output (${runWarnings.length} warning(s)); leaving the sidecar untouched so it retries.`,
+ );
+ record = await writeDigestFailure(videoDir, {
+ section,
+ at: new Date().toISOString(),
+ appId: app.id,
+ model: reportedModel,
+ promptVersion: PROMPT_VERSION,
+ reason: "no-output",
+ chunks: chunks.length,
+ chunksOk: outputs.length,
+ warnings: runWarnings,
+ });
+ warningCount += runWarnings.length;
+ continue;
+ }
+
+ const parsed =
+ section === "chapters"
+ ? parseChapters(outputs, cues)
+ : parseTags(outputs);
+ const items: DigestItem[] =
+ "chapters" in parsed ? parsed.chapters : parsed.tags;
+ const warnings = [...runWarnings, ...parsed.warnings];
+
+ if (items.length === 0) {
+ // The model DID propose content and every item failed a guard — a
+ // different failure from "produced nothing", and one the warnings can
+ // actually explain (the validation run's worst video had 13 chapters
+ // clamped away as out-of-range). Recording which of the two happened is
+ // the difference between a review queue that can act and one that can only
+ // report a blank.
+ log(
+ `${opts.channelSlug}/${opts.videoId}: every ${section} entry was rejected (${warnings.length} warning(s)); leaving the sidecar untouched so it retries.`,
+ );
+ record = await writeDigestFailure(videoDir, {
+ section,
+ at: new Date().toISOString(),
+ appId: app.id,
+ model: reportedModel,
+ promptVersion: PROMPT_VERSION,
+ reason: "all-rejected",
+ chunks: chunks.length,
+ chunksOk: outputs.length,
+ warnings,
+ });
+ warningCount += warnings.length;
+ continue;
+ }
+
+ const provenance: DigestProvenance = {
+ appId: app.id,
+ model: reportedModel,
+ modelRequested,
+ lane: app.lane,
+ generatedAt: new Date().toISOString(),
+ promptVersion: PROMPT_VERSION,
+ ...(promptVariant ? { promptVariant } : {}),
+ contextHash: context.hash,
+ chunks: chunks.length,
+ chunksOk: outputs.length,
+ ...(app.metered && sectionCost > 0 ? { costUsd: sectionCost } : {}),
+ };
+ record = await writeDigestSection(videoDir, {
+ section,
+ items,
+ provenance,
+ warnings,
+ });
+ itemCount += items.length;
+ warningCount += warnings.length;
+ costUsd += sectionCost;
+ log(
+ `${opts.channelSlug}/${opts.videoId}: ${section} → ${items.length} item(s), ${warnings.length} warning(s)` +
+ (sectionCost > 0 ? `, $${sectionCost.toFixed(4)}` : ""),
+ );
+ }
+
+ if (!record || itemCount === 0) {
+ // Every requested section failed. Report it as a skip rather than a write so
+ // the batch's counters stay honest.
+ return { status: "skipped", reason: "no-cues" };
+ }
+ return {
+ status: "wrote",
+ record,
+ sections: stale,
+ itemCount,
+ warningCount,
+ costUsd,
+ engineCalls,
+ };
+}
+
+// Render one chunk as the text the engine sees. transcriptToMarkdown is the
+// declared single source of truth for "transcript -> text for an AI"; the
+// stampForCue override makes every line carry a HH:MM:SS marker, which is what
+// the prompt tells the model to copy its `start` values from.
+//
+// toHms — NOT aiHandoff's hms() — because hms drops the hour field below an hour
+// ("2:36") and abbreviates it above one ("1:00:00"), while the schema pattern
+// requires two digits in all three fields. Feeding the model markers it cannot
+// legally echo would reintroduce the malformed-timestamp failure the pin fixed.
+//
+// The description is omitted: it is the uploader's own promotional copy and
+// biases titles toward it.
+//
+// `offsetSeconds` is subtracted from every marker. It is 0 in absolute mode and
+// the chunk's own start in chunk-local mode, which is the whole of what
+// "re-basing" means — the cue list itself is never modified, only how it is
+// stamped, so the parser's cue snap still works against real cue times.
+function renderChunk(
+ transcript: { id: string; title: string; channel?: string; duration?: number },
+ cues: Cue[],
+ offsetSeconds = 0,
+): string {
+ return transcriptToMarkdown(
+ {
+ id: transcript.id,
+ title: transcript.title,
+ channel: transcript.channel,
+ duration: transcript.duration,
+ cues,
+ },
+ {
+ timestamps: true,
+ includeDescription: false,
+ includeTags: false,
+ stampForCue: (_clock, seconds) =>
+ toHms(Math.max(0, seconds - offsetSeconds)),
+ },
+ );
+}
diff --git a/common/controller/digestYield.ts b/common/controller/digestYield.ts
@@ -0,0 +1,68 @@
+// GPU arbitration between the digest lane and the transcription lane.
+//
+// `queueKeys.ts` deliberately puts `digest:local` on a DIFFERENT registry queue
+// from TRANSCRIPTION_QUEUE, so the two never serialize against each other. That
+// reasoning is right for digest-local vs digest-remote (GPU-bound vs
+// network-bound), but its consequence for digest-local vs transcription is that
+// **ollama and whisper run concurrently on the same 8 GB card**.
+//
+// That is not hypothetical: the bake-off projected 27 s per audio-hour on an
+// idle box and the 102-video validation run measured **90 s** on a box also
+// running auto-transcribe. Over a 77,000-audio-hour sweep the difference is
+// roughly 80 days versus 24 — far and away the largest cost lever in the
+// backfill, and about fifty times bigger than everything duplicate sharing can
+// save (see bin/digest-plan.ts).
+//
+// The fix is deliberately NOT a scheduler. `digestBatch`'s limit() already
+// returns 0 to idle-wait on a pause, and runPool treats a zero limit as "hold,
+// don't finish" — so the digest lane can step aside using machinery that is
+// already proven, with no priority system, no new queue and no boot hook.
+//
+// This is a YIELD, not a lock. A digest already in flight is not interrupted, so
+// the two can still overlap for the length of one generation if transcription
+// starts just after a dispatch. Interrupting would waste the partial work; the
+// next dispatch sees the busy lane and holds.
+
+import { getWorkerPool } from "../jobs/workerPool";
+import { getRegistry } from "../jobs/registry";
+import { TRANSCRIPTION_QUEUE } from "../lib/queueKeys";
+
+export type TranscriptionActivity = {
+ busy: boolean;
+ // Which signal fired, for the log line. An operator watching a sweep sit at
+ // zero throughput needs to be told it is yielding rather than wedged.
+ reason: "local-worker" | "queued-job" | null;
+};
+
+// Is the local transcription lane using the GPU right now?
+//
+// Two signals, because one alone leaves a hole:
+// - a busy LOCAL worker is whisper actually running on this box's card
+// (`remote` workers delegate to another machine and compete for nothing
+// here, so they are deliberately excluded);
+// - a RUNNING job on TRANSCRIPTION_QUEUE covers the gaps between worker
+// acquisitions — audio extraction, model load, the moment between two
+// videos in a batch. Those gaps are exactly where a multi-minute digest
+// generation would otherwise slip in and hold the card.
+export function transcriptionActivity(): TranscriptionActivity {
+ try {
+ for (const w of getWorkerPool().summary()) {
+ if (w.busy && w.kind === "local") {
+ return { busy: true, reason: "local-worker" };
+ }
+ }
+ } catch {
+ // A pool that cannot be read must not wedge the digest lane: fail OPEN, i.e.
+ // keep generating. The cost of being wrong here is contention, not deadlock.
+ }
+ try {
+ for (const job of getRegistry().list()) {
+ if (job.queueKey === TRANSCRIPTION_QUEUE && job.status === "running") {
+ return { busy: true, reason: "queued-job" };
+ }
+ }
+ } catch {
+ /* same: fail open */
+ }
+ return { busy: false, reason: null };
+}
diff --git a/common/controller/duplicateShorts.ts b/common/controller/duplicateShorts.ts
@@ -1,22 +1,41 @@
-// Cross-platform duplicate "shorts" detection — a global pass over every
-// channel's videos (NOT per-channel, since duplicates are cross-channel and
-// cross-platform by definition).
+// Cross-platform duplicate detection — a global pass over every channel's videos
+// (NOT per-channel, since duplicates are cross-channel and cross-platform by
+// definition).
//
-// Hybrid cascade:
-// Phase 1 (cheap) — pre-cluster candidates by rounded duration, the only
-// signal stable across platforms and re-titles. Consumes
-// the already-aggregated VideoStat index (LMDB statsByPath,
-// populated by buildStats) rather than re-walking dirs.
-// Phase 2 (confirm)— for each candidate pair, compare transcript content
-// (read from the LMDB `cues` sub-db that buildIndex
-// populates): exact text hash → 5-gram Jaccard →
-// (all-durations only) containment for "short is a clip of
-// a longer video". Falls back to a metadata-only match when
-// a transcript is missing/empty.
+// THE PRE-FILTER PROPOSES; THE TRANSCRIPT DISPOSES.
//
-// Output is a single global transcripts/duplicates.json. Flag-only: nothing is
-// merged or deleted. Requires build:index (cues) + build:stats (statsByPath) to
-// have run first.
+// A blocking strategy (§ "Phase 1" below) only NOMINATES pairs. It is never
+// evidence on its own. Every nominated pair then goes through the transcript
+// cascade, which has three outcomes:
+//
+// both sides have a transcript, content matches → transcript-exact/near.
+// Confirmed; may share.
+// both sides have a transcript, content does NOT → REJECTED, no cluster.
+// This is the whole value of
+// testing the suspects.
+// one side has no transcript → untestable. Kept only when
+// the pre-filter's own claim
+// is strong enough to be
+// worth a human's attention
+// (title + near-identical
+// runtime), as a needsReview
+// `title-duration` suspect
+// that shares nothing.
+//
+// STREAMING, NOT BATCHING. The first implementation built one global candidate
+// array, then fingerprinted every video it referenced, then evaluated. Both
+// halves fail corpus-wide: the containment pass alone was an unblocked cartesian
+// product (498 M pairs — V8 throws RangeError building the array), and holding
+// 5-word shingle sets for the whole corpus at once (a 3.5 h video is ~40 k
+// strings) exhausted a 4 GB heap. Neither is inherent to corpus-wide detection.
+// Blocking fixes the first; this file's block-at-a-time iteration fixes the
+// second: fingerprint one block, evaluate its pairs, keep only the confirmed
+// matches, release. Peak memory is O(largest block), not O(corpus), and no
+// multi-million-element pair array is ever materialised.
+//
+// Inputs are the already-aggregated VideoStat index (LMDB statsByPath, from
+// buildStats) plus the `cues` sub-db (from buildIndex). Output is a single
+// global transcripts/duplicates.json. Flag-only: nothing is merged or deleted.
import path from "node:path";
import { mkdir, rename, writeFile, readFile } from "node:fs/promises";
@@ -30,19 +49,31 @@ import type { Cue } from "../lib/vtt";
import { parseTranscriptJson } from "../lib/whisper";
import { readVideoFiles, pickIndexTranscript } from "../lib/videoStatus";
import {
+ DEFAULT_ALIGNMENT_TOLERANCE_SECONDS,
DEFAULT_CONTAINMENT_THRESHOLD,
DEFAULT_DURATION_TOLERANCE_SECONDS,
+ DEFAULT_TITLE_DURATION_RATIO,
+ MAX_DURATION_BLOCK_SIZE,
+ MAX_TITLE_GROUP_SIZE,
+ durationsCompatible,
+ measureAlignment,
+ normalizeTitleKey,
DEFAULT_NEAR_THRESHOLD,
DEFAULT_SHINGLE_SIZE,
DEFAULT_SHORT_THRESHOLD_SECONDS,
DUPLICATES_FILENAME,
DUPLICATE_REPORT_VERSION,
+ DUPLICATE_OVERRIDES_FILENAME,
UnionFind,
comparisonText,
containment,
jaccard,
+ pickCanonicalSlug,
+ sanitizeDuplicateOverrides,
shingles,
strongerMatch,
+ type DuplicateBlocking,
+ type DuplicateOverrides,
type DuplicateCluster,
type DuplicateMatchKind,
type DuplicateReport,
@@ -50,6 +81,17 @@ import {
type DuplicateVideoRef,
} from "../lib/duplicates";
+// A normalized title shorter than this carries no blocking information —
+// pairing on it would rebuild the cartesian product the pass exists to avoid.
+const MIN_TITLE_KEY_LENGTH = 8;
+
+// Hard budget for the shorts-vs-longer containment sweep, which is a genuine
+// cartesian product and the one pass blocking cannot rescue. Under the budget it
+// runs (and finds clip-of-longer matches, which nothing else can); over it, it
+// is SKIPPED AND REPORTED rather than silently truncated or left to die. Scaling
+// containment properly wants MinHash/LSH signatures — see STATE.md.
+const MAX_CONTAINMENT_PAIRS = 5_000_000;
+
type PathKey = [string, string];
type IndexKey = [string, string, string];
type StatsRecord = { metaMs: number; stat: VideoStat };
@@ -63,11 +105,21 @@ export type DetectDuplicateShortsOptions = {
nearThreshold?: number;
containmentThreshold?: number;
shingleSize?: number;
+ // Candidate-nomination strategy. Defaults to "title" corpus-wide (where a
+ // duration sweep is affordable but far slower) and "duration" in shorts mode,
+ // which preserves today's proven behaviour at that scale.
+ blocking?: DuplicateBlocking;
+ titleDurationRatio?: number;
+ // Tolerance for the per-cluster timing measurement written onto each ref.
+ alignmentToleranceSeconds?: number;
onLog?: (msg: string) => void;
signal?: AbortSignal;
};
-// Per-participant transcript fingerprint, computed once and reused across pairs.
+// Per-participant transcript fingerprint, computed once per BLOCK and released
+// when the block is done. Holding these for the whole corpus is what exhausted
+// the heap; the shingle set is the expensive part (~40 k strings for a 3.5 h
+// video), which is why it never outlives its block.
type Fingerprint = {
stat: VideoStat;
hasTranscript: boolean;
@@ -77,6 +129,24 @@ type Fingerprint = {
type Entry = { stat: VideoStat; videoDir: string };
+// One nominated pair. `suspectKind` is what this pair's evidence is worth when
+// the content CANNOT be compared because a side has no transcript: a shared
+// title AND a near-identical runtime is a claim worth a human's attention, so it
+// survives as a `title-duration` suspect; a shared duration alone is not, and
+// treating it as one is exactly what produced enormous false clusters of
+// unrelated same-length videos before, so it is null and the pair is dropped.
+type Nomination = {
+ a: Entry;
+ b: Entry;
+ // Skip the similarity test and go straight to containment — a clip can never
+ // pass Jaccard against the full recording it was cut from.
+ containmentOnly: boolean;
+ suspectKind: DuplicateMatchKind | null;
+};
+
+// A unit of work: fingerprint these members, evaluate these pairs, release.
+type Block = { members: Entry[]; nominations: Nomination[] };
+
type PairResult = {
a: string; // slug
b: string; // slug
@@ -85,10 +155,29 @@ type PairResult = {
contained: boolean;
};
+// Ordered slug-pair key, so "both" can union two nomination streams without
+// evaluating the overlap twice.
+function pairKey(a: string, b: string): string {
+ return a < b ? `${a}\n${b}` : `${b}\n${a}`;
+}
+
+function membersOf(nominations: Nomination[]): Entry[] {
+ const m = new Map<string, Entry>();
+ for (const n of nominations) {
+ m.set(n.a.stat.slug, n.a);
+ m.set(n.b.stat.slug, n.b);
+ }
+ return [...m.values()];
+}
+
export async function detectDuplicateShorts(
opts: DetectDuplicateShortsOptions,
): Promise<DuplicateReport> {
const log = opts.onLog ?? ((m: string) => console.log(m));
+ // Corpus-wide (thresholdSeconds: null) is untested at 74k videos across the
+ // long tail, so the run reports its own wall-clock: the cost of the pass has to
+ // be a measured number before anything is wired to it on a schedule.
+ const startedAt = Date.now();
const thresholdSeconds =
opts.thresholdSeconds === undefined
? DEFAULT_SHORT_THRESHOLD_SECONDS
@@ -99,12 +188,24 @@ export async function detectDuplicateShorts(
opts.containmentThreshold ?? DEFAULT_CONTAINMENT_THRESHOLD;
const shingleSize = opts.shingleSize ?? DEFAULT_SHINGLE_SIZE;
+ // Corpus-wide defaults to TITLE blocking: it is effectively linear and it is
+ // the signal that finds cross-platform re-uploads, which keep their name.
+ // Duration blocking is viable corpus-wide too since the streaming rewrite (it
+ // is the only strategy that catches a RE-TITLED mirror) but it nominates
+ // millions of pairs where title nominates thousands, so it is opt-in.
+ const blocking =
+ opts.blocking ?? (thresholdSeconds === null ? "title" : "duration");
+ const titleDurationRatio =
+ opts.titleDurationRatio ?? DEFAULT_TITLE_DURATION_RATIO;
+ const alignmentTolerance =
+ opts.alignmentToleranceSeconds ?? DEFAULT_ALIGNMENT_TOLERANCE_SECONDS;
const runConfig: DuplicateRunConfig = {
thresholdSeconds,
durationToleranceSeconds: W,
nearThreshold,
containmentThreshold,
shingleSize,
+ blocking,
};
// ---- Open the index: statsByPath (Phase-1 input) + cues (Phase-2 input) ---
@@ -151,113 +252,248 @@ export async function detectDuplicateShorts(
`(threshold=${thresholdSeconds === null ? "all" : `${thresholdSeconds}s`}).`,
);
- // ---- Phase 1: bucket by rounded duration ---------------------------------
- const bucketOf = (d: number) => Math.round(d / W);
- const buckets = new Map<number, Entry[]>();
- for (const e of participants) {
- const b = bucketOf(e.stat.duration);
- const arr = buckets.get(b);
- if (arr) arr.push(e);
- else buckets.set(b, [e]);
- }
+ // ---- Phase 1 + 2, interleaved: nominate a block, judge it, release it -----
+ //
+ // The two phases are no longer separate passes. Each block is fingerprinted,
+ // evaluated and dropped before the next is built, so peak memory is O(largest
+ // block) and only the (few, by construction) confirmed matches survive.
+ const stats = {
+ nominated: 0,
+ evaluated: 0,
+ confirmed: 0,
+ rejected: 0,
+ suspects: 0,
+ titleGroups: 0,
+ oversizedGroups: 0,
+ oversizedGroupVideos: 0,
+ oversizedBlocks: 0,
+ oversizedBlockVideos: 0,
+ untitled: 0,
+ fingerprinted: 0,
+ rawFallbacks: 0,
+ rawHits: 0,
+ peakRssMb: 0,
+ };
+ const noteRss = () => {
+ const mb = Math.round(process.memoryUsage().rss / 1_048_576);
+ if (mb > stats.peakRssMb) stats.peakRssMb = mb;
+ };
- // Candidate pairs. Same or adjacent duration bucket + relative-duration guard,
- // or (all-durations mode) short/longer pairs eligible for containment.
- const maxDelta = (d: number) => Math.max(W, Math.ceil(0.02 * d));
- type Candidate = { a: Entry; b: Entry; containmentOnly: boolean };
- const candidates: Candidate[] = [];
- const sortedBuckets = [...buckets.keys()].sort((x, y) => x - y);
- for (const b of sortedBuckets) {
- const here = buckets.get(b) as Entry[];
- const next = buckets.get(b + 1) ?? [];
- for (let i = 0; i < here.length; i++) {
- for (let j = i + 1; j < here.length; j++) {
- if (durationsClose(here[i], here[j], maxDelta))
- candidates.push({ a: here[i], b: here[j], containmentOnly: false });
+ // Fingerprint one block's members, re-using anything the previous block
+ // already computed. Primary source is the LMDB `cues` sub-db — but parseVtt
+ // only extracts text from YouTube karaoke-tagged cues, so plain-VTT
+ // transcripts (most non-YouTube captions, whisper-as-vtt, manual subs) store
+ // as zero cues and would look transcript-less. For those, fall back to reading
+ // the raw transcript file off disk and extracting plain text, for comparison
+ // only — nothing is persisted.
+ const fingerprintBlock = async (
+ members: Entry[],
+ reuse: ReadonlyMap<string, Fingerprint>,
+ ): Promise<Map<string, Fingerprint>> => {
+ const out = new Map<string, Fingerprint>();
+ const rawFallbacks: Entry[] = [];
+ for (const e of members) {
+ const s = e.stat;
+ if (out.has(s.slug)) continue;
+ const cached = reuse.get(s.slug);
+ if (cached) {
+ out.set(s.slug, cached);
+ continue;
+ }
+ const cues = cuesDb.get([s.uploadDate, s.channelSlug, s.id]);
+ if (cues && cues.length > 0) {
+ stats.fingerprinted++;
+ out.set(s.slug, fingerprintFrom(s, cues, shingleSize));
+ } else {
+ rawFallbacks.push(e);
}
}
- for (const x of here) {
- for (const y of next) {
- if (durationsClose(x, y, maxDelta))
- candidates.push({ a: x, b: y, containmentOnly: false });
+ if (rawFallbacks.length > 0) {
+ stats.rawFallbacks += rawFallbacks.length;
+ const limit = pLimit(16);
+ await Promise.all(
+ rawFallbacks.map((e) =>
+ limit(async () => {
+ opts.signal?.throwIfAborted();
+ const cues = await readRawCues(opts.paths.channelsDir, e);
+ if (cues && cues.length > 0) stats.rawHits++;
+ stats.fingerprinted++;
+ out.set(e.stat.slug, fingerprintFrom(e.stat, cues, shingleSize));
+ }),
+ ),
+ );
+ }
+ return out;
+ };
+
+ // Cues WITH real start times, for the alignment measurement only. Deliberately
+ // LMDB-only: the raw-VTT fallback stamps every cue at start 0, which is fine
+ // for a set-based content fingerprint and worthless — actively misleading —
+ // for measuring an offset. No timed cues → no measurement → `aligned` stays
+ // undefined, which every consumer must read as "not aligned".
+ const timedCues = (e: Entry | undefined): Cue[] | null => {
+ if (!e) return null;
+ const s = e.stat;
+ const cues = cuesDb.get([s.uploadDate, s.channelSlug, s.id]);
+ return cues && cues.length > 0 ? cues : null;
+ };
+
+ const streams: Generator<Block>[] = [];
+ if (blocking === "title" || blocking === "both") {
+ streams.push(titleBlocks(participants, titleDurationRatio, stats, log));
+ }
+ if (blocking === "duration" || blocking === "both") {
+ streams.push(durationBlocks(participants, W, stats, log));
+ }
+ // Only "both" needs cross-stream dedup. It is deliberately one-directional:
+ // the FIRST stream records its nominations and later streams consult them.
+ // Title blocking is pushed first precisely so the remembered set is the small
+ // one — recording the duration stream's millions of pair keys as well would
+ // reintroduce, in the dedup set, exactly the whole-corpus retention the
+ // streaming rewrite exists to remove.
+ const seenPairs = blocking === "both" ? new Set<string>() : null;
+
+ const matches: PairResult[] = [];
+ const hasTranscriptBySlug = new Map<string, boolean>();
+
+ // 2-block sliding fingerprint cache: a duration block is bucket[b] ∪
+ // bucket[b+1], so every member except the last bucket is re-used by the very
+ // next block. Without this, each bucket would be read and shingled twice.
+ let prevFps = new Map<string, Fingerprint>();
+ let blocksDone = 0;
+
+ for (const [streamIndex, stream] of streams.entries()) {
+ const recordPairs = seenPairs !== null && streamIndex === 0;
+ const skipSeenPairs = seenPairs !== null && streamIndex > 0;
+ for (const block of stream) {
+ opts.signal?.throwIfAborted();
+ let nominations = block.nominations;
+ if (recordPairs) {
+ for (const n of nominations) {
+ seenPairs.add(pairKey(n.a.stat.slug, n.b.stat.slug));
+ }
+ } else if (skipSeenPairs) {
+ nominations = nominations.filter(
+ (n) => !seenPairs.has(pairKey(n.a.stat.slug, n.b.stat.slug)),
+ );
+ }
+ stats.nominated += nominations.length;
+ if (nominations.length === 0) continue;
+
+ // When dedup narrowed the nominations, re-derive the members from what is
+ // actually left rather than fingerprinting videos no surviving pair needs.
+ const fps = await fingerprintBlock(
+ skipSeenPairs ? membersOf(nominations) : block.members,
+ prevFps,
+ );
+ for (const [slug, fp] of fps) hasTranscriptBySlug.set(slug, fp.hasTranscript);
+
+ for (const n of nominations) {
+ const fa = fps.get(n.a.stat.slug);
+ const fb = fps.get(n.b.stat.slug);
+ if (!fa || !fb) continue;
+ stats.evaluated++;
+ const res = n.containmentOnly
+ ? evalContainment(fa, fb, containmentThreshold)
+ : evalBlocked(fa, fb, {
+ nearThreshold,
+ containmentThreshold,
+ allowContainment: includeContainment,
+ suspectKind: n.suspectKind,
+ });
+ if (!res) {
+ stats.rejected++;
+ continue;
+ }
+ if (res.kind === "title-duration") stats.suspects++;
+ else stats.confirmed++;
+ matches.push({ a: n.a.stat.slug, b: n.b.stat.slug, ...res });
+ }
+ prevFps = fps;
+ noteRss();
+ // A corpus-wide duration run is long enough that silence is
+ // indistinguishable from a hang. Heartbeat with the numbers that matter.
+ if (++blocksDone % 100 === 0) {
+ log(
+ ` …${blocksDone} block(s), ${stats.evaluated} pair(s) tested, ` +
+ `${stats.confirmed} confirmed, RSS ${Math.round(process.memoryUsage().rss / 1_048_576)} MB`,
+ );
}
}
+ prevFps = new Map();
}
- // Containment candidates: a short paired with any meaningfully-longer video.
+ // ---- The containment sweep: shorts vs every meaningfully-longer video -----
+ //
+ // The one pass blocking cannot rescue — a clip and its parent share neither a
+ // duration nor, usually, a title, so nothing nominates them but a cartesian
+ // product. Bounded by an explicit pair budget and skipped-with-a-log when it
+ // does not fit, rather than silently truncated. Memory stays bounded because
+ // the budget implicitly bounds the short side: at 76 k longer videos, fitting
+ // under the budget means only a handful of shorts.
if (includeContainment) {
- const shortsForContainment = participants.filter(
+ const shorts = participants.filter(
(e) =>
- e.stat.duration > 0 &&
- e.stat.duration <= DEFAULT_SHORT_THRESHOLD_SECONDS,
+ e.stat.duration > 0 && e.stat.duration <= DEFAULT_SHORT_THRESHOLD_SECONDS,
);
- for (const s of shortsForContainment) {
+ const budget = shorts.length * participants.length;
+ if (shorts.length === 0) {
+ // nothing to sweep
+ } else if (budget > MAX_CONTAINMENT_PAIRS) {
+ log(
+ `Containment sweep SKIPPED: ${shorts.length} short(s) × ${participants.length} video(s) ` +
+ `= ~${budget.toLocaleString("en-US")} pairs, over the ${MAX_CONTAINMENT_PAIRS.toLocaleString("en-US")} budget. ` +
+ `Clip-of-longer duplicates are NOT covered by this run.`,
+ );
+ } else {
+ const shortFps = await fingerprintBlock(shorts, new Map());
+ for (const [slug, fp] of shortFps) {
+ hasTranscriptBySlug.set(slug, fp.hasTranscript);
+ }
+ const shortSlugs = new Set(shorts.map((e) => e.stat.slug));
+ let swept = 0;
for (const v of participants) {
- if (v === s) continue;
- if (v.stat.duration >= s.stat.duration * 1.5)
- candidates.push({ a: s, b: v, containmentOnly: true });
+ opts.signal?.throwIfAborted();
+ // One long video at a time: fingerprint, compare against every short,
+ // release. O(shorts + 1) fingerprints held.
+ const relevant = shorts.filter(
+ (s) =>
+ s.stat.slug !== v.stat.slug &&
+ v.stat.duration >= s.stat.duration * 1.5,
+ );
+ if (relevant.length === 0) continue;
+ const fv = shortSlugs.has(v.stat.slug)
+ ? shortFps.get(v.stat.slug)
+ : (await fingerprintBlock([v], new Map())).get(v.stat.slug);
+ if (!fv) continue;
+ hasTranscriptBySlug.set(v.stat.slug, fv.hasTranscript);
+ for (const s of relevant) {
+ const fs = shortFps.get(s.stat.slug);
+ if (!fs) continue;
+ swept++;
+ stats.evaluated++;
+ const res = evalContainment(fs, fv, containmentThreshold);
+ if (!res) {
+ stats.rejected++;
+ continue;
+ }
+ stats.confirmed++;
+ matches.push({ a: s.stat.slug, b: v.stat.slug, ...res });
+ }
+ noteRss();
}
+ stats.nominated += swept;
+ log(`Containment sweep: ${swept} pair(s) evaluated.`);
}
}
- log(
- `Phase 1: ${buckets.size} duration buckets → ${candidates.length} candidate pairs.`,
- );
-
- // ---- Build fingerprints for every video in a candidate ------------------
- // Primary source is the LMDB `cues` sub-db. But parseVtt only extracts text
- // from YouTube karaoke-tagged cues, so plain-VTT transcripts (most non-YouTube
- // captions, whisper-as-vtt, manual subs) store as zero cues and would look
- // transcript-less. For those, fall back to reading the raw transcript file off
- // disk and extracting plain text — for comparison only, not stored anywhere.
- const needed = new Map<string, Entry>();
- for (const c of candidates) {
- needed.set(c.a.stat.slug, c.a);
- needed.set(c.b.stat.slug, c.b);
- }
- const fingerprints = new Map<string, Fingerprint>();
- const rawFallbacks: Entry[] = [];
- for (const e of needed.values()) {
- opts.signal?.throwIfAborted();
- const s = e.stat;
- const cues = cuesDb.get([s.uploadDate, s.channelSlug, s.id]);
- if (cues && cues.length > 0) {
- fingerprints.set(s.slug, fingerprintFrom(s, cues, shingleSize));
- } else {
- rawFallbacks.push(e);
- }
- }
- await root.close();
- let rawHits = 0;
- if (rawFallbacks.length > 0) {
- const limit = pLimit(16);
- await Promise.all(
- rawFallbacks.map((e) =>
- limit(async () => {
- opts.signal?.throwIfAborted();
- const cues = await readRawCues(opts.paths.channelsDir, e);
- if (cues && cues.length > 0) rawHits++;
- fingerprints.set(e.stat.slug, fingerprintFrom(e.stat, cues, shingleSize));
- }),
- ),
- );
- }
log(
- `Phase 2: fingerprinted ${fingerprints.size} candidate videos ` +
- `(${rawFallbacks.length} missing LMDB cues; ${rawHits} recovered from raw transcripts).`,
+ `Phase 2: ${stats.evaluated} nominated pair(s) tested → ${stats.confirmed} confirmed, ` +
+ `${stats.rejected} rejected by content, ${stats.suspects} untestable suspect(s). ` +
+ `Fingerprinted ${stats.fingerprinted} video(s) (${stats.rawFallbacks} missing LMDB cues; ` +
+ `${stats.rawHits} recovered from raw transcripts).`,
);
-
- // ---- Phase 2: evaluate each candidate pair -------------------------------
- const matches: PairResult[] = [];
- for (const c of candidates) {
- opts.signal?.throwIfAborted();
- const fa = fingerprints.get(c.a.stat.slug) as Fingerprint;
- const fb = fingerprints.get(c.b.stat.slug) as Fingerprint;
- const res = c.containmentOnly
- ? evalContainment(fa, fb, containmentThreshold)
- : evalSimilar(fa, fb, nearThreshold);
- if (res) matches.push({ a: c.a.stat.slug, b: c.b.stat.slug, ...res });
- }
+ prevFps = new Map();
// ---- Cluster matched pairs (union-find) ----------------------------------
const uf = new UnionFind();
@@ -289,15 +525,26 @@ export async function detectDuplicateShorts(
}
}
+ // Refs are rebuilt from the participant index, NOT from fingerprints — those
+ // were released with their block. VideoStat is small metadata and the index is
+ // already held for the whole run, so this costs nothing.
+ const bySlug = new Map<string, Entry>();
+ for (const e of participants) bySlug.set(e.stat.slug, e);
+
const clusters: DuplicateCluster[] = components.map((slugs) => {
const agg = aggByRoot.get(slugs[0]) as Agg;
const refs = slugs
- .map((slug) => toRef(fingerprints.get(slug) as Fingerprint))
+ .map((slug) =>
+ toRef(
+ bySlug.get(slug) as Entry,
+ hasTranscriptBySlug.get(slug) ?? false,
+ ),
+ )
.sort((x, y) => x.slug.localeCompare(y.slug));
const platforms = new Set(refs.map((r) => r.platform));
const channels = new Set(refs.map((r) => r.channelSlug));
const minDuration = Math.min(...refs.map((r) => r.duration));
- return {
+ const cluster: DuplicateCluster = {
clusterId: sha1(refs.map((r) => r.slug).join("\n")),
matchKind: agg.matchKind,
score: agg.score,
@@ -305,13 +552,71 @@ export async function detectDuplicateShorts(
durationBucket: Math.round(minDuration),
crossPlatform: platforms.size > 1,
crossChannel: channels.size > 1,
+ // Nothing compared this cluster's CONTENT — every pair in it was
+ // untestable. It is a suspect for human review and shares nothing until
+ // someone records `confirmed`. Written only when true so that a
+ // content-confirmed cluster serialises exactly as it did before.
+ ...(agg.matchKind === "title-duration" ? { needsReview: true } : {}),
videoRefs: refs,
};
+ // Record the rule's choice of canonical member: the one derived work (an AI
+ // digest today, attribution later) is generated for and shared FROM. A human
+ // can override it in duplicates.overrides.json — which is a separate file
+ // precisely because this report is rewritten wholesale on every run.
+ return { ...cluster, canonicalSlug: pickCanonicalSlug(cluster) };
});
+ // ---- Measure timing alignment against each cluster's canonical member -----
+ //
+ // Content similarity says NOTHING about timing: a mirror with a longer intro
+ // matches on text at shifted times. Anything that seeks into a sibling — a
+ // shared digest's chapters, the search-result "jump to this moment" — is wrong
+ // without this, and wrong in the way that looks right. digestSharing measured
+ // it already but threw the result away; persisting it is what lets the viewer
+ // be honest about when a jump is trustworthy.
+ //
+ // Confirmed clusters only, and matches are few by construction, so re-reading
+ // those cue lists is cheap. Suspects are skipped: they share nothing anyway.
+ let aligned = 0;
+ let alignmentPairs = 0;
+ for (const cluster of clusters) {
+ opts.signal?.throwIfAborted();
+ if (cluster.needsReview || cluster.contained) continue;
+ const canonicalSlug = cluster.canonicalSlug;
+ if (!canonicalSlug) continue;
+ const canonicalCues = timedCues(bySlug.get(canonicalSlug));
+ if (!canonicalCues) continue;
+ for (const ref of cluster.videoRefs) {
+ if (ref.slug === canonicalSlug) {
+ ref.offsetSeconds = 0;
+ ref.aligned = true;
+ continue;
+ }
+ const cues = timedCues(bySlug.get(ref.slug));
+ if (!cues) continue;
+ alignmentPairs++;
+ const a = measureAlignment(canonicalCues, cues, {
+ toleranceSeconds: alignmentTolerance,
+ });
+ ref.aligned = a.aligned;
+ ref.offsetSeconds = Number.isFinite(a.maxOffsetSeconds)
+ ? Math.round(a.maxOffsetSeconds * 100) / 100
+ : null;
+ if (a.aligned) aligned++;
+ }
+ }
+ if (alignmentPairs > 0) {
+ log(
+ `Alignment: ${aligned}/${alignmentPairs} mirror(s) within ${alignmentTolerance}s of their canonical member.`,
+ );
+ }
+ await root.close();
+
const rank: Record<DuplicateMatchKind, number> = {
"transcript-exact": 2,
"transcript-near": 1,
+ // Suspects sort last: they are a review queue, not a result.
+ "title-duration": 0,
};
clusters.sort(
(a, b) =>
@@ -335,8 +640,15 @@ export async function detectDuplicateShorts(
const tmp = `${outPath}.tmp-${process.pid}`;
await writeFile(tmp, JSON.stringify(report));
await rename(tmp, outPath);
+ noteRss();
+ const suspectClusters = clusters.filter((c) => c.needsReview).length;
log(
- `Done: ${clusters.length} duplicate cluster(s) over ${videosInClusters} video(s) → ${DUPLICATES_FILENAME}.`,
+ `Done: ${clusters.length} duplicate cluster(s) over ${videosInClusters} video(s) ` +
+ `(${suspectClusters} awaiting review) → ${DUPLICATES_FILENAME} ` +
+ `in ${Math.round((Date.now() - startedAt) / 1000)}s ` +
+ `(blocking=${blocking}, ${videosScanned} scanned, ${stats.nominated} nominated pair(s), ` +
+ `${stats.confirmed} confirmed, ${stats.rejected} rejected by content, ${stats.suspects} suspect(s), ` +
+ `peak RSS ${stats.peakRssMb} MB).`,
);
return report;
}
@@ -357,6 +669,261 @@ export async function readDuplicateReport(
}
}
+// ---------------------------------------------------------------------------
+// Human review decisions (canonical choice / not-a-duplicate)
+// ---------------------------------------------------------------------------
+
+export function duplicateOverridesPath(paths: Paths): string {
+ return path.join(paths.transcriptsDir, DUPLICATE_OVERRIDES_FILENAME);
+}
+
+// Never throws: an unreadable or malformed overrides file reads as "no decisions
+// recorded", so the duplicates page still renders.
+export async function readDuplicateOverrides(
+ paths: Paths,
+): Promise<DuplicateOverrides> {
+ try {
+ const raw = await readFile(duplicateOverridesPath(paths), "utf8");
+ return sanitizeDuplicateOverrides(JSON.parse(raw));
+ } catch {
+ return sanitizeDuplicateOverrides(null);
+ }
+}
+
+// Record one cluster's decision, preserving every other cluster's. Read-modify-
+// write with the same atomic tmp+rename the report itself uses.
+export async function updateDuplicateOverride(
+ paths: Paths,
+ clusterId: string,
+ patch: {
+ canonicalSlug?: string;
+ notDuplicate?: boolean;
+ // "I looked, and these really are the same video." The ONLY thing that lets
+ // a needsReview (title+duration) cluster share derived work.
+ confirmed?: boolean;
+ note?: string;
+ },
+): Promise<DuplicateOverrides> {
+ const current = await readDuplicateOverrides(paths);
+ const existing = current.clusters[clusterId] ?? {};
+ const next = {
+ ...existing,
+ ...(patch.canonicalSlug !== undefined
+ ? { canonicalSlug: patch.canonicalSlug }
+ : {}),
+ ...(patch.notDuplicate !== undefined
+ ? { notDuplicate: patch.notDuplicate }
+ : {}),
+ ...(patch.confirmed !== undefined ? { confirmed: patch.confirmed } : {}),
+ ...(patch.note !== undefined ? { note: patch.note } : {}),
+ decidedAt: new Date().toISOString(),
+ };
+ // An empty patch clears the decision (back to "awaiting review") rather than
+ // leaving a decidedAt-only stub that would read as reviewed. `confirmed` has
+ // to be part of that test: without it, clearing a canonical choice on a
+ // confirmed suspect would DELETE the confirmation and silently un-share the
+ // cluster's derived work.
+ if (!next.canonicalSlug && next.notDuplicate !== true && next.confirmed !== true) {
+ delete current.clusters[clusterId];
+ } else {
+ current.clusters[clusterId] = next;
+ }
+ const out = sanitizeDuplicateOverrides(current);
+ const file = duplicateOverridesPath(paths);
+ await mkdir(path.dirname(file), { recursive: true });
+ const tmp = `${file}.tmp-${process.pid}`;
+ await writeFile(tmp, JSON.stringify(out, null, 2) + "\n");
+ await rename(tmp, file);
+ return out;
+}
+
+// ---------------------------------------------------------------------------
+// Blocking strategies — they NOMINATE, they never decide
+// ---------------------------------------------------------------------------
+
+// Counters the generators fill in as they run, so the caller can report what was
+// covered AND what was skipped. Skipping without reporting reads as "covered
+// everything" when it did not.
+type BlockStats = {
+ titleGroups: number;
+ oversizedGroups: number;
+ oversizedGroupVideos: number;
+ oversizedBlocks: number;
+ oversizedBlockVideos: number;
+ untitled: number;
+};
+
+// TITLE BLOCKING. One pass to group by exact normalized title, then pair within
+// each group when the runtimes agree. Quadratic only INSIDE a group, and a group
+// is a handful of videos, so the whole pass is effectively linear.
+//
+// The key is deliberately EXACT rather than fuzzy: a looser key merges "Episode
+// 12" with "Episode 13", and every such merge is a false cluster a human then
+// has to reject. The cost is recall on re-titled mirrors, which duration
+// blocking covers instead — see STATE.md for the deferred middle grounds.
+function* titleBlocks(
+ participants: Entry[],
+ ratio: number,
+ stats: BlockStats,
+ log: (msg: string) => void,
+): Generator<Block> {
+ const byTitle = new Map<string, Entry[]>();
+ for (const e of participants) {
+ const key = normalizeTitleKey(e.stat.title ?? "");
+ if (key.length < MIN_TITLE_KEY_LENGTH) {
+ stats.untitled++;
+ continue;
+ }
+ const arr = byTitle.get(key);
+ if (arr) arr.push(e);
+ else byTitle.set(key, [e]);
+ }
+
+ // Decide the whole work list BEFORE yielding any of it, so the summary — and
+ // in particular what was SKIPPED — is reported up front rather than after the
+ // long evaluation it describes.
+ const accepted: Entry[][] = [];
+ for (const [key, group] of byTitle) {
+ if (group.length < 2) continue;
+ if (group.length > MAX_TITLE_GROUP_SIZE) {
+ // A FORMAT, not a title — "live stream", "untitled", a daily show's
+ // date-less name. Reported, never silently truncated.
+ stats.oversizedGroups++;
+ stats.oversizedGroupVideos += group.length;
+ log(
+ ` skipping title group of ${group.length} (over ${MAX_TITLE_GROUP_SIZE}): "${key.slice(0, 60)}"`,
+ );
+ continue;
+ }
+ accepted.push(group);
+ }
+ log(
+ `Title blocking: ${byTitle.size} distinct title(s), ${accepted.length} group(s) with 2+ members ` +
+ `(${stats.untitled} video(s) with no usable title; ${stats.oversizedGroups} group(s) over ` +
+ `${MAX_TITLE_GROUP_SIZE} skipped, covering ${stats.oversizedGroupVideos} video(s)).`,
+ );
+
+ for (const group of accepted) {
+ const nominations: Nomination[] = [];
+ for (let i = 0; i < group.length; i++) {
+ for (let j = i + 1; j < group.length; j++) {
+ if (
+ durationsCompatible(
+ group[i].stat.duration,
+ group[j].stat.duration,
+ ratio,
+ )
+ ) {
+ nominations.push({
+ a: group[i],
+ b: group[j],
+ containmentOnly: false,
+ // Same title AND a near-identical runtime: strong enough to be
+ // worth a human's attention when no transcript can settle it.
+ suspectKind: "title-duration",
+ });
+ }
+ }
+ }
+ if (nominations.length === 0) continue;
+ stats.titleGroups++;
+ yield { members: membersOf(nominations), nominations };
+ }
+}
+
+// DURATION BLOCKING. One block per rounded-duration bucket, unioned with the
+// next so a pair straddling a bucket edge is still nominated. Emitting
+// here×here and here×next (never next×next) makes every pair appear exactly
+// once, and yielding bucket b+1's members as part of block b is what lets the
+// caller's 2-block fingerprint cache halve the transcript reads.
+//
+// This is the only strategy that catches a RE-TITLED mirror, and it is viable
+// corpus-wide only because blocks are evaluated and released one at a time.
+function* durationBlocks(
+ participants: Entry[],
+ W: number,
+ stats: BlockStats,
+ log: (msg: string) => void,
+): Generator<Block> {
+ const bucketOf = (d: number) => Math.round(d / W);
+ const buckets = new Map<number, Entry[]>();
+ for (const e of participants) {
+ const b = bucketOf(e.stat.duration);
+ const arr = buckets.get(b);
+ if (arr) arr.push(e);
+ else buckets.set(b, [e]);
+ }
+ const maxDelta = (d: number) => Math.max(W, Math.ceil(0.02 * d));
+ const sorted = [...buckets.keys()].sort((x, y) => x - y);
+
+ // As with title blocking: decide the work list, report it (skips included),
+ // then do it. A summary that only arrives after an hour of evaluation is not a
+ // summary anyone can act on.
+ const accepted: number[] = [];
+ for (const b of sorted) {
+ const size =
+ (buckets.get(b) as Entry[]).length + (buckets.get(b + 1)?.length ?? 0);
+ // A bucket can itself be pathological — round numbers attract videos, and a
+ // corpus can hold thousands that are exactly 60 s. Same cap-and-report
+ // treatment as an oversized title group.
+ if (size > MAX_DURATION_BLOCK_SIZE) {
+ stats.oversizedBlocks++;
+ stats.oversizedBlockVideos += (buckets.get(b) as Entry[]).length;
+ log(
+ ` skipping duration block ~${b * W}s of ${size} video(s) ` +
+ `(over ${MAX_DURATION_BLOCK_SIZE}).`,
+ );
+ continue;
+ }
+ accepted.push(b);
+ }
+ log(
+ `Duration blocking: ${buckets.size} bucket(s) of ${W}s, ${accepted.length} block(s) to evaluate` +
+ (stats.oversizedBlocks > 0
+ ? ` (${stats.oversizedBlocks} block(s) over ${MAX_DURATION_BLOCK_SIZE} skipped, ` +
+ `covering ${stats.oversizedBlockVideos} video(s))`
+ : "") +
+ ".",
+ );
+
+ for (const b of accepted) {
+ const here = buckets.get(b) as Entry[];
+ const next = buckets.get(b + 1) ?? [];
+ const nominations: Nomination[] = [];
+ for (let i = 0; i < here.length; i++) {
+ for (let j = i + 1; j < here.length; j++) {
+ if (durationsClose(here[i], here[j], maxDelta)) {
+ // A shared duration ALONE is not evidence — it produced enormous
+ // false clusters of unrelated same-length videos. If the content
+ // cannot be compared, this pair is worth nothing.
+ nominations.push({
+ a: here[i],
+ b: here[j],
+ containmentOnly: false,
+ suspectKind: null,
+ });
+ }
+ }
+ }
+ for (const x of here) {
+ for (const y of next) {
+ if (durationsClose(x, y, maxDelta)) {
+ nominations.push({
+ a: x,
+ b: y,
+ containmentOnly: false,
+ suspectKind: null,
+ });
+ }
+ }
+ }
+ if (nominations.length === 0) continue;
+ // Members are here ∪ next, so the caller's sliding cache carries bucket b+1
+ // straight into the next block.
+ yield { members: membersOf(nominations), nominations };
+ }
+}
+
// --- helpers ---------------------------------------------------------------
function durationsClose(
@@ -469,6 +1036,16 @@ function evalSimilar(
if (a.hash && a.hash === b.hash) {
return { kind: "transcript-exact", score: 1, contained: false };
}
+ // Size prune. Exact, not a heuristic: |A∩B| ≤ min and |A∪B| ≥ max, so
+ // J ≤ min/max. A pair whose shingle counts are further apart than the
+ // threshold CANNOT reach it, and skipping it changes no result. Worth having
+ // because it is O(1) where the Jaccard it replaces is O(min set size), and
+ // corpus-wide duration blocking evaluates millions of pairs whose members
+ // share a runtime but not a word count (a music video and a lecture can both
+ // be ten minutes long).
+ const small = Math.min(a.shingleSet.size, b.shingleSet.size);
+ const large = Math.max(a.shingleSet.size, b.shingleSet.size);
+ if (large === 0 || small / large < nearThreshold) return null;
const j = jaccard(a.shingleSet, b.shingleSet);
if (j >= nearThreshold) {
return { kind: "transcript-near", score: round3(j), contained: false };
@@ -476,6 +1053,46 @@ function evalSimilar(
return null;
}
+// The verdict on one nominated pair — the point where the pre-filter's proposal
+// meets the transcript's disposal. Three outcomes, and the middle one is the
+// reason nominating aggressively is safe:
+//
+// both transcripts, content agrees → confirmed (may share)
+// both transcripts, content differs → null. REJECTED. Two episodes of a daily
+// show can share a title and a runtime and
+// be entirely different material; testing
+// them is what keeps that out.
+// a transcript is missing → nothing can compare the content, so the
+// pre-filter's own claim is all there is.
+// Kept as a needsReview suspect when that
+// claim is strong (title + runtime), and
+// dropped when it is not (duration alone).
+function evalBlocked(
+ a: Fingerprint,
+ b: Fingerprint,
+ opts: {
+ nearThreshold: number;
+ containmentThreshold: number;
+ allowContainment: boolean;
+ suspectKind: DuplicateMatchKind | null;
+ },
+): Omit<PairResult, "a" | "b"> | null {
+ if (a.hasTranscript && b.hasTranscript) {
+ const similar = evalSimilar(a, b, opts.nearThreshold);
+ if (similar) return similar;
+ // Containment as a FALLBACK inside a pair we already nominated: it costs one
+ // more set intersection over shingle sets that are already in hand, so it is
+ // free relative to the nomination that got us here.
+ if (opts.allowContainment) {
+ const contained = evalContainment(a, b, opts.containmentThreshold);
+ if (contained) return contained;
+ }
+ return null;
+ }
+ if (!opts.suspectKind) return null;
+ return { kind: opts.suspectKind, score: null, contained: false };
+}
+
// Containment pair: the shorter transcript's shingles are largely a subset of
// the longer one's. Requires both transcripts (no metadata fallback here).
function evalContainment(
@@ -491,18 +1108,25 @@ function evalContainment(
return null;
}
-function toRef(fp: Fingerprint): DuplicateVideoRef {
- const s = fp.stat;
+// Built from the participant index rather than a fingerprint: fingerprints are
+// released with their block, and VideoStat is the small half of what they held.
+// `offsetSeconds` / `aligned` are filled in afterwards by the alignment pass.
+function toRef(e: Entry, hasTranscript: boolean): DuplicateVideoRef {
+ const s = e.stat;
return {
slug: s.slug,
channelSlug: s.channelSlug,
channel: s.channel,
platform: s.platform,
id: s.id,
+ // Only when it differs — see DuplicateVideoRef.videoDir. The detector is the
+ // one place that has both halves in hand, and dropping the directory here is
+ // what silently broke digest cluster-sharing for every Rumble mirror.
+ ...(e.videoDir !== s.id ? { videoDir: e.videoDir } : {}),
title: s.title,
duration: s.duration,
uploadDate: s.uploadDate,
- hasTranscript: fp.hasTranscript,
+ hasTranscript,
};
}
diff --git a/common/jobs/jobKinds.ts b/common/jobs/jobKinds.ts
@@ -99,6 +99,37 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
bookmarkable: true,
queueKeyStrategy: "custom",
},
+ // AI digest sweep, local (ollama) lane — the one that carries the corpus. Both
+ // digest kinds are drainable (the batch honors the drain signal: it stops
+ // pulling new videos and lets the in-flight one finish) and bookmarkable, since
+ // a channel-scoped sweep is exactly the kind of thing an operator re-launches.
+ "digest-channel-local": {
+ kind: "digest-channel-local",
+ label: "Digest channel (local)",
+ drainable: true,
+ bookmarkable: true,
+ queueKeyStrategy: "custom",
+ },
+ // Same batch, metered lane. Off unless settings.digest.remoteEnabled is true,
+ // and it lands on its own queue key so it runs CONCURRENTLY with the local lane
+ // rather than behind it.
+ "digest-channel-remote": {
+ kind: "digest-channel-remote",
+ label: "Digest channel (metered)",
+ drainable: true,
+ bookmarkable: true,
+ queueKeyStrategy: "custom",
+ },
+ // Copy a duplicate cluster's canonical digest onto its aligned mirrors. A fast
+ // file operation gated by the timestamp-alignment check, so it is not drainable
+ // but is bookmarkable.
+ "digest-share-cluster": {
+ kind: "digest-share-cluster",
+ label: "Share cluster digest",
+ drainable: false,
+ bookmarkable: false,
+ queueKeyStrategy: "parallel",
+ },
"redownload-incomplete-bucket": {
kind: "redownload-incomplete-bucket",
label: "Re-download truncated transcripts",
diff --git a/common/jobs/registry.ts b/common/jobs/registry.ts
@@ -11,15 +11,37 @@ export type JobStatus =
| "failed"
| "cancelled";
-export type JobProgressMetric = "downloads" | "transcripts";
+// Extended for the digest sweep. NOTE: this union has one re-spelled copy in
+// RunningJobsList.tsx's RunningJobsListItem — kept as an IMPORT there now, so
+// TypeScript actually flags the next member added here.
+export type JobProgressMetric = "downloads" | "transcripts" | "digests";
export type JobProgress = {
metric: JobProgressMetric;
initial: number;
target: number;
+ // Progress as counted BY THE RUNNER, when the runner knows better than the
+ // disk does.
+ //
+ // The default is to re-count `current` from on-disk channel stats, which is
+ // right for downloads and transcripts: a file appears, the count goes up.
+ // It is WRONG for digests under regeneration — a regenerated digest is
+ // rewritten in place, so the file count never moves and the bar sits at 0%
+ // for the whole job (observed: {initial:39, current:39, target:72, pct:0}).
+ // Cosmetic for a five-minute job; over an 81-day sweep it makes working work
+ // look wedged.
+ //
+ // Only set it where the runner has a genuinely better number. Absent → the
+ // disk re-count, unchanged.
+ current?: number;
+ // Audio-seconds still to process. Paired with completedTaskAudioSeconds it
+ // gives an ETA in the unit the work is actually priced in; without it the
+ // estimate falls back to averaging TASKS, which for digests is wrong by more
+ // than an order of magnitude between a VOD channel and a shorts channel.
+ remainingAudioSeconds?: number;
};
-export type JobTaskKind = "download" | "transcribe";
+export type JobTaskKind = "download" | "transcribe" | "digest";
// A single in-flight sub-operation within a job (one video download or one
// transcription). Only currently-running tasks are kept on the record — they
@@ -80,6 +102,16 @@ export type JobRecord = {
// reflects the whole batch. Both undefined until the first task completes.
completedTaskCount?: number;
completedTaskMs?: number;
+ // Audio-seconds covered by completed sub-operations. Only digest/transcribe
+ // tasks report it, because only they consume work proportional to a video's
+ // LENGTH — a digest of a 3-hour VOD is not one task's worth of anything.
+ //
+ // It is what makes a sweep ETA expressible: `completedTaskMs` over this gives
+ // seconds-per-audio-hour, the unit the whole backfill is estimated in, and
+ // the only unit in which "how long is this going to take" has an answer. A
+ // task-count average cannot express it — the corpus is 77k videos and 77k
+ // audio-hours, and those distribute completely differently.
+ completedTaskAudioSeconds?: number;
};
export type QueueSnapshot = {
@@ -302,11 +334,19 @@ class JobRegistry {
// Fold one finished sub-operation's wall-clock duration into the job's
// running totals (see JobRecord.completedTaskCount/Ms). No-op if the job is
// gone or the duration is nonsensical.
- recordTaskDuration(jobId: string, durationMs: number): void {
+ recordTaskDuration(
+ jobId: string,
+ durationMs: number,
+ audioSeconds?: number,
+ ): void {
const job = this.jobs.get(jobId);
if (!job || durationMs < 0) return;
job.completedTaskCount = (job.completedTaskCount ?? 0) + 1;
job.completedTaskMs = (job.completedTaskMs ?? 0) + durationMs;
+ if (typeof audioSeconds === "number" && audioSeconds > 0) {
+ job.completedTaskAudioSeconds =
+ (job.completedTaskAudioSeconds ?? 0) + audioSeconds;
+ }
}
// Snapshot of queues for UI, resolved from the scheduler's id views back to
diff --git a/common/jobs/snapshotScheduler.ts b/common/jobs/snapshotScheduler.ts
@@ -27,7 +27,20 @@ import { drainStream } from "./drainStream";
// — including it would make the central runManagedFunction hook re-arm the timer
// from inside the regen job, an infinite loop. "detect-duplicates" is a global
// read-only scan with no per-channel snapshot impact.
-const NO_REGEN_KINDS = new Set<string>(["refresh-report", "detect-duplicates"]);
+//
+// The digest kinds are here for a different, load-bearing reason: SCALE.
+// ctx.recordTaskDone arms a channel-snapshot regen after EVERY completed
+// sub-operation, and each regen is a full 16-way per-video fan-out over the whole
+// channel. A digest sweep is ~119,600 sub-operations corpus-wide, so leaving them
+// in would spend more machine time regenerating snapshots than generating
+// digests. The digest actions instead regenerate ONCE per channel, at job end.
+const NO_REGEN_KINDS = new Set<string>([
+ "refresh-report",
+ "detect-duplicates",
+ "digest-channel-local",
+ "digest-channel-remote",
+ "digest-share-cluster",
+]);
export function shouldRequestSnapshot(kind: string): boolean {
return !NO_REGEN_KINDS.has(kind);
diff --git a/common/jobs/streamCommand.ts b/common/jobs/streamCommand.ts
@@ -91,7 +91,7 @@ export type JobRunContext = {
},
) => void;
removeTask: (taskId: string) => void;
- recordTaskDone: (durationMs: number) => void;
+ recordTaskDone: (durationMs: number, audioSeconds?: number) => void;
};
export type RunManagedFunctionOpts = CommonOpts & {
@@ -324,8 +324,8 @@ export async function runManagedFunction(
addTask: (task) => registry.addTask(id, task),
updateTask: (taskId, patch) => registry.updateTask(id, taskId, patch),
removeTask: (taskId) => registry.removeTask(id, taskId),
- recordTaskDone: (ms) => {
- registry.recordTaskDuration(id, ms);
+ recordTaskDone: (ms, audioSeconds) => {
+ registry.recordTaskDuration(id, ms, audioSeconds);
// Refresh the report after EACH completed sub-operation (each video
// downloaded/transcribed in a batch), not only when the whole batch
// finishes — so a long batch updates incrementally. The global debounce
diff --git a/common/jobs/taskHooks.ts b/common/jobs/taskHooks.ts
@@ -28,6 +28,11 @@ export type TaskTracker = {
appId?: string;
// The worker that owns this task, surfaced on the Workers page.
workerId?: string;
+ // The video's duration. Recorded against the task's measured wall clock so
+ // the job can report SECONDS PER AUDIO-HOUR — the unit a digest sweep is
+ // estimated in, and one a task-count average cannot express. Omit where the
+ // work is not proportional to length (a download is bytes, not minutes).
+ audioSeconds?: number;
}) => TaskHandle;
};
@@ -44,7 +49,7 @@ export function makeTaskTracker(
forwardLog: (line: string) => void,
): TaskTracker {
return {
- start({ id, label, kind, appId, workerId }) {
+ start({ id, label, kind, appId, workerId, audioSeconds }) {
if (!ctx) {
return { onLog: forwardLog, update: () => {}, end: () => {} };
}
@@ -63,7 +68,16 @@ export function makeTaskTracker(
// (DLOM_PROBE) drive the probe phase and probe-aware ETA. Transcribe
// tasks without an appId (remote workers stream pre-parsed progress) fall
// through to it too — harmless for non-matching lines.
- const downloadParser = transcribeParser ? null : createDownloadProgressParser();
+ //
+ // Digest tasks get NO parser: a digest's per-chunk progress is discrete
+ // (chunk 3 of 7), not a byte/second rate, and its log lines carry model
+ // token counts that the download parser would happily misread as a
+ // percentage. The task row shows an indeterminate bar instead, which is
+ // honest.
+ const downloadParser =
+ transcribeParser || kind === "digest"
+ ? null
+ : createDownloadProgressParser();
let ended = false;
const onLog = (line: string) => {
forwardLog(line);
@@ -71,9 +85,14 @@ export function makeTaskTracker(
// split on both \r and \n to see each discrete update.
for (const part of line.split(/[\r\n]+/)) {
if (!part) continue;
+ // Both parsers are null for a digest task (see above), so this must
+ // tolerate having neither. The `!` this replaces was a lie that threw
+ // on the FIRST log line of every digest job — the batch caught it as a
+ // per-video failure, so a whole sweep would have reported "0
+ // generated, N failed" while looking like an engine problem.
const update = transcribeParser
? transcribeParser.feed(part)
- : downloadParser!.feed(part);
+ : downloadParser?.feed(part);
if (update) ctx.updateTask(id, update);
}
};
@@ -83,7 +102,7 @@ export function makeTaskTracker(
const end = () => {
if (ended) return;
ended = true;
- ctx.recordTaskDone(Date.now() - startedAt);
+ ctx.recordTaskDone(Date.now() - startedAt, audioSeconds);
ctx.removeTask(id);
};
return { onLog, update, end };
diff --git a/common/lib/channelSignature.ts b/common/lib/channelSignature.ts
@@ -20,10 +20,15 @@ import { existsSync } from "node:fs";
import { open } from "lmdb";
import type { Paths } from "./paths";
+// Structural subset of buildIndex.ts's MtimeRecord — only the fields that feed
+// the signature. `digestMs` is the newest mtime across BOTH digest sidecars
+// (machine + human overrides); without it in the hash below, a digest-only
+// change would leave every archive believing the channel was unchanged.
type MtimeRecord = {
metaMs: number;
transcriptMs: number | null;
subsMs: number | null;
+ digestMs: number | null;
};
type PathKey = [string, string];
@@ -74,7 +79,7 @@ export function openChannelSigner(paths: Paths): ChannelSigner {
if (k[0] !== slug) break;
sawAny = true;
h.update(
- `${k[1]}\t${value.metaMs}\t${value.transcriptMs ?? ""}\t${value.subsMs ?? ""}\n`,
+ `${k[1]}\t${value.metaMs}\t${value.transcriptMs ?? ""}\t${value.subsMs ?? ""}\t${value.digestMs ?? ""}\n`,
);
}
// A channel with zero indexed videos still gets a stable signature (schema
diff --git a/common/lib/corpus.test.ts b/common/lib/corpus.test.ts
@@ -80,6 +80,48 @@ test("renderSiteLlmsTxt: title, corpus link, channels", () => {
assert.match(txt, /Bulk archives/);
});
+test("buildSiteCorpus: digest layer is advertised only where it exists", () => {
+ const corpus = buildSiteCorpus(descriptor(), {
+ hasArchives: false,
+ digestCounts: { alice: 7, bob: 0 },
+ });
+ // Advertised on the channel that has digests…
+ assert.equal(corpus.channels[0].digestCount, 7);
+ assert.equal(
+ corpus.channels[0].manifests.digests,
+ "/digests/alice/manifest.json",
+ );
+ // …and absent, not zero-valued, on the one that doesn't. A `digests` pointer
+ // to a manifest that was never written would send clients to a 404.
+ assert.equal(corpus.channels[1].digestCount, undefined);
+ assert.equal(corpus.channels[1].manifests.digests, undefined);
+ assert.ok(corpus.digestScheme, "some digests → digestScheme block");
+ // The sparsity contract is the part a client must not get wrong.
+ assert.match(corpus.digestScheme!.description, /SPARSE/);
+ assert.match(corpus.digestScheme!.chapterTiming, /SECONDS/);
+});
+
+test("buildSiteCorpus: no digests → no digestScheme, unchanged channels", () => {
+ const corpus = buildSiteCorpus(descriptor(), { hasArchives: false });
+ assert.equal(corpus.digestScheme, undefined);
+ assert.equal(corpus.channels[0].manifests.digests, undefined);
+ assert.equal(corpus.channels[0].digestCount, undefined);
+});
+
+test("renderSiteLlmsTxt: names the digest layer only when present", () => {
+ const withDigests = renderSiteLlmsTxt(
+ buildSiteCorpus(descriptor({ siteUrl: "https://demo.example" }), {
+ hasArchives: false,
+ digestCounts: { alice: 7 },
+ }),
+ );
+ assert.match(withDigests, /AI digests: 7 of these transcripts/);
+ const without = renderSiteLlmsTxt(
+ buildSiteCorpus(descriptor(), { hasArchives: false }),
+ );
+ assert.doesNotMatch(without, /AI digests/);
+});
+
test("buildHubCorpus: drops members with no siteUrl, links each corpus", () => {
const hub = buildHubCorpus(
[
diff --git a/common/lib/corpus.ts b/common/lib/corpus.ts
@@ -13,7 +13,9 @@ import type { PublicSiteDescriptor } from "./siteDescriptor";
// v2: the corpus now also describes the parallel social-post layer
// (postScheme + per-channel posts manifest pointers).
-export const CORPUS_SPEC_VERSION = 2;
+// v3: …and the DERIVED layer — AI digests (chapters + topic tags) served under
+// the same shard scheme (digestScheme + per-channel digests manifest pointers).
+export const CORPUS_SPEC_VERSION = 3;
// How to resolve a single transcript from the paginated shards, described once
// and embedded in every corpus.json so any HTTP client can navigate without
@@ -63,6 +65,39 @@ const POST_SCHEME = {
permalink: "each post carries its own canonical `url`; no timestamp fragment applies",
} as const;
+// The AI-digest corpus: a DERIVED layer over transcripts, not a parallel source
+// like posts. Same paginated-shard scheme, but sparse — a video absent from a
+// digests manifest simply has not been digested, which is the normal case.
+const DIGEST_SCHEME = {
+ description:
+ "AI digests are chapters and topic tags DERIVED from a video's transcript " +
+ "by a local model, composed with any human corrections before publication. " +
+ "They are served as paginated JSON shards under the same scheme: (1) GET " +
+ "the channel's digests manifest; (2) look up the video id in its " +
+ "`slugToPage` map to get a page number N; (3) GET page-<NNNN>.json and take " +
+ "the record whose `id` matches. THE LAYER IS SPARSE: a video id absent from " +
+ "`slugToPage` has no digest, and a channel with no digests has no manifest " +
+ "at all. Absence means 'not yet generated', never 'nothing to say'.",
+ digestsManifest:
+ "<channel.manifests.digests> -> { pageCount, slugToPage: { <videoId>: <pageNumber> } }",
+ digestPage:
+ "/digests/<slug>/page-<NNNN>.json -> array of { id, slug, generatedAt, " +
+ "chapters: [{ id, start, clock, title, decidedBy }], tags: [{ id, tag, " +
+ "decidedBy }], provenance, derivedFrom? }",
+ chapterTiming:
+ "`start` is SECONDS, already snapped to a real transcript cue boundary — use " +
+ "it to seek. `clock` is the raw HH:MM:SS string the model emitted, kept for " +
+ "auditing; do not parse it for timing.",
+ attribution:
+ "`decidedBy` is \"ai\" or \"human\" per item, so machine output and human " +
+ "corrections stay distinguishable after composition.",
+ derivedFrom:
+ "When present, this digest was generated for a DIFFERENT video (the canonical " +
+ "member of a duplicate cluster) and shared onto this one; it names that video " +
+ "and the measured cue-timing offset in seconds. Treat its chapter titles as " +
+ "describing the canonical upload.",
+} as const;
+
export type CorpusChannel = {
slug: string;
name: string;
@@ -73,9 +108,16 @@ export type CorpusChannel = {
subs: string;
// Present only for social channels (the posts corpus).
posts?: string;
+ // Present only for channels with at least one digested video.
+ digests?: string;
};
// Present only for social channels.
postCount?: number;
+ // Number of this channel's videos that carry a digest. Present only when
+ // non-zero, and deliberately reported ALONGSIDE videoCount rather than
+ // instead of it: the ratio is the coverage of the derived layer, which is
+ // what tells a client whether to expect a digest for an arbitrary video.
+ digestCount?: number;
};
export type SiteCorpus = {
@@ -94,6 +136,8 @@ export type SiteCorpus = {
shardScheme: typeof SHARD_SCHEME;
// Present when this site includes at least one social channel.
postScheme?: typeof POST_SCHEME;
+ // Present when this site ships at least one digested video.
+ digestScheme?: typeof DIGEST_SCHEME;
// Present when this build ships bulk-download archives (whole-channel zips).
bulkArchives?: { manifest: string; note: string };
// Pointer to the human page and BYO-key chat.
@@ -144,24 +188,33 @@ export function buildSiteCorpus(
// slug -> archived post count, for the social channels in this site. Absent
// / empty means the site has no posts corpus and postScheme is omitted.
postCounts?: Record<string, number>;
+ // slug -> digested video count. Absent / empty means the site ships no
+ // derived layer and digestScheme is omitted.
+ digestCounts?: Record<string, number>;
},
): SiteCorpus {
const base = descriptor.siteUrl;
const postCounts = opts.postCounts ?? {};
+ const digestCounts = opts.digestCounts ?? {};
const channels: CorpusChannel[] = descriptor.channels.map((c) => {
const postCount = postCounts[c.slug];
+ const digestCount = digestCounts[c.slug];
return {
slug: c.slug,
name: c.name,
videoCount: c.count,
...(c.groupId ? { groupId: c.groupId } : {}),
...(postCount ? { postCount } : {}),
+ ...(digestCount ? { digestCount } : {}),
manifests: {
transcripts: join(base, `/transcripts/${c.slug}/manifest.json`),
subs: join(base, `/subs/${c.slug}/manifest.json`),
...(postCount
? { posts: join(base, `/posts/${c.slug}/manifest.json`) }
: {}),
+ ...(digestCount
+ ? { digests: join(base, `/digests/${c.slug}/manifest.json`) }
+ : {}),
},
};
});
@@ -188,6 +241,11 @@ export function buildSiteCorpus(
if (Object.values(postCounts).some((n) => n > 0)) {
corpus.postScheme = POST_SCHEME;
}
+ // Likewise for the derived layer: a site with no digests is unchanged apart
+ // from the spec bump.
+ if (Object.values(digestCounts).some((n) => n > 0)) {
+ corpus.digestScheme = DIGEST_SCHEME;
+ }
if (opts.hasArchives) {
corpus.bulkArchives = {
manifest: join(base, "/archives/manifest.json"),
@@ -263,6 +321,17 @@ export function renderSiteLlmsTxt(corpus: SiteCorpus): string {
`- [corpus.json](${join(base, "/corpus.json")}): machine-readable index — ` +
`channels and how to fetch any transcript from the paginated JSON shards.`,
);
+ if (corpus.digestScheme) {
+ const digested = corpus.channels.reduce(
+ (n, c) => n + (c.digestCount ?? 0),
+ 0,
+ );
+ out.push(
+ `- AI digests: ${digested.toLocaleString()} of these transcripts also carry ` +
+ `machine-generated chapters and topic tags, served under /digests/ — ` +
+ `see corpus.json's digestScheme. Coverage is partial and growing.`,
+ );
+ }
if (corpus.bulkArchives) {
out.push(
`- [Bulk archives](${join(base, "/downloads")}): whole-channel transcript ` +
diff --git a/common/lib/digest-server.test.ts b/common/lib/digest-server.test.ts
@@ -0,0 +1,221 @@
+import { test } from "node:test";
+import assert from "node:assert/strict";
+import { mkdtemp, rm } from "node:fs/promises";
+import { tmpdir } from "node:os";
+import path from "node:path";
+import {
+ DIGEST_SCHEMA_VERSION,
+ isSectionFresh,
+ type DigestFreshnessTarget,
+ type DigestProvenance,
+} from "./digest";
+import {
+ loadDigest,
+ writeDigestFailure,
+ writeDigestSection,
+} from "./digest-server";
+
+// Run with:
+// pnpm --filter yt-dlp-transcript-common exec tsx --test common/lib/digest-server.test.ts
+
+async function withVideoDir(fn: (dir: string) => Promise<void>): Promise<void> {
+ const dir = await mkdtemp(path.join(tmpdir(), "ttb-digest-server-"));
+ try {
+ await fn(dir);
+ } finally {
+ await rm(dir, { recursive: true, force: true });
+ }
+}
+
+const PROVENANCE: DigestProvenance = {
+ appId: "ollama-direct",
+ model: "qwen2.5:7b",
+ modelRequested: "qwen2.5:7b",
+ lane: "local-gpu",
+ generatedAt: "2026-07-29T00:00:00.000Z",
+ promptVersion: 2,
+ contextHash: "ctx",
+ chunks: 3,
+ chunksOk: 3,
+};
+
+const TARGET: DigestFreshnessTarget = {
+ appId: "ollama-direct",
+ model: "qwen2.5:7b",
+ promptVersion: 2,
+ contextHash: "ctx",
+};
+
+// The whole point of `failures`: the two total-failure paths in digestVideo
+// deliberately write no section so the video RETRIES, which used to mean the
+// warnings from the worst outputs survived nowhere but a rotating job log.
+test("a recorded failure persists its warnings without making the video look fresh", async () => {
+ await withVideoDir(async (dir) => {
+ await writeDigestFailure(dir, {
+ section: "chapters",
+ at: "2026-07-29T00:00:00.000Z",
+ appId: "ollama-direct",
+ model: "qwen2.5:7b",
+ promptVersion: 2,
+ reason: "all-rejected",
+ chunks: 3,
+ chunksOk: 3,
+ warnings: [
+ { code: "out-of-range", section: "chapters", detail: "13 clamped" },
+ ],
+ });
+
+ const record = await loadDigest(dir);
+ assert.equal(record?.failures?.length, 1);
+ assert.equal(record?.failures?.[0].reason, "all-rejected");
+ assert.equal(record?.failures?.[0].warnings.length, 1);
+ // The retry behaviour is the constraint this design exists to preserve.
+ assert.equal(record?.sections.chapters, undefined);
+ assert.equal(
+ isSectionFresh(record, "chapters", TARGET),
+ false,
+ "a failure must never read as fresh, or the video is never retried",
+ );
+ });
+});
+
+// "the model proposed nothing" and "the model proposed and every item was
+// rejected by a guard" look identical from outside and need different fixes.
+test("the two failure reasons are recorded distinctly", async () => {
+ await withVideoDir(async (dir) => {
+ await writeDigestFailure(dir, {
+ section: "chapters",
+ at: "2026-07-29T00:00:00.000Z",
+ appId: "a",
+ model: "m",
+ promptVersion: 2,
+ reason: "no-output",
+ chunks: 5,
+ chunksOk: 0,
+ warnings: [],
+ });
+ assert.equal((await loadDigest(dir))?.failures?.[0].reason, "no-output");
+ assert.equal((await loadDigest(dir))?.failures?.[0].chunksOk, 0);
+ });
+});
+
+// A sweep retries. Appending would grow the sidecar without bound on a video
+// that fails every single pass over 81 days.
+test("only the latest failure per section is kept", async () => {
+ await withVideoDir(async (dir) => {
+ for (const at of ["2026-07-01T00:00:00.000Z", "2026-07-29T00:00:00.000Z"]) {
+ await writeDigestFailure(dir, {
+ section: "chapters",
+ at,
+ appId: "a",
+ model: "m",
+ promptVersion: 2,
+ reason: "no-output",
+ chunks: 1,
+ chunksOk: 0,
+ warnings: [],
+ });
+ }
+ const record = await loadDigest(dir);
+ assert.equal(record?.failures?.length, 1);
+ assert.equal(record?.failures?.[0].at, "2026-07-29T00:00:00.000Z");
+ });
+});
+
+test("a failure in one section does not disturb another section's", async () => {
+ await withVideoDir(async (dir) => {
+ for (const section of ["chapters", "tags"] as const) {
+ await writeDigestFailure(dir, {
+ section,
+ at: "2026-07-29T00:00:00.000Z",
+ appId: "a",
+ model: "m",
+ promptVersion: 2,
+ reason: "no-output",
+ chunks: 1,
+ chunksOk: 0,
+ warnings: [],
+ });
+ }
+ const record = await loadDigest(dir);
+ assert.equal(record?.failures?.length, 2);
+ assert.deepEqual(
+ record?.failures?.map((f) => f.section).sort(),
+ ["chapters", "tags"],
+ );
+ });
+});
+
+test("a section that later succeeds clears its own failure, not the other's", async () => {
+ await withVideoDir(async (dir) => {
+ for (const section of ["chapters", "tags"] as const) {
+ await writeDigestFailure(dir, {
+ section,
+ at: "2026-07-29T00:00:00.000Z",
+ appId: "a",
+ model: "m",
+ promptVersion: 2,
+ reason: "all-rejected",
+ chunks: 1,
+ chunksOk: 1,
+ warnings: [],
+ });
+ }
+ await writeDigestSection(dir, {
+ section: "chapters",
+ items: [
+ { id: "c0", start: 0, clock: "0:00", title: "Intro", decidedBy: "ai" },
+ ],
+ provenance: PROVENANCE,
+ warnings: [],
+ });
+
+ const record = await loadDigest(dir);
+ assert.deepEqual(
+ record?.failures?.map((f) => f.section),
+ ["tags"],
+ "the succeeding section's failure is history; the other's is not ours to clear",
+ );
+ assert.equal(isSectionFresh(record, "chapters", TARGET), true);
+ assert.equal(record?.digestSchemaVersion, DIGEST_SCHEMA_VERSION);
+ });
+});
+
+// A mirror carrying a borrowed digest whose own regeneration fails must not
+// lose its attribution — the viewer labels borrowed content as borrowed.
+test("recording a failure preserves derivedFrom", async () => {
+ await withVideoDir(async (dir) => {
+ const { writeSharedDigest } = await import("./digest-server");
+ await writeSharedDigest(
+ dir,
+ {
+ digestSchemaVersion: DIGEST_SCHEMA_VERSION,
+ promptVersion: 2,
+ contextHash: "ctx",
+ warnings: [],
+ sections: {
+ chapters: { provenance: PROVENANCE, items: [] },
+ },
+ },
+ {
+ slug: "chan/canon",
+ clusterId: "c1",
+ sharedAt: "2026-07-29T00:00:00.000Z",
+ offsetSeconds: 0.2,
+ },
+ );
+ await writeDigestFailure(dir, {
+ section: "tags",
+ at: "2026-07-29T00:00:00.000Z",
+ appId: "a",
+ model: "m",
+ promptVersion: 2,
+ reason: "no-output",
+ chunks: 1,
+ chunksOk: 0,
+ warnings: [],
+ });
+ const record = await loadDigest(dir);
+ assert.equal(record?.derivedFrom?.slug, "chan/canon");
+ });
+});
diff --git a/common/lib/digest-server.ts b/common/lib/digest-server.ts
@@ -0,0 +1,350 @@
+// Server-only disk I/O for the two digest sidecars. Modelled on
+// availability-server.ts, which is the only per-video sidecar writer in the repo
+// with all four properties this needs: atomic tmp+rename, a tolerant validated
+// read that degrades to null, append-on-change history so re-runs are auditable,
+// and a PARTIAL-OWNERSHIP merge that preserves fields another writer owns.
+//
+// The partial-ownership property is the load-bearing one here: writeDigestSection
+// replaces exactly one section and leaves the other alone, so a metered tags run
+// never clobbers local chapters (and vice versa).
+
+import path from "node:path";
+import { readFile, rename, rm, writeFile } from "node:fs/promises";
+import {
+ DIGEST_FILENAME,
+ DIGEST_OVERRIDES_FILENAME,
+ DIGEST_OVERRIDES_VERSION,
+ DIGEST_SCHEMA_VERSION,
+ isDigestSectionKind,
+ type DigestChapter,
+ type DigestDerivedFrom,
+ type DigestHistoryEntry,
+ type DigestItem,
+ type DigestOverrides,
+ type DigestProvenance,
+ type DigestRecord,
+ type DigestSectionFailure,
+ type DigestSectionKind,
+ type DigestTag,
+ type DigestWarning,
+} from "./digest";
+
+// Cap the audit trail so a video that is regenerated across many prompt
+// iterations during Stage B tuning cannot grow an unbounded sidecar.
+const MAX_HISTORY_ENTRIES = 40;
+
+export function digestPath(videoDir: string): string {
+ return path.join(videoDir, DIGEST_FILENAME);
+}
+
+export function digestOverridesPath(videoDir: string): string {
+ return path.join(videoDir, DIGEST_OVERRIDES_FILENAME);
+}
+
+// ---------------------------------------------------------------------------
+// Reads (tolerant: anything unparseable or structurally wrong reads as absent,
+// so one corrupt sidecar can never fail a channel-wide sweep)
+// ---------------------------------------------------------------------------
+
+export async function loadDigest(
+ videoDir: string,
+): Promise<DigestRecord | null> {
+ try {
+ const raw = await readFile(digestPath(videoDir), "utf8");
+ const parsed = JSON.parse(raw) as Partial<DigestRecord>;
+ if (typeof parsed?.digestSchemaVersion !== "number") return null;
+ if (!parsed.sections || typeof parsed.sections !== "object") return null;
+ return {
+ digestSchemaVersion: parsed.digestSchemaVersion,
+ promptVersion:
+ typeof parsed.promptVersion === "number" ? parsed.promptVersion : 0,
+ contextHash:
+ typeof parsed.contextHash === "string" ? parsed.contextHash : "",
+ warnings: Array.isArray(parsed.warnings)
+ ? (parsed.warnings as DigestWarning[])
+ : [],
+ sections: parsed.sections,
+ // This reader rebuilds the record field by field rather than spreading,
+ // so EVERY new field has to be added here or it round-trips to nothing —
+ // silently, since the write succeeds and the read just omits it.
+ ...(Array.isArray(parsed.failures)
+ ? { failures: parsed.failures as DigestSectionFailure[] }
+ : {}),
+ ...(Array.isArray(parsed.history)
+ ? { history: parsed.history as DigestHistoryEntry[] }
+ : {}),
+ ...(parsed.derivedFrom
+ ? { derivedFrom: parsed.derivedFrom as DigestDerivedFrom }
+ : {}),
+ };
+ } catch {
+ return null;
+ }
+}
+
+// Cheap existence/coverage check for the snapshot's noDigest bucket and the
+// channel digestCount, which run over every video dir in a channel. Reads the
+// file (a few KB) rather than statting, because a digest whose sections are all
+// empty is not coverage.
+export async function hasDigest(videoDir: string): Promise<boolean> {
+ const record = await loadDigest(videoDir);
+ if (!record) return false;
+ return (
+ (record.sections.chapters?.items.length ?? 0) > 0 ||
+ (record.sections.tags?.items.length ?? 0) > 0
+ );
+}
+
+export async function loadDigestOverrides(
+ videoDir: string,
+): Promise<DigestOverrides | null> {
+ try {
+ const raw = await readFile(digestOverridesPath(videoDir), "utf8");
+ const parsed = JSON.parse(raw) as Partial<DigestOverrides>;
+ // A hand-authored file may omit `version`; treat it as current rather than
+ // discarding human work over a missing scalar.
+ const chapters = sanitizeChapters(parsed.chapters);
+ const tags = sanitizeTags(parsed.tags);
+ if (chapters.length === 0 && tags.length === 0 && !parsed.note) return null;
+ return {
+ version:
+ typeof parsed.version === "number"
+ ? parsed.version
+ : DIGEST_OVERRIDES_VERSION,
+ ...(chapters.length > 0 ? { chapters } : {}),
+ ...(tags.length > 0 ? { tags } : {}),
+ ...(typeof parsed.note === "string" ? { note: parsed.note } : {}),
+ ...(typeof parsed.updatedAt === "string"
+ ? { updatedAt: parsed.updatedAt }
+ : {}),
+ };
+ } catch {
+ return null;
+ }
+}
+
+// Hand-authored overrides are coerced, not trusted: an entry missing an id or a
+// body is dropped rather than poisoning the merge.
+function sanitizeChapters(value: unknown): DigestChapter[] {
+ if (!Array.isArray(value)) return [];
+ const out: DigestChapter[] = [];
+ for (const raw of value) {
+ if (!raw || typeof raw !== "object") continue;
+ const r = raw as Record<string, unknown>;
+ if (typeof r.id !== "string" || !r.id.trim()) continue;
+ const title = typeof r.title === "string" ? r.title.trim() : "";
+ const start =
+ typeof r.start === "number" && Number.isFinite(r.start)
+ ? Math.max(0, Math.floor(r.start))
+ : undefined;
+ // An override may be a patch (id + enabled:false) with no title at all.
+ if (!title && r.enabled !== false) continue;
+ out.push({
+ id: r.id,
+ ...(start !== undefined ? { start } : {}),
+ ...(typeof r.clock === "string" ? { clock: r.clock } : {}),
+ title,
+ decidedBy: "human",
+ ...(r.enabled === false ? { enabled: false } : {}),
+ } as DigestChapter);
+ }
+ return out;
+}
+
+function sanitizeTags(value: unknown): DigestTag[] {
+ if (!Array.isArray(value)) return [];
+ const out: DigestTag[] = [];
+ for (const raw of value) {
+ if (!raw || typeof raw !== "object") continue;
+ const r = raw as Record<string, unknown>;
+ if (typeof r.id !== "string" || !r.id.trim()) continue;
+ const tag = typeof r.tag === "string" ? r.tag.trim() : "";
+ if (!tag && r.enabled !== false) continue;
+ out.push({
+ id: r.id,
+ tag,
+ decidedBy: "human",
+ ...(r.enabled === false ? { enabled: false } : {}),
+ });
+ }
+ return out;
+}
+
+// ---------------------------------------------------------------------------
+// Writes
+// ---------------------------------------------------------------------------
+
+async function writeJsonAtomic(file: string, value: unknown): Promise<void> {
+ const tmp = `${file}.tmp-${process.pid}`;
+ await writeFile(tmp, JSON.stringify(value, null, 2) + "\n");
+ await rename(tmp, file);
+}
+
+export async function writeDigest(
+ videoDir: string,
+ record: DigestRecord,
+): Promise<void> {
+ await writeJsonAtomic(digestPath(videoDir), record);
+}
+
+export type WriteDigestSectionInput = {
+ section: DigestSectionKind;
+ items: DigestItem[];
+ provenance: DigestProvenance;
+ // Warnings from THIS pass. Warnings belonging to the section being rewritten
+ // are replaced; another section's warnings are preserved.
+ warnings: DigestWarning[];
+};
+
+// Read-modify-write ONE section, preserving everything another writer owns.
+// This is the only path the generator uses, and it never touches
+// ai-digest.overrides.json — that file is human-owned, full stop.
+export async function writeDigestSection(
+ videoDir: string,
+ input: WriteDigestSectionInput,
+): Promise<DigestRecord> {
+ const existing = await loadDigest(videoDir);
+ const sections = { ...(existing?.sections ?? {}) };
+ // The cast is contained here: DigestSectionKind and the item type are paired
+ // by construction at every call site (digestVideo builds both together).
+ (sections as Record<string, unknown>)[input.section] = {
+ provenance: input.provenance,
+ items: input.items,
+ };
+
+ // Drop the prior pass's warnings for THIS section only.
+ const keptWarnings = (existing?.warnings ?? []).filter(
+ (w) => w.section !== input.section,
+ );
+
+ // This section just succeeded, so its recorded total failure is history.
+ // Another section's failure is not ours to clear.
+ const keptFailures = (existing?.failures ?? []).filter(
+ (f) => f.section !== input.section,
+ );
+
+ const entry: DigestHistoryEntry = {
+ section: input.section,
+ generatedAt: input.provenance.generatedAt,
+ appId: input.provenance.appId,
+ model: input.provenance.model,
+ promptVersion: input.provenance.promptVersion,
+ contextHash: input.provenance.contextHash,
+ itemCount: input.items.length,
+ warningCount: input.warnings.length,
+ };
+ const history = [...(existing?.history ?? []), entry].slice(
+ -MAX_HISTORY_ENTRIES,
+ );
+
+ const record: DigestRecord = {
+ digestSchemaVersion: DIGEST_SCHEMA_VERSION,
+ promptVersion: input.provenance.promptVersion,
+ contextHash: input.provenance.contextHash,
+ warnings: [...keptWarnings, ...input.warnings],
+ sections,
+ ...(keptFailures.length > 0 ? { failures: keptFailures } : {}),
+ history,
+ };
+ // A freshly generated section makes the record this video's own again: it is
+ // no longer a copy of a cluster's canonical member.
+ await writeDigest(videoDir, record);
+ return record;
+}
+
+// Record a generation pass that produced nothing usable.
+//
+// Writes NO section, on purpose — see DigestSectionFailure. The record it
+// leaves is what makes the failure reviewable at all: without it the only trace
+// is a job log that rotates, and the videos the model does worst on are exactly
+// the ones a review queue most needs to surface.
+//
+// At most one failure per section is kept: a sweep retries, and appending would
+// grow the sidecar without bound on a video that fails every pass.
+export async function writeDigestFailure(
+ videoDir: string,
+ failure: DigestSectionFailure,
+): Promise<DigestRecord> {
+ const existing = await loadDigest(videoDir);
+ const failures = [
+ ...(existing?.failures ?? []).filter((f) => f.section !== failure.section),
+ failure,
+ ];
+ const record: DigestRecord = {
+ digestSchemaVersion: DIGEST_SCHEMA_VERSION,
+ // A failed pass must NOT claim the record's top-level identity — those
+ // mirror the last pass that actually wrote a section, and overwriting them
+ // here would make a corpus survey read a failure as a generation.
+ promptVersion: existing?.promptVersion ?? failure.promptVersion,
+ contextHash: existing?.contextHash ?? "",
+ warnings: existing?.warnings ?? [],
+ sections: existing?.sections ?? {},
+ failures,
+ ...(existing?.history ? { history: existing.history } : {}),
+ // Preserved: a mirror whose own regeneration failed is still carrying the
+ // canonical member's digest, and dropping this would silently un-attribute
+ // borrowed content the viewer labels as borrowed.
+ ...(existing?.derivedFrom ? { derivedFrom: existing.derivedFrom } : {}),
+ };
+ await writeDigest(videoDir, record);
+ return record;
+}
+
+// Copy a canonical member's digest onto an aligned duplicate, stamped with the
+// provenance of where it came from and the measured timing offset that made
+// sharing safe. The receiving video's own overrides are untouched, so a human
+// correction on a mirror still wins over the shared machine content.
+export async function writeSharedDigest(
+ videoDir: string,
+ source: DigestRecord,
+ derivedFrom: DigestDerivedFrom,
+): Promise<DigestRecord> {
+ const record: DigestRecord = {
+ ...source,
+ // The share is an event in the receiving video's history too.
+ history: [
+ ...(source.history ?? []),
+ ...Object.entries(source.sections)
+ .filter(([kind]) => isDigestSectionKind(kind))
+ .map(([kind, section]) => ({
+ section: kind as DigestSectionKind,
+ generatedAt: derivedFrom.sharedAt,
+ appId: section!.provenance.appId,
+ model: section!.provenance.model,
+ promptVersion: section!.provenance.promptVersion,
+ contextHash: section!.provenance.contextHash,
+ itemCount: section!.items.length,
+ warningCount: 0,
+ })),
+ ].slice(-MAX_HISTORY_ENTRIES),
+ derivedFrom,
+ };
+ await writeDigest(videoDir, record);
+ return record;
+}
+
+// Human-authored write. Separate function, separate file, so nothing in the
+// generation path can reach it.
+export async function writeDigestOverrides(
+ videoDir: string,
+ overrides: DigestOverrides,
+): Promise<void> {
+ const hasContent =
+ (overrides.chapters?.length ?? 0) > 0 ||
+ (overrides.tags?.length ?? 0) > 0 ||
+ Boolean(overrides.note);
+ const file = digestOverridesPath(videoDir);
+ if (!hasContent) {
+ // Emptying the override list means "revert to machine output" — remove the
+ // file rather than leaving an empty shadow behind.
+ await rm(file, { force: true });
+ return;
+ }
+ await writeJsonAtomic(file, {
+ version: DIGEST_OVERRIDES_VERSION,
+ ...(overrides.chapters?.length ? { chapters: overrides.chapters } : {}),
+ ...(overrides.tags?.length ? { tags: overrides.tags } : {}),
+ ...(overrides.note ? { note: overrides.note } : {}),
+ updatedAt: overrides.updatedAt ?? new Date().toISOString(),
+ });
+}
diff --git a/common/lib/digest.test.ts b/common/lib/digest.test.ts
@@ -0,0 +1,358 @@
+import { test } from "node:test";
+import assert from "node:assert/strict";
+import {
+ DIGEST_SCHEMA_VERSION,
+ chapterId,
+ digestPromptVariant,
+ effectiveDigest,
+ isSectionFresh,
+ isSharedFrom,
+ tagId,
+ type DigestOverrides,
+ type DigestRecord,
+} from "./digest";
+
+// The override model is the single most important thing in the digest layer: a
+// full sweep is weeks of wall-clock, so a human correction that a regeneration
+// destroys is work that can never be affordably redone.
+
+function machineRecord(over: Partial<DigestRecord> = {}): DigestRecord {
+ return {
+ digestSchemaVersion: DIGEST_SCHEMA_VERSION,
+ promptVersion: 1,
+ contextHash: "abc123",
+ warnings: [],
+ sections: {
+ chapters: {
+ provenance: {
+ appId: "ollama-direct",
+ model: "qwen2.5:7b",
+ modelRequested: "qwen2.5:7b",
+ lane: "local-gpu",
+ generatedAt: "2026-07-26T00:00:00.000Z",
+ promptVersion: 1,
+ contextHash: "abc123",
+ },
+ items: [
+ { id: "c0", start: 0, clock: "00:00:00", title: "Intro", decidedBy: "ai" },
+ {
+ id: "c300",
+ start: 300,
+ clock: "00:05:00",
+ title: "Court filing deadlnes",
+ decidedBy: "ai",
+ },
+ ],
+ },
+ },
+ ...over,
+ };
+}
+
+test("effectiveDigest returns machine content when there are no overrides", () => {
+ const e = effectiveDigest(machineRecord(), null);
+ assert.deepEqual(
+ e.chapters.map((c) => c.title),
+ ["Intro", "Court filing deadlnes"],
+ );
+ assert.equal(e.hasOverrides, false);
+ assert.ok(e.chapters.every((c) => c.decidedBy === "ai"));
+});
+
+test("a human override REPLACES the machine item and reports decidedBy: human", () => {
+ const overrides: DigestOverrides = {
+ version: 1,
+ chapters: [
+ { id: "c300", start: 300, clock: "00:05:00", title: "Court filing deadlines", decidedBy: "human" },
+ ],
+ };
+ const e = effectiveDigest(machineRecord(), overrides);
+ const fixed = e.chapters.find((c) => c.id === "c300");
+ assert.equal(fixed?.title, "Court filing deadlines", "the typo is corrected");
+ assert.equal(fixed?.decidedBy, "human");
+ assert.equal(e.chapters.length, 2, "the untouched machine chapter survives");
+ assert.equal(e.hasOverrides, true);
+});
+
+test("an override authored as a patch does not blank the machine's start", () => {
+ const overrides: DigestOverrides = {
+ version: 1,
+ // Title-only patch: no `start`.
+ chapters: [{ id: "c300", title: "Retitled", decidedBy: "human" }] as never,
+ };
+ const e = effectiveDigest(machineRecord(), overrides);
+ const patched = e.chapters.find((c) => c.id === "c300");
+ assert.equal(patched?.start, 300, "the machine's start is preserved");
+ assert.equal(patched?.title, "Retitled");
+});
+
+test("enabled: false SUPPRESSES a generated item without deleting it", () => {
+ const overrides: DigestOverrides = {
+ version: 1,
+ chapters: [{ id: "c0", title: "", decidedBy: "human", enabled: false }] as never,
+ };
+ const e = effectiveDigest(machineRecord(), overrides);
+ assert.deepEqual(
+ e.chapters.map((c) => c.id),
+ ["c300"],
+ );
+ // The suppression is a recorded decision, so re-generating the same item does
+ // not resurrect something a human rejected.
+ assert.equal(e.hasOverrides, true);
+});
+
+test("an override with a new id APPENDS a human-authored chapter", () => {
+ const overrides: DigestOverrides = {
+ version: 1,
+ chapters: [
+ { id: "c600", start: 600, clock: "00:10:00", title: "Missed topic", decidedBy: "human" },
+ ],
+ };
+ const e = effectiveDigest(machineRecord(), overrides);
+ assert.deepEqual(
+ e.chapters.map((c) => c.start),
+ [0, 300, 600],
+ "composed list stays sorted by start",
+ );
+});
+
+test("overrides survive a regeneration that rewrites the machine file", () => {
+ const overrides: DigestOverrides = {
+ version: 1,
+ chapters: [{ id: "c300", title: "Court filing deadlines", decidedBy: "human" }] as never,
+ };
+ // A later run at a new promptVersion re-emits a DIFFERENT title at the same
+ // snapped start — the id is derived from the start, so the human's correction
+ // still shadows it.
+ const regenerated = machineRecord({
+ promptVersion: 2,
+ sections: {
+ chapters: {
+ provenance: {
+ appId: "ollama-direct",
+ model: "qwen2.5:7b",
+ modelRequested: "qwen2.5:7b",
+ lane: "local-gpu",
+ generatedAt: "2026-08-01T00:00:00.000Z",
+ promptVersion: 2,
+ contextHash: "abc123",
+ },
+ items: [
+ {
+ id: "c300",
+ start: 300,
+ clock: "00:05:00",
+ title: "Some new machine title",
+ decidedBy: "ai",
+ },
+ ],
+ },
+ },
+ });
+ const e = effectiveDigest(regenerated, overrides);
+ assert.equal(e.chapters[0].title, "Court filing deadlines");
+ assert.equal(e.chapters[0].decidedBy, "human");
+});
+
+test("chapterId is derived from the snapped start, so it is stable", () => {
+ assert.equal(chapterId(300), "c300");
+ assert.equal(chapterId(300.9), "c300");
+});
+
+test("tagId normalizes so a re-generated tag lands on the same id", () => {
+ assert.equal(tagId("Court Filings"), "tcourt-filings");
+ assert.equal(tagId("court filings!"), "tcourt-filings");
+});
+
+// ---------------------------------------------------------------------------
+// Freshness — what keeps a re-run from becoming a second multi-week sweep
+// ---------------------------------------------------------------------------
+
+const TARGET = {
+ appId: "ollama-direct",
+ model: "qwen2.5:7b",
+ promptVersion: 1,
+ contextHash: "abc123",
+};
+
+test("an unchanged identity is fresh (a re-run is a no-op)", () => {
+ assert.equal(isSectionFresh(machineRecord(), "chapters", TARGET), true);
+});
+
+test("a missing record or section is stale", () => {
+ assert.equal(isSectionFresh(null, "chapters", TARGET), false);
+ assert.equal(isSectionFresh(machineRecord(), "tags", TARGET), false);
+});
+
+test("a bumped promptVersion makes the section stale", () => {
+ assert.equal(
+ isSectionFresh(machineRecord(), "chapters", { ...TARGET, promptVersion: 2 }),
+ false,
+ );
+});
+
+test("a changed contextHash makes the section stale", () => {
+ // This is why contextHash is plumbed BEFORE the sweep even with empty notes:
+ // adding a channel note later must invalidate that channel, not the corpus.
+ assert.equal(
+ isSectionFresh(machineRecord(), "chapters", { ...TARGET, contextHash: "zzz" }),
+ false,
+ );
+});
+
+test("a different engine or model makes the section stale", () => {
+ assert.equal(
+ isSectionFresh(machineRecord(), "chapters", { ...TARGET, appId: "claude-code" }),
+ false,
+ );
+ assert.equal(
+ isSectionFresh(machineRecord(), "chapters", { ...TARGET, model: "llama3" }),
+ false,
+ );
+});
+
+test("freshness compares the REQUESTED model, not the resolved one", () => {
+ // An alias resolving to a full tag is not a model change; treating it as one
+ // would re-run the entire corpus.
+ const record = machineRecord({
+ sections: {
+ chapters: {
+ provenance: {
+ appId: "ollama-direct",
+ model: "qwen2.5:7b",
+ modelRequested: "qwen2.5",
+ lane: "local-gpu",
+ generatedAt: "2026-07-26T00:00:00.000Z",
+ promptVersion: 1,
+ contextHash: "abc123",
+ },
+ items: [],
+ },
+ },
+ });
+ assert.equal(
+ isSectionFresh(record, "chapters", { ...TARGET, model: "qwen2.5" }),
+ true,
+ );
+});
+
+test("a schema-version bump invalidates every digest on disk", () => {
+ const record = machineRecord({ digestSchemaVersion: 0 });
+ assert.equal(isSectionFresh(record, "chapters", TARGET), false);
+});
+
+test("isSharedFrom recognizes a digest copied from a cluster's canonical member", () => {
+ const shared = machineRecord({
+ derivedFrom: {
+ slug: "Quartering/abc",
+ clusterId: "cluster1",
+ sharedAt: "2026-07-26T00:00:00.000Z",
+ offsetSeconds: 0.4,
+ },
+ });
+ assert.equal(isSharedFrom(shared, "Quartering/abc"), true);
+ assert.equal(isSharedFrom(shared, "Quartering/other"), false);
+ assert.equal(isSharedFrom(machineRecord(), "Quartering/abc"), false);
+});
+
+// ---------------------------------------------------------------------------
+// promptVariant — the identity field that lets several prompt shapes coexist
+// ---------------------------------------------------------------------------
+
+test("digestPromptVariant: the default configuration has no variant", () => {
+ // The derivation is RELATIVE TO THE DEFAULTS, so this tracks whatever they
+ // currently are — today chunk-local at 600 cues. Pinned explicitly (rather
+ // than via the constants) so that flipping a default without bumping
+ // PROMPT_VERSION fails here loudly instead of silently re-marking every
+ // existing digest as fresh.
+ assert.equal(digestPromptVariant({}), undefined);
+ assert.equal(digestPromptVariant({ timestampMode: "chunk-local" }), undefined);
+ assert.equal(digestPromptVariant({ maxCues: 600 }), undefined);
+ assert.equal(
+ digestPromptVariant({ timestampMode: "chunk-local", maxCues: 600 }),
+ undefined,
+ );
+ assert.equal(digestPromptVariant({ promptVariant: " " }), undefined);
+});
+
+test("digestPromptVariant: a non-default timestampMode is folded in", () => {
+ // Folded in rather than left to the operator to remember: a knob that changes
+ // the output but not the identity would let a re-run under a different mode
+ // skip every video as "fresh". Now that chunk-local is the default, it is
+ // ABSOLUTE that must carry an identity of its own.
+ assert.equal(digestPromptVariant({ timestampMode: "absolute" }), "absolute");
+ assert.equal(
+ digestPromptVariant({ promptVariant: "dense", timestampMode: "absolute" }),
+ "dense+absolute",
+ );
+ assert.equal(digestPromptVariant({ promptVariant: "dense" }), "dense");
+});
+
+test("isSectionFresh: a record with no promptVariant stays fresh by default", () => {
+ // The compatibility guarantee. Every digest written before the field existed
+ // must keep matching, or adding it would have invalidated the whole corpus.
+ assert.equal(isSectionFresh(machineRecord(), "chapters", TARGET), true);
+});
+
+test("isSectionFresh: a variant change invalidates the section", () => {
+ assert.equal(
+ isSectionFresh(machineRecord(), "chapters", {
+ ...TARGET,
+ promptVariant: "chunk-local",
+ }),
+ false,
+ "a chunk-local re-run must not skip absolute-mode records as fresh",
+ );
+});
+
+test("isSectionFresh: a variant record is stale against the default target", () => {
+ const record = machineRecord({
+ sections: {
+ chapters: {
+ provenance: {
+ appId: "ollama-direct",
+ model: "qwen2.5:7b",
+ modelRequested: "qwen2.5:7b",
+ lane: "local-gpu",
+ generatedAt: "2026-07-26T00:00:00.000Z",
+ promptVersion: 1,
+ promptVariant: "chunk-local",
+ contextHash: "abc123",
+ },
+ items: [],
+ },
+ },
+ });
+ assert.equal(isSectionFresh(record, "chapters", TARGET), false);
+ assert.equal(
+ isSectionFresh(record, "chapters", {
+ ...TARGET,
+ promptVariant: "chunk-local",
+ }),
+ true,
+ );
+});
+
+test("digestPromptVariant: the default chunk size contributes nothing", () => {
+ assert.equal(digestPromptVariant({ maxCues: 600 }), undefined);
+ assert.equal(digestPromptVariant({ maxCues: 1200 }), "c1200");
+ assert.equal(
+ digestPromptVariant({ timestampMode: "absolute", maxCues: 1200 }),
+ "absolute+c1200",
+ );
+});
+
+test("isSectionFresh: a chunk-size change invalidates the section", () => {
+ // Measured: halving the chunk took qwen2.5:7b from 9.34 to 13.37 chapters per
+ // hour and its worst coverage gap from 1:27:48 to 24:13. A change that large
+ // must not be able to hide behind an unchanged identity. 600 is now the
+ // default (and so contributes nothing), so the departure tested here is a
+ // move BACK UP to the old 1200.
+ assert.equal(
+ isSectionFresh(machineRecord(), "chapters", {
+ ...TARGET,
+ promptVariant: digestPromptVariant({ maxCues: 1200 }),
+ }),
+ false,
+ );
+});
diff --git a/common/lib/digest.ts b/common/lib/digest.ts
@@ -0,0 +1,525 @@
+// Client-safe types, constants and PURE merge helpers for the per-video AI
+// digest (chapters + topic tags). No node-only imports — the editor's video
+// panel is a "use client" file that pulls this in for the review UI. Server-only
+// disk I/O lives in digest-server.ts. Same split, for the same reason, as
+// doNotClean.ts / doNotClean-server.ts.
+//
+// TWO SIDECARS PER VIDEO DIR, and the split is the whole point:
+//
+// ai-digest.json machine-generated; freely overwritten by a re-run
+// ai-digest.overrides.json human-authored; regeneration NEVER writes it
+//
+// A sweep over 74k transcripts takes weeks, so a redo is unaffordable and human
+// corrections must survive one. Keeping the two in separate files means no merge
+// bug in the generator can destroy hand-written work: the generator only ever
+// opens the machine file. Readers compose the two with effectiveDigest().
+//
+// NEVER rename these to `transcript.<x>.<y>` — SUB_FILE_RE in videoStatus.ts
+// would claim such a file as a subtitle track.
+
+export const DIGEST_FILENAME = "ai-digest.json";
+export const DIGEST_OVERRIDES_FILENAME = "ai-digest.overrides.json";
+
+// Shape of ai-digest.json itself. Bumping this invalidates every digest on
+// disk, so it changes only when the FILE LAYOUT changes — not when a prompt
+// changes (that is promptVersion, which invalidates per section).
+export const DIGEST_SCHEMA_VERSION = 1;
+export const DIGEST_OVERRIDES_VERSION = 1;
+
+// Which engine lane produced a section. "local-gpu" is the default and carries
+// the corpus; "remote-api" is the opt-in metered overflow.
+//
+// The lane type, the app ids and the per-app config live HERE rather than in
+// digestApps.ts because settings.ts needs them and digestApps.ts imports execa —
+// a node-only dependency that must never be reachable from a client bundle.
+// digestApps.ts re-exports them so the registry still reads as one unit.
+export type DigestLane = "local-gpu" | "remote-api";
+
+// Default ollama context, and the chunk size sized to it. Both live HERE rather
+// than in digestApps.ts / digestPrompt.ts so the pure prompt module can size a
+// chunk against them without reaching a module that imports execa, and so the
+// identity helper below can compare against the default without a cycle.
+//
+// 8192, ON MEASUREMENT, not on comfort. Round 2 of the bake-off
+// (plans/bakeoff/round2.md) ran qwen2.5:7b at both sizes over the same 6.96
+// audio-hours, chunk-local in both cases:
+//
+// @16384 11.1% zero-yield, 9.34 chapters/h, worst gap 1:27:48, 25.1 days
+// @8192 11.8% zero-yield, 13.37 chapters/h, worst gap 0:24:13, 24.2 days
+//
+// Halving the window nearly halves the worst coverage gap and raises the
+// segmentation rate by 43% at no throughput cost — twice as many calls each
+// carry half the prompt, so the projected sweep is if anything shorter. 16k was
+// the cautious choice and it measured worse.
+//
+// A 7B model at 8k is also well within an 8 GB card (qwen2.5:7b KV cache
+// ~56 KB/token -> ~0.45 GB at 8k, atop 4.7 GB of weights). digestPrompt.ts
+// re-exports the cue count as DIGEST_MAX_CUES_PER_CHUNK and derives other sizes
+// from it; maxCuesForContext() scales the cue count with whatever numCtx an app
+// is actually configured for, so these two must stay in proportion (~10
+// tokens/cue: 600 cues ~= 6k tokens of transcript inside an 8k window, leaving
+// room for the prompt and the response).
+export const DEFAULT_DIGEST_NUM_CTX = 8192;
+export const DEFAULT_DIGEST_MAX_CUES_PER_CHUNK = 600;
+
+// How the transcript markers inside ONE chunk are numbered, and therefore what
+// the model is asked to copy.
+//
+// absolute the chunk's cues carry their real video times, and the prompt
+// states the chunk's real range ("01:31:43 to 02:16:09").
+// chunk-local the chunk is re-based to 00:00:00 and the prompt states
+// "00:00:00 to 00:44:26". The parser adds the offset back before
+// any guard runs, so the range clamp still checks the chunk's
+// REAL range and warnings still report real video times.
+//
+// This exists because of a measured failure, not a hunch. On a 2.3 h video
+// (community-notes/v2chrch, 3505 cues → 3 chunks) the third chunk — range
+// 01:31:43–02:16:09 — came back with nine starts of 00:00:00, 00:03:54,
+// 00:12:26 …: the model had reverted to counting from zero. The per-chunk clamp
+// caught all nine, which is exactly its job, but a caught error is still a lost
+// chunk, and >4 h videos are 8.2% of the corpus by count and 46% of its tokens.
+// chunk-local removes the large offset the model has to hold. Which mode is
+// actually better was a BAKE-OFF QUESTION (common/bin/digest-bakeoff.ts), which
+// is why both were shipped rather than one being pre-applied as a fix.
+//
+// IT HAS BEEN ANSWERED. Round 2 (plans/bakeoff/round2.md), qwen2.5:7b@8192 over
+// the same 6.96 audio-hours:
+//
+// absolute 29.4% zero-yield chunks, 10.35 chapters/h, worst gap 1:05:16,
+// 47.5% of items rejected (65 of them out-of-range)
+// chunk-local 11.8% zero-yield chunks, 13.37 chapters/h, worst gap 0:24:13,
+// 19.1% rejected (13 out-of-range)
+//
+// gemma2:9b ranks the two the same way, so the effect is not model-specific.
+// chunk-local is now the DEFAULT. Changing it changes generated output, so it
+// was paired with a PROMPT_VERSION bump (digestPrompt.ts) rather than left to
+// digestPromptVariant's relative-to-default derivation — see the note there.
+export const DIGEST_TIMESTAMP_MODES = ["absolute", "chunk-local"] as const;
+export type DigestTimestampMode = (typeof DIGEST_TIMESTAMP_MODES)[number];
+export const DEFAULT_DIGEST_TIMESTAMP_MODE: DigestTimestampMode = "chunk-local";
+
+export function isDigestTimestampMode(v: unknown): v is DigestTimestampMode {
+ return (
+ typeof v === "string" &&
+ (DIGEST_TIMESTAMP_MODES as readonly string[]).includes(v)
+ );
+}
+
+// The ONE place the recorded promptVariant string is derived, so every writer
+// (the controller, the bake-off harness) produces the same identity for the same
+// configuration.
+//
+// timestampMode is folded in rather than left to the operator to remember: a
+// knob that changes the output but not the identity would let a re-run under a
+// different mode skip every video as "fresh", which is precisely the failure
+// this identity exists to prevent. The default configuration maps to `undefined`
+// so pre-existing records stay fresh.
+export function digestPromptVariant(input: {
+ promptVariant?: string;
+ timestampMode?: DigestTimestampMode;
+ // Cues per chunk. Folded in for the same reason as timestampMode, and on the
+ // same evidence: halving it took qwen2.5:7b from 9.34 to 13.37 chapters/hour
+ // and its worst coverage gap from 1:27:48 to 24:13 on a 3.5 h video. A knob
+ // that changes the output that much cannot sit outside the identity, or a
+ // re-run at a new chunk size would skip the whole corpus as "fresh".
+ maxCues?: number;
+}): string | undefined {
+ const parts: string[] = [];
+ const named = input.promptVariant?.trim();
+ if (named) parts.push(named);
+ if (input.timestampMode && input.timestampMode !== DEFAULT_DIGEST_TIMESTAMP_MODE) {
+ parts.push(input.timestampMode);
+ }
+ // The default chunk size contributes NOTHING, so every record written before
+ // this field existed still compares equal.
+ if (input.maxCues && input.maxCues !== DEFAULT_DIGEST_MAX_CUES_PER_CHUNK) {
+ parts.push(`c${Math.floor(input.maxCues)}`);
+ }
+ return parts.length > 0 ? parts.join("+") : undefined;
+}
+
+export const OLLAMA_DIGEST_APP_ID = "ollama-direct";
+export const CLAUDE_DIGEST_APP_ID = "claude-code";
+export const DEFAULT_DIGEST_APP_ID = OLLAMA_DIGEST_APP_ID;
+
+// Per-app configuration persisted under settings.digest.apps[id]. Every field is
+// optional; an app falls back to its own defaults.
+export type DigestAppConfig = {
+ // Binary path/name override (process-based apps only).
+ bin?: string;
+ // Base URL override (HTTP apps only).
+ baseUrl?: string;
+ // Model id, e.g. "qwen2.5:7b" or "haiku".
+ model?: string;
+ // Context window in tokens. MUST reach the engine explicitly for ollama: its
+ // 4096 default silently truncates the input and the model then summarizes
+ // whatever fragment survived — measured, and the single easiest way to get
+ // quietly-wrong output at scale.
+ numCtx?: number;
+ // Sampling temperature. 0 for a structured extraction task.
+ temperature?: number;
+ // Reasoning-model toggle (ollama's top-level `think`). Only sent when set, so
+ // a model that does not support thinking is never handed a field it rejects.
+ //
+ // It matters for throughput, not correctness: measured on this box, qwen3:8b
+ // with thinking on spends most of its output budget on a `thinking` block
+ // before the JSON body the schema constrains. For an extraction task with a
+ // pinned schema that reasoning buys little and costs a multiple of the tokens,
+ // and tokens are what a multi-week sweep is priced in.
+ think?: boolean;
+ // Per-request wall-clock ceiling (ms). A wedged engine must not stall a sweep.
+ timeoutMs?: number;
+};
+
+// The two generated sections. Per-SECTION provenance (not per-file) because the
+// controller does a read-modify-write merge, so one video can legitimately hold
+// local chapters and metered tags.
+export const DIGEST_SECTION_KINDS = ["chapters", "tags"] as const;
+export type DigestSectionKind = (typeof DIGEST_SECTION_KINDS)[number];
+
+export function isDigestSectionKind(v: unknown): v is DigestSectionKind {
+ return (
+ typeof v === "string" &&
+ (DIGEST_SECTION_KINDS as readonly string[]).includes(v)
+ );
+}
+
+// Who decided this item. Every AI decision is marked as such and is
+// human-overridable — the same mechanism Phase 9 attribution will reuse rather
+// than inventing a second one.
+export type DecidedBy = "ai" | "human";
+
+// A chapter: a titled moment. `start` is SECONDS (the cue unit — see vtt.ts),
+// snapped to a real cue boundary by the parser. `clock` is the HH:MM:SS form the
+// model emitted, kept for auditing what the model actually said.
+export type DigestChapter = {
+ // Stable id so an override can shadow exactly one generated item. Derived from
+ // the snapped start (see chapterId) — deterministic, no Date.now/random.
+ id: string;
+ start: number;
+ clock: string;
+ title: string;
+ decidedBy: DecidedBy;
+ // Defaults true. An override sets false to SUPPRESS a generated item without
+ // deleting it (the searchAliases.ts `enabled` idiom), so a regeneration that
+ // re-emits the same item does not resurrect something a human rejected.
+ enabled?: boolean;
+};
+
+export type DigestTag = {
+ id: string;
+ tag: string;
+ decidedBy: DecidedBy;
+ enabled?: boolean;
+};
+
+export type DigestItem = DigestChapter | DigestTag;
+
+// Why an item or a whole chunk was rejected. Recorded, never silently dropped —
+// a multi-week sweep is only tunable if its failures are inspectable, and the
+// Phase 11a review queue is built entirely on this array.
+export type DigestWarningCode =
+ | "malformed-timestamp"
+ | "out-of-range"
+ | "language-drift"
+ | "non-monotonic"
+ | "seam-duplicate"
+ | "empty-title"
+ | "empty-output"
+ | "parse-failed"
+ | "chunk-failed";
+
+export type DigestWarning = {
+ code: DigestWarningCode;
+ // Which section's generation produced it, so warnings stay attributable after
+ // a read-modify-write merge of two separately-generated sections.
+ section: DigestSectionKind;
+ // Zero-based index of the transcript chunk that produced it, when applicable.
+ chunk?: number;
+ // The offending value, verbatim, so a prompt regression is diagnosable from
+ // the artifact alone.
+ value?: string;
+ detail?: string;
+};
+
+// What a regeneration compares against to decide "skip". This is what makes a
+// re-run targeted instead of a second multi-week sweep.
+export type DigestProvenance = {
+ appId: string;
+ // The model that ACTUALLY ran, as the engine reported it (e.g. "qwen2.5:7b").
+ // The audit record.
+ model: string;
+ // What the config ASKED for (e.g. "qwen2.5"). Freshness compares this, not
+ // `model`: an alias resolving to a full tag is not a model change, and treating
+ // it as one would re-run the entire corpus. Older records lack it, so readers
+ // fall back to `model`.
+ modelRequested?: string;
+ lane: DigestLane;
+ generatedAt: string;
+ // Bumped when the prompt/schema changes. Invalidates this section only.
+ promptVersion: number;
+ // Which prompt SHAPE produced this section, when it was not the default one.
+ // promptVersion answers "has the prompt changed since?"; this answers "which
+ // of several concurrently-supported shapes was used?" — the two are different
+ // questions and a bake-off needs both. Absent means the default shape
+ // (timestampMode "absolute", no named variant), so every record written before
+ // this field existed stays valid: isSectionFresh treats absent-on-both as
+ // equal. Same "one identity field per thing that can change the output" rule
+ // the rest of this record already follows.
+ promptVariant?: string;
+ // Hash of the channel-context inputs. Plumbed from the start even while
+ // context files are empty, so adding them later doesn't invalidate the corpus.
+ contextHash: string;
+ // Number of transcript chunks the engine was asked to process, and how many
+ // came back usable. A section generated from 3 of 5 chunks is suspect.
+ chunks?: number;
+ chunksOk?: number;
+ // Metered lanes only: what this section cost.
+ costUsd?: number;
+};
+
+export type DigestSection<T extends DigestItem> = {
+ provenance: DigestProvenance;
+ items: T[];
+};
+
+export type DigestSections = {
+ chapters?: DigestSection<DigestChapter>;
+ tags?: DigestSection<DigestTag>;
+};
+
+// One entry per generation pass, appended on change so a re-run is auditable
+// (the availability-server.ts history idiom).
+export type DigestHistoryEntry = {
+ section: DigestSectionKind;
+ generatedAt: string;
+ appId: string;
+ model: string;
+ promptVersion: number;
+ contextHash: string;
+ itemCount: number;
+ warningCount: number;
+};
+
+// A generation pass that produced NOTHING usable, recorded so the failure
+// survives.
+//
+// The two total-failure paths in digestVideo deliberately do not write a
+// section: an empty section carrying current provenance would read as FRESH and
+// the video would never be retried. The consequence was that the WORST failures
+// persisted zero evidence — they existed only in a job log that rotates at 500
+// records / 30 days — so a warnings-driven review queue would have been
+// systematically blind to exactly the videos it exists to catch.
+//
+// This is a sibling of `sections`, not a member of it, which is what keeps the
+// retry behaviour intact: isSectionFresh reads `sections[section].provenance`
+// and nothing else, so nothing here can make a failed video look done.
+export type DigestSectionFailure = {
+ section: DigestSectionKind;
+ at: string; // ISO
+ appId: string;
+ model: string;
+ promptVersion: number;
+ // WHY nothing survived, and the distinction is the whole point:
+ // "no-output" — every chunk failed or came back empty. The model never
+ // proposed anything: an engine, prompt or context problem.
+ // "all-rejected" — the model proposed items and every one failed a guard.
+ // The content exists and the guard threw it away.
+ // The validation run's worst video was the second kind — 13 chapters clamped
+ // away as out-of-range — and it looked identical to the first from outside.
+ // They need different fixes, and only a recorded artifact tells them apart.
+ reason: "no-output" | "all-rejected";
+ chunks: number;
+ chunksOk: number;
+ warnings: DigestWarning[];
+};
+
+export type DigestRecord = {
+ digestSchemaVersion: number;
+ // The most recent generation pass's prompt/context identity. Per-section
+ // provenance is authoritative for freshness (a video can hold sections made
+ // at different versions); these mirror whichever pass wrote last, so a
+ // corpus-wide sweep can be surveyed without opening every section.
+ promptVersion: number;
+ contextHash: string;
+ warnings: DigestWarning[];
+ sections: DigestSections;
+ // Latest total failure per section, if any. At most one entry per section —
+ // an 81-day sweep retries, and an append-only list would grow without bound
+ // on a video that fails every time. Cleared for a section that later
+ // succeeds. See DigestSectionFailure.
+ failures?: DigestSectionFailure[];
+ history?: DigestHistoryEntry[];
+ // Set when this digest was SHARED from another video (a duplicate cluster's
+ // canonical member) rather than generated for this one. Keeps the sharing
+ // honest and lets a later correction propagate to the whole cluster.
+ derivedFrom?: DigestDerivedFrom;
+};
+
+export type DigestDerivedFrom = {
+ // `${channelSlug}/${id}` of the canonical member the digest was generated for.
+ slug: string;
+ clusterId: string;
+ sharedAt: string;
+ // Measured max cue-timing offset (seconds) between the two videos at the
+ // sampled anchors. Sharing only happens at near-zero offset, so this records
+ // WHY it was considered safe.
+ offsetSeconds: number;
+};
+
+// ---------------------------------------------------------------------------
+// Overrides
+// ---------------------------------------------------------------------------
+
+// Human-authored shadow of the machine file, id-keyed exactly like a per-site
+// search-alias list shadows the global one (mergeAliases: later source wins).
+// An override entry with an id that matches a generated item REPLACES it; an id
+// with no match APPENDS. `enabled: false` suppresses without deleting.
+export type DigestOverrides = {
+ version: number;
+ chapters?: DigestChapter[];
+ tags?: DigestTag[];
+ // Free-text operator note (why this was corrected). Never read by code.
+ note?: string;
+ updatedAt?: string;
+};
+
+// ---------------------------------------------------------------------------
+// Pure helpers
+// ---------------------------------------------------------------------------
+
+// Deterministic per-item ids. A chapter's identity is its (snapped) start
+// second: that is what a human is correcting when they retitle a chapter, and it
+// stays stable across a regeneration that produces the same segmentation.
+export function chapterId(startSeconds: number): string {
+ return `c${Math.max(0, Math.floor(startSeconds))}`;
+}
+
+// A tag's identity is its normalized text, so re-generating the same tag lands
+// on the same id (and thus keeps honoring a human's `enabled: false`).
+export function tagId(tag: string): string {
+ const slug = tag
+ .toLowerCase()
+ .replace(/[^\p{L}\p{N}]+/gu, "-")
+ .replace(/^-+|-+$/g, "");
+ return `t${slug || "tag"}`;
+}
+
+// Merge a machine-generated item list with its human shadow. Later source wins
+// per id (mergeAliases), an overridden item reports decidedBy: "human", and
+// suppressed items (enabled === false) are dropped from the effective list.
+// `sort` keeps the composed list in the caller's canonical order.
+function mergeItems<T extends DigestItem>(
+ machine: T[],
+ overrides: T[] | undefined,
+ sort: (a: T, b: T) => number,
+): T[] {
+ const byId = new Map<string, T>();
+ for (const item of machine) byId.set(item.id, item);
+ for (const item of overrides ?? []) {
+ const existing = byId.get(item.id);
+ // A human edit is authoritative for the fields it names, but an override
+ // authored as a patch (id + title only) must not blank the machine's start.
+ byId.set(item.id, { ...(existing ?? {}), ...item, decidedBy: "human" } as T);
+ }
+ return Array.from(byId.values())
+ .filter((item) => item.enabled !== false)
+ .sort(sort);
+}
+
+const byStart = (a: DigestChapter, b: DigestChapter): number =>
+ a.start - b.start || a.title.localeCompare(b.title);
+const byTag = (a: DigestTag, b: DigestTag): number => a.tag.localeCompare(b.tag);
+
+export type EffectiveDigest = {
+ chapters: DigestChapter[];
+ tags: DigestTag[];
+ // True when any effective item came from the override file — the signal the
+ // editor uses to show "edited by hand".
+ hasOverrides: boolean;
+ warnings: DigestWarning[];
+ sections: DigestSections;
+ derivedFrom?: DigestDerivedFrom;
+};
+
+// Compose the machine file and the human shadow into what a reader should see.
+// Either side may be null (no digest yet / no corrections yet).
+export function effectiveDigest(
+ machine: DigestRecord | null,
+ overrides: DigestOverrides | null,
+): EffectiveDigest {
+ const chapters = mergeItems(
+ machine?.sections.chapters?.items ?? [],
+ overrides?.chapters,
+ byStart,
+ );
+ const tags = mergeItems(
+ machine?.sections.tags?.items ?? [],
+ overrides?.tags,
+ byTag,
+ );
+ return {
+ chapters,
+ tags,
+ hasOverrides:
+ (overrides?.chapters?.length ?? 0) > 0 ||
+ (overrides?.tags?.length ?? 0) > 0,
+ warnings: machine?.warnings ?? [],
+ sections: machine?.sections ?? {},
+ ...(machine?.derivedFrom ? { derivedFrom: machine.derivedFrom } : {}),
+ };
+}
+
+// The freshness test that keeps a re-run from becoming a second sweep: a section
+// is regenerated ONLY when its recorded identity differs from what we would
+// produce now. Missing section → stale (never generated).
+export type DigestFreshnessTarget = {
+ appId: string;
+ model: string;
+ promptVersion: number;
+ contextHash: string;
+ // Absent (or empty) means the default prompt shape — see DigestProvenance.
+ promptVariant?: string;
+ schemaVersion?: number;
+};
+
+// Absent and "" are the same thing (the default shape), so that a record written
+// before promptVariant existed compares equal to one written today with the
+// default settings. Without this normalization, adding the field would have
+// invalidated every digest on disk.
+function sameVariant(a: string | undefined, b: string | undefined): boolean {
+ return (a ?? "") === (b ?? "");
+}
+
+export function isSectionFresh(
+ record: DigestRecord | null,
+ section: DigestSectionKind,
+ target: DigestFreshnessTarget,
+): boolean {
+ if (!record) return false;
+ if (
+ record.digestSchemaVersion !==
+ (target.schemaVersion ?? DIGEST_SCHEMA_VERSION)
+ ) {
+ return false;
+ }
+ const p = record.sections[section]?.provenance;
+ if (!p) return false;
+ return (
+ p.appId === target.appId &&
+ (p.modelRequested ?? p.model) === target.model &&
+ p.promptVersion === target.promptVersion &&
+ p.contextHash === target.contextHash &&
+ sameVariant(p.promptVariant, target.promptVariant)
+ );
+}
+
+// A shared digest is fresh for the receiving video as long as it still points at
+// the same canonical member; the canonical member's own freshness is what drives
+// regeneration, and the share is re-applied from it.
+export function isSharedFrom(
+ record: DigestRecord | null,
+ canonicalSlug: string,
+): boolean {
+ return record?.derivedFrom?.slug === canonicalSlug;
+}
diff --git a/common/lib/digestApps.ts b/common/lib/digestApps.ts
@@ -0,0 +1,382 @@
+// First-class registry of DIGEST apps — the engines that turn a transcript into
+// derived content (chapters, topic tags). Deliberately mirrors
+// transcriptionApps.ts: each app owns how it is invoked, how its config maps onto
+// that invocation, and how its raw output becomes a parsed JSON body. The editor
+// selects one app per section, storing a small per-app config block
+// (settings.digest.apps[id]); digestVideo resolves the app and runs it.
+//
+// LOCAL-FIRST, BY DECISION. Hosted AI is not a dependency this project can rely
+// on given the nature of the archived content, so `ollama-direct` carries the
+// corpus and making it good enough IS the goal. `claude-code` is built but OFF by
+// default (settings.digest.remoteEnabled === false) — an opt-in overflow for the
+// >4h tail or a channel where local quality is poor, never the default.
+//
+// Pure definitions only (no I/O beyond reading process.env / getPaths inside
+// builders) so both the `common` controllers and the editor UI can import this.
+// Dependency direction is one-way: settings.ts -> digestApps.ts -> { paths,
+// parseStdoutJson }. Keep it that way to avoid import cycles.
+
+import { execa } from "execa";
+import { getPaths } from "./paths";
+import { extractJsonObject, parseStdoutJson } from "./parseStdoutJson";
+import {
+ CLAUDE_DIGEST_APP_ID,
+ DEFAULT_DIGEST_APP_ID,
+ DEFAULT_DIGEST_NUM_CTX,
+ OLLAMA_DIGEST_APP_ID,
+ type DigestAppConfig,
+ type DigestLane,
+} from "./digest";
+
+// The lane type, the app ids and DigestAppConfig are defined in the client-safe
+// digest.ts and re-exported here, so settings.ts can consume them without
+// reaching this module (which imports execa). `lane` is also WHY the two engines
+// get SEPARATE queue keys: registry.ts hardcodes concurrency 1 per key, so one
+// shared key would serialize a GPU-bound lane behind a network-bound one and
+// waste half the throughput of a multi-week sweep.
+export { CLAUDE_DIGEST_APP_ID, DEFAULT_DIGEST_APP_ID, OLLAMA_DIGEST_APP_ID };
+export type { DigestAppConfig, DigestLane };
+
+export type DigestRunInput = {
+ system: string;
+ prompt: string;
+ // JSON schema for the expected body. Engines with constrained decoding
+ // (ollama's `format`) enforce it; CLI lanes get it embedded in the prompt.
+ schema: Record<string, unknown>;
+ config: DigestAppConfig;
+ signal?: AbortSignal;
+ onLog?: (msg: string) => void;
+};
+
+export type DigestRunResult = {
+ // The parsed JSON body. Shape validation is the parser's job (digestParse.ts),
+ // not the engine's — an engine only guarantees "this is JSON".
+ data: unknown;
+ // The model that ACTUALLY ran, as reported by the engine when it says so.
+ // Recorded in provenance, so a config of "qwen2.5" resolving to "qwen2.5:7b"
+ // doesn't later look like a model change and trigger a needless regeneration.
+ model: string;
+ durationMs: number;
+ // Metered lanes only.
+ costUsd?: number;
+ inputTokens?: number;
+ outputTokens?: number;
+};
+
+export type DigestApp = {
+ id: string;
+ label: string;
+ lane: DigestLane;
+ // True when a run costs money. Metered apps are opt-in, never auto-queued, and
+ // their per-call count + cumulative cost is logged by the batch.
+ metered: boolean;
+ // Which config fields this app surfaces in the settings UI / consumes.
+ fields: {
+ bin?: boolean;
+ baseUrl?: boolean;
+ model?: boolean;
+ numCtx?: boolean;
+ temperature?: boolean;
+ };
+ // Model used when the per-app `model` override is empty.
+ defaultModel: () => string;
+ // Cheap reachability probe, so the editor can say "ollama is down" instead of
+ // failing 74k items one at a time. Mirrors pingRemoteHealth's contract: true
+ // only on a positive response, never throws.
+ probe: (config: DigestAppConfig) => Promise<boolean>;
+ run: (input: DigestRunInput) => Promise<DigestRunResult>;
+};
+
+// Fallbacks shared by both apps. The context default lives in digest.ts, so the
+// pure prompt module can size a chunk against the same number this resolves to —
+// two copies would let the chunker and the engine disagree, and ollama truncates
+// the excess SILENTLY.
+const DEFAULT_NUM_CTX = DEFAULT_DIGEST_NUM_CTX;
+const DEFAULT_TEMPERATURE = 0;
+const DEFAULT_TIMEOUT_MS = 10 * 60_000;
+
+function resolveNumCtx(config: DigestAppConfig): number {
+ return typeof config.numCtx === "number" && config.numCtx > 0
+ ? Math.floor(config.numCtx)
+ : DEFAULT_NUM_CTX;
+}
+
+function resolveTemperature(config: DigestAppConfig): number {
+ return typeof config.temperature === "number" && config.temperature >= 0
+ ? config.temperature
+ : DEFAULT_TEMPERATURE;
+}
+
+function resolveTimeoutMs(config: DigestAppConfig): number {
+ return typeof config.timeoutMs === "number" && config.timeoutMs > 0
+ ? Math.floor(config.timeoutMs)
+ : DEFAULT_TIMEOUT_MS;
+}
+
+// ---------------------------------------------------------------------------
+// ollama-direct — the local GPU lane that carries the corpus
+// ---------------------------------------------------------------------------
+
+function ollamaBase(config: DigestAppConfig): string {
+ const override = config.baseUrl?.trim();
+ // Strip trailing slashes the way remoteTranscribe.ts normalizes a worker base,
+ // so callers can always concatenate a leading-slash path.
+ return (override ? override.replace(/\/+$/, "") : getPaths().ollamaUrl) || "";
+}
+
+const ollamaDirect: DigestApp = {
+ id: OLLAMA_DIGEST_APP_ID,
+ label: "ollama (local, JSON schema)",
+ lane: "local-gpu",
+ metered: false,
+ fields: { baseUrl: true, model: true, numCtx: true, temperature: true },
+ defaultModel: () => process.env.OLLAMA_DIGEST_MODEL ?? "qwen2.5:7b",
+ async probe(config) {
+ const base = ollamaBase(config);
+ if (!base) return false;
+ try {
+ const res = await fetch(`${base}/api/tags`, {
+ signal: AbortSignal.timeout(3000),
+ });
+ return res.ok;
+ } catch {
+ return false;
+ }
+ },
+ async run({ system, prompt, schema, config, signal, onLog }) {
+ const base = ollamaBase(config);
+ if (!base) throw new Error("ollama URL is not configured (set OLLAMA_URL)");
+ const model = config.model?.trim() || ollamaDirect.defaultModel();
+ const numCtx = resolveNumCtx(config);
+ const startedAt = Date.now();
+
+ // The timeout is combined with the caller's cancel signal so a drain/cancel
+ // stops a long generation promptly AND a wedged engine can't stall forever.
+ const timeout = AbortSignal.timeout(resolveTimeoutMs(config));
+ const composed = signal ? AbortSignal.any([signal, timeout]) : timeout;
+
+ let res: Response;
+ try {
+ res = await fetch(`${base}/api/chat`, {
+ method: "POST",
+ headers: { "content-type": "application/json" },
+ signal: composed,
+ body: JSON.stringify({
+ model,
+ stream: false,
+ // Omitted entirely unless configured: ollama rejects `think` for
+ // models that have no reasoning mode, so sending a default would
+ // break every non-reasoning engine.
+ ...(typeof config.think === "boolean" ? { think: config.think } : {}),
+ // Schema-constrained decoding. This — not prompt wording — is what
+ // made a 7B model emit well-formed timestamps.
+ format: schema,
+ options: {
+ // Explicit, always. See DigestAppConfig.numCtx.
+ num_ctx: numCtx,
+ temperature: resolveTemperature(config),
+ },
+ messages: [
+ { role: "system", content: system },
+ { role: "user", content: prompt },
+ ],
+ }),
+ });
+ } catch (err) {
+ const message = (err as Error)?.message ?? String(err);
+ if (signal?.aborted) throw err;
+ throw new Error(
+ `ollama request failed (${base}): ${message}. Is the ollama service running?`,
+ );
+ }
+ if (!res.ok) {
+ const body = await res.text().catch(() => "");
+ throw new Error(
+ `ollama returned ${res.status}: ${body.slice(0, 300) || res.statusText}`,
+ );
+ }
+ const body = (await res.json()) as {
+ model?: string;
+ message?: { content?: string };
+ prompt_eval_count?: number;
+ eval_count?: number;
+ };
+ const content = body.message?.content ?? "";
+ const data = extractJsonObject(content);
+ if (data === null) {
+ throw new Error(
+ `ollama returned unparseable content: ${content.slice(0, 300)}`,
+ );
+ }
+ onLog?.(
+ `ollama ${body.model ?? model}: ${body.prompt_eval_count ?? "?"} in / ${
+ body.eval_count ?? "?"
+ } out tokens in ${Math.round((Date.now() - startedAt) / 100) / 10}s (num_ctx ${numCtx})`,
+ );
+ return {
+ data,
+ model: body.model ?? model,
+ durationMs: Date.now() - startedAt,
+ inputTokens: body.prompt_eval_count,
+ outputTokens: body.eval_count,
+ };
+ },
+};
+
+// ---------------------------------------------------------------------------
+// claude-code — the metered overflow lane, OFF by default
+// ---------------------------------------------------------------------------
+
+const claudeCode: DigestApp = {
+ id: CLAUDE_DIGEST_APP_ID,
+ label: "Claude Code CLI (metered)",
+ lane: "remote-api",
+ metered: true,
+ fields: { bin: true, model: true },
+ defaultModel: () => process.env.CLAUDE_DIGEST_MODEL ?? "",
+ async probe(config) {
+ const bin = config.bin?.trim() || getPaths().claudeBin;
+ try {
+ const res = await execa(bin, ["--version"], {
+ buffer: true,
+ reject: false,
+ timeout: 10_000,
+ });
+ return res.exitCode === 0;
+ } catch {
+ return false;
+ }
+ },
+ async run({ system, prompt, schema, config, signal, onLog }) {
+ const bin = config.bin?.trim() || getPaths().claudeBin;
+ const model = config.model?.trim() || claudeCode.defaultModel();
+ const startedAt = Date.now();
+ const argv = ["-p", "--output-format", "json"];
+ if (model) argv.push("--model", model);
+
+ // No constrained decoding over the CLI, so the contract goes in the prompt
+ // and the parser's guards do the enforcing. The prompt is passed on stdin
+ // rather than argv: a 12k-token chunk is ~50 KB and does not belong in an
+ // argument list.
+ const stdin = [
+ system,
+ "",
+ "Reply with a single JSON object and nothing else — no prose, no code fence.",
+ "It must validate against this JSON schema:",
+ JSON.stringify(schema),
+ "",
+ prompt,
+ ].join("\n");
+
+ let result: { exitCode: number | null; stdout: string; stderr: string };
+ try {
+ result = (await execa(bin, argv, {
+ input: stdin,
+ buffer: true,
+ reject: false,
+ cancelSignal: signal,
+ timeout: resolveTimeoutMs(config),
+ })) as typeof result;
+ } catch (err) {
+ const message = (err as Error)?.message ?? String(err);
+ // Same ENOENT-to-friendly-message treatment as the gallery-dl fetcher: a
+ // missing optional binary is a configuration problem, not a crash.
+ throw new Error(
+ /ENOENT/.test(message)
+ ? `claude CLI not found (set CLAUDE_BIN)`
+ : message,
+ );
+ }
+ if (result.exitCode !== 0) {
+ throw new Error(
+ `claude exited ${result.exitCode}: ${(result.stderr || result.stdout).slice(0, 300)}`,
+ );
+ }
+ // Two layers: the CLI's own JSON wrapper, then the model's JSON body inside
+ // the wrapper's `result` string.
+ const wrapper = parseStdoutJson(result.stdout) as {
+ result?: unknown;
+ total_cost_usd?: unknown;
+ usage?: { input_tokens?: number; output_tokens?: number };
+ is_error?: unknown;
+ } | null;
+ if (!wrapper || typeof wrapper !== "object") {
+ throw new Error(
+ `claude produced no parseable JSON wrapper: ${result.stdout.slice(0, 300)}`,
+ );
+ }
+ if (wrapper.is_error === true) {
+ throw new Error(
+ `claude reported an error: ${String(wrapper.result).slice(0, 300)}`,
+ );
+ }
+ const inner = typeof wrapper.result === "string" ? wrapper.result : "";
+ const data = extractJsonObject(inner);
+ if (data === null) {
+ throw new Error(
+ `claude returned unparseable content: ${inner.slice(0, 300)}`,
+ );
+ }
+ const costUsd =
+ typeof wrapper.total_cost_usd === "number"
+ ? wrapper.total_cost_usd
+ : undefined;
+ onLog?.(
+ `claude ${model || "(default model)"}: ${
+ costUsd !== undefined ? `$${costUsd.toFixed(4)}` : "cost unknown"
+ } in ${Math.round((Date.now() - startedAt) / 100) / 10}s`,
+ );
+ return {
+ data,
+ model: model || "claude-default",
+ durationMs: Date.now() - startedAt,
+ ...(costUsd !== undefined ? { costUsd } : {}),
+ inputTokens: wrapper.usage?.input_tokens,
+ outputTokens: wrapper.usage?.output_tokens,
+ };
+ },
+};
+
+// ---------------------------------------------------------------------------
+// Registry
+// ---------------------------------------------------------------------------
+
+export const DIGEST_APPS: Record<string, DigestApp> = {
+ [ollamaDirect.id]: ollamaDirect,
+ [claudeCode.id]: claudeCode,
+};
+
+// Total by construction — an unknown id falls back to the local default rather
+// than throwing, so a hand-edited settings.json can never crash a sweep.
+export function getDigestApp(id: string | undefined): DigestApp {
+ return (
+ (id ? DIGEST_APPS[id] : undefined) ?? DIGEST_APPS[DEFAULT_DIGEST_APP_ID]
+ );
+}
+
+export function isDigestAppId(id: unknown): id is string {
+ return typeof id === "string" && Boolean(DIGEST_APPS[id]);
+}
+
+// A client-safe view of an app (no functions). Build it on the server and pass it
+// to the settings form so this module — which reaches getPaths()/process.env —
+// never ends up in the client bundle.
+export type DigestAppDescriptor = {
+ id: string;
+ label: string;
+ lane: DigestLane;
+ metered: boolean;
+ fields: DigestApp["fields"];
+ defaultModel: string;
+};
+
+export function listDigestApps(): DigestAppDescriptor[] {
+ return Object.values(DIGEST_APPS).map((a) => ({
+ id: a.id,
+ label: a.label,
+ lane: a.lane,
+ metered: a.metered,
+ fields: a.fields,
+ defaultModel: a.defaultModel(),
+ }));
+}
diff --git a/common/lib/digestContext-server.ts b/common/lib/digestContext-server.ts
@@ -0,0 +1,63 @@
+// Per-channel context notes and the `contextHash` every generated section
+// records.
+//
+// This is the MINIMUM of PLAN.md's Phase 1.5, deliberately: the full
+// channel-context feature (structured frontmatter, per-channel entity lists) is
+// not built here, but the KEY is plumbed now. PLAN.md's "do not backfill before
+// 1.5" trap is exactly this — a digest generated without a contextHash cannot
+// tell whether it predates a channel's context note, so adding notes later would
+// invalidate the whole corpus. Hashing the (currently usually empty) note from
+// day one means adding a note later invalidates only that channel.
+//
+// The note is a plain Markdown file in the channel dir, hand-authored. Stage B's
+// compounding step is writing a correction here — "the co-host is Sam, not Sand"
+// — rather than patching individual videos, because a note improves every future
+// generation for that channel.
+
+import path from "node:path";
+import { readFile } from "node:fs/promises";
+import { createHash } from "node:crypto";
+import type { Paths } from "./paths";
+
+export const DIGEST_CONTEXT_FILENAME = "digest-context.md";
+
+// Cap what reaches the prompt: a note is guidance, not a second transcript, and
+// an unbounded one would eat the context window the transcript needs.
+const MAX_CONTEXT_CHARS = 4000;
+
+export type DigestContext = {
+ // The note text to embed in the prompt. "" when the channel has none.
+ note: string;
+ // Stable hash of every context input. Recorded in each section's provenance and
+ // compared on re-run. The empty-note hash is a real, stable value — NOT "" —
+ // so "no note" and "note removed" are the same state and neither is confused
+ // with "generated before contextHash existed" (which reads as "" and is stale).
+ hash: string;
+};
+
+export function digestContextPath(paths: Paths, channelSlug: string): string {
+ return path.join(paths.channelsDir, channelSlug, DIGEST_CONTEXT_FILENAME);
+}
+
+// Hash the context inputs. Versioned by a literal prefix so the hashing scheme
+// itself can change later without colliding with old values.
+export function hashDigestContext(note: string): string {
+ return createHash("sha1")
+ .update(`digest-context-v1\n${note}`)
+ .digest("hex")
+ .slice(0, 16);
+}
+
+export async function readDigestContext(
+ paths: Paths,
+ channelSlug: string,
+): Promise<DigestContext> {
+ let note = "";
+ try {
+ const raw = await readFile(digestContextPath(paths, channelSlug), "utf8");
+ note = raw.trim().slice(0, MAX_CONTEXT_CHARS);
+ } catch {
+ note = "";
+ }
+ return { note, hash: hashDigestContext(note) };
+}
diff --git a/common/lib/digestParse.test.ts b/common/lib/digestParse.test.ts
@@ -0,0 +1,422 @@
+import { test } from "node:test";
+import assert from "node:assert/strict";
+import { parseChapters, parseTags, snapToCueStart } from "./digestParse";
+import type { DigestChunkOutput } from "./digestParse";
+import type { Cue } from "./vtt";
+
+// Each guard here corresponds to a failure MEASURED on qwen2.5:7b before the
+// prompt/schema were hardened, so these tests are regression pins for real
+// output, not hypotheticals.
+
+function cue(start: number, text = `t${start}`): Cue {
+ return { start, end: start + 5, text };
+}
+
+// A 20-minute transcript with a cue every 10s.
+const CUES: Cue[] = Array.from({ length: 120 }, (_, i) => cue(i * 10));
+
+function chunk(
+ chapters: unknown[],
+ opts: { index?: number; start?: number; end?: number } = {},
+): DigestChunkOutput {
+ return {
+ index: opts.index ?? 0,
+ startSeconds: opts.start ?? 0,
+ endSeconds: opts.end ?? 1200,
+ data: { chapters },
+ };
+}
+
+test("guard 1: rejects a malformed timestamp and records it verbatim", () => {
+ // ":00:27" is the exact shape the naive prompt produced.
+ const { chapters, warnings } = parseChapters(
+ [
+ chunk([
+ { start: ":00:27", title: "Bad stamp" },
+ { start: "00:01:00", title: "Good stamp" },
+ ]),
+ ],
+ CUES,
+ );
+ assert.deepEqual(
+ chapters.map((c) => c.title),
+ ["Good stamp"],
+ );
+ const w = warnings.find((x) => x.code === "malformed-timestamp");
+ assert.ok(w, "the rejection is recorded, never silently dropped");
+ assert.equal(w?.value, ":00:27");
+ assert.equal(w?.section, "chapters");
+ assert.equal(w?.chunk, 0);
+});
+
+test("guard 1: rejects single-digit fields the regex pin forbids", () => {
+ const { chapters, warnings } = parseChapters(
+ [chunk([{ start: "1:2:3", title: "Loose stamp" }])],
+ CUES,
+ );
+ assert.equal(chapters.length, 0);
+ assert.equal(warnings[0].code, "malformed-timestamp");
+});
+
+test("guard 1: rejects an in-shape stamp with an impossible field", () => {
+ const { chapters, warnings } = parseChapters(
+ [chunk([{ start: "00:99:00", title: "Ninety-nine minutes" }])],
+ CUES,
+ );
+ assert.equal(chapters.length, 0);
+ assert.equal(warnings[0].code, "malformed-timestamp");
+});
+
+test("guard 2: flags a title that drifted out of English", () => {
+ // The measured drift was into Chinese on an 8.2k-token chunk.
+ const { chapters, warnings } = parseChapters(
+ [
+ chunk([
+ { start: "00:00:10", title: "法庭文件截止日期" },
+ { start: "00:02:00", title: "Court filing deadlines" },
+ ]),
+ ],
+ CUES,
+ );
+ assert.deepEqual(
+ chapters.map((c) => c.title),
+ ["Court filing deadlines"],
+ );
+ const w = warnings.find((x) => x.code === "language-drift");
+ assert.ok(w);
+ assert.equal(w?.value, "法庭文件截止日期");
+});
+
+test("guard 3: clamps to the CHUNK's range, not the video's", () => {
+ // The measured failure: 01:10:29 emitted for an input spanning 00:04:45 to
+ // 00:15:36. The video was 176 minutes long, so a whole-video range check would
+ // have ACCEPTED it. This is why the clamp is per-chunk.
+ const { chapters, warnings } = parseChapters(
+ [
+ chunk([{ start: "01:10:29", title: "Way past the end" }], {
+ start: 285,
+ end: 936,
+ }),
+ ],
+ CUES,
+ );
+ assert.equal(chapters.length, 0);
+ const w = warnings.find((x) => x.code === "out-of-range");
+ assert.ok(w);
+ assert.equal(w?.value, "01:10:29");
+ assert.match(w?.detail ?? "", /00:04:45/);
+});
+
+test("guard 3: accepts a start inside the chunk's own range", () => {
+ const { chapters } = parseChapters(
+ [
+ chunk([{ start: "00:05:00", title: "Inside the window" }], {
+ start: 285,
+ end: 936,
+ }),
+ ],
+ CUES,
+ );
+ assert.equal(chapters.length, 1);
+ assert.equal(chapters[0].title, "Inside the window");
+});
+
+test("guard 4a: drops a non-monotonic entry within a chunk", () => {
+ const { chapters, warnings } = parseChapters(
+ [
+ chunk([
+ { start: "00:01:00", title: "First" },
+ { start: "00:00:30", title: "Backwards" },
+ { start: "00:02:00", title: "Third" },
+ ]),
+ ],
+ CUES,
+ );
+ assert.deepEqual(
+ chapters.map((c) => c.title),
+ ["First", "Third"],
+ );
+ assert.ok(warnings.some((w) => w.code === "non-monotonic"));
+});
+
+test("guard 4b: de-dups equivalent chapters across a chunk seam", () => {
+ const { chapters, warnings } = parseChapters(
+ [
+ chunk([{ start: "00:05:00", title: "Court filing deadlines" }], {
+ index: 0,
+ start: 0,
+ end: 400,
+ }),
+ chunk([{ start: "00:05:10", title: "court filing deadlines." }], {
+ index: 1,
+ start: 300,
+ end: 700,
+ }),
+ ],
+ CUES,
+ );
+ assert.equal(chapters.length, 1, "the overlap's duplicate view is collapsed");
+ assert.ok(warnings.some((w) => w.code === "seam-duplicate"));
+});
+
+test("guard 4b: keeps a genuinely different topic near a seam", () => {
+ const { chapters } = parseChapters(
+ [
+ chunk([{ start: "00:05:00", title: "Court filing deadlines" }], {
+ index: 0,
+ start: 0,
+ end: 400,
+ }),
+ chunk([{ start: "00:05:20", title: "Jury selection" }], {
+ index: 1,
+ start: 300,
+ end: 700,
+ }),
+ ],
+ CUES,
+ );
+ assert.equal(chapters.length, 2);
+});
+
+test("guard 4b: collapses two chunks that snap onto the same cue", () => {
+ const { chapters, warnings } = parseChapters(
+ [
+ chunk([{ start: "00:05:00", title: "Alpha" }], { index: 0, end: 700 }),
+ chunk([{ start: "00:05:02", title: "Completely unrelated beta" }], {
+ index: 1,
+ end: 700,
+ }),
+ ],
+ CUES,
+ );
+ assert.equal(chapters.length, 1);
+ const w = warnings.find((w) => w.code === "seam-duplicate");
+ assert.match(w?.detail ?? "", /same start/);
+});
+
+test("guard 5: snaps a start onto the nearest cue boundary", () => {
+ // 00:05:04 (304s) sits between cues at 300s and 310s; 300 is nearer.
+ const { chapters } = parseChapters(
+ [chunk([{ start: "00:05:04", title: "Between cues" }])],
+ CUES,
+ );
+ assert.equal(chapters[0].start, 300);
+ assert.equal(chapters[0].clock, "00:05:04", "the model's own stamp is kept for audit");
+ assert.equal(chapters[0].id, "c300", "the id derives from the SNAPPED start");
+});
+
+test("guard 5: snaps forward when the later cue is nearer", () => {
+ const { chapters } = parseChapters(
+ [chunk([{ start: "00:05:08", title: "Nearer the next cue" }])],
+ CUES,
+ );
+ assert.equal(chapters[0].start, 310);
+});
+
+test("snapToCueStart handles a time before the first cue", () => {
+ assert.equal(snapToCueStart([cue(40), cue(50)], 5), 40);
+});
+
+test("snapToCueStart handles an empty cue list", () => {
+ assert.equal(snapToCueStart([], 42.7), 42);
+});
+
+test("every kept chapter is marked decidedBy: ai", () => {
+ const { chapters } = parseChapters(
+ [chunk([{ start: "00:01:00", title: "Machine-decided" }])],
+ CUES,
+ );
+ assert.equal(chapters[0].decidedBy, "ai");
+});
+
+test("a chunk with no chapters array is recorded as parse-failed", () => {
+ const { chapters, warnings } = parseChapters(
+ [{ index: 0, startSeconds: 0, endSeconds: 600, data: { nope: true } }],
+ CUES,
+ );
+ assert.equal(chapters.length, 0);
+ assert.equal(warnings[0].code, "parse-failed");
+});
+
+test("an empty chapters array is recorded as empty-output", () => {
+ const { warnings } = parseChapters([chunk([])], CUES);
+ assert.equal(warnings[0].code, "empty-output");
+});
+
+test("an entry with no title is recorded, not silently dropped", () => {
+ const { chapters, warnings } = parseChapters(
+ [chunk([{ start: "00:01:00", title: " " }])],
+ CUES,
+ );
+ assert.equal(chapters.length, 0);
+ assert.equal(warnings[0].code, "empty-title");
+});
+
+// ---------------------------------------------------------------------------
+// Tags
+// ---------------------------------------------------------------------------
+
+test("parseTags lowercases, dedups across chunks, and keeps ids stable", () => {
+ const { tags } = parseTags([
+ { index: 0, startSeconds: 0, endSeconds: 600, data: { tags: ["Court Filings", "appeals"] } },
+ { index: 1, startSeconds: 500, endSeconds: 1200, data: { tags: ["court filings", "sentencing"] } },
+ ]);
+ assert.deepEqual(
+ tags.map((t) => t.tag),
+ ["appeals", "court filings", "sentencing"],
+ );
+ assert.equal(tags[1].id, "tcourt-filings");
+});
+
+test("parseTags does NOT warn about an expected overlap duplicate", () => {
+ const { warnings } = parseTags([
+ { index: 0, startSeconds: 0, endSeconds: 600, data: { tags: ["appeals"] } },
+ { index: 1, startSeconds: 500, endSeconds: 1200, data: { tags: ["appeals"] } },
+ ]);
+ assert.equal(warnings.length, 0);
+});
+
+test("parseTags flags a drifted tag", () => {
+ const { tags, warnings } = parseTags([
+ { index: 0, startSeconds: 0, endSeconds: 600, data: { tags: ["上訴", "appeals"] } },
+ ]);
+ assert.deepEqual(
+ tags.map((t) => t.tag),
+ ["appeals"],
+ );
+ assert.equal(warnings[0].code, "language-drift");
+});
+
+// ---------------------------------------------------------------------------
+// chunk-local timestamps
+//
+// These pin the EXACT failure that motivated the mode. Digesting
+// community-notes/v2chrch (2.3 h, 3505 cues -> 3 chunks), the third chunk —
+// range 01:31:43-02:16:09 — came back with nine starts of 00:00:00, 00:03:54,
+// 00:12:26 ...: the model had reverted to counting from zero. The per-chunk
+// clamp rejected all nine, so the chunk yielded nothing.
+// ---------------------------------------------------------------------------
+
+// The real third chunk's range, to the second.
+const CHUNK3_START = 91 * 60 + 43; // 01:31:43 = 5503
+const CHUNK3_END = 2 * 3600 + 16 * 60 + 9; // 02:16:09 = 8169
+
+// Cues every 10s across the whole 2.3h video, so the snap has real boundaries.
+const LONG_CUES: Cue[] = Array.from({ length: 830 }, (_, i) => cue(i * 10));
+
+function chunk3(chapters: unknown[], mode?: "absolute" | "chunk-local"): DigestChunkOutput {
+ return {
+ index: 2,
+ startSeconds: CHUNK3_START,
+ endSeconds: CHUNK3_END,
+ data: { chapters },
+ ...(mode ? { timestampMode: mode } : {}),
+ };
+}
+
+test("chunk-local: a 00:03:00 start in the third chunk resolves to 01:34:43", () => {
+ const { chapters, warnings } = parseChapters(
+ [chunk3([{ start: "00:03:00", title: "Filing deadlines" }], "chunk-local")],
+ LONG_CUES,
+ );
+ assert.equal(chapters.length, 1, "the offset is added back, so it is in range");
+ // 5503 + 180 = 5683, snapped to the nearest cue boundary (5680).
+ assert.equal(chapters[0].start, 5680);
+ // `clock` is stored in REAL video time, never in the numbering the model used.
+ assert.equal(chapters[0].clock, "01:34:43");
+ assert.equal(
+ warnings.length,
+ 0,
+ "nothing is rejected: this is exactly the output the mode exists to accept",
+ );
+});
+
+test("absolute: the same 00:03:00 is rejected as out-of-range", () => {
+ const { chapters, warnings } = parseChapters(
+ [chunk3([{ start: "00:03:00", title: "Filing deadlines" }])],
+ LONG_CUES,
+ );
+ assert.equal(chapters.length, 0);
+ const w = warnings.find((x) => x.code === "out-of-range");
+ assert.ok(w, "the per-chunk clamp is what caught the measured failure");
+ assert.equal(w?.value, "00:03:00");
+ assert.match(String(w?.detail), /01:31:43/);
+});
+
+test("chunk-local: the measured nine-start chunk yields chapters instead of nothing", () => {
+ // The first three of the nine starts the model actually emitted.
+ const emitted = ["00:00:00", "00:03:54", "00:12:26"];
+ const absolute = parseChapters(
+ [chunk3(emitted.map((start, i) => ({ start, title: `Topic ${i}` })))],
+ LONG_CUES,
+ );
+ assert.equal(absolute.chapters.length, 0, "the measured outcome: a lost chunk");
+ assert.equal(
+ absolute.warnings.filter((w) => w.code === "out-of-range").length,
+ 3,
+ );
+
+ const local = parseChapters(
+ [
+ chunk3(
+ emitted.map((start, i) => ({ start, title: `Topic ${i}` })),
+ "chunk-local",
+ ),
+ ],
+ LONG_CUES,
+ );
+ assert.equal(local.chapters.length, 3);
+ assert.deepEqual(
+ local.chapters.map((c) => c.clock),
+ ["01:31:43", "01:35:37", "01:44:09"],
+ );
+});
+
+test("chunk-local does not weaken the range clamp", () => {
+ // 00:50:00 local is 02:21:43 absolute — past the chunk's real end, so it is
+ // still rejected. The guard checks the chunk's REAL range either way.
+ const { chapters, warnings } = parseChapters(
+ [chunk3([{ start: "00:50:00", title: "Past the end" }], "chunk-local")],
+ LONG_CUES,
+ );
+ assert.equal(chapters.length, 0);
+ const w = warnings.find((x) => x.code === "out-of-range");
+ assert.ok(w);
+ // The warning reports REAL video time, with the model's own value recorded
+ // alongside it so a prompt regression stays diagnosable from the artifact.
+ assert.equal(w?.value, "02:21:43");
+ assert.match(String(w?.detail), /model emitted 00:50:00/);
+});
+
+test("chunk-local: monotonicity is still checked, in real video time", () => {
+ const { chapters, warnings } = parseChapters(
+ [
+ chunk3(
+ [
+ { start: "00:10:00", title: "Second thing" },
+ { start: "00:02:00", title: "Backwards" },
+ ],
+ "chunk-local",
+ ),
+ ],
+ LONG_CUES,
+ );
+ assert.deepEqual(
+ chapters.map((c) => c.title),
+ ["Second thing"],
+ );
+ const w = warnings.find((x) => x.code === "non-monotonic");
+ assert.ok(w);
+ assert.equal(w?.value, "01:33:43");
+});
+
+test("a chunk starting at 0 is identical under both modes", () => {
+ const entries = [{ start: "00:01:00", title: "Opening" }];
+ const abs = parseChapters([chunk(entries)], CUES);
+ const local = parseChapters(
+ [{ ...chunk(entries), timestampMode: "chunk-local" as const }],
+ CUES,
+ );
+ assert.deepEqual(abs.chapters, local.chapters);
+ assert.deepEqual(abs.warnings, local.warnings);
+});
diff --git a/common/lib/digestParse.ts b/common/lib/digestParse.ts
@@ -0,0 +1,354 @@
+// The correctness layer between a local 7B model and the artifact on disk.
+//
+// Every guard here is traceable to a MEASURED failure, not to defensive
+// instinct. On a real transcript, before the prompt/schema were hardened,
+// qwen2.5:7b produced: a malformed stamp (":00:27"), 11 of 11 starts outside the
+// input's range (one 55 minutes past the end), and titles that drifted into
+// Chinese. The hardened schema eliminated all three at the decoder — but a parser
+// that trusts the decoder is a parser that breaks the day an engine ignores the
+// schema, so each rule is enforced here too.
+//
+// Nothing is ever silently dropped. Every rejection lands in `warnings[]` with
+// the offending value, because a multi-week sweep is only tunable if its failures
+// are inspectable, and the Phase 11a review queue is built entirely on that array.
+//
+// Pure (no I/O) — unit-tested in digestParse.test.ts.
+
+import {
+ chapterId,
+ tagId,
+ type DigestChapter,
+ type DigestTag,
+ type DigestTimestampMode,
+ type DigestWarning,
+} from "./digest";
+import {
+ HMS_RE,
+ hmsToSeconds,
+ promptOffsetSeconds,
+ toHms,
+} from "./digestPrompt";
+import type { Cue } from "./vtt";
+
+// One engine call's output, with the range that call was RESPONSIBLE for. The
+// range is per-chunk on purpose: a whole-video check would have accepted the
+// measured 01:10:29 for a 15-minute input, because the video was 176 minutes long.
+export type DigestChunkOutput = {
+ index: number;
+ startSeconds: number;
+ endSeconds: number;
+ // Whatever the engine returned. Unvalidated by construction.
+ data: unknown;
+ // Which numbering the engine was PROMPTED in. Must match what the caller
+ // rendered the chunk with; digestVideo passes the same value to both. Defaults
+ // to "absolute", so every existing caller is unaffected.
+ //
+ // Under "chunk-local" the model counted from 00:00:00 and the parser adds
+ // startSeconds back BEFORE any guard runs. That ordering is deliberate: the
+ // range clamp then still checks the chunk's real range (unchanged in
+ // strength), monotonicity is still compared in real video time, the cue snap
+ // still lands on a real cue, and every warning still names a real video time
+ // rather than an offset the reader would have to undo by hand.
+ timestampMode?: DigestTimestampMode;
+};
+
+// How far outside its own range a chunk's start may fall before it is rejected.
+// Small and non-zero: a cue that begins a hair before the slice's first cue start
+// is a rounding artifact, not a hallucination.
+const RANGE_TOLERANCE_SECONDS = 2;
+
+// Two chapters closer than this, with equivalent titles, are the same chapter
+// seen through the chunk overlap. Sized to the overlap (40 cues ≈ 2-4 minutes of
+// speech) so a genuine topic change at a 3-minute gap survives.
+const SEAM_DEDUP_SECONDS = 90;
+
+// Scripts that mean the model stopped writing English. The prompt and system
+// message both demand English titles, so ANY character from one of these is
+// drift, not a loanword — and drift is the signal that a chunk confused the
+// model, which makes the rest of its output for that chunk suspect too.
+const NON_LATIN_RE =
+ /[\p{Script=Han}\p{Script=Hiragana}\p{Script=Katakana}\p{Script=Hangul}\p{Script=Cyrillic}\p{Script=Arabic}\p{Script=Hebrew}\p{Script=Devanagari}\p{Script=Thai}\p{Script=Greek}]/u;
+
+export type ParsedChapters = {
+ chapters: DigestChapter[];
+ warnings: DigestWarning[];
+};
+
+export type ParsedTags = {
+ tags: DigestTag[];
+ warnings: DigestWarning[];
+};
+
+// GUARD 5 — snap a model-emitted second to the nearest real cue boundary.
+//
+// A chapter must start where someone actually starts speaking, or the viewer
+// jumps into the middle of a sentence. Hand-rolled binary search over cue starts,
+// the same shape as findActiveIndex in TranscriptModal.tsx. (resolveCitationSeconds
+// in the viewer scans SNIPPETS — the ~5 matched lines per video — not cues, so it
+// cannot be reused here.)
+export function snapToCueStart(cues: Cue[], seconds: number): number {
+ if (cues.length === 0) return Math.max(0, Math.floor(seconds));
+ let lo = 0;
+ let hi = cues.length - 1;
+ let found = -1;
+ while (lo <= hi) {
+ const mid = (lo + hi) >> 1;
+ if (cues[mid].start <= seconds) {
+ found = mid;
+ lo = mid + 1;
+ } else {
+ hi = mid - 1;
+ }
+ }
+ // Before the first cue: the first cue is the only sensible boundary.
+ if (found < 0) return Math.max(0, Math.floor(cues[0].start));
+ const before = cues[found].start;
+ const after = found + 1 < cues.length ? cues[found + 1].start : null;
+ if (after !== null && after - seconds < seconds - before) {
+ return Math.max(0, Math.floor(after));
+ }
+ return Math.max(0, Math.floor(before));
+}
+
+// Titles are compared for seam de-dup after this normalization, so "Court filing
+// deadlines" and "court filing deadlines." collapse.
+function normalizeTitle(title: string): string {
+ return title
+ .toLowerCase()
+ .replace(/[^\p{L}\p{N}]+/gu, " ")
+ .trim();
+}
+
+function titlesEquivalent(a: string, b: string): boolean {
+ const na = normalizeTitle(a);
+ const nb = normalizeTitle(b);
+ if (!na || !nb) return false;
+ if (na === nb) return true;
+ // The overlap frequently yields one call's fuller phrasing of the other's.
+ return na.includes(nb) || nb.includes(na);
+}
+
+type RawChapter = { start: unknown; title: unknown };
+
+function readRawChapters(data: unknown): RawChapter[] | null {
+ if (!data || typeof data !== "object") return null;
+ const arr = (data as { chapters?: unknown }).chapters;
+ if (!Array.isArray(arr)) return null;
+ return arr as RawChapter[];
+}
+
+// Parse ONE chunk's chapters, applying guards 1-4 in the order that keeps the
+// warnings legible: shape, then range, then language, then monotonicity.
+function parseChapterChunk(
+ chunk: DigestChunkOutput,
+ cues: Cue[],
+): ParsedChapters {
+ const warnings: DigestWarning[] = [];
+ const warn = (
+ code: DigestWarning["code"],
+ value?: string,
+ detail?: string,
+ ): void => {
+ warnings.push({
+ code,
+ section: "chapters",
+ chunk: chunk.index,
+ ...(value !== undefined ? { value } : {}),
+ ...(detail !== undefined ? { detail } : {}),
+ });
+ };
+
+ const raw = readRawChapters(chunk.data);
+ if (raw === null) {
+ warn("parse-failed", JSON.stringify(chunk.data ?? null).slice(0, 200));
+ return { chapters: [], warnings };
+ }
+ if (raw.length === 0) {
+ warn("empty-output");
+ return { chapters: [], warnings };
+ }
+
+ // Added back to every emitted stamp before the guards run. Zero unless the
+ // chunk was prompted in chunk-local numbering.
+ const offset = promptOffsetSeconds(chunk);
+ const lo = chunk.startSeconds - RANGE_TOLERANCE_SECONDS;
+ const hi = chunk.endSeconds + RANGE_TOLERANCE_SECONDS;
+ const kept: DigestChapter[] = [];
+ // Monotonicity is checked against the model's EMITTED order, not sorted order:
+ // a chunk that jumps backwards has lost track of where it is, and that entry is
+ // the suspect one.
+ let lastStart = -1;
+
+ for (const entry of raw) {
+ const rawStart = typeof entry?.start === "string" ? entry.start : "";
+ const rawTitle = typeof entry?.title === "string" ? entry.title.trim() : "";
+
+ // GUARD 1 — timestamp shape. The schema pins this, so a hit here means an
+ // engine ignored the schema; recording it is how we'd find that out.
+ if (!HMS_RE.test(rawStart)) {
+ warn("malformed-timestamp", rawStart || String(entry?.start));
+ continue;
+ }
+ const emitted = hmsToSeconds(rawStart);
+ if (emitted === null) {
+ warn("malformed-timestamp", rawStart);
+ continue;
+ }
+ // From here on everything is in REAL VIDEO TIME. `clock` follows suit, so a
+ // stored chapter never carries a stamp in one numbering next to a `start` in
+ // the other. In absolute mode both are identical to what the model emitted.
+ const seconds = emitted + offset;
+ const clock = toHms(seconds);
+ // Only worth saying when the two differ, i.e. chunk-local.
+ const emittedNote = clock === rawStart ? "" : ` (model emitted ${rawStart})`;
+
+ if (!rawTitle) {
+ warn("empty-title", clock);
+ continue;
+ }
+
+ // GUARD 2 — language drift.
+ if (NON_LATIN_RE.test(rawTitle)) {
+ warn("language-drift", rawTitle);
+ continue;
+ }
+
+ // GUARD 3 — per-chunk range clamp, always against the chunk's REAL range.
+ // Under chunk-local this is the guard the offset is designed to let the
+ // model pass; it is not weakened to do so.
+ if (seconds < lo || seconds > hi) {
+ warn(
+ "out-of-range",
+ clock,
+ `outside ${toHms(chunk.startSeconds)}–${toHms(chunk.endSeconds)}${emittedNote}`,
+ );
+ continue;
+ }
+
+ // GUARD 4a — monotonic starts within the chunk.
+ if (seconds <= lastStart) {
+ warn("non-monotonic", clock, `after ${toHms(lastStart)}${emittedNote}`);
+ continue;
+ }
+ lastStart = seconds;
+
+ // GUARD 5 — snap onto a real cue boundary.
+ const snapped = snapToCueStart(cues, seconds);
+ kept.push({
+ id: chapterId(snapped),
+ start: snapped,
+ clock,
+ title: rawTitle,
+ decidedBy: "ai",
+ });
+ }
+
+ return { chapters: kept, warnings };
+}
+
+// Parse every chunk and merge, applying guard 4b (seam de-dup) across chunk
+// boundaries. Chunks are processed in `index` order so "first wins" is stable.
+export function parseChapters(
+ chunks: DigestChunkOutput[],
+ cues: Cue[],
+): ParsedChapters {
+ const warnings: DigestWarning[] = [];
+ const all: DigestChapter[] = [];
+ for (const chunk of [...chunks].sort((a, b) => a.index - b.index)) {
+ const parsed = parseChapterChunk(chunk, cues);
+ warnings.push(...parsed.warnings);
+ all.push(...parsed.chapters);
+ }
+
+ all.sort((a, b) => a.start - b.start || a.title.localeCompare(b.title));
+ const kept: DigestChapter[] = [];
+ for (const chapter of all) {
+ const prev = kept[kept.length - 1];
+ if (prev) {
+ // Snapping can land two chunks' views of one moment on the same cue.
+ if (prev.start === chapter.start) {
+ warnings.push({
+ code: "seam-duplicate",
+ section: "chapters",
+ value: chapter.title,
+ detail: `same start as "${prev.title}"`,
+ });
+ continue;
+ }
+ if (
+ chapter.start - prev.start <= SEAM_DEDUP_SECONDS &&
+ titlesEquivalent(prev.title, chapter.title)
+ ) {
+ warnings.push({
+ code: "seam-duplicate",
+ section: "chapters",
+ value: chapter.title,
+ detail: `equivalent to "${prev.title}" ${chapter.start - prev.start}s earlier`,
+ });
+ continue;
+ }
+ }
+ kept.push(chapter);
+ }
+
+ return { chapters: kept, warnings };
+}
+
+// ---------------------------------------------------------------------------
+// Tags
+// ---------------------------------------------------------------------------
+
+export function parseTags(chunks: DigestChunkOutput[]): ParsedTags {
+ const warnings: DigestWarning[] = [];
+ const byId = new Map<string, DigestTag>();
+ for (const chunk of [...chunks].sort((a, b) => a.index - b.index)) {
+ const data = chunk.data;
+ const arr =
+ data && typeof data === "object"
+ ? (data as { tags?: unknown }).tags
+ : undefined;
+ if (!Array.isArray(arr)) {
+ warnings.push({
+ code: "parse-failed",
+ section: "tags",
+ chunk: chunk.index,
+ value: JSON.stringify(data ?? null).slice(0, 200),
+ });
+ continue;
+ }
+ if (arr.length === 0) {
+ warnings.push({ code: "empty-output", section: "tags", chunk: chunk.index });
+ continue;
+ }
+ for (const raw of arr) {
+ const tag = typeof raw === "string" ? raw.trim().toLowerCase() : "";
+ if (!tag) {
+ warnings.push({
+ code: "empty-title",
+ section: "tags",
+ chunk: chunk.index,
+ value: String(raw),
+ });
+ continue;
+ }
+ if (NON_LATIN_RE.test(tag)) {
+ warnings.push({
+ code: "language-drift",
+ section: "tags",
+ chunk: chunk.index,
+ value: tag,
+ });
+ continue;
+ }
+ const id = tagId(tag);
+ // Chunks overlap, so the same tag arrives repeatedly — that is expected,
+ // not a failure, so it is deduped WITHOUT a warning (unlike a chapter seam
+ // duplicate, which indicates a real segmentation ambiguity).
+ if (!byId.has(id)) byId.set(id, { id, tag, decidedBy: "ai" });
+ }
+ }
+ return {
+ tags: Array.from(byId.values()).sort((a, b) => a.tag.localeCompare(b.tag)),
+ warnings,
+ };
+}
diff --git a/common/lib/digestPrompt.test.ts b/common/lib/digestPrompt.test.ts
@@ -0,0 +1,130 @@
+import { test } from "node:test";
+import assert from "node:assert/strict";
+import {
+ DIGEST_MAX_CUES_PER_CHUNK,
+ PROMPT_VERSION,
+ buildChapterPrompt,
+ buildTagPrompt,
+ maxCuesForContext,
+ promptOffsetSeconds,
+} from "./digestPrompt";
+import {
+ DEFAULT_DIGEST_MAX_CUES_PER_CHUNK,
+ DEFAULT_DIGEST_NUM_CTX,
+ DEFAULT_DIGEST_TIMESTAMP_MODE,
+} from "./digest";
+
+// The chunk size MUST track the configured context. Lowering num_ctx without
+// lowering it feeds ollama more transcript than its window holds, and ollama
+// truncates SILENTLY — the model then summarizes a fragment and the result reads
+// as a bad model rather than as a misconfiguration.
+
+test("the two copies of the default chunk size agree", () => {
+ // digestPrompt.ts re-exports digest.ts's constant; this pins them equal so the
+ // identity helper (which compares against digest.ts's copy) and the chunker
+ // (which uses this one) can never drift apart.
+ assert.equal(DIGEST_MAX_CUES_PER_CHUNK, DEFAULT_DIGEST_MAX_CUES_PER_CHUNK);
+});
+
+test("maxCuesForContext scales the slice with the window", () => {
+ // Ratio-pinned to the 8192 -> 600 default pair.
+ assert.equal(maxCuesForContext(8192), 600);
+ assert.equal(maxCuesForContext(16384), 1200);
+ assert.equal(maxCuesForContext(32768), 2400);
+ assert.equal(maxCuesForContext(4096), 300);
+ // Unset or nonsense falls back to the default window, never to zero cues.
+ assert.equal(maxCuesForContext(undefined), DEFAULT_DIGEST_MAX_CUES_PER_CHUNK);
+ assert.equal(maxCuesForContext(0), DEFAULT_DIGEST_MAX_CUES_PER_CHUNK);
+ // Floored, so a tiny window still yields a usable slice rather than one cue.
+ assert.equal(maxCuesForContext(256), 100);
+});
+
+test("the shipped defaults are the configuration the bake-off measured best", () => {
+ // Round 2 (plans/bakeoff/round2.md) scored qwen2.5:7b@8192/chunk-local best on
+ // zero-yield rate, chapters/hour and worst coverage gap. The decision was made
+ // and then not shipped for a while — the defaults stayed at the losing
+ // 16384/absolute pair — so this pins the two constants that carry it.
+ assert.equal(DEFAULT_DIGEST_NUM_CTX, 8192);
+ assert.equal(DEFAULT_DIGEST_MAX_CUES_PER_CHUNK, 600);
+ assert.equal(DEFAULT_DIGEST_TIMESTAMP_MODE, "chunk-local");
+});
+
+test("PROMPT_VERSION was bumped with the defaults change", () => {
+ // The pairing is the whole point. digestPromptVariant derives its string
+ // RELATIVE to the defaults, so a record written under the old defaults carries
+ // promptVariant: undefined and would compare EQUAL under the new ones —
+ // freezing bake-off-losing output into the corpus as permanently "fresh".
+ // isSectionFresh compares promptVersion, so the bump is what invalidates them.
+ assert.ok(
+ PROMPT_VERSION >= 2,
+ "changing DEFAULT_DIGEST_* without bumping PROMPT_VERSION silently skips every existing digest as fresh",
+ );
+});
+
+test("chunk-local re-bases the range the chapter prompt states", () => {
+ const base = {
+ title: "T",
+ channel: "C",
+ startSeconds: 5503, // 01:31:43 — the measured third chunk
+ endSeconds: 8169, // 02:16:09
+ transcript: "[00:00:00] hello",
+ };
+ const absolute = buildChapterPrompt(base);
+ assert.match(absolute, /This section covers 01:31:43 to 02:16:09/);
+ assert.match(absolute, /between 01:31:43 and 02:16:09 inclusive/);
+
+ const local = buildChapterPrompt({ ...base, timestampMode: "chunk-local" });
+ assert.match(local, /This section covers 00:00:00 to 00:44:26/);
+ assert.match(local, /between 00:00:00 and 00:44:26 inclusive/);
+
+ // The SPAN is unchanged, so the density target and the schema's minItems floor
+ // are identical in both modes — only the numbering moves.
+ assert.match(absolute, /44 minutes of material/);
+ assert.match(local, /44 minutes of material/);
+});
+
+test("the tag prompt is re-based the same way", () => {
+ // It sees the same rendered transcript, so a range stated in the other
+ // numbering would contradict the markers in front of it.
+ const base = {
+ title: "T",
+ channel: "C",
+ startSeconds: 5503,
+ endSeconds: 8169,
+ transcript: "[00:00:00] hello",
+ };
+ assert.match(buildTagPrompt(base), /covering 01:31:43 to 02:16:09/);
+ assert.match(
+ buildTagPrompt({ ...base, timestampMode: "chunk-local" }),
+ /covering 00:00:00 to 00:44:26/,
+ );
+});
+
+test("absolute mode renders byte-identically to the pre-timestampMode prompt", () => {
+ // The compatibility guarantee that let PROMPT_VERSION stay at 1: adding the
+ // mode must not change what the default configuration sends, or every digest
+ // on disk would have been invalidated.
+ const base = {
+ title: "T",
+ channel: "C",
+ startSeconds: 100,
+ endSeconds: 700,
+ transcript: "[00:01:40] hello",
+ };
+ assert.equal(
+ buildChapterPrompt(base),
+ buildChapterPrompt({ ...base, timestampMode: "absolute" }),
+ );
+});
+
+test("promptOffsetSeconds is zero in absolute mode by construction", () => {
+ assert.equal(promptOffsetSeconds({ startSeconds: 5503 }), 0);
+ assert.equal(
+ promptOffsetSeconds({ startSeconds: 5503, timestampMode: "absolute" }),
+ 0,
+ );
+ assert.equal(
+ promptOffsetSeconds({ startSeconds: 5503, timestampMode: "chunk-local" }),
+ 5503,
+ );
+});
diff --git a/common/lib/digestPrompt.ts b/common/lib/digestPrompt.ts
@@ -0,0 +1,294 @@
+// The prompt + JSON schema the digest engines are driven with, and the ONE
+// constant (PROMPT_VERSION) that invalidates generated sections when either
+// changes. Pure — no I/O — so it can be unit-tested and imported anywhere.
+//
+// WHY THE SCHEMA IS THIS STRICT. Measured on a real transcript with qwen2.5:7b:
+//
+// naive prompt, loose schema → malformed stamps (":00:27"), 11 of 11 starts
+// out of range, output drifted to Chinese
+// hardened prompt + this → 0 malformed, 0 out of range, 0 drift
+//
+// Pinning `start` to a full HH:MM:SS regex, stating the chunk's own time range in
+// the prompt, and demanding English titles eliminated EVERY correctness failure.
+// A 7B model will not voluntarily honor a text contract; schema-constrained
+// decoding is what makes the local lane usable. Do not loosen the pattern.
+//
+// Bump PROMPT_VERSION for any change to the prompt text, the schema, or the
+// chunking constants below — a section whose recorded promptVersion differs is
+// regenerated, and a section whose version matches is skipped. That is what
+// keeps a re-run minutes long instead of weeks.
+import {
+ DEFAULT_DIGEST_MAX_CUES_PER_CHUNK,
+ DEFAULT_DIGEST_NUM_CTX,
+ type DigestTimestampMode,
+} from "./digest";
+
+export const PROMPT_VERSION = 2;
+
+// Version 1 -> 2 is a DEFAULTS change, and the bump is what makes it honest.
+//
+// The measured-best configuration (chunk-local timestamps, an 8192 context and
+// the 600-cue chunk sized to it) is now the default. digestPromptVariant()
+// derives its string RELATIVE TO THE DEFAULT CONSTANTS, so a value equal to the
+// default contributes nothing and the variant comes out `undefined`. That is
+// correct and stays as it is — but it means flipping the defaults alone would
+// have been silently destructive: the 39 pilot digests were generated under
+// absolute/1200 and recorded `promptVariant: undefined`, so under the new
+// defaults they would compare EQUAL and be skipped as fresh forever. Output from
+// the configuration the bake-off measured as worst would have been frozen into
+// the corpus, indistinguishable from output of the best one.
+//
+// isSectionFresh compares promptVersion, so bumping it invalidates every section
+// explicitly. That is the right tool for a default change; promptVariant is the
+// tool for comparing several shapes concurrently WITHOUT invalidating the
+// corpus, which is what a bake-off round needs. Cost here is 39 regenerations.
+//
+// (Version 1 also predates the addition of timestampMode. In "absolute" mode the
+// rendered prompt was byte-identical to what version 1 always produced, which is
+// why adding the knob did not itself require a bump.)
+
+// Chunking. ~6k usable tokens of transcript per call inside the default 8k
+// context leaves room for the prompt and the response. Expressed in CUES because
+// that is what the chunker slices; ~10 tokens/cue is the corpus average, so 600
+// cues ≈ 6k tokens. The overlap exists so a topic straddling a seam is visible
+// whole to at least one call; the parser de-dups the resulting near-identical
+// chapters.
+export const DIGEST_MAX_CUES_PER_CHUNK = DEFAULT_DIGEST_MAX_CUES_PER_CHUNK;
+
+// Cues per chunk, SIZED TO THE CONFIGURED CONTEXT.
+//
+// The 600-cue default is sized for the default 8k window (~10 tokens/cue -> ~6k
+// tokens of transcript, leaving room for the prompt and the response). Lowering
+// `numCtx` without lowering this feeds the engine more transcript than its
+// window holds, and ollama TRUNCATES SILENTLY — the model then summarizes
+// whatever fragment survived and the result looks like a bad model rather than a
+// misconfiguration. That failure is measured (see FACTS.md, smoke test 1) and it
+// is why this derivation exists rather than a bare constant.
+//
+// Halving the context is close to free on throughput — measured 24.2 vs 25.1
+// projected sweep days — because twice as many calls each carry half the prompt.
+export function maxCuesForContext(numCtx?: number): number {
+ const ctx = numCtx && numCtx > 0 ? numCtx : DEFAULT_DIGEST_NUM_CTX;
+ return Math.max(
+ 100,
+ Math.round((DIGEST_MAX_CUES_PER_CHUNK * ctx) / DEFAULT_DIGEST_NUM_CTX),
+ );
+}
+export const DIGEST_OVERLAP_CUES = 40;
+
+// Segmentation density target. 3 chapters for 22 minutes (the measured
+// hardened-prompt result) is too thin to be useful, so the prompt states an
+// explicit rate and the schema carries a matching minItems floor. This is the
+// open tuning question Stage B exists to settle — it is a knob, not a fact.
+export const DIGEST_MINUTES_PER_CHAPTER = 4;
+// Never demand more than this from one chunk, however long it is: an unreachable
+// minItems floor makes a constrained decoder pad with junk.
+export const DIGEST_MAX_CHAPTERS_PER_CHUNK = 24;
+
+// The regex that eliminated every malformed timestamp. Two digits per field, so
+// ":00:27" and "1:2:3" are both rejected by the decoder itself.
+export const HMS_PATTERN = "^[0-9][0-9]:[0-9][0-9]:[0-9][0-9]$";
+// The parser's own copy of the same rule (guard 1): the schema pins it, but the
+// parser must still reject a bad stamp in case an engine ignores the schema.
+export const HMS_RE = /^[0-9][0-9]:[0-9][0-9]:[0-9][0-9]$/;
+
+export function hmsToSeconds(clock: string): number | null {
+ if (!HMS_RE.test(clock)) return null;
+ const [h, m, s] = clock.split(":").map(Number);
+ if (m > 59 || s > 59) return null;
+ return h * 3600 + m * 60 + s;
+}
+
+// Zero-padded HH:MM:SS. Deliberately NOT aiHandoff's hms(), which drops the hour
+// field for short videos ("2:36") — the schema pattern requires all three fields.
+export function toHms(totalSeconds: number): string {
+ const n = Math.max(0, Math.floor(totalSeconds));
+ const h = Math.floor(n / 3600);
+ const m = Math.floor((n % 3600) / 60);
+ const s = n % 60;
+ return [h, m, s].map((v) => String(v).padStart(2, "0")).join(":");
+}
+
+export function minChaptersForSpan(spanSeconds: number): number {
+ const byRate = Math.floor(spanSeconds / 60 / DIGEST_MINUTES_PER_CHAPTER);
+ return Math.max(1, Math.min(DIGEST_MAX_CHAPTERS_PER_CHUNK, byRate));
+}
+
+export function maxChaptersForSpan(spanSeconds: number): number {
+ return Math.max(
+ minChaptersForSpan(spanSeconds) + 2,
+ Math.min(
+ DIGEST_MAX_CHAPTERS_PER_CHUNK,
+ Math.ceil(spanSeconds / 60 / Math.max(1, DIGEST_MINUTES_PER_CHAPTER - 2)),
+ ),
+ );
+}
+
+// ---------------------------------------------------------------------------
+// Chapters
+// ---------------------------------------------------------------------------
+
+export type ChapterPromptInput = {
+ title: string;
+ channel: string;
+ // The chunk's OWN range, in seconds, in REAL video time. Stating it in the
+ // prompt is one of the three changes that fixed out-of-range output.
+ startSeconds: number;
+ endSeconds: number;
+ // The transcript slice, already rendered as `[HH:MM:SS] text` lines by
+ // transcriptToMarkdown with stampForCue.
+ transcript: string;
+ // Optional per-channel context note (Phase 1.5). Plumbed from the start so
+ // adding notes later doesn't invalidate the corpus — see contextHash.
+ contextNote?: string;
+ // Which numbering the caller rendered `transcript` with. Defaults to
+ // "absolute". Under "chunk-local" the caller has re-based every marker to
+ // 00:00:00, so the range this prompt states must be re-based to match — a
+ // prompt that says "01:31:43 to 02:16:09" over markers that start at 00:00:00
+ // would contradict itself and is worse than either mode alone.
+ timestampMode?: DigestTimestampMode;
+};
+
+// The offset the caller subtracted from every marker, and that the parser must
+// add back. Zero in absolute mode by construction.
+export function promptOffsetSeconds(input: {
+ startSeconds: number;
+ timestampMode?: DigestTimestampMode;
+}): number {
+ return input.timestampMode === "chunk-local"
+ ? Math.max(0, Math.floor(input.startSeconds))
+ : 0;
+}
+
+export function chapterSchema(spanSeconds: number): Record<string, unknown> {
+ return {
+ type: "object",
+ properties: {
+ chapters: {
+ type: "array",
+ minItems: minChaptersForSpan(spanSeconds),
+ maxItems: maxChaptersForSpan(spanSeconds),
+ items: {
+ type: "object",
+ properties: {
+ // The pin. Every correctness failure measured before this existed.
+ start: { type: "string", pattern: HMS_PATTERN },
+ title: { type: "string", minLength: 3, maxLength: 90 },
+ },
+ required: ["start", "title"],
+ },
+ },
+ },
+ required: ["chapters"],
+ };
+}
+
+export const CHAPTER_SYSTEM_PROMPT = [
+ "You segment transcripts into chapters.",
+ "You reply with JSON only, matching the provided schema exactly.",
+ "Every title you write is in ENGLISH, regardless of the transcript's language.",
+ "You never invent a timestamp: every start you emit is copied from a",
+ "[HH:MM:SS] marker that appears in the transcript you were given.",
+].join(" ");
+
+export function buildChapterPrompt(input: ChapterPromptInput): string {
+ const offset = promptOffsetSeconds(input);
+ const from = toHms(input.startSeconds - offset);
+ const to = toHms(input.endSeconds - offset);
+ const span = Math.max(0, input.endSeconds - input.startSeconds);
+ const minItems = minChaptersForSpan(span);
+ const lines: string[] = [];
+
+ lines.push(
+ `Below is one section of the transcript of "${input.title}" (${input.channel}).`,
+ );
+ lines.push("");
+ lines.push(
+ `This section covers ${from} to ${to} — ${Math.round(span / 60)} minutes of material.`,
+ );
+ lines.push(
+ `EVERY start you emit MUST be between ${from} and ${to} inclusive. A start outside`,
+ `that range is wrong even if the topic is real. Copy starts from the [HH:MM:SS]`,
+ "markers in the transcript; do not compute or estimate them.",
+ );
+ lines.push("");
+ lines.push(
+ `Aim for roughly one chapter per ${DIGEST_MINUTES_PER_CHAPTER} minutes of material —`,
+ `at least ${minItems} for this section. A chapter marks where the subject genuinely`,
+ "changes; do not split one continuous discussion into several chapters, and do not",
+ "merge unrelated subjects into one.",
+ );
+ lines.push("");
+ lines.push(
+ "Each title is a specific, concrete English noun phrase naming what is discussed",
+ '(e.g. "Court filing deadlines" — not "Discussion" or "Part two"). Do not use the',
+ "speaker's own words as a quote, and do not editorialize.",
+ );
+ if (input.contextNote?.trim()) {
+ lines.push("");
+ lines.push("Context for this channel (use it for names and recurring topics):");
+ lines.push(input.contextNote.trim());
+ }
+ lines.push("");
+ lines.push("Transcript section:");
+ lines.push("");
+ lines.push(input.transcript);
+ return lines.join("\n");
+}
+
+// ---------------------------------------------------------------------------
+// Tags
+// ---------------------------------------------------------------------------
+
+export const TAG_MIN_ITEMS = 3;
+export const TAG_MAX_ITEMS = 12;
+
+export function tagSchema(): Record<string, unknown> {
+ return {
+ type: "object",
+ properties: {
+ tags: {
+ type: "array",
+ minItems: TAG_MIN_ITEMS,
+ maxItems: TAG_MAX_ITEMS,
+ items: { type: "string", minLength: 2, maxLength: 40 },
+ },
+ },
+ required: ["tags"],
+ };
+}
+
+export const TAG_SYSTEM_PROMPT = [
+ "You extract topic tags from transcripts.",
+ "You reply with JSON only, matching the provided schema exactly.",
+ "Every tag is in ENGLISH, lowercase, and is a topic — not a sentence,",
+ "not a summary, and not a person's opinion of the topic.",
+].join(" ");
+
+export function buildTagPrompt(input: ChapterPromptInput): string {
+ // Re-based the same way as the chapter prompt: the tag prompt sees the SAME
+ // rendered transcript, so a range stated in the other numbering would
+ // contradict the markers in front of it.
+ const offset = promptOffsetSeconds(input);
+ const lines: string[] = [];
+ lines.push(
+ `Below is one section of the transcript of "${input.title}" (${input.channel}),`,
+ `covering ${toHms(input.startSeconds - offset)} to ${toHms(input.endSeconds - offset)}.`,
+ );
+ lines.push("");
+ lines.push(
+ `List between ${TAG_MIN_ITEMS} and ${TAG_MAX_ITEMS} lowercase English topic tags for`,
+ "what this section is ABOUT. Prefer the specific over the generic: name the",
+ "subject, event, or field, not the format of the video.",
+ );
+ if (input.contextNote?.trim()) {
+ lines.push("");
+ lines.push("Context for this channel:");
+ lines.push(input.contextNote.trim());
+ }
+ lines.push("");
+ lines.push("Transcript section:");
+ lines.push("");
+ lines.push(input.transcript);
+ return lines.join("\n");
+}
diff --git a/common/lib/digests.test.ts b/common/lib/digests.test.ts
@@ -0,0 +1,180 @@
+import { test } from "node:test";
+import assert from "node:assert/strict";
+import {
+ digestPageFileName,
+ manifestHasDigest,
+ newestGeneratedAt,
+ type ChannelDigestsManifest,
+ type VideoDigest,
+} from "./digests";
+import {
+ effectiveDigest,
+ type DigestOverrides,
+ type DigestRecord,
+} from "./digest";
+
+// Run with:
+// pnpm --filter yt-dlp-transcript-common exec tsx --test common/lib/digests.test.ts
+
+test("digestPageFileName: zero-padded to 4, matching every other page tree", () => {
+ assert.equal(digestPageFileName(0), "page-0000.json");
+ assert.equal(digestPageFileName(7), "page-0007.json");
+ assert.equal(digestPageFileName(1234), "page-1234.json");
+});
+
+function manifest(
+ slugToPage: Record<string, number>,
+): ChannelDigestsManifest {
+ return {
+ version: 1,
+ channelSlug: "alice",
+ pageCount: 1,
+ maxPageBytes: 1_000_000,
+ generatedAt: "2026-07-29T00:00:00.000Z",
+ slugToPage,
+ };
+}
+
+test("manifestHasDigest: slugToPage IS the existence check", () => {
+ const m = manifest({ vaaa: 0, vbbb: 0 });
+ assert.equal(manifestHasDigest(m, "vaaa"), true);
+ // Page 0 is a real page — a falsy page index must not read as "absent".
+ assert.equal(manifestHasDigest(m, "vbbb"), true);
+ assert.equal(manifestHasDigest(m, "vccc"), false);
+ // No manifest at all (channel has zero digests) is simply "no".
+ assert.equal(manifestHasDigest(null, "vaaa"), false);
+});
+
+test("newestGeneratedAt: max across sections, tolerant of missing ones", () => {
+ assert.equal(
+ newestGeneratedAt({
+ chapters: { generatedAt: "2026-07-01T00:00:00.000Z" },
+ tags: { generatedAt: "2026-07-20T00:00:00.000Z" },
+ }),
+ "2026-07-20T00:00:00.000Z",
+ );
+ // One section only.
+ assert.equal(
+ newestGeneratedAt({ chapters: { generatedAt: "2026-07-01T00:00:00.000Z" } }),
+ "2026-07-01T00:00:00.000Z",
+ );
+ // Nothing to date — the cache treats "" as "always a miss", which is the safe
+ // direction: it re-fetches rather than serving something it can't version.
+ assert.equal(newestGeneratedAt({}), "");
+});
+
+// --- the shipped record is the COMPOSED digest, not the raw sidecar ---------
+// This is the property the whole build edge depends on: what reaches a page
+// file is effectiveDigest(machine, overrides), so a human correction ships and
+// a human rejection does not.
+
+function record(): DigestRecord {
+ return {
+ digestSchemaVersion: 1,
+ promptVersion: 2,
+ contextHash: "",
+ warnings: [
+ { code: "out-of-range", section: "chapters", chunk: 2, value: "03:11:00" },
+ ],
+ sections: {
+ chapters: {
+ provenance: {
+ appId: "ollama-direct",
+ model: "qwen2.5:7b",
+ lane: "local-gpu",
+ generatedAt: "2026-07-20T00:00:00.000Z",
+ promptVersion: 2,
+ contextHash: "",
+ },
+ items: [
+ { id: "c300", start: 300, clock: "00:05:00", title: "Second", decidedBy: "ai" },
+ { id: "c0", start: 0, clock: "00:00:00", title: "Machine title", decidedBy: "ai" },
+ { id: "c900", start: 900, clock: "00:15:00", title: "Rejected", decidedBy: "ai" },
+ ],
+ },
+ },
+ };
+}
+
+// Mirror of the projection buildIndex.ts performs when it stores a record.
+function toVideoDigest(slug: string, id: string): VideoDigest {
+ const eff = effectiveDigest(record(), {
+ version: 1,
+ // Authored as PATCHES (id + title, no start/clock) — the shape
+ // digest-server.ts's sanitizeChapters actually emits for a hand-written
+ // overrides file, and which it casts the same way. mergeItems is built to
+ // spread the machine item underneath, so the snapped start survives.
+ chapters: [
+ // Replaces a generated item by id (a retitle)…
+ { id: "c0", title: "Human title", decidedBy: "human" },
+ // …and suppresses another without deleting it.
+ { id: "c900", title: "", decidedBy: "human", enabled: false },
+ ] as DigestOverrides["chapters"],
+ });
+ const provenance = {
+ ...(eff.sections.chapters
+ ? { chapters: eff.sections.chapters.provenance }
+ : {}),
+ ...(eff.sections.tags ? { tags: eff.sections.tags.provenance } : {}),
+ };
+ return {
+ slug,
+ id,
+ chapters: eff.chapters.map((c) => ({
+ id: c.id,
+ start: c.start,
+ clock: c.clock,
+ title: c.title,
+ decidedBy: c.decidedBy,
+ })),
+ tags: eff.tags.map((t) => ({ id: t.id, tag: t.tag, decidedBy: t.decidedBy })),
+ generatedAt: newestGeneratedAt(provenance),
+ provenance,
+ ...(eff.derivedFrom ? { derivedFrom: eff.derivedFrom } : {}),
+ };
+}
+
+test("shipped digest: overrides applied, suppressed items dropped, sorted by start", () => {
+ const d = toVideoDigest("alice/vaaa", "vaaa");
+ assert.deepEqual(
+ d.chapters.map((c) => c.title),
+ ["Human title", "Second"],
+ );
+ // Sorted by start, not by the order the model emitted them.
+ assert.deepEqual(d.chapters.map((c) => c.start), [0, 300]);
+ // A retitled chapter keeps the machine's snapped start — an override authored
+ // as a patch (id + title) must not blank it, or the seek target is lost.
+ assert.equal(d.chapters[0].start, 0);
+ assert.equal(d.chapters[0].decidedBy, "human");
+ assert.equal(d.chapters[1].decidedBy, "ai");
+});
+
+test("shipped digest: warnings are operator telemetry and never ship", () => {
+ const d = toVideoDigest("alice/vaaa", "vaaa") as VideoDigest &
+ Record<string, unknown>;
+ assert.equal(d.warnings, undefined);
+ assert.equal(d.history, undefined);
+ assert.equal(JSON.stringify(d).includes("out-of-range"), false);
+});
+
+test("shipped digest: provenance and the cache version come through", () => {
+ const d = toVideoDigest("alice/vaaa", "vaaa");
+ assert.equal(d.provenance.chapters?.model, "qwen2.5:7b");
+ assert.equal(d.provenance.chapters?.lane, "local-gpu");
+ assert.equal(d.generatedAt, "2026-07-20T00:00:00.000Z");
+});
+
+test("shipped digest: derivedFrom survives so a borrowed digest stays borrowed", () => {
+ const shared: DigestRecord = {
+ ...record(),
+ derivedFrom: {
+ slug: "bob/vzzz",
+ clusterId: "cl1",
+ sharedAt: "2026-07-21T00:00:00.000Z",
+ offsetSeconds: 0.4,
+ },
+ };
+ const eff = effectiveDigest(shared, null);
+ assert.equal(eff.derivedFrom?.slug, "bob/vzzz");
+ assert.equal(eff.derivedFrom?.offsetSeconds, 0.4);
+});
diff --git a/common/lib/digests.ts b/common/lib/digests.ts
@@ -0,0 +1,160 @@
+// The SHIPPED shape of the AI-digest corpus: a per-channel paginated page tree
+// at /digests/<slug>/{manifest,page-NNNN}.json, mirroring common/lib/manifest.ts
+// (transcripts) and common/lib/posts.ts (posts).
+//
+// This is deliberately NOT DigestRecord (common/lib/digest.ts). That type is the
+// on-disk sidecar the generator owns: two files per video, machine output plus a
+// human shadow, carrying operator telemetry. What ships is the COMPOSED result
+// of effectiveDigest() — human overrides applied, suppressed items dropped,
+// sorted — which is the only version a reader should ever see.
+//
+// Two things are deliberately absent from the shipped record:
+//
+// warnings[] operator telemetry for the editor panel and the Phase 11a review
+// queue. A reader has no use for "the model proposed 13 chapters
+// that were clamped away"; shipping it would put the corpus's
+// failures in front of the audience rather than the operator.
+// history[] the regeneration audit trail. Same reasoning, plus it is the
+// single largest field in a mature sidecar.
+//
+// SPARSE BY DESIGN. A channel's manifest lists only the videos that actually
+// have a digest — 102 of ~76,000 at the time of writing. `slugToPage` is
+// therefore also the existence check: a videoId absent from it has no digest,
+// which is what lets the viewer decide whether to render its Digest control
+// without a per-video fetch (the export/app/lib/duplicates.ts hasDuplicates()
+// idiom). A channel with zero digests gets no manifest at all.
+
+import type {
+ DecidedBy,
+ DigestDerivedFrom,
+ DigestProvenance,
+} from "./digest";
+
+// v1: the initial shipped shape.
+export const DIGESTS_MANIFEST_VERSION = 1;
+
+// Zero-padded to 4 digits, the convention every other page tree uses
+// (pageFileName / transcriptPageFileName / subsPageFileName / postsPageFileName).
+export function digestPageFileName(index: number): string {
+ return `page-${String(index).padStart(4, "0")}.json`;
+}
+
+// A titled moment. `start` is SECONDS and is what a seek uses: the parser has
+// already snapped it to a real cue boundary. `clock` is the raw HH:MM:SS string
+// the model emitted, carried for auditing — NEVER re-parse it to seek, or a
+// model that miscounted gets to move the playhead.
+export type DigestPageChapter = {
+ id: string;
+ start: number;
+ clock: string;
+ title: string;
+ // "human" when an override replaced or added this item, so the viewer can be
+ // honest about which chapters a person wrote.
+ decidedBy: DecidedBy;
+};
+
+export type DigestPageTag = {
+ id: string;
+ tag: string;
+ decidedBy: DecidedBy;
+};
+
+// One video's shipped digest — the unit a page file holds an array of.
+export type VideoDigest = {
+ // "channelSlug/videoId", matching TranscriptDetail.slug.
+ slug: string;
+ // The bare video id, matching the manifest's slugToPage key.
+ id: string;
+ chapters: DigestPageChapter[];
+ tags: DigestPageTag[];
+ // Newest section generatedAt in this record (ISO-8601), or "" when no section
+ // carries one. THE PER-ENTRY VERSION: digests are regenerated in place, so a
+ // client's cached copy stays structurally valid while going stale, and a
+ // structure sniff (what transcriptStore.ts gets away with) cannot detect that.
+ // The client cache compares this field and treats a mismatch as a miss —
+ // otherwise a corrected digest never reaches a returning reader.
+ generatedAt: string;
+ // Per-SECTION provenance, because one video can legitimately hold chapters
+ // from the local lane and tags from the metered one.
+ provenance: {
+ chapters?: DigestProvenance;
+ tags?: DigestProvenance;
+ };
+ // Set when this digest was generated for a DIFFERENT video (a duplicate
+ // cluster's canonical member) and shared onto this one. Presenting a borrowed
+ // digest as native is the failure mode that looks like success: every chapter
+ // reads plausibly while describing another upload. The viewer renders this.
+ derivedFrom?: DigestDerivedFrom;
+};
+
+export type DigestPage = VideoDigest[];
+
+export type ChannelDigestsManifest = {
+ version: number;
+ channelSlug: string;
+ pageCount: number;
+ maxPageBytes: number;
+ generatedAt: string;
+ // videoId -> page index. Only DIGESTED videos appear; see the sparse note above.
+ slugToPage: Record<string, number>;
+ // Content hash of each page, indexed by page number — the CACHE VERSION.
+ //
+ // A client cannot know whether its cached copy of one video's digest is
+ // current without something to compare, and the record's own generatedAt is
+ // only legible once the page has been fetched (which is the cost the cache
+ // exists to avoid). The build already computes these hashes for its own
+ // page-skip logic, so surfacing them is free and gives exact per-page
+ // versioning: any regeneration, retitle or suppression changes the hash of
+ // the page it lands on and nothing else.
+ //
+ // Optional so a manifest written without it still parses; a client that
+ // finds it absent simply treats every cached entry as a miss.
+ pageHashes?: string[];
+};
+
+// --- site level ------------------------------------------------------------
+// The per-SITE manifest at /digests/manifest.json: which member channels carry
+// digests, and how many. Mirrors the site posts manifest, and exists for the
+// same reason — compose-site.ts reads it to populate corpus.json's per-channel
+// digests pointer without walking every channel's page tree.
+export const SITE_DIGESTS_MANIFEST_VERSION = 1;
+
+export type DigestsChannelEntry = {
+ name: string;
+ slug: string;
+ digestCount: number;
+ groupId?: string;
+};
+
+export type DigestsManifest = {
+ version: number;
+ channels: DigestsChannelEntry[];
+ totalCount: number;
+ generatedAt: string;
+ siteId?: string;
+};
+
+// Whether a manifest covers a given video. The whole existence check, in one
+// place, so the viewer's control-gating and the fetch path agree on what
+// "has a digest" means.
+export function manifestHasDigest(
+ manifest: ChannelDigestsManifest | null,
+ videoId: string,
+): boolean {
+ return manifest?.slugToPage[videoId] !== undefined;
+}
+
+// The newest generatedAt across a record's sections — the value that lands in
+// VideoDigest.generatedAt. Pure, so both the builder and any test can call it.
+export function newestGeneratedAt(provenance: {
+ chapters?: Pick<DigestProvenance, "generatedAt">;
+ tags?: Pick<DigestProvenance, "generatedAt">;
+}): string {
+ const stamps = [
+ provenance.chapters?.generatedAt,
+ provenance.tags?.generatedAt,
+ ].filter((s): s is string => typeof s === "string" && s !== "");
+ if (stamps.length === 0) return "";
+ // ISO-8601 sorts lexicographically, so no Date parsing is needed.
+ return stamps.reduce((a, b) => (a > b ? a : b));
+}
diff --git a/common/lib/duplicates.test.ts b/common/lib/duplicates.test.ts
@@ -0,0 +1,564 @@
+import { test } from "node:test";
+import assert from "node:assert/strict";
+import {
+ DEFAULT_TITLE_DURATION_RATIO,
+ TITLE_DURATION_MIN_TOLERANCE_SECONDS,
+ clusterIsPublishable,
+ clusterMaySharePartial,
+ durationsCompatible,
+ filterClusterToChannels,
+ isClusterReviewed,
+ measureAlignment,
+ normalizeTitleKey,
+ pickCanonicalSlug,
+ resolveCanonicalSlug,
+ sanitizeDuplicateOverrides,
+ strongerMatch,
+ type DuplicateCluster,
+ type DuplicateVideoRef,
+} from "./duplicates";
+import type { Cue } from "./vtt";
+
+function ref(over: Partial<DuplicateVideoRef> = {}): DuplicateVideoRef {
+ return {
+ slug: "chan/vid",
+ channelSlug: "chan",
+ channel: "Chan",
+ platform: "youtube",
+ id: "vid",
+ title: "A video",
+ duration: 600,
+ uploadDate: "20250101",
+ hasTranscript: true,
+ ...over,
+ };
+}
+
+function cluster(
+ refs: DuplicateVideoRef[],
+ over: Partial<DuplicateCluster> = {},
+): DuplicateCluster {
+ return {
+ clusterId: "cluster1",
+ matchKind: "transcript-exact",
+ score: 1,
+ contained: false,
+ durationBucket: 600,
+ crossPlatform: true,
+ crossChannel: false,
+ videoRefs: refs,
+ ...over,
+ };
+}
+
+// ---------------------------------------------------------------------------
+// Canonical selection
+// ---------------------------------------------------------------------------
+
+test("a member with a transcript beats one without", () => {
+ const c = cluster([
+ ref({ slug: "a/1", hasTranscript: false, duration: 900 }),
+ ref({ slug: "b/2", hasTranscript: true, duration: 600 }),
+ ]);
+ assert.equal(pickCanonicalSlug(c), "b/2");
+});
+
+test("the longest (most complete) member wins next", () => {
+ // Load-bearing: a mirror that cut the intro would place every shared chapter
+ // wrong, so the fullest artifact owns the digest.
+ const c = cluster([
+ ref({ slug: "a/1", duration: 600 }),
+ ref({ slug: "b/2", duration: 640 }),
+ ]);
+ assert.equal(pickCanonicalSlug(c), "b/2");
+});
+
+test("the preferred platform breaks a duration tie", () => {
+ const c = cluster([
+ ref({ slug: "a/1", platform: "rumble" }),
+ ref({ slug: "b/2", platform: "youtube" }),
+ ]);
+ assert.equal(pickCanonicalSlug(c), "b/2");
+});
+
+test("the earliest upload breaks a platform tie", () => {
+ const c = cluster([
+ ref({ slug: "a/1", uploadDate: "20250301" }),
+ ref({ slug: "b/2", uploadDate: "20250101" }),
+ ]);
+ assert.equal(pickCanonicalSlug(c), "b/2");
+});
+
+test("selection is deterministic when everything ties", () => {
+ const refs = [ref({ slug: "b/2" }), ref({ slug: "a/1" })];
+ assert.equal(pickCanonicalSlug(cluster(refs)), "a/1");
+ assert.equal(pickCanonicalSlug(cluster([...refs].reverse())), "a/1");
+});
+
+test("a human canonical override WINS over the rule", () => {
+ const c = cluster([
+ ref({ slug: "a/1", duration: 600 }),
+ ref({ slug: "b/2", duration: 900 }),
+ ]);
+ assert.equal(pickCanonicalSlug(c), "b/2", "the rule prefers the longer one");
+ const overrides = sanitizeDuplicateOverrides({
+ clusters: { cluster1: { canonicalSlug: "a/1" } },
+ });
+ assert.equal(resolveCanonicalSlug(c, overrides), "a/1");
+});
+
+test("an override naming a non-member is ignored, not obeyed", () => {
+ const c = cluster([ref({ slug: "a/1" }), ref({ slug: "b/2" })]);
+ const overrides = sanitizeDuplicateOverrides({
+ clusters: { cluster1: { canonicalSlug: "someone/else" } },
+ });
+ assert.equal(resolveCanonicalSlug(c, overrides), pickCanonicalSlug(c));
+});
+
+test("notDuplicate suppresses the cluster entirely (nothing is shared)", () => {
+ const c = cluster([ref({ slug: "a/1" }), ref({ slug: "b/2" })]);
+ const overrides = sanitizeDuplicateOverrides({
+ clusters: { cluster1: { notDuplicate: true } },
+ });
+ assert.equal(resolveCanonicalSlug(c, overrides), null);
+});
+
+test("a recorded canonicalSlug on the report is honored over the rule", () => {
+ const c = cluster(
+ [ref({ slug: "a/1", duration: 600 }), ref({ slug: "b/2", duration: 900 })],
+ { canonicalSlug: "a/1" },
+ );
+ assert.equal(resolveCanonicalSlug(c, null), "a/1");
+});
+
+test("isClusterReviewed distinguishes a decision from a bare stub", () => {
+ const c = cluster([ref({ slug: "a/1" }), ref({ slug: "b/2" })]);
+ assert.equal(isClusterReviewed(c, null), false);
+ assert.equal(
+ isClusterReviewed(
+ c,
+ sanitizeDuplicateOverrides({ clusters: { cluster1: { canonicalSlug: "a/1" } } }),
+ ),
+ true,
+ );
+ // A decidedAt-only entry is dropped by the sanitizer, so it never reads as
+ // reviewed.
+ assert.equal(
+ isClusterReviewed(
+ c,
+ sanitizeDuplicateOverrides({ clusters: { cluster1: { decidedAt: "now" } } }),
+ ),
+ false,
+ );
+});
+
+test("sanitizeDuplicateOverrides drops malformed entries rather than throwing", () => {
+ const o = sanitizeDuplicateOverrides({
+ clusters: {
+ good: { canonicalSlug: "a/1" },
+ blank: { canonicalSlug: " " },
+ junk: 42,
+ arrayish: [],
+ },
+ });
+ assert.deepEqual(Object.keys(o.clusters), ["good"]);
+});
+
+test("sanitizeDuplicateOverrides survives a non-object file", () => {
+ assert.deepEqual(sanitizeDuplicateOverrides(null).clusters, {});
+ assert.deepEqual(sanitizeDuplicateOverrides([1, 2, 3]).clusters, {});
+});
+
+// `confirmed` is the ONLY thing that lets a title+duration suspect ship or share
+// derived work. If the sanitizer dropped it, every human confirmation would be
+// silently discarded on the next read and the cluster would quietly revert to
+// "unreviewed" — a failure that leaves no trace anywhere.
+test("sanitizeDuplicateOverrides preserves confirmed", () => {
+ const o = sanitizeDuplicateOverrides({
+ clusters: { c1: { confirmed: true }, c2: { confirmed: "yes" } },
+ });
+ assert.equal(o.clusters.c1?.confirmed, true);
+ // Only a real boolean true counts; a truthy string is not a decision.
+ assert.equal(o.clusters.c2, undefined);
+});
+
+test("isClusterReviewed accepts a confirmed-only entry", () => {
+ // A confirmation with no canonical override IS a decision — it is the whole
+ // point of the review queue — so it must clear the cluster off the worklist.
+ const c = cluster([ref({ slug: "a/1" }), ref({ slug: "b/2" })]);
+ assert.equal(
+ isClusterReviewed(
+ c,
+ sanitizeDuplicateOverrides({ clusters: { cluster1: { confirmed: true } } }),
+ ),
+ true,
+ );
+});
+
+// ---------------------------------------------------------------------------
+// Suspects: needsReview gates sharing until a human confirms
+// ---------------------------------------------------------------------------
+
+test("a needsReview cluster shares nothing without a confirmation", () => {
+ // Title + near-identical runtime is a suspicion, not a comparison. Two
+ // episodes of a daily show can share both and be entirely different material,
+ // and a shared digest would then describe the wrong video convincingly.
+ const c = cluster([ref({ slug: "a/1" }), ref({ slug: "b/2" })], {
+ matchKind: "title-duration",
+ score: null,
+ needsReview: true,
+ });
+ assert.equal(clusterMaySharePartial(c, null), false);
+ assert.equal(
+ clusterMaySharePartial(
+ c,
+ sanitizeDuplicateOverrides({ clusters: { cluster1: { canonicalSlug: "a/1" } } }),
+ ),
+ false,
+ "picking a canonical member is not the same as confirming the duplicate",
+ );
+});
+
+test("confirmed unblocks a needsReview cluster", () => {
+ const c = cluster([ref({ slug: "a/1" }), ref({ slug: "b/2" })], {
+ matchKind: "title-duration",
+ score: null,
+ needsReview: true,
+ });
+ assert.equal(
+ clusterMaySharePartial(
+ c,
+ sanitizeDuplicateOverrides({ clusters: { cluster1: { confirmed: true } } }),
+ ),
+ true,
+ );
+});
+
+test("confirmed does NOT unblock a contained (clip-of-longer) cluster", () => {
+ // A clip is a different artifact from the recording it was cut from, however
+ // confident a human is that they are related.
+ const c = cluster([ref({ slug: "a/1" }), ref({ slug: "b/2" })], {
+ contained: true,
+ });
+ assert.equal(
+ clusterMaySharePartial(
+ c,
+ sanitizeDuplicateOverrides({ clusters: { cluster1: { confirmed: true } } }),
+ ),
+ false,
+ );
+});
+
+// ---------------------------------------------------------------------------
+// What reaches a built site
+// ---------------------------------------------------------------------------
+
+test("an unconfirmed suspect does NOT ship to a built site", () => {
+ // compose-site applies this. A suspect asserts a relationship no machine and
+ // no human has checked, so it stays an internal review queue.
+ const c = cluster([ref({ slug: "a/1" }), ref({ slug: "b/2" })], {
+ matchKind: "title-duration",
+ score: null,
+ needsReview: true,
+ });
+ assert.equal(clusterIsPublishable(c, null), false);
+ assert.equal(
+ clusterIsPublishable(
+ c,
+ sanitizeDuplicateOverrides({ clusters: { cluster1: { canonicalSlug: "a/1" } } }),
+ ),
+ false,
+ );
+ assert.equal(
+ clusterIsPublishable(
+ c,
+ sanitizeDuplicateOverrides({ clusters: { cluster1: { confirmed: true } } }),
+ ),
+ true,
+ );
+});
+
+test("a content-confirmed cluster ships with no overrides at all", () => {
+ const c = cluster([ref({ slug: "a/1" }), ref({ slug: "b/2" })]);
+ assert.equal(clusterIsPublishable(c, null), true);
+});
+
+test("a contained cluster ships even though it shares nothing", () => {
+ // The two predicates deliberately disagree here: "this is a clip of that" is a
+ // real relationship worth showing a viewer, and simultaneously a reason never
+ // to copy the longer video's derived work onto the clip.
+ const c = cluster([ref({ slug: "a/1" }), ref({ slug: "b/2" })], {
+ contained: true,
+ });
+ assert.equal(clusterIsPublishable(c, null), true);
+ assert.equal(clusterMaySharePartial(c, null), false);
+});
+
+// ---------------------------------------------------------------------------
+// Match tiers
+// ---------------------------------------------------------------------------
+
+test("strongerMatch ranks exact > near > title-duration", () => {
+ assert.equal(strongerMatch("transcript-near", "transcript-exact"), "transcript-exact");
+ assert.equal(strongerMatch("transcript-exact", "transcript-near"), "transcript-exact");
+ assert.equal(strongerMatch("title-duration", "transcript-near"), "transcript-near");
+ assert.equal(strongerMatch("transcript-near", "title-duration"), "transcript-near");
+ assert.equal(strongerMatch("title-duration", "title-duration"), "title-duration");
+ // Load-bearing for needsReview: a cluster is only a suspect when EVERY pair in
+ // it was untestable, so one content-confirmed pair must dominate.
+ assert.equal(strongerMatch("title-duration", "transcript-exact"), "transcript-exact");
+});
+
+// ---------------------------------------------------------------------------
+// The title blocking key
+// ---------------------------------------------------------------------------
+
+test("normalizeTitleKey folds case, punctuation and diacritics", () => {
+ assert.equal(normalizeTitleKey("Pokémon: The First Movie!"), "pokemon the first movie");
+ assert.equal(normalizeTitleKey(" Multiple spaces "), "multiple spaces");
+ assert.equal(
+ normalizeTitleKey("The Show — Episode 4"),
+ normalizeTitleKey("the show episode 4"),
+ );
+});
+
+test("normalizeTitleKey drops platform re-upload suffixes", () => {
+ // These are what a mirror ADDS to an otherwise identical title, so folding
+ // them is what makes the mirror block with its original.
+ const base = normalizeTitleKey("Weekly Roundup");
+ assert.equal(normalizeTitleKey("Weekly Roundup (reupload)"), base);
+ assert.equal(normalizeTitleKey("Weekly Roundup [mirror]"), base);
+ assert.equal(normalizeTitleKey("Weekly Roundup #shorts"), base);
+});
+
+test("normalizeTitleKey does NOT merge a series", () => {
+ // The whole reason the key is exact rather than fuzzy: a looser key merges
+ // consecutive episodes, and every such merge is a false cluster a human then
+ // has to reject.
+ assert.notEqual(normalizeTitleKey("Episode 12"), normalizeTitleKey("Episode 13"));
+ assert.notEqual(
+ normalizeTitleKey("Morning Show Jan 4"),
+ normalizeTitleKey("Morning Show Jan 5"),
+ );
+});
+
+test("durationsCompatible allows a re-encode's drift but not a different cut", () => {
+ // 2% of the longer runtime, floored at 2s.
+ assert.equal(durationsCompatible(3600, 3620), true, "20s on an hour is 0.6%");
+ assert.equal(durationsCompatible(3600, 3800), false, "200s on an hour is 5.6%");
+ // The floor keeps sub-second rounding from splitting two short clips: 2% of
+ // 30s is 0.6s, which nothing survives.
+ assert.equal(durationsCompatible(30, 31), true);
+ assert.equal(durationsCompatible(30, 40), false);
+ assert.equal(
+ durationsCompatible(1000, 1000 + 1000 * DEFAULT_TITLE_DURATION_RATIO),
+ true,
+ "exactly at the ratio is compatible",
+ );
+ assert.equal(
+ durationsCompatible(100, 100 + TITLE_DURATION_MIN_TOLERANCE_SECONDS),
+ true,
+ "exactly at the floor is compatible",
+ );
+});
+
+test("durationsCompatible refuses a missing or zero duration", () => {
+ // An unknown runtime is not a match — treating 0 as "close to 0" would block
+ // every metadata-less video together.
+ assert.equal(durationsCompatible(0, 0), false);
+ assert.equal(durationsCompatible(600, 0), false);
+ assert.equal(durationsCompatible(-5, -5), false);
+});
+
+// ---------------------------------------------------------------------------
+// Narrowing a global cluster to one site's channels
+// ---------------------------------------------------------------------------
+
+test("filterClusterToChannels keeps in-site members and recomputes the flags", () => {
+ const c = cluster(
+ [
+ ref({ slug: "a/1", channelSlug: "a", platform: "youtube" }),
+ ref({ slug: "b/2", channelSlug: "b", platform: "rumble" }),
+ ref({ slug: "c/3", channelSlug: "c", platform: "odysee" }),
+ ],
+ { crossPlatform: true, crossChannel: true },
+ );
+ const out = filterClusterToChannels(c, new Set(["a", "b"]));
+ assert.ok(out);
+ assert.deepEqual(out.videoRefs.map((r) => r.slug), ["a/1", "b/2"]);
+ assert.equal(out.crossPlatform, true);
+ assert.equal(out.crossChannel, true);
+});
+
+test("filterClusterToChannels drops a cluster that falls below two members", () => {
+ // One video is not a visible duplicate — there is nothing to switch to.
+ const c = cluster([
+ ref({ slug: "a/1", channelSlug: "a" }),
+ ref({ slug: "b/2", channelSlug: "b" }),
+ ]);
+ assert.equal(filterClusterToChannels(c, new Set(["a"])), null);
+ assert.equal(filterClusterToChannels(c, new Set(["z"])), null);
+});
+
+test("filterClusterToChannels clears cross-* flags the survivors no longer earn", () => {
+ const c = cluster(
+ [
+ ref({ slug: "a/1", channelSlug: "a", platform: "youtube" }),
+ ref({ slug: "a/2", channelSlug: "a", platform: "youtube" }),
+ ref({ slug: "b/3", channelSlug: "b", platform: "rumble" }),
+ ],
+ { crossPlatform: true, crossChannel: true },
+ );
+ const out = filterClusterToChannels(c, new Set(["a"]));
+ assert.ok(out);
+ assert.equal(out.crossPlatform, false);
+ assert.equal(out.crossChannel, false);
+});
+
+test("filterClusterToChannels carries needsReview and the alignment fields through", () => {
+ // The site-narrowing step must not launder a suspect into a shipped cluster,
+ // and it must not lose the per-member alignment the viewer's jump depends on.
+ const c = cluster(
+ [
+ ref({ slug: "a/1", channelSlug: "a", aligned: true, offsetSeconds: 0 }),
+ ref({ slug: "b/2", channelSlug: "b", aligned: false, offsetSeconds: 41.5 }),
+ ref({ slug: "c/3", channelSlug: "c" }),
+ ],
+ { matchKind: "title-duration", score: null, needsReview: true, canonicalSlug: "a/1" },
+ );
+ const out = filterClusterToChannels(c, new Set(["a", "b"]));
+ assert.ok(out);
+ assert.equal(out.needsReview, true);
+ assert.equal(out.canonicalSlug, "a/1");
+ assert.equal(out.videoRefs[0].aligned, true);
+ assert.equal(out.videoRefs[1].aligned, false);
+ assert.equal(out.videoRefs[1].offsetSeconds, 41.5);
+});
+
+test("filterClusterToChannels leaves matchKind and score as the detector reported them", () => {
+ // Intentional and documented: they describe the strongest pair in the FULL
+ // cluster. Recomputing them here would change the meaning of clusters already
+ // shipped, and there is no similarity data at this point to recompute from.
+ const c = cluster(
+ [
+ ref({ slug: "a/1", channelSlug: "a" }),
+ ref({ slug: "b/2", channelSlug: "b" }),
+ ref({ slug: "c/3", channelSlug: "c" }),
+ ],
+ { matchKind: "transcript-exact", score: 1 },
+ );
+ const out = filterClusterToChannels(c, new Set(["a", "b"]));
+ assert.ok(out);
+ assert.equal(out.matchKind, "transcript-exact");
+ assert.equal(out.score, 1);
+});
+
+// ---------------------------------------------------------------------------
+// The alignment gate
+// ---------------------------------------------------------------------------
+
+// A synthetic transcript whose sentences do NOT share long phrases, so the
+// 8-word anchors are unambiguous — which is what the gate requires and what a
+// real transcript mostly provides. (An earlier fixture repeated the same nine
+// words in every cue; the gate correctly refused to measure it, which is how the
+// ambiguous-anchor rule got written.)
+const VOCAB = [
+ "filing", "deadline", "jury", "selection", "witness", "testimony", "exhibit",
+ "objection", "sustained", "overruled", "docket", "motion", "dismissal",
+ "appeal", "verdict", "sentencing", "transcript", "counsel", "recess",
+ "subpoena", "affidavit", "discovery", "deposition", "settlement", "mediation",
+ "injunction", "damages", "liability", "negligence", "statute", "precedent",
+ "jurisdiction", "venue", "indictment", "arraignment", "plea", "bail",
+ "custody", "warrant", "evidence",
+];
+
+function makeCues(shiftSeconds = 0, count = 40): Cue[] {
+ return Array.from({ length: count }, (_, i) => {
+ // Each cue draws a distinct rotation of the vocabulary, so no 8-word window
+ // recurs anywhere in the transcript.
+ const words = Array.from(
+ { length: 12 },
+ (_, j) => VOCAB[(i * 7 + j * 3) % VOCAB.length],
+ );
+ return {
+ start: i * 15 + shiftSeconds,
+ end: i * 15 + 14 + shiftSeconds,
+ text: `${words.join(" ")} marker${i}`,
+ };
+ });
+}
+
+test("identical timings are aligned at a zero offset", () => {
+ const a = makeCues(0);
+ const result = measureAlignment(a, makeCues(0));
+ assert.equal(result.aligned, true);
+ assert.equal(result.maxOffsetSeconds, 0);
+ assert.equal(result.matchedAnchors, result.totalAnchors);
+});
+
+test("a mirror with a shifted intro is REFUSED", () => {
+ // The failure mode that looks like success: identical text, every chapter
+ // placed 40s wrong. Content similarity cannot see this — shingles are a set.
+ const result = measureAlignment(makeCues(0), makeCues(40));
+ assert.equal(result.aligned, false);
+ assert.equal(result.reason, "offset-exceeded");
+ assert.ok(result.maxOffsetSeconds >= 39 && result.maxOffsetSeconds <= 41);
+});
+
+test("a sub-tolerance shift is still aligned", () => {
+ const result = measureAlignment(makeCues(0), makeCues(2));
+ assert.equal(result.aligned, true);
+});
+
+test("the tolerance is configurable", () => {
+ assert.equal(
+ measureAlignment(makeCues(0), makeCues(8), { toleranceSeconds: 10 }).aligned,
+ true,
+ );
+ assert.equal(
+ measureAlignment(makeCues(0), makeCues(8), { toleranceSeconds: 3 }).aligned,
+ false,
+ );
+});
+
+test("unrelated transcripts are refused for want of anchors", () => {
+ const other: Cue[] = Array.from({ length: 40 }, (_, i) => ({
+ start: i * 15,
+ end: i * 15 + 14,
+ text: `zebra${i} quartz${i} lantern${i} beacon${i} pumice${i} sorrel${i} thicket${i} vellum${i} wicket${i}`,
+ }));
+ const result = measureAlignment(makeCues(0), other);
+ assert.equal(result.aligned, false);
+ assert.equal(result.reason, "too-few-anchors");
+});
+
+test("a mirror aligned at the start but drifting mid-way is refused", () => {
+ // Sampling SEVERAL anchors is what catches an ad break inserted in the middle.
+ const canonical = makeCues(0, 40);
+ const drifting = canonical.map((c, i) => ({
+ ...c,
+ start: i < 20 ? c.start : c.start + 45,
+ end: i < 20 ? c.end : c.end + 45,
+ }));
+ const result = measureAlignment(canonical, drifting);
+ assert.equal(result.aligned, false);
+ assert.equal(result.reason, "offset-exceeded");
+});
+
+test("an empty transcript is refused, not treated as aligned", () => {
+ assert.equal(measureAlignment([], makeCues(0)).aligned, false);
+ assert.equal(measureAlignment(makeCues(0), []).reason, "empty-transcript");
+});
+
+test("a contained cluster (clip of a longer video) never shares", () => {
+ // Correct even at a perfect zero offset: a clip is a different artifact, and
+ // the longer video's chapters describe material it does not contain.
+ assert.equal(
+ clusterMaySharePartial(cluster([ref()], { contained: true })),
+ false,
+ );
+ assert.equal(
+ clusterMaySharePartial(cluster([ref()], { contained: false })),
+ true,
+ );
+});
diff --git a/common/lib/duplicates.ts b/common/lib/duplicates.ts
@@ -16,7 +16,33 @@ export const DEFAULT_SHORT_THRESHOLD_SECONDS = 180;
// land in a shared comparison set (see controller).
export const DEFAULT_DURATION_TOLERANCE_SECONDS = 2;
// Phase-2 near-duplicate threshold (5-gram Jaccard).
-export const DEFAULT_NEAR_THRESHOLD = 0.6;
+//
+// 0.35, lowered from 0.6 on measurement (2026-07-29). 0.6 was tuned for
+// same-engine text and structurally failed the case detection exists to catch:
+// the two sides of a cross-platform mirror are transcribed by DIFFERENT ASR
+// engines, and a 5-gram Jaccard is unforgiving of word-level disagreement, so
+// two transcripts of the same audio land at ~0.35–0.60 rather than ≥ 0.6.
+//
+// Bracketed corpus-wide at 0.6 / 0.45 / 0.35 / 0.25 on identical inputs. Each
+// step down is a strict superset — zero videos are lost — and the returns fall
+// off a cliff: 0.6→0.45 adds 4,264 clusters, 0.45→0.35 adds 324, 0.35→0.25 adds
+// 59. The marginal band was read, not just counted: of the 4,264 admitted at
+// 0.45, 96.2% are byte-identical titles, 99.5% cross-platform, and 3 (0.07%)
+// are same-channel. The 324 admitted at 0.35 are the same shape (97.2%
+// byte-identical titles) and the riskiest 11 were inspected individually — all
+// same recording, same runtime to the second, mirrored platform.
+//
+// 0.45 is where the RECALL knee is, if a more conservative value is ever
+// wanted; 0.35 is chosen because its marginal band is still clean and the
+// report's job is to surface real mirrors to readers. 0.25 is the flat tail.
+//
+// This is NOT a meaningful digest-sweep cost lever, whatever the sharing code's
+// header says: cluster members are 19% of the corpus by video count but only
+// 4.3% of its AUDIO-HOURS (mirrors skew short, long-form VODs are unclustered),
+// so 0.6→0.35 moves the sweep from 80.0 to 78.8 days. It is a publishing change
+// — it is what the archive asserts to readers — and it is justified on that.
+// See bin/digest-plan.ts for the measurement.
+export const DEFAULT_NEAR_THRESHOLD = 0.35;
// Containment threshold for the "short is a clip of a longer video" case.
export const DEFAULT_CONTAINMENT_THRESHOLD = 0.8;
// Shingle (word n-gram) size for similarity. 5 tolerates word-level ASR
@@ -24,10 +50,41 @@ export const DEFAULT_CONTAINMENT_THRESHOLD = 0.8;
// unigrams.
export const DEFAULT_SHINGLE_SIZE = 5;
-// Only content-confirmed tiers form clusters. Duration coincidence alone is
-// never treated as a match (it produced enormous false clusters of unrelated
-// same-length videos).
-export type DuplicateMatchKind = "transcript-exact" | "transcript-near";
+// Match tiers, weakest first.
+//
+// `title-duration` is a SUSPECT, not a confirmation. Duration coincidence alone
+// was never a match — it produced enormous false clusters of unrelated
+// same-length videos — but the same title AND a near-identical runtime is a
+// different claim entirely, and it is the only signal available when one side
+// has no transcript to compare. Such a cluster is flagged `needsReview` and
+// shares nothing until a human confirms it (see clusterMaySharePartial).
+export type DuplicateMatchKind =
+ | "title-duration"
+ | "transcript-near"
+ | "transcript-exact";
+
+// Two videos with the same normalized title are candidates when their runtimes
+// agree within this fraction of the longer one. A mirror re-encode drifts by a
+// second or two; 2% also absorbs an ad-break difference on a long video without
+// admitting a genuinely different cut.
+export const DEFAULT_TITLE_DURATION_RATIO = 0.02;
+// ...but never demand tighter than this, so two 30-second clips are not split by
+// sub-second rounding.
+export const TITLE_DURATION_MIN_TOLERANCE_SECONDS = 2;
+
+// A title group larger than this is a FORMAT, not a title — "live stream",
+// "untitled", a daily show's date-less name. Pairing inside it is quadratic and
+// the matches would be noise, so the group is skipped and reported rather than
+// silently truncated.
+export const MAX_TITLE_GROUP_SIZE = 40;
+
+// The same treatment for a pathological DURATION block. Round numbers attract
+// videos — a corpus can hold thousands of videos that are exactly 60s — and one
+// such block is quadratic on its own. Blocks over this are skipped and reported,
+// for the same reason: a silent truncation reads as "covered everything".
+// Generous, because the proven shorts-mode run has legitimately large buckets
+// (34,915 shorts over ~90 two-second buckets) and must keep behaving as it does.
+export const MAX_DURATION_BLOCK_SIZE = 2000;
export type DuplicateVideoRef = {
slug: string; // `${channelSlug}/${id}`
@@ -35,10 +92,35 @@ export type DuplicateVideoRef = {
channel: string; // display name
platform: Platform;
id: string; // canonical id (may differ from the on-disk dir name)
+ // The on-disk directory under `<channel>/data/`, recorded ONLY when it differs
+ // from `id` — which it does for 14.5% of the corpus (every Rumble re-upload:
+ // 7,870 of the 7,870 `the-quartering-rumble` videos alone).
+ //
+ // It is here because `slug` is `${channelSlug}/${id}` and is therefore NOT a
+ // path. Anything that resolves a member back to its files has to be told the
+ // difference, and the detector is the only place that cheaply knows it (it
+ // keys its own scan by directory). Omitted when dir === id so the shipped
+ // report does not grow a redundant field on the other 85%; readers fall back
+ // to `id`, which is what every reader assumed before this field existed.
+ videoDir?: string;
title: string;
duration: number; // seconds
uploadDate: string; // YYYYMMDD
hasTranscript: boolean;
+ // Timing alignment against the cluster's canonical member, measured by
+ // measureAlignment() at detection time (confirmed clusters only — see the
+ // detector). Both are optional: reports written before these fields existed
+ // lack them, and absent must be read as "not measured", i.e. NOT aligned.
+ //
+ // `offsetSeconds` is the largest |offset| observed across the matched anchors,
+ // not a signed shift — the same quantity writeSharedDigest already records.
+ // It is informational; `aligned` is the gate. Anything that seeks INTO a
+ // sibling (the search-result duplicate badge, a shared digest's chapters) must
+ // key off `aligned`, because a mirror with a longer intro matches on text at
+ // shifted times and would otherwise land in the wrong place while looking
+ // perfectly plausible.
+ offsetSeconds?: number | null;
+ aligned?: boolean;
};
export type DuplicateCluster = {
@@ -49,15 +131,44 @@ export type DuplicateCluster = {
durationBucket: number; // representative rounded duration (seconds)
crossPlatform: boolean; // members span more than one platform
crossChannel: boolean; // members span more than one channelSlug
+ // True when the cluster's strongest evidence is title+duration only — nothing
+ // compared the actual content. It is offered for human review and shares no
+ // derived work until confirmed. Optional: reports written before this field
+ // existed lack it, and absent means "content-confirmed", which is what those
+ // reports only ever contained.
+ needsReview?: boolean;
videoRefs: DuplicateVideoRef[];
+ // The member that OWNS derived work for this cluster: the one an AI digest is
+ // generated for, and the one aligned mirrors copy it from. Chosen by
+ // pickCanonicalSlug() at detection time and human-overridable afterwards
+ // (see DuplicateOverrides). Optional: reports written before this field
+ // existed lack it, so readers fall back to pickCanonicalSlug().
+ canonicalSlug?: string;
};
+// How candidate pairs are NOMINATED. A blocking strategy never decides anything
+// on its own — it only proposes pairs that the transcript cascade then confirms
+// or rejects (see the controller's evalBlocked). The choice is therefore about
+// recall and cost, not correctness.
+//
+// "duration" — same/adjacent rounded-duration bucket. The only strategy that
+// catches a RE-TITLED mirror. Corpus-viable only since the
+// streaming rewrite; still much the more expensive of the two.
+// "title" — exact normalized-title groups, paired when the runtimes agree.
+// Effectively linear, and it is what actually finds cross-platform
+// re-uploads, which keep their name.
+// "both" — the union, deduped.
+export type DuplicateBlocking = "duration" | "title" | "both";
+
export type DuplicateRunConfig = {
thresholdSeconds: number | null; // null === all durations
durationToleranceSeconds: number;
nearThreshold: number; // Jaccard
containmentThreshold: number;
shingleSize: number;
+ // Optional: reports written before blocking was configurable lack it, and
+ // those were all duration-blocked.
+ blocking?: DuplicateBlocking;
};
export type DuplicateReport = {
@@ -75,6 +186,419 @@ export type DuplicateReport = {
export const DUPLICATES_FILENAME = "duplicates.json";
// ---------------------------------------------------------------------------
+// Human review: canonical choice + not-a-duplicate, kept OUT of the report
+// ---------------------------------------------------------------------------
+
+// duplicates.json is regenerated wholesale by every detection run, so a human
+// decision recorded in it would be destroyed on the next run. It therefore lives
+// in a sibling override file — the same separation, for the same reason, as
+// ai-digest.overrides.json vs ai-digest.json.
+export const DUPLICATE_OVERRIDES_FILENAME = "duplicates.overrides.json";
+export const DUPLICATE_OVERRIDES_VERSION = 1;
+
+export type DuplicateClusterOverride = {
+ // Operator's choice of canonical member (a `${channelSlug}/${id}` slug). Wins
+ // over the rule in pickCanonicalSlug. Ignored when the slug is not a member.
+ canonicalSlug?: string;
+ // "These are not the same video." Suppresses the cluster entirely: it stops
+ // being offered for review AND stops sharing derived work. The detector will
+ // keep finding it (content really is similar), which is exactly why the
+ // decision has to be recorded outside the report.
+ notDuplicate?: boolean;
+ // "I looked, and these really are the same video." The positive counterpart of
+ // notDuplicate, and the ONLY thing that lets a title-duration cluster share
+ // derived work. Content-confirmed clusters do not need it.
+ confirmed?: boolean;
+ decidedAt?: string;
+ note?: string;
+};
+
+export type DuplicateOverrides = {
+ version: number;
+ // Keyed by clusterId — a stable sha1 over the sorted member slugs, so the key
+ // survives re-detection as long as the membership does. A cluster that GAINS a
+ // member gets a new id and returns to review, which is the honest behavior:
+ // the canonical choice was made over a different set of videos.
+ clusters: Record<string, DuplicateClusterOverride>;
+};
+
+export function emptyDuplicateOverrides(): DuplicateOverrides {
+ return { version: DUPLICATE_OVERRIDES_VERSION, clusters: {} };
+}
+
+// Coerce a raw overrides file, dropping ill-typed entries rather than throwing —
+// one hand-edit typo must not break the duplicates page.
+export function sanitizeDuplicateOverrides(value: unknown): DuplicateOverrides {
+ const out = emptyDuplicateOverrides();
+ if (!value || typeof value !== "object") return out;
+ const raw = (value as { clusters?: unknown }).clusters;
+ if (!raw || typeof raw !== "object" || Array.isArray(raw)) return out;
+ for (const [clusterId, entry] of Object.entries(raw as Record<string, unknown>)) {
+ if (!entry || typeof entry !== "object") continue;
+ const e = entry as Record<string, unknown>;
+ const override: DuplicateClusterOverride = {};
+ if (typeof e.canonicalSlug === "string" && e.canonicalSlug.trim()) {
+ override.canonicalSlug = e.canonicalSlug.trim();
+ }
+ if (e.notDuplicate === true) override.notDuplicate = true;
+ if (e.confirmed === true) override.confirmed = true;
+ if (typeof e.decidedAt === "string") override.decidedAt = e.decidedAt;
+ if (typeof e.note === "string" && e.note.trim()) override.note = e.note.trim();
+ if (Object.keys(override).length === 0) continue;
+ out.clusters[clusterId] = override;
+ }
+ return out;
+}
+
+// ---------------------------------------------------------------------------
+// Canonical member selection
+// ---------------------------------------------------------------------------
+
+// Platform preference for the canonical member, most-preferred first. YouTube
+// leads because its videos carry the richest metadata and the most reliable
+// caption tracks, so a digest generated there is the best one to share.
+const PLATFORM_PREFERENCE: ReadonlyArray<string> = [
+ "youtube",
+ "rumble",
+ "odysee",
+ "kick",
+ "twitch",
+];
+
+function platformRank(platform: string): number {
+ const i = PLATFORM_PREFERENCE.indexOf(platform);
+ return i < 0 ? PLATFORM_PREFERENCE.length : i;
+}
+
+// Pick the cluster member that should own derived work, by rule. Ordered by
+// what actually makes a shared digest good:
+// 1. has a transcript — you cannot digest a video without one
+// 2. longest duration — the most complete artifact; a mirror that cuts
+// the intro would place every shared chapter wrong
+// 3. preferred platform — richest metadata
+// 4. earliest upload — the original, where the same content appears twice
+// 5. slug — a total order, so the choice is deterministic
+// Deterministic and pure: no Date.now(), no Math.random(), so re-detection over
+// unchanged inputs picks the same member.
+export function pickCanonicalSlug(cluster: DuplicateCluster): string {
+ const refs = cluster.videoRefs;
+ if (refs.length === 0) return "";
+ const best = refs.reduce((a, b) => (canonicalBetter(b, a) ? b : a));
+ return best.slug;
+}
+
+function canonicalBetter(
+ candidate: DuplicateVideoRef,
+ incumbent: DuplicateVideoRef,
+): boolean {
+ if (candidate.hasTranscript !== incumbent.hasTranscript) {
+ return candidate.hasTranscript;
+ }
+ if (candidate.duration !== incumbent.duration) {
+ return candidate.duration > incumbent.duration;
+ }
+ const pc = platformRank(candidate.platform);
+ const pi = platformRank(incumbent.platform);
+ if (pc !== pi) return pc < pi;
+ if (candidate.uploadDate !== incumbent.uploadDate) {
+ // Empty upload dates sort last so a dated member wins over an undated one.
+ if (!candidate.uploadDate) return false;
+ if (!incumbent.uploadDate) return true;
+ return candidate.uploadDate < incumbent.uploadDate;
+ }
+ return candidate.slug.localeCompare(incumbent.slug) < 0;
+}
+
+// The effective canonical member: the human choice when it names a real member,
+// otherwise the recorded one, otherwise the rule. Returns null for a cluster a
+// human marked not-a-duplicate — such a cluster shares nothing.
+export function resolveCanonicalSlug(
+ cluster: DuplicateCluster,
+ overrides?: DuplicateOverrides | null,
+): string | null {
+ const override = overrides?.clusters[cluster.clusterId];
+ if (override?.notDuplicate) return null;
+ const members = new Set(cluster.videoRefs.map((r) => r.slug));
+ if (override?.canonicalSlug && members.has(override.canonicalSlug)) {
+ return override.canonicalSlug;
+ }
+ if (cluster.canonicalSlug && members.has(cluster.canonicalSlug)) {
+ return cluster.canonicalSlug;
+ }
+ return pickCanonicalSlug(cluster) || null;
+}
+
+// Whether a human has recorded a decision for this cluster. Drives the
+// /actionable "awaiting review" list — a cluster with no decision is work.
+export function isClusterReviewed(
+ cluster: DuplicateCluster,
+ overrides?: DuplicateOverrides | null,
+): boolean {
+ const o = overrides?.clusters[cluster.clusterId];
+ if (!o) return false;
+ return (
+ o.notDuplicate === true || o.confirmed === true || Boolean(o.canonicalSlug)
+ );
+}
+
+// ---------------------------------------------------------------------------
+// The timestamp-alignment gate — the correctness crux of digest sharing
+// ---------------------------------------------------------------------------
+
+// Content similarity does NOT imply timing alignment, and this is the failure
+// mode that looks like success: a mirror with a 40-second-longer intro has
+// matching text at shifted times, so a shared digest places EVERY chapter wrong
+// while looking perfectly plausible. Nothing in the 5-gram Jaccard cascade
+// notices, because shingles are a set — order and position are discarded.
+//
+// So before sharing we measure it directly: sample several anchor phrases spread
+// through the canonical transcript, find where each occurs in the mirror, and
+// require every offset to be near zero.
+
+// Max |offset| (seconds) at any anchor for two transcripts to count as aligned.
+// Tight on purpose: a chapter placed 5s early still lands in the right sentence,
+// 15s does not.
+export const DEFAULT_ALIGNMENT_TOLERANCE_SECONDS = 5;
+// How many anchors to sample. Several, spread out, because a mirror can share a
+// start time and then diverge at an ad break in the middle.
+export const DEFAULT_ALIGNMENT_ANCHORS = 5;
+// An anchor must be this many words long to be a reliable locator; shorter
+// phrases recur.
+const ANCHOR_WORDS = 8;
+// How far forward to look for a USABLE anchor when the phrase at the sampled
+// position is ambiguous (occurs more than once on either side). Real transcripts
+// repeat themselves — intros, catchphrases, ad reads — so without this the gate
+// refuses perfectly aligned mirrors and the sharing optimisation never fires.
+const ANCHOR_SEARCH_WORDS = 400;
+
+export type AlignmentResult = {
+ aligned: boolean;
+ // Largest |offset| observed across the matched anchors, in seconds.
+ maxOffsetSeconds: number;
+ // How many anchors were located in the other transcript. Too few and the
+ // measurement is not trustworthy, so `aligned` is false regardless of offset.
+ matchedAnchors: number;
+ totalAnchors: number;
+ reason?:
+ | "too-few-anchors"
+ | "offset-exceeded"
+ | "empty-transcript"
+ | "contained";
+};
+
+function normalizeWords(cues: Cue[]): { word: string; start: number }[] {
+ const out: { word: string; start: number }[] = [];
+ for (const cue of cues) {
+ const words = cue.text
+ .toLowerCase()
+ .replace(/[^\p{L}\p{N}\s]/gu, " ")
+ .split(/\s+/)
+ .filter(Boolean);
+ for (const word of words) out.push({ word, start: cue.start });
+ }
+ return out;
+}
+
+// Word n-gram -> { first start time, occurrence count }. The count is what makes
+// an ambiguous phrase skippable instead of silently mismatched.
+function buildAnchorIndex(
+ words: { word: string; start: number }[],
+): Map<string, { start: number; count: number }> {
+ const index = new Map<string, { start: number; count: number }>();
+ for (let i = 0; i + ANCHOR_WORDS <= words.length; i++) {
+ const key = words
+ .slice(i, i + ANCHOR_WORDS)
+ .map((w) => w.word)
+ .join(" ");
+ const existing = index.get(key);
+ if (existing) existing.count++;
+ else index.set(key, { start: words[i].start, count: 1 });
+ }
+ return index;
+}
+
+// Measure timing alignment between two cue lists. Pure, so the gate is directly
+// unit-testable (and it IS tested: a shifted-intro mirror must fail).
+export function measureAlignment(
+ canonical: Cue[],
+ other: Cue[],
+ opts: {
+ toleranceSeconds?: number;
+ anchors?: number;
+ } = {},
+): AlignmentResult {
+ const tolerance =
+ opts.toleranceSeconds ?? DEFAULT_ALIGNMENT_TOLERANCE_SECONDS;
+ const anchorCount = Math.max(1, opts.anchors ?? DEFAULT_ALIGNMENT_ANCHORS);
+
+ const a = normalizeWords(canonical);
+ const b = normalizeWords(other);
+ if (a.length < ANCHOR_WORDS || b.length < ANCHOR_WORDS) {
+ return {
+ aligned: false,
+ maxOffsetSeconds: Infinity,
+ matchedAnchors: 0,
+ totalAnchors: 0,
+ reason: "empty-transcript",
+ };
+ }
+
+ // Index BOTH sides' word n-grams, counting occurrences — not just recording the
+ // first. A phrase that occurs twice is not a locator: matching it to whichever
+ // copy came first would report a bogus offset, which fails in both directions
+ // (a false refusal wastes a generation; a false match shares a misplaced
+ // digest). So an ambiguous phrase is skipped rather than guessed at.
+ const indexB = buildAnchorIndex(b);
+ const indexA = buildAnchorIndex(a);
+
+ // Anchors spread evenly through the canonical transcript, skipping the very
+ // start and end (intros/outros are exactly where mirrors differ).
+ const usable = a.length - ANCHOR_WORDS;
+ const positions: number[] = [];
+ for (let k = 1; k <= anchorCount; k++) {
+ positions.push(Math.floor((usable * k) / (anchorCount + 1)));
+ }
+
+ let matched = 0;
+ let maxOffset = 0;
+ const usedTargets = new Set<number>();
+ for (const pos of positions) {
+ // Walk forward from the sampled position until a phrase is unique on BOTH
+ // sides. Bounded, so a pathologically repetitive stretch just yields no
+ // anchor here rather than scanning the whole transcript.
+ const limit = Math.min(usable, pos + ANCHOR_SEARCH_WORDS);
+ for (let i = pos; i <= limit; i++) {
+ const key = a
+ .slice(i, i + ANCHOR_WORDS)
+ .map((w) => w.word)
+ .join(" ");
+ if ((indexA.get(key)?.count ?? 0) !== 1) continue;
+ const hit = indexB.get(key);
+ if (!hit || hit.count !== 1) continue;
+ // Don't let two sampled positions collapse onto the same anchor — that
+ // would report "2 anchors matched" from one measurement.
+ if (usedTargets.has(hit.start)) break;
+ usedTargets.add(hit.start);
+ matched++;
+ maxOffset = Math.max(maxOffset, Math.abs(hit.start - a[i].start));
+ break;
+ }
+ }
+
+ // Require a majority of anchors to be located: a mirror we can only match in
+ // one place is not a mirror we can trust timings from.
+ const enough = matched >= Math.ceil(positions.length / 2);
+ if (!enough) {
+ return {
+ aligned: false,
+ maxOffsetSeconds: matched === 0 ? Infinity : maxOffset,
+ matchedAnchors: matched,
+ totalAnchors: positions.length,
+ reason: "too-few-anchors",
+ };
+ }
+ if (maxOffset > tolerance) {
+ return {
+ aligned: false,
+ maxOffsetSeconds: maxOffset,
+ matchedAnchors: matched,
+ totalAnchors: positions.length,
+ reason: "offset-exceeded",
+ };
+ }
+ return {
+ aligned: true,
+ maxOffsetSeconds: maxOffset,
+ matchedAnchors: matched,
+ totalAnchors: positions.length,
+ };
+}
+
+// Whether a cluster may share derived work AT ALL, before any timing is measured.
+// A `contained` cluster matched by containment — one member is a CLIP of a longer
+// video, not a mirror of it. A clip is a different artifact: the longer video's
+// chapters describe material the clip does not contain, so sharing wholesale
+// would be wrong even at a perfect zero offset.
+//
+// A `needsReview` cluster (title+duration only) shares nothing either, until a
+// human records `confirmed: true`. The alignment gate would still protect the
+// TIMING, but nothing here has compared the CONTENT — two episodes of a daily
+// show can share a title and a runtime and be entirely different material, and
+// a shared digest would then describe the wrong video convincingly.
+export function clusterMaySharePartial(
+ cluster: DuplicateCluster,
+ overrides?: DuplicateOverrides | null,
+): boolean {
+ if (cluster.contained) return false;
+ if (!cluster.needsReview) return true;
+ return overrides?.clusters[cluster.clusterId]?.confirmed === true;
+}
+
+// Whether a cluster may be SHIPPED to a built site at all.
+//
+// Distinct from clusterMaySharePartial, which asks whether derived work may flow
+// between members. The two disagree in both directions, on purpose:
+//
+// a `contained` cluster SHIPS (a clip of a longer video is a real, useful
+// relationship for a viewer to see) but SHARES NOTHING (the longer video's
+// chapters describe material the clip does not contain);
+//
+// an unconfirmed `needsReview` cluster does NEITHER — nothing compared its
+// members' content, so asserting the relationship to a viewer would be
+// claiming something no machine and no human has actually checked. It stays an
+// internal review queue until someone records `confirmed`.
+//
+// Fails closed: no overrides means no confirmations means no suspects ship.
+export function clusterIsPublishable(
+ cluster: DuplicateCluster,
+ overrides?: DuplicateOverrides | null,
+): boolean {
+ if (!cluster.needsReview) return true;
+ return overrides?.clusters[cluster.clusterId]?.confirmed === true;
+}
+
+// The blocking key for title-based candidate generation.
+//
+// This is what makes corpus-wide detection tractable. The alternative the first
+// implementation used — pairing every short with every longer video — is
+// quadratic and measured at 498 MILLION pairs on this corpus. Grouping by an
+// exact normalized title instead is one pass and a hash lookup, and it is the
+// signal that actually finds cross-platform mirrors, which are re-uploads of the
+// same file under the same name.
+//
+// Deliberately conservative: lowercase, strip punctuation and diacritics, drop a
+// few platform suffixes, collapse whitespace. It does NOT stem, fuzzy-match or
+// drop stopwords — a looser key merges a series ("Episode 12" vs "Episode 13")
+// and every such merge is a false cluster a human then has to reject.
+const TITLE_NOISE_RE =
+ /\s*(?:#shorts?|\(official(?: video| audio)?\)|\[official\]|\|\s*full episode|\(full episode\)|\(reupload\)|\[reupload\]|\(mirror\)|\[mirror\])\s*/gi;
+
+export function normalizeTitleKey(title: string): string {
+ return title
+ .normalize("NFKD")
+ // Strip combining marks so "Pokémon" and "Pokemon" block together.
+ .replace(/[\u0300-\u036f]/g, "")
+ .toLowerCase()
+ .replace(TITLE_NOISE_RE, " ")
+ .replace(/[^\p{L}\p{N}]+/gu, " ")
+ .trim();
+}
+
+// Are two same-titled videos close enough in runtime to be candidates?
+export function durationsCompatible(
+ a: number,
+ b: number,
+ ratio = DEFAULT_TITLE_DURATION_RATIO,
+): boolean {
+ if (!(a > 0) || !(b > 0)) return false;
+ const tolerance = Math.max(
+ TITLE_DURATION_MIN_TOLERANCE_SECONDS,
+ Math.max(a, b) * ratio,
+ );
+ return Math.abs(a - b) <= tolerance;
+}
+
+// ---------------------------------------------------------------------------
// Pure helpers (no I/O) — exported for direct testing.
// ---------------------------------------------------------------------------
@@ -175,6 +699,7 @@ export class UnionFind {
// Tier ranking so a cluster reports its strongest evidence.
const MATCH_RANK: Record<DuplicateMatchKind, number> = {
+ "title-duration": 0,
"transcript-near": 1,
"transcript-exact": 2,
};
diff --git a/common/lib/parseStdoutJson.ts b/common/lib/parseStdoutJson.ts
@@ -0,0 +1,58 @@
+// Parse the JSON a CLI printed on stdout, scanning lines LAST-TO-FIRST.
+//
+// Lifted from checkAvailability.ts (where it read yt-dlp --dump-json) because
+// the digest registry needs exactly the same tolerance for `claude -p
+// --output-format json`: both tools may print warnings, deprecation notices, or
+// progress to stdout BEFORE the payload, so the JSON is the last parseable line,
+// not the first. Reading forward finds the warning; reading backward finds the
+// answer.
+export function parseStdoutJson(stdout: string): unknown | null {
+ const lines = stdout
+ .split("\n")
+ .map((l) => l.trim())
+ .filter(Boolean);
+ for (let i = lines.length - 1; i >= 0; i--) {
+ try {
+ return JSON.parse(lines[i]);
+ } catch {
+ continue;
+ }
+ }
+ // Fall back to the whole buffer: `--output-format json` pretty-prints across
+ // several lines, so no single line parses on its own.
+ try {
+ return JSON.parse(stdout);
+ } catch {
+ return null;
+ }
+}
+
+// Pull a JSON object out of model prose. A CLI lane has no constrained decoding,
+// so the body may arrive fenced (```json … ```) or with a sentence in front of
+// it. Tries the whole string, then the fenced block, then the outermost braces.
+export function extractJsonObject(text: string): unknown | null {
+ const trimmed = text.trim();
+ try {
+ return JSON.parse(trimmed);
+ } catch {
+ // fall through
+ }
+ const fence = trimmed.match(/```(?:json)?\s*([\s\S]*?)```/);
+ if (fence) {
+ try {
+ return JSON.parse(fence[1].trim());
+ } catch {
+ // fall through
+ }
+ }
+ const first = trimmed.indexOf("{");
+ const last = trimmed.lastIndexOf("}");
+ if (first >= 0 && last > first) {
+ try {
+ return JSON.parse(trimmed.slice(first, last + 1));
+ } catch {
+ return null;
+ }
+ }
+ return null;
+}
diff --git a/common/lib/paths.ts b/common/lib/paths.ts
@@ -61,6 +61,13 @@ export type Paths = {
// Served per-channel social-post page tree (/posts/<slug>/{manifest,page-NNNN}.json).
// The parallel corpus to transcripts/subs — see common/lib/posts.ts.
exportPostsDir: string;
+ // Served per-channel AI-digest page tree
+ // (/digests/<slug>/{manifest,page-NNNN}.json). The DERIVED corpus: chapters
+ // and topic tags generated from a transcript, composed with human overrides
+ // at index time. Sparse by design — only videos that have been digested
+ // appear in a manifest's slugToPage, which is what lets the viewer hide its
+ // Digest control without a per-video fetch. See common/lib/digests.ts.
+ exportDigestsDir: string;
exportStatsDir: string;
// Staging area (NOT served) where the shared index + per-site aggregates are
// built before composition. Shared per-channel transcript/subs pages are
@@ -71,6 +78,7 @@ export type Paths = {
exportSharedTranscriptsDir: string;
exportSharedSubsDir: string;
exportSharedPostsDir: string;
+ exportSharedDigestsDir: string;
exportSitesIndexDir: string;
// Per-site extracted static output (`out/`) from an isolated (Docker) build,
// keyed exportBuildsDir/<siteId>. Sibling of .export-index. The wrangler deploy
@@ -108,6 +116,16 @@ export type Paths = {
parakeetBin: string;
parakeetCliBin: string;
parakeetModel: string;
+ // Base URL of the local ollama server, the local-GPU digest lane
+ // (common/lib/digestApps.ts posts to `${ollamaUrl}/api/chat`). The FIRST
+ // URL-valued entry in Paths, so it is normalized here the way
+ // remoteTranscribe.ts normalizes a worker base: trailing slashes stripped, so
+ // callers can always concatenate a leading-slash path.
+ ollamaUrl: string;
+ // The `claude` CLI, driving the opt-in metered digest lane (off by default —
+ // settings.digest.remoteEnabled). Not bundled; install it separately and point
+ // CLAUDE_BIN at it if it isn't on PATH.
+ claudeBin: string;
};
let cached: Paths | null = null;
@@ -152,12 +170,14 @@ export function getPaths(): Paths {
exportTranscriptsDir: path.join(exportPublicDir, "transcripts"),
exportSubsDir: path.join(exportPublicDir, "subs"),
exportPostsDir: path.join(exportPublicDir, "posts"),
+ exportDigestsDir: path.join(exportPublicDir, "digests"),
exportStatsDir: path.join(exportPublicDir, "stats"),
exportIndexDir,
exportSharedDir,
exportSharedTranscriptsDir: path.join(exportSharedDir, "transcripts"),
exportSharedSubsDir: path.join(exportSharedDir, "subs"),
exportSharedPostsDir: path.join(exportSharedDir, "posts"),
+ exportSharedDigestsDir: path.join(exportSharedDir, "digests"),
exportSitesIndexDir: path.join(exportIndexDir, "sites"),
exportBuildsDir:
process.env.EXPORT_BUILDS_DIR ??
@@ -190,6 +210,11 @@ export function getPaths(): Paths {
path.join(monorepoRoot, "scripts", "parakeet-stitch.mjs"),
parakeetCliBin: process.env.PARAKEET_CLI ?? "parakeet-cli",
parakeetModel: process.env.PARAKEET_MODEL ?? "",
+ ollamaUrl: (process.env.OLLAMA_URL ?? "http://127.0.0.1:11434").replace(
+ /\/+$/,
+ "",
+ ),
+ claudeBin: process.env.CLAUDE_BIN ?? "claude",
};
return cached;
}
diff --git a/common/lib/queueKeys.ts b/common/lib/queueKeys.ts
@@ -7,6 +7,17 @@ import {
export { TRANSCRIPTION_QUEUE };
+// The two digest lanes get SEPARATE queue keys, and that separation is the whole
+// point. registry.ts submits every non-empty queueKey with concurrency 1, so:
+// - one shared digest key would serialize the lanes, wasting the network lane
+// while the GPU works (and vice versa) across a multi-week sweep;
+// - putting the local lane on TRANSCRIPTION_QUEUE would let ollama and the
+// transcription engine thrash the same 8 GB of VRAM.
+// (A queueKey of "" means "run immediately, untracked" — never right for either
+// of these, both of which must be serialized against themselves.)
+export const DIGEST_LOCAL_QUEUE = "digest:local";
+export const DIGEST_REMOTE_QUEUE = "digest:remote";
+
// Per-channel queue for channel-local bookkeeping jobs (clean/clear/verify).
export function channelQueueKey(slug: string): string {
return `channel:${slug}`;
diff --git a/common/lib/settings.ts b/common/lib/settings.ts
@@ -29,6 +29,20 @@ import {
isCookieMode,
type CookieMode,
} from "./cookiePolicy";
+// From the CLIENT-SAFE digest module, deliberately — digestApps.ts imports execa,
+// and settings.ts must stay reachable from anywhere.
+import {
+ CLAUDE_DIGEST_APP_ID,
+ DEFAULT_DIGEST_APP_ID,
+ DEFAULT_DIGEST_TIMESTAMP_MODE,
+ DIGEST_SECTION_KINDS,
+ DIGEST_TIMESTAMP_MODES,
+ isDigestSectionKind,
+ isDigestTimestampMode,
+ type DigestAppConfig,
+ type DigestSectionKind,
+ type DigestTimestampMode,
+} from "./digest";
export type { Worker } from "./workers";
export type { AutoQueueSettings } from "../jobs/autoQueuePolicy";
@@ -175,6 +189,68 @@ export type SiteSettings = {
// pipeline itself is a follow-up; this block persists the chosen mode plus the
// container/concurrency knobs the deploy page and the future orchestrator read.
buildPipeline: BuildPipelineSettings;
+ // AI digest generation (chapters + topic tags over the existing transcripts).
+ // Local-first: the metered lane is off by default. See DigestSettings.
+ digest: DigestSettings;
+};
+
+// Configuration for the derived-corpus digest layer. Local-first by decision:
+// `remoteEnabled` gates the metered lane and defaults to false, so nothing here
+// can spend money until it is explicitly turned on.
+export type DigestSettings = {
+ // Master switch for the metered (remote-api) lane. OFF by default — an opt-in
+ // overflow for the long tail or a channel where local quality is poor, never
+ // the default path.
+ remoteEnabled: boolean;
+ // Videos longer than this are "long tail": 8.2% of the corpus by count, 46% of
+ // all transcript tokens. The batch's duration-aware ordering and the optional
+ // remote overflow both key off it.
+ longTailSeconds: number;
+ // The engine each lane uses (ids from common/lib/digestApps.ts).
+ localAppId: string;
+ remoteAppId: string;
+ // Per-app config, keyed by app id — the same id-keyed sub-record shape as
+ // transcriptionApps.
+ apps: Record<string, DigestAppConfig>;
+ // Global pause. Read at DISPATCH time by the batch (the downloadsPaused
+ // pattern), so a pause survives a restart with no boot hook — unlike
+ // transcriptionsPaused, which needs editor/instrumentation.ts to re-apply it.
+ digestsPaused: boolean;
+ // Yield the GPU to the transcription lane: while whisper is working, the
+ // digest batch's limit() returns 0 and the pool idle-waits. ON by default,
+ // because `digest:local` is deliberately on a different queue from
+ // TRANSCRIPTION_QUEUE and so would otherwise run ollama and whisper on the
+ // same 8 GB card — measured at 90 s per audio-hour against the 27 s an idle
+ // box projected. See controller/digestYield.ts.
+ yieldToTranscription: boolean;
+ // The corpus-wide sweep is armed. Read at boot by the editor's instrumentation
+ // hook, the same way the auto-transcribe/auto-download runners are, so a sweep
+ // survives a server restart. It is persisted INTENT, not a cursor: the batch
+ // re-derives eligibility from disk on every pull, so a resumed sweep does zero
+ // rework and needs nothing else remembered.
+ sweepEnabled: boolean;
+ // Channels the armed sweep covers. Empty = the whole corpus.
+ //
+ // Persisted alongside `sweepEnabled` because the boot hook re-launches from
+ // settings alone: without it, a sweep deliberately scoped to two channels
+ // would come back after a restart as an unscoped corpus-wide run — silently
+ // widening GPU-weeks of work that an operator had bounded on purpose.
+ sweepChannels: string[];
+ // Hard ceiling on cumulative metered spend per job, USD. 0 = no cap. Only ever
+ // consulted for a metered app.
+ spendCapUsd: number;
+ // Which sections a sweep generates. Chapters alone is the default: tags double
+ // the call count for a smaller payoff.
+ sections: DigestSectionKind[];
+ // How each chunk's transcript markers are numbered — see DigestTimestampMode.
+ // Was a scored variable in the bake-off rather than a pre-applied fix; the
+ // measurement is in and "chunk-local" is now the shipped default.
+ timestampMode: DigestTimestampMode;
+ // A free-text label for a non-default prompt shape, folded into the recorded
+ // provenance by digestPromptVariant(). Setting it invalidates every digest
+ // generated under a different label, which is exactly what makes a bake-off
+ // round re-run its sample instead of skipping it as fresh. Empty = default.
+ promptVariant: string;
};
// "basic" — `pnpm run build` in export/, serialized on the build queue (shared
@@ -427,6 +503,133 @@ export function sanitizeSavedVideoBackup(
};
}
+// 4 hours. Measured: videos over this are 8.2% of the corpus by count but hold
+// 46% of all transcript tokens, so they are where a sweep's wall-clock actually
+// goes and where chunk-seam bugs live.
+export const DIGEST_LONG_TAIL_DEFAULT_SECONDS = 4 * 3600;
+export const DIGEST_LONG_TAIL_MAX_SECONDS = 24 * 3600;
+
+export function defaultDigest(): DigestSettings {
+ return {
+ // OFF. The metered lane is built but never the default — see PLAN.md.
+ remoteEnabled: false,
+ longTailSeconds: DIGEST_LONG_TAIL_DEFAULT_SECONDS,
+ localAppId: DEFAULT_DIGEST_APP_ID,
+ remoteAppId: CLAUDE_DIGEST_APP_ID,
+ // Empty on purpose: every per-app knob falls through to its own default
+ // constant (resolveNumCtx -> DEFAULT_DIGEST_NUM_CTX, now 8192, and
+ // maxCuesForContext sizes the chunk to it). Seeding a copy of those values
+ // here would give the same number two homes and let them drift.
+ apps: {},
+ digestsPaused: false,
+ // ON. Contention with whisper is the single largest cost in the backfill
+ // (90 s/audio-hour measured vs 27 projected on an idle box), so the safe
+ // default is to step aside; turning it off is the deliberate choice.
+ yieldToTranscription: true,
+ // OFF. A corpus-wide sweep is GPU-weeks of work and is never armed by
+ // default — an operator starts it.
+ sweepEnabled: false,
+ sweepChannels: [],
+ spendCapUsd: 0,
+ sections: ["chapters"],
+ timestampMode: DEFAULT_DIGEST_TIMESTAMP_MODE,
+ promptVariant: "",
+ };
+}
+
+// Coerce a raw settings.digest.apps value into a clean keyed map of
+// DigestAppConfig. Mirrors sanitizeTranscriptionApps — INCLUDING its
+// Array.isArray guard, without which a JSON array would pass the typeof check and
+// produce numeric-keyed garbage.
+export function sanitizeDigestApps(
+ value: unknown,
+): Record<string, DigestAppConfig> {
+ if (!value || typeof value !== "object" || Array.isArray(value)) return {};
+ const out: Record<string, DigestAppConfig> = {};
+ for (const [id, raw] of Object.entries(value as Record<string, unknown>)) {
+ if (!raw || typeof raw !== "object") continue;
+ const r = raw as Record<string, unknown>;
+ const cfg: DigestAppConfig = {};
+ if (typeof r.bin === "string" && r.bin.trim()) cfg.bin = r.bin.trim();
+ if (typeof r.baseUrl === "string" && r.baseUrl.trim()) {
+ cfg.baseUrl = r.baseUrl.trim();
+ }
+ if (typeof r.model === "string" && r.model.trim()) cfg.model = r.model.trim();
+ if (typeof r.numCtx === "number" && r.numCtx > 0) {
+ cfg.numCtx = Math.floor(r.numCtx);
+ }
+ if (typeof r.temperature === "number" && r.temperature >= 0) {
+ cfg.temperature = r.temperature;
+ }
+ if (typeof r.timeoutMs === "number" && r.timeoutMs > 0) {
+ cfg.timeoutMs = Math.floor(r.timeoutMs);
+ }
+ // Only carried when explicitly set — see DigestAppConfig.think.
+ if (typeof r.think === "boolean") cfg.think = r.think;
+ out[id] = cfg;
+ }
+ return out;
+}
+
+export function sanitizeDigest(value: unknown): DigestSettings {
+ const d = defaultDigest();
+ if (!value || typeof value !== "object") return d;
+ const r = value as Record<string, unknown>;
+ const sections = Array.isArray(r.sections)
+ ? (r.sections.filter(isDigestSectionKind) as DigestSectionKind[])
+ : [];
+ return {
+ remoteEnabled: r.remoteEnabled === true,
+ longTailSeconds: clampPositiveInt(
+ r.longTailSeconds,
+ d.longTailSeconds,
+ DIGEST_LONG_TAIL_MAX_SECONDS,
+ ),
+ // Unknown app ids are not rejected here: getDigestApp() is total and falls
+ // back to the local default, so a stale id degrades rather than breaking.
+ localAppId:
+ typeof r.localAppId === "string" && r.localAppId.trim()
+ ? r.localAppId.trim()
+ : d.localAppId,
+ remoteAppId:
+ typeof r.remoteAppId === "string" && r.remoteAppId.trim()
+ ? r.remoteAppId.trim()
+ : d.remoteAppId,
+ apps: sanitizeDigestApps(r.apps),
+ digestsPaused: r.digestsPaused === true,
+ // Defaults to ON when absent — `=== false` rather than `!== true`, so a
+ // settings file written before this field existed keeps the GPU-safe
+ // behaviour instead of silently opting into contention.
+ yieldToTranscription: r.yieldToTranscription !== false,
+ sweepEnabled: r.sweepEnabled === true,
+ sweepChannels: Array.isArray(r.sweepChannels)
+ ? r.sweepChannels.filter(
+ (v): v is string => typeof v === "string" && v.trim().length > 0,
+ )
+ : [],
+ spendCapUsd:
+ typeof r.spendCapUsd === "number" && r.spendCapUsd > 0
+ ? Math.round(r.spendCapUsd * 100) / 100
+ : 0,
+ // An empty/garbage list would silently generate nothing, so fall back to the
+ // default rather than honoring it.
+ sections: sections.length > 0 ? sections : d.sections,
+ timestampMode: isDigestTimestampMode(r.timestampMode)
+ ? r.timestampMode
+ : d.timestampMode,
+ // Trimmed and length-capped: it goes into provenance on every record, and a
+ // runaway value would bloat 119k sidecars.
+ promptVariant:
+ typeof r.promptVariant === "string"
+ ? r.promptVariant.trim().slice(0, 40)
+ : d.promptVariant,
+ };
+}
+
+// Every known section kind, for the settings UI's checkbox list.
+export const DIGEST_SECTION_OPTIONS = DIGEST_SECTION_KINDS;
+export const DIGEST_TIMESTAMP_MODE_OPTIONS = DIGEST_TIMESTAMP_MODES;
+
export const BUILD_MAX_PARALLEL_DEFAULT = 2;
export const BUILD_MAX_PARALLEL_MAX = 16;
export const DEFAULT_BUILD_IMAGE = "yt-dlp-transcript-browser-build";
@@ -498,6 +701,9 @@ function defaults(): SiteSettings {
homepageUrl: "",
savedVideoBackup: defaultSavedVideoBackup(),
buildPipeline: defaultBuildPipeline(),
+ // Must be listed here or the allowlist loop in getSettings() drops the key
+ // entirely and the whole section is never read from disk.
+ digest: defaultDigest(),
};
}
@@ -712,6 +918,7 @@ export function getSettings(): SiteSettings {
merged.homepageUrl = normalizeHomepageUrl(merged.homepageUrl);
merged.savedVideoBackup = sanitizeSavedVideoBackup(merged.savedVideoBackup);
merged.buildPipeline = sanitizeBuildPipeline(merged.buildPipeline);
+ merged.digest = sanitizeDigest(merged.digest);
// Workers. When the file predates the worker model (no `workers` key),
// synthesize a default list from the (now-settled) active app + per-app
// configs so existing installs behave identically. Otherwise sanitize the
@@ -904,6 +1111,7 @@ export async function writeSettings(next: SiteSettings): Promise<void> {
homepageUrl: normalizeHomepageUrl(next.homepageUrl),
savedVideoBackup: sanitizeSavedVideoBackup(next.savedVideoBackup),
buildPipeline: sanitizeBuildPipeline(next.buildPipeline),
+ digest: sanitizeDigest(next.digest),
};
const tmp = `${file}.tmp-${process.pid}`;
await fs.promises.writeFile(tmp, JSON.stringify(merged, null, 2) + "\n");
diff --git a/common/lib/site.ts b/common/lib/site.ts
@@ -150,6 +150,10 @@ export function sitePostsDir(paths: Paths, siteId: string): string {
return path.join(siteIndexDir(paths, siteId), "posts");
}
+export function siteDigestsDir(paths: Paths, siteId: string): string {
+ return path.join(siteIndexDir(paths, siteId), "digests");
+}
+
export function siteStatsDir(paths: Paths, siteId: string): string {
return path.join(siteIndexDir(paths, siteId), "stats");
}
diff --git a/common/lib/transcriptWindow.test.ts b/common/lib/transcriptWindow.test.ts
@@ -1,6 +1,11 @@
import { test } from "node:test";
import assert from "node:assert/strict";
-import { windowCues, cuesToSnippets, mergeSnippets } from "./transcriptWindow";
+import {
+ chunkCuesForContext,
+ cuesToSnippets,
+ mergeSnippets,
+ windowCues,
+} from "./transcriptWindow";
import type { Cue } from "./vtt";
function cue(start: number, text = `t${start}`): Cue {
@@ -82,3 +87,67 @@ test("mergeSnippets never drops existing hits and stops adding at cap", () => {
[1, 2, 3],
);
});
+
+// ---------------------------------------------------------------------------
+// chunkCuesForContext — the digest chunker. windowCues cannot do this job (it is
+// center-based and measured in seconds), so these cases pin the contract the
+// digest generator depends on: full coverage, in order, with a shared overlap.
+// ---------------------------------------------------------------------------
+
+test("chunkCuesForContext returns one chunk when the transcript fits", () => {
+ const cues = [cue(0), cue(5), cue(10)];
+ const chunks = chunkCuesForContext(cues, { maxCues: 10, overlapCues: 2 });
+ assert.equal(chunks.length, 1);
+ assert.deepEqual(chunks[0], cues);
+});
+
+test("chunkCuesForContext covers every cue at least once", () => {
+ const cues = Array.from({ length: 25 }, (_, i) => cue(i * 10));
+ const chunks = chunkCuesForContext(cues, { maxCues: 10, overlapCues: 3 });
+ const seen = new Set(chunks.flat().map((c) => c.start));
+ assert.equal(seen.size, cues.length, "every cue appears in some chunk");
+});
+
+test("chunkCuesForContext overlaps consecutive chunks by overlapCues", () => {
+ const cues = Array.from({ length: 25 }, (_, i) => cue(i * 10));
+ const chunks = chunkCuesForContext(cues, { maxCues: 10, overlapCues: 3 });
+ for (let i = 1; i < chunks.length; i++) {
+ const prevTail = chunks[i - 1].slice(-3).map((c) => c.start);
+ const head = chunks[i].slice(0, 3).map((c) => c.start);
+ assert.deepEqual(head, prevTail, `chunk ${i} shares its head with the previous tail`);
+ }
+});
+
+test("chunkCuesForContext emits chunks in chronological order", () => {
+ const cues = Array.from({ length: 40 }, (_, i) => cue(i * 10));
+ const chunks = chunkCuesForContext(cues, { maxCues: 12, overlapCues: 4 });
+ for (let i = 1; i < chunks.length; i++) {
+ assert.ok(
+ chunks[i][0].start > chunks[i - 1][0].start,
+ "each chunk starts later than the previous one",
+ );
+ }
+});
+
+test("chunkCuesForContext never emits a trailing chunk that is pure overlap", () => {
+ // 13 cues at maxCues 10 / overlap 3 steps by 7: [0..9], [7..12]. A third chunk
+ // starting at 14 would be empty, and one starting at 12 would be re-work.
+ const cues = Array.from({ length: 13 }, (_, i) => cue(i * 10));
+ const chunks = chunkCuesForContext(cues, { maxCues: 10, overlapCues: 3 });
+ assert.equal(chunks.length, 2);
+ assert.equal(chunks[1][chunks[1].length - 1].start, 120);
+});
+
+test("chunkCuesForContext clamps an overlap that would stall the loop", () => {
+ // overlapCues >= maxCues would make step 0 and loop forever; it is clamped so
+ // progress is guaranteed.
+ const cues = Array.from({ length: 30 }, (_, i) => cue(i * 10));
+ const chunks = chunkCuesForContext(cues, { maxCues: 5, overlapCues: 99 });
+ assert.ok(chunks.length > 1 && chunks.length < 40);
+ const seen = new Set(chunks.flat().map((c) => c.start));
+ assert.equal(seen.size, cues.length);
+});
+
+test("chunkCuesForContext handles an empty transcript", () => {
+ assert.deepEqual(chunkCuesForContext([], { maxCues: 10 }), []);
+});
diff --git a/common/lib/transcriptWindow.ts b/common/lib/transcriptWindow.ts
@@ -40,6 +40,47 @@ export function windowCues(
.sort((a, b) => a.start - b.start);
}
+// Split a whole transcript into SEQUENTIAL, overlapping, context-sized slices —
+// the chunker the digest generator drives its engine with, one call per slice.
+//
+// This is a different job from windowCues() above and cannot be expressed with
+// it: that one is CENTER-based (a window around a search hit, measured in
+// seconds) and has no notion of covering a transcript exactly once. Here the
+// contract is coverage: every cue appears in at least one slice, slices are in
+// order, and consecutive slices share `overlapCues` cues.
+//
+// The overlap exists because a topic that straddles a seam is otherwise invisible
+// to both calls — each sees half a discussion and titles it wrong. With the
+// overlap, at least one call sees it whole; the parser's seam de-dup then drops
+// the resulting near-identical chapters.
+//
+// `maxCues` is measured in cues rather than tokens deliberately: cues are what we
+// slice, and the token estimate that sizes them belongs with the prompt (see
+// DIGEST_MAX_CUES_PER_CHUNK), not here.
+export function chunkCuesForContext(
+ cues: Cue[],
+ opts: { maxCues?: number; overlapCues?: number } = {},
+): Cue[][] {
+ const maxCues = Math.max(1, Math.floor(opts.maxCues ?? 1200));
+ // Overlap must leave forward progress, or the loop never advances.
+ const overlapCues = Math.max(
+ 0,
+ Math.min(Math.floor(opts.overlapCues ?? 0), maxCues - 1),
+ );
+ if (cues.length === 0) return [];
+ if (cues.length <= maxCues) return [cues.slice()];
+
+ const step = maxCues - overlapCues;
+ const out: Cue[][] = [];
+ for (let start = 0; start < cues.length; start += step) {
+ out.push(cues.slice(start, start + maxCues));
+ // Stop once this slice reached the end, so we never emit a trailing slice
+ // that is pure overlap (it would be re-processed for nothing).
+ if (start + maxCues >= cues.length) break;
+ }
+ return out;
+}
+
// Render cues as snippet objects (clock + seconds + collapsed, capped text).
// Empty cues are dropped.
export function cuesToSnippets(cues: Cue[]): WindowSnippet[] {
diff --git a/common/package.json b/common/package.json
@@ -3,6 +3,9 @@
"version": "0.1.0",
"private": true,
"type": "module",
+ "scripts": {
+ "test": "tsx --test \"{lib,controller,jobs,social,ytdlp,components}/*.test.ts\""
+ },
"dependencies": {
"@sindresorhus/slugify": "^3.0.0",
"@tanstack/react-query": "^5.99.1",
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -1,159 +1,7 @@
# Changelog
## [Unreleased]
-- **Replace YouTube's auto-captions with transcripts of our own.** Most of the corpus rides on YouTube ASR captions, which are noticeably worse than what the transcription workers produce — no punctuation, rolling duplicate cues, `[Music]` filler — and they were *sticky*: `isVideoTranscribed()` counts any English VTT, so a video with only auto-captions was permanently invisible to every transcribe bucket and every transcribe job. There is now an opt-in, strictly-lowest-priority lane that finds those videos, downloads their audio, and transcribes them properly; whisper's `transcript.json` then wins the index pick automatically. Provenance is decided by a 4 KB sniff of the VTT itself (YouTube ASR marks ~96–100% of cues with `align:start position:N%` plus inline word timings; manual tracks mark 0%), with the 490 KB `metadata.info.json` parse kept only as a tie-breaker — so the per-regen cost is one small read per English-VTT-having video. Three snapshot buckets carry it: `autoSubsOnly` (needs audio) → `downloadedAutoSubsOnly` (needs whisper) → `supersededAutoSubs` (done, old VTT kept as a backup). Nothing is automatic by default: the auto-queue gains a per-runner **Replace YouTube auto-captions** switch that appends the bucket to the *tail* of the default union (real work always drains first), and a leaf can target the bucket by name for per-channel opt-in — the default unions are byte-identical to before, so existing setups are untouched. Manually, the Transcribe stage gains a two-step "YouTube auto-captions only" section and a single-video *Replace auto-captions* action, and each subtitle track is now labelled *YouTube auto-captions* / *manual captions*. A caption track whose provenance can't be proven machine-generated is **never** a candidate, and the original VTT is never deleted automatically — the Cleanup stage's purge button is the only thing that removes it (English ASR tracks only, honouring `do-not-clean`), so an AI-vs-YouTube comparison stays possible. See `common/lib/subtitleProvenance.ts`, `common/controller/purgeSupersededAutoSubs.ts`, `common/jobs/autoQueuePolicy.ts`, `editor/e2e/auto-subs-replace.spec.ts`.
-- **Social posts are archived as a parallel corpus to video transcripts.** The archive can now ingest X/Twitter and Bluesky accounts from the same commentators and search them *together* with video transcripts — one corpus, one set of searches. A channel gains `sourceKind: "social"` (a separate axis from `handling`, so every existing `handling === "transcribe" ? … : …` branch stays binary and can never misroute), plus `postFetcher` and `socialHandle`. Posts are modelled on the live-chat layer, not the video layer: their own month-sharded JSONL on disk (`channels/<slug>/posts/YYYY-MM.jsonl` + a `posts-archive` of seen ids, so a re-run is a no-op), their own LMDB sub-DB keyed `[createdAt, channelSlug, id]` (ISO-8601 sorts chronologically, fixing the intra-day ordering the `YYYYMMDD` video key has), and their own `/posts/<slug>/{manifest,page-NNNN}.json` page tree. Ingest is a pluggable `SocialFetcher` registry (`common/social/fetchers.ts`) modelled on the transcription-app registry: **`bluesky-atproto`** (pure `fetch` against the public AT Protocol — no auth, no binary, verified end-to-end against a live account), **`x-gallery-dl`** (the primary X path: a light headless subprocess, cookies via the existing `cookiePolicy.ts`), and **`x-playwright`** (the fallback, immune to the GraphQL query-id rotations that periodically break gallery-dl). One `fetch-posts` job kind carries it, drainable and bookmarkable, routed to `platform:x.com` / `platform:bsky.app` by the existing queue keys. A social channel gets a minimal two-stage rail (Fetch → Index) instead of the six video stages, and the channel form hides every video-only control (audio format, download format, keep-source-video, extraction mode, saved-video dir). Auto-sync works unchanged — but the scheduler's *dispatch* now routes social channels to a post fetch rather than a yt-dlp video sync. See `common/lib/posts.ts`, `common/social/*`, `common/controller/fetchPosts.ts`, `editor/e2e/social-channel.spec.ts`.
-- **Connect an X account once, instead of re-supplying cookies every few days.** X session cookies expire within days, which is gallery-dl's worst flaw as an archiving path. Settings gains an X-session broker: a headed "Connect X account" flow opens a browser on the editor host so the operator logs in by hand (2FA and captcha included — the login is deliberately never automated), storing a persistent browser profile. The fetcher then re-exports a fresh `cookies.txt` from that profile on demand and prefers it over `--cookies-from-browser`, so the session stops being the thing that breaks. Only X's own cookies are exported. See `common/social/xSessionBroker.ts`, `editor/app/settings/components/XSessionSection.tsx`.
-- **The dashboard is now a live mission-control cockpit.** The home page used to be a static SSR card stack (four stat tiles, a "needs attention" link, a bare channels table) — all the *live* operational density lived only in the opt-in `/widget` monitor. The dashboard is rebuilt as an information-first operations surface that updates in place, following the widget's proven pattern: `page.tsx` stays a server component that SSRs initial payloads and hands them to a client shell (`DashboardCockpit`) that polls the same `/api` endpoints (`/api/jobs/active`, `/api/workers`, `/api/widget/actionable`, `/api/widget/sync`). The hero is a full-width **Pipeline band** — a single instrument readout of running/queued jobs, worker-pool busy/total + pause state, sync heartbeat + scheduler on/off, the live job rows with progress bars, and the global controls (Pause transcriptions, Pause downloads, Sync all, + Add). Below it a **Needs-work** panel (top channels with ↓/✎ counts and inline Download/Transcribe actions, linking to `/actionable`) sits beside a **Quick-add / recent-changes** column, over an **enriched channels table** (relative "last sync" that ticks live, plus per-row inline Sync / Download-missing / Top-of-queue actions). Everything stays on the existing semantic "base" tokens (no new palette) and respects `prefers-reduced-motion`. The two ~1s poll hooks are extracted from the widget into `editor/app/widget/lib/usePolledPayload.ts` and the relative-time helpers into `.../relativeTime.ts`, shared by both surfaces. See `editor/app/page.tsx`, `editor/app/components/dashboard/*`, and `editor/e2e/dashboard.spec.ts`.
-- **URL-first channel onboarding.** The New-channel form now derives what it can from a pasted URL with **no network call** — platform (`detectPlatform`), handling (YouTube → subs, else transcribe), and a slug candidate (`@Veritasium` → `veritasium`, `/c/Some Name` → `some-name`) — and surfaces the **derived download queue** as a read-only hint (e.g. `Queue: platform:vimeo.com (new)` for a host the app doesn't recognize). An opt-in **Fetch details** button runs a lightweight single-entry yt-dlp probe (`probeChannelMeta` → `probeChannelUrlAction`) that fills the name and, crucially, lets a site **unknown to the app but known to yt-dlp** be created with its own `platform:<domain>` serial queue in the same action (no change to the closed `Platform` union). On create it now **always stores the playlist** ("Fetch playlist now", default on) so pending downloads populate immediately, with an optional **Add to top of auto-queue** that prepends a channel leaf at the head of the download policy tree and starts the runner (`prioritizeChannelDownloadAction`, also a one-click "Top of queue" on the dashboard channels table). Also fixes the missing **Kick** option in the platform select. See `editor/app/channels/{components/ChannelForm.tsx,actions.ts}`, `editor/app/auto-queue/actions.ts`, `common/ytdlp/runYtdlp.ts`, and `editor/e2e/new-channel-onboarding.spec.ts`.
-- **Pause is now first-class and survives restarts — for transcriptions and downloads.** Transcription pause was runtime-only (lost on restart) and buried on the Workers page; downloads had no pause at all. Both are now persisted in `settings.json` (`transcriptionsPaused` / `downloadsPaused`) and toggled from the dashboard Pipeline band (and, for parity, the monitor widget's controls). Pausing transcriptions persists the flag and re-pauses the worker pool at boot (`editor/instrumentation.ts`); pausing downloads idles the auto-download runner on its next loop iteration (re-read each tick, like the `enabled` flag) **and** makes manual download-bearing pipeline actions (`sync`, `download-from-playlist`, `download-missing`, `download-missing-subs`, `retry-bucket`) return a friendly "Downloads are paused" notice — store-playlist/enumeration stay allowed since they write no media. See `common/lib/settings.ts`, `editor/app/workers/actions.ts`, `editor/app/jobs/actions.ts`, `common/controller/autoRunner.ts`, and `editor/app/channels/[slug]/pipelineActions.ts`.
-- **Channel groups gain a "Show channels as individual chips" checkbox on the Site form.** Toggling it on marks the group `inline` in site.json, so the viewer renders each member channel as its own loose chip in the filter row instead of a collapsible group box. New sites' seeded default group ("All channels") ships with it **on**; groups added manually on the Site form or created inline from a channel form default **off**. See `editor/app/sites/components/SiteForm.tsx`, `common/lib/channelGroups.ts`, and `editor/e2e/sites-crud.spec.ts`.
-- **Cookies-from-browser is now configurable: three cookie modes, per-channel overrides, and a "Needs cookies" bucket.** The single global cookies value used to be hard-wired to one behavior (passed only on the download auth-retry attempt). Settings now carry a **Cookie mode** next to the value — **When required** (default; the old behavior, now also cookie-retrying a failed *metadata prefetch*, so an age-gated video succeeds on its first real attempt and reports `ok-with-cookies` with a `metadata-prefetch-auth-retry` attempt in `download-outcome.json`), **Always** (every yt-dlp invocation carries `--cookies-from-browser`: enumeration, prefetch, availability probes, downloads — for channels whose listing itself needs auth; a clean success still reports plain `ok`), and **Defer** (normal runs never use cookies: known `needs_auth` videos are excluded from sync/download-missing batches — previously they were re-attempted and re-failed on *every* run — and wait for a manual cookie run). Both the value and the mode can be overridden per channel in the channel form's Advanced section (blank/Inherit = use global). Every channel's Download stage gains a **Needs cookies (N)** card listing undownloaded videos whose recorded availability is cookie-recoverable (`needs_auth`, and also `members_only`/`private`, which batches always excluded but a subscribed/owning account's cookies can fetch) with a **Download with cookies** button that re-runs them with cookies forced on every invocation (bookmarkable, like the other bucket jobs; it warns if no cookie value is configured anywhere). Caveats: in defer mode a *manual single-video* download intentionally gets no cookies and no auth retry — the bucket button is the explicit cookie path; availability probes only attach cookies in Always mode, so needs_auth keeps being observed (defer depends on that signal). Back-compat: an existing `settings.json` without `cookieMode` behaves exactly as before (`when-required`). See `common/lib/cookiePolicy.ts`, `common/ytdlp/{downloadOneManaged,runYtdlp}.ts`, `common/controller/{channelSnapshot,checkAvailability}.ts`, and `editor/e2e/cookies-mode.spec.ts`.
-- **Channel forms now edit site membership directly — pick sites, groups, and create groups inline.** A channel's site membership used to be editable only from the Site form (`/sites/<id>`), and creating a channel silently appended it to the active site (a hidden `activeSite` field) with no group choice and no visibility. Both the **New channel** and channel **Configure** forms now carry a **Sites** section: every configured site listed with a membership checkbox and a compact group dropdown — `(default)`, any existing group, or **"+ New group…"**, which reveals a name input and creates the group on that site (id slugified from the name, reused if it already exists) as part of the save. On create, the active site (`?site=`) is pre-checked, reproducing the old behavior but visibly and overridably; on edit, current memberships pre-check with their groups, and unchecking removes the membership. The checked sites serialize into one hidden `siteMembershipsJson` field (the SiteForm hidden-JSON precedent); the server plans all writes up front — validating group ids, preserving other channels' entries and this channel's `order`, erroring clearly on a since-deleted group, skipping since-deleted sites, and rewriting only sites that actually changed. See the new `editor/app/channels/{lib/siteMemberships.ts,components/SiteMembershipsSection.tsx}`, `editor/app/channels/{actions.ts,components/{ChannelForm,ChannelFormClient}.tsx,new/page.tsx,[slug]/page.tsx}`, and `editor/e2e/channel-site-membership.spec.ts`.
-- **The monitor widget can now start a sync and shows sync freshness + scheduler health — each behind its own flag.** The widget was start-a-sync-less and said nothing about how fresh your channels were. Five new opt-in URL flags, all toggleable in the builder and the in-widget gear (all default off, so existing links are unchanged): **(1) a channel-aware Sync button** (`sync=1`) — pinned to one channel (`channel=X`) it runs that channel's streaming sync; otherwise it sweeps all channels (`syncAllChannelsAction`), briefly showing `Queued X · skipped Y`. **(2) A last-sync readout** (`lastsync=1`) — `Last full sync: …` from a new persisted `lastSyncAllAt` marker written at the end of each Sync-all sweep, plus a `Last channel sync: …` line when an individual channel synced more recently. **(3) A scheduler-status strip** (`sched=1`) — `Auto-sync on/off · next … · last run …`, derived from the same `buildScheduleView` the `/scheduler` page uses. **(4) An absolute-time toggle** (`abstime=1`) — readouts show locale timestamps instead of relative "5m ago". **(5) A Sync-all confirm** (`syncask=1`) — a `window.confirm` before a full sweep. The two readouts poll a new lightweight `/api/widget/sync` route (a few scalars, 15s floor) and, with the Sync button/controls, keep the widget from collapsing to "Idle" — sync health is exactly what you check when nothing's running. See `editor/app/widget/lib/config.ts`, `editor/app/widget/components/{MonitorWidget,WidgetControls,WidgetConfigForm}.tsx`, the new `editor/app/api/widget/sync/route.ts`, `common/jobs/syncSchedulerState.ts` (`lastSyncAllAt`), `editor/app/channels/actions.ts`, and `editor/e2e/widget.spec.ts`.
-- **The audio-integrity check now tightens its probe interval when a source starts serving corruption, then relaxes as it stabilises.** Previously the integrity probe ran on a fixed cadence (default 60s) for the whole download, so up to ~60s of bytes were downloaded — and discarded — between a corruption event and the checkpoint that caught it. The interval is now adaptive (AIMD, like TCP congestion control, inverted): each **malformed** checkpoint **halves** the live interval (60→30→15→10s, floored at the existing `AUDIO_CHECK_INTERVAL_MIN_SECONDS` of 10s), so a misbehaving source gets probed more aggressively and wastes fewer bytes per rollback; a run of clean checkpoints then **steps it back up** additively (+15s after every 2 clean probes) toward the configured interval. The reduced cadence persists across yt-dlp relaunches for the rest of the download run. Fully backward compatible — a clean download never leaves the configured interval. Tunable via constants in `common/lib/channelConfig.ts` (`AUDIO_CHECK_INTERVAL_BACKOFF_FACTOR_DEFAULT`, `AUDIO_CHECK_INTERVAL_RECOVER_STEP_SECONDS`, `AUDIO_CHECK_INTERVAL_RECOVER_AFTER_CLEAN`) plus test-only env overrides. See `common/ytdlp/audioCheckCadence.ts` (pure AIMD math + `audioCheckCadence.test.ts`) and `common/ytdlp/audioCheckedDownload.ts` (`resolveKnobs`, the watcher loop, and the advance/malformed checkpoint branches).
-- **Docker build mode is now real: build every site in parallel, then deploy them serially.** The `Docker` build mode (Settings → Build pipeline) was previously a stub that fell back to the basic build. It now runs a proper pipeline, driven by a new **Build all sites** control on the Deploy page (one job, one log, one Cancel). The shared, corpus-scale work — the search index, the per-site staging, and the downloadable archive zips — runs **once on the host**; then each site's `compose + next build` runs in its **own container in parallel** (capped by the **Max parallel builds** setting), each writing an isolated per-site `out/` under `export/.export-builds/<siteId>/`; then the built sites **deploy serially** on the host (R2 upload + `wrangler pages deploy`), tolerant of a single site failing. Containers are read-only over the shared corpus/index/archive cache and run as your host user so outputs aren't root-owned. The image (`Dockerfile.build`, tag from **Build image**) is built/reused via Docker layer caching; when no container engine is available the action falls back to a serial host build+deploy. New env knobs: `DOCKER_BIN` (e.g. `podman`), `DOCKER_BUILD_MEMORY`/`DOCKER_BUILD_CPUS` (per-container caps). See `editor/app/deploy/buildDeployCore.ts` (`runDockerBuildAllPhase`/`runDockerDeployAllPhase`), `editor/app/build/buildAction.ts` (`buildAndDeployAllSitesAction`/`buildAllSitesAction`), `editor/app/deploy/components/BuildAllSitesButton.tsx`, `Dockerfile.build`, `docker/build-site.sh`, `common/bin/build-archives.ts`, and **[DEPLOY_DOCKER.md](../DEPLOY_DOCKER.md)**.
+- test bullet for cut-release spec
-## [0.7.3] - 2026-07-07
-- **Deploys no longer re-upload unchanged oversize archives to R2.** The deploy step used to stream every over-cap archive to R2 on every deploy, even ones byte-identical to what was already there. Because the export build now reuses an unchanged channel's cached zip verbatim, the upload step first does a cheap `HeadObject` and skips any archive whose R2 object already has the same size — so a redeploy after changing one channel only re-uploads that channel's oversize bundle. See `editor/app/deploy/buildDeployCore.ts`.
-- **New "Search aliases" page for authoring known-term suggestions.** A concept is often spelled many ways — transcription in particular mangles them (AI expands `loli` to `lolly`/`loly`) — so searching one spelling silently misses the rest. This page authors a curated dictionary where one entry groups all the trigger spellings of a concept with a single robust replacement (e.g. `\blol(i|ly)`). A **Global** section applies to every site and ships a small seeded default set you can edit or remove; a **per-site** section (shown when a site is selected in the sidebar) adds to or overrides the global list by id — reuse a global id and mark it disabled to hide that alias for just that site. Nothing is forced: the entries only surface a *suggestion* in the viewer's search box, which the searcher can apply or ignore. Stored as `search-aliases.json` (global under the data dir, per-site under `sites/<id>/`) and merged into each site's bundle at build time. See `editor/app/aliases/{page.tsx,actions.ts,EditorAliasesClient.tsx}`, `common/lib/{searchAliases,aliasesStore}.ts`, `editor/app/lib/nav.ts`, and `editor/e2e/aliases.spec.ts`.
-- **Fixed: the editor (and every public site) always loaded in light mode until you toggled the theme.** Dark styling is driven purely by a `.dark` class on `<html>`; the pre-paint `<ThemeScript>` set it correctly before first paint, but `<html>` is server-rendered with a static class that omits `dark`, so that class was lost across React's hydration boundary and nothing put it back — the page fell to the light palette on every load until a manual toggle re-applied it directly (which is why toggling then "stuck"). `ThemeProvider` now re-asserts the persisted, resolved family + mode to `<html>` on mount via `useLayoutEffect` (before paint) — idempotent with the script, so there's no flash and no toggle needed. Explicit dark, `system` on a dark OS, and non-Base families all now survive a refresh. See `common/components/ThemeProvider.tsx` and `editor/e2e/theme.spec.ts`.
-- **Search results scroll smoothly again on large result sets.** The results list windows one card per matching video (only the on-screen cards are mounted), but each visible card and every one of its hit rows was re-rendering on *every* scroll frame — and each hit row re-ran its `<mark>` highlighting, so a single video with hundreds of hits meant hundreds of redundant highlight passes per frame while scrolling. The result cards and individual hit rows are now memoized so an unchanged card/row is skipped during scroll, and opening the modal on a hit only re-renders the two rows whose highlight state actually changes. No visible/behavioral change — same DOM, same results, just far less work per frame. See `common/components/TranscriptSearch.tsx` (`ResultCard`/`HitRow` memoization, `openWithMode` stabilized via `useCallback`).
-- **The monitor widget gains a needs-work channel list, more interaction buttons, and an in-place settings gear.** Three additions, all driveable from the widget builder. **(1) A "Needs work" list** (URL flag `act=1`) — a compact, per-channel worklist of videos to download (`↓ N`) or transcribe (`✎ N`), reusing the same `loadActionableSummary` that powers the `/actionable` page via a new `/api/widget/actionable` route; it polls on a 15s floor (the backlog changes on job completions, not seconds) and caps at 6 channels with a `+N more` line. **(2) More interactions** behind the existing `controls=1` switch: each needs-work row gains the same per-channel **Download missing** / **Transcribe pending** buttons as the actionable page (reusing `InlineActionButton`), and the controls row adds **Retry all failed** alongside Pause/Resume + Drain. **(3) An in-place settings gear** (on by default; URL flag `gear=0` to hide, or a **Show settings gear** builder checkbox) — clicking it opens the builder's own form *inside the widget window*, so a pinned widget can be reconfigured live without opening the builder page; edits apply immediately and mirror into the address bar via `history.replaceState`, so a reload preserves them and the link stays copyable. The builder form is extracted into a shared `WidgetConfigForm` used by both the builder and the overlay, and the widget's poller now fetches immediately on (re)subscribe instead of after one interval, so newly-enabled sections render at once. Existing links render unchanged (the two new flags default to their old behavior; the gear is the one new default-visible affordance and is read-only — it mutates no server state). See `editor/app/widget/lib/config.ts`, the new `editor/app/widget/components/WidgetConfigForm.tsx` and `editor/app/api/widget/actionable/route.ts`, `editor/app/widget/components/{MonitorWidget,WidgetControls}.tsx`, `editor/app/widget/builder/components/WidgetBuilder.tsx`, and `editor/e2e/widget.spec.ts`.
-- **Every site build now bundles downloadable per-channel transcript & live-chat archive zips.** The archive builders (per-channel `<slug>.zip` / `<slug>.live_chat.zip`) previously only ran as standalone actions that wrote to a non-served directory; now `compose-site` generates them for the site's own channels straight into the served `public/archives/` and writes a `manifest.json` (sizes + counts) that the site's new **Downloads** page reads. **`zip` is now the default archive format** everywhere (was `tar.gz`), and the Build page's format help text tracks the selected format. Generation is **on by default with three opt-out levels**: a global **Generate archive zips on build** toggle in Settings, a per-site **Generate archive zips** toggle (plus an optional **Archive size cap (MB)**) on the site's page, and a per-build **Skip archive zips** checkbox on the Build and Build & Deploy controls (`BUILD_ARCHIVES=0`). See `common/bin/compose-site.ts` (`composeArchives`), `common/controller/archive{Transcripts,LiveChat}.ts` (new `outDir` option), `common/lib/archiveOptions.ts` (default + manifest types), `common/lib/{site,settings}.ts` (opt-out flags), and `editor/app/{deploy/buildDeployCore.ts,build/buildAction.ts,deploy/components/Build{Export,Deploy}Button.tsx,sites/components/SiteForm.tsx,settings/components/SettingsForm.tsx}`.
-- **Oversize archives now overflow to Cloudflare R2 instead of being dropped.** A single file over 25 MB breaks a Cloudflare Pages deploy, so a channel zip over the cap (default 25 MB; `0` = no cap) used to be removed from what's served and flagged `oversize`. Now, when **Archive overflow storage** is configured in Settings (an R2 **bucket** + its **public URL**), `compose-site` stages each oversize archive to `export/.r2-staging/<siteId>/` and records its future public URL in the manifest; the deploy step then uploads it to `<bucket>/<siteId>/archives/<file>` **before** the Pages deploy, so the Downloads page links straight to R2. Uploads go over R2's **S3 API** using the AWS SDK's multipart uploader (`@aws-sdk/lib-storage`) — `wrangler r2 object put` caps a single upload at 300 MiB and real live-chat archives are larger, whereas multipart streams any size. This needs R2 **S3 credentials** in the environment (`R2_ACCESS_KEY_ID`, `R2_SECRET_ACCESS_KEY`, `CLOUDFLARE_ACCOUNT_ID`); if they're missing when there's something to upload, the deploy fails before the Pages step (so the site never links to a missing file). With no bucket configured the old drop-and-flag behavior is unchanged (uploads run only in the editor's Deploy / Build & deploy actions, not a raw `pnpm deploy`). The **combined "whole site" archives were removed** — they duplicated the per-channel content and were always the first to blow the cap. Each R2 upload also now sets `Cache-Control: public, max-age=3600` so a Cloudflare custom domain caches downloads at the edge — the main defense against download-abuse cost (R2 egress is free; only origin reads are billable, and cached hits skip the origin). New **[DEPLOY_CLOUDFLARE.md](../DEPLOY_CLOUDFLARE.md)** documents the full R2 setup plus the Cloudflare custom-domain / caching / rate-limiting / bot config for instance operators. See `common/lib/settings.ts` (`archiveStorage`), `common/bin/compose-site.ts` (staging + manifest `url`), `common/lib/archiveOptions.ts` (`ArchiveManifestEntry.url`), and `editor/app/deploy/buildDeployCore.ts` (`runArchiveUploadIntoLog`, `ARCHIVE_CACHE_CONTROL`) / `deploy/deployAction.ts` / `build/buildAction.ts`.
-- **Archive downloads can now be served securely without owning a domain (`r2-proxy/` Worker).** Serving oversize R2 archives with cost/abuse protection previously implied a Cloudflare **custom domain** (for CDN caching + rate-limiting rules). New self-contained Cloudflare Worker at `r2-proxy/` serves the bucket on a free `*.workers.dev` subdomain instead — the same domain-free model as Pages' `*.pages.dev`, addressing the privacy/expense of registering a domain. It's a pure passthrough (request path `<siteId>/archives/<file>.zip` → bucket key), so **one Worker serves every site** — deploy it once, not per site. It adds edge caching (Cache API, honoring the object's `Cache-Control`), native per-IP rate limiting (free binding), path allow-listing (`*/archives/*.zip` only), and `Range`/resumable-download support. Point the editor's **Archive overflow public URL** at the `workers.dev` URL and nothing else changes (uploads/manifest are identical). `DEPLOY_CLOUDFLARE.md` now documents both paths (Worker vs custom domain), with the Worker as the recommended no-domain option. See `r2-proxy/{src/index.ts,wrangler.toml,package.json,README.md}` and `editor/app/settings/components/SettingsForm.tsx` (public-URL hint).
-- **The Duplicates page is now a per-site toggle and hides itself when empty.** Each site's editor page gains a **Show the Duplicates page** checkbox (on by default). `compose-site` writes the site-filtered `duplicates.json` only when the toggle is on *and* there's at least one in-scope cluster, and the export Header keys its Duplicates nav link off a new `hasDuplicates()` — so the link and page disappear both when a site opts out and when it simply has no detected duplicates. See `common/lib/site.ts` (`duplicates` flag), `common/bin/compose-site.ts` (gated write), `export/app/lib/duplicates.ts` (new), `export/app/components/Header.tsx`, and `editor/app/sites/{components/SiteForm.tsx,actions.ts}`.
-- **You can now change a channel's slug (its id) — deliberately, from the Danger zone.** A channel's slug *is* its on-disk directory name (`transcripts/channels/<slug>/`), so it used to be fixed at creation ("Slug is fixed once a channel is created"). A new **Rename** form in the channel's Danger zone lifts that: enter a new slug and **type the current slug to confirm** (same friction as delete), and the rename is blocked while the channel has running/queued jobs (the in-memory registry keys by slug). Because the slug is a directory name, the rename does a **full migration** of every slug-keyed store so nothing silently breaks: it moves the channel dir (config, data, playlist, snapshot, shards, failed lists) **and** the saved-video store dir — rewriting each `saved-video.json` pointer's absolute `dir` so persisted source videos still resolve — then retargets every site.json membership, the sync scheduler's per-channel backoff state, and any job bookmarks. The two filesystem moves run first and roll back on failure; the metadata updates that follow are atomic and best-effort (surfaced as warnings). Renaming **changes the channel's public URL** (the old one 404s), which the form warns about. The slug grammar is also now validated on create. See `common/controller/renameChannel.ts`, `common/controller/channels.ts` (`isValidChannelSlug`), `common/lib/savedVideo-server.ts` (`rewriteSavedVideoDir`), `common/jobs/bookmarks.ts` (`renameChannelInBookmarks`), `editor/app/channels/{actions.ts,components/RenameChannelForm.tsx,[slug]/page.tsx}`, and `editor/e2e/channel-rename.spec.ts`.
-- **New Queue diagnostics page (`/jobs/queue`): see & force-release stuck jobs.** The job system has two sources of truth that can drift — the registry owns each job's `status`, the scheduler owns the running SLOT per queue. A cancel that never finalizes (a child that ignored SIGTERM, a crashed finalizer) leaves a job "cancelled" in the registry while the scheduler still marks its slot running, silently blocking every job behind it on that queue — and the Active Jobs page hides it (it filters to running/queued). The new **Queue** page reconciles the two: it builds from the **scheduler** as the source of truth for slots, cross-checks each against its registry record, and flags a running head as **stuck** when the record is terminal-but-holding-slot, evicted, or (softer) a live job idle past 10 minutes. It **auto-heals** the hard cases on every view/poll (frees terminal/evicted slots), shows a health strip (active queues, running, queued, **stuck**, workers), per-queue cards with the held-for duration / PID (`kill -9` hint) / last log line, and a **Force-release** button per slot (SIGKILLs the child and frees the slot unconditionally) plus a **Reap all stuck** action. Force-release is also available on any running job in Active Jobs, and Active Jobs links to Queue with a stuck-count badge. See `common/jobs/registry.ts` (`forceRelease`), `editor/app/jobs/queue/*`, `editor/app/jobs/{actions.ts,components/ForceReleaseJobButton.tsx}`, and `editor/e2e/queue.spec.ts`.
-- **Jobs page: real log retention + pagination (replaces the dead "Clear archived logs" button).** The old button only deleted logs absent from the in-memory registry — which, since the registry keeps the 100 newest finished jobs and sidecars preserve their real status, was almost never anything, so it did nothing. It's replaced by a **Clear logs** dropdown that prunes finished-job logs by age (older than 7 / 30 / 90 days) or all at once; running/queued jobs are never deleted. The `.jobs` directory also **self-trims on job finish** (throttled; keep newest 500, drop >30 days) so it can't grow unbounded. Job ids are now **ULIDs** (lexicographically time-sortable, timestamp decodable from the id), letting the list **paginate** — `listAllJobs` returns one page (default 50, grown by a **Load more** link) and only `stat`s/reads the sidecar for the shown page instead of every file on every load. `jobIdTime()` decodes both ULID and the legacy `<t36>-<rand>` ids, so existing on-disk logs still sort/read correctly. See `common/jobs/{ulid,listJobs,registry,streamCommand}.ts`, `editor/app/jobs/{page.tsx,actions.ts,components/ClearLogsMenu.tsx,[id]/page.tsx}`, and `editor/e2e/jobs.spec.ts`.
-- **"Move to top" button on the auto-queue policy editor.** Each reorderable rule/group in the auto-queue policy tree gains a **⤒** button beside the existing ↑/↓ swap controls that jumps the node straight to the front of its sibling list in one click (disabled on the first row, like ↑). Reordering stays local until **Save policy**, matching the swap buttons. See `editor/app/auto-queue/components/PolicyTreeEditor.tsx` and `editor/e2e/auto-queue.spec.ts`.
-- **Kick VOD playback + VOD-expiry indicators.** Kick becomes a first-class platform (`Platform` union, `detectPlatform`, `platformFromMetadata` `/^kick/i`, `extractVideoId` kick branch, `defaultWebpageUrl`). Kick VODs have no iframe embed, so playback streams the HLS manifest yt-dlp resolves at download time: `summarize()` persists `manifest_url` → `hlsUrl` on the transcript summary/detail, and a new client-only `common/components/KickPlayer.tsx` plays it in a native `<video>` via the bundled **hls.js** (not react-player's file player, which loads hls.js from a CDN and would break the offline export). It exposes the same `seekTo`/`onReady`/`onProgress` handle as the YouTube player, so Kick gets full scrubbing + cue highlighting; on a fatal manifest error it falls back to an expiry notice + source link. Separately, a shared `common/lib/vodExpiry.ts` (retention: Kick 30d, Twitch 14d, tunable) drives a new `VodExpiredBadge` on search result cards for likely-deleted Kick/Twitch VODs, with a `title=` tooltip explaining each platform's retention. Cache versions bumped so stale data re-derives (`transcriptStore` `DB_VERSION` 4, `normalizeTranscript` `CUES_FILE_VERSION` 2). See `common/lib/{platform,transcripts,transcripts-server,vodExpiry,format}.ts`, `common/components/{KickPlayer,PlayerProvider,badges,TranscriptSearch}.tsx`, `common/ytdlp/runYtdlp.ts`, and `export/e2e/kick-vod.spec.ts`.
-- **Clicking a site on the Sites list now opens its edit page instead of bouncing back to the list.** The sidebar site selector seeds the active site into the URL (`?site=`) on mount so the scoped server pages (Dashboard, Channels, Charts, Deploy) can read it — but it was also firing on the `/sites` CRUD pages, where a mount-time `router.replace("/sites?site=<id>")` raced and clobbered the in-flight navigation to `/sites/<id>` from a list link, dumping you back on the list. The seed is now skipped on `/sites*` routes (which never consume `?site=`), so site links navigate straight to the editor; scoped-page seeding is unchanged. See `editor/app/components/SiteScopeSelect.tsx`.
-- **"Build static export" can now skip the data rebuild and compose from existing staging.** The `export/` build normally regenerates the pool-wide data first (its npm `prebuild` hook runs `build:data` = `build:index && build:stats && build:templates`), and the index step (heavy transcript → page-tree + LMDB processing) dominates build time. A new **Skip data rebuild (index, stats, charts)** checkbox on the Build static export control lets you rebuild a site *without* that work — it runs only `compose:site && next build` against the current `.export-index/` staging, which is exactly what you want when re-composing after a code/theme/template change or building a different site from already-staged data. It routes to a new `build:nodata` export script (same body as `build`, but a distinct name so npm's `prebuild` hook doesn't fire); the checkbox reuses the streamed-log/queue/cancel machinery of the existing managed build. Assumes a prior full build produced the staging (do a full build first if the data is stale). Only the build-only control is affected — the one-click **Build & deploy** button always does a full build. See `export/package.json` (`build:nodata`), `editor/app/deploy/buildDeployCore.ts` (`runBuildPhase` `skipData`), `editor/app/build/buildAction.ts` (`buildExportAction`), and `editor/app/deploy/components/BuildExportButton.tsx`.
-- **The neutral Base theme is now all-sans.** The default Base family previously inherited the shared Source **Serif** display face for the wordmark and page headings; it now uses the sans face instead, matching the pre-theme all-sans look. This is a shared token change (`html:not([data-theme])` in `common/styles/tokens.css`), so any surface on the Base family — including the editor's default — reads sans; the Archive/Selenized/Swiss families keep their own type voices.
-- **A dedicated Cleanup page with a live "reclaimable" sidebar badge and per-channel include/exclude.** Space-reclaim cleaning now has its own home (`/cleanup`, in the Pool nav) instead of being scattered across each channel's detail page. It opens with a serif **reclamation console** readout — the total reclaimable audio across channels, a token-colored breakdown bar (transcribed audio / extra formats / wrong-format), and a per-channel ledger that runs the same `cleanAudioAction` / `cleanExtraAudioFormatsAction` / `removeWrongFormatAudioAction` sweeps (plus failed-list housekeeping) the channel page does — the per-channel `CleanupStage` stays put. The sidebar **Cleanup** item carries an amber badge with the running reclaimable total (e.g. `12.4 GB`), recomputed on each auto-refresh tick. A per-channel **Counted / Excluded** toggle (new `ChannelConfig.excludeFromCleanup`, persisted in `config.json` like `excludeFromBuild`/`excludeFromSync`) holds a channel's space back from the total without disabling its sweeps — e.g. keep a finicky-to-redownload channel's audio around for now. The headline total uses the primary "clean audio" reclaim only (the three sweeps overlap, so they aren't summed). See `editor/app/cleanup/{page.tsx,lib/loadCleanup.ts,components/{ChannelCleanupCard,ChannelCleanupToggle}.tsx}`, `editor/app/channels/actions.ts` (`toggleChannelCleanupInclusionAction`), `editor/app/{lib/nav.ts,layout.tsx}`, and `common/lib/channelConfig.ts`.
-- **Badges and alerts are now colorful and theme-aware in every theme.** The design system only tokenized red (`--destructive`); every green/amber/blue status was hardcoded Tailwind palette that ignored the four theme families. New semantic tokens — **`--success` / `--warning` / `--info`** (each with `-foreground` + a soft fill partner) plus `--destructive-soft` — are defined across all eight family blocks (base/archive/terminal/swiss × light/dark) and registered in `@theme inline`. The shared kit `Badge` gains `success` / `warning` / `info` / `brand` soft-filled variants, and a **new `Alert` component** (`common/components/ui/alert.tsx`) replaces ad-hoc notice divs. High-traffic status UI is migrated onto the tokens: channel stage badges (`StageBadge`), worker state badges + dots (`WorkersView`, `MonitorWidget`), the scheduler status column (`SchedulerView`), and the widget disk strip — so each recolors correctly in Archive, Terminal, and Swiss, light and dark. See `common/styles/tokens.css`, `common/components/ui/{badge,alert}.tsx`.
-- **Per-video "Exclude from truncated check".** Videos that legitimately have no speech for their back half (e.g. long ambient/music tails) kept getting false-flagged as truncated/incomplete transcripts. A new per-video toggle on the video panel writes an `exclude-truncated-check.json` marker (mirroring the `do-not-clean.json` pattern) that suppresses the flag everywhere it surfaces — the snapshot's `incompleteTranscript` + `shortAudio` buckets (so bulk actions, the `/actionable` lists, and channel row dots all skip it) **and** the per-video panel's "looks truncated" banner, which computes independently of the snapshot. The detection thresholds in `transcriptCoverage.ts` are unchanged (other videos unaffected). See `common/lib/excludeTruncatedCheck{,-server}.ts`, `common/controller/channelSnapshot.ts`, `editor/app/channels/[slug]/videos/[id]/{videoActions.ts,components/VideoPanel.tsx,page.tsx}`, and `editor/app/channels/[slug]/page.tsx`.
-- **Monitor widget: an optional cleanable-data indicator, and the builder remembers your last config.** The widget builder gains a **Show cleanable indicator** option (URL flag `clean=1`) that adds a strip showing total reclaimable audio (polled from a new `/api/widget/cleanable` route, honoring `excludeFromCleanup`), styled like the disk strip. The builder also now **persists its form to localStorage** (`ytdlp-tb:widget-config`, versioned, SSR-guarded — mirroring `jobsFilterStorage`), so it reopens with your last configuration; the shared/embedded `/widget` URL stays authoritative for what actually renders. See `editor/app/widget/{lib/config.ts,builder/{widgetConfigStorage.ts,components/WidgetBuilder.tsx},components/MonitorWidget.tsx}` and `editor/app/api/widget/cleanable/route.ts`.
-- **The whole editor now follows the selected theme (not just light/dark).** Dozens of editor screens — the Active Jobs cards, channel pipeline/stage views, the jobs table, deploy/scheduler/workers panels, settings/sites forms, the widget builder, and more — hardcoded Tailwind palette colors (`bg-white`/`dark:bg-zinc-900`, `text-zinc-500`, red/amber/green/blue status colors) that tracked light/dark but **ignored the theme family**, so they stayed zinc/white in Archive, Selenized, and Swiss. Every one of these (~100 components across `editor/app` + shared `common/components`) is migrated onto the semantic tokens (`bg-card`, `border-border`, `text-muted-foreground`, `text-foreground`, and `destructive`/`warning`/`success`/`info` for status), so they recolor correctly in every family. The one deliberate exception is the audio-probe "scanning" fill (violet) and chart data-series colors, which are intentional encodings.
-- **Two more selectable themes + a theme picker, and the families now feel distinct.** A palette menu beside the light/dark toggle (in every app — editor, export, homepage) switches the theme family between **Archive** (warm reading room), **Selenized** (the Solarized successor — teal-slate/warm-tan), **Swiss** (red/black/white editorial), and **Base** (neutral), each in light/dark/system, persisted client-side. Beyond color, families now also carry their own **corner radius and type voice**: Swiss is hard-cornered (0px) in a neo-grotesque (Archivo), Selenized uses calm rounding with a code voice (JetBrains Mono display + IBM Plex Sans body), Archive/Base keep soft corners and the Source serif family. The token vocabulary also grew — `--surface`, `--border-strong`, `--faint`, `--brand-ink` are now defined for **every** family and exposed as utilities (`bg-surface`, `border-border-strong`, `text-faint`, `text-brand-ink`), so more chrome can be themed. Each family is still a pure CSS token swap in `common/styles/tokens.css` — no markup changes — wired through `common/components/{themeConfig.ts,ThemeMenu.tsx}` and `common/styles/fonts.ts`. (A previously-selected "Terminal" family migrates automatically to "Selenized".)
-- **Per-site brand accent.** Each site's config (under Sites) gains an optional **Brand accent** field — a hex color that overrides the family brass on that site's public build. Validated on save; blank inherits the family accent. The long-running action logs (download, transcribe, build, deploy, …) are also restyled onto the shared kit (Button + design tokens), keeping their exact accessible labels. See `common/lib/accent.ts`, `common/lib/site.ts`, `common/components/StreamActionLog.tsx`, `editor/app/sites/{components/SiteForm.tsx,actions.ts}`, and `export/app/layout.tsx`.
-- **Command-first cockpit: a ⌘K command palette that runs actions, plus a live dashboard.** Press **⌘K / Ctrl-K** anywhere to open a fast command palette (rebuilt on cmdk) that jumps to any page **and runs pool/site actions inline** — Sync all channels, Refresh all reports, Detect duplicate shorts, Retry all failed jobs, Drain the queue, Pause/Resume all workers — each firing the server action and reporting the result as a **toast** (Sonner). On a channel page it also offers the channel's status filters. The sidebar and palette now read from one shared nav source (`editor/app/lib/nav.ts`) so they never drift (the old palette listed a stale 5-route subset), and every nav item gains an icon. The **Dashboard** gains a **Live jobs** panel that polls running jobs in real time (reusing the active-jobs monitor), and the whole editor shell is restyled compact onto the shared design tokens/kit with the Source type family. See `editor/app/{layout.tsx,page.tsx,lib/nav.ts,components/{CommandPalette.tsx,ChangelogNavLink.tsx}}` and `common/components/ui/sonner.tsx`.
-- **New "Family hub URL" setting.** Settings gains an optional hub URL (e.g. `https://archilyzer.com`); every export site renders a link back to it in the header/footer, tying the family of public sites together. Leave it blank for no hub link. Normalized to an absolute http(s) URL on save (`settings.homepageUrl`; see `common/lib/settings.ts`, `editor/app/settings/{components/SettingsForm.tsx,actions.ts}`). Part of the family-wide UI redesign that also restyles the public export sites.
-- **Source-truncated downloads are now caught at download time, and the yt-dlp download format is selectable (Odysee defaults to `original`).** A download that completed (yt-dlp exit 0) but whose audio is far shorter than the video — e.g. an Odysee/LBRY video whose every HLS rung is CDN-truncated to a few minutes while the full audio lives only in the multi-GB `original` format — used to pass the audio-check (the short file is structurally *clean*) and get transcribed, so only the **post-transcription** coverage detector caught it. Two changes fix this. **(1) A download-time duration guard:** after each managed download the produced `audio.<fmt>` is probed with **ffprobe** (new `common/ytdlp/ffprobeDuration.ts`, `paths.ffprobeBin` / `FFPROBE_BIN`) and compared to the metadata duration using the same thresholds as the transcript-coverage detector (`isShortAudio` in `common/lib/transcriptCoverage.ts`: non-livestream, ≥10min, <50% covered). A large shortfall is recorded as a new terminal **`failed-short-audio`** status with the measured `{audioDurationSec, expectedDurationSec, coverage}` — caught **before** transcription (the inline no-subs-fallback whisper is pre-checked too). The short stub is **kept on disk** (so it isn't silently re-downloaded into a loop) and is **not** classified for platform backoff (a truncated stream isn't a transient error). Surfaced as a new **`shortAudio`** snapshot bucket (mirroring `corrupt-full-source`: excluded from `downloadedNoTranscript` so it's never auto-transcribed, and from the wrong-format cleanup so the kept file isn't deleted) → a **`short_audio`** channel-list filter + **Select short-audio** quick-select + bulk **Re-download truncated** action, a **"Download was truncated at the source"** banner on the video page with a one-click **Re-download as Original**, a **"Channels with truncated downloads (short audio)"** `/actionable` section, and a **"Truncated download — audio far shorter than the video (file kept)"** download badge. The auto-download runner lands a `failed-short-audio` unit as a failed job. **(2) Selectable, platform-aware download format:** the previously-hardcoded `-f bestaudio/worst` is now a preset (`auto` | `original` | `bestaudio` | `bestvideo_audio`; `common/ytdlp/downloadFormat.ts`) resolved with the same override chain as `audioFormat` — per-download (the redownload **Format** picker) > per-channel (`ChannelConfig.downloadFormat`, a select in the channel form) > global (`SiteSettings.downloadFormat`, a select in **Settings**). **`auto` is platform-aware:** Odysee/LBRY resolves to `original/bestaudio/worst` (so the truncation is avoided at the source); everything else stays `bestaudio/worst`. The source platform is taken from the prefetched metadata extractor. The short-audio bucket's re-download reuses the per-video fixer (delete the stub → re-fetch with the per-source default → re-transcribe) via the existing `redownload-incomplete-bucket` job (now also driving `ReplayBucket: "shortAudio"`). See `common/ytdlp/{downloadFormat.ts,ffprobeDuration.ts,downloadOneManaged.ts,runYtdlp.ts}`, `common/lib/{transcriptCoverage.ts,downloadOutcome.ts,settings.ts,channelConfig.ts,paths.ts,transcripts-server.ts}`, `common/controller/{channelSnapshot.ts,autoRunner.ts}`, `common/jobs/jobSpec.ts`, `editor/app/channels/[slug]/{incompleteTranscriptActions.ts,bulkVideoActions.ts,lib/{fixIncompleteTranscript.ts,videoRows.ts,videoRowsServer.ts,stageStatus.ts},components/VideoListPane.tsx,videos/[id]/{videoActions.ts,components/VideoPanel.tsx}}`, `editor/app/{settings/{actions.ts,components/SettingsForm.tsx},channels/components/{ChannelForm.tsx,parseChannelForm.ts},actionable/{lib/loadActionable.ts,page.tsx,components/InlineActionButton.tsx},jobs/jobReplayRegistry.ts}`, and `editor/e2e/download-format-guard.spec.ts`.
-- **A complete-but-corrupt audio download is no longer re-downloaded forever — it's kept and flagged instead.** When an audio-checked download finished (yt-dlp exit 0, all bytes) but its final integrity probe came back `malformed`, the orchestrator rolled the whole file back and re-downloaded it — repeatedly. For a fast download that completes inside one checkpoint interval no `.good` baseline ever exists, so each "rollback" discarded the entire file and re-fetched it from scratch (a 2.46 GB Odysee video looped until the rollback cap, leaving no audio behind), and a cancel landing mid-loop was deferred behind the next full re-download. Now the final-probe path is bounded: a malformed final probe triggers **exactly one** re-download; if it's still malformed, the orchestrator **keeps the downloaded container on disk** (for inspection) and records a new terminal **`corrupt-full-source`** status — re-downloading a complete file can't change a deterministic verdict. This is distinct from `failed-corrupt-source` (a download that never finished, driven by the checkpoint rollback cap). Surfaced as a new **`corruptFullSource`** channel-snapshot bucket → a Download-stage summary line ("corrupt full source (file kept)"), an **amber download-dot** + `corrupt_full_source` row status in the per-channel video list (folded into the *No audio* filter), and a **"Corrupt full source — download completed but audio is malformed (file kept)"** badge on the video page. It's terminal and **not** retried (auto-runner treats it as skipped; the kept file is an artifact, so it isn't re-queued) and **not** counted as a failed transcription or a transcode-pending video. Cancellation is also fixed: a pending abort now wins over an in-flight rollback decision, and the loop re-checks the abort signal after the final probe so Cancel stops it promptly. See `common/ytdlp/audioCheckedDownload.ts`, `common/ytdlp/downloadOneManaged.ts`, `common/lib/downloadOutcome.ts`, `common/controller/{autoRunner.ts,channelSnapshot.ts}`, `common/ytdlp/runYtdlp.ts`, `editor/app/channels/[slug]/{lib/videoRows.ts,lib/videoRowsServer.ts,lib/stageStatus.ts,components/VideoListPane.tsx,videos/[id]/components/VideoPanel.tsx}`, and `editor/e2e/audio-check-scenarios.spec.ts`.
-- **The Deploy page is reworked around a clearer build/deploy lifecycle, with one-click build-then-deploy and batch multi-site builds.** The page now reads top-to-bottom as you'd actually ship: **Release notes** (the `## [Unreleased]` changelog preview + Cut release) → **Build & deploy** → optional **Individual steps** → **Build multiple sites**. A new **Build & deploy** button runs the build and, only if it succeeds (and wasn't cancelled), deploys it — as a single managed job with one combined streamed log and one Cancel (`buildAndDeployAction`, a composite `runManagedFunction`; cancelling mid-build skips the deploy). The new **Build multiple sites** panel kicks off a build (optionally build+deploy) for several sites at once, each rendered as its own live status-chipped log lane (`BuildSitesPanel` + `JobLane`); in Basic mode the jobs serialize on the shared build/deploy queue (the `export/` output tree is shared), with a note that true parallelism arrives with Docker mode. A **Build mode** toggle (Basic | Docker) on the page persists the choice as the default (`settings.buildPipeline`, also editable on Settings); Docker mode is a follow-up and currently falls back to a basic build with an inline notice. The build/deploy commands now share a child-streaming helper (`common/jobs/runChild.ts`) and mode-routing core (`editor/app/deploy/buildDeployCore.ts`). See `editor/app/deploy/{page.tsx,buildAction usage,components/*}`, `editor/app/build/buildAction.ts`, and `common/lib/settings.ts`.
-- **Truncated transcripts are now detected and flagged for re-download.** When an audio download silently stops early (yt-dlp exits `ok`, `download-outcome.json` records success), whisper transcribes only the few minutes that landed — so a 2h22m video ends up with a ~7-minute transcript and nothing warns you. A new coverage check (last cue end ÷ video duration) flags any non-livestream video ≥10min whose transcript covers <50% of its runtime. The single source of truth is `common/lib/transcriptCoverage.ts` (`transcriptCoverage` + `isIncompleteTranscript`, with named thresholds), read from each video's `transcript.cues.json` so the existing corpus is flagged with no migration. Surfaced everywhere: a new **`incompleteTranscript`** channel-snapshot bucket → an **"Incomplete transcript"** filter chip and an **amber transcribed-dot** in the per-channel video list; a warning banner on the video page ("Transcript covers 6:52 of 2:22:21 (4.8%)…") with a one-click **Re-download & re-transcribe** button; and an **"Channels with incomplete (truncated) transcripts"** section on `/actionable`. The fix action (`redownloadIncompleteTranscriptAction`) deletes the truncated audio first, then re-downloads and re-transcribes — re-running whisper alone would just reproduce the short transcript. See `common/controller/channelSnapshot.ts`, `editor/app/channels/[slug]/{lib/videoRows.ts,lib/videoRowsServer.ts,lib/stageStatus.ts,components/VideoListPane.tsx,videos/[id]/{components/VideoPanel.tsx,videoActions.ts,page.tsx},page.tsx}`, and `editor/app/actionable/{lib/loadActionable.ts,page.tsx}`.
-- **Fix truncated transcripts in bulk — two buttons, in three places.** The per-video fix now has channel-wide and cross-channel counterparts, each offered as a **batch re-fix** (queues one job that removes the truncated audio → re-downloads → re-transcribes every flagged video in place; the transcript is never gapped) **and** a **clear & re-queue** (deletes the truncated audio + transcript so the videos drop back into the normal *undownloaded → needs-transcript* pipeline, then enables + starts the auto-download/auto-transcribe runners so they reprocess automatically). Both appear on the **`/actionable`** "incomplete transcripts" section — per-channel **Re-download & re-transcribe** / **Clear & re-queue** buttons (replacing the old "Review"-only link) plus a section-header **Re-fix all** / **Clear & re-queue all** that acts across every affected channel — and on the **channel page bulk bar** as two new Action options with a new **Select incomplete** quick-select. The clear path needs no archive pruning: `undownloadedIds` is derived purely from on-disk artifacts, and a single-video re-download isn't archive-gated. Destructive clears are confirm-gated everywhere; enabling the runners is disclosed in the confirm (note: the auto-queue policy must cover the channel for auto-reprocessing — cleared videos also surface in the existing "Download missing" / "Transcribe pending" sections as a fallback). New shared helper `editor/app/channels/[slug]/lib/fixIncompleteTranscript.ts` is the single source of truth for the per-video fix/clear, reused by the per-video action, the new `redownload-incomplete-bucket` batch job (bookmarkable; re-derives the live `incompleteTranscript` bucket), the bulk-bar wrappers, and the global actions. See `editor/app/channels/[slug]/{incompleteTranscriptActions.ts,bulkVideoActions.ts,components/VideoListPane.tsx}`, `editor/app/actionable/{actions.ts,page.tsx,components/{InlineActionButton.tsx,FixAllIncompleteButton.tsx}}`, `common/jobs/{jobKinds.ts,jobSpec.ts}`, `editor/app/jobs/jobReplayRegistry.ts`, and `editor/e2e/incomplete-transcript.spec.ts`.
-- **Auto-queue rules with no bucket now draw from *all* of a runner's buckets, and auto-download can resume partial downloads.** A policy-tree rule left at the **"all buckets (default)"** setting (previously just labeled *default*) now draws from the **union** of every bucket that runner kind tracks — deduped, in priority order — instead of only the single primary bucket. This fixes channels (e.g. an Odysee channel mid-download) that quietly stopped being auto-downloaded once their remaining work drifted entirely into **partially-downloaded** videos: those have a `.part` file but no completed audio, so they live in the `partialDownloads` bucket and were **absent from `undownloadedIds`** — the only bucket auto-download used to load. The download runner now loads `partialDownloads` alongside `undownloadedIds` (partials first, so in-progress downloads resume via `downloadOneManaged` before fresh ones start), and exposes `partialDownloads` as a selectable bucket in the policy editor so you can dedicate a high-priority rule to resuming partials. The per-kind bucket lists are consolidated behind a single `bucketsForKind` source of truth shared by the runner, the per-rule pending-count helper, and the editor's bucket picker (so they can't drift). Note: a *bucketless* auto-transcribe rule now also drains `failedListed` after `downloadedNoTranscript` (it already loaded both); platform rate-limit backoff is unchanged and remains an independent reason a throttled platform may pause. See `common/jobs/autoQueuePolicy.ts` (`buildPendingByLeaf` + `bucketsForKind` + unit tests), `common/controller/autoRunner.ts`, and `editor/app/auto-queue/{page.tsx,components/PolicyTreeEditor.tsx}`.
-- **Hub homepage redesigned into a cross-site landing; the homepage page-creator is removed.** The hub's home page is now a single mobile-first cross-site landing (headline KPIs and one stacked activity chart with Metric [Transcribed/Downloaded] · Breakdown [By site/By channel] · Bucket [Week/Month/Cumulative] · Range [90d/12mo/All] · Display [Share/Counts] controls, plus a metric-aware site-links grid with sparklines and a "#1 this month" badge), built from a small `homepage-summary.json` pre-computed by `compose-homepage`. The separate `/stats` dashboard route folds into it. Consequently the hub's **Markdown-pages subsystem is dropped**: **Manage → Homepage** now edits only branding (the Pages list, New-page, and the page editor are gone), and the homepage config no longer carries a `nav`. The page server actions (`saveHomepagePageAction`/`deleteHomepagePageAction`), `editor/app/homepage/pages/*`, `PageEditor.tsx`, `common/lib/{homepagePages,homepageConstants}.ts`, and `paths.homepagePagesDir` are removed. See `editor/app/homepage/{page.tsx,actions.ts}`, `common/bin/compose-homepage.ts`, `common/lib/{homepageSummary,homepageChart}.ts`, and the `homepage/` package. (Re-addable later if needed.)
-- **One-click Retry for failed jobs (plus "Retry all failed").** A failed job that carries a replay descriptor (any bookmarkable kind — sync, download-missing, transcribe-all, retry-bucket, …) now shows a **Retry** button on the Jobs history table, and the page header gains a **Retry all failed** button whenever at least one such job is listed. Retry re-runs the job from its stored spec exactly like a bookmark re-run (so bucket jobs re-derive from the channel's *current* state), and the re-run **jumps ahead of other queued work** (it's promoted to the front of its queue, reusing the new reorder machinery) so a fix-and-retry runs next rather than at the back of the line. The spec is resolved from the live registry or, for an evicted/archived job, from its on-disk `<id>.meta.json` sidecar — so even a failure the 100-job cap has dropped is still retryable. Kinds with no replay descriptor (e.g. `import-one`) intentionally offer no Retry. See `editor/app/jobs/actions.ts` (`retryJobAction` / `retryAllFailedAction`), the new `RetryJobButton` / `RetryAllFailedButton`, and `editor/e2e/jobs-retry.spec.ts`.
-- **Reorder and promote queued jobs from Active Jobs.** A queued job's row now carries **Promote / ↑ / ↓** controls (mirroring the auto-queue policy editor's move buttons) to change its order within its queue — Promote sends it to the front so it runs next, ↑/↓ nudge it one slot. Only actionable moves render (the first-queued job shows no up/promote, the last no down), and a running job is never displaced. See `editor/app/jobs/components/ReorderJobButtons.tsx`, the `reorderJobAction` / `promoteJobAction` server actions, and `editor/e2e/jobs-reorder.spec.ts`.
-- **Job-queue internals unified onto one scheduler + one concurrency primitive (foundational refactor; no behavior change beyond the two features above).** The "foreground preempts background" priority was previously implemented twice (once for job queue ordering, once for worker-slot waiters); both now share a single `compareTier` comparator in a new `common/jobs/scheduler.ts` (priority tiers urgent/foreground/background, per-queue concurrency, and the reorder/promote operations the UI uses), which the registry delegates its queue ordering to. The auto-transcribe/-download runner's hand-written fill-to-capacity loop **and** the whisper batch's `Promise.all` are both replaced by one audited `runPool()` primitive (`common/jobs/concurrentRunner.ts`) that encodes the no-event-loop-spin wait once — structurally eliminating the "Drain all hangs" bug class rather than patching each loop. Per-kind metadata (label, drainability, bookmarkability) is consolidated into one `common/jobs/jobKinds.ts` table, and the bookmark/replay dispatch is now a data-driven lookup, so adding a job kind touches ~1 file instead of ~5. Covered by new unit tests (`scheduler`, `concurrentRunner`, `jobKinds`) and the existing queue/drain e2e suites.
-- **"Drain all" no longer hangs the server when an auto-transcribe/-download unit is actively running.** A second, distinct drain hang remained after the earlier parked-unit fix: the runner's internal `waitNext()` helper short-circuited to an *immediately-resolved* promise whenever a signal was *already* aborted — so once a soft drain fired (and `drainSignal` stays aborted for the rest of the run), every wait while in-flight units were still finishing returned with no delay. That turned the runner's poll loops into a timer-less **microtask spin** that starved the Node event loop (the spin was reached first in the main fill loop's at-capacity wait once the drain target dropped to 0, before the terminal drain-wait was ever hit). A transcription that drain deliberately lets finish completes via a child-process `exit` event — a *macrotask* — which the spin never let run, so the in-flight count never reached zero, a CPU core pegged, and the whole app appeared frozen. (The earlier fix only covered *parked* units, which settle via microtasks and so cleared even under the spin; a genuinely *running* unit depends on a macrotask and didn't.) `waitNext()` now only fast-paths a real pending `wake()`; an already-aborted signal falls through to a real timer, so all three wait sites pace instead of spinning while still being woken promptly by a finishing unit or by an abort firing mid-wait. The `whisper-all` batch was never affected (it `await Promise.all(...)` with no manual poll loop). See `common/controller/autoRunner.ts` (`waitNext`) and the new running-unit drain regression test in `editor/e2e/auto-queue.spec.ts`.
-- **First-class video-persistence UI (phase 5, the final phase): a Saved Videos area, per-channel retention controls, and per-video persist/unpersist.** The video-persistence subsystem built up over phases 1–4 is now driveable end to end from the editor. A new top-level **Saved videos** page (`/saved-videos`, in the Pool nav) summarizes the whole saved-video store — total count and size, per-channel breakdown (count, size, how many carry a backup checksum), the default store location, and the last backup time — and hosts the **backup configuration** (destination, scheduled on/off, interval) plus **Back up now** / **Verify backup** buttons. Each channel's **Cleanup stage** gains a **Retention & persistence** section (shown whenever keep-latest is on or the channel has saved videos) with live counts and three buttons: **Check kept videos** (re-probe the window for deleted-from-source videos and pin them), **Persist kept now** (a new bulk catch-up pass that re-fetches the source container for any in-window video whose source isn't saved yet — `persistKeptAction` / `common/controller/persistKept.ts`), and **Back up saved videos**. The **channel settings form** adds a Retention & persistence section: **keep latest** (window size), **extraction mode** (yt-dlp vs app-side ffmpeg), and a per-channel **saved-video store dir** override. Each **video page** gains a **Source video** card showing persisted status (file, size, stored time, keep reason, sha256, location) with an **Unpersist** control that moves the container back into the data dir, or a **Persist source video** button (re-fetch + archive) when it isn't saved yet. New job-kind label for `persist-kept`; `persist-kept` is re-runnable from bookmarks. The saved-video actions moved from `editor/app/savedVideos/` to `editor/app/saved-videos/` to match the route. Covered by `editor/e2e/saved-videos.spec.ts` and `common/controller/persistKept.test.ts`. (Deferred: surfacing kept-check/persist as Actionable-page rows, and streaming the player directly from the store — unpersist brings the container back to the data dir to play it.)
-- **Backups for the saved-video store: rsync mirror + per-backup manifest + drift verification (phase 4 of the video-persistence subsystem; backend + scheduler, UI lands later).** The (large, often irreplaceable) saved source videos can now be **backed up to a configured destination**. A backup walks every saved-video pointer across all channels (so per-channel store overrides are covered automatically) and **`rsync`-mirrors each container** into `<dest>/<slug>/<videoId>/` — incremental and resumable (`-a --partial`), additive (no deletes), so re-running only transfers changed or new files. It writes a **`backup-manifest.json`** at the destination root recording each container's canonical location, byte size, and a **streamed sha256**, and caches that hash back onto the live pointer. A **verify** step reads the manifest back and reports drift in four buckets — `missing`, `sizeMismatch`, `checksumMismatch` (re-hashing each present file), and `extra` (containers at the destination the manifest doesn't know about). New global settings block **`savedVideoBackup`** (`{ enabled, dest, intervalMinutes }`; a blank `dest` forces `enabled` off) plus a **`RSYNC_BIN`** env override. When enabled with a destination, the **sync scheduler** runs the backup automatically on its own cadence (a global, not per-channel, job — suppressed during quiet hours, tracked via `lastSavedVideoBackupAt`). Backups can also be run/verified manually via `backupSavedVideosAction` / `verifySavedVideoBackupAction` (managed jobs on a dedicated `saved-videos` queue). The destination is treated as a local filesystem path (a mounted backup disk). See the new `common/lib/savedVideoBackup.ts` (manifest types/parse), `common/controller/{backupSavedVideos,savedVideoInventory}.ts` (+ tests), `common/lib/paths.ts` (`rsyncBin`), `common/lib/settings.ts` (`savedVideoBackup`), `common/jobs/syncSchedulerState.ts`, `editor/app/savedVideos/backupActions.ts`, and `editor/app/scheduler/runTick.ts`.
-- **Saved-video store: persisted source videos move to a separate dir/disk, with retention pruning (phase 3 of the video-persistence subsystem; backend, UI lands later).** When the per-download persistence rule (phase 2) keeps a source video, the downloaded container is now **moved out of the per-video data dir into a separate saved-video store** — leaving only a small `saved-video.json` pointer behind — so the main data volume holds just audio + transcripts while the (large) source videos can live on another disk. The store root defaults to `<transcripts>/saved-videos`, is overridable globally via the **`SAVED_VIDEOS_DIR`** env var, and can be further overridden **per channel** (`savedVideosDir` in `config.json`); a video's container lands under `<root>/<slug>/<videoId>/`. The move is **cross-device-safe** (rename within a disk, copy-to-temp + atomic rename + unlink across disks) and **best-effort** — a failed move leaves the container in the data dir as `source-media.<ext>` (still persisted, just not relocated) rather than failing the download. **Transcription resolves from the store**: when no extracted `audio.*` exists, the transcribe fallback follows the pointer to the stored container (returned as a path relative to the video dir so both the local engine and the remote uploader read it correctly), so a kept-but-cleaned or archive-only video still transcribes. **Retention pruning** (the phase-2 follow-up) now bounds the store: the Clean-audio sweep also evicts any *keep-latest* container that has rolled out of the window — but **never** a manually-archived (`override`) or pinned/irreplaceable (`pin`/do-not-clean) one, distinguished by a `keepReason` recorded on each pointer. Reversible helpers ship for the upcoming UI: `unpersistSavedVideo` (move the container back) and `dropSavedVideo` (delete it). A reusable `checkDiskSpaceFor(dir, …)` lands so disk gating can target the store filesystem (used by the UI/backup phases). Note: source-video persistence is still skipped for audio-check channels (deferred), and saved-store counts aren't yet surfaced in the channel snapshot (lands with the phase-5 UI). See the new `common/lib/savedVideo.ts` (+ `savedVideo-server.ts` + tests), `common/controller/pruneSavedVideos.ts` (+ tests), `common/lib/paths.ts` (`savedVideosDir`), `common/lib/channelConfig.ts` (`savedVideosDir`), `common/lib/diskSpace.ts`, `common/ytdlp/{persistencePlan,downloadOneManaged}.ts`, `common/controller/{transcribeOne,cleanAudioFromTranscribed}.ts`.
-- **Per-download persistence rule + app-side audio extraction (phase 2 of the video-persistence subsystem; backend, UI lands later).** Each individual download now consults the channel's keep-latest rule (plus any per-run overrides) to decide *what to keep*: a video inside the keep-latest window — or one carrying a `do-not-clean` pin — downloads its **full source video** (`bestvideo*+bestaudio/best`) and the app extracts `audio.<fmt>` from it with ffmpeg, keeping the container as `source-media.<ext>` (a deliberately distinct name from `audio.<ext>` so it's never mistaken for cleanable audio); everything else stays audio-only as before. Crucially the keep decision is made per video by its **upload date against the channel's Nth-newest cutoff** (computed once per run via the new `computeKeepWindow`/`isInKeepWindow`), so the newest videos — which aren't on disk yet at download time — are correctly persisted. A new **`extractionMode`** channel setting (`"ytdlp"` default | `"app"`) selects who extracts audio for audio-only downloads; persisting always forces app-side extraction. **Per-run overrides** thread through `download-missing` (and the shared managed-download path): `keepSourceVideoOverride` (force keep/discard), `extractImmediately` (extract now + discard the container even on a keep channel — the disk-saving backfill case), and `audioFormatOverride`. **Transcription falls back to the source container** when no extracted `audio.*` exists (parakeet ffmpeg-slices any container), so a kept-but-cleaned video or an archive-only download is still transcribable. New per-video **"Archive source video"** action (`redownloadToArchiveAction`) re-fetches an existing video purely to grab + keep its source container without disturbing the transcript. Legacy `"ytdlp"`-mode downloads produce byte-identical yt-dlp args to before (no behavior change for existing channels). Note: persisted source containers currently remain in the data dir and are not auto-pruned when they roll out of the window, and source-video persistence is skipped (with a log note) for audio-check channels — both addressed by the saved-video store in phase 3. See `common/ytdlp/persistencePlan.ts` (+ tests), `common/controller/keptVideos.ts` (keep-window), `common/ytdlp/downloadOneManaged.ts`, `common/ytdlp/runYtdlp.ts`, `common/controller/transcribeOne.ts`, `common/lib/videoStatus.ts` (`source-media`/`isVideoContainer`), and `editor/app/channels/[slug]/pipelineActions.ts` / `videos/[id]/videoActions.ts`.
-- **New per-channel "keep latest N" retention rule that protects recent videos from the Clean-audio sweep and pins any that get deleted from their source (backend; UI lands in a later change).** A channel can set `keepLatest` (in `config.json` for now) to shield its newest N videos — by upload date, a rolling window — from the **Clean audio** cleanup: those dirs are skipped just like a `do-not-clean` marker, and the snapshot's reclaim estimates/cleanup buckets exclude them (a new `keptCount` is recorded). Because a kept video can later be **deleted from its source** (YouTube etc.) and become irreplaceable, a new **kept-deletion check** re-probes just the kept window's availability (reusing `runAvailabilityCheck` with `onlyIds` + `recheck-non-deleted`) and **permanently pins** any video found `deleted`/`private`/`members_only` with a `do-not-clean` marker, so it survives even after it rolls out of the window. The check runs on demand via the new `check-kept-deleted` managed job (`checkKeptDeletedAction`, re-runnable/bookmarkable) and automatically from the sync scheduler on its own cadence (new `syncScheduler.keepLatestCheckIntervalMinutes`, default daily; per-channel `lastKeptCheckAt` state; suppressed during quiet hours, capped per tick, and skipped for a channel just synced this tick). This is phase 1 of a larger **video-persistence** subsystem (per-download persistence rules, a separate saved-video store, and backups follow). See `common/lib/channelConfig.ts` (`keepLatest`), the new `common/controller/keptVideos.ts` (`computeKeptVideoIds` + tests) and `common/controller/checkKeptDeleted.ts`, `common/controller/cleanAudioFromTranscribed.ts`, `common/controller/channelSnapshot.ts`, and the scheduler wiring in `common/lib/settings.ts`, `common/jobs/syncSchedulerState.ts`, and `editor/app/scheduler/runTick.ts`.
-- **The monitor widget can now carry optional control buttons (Pause/Resume Transcriptions, Drain all).** The `/widget` view stays read-only by default, but a new **Show control buttons** option in the widget builder (URL flag `controls=1`) adds an interactive row at the top with the same **Pause Transcriptions** toggle (between-segment GPU release) and **Drain all** as the full app — so a pinned/iframe monitor can pause the GPU or wind work down without opening the editor. The controls stay visible even when the widget is otherwise idle (so you can pause preemptively), and the widget refetches its worker payload on a pause/resume so the toggle flips immediately instead of waiting for the next poll. Defaults keep the widget control-free, so existing links render unchanged. See `editor/app/widget/lib/config.ts`, the new `editor/app/widget/components/WidgetControls.tsx`, `editor/app/widget/components/MonitorWidget.tsx`, and `editor/app/widget/builder/components/WidgetBuilder.tsx`.
-- **"Pause Transcriptions" now frees the GPU between parakeet segments instead of running the in-flight video to completion.** The global pause control (renamed from "Pause all" → **Pause Transcriptions** / **Resume Transcriptions**) used to only stop handing out new worker slots — any in-flight transcription kept running until its whole file was done, so the GPU stayed busy. Pausing now *also* sends a graceful between-segment stop to any partial-capable in-flight job: a **parakeet** worker finishes the current window, caches it (`win-NNNN.json`), and exits `paused` (a skip, not a failure), so the GPU frees within one segment and the video resumes from its cached windows on the next run — the same mechanism as the per-worker **Stop & keep progress** button, now wired into the global pause. Non-parakeet engines (whisper-cpp, chough) keep prior behavior: they stop taking new work but run their in-flight file to completion. The button is also surfaced on the **Active Jobs** screen (`/jobs/active`) next to **Drain all**, not just the Workers page. `resumeAll()` restores each worker's pre-pause state as before. See `common/jobs/workerPool.ts` (`pauseAll`), the new `editor/app/jobs/components/PauseTranscriptionsButton.tsx` (shared by both screens), `editor/app/workers/components/WorkersView.tsx`, and `editor/app/jobs/active/page.tsx`.
-- **"Drain all" (and the per-runner Drain) no longer hangs the auto-transcribe runner.** Draining could leave the runner stuck "draining" forever — appearing to freeze the app — whenever one of its per-video transcription units was *parked* in the worker pool waiting for a free slot at the moment drain fired (much more likely now that auto units yield slots to manual transcriptions). The runner forwarded only its hard-cancel signal — never the soft `drainSignal` — to a parked unit's `pool.acquire()`, so a soft drain could never unblock it: the unit's promise never settled, the runner's in-flight count never reached zero, and its drain loop spun indefinitely (hard **Cancel**/**Stop** always worked, since that signal *was* forwarded). The runner now threads `ctx.drainSignal` into each transcription unit, so a parked (not-yet-started) unit unblocks and is skipped on drain while a unit already transcribing finishes normally — correct drain semantics, and the runner finalizes promptly. See `common/controller/autoRunner.ts` (the `launchUnit` drainSignal wiring) and the new parked-unit drain regression test in `editor/e2e/auto-queue.spec.ts`.
-- **Manual transcriptions now preempt auto-queued ones — click "Transcribe missing" (or any non-auto transcribe) while auto-transcribe is running and yours goes next.** The auto-transcribe runner shares the global transcription **worker pool** with every manual transcribe (channel batch, bucket retry, single-video), so they already serialized — but the pool granted freed worker slots in plain arrival order, so a manual transcribe could wait behind the runner's next auto pick. The pool's waiter queue is now **priority-ordered**: a manual (foreground) acquire is served before any parked auto (background) one, with FIFO preserved within each class — the same foreground-vs-background priority the per-platform sync queue already uses, now applied to the pool too. The auto-runner acquires its per-video transcription slots at **background priority**, so a manual transcribe jumps ahead: the in-flight auto transcription is **not interrupted** (it finishes — "drains"), then every freed worker goes to the manual work until it's exhausted, then auto resumes. Worker N-way parallelism is untouched (only *parked* waiters are reordered). A future "auto-sync queues its own transcriptions" path gets the same yield for free by acquiring as background. See `common/jobs/workerPool.ts` (waiter priority), `common/controller/transcribeOne.ts`, `common/controller/transcribeOneFromQueue.ts`, and `common/controller/autoRunner.ts`.
-- **Auto-download now shares the same per-platform queue as a manual Sync, so you can click Sync while auto-download is running without risking a 429.** Each auto-download video is now launched as a real job on the channel's platform queue (e.g. `platform:youtube`) — the same queue Sync uses — instead of running out-of-band. The job registry serializes them, so the two never spawn yt-dlp against one platform at the same time (in either direction), and each auto-download unit appears as its own row in Active Jobs with progress. A manually clicked Sync **preempts** the runner's *queued* auto-download units on that platform (it doesn't interrupt one already downloading), and a Sync is **refused with a clear message** (showing remaining seconds) while that platform is in a rate-limit cooldown. A 429 hit by either path now records the shared cooldown, so manual and automatic downloads back off together. See `common/jobs/registry.ts` (queue priority), `common/controller/autoRunner.ts`, `common/jobs/downloadBackoff.ts`, and `common/jobs/streamCommand.ts`.
-- **Auto-queue downloads now back off per-platform on rate limits instead of hammering the source.** When a managed auto-download hits an HTTP 429 / "too many requests" or a network error (most often on Odysee), the runner pauses *that platform* for an exponential cooldown (1 min, doubling up to 30 min, with jitter) while other platforms keep flowing, and the affected video is retried after the cooldown rather than being burned as a false success. Previously the runner discarded the download outcome, counted the rate-limited video as done, and immediately re-hit the same platform — so it never actually downloaded and re-stormed the source on every restart. Failures are now classified against the full yt-dlp stderr (not just the last few lines), so a 429 that yt-dlp logs as a WARNING before failing with a different final error still triggers the backoff. Cooldowns persist across restarts (`.auto-queue` state), and a successful download clears the platform's backoff. To stop a runner immediately, use Cancel in Active Jobs (Drain still waits for the in-flight video to finish, by design). See `common/jobs/platformBackoff.ts` and `common/controller/autoRunner.ts`.
-- **Per-video progress bar now advances for subtitle-only (YouTube-handling) downloads.** YouTube-handling channels fetch only subtitles (`--skip-download`), which yt-dlp reports with no byte total — so the Active Jobs progress bar sat empty and jumped straight to 100%. It now steps once per subtitle track (e.g. `subs 1/2`) using the track list yt-dlp announces, while real media downloads (Odysee/transcribe) keep their byte-based bar. See `createDownloadProgressParser` in `common/jobs/progressParsers.ts`.
-- **New Archilyzer homepage (hub): a standalone marketing/docs site, configurable here.** A fourth workspace package, `homepage`, builds a single instance-level static site (`output: "export"`) that sits above the per-content export sites — for info/docs pages and cross-site charts that don't belong on any one content site. It's managed from the new **Manage → Homepage** page: edit branding (title/header/description/tagline/public URL/Cloudflare project) and author **Markdown pages** (a slug + nav label + body, with a live markdown-to-jsx preview), which the hub renders server-side at build (the `index` page is the home body; every other slug gets a `/<slug>` route). Page content lives in the data dir (`sites/_homepage/`, a reserved id `listSiteIds()` ignores), so copy changes need no code deploy. The hub's **/stats** dashboard charts downloads/transcriptions completed across **all** content sites, leading with a per-site breakdown — the chart engine gains a **Site** grouping option (`groupBy: "site"`) that fans each video out to every site exposing its channel, resolved through a channel→sites map the hub supplies; combined whole-pool totals remain one series. Build with `pnpm build:homepage` (a `compose-homepage` step stages whole-pool stats + the channel→sites map ahead of `next build`). See `common/lib/{homepage,homepagePages}.ts`, `common/bin/compose-homepage.ts`, `common/components/charts/channelSites.tsx`, the `homepage/` package, and `editor/app/homepage/*`.
-- **"Cut release" can now create the release commit for you.** After turning `## [Unreleased]` into a dated semver heading, cutting a release used to leave the changelog edit sitting in your working tree to `git commit` by hand. A **Commit changelog** checkbox now sits next to the **Cut release** button (on `/changelog` for the editor and `/deploy` for the export), checked by default — leave it on and the cut is followed by a path-limited `git commit` of just that one CHANGELOG.md, with the message `Release <workspace> <version>` (e.g. `Release export 0.4.1`). To keep the release commit clean it commits *only* the changelog: if the working tree has any *other* uncommitted change, the cut is refused up front (nothing is written) with an error telling you to commit or stash those first — a dirty changelog itself is fine, so uncommitted `[Unreleased]` bullets get folded into the release commit. Uncheck the box to cut without committing, exactly as before. See `editor/app/deploy/cutReleaseAction.ts`, `editor/app/deploy/components/CutReleaseForm.tsx`, and the new `common/lib/git.ts`.
-- **"Stop & keep progress" no longer mislabels the paused video as a failed transcription.** Using **Stop & keep progress** on a busy parakeet worker (or any partial-capable engine) sends the engine a graceful SIGTERM so it stops after the current window and the video resumes next run. But if the engine took longer than execa's 5-second force-kill window to exit — which a parakeet window routinely does, since finishing/stitching one ~480s window outlasts 5s — execa force-SIGKILLed it and the resulting "Command was killed with SIGTERM … forcefully terminated after 5000 milliseconds" error escaped the pause handling: it was treated as a genuine transcription failure and the video was written to the channel's `failed-transcriptions` file *permanently* (so even though its completed windows were cached for resume, it was skipped as "failed" on every later run). The transcribe path now recognizes that a force-killed **requested pause** is still a pause, not a failure — it returns the `paused` outcome (a skip, not a failure), so nothing lands in `failed-transcriptions` and the next "Transcribe missing" resumes it from the cached windows. Hard **Cancel** and **Drain** were never affected (their abort signal already classifies the kill as a skip). A video wrongly blacklisted by the old behavior won't auto-prune (it has real audio) — clear it with the channel's **Clear failed transcriptions** action to retry. See `common/controller/transcribeOne.ts`.
-- **New Auto-queue: automatically transcribe (and download) across all channels by a configurable priority policy, instead of running one channel batch at a time.** Previously the only way to process pending work was to manually fire a per-channel batch (e.g. *Transcribe missing* on one channel), and since every transcription job serialized on a single queue, a batch ran to completion before any other channel got a turn — there was no way to say "do cornbreadman first, then fall back to hasanabi." The new **Auto-queue** page (under Pool → Auto-queue) adds two always-on runners, **auto-transcribe** and **auto-download**, each driven by a **policy tree**: order rules top-to-bottom for **strict** priority, or wrap rules in a group set to **round-robin** or **weighted-fair** (smooth weighted round-robin) to *alternate* between rulesets. A rule (leaf) matches a **channel**, a whole **platform**, or **all** channels, optionally narrowed to a snapshot **bucket** (e.g. prioritize `failedListed` retries over fresh `downloadedNoTranscript`), and any rule or group can carry a **max-workers** cap (a saturated subtree falls through to the next-priority sibling, like an HTB ceil). The highest-priority channel with available work claims the **next freed worker slot** — non-destructive, so a higher-priority video never kills an in-flight transcription, it just wins the next slot; when a channel's work runs out the runner falls back automatically. Transcription concurrency is bounded by the worker pool's eligible slots (so policy decides *which* video runs, the pool decides *how many*); downloads have no pool, so the runner gates to **one download per platform at a time**, matching the per-platform serial queue's politeness. Each runner is a real, drainable/cancellable job (visible on the Jobs pages), and the Auto-queue page shows live per-rule pending counts and a recent-pick log. Independent of the sync **Schedule** (which only decides *when* to re-fetch a channel) — manual batches keep working alongside it. Policies live in `settings.json` under `autoQueue` (defensively sanitized like `syncScheduler`); fairness cursors persist in `transcripts/.auto-queue/state.json`. See `common/jobs/autoQueuePolicy.ts` (pure selection engine + unit tests), `common/controller/autoRunner.ts`, `common/controller/transcribeOneFromQueue.ts` (shared per-video gating, also used by the existing whisper batch), and `editor/app/auto-queue/*`.
- - **Auto-queue refinements:** (1) the auto-transcribe runner now honors disabled workers — it no longer used CPU workers you'd turned off via Workers → *Set as default*. The runner started at boot and called the worker pool's `reconfigure()` before anything triggered the pool's lazy init, which set the pool's `initialized` flag *without* applying the saved `.worker-defaults.json` arrangement, so disabled workers came back enabled. The runner no longer pre-empts that init (it relies on the pool's own first-use init, which applies both settings and the saved default). (2) The runners no longer show up under a generic **"Other"** group on the **Active Jobs** screen — channel-less jobs are now grouped into their own labeled sections ("Auto-transcribe" / "Auto-download") via a shared `jobKindLabel` map (also used for friendlier kind labels in the running-jobs list), and each in-flight item is labeled `channel/videoId` so you can see which channel it's on. See `common/controller/autoRunner.ts`, `editor/app/jobs/jobKindLabels.ts`, and `editor/app/jobs/components/ActiveJobsLive.tsx`. (3) The Auto-queue page now has a **Drain** control (graceful stop: finish in-flight items, start no new ones, then stop) alongside **Stop** (hard stop, aborts in-flight); both the **Drain** and **Cancel** buttons on the **Active Jobs** screen also act on the runner jobs.
-- **Charts can now track content *added to the sites* over time, not just when creators uploaded it.** The chart engine previously only binned the time axis on a video's **upload date**. Two acquisition dates are now recorded per video — when *we* downloaded it and when *we* transcribed it — and the chart editor's X-axis gains a **Date field** selector (Uploaded / Downloaded / Transcribed) alongside the bin. A new **Content added** preset group ships three ready charts (cumulative *Library growth (added)*, cumulative *Transcribed over time*, and *Added per month* stacked by channel), and the default dashboard now includes the cumulative library-growth-by-acquisition chart so the progress view is present out of the box. Everything reuses the existing charts UI, so per-channel filtering, cumulative curves, CSV/PNG export, and shareable URLs all work unchanged. Acquisition dates come from the per-video `download-outcome.json` (`finishedAt`) and a new `transcribe-outcome.json` sidecar written when a transcript is finalized; the stats build falls back to file mtimes for content added before the sidecars existed. Requires a one-time data rebuild (`build:index` + `build:stats`) on the bumped `STATS_SCHEMA_VERSION`. See `common/lib/{stats,chartConfig,chartAggregate,chartShare,transcribeOutcome}.ts`, `common/controller/{buildStats,transcribeOne}.ts`, and `common/components/charts/ChartConfigEditor.tsx`.
-- **The Schedule page is now a one-stop editor for per-channel sync cadence.** The `/scheduler` page used to be read-only — you could see each channel's interval, last sync, and next-due time, but to *change* a cadence you had to open that channel's editor (Source → Auto-sync), one channel at a time, and the headline global toggles lived only in Settings. Now each row's **Interval** cell is an inline editor: pick a preset (Default / Off / Every 10–30m / Hourly / 6h / 12h / Daily / Weekly) **or** choose **Custom (minutes)…** and type an exact minute count, then **Save** — writing just `syncIntervalMinutes` to that channel's `config.json` and leaving every other field untouched (it does *not* go through the full channel-form merge). The page also gained a **Global controls** block to toggle the master **enable**, the **default interval**, and the **internal heartbeat** right there (the advanced knobs — concurrency, quiet hours, backoff — still link out to Settings). The underlying due logic is unchanged: a channel auto-syncs on the next heartbeat once `now − lastSyncedAt ≥ its interval`. The live status columns keep polling every 5s, but each row's editor holds its own state seeded once from the stored value, so a refresh can't clobber an in-progress edit. The preset list is shared with the channel editor (`editor/app/scheduler/intervalPresets.ts`), and `GET /api/scheduler/status` now carries each channel's raw `configuredIntervalMinutes` so the editor can tell *inherit-default* from *explicit-off* from *explicit-minutes*. See `editor/app/scheduler/actions.ts` and `editor/app/scheduler/components/{ChannelIntervalEditor,SchedulerSettingsForm}.tsx`.
-- **Scheduled sync can now run without an external cron job.** The sync scheduler previously only fired when an OS cron entry POSTed to `/api/scheduler/tick` (via `pnpm sync:tick`) — fine on a server, but a chore to set up just to call a function the editor already hosts in-process. The editor can now drive its own heartbeat through a **Next.js instrumentation hook** (`editor/instrumentation.ts`): on server startup it arms a single in-process timer that calls `runSchedulerTick()` directly — no HTTP, no cron, no token. Turn it on with **Settings → Sync scheduler → Internal heartbeat (seconds)**: `0` = off (keep using an external cron heartbeat), any positive value is clamped to `[15, 3600]`s and is the cadence the editor ticks itself at; the `SYNC_HEARTBEAT_SECONDS` env var overrides the setting at runtime. The timer is a self-rescheduling, `unref`'d `setTimeout` loop (so it never holds the process open and never overlaps a tick), re-reading the cadence each fire so a change takes effect on the next tick — though turning it on *from 0* needs a restart, since the timer is armed once at boot. It's modeled on the existing snapshot-scheduler timer and reuses the already-overlap-guarded `runSchedulerTick()`, so internal and external heartbeats are interchangeable and may even coexist. The **Schedule** page header now reports how ticks are driven ("internal heartbeat every N" vs. "external heartbeat (cron)"), and `GET /api/scheduler/status` carries the effective `heartbeatSeconds`. Defaults to off, so dev/test and existing cron installs are unchanged. One caveat for multi-instance deployments: the timer runs once *per server instance*, so a cluster against one data dir should set `SYNC_HEARTBEAT_SECONDS=0` on all but one — see `SCHEDULED_SYNC.md`.
-- **The channel video-list filters are now combinable (intersection), with a new Partial filter.** The filter chips above a channel's video list used to be single-select — clicking one replaced the last — and there was no way to filter for partial downloads at all. Each chip is now an independent toggle, and selecting several narrows to videos matching **all** of them (an intersection). A new **Partial** chip surfaces videos with a leftover `.part` download. The headline use: **Transcribed + Partial** finds videos that are already transcribed but still carry an orphaned `audio.<ext>.part` (e.g. cornbreadman's, where the transcript is done but the partial lingers as cruft) — previously unreachable because a video's status is a single mutually-exclusive value where `partial_download` outranks `transcribed`. To make combinations meaningful the **Transcribed** and **Partial** chips now key off the row's independent `transcribed` / `partial` flags rather than that status enum, so "Transcribed" alone now also includes transcribed videos that happen to carry a `.part` or a failed-transcoding marker (they *are* transcribed). The active set is mirrored to the URL as a comma-joined `?filter=a,b` via `history.replaceState` (so reload/share preserves it without an RSC refetch per toggle, and without racing the global auto-refresh), and the **All** chip clears the selection. See `editor/app/channels/[slug]/lib/videoRows.ts` and `editor/app/channels/[slug]/components/VideoListPane.tsx`.
-- **Bookmarks can be reordered on the management page, and the order carries to the compact menu.** Both the `/jobs/bookmarks` management list and the compact one-click menu atop `/jobs` and `/jobs/active` render bookmarks in stored order, but there was no way to change it — new bookmarks just landed on top. Each row on the management page now has **↑ / ↓** buttons that move the bookmark one slot (disabled at the ends), persisting the new order immediately to `transcripts/.bookmarks/bookmarks.json`. Because both views read the same array in order, reordering on the management page is reflected in the compact quick-run menu too, so you can put your most-used job first. See `moveBookmark` in `common/jobs/bookmarks.ts`, `moveBookmarkAction` in `editor/app/jobs/bookmarkActions.ts`, and `editor/app/jobs/components/BookmarksList.tsx`.
-- **"Retry partial downloads" no longer skips partials that carry an audio-check snapshot.** A transcribe channel's retry-bucket prefilter decided a video was "already complete" by looking for any `audio.*` file not ending in `.part` — which wrongly matched the audio-integrity snapshots (`audio.<ext>.part.good` / `.part.testing`) and sidecars (`audio.info.json`, `audio.live_chat.json`, `audio.*.tmp-*`) left in a partial video's dir. So a genuine `audio.<ext>.part` that happened to sit next to a `.part.good` snapshot got prefiltered out (`Prefilter: 0 missing destination files, N already complete` → `Nothing to fetch`), even though that same snapshot is excluded when the video is placed in the **Partial downloads** bucket. The prefilter (`destinationExists`) now reuses the same `isRealAudioFile` predicate the bucket uses, so the two agree and genuine partials resume. See `common/ytdlp/runYtdlp.ts` and `common/lib/videoStatus.ts`.
-- **One-click "Drain all" to spin work down before a server restart.** The `/jobs/active` header gained a **Drain all** button that, in a single confirmed action, **drains every running job** (lets in-flight sub-operations finish, starts no new ones) and **cancels every queued job** — so you can wind the queue down gracefully before restarting the server instead of draining/cancelling each job row by hand. It reuses the existing per-job drain path (`requestDrain` drains running jobs and cancels queued ones), so semantics match the per-row **Drain**/**Cancel** buttons exactly; non-drainable running kinds are marked draining and finish on their own. A confirm prompt guards the bulk action. See `editor/app/jobs/components/DrainAllButton.tsx` and `drainAllAction` in `editor/app/jobs/actions.ts`.
-- **Queued jobs now have a Cancel button right on the row.** On `/jobs/active`, a queued (not-yet-started) job previously offered no inline way to cancel it — you had to expand its log to find a control. Each active row now shows a **Cancel** button directly (queued *and* running), so you can drop a queued job without opening anything. The running-only **Drain** button is unchanged. See `editor/app/jobs/components/RunningJobsList.tsx`.
-- **The Bookmarks panel is now a compact one-click run menu, with management moved to its own page.** The bookmarks block atop `/jobs` and `/jobs/active` used to render a heavy multi-row card per bookmark — inline rename, tags, created/last-run metadata, a confirm-gated delete, and a full streaming **Run again** button that popped a tall inline log box on every run — which buried the live job list below it. It's now a tight **wrap of quick-access buttons**, one per bookmark, labeled with the bookmark's name (the full kind/bucket/channel shows on hover). Clicking a button **launches the job fire-and-forget** — no inline shell, no expanding log — and the relaunched job simply appears in the live list below (it runs to completion server-side whether or not the page watches its stream); the button briefly reads *Launching… → Launched ✓*, an empty re-derived bucket still reads as a neutral *"Nothing to retry right now"* line, and a bookmark whose channel was deleted is disabled with a *channel missing* hint. The heavier controls — rename, delete, created/last-run metadata, and the full streaming **Run again** with its inline log — move to a new **`/jobs/bookmarks`** management page, reachable via the **Manage** link in the menu header. See `editor/app/jobs/components/BookmarksMenu.tsx`, `editor/app/jobs/components/BookmarkRunButton.tsx`, and `editor/app/jobs/bookmarks/page.tsx`.
-- **ETAs are smarter in two edge cases, and audio-integrity probes now show their own progress bar.** Batch ETAs (downloads/transcripts on `/jobs/active` and the widget) now round the remaining work up to whole **parallel waves** instead of dividing straight through, so the tail no longer reads too low — e.g. *2 videos left across 4 workers* now estimates ~one full task, not "half a task." For **audio-checked downloads**, the per-task ETA used to come straight from yt-dlp, which only counts active download time and is blind to the periodic ffmpeg integrity **probe** pauses (each one transcodes the whole, growing `.part`, so later probes cost more). The download parser now measures each probe and **projects the remaining probe overhead** — extrapolating the per-probe duration as an arithmetic progression — onto yt-dlp's ETA, so a long audio-checked download no longer counts down faster than wall-clock. And while a probe runs (yt-dlp is paused, so the download bar would otherwise just **freeze**), the task now shows a distinct violet **"probing audio"** bar that fills against the estimated probe duration. See `editor/app/jobs/active/buildActiveJobs.ts`, `common/jobs/progressParsers.ts`, and `common/ytdlp/audioCheckedDownload.ts`.
-- **Bookmark a job to re-run it with one click.** Any bookmarkable job now shows a **Bookmark** button — on the `/jobs` table rows, the job detail page, and the `/jobs/active` rows — and saved jobs appear in a **Bookmarks** panel atop both `/jobs` and `/jobs/active`, each with a **Run again** button (which streams the relaunched job's log) and **Delete**. So bookmarking a *whisper-all on HasanAbiVODs3* or a *retry partial downloads on cornbreadman* gives you a button that re-launches the same job on the same channel later. Re-running **re-derives the work from the channel's current state** rather than replaying a frozen list: bucket jobs (retry partial downloads, transcribe downloaded-no-transcript, missing-transcript-and-no-download) store only the snapshot bucket category, so "Run again" always acts on whatever's in that bucket *now* (and reports "Nothing to retry right now" when it's empty) — matching how *Transcribe missing* already re-scans. Every channel-scoped pipeline, whisper, transcode, and cleanup job kind is bookmarkable; ad-hoc checkbox selections and one-off URL imports are not (there's no stable set to re-derive). A bookmark captures the job's kind, channel, and its flags (queue, audio format, abort-on-error, etc.) as a small replay descriptor persisted on the job (and its `<id>.meta.json` sidecar, so even an archived job can be bookmarked) and saved to `transcripts/.bookmarks/bookmarks.json`. Each bookmark can be **renamed** (auto-named `kind · channel` by default — handy when two bookmarks differ only in their options), shows its **created / last-run** times, and **delete is confirm-gated**. An empty re-derived bucket reads as a **neutral notice** ("Nothing to retry right now") rather than a red error, and a bookmark whose channel was since deleted is **flagged and its Run-again disabled**. See `common/jobs/jobSpec.ts`, `common/jobs/bookmarks.ts`, and `editor/app/jobs/runJobSpec.ts`.
-- **In-progress bars with no percentage now read clearly as "working, no number."** A task that hasn't reported a parseable percent yet (a download before yt-dlp's first percent, an engine before its header line) used to show a static 1/3-width bar in the same fill color as real progress — easy to misread as "stalled at ~33%." It's now a **full-width amber pulse**, visually distinct from the emerald/blue progress fill, so the indeterminate state is obvious. Applied consistently on the monitor widget, `/jobs/active`, and `/workers`.
-- **The monitor widget's elements are now individually toggleable.** Each piece of the `/widget` view has its own GET param (and builder checkbox), so you can compose exactly the at-a-glance view you want and hand off the link. New flags: hide the per-job **batch/channel progress bar** (`jobbar=0`) to show only the individual per-task bars; optionally carry the batch progress (e.g. `5/10 · ~2m left`) as text in each job's **one-line heading** (`headtext=1`) so you keep the count/ETA after hiding the bar; hide all **time estimates** (`eta=0`, counts only); hide the **disk indicator** strip (`disk=0`); and collapse the Workers strip to **colored dots only** by hiding worker names (`wnames=0`). The combination `jobbar=0&headtext=1` yields the "only individual task bars under a text heading" layout. All are additive and back-compatible (existing `compact`/`titles`/`idle` links render identically) and omitted from shared URLs at their defaults. See `editor/app/widget/lib/config.ts`.
-- **Audio-integrity check no longer copies the `.part` just to probe it.** In the default paused mode the download is held SIGSTOPped across each integrity probe, so the in-progress `.part` can't change underfoot — the probe now reads it in place instead of first reflink-copying it to a `.part.testing` snapshot. A copy is only taken when a checkpoint passes, to freeze the known-good `.part.good` baseline (and on malformed verdicts, no copy is taken at all). Legacy `resumeDuringProbe` mode is unchanged — it still snapshots up front since the child keeps writing during the probe. See `common/ytdlp/audioCheckedDownload.ts`.
-- **Corrupt-source / undownloaded videos are no longer counted as failed transcriptions.** A video the downloader couldn't produce a real audio file for — most often a malformed source the audio-integrity check gave up on (`download-outcome.json` status `failed-corrupt-source`), or a download whose extract-to-`mp3` left no usable audio — used to get picked up by **Transcribe missing**, fail with "no audio file found", and be written to the channel's `failed-transcriptions` file *permanently* (so even a later clean re-download was skipped forever). These were download/source problems mislabeled as transcription failures, and they piled up. Now the transcribe pass **skips any video with no real audio file** instead of attempting it (the missing-audio case is classed `no-audio` and treated as a skip, not a failure), so it never lands in `failed-transcriptions`; it transcribes automatically once a good audio file exists. The pass also **auto-prunes** existing bogus entries — on each run, any id on the failed list whose dir currently has no audio is removed (it'll retry once re-downloaded). These videos are now surfaced *distinctly* rather than as failures: a new **corrupt-source** channel-report bucket feeds the **Download** stage summary ("N corrupt source (needs re-download)"), and in the video list they show a **corrupt_source** status (amber download dot, no red failure dot) and appear under the **No audio** filter. See `common/controller/whisperBatch.ts`, `common/controller/transcribeOne.ts`, and the new `corruptSource` snapshot bucket.
-- **Quick availability check: spot videos that vanished from a channel with one cheap yt-dlp call.** A channel's **Diagnostics → Availability checks** stage gained a **Quick check (flat playlist)** button that fetches the channel's *current* flat playlist (one `--flat-playlist --print url` pass, no per-video metadata) and diffs it against the videos already on disk. Any known video absent from the fresh listing is flagged **maybe-missing** — it was deleted, made private, or **unlisted** (unlisted videos legitimately drop out of channel listings, so a flag is a candidate, not a verdict). The set is persisted to `channels/<slug>/maybe-missing.json` and re-surfaced through the channel snapshot (intersected with on-disk dirs), so it survives report regens. This is the fast first pass that avoids running the full per-video `--dump-json` check across a whole channel just to notice the handful that disappeared.
-- **"Full-check unexpected": resolve maybe-missing videos without re-probing the whole channel.** Alongside the maybe-missing list, a **Full-check unexpected** button (with its own concurrency control) runs the full availability check on *only* the maybe-missing ids, skipping any already known permanently gone (`deleted` / `private` / `members_only`) and re-probing the rest (`unlisted` / `needs_auth` / `public` / `error` / never-checked, since their status may have changed). This efficiently resolves e.g. deleted-vs-unlisted for just the suspect videos.
-- **Per-video availability history.** Each video's `availability.json` now keeps a `history[]` timeline that appends an entry only when the observed availability *changes* — so a flip like `public → deleted` is preserved rather than overwritten. Entries are tagged by source (`check`, `backfill`, or `download` — the last captures a public video the uploader deleted, observed at download time, without disturbing the explicit-check fields). The video detail page renders this as an **Availability history** card (reverse-chronological, "No availability changes recorded" when empty).
-- **The Jobs screen has persistent filters, and archived jobs keep their details.** A filter bar atop `/jobs` lets you hide jobs by **kind** and **status** and **search** by id / channel / video; the choices persist in `localStorage` and, by default, hide the noisy `refresh-report` report-regen job so the list shows the work you actually triggered (a **Reset filters** button restores the defaults, and a `Showing N of M · K hidden` summary makes the active filtering obvious). Separately, each job now writes a small `<id>.meta.json` sidecar next to its log capturing its kind, channel, video, status, and timings — so once the in-memory registry evicts it (it keeps only the 100 most-recent finished jobs) or the server restarts, an **archived** job still shows that metadata instead of a bare id. The sidecar writes are best-effort and never delay or break a job, and **Clear archived** removes the sidecars along with the logs. See `common/jobs/jobMeta.ts`.
-- **The Workers page can save the current arrangement as a launch default, and "Stop & keep progress" now takes the worker out of rotation.** A **Set as default** button (top of `/workers`) snapshots which workers are enabled right now; on the next server launch the pool starts exactly those workers enabled and **every other worker disabled** — including workers added later that aren't in the saved set. This persists the otherwise-transient runtime on/off state across restarts without touching `settings.json` (it's a small `transcripts/.workers/defaults.json` the pool reads on first use). Once a default is saved the button reads **Update default** and each included worker shows a small **default** badge; the default governs the launch-time seed only, so editing a worker's Enabled flag in Settings still takes effect as before. Separately, **Stop & keep progress** now drains the worker after finishing the current window (it ends *disabled*) instead of leaving it enabled to immediately grab the next video — matching **Drain**, which already ended disabled. See `common/jobs/workerDefaults.ts`.
-- **Downloads stop and stay blocked when free disk space runs low.** A new **Settings → Minimum free disk space (GB)** floor (default **5 GB**; set **0** to disable) guards every download against filling the disk. When free space on the transcripts data directory is at or below the floor, a download job is **prevented from starting** — the channel/import action returns a clear "Low disk space: X free, Y required" error instead of queuing — and a **running batch stops between videos**: the in-flight download finishes, no new ones start, and the batch ends cleanly (status *done*, partial progress preserved) rather than crashing into an out-of-space error mid-file. Free space is measured natively (`statfs`, no new dependency) and the check **fails open** — if it can't read the filesystem, downloads proceed rather than being wrongly blocked. The monitor widget (and `/api/jobs/active`) gained a compact disk indicator showing free space, which turns red and reads "downloads paused" when below the floor (and stays visible even when idle, so it explains why nothing is downloading). `store-playlist` (which writes no media) is not gated. See `common/lib/diskSpace.ts`.
-- **New read-only monitor widget (`/widget`) plus a builder to compose and embed it.** A compact, chrome-less page shows worker status (a colored idle/busy/draining/disabled/degraded dot per worker) and active-job progress bars at a glance — no sidebar, no command palette, and no action controls — so it fits in a small pinned window or an `<iframe>` for at-a-glance monitoring. It reuses the existing `/api/jobs/active` and `/api/workers` endpoints (polled live), and its initial paint is server-rendered for no flicker. What it shows is driven entirely by GET params: `jobs`/`workers` (toggle each section), `channel` (filter active jobs to one slug), `poll` (refresh seconds), `compact` (drop per-task detail), `titles` (section headers), and `idle=hide` (collapse to a tiny "Idle" line when nothing is active). A new **Monitor** page under the sidebar's **Pool** group (`/widget/builder`) exposes all of those as form controls, builds the shareable link with a **Copy** button, an **Open popup** button that launches the widget in a chrome-less `window.open` popup (a tab-less window) at the selected preview size, and live-previews the real widget in a sized iframe. To strip the app shell on exactly the widget route, the root layout now renders its sidebar/command-palette/auto-refresh through a small `AppFrame` client wrapper that hides them when the path is `/widget` (the builder keeps the normal shell). The per-worker payload builder shared by the Workers page and `/api/workers` was extracted to `buildWorkersPayload()` so the widget reuses it too.
-- **The channel video selector can bulk-remove audio files and wrong-format audio, and a new channel-wide sweep clears wrong-format audio in one click.** The selector pane's bulk action picker (channel page → video list) gained two operations, both pure filesystem ops that queue no job (so they never trigger a transcode). **Remove audio files** deletes each checked video's finalized `audio.<ext>` files while keeping transcripts, metadata, and any in-progress `.part` download (which can still resume). **Remove wrong-format audio** deletes only audio files that aren't the channel's target format — e.g. the `audio.m4a` / `audio.mp4` leftovers from downloads that failed yt-dlp's extract-to-`mp3` step *before* audio-integrity checking existed — even when that's a video's only audio, so it re-downloads cleanly. A matching **Select wrong-format** quick-select (shown when any such videos exist) checks exactly those videos, so the cleanup is one flow: **Select wrong-format** → action **Remove wrong-format audio** → **Apply** (each is confirmed first). For whole-channel cleanup, the channel page's **Cleanup** stage gained a **Remove wrong-format audio** section: a `type "remove" to confirm` sweep that walks every video dir and deletes all non-target audio — including the failed-extract orphans the existing **Clean extra audio formats** deliberately skips (it only de-dupes extras when the target file already exists). The sweep respects per-video **do not clean** markers and shows a reclaim estimate; the bulk action, being an explicit selection, removes regardless of the marker.
-- **The channel video selector can bulk-delete directories and clear failure markers, and its action bar is now an action picker.** The selector pane's bulk bar (channel page → video list) previously stacked separate Transcribe / Retry / Mark-untranscribable buttons; it's now a single **Action** dropdown + **Apply** button that also exposes two new operations. **Delete directories** removes each checked video's directory outright (`fs.rm` recursive) — a pure filesystem op that queues no job, so it never triggers a transcode the way leaving failed downloads in place can; this is the quick way to clear a batch of failed downloads. It's gated by an inline *type `delete` to confirm* box (the Apply button stays disabled until matched), mirroring the Clean-extra-formats pattern. **Clear failed markers** prunes the selected ids from the channel's `failed-transcriptions` and `failed-transcodings` files so they're retried on the next pass. Two quick-select helpers, **Select failed** (every video listed in either failure file) and **Select filtered** (every row matching the current filter + search), make the cleanup one flow: filter **Failed** → **Select failed** → action **Delete directories** → type `delete` → **Apply**. Both new actions report a per-id success/failure summary and refresh the channel report like the existing bulk actions.
-- **Audio-integrity checks now pause the download while they run, cutting re-downloaded bytes and HTTP 429 risk.** With audio-integrity checking enabled, the downloader periodically snapshots the in-progress `.part` and validates it with ffmpeg. Previously yt-dlp was only paused for the brief *copy* of that snapshot and then resumed immediately, so it kept downloading throughout the (longer) ffmpeg probe — and if the probe came back malformed, every byte pulled during the probe, plus everything back to the last good checkpoint, was discarded and had to be re-fetched. That wasted, repeated fetching is a prime driver of rate-limit (429) responses. Now yt-dlp stays suspended (SIGSTOP) across the whole probe and only resumes on a clean verdict; on a corrupt verdict it's killed while still stopped and rolled back, having downloaded zero throwaway bytes. The trade-off is a briefly idle source connection during each probe (probes are seconds; if a held connection is ever dropped, yt-dlp's own `-c` resume recovers on the next launch). This is the new default; a per-channel **Resume during probe (legacy)** checkbox (channel editor → Audio-integrity checking, `audioCheck.resumeDuringProbe` in `config.json`) restores the old resume-immediately behavior for comparison, and the `AUDIO_CHECK_RESUME_DURING_PROBE` env var overrides it for one-off runs.
-- **Channels can sync automatically on a schedule.** Each channel gained an **Auto-sync** setting (channel editor → Source): *Default* (inherit the global cadence), *Off*, or a concrete interval (every 10m / 30m / hourly / 6h / 12h / daily / weekly), stored as `syncIntervalMinutes` in `config.json`. Inspired by the Laravel scheduler, a single lightweight cron heartbeat (`pnpm sync:tick`, an ~30-line client) POSTs to the editor's new `/api/scheduler/tick`, and the **server** decides which channels are due — a channel is due when `now − lastSyncedAt ≥ its interval`, so a missed tick (server down, machine asleep) simply runs at the next one with no catch-up storm. All work runs **inside the editor** through the existing job queue and per-channel lock, so a scheduled sync can't collide with a manual **Sync** click, shows up live on `/jobs`, and feeds the same transcription worker pool — no second process, no new file locks. Global controls live in **Settings → Sync scheduler**: a master **enable** (off by default), a **default interval**, a **max concurrent syncs** cap (a tick queues at most `cap − running` channels, most-overdue first, rolling the rest to the next tick — which both bounds load and staggers a large due-batch so it doesn't hit the source all at once), an optional **quiet-hours** window, and **failure backoff** (after N consecutive failures a channel waits `base·2^(N-1)` minutes, capped, before retrying). Channels already marked **Exclude from sync** never auto-sync. A new **Schedule** page (`/scheduler`) shows each channel's interval, last sync, next-due time, last outcome, and any active backoff, plus a recent-ticks log and a **Run scheduler now** button; the same data is at `GET /api/scheduler/status`. The cron client targets the editor's port (3001) by default and is hardenable with a `SYNC_TICK_TOKEN` bearer token for installs that expose the editor — see `SCHEDULED_SYNC.md`.
-- **Sites can link to each other.** A site's form gained a **Public URL** field (the absolute URL it's served at, e.g. `https://jeralyzer.com`) and a **Related sites** section. The export footer automatically links to every *other* site that has a Public URL, so filling these in is all that's needed for cross-site links; a site left without a URL is simply omitted from the lists. The **Related sites** editor lets a site pull closely-related siblings to the front under named groups (e.g. Jeralyzer featuring Rekietalyzer under "MTG drama") — add a group, give it an optional heading, and check which sibling sites belong; everything you don't feature falls into a trailing "Other sites" group on its own. Groups reorder with ↑/↓. The picker only lists sites that actually exist, and featured ids for sites that were since deleted are dropped on save (with a heads-up note). It's a subtle, secondary feature — see the matching note in the export changelog for how it renders.
-- **The editor refreshes itself on a timer so its data stays live without a manual reload.** Every page now passively re-fetches its own server-rendered data on a configurable interval — so the sidebar badges (active/running job counts, changelog dot), channel reports, and any other on-screen figures keep up to date on their own. It uses Next's `router.refresh()` (the same mechanism the jobs list already used) mounted once globally in the root layout, so it covers every page and the shared sidebar with no per-page wiring. To avoid wasting work when you're not looking, it **pauses entirely while the browser tab is hidden** and does **one immediate refresh the moment you return** to the tab (rather than waiting out the interval); it also skips a tick while a previous refresh is still settling, so refreshes can't pile up. The cadence is set in **Settings → Auto-refresh interval (seconds)**: default **5s** (clamped 1–600), or **0 to disable** passive refresh completely. This replaces the jobs page's old bespoke 2.5s auto-refresh (the `/jobs/active` page keeps its faster 1s progress-bar polling, which animates per-task bars without a full re-render).
-- **Transcription is now driven by configurable workers instead of one global engine.** The old single **App** dropdown in **Settings → Transcription** is replaced by a **Transcription workers** list. Each worker is **one processing slot** — one transcription at a time — with its own engine (whisper.cpp / chough / parakeet) and config, and a priority given by its position in the list (top = preferred). To run several in parallel, add more workers; a **Copy** button duplicates one (e.g. point two copies at the same chough `--server` for two togglable server slots). A batch ("Transcribe missing", bucket, bulk, single-video) hands each video — per task — to the highest-priority free worker, so a fast GPU worker and a slower CPU worker (e.g. parakeet on the GPU + chough on the CPU) run side by side instead of one engine doing everything. Total parallelism is the number of enabled workers; the old per-run **Concurrency** control and the global **Parallel transcriptions** setting are gone (add/remove workers, or disable/drain one, to change load). A pre-worker `settings.json` migrates automatically to one worker per slot of the previously-selected app (the old parallel-transcriptions count becomes that many enabled copies), plus a disabled worker for any other engine you had configured, so existing installs keep their parallelism. Scheduling is a single process-wide pool, so two batches can't oversubscribe the same GPU. One-slot-per-worker also means you can disable a single slot to free *some* of a CPU/GPU while the rest keep transcribing.
-- **Remote workers: offload transcription to another instance of this app on your LAN.** Add a **remote** worker in the Settings list with the base URL of another instance (e.g. `http://gpu-box.lan:3001`) and a shared token. When a video is dispatched to it, this instance uploads the audio over HTTP, the remote transcribes it through *its own* worker pool (picking among its local engines), streams progress and log back, and this instance pulls the finished `transcript.json` and normalizes it locally — so the remote needs no knowledge of your channels, just CPU/GPU. The protocol lives under `/api/worker/*` and is **disabled unless `WORKER_TOKEN` is set** in the environment, so an instance is never an open transcription server by accident; every request carries `Authorization: Bearer <token>`, validated with a constant-time compare against the accepting instance's own `WORKER_TOKEN` (never against settings). Uploaded audio and the produced transcript live in a scratch dir that's cleaned up once the result is pulled (or the job is cancelled). If a remote returns a transport error mid-job, the video is automatically retried on another worker; a genuine transcription failure on the remote is not retried. On a transport failure the remote's `GET /api/worker/health` is probed, and a remote confirmed **down** is auto-disabled (shown "degraded" on the Workers page, with **Enable** to retry once it's back) so neither the current video nor later ones keep burning attempts on it — they fail over to a healthy worker. A worker that racks up repeated failures while still reachable is auto-disabled after a few strikes.
-- **New Workers page (`/workers`) with live status and runtime controls.** Lists every worker with its state (idle / busy / draining / disabled / degraded) and the video it's currently transcribing with per-task progress. Each worker can be **disabled** (stop taking new work immediately; in-flight transcriptions keep running), **drained** (stop taking new work but let the current video finish — the graceful "free up the GPU when it's done" path), or **enabled** again — without editing settings, so you can hand a CPU/GPU back to other programs and reclaim it later. A **Pause all** button disables every worker at once and remembers each one's state; **Resume all** restores them exactly. These runtime controls are transient (a restart returns workers to their configured enabled state, unless you capture the current arrangement with **Set as default**); the Settings list is where the persisted per-worker config lives.
-- **parakeet.cpp transcriptions are resumable, and can be paused mid-run.** The overlapping-segment wrapper now writes each window's raw parakeet-cli JSON to a per-audio work dir (`.<audio>.parakeet/`) as it finishes, and stitches the final transcript only once *all* windows are done (then removes the work dir). Re-running the same transcription picks up the cached windows and only does what's missing — yt-dlp-style resume, so a crash, cancel, or pause never loses completed windows. A busy parakeet worker on the Workers page shows a **Stop & keep progress** button: it finishes the in-flight window, stops and frees the worker (no transcript written yet), and the next "Transcribe missing" resumes from the cached windows and completes. Handy to reclaim a GPU mid-run. (whisper.cpp/chough run as a single pass and don't offer this.)
-- **Selectable compute device for parakeet.cpp.** A parakeet worker gained a **Device** field that forces the compute device — `cpu` to run on CPU, or a specific GPU like `CUDA0` / `Vulkan1`. parakeet.cpp's `parakeet-cli` has no `--device` flag and otherwise auto-grabs the first GPU the ggml registry reports, so the wrapper now exports the choice as the **`PARAKEET_DEVICE`** environment variable to the CLI (previously it was passed as a non-existent `--device` flag, which the CLI ignored — so a worker set to `cpu` still ran on the GPU). Combined with one-worker-per-slot and Copy, you can pin different parakeet workers to different devices.
-- **The Workers page and Active Jobs page cross-reference each other.** Each busy worker on `/workers` now shows what it's transcribing right now — the video (linked), the channel it's in, a live elapsed timer, percent, and the engine's progress detail — not just a bare bar. Conversely, every in-flight transcription on `/jobs/active` now says which worker it's running **on** (e.g. "Transcribing <id> on GPU"), so you can see how a batch is spread across your workers at a glance.
-- **Batches pause instead of failing when no worker is available.** If every worker is disabled (or you hit **Pause all**) while a transcription batch is running, the batch parks — it keeps its in-flight video to completion, starts no new ones, and stays **running** on `/jobs/active` rather than failing the remaining videos. Re-enabling any worker (or **Resume all**) immediately resumes it where it left off. A video whose worker fails for a transport reason (e.g. a remote worker that went away) is automatically retried on another worker before being recorded as failed.
-- **Git worktrees can run in parallel on non-colliding ports (dev tooling).** Two checkouts of the repo (via `git worktree`) can now run their dev servers and e2e suites at the same time without port clashes. A new helper, `scripts/worktree.mjs` (exposed as `pnpm wt`), assigns each worktree a port block offset by `index * 100` based on its position in `git worktree list` — the main worktree keeps the original defaults (editor 3001, test 3011, export 3010/3000/3020), worktree #1 gets 31xx, and so on. `pnpm dev:editor`, `pnpm dev:export`, `pnpm start:export`, and `pnpm e2e` route through `wt run`, which injects the assigned ports, so they "just work" per worktree; the editor/export `package.json` port flags and the Playwright configs now honor these env vars (previously `pnpm dev:test` hardcoded 3011, so a custom `PORT` only moved the URL Playwright waited on, not the server). E2E specs that hit the editor's test API now derive the base URL from `PLAYWRIGHT_BASE_URL` (centralized in `editor/e2e/baseUrl.ts`) instead of hardcoding `localhost:3011`. `pnpm wt add <branch>` creates a sibling worktree pre-seeded with `settings.json` and prints its ports; `--share-data` links it to the main worktree's downloaded `transcripts/` for read-mostly reuse (with an LMDB concurrent-write caveat). See `WORKTREES.md`.
-- **Channel reports refresh themselves automatically, on a global debounce.** Any action that changes what a channel report (the per-channel snapshot powering the Channels list, `/actionable`, and the channel page) would say now regenerates that report on its own when it finishes — no more manually clicking **Refresh report** after transcribing, transcoding, cleaning audio, checking availability, editing channel config, or the per-video file operations (delete file, set primary transcript, delete dir, mark untranscribable, archive/do-not-clean). Every report-changing action marks its channel "dirty" and re-arms one shared debounce timer; when activity settles, the scheduler regenerates each dirty channel's snapshot in parallel (reusing the existing `refresh-report` job, deduped against any refresh already running) and revalidates the affected pages. This happens **after each completed sub-operation within a batch**, not only when the whole batch finishes — so a long download or transcription run updates its report incrementally as each video lands, rather than staying stale until the end. The debounce coalesces bursts — videos that finish within the same window collapse into a single regen pass instead of one rewrite per operation. The window is configurable in **Settings → Report refresh debounce**: **Fast** (~1s after activity settles, no cap — the default), **Balanced** (~3s, 30s max), or **Lazy** (~10s, 60s max). The download/sync pipeline's previous behavior of regenerating its report inline is removed in favor of this one uniform mechanism (so its report now lags by the debounce window — ~1s by default — rather than being written synchronously). The manual **Refresh report** / **Update all reports** buttons are unchanged.
-- **Parakeet transcriptions show a per-video ETA.** While a video is being transcribed with the parakeet app, its per-task progress bar on `/jobs/active` now shows an estimate of how long that single video has left (e.g. `segment 3/12 · ETA 6:10`), alongside the existing segment count and percent. The wrapper (`scripts/parakeet-stitch.mjs`) measures each segment's real transcription wall-time and projects the remaining time as **average time per completed segment × remaining segments**, emitting it on its progress lines; the parser surfaces that ETA in the task detail. It appears from the second segment onward (the first segment has no average to project from yet). This is distinct from the batch-level "~4:30 left" estimate across all videos.
-- **Each active sub-operation shows how long it has been running.** Every per-task bar on `/jobs/active` (each in-flight download or transcription) now leads with a live `m:ss` timer of its own elapsed wall-time, ticking once a second — e.g. `1:42 · 25% · segment 3/12 · ETA 6:10`. The start time comes from the job registry's per-task `startedAt`, so the timer survives page reloads and reflects the real runtime, not time-since-open.
-- **The single-video download pipeline reconstructs the source URL when a video has no `metadata.info.json`.** Running the per-video **Download** / **Audio + Whisper** action resolves the video's URL by reading its `metadata.info.json` (`webpage_url`), then by matching the canonical id against the channel's stored `playlist`. When a video directory exists but has neither — e.g. a manually-placed video, or a channel whose playlist was never stored — the pipeline previously gave up with "Could not determine the video URL". It now re-creates the URL from the video's canonical id (its directory name) and the channel's platform (from `config.platform`, falling back to detecting it from `config.url`), so the download proceeds. Reconstruction is a last resort — the exact URL from metadata or the playlist always wins when present — and only fires when the platform is known (no blind guess); note a reconstructed Odysee URL drops the channel-name prefix, so it may not resolve.
-- **New transcription app: parakeet.cpp with overlapping-segment stitching.** A third **App** option (alongside whisper.cpp and chough) in **Settings → Transcription** transcribes via `parakeet-cli`, working around its two long-audio limitations: it only accepts a single WAV, and a >4GB PCM stream silently yields an empty transcript (dr_wav's data-chunk size is 32-bit). A standalone wrapper (`scripts/parakeet-stitch.mjs`) — usable both as the app's binary and directly on the CLI — slices the source into **overlapping** 16kHz-mono windows (the format parakeet-cli resamples to internally anyway), runs `parakeet-cli --json --timestamps` on each, then stitches the per-window word timestamps back into one transcript. The overlap guarantees a word clipped at one window's boundary is captured whole by the neighbour; a cut point in the middle of each overlap region decides which window owns each word, so nothing is dropped or duplicated across seams. Output is chough-native JSON (seconds-based `chunk_data`), so it flows through the existing format-aware normalize/index path with no downstream changes and is tagged `chough-json` in `transcript.cues.json`. The settings fields are **Model** (the `.gguf` path) and **Chunk size** (per-window length, default 480s); the underlying `parakeet-cli` binary and default model resolve from `PARAKEET_CLI` / `PARAKEET_MODEL` (overlap is tunable via `PARAKEET_OVERLAP_SEC`). Run on the CLI with `scripts/parakeet-stitch.mjs --model <gguf> <audio> [out.json]` (prints JSON to stdout when no output path is given). Per-window progress drives the job progress bars.
-- **Batch jobs estimate the time remaining.** A running batch's overall progress bar on `/jobs/active` now shows an estimate of how long is left (e.g. `~4:30 left`), alongside the existing `Transcripts: 12 / 50` count. The estimate is **remaining tasks × the measured average time per task**: each download/transcription's real wall-clock duration is folded into per-job running totals as it finishes, and that average is converted to a wall-clock figure using the parallelism observed so far (so a 4-way parallel transcribe batch isn't estimated as if it ran one at a time). It reads as **estimating…** until the first sub-operation completes (no average yet), and disappears once no work remains. Jobs without per-task tracking (e.g. storing a playlist) show no estimate.
-- **Audio-checked downloads recover correctly when an interrupted attempt left a malformed `.part` and the video id isn't derivable from its URL.** For sources where the canonical id only appears after metadata (e.g. Odysee), the integrity-checked downloader's pre-check would correctly roll back a corrupt leftover `.part` to its `.good` snapshot (or discard it), but the post-launch file discovery then ignored that same directory as a "pre-existing" one — so the freshly re-downloaded audio was never found and the download was recorded as failed. The pre-check now reports the directory it acted on, and discovery scopes to it, so the resumed download finalizes as `ok-audio-checked`. Unrelated stale `.part`s from other videos' interrupted attempts are still ignored (clean pre-existing parts aren't reported), so the cross-video protection is unchanged.
-- **First-class support for multiple transcription apps (chough + whisper.cpp).** Transcription is no longer hardcoded to whisper.cpp. **Settings → Transcription** now has an **App** dropdown (whisper.cpp / chough) with per-app fields, replacing the old flat Binary / Model / Args command. Each app owns how it builds its command line, what file it writes, how its output JSON is parsed, and how its progress output is read — so adding another tool is a small code module (`common/lib/transcriptionApps.ts`). **chough** is supported in both **local** and **remote** modes (set a **Remote URL** to transcribe via a `chough --server`, passing `CHOUGH_URL`; leave it blank for local), with optional **Chunk size** (`-c`) and **Model** (`CHOUGH_MODEL`) fields; its binary defaults to the `CHOUGH_BIN` env var. whisper.cpp keeps its Binary / Model / custom-args template. Because chough writes its output to the exact `-o` path (no `.json` appended, unlike whisper-cli's `-of`), the runner now renames the app-declared output file — fixing transcripts that previously failed to materialize under chough. Transcript parsing is **format-aware and back-compatible**: existing whisper.cpp `transcript.json` files and new chough files coexist, each parsed correctly by content sniff (chough's seconds-based `chunk_data` vs whisper's millisecond `transcription` offsets), with the detected format recorded per video in `transcript.cues.json` (`transcriptFormat`) and a fallback to whisper.cpp for anything unrecognized — so re-indexing a mixed corpus (including across shard machines running different tools) just works. An existing `settings.json` migrates automatically: legacy `transcribeBin`/`transcribeArgs`/`transcribeModel` map onto the matching app (a binary named `chough` adopts the chough app; everything else becomes whisper.cpp, preserving a customized args template).
-- **Global default for parallel transcriptions.** A new **Settings → Parallel transcriptions** field sets how many videos a "Transcribe missing" / bucket run transcribes at once when its per-run Concurrency input is left blank (default **2**, clamped 1–16). This replaces the old `PARALLEL_TRANSCRIBE_LIMIT` env-var default of 4 as the source of the default: the value is now persisted in `settings.json`, shown as the placeholder in each channel's Concurrency input, and used as the server-side fallback. The per-run Concurrency input still overrides it for a single run.
-- **Managed downloads skip live and upcoming videos by default.** Before downloading each video, the managed downloader now runs a quick metadata-only pass, then evaluates an app-level filter: videos that are currently live or scheduled/upcoming are skipped (their finished VODs still download normally — a skip isn't archived, so the next **Sync** / **Download missing** retries the video once the stream ends). A skip is recorded in the video's `download-outcome.json` (`status: "skipped-filtered"` with the filter name and reason) and logged to `download.log`; it never counts as a failure or aborts the batch, and the run's summary line reports how many were skipped. Skipped-live videos are surfaced in a new **Skipped: live or upcoming** bucket in the channel's Diagnostics (with a **Retry** to force an attempt now). On by default for every channel via a new **Settings → Skip live and upcoming videos** toggle, with a per-channel override (`config.json` `skipLiveDownloads`). To support the filter (and as a reusable building block for future per-video decisions), each managed download is now split into a metadata fetch followed by the real download, which reuses that metadata via `--load-info-json` so the second pass doesn't re-extract; the audio-integrity-checked download path re-extracts as before. Turn the toggle off for a channel that streams nothing to allow live captures.
-- **Global default social links, overridable per site.** Footer social links can now be defined once in **Settings → Social links** and apply to every site by default. Each site's form gained a **Use global default social links** checkbox (on by default): leave it checked to inherit the global list, or uncheck it to give that site its own links (an empty list shows none). Existing sites keep their current footer — sites that already had links stay as explicit overrides until you flip the toggle. Validation and SVG sanitization are unchanged and shared by both the global and per-site editors.
-- **Import a single off-playlist video by URL.** A channel's **Playlist** stage gained an **Import single video** panel: paste a video URL, click **Import video**, and that one video is downloaded into the channel using its normal handling (subtitles + metadata for YouTube channels, audio for transcribe channels) and appended to the archive — no need to add it to the stored playlist first. It reuses the same single-video downloader as the per-video "Download" button (including the auth-retry and no-subs→audio fallbacks), streams yt-dlp's output live, and runs on the channel's download queue. Transcription stays a separate step (use the per-video **Transcribe** button afterward), same as a normal download.
-- **Duplicate shorts detection on `/actionable`.** A new **Duplicate shorts** section finds the same short re-uploaded under a new id, re-titled, posted on another channel, or cross-posted to another platform. It runs a global, cross-platform pass: a cheap duration pre-cluster narrows candidates, then transcript content is compared (exact text hash → near-duplicate similarity → containment for "this short is a clip of a longer video"). Matches are **content-confirmed only** — a shared duration alone never clusters videos (that produced huge false clusters of unrelated same-length videos), and plain (non-karaoke) VTT captions that the index stores as zero cues are recovered by reading the raw transcript. Pick a scope — **Shorts only** (default ≤ 180s) or **All durations** (a heavier one-off that also surfaces clip-of-longer matches) — and click **Detect duplicates**; results list each cluster with its members (links), match strength, score, and cross-platform/cross-channel/clip-of-longer badges. It's **flag-only** — nothing is merged or deleted. Reads the cues + per-video stats written by **Build index** + **Build stats dataset**, so run those first; the report is written to `transcripts/duplicates.json` (also producible from the CLI via `pnpm --filter export run detect:duplicates [-- --all-durations]`).
-- **Cleanup sections now estimate how much disk space they'll reclaim.** Both the `/actionable` cleanup rows and a channel's **Cleanup** stage show an estimated size next to each cleanup operation — the **Actionable** page gained an **Est. reclaim** column for "cleanable transcribed audio" and "extra audio formats", and the channel Cleanup stage shows an "Estimated space to reclaim: ~N" line under each of its two actions. The estimate is computed when a channel's report is generated (it sums the `audio.*` files each cleanup would delete — all audio for transcribed-audio cleanup, every non-target format for extra-format cleanup — and respects "do not clean"), so refresh a channel's report to populate it. Older reports without the figure show `~0 B` until refreshed.
-- **Videos whose English captions only exist under a regional/auto code are now indexed.** YouTube occasionally serves a video's English subtitles only as `en-US`, `en-en-US`, or `en-orig` with no plain `en` track, so yt-dlp wrote e.g. `transcript.en-US.vtt` but never `transcript.en.vtt`. The index only recognized the literal `transcript.en.vtt`, so such a video looked untranscribed and never appeared in search. Build index now falls back to the best available English VTT (preferring `en`, then `en-orig`, then regional `en-US`/`en-GB`, then auto-translated `en-en-*`) while ignoring true translation tracks like `es-en-US`; whisper also treats these as already-transcribed. Re-run **Build index** to pick up affected videos already on disk. The video detail page now reflects the same fallback (it previously hardcoded `transcript.en.vtt`, so a regional-only video showed as untranscribed there).
-- **Pick which subtitle track is a video's transcript.** The video page has a new **Transcript source** section listing every `transcript.<lang>.vtt` track, marking the current primary, with a **Set as transcript** button that promotes any track to the canonical `transcript.en.vtt` (the chosen track is copied, so the original stays and the choice is reversible — delete `transcript.en.vtt` to fall back to the automatic English pick, or pick another track to switch). Useful when the auto-picked track isn't the one you want, or when a video's only captions are a non-English track.
-- **Diagnostics: "Non-standard transcript VTT name" bucket.** A channel's Diagnostics now lists videos whose transcript rides on a non-canonical VTT name (e.g. only `transcript.en-US.vtt`, or a foreign-language track) rather than the standard `transcript.en.vtt` — each links to the video so you can normalize/switch the primary transcript. Refresh the channel snapshot to populate it.
-- **Managed downloads stream yt-dlp's full output again, and progress bars now read a structured progress template.** The per-video archive marker is captured with yt-dlp's `--print`, which silently implies `--quiet` — so managed downloads ran nearly silent: `download.log` held little more than the archive line, and the per-video progress bars on `/jobs/active` never advanced (the `[download]` lines they parsed were suppressed). Managed downloads now re-enable full logging (`--no-quiet`) and emit a machine-readable `--progress-template` line (throttled to ~1/sec) that the progress parser reads directly instead of scraping the human progress text — so the bars advance reliably during real downloads (with speed/ETA detail) and `download.log` captures the whole run. The legacy `[download]` line parser is kept as a fallback for non-managed single-video/subs-only downloads.
-- **One global site selector scopes the editor to a single site.** The sidebar gained a single **site** dropdown (with an **All sites** option) that scopes the **Dashboard**, **Channels**, **Charts**, and **Deploy** views to the chosen site: Dashboard stats and the channel table, and the Channels list, now show only that site's member channels; Charts edits that site's dashboard; and Build static export / Deploy target it. Views that act on the shared pool — **Jobs**, **Active**, **Build**, and **Actionable** — always show every site regardless of the selector. The nav is regrouped to match: a **Site** group first, then **Pool**, then **Manage**. The selection persists in `localStorage` and is mirrored into the `?site=` URL param. Under "All sites", Charts and Deploy ask you to pick a specific site; creating a channel while a specific site is active also adds it to that site's membership. This replaces the per-page site tabs on Charts and the per-button site dropdowns on Deploy.
-- **Date-range filter on the public site's search.** The exported site's search filters gained a **Date** row (From / To pickers) that restricts results to videos uploaded within an inclusive range. The range saves in profiles and shareable search links and carries over to "Chart this search", reusing the same date mechanism the charts dashboard already uses. No editor-chrome change; it ships in every built site.
-- **Bulk checkbox transcribe/retry now queue like every other batch.** Selecting videos in the channel's list and clicking **Transcribe** (or **Retry download**) used to fire one queued single-video job per selection onto the channel's *platform* queue — diverging from "Transcribe missing"/"Download", which submit one batch job on the shared `transcription`/platform queue. The checkbox actions now submit a **single** batch job through the same path (so bulk and the stage buttons serialize together instead of contending), and the resulting job shows up in the page's running-jobs list. The selection bar gained the matching controls: a **queue** selector for each action (defaulting to `transcription` for transcribe and the platform queue for retry), a **Parallel** concurrency input for transcribe, and an **Abort on error** toggle for retry. Default-queue resolution for all batch features now lives in one shared helper (`common/lib/queueKeys.ts`) so they can't drift apart again. "Mark untranscribable" is unchanged (instant metadata write).
-- **Shard controls overhauled: "Save" button, 1-based `i / N` inputs, remaining-only transcribe slices, and always-on logging.** The shard inputs now read **`i / N`** (this shard's number first, 1-based — a 2-way split is `1/2` and `2/2`) and each box has its own hover title so they're no longer easy to mix up. A new **Save** button persists a slice to `shard-<op>.json` *without* starting the job — handy for dividing the remaining videos across machines (that share an identical synced copy) up front, and for confirming the slice actually saved; the saved-shard pill updates immediately. **Transcribe** sharding now splits only the videos that still need a transcript (e.g. 6 untranscribed videos split 2 ways → 3 each) instead of slicing the whole channel's video list. And every Download/Transcribe/Availability run now logs whether it's running a computed/saved shard slice or the full set, so a run that silently ignored shard inputs is no longer indistinguishable from one that honored it.
-- **Shard configs now show up immediately, and a shard job's progress bar totals only its slice.** Previously, setting a shard (N/i) and running **Download videos** / **Transcribe missing** did persist the slice to `shard-*.json`, but the saved-shard indicator never updated until you manually reloaded the whole page — so it looked like nothing saved. The indicator on the Download/Transcribe/Diagnostics controls (and channel counts, buckets, and failed-lists generally) now refresh as soon as a run finishes — including after a cancel — without a reload. The download slice is still taken over the *missing* set and snapshotted to the file, so a resume run with the same N/i reuses that exact slice instead of re-slicing the now-smaller set. A sharded job's progress bar on `/jobs/active` now totals only its own slice rather than the whole channel.
-- **Multi-site support (major change).** One editor instance can now power several public sites (e.g. "Jeralyzer" and "Rekietalyzer") over a single shared channel pool — a channel's downloads are stored once and reused by every site that includes it. A new **Sites** area (sidebar) manages each site's `sites/<id>/site.json`: its branding (title, header, description, tagline), social links, channel-group layout, and **which channels it exposes** (with a per-site group for each member, so the same channel can sit in different groups on different sites) plus its Cloudflare Pages project. Site-level branding/groups/social have moved **out** of global Settings, which now holds only operational config plus a new **Admin title** for the editor's own shell; the per-channel "Group" field is gone (grouping is configured per site). The **Charts** tab is per-site (pick the site at the top; each site has its own dashboard at `sites/<id>/chart-templates.json`). **Build index** / **Build stats dataset** now build the shared per-channel data once and a filtered bundle per site; the Deploy page's **Build static export** and **Deploy** gained a site selector and deploy each site to its own Cloudflare project. Upgrading an existing single-site install: the Sites page shows a one-click **Migrate** button (also a `migrate-to-sites` CLI) that lifts your current branding, channel groups, per-channel group assignments, and chart dashboard into a first site and reduces `settings.json` to operational config.
-- **"Audio + Whisper (skip pipeline)" no longer mistakes a live-chat sidecar for audio.** yt-dlp writes the live-chat track as `audio.live_chat.json` (it ignores the subtitle output path), so an interrupted download could leave an `audio.live_chat.json.part` behind. One-click whisper saw the `audio.` prefix, decided audio was already on disk, skipped the download, and then failed with "no audio file found". Live-chat sidecars (`.part` or completed) are now excluded everywhere audio is detected, so whisper downloads the real audio and transcribes as expected.
-- **Cleanup tasks on `/actionable`.** The global Actionable view now surfaces two housekeeping sections alongside the download/transcribe backlog: **channels with cleanable transcribed audio** (videos that have a whisper transcript but still keep their audio on disk) and **channels with extra audio formats** (videos with leftover audio files beside the channel's configured format), each with a per-channel count and an inline button. Because these delete files, the buttons pop a confirm dialog before queuing. The dashboard "Needs attention" card is unchanged — it still tracks only download/transcribe work.
-- **Per-video "do not clean" archive toggle.** Each video detail page has a new **Archive media** section to mark a video "do not clean". Marked videos are skipped by both cleanup jobs (their audio is preserved for archiving) and excluded from the cleanup counts on `/actionable`. The marker is reversible from the same toggle, and an "archived" badge shows on the video page while it's set.
-- **Twitch.tv support.** Twitch is now a first-class platform: Twitch channel/VOD URLs are auto-detected, "Twitch" is selectable in the channel form's platform dropdown and the charts platform filter, videos play via an in-browser Twitch embed, and downloads are queued on a `platform:twitch` queue like the other platforms.
-- **Downloads always land in the canonical `data/<id>/` dir.** Previously yt-dlp chose the directory from its own extractor id, which diverges from the app's URL-derived id on Twitch (`v<id>` prefix), Rumble, and Odysee — so a video's audio/metadata and its `download-outcome.json` could end up split across two sibling dirs and the editor couldn't resolve the video. The app now pins yt-dlp's output path per video, and a reconcile pass (run automatically when a channel's report regenerates) merges any pre-existing split dirs back together. A one-time migration command, `reconcile-video-dirs`, repairs all existing channels (`--dry-run` to preview). Sync and "download missing subtitles" now fetch one video per yt-dlp run (sync walks the channel a page at a time, downloading the diff against the archive and stopping once it reaches already-synced videos).
-- **Per-video download logs.** Each managed download now writes a `download.log` alongside the video's data, capturing that run's full yt-dlp output for after-the-fact debugging.
-- **Smarter queues for unrecognized sources.** When a channel's platform can't be detected, its jobs are now queued per-domain (e.g. `platform:vimeo.com`, with subdomains stripped) instead of all sharing the single `platform:unknown` queue — so unrelated unknown sources no longer block one another.
-- **Charts authoring + stats dataset.** A new **Charts** tab lets you author the default chart dashboard that ships to the viewer — add/edit/remove charts and configure axes, metrics, series, filters, and search-derived series with a live preview; edits save automatically and bake into the next export build. Search-derived charts now use the full layered query builder (AND/OR/NOT, nesting, per-layer scope/regex), matching the search page. A new **Build stats dataset** action on the Build page extracts per-video stats (views, likes, comments, follower count, duration, categories, language, cue counts) into `export/public/stats/` for the charts to read; it's incremental (mtime short-circuit) and also runs automatically as part of the static export build.
-- **Per-operation progress bars on `/jobs/active`.** Each running download/transcription now shows its own live progress bar parsed from the tool's shell output — yt-dlp's download percent (and fragment count) and whisper's transcribed position against the audio length. Batch jobs (Transcribe missing, Download from playlist, Sync, …) list a bar per in-flight video underneath the batch's overall bar; standalone single-video jobs get one too. The screen now polls about once a second so the bars advance live.
-- **"Drain" (soft-cancel) for batch jobs.** Alongside the existing Cancel (which immediately kills everything), running batches now offer **Drain**: it lets the in-flight operations finish, starts no new ones, then completes the job normally and releases the queue so the next batch can run. Hard Cancel still works at any time, including mid-drain.
-- **"Build static export" moved from the Build page to the top of the Deploy page**, co-locating it with the export changelog preview, "Cut release", and "Deploy static export" so the whole build → review → cut → deploy sequence lives on one page.
-- **The embedded single-video view on the channel page is now collapsed by default**, keeping the channel view compact. Clicking a video in the list expands the panel and scrolls to it; navigating away from a video (no video selected) collapses it again. A manual collapse/expand toggle is available on the panel, mirroring the video list's existing collapse.
-- **Unlisted videos now get their own clickable list in the channel availability diagnostics**, alongside Deleted/Private/Members-only/Needs-auth/Error. Previously unlisted videos were only shown as a count, so there was no way to jump to the specific videos.
-- **Cutting a release now clears `## [Unreleased]` entirely** instead of leaving an empty heading sitting above the new dated section. The `/changelog` and `/deploy` pages stop showing a blank Unreleased block after a release; the heading reappears the next time somebody hand-adds a pending bullet to the file.
-
-## [0.2.0] - 2026-05-22
-- **`/actionable` page** aggregating pending work across all channels — three sections (channels with undownloaded videos, channels with downloaded videos awaiting transcription, and channels with stale or missing reports), each row linking to the channel with inline "Download missing" / "Transcribe pending" / "Refresh report" buttons. An "Update all reports" button queues a refresh per channel. A "Needs attention" card on the dashboard links here whenever either of the first two lists has anything in it.
-- **Pipeline jobs regenerate the channel's report when they finish** (sync, download-from-playlist, download-missing, download-missing-subs, retry-bucket, store-playlist), so reports stay current without a manual refresh.
-- **`/changelog` page** rendering this changelog, linked from the sidebar nav. A subtle blue dot appears on the **Changelog** nav item when there's an entry you haven't viewed yet, clearing once the page is opened. Per-heading copy-link buttons give permalinks to any section, and a "Cut release" form at the bottom turns `## [Unreleased]` into a dated heading. The dashboard's "Recent changes" card shows the most recent entry, clamped to a fixed height with a gradient fade, and the whole card links to the full changelog.
-- **`/deploy` page** that previews `export/CHANGELOG.md` pending changes, offers a "Cut release" form for the export side, and a "Deploy" button that runs the export deploy with live log streaming.
-- **Composable layered search** in the export viewer — see `export/CHANGELOG.md` for the user-visible details. Old `?q=…&m=…&re=…` share-links keep working.
-- **Per-channel "Include in Sync all" toggle** on the `/channels` list. Excluded channels are skipped by **Sync all** with reason "excluded from sync all" in the skipped tooltip; their rows are dimmed in the table.
-- **Sortable column headers** on the `/channels` table — click any column to sort, click again to flip direction. Sort state doesn't persist across navigation.
-- **`/jobs/active` overhaul.** Each running batch job gets its own progress bar scaled to that job's own work, replacing the old channel-wide bar. Running jobs now appear above queued ones (both within each channel section and across channels), so live work stays at the top of the page. The sidebar **Jobs** badge still counts running + queued; the new **Active** badge counts only running.
-- **Deleted, private, and members-only videos no longer stick around as "awaiting transcription".** A failed download attempt that finds a video unavailable is now respected even if an older availability sweep had it as public. `/actionable` counts, the dashboard "Needs attention" summary, the channel page's Transcribe stage, and the Download stage's "dirs missing transcript & audio" list all exclude these videos. The video's row in the channel list still shows its true on-disk state.
-- **Channel page navigation fixes.** Clicking a pipeline stage in the side rail (or a mobile stage badge) no longer scrolls up into the selected video's viewer panel, sticky offsets now clear the header on every width, and clicking a stage when the "Channel pipeline & settings" wrapper is collapsed auto-expands it before scrolling into view.
-- **Audio download integrity fixes.** `.part` files with mid-stream codec corruption (hundreds of decoder errors that ffmpeg still treated as a clean exit) are now flagged as malformed instead of finalising as a finished download. A single download on a channel with many leftover `.part` files no longer stalls for minutes pre-checking unrelated videos, and no longer occasionally finalises the wrong file by latching onto a stale neighbour's `.part`.
-- **Editor's `/changelog` page no longer renders unstyled** — Tailwind now scans the shared changelog component.
+## [9.9.9] - 2024-01-01
+- old released bullet
diff --git a/editor/app/actionable/actions.ts b/editor/app/actionable/actions.ts
@@ -10,7 +10,12 @@ import {
import { getRegistry } from "yt-dlp-transcript-common/jobs/registry";
import { runManagedFunction } from "yt-dlp-transcript-common/jobs/streamCommand";
import { drainStream } from "yt-dlp-transcript-common/jobs/drainStream";
-import { detectDuplicateShorts } from "yt-dlp-transcript-common/controller/duplicateShorts";
+import {
+ detectDuplicateShorts,
+ readDuplicateReport,
+ updateDuplicateOverride,
+} from "yt-dlp-transcript-common/controller/duplicateShorts";
+import { shareClusterFromCanonical } from "yt-dlp-transcript-common/controller/digestSharing";
import {
clearIncompleteTranscriptsAction,
enableAutoRunners,
@@ -130,6 +135,69 @@ export async function runDuplicateDetectionAction(
return { ok: true, clusters, videosInClusters };
}
+// What a human can say about a cluster the detector could not decide.
+// confirmed — "I looked; these really are the same video." Unblocks
+// sharing AND publication for a title+duration suspect.
+// not-duplicate — "They are not." Suppresses the cluster entirely.
+// clear — undo, back to awaiting review.
+export type DuplicateClusterDecision = "confirmed" | "not-duplicate" | "clear";
+
+export type ReviewDuplicateClusterResult =
+ | { ok: true; shared: number; misaligned: number }
+ | { ok: false; error: string };
+
+// Record a review decision for one cluster.
+//
+// `updateDuplicateOverride` has existed — with the `confirmed` flag, the
+// read-modify-write, the atomic rename and the "an empty patch clears the
+// decision" rule — since duplicate review was built, and until now **nothing in
+// the repo called it**. `clusterMaySharePartial` fails closed, so every
+// needsReview cluster shared nothing and there was no way for a human to change
+// that. This is that missing caller.
+//
+// Confirming also attempts the share immediately, via the equally-uncalled
+// `shareClusterFromCanonical`: the point of confirming is to let derived work
+// flow, and making the operator wait for the canonical member's next sweep to
+// find out whether it would have would make the button feel inert. It is a
+// no-op when the canonical has no digest yet, which today is almost always.
+export async function reviewDuplicateClusterAction(
+ clusterId: string,
+ decision: DuplicateClusterDecision,
+): Promise<ReviewDuplicateClusterResult> {
+ const paths = getPaths();
+ try {
+ const overrides = await updateDuplicateOverride(
+ paths,
+ clusterId,
+ decision === "confirmed"
+ ? { confirmed: true, notDuplicate: false }
+ : decision === "not-duplicate"
+ ? { notDuplicate: true, confirmed: false }
+ : { confirmed: false, notDuplicate: false },
+ );
+
+ let shared = 0;
+ let misaligned = 0;
+ if (decision === "confirmed") {
+ const report = await readDuplicateReport(paths);
+ const cluster = report?.clusters.find((c) => c.clusterId === clusterId);
+ if (cluster) {
+ for (const outcome of await shareClusterFromCanonical(paths, cluster, {
+ overrides,
+ })) {
+ if (outcome.status === "shared") shared++;
+ else if (outcome.status === "misaligned") misaligned++;
+ }
+ }
+ }
+
+ revalidatePath("/actionable");
+ return { ok: true, shared, misaligned };
+ } catch (e) {
+ return { ok: false, error: (e as Error)?.message ?? String(e) };
+ }
+}
+
export type GlobalIncompleteResult =
| { ok: true; channels: number; affected: number }
| { ok: false; error: string };
diff --git a/editor/app/actionable/components/DuplicateClusterReview.tsx b/editor/app/actionable/components/DuplicateClusterReview.tsx
@@ -0,0 +1,122 @@
+"use client";
+
+import { useEffect, useState } from "react";
+import {
+ reviewDuplicateClusterAction,
+ type DuplicateClusterDecision,
+ type ReviewDuplicateClusterResult,
+} from "../actions";
+import type { DuplicateClusterOverride } from "yt-dlp-transcript-common/lib/duplicates";
+
+type Status =
+ | { kind: "idle" }
+ | { kind: "running"; decision: DuplicateClusterDecision }
+ | { kind: "done"; result: ReviewDuplicateClusterResult }
+ | { kind: "error"; message: string };
+
+// NOTE ON THE aria-labels BELOW: none of them may contain the substring
+// "duplicate cluster <id>". That is the CARD's own label, and Playwright's
+// getByLabel matches by SUBSTRING, so a button labelled "confirm duplicate
+// cluster <id>" makes the card's own locator resolve to three elements and
+// every existing assertion on it dies with a strict-mode violation. That is
+// exactly what happened here, and this repo has now hit the same collision
+// three times (deploy-page's "Build & deploy" vs "Build & deploy all sites",
+// the settings placeholder collision, and this).
+export function DuplicateClusterReview({
+ clusterId,
+ override,
+}: {
+ clusterId: string;
+ override: DuplicateClusterOverride | undefined;
+}) {
+ const [status, setStatus] = useState<Status>({ kind: "idle" });
+ // Disabled until mounted, deliberately. A server-rendered button has no
+ // handler until React hydrates, so a click before then fires NOTHING — no
+ // request, no error, nothing to debug. That is the exact failure that made
+ // the digest pilot's job look "un-created", and /actionable renders a page
+ // heavy enough to hit the same window. See StreamActionLog.
+ const [mounted, setMounted] = useState(false);
+ useEffect(() => setMounted(true), []);
+
+ async function decide(decision: DuplicateClusterDecision) {
+ setStatus({ kind: "running", decision });
+ try {
+ setStatus({
+ kind: "done",
+ result: await reviewDuplicateClusterAction(clusterId, decision),
+ });
+ } catch (e) {
+ setStatus({ kind: "error", message: (e as Error).message });
+ }
+ }
+
+ const busy = status.kind === "running";
+ const disabled = busy || !mounted;
+ const confirmed = override?.confirmed === true;
+ const rejected = override?.notDuplicate === true;
+
+ return (
+ <div className="flex items-center gap-2 flex-wrap">
+ <button
+ type="button"
+ onClick={() => decide("confirmed")}
+ disabled={disabled || confirmed}
+ aria-label={`confirm cluster ${clusterId}`}
+ className="px-2 py-1 rounded-md border border-border text-xs font-medium hover:bg-muted disabled:opacity-50"
+ >
+ {busy && status.decision === "confirmed" ? "Confirming…" : "Confirm"}
+ </button>
+ <button
+ type="button"
+ onClick={() => decide("not-duplicate")}
+ disabled={disabled || rejected}
+ aria-label={`reject cluster ${clusterId}`}
+ className="px-2 py-1 rounded-md border border-border text-xs font-medium hover:bg-muted disabled:opacity-50"
+ >
+ {busy && status.decision === "not-duplicate" ? "Rejecting…" : "Not a duplicate"}
+ </button>
+ {(confirmed || rejected) && (
+ <button
+ type="button"
+ onClick={() => decide("clear")}
+ disabled={disabled}
+ aria-label={`clear cluster decision ${clusterId}`}
+ className="px-2 py-1 rounded-md text-xs text-muted-foreground underline hover:text-foreground disabled:opacity-50"
+ >
+ {busy && status.decision === "clear" ? "Clearing…" : "Undo"}
+ </button>
+ )}
+ {status.kind === "done" && status.result.ok && (
+ <span
+ aria-label={`cluster review result ${clusterId}`}
+ className="text-xs text-muted-foreground"
+ >
+ {status.result.shared > 0 || status.result.misaligned > 0
+ ? `${status.result.shared} digest(s) shared` +
+ (status.result.misaligned > 0
+ ? `, ${status.result.misaligned} misaligned`
+ : "")
+ : "Saved."}
+ </span>
+ )}
+ {status.kind === "done" && !status.result.ok && (
+ <span
+ role="alert"
+ aria-label={`cluster review error ${clusterId}`}
+ className="text-xs text-destructive"
+ >
+ {status.result.error}
+ </span>
+ )}
+ {status.kind === "error" && (
+ <span
+ role="alert"
+ aria-label={`cluster review error ${clusterId}`}
+ className="text-xs text-destructive"
+ >
+ {status.message}
+ </span>
+ )}
+ </div>
+ );
+}
diff --git a/editor/app/actionable/lib/loadActionable.ts b/editor/app/actionable/lib/loadActionable.ts
@@ -8,8 +8,14 @@ import {
readChannelSnapshot,
type ChannelSnapshot,
} from "yt-dlp-transcript-common/controller/channelSnapshot";
-import { readDuplicateReport } from "yt-dlp-transcript-common/controller/duplicateShorts";
-import type { DuplicateReport } from "yt-dlp-transcript-common/lib/duplicates";
+import {
+ readDuplicateOverrides,
+ readDuplicateReport,
+} from "yt-dlp-transcript-common/controller/duplicateShorts";
+import type {
+ DuplicateOverrides,
+ DuplicateReport,
+} from "yt-dlp-transcript-common/lib/duplicates";
export type ActionableRow = {
channel: ChannelStat;
@@ -25,7 +31,14 @@ export type ActionableSummary = {
cleanTranscribedAudio: ActionableRow[];
cleanExtraFormats: ActionableRow[];
staleOrMissing: ActionableRow[];
+ digestWarnings: ActionableRow[];
duplicates: DuplicateReport | null;
+ // The human decisions kept alongside the report — a cluster's canonical
+ // choice, "not a duplicate", and the `confirmed` flag that is the only thing
+ // letting a needsReview cluster share derived work. Loaded here because the
+ // review UI cannot show what has already been decided without it, and a
+ // review queue that forgets its own answers re-asks every question.
+ duplicateOverrides: DuplicateOverrides;
};
export function isStaleOrMissing(row: ActionableRow): boolean {
@@ -86,6 +99,18 @@ export function actionableCleanExtraFormatsCount(row: ActionableRow): number {
return row.snapshot?.buckets.multipleAudioFormats?.length ?? 0;
}
+// The digest layer's two work lists. `noDigest` is "has no digest at the
+// current identity" — the backfill's denominator. `digestWarnings` is "the
+// model produced something a human should look at", which includes the total
+// failures that write no section and so are invisible to any count of files.
+export function actionableNoDigestCount(row: ActionableRow): number {
+ return row.snapshot?.buckets.noDigest?.length ?? 0;
+}
+
+export function actionableDigestWarningsCount(row: ActionableRow): number {
+ return row.snapshot?.buckets.digestWarnings?.length ?? 0;
+}
+
// Estimated bytes each cleanup would reclaim (default 0 for snapshots written
// before cleanupBytes existed).
export function actionableCleanTranscribedBytes(row: ActionableRow): number {
@@ -100,7 +125,7 @@ export async function loadActionableSummary(
paths: Paths,
): Promise<ActionableSummary> {
const channels = await listChannels(paths);
- const [rows, duplicates] = await Promise.all([
+ const [rows, duplicates, duplicateOverrides] = await Promise.all([
Promise.all(
channels.map(async (channel) => ({
channel,
@@ -108,6 +133,7 @@ export async function loadActionableSummary(
})),
),
readDuplicateReport(paths),
+ readDuplicateOverrides(paths),
]);
const undownloaded = rows
@@ -149,6 +175,13 @@ export async function loadActionableSummary(
actionableCleanExtraFormatsCount(a),
);
+ const digestWarnings = rows
+ .filter((r) => actionableDigestWarningsCount(r) > 0)
+ .sort(
+ (a, b) =>
+ actionableDigestWarningsCount(b) - actionableDigestWarningsCount(a),
+ );
+
const staleOrMissing = rows
.filter(isStaleOrMissing)
.sort((a, b) => a.channel.slug.localeCompare(b.channel.slug));
@@ -162,6 +195,8 @@ export async function loadActionableSummary(
cleanTranscribedAudio,
cleanExtraFormats,
staleOrMissing,
+ digestWarnings,
duplicates,
+ duplicateOverrides,
};
}
diff --git a/editor/app/actionable/page.tsx b/editor/app/actionable/page.tsx
@@ -7,6 +7,7 @@ import {
actionableCleanExtraFormatsCount,
actionableCleanTranscribedBytes,
actionableCleanTranscribedCount,
+ actionableDigestWarningsCount,
actionableIncompleteTranscriptCount,
actionableShortAudioCount,
actionableUndownloadedCount,
@@ -18,8 +19,10 @@ import { InlineActionButton } from "./components/InlineActionButton";
import { FixAllIncompleteButton } from "./components/FixAllIncompleteButton";
import { RefreshAllReportsButton } from "./components/RefreshAllReportsButton";
import { RunDuplicateDetectionButton } from "./components/RunDuplicateDetectionButton";
+import { DuplicateClusterReview } from "./components/DuplicateClusterReview";
import type {
DuplicateCluster,
+ DuplicateOverrides,
DuplicateReport,
} from "yt-dlp-transcript-common/lib/duplicates";
@@ -55,7 +58,8 @@ export default async function ActionablePage() {
summary.shortAudio.length === 0 &&
summary.cleanTranscribedAudio.length === 0 &&
summary.cleanExtraFormats.length === 0 &&
- summary.staleOrMissing.length === 0;
+ summary.staleOrMissing.length === 0 &&
+ summary.digestWarnings.length === 0;
const sections: { config: SectionConfig; rows: ActionableRow[] }[] = [
{
@@ -128,6 +132,27 @@ export default async function ActionablePage() {
},
{
config: {
+ id: "digest-warnings",
+ title: "Channels with digest passes that need a look",
+ description:
+ "Videos where the AI digest pass recorded warnings, or produced nothing usable at all. The second kind is the one worth opening: a total failure deliberately writes no section so the video retries, and “the model proposed nothing” and “the model proposed chapters and every one was rejected by a guard” look identical from outside but need different fixes.",
+ countLabel: "with warnings",
+ emptyLabel: "None recorded.",
+ getCount: actionableDigestWarningsCount,
+ primaryAction: (r) => (
+ <Link
+ href={`/channels/${r.channel.slug}?filter=digest_warnings`}
+ aria-label={`review digest warnings for ${r.channel.slug}`}
+ className="inline-flex items-center px-2.5 py-1 rounded-md border border-border text-xs font-medium hover:bg-muted whitespace-nowrap"
+ >
+ Review
+ </Link>
+ ),
+ },
+ rows: summary.digestWarnings,
+ },
+ {
+ config: {
id: "short-audio",
title: "Channels with truncated downloads (short audio)",
description:
@@ -232,12 +257,21 @@ export default async function ActionablePage() {
<Section key={config.id} config={config} rows={rows} />
))
)}
- <DuplicatesSection report={summary.duplicates} />
+ <DuplicatesSection
+ report={summary.duplicates}
+ overrides={summary.duplicateOverrides}
+ />
</div>
);
}
-function DuplicatesSection({ report }: { report: DuplicateReport | null }) {
+function DuplicatesSection({
+ report,
+ overrides,
+}: {
+ report: DuplicateReport | null;
+ overrides: DuplicateOverrides;
+}) {
const clusters = report?.clusters ?? [];
return (
<section aria-label="duplicate-shorts" className="flex flex-col gap-2">
@@ -275,7 +309,11 @@ function DuplicatesSection({ report }: { report: DuplicateReport | null }) {
) : (
<ul className="flex flex-col gap-3">
{clusters.map((cluster) => (
- <DuplicateClusterCard key={cluster.clusterId} cluster={cluster} />
+ <DuplicateClusterCard
+ key={cluster.clusterId}
+ cluster={cluster}
+ override={overrides.clusters[cluster.clusterId]}
+ />
))}
</ul>
)}
@@ -283,10 +321,17 @@ function DuplicatesSection({ report }: { report: DuplicateReport | null }) {
);
}
-function DuplicateClusterCard({ cluster }: { cluster: DuplicateCluster }) {
+function DuplicateClusterCard({
+ cluster,
+ override,
+}: {
+ cluster: DuplicateCluster;
+ override: DuplicateOverrides["clusters"][string] | undefined;
+}) {
const matchLabel: Record<DuplicateCluster["matchKind"], string> = {
"transcript-exact": "exact transcript",
"transcript-near": "near transcript",
+ "title-duration": "title + runtime",
};
return (
<li
@@ -295,6 +340,14 @@ function DuplicateClusterCard({ cluster }: { cluster: DuplicateCluster }) {
>
<div className="flex items-center gap-2 flex-wrap text-xs">
<Badge>{matchLabel[cluster.matchKind]}</Badge>
+ {/* Nothing compared these videos' content — one side has no transcript.
+ The cluster is a suspect for a human, stays out of the built site,
+ and shares no derived work until someone confirms it. */}
+ {cluster.needsReview && !override?.confirmed && !override?.notDuplicate && (
+ <Badge>needs review</Badge>
+ )}
+ {override?.confirmed && <Badge>confirmed</Badge>}
+ {override?.notDuplicate && <Badge>not a duplicate</Badge>}
{cluster.score !== null && (
<Badge>score {cluster.score.toFixed(2)}</Badge>
)}
@@ -304,6 +357,17 @@ function DuplicateClusterCard({ cluster }: { cluster: DuplicateCluster }) {
<span className="text-muted-foreground">
{cluster.videoRefs.length} videos · ~{cluster.durationBucket}s
</span>
+ {/* The confirm path exists precisely for `needsReview` clusters:
+ clusterMaySharePartial fails closed, so until a human says
+ "confirmed" these share no digest and never reach a built site.
+ Content-confirmed clusters need no confirmation — but they can still
+ be rejected, which is the only way to un-assert a wrong one. */}
+ <span className="ml-auto">
+ <DuplicateClusterReview
+ clusterId={cluster.clusterId}
+ override={override}
+ />
+ </span>
</div>
<ul className="flex flex-col gap-1">
{cluster.videoRefs.map((ref) => (
diff --git a/editor/app/api/widget/actionable/route.ts b/editor/app/api/widget/actionable/route.ts
@@ -1,6 +1,7 @@
import { NextResponse } from "next/server";
import { getPaths } from "yt-dlp-transcript-common/lib/paths";
import {
+ actionableNoDigestCount,
actionableUndownloadedCount,
actionableUntranscribedCount,
loadActionableSummary,
@@ -12,6 +13,12 @@ export type WidgetActionableChannel = {
slug: string;
undownloaded: number;
untranscribed: number;
+ // Videos with no digest at the CURRENT identity — the backfill's per-channel
+ // work list. Reported but NOT used to decide whether a channel "needs work":
+ // during the backfill that is 99.87% of the corpus, so counting it would put
+ // every channel in the list forever and drown the two buckets a human can
+ // actually act on today.
+ noDigest: number;
};
export type WidgetActionablePayload = {
@@ -29,6 +36,7 @@ export async function GET() {
slug: row.channel.slug,
undownloaded: actionableUndownloadedCount(row),
untranscribed: actionableUntranscribedCount(row),
+ noDigest: actionableNoDigestCount(row),
}))
.filter((c) => c.undownloaded > 0 || c.untranscribed > 0)
.sort(
diff --git a/editor/app/api/widget/sync/route.ts b/editor/app/api/widget/sync/route.ts
@@ -23,6 +23,20 @@ export type WidgetSyncPayload = {
nextRunAt: number | null; // soonest eligible channel's nextDueAt
overdue: boolean; // some eligible channel is due now
};
+ // Corpus-wide digest coverage. SCALARS ONLY, honouring this payload's stated
+ // constraint — three numbers and two booleans, not a per-channel breakdown.
+ //
+ // It belongs on the poll rather than a page load because the number it
+ // reports moves over WEEKS: an operator watching an 80-day backfill needs to
+ // see it move at all, and "0.13% of 77,207" is not a figure any single
+ // channel page can show.
+ digest: {
+ digested: number; // videos carrying a non-empty ai-digest.json
+ videos: number; // videos in the corpus (the denominator)
+ channelsWithAny: number; // channels the layer has reached at all
+ paused: boolean; // settings.digest.digestsPaused
+ sweeping: boolean; // a corpus-wide sweep is armed
+ };
};
// Build the widget sync payload. Exported so the dashboard cockpit can seed its
@@ -61,6 +75,18 @@ export async function buildWidgetSyncPayload(): Promise<WidgetSyncPayload> {
}
}
+ // Free: listChannels already counted these, it just used to throw the digest
+ // half away (see controller/channels.ts).
+ let digested = 0;
+ let videos = 0;
+ let channelsWithAny = 0;
+ for (const c of channels) {
+ videos += c.videoCount;
+ const n = c.digestCount ?? 0;
+ digested += n;
+ if (n > 0) channelsWithAny++;
+ }
+
return {
lastSyncAllAt: state.lastSyncAllAt,
lastIndividualSyncAt,
@@ -70,6 +96,13 @@ export async function buildWidgetSyncPayload(): Promise<WidgetSyncPayload> {
nextRunAt,
overdue,
},
+ digest: {
+ digested,
+ videos,
+ channelsWithAny,
+ paused: settings.digest.digestsPaused,
+ sweeping: settings.digest.sweepEnabled,
+ },
} satisfies WidgetSyncPayload;
}
diff --git a/editor/app/channels/[slug]/components/StatusHeader.tsx b/editor/app/channels/[slug]/components/StatusHeader.tsx
@@ -8,6 +8,7 @@ const BADGE_ORDER: StageId[] = [
"download",
"transcode",
"transcribe",
+ "digest",
"cleanup",
"diagnostics",
];
diff --git a/editor/app/channels/[slug]/components/VideoListPane.tsx b/editor/app/channels/[slug]/components/VideoListPane.tsx
@@ -68,6 +68,7 @@ const FILTER_OPTIONS: { value: VideoFilter; label: string }[] = [
{ value: "downloaded_no_transcript", label: "No transcript" },
{ value: "partial", label: "Partial" },
{ value: "incomplete_transcript", label: "Incomplete transcript" },
+ { value: "digest_warnings", label: "Digest warnings" },
{ value: "short_audio", label: "Short audio" },
{ value: "untranscribable", label: "Untranscribable" },
{ value: "running", label: "Running" },
diff --git a/editor/app/channels/[slug]/components/stages/DigestStage.tsx b/editor/app/channels/[slug]/components/stages/DigestStage.tsx
@@ -0,0 +1,165 @@
+"use client";
+
+// The minimum surface needed to RUN a digest sweep on a channel: a count, a lane
+// picker, and a button. Deliberately not a review UI — the per-video review
+// panel, the /actionable cluster section and the settings form are a separate
+// piece of work. What this exists for is that the sweep has to be startable from
+// a browser at all, both for an operator and for the e2e suite that proves the
+// job path end to end.
+//
+// Modelled on TranscribeStage: same StreamActionLog + QueueControl shape, same
+// bucket-count heading, so the stage rail reads as one system.
+
+import { useState } from "react";
+import { StreamActionLog } from "yt-dlp-transcript-common/components/StreamActionLog";
+import { QueueControl } from "../../../../components/QueueControl";
+import { cancelJobAction } from "../../../../jobs/actions";
+import {
+ digestChannelAction,
+ type DigestLaneChoice,
+} from "../../digestActions";
+import { VideoIdList } from "../VideoIdList";
+
+type Props = {
+ slug: string;
+ existingQueues: string[];
+ // Videos with a transcript but no digest at the current identity.
+ noDigestIds: string[];
+ // The two lanes' default queue keys. They are DIFFERENT on purpose: the local
+ // lane is GPU-bound and the metered lane is network-bound, and the registry
+ // hardcodes concurrency 1 per key, so a shared key would serialize them.
+ localQueueKey: string;
+ remoteQueueKey: string;
+ // settings.digest.remoteEnabled. False disables the metered option outright
+ // rather than letting the action reject it after the fact.
+ remoteEnabled: boolean;
+ // Longest-first is worth offering because the >4h tail is 8.2% of videos but
+ // 46% of the tokens — the half where a prompt problem is most expensive to
+ // discover late.
+ defaultOrder?: DigestOrder;
+};
+
+type DigestOrder = "shortest-first" | "longest-first";
+
+export function DigestStage({
+ slug,
+ existingQueues,
+ noDigestIds,
+ localQueueKey,
+ remoteQueueKey,
+ remoteEnabled,
+ defaultOrder = "shortest-first",
+}: Props) {
+ const [lane, setLane] = useState<DigestLaneChoice>("local");
+ const [queue, setQueue] = useState(localQueueKey);
+ const [order, setOrder] = useState<DigestOrder>(defaultOrder);
+ const [limit, setLimit] = useState("");
+
+ const onLaneChange = (next: DigestLaneChoice): void => {
+ setLane(next);
+ // Follow the lane's own default key unless the operator has typed a custom
+ // one — picking "metered" and silently keeping the local key would put both
+ // lanes on one key and serialize them, which is the exact failure the two
+ // keys exist to avoid.
+ setQueue((current) =>
+ current === localQueueKey || current === remoteQueueKey
+ ? next === "remote"
+ ? remoteQueueKey
+ : localQueueKey
+ : current,
+ );
+ };
+
+ const parsedLimit = Number.parseInt(limit, 10);
+ const limitCount =
+ Number.isFinite(parsedLimit) && parsedLimit > 0 ? parsedLimit : undefined;
+
+ return (
+ <div
+ aria-label="digest section"
+ className="flex flex-col gap-3"
+ >
+ <div>
+ <h3 className="text-base font-semibold">
+ Generate digests ({noDigestIds.length})
+ </h3>
+ <p className="text-sm text-muted-foreground">
+ Chapters (and optionally topic tags) derived from each transcript by a
+ local model, written to <code>ai-digest.json</code> next to it. A
+ re-run regenerates only what the current model and prompt version have
+ not already produced, so running this twice costs nothing the second
+ time. Hand corrections live in a separate{" "}
+ <code>ai-digest.overrides.json</code> and are never overwritten.
+ </p>
+ </div>
+
+ <VideoIdList
+ slug={slug}
+ ids={noDigestIds}
+ ariaLabel="videos without a digest list"
+ emptyAriaLabel="videos without a digest empty"
+ emptyMessage="Every transcript has a current digest."
+ itemAriaLabel={(id) => `video without a digest ${id}`}
+ />
+
+ <StreamActionLog
+ trigger={() =>
+ digestChannelAction(slug, lane, queue, order, limitCount)
+ }
+ cancelAction={cancelJobAction}
+ buttonLabel="Digest channel"
+ runningLabel="Digesting…"
+ label="Digest channel"
+ extraControls={
+ <>
+ <label className="flex items-center gap-1 text-xs text-muted-foreground">
+ Lane
+ <select
+ value={lane}
+ onChange={(e) => onLaneChange(e.target.value as DigestLaneChoice)}
+ aria-label="lane for Digest channel"
+ className="rounded border border-border bg-card px-1 py-0.5"
+ >
+ <option value="local">local (ollama)</option>
+ <option value="remote" disabled={!remoteEnabled}>
+ metered{remoteEnabled ? "" : " — disabled in Settings"}
+ </option>
+ </select>
+ </label>
+ <label className="flex items-center gap-1 text-xs text-muted-foreground">
+ Order
+ <select
+ value={order}
+ onChange={(e) => setOrder(e.target.value as DigestOrder)}
+ aria-label="order for Digest channel"
+ className="rounded border border-border bg-card px-1 py-0.5"
+ >
+ <option value="shortest-first">shortest first</option>
+ <option value="longest-first">longest first</option>
+ </select>
+ </label>
+ <label className="flex items-center gap-1 text-xs text-muted-foreground">
+ Limit
+ <input
+ type="text"
+ inputMode="numeric"
+ value={limit}
+ onChange={(e) => setLimit(e.target.value)}
+ placeholder="all"
+ aria-label="limit for Digest channel"
+ className="w-16 rounded border border-border bg-card px-1 py-0.5"
+ />
+ </label>
+ <QueueControl
+ value={queue}
+ onChange={setQueue}
+ defaultQueueKey={lane === "remote" ? remoteQueueKey : localQueueKey}
+ existingQueues={existingQueues}
+ actionLabel="Digest channel"
+ />
+ </>
+ }
+ />
+ </div>
+ );
+}
diff --git a/editor/app/channels/[slug]/digestActions.ts b/editor/app/channels/[slug]/digestActions.ts
@@ -0,0 +1,247 @@
+"use server";
+
+import { revalidatePath } from "next/cache";
+import { getPaths } from "yt-dlp-transcript-common/lib/paths";
+import { getSettings } from "yt-dlp-transcript-common/lib/settings";
+import {
+ DIGEST_LOCAL_QUEUE,
+ DIGEST_REMOTE_QUEUE,
+ resolveQueueKey,
+} from "yt-dlp-transcript-common/lib/queueKeys";
+import {
+ countMissingDigests,
+ runDigestBatch,
+ type DigestOrder,
+} from "yt-dlp-transcript-common/controller/digestBatch";
+import {
+ runManagedFunction,
+ type StreamActionResult,
+} from "yt-dlp-transcript-common/jobs/streamCommand";
+import { makeTaskTracker } from "yt-dlp-transcript-common/jobs/taskHooks";
+import { requestChannelSnapshot } from "yt-dlp-transcript-common/jobs/snapshotScheduler";
+import {
+ loadDigest,
+ loadDigestOverrides,
+ writeDigestOverrides,
+} from "yt-dlp-transcript-common/lib/digest-server";
+import {
+ effectiveDigest,
+ type DigestChapter,
+ type DigestOverrides,
+ type DigestTag,
+ type EffectiveDigest,
+} from "yt-dlp-transcript-common/lib/digest";
+import path from "node:path";
+
+function isDigestOrder(v: unknown): v is DigestOrder {
+ return v === "shortest-first" || v === "longest-first";
+}
+
+export type DigestLaneChoice = "local" | "remote";
+
+// Run the digest sweep over one channel. ONE JOB PER CHANNEL, deliberately:
+// job logs keep only the newest 500 (30 days) and the in-memory registry keeps
+// 100 records, so a job per video would evict the whole history of a sweep —
+// including the running jobs' own logs.
+export async function digestChannelAction(
+ slug: string,
+ lane: DigestLaneChoice = "local",
+ queueKey?: string,
+ order?: string,
+ limitCount?: number,
+ force?: boolean,
+): Promise<StreamActionResult> {
+ const paths = getPaths();
+ const settings = getSettings();
+ if (lane === "remote" && !settings.digest.remoteEnabled) {
+ return {
+ ok: false,
+ error:
+ "The metered digest lane is off. Enable it in Settings → Digest before running it.",
+ };
+ }
+ const remote = lane === "remote";
+ const kind = remote ? "digest-channel-remote" : "digest-channel-local";
+ return runManagedFunction({
+ kind,
+ // Separate keys per lane so the two run CONCURRENTLY — the local lane is
+ // GPU-bound and the metered lane is network-bound, so serializing them would
+ // waste half the throughput of a multi-week sweep.
+ queueKey: resolveQueueKey(
+ remote ? DIGEST_REMOTE_QUEUE : DIGEST_LOCAL_QUEUE,
+ queueKey,
+ ),
+ paths,
+ channelSlug: slug,
+ spec: {
+ kind,
+ slug,
+ params: { queueKey, lane, order, limitCount, force },
+ },
+ fn: async (onLog, signal, setProgress, ctx) => {
+ // The LANE matters: the progress target must be computed against the
+ // engine that is actually about to run, not whichever one the local
+ // setting names.
+ const missing = await countMissingDigests(paths, slug, undefined, lane);
+ // The bar measures THIS RUN, from zero, and the batch reports its own
+ // `current`. It used to be seeded with the count of digests already on
+ // disk and left to the UI's disk re-count — which cannot move during a
+ // REGENERATION, because a regenerated digest is rewritten in place. The
+ // bar sat at 0% for whole jobs ({initial:39, current:39, target:72}).
+ const result = await runDigestBatch({
+ channelSlug: slug,
+ paths,
+ lane,
+ order: isDigestOrder(order) ? order : undefined,
+ // The local lane takes everything up to the long-tail cutoff; the metered
+ // lane exists for the tail above it. Passing no window (the default) runs
+ // the whole channel on one lane, which is what a single-lane sweep wants.
+ ...(remote
+ ? { minDurationSeconds: settings.digest.longTailSeconds }
+ : {}),
+ limitCount:
+ typeof limitCount === "number" && limitCount > 0 ? limitCount : undefined,
+ force: force === true,
+ setProgress,
+ progressBaseline: 0,
+ progressTarget: missing,
+ onLog,
+ signal,
+ drainSignal: ctx.drainSignal,
+ tracker: makeTaskTracker(ctx, onLog),
+ });
+ onLog(
+ `Digest batch: ${result.succeeded} generated, ${result.fresh} already current, ` +
+ `${result.shared} shared to mirrors, ${result.misaligned} mirror(s) refused by the alignment gate, ` +
+ `${result.skipped} skipped, ${result.failed} failed; ` +
+ `${result.engineCalls} model call(s), ${result.warnings} warning(s)` +
+ (result.costUsd > 0 ? `, $${result.costUsd.toFixed(4)}` : "") +
+ (result.spendCapped ? " (stopped at the spend cap)" : "") +
+ ".",
+ );
+ // ONCE, at job end — the digest kinds are in NO_REGEN_KINDS precisely so
+ // that per-video regens don't fire. See snapshotScheduler.ts.
+ requestChannelSnapshot(paths, slug);
+ revalidatePath(`/channels/${slug}`);
+ },
+ });
+}
+
+// Digest an explicit id set (a bucket selection, or the pilot's single channel
+// slice). Shares the batch machinery; scoped ids are intersected with disk.
+export async function digestBucketAction(
+ slug: string,
+ ids: string[],
+ lane: DigestLaneChoice = "local",
+ queueKey?: string,
+): Promise<StreamActionResult> {
+ const paths = getPaths();
+ const settings = getSettings();
+ const cleaned = Array.from(new Set(ids.map((id) => id.trim()).filter(Boolean)));
+ if (cleaned.length === 0) {
+ return { ok: false, error: "No video ids supplied" };
+ }
+ if (lane === "remote" && !settings.digest.remoteEnabled) {
+ return { ok: false, error: "The metered digest lane is off." };
+ }
+ const remote = lane === "remote";
+ const kind = remote ? "digest-channel-remote" : "digest-channel-local";
+ return runManagedFunction({
+ kind,
+ queueKey: resolveQueueKey(
+ remote ? DIGEST_REMOTE_QUEUE : DIGEST_LOCAL_QUEUE,
+ queueKey,
+ ),
+ paths,
+ channelSlug: slug,
+ fn: async (onLog, signal, setProgress, ctx) => {
+ // Same shape as the channel action: the bar measures this run from zero
+ // and the batch reports its own `current`, because the disk re-count it
+ // would otherwise use cannot see a digest rewritten in place.
+ const missing = await countMissingDigests(paths, slug, cleaned, lane);
+ const result = await runDigestBatch({
+ channelSlug: slug,
+ paths,
+ lane,
+ ids: cleaned,
+ setProgress,
+ progressBaseline: 0,
+ progressTarget: missing,
+ onLog,
+ signal,
+ drainSignal: ctx.drainSignal,
+ tracker: makeTaskTracker(ctx, onLog),
+ });
+ onLog(
+ `Digest bucket: ${result.succeeded} generated, ${result.fresh} already current, ${result.failed} failed.`,
+ );
+ requestChannelSnapshot(paths, slug);
+ revalidatePath(`/channels/${slug}`);
+ },
+ });
+}
+
+// ---------------------------------------------------------------------------
+// Per-video review: read the composed digest, write human corrections
+// ---------------------------------------------------------------------------
+
+export type VideoDigestView = {
+ digest: EffectiveDigest;
+ hasMachineDigest: boolean;
+ overrides: DigestOverrides | null;
+};
+
+function videoDir(slug: string, id: string): string {
+ return path.join(getPaths().channelsDir, slug, "data", id);
+}
+
+export async function readVideoDigestAction(
+ slug: string,
+ id: string,
+): Promise<VideoDigestView> {
+ const dir = videoDir(slug, id);
+ const [machine, overrides] = await Promise.all([
+ loadDigest(dir),
+ loadDigestOverrides(dir),
+ ]);
+ return {
+ digest: effectiveDigest(machine, overrides),
+ hasMachineDigest: machine !== null,
+ overrides,
+ };
+}
+
+export type SaveDigestOverridesResult =
+ | { ok: true; digest: EffectiveDigest }
+ | { ok: false; error: string };
+
+// Write human corrections. This ONLY ever touches ai-digest.overrides.json —
+// never ai-digest.json — which is what makes a correction survive the next
+// regeneration of a corpus too large to re-generate twice.
+export async function saveDigestOverridesAction(
+ slug: string,
+ id: string,
+ input: {
+ chapters?: DigestChapter[];
+ tags?: DigestTag[];
+ note?: string;
+ },
+): Promise<SaveDigestOverridesResult> {
+ const dir = videoDir(slug, id);
+ try {
+ await writeDigestOverrides(dir, {
+ version: 1,
+ ...(input.chapters ? { chapters: input.chapters } : {}),
+ ...(input.tags ? { tags: input.tags } : {}),
+ ...(input.note ? { note: input.note } : {}),
+ });
+ } catch (e) {
+ return { ok: false, error: (e as Error).message };
+ }
+ const [machine, overrides] = await Promise.all([
+ loadDigest(dir),
+ loadDigestOverrides(dir),
+ ]);
+ revalidatePath(`/channels/${slug}/videos/${id}`);
+ return { ok: true, digest: effectiveDigest(machine, overrides) };
+}
diff --git a/editor/app/channels/[slug]/lib/stageStatus.ts b/editor/app/channels/[slug]/lib/stageStatus.ts
@@ -34,6 +34,7 @@ export function normalizeBuckets(
downloadedAutoSubsOnly: raw?.downloadedAutoSubsOnly ?? [],
supersededAutoSubs: raw?.supersededAutoSubs ?? [],
needsCookies: raw?.needsCookies ?? [],
+ noDigest: raw?.noDigest ?? [],
};
}
@@ -43,6 +44,7 @@ export type StageId =
| "download"
| "transcode"
| "transcribe"
+ | "digest"
| "cleanup"
| "diagnostics"
| "danger";
@@ -75,6 +77,9 @@ const JOB_KIND_TO_STAGE: Record<string, StageId> = {
"transcribe-one": "transcribe",
"whisper-video": "transcribe",
"whisper-bucket-auto-subs": "transcribe",
+ "digest-channel-local": "digest",
+ "digest-channel-remote": "digest",
+ "digest-share-cluster": "digest",
"clean-audio-transcribed": "cleanup",
"purge-superseded-auto-subs": "cleanup",
"clean-extra-audio-formats": "cleanup",
@@ -327,6 +332,38 @@ export function computeStageStatuses(
}),
};
+ // Videos with a transcript whose digest is missing OR stale against the local
+ // lane's current identity — the snapshot's noDigest bucket now computes that
+ // full freshness target (channelSnapshot.ts), so this label and that bucket
+ // finally mean the same thing. Counted as pending work rather than merely
+ // informational: unlike the auto-captions lane, every transcribed video is
+ // eventually meant to have one.
+ const digestRunning = runningByStage.has("digest");
+ const digestPending = buckets.noDigest.length;
+ const digest: StageStatus = {
+ id: "digest",
+ title: "Digest",
+ pending: digestPending,
+ failed: 0,
+ running: digestRunning,
+ defaultOpen: true,
+ summary: digestRunning
+ ? "Running…"
+ : digestPending > 0
+ ? pluralize(
+ digestPending,
+ "transcript needs a digest",
+ "transcripts need a digest",
+ )
+ : "All digested at the current settings.",
+ tone: pickTone({
+ running: digestRunning,
+ pending: digestPending,
+ failed: 0,
+ fallback: "ok",
+ }),
+ };
+
const cleanupRunning = runningByStage.has("cleanup");
const cleanupParts: string[] = [];
if (cleanupPending > 0) {
@@ -407,6 +444,7 @@ export function computeStageStatuses(
download,
transcode,
transcribe,
+ digest,
cleanup,
diagnostics,
danger,
diff --git a/editor/app/channels/[slug]/lib/videoRows.ts b/editor/app/channels/[slug]/lib/videoRows.ts
@@ -39,6 +39,11 @@ export type VideoRow = {
// duration — the audio download truncated silently. Independent flag (not a
// `status`) so it composes with `transcribed`. See transcriptCoverage.
incompleteTranscript: boolean;
+ // The digest pass recorded warnings, or failed outright and wrote no section
+ // at all. Independent flag (it composes with `transcribed`, and a total
+ // failure is also in noDigest) so the review queue can ask "what did the model
+ // do badly here?" separately from "what has no digest yet?".
+ digestWarnings: boolean;
// Download completed but the audio was far shorter than the video — the source
// served a truncated stream (download-outcome "failed-short-audio"). The stub
// is kept on disk; re-download with a different format to fix. Independent flag
@@ -65,6 +70,7 @@ export type VideoFilter =
| "incomplete_transcript"
| "short_audio"
| "untranscribable"
+ | "digest_warnings"
| "running";
const VIDEO_FILTERS: readonly VideoFilter[] = [
@@ -77,6 +83,7 @@ const VIDEO_FILTERS: readonly VideoFilter[] = [
"incomplete_transcript",
"short_audio",
"untranscribable",
+ "digest_warnings",
"running",
];
@@ -111,6 +118,8 @@ function matchesFilter(r: VideoRow, filter: VideoFilter): boolean {
return r.shortAudio;
case "untranscribable":
return r.untranscribable;
+ case "digest_warnings":
+ return r.digestWarnings;
case "running":
return r.running;
}
diff --git a/editor/app/channels/[slug]/lib/videoRowsServer.ts b/editor/app/channels/[slug]/lib/videoRowsServer.ts
@@ -42,6 +42,7 @@ export function computeVideoRows(input: ComputeRowsInput): VideoRow[] {
const partial = new Set(buckets.partialDownloads);
const incompleteTranscript = new Set(buckets.incompleteTranscript);
const shortAudioSet = new Set(buckets.shortAudio ?? []);
+ const digestWarningsSet = new Set(buckets.digestWarnings ?? []);
const corruptSourceSet = new Set(buckets.corruptSource);
const corruptFullSourceSet = new Set(buckets.corruptFullSource);
const failedTranscription = new Set(input.failedTranscriptionIds);
@@ -113,6 +114,7 @@ export function computeVideoRows(input: ComputeRowsInput): VideoRow[] {
excluded,
incompleteTranscript: incompleteTranscript.has(id),
shortAudio: shortAudioSet.has(id),
+ digestWarnings: digestWarningsSet.has(id),
running: runningIds.has(id),
status,
});
diff --git a/editor/app/channels/[slug]/page.tsx b/editor/app/channels/[slug]/page.tsx
@@ -45,6 +45,10 @@ import {
queueKeyForUrl,
TRANSCRIPTION_QUEUE,
} from "yt-dlp-transcript-common/lib/platform";
+import {
+ DIGEST_LOCAL_QUEUE,
+ DIGEST_REMOTE_QUEUE,
+} from "yt-dlp-transcript-common/lib/queueKeys";
import { getRegistry } from "yt-dlp-transcript-common/jobs/registry";
import { listSites } from "yt-dlp-transcript-common/lib/site";
import { sortGroups } from "yt-dlp-transcript-common/lib/channelGroups";
@@ -64,6 +68,7 @@ import { DownloadStage } from "./components/stages/DownloadStage";
import { PlaylistStage } from "./components/stages/PlaylistStage";
import { TranscodeStage } from "./components/stages/TranscodeStage";
import { TranscribeStage } from "./components/stages/TranscribeStage";
+import { DigestStage } from "./components/stages/DigestStage";
import { VideoListPane } from "./components/VideoListPane";
import { VideoListPaneSection } from "./components/VideoListPaneSection";
import { VideoPanel, type VideoFile } from "./videos/[id]/components/VideoPanel";
@@ -223,7 +228,8 @@ export default async function ChannelDetailPage({
// Saved-video store summary for this channel + whether backups are configured,
// for the Cleanup stage's Retention & persistence section (Phase 5).
const savedTotals = await savedVideoTotals({ paths, channelSlug: slug });
- const backupConfigured = getSettings().savedVideoBackup.dest.trim() !== "";
+ const settings = getSettings();
+ const backupConfigured = settings.savedVideoBackup.dest.trim() !== "";
// The retry-failures panel is interactive (a click mutates the file), so
// read it fresh on every render — the snapshot bucket only reflects state
// at refresh time.
@@ -282,6 +288,7 @@ export default async function ChannelDetailPage({
"download",
...(transcodeApplies ? (["transcode"] as const) : []),
"transcribe",
+ "digest",
"cleanup",
"diagnostics",
"danger",
@@ -356,6 +363,16 @@ export default async function ChannelDetailPage({
missingShard={transcribeMissingShard}
/>
),
+ digest: (
+ <DigestStage
+ slug={slug}
+ existingQueues={existingQueues}
+ noDigestIds={buckets.noDigest}
+ localQueueKey={DIGEST_LOCAL_QUEUE}
+ remoteQueueKey={DIGEST_REMOTE_QUEUE}
+ remoteEnabled={settings.digest.remoteEnabled}
+ />
+ ),
cleanup: (
<CleanupStage
slug={slug}
diff --git a/editor/app/channels/[slug]/videos/[id]/components/DigestPanel.tsx b/editor/app/channels/[slug]/videos/[id]/components/DigestPanel.tsx
@@ -0,0 +1,514 @@
+"use client";
+
+// The per-video digest review surface — the thing that makes a validation run
+// readable.
+//
+// Until this existed the digest layer had nowhere for a human to LOOK at what it
+// produced: the channel Digest stage is a queue-and-count card by its own
+// admission, and the per-video page had no digest reference at all. A ~100-video
+// validation run cannot mean anything while its output is invisible, and
+// STATE.md's own lesson from this layer is that "the digest layer's failure
+// modes live at the wiring, not in the pure functions."
+//
+// So this panel leads with the things that are wrong rather than the things that
+// are right: warnings first, then provenance (WHICH configuration produced
+// this), then the items. During a bake-off-driven validation the provenance line
+// is the entire point — two videos can carry visually similar chapters from
+// completely different engines, models or prompt versions.
+
+import { useState, useTransition } from "react";
+import { StreamActionLog } from "yt-dlp-transcript-common/components/StreamActionLog";
+import {
+ digestBucketAction,
+ saveDigestOverridesAction,
+ type DigestLaneChoice,
+} from "../../../digestActions";
+import type {
+ DigestChapter,
+ DigestProvenance,
+ DigestSectionKind,
+ DigestWarning,
+ EffectiveDigest,
+} from "yt-dlp-transcript-common/lib/digest";
+import { cancelJobAction } from "../../../../../jobs/actions";
+
+// One row per section KIND the current settings ask for, so a section that was
+// never generated is as visible as one that is stale. Computed on the server —
+// freshness needs the resolved target, which reaches settings and disk.
+export type DigestSectionState = {
+ kind: DigestSectionKind;
+ present: boolean;
+ fresh: boolean;
+ provenance?: DigestProvenance;
+};
+
+export type DigestPanelData = {
+ digest: EffectiveDigest;
+ hasMachineDigest: boolean;
+ sections: DigestSectionState[];
+ // What a regeneration would produce right now, for the "stale because…" line.
+ target: {
+ appId: string;
+ model: string;
+ promptVersion: number;
+ promptVariant?: string;
+ contextHash: string;
+ };
+ hasOverrides: boolean;
+ note: string;
+ remoteEnabled: boolean;
+};
+
+export function DigestPanel({
+ slug,
+ videoId,
+ data,
+}: {
+ slug: string;
+ videoId: string;
+ data: DigestPanelData;
+}) {
+ const [lane, setLane] = useState<DigestLaneChoice>("local");
+ const { digest, sections, target } = data;
+ const anyStale = sections.some((s) => !s.fresh);
+
+ return (
+ <section
+ aria-label="digest panel"
+ className="flex flex-col gap-4 rounded border border-border p-3"
+ >
+ <div className="flex flex-wrap items-baseline justify-between gap-2">
+ <h2 className="text-lg font-semibold">AI digest</h2>
+ <span
+ aria-label="digest freshness"
+ className={
+ !data.hasMachineDigest
+ ? "text-xs text-muted-foreground"
+ : anyStale
+ ? "text-xs font-medium text-warning"
+ : "text-xs font-medium text-success"
+ }
+ >
+ {!data.hasMachineDigest
+ ? "not generated"
+ : anyStale
+ ? "stale — regenerating would replace this"
+ : "current"}
+ </span>
+ </div>
+
+ {data.digest.derivedFrom && (
+ <SharedFrom derivedFrom={data.digest.derivedFrom} />
+ )}
+
+ {!data.hasMachineDigest ? (
+ <p aria-label="digest empty" className="text-sm text-muted-foreground">
+ No digest has been generated for this video yet. Run one below, or
+ sweep the whole channel from the channel page's Digest stage.
+ </p>
+ ) : (
+ <>
+ <Warnings warnings={digest.warnings} />
+ <SectionProvenance sections={sections} target={target} />
+ <Chapters chapters={digest.chapters} />
+ <Tags tags={digest.tags} />
+ </>
+ )}
+
+ <OverridesEditor
+ slug={slug}
+ videoId={videoId}
+ chapters={digest.chapters}
+ note={data.note}
+ hasOverrides={data.hasOverrides}
+ />
+
+ <div className="flex flex-col gap-2 border-t border-border pt-3">
+ <StreamActionLog
+ trigger={() => digestBucketAction(slug, [videoId], lane)}
+ cancelAction={cancelJobAction}
+ buttonLabel="Regenerate digest"
+ runningLabel="Generating…"
+ label="Regenerate digest"
+ extraControls={
+ <label className="flex items-center gap-1 text-xs text-muted-foreground">
+ Lane
+ <select
+ value={lane}
+ onChange={(e) => setLane(e.target.value as DigestLaneChoice)}
+ aria-label="lane for Regenerate digest"
+ className="rounded border border-border bg-card px-1 py-0.5"
+ >
+ <option value="local">local (ollama)</option>
+ <option value="remote" disabled={!data.remoteEnabled}>
+ metered{data.remoteEnabled ? "" : " — disabled in Settings"}
+ </option>
+ </select>
+ </label>
+ }
+ />
+ <p className="text-xs text-muted-foreground">
+ Regeneration only ever rewrites <code>ai-digest.json</code>. Hand
+ corrections live in <code>ai-digest.overrides.json</code> and are never
+ touched by it — which is what lets a correction survive a re-sweep of a
+ corpus too large to generate twice.
+ </p>
+ </div>
+ </section>
+ );
+}
+
+// A borrowed digest must NEVER be presented as native. Sharing only happens at
+// near-zero measured offset, so the offset is shown as the reason it was
+// considered safe rather than left implicit.
+function SharedFrom({
+ derivedFrom,
+}: {
+ derivedFrom: NonNullable<EffectiveDigest["derivedFrom"]>;
+}) {
+ const [channelSlug, id] = derivedFrom.slug.split("/");
+ return (
+ <div
+ aria-label="digest shared from"
+ className="rounded border border-info/30 bg-info-soft p-2 text-sm"
+ >
+ <p>
+ <strong>Shared from a duplicate.</strong> This digest was generated for{" "}
+ <a
+ href={`/channels/${channelSlug}/videos/${id}`}
+ className="underline font-mono"
+ >
+ {derivedFrom.slug}
+ </a>
+ , the canonical member of its cluster, and copied here.
+ </p>
+ <p className="text-xs text-muted-foreground mt-1">
+ Cue timings were measured{" "}
+ <strong>{derivedFrom.offsetSeconds.toFixed(2)}s</strong> apart at the
+ sampled anchors, which is why placing it here was considered safe.
+ Regenerating the canonical member re-applies the share.
+ </p>
+ </div>
+ );
+}
+
+// Warnings lead. This is the array Phase 11a's review queue is built on, and
+// showing it per-video first is what makes a validation run readable: a section
+// generated from 3 of 5 chunks looks perfectly healthy until you read these.
+function Warnings({ warnings }: { warnings: DigestWarning[] }) {
+ if (warnings.length === 0) {
+ return (
+ <p aria-label="digest warnings empty" className="text-xs text-success">
+ No warnings — every chunk parsed and every item passed the guards.
+ </p>
+ );
+ }
+ const byCode = new Map<string, DigestWarning[]>();
+ for (const w of warnings) {
+ const list = byCode.get(w.code) ?? [];
+ list.push(w);
+ byCode.set(w.code, list);
+ }
+ return (
+ <div
+ aria-label="digest warnings"
+ className="rounded border border-warning/30 bg-warning-soft p-2"
+ >
+ <p className="text-sm font-medium">
+ {warnings.length} warning{warnings.length === 1 ? "" : "s"}
+ </p>
+ <ul className="mt-1 flex flex-col gap-1 text-xs">
+ {Array.from(byCode.entries()).map(([code, list]) => (
+ <li key={code} aria-label={`digest warning ${code}`}>
+ <code className="font-medium">{code}</code> × {list.length}
+ <span className="text-muted-foreground">
+ {" "}
+ ({list[0].section}
+ {list[0].chunk !== undefined ? `, chunk ${list[0].chunk}` : ""})
+ </span>
+ {list[0].value && (
+ <span className="text-muted-foreground">
+ {" "}
+ — first: <code>{list[0].value}</code>
+ </span>
+ )}
+ {list[0].detail && (
+ <span className="text-muted-foreground"> — {list[0].detail}</span>
+ )}
+ </li>
+ ))}
+ </ul>
+ </div>
+ );
+}
+
+// WHICH configuration produced this. During a bake-off-driven validation, two
+// videos can carry similar-looking chapters from entirely different engines.
+function SectionProvenance({
+ sections,
+ target,
+}: {
+ sections: DigestSectionState[];
+ target: DigestPanelData["target"];
+}) {
+ return (
+ <div aria-label="digest provenance" className="flex flex-col gap-2">
+ {sections.map((s) => (
+ <div
+ key={s.kind}
+ aria-label={`digest section ${s.kind}`}
+ className="rounded border border-border p-2 text-xs"
+ >
+ <div className="flex flex-wrap items-baseline gap-2">
+ <span className="text-sm font-medium">{s.kind}</span>
+ {!s.present ? (
+ <span className="text-muted-foreground">never generated</span>
+ ) : s.fresh ? (
+ <span className="text-success">current</span>
+ ) : (
+ <span className="text-warning">stale</span>
+ )}
+ </div>
+ {s.provenance ? (
+ <dl className="mt-1 grid grid-cols-[auto_1fr] gap-x-3 gap-y-0.5 font-mono">
+ <Row label="engine" value={s.provenance.appId} />
+ <Row
+ label="model"
+ value={
+ s.provenance.modelRequested &&
+ s.provenance.modelRequested !== s.provenance.model
+ ? `${s.provenance.modelRequested} → ${s.provenance.model}`
+ : s.provenance.model
+ }
+ />
+ <Row label="lane" value={s.provenance.lane} />
+ <Row
+ label="promptVersion"
+ value={String(s.provenance.promptVersion)}
+ />
+ <Row
+ label="promptVariant"
+ value={s.provenance.promptVariant ?? "(default)"}
+ />
+ <Row label="contextHash" value={s.provenance.contextHash} />
+ {s.provenance.chunks !== undefined && (
+ <Row
+ label="chunks"
+ value={`${s.provenance.chunksOk ?? "?"} of ${s.provenance.chunks} usable`}
+ />
+ )}
+ <Row label="generated" value={s.provenance.generatedAt} />
+ {s.provenance.costUsd !== undefined && (
+ <Row label="cost" value={`$${s.provenance.costUsd.toFixed(4)}`} />
+ )}
+ </dl>
+ ) : null}
+ {s.present && !s.fresh && (
+ <p className="mt-1 text-muted-foreground">
+ A regeneration would produce{" "}
+ <code>
+ {target.appId}/{target.model} v{target.promptVersion}
+ {target.promptVariant ? ` (${target.promptVariant})` : ""}
+ </code>
+ .
+ </p>
+ )}
+ </div>
+ ))}
+ </div>
+ );
+}
+
+function Row({ label, value }: { label: string; value: string }) {
+ return (
+ <>
+ <dt className="text-muted-foreground">{label}</dt>
+ <dd className="break-all">{value}</dd>
+ </>
+ );
+}
+
+function Chapters({ chapters }: { chapters: DigestChapter[] }) {
+ if (chapters.length === 0) {
+ return (
+ <p aria-label="digest chapters empty" className="text-sm text-muted-foreground">
+ No chapters.
+ </p>
+ );
+ }
+ return (
+ <div>
+ <h3 className="text-sm font-medium">Chapters ({chapters.length})</h3>
+ <ol aria-label="digest chapters" className="mt-1 flex flex-col gap-0.5">
+ {chapters.map((c) => (
+ <li
+ key={c.id}
+ aria-label={`digest chapter ${c.id}`}
+ className="flex gap-2 text-sm"
+ >
+ <span className="font-mono text-muted-foreground">{c.clock}</span>
+ <span>{c.title}</span>
+ {c.decidedBy === "human" && (
+ <span className="text-xs text-info">edited</span>
+ )}
+ </li>
+ ))}
+ </ol>
+ </div>
+ );
+}
+
+function Tags({ tags }: { tags: EffectiveDigest["tags"] }) {
+ if (tags.length === 0) return null;
+ return (
+ <div>
+ <h3 className="text-sm font-medium">Tags ({tags.length})</h3>
+ <ul aria-label="digest tags" className="mt-1 flex flex-wrap gap-1">
+ {tags.map((t) => (
+ <li
+ key={t.id}
+ className="rounded bg-muted px-1.5 py-0.5 text-xs"
+ >
+ {t.tag}
+ {t.decidedBy === "human" && <span className="text-info"> ✎</span>}
+ </li>
+ ))}
+ </ul>
+ </div>
+ );
+}
+
+// Corrections. Deliberately minimal: retitle a chapter, or suppress it. Those
+// are the two things a reviewer actually does, and both are id-keyed so a
+// regeneration that re-emits the same segmentation keeps honoring them.
+function OverridesEditor({
+ slug,
+ videoId,
+ chapters,
+ note,
+ hasOverrides,
+}: {
+ slug: string;
+ videoId: string;
+ chapters: DigestChapter[];
+ note: string;
+ hasOverrides: boolean;
+}) {
+ const [rows, setRows] = useState(() =>
+ chapters.map((c) => ({ id: c.id, title: c.title, enabled: true })),
+ );
+ const [noteText, setNoteText] = useState(note);
+ const [pending, startTransition] = useTransition();
+ const [result, setResult] = useState<string | null>(null);
+
+ // The panel re-renders from the server after a save (revalidatePath), so the
+ // row list can go stale against a regeneration that happened in between.
+ const chapterIds = chapters.map((c) => c.id).join(",");
+ const rowIds = rows.map((r) => r.id).join(",");
+ if (chapterIds !== rowIds) {
+ setRows(chapters.map((c) => ({ id: c.id, title: c.title, enabled: true })));
+ }
+
+ if (chapters.length === 0 && !hasOverrides) return null;
+
+ const save = (): void => {
+ setResult(null);
+ startTransition(async () => {
+ const edited = rows.filter((row, i) => {
+ const original = chapters[i];
+ return (
+ !original || row.title !== original.title || row.enabled === false
+ );
+ });
+ const res = await saveDigestOverridesAction(slug, videoId, {
+ chapters: edited.map((row) => ({
+ id: row.id,
+ // start/clock are patched in from the machine item by mergeItems — an
+ // override authored as (id + title) must not blank them.
+ start: chapters.find((c) => c.id === row.id)?.start ?? 0,
+ clock: chapters.find((c) => c.id === row.id)?.clock ?? "00:00:00",
+ title: row.title,
+ decidedBy: "human" as const,
+ ...(row.enabled ? {} : { enabled: false }),
+ })),
+ ...(noteText.trim() ? { note: noteText.trim() } : {}),
+ });
+ setResult(res.ok ? "Saved." : res.error);
+ });
+ };
+
+ return (
+ <details className="text-sm">
+ <summary className="cursor-pointer font-medium">
+ Correct this digest{hasOverrides ? " (edited)" : ""}
+ </summary>
+ <div className="mt-2 flex flex-col gap-2">
+ {rows.map((row, i) => (
+ <div key={row.id} className="flex items-center gap-2">
+ <input
+ type="checkbox"
+ checked={row.enabled}
+ aria-label={`keep chapter ${row.id}`}
+ onChange={(e) =>
+ setRows((prev) =>
+ prev.map((r, j) =>
+ j === i ? { ...r, enabled: e.target.checked } : r,
+ ),
+ )
+ }
+ />
+ <span className="font-mono text-xs text-muted-foreground">
+ {chapters[i]?.clock}
+ </span>
+ <input
+ type="text"
+ value={row.title}
+ aria-label={`title for chapter ${row.id}`}
+ onChange={(e) =>
+ setRows((prev) =>
+ prev.map((r, j) =>
+ j === i ? { ...r, title: e.target.value } : r,
+ ),
+ )
+ }
+ className="flex-1 rounded border border-border bg-card px-2 py-1 text-sm"
+ />
+ </div>
+ ))}
+ <label className="flex flex-col gap-1">
+ <span className="text-xs font-medium">Note (why)</span>
+ <input
+ type="text"
+ value={noteText}
+ aria-label="digest override note"
+ onChange={(e) => setNoteText(e.target.value)}
+ className="rounded border border-border bg-card px-2 py-1 text-sm"
+ />
+ </label>
+ <div className="flex items-center gap-3">
+ <button
+ type="button"
+ onClick={save}
+ disabled={pending}
+ className="px-3 py-1.5 rounded-md bg-primary text-primary-foreground text-sm font-medium hover:opacity-90 disabled:opacity-50"
+ >
+ {pending ? "Saving…" : "Save corrections"}
+ </button>
+ {result && (
+ <span
+ role="status"
+ className="text-sm text-muted-foreground"
+ >
+ {result}
+ </span>
+ )}
+ </div>
+ <p className="text-xs text-muted-foreground">
+ Unchecking a chapter suppresses it without deleting it, so a
+ regeneration that re-emits the same item does not resurrect something
+ you rejected.
+ </p>
+ </div>
+ </details>
+ );
+}
diff --git a/editor/app/channels/[slug]/videos/[id]/page.tsx b/editor/app/channels/[slug]/videos/[id]/page.tsx
@@ -26,8 +26,23 @@ import {
queueKeyForUrl,
} from "yt-dlp-transcript-common/lib/platform";
import { getRegistry } from "yt-dlp-transcript-common/jobs/registry";
+import { getSettings } from "yt-dlp-transcript-common/lib/settings";
+import {
+ loadDigest,
+ loadDigestOverrides,
+} from "yt-dlp-transcript-common/lib/digest-server";
+import {
+ effectiveDigest,
+ isSectionFresh,
+} from "yt-dlp-transcript-common/lib/digest";
+import { resolveDigestTarget } from "yt-dlp-transcript-common/controller/digestTarget";
import { RunningJobsList } from "../../../../jobs/components/RunningJobsList";
import { VideoPanel, type VideoFile } from "./components/VideoPanel";
+import {
+ DigestPanel,
+ type DigestPanelData,
+ type DigestSectionState,
+} from "./components/DigestPanel";
export const dynamic = "force-dynamic";
@@ -137,6 +152,37 @@ export default async function VideoDetailPage({
}
: null;
+ // The digest artifact, its human shadow, and whether each section still
+ // matches what a regeneration would produce. Freshness is computed HERE
+ // rather than in the panel because it needs the resolved target, which reads
+ // settings, the engine registry and the channel's context note.
+ const [digestRecord, digestOverrides, digestTarget] = await Promise.all([
+ loadDigest(videoDir),
+ loadDigestOverrides(videoDir),
+ resolveDigestTarget({ paths: getPaths(), channelSlug: slug, lane: "local" }),
+ ]);
+ const digestData: DigestPanelData = {
+ digest: effectiveDigest(digestRecord, digestOverrides),
+ hasMachineDigest: digestRecord !== null,
+ sections: digestTarget.sections.map(
+ (kind): DigestSectionState => ({
+ kind,
+ present: digestRecord?.sections[kind] !== undefined,
+ // A shared digest is fresh for the receiving video as long as it still
+ // points at its canonical member — that member's own freshness is what
+ // drives regeneration (isSharedFrom's contract).
+ fresh:
+ digestRecord?.derivedFrom != null ||
+ isSectionFresh(digestRecord, kind, digestTarget.target),
+ provenance: digestRecord?.sections[kind]?.provenance,
+ }),
+ ),
+ target: digestTarget.target,
+ hasOverrides: digestOverrides !== null,
+ note: digestOverrides?.note ?? "",
+ remoteEnabled: getSettings().digest.remoteEnabled,
+ };
+
const registry = getRegistry();
const existingQueues = registry.activeQueueNames();
const activeJobs = registry
@@ -219,6 +265,8 @@ export default async function VideoDetailPage({
savedVideo={savedVideo}
coverage={coverage}
/>
+
+ <DigestPanel slug={slug} videoId={id} data={digestData} />
</div>
);
}
diff --git a/editor/app/components/dashboard/ChannelsTable.tsx b/editor/app/components/dashboard/ChannelsTable.tsx
@@ -32,6 +32,12 @@ export function ChannelsTable({
<th className="text-left font-medium px-3 py-2">Slug</th>
<th className="text-left font-medium px-3 py-2">Handling</th>
<th className="text-right font-medium px-3 py-2">Videos</th>
+ <th
+ className="text-right font-medium px-3 py-2 whitespace-nowrap"
+ title="Videos with no AI digest at the current prompt/model identity"
+ >
+ No digest
+ </th>
<th className="text-left font-medium px-3 py-2 whitespace-nowrap">
Last sync
</th>
@@ -57,6 +63,11 @@ export function ChannelsTable({
<td className="px-3 py-2 text-right tabular-nums">
{c.videoCount}
</td>
+ {/* Muted: during the backfill this is nearly every video, so
+ it is a coverage readout rather than a call to action. */}
+ <td className="px-3 py-2 text-right tabular-nums text-xs text-muted-foreground">
+ {c.noDigest}
+ </td>
<td className="px-3 py-2 text-xs text-muted-foreground whitespace-nowrap tabular-nums">
{c.lastSyncedAt == null
? "never"
diff --git a/editor/app/components/dashboard/NeedsWorkPanel.tsx b/editor/app/components/dashboard/NeedsWorkPanel.tsx
@@ -68,6 +68,17 @@ export function NeedsWorkPanel({
✎ {c.untranscribed}
</span>
)}
+ {/* Muted, not coloured: during the backfill this is nearly
+ every video in the channel, so it is context for the row
+ rather than a call to action like the two above. */}
+ {c.noDigest > 0 && (
+ <span
+ title={`${c.noDigest} video(s) with no digest at the current identity`}
+ className="rounded px-1.5 py-0.5 text-[10px] tabular-nums bg-muted text-muted-foreground"
+ >
+ ◆ {c.noDigest}
+ </span>
+ )}
</span>
</div>
<div className="flex flex-wrap items-center gap-1.5">
diff --git a/editor/app/components/dashboard/PipelineBand.tsx b/editor/app/components/dashboard/PipelineBand.tsx
@@ -8,6 +8,7 @@ import type { WidgetSyncPayload } from "../../api/widget/sync/route";
import { ActiveJobsLive } from "../../jobs/components/ActiveJobsLive";
import { PauseTranscriptionsButton } from "../../jobs/components/PauseTranscriptionsButton";
import { PauseDownloadsButton } from "../../jobs/components/PauseDownloadsButton";
+import { DigestSweepControls } from "../../jobs/components/DigestSweepControls";
import { syncAllChannelsAction, type SyncAllResult } from "../../channels/actions";
import { fmtTime } from "../../widget/lib/relativeTime";
@@ -37,6 +38,13 @@ export function PipelineBand({
const busy = workerList.filter((w) => w.busy).length;
const paused = workers?.paused ?? false;
const downloadsPaused = workers?.downloadsPaused ?? false;
+ const digest = sync?.digest ?? null;
+ // Deliberately not rounded up. At 0.13% a "1%" would be a lie of the kind
+ // that makes an 80-day backfill look nearly begun.
+ const digestPct =
+ digest && digest.videos > 0
+ ? (digest.digested / digest.videos) * 100
+ : null;
const lastSyncText =
sync == null
@@ -112,6 +120,36 @@ export function PipelineBand({
label={<span className="text-destructive">downloads paused</span>}
/>
)}
+ {digest && (
+ <Instrument
+ dotClass={
+ digest.paused
+ ? "bg-destructive"
+ : digest.sweeping
+ ? "bg-success"
+ : "bg-muted-foreground/40"
+ }
+ label={
+ <>
+ ◆ digests{" "}
+ <span className="font-medium">
+ {digest.digested.toLocaleString()}
+ </span>
+ {digestPct !== null && (
+ <span className="text-muted-foreground">
+ {" "}
+ ({digestPct < 1 ? digestPct.toFixed(2) : digestPct.toFixed(1)}%
+ of {digest.videos.toLocaleString()})
+ </span>
+ )}
+ {digest.paused && (
+ <span className="text-destructive"> · paused</span>
+ )}
+ {!digest.paused && digest.sweeping && " · sweeping"}
+ </>
+ }
+ />
+ )}
</div>
</div>
@@ -121,6 +159,11 @@ export function PipelineBand({
paused={downloadsPaused}
onChange={onWorkersChange}
/>
+ <DigestSweepControls
+ sweeping={digest?.sweeping ?? false}
+ paused={digest?.paused ?? false}
+ onChange={onSynced}
+ />
<SyncAllButton onSynced={onSynced} />
<Link
href="/channels/new"
diff --git a/editor/app/components/dashboard/types.ts b/editor/app/components/dashboard/types.ts
@@ -11,4 +11,8 @@ export type DashboardChannel = {
hasUrl: boolean;
undownloaded: number;
untranscribed: number;
+ // Videos with no digest at the CURRENT identity. The backfill's per-channel
+ // denominator, and the only per-channel number that makes corpus coverage
+ // legible while a multi-week sweep is running.
+ noDigest: number;
};
diff --git a/editor/app/jobs/actions.ts b/editor/app/jobs/actions.ts
@@ -8,6 +8,10 @@ import { getRegistry } from "yt-dlp-transcript-common/jobs/registry";
import { readJobMeta } from "yt-dlp-transcript-common/jobs/jobMeta";
import type { JobSpec } from "yt-dlp-transcript-common/jobs/jobSpec";
import type { StreamActionResult } from "yt-dlp-transcript-common/jobs/streamCommand";
+import {
+ startDigestSweep,
+ stopDigestSweep,
+} from "yt-dlp-transcript-common/controller/digestSweep";
import { runJobSpec } from "./runJobSpec";
import { buildQueueView } from "./queue/buildQueueView";
@@ -169,6 +173,73 @@ export async function resumeDownloadsAction(): Promise<DownloadsPauseResult> {
return setDownloadsPaused(false);
}
+// Global, restart-surviving digest pause — the counterpart of the downloads one
+// and, until now, a flag with no writer. `digestBatch`'s limit() has always
+// re-read `digestsPaused` at dispatch time and returned 0 to make the pool
+// idle-wait (a real pause: the job HOLDS rather than ending, so nothing has to
+// be re-derived on resume), and settings/actions.ts passed the field through
+// untouched with a comment saying "the dashboard/channel controls own the
+// pause". Those controls did not exist, so the pause could not be set from
+// anywhere. This is them.
+//
+// It needs no boot hook, for the same reason downloadsPaused doesn't: the flag
+// is consulted at dispatch, not applied to a live pool.
+export type DigestPauseResult = { ok: boolean; error?: string };
+
+async function setDigestsPaused(paused: boolean): Promise<DigestPauseResult> {
+ try {
+ const current = getSettings();
+ if (current.digest.digestsPaused !== paused) {
+ await writeSettings({
+ ...current,
+ digest: { ...current.digest, digestsPaused: paused },
+ });
+ }
+ } catch (e) {
+ return { ok: false, error: (e as Error).message };
+ }
+ revalidatePath("/");
+ revalidatePath("/jobs");
+ return { ok: true };
+}
+
+export async function pauseDigestsAction(): Promise<DigestPauseResult> {
+ return setDigestsPaused(true);
+}
+
+export async function resumeDigestsAction(): Promise<DigestPauseResult> {
+ return setDigestsPaused(false);
+}
+
+// Arm / disarm the corpus-wide sweep. Distinct from the pause: a pause holds a
+// running sweep at zero throughput, this decides whether there is a sweep at
+// all — and it persists, so the boot hook resumes it.
+export type DigestSweepResult = { ok: boolean; jobId?: string; error?: string };
+
+export async function startDigestSweepAction(): Promise<DigestSweepResult> {
+ try {
+ const jobId = await startDigestSweep();
+ revalidatePath("/");
+ revalidatePath("/jobs");
+ return jobId
+ ? { ok: true, jobId }
+ : { ok: false, error: "The sweep could not be started (see job logs)." };
+ } catch (e) {
+ return { ok: false, error: (e as Error).message };
+ }
+}
+
+export async function stopDigestSweepAction(): Promise<DigestSweepResult> {
+ try {
+ await stopDigestSweep();
+ revalidatePath("/");
+ revalidatePath("/jobs");
+ return { ok: true };
+ } catch (e) {
+ return { ok: false, error: (e as Error).message };
+ }
+}
+
// Retention scopes offered by the ClearLogsMenu. "all" clears every finished
// job's log; the day-scopes clear anything older than that. Running/queued jobs
// are never deleted (see pruneJobLogs).
diff --git a/editor/app/jobs/active/buildActiveJobs.ts b/editor/app/jobs/active/buildActiveJobs.ts
@@ -39,12 +39,27 @@ function computeEtaSeconds(
job: JobRecord,
remaining: number,
now: number,
+ remainingAudioSeconds?: number,
): number | undefined {
const count = job.completedTaskCount ?? 0;
const totalMs = job.completedTaskMs ?? 0;
if (count < 1 || remaining <= 0 || job.startedAt === undefined) {
return undefined;
}
+ // Prefer an AUDIO-HOUR estimate where the work is proportional to length.
+ // Averaging tasks assumes every unit costs about the same, which is true for
+ // downloads and wildly false for digests: this corpus is ~77k videos and ~77k
+ // audio-hours, and a channel of 9-hour VODs and a channel of 10-minute clips
+ // have the same task count and a 50x difference in cost. A task average would
+ // therefore quote an ETA that is wrong by more than an order of magnitude at
+ // exactly the moment an operator most needs it — the start of an 80-day run.
+ const doneAudio = job.completedTaskAudioSeconds ?? 0;
+ if (doneAudio > 0 && remainingAudioSeconds && remainingAudioSeconds > 0) {
+ const secondsPerAudioSecond = totalMs / 1000 / doneAudio;
+ const elapsedMs = Math.max(1, now - job.startedAt);
+ const concurrency = Math.max(1, totalMs / elapsedMs);
+ return (remainingAudioSeconds * secondsPerAudioSecond) / concurrency;
+ }
const avgProcMs = totalMs / count; // measured average per task
const elapsedMs = Math.max(1, now - job.startedAt);
const concurrency = Math.max(1, totalMs / elapsedMs); // effective parallelism
@@ -62,8 +77,21 @@ function computeJobProgressView(
): RunningJobsListItem["progress"] {
const snap = job.progress;
if (!snap || !stat) return undefined;
+ // `current` is RE-COUNTED from disk (readChannelStat), never reported by the
+ // runner — which is why the digest metric needed its own on-disk counter
+ // (digestCount) rather than a number the batch could have just told us.
+ // A runner-reported `current` WINS. The disk re-count below cannot see a
+ // regeneration — a regenerated digest is rewritten in place, so the file
+ // count never moves and the bar sits at 0% for the whole job. Only the runner
+ // knows it did the work. Downloads and transcripts report nothing and keep
+ // the disk re-count, unchanged.
const current =
- snap.metric === "downloads" ? stat.downloadCount : stat.transcriptCount;
+ snap.current ??
+ (snap.metric === "downloads"
+ ? stat.downloadCount
+ : snap.metric === "digests"
+ ? (stat.digestCount ?? 0)
+ : stat.transcriptCount);
const range = Math.max(0, snap.target - snap.initial);
const advance = Math.max(0, current - snap.initial);
const pct =
@@ -75,7 +103,12 @@ function computeJobProgressView(
current,
target: snap.target,
pct,
- etaSeconds: computeEtaSeconds(job, remaining, now),
+ etaSeconds: computeEtaSeconds(
+ job,
+ remaining,
+ now,
+ snap.remainingAudioSeconds,
+ ),
};
}
diff --git a/editor/app/jobs/components/DigestSweepControls.tsx b/editor/app/jobs/components/DigestSweepControls.tsx
@@ -0,0 +1,104 @@
+"use client";
+
+import { useEffect, useState, useTransition } from "react";
+import { useRouter } from "next/navigation";
+import {
+ pauseDigestsAction,
+ resumeDigestsAction,
+ startDigestSweepAction,
+ stopDigestSweepAction,
+} from "../actions";
+
+// The two digest controls, side by side, because they are NOT the same control
+// and conflating them is how an operator loses a week of GPU time:
+//
+// Sweep on/off — is there a corpus-wide backfill at all. Persisted, so a
+// server restart resumes it (editor/instrumentation.ts).
+// Pause/resume — hold a running sweep at zero throughput without ending it.
+// The batch's limit() returns 0, which makes the pool idle-WAIT rather than
+// finish, so resuming costs nothing and re-derives nothing.
+//
+// Stopping the sweep drains rather than cancels: the channel in flight finishes
+// instead of losing a part-generated video.
+export function DigestSweepControls({
+ sweeping,
+ paused,
+ onChange,
+}: {
+ sweeping: boolean;
+ paused: boolean;
+ onChange?: () => void | Promise<void>;
+}) {
+ const [pending, startTransition] = useTransition();
+ const [error, setError] = useState<string | null>(null);
+ const router = useRouter();
+ // Disabled until hydrated. A click on a server-rendered button before React
+ // attaches fires nothing at all — no request, no job, no error — which is the
+ // recorded root cause of the digest pilot's "un-created job".
+ const [mounted, setMounted] = useState(false);
+ useEffect(() => setMounted(true), []);
+
+ function run(fn: () => Promise<{ ok: boolean; error?: string }>) {
+ setError(null);
+ startTransition(async () => {
+ const result = await fn();
+ if (!result.ok) setError(result.error ?? "Failed.");
+ if (onChange) await onChange();
+ else router.refresh();
+ });
+ }
+
+ const disabled = pending || !mounted;
+
+ return (
+ <div className="flex items-center gap-2">
+ <button
+ type="button"
+ disabled={disabled}
+ aria-label={sweeping ? "stop digest sweep" : "start digest sweep"}
+ onClick={() =>
+ run(sweeping ? stopDigestSweepAction : startDigestSweepAction)
+ }
+ title={
+ sweeping
+ ? "Stop the corpus-wide digest backfill. The channel in flight finishes first; nothing already generated is lost."
+ : "Start the corpus-wide digest backfill: every channel in turn, heaviest first by remaining audio-hours. Survives a restart."
+ }
+ className={
+ sweeping
+ ? "px-3 py-1.5 rounded-md bg-warning text-warning-foreground text-sm font-medium hover:bg-warning/90 disabled:opacity-50"
+ : "px-3 py-1.5 rounded-md border border-border text-sm hover:bg-muted disabled:opacity-50"
+ }
+ >
+ {sweeping ? "Stop Digest Sweep" : "Start Digest Sweep"}
+ </button>
+ {(sweeping || paused) && (
+ <button
+ type="button"
+ disabled={disabled}
+ aria-label={paused ? "resume digests" : "pause digests"}
+ onClick={() =>
+ run(paused ? resumeDigestsAction : pauseDigestsAction)
+ }
+ title={
+ paused
+ ? "Resume digest generation. The running job picks up where it left off — it was holding, not stopped."
+ : "Hold digest generation without ending the sweep. The running job idles at zero and resumes instantly."
+ }
+ className={
+ paused
+ ? "px-3 py-1.5 rounded-md bg-warning text-warning-foreground text-sm font-medium hover:bg-warning/90 disabled:opacity-50"
+ : "px-3 py-1.5 rounded-md border border-border text-sm hover:bg-muted disabled:opacity-50"
+ }
+ >
+ {paused ? "Resume Digests" : "Pause Digests"}
+ </button>
+ )}
+ {error && (
+ <span role="alert" className="text-xs text-destructive">
+ {error}
+ </span>
+ )}
+ </div>
+ );
+}
diff --git a/editor/app/jobs/components/RunningJobsList.tsx b/editor/app/jobs/components/RunningJobsList.tsx
@@ -3,6 +3,10 @@
import Link from "next/link";
import { useEffect, useState } from "react";
import { formatDuration } from "yt-dlp-transcript-common/lib/format";
+import type {
+ JobProgressMetric,
+ JobTaskKind,
+} from "yt-dlp-transcript-common/jobs/registry";
import { JobLogTail } from "../[id]/components/JobLogTail";
import { jobKindLabel } from "../jobKindLabels";
import { DrainJobButton } from "./DrainJobButton";
@@ -14,7 +18,9 @@ import { ReorderJobButtons } from "./ReorderJobButtons";
export type RunningJobsTask = {
id: string;
label: string;
- kind: "download" | "transcribe";
+ // Imported for the same reason as `metric` below: a re-spelled literal here
+ // would not fail the build when JobTaskKind grew a member.
+ kind: JobTaskKind;
fraction?: number;
detail?: string;
// Epoch ms when this sub-operation started, for the live "running for" timer.
@@ -38,7 +44,9 @@ export type RunningJobsListItem = {
channelSlug?: string;
videoId?: string;
progress?: {
- metric: "downloads" | "transcripts";
+ // Imported, NOT re-spelled: a literal copy here silently drifted from
+ // JobProgressMetric and would not fail the build when the union grew.
+ metric: JobProgressMetric;
initial: number;
current: number;
target: number;
@@ -214,16 +222,25 @@ function useNow(): number | null {
return now;
}
+// Records over JobTaskKind, so adding a kind is a compile error here.
+const TASK_KIND_VERB: Record<JobTaskKind, string> = {
+ download: "Downloading",
+ transcribe: "Transcribing",
+ digest: "Digesting",
+};
+
+const METRIC_FILL_BY_TASK: Record<JobTaskKind, string> = {
+ download: "bg-success/60",
+ transcribe: "bg-success",
+ digest: "bg-info",
+};
+
function TaskProgressBar({ task }: { task: RunningJobsTask }) {
// How long this task has been running. Null until mounted (see useNow);
// formatDuration returns "" for 0, so the just-started case shows "0:00".
const now = useNow();
const probing = task.phase === "probing";
- const verb = probing
- ? "Probing audio"
- : task.kind === "download"
- ? "Downloading"
- : "Transcribing";
+ const verb = probing ? "Probing audio" : TASK_KIND_VERB[task.kind];
// While probing, fill against the estimated probe duration (a distinct violet
// "scanning" bar) rather than the frozen download fraction. yt-dlp is paused,
// so the download fraction wouldn't advance anyway. Falls back to an
@@ -246,9 +263,7 @@ function TaskProgressBar({ task }: { task: RunningJobsTask }) {
const pct = hasFraction ? Math.round((fraction as number) * 100) : 0;
const fillClass = probing
? "bg-violet-400 dark:bg-violet-500"
- : task.kind === "download"
- ? "bg-success/60"
- : "bg-success";
+ : METRIC_FILL_BY_TASK[task.kind];
const pulseClass = probing
? "bg-violet-400 dark:bg-violet-500"
: "bg-warning";
@@ -304,15 +319,26 @@ function TaskProgressBar({ task }: { task: RunningJobsTask }) {
);
}
+// One place per metric, so adding a metric to JobProgressMetric is a compile
+// error here (Record over the union) rather than a silently-wrong label.
+const METRIC_LABELS: Record<JobProgressMetric, string> = {
+ downloads: "Downloads",
+ transcripts: "Transcripts",
+ digests: "Digests",
+};
+
+const METRIC_FILL: Record<JobProgressMetric, string> = {
+ downloads: "bg-success/60",
+ transcripts: "bg-success",
+ digests: "bg-info",
+};
+
function JobProgressBar({
progress,
}: {
progress: NonNullable<RunningJobsListItem["progress"]>;
}) {
- const label =
- progress.metric === "downloads"
- ? `Downloads: ${progress.current} / ${progress.target}`
- : `Transcripts: ${progress.current} / ${progress.target}`;
+ const label = `${METRIC_LABELS[progress.metric]}: ${progress.current} / ${progress.target}`;
// Append an ETA once the batch has a measured average. formatDuration returns
// "" for 0/falsy, so guard against printing a bare "·".
const remaining = progress.target - progress.current;
@@ -322,10 +348,7 @@ function JobProgressBar({
: typeof progress.etaSeconds === "number"
? `~${formatDuration(Math.max(1, Math.round(progress.etaSeconds)))} left`
: "estimating…";
- const fillClass =
- progress.metric === "downloads"
- ? "bg-success/60"
- : "bg-success";
+ const fillClass = METRIC_FILL[progress.metric];
return (
<div className="flex flex-col gap-1">
<div
diff --git a/editor/app/jobs/jobReplayRegistry.ts b/editor/app/jobs/jobReplayRegistry.ts
@@ -35,6 +35,10 @@ import {
transcribeBucketAction,
transcribeMissingAction,
} from "../channels/[slug]/whisperActions";
+import {
+ digestChannelAction,
+ type DigestLaneChoice,
+} from "../channels/[slug]/digestActions";
import { persistKeptAction } from "../channels/[slug]/persistActions";
import { fetchPostsAction } from "../channels/[slug]/socialActions";
import {
@@ -73,6 +77,31 @@ function params(spec: JobSpec): {
}
export const JOB_REPLAY_HANDLERS: Record<string, ReplayHandler> = {
+ // Both digest lanes replay through one action; the lane comes from params so a
+ // bookmarked metered run stays metered (and is refused if the lane has since
+ // been turned off, rather than quietly falling back to local).
+ "digest-channel-local": (spec) => {
+ const { p, queueKey } = params(spec);
+ return digestChannelAction(
+ spec.slug,
+ (str(p.lane) as DigestLaneChoice | undefined) ?? "local",
+ queueKey,
+ str(p.order),
+ num(p.limitCount),
+ bool(p.force),
+ );
+ },
+ "digest-channel-remote": (spec) => {
+ const { p, queueKey } = params(spec);
+ return digestChannelAction(
+ spec.slug,
+ (str(p.lane) as DigestLaneChoice | undefined) ?? "remote",
+ queueKey,
+ str(p.order),
+ num(p.limitCount),
+ bool(p.force),
+ );
+ },
"whisper-all": (spec) => {
const { p, queueKey } = params(spec);
return transcribeMissingAction(
diff --git a/editor/app/page.tsx b/editor/app/page.tsx
@@ -9,6 +9,7 @@ import {
siteChannelSlugs,
} from "yt-dlp-transcript-common/lib/site";
import {
+ actionableNoDigestCount,
actionableUndownloadedCount,
actionableUntranscribedCount,
loadActionableSummary,
@@ -75,6 +76,7 @@ export default async function Dashboard({
hasUrl: Boolean(r.channel.config.url),
undownloaded: actionableUndownloadedCount(r),
untranscribed: actionableUntranscribedCount(r),
+ noDigest: actionableNoDigestCount(r),
}));
// "Needs work" payload, same shape the widget endpoint the cockpit polls
@@ -86,6 +88,7 @@ export default async function Dashboard({
slug: c.slug,
undownloaded: c.undownloaded,
untranscribed: c.untranscribed,
+ noDigest: c.noDigest,
}))
.sort(
(a, b) =>
diff --git a/editor/app/settings/actions.ts b/editor/app/settings/actions.ts
@@ -21,6 +21,7 @@ import {
type SocialLink,
} from "yt-dlp-transcript-common/lib/settings";
import { DEFAULT_TRANSCRIPTION_APP_ID } from "yt-dlp-transcript-common/lib/transcriptionApps";
+import { isDigestTimestampMode } from "yt-dlp-transcript-common/lib/digest";
import {
DEFAULT_COOKIE_MODE,
isCookieMode,
@@ -201,6 +202,65 @@ export async function saveSettingsAction(
// Build pipeline. Values are clamped/coerced by sanitizeBuildPipeline inside
// writeSettings, so we only read the form here (NaN/blank → default). The
// deploy-page toggle also writes `mode`; whichever saves last wins.
+ // Digest. The metered lane's switch is a checkbox like any other, but note the
+ // asymmetry: everything else here defaults to the CURRENT value on a partial
+ // save, while `remoteEnabled` is read straight from the form so it can never be
+ // turned on by an unrelated save. writeSettings re-sanitizes the whole block.
+ const dD = getSettings().digest;
+ const digestAppsRaw = String(formData.get("digestAppsJson") ?? "").trim();
+ let digestApps: unknown = dD.apps;
+ if (digestAppsRaw) {
+ try {
+ digestApps = JSON.parse(digestAppsRaw);
+ } catch {
+ return { ok: false, error: "Digest app config payload is malformed" };
+ }
+ }
+ const digestSectionsRaw = formData.getAll("digestSections").map(String);
+ // A hidden marker, because unchecked checkboxes are simply ABSENT from a
+ // FormData: without it, a submit from any form that lacks the digest fields
+ // would read remoteEnabled as false and silently reset the block.
+ const digestFormPresent = formData.get("digestFormPresent") === "1";
+ const digestTimestampModeRaw = String(
+ formData.get("digestTimestampMode") ?? "",
+ ).trim();
+ const digestSettings = (
+ digestFormPresent
+ ? {
+ remoteEnabled: formData.get("digestRemoteEnabled") === "on",
+ longTailSeconds: Number.parseInt(
+ String(formData.get("digestLongTailSeconds") ?? "").trim(),
+ 10,
+ ),
+ localAppId:
+ String(formData.get("digestLocalAppId") ?? "").trim() ||
+ dD.localAppId,
+ remoteAppId:
+ String(formData.get("digestRemoteAppId") ?? "").trim() ||
+ dD.remoteAppId,
+ apps: digestApps,
+ // Not edited by this form — the dashboard/channel controls own the pause.
+ digestsPaused: dD.digestsPaused,
+ spendCapUsd: Number.parseFloat(
+ String(formData.get("digestSpendCapUsd") ?? "").trim(),
+ ),
+ sections: digestSectionsRaw.length > 0 ? digestSectionsRaw : dD.sections,
+ // Prompt SHAPE — both freshness-affecting (see digestPromptVariant),
+ // so a silent reset here would invalidate every digest generated
+ // under a non-default shape. Each still FALLS BACK to the current
+ // value rather than to the default when the field is absent: a form
+ // that rebuilds this block but omits a field is exactly how the reset
+ // bug happens, and the fallback is what makes omission harmless.
+ timestampMode: isDigestTimestampMode(digestTimestampModeRaw)
+ ? digestTimestampModeRaw
+ : dD.timestampMode,
+ promptVariant: formData.has("digestPromptVariant")
+ ? String(formData.get("digestPromptVariant") ?? "").trim()
+ : dD.promptVariant,
+ }
+ : dD
+ ) as SiteSettings["digest"];
+
const dB = defaultBuildPipeline();
const buildModeRaw = String(formData.get("buildMode") ?? "").trim();
const buildPipeline = {
@@ -247,6 +307,9 @@ export async function saveSettingsAction(
// Saved Videos page edits it). writeSettings re-sanitizes it regardless.
savedVideoBackup: getSettings().savedVideoBackup,
buildPipeline,
+ // Same: the Digest section of this form owns these fields, but an unrelated
+ // save must not reset them (and must never silently flip remoteEnabled on).
+ digest: digestSettings,
};
try {
await writeSettings(next);
diff --git a/editor/app/settings/components/DigestAppsField.tsx b/editor/app/settings/components/DigestAppsField.tsx
@@ -0,0 +1,185 @@
+"use client";
+
+// Per-engine digest configuration, serialized to a hidden JSON field — the
+// WorkersField idiom, for the same reason: `settings.digest.apps` is a keyed map
+// of optional fields, and a flat set of form inputs cannot express "this app
+// overrides only the model" without inventing a naming scheme per app.
+//
+// Only the fields an engine actually supports are rendered (DigestApp.fields),
+// so the metered lane never shows a `num_ctx` box it would ignore.
+//
+// A blank input means "unset" and is OMITTED from the JSON rather than written
+// as "" or 0 — sanitizeDigestApps drops empty values anyway, but writing them
+// would make an untouched app look configured, and `numCtx: 0` in particular
+// would read as a real (invalid) context size rather than as absence.
+
+import { useState } from "react";
+import type { DigestAppDescriptor } from "yt-dlp-transcript-common/lib/digestApps";
+import type { DigestAppConfig } from "yt-dlp-transcript-common/lib/digest";
+import { DEFAULT_DIGEST_NUM_CTX } from "yt-dlp-transcript-common/lib/digest";
+
+type Row = {
+ bin: string;
+ baseUrl: string;
+ model: string;
+ numCtx: string;
+ temperature: string;
+};
+
+function toRow(cfg: DigestAppConfig | undefined): Row {
+ return {
+ bin: cfg?.bin ?? "",
+ baseUrl: cfg?.baseUrl ?? "",
+ model: cfg?.model ?? "",
+ numCtx: cfg?.numCtx != null ? String(cfg.numCtx) : "",
+ temperature: cfg?.temperature != null ? String(cfg.temperature) : "",
+ };
+}
+
+function serialize(rows: Record<string, Row>): string {
+ const out: Record<string, DigestAppConfig> = {};
+ for (const [id, row] of Object.entries(rows)) {
+ const cfg: DigestAppConfig = {};
+ if (row.bin.trim()) cfg.bin = row.bin.trim();
+ if (row.baseUrl.trim()) cfg.baseUrl = row.baseUrl.trim();
+ if (row.model.trim()) cfg.model = row.model.trim();
+ const n = Number.parseInt(row.numCtx.trim(), 10);
+ if (Number.isFinite(n) && n > 0) cfg.numCtx = n;
+ const t = Number.parseFloat(row.temperature.trim());
+ if (Number.isFinite(t) && t >= 0) cfg.temperature = t;
+ if (Object.keys(cfg).length > 0) out[id] = cfg;
+ }
+ return JSON.stringify(out);
+}
+
+export function DigestAppsField({
+ apps,
+ initial,
+ name,
+}: {
+ apps: DigestAppDescriptor[];
+ initial: Record<string, DigestAppConfig>;
+ name: string;
+}) {
+ const [rows, setRows] = useState<Record<string, Row>>(() =>
+ Object.fromEntries(apps.map((a) => [a.id, toRow(initial[a.id])])),
+ );
+
+ const set = (id: string, key: keyof Row, value: string): void =>
+ setRows((prev) => ({ ...prev, [id]: { ...prev[id], [key]: value } }));
+
+ return (
+ <div className="flex flex-col gap-3">
+ {apps.map((app) => {
+ const row = rows[app.id] ?? toRow(undefined);
+ return (
+ <div
+ key={app.id}
+ aria-label={`digest engine ${app.id}`}
+ className="flex flex-col gap-2 rounded border border-border p-2"
+ >
+ <div className="flex items-baseline gap-2">
+ <span className="text-sm font-medium">{app.label}</span>
+ <code className="text-xs text-muted-foreground">{app.id}</code>
+ <span className="text-xs text-muted-foreground">
+ {app.metered ? "metered" : "local"}
+ </span>
+ </div>
+ {app.fields.model && (
+ <Cell
+ label="Model"
+ id={app.id}
+ field="model"
+ value={row.model}
+ onChange={set}
+ placeholder={app.defaultModel}
+ hint={`Blank uses ${app.defaultModel}. Freshness compares what you ask for, so an alias resolving to a full tag is not a model change.`}
+ />
+ )}
+ {app.fields.numCtx && (
+ <Cell
+ label="Context window (num_ctx)"
+ id={app.id}
+ field="numCtx"
+ value={row.numCtx}
+ onChange={set}
+ placeholder={String(DEFAULT_DIGEST_NUM_CTX)}
+ type="number"
+ hint={`Blank uses ${DEFAULT_DIGEST_NUM_CTX}, the size the bake-off measured best. The transcript slice is sized to this automatically — ollama truncates an over-long input SILENTLY, so the two must move together. Changing it changes the recorded identity and re-runs the corpus.`}
+ />
+ )}
+ {app.fields.temperature && (
+ <Cell
+ label="Temperature"
+ id={app.id}
+ field="temperature"
+ value={row.temperature}
+ onChange={set}
+ placeholder="0"
+ hint="0 for a structured extraction task."
+ />
+ )}
+ {app.fields.baseUrl && (
+ <Cell
+ label="Base URL"
+ id={app.id}
+ field="baseUrl"
+ value={row.baseUrl}
+ onChange={set}
+ placeholder="http://127.0.0.1:11434"
+ hint="Blank uses OLLAMA_URL, then the local default."
+ />
+ )}
+ {app.fields.bin && (
+ <Cell
+ label="Binary"
+ id={app.id}
+ field="bin"
+ value={row.bin}
+ onChange={set}
+ placeholder="claude"
+ hint="Blank uses CLAUDE_BIN, then PATH."
+ />
+ )}
+ </div>
+ );
+ })}
+ <input type="hidden" name={name} value={serialize(rows)} readOnly />
+ </div>
+ );
+}
+
+function Cell({
+ label,
+ id,
+ field,
+ value,
+ onChange,
+ placeholder,
+ hint,
+ type = "text",
+}: {
+ label: string;
+ id: string;
+ field: keyof Row;
+ value: string;
+ onChange: (id: string, key: keyof Row, value: string) => void;
+ placeholder?: string;
+ hint?: string;
+ type?: string;
+}) {
+ return (
+ <label className="flex flex-col gap-1 text-sm">
+ <span className="font-medium">{label}</span>
+ <input
+ type={type}
+ value={value}
+ placeholder={placeholder}
+ aria-label={`${label} for ${id}`}
+ onChange={(e) => onChange(id, field, e.target.value)}
+ className="rounded border border-border bg-card px-2 py-1 text-sm"
+ />
+ {hint && <span className="text-xs text-muted-foreground">{hint}</span>}
+ </label>
+ );
+}
diff --git a/editor/app/settings/components/SettingsForm.tsx b/editor/app/settings/components/SettingsForm.tsx
@@ -6,11 +6,22 @@ import {
type SaveResult,
} from "../actions";
import type { SiteSettings } from "yt-dlp-transcript-common/lib/settings";
+// From the CLIENT-SAFE digest module, NOT settings.ts's re-exports of them:
+// settings.ts opens with `import fs from "node:fs"`, so pulling the option
+// lists through it drags node:fs into the client bundle and the page dies at
+// runtime. Same rule the digest modules are already split along.
+import {
+ DEFAULT_DIGEST_TIMESTAMP_MODE,
+ DIGEST_SECTION_KINDS as DIGEST_SECTION_OPTIONS,
+ DIGEST_TIMESTAMP_MODES as DIGEST_TIMESTAMP_MODE_OPTIONS,
+} from "yt-dlp-transcript-common/lib/digest";
import {
DOWNLOAD_FORMAT_LABELS,
DOWNLOAD_FORMAT_PRESETS,
} from "yt-dlp-transcript-common/ytdlp/downloadFormat";
import type { TranscriptionAppDescriptor } from "yt-dlp-transcript-common/lib/transcriptionApps";
+import type { DigestAppDescriptor } from "yt-dlp-transcript-common/lib/digestApps";
+import { DigestAppsField } from "./DigestAppsField";
import {
SocialLinksField,
toSocialRow,
@@ -21,9 +32,12 @@ import { WorkersField } from "./WorkersField";
type Props = {
initial: SiteSettings;
apps: TranscriptionAppDescriptor[];
+ // Built on the server: digestApps.ts reaches process.env and imports execa, so
+ // it must never end up in the client bundle. Same reason as `apps`.
+ digestApps: DigestAppDescriptor[];
};
-export function SettingsForm({ initial, apps }: Props) {
+export function SettingsForm({ initial, apps, digestApps }: Props) {
const [state, formAction] = useActionState<SaveResult | undefined, FormData>(
saveSettingsAction,
undefined,
@@ -403,6 +417,160 @@ export function SettingsForm({ initial, apps }: Props) {
/>
</fieldset>
<fieldset className="flex flex-col gap-3 border border-border rounded p-3">
+ <legend className="px-1 text-sm font-medium">Digest</legend>
+ {/*
+ THE MARKER THAT ARMS THE WHOLE BLOCK. Unchecked checkboxes are simply
+ absent from a FormData, so without it a submit from any other form
+ would read `digestRemoteEnabled` as false and silently reset this
+ section. saveSettingsAction gates every field below on it.
+ */}
+ <input type="hidden" name="digestFormPresent" value="1" readOnly />
+ <p className="text-xs text-muted-foreground">
+ AI chapters and topic tags generated from existing transcripts by a
+ local model, written to <code>ai-digest.json</code> beside each one.
+ Run a sweep from a channel's <strong>Digest</strong> stage. Every
+ section records the exact engine, model, prompt version and prompt
+ shape that produced it, and a re-run regenerates only what no longer
+ matches — so changing anything here is what makes the next run redo
+ work, and leaving it alone is what makes a re-run nearly free.
+ </p>
+ <label className="flex flex-col gap-1 text-sm">
+ <span className="font-medium">Local engine</span>
+ <select
+ name="digestLocalAppId"
+ defaultValue={initial.digest.localAppId}
+ className="rounded border border-border bg-card px-2 py-1 text-sm"
+ >
+ {digestApps
+ .filter((a) => !a.metered)
+ .map((a) => (
+ <option key={a.id} value={a.id}>
+ {a.label}
+ </option>
+ ))}
+ </select>
+ <span className="text-xs text-muted-foreground">
+ The lane that carries the corpus. Runs on your own hardware; nothing
+ leaves the machine.
+ </span>
+ </label>
+ <label className="flex flex-col gap-1 text-sm">
+ <span className="font-medium">Sections to generate</span>
+ <span className="flex flex-wrap gap-3">
+ {DIGEST_SECTION_OPTIONS.map((section) => (
+ <label key={section} className="flex items-center gap-1 text-sm">
+ <input
+ type="checkbox"
+ name="digestSections"
+ value={section}
+ defaultChecked={initial.digest.sections.includes(section)}
+ />
+ {section}
+ </label>
+ ))}
+ </span>
+ <span className="text-xs text-muted-foreground">
+ Chapters alone is the default: tags roughly double the model calls
+ for a smaller payoff. Unchecking everything falls back to chapters
+ rather than generating nothing.
+ </span>
+ </label>
+ <label className="flex flex-col gap-1 text-sm">
+ <span className="font-medium">Timestamp mode</span>
+ <select
+ name="digestTimestampMode"
+ defaultValue={initial.digest.timestampMode}
+ className="rounded border border-border bg-card px-2 py-1 text-sm"
+ >
+ {DIGEST_TIMESTAMP_MODE_OPTIONS.map((mode) => (
+ <option key={mode} value={mode}>
+ {mode}
+ {mode === DEFAULT_DIGEST_TIMESTAMP_MODE ? " (recommended)" : ""}
+ </option>
+ ))}
+ </select>
+ <span className="text-xs text-muted-foreground">
+ How each chunk's transcript markers are numbered.{" "}
+ <strong>chunk-local</strong> re-bases every chunk to 00:00:00 and
+ adds the offset back before any guard runs; on the long tail it cut
+ wasted chunks from 29.4% to 11.8% and the worst coverage gap from
+ 1:05:16 to 24:13, because the model stops having to hold a large
+ offset. <strong>absolute</strong> is kept for comparison. This
+ changes the recorded identity, so switching it re-runs the corpus
+ instead of silently mixing two shapes.
+ </span>
+ </label>
+ <Field
+ label="Prompt variant label"
+ name="digestPromptVariant"
+ defaultValue={initial.digest.promptVariant}
+ hint="Free-text label for a non-default prompt shape, recorded in every section's provenance (trimmed, max 40 chars). Setting or changing it invalidates digests generated under a different label — which is exactly what makes a bake-off round re-run its sample instead of skipping it as fresh. Leave blank unless you are running one."
+ />
+ <details className="text-sm">
+ <summary className="cursor-pointer font-medium">
+ Per-engine configuration
+ </summary>
+ <div className="mt-2">
+ <DigestAppsField
+ apps={digestApps}
+ initial={initial.digest.apps}
+ name="digestAppsJson"
+ />
+ </div>
+ </details>
+ <label className="flex items-start gap-2 text-sm">
+ <input
+ type="checkbox"
+ name="digestRemoteEnabled"
+ defaultChecked={initial.digest.remoteEnabled}
+ className="mt-1"
+ />
+ <span className="flex flex-col gap-1">
+ <span className="font-medium">
+ Enable the metered (paid) digest lane
+ </span>
+ <span className="text-xs text-muted-foreground">
+ Off by default, and deliberately: this lane sends transcripts to a
+ paid API. It exists for the long tail of very long videos, which
+ is a small fraction of the corpus by count and a large one by
+ tokens. With it off, the metered option is disabled in every
+ channel's Digest stage rather than failing after the fact.
+ </span>
+ </span>
+ </label>
+ <label className="flex flex-col gap-1 text-sm">
+ <span className="font-medium">Metered engine</span>
+ <select
+ name="digestRemoteAppId"
+ defaultValue={initial.digest.remoteAppId}
+ className="rounded border border-border bg-card px-2 py-1 text-sm"
+ >
+ {digestApps
+ .filter((a) => a.metered)
+ .map((a) => (
+ <option key={a.id} value={a.id}>
+ {a.label}
+ </option>
+ ))}
+ </select>
+ </label>
+ <Field
+ label="Long-tail cutoff (seconds)"
+ name="digestLongTailSeconds"
+ defaultValue={String(initial.digest.longTailSeconds)}
+ type="number"
+ hint="Videos at least this long are what the metered lane takes when you run it. Default 14400 (4 h)."
+ />
+ <Field
+ label="Spend cap (USD per job)"
+ name="digestSpendCapUsd"
+ defaultValue={String(initial.digest.spendCapUsd)}
+ type="number"
+ step="0.01"
+ hint="Hard ceiling on cumulative metered spend within one job; the lane parks itself when it is reached. 0 means no cap. Only ever consulted for a metered engine."
+ />
+ </fieldset>
+ <fieldset className="flex flex-col gap-3 border border-border rounded p-3">
<legend className="px-1 text-sm font-medium">Social links</legend>
<p className="text-xs text-muted-foreground">
Default social links shown in every site's footer. Each site can
@@ -449,6 +617,7 @@ function Field({
hint,
required,
type = "text",
+ step,
}: {
label: string;
name: string;
@@ -456,6 +625,11 @@ function Field({
hint?: string;
required?: boolean;
type?: string;
+ // `type="number"` carries an IMPLICIT step of 1, so a decimal value fails
+ // constraint validation and the browser blocks the submit SILENTLY — no
+ // error, no request, the save just never happens. Any numeric field that
+ // accepts fractions must pass a step.
+ step?: string;
}) {
return (
<label className="flex flex-col gap-1 text-sm">
@@ -465,6 +639,7 @@ function Field({
name={name}
defaultValue={defaultValue}
required={required}
+ step={step}
className="rounded border border-border bg-card px-2 py-1 text-sm"
/>
{hint && <span className="text-xs text-muted-foreground">{hint}</span>}
diff --git a/editor/app/settings/page.tsx b/editor/app/settings/page.tsx
@@ -2,6 +2,7 @@ import type { Metadata } from "next";
import { getPaths } from "yt-dlp-transcript-common/lib/paths";
import { getSettings } from "yt-dlp-transcript-common/lib/settings";
import { listTranscriptionApps } from "yt-dlp-transcript-common/lib/transcriptionApps";
+import { listDigestApps } from "yt-dlp-transcript-common/lib/digestApps";
import { readXSessionStatus } from "yt-dlp-transcript-common/social/xSessionBroker";
import { SettingsForm } from "./components/SettingsForm";
import { XSessionSection } from "./components/XSessionSection";
@@ -42,7 +43,11 @@ export default async function SettingsPage() {
Used by the export build to title the static site.
</p>
</div>
- <SettingsForm initial={settings} apps={listTranscriptionApps()} />
+ <SettingsForm
+ initial={settings}
+ apps={listTranscriptionApps()}
+ digestApps={listDigestApps()}
+ />
</section>
<section className="flex flex-col gap-3 border-t border-border pt-6">
diff --git a/editor/app/widget/components/MonitorWidget.tsx b/editor/app/widget/components/MonitorWidget.tsx
@@ -9,6 +9,10 @@ import type {
DiskStatusView,
} from "../../jobs/active/buildActiveJobs";
import type { RunningJobsListItem } from "../../jobs/components/RunningJobsList";
+import type {
+ JobProgressMetric,
+ JobTaskKind,
+} from "yt-dlp-transcript-common/jobs/registry";
import { jobKindLabel } from "../../jobs/jobKindLabels";
import type { WorkersPayload, WorkerView } from "../../workers/components/WorkersView";
import { InlineActionButton } from "../../actionable/components/InlineActionButton";
@@ -604,6 +608,29 @@ function JobRow({
);
}
+// Per-task-kind glyph/fill, as Records over JobTaskKind so a new kind is a
+// compile error rather than silently rendering as a transcription.
+const TASK_KIND_VERB: Record<JobTaskKind, string> = {
+ download: "\u2193",
+ transcribe: "\u270e",
+ digest: "\u00b6",
+};
+
+const TASK_KIND_FILL: Record<JobTaskKind, string> = {
+ download: "bg-success/60",
+ transcribe: "bg-success",
+ digest: "bg-info",
+};
+
+// Per-metric glyph for the compact widget line. A Record over JobProgressMetric
+// so a new metric is a compile error, not a mislabelled bar (the two copies of
+// this ternary previously had to be kept in sync by hand).
+const METRIC_PREFIX: Record<JobProgressMetric, string> = {
+ downloads: "\u2193 ",
+ transcripts: "",
+ digests: "\u00b6 ",
+};
+
// One-line textual summary of a job's batch progress, e.g. "↓ 5/10 · ~2m left".
// Shared by the job heading (headingProgress) and the job bar's caption so the
// label/ETA formatting stays in one place.
@@ -611,10 +638,7 @@ function jobProgressText(
progress: NonNullable<RunningJobsListItem["progress"]>,
showEta: boolean,
): string {
- const label =
- progress.metric === "downloads"
- ? `↓ ${progress.current}/${progress.target}`
- : `${progress.current}/${progress.target}`;
+ const label = `${METRIC_PREFIX[progress.metric]}${progress.current}/${progress.target}`;
const remaining = progress.target - progress.current;
const etaText =
!showEta || remaining <= 0
@@ -632,10 +656,7 @@ function JobProgressBar({
progress: NonNullable<RunningJobsListItem["progress"]>;
showEta: boolean;
}) {
- const label =
- progress.metric === "downloads"
- ? `↓ ${progress.current}/${progress.target}`
- : `${progress.current}/${progress.target}`;
+ const label = `${METRIC_PREFIX[progress.metric]}${progress.current}/${progress.target}`;
const remaining = progress.target - progress.current;
const etaText =
!showEta || remaining <= 0
@@ -673,7 +694,7 @@ function TaskBar({
}) {
const now = useNow();
const probing = task.phase === "probing";
- const verb = probing ? "🔍" : task.kind === "download" ? "↓" : "✎";
+ const verb = probing ? "🔍" : TASK_KIND_VERB[task.kind];
// While probing, fill against the estimated probe duration (violet "scanning"
// bar) rather than the frozen download fraction. See TaskProgressBar.
const probeFraction =
@@ -694,9 +715,7 @@ function TaskBar({
const pct = hasFraction ? Math.round((fraction as number) * 100) : 0;
const fillClass = probing
? "bg-violet-400 dark:bg-violet-500"
- : task.kind === "download"
- ? "bg-success/60"
- : "bg-success";
+ : TASK_KIND_FILL[task.kind];
const pulseClass = probing
? "bg-violet-400 dark:bg-violet-500"
: "bg-warning";
diff --git a/editor/e2e/digest.spec.ts b/editor/e2e/digest.spec.ts
@@ -0,0 +1,668 @@
+import { test, expect } from "@playwright/test";
+import { writeFile, readFile, rm } from "node:fs/promises";
+import { join } from "node:path";
+// Imported by RELATIVE path, not by package name. The other specs only ever
+// import types from `yt-dlp-transcript-common/...`, which are erased at compile
+// time; `common` publishes no `exports` map, so a runtime import of the package
+// specifier does not resolve under playwright's loader. The relative path does,
+// the same way ./helpers does — and pulling in the REAL effectiveDigest matters
+// here: asserting override survival against a reimplementation of the merge
+// would prove nothing about the merge that ships.
+import {
+ effectiveDigest,
+ type DigestOverrides,
+ type DigestRecord,
+} from "../../common/lib/digest";
+import {
+ pathExists,
+ readJson,
+ resetData,
+ resolvePath,
+ writeChannelConfig,
+ writeDigestVideo,
+ writeSettings,
+} from "./helpers";
+
+// The digest lane, end to end through the real job path.
+//
+// The local engine is an HTTP stub (e2e/fixtures/ollama-stub.mjs, a third
+// playwright webServer) rather than a fake binary, because ollama-direct POSTs
+// to ${ollamaUrl}/api/chat and has no binary to replace. The metered lane's
+// engine IS a subprocess and so uses the ordinary fake-binary idiom
+// (e2e/fixtures/bin/fake-claude.mjs, wired in via CLAUDE_BIN).
+//
+// Post-mutation assertions poll with reload where they read a rendered page:
+// the channel page serves a persisted snapshot regenerated on a ~1s debounce.
+// Assertions that read the SIDECARS go straight to disk and need no polling —
+// the batch"s closing summary line is the happens-before edge.
+
+const CHANNEL = "digest-channel";
+const VIDEO = "digestvid0001";
+
+function digestPath(channel: string, video: string): string {
+ return join("test-transcripts", "channels", channel, "data", video, "ai-digest.json");
+}
+
+function overridesPath(channel: string, video: string): string {
+ return join(
+ "test-transcripts",
+ "channels",
+ channel,
+ "data",
+ video,
+ "ai-digest.overrides.json",
+ );
+}
+
+// Settings with the digest block spelled out. sanitizeDigest fills the rest.
+function digestSettings(over: Record<string, unknown> = {}) {
+ return {
+ adminTitle: "Test Admin",
+ maxTranscriptPageBytes: 8388608,
+ sleepBetweenDownloadsSeconds: 0,
+ minFreeDiskGB: 0,
+ digest: {
+ localAppId: "ollama-direct",
+ remoteAppId: "claude-code",
+ sections: ["chapters"],
+ ...over,
+ },
+ };
+}
+
+// The batch's closing summary line. Waiting on THAT rather than on a generic
+// "succeeded" is what makes these assertions a happens-before edge for the
+// sidecar reads below: the line is emitted after the last write.
+const BATCH_DONE = "Digest batch:";
+
+async function runDigest(page: import("@playwright/test").Page, slug: string) {
+ await page.goto(`/channels/${slug}`);
+ await page.getByRole("button", { name: "Digest channel" }).click();
+ await expect(page.getByLabel("Digest channel output")).toContainText(
+ BATCH_DONE,
+ { timeout: 60_000 },
+ );
+}
+
+// ---------------------------------------------------------------------------
+// 1. Override survival — the single most important test in this file.
+//
+// A full sweep is weeks of wall-clock, so a hand correction that regeneration
+// destroys is work that can never be affordably redone. The two sidecars exist
+// for exactly this, and nothing else in the suite would notice if the generator
+// started writing through to the human file.
+// ---------------------------------------------------------------------------
+
+test("a human override survives a regeneration and reads back as human", async ({
+ page,
+}) => {
+ await resetData(null);
+ await writeSettings(digestSettings());
+ await writeChannelConfig(CHANNEL);
+ await writeDigestVideo({ channelSlug: CHANNEL, videoId: VIDEO });
+
+ await runDigest(page, CHANNEL);
+
+ const first = await readJson<DigestRecord>(digestPath(CHANNEL, VIDEO));
+ const generated = first.sections.chapters?.items ?? [];
+ expect(generated.length).toBeGreaterThan(0);
+ expect(generated[0].decidedBy).toBe("ai");
+
+ // A human retitles the first chapter and suppresses the second, using the
+ // ids the generator produced.
+ const overrides: DigestOverrides = {
+ version: 1,
+ chapters: [
+ {
+ ...generated[0],
+ title: "Hand-written chapter title",
+ decidedBy: "human",
+ },
+ ...(generated[1]
+ ? [{ ...generated[1], decidedBy: "human" as const, enabled: false }]
+ : []),
+ ],
+ note: "corrected by hand in e2e",
+ };
+ await writeFile(
+ resolvePath(overridesPath(CHANNEL, VIDEO)),
+ JSON.stringify(overrides, null, 2),
+ );
+ const overridesBefore = await readFile(
+ resolvePath(overridesPath(CHANNEL, VIDEO)),
+ "utf8",
+ );
+
+ // Force a genuine regeneration by changing the requested model — that is a
+ // real identity change, not a test-only escape hatch, so this exercises the
+ // same path a prompt or model change would take in production.
+ await writeSettings(
+ digestSettings({ apps: { "ollama-direct": { model: "qwen2.5:7b-v2" } } }),
+ );
+ await runDigest(page, CHANNEL);
+
+ const second = await readJson<DigestRecord>(digestPath(CHANNEL, VIDEO));
+ expect(second.sections.chapters?.provenance.modelRequested).toBe(
+ "qwen2.5:7b-v2",
+ );
+
+ // The human file is untouched, byte for byte. The generator must not be able
+ // to reach it at all.
+ const overridesAfter = await readFile(
+ resolvePath(overridesPath(CHANNEL, VIDEO)),
+ "utf8",
+ );
+ expect(overridesAfter).toBe(overridesBefore);
+
+ // And the composed view a reader would see honors it: the retitle wins and
+ // reports decidedBy "human"; the suppressed item is gone without being
+ // deleted from either file.
+ const composed = effectiveDigest(second, overrides);
+ const kept = composed.chapters.find(
+ (c) => c.title === "Hand-written chapter title",
+ );
+ expect(kept).toBeTruthy();
+ expect(kept?.decidedBy).toBe("human");
+ expect(composed.hasOverrides).toBe(true);
+ if (generated[1]) {
+ expect(composed.chapters.some((c) => c.id === generated[1].id)).toBe(false);
+ }
+});
+
+// ---------------------------------------------------------------------------
+// 2. Dedup sharing — the correctness crux.
+//
+// Content similarity says nothing about TIMING. A mirror with a longer intro
+// matches on text at shifted times, so sharing a digest onto it would place
+// every chapter wrong while the artifact looked perfectly healthy.
+// ---------------------------------------------------------------------------
+
+type ClusterFixture = {
+ clusterId: string;
+ members: string[];
+ contained?: boolean;
+ // Pinned explicitly in these fixtures rather than left to pickCanonicalSlug.
+ // The rule's tie-break for equal-duration, equally-transcribed members is
+ // LEXICOGRAPHIC on the slug, which silently made the intended mirror the
+ // canonical and inverted the assertion. Naming it here keeps each test about
+ // the thing it is testing — the alignment gate — instead of about the
+ // tie-break.
+ canonicalSlug?: string;
+};
+
+async function writeDuplicateReport(clusters: ClusterFixture[]) {
+ const report = {
+ version: 1,
+ generatedAt: new Date().toISOString(),
+ runConfig: {
+ thresholdSeconds: null,
+ durationToleranceSeconds: 2,
+ nearThreshold: 0.6,
+ containmentThreshold: 0.8,
+ shingleSize: 5,
+ },
+ totals: {
+ videosScanned: clusters.reduce((a, c) => a + c.members.length, 0),
+ clusters: clusters.length,
+ videosInClusters: clusters.reduce((a, c) => a + c.members.length, 0),
+ },
+ clusters: clusters.map((c) => ({
+ clusterId: c.clusterId,
+ ...(c.canonicalSlug ? { canonicalSlug: c.canonicalSlug } : {}),
+ matchKind: "transcript-exact",
+ score: 1,
+ contained: c.contained ?? false,
+ durationBucket: 300,
+ crossPlatform: true,
+ crossChannel: true,
+ videoRefs: c.members.map((slug) => {
+ const [channelSlug, id] = slug.split("/");
+ return {
+ slug,
+ channelSlug,
+ channel: channelSlug,
+ platform: "youtube",
+ id,
+ title: `Mirror of ${id}`,
+ duration: 600,
+ uploadDate: "20240101",
+ hasTranscript: true,
+ };
+ }),
+ })),
+ };
+ await writeFile(
+ resolvePath(join("test-transcripts", "duplicates.json")),
+ JSON.stringify(report, null, 2),
+ );
+}
+
+test("a digest is shared to an aligned mirror and refused to a shifted one", async ({
+ page,
+}) => {
+ await resetData(null);
+ await writeSettings(digestSettings());
+ await writeChannelConfig(CHANNEL);
+ // canonical + a byte-aligned mirror + a mirror whose cues are shifted 40s by
+ // a longer intro. All three carry identical TEXT — only the timings differ,
+ // which is precisely the case a text-similarity check cannot distinguish.
+ await writeDigestVideo({ channelSlug: CHANNEL, videoId: "canon00000001" });
+ await writeDigestVideo({ channelSlug: CHANNEL, videoId: "aligned000001" });
+ await writeDigestVideo({ channelSlug: CHANNEL, videoId: "canon00000002" });
+ await writeDigestVideo({
+ channelSlug: CHANNEL,
+ videoId: "shifted000001",
+ startOffsetSeconds: 40,
+ });
+
+ // Each cluster gets its OWN canonical member. Clusters partition the corpus,
+ // and DigestClusterPlan.bySlug is keyed by slug — so a video listed in two
+ // clusters would have its role silently overwritten by whichever cluster is
+ // read last, and only one of the two outcomes would ever be observed.
+ await writeDuplicateReport([
+ {
+ clusterId: "cluster-aligned",
+ members: [`${CHANNEL}/canon00000001`, `${CHANNEL}/aligned000001`],
+ canonicalSlug: `${CHANNEL}/canon00000001`,
+ },
+ {
+ clusterId: "cluster-shifted",
+ members: [`${CHANNEL}/canon00000002`, `${CHANNEL}/shifted000001`],
+ canonicalSlug: `${CHANNEL}/canon00000002`,
+ },
+ ]);
+
+ await runDigest(page, CHANNEL);
+
+ const aligned = await readJson<DigestRecord>(
+ digestPath(CHANNEL, "aligned000001"),
+ );
+ expect(aligned.derivedFrom?.slug).toBe(`${CHANNEL}/canon00000001`);
+ // Sharing only happens at near-zero offset, and the record says why it was
+ // considered safe.
+ expect(Math.abs(aligned.derivedFrom?.offsetSeconds ?? 99)).toBeLessThanOrEqual(5);
+ expect(aligned.sections.chapters?.items.length).toBeGreaterThan(0);
+
+ // The shifted mirror was refused the share. It may still have been generated
+ // for in its own right — what must never happen is it CARRYING the canonical
+ // member's digest.
+ const shifted = await pathExists(digestPath(CHANNEL, "shifted000001"));
+ if (shifted) {
+ const rec = await readJson<DigestRecord>(digestPath(CHANNEL, "shifted000001"));
+ expect(rec.derivedFrom).toBeUndefined();
+ }
+});
+
+test("a contained cluster never shares, however well it aligns", async ({
+ page,
+}) => {
+ await resetData(null);
+ await writeSettings(digestSettings());
+ await writeChannelConfig(CHANNEL);
+ await writeDigestVideo({ channelSlug: CHANNEL, videoId: "canon00000001" });
+ await writeDigestVideo({ channelSlug: CHANNEL, videoId: "clipped000001" });
+
+ // Identical timings — this pair would sail through the alignment gate. It is
+ // refused earlier, because containment means one member is a CLIP: the long
+ // video's chapters describe material the clip does not contain.
+ await writeDuplicateReport([
+ {
+ clusterId: "cluster-contained",
+ members: [`${CHANNEL}/canon00000001`, `${CHANNEL}/clipped000001`],
+ canonicalSlug: `${CHANNEL}/canon00000001`,
+ contained: true,
+ },
+ ]);
+
+ await runDigest(page, CHANNEL);
+
+ const clip = await readJson<DigestRecord>(digestPath(CHANNEL, "clipped000001"));
+ expect(clip.derivedFrom).toBeUndefined();
+ // It got its own digest instead of being skipped — a clip is a different
+ // artifact and deserves one.
+ expect(clip.sections.chapters?.items.length).toBeGreaterThan(0);
+});
+
+test("a human canonical override beats the rule", async ({ page }) => {
+ await resetData(null);
+ await writeSettings(digestSettings());
+ await writeChannelConfig(CHANNEL);
+ // Equal duration and both transcribed, so the rule falls through to the
+ // lexicographic tie-break and would pick "aaa…". The human names the other.
+ await writeDigestVideo({ channelSlug: CHANNEL, videoId: "aaa000000001" });
+ await writeDigestVideo({ channelSlug: CHANNEL, videoId: "zzz000000001" });
+
+ await writeDuplicateReport([
+ {
+ clusterId: "cluster-human",
+ members: [`${CHANNEL}/aaa000000001`, `${CHANNEL}/zzz000000001`],
+ },
+ ]);
+ await writeFile(
+ resolvePath(join("test-transcripts", "duplicates.overrides.json")),
+ JSON.stringify({
+ version: 1,
+ clusters: {
+ "cluster-human": {
+ canonicalSlug: `${CHANNEL}/zzz000000001`,
+ decidedAt: new Date().toISOString(),
+ note: "e2e: operator picked the later id",
+ },
+ },
+ }),
+ );
+
+ await runDigest(page, CHANNEL);
+
+ // The human's choice generated; the other received the share.
+ const shared = await readJson<DigestRecord>(digestPath(CHANNEL, "aaa000000001"));
+ expect(shared.derivedFrom?.slug).toBe(`${CHANNEL}/zzz000000001`);
+ const canonical = await readJson<DigestRecord>(
+ digestPath(CHANNEL, "zzz000000001"),
+ );
+ expect(canonical.derivedFrom).toBeUndefined();
+});
+
+// ---------------------------------------------------------------------------
+// 3. The job path
+// ---------------------------------------------------------------------------
+
+test("the Digest stage queues a job and writes the sidecar", async ({ page }) => {
+ await resetData(null);
+ await writeSettings(digestSettings());
+ await writeChannelConfig(CHANNEL);
+ await writeDigestVideo({ channelSlug: CHANNEL, videoId: VIDEO });
+
+ await runDigest(page, CHANNEL);
+
+ await page.goto("/jobs");
+ const firstRow = page.getByRole("row").nth(1);
+ await expect(firstRow).toContainText("digest:local");
+ await expect(firstRow).toContainText("Digest channel (local)");
+
+ const record = await readJson<DigestRecord>(digestPath(CHANNEL, VIDEO));
+ expect(record.sections.chapters?.provenance.appId).toBe("ollama-direct");
+ expect(record.sections.chapters?.provenance.lane).toBe("local-gpu");
+ expect(record.digestSchemaVersion).toBe(1);
+ // A default-configuration record carries NO promptVariant, which is what
+ // keeps every digest written before that field existed fresh.
+ expect(record.sections.chapters?.provenance.promptVariant).toBeUndefined();
+});
+
+test("a second run with unchanged versions is a no-op", async ({ page }) => {
+ await resetData(null);
+ await writeSettings(digestSettings());
+ await writeChannelConfig(CHANNEL);
+ await writeDigestVideo({ channelSlug: CHANNEL, videoId: VIDEO });
+
+ await runDigest(page, CHANNEL);
+ const first = await readJson<DigestRecord>(digestPath(CHANNEL, VIDEO));
+ const generatedAt = first.sections.chapters?.provenance.generatedAt;
+
+ await runDigest(page, CHANNEL);
+ await expect(page.getByLabel("Digest channel output")).toContainText(
+ "1 already current",
+ );
+
+ // The freshness skip is what makes a re-run minutes instead of a second
+ // multi-week sweep, so "nothing was rewritten" is the assertion.
+ const second = await readJson<DigestRecord>(digestPath(CHANNEL, VIDEO));
+ expect(second.sections.chapters?.provenance.generatedAt).toBe(generatedAt);
+});
+
+// ---------------------------------------------------------------------------
+// 4. Two lanes, two queue keys
+// ---------------------------------------------------------------------------
+
+test("the two lanes land on different queue keys", async ({ page }) => {
+ await resetData(null);
+ // longTailSeconds is lowered so the fixture video falls INSIDE the metered
+ // lane's window. digestChannelAction routes the remote lane with
+ // minDurationSeconds: settings.digest.longTailSeconds, because the metered
+ // lane exists for the >4h tail — at the default the 10-minute fixture would be
+ // out of scope and the run would legitimately find nothing to do.
+ await writeSettings(digestSettings({ remoteEnabled: true, longTailSeconds: 60 }));
+ await writeChannelConfig(CHANNEL);
+ await writeDigestVideo({ channelSlug: CHANNEL, videoId: VIDEO });
+
+ await page.goto(`/channels/${CHANNEL}`);
+ await page.getByRole("button", { name: "Digest channel" }).click();
+ await expect(page.getByLabel("Digest channel output")).toContainText(
+ BATCH_DONE,
+ { timeout: 60_000 },
+ );
+
+ // Switch to the metered lane. The control follows the lane's own default key —
+ // if it did not, both lanes would share one key and the registry's
+ // concurrency-1-per-key rule would serialize a GPU lane behind a network one.
+ await page.goto(`/channels/${CHANNEL}`);
+ await page.getByLabel("lane for Digest channel").selectOption("remote");
+ await expect(page.getByLabel("queue for Digest channel")).toHaveValue(
+ "digest:remote",
+ );
+ // The metered lane must actually have work to do, so invalidate the identity.
+ await writeSettings(
+ digestSettings({
+ remoteEnabled: true,
+ longTailSeconds: 60,
+ apps: { "claude-code": { model: "haiku" } },
+ }),
+ );
+ await page.reload();
+ await page.getByLabel("lane for Digest channel").selectOption("remote");
+ await page.getByRole("button", { name: "Digest channel" }).click();
+ await expect(page.getByLabel("Digest channel output")).toContainText(
+ BATCH_DONE,
+ { timeout: 60_000 },
+ );
+
+ await page.goto("/jobs");
+ const rows = page.getByRole("row");
+ await expect(rows.nth(1)).toContainText("digest:remote");
+ await expect(rows.nth(2)).toContainText("digest:local");
+
+ // And the metered lane's own engine really ran: the record carries its lane
+ // and the cost the CLI wrapper reported.
+ const record = await readJson<DigestRecord>(digestPath(CHANNEL, VIDEO));
+ expect(record.sections.chapters?.provenance.lane).toBe("remote-api");
+ expect(record.sections.chapters?.provenance.appId).toBe("claude-code");
+ expect(record.sections.chapters?.provenance.costUsd).toBeGreaterThan(0);
+});
+
+test("the metered lane is refused while it is disabled in settings", async ({
+ page,
+}) => {
+ await resetData(null);
+ await writeSettings(digestSettings({ remoteEnabled: false }));
+ await writeChannelConfig(CHANNEL);
+ await writeDigestVideo({ channelSlug: CHANNEL, videoId: VIDEO });
+
+ await page.goto(`/channels/${CHANNEL}`);
+ const lane = page.getByLabel("lane for Digest channel");
+ // Nothing can spend money until it is explicitly turned on, so the option is
+ // disabled at the control rather than rejected after the click.
+ //
+ // Asserted via the ATTRIBUTE: playwright's toBeDisabled() reports an <option>
+ // as enabled regardless of its disabled attribute (it models form-control
+ // disabled state, and an <option> is not a form control).
+ await expect(lane.locator("option[value='remote']")).toHaveAttribute(
+ "disabled",
+ "",
+ );
+});
+
+// ---------------------------------------------------------------------------
+// 5. The guards, through the real pipeline
+//
+// The unit tests pin each guard in isolation; this pins that the guards are
+// actually WIRED — that a bad engine response reaches warnings[] on disk and
+// that nothing bogus reaches items[].
+// ---------------------------------------------------------------------------
+
+test("bad engine output is recorded in warnings and never reaches items", async ({
+ page,
+}) => {
+ await resetData(null);
+ await writeSettings(digestSettings());
+ await writeChannelConfig(CHANNEL);
+ // The stub keys its bad-output mode off this sentinel in the request body —
+ // the SLOWOP convention the other fakes use.
+ await writeDigestVideo({ channelSlug: CHANNEL, videoId: "badout000001" });
+
+ await runDigest(page, CHANNEL);
+
+ const record = await readJson<DigestRecord>(
+ digestPath(CHANNEL, "badout000001"),
+ );
+ const codes = record.warnings.map((w) => w.code);
+ expect(codes).toContain("malformed-timestamp");
+ expect(codes).toContain("language-drift");
+ expect(codes).toContain("out-of-range");
+
+ // Every rejection names the offending value verbatim, because a multi-week
+ // sweep is only tunable if its failures are inspectable from the artifact.
+ const malformed = record.warnings.find(
+ (w) => w.code === "malformed-timestamp",
+ );
+ expect(malformed?.value).toBe(":00:27");
+ expect(malformed?.section).toBe("chapters");
+
+ const items = record.sections.chapters?.items ?? [];
+ expect(items.length).toBeGreaterThan(0);
+ for (const item of items) {
+ expect(item.clock).toMatch(/^\d\d:\d\d:\d\d$/);
+ // Nothing that failed a guard survived.
+ expect(item.title).not.toMatch(/Malformed stamp|Out of range topic/);
+ expect(item.title).toMatch(/^[\p{L}\p{N}\p{P}\p{Zs}]+$/u);
+ expect(item.start).toBeLessThanOrEqual(600);
+ }
+});
+
+test.afterAll(async () => {
+ // duplicates.json / duplicates.overrides.json live at the transcripts ROOT, so
+ // they survive a channel-scoped reset and would leak into unrelated specs.
+ await rm(resolvePath(join("test-transcripts", "duplicates.json")), {
+ force: true,
+ });
+ await rm(resolvePath(join("test-transcripts", "duplicates.overrides.json")), {
+ force: true,
+ });
+});
+
+// ---------------------------------------------------------------------------
+// 10. The per-video review panel.
+//
+// The backend for this (readVideoDigestAction, saveDigestOverridesAction,
+// digestBucketAction) shipped with Stage B1 and had NO caller — the digest layer
+// produced artifacts nothing could display. A validation run is unreadable while
+// its output is invisible, so these cover the three things the panel exists to
+// show: what was produced, WHICH configuration produced it, and whether that is
+// still the current one.
+// ---------------------------------------------------------------------------
+
+test("the video panel renders a digest with its provenance and warnings", async ({
+ page,
+}) => {
+ await resetData(null);
+ await writeSettings(digestSettings());
+ await writeChannelConfig(CHANNEL);
+ await writeDigestVideo({ channelSlug: CHANNEL, videoId: VIDEO });
+
+ // Empty state first — it is what an operator sees for most of a sweep.
+ await page.goto(`/channels/${CHANNEL}/videos/${VIDEO}`);
+ await expect(page.getByLabel("digest empty")).toBeVisible();
+ await expect(page.getByLabel("digest freshness")).toHaveText("not generated");
+
+ await runDigest(page, CHANNEL);
+
+ await page.goto(`/channels/${CHANNEL}/videos/${VIDEO}`);
+ await expect(page.getByLabel("digest freshness")).toHaveText("current");
+
+ // Chapters, with the clock the model actually emitted.
+ const chapters = page.getByLabel("digest chapters").locator("li");
+ expect(await chapters.count()).toBeGreaterThan(0);
+
+ // Provenance — the whole point during a bake-off-driven validation.
+ const section = page.getByLabel("digest section chapters");
+ await expect(section).toContainText("ollama-direct");
+ await expect(section).toContainText("promptVersion");
+ await expect(section).toContainText("contextHash");
+ await expect(section).toContainText("current");
+
+ // Warnings are always accounted for — either a list or an explicit "none".
+ // `exact` matters: getByLabel is substring-matched by default, so the plain
+ // form would also match "digest warnings empty" and count both states.
+ const hasWarnings = await page
+ .getByLabel("digest warnings", { exact: true })
+ .count();
+ const noWarnings = await page
+ .getByLabel("digest warnings empty", { exact: true })
+ .count();
+ expect(hasWarnings + noWarnings).toBe(1);
+});
+
+test("the video panel shows the digest as stale after a config change", async ({
+ page,
+}) => {
+ await resetData(null);
+ await writeSettings(digestSettings());
+ await writeChannelConfig(CHANNEL);
+ await writeDigestVideo({ channelSlug: CHANNEL, videoId: VIDEO });
+ await runDigest(page, CHANNEL);
+
+ await page.goto(`/channels/${CHANNEL}/videos/${VIDEO}`);
+ await expect(page.getByLabel("digest freshness")).toHaveText("current");
+
+ // A real identity change — the same path a prompt or model change takes.
+ await writeSettings(
+ digestSettings({ apps: { "ollama-direct": { model: "qwen2.5:7b-v2" } } }),
+ );
+
+ await page.goto(`/channels/${CHANNEL}/videos/${VIDEO}`);
+ await expect(page.getByLabel("digest freshness")).toContainText("stale");
+ await expect(page.getByLabel("digest section chapters")).toContainText(
+ "qwen2.5:7b-v2",
+ );
+});
+
+test("corrections saved from the panel land in the overrides sidecar only", async ({
+ page,
+}) => {
+ await resetData(null);
+ await writeSettings(digestSettings());
+ await writeChannelConfig(CHANNEL);
+ await writeDigestVideo({ channelSlug: CHANNEL, videoId: VIDEO });
+ await runDigest(page, CHANNEL);
+
+ const before = await readJson<DigestRecord>(digestPath(CHANNEL, VIDEO));
+ const firstId = (before.sections.chapters?.items ?? [])[0]?.id;
+ expect(firstId).toBeTruthy();
+
+ await page.goto(`/channels/${CHANNEL}/videos/${VIDEO}`);
+ await page.getByText("Correct this digest").click();
+ await page
+ .getByLabel(`title for chapter ${firstId}`)
+ .fill("Renamed from the panel");
+ await page.getByLabel("digest override note").fill("panel e2e");
+ await page.getByRole("button", { name: "Save corrections" }).click();
+ await expect(page.getByRole("status").filter({ hasText: "Saved" })).toBeVisible();
+
+ const overrides = await readJson<DigestOverrides>(
+ overridesPath(CHANNEL, VIDEO),
+ );
+ expect(overrides.chapters?.[0].title).toBe("Renamed from the panel");
+ expect(overrides.chapters?.[0].decidedBy).toBe("human");
+ expect(overrides.note).toBe("panel e2e");
+
+ // The machine file is untouched — the invariant the two sidecars exist for.
+ const after = await readJson<DigestRecord>(digestPath(CHANNEL, VIDEO));
+ expect(after.sections.chapters?.items[0].title).toBe(
+ before.sections.chapters?.items[0].title,
+ );
+
+ // And the composed view reports the edit as human.
+ await page.goto(`/channels/${CHANNEL}/videos/${VIDEO}`);
+ await expect(page.getByLabel("digest chapters")).toContainText(
+ "Renamed from the panel",
+ );
+});
diff --git a/editor/e2e/duplicate-shorts.spec.ts b/editor/e2e/duplicate-shorts.spec.ts
@@ -9,17 +9,26 @@ import { readJson, resetData, resolvePath, writeSite } from "./helpers";
// Drives Build index + Build stats, then runs detection from /actionable and
// asserts on transcripts/duplicates.json.
-type DuplicateRef = { slug: string; channelSlug: string; platform: string };
+type DuplicateRef = {
+ slug: string;
+ channelSlug: string;
+ platform: string;
+ aligned?: boolean;
+ offsetSeconds?: number | null;
+};
type DuplicateCluster = {
+ clusterId: string;
matchKind: string;
score: number | null;
contained: boolean;
crossPlatform: boolean;
crossChannel: boolean;
+ needsReview?: boolean;
+ canonicalSlug?: string;
videoRefs: DuplicateRef[];
};
type DuplicateReport = {
- runConfig: { thresholdSeconds: number | null };
+ runConfig: { thresholdSeconds: number | null; blocking?: string };
totals: { clusters: number };
clusters: DuplicateCluster[];
};
@@ -97,6 +106,9 @@ type Seed = {
duration: number;
transcript: string | null;
plain?: boolean; // emit plain (non-karaoke) VTT
+ // Defaults to a per-id unique title, so the duration-blocking seeds below
+ // never accidentally block by title too. The title-blocking seeds set it.
+ title?: string;
};
const SEEDS: Seed[] = [
@@ -125,8 +137,8 @@ const PLATFORM_KEY: Record<Seed["platform"], string> = {
rumble: "Rumble",
};
-async function seed(): Promise<void> {
- const channels = new Set(SEEDS.map((s) => s.channel));
+async function seed(seeds: Seed[] = SEEDS): Promise<void> {
+ const channels = new Set(seeds.map((s) => s.channel));
for (const channel of channels) {
const dir = resolvePath(`test-transcripts/channels/${channel}`);
await mkdir(dir, { recursive: true });
@@ -135,14 +147,14 @@ async function seed(): Promise<void> {
JSON.stringify({ handling: "transcribe", name: channel }),
);
}
- for (const s of SEEDS) {
+ for (const s of seeds) {
const dir = resolvePath(`test-transcripts/channels/${s.channel}/data/${s.id}`);
await mkdir(dir, { recursive: true });
await writeFile(
`${dir}/metadata.info.json`,
JSON.stringify({
id: s.id,
- title: `Title ${s.id}`,
+ title: s.title ?? `Title ${s.id}`,
channel: s.channel,
upload_date: "20240101",
duration: s.duration,
@@ -235,11 +247,22 @@ test("clusters cross-platform near-duplicate shorts (incl. plain VTT) and ignore
expect(plain!.videoRefs.map((r) => r.slug)).not.toContain("yt-a/plainx");
expect(clusterWith(report, "yt-a/plainx")).toBeUndefined();
- // Every reported cluster is content-confirmed (no metadata-only matches).
+ // Every cluster that would SHIP is content-confirmed. (The original invariant
+ // was "every reported cluster", which was true when duration coincidence
+ // produced nothing at all. Title + near-identical runtime is a far stronger
+ // claim and now does produce clusters — but they are quarantined behind
+ // needsReview, never auto-share, and never reach a built site, so the
+ // guarantee is kept exactly where it matters. These seeds have distinct titles
+ // anyway, so nothing here is a suspect.)
for (const c of report.clusters) {
+ if (c.needsReview) {
+ expect(c.matchKind).toBe("title-duration");
+ continue;
+ }
expect(["transcript-exact", "transcript-near"]).toContain(c.matchKind);
expect(c.score).not.toBeNull();
}
+ expect(report.clusters.filter((c) => c.needsReview)).toHaveLength(0);
// The detected cluster renders on the page after a refresh.
await page.reload();
@@ -276,3 +299,95 @@ test("all-durations run pulls a long video into the short's cluster via containm
expect(cluster!.contained).toBe(true);
expect(cluster!.videoRefs.map((r) => r.slug)).toContain("yt-d/longvid");
});
+
+// ---------------------------------------------------------------------------
+// Title blocking: the pre-filter proposes, the transcript disposes
+// ---------------------------------------------------------------------------
+
+// Three same-title, compatible-runtime pairs that differ only in what the
+// TRANSCRIPTS say. The point of the whole design is that the pre-filter treats
+// all three identically and the content cascade then splits them three ways.
+const TITLE_SEEDS: Seed[] = [
+ // (1) Same title, same content → CONFIRMED. A cross-platform mirror.
+ { channel: "t-a", id: "mirror1", platform: "youtube", duration: 300, transcript: BASE_WORDS, title: "The Weekly Roundup" },
+ { channel: "t-b", id: "mirror2", platform: "rumble", duration: 303, transcript: variant(5, "foxtrotx"), title: "The Weekly Roundup" },
+ // (2) Same title, same runtime, DIFFERENT content → REJECTED. Two episodes of
+ // a daily show. This is the case that makes nominating aggressively safe.
+ { channel: "t-a", id: "epis1", platform: "youtube", duration: 240, transcript: BASE_WORDS, title: "Daily Show Recap" },
+ { channel: "t-b", id: "epis2", platform: "rumble", duration: 241, transcript: UNIQUE_WORDS, title: "Daily Show Recap" },
+ // (3) Same title, same runtime, one side has NO transcript → untestable, so a
+ // needsReview SUSPECT rather than a match or a silent drop.
+ { channel: "t-a", id: "susp1", platform: "youtube", duration: 200, transcript: BASE_WORDS, title: "Archive Upload" },
+ { channel: "t-b", id: "susp2", platform: "rumble", duration: 200, transcript: null, title: "Archive Upload" },
+ // (4) Same title but a genuinely different cut (2x the runtime) → not even
+ // nominated, so the durations gate is doing its job.
+ { channel: "t-a", id: "cut1", platform: "youtube", duration: 300, transcript: BASE_WORDS, title: "Extended Interview" },
+ { channel: "t-b", id: "cut2", platform: "rumble", duration: 700, transcript: BASE_WORDS, title: "Extended Interview" },
+];
+
+test("title blocking confirms matching content, rejects differing content, and quarantines the untestable", async ({
+ page,
+}) => {
+ await resetData(null);
+ await seed(TITLE_SEEDS);
+ await writeSite("testsite", {
+ channels: [...new Set(TITLE_SEEDS.map((s) => s.channel))].map((slug) => ({
+ slug,
+ groupId: "default",
+ })),
+ });
+ await buildData(page);
+
+ // Scope "all" runs corpus-wide, where the default blocking strategy is title.
+ await page.goto("/actionable");
+ await page.getByLabel("duplicate detection scope").selectOption("all");
+ await page.getByRole("button", { name: "detect duplicate shorts" }).click();
+ await expect(page.getByLabel("detect duplicate shorts result")).toBeVisible({
+ timeout: 30_000,
+ });
+
+ const report = await readJson<DuplicateReport>("test-transcripts/duplicates.json");
+ expect(report.runConfig.blocking).toBe("title");
+
+ // (1) CONFIRMED — the content agreed, so this is a real cluster that may share.
+ const mirror = clusterWith(report, "t-a/mirror1");
+ expect(mirror).toBeTruthy();
+ expect(mirror!.needsReview).toBeFalsy();
+ expect(["transcript-exact", "transcript-near"]).toContain(mirror!.matchKind);
+ expect(mirror!.videoRefs.map((r) => r.slug).sort()).toEqual([
+ "t-a/mirror1",
+ "t-b/mirror2",
+ ]);
+ // Alignment is measured and persisted for confirmed clusters — without it the
+ // viewer's "jump to this moment" has nothing honest to key off.
+ expect(mirror!.canonicalSlug).toBeTruthy();
+ for (const ref of mirror!.videoRefs) {
+ expect(typeof ref.aligned).toBe("boolean");
+ }
+
+ // (2) REJECTED — same title, same runtime, different words. No cluster at all.
+ expect(clusterWith(report, "t-a/epis1")).toBeUndefined();
+ expect(clusterWith(report, "t-b/epis2")).toBeUndefined();
+
+ // (3) SUSPECT — nothing could compare the content, so it is a review item.
+ const suspect = clusterWith(report, "t-a/susp1");
+ expect(suspect).toBeTruthy();
+ expect(suspect!.needsReview).toBe(true);
+ expect(suspect!.matchKind).toBe("title-duration");
+ expect(suspect!.score).toBeNull();
+ expect(suspect!.videoRefs.map((r) => r.slug).sort()).toEqual([
+ "t-a/susp1",
+ "t-b/susp2",
+ ]);
+
+ // (4) NOT NOMINATED — a 300s and a 700s video are a different cut, not a
+ // mirror, however identical their titles.
+ expect(clusterWith(report, "t-a/cut1")).toBeUndefined();
+ expect(clusterWith(report, "t-b/cut2")).toBeUndefined();
+
+ // The suspect is visibly flagged in the editor's review list.
+ await page.reload();
+ const card = page.getByLabel(`duplicate cluster ${suspect!.clusterId}`);
+ await expect(card).toContainText("needs review");
+ await expect(card).toContainText("title + runtime");
+});
diff --git a/editor/e2e/fixtures/bin/fake-claude.mjs b/editor/e2e/fixtures/bin/fake-claude.mjs
@@ -0,0 +1,86 @@
+#!/usr/bin/env node
+// E2E fake `claude` CLI — the metered digest lane's engine.
+//
+// Unlike the ollama stub (an HTTP server, because ollama-direct talks HTTP),
+// claude-code IS a subprocess, so this follows the existing fake-binary idiom:
+// a .mjs in e2e/fixtures/bin/ wired in through an env var (CLAUDE_BIN).
+//
+// Mimics what digestApps.ts's claudeCode.run() actually invokes and parses:
+// claude -p --output-format json [--model <m>] with the prompt on STDIN
+// and expects on stdout the CLI's own JSON WRAPPER, whose `result` field is a
+// STRING containing the model's JSON body. Reproducing both layers matters —
+// the double parse is exactly where a real integration breaks.
+//
+// There is no constrained decoding over the CLI, so this deliberately emits the
+// contract by hand, the same way the real lane depends on the parser's guards
+// rather than on a schema.
+
+import { readFileSync } from "node:fs";
+
+const argv = process.argv.slice(2);
+function flag(name) {
+ const i = argv.indexOf(name);
+ return i < 0 ? undefined : argv[i + 1];
+}
+
+// The reachability probe digestBatch runs before touching 74k videos — it
+// shells out to `claude --version` and treats a non-zero exit as "engine down".
+// Answering it is not optional: without this the metered lane fails at the probe
+// and never reaches the code the rest of this fake exists to exercise.
+if (argv.includes("--version")) {
+ process.stdout.write("1.0.0 (fake-claude)\n");
+ process.exit(0);
+}
+
+// Fail loudly if the invocation drifts: a fake that accepts anything stops
+// testing the thing it exists to test.
+if (!argv.includes("-p") || flag("--output-format") !== "json") {
+ process.stderr.write(
+ `[fake-claude] unexpected argv: ${JSON.stringify(argv)}\n`,
+ );
+ process.exit(2);
+}
+
+let stdin = "";
+try {
+ stdin = readFileSync(0, "utf8");
+} catch {
+ stdin = "";
+}
+
+function hms(total) {
+ const n = Math.max(0, Math.floor(total));
+ return [Math.floor(n / 3600), Math.floor((n % 3600) / 60), n % 60]
+ .map((v) => String(v).padStart(2, "0"))
+ .join(":");
+}
+
+function toSeconds(clock) {
+ const m = /^(\d\d):(\d\d):(\d\d)$/.exec(clock);
+ return m ? Number(m[1]) * 3600 + Number(m[2]) * 60 + Number(m[3]) : null;
+}
+
+const m = /This section covers (\d\d:\d\d:\d\d) to (\d\d:\d\d:\d\d)/.exec(stdin);
+const start = m ? (toSeconds(m[1]) ?? 0) : 0;
+const end = m ? (toSeconds(m[2]) ?? 600) : 600;
+const span = Math.max(1, end - start);
+
+const body = /"tags"/.test(stdin)
+ ? { tags: ["metered lane", "court filings", "podcast"] }
+ : {
+ chapters: [
+ { start: hms(start), title: "Metered lane opening remarks" },
+ { start: hms(start + Math.floor(span / 3)), title: "Ruling analysis" },
+ ],
+ };
+
+process.stdout.write(
+ JSON.stringify({
+ type: "result",
+ is_error: false,
+ // The inner body is a STRING, exactly as the real CLI reports it.
+ result: JSON.stringify(body),
+ total_cost_usd: 0.0123,
+ usage: { input_tokens: 1000, output_tokens: 120 },
+ }),
+);
diff --git a/editor/e2e/fixtures/ollama-stub.mjs b/editor/e2e/fixtures/ollama-stub.mjs
@@ -0,0 +1,152 @@
+#!/usr/bin/env node
+// A stand-in for the ollama HTTP API, for e2e.
+//
+// WHY THIS IS A SERVER AND NOT A FAKE BINARY. Every other engine in this repo is
+// a subprocess, so the e2e idiom is a fake executable in e2e/fixtures/bin/ wired
+// in through an env var. `ollama-direct` is different: it POSTs to
+// `${ollamaUrl}/api/chat` (common/lib/digestApps.ts), so there is no binary to
+// replace. It therefore runs as a third playwright `webServer` and the editor is
+// pointed at it with OLLAMA_URL.
+//
+// It answers from the PROMPT, not from a canned script, so it exercises the real
+// contract: the chunk range the model is told to stay inside is parsed back out
+// of the prompt text and every emitted start is placed within it. That means the
+// stub works unchanged for both timestamp modes — in chunk-local mode the prompt
+// states 00:00:00–00:44:26 and the stub answers in that numbering, which is
+// exactly what a real model does and what the parser must offset back.
+//
+// DETERMINISTIC BAD-OUTPUT MODE. When the request carries the BADOUT sentinel
+// (the SLOWOP convention from fake-whisper.mjs et al — matched case-insensitively
+// anywhere in the request body, so it can be carried by a video id or a title),
+// the stub emits one valid chapter plus one of each measured failure: a malformed
+// stamp, a non-Latin title, and an out-of-range start. One VALID chapter is
+// included on purpose: digestVideo deliberately leaves the sidecar untouched when
+// a section yields nothing (so the video retries), so an all-bad response would
+// write no file and there would be no warnings[] on disk to assert against.
+
+import { createServer } from "node:http";
+
+const PORT = Number(process.env.OLLAMA_STUB_PORT ?? 11435);
+const MODEL = process.env.OLLAMA_STUB_MODEL ?? "qwen2.5:7b";
+const SENTINEL = /badout/i;
+
+function hms(total) {
+ const n = Math.max(0, Math.floor(total));
+ return [Math.floor(n / 3600), Math.floor((n % 3600) / 60), n % 60]
+ .map((v) => String(v).padStart(2, "0"))
+ .join(":");
+}
+
+function toSeconds(clock) {
+ const m = /^(\d\d):(\d\d):(\d\d)$/.exec(clock);
+ if (!m) return null;
+ return Number(m[1]) * 3600 + Number(m[2]) * 60 + Number(m[3]);
+}
+
+// The prompt states its own range ("This section covers HH:MM:SS to HH:MM:SS"),
+// which is the one piece of state the stub needs to answer plausibly.
+function rangeFromPrompt(prompt) {
+ const m = /This section covers (\d\d:\d\d:\d\d) to (\d\d:\d\d:\d\d)/.exec(
+ prompt ?? "",
+ );
+ if (!m) return { start: 0, end: 600 };
+ return { start: toSeconds(m[1]) ?? 0, end: toSeconds(m[2]) ?? 600 };
+}
+
+// Titles are concrete noun phrases, not "Discussion" — the stub should not be
+// the thing that fails a generic-title assertion.
+const TITLES = [
+ "Court filing deadlines",
+ "Sponsor read and housekeeping",
+ "Audience questions on the ruling",
+ "Closing arguments recap",
+];
+
+function goodChapters({ start, end }) {
+ const span = Math.max(1, end - start);
+ const out = [];
+ // Three evenly spaced starts, the last comfortably inside the range so a
+ // rounding difference can never push it past the clamp.
+ for (let i = 0; i < 3; i++) {
+ const at = start + Math.floor((span * i) / 4);
+ out.push({ start: hms(at), title: TITLES[i % TITLES.length] });
+ }
+ return out;
+}
+
+function badChapters(range) {
+ const [first] = goodChapters(range);
+ return [
+ // Survives every guard, so the section is written and warnings[] lands on
+ // disk where a spec can read it.
+ first,
+ // GUARD 1 — the exact malformed shape the naive prompt produced.
+ { start: ":00:27", title: "Malformed stamp" },
+ // GUARD 2 — language drift.
+ { start: hms(range.start + 5), title: "法廷の締め切り" },
+ // GUARD 3 — an hour past the end of this chunk's range.
+ { start: hms(range.end + 3600), title: "Out of range topic" },
+ ];
+}
+
+function readBody(req) {
+ return new Promise((resolve, reject) => {
+ let raw = "";
+ req.on("data", (c) => {
+ raw += c;
+ });
+ req.on("end", () => resolve(raw));
+ req.on("error", reject);
+ });
+}
+
+const server = createServer(async (req, res) => {
+ const url = req.url ?? "/";
+
+ // The reachability probe digestBatch runs before touching 74k videos.
+ if (req.method === "GET" && url.startsWith("/api/tags")) {
+ res.writeHead(200, { "content-type": "application/json" });
+ res.end(JSON.stringify({ models: [{ name: MODEL, model: MODEL }] }));
+ return;
+ }
+
+ if (req.method === "POST" && url.startsWith("/api/chat")) {
+ const raw = await readBody(req);
+ let body = {};
+ try {
+ body = JSON.parse(raw);
+ } catch {
+ /* fall through to the default range */
+ }
+ const messages = Array.isArray(body.messages) ? body.messages : [];
+ const prompt = messages.map((m) => m?.content ?? "").join("\n");
+ const range = rangeFromPrompt(prompt);
+ const bad = SENTINEL.test(raw);
+
+ // Tags ask for a `tags` array; chapters ask for `chapters`. Answer whichever
+ // the caller's schema names, so the stub covers both sections.
+ const wantsTags = Boolean(body.format?.properties?.tags);
+ const data = wantsTags
+ ? { tags: ["court filings", "podcast", "legal news"] }
+ : { chapters: bad ? badChapters(range) : goodChapters(range) };
+
+ res.writeHead(200, { "content-type": "application/json" });
+ res.end(
+ JSON.stringify({
+ model: body.model || MODEL,
+ message: { role: "assistant", content: JSON.stringify(data) },
+ done: true,
+ prompt_eval_count: 100,
+ eval_count: 50,
+ }),
+ );
+ return;
+ }
+
+ res.writeHead(404, { "content-type": "application/json" });
+ res.end(JSON.stringify({ error: `no stub route for ${req.method} ${url}` }));
+});
+
+server.listen(PORT, "127.0.0.1", () => {
+ console.log(`ollama stub listening on http://127.0.0.1:${PORT}`);
+});
diff --git a/editor/e2e/helpers.ts b/editor/e2e/helpers.ts
@@ -6,6 +6,7 @@ import {
readFile,
rm,
stat,
+ utimes,
writeFile,
} from "node:fs/promises";
import { baseUrl } from "./baseUrl";
@@ -107,3 +108,157 @@ export async function copyFixture(name: string) {
await mkdir(dst, { recursive: true });
await cp(testTranscriptsDir, dst, { recursive: true });
}
+
+// ---------------------------------------------------------------------------
+// Digest fixtures
+// ---------------------------------------------------------------------------
+
+// Write a video that the digest lane will actually accept.
+//
+// digestVideo refuses to run unless transcript.cues.json is FRESH — at least as
+// new as both metadata.info.json and the raw transcript (isCuesJsonFresh) — so a
+// digest never describes text that is about to be rewritten. Copying a fixture
+// tree can't guarantee that ordering, because `cp` stamps every file with the
+// time of the copy and the resulting order is whatever the walk produced. So the
+// mtimes are set EXPLICITLY here: the sources are backdated and cues.json is
+// left at "now". Without this the whole digest suite fails intermittently with
+// "stale-cues", which reads like a product bug and isn't one.
+export async function writeDigestVideo(opts: {
+ channelSlug: string;
+ videoId: string;
+ title?: string;
+ // Total video length in seconds. Cues are laid down every 5s across it.
+ durationSeconds?: number;
+ channelName?: string;
+ // Shift every cue by this many seconds while keeping the TEXT identical — a
+ // mirror with a longer intro. The alignment gate must refuse to share a digest
+ // onto one of these: the content matches, so a text-only check would pass it,
+ // and every shared chapter would then land at the wrong moment while the
+ // artifact looked perfectly healthy.
+ startOffsetSeconds?: number;
+}) {
+ const {
+ channelSlug,
+ videoId,
+ title = "Synthetic Digest Video",
+ durationSeconds = 600,
+ channelName = channelSlug,
+ startOffsetSeconds = 0,
+ } = opts;
+ const dir = join(testTranscriptsDir, "channels", channelSlug, "data", videoId);
+ await mkdir(dir, { recursive: true });
+
+ // Every cue's words are GLOBALLY UNIQUE. measureAlignment anchors on 8-word
+ // n-grams that occur exactly once on each side, so the obvious fixture — the
+ // same sentence in every cue with only the number changed — yields no unique
+ // anchors at all and reports "too-few-anchors" instead of the offset result a
+ // sharing test is actually trying to observe.
+ const cues = [];
+ let i = 0;
+ for (let t = 0; t + 5 <= durationSeconds; t += 5, i++) {
+ const words = [
+ "segment",
+ "filing",
+ "deadline",
+ "schedule",
+ "ruling",
+ "hearing",
+ "motion",
+ "brief",
+ "docket",
+ "counsel",
+ "exhibit",
+ "transcript",
+ ].map((w) => `${w}${i}`);
+ cues.push({
+ start: t + startOffsetSeconds,
+ end: t + 5 + startOffsetSeconds,
+ text: words.join(" "),
+ });
+ }
+
+ const meta = {
+ id: videoId,
+ title,
+ channel: channelName,
+ channel_id: `UC${channelSlug}`,
+ uploader: channelName,
+ upload_date: "20240101",
+ duration: durationSeconds + startOffsetSeconds,
+ description: "",
+ is_live: false,
+ was_live: false,
+ live_status: "not_live",
+ age_limit: 0,
+ extractor_key: "Youtube",
+ webpage_url: `https://www.youtube.com/watch?v=${videoId}`,
+ };
+ const vtt = [
+ "WEBVTT",
+ "Kind: captions",
+ "Language: en",
+ "",
+ ...cues.flatMap((c) => [
+ `${vttStamp(c.start)} --> ${vttStamp(c.end)}`,
+ c.text,
+ "",
+ ]),
+ ].join("\n");
+
+ const detail = {
+ version: 2,
+ source: "vtt",
+ transcriptFormat: "vtt",
+ id: videoId,
+ slug: `${channelSlug}/${videoId}`,
+ channelSlug,
+ channel: channelName,
+ title,
+ uploadDate: "20240101",
+ duration: durationSeconds + startOffsetSeconds,
+ isLivestream: false,
+ cues,
+ };
+
+ const metaPath = join(dir, "metadata.info.json");
+ const vttPath = join(dir, "transcript.en.vtt");
+ const cuesPath = join(dir, "transcript.cues.json");
+ await writeFile(metaPath, JSON.stringify(meta, null, 2));
+ await writeFile(vttPath, vtt);
+ await writeFile(cuesPath, JSON.stringify(detail));
+
+ const older = new Date(Date.now() - 60_000);
+ await utimes(metaPath, older, older);
+ await utimes(vttPath, older, older);
+ return { dir, cuesPath, cues };
+}
+
+function vttStamp(seconds: number): string {
+ const n = Math.max(0, Math.floor(seconds));
+ const h = String(Math.floor(n / 3600)).padStart(2, "0");
+ const m = String(Math.floor((n % 3600) / 60)).padStart(2, "0");
+ const s = String(n % 60).padStart(2, "0");
+ return `${h}:${m}:${s}.000`;
+}
+
+// Minimal channel config so the channel page renders and the batch can run.
+export async function writeChannelConfig(
+ channelSlug: string,
+ config: Record<string, unknown> = {},
+) {
+ const dir = join(testTranscriptsDir, "channels", channelSlug);
+ await mkdir(dir, { recursive: true });
+ await writeFile(
+ join(dir, "config.json"),
+ JSON.stringify(
+ {
+ handling: config.handling ?? "youtube",
+ name: config.name ?? channelSlug,
+ url: config.url ?? "https://www.youtube.com/@example/videos",
+ ...config,
+ },
+ null,
+ 2,
+ ),
+ );
+}
diff --git a/editor/e2e/settings.spec.ts b/editor/e2e/settings.spec.ts
@@ -165,3 +165,92 @@ test("saves build pipeline settings (mode, concurrency, image)", async ({
.getByRole("button", { name: "Docker" }),
).toHaveAttribute("aria-pressed", "true");
});
+
+// ---------------------------------------------------------------------------
+// Digest
+// ---------------------------------------------------------------------------
+//
+// Every digest knob used to be unreachable from a browser: saveSettingsAction
+// had a full digest branch gated on a `digestFormPresent` marker that NOTHING in
+// the repo emitted, while two places told the operator to "enable it in
+// Settings → Digest" — a section that did not exist.
+
+test("digest settings round-trip through the form", async ({ page }) => {
+ await page.goto("/settings");
+
+ await page.getByLabel("Timestamp mode").selectOption("absolute");
+ await page.getByLabel(/prompt variant label/i).fill("bakeoff-r3");
+ await page.getByLabel(/spend cap/i).fill("12.5");
+ await page.getByLabel(/long-tail cutoff/i).fill("7200");
+ await page
+ .getByRole("checkbox", { name: /enable the metered/i })
+ .check();
+ // Per-app numCtx rides in the hidden digestAppsJson payload.
+ await page.getByRole("group", { name: "Digest" }).getByText(
+ "Per-engine configuration",
+ ).click();
+ await page.getByLabel("Context window (num_ctx) for ollama-direct").fill("4096");
+
+ await page.getByRole("button", { name: /save settings/i }).click();
+ await expect(
+ page.getByRole("status").filter({ hasText: "Saved" }),
+ ).toBeVisible();
+
+ const saved = await readJson<{
+ digest?: {
+ timestampMode: string;
+ promptVariant: string;
+ spendCapUsd: number;
+ longTailSeconds: number;
+ remoteEnabled: boolean;
+ sections: string[];
+ apps: Record<string, { numCtx?: number }>;
+ };
+ }>("test-settings.json");
+ expect(saved.digest?.timestampMode).toBe("absolute");
+ expect(saved.digest?.promptVariant).toBe("bakeoff-r3");
+ expect(saved.digest?.spendCapUsd).toBe(12.5);
+ expect(saved.digest?.longTailSeconds).toBe(7200);
+ expect(saved.digest?.remoteEnabled).toBe(true);
+ expect(saved.digest?.apps["ollama-direct"]?.numCtx).toBe(4096);
+ // Never left empty — an empty section list would generate nothing.
+ expect(saved.digest?.sections.length).toBeGreaterThan(0);
+});
+
+test("an unrelated settings save does not reset the digest prompt shape", async ({
+ page,
+}) => {
+ // THE BUG THIS EXISTS FOR. The digest branch REBUILDS the whole block on every
+ // save, so any freshness-affecting field it fails to carry through is silently
+ // reset — and timestampMode/promptVariant are folded into the recorded
+ // identity, so a reset would invalidate every digest generated under the
+ // non-default shape while looking like a no-op.
+ await page.goto("/settings");
+ await page.getByLabel("Timestamp mode").selectOption("absolute");
+ await page.getByLabel(/prompt variant label/i).fill("keepme");
+ await page.getByRole("button", { name: /save settings/i }).click();
+ await expect(
+ page.getByRole("status").filter({ hasText: "Saved" }),
+ ).toBeVisible();
+
+ // Reload rather than submitting straight again: React 19 resets <form action>
+ // inputs to the defaultValue of the render they were in, so a second submit
+ // from the same render would re-post the digest values from BEFORE the first
+ // save and prove nothing. A reload is also the real scenario — the operator
+ // sets the digest up, comes back later, and changes something unrelated.
+ await page.reload();
+ await expect(page.getByLabel("Timestamp mode")).toHaveValue("absolute");
+ await page.getByLabel(/admin title/i).fill("Unrelated Change");
+ await page.getByRole("button", { name: /save settings/i }).click();
+ await expect(
+ page.getByRole("status").filter({ hasText: "Saved" }),
+ ).toBeVisible();
+
+ const saved = await readJson<{
+ adminTitle: string;
+ digest?: { timestampMode: string; promptVariant: string };
+ }>("test-settings.json");
+ expect(saved.adminTitle).toBe("Unrelated Change");
+ expect(saved.digest?.timestampMode).toBe("absolute");
+ expect(saved.digest?.promptVariant).toBe("keepme");
+});
diff --git a/editor/instrumentation.ts b/editor/instrumentation.ts
@@ -46,4 +46,24 @@ export async function register() {
} catch {
/* a runner that fails to start must not block server readiness */
}
+
+ // Resume the corpus-wide digest sweep, if one is armed. A sweep is GPU-WEEKS
+ // long, so it will outlive several restarts by construction — and before this
+ // hook a restart silently ended one that had been running for days, with
+ // nothing to say so.
+ //
+ // Re-launching is safe rather than merely convenient: the sweep stores no
+ // cursor and the batch re-derives eligibility from disk on every pull, so a
+ // resumed sweep re-does exactly zero work. The digest PAUSE needs no hook
+ // here for the same reason downloadsPaused doesn't — it is read at dispatch
+ // time — but "is a sweep running" is process state, and process state is what
+ // a restart destroys.
+ try {
+ const { resumeDigestSweepIfEnabled } = await import(
+ "yt-dlp-transcript-common/controller/digestSweep"
+ );
+ await resumeDigestSweepIfEnabled();
+ } catch {
+ /* a sweep that fails to resume must not block server readiness */
+ }
}
diff --git a/editor/package.json b/editor/package.json
@@ -5,8 +5,8 @@
"type": "module",
"scripts": {
"dev": "next dev --port ${EDITOR_PORT:-3001}",
- "dev:test": "WORKER_TOKEN=test-worker-token TRANSCRIPTS_DIR=$(pwd)/test-transcripts EXPORT_PUBLIC_DIR=$(pwd)/test-transcripts/.export-public SETTINGS_FILE=$(pwd)/test-settings.json YTDLP_BIN=$(pwd)/e2e/fixtures/bin/fake-ytdlp.mjs GALLERY_DL_BIN=$(pwd)/e2e/fixtures/bin/fake-gallery-dl.mjs WHISPER_BIN=$(pwd)/e2e/fixtures/bin/fake-whisper.mjs WHISPER_MODEL=/dev/null CHOUGH_BIN=$(pwd)/e2e/fixtures/bin/fake-chough.mjs CHOUGH_MODEL=/dev/null PARAKEET_STITCH_BIN=$(pwd)/e2e/fixtures/bin/fake-parakeet-stitch.mjs PARAKEET_CLI=/dev/null PARAKEET_MODEL=/dev/null FFMPEG_BIN=$(pwd)/e2e/fixtures/bin/fake-ffmpeg.mjs FFPROBE_BIN=$(pwd)/e2e/fixtures/bin/fake-ffprobe.mjs AUDIO_CHECK_INTERVAL_MS_OVERRIDE=300 AUDIO_CHECK_SIZE_GATE_OVERRIDE=4096 AUDIO_CHECK_INTERVAL_FLOOR_MS_OVERRIDE=50 AUDIO_CHECK_RECOVER_STEP_MS_OVERRIDE=100 AUDIO_CHECK_RECOVER_AFTER_OVERRIDE=2 next dev --port ${PORT:-3011}",
- "start:test": "WORKER_TOKEN=test-worker-token TRANSCRIPTS_DIR=$(pwd)/test-transcripts EXPORT_PUBLIC_DIR=$(pwd)/test-transcripts/.export-public SETTINGS_FILE=$(pwd)/test-settings.json YTDLP_BIN=$(pwd)/e2e/fixtures/bin/fake-ytdlp.mjs GALLERY_DL_BIN=$(pwd)/e2e/fixtures/bin/fake-gallery-dl.mjs WHISPER_BIN=$(pwd)/e2e/fixtures/bin/fake-whisper.mjs WHISPER_MODEL=/dev/null CHOUGH_BIN=$(pwd)/e2e/fixtures/bin/fake-chough.mjs CHOUGH_MODEL=/dev/null PARAKEET_STITCH_BIN=$(pwd)/e2e/fixtures/bin/fake-parakeet-stitch.mjs PARAKEET_CLI=/dev/null PARAKEET_MODEL=/dev/null FFMPEG_BIN=$(pwd)/e2e/fixtures/bin/fake-ffmpeg.mjs FFPROBE_BIN=$(pwd)/e2e/fixtures/bin/fake-ffprobe.mjs AUDIO_CHECK_INTERVAL_MS_OVERRIDE=300 AUDIO_CHECK_SIZE_GATE_OVERRIDE=4096 AUDIO_CHECK_INTERVAL_FLOOR_MS_OVERRIDE=50 AUDIO_CHECK_RECOVER_STEP_MS_OVERRIDE=100 AUDIO_CHECK_RECOVER_AFTER_OVERRIDE=2 next start --port ${PORT:-3011}",
+ "dev:test": "WORKER_TOKEN=test-worker-token TRANSCRIPTS_DIR=$(pwd)/test-transcripts EXPORT_PUBLIC_DIR=$(pwd)/test-transcripts/.export-public SETTINGS_FILE=$(pwd)/test-settings.json YTDLP_BIN=$(pwd)/e2e/fixtures/bin/fake-ytdlp.mjs GALLERY_DL_BIN=$(pwd)/e2e/fixtures/bin/fake-gallery-dl.mjs WHISPER_BIN=$(pwd)/e2e/fixtures/bin/fake-whisper.mjs WHISPER_MODEL=/dev/null CHOUGH_BIN=$(pwd)/e2e/fixtures/bin/fake-chough.mjs CHOUGH_MODEL=/dev/null PARAKEET_STITCH_BIN=$(pwd)/e2e/fixtures/bin/fake-parakeet-stitch.mjs PARAKEET_CLI=/dev/null PARAKEET_MODEL=/dev/null FFMPEG_BIN=$(pwd)/e2e/fixtures/bin/fake-ffmpeg.mjs FFPROBE_BIN=$(pwd)/e2e/fixtures/bin/fake-ffprobe.mjs OLLAMA_URL=http://127.0.0.1:${OLLAMA_STUB_PORT:-11435} CLAUDE_BIN=$(pwd)/e2e/fixtures/bin/fake-claude.mjs AUDIO_CHECK_INTERVAL_MS_OVERRIDE=300 AUDIO_CHECK_SIZE_GATE_OVERRIDE=4096 AUDIO_CHECK_INTERVAL_FLOOR_MS_OVERRIDE=50 AUDIO_CHECK_RECOVER_STEP_MS_OVERRIDE=100 AUDIO_CHECK_RECOVER_AFTER_OVERRIDE=2 next dev --port ${PORT:-3011}",
+ "start:test": "WORKER_TOKEN=test-worker-token TRANSCRIPTS_DIR=$(pwd)/test-transcripts EXPORT_PUBLIC_DIR=$(pwd)/test-transcripts/.export-public SETTINGS_FILE=$(pwd)/test-settings.json YTDLP_BIN=$(pwd)/e2e/fixtures/bin/fake-ytdlp.mjs GALLERY_DL_BIN=$(pwd)/e2e/fixtures/bin/fake-gallery-dl.mjs WHISPER_BIN=$(pwd)/e2e/fixtures/bin/fake-whisper.mjs WHISPER_MODEL=/dev/null CHOUGH_BIN=$(pwd)/e2e/fixtures/bin/fake-chough.mjs CHOUGH_MODEL=/dev/null PARAKEET_STITCH_BIN=$(pwd)/e2e/fixtures/bin/fake-parakeet-stitch.mjs PARAKEET_CLI=/dev/null PARAKEET_MODEL=/dev/null FFMPEG_BIN=$(pwd)/e2e/fixtures/bin/fake-ffmpeg.mjs FFPROBE_BIN=$(pwd)/e2e/fixtures/bin/fake-ffprobe.mjs OLLAMA_URL=http://127.0.0.1:${OLLAMA_STUB_PORT:-11435} CLAUDE_BIN=$(pwd)/e2e/fixtures/bin/fake-claude.mjs AUDIO_CHECK_INTERVAL_MS_OVERRIDE=300 AUDIO_CHECK_SIZE_GATE_OVERRIDE=4096 AUDIO_CHECK_INTERVAL_FLOOR_MS_OVERRIDE=50 AUDIO_CHECK_RECOVER_STEP_MS_OVERRIDE=100 AUDIO_CHECK_RECOVER_AFTER_OVERRIDE=2 next start --port ${PORT:-3011}",
"build": "next build",
"start": "next start --port ${EDITOR_PORT:-3001}",
"lint": "eslint",
diff --git a/editor/playwright.config.ts b/editor/playwright.config.ts
@@ -10,6 +10,13 @@ process.env.PLAYWRIGHT_BASE_URL = baseURL;
const webServerCommand =
process.env.E2E_MODE === "start" ? "pnpm start:test" : "pnpm dev:test";
+// The digest lane's local engine is reached over HTTP, not spawned, so it gets a
+// webServer entry instead of a fake binary in e2e/fixtures/bin/. Its port is
+// exported so the editor's dev:test / start:test scripts point OLLAMA_URL at the
+// same place, and so an offset-port worktree does not collide.
+const OLLAMA_STUB_PORT = Number(process.env.OLLAMA_STUB_PORT ?? 11435);
+process.env.OLLAMA_STUB_PORT = String(OLLAMA_STUB_PORT);
+
const EXPORT_PORT = Number(process.env.EXPORT_PORT ?? 3010);
const exportBaseURL = `http://localhost:${EXPORT_PORT}`;
// Share the editor's test-settings.json with the export server so the
@@ -54,6 +61,15 @@ export default defineConfig({
SITE_ID: "testsite",
},
},
+ {
+ command: `node ${path.resolve(process.cwd(), "e2e", "fixtures", "ollama-stub.mjs")}`,
+ // /api/tags is the same endpoint digestApps.ts probes, so playwright's
+ // readiness check and the app's own reachability check agree.
+ url: `http://127.0.0.1:${OLLAMA_STUB_PORT}/api/tags`,
+ timeout: 30_000,
+ reuseExistingServer: !process.env.CI,
+ env: { OLLAMA_STUB_PORT: String(OLLAMA_STUB_PORT) },
+ },
],
use: {
baseURL,
diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md
@@ -1,6 +1,9 @@
# Changelog
## [Unreleased]
+- **The archive now recognises far more cross-platform re-uploads as the same video.** Duplicate detection compared two transcripts and called them the same only above a similarity of 0.6 — a threshold tuned for two transcripts of the same *text*, which quietly failed the case detection exists for. The two sides of a YouTube↔Rumble mirror are transcribed by **different speech-recognition engines**, and word-level disagreement between them lands a five-word-window comparison at roughly 0.35–0.60, i.e. just *under* the old cutoff. The threshold is now **0.35**, which takes the archive from 2,846 duplicate clusters over 5,746 videos to **7,434 clusters over 14,997 videos** — so a search result is far more likely to tell you the same recording exists elsewhere, and to offer you the jump. The change was bracketed at four values on the full archive before being made, and every step is a **strict superset**: no video that was previously flagged as a duplicate stopped being one. What it newly admits was inspected rather than counted — 96–97% have byte-identical titles, 99% span two platforms, and the handful of same-channel cases were read individually. Nothing about *how* a duplicate is decided changed: pairs are still confirmed by comparing real transcripts, a pair that merely shares a title and a runtime is still an internal review item that never reaches you, and the "jump to this moment in the other copy" button still only carries your timestamp when the two were *measured* as aligned.
+- **AI chapters: jump straight to the part of a video you want.** Videos that have been through the local-AI digest pass now ship their derived **chapters and topic tags** to the site, and the player gains a third panel beside Transcript and Live chat. Open it and you get a titled list of moments — click one and the player **seeks there**; the chapter you're currently inside stays marked as the video plays. The layer is **sparse on purpose and honest about it**: only a small, growing fraction of the archive has been digested (generation is a multi-week GPU pass), so the control simply **isn't shown** on a video that has no digest, rather than offering a button that opens an empty panel — and if a digest can't be loaded, the panel snaps back to the transcript with a notice instead of stranding you. What ships is the **composed** digest: any human correction is applied and any chapter a human rejected is dropped, so you see what a person approved rather than raw model output, with hand-edited chapters marked **edited** and a provenance line naming the model that wrote the rest. **A digest borrowed from a duplicate upload says so, prominently** — when the same recording exists twice in the archive, one copy's chapters can be shared onto the other, and the panel names the source video and the measured timing offset rather than passing them off as native (plausible chapters describing a *different* upload is the failure that looks like success). Digests live at `/digests/<channel>/` under the same paginated-shard scheme as transcripts and posts, are offline-cached by the service worker, and are described in `corpus.json` — which bumps to **spec 3** with a `digestScheme` and per-channel `digests` manifest pointers, so an AI tool reading the corpus can navigate them and knows that an absent video means "not yet generated" rather than "nothing to say". Deep links carry the panel (`?vm=digest`), and share links reopen on it. Operator telemetry (why the model's proposals were rejected, the regeneration history) is deliberately **not** shipped — that stays in the editor. See `common/lib/digests.ts`, `common/components/{digestCache,digestStore}.ts`, `common/components/{PlayerProvider,TranscriptModal,urlState}.tsx/ts`, and `export/e2e/modal-digest.spec.ts`.
+- **Search results tell you when a video exists elsewhere in the archive, and take you there.** A result that belongs to a duplicate cluster now carries a **Dupe** badge, and a strip under the card header offers one button per other copy — the same recording mirrored to another platform, or re-uploaded on another channel. Clicking one opens that copy in the player. **The jump is honest about what it knows:** matching content does not imply matching timings (a mirror with a longer intro carries the same words at shifted times), so a button only carries your current timestamp when detection *measured* the two as aligned; otherwise it says so and opens the other copy from the start. A copy whose alignment was never measured is treated as not aligned. Only clusters whose transcripts were actually compared reach the site — a pair that merely shares a title and a runtime stays an internal review item and is never asserted to you. In hub mode the badge is limited to same-origin results, since the duplicate index is per-site. Sites with no duplicate report are entirely unaffected.
- **One search now covers video transcripts *and* social posts.** Archived X/Twitter and Bluesky posts ship as a parallel corpus beside transcripts and live chat, and compose into the same boolean query tree — so `(transcripts:"foo" OR posts:"foo")` returns both kinds in one ranked, newest-first result set. `LayerScope` gains `"posts"` (whitelisted in `qt=` deserialization, so a shared link round-trips a posts leaf), the leaf scope selector gains **Posts**, and the filter row gains a **Posts** media kind beside Videos and Livestreams — a third kind, because a post is neither, and folding it into the video toggle would silently drop the whole corpus. Post and video slugs live in disjoint namespaces, partitioned per-leaf by the eval engine so a transcripts leaf never fetches a post and an AND across the two can't collapse to nothing. Post result cards drop what doesn't apply (no seek gutter, no livestream/age badges, no VOD expiry) and lead with the post body; opening one shows a new **PostModal** — a sibling of the transcript reader, not a generalization of it — with the post, its archived thread, its outbound links and its engagement counts. Date filters work unchanged: every post carries a derived `uploadDate`. Posts are cached and served under `/posts/`, offline-cached by the service worker, and CORS-readable so a federating hub merges them across origins.
- **"Ask AI" is grounded in posts as well as transcripts.** Retrieval adds a posts leaf per keyword alongside the transcript and metadata leaves, sharing the same term key so ranking still counts a keyword once rather than three times. Post excerpts render without a `[clock]` line and are cited as a bare `[n]` (never `[n @ mm:ss]` — a post has no timeline), the source list shows a date instead of a meaningless 0:00 seek button, and "load more context" on a post returns its **thread** rather than a time window.
- **MCP: posts are first-class.** `search_transcripts` gains `content_types` (defaulting to **both**, so existing agent flows pick posts up automatically) and renders post hits without moment links; new `get_post` and `get_thread` tools read one post or a whole conversation; `open_link` accepts a `posts` query scope; and the sweep prompt teaches the post citation form. All of it lands on the single `ShardSource` boundary, so local, remote and hub transports gain it at once. `corpus.json` bumps to spec 2 with a `postScheme` describing the new shards.
diff --git a/export/app/duplicates/DuplicatesClient.tsx b/export/app/duplicates/DuplicatesClient.tsx
@@ -28,11 +28,20 @@ const PLATFORM_LABEL: Record<string, string> = {
const MATCH_LABEL: Record<DuplicateCluster["matchKind"], string> = {
"transcript-exact": "exact transcript",
"transcript-near": "near transcript",
+ "title-duration": "title + runtime",
};
+// MATCH_KINDS is not just a label list — it seeds the default filter state, and
+// the visible-cluster filter tests membership in it. A kind missing from here is
+// invisibly dropped from the page with no checkbox to turn it back on, so every
+// variant of DuplicateMatchKind must appear. (`title-duration` clusters are
+// filtered out at compose time unless a human confirmed them, so in practice
+// this checkbox only has anything to show once that happens — which is exactly
+// the case that must not silently vanish.)
const MATCH_KINDS: DuplicateCluster["matchKind"][] = [
"transcript-exact",
"transcript-near",
+ "title-duration",
];
const RELATIONSHIPS = [
diff --git a/export/e2e/fixtures/data.ts b/export/e2e/fixtures/data.ts
@@ -269,6 +269,93 @@ export function transcriptPage() {
];
}
+// ─── AI-digest fixtures ───
+// The derived layer is SPARSE: only VIDEO_TRANSCRIPT_ONLY carries a digest, so
+// the same fixture set covers both "has a digest" and "does not" without a
+// second channel. VIDEO_CHAT_SMALL deliberately has none — that is what proves
+// the Digest control stays hidden rather than becoming a dead end.
+//
+// Chapter starts are chosen to sit ON the transcript cue starts above (5 / 50 /
+// 100), which is what the real parser guarantees by snapping to cue boundaries.
+export function channelDigestsManifest() {
+ return {
+ version: 1,
+ channelSlug: CHANNEL_SLUG,
+ pageCount: 1,
+ maxPageBytes: 8388608,
+ generatedAt: new Date().toISOString(),
+ slugToPage: { [VIDEO_TRANSCRIPT_ONLY]: 0 },
+ pageHashes: ["fixture-hash-0"],
+ };
+}
+
+export function digestPage() {
+ return [
+ {
+ slug: slug(VIDEO_TRANSCRIPT_ONLY),
+ id: VIDEO_TRANSCRIPT_ONLY,
+ generatedAt: "2026-07-20T00:00:00.000Z",
+ chapters: [
+ {
+ id: "c5",
+ start: 5,
+ clock: "00:00:05",
+ title: "Opening remarks",
+ decidedBy: "ai" as const,
+ },
+ {
+ id: "c50",
+ start: 50,
+ clock: "00:00:50",
+ title: "The main argument",
+ decidedBy: "ai" as const,
+ },
+ {
+ id: "c100",
+ start: 100,
+ clock: "00:01:40",
+ // A human-corrected title, so the "edited" marker has something to
+ // render and the ai/human distinction is exercised.
+ title: "Corrected by hand",
+ decidedBy: "human" as const,
+ },
+ ],
+ tags: [
+ { id: "tnews", tag: "news", decidedBy: "ai" as const },
+ { id: "tpolicy", tag: "policy", decidedBy: "ai" as const },
+ ],
+ provenance: {
+ chapters: {
+ appId: "ollama-direct",
+ model: "qwen2.5:7b",
+ lane: "local-gpu" as const,
+ generatedAt: "2026-07-20T00:00:00.000Z",
+ promptVersion: 2,
+ contextHash: "",
+ },
+ },
+ },
+ ];
+}
+
+// The same digest, marked as borrowed from a duplicate cluster's canonical
+// member. Used to prove a shared digest is presented AS shared — the failure
+// mode here looks like success, so it needs its own assertion.
+export function borrowedDigestPage() {
+ const [entry] = digestPage();
+ return [
+ {
+ ...entry,
+ derivedFrom: {
+ slug: "other-channel/vid-canonical",
+ clusterId: "cluster-1",
+ sharedAt: "2026-07-21T00:00:00.000Z",
+ offsetSeconds: 0.4,
+ },
+ },
+ ];
+}
+
export function channelTranscriptsManifest() {
return {
version: 1,
diff --git a/export/e2e/helpers.ts b/export/e2e/helpers.ts
@@ -66,12 +66,34 @@ export async function installRoutes(page: Page) {
await page.route(/\/posts\/[^/]+\/page-\d+\.json$/, async (route) => {
await fulfillJson(route, postsPage());
});
+ // AI digests — absent by default, which is also the real default: only ~0.1%
+ // of the corpus is digested, so most channels ship no digests tree at all and
+ // the manifest 404s. Tests that want the Digest control register their own
+ // route afterwards, which takes precedence.
+ await page.route(/\/digests\/[^/]+\/manifest\.json$/, async (route) => {
+ await route.fulfill({
+ status: 404,
+ contentType: "application/json",
+ body: "{}",
+ });
+ });
// Search-alias dictionary — empty by default; alias-suggestion.spec overrides
// this with a populated list. Kept here so other search tests get a clean
// intercept instead of a real 404.
await page.route("**/search-aliases.json", async (route) => {
await fulfillJson(route, { aliases: [] });
});
+ // Duplicate clusters — absent by default, which is also the real default:
+ // compose-site writes /duplicates.json only when a site has at least one
+ // shippable cluster. Tests that want the search-result dupe badge register
+ // their own route afterwards, which takes precedence.
+ await page.route("**/duplicates.json", async (route) => {
+ await route.fulfill({
+ status: 404,
+ contentType: "application/json",
+ body: "{}",
+ });
+ });
}
// Charts page fetches: stats dataset + baked templates. Also installs the
diff --git a/export/e2e/modal-digest.spec.ts b/export/e2e/modal-digest.spec.ts
@@ -0,0 +1,198 @@
+import { expect, test, type Page } from "@playwright/test";
+import {
+ CHANNEL_SLUG,
+ VIDEO_CHAT_SMALL,
+ VIDEO_TRANSCRIPT_ONLY,
+ borrowedDigestPage,
+ channelDigestsManifest,
+ digestPage,
+} from "./fixtures/data";
+import { expectModalOpen, installRoutes, urlParams } from "./helpers";
+
+test.use({
+ permissions: ["clipboard-read", "clipboard-write"],
+});
+
+// Register the digest tree AFTER installRoutes so these take precedence over
+// its default 404s (the duplicates.json idiom).
+async function installDigestRoutes(
+ page: Page,
+ page0: unknown = digestPage(),
+): Promise<void> {
+ await page.route(/\/digests\/[^/]+\/manifest\.json$/, async (route) => {
+ await route.fulfill({
+ status: 200,
+ contentType: "application/json",
+ body: JSON.stringify(channelDigestsManifest()),
+ });
+ });
+ await page.route(/\/digests\/[^/]+\/page-\d+\.json$/, async (route) => {
+ await route.fulfill({
+ status: 200,
+ contentType: "application/json",
+ body: JSON.stringify(page0),
+ });
+ });
+}
+
+test.describe("video modal — AI digest panel", () => {
+ test.beforeEach(async ({ page }) => {
+ await installRoutes(page);
+ });
+
+ test("control is hidden on a video with no digest", async ({ page }) => {
+ await installDigestRoutes(page);
+ // VIDEO_CHAT_SMALL is absent from the manifest's slugToPage.
+ await page.goto(`/?v=${CHANNEL_SLUG}/${VIDEO_CHAT_SMALL}`);
+ await expectModalOpen(page);
+ await expect(page.getByText(/alpha line/)).toBeVisible();
+ await expect(
+ page.getByRole("button", { name: "Show AI chapters" }),
+ ).toHaveCount(0);
+ });
+
+ test("control is hidden when the channel ships no digests at all", async ({
+ page,
+ }) => {
+ // No installDigestRoutes → the manifest 404s, as for most of the corpus.
+ await page.goto(`/?v=${CHANNEL_SLUG}/${VIDEO_TRANSCRIPT_ONLY}`);
+ await expectModalOpen(page);
+ await expect(
+ page.getByRole("button", { name: "Show AI chapters" }),
+ ).toHaveCount(0);
+ });
+
+ test("toggle sets vm=digest and back, touching nothing else", async ({
+ page,
+ }) => {
+ await installDigestRoutes(page);
+ await page.goto(
+ `/?m=subs&tk=foo&v=${CHANNEL_SLUG}/${VIDEO_TRANSCRIPT_ONLY}`,
+ );
+ await expectModalOpen(page);
+
+ const toggle = page.getByRole("button", { name: "Show AI chapters" });
+ await expect(toggle).toBeVisible();
+ await toggle.click();
+
+ await expect(page.getByText("The main argument")).toBeVisible();
+ let params = await urlParams(page);
+ expect(params.get("vm")).toBe("digest");
+ expect(params.get("m")).toBe("subs");
+ expect(params.get("tk")).toBe("foo");
+
+ // Back to the transcript — vm is DELETED, not set to "transcript".
+ await page.getByRole("button", { name: "Show transcript" }).click();
+ await expect(page.getByText(/alpha line/)).toBeVisible();
+ params = await urlParams(page);
+ expect(params.get("vm")).toBeNull();
+ });
+
+ test("deep link with ?vm=digest opens straight into the panel", async ({
+ page,
+ }) => {
+ await installDigestRoutes(page);
+ await page.goto(
+ `/?v=${CHANNEL_SLUG}/${VIDEO_TRANSCRIPT_ONLY}&vm=digest`,
+ );
+ await expectModalOpen(page);
+ await expect(page.getByText("Opening remarks")).toBeVisible();
+ await expect(page.getByText("The main argument")).toBeVisible();
+ // Topic tags and the provenance line ship alongside the chapters.
+ // Exact, because the provenance sentence also contains the word "topics".
+ await expect(page.getByText("Topics", { exact: true })).toBeVisible();
+ await expect(page.getByText(/qwen2\.5:7b/)).toBeVisible();
+ });
+
+ test("a chapter click seeks the player to that moment", async ({ page }) => {
+ await installDigestRoutes(page);
+ await page.goto(`/?v=${CHANNEL_SLUG}/${VIDEO_TRANSCRIPT_ONLY}&vm=digest`);
+ await expectModalOpen(page);
+
+ // The playhead starts at 0, before the first chapter's 00:05 — so nothing
+ // is active yet. (A chapter list need not start at zero.)
+ await expect(
+ page.getByRole("button", { name: /Opening remarks/ }),
+ ).not.toHaveAttribute("aria-current", "true");
+
+ // Click the 00:50 chapter; the active-chapter marker follows the playhead,
+ // which is what proves the seek took effect.
+ await page.getByRole("button", { name: /The main argument/ }).click();
+ await expect(
+ page.getByRole("button", { name: /The main argument/ }),
+ ).toHaveAttribute("aria-current", "true");
+
+ // And back — the marker tracks the player, it isn't just a click highlight.
+ await page.getByRole("button", { name: /Opening remarks/ }).click();
+ await expect(
+ page.getByRole("button", { name: /Opening remarks/ }),
+ ).toHaveAttribute("aria-current", "true");
+ await expect(
+ page.getByRole("button", { name: /The main argument/ }),
+ ).not.toHaveAttribute("aria-current", "true");
+ });
+
+ test("share link carries the digest panel", async ({ page, context }) => {
+ await context.grantPermissions(["clipboard-read", "clipboard-write"]);
+ await installDigestRoutes(page);
+ await page.goto(
+ `/?v=${CHANNEL_SLUG}/${VIDEO_TRANSCRIPT_ONLY}&vm=digest&t=50`,
+ );
+ await expectModalOpen(page);
+ await expect(page.getByText("The main argument")).toBeVisible();
+
+ await page
+ .getByRole("button", { name: "Copy share link at current time" })
+ .click();
+ const url = new URL(
+ await page.evaluate(() => navigator.clipboard.readText()),
+ );
+ expect(url.searchParams.get("vm")).toBe("digest");
+ expect(url.searchParams.get("t")).toBe("50");
+ });
+
+ test("warnings never reach the viewer", async ({ page }) => {
+ await installDigestRoutes(page);
+ await page.goto(`/?v=${CHANNEL_SLUG}/${VIDEO_TRANSCRIPT_ONLY}&vm=digest`);
+ await expectModalOpen(page);
+ await expect(page.getByText("Opening remarks")).toBeVisible();
+ // Operator telemetry (out-of-range clamps etc.) is for the editor panel and
+ // the review queue, not for readers.
+ await expect(page.getByText(/out-of-range/)).toHaveCount(0);
+ });
+
+ test("a borrowed digest is presented as borrowed", async ({ page }) => {
+ await installDigestRoutes(page, borrowedDigestPage());
+ await page.goto(`/?v=${CHANNEL_SLUG}/${VIDEO_TRANSCRIPT_ONLY}&vm=digest`);
+ await expectModalOpen(page);
+ // The failure mode this guards against looks like success: plausible
+ // chapters that describe a different upload.
+ await expect(page.getByText("Borrowed from a duplicate upload")).toBeVisible();
+ await expect(page.getByText(/other-channel\/vid-canonical/)).toBeVisible();
+ });
+
+ test("snaps back to the transcript when the page is unreachable", async ({
+ page,
+ }) => {
+ // Manifest says the video is digested, but the page fetch fails.
+ await page.route(/\/digests\/[^/]+\/manifest\.json$/, async (route) => {
+ await route.fulfill({
+ status: 200,
+ contentType: "application/json",
+ body: JSON.stringify(channelDigestsManifest()),
+ });
+ });
+ await page.route(/\/digests\/[^/]+\/page-\d+\.json$/, async (route) => {
+ await route.fulfill({ status: 500, body: "boom" });
+ });
+
+ await page.goto(`/?v=${CHANNEL_SLUG}/${VIDEO_TRANSCRIPT_ONLY}&vm=digest`);
+ await expectModalOpen(page);
+ await expect(page.getByRole("status")).toContainText(
+ /Couldn't load the AI digest/,
+ );
+ await expect(page.getByText(/alpha line/)).toBeVisible();
+ const params = await urlParams(page);
+ expect(params.get("vm")).toBeNull();
+ });
+});
diff --git a/export/e2e/search-duplicates.spec.ts b/export/e2e/search-duplicates.spec.ts
@@ -0,0 +1,199 @@
+import { expect, test, type Page } from "@playwright/test";
+import { installRoutes, urlParams } from "./helpers";
+import {
+ CHANNEL,
+ CHANNEL_SLUG,
+ VIDEO_CHAT_SMALL,
+ VIDEO_TRANSCRIPT_ONLY,
+} from "./fixtures/data";
+
+// The search-result duplicate affordance: a badge saying "this exists elsewhere
+// in the archive", plus a strip of jump buttons to the other copies.
+//
+// The load-bearing assertion here is the HONESTY RULE. Matching content does not
+// imply matching timings — a mirror with a longer intro carries the same words at
+// shifted times — so a jump only carries the current timestamp when the detector
+// MEASURED the two as aligned. An unaligned (or unmeasured) sibling opens at 0.
+// Get this wrong and the viewer lands mid-sentence in the wrong place on a jump
+// that looks like it worked, which is the exact failure measureAlignment exists
+// to prevent.
+
+const ALIGNED_SIBLING = "mirror-chan/aligned-copy";
+const UNALIGNED_SIBLING = "mirror-chan/drifted-copy";
+
+// "gamma" matches one transcript cue per fixture video, at start=100s.
+const HIT_SECONDS = 100;
+
+function dupRef(
+ slug: string,
+ over: Record<string, unknown> = {},
+): Record<string, unknown> {
+ const [channelSlug, id] = slug.split("/");
+ return {
+ slug,
+ channelSlug,
+ channel: channelSlug === CHANNEL_SLUG ? CHANNEL : "Mirror Chan",
+ platform: channelSlug === CHANNEL_SLUG ? "youtube" : "rumble",
+ id,
+ title: `Title ${id}`,
+ duration: 300,
+ uploadDate: "20260101",
+ hasTranscript: true,
+ ...over,
+ };
+}
+
+function duplicatesReport() {
+ return {
+ version: 1,
+ generatedAt: "2026-07-01T00:00:00.000Z",
+ runConfig: {
+ thresholdSeconds: null,
+ durationToleranceSeconds: 2,
+ nearThreshold: 0.6,
+ containmentThreshold: 0.8,
+ shingleSize: 5,
+ blocking: "title",
+ },
+ totals: { videosScanned: 4, clusters: 2, videosInClusters: 4 },
+ clusters: [
+ {
+ // Measured as aligned: the jump may carry the moment.
+ clusterId: "cluster-aligned",
+ matchKind: "transcript-exact",
+ score: 1,
+ contained: false,
+ durationBucket: 300,
+ crossPlatform: true,
+ crossChannel: true,
+ canonicalSlug: `${CHANNEL_SLUG}/${VIDEO_TRANSCRIPT_ONLY}`,
+ videoRefs: [
+ dupRef(`${CHANNEL_SLUG}/${VIDEO_TRANSCRIPT_ONLY}`, {
+ aligned: true,
+ offsetSeconds: 0,
+ }),
+ dupRef(ALIGNED_SIBLING, { aligned: true, offsetSeconds: 1.2 }),
+ ],
+ },
+ {
+ // Same content, DIFFERENT timings — a 41s intro difference. The jump
+ // must fall back to the start of the sibling.
+ clusterId: "cluster-drifted",
+ matchKind: "transcript-near",
+ score: 0.91,
+ contained: false,
+ durationBucket: 300,
+ crossPlatform: true,
+ crossChannel: true,
+ canonicalSlug: `${CHANNEL_SLUG}/${VIDEO_CHAT_SMALL}`,
+ videoRefs: [
+ dupRef(`${CHANNEL_SLUG}/${VIDEO_CHAT_SMALL}`, {
+ aligned: true,
+ offsetSeconds: 0,
+ }),
+ dupRef(UNALIGNED_SIBLING, { aligned: false, offsetSeconds: 41.5 }),
+ ],
+ },
+ ],
+ };
+}
+
+async function installDuplicates(page: Page, report: unknown) {
+ await page.route("**/duplicates.json", async (route) => {
+ await route.fulfill({
+ status: 200,
+ contentType: "application/json",
+ body: JSON.stringify(report),
+ });
+ });
+}
+
+function qt(query: string): string {
+ return encodeURIComponent(
+ JSON.stringify({
+ k: "g",
+ o: "AND",
+ c: [{ k: "l", q: query, s: "transcripts" }],
+ }),
+ );
+}
+
+function card(page: Page, slug: string) {
+ return page.locator(`[data-result-slug="${slug}"]`);
+}
+
+test.describe("search result duplicates", () => {
+ test("badges a result that exists elsewhere and jumps to the moment when aligned", async ({
+ page,
+ }) => {
+ await installRoutes(page);
+ await installDuplicates(page, duplicatesReport());
+ await page.goto(`/?qt=${qt("gamma")}`);
+
+ const alignedCard = card(page, `${CHANNEL_SLUG}/${VIDEO_TRANSCRIPT_ONLY}`);
+ await expect(alignedCard).toBeVisible();
+ await expect(alignedCard.getByText("Dupe", { exact: true })).toBeVisible();
+
+ // The jump button states where it will land, and it lands there.
+ const jump = alignedCard.locator(
+ `[data-duplicate-slug="${ALIGNED_SIBLING}"]`,
+ );
+ await expect(jump).toHaveAttribute("data-duplicate-aligned", "true");
+ await expect(jump).toContainText("01:40");
+ await jump.click();
+
+ const params = await urlParams(page);
+ expect(params.get("v")).toBe(ALIGNED_SIBLING);
+ expect(params.get("t")).toBe(String(HIT_SECONDS));
+ });
+
+ test("an unaligned sibling opens at the start, not at the current moment", async ({
+ page,
+ }) => {
+ await installRoutes(page);
+ await installDuplicates(page, duplicatesReport());
+ await page.goto(`/?qt=${qt("gamma")}`);
+
+ const driftedCard = card(page, `${CHANNEL_SLUG}/${VIDEO_CHAT_SMALL}`);
+ await expect(driftedCard).toBeVisible();
+ await expect(driftedCard.getByText("Dupe", { exact: true })).toBeVisible();
+
+ const jump = driftedCard.locator(
+ `[data-duplicate-slug="${UNALIGNED_SIBLING}"]`,
+ );
+ await expect(jump).toHaveAttribute("data-duplicate-aligned", "false");
+ // No timestamp is offered, because none was measured.
+ await expect(jump).not.toContainText("@");
+ await jump.click();
+
+ const params = await urlParams(page);
+ expect(params.get("v")).toBe(UNALIGNED_SIBLING);
+ expect(params.get("t")).toBe("0");
+ });
+
+ test("a result in no cluster carries no badge and no jump strip", async ({
+ page,
+ }) => {
+ await installRoutes(page);
+ await installDuplicates(page, duplicatesReport());
+ await page.goto(`/?qt=${qt("gamma")}`);
+
+ const lone = card(page, `${CHANNEL_SLUG}/vid-chat-large`);
+ await expect(lone).toBeVisible();
+ await expect(lone.getByText("Dupe", { exact: true })).toHaveCount(0);
+ await expect(lone.locator("[data-duplicate-strip]")).toHaveCount(0);
+ });
+
+ test("no duplicates.json means no duplicate affordance at all", async ({
+ page,
+ }) => {
+ // The overwhelmingly common case: the file is absent, and search must be
+ // completely unaffected rather than erroring on the 404.
+ await installRoutes(page);
+ await page.goto(`/?qt=${qt("gamma")}`);
+
+ const anyCard = card(page, `${CHANNEL_SLUG}/${VIDEO_TRANSCRIPT_ONLY}`);
+ await expect(anyCard).toBeVisible();
+ await expect(page.locator("[data-duplicate-strip]")).toHaveCount(0);
+ });
+});
diff --git a/export/service-worker/site-sw.js b/export/service-worker/site-sw.js
@@ -24,7 +24,7 @@ const PAGES = `pages-${VERSION}`;
const META = `meta-${VERSION}`;
// Matches /transcripts/<slug>/... and the parallel /subs, /summaries trees.
-const SHARD_RE = /^\/(transcripts|subs|posts|summaries)\/([^/]+)\/(.+)$/;
+const SHARD_RE = /^\/(transcripts|subs|posts|digests|summaries)\/([^/]+)\/(.+)$/;
self.addEventListener("install", () => {
// Activate immediately — no precache list (corpus is too large to bundle).
diff --git a/plans/FACTS.md b/plans/FACTS.md
@@ -0,0 +1,880 @@
+# Verified codebase facts
+
+Facts established by reading the tree, each with a `file:line` anchor. **Trust these over
+re-deriving them** — that is the entire point of this file. If one turns out to be stale,
+correct it in place and note the correction in `STATE.md`.
+
+Append as you verify new things. Do not add anything here you have not actually checked.
+
+Last full verification pass: **2026-07-26**, against `main` @ `18c5a7a`.
+
+---
+
+## Corrections to the original draft plan
+
+These were confidently stated and **wrong**. They are the most valuable entries here.
+
+### `windowCues()` is not a chunker
+
+`common/lib/transcriptWindow.ts:21` —
+`windowCues(cues, centerSeconds, { before = 45, after = 45, maxCues = 60 })`. It is a
+**center-based** window around a search hit, with `before`/`after` in **seconds**. It
+cannot split a transcript into sequential context-sized chunks. Callers:
+`export/app/ask/useAskChat.ts:1333`, `export/app/lib/searchAgent.ts:509,742`.
+
+→ A sequential `chunkCuesForContext()` must be written. Put it in the same file; it is pure
+and `transcriptWindow.test.ts` already exists.
+
+### No LMDB `maxDbs` bump is needed
+
+`maxDbs` caps the named DBs opened *by one handle*, not the DBs present in the environment.
+
+- `common/controller/buildIndex.ts` — `maxDbs: 17`, now opens **15**: the original 12
+ (`sums`, `cues`, `subs`, `mtimes`, `byChannel`, `pageHashes`, `subPageHashes`,
+ `channelStats`, `posts`, `postPageHashes`, `channelPostsStats`, `meta`) plus the three
+ Phase 2 added (`digests`, `digestPageHashes`, `channelDigestStats`). Still under.
+- `common/lib/channelSignature.ts:48` — `maxDbs: 14`, opens **2** (`mtimes`, `meta`).
+- `common/controller/duplicateShorts.ts:111` — `maxDbs: 12`, opens **3**.
+
+None of these break. The "17 vs 14 inconsistency" is a red herring.
+
+→ What **was** required and easy to miss: the schema-invalidation block enumerates every
+sub-DB by hand with `clearAsync()`. **Done (2026-07-29)** — the three digest sub-DBs are
+listed there and `SCHEMA_VERSION` is **13**. The next person to add a sub-DB still has to
+add its `clearAsync()` line by hand; nothing enforces it.
+
+### `JobProgressMetric` is not a closed union in practice — **RESOLVED, table was stale**
+
+**This entry is kept only to stop the next session re-doing it. Verified 2026-07-29: all of
+it is already done.** The union is extended and every site is correct:
+
+- `common/jobs/registry.ts:17` — `JobProgressMetric` is `"downloads" | "transcripts" |
+ "digests"`, and `:25` `JobTaskKind` is `"download" | "transcribe" | "digest"`.
+- `editor/app/jobs/components/RunningJobsList.tsx:6-9` **imports** `JobProgressMetric` and
+ `JobTaskKind` rather than re-spelling them, so TypeScript now does flag the next member.
+ A comment on the union at `registry.ts:14` records that this is why.
+- `editor/app/widget/components/MonitorWidget.tsx:628` is a `METRIC_PREFIX` lookup table
+ (`Record<JobProgressMetric, string>`), not the two copies of a `metric === "downloads" ? …`
+ binary the old table described — so it is exhaustiveness-checked and there is nothing
+ left to duplicate.
+
+PLAN.md's Phase 2.5 opens by demanding this fix "first". It is done; start 2.5 at its
+second item.
+
+### Queue concurrency is not declared per key
+
+`common/lib/queueKeys.ts` has **no** concurrency map. `common/jobs/registry.ts:158`
+hardcodes `concurrency: 1` for every non-empty `queueKey`. A `queueKey` of `""` means "run
+immediately, unserialized".
+
+`TRANSCRIPTION_QUEUE` is defined at `common/lib/platform.ts:112` and merely re-exported from
+`queueKeys.ts:2-8`. `resolveQueueKey(defaultKey, override)` is at `queueKeys.ts:27-32`:
+`undefined` → default, `""` → immediate, else trimmed.
+
+### `resolveCitationSeconds` snaps to snippets, not cues
+
+`export/app/ask/citations.tsx:76` scans `src.snippets` — only the ~5 matched lines per
+video (`export/app/lib/askRetrieval.ts:132`). It never sees a cue list, so it cannot be
+reused for cue snapping.
+
+→ The right reference for a cue-boundary snap is `findActiveIndex`
+(`common/components/TranscriptModal.tsx:473-491`), a hand-rolled binary search over cue
+starts.
+
+### There is no vitest or jest
+
+Unit tests are `node:test` run via `tsx --test`. Only `mcp/package.json:12` has a `test`
+script (`tsx --test src/*.test.ts`). `common/package.json` has **no `scripts` key at all**;
+neither `editor` nor `export` has a `test` script. ~46 `*.test.ts` files across the repo are
+unrunnable as a suite today.
+
+### `subsCache.ts` has no IndexedDB
+
+- `common/components/subsCache.ts` — memory + react-query only.
+- `common/components/summariesCache.ts` — pure react-query, eager over all pages, no IDB, no
+ `slugToPage`.
+- `common/components/transcriptCache.ts` — the lazy fetch/memo layer (manifest →
+ `slugToPage` → page), and it opportunistically warms every entry in a fetched page.
+- `common/components/transcriptStore.ts` — the IDB layer. `DB_NAME =
+ "yt-dlp-transcript-browser"`, **`DB_VERSION = 4`** (`:15`), `STORE = "transcripts"`. Its
+ `onupgradeneeded` (`:43-47`) **deletes and recreates** the store, so a version bump wipes
+ the cache.
+
+→ Copy the `transcriptCache` + `transcriptStore` **pair**.
+→ `editor/e2e/export-player-platform-cache.spec.ts:115` hardcodes DB version **3** and is
+already stale relative to `DB_VERSION = 4`.
+
+### Social-posts incrementality cannot be copied for a per-video artifact
+
+`buildIndex.ts:233` — `scanSource()` does `if (isSocialChannel(cfg)) continue;`. Social
+channels never enter the mtime scan at all, which is *why* posts use a per-channel
+shard-signature (`buildIndex.ts:1088-1118`) instead of `MtimeRecord`.
+
+→ For a per-video artifact use the mtime mechanism (`digestMs` in `MtimeRecord`). Follow the
+posts precedent (`buildIndex.ts:1065-1240`) only for the **page-tree emission shape**.
+
+### Archive gating already has a seam
+
+`common/controller/archiveTranscripts.ts:177` computes
+`const requiresRewrite = !build.includeMetadata || build.prettyPrint;` and branches at
+`:205-209` to `writeTransformedCues(...)` vs `link(...)` (a hard link straight from the
+source tree).
+
+→ Phase 8 extends that predicate and that transform, rather than building a parallel gate.
+
+---
+
+## Naming hazards
+
+**`summaries` is taken.** `exportSummariesDir` (`common/lib/paths.ts:58,151`) →
+`/summaries/{manifest,page-NNNN}.json` is the cross-channel browse index of video *metadata*
+(`DisplaySummary` records), documented in `common/lib/corpus.ts:35-36` and served by
+`common/components/summariesCache.ts`. Nothing to do with AI summaries. Use **`digest`**.
+
+**Never name a sidecar `transcript.<x>.<y>`.** `common/lib/videoStatus.ts:137` —
+`const SUB_FILE_RE = /^transcript\.([^.]+)\.([^.]+)$/;` treats any such file as a subtitle
+track. The exclusion list in `readSubTracks` (`:238-256`) is hardcoded. Use
+`ai-digest.json`, `diarization.json`, `attribution.json`.
+
+**`report` already means three things**: `reportDebouncePreset` in settings, the
+`refresh-report` job kind, and `/ask`'s `ReportPanel`. Use `feedback` / `review` instead.
+
+---
+
+## Reusable helpers (do not rewrite these)
+
+| What | Where | Notes |
+| --- | --- | --- |
+| Transcript → text for an AI | `common/lib/transcriptToMarkdown.ts:60` | Declared single source of truth. `stampForCue?: (clock, seconds) => string` at `:43` takes precedence over `linkForCue`. |
+| Zero-padded `HH:MM:SS` | `common/lib/aiHandoff.ts:54` — `hms(s)` | So the stamper is `(_clock, s) => hms(s)`. No new formatter needed. |
+| Read a normalized transcript | `common/controller/normalizeTranscript.ts:161` | `readNormalizedTranscript(cuesPath)`. |
+| Freshness check | `normalizeTranscript.ts:193` | `isCuesJsonFresh(videoDir)`. |
+| tmp+rename write idiom | `normalizeTranscript.ts:139-141` | `${path}.tmp-${process.pid}` then `rename`. |
+| Engine registry + client-safe listing | `common/lib/transcriptionApps.ts:62-81, 248-278` | The pattern for `digestApps.ts`. |
+| Inheritance resolve (global → channel) | `common/lib/cookiePolicy.ts:53-66` | Two rules: truthiness-after-trim for free text, type-guard for enums. |
+| Child process into a job log | `common/jobs/runChild.ts:22` | `runChildIntoLog(onLog, signal, opts)`. |
+| Bounded concurrent pool | `common/jobs/concurrentRunner.ts:29` | `runPool()`; honors `signal` + `drainSignal`. |
+| Managed server action | `editor/app/channels/[slug]/whisperActions.ts:45-103` | The `runManagedFunction` template. |
+| Paginated page writer | `common/controller/buildIndex.ts:694` | `createPageWriter()` — size budget + sha1 skip-if-unchanged. |
+| Health probe | `common/controller/remoteTranscribe.ts:51` | `pingRemoteHealth(worker, timeoutMs)`. |
+| ENOENT → friendly message | `common/social/xGalleryDlFetcher.ts:207-213` | `/ENOENT/.test(message) ? "X not found (set X_BIN)" : message`. |
+| ETA estimator | `editor/app/jobs/active/buildActiveJobs.ts:38` | `computeEtaSeconds`. |
+| Binary search over cue starts | `common/components/TranscriptModal.tsx:473-491` | `findActiveIndex`. |
+
+---
+
+## Transcription engines (Phase 0)
+
+All three are **already selectable** — `common/lib/transcriptionApps.ts:248-252`:
+
+| id | Label | Line | Fields | Notes |
+| --- | --- | --- | --- | --- |
+| `whisper-cpp` | whisper.cpp (whisper-cli) | 157 | model, customArgs | `DEFAULT_TRANSCRIPTION_APP_ID` (`:254`) |
+| `chough` | chough | 179 | model, remoteUrl, chunkSize | writes exactly the `-o` path |
+| `parakeet` | parakeet.cpp (overlapping segments) | 216 | model, chunkSize, device | **`supportsPartialStop: true`** (`:222`) — stops after the current window and stitches a partial transcript on SIGTERM, so a long run is interruptible |
+
+Exposed to the settings form via `listTranscriptionApps()` (`:272`), consumed at
+`editor/app/settings/page.tsx:45`. Per-worker `appId` means a mixed fleet already works
+(`common/lib/settings.ts:837`).
+
+**Nothing measures relative speed.** Phase 0 fills that gap — record the numbers below.
+
+> _Benchmark results: not yet measured._
+
+---
+
+## Build pipeline constants
+
+| Constant | Where | Value |
+| --- | --- | --- |
+| `SCHEMA_VERSION` | `buildIndex.ts:116` | 12 |
+| `CUES_FILE_VERSION` | `normalizeTranscript.ts:33` | 2 |
+| `CORPUS_SPEC_VERSION` | `common/lib/corpus.ts:16` | 2 |
+| `MANIFEST_VERSION` (site summaries) | `common/lib/manifest.ts:36` | 3 |
+| `TRANSCRIPTS_MANIFEST_VERSION` | `manifest.ts:43` | 1 |
+| `SUBS_MANIFEST_VERSION` | `manifest.ts:59` | 4 |
+| `SUMMARIES_PAGE_SIZE` | `manifest.ts:37` | 1000 |
+| `DEFAULT_ARCHIVE_MAX_BYTES` | `common/lib/archiveOptions.ts:72` | 25 MB — over it, **all** archives are dropped |
+
+`MtimeRecord` (`buildIndex.ts:146-154`): `metaMs`, `transcriptMs`, `subsMs`,
+`availabilityMs`, `isDeleted`, `isUnlisted`, `indexKey`.
+Changed-detection comparison at `:455-468`; the single `mtimes.put` at `:631-639`.
+
+`channelSignature.ts` uses a **narrower structural subset** (`:23-27`): only `metaMs`,
+`transcriptMs`, `subsMs`, and its hash (`:65-83`) covers only those three. **A new per-video
+artifact that does not move one of those three will not invalidate archive/compose caches.**
+
+`compose-site.ts`: `ComposeCache` at `:486-495`; `reconcileChannelTree()` at `:579`, called
+three times (`:667` transcripts, `:675` subs, `:686` posts); it passes
+`ignoreBasename: "manifest.json"` internally at `:597` because `generatedAt` churns.
+
+Transcript shards ship the **full `cues` array inline** — `buildIndex.ts:830-832`.
+
+---
+
+## Digest bake-off (measured 2026-07-26)
+
+Harness: `common/bin/digest-bakeoff.ts`. Fixed stratified sample in
+`plans/bakeoff/sample.json` (8 videos, 7 channels, 13 min - 8 h), never written
+to the corpus. Full tables in `plans/bakeoff/round{1,2}.{json,md}`.
+
+Corpus totals captured by the same scan that picked the sample:
+**76,354 videos, 73,367 with transcripts, 77,298 audio-hours**; 6,038 videos over
+4 h holding 40,153 h (8.2% of videos, 52% of audio). Sweep days below = measured
+seconds-per-audio-hour x 77,298 h, one lane, no parallelism.
+
+### Round 1 — screening (short+medium, absolute)
+
+| Candidate | Zero-yield | Chapters/h | Rejection | Max gap | Generic | Sweep days |
+| --- | --- | --- | --- | --- | --- | --- |
+| `gemma2:9b@8192` | 0/10 | 15.78 | 0% | 14:59 | 2.3% | 288.1 |
+| `qwen3:8b@16384` | 1/7 | 15.42 | 23.6% | 39:57 | 11.9% | 204.9 |
+| `mistral-nemo:12b@16384` | 1/7 | 10.28 | 28.2% | 31:20 | 7.1% | 447.3 |
+| `qwen3:14b@8192` | 2/10 | 16.52 | 4.3% | 47:51 | 6.7% | 541.3 |
+| `qwen2.5:7b@16384` (incumbent) | 2/7 | 6.61 | 55% | 41:37 | 11.1% | 59.0 |
+
+qwen3 candidates were run with `think: false`; with thinking on they spend most
+of their output budget on a reasoning block the pinned schema then discards.
+Two `qwen3:14b` chunks failed with `fetch failed` (ollama dropped the connection
+under memory pressure — a 9.3 GB model on an 8 GB card) and are counted as
+zero-yield, which inflates that row. It does not change its exclusion at 541 days.
+
+**The incumbent is the fastest by 3.5x and the worst on quality**: it discards 55%
+of what it generates and lands at 6.6 chapters/h against a ~15 target.
+
+### Round 2 — the long tail (2 x ~3.5 h videos, both timestamp modes)
+
+| Candidate | Zero-yield | Chapters/h | Rejection | Max gap | Generic | Sweep days |
+| --- | --- | --- | --- | --- | --- | --- |
+| `gemma2:9b@8192` **chunk-local** | 0/17 (0%) | 12.93 | 14.3% | 24:04 | 10% | 132.4 |
+| `qwen2.5:7b@8192` **chunk-local** | 2/17 (11.8%) | 13.37 | 19.1% | 24:13 | 31.2% | **24.2** |
+| `qwen2.5:7b@16384` chunk-local | 1/9 (11.1%) | 9.34 | 41.4% | 1:27:48 | 20% | 25.1 |
+| `gemma2:9b@8192` absolute | 3/17 (17.7%) | 11.07 | 28% | 1:04:07 | 5.2% | 127.9 |
+| `qwen2.5:7b@8192` absolute | 5/17 (29.4%) | 10.35 | 47.5% | 1:05:16 | 18.1% | 27.7 |
+| `qwen2.5:7b@16384` absolute (today's default) | 3/9 (33.3%) | 8.19 | 50.4% | 1:27:48 | 17.5% | 26.8 |
+
+**Three findings, all measured, none assumed:**
+
+1. **chunk-local beats absolute for every model tested**, on the metric this stage
+ exists to fix. Zero-yield chunks: 33.3% -> 11.1% (qwen2.5@16k), 29.4% -> 11.8%
+ (qwen2.5@8k), 17.7% -> 0% (gemma2). Rejection rate falls with it in every case.
+ The hypothesis — that the model cannot hold a large absolute offset across a
+ long chunk and reverts to counting from zero — is confirmed.
+
+2. **Smaller chunks are a second, independent win, and they are nearly free.**
+ Halving the context (16k -> 8k, 1200 -> 600 cues) took qwen2.5 from 9.34 to
+ 13.37 chapters/h and its max coverage gap from **1:27:48 to 24:13**. The
+ expected cost did not materialise: 24.2 vs 25.1 projected days. Twice the calls
+ at half the prompt each is the same seconds-per-audio-hour. The plan's working
+ assumption that a smaller context roughly doubles sweep cost is **wrong** —
+ it is flat.
+
+3. **Sweep days are far lower than the short-video estimate suggested.** Round 1
+ measured 59 days on short+medium; Round 2 measures ~25 on the long tail, because
+ long videos amortise the fixed per-call overhead. The 77,298-hour corpus is
+ dominated by long videos, so ~25 days is the number to plan against.
+
+**Recommendation: `qwen2.5:7b` at `num_ctx` 8192 with `timestampMode:
+"chunk-local"`.** It is a strict improvement over today's default on every axis
+measured, including throughput (24.2 vs 26.8 days). Its one weakness is title
+quality — 31.2% generic, the worst in the table; by hand its titles read
+"Introduction and Context" / "Conclusion and Final Thoughts" where gemma2 writes
+"Andrew Wilson and His Comparison to a Serial Killer".
+
+**`gemma2:9b@8192` + chunk-local is the quality option**: 0% zero-yield, 10%
+generic, comparable density — at 5.5x the wall-clock (132 days). It is the right
+engine for a targeted re-run of high-value channels, not for the first full sweep.
+Note gemma2 is hard-capped at an 8192 context by the model itself, so it has no
+16k row to compare.
+
+Not yet measured: the >4 h `verylong` stratum (Round 2 used the `long` bucket),
+and the ~100-video validation run.
+
+---
+
+## Pilot: community-notes (measured 2026-07-26)
+
+39 videos, local lane, `qwen2.5:7b` / `num_ctx` 8192 / `chunk-local`, launched from the
+channel page's Digest stage through the real managed-job path.
+
+| | |
+| --- | --- |
+| Result | **39 generated, 0 failed**, 90 model calls, 140 warnings |
+| Chunks | **90 / 90 usable (100%)** |
+| Chapters | 435 across 39 videos |
+| Warning codes | `out-of-range` 130, `non-monotonic` 9, `seam-duplicate` 1 |
+| Generic titles | 103 / 435 (23.7%) |
+| Re-run | `countMissingDigests` = 0; batch reports **0 generated, 39 fresh, 0 engine calls, 0.1 s** |
+
+**The defect video is fixed.** `community-notes/v2chrch` (3505 cues) is the 2.3 h video whose
+third chunk previously produced NOTHING — nine of eleven starts rejected as out-of-range
+because the model had reverted to counting from zero. At 8 k / chunk-local it is 7 chunks,
+**7/7 usable, 42 chapters, one out-of-range warning**, covering 00:16:59 → 02:12:55 with no
+gap larger than the leading one.
+
+The other two 6-chunk videos behave the same: `v2on7b3` 36 chapters 6/6, `v2apmfn` 20
+chapters 6/6.
+
+**Two honest caveats:**
+
+- 130 `out-of-range` rejections remain across the channel. The guard is catching them and
+ the chunks still yield, so this is degraded output rather than lost output — but it is not
+ zero, and it is the metric to watch in the validation run.
+- **Coverage of the opening minutes is weak on long videos.** The first chapter lands at
+ 00:16:59 (`v2chrch`), 00:15:30 (`v2on7b3`) and **00:41:47** (`v2apmfn`) — the early chunks
+ yielded little or nothing, which is the mirror image of the original defect. Not
+ investigated in this stage. Worth checking whether these channels open with long
+ pre-shows, or whether the first chunk is systematically weaker.
+
+Generic-title rate on real data (23.7%) sits between the bake-off's long-tail sample figure
+for this config (31.2%) and gemma2's (10%), which is consistent rather than surprising.
+
+---
+
+## Duplicate detection at corpus scale (measured 2026-07-26)
+
+### Why the first attempt failed (the counts here are correct and are the reason blocking exists)
+
+`duplicate-shorts.ts --all-durations` originally died two different ways:
+
+| Heap | Outcome | Elapsed |
+| --- | --- | --- |
+| default (~4 GB) | `FATAL ERROR: Ineffective mark-compacts near heap limit` | 87 s |
+| `--max-old-space-size=24576` | `RangeError: Invalid array length` at the containment `candidates.push` | 34 s |
+
+Counting the pairs the old generator would build, without building them:
+
+| | Count |
+| --- | --- |
+| Videos scanned | 76,354 |
+| Participants (`duration > 0`) | 76,318 |
+| Shorts (≤ 180 s) | 6,970 |
+| Duration buckets (2 s) | 11,008 |
+| Duration-bucket pairs (same + adjacent) | 9,131,916 |
+| **Containment pairs** (short × every video ≥ 1.5× its length) | **497,860,972** |
+| Total candidate pairs | 506,992,888 |
+
+Two independent defects, not one:
+
+1. **The containment pass was an unblocked cartesian product** — 98.2% of all
+ candidates. V8 throws `RangeError` because a fast-mode object array caps well
+ below 500 M elements.
+2. **Fingerprints were held for the whole corpus at once.** `fingerprintFrom`
+ builds a `Set` of 5-word shingles over the entire transcript (~40 k strings
+ for a 3.5 h video), and the duration pass alone made `needed` ≈ the corpus.
+ That is what killed the 4 GB run before the array was even reached.
+
+### The conclusion that was wrong
+
+The earlier note here said corpus-wide detection "does not complete on this
+corpus" and recorded that "as a cost, not worked around". **That was wrong.**
+Neither failure is inherent to corpus-wide detection: blocking fixes (1) and
+block-at-a-time streaming fixes (2). Both are now implemented, and corpus-wide
+detection completes comfortably in minutes on a default heap.
+
+### Corpus-wide runs, measured (same corpus, 76,354 videos)
+
+`detectDuplicateShorts` now iterates blocks — fingerprint a block, evaluate its
+pairs, keep the confirmed matches, release — so peak memory is O(largest block),
+and no multi-million-element pair array is ever materialised. `--blocking`
+selects the nomination strategy.
+
+| | `--blocking title` | `--blocking duration` |
+| --- | --- | --- |
+| Wall clock | **206 s** | **2,889 s** (48 min) |
+| Peak RSS | **2,001 MB** | **2,450 MB** |
+| Blocks | 7,551 title groups (of 67,064 distinct titles) | 11,008 buckets of 2 s |
+| Oversized blocks skipped | **0** (cap 40) | **0** (cap 2,000) |
+| Nominated pairs | 8,352 | 9,131,476 |
+| Confirmed by content | 2,717 | **5,708** |
+| Rejected by content | 5,250 | 9,125,768 |
+| Untestable suspects | 385 | **0** (by design — see below) |
+| Videos fingerprinted | 15,479 | 74,943 |
+| Clusters written | 2,994 / 6,044 videos (335 `needsReview`) | 2,917 / 5,955 videos (0 `needsReview`) |
+| Cluster kinds | 79 exact, 2,580 near, 335 suspect | 108 exact, 2,809 near |
+| Alignment | 1,599 of 2,689 within 5 s | 1,818 of 3,028 within 5 s |
+
+The duration run's 9,131,476 nominated pairs land within 0.005% of the 9,131,916
+predicted above — the old counting was right, only the conclusion drawn from it
+was wrong.
+
+**Duration blocking yields no suspects, and that is deliberate.** A shared runtime
+alone was never evidence (it produced enormous false clusters of unrelated
+same-length videos), so when a duration-nominated pair cannot be content-tested it
+is dropped, not queued for review. Only title blocking produces suspects, because
+only *title + near-identical runtime* is a claim worth a human's attention.
+
+Title-run cluster sizes are overwhelmingly pairs: 2,950×2, 39×3, 3×4, 1×6, 1×9.
+2,728 cross-platform, 2,891 cross-channel.
+
+### Measured recall: what title blocking actually misses
+
+Comparing the two runs' clusters directly (title's suspects excluded, since they
+are not confirmed):
+
+| | Count |
+| --- | --- |
+| Clusters with identical membership in both | 2,613 |
+| Title-only clusters | 46 |
+| Duration-only clusters | 304 |
+| Videos in title clusters / in duration clusters | 5,351 / 5,955 |
+| **Videos found only by duration** (re-titled mirrors) | **672** |
+| Videos found only by title (runtime drift past the bucket) | 68 |
+
+So **title blocking has ~89% of duration blocking's video-level recall at 1/14th
+the wall clock**, and the 672 it misses are exactly the re-titled mirrors it
+structurally cannot see. `--blocking both` exists for when that 11% matters; the
+title default is the right routine choice.
+
+The containment sweep is now budget-capped (`MAX_CONTAINMENT_PAIRS`, 5 M) and
+reports itself skipped corpus-wide: `6,970 shorts × 76,318 videos ≈ 531,936,460
+pairs`. It still runs, unchanged, at shorts scale. Making it scale is the one
+piece genuinely deferred — see STATE.md.
+
+### 63% of title matches were rejected at 0.6 — and most of those rejections were wrong
+
+Of 8,352 pairs nominated by an exact normalized title *and* a compatible runtime,
+**5,250 were rejected once their transcripts were compared** at the default 0.6
+5-gram Jaccard. Only 2,717 held up.
+
+That was read as "the title-level count was right; the inference that they are
+duplicates was not." **Measurement has since shown the opposite: the threshold was
+wrong, not the inference.** See the next section.
+
+### MEASURED: the 0.6 near threshold was rejecting real cross-platform mirrors
+
+Re-run corpus-wide, title blocking, everything else identical, only `--near`
+changed (2026-07-27):
+
+| | `--near 0.6` (default) | `--near 0.35` |
+| --- | --- | --- |
+| Nominated pairs | 8,352 | 8,352 |
+| **Confirmed by content** | 2,717 | **7,329** |
+| Rejected by content | 5,250 | **638** |
+| Untestable suspects | 385 | 385 |
+| Clusters written | 2,994 / 6,044 videos | **7,450 / 15,030 videos** |
+| `needsReview` | 335 | 334 |
+| **Timing-aligned mirrors** | 1,599 of 2,689 | **3,952 of 7,221** |
+| Wall clock | 206 s | 249 s |
+
+The confirmed set grows **+170%** and is a strict superset: **0 videos** present in
+the 0.6 clusters are absent from the 0.35 clusters.
+
+**The 4,456 entirely-new clusters were inspected, not just counted**, and they are
+real mirrors. Evidence, in descending order of strength:
+
+- **4,423 of 4,456 (99.3%) are cross-platform AND cross-channel.** 3,954 pair a
+ channel with its own platform-suffixed sibling (`the-quartering` ↔
+ `the-quartering-rumble`); of the 502 that don't, the sampled ones pair a
+ creator's alternate channels (`jeremy-hambly` / `quartering-live` /
+ `the-quartering`, `HasanAbiVODs` / `HasanAbiVODsbackup`).
+- **4,290 of 4,456 have a byte-identical title** across every copy; 4,023 agree on
+ runtime within 2 s; 3,572 share an upload date.
+- The score histogram is a tight band at **0.45–0.60** (1,995 at 0.55–0.60, 1,481
+ at 0.50–0.55, 665 at 0.45–0.50) — i.e. they pile up *just under* the old cutoff,
+ which is the signature of a systematic offset, not of noise.
+- The **138 riskiest** (score < 0.42) and the **502 non-sibling** clusters were
+ sampled across the whole band. Every one examined is the same recording: same
+ title, same runtime to the second, same upload date, mirrored platform.
+
+The mechanism is the one the caveat predicted: the two sides of a cross-platform
+mirror are transcribed by **different ASR engines**, and a 5-gram Jaccard is
+unforgiving of word-level disagreement. Two transcripts of the same audio from
+different engines land at **~0.35–0.60**, not ≥ 0.6. The old default was set for
+same-engine text and silently failed the cross-platform case, which is precisely
+the case duplicate detection exists to catch.
+
+**Consequence for the digest backfill.** The ceiling on shared-digest savings is
+not 2,717 pairs. At 0.35 it is **7,329 confirmed pairs, of which 3,952 are
+timing-aligned** enough for a shared digest to be placed correctly — **2.5× the
+1,599 previously banked**. Against a corpus of 76,318 videos that is ~5.2% of the
+sweep avoidable by sharing rather than ~2.1%.
+
+**CHANGED 2026-07-29: `DEFAULT_NEAR_THRESHOLD` now ships at 0.35**, in its own
+commit, after bracketing it at four values — see the next section. Note that the
+"consequence for the digest backfill" above was **overstated**: the saving is
+real in pair count but nearly absent in audio-hours, which is the unit the sweep
+is actually priced in. ~5.2% of pairs is ~1.7 sweep days of 80.
+
+Reproduce with:
+`pnpm exec tsx bin/duplicate-shorts.ts --all-durations --blocking title --near 0.35`
+(back up `transcripts/duplicates.json` first — the run overwrites it).
+
+### Bracketing the near threshold — 4 runs, identical inputs (measured 2026-07-29)
+
+Corpus-wide, `--all-durations --blocking title`, only `--near` varied. 76,354 videos
+scanned, 8,352 pairs nominated in every run.
+
+| `--near` | confirmed | rejected | clusters | videos | aligned | needsReview | wall |
+| --- | --- | --- | --- | --- | --- | --- | --- |
+| 0.6 | 2,736 | 5,404 | 2,846 | 5,746 | 4,273 | 168 | 124 s |
+| 0.45 | 7,147 | 993 | 7,110 | 14,344 | 10,747 | 166 | 199 s |
+| **0.35** | **7,483** | **657** | **7,434** | **14,997** | **11,217** | 166 | 225 s |
+| 0.25 | 7,542 | 598 | 7,493 | 15,115 | 11,297 | 166 | ~230 s |
+
+Every step down is a **strict superset** — 0 videos lost at any step. Returns collapse:
+0.6→0.45 admits **4,264** entirely-new clusters, 0.45→0.35 admits **324**, 0.35→0.25 admits
+**59**.
+
+The marginal bands were read, not counted:
+
+| band | new clusters | byte-identical title | runtime ±2 s | cross-platform | SAME-channel |
+| --- | --- | --- | --- | --- | --- |
+| 0.6 → 0.45 | 4,264 | 96.2% | 91.1% | 99.5% | **3** |
+| 0.45 → 0.35 | 324 | 97.2% | 83.3% | 96.0% | **2** |
+| 0.35 → 0.25 | 59 | 94.9% | 86.4% | 93.2% | **0** |
+
+The riskiest cases in each band were printed individually. In the 0.6→0.45 band they are all
+YouTube↔Rumble pairs agreeing on title (modulo whitespace), runtime to the second, and upload
+date. The 0.45→0.35 band contains the only two arguable cases in the whole sweep: a
+`HasanAbiVODs` pair with the same title but runtimes 9 minutes apart, and an `omnibased` pair
+with identical title/duration/date. Neither is a clear false positive; the first would be
+refused by the alignment gate regardless.
+
+**Adopted: 0.35.** The recall knee is at **0.45** — the value to take if a more conservative
+assertion is ever wanted.
+
+### The digest sweep, priced in AUDIO-HOURS (measured 2026-07-29)
+
+`common/bin/digest-plan.ts --no-freshness`, at the measured 90 s/audio-hour. Total digestable
+corpus: **77,298 audio-hours** over 73,367 videos (2,987 have no transcript), which reproduces
+the roadmap's ~81-day headline independently.
+
+| `--near` | canonical | mirror-aligned (free) | mirror-unaligned | unclustered | **to generate** | **days** |
+| --- | --- | --- | --- | --- | --- | --- |
+| 0.6 | 853 h | 494 h | 344 h | 75,607 h | 76,804 h | 80.0 |
+| 0.45 | 3,472 h | 1,448 h | 1,866 h | 70,513 h | 75,851 h | 79.0 |
+| 0.35 | 4,070 h | 1,661 h | 2,206 h | 69,361 h | 75,638 h | 78.8 |
+| 0.25 | 4,137 h | 1,683 h | 2,242 h | 69,236 h | 75,616 h | 78.8 |
+
+**Duplicate sharing is not a meaningful cost lever, and `digestSharing.ts`'s "~11% of the
+sweep" header is wrong.** Cluster members are 19% of the corpus by video count but **4.3% of
+its audio-hours**: mirrors skew SHORT while the sweep is dominated by unclustered long-form
+VODs. Loosening the threshold as far as it goes moves the sweep from 80.0 to 78.8 days.
+
+Note the `mirror-unaligned` column: only ~55% of mirrors pass the alignment gate, so counting
+all mirrors as free would overstate the saving by nearly half. The plan splits them.
+
+Heaviest channels (the sweep's work-queue order), audio-hours to generate at 0.35:
+
+| channel | h | days |
+| --- | --- | --- |
+| HasanAbiVODs3 | 8,331 | 8.7 |
+| omnibased | 7,783 | 8.1 |
+| HasanAbiVODs | 5,380 | 5.6 |
+| rekietalaw | 4,993 | 5.2 |
+| destiny | 4,020 | 4.2 |
+| shondo-vods | 3,332 | 3.5 |
+
+`HasanAbiVODs3` alone outweighs every duplicate mirror in the corpus combined.
+
+### The id-vs-directory mismatch (measured 2026-07-29)
+
+A slug is `${channelSlug}/${id}` and an id is **not** the on-disk directory name for
+**11,175 of 77,106 indexed videos (14.5%)**, concentrated almost entirely in Rumble
+re-uploads:
+
+| channel | dir !== id |
+| --- | --- |
+| the-quartering-rumble | 7,870 (of 7,870 — every video) |
+| leaflit-rumble | 838 |
+| midwestly | 758 |
+| rekietalaw-rumble | 641 |
+| omnimirror | 522 |
+| cornbreadman | 337 |
+
+Verified directly: for all 7,870 `the-quartering-rumble` entries, `stat.slug ===
+channel/id` and `stat.slug !== channel/dir`. `digestBatch` looked the cluster plan up by
+DIRECTORY, so every one of those lookups missed silently and the mirror regenerated. See the
+`videoDir` field on `DuplicateVideoRef`.
+
+### What this unblocks
+
+Corpus-wide detection is no longer off. `buildDigestClusterPlan` reads
+`duplicates.json` and returns an empty plan when it is absent, so the digest batch
+stays correct either way — it just re-generates mirrors instead of sharing them.
+
+---
+
+## Digest validation run — 102 videos, real corpus (measured 2026-07-27)
+
+The first digest run at corpus scale rather than on a hand-picked bake-off
+sample. Shipped defaults (`qwen2.5:7b` @ 8192, chunk-local, 600 cues,
+`PROMPT_VERSION` 2), production build, driven through the channel page's
+**Digest channel** button. `community-notes` (39 videos, 33.5 audio-h — the
+pilot channel, every digest invalidated by the version bump) plus `friendofrc`
+(63 videos, 8.7 audio-h, never digested). Score with:
+
+```
+cd common && pnpm exec tsx bin/digest-validate.ts community-notes friendofrc
+```
+
+**102 digests, 42.2 audio-hours, 153 chunks, 622 chapters, 0 failures.**
+
+### Quality: better than projected on everything except coverage gap
+
+| Metric | round 2 projection | **validation (102 videos)** | |
+| --- | --- | --- | --- |
+| Zero-yield chunks | 11.8% (2 of 17) | **0.0% (0 of 153)** | far better |
+| Chapters/hour | 13.37 | **14.74** | better |
+| Rejection rate | 19.1% | **18.4%** | matches |
+| Generic titles | 31.2% | **27.0%** | better |
+| Duplicate titles | 3.2% | **0.3%** | far better |
+| Worst coverage gap | 00:24:13 | **01:09:48** | **2.9× worse** |
+
+**Zero-yield went to actually zero.** Not one of 153 chunks came back empty.
+Round 2's 11.8% came from 2 chunks in a 17-chunk sample; at 9× the chunks the
+rate is 0. The chunk-local + 8192 pairing does not produce dead chunks.
+
+**Rejection rate landed within 0.7 points of a 2-video projection**, which is
+the strongest evidence available that the bake-off sample was representative on
+quality.
+
+**Rejections are now almost entirely one guard:**
+
+| Guard | Count | round 2 |
+| --- | --- | --- |
+| `out-of-range` | **130** (93%) | 13 |
+| `non-monotonic` | 9 | 2 |
+| `seam-duplicate` | 1 | 0 |
+| `language-drift` | **0** | 7 |
+
+Language drift disappeared across 42 audio-hours (round 2 saw 7 in 7). The range
+clamp is doing effectively all of the work, which means it is the guard whose
+removal would silently corrupt the corpus.
+
+**The worst coverage gap is the one metric that got worse at scale, and the
+per-video panel explains it.** `community-notes/v2fkbw7` (1:58:42, 20 chapters,
+gap 1:09:48) recorded **13 `out-of-range` rejections** — the model proposed
+chapters for that hour and every one was thrown away by the clamp. So the gap is
+not "the model ignored an hour of video", it is "the model's output for that hour
+was unusable". Those are different problems with different fixes, and the
+distinction is only visible because the artifact records rejections instead of
+dropping them. **The median gap is 00:05:19**, so this is a tail, not the norm —
+report the median alongside the max.
+
+### Throughput: 3.3× worse than projected, and video length is why
+
+| | audio-h | wall | s per audio-h |
+| --- | --- | --- | --- |
+| round 2 (long videos, idle box) | 6.96 | — | **27** |
+| `community-notes` (avg 51 min/video) | 33.5 | 42.1 min | **75** |
+| `friendofrc` (avg 8 min/video) | 8.7 | 20.8 min | **143** |
+| **combined** | **42.2** | **63.4 min** | **90** |
+
+Projected one-lane sweep over the corpus's 77,298 audio-hours: **~81 days, not
+the 24.2 round 2 projected.**
+
+Two compounding causes, and both matter for planning:
+
+1. **Short videos cost ~2× per audio-hour** (143 vs 75 s). A short video is one
+ chunk, so per-call overhead is amortized over minutes instead of hours. Round 2
+ measured only the `long` bucket, so its seconds-per-audio-hour is the corpus's
+ *best* case, applied to the whole corpus. The head of the corpus by count is
+ short videos.
+2. **The box was not idle.** `auto-download` and `auto-transcribe` ran throughout
+ (load average 10–14, whisper contending for the same GPU). This is
+ representative of production but not comparable to the bake-off's conditions.
+
+Even the long-video channel measured 75 s/audio-h against a 27 s projection, so
+contention alone accounts for roughly a 2.8× factor and video-length mix for the
+rest. **A sweep plan should assume ~80–90 days on a shared box, and should
+schedule against the transcription lanes rather than beside them.**
+
+### Two things the run confirmed about the plumbing
+
+- **The counters agree now.** Before the sweep, `noDigest` reported **39** and
+ `countMissingDigests` independently reported **39** for `community-notes` — the
+ version bump correctly invalidating every pilot digest. Under the old bucket the
+ Digest stage read "All digested" on the same corpus. Snapshot regeneration with
+ the added freshness target takes **0.1 s** for a 39-video channel, so the hot
+ path did not get meaningfully slower.
+- **Progress bars cannot move during a REGENERATION.** `setProgress` uses
+ `initial: stat.digestCount` and `target: initial + missing`, but `digestCount`
+ counts digests that *exist* — and a regeneration rewrites in place, so the count
+ never grows and `pct` sits at 0 for the whole job (observed:
+ `{initial:39,current:39,target:72,pct:0}`). Cosmetic, but it makes a long
+ re-sweep look wedged. Fixing it means counting digests *at the current identity*
+ rather than digests-on-disk.
+
+## Editor surfaces
+
+| Surface | File | Note |
+| --- | --- | --- |
+| System paths table | `editor/app/settings/page.tsx:17-28` | Hand-curated tuple list, **not** `Object.entries(paths)`. The blurb at `:57-64` is already stale (missing parakeet/gallery-dl). |
+| Job kinds | `common/jobs/jobKinds.ts:21-40, 44-255` | ~30 entries. **No entry sets `defaultTier`** — declared but unused so far. |
+| Replay handlers | `editor/app/jobs/jobReplayRegistry.ts:75-236` | Bucket kinds re-derive ids from the live snapshot (`:50-57`), never a frozen list. |
+| Snapshot buckets | `common/controller/channelSnapshot.ts:57-143` | 21 buckets, all `string[]` of video ids. Readers default `?.length ?? 0`. `undownloadedIds` is deliberately **outside** `buckets` (`:144`). |
+| Actionable sections | `editor/app/actionable/page.tsx:30-46` | `SectionConfig` is module-private. Counters live in `actionable/lib/loadActionable.ts`. |
+| Widget sync payload | `editor/app/api/widget/sync/route.ts:15-26` | Comment at `:10-14` states it deliberately stays "a handful of scalars". `buildWidgetSyncPayload()` (`:30`) is exported for SSR seeding. |
+| Pipeline band | `editor/app/components/dashboard/PipelineBand.tsx:63-114` | Pure props, no fetching. `<Instrument dotClass=…>` encodes state color. |
+
+---
+
+## Viewer surfaces
+
+`ModalMode` — `common/components/urlState.ts:13`, currently
+`"transcript" | "chat" | "post"`. Parsed at `:63-65` (unrecognized → `"transcript"`),
+written at `:126-130`. Referenced in **only three files**: `urlState.ts`,
+`PlayerProvider.tsx`, `TranscriptModal.tsx`.
+
+Lazy chat fetch effect — `PlayerProvider.tsx:590-640`, deps `[urlSlug, modalMode]`, with a
+ref-based in-flight guard and a comment at `:585-589` explaining why reducer state in the
+dep array drops results. It snaps back via `writeUrlParams({ vm: "transcript" })` on missing.
+`seekTo` at `:522-530`.
+
+**Styling:** `TranscriptModal.tsx` and `PlayerProvider.tsx` are **not** on semantic theme
+tokens. Verified hard-coded: `bg-black/70` (`:199`), `bg-zinc-900/80 ring-white/10` (`:210`),
+`text-white` (`:331`), `text-zinc-300` (`:336`), `border-zinc-700 bg-zinc-900` (`:357`),
+`bg-blue-950/50` (`:419`). By contrast `export/app/ask/citations.tsx` uses semantic tokens
+throughout — the two surfaces are on different systems. Do not mix them.
+
+`export/app/lib/askProvider.ts`: `Provider` union at `:19`, `PROVIDERS` at `:34-59`,
+`askStream` switch at `:128-137`, `askOpenAI` at `:336-397`. Provider settings live at
+`export/app/ask/ProviderSettings.tsx` (**not** `app/components/`); its `Props` is a 21-field
+flat prop bag (`:13-36`).
+
+Two independent search index paths: `common/components/searchIndex.worker.ts` (FlexSearch,
+client) and `mcp/src/search.ts` (server). MCP has 11 tools (`mcp/src/server.ts:132`) and a
+14-member `ShardSource` interface (`mcp/src/source.ts:119`) implemented three times —
+`LocalSource` (`:185`), `RemoteSource` (`:353`), `HubSource` (`:477`).
+
+---
+
+## E2E environment
+
+- `editor/playwright.config.ts` — editor on `PORT ?? 3011`, export on `EXPORT_PORT ?? 3010`.
+ `E2E_MODE=start` → `pnpm start:test`, otherwise `pnpm dev:test` (`:10-11`).
+ `fullyParallel: false`, `workers: 1`, `timeout: 30_000`.
+- `editor/e2e/helpers.ts` — `resetData` (`:32`), `writeSettings` (`:49`), `readJson` (`:90`).
+- Fake binaries live in `editor/e2e/fixtures/bin/`. The **`SLOWOP` sentinel** (checked
+ case-insensitively against cwd or video id) makes an instant fake emit paced progress —
+ three independent copies at `fake-whisper.mjs:29-40`, `fake-chough.mjs:25`,
+ `fake-ytdlp.mjs:173-183`.
+- Binary env vars are wired in `editor/package.json`'s `dev:test` **and** `start:test`
+ scripts — a single line duplicated verbatim, differing only in `next dev` vs `next start`.
+ **Adding one env var means editing both.**
+- Route-stubbing scaffold: `editor/e2e/export-player-platform-cache.spec.ts:39-111`.
+
+### Known-failing on base — not regressions
+
+**Re-verified 2026-07-26** by stashing all working changes and re-running the suspect specs
+against `8c041fd`. Full suite in default dev mode: **21 failed / 363 passed of 384**, and
+every one of the 21 is accounted for below. Do not chase these.
+
+| Spec | Count | Evidence |
+| --- | --- | --- |
+| `actionable.spec` — 2 "queues a job" + the "Needs attention" dashboard card | 3 | fails on base |
+| `bulk-actions.spec` raw-job-kind tests | 3 | fails on base (known) |
+| `channels-actions.spec:42` | 1 | fails on base (known) |
+| `cleanup-actionable.spec` inline "Clean audio" | 1 | fails on base |
+| `deploy-page.spec` (all three) | 3 | fails on base — `getByRole('button', {name:'Build & deploy'})` also matches "Build & deploy all sites"; `getByLabel('Deploy after build')` matches 2 checkboxes |
+| `new-channel-onboarding.spec` (all five) | 5 | fails on base — `getByLabel('URL')` resolves to 4 elements (a `<label>` wrapping several controls) |
+| `settings.spec` social-links | 1 | fails on base (known placeholder collision) |
+| `site-scope.spec` dashboard stats | 1 | fails on base |
+| `cut-release.spec` (both) | 2 | expected with a dirty working tree — the spec rewrites the real `editor/CHANGELOG.md` and makes a git commit |
+| `scheduler.spec:28` — queues due channels | 1 | fails on base (verified 2026-07-29 against `bbaa887`) |
+| `sync-break-on-existing-tree.spec` | 1 | **passes in isolation** — order-dependent flake, not a base failure |
+| `undownloaded.spec:53` | 1 | **passes in isolation** (all 4 in 22.6 s) — the CPU-contention flake already recorded below, seen again 2026-07-27 |
+
+**Re-verified again 2026-07-27** after Stage B2: **369 passed / 21 failed of 390** (the suite
+grew by the 5 new digest specs, all passing). The 21 are the same set, with two differences
+worth noting: `cut-release` failed **once** rather than twice, and `undownloaded.spec:53`
+appeared in place of the second `cut-release` — it passes in isolation. Export suite:
+**141/141**.
+
+**Re-verified again 2026-07-29** after Phases 2/3: **366 passed / 24 failed of 390**. The 24
+are the 22 in the table above (`cut-release` failed **both** this time — the working tree was
+dirty, which is exactly when it does) plus **two more, each verified passing in isolation**:
+
+| Spec | Isolation result |
+| --- | --- |
+| `auto-queue.spec:209` — strict priority drains the high-priority channel first | passes (13/13 in 1.8 min) |
+| `scheduler.spec:239` — schedule page edits the global controls | passes (all 8 scheduler tests) |
+
+Both are the **CPU-contention flake** already recorded for `undownloaded.spec:53` and
+`sync-break-on-existing`: this run shared the box with an unrelated containerized playwright
+suite at load ~10. Add them to the "passes in isolation" set rather than the base-failure
+set. Export suite the same day: **150/150** (141 + 9 new digest specs). `digest.spec.ts`
+(12) and the `settings.spec` digest tests all passed.
+
+The `deploy-page` (×3), `new-channel-onboarding` (×5) and `site-scope` failures — **9 specs**
+— are all *strict-mode locator violations* ("resolved to N elements"): specs written against
+an earlier form structure that has since gained sibling controls with overlapping accessible
+names. They are **a real backlog item, not a regression**, and cheap to fix independently of
+any feature work: each needs a tightened locator (`exact: true`, a `getByRole` scoped to the
+form, or a `data-testid`), not a behaviour change. Fixing them is worth doing precisely
+because 9 permanently-red specs train everyone to ignore the suite's exit code.
+
+`widget.spec`'s "no buttons" assertion fails under `dev:test` because `next dev` injects a
+Dev Tools button; it passes under `E2E_MODE=start`. (It passed in this run.)
+
+The `digest.spec.ts` tests all pass — now **12**, the original 9 plus three for the per-video
+review panel (renders provenance + warnings, flips to stale after an identity change,
+corrections land in the overrides sidecar with the machine file untouched). `settings.spec`
+gains two more that also pass: the Digest fieldset round-trips, and an unrelated save does
+not reset `timestampMode`/`promptVariant`.
+
+Run `pnpm e2e` in **default dev mode** — `E2E_MODE=start` serves a stale build. Kill stale
+dev servers by port between runs.
+
+---
+
+## Dependencies present / absent
+
+Present: `execa`, `lmdb`, `flexsearch`, `p-limit`, `fs-extra`, `@tanstack/react-query`,
+`@tanstack/react-virtual`, `markdown-to-jsx`, `tsx`.
+
+**Absent — do not assume:** any YAML parser, `gray-matter`, `zod`, `vitest`, `jest`.
+A frontmatter parser for Phase 1.5 must be hand-rolled (three scalar/list field types).
+
+---
+
+**Re-verified 2026-07-29** after the backfill-readiness work: **368 passed / 22
+failed of 390** — one better than the previous run. All 22 are in the table
+above. Two were isolated rather than assumed:
+
+| Spec | Isolation result |
+| --- | --- |
+| `duplicate-shorts.spec:328` | **was a real regression, now fixed.** New review buttons carried `aria-label="confirm duplicate cluster <id>"`, and `getByLabel` matches by SUBSTRING, so the card's own `duplicate cluster <id>` locator resolved to 3 elements. Renamed; 3/3 green. **Third occurrence of this collision in this repo.** |
+| `scheduler.spec:28` | fails on base too — verified by checking out `bbaa887` and re-running. Add it to the base-failure set. |
+
+Common tests **374** (was 356). Export playwright **150/150**. `tsc` clean in all
+three packages.
+
+**`cut-release.spec` mutates the repo, and it did again**: it rewrote
+`editor/CHANGELOG.md` AND created a real `Release editor 9.9.10` commit. Reset
+after the run. Budget for this every time the full suite is run with a dirty or
+in-progress changelog.
+
+## Smoke test results (2026-07-25)
+
+Run against a real 176-minute podcast transcript (`UndertheTeaVT/data/5JRFDQ7TtZ8`, 3877
+cues, ~34k tokens) on the RX 6600.
+
+1. **Ollama silently truncates.** `ollama ps` reported `CONTEXT 4096` — the default. A 10k
+ token excerpt was cut without warning and the model summarized whatever fragment
+ survived. **`num_ctx` must be set explicitly per request.** 16k is comfortable on this
+ GPU (qwen2.5:7b KV cache ≈ 56 KB/token → ~0.9 GB at 16k, atop 4.7 GB of weights).
+2. **fabric's stock pattern fails at this model size.** With `create_video_chapters`
+ verified applied (via `--dry-run`) *and* an input small enough to fit, qwen2.5:7b ignored
+ the `HH:MM:SS TOPIC` contract entirely and returned conversational prose. Twice, at two
+ input sizes. It is a long discursive prompt written for frontier models.
+3. **Schema-constrained decoding fixes it.** POSTing to `/api/chat` with a JSON schema in
+ `format` returned clean parseable output in **15.5 s**.
+4. **It still hallucinates timestamps.** That same run emitted `01:10:29` for an input
+ spanning only `00:04:45 → 00:15:36` — 55 minutes past the end. **The parser's range,
+ monotonicity, and cue-snapping guards are load-bearing, not defensive polish**, and
+ clamping must be per-chunk: a whole-video range check would have accepted that value.
+5. **GPU offload works** — `100% GPU` on the RX 6600 via Vulkan.
+
+### Environment state as of that run
+
+`ollama-vulkan` 0.32.4 and `fabric-ai` 1.4.375 installed. The ollama systemd unit ships
+**inactive and disabled**; no `~/.ollama`, no models. `~/.config/fabric` has no `patterns/`
+and no `.env`, so `create_video_chapters` does not exist on disk until `fabric-ai --setup`
+clones the pattern repo.
+
+Note: Arch names the binary **`fabric-ai`**, not `fabric`, to avoid colliding with the
+unrelated Python `fabric` deployment tool. Moot while fabric is deferred, but relevant if it
+is ever added as an engine.
diff --git a/plans/README.md b/plans/README.md
@@ -0,0 +1,45 @@
+# `plans/` — how this work remembers itself
+
+The local-AI derived-corpus work spans many phases and months, and **context is cleared
+between phases** to keep token cost down. That only works if a cold agent can resume without
+re-deriving anything. Four artifacts with deliberately different lifetimes:
+
+| File | Lifetime | Contents |
+| --- | --- | --- |
+| [`../PLAN.md`](../PLAN.md) | Rarely changes | The durable roadmap: why, architecture, all phases, sequencing. |
+| [`FACTS.md`](FACTS.md) | Append-only | Verified codebase facts with `file:line` anchors. |
+| [`STATE.md`](STATE.md) | Rewritten each session | Phase status, decisions log, open questions, commands known to pass. |
+| `phase-N-<slug>.md` | Written at phase start, archived at merge | File-level detail for the phase in flight. |
+
+`FACTS.md` is the most valuable file here. It holds exactly the material that is expensive to
+establish and cheap to get wrong — the original draft of this plan contained several
+confident, wrong claims about how the build pipeline works, and each one would have cost real
+implementation time. The roadmap is stable prose; the ledger grows every time something is
+checked against the tree.
+
+## Context-clear protocol
+
+**Before clearing context:**
+
+1. Update `STATE.md` — phase status, decisions made and *why*, anything surprising.
+2. Append newly verified facts to `FACTS.md`, each with a `file:line` anchor.
+3. **Commit.** An uncommitted memory file does not survive a container reclaim.
+
+**After clearing context:**
+
+`AGENTS.md` arrives automatically (`CLAUDE.md` is just `@AGENTS.md`), and it points here.
+Read `STATE.md`, then `FACTS.md`, then the current `phase-N-*.md`. That is the whole warm-up
+and it should cost a few thousand tokens rather than a re-exploration.
+
+**Never re-verify something already in `FACTS.md`** unless the surrounding code changed. If a
+fact turns out to be stale, correct it in place and note the correction in `STATE.md`.
+
+## Writing a phase plan
+
+Start it when the phase starts, not before — a plan written three phases early is written
+against a tree that no longer exists. Keep it to what `PLAN.md` deliberately omits: exact
+files, exact edits, the order to make them in, and how to verify. Everything general belongs
+in `PLAN.md`; everything verified belongs in `FACTS.md`.
+
+Archive rather than delete on merge — `git mv plans/phase-N-*.md plans/done/` — so the
+reasoning behind a shipped phase stays findable.
diff --git a/plans/STATE.md b/plans/STATE.md
@@ -0,0 +1,529 @@
+# Running state
+
+The working memory for the local-AI derived-corpus work. Rewritten at the end of every
+session, before context is cleared. See [`README.md`](README.md) for the protocol.
+
+**Last updated:** 2026-07-29 — **the backfill is at the starting line.** A corpus-wide sweep
+can now be started from one control, survives a server restart, can be paused without being
+lost, yields the GPU to transcription, and reports coverage and an ETA in the unit the work
+is actually priced in. Stages A–D of the "cheapest-first" plan are all in; see the decisions
+below for what the measurements changed.
+
+**The headline is a negative result, and it should be read before planning any more cost
+work.** The plan's own premise — shrink the work-list via duplicate sharing — does not pay.
+Priced in audio-hours by the new `common/bin/digest-plan.ts`:
+
+| `--near` | clusters | to generate | sweep days | saved by sharing |
+| --- | --- | --- | --- | --- |
+| 0.6 (was) | 2,846 | 76,804 h | 80.0 | 0.5 d |
+| 0.45 | 7,110 | 75,851 h | 79.0 | 1.5 d |
+| **0.35 (now)** | **7,434** | **75,638 h** | **78.8** | **1.7 d** |
+| 0.25 | 7,493 | 75,616 h | 78.8 | 1.8 d |
+
+Loosening the threshold as far as it goes buys **1.2 days of 80**. The plan assumed mirrors
+skew long; **they skew short**. Cluster members are 19% of the corpus by video count and
+**4.3% of its audio-hours** — the sweep is dominated by unclustered long-form VODs, and
+`HasanAbiVODs3` alone (8,331 audio-hours) outweighs every mirror in the corpus combined.
+Only ~55% of mirrors pass the alignment gate, so half the nominal saving is refused anyway.
+
+**The lever that does pay is GPU arbitration**: 90 s/audio-hour measured against 27 s
+projected on an idle box is ~80 days versus ~24 — about fifty times every duplicate lever
+put together. It is now implemented (`controller/digestYield.ts`) and **still unmeasured in
+production**; measuring it is the single most valuable next action.
+
+**The sweep was run for real, not just built.** Scoped to `teamrcn` (7 videos,
+0.27 audio-hours) against ollama and the real corpus: it armed with its scope
+persisted, launched the per-channel job, generated 7 digests with correct
+provenance at `promptVersion` 2, refreshed the snapshot and ended clean. A second
+process then resumed **from settings alone**, found nothing to do, launched
+**zero** jobs and exited — the B2 claim proven, since the correct amount of
+rework after a restart is none.
+
+Doing that found two bugs no test would have. Every channel's progress bar would
+have stalled one short forever (`countMissingDigests` counted a transcript-less
+directory the batch correctly drops — 2,987 such videos corpus-wide), and
+stopping a sweep left its channel scope behind, so the next sweep would silently
+inherit a weeks-old list and report itself finished. Both fixed and re-verified.
+
+Throughput datapoint, uncontended: **~151 s per audio-hour** on 2.4-minute
+videos. Consistent with "short videos cost ~2x" against the 90 s corpus average.
+It is **not** the measurement that matters — that is still contended vs
+uncontended, and it is still not done.
+
+**Previous entry —** 2026-07-29 — **Phases 2 and 3: the digest layer SHIPS.** The 102 digests
+already on disk now travel generation → LMDB → a shared `/digests/<slug>/` page tree →
+compose → the viewer, where a reader can open a Digest panel, click a chapter and seek to it.
+Proven end to end against the real corpus, not fixtures: `community-notes/v2cywen` renders its
+20 real chapters (first at 01:09:48 — the validation run's worst coverage gap) and the
+playhead follows a click. The control is hidden on the 99.9% of videos with no digest, and a
+borrowed digest is labelled as borrowed.
+
+---
+
+## Phase status
+
+| Phase | Status | Notes |
+| --- | --- | --- |
+| 0 · Benchmark transcription engines | not started | Still unmeasured. The DIGEST bake-off is done; this is the separate transcription one. |
+| 1 · Generation harness | **done + validated** | Spine in `8c041fd`; correctness + bake-off + e2e in Stage B1; operator surfaces + measured-best defaults + 102-video validation run in Stage B2. |
+| 1.5 · Channel context | not started | **Not a backfill gate** — the empty note hashes stably, so a note added later invalidates only its own channel. The gate is the `digest-context-v1` prefix decision. |
+| 2 · Digest corpus in build | **done** | Shared `/digests/<slug>/` page tree, `digests` + `digestPageHashes` + `channelDigestStats` sub-DBs at `SCHEMA_VERSION` 13, `digestMs` in the mtime record and the channel signature, compose reconcile, `CORPUS_SPEC_VERSION` 3. |
+| 2.5 · Observability | **done** | Corpus coverage on the dashboard + widget, `noDigest` counters, a digest instrument, audio-hour ETAs, and the two progress bugs fixed. |
+| 3 · Viewer `?vm=digest` | **done** | Not `?vm=summary` — see the naming decision below. |
+| 4 · Search indexing | not started | |
+| 5 · Auto-queue | not started | |
+| 6 · Ollama `/ask` provider | not started | **No dependencies.** |
+| 7 · Tags + chat highlights | not started | Tag generation works; nothing consumes it. |
+| 8 · Visibility policy | not started | |
+| 9 · Attribution + quote filtering | not started | |
+| 10 · Lead with the derived corpus | not started | |
+| 11a · Review queue | **minimum landed** | Total failures now persist warnings (they used to persist none), `buckets.digestWarnings` + a `digest_warnings` filter + an /actionable section, and duplicate-cluster confirm/reject. Approve/dismiss state deferred — needs new persistence. |
+| 11b · Viewer feedback | not started | |
+
+**Three corrections to this document's own claims, verified against the code 2026-07-29.**
+Exploration for Phase 2 found the planning docs describing more unbuilt work than exists:
+
+- **Phase 2.5 is not "not started".** `JobProgressMetric` already carries `"digests"` and
+ `JobTaskKind` already carries `"digest"` (`common/jobs/registry.ts:17,25`).
+- **PLAN.md's Phase 2.5 opens by demanding the duplicated-metric-union fix "first". It is
+ done.** `RunningJobsList.tsx:6-9` imports the union instead of re-spelling it, and
+ `MonitorWidget.tsx:628` is a `METRIC_PREFIX` lookup table, not the two copies of a binary
+ ternary the old FACTS.md table described. That table is now marked RESOLVED there.
+- **`listChannels` already pays for a digest count and then throws it away.**
+ `common/controller/channels.ts:192` calls `countDataFiles`, which computes `counts.digests`
+ — but unlike `readChannelStat` (`:139`), `listChannels` omits `digestCount` from the
+ `ChannelStat` it emits. The field is already optional on the type, so recovering it is one
+ line and every cross-channel surface gets a counter it is already paying for. Left
+ unchanged here deliberately: it is Phase 2.5's to take, not a digest-corpus side effect.
+
+**Recommended next**, in the order the measurements argue for:
+
+1. **MEASURE THE GPU YIELD IN PRODUCTION.** Everything else is second by a factor of fifty.
+ Start a transcription job during a digest sweep, confirm the digest lane idles rather than
+ contending, and measure seconds-per-audio-hour with `yieldToTranscription` on and off. If
+ the 2.8x is real the sweep is ~24 days, not ~80, and every other estimate in these docs
+ changes. If it is NOT real, the contention hypothesis is wrong and the 90 s/audio-hour
+ needs a different explanation — which is just as valuable to know before spending 80 days.
+2. **Run the sweep for a day and watch it.** The launcher, resume, pause and coverage
+ readouts are all built and unit-tested but have not driven a real multi-channel run. Kill
+ the editor mid-run and confirm it resumes with zero rework (eligibility is re-derived from
+ disk, so the correct result is zero).
+3. **Decide the `digest-context-v1` hash-prefix question before sweeping.** This is the real
+ 1.5 gate, and it is a decision, not work: bumping the prefix invalidates the entire
+ corpus, so it must happen before 80 GPU-days go in, or not at all.
+4. **Then Phase 4 search indexing / Phase 7 tags,** which are starved until coverage exists.
+
+Deferred deliberately: Phase 11a approve/dismiss state (needs a new sidecar field or sibling
+file), Phase 1.5 as fully specced, Phase 6 Ollama `/ask` (genuinely independent — a good
+parallel task), Phase 0 transcription benchmark (a separate bottleneck).
+
+---
+
+## Decisions log
+
+Recorded with reasons, because these are exactly what a cold agent would otherwise
+relitigate. Earlier entries (fabric deferred, ollama-direct as the structured default,
+Claude Code on a second queue key, digests in their own page tree, per-section provenance,
+transcripts never rewritten) still stand and are unchanged.
+
+**Cluster sharing was DEAD for 14.5% of the corpus — the half it was written for
+(2026-07-29).** `digestBatch` looked the plan up with `${channelSlug}/${directoryName}`; the
+plan is keyed `${channelSlug}/${metadataId}`. Those differ for **11,175 of 77,106 videos**,
+and not evenly: **7,870 are `the-quartering-rumble`, where dir !== id for every single
+video** — precisely the mirror set `digestSharing.ts`'s header cites as the win. A miss is
+silent and looks like success (the mirror is simply not recognised, so it regenerates), which
+is why it survived a bake-off, a pilot and a 102-video validation run. `videoDirForSlug` had
+the same bug from the other side. Fixed by recording `DuplicateVideoRef.videoDir` — which the
+detector already knew and threw away — only where it differs from `id`.
+
+**A cost lever is only real in AUDIO-HOURS (2026-07-29).** See the table at the top. The
+general lesson, and it is the same one the validation run taught about throughput: **the
+corpus's video count and its audio-hours are differently distributed, so any estimate that
+counts videos is measuring the wrong thing.** `bin/digest-plan.ts` exists so the next lever
+gets priced before it gets built.
+
+**`NEAR_THRESHOLD_DEFAULT` is 0.35 (2026-07-29), and it is a PUBLISHING change.** Bracketed
+at 0.6/0.45/0.35/0.25 on identical inputs; every step down is a strict superset. Marginal
+bands were read, not counted: of the 4,264 clusters admitted at 0.45, 96.2% have
+byte-identical titles, 99.5% are cross-platform, 3 are same-channel. The recall knee is at
+**0.45** — take that if a more conservative assertion is ever wanted. 0.35 was chosen because
+its band is still clean and the report's job is to surface real mirrors to readers. It was
+changed in its own commit because it changes what every built site asserts, NOT because it
+saves sweep time (it saves 1.2 days of 80).
+
+**`updateDuplicateOverride` had ZERO CALLERS before this session.** It shipped complete —
+the `confirmed` flag, the read-modify-write, the atomic rename, and a careful rule that an
+empty patch clears a decision without deleting a confirmation — and nothing in the repo ever
+invoked it. Because `clusterMaySharePartial` fails CLOSED, that meant ~166 `needsReview`
+clusters could never be confirmed by anyone, and the review queue asked a question with no
+way to answer. `shareClusterFromCanonical` was in the same state. Both are now wired from
+/actionable. **Worth generalising: "the extension point exists" is not the same as "the
+feature works", and a fails-closed gate with no way to open it is indistinguishable from a
+missing feature.**
+
+**Total digest failures persisted NOTHING, and that was the review queue's data source
+(2026-07-29).** Both total-failure paths in `digestVideo` skipped `writeDigestSection` for a
+good reason — an empty section with current provenance reads as fresh and the video is never
+retried — and the consequence was that the worst outputs left evidence only in a job log that
+rotates. `DigestRecord.failures` is a SIBLING of `sections`, so `isSectionFresh` cannot see
+it and the retry is preserved. It records `no-output` vs `all-rejected` separately, because
+"the model proposed nothing" and "the model proposed 13 chapters and a guard clamped them
+all" look identical from outside and need opposite fixes.
+
+**Writing the test for that found a second bug, of a shape this repo has now hit twice.**
+`loadDigest` rebuilds the record field by field rather than spreading, so `failures`
+round-tripped to nothing — the write succeeded and only the read omitted it. Structurally
+identical to Phase 2's `pageHashes: [""]`, and caught the same way: by asserting on what came
+BACK, not on what went in. Anything added to `DigestRecord` must be added to that reader.
+
+**The viewer mode is `?vm=digest`, NOT `?vm=summary` (2026-07-29).** "Summary" was already
+taken twice — `DisplaySummary` is a video *listing card* (`common/lib/transcripts.ts`) and
+`/summaries/` is the browse-index page tree the search index serves. PLAN.md's own naming
+rule ("`report` already means three different things — do not add a fourth") applies exactly
+as well to a third meaning of "summary". `digest` already names the artifact, the stage, the
+settings section, the editor panel and the page tree.
+
+**What ships is `effectiveDigest()`, and `warnings[]` is deliberately NOT in it
+(2026-07-29).** The page tree carries the COMPOSED digest — human overrides applied,
+`enabled: false` items dropped, chapters sorted — so a reader sees what a human approved
+rather than raw model output. `warnings[]` and `history[]` stay behind: they are operator
+telemetry for the editor panel and the Phase 11a review queue. Shipping "the model proposed
+13 chapters that were clamped away" puts the corpus's failures in front of the audience
+instead of the operator, and `history[]` is the largest field in a mature sidecar.
+
+**The digest control is HIDDEN where there is no digest, and the manifest is what says so
+(2026-07-29).** Coverage is 102 of ~77,108 videos — 0.1% — so an always-present button is a
+dead end almost everywhere. `slugToPage` doubles as the existence check (a channel with no
+digests ships no manifest at all), so gating costs one small cached request per channel and
+never a per-video fetch. Same shape as `hasDuplicates()` gating the Duplicates nav link.
+
+**The client cache is versioned by PAGE CONTENT HASH, not `provenance.generatedAt`
+(2026-07-29).** Digests are regenerated in place, so a cached entry stays structurally valid
+while going stale — the shape sniff `transcriptStore.ts` gets away with cannot see that.
+The plan called for storing `generatedAt`, but that is only legible *after* fetching the
+page, which is the cost the cache exists to avoid. The build already computes per-page
+hashes for its own skip logic, so `ChannelDigestsManifest.pageHashes` publishes them: exact
+versioning, knowable from the manifest alone, and it also catches an overrides-only retitle
+that leaves `generatedAt` untouched. **This was not theoretical** — the first build emitted
+`pageHashes: [""]` because an LMDB read-back immediately after `put()` returned nothing.
+The hashes are now returned directly by `createPageWriter` (recorded on BOTH the written and
+the skipped branch), which is correct by construction rather than by LMDB semantics.
+
+**Digests get their own IndexedDB DATABASE, not another store in the transcript one
+(2026-07-29).** `DB_VERSION` is a property of the database and `transcriptStore.ts`'s
+`onupgradeneeded` does `deleteObjectStore` + `createObjectStore`, so adding a store there
+would force a version bump and wipe every client's transcript cache. `searchLayerCache.ts`
+already establishes the separate-`DB_NAME` idiom; digests follow it.
+
+**`chunk-local` timestamps are adopted, on measurement — and (2026-07-27) SHIPPED as the
+default.** Each chunk's transcript is re-based to `00:00:00` and the parser adds the offset
+back before any guard runs. Measured on the long tail it cut zero-yield chunks from 33.3% to
+11.1% (qwen2.5@16k), 29.4% to 11.8% (qwen2.5@8k) and 17.7% to 0% (gemma2), with the
+rejection rate falling in every case. The working hypothesis — the model cannot hold a large
+absolute offset across a long chunk and reverts to counting from zero — is confirmed. It was
+shipped as a *scored variable*, not a pre-applied fix, which is why there are numbers to
+quote.
+
+For a while it was adopted *on paper only*: `DEFAULT_DIGEST_TIMESTAMP_MODE` stayed
+`"absolute"` and `DEFAULT_DIGEST_NUM_CTX` stayed 16384, so a fresh install — and the pilot —
+ran the configuration round 2 measured as **worst**. Both defaults now carry the decision
+(`chunk-local`, 8192, 600 cues).
+
+**The mechanism for that change was a `PROMPT_VERSION` bump (1 → 2), and it had to be.**
+`digestPromptVariant()` derives the variant *relative to the default constants*, so a value
+equal to the default contributes nothing and the variant is `undefined`. Flipping the
+defaults alone would therefore have made the 39 pilot digests — written under
+absolute/1200 with `promptVariant: undefined` — compare **fresh** under the new defaults:
+bake-off-losing output frozen into the corpus, indistinguishable from the best config's.
+`isSectionFresh` compares `promptVersion`, so bumping it invalidates every section
+explicitly. That is the right division of labour: **`promptVersion` for a default change,
+`promptVariant` for several shapes coexisting during a bake-off round.** Cost: 39
+regenerations. `digestPrompt.test.ts` pins both the new defaults and the bump so the pairing
+cannot be broken silently.
+
+**Chunk size is sized to `num_ctx`, and it is a bigger lever than expected.** Halving the
+context (16k → 8k, 1200 → 600 cues) took qwen2.5:7b from 9.34 to 13.37 chapters/hour and its
+worst coverage gap from 1:27:48 to **24:13**. The cost the plan assumed did not appear:
+24.2 vs 25.1 projected sweep days. Twice the calls at half the prompt each is the same
+seconds-per-audio-hour. Production previously hardcoded 1200 cues regardless of `num_ctx`,
+so lowering the context would have silently truncated every call — `maxCuesForContext()`
+now derives it.
+
+**Anything that changes the output is in the freshness identity.** `promptVariant` folds in
+`timestampMode` and a non-default chunk size, derived in ONE place
+(`digestPromptVariant`). The default configuration maps to `undefined`, so every digest
+written before the field existed still compares fresh. The alternative — an operator
+remembering to bump a label — fails silently by skipping the whole corpus as "fresh".
+
+**One freshness target, three callers (2026-07-27).** `resolveDigestTarget()`
+(`common/controller/digestTarget.ts`) is now the only place a target is derived, and the
+batch runner, `countMissingDigests` and the channel snapshot's `noDigest` bucket all use it.
+Two counters were previously lying:
+
+- `noDigest` asked only `engine && hasItems` — "is there an ai-digest.json with something in
+ it?" — while its own label and `stageStatus.ts` claimed an identity check. After any config
+ change the Digest stage read **"All digested"** while `countMissingDigests` reported the
+ whole channel stale. The two could disagree by an entire channel, which is exactly how a
+ sweep reports "nothing to do" on a corpus that needs redoing. It now computes the real
+ target; the per-video cost is zero extra I/O (the sidecar is already loaded and
+ `isSectionFresh` is pure), and only the target itself is new work, resolved once per
+ channel. Coverage-by-engine (`digestEngines`) stays deliberately identity-**blind** — a
+ section made by an older prompt version was still made by that engine.
+- `countMissingDigests` always resolved `settings.localAppId`, so a metered-lane job's
+ progress bar was sized against the *local* engine's identity (and, once `numCtx` joined the
+ identity, the local app's context size). It takes the lane as a parameter now; both call
+ sites in `digestActions.ts` already knew which lane they were.
+
+A digest shared from a duplicate cluster's canonical member counts as done in the bucket —
+the canonical member's own freshness drives regeneration and the share is re-applied from it.
+
+**The 102-video validation run says the bake-off was right about quality and wrong about
+time (2026-07-27).** Full table in FACTS.md. Quality came in at or better than round 2's
+projection on every metric — **zero-yield went to literally 0 of 153 chunks**, chapters/hour
+14.74 vs 13.37, rejection rate 18.4% vs 19.1% predicted, generic titles 27.0% vs 31.2% — so a
+2-video sample turned out to be representative of quality. Throughput was not: **90 s per
+audio-hour vs 27**, projecting **~81 sweep days rather than 24.2**, because round 2 measured
+only long videos on an idle box and the corpus is mostly short videos on a box also running
+whisper. The lesson to carry: **a bake-off sample stratified by content generalizes; one
+stratified by length does not generalize to throughput.**
+
+The one metric that degraded is the worst coverage gap (1:09:48 vs 24:13 projected), and the
+new per-video panel explains it rather than leaving it a mystery: the worst video recorded 13
+`out-of-range` rejections, so the model *did* propose chapters for the missing hour and the
+clamp threw all of them away. "Never proposed" and "proposed and rejected" need different
+fixes, and only a recorded-warnings artifact can tell them apart. Median gap is 5:19 — quote
+the median with the max.
+
+**The sweep is ~25 days, not ~64.** Round 1 measured 59 days on short+medium; Round 2
+measured ~25 on the long tail, because long videos amortise the fixed per-call overhead and
+the corpus is dominated by them. Throughput is no longer the binding constraint it was
+assumed to be.
+
+**Corpus-wide duplicate detection is ON — the earlier "it does not complete" conclusion was
+wrong and has been reversed.** The 507 M-pair measurement was correct; the inference from it
+was not. Two independent defects, neither inherent to corpus-wide detection: an unblocked
+containment cartesian (498 M pairs, 98.2% of the work) and holding every video's shingle set
+at once. Blocking fixes the first, block-at-a-time streaming fixes the second. **Both
+strategies now complete corpus-wide on a default heap**: title in **206 s at 2.0 GB peak
+RSS**, duration in **2,889 s at 2.45 GB** — the latter nominating 9,131,476 pairs, within
+0.005% of the 9,131,916 the old counting predicted. The measurement was right; the inference
+drawn from it was not. See
+[`FACTS.md`](FACTS.md#duplicate-detection-at-corpus-scale-measured-2026-07-26) for the full
+table and the measured recall comparison.
+
+**A blocking strategy NOMINATES; the transcript cascade DECIDES.** `--blocking
+title|duration|both` chooses only how pairs are proposed, never what counts as a duplicate,
+so choosing one trades recall against runtime rather than correctness. This is what makes
+nominating aggressively safe — and it earns its keep: **5,250 of 8,352 title-nominated pairs
+(63%) were rejected once their transcripts were compared.**
+
+**SUPERSEDED (2026-07-27) — that 63% was mostly the threshold, not the corpus.** This section
+originally read "the projected digest saving was overstated by ~5x": of ~8,204 title-level
+redundancies only 2,717 survived content comparison, and only 1,599 of those were
+timing-aligned. Measuring `--near` (the caveat that was attached to it) showed the rejections
+were overwhelmingly **real cross-platform mirrors transcribed by different ASR engines**. At
+`--near 0.35` the confirmed set is **7,329 pairs / 3,952 aligned** — 2.5× the banked saving,
+and inspection of all 4,456 new clusters found no false positives. Plan the backfill on
+**7,329 / 3,952**, and see the near-threshold entry under Open questions before treating 0.6
+as settled. Full evidence in FACTS.md.
+
+**A previous decision is deliberately reversed: title+duration clusters DO form.** The old
+rule was that only content-confirmed tiers cluster, because duration coincidence alone
+produced enormous false clusters. Title **and** near-identical runtime is a far stronger
+claim than duration alone, and such clusters are quarantined behind `needsReview`: they never
+auto-share derived work (`clusterMaySharePartial`) and never reach a built site
+(`clusterIsPublishable`) until a human records `confirmed`. The old invariant is kept exactly
+where it mattered — every *shipped* cluster is still content-confirmed — and the e2e
+assertion was narrowed to say so rather than deleted.
+
+**Alignment is measured at detection time and persisted.** `measureAlignment` already
+existed but ran only in `digestSharing` and threw its result away. `DuplicateVideoRef` now
+carries `offsetSeconds` + `aligned`, computed against the cluster's canonical member for
+confirmed clusters only. Nothing that seeks INTO a sibling can be honest without it: a mirror
+with a longer intro carries the same words at shifted times, so "jump to this moment" would
+land wrong while looking right. Absent (older report, or no timed cues) reads as NOT aligned.
+
+**The containment sweep is budget-capped, not deleted.** It is the one pass blocking cannot
+rescue — a clip and its parent share neither duration nor title. Under
+`MAX_CONTAINMENT_PAIRS` (5 M) it runs unchanged; over it, it logs that it was skipped and
+that clip-of-longer duplicates are not covered. Corpus-wide it reports ~531.9 M pairs and
+skips. Silent truncation was never an option: it reads as "covered everything".
+
+**The ollama e2e stub is an HTTP server, not a fake binary.** `ollama-direct` POSTs to
+`${ollamaUrl}/api/chat`, so the `e2e/fixtures/bin/` idiom does not apply; it is a third
+playwright `webServer` with `OLLAMA_URL` pointed at it in **both** `dev:test` and
+`start:test`. `fake-claude.mjs` IS a subprocess and follows the normal idiom.
+
+---
+
+## Open questions
+
+- **Title quality vs throughput.** `qwen2.5:7b@8192` + chunk-local is the adopted default at
+ ~24 sweep days, but 31.2% of its titles are generic ("Introduction and Context",
+ "Conclusion and Final Thoughts"). `gemma2:9b@8192` + chunk-local writes markedly better
+ titles (10% generic, 0% zero-yield) at 5.5x the wall-clock (132 days). Worth revisiting
+ once Phase 11a can measure review effort: better titles may be cheaper than the review
+ they save. gemma2 is hard-capped at an 8192 context by the model.
+- **The `verylong` (>6 h) stratum is unmeasured.** Round 2 used the `long` bucket (~3.5 h).
+ The sample already contains an 8 h video for it.
+- **ANSWERED (2026-07-27): yes — the 0.6 near threshold was rejecting real mirrors, in bulk.**
+ Re-running `--blocking title --near 0.35` corpus-wide takes confirmed pairs from **2,717 to
+ 7,329** (+170%) and timing-aligned mirrors from **1,599 to 3,952**, a strict superset (no
+ video in the 0.6 clusters is absent from the 0.35 ones). The 4,456 new clusters were
+ inspected, not merely counted: **99.3% are cross-platform *and* cross-channel**, 96% have a
+ byte-identical title, 90% agree on runtime within 2 s, and their scores pile into a tight
+ **0.45–0.60** band just under the old cutoff. Sampling the 138 riskiest (< 0.42) and all
+ 502 non-sibling cases found **no false positives**. Mechanism confirmed as hypothesised:
+ two *different ASR engines* transcribing the same audio agree at ~0.35–0.60 on 5-gram
+ Jaccard, so 0.6 was tuned for same-engine text and structurally failed the cross-platform
+ case. Full table + reproduction in FACTS.md. **The caveat is deleted, not carried.**
+
+ **Now a decision, not a question: should `NEAR_THRESHOLD_DEFAULT` drop from 0.6?** The
+ measurement says yes and the digest-saving ceiling rises 2.5× if it does. It was *not*
+ changed here, deliberately — detection output feeds every built site's Duplicates page and
+ search badges, so loosening it changes what the archive asserts to readers, and that is its
+ own reviewable change rather than a side effect of a digest task. What is still unmeasured
+ is where the floor is: 0.35 was the single probe, and the 638 pairs it still rejects have
+ not been examined. Recommended next: probe 0.25 / 0.45 to bracket the knee, eyeball the
+ remaining rejections, then change the default in its own commit.
+- **Recall of the title pre-filter — now measured, and the gap is 11%.** Running both
+ strategies corpus-wide and diffing the clusters: duration blocking finds **672 videos title
+ blocking misses** (the re-titled mirrors it structurally cannot see), title finds 68 that
+ duration misses (runtime drift past the bucket), and 2,613 clusters come out identical.
+ Title has **~89% of duration's video-level recall at 1/14th the wall clock** (206 s vs
+ 2,889 s). So the routine default is settled: title, with `--blocking both` when that 11%
+ matters. What remains open is whether to close the gap *cheaply* rather than by paying 14×:
+ (a) token-set / Jaccard over title *words* — cheap, but every loosening merges series
+ episodes and each merge is a false cluster a human must reject; (b) MinHash / LSH
+ signatures over transcript shingles — the general fix, and the same machinery that would
+ let the containment sweep scale. Neither is worth building until something needs better
+ than 89%.
+- **`duplicates.json` has no offline story.** `export/service-worker/site-sw.js` keys its
+ runtime strategies off `manifest.json` / `page-NNNN.json` paths (`:55`, `:147`, `:205`), so
+ a root-level JSON is uncached. Affects the search-result duplicate badge in PWA mode only —
+ it silently does not appear offline, which is the correct failure but an unstated one.
+- **ANSWERED (2026-07-27): the pilot's un-created job was a pre-hydration lost click.**
+ Reproduced against the real corpus on a production build. The click fires **no POST at
+ all**: the server-rendered button has no handler until React hydrates, so a click before
+ that is silently discarded — no request, no job, no log, no error. Instrumented, the first
+ click produced 0 POSTs and a click five seconds later started the job normally. There is no
+ dedupe guard in `runManagedFunction` to explain it because there was nothing to dedupe —
+ the action was never invoked. The window is worst exactly where the pilot hit it: a heavy
+ channel page under a dev server (30 s – 9.9 min to render). **Fixed** — `StreamActionLog`
+ disables its button until mounted, which also closes the same latent race for sync,
+ download and transcribe, and makes e2e clicks wait for enablement instead of losing them
+ (the same failure mode that made the charts specs flaky). Verified after the fix: one
+ click, one POST, job runs. The button is now trusted to drive a long sweep.
+- **Diarizer choice** — unchanged, decide at Phase 9 on real audio.
+- **Phase 11b collector** — **answered: no, not `r2-proxy/`.** `src/index.ts:35-41` rejects
+ every non-GET/HEAD method and `KEY_RE` at `:27` carries a comment stating the proxy must
+ never become a general oracle over the bucket. Making it a collector works directly against
+ that design intent. Its one reusable piece is the `RATE_LIMITER` binding. Use the
+ zero-infrastructure fallback the phase already names (copy-to-clipboard / download-JSON →
+ paste into an editor action), which needs no deployment and matches the existing idiom at
+ `PlayerProvider.tsx:131-136`.
+- **Transcription engine default** — still pending Phase 0 numbers.
+
+---
+
+## Commands known to pass
+
+Established against this branch:
+
+```bash
+pnpm --filter yt-dlp-transcript-common run test # 356 tests (was 344)
+pnpm --filter editor build
+pnpm e2e # default dev mode; E2E_MODE=start serves a stale build
+pnpm --filter export exec playwright test # 150 tests, separate suite on :3020
+```
+
+Run `tsc` per package — the brace form is mangled by fish:
+
+```bash
+cd common && pnpm exec tsc --noEmit
+cd editor && pnpm exec tsc --noEmit
+cd export && pnpm exec tsc --noEmit
+```
+
+A `SCHEMA_VERSION` bump means a FULL index rebuild: 77k videos, **~35 min** and it re-reads
+every transcript. Subsequent incremental builds are ~6 min (the scan of 77k dirs dominates;
+digest pages skip on content hash). Budget for it before bumping.
+
+**`pkill -f <pattern>` can kill its own shell.** `pkill -f 'next dev --port 3020'` matches
+the `bash -c` process whose command line *contains that string* — so chaining it before a
+test run silently kills the run (exit 144, no output). Bracket one character
+(`'next [d]ev --port 3020'`), or run the pkill as its own command. This is the same bracket
+trick already noted for `fixtures/bin/fake[-]`, for a different reason.
+
+Duplicate detection (writes `transcripts/duplicates.json`; back it up first if the current
+one matters):
+
+```bash
+cd common
+pnpm exec tsx bin/duplicate-shorts.ts --all-durations # title, ~3.5 min
+pnpm exec tsx bin/duplicate-shorts.ts --all-durations --blocking duration # ~48 min
+pnpm exec tsx bin/duplicate-shorts.ts --all-durations --blocking both
+```
+
+**Do not run a corpus-wide detection pass and `pnpm e2e` at the same time.** The detector
+pins a core for its whole run and the contention alone failed
+`partial-downloads-bucket`, `undownloaded` and `video-page` — all three pass in isolation.
+That is three spurious entries to chase in a suite that already has 21 known-red ones.
+
+Bake-off (never writes to the corpus):
+
+```bash
+cd common
+pnpm exec tsx bin/digest-bakeoff.ts --pick # ONCE — fixes the sample
+pnpm exec tsx bin/digest-bakeoff.ts --label rN --buckets long \
+ --modes absolute,chunk-local --candidates 'qwen2.5:7b@8192'
+```
+
+**Kill stale dev servers AND orphaned fixture binaries between e2e runs.** Killing a
+playwright run mid-flight leaves `e2e/fixtures/bin/fake-*.mjs` children orphaned; they spin
+at ~24% CPU each and dozens of them will drive the load average past 35 and make every
+subsequent HTTP request time out while the server still logs 200s. `pkill -9 -f
+"fixtures/bin/fake-"` is part of cleanup, not just killing by port.
+
+For anything driving the editor against the REAL corpus (the pilot), use **production
+mode** (`pnpm --filter editor build` then `pnpm start`): the dev server renders the
+39-video channel page in 30 s - 9.9 min and holds ~21% of RAM, while the production build
+serves the same page in 7.6 s.
+
+The known-failing-on-base list lives in
+[`FACTS.md`](FACTS.md#known-failing-on-base--not-regressions) — do not chase those.
+
+---
+
+## Surprises hit so far
+
+`windowCues()` is not a chunker (corrected in FACTS.md, and `chunkCuesForContext` was
+written for the job). Still true, plus:
+
+**Two bugs the unit tests could not have caught, both found by the new e2e.**
+`taskHooks.ts` deliberately gives digest tasks no progress parser and then called
+`downloadParser!.feed()` unconditionally — every digest job died on its first log line, and
+a sweep would have reported "0 generated, N failed" looking like an engine fault. And
+`editor/app/settings/actions.ts` rebuilds the digest settings block field by field, so the
+first save after the settings form lands would have silently reset `timestampMode` and
+invalidated every digest made under it. Both are fixed. The lesson is that the digest
+layer's failure modes live at the *wiring*, not in the pure functions.
+
+**Phase 2/3 (2026-07-29): the two things that would have shipped broken and looked fine.**
+Both were caught by building the real thing and looking at the output, not by types or tests:
+
+- **`pageHashes: [""]`.** Reading `digestPageHashes` back immediately after `put()` in the
+ same build returned nothing. The manifest was structurally valid and the site worked — the
+ cache version was just empty, which fails *open* (every read a miss) rather than loudly.
+ Had it failed the other way, clients would have pinned stale digests forever.
+- **A video's directory name is not its video id.** `community-notes/v2fkbw7` — the id the
+ plan named, and the one the editor's URL shows — holds video `v2cywen`. Page trees key
+ `slugToPage` by `summary.id`, so the digest manifest correctly contains `v2cywen` and not
+ `v2fkbw7`; the transcripts manifest has always done the same. Anyone hand-checking a
+ digest manifest against a `data/` listing will "find" a missing video that is not missing.
+
+**The bake-off's most useful finding was one nobody asked for.** The plan framed the choice
+as model-vs-model with chunk-local as the variable. The data showed gemma2's clean sweep was
+confounded with its forced 8k context, and testing chunk size directly turned out to matter
+as much as the timestamp mode — and to be free. A comparison worth running is one that can
+surprise you.
diff --git a/plans/bakeoff/round1.json b/plans/bakeoff/round1.json
@@ -0,0 +1,771 @@
+{
+ "label": "round1",
+ "sample": "/home/user/Projects/yt-dlp-transcript-browser/plans/bakeoff/sample.json",
+ "corpus": {
+ "videosScanned": 76354,
+ "videosWithTranscript": 73367,
+ "audioHours": 77298,
+ "longTailVideos": 6038,
+ "longTailAudioHours": 40153
+ },
+ "buckets": [
+ "short",
+ "medium"
+ ],
+ "scores": [
+ {
+ "candidate": {
+ "key": "gemma2:9b@8192/absolute",
+ "model": "gemma2:9b",
+ "numCtx": 8192,
+ "maxCues": 600,
+ "timestampMode": "absolute"
+ },
+ "videos": [
+ {
+ "slug": "the-quartering-rumble/v6ve4n0",
+ "bucket": "short",
+ "durationSeconds": 791,
+ "chunks": 1,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 3,
+ "maxGapSeconds": 360,
+ "engineSeconds": 78.767,
+ "inputTokens": 4644,
+ "outputTokens": 131,
+ "warningsByCode": {},
+ "genericTitles": 0,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:00:00 Trump Trolls Colbert While Setting Up a Trap for Democrats",
+ "00:04:00 Colbert's Cancellation and the Changing Landscape of Late Night Television",
+ "00:10:00 Trump's DC Move Forces Democrats to Address Crime"
+ ]
+ },
+ {
+ "slug": "chibi-reviews/VQykVuHd9xQ",
+ "bucket": "short",
+ "durationSeconds": 433,
+ "chunks": 1,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 2,
+ "maxGapSeconds": 248,
+ "engineSeconds": 51.266,
+ "inputTokens": 3630,
+ "outputTokens": 81,
+ "warningsByCode": {},
+ "genericTitles": 0,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:00:02 Crunchyroll's Translation Issues",
+ "00:04:10 Impact on Viewership Experience"
+ ]
+ },
+ {
+ "slug": "the-quartering/J0ySGwzP4Nw",
+ "bucket": "short",
+ "durationSeconds": 785,
+ "chunks": 1,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 3,
+ "maxGapSeconds": 331,
+ "engineSeconds": 104.735,
+ "inputTokens": 6119,
+ "outputTokens": 122,
+ "warningsByCode": {},
+ "genericTitles": 0,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:00:04 Tim Pool Discusses Censorship on Social Media",
+ "00:04:01 Twitter's Handling of the 'G Word'",
+ "00:09:32 Concerns About LGBTQ+ Content in Schools"
+ ]
+ },
+ {
+ "slug": "destiny/5nmDzKB23OU",
+ "bucket": "medium",
+ "durationSeconds": 4419,
+ "chunks": 4,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 20,
+ "maxGapSeconds": 804,
+ "engineSeconds": 408.8299999999999,
+ "inputTokens": 20127,
+ "outputTokens": 724,
+ "warningsByCode": {},
+ "genericTitles": 0,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:13:24 The Pure Concept of Religion",
+ "00:16:53 Assumptions and Definitions",
+ "00:20:02 Beyond Christianity",
+ "00:20:41 Reconciling Perspectives",
+ "00:20:56 The Importance of Openness",
+ "00:32:58 Methodological Approach to Incorporating Evidence for Moral Facts",
+ "00:34:02 Determining Which Brain States are 'Moral'",
+ "00:37:01 Skepticism and the Burden of Proof"
+ ]
+ },
+ {
+ "slug": "angryjoeshow/RPJqkewZP5I",
+ "bucket": "medium",
+ "durationSeconds": 3380,
+ "chunks": 3,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 15,
+ "maxGapSeconds": 899,
+ "engineSeconds": 233.37699999999998,
+ "inputTokens": 12297,
+ "outputTokens": 552,
+ "warningsByCode": {},
+ "genericTitles": 1,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:14:59 Game's Defense and Blame Shifting",
+ "00:15:02 Content Creators and Gamers Blamed",
+ "00:15:38 One Peg Response - Game's Issues Highlighted",
+ "00:16:21 Mixed Opinions on the Game",
+ "00:17:34 Studio's Responsibility and No Man's Sky Comparison",
+ "00:18:25 Microtransactions and Lack of Appeal",
+ "00:20:39 Jeff Keley's Role and Personal Taste",
+ "00:34:25 Twitter and Patch Notes Struggles"
+ ]
+ }
+ ],
+ "totals": {
+ "videos": 5,
+ "audioHours": 2.72,
+ "chunks": 10,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "zeroYieldRate": 0,
+ "kept": 43,
+ "chaptersPerHour": 15.78,
+ "maxGapSeconds": 899,
+ "meanGapSeconds": 528,
+ "genericTitleRate": 0.0233,
+ "duplicateTitleRate": 0,
+ "warningsByCode": {},
+ "rejectionRate": 0,
+ "engineSeconds": 877,
+ "tokensPerSecond": 55.2,
+ "secondsPerAudioHour": 322,
+ "projectedSweepDays": 288.1
+ }
+ },
+ {
+ "candidate": {
+ "key": "qwen3:8b@16384/absolute",
+ "model": "qwen3:8b",
+ "numCtx": 16384,
+ "maxCues": 1200,
+ "timestampMode": "absolute",
+ "think": false
+ },
+ "videos": [
+ {
+ "slug": "the-quartering-rumble/v6ve4n0",
+ "bucket": "short",
+ "durationSeconds": 791,
+ "chunks": 1,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 6,
+ "maxGapSeconds": 246,
+ "engineSeconds": 62.781,
+ "inputTokens": 4523,
+ "outputTokens": 215,
+ "warningsByCode": {},
+ "genericTitles": 0,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:00:00 Trump Trolls Colbert by Taking His Prime Time Job",
+ "00:02:03 Speculation About Secret Side Deals and Late Night Television",
+ "00:06:09 Paramount's Merger and the Late Night Show Format Crisis",
+ "00:08:31 Democrats Grapple with Crime and Trump's Strategy",
+ "00:10:06 Trump's Impact on Crime Perception and Democratic Vulnerabilities",
+ "00:12:01 Woke DAs and the Changing Crime Statistics"
+ ]
+ },
+ {
+ "slug": "chibi-reviews/VQykVuHd9xQ",
+ "bucket": "short",
+ "durationSeconds": 433,
+ "chunks": 1,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 4,
+ "maxGapSeconds": 230,
+ "engineSeconds": 34.821,
+ "inputTokens": 3576,
+ "outputTokens": 91,
+ "warningsByCode": {},
+ "genericTitles": 0,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:00:02 Crunchyroll's Translation Issues",
+ "00:01:00 Translation Errors in Gundam Episode",
+ "00:02:33 Unwatchable Episode Due to Subtitle Delays",
+ "00:03:23 Translation Problems in Overlord Series},{"
+ ]
+ },
+ {
+ "slug": "the-quartering/J0ySGwzP4Nw",
+ "bucket": "short",
+ "durationSeconds": 785,
+ "chunks": 1,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 7,
+ "maxGapSeconds": 168,
+ "engineSeconds": 76.327,
+ "inputTokens": 6067,
+ "outputTokens": 223,
+ "warningsByCode": {},
+ "genericTitles": 0,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:00:04 Tim Pool's War With Twitter",
+ "00:01:13 The G Word and Banning",
+ "00:04:01 Free Speech and Platform Rules",
+ "00:06:00 Twitter's Alleged Protection of Creeps",
+ "00:08:01 Twitter's Uneven Enforcement",
+ "00:10:00 Concerns About Teachers and Young Ones",
+ "00:12:05 Criticizing the Platform's Policies"
+ ]
+ },
+ {
+ "slug": "destiny/5nmDzKB23OU",
+ "bucket": "medium",
+ "durationSeconds": 4419,
+ "chunks": 2,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 1,
+ "kept": 8,
+ "maxGapSeconds": 1900,
+ "engineSeconds": 231.563,
+ "inputTokens": 16388,
+ "outputTokens": 527,
+ "warningsByCode": {
+ "out-of-range": 12
+ },
+ "genericTitles": 3,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:02:00 Exploring Moral Systems and Preferences",
+ "00:06:00 Meta Ethics and the Nature of Moral Truth",
+ "00:12:00 The Role of Neuroscience in Understanding Morality",
+ "00:18:00 Skepticism and the Burden of Proof",
+ "00:24:00 Normativity and Its Interpretations",
+ "00:30:00 Conclusion and Reflection on Moral Frameworks",
+ "00:36:00 Final Thoughts and Open Questions",
+ "00:42:00 Closing Remarks and Further Exploration"
+ ]
+ },
+ {
+ "slug": "angryjoeshow/RPJqkewZP5I",
+ "bucket": "medium",
+ "durationSeconds": 3380,
+ "chunks": 2,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 17,
+ "maxGapSeconds": 2397,
+ "engineSeconds": 218.47199999999998,
+ "inputTokens": 15792,
+ "outputTokens": 499,
+ "warningsByCode": {
+ "out-of-range": 1
+ },
+ "genericTitles": 2,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:00:28 Silent Hill and New Game",
+ "00:00:30 John Wick and God of War",
+ "00:00:40 Marathon and High Guard",
+ "00:00:50 Crimson Desert and AI",
+ "00:01:00 State of Play and AI Controversy",
+ "00:01:10 Conclusion and Final Thoughts",
+ "00:01:20 Wrap-Up and Farewell",
+ "00:01:30 End of Episode"
+ ]
+ }
+ ],
+ "totals": {
+ "videos": 5,
+ "audioHours": 2.72,
+ "chunks": 7,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 1,
+ "zeroYieldRate": 0.1429,
+ "kept": 42,
+ "chaptersPerHour": 15.42,
+ "maxGapSeconds": 2397,
+ "meanGapSeconds": 988,
+ "genericTitleRate": 0.119,
+ "duplicateTitleRate": 0,
+ "warningsByCode": {
+ "out-of-range": 13
+ },
+ "rejectionRate": 0.2364,
+ "engineSeconds": 624,
+ "tokensPerSecond": 76.8,
+ "secondsPerAudioHour": 229,
+ "projectedSweepDays": 204.9
+ }
+ },
+ {
+ "candidate": {
+ "key": "mistral-nemo:12b@16384/absolute",
+ "model": "mistral-nemo:12b",
+ "numCtx": 16384,
+ "maxCues": 1200,
+ "timestampMode": "absolute"
+ },
+ "videos": [
+ {
+ "slug": "the-quartering-rumble/v6ve4n0",
+ "bucket": "short",
+ "durationSeconds": 791,
+ "chunks": 1,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 3,
+ "maxGapSeconds": 297,
+ "engineSeconds": 106.735,
+ "inputTokens": 4552,
+ "outputTokens": 110,
+ "warningsByCode": {},
+ "genericTitles": 0,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:00:00 Trump's Ruthless Troll Against Colbert",
+ "00:04:09 Colbert's Late Show Cancellation and Trump's Involvement",
+ "00:08:14 Economic Challenges of Late Night Shows"
+ ]
+ },
+ {
+ "slug": "chibi-reviews/VQykVuHd9xQ",
+ "bucket": "short",
+ "durationSeconds": 433,
+ "chunks": 1,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 2,
+ "maxGapSeconds": 364,
+ "engineSeconds": 68.864,
+ "inputTokens": 3590,
+ "outputTokens": 73,
+ "warningsByCode": {},
+ "genericTitles": 0,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:00:02 Crunchyroll's translations issues",
+ "00:01:09 Gundam episode with subtitle delays"
+ ]
+ },
+ {
+ "slug": "the-quartering/J0ySGwzP4Nw",
+ "bucket": "short",
+ "durationSeconds": 785,
+ "chunks": 1,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 3,
+ "maxGapSeconds": 609,
+ "engineSeconds": 138.169,
+ "inputTokens": 6093,
+ "outputTokens": 107,
+ "warningsByCode": {},
+ "genericTitles": 0,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:00:04 Word meanings and political aisles",
+ "00:01:13 Tim Pool's Twitter ban for using the 'G word'",
+ "00:02:56 Twitter's advertising restrictions on Tim Pool"
+ ]
+ },
+ {
+ "slug": "destiny/5nmDzKB23OU",
+ "bucket": "medium",
+ "durationSeconds": 4419,
+ "chunks": 2,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 1,
+ "kept": 8,
+ "maxGapSeconds": 1880,
+ "engineSeconds": 550.203,
+ "inputTokens": 16390,
+ "outputTokens": 504,
+ "warningsByCode": {
+ "out-of-range": 10
+ },
+ "genericTitles": 2,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:26:10 Segment Transcript",
+ "00:27:58 Example of Moral System",
+ "00:34:30 General Claim about Psychology",
+ "00:36:17 Claim about Matrix Facts",
+ "00:40:22 Default Position on Normativity",
+ "00:41:04 Definition of Normativity",
+ "00:41:36 Christine Korsgaard's Theory of Normativity",
+ "00:42:19 Step Made by Christine Korsgaard"
+ ]
+ },
+ {
+ "slug": "angryjoeshow/RPJqkewZP5I",
+ "bucket": "medium",
+ "durationSeconds": 3380,
+ "chunks": 2,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 12,
+ "maxGapSeconds": 1707,
+ "engineSeconds": 497.004,
+ "inputTokens": 15847,
+ "outputTokens": 355,
+ "warningsByCode": {
+ "out-of-range": 1
+ },
+ "genericTitles": 0,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:28:27 Silent Hill",
+ "00:29:50 High Guard",
+ "00:34:41 Crimson Desert",
+ "00:36:02 Chicken Man",
+ "00:37:19 God of War Remake",
+ "00:38:40 AI Deepfakes",
+ "00:41:37 Copyright Infringement in AI-Generated Content",
+ "00:42:23 Cance 2.0"
+ ]
+ }
+ ],
+ "totals": {
+ "videos": 5,
+ "audioHours": 2.72,
+ "chunks": 7,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 1,
+ "zeroYieldRate": 0.1429,
+ "kept": 28,
+ "chaptersPerHour": 10.28,
+ "maxGapSeconds": 1880,
+ "meanGapSeconds": 971,
+ "genericTitleRate": 0.0714,
+ "duplicateTitleRate": 0,
+ "warningsByCode": {
+ "out-of-range": 11
+ },
+ "rejectionRate": 0.2821,
+ "engineSeconds": 1361,
+ "tokensPerSecond": 35,
+ "secondsPerAudioHour": 500,
+ "projectedSweepDays": 447.3
+ }
+ },
+ {
+ "candidate": {
+ "key": "qwen3:14b@8192/absolute",
+ "model": "qwen3:14b",
+ "numCtx": 8192,
+ "maxCues": 600,
+ "timestampMode": "absolute",
+ "think": false
+ },
+ "videos": [
+ {
+ "slug": "the-quartering-rumble/v6ve4n0",
+ "bucket": "short",
+ "durationSeconds": 791,
+ "chunks": 1,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 3,
+ "maxGapSeconds": 344,
+ "engineSeconds": 168.879,
+ "inputTokens": 4523,
+ "outputTokens": 119,
+ "warningsByCode": {},
+ "genericTitles": 0,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:00:00 Trump's Troll on Colbert and the Kennedy Center Honors",
+ "00:03:19 Late Night Television's Decline and the Impact of Trump Derangement Syndrome",
+ "00:09:03 Democrats' Struggle with Crime and Trump's Political Strategy"
+ ]
+ },
+ {
+ "slug": "chibi-reviews/VQykVuHd9xQ",
+ "bucket": "short",
+ "durationSeconds": 433,
+ "chunks": 1,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 1,
+ "maxGapSeconds": 431,
+ "engineSeconds": 92.794,
+ "inputTokens": 3576,
+ "outputTokens": 41,
+ "warningsByCode": {},
+ "genericTitles": 0,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:00:02 Translation Errors on Crunchyroll"
+ ]
+ },
+ {
+ "slug": "the-quartering/J0ySGwzP4Nw",
+ "bucket": "short",
+ "durationSeconds": 785,
+ "chunks": 1,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 3,
+ "maxGapSeconds": 301,
+ "engineSeconds": 207.878,
+ "inputTokens": 6067,
+ "outputTokens": 110,
+ "warningsByCode": {},
+ "genericTitles": 2,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:00:04 Introduction to the Controversy Over Word Usage",
+ "00:03:06 Tim Pool's Twitter Ban and His Response",
+ "00:08:07 Discussion on Twitter's Enforcement and LGBTQ+ Terminology"
+ ]
+ },
+ {
+ "slug": "destiny/5nmDzKB23OU",
+ "bucket": "medium",
+ "durationSeconds": 4419,
+ "chunks": 4,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 30,
+ "maxGapSeconds": 816,
+ "engineSeconds": 977.384,
+ "inputTokens": 19990,
+ "outputTokens": 954,
+ "warningsByCode": {
+ "seam-duplicate": 1
+ },
+ "genericTitles": 0,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:13:15 Different Reasoning Processes and the Concept of a Higher Power",
+ "00:13:45 Faith vs. Science in Practical Applications",
+ "00:15:05 The Role of Faith in Metaphysical vs. Material Claims",
+ "00:16:04 Philosophical vs. Societal Definitions of Religion",
+ "00:17:56 Assumptions About Religion and Their Implications",
+ "00:19:17 The Frustration with Narrow Interpretations of Religion",
+ "00:20:32 Conceding the Common Vernacular of Religion",
+ "00:32:50 The Challenge of Moral Facts and Neuroscience"
+ ]
+ },
+ {
+ "slug": "angryjoeshow/RPJqkewZP5I",
+ "bucket": "medium",
+ "durationSeconds": 3380,
+ "chunks": 3,
+ "chunksFailed": 2,
+ "zeroYieldChunks": 2,
+ "kept": 8,
+ "maxGapSeconds": 2871,
+ "engineSeconds": 201.428,
+ "inputTokens": 4098,
+ "outputTokens": 196,
+ "warningsByCode": {
+ "chunk-failed": 2
+ },
+ "genericTitles": 1,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:47:51 AI in Healthcare and Accuracy Concerns",
+ "00:48:10 AI's Role in Managing Health Records",
+ "00:49:00 Escaping AI Algorithms and Content",
+ "00:50:00 AI Memes and Cultural Impact",
+ "00:51:00 AI and the Future of Entertainment",
+ "00:52:00 AI's Environmental Impact",
+ "00:54:00 Water Usage and AI",
+ "00:55:00 Conclusion and Advertisements"
+ ]
+ }
+ ],
+ "totals": {
+ "videos": 5,
+ "audioHours": 2.72,
+ "chunks": 10,
+ "chunksFailed": 2,
+ "zeroYieldChunks": 2,
+ "zeroYieldRate": 0.2,
+ "kept": 45,
+ "chaptersPerHour": 16.52,
+ "maxGapSeconds": 2871,
+ "meanGapSeconds": 953,
+ "genericTitleRate": 0.0667,
+ "duplicateTitleRate": 0,
+ "warningsByCode": {
+ "seam-duplicate": 1,
+ "chunk-failed": 2
+ },
+ "rejectionRate": 0.0426,
+ "engineSeconds": 1648,
+ "tokensPerSecond": 24.1,
+ "secondsPerAudioHour": 605,
+ "projectedSweepDays": 541.3
+ }
+ },
+ {
+ "candidate": {
+ "key": "qwen2.5:7b@16384/absolute",
+ "model": "qwen2.5:7b",
+ "numCtx": 16384,
+ "maxCues": 1200,
+ "timestampMode": "absolute"
+ },
+ "videos": [
+ {
+ "slug": "the-quartering-rumble/v6ve4n0",
+ "bucket": "short",
+ "durationSeconds": 791,
+ "chunks": 1,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 4,
+ "maxGapSeconds": 297,
+ "engineSeconds": 6.85,
+ "inputTokens": 4515,
+ "outputTokens": 137,
+ "warningsByCode": {},
+ "genericTitles": 1,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:00:00 Introduction and Discord Promotion",
+ "00:01:49 Trump's Trolling of Stephen Colbert",
+ "00:03:56 Announcement of Trump Hosting Kennedy Center Honors",
+ "00:08:14 Late Night Television Challenges and Viewership Shifts"
+ ]
+ },
+ {
+ "slug": "chibi-reviews/VQykVuHd9xQ",
+ "bucket": "short",
+ "durationSeconds": 433,
+ "chunks": 1,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 2,
+ "maxGapSeconds": 224,
+ "engineSeconds": 3.725,
+ "inputTokens": 3568,
+ "outputTokens": 71,
+ "warningsByCode": {},
+ "genericTitles": 0,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:00:05 Crunchyroll's Translation Issues",
+ "00:03:49 Multiple Shows with Translation Errors"
+ ]
+ },
+ {
+ "slug": "the-quartering/J0ySGwzP4Nw",
+ "bucket": "short",
+ "durationSeconds": 785,
+ "chunks": 1,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 3,
+ "maxGapSeconds": 516,
+ "engineSeconds": 5.496,
+ "inputTokens": 6059,
+ "outputTokens": 109,
+ "warningsByCode": {},
+ "genericTitles": 1,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:00:48 Discussion on the redefinition of words",
+ "00:01:35 Tim Pool's ban from Twitter and his response",
+ "00:04:29 Journalist Tim Pool's criticism of Twitter's policies"
+ ]
+ },
+ {
+ "slug": "destiny/5nmDzKB23OU",
+ "bucket": "medium",
+ "durationSeconds": 4419,
+ "chunks": 2,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 1,
+ "kept": 6,
+ "maxGapSeconds": 1882,
+ "engineSeconds": 85.707,
+ "inputTokens": 16388,
+ "outputTokens": 507,
+ "warningsByCode": {
+ "out-of-range": 12
+ },
+ "genericTitles": 0,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:26:11 The easiest way to prove that ever exists",
+ "00:27:05 Does conceding the possibility of metaphysical truth change anything?",
+ "00:30:02 Is it possible to find moral facts by analyzing brain states?",
+ "00:40:00 Arguments against the existence of moral facts",
+ "00:41:03 Definition and explanation of normativity",
+ "00:42:17 Normativity in practical problems"
+ ]
+ },
+ {
+ "slug": "angryjoeshow/RPJqkewZP5I",
+ "bucket": "medium",
+ "durationSeconds": 3380,
+ "chunks": 2,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 1,
+ "kept": 3,
+ "maxGapSeconds": 2497,
+ "engineSeconds": 77.36699999999999,
+ "inputTokens": 15784,
+ "outputTokens": 363,
+ "warningsByCode": {
+ "out-of-range": 10
+ },
+ "genericTitles": 0,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:41:37 Copyright Issues with Unauthorized AI Creations",
+ "00:45:24 Sony Considering Delaying PlayStation Console Launch Due to AI Demand",
+ "00:48:10 AI Accuracy in Healthcare Records Management"
+ ]
+ }
+ ],
+ "totals": {
+ "videos": 5,
+ "audioHours": 2.72,
+ "chunks": 7,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 2,
+ "zeroYieldRate": 0.2857,
+ "kept": 18,
+ "chaptersPerHour": 6.61,
+ "maxGapSeconds": 2497,
+ "meanGapSeconds": 1083,
+ "genericTitleRate": 0.1111,
+ "duplicateTitleRate": 0,
+ "warningsByCode": {
+ "out-of-range": 22
+ },
+ "rejectionRate": 0.55,
+ "engineSeconds": 179,
+ "tokensPerSecond": 265.2,
+ "secondsPerAudioHour": 66,
+ "projectedSweepDays": 59
+ }
+ }
+ ]
+}
diff --git a/plans/bakeoff/round1.md b/plans/bakeoff/round1.md
@@ -0,0 +1,262 @@
+# Digest bake-off — round1
+
+Sample: 5 video(s) from `plans/bakeoff/sample.json` (buckets: short, medium), 2.72 audio-hours.
+Sweep days are projected as measured seconds-per-audio-hour x 77298 corpus audio-hours, one lane, no parallelism.
+
+| Candidate | Zero-yield chunks | Chapters/h | Rejection rate | Max gap | Generic | Dup | tok/s | s per audio-h | **Sweep days** |
+| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
+| `gemma2:9b@8192/absolute` | 0/10 (0%) | 15.78 | 0% | 00:14:59 | 2.3% | 0% | 55.2 | 322 | **288.1** |
+| `qwen3:8b@16384/absolute` | 1/7 (14.3%) | 15.42 | 23.6% | 00:39:57 | 11.9% | 0% | 76.8 | 229 | **204.9** |
+| `mistral-nemo:12b@16384/absolute` | 1/7 (14.3%) | 10.28 | 28.2% | 00:31:20 | 7.1% | 0% | 35 | 500 | **447.3** |
+| `qwen3:14b@8192/absolute` | 2/10 (20%) | 16.52 | 4.3% | 00:47:51 | 6.7% | 0% | 24.1 | 605 | **541.3** |
+| `qwen2.5:7b@16384/absolute` | 2/7 (28.6%) | 6.61 | 55% | 00:41:37 | 11.1% | 0% | 265.2 | 66 | **59** |
+
+## Rejections by guard
+
+| Candidate | chunk-failed | out-of-range | seam-duplicate |
+| --- | --- | --- | --- |
+| `gemma2:9b@8192/absolute` | 0 | 0 | 0 |
+| `qwen3:8b@16384/absolute` | 0 | 13 | 0 |
+| `mistral-nemo:12b@16384/absolute` | 0 | 11 | 0 |
+| `qwen3:14b@8192/absolute` | 2 | 0 | 1 |
+| `qwen2.5:7b@16384/absolute` | 0 | 22 | 0 |
+
+## Per-video
+
+| Candidate | Video | Bucket | Chunks | Zero-yield | Chapters | Max gap | Engine s |
+| --- | --- | --- | --- | --- | --- | --- | --- |
+| `gemma2:9b@8192/absolute` | `the-quartering-rumble/v6ve4n0` | short | 1 | 0 | 3 | 00:06:00 | 79 |
+| `gemma2:9b@8192/absolute` | `chibi-reviews/VQykVuHd9xQ` | short | 1 | 0 | 2 | 00:04:08 | 51 |
+| `gemma2:9b@8192/absolute` | `the-quartering/J0ySGwzP4Nw` | short | 1 | 0 | 3 | 00:05:31 | 105 |
+| `gemma2:9b@8192/absolute` | `destiny/5nmDzKB23OU` | medium | 4 | 0 | 20 | 00:13:24 | 409 |
+| `gemma2:9b@8192/absolute` | `angryjoeshow/RPJqkewZP5I` | medium | 3 | 0 | 15 | 00:14:59 | 233 |
+| `qwen3:8b@16384/absolute` | `the-quartering-rumble/v6ve4n0` | short | 1 | 0 | 6 | 00:04:06 | 63 |
+| `qwen3:8b@16384/absolute` | `chibi-reviews/VQykVuHd9xQ` | short | 1 | 0 | 4 | 00:03:50 | 35 |
+| `qwen3:8b@16384/absolute` | `the-quartering/J0ySGwzP4Nw` | short | 1 | 0 | 7 | 00:02:48 | 76 |
+| `qwen3:8b@16384/absolute` | `destiny/5nmDzKB23OU` | medium | 2 | 1 | 8 | 00:31:40 | 232 |
+| `qwen3:8b@16384/absolute` | `angryjoeshow/RPJqkewZP5I` | medium | 2 | 0 | 17 | 00:39:57 | 218 |
+| `mistral-nemo:12b@16384/absolute` | `the-quartering-rumble/v6ve4n0` | short | 1 | 0 | 3 | 00:04:57 | 107 |
+| `mistral-nemo:12b@16384/absolute` | `chibi-reviews/VQykVuHd9xQ` | short | 1 | 0 | 2 | 00:06:04 | 69 |
+| `mistral-nemo:12b@16384/absolute` | `the-quartering/J0ySGwzP4Nw` | short | 1 | 0 | 3 | 00:10:09 | 138 |
+| `mistral-nemo:12b@16384/absolute` | `destiny/5nmDzKB23OU` | medium | 2 | 1 | 8 | 00:31:20 | 550 |
+| `mistral-nemo:12b@16384/absolute` | `angryjoeshow/RPJqkewZP5I` | medium | 2 | 0 | 12 | 00:28:27 | 497 |
+| `qwen3:14b@8192/absolute` | `the-quartering-rumble/v6ve4n0` | short | 1 | 0 | 3 | 00:05:44 | 169 |
+| `qwen3:14b@8192/absolute` | `chibi-reviews/VQykVuHd9xQ` | short | 1 | 0 | 1 | 00:07:11 | 93 |
+| `qwen3:14b@8192/absolute` | `the-quartering/J0ySGwzP4Nw` | short | 1 | 0 | 3 | 00:05:01 | 208 |
+| `qwen3:14b@8192/absolute` | `destiny/5nmDzKB23OU` | medium | 4 | 0 | 30 | 00:13:36 | 977 |
+| `qwen3:14b@8192/absolute` | `angryjoeshow/RPJqkewZP5I` | medium | 3 | 2 | 8 | 00:47:51 | 201 |
+| `qwen2.5:7b@16384/absolute` | `the-quartering-rumble/v6ve4n0` | short | 1 | 0 | 4 | 00:04:57 | 7 |
+| `qwen2.5:7b@16384/absolute` | `chibi-reviews/VQykVuHd9xQ` | short | 1 | 0 | 2 | 00:03:44 | 4 |
+| `qwen2.5:7b@16384/absolute` | `the-quartering/J0ySGwzP4Nw` | short | 1 | 0 | 3 | 00:08:36 | 5 |
+| `qwen2.5:7b@16384/absolute` | `destiny/5nmDzKB23OU` | medium | 2 | 1 | 6 | 00:31:22 | 86 |
+| `qwen2.5:7b@16384/absolute` | `angryjoeshow/RPJqkewZP5I` | medium | 2 | 1 | 3 | 00:41:37 | 77 |
+
+## Sample output (first chapters per video)
+
+### `gemma2:9b@8192/absolute`
+
+**the-quartering-rumble/v6ve4n0** (short, 00:13:11)
+
+- 00:00:00 Trump Trolls Colbert While Setting Up a Trap for Democrats
+- 00:04:00 Colbert's Cancellation and the Changing Landscape of Late Night Television
+- 00:10:00 Trump's DC Move Forces Democrats to Address Crime
+
+**chibi-reviews/VQykVuHd9xQ** (short, 00:07:13)
+
+- 00:00:02 Crunchyroll's Translation Issues
+- 00:04:10 Impact on Viewership Experience
+
+**the-quartering/J0ySGwzP4Nw** (short, 00:13:05)
+
+- 00:00:04 Tim Pool Discusses Censorship on Social Media
+- 00:04:01 Twitter's Handling of the 'G Word'
+- 00:09:32 Concerns About LGBTQ+ Content in Schools
+
+**destiny/5nmDzKB23OU** (medium, 01:13:39)
+
+- 00:13:24 The Pure Concept of Religion
+- 00:16:53 Assumptions and Definitions
+- 00:20:02 Beyond Christianity
+- 00:20:41 Reconciling Perspectives
+- 00:20:56 The Importance of Openness
+- 00:32:58 Methodological Approach to Incorporating Evidence for Moral Facts
+- 00:34:02 Determining Which Brain States are 'Moral'
+- 00:37:01 Skepticism and the Burden of Proof
+
+**angryjoeshow/RPJqkewZP5I** (medium, 00:56:20)
+
+- 00:14:59 Game's Defense and Blame Shifting
+- 00:15:02 Content Creators and Gamers Blamed
+- 00:15:38 One Peg Response - Game's Issues Highlighted
+- 00:16:21 Mixed Opinions on the Game
+- 00:17:34 Studio's Responsibility and No Man's Sky Comparison
+- 00:18:25 Microtransactions and Lack of Appeal
+- 00:20:39 Jeff Keley's Role and Personal Taste
+- 00:34:25 Twitter and Patch Notes Struggles
+
+### `qwen3:8b@16384/absolute`
+
+**the-quartering-rumble/v6ve4n0** (short, 00:13:11)
+
+- 00:00:00 Trump Trolls Colbert by Taking His Prime Time Job
+- 00:02:03 Speculation About Secret Side Deals and Late Night Television
+- 00:06:09 Paramount's Merger and the Late Night Show Format Crisis
+- 00:08:31 Democrats Grapple with Crime and Trump's Strategy
+- 00:10:06 Trump's Impact on Crime Perception and Democratic Vulnerabilities
+- 00:12:01 Woke DAs and the Changing Crime Statistics
+
+**chibi-reviews/VQykVuHd9xQ** (short, 00:07:13)
+
+- 00:00:02 Crunchyroll's Translation Issues
+- 00:01:00 Translation Errors in Gundam Episode
+- 00:02:33 Unwatchable Episode Due to Subtitle Delays
+- 00:03:23 Translation Problems in Overlord Series},{
+
+**the-quartering/J0ySGwzP4Nw** (short, 00:13:05)
+
+- 00:00:04 Tim Pool's War With Twitter
+- 00:01:13 The G Word and Banning
+- 00:04:01 Free Speech and Platform Rules
+- 00:06:00 Twitter's Alleged Protection of Creeps
+- 00:08:01 Twitter's Uneven Enforcement
+- 00:10:00 Concerns About Teachers and Young Ones
+- 00:12:05 Criticizing the Platform's Policies
+
+**destiny/5nmDzKB23OU** (medium, 01:13:39)
+
+- 00:02:00 Exploring Moral Systems and Preferences
+- 00:06:00 Meta Ethics and the Nature of Moral Truth
+- 00:12:00 The Role of Neuroscience in Understanding Morality
+- 00:18:00 Skepticism and the Burden of Proof
+- 00:24:00 Normativity and Its Interpretations
+- 00:30:00 Conclusion and Reflection on Moral Frameworks
+- 00:36:00 Final Thoughts and Open Questions
+- 00:42:00 Closing Remarks and Further Exploration
+
+**angryjoeshow/RPJqkewZP5I** (medium, 00:56:20)
+
+- 00:00:28 Silent Hill and New Game
+- 00:00:30 John Wick and God of War
+- 00:00:40 Marathon and High Guard
+- 00:00:50 Crimson Desert and AI
+- 00:01:00 State of Play and AI Controversy
+- 00:01:10 Conclusion and Final Thoughts
+- 00:01:20 Wrap-Up and Farewell
+- 00:01:30 End of Episode
+
+### `mistral-nemo:12b@16384/absolute`
+
+**the-quartering-rumble/v6ve4n0** (short, 00:13:11)
+
+- 00:00:00 Trump's Ruthless Troll Against Colbert
+- 00:04:09 Colbert's Late Show Cancellation and Trump's Involvement
+- 00:08:14 Economic Challenges of Late Night Shows
+
+**chibi-reviews/VQykVuHd9xQ** (short, 00:07:13)
+
+- 00:00:02 Crunchyroll's translations issues
+- 00:01:09 Gundam episode with subtitle delays
+
+**the-quartering/J0ySGwzP4Nw** (short, 00:13:05)
+
+- 00:00:04 Word meanings and political aisles
+- 00:01:13 Tim Pool's Twitter ban for using the 'G word'
+- 00:02:56 Twitter's advertising restrictions on Tim Pool
+
+**destiny/5nmDzKB23OU** (medium, 01:13:39)
+
+- 00:26:10 Segment Transcript
+- 00:27:58 Example of Moral System
+- 00:34:30 General Claim about Psychology
+- 00:36:17 Claim about Matrix Facts
+- 00:40:22 Default Position on Normativity
+- 00:41:04 Definition of Normativity
+- 00:41:36 Christine Korsgaard's Theory of Normativity
+- 00:42:19 Step Made by Christine Korsgaard
+
+**angryjoeshow/RPJqkewZP5I** (medium, 00:56:20)
+
+- 00:28:27 Silent Hill
+- 00:29:50 High Guard
+- 00:34:41 Crimson Desert
+- 00:36:02 Chicken Man
+- 00:37:19 God of War Remake
+- 00:38:40 AI Deepfakes
+- 00:41:37 Copyright Infringement in AI-Generated Content
+- 00:42:23 Cance 2.0
+
+### `qwen3:14b@8192/absolute`
+
+**the-quartering-rumble/v6ve4n0** (short, 00:13:11)
+
+- 00:00:00 Trump's Troll on Colbert and the Kennedy Center Honors
+- 00:03:19 Late Night Television's Decline and the Impact of Trump Derangement Syndrome
+- 00:09:03 Democrats' Struggle with Crime and Trump's Political Strategy
+
+**chibi-reviews/VQykVuHd9xQ** (short, 00:07:13)
+
+- 00:00:02 Translation Errors on Crunchyroll
+
+**the-quartering/J0ySGwzP4Nw** (short, 00:13:05)
+
+- 00:00:04 Introduction to the Controversy Over Word Usage
+- 00:03:06 Tim Pool's Twitter Ban and His Response
+- 00:08:07 Discussion on Twitter's Enforcement and LGBTQ+ Terminology
+
+**destiny/5nmDzKB23OU** (medium, 01:13:39)
+
+- 00:13:15 Different Reasoning Processes and the Concept of a Higher Power
+- 00:13:45 Faith vs. Science in Practical Applications
+- 00:15:05 The Role of Faith in Metaphysical vs. Material Claims
+- 00:16:04 Philosophical vs. Societal Definitions of Religion
+- 00:17:56 Assumptions About Religion and Their Implications
+- 00:19:17 The Frustration with Narrow Interpretations of Religion
+- 00:20:32 Conceding the Common Vernacular of Religion
+- 00:32:50 The Challenge of Moral Facts and Neuroscience
+
+**angryjoeshow/RPJqkewZP5I** (medium, 00:56:20)
+
+- 00:47:51 AI in Healthcare and Accuracy Concerns
+- 00:48:10 AI's Role in Managing Health Records
+- 00:49:00 Escaping AI Algorithms and Content
+- 00:50:00 AI Memes and Cultural Impact
+- 00:51:00 AI and the Future of Entertainment
+- 00:52:00 AI's Environmental Impact
+- 00:54:00 Water Usage and AI
+- 00:55:00 Conclusion and Advertisements
+
+### `qwen2.5:7b@16384/absolute`
+
+**the-quartering-rumble/v6ve4n0** (short, 00:13:11)
+
+- 00:00:00 Introduction and Discord Promotion
+- 00:01:49 Trump's Trolling of Stephen Colbert
+- 00:03:56 Announcement of Trump Hosting Kennedy Center Honors
+- 00:08:14 Late Night Television Challenges and Viewership Shifts
+
+**chibi-reviews/VQykVuHd9xQ** (short, 00:07:13)
+
+- 00:00:05 Crunchyroll's Translation Issues
+- 00:03:49 Multiple Shows with Translation Errors
+
+**the-quartering/J0ySGwzP4Nw** (short, 00:13:05)
+
+- 00:00:48 Discussion on the redefinition of words
+- 00:01:35 Tim Pool's ban from Twitter and his response
+- 00:04:29 Journalist Tim Pool's criticism of Twitter's policies
+
+**destiny/5nmDzKB23OU** (medium, 01:13:39)
+
+- 00:26:11 The easiest way to prove that ever exists
+- 00:27:05 Does conceding the possibility of metaphysical truth change anything?
+- 00:30:02 Is it possible to find moral facts by analyzing brain states?
+- 00:40:00 Arguments against the existence of moral facts
+- 00:41:03 Definition and explanation of normativity
+- 00:42:17 Normativity in practical problems
+
+**angryjoeshow/RPJqkewZP5I** (medium, 00:56:20)
+
+- 00:41:37 Copyright Issues with Unauthorized AI Creations
+- 00:45:24 Sony Considering Delaying PlayStation Console Launch Due to AI Demand
+- 00:48:10 AI Accuracy in Healthcare Records Management
+
diff --git a/plans/bakeoff/round2.json b/plans/bakeoff/round2.json
@@ -0,0 +1,562 @@
+{
+ "label": "round2",
+ "sample": "/home/user/Projects/yt-dlp-transcript-browser/plans/bakeoff/sample.json",
+ "corpus": {
+ "videosScanned": 76354,
+ "videosWithTranscript": 73367,
+ "audioHours": 77298,
+ "longTailVideos": 6038,
+ "longTailAudioHours": 40153
+ },
+ "buckets": [
+ "long"
+ ],
+ "scores": [
+ {
+ "candidate": {
+ "key": "gemma2:9b@8192/chunk-local",
+ "model": "gemma2:9b",
+ "numCtx": 8192,
+ "maxCues": 600,
+ "timestampMode": "chunk-local"
+ },
+ "videos": [
+ {
+ "slug": "rekietalaw/EsZhaCfc8HQ",
+ "bucket": "long",
+ "durationSeconds": 12497,
+ "chunks": 8,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 44,
+ "maxGapSeconds": 1228,
+ "engineSeconds": 502.5470000000001,
+ "inputTokens": 30597,
+ "outputTokens": 1816,
+ "warningsByCode": {
+ "out-of-range": 4,
+ "non-monotonic": 3
+ },
+ "genericTitles": 7,
+ "duplicateTitles": 1,
+ "sampleTitles": [
+ "00:19:17 Introduction and Bronca's Take",
+ "00:20:36 Discussing the Shooting Incident",
+ "00:24:18 Andrew Wilson and His Comparison to a Serial Killer",
+ "00:24:54 Analyzing the Lawyer's Argument",
+ "00:26:32 The Fourth Amendment and Reasonable Belief",
+ "00:27:37 Minnesota Law and Police Shootings",
+ "00:42:50 ICE Jurisdiction and State Laws",
+ "00:48:03 The Incident: Vehicle Approach and Officer Orders"
+ ]
+ },
+ {
+ "slug": "chrissie-mayr/cfLF2o2-0BA",
+ "bucket": "long",
+ "durationSeconds": 12553,
+ "chunks": 9,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 46,
+ "maxGapSeconds": 1444,
+ "engineSeconds": 527.155,
+ "inputTokens": 37239,
+ "outputTokens": 1958,
+ "warningsByCode": {
+ "non-monotonic": 7,
+ "out-of-range": 1
+ },
+ "genericTitles": 2,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:20:33 Initial Reaction and Expectations",
+ "00:20:57 Humor and Critique of Masculinity",
+ "00:22:04 Addressing the 'Woke' Label",
+ "00:24:51 Comparison to Past Chick Flicks",
+ "00:26:37 The Impact of Online Movie Reviews",
+ "00:27:33 Ryan Gosling's Performance and Humor",
+ "00:44:41 Barbie's Throwaway Boyfriend and Teresa's Absence",
+ "00:51:40 The Appeal of Barbie History and Gender Differences"
+ ]
+ }
+ ],
+ "totals": {
+ "videos": 2,
+ "audioHours": 6.96,
+ "chunks": 17,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "zeroYieldRate": 0,
+ "kept": 90,
+ "chaptersPerHour": 12.93,
+ "maxGapSeconds": 1444,
+ "meanGapSeconds": 1336,
+ "genericTitleRate": 0.1,
+ "duplicateTitleRate": 0.0111,
+ "warningsByCode": {
+ "out-of-range": 5,
+ "non-monotonic": 10
+ },
+ "rejectionRate": 0.1429,
+ "engineSeconds": 1030,
+ "tokensPerSecond": 69.5,
+ "secondsPerAudioHour": 148,
+ "projectedSweepDays": 132.4
+ }
+ },
+ {
+ "candidate": {
+ "key": "qwen2.5:7b@16384/chunk-local",
+ "model": "qwen2.5:7b",
+ "numCtx": 16384,
+ "maxCues": 1200,
+ "timestampMode": "chunk-local"
+ },
+ "videos": [
+ {
+ "slug": "rekietalaw/EsZhaCfc8HQ",
+ "bucket": "long",
+ "durationSeconds": 12497,
+ "chunks": 4,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 36,
+ "maxGapSeconds": 2981,
+ "engineSeconds": 95.23400000000001,
+ "inputTokens": 34715,
+ "outputTokens": 1360,
+ "warningsByCode": {
+ "out-of-range": 17,
+ "seam-duplicate": 1
+ },
+ "genericTitles": 4,
+ "duplicateTitles": 1,
+ "sampleTitles": [
+ "00:00:19 Introduction and Context",
+ "00:01:42 The Incident and Initial Analysis",
+ "00:05:15 Legal Justification for Use of Force",
+ "00:10:46 Use of Deadly Force and Self-Defense",
+ "00:14:18 Tactical Considerations and Officer Safety",
+ "00:19:05 The Role of Bystanders and Video Evidence",
+ "00:23:16 Public Perception and Media Coverage",
+ "00:30:48 Legal Analysis and Case Law"
+ ]
+ },
+ {
+ "slug": "chrissie-mayr/cfLF2o2-0BA",
+ "bucket": "long",
+ "durationSeconds": 12553,
+ "chunks": 5,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 1,
+ "kept": 29,
+ "maxGapSeconds": 5268,
+ "engineSeconds": 100.58200000000001,
+ "inputTokens": 34157,
+ "outputTokens": 1592,
+ "warningsByCode": {
+ "out-of-range": 29
+ },
+ "genericTitles": 9,
+ "duplicateTitles": 1,
+ "sampleTitles": [
+ "01:27:49 Opinions on the Barbie movie",
+ "01:31:24 Trust in opinions based on experience",
+ "01:33:28 Comparison to other movies and remakes",
+ "01:35:38 Amy Schumer's involvement with the Barbie project",
+ "01:40:01 Discussion on feminism in the Barbie movie",
+ "01:43:27 Dreams of childhood and their impact",
+ "01:45:52 Barbie movies for a different age group",
+ "01:46:35 Introduction and Opening Remarks"
+ ]
+ }
+ ],
+ "totals": {
+ "videos": 2,
+ "audioHours": 6.96,
+ "chunks": 9,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 1,
+ "zeroYieldRate": 0.1111,
+ "kept": 65,
+ "chaptersPerHour": 9.34,
+ "maxGapSeconds": 5268,
+ "meanGapSeconds": 4125,
+ "genericTitleRate": 0.2,
+ "duplicateTitleRate": 0.0308,
+ "warningsByCode": {
+ "out-of-range": 46,
+ "seam-duplicate": 1
+ },
+ "rejectionRate": 0.4144,
+ "engineSeconds": 196,
+ "tokensPerSecond": 366.8,
+ "secondsPerAudioHour": 28,
+ "projectedSweepDays": 25.1
+ }
+ },
+ {
+ "candidate": {
+ "key": "qwen2.5:7b@8192/chunk-local",
+ "model": "qwen2.5:7b",
+ "numCtx": 8192,
+ "maxCues": 600,
+ "timestampMode": "chunk-local"
+ },
+ "videos": [
+ {
+ "slug": "rekietalaw/EsZhaCfc8HQ",
+ "bucket": "long",
+ "durationSeconds": 12497,
+ "chunks": 8,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 0,
+ "kept": 47,
+ "maxGapSeconds": 1382,
+ "engineSeconds": 85.91499999999999,
+ "inputTokens": 30565,
+ "outputTokens": 1441,
+ "warningsByCode": {
+ "out-of-range": 4,
+ "non-monotonic": 2
+ },
+ "genericTitles": 16,
+ "duplicateTitles": 3,
+ "sampleTitles": [
+ "00:00:19 Introduction and Context",
+ "00:02:15 Overview of the Incident",
+ "00:04:42 Legal Analysis of Use of Force",
+ "00:07:57 Discussion on Objective Reasonable Fear",
+ "00:11:15 Subjective Fear and Its Role",
+ "00:14:48 Minnesota's Specific Law on Use of Force",
+ "00:17:57 Impact of the Lawsuit on Objective Reasonable Fear Standard",
+ "00:21:15 Conclusion and Final Thoughts"
+ ]
+ },
+ {
+ "slug": "chrissie-mayr/cfLF2o2-0BA",
+ "bucket": "long",
+ "durationSeconds": 12553,
+ "chunks": 9,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 2,
+ "kept": 46,
+ "maxGapSeconds": 1453,
+ "engineSeconds": 104.55800000000002,
+ "inputTokens": 37164,
+ "outputTokens": 1785,
+ "warningsByCode": {
+ "out-of-range": 9,
+ "language-drift": 7
+ },
+ "genericTitles": 13,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:20:21 Personal enjoyment and expectations",
+ "00:20:57 Anti-male themes in the movie",
+ "00:22:04 Balancing empowerment with criticism",
+ "00:24:17 Target audience and gender dynamics",
+ "00:25:49 Impact of political correctness on entertainment",
+ "00:27:33 Ryan Gosling's humor in the film",
+ "00:51:45 Introduction and Personal Background",
+ "00:58:20 Discussion on the impact of moving to Alabama"
+ ]
+ }
+ ],
+ "totals": {
+ "videos": 2,
+ "audioHours": 6.96,
+ "chunks": 17,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 2,
+ "zeroYieldRate": 0.1176,
+ "kept": 93,
+ "chaptersPerHour": 13.37,
+ "maxGapSeconds": 1453,
+ "meanGapSeconds": 1418,
+ "genericTitleRate": 0.3118,
+ "duplicateTitleRate": 0.0323,
+ "warningsByCode": {
+ "out-of-range": 13,
+ "non-monotonic": 2,
+ "language-drift": 7
+ },
+ "rejectionRate": 0.1913,
+ "engineSeconds": 190,
+ "tokensPerSecond": 372.5,
+ "secondsPerAudioHour": 27,
+ "projectedSweepDays": 24.2
+ }
+ },
+ {
+ "candidate": {
+ "key": "gemma2:9b@8192/absolute",
+ "model": "gemma2:9b",
+ "numCtx": 8192,
+ "maxCues": 600,
+ "timestampMode": "absolute"
+ },
+ "videos": [
+ {
+ "slug": "rekietalaw/EsZhaCfc8HQ",
+ "bucket": "long",
+ "durationSeconds": 12497,
+ "chunks": 8,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 1,
+ "kept": 37,
+ "maxGapSeconds": 1634,
+ "engineSeconds": 456.12999999999994,
+ "inputTokens": 30597,
+ "outputTokens": 1810,
+ "warningsByCode": {
+ "non-monotonic": 7,
+ "out-of-range": 7
+ },
+ "genericTitles": 3,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:19:17 Introduction and Bronca's Take",
+ "00:20:36 Discussing the Shooting Video",
+ "00:24:18 Andrew Wilson and Scarlett Hampton",
+ "00:25:00 Legal Analysis of the Shooting",
+ "00:27:32 Objective vs. Subjective Fear",
+ "00:27:58 Continuing the Legal Analysis",
+ "00:42:50 ICE Jurisdiction and State Laws",
+ "00:43:12 State Charges Against Federal Agents"
+ ]
+ },
+ {
+ "slug": "chrissie-mayr/cfLF2o2-0BA",
+ "bucket": "long",
+ "durationSeconds": 12553,
+ "chunks": 9,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 2,
+ "kept": 40,
+ "maxGapSeconds": 3847,
+ "engineSeconds": 535.993,
+ "inputTokens": 37239,
+ "outputTokens": 2021,
+ "warningsByCode": {
+ "non-monotonic": 1,
+ "out-of-range": 15
+ },
+ "genericTitles": 1,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:20:23 Barbie Movie Review - A Divided Opinion",
+ "00:21:48 The Pressure of Manhood and Womanhood",
+ "00:23:57 Barbie vs. Fight Club - A Comparison",
+ "00:26:19 The Political Brain Rot of Entertainment",
+ "00:27:48 Ryan Gosling's Ken - A Hilarious Performance",
+ "00:28:19 The Making of a Ken Doll - Ryan Gosling's Daughter",
+ "00:44:41 Barbie's Throwaway Boyfriend and Teresa's Absence",
+ "00:45:21 Greta Gerwig and the Brunettes of the World"
+ ]
+ }
+ ],
+ "totals": {
+ "videos": 2,
+ "audioHours": 6.96,
+ "chunks": 17,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 3,
+ "zeroYieldRate": 0.1765,
+ "kept": 77,
+ "chaptersPerHour": 11.07,
+ "maxGapSeconds": 3847,
+ "meanGapSeconds": 2741,
+ "genericTitleRate": 0.0519,
+ "duplicateTitleRate": 0,
+ "warningsByCode": {
+ "non-monotonic": 8,
+ "out-of-range": 22
+ },
+ "rejectionRate": 0.2804,
+ "engineSeconds": 992,
+ "tokensPerSecond": 72.2,
+ "secondsPerAudioHour": 143,
+ "projectedSweepDays": 127.9
+ }
+ },
+ {
+ "candidate": {
+ "key": "qwen2.5:7b@8192/absolute",
+ "model": "qwen2.5:7b",
+ "numCtx": 8192,
+ "maxCues": 600,
+ "timestampMode": "absolute"
+ },
+ "videos": [
+ {
+ "slug": "rekietalaw/EsZhaCfc8HQ",
+ "bucket": "long",
+ "durationSeconds": 12497,
+ "chunks": 8,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 2,
+ "kept": 32,
+ "maxGapSeconds": 3291,
+ "engineSeconds": 104.74499999999999,
+ "inputTokens": 30565,
+ "outputTokens": 1955,
+ "warningsByCode": {
+ "out-of-range": 37
+ },
+ "genericTitles": 8,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:01:36 Overview of the Incident",
+ "00:04:25 Use of Force Analysis",
+ "00:07:38 Legal Framework and Reasonable Fear",
+ "00:11:29 Subjective vs. Objective Reasonable Fear",
+ "00:14:56 Minnesota's Use of Force Law",
+ "00:18:37 Discussion with Bronca",
+ "00:22:41 Critique of Andrew Wilson's Analysis",
+ "00:25:34 Conclusion and Legal Implications"
+ ]
+ },
+ {
+ "slug": "chrissie-mayr/cfLF2o2-0BA",
+ "bucket": "long",
+ "durationSeconds": 12553,
+ "chunks": 9,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 3,
+ "kept": 40,
+ "maxGapSeconds": 3916,
+ "engineSeconds": 108.733,
+ "inputTokens": 37164,
+ "outputTokens": 1969,
+ "warningsByCode": {
+ "out-of-range": 28
+ },
+ "genericTitles": 5,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:19:56 Review of the movie",
+ "00:20:07 Enjoyment and target audience",
+ "00:20:34 Criticism of the review culture",
+ "00:21:55 Discussion on gender roles and pressures",
+ "00:24:07 Comparison with other movies",
+ "00:26:58 Sensationalism in movie reviews",
+ "01:08:05 Personal Background and Travel Plans",
+ "01:08:34 East Coast vs. Southern East Coast"
+ ]
+ }
+ ],
+ "totals": {
+ "videos": 2,
+ "audioHours": 6.96,
+ "chunks": 17,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 5,
+ "zeroYieldRate": 0.2941,
+ "kept": 72,
+ "chaptersPerHour": 10.35,
+ "maxGapSeconds": 3916,
+ "meanGapSeconds": 3604,
+ "genericTitleRate": 0.1806,
+ "duplicateTitleRate": 0,
+ "warningsByCode": {
+ "out-of-range": 65
+ },
+ "rejectionRate": 0.4745,
+ "engineSeconds": 213,
+ "tokensPerSecond": 335.6,
+ "secondsPerAudioHour": 31,
+ "projectedSweepDays": 27.7
+ }
+ },
+ {
+ "candidate": {
+ "key": "qwen2.5:7b@16384/absolute",
+ "model": "qwen2.5:7b",
+ "numCtx": 16384,
+ "maxCues": 1200,
+ "timestampMode": "absolute"
+ },
+ "videos": [
+ {
+ "slug": "rekietalaw/EsZhaCfc8HQ",
+ "bucket": "long",
+ "durationSeconds": 12497,
+ "chunks": 4,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 1,
+ "kept": 22,
+ "maxGapSeconds": 3527,
+ "engineSeconds": 102.381,
+ "inputTokens": 34715,
+ "outputTokens": 1322,
+ "warningsByCode": {
+ "out-of-range": 31
+ },
+ "genericTitles": 3,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "00:01:23 The Incident and Initial Analysis",
+ "00:05:46 Legal Justification for Use of Force",
+ "00:17:28 Analysis of the Video Evidence",
+ "00:35:15 Police Authority and Self-Defense",
+ "00:44:12 Discussion on Legal Standards and Analysis",
+ "00:48:05 Video Evidence and Analysis",
+ "00:53:52 Imminence of Threat",
+ "00:56:02 Conclusion and Final Thoughts"
+ ]
+ },
+ {
+ "slug": "chrissie-mayr/cfLF2o2-0BA",
+ "bucket": "long",
+ "durationSeconds": 12553,
+ "chunks": 5,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 2,
+ "kept": 35,
+ "maxGapSeconds": 5268,
+ "engineSeconds": 104.44,
+ "inputTokens": 34157,
+ "outputTokens": 1783,
+ "warningsByCode": {
+ "out-of-range": 27
+ },
+ "genericTitles": 7,
+ "duplicateTitles": 0,
+ "sampleTitles": [
+ "01:27:49 Disagreement about the impact of your opinion on others",
+ "01:28:06 Discussion on a potential conservative Barbie movie",
+ "01:28:54 Criticism of the original Barbie movie's content and tone",
+ "01:30:09 Comparison between different opinions on the same topic",
+ "01:31:00 Defense of a male reviewer's opinion on Barbie",
+ "01:32:05 Criticism of Ben Shapiro's review as overly negative and unbalanced",
+ "01:34:06 Discussion about potential Ken movies with Ryan Gosling",
+ "01:37:03 Exploration of different interpretations of feminism"
+ ]
+ }
+ ],
+ "totals": {
+ "videos": 2,
+ "audioHours": 6.96,
+ "chunks": 9,
+ "chunksFailed": 0,
+ "zeroYieldChunks": 3,
+ "zeroYieldRate": 0.3333,
+ "kept": 57,
+ "chaptersPerHour": 8.19,
+ "maxGapSeconds": 5268,
+ "meanGapSeconds": 4398,
+ "genericTitleRate": 0.1754,
+ "duplicateTitleRate": 0,
+ "warningsByCode": {
+ "out-of-range": 58
+ },
+ "rejectionRate": 0.5043,
+ "engineSeconds": 207,
+ "tokensPerSecond": 348,
+ "secondsPerAudioHour": 30,
+ "projectedSweepDays": 26.8
+ }
+ }
+ ]
+}
diff --git a/plans/bakeoff/round2.md b/plans/bakeoff/round2.md
@@ -0,0 +1,188 @@
+# Digest bake-off — round2
+
+Sample: 2 video(s) from `plans/bakeoff/sample.json` (buckets: long), 6.96 audio-hours.
+Sweep days are projected as measured seconds-per-audio-hour x 77298 corpus audio-hours, one lane, no parallelism.
+
+| Candidate | Zero-yield chunks | Chapters/h | Rejection rate | Max gap | Generic | Dup | tok/s | s per audio-h | **Sweep days** |
+| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
+| `gemma2:9b@8192/chunk-local` | 0/17 (0%) | 12.93 | 14.3% | 00:24:04 | 10% | 1.1% | 69.5 | 148 | **132.4** |
+| `qwen2.5:7b@16384/chunk-local` | 1/9 (11.1%) | 9.34 | 41.4% | 01:27:48 | 20% | 3.1% | 366.8 | 28 | **25.1** |
+| `qwen2.5:7b@8192/chunk-local` | 2/17 (11.8%) | 13.37 | 19.1% | 00:24:13 | 31.2% | 3.2% | 372.5 | 27 | **24.2** |
+| `gemma2:9b@8192/absolute` | 3/17 (17.7%) | 11.07 | 28% | 01:04:07 | 5.2% | 0% | 72.2 | 143 | **127.9** |
+| `qwen2.5:7b@8192/absolute` | 5/17 (29.4%) | 10.35 | 47.5% | 01:05:16 | 18.1% | 0% | 335.6 | 31 | **27.7** |
+| `qwen2.5:7b@16384/absolute` | 3/9 (33.3%) | 8.19 | 50.4% | 01:27:48 | 17.5% | 0% | 348 | 30 | **26.8** |
+
+## Rejections by guard
+
+| Candidate | language-drift | non-monotonic | out-of-range | seam-duplicate |
+| --- | --- | --- | --- | --- |
+| `gemma2:9b@8192/chunk-local` | 0 | 10 | 5 | 0 |
+| `qwen2.5:7b@16384/chunk-local` | 0 | 0 | 46 | 1 |
+| `qwen2.5:7b@8192/chunk-local` | 7 | 2 | 13 | 0 |
+| `gemma2:9b@8192/absolute` | 0 | 8 | 22 | 0 |
+| `qwen2.5:7b@8192/absolute` | 0 | 0 | 65 | 0 |
+| `qwen2.5:7b@16384/absolute` | 0 | 0 | 58 | 0 |
+
+## Per-video
+
+| Candidate | Video | Bucket | Chunks | Zero-yield | Chapters | Max gap | Engine s |
+| --- | --- | --- | --- | --- | --- | --- | --- |
+| `gemma2:9b@8192/chunk-local` | `rekietalaw/EsZhaCfc8HQ` | long | 8 | 0 | 44 | 00:20:28 | 503 |
+| `gemma2:9b@8192/chunk-local` | `chrissie-mayr/cfLF2o2-0BA` | long | 9 | 0 | 46 | 00:24:04 | 527 |
+| `qwen2.5:7b@16384/chunk-local` | `rekietalaw/EsZhaCfc8HQ` | long | 4 | 0 | 36 | 00:49:41 | 95 |
+| `qwen2.5:7b@16384/chunk-local` | `chrissie-mayr/cfLF2o2-0BA` | long | 5 | 1 | 29 | 01:27:48 | 101 |
+| `qwen2.5:7b@8192/chunk-local` | `rekietalaw/EsZhaCfc8HQ` | long | 8 | 0 | 47 | 00:23:02 | 86 |
+| `qwen2.5:7b@8192/chunk-local` | `chrissie-mayr/cfLF2o2-0BA` | long | 9 | 2 | 46 | 00:24:13 | 105 |
+| `gemma2:9b@8192/absolute` | `rekietalaw/EsZhaCfc8HQ` | long | 8 | 1 | 37 | 00:27:14 | 456 |
+| `gemma2:9b@8192/absolute` | `chrissie-mayr/cfLF2o2-0BA` | long | 9 | 2 | 40 | 01:04:07 | 536 |
+| `qwen2.5:7b@8192/absolute` | `rekietalaw/EsZhaCfc8HQ` | long | 8 | 2 | 32 | 00:54:51 | 105 |
+| `qwen2.5:7b@8192/absolute` | `chrissie-mayr/cfLF2o2-0BA` | long | 9 | 3 | 40 | 01:05:16 | 109 |
+| `qwen2.5:7b@16384/absolute` | `rekietalaw/EsZhaCfc8HQ` | long | 4 | 1 | 22 | 00:58:47 | 102 |
+| `qwen2.5:7b@16384/absolute` | `chrissie-mayr/cfLF2o2-0BA` | long | 5 | 2 | 35 | 01:27:48 | 104 |
+
+## Sample output (first chapters per video)
+
+### `gemma2:9b@8192/chunk-local`
+
+**rekietalaw/EsZhaCfc8HQ** (long, 03:28:17)
+
+- 00:19:17 Introduction and Bronca's Take
+- 00:20:36 Discussing the Shooting Incident
+- 00:24:18 Andrew Wilson and His Comparison to a Serial Killer
+- 00:24:54 Analyzing the Lawyer's Argument
+- 00:26:32 The Fourth Amendment and Reasonable Belief
+- 00:27:37 Minnesota Law and Police Shootings
+- 00:42:50 ICE Jurisdiction and State Laws
+- 00:48:03 The Incident: Vehicle Approach and Officer Orders
+
+**chrissie-mayr/cfLF2o2-0BA** (long, 03:29:13)
+
+- 00:20:33 Initial Reaction and Expectations
+- 00:20:57 Humor and Critique of Masculinity
+- 00:22:04 Addressing the 'Woke' Label
+- 00:24:51 Comparison to Past Chick Flicks
+- 00:26:37 The Impact of Online Movie Reviews
+- 00:27:33 Ryan Gosling's Performance and Humor
+- 00:44:41 Barbie's Throwaway Boyfriend and Teresa's Absence
+- 00:51:40 The Appeal of Barbie History and Gender Differences
+
+### `qwen2.5:7b@16384/chunk-local`
+
+**rekietalaw/EsZhaCfc8HQ** (long, 03:28:17)
+
+- 00:00:19 Introduction and Context
+- 00:01:42 The Incident and Initial Analysis
+- 00:05:15 Legal Justification for Use of Force
+- 00:10:46 Use of Deadly Force and Self-Defense
+- 00:14:18 Tactical Considerations and Officer Safety
+- 00:19:05 The Role of Bystanders and Video Evidence
+- 00:23:16 Public Perception and Media Coverage
+- 00:30:48 Legal Analysis and Case Law
+
+**chrissie-mayr/cfLF2o2-0BA** (long, 03:29:13)
+
+- 01:27:49 Opinions on the Barbie movie
+- 01:31:24 Trust in opinions based on experience
+- 01:33:28 Comparison to other movies and remakes
+- 01:35:38 Amy Schumer's involvement with the Barbie project
+- 01:40:01 Discussion on feminism in the Barbie movie
+- 01:43:27 Dreams of childhood and their impact
+- 01:45:52 Barbie movies for a different age group
+- 01:46:35 Introduction and Opening Remarks
+
+### `qwen2.5:7b@8192/chunk-local`
+
+**rekietalaw/EsZhaCfc8HQ** (long, 03:28:17)
+
+- 00:00:19 Introduction and Context
+- 00:02:15 Overview of the Incident
+- 00:04:42 Legal Analysis of Use of Force
+- 00:07:57 Discussion on Objective Reasonable Fear
+- 00:11:15 Subjective Fear and Its Role
+- 00:14:48 Minnesota's Specific Law on Use of Force
+- 00:17:57 Impact of the Lawsuit on Objective Reasonable Fear Standard
+- 00:21:15 Conclusion and Final Thoughts
+
+**chrissie-mayr/cfLF2o2-0BA** (long, 03:29:13)
+
+- 00:20:21 Personal enjoyment and expectations
+- 00:20:57 Anti-male themes in the movie
+- 00:22:04 Balancing empowerment with criticism
+- 00:24:17 Target audience and gender dynamics
+- 00:25:49 Impact of political correctness on entertainment
+- 00:27:33 Ryan Gosling's humor in the film
+- 00:51:45 Introduction and Personal Background
+- 00:58:20 Discussion on the impact of moving to Alabama
+
+### `gemma2:9b@8192/absolute`
+
+**rekietalaw/EsZhaCfc8HQ** (long, 03:28:17)
+
+- 00:19:17 Introduction and Bronca's Take
+- 00:20:36 Discussing the Shooting Video
+- 00:24:18 Andrew Wilson and Scarlett Hampton
+- 00:25:00 Legal Analysis of the Shooting
+- 00:27:32 Objective vs. Subjective Fear
+- 00:27:58 Continuing the Legal Analysis
+- 00:42:50 ICE Jurisdiction and State Laws
+- 00:43:12 State Charges Against Federal Agents
+
+**chrissie-mayr/cfLF2o2-0BA** (long, 03:29:13)
+
+- 00:20:23 Barbie Movie Review - A Divided Opinion
+- 00:21:48 The Pressure of Manhood and Womanhood
+- 00:23:57 Barbie vs. Fight Club - A Comparison
+- 00:26:19 The Political Brain Rot of Entertainment
+- 00:27:48 Ryan Gosling's Ken - A Hilarious Performance
+- 00:28:19 The Making of a Ken Doll - Ryan Gosling's Daughter
+- 00:44:41 Barbie's Throwaway Boyfriend and Teresa's Absence
+- 00:45:21 Greta Gerwig and the Brunettes of the World
+
+### `qwen2.5:7b@8192/absolute`
+
+**rekietalaw/EsZhaCfc8HQ** (long, 03:28:17)
+
+- 00:01:36 Overview of the Incident
+- 00:04:25 Use of Force Analysis
+- 00:07:38 Legal Framework and Reasonable Fear
+- 00:11:29 Subjective vs. Objective Reasonable Fear
+- 00:14:56 Minnesota's Use of Force Law
+- 00:18:37 Discussion with Bronca
+- 00:22:41 Critique of Andrew Wilson's Analysis
+- 00:25:34 Conclusion and Legal Implications
+
+**chrissie-mayr/cfLF2o2-0BA** (long, 03:29:13)
+
+- 00:19:56 Review of the movie
+- 00:20:07 Enjoyment and target audience
+- 00:20:34 Criticism of the review culture
+- 00:21:55 Discussion on gender roles and pressures
+- 00:24:07 Comparison with other movies
+- 00:26:58 Sensationalism in movie reviews
+- 01:08:05 Personal Background and Travel Plans
+- 01:08:34 East Coast vs. Southern East Coast
+
+### `qwen2.5:7b@16384/absolute`
+
+**rekietalaw/EsZhaCfc8HQ** (long, 03:28:17)
+
+- 00:01:23 The Incident and Initial Analysis
+- 00:05:46 Legal Justification for Use of Force
+- 00:17:28 Analysis of the Video Evidence
+- 00:35:15 Police Authority and Self-Defense
+- 00:44:12 Discussion on Legal Standards and Analysis
+- 00:48:05 Video Evidence and Analysis
+- 00:53:52 Imminence of Threat
+- 00:56:02 Conclusion and Final Thoughts
+
+**chrissie-mayr/cfLF2o2-0BA** (long, 03:29:13)
+
+- 01:27:49 Disagreement about the impact of your opinion on others
+- 01:28:06 Discussion on a potential conservative Barbie movie
+- 01:28:54 Criticism of the original Barbie movie's content and tone
+- 01:30:09 Comparison between different opinions on the same topic
+- 01:31:00 Defense of a male reviewer's opinion on Barbie
+- 01:32:05 Criticism of Ben Shapiro's review as overly negative and unbalanced
+- 01:34:06 Discussion about potential Ken movies with Ryan Gosling
+- 01:37:03 Exploration of different interpretations of feminism
+
diff --git a/plans/bakeoff/sample.json b/plans/bakeoff/sample.json
@@ -0,0 +1,93 @@
+{
+ "version": 1,
+ "pickedAt": "2026-07-26T17:58:35.059Z",
+ "corpus": {
+ "videosScanned": 76354,
+ "videosWithTranscript": 73367,
+ "audioHours": 77298,
+ "longTailVideos": 6038,
+ "longTailAudioHours": 40153
+ },
+ "videos": [
+ {
+ "slug": "the-quartering-rumble/v6ve4n0",
+ "channelSlug": "the-quartering-rumble",
+ "videoId": "v6ve4n0",
+ "videoDir": "v6xl0vu",
+ "title": "Donald Trump TROLLS Stephen Colbert While Setting Up GENIUS Trap That Woke Leftists Immediately Fall",
+ "bucket": "short",
+ "durationSeconds": 791,
+ "cueCount": 134
+ },
+ {
+ "slug": "chibi-reviews/VQykVuHd9xQ",
+ "channelSlug": "chibi-reviews",
+ "videoId": "VQykVuHd9xQ",
+ "videoDir": "VQykVuHd9xQ",
+ "title": "Crunchyroll This is Unacceptable! Your Service is LITERALLY Unwatchable",
+ "bucket": "short",
+ "durationSeconds": 433,
+ "cueCount": 173
+ },
+ {
+ "slug": "the-quartering/J0ySGwzP4Nw",
+ "channelSlug": "the-quartering",
+ "videoId": "J0ySGwzP4Nw",
+ "videoDir": "J0ySGwzP4Nw",
+ "title": "Tim Pool Goes To WAR With Twitter After Being Banned For Saying The FORBIDDEN Word On Timcast IRL",
+ "bucket": "short",
+ "durationSeconds": 785,
+ "cueCount": 315
+ },
+ {
+ "slug": "destiny/5nmDzKB23OU",
+ "channelSlug": "destiny",
+ "videoId": "5nmDzKB23OU",
+ "videoDir": "5nmDzKB23OU",
+ "title": "Can Science Answer All Questions?",
+ "bucket": "medium",
+ "durationSeconds": 4419,
+ "cueCount": 2072
+ },
+ {
+ "slug": "angryjoeshow/RPJqkewZP5I",
+ "channelSlug": "angryjoeshow",
+ "videoId": "RPJqkewZP5I",
+ "videoDir": "RPJqkewZP5I",
+ "title": "AJSN WK8A- AngryJoe Eye Surgery, HIGHGUARD Laysoff 80%, Riot Laysoff 2XKO Devs, Sony State of Play!",
+ "bucket": "medium",
+ "durationSeconds": 3380,
+ "cueCount": 1526
+ },
+ {
+ "slug": "rekietalaw/EsZhaCfc8HQ",
+ "channelSlug": "rekietalaw",
+ "videoId": "EsZhaCfc8HQ",
+ "videoDir": "EsZhaCfc8HQ",
+ "title": "ICE Agent Kills Woman In Minneapolis: Twitter Most Affected",
+ "bucket": "long",
+ "durationSeconds": 12497,
+ "cueCount": 3997
+ },
+ {
+ "slug": "chrissie-mayr/cfLF2o2-0BA",
+ "channelSlug": "chrissie-mayr",
+ "videoId": "cfLF2o2-0BA",
+ "videoDir": "cfLF2o2-0BA",
+ "title": "SimpCast 85! Chrissie Mayr, Wicked Virtue, Monica Paige, Brittany Venti, Lauren, - Barbie, Whatever",
+ "bucket": "long",
+ "durationSeconds": 12553,
+ "cueCount": 4692
+ },
+ {
+ "slug": "HasanAbiVODs/GjPX_ueTdfc",
+ "channelSlug": "HasanAbiVODs",
+ "videoId": "GjPX_ueTdfc",
+ "videoDir": "GjPX_ueTdfc",
+ "title": "2/2 HasanAbi April 7, 2021 - Sunburn OMEGALUL, Basketball Memes, OKBUDDY, 🎮GTA NoPixel🎮 FULL VOD",
+ "bucket": "verylong",
+ "durationSeconds": 28965,
+ "cueCount": 7223
+ }
+ ]
+}