# Roadmap: the local-AI transformative layer Long-running work adding a derived, transformative layer over the verbatim transcript archive: timestamped chapters, topic tags, speaker attribution, a graduated visibility policy, and eventually a site that leads with the derived corpus rather than raw reproduction. This file is the **durable roadmap** and changes rarely. Three companions: - [`plans/FACTS.md`](plans/FACTS.md) — verified codebase facts. Trust it over re-deriving. - [`plans/STATE.md`](plans/STATE.md) — current status, decisions log, open questions. - `plans/phase-N-*.md` — file-level detail for the phase in flight, written when it starts. See [`plans/README.md`](plans/README.md) for the context-clear protocol. ## Why Two goals drive the work. **Utility.** A timestamped summary per video lets a reader navigate a 3-hour livestream without scrubbing, and gives search something to match beyond raw speech. **Copyright posture.** The archive is more defensible when it publishes new expression *about* works alongside — and eventually instead of — full reproduction of them. Chapters, topic tags, and analysis are genuinely transformative; a graduated per-channel visibility policy replaces all-or-nothing takedown handling; and filtering out embedded third-party clips reduces reproduction of works the archive has the weakest claim to. ## Non-goal: transcripts are never rewritten Verbatim accuracy is what makes search, `[n @ mm:ss]` citations, and the MCP corpus sweeps trustworthy. A full-length paraphrase would still be a derivative work while adding hallucination risk. Transformative value is layered **on top of** verbatim cues, and reduction of exposure happens by **withholding** cues, not by altering them. ## Architecture One derived sidecar per video directory, mirroring the `transcript.cues.json` precedent: versioned, self-describing, written tmp+rename, with an mtime freshness check. ``` transcripts/channels//data// transcript.cues.json # verbatim source, never rewritten ai-digest.json # chapters + tags + per-section provenance diarization.json # phase 9 attribution.json # phase 9 ``` **Never name a sidecar `transcript..`** — `SUB_FILE_RE` claims any such file as a subtitle track. Likewise **never use the name `summaries`** for the new artifact; that is already the cross-channel metadata browse index. Both hazards are detailed in `FACTS.md`. ```mermaid flowchart LR cues["transcript.cues.json
(verbatim, immutable)"] --> md["transcriptToMarkdown
stampForCue = hms"] ctx["AI-CONTEXT.md → channel → video
(phase 1.5)"] --> md md --> chunk["chunkCuesForContext
(NEW — windowCues cannot do this)"] chunk --> eng{"digestApps registry"} eng -->|"local-gpu lane"| oll["ollama-direct
/api/chat + JSON schema"] eng -->|"remote-api lane"| cc["claude-code
claude -p --output-format json"] oll --> parse["digestParse
range · monotonic · snap · warnings"] cc --> parse parse --> dig["ai-digest.json
+ per-section provenance"] dig --> bi["buildIndex: digestMs → digests page tree"] bi --> pub["/digests/<slug>/page-NNNN.json"] pub --> ui["viewer ?vm=summary · search · MCP · corpus.json"] ``` ### Engine registries, and why there are two lanes The repo's established idiom for "swappable engine" is a **registry of descriptors with a client-safe listing function** — `TRANSCRIPTION_APPS` / `listTranscriptionApps()`. The digest layer copies that idiom exactly rather than inventing a second pattern. Transcription already has three engines; digests get two. | Lane | Engine | Queue key | Runs alongside | Cost | | --- | --- | --- | --- | --- | | `local-gpu` | `ollama-direct` (default) | `TRANSCRIPTION_QUEUE` | nothing — the 8 GB card is shared with whisper/parakeet | free | | `remote-api` | `claude-code` | `DIGEST_REMOTE_QUEUE` | the GPU lane | metered | Two lanes on **different queue keys** is the structural idea that makes Claude Code useful rather than merely available: a backfill can drive both at once, Ollama saturating the GPU while Claude Code works the same backlog over the network. Putting both on one key would serialize them and waste the network lane. Putting the local engine on its own key would let Ollama and the transcription engine thrash the same 8 GB of VRAM. Metered engines are **opt-in** (`digest.remoteEnabled`, default `false`) and are never the auto-queue default by inheritance. ### Provenance — which AI was used Authoritative per **section**, because the digest controller does a read–modify–write merge and a video can legitimately carry Ollama chapters and Claude Code tags: ```jsonc { "version": 1, "contextHash": "sha1…", // phase 1.5 — invalidates on channel-context edits "chapters": [{ "start": 0, "title": "COLD OPEN" }], "tags": ["game review"], "warnings": [{ "section": "chapters", "reason": "timestamp-out-of-range", "raw": "01:10:29 …" }], "sections": { "chapters": { "appId": "ollama-direct", "model": "qwen2.5:7b", "lane": "local-gpu", "generatedAt": "2026-07-25T…", "promptVersion": 1 }, "tags": { "appId": "claude-code", "model": "claude-opus-5", "lane": "remote-api", "generatedAt": "2026-07-26T…", "promptVersion": 1 } } } ``` `engines: string[]` (the distinct `appId`s) is derived on read for cheap display and filtering — do not persist it as a second source of truth. Provenance must reach every surface, not just disk: the per-video editor badge, a `digestEngines: Record` stat on `ChannelSnapshot` (a sibling of `totals`, **not** a `buckets` entry — buckets are work-lanes of video ids), the digest record shipped to the export, and a viewer badge. The Phase 10 honesty requirement extends to *which* model, not merely *that* it was generated. ## Phases ```mermaid flowchart TD P0["0 · Benchmark engines"] --> P1["1 · Generation harness"] P1 --> P15["1.5 · Channel context"] P15 --> P2["2 · Digest corpus in build"] P1 --> P25["2.5 · Observability"] P25 --> BF(["BACKFILL"]) P15 --> BF P11a["11a · Review queue"] --> BF P2 --> P3["3 · Viewer ?vm=summary"] P2 --> P4["4 · Search indexing"] P25 --> P5["5 · Auto-queue"] P3 --> P10["10 · Lead with the corpus"] P4 --> P10 P8["8 · Visibility policy"] --> P10 P2 --> P9["9 · Attribution + quote filter"] P6["6 · Ollama /ask provider
(no dependencies)"] P1 --> P7["7 · Tags + chat highlights"] P10 --> P11b["11b · Viewer feedback"] ``` ### Phase 0 — Benchmark the transcription engines Small, and it belongs first because **transcription is the upstream bottleneck for the digest backfill**: a video cannot be digested until it has a transcript. `parakeet` is **already** a fully selectable engine — the third entry in `TRANSCRIPTION_APPS`, exposed in settings through `listTranscriptionApps()`, with a `device` field and `supportsPartialStop: true` (it stops after the current window and stitches a partial transcript on SIGTERM, so a long run is interruptible — directly relevant to a multi-day sweep). The default is `whisper-cpp`. Nothing is missing structurally. What *is* missing is evidence. Nothing in the repo measures relative throughput on this hardware, so "if parakeet is faster" is currently unanswerable. - Time `whisper-cpp`, `chough`, and `parakeet` over the same handful of real videos of varying length, on this GPU. Record wall-clock per audio-minute, VRAM, and transcript quality spot-checks. - Write the numbers into `plans/FACTS.md`. This is a fact that will be re-asked for months and should never need re-measuring. - If parakeet wins, change `DEFAULT_TRANSCRIPTION_APP_ID` and note it in `STATE.md`. Because workers already carry a per-worker `appId`, a mixed fleet is also possible without new machinery. - Deliverable is measurements and possibly a one-line default change — **not** a new benchmark harness. Do not build infrastructure for a question asked once. ### Phase 1 — Generation harness Modelled on the transcription-app registry, **not** on `askProvider.ts` — that runs in the visitor's browser in the static export and cannot spawn a process. - `common/lib/paths.ts` — add `ollamaUrl` (`OLLAMA_URL`) and `claudeBin` (`CLAUDE_BIN`), following the `process.env.X ?? default` idiom. Add both to the hand-curated `pathRows` in the settings page's System paths table. - `common/lib/digestApps.ts` *(new)* — mirror `transcriptionApps.ts`: a `DigestApp` type, a `DIGEST_APPS` record, a `getDigestApp(id)` that never throws, and a client-safe `listDigestApps()` descriptor export so the registry never reaches the client bundle. Each app carries `lane` and `metered`. - **`ollama-direct`** — POST `${ollamaUrl}/api/chat` with `format: `, `options: { num_ctx: 16384, temperature: 0 }`, `stream: false`. **`num_ctx` must be explicit**: the 4096 default silently truncates, the single easiest way to get quietly-wrong output at scale. - **`claude-code`** — shell out to `claudeBin -p --output-format json`, parse the wrapper's `result` string, then the JSON body. No context-window concern, but keep the same chunking so output shape stays engine-independent. - **`fabric`** — deferred. Measured to ignore the `HH:MM:SS TOPIC` contract at 7B, twice, at two input sizes. Its 256-pattern library stays valuable as prompt source material; it becomes a third registry entry when a prose-shaped output needs it, with no rework. - `common/lib/transcriptWindow.ts` — add `chunkCuesForContext(cues, { maxCues, overlapCues })` returning sequential overlapping slices. **`windowCues()` cannot do this** — it is a center-based helper for search hits (see `FACTS.md`). - `common/lib/digestParse.ts` *(new)* — parse engine output into `{ start, title }[]`. Must reject timestamps beyond the video duration, **clamp each chunk's output to that chunk's own range**, enforce monotonic starts, drop malformed entries, snap starts to the nearest cue boundary, and de-duplicate across chunk seams. Every rejection is **recorded in a `warnings` array, never silently dropped** — Phase 11's review queue is built entirely on this. Unit-tested; it is the layer absorbing local-model sloppiness. - `common/controller/digestVideo.ts` *(new)* — freshness check → `readNormalizedTranscript` → `transcriptToMarkdown` with `stampForCue: (_c, s) => hms(s)` → chunk → engine → parse → read-modify-write merge → tmp+rename. Reachability probe for the Ollama lane; a friendly ENOENT message for the Claude lane. - `common/controller/digestBatch.ts` *(new)* — `runPool()`, honoring `signal` and `ctx.drainSignal`. Logs a running count of **metered** calls so a remote backfill's cost is visible in the job log as it happens. - `common/lib/queueKeys.ts` — `DIGEST_REMOTE_QUEUE` plus `digestQueueKey(appId)` mapping lane → key. No concurrency declaration; the registry hardcodes 1 per key. - One `jobKinds.ts` entry (`digest-channel`), one server action modelled on `whisperActions.ts`, one replay handler. - Settings: `defaultDigest()` + `sanitizeDigest()` wired into `getSettings()`'s normalize chain, with per-engine config keyed by `appId` and `remoteEnabled: false`. - UI: a channel stage component, an engine picker, a per-video button, and a **provenance badge** (appId + model + generated-at) on the video page. - `common/package.json` — add a `test` script; there is currently none, and ~46 unit test files are unrunnable as a suite. ### Phase 1.5 — Channel context layers Cascading human-authored notes injected into every prompt, modelled on this repo's own `CLAUDE.md` → `AGENTS.md`. Lands early because every later phase improves from it, and Phase 9's attribution is close to unusable without it. ``` transcripts/AI-CONTEXT.md # corpus-wide transcripts/channels//context.md # the main one transcripts/channels//data//context.md # rare, for oddities ``` Markdown with optional frontmatter: `hosts`, `recurring_guests`, `plays_third_party_media`, `boilerplate`. - `common/lib/aiContext.ts` *(new)* — `resolveAiContext(paths, slug, videoId?)` follows the `resolveCookiePolicy` inheritance shape. Caps total tokens (~1500) so context never crowds out transcript. **There is no YAML dependency in the repo** — hand-roll a minimal scalar/list frontmatter parser rather than adding one for three field types. - `boilerplate` earns its keep with no model involvement: a deterministic filter dropping recurring sponsor/subscribe cues from digest input and chapter candidates. Cheap, exact, and it removes the most common junk chapter. - `plays_third_party_media: false` lets Phase 9 skip attribution entirely for that channel — a large accuracy and compute win. - Editing UI on the channel and settings pages (tmp+rename). Frontmatter errors surface inline, never at job time. - A `propose-channel-context` job samples N transcripts and drafts `context.suggested.md`. It **never** writes `context.md` — auto-applying model-authored context would let one bad inference silently degrade every downstream summary for a channel, invisibly, because the output would still look plausible. - Hash the resolved context into `ai-digest.json.contextHash` so editing notes correctly marks digests stale. **Corrected 2026-07-29:** this is NOT why backfill must wait. The empty note already hashes to a stable value, so a note added later invalidates only its own channel. The corpus-wide risk is the deliberate one — bumping the `digest-context-v1` hash prefix — so what must precede a sweep is that DECISION, not this phase. ### Phase 2 — Digests as a first-class corpus — **DONE (2026-07-29)** > Shipped as specified, plus three things this spec omitted: `exportDigestsDir` / > `exportSharedDigestsDir` in `paths.ts` (compose cannot be written without them), a > per-SITE digests manifest so `corpus.json` can source per-channel counts the way it does > for posts, and `digests` added to the service worker's `SHARD_RE` (without it the tree has > no offline story — the gap already recorded for `duplicates.json`). Note `?vm=digest`, not > `?vm=summary`; see the naming decision in `plans/STATE.md`. An earlier draft inlined the digest into `TranscriptDetail`. That is wrong: under `summary-only` visibility a transcript page ships **no cues**, so a digest must never be bundled inside a payload that policy may withhold. Digests get their own page tree. - `buildIndex.ts` — three new LMDB sub-DBs (`digests`, `digestPageHashes`, `channelDigestStats`); **no `maxDbs` change needed**, but three new `clearAsync()` lines in the hand-enumerated schema-invalidation block; `SCHEMA_VERSION` 12 → 13; `digestMs` (later `attributionMs`) added to `MtimeRecord`, `LiveEntry`, the `scanSource()` stat loop, the changed-detection comparison, and the `mtimes.put` call; emit `.export-index/shared/digests//{manifest,page-NNNN}.json` via `createPageWriter()`. Each record carries its `sections` provenance. - `common/lib/digests.ts` *(new)* — `VideoDigest` / `DigestPage` / `ChannelDigestsManifest` with `slugToPage`, mirroring `manifest.ts`. - `channelSignature.ts` — add `digestMs` to its three-field structural-subset `MtimeRecord` **and** to the hash input, or archives will not rebuild on a digest-only change. - `compose-site.ts` — a `digests?` key in `ComposeCache` and a fourth `reconcileChannelTree()` call. It already passes `ignoreBasename: "manifest.json"` internally; nothing to add there. - `common/components/digestCache.ts` + `digestStore.ts` *(new)* — copy the `transcriptCache.ts` + `transcriptStore.ts` pair (**not** `subsCache.ts`, which has no IndexedDB). Use a **separate IDB store** so the transcript store's `DB_VERSION` need not bump and existing transcript caches survive. - `corpus.ts` — `CORPUS_SPEC_VERSION` 2 → 3, a `DIGEST_SCHEME` modelled on `POST_SCHEME`, and a digest manifest pointer on `CorpusChannel.manifests`. Update `corpus.test.ts`. - **Carry `derivedFrom` through to the client.** *(Added by the duplicate-detection work.)* Once digests are shared across a duplicate cluster, a video's digest may have been generated for a *different* video. `DigestRecord.derivedFrom` already records that (`{slug, clusterId, sharedAt, offsetSeconds}`, written by `writeSharedDigest`), but it stops at the record — the page-tree schema above must include it, or the viewer cannot tell a native digest from a borrowed one. Sharing that is invisible is sharing that is indistinguishable from a claim, which Phase 3 then has no way to be honest about. This is a dependency Phase 2 creates, not an optional extra. ### Phase 2.5 — Observability for the backfill Build this **before** backfilling. A multi-day sweep you cannot observe is one you cannot tune or safely interrupt. - **First**, fix the duplicated literal union in `RunningJobsList.tsx` to import `JobProgressMetric`. Until then the compiler-driven audit the rest of this phase relies on has a silent hole. Six sites total — see `FACTS.md`. - Extend `JobProgressMetric` with `"digests"` and `JobTaskKind` with `"digest"`. - The two copies of the `metric === "downloads" ? … : …` binary in `MonitorWidget.tsx` become **one lookup table keyed by metric** (label, glyph, color), rather than growing a third branch — otherwise the next metric repeats the bug. - **Coverage, not just progress.** Per-job bars answer "how's this job"; during backfill the question is "how much of the corpus is done" — `digested / total`, **split by engine**. That is the single most useful number during a multi-day sweep, and with two lanes running it is also how you see whether the remote lane is pulling its weight. - Widget sync payload gains coverage **scalars only** — honor the file's stated design constraint that it stays a handful of scalars rather than shipping per-channel breakdowns on every poll. Its builder is already reused for dashboard SSR seeding, so both surfaces get it free. - A digest instrument in `PipelineBand`, `noDigest` counts in `NeedsWorkPanel` / `ChannelsTable`, command-palette actions, and ETA via the existing `computeEtaSeconds`. - `channelSnapshot.ts` gains `noDigest` **here**, not in Phase 5; Phase 5 then consumes what already exists. ### Phase 3 — Viewer — **DONE (2026-07-29), as `?vm=digest`** > Every `"summary"` below reads `"digest"` in the shipped code: "summary" already means a > video listing card AND the `/summaries/` page tree, and this file's own naming-hazard rule > forbids a third meaning. Also note the control is HIDDEN where no digest exists (0.1% > coverage makes an always-present button a dead end), and the client cache is versioned by > page content hash rather than `generatedAt` — reasons for both in `plans/STATE.md`. - `urlState.ts` — add `"summary"` to `ModalMode`, the parse chain, and the `writeUrlParams` `vm` branch. `ModalMode` appears in only three files. - `PlayerProvider.tsx` — lazy-fetch the digest when `modalMode === "summary"`, mirroring the chat effect including its ref-based in-flight guard (there is a comment explaining why reducer state in the dep array drops results) and its snap-back-to-transcript on missing. - `TranscriptModal.tsx` — a toolbar `ControlButton` beside the transcript/chat toggle; chapters as buttons calling `seekTo(start)` with `scrollKindRef.current = "smooth"`; active-chapter highlight off `currentTime`; a provenance line; an empty state. - **Styling:** this modal and `PlayerProvider` are **not** on semantic theme tokens — they use hard-coded `bg-zinc-900/80`, `ring-white/10`, `text-white` throughout. Match that local palette; do not introduce `bg-card`/`text-foreground` here. Migrating the modal chrome is out of scope. - **Show a borrowed digest as borrowed.** *(Added by the duplicate-detection work.)* When `derivedFrom` is set (Phase 2), the provenance line must name the member the digest was generated *for* and the measured offset — not present it as native to the video being watched. The chapters are placed by the canonical member's timeline; the detector's `aligned` verdict is why sharing was allowed at all, and its `offsetSeconds` is the residual error the viewer is looking at. A shared digest presented as native is the failure mode that looks like success: every chapter plausible, all of them describing a different upload. ### Phase 4 — Search indexing Fold chapter titles and tags into the search corpus with a `source: "ai"` marker and the producing `appId`. **Two independent index paths exist** and both need it: the client-side FlexSearch worker and the server-side MCP search. Verbatim and generated hits must stay visually distinguishable — that honesty requirement is what makes indexing generated text acceptable at all. ### Phase 5 — Auto-queue A third `AutoQueuePolicy` alongside `transcription`/`download`, a dispatch branch in `autoRunner.ts` re-reading settings per iteration so a pause flag takes effect live, exposure in the policy tree editor, a `digestsPaused` flag, and a pause button copied from the downloads one. **The policy carries an explicit `engineId` defaulting to the local lane** — a metered engine must never become the auto-queue default by inheritance. ### Phase 6 — Ollama provider for `/ask` Zero dependencies on any other phase; a good early win. Widen the `Provider` union, add a `PROVIDERS` entry (no key required), a `switch` arm, and an `askOllama` using Ollama's OpenAI-compatible `/v1/chat/completions` SSE shape — the existing `askOpenAI` is a near-template. Make the base URL editable in the provider settings panel. **Caveat to surface in the UI:** the export is static and runs in the visitor's browser, so this only works when the visitor can reach an Ollama instance, and requires `OLLAMA_ORIGINS` for CORS. Local/LAN use only. Knock-on benefit: the MCP corpus sweep's extraction step can then run locally. ### Phase 7 — Tags and live-chat highlights Tags come free with Phase 1 into the same `ai-digest.json`. Live-chat highlights reuse the whole harness against `live_chat.cues.json`, writing `ai-chat-digest.json` and rendering in the existing `?vm=chat` view. ### Phase 8 — Transcript visibility policy Mandatory, not optional — leading with derived data means the published site's default posture is set here. - `common/lib/visibilityPolicy.ts` *(new)* — copy `cookiePolicy.ts` exactly: `"full" | "excerpt" | "summary-only"`, a values array, a default, a type guard, and `resolveVisibility(settings, channelConfig)`. Note cookiePolicy uses **two** inheritance rules — truthiness-after-trim for free-text values, type-guard for enums; this is the enum case. - Wire through settings, `parseChannelConfig`, both forms — and **list the key in `CHANNEL_FORM_FIELDS`** or clearing back to inherit silently won't work. - **Enforced at build time** in the `buildIndex.ts` page writer, the single place that decides what `cues` array ships. One enforcement point covers viewer, MCP, `/ask`, and search, because all four read the same shards. - **Archives bypass the page writer** — `archiveTranscripts.ts` hard-links `transcript.cues.json` straight from the source tree. Extend its existing `requiresRewrite` predicate and `writeTransformedCues` path. This is the highest-exposure surface and inherits nothing for free. ### Phase 9 — Attribution (two lanes) + quote filtering > **CAPTURE AND BOTH ATTRIBUTION LANES LANDED 2026-08-07. Nothing is armed.** > `diarization.json` per video with the cleanup guard, then `attribution.json` with both > lanes, `attributionStatus.ts`, and the freshness comparators — all off by default, under > an off-by-default master switch. What is still to do: `attributionMs` through the build, > viewer badges, per-channel export counts, and quote filtering (which additionally > depends on Phase 8). Measurements that changed the design are in > [`plans/diarization-spike-results.md`](plans/diarization-spike-results.md); the > decisions are in `plans/STATE.md`. > > **Four corrections to what this section says.** > > **1. The sequencing below is inverted.** "Run the text-only pass over the whole backlog > immediately, then upgrade selectively" was written before the numbers existed. Text-only > attribution must read the transcript to find speaker changes at all, so it costs roughly > the digest sweep's chunk count — ~194,000 model calls, another 25–55 GPU-days — and it > would spend that while the digest sweep has completed **0.17%** of its own (122 > `ai-digest.json` of 73,367). Both want the same 8 GB card and nothing arbitrates between > them. The diarized lane costs about **one call per video** and is better, and its only > problem is input coverage — which the backfill lane now fixes. So: both built, neither > armed, and a measured pilot picks the default. > > **2. The upgrade job does not need building.** This section describes a bespoke job that > "re-downloads audio, diarizes, attributes, then removes the re-fetched audio in a > `finally`". That job now exists generically: `attribution-diarized` reports a text-only > record as `missing` work, so the lane, the sweep and the indicators queue the upgrade > with no new machinery, and the re-acquire-and-delete-in-a-`finally` half is > [`common/controller/backfillReacquire.ts`](common/controller/backfillReacquire.ts). > > **3. Audio deletion is not automatic** — `cleanAudioFromTranscribed` has one call site, > an operator-triggered action, so the deadline is disk pressure, not the transcribe job. > > **4. Diarization does not run "right after transcription" by default**: it is 2–3× > slower than the transcription it would follow (221 s/audio-hour on the GPU vs ~500–680 on > the CPU), so the inline hook ships off and the intended shape for a batch is capture-on, > inline-off, backfill after. > > **And one thing NOT to re-propose: attribution cannot be a digest section.** Folding > `speakers` into the digest prompt looks like it halves the bill. Digests are generated > **chunk-local**, so each chunk is labelled with no knowledge of the others — and the one > property attribution needs above all is that speaker 0 in chunk 1 is the same person in > chunk 30. > Largest and least certain. Ship 1–8 first; keep it off by default. **Blocking constraint:** audio is deleted once a video is transcribed, except for dirs holding `do-not-clean.json` or entries in the saved-video store. **Most of the existing archive has no audio to diarize.** Two first-class lanes producing the same artifact at different quality tiers: | Lane | When | Input | Recorded as | | --- | --- | --- | --- | | Diarization-assisted | Going forward; on-demand upgrade | audio → speaker turns → LLM labels the turns | `method: "diarized"` | | Text-only | Legacy videos with no audio | cues alone → LLM segments and labels from content | `method: "text-only"` | Neither is a fallback for the other in code. Going forward, capture diarization while audio is still on disk and **before** the cleanup sweep. (The original text said to run the text-only pass over the whole backlog immediately and upgrade selectively — see correction 1 above: that is 25–55 GPU-days spent on the worse lane, and the sequencing is inverted.) - `scripts/diarize.mjs` behind a `DIARIZE_BIN` path, exactly as the parakeet wrapper works. Keeps the engine swappable: evaluate `sherpa-onnx` (CPU-friendly, no HF token) against `pyannote` (better, needs token + GPU) **on real audio before committing** — quality here cannot be judged from code. - ~~`attributeCues.ts`~~ **`common/controller/attributeOne.ts`** writes `attribution.json` with ranges, confidence, method, and engine provenance. Both lanes live in it because they share every guard around the artifact; only the middle differs. - Filter in the page writer alongside Phase 8. **Bias toward dropping on uncertainty** — the cost is asymmetric. Make the threshold policy-driven so a channel can filter on `diarized` only and ignore `text-only` labels it doesn't trust. - Status everywhere: a client-safe `attributionStatus.ts` (`"none" | "text-only" | "diarized" | "stale"`), new snapshot buckets, a per-video badge, per-channel counts. Mark a video **ineligible** when availability says deleted/private and no audio is retained — a distinct state, not an upgrade button that can only fail. - ~~Re-processing: per-video re-attribute / upgrade, per-channel bulk over the text-only bucket. The upgrade job re-downloads audio, diarizes, attributes, then removes the re-fetched audio **in a `finally`**…~~ **Done generically — see correction 2 above.** `attribution-diarized` reports a text-only record as `missing`, so the backfill lane, the corpus sweep and the per-channel button all queue the upgrade already; the fetch-use-delete-in-a-`finally` half is `controller/backfillReacquire.ts`, which honours `do-not-clean.json` and a free-disk floor and has e2e tests for the leak paths. - `attributionMs` threads through the build the same way `digestMs` does. - **Set expectations honestly:** this misfires on rapid back-and-forth, and auto-caption channels have no speaker turns at all with cue boundaries that don't align to them. A good-faith reduction in reproduction, not a guarantee; the UI must not claim otherwise. ### Phase 10 — Leading with the derived corpus Depends on 2, 3, 4, 8. - **MCP** — a `get_digest` tool and digest-aware search, plus digest members on the `ShardSource` interface implemented in **all three** classes (local, remote, hub). The hub path needs it too, or federated sites silently lack digests. - **Archives** — a digests zip. Remember the 25 MB default cap drops *all* archives for large sites. - **UI inversion** — search results lead with chapters and tags, verbatim cue matches secondary and badged; listings surface chapter counts and topic tags; the modal defaults to `?vm=summary` for `summary-only` channels (a default-selection change — the mode is already URL-driven). - **Graceful fallback** where a digest is absent. A partially-digested corpus is the normal state for a long time and must not look broken. - **Honesty:** generated content visually distinct from verbatim everywhere, labelled with the producing model. This matters more once generated text is the primary thing users see. ### Phase 11 — Human review queue & viewer feedback Without this, a corpus-wide backfill is a one-shot gamble on prompt quality. **Naming hazard:** `report` already means three different things in this repo. Use **`feedback`** for viewer-submitted items and **`review`** for triage state. Do not add a fourth meaning of "report". **11a — review queue (land before backfilling).** **REVISED 2026-08-25 by [`plans/editor-operations-ia.md`](plans/editor-operations-ia.md), which dissolves `/actionable` in its slice 4.** The instruction here used to be "extend `/actionable` rather than building a parallel one", and the reasoning behind it still holds — do not build a parallel page — but the destination moved: `/actionable` was one of four different answers to "what needs doing", and the operation is the noun those answers are folding into. So the same sections, filed by what they are about: - **Per-operation review** — digests needing review (driven by the `warnings` array Phase 1 persists) and uncertain attribution — becomes the *attention* section of `/operations/`, beside that operation's own lane, sweep and band. - **Corpus review** — duplicate clusters awaiting confirmation, viewer feedback, and proposed channel context awaiting promotion — becomes a small `/review` page under Corpus. These are judgements about the ARCHIVE, not about one operation's output. The `SectionConfig` array with its counters in `loadActionable.ts` is still the designed extension point and still the thing to extend; it **moves with the sections** rather than disappearing (the widget payload reads `loadActionable.ts` directly and must keep working). **Duplicate clusters awaiting confirmation** *(added by the duplicate-detection work — UI only, the data and the write path already exist).* A `title-duration` cluster is a suspect: two videos share a title and a near-identical runtime, and nothing compared their content because at least one side has no transcript. It ships nowhere (`clusterIsPublishable`) and shares no derived work (`clusterMaySharePartial`) until a human records `confirmed: true`. Corpus-wide that is currently **335 clusters**. Everything needed is in place: - the queue is `report.clusters.filter(c => c.needsReview)` minus `isClusterReviewed(...)`; - `updateDuplicateOverride(paths, clusterId, {confirmed})` already accepts and persists it, and its clear-heuristic already refuses to delete a confirmation; - `/actionable` already renders cluster cards with a `needs review` badge. The missing piece is two buttons — *these are the same* / *these are not* — on that card. Actions per item: approve · edit inline · regenerate · dismiss · **"add note to channel context"**. The last is the one that compounds — a correction applied to one video fixes one video; the same correction written into `context.md` fixes every future generation for that channel. **Regenerate should offer the other engine**: "this looks wrong, redo it on Claude Code" is the most natural use of a second lane, and the provenance record makes it obvious when one engine is systematically weaker on a given channel. **11b — viewer feedback (can follow the corpus going public).** A "flag this" control per chapter, with categories including **"misrepresents what was said"**. Transport in preference order: an optional per-site `feedbackUrl` POST, else copy-to-clipboard / download-JSON. **Collector evaluated — the answer is no, not `r2-proxy/`.** *(Recorded by the duplicate-detection work so the question is not re-opened.)* The proxy is **read-only by explicit design**: `src/index.ts:35-41` rejects every non-GET/HEAD method, and `KEY_RE` at `:27` carries a comment stating the proxy must never become a general oracle over the bucket. Turning it into a write endpoint works directly against that intent, for a feature that does not need a server at all. Its one genuinely reusable piece is the `RATE_LIMITER` binding. **Use the zero-infrastructure fallback this phase already names** — copy-to-clipboard / download-JSON, pasted into an editor action. It needs no deployment and matches the existing idiom at `PlayerProvider.tsx:131-136`. **New trigger for this phase:** sharing one digest across a duplicate pair is precisely the case a viewer needs to be able to flag. A borrowed digest can be misplaced (the mirror drifts after the anchors that were measured) or simply wrong for that upload, and the viewer is the only party who will ever notice — the detector already believes the two are the same video. Ingestion is an editor action accepting pasted JSON. Treat submissions as **untrusted input rendered in an admin UI**: escape it, never feed it into a prompt unreviewed. Store under `transcripts/.feedback/`, a sibling of `.jobs` and `.scheduler`, outside the build trees. **Why that category is not boilerplate:** the smoke test summarized allegation-heavy content about named individuals and restated those allegations as plain fact. A published AI summary that misstates what a real person said or did is the highest-risk output this system can produce, and the one failure mode no technical guard catches. Wire `dismiss` so it can **suppress a digest from the next build**, not merely flag it for later. ## Verification strategy Per the repo's `verification-playwright-first` convention: verify with Playwright, open a browser only to diagnose failures. 1. **Stub Ollama** — the default engine is an HTTP call from the Next server, so the fake-binary trick does not cover it. Add an `ollama-stub.mjs` fixture launched from `playwright.config.ts` alongside the existing web servers, with `OLLAMA_URL` pointed at it. Give it a **deterministic bad-output mode** (out-of-range, non-monotonic) so the parser guards are exercised by tests rather than by luck in production. 2. **Fake `claude`** — a `fake-claude.mjs` echoing the `--output-format json` wrapper, plus the `SLOWOP` paced-output trick so drain/progress specs can observe it. Wire `CLAUDE_BIN` into **both** the `dev:test` and `start:test` lines; they are duplicated verbatim. 3. **Unit** — `digestParse.test.ts` and the chunker cases, using `node:test`, runnable via the new `common` test script. 4. **Editor e2e** — queue a job, assert the row label, assert `ai-digest.json` lands with the right `sections.chapters.appId`. Post-mutation assertions must **poll-with-reload**; the channel page serves a snapshot regenerated on a ~1 s debounce. 5. **Two-lane e2e** — queue both engines and assert they occupy **different queue keys** and run concurrently. This is the behavior that makes the Claude Code lane worth having, so it needs a test, not a comment. 6. **Export e2e** — clone the route-stubbing scaffold in `export-player-platform-cache.spec.ts`: stub manifest/pages with a digest-bearing fixture, open `?v=&vm=summary`, assert chapters render, clicking one seeks, and the engine badge shows. 7. **Glance surfaces** — assert the sync payload carries coverage split by engine and the pipeline band shows the digest chip and paused state. Run widget assertions under `E2E_MODE=start` (see the Dev Tools caveat in `FACTS.md`). 8. **Full suite** — `pnpm e2e` in **default dev mode**; `E2E_MODE=start` serves a stale build. Kill stale dev servers by port between runs. The known-failing-on-base list is in `FACTS.md` — do not chase those as regressions. 9. **Real end-to-end** — one real channel against live Ollama, then `pnpm build:index && pnpm --filter export run build`; confirm the digest survives into `export/public/digests//page-*.json`. For Phase 9, confirm a `summary-only` channel ships **zero** cues by inspecting the shard directly, not the UI. 10. **Changelogs** — `editor/CHANGELOG.md` and `export/CHANGELOG.md` per repo convention. ## Sequencing | Group | Why they group | | --- | --- | | 0 | Measurement only. Answers the transcription-speed question once, permanently. | | 1, 1.5, 2, 2.5, 3 | The shippable core: generate → context → build → observe → display. | | 4, 5, 6, 7 | Each independently shippable. **6 has no dependencies** — good early win. | | 8 | Self-contained; prerequisite for 10. | | 9 | Largest and least certain. Off by default. | | 10 | Depends on 2, 3, 4, 8. | | 11a / 11b | Review queue before backfill; viewer feedback after the corpus is public. | ### Ordering traps - **CORRECTED (2026-07-29): 1.5 is not a backfill gate — a hash-prefix decision is.** `digestContext-server.ts:31-35` hashes the EMPTY note to a stable value, so adding a channel note later invalidates only that channel, not the corpus. What would invalidate everything is bumping the `digest-context-v1` prefix (`:42`), which 1.5-as-specced does. So the gate is a decision to take before sweeping, not work to do first. - **CORRECTED (2026-07-29): "before 2.5 and 11a" named the wrong things.** Both were largely landed — the metric union, `METRIC_PREFIX`, per-job progress and the `noDigest` bucket all existed. The things actually missing were never on this list: **no corpus-wide launcher, no boot-time resume, no pause writer, and no GPU arbitration**. All four are now built (`controller/digestSweep.ts`, `controller/digestYield.ts`, `instrumentation.ts`, `DigestSweepControls`). The remaining honest gate is the review path, and its minimum — persisted failure warnings plus a `digestWarnings` bucket — is in. - **Benchmark before committing to backfill.** The smoke test measured 15.5 s for a ~2.7k token chunk; a 176-minute podcast needs roughly a dozen chunks. Budget minutes per long video and multiply by corpus size. The remote lane changes this arithmetic — measure both, then split the backlog between them. - **RESOLVED: the duplicated metric union is fixed.** `RunningJobsList` imports the union and `MonitorWidget` uses a `METRIC_PREFIX` lookup table, so extending it is now compiler-checked. - **Price a cost lever in AUDIO-HOURS before believing it** (`common/bin/digest-plan.ts`). Measured 2026-07-29: duplicate-cluster sharing is worth **~1.7 sweep days of 80**, not the "~11%" `digestSharing.ts`'s header claims — cluster members are 19% of the corpus by video count but **4.3% of its audio-hours**, because mirrors skew SHORT and the sweep is dominated by unclustered long-form VODs. GPU contention with whisper is worth ~56 days on the same measurement, roughly fifty times more than every duplicate lever combined. ### Hardware and prerequisites Radeon RX 6600, 8 GB VRAM, 15 GB system RAM. This is the binding constraint on model choice: target a 7–8B at Q4 (~5 GB), leaving ~3 GB for KV cache. A 14B at Q4 (~9 GB) spills to CPU and makes corpus-wide generation impractical. `ollama-vulkan` 0.32.4 verified with `100% GPU` offload. The systemd unit ships **inactive and disabled**, with no `~/.ollama` and no models: ```bash sudo systemctl enable --now ollama ollama pull qwen2.5:7b # ~4.7 GB Q4; llama3.1:8b is the alternative ``` The Claude Code lane needs the `claude` CLI installed and authenticated on the host, plus `digest.remoteEnabled` turned on in settings. ## Beyond the AI track: releases The phases above are the AI track. Later work is planned and recorded one release per file under `plans/` (releases 5–16: `plans/release-5.md` … `plans/release-16.md`). From release 17: | release | what | file | status (2026-10-09) | |---|---|---|---| | 17 | the media tier: a channel's big files on another drive, its text always on the corpus disk | [`plans/release-17.md`](plans/release-17.md) | complete, LIVE 2026-10-02 | | 18 | publishing as queueable stages: one index, per-site bundles, serial builds, deploys checked live, the publish lane | [`plans/release-18.md`](plans/release-18.md) | complete on `r18/integration` (`main` merged in, `4cffda3f`); fast-forward + rollout owed | | 19 | agents run the archive: ops/CLI/MCP (Track A), machine safety and tooling (Track B), OPERATING.md and the docs (Track C) | [`plans/release-19.md`](plans/release-19.md) | in flight | | 20 | the data model and the index: recorded dates, Twitch ids, the caption-track bug closed | [`plans/release-20.md`](plans/release-20.md) | planned, after 19 | | 21 | playable archives: local media attached to held videos (clips cut locally), per-video torrents played in the page, a home seeder of last resort behind a VPN; pilot TISM on jeralyzer-private | [`plans/release-21.md`](plans/release-21.md) | planned 2026-10-09 | | 22 | one article, two shapes: slides from the same report.json (authoring fields + a derived default), a reader switch Article / Slides / Overview, the isometric overview pairing sections with slides, `slides.html`/`slides.pdf` exports | [`plans/release-22.md`](plans/release-22.md) | planned 2026-10-09 | Work that landed on `main` without a plan of its own: [`plans/landed-2026-10.md`](plans/landed-2026-10.md).