Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 483631920159b2cd8c3725ef8f07eb799f2b9afe
parent dd7f18d2c5bde8350eb76df2a77f1cf5253a6c31
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sun, 26 Jul 2026 01:21:24 -0400

Add the local-AI derived-corpus roadmap and its memory system

Documentation only; no source changes.

The work of adding a transformative layer over the verbatim archive —
timestamped chapters, topic tags, attribution, a graduated visibility
policy — spans many phases and months, with context cleared between
them. That only works if a cold session can resume without re-deriving
what was already established, so this lands a roadmap plus a small
memory system rather than a single plan document.

  PLAN.md         the durable roadmap: why, architecture, phases 0-11
  plans/FACTS.md  verified codebase facts, every claim file:line anchored
  plans/STATE.md  running status, decisions log, open questions
  plans/README.md the context-clear protocol

FACTS.md carries the most weight. An earlier draft of this plan named
windowCues() as the transcript chunker; it is a center-based search-hit
helper and cannot do that. Several other confident claims were wrong the
same way — the LMDB maxDbs bump that isn't needed, a "closed" union with
a duplicated literal the compiler won't flag, a per-key concurrency map
that doesn't exist. Those corrections open the ledger.

Two engine lanes on separate queue keys let a backfill drive the local
GPU and a network model concurrently, and per-section provenance records
which model produced what, end to end.

AGENTS.md gains a four-line Roadmap pointer. It loads on every turn, so
a cleared-context session finds this without being told.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Diffstat:
MAGENTS.md | 7+++++++
APLAN.md | 543+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aplans/FACTS.md | 326+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aplans/README.md | 45+++++++++++++++++++++++++++++++++++++++++++++
Aplans/STATE.md | 112+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
5 files changed, 1033 insertions(+), 0 deletions(-)

diff --git a/AGENTS.md b/AGENTS.md @@ -10,3 +10,10 @@ To run more than one checkout at once (parallel dev servers / e2e), use git work per-worktree non-colliding ports. `pnpm wt add <branch>` creates one; `pnpm wt list` shows each worktree's port block. `pnpm dev:editor` / `pnpm e2e` auto-assign ports per worktree. See [WORKTREES.md](WORKTREES.md) for the port scheme and the shared-data caveat. + +# Roadmap + +Long-running work on the local-AI derived corpus is tracked in [PLAN.md](PLAN.md). +Before doing any phase work, read `plans/STATE.md` (current status + decisions) and +`plans/FACTS.md` (verified codebase facts — trust these over re-deriving them). +The phase in flight has a detailed plan at `plans/phase-N-*.md`. diff --git a/PLAN.md b/PLAN.md @@ -0,0 +1,543 @@ +# Roadmap: the local-AI transformative layer + +Long-running work adding a derived, transformative layer over the verbatim transcript +archive: timestamped chapters, topic tags, speaker attribution, a graduated visibility +policy, and eventually a site that leads with the derived corpus rather than raw +reproduction. + +This file is the **durable roadmap** and changes rarely. Three companions: + +- [`plans/FACTS.md`](plans/FACTS.md) — verified codebase facts. Trust it over re-deriving. +- [`plans/STATE.md`](plans/STATE.md) — current status, decisions log, open questions. +- `plans/phase-N-*.md` — file-level detail for the phase in flight, written when it starts. + +See [`plans/README.md`](plans/README.md) for the context-clear protocol. + +## Why + +Two goals drive the work. + +**Utility.** A timestamped summary per video lets a reader navigate a 3-hour livestream +without scrubbing, and gives search something to match beyond raw speech. + +**Copyright posture.** The archive is more defensible when it publishes new expression +*about* works alongside — and eventually instead of — full reproduction of them. Chapters, +topic tags, and analysis are genuinely transformative; a graduated per-channel visibility +policy replaces all-or-nothing takedown handling; and filtering out embedded third-party +clips reduces reproduction of works the archive has the weakest claim to. + +## Non-goal: transcripts are never rewritten + +Verbatim accuracy is what makes search, `[n @ mm:ss]` citations, and the MCP corpus sweeps +trustworthy. A full-length paraphrase would still be a derivative work while adding +hallucination risk. Transformative value is layered **on top of** verbatim cues, and +reduction of exposure happens by **withholding** cues, not by altering them. + +## Architecture + +One derived sidecar per video directory, mirroring the `transcript.cues.json` precedent: +versioned, self-describing, written tmp+rename, with an mtime freshness check. + +``` +transcripts/channels/<slug>/data/<videoId>/ + transcript.cues.json # verbatim source, never rewritten + ai-digest.json # chapters + tags + per-section provenance + diarization.json # phase 9 + attribution.json # phase 9 +``` + +**Never name a sidecar `transcript.<x>.<y>`** — `SUB_FILE_RE` claims any such file as a +subtitle track. Likewise **never use the name `summaries`** for the new artifact; that is +already the cross-channel metadata browse index. Both hazards are detailed in `FACTS.md`. + +```mermaid +flowchart LR + cues["transcript.cues.json<br/>(verbatim, immutable)"] --> md["transcriptToMarkdown<br/>stampForCue = hms"] + ctx["AI-CONTEXT.md → channel → video<br/>(phase 1.5)"] --> md + md --> chunk["chunkCuesForContext<br/>(NEW — windowCues cannot do this)"] + chunk --> eng{"digestApps registry"} + eng -->|"local-gpu lane"| oll["ollama-direct<br/>/api/chat + JSON schema"] + eng -->|"remote-api lane"| cc["claude-code<br/>claude -p --output-format json"] + oll --> parse["digestParse<br/>range · monotonic · snap · warnings"] + cc --> parse + parse --> dig["ai-digest.json<br/>+ per-section provenance"] + dig --> bi["buildIndex: digestMs → digests page tree"] + bi --> pub["/digests/&lt;slug&gt;/page-NNNN.json"] + pub --> ui["viewer ?vm=summary · search · MCP · corpus.json"] +``` + +### Engine registries, and why there are two lanes + +The repo's established idiom for "swappable engine" is a **registry of descriptors with a +client-safe listing function** — `TRANSCRIPTION_APPS` / `listTranscriptionApps()`. The +digest layer copies that idiom exactly rather than inventing a second pattern. Transcription +already has three engines; digests get two. + +| Lane | Engine | Queue key | Runs alongside | Cost | +| --- | --- | --- | --- | --- | +| `local-gpu` | `ollama-direct` (default) | `TRANSCRIPTION_QUEUE` | nothing — the 8 GB card is shared with whisper/parakeet | free | +| `remote-api` | `claude-code` | `DIGEST_REMOTE_QUEUE` | the GPU lane | metered | + +Two lanes on **different queue keys** is the structural idea that makes Claude Code useful +rather than merely available: a backfill can drive both at once, Ollama saturating the GPU +while Claude Code works the same backlog over the network. Putting both on one key would +serialize them and waste the network lane. Putting the local engine on its own key would let +Ollama and the transcription engine thrash the same 8 GB of VRAM. + +Metered engines are **opt-in** (`digest.remoteEnabled`, default `false`) and are never the +auto-queue default by inheritance. + +### Provenance — which AI was used + +Authoritative per **section**, because the digest controller does a read–modify–write merge +and a video can legitimately carry Ollama chapters and Claude Code tags: + +```jsonc +{ + "version": 1, + "contextHash": "sha1…", // phase 1.5 — invalidates on channel-context edits + "chapters": [{ "start": 0, "title": "COLD OPEN" }], + "tags": ["game review"], + "warnings": [{ "section": "chapters", "reason": "timestamp-out-of-range", "raw": "01:10:29 …" }], + "sections": { + "chapters": { "appId": "ollama-direct", "model": "qwen2.5:7b", "lane": "local-gpu", + "generatedAt": "2026-07-25T…", "promptVersion": 1 }, + "tags": { "appId": "claude-code", "model": "claude-opus-5", "lane": "remote-api", + "generatedAt": "2026-07-26T…", "promptVersion": 1 } + } +} +``` + +`engines: string[]` (the distinct `appId`s) is derived on read for cheap display and +filtering — do not persist it as a second source of truth. + +Provenance must reach every surface, not just disk: the per-video editor badge, a +`digestEngines: Record<string, number>` stat on `ChannelSnapshot` (a sibling of `totals`, +**not** a `buckets` entry — buckets are work-lanes of video ids), the digest record shipped +to the export, and a viewer badge. The Phase 10 honesty requirement extends to *which* +model, not merely *that* it was generated. + +## Phases + +```mermaid +flowchart TD + P0["0 · Benchmark engines"] --> P1["1 · Generation harness"] + P1 --> P15["1.5 · Channel context"] + P15 --> P2["2 · Digest corpus in build"] + P1 --> P25["2.5 · Observability"] + P25 --> BF(["BACKFILL"]) + P15 --> BF + P11a["11a · Review queue"] --> BF + P2 --> P3["3 · Viewer ?vm=summary"] + P2 --> P4["4 · Search indexing"] + P25 --> P5["5 · Auto-queue"] + P3 --> P10["10 · Lead with the corpus"] + P4 --> P10 + P8["8 · Visibility policy"] --> P10 + P2 --> P9["9 · Attribution + quote filter"] + P6["6 · Ollama /ask provider<br/>(no dependencies)"] + P1 --> P7["7 · Tags + chat highlights"] + P10 --> P11b["11b · Viewer feedback"] +``` + +### Phase 0 — Benchmark the transcription engines + +Small, and it belongs first because **transcription is the upstream bottleneck for the +digest backfill**: a video cannot be digested until it has a transcript. + +`parakeet` is **already** a fully selectable engine — the third entry in +`TRANSCRIPTION_APPS`, exposed in settings through `listTranscriptionApps()`, with a +`device` field and `supportsPartialStop: true` (it stops after the current window and +stitches a partial transcript on SIGTERM, so a long run is interruptible — directly +relevant to a multi-day sweep). The default is `whisper-cpp`. Nothing is missing +structurally. + +What *is* missing is evidence. Nothing in the repo measures relative throughput on this +hardware, so "if parakeet is faster" is currently unanswerable. + +- Time `whisper-cpp`, `chough`, and `parakeet` over the same handful of real videos of + varying length, on this GPU. Record wall-clock per audio-minute, VRAM, and transcript + quality spot-checks. +- Write the numbers into `plans/FACTS.md`. This is a fact that will be re-asked for months + and should never need re-measuring. +- If parakeet wins, change `DEFAULT_TRANSCRIPTION_APP_ID` and note it in `STATE.md`. + Because workers already carry a per-worker `appId`, a mixed fleet is also possible + without new machinery. +- Deliverable is measurements and possibly a one-line default change — **not** a new + benchmark harness. Do not build infrastructure for a question asked once. + +### Phase 1 — Generation harness + +Modelled on the transcription-app registry, **not** on `askProvider.ts` — that runs in the +visitor's browser in the static export and cannot spawn a process. + +- `common/lib/paths.ts` — add `ollamaUrl` (`OLLAMA_URL`) and `claudeBin` (`CLAUDE_BIN`), + following the `process.env.X ?? default` idiom. Add both to the hand-curated `pathRows` + in the settings page's System paths table. +- `common/lib/digestApps.ts` *(new)* — mirror `transcriptionApps.ts`: a `DigestApp` type, a + `DIGEST_APPS` record, a `getDigestApp(id)` that never throws, and a client-safe + `listDigestApps()` descriptor export so the registry never reaches the client bundle. + Each app carries `lane` and `metered`. + - **`ollama-direct`** — POST `${ollamaUrl}/api/chat` with `format: <JSON schema>`, + `options: { num_ctx: 16384, temperature: 0 }`, `stream: false`. **`num_ctx` must be + explicit**: the 4096 default silently truncates, the single easiest way to get + quietly-wrong output at scale. + - **`claude-code`** — shell out to `claudeBin -p <prompt> --output-format json`, parse the + wrapper's `result` string, then the JSON body. No context-window concern, but keep the + same chunking so output shape stays engine-independent. + - **`fabric`** — deferred. Measured to ignore the `HH:MM:SS TOPIC` contract at 7B, twice, + at two input sizes. Its 256-pattern library stays valuable as prompt source material; + it becomes a third registry entry when a prose-shaped output needs it, with no rework. +- `common/lib/transcriptWindow.ts` — add `chunkCuesForContext(cues, { maxCues, overlapCues })` + returning sequential overlapping slices. **`windowCues()` cannot do this** — it is a + center-based helper for search hits (see `FACTS.md`). +- `common/lib/digestParse.ts` *(new)* — parse engine output into `{ start, title }[]`. Must + reject timestamps beyond the video duration, **clamp each chunk's output to that chunk's + own range**, enforce monotonic starts, drop malformed entries, snap starts to the nearest + cue boundary, and de-duplicate across chunk seams. Every rejection is **recorded in a + `warnings` array, never silently dropped** — Phase 11's review queue is built entirely on + this. Unit-tested; it is the layer absorbing local-model sloppiness. +- `common/controller/digestVideo.ts` *(new)* — freshness check → `readNormalizedTranscript` + → `transcriptToMarkdown` with `stampForCue: (_c, s) => hms(s)` → chunk → engine → parse → + read-modify-write merge → tmp+rename. Reachability probe for the Ollama lane; a friendly + ENOENT message for the Claude lane. +- `common/controller/digestBatch.ts` *(new)* — `runPool()`, honoring `signal` and + `ctx.drainSignal`. Logs a running count of **metered** calls so a remote backfill's cost + is visible in the job log as it happens. +- `common/lib/queueKeys.ts` — `DIGEST_REMOTE_QUEUE` plus `digestQueueKey(appId)` mapping + lane → key. No concurrency declaration; the registry hardcodes 1 per key. +- One `jobKinds.ts` entry (`digest-channel`), one server action modelled on + `whisperActions.ts`, one replay handler. +- Settings: `defaultDigest()` + `sanitizeDigest()` wired into `getSettings()`'s normalize + chain, with per-engine config keyed by `appId` and `remoteEnabled: false`. +- UI: a channel stage component, an engine picker, a per-video button, and a **provenance + badge** (appId + model + generated-at) on the video page. +- `common/package.json` — add a `test` script; there is currently none, and ~46 unit test + files are unrunnable as a suite. + +### Phase 1.5 — Channel context layers + +Cascading human-authored notes injected into every prompt, modelled on this repo's own +`CLAUDE.md` → `AGENTS.md`. Lands early because every later phase improves from it, and +Phase 9's attribution is close to unusable without it. + +``` +transcripts/AI-CONTEXT.md # corpus-wide +transcripts/channels/<slug>/context.md # the main one +transcripts/channels/<slug>/data/<id>/context.md # rare, for oddities +``` + +Markdown with optional frontmatter: `hosts`, `recurring_guests`, +`plays_third_party_media`, `boilerplate`. + +- `common/lib/aiContext.ts` *(new)* — `resolveAiContext(paths, slug, videoId?)` follows the + `resolveCookiePolicy` inheritance shape. Caps total tokens (~1500) so context never crowds + out transcript. **There is no YAML dependency in the repo** — hand-roll a minimal + scalar/list frontmatter parser rather than adding one for three field types. +- `boilerplate` earns its keep with no model involvement: a deterministic filter dropping + recurring sponsor/subscribe cues from digest input and chapter candidates. Cheap, exact, + and it removes the most common junk chapter. +- `plays_third_party_media: false` lets Phase 9 skip attribution entirely for that channel — + a large accuracy and compute win. +- Editing UI on the channel and settings pages (tmp+rename). Frontmatter errors surface + inline, never at job time. +- A `propose-channel-context` job samples N transcripts and drafts `context.suggested.md`. + It **never** writes `context.md` — auto-applying model-authored context would let one bad + inference silently degrade every downstream summary for a channel, invisibly, because the + output would still look plausible. +- Hash the resolved context into `ai-digest.json.contextHash` so editing notes correctly + marks digests stale. **This is why backfill must not start before 1.5 exists.** + +### Phase 2 — Digests as a first-class corpus + +An earlier draft inlined the digest into `TranscriptDetail`. That is wrong: under +`summary-only` visibility a transcript page ships **no cues**, so a digest must never be +bundled inside a payload that policy may withhold. Digests get their own page tree. + +- `buildIndex.ts` — three new LMDB sub-DBs (`digests`, `digestPageHashes`, + `channelDigestStats`); **no `maxDbs` change needed**, but three new `clearAsync()` lines in + the hand-enumerated schema-invalidation block; `SCHEMA_VERSION` 12 → 13; `digestMs` (later + `attributionMs`) added to `MtimeRecord`, `LiveEntry`, the `scanSource()` stat loop, the + changed-detection comparison, and the `mtimes.put` call; emit + `.export-index/shared/digests/<slug>/{manifest,page-NNNN}.json` via `createPageWriter()`. + Each record carries its `sections` provenance. +- `common/lib/digests.ts` *(new)* — `VideoDigest` / `DigestPage` / `ChannelDigestsManifest` + with `slugToPage`, mirroring `manifest.ts`. +- `channelSignature.ts` — add `digestMs` to its three-field structural-subset `MtimeRecord` + **and** to the hash input, or archives will not rebuild on a digest-only change. +- `compose-site.ts` — a `digests?` key in `ComposeCache` and a fourth + `reconcileChannelTree()` call. It already passes `ignoreBasename: "manifest.json"` + internally; nothing to add there. +- `common/components/digestCache.ts` + `digestStore.ts` *(new)* — copy the + `transcriptCache.ts` + `transcriptStore.ts` pair (**not** `subsCache.ts`, which has no + IndexedDB). Use a **separate IDB store** so the transcript store's `DB_VERSION` need not + bump and existing transcript caches survive. +- `corpus.ts` — `CORPUS_SPEC_VERSION` 2 → 3, a `DIGEST_SCHEME` modelled on `POST_SCHEME`, + and a digest manifest pointer on `CorpusChannel.manifests`. Update `corpus.test.ts`. + +### Phase 2.5 — Observability for the backfill + +Build this **before** backfilling. A multi-day sweep you cannot observe is one you cannot +tune or safely interrupt. + +- **First**, fix the duplicated literal union in `RunningJobsList.tsx` to import + `JobProgressMetric`. Until then the compiler-driven audit the rest of this phase relies on + has a silent hole. Six sites total — see `FACTS.md`. +- Extend `JobProgressMetric` with `"digests"` and `JobTaskKind` with `"digest"`. +- The two copies of the `metric === "downloads" ? … : …` binary in `MonitorWidget.tsx` become + **one lookup table keyed by metric** (label, glyph, color), rather than growing a third + branch — otherwise the next metric repeats the bug. +- **Coverage, not just progress.** Per-job bars answer "how's this job"; during backfill the + question is "how much of the corpus is done" — `digested / total`, **split by engine**. + That is the single most useful number during a multi-day sweep, and with two lanes running + it is also how you see whether the remote lane is pulling its weight. +- Widget sync payload gains coverage **scalars only** — honor the file's stated design + constraint that it stays a handful of scalars rather than shipping per-channel breakdowns + on every poll. Its builder is already reused for dashboard SSR seeding, so both surfaces + get it free. +- A digest instrument in `PipelineBand`, `noDigest` counts in `NeedsWorkPanel` / + `ChannelsTable`, command-palette actions, and ETA via the existing `computeEtaSeconds`. +- `channelSnapshot.ts` gains `noDigest` **here**, not in Phase 5; Phase 5 then consumes what + already exists. + +### Phase 3 — Viewer + +- `urlState.ts` — add `"summary"` to `ModalMode`, the parse chain, and the `writeUrlParams` + `vm` branch. `ModalMode` appears in only three files. +- `PlayerProvider.tsx` — lazy-fetch the digest when `modalMode === "summary"`, mirroring the + chat effect including its ref-based in-flight guard (there is a comment explaining why + reducer state in the dep array drops results) and its snap-back-to-transcript on missing. +- `TranscriptModal.tsx` — a toolbar `ControlButton` beside the transcript/chat toggle; + chapters as buttons calling `seekTo(start)` with `scrollKindRef.current = "smooth"`; + active-chapter highlight off `currentTime`; a provenance line; an empty state. +- **Styling:** this modal and `PlayerProvider` are **not** on semantic theme tokens — they + use hard-coded `bg-zinc-900/80`, `ring-white/10`, `text-white` throughout. Match that + local palette; do not introduce `bg-card`/`text-foreground` here. Migrating the modal + chrome is out of scope. + +### Phase 4 — Search indexing + +Fold chapter titles and tags into the search corpus with a `source: "ai"` marker and the +producing `appId`. **Two independent index paths exist** and both need it: the client-side +FlexSearch worker and the server-side MCP search. Verbatim and generated hits must stay +visually distinguishable — that honesty requirement is what makes indexing generated text +acceptable at all. + +### Phase 5 — Auto-queue + +A third `AutoQueuePolicy` alongside `transcription`/`download`, a dispatch branch in +`autoRunner.ts` re-reading settings per iteration so a pause flag takes effect live, exposure +in the policy tree editor, a `digestsPaused` flag, and a pause button copied from the +downloads one. **The policy carries an explicit `engineId` defaulting to the local lane** — a +metered engine must never become the auto-queue default by inheritance. + +### Phase 6 — Ollama provider for `/ask` + +Zero dependencies on any other phase; a good early win. Widen the `Provider` union, add a +`PROVIDERS` entry (no key required), a `switch` arm, and an `askOllama` using Ollama's +OpenAI-compatible `/v1/chat/completions` SSE shape — the existing `askOpenAI` is a +near-template. Make the base URL editable in the provider settings panel. + +**Caveat to surface in the UI:** the export is static and runs in the visitor's browser, so +this only works when the visitor can reach an Ollama instance, and requires `OLLAMA_ORIGINS` +for CORS. Local/LAN use only. + +Knock-on benefit: the MCP corpus sweep's extraction step can then run locally. + +### Phase 7 — Tags and live-chat highlights + +Tags come free with Phase 1 into the same `ai-digest.json`. Live-chat highlights reuse the +whole harness against `live_chat.cues.json`, writing `ai-chat-digest.json` and rendering in +the existing `?vm=chat` view. + +### Phase 8 — Transcript visibility policy + +Mandatory, not optional — leading with derived data means the published site's default +posture is set here. + +- `common/lib/visibilityPolicy.ts` *(new)* — copy `cookiePolicy.ts` exactly: + `"full" | "excerpt" | "summary-only"`, a values array, a default, a type guard, and + `resolveVisibility(settings, channelConfig)`. Note cookiePolicy uses **two** inheritance + rules — truthiness-after-trim for free-text values, type-guard for enums; this is the enum + case. +- Wire through settings, `parseChannelConfig`, both forms — and **list the key in + `CHANNEL_FORM_FIELDS`** or clearing back to inherit silently won't work. +- **Enforced at build time** in the `buildIndex.ts` page writer, the single place that + decides what `cues` array ships. One enforcement point covers viewer, MCP, `/ask`, and + search, because all four read the same shards. +- **Archives bypass the page writer** — `archiveTranscripts.ts` hard-links + `transcript.cues.json` straight from the source tree. Extend its existing `requiresRewrite` + predicate and `writeTransformedCues` path. This is the highest-exposure surface and + inherits nothing for free. + +### Phase 9 — Attribution (two lanes) + quote filtering + +Largest and least certain. Ship 1–8 first; keep it off by default. + +**Blocking constraint:** audio is deleted once a video is transcribed, except for dirs +holding `do-not-clean.json` or entries in the saved-video store. **Most of the existing +archive has no audio to diarize.** + +Two first-class lanes producing the same artifact at different quality tiers: + +| Lane | When | Input | Recorded as | +| --- | --- | --- | --- | +| Diarization-assisted | Going forward; on-demand upgrade | audio → speaker turns → LLM labels the turns | `method: "diarized"` | +| Text-only | Legacy videos with no audio | cues alone → LLM segments and labels from content | `method: "text-only"` | + +Neither is a fallback for the other in code. Going forward, run diarization right after +transcription while audio is still on disk and **before** the cleanup sweep. For legacy, run +the text-only pass over the whole backlog immediately, then upgrade selectively. + +- `scripts/diarize.mjs` behind a `DIARIZE_BIN` path, exactly as the parakeet wrapper works. + Keeps the engine swappable: evaluate `sherpa-onnx` (CPU-friendly, no HF token) against + `pyannote` (better, needs token + GPU) **on real audio before committing** — quality here + cannot be judged from code. +- `attributeCues.ts` writes `attribution.json` with ranges, confidence, method, and engine + provenance. +- Filter in the page writer alongside Phase 8. **Bias toward dropping on uncertainty** — the + cost is asymmetric. Make the threshold policy-driven so a channel can filter on `diarized` + only and ignore `text-only` labels it doesn't trust. +- Status everywhere: a client-safe `attributionStatus.ts` + (`"none" | "text-only" | "diarized" | "stale"`), new snapshot buckets, a per-video badge, + per-channel counts. Mark a video **ineligible** when availability says deleted/private and + no audio is retained — a distinct state, not an upgrade button that can only fail. +- Re-processing: per-video re-attribute / upgrade, per-channel bulk over the text-only + bucket. The upgrade job re-downloads audio, diarizes, attributes, then removes the + re-fetched audio **in a `finally`** unless `do-not-clean.json` is present — this job can + pull gigabytes and a cancelled run must not silently fill the disk. +- `attributionMs` threads through the build the same way `digestMs` does. +- **Set expectations honestly:** this misfires on rapid back-and-forth, and auto-caption + channels have no speaker turns at all with cue boundaries that don't align to them. A + good-faith reduction in reproduction, not a guarantee; the UI must not claim otherwise. + +### Phase 10 — Leading with the derived corpus + +Depends on 2, 3, 4, 8. + +- **MCP** — a `get_digest` tool and digest-aware search, plus digest members on the + `ShardSource` interface implemented in **all three** classes (local, remote, hub). The hub + path needs it too, or federated sites silently lack digests. +- **Archives** — a digests zip. Remember the 25 MB default cap drops *all* archives for large + sites. +- **UI inversion** — search results lead with chapters and tags, verbatim cue matches + secondary and badged; listings surface chapter counts and topic tags; the modal defaults to + `?vm=summary` for `summary-only` channels (a default-selection change — the mode is already + URL-driven). +- **Graceful fallback** where a digest is absent. A partially-digested corpus is the normal + state for a long time and must not look broken. +- **Honesty:** generated content visually distinct from verbatim everywhere, labelled with + the producing model. This matters more once generated text is the primary thing users see. + +### Phase 11 — Human review queue & viewer feedback + +Without this, a corpus-wide backfill is a one-shot gamble on prompt quality. + +**Naming hazard:** `report` already means three different things in this repo. Use +**`feedback`** for viewer-submitted items and **`review`** for triage state. Do not add a +fourth meaning of "report". + +**11a — review queue (land before backfilling).** Extend the existing `/actionable` page +rather than building a parallel one: its `SectionConfig` array with counters in +`loadActionable.ts` is the designed extension point. New sections: digests needing review +(driven by the `warnings` array Phase 1 persists), proposed channel context awaiting +promotion, uncertain attribution, viewer feedback. + +Actions per item: approve · edit inline · regenerate · dismiss · **"add note to channel +context"**. The last is the one that compounds — a correction applied to one video fixes one +video; the same correction written into `context.md` fixes every future generation for that +channel. **Regenerate should offer the other engine**: "this looks wrong, redo it on Claude +Code" is the most natural use of a second lane, and the provenance record makes it obvious +when one engine is systematically weaker on a given channel. + +**11b — viewer feedback (can follow the corpus going public).** A "flag this" control per +chapter, with categories including **"misrepresents what was said"**. Transport in preference +order: an optional per-site `feedbackUrl` POST, else copy-to-clipboard / download-JSON. The +existing `r2-proxy/` package is a plausible minimal collector — evaluate it before standing +up new infrastructure. Ingestion is an editor action accepting pasted JSON. Treat submissions +as **untrusted input rendered in an admin UI**: escape it, never feed it into a prompt +unreviewed. Store under `transcripts/.feedback/`, a sibling of `.jobs` and `.bookmarks`, +outside the build trees. + +**Why that category is not boilerplate:** the smoke test summarized allegation-heavy content +about named individuals and restated those allegations as plain fact. A published AI summary +that misstates what a real person said or did is the highest-risk output this system can +produce, and the one failure mode no technical guard catches. Wire `dismiss` so it can +**suppress a digest from the next build**, not merely flag it for later. + +## Verification strategy + +Per the repo's `verification-playwright-first` convention: verify with Playwright, open a +browser only to diagnose failures. + +1. **Stub Ollama** — the default engine is an HTTP call from the Next server, so the + fake-binary trick does not cover it. Add an `ollama-stub.mjs` fixture launched from + `playwright.config.ts` alongside the existing web servers, with `OLLAMA_URL` pointed at + it. Give it a **deterministic bad-output mode** (out-of-range, non-monotonic) so the + parser guards are exercised by tests rather than by luck in production. +2. **Fake `claude`** — a `fake-claude.mjs` echoing the `--output-format json` wrapper, plus + the `SLOWOP` paced-output trick so drain/progress specs can observe it. Wire `CLAUDE_BIN` + into **both** the `dev:test` and `start:test` lines; they are duplicated verbatim. +3. **Unit** — `digestParse.test.ts` and the chunker cases, using `node:test`, runnable via + the new `common` test script. +4. **Editor e2e** — queue a job, assert the row label, assert `ai-digest.json` lands with the + right `sections.chapters.appId`. Post-mutation assertions must **poll-with-reload**; the + channel page serves a snapshot regenerated on a ~1 s debounce. +5. **Two-lane e2e** — queue both engines and assert they occupy **different queue keys** and + run concurrently. This is the behavior that makes the Claude Code lane worth having, so it + needs a test, not a comment. +6. **Export e2e** — clone the route-stubbing scaffold in + `export-player-platform-cache.spec.ts`: stub manifest/pages with a digest-bearing fixture, + open `?v=<slug>&vm=summary`, assert chapters render, clicking one seeks, and the engine + badge shows. +7. **Glance surfaces** — assert the sync payload carries coverage split by engine and the + pipeline band shows the digest chip and paused state. Run widget assertions under + `E2E_MODE=start` (see the Dev Tools caveat in `FACTS.md`). +8. **Full suite** — `pnpm e2e` in **default dev mode**; `E2E_MODE=start` serves a stale build. + Kill stale dev servers by port between runs. The known-failing-on-base list is in + `FACTS.md` — do not chase those as regressions. +9. **Real end-to-end** — one real channel against live Ollama, then `pnpm build:index && + pnpm --filter export run build`; confirm the digest survives into + `export/public/digests/<slug>/page-*.json`. For Phase 9, confirm a `summary-only` channel + ships **zero** cues by inspecting the shard directly, not the UI. +10. **Changelogs** — `editor/CHANGELOG.md` and `export/CHANGELOG.md` per repo convention. + +## Sequencing + +| Group | Why they group | +| --- | --- | +| 0 | Measurement only. Answers the transcription-speed question once, permanently. | +| 1, 1.5, 2, 2.5, 3 | The shippable core: generate → context → build → observe → display. | +| 4, 5, 6, 7 | Each independently shippable. **6 has no dependencies** — good early win. | +| 8 | Self-contained; prerequisite for 10. | +| 9 | Largest and least certain. Off by default. | +| 10 | Depends on 2, 3, 4, 8. | +| 11a / 11b | Review queue before backfill; viewer feedback after the corpus is public. | + +### Ordering traps + +- **Do not backfill before 1.5.** Channel context feeds `contextHash`, which correctly + invalidates every digest generated without it. At 15–60 s per video that is a costly redo. +- **Do not backfill before 2.5 and 11a.** An unobservable multi-day sweep cannot be tuned or + safely interrupted, and one with no review path is a one-shot gamble on prompt quality. +- **Benchmark before committing to backfill.** The smoke test measured 15.5 s for a ~2.7k + token chunk; a 176-minute podcast needs roughly a dozen chunks. Budget minutes per long + video and multiply by corpus size. The remote lane changes this arithmetic — measure both, + then split the backlog between them. +- **Fix the duplicated metric union before extending `JobProgressMetric`**, or Phase 2.5's + compiler-driven audit silently misses a consumer. + +### Hardware and prerequisites + +Radeon RX 6600, 8 GB VRAM, 15 GB system RAM. This is the binding constraint on model choice: +target a 7–8B at Q4 (~5 GB), leaving ~3 GB for KV cache. A 14B at Q4 (~9 GB) spills to CPU +and makes corpus-wide generation impractical. `ollama-vulkan` 0.32.4 verified with `100% GPU` +offload. The systemd unit ships **inactive and disabled**, with no `~/.ollama` and no models: + +```bash +sudo systemctl enable --now ollama +ollama pull qwen2.5:7b # ~4.7 GB Q4; llama3.1:8b is the alternative +``` + +The Claude Code lane needs the `claude` CLI installed and authenticated on the host, plus +`digest.remoteEnabled` turned on in settings. diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -0,0 +1,326 @@ +# Verified codebase facts + +Facts established by reading the tree, each with a `file:line` anchor. **Trust these over +re-deriving them** — that is the entire point of this file. If one turns out to be stale, +correct it in place and note the correction in `STATE.md`. + +Append as you verify new things. Do not add anything here you have not actually checked. + +Last full verification pass: **2026-07-26**, against `main` @ `18c5a7a`. + +--- + +## Corrections to the original draft plan + +These were confidently stated and **wrong**. They are the most valuable entries here. + +### `windowCues()` is not a chunker + +`common/lib/transcriptWindow.ts:21` — +`windowCues(cues, centerSeconds, { before = 45, after = 45, maxCues = 60 })`. It is a +**center-based** window around a search hit, with `before`/`after` in **seconds**. It +cannot split a transcript into sequential context-sized chunks. Callers: +`export/app/ask/useAskChat.ts:1333`, `export/app/lib/searchAgent.ts:509,742`. + +→ A sequential `chunkCuesForContext()` must be written. Put it in the same file; it is pure +and `transcriptWindow.test.ts` already exists. + +### No LMDB `maxDbs` bump is needed + +`maxDbs` caps the named DBs opened *by one handle*, not the DBs present in the environment. + +- `common/controller/buildIndex.ts:365` — `maxDbs: 17`, opens **12** (`sums`, `cues`, + `subs`, `mtimes`, `byChannel`, `pageHashes`, `subPageHashes`, `channelStats`, `posts`, + `postPageHashes`, `channelPostsStats`, `meta`). Three more → 15, still under. +- `common/lib/channelSignature.ts:48` — `maxDbs: 14`, opens **2** (`mtimes`, `meta`). +- `common/controller/duplicateShorts.ts:111` — `maxDbs: 12`, opens **3**. + +None of these break. The "17 vs 14 inconsistency" is a red herring. + +→ What **is** required and easy to miss: the schema-invalidation block at +`buildIndex.ts:425-439` enumerates every sub-DB by hand with `clearAsync()`. Three new lines +there, plus `SCHEMA_VERSION` (`buildIndex.ts:116`, currently **12**) → 13. + +### `JobProgressMetric` is not a closed union in practice + +`editor/app/jobs/components/RunningJobsList.tsx:41` re-spells +`metric: "downloads" | "transcripts"` as a **literal**, not an import of +`JobProgressMetric`. TypeScript will therefore **not** flag it when the union is extended. + +Fix that line first, then extend. All six sites: + +| File | Line | +| --- | --- | +| `common/jobs/registry.ts` | 14 (the union), 22 (`JobTaskKind`) | +| `editor/app/jobs/components/RunningJobsList.tsx` | 41 (literal), 313, 326 | +| `editor/app/widget/components/MonitorWidget.tsx` | 615, 636 | +| `editor/app/jobs/active/buildActiveJobs.ts` | 66 | + +`MonitorWidget.tsx:615` and `:636` are two copies of the same +`metric === "downloads" ? … : …` binary, in `jobProgressText()` and `JobProgressBar()`. + +### Queue concurrency is not declared per key + +`common/lib/queueKeys.ts` has **no** concurrency map. `common/jobs/registry.ts:158` +hardcodes `concurrency: 1` for every non-empty `queueKey`. A `queueKey` of `""` means "run +immediately, unserialized". + +`TRANSCRIPTION_QUEUE` is defined at `common/lib/platform.ts:112` and merely re-exported from +`queueKeys.ts:2-8`. `resolveQueueKey(defaultKey, override)` is at `queueKeys.ts:27-32`: +`undefined` → default, `""` → immediate, else trimmed. + +### `resolveCitationSeconds` snaps to snippets, not cues + +`export/app/ask/citations.tsx:76` scans `src.snippets` — only the ~5 matched lines per +video (`export/app/lib/askRetrieval.ts:132`). It never sees a cue list, so it cannot be +reused for cue snapping. + +→ The right reference for a cue-boundary snap is `findActiveIndex` +(`common/components/TranscriptModal.tsx:473-491`), a hand-rolled binary search over cue +starts. + +### There is no vitest or jest + +Unit tests are `node:test` run via `tsx --test`. Only `mcp/package.json:12` has a `test` +script (`tsx --test src/*.test.ts`). `common/package.json` has **no `scripts` key at all**; +neither `editor` nor `export` has a `test` script. ~46 `*.test.ts` files across the repo are +unrunnable as a suite today. + +### `subsCache.ts` has no IndexedDB + +- `common/components/subsCache.ts` — memory + react-query only. +- `common/components/summariesCache.ts` — pure react-query, eager over all pages, no IDB, no + `slugToPage`. +- `common/components/transcriptCache.ts` — the lazy fetch/memo layer (manifest → + `slugToPage` → page), and it opportunistically warms every entry in a fetched page. +- `common/components/transcriptStore.ts` — the IDB layer. `DB_NAME = + "yt-dlp-transcript-browser"`, **`DB_VERSION = 4`** (`:15`), `STORE = "transcripts"`. Its + `onupgradeneeded` (`:43-47`) **deletes and recreates** the store, so a version bump wipes + the cache. + +→ Copy the `transcriptCache` + `transcriptStore` **pair**. +→ `editor/e2e/export-player-platform-cache.spec.ts:115` hardcodes DB version **3** and is +already stale relative to `DB_VERSION = 4`. + +### Social-posts incrementality cannot be copied for a per-video artifact + +`buildIndex.ts:233` — `scanSource()` does `if (isSocialChannel(cfg)) continue;`. Social +channels never enter the mtime scan at all, which is *why* posts use a per-channel +shard-signature (`buildIndex.ts:1088-1118`) instead of `MtimeRecord`. + +→ For a per-video artifact use the mtime mechanism (`digestMs` in `MtimeRecord`). Follow the +posts precedent (`buildIndex.ts:1065-1240`) only for the **page-tree emission shape**. + +### Archive gating already has a seam + +`common/controller/archiveTranscripts.ts:177` computes +`const requiresRewrite = !build.includeMetadata || build.prettyPrint;` and branches at +`:205-209` to `writeTransformedCues(...)` vs `link(...)` (a hard link straight from the +source tree). + +→ Phase 8 extends that predicate and that transform, rather than building a parallel gate. + +--- + +## Naming hazards + +**`summaries` is taken.** `exportSummariesDir` (`common/lib/paths.ts:58,151`) → +`/summaries/{manifest,page-NNNN}.json` is the cross-channel browse index of video *metadata* +(`DisplaySummary` records), documented in `common/lib/corpus.ts:35-36` and served by +`common/components/summariesCache.ts`. Nothing to do with AI summaries. Use **`digest`**. + +**Never name a sidecar `transcript.<x>.<y>`.** `common/lib/videoStatus.ts:137` — +`const SUB_FILE_RE = /^transcript\.([^.]+)\.([^.]+)$/;` treats any such file as a subtitle +track. The exclusion list in `readSubTracks` (`:238-256`) is hardcoded. Use +`ai-digest.json`, `diarization.json`, `attribution.json`. + +**`report` already means three things**: `reportDebouncePreset` in settings, the +`refresh-report` job kind, and `/ask`'s `ReportPanel`. Use `feedback` / `review` instead. + +--- + +## Reusable helpers (do not rewrite these) + +| What | Where | Notes | +| --- | --- | --- | +| Transcript → text for an AI | `common/lib/transcriptToMarkdown.ts:60` | Declared single source of truth. `stampForCue?: (clock, seconds) => string` at `:43` takes precedence over `linkForCue`. | +| Zero-padded `HH:MM:SS` | `common/lib/aiHandoff.ts:54` — `hms(s)` | So the stamper is `(_clock, s) => hms(s)`. No new formatter needed. | +| Read a normalized transcript | `common/controller/normalizeTranscript.ts:161` | `readNormalizedTranscript(cuesPath)`. | +| Freshness check | `normalizeTranscript.ts:193` | `isCuesJsonFresh(videoDir)`. | +| tmp+rename write idiom | `normalizeTranscript.ts:139-141` | `${path}.tmp-${process.pid}` then `rename`. | +| Engine registry + client-safe listing | `common/lib/transcriptionApps.ts:62-81, 248-278` | The pattern for `digestApps.ts`. | +| Inheritance resolve (global → channel) | `common/lib/cookiePolicy.ts:53-66` | Two rules: truthiness-after-trim for free text, type-guard for enums. | +| Child process into a job log | `common/jobs/runChild.ts:22` | `runChildIntoLog(onLog, signal, opts)`. | +| Bounded concurrent pool | `common/jobs/concurrentRunner.ts:29` | `runPool()`; honors `signal` + `drainSignal`. | +| Managed server action | `editor/app/channels/[slug]/whisperActions.ts:45-103` | The `runManagedFunction` template. | +| Paginated page writer | `common/controller/buildIndex.ts:694` | `createPageWriter()` — size budget + sha1 skip-if-unchanged. | +| Health probe | `common/controller/remoteTranscribe.ts:51` | `pingRemoteHealth(worker, timeoutMs)`. | +| ENOENT → friendly message | `common/social/xGalleryDlFetcher.ts:207-213` | `/ENOENT/.test(message) ? "X not found (set X_BIN)" : message`. | +| ETA estimator | `editor/app/jobs/active/buildActiveJobs.ts:38` | `computeEtaSeconds`. | +| Binary search over cue starts | `common/components/TranscriptModal.tsx:473-491` | `findActiveIndex`. | + +--- + +## Transcription engines (Phase 0) + +All three are **already selectable** — `common/lib/transcriptionApps.ts:248-252`: + +| id | Label | Line | Fields | Notes | +| --- | --- | --- | --- | --- | +| `whisper-cpp` | whisper.cpp (whisper-cli) | 157 | model, customArgs | `DEFAULT_TRANSCRIPTION_APP_ID` (`:254`) | +| `chough` | chough | 179 | model, remoteUrl, chunkSize | writes exactly the `-o` path | +| `parakeet` | parakeet.cpp (overlapping segments) | 216 | model, chunkSize, device | **`supportsPartialStop: true`** (`:222`) — stops after the current window and stitches a partial transcript on SIGTERM, so a long run is interruptible | + +Exposed to the settings form via `listTranscriptionApps()` (`:272`), consumed at +`editor/app/settings/page.tsx:45`. Per-worker `appId` means a mixed fleet already works +(`common/lib/settings.ts:837`). + +**Nothing measures relative speed.** Phase 0 fills that gap — record the numbers below. + +> _Benchmark results: not yet measured._ + +--- + +## Build pipeline constants + +| Constant | Where | Value | +| --- | --- | --- | +| `SCHEMA_VERSION` | `buildIndex.ts:116` | 12 | +| `CUES_FILE_VERSION` | `normalizeTranscript.ts:33` | 2 | +| `CORPUS_SPEC_VERSION` | `common/lib/corpus.ts:16` | 2 | +| `MANIFEST_VERSION` (site summaries) | `common/lib/manifest.ts:36` | 3 | +| `TRANSCRIPTS_MANIFEST_VERSION` | `manifest.ts:43` | 1 | +| `SUBS_MANIFEST_VERSION` | `manifest.ts:59` | 4 | +| `SUMMARIES_PAGE_SIZE` | `manifest.ts:37` | 1000 | +| `DEFAULT_ARCHIVE_MAX_BYTES` | `common/lib/archiveOptions.ts:72` | 25 MB — over it, **all** archives are dropped | + +`MtimeRecord` (`buildIndex.ts:146-154`): `metaMs`, `transcriptMs`, `subsMs`, +`availabilityMs`, `isDeleted`, `isUnlisted`, `indexKey`. +Changed-detection comparison at `:455-468`; the single `mtimes.put` at `:631-639`. + +`channelSignature.ts` uses a **narrower structural subset** (`:23-27`): only `metaMs`, +`transcriptMs`, `subsMs`, and its hash (`:65-83`) covers only those three. **A new per-video +artifact that does not move one of those three will not invalidate archive/compose caches.** + +`compose-site.ts`: `ComposeCache` at `:486-495`; `reconcileChannelTree()` at `:579`, called +three times (`:667` transcripts, `:675` subs, `:686` posts); it passes +`ignoreBasename: "manifest.json"` internally at `:597` because `generatedAt` churns. + +Transcript shards ship the **full `cues` array inline** — `buildIndex.ts:830-832`. + +--- + +## Editor surfaces + +| Surface | File | Note | +| --- | --- | --- | +| System paths table | `editor/app/settings/page.tsx:17-28` | Hand-curated tuple list, **not** `Object.entries(paths)`. The blurb at `:57-64` is already stale (missing parakeet/gallery-dl). | +| Job kinds | `common/jobs/jobKinds.ts:21-40, 44-255` | ~30 entries. **No entry sets `defaultTier`** — declared but unused so far. | +| Replay handlers | `editor/app/jobs/jobReplayRegistry.ts:75-236` | Bucket kinds re-derive ids from the live snapshot (`:50-57`), never a frozen list. | +| Snapshot buckets | `common/controller/channelSnapshot.ts:57-143` | 21 buckets, all `string[]` of video ids. Readers default `?.length ?? 0`. `undownloadedIds` is deliberately **outside** `buckets` (`:144`). | +| Actionable sections | `editor/app/actionable/page.tsx:30-46` | `SectionConfig` is module-private. Counters live in `actionable/lib/loadActionable.ts`. | +| Widget sync payload | `editor/app/api/widget/sync/route.ts:15-26` | Comment at `:10-14` states it deliberately stays "a handful of scalars". `buildWidgetSyncPayload()` (`:30`) is exported for SSR seeding. | +| Pipeline band | `editor/app/components/dashboard/PipelineBand.tsx:63-114` | Pure props, no fetching. `<Instrument dotClass=…>` encodes state color. | + +--- + +## Viewer surfaces + +`ModalMode` — `common/components/urlState.ts:13`, currently +`"transcript" | "chat" | "post"`. Parsed at `:63-65` (unrecognized → `"transcript"`), +written at `:126-130`. Referenced in **only three files**: `urlState.ts`, +`PlayerProvider.tsx`, `TranscriptModal.tsx`. + +Lazy chat fetch effect — `PlayerProvider.tsx:590-640`, deps `[urlSlug, modalMode]`, with a +ref-based in-flight guard and a comment at `:585-589` explaining why reducer state in the +dep array drops results. It snaps back via `writeUrlParams({ vm: "transcript" })` on missing. +`seekTo` at `:522-530`. + +**Styling:** `TranscriptModal.tsx` and `PlayerProvider.tsx` are **not** on semantic theme +tokens. Verified hard-coded: `bg-black/70` (`:199`), `bg-zinc-900/80 ring-white/10` (`:210`), +`text-white` (`:331`), `text-zinc-300` (`:336`), `border-zinc-700 bg-zinc-900` (`:357`), +`bg-blue-950/50` (`:419`). By contrast `export/app/ask/citations.tsx` uses semantic tokens +throughout — the two surfaces are on different systems. Do not mix them. + +`export/app/lib/askProvider.ts`: `Provider` union at `:19`, `PROVIDERS` at `:34-59`, +`askStream` switch at `:128-137`, `askOpenAI` at `:336-397`. Provider settings live at +`export/app/ask/ProviderSettings.tsx` (**not** `app/components/`); its `Props` is a 21-field +flat prop bag (`:13-36`). + +Two independent search index paths: `common/components/searchIndex.worker.ts` (FlexSearch, +client) and `mcp/src/search.ts` (server). MCP has 11 tools (`mcp/src/server.ts:132`) and a +14-member `ShardSource` interface (`mcp/src/source.ts:119`) implemented three times — +`LocalSource` (`:185`), `RemoteSource` (`:353`), `HubSource` (`:477`). + +--- + +## E2E environment + +- `editor/playwright.config.ts` — editor on `PORT ?? 3011`, export on `EXPORT_PORT ?? 3010`. + `E2E_MODE=start` → `pnpm start:test`, otherwise `pnpm dev:test` (`:10-11`). + `fullyParallel: false`, `workers: 1`, `timeout: 30_000`. +- `editor/e2e/helpers.ts` — `resetData` (`:32`), `writeSettings` (`:49`), `readJson` (`:90`). +- Fake binaries live in `editor/e2e/fixtures/bin/`. The **`SLOWOP` sentinel** (checked + case-insensitively against cwd or video id) makes an instant fake emit paced progress — + three independent copies at `fake-whisper.mjs:29-40`, `fake-chough.mjs:25`, + `fake-ytdlp.mjs:173-183`. +- Binary env vars are wired in `editor/package.json`'s `dev:test` **and** `start:test` + scripts — a single line duplicated verbatim, differing only in `next dev` vs `next start`. + **Adding one env var means editing both.** +- Route-stubbing scaffold: `editor/e2e/export-player-platform-cache.spec.ts:39-111`. + +### Known-failing on base — not regressions + +- The two `actionable.spec` "queues a job" tests +- The `settings.spec` social-links test +- `channels-actions.spec:42` +- The 4 raw-job-kind tests in bulk-actions / cleanup-actionable +- `widget.spec`'s "no buttons" assertion fails under `dev:test` because `next dev` injects a + Dev Tools button; it passes under `E2E_MODE=start` + +Run `pnpm e2e` in **default dev mode** — `E2E_MODE=start` serves a stale build. Kill stale +dev servers by port between runs. + +--- + +## Dependencies present / absent + +Present: `execa`, `lmdb`, `flexsearch`, `p-limit`, `fs-extra`, `@tanstack/react-query`, +`@tanstack/react-virtual`, `markdown-to-jsx`, `tsx`. + +**Absent — do not assume:** any YAML parser, `gray-matter`, `zod`, `vitest`, `jest`. +A frontmatter parser for Phase 1.5 must be hand-rolled (three scalar/list field types). + +--- + +## Smoke test results (2026-07-25) + +Run against a real 176-minute podcast transcript (`UndertheTeaVT/data/5JRFDQ7TtZ8`, 3877 +cues, ~34k tokens) on the RX 6600. + +1. **Ollama silently truncates.** `ollama ps` reported `CONTEXT 4096` — the default. A 10k + token excerpt was cut without warning and the model summarized whatever fragment + survived. **`num_ctx` must be set explicitly per request.** 16k is comfortable on this + GPU (qwen2.5:7b KV cache ≈ 56 KB/token → ~0.9 GB at 16k, atop 4.7 GB of weights). +2. **fabric's stock pattern fails at this model size.** With `create_video_chapters` + verified applied (via `--dry-run`) *and* an input small enough to fit, qwen2.5:7b ignored + the `HH:MM:SS TOPIC` contract entirely and returned conversational prose. Twice, at two + input sizes. It is a long discursive prompt written for frontier models. +3. **Schema-constrained decoding fixes it.** POSTing to `/api/chat` with a JSON schema in + `format` returned clean parseable output in **15.5 s**. +4. **It still hallucinates timestamps.** That same run emitted `01:10:29` for an input + spanning only `00:04:45 → 00:15:36` — 55 minutes past the end. **The parser's range, + monotonicity, and cue-snapping guards are load-bearing, not defensive polish**, and + clamping must be per-chunk: a whole-video range check would have accepted that value. +5. **GPU offload works** — `100% GPU` on the RX 6600 via Vulkan. + +### Environment state as of that run + +`ollama-vulkan` 0.32.4 and `fabric-ai` 1.4.375 installed. The ollama systemd unit ships +**inactive and disabled**; no `~/.ollama`, no models. `~/.config/fabric` has no `patterns/` +and no `.env`, so `create_video_chapters` does not exist on disk until `fabric-ai --setup` +clones the pattern repo. + +Note: Arch names the binary **`fabric-ai`**, not `fabric`, to avoid colliding with the +unrelated Python `fabric` deployment tool. Moot while fabric is deferred, but relevant if it +is ever added as an engine. diff --git a/plans/README.md b/plans/README.md @@ -0,0 +1,45 @@ +# `plans/` — how this work remembers itself + +The local-AI derived-corpus work spans many phases and months, and **context is cleared +between phases** to keep token cost down. That only works if a cold agent can resume without +re-deriving anything. Four artifacts with deliberately different lifetimes: + +| File | Lifetime | Contents | +| --- | --- | --- | +| [`../PLAN.md`](../PLAN.md) | Rarely changes | The durable roadmap: why, architecture, all phases, sequencing. | +| [`FACTS.md`](FACTS.md) | Append-only | Verified codebase facts with `file:line` anchors. | +| [`STATE.md`](STATE.md) | Rewritten each session | Phase status, decisions log, open questions, commands known to pass. | +| `phase-N-<slug>.md` | Written at phase start, archived at merge | File-level detail for the phase in flight. | + +`FACTS.md` is the most valuable file here. It holds exactly the material that is expensive to +establish and cheap to get wrong — the original draft of this plan contained several +confident, wrong claims about how the build pipeline works, and each one would have cost real +implementation time. The roadmap is stable prose; the ledger grows every time something is +checked against the tree. + +## Context-clear protocol + +**Before clearing context:** + +1. Update `STATE.md` — phase status, decisions made and *why*, anything surprising. +2. Append newly verified facts to `FACTS.md`, each with a `file:line` anchor. +3. **Commit.** An uncommitted memory file does not survive a container reclaim. + +**After clearing context:** + +`AGENTS.md` arrives automatically (`CLAUDE.md` is just `@AGENTS.md`), and it points here. +Read `STATE.md`, then `FACTS.md`, then the current `phase-N-*.md`. That is the whole warm-up +and it should cost a few thousand tokens rather than a re-exploration. + +**Never re-verify something already in `FACTS.md`** unless the surrounding code changed. If a +fact turns out to be stale, correct it in place and note the correction in `STATE.md`. + +## Writing a phase plan + +Start it when the phase starts, not before — a plan written three phases early is written +against a tree that no longer exists. Keep it to what `PLAN.md` deliberately omits: exact +files, exact edits, the order to make them in, and how to verify. Everything general belongs +in `PLAN.md`; everything verified belongs in `FACTS.md`. + +Archive rather than delete on merge — `git mv plans/phase-N-*.md plans/done/` — so the +reasoning behind a shipped phase stays findable. diff --git a/plans/STATE.md b/plans/STATE.md @@ -0,0 +1,112 @@ +# Running state + +The working memory for the local-AI derived-corpus work. Rewritten at the end of every +session, before context is cleared. See [`README.md`](README.md) for the protocol. + +**Last updated:** 2026-07-26 — roadmap written, no code yet. + +--- + +## Phase status + +| Phase | Status | Notes | +| --- | --- | --- | +| 0 · Benchmark transcription engines | not started | Measurement only. Parakeet is already selectable; what's missing is numbers. | +| 1 · Generation harness | not started | | +| 1.5 · Channel context | not started | **Blocks backfill** — `contextHash` invalidates digests made without it. | +| 2 · Digest corpus in build | not started | | +| 2.5 · Observability | not started | **Blocks backfill** — an unobservable sweep can't be tuned. | +| 3 · Viewer `?vm=summary` | not started | | +| 4 · Search indexing | not started | | +| 5 · Auto-queue | not started | | +| 6 · Ollama `/ask` provider | not started | **No dependencies.** Good first commit. | +| 7 · Tags + chat highlights | not started | | +| 8 · Visibility policy | not started | | +| 9 · Attribution + quote filtering | not started | Largest, least certain. Off by default. | +| 10 · Lead with the derived corpus | not started | | +| 11a · Review queue | not started | **Land before backfilling.** | +| 11b · Viewer feedback | not started | After the corpus is public. | + +**Recommended next:** Phase 6 (small, self-contained, zero dependencies) for an early win, +or Phase 0 → 1 to start the real spine. + +**Nothing has been implemented.** The only changes are `PLAN.md`, `plans/*`, and a +four-line `# Roadmap` block in `AGENTS.md`. + +--- + +## Decisions log + +Recorded with reasons, because these are exactly what a cold agent would otherwise +relitigate. + +**Fabric is deferred as a digest engine.** Measured to ignore the `HH:MM:SS TOPIC` output +contract at 7B — twice, at two input sizes, with the pattern verified applied via +`--dry-run` and an input small enough to fit the window. It is a long discursive prompt +written for frontier models. Its 256-pattern library remains valuable as *prompt source +material*, and it can become a third registry entry later with no rework. + +**Ollama-direct with JSON-schema `format` is the structured default.** Measured clean and +parseable in 15.5 s on the same input fabric failed. Nothing structured should depend on a +7B voluntarily honoring a text format — that is now measured, not assumed. + +**Claude Code is a second engine on a separate queue key.** Both lanes are GPU/network +disjoint, so a backfill can drive them concurrently — Ollama saturating the RX 6600 while +Claude Code works the same backlog over the network. One shared key would serialize them and +waste the network lane; giving the local engine its own key would let Ollama and the +transcription engine thrash the same 8 GB of VRAM. Metered engines are opt-in and never the +auto-queue default. + +**Digests get their own page tree** rather than being inlined into `TranscriptDetail`. +Under `summary-only` visibility a transcript page ships no cues, so a digest must never live +inside a payload that policy may withhold. + +**Provenance is per-section, not per-file.** The digest controller does a read–modify–write +merge, so a video can legitimately carry Ollama chapters and Claude Code tags. A top-level +`engines` list is derived on read, never persisted as a second source of truth. + +**Transcripts are never rewritten.** Verbatim accuracy is what makes search, citations, and +MCP sweeps trustworthy; a paraphrase would still be derivative while adding hallucination +risk. Exposure is reduced by *withholding* cues, not altering them. + +--- + +## Open questions + +- **Diarizer choice** — `sherpa-onnx` (ONNX, CPU-friendly, no HF token) vs `pyannote` + (higher quality, needs a HF token and GPU). Decide at Phase 9, on real audio. Quality here + cannot be assessed from code. +- **Phase 11b collector** — can the existing `r2-proxy/` package host the minimal feedback + POST endpoint? Worth evaluating before standing up new infrastructure. +- **Transcription engine default** — pending Phase 0 numbers. If parakeet wins, change + `DEFAULT_TRANSCRIPTION_APP_ID` and record it here. + +--- + +## Commands known to pass + +Established against `main` @ `18c5a7a`: + +```bash +pnpm e2e # default dev mode; E2E_MODE=start serves a stale build +pnpm build:index +pnpm --filter export run build +pnpm --filter yt-dlp-transcript-mcp run test +``` + +Kill stale dev servers by port between e2e runs. The known-failing-on-base list lives in +[`FACTS.md`](FACTS.md#known-failing-on-base--not-regressions) — do not chase those as +regressions. + +Note there is currently **no root or `common` test script**; ~46 unit test files are +unrunnable as a suite. Phase 1 adds one to `common/package.json`. + +--- + +## Surprises hit so far + +The original draft plan named `windowCues()` as the transcript chunker. It is a center-based +search-hit helper and cannot do that — the chunker has to be written. Several other +confidently-stated claims in that draft were also wrong; all are corrected at the top of +[`FACTS.md`](FACTS.md#corrections-to-the-original-draft-plan). Treat plan prose written +without a `file:line` anchor as unverified.