commit cf15ea46a152a45c0a96d718399bf906db011b5b
parent 348813a3afa0d7df4a1a2b75c8417d16939ed5e9
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Tue, 18 Aug 2026 20:59:59 -0400
report-to-video: adopt the pipeline, and stop it re-downloading its own cache
1,522 lines of working pipeline that turns a cited sweep report into a finished
video, untracked in git while six real reports already depend on it. Committed
as-is, then given the three things umtool needs to drive it.
Containing-window cache reuse. A raw clip's window is in its name, which makes
the cache content-addressed -- but the lookup was for the EXACT name, so every
window edit re-downloaded material already on disk. ferret-rescue/out/clips-raw
holds 31 files for 10 clips; one source is there four times over overlapping
windows. Now any cached file containing the request satisfies it, and the
TIGHTEST container wins: silence detection decodes the whole file, so a 40s file
costs more than the 24s one that would also have done. Verified by nudging c01's
start 0.5s -- previously a fresh yt-dlp run, now it reuses 40.12-64.48 and
touches the network not at all.
--progress ndjson, so umtool's driver reads events instead of scraping prose.
The event set is exactly what was already being printed; making it a format
switch is what stops a wording change from breaking the driver.
--continue-on-error, because a dead source at clip 14 of 19 currently throws
away thirteen fetches already paid for. It finishes everything buildable and
then REFUSES to concatenate, exiting non-zero: a finished file that quietly
lost a citation looks complete, which is worse than no file at all.
Also: main() split out into an exported buildVideo(), --fetch-only <id> --pad <s>
so the clip bench's wide fetch runs the build's own fetch code (same naming, same
VP9 pin, same Rumble HLS retry), and check-availability.mjs -- yt-dlp --simulate
per distinct (channel, video), no bytes downloaded, classifying deleted/private/
restricted/no-cues. It belongs at the FRONT of a build: it is the one fact about
a manifest that goes stale in both directions.
resolve-windows.mjs still reports "0 window(s) changed" on all six real
manifests -- the fixed point holds.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Diffstat:
10 files changed, 2308 insertions(+), 0 deletions(-)
diff --git a/pnpm-workspace.yaml b/pnpm-workspace.yaml
@@ -4,6 +4,7 @@ packages:
- export
- homepage
- mcp
+ - scripts/report-to-video
- umtool
allowBuilds:
diff --git a/scripts/report-to-video/README.md b/scripts/report-to-video/README.md
@@ -0,0 +1,324 @@
+# report-to-video
+
+Turns a cited sweep report into a video: the clips run in chronological order and
+let the source speak for itself, with thin chrome carrying the citation and a
+timeline of where you are. A video rendering of the reports we already write.
+
+Two scripts and a manifest:
+
+| file | lifetime | what it is |
+|---|---|---|
+| `build-video.mjs` | stable | manifest → mp4. Fetches clips, snaps cuts to silence, letterboxes them into the chrome, crossfades. |
+| `render-cards.mjs` | stable | draws the timeline footer and marker, plus optional card stills. Imported by `build-video.mjs`. |
+| `resolve-windows.mjs` | stable | widens clip windows from cue spans to whole sentences. Run once after authoring a manifest. |
+| `<report>/video.manifest.json` | per report | the edit decision list. **This is the regeneration source of truth**, not the report. |
+
+```
+node scripts/report-to-video/resolve-windows.mjs ~/reports/<slug>/video.manifest.json --write
+node scripts/report-to-video/build-video.mjs ~/reports/<slug>/video.manifest.json
+```
+
+Output lands in `<report dir>/out/`: `cards/`, `clips-raw/`, `segments/`, and the
+finished `<slug>.mp4`. First worked example: `~/reports/ferret-rescue/`.
+
+## Why there is a manifest at all
+
+**A sweep report does not contain enough information to cut a video from.** Its
+citations carry a single start second and nothing else — `momentUrl()`
+(`common/lib/momentUrl.ts:106`) takes one `seconds` and floors it, and the MCP
+`Snippet` type (`mcp/src/search.ts:41`) has no `end` field. There is no clip
+length anywhere in a report.
+
+The end times do exist, they are just never rendered: every cue in
+`transcripts/channels/<slug>/data/<id>/transcript.cues.json` is `{start, end, text}`.
+So the manifest is built by matching each quote back to its covering cues and
+recording the real window. That is also what handles ellipsis-joined citations —
+a report quote like `"…" … "…"` is often two separate cue spans presented as one.
+
+The manifest additionally pins the provenance (share link, corpus handle, match
+counts, the narrowing queries) so a rebuild months later is reproducible and the
+video's own claims about its coverage can be checked.
+
+## Where a clip actually gets cut
+
+Three stages, because a cue span is the wrong answer twice over.
+
+1. **Cue span** — the raw window covering the quote, from `transcript.cues.json`.
+2. **Sentence widening** (`resolve-windows.mjs`) — walk outward to the nearest cue
+ ending in `.`, `?` or `!`. A cue boundary is where the *caption line wrapped*,
+ so cutting there drops the run-up that makes a quote intelligible.
+
+ Two asymmetries matter. The **start** takes a lead-in only if a real sentence
+ opening sits within `--max-lead` (8 s); otherwise it takes none, because a
+ half-sentence run-up is the irrelevant context you were trying to avoid, not
+ context. The **end** is never clamped to a budget — stopping partway through a
+ sentence is the exact mid-thought ending this removes — so `--max-tail` (12 s)
+ only bounds how far it looks before giving up and using the cue end.
+
+ **Some uploads have no punctuation at all.** Older ASR in this corpus emits
+ unpunctuated cue text for whole videos, and sentence detection then has nothing
+ to find: widening degrades to the raw cue span at both ends. That is not a
+ silent failure you can ignore — it is what produced a clip opening mid-thought
+ on "higher than they can afford and because", and what let another clip's tail
+ run through its neighbour. For those videos, pick the window by reading the
+ cues and set it by hand; the de-overlap and silence passes still apply.
+
+ **`lockStart` / `lockEnd` pin an edge** to exactly what the manifest says, and
+ neither widening nor de-overlap will move it. Reach for it when the utterance's
+ real trailing pause does not line up with its last cue's end — the snap picks
+ the *nearest* silence, and in speech over game audio the nearest one is often a
+ gap between syllables rather than the pause at the end of the thought.
+3. **Silence snapping** (`build-video.mjs`) — a sentence boundary in the
+ *transcript* still isn't a boundary in the *audio*, so clips clip words in
+ half. Fetch `fetchPad` seconds wider than needed, run `silencedetect` over the
+ result, and move each cut to the nearest silence within `snapWindow`. Starts
+ land on a silence's END (just before speech resumes), ends on a silence's START
+ (just after speech stops). No silence close enough → keep the exact point; a
+ tight cut beats a cut in the wrong place. The build logs `start✓ end✓` per clip
+ so you can see which snapped.
+
+ **The silence threshold is relative, and it has to be.** These are game
+ streams: the gaps between words are full of game audio and music — quiet, but
+ nowhere near silent. A fixed absolute threshold sits below the noise floor and
+ finds nothing. On one measured clip: mean volume −21 dB, **0** silences at
+ −32 dB, **25** at −26 dB. So each clip is measured with `volumedetect` first
+ and the threshold set `silenceRelDb` (default 6) below its own mean. If a rebuild
+ suddenly reports mostly `start– end–`, this is the knob.
+
+The trim happens during the burn-in encode, so snapping costs nothing extra.
+
+4. **De-overlap** (`resolve-windows.mjs`) — widening is per-clip and blind to its
+ neighbours, so two clips cut from the *same* video can end up overlapping, and
+ the overlap plays as the same footage twice. Any earlier clip whose tail runs
+ into a later clip's start is trimmed back to that start. This is not a rare
+ edge case: it fired on the first report, where a 2024 upload has **no
+ punctuation at all** in the relevant stretch, so sentence detection found
+ nothing and the tail ran the full budget straight through the next clip.
+
+## Regenerating and changing a video
+
+Everything is cached by content, so iteration is cheap:
+
+- **Reorder, drop or add clips** — edit `timeline`, re-run. Cached clips are not
+ refetched, so a re-cut costs an encode, not a download.
+- **Change a clip's window** — edit `start`/`end`, re-run. A raw file's window is
+ in its name, so the cache is content-addressed; and since a request is satisfied
+ by any cached file that **contains** it, a nudge inside the existing pad costs
+ nothing at all. Only a window that escapes every cached file downloads again.
+- **Preview one entry** — `--only <id>` builds a single segment and stops.
+- **Work offline** — `--skip-fetch` fails loudly instead of downloading, so you
+ can confirm you are working entirely from cache.
+- **Fetch one clip, wide** — `--fetch-only <id> --pad 20` puts a generous window
+ in the cache without building anything. This is what the umtool clip bench runs,
+ and containing-window reuse is what makes that fetch double as the build's cache.
+- **Force a refetch** — delete `out/clips-raw/`, or pass `--no-reuse` to require an
+ exact-window file.
+
+Containing-window reuse was retrofitted, and the waste it removes is measurable:
+`ferret-rescue/out/clips-raw` holds **31 files for 10 clips** because every window
+edit before this downloaded the same material again — one source is there four
+times over overlapping windows. The tightest containing file wins, not the widest,
+because silence detection decodes the whole file and a 40 s file costs more than
+the 24 s one that would also have done.
+
+After changing any window, re-run `resolve-windows.mjs --write` before building.
+It is a **fixed point**: running it on an already-resolved manifest reports
+`0 window(s) changed` and rewrites nothing. That property is load-bearing and was
+not free — the manifest stores times rounded to 2 dp, so a value read back can sit
+a hair below the cue end it came from, which lands the end lookup on the previous
+cue and runs the search on to the *next* sentence. Left alone, every re-run grew
+the same clip. Hence `EPS` in the end lookup and the 0.05 s deadband on applying a
+change.
+
+## Driven from umtool
+
+The three CLIs are the source of truth and stay usable on their own; umtool drives
+them rather than reimplementing them, so the UI and the terminal can never disagree
+about a window, a format string or the Rumble retry. Three additions exist for that:
+
+- **`--progress ndjson`** — one JSON object per line instead of prose:
+ `start`, `card`, `clip`, `fetch`, `snap`, `segment`, `entry-failed`, `concat`,
+ `chapters`, `note`, `done`, `error`. The event set is exactly what was already
+ being printed; making it a format switch is what stops a wording change from
+ breaking the driver.
+- **`--continue-on-error`** — record a failed entry and carry on. A dead source at
+ clip 14 of 19 otherwise throws away thirteen fetches already paid for. The run
+ still **refuses to concatenate** and exits non-zero: a finished file that quietly
+ lost a citation is worse than no file.
+- **`buildVideo({manifestPath, opts, out, only, fetchOnly})`** is exported, and
+ `widen()` from `resolve-windows.mjs` already was. umtool imports `widen()` so the
+ bench's "extend to sentence end" is the CLI's own function, and **spawns** the
+ build — a 40-minute chain of yt-dlp and ffmpeg inside a request handler has no
+ cancellation story.
+
+`check-availability.mjs` is the fourth CLI and belongs at the *front* of a build:
+
+```
+node scripts/report-to-video/check-availability.mjs <manifest.json>
+```
+
+It runs `yt-dlp --simulate` once per distinct `(channel, video)` — no bytes
+downloaded — classifies each failure (`deleted`, `private`, `restricted`,
+`members-only`, `geo-blocked`, `no-cues`, `maybe_missing`), and writes
+`out/availability.json` with a timestamp. This is the one fact about a manifest
+that goes stale in *both* directions: a source can die after the manifest is
+written, and a source annotated "gone" can come back. `no-cues` is called out
+separately because it is a different bug — usually the Rumble two-ids trap, where
+the manifest names the MCP video id while the cue file lives under the URL slug.
+
+## Manifest shape
+
+`timeline` is an ordered list; entries are `card` or `clip`.
+
+```jsonc
+{ "type": "card", "id": "ch3", "style": "chapter", "seconds": 4.0,
+ "kicker": "March – November 2025", "heading": "Then: the county",
+ "sub": "Six months for the first approval" }
+
+{ "type": "clip", "id": "c04", "video": "uyz1_FIqIEk",
+ "start": 32989.56, "end": 32994.19, // cue-accurate, from transcript.cues.json
+ "cite": 32989, // the second shown in the attribution line
+ "quote": "The pre-application screening was approved by the county, dude." }
+```
+
+Two per-clip fields exist for compilations that span sources or need a hand-cut
+window:
+
+- **`channel`** — the archived channel this clip's cue file lives under, overriding
+ `provenance.channelSlug`. A compilation about one person routinely spans several
+ mirror channels (`HasanAbiVODs` / `…VODs3` / `…VODsbackup`), and cue files are
+ keyed by channel, so a single manifest-wide slug cannot find them all.
+- **`lock`** — exempt this clip from `resolve-windows`. Sentence-widening exists to
+ stop clips ending mid-thought, but that is exactly wrong when the author has
+ deliberately cut a quote short: a single cue often holds a whole paragraph, so
+ trimming to "I hate this country so much sometimes" and dropping the rest of the
+ sentence is an editorial decision that widening would silently undo. `lock` also
+ handles the reverse case — a clip whose lead-in would drag in seconds of some
+ *other* audio (a news package playing before the speaker starts).
+
+Clips also carry `section` and (auto-set) `sectionEnter`. Card styles — `title`,
+`timeline`, `status`, `bullets`, `sources` — still work, but the ferret-rescue cut
+uses none of them. `render` holds resolution, fps, fonts, palette and the knobs
+(`fetchPad`, `snapWindow`, `silenceRelDb`, `transition`, `slideSeconds`,
+`headerHeight`, `footerHeight`); `provenance` holds the sweep's scope and counts.
+
+## Chrome, not cards
+
+**The ferret-rescue cut has no cards at all** — no title, no chapter breaks, no
+closing slate. It is a cited timeline and nothing else: the clips run in
+chronological order and the source material carries the argument. Cards remain
+supported for reports that want them, but the default posture is that anything
+drawn is an interruption which has to earn its place.
+
+Nothing is drawn *over* the picture either. The video is **letterboxed between**
+thin chrome rather than overlaid by it:
+
+- **Header (`headerHeight`, 56 px).** The citation only — cleaned title · upload
+ date · timestamp. No quote: the clip is already saying it, and burning in a
+ transcription of speech you can hear is noise.
+- **Footer (`footerHeight`, 100 px).** The timeline: one node per milestone, each
+ with a label and a month/year stamp beneath it. Drawn once by
+ `renderFooterAssets()`.
+- **The marker slides.** On the first clip of each section (`sectionEnter`, set
+ automatically), the fill bar and the amber marker animate from the previous node
+ to the current one over `slideSeconds`. Everywhere else they hold position. The
+ motion is ffmpeg expressions on `drawbox`/`overlay`, so it costs nothing beyond
+ the encode that was happening anyway.
+
+Clips carry `section` (index into `timelineNodes`); nodes supply `label` and
+`date`. Ordering clips chronologically is the author's job — the manifest plays in
+the order it is written.
+
+A consequence worth knowing: 16:9 source into the reduced height leaves narrow
+pillarbox bars. That is the price of never covering the picture, and it is why the
+chrome is kept as thin as it is.
+
+**Both bands are optional, and turning them off is a real setting, not a hack.**
+`footerHeight: 0` (or an empty `timelineNodes`) drops the timeline; `headerHeight: 0`
+drops the citation line. With both at zero the clips fill the whole frame and
+nothing is drawn at all — the filtergraph loses the overlays rather than compositing
+invisible ones, and `renderFooterAssets` returns early instead of drawing PNGs
+nobody uses. Two reasons this comes up:
+
+- A timeline footer only means something if the clips *are* a progression through
+ time. A cut ordered by argument rather than by date should not draw one.
+- The header prints the **upload date of the archived copy**, which for a VOD-mirror
+ channel is often years after the stream (a Nov 2019 stream re-uploaded in Apr 2023
+ reads `… November 6, 2019 … · 2023-04-06`). The title usually carries the true
+ date, so nothing is false, but on a cut spanning many re-uploads it reads badly.
+
+The `hasan-hate-america` cut runs with both off. Restoring them is a two-value edit.
+
+Note: commas inside an ffmpeg filter expression have to survive filtergraph
+parsing — wrapping the expression in single quotes is what protects them.
+
+- **Segments crossfade** (`transition`, default 0.5 s). This forces a full
+ re-encode of the timeline via `xfade`/`acrossfade` — the concat demuxer can only
+ stream-copy hard cuts. Pass `--no-xfade` for a fast hard-cut build while
+ iterating; the last pass can add the transitions back.
+
+## Things that cost time to find out
+
+**yt-dlp picks VP9 at `height<=720`, and that is a trap.** `--download-sections`
+combined with `--force-keyframes-at-cuts` re-encodes, so a VP9 pick means
+libvpx-vp9 — 27 seconds to cut a 5-second clip. It also writes a `.webm` and
+appends that extension to whatever `-o` you gave, so the file never lands where
+you asked and the run fails looking for it. Pin H.264/AAC in mp4 and the same cut
+takes ~12 seconds. `build-video.mjs` does this and keeps a rename fallback for
+the case where a fallback format still forces another container.
+
+**`--force-keyframes-at-cuts` is not optional here.** Without it the cut snaps to
+the nearest preceding keyframe and can start seconds early. That is fine for a
+human scrubbing a VOD; it is not fine when the clip *is* the citation.
+`PlayerProvider.tsx:478` builds the copyable clip command without this flag —
+correct for its purpose, wrong for ours.
+
+**`--ignore-config` is mandatory.** The operator's own yt-dlp config redirects
+output to `~/Podcasts` and attaches thumbnail/metadata post-processors. Every
+managed yt-dlp call in this repo passes `--ignore-config` for the same reason.
+
+**yt-dlp exit 101 is success**, not failure — it means a clean early stop. The
+repo encodes this at `common/ytdlp/downloadOneManaged.ts:355`; a new caller has to
+replicate it.
+
+**Local media will not help you.** Only 5 of 1,804 PirateSoftware video dirs hold
+any media at all, and none are ones a report is likely to cite. Clips are a
+network fetch. `metadata.info.json` and `transcript.cues.json` *are* present for
+every video, so titles, dates, durations, webpage URLs and cue timings all come
+from disk with no probe.
+
+**Upstream availability is load-bearing.** A clip can only be fetched while the
+source is still up. Check `platform state` via `get_video_metadata` before
+committing to a clip — a `deleted`/`maybe_missing` source needs a quote card
+instead of footage. The archive outlives its sources, so a video built from an
+old report will be *less* complete than the report unless this is handled
+deliberately.
+
+**ImageMagick's `-size` leaks into the Pango group.** `-size 1920x1080 xc:BG`
+followed by `( ... pango:@file )` renders the text into a full-frame box, which
+pins it to the top and wraps at the frame edge instead of the text column. Reset
+`-size` inside the parens.
+
+**ffmpeg `drawtext` does not wrap and hates punctuation.** Both are solved the
+same way: wrap to a column count in JS, write to a file, and use
+`textfile=`. Nothing then needs escaping. Stream titles also need emoji and
+`!command` suffixes stripped or they render as tofu in the attribution line.
+
+**Segments are encoded to identical parameters on purpose** so the final
+concatenation is a stream copy via the concat demuxer. Mismatched streams are the
+usual reason a naive concat produces a broken or audio-desynced file.
+
+## Not done yet
+
+- **Narration is silent by design.** Cards carry the connective text; the only
+ audio is the clips'. A TTS layer would attach per card (`seconds` already gives
+ it a duration to fill) — deliberately deferred rather than designed out.
+- **Snapping is silence-based, not word-based.** It finds gaps in the audio, which
+ is usually the same thing as a word boundary but is not guaranteed to be —
+ a speaker who does not pause gets the unsnapped cut. Forced alignment against
+ the transcript would be exact; `silencedetect` is a tenth of the work and
+ handles the cases that were actually audible.
+- **Manifests are written by hand** from verified cue data. Deriving a first-draft
+ manifest automatically from a report's citations is the obvious next step; the
+ report parse is straightforward (`> "quote"` followed by
+ `— [title @ h:mm:ss](…?v=slug%2Fid&t=sec)`), the cue-matching is the real work.
diff --git a/scripts/report-to-video/build-video.mjs b/scripts/report-to-video/build-video.mjs
@@ -0,0 +1,839 @@
+#!/usr/bin/env node
+// build-video.mjs — render a cited sweep report into a narrated-by-text video.
+//
+// Takes a video manifest (see README.md next to this file) and produces one mp4:
+// text cards state the findings, clips let the source say it in their own voice,
+// and every clip carries a burned-in quote plus its attribution.
+//
+// Pipeline, per manifest entry:
+// card -> still PNG (render-cards.mjs) -> N seconds of video + silent audio
+// clip -> yt-dlp --download-sections (WIDE) -> silence-snap -> trim + burn
+// then the segments are crossfaded together into the finished file.
+//
+// Three things worth knowing about how clips are cut:
+//
+// 1. Windows come from the manifest as absolute [start, end] seconds, derived
+// from transcript.cues.json (which carries an END per cue). A sweep report
+// only ever records a single start second, so windows cannot be recovered
+// from the report alone.
+// 2. Those windows are widened to sentence boundaries by resolve-windows.mjs,
+// so a clip carries the run-up that makes the quote make sense.
+// 3. A cue boundary is still not a *speech* boundary — cutting there clips
+// words in half. So we fetch wider than needed and snap the real cut to a
+// silence found in the audio. That is what makes clips start and end
+// between words rather than through them.
+//
+// Fetched clips are cached by (video, start, end); re-running is cheap and only
+// changed entries re-download. Delete out/clips-raw to force a refetch.
+//
+// In the app: not used. On the CLI:
+// node scripts/report-to-video/build-video.mjs <manifest.json> [options]
+//
+// Options:
+// --out <dir> Output root (default: manifest dir + /out)
+// --skip-fetch Fail instead of downloading anything not already cached
+// --only <id> Build a single entry's segment and stop (for iterating)
+// --no-xfade Hard cuts instead of crossfades (much faster; concat copy)
+// --progress ndjson One JSON event per line instead of prose (for umtool)
+// --continue-on-error Record a failed entry and carry on, instead of aborting
+// --fetch-only <id> Fetch one clip's window into clips-raw and stop
+// --pad <s> Override render.fetchPad (the clip bench fetches wide)
+//
+// Requires: yt-dlp, ffmpeg/ffprobe, ImageMagick with Pango.
+
+import { execFile } from "node:child_process";
+import { promisify } from "node:util";
+import { mkdir, writeFile, readFile, access, readdir, rename } from "node:fs/promises";
+import path from "node:path";
+
+import { renderCard, renderFooterAssets } from "./render-cards.mjs";
+
+const execFileP = promisify(execFile);
+
+const YTDLP = process.env.YTDLP_BIN ?? "yt-dlp";
+const FFMPEG = process.env.FFMPEG_BIN ?? "ffmpeg";
+const FFPROBE = process.env.FFPROBE_BIN ?? "ffprobe";
+const QRENCODE = process.env.QRENCODE_BIN ?? "qrencode";
+
+const CHANNELS_DIR =
+ process.env.CHANNELS_DIR ??
+ "/home/user/Projects/yt-dlp-transcript-browser/transcripts/channels";
+
+const exists = (p) => access(p).then(() => true, () => false);
+
+// ---- progress protocol ---------------------------------------------------
+// This has two audiences: a human watching a terminal, and umtool's build driver
+// reading the pipe. Rather than have the driver scrape prose (which would make
+// every wording change a breaking change), `--progress ndjson` switches every
+// line to one JSON object. The event set is exactly what was already being
+// printed -- this is a formatting switch, not new instrumentation.
+//
+// Events: start, card, clip, fetch, snap, segment, entry-failed, concat,
+// chapters, note, done.
+const HUMAN = {
+ start: (e) => `${e.title} — ${e.entries} entr(ies)`,
+ card: (e) => `card ${e.id}`,
+ clip: (e) => `clip ${e.id} (${e.video}) §${e.section}${e.sectionEnter ? " ⟶" : ""}`,
+ fetch: (e) =>
+ e.reuse
+ ? ` fetch ${e.id}: ${e.reuse} already covers ${hms(e.from)}–${hms(e.to)} — no download`
+ : e.cached
+ ? null
+ : ` fetch ${e.id}: ${e.video} ${hms(e.from)}–${hms(e.to)}`,
+ snap: (e) =>
+ ` snap ${e.id}: ${e.start ? "start✓" : "start–"} ${e.end ? "end✓" : "end–"} ` +
+ `(${Number(e.seconds).toFixed(1)}s)`,
+ segment: () => null,
+ "entry-failed": (e) => ` ** ${e.id} failed: ${e.message}`,
+ concat: (e) => `${e.mode === "xfade" ? "crossfading" : "hard-cutting"} ${e.n} segments…`,
+ chapters: (e) => `chapters: ${e.n} marker(s) -> ${e.file}`,
+ note: (e) => e.message,
+ done: (e) =>
+ e.duration === undefined
+ ? `built ${e.out}`
+ : `\n${e.out}\nduration=${e.duration}\nsize=${e.size}`,
+};
+
+let EMIT = (ev, fields = {}) => {
+ const line = HUMAN[ev]?.({ ev, ...fields });
+ if (line) console.log(line);
+};
+
+export function setProgressMode(mode) {
+ EMIT =
+ mode === "ndjson"
+ ? (ev, fields = {}) => process.stdout.write(JSON.stringify({ ev, ...fields }) + "\n")
+ : (ev, fields = {}) => {
+ const line = HUMAN[ev]?.({ ev, ...fields });
+ if (line) console.log(line);
+ };
+}
+
+function hms(total) {
+ const s = Math.floor(total);
+ const h = Math.floor(s / 3600);
+ const m = Math.floor((s % 3600) / 60);
+ const sec = s % 60;
+ return h > 0
+ ? `${h}:${String(m).padStart(2, "0")}:${String(sec).padStart(2, "0")}`
+ : `${m}:${String(sec).padStart(2, "0")}`;
+}
+
+// Stream titles here are full of emoji and !commands. drawtext renders them as
+// tofu with a text font, and they add nothing to an attribution line.
+function cleanTitle(title) {
+ return title
+ .replace(/[\u{1F000}-\u{1FFFF}\u{2600}-\u{27BF}\u{FE0F}]/gu, "")
+ .replace(/\s*[!@]\S+/g, "")
+ .replace(/\s{2,}/g, " ")
+ .replace(/[\s·|-]+$/, "")
+ .trim();
+}
+
+// drawtext does not wrap. Break to a character budget, write to a file, and use
+// textfile= so nothing needs shell or filter escaping.
+function wrap(text, cols) {
+ const words = text.split(/\s+/);
+ const lines = [];
+ let line = "";
+ for (const w of words) {
+ if (line && (line + " " + w).length > cols) {
+ lines.push(line);
+ line = w;
+ } else {
+ line = line ? line + " " + w : w;
+ }
+ }
+ if (line) lines.push(line);
+ return lines.join("\n");
+}
+
+async function videoMeta(videoId, channelSlug) {
+ const p = path.join(CHANNELS_DIR, channelSlug, "data", videoId, "transcript.cues.json");
+ const d = JSON.parse(await readFile(p, "utf8"));
+ return { title: d.title, uploadDate: d.uploadDate, webpageUrl: d.webpageUrl, duration: d.duration };
+}
+
+async function probeDuration(file) {
+ const { stdout } = await execFileP(FFPROBE, [
+ "-v", "error", "-show_entries", "format=duration",
+ "-of", "default=nw=1:nk=1", file,
+ ]);
+ return Number(stdout.trim());
+}
+
+// yt-dlp exits 101 on a clean early stop (break-on-existing / max-downloads).
+// The repo treats that as success everywhere else; do the same here.
+const ytdlpOk = (err) => err?.code === 101;
+
+// ---- the clip cache ------------------------------------------------------
+// A raw clip's window is IN ITS NAME, which makes the file immutable and the
+// cache content-addressed. The original lookup was for the exact name, so any
+// change to a window -- a hand edit, a widen, a nudge in the clip bench -- was a
+// fresh download of material already on disk. Measured on ferret-rescue: 31
+// files for 10 clips, one source fetched four times over overlapping windows.
+//
+// So: satisfy a request from ANY cached file that contains it. The TIGHTEST
+// container wins, because detectSilence decodes the whole file and a 40s file
+// costs more than the 14s one that would also have done. The clip bench fetches
+// deliberately wide, and this is what makes that generous fetch become the
+// build's cache rather than a second one.
+const WINDOW_RE = /^(\d+(?:\.\d+)?)-(\d+(?:\.\d+)?)$/;
+
+// A window read back from a 2 dp manifest can sit a hair outside the file that
+// produced it; the same tolerance resolve-windows.mjs uses for the same reason.
+const WIN_EPS = 0.02;
+
+export async function cachedWindowsFor(rawDir, video) {
+ let names;
+ try {
+ names = await readdir(rawDir);
+ } catch {
+ return [];
+ }
+ const prefix = `${video}_`;
+ const out = [];
+ for (const name of names) {
+ if (!name.startsWith(prefix) || !name.endsWith(".mp4")) continue;
+ // The remainder must be exactly `a-b`, which is what stops a video id that
+ // is a prefix of another (or one containing `_`) from claiming its files.
+ const m = WINDOW_RE.exec(name.slice(prefix.length, -4));
+ if (!m) continue;
+ out.push({ name, path: path.join(rawDir, name), from: Number(m[1]), to: Number(m[2]) });
+ }
+ return out;
+}
+
+/** The tightest cached file containing [from, to], or null. */
+export async function findContainingWindow(rawDir, video, from, to) {
+ const windows = await cachedWindowsFor(rawDir, video);
+ let best = null;
+ for (const w of windows) {
+ if (w.from > from + WIN_EPS || w.to < to - WIN_EPS) continue;
+ if (!best || w.to - w.from < best.to - best.from) best = w;
+ }
+ return best;
+}
+
+async function fetchClip(entry, meta, render, outDir, opts) {
+ // Deliberately over-fetch: the snapping pass below needs room on both sides to
+ // find a silence, and a clip that has no slack can only be cut where the cue
+ // happened to break — which is what put words in half in the first place.
+ const pad = opts.pad ?? render.fetchPad ?? 3.0;
+ const from = Math.max(0, entry.start - pad);
+ const to = entry.end + pad;
+
+ const rawDir = path.join(outDir, "clips-raw");
+ const name = `${entry.video}_${from.toFixed(2)}-${to.toFixed(2)}.mp4`;
+ const dest = path.join(rawDir, name);
+ if (await exists(dest)) {
+ EMIT("fetch", { id: entry.id, video: entry.video, from, to, cached: true });
+ return { path: dest, fetchStart: from, cached: true };
+ }
+ if (!opts.noReuse) {
+ const hit = await findContainingWindow(rawDir, entry.video, from, to);
+ if (hit) {
+ EMIT("fetch", {
+ id: entry.id, video: entry.video, from, to, cached: true, reuse: hit.name,
+ });
+ // fetchStart is the CACHED file's start, not the requested one -- every cut
+ // downstream is expressed relative to it, so reuse is transparent.
+ return { path: hit.path, fetchStart: hit.from, cached: true };
+ }
+ }
+ if (opts.skipFetch) throw new Error(`--skip-fetch set and no cached window covers ${name}`);
+
+ const maxH = render.maxHeightSource;
+ const fmt = [
+ `bv*[vcodec^=avc1][height<=${maxH}]+ba[acodec^=mp4a]`,
+ `bv*[ext=mp4][height<=${maxH}]+ba[ext=m4a]`,
+ `b[ext=mp4][height<=${maxH}]`,
+ `b[height<=${maxH}]`,
+ ].join("/");
+
+ const argsWith = (extra) => [
+ // The operator's own yt-dlp config redirects output and attaches thumbnail
+ // and metadata post-processors; without this the clips land elsewhere.
+ "--ignore-config",
+ "--no-playlist",
+ "--download-sections", `*${from.toFixed(2)}-${to.toFixed(2)}`,
+ // Without this the cut snaps to the nearest preceding keyframe, which can be
+ // seconds early — fine for scrubbing, not fine when the clip IS the citation.
+ "--force-keyframes-at-cuts",
+ ...extra,
+ // Pin H.264/AAC in mp4. Left alone yt-dlp picks VP9+Opus at these heights,
+ // and since --force-keyframes-at-cuts re-encodes, that means libvpx-vp9 —
+ // 27s to cut a 5s clip. It also writes .webm and appends that to -o.
+ "-f", fmt,
+ "--merge-output-format", "mp4",
+ "-o", dest,
+ "--", meta.webpageUrl,
+ ];
+
+ const attempt = async (extra) => {
+ try {
+ await execFileP(YTDLP, argsWith(extra), { maxBuffer: 1 << 26 });
+ return null;
+ } catch (err) {
+ return ytdlpOk(err) ? null : err;
+ }
+ };
+
+ EMIT("fetch", { id: entry.id, video: entry.video, from, to, cached: false });
+ let err = await attempt([]);
+
+ // Rumble delivers HLS whose segments are named `.tar`, and ffmpeg 8's picky
+ // extension check rejects those outright — "URL ... is not in
+ // allowed_segment_extensions" — killing the fetch with exit 183. Rumble ships
+ // no progressive format to fall back to, so without this every Rumble-sourced
+ // clip is unbuildable.
+ //
+ // It has to be a RETRY, not a default: -extension_picky lives on the HLS
+ // demuxer, so passing it against a progressive URL (YouTube's googlevideo mp4)
+ // makes ffmpeg abort with "Option extension_picky not found" — i.e. adding it
+ // unconditionally trades a Rumble failure for a YouTube one.
+ if (err && /allowed_segment_extensions|allowed_extensions/.test(String(err.stderr ?? err.message ?? ""))) {
+ EMIT("note", { id: entry.id, message: ` ${entry.id}: HLS segment extension rejected, retrying with -extension_picky 0` });
+ err = await attempt(["--downloader-args", "ffmpeg_i:-extension_picky 0"]);
+ }
+ if (err) {
+ throw new Error(`yt-dlp failed for ${entry.id} (${entry.video}): ${err.stderr ?? err.message}`);
+ }
+ if (!(await exists(dest))) {
+ // If a fallback format still forced another container, yt-dlp writes
+ // "<dest>.<realext>". Adopt it rather than failing the run.
+ const dir = path.dirname(dest);
+ const base = path.basename(dest);
+ const stray = (await readdir(dir)).find((f) => f.startsWith(base + "."));
+ if (!stray) throw new Error(`yt-dlp reported success but produced no file for ${entry.id}`);
+ await rename(path.join(dir, stray), dest);
+ }
+ return { path: dest, fetchStart: from, cached: false };
+}
+
+// Parse ffmpeg's silencedetect output into [{s, e}] intervals, in seconds
+// relative to the start of the given file.
+async function detectSilence(file, render) {
+ const minDur = render.silenceMinDur ?? 0.09;
+
+ // The threshold has to be RELATIVE to the clip, not absolute. These are game
+ // streams: the gaps between words are full of game audio and music, so they
+ // are quiet but nowhere near silent. A fixed -32 dB sits below the noise floor
+ // of a typical clip here and finds literally zero silences (measured: mean
+ // volume -21 dB, 0 hits at -32 dB, 25 hits at -26 dB). Measure the clip first
+ // and cut a few dB under its own mean instead.
+ const { stderr: volLog } = await execFileP(
+ FFMPEG,
+ ["-nostdin", "-i", file, "-af", "volumedetect", "-f", "null", "-"],
+ { maxBuffer: 1 << 26 },
+ ).catch((e) => ({ stderr: e.stderr ?? "" }));
+ const meanMatch = (volLog ?? "").match(/mean_volume:\s*(-?[\d.]+) dB/);
+ const mean = meanMatch ? Number(meanMatch[1]) : -24;
+ const noise = Math.max(-45, Math.min(-18, mean - (render.silenceRelDb ?? 6)));
+
+ // ffmpeg exits 0 here, so stderr comes back on the resolved result.
+ const { stderr } = await execFileP(
+ FFMPEG,
+ ["-nostdin", "-i", file, "-af", `silencedetect=noise=${noise.toFixed(1)}dB:d=${minDur}`, "-f", "null", "-"],
+ { maxBuffer: 1 << 26 },
+ ).catch((e) => ({ stderr: e.stderr ?? "" }));
+ const log = stderr ?? "";
+
+ const out = [];
+ let open = null;
+ for (const line of log.split("\n")) {
+ const s = line.match(/silence_start:\s*(-?[\d.]+)/);
+ if (s) open = Number(s[1]);
+ const e = line.match(/silence_end:\s*(-?[\d.]+)/);
+ if (e && open !== null) {
+ out.push({ s: open, e: Number(e[1]) });
+ open = null;
+ }
+ }
+ return out;
+}
+
+// Snap a desired cut to the nearest silence, so the clip begins and ends between
+// words instead of through one. Returns the desired point unchanged when no
+// silence is close enough — better a tight cut than a cut in the wrong place.
+function snap(desired, intervals, kind, window) {
+ let best = null;
+ for (const iv of intervals) {
+ // Starting: we want to resume just before speech does -> the silence's END.
+ // Ending: we want to stop just after speech does -> the silence's START.
+ const point = kind === "start" ? iv.e : iv.s;
+ const d = Math.abs(point - desired);
+ if (d > window) continue;
+ if (!best || d < best.d) best = { d, point };
+ }
+ if (!best) return { at: desired, snapped: false };
+ const lead = kind === "start" ? -0.10 : 0.18;
+ return { at: Math.max(0, best.point + lead), snapped: true };
+}
+
+// Encoder quality is manifest-driven so a cut can trade size for fidelity without
+// editing this file. Defaults reproduce the original hardcoded settings exactly.
+const encodeArgs = (render) => [
+ "-c:v", "libx264",
+ "-preset", render.preset ?? "medium",
+ "-crf", String(render.crf ?? 20),
+ "-pix_fmt", "yuv420p",
+ "-r", String(render.fps),
+ "-c:a", "aac",
+ "-b:a", render.audioBitrate ?? "160k",
+ "-ar", String(render.audioRate),
+ "-ac", String(render.audioChannels),
+ "-movflags", "+faststart",
+];
+
+async function buildClipSegment(entry, meta, render, outDir, opts, chrome, nodes, provenance) {
+ const { path: raw, fetchStart } = await fetchClip(entry, meta, render, outDir, opts);
+ const seg = path.join(outDir, "segments", `${entry.id}.mp4`);
+ const pal = render.palette;
+ const { width, height } = render;
+
+ // Desired cut points, expressed relative to the over-fetched file.
+ const wantA = entry.start - fetchStart;
+ const wantB = entry.end - fetchStart;
+ const win = render.snapWindow ?? 1.6;
+
+ const sil = await detectSilence(raw, render);
+ const a = snap(wantA, sil, "start", win);
+ const b = snap(wantB, sil, "end", win);
+ // Never let snapping invert or collapse the window.
+ const cutA = Math.min(a.at, wantB - 1);
+ const cutB = Math.max(b.at, cutA + 1);
+ EMIT("snap", { id: entry.id, start: a.snapped, end: b.snapped, seconds: cutB - cutA });
+
+ const quotePath = path.join(outDir, "segments", `${entry.id}.quote.txt`);
+ const attribPath = path.join(outDir, "segments", `${entry.id}.attrib.txt`);
+ // Written for reference/diffing only — the quote is no longer drawn on screen.
+ await writeFile(quotePath, wrap(`“${entry.quote}”`, 92), "utf8");
+
+ const d = meta.uploadDate;
+ const date = `${d.slice(0, 4)}-${d.slice(4, 6)}-${d.slice(6, 8)}`;
+ await writeFile(
+ attribPath,
+ `${cleanTitle(meta.title)} · ${date} @ ${hms(entry.cite ?? entry.start)}`,
+ "utf8",
+ );
+
+ // The picture is the point. Nothing is drawn over it: the video is letterboxed
+ // between a thin citation header and a thin timeline footer, so the source
+ // material plays unobstructed and the additions stay subtle.
+ const HH = render.headerHeight ?? 56;
+ const FH = chrome.footerHeight;
+ const hasFooter = FH > 0 && chrome.footer;
+ // headerHeight:0 drops the citation line too, leaving the clips alone on screen.
+ // Worth having: a cut whose sources are listed elsewhere does not need to carry
+ // its own attribution burnt into every frame.
+ const hasHeader = HH > 0;
+ const VH = height - HH - FH;
+ const trackAbsY = height - FH + chrome.trackY;
+
+ // Where the progress marker travels this clip. Only the first clip of a
+ // section moves it; the rest hold it in place.
+ const T = render.slideSeconds ?? 0.9;
+ const xTo = chrome.xs[entry.section];
+ const xFrom = entry.sectionEnter ? chrome.xs[Math.max(0, entry.section - 1)] : xTo;
+ // Commas inside a filter option have to survive filtergraph parsing; single
+ // quotes around the expression is what protects them.
+ const ramp = (a, b) =>
+ a === b ? String(b) : `'if(lt(t,${T}),${a}+(${b}-${a})*t/${T},${b})'`;
+ const fillW = ramp(xFrom - chrome.x0, xTo - chrome.x0);
+ const markX = ramp(xFrom - chrome.markerRadius, xTo - chrome.markerRadius);
+
+ const base = [
+ `scale=${width}:${VH}:force_original_aspect_ratio=decrease`,
+ `pad=${width}:${VH}:(ow-iw)/2:(oh-ih)/2:color=${pal.bg}`,
+ `pad=${width}:${height}:0:${HH}:color=${pal.bg}`,
+ "setsar=1",
+ `fps=${render.fps}`,
+ ...(hasHeader
+ ? [
+ `drawbox=x=90:y=${Math.round((HH - 24) / 2)}:w=4:h=24:color=${pal.accent}:t=fill`,
+ [
+ `drawtext=textfile='${attribPath}'`,
+ `fontfile='${render.fontRegular}'`,
+ "fontsize=22",
+ `fontcolor=${pal.muted}`,
+ "x=118",
+ `y=${Math.round((HH - 26) / 2)}`,
+ ].join(":"),
+ ]
+ : []),
+ ].join(",");
+
+ const qr = render.qr === false ? null : await qrForEntry(entry, provenance, render, outDir);
+ const qrIdx = hasFooter ? 3 : 1;
+ const qrM = render.qr?.margin ?? 28;
+
+ const parts = hasFooter
+ ? [
+ `[0:v]${base}[b]`,
+ `[b][1:v]overlay=0:${height - FH}[f]`,
+ `[f]drawbox=x=${chrome.x0}:y=${trackAbsY - 1}:w=${fillW}:h=3:color=${pal.accent}:t=fill[g]`,
+ `[g][2:v]overlay=x=${markX}:y=${trackAbsY - chrome.markerRadius}[q]`,
+ ]
+ : [`[0:v]${base}[q]`];
+
+ // Sit above the footer when there is one, so the code never straddles the chrome.
+ parts.push(
+ qr
+ ? `[q][${qrIdx}:v]overlay=x=W-w-${qrM}:y=H-h-${FH + qrM}[v]`
+ : `[q]null[v]`,
+ );
+
+ await execFileP(
+ FFMPEG,
+ [
+ "-nostdin", "-v", "error", "-y",
+ "-ss", cutA.toFixed(3), "-to", cutB.toFixed(3), "-i", raw,
+ ...(hasFooter ? ["-i", chrome.footer, "-i", chrome.marker] : []),
+ ...(qr ? ["-i", qr.png] : []),
+ "-filter_complex", parts.join(";"),
+ "-map", "[v]", "-map", "0:a",
+ ...encodeArgs(render),
+ seg,
+ ],
+ { maxBuffer: 1 << 24 },
+ );
+ return seg;
+}
+
+// ---- QR provenance code --------------------------------------------------
+// A compilation asks the viewer to take the edit on trust. The QR is the antidote:
+// it resolves to this clip's exact START in the archive's own viewer, so anyone can
+// pull up the surrounding hour and check that the cut is fair. Per clip, because a
+// single code for the whole video would send everyone to the first citation.
+//
+// Two rules learned the hard way: it must be FULLY OPAQUE (a translucent QR will
+// not scan) and it must keep its quiet zone (the white border is part of the
+// symbol, not decoration).
+async function qrForEntry(entry, provenance, render, outDir) {
+ const q = render.qr ?? {};
+ // A mirror's LOCAL slug is not the id the site serves, and a clip taken from a
+ // copy whose archived transcript is broken should point at the copy that reads —
+ // so an explicit per-clip citeUrl always wins over the derived one.
+ const url =
+ entry.citeUrl ??
+ `${provenance.siteOrigin}/?v=${encodeURIComponent(
+ `${entry.channel ?? provenance.channelSlug}/${entry.video}`,
+ )}&t=${Math.floor(entry.start)}`;
+ const png = path.join(outDir, "qr", `${entry.id}.png`);
+ await execFileP(QRENCODE, [
+ "-o", png,
+ "-s", String(q.scale ?? 4),
+ "-m", String(q.quiet ?? 3),
+ "-l", q.ecc ?? "M",
+ url,
+ ]);
+ return { png, url };
+}
+
+async function buildCardSegment(card, render, outDir, nodes) {
+ const png = await renderCard(card, render, outDir, nodes);
+ const seg = path.join(outDir, "segments", `${card.id}.mp4`);
+ const dur = String(card.seconds);
+
+ await execFileP(
+ FFMPEG,
+ [
+ "-nostdin", "-v", "error", "-y",
+ "-loop", "1", "-t", dur, "-i", png,
+ "-f", "lavfi", "-t", dur,
+ "-i", `anullsrc=channel_layout=stereo:sample_rate=${render.audioRate}`,
+ "-vf", `fps=${render.fps},setsar=1`,
+ ...encodeArgs(render),
+ "-shortest",
+ seg,
+ ],
+ { maxBuffer: 1 << 24 },
+ );
+ return seg;
+}
+
+// Crossfade every segment into the next. This is a full re-encode of the
+// timeline — the concat demuxer can only stream-copy hard cuts — so --no-xfade
+// stays available for quick iteration.
+async function concatWithXfade(segments, render, outPath) {
+ const D = render.transition ?? 0.5;
+ const durs = [];
+ for (const s of segments) durs.push(await probeDuration(s));
+
+ const inputs = segments.flatMap((s) => ["-i", s]);
+ const parts = [];
+ let vlab = "[0:v]";
+ let alab = "[0:a]";
+ let acc = durs[0];
+
+ for (let i = 1; i < segments.length; i += 1) {
+ const off = acc - D;
+ parts.push(`${vlab}[${i}:v]xfade=transition=fade:duration=${D}:offset=${off.toFixed(3)}[v${i}]`);
+ parts.push(`${alab}[${i}:a]acrossfade=d=${D}:c1=tri:c2=tri[a${i}]`);
+ vlab = `[v${i}]`;
+ alab = `[a${i}]`;
+ acc = acc + durs[i] - D;
+ }
+
+ await execFileP(
+ FFMPEG,
+ [
+ "-nostdin", "-v", "error", "-y",
+ ...inputs,
+ "-filter_complex", parts.join(";"),
+ "-map", vlab, "-map", alab,
+ ...encodeArgs(render),
+ outPath,
+ ],
+ { maxBuffer: 1 << 26 },
+ );
+}
+
+// ---- chapter markers -----------------------------------------------------
+// A compilation like this is a reference document as much as a video: the report
+// cites moments, and a viewer wants to jump to them. Every clip therefore becomes
+// a chapter. Offsets are derived exactly the way concatWithXfade derives its xfade
+// offsets, so they stay correct for both crossfaded and hard-cut timelines.
+//
+// ffmetadata is a line-based format where =, ;, # and \ are structural, so a
+// title carrying any of them has to be escaped or the file silently mis-parses.
+const ffmetaEscape = (s) => String(s).replace(/([=;#\\])/g, "\\$1").replace(/\n/g, " ");
+
+async function segmentOffsets(segments, D) {
+ const durs = [];
+ for (const s of segments) durs.push(await probeDuration(s));
+ const starts = [];
+ let acc = 0;
+ for (let i = 0; i < durs.length; i += 1) {
+ starts.push(acc);
+ acc += durs[i] - (i < durs.length - 1 ? D : 0);
+ }
+ return { starts, total: acc };
+}
+
+async function chapterTitle(entry, index, provenance) {
+ if (entry.chapter) return entry.chapter;
+ if (entry.type === "card") return entry.title ?? `Card ${index + 1}`;
+ try {
+ const meta = await videoMeta(entry.video, entry.channel ?? provenance.channelSlug);
+ const d = String(meta.uploadDate ?? "");
+ const date = /^\d{8}$/.test(d) ? `${d.slice(0, 4)}-${d.slice(4, 6)}-${d.slice(6, 8)}` : d;
+ const title = String(meta.title ?? entry.video);
+ return `${date} — ${title.length > 60 ? `${title.slice(0, 57)}…` : title}`.trim();
+ } catch {
+ return `${index + 1}. ${entry.video}`;
+ }
+}
+
+async function muxChapters(finalPath, entries, segments, D, outDir, provenance) {
+ if (segments.length < 2) return;
+ const { starts, total } = await segmentOffsets(segments, D);
+ const lines = [";FFMETADATA1", ""];
+ for (let i = 0; i < entries.length; i += 1) {
+ // Land just PAST the crossfade, so the marker opens on the incoming clip
+ // rather than on the outgoing one mid-dissolve.
+ const start = i === 0 ? 0 : starts[i] + D;
+ const end = i === entries.length - 1 ? total : starts[i + 1] + D;
+ lines.push(
+ "[CHAPTER]",
+ "TIMEBASE=1/1000",
+ `START=${Math.round(start * 1000)}`,
+ `END=${Math.round(end * 1000)}`,
+ `title=${ffmetaEscape(await chapterTitle(entries[i], i, provenance))}`,
+ "",
+ );
+ }
+ const metaPath = path.join(outDir, "chapters.ffmeta");
+ await writeFile(metaPath, lines.join("\n"), "utf8");
+
+ // Stream copy — adding chapters must never re-encode the finished timeline.
+ const tmp = finalPath.replace(/\.mp4$/, ".chapters.mp4");
+ await execFileP(
+ FFMPEG,
+ ["-nostdin", "-v", "error", "-y", "-i", finalPath, "-i", metaPath,
+ "-map", "0", "-map_metadata", "0", "-map_chapters", "1", "-c", "copy", tmp],
+ { maxBuffer: 1 << 24 },
+ );
+ await rename(tmp, finalPath);
+ EMIT("chapters", { n: entries.length, file: path.basename(metaPath) });
+}
+
+async function concatHardCut(segments, outDir, outPath) {
+ const listPath = path.join(outDir, "concat.txt");
+ await writeFile(listPath, segments.map((s) => `file '${s}'`).join("\n") + "\n", "utf8");
+ await execFileP(
+ FFMPEG,
+ ["-nostdin", "-v", "error", "-y", "-f", "concat", "-safe", "0", "-i", listPath, "-c", "copy", outPath],
+ { maxBuffer: 1 << 24 },
+ );
+}
+
+/**
+ * Build a manifest into a video.
+ *
+ * Exported so umtool's driver runs the SAME code the CLI does. It is still
+ * SPAWNED rather than imported by the app: a 40-minute chain of yt-dlp and
+ * ffmpeg inside a request handler has no cancellation story, and a runaway
+ * grandchild would outlive the request that started it.
+ */
+export async function buildVideo({ manifestPath, opts = {}, out, only, fetchOnly } = {}) {
+ const manifest = JSON.parse(await readFile(manifestPath, "utf8"));
+ const { render, provenance } = manifest;
+ const outDir = out ?? path.join(path.dirname(path.resolve(manifestPath)), "out");
+
+ for (const d of ["cards", "clips-raw", "segments", "qr"]) {
+ await mkdir(path.join(outDir, d), { recursive: true });
+ }
+
+ // Fetch one clip's window and stop. This is what the clip bench's "fetch 20s
+ // more" runs, so a bench fetch and a build fetch can never disagree about
+ // naming, format selection, the VP9 trap or the Rumble HLS retry.
+ if (fetchOnly) {
+ const entry = manifest.timeline.find((e) => e.id === fetchOnly);
+ if (!entry) throw new Error(`no timeline entry with id ${fetchOnly}`);
+ if (entry.type === "card") throw new Error(`${fetchOnly} is a card, not a clip`);
+ const meta = await videoMeta(entry.video, entry.channel ?? provenance.channelSlug);
+ const r = await fetchClip(entry, meta, render, outDir, opts);
+ EMIT("done", { out: r.path, fetchStart: r.fetchStart, cached: r.cached });
+ return { out: r.path, failures: [] };
+ }
+
+ // Footer chrome is shared by every clip, so build it once up front.
+ const chrome = await renderFooterAssets(render, manifest.timelineNodes, outDir);
+
+ const entries = manifest.timeline.filter((e) => !only || e.id === only);
+ if (only && !entries.length) throw new Error(`no timeline entry with id ${only}`);
+ const segments = [];
+ const failures = [];
+
+ const D = opts.noXfade || (render.transition ?? 0.5) === 0 ? 0 : render.transition ?? 0.5;
+
+ // Retro-fit chapters onto an already-built file without re-encoding it. The
+ // per-clip segments are still on disk, which is all the offsets need.
+ if (opts.chaptersOnly) {
+ const finalPath = path.join(outDir, `${manifest.slug}.mp4`);
+ const segs = entries.map((e) => path.join(outDir, "segments", `${e.id}.mp4`));
+ for (const seg of segs) {
+ if (!(await exists(seg)))
+ throw new Error(`--chapters-only needs ${seg}, which is missing — run a full build first`);
+ }
+ await muxChapters(finalPath, entries, segs, D, outDir, provenance);
+ return { out: finalPath, failures: [] };
+ }
+
+ EMIT("start", { title: manifest.title, entries: entries.length, out: outDir });
+ for (let i = 0; i < entries.length; i += 1) {
+ const entry = entries[i];
+ try {
+ if (entry.type === "card") {
+ EMIT("card", { id: entry.id, i, n: entries.length });
+ segments.push(await buildCardSegment(entry, render, outDir, manifest.timelineNodes));
+ } else {
+ const meta = await videoMeta(entry.video, entry.channel ?? provenance.channelSlug);
+ EMIT("clip", {
+ id: entry.id, i, n: entries.length, video: entry.video,
+ section: entry.section, sectionEnter: !!entry.sectionEnter,
+ });
+ segments.push(
+ await buildClipSegment(
+ entry, meta, render, outDir, opts, chrome, manifest.timelineNodes, provenance,
+ ),
+ );
+ }
+ EMIT("segment", { id: entry.id, path: segments[segments.length - 1] });
+ } catch (err) {
+ // Without --continue-on-error a dead source at entry 14 of 19 throws away
+ // the thirteen fetches already paid for. With it, everything buildable is
+ // built and the run reports what was not.
+ if (!opts.continueOnError) throw err;
+ const message = err?.message ?? String(err);
+ failures.push({ id: entry.id, message });
+ EMIT("entry-failed", { id: entry.id, message });
+ }
+ }
+
+ if (only) {
+ EMIT("done", { out: segments[0], failures });
+ return { out: segments[0], failures };
+ }
+
+ // A timeline that silently lost a clip is a worse outcome than no file at all:
+ // the finished video would look complete and be missing a citation. So the
+ // segments are kept (they cost the fetches) and the concat is refused.
+ if (failures.length) {
+ EMIT("note", {
+ message: `refusing to concat: ${failures.length} of ${entries.length} entries failed ` +
+ `(${failures.map((f) => f.id).join(", ")})`,
+ });
+ return { out: null, failures };
+ }
+
+ const final = path.join(outDir, `${manifest.slug}.mp4`);
+ // `transition: 0` is a real editorial choice, not just a speed knob: hard cuts
+ // hit harder on a compilation whose point is repetition. Honouring it here keeps
+ // the manifest the source of truth, so a rebuild does not silently re-add fades.
+ EMIT("concat", { mode: D === 0 ? "hardcut" : "xfade", n: segments.length });
+ if (D === 0) await concatHardCut(segments, outDir, final);
+ else await concatWithXfade(segments, render, final);
+
+ if (!opts.noChapters) await muxChapters(final, entries, segments, D, outDir, provenance);
+
+ const { stdout } = await execFileP(FFPROBE, [
+ "-v", "error", "-show_entries", "format=duration,size",
+ "-of", "default=noprint_wrappers=1", final,
+ ]);
+ const probe = Object.fromEntries(
+ stdout.trim().split("\n").map((l) => l.split("=")),
+ );
+ EMIT("done", { out: final, duration: Number(probe.duration), size: Number(probe.size) });
+ return { out: final, failures };
+}
+
+async function main() {
+ const argv = process.argv.slice(2);
+ const manifestPath = argv.find((a) => !a.startsWith("--"));
+ if (!manifestPath) {
+ console.error(
+ "usage: build-video.mjs <manifest.json> [--out <dir>] [--only <id>] [--fetch-only <id>]\n" +
+ " [--pad <s>] [--skip-fetch] [--no-xfade] [--no-chapters] [--chapters-only]\n" +
+ " [--progress ndjson] [--continue-on-error] [--no-reuse]",
+ );
+ process.exit(2);
+ }
+ const flag = (n) => {
+ const i = argv.indexOf(n);
+ return i >= 0 ? argv[i + 1] : undefined;
+ };
+ setProgressMode(flag("--progress") ?? "human");
+
+ const padArg = flag("--pad");
+ const opts = {
+ skipFetch: argv.includes("--skip-fetch"),
+ continueOnError: argv.includes("--continue-on-error"),
+ noXfade: argv.includes("--no-xfade"),
+ noChapters: argv.includes("--no-chapters"),
+ chaptersOnly: argv.includes("--chapters-only"),
+ noReuse: argv.includes("--no-reuse"),
+ pad: padArg === undefined ? undefined : Number(padArg),
+ };
+
+ const { failures } = await buildVideo({
+ manifestPath,
+ opts,
+ out: flag("--out"),
+ only: flag("--only"),
+ fetchOnly: flag("--fetch-only"),
+ });
+ // Non-zero on a partial run, so a caller that ignores the events still learns
+ // the build did not produce what was asked for.
+ if (failures.length) process.exit(1);
+}
+
+if (import.meta.url === `file://${process.argv[1]}`) {
+ main().catch((err) => {
+ EMIT("error", { message: err?.message ?? String(err) });
+ console.error(err.message ?? err);
+ process.exit(1);
+ });
+}
diff --git a/scripts/report-to-video/check-availability.mjs b/scripts/report-to-video/check-availability.mjs
@@ -0,0 +1,169 @@
+#!/usr/bin/env node
+// check-availability.mjs — is every source this manifest cites still fetchable?
+//
+// This is the one fact about a report video that goes stale in BOTH directions
+// and that nothing on disk records. A source can be deleted between writing the
+// manifest and building it (so a 40-minute build dies at clip 14 having paid for
+// thirteen fetches), and a source can come back (so a manifest annotated "gone"
+// stays wrong). Neither is visible until a build runs.
+//
+// So it runs first, it runs cheap, and it writes down when it ran. `--simulate`
+// resolves formats without downloading a byte: a few seconds for a whole
+// manifest against twenty-odd minutes for the build it protects.
+//
+// On the CLI:
+// node scripts/report-to-video/check-availability.mjs <manifest.json> [--json]
+//
+// Options:
+// --json Print the report as JSON instead of a table
+// --out <dir> Output root (default: manifest dir + /out)
+// --allow-missing Exit 0 even when a source is gone (report only)
+// --max-age <days> Reuse a recorded verdict younger than this (default: 0)
+
+import { execFile } from "node:child_process";
+import { promisify } from "node:util";
+import { mkdir, readFile, writeFile } from "node:fs/promises";
+import path from "node:path";
+
+const execFileP = promisify(execFile);
+
+const YTDLP = process.env.YTDLP_BIN ?? "yt-dlp";
+const CHANNELS_DIR =
+ process.env.CHANNELS_DIR ??
+ "/home/user/Projects/yt-dlp-transcript-browser/transcripts/channels";
+
+// yt-dlp says why in prose, and the distinction matters editorially: a private
+// or removed video needs the clip converting to a quote card, while a network
+// blip needs a retry. Anything unrecognised stays `maybe_missing` rather than
+// being called deleted -- claiming a source is gone when it is not is the more
+// expensive mistake, because it invites deleting a citation.
+function classify(stderr) {
+ const t = String(stderr ?? "");
+ if (/Private video|private/i.test(t)) return "private";
+ if (/removed by the uploader|has been removed|no longer available|Video unavailable|does not exist|410/i.test(t))
+ return "deleted";
+ if (/age.?restrict|Sign in to confirm|confirm your age/i.test(t)) return "restricted";
+ if (/members-only|join this channel/i.test(t)) return "members-only";
+ if (/geo|not available in your country/i.test(t)) return "geo-blocked";
+ return "maybe_missing";
+}
+
+async function cueMeta(videoId, channelSlug) {
+ const p = path.join(CHANNELS_DIR, channelSlug, "data", videoId, "transcript.cues.json");
+ const d = JSON.parse(await readFile(p, "utf8"));
+ return { title: d.title, webpageUrl: d.webpageUrl, duration: d.duration };
+}
+
+export async function checkAvailability(manifestPath, { outDir, maxAgeDays = 0 } = {}) {
+ const manifest = JSON.parse(await readFile(manifestPath, "utf8"));
+ const slug = manifest.provenance?.channelSlug;
+ const dir = outDir ?? path.join(path.dirname(path.resolve(manifestPath)), "out");
+ const file = path.join(dir, "availability.json");
+
+ // A clip may name its own channel: the same streamer's VODs are mirrored
+ // across more than one archive, and the same id under a different slug is a
+ // different file. So the unit of work is (channel, video), never video alone.
+ const wanted = new Map();
+ for (const e of manifest.timeline ?? []) {
+ if (e.type !== "clip") continue;
+ const channel = e.channel ?? slug;
+ const key = `${channel}/${e.video}`;
+ if (!wanted.has(key)) wanted.set(key, { key, channel, video: e.video, clips: [] });
+ wanted.get(key).clips.push(e.id);
+ }
+
+ const prev = await readFile(file, "utf8").then(
+ (s) => JSON.parse(s),
+ () => ({ sources: [] }),
+ );
+ const prevBy = new Map((prev.sources ?? []).map((s) => [s.key, s]));
+ const freshMs = maxAgeDays * 86400_000;
+
+ const sources = [];
+ for (const w of wanted.values()) {
+ const was = prevBy.get(w.key);
+ if (freshMs > 0 && was?.checkedAt && Date.now() - Date.parse(was.checkedAt) < freshMs) {
+ sources.push({ ...was, clips: w.clips, reused: true });
+ continue;
+ }
+
+ let meta;
+ try {
+ meta = await cueMeta(w.video, w.channel);
+ } catch {
+ // No cue file is a DIFFERENT failure from a dead source, and it is the one
+ // the Rumble two-ids trap produces: the manifest names the MCP video id
+ // while the cues live under the URL slug. Build would die here too, so it
+ // is reported here rather than discovered twenty minutes in.
+ sources.push({
+ ...w, ok: false, state: "no-cues", checkedAt: new Date().toISOString(),
+ error: `no transcript.cues.json under ${w.channel}/data/${w.video}`,
+ });
+ continue;
+ }
+
+ try {
+ await execFileP(
+ YTDLP,
+ ["--ignore-config", "--no-playlist", "--simulate", "--quiet", "--no-warnings", "--", meta.webpageUrl],
+ { maxBuffer: 1 << 24 },
+ );
+ sources.push({
+ ...w, ok: true, state: "ok", title: meta.title, url: meta.webpageUrl,
+ checkedAt: new Date().toISOString(), error: null,
+ });
+ } catch (err) {
+ const stderr = err?.stderr ?? err?.message ?? "";
+ sources.push({
+ ...w, ok: false, state: classify(stderr), title: meta.title, url: meta.webpageUrl,
+ checkedAt: new Date().toISOString(), error: String(stderr).trim().split("\n").slice(-3).join(" "),
+ });
+ }
+ }
+
+ const report = { manifest: path.resolve(manifestPath), checkedAt: new Date().toISOString(), sources };
+ await mkdir(dir, { recursive: true });
+ await writeFile(file, JSON.stringify(report, null, 2) + "\n", "utf8");
+ return { ...report, file };
+}
+
+async function main() {
+ const argv = process.argv.slice(2);
+ const manifestPath = argv.find((a) => !a.startsWith("--"));
+ if (!manifestPath) {
+ console.error("usage: check-availability.mjs <manifest.json> [--json] [--out <dir>] [--allow-missing]");
+ process.exit(2);
+ }
+ const flag = (n) => {
+ const i = argv.indexOf(n);
+ return i >= 0 ? argv[i + 1] : undefined;
+ };
+ const report = await checkAvailability(manifestPath, {
+ outDir: flag("--out"),
+ maxAgeDays: Number(flag("--max-age") ?? 0),
+ });
+
+ if (argv.includes("--json")) {
+ console.log(JSON.stringify(report, null, 2));
+ } else {
+ for (const s of report.sources) {
+ const mark = s.ok ? "ok " : "GONE";
+ console.log(
+ `${mark} ${s.key.padEnd(40)} ${String(s.state).padEnd(14)} ` +
+ `${s.clips.length} clip(s)${s.reused ? " (cached)" : ""}`,
+ );
+ if (!s.ok && s.error) console.log(` ${s.error}`);
+ }
+ const bad = report.sources.filter((s) => !s.ok).length;
+ console.log(`\n${report.sources.length} source(s), ${bad} unavailable -> ${report.file}`);
+ }
+
+ if (report.sources.some((s) => !s.ok) && !argv.includes("--allow-missing")) process.exit(1);
+}
+
+if (import.meta.url === `file://${process.argv[1]}`) {
+ main().catch((err) => {
+ console.error(err.message ?? err);
+ process.exit(1);
+ });
+}
diff --git a/scripts/report-to-video/package.json b/scripts/report-to-video/package.json
@@ -0,0 +1,19 @@
+{
+ "name": "report-to-video",
+ "version": "0.1.0",
+ "private": true,
+ "type": "module",
+ "description": "Turn a cited sweep report into a narrated-by-text video.",
+ "bin": {
+ "report-build-video": "./build-video.mjs",
+ "report-resolve-windows": "./resolve-windows.mjs",
+ "report-check-availability": "./check-availability.mjs"
+ },
+ "exports": {
+ "./resolve-windows": "./resolve-windows.mjs",
+ "./build-video": "./build-video.mjs",
+ "./render-cards": "./render-cards.mjs",
+ "./check-availability": "./check-availability.mjs",
+ "./package.json": "./package.json"
+ }
+}
diff --git a/scripts/report-to-video/render-cards.mjs b/scripts/report-to-video/render-cards.mjs
@@ -0,0 +1,379 @@
+#!/usr/bin/env node
+// render-cards.mjs — turn a video manifest's `card` entries into PNG stills.
+//
+// One PNG per card, written to <outDir>/cards/<id>.png at the manifest's render
+// resolution. Text is laid out by ImageMagick's Pango delegate rather than
+// ffmpeg's drawtext: Pango wraps, kerns and takes inline markup, so a card is a
+// single markup string instead of a stack of hand-positioned drawtext filters.
+//
+// Card styles (manifest `style` field):
+// title — the opening card: big heading, subtitle, provenance footer
+// chapter — an act break: small amber kicker over a large heading
+// status — the bottom-line card: kicker, amber heading, subtitle
+// bullets — heading plus a list of caveats
+// sources — closing attribution
+//
+// In the app: not used. On the CLI:
+// node scripts/report-to-video/render-cards.mjs <manifest.json> [--out <dir>]
+//
+// Options:
+// --out <dir> Output root (default: the manifest's directory + /out)
+// --only <id> Render just one card, by manifest id
+//
+// Requires: ImageMagick built with Pango (magick -list format | grep PANGO).
+
+import { execFile } from "node:child_process";
+import { promisify } from "node:util";
+import { mkdir, writeFile, readFile } from "node:fs/promises";
+import path from "node:path";
+
+const execFileP = promisify(execFile);
+
+// Pango markup is XML-ish, so anything we interpolate has to be escaped first.
+// Curly quotes and the ellipsis pass through fine; only these five matter.
+function esc(s) {
+ return String(s)
+ .replace(/&/g, "&")
+ .replace(/</g, "<")
+ .replace(/>/g, ">")
+ .replace(/"/g, """)
+ .replace(/'/g, "'");
+}
+
+function span(text, { size, color, weight, family = "Fira Sans" }) {
+ const attrs = [`font_family="${family}"`, `size="${Math.round(size * 1024)}"`];
+ if (color) attrs.push(`foreground="${color}"`);
+ if (weight) attrs.push(`weight="${weight}"`);
+ return `<span ${attrs.join(" ")}>${text}</span>`;
+}
+
+// Each style returns Pango markup for the whole card body. Blank lines are real
+// newlines in the markup — Pango honours them, which is how vertical rhythm is
+// set without positioning each run separately.
+function markupFor(card, pal) {
+ const H = (t, size = 62) =>
+ span(esc(t), { size, color: pal.fg, weight: "bold" });
+ const KICKER = (t) =>
+ span(esc(t.toUpperCase()), { size: 24, color: pal.amber, weight: "bold" });
+ const SUB = (t, size = 30) => span(esc(t), { size, color: pal.muted });
+
+ switch (card.style) {
+ case "title":
+ return [
+ span(esc(card.heading), { size: 82, color: pal.fg, weight: "bold" }),
+ "",
+ span(esc(card.sub), { size: 38, color: pal.accent }),
+ "",
+ "",
+ SUB(card.foot, 24),
+ ].join("\n");
+
+ case "chapter":
+ return [
+ card.kicker ? KICKER(card.kicker) : null,
+ card.kicker ? "" : null,
+ H(card.heading),
+ card.sub ? "" : null,
+ card.sub ? SUB(card.sub, 32) : null,
+ ]
+ .filter((l) => l !== null)
+ .join("\n");
+
+ case "status":
+ return [
+ card.kicker ? KICKER(card.kicker) : null,
+ card.kicker ? "" : null,
+ span(esc(card.heading), { size: 58, color: pal.amber, weight: "bold" }),
+ card.sub ? "" : null,
+ card.sub ? SUB(card.sub, 32) : null,
+ ]
+ .filter((l) => l !== null)
+ .join("\n");
+
+ case "bullets": {
+ const items = (card.bullets ?? []).flatMap((b) => [
+ `${span("— ", { size: 30, color: pal.accent, weight: "bold" })}${span(
+ esc(b),
+ { size: 30, color: pal.fg },
+ )}`,
+ "",
+ ]);
+ return [H(card.heading, 52), "", ...items].join("\n");
+ }
+
+ case "sources":
+ return [
+ H(card.heading, 52),
+ "",
+ card.sub ? SUB(card.sub, 32) : null,
+ card.foot ? "" : null,
+ card.foot
+ ? card.foot
+ .split("\n")
+ .map((l) => SUB(l, 24))
+ .join("\n")
+ : null,
+ ]
+ .filter((l) => l !== null)
+ .join("\n");
+
+ default:
+ return H(card.heading ?? card.id);
+ }
+}
+
+// A chapter break that just states a heading is dead air — it stops the video to
+// say something the next clip is about to say anyway. This draws the whole
+// project arc instead, with the current step lit and everything before it
+// filled, so the pause carries information: where we are and how far is left.
+//
+// Nodes come from the manifest's top-level `timelineNodes`; the card names its
+// position with `step` (0-based).
+async function renderTimelineCard(card, render, nodes, outDir) {
+ const pal = render.palette;
+ const { width, height } = render;
+ const outPath = path.join(outDir, "cards", `${card.id}.png`);
+ const dir = path.join(outDir, "cards");
+
+ const x0 = 260;
+ const x1 = width - 260;
+ const axisY = Math.round(height * 0.56);
+ const gap = (x1 - x0) / (nodes.length - 1);
+ const xs = nodes.map((_, i) => Math.round(x0 + i * gap));
+ const cur = card.step;
+
+ const args = ["-size", `${width}x${height}`, `xc:${pal.bg}`, "-strokewidth", "4"];
+
+ // Track: filled up to the current node, dim beyond it.
+ args.push(
+ "-stroke", pal.muted, "-fill", "none",
+ "-draw", `line ${xs[0]},${axisY} ${xs[xs.length - 1]},${axisY}`,
+ );
+ if (cur > 0) {
+ args.push("-stroke", pal.accent, "-draw", `line ${xs[0]},${axisY} ${xs[cur]},${axisY}`);
+ }
+
+ // Nodes: past and present filled, future hollow. The current one is larger and
+ // amber so the eye lands on it without needing a label to say "you are here".
+ nodes.forEach((_, i) => {
+ const r = i === cur ? 19 : 11;
+ const color = i === cur ? pal.amber : i < cur ? pal.accent : pal.bg;
+ args.push(
+ "-stroke", i <= cur ? (i === cur ? pal.amber : pal.accent) : pal.muted,
+ "-fill", color,
+ "-draw", `circle ${xs[i]},${axisY} ${xs[i] + r},${axisY}`,
+ );
+ });
+
+ args.push("-stroke", "none");
+
+ // Heading, centred over the whole card.
+ const headMarkup = [
+ span(esc((card.kicker ?? nodes[cur].label).toUpperCase()), {
+ size: 26, color: pal.amber, weight: "bold",
+ }),
+ "",
+ span(esc(card.heading ?? nodes[cur].title), { size: 62, color: pal.fg, weight: "bold" }),
+ card.sub ? "" : null,
+ card.sub ? span(esc(card.sub), { size: 30, color: pal.muted }) : null,
+ ]
+ .filter((l) => l !== null)
+ .join("\n");
+
+ // Heading, left-aligned on the same margin the other card styles use.
+ const headPath = path.join(dir, `${card.id}.head.pango`);
+ await writeFile(headPath, headMarkup, "utf8");
+ args.push(
+ "(", "-size", `${width - 460}x`, "-background", "none",
+ "-define", `pango:width=${width - 460}`,
+ `pango:@${headPath}`, ")",
+ "-gravity", "NorthWest",
+ "-geometry", `+${Math.round(width * 0.09) + 58}+${Math.round(height * 0.19)}`,
+ "-composite",
+ );
+
+ // Per-node date labels, centred under their dot. ImageMagick's
+ // `pango:alignment` define does not actually centre the text inside the box,
+ // so measure each rendered label and place it by hand instead of trusting it.
+ for (let i = 0; i < nodes.length; i += 1) {
+ const isCur = i === cur;
+ const labMarkup = span(esc(nodes[i].label), {
+ size: isCur ? 24 : 21,
+ color: isCur ? pal.fg : i < cur ? pal.muted : "#5c5570",
+ weight: isCur ? "bold" : "normal",
+ });
+ const labPath = path.join(dir, `${card.id}.n${i}.pango`);
+ const labPng = path.join(dir, `${card.id}.n${i}.png`);
+ await writeFile(labPath, labMarkup, "utf8");
+ await execFileP("magick", ["-background", "none", `pango:@${labPath}`, labPng]);
+ const { stdout } = await execFileP("magick", ["identify", "-format", "%w", labPng]);
+ const w = Number(stdout.trim());
+ args.push(labPng, "-geometry", `+${xs[i] - Math.round(w / 2)}+${axisY + 44}`, "-composite");
+ }
+
+ args.push(outPath);
+ await execFileP("magick", args, { maxBuffer: 1 << 24 });
+ return outPath;
+}
+
+// Footer chrome, drawn once and overlaid on every clip: a track, a dot per
+// section and its label. The *progress* along it is not baked in here — the fill
+// bar and the amber marker are drawn by ffmpeg at encode time so they can slide
+// between sections instead of cutting. See buildClipSegment in build-video.mjs.
+//
+// Returns the geometry the encoder needs to place those moving parts.
+export async function renderFooterAssets(render, nodes, outDir) {
+ const pal = render.palette;
+ const { width } = render;
+ const FH = render.footerHeight ?? 92;
+ const dir = path.join(outDir, "cards");
+
+ // A cut whose clips are not a progression through time has nothing for a
+ // timeline to say, and a footer drawn anyway is chrome that has not earned its
+ // place. No nodes (or an explicit zero height) means no footer at all — the
+ // caller letterboxes against the header alone.
+ if (!nodes?.length || FH === 0) {
+ return { footer: null, marker: null, footerHeight: 0, trackY: 0, xs: [], x0: 0, markerRadius: 0 };
+ }
+
+ const x0 = 200;
+ const x1 = width - 200;
+ const trackY = 26;
+ const gap = (x1 - x0) / (nodes.length - 1);
+ const xs = nodes.map((_, i) => Math.round(x0 + i * gap));
+
+ const footer = path.join(dir, "_footer.png");
+ const args = [
+ "-size", `${width}x${FH}`, `xc:${pal.bg}`,
+ "-strokewidth", "3",
+ "-stroke", "#3a3450", "-fill", "none",
+ "-draw", `line ${xs[0]},${trackY} ${xs[xs.length - 1]},${trackY}`,
+ "-stroke", "none",
+ ];
+ for (const x of xs) {
+ args.push("-fill", "#4a4363", "-draw", `circle ${x},${trackY} ${x + 6},${trackY}`);
+ }
+
+ // Two lines per node: what happened, then when. Both are measured and placed
+ // by hand — ImageMagick's `pango:alignment` define does not actually centre
+ // text inside its box.
+ for (let i = 0; i < nodes.length; i += 1) {
+ const lines = [
+ { text: nodes[i].label, size: 18, color: pal.fg, dy: 18 },
+ { text: nodes[i].date, size: 16, color: pal.muted, dy: 42 },
+ ];
+ for (const [k, ln] of lines.entries()) {
+ const pPath = path.join(dir, `_footer.n${i}.l${k}.pango`);
+ const pPng = path.join(dir, `_footer.n${i}.l${k}.png`);
+ await writeFile(pPath, span(esc(ln.text), { size: ln.size, color: ln.color }), "utf8");
+ await execFileP("magick", ["-background", "none", `pango:@${pPath}`, pPng]);
+ const { stdout } = await execFileP("magick", ["identify", "-format", "%w", pPng]);
+ args.push(
+ pPng,
+ "-geometry", `+${xs[i] - Math.round(Number(stdout.trim()) / 2)}+${trackY + ln.dy}`,
+ "-composite",
+ );
+ }
+ }
+ args.push(footer);
+ await execFileP("magick", args, { maxBuffer: 1 << 24 });
+
+ // The travelling marker.
+ const marker = path.join(dir, "_marker.png");
+ const r = 11;
+ await execFileP("magick", [
+ "-size", `${r * 2 + 2}x${r * 2 + 2}`, "xc:none",
+ "-fill", pal.amber, "-stroke", "none",
+ "-draw", `circle ${r + 1},${r + 1} ${r * 2 + 1},${r + 1}`,
+ marker,
+ ]);
+
+ return { footer, marker, footerHeight: FH, trackY, xs, x0, markerRadius: r };
+}
+
+export async function renderCard(card, render, outDir, nodes) {
+ if (card.style === "timeline") {
+ if (!nodes?.length) throw new Error(`card ${card.id} is style:timeline but no timelineNodes given`);
+ return renderTimelineCard(card, render, nodes, outDir);
+ }
+ return renderPlainCard(card, render, outDir);
+}
+
+async function renderPlainCard(card, render, outDir) {
+ const pal = render.palette;
+ const { width, height } = render;
+ const textWidth = Math.round(width * 0.74);
+ const outPath = path.join(outDir, "cards", `${card.id}.png`);
+
+ // Pango reads its markup from a file to keep it clear of shell/argv quoting.
+ const markupPath = path.join(outDir, "cards", `${card.id}.pango`);
+ await writeFile(markupPath, markupFor(card, pal), "utf8");
+
+ // One magick invocation: solid ground, an accent rule down the left margin,
+ // then the Pango block composited over it. The rule is what keeps the cards
+ // recognisably one family across styles.
+ const barX = Math.round(width * 0.09);
+ const barTop = Math.round(height * 0.28);
+ const barBottom = Math.round(height * 0.72);
+
+ const args = [
+ "-size", `${width}x${height}`,
+ `xc:${pal.bg}`,
+ "-fill", pal.accent,
+ "-draw", `rectangle ${barX},${barTop} ${barX + 6},${barBottom}`,
+ "(",
+ // `-size` is still set to the full frame from the canvas above, and the
+ // pango delegate honours it — leaving it alone renders the text into a
+ // 1920x1080 box, which pins the block to the top and wraps at the frame
+ // edge instead of the margin. Reset it to the text column, height auto.
+ "-size", `${textWidth}x`,
+ "-background", "none",
+ "-define", `pango:width=${textWidth}`,
+ "-define", "pango:alignment=left",
+ "-define", "pango:wrap=word",
+ `pango:@${markupPath}`,
+ ")",
+ "-gravity", "West",
+ "-geometry", `+${barX + 58}+0`,
+ "-composite",
+ outPath,
+ ];
+
+ await execFileP("magick", args, { maxBuffer: 1 << 24 });
+ return outPath;
+}
+
+async function main() {
+ const argv = process.argv.slice(2);
+ const manifestPath = argv.find((a) => !a.startsWith("--"));
+ if (!manifestPath) {
+ console.error("usage: render-cards.mjs <manifest.json> [--out <dir>] [--only <id>]");
+ process.exit(2);
+ }
+ const flag = (name) => {
+ const i = argv.indexOf(name);
+ return i >= 0 ? argv[i + 1] : undefined;
+ };
+
+ const manifest = JSON.parse(await readFile(manifestPath, "utf8"));
+ const outDir = flag("--out") ?? path.join(path.dirname(path.resolve(manifestPath)), "out");
+ const only = flag("--only");
+
+ await mkdir(path.join(outDir, "cards"), { recursive: true });
+
+ const cards = manifest.timeline.filter(
+ (e) => e.type === "card" && (!only || e.id === only),
+ );
+ for (const card of cards) {
+ const p = await renderCard(card, manifest.render, outDir, manifest.timelineNodes);
+ console.log(`card ${card.id} -> ${p}`);
+ }
+ console.log(`${cards.length} card(s) rendered`);
+}
+
+if (import.meta.url === `file://${process.argv[1]}`) {
+ main().catch((err) => {
+ console.error(err);
+ process.exit(1);
+ });
+}
diff --git a/scripts/report-to-video/resolve-windows.mjs b/scripts/report-to-video/resolve-windows.mjs
@@ -0,0 +1,220 @@
+#!/usr/bin/env node
+// resolve-windows.mjs — widen a manifest's clip windows to whole sentences.
+//
+// A manifest window starts life as the cue span covering a quote, and a cue
+// boundary is a bad place to cut: ASR breaks cues where the caption line wrapped,
+// which is routinely mid-sentence and often mid-word. Cutting there drops the
+// lead-in that makes a quote make sense, and clips audibly start and stop in the
+// middle of speech.
+//
+// This walks outward from the cue span to the nearest sentence boundary in the
+// transcript — a cue whose text ends in . ? or ! — so the clip carries the whole
+// thought. Word-level alignment is a separate, audio-side problem: build-video.mjs
+// snaps the actual cut to a silence (see --fetch-pad / snapping there).
+//
+// Expansion is capped so a run-on passage can't drag a clip out to a minute.
+//
+// In the app: not used. On the CLI:
+// node scripts/report-to-video/resolve-windows.mjs <manifest.json> [--write]
+//
+// Options:
+// --write Rewrite the manifest in place (default: dry run, print a table)
+// --max-lead <s> Max seconds to expand backwards (default 9)
+// --max-tail <s> Max seconds to expand forwards (default 12)
+//
+// A clip entry may set `lockStart` / `lockEnd` to pin that edge exactly.
+
+import { readFile, writeFile } from "node:fs/promises";
+import path from "node:path";
+
+const CHANNELS_DIR =
+ process.env.CHANNELS_DIR ??
+ "/home/user/Projects/yt-dlp-transcript-browser/transcripts/channels";
+
+const ENDS_SENTENCE = /[.!?]["'”’)\]]*\s*$/;
+
+// A cue that is only "[music]" or "[ __ ]" (the profanity bleep) carries no
+// sentence signal; treat it as transparent so expansion walks past it.
+const IS_FILLER = /^\s*(\[[^\]]*\]|>>|♪|—|-)*\s*$/;
+
+// The manifest stores times rounded to 2 dp, so a value read back from it can sit
+// a hair BELOW the cue end it came from. Without a tolerance the end lookup then
+// lands on the previous cue, the forward search runs on to the next sentence, and
+// the clip grows a little every time this is run — it has to be a fixed point.
+const EPS = 0.02;
+
+function loadCues(videoId, channelSlug) {
+ const p = path.join(CHANNELS_DIR, channelSlug, "data", videoId, "transcript.cues.json");
+ return readFile(p, "utf8").then((s) => JSON.parse(s).cues);
+}
+
+function indexAt(cues, t, which) {
+ // First cue whose span contains t, else the nearest one on the right side.
+ let idx = cues.findIndex((c) => c.end > t);
+ if (idx < 0) idx = cues.length - 1;
+ if (which === "end") {
+ let j = cues.findIndex((c) => c.end >= t - EPS);
+ if (j < 0) j = cues.length - 1;
+ idx = j;
+ }
+ return idx;
+}
+
+export function widen(cues, start, end, { maxLead = 8, maxTail = 12 } = {}) {
+ const isBoundary = (c) => ENDS_SENTENCE.test(c.text) && !IS_FILLER.test(c.text);
+ const i0 = indexAt(cues, start, "start");
+ const i1 = indexAt(cues, end, "end");
+
+ // START: the latest cue that OPENS a sentence (i.e. its predecessor closes
+ // one) at or before the quote, within the lead budget. Finding no such cue
+ // means every candidate lead-in is a sentence fragment, so take none at all —
+ // a fragment is the irrelevant context we are trying to avoid, not context.
+ let si = null;
+ for (let i = i0; i > 0; i -= 1) {
+ if (start - cues[i].start > maxLead) break;
+ if (isBoundary(cues[i - 1])) {
+ si = i;
+ break;
+ }
+ }
+ if (si === null) si = i0;
+
+ // END: the first cue that CLOSES a sentence at or after the quote. Never
+ // clamp to a budget here — stopping partway through a sentence is exactly the
+ // mid-thought ending this is meant to remove, so the budget only decides how
+ // far to look, and failing to find one falls back to the original cue end.
+ let ei = null;
+ for (let j = i1; j < cues.length; j += 1) {
+ if (cues[j].end - end > maxTail) break;
+ if (isBoundary(cues[j])) {
+ ei = j;
+ break;
+ }
+ }
+ if (ei === null) ei = i1;
+
+ return {
+ start: cues[si].start,
+ end: cues[ei].end,
+ leadCues: i0 - si,
+ tailCues: ei - i1,
+ };
+}
+
+async function main() {
+ const argv = process.argv.slice(2);
+ const manifestPath = argv.find((a) => !a.startsWith("--"));
+ if (!manifestPath) {
+ console.error("usage: resolve-windows.mjs <manifest.json> [--write]");
+ process.exit(2);
+ }
+ const num = (name, dflt) => {
+ const i = argv.indexOf(name);
+ return i >= 0 ? Number(argv[i + 1]) : dflt;
+ };
+ // Lead is where the context lives — it is the run-up that makes a quote make
+ // sense. Tail only needs to finish the sentence, so it gets a smaller budget.
+ const opts = { maxLead: num("--max-lead", 8), maxTail: num("--max-tail", 12) };
+
+ const manifest = JSON.parse(await readFile(manifestPath, "utf8"));
+ const slug = manifest.provenance.channelSlug;
+ const cache = new Map();
+
+ let changed = 0;
+ for (const e of manifest.timeline) {
+ if (e.type !== "clip") continue;
+ // A compilation can span several archived channels (the same streamer's VODs
+ // are mirrored across more than one), so a clip may name its own. Key the
+ // cache by channel too — the same id under a different slug is a different file.
+ // An author can trim a clip to land mid-cue on purpose — a cue often carries
+ // a whole paragraph, and cutting a quote short is an editorial decision.
+ // Widening would undo exactly that, so `lock` opts the clip out.
+ if (e.lock) {
+ console.log(`${e.id.padEnd(4)} ${e.video.padEnd(12)} locked, left at ${e.start.toFixed(1)}–${e.end.toFixed(1)}`);
+ continue;
+ }
+ const chan = e.channel ?? slug;
+ const key = `${chan}/${e.video}`;
+ if (!cache.has(key)) cache.set(key, await loadCues(e.video, chan));
+ const cues = cache.get(key);
+
+ const before = { start: e.start, end: e.end };
+ const w = widen(cues, e.start, e.end, opts);
+
+ // `lockStart` / `lockEnd` pin an edge to exactly what the author wrote. The
+ // escape hatch exists because sentence detection is only as good as the ASR's
+ // punctuation, and some uploads have none at all — and because an utterance's
+ // real trailing pause does not always line up with its last cue's end.
+ if (e.lockStart) w.start = before.start;
+ if (e.lockEnd) w.end = before.end;
+ const dLead = (before.start - w.start).toFixed(1);
+ const dTail = (w.end - before.end).toFixed(1);
+ const dur = (w.end - w.start).toFixed(1);
+
+ // Ignore sub-frame drift so a re-run on an already-resolved manifest is a
+ // genuine no-op rather than a rewrite that nudges every window.
+ const moved =
+ Math.abs(w.start - before.start) > 0.05 || Math.abs(w.end - before.end) > 0.05;
+ if (moved) changed += 1;
+ console.log(
+ `${e.id.padEnd(4)} ${e.video.padEnd(12)} ` +
+ `${before.start.toFixed(1)}–${before.end.toFixed(1)} -> ` +
+ `${w.start.toFixed(1)}–${w.end.toFixed(1)} (+${dLead}s lead, +${dTail}s tail, ${dur}s)`,
+ );
+
+ if (moved) {
+ e.start = Number(w.start.toFixed(2));
+ e.end = Number(w.end.toFixed(2));
+ }
+ }
+
+ // De-overlap clips that come from the SAME video. Widening is per-clip and
+ // blind to its neighbours, so a tail that finds no sentence boundary runs to
+ // the budget and can swallow the next clip's material — which plays as the
+ // same footage twice. (Real case: a 2024 upload whose ASR carries no
+ // punctuation at all in that stretch, so nothing stopped the search.)
+ // The later clip's start is the deliberate one, so trim the earlier clip's tail.
+ const byVideo = new Map();
+ for (const e of manifest.timeline) {
+ if (e.type !== "clip") continue;
+ if (!byVideo.has(e.video)) byVideo.set(e.video, []);
+ byVideo.get(e.video).push(e);
+ }
+ for (const [video, list] of byVideo) {
+ if (list.length < 2) continue;
+ list.sort((a, b) => a.start - b.start);
+ for (let i = 0; i < list.length - 1; i += 1) {
+ const a = list[i];
+ const b = list[i + 1];
+ if (a.end <= b.start) continue;
+ const overlap = a.end - b.start;
+ if (a.lockEnd) {
+ console.log(` ⚠ ${a.id} overlaps ${b.id} by ${overlap.toFixed(1)}s but has lockEnd — not trimmed`);
+ continue;
+ }
+ a.end = Number(b.start.toFixed(2));
+ changed += 1;
+ console.log(
+ ` de-overlap ${video}: ${a.id} trimmed ${overlap.toFixed(1)}s off its tail ` +
+ `(it ran into ${b.id})`,
+ );
+ if (a.end - a.start < 3) {
+ console.log(` ⚠ ${a.id} is now only ${(a.end - a.start).toFixed(1)}s — check it`);
+ }
+ }
+ }
+
+ if (argv.includes("--write")) {
+ await writeFile(manifestPath, JSON.stringify(manifest, null, 2) + "\n", "utf8");
+ console.log(`\nwrote ${manifestPath} (${changed} window(s) changed)`);
+ } else {
+ console.log(`\ndry run — ${changed} window(s) would change; pass --write to apply`);
+ }
+}
+
+if (import.meta.url === `file://${process.argv[1]}`) {
+ main().catch((err) => {
+ console.error(err);
+ process.exit(1);
+ });
+}
diff --git a/umtool/docs/README.md b/umtool/docs/README.md
@@ -0,0 +1,68 @@
+# umtool docs
+
+umtool judges and drives the things this repo makes videos out of. There are two
+kinds of work in it today and the tool treats them the same way: as **projects**
+in a folder tree, each with a state, a set of open decisions, and a build.
+
+These sheets live in the repo because they describe code and have to move with it.
+Notes about *one particular video* belong beside that video, in its own `README.md`.
+
+## You have been asked for an umtool video
+
+Work out which kind you are making first — the rest follows from it.
+
+| What you were handed | Kind | Start here |
+|---|---|---|
+| A cited sweep report (`*sweep-report.md`) and "make this a video" | `report-video` | [report-video.md](report-video.md), then [authoring.md](authoring.md) |
+| A song's clips and "judge these" | `song` (template `um-song`) | `~/reports/quartering-uh-song/specs/` |
+| A report with no manifest yet | `sweep-report` | [authoring.md](authoring.md) |
+
+The short path from a cited report to a built video:
+
+```sh
+# 1. write video.manifest.json beside the report (authoring.md — this is the work)
+# 2. is every source still fetchable, and is every citation wired up?
+umtool check ~/reports/<slug>
+# 3. widen windows to whole sentences (dry first, then apply)
+node scripts/report-to-video/resolve-windows.mjs ~/reports/<slug>/video.manifest.json
+node scripts/report-to-video/resolve-windows.mjs ~/reports/<slug>/video.manifest.json --write
+# 4. a fast pass to look at, then the real one
+umtool build <slug> --preset fast
+umtool build <slug> --preset final
+```
+
+**Run step 2 before step 4, always.** It is a few seconds and it catches the two
+defects that have already shipped in real videos: a manifest with no `siteOrigin`
+(19 QR codes encoding `undefined/?v=…`) and one pointing at `http://localhost:3000`
+(QR codes that resolve to nothing on anybody's phone).
+
+## The sheets
+
+| Sheet | What it covers |
+|---|---|
+| [projects.md](projects.md) | Kinds vs templates, marker files, ids, how to add a kind |
+| [folders.md](folders.md) | The walk, `REPORTS_ROOT`, read roots vs write roots |
+| [browse.md](browse.md) | The project index, the four filters, cards per kind |
+| [report-video.md](report-video.md) | The manifest as an EDL, the three-stage window model, `lock` |
+| [clip-bench.md](clip-bench.md) | Editing a clip's window against the waveform and the cues |
+| [build.md](build.md) | The four-step chain, presets, cancellation, overwrite |
+| [decisions.md](decisions.md) | What earns a severity, and how a kind contributes |
+| [mix-from-a-project.md](mix-from-a-project.md) | Deep-linking a clip into `/mix` |
+| [index.md](index.md) | The LMDB index, and why the filesystem stays the model |
+| [cli.md](cli.md) | `umtool ls / show / check / build / window / …` |
+| [authoring.md](authoring.md) | Writing a manifest from a sweep report |
+| [e2e.md](e2e.md) | The fixture, the stubs, the global queue |
+| [quirks.md](quirks.md) | Everything that cost time to find out |
+
+## The rules that outrank convenience
+
+1. **The filesystem is the model.** An index may cache what the tree says; if a
+ value exists *only* in the index, that is a bug. See [index.md](index.md).
+2. **The index must not shell out.** Listing projects never probes, never runs
+ ffprobe, never runs yt-dlp. Measuring is what a project page and a job do.
+3. **Judgements travel with the tree.** `verdicts.json`, `notes.json` and
+ `video.manifest.json` live *inside* the project directory, so copying the
+ directory copies the decisions.
+4. **Never guess at something you cannot read.** A directory whose name does not
+ route, a project two kinds match, a citation with no cue file — each is
+ *reported*, never silently dropped or resolved by picking one.
diff --git a/umtool/docs/quirks.md b/umtool/docs/quirks.md
@@ -0,0 +1,129 @@
+# Quirks
+
+Things that cost time to find out. Each one is here because it was discovered by
+getting it wrong, and none of them is guessable from the code.
+
+## Fetching source clips
+
+**yt-dlp picks VP9 + Opus at these heights unless you pin the format.** Since
+`--force-keyframes-at-cuts` re-encodes, that means `libvpx-vp9`: 27 seconds to cut
+a 5-second clip. It also writes `.webm` and appends that to `-o`, so the file you
+asked for is not the file on disk. Pin `bv*[vcodec^=avc1][height<=N]+ba[acodec^=mp4a]`
+and `--merge-output-format mp4`.
+
+**`--ignore-config` is not optional.** The operator's own yt-dlp config redirects
+output and attaches thumbnail and metadata post-processors. Without it the clips
+land somewhere else entirely and the build reports success having produced nothing
+where it was looking.
+
+**yt-dlp exit 101 is success.** It is the clean early stop (`break-on-existing`,
+`--max-downloads`). Treat it as success, as the rest of the repo does.
+
+**Rumble HLS needs `-extension_picky 0`, as a RETRY and never as a default.**
+Rumble serves HLS whose segments are named `.tar`, which ffmpeg 8 rejects outright
+("URL … is not in allowed_segment_extensions"), killing the fetch with exit 183 —
+and Rumble ships no progressive fallback, so every Rumble clip is unbuildable
+without it. But the option lives on the **HLS demuxer**: pass it against a
+progressive URL (YouTube's googlevideo mp4) and ffmpeg aborts with "Option
+extension_picky not found". Adding it unconditionally trades a Rumble failure for
+a YouTube one.
+
+**`--force-keyframes-at-cuts` matters because the clip IS the citation.** Without
+it the cut snaps to the nearest preceding keyframe, which can be seconds early.
+Fine for scrubbing; not fine when someone is checking your quote.
+
+## Cutting
+
+**The silence threshold has to be relative to the clip, not absolute.** These are
+game streams: the gaps between words are full of game audio and music. Measured on
+a typical clip — mean volume −21 dB, **0** silences found at −32 dB, **25** at
+−26 dB. Measure with `volumedetect` first and cut a few dB under the clip's own
+mean.
+
+**Snapping only proves the edges are quiet.** It finds gaps in audio, which is
+usually a word boundary and is not guaranteed to be. A speaker who does not pause
+gets the unsnapped cut.
+
+**ASR cue boundaries are line-wrap boundaries, not sentence boundaries.** They land
+mid-sentence routinely and mid-word often. That is the whole reason
+`resolve-windows.mjs` exists — and the reason a clip bench that shows you the cues
+is worth more than one that only shows you a waveform.
+
+**Some uploads carry no punctuation at all.** Sentence widening then has nothing to
+find and silently does nothing. The clip bench says so explicitly rather than
+leaving you to wonder why "extend to sentence end" is inert; set the edge by ear
+and lock it.
+
+## Manifests
+
+**`resolve-windows.mjs` is a fixed point, and that was not free.** The manifest
+stores times at 2 dp, so a value read back can sit a hair below the cue end it came
+from — which lands the end lookup on the *previous* cue, runs the forward search on
+to the *next* sentence, and grows the same clip a little on every run. Hence `EPS`
+in the lookup and the 0.05 s deadband on applying a change. **Anything that writes a
+window must round to 2 dp**, or it reintroduces exactly this bug.
+
+**`writeJsonAtomic` would vandalise a manifest.** It writes `JSON.stringify(v, null, 1)`;
+`resolve-windows.mjs` writes `null, 2` plus a trailing newline. Saving one window
+edit through the default writer reformats 600 lines and makes the diff unreadable.
+Manifest writes pass `{space: 2, newline: true}`.
+
+**`lock: true` is the norm, not the exception.** Measured across the six real
+manifests: ferret-rescue locks 1 of 10, every other manifest locks **100%**. A
+human-chosen window usually *is* the truth, and widening it would undo an editorial
+decision — cutting a quote short is a choice, and a single ASR cue often carries a
+whole paragraph. So `lock` is a first-class explained control in the bench, not an
+advanced toggle, and moving an edge to somewhere `widen()` would not produce offers
+to set the matching lock.
+
+**`siteOrigin` is unvalidated and has already shipped broken twice.**
+`quartering-employee-count` has none — 19 QR codes encoding `undefined/?v=…` — and
+`ferret-rescue` has `http://localhost:3000`, a shipped video whose codes resolve to
+nothing on anyone's phone. `umtool check` exists largely for this.
+
+**A clip may name its own `channel`.** The same streamer's VODs are mirrored across
+several archived channels, and cue files are keyed by channel, so one manifest-wide
+slug cannot find them all.
+
+**Rumble ids: the manifest's `video` must be the local directory slug.** Cue files
+live under the URL slug, not the MCP video id. A manifest using the id finds
+nothing — and finds it twenty minutes into a build, unless `umtool check` ran first.
+
+## Rendering
+
+**ImageMagick's `-size` leaks into Pango.** It applies to the *next* image
+operation, so a stale `-size` silently changes the text raster.
+
+**drawtext does not wrap, and commas are structural in a filtergraph.** Wrap to a
+character budget, write the text to a file and use `textfile=`, so nothing needs
+shell or filter escaping. Single-quote any expression containing a comma.
+
+**Stream titles need emoji and `!command` suffixes stripped** or they render as
+tofu in the attribution line.
+
+**Segments are encoded to identical parameters on purpose**, so the final concat is
+a stream copy. Mismatched streams are the usual reason a naive concat produces a
+broken or audio-desynced file.
+
+**A QR must be fully opaque and must keep its quiet zone.** A translucent QR will
+not scan, and the white border is part of the symbol, not decoration.
+
+**ffmetadata is line-based and `=`, `;`, `#`, `\` are structural.** A chapter title
+carrying any of them has to be escaped or the file silently mis-parses.
+
+## The tool itself
+
+**`node -e console.log(<number>)` emits ANSI escapes on a TTY**, which corrupt
+ffmpeg filtergraphs and shell tests silently.
+
+**`pnpm lint` in `editor/` always fails** — there is no eslint config there. Use
+`pnpm exec tsc --noEmit`.
+
+**A value import that drags `node:fs` into a client component 500s every page** and
+passes typecheck. `pnpm build`, not just `tsc --noEmit`, is what catches it — which
+is why the registry's *types* live in `lib/project-types.ts`, separate from the
+`.mjs` that reads the disk.
+
+**e2e is serialized machine-wide.** A "waiting for the e2e queue" banner is normal,
+not a hang; the serial suite is long. A run that wins the lock but finds its ports
+bound aborts and names the offending pid.
diff --git a/umtool/docs/report-video.md b/umtool/docs/report-video.md
@@ -0,0 +1,160 @@
+# The `report-video` kind
+
+A report video is a **cited timeline**: a sweep report's findings, said by the
+sources in their own voice, in order, each clip carrying a burned-in attribution
+and a QR that resolves to that exact moment in the archive's own viewer. The point
+is that a viewer does not have to take the edit on trust.
+
+The pipeline lives at `scripts/report-to-video/` and its own
+[README](../../scripts/report-to-video/README.md) is the source of truth for how it
+renders. This sheet covers what that README cannot know: what umtool reads, what
+umtool **writes**, and the rules a UI has to honour so the CLI and the app can never
+disagree.
+
+## The manifest is an EDL, and it is the project
+
+`<project>/video.manifest.json` is both the marker file for the kind and the edit
+decision list. Nothing else needs to exist for a directory to be a report video.
+
+```jsonc
+{
+ "schemaVersion": 1,
+ "slug": "quartering-gout", // out/<slug>.mp4 is the deliverable
+ "title": "…", "subtitle": "…", "generatedOn": "2026-08-12",
+ "provenance": { … }, // how the sweep was done, and its caveats
+ "render": { … }, // resolution, fonts, palette, the cut knobs
+ "timelineNodes": [ … ], // the footer's progress track, if any
+ "timeline": [ … ] // the ordered cut. Array order IS the cut.
+}
+```
+
+**Ordering is array order.** There is no `index` field and no sort. A reorder is a
+move within the array, and it has to recompute `sectionEnter` — the flag that says
+"this clip is the first of its section", which is what makes the footer marker
+travel. umtool never reorders as a side effect of a window edit.
+
+### A clip entry
+
+```jsonc
+{ "type": "clip", "id": "c04", "video": "uyz1_FIqIEk",
+ "channel": "hasanabi-vods3", // optional: which archived channel's cues
+ "start": 32980.24, "end": 32994.19, // absolute source seconds, 2 dp
+ "cite": 32989, // the second shown in the attribution line
+ "citeUrl": "https://…", // optional: overrides the derived QR target
+ "quote": "…", // the words this clip exists for
+ "note": "…", // why it is in the cut (editorial, for humans)
+ "chapter": "…", // chapter title; falls back to date + title
+ "section": 2, "sectionEnter": true,
+ "lock": true, "lockStart": true, "lockEnd": true }
+```
+
+A card entry is `{"type":"card", "id", "style", "seconds", …}` — see the pipeline
+README for the styles. Cards have no window and no source.
+
+## The three-stage window model
+
+A clip's window passes through three different notions of "where the cut is", and
+confusing them is the source of most window bugs.
+
+| Stage | Where it comes from | Whose problem |
+|---|---|---|
+| **1. Cue span** | `transcript.cues.json` — the cues covering the quote | The sweep. A report only records a single start second, so a window cannot be recovered from the report alone. |
+| **2. Sentence window** | `resolve-windows.mjs` walks outward to a cue ending in `.?!` | Meaning. A cue boundary is a *line-wrap* boundary; cutting there drops the lead-in that makes a quote make sense. |
+| **3. Audio cut** | `build-video.mjs` snaps to a silence found in the fetched audio | Sound. A sentence boundary is still not a *speech* boundary; snapping is what stops clips cutting through a word. |
+
+Only stage 2 is stored. Stage 1 is recoverable from the cue file, and stage 3 is
+recomputed on every build from the audio — which is why a build is reproducible
+from the manifest alone, and why the clip bench edits stage 2 and shows you stage 1.
+
+**Everything is in absolute source seconds**, from the manifest through the bench
+to the API. The fetched file's own start (`fetchStart`) is the only place a relative
+number appears, and it is always derived, never stored.
+
+## Writing a window: the four rules
+
+umtool is a *second* writer of a file the CLI also writes. So:
+
+1. **Round to 2 dp.** Non-negotiable. `resolve-windows.mjs` is a fixed point and
+ both its `EPS` lookup tolerance and its 0.05 s deadband assume 2 dp storage.
+ Writing 4 dp makes the widener grow the clip on every subsequent run.
+2. **Re-read before writing, under `withStateLock`, and write atomically.** The
+ file has two writers; a lost update here is lost human judgement.
+3. **Preserve the CLI's formatting** — `JSON.stringify(m, null, 2) + "\n"`. The
+ default `writeJsonAtomic` uses indent 1, which turns a two-number edit into a
+ 600-line diff.
+4. **Guard with an mtime token.** A `PUT` carrying a stale token is a 409, not a
+ silent overwrite — an agent or a CLI run may have written in between.
+
+A `.bak` of the last hand-authored state is kept (rate-limited, not a rolling
+stack); `ferret-rescue/video.manifest.json.bak` already set that precedent.
+
+## `lock`, `lockStart`, `lockEnd`
+
+These are **editorial acknowledgements**, not advanced options, and the data says
+so: across the six real manifests ferret-rescue locks 1 clip of 10 and *every other
+manifest locks 100%*.
+
+- **`lock`** — resolve-windows must not touch this clip at all. The two cases:
+ the author deliberately cut a quote short (a single ASR cue often carries a whole
+ paragraph, so trimming to the sentence that matters is a real decision that
+ widening would undo), or the lead-in would drag in seconds of some *other* audio.
+ A live case from `quartering-gout`: two clips whose source cue file has two cues
+ spanning almost the whole runtime, where an unlocked widening pass would expand a
+ 35-second window to a 2,488-second one.
+- **`lockStart` / `lockEnd`** — pin one edge exactly. Needed because sentence
+ detection is only as good as the ASR's punctuation, and some uploads have none.
+
+`lockEnd` is also how the "ends mid-sentence" warning is *acknowledged*: setting it
+is the author saying "I meant to cut here", and it suppresses the warning. This
+matters — one 14-clip cut shipped with 8 clips ending mid-thought.
+
+**The rule that prevents silent loss:** if you move an edge in the bench to a value
+`widen()` would not produce, the bench offers to set the matching lock, defaulted
+on. Without it, the next `resolve-windows --write` reverts your edit.
+
+## Provenance, and the two defects that shipped
+
+`provenance` is prose written by the sweep, and most of it is for humans. Three
+fields are load-bearing:
+
+- **`siteOrigin`** — the archive origin every QR is built from. **Nothing validated
+ this**, and two real videos shipped broken: `quartering-employee-count` has no
+ `siteOrigin` at all (19 QR codes encoding `undefined/?v=…`) and `ferret-rescue`
+ has `http://localhost:3000` (codes that resolve to nothing on a phone). This is
+ the single best reason to run `umtool check` before every build.
+- **`channelSlug`** — the default archived channel for cue lookups, overridable per
+ clip.
+- **`availabilityCheckedOn`** — when sources were last confirmed fetchable. Stale in
+ both directions; `check-availability.mjs` refreshes it.
+
+## The Rumble two-ids rule
+
+A Rumble video has **two** ids: the site/MCP `video_id` is the *embed* id, while the
+local cue directory is named for the **URL slug**. The manifest's `video` field must
+be the local slug; the report cites the site id. A manifest that uses the site id
+finds no cue file — and finds that out twenty minutes into a build, unless
+`umtool check` ran first. `citeUrl` exists for the mirror case: a clip taken from a
+copy whose archived transcript is broken should point the QR at the copy that reads.
+
+## Deleted sources
+
+A source can be deleted from YouTube after the manifest is written, and clips are a
+**live network fetch** — there is no local media. The archive's own availability
+data is a *snapshot*, so it goes stale in both directions. Options, in order of
+preference: cut the same moment from a live mirror (via a `CHANNELS_DIR` shadow
+directory — but check the archive's alignment statement first, because mirrors are
+not assumed to share a clock), or convert the clip to a quote card.
+
+## Discovered by getting it wrong once
+
+- **A build that loses a clip is worse than a build that fails.** `--continue-on-error`
+ finishes everything buildable and then *refuses to concatenate*, because a
+ finished file quietly missing a citation looks complete.
+- **The cache is content-addressed by window, and that was being wasted.**
+ Before containing-window reuse, every window edit was a fresh download:
+ `ferret-rescue/out/clips-raw` holds 31 files for 10 clips, one source four times
+ over overlapping windows.
+- **Mirrors are not assumed to share a clock.** Timestamps mapped from one mirror to
+ another have to be verified against the mirror's own cue text, not assumed.
+- **`--chapters-only` is cheap and separate.** Retitling chapters does not need a
+ re-encode; the per-clip segments on disk are all the offsets need.