Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit cf15ea46a152a45c0a96d718399bf906db011b5b
parent 348813a3afa0d7df4a1a2b75c8417d16939ed5e9
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Tue, 18 Aug 2026 20:59:59 -0400

report-to-video: adopt the pipeline, and stop it re-downloading its own cache

1,522 lines of working pipeline that turns a cited sweep report into a finished
video, untracked in git while six real reports already depend on it. Committed
as-is, then given the three things umtool needs to drive it.

Containing-window cache reuse. A raw clip's window is in its name, which makes
the cache content-addressed -- but the lookup was for the EXACT name, so every
window edit re-downloaded material already on disk. ferret-rescue/out/clips-raw
holds 31 files for 10 clips; one source is there four times over overlapping
windows. Now any cached file containing the request satisfies it, and the
TIGHTEST container wins: silence detection decodes the whole file, so a 40s file
costs more than the 24s one that would also have done. Verified by nudging c01's
start 0.5s -- previously a fresh yt-dlp run, now it reuses 40.12-64.48 and
touches the network not at all.

--progress ndjson, so umtool's driver reads events instead of scraping prose.
The event set is exactly what was already being printed; making it a format
switch is what stops a wording change from breaking the driver.

--continue-on-error, because a dead source at clip 14 of 19 currently throws
away thirteen fetches already paid for. It finishes everything buildable and
then REFUSES to concatenate, exiting non-zero: a finished file that quietly
lost a citation looks complete, which is worse than no file at all.

Also: main() split out into an exported buildVideo(), --fetch-only <id> --pad <s>
so the clip bench's wide fetch runs the build's own fetch code (same naming, same
VP9 pin, same Rumble HLS retry), and check-availability.mjs -- yt-dlp --simulate
per distinct (channel, video), no bytes downloaded, classifying deleted/private/
restricted/no-cues. It belongs at the FRONT of a build: it is the one fact about
a manifest that goes stale in both directions.

resolve-windows.mjs still reports "0 window(s) changed" on all six real
manifests -- the fixed point holds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Diffstat:
Mpnpm-workspace.yaml | 1+
Ascripts/report-to-video/README.md | 324+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Ascripts/report-to-video/build-video.mjs | 839+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Ascripts/report-to-video/check-availability.mjs | 169+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Ascripts/report-to-video/package.json | 19+++++++++++++++++++
Ascripts/report-to-video/render-cards.mjs | 379+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Ascripts/report-to-video/resolve-windows.mjs | 220+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aumtool/docs/README.md | 68++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aumtool/docs/quirks.md | 129+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aumtool/docs/report-video.md | 160+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
10 files changed, 2308 insertions(+), 0 deletions(-)

diff --git a/pnpm-workspace.yaml b/pnpm-workspace.yaml @@ -4,6 +4,7 @@ packages: - export - homepage - mcp + - scripts/report-to-video - umtool allowBuilds: diff --git a/scripts/report-to-video/README.md b/scripts/report-to-video/README.md @@ -0,0 +1,324 @@ +# report-to-video + +Turns a cited sweep report into a video: the clips run in chronological order and +let the source speak for itself, with thin chrome carrying the citation and a +timeline of where you are. A video rendering of the reports we already write. + +Two scripts and a manifest: + +| file | lifetime | what it is | +|---|---|---| +| `build-video.mjs` | stable | manifest → mp4. Fetches clips, snaps cuts to silence, letterboxes them into the chrome, crossfades. | +| `render-cards.mjs` | stable | draws the timeline footer and marker, plus optional card stills. Imported by `build-video.mjs`. | +| `resolve-windows.mjs` | stable | widens clip windows from cue spans to whole sentences. Run once after authoring a manifest. | +| `<report>/video.manifest.json` | per report | the edit decision list. **This is the regeneration source of truth**, not the report. | + +``` +node scripts/report-to-video/resolve-windows.mjs ~/reports/<slug>/video.manifest.json --write +node scripts/report-to-video/build-video.mjs ~/reports/<slug>/video.manifest.json +``` + +Output lands in `<report dir>/out/`: `cards/`, `clips-raw/`, `segments/`, and the +finished `<slug>.mp4`. First worked example: `~/reports/ferret-rescue/`. + +## Why there is a manifest at all + +**A sweep report does not contain enough information to cut a video from.** Its +citations carry a single start second and nothing else — `momentUrl()` +(`common/lib/momentUrl.ts:106`) takes one `seconds` and floors it, and the MCP +`Snippet` type (`mcp/src/search.ts:41`) has no `end` field. There is no clip +length anywhere in a report. + +The end times do exist, they are just never rendered: every cue in +`transcripts/channels/<slug>/data/<id>/transcript.cues.json` is `{start, end, text}`. +So the manifest is built by matching each quote back to its covering cues and +recording the real window. That is also what handles ellipsis-joined citations — +a report quote like `"…" … "…"` is often two separate cue spans presented as one. + +The manifest additionally pins the provenance (share link, corpus handle, match +counts, the narrowing queries) so a rebuild months later is reproducible and the +video's own claims about its coverage can be checked. + +## Where a clip actually gets cut + +Three stages, because a cue span is the wrong answer twice over. + +1. **Cue span** — the raw window covering the quote, from `transcript.cues.json`. +2. **Sentence widening** (`resolve-windows.mjs`) — walk outward to the nearest cue + ending in `.`, `?` or `!`. A cue boundary is where the *caption line wrapped*, + so cutting there drops the run-up that makes a quote intelligible. + + Two asymmetries matter. The **start** takes a lead-in only if a real sentence + opening sits within `--max-lead` (8 s); otherwise it takes none, because a + half-sentence run-up is the irrelevant context you were trying to avoid, not + context. The **end** is never clamped to a budget — stopping partway through a + sentence is the exact mid-thought ending this removes — so `--max-tail` (12 s) + only bounds how far it looks before giving up and using the cue end. + + **Some uploads have no punctuation at all.** Older ASR in this corpus emits + unpunctuated cue text for whole videos, and sentence detection then has nothing + to find: widening degrades to the raw cue span at both ends. That is not a + silent failure you can ignore — it is what produced a clip opening mid-thought + on "higher than they can afford and because", and what let another clip's tail + run through its neighbour. For those videos, pick the window by reading the + cues and set it by hand; the de-overlap and silence passes still apply. + + **`lockStart` / `lockEnd` pin an edge** to exactly what the manifest says, and + neither widening nor de-overlap will move it. Reach for it when the utterance's + real trailing pause does not line up with its last cue's end — the snap picks + the *nearest* silence, and in speech over game audio the nearest one is often a + gap between syllables rather than the pause at the end of the thought. +3. **Silence snapping** (`build-video.mjs`) — a sentence boundary in the + *transcript* still isn't a boundary in the *audio*, so clips clip words in + half. Fetch `fetchPad` seconds wider than needed, run `silencedetect` over the + result, and move each cut to the nearest silence within `snapWindow`. Starts + land on a silence's END (just before speech resumes), ends on a silence's START + (just after speech stops). No silence close enough → keep the exact point; a + tight cut beats a cut in the wrong place. The build logs `start✓ end✓` per clip + so you can see which snapped. + + **The silence threshold is relative, and it has to be.** These are game + streams: the gaps between words are full of game audio and music — quiet, but + nowhere near silent. A fixed absolute threshold sits below the noise floor and + finds nothing. On one measured clip: mean volume −21 dB, **0** silences at + −32 dB, **25** at −26 dB. So each clip is measured with `volumedetect` first + and the threshold set `silenceRelDb` (default 6) below its own mean. If a rebuild + suddenly reports mostly `start– end–`, this is the knob. + +The trim happens during the burn-in encode, so snapping costs nothing extra. + +4. **De-overlap** (`resolve-windows.mjs`) — widening is per-clip and blind to its + neighbours, so two clips cut from the *same* video can end up overlapping, and + the overlap plays as the same footage twice. Any earlier clip whose tail runs + into a later clip's start is trimmed back to that start. This is not a rare + edge case: it fired on the first report, where a 2024 upload has **no + punctuation at all** in the relevant stretch, so sentence detection found + nothing and the tail ran the full budget straight through the next clip. + +## Regenerating and changing a video + +Everything is cached by content, so iteration is cheap: + +- **Reorder, drop or add clips** — edit `timeline`, re-run. Cached clips are not + refetched, so a re-cut costs an encode, not a download. +- **Change a clip's window** — edit `start`/`end`, re-run. A raw file's window is + in its name, so the cache is content-addressed; and since a request is satisfied + by any cached file that **contains** it, a nudge inside the existing pad costs + nothing at all. Only a window that escapes every cached file downloads again. +- **Preview one entry** — `--only <id>` builds a single segment and stops. +- **Work offline** — `--skip-fetch` fails loudly instead of downloading, so you + can confirm you are working entirely from cache. +- **Fetch one clip, wide** — `--fetch-only <id> --pad 20` puts a generous window + in the cache without building anything. This is what the umtool clip bench runs, + and containing-window reuse is what makes that fetch double as the build's cache. +- **Force a refetch** — delete `out/clips-raw/`, or pass `--no-reuse` to require an + exact-window file. + +Containing-window reuse was retrofitted, and the waste it removes is measurable: +`ferret-rescue/out/clips-raw` holds **31 files for 10 clips** because every window +edit before this downloaded the same material again — one source is there four +times over overlapping windows. The tightest containing file wins, not the widest, +because silence detection decodes the whole file and a 40 s file costs more than +the 24 s one that would also have done. + +After changing any window, re-run `resolve-windows.mjs --write` before building. +It is a **fixed point**: running it on an already-resolved manifest reports +`0 window(s) changed` and rewrites nothing. That property is load-bearing and was +not free — the manifest stores times rounded to 2 dp, so a value read back can sit +a hair below the cue end it came from, which lands the end lookup on the previous +cue and runs the search on to the *next* sentence. Left alone, every re-run grew +the same clip. Hence `EPS` in the end lookup and the 0.05 s deadband on applying a +change. + +## Driven from umtool + +The three CLIs are the source of truth and stay usable on their own; umtool drives +them rather than reimplementing them, so the UI and the terminal can never disagree +about a window, a format string or the Rumble retry. Three additions exist for that: + +- **`--progress ndjson`** — one JSON object per line instead of prose: + `start`, `card`, `clip`, `fetch`, `snap`, `segment`, `entry-failed`, `concat`, + `chapters`, `note`, `done`, `error`. The event set is exactly what was already + being printed; making it a format switch is what stops a wording change from + breaking the driver. +- **`--continue-on-error`** — record a failed entry and carry on. A dead source at + clip 14 of 19 otherwise throws away thirteen fetches already paid for. The run + still **refuses to concatenate** and exits non-zero: a finished file that quietly + lost a citation is worse than no file. +- **`buildVideo({manifestPath, opts, out, only, fetchOnly})`** is exported, and + `widen()` from `resolve-windows.mjs` already was. umtool imports `widen()` so the + bench's "extend to sentence end" is the CLI's own function, and **spawns** the + build — a 40-minute chain of yt-dlp and ffmpeg inside a request handler has no + cancellation story. + +`check-availability.mjs` is the fourth CLI and belongs at the *front* of a build: + +``` +node scripts/report-to-video/check-availability.mjs <manifest.json> +``` + +It runs `yt-dlp --simulate` once per distinct `(channel, video)` — no bytes +downloaded — classifies each failure (`deleted`, `private`, `restricted`, +`members-only`, `geo-blocked`, `no-cues`, `maybe_missing`), and writes +`out/availability.json` with a timestamp. This is the one fact about a manifest +that goes stale in *both* directions: a source can die after the manifest is +written, and a source annotated "gone" can come back. `no-cues` is called out +separately because it is a different bug — usually the Rumble two-ids trap, where +the manifest names the MCP video id while the cue file lives under the URL slug. + +## Manifest shape + +`timeline` is an ordered list; entries are `card` or `clip`. + +```jsonc +{ "type": "card", "id": "ch3", "style": "chapter", "seconds": 4.0, + "kicker": "March – November 2025", "heading": "Then: the county", + "sub": "Six months for the first approval" } + +{ "type": "clip", "id": "c04", "video": "uyz1_FIqIEk", + "start": 32989.56, "end": 32994.19, // cue-accurate, from transcript.cues.json + "cite": 32989, // the second shown in the attribution line + "quote": "The pre-application screening was approved by the county, dude." } +``` + +Two per-clip fields exist for compilations that span sources or need a hand-cut +window: + +- **`channel`** — the archived channel this clip's cue file lives under, overriding + `provenance.channelSlug`. A compilation about one person routinely spans several + mirror channels (`HasanAbiVODs` / `…VODs3` / `…VODsbackup`), and cue files are + keyed by channel, so a single manifest-wide slug cannot find them all. +- **`lock`** — exempt this clip from `resolve-windows`. Sentence-widening exists to + stop clips ending mid-thought, but that is exactly wrong when the author has + deliberately cut a quote short: a single cue often holds a whole paragraph, so + trimming to "I hate this country so much sometimes" and dropping the rest of the + sentence is an editorial decision that widening would silently undo. `lock` also + handles the reverse case — a clip whose lead-in would drag in seconds of some + *other* audio (a news package playing before the speaker starts). + +Clips also carry `section` and (auto-set) `sectionEnter`. Card styles — `title`, +`timeline`, `status`, `bullets`, `sources` — still work, but the ferret-rescue cut +uses none of them. `render` holds resolution, fps, fonts, palette and the knobs +(`fetchPad`, `snapWindow`, `silenceRelDb`, `transition`, `slideSeconds`, +`headerHeight`, `footerHeight`); `provenance` holds the sweep's scope and counts. + +## Chrome, not cards + +**The ferret-rescue cut has no cards at all** — no title, no chapter breaks, no +closing slate. It is a cited timeline and nothing else: the clips run in +chronological order and the source material carries the argument. Cards remain +supported for reports that want them, but the default posture is that anything +drawn is an interruption which has to earn its place. + +Nothing is drawn *over* the picture either. The video is **letterboxed between** +thin chrome rather than overlaid by it: + +- **Header (`headerHeight`, 56 px).** The citation only — cleaned title · upload + date · timestamp. No quote: the clip is already saying it, and burning in a + transcription of speech you can hear is noise. +- **Footer (`footerHeight`, 100 px).** The timeline: one node per milestone, each + with a label and a month/year stamp beneath it. Drawn once by + `renderFooterAssets()`. +- **The marker slides.** On the first clip of each section (`sectionEnter`, set + automatically), the fill bar and the amber marker animate from the previous node + to the current one over `slideSeconds`. Everywhere else they hold position. The + motion is ffmpeg expressions on `drawbox`/`overlay`, so it costs nothing beyond + the encode that was happening anyway. + +Clips carry `section` (index into `timelineNodes`); nodes supply `label` and +`date`. Ordering clips chronologically is the author's job — the manifest plays in +the order it is written. + +A consequence worth knowing: 16:9 source into the reduced height leaves narrow +pillarbox bars. That is the price of never covering the picture, and it is why the +chrome is kept as thin as it is. + +**Both bands are optional, and turning them off is a real setting, not a hack.** +`footerHeight: 0` (or an empty `timelineNodes`) drops the timeline; `headerHeight: 0` +drops the citation line. With both at zero the clips fill the whole frame and +nothing is drawn at all — the filtergraph loses the overlays rather than compositing +invisible ones, and `renderFooterAssets` returns early instead of drawing PNGs +nobody uses. Two reasons this comes up: + +- A timeline footer only means something if the clips *are* a progression through + time. A cut ordered by argument rather than by date should not draw one. +- The header prints the **upload date of the archived copy**, which for a VOD-mirror + channel is often years after the stream (a Nov 2019 stream re-uploaded in Apr 2023 + reads `… November 6, 2019 … · 2023-04-06`). The title usually carries the true + date, so nothing is false, but on a cut spanning many re-uploads it reads badly. + +The `hasan-hate-america` cut runs with both off. Restoring them is a two-value edit. + +Note: commas inside an ffmpeg filter expression have to survive filtergraph +parsing — wrapping the expression in single quotes is what protects them. + +- **Segments crossfade** (`transition`, default 0.5 s). This forces a full + re-encode of the timeline via `xfade`/`acrossfade` — the concat demuxer can only + stream-copy hard cuts. Pass `--no-xfade` for a fast hard-cut build while + iterating; the last pass can add the transitions back. + +## Things that cost time to find out + +**yt-dlp picks VP9 at `height<=720`, and that is a trap.** `--download-sections` +combined with `--force-keyframes-at-cuts` re-encodes, so a VP9 pick means +libvpx-vp9 — 27 seconds to cut a 5-second clip. It also writes a `.webm` and +appends that extension to whatever `-o` you gave, so the file never lands where +you asked and the run fails looking for it. Pin H.264/AAC in mp4 and the same cut +takes ~12 seconds. `build-video.mjs` does this and keeps a rename fallback for +the case where a fallback format still forces another container. + +**`--force-keyframes-at-cuts` is not optional here.** Without it the cut snaps to +the nearest preceding keyframe and can start seconds early. That is fine for a +human scrubbing a VOD; it is not fine when the clip *is* the citation. +`PlayerProvider.tsx:478` builds the copyable clip command without this flag — +correct for its purpose, wrong for ours. + +**`--ignore-config` is mandatory.** The operator's own yt-dlp config redirects +output to `~/Podcasts` and attaches thumbnail/metadata post-processors. Every +managed yt-dlp call in this repo passes `--ignore-config` for the same reason. + +**yt-dlp exit 101 is success**, not failure — it means a clean early stop. The +repo encodes this at `common/ytdlp/downloadOneManaged.ts:355`; a new caller has to +replicate it. + +**Local media will not help you.** Only 5 of 1,804 PirateSoftware video dirs hold +any media at all, and none are ones a report is likely to cite. Clips are a +network fetch. `metadata.info.json` and `transcript.cues.json` *are* present for +every video, so titles, dates, durations, webpage URLs and cue timings all come +from disk with no probe. + +**Upstream availability is load-bearing.** A clip can only be fetched while the +source is still up. Check `platform state` via `get_video_metadata` before +committing to a clip — a `deleted`/`maybe_missing` source needs a quote card +instead of footage. The archive outlives its sources, so a video built from an +old report will be *less* complete than the report unless this is handled +deliberately. + +**ImageMagick's `-size` leaks into the Pango group.** `-size 1920x1080 xc:BG` +followed by `( ... pango:@file )` renders the text into a full-frame box, which +pins it to the top and wraps at the frame edge instead of the text column. Reset +`-size` inside the parens. + +**ffmpeg `drawtext` does not wrap and hates punctuation.** Both are solved the +same way: wrap to a column count in JS, write to a file, and use +`textfile=`. Nothing then needs escaping. Stream titles also need emoji and +`!command` suffixes stripped or they render as tofu in the attribution line. + +**Segments are encoded to identical parameters on purpose** so the final +concatenation is a stream copy via the concat demuxer. Mismatched streams are the +usual reason a naive concat produces a broken or audio-desynced file. + +## Not done yet + +- **Narration is silent by design.** Cards carry the connective text; the only + audio is the clips'. A TTS layer would attach per card (`seconds` already gives + it a duration to fill) — deliberately deferred rather than designed out. +- **Snapping is silence-based, not word-based.** It finds gaps in the audio, which + is usually the same thing as a word boundary but is not guaranteed to be — + a speaker who does not pause gets the unsnapped cut. Forced alignment against + the transcript would be exact; `silencedetect` is a tenth of the work and + handles the cases that were actually audible. +- **Manifests are written by hand** from verified cue data. Deriving a first-draft + manifest automatically from a report's citations is the obvious next step; the + report parse is straightforward (`> "quote"` followed by + `— [title @ h:mm:ss](…?v=slug%2Fid&t=sec)`), the cue-matching is the real work. diff --git a/scripts/report-to-video/build-video.mjs b/scripts/report-to-video/build-video.mjs @@ -0,0 +1,839 @@ +#!/usr/bin/env node +// build-video.mjs — render a cited sweep report into a narrated-by-text video. +// +// Takes a video manifest (see README.md next to this file) and produces one mp4: +// text cards state the findings, clips let the source say it in their own voice, +// and every clip carries a burned-in quote plus its attribution. +// +// Pipeline, per manifest entry: +// card -> still PNG (render-cards.mjs) -> N seconds of video + silent audio +// clip -> yt-dlp --download-sections (WIDE) -> silence-snap -> trim + burn +// then the segments are crossfaded together into the finished file. +// +// Three things worth knowing about how clips are cut: +// +// 1. Windows come from the manifest as absolute [start, end] seconds, derived +// from transcript.cues.json (which carries an END per cue). A sweep report +// only ever records a single start second, so windows cannot be recovered +// from the report alone. +// 2. Those windows are widened to sentence boundaries by resolve-windows.mjs, +// so a clip carries the run-up that makes the quote make sense. +// 3. A cue boundary is still not a *speech* boundary — cutting there clips +// words in half. So we fetch wider than needed and snap the real cut to a +// silence found in the audio. That is what makes clips start and end +// between words rather than through them. +// +// Fetched clips are cached by (video, start, end); re-running is cheap and only +// changed entries re-download. Delete out/clips-raw to force a refetch. +// +// In the app: not used. On the CLI: +// node scripts/report-to-video/build-video.mjs <manifest.json> [options] +// +// Options: +// --out <dir> Output root (default: manifest dir + /out) +// --skip-fetch Fail instead of downloading anything not already cached +// --only <id> Build a single entry's segment and stop (for iterating) +// --no-xfade Hard cuts instead of crossfades (much faster; concat copy) +// --progress ndjson One JSON event per line instead of prose (for umtool) +// --continue-on-error Record a failed entry and carry on, instead of aborting +// --fetch-only <id> Fetch one clip's window into clips-raw and stop +// --pad <s> Override render.fetchPad (the clip bench fetches wide) +// +// Requires: yt-dlp, ffmpeg/ffprobe, ImageMagick with Pango. + +import { execFile } from "node:child_process"; +import { promisify } from "node:util"; +import { mkdir, writeFile, readFile, access, readdir, rename } from "node:fs/promises"; +import path from "node:path"; + +import { renderCard, renderFooterAssets } from "./render-cards.mjs"; + +const execFileP = promisify(execFile); + +const YTDLP = process.env.YTDLP_BIN ?? "yt-dlp"; +const FFMPEG = process.env.FFMPEG_BIN ?? "ffmpeg"; +const FFPROBE = process.env.FFPROBE_BIN ?? "ffprobe"; +const QRENCODE = process.env.QRENCODE_BIN ?? "qrencode"; + +const CHANNELS_DIR = + process.env.CHANNELS_DIR ?? + "/home/user/Projects/yt-dlp-transcript-browser/transcripts/channels"; + +const exists = (p) => access(p).then(() => true, () => false); + +// ---- progress protocol --------------------------------------------------- +// This has two audiences: a human watching a terminal, and umtool's build driver +// reading the pipe. Rather than have the driver scrape prose (which would make +// every wording change a breaking change), `--progress ndjson` switches every +// line to one JSON object. The event set is exactly what was already being +// printed -- this is a formatting switch, not new instrumentation. +// +// Events: start, card, clip, fetch, snap, segment, entry-failed, concat, +// chapters, note, done. +const HUMAN = { + start: (e) => `${e.title} — ${e.entries} entr(ies)`, + card: (e) => `card ${e.id}`, + clip: (e) => `clip ${e.id} (${e.video}) §${e.section}${e.sectionEnter ? " ⟶" : ""}`, + fetch: (e) => + e.reuse + ? ` fetch ${e.id}: ${e.reuse} already covers ${hms(e.from)}–${hms(e.to)} — no download` + : e.cached + ? null + : ` fetch ${e.id}: ${e.video} ${hms(e.from)}–${hms(e.to)}`, + snap: (e) => + ` snap ${e.id}: ${e.start ? "start✓" : "start–"} ${e.end ? "end✓" : "end–"} ` + + `(${Number(e.seconds).toFixed(1)}s)`, + segment: () => null, + "entry-failed": (e) => ` ** ${e.id} failed: ${e.message}`, + concat: (e) => `${e.mode === "xfade" ? "crossfading" : "hard-cutting"} ${e.n} segments…`, + chapters: (e) => `chapters: ${e.n} marker(s) -> ${e.file}`, + note: (e) => e.message, + done: (e) => + e.duration === undefined + ? `built ${e.out}` + : `\n${e.out}\nduration=${e.duration}\nsize=${e.size}`, +}; + +let EMIT = (ev, fields = {}) => { + const line = HUMAN[ev]?.({ ev, ...fields }); + if (line) console.log(line); +}; + +export function setProgressMode(mode) { + EMIT = + mode === "ndjson" + ? (ev, fields = {}) => process.stdout.write(JSON.stringify({ ev, ...fields }) + "\n") + : (ev, fields = {}) => { + const line = HUMAN[ev]?.({ ev, ...fields }); + if (line) console.log(line); + }; +} + +function hms(total) { + const s = Math.floor(total); + const h = Math.floor(s / 3600); + const m = Math.floor((s % 3600) / 60); + const sec = s % 60; + return h > 0 + ? `${h}:${String(m).padStart(2, "0")}:${String(sec).padStart(2, "0")}` + : `${m}:${String(sec).padStart(2, "0")}`; +} + +// Stream titles here are full of emoji and !commands. drawtext renders them as +// tofu with a text font, and they add nothing to an attribution line. +function cleanTitle(title) { + return title + .replace(/[\u{1F000}-\u{1FFFF}\u{2600}-\u{27BF}\u{FE0F}]/gu, "") + .replace(/\s*[!@]\S+/g, "") + .replace(/\s{2,}/g, " ") + .replace(/[\s·|-]+$/, "") + .trim(); +} + +// drawtext does not wrap. Break to a character budget, write to a file, and use +// textfile= so nothing needs shell or filter escaping. +function wrap(text, cols) { + const words = text.split(/\s+/); + const lines = []; + let line = ""; + for (const w of words) { + if (line && (line + " " + w).length > cols) { + lines.push(line); + line = w; + } else { + line = line ? line + " " + w : w; + } + } + if (line) lines.push(line); + return lines.join("\n"); +} + +async function videoMeta(videoId, channelSlug) { + const p = path.join(CHANNELS_DIR, channelSlug, "data", videoId, "transcript.cues.json"); + const d = JSON.parse(await readFile(p, "utf8")); + return { title: d.title, uploadDate: d.uploadDate, webpageUrl: d.webpageUrl, duration: d.duration }; +} + +async function probeDuration(file) { + const { stdout } = await execFileP(FFPROBE, [ + "-v", "error", "-show_entries", "format=duration", + "-of", "default=nw=1:nk=1", file, + ]); + return Number(stdout.trim()); +} + +// yt-dlp exits 101 on a clean early stop (break-on-existing / max-downloads). +// The repo treats that as success everywhere else; do the same here. +const ytdlpOk = (err) => err?.code === 101; + +// ---- the clip cache ------------------------------------------------------ +// A raw clip's window is IN ITS NAME, which makes the file immutable and the +// cache content-addressed. The original lookup was for the exact name, so any +// change to a window -- a hand edit, a widen, a nudge in the clip bench -- was a +// fresh download of material already on disk. Measured on ferret-rescue: 31 +// files for 10 clips, one source fetched four times over overlapping windows. +// +// So: satisfy a request from ANY cached file that contains it. The TIGHTEST +// container wins, because detectSilence decodes the whole file and a 40s file +// costs more than the 14s one that would also have done. The clip bench fetches +// deliberately wide, and this is what makes that generous fetch become the +// build's cache rather than a second one. +const WINDOW_RE = /^(\d+(?:\.\d+)?)-(\d+(?:\.\d+)?)$/; + +// A window read back from a 2 dp manifest can sit a hair outside the file that +// produced it; the same tolerance resolve-windows.mjs uses for the same reason. +const WIN_EPS = 0.02; + +export async function cachedWindowsFor(rawDir, video) { + let names; + try { + names = await readdir(rawDir); + } catch { + return []; + } + const prefix = `${video}_`; + const out = []; + for (const name of names) { + if (!name.startsWith(prefix) || !name.endsWith(".mp4")) continue; + // The remainder must be exactly `a-b`, which is what stops a video id that + // is a prefix of another (or one containing `_`) from claiming its files. + const m = WINDOW_RE.exec(name.slice(prefix.length, -4)); + if (!m) continue; + out.push({ name, path: path.join(rawDir, name), from: Number(m[1]), to: Number(m[2]) }); + } + return out; +} + +/** The tightest cached file containing [from, to], or null. */ +export async function findContainingWindow(rawDir, video, from, to) { + const windows = await cachedWindowsFor(rawDir, video); + let best = null; + for (const w of windows) { + if (w.from > from + WIN_EPS || w.to < to - WIN_EPS) continue; + if (!best || w.to - w.from < best.to - best.from) best = w; + } + return best; +} + +async function fetchClip(entry, meta, render, outDir, opts) { + // Deliberately over-fetch: the snapping pass below needs room on both sides to + // find a silence, and a clip that has no slack can only be cut where the cue + // happened to break — which is what put words in half in the first place. + const pad = opts.pad ?? render.fetchPad ?? 3.0; + const from = Math.max(0, entry.start - pad); + const to = entry.end + pad; + + const rawDir = path.join(outDir, "clips-raw"); + const name = `${entry.video}_${from.toFixed(2)}-${to.toFixed(2)}.mp4`; + const dest = path.join(rawDir, name); + if (await exists(dest)) { + EMIT("fetch", { id: entry.id, video: entry.video, from, to, cached: true }); + return { path: dest, fetchStart: from, cached: true }; + } + if (!opts.noReuse) { + const hit = await findContainingWindow(rawDir, entry.video, from, to); + if (hit) { + EMIT("fetch", { + id: entry.id, video: entry.video, from, to, cached: true, reuse: hit.name, + }); + // fetchStart is the CACHED file's start, not the requested one -- every cut + // downstream is expressed relative to it, so reuse is transparent. + return { path: hit.path, fetchStart: hit.from, cached: true }; + } + } + if (opts.skipFetch) throw new Error(`--skip-fetch set and no cached window covers ${name}`); + + const maxH = render.maxHeightSource; + const fmt = [ + `bv*[vcodec^=avc1][height<=${maxH}]+ba[acodec^=mp4a]`, + `bv*[ext=mp4][height<=${maxH}]+ba[ext=m4a]`, + `b[ext=mp4][height<=${maxH}]`, + `b[height<=${maxH}]`, + ].join("/"); + + const argsWith = (extra) => [ + // The operator's own yt-dlp config redirects output and attaches thumbnail + // and metadata post-processors; without this the clips land elsewhere. + "--ignore-config", + "--no-playlist", + "--download-sections", `*${from.toFixed(2)}-${to.toFixed(2)}`, + // Without this the cut snaps to the nearest preceding keyframe, which can be + // seconds early — fine for scrubbing, not fine when the clip IS the citation. + "--force-keyframes-at-cuts", + ...extra, + // Pin H.264/AAC in mp4. Left alone yt-dlp picks VP9+Opus at these heights, + // and since --force-keyframes-at-cuts re-encodes, that means libvpx-vp9 — + // 27s to cut a 5s clip. It also writes .webm and appends that to -o. + "-f", fmt, + "--merge-output-format", "mp4", + "-o", dest, + "--", meta.webpageUrl, + ]; + + const attempt = async (extra) => { + try { + await execFileP(YTDLP, argsWith(extra), { maxBuffer: 1 << 26 }); + return null; + } catch (err) { + return ytdlpOk(err) ? null : err; + } + }; + + EMIT("fetch", { id: entry.id, video: entry.video, from, to, cached: false }); + let err = await attempt([]); + + // Rumble delivers HLS whose segments are named `.tar`, and ffmpeg 8's picky + // extension check rejects those outright — "URL ... is not in + // allowed_segment_extensions" — killing the fetch with exit 183. Rumble ships + // no progressive format to fall back to, so without this every Rumble-sourced + // clip is unbuildable. + // + // It has to be a RETRY, not a default: -extension_picky lives on the HLS + // demuxer, so passing it against a progressive URL (YouTube's googlevideo mp4) + // makes ffmpeg abort with "Option extension_picky not found" — i.e. adding it + // unconditionally trades a Rumble failure for a YouTube one. + if (err && /allowed_segment_extensions|allowed_extensions/.test(String(err.stderr ?? err.message ?? ""))) { + EMIT("note", { id: entry.id, message: ` ${entry.id}: HLS segment extension rejected, retrying with -extension_picky 0` }); + err = await attempt(["--downloader-args", "ffmpeg_i:-extension_picky 0"]); + } + if (err) { + throw new Error(`yt-dlp failed for ${entry.id} (${entry.video}): ${err.stderr ?? err.message}`); + } + if (!(await exists(dest))) { + // If a fallback format still forced another container, yt-dlp writes + // "<dest>.<realext>". Adopt it rather than failing the run. + const dir = path.dirname(dest); + const base = path.basename(dest); + const stray = (await readdir(dir)).find((f) => f.startsWith(base + ".")); + if (!stray) throw new Error(`yt-dlp reported success but produced no file for ${entry.id}`); + await rename(path.join(dir, stray), dest); + } + return { path: dest, fetchStart: from, cached: false }; +} + +// Parse ffmpeg's silencedetect output into [{s, e}] intervals, in seconds +// relative to the start of the given file. +async function detectSilence(file, render) { + const minDur = render.silenceMinDur ?? 0.09; + + // The threshold has to be RELATIVE to the clip, not absolute. These are game + // streams: the gaps between words are full of game audio and music, so they + // are quiet but nowhere near silent. A fixed -32 dB sits below the noise floor + // of a typical clip here and finds literally zero silences (measured: mean + // volume -21 dB, 0 hits at -32 dB, 25 hits at -26 dB). Measure the clip first + // and cut a few dB under its own mean instead. + const { stderr: volLog } = await execFileP( + FFMPEG, + ["-nostdin", "-i", file, "-af", "volumedetect", "-f", "null", "-"], + { maxBuffer: 1 << 26 }, + ).catch((e) => ({ stderr: e.stderr ?? "" })); + const meanMatch = (volLog ?? "").match(/mean_volume:\s*(-?[\d.]+) dB/); + const mean = meanMatch ? Number(meanMatch[1]) : -24; + const noise = Math.max(-45, Math.min(-18, mean - (render.silenceRelDb ?? 6))); + + // ffmpeg exits 0 here, so stderr comes back on the resolved result. + const { stderr } = await execFileP( + FFMPEG, + ["-nostdin", "-i", file, "-af", `silencedetect=noise=${noise.toFixed(1)}dB:d=${minDur}`, "-f", "null", "-"], + { maxBuffer: 1 << 26 }, + ).catch((e) => ({ stderr: e.stderr ?? "" })); + const log = stderr ?? ""; + + const out = []; + let open = null; + for (const line of log.split("\n")) { + const s = line.match(/silence_start:\s*(-?[\d.]+)/); + if (s) open = Number(s[1]); + const e = line.match(/silence_end:\s*(-?[\d.]+)/); + if (e && open !== null) { + out.push({ s: open, e: Number(e[1]) }); + open = null; + } + } + return out; +} + +// Snap a desired cut to the nearest silence, so the clip begins and ends between +// words instead of through one. Returns the desired point unchanged when no +// silence is close enough — better a tight cut than a cut in the wrong place. +function snap(desired, intervals, kind, window) { + let best = null; + for (const iv of intervals) { + // Starting: we want to resume just before speech does -> the silence's END. + // Ending: we want to stop just after speech does -> the silence's START. + const point = kind === "start" ? iv.e : iv.s; + const d = Math.abs(point - desired); + if (d > window) continue; + if (!best || d < best.d) best = { d, point }; + } + if (!best) return { at: desired, snapped: false }; + const lead = kind === "start" ? -0.10 : 0.18; + return { at: Math.max(0, best.point + lead), snapped: true }; +} + +// Encoder quality is manifest-driven so a cut can trade size for fidelity without +// editing this file. Defaults reproduce the original hardcoded settings exactly. +const encodeArgs = (render) => [ + "-c:v", "libx264", + "-preset", render.preset ?? "medium", + "-crf", String(render.crf ?? 20), + "-pix_fmt", "yuv420p", + "-r", String(render.fps), + "-c:a", "aac", + "-b:a", render.audioBitrate ?? "160k", + "-ar", String(render.audioRate), + "-ac", String(render.audioChannels), + "-movflags", "+faststart", +]; + +async function buildClipSegment(entry, meta, render, outDir, opts, chrome, nodes, provenance) { + const { path: raw, fetchStart } = await fetchClip(entry, meta, render, outDir, opts); + const seg = path.join(outDir, "segments", `${entry.id}.mp4`); + const pal = render.palette; + const { width, height } = render; + + // Desired cut points, expressed relative to the over-fetched file. + const wantA = entry.start - fetchStart; + const wantB = entry.end - fetchStart; + const win = render.snapWindow ?? 1.6; + + const sil = await detectSilence(raw, render); + const a = snap(wantA, sil, "start", win); + const b = snap(wantB, sil, "end", win); + // Never let snapping invert or collapse the window. + const cutA = Math.min(a.at, wantB - 1); + const cutB = Math.max(b.at, cutA + 1); + EMIT("snap", { id: entry.id, start: a.snapped, end: b.snapped, seconds: cutB - cutA }); + + const quotePath = path.join(outDir, "segments", `${entry.id}.quote.txt`); + const attribPath = path.join(outDir, "segments", `${entry.id}.attrib.txt`); + // Written for reference/diffing only — the quote is no longer drawn on screen. + await writeFile(quotePath, wrap(`“${entry.quote}”`, 92), "utf8"); + + const d = meta.uploadDate; + const date = `${d.slice(0, 4)}-${d.slice(4, 6)}-${d.slice(6, 8)}`; + await writeFile( + attribPath, + `${cleanTitle(meta.title)} · ${date} @ ${hms(entry.cite ?? entry.start)}`, + "utf8", + ); + + // The picture is the point. Nothing is drawn over it: the video is letterboxed + // between a thin citation header and a thin timeline footer, so the source + // material plays unobstructed and the additions stay subtle. + const HH = render.headerHeight ?? 56; + const FH = chrome.footerHeight; + const hasFooter = FH > 0 && chrome.footer; + // headerHeight:0 drops the citation line too, leaving the clips alone on screen. + // Worth having: a cut whose sources are listed elsewhere does not need to carry + // its own attribution burnt into every frame. + const hasHeader = HH > 0; + const VH = height - HH - FH; + const trackAbsY = height - FH + chrome.trackY; + + // Where the progress marker travels this clip. Only the first clip of a + // section moves it; the rest hold it in place. + const T = render.slideSeconds ?? 0.9; + const xTo = chrome.xs[entry.section]; + const xFrom = entry.sectionEnter ? chrome.xs[Math.max(0, entry.section - 1)] : xTo; + // Commas inside a filter option have to survive filtergraph parsing; single + // quotes around the expression is what protects them. + const ramp = (a, b) => + a === b ? String(b) : `'if(lt(t,${T}),${a}+(${b}-${a})*t/${T},${b})'`; + const fillW = ramp(xFrom - chrome.x0, xTo - chrome.x0); + const markX = ramp(xFrom - chrome.markerRadius, xTo - chrome.markerRadius); + + const base = [ + `scale=${width}:${VH}:force_original_aspect_ratio=decrease`, + `pad=${width}:${VH}:(ow-iw)/2:(oh-ih)/2:color=${pal.bg}`, + `pad=${width}:${height}:0:${HH}:color=${pal.bg}`, + "setsar=1", + `fps=${render.fps}`, + ...(hasHeader + ? [ + `drawbox=x=90:y=${Math.round((HH - 24) / 2)}:w=4:h=24:color=${pal.accent}:t=fill`, + [ + `drawtext=textfile='${attribPath}'`, + `fontfile='${render.fontRegular}'`, + "fontsize=22", + `fontcolor=${pal.muted}`, + "x=118", + `y=${Math.round((HH - 26) / 2)}`, + ].join(":"), + ] + : []), + ].join(","); + + const qr = render.qr === false ? null : await qrForEntry(entry, provenance, render, outDir); + const qrIdx = hasFooter ? 3 : 1; + const qrM = render.qr?.margin ?? 28; + + const parts = hasFooter + ? [ + `[0:v]${base}[b]`, + `[b][1:v]overlay=0:${height - FH}[f]`, + `[f]drawbox=x=${chrome.x0}:y=${trackAbsY - 1}:w=${fillW}:h=3:color=${pal.accent}:t=fill[g]`, + `[g][2:v]overlay=x=${markX}:y=${trackAbsY - chrome.markerRadius}[q]`, + ] + : [`[0:v]${base}[q]`]; + + // Sit above the footer when there is one, so the code never straddles the chrome. + parts.push( + qr + ? `[q][${qrIdx}:v]overlay=x=W-w-${qrM}:y=H-h-${FH + qrM}[v]` + : `[q]null[v]`, + ); + + await execFileP( + FFMPEG, + [ + "-nostdin", "-v", "error", "-y", + "-ss", cutA.toFixed(3), "-to", cutB.toFixed(3), "-i", raw, + ...(hasFooter ? ["-i", chrome.footer, "-i", chrome.marker] : []), + ...(qr ? ["-i", qr.png] : []), + "-filter_complex", parts.join(";"), + "-map", "[v]", "-map", "0:a", + ...encodeArgs(render), + seg, + ], + { maxBuffer: 1 << 24 }, + ); + return seg; +} + +// ---- QR provenance code -------------------------------------------------- +// A compilation asks the viewer to take the edit on trust. The QR is the antidote: +// it resolves to this clip's exact START in the archive's own viewer, so anyone can +// pull up the surrounding hour and check that the cut is fair. Per clip, because a +// single code for the whole video would send everyone to the first citation. +// +// Two rules learned the hard way: it must be FULLY OPAQUE (a translucent QR will +// not scan) and it must keep its quiet zone (the white border is part of the +// symbol, not decoration). +async function qrForEntry(entry, provenance, render, outDir) { + const q = render.qr ?? {}; + // A mirror's LOCAL slug is not the id the site serves, and a clip taken from a + // copy whose archived transcript is broken should point at the copy that reads — + // so an explicit per-clip citeUrl always wins over the derived one. + const url = + entry.citeUrl ?? + `${provenance.siteOrigin}/?v=${encodeURIComponent( + `${entry.channel ?? provenance.channelSlug}/${entry.video}`, + )}&t=${Math.floor(entry.start)}`; + const png = path.join(outDir, "qr", `${entry.id}.png`); + await execFileP(QRENCODE, [ + "-o", png, + "-s", String(q.scale ?? 4), + "-m", String(q.quiet ?? 3), + "-l", q.ecc ?? "M", + url, + ]); + return { png, url }; +} + +async function buildCardSegment(card, render, outDir, nodes) { + const png = await renderCard(card, render, outDir, nodes); + const seg = path.join(outDir, "segments", `${card.id}.mp4`); + const dur = String(card.seconds); + + await execFileP( + FFMPEG, + [ + "-nostdin", "-v", "error", "-y", + "-loop", "1", "-t", dur, "-i", png, + "-f", "lavfi", "-t", dur, + "-i", `anullsrc=channel_layout=stereo:sample_rate=${render.audioRate}`, + "-vf", `fps=${render.fps},setsar=1`, + ...encodeArgs(render), + "-shortest", + seg, + ], + { maxBuffer: 1 << 24 }, + ); + return seg; +} + +// Crossfade every segment into the next. This is a full re-encode of the +// timeline — the concat demuxer can only stream-copy hard cuts — so --no-xfade +// stays available for quick iteration. +async function concatWithXfade(segments, render, outPath) { + const D = render.transition ?? 0.5; + const durs = []; + for (const s of segments) durs.push(await probeDuration(s)); + + const inputs = segments.flatMap((s) => ["-i", s]); + const parts = []; + let vlab = "[0:v]"; + let alab = "[0:a]"; + let acc = durs[0]; + + for (let i = 1; i < segments.length; i += 1) { + const off = acc - D; + parts.push(`${vlab}[${i}:v]xfade=transition=fade:duration=${D}:offset=${off.toFixed(3)}[v${i}]`); + parts.push(`${alab}[${i}:a]acrossfade=d=${D}:c1=tri:c2=tri[a${i}]`); + vlab = `[v${i}]`; + alab = `[a${i}]`; + acc = acc + durs[i] - D; + } + + await execFileP( + FFMPEG, + [ + "-nostdin", "-v", "error", "-y", + ...inputs, + "-filter_complex", parts.join(";"), + "-map", vlab, "-map", alab, + ...encodeArgs(render), + outPath, + ], + { maxBuffer: 1 << 26 }, + ); +} + +// ---- chapter markers ----------------------------------------------------- +// A compilation like this is a reference document as much as a video: the report +// cites moments, and a viewer wants to jump to them. Every clip therefore becomes +// a chapter. Offsets are derived exactly the way concatWithXfade derives its xfade +// offsets, so they stay correct for both crossfaded and hard-cut timelines. +// +// ffmetadata is a line-based format where =, ;, # and \ are structural, so a +// title carrying any of them has to be escaped or the file silently mis-parses. +const ffmetaEscape = (s) => String(s).replace(/([=;#\\])/g, "\\$1").replace(/\n/g, " "); + +async function segmentOffsets(segments, D) { + const durs = []; + for (const s of segments) durs.push(await probeDuration(s)); + const starts = []; + let acc = 0; + for (let i = 0; i < durs.length; i += 1) { + starts.push(acc); + acc += durs[i] - (i < durs.length - 1 ? D : 0); + } + return { starts, total: acc }; +} + +async function chapterTitle(entry, index, provenance) { + if (entry.chapter) return entry.chapter; + if (entry.type === "card") return entry.title ?? `Card ${index + 1}`; + try { + const meta = await videoMeta(entry.video, entry.channel ?? provenance.channelSlug); + const d = String(meta.uploadDate ?? ""); + const date = /^\d{8}$/.test(d) ? `${d.slice(0, 4)}-${d.slice(4, 6)}-${d.slice(6, 8)}` : d; + const title = String(meta.title ?? entry.video); + return `${date} — ${title.length > 60 ? `${title.slice(0, 57)}…` : title}`.trim(); + } catch { + return `${index + 1}. ${entry.video}`; + } +} + +async function muxChapters(finalPath, entries, segments, D, outDir, provenance) { + if (segments.length < 2) return; + const { starts, total } = await segmentOffsets(segments, D); + const lines = [";FFMETADATA1", ""]; + for (let i = 0; i < entries.length; i += 1) { + // Land just PAST the crossfade, so the marker opens on the incoming clip + // rather than on the outgoing one mid-dissolve. + const start = i === 0 ? 0 : starts[i] + D; + const end = i === entries.length - 1 ? total : starts[i + 1] + D; + lines.push( + "[CHAPTER]", + "TIMEBASE=1/1000", + `START=${Math.round(start * 1000)}`, + `END=${Math.round(end * 1000)}`, + `title=${ffmetaEscape(await chapterTitle(entries[i], i, provenance))}`, + "", + ); + } + const metaPath = path.join(outDir, "chapters.ffmeta"); + await writeFile(metaPath, lines.join("\n"), "utf8"); + + // Stream copy — adding chapters must never re-encode the finished timeline. + const tmp = finalPath.replace(/\.mp4$/, ".chapters.mp4"); + await execFileP( + FFMPEG, + ["-nostdin", "-v", "error", "-y", "-i", finalPath, "-i", metaPath, + "-map", "0", "-map_metadata", "0", "-map_chapters", "1", "-c", "copy", tmp], + { maxBuffer: 1 << 24 }, + ); + await rename(tmp, finalPath); + EMIT("chapters", { n: entries.length, file: path.basename(metaPath) }); +} + +async function concatHardCut(segments, outDir, outPath) { + const listPath = path.join(outDir, "concat.txt"); + await writeFile(listPath, segments.map((s) => `file '${s}'`).join("\n") + "\n", "utf8"); + await execFileP( + FFMPEG, + ["-nostdin", "-v", "error", "-y", "-f", "concat", "-safe", "0", "-i", listPath, "-c", "copy", outPath], + { maxBuffer: 1 << 24 }, + ); +} + +/** + * Build a manifest into a video. + * + * Exported so umtool's driver runs the SAME code the CLI does. It is still + * SPAWNED rather than imported by the app: a 40-minute chain of yt-dlp and + * ffmpeg inside a request handler has no cancellation story, and a runaway + * grandchild would outlive the request that started it. + */ +export async function buildVideo({ manifestPath, opts = {}, out, only, fetchOnly } = {}) { + const manifest = JSON.parse(await readFile(manifestPath, "utf8")); + const { render, provenance } = manifest; + const outDir = out ?? path.join(path.dirname(path.resolve(manifestPath)), "out"); + + for (const d of ["cards", "clips-raw", "segments", "qr"]) { + await mkdir(path.join(outDir, d), { recursive: true }); + } + + // Fetch one clip's window and stop. This is what the clip bench's "fetch 20s + // more" runs, so a bench fetch and a build fetch can never disagree about + // naming, format selection, the VP9 trap or the Rumble HLS retry. + if (fetchOnly) { + const entry = manifest.timeline.find((e) => e.id === fetchOnly); + if (!entry) throw new Error(`no timeline entry with id ${fetchOnly}`); + if (entry.type === "card") throw new Error(`${fetchOnly} is a card, not a clip`); + const meta = await videoMeta(entry.video, entry.channel ?? provenance.channelSlug); + const r = await fetchClip(entry, meta, render, outDir, opts); + EMIT("done", { out: r.path, fetchStart: r.fetchStart, cached: r.cached }); + return { out: r.path, failures: [] }; + } + + // Footer chrome is shared by every clip, so build it once up front. + const chrome = await renderFooterAssets(render, manifest.timelineNodes, outDir); + + const entries = manifest.timeline.filter((e) => !only || e.id === only); + if (only && !entries.length) throw new Error(`no timeline entry with id ${only}`); + const segments = []; + const failures = []; + + const D = opts.noXfade || (render.transition ?? 0.5) === 0 ? 0 : render.transition ?? 0.5; + + // Retro-fit chapters onto an already-built file without re-encoding it. The + // per-clip segments are still on disk, which is all the offsets need. + if (opts.chaptersOnly) { + const finalPath = path.join(outDir, `${manifest.slug}.mp4`); + const segs = entries.map((e) => path.join(outDir, "segments", `${e.id}.mp4`)); + for (const seg of segs) { + if (!(await exists(seg))) + throw new Error(`--chapters-only needs ${seg}, which is missing — run a full build first`); + } + await muxChapters(finalPath, entries, segs, D, outDir, provenance); + return { out: finalPath, failures: [] }; + } + + EMIT("start", { title: manifest.title, entries: entries.length, out: outDir }); + for (let i = 0; i < entries.length; i += 1) { + const entry = entries[i]; + try { + if (entry.type === "card") { + EMIT("card", { id: entry.id, i, n: entries.length }); + segments.push(await buildCardSegment(entry, render, outDir, manifest.timelineNodes)); + } else { + const meta = await videoMeta(entry.video, entry.channel ?? provenance.channelSlug); + EMIT("clip", { + id: entry.id, i, n: entries.length, video: entry.video, + section: entry.section, sectionEnter: !!entry.sectionEnter, + }); + segments.push( + await buildClipSegment( + entry, meta, render, outDir, opts, chrome, manifest.timelineNodes, provenance, + ), + ); + } + EMIT("segment", { id: entry.id, path: segments[segments.length - 1] }); + } catch (err) { + // Without --continue-on-error a dead source at entry 14 of 19 throws away + // the thirteen fetches already paid for. With it, everything buildable is + // built and the run reports what was not. + if (!opts.continueOnError) throw err; + const message = err?.message ?? String(err); + failures.push({ id: entry.id, message }); + EMIT("entry-failed", { id: entry.id, message }); + } + } + + if (only) { + EMIT("done", { out: segments[0], failures }); + return { out: segments[0], failures }; + } + + // A timeline that silently lost a clip is a worse outcome than no file at all: + // the finished video would look complete and be missing a citation. So the + // segments are kept (they cost the fetches) and the concat is refused. + if (failures.length) { + EMIT("note", { + message: `refusing to concat: ${failures.length} of ${entries.length} entries failed ` + + `(${failures.map((f) => f.id).join(", ")})`, + }); + return { out: null, failures }; + } + + const final = path.join(outDir, `${manifest.slug}.mp4`); + // `transition: 0` is a real editorial choice, not just a speed knob: hard cuts + // hit harder on a compilation whose point is repetition. Honouring it here keeps + // the manifest the source of truth, so a rebuild does not silently re-add fades. + EMIT("concat", { mode: D === 0 ? "hardcut" : "xfade", n: segments.length }); + if (D === 0) await concatHardCut(segments, outDir, final); + else await concatWithXfade(segments, render, final); + + if (!opts.noChapters) await muxChapters(final, entries, segments, D, outDir, provenance); + + const { stdout } = await execFileP(FFPROBE, [ + "-v", "error", "-show_entries", "format=duration,size", + "-of", "default=noprint_wrappers=1", final, + ]); + const probe = Object.fromEntries( + stdout.trim().split("\n").map((l) => l.split("=")), + ); + EMIT("done", { out: final, duration: Number(probe.duration), size: Number(probe.size) }); + return { out: final, failures }; +} + +async function main() { + const argv = process.argv.slice(2); + const manifestPath = argv.find((a) => !a.startsWith("--")); + if (!manifestPath) { + console.error( + "usage: build-video.mjs <manifest.json> [--out <dir>] [--only <id>] [--fetch-only <id>]\n" + + " [--pad <s>] [--skip-fetch] [--no-xfade] [--no-chapters] [--chapters-only]\n" + + " [--progress ndjson] [--continue-on-error] [--no-reuse]", + ); + process.exit(2); + } + const flag = (n) => { + const i = argv.indexOf(n); + return i >= 0 ? argv[i + 1] : undefined; + }; + setProgressMode(flag("--progress") ?? "human"); + + const padArg = flag("--pad"); + const opts = { + skipFetch: argv.includes("--skip-fetch"), + continueOnError: argv.includes("--continue-on-error"), + noXfade: argv.includes("--no-xfade"), + noChapters: argv.includes("--no-chapters"), + chaptersOnly: argv.includes("--chapters-only"), + noReuse: argv.includes("--no-reuse"), + pad: padArg === undefined ? undefined : Number(padArg), + }; + + const { failures } = await buildVideo({ + manifestPath, + opts, + out: flag("--out"), + only: flag("--only"), + fetchOnly: flag("--fetch-only"), + }); + // Non-zero on a partial run, so a caller that ignores the events still learns + // the build did not produce what was asked for. + if (failures.length) process.exit(1); +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main().catch((err) => { + EMIT("error", { message: err?.message ?? String(err) }); + console.error(err.message ?? err); + process.exit(1); + }); +} diff --git a/scripts/report-to-video/check-availability.mjs b/scripts/report-to-video/check-availability.mjs @@ -0,0 +1,169 @@ +#!/usr/bin/env node +// check-availability.mjs — is every source this manifest cites still fetchable? +// +// This is the one fact about a report video that goes stale in BOTH directions +// and that nothing on disk records. A source can be deleted between writing the +// manifest and building it (so a 40-minute build dies at clip 14 having paid for +// thirteen fetches), and a source can come back (so a manifest annotated "gone" +// stays wrong). Neither is visible until a build runs. +// +// So it runs first, it runs cheap, and it writes down when it ran. `--simulate` +// resolves formats without downloading a byte: a few seconds for a whole +// manifest against twenty-odd minutes for the build it protects. +// +// On the CLI: +// node scripts/report-to-video/check-availability.mjs <manifest.json> [--json] +// +// Options: +// --json Print the report as JSON instead of a table +// --out <dir> Output root (default: manifest dir + /out) +// --allow-missing Exit 0 even when a source is gone (report only) +// --max-age <days> Reuse a recorded verdict younger than this (default: 0) + +import { execFile } from "node:child_process"; +import { promisify } from "node:util"; +import { mkdir, readFile, writeFile } from "node:fs/promises"; +import path from "node:path"; + +const execFileP = promisify(execFile); + +const YTDLP = process.env.YTDLP_BIN ?? "yt-dlp"; +const CHANNELS_DIR = + process.env.CHANNELS_DIR ?? + "/home/user/Projects/yt-dlp-transcript-browser/transcripts/channels"; + +// yt-dlp says why in prose, and the distinction matters editorially: a private +// or removed video needs the clip converting to a quote card, while a network +// blip needs a retry. Anything unrecognised stays `maybe_missing` rather than +// being called deleted -- claiming a source is gone when it is not is the more +// expensive mistake, because it invites deleting a citation. +function classify(stderr) { + const t = String(stderr ?? ""); + if (/Private video|private/i.test(t)) return "private"; + if (/removed by the uploader|has been removed|no longer available|Video unavailable|does not exist|410/i.test(t)) + return "deleted"; + if (/age.?restrict|Sign in to confirm|confirm your age/i.test(t)) return "restricted"; + if (/members-only|join this channel/i.test(t)) return "members-only"; + if (/geo|not available in your country/i.test(t)) return "geo-blocked"; + return "maybe_missing"; +} + +async function cueMeta(videoId, channelSlug) { + const p = path.join(CHANNELS_DIR, channelSlug, "data", videoId, "transcript.cues.json"); + const d = JSON.parse(await readFile(p, "utf8")); + return { title: d.title, webpageUrl: d.webpageUrl, duration: d.duration }; +} + +export async function checkAvailability(manifestPath, { outDir, maxAgeDays = 0 } = {}) { + const manifest = JSON.parse(await readFile(manifestPath, "utf8")); + const slug = manifest.provenance?.channelSlug; + const dir = outDir ?? path.join(path.dirname(path.resolve(manifestPath)), "out"); + const file = path.join(dir, "availability.json"); + + // A clip may name its own channel: the same streamer's VODs are mirrored + // across more than one archive, and the same id under a different slug is a + // different file. So the unit of work is (channel, video), never video alone. + const wanted = new Map(); + for (const e of manifest.timeline ?? []) { + if (e.type !== "clip") continue; + const channel = e.channel ?? slug; + const key = `${channel}/${e.video}`; + if (!wanted.has(key)) wanted.set(key, { key, channel, video: e.video, clips: [] }); + wanted.get(key).clips.push(e.id); + } + + const prev = await readFile(file, "utf8").then( + (s) => JSON.parse(s), + () => ({ sources: [] }), + ); + const prevBy = new Map((prev.sources ?? []).map((s) => [s.key, s])); + const freshMs = maxAgeDays * 86400_000; + + const sources = []; + for (const w of wanted.values()) { + const was = prevBy.get(w.key); + if (freshMs > 0 && was?.checkedAt && Date.now() - Date.parse(was.checkedAt) < freshMs) { + sources.push({ ...was, clips: w.clips, reused: true }); + continue; + } + + let meta; + try { + meta = await cueMeta(w.video, w.channel); + } catch { + // No cue file is a DIFFERENT failure from a dead source, and it is the one + // the Rumble two-ids trap produces: the manifest names the MCP video id + // while the cues live under the URL slug. Build would die here too, so it + // is reported here rather than discovered twenty minutes in. + sources.push({ + ...w, ok: false, state: "no-cues", checkedAt: new Date().toISOString(), + error: `no transcript.cues.json under ${w.channel}/data/${w.video}`, + }); + continue; + } + + try { + await execFileP( + YTDLP, + ["--ignore-config", "--no-playlist", "--simulate", "--quiet", "--no-warnings", "--", meta.webpageUrl], + { maxBuffer: 1 << 24 }, + ); + sources.push({ + ...w, ok: true, state: "ok", title: meta.title, url: meta.webpageUrl, + checkedAt: new Date().toISOString(), error: null, + }); + } catch (err) { + const stderr = err?.stderr ?? err?.message ?? ""; + sources.push({ + ...w, ok: false, state: classify(stderr), title: meta.title, url: meta.webpageUrl, + checkedAt: new Date().toISOString(), error: String(stderr).trim().split("\n").slice(-3).join(" "), + }); + } + } + + const report = { manifest: path.resolve(manifestPath), checkedAt: new Date().toISOString(), sources }; + await mkdir(dir, { recursive: true }); + await writeFile(file, JSON.stringify(report, null, 2) + "\n", "utf8"); + return { ...report, file }; +} + +async function main() { + const argv = process.argv.slice(2); + const manifestPath = argv.find((a) => !a.startsWith("--")); + if (!manifestPath) { + console.error("usage: check-availability.mjs <manifest.json> [--json] [--out <dir>] [--allow-missing]"); + process.exit(2); + } + const flag = (n) => { + const i = argv.indexOf(n); + return i >= 0 ? argv[i + 1] : undefined; + }; + const report = await checkAvailability(manifestPath, { + outDir: flag("--out"), + maxAgeDays: Number(flag("--max-age") ?? 0), + }); + + if (argv.includes("--json")) { + console.log(JSON.stringify(report, null, 2)); + } else { + for (const s of report.sources) { + const mark = s.ok ? "ok " : "GONE"; + console.log( + `${mark} ${s.key.padEnd(40)} ${String(s.state).padEnd(14)} ` + + `${s.clips.length} clip(s)${s.reused ? " (cached)" : ""}`, + ); + if (!s.ok && s.error) console.log(` ${s.error}`); + } + const bad = report.sources.filter((s) => !s.ok).length; + console.log(`\n${report.sources.length} source(s), ${bad} unavailable -> ${report.file}`); + } + + if (report.sources.some((s) => !s.ok) && !argv.includes("--allow-missing")) process.exit(1); +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main().catch((err) => { + console.error(err.message ?? err); + process.exit(1); + }); +} diff --git a/scripts/report-to-video/package.json b/scripts/report-to-video/package.json @@ -0,0 +1,19 @@ +{ + "name": "report-to-video", + "version": "0.1.0", + "private": true, + "type": "module", + "description": "Turn a cited sweep report into a narrated-by-text video.", + "bin": { + "report-build-video": "./build-video.mjs", + "report-resolve-windows": "./resolve-windows.mjs", + "report-check-availability": "./check-availability.mjs" + }, + "exports": { + "./resolve-windows": "./resolve-windows.mjs", + "./build-video": "./build-video.mjs", + "./render-cards": "./render-cards.mjs", + "./check-availability": "./check-availability.mjs", + "./package.json": "./package.json" + } +} diff --git a/scripts/report-to-video/render-cards.mjs b/scripts/report-to-video/render-cards.mjs @@ -0,0 +1,379 @@ +#!/usr/bin/env node +// render-cards.mjs — turn a video manifest's `card` entries into PNG stills. +// +// One PNG per card, written to <outDir>/cards/<id>.png at the manifest's render +// resolution. Text is laid out by ImageMagick's Pango delegate rather than +// ffmpeg's drawtext: Pango wraps, kerns and takes inline markup, so a card is a +// single markup string instead of a stack of hand-positioned drawtext filters. +// +// Card styles (manifest `style` field): +// title — the opening card: big heading, subtitle, provenance footer +// chapter — an act break: small amber kicker over a large heading +// status — the bottom-line card: kicker, amber heading, subtitle +// bullets — heading plus a list of caveats +// sources — closing attribution +// +// In the app: not used. On the CLI: +// node scripts/report-to-video/render-cards.mjs <manifest.json> [--out <dir>] +// +// Options: +// --out <dir> Output root (default: the manifest's directory + /out) +// --only <id> Render just one card, by manifest id +// +// Requires: ImageMagick built with Pango (magick -list format | grep PANGO). + +import { execFile } from "node:child_process"; +import { promisify } from "node:util"; +import { mkdir, writeFile, readFile } from "node:fs/promises"; +import path from "node:path"; + +const execFileP = promisify(execFile); + +// Pango markup is XML-ish, so anything we interpolate has to be escaped first. +// Curly quotes and the ellipsis pass through fine; only these five matter. +function esc(s) { + return String(s) + .replace(/&/g, "&amp;") + .replace(/</g, "&lt;") + .replace(/>/g, "&gt;") + .replace(/"/g, "&quot;") + .replace(/'/g, "&apos;"); +} + +function span(text, { size, color, weight, family = "Fira Sans" }) { + const attrs = [`font_family="${family}"`, `size="${Math.round(size * 1024)}"`]; + if (color) attrs.push(`foreground="${color}"`); + if (weight) attrs.push(`weight="${weight}"`); + return `<span ${attrs.join(" ")}>${text}</span>`; +} + +// Each style returns Pango markup for the whole card body. Blank lines are real +// newlines in the markup — Pango honours them, which is how vertical rhythm is +// set without positioning each run separately. +function markupFor(card, pal) { + const H = (t, size = 62) => + span(esc(t), { size, color: pal.fg, weight: "bold" }); + const KICKER = (t) => + span(esc(t.toUpperCase()), { size: 24, color: pal.amber, weight: "bold" }); + const SUB = (t, size = 30) => span(esc(t), { size, color: pal.muted }); + + switch (card.style) { + case "title": + return [ + span(esc(card.heading), { size: 82, color: pal.fg, weight: "bold" }), + "", + span(esc(card.sub), { size: 38, color: pal.accent }), + "", + "", + SUB(card.foot, 24), + ].join("\n"); + + case "chapter": + return [ + card.kicker ? KICKER(card.kicker) : null, + card.kicker ? "" : null, + H(card.heading), + card.sub ? "" : null, + card.sub ? SUB(card.sub, 32) : null, + ] + .filter((l) => l !== null) + .join("\n"); + + case "status": + return [ + card.kicker ? KICKER(card.kicker) : null, + card.kicker ? "" : null, + span(esc(card.heading), { size: 58, color: pal.amber, weight: "bold" }), + card.sub ? "" : null, + card.sub ? SUB(card.sub, 32) : null, + ] + .filter((l) => l !== null) + .join("\n"); + + case "bullets": { + const items = (card.bullets ?? []).flatMap((b) => [ + `${span("— ", { size: 30, color: pal.accent, weight: "bold" })}${span( + esc(b), + { size: 30, color: pal.fg }, + )}`, + "", + ]); + return [H(card.heading, 52), "", ...items].join("\n"); + } + + case "sources": + return [ + H(card.heading, 52), + "", + card.sub ? SUB(card.sub, 32) : null, + card.foot ? "" : null, + card.foot + ? card.foot + .split("\n") + .map((l) => SUB(l, 24)) + .join("\n") + : null, + ] + .filter((l) => l !== null) + .join("\n"); + + default: + return H(card.heading ?? card.id); + } +} + +// A chapter break that just states a heading is dead air — it stops the video to +// say something the next clip is about to say anyway. This draws the whole +// project arc instead, with the current step lit and everything before it +// filled, so the pause carries information: where we are and how far is left. +// +// Nodes come from the manifest's top-level `timelineNodes`; the card names its +// position with `step` (0-based). +async function renderTimelineCard(card, render, nodes, outDir) { + const pal = render.palette; + const { width, height } = render; + const outPath = path.join(outDir, "cards", `${card.id}.png`); + const dir = path.join(outDir, "cards"); + + const x0 = 260; + const x1 = width - 260; + const axisY = Math.round(height * 0.56); + const gap = (x1 - x0) / (nodes.length - 1); + const xs = nodes.map((_, i) => Math.round(x0 + i * gap)); + const cur = card.step; + + const args = ["-size", `${width}x${height}`, `xc:${pal.bg}`, "-strokewidth", "4"]; + + // Track: filled up to the current node, dim beyond it. + args.push( + "-stroke", pal.muted, "-fill", "none", + "-draw", `line ${xs[0]},${axisY} ${xs[xs.length - 1]},${axisY}`, + ); + if (cur > 0) { + args.push("-stroke", pal.accent, "-draw", `line ${xs[0]},${axisY} ${xs[cur]},${axisY}`); + } + + // Nodes: past and present filled, future hollow. The current one is larger and + // amber so the eye lands on it without needing a label to say "you are here". + nodes.forEach((_, i) => { + const r = i === cur ? 19 : 11; + const color = i === cur ? pal.amber : i < cur ? pal.accent : pal.bg; + args.push( + "-stroke", i <= cur ? (i === cur ? pal.amber : pal.accent) : pal.muted, + "-fill", color, + "-draw", `circle ${xs[i]},${axisY} ${xs[i] + r},${axisY}`, + ); + }); + + args.push("-stroke", "none"); + + // Heading, centred over the whole card. + const headMarkup = [ + span(esc((card.kicker ?? nodes[cur].label).toUpperCase()), { + size: 26, color: pal.amber, weight: "bold", + }), + "", + span(esc(card.heading ?? nodes[cur].title), { size: 62, color: pal.fg, weight: "bold" }), + card.sub ? "" : null, + card.sub ? span(esc(card.sub), { size: 30, color: pal.muted }) : null, + ] + .filter((l) => l !== null) + .join("\n"); + + // Heading, left-aligned on the same margin the other card styles use. + const headPath = path.join(dir, `${card.id}.head.pango`); + await writeFile(headPath, headMarkup, "utf8"); + args.push( + "(", "-size", `${width - 460}x`, "-background", "none", + "-define", `pango:width=${width - 460}`, + `pango:@${headPath}`, ")", + "-gravity", "NorthWest", + "-geometry", `+${Math.round(width * 0.09) + 58}+${Math.round(height * 0.19)}`, + "-composite", + ); + + // Per-node date labels, centred under their dot. ImageMagick's + // `pango:alignment` define does not actually centre the text inside the box, + // so measure each rendered label and place it by hand instead of trusting it. + for (let i = 0; i < nodes.length; i += 1) { + const isCur = i === cur; + const labMarkup = span(esc(nodes[i].label), { + size: isCur ? 24 : 21, + color: isCur ? pal.fg : i < cur ? pal.muted : "#5c5570", + weight: isCur ? "bold" : "normal", + }); + const labPath = path.join(dir, `${card.id}.n${i}.pango`); + const labPng = path.join(dir, `${card.id}.n${i}.png`); + await writeFile(labPath, labMarkup, "utf8"); + await execFileP("magick", ["-background", "none", `pango:@${labPath}`, labPng]); + const { stdout } = await execFileP("magick", ["identify", "-format", "%w", labPng]); + const w = Number(stdout.trim()); + args.push(labPng, "-geometry", `+${xs[i] - Math.round(w / 2)}+${axisY + 44}`, "-composite"); + } + + args.push(outPath); + await execFileP("magick", args, { maxBuffer: 1 << 24 }); + return outPath; +} + +// Footer chrome, drawn once and overlaid on every clip: a track, a dot per +// section and its label. The *progress* along it is not baked in here — the fill +// bar and the amber marker are drawn by ffmpeg at encode time so they can slide +// between sections instead of cutting. See buildClipSegment in build-video.mjs. +// +// Returns the geometry the encoder needs to place those moving parts. +export async function renderFooterAssets(render, nodes, outDir) { + const pal = render.palette; + const { width } = render; + const FH = render.footerHeight ?? 92; + const dir = path.join(outDir, "cards"); + + // A cut whose clips are not a progression through time has nothing for a + // timeline to say, and a footer drawn anyway is chrome that has not earned its + // place. No nodes (or an explicit zero height) means no footer at all — the + // caller letterboxes against the header alone. + if (!nodes?.length || FH === 0) { + return { footer: null, marker: null, footerHeight: 0, trackY: 0, xs: [], x0: 0, markerRadius: 0 }; + } + + const x0 = 200; + const x1 = width - 200; + const trackY = 26; + const gap = (x1 - x0) / (nodes.length - 1); + const xs = nodes.map((_, i) => Math.round(x0 + i * gap)); + + const footer = path.join(dir, "_footer.png"); + const args = [ + "-size", `${width}x${FH}`, `xc:${pal.bg}`, + "-strokewidth", "3", + "-stroke", "#3a3450", "-fill", "none", + "-draw", `line ${xs[0]},${trackY} ${xs[xs.length - 1]},${trackY}`, + "-stroke", "none", + ]; + for (const x of xs) { + args.push("-fill", "#4a4363", "-draw", `circle ${x},${trackY} ${x + 6},${trackY}`); + } + + // Two lines per node: what happened, then when. Both are measured and placed + // by hand — ImageMagick's `pango:alignment` define does not actually centre + // text inside its box. + for (let i = 0; i < nodes.length; i += 1) { + const lines = [ + { text: nodes[i].label, size: 18, color: pal.fg, dy: 18 }, + { text: nodes[i].date, size: 16, color: pal.muted, dy: 42 }, + ]; + for (const [k, ln] of lines.entries()) { + const pPath = path.join(dir, `_footer.n${i}.l${k}.pango`); + const pPng = path.join(dir, `_footer.n${i}.l${k}.png`); + await writeFile(pPath, span(esc(ln.text), { size: ln.size, color: ln.color }), "utf8"); + await execFileP("magick", ["-background", "none", `pango:@${pPath}`, pPng]); + const { stdout } = await execFileP("magick", ["identify", "-format", "%w", pPng]); + args.push( + pPng, + "-geometry", `+${xs[i] - Math.round(Number(stdout.trim()) / 2)}+${trackY + ln.dy}`, + "-composite", + ); + } + } + args.push(footer); + await execFileP("magick", args, { maxBuffer: 1 << 24 }); + + // The travelling marker. + const marker = path.join(dir, "_marker.png"); + const r = 11; + await execFileP("magick", [ + "-size", `${r * 2 + 2}x${r * 2 + 2}`, "xc:none", + "-fill", pal.amber, "-stroke", "none", + "-draw", `circle ${r + 1},${r + 1} ${r * 2 + 1},${r + 1}`, + marker, + ]); + + return { footer, marker, footerHeight: FH, trackY, xs, x0, markerRadius: r }; +} + +export async function renderCard(card, render, outDir, nodes) { + if (card.style === "timeline") { + if (!nodes?.length) throw new Error(`card ${card.id} is style:timeline but no timelineNodes given`); + return renderTimelineCard(card, render, nodes, outDir); + } + return renderPlainCard(card, render, outDir); +} + +async function renderPlainCard(card, render, outDir) { + const pal = render.palette; + const { width, height } = render; + const textWidth = Math.round(width * 0.74); + const outPath = path.join(outDir, "cards", `${card.id}.png`); + + // Pango reads its markup from a file to keep it clear of shell/argv quoting. + const markupPath = path.join(outDir, "cards", `${card.id}.pango`); + await writeFile(markupPath, markupFor(card, pal), "utf8"); + + // One magick invocation: solid ground, an accent rule down the left margin, + // then the Pango block composited over it. The rule is what keeps the cards + // recognisably one family across styles. + const barX = Math.round(width * 0.09); + const barTop = Math.round(height * 0.28); + const barBottom = Math.round(height * 0.72); + + const args = [ + "-size", `${width}x${height}`, + `xc:${pal.bg}`, + "-fill", pal.accent, + "-draw", `rectangle ${barX},${barTop} ${barX + 6},${barBottom}`, + "(", + // `-size` is still set to the full frame from the canvas above, and the + // pango delegate honours it — leaving it alone renders the text into a + // 1920x1080 box, which pins the block to the top and wraps at the frame + // edge instead of the margin. Reset it to the text column, height auto. + "-size", `${textWidth}x`, + "-background", "none", + "-define", `pango:width=${textWidth}`, + "-define", "pango:alignment=left", + "-define", "pango:wrap=word", + `pango:@${markupPath}`, + ")", + "-gravity", "West", + "-geometry", `+${barX + 58}+0`, + "-composite", + outPath, + ]; + + await execFileP("magick", args, { maxBuffer: 1 << 24 }); + return outPath; +} + +async function main() { + const argv = process.argv.slice(2); + const manifestPath = argv.find((a) => !a.startsWith("--")); + if (!manifestPath) { + console.error("usage: render-cards.mjs <manifest.json> [--out <dir>] [--only <id>]"); + process.exit(2); + } + const flag = (name) => { + const i = argv.indexOf(name); + return i >= 0 ? argv[i + 1] : undefined; + }; + + const manifest = JSON.parse(await readFile(manifestPath, "utf8")); + const outDir = flag("--out") ?? path.join(path.dirname(path.resolve(manifestPath)), "out"); + const only = flag("--only"); + + await mkdir(path.join(outDir, "cards"), { recursive: true }); + + const cards = manifest.timeline.filter( + (e) => e.type === "card" && (!only || e.id === only), + ); + for (const card of cards) { + const p = await renderCard(card, manifest.render, outDir, manifest.timelineNodes); + console.log(`card ${card.id} -> ${p}`); + } + console.log(`${cards.length} card(s) rendered`); +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main().catch((err) => { + console.error(err); + process.exit(1); + }); +} diff --git a/scripts/report-to-video/resolve-windows.mjs b/scripts/report-to-video/resolve-windows.mjs @@ -0,0 +1,220 @@ +#!/usr/bin/env node +// resolve-windows.mjs — widen a manifest's clip windows to whole sentences. +// +// A manifest window starts life as the cue span covering a quote, and a cue +// boundary is a bad place to cut: ASR breaks cues where the caption line wrapped, +// which is routinely mid-sentence and often mid-word. Cutting there drops the +// lead-in that makes a quote make sense, and clips audibly start and stop in the +// middle of speech. +// +// This walks outward from the cue span to the nearest sentence boundary in the +// transcript — a cue whose text ends in . ? or ! — so the clip carries the whole +// thought. Word-level alignment is a separate, audio-side problem: build-video.mjs +// snaps the actual cut to a silence (see --fetch-pad / snapping there). +// +// Expansion is capped so a run-on passage can't drag a clip out to a minute. +// +// In the app: not used. On the CLI: +// node scripts/report-to-video/resolve-windows.mjs <manifest.json> [--write] +// +// Options: +// --write Rewrite the manifest in place (default: dry run, print a table) +// --max-lead <s> Max seconds to expand backwards (default 9) +// --max-tail <s> Max seconds to expand forwards (default 12) +// +// A clip entry may set `lockStart` / `lockEnd` to pin that edge exactly. + +import { readFile, writeFile } from "node:fs/promises"; +import path from "node:path"; + +const CHANNELS_DIR = + process.env.CHANNELS_DIR ?? + "/home/user/Projects/yt-dlp-transcript-browser/transcripts/channels"; + +const ENDS_SENTENCE = /[.!?]["'”’)\]]*\s*$/; + +// A cue that is only "[music]" or "[ __ ]" (the profanity bleep) carries no +// sentence signal; treat it as transparent so expansion walks past it. +const IS_FILLER = /^\s*(\[[^\]]*\]|>>|♪|—|-)*\s*$/; + +// The manifest stores times rounded to 2 dp, so a value read back from it can sit +// a hair BELOW the cue end it came from. Without a tolerance the end lookup then +// lands on the previous cue, the forward search runs on to the next sentence, and +// the clip grows a little every time this is run — it has to be a fixed point. +const EPS = 0.02; + +function loadCues(videoId, channelSlug) { + const p = path.join(CHANNELS_DIR, channelSlug, "data", videoId, "transcript.cues.json"); + return readFile(p, "utf8").then((s) => JSON.parse(s).cues); +} + +function indexAt(cues, t, which) { + // First cue whose span contains t, else the nearest one on the right side. + let idx = cues.findIndex((c) => c.end > t); + if (idx < 0) idx = cues.length - 1; + if (which === "end") { + let j = cues.findIndex((c) => c.end >= t - EPS); + if (j < 0) j = cues.length - 1; + idx = j; + } + return idx; +} + +export function widen(cues, start, end, { maxLead = 8, maxTail = 12 } = {}) { + const isBoundary = (c) => ENDS_SENTENCE.test(c.text) && !IS_FILLER.test(c.text); + const i0 = indexAt(cues, start, "start"); + const i1 = indexAt(cues, end, "end"); + + // START: the latest cue that OPENS a sentence (i.e. its predecessor closes + // one) at or before the quote, within the lead budget. Finding no such cue + // means every candidate lead-in is a sentence fragment, so take none at all — + // a fragment is the irrelevant context we are trying to avoid, not context. + let si = null; + for (let i = i0; i > 0; i -= 1) { + if (start - cues[i].start > maxLead) break; + if (isBoundary(cues[i - 1])) { + si = i; + break; + } + } + if (si === null) si = i0; + + // END: the first cue that CLOSES a sentence at or after the quote. Never + // clamp to a budget here — stopping partway through a sentence is exactly the + // mid-thought ending this is meant to remove, so the budget only decides how + // far to look, and failing to find one falls back to the original cue end. + let ei = null; + for (let j = i1; j < cues.length; j += 1) { + if (cues[j].end - end > maxTail) break; + if (isBoundary(cues[j])) { + ei = j; + break; + } + } + if (ei === null) ei = i1; + + return { + start: cues[si].start, + end: cues[ei].end, + leadCues: i0 - si, + tailCues: ei - i1, + }; +} + +async function main() { + const argv = process.argv.slice(2); + const manifestPath = argv.find((a) => !a.startsWith("--")); + if (!manifestPath) { + console.error("usage: resolve-windows.mjs <manifest.json> [--write]"); + process.exit(2); + } + const num = (name, dflt) => { + const i = argv.indexOf(name); + return i >= 0 ? Number(argv[i + 1]) : dflt; + }; + // Lead is where the context lives — it is the run-up that makes a quote make + // sense. Tail only needs to finish the sentence, so it gets a smaller budget. + const opts = { maxLead: num("--max-lead", 8), maxTail: num("--max-tail", 12) }; + + const manifest = JSON.parse(await readFile(manifestPath, "utf8")); + const slug = manifest.provenance.channelSlug; + const cache = new Map(); + + let changed = 0; + for (const e of manifest.timeline) { + if (e.type !== "clip") continue; + // A compilation can span several archived channels (the same streamer's VODs + // are mirrored across more than one), so a clip may name its own. Key the + // cache by channel too — the same id under a different slug is a different file. + // An author can trim a clip to land mid-cue on purpose — a cue often carries + // a whole paragraph, and cutting a quote short is an editorial decision. + // Widening would undo exactly that, so `lock` opts the clip out. + if (e.lock) { + console.log(`${e.id.padEnd(4)} ${e.video.padEnd(12)} locked, left at ${e.start.toFixed(1)}–${e.end.toFixed(1)}`); + continue; + } + const chan = e.channel ?? slug; + const key = `${chan}/${e.video}`; + if (!cache.has(key)) cache.set(key, await loadCues(e.video, chan)); + const cues = cache.get(key); + + const before = { start: e.start, end: e.end }; + const w = widen(cues, e.start, e.end, opts); + + // `lockStart` / `lockEnd` pin an edge to exactly what the author wrote. The + // escape hatch exists because sentence detection is only as good as the ASR's + // punctuation, and some uploads have none at all — and because an utterance's + // real trailing pause does not always line up with its last cue's end. + if (e.lockStart) w.start = before.start; + if (e.lockEnd) w.end = before.end; + const dLead = (before.start - w.start).toFixed(1); + const dTail = (w.end - before.end).toFixed(1); + const dur = (w.end - w.start).toFixed(1); + + // Ignore sub-frame drift so a re-run on an already-resolved manifest is a + // genuine no-op rather than a rewrite that nudges every window. + const moved = + Math.abs(w.start - before.start) > 0.05 || Math.abs(w.end - before.end) > 0.05; + if (moved) changed += 1; + console.log( + `${e.id.padEnd(4)} ${e.video.padEnd(12)} ` + + `${before.start.toFixed(1)}–${before.end.toFixed(1)} -> ` + + `${w.start.toFixed(1)}–${w.end.toFixed(1)} (+${dLead}s lead, +${dTail}s tail, ${dur}s)`, + ); + + if (moved) { + e.start = Number(w.start.toFixed(2)); + e.end = Number(w.end.toFixed(2)); + } + } + + // De-overlap clips that come from the SAME video. Widening is per-clip and + // blind to its neighbours, so a tail that finds no sentence boundary runs to + // the budget and can swallow the next clip's material — which plays as the + // same footage twice. (Real case: a 2024 upload whose ASR carries no + // punctuation at all in that stretch, so nothing stopped the search.) + // The later clip's start is the deliberate one, so trim the earlier clip's tail. + const byVideo = new Map(); + for (const e of manifest.timeline) { + if (e.type !== "clip") continue; + if (!byVideo.has(e.video)) byVideo.set(e.video, []); + byVideo.get(e.video).push(e); + } + for (const [video, list] of byVideo) { + if (list.length < 2) continue; + list.sort((a, b) => a.start - b.start); + for (let i = 0; i < list.length - 1; i += 1) { + const a = list[i]; + const b = list[i + 1]; + if (a.end <= b.start) continue; + const overlap = a.end - b.start; + if (a.lockEnd) { + console.log(` ⚠ ${a.id} overlaps ${b.id} by ${overlap.toFixed(1)}s but has lockEnd — not trimmed`); + continue; + } + a.end = Number(b.start.toFixed(2)); + changed += 1; + console.log( + ` de-overlap ${video}: ${a.id} trimmed ${overlap.toFixed(1)}s off its tail ` + + `(it ran into ${b.id})`, + ); + if (a.end - a.start < 3) { + console.log(` ⚠ ${a.id} is now only ${(a.end - a.start).toFixed(1)}s — check it`); + } + } + } + + if (argv.includes("--write")) { + await writeFile(manifestPath, JSON.stringify(manifest, null, 2) + "\n", "utf8"); + console.log(`\nwrote ${manifestPath} (${changed} window(s) changed)`); + } else { + console.log(`\ndry run — ${changed} window(s) would change; pass --write to apply`); + } +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main().catch((err) => { + console.error(err); + process.exit(1); + }); +} diff --git a/umtool/docs/README.md b/umtool/docs/README.md @@ -0,0 +1,68 @@ +# umtool docs + +umtool judges and drives the things this repo makes videos out of. There are two +kinds of work in it today and the tool treats them the same way: as **projects** +in a folder tree, each with a state, a set of open decisions, and a build. + +These sheets live in the repo because they describe code and have to move with it. +Notes about *one particular video* belong beside that video, in its own `README.md`. + +## You have been asked for an umtool video + +Work out which kind you are making first — the rest follows from it. + +| What you were handed | Kind | Start here | +|---|---|---| +| A cited sweep report (`*sweep-report.md`) and "make this a video" | `report-video` | [report-video.md](report-video.md), then [authoring.md](authoring.md) | +| A song's clips and "judge these" | `song` (template `um-song`) | `~/reports/quartering-uh-song/specs/` | +| A report with no manifest yet | `sweep-report` | [authoring.md](authoring.md) | + +The short path from a cited report to a built video: + +```sh +# 1. write video.manifest.json beside the report (authoring.md — this is the work) +# 2. is every source still fetchable, and is every citation wired up? +umtool check ~/reports/<slug> +# 3. widen windows to whole sentences (dry first, then apply) +node scripts/report-to-video/resolve-windows.mjs ~/reports/<slug>/video.manifest.json +node scripts/report-to-video/resolve-windows.mjs ~/reports/<slug>/video.manifest.json --write +# 4. a fast pass to look at, then the real one +umtool build <slug> --preset fast +umtool build <slug> --preset final +``` + +**Run step 2 before step 4, always.** It is a few seconds and it catches the two +defects that have already shipped in real videos: a manifest with no `siteOrigin` +(19 QR codes encoding `undefined/?v=…`) and one pointing at `http://localhost:3000` +(QR codes that resolve to nothing on anybody's phone). + +## The sheets + +| Sheet | What it covers | +|---|---| +| [projects.md](projects.md) | Kinds vs templates, marker files, ids, how to add a kind | +| [folders.md](folders.md) | The walk, `REPORTS_ROOT`, read roots vs write roots | +| [browse.md](browse.md) | The project index, the four filters, cards per kind | +| [report-video.md](report-video.md) | The manifest as an EDL, the three-stage window model, `lock` | +| [clip-bench.md](clip-bench.md) | Editing a clip's window against the waveform and the cues | +| [build.md](build.md) | The four-step chain, presets, cancellation, overwrite | +| [decisions.md](decisions.md) | What earns a severity, and how a kind contributes | +| [mix-from-a-project.md](mix-from-a-project.md) | Deep-linking a clip into `/mix` | +| [index.md](index.md) | The LMDB index, and why the filesystem stays the model | +| [cli.md](cli.md) | `umtool ls / show / check / build / window / …` | +| [authoring.md](authoring.md) | Writing a manifest from a sweep report | +| [e2e.md](e2e.md) | The fixture, the stubs, the global queue | +| [quirks.md](quirks.md) | Everything that cost time to find out | + +## The rules that outrank convenience + +1. **The filesystem is the model.** An index may cache what the tree says; if a + value exists *only* in the index, that is a bug. See [index.md](index.md). +2. **The index must not shell out.** Listing projects never probes, never runs + ffprobe, never runs yt-dlp. Measuring is what a project page and a job do. +3. **Judgements travel with the tree.** `verdicts.json`, `notes.json` and + `video.manifest.json` live *inside* the project directory, so copying the + directory copies the decisions. +4. **Never guess at something you cannot read.** A directory whose name does not + route, a project two kinds match, a citation with no cue file — each is + *reported*, never silently dropped or resolved by picking one. diff --git a/umtool/docs/quirks.md b/umtool/docs/quirks.md @@ -0,0 +1,129 @@ +# Quirks + +Things that cost time to find out. Each one is here because it was discovered by +getting it wrong, and none of them is guessable from the code. + +## Fetching source clips + +**yt-dlp picks VP9 + Opus at these heights unless you pin the format.** Since +`--force-keyframes-at-cuts` re-encodes, that means `libvpx-vp9`: 27 seconds to cut +a 5-second clip. It also writes `.webm` and appends that to `-o`, so the file you +asked for is not the file on disk. Pin `bv*[vcodec^=avc1][height<=N]+ba[acodec^=mp4a]` +and `--merge-output-format mp4`. + +**`--ignore-config` is not optional.** The operator's own yt-dlp config redirects +output and attaches thumbnail and metadata post-processors. Without it the clips +land somewhere else entirely and the build reports success having produced nothing +where it was looking. + +**yt-dlp exit 101 is success.** It is the clean early stop (`break-on-existing`, +`--max-downloads`). Treat it as success, as the rest of the repo does. + +**Rumble HLS needs `-extension_picky 0`, as a RETRY and never as a default.** +Rumble serves HLS whose segments are named `.tar`, which ffmpeg 8 rejects outright +("URL … is not in allowed_segment_extensions"), killing the fetch with exit 183 — +and Rumble ships no progressive fallback, so every Rumble clip is unbuildable +without it. But the option lives on the **HLS demuxer**: pass it against a +progressive URL (YouTube's googlevideo mp4) and ffmpeg aborts with "Option +extension_picky not found". Adding it unconditionally trades a Rumble failure for +a YouTube one. + +**`--force-keyframes-at-cuts` matters because the clip IS the citation.** Without +it the cut snaps to the nearest preceding keyframe, which can be seconds early. +Fine for scrubbing; not fine when someone is checking your quote. + +## Cutting + +**The silence threshold has to be relative to the clip, not absolute.** These are +game streams: the gaps between words are full of game audio and music. Measured on +a typical clip — mean volume −21 dB, **0** silences found at −32 dB, **25** at +−26 dB. Measure with `volumedetect` first and cut a few dB under the clip's own +mean. + +**Snapping only proves the edges are quiet.** It finds gaps in audio, which is +usually a word boundary and is not guaranteed to be. A speaker who does not pause +gets the unsnapped cut. + +**ASR cue boundaries are line-wrap boundaries, not sentence boundaries.** They land +mid-sentence routinely and mid-word often. That is the whole reason +`resolve-windows.mjs` exists — and the reason a clip bench that shows you the cues +is worth more than one that only shows you a waveform. + +**Some uploads carry no punctuation at all.** Sentence widening then has nothing to +find and silently does nothing. The clip bench says so explicitly rather than +leaving you to wonder why "extend to sentence end" is inert; set the edge by ear +and lock it. + +## Manifests + +**`resolve-windows.mjs` is a fixed point, and that was not free.** The manifest +stores times at 2 dp, so a value read back can sit a hair below the cue end it came +from — which lands the end lookup on the *previous* cue, runs the forward search on +to the *next* sentence, and grows the same clip a little on every run. Hence `EPS` +in the lookup and the 0.05 s deadband on applying a change. **Anything that writes a +window must round to 2 dp**, or it reintroduces exactly this bug. + +**`writeJsonAtomic` would vandalise a manifest.** It writes `JSON.stringify(v, null, 1)`; +`resolve-windows.mjs` writes `null, 2` plus a trailing newline. Saving one window +edit through the default writer reformats 600 lines and makes the diff unreadable. +Manifest writes pass `{space: 2, newline: true}`. + +**`lock: true` is the norm, not the exception.** Measured across the six real +manifests: ferret-rescue locks 1 of 10, every other manifest locks **100%**. A +human-chosen window usually *is* the truth, and widening it would undo an editorial +decision — cutting a quote short is a choice, and a single ASR cue often carries a +whole paragraph. So `lock` is a first-class explained control in the bench, not an +advanced toggle, and moving an edge to somewhere `widen()` would not produce offers +to set the matching lock. + +**`siteOrigin` is unvalidated and has already shipped broken twice.** +`quartering-employee-count` has none — 19 QR codes encoding `undefined/?v=…` — and +`ferret-rescue` has `http://localhost:3000`, a shipped video whose codes resolve to +nothing on anyone's phone. `umtool check` exists largely for this. + +**A clip may name its own `channel`.** The same streamer's VODs are mirrored across +several archived channels, and cue files are keyed by channel, so one manifest-wide +slug cannot find them all. + +**Rumble ids: the manifest's `video` must be the local directory slug.** Cue files +live under the URL slug, not the MCP video id. A manifest using the id finds +nothing — and finds it twenty minutes into a build, unless `umtool check` ran first. + +## Rendering + +**ImageMagick's `-size` leaks into Pango.** It applies to the *next* image +operation, so a stale `-size` silently changes the text raster. + +**drawtext does not wrap, and commas are structural in a filtergraph.** Wrap to a +character budget, write the text to a file and use `textfile=`, so nothing needs +shell or filter escaping. Single-quote any expression containing a comma. + +**Stream titles need emoji and `!command` suffixes stripped** or they render as +tofu in the attribution line. + +**Segments are encoded to identical parameters on purpose**, so the final concat is +a stream copy. Mismatched streams are the usual reason a naive concat produces a +broken or audio-desynced file. + +**A QR must be fully opaque and must keep its quiet zone.** A translucent QR will +not scan, and the white border is part of the symbol, not decoration. + +**ffmetadata is line-based and `=`, `;`, `#`, `\` are structural.** A chapter title +carrying any of them has to be escaped or the file silently mis-parses. + +## The tool itself + +**`node -e console.log(<number>)` emits ANSI escapes on a TTY**, which corrupt +ffmpeg filtergraphs and shell tests silently. + +**`pnpm lint` in `editor/` always fails** — there is no eslint config there. Use +`pnpm exec tsc --noEmit`. + +**A value import that drags `node:fs` into a client component 500s every page** and +passes typecheck. `pnpm build`, not just `tsc --noEmit`, is what catches it — which +is why the registry's *types* live in `lib/project-types.ts`, separate from the +`.mjs` that reads the disk. + +**e2e is serialized machine-wide.** A "waiting for the e2e queue" banner is normal, +not a hang; the serial suite is long. A run that wins the lock but finds its ports +bound aborts and names the offending pid. diff --git a/umtool/docs/report-video.md b/umtool/docs/report-video.md @@ -0,0 +1,160 @@ +# The `report-video` kind + +A report video is a **cited timeline**: a sweep report's findings, said by the +sources in their own voice, in order, each clip carrying a burned-in attribution +and a QR that resolves to that exact moment in the archive's own viewer. The point +is that a viewer does not have to take the edit on trust. + +The pipeline lives at `scripts/report-to-video/` and its own +[README](../../scripts/report-to-video/README.md) is the source of truth for how it +renders. This sheet covers what that README cannot know: what umtool reads, what +umtool **writes**, and the rules a UI has to honour so the CLI and the app can never +disagree. + +## The manifest is an EDL, and it is the project + +`<project>/video.manifest.json` is both the marker file for the kind and the edit +decision list. Nothing else needs to exist for a directory to be a report video. + +```jsonc +{ + "schemaVersion": 1, + "slug": "quartering-gout", // out/<slug>.mp4 is the deliverable + "title": "…", "subtitle": "…", "generatedOn": "2026-08-12", + "provenance": { … }, // how the sweep was done, and its caveats + "render": { … }, // resolution, fonts, palette, the cut knobs + "timelineNodes": [ … ], // the footer's progress track, if any + "timeline": [ … ] // the ordered cut. Array order IS the cut. +} +``` + +**Ordering is array order.** There is no `index` field and no sort. A reorder is a +move within the array, and it has to recompute `sectionEnter` — the flag that says +"this clip is the first of its section", which is what makes the footer marker +travel. umtool never reorders as a side effect of a window edit. + +### A clip entry + +```jsonc +{ "type": "clip", "id": "c04", "video": "uyz1_FIqIEk", + "channel": "hasanabi-vods3", // optional: which archived channel's cues + "start": 32980.24, "end": 32994.19, // absolute source seconds, 2 dp + "cite": 32989, // the second shown in the attribution line + "citeUrl": "https://…", // optional: overrides the derived QR target + "quote": "…", // the words this clip exists for + "note": "…", // why it is in the cut (editorial, for humans) + "chapter": "…", // chapter title; falls back to date + title + "section": 2, "sectionEnter": true, + "lock": true, "lockStart": true, "lockEnd": true } +``` + +A card entry is `{"type":"card", "id", "style", "seconds", …}` — see the pipeline +README for the styles. Cards have no window and no source. + +## The three-stage window model + +A clip's window passes through three different notions of "where the cut is", and +confusing them is the source of most window bugs. + +| Stage | Where it comes from | Whose problem | +|---|---|---| +| **1. Cue span** | `transcript.cues.json` — the cues covering the quote | The sweep. A report only records a single start second, so a window cannot be recovered from the report alone. | +| **2. Sentence window** | `resolve-windows.mjs` walks outward to a cue ending in `.?!` | Meaning. A cue boundary is a *line-wrap* boundary; cutting there drops the lead-in that makes a quote make sense. | +| **3. Audio cut** | `build-video.mjs` snaps to a silence found in the fetched audio | Sound. A sentence boundary is still not a *speech* boundary; snapping is what stops clips cutting through a word. | + +Only stage 2 is stored. Stage 1 is recoverable from the cue file, and stage 3 is +recomputed on every build from the audio — which is why a build is reproducible +from the manifest alone, and why the clip bench edits stage 2 and shows you stage 1. + +**Everything is in absolute source seconds**, from the manifest through the bench +to the API. The fetched file's own start (`fetchStart`) is the only place a relative +number appears, and it is always derived, never stored. + +## Writing a window: the four rules + +umtool is a *second* writer of a file the CLI also writes. So: + +1. **Round to 2 dp.** Non-negotiable. `resolve-windows.mjs` is a fixed point and + both its `EPS` lookup tolerance and its 0.05 s deadband assume 2 dp storage. + Writing 4 dp makes the widener grow the clip on every subsequent run. +2. **Re-read before writing, under `withStateLock`, and write atomically.** The + file has two writers; a lost update here is lost human judgement. +3. **Preserve the CLI's formatting** — `JSON.stringify(m, null, 2) + "\n"`. The + default `writeJsonAtomic` uses indent 1, which turns a two-number edit into a + 600-line diff. +4. **Guard with an mtime token.** A `PUT` carrying a stale token is a 409, not a + silent overwrite — an agent or a CLI run may have written in between. + +A `.bak` of the last hand-authored state is kept (rate-limited, not a rolling +stack); `ferret-rescue/video.manifest.json.bak` already set that precedent. + +## `lock`, `lockStart`, `lockEnd` + +These are **editorial acknowledgements**, not advanced options, and the data says +so: across the six real manifests ferret-rescue locks 1 clip of 10 and *every other +manifest locks 100%*. + +- **`lock`** — resolve-windows must not touch this clip at all. The two cases: + the author deliberately cut a quote short (a single ASR cue often carries a whole + paragraph, so trimming to the sentence that matters is a real decision that + widening would undo), or the lead-in would drag in seconds of some *other* audio. + A live case from `quartering-gout`: two clips whose source cue file has two cues + spanning almost the whole runtime, where an unlocked widening pass would expand a + 35-second window to a 2,488-second one. +- **`lockStart` / `lockEnd`** — pin one edge exactly. Needed because sentence + detection is only as good as the ASR's punctuation, and some uploads have none. + +`lockEnd` is also how the "ends mid-sentence" warning is *acknowledged*: setting it +is the author saying "I meant to cut here", and it suppresses the warning. This +matters — one 14-clip cut shipped with 8 clips ending mid-thought. + +**The rule that prevents silent loss:** if you move an edge in the bench to a value +`widen()` would not produce, the bench offers to set the matching lock, defaulted +on. Without it, the next `resolve-windows --write` reverts your edit. + +## Provenance, and the two defects that shipped + +`provenance` is prose written by the sweep, and most of it is for humans. Three +fields are load-bearing: + +- **`siteOrigin`** — the archive origin every QR is built from. **Nothing validated + this**, and two real videos shipped broken: `quartering-employee-count` has no + `siteOrigin` at all (19 QR codes encoding `undefined/?v=…`) and `ferret-rescue` + has `http://localhost:3000` (codes that resolve to nothing on a phone). This is + the single best reason to run `umtool check` before every build. +- **`channelSlug`** — the default archived channel for cue lookups, overridable per + clip. +- **`availabilityCheckedOn`** — when sources were last confirmed fetchable. Stale in + both directions; `check-availability.mjs` refreshes it. + +## The Rumble two-ids rule + +A Rumble video has **two** ids: the site/MCP `video_id` is the *embed* id, while the +local cue directory is named for the **URL slug**. The manifest's `video` field must +be the local slug; the report cites the site id. A manifest that uses the site id +finds no cue file — and finds that out twenty minutes into a build, unless +`umtool check` ran first. `citeUrl` exists for the mirror case: a clip taken from a +copy whose archived transcript is broken should point the QR at the copy that reads. + +## Deleted sources + +A source can be deleted from YouTube after the manifest is written, and clips are a +**live network fetch** — there is no local media. The archive's own availability +data is a *snapshot*, so it goes stale in both directions. Options, in order of +preference: cut the same moment from a live mirror (via a `CHANNELS_DIR` shadow +directory — but check the archive's alignment statement first, because mirrors are +not assumed to share a clock), or convert the clip to a quote card. + +## Discovered by getting it wrong once + +- **A build that loses a clip is worse than a build that fails.** `--continue-on-error` + finishes everything buildable and then *refuses to concatenate*, because a + finished file quietly missing a citation looks complete. +- **The cache is content-addressed by window, and that was being wasted.** + Before containing-window reuse, every window edit was a fresh download: + `ferret-rescue/out/clips-raw` holds 31 files for 10 clips, one source four times + over overlapping windows. +- **Mirrors are not assumed to share a clock.** Timestamps mapped from one mirror to + another have to be verified against the mirror's own cue text, not assumed. +- **`--chapters-only` is cheap and separate.** Retitling chapters does not need a + re-encode; the per-clip segments on disk are all the offsets need.