# report-to-video Turns a cited sweep report into a video: the clips run in chronological order and let the source speak for itself, with thin chrome carrying the citation and a timeline of where you are. A video rendering of the reports we already write. Two scripts and a manifest: | file | lifetime | what it is | |---|---|---| | `build-video.mjs` | stable | manifest → mp4. Fetches clips, snaps cuts to silence, letterboxes them into the chrome, crossfades. | | `render-cards.mjs` | stable | draws the timeline footer and marker, plus optional card stills. Imported by `build-video.mjs`. | | `resolve-windows.mjs` | stable | widens clip windows from cue spans to whole sentences. Run once after authoring a manifest. | | `brand.mjs`, `brand-cards.mjs` | stable | the opt-in brand preset (`render.brand`); see [its section](#the-archilyzer-media-preset-renderbrand). | | `/video.manifest.json` | per report | the edit decision list. **This is the regeneration source of truth**, not the report. | ``` node umtool/report-to-video/resolve-windows.mjs ~/reports//video.manifest.json --write node umtool/report-to-video/build-video.mjs ~/reports//video.manifest.json ``` Output lands in `/out/`: `cards/`, `clips-raw/`, `segments/`, and the finished `.mp4`. First worked example: `~/reports/ferret-rescue/`. ## Why there is a manifest at all **A sweep report does not contain enough information to cut a video from.** Its citations carry a single start second and nothing else — `momentUrl()` (`common/lib/momentUrl.ts:106`) takes one `seconds` and floors it, and the MCP `Snippet` type (`mcp/src/search.ts:41`) has no `end` field. There is no clip length anywhere in a report. The end times do exist, they are just never rendered: every cue in `transcripts/channels//data//transcript.cues.json` is `{start, end, text}`. So the manifest is built by matching each quote back to its covering cues and recording the real window. That is also what handles ellipsis-joined citations — a report quote like `"…" … "…"` is often two separate cue spans presented as one. The manifest additionally pins the provenance (share link, corpus handle, match counts, the narrowing queries) so a rebuild months later is reproducible and the video's own claims about its coverage can be checked. ## Where a clip actually gets cut Three stages, because a cue span is the wrong answer twice over. 1. **Cue span** — the raw window covering the quote, from `transcript.cues.json`. 2. **Sentence widening** (`resolve-windows.mjs`) — walk outward to the nearest cue ending in `.`, `?` or `!`. A cue boundary is where the *caption line wrapped*, so cutting there drops the run-up that makes a quote intelligible. Two asymmetries matter. The **start** takes a lead-in only if a real sentence opening sits within `--max-lead` (8 s); otherwise it takes none, because a half-sentence run-up is the irrelevant context you were trying to avoid, not context. The **end** is never clamped to a budget — stopping partway through a sentence is the exact mid-thought ending this removes — so `--max-tail` (12 s) only bounds how far it looks before giving up and using the cue end. **Some uploads have no punctuation at all.** Older ASR in this corpus emits unpunctuated cue text for whole videos, and sentence detection then has nothing to find: widening degrades to the raw cue span at both ends. That is not a silent failure you can ignore — it is what produced a clip opening mid-thought on "higher than they can afford and because", and what let another clip's tail run through its neighbour. For those videos, pick the window by reading the cues and set it by hand; the de-overlap and silence passes still apply. **`lockStart` / `lockEnd` pin an edge** to exactly what the manifest says, and neither widening nor de-overlap will move it. Reach for it when the utterance's real trailing pause does not line up with its last cue's end — the snap picks the *nearest* silence, and in speech over game audio the nearest one is often a gap between syllables rather than the pause at the end of the thought. 3. **Silence snapping** (`build-video.mjs`) — a sentence boundary in the *transcript* still isn't a boundary in the *audio*, so clips clip words in half. Fetch `fetchPad` seconds wider than needed, run `silencedetect` over the result, and move each cut to the nearest silence within `snapWindow`. Starts land on a silence's END (just before speech resumes), ends on a silence's START (just after speech stops). No silence close enough → keep the exact point; a tight cut beats a cut in the wrong place. The build logs `start✓ end✓` per clip so you can see which snapped. How much of the pause is kept: 0.10 s before speech resumes and 0.18 s after it stops by default. `render.snapLead` / `render.snapTail` (seconds) set them, and then never more than the pause holds — so with a crossfade, a tail as long as the `transition` fades out over the breath instead of the last word, and a long lead lets the next clip fade in before its first word. Pair them with a `silenceMinDur` that finds real pauses (~0.3 s), not the gaps between words. **The silence threshold is relative, and it has to be.** These are game streams: the gaps between words are full of game audio and music — quiet, but nowhere near silent. A fixed absolute threshold sits below the noise floor and finds nothing. On one measured clip: mean volume −21 dB, **0** silences at −32 dB, **25** at −26 dB. So each clip is measured with `volumedetect` first and the threshold set `silenceRelDb` (default 6) below its own mean. If a rebuild suddenly reports mostly `start– end–`, this is the knob. The trim happens during the burn-in encode, so snapping costs nothing extra. 4. **De-overlap** (`resolve-windows.mjs`) — widening is per-clip and blind to its neighbours, so two clips cut from the *same* video can end up overlapping, and the overlap plays as the same footage twice. Any earlier clip whose tail runs into a later clip's start is trimmed back to that start. This is not a rare edge case: it fired on the first report, where a 2024 upload has **no punctuation at all** in the relevant stretch, so sentence detection found nothing and the tail ran the full budget straight through the next clip. ## Regenerating and changing a video Everything is cached by content, so iteration is cheap: - **Reorder, drop or add clips** — edit `timeline`, re-run. Cached clips are not refetched, so a re-cut costs an encode, not a download. - **Change a clip's window** — edit `start`/`end`, re-run. A raw file's window is in its name, so the cache is content-addressed; and since a request is satisfied by any cached file that **contains** it, a nudge inside the existing pad costs nothing at all. Only a window that escapes every cached file downloads again. - **Preview one entry** — `--only ` builds a single segment and stops. - **Work offline** — `--skip-fetch` fails loudly instead of downloading, so you can confirm you are working entirely from cache. `--no-network` is the stronger form: every clip's source is found on disk **before** anything is rendered, and if even one clip would need a download the build refuses at once, listing each such clip as `timeline[] /