Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 69d0e9a6ca1495b8a4d01c6e5afe665ebf308b1a
parent 17db195a27e03e7315e39048d11bbffc71cadbe7
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon,  7 Sep 2026 15:44:44 -0400

umtool: report-to-video moves under the umtool umbrella

`git mv scripts/report-to-video umtool/report-to-video`, and the package is
renamed `report-to-video` -> `umtool-report-to-video`. umtool is the home of
every project kind (one-core plan, phase 5); this is the rename that makes the
later move a no-op.

A rename only — no behaviour, no bins, no exports change. `cues.mjs`'s
REPO_ROOT is still two levels up, so `DEFAULT_CHANNELS_DIR` is unchanged.

Updated: pnpm-workspace.yaml, the root `test:scripts` glob, umtool's
dependency, 8 import sites in umtool (`report-to-video/...` ->
`umtool-report-to-video/...`: app/api/report/build/route.ts,
components/projects/ClaimBench{,Page}.tsx, lib/projects/report.mjs,
lib/report/{export,manifest,serve}.mjs), 3 path references
(lib/report/driver.mjs PIPELINE_DIR, e2e/clip-bench.spec.ts,
bin/umtool.mjs's printed command), Dockerfile's manifest COPY, and the docs
(AGENTS.md, README.md, CONTRIBUTING.md, RUNNING_IN_DOCKER.md,
umtool/docs/{README,authoring,report-video}.md, the package README and the
six script header comments).

Verified: tree grep for `scripts/report-to-video` is empty apart from the
plan document's own prose. `pnpm install` relinks 8 workspace projects.
`pnpm test:scripts` 72 tests, 71 pass / 0 fail / 1 skipped.
`pnpm --filter umtool run typecheck` clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
MAGENTS.md | 4++--
MCONTRIBUTING.md | 2+-
MDockerfile | 2+-
MREADME.md | 6+++---
MRUNNING_IN_DOCKER.md | 2+-
Mpackage.json | 2+-
Mpnpm-lock.yaml | 46+++++++++++-----------------------------------
Mpnpm-workspace.yaml | 2+-
Dscripts/report-to-video/README.md | 786-------------------------------------------------------------------------------
Dscripts/report-to-video/build-video.mjs | 1838-------------------------------------------------------------------------------
Dscripts/report-to-video/check-availability.mjs | 182-------------------------------------------------------------------------------
Dscripts/report-to-video/cues.mjs | 270-------------------------------------------------------------------------------
Dscripts/report-to-video/package.json | 25-------------------------
Dscripts/report-to-video/render-cards.mjs | 1530-------------------------------------------------------------------------------
Dscripts/report-to-video/resolve-windows.mjs | 241-------------------------------------------------------------------------------
Dscripts/report-to-video/verify-build.mjs | 110-------------------------------------------------------------------------------
Mumtool/app/api/report/build/route.ts | 2+-
Mumtool/bin/umtool.mjs | 2+-
Mumtool/components/projects/ClaimBench.tsx | 2+-
Mumtool/components/projects/ClaimBenchPage.tsx | 2+-
Mumtool/docs/README.md | 4++--
Mumtool/docs/authoring.md | 4++--
Mumtool/docs/report-video.md | 6+++---
Mumtool/e2e/clip-bench.spec.ts | 2+-
Mumtool/lib/decisions.ts | 2+-
Mumtool/lib/projects/report.mjs | 6+++---
Mumtool/lib/report/driver.mjs | 2+-
Mumtool/lib/report/export.mjs | 2+-
Mumtool/lib/report/manifest.mjs | 2+-
Mumtool/lib/report/serve.mjs | 2+-
Mumtool/package.json | 2+-
Aumtool/report-to-video/README.md | 786+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aumtool/report-to-video/build-video.mjs | 1838+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aumtool/report-to-video/check-availability.mjs | 182+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Rscripts/report-to-video/compose-chrome.mjs -> umtool/report-to-video/compose-chrome.mjs | 0
Aumtool/report-to-video/cues.mjs | 270+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Rscripts/report-to-video/cues.test.mjs -> umtool/report-to-video/cues.test.mjs | 0
Rscripts/report-to-video/ledger-totals.mjs -> umtool/report-to-video/ledger-totals.mjs | 0
Rscripts/report-to-video/ledger-totals.test.mjs -> umtool/report-to-video/ledger-totals.test.mjs | 0
Aumtool/report-to-video/package.json | 25+++++++++++++++++++++++++
Aumtool/report-to-video/render-cards.mjs | 1530+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aumtool/report-to-video/resolve-windows.mjs | 241+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aumtool/report-to-video/verify-build.mjs | 110+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
43 files changed, 5024 insertions(+), 5048 deletions(-)

diff --git a/AGENTS.md b/AGENTS.md @@ -56,7 +56,7 @@ The high-value loop for a repo with no corpus: point the MCP at a public instanc about a subject, then use `yt-dlp --download-sections` to pull **just the cited seconds** rather than whole videos. Searching text first is what makes fetching cheap. -`scripts/report-to-video/` renders a cited sweep report to an mp4. What it needs: +`umtool/report-to-video/` renders a cited sweep report to an mp4. What it needs: - **yt-dlp** — clip media is fetched over the network per clip. Local media does not help: most archived video dirs hold captions and metadata, not video. @@ -65,7 +65,7 @@ seconds** rather than whole videos. Searching text first is what makes fetching boundary is where the caption line wrapped, so cutting there ends mid-thought. Neither a report nor an MCP snippet carries an end. -Cues resolve through `scripts/report-to-video/cues.mjs`: a local corpus when there is +Cues resolve through `umtool/report-to-video/cues.mjs`: a local corpus when there is one (`CHANNELS_DIR`, now resolved relative to the repo), else the published archive the manifest names in `provenance.siteOrigin`. The shard walk is the contract in `/corpus.json`: corpus → channel transcripts manifest → `slugToPage` → `page-<NNNN>.json` diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md @@ -21,7 +21,7 @@ A pnpm-workspace monorepo. Every package consumes `common/` via `workspace:*`. | `export/` | The public static site (`output: "export"`). Read-only, no runtime. | | `homepage/` | The project's own site: marketing, docs, downloads, cross-site stats at `/stats/`. | | `mcp/` | The MCP server that exposes a published archive to LLM clients. See [mcp/README.md](mcp/README.md). | -| `umtool/` | A local bench for the report-to-video work (port 3050). Not part of an archive install. | +| `umtool/` | A local bench for the report-to-video work (port 3050), including the `umtool/report-to-video/` pipeline package. Not part of an archive install. | Your corpus lives at `<repo>/transcripts/` and is **its own git repo**, untouched by the workspace. That separation is deliberate: updating the software never touches diff --git a/Dockerfile b/Dockerfile @@ -249,7 +249,7 @@ COPY editor/package.json editor/package.json COPY export/package.json export/package.json COPY homepage/package.json homepage/package.json COPY mcp/package.json mcp/package.json -COPY scripts/report-to-video/package.json scripts/report-to-video/package.json +COPY umtool/report-to-video/package.json umtool/report-to-video/package.json COPY umtool/package.json umtool/package.json # `allowBuilds` / `onlyBuiltDependencies` in pnpm-workspace.yaml rebuild the diff --git a/README.md b/README.md @@ -359,15 +359,15 @@ report is minutes of video rather than years of it. ### Rendering a report to video -`scripts/report-to-video/` turns a cited sweep report into an mp4: clips in +`umtool/report-to-video/` turns a cited sweep report into an mp4: clips in chronological order, thin chrome carrying the citation, and a timeline of where you are. The report is *not* the regeneration source — a per-report `video.manifest.json` is, because a report's citations carry a start second and no clip length, so the real windows are recovered by matching each quote back to its covering caption cues. ```sh -node scripts/report-to-video/resolve-windows.mjs <report>/video.manifest.json --write -node scripts/report-to-video/build-video.mjs <report>/video.manifest.json +node umtool/report-to-video/resolve-windows.mjs <report>/video.manifest.json --write +node umtool/report-to-video/build-video.mjs <report>/video.manifest.json ``` **Clip boundaries, and why this runs without a corpus.** Cutting on the raw cue span diff --git a/RUNNING_IN_DOCKER.md b/RUNNING_IN_DOCKER.md @@ -404,7 +404,7 @@ silently loses formats), `ffmpeg`/`ffprobe`, `whisper-cli` (statically linked), - **whisper models.** 142 MB to 3 GB, and the choice is yours. Fetched on boot. - **The export site build.** It is a static render *of a corpus*, and there is no corpus at image-build time. `docker/publish-site.sh` makes it at run time. -- **ImageMagick with Pango, and `qrencode`.** Only `scripts/report-to-video/` +- **ImageMagick with Pango, and `qrencode`.** Only `umtool/report-to-video/` needs them, and rendering a report to video is a workstation task, not something a server does. Run that part on a host checkout. - **gallery-dl.** The X/Twitter post fetcher. Install it and point diff --git a/package.json b/package.json @@ -21,7 +21,7 @@ "e2e": "node scripts/worktree.mjs run -- pnpm --filter editor run e2e", "wt": "node scripts/worktree.mjs", "e2e:sharded": "node scripts/run-sharded-e2e.mjs", - "test:scripts": "node --test scripts/*.test.mjs scripts/report-to-video/*.test.mjs", + "test:scripts": "node --test scripts/*.test.mjs umtool/report-to-video/*.test.mjs", "lint": "pnpm --filter export run lint" }, "devDependencies": { diff --git a/pnpm-lock.yaml b/pnpm-lock.yaml @@ -302,8 +302,6 @@ importers: specifier: ^5.6.0 version: 5.9.3 - scripts/report-to-video: {} - umtool: dependencies: class-variance-authority: @@ -324,12 +322,12 @@ importers: react-dom: specifier: 19.2.4 version: 19.2.4(react@19.2.4) - report-to-video: - specifier: workspace:* - version: link:../scripts/report-to-video tailwind-merge: specifier: ^3.6.0 version: 3.6.0 + umtool-report-to-video: + specifier: workspace:* + version: link:report-to-video yt-dlp-transcript-common: specifier: workspace:* version: link:../common @@ -356,6 +354,8 @@ importers: specifier: ^5.9.3 version: 5.9.3 + umtool/report-to-video: {} + packages: '@alloc/quick-lru@5.2.0': @@ -2670,6 +2670,7 @@ packages: eslint@9.39.4: resolution: {integrity: sha512-XoMjdBOwe/esVgEvLmNsD3IRHkm7fbKIUGvrleloJXUZgDHig2IPWNniv+GwjyJXzuNqVjlr5+4yVUZjycJwfQ==} engines: {node: ^18.18.0 || ^20.9.0 || >=21.1.0} + deprecated: This version is no longer supported. Please see https://eslint.org/version-support for other options. hasBin: true peerDependencies: jiti: '*' @@ -6360,7 +6361,7 @@ snapshots: '@next/eslint-plugin-next': 16.2.3 eslint: 9.39.4(jiti@2.6.1) eslint-import-resolver-node: 0.3.10 - eslint-import-resolver-typescript: 3.10.1(eslint-plugin-import@2.32.0(@typescript-eslint/parser@8.59.0(eslint@9.39.4(jiti@2.6.1))(typescript@5.9.3))(eslint@9.39.4(jiti@2.6.1)))(eslint@9.39.4(jiti@2.6.1)) + eslint-import-resolver-typescript: 3.10.1(eslint-plugin-import@2.32.0)(eslint@9.39.4(jiti@2.6.1)) eslint-plugin-import: 2.32.0(@typescript-eslint/parser@8.59.0(eslint@9.39.4(jiti@2.6.1))(typescript@5.9.3))(eslint-import-resolver-typescript@3.10.1)(eslint@9.39.4(jiti@2.6.1)) eslint-plugin-jsx-a11y: 6.10.2(eslint@9.39.4(jiti@2.6.1)) eslint-plugin-react: 7.37.5(eslint@9.39.4(jiti@2.6.1)) @@ -6403,21 +6404,6 @@ snapshots: transitivePeerDependencies: - supports-color - eslint-import-resolver-typescript@3.10.1(eslint-plugin-import@2.32.0(@typescript-eslint/parser@8.59.0(eslint@9.39.4(jiti@2.6.1))(typescript@5.9.3))(eslint@9.39.4(jiti@2.6.1)))(eslint@9.39.4(jiti@2.6.1)): - dependencies: - '@nolyfill/is-core-module': 1.0.39 - debug: 4.4.3 - eslint: 9.39.4(jiti@2.6.1) - get-tsconfig: 4.14.0 - is-bun-module: 2.0.0 - stable-hash: 0.0.5 - tinyglobby: 0.2.16 - unrs-resolver: 1.11.1 - optionalDependencies: - eslint-plugin-import: 2.32.0(@typescript-eslint/parser@8.59.0(eslint@9.39.4(jiti@2.6.1))(typescript@5.9.3))(eslint-import-resolver-typescript@3.10.1)(eslint@9.39.4(jiti@2.6.1)) - transitivePeerDependencies: - - supports-color - eslint-import-resolver-typescript@3.10.1(eslint-plugin-import@2.32.0)(eslint@9.39.4(jiti@2.6.1)): dependencies: '@nolyfill/is-core-module': 1.0.39 @@ -6429,27 +6415,17 @@ snapshots: tinyglobby: 0.2.16 unrs-resolver: 1.11.1 optionalDependencies: - eslint-plugin-import: 2.32.0(eslint-import-resolver-typescript@3.10.1)(eslint@9.39.4(jiti@2.6.1)) + eslint-plugin-import: 2.32.0(@typescript-eslint/parser@8.59.0(eslint@9.39.4(jiti@2.6.1))(typescript@5.9.3))(eslint-import-resolver-typescript@3.10.1)(eslint@9.39.4(jiti@2.6.1)) transitivePeerDependencies: - supports-color - eslint-module-utils@2.12.1(@typescript-eslint/parser@8.59.0(eslint@9.39.4(jiti@2.6.1))(typescript@5.9.3))(eslint-import-resolver-node@0.3.10)(eslint-import-resolver-typescript@3.10.1(eslint-plugin-import@2.32.0(@typescript-eslint/parser@8.59.0(eslint@9.39.4(jiti@2.6.1))(typescript@5.9.3))(eslint@9.39.4(jiti@2.6.1)))(eslint@9.39.4(jiti@2.6.1)))(eslint@9.39.4(jiti@2.6.1)): + eslint-module-utils@2.12.1(@typescript-eslint/parser@8.59.0(eslint@9.39.4(jiti@2.6.1))(typescript@5.9.3))(eslint-import-resolver-node@0.3.10)(eslint-import-resolver-typescript@3.10.1)(eslint@9.39.4(jiti@2.6.1)): dependencies: debug: 3.2.7 optionalDependencies: '@typescript-eslint/parser': 8.59.0(eslint@9.39.4(jiti@2.6.1))(typescript@5.9.3) eslint: 9.39.4(jiti@2.6.1) eslint-import-resolver-node: 0.3.10 - eslint-import-resolver-typescript: 3.10.1(eslint-plugin-import@2.32.0(@typescript-eslint/parser@8.59.0(eslint@9.39.4(jiti@2.6.1))(typescript@5.9.3))(eslint@9.39.4(jiti@2.6.1)))(eslint@9.39.4(jiti@2.6.1)) - transitivePeerDependencies: - - supports-color - - eslint-module-utils@2.12.1(eslint-import-resolver-node@0.3.10)(eslint-import-resolver-typescript@3.10.1)(eslint@9.39.4(jiti@2.6.1)): - dependencies: - debug: 3.2.7 - optionalDependencies: - eslint: 9.39.4(jiti@2.6.1) - eslint-import-resolver-node: 0.3.10 eslint-import-resolver-typescript: 3.10.1(eslint-plugin-import@2.32.0)(eslint@9.39.4(jiti@2.6.1)) transitivePeerDependencies: - supports-color @@ -6465,7 +6441,7 @@ snapshots: doctrine: 2.1.0 eslint: 9.39.4(jiti@2.6.1) eslint-import-resolver-node: 0.3.10 - eslint-module-utils: 2.12.1(@typescript-eslint/parser@8.59.0(eslint@9.39.4(jiti@2.6.1))(typescript@5.9.3))(eslint-import-resolver-node@0.3.10)(eslint-import-resolver-typescript@3.10.1(eslint-plugin-import@2.32.0(@typescript-eslint/parser@8.59.0(eslint@9.39.4(jiti@2.6.1))(typescript@5.9.3))(eslint@9.39.4(jiti@2.6.1)))(eslint@9.39.4(jiti@2.6.1)))(eslint@9.39.4(jiti@2.6.1)) + eslint-module-utils: 2.12.1(@typescript-eslint/parser@8.59.0(eslint@9.39.4(jiti@2.6.1))(typescript@5.9.3))(eslint-import-resolver-node@0.3.10)(eslint-import-resolver-typescript@3.10.1)(eslint@9.39.4(jiti@2.6.1)) hasown: 2.0.3 is-core-module: 2.16.1 is-glob: 4.0.3 @@ -6494,7 +6470,7 @@ snapshots: doctrine: 2.1.0 eslint: 9.39.4(jiti@2.6.1) eslint-import-resolver-node: 0.3.10 - eslint-module-utils: 2.12.1(eslint-import-resolver-node@0.3.10)(eslint-import-resolver-typescript@3.10.1)(eslint@9.39.4(jiti@2.6.1)) + eslint-module-utils: 2.12.1(@typescript-eslint/parser@8.59.0(eslint@9.39.4(jiti@2.6.1))(typescript@5.9.3))(eslint-import-resolver-node@0.3.10)(eslint-import-resolver-typescript@3.10.1)(eslint@9.39.4(jiti@2.6.1)) hasown: 2.0.3 is-core-module: 2.16.1 is-glob: 4.0.3 diff --git a/pnpm-workspace.yaml b/pnpm-workspace.yaml @@ -4,7 +4,7 @@ packages: - export - homepage - mcp - - scripts/report-to-video + - umtool/report-to-video - umtool allowBuilds: diff --git a/scripts/report-to-video/README.md b/scripts/report-to-video/README.md @@ -1,786 +0,0 @@ -# report-to-video - -Turns a cited sweep report into a video: the clips run in chronological order and -let the source speak for itself, with thin chrome carrying the citation and a -timeline of where you are. A video rendering of the reports we already write. - -Two scripts and a manifest: - -| file | lifetime | what it is | -|---|---|---| -| `build-video.mjs` | stable | manifest → mp4. Fetches clips, snaps cuts to silence, letterboxes them into the chrome, crossfades. | -| `render-cards.mjs` | stable | draws the timeline footer and marker, plus optional card stills. Imported by `build-video.mjs`. | -| `resolve-windows.mjs` | stable | widens clip windows from cue spans to whole sentences. Run once after authoring a manifest. | -| `<report>/video.manifest.json` | per report | the edit decision list. **This is the regeneration source of truth**, not the report. | - -``` -node scripts/report-to-video/resolve-windows.mjs ~/reports/<slug>/video.manifest.json --write -node scripts/report-to-video/build-video.mjs ~/reports/<slug>/video.manifest.json -``` - -Output lands in `<report dir>/out/`: `cards/`, `clips-raw/`, `segments/`, and the -finished `<slug>.mp4`. First worked example: `~/reports/ferret-rescue/`. - -## Why there is a manifest at all - -**A sweep report does not contain enough information to cut a video from.** Its -citations carry a single start second and nothing else — `momentUrl()` -(`common/lib/momentUrl.ts:106`) takes one `seconds` and floors it, and the MCP -`Snippet` type (`mcp/src/search.ts:41`) has no `end` field. There is no clip -length anywhere in a report. - -The end times do exist, they are just never rendered: every cue in -`transcripts/channels/<slug>/data/<id>/transcript.cues.json` is `{start, end, text}`. -So the manifest is built by matching each quote back to its covering cues and -recording the real window. That is also what handles ellipsis-joined citations — -a report quote like `"…" … "…"` is often two separate cue spans presented as one. - -The manifest additionally pins the provenance (share link, corpus handle, match -counts, the narrowing queries) so a rebuild months later is reproducible and the -video's own claims about its coverage can be checked. - -## Where a clip actually gets cut - -Three stages, because a cue span is the wrong answer twice over. - -1. **Cue span** — the raw window covering the quote, from `transcript.cues.json`. -2. **Sentence widening** (`resolve-windows.mjs`) — walk outward to the nearest cue - ending in `.`, `?` or `!`. A cue boundary is where the *caption line wrapped*, - so cutting there drops the run-up that makes a quote intelligible. - - Two asymmetries matter. The **start** takes a lead-in only if a real sentence - opening sits within `--max-lead` (8 s); otherwise it takes none, because a - half-sentence run-up is the irrelevant context you were trying to avoid, not - context. The **end** is never clamped to a budget — stopping partway through a - sentence is the exact mid-thought ending this removes — so `--max-tail` (12 s) - only bounds how far it looks before giving up and using the cue end. - - **Some uploads have no punctuation at all.** Older ASR in this corpus emits - unpunctuated cue text for whole videos, and sentence detection then has nothing - to find: widening degrades to the raw cue span at both ends. That is not a - silent failure you can ignore — it is what produced a clip opening mid-thought - on "higher than they can afford and because", and what let another clip's tail - run through its neighbour. For those videos, pick the window by reading the - cues and set it by hand; the de-overlap and silence passes still apply. - - **`lockStart` / `lockEnd` pin an edge** to exactly what the manifest says, and - neither widening nor de-overlap will move it. Reach for it when the utterance's - real trailing pause does not line up with its last cue's end — the snap picks - the *nearest* silence, and in speech over game audio the nearest one is often a - gap between syllables rather than the pause at the end of the thought. -3. **Silence snapping** (`build-video.mjs`) — a sentence boundary in the - *transcript* still isn't a boundary in the *audio*, so clips clip words in - half. Fetch `fetchPad` seconds wider than needed, run `silencedetect` over the - result, and move each cut to the nearest silence within `snapWindow`. Starts - land on a silence's END (just before speech resumes), ends on a silence's START - (just after speech stops). No silence close enough → keep the exact point; a - tight cut beats a cut in the wrong place. The build logs `start✓ end✓` per clip - so you can see which snapped. - - **The silence threshold is relative, and it has to be.** These are game - streams: the gaps between words are full of game audio and music — quiet, but - nowhere near silent. A fixed absolute threshold sits below the noise floor and - finds nothing. On one measured clip: mean volume −21 dB, **0** silences at - −32 dB, **25** at −26 dB. So each clip is measured with `volumedetect` first - and the threshold set `silenceRelDb` (default 6) below its own mean. If a rebuild - suddenly reports mostly `start– end–`, this is the knob. - -The trim happens during the burn-in encode, so snapping costs nothing extra. - -4. **De-overlap** (`resolve-windows.mjs`) — widening is per-clip and blind to its - neighbours, so two clips cut from the *same* video can end up overlapping, and - the overlap plays as the same footage twice. Any earlier clip whose tail runs - into a later clip's start is trimmed back to that start. This is not a rare - edge case: it fired on the first report, where a 2024 upload has **no - punctuation at all** in the relevant stretch, so sentence detection found - nothing and the tail ran the full budget straight through the next clip. - -## Regenerating and changing a video - -Everything is cached by content, so iteration is cheap: - -- **Reorder, drop or add clips** — edit `timeline`, re-run. Cached clips are not - refetched, so a re-cut costs an encode, not a download. -- **Change a clip's window** — edit `start`/`end`, re-run. A raw file's window is - in its name, so the cache is content-addressed; and since a request is satisfied - by any cached file that **contains** it, a nudge inside the existing pad costs - nothing at all. Only a window that escapes every cached file downloads again. -- **Preview one entry** — `--only <id>` builds a single segment and stops. -- **Work offline** — `--skip-fetch` fails loudly instead of downloading, so you - can confirm you are working entirely from cache. -- **Fetch one clip, wide** — `--fetch-only <id> --pad 20` puts a generous window - in the cache without building anything. This is what the umtool clip bench runs, - and containing-window reuse is what makes that fetch double as the build's cache. -- **Force a refetch** — delete `out/clips-raw/`, or pass `--no-reuse` to require an - exact-window file. -- **Iterate on the rail** — `--rail-only` re-runs only the rail chain over a cached - `out/<slug>.prerail.mp4`; `--preview <start> <dur>` does the same over a window. - `--no-rail` builds the cut without one. See - [The claim rail](#the-claim-rail-renderrail). - -Containing-window reuse was retrofitted, and the waste it removes is measurable: -`ferret-rescue/out/clips-raw` holds **31 files for 10 clips** because every window -edit before this downloaded the same material again — one source is there four -times over overlapping windows. The tightest containing file wins, not the widest, -because silence detection decodes the whole file and a 40 s file costs more than -the 24 s one that would also have done. - -After changing any window, re-run `resolve-windows.mjs --write` before building. -It is a **fixed point**: running it on an already-resolved manifest reports -`0 window(s) changed` and rewrites nothing. That property is load-bearing and was -not free — the manifest stores times rounded to 2 dp, so a value read back can sit -a hair below the cue end it came from, which lands the end lookup on the previous -cue and runs the search on to the *next* sentence. Left alone, every re-run grew -the same clip. Hence `EPS` in the end lookup and the 0.05 s deadband on applying a -change. - -## Driven from umtool - -The three CLIs are the source of truth and stay usable on their own; umtool drives -them rather than reimplementing them, so the UI and the terminal can never disagree -about a window, a format string or the Rumble retry. Three additions exist for that: - -- **`--progress ndjson`** — one JSON object per line instead of prose: - `start`, `card`, `clip`, `fetch`, `snap`, `segment`, `entry-failed`, `concat`, - `chapters`, `note`, `done`, `error`. The event set is exactly what was already - being printed; making it a format switch is what stops a wording change from - breaking the driver. -- **`--continue-on-error`** — record a failed entry and carry on. A dead source at - clip 14 of 19 otherwise throws away thirteen fetches already paid for. The run - still **refuses to concatenate** and exits non-zero: a finished file that quietly - lost a citation is worse than no file. -- **`buildVideo({manifestPath, opts, out, only, fetchOnly})`** is exported, and - `widen()` from `resolve-windows.mjs` already was. umtool imports `widen()` so the - bench's "extend to sentence end" is the CLI's own function, and **spawns** the - build — a 40-minute chain of yt-dlp and ffmpeg inside a request handler has no - cancellation story. - -`check-availability.mjs` is the fourth CLI and belongs at the *front* of a build: - -``` -node scripts/report-to-video/check-availability.mjs <manifest.json> -``` - -It runs `yt-dlp --simulate` once per distinct `(channel, video)` — no bytes -downloaded — classifies each failure (`deleted`, `private`, `restricted`, -`members-only`, `geo-blocked`, `no-cues`, `maybe_missing`), and writes -`out/availability.json` with a timestamp. This is the one fact about a manifest -that goes stale in *both* directions: a source can die after the manifest is -written, and a source annotated "gone" can come back. `no-cues` is called out -separately because it is a different bug — usually the Rumble two-ids trap, where -the manifest names the MCP video id while the cue file lives under the URL slug. - -## Manifest shape - -`timeline` is an ordered list; entries are `card` or `clip`. - -```jsonc -{ "type": "card", "id": "ch3", "style": "chapter", "seconds": 4.0, - "kicker": "March – November 2025", "heading": "Then: the county", - "sub": "Six months for the first approval" } - -{ "type": "clip", "id": "c04", "video": "uyz1_FIqIEk", - "start": 32989.56, "end": 32994.19, // cue-accurate, from transcript.cues.json - "cite": 32989, // the second shown in the attribution line - "quote": "The pre-application screening was approved by the county, dude." } -``` - -Two per-clip fields exist for compilations that span sources or need a hand-cut -window: - -- **`channel`** — the archived channel this clip's cue file lives under, overriding - `provenance.channelSlug`. A compilation about one person routinely spans several - mirror channels (`HasanAbiVODs` / `…VODs3` / `…VODsbackup`), and cue files are - keyed by channel, so a single manifest-wide slug cannot find them all. -- **`lock`** — exempt this clip from `resolve-windows`. Sentence-widening exists to - stop clips ending mid-thought, but that is exactly wrong when the author has - deliberately cut a quote short: a single cue often holds a whole paragraph, so - trimming to "I hate this country so much sometimes" and dropping the rest of the - sentence is an editorial decision that widening would silently undo. `lock` also - handles the reverse case — a clip whose lead-in would drag in seconds of some - *other* audio (a news package playing before the speaker starts). - -Clips also carry `section` and (auto-set) `sectionEnter`. Card styles — `title`, -`timeline`, `status`, `bullets`, `sources` — still work, but the ferret-rescue cut -uses none of them. `render` holds resolution, fps, fonts, palette and the knobs -(`fetchPad`, `snapWindow`, `silenceRelDb`, `transition`, `slideSeconds`, -`headerHeight`, `footerHeight`, and the optional `rail`); `provenance` holds the -sweep's scope and counts. Two further entry `type`s, `scroll` and `chart`, close a -cut off a top-level `ledger[]` — see [The claim rail](#the-claim-rail-renderrail). - -## Chrome, not cards - -**The ferret-rescue cut has no cards at all** — no title, no chapter breaks, no -closing slate. It is a cited timeline and nothing else: the clips run in -chronological order and the source material carries the argument. Cards remain -supported for reports that want them, but the default posture is that anything -drawn is an interruption which has to earn its place. - -Nothing is drawn *over* the picture either. The video is **letterboxed between** -thin chrome rather than overlaid by it: - -- **Header (`headerHeight`, 56 px).** The citation only — cleaned title · upload - date · timestamp. No quote: the clip is already saying it, and burning in a - transcription of speech you can hear is noise. -- **Footer (`footerHeight`, 100 px).** The timeline: one node per milestone, each - with a label and a month/year stamp beneath it. Drawn once by - `renderFooterAssets()`. -- **The marker slides.** On the first clip of each section (`sectionEnter`, set - automatically), the fill bar and the amber marker animate from the previous node - to the current one over `slideSeconds`. Everywhere else they hold position. The - motion is ffmpeg expressions on `crop`/`overlay`, so it costs nothing beyond the - encode that was happening anyway. - - > **This used to be half true.** Until 2026-08-19 only the marker moved. The - > fill bar was a `drawbox` whose width was `if(lt(t,0.9),…)` — but **`drawbox` - > has no time variable**: its `t` is the box *thickness*, and with `t=fill` - > that is effectively `INT_MAX`, so the comparison was always false and the - > width expression collapsed to its end value. (`drawbox=w='t*10':h=8:t=2` and - > `drawbox=w=20:h=8:t=2` produce an identical YAVG.) - > - > It was worse than "always full", because **`drawbox` reads `w=0` as *the - > input width***. Section 0's fill is 0 px, so every clip in the first section - > drew the bar across the **whole frame** — the progress track read 100 % - > complete on the opening clip of every cut that has a footer. Measured on - > `quartering-flagging-takedowns/n02`: 1708 accent px on the track row before, - > 0 after (the correct value), with the amber marker parked on node 0. - > - > It is now a `_bar.png` strip of `2·trackLen × 3` — accent on the left half, - > transparent on the right — translated under a fixed-width `crop`, which - > *does* evaluate `x` per frame. Verified on `ferret-rescue/c07`: 608 → 698 → - > 798 → 912 px across t = 0.0 … 0.9 s, then parked. - -Clips carry `section` (index into `timelineNodes`); nodes supply `label` and -`date`. Ordering clips chronologically is the author's job — the manifest plays in -the order it is written. - -A consequence worth knowing: 16:9 source into the reduced height leaves narrow -pillarbox bars. That is the price of never covering the picture, and it is why the -chrome is kept as thin as it is. - -**Both bands are optional, and turning them off is a real setting, not a hack.** -`footerHeight: 0` (or an empty `timelineNodes`) drops the timeline; `headerHeight: 0` -drops the citation line. With both at zero the clips fill the whole frame and -nothing is drawn at all — the filtergraph loses the overlays rather than compositing -invisible ones, and `renderFooterAssets` returns early instead of drawing PNGs -nobody uses. Two reasons this comes up: - -- A timeline footer only means something if the clips *are* a progression through - time. A cut ordered by argument rather than by date should not draw one. -- The header prints the **upload date of the archived copy**, which for a VOD-mirror - channel is often years after the stream (a Nov 2019 stream re-uploaded in Apr 2023 - reads `… November 6, 2019 … · 2023-04-06`). The title usually carries the true - date, so nothing is false, but on a cut spanning many re-uploads it reads badly. - -The `hasan-hate-america` cut runs with both off. Restoring them is a two-value edit. - -Note: commas inside an ffmpeg filter expression have to survive filtergraph -parsing — wrapping the expression in single quotes is what protects them. - -- **Segments crossfade** (`transition`, default 0.5 s). This forces a full - re-encode of the timeline via `xfade`/`acrossfade` — the concat demuxer can only - stream-copy hard cuts. Pass `--no-xfade` for a fast hard-cut build while - iterating; the last pass can add the transitions back. - -## Two cuts from one manifest (`--variant`) - -A sweep finds more claims than a cut can show footage for. There are two -defensible answers to that and they make different videos, so the manifest -describes both and one filter picks between them. - -| variant | output | ledger | what the viewer sees | -|---|---|---|---| -| `sourced` (default) | `out/<slug>.mp4` | claims with a clip | every row on screen has footage behind it | -| `full` | `out/<slug>-full.mp4` | every claim | the unclipped ones are stacked onto `ledger` cards | - -`selectVariant(manifest, variant)` runs **immediately after the manifest is read** -and is the whole mechanism. Three rules, in this order: - -1. a timeline entry tagged `"variant": "full"` survives only in that variant; -2. a claim survives only if the entry its `entryId` names survived — which is what - makes `sourced` a sourced-only ledger, because in `full` every claim is pinned, - to a clip or to a `ledger` card; -3. `card.variants[<name>]` field overrides are merged in and the key dropped. - -Nothing downstream learns about variants. `ledgerTotals`, the rail, the chart -band, `scheduleClaims`, the chapters and the closing cards already take the -ledger and the timeline as inputs. - -**Rule 3 exists because copy can be false in one cut.** A title card saying "48 -dated claims" is a lie in a cut that shows nineteen, and the closing sources card -says two claims stayed ambiguous — both of which happen to be unclipped, so in -`sourced` there are none. - -### Output layout - -``` -out/ - clips-raw/ SHARED — the only expensive thing in a build - availability.json SHARED — a fact about the manifest, not about a cut - <slug>.mp4 sourced - <slug>-full.mp4 full - sourced/{cards,segments,qr,chrome,schedule.json} - full/{cards,segments,qr,chrome,schedule.json} -``` - -`clips-raw` is shared deliberately: `sourced`'s clips are a subset of `full`'s, so -no clip is ever fetched twice. `sourced` writes `out/<slug>.mp4` because that is -the path umtool's build probe already looks for. - -`compose-chrome.mjs` and `verify-build.mjs` both take `--variant` for the same -reason: the band plots the ledger the cut carries, and verifying the whole -manifest against one variant's file would report a missing chapter for every -entry the other cut has. - -### The `ledger` entry type - -`full`'s answer to the claims no clip covers. One card per **run of consecutive -unclipped claims**, so 29 claims cost 13 cards and about 66 seconds. - -```jsonc -{ "type": "ledger", "id": "L10", "variant": "full", "seconds": 4.8, - "kicker": "November 2024", - "heading": "Ten and eight, named separately, in one breath", - "sub": "…", // optional - "claims": ["c12", "m14"] } // ledger ids, in ledger order -``` - -Each row draws `date · scope pill · why it is not footage · the quote · his -figure`, plus a right-hand **arithmetic column**: `media / coffee / publica` as -they stand, the implied total, the layer this claim just moved lit, and its delta. -The arithmetic is **read from `ledgerTotals`, never recomputed** — one walk, or -the card and the band disagree about the same sum. - -Rows **reveal in sequence** behind an opaque `pal.bg` rectangle walking down the -card: the rail curtain's device, exact because the card ground is flat. -`seconds` is derived (`2.2 + 1.3·rows`) rather than authored, because the rail -pins to the same clock — see below. - -**"Why it is not footage" comes from a probe, not from a hand-written kicker.** -`check-availability.mjs` now probes every **ledger** source as well as every clip -source, so a row says `source deleted`, `source unreachable` or `not clipped` -because `yt-dlp --simulate` said so on a recorded date. Several unclipped claims -are cut from videos this cut clips elsewhere, i.e. demonstrably live; saying "the -upload is gone" about one of those is the kind of error that discredits the whole -compilation. - -**Pins gain a within-segment offset.** A claim on a `ledger` card is pinned to its -own row's reveal (`ledgerRevealAt(r)`), not to the segment's mid-dissolve. -Otherwise four rail rows land on one frame, and the pin-order guard's strict -monotonicity breaks for no reason. In `full` this pins the rail almost exactly: -every claim has a segment, so `scheduleClaims` interpolates almost nothing. - -**`status` is retired.** It existed to quote a claim whose source had gone; a -`ledger` card does the same thing better, alongside the arithmetic the claim moves -and with the reason coming from the probe. - -## The claim rail (`render.rail`) - -A cut whose whole point is *which company a number was about* has a problem: the -dates and the figures are **spoken**, and shown only in the header's citation -line. A viewer can hear "nearly ten" three times without ever seeing that the -three refer to three different payrolls. - -`render.rail` adds a persistent **vertical ledger down the right edge**. It -appends one row per claim as the video runs, keeps a live per-company tally -beside it, and lists **every claim the sweep found** — not just the ones with a -clip behind them. Claims with no clip are dimmed (muted ink, hollow dot) and pass -with no audio; they are what stops the rail from implying the cut is the corpus. - -It is **entirely opt-in**. With no `render.rail` key the filtergraph is the one -that was there before, and output is byte-for-byte unchanged. - -```jsonc -"render": { - "rail": { - "width": 420, // picture shrinks to width - 420 - "rowHeight": 46, // one claim row; window height is a whole multiple - "tallyRowHeight": 44, - "tallyTop": 96, // y of the tally block inside the rail column - "pad": 22, - "slide": 0.55, // seconds per row-change ease - "rule": "#2A322F", - "tracks": [ // one per company; ORDER is the rail/legend order - { "key": "media", "label": "The Quartering · media", "color": "#22AB83" } - ] - } -} -``` - -and a top-level `ledger[]`, in **playback order** — each track's claims contiguous -and date-sorted within the track: - -```jsonc -{ "id": "m02", "date": "2022-09-17", "company": "media", - "value": 4, // null for a qualitative claim; the tally ignores those - "display": "4", // the badge - "label": "counts them out: one, two, three, four", - "quote": "…", "src": "…", - "hedged": false, // a hedge word, not a figure -> hollow dot on the chart - "plotted": true, // appears in the step chart - "entryId": "a01" } // pins the row to that timeline entry's segment -``` - -Pinned entries carry a `claim` back-reference so the link reads both ways. - -### How it is put together - -Everything that moves is **one tall strip walked by a fixed-size `crop`**, not a -per-state still, because swapping stills can only cut and a crop can ease. Five -strips, all bounded `-loop 1 -framerate <fps> -t <total+2>`: - -| strip | size | what it is | -|---|---|---| -| `_rail_chrome.png` | `RW × RHGT` | opaque panel, title, rules, **and the tally swatches and labels**. Runs the whole video — no `enable=` gates | -| `_rail_log.png` | `RW × N·rowHeight` | every claim, stacked, no padding | -| `_rail_curtain.png` | `RW × LOGH` | opaque `pal.bg` | -| `_rail_hl.png` | `RW × rowHeight` | the amber current-row marker | -| `_rail_tally.png` | `ΣlaneW × rows·cellH` | one COLUMN per lane — four rolling cells and the roster line | -| `_rail_qr.png` | `TILEW × segments·TILEH` | one provenance tile per segment | - -### The tally rolls one number at a time - -It used to be a column of four-row slabs walked by one `crop`: when the coffee -company's number changed, all four rows moved, and "The Quartering" slid up the -screen for a reason that had nothing to do with it. Text that has not changed -must not move. - -So the **swatch and the company label went into the static chrome**, and each -track got its own **rolling cell** — number, delta triangle, `as of <date>` and a -population chip, right-aligned in a ~170 px column. All the lanes live side by -side in **one** PNG, so it is still one input: five `crop`s at different `x`, five -overlays. - -**The direction of a roll is decided by the strip's LAYOUT, not by the ramp.** - -| | rows | the crop | what you see | -|---|---|---|---| -| rise | `[old, new]` | walks **down** | content moves **up** | -| fall | `[new, old]` | walks **up** | content moves **down** | - -Between transitions a one-frame `gte()` step repositions to the next pair's -starting row. That step is invisible **because both endpoint rows hold identical -content** — which is why every pair repeats the value it starts from instead of -sharing a row with its neighbour, and why the delta chip rides on both rows and -therefore stays on screen until the next change. - -``` -y_j(t) = r_j0 + Σ_k [ (a_k − b_{k−1})·gte(t,t_k) + (b_k − a_k)·ease(t_k) ] -``` - -Every term is cumulative and saturating — the rail's hard rule. - -A repeated identical figure still rolls, upward: he said it again on a new date, -and the `as of` line underneath is what changed. - -### The roster line - -A fifth lane under the tally, `2 editors · 1 designer`, in `pal.muted`. It moves -**only when the rendered line changes**, which in this corpus means it stands -still through October and December 2023 while the total above it goes from three -to four. That is the finding, drawn rather than asserted. - -It comes from an optional `roles` field on a ledger claim, and `rosterAt()` / -`rosterLine()` in `ledger-totals.mjs` are the one implementation, because a -chapter card states the same thing in words. - -```jsonc -"roles": [{ "role": "video editor", "count": 2, "verbatim": "two video editors" }] -``` - -`verbatim` is his words; `role` and `count` are our reading. `roles` is **not** one -of the six adjudication fields — most claims are a number and nothing else, and -gating the inbox on a field a handful of entries can carry would leave it -permanently red. - -### The QR moved into the rail's foot - -It used to float over the bottom-right of the **picture**, which is the one part -of the frame this cut promises never to draw on. It is now a bordered tile parked -at the foot of the rail column — `pal.accent` rule, `SCAN → JERALYZER` above the -code, what it opens below it — and one more strip: one tile per segment, -crop-walked with **instantaneous `gte()` steps** at segment mid-dissolves. A code -that eased into place would spend the ease unscannable. - -The tile overlays **after the curtain**, which is what stops the parked curtain -painting over it. - -> **A card cannot carry the report's share link.** That link carries all 23 -> channel filters and is ~1.4 k characters: a version-40 symbol, 177 modules in a -> 132 px tile, about 0.7 px per module. Cards get `provenance.qrLink` — the same -> query without the channel list, ~200 chars, 63 modules, verified scannable at -> this size — and fall back to `provenance.siteOrigin`. Manifests with **no** -> rail keep the old per-clip overlay in `buildClipSegment`, byte for byte. - -> **`railGeometry` used to derive its height from `render.footerHeight`.** With -> the chart band on, the band reserves 200 px and `footerHeight` says 100, so the -> rail column ran a hundred pixels — about two rows of its log window — past the -> line every other renderer letterboxes to. It reads `reservedFooterHeight(render)` -> now, the same fix the closing cards needed for the same reason. - -**The curtain is why one strip is enough.** With the window parked at the top -while the list is still filling, rows `i+1 … K-1` would show claims the video has -not made yet. The curtain is an opaque rectangle riding just below the last -revealed row; once the list is full it parks exactly one window-height down, -which is the bottom of the rail column — permanently outside the window. It is -`pal.bg` precisely so that parking there is invisible against the footer band. -Curtain and log **must share the same eased `P`**, or the curtain lags the rows -mid-slide and unrevealed claims flash into view. - -**Ramps are cumulative and saturating, never gated.** A piecewise sum of -`gte(t,sᵢ)·lt(t,sᵢ₊₁)·…` terms flashes to `y=0` for one frame at any boundary -gap, because every gate evaluates false at once and the sum collapses. Terms that -rise to their delta and stay cannot do that. - -**Scheduling.** Playback is ONE chronology across every company, and the ledger is -sorted the same way, so a claim's position in the rail *is* its position in time. -A claim with a clip behind it is pinned to that clip's segment; a claim on a -`ledger` card is pinned to its own row's reveal; the rest are spread evenly -between their neighbouring pins. A pin that runs backwards is refused — the -ledger and the timeline disagreeing about the order of events is a manifest bug, -and the whole cut rests on the two agreeing. - -Segment-level pins land at **`starts[i] + D/2`** — mid-dissolve, where the picture -is already crossfading and a ±3-frame error is invisible. - -The chain attaches **after the last `xfade` node**, inside the concat pass. `t` -there is absolute and continuous from 0, and nothing downstream of the last xfade -is dissolved — so it already has post-pass semantics without a second encode, -which would re-quantize crf-20 output at exactly the content that hurts most -(antialiased text on flat colour). - -### `--rail-only` and `--preview` - -`--rail-only` re-runs just the rail over a cached `out/<slug>.prerail.mp4`, -building that file from the existing segments the first time. Seconds instead of -the full concat. The hard-cut base is a separate file -(`<slug>.prerail-hardcut.mp4`), because the two timelines are different lengths -and a cached base from the wrong mode is a stale-cache trap the length assertion -would otherwise have to explain. - -**It is mandatory for `--no-xfade`**, not an optimisation: `concatHardCut` is -`-c copy` and a stream-copy mux cannot host a filtergraph at all. That path -concats to `.prerail.mp4`, asserts its length against `segmentOffsets().total`, -and then applies the rail. - -`--preview <start> <dur>` renders a window. `-ss` restarts `t` near zero, which -would put every absolute-time ramp in the wrong place — so the preview path -inserts `setpts=PTS+<start>/TB` before the rail chain and rebases afterwards. -Getting that wrong makes a working rail look broken. - -## The end sequence: `scroll` and `chart` - -Two timeline entry `type`s that exist to close a cut, both driven off the same -`ledger[]`: - -- **`scroll`** — the whole ledger as one tall PNG, walked by an animated `crop`. - **One chronological line, a column per company**: it used to group by company, - which re-told the cut's own order backwards and hid the only thing worth seeing - there — that the four payrolls were being described in the same weeks. Company - is read from COLUMN POSITION, so colour is the secondary encoding. - `hold` (default 2 s) buys a still moment at both ends; - `clip()` in the expression provides it for free, and crop's own clamping - degrades an off-by-a-few content height into a static last frame, not an error. -- **`chart`** — the four-series step chart over the `plotted` claims, authored as - SVG and rasterized with `rsvg-convert` (deterministic about output size in a - way ImageMagick's RSVG delegate is not). Fonts inside the SVG resolve through - **fontconfig, not `render.fontRegular`** — use the family name the Pango cards - use. - -The chart's wipe **cannot** be `crop=w='<ramp>'`: crop's `w` is config-time and -`t` is undefined there. It is a curtain instead — an opaque `pal.bg` rectangle -slid rightwards off the plot, which is exact because the card ground is flat. - -**Colour is not the only encoding on that chart, and that is a requirement.** No -four-colour categorical palette clears the data-viz all-pairs CVD gate (three -slots is the documented ceiling), so every series also carries a distinct dash -pattern and a direct end-of-line label with a leader elbow. The four hues are the -published artifact's, re-validated against this video's darker ground (`#0F1312`) -on the *adjacent* pairlist — the pairlist for line charts — where all five checks -pass (worst adjacent CVD ΔE 10.2 against a ≥8 target; normal-vision ΔE 17.4 -against a ≥15 floor). - -**`hideRail: true`** on a closing card slides the whole rail column off to the -right over that card's dissolve — one offset expression shared by every rail -overlay, so the column moves as one object — and lets the card render at the full -`width` instead of `contentWidth`. Not an `enable=` pop: a column that vanishes -between two frames reads as a dropped frame. The slide is cumulative and -saturating like every other ramp, so the rail does not come back; every card -after the first `hideRail` one should carry the flag too, or it lays out inside a -content width whose rail is no longer there. - -Both kinds go through the same `fps=,setsar=1` and the same `encodeArgs` as every -other segment. They have to: `xfade` rejects a mismatched link with *"First input -link parameters do not match"*, and that surfaces at concat time, after every -fetch has been paid for. - -## The ledger is adjudicated, and both totals depend on it - -`ledger[].company` used to be an **undocumented interpretation**, and four -different hazards were riding on it: - -| hazard | example | why it mattered | -|---|---|---| -| **scope ambiguity** | *"I have 10 employees, my coffee company employees… my editors"* | all-companies or coffee-only, depending on where the comma falls | -| **derived, not stated** | *"10 at coffee brand coffee, I've got eight staff for the live stream"* recorded as **18** | he never says 18 | -| **population drift** | 5 *"full-time salaried"*, 10 *"employees"*, 10 *"all basically contractors"* | different denominators, one series | -| **synthetic values** | 10.5 for *"about 10 people, 11 people"* | a midpoint we invented and attributed to him | - -So every ledger entry now carries six adjudicated fields — `scope`, -`scopeBasis`, `scopeConfidence`, `population`, `valueKind`, `flags` — settled -against **±90 s of surrounding context, never the quote alone**. A first-person -quote is routinely the host reading someone else's words or being sarcastic, and -neither is visible inside the quote. One claim in this corpus is a guest's -payroll rather than his, and it reads identically until you listen either side. - -`scopeConfidence: "unresolved"` is a legitimate outcome and **feeds neither -total**. - -**The rule that follows:** the *stated* series may contain only **a figure he -utters as a single number for a named scope**. Sums and midpoints are ours, and -live in the *implied* series, which says so on screen. - -`ledger-totals.mjs` is the one implementation of that arithmetic — the umtool -inbox, the chart band and the closing card all import it, so none of them can -disagree. It **refuses to run on an unadjudicated ledger**, because both totals -lie if you act on one. **Six** named predicates compute incoherence rather than -asserting it (`contradicts_component`, `same_day_conflict`, `self_negating`, -`population_mismatch`, `not_his_number`, `status_flip`). Deliberately **not** a -predicate: a large rise or fall between claims. Fluctuation is the subject, not a -defect. - -`status_flip` is the sixth: the same people described as staff and then as -contractors, or the reverse. Three details in it are load-bearing. - -- **`employees` and `people` are in NEITHER camp.** They are what he says when he - is not making a claim about status at all, and reading them as one side or the - other manufactures a reversal out of a change of vocabulary. -- **An `all` claim is comparable with any company; two companies are not - comparable with each other.** Without that asymmetry the corpus's clearest - reversal is invisible: December 2024's ten are the *channel's*, May 2025's ten - or eleven are *everything's*. -- **Only against the most recent comparable claim that carries a camp.** Fire on - every earlier pair and one 2022 "all 1099 and not full-time" flags each of the - next seven claims in turn — seven findings where there is one. Bounded this - way it fires at the TRANSITIONS, which is what a flip-flop is. - -It is not gated on the claim having a figure. *"That's why all my workers are -contract workers"* names no number and is the single clearest status claim here. - -Work the adjudication in umtool, at `/browse/<project>/claim/<id>`. Sign-off is -"the inbox is empty": `claim-unadjudicated` is **blocking**. - -## `render.chromeEngine: "hyperframes"` — the chart band - -Opt-in, and absent it the ffmpeg chrome path is byte-for-byte unchanged. - -The chart band **replaces** the footer node track, which only moved at section -handovers — precisely the fault it exists to fix. It takes the footer's ground -and 100 px more, and the picture loses that height (1500×924 → 1500×824). - -`compose-chrome.mjs` emits a HyperFrames project per region and renders it to a -**lossless RGBA PNG sequence**; `build-video.mjs` overlays the frames. Three -things about that are load-bearing: - -- **Footage never enters Chrome.** HyperFrames pre-extracts source video to JPEG - q95, which is unacceptable when the picture *is* the cited evidence. Only - chrome is composed there — about 37 % of full-frame pixels rather than 100 %. -- **The PNG regions overlay BEFORE the rail chain, not after.** The rail chain - ends in `format=yuv420p`, and overlaying an alpha sequence onto yuv420p is the - same alpha-subsampling trap the rail already documents, one layer later. -- **The playhead is driven by `out/schedule.json`**, which the build writes. - Recomputing claim times here would be a second implementation of - `segmentOffsets()` and would drift the first time `transition` changed. Same - rule as `widen()`: imported, never reimplemented. - -One clip-path sweeps the whole plot rather than a `stroke-dashoffset` per series. -The obvious build animates each path's dash offset, and it looks right for the -strokes and wrong for everything else: the gap band between the two totals is a -filled polygon with no stroke to offset, so it appears whole the moment it fades -in and the chart is showing an answer the playhead has not reached. - -**The flag colour is not the palette's amber.** `#E8A33F` sits at ΔE 12.4 from -the coffee series' `#D2732F` at *normal* vision — below the 15 floor — so a flag -badge beside a coffee mark was hard to tell from the coffee mark. `#E0E24A` -replaces it and adds **no new worst pair**: the worst CVD pair -(`#C55F9C↔#22AB83`, ΔE 5.4 deutan) and the worst normal-vision pair -(`#C55F9C↔#D2732F`, ΔE 16.3) are identical with and without it. The implied -total's `#EDF0EC` fails the categorical lightness and chroma checks *by design* — -it is an aggregate, not a categorical peer, so it is encoded by weight and -consumes no palette slot. - -## Things that cost time to find out - -**yt-dlp picks VP9 at `height<=720`, and that is a trap.** `--download-sections` -combined with `--force-keyframes-at-cuts` re-encodes, so a VP9 pick means -libvpx-vp9 — 27 seconds to cut a 5-second clip. It also writes a `.webm` and -appends that extension to whatever `-o` you gave, so the file never lands where -you asked and the run fails looking for it. Pin H.264/AAC in mp4 and the same cut -takes ~12 seconds. `build-video.mjs` does this and keeps a rename fallback for -the case where a fallback format still forces another container. - -**`--force-keyframes-at-cuts` is not optional here.** Without it the cut snaps to -the nearest preceding keyframe and can start seconds early. That is fine for a -human scrubbing a VOD; it is not fine when the clip *is* the citation. -`PlayerProvider.tsx:478` builds the copyable clip command without this flag — -correct for its purpose, wrong for ours. - -**`--ignore-config` is mandatory.** The operator's own yt-dlp config redirects -output to `~/Podcasts` and attaches thumbnail/metadata post-processors. Every -managed yt-dlp call in this repo passes `--ignore-config` for the same reason. - -**yt-dlp exit 101 is success**, not failure — it means a clean early stop. The -repo encodes this at `common/ytdlp/downloadOneManaged.ts:355`; a new caller has to -replicate it. - -**Local media will not help you.** Only 5 of 1,804 PirateSoftware video dirs hold -any media at all, and none are ones a report is likely to cite. Clips are a -network fetch. `metadata.info.json` and `transcript.cues.json` *are* present for -every video, so titles, dates, durations, webpage URLs and cue timings all come -from disk with no probe. - -**Upstream availability is load-bearing.** A clip can only be fetched while the -source is still up. Check `platform state` via `get_video_metadata` before -committing to a clip — a `deleted`/`maybe_missing` source needs a quote card -instead of footage. The archive outlives its sources, so a video built from an -old report will be *less* complete than the report unless this is handled -deliberately. - -**ImageMagick's `-size` leaks into the Pango group.** `-size 1920x1080 xc:BG` -followed by `( ... pango:@file )` renders the text into a full-frame box, which -pins it to the top and wraps at the frame edge instead of the text column. Reset -`-size` inside the parens. - -**ffmpeg `drawtext` does not wrap and hates punctuation.** Both are solved the -same way: wrap to a column count in JS, write to a file, and use -`textfile=`. Nothing then needs escaping. Stream titles also need emoji and -`!command` suffixes stripped or they render as tofu in the attribution line. - -**Segments are encoded to identical parameters on purpose** so the final -concatenation is a stream copy via the concat demuxer. Mismatched streams are the -usual reason a naive concat produces a broken or audio-desynced file. - -## Not done yet - -- **Narration is silent by design.** Cards carry the connective text; the only - audio is the clips'. A TTS layer would attach per card (`seconds` already gives - it a duration to fill) — deliberately deferred rather than designed out. -- **Snapping is silence-based, not word-based.** It finds gaps in the audio, which - is usually the same thing as a word boundary but is not guaranteed to be — - a speaker who does not pause gets the unsnapped cut. Forced alignment against - the transcript would be exact; `silencedetect` is a tenth of the work and - handles the cases that were actually audible. -- **Only the chart band is a HyperFrames region.** `chromeRegions()` returns one - entry. The rail and the header are still drawn by the ffmpeg chain, and porting - them is the rest of the job — the rail needs the ghost/hop convention (a row - with no clip fades in dimmed under a dashed rule and the amber highlight *hops - over* it to the next cited row) which the strip builders cannot express. Until - then `railFilterChain` and `renderFooterAssets` stay; they must not be retired - on the strength of the band alone. -- **`timelineNodes` / `section` / `sectionEnter` are still live.** They lose their - only consumer when the ffmpeg footer goes, not when the band arrives — so they - retire with `renderFooterAssets`, in that same commit. -- **Manifests are written by hand** from verified cue data. Deriving a first-draft - manifest automatically from a report's citations is the obvious next step; the - report parse is straightforward (`> "quote"` followed by - `— [title @ h:mm:ss](…?v=slug%2Fid&t=sec)`), the cue-matching is the real work. diff --git a/scripts/report-to-video/build-video.mjs b/scripts/report-to-video/build-video.mjs @@ -1,1838 +0,0 @@ -#!/usr/bin/env node -// build-video.mjs — render a cited sweep report into a narrated-by-text video. -// -// Takes a video manifest (see README.md next to this file) and produces one mp4: -// text cards state the findings, clips let the source say it in their own voice, -// and every clip carries a burned-in quote plus its attribution. -// -// Pipeline, per manifest entry: -// card -> still PNG (render-cards.mjs) -> N seconds of video + silent audio -// clip -> yt-dlp --download-sections (WIDE) -> silence-snap -> trim + burn -// then the segments are crossfaded together into the finished file. -// -// Three things worth knowing about how clips are cut: -// -// 1. Windows come from the manifest as absolute [start, end] seconds, derived -// from transcript.cues.json (which carries an END per cue). A sweep report -// only ever records a single start second, so windows cannot be recovered -// from the report alone. -// 2. Those windows are widened to sentence boundaries by resolve-windows.mjs, -// so a clip carries the run-up that makes the quote make sense. -// 3. A cue boundary is still not a *speech* boundary — cutting there clips -// words in half. So we fetch wider than needed and snap the real cut to a -// silence found in the audio. That is what makes clips start and end -// between words rather than through them. -// -// Fetched clips are cached by (video, start, end); re-running is cheap and only -// changed entries re-download. Delete out/clips-raw to force a refetch. -// -// In the app: not used. On the CLI: -// node scripts/report-to-video/build-video.mjs <manifest.json> [options] -// -// Options: -// --out <dir> Output root (default: manifest dir + /out) -// --variant <name> Which cut to build (sourced | full; default sourced) -// --skip-fetch Fail instead of downloading anything not already cached -// --only <id> Build a single entry's segment and stop (for iterating) -// --no-xfade Hard cuts instead of crossfades (much faster; concat copy) -// --progress ndjson One JSON event per line instead of prose (for umtool) -// --continue-on-error Record a failed entry and carry on, instead of aborting -// --fetch-only <id> Fetch one clip's window into clips-raw and stop -// --pad <s> Override render.fetchPad (the clip bench fetches wide) -// --site-origin <url> Archive to read cue windows from when there is no local -// corpus (defaults to the manifest's provenance.siteOrigin) -// --resolve-site-ids On a published-id miss, find the record by scanning the -// channel's shards. Slow; see cues.mjs. -// --cue-source <which> auto (default) | local | http. The two can disagree -// once a corpus moves past its last publish — see cues.mjs. -// --no-rail Skip the claim rail even when the manifest configures one -// --rail-only Re-run just the rail over out/<slug>.prerail.mp4 -// --preview <s> <d> Rail-only, over a <d>-second window starting at <s> -// -// Requires: yt-dlp, ffmpeg/ffprobe, ImageMagick with Pango. - -import { execFile } from "node:child_process"; -import { promisify } from "node:util"; -import { mkdir, writeFile, readFile, access, readdir, rename } from "node:fs/promises"; -import path from "node:path"; - -import { - renderCard, renderFooterAssets, renderRailAssets, renderScrollCard, renderChartCard, - renderLedgerCard, ledgerRevealAt, ledgerSeconds, - cardWidth, contentWidth, reservedFooterHeight, -} from "./render-cards.mjs"; -import { createCueSource, siteOriginFromManifest } from "./cues.mjs"; - -const execFileP = promisify(execFile); - -const YTDLP = process.env.YTDLP_BIN ?? "yt-dlp"; -const FFMPEG = process.env.FFMPEG_BIN ?? "ffmpeg"; -const FFPROBE = process.env.FFPROBE_BIN ?? "ffprobe"; -const QRENCODE = process.env.QRENCODE_BIN ?? "qrencode"; - -// Cue windows and per-video metadata come from a local corpus when there is one -// and from the published archive otherwise, so this runs in a clone with no -// `transcripts/` directory. Built once main() has the manifest (it carries the -// archive origin); see cues.mjs. -let CUES = null; - -const exists = (p) => access(p).then(() => true, () => false); - -// ---- variants ------------------------------------------------------------ -// ONE manifest, two cuts, one filter, applied once. -// -// The question the two variants answer differently is what to do with a claim -// the sweep found but no clip covers. `sourced` refuses to put it on screen at -// all -- every row the viewer sees has footage behind it. `full` gives each one -// a slot on a stacked ledger card, so nothing is dropped and the arithmetic of -// each layer is shown rather than asserted. -// -// Both end on the same three numbers. That is the point of shipping both: if -// the totals moved when the unsourced rows came off, the thesis would rest on -// rows nobody can check. -// -// The filter runs IMMEDIATELY after the manifest is read, and nothing -// downstream learns about variants. `ledgerTotals`, the rail, the chart band, -// `scheduleClaims`, the chapters and the scroll already take the ledger and the -// timeline as inputs, so selecting is the whole of the mechanism. -export const VARIANTS = ["sourced", "full"]; -/** The cut a caller means when it does not say. `out/<slug>.mp4`. */ -export const DEFAULT_VARIANT = "sourced"; - -/** - * The manifest as one variant sees it. - * - * Three things happen, in this order: - * - * 1. A timeline entry tagged `variant` survives only in that variant. (The - * stacked ledger cards are `variant: "full"`.) - * 2. A claim survives only if the entry it is pinned to survived. That single - * rule is what makes `sourced` a sourced-only ledger: in `full` every claim - * is pinned -- to a clip or to a ledger card -- so nothing is dropped. - * 3. `card.variants[<name>]` field overrides are merged in. The title and - * sources cards have to state their own scope honestly, and "50 dated - * claims" is simply false in `sourced`. - */ -export function selectVariant(manifest, variant = DEFAULT_VARIANT) { - if (!VARIANTS.includes(variant)) { - throw new Error(`unknown variant \`${variant}\` — one of ${VARIANTS.join(", ")}`); - } - const timeline = (manifest.timeline ?? []) - .filter((e) => !e.variant || e.variant === variant) - .map((e) => { - if (!e.variants) return e; - const { variants, ...rest } = e; - return { ...rest, ...(variants[variant] ?? {}) }; - }); - const kept = new Set(timeline.map((e) => e.id)); - const ledger = (manifest.ledger ?? []).filter((c) => c.entryId && kept.has(c.entryId)); - return { ...manifest, variant, timeline, ledger }; -} - -/** - * Where a variant's own working files live. - * - * `clips-raw` stays at the ROOT and is shared: it holds the only expensive - * thing in the build (network fetches), and `sourced`'s clips are a subset of - * `full`'s, so a shared cache means no clip is ever fetched twice. Everything - * else is per-variant, because every one of them differs between the two cuts. - * - * `sourced` writes `out/<slug>.mp4` -- the path umtool's build probe already - * looks for -- and `full` writes `out/<slug>-full.mp4` beside it. - */ -export function variantPaths(outRoot, slug, variant) { - return { - root: outRoot, - dir: path.join(outRoot, variant), - rawDir: path.join(outRoot, "clips-raw"), - final: path.join(outRoot, variant === "sourced" ? `${slug}.mp4` : `${slug}-${variant}.mp4`), - }; -} - -// ---- progress protocol --------------------------------------------------- -// This has two audiences: a human watching a terminal, and umtool's build driver -// reading the pipe. Rather than have the driver scrape prose (which would make -// every wording change a breaking change), `--progress ndjson` switches every -// line to one JSON object. The event set is exactly what was already being -// printed -- this is a formatting switch, not new instrumentation. -// -// Events: start, card, clip, fetch, snap, segment, entry-failed, concat, -// chapters, note, done. -const HUMAN = { - start: (e) => `${e.title} — ${e.entries} entr(ies)`, - card: (e) => `card ${e.id}`, - clip: (e) => `clip ${e.id} (${e.video}) §${e.section}${e.sectionEnter ? " ⟶" : ""}`, - fetch: (e) => - e.reuse - ? ` fetch ${e.id}: ${e.reuse} already covers ${hms(e.from)}–${hms(e.to)} — no download` - : e.cached - ? null - : ` fetch ${e.id}: ${e.video} ${hms(e.from)}–${hms(e.to)}`, - snap: (e) => - ` snap ${e.id}: ${e.start ? "start✓" : "start–"} ${e.end ? "end✓" : "end–"} ` + - `(${Number(e.seconds).toFixed(1)}s)`, - segment: () => null, - "entry-failed": (e) => ` ** ${e.id} failed: ${e.message}`, - concat: (e) => `${e.mode === "xfade" ? "crossfading" : "hard-cutting"} ${e.n} segments…`, - chapters: (e) => `chapters: ${e.n} marker(s) -> ${e.file}`, - note: (e) => e.message, - done: (e) => - e.duration === undefined - ? `built ${e.out}` - : `\n${e.out}\nduration=${e.duration}\nsize=${e.size}`, -}; - -let EMIT = (ev, fields = {}) => { - const line = HUMAN[ev]?.({ ev, ...fields }); - if (line) console.log(line); -}; - -export function setProgressMode(mode) { - EMIT = - mode === "ndjson" - ? (ev, fields = {}) => process.stdout.write(JSON.stringify({ ev, ...fields }) + "\n") - : (ev, fields = {}) => { - const line = HUMAN[ev]?.({ ev, ...fields }); - if (line) console.log(line); - }; -} - -function hms(total) { - const s = Math.floor(total); - const h = Math.floor(s / 3600); - const m = Math.floor((s % 3600) / 60); - const sec = s % 60; - return h > 0 - ? `${h}:${String(m).padStart(2, "0")}:${String(sec).padStart(2, "0")}` - : `${m}:${String(sec).padStart(2, "0")}`; -} - -// Stream titles here are full of emoji and !commands. drawtext renders them as -// tofu with a text font, and they add nothing to an attribution line. -function cleanTitle(title) { - return title - .replace(/[\u{1F000}-\u{1FFFF}\u{2600}-\u{27BF}\u{FE0F}]/gu, "") - .replace(/\s*[!@]\S+/g, "") - .replace(/\s{2,}/g, " ") - .replace(/[\s·|-]+$/, "") - .trim(); -} - -// drawtext does not wrap. Break to a character budget, write to a file, and use -// textfile= so nothing needs shell or filter escaping. -function wrap(text, cols) { - const words = text.split(/\s+/); - const lines = []; - let line = ""; - for (const w of words) { - if (line && (line + " " + w).length > cols) { - lines.push(line); - line = w; - } else { - line = line ? line + " " + w : w; - } - } - if (line) lines.push(line); - return lines.join("\n"); -} - -// The published shard record carries the same fields as a local cue file, so this -// reads identically whichever source answered. -async function videoMeta(videoId, channelSlug, hints = {}) { - const d = await CUES.load(channelSlug, videoId, hints); - return { title: d.title, uploadDate: d.uploadDate, webpageUrl: d.webpageUrl, duration: d.duration }; -} - -// The CONTAINER's duration is max(video, audio), and the audio is longer: the -// AAC encoder pads the front with ~21 ms of decoder delay, and a video duration -// is rarely an exact multiple of the frame interval. Either way the excess is -// small — and it ACCUMULATES through segmentOffsets, which subtracts one -// transition per segment and hands the result to xfade, the chapter marks and -// (now) the rail. A few hundred ms of drift by segment 20 is enough to land a -// rail row-change on the wrong side of a cut. -// -// The video stream's frame COUNT is the number the timeline actually runs on, -// so derive the duration from it. nb_frames is absent on some demuxers; fall -// back to the container rather than failing a build over a probe. -async function probeDuration(file, fps) { - if (fps) { - const { stdout } = await execFileP(FFPROBE, [ - "-v", "error", "-select_streams", "v:0", "-show_entries", "stream=nb_frames", - "-of", "default=nw=1:nk=1", file, - ]); - const n = Number(stdout.trim()); - if (Number.isFinite(n) && n > 0) return n / fps; - } - const { stdout } = await execFileP(FFPROBE, [ - "-v", "error", "-show_entries", "format=duration", - "-of", "default=nw=1:nk=1", file, - ]); - return Number(stdout.trim()); -} - -// yt-dlp exits 101 on a clean early stop (break-on-existing / max-downloads). -// The repo treats that as success everywhere else; do the same here. -const ytdlpOk = (err) => err?.code === 101; - -// ---- the clip cache ------------------------------------------------------ -// A raw clip's window is IN ITS NAME, which makes the file immutable and the -// cache content-addressed. The original lookup was for the exact name, so any -// change to a window -- a hand edit, a widen, a nudge in the clip bench -- was a -// fresh download of material already on disk. Measured on ferret-rescue: 31 -// files for 10 clips, one source fetched four times over overlapping windows. -// -// So: satisfy a request from ANY cached file that contains it. The TIGHTEST -// container wins, because detectSilence decodes the whole file and a 40s file -// costs more than the 14s one that would also have done. The clip bench fetches -// deliberately wide, and this is what makes that generous fetch become the -// build's cache rather than a second one. -const WINDOW_RE = /^(\d+(?:\.\d+)?)-(\d+(?:\.\d+)?)$/; - -// A window read back from a 2 dp manifest can sit a hair outside the file that -// produced it; the same tolerance resolve-windows.mjs uses for the same reason. -const WIN_EPS = 0.02; - -export async function cachedWindowsFor(rawDir, video) { - let names; - try { - names = await readdir(rawDir); - } catch { - return []; - } - const prefix = `${video}_`; - const out = []; - for (const name of names) { - if (!name.startsWith(prefix) || !name.endsWith(".mp4")) continue; - // The remainder must be exactly `a-b`, which is what stops a video id that - // is a prefix of another (or one containing `_`) from claiming its files. - const m = WINDOW_RE.exec(name.slice(prefix.length, -4)); - if (!m) continue; - out.push({ name, path: path.join(rawDir, name), from: Number(m[1]), to: Number(m[2]) }); - } - return out; -} - -/** The tightest cached file containing [from, to], or null. */ -export async function findContainingWindow(rawDir, video, from, to) { - const windows = await cachedWindowsFor(rawDir, video); - let best = null; - for (const w of windows) { - if (w.from > from + WIN_EPS || w.to < to - WIN_EPS) continue; - if (!best || w.to - w.from < best.to - best.from) best = w; - } - return best; -} - -async function fetchClip(entry, meta, render, rawDir, opts) { - // Deliberately over-fetch: the snapping pass below needs room on both sides to - // find a silence, and a clip that has no slack can only be cut where the cue - // happened to break — which is what put words in half in the first place. - const pad = opts.pad ?? render.fetchPad ?? 3.0; - const from = Math.max(0, entry.start - pad); - const to = entry.end + pad; - - // Shared across variants, and deliberately so: this is the only expensive - // thing in a build, and the two cuts overlap almost entirely. - const name = `${entry.video}_${from.toFixed(2)}-${to.toFixed(2)}.mp4`; - const dest = path.join(rawDir, name); - if (await exists(dest)) { - EMIT("fetch", { id: entry.id, video: entry.video, from, to, cached: true }); - return { path: dest, fetchStart: from, cached: true }; - } - if (!opts.noReuse) { - const hit = await findContainingWindow(rawDir, entry.video, from, to); - if (hit) { - EMIT("fetch", { - id: entry.id, video: entry.video, from, to, cached: true, reuse: hit.name, - }); - // fetchStart is the CACHED file's start, not the requested one -- every cut - // downstream is expressed relative to it, so reuse is transparent. - return { path: hit.path, fetchStart: hit.from, cached: true }; - } - } - if (opts.skipFetch) throw new Error(`--skip-fetch set and no cached window covers ${name}`); - - const maxH = render.maxHeightSource; - const fmt = [ - `bv*[vcodec^=avc1][height<=${maxH}]+ba[acodec^=mp4a]`, - `bv*[ext=mp4][height<=${maxH}]+ba[ext=m4a]`, - `b[ext=mp4][height<=${maxH}]`, - `b[height<=${maxH}]`, - ].join("/"); - - const argsWith = (extra) => [ - // The operator's own yt-dlp config redirects output and attaches thumbnail - // and metadata post-processors; without this the clips land elsewhere. - "--ignore-config", - "--no-playlist", - "--download-sections", `*${from.toFixed(2)}-${to.toFixed(2)}`, - // Without this the cut snaps to the nearest preceding keyframe, which can be - // seconds early — fine for scrubbing, not fine when the clip IS the citation. - "--force-keyframes-at-cuts", - ...extra, - // Pin H.264/AAC in mp4. Left alone yt-dlp picks VP9+Opus at these heights, - // and since --force-keyframes-at-cuts re-encodes, that means libvpx-vp9 — - // 27s to cut a 5s clip. It also writes .webm and appends that to -o. - "-f", fmt, - "--merge-output-format", "mp4", - "-o", dest, - "--", meta.webpageUrl, - ]; - - const attempt = async (extra) => { - try { - await execFileP(YTDLP, argsWith(extra), { maxBuffer: 1 << 26 }); - return null; - } catch (err) { - return ytdlpOk(err) ? null : err; - } - }; - - EMIT("fetch", { id: entry.id, video: entry.video, from, to, cached: false }); - let err = await attempt([]); - - // Rumble delivers HLS whose segments are named `.tar`, and ffmpeg 8's picky - // extension check rejects those outright — "URL ... is not in - // allowed_segment_extensions" — killing the fetch with exit 183. Rumble ships - // no progressive format to fall back to, so without this every Rumble-sourced - // clip is unbuildable. - // - // It has to be a RETRY, not a default: -extension_picky lives on the HLS - // demuxer, so passing it against a progressive URL (YouTube's googlevideo mp4) - // makes ffmpeg abort with "Option extension_picky not found" — i.e. adding it - // unconditionally trades a Rumble failure for a YouTube one. - if (err && /allowed_segment_extensions|allowed_extensions/.test(String(err.stderr ?? err.message ?? ""))) { - EMIT("note", { id: entry.id, message: ` ${entry.id}: HLS segment extension rejected, retrying with -extension_picky 0` }); - err = await attempt(["--downloader-args", "ffmpeg_i:-extension_picky 0"]); - } - if (err) { - throw new Error(`yt-dlp failed for ${entry.id} (${entry.video}): ${err.stderr ?? err.message}`); - } - if (!(await exists(dest))) { - // If a fallback format still forced another container, yt-dlp writes - // "<dest>.<realext>". Adopt it rather than failing the run. - const dir = path.dirname(dest); - const base = path.basename(dest); - const stray = (await readdir(dir)).find((f) => f.startsWith(base + ".")); - if (!stray) throw new Error(`yt-dlp reported success but produced no file for ${entry.id}`); - await rename(path.join(dir, stray), dest); - } - return { path: dest, fetchStart: from, cached: false }; -} - -// Parse ffmpeg's silencedetect output into [{s, e}] intervals, in seconds -// relative to the start of the given file. -async function detectSilence(file, render) { - const minDur = render.silenceMinDur ?? 0.09; - - // The threshold has to be RELATIVE to the clip, not absolute. These are game - // streams: the gaps between words are full of game audio and music, so they - // are quiet but nowhere near silent. A fixed -32 dB sits below the noise floor - // of a typical clip here and finds literally zero silences (measured: mean - // volume -21 dB, 0 hits at -32 dB, 25 hits at -26 dB). Measure the clip first - // and cut a few dB under its own mean instead. - const { stderr: volLog } = await execFileP( - FFMPEG, - ["-nostdin", "-i", file, "-af", "volumedetect", "-f", "null", "-"], - { maxBuffer: 1 << 26 }, - ).catch((e) => ({ stderr: e.stderr ?? "" })); - const meanMatch = (volLog ?? "").match(/mean_volume:\s*(-?[\d.]+) dB/); - const mean = meanMatch ? Number(meanMatch[1]) : -24; - const noise = Math.max(-45, Math.min(-18, mean - (render.silenceRelDb ?? 6))); - - // ffmpeg exits 0 here, so stderr comes back on the resolved result. - const { stderr } = await execFileP( - FFMPEG, - ["-nostdin", "-i", file, "-af", `silencedetect=noise=${noise.toFixed(1)}dB:d=${minDur}`, "-f", "null", "-"], - { maxBuffer: 1 << 26 }, - ).catch((e) => ({ stderr: e.stderr ?? "" })); - const log = stderr ?? ""; - - const out = []; - let open = null; - for (const line of log.split("\n")) { - const s = line.match(/silence_start:\s*(-?[\d.]+)/); - if (s) open = Number(s[1]); - const e = line.match(/silence_end:\s*(-?[\d.]+)/); - if (e && open !== null) { - out.push({ s: open, e: Number(e[1]) }); - open = null; - } - } - return out; -} - -// Snap a desired cut to the nearest silence, so the clip begins and ends between -// words instead of through one. Returns the desired point unchanged when no -// silence is close enough — better a tight cut than a cut in the wrong place. -function snap(desired, intervals, kind, window) { - let best = null; - for (const iv of intervals) { - // Starting: we want to resume just before speech does -> the silence's END. - // Ending: we want to stop just after speech does -> the silence's START. - const point = kind === "start" ? iv.e : iv.s; - const d = Math.abs(point - desired); - if (d > window) continue; - if (!best || d < best.d) best = { d, point }; - } - if (!best) return { at: desired, snapped: false }; - const lead = kind === "start" ? -0.10 : 0.18; - return { at: Math.max(0, best.point + lead), snapped: true }; -} - -// Encoder quality is manifest-driven so a cut can trade size for fidelity without -// editing this file. Defaults reproduce the original hardcoded settings exactly. -const encodeArgs = (render) => [ - "-c:v", "libx264", - "-preset", render.preset ?? "medium", - "-crf", String(render.crf ?? 20), - "-pix_fmt", "yuv420p", - "-r", String(render.fps), - "-c:a", "aac", - "-b:a", render.audioBitrate ?? "160k", - "-ar", String(render.audioRate), - "-ac", String(render.audioChannels), - "-movflags", "+faststart", -]; - -// The rail pass runs over an ALREADY ENCODED file, so its audio is already the -// finished AAC. Re-encoding it would cost a whole generation for nothing — and -// would make "the rail does not touch the audio" untrue. -const encodeArgsVideoOnly = (render) => [ - "-c:v", "libx264", - "-preset", render.preset ?? "medium", - "-crf", String(render.crf ?? 20), - "-pix_fmt", "yuv420p", - "-r", String(render.fps), - "-c:a", "copy", - "-movflags", "+faststart", -]; - -async function buildClipSegment(entry, meta, render, dirs, opts, chrome, nodes, provenance) { - const outDir = dirs.dir; - const { path: raw, fetchStart } = await fetchClip(entry, meta, render, dirs.rawDir, opts); - const seg = path.join(outDir, "segments", `${entry.id}.mp4`); - const pal = render.palette; - const { width, height } = render; - - // Desired cut points, expressed relative to the over-fetched file. - const wantA = entry.start - fetchStart; - const wantB = entry.end - fetchStart; - const win = render.snapWindow ?? 1.6; - - const sil = await detectSilence(raw, render); - const a = snap(wantA, sil, "start", win); - const b = snap(wantB, sil, "end", win); - // Never let snapping invert or collapse the window. - const cutA = Math.min(a.at, wantB - 1); - const cutB = Math.max(b.at, cutA + 1); - EMIT("snap", { id: entry.id, start: a.snapped, end: b.snapped, seconds: cutB - cutA }); - - const quotePath = path.join(outDir, "segments", `${entry.id}.quote.txt`); - const attribPath = path.join(outDir, "segments", `${entry.id}.attrib.txt`); - // Written for reference/diffing only — the quote is no longer drawn on screen. - await writeFile(quotePath, wrap(`“${entry.quote}”`, 92), "utf8"); - - const d = meta.uploadDate; - const date = `${d.slice(0, 4)}-${d.slice(4, 6)}-${d.slice(6, 8)}`; - await writeFile( - attribPath, - `${cleanTitle(meta.title)} · ${date} @ ${hms(entry.cite ?? entry.start)}`, - "utf8", - ); - - // The picture is the point. Nothing is drawn over it: the video is letterboxed - // between a thin citation header and a thin timeline footer, so the source - // material plays unobstructed and the additions stay subtle. - const HH = render.headerHeight ?? 56; - // The picture lives left of the rail column; the rail's own pixels are painted - // by the rail chain at concat time, over ground this pad leaves for it. - const VW = contentWidth(render); - const FH = chrome.footerHeight; - const hasFooter = FH > 0 && chrome.footer; - // headerHeight:0 drops the citation line too, leaving the clips alone on screen. - // Worth having: a cut whose sources are listed elsewhere does not need to carry - // its own attribution burnt into every frame. - const hasHeader = HH > 0; - const VH = height - HH - FH; - const trackAbsY = height - FH + chrome.trackY; - - // Where the progress marker travels this clip. Only the first clip of a - // section moves it; the rest hold it in place. - const T = render.slideSeconds ?? 0.9; - const xTo = chrome.xs[entry.section]; - const xFrom = entry.sectionEnter ? chrome.xs[Math.max(0, entry.section - 1)] : xTo; - // Commas inside a filter option have to survive filtergraph parsing; single - // quotes around the expression is what protects them. - const ramp = (a, b) => - a === b ? String(b) : `'if(lt(t,${T}),${a}+(${b}-${a})*t/${T},${b})'`; - const markX = ramp(xFrom - chrome.markerRadius, xTo - chrome.markerRadius); - - // The fill bar CANNOT be a drawbox with a `t`-dependent width. drawbox has no - // time variable at all: its `t` is the box THICKNESS, and with `t=fill` that - // is effectively INT_MAX, so the old `if(lt(t,0.9),…)` was always false and - // the bar was always drawn at its final width. (Proof: `drawbox=w='t*10'` and - // `drawbox=w=20` produce an identical YAVG.) Only the amber marker ever moved. - // - // So do it the way the rail does: a 2*LEN-wide strip, accent on the left half - // and transparent on the right, translated under a fixed-width crop. crop's - // x IS per-frame in `t`, and it clamps, so the ends are self-parking. - const fillA = xFrom - chrome.x0; - const fillB = xTo - chrome.x0; - const fillExpr = fillA === fillB - ? String(fillB) - : `${fillA}+(${fillB - fillA})*clip(t/${T},0,1)`; - - const base = [ - `scale=${VW}:${VH}:force_original_aspect_ratio=decrease`, - `pad=${VW}:${VH}:(ow-iw)/2:(oh-ih)/2:color=${pal.bg}`, - // Widen back to the full frame, leaving the rail column (if any) as ground. - `pad=${width}:${VH}:0:0:color=${pal.bg}`, - `pad=${width}:${height}:0:${HH}:color=${pal.bg}`, - "setsar=1", - `fps=${render.fps}`, - ...(hasHeader - ? [ - `drawbox=x=90:y=${Math.round((HH - 24) / 2)}:w=4:h=24:color=${pal.accent}:t=fill`, - [ - `drawtext=textfile='${attribPath}'`, - `fontfile='${render.fontRegular}'`, - "fontsize=22", - `fontcolor=${pal.muted}`, - "x=118", - `y=${Math.round((HH - 26) / 2)}`, - ].join(":"), - ] - : []), - ].join(","); - - // A manifest with a RAIL draws the code in the rail's foot instead, as one - // more strip: bottom-right of the frame, bordered, one per clip. It used to - // float over the bottom-right of the picture — the one part of the frame this - // cut promises never to draw on. Manifests with no rail keep the old overlay, - // byte for byte. - const qr = - render.qr === false || render.rail ? null : await qrForEntry(entry, provenance, render, outDir); - const qrM = render.qr?.margin ?? 28; - - // Bound the bar strip SHORTER than the clip. An overlay secondary that outruns - // the main extends the output, and the fix for that (shortest=1) would instead - // truncate the clip to the strip. Ending early is free: overlay's default - // eof_action=repeat holds the strip's last frame, which is the parked bar. - const barT = Math.max(0.2, cutB - cutA - 0.25); - - const inputs = ["-ss", cutA.toFixed(3), "-to", cutB.toFixed(3), "-i", raw]; - let nextIdx = 1; - let footerIdx, markerIdx, barIdx, qrIdx; - if (hasFooter) { - footerIdx = nextIdx++; inputs.push("-i", chrome.footer); - markerIdx = nextIdx++; inputs.push("-i", chrome.marker); - // A PNG fed with a plain -i through an ANIMATED crop is frozen: the crop - // sees one frame at t=0 and repeatlast repeats the already-cropped result. - // -loop 1 -framerate is what makes the strip a video the crop can walk. - barIdx = nextIdx++; - inputs.push( - "-loop", "1", "-framerate", String(render.fps), "-t", barT.toFixed(3), - "-i", chrome.bar, - ); - } - if (qr) { qrIdx = nextIdx++; inputs.push("-i", qr.png); } - - const parts = hasFooter - ? [ - `[0:v]${base}[b]`, - `[b][${footerIdx}:v]overlay=0:${height - FH}[f]`, - `[${barIdx}:v]crop=w=${chrome.trackLen}:h=3:x='${chrome.trackLen}-(${fillExpr})':y=0[bar]`, - `[f][bar]overlay=x=${chrome.x0}:y=${trackAbsY - 1}[g]`, - `[g][${markerIdx}:v]overlay=x=${markX}:y=${trackAbsY - chrome.markerRadius}[q]`, - ] - : [`[0:v]${base}[q]`]; - - // Sit above the footer when there is one, so the code never straddles the chrome. - parts.push( - qr - ? `[q][${qrIdx}:v]overlay=x=${VW}-w-${qrM}:y=H-h-${FH + qrM}[v]` - : `[q]null[v]`, - ); - - await execFileP( - FFMPEG, - [ - "-nostdin", "-v", "error", "-y", - ...inputs, - "-filter_complex", parts.join(";"), - "-map", "[v]", "-map", "0:a", - ...encodeArgs(render), - seg, - ], - { maxBuffer: 1 << 24 }, - ); - return seg; -} - -// ---- QR provenance code -------------------------------------------------- -// A compilation asks the viewer to take the edit on trust. The QR is the antidote: -// it resolves to this clip's exact START in the archive's own viewer, so anyone can -// pull up the surrounding hour and check that the cut is fair. Per clip, because a -// single code for the whole video would send everyone to the first citation. -// -// Two rules learned the hard way: it must be FULLY OPAQUE (a translucent QR will -// not scan) and it must keep its quiet zone (the white border is part of the -// symbol, not decoration). -async function qrForEntry(entry, provenance, render, outDir) { - const q = render.qr ?? {}; - // A mirror's LOCAL slug is not the id the site serves, and a clip taken from a - // copy whose archived transcript is broken should point at the copy that reads — - // so an explicit per-clip citeUrl always wins over the derived one. - const url = - entry.citeUrl ?? - `${provenance.siteOrigin}/?v=${encodeURIComponent( - `${entry.channel ?? provenance.channelSlug}/${entry.video}`, - )}&t=${Math.floor(entry.start)}`; - const png = path.join(outDir, "qr", `${entry.id}.png`); - await execFileP(QRENCODE, [ - "-o", png, - "-s", String(q.scale ?? 4), - "-m", String(q.quiet ?? 3), - "-l", q.ecc ?? "M", - url, - ]); - return { png, url }; -} - -async function buildCardSegment(card, render, outDir, nodes) { - const png = await renderCard(card, render, outDir, nodes); - const seg = path.join(outDir, "segments", `${card.id}.mp4`); - const dur = String(card.seconds); - - await execFileP( - FFMPEG, - [ - "-nostdin", "-v", "error", "-y", - // Without -framerate the image demuxer runs at its 25 fps default and the - // `-vf fps=30` below DUPLICATES a frame — at the segment's first frame, - // which is exactly where the next xfade seam lands. - "-loop", "1", "-framerate", String(render.fps), "-t", dur, "-i", png, - "-f", "lavfi", "-t", dur, - "-i", `anullsrc=channel_layout=stereo:sample_rate=${render.audioRate}`, - "-vf", `fps=${render.fps},setsar=1`, - ...encodeArgs(render), - "-shortest", - seg, - ], - { maxBuffer: 1 << 24 }, - ); - return seg; -} - -// =========================================================================== -// The claim rail -// =========================================================================== -// A persistent vertical ledger down the right edge, appending one row per claim -// as the video runs. It is folded into the concat pass rather than added as a -// second encode: the chain attaches AFTER the final xfade node, which already -// has post-pass semantics (nothing downstream of the last xfade is dissolved, -// and `t` there is absolute and continuous from 0). A separate pass would -// re-quantize crf-20 output, and antialiased text on flat colour is exactly the -// content that costs most. -// -// Five hard-won rules are load-bearing here; breaking any one produces a hang, -// a silently wrong-length file or a frozen overlay: -// -// 1. crop's w/h are CONFIG-TIME (`t` is undefined there) but x/y are -// per-frame. So every moving part is a fixed-size window walking a strip. -// 2. crop clamps x/y into range, so over-scroll is safe and self-parking. -// 3. A PNG on a plain -i through an animated crop is FROZEN. Every strip -// needs `-loop 1 -framerate <fps> -t <bound>`. -// 4. An UNBOUNDED `-loop 1` input deadlocks ffmpeg once several are chained. -// Hence `-t` on all five. -// 5. An overlay secondary longer than the main EXTENDS the output. The strips -// are deliberately longer (bound = total + 2), so `shortest=1` is required -// on EVERY rail overlay, not just the first. -// -// And two rendering ones: overlay's default `format=yuv420` subsamples alpha as -// well as chroma, which fringes 14–23 px rail text — so every rail overlay is -// `format=yuv444`, with a single `format=yuv420p` before the encoder. - -/** - * Fill an evenly-spaced schedule between known anchors. - * - * `known` holds the entries that are pinned to a segment; everything else is - * distributed linearly between its neighbouring pins, with `lo`/`hi` acting as - * virtual anchors just outside the run. - */ -function distribute(known, n, lo, hi) { - const at = new Array(n).fill(null); - for (const [i, t] of known) at[i] = t; - const pts = [[-1, lo], ...[...known].sort((a, b) => a[0] - b[0]), [n, hi]]; - for (let k = 0; k < pts.length - 1; k += 1) { - const [a, ta] = pts[k]; - const [b, tb] = pts[k + 1]; - for (let j = a + 1; j < b; j += 1) at[j] = ta + ((tb - ta) * (j - a)) / (b - a); - } - return at; -} - -/** - * When each ledger row appears, in finished-timeline seconds. - * - * ONE CHRONOLOGY. The cut plays in date order across every company, and the - * ledger is sorted the same way, so a claim's position in the rail IS its - * position in time. That collapses what used to live here: there is no longer a - * per-company span to bound a claim to, no contiguity rule to enforce, and no - * risk of scheduling a media claim over a coffee clip -- because "over a coffee - * clip" now means "later in the same chronology", which is exactly right. - * - * What remains is the part that was always doing the work: a claim WITH a clip - * behind it is pinned to that clip's segment, and the rest are spread evenly - * between their neighbouring pins. Monotonicity holds by construction, since - * both the pins and the rows are in date order. - * - * Every state change lands at `starts[i] + D/2` -- MID-DISSOLVE -- where the - * picture is already crossfading and a ±3-frame error is invisible. - */ -export function scheduleClaims(ledger, entries, starts, D, endBound) { - const segOf = new Map(entries.map((e, i) => [e.id, i])); - const mid = (seg) => starts[seg] + D / 2; - - // A stacked ledger card carries several claims, and each one has a MOMENT - // inside that card: the reveal of its own row. Pinning all of them to the - // segment's mid-dissolve would land four rail rows on one frame and, worse, - // break the pin-order guard's strict monotonicity for no reason. So a claim - // on such a card is pinned to its own row's reveal. - const within = new Map(); - for (const e of entries) { - if (e.type !== "ledger") continue; - (e.claims ?? []).forEach((cid, r) => within.set(`${e.id}|${cid}`, ledgerRevealAt(r))); - } - - const known = new Map(); - ledger.forEach((c, i) => { - const seg = c.entryId ? segOf.get(c.entryId) : undefined; - if (seg === undefined) return; - const off = within.get(`${c.entryId}|${c.id}`); - known.set(i, off === undefined ? mid(seg) : starts[seg] + off); - }); - if (!known.size) throw new Error("ledger: no claim is pinned to a clip, so nothing anchors the rail"); - - // A pin that runs backwards means the ledger and the timeline disagree about - // the order of events, which is a manifest bug rather than something to - // silently smooth over -- the whole cut rests on the two agreeing. - const pins = [...known].sort((a, b) => a[0] - b[0]); - for (let i = 1; i < pins.length; i += 1) { - if (pins[i][1] <= pins[i - 1][1]) { - throw new Error( - `ledger: ${ledger[pins[i][0]].id} is pinned to ${ledger[pins[i][0]].entryId}, which plays ` + - `before ${ledger[pins[i - 1][0]].id}'s clip — the ledger is not in the cut's order`, - ); - } - } - - // The first card is the title; the rail's own run opens just after it. - const times = distribute(known, ledger.length, mid(0), endBound); - - for (let i = 1; i < times.length; i += 1) { - if (times[i] <= times[i - 1]) times[i] = times[i - 1] + 1 / 30; - } - return times; -} - -/** - * The rail's filtergraph, as one builder with two call sites — the concat pass - * and `--rail-only` — so the two paths cannot drift. - * - * Ramps are CUMULATIVE AND SATURATING, never gated. A piecewise sum of - * `gte(t,s)*lt(t,s')*…` terms flashes to y=0 for one frame at any boundary gap, - * because every gate evaluates false at once and the sum collapses. Terms that - * rise to their delta and stay there cannot do that. - */ -export function railFilterChain(rail, assets, times, render, inLabel, firstInputIdx, bound, opts = {}) { - const g = assets.geom; - const SLIDE = rail.slide ?? 0.55; - const fps = render.fps; - - const P = (s) => `clip((t-${s.toFixed(3)})/${SLIDE},0,1)`; - // smoothstep() does not exist in ffmpeg's expression language. This is it. - const ease = (s) => { const p = P(s); return `${p}*${p}*(3-2*${p})`; }; - // Signed, and explicitly so. Joining terms with "+" was fine while every - // delta was a positive row height; a rolling cell FALLS as often as it rises, - // and `…+-40*x` is at best relying on ffmpeg's unary minus. - const sum = (y0, terms) => - terms.reduce( - (acc, t) => `${acc}${t.d < 0 ? "-" : "+"}${Math.abs(t.d)}*${t.f}`, - String(y0), - ); - const ramp = (y0, steps) => - sum(y0, steps.filter((st) => st.delta !== 0).map((st) => ({ d: st.delta, f: ease(st.at) }))); - - const { K, ROWH, RW, RX, RTOP, LOGH, LOGTOP, TALLYTOP } = g; - - const logY = ramp(0, times.map((t, i) => ({ at: t, delta: i + 1 > K ? ROWH : 0 }))); - // The curtain and the log MUST share the same eased P, or the curtain visibly - // lags the rows mid-slide and unrevealed claims flash into view. - const curtainY = ramp(LOGTOP, times.map((t, i) => ({ at: t, delta: i + 1 <= K ? ROWH : 0 }))); - const hlY = ramp(LOGTOP, times.map((t, i) => ({ at: t, delta: i > 0 && i < K ? ROWH : 0 }))); - - /** - * One lane's y, in the strip's own pixels. - * - * y(t) = r0 + Σ_k [ (a_k − b_{k−1})·gte(t,t_k) + (b_k − a_k)·ease(t_k) ] - * - * The first term is the instantaneous reposition to the next pair's starting - * row; the second is the roll itself. Both are CUMULATIVE AND SATURATING, - * which is the rail's hard rule: a gated piecewise sum flashes to y=0 for one - * frame at any boundary gap, because every gate goes false at once. - */ - const laneY = (lane) => { - const terms = []; - let prevB = 0; - lane.steps.forEach((st, i) => { - if (!st) return; - const at = times[i]; - const jump = (st.a - prevB) * lane.cellH; - const roll = (st.b - st.a) * lane.cellH; - if (jump !== 0) terms.push({ d: jump, f: `gte(t,${at.toFixed(3)})` }); - if (roll !== 0) terms.push({ d: roll, f: ease(at) }); - prevB = st.b; - }); - return sum(0, terms); - }; - - // The rail leaves by SLIDING OFF to the right, not by an enable= pop. One - // offset expression shared by every overlay, so the column moves as one - // object; `overlay`'s x is per-frame in `t`, which is what makes that - // possible at all. Cumulative and saturating, like everything else here. - const hideAt = opts.hideAt ?? null; - const OFF = hideAt == null ? "" : `+${RW + 8}*${ease(hideAt)}`; - const X = (x) => (OFF ? `'${x}${OFF}'` : String(x)); - - const i0 = firstInputIdx; - const files = [assets.chrome, assets.log, assets.curtain, assets.hl, assets.tally]; - if (assets.qr) files.push(assets.qr.path); - const inputs = files.flatMap((f) => [ - "-loop", "1", "-framerate", String(fps), "-t", bound.toFixed(3), "-i", f, - ]); - - const chain = [ - `[${i0 + 1}:v]crop=w=${RW}:h=${LOGH}:x=0:y='${logY}'[rlog]`, - `${inLabel}[${i0}:v]overlay=x=${X(RX)}:y=${RTOP}:format=yuv444:shortest=1[rr0]`, - `[rr0][rlog]overlay=x=${X(RX)}:y=${LOGTOP}:format=yuv444:shortest=1[rr1]`, - // The highlight goes UNDER the curtain: while the list is still filling, the - // row it marks has not been revealed yet, and the curtain is what hides it. - `[rr1][${i0 + 3}:v]overlay=x=${X(RX)}:y='${hlY}':format=yuv444:shortest=1[rr2]`, - `[rr2][${i0 + 2}:v]overlay=x=${X(RX)}:y='${curtainY}':format=yuv444:shortest=1[rr3]`, - ]; - - // One crop per lane out of the SINGLE tally strip. Four numbers that roll - // independently and a roster line that mostly does not, for one more input - // than the slab cost. - // - // `split` first, and it is NOT optional: a filtergraph link may be consumed - // exactly once, so five crops reading `[N:v]` is a parse error, not a - // shortcut. This is the whole reason the lanes share one PNG and still cost - // one input. - chain.push( - `[${i0 + 4}:v]split=${assets.lanes.length}${assets.lanes.map((_, j) => `[ts${j}]`).join("")}`, - ); - let lab = "[rr3]"; - assets.lanes.forEach((lane, j) => { - const isRoster = lane.kind === "roster"; - const h = isRoster ? g.ROSTERH : g.TALLYROWH; - const y = isRoster ? g.ROSTERTOP : TALLYTOP + j * g.TALLYROWH; - const x = isRoster ? RX + g.ROSTERX : RX + g.CELLX; - chain.push( - `[ts${j}]crop=w=${lane.w}:h=${h}:x=${lane.x}:y='${laneY(lane)}'[rc${j}]`, - `${lab}[rc${j}]overlay=x=${X(x)}:y=${y}:format=yuv444:shortest=1[rt${j}]`, - ); - lab = `[rt${j}]`; - }); - - // The provenance tile LAST, so the parked curtain cannot paint over it. - if (assets.qr) { - const qrY = sum(0, assets.qr.steps.map((st) => ({ d: st.delta, f: `gte(t,${st.at.toFixed(3)})` }))); - chain.push( - `[${i0 + 5}:v]crop=w=${g.TILEW}:h=${g.TILEH}:x=0:y='${qrY}'[rqr]`, - `${lab}[rqr]overlay=x=${X(RX + g.PAD)}:y=${g.TILETOP}:format=yuv444:shortest=1[rq]`, - ); - lab = "[rq]"; - } - - chain.push(`${lab}format=yuv420p[vout]`); - - return { inputs, chain: chain.join(";"), outLabel: "[vout]" }; -} - -/** - * The chrome as PNG-sequence overlays, for `render.chromeEngine: "hyperframes"`. - * - * OPT-IN, and absent it nothing below runs -- the ffmpeg chrome path is left - * byte-for-byte alone, which is the same bargain the rail was added under. - * - * The five ffmpeg traps the rail documents apply here unchanged, and two of them - * bite harder with an image sequence: - * - * * `format=yuv444` on EVERY overlay. overlay's default yuv420 subsamples - * ALPHA as well as chroma, which fringes small text -- and the band is - * nothing but small text. - * * `shortest=1` on EVERY overlay. A secondary longer than the main EXTENDS - * the output; the sequence is rendered to the same length as the concat, but - * a one-frame rounding difference either way must not change the duration. - * * One `format=yuv420p` before the encoder, once, at the end. - * - * A finite image sequence needs no `-t`: unlike `-loop 1` it ends by itself, so - * the deadlock the rail's five chained loops hit cannot happen here. - */ -export function chromeOverlayChain(render, regions, inLabel, firstInputIdx, opts = {}) { - const { outLabel = "[hfout]", final = true } = opts; - const inputs = []; - const parts = []; - let lab = inLabel; - regions.forEach((r, i) => { - inputs.push( - "-framerate", String(render.fps), - "-start_number", "1", - "-i", path.join(r.frames, "frame_%06d.png"), - ); - const idx = firstInputIdx + i; - const last = i === regions.length - 1; - const out = last && !final ? outLabel : `[hf${i}]`; - parts.push(`${lab}[${idx}:v]overlay=x=${r.x}:y=${r.y}:format=yuv444:shortest=1${out}`); - lab = out; - }); - if (final) parts.push(`${lab}format=yuv420p[vout]`); - return { - inputs, - chain: parts.join(";"), - outLabel: final ? "[vout]" : outLabel, - count: regions.length, - }; -} - -/** - * Where each rendered chrome region sits in the frame. - * - * The chart band REPLACES the footer node track rather than joining it: the - * track only moved at section handovers, which is precisely the fault the band - * exists to fix. So it takes the footer's ground and 100px more of it, and the - * picture loses that height. - */ -export function chromeRegions(render, outDir) { - const H = render.chart?.height ?? 200; - return [ - { - name: "chart", - frames: path.join(outDir, "chrome", "chart-frames"), - x: 0, - y: render.height - H, - width: contentWidth(render), - height: H, - }, - ]; -} - -/** - * The footer's stand-in when the chrome is drawn in a browser. - * - * It reserves the band's HEIGHT and draws nothing, so every segment letterboxes - * to the same picture box the overlay expects and the ground under the band is - * the palette background. `footer: null` is what switches the whole ffmpeg - * footer -- image, marker and fill bar -- off; the degenerate shape is the one - * renderFooterAssets already returns for a manifest with no nodes, so this path - * is not new. - */ -function reservedFooter(render) { - return { - footer: null, marker: null, bar: null, trackLen: 0, - footerHeight: render.chart?.height ?? 200, - trackY: 0, xs: [], x0: 0, markerRadius: 0, - }; -} - -/** - * Everything the rail chain needs that depends on the built segments. Returns - * null when the manifest does not ask for a rail — which is what keeps this - * whole feature opt-in and every existing report byte-for-byte unchanged. - */ -async function buildRailPlan(manifest, render, entries, segments, D, outDir) { - const rail = render.rail; - if (!rail || !manifest.ledger?.length) return null; - const { starts, total } = await segmentOffsets(segments, D, render.fps); - const assets = await renderRailAssets( - render, manifest.ledger, outDir, entries, manifest.provenance, - ); - // Every claim must be on the board before the closing ledger scroll reads it - // back, so the last section's spare rows are spread up to that segment. - const endIdx = entries.findIndex((e) => e.type === "scroll" || e.type === "chart"); - const endBound = endIdx > 0 ? starts[endIdx] : total; - const times = scheduleClaims(manifest.ledger, entries, starts, D, endBound); - - // The QR tile changes at the MID-DISSOLVE of every segment, instantaneously - // — a code that eased into place would spend the ease unscannable, and the - // picture is already crossfading there. - if (assets.qr) { - assets.qr.steps = entries.slice(1).map((_, i) => ({ - at: starts[i + 1] + D / 2, - delta: assets.geom.TILEH, - })); - } - - // Where the rail leaves. The closing ledger is a full-width card and the rail - // is the one thing on screen it would have to be read around, so the column - // slides off over that card's dissolve and does not come back. - const hideIdx = entries.findIndex((e) => e.hideRail); - const hideAt = hideIdx > 0 ? starts[hideIdx] : null; - - // The schedule, written down. - // - // The chart band has to sweep in step with the rail -- a playhead that tracks - // the current moment is the whole point of it -- and the only way it and the - // rail can be guaranteed to agree is for one of them to compute the schedule - // and the other to READ it. Recomputing from segment durations would be a - // second implementation of segmentOffsets(), and it would drift the first time - // the crossfade changed. Same rule as widen(): imported, never reimplemented. - await writeFile( - path.join(outDir, "schedule.json"), - JSON.stringify( - { - fps: render.fps, - transition: D, - total, - endBound, - segments: entries.map((e, i) => ({ id: e.id, type: e.type, start: starts[i] })), - claims: manifest.ledger.map((c, i) => ({ id: c.id, at: times[i], entryId: c.entryId ?? null })), - }, - null, - 2, - ) + "\n", - ); - - EMIT("note", { - message: `rail: ${manifest.ledger.length} claims, ${times.filter((_, i) => manifest.ledger[i].entryId).length} pinned, ` + - `window ${assets.geom.K} rows` + (hideAt == null ? "" : `, hides at ${hideAt.toFixed(1)}s`), - }); - return { assets, times, total, hideAt }; -} - -// ---- the end sequence ---------------------------------------------------- -// Two segment kinds that exist only to close the cut: the whole ledger read -// back in one scroll, then the same claims plotted. Both go through the SAME -// `fps=,setsar=1` and the SAME encodeArgs as every other segment — xfade -// rejects a mismatched link with "First input link parameters do not match", -// which would surface only at concat time, after every fetch has been paid for. - -async function buildScrollSegment(card, render, outDir, ledger) { - const { path: png, contentHeight, width: VW } = await renderScrollCard(card, render, ledger, outDir); - const seg = path.join(outDir, "segments", `${card.id}.mp4`); - const pal = render.palette; - const { width, height } = render; - const HH = render.headerHeight ?? 56; - // The band's height, not the manifest's footerHeight — otherwise the last - // 100px of the scroll play underneath the chart. - const FH = reservedFooterHeight(render); - const winH = height - HH - FH; - const dur = String(card.seconds); - - // clip() buys a free hold at BOTH ends, and crop's own clamping degrades an - // off-by-a-few contentHeight into a static last frame rather than an error. - const hold = card.hold ?? 2.0; - const travel = Math.max(0, contentHeight - winH); - const denom = Math.max(0.1, card.seconds - 2 * hold); - - await execFileP( - FFMPEG, - [ - "-nostdin", "-v", "error", "-y", - "-loop", "1", "-framerate", String(render.fps), "-t", dur, "-i", png, - "-f", "lavfi", "-t", dur, "-i", `color=c=${pal.bg}:s=${width}x${height}:r=${render.fps}`, - "-f", "lavfi", "-t", dur, - "-i", `anullsrc=channel_layout=stereo:sample_rate=${render.audioRate}`, - "-filter_complex", [ - `[0:v]crop=w=${VW}:h=${winH}:x=0:y='${travel}*clip((t-${hold})/${denom.toFixed(3)},0,1)'[win]`, - `[1:v][win]overlay=x=0:y=${HH}:shortest=1,fps=${render.fps},setsar=1[v]`, - ].join(";"), - "-map", "[v]", "-map", "2:a", - ...encodeArgs(render), - "-shortest", - seg, - ], - { maxBuffer: 1 << 24 }, - ); - return seg; -} - -/** - * A stacked ledger card: rows revealed in sequence by a walking curtain. - * - * The curtain is an opaque `pal.bg` rectangle that starts covering every row - * and steps down one row-height per reveal. Same device as the rail's, and for - * the same reason: the card ground is flat, so an opaque rectangle over it is - * an exact in-place wipe with no per-pixel filter. - * - * The ramp is CUMULATIVE AND SATURATING, like every other ramp here. - */ -async function buildLedgerSegment(card, render, outDir, ledger, avail) { - const geo = await renderLedgerCard(card, render, ledger, outDir, avail); - const seg = path.join(outDir, "segments", `${card.id}.mp4`); - const pal = render.palette; - const { width, height } = render; - // `seconds` is DERIVED, not authored: the pins that land claims on their own - // rows read the same clock, so a hand-set duration would silently move them. - const dur = String(card.seconds ?? ledgerSeconds(geo.rows)); - - const curtainH = height; - const curtain = path.join(outDir, "cards", `${card.id}.curtain.png`); - await execFileP("magick", [ - "-size", `${geo.width}x${curtainH}`, `xc:${pal.bg}`, curtain, - ]); - - const SLIDE = render.rail?.slide ?? 0.55; - const ease = (at) => { - const p = `clip((t-${at.toFixed(3)})/${SLIDE},0,1)`; - return `${p}*${p}*(3-2*${p})`; - }; - const y = [ - String(geo.rowsTop), - ...Array.from({ length: geo.rows }, (_, r) => `${geo.rowHeight}*${ease(ledgerRevealAt(r))}`), - ].join("+"); - - await execFileP( - FFMPEG, - [ - "-nostdin", "-v", "error", "-y", - "-f", "lavfi", "-t", dur, "-i", `color=c=${pal.bg}:s=${width}x${height}:r=${render.fps}`, - "-loop", "1", "-framerate", String(render.fps), "-t", dur, "-i", geo.path, - "-loop", "1", "-framerate", String(render.fps), "-t", dur, "-i", curtain, - "-f", "lavfi", "-t", dur, - "-i", `anullsrc=channel_layout=stereo:sample_rate=${render.audioRate}`, - "-filter_complex", [ - `[0:v][1:v]overlay=x=0:y=0:shortest=1[a]`, - `[a][2:v]overlay=x=0:y='${y}':shortest=1,fps=${render.fps},setsar=1[v]`, - ].join(";"), - "-map", "[v]", "-map", "3:a", - ...encodeArgs(render), - "-shortest", - seg, - ], - { maxBuffer: 1 << 24 }, - ); - return seg; -} - -async function buildChartSegment(card, render, outDir, ledger) { - const chart = await renderChartCard(card, render, ledger, outDir); - const seg = path.join(outDir, "segments", `${card.id}.mp4`); - const pal = render.palette; - const { width, height } = render; - const VW = cardWidth(card, render); - const dur = String(card.seconds); - - // The wipe CANNOT be `crop=w='<ramp>'` — crop's w is config-time and `t` is - // undefined there ("Error when evaluating the expression"). So: overlay the - // finished chart, then slide an opaque pal.bg rectangle rightwards off it. - // The card ground is flat pal.bg, so this is an exact in-place wipe with no - // per-pixel filter, and it draws the plot in like a plotter. - // `hold: true` -- the closing chart is a HOLD, not a reveal. - // - // The wipe existed because this card was the first and only time the viewer - // saw the numbers plotted. With the chart band drawing live under the whole - // cut, wiping it in again would re-tell a story the viewer has just watched - // happen. So the card opens on the finished plot and the seconds go to - // reading the final gap and its flags instead. - if (card.hold) { - await execFileP( - FFMPEG, - [ - "-nostdin", "-v", "error", "-y", - "-f", "lavfi", "-t", dur, "-i", `color=c=${pal.bg}:s=${width}x${height}:r=${render.fps}`, - "-loop", "1", "-framerate", String(render.fps), "-t", dur, "-i", chart.path, - "-f", "lavfi", "-t", dur, - "-i", `anullsrc=channel_layout=stereo:sample_rate=${render.audioRate}`, - "-filter_complex", - `[0:v][1:v]overlay=x=0:y=0:shortest=1,fps=${render.fps},setsar=1[v]`, - "-map", "[v]", "-map", "2:a", - ...encodeArgs(render), - "-shortest", - seg, - ], - { maxBuffer: 1 << 24 }, - ); - return seg; - } - - const wipeW = VW - chart.plotX; - const wipe = path.join(outDir, "cards", `${card.id}.wipe.png`); - await execFileP("magick", [ - "-size", `${wipeW}x${Math.round(chart.plotH)}`, `xc:${pal.bg}`, wipe, - ]); - - const wipeStart = card.wipeStart ?? 0.8; - const wipeDur = card.wipeSeconds ?? Math.max(1, card.seconds - wipeStart - 3.0); - - await execFileP( - FFMPEG, - [ - "-nostdin", "-v", "error", "-y", - "-f", "lavfi", "-t", dur, "-i", `color=c=${pal.bg}:s=${width}x${height}:r=${render.fps}`, - "-loop", "1", "-framerate", String(render.fps), "-t", dur, "-i", chart.path, - "-loop", "1", "-framerate", String(render.fps), "-t", dur, "-i", wipe, - "-f", "lavfi", "-t", dur, - "-i", `anullsrc=channel_layout=stereo:sample_rate=${render.audioRate}`, - "-filter_complex", [ - `[0:v][1:v]overlay=x=0:y=0:shortest=1[a]`, - `[a][2:v]overlay=x='${chart.plotX}+${wipeW}*clip((t-${wipeStart})/${wipeDur.toFixed(3)},0,1)'` + - `:y=${Math.round(chart.plotY)}:shortest=1,fps=${render.fps},setsar=1[v]`, - ].join(";"), - "-map", "[v]", "-map", "3:a", - ...encodeArgs(render), - "-shortest", - seg, - ], - { maxBuffer: 1 << 24 }, - ); - return seg; -} - -// Crossfade every segment into the next. This is a full re-encode of the -// timeline — the concat demuxer can only stream-copy hard cuts — so --no-xfade -// stays available for quick iteration. -async function concatWithXfade(segments, render, outPath, railPlan, chrome = null) { - const D = render.transition ?? 0.5; - const durs = []; - for (const s of segments) durs.push(await probeDuration(s, render.fps)); - - const inputs = segments.flatMap((s) => ["-i", s]); - const parts = []; - let vlab = "[0:v]"; - let alab = "[0:a]"; - let acc = durs[0]; - - for (let i = 1; i < segments.length; i += 1) { - const off = acc - D; - parts.push(`${vlab}[${i}:v]xfade=transition=fade:duration=${D}:offset=${off.toFixed(3)}[v${i}]`); - parts.push(`${alab}[${i}:a]acrossfade=d=${D}:c1=tri:c2=tri[a${i}]`); - vlab = `[v${i}]`; - alab = `[a${i}]`; - acc = acc + durs[i] - D; - } - - // The rail attaches to the LAST xfade node, so it runs after every dissolve - // and sees an absolute, continuous `t`. One encode, not two. - // - // When the chrome is rendered rather than drawn, the PNG regions go on FIRST - // and the rail chain reads their output. Not the other way round: the rail - // chain ends in `format=yuv420p`, and overlaying an alpha sequence onto - // yuv420p is the fringing trap the rail already documents, one layer later. - const chromeIn = chrome ?? null; - const railIn = chromeIn ? chromeIn.outLabel : vlab; - const rc = railPlan - ? railFilterChain( - render.rail, railPlan.assets, railPlan.times, render, - railIn, segments.length, railPlan.total + 2, - { hideAt: railPlan.hideAt }, - ) - : null; - const railInputs = rc ? rc.inputs.filter((a) => a === "-i").length : 0; - const hf = chromeIn - ? chromeOverlayChain(render, chromeIn.regions, vlab, segments.length + railInputs, { - outLabel: chromeIn.outLabel, - final: !rc, - }) - : null; - if (hf) parts.push(hf.chain); - if (rc) parts.push(rc.chain); - - const tail = rc ? rc.outLabel : hf ? hf.outLabel : vlab; - - await execFileP( - FFMPEG, - [ - "-nostdin", "-v", "error", "-y", - ...inputs, - ...(rc ? rc.inputs : []), - ...(hf ? hf.inputs : []), - "-filter_complex", parts.join(";"), - "-map", tail, "-map", alab, - ...encodeArgs(render), - outPath, - ], - { maxBuffer: 1 << 26 }, - ); -} - -/** - * Run the rail chain over an already-concatenated file. - * - * Two callers need this. `--rail-only` iterates on the rail in seconds instead - * of re-running the whole concat; and `--no-xfade` has no choice, because - * concatHardCut is `-c copy` and a stream-copy mux cannot host a filtergraph - * at all. - */ -async function applyRail(inPath, outPath, render, railPlan, preview) { - const rc = railFilterChain( - render.rail, railPlan.assets, railPlan.times, render, - preview ? "[base]" : "[0:v]", 1, railPlan.total + 2, - { hideAt: railPlan.hideAt }, - ); - const parts = []; - if (preview) { - // -ss restarts `t` near zero, which would put every absolute-time ramp in - // the wrong place — the rail would look broken while being correct. Shift - // the timestamps back to where the expressions think they are, then rebase - // them so the preview file still starts at 0. - parts.push(`[0:v]setpts=PTS+${preview.start.toFixed(3)}/TB[base]`); - } - parts.push(rc.chain); - const tail = preview ? "[vshift]" : rc.outLabel; - if (preview) parts.push(`${rc.outLabel}setpts=PTS-STARTPTS[vshift]`); - - await execFileP( - FFMPEG, - [ - "-nostdin", "-v", "error", "-y", - ...(preview ? ["-ss", String(preview.start), "-t", String(preview.dur)] : []), - "-i", inPath, - ...rc.inputs, - "-filter_complex", parts.join(";"), - "-map", tail, "-map", "0:a", - ...encodeArgsVideoOnly(render), - outPath, - ], - { maxBuffer: 1 << 26 }, - ); -} - -// ---- chapter markers ----------------------------------------------------- -// A compilation like this is a reference document as much as a video: the report -// cites moments, and a viewer wants to jump to them. Every clip therefore becomes -// a chapter. Offsets are derived exactly the way concatWithXfade derives its xfade -// offsets, so they stay correct for both crossfaded and hard-cut timelines. -// -// ffmetadata is a line-based format where =, ;, # and \ are structural, so a -// title carrying any of them has to be escaped or the file silently mis-parses. -const ffmetaEscape = (s) => String(s).replace(/([=;#\\])/g, "\\$1").replace(/\n/g, " "); - -export async function segmentOffsets(segments, D, fps) { - const durs = []; - for (const s of segments) durs.push(await probeDuration(s, fps)); - const starts = []; - let acc = 0; - for (let i = 0; i < durs.length; i += 1) { - starts.push(acc); - acc += durs[i] - (i < durs.length - 1 ? D : 0); - } - return { starts, total: acc }; -} - -async function chapterTitle(entry, index, provenance) { - if (entry.chapter) return entry.chapter; - if (entry.type !== "clip") return entry.title ?? entry.heading ?? `Card ${index + 1}`; - try { - const meta = await videoMeta(entry.video, entry.channel ?? provenance.channelSlug, { siteChannel: entry.siteChannel, siteVideo: entry.siteVideo }); - const d = String(meta.uploadDate ?? ""); - const date = /^\d{8}$/.test(d) ? `${d.slice(0, 4)}-${d.slice(4, 6)}-${d.slice(6, 8)}` : d; - const title = String(meta.title ?? entry.video); - return `${date} — ${title.length > 60 ? `${title.slice(0, 57)}…` : title}`.trim(); - } catch { - return `${index + 1}. ${entry.video}`; - } -} - -async function muxChapters(finalPath, entries, segments, D, outDir, provenance, fps) { - if (segments.length < 2) return; - const { starts, total } = await segmentOffsets(segments, D, fps); - const lines = [";FFMETADATA1", ""]; - for (let i = 0; i < entries.length; i += 1) { - // Land just PAST the crossfade, so the marker opens on the incoming clip - // rather than on the outgoing one mid-dissolve. - const start = i === 0 ? 0 : starts[i] + D; - const end = i === entries.length - 1 ? total : starts[i + 1] + D; - lines.push( - "[CHAPTER]", - "TIMEBASE=1/1000", - `START=${Math.round(start * 1000)}`, - `END=${Math.round(end * 1000)}`, - `title=${ffmetaEscape(await chapterTitle(entries[i], i, provenance))}`, - "", - ); - } - const metaPath = path.join(outDir, "chapters.ffmeta"); - await writeFile(metaPath, lines.join("\n"), "utf8"); - - // Stream copy — adding chapters must never re-encode the finished timeline. - const tmp = finalPath.replace(/\.mp4$/, ".chapters.mp4"); - await execFileP( - FFMPEG, - ["-nostdin", "-v", "error", "-y", "-i", finalPath, "-i", metaPath, - "-map", "0", "-map_metadata", "0", "-map_chapters", "1", "-c", "copy", tmp], - { maxBuffer: 1 << 24 }, - ); - await rename(tmp, finalPath); - EMIT("chapters", { n: entries.length, file: path.basename(metaPath) }); -} - -async function concatHardCut(segments, outDir, outPath) { - const listPath = path.join(outDir, "concat.txt"); - await writeFile(listPath, segments.map((s) => `file '${s}'`).join("\n") + "\n", "utf8"); - await execFileP( - FFMPEG, - ["-nostdin", "-v", "error", "-y", "-f", "concat", "-safe", "0", - // The concat demuxer stitches per-file timestamps; without generated PTS a - // stream copy can hand the next stage a discontinuous timeline, which the - // rail's absolute-time expressions would then read off by that much. - "-fflags", "+genpts", - "-i", listPath, "-c", "copy", outPath], - { maxBuffer: 1 << 24 }, - ); -} - -// A hard-cut concat and a crossfaded one are different lengths, so a cached -// prerail from one is a wrong base for the other. Keeping them in separate files -// means the mode can be switched without a stale-cache trap -- and without the -// length assertion below having to be the thing that explains it. -const prerailPath = (outDir, slug, D) => - path.join(outDir, `${slug}.prerail${D === 0 ? "-hardcut" : ""}.mp4`); - -// The finished timeline must be exactly as long as segmentOffsets says. Anything -// else means a filter changed the length behind our backs. -async function assertConcatLength(file, expected, fps, what) { - const got = await probeDuration(file, fps); - if (Math.abs(got - expected) > 1.5 / fps) { - throw new Error( - `${what}: duration ${got.toFixed(3)}s but the timeline is ${expected.toFixed(3)}s ` + - `(${((got - expected) * fps).toFixed(1)} frames out)` + - (/prerail/.test(what) ? " — delete it and let this rebuild it" : ""), - ); - } -} - -/** - * Build a manifest into a video. - * - * Exported so umtool's driver runs the SAME code the CLI does. It is still - * SPAWNED rather than imported by the app: a 40-minute chain of yt-dlp and - * ffmpeg inside a request handler has no cancellation story, and a runaway - * grandchild would outlive the request that started it. - */ -export async function buildVideo({ manifestPath, opts = {}, out, only, fetchOnly } = {}) { - const variant = opts.variant ?? "sourced"; - const whole = JSON.parse(await readFile(manifestPath, "utf8")); - const manifest = selectVariant(whole, variant); - const { render, provenance } = manifest; - - // The manifest already records which archive it was built against, so a clone - // with no corpus needs no extra configuration to read cue windows. - CUES = createCueSource({ - siteOrigin: opts.siteOrigin ?? process.env.SITE_ORIGIN ?? siteOriginFromManifest(whole), - resolveSiteIds: opts.resolveSiteIds === true, - prefer: opts.cueSource ?? "auto", - log: (m) => EMIT("log", { message: m }), - }); - const outRoot = out ?? path.join(path.dirname(path.resolve(manifestPath)), "out"); - const dirs = variantPaths(outRoot, manifest.slug, variant); - const outDir = dirs.dir; - - await mkdir(dirs.rawDir, { recursive: true }); - for (const d of ["cards", "segments", "qr"]) { - await mkdir(path.join(outDir, d), { recursive: true }); - } - - // Fetch one clip's window and stop. This is what the clip bench's "fetch 20s - // more" runs, so a bench fetch and a build fetch can never disagree about - // naming, format selection, the VP9 trap or the Rumble HLS retry. - // Reads the WHOLE manifest, not the variant's view of it: a clip bench fetch - // is about a moment in the corpus, and which cut happens to carry it is - // beside the point. - if (fetchOnly) { - let entry = whole.timeline.find((e) => e.id === fetchOnly); - // `!== "clip"`, not `=== "card"`. The timeline's vocabulary is OPEN -- one - // real manifest carries `scroll` and `chart` entries -- and the card-only - // check sent `undefined` into the fetcher for either of those. - if (entry && entry.type !== "clip") { - throw new Error(`${fetchOnly} is a ${entry.type ?? "non-clip"} entry, not a clip`); - } - if (!entry) { - // A LEDGER CLAIM. Adjudicating one means listening around the moment, and - // most of the ledger is cited by no clip at all -- so the claim page asks - // for a window the timeline has no entry for. It is fetched through this - // same path so the file lands in clips-raw under the build's own naming, - // inherits the format pin and the Rumble HLS retry, and is REUSED by a - // later build rather than fetched a second time. - const claim = (whole.ledger ?? []).find((e) => e.id === fetchOnly); - if (!claim) throw new Error(`no timeline entry or ledger claim with id ${fetchOnly}`); - if (!claim.video) throw new Error(`ledger claim ${fetchOnly} has no \`video\` to fetch`); - const at = Number(claim.cite); - if (!Number.isFinite(at)) throw new Error(`ledger claim ${fetchOnly} has no \`cite\` second`); - // A claim is a MOMENT, not a window: the pad is the whole point, so the - // entry is a hair either side of the cite and --pad does the rest. - entry = { - id: claim.id, - video: claim.video, - channel: claim.channel ?? null, - start: Math.max(0, at - 1), - end: at + 1, - }; - } - const meta = await videoMeta(entry.video, entry.channel ?? provenance.channelSlug, { siteChannel: entry.siteChannel, siteVideo: entry.siteVideo }); - const r = await fetchClip(entry, meta, render, dirs.rawDir, opts); - EMIT("done", { out: r.path, fetchStart: r.fetchStart, cached: r.cached }); - return { out: r.path, failures: [] }; - } - - // Footer chrome is shared by every clip, so build it once up front. - const hyper = render.chromeEngine === "hyperframes"; - const chrome = hyper - ? reservedFooter(render) - : await renderFooterAssets(render, manifest.timelineNodes, outDir); - - const entries = manifest.timeline.filter((e) => !only || e.id === only); - if (only && !entries.length) throw new Error(`no timeline entry with id ${only}`); - - // Read once, at the ROOT: a source's state is a fact about the manifest, not - // about a variant. A stacked ledger card says why each claim is text rather - // than footage, and this is where that answer comes from. - const availability = new Map( - ( - await readFile(path.join(dirs.root, "availability.json"), "utf8").then( - (j) => JSON.parse(j).sources ?? [], - () => [], - ) - ).flatMap((src) => (src.claims ?? []).map((id) => [id, src.state])), - ); - const segments = []; - const failures = []; - - const D = opts.noXfade || (render.transition ?? 0.5) === 0 ? 0 : render.transition ?? 0.5; - - // Retro-fit chapters onto an already-built file without re-encoding it. The - // per-clip segments are still on disk, which is all the offsets need. - if (opts.chaptersOnly) { - const finalPath = dirs.final; - const segs = entries.map((e) => path.join(outDir, "segments", `${e.id}.mp4`)); - for (const seg of segs) { - if (!(await exists(seg))) - throw new Error(`--chapters-only needs ${seg}, which is missing — run a full build first`); - } - await muxChapters(finalPath, entries, segs, D, outDir, provenance, render.fps); - return { out: finalPath, failures: [] }; - } - - // Re-run the rail over a cached concat instead of rebuilding the timeline. - // The rail is the part that gets iterated on; the 40-minute concat is not. - if (opts.railOnly) { - if (!render.rail) throw new Error("--rail-only needs render.rail in the manifest"); - const finalPath = dirs.final; - const prerail = prerailPath(outDir, manifest.slug, D); - const segs = entries.map((e) => path.join(outDir, "segments", `${e.id}.mp4`)); - for (const seg of segs) { - if (!(await exists(seg))) - throw new Error(`--rail-only needs ${seg}, which is missing — run a full build first`); - } - if (!(await exists(prerail))) { - EMIT("concat", { mode: D === 0 ? "hardcut" : "xfade", n: segs.length }); - if (D === 0) await concatHardCut(segs, outDir, prerail); - else await concatWithXfade(segs, render, prerail, null); - } - const railPlan = await buildRailPlan(manifest, render, entries, segs, D, outDir); - await assertConcatLength(prerail, railPlan.total, render.fps, - `cached ${path.basename(prerail)}`); - const out = opts.preview - ? path.join(outDir, `${manifest.slug}.preview.mp4`) - : finalPath; - await applyRail(prerail, out, render, railPlan, opts.preview ?? null); - if (!opts.preview) { - await assertConcatLength(out, railPlan.total, render.fps, "rail build"); - // applyRail re-encodes, so the chapters muxed onto the previous final are - // gone. Put them back, or --rail-only quietly ships a chapterless cut. - if (!opts.noChapters) { - await muxChapters(out, entries, segs, D, outDir, provenance, render.fps); - } - } - EMIT("done", { out, failures: [] }); - return { out, failures: [] }; - } - - EMIT("start", { title: manifest.title, entries: entries.length, out: outDir }); - for (let i = 0; i < entries.length; i += 1) { - const entry = entries[i]; - try { - if (entry.type === "card") { - EMIT("card", { id: entry.id, i, n: entries.length }); - segments.push(await buildCardSegment(entry, render, outDir, manifest.timelineNodes)); - } else if (entry.type === "scroll" || entry.type === "chart" || entry.type === "ledger") { - EMIT("card", { id: entry.id, i, n: entries.length }); - if (!manifest.ledger?.length) - throw new Error(`${entry.id} is type:${entry.type} but the manifest has no ledger[]`); - segments.push( - entry.type === "scroll" - ? await buildScrollSegment(entry, render, outDir, manifest.ledger) - : entry.type === "chart" - ? await buildChartSegment(entry, render, outDir, manifest.ledger) - : await buildLedgerSegment(entry, render, outDir, manifest.ledger, availability), - ); - } else { - const meta = await videoMeta(entry.video, entry.channel ?? provenance.channelSlug, { siteChannel: entry.siteChannel, siteVideo: entry.siteVideo }); - EMIT("clip", { - id: entry.id, i, n: entries.length, video: entry.video, - section: entry.section, sectionEnter: !!entry.sectionEnter, - }); - segments.push( - await buildClipSegment( - entry, meta, render, dirs, opts, chrome, manifest.timelineNodes, provenance, - ), - ); - } - EMIT("segment", { id: entry.id, path: segments[segments.length - 1] }); - } catch (err) { - // Without --continue-on-error a dead source at entry 14 of 19 throws away - // the thirteen fetches already paid for. With it, everything buildable is - // built and the run reports what was not. - if (!opts.continueOnError) throw err; - const message = err?.message ?? String(err); - failures.push({ id: entry.id, message }); - EMIT("entry-failed", { id: entry.id, message }); - } - } - - if (only) { - EMIT("done", { out: segments[0], failures }); - return { out: segments[0], failures }; - } - - // A timeline that silently lost a clip is a worse outcome than no file at all: - // the finished video would look complete and be missing a citation. So the - // segments are kept (they cost the fetches) and the concat is refused. - if (failures.length) { - EMIT("note", { - message: `refusing to concat: ${failures.length} of ${entries.length} entries failed ` + - `(${failures.map((f) => f.id).join(", ")})`, - }); - return { out: null, failures }; - } - - const final = dirs.final; - const railPlan = opts.noRail ? null : await buildRailPlan(manifest, render, entries, segments, D, outDir); - - // The rendered chrome, if this manifest asks for it. Absent, `chromePlan` is - // null and every line below behaves exactly as it did -- which is the claim - // the MD5 check tests. - let chromePlan = null; - if (hyper) { - const regions = chromeRegions(render, outDir); - for (const r of regions) { - if (!(await exists(path.join(r.frames, "frame_000001.png")))) { - throw new Error( - `render.chromeEngine is "hyperframes" but ${r.name} has no frames at ${r.frames}. ` + - `Run compose-chrome.mjs --region ${r.name} --render first.`, - ); - } - } - chromePlan = { regions, outLabel: "[hfout]" }; - EMIT("note", { message: `chrome: ${regions.map((r) => `${r.name} ${r.width}x${r.height}`).join(", ")} as png-sequence` }); - } - - // `transition: 0` is a real editorial choice, not just a speed knob: hard cuts - // hit harder on a compilation whose point is repetition. Honouring it here keeps - // the manifest the source of truth, so a rebuild does not silently re-add fades. - EMIT("concat", { mode: D === 0 ? "hardcut" : "xfade", n: segments.length }); - if (D === 0) { - if (chromePlan) { - throw new Error( - 'render.chromeEngine "hyperframes" needs a filtergraph, and `transition: 0` concatenates with ' + - "-c copy, which cannot host one. Give the manifest a transition, or drop chromeEngine.", - ); - } - // concatHardCut is `-c copy`, which cannot host a filtergraph, so the rail - // has to be a second pass here whether we like it or not. - const prerail = prerailPath(outDir, manifest.slug, D); - await concatHardCut(segments, outDir, railPlan ? prerail : final); - if (railPlan) { - await assertConcatLength(prerail, railPlan.total, render.fps, "hard-cut concat"); - await applyRail(prerail, final, render, railPlan, null); - } - } else { - await concatWithXfade(segments, render, final, railPlan, chromePlan); - } - - // Length is the canary for the two ways a rail input can go wrong: a file - // LONGER than the timeline means a strip outran the main (a missing - // shortest=1), and a hang means an unbounded -loop 1. - if (railPlan) await assertConcatLength(final, railPlan.total, render.fps, "rail build"); - - if (!opts.noChapters) await muxChapters(final, entries, segments, D, outDir, provenance, render.fps); - - const { stdout } = await execFileP(FFPROBE, [ - "-v", "error", "-show_entries", "format=duration,size", - "-of", "default=noprint_wrappers=1", final, - ]); - const probe = Object.fromEntries( - stdout.trim().split("\n").map((l) => l.split("=")), - ); - EMIT("done", { out: final, duration: Number(probe.duration), size: Number(probe.size) }); - return { out: final, failures }; -} - -async function main() { - const argv = process.argv.slice(2); - const manifestPath = argv.find((a) => !a.startsWith("--")); - if (!manifestPath) { - console.error( - "usage: build-video.mjs <manifest.json> [--out <dir>] [--variant sourced|full]\n" + - " [--only <id>] [--fetch-only <id>]\n" + - " [--pad <s>] [--skip-fetch] [--no-xfade] [--no-chapters] [--chapters-only]\n" + - " [--progress ndjson] [--continue-on-error] [--no-reuse]\n" + - " [--no-rail] [--rail-only] [--preview <start> <dur>]\n" + - " [--site-origin <url>] [--resolve-site-ids] [--cue-source auto|local|http]", - ); - process.exit(2); - } - const flag = (n) => { - const i = argv.indexOf(n); - return i >= 0 ? argv[i + 1] : undefined; - }; - setProgressMode(flag("--progress") ?? "human"); - - const padArg = flag("--pad"); - const opts = { - variant: flag("--variant") ?? "sourced", - skipFetch: argv.includes("--skip-fetch"), - continueOnError: argv.includes("--continue-on-error"), - noXfade: argv.includes("--no-xfade"), - noChapters: argv.includes("--no-chapters"), - chaptersOnly: argv.includes("--chapters-only"), - noReuse: argv.includes("--no-reuse"), - noRail: argv.includes("--no-rail"), - railOnly: argv.includes("--rail-only"), - pad: padArg === undefined ? undefined : Number(padArg), - siteOrigin: flag("--site-origin"), - resolveSiteIds: argv.includes("--resolve-site-ids"), - cueSource: flag("--cue-source"), - }; - const pv = argv.indexOf("--preview"); - if (pv >= 0) { - opts.preview = { start: Number(argv[pv + 1]), dur: Number(argv[pv + 2]) }; - opts.railOnly = true; - if (!Number.isFinite(opts.preview.start) || !Number.isFinite(opts.preview.dur)) { - console.error("--preview takes <start> <dur> in seconds"); - process.exit(2); - } - } - - const { failures } = await buildVideo({ - manifestPath, - opts, - out: flag("--out"), - only: flag("--only"), - fetchOnly: flag("--fetch-only"), - }); - // Non-zero on a partial run, so a caller that ignores the events still learns - // the build did not produce what was asked for. - if (failures.length) process.exit(1); -} - -if (import.meta.url === `file://${process.argv[1]}`) { - main().catch((err) => { - EMIT("error", { message: err?.message ?? String(err) }); - console.error(err.message ?? err); - process.exit(1); - }); -} diff --git a/scripts/report-to-video/check-availability.mjs b/scripts/report-to-video/check-availability.mjs @@ -1,182 +0,0 @@ -#!/usr/bin/env node -// check-availability.mjs — is every source this manifest cites still fetchable? -// -// This is the one fact about a report video that goes stale in BOTH directions -// and that nothing on disk records. A source can be deleted between writing the -// manifest and building it (so a 40-minute build dies at clip 14 having paid for -// thirteen fetches), and a source can come back (so a manifest annotated "gone" -// stays wrong). Neither is visible until a build runs. -// -// So it runs first, it runs cheap, and it writes down when it ran. `--simulate` -// resolves formats without downloading a byte: a few seconds for a whole -// manifest against twenty-odd minutes for the build it protects. -// -// On the CLI: -// node scripts/report-to-video/check-availability.mjs <manifest.json> [--json] -// -// Options: -// --json Print the report as JSON instead of a table -// --out <dir> Output root (default: manifest dir + /out) -// --allow-missing Exit 0 even when a source is gone (report only) -// --max-age <days> Reuse a recorded verdict younger than this (default: 0) - -import { execFile } from "node:child_process"; -import { promisify } from "node:util"; -import { mkdir, readFile, writeFile } from "node:fs/promises"; -import path from "node:path"; - -import { DEFAULT_CHANNELS_DIR } from "./cues.mjs"; - -const execFileP = promisify(execFile); - -const YTDLP = process.env.YTDLP_BIN ?? "yt-dlp"; -// Resolved relative to the repo (see cues.mjs) rather than an absolute path in -// one machine's home directory, which every other clone would miss. -const CHANNELS_DIR = DEFAULT_CHANNELS_DIR; - -// yt-dlp says why in prose, and the distinction matters editorially: a private -// or removed video needs the clip converting to a quote card, while a network -// blip needs a retry. Anything unrecognised stays `maybe_missing` rather than -// being called deleted -- claiming a source is gone when it is not is the more -// expensive mistake, because it invites deleting a citation. -function classify(stderr) { - const t = String(stderr ?? ""); - if (/Private video|private/i.test(t)) return "private"; - if (/removed by the uploader|has been removed|no longer available|Video unavailable|does not exist|410/i.test(t)) - return "deleted"; - if (/age.?restrict|Sign in to confirm|confirm your age/i.test(t)) return "restricted"; - if (/members-only|join this channel/i.test(t)) return "members-only"; - if (/geo|not available in your country/i.test(t)) return "geo-blocked"; - return "maybe_missing"; -} - -async function cueMeta(videoId, channelSlug) { - const p = path.join(CHANNELS_DIR, channelSlug, "data", videoId, "transcript.cues.json"); - const d = JSON.parse(await readFile(p, "utf8")); - return { title: d.title, webpageUrl: d.webpageUrl, duration: d.duration }; -} - -export async function checkAvailability(manifestPath, { outDir, maxAgeDays = 0 } = {}) { - const manifest = JSON.parse(await readFile(manifestPath, "utf8")); - const slug = manifest.provenance?.channelSlug; - const dir = outDir ?? path.join(path.dirname(path.resolve(manifestPath)), "out"); - const file = path.join(dir, "availability.json"); - - // A clip may name its own channel: the same streamer's VODs are mirrored - // across more than one archive, and the same id under a different slug is a - // different file. So the unit of work is (channel, video), never video alone. - const wanted = new Map(); - const want = (channel, video) => { - const key = `${channel}/${video}`; - if (!wanted.has(key)) wanted.set(key, { key, channel, video, clips: [], claims: [] }); - return wanted.get(key); - }; - for (const e of manifest.timeline ?? []) { - if (e.type !== "clip") continue; - want(e.channel ?? slug, e.video).clips.push(e.id); - } - // LEDGER SOURCES TOO, not just the clipped ones. A cut that stacks its - // unclipped claims onto cards has to say WHY each one is a line of text - // rather than footage, and "the upload is gone" and "we did not cut it" are - // different sentences. Guessing which is which is how a live source ends up - // labelled deleted on screen. - for (const c of manifest.ledger ?? []) { - if (!c.video) continue; - want(c.channel ?? slug, c.video).claims.push(c.id); - } - - const prev = await readFile(file, "utf8").then( - (s) => JSON.parse(s), - () => ({ sources: [] }), - ); - const prevBy = new Map((prev.sources ?? []).map((s) => [s.key, s])); - const freshMs = maxAgeDays * 86400_000; - - const sources = []; - for (const w of wanted.values()) { - const was = prevBy.get(w.key); - if (freshMs > 0 && was?.checkedAt && Date.now() - Date.parse(was.checkedAt) < freshMs) { - sources.push({ ...was, clips: w.clips, claims: w.claims, reused: true }); - continue; - } - - let meta; - try { - meta = await cueMeta(w.video, w.channel); - } catch { - // No cue file is a DIFFERENT failure from a dead source, and it is the one - // the Rumble two-ids trap produces: the manifest names the MCP video id - // while the cues live under the URL slug. Build would die here too, so it - // is reported here rather than discovered twenty minutes in. - sources.push({ - ...w, ok: false, state: "no-cues", checkedAt: new Date().toISOString(), - error: `no transcript.cues.json under ${w.channel}/data/${w.video}`, - }); - continue; - } - - try { - await execFileP( - YTDLP, - ["--ignore-config", "--no-playlist", "--simulate", "--quiet", "--no-warnings", "--", meta.webpageUrl], - { maxBuffer: 1 << 24 }, - ); - sources.push({ - ...w, ok: true, state: "ok", title: meta.title, url: meta.webpageUrl, - checkedAt: new Date().toISOString(), error: null, - }); - } catch (err) { - const stderr = err?.stderr ?? err?.message ?? ""; - sources.push({ - ...w, ok: false, state: classify(stderr), title: meta.title, url: meta.webpageUrl, - checkedAt: new Date().toISOString(), error: String(stderr).trim().split("\n").slice(-3).join(" "), - }); - } - } - - const report = { manifest: path.resolve(manifestPath), checkedAt: new Date().toISOString(), sources }; - await mkdir(dir, { recursive: true }); - await writeFile(file, JSON.stringify(report, null, 2) + "\n", "utf8"); - return { ...report, file }; -} - -async function main() { - const argv = process.argv.slice(2); - const manifestPath = argv.find((a) => !a.startsWith("--")); - if (!manifestPath) { - console.error("usage: check-availability.mjs <manifest.json> [--json] [--out <dir>] [--allow-missing]"); - process.exit(2); - } - const flag = (n) => { - const i = argv.indexOf(n); - return i >= 0 ? argv[i + 1] : undefined; - }; - const report = await checkAvailability(manifestPath, { - outDir: flag("--out"), - maxAgeDays: Number(flag("--max-age") ?? 0), - }); - - if (argv.includes("--json")) { - console.log(JSON.stringify(report, null, 2)); - } else { - for (const s of report.sources) { - const mark = s.ok ? "ok " : "GONE"; - console.log( - `${mark} ${s.key.padEnd(40)} ${String(s.state).padEnd(14)} ` + - `${s.clips.length} clip(s)${s.reused ? " (cached)" : ""}`, - ); - if (!s.ok && s.error) console.log(` ${s.error}`); - } - const bad = report.sources.filter((s) => !s.ok).length; - console.log(`\n${report.sources.length} source(s), ${bad} unavailable -> ${report.file}`); - } - - if (report.sources.some((s) => !s.ok) && !argv.includes("--allow-missing")) process.exit(1); -} - -if (import.meta.url === `file://${process.argv[1]}`) { - main().catch((err) => { - console.error(err.message ?? err); - process.exit(1); - }); -} diff --git a/scripts/report-to-video/cues.mjs b/scripts/report-to-video/cues.mjs @@ -1,270 +0,0 @@ -// Where a clip's caption cues come from. -// -// A clip window is widened from a cue span to a whole sentence, which needs cue -// END times. Nothing else in the pipeline carries them: a report citation is a -// single start second, and the MCP `Snippet` type has no `end` field. So this is -// the one place that answers "what are the real cue boundaries for this video". -// -// TWO SOURCES, SAME SHAPE. A local corpus stores each video as -// `<CHANNELS_DIR>/<slug>/data/<id>/transcript.cues.json`, and a *published* -// archive serves the same record inside a paginated shard. The two carry the -// same fields — `{ slug, id, channelSlug, title, uploadDate, duration, channel, -// description, platform, webpageUrl, cues: [{start, end, text}] }` — so one -// resolver serves both `loadCues` and `videoMeta`, and a caller cannot tell -// which it got beyond the `from` marker. -// -// That parity is what makes a corpus optional. Clone the repo, point a manifest -// at a public instance, and the video pipeline can cut clips without mirroring a -// single channel: the cue windows come over HTTP, and the media itself was -// always a network fetch (`yt-dlp --download-sections`). -// -// The shard walk is the contract published at `/corpus.json` under `shardScheme`: -// 1. GET <origin>/corpus.json -> channels[].manifests.transcripts -// 2. GET that manifest -> { pageCount, slugToPage: { <id>: N } } -// 3. GET page-<NNNN>.json (N zero-padded to 4) -> array of records -// 4. take the record whose `id` matches -// -// Local wins when present: it is faster, works offline, and is the operator's own -// data. HTTP is the fallback, not a preference. -// -// THE TWO SOURCES CAN DISAGREE, AND IT IS NOT ROUNDING. A published archive is a -// snapshot; a live corpus keeps moving. Re-synced platform captions, an -// auto-caption replacement or a re-transcription all rewrite a video's cues in -// place, and the archive keeps the text it was built from until it is rebuilt. -// Measured on this corpus (local 2026-08-13 against a 2026-08-07 publish): of -// four videos checked, three were byte-identical and one had 65 of its 84 cue -// texts changed with timings shifted by up to **2.24 s** — enough to cut a clip -// in the wrong place. -// -// So `prefer` is a real decision, not a micro-optimisation: -// "auto" (default) local when present, else HTTP. Right for an operator. -// "local" never fall back. Fail loudly instead of silently cutting from -// different cues than the ones a window was authored against. -// "http" always the archive. Right when you want the windows to match what a -// reader following the citation will actually see, and the only -// option that is reproducible on a machine with no corpus. -// Whatever answers, the returned record carries `from` so a caller can record it. - -import { readFile, writeFile, mkdir } from "node:fs/promises"; -import path from "node:path"; -import os from "node:os"; -import { createHash } from "node:crypto"; -import { fileURLToPath } from "node:url"; - -// This file lives at <repo>/scripts/report-to-video/, so the corpus a plain -// checkout would have is two levels up. Previously this defaulted to an absolute -// path inside the original author's home directory, which meant every other -// clone silently looked in a directory that does not exist. -const REPO_ROOT = path.resolve(path.dirname(fileURLToPath(import.meta.url)), "..", ".."); - -export const DEFAULT_CHANNELS_DIR = - process.env.CHANNELS_DIR ?? path.join(REPO_ROOT, "transcripts", "channels"); - -const DEFAULT_CACHE_DIR = - process.env.REPORT_CACHE_DIR ?? - path.join(os.homedir(), ".cache", "archilyzer-report-to-video"); - -// A shard page is capped at 8 MB and holds ~100 videos, so refetching one per -// clip — across two separate processes, resolve-windows then build-video — is -// the difference between usable and painful. Cached by URL on disk; archives are -// rebuilt rarely and a stale page only matters if the cues themselves changed. -function cacheKey(url) { - return createHash("sha1").update(url).digest("hex") + ".json"; -} - -export function pageFileName(pageNumber) { - return `page-${String(pageNumber).padStart(4, "0")}.json`; -} - -export function pageUrlFrom(manifestUrl, pageNumber) { - const u = new URL(manifestUrl); - u.pathname = u.pathname.replace(/[^/]+$/, pageFileName(pageNumber)); - return u.toString(); -} - -// The origin of the archive a manifest was built against. Every manifest already -// records this — `siteOrigin` explicitly, and `corpus` as `remote:<url>` or -// `local:<path>` — so the common case needs no configuration at all. -export function siteOriginFromManifest(manifest) { - const p = manifest?.provenance ?? {}; - if (typeof p.siteOrigin === "string" && p.siteOrigin.trim()) { - return p.siteOrigin.replace(/\/+$/, ""); - } - if (typeof p.corpus === "string" && p.corpus.startsWith("remote:")) { - return p.corpus.slice("remote:".length).replace(/\/+$/, ""); - } - if (typeof p.shareLink === "string" && /^https?:/.test(p.shareLink)) { - try { - return new URL(p.shareLink).origin; - } catch { - /* fall through */ - } - } - return null; -} - -export class CueLookupError extends Error { - constructor(message, { channelSlug, videoId, tried }) { - super(message); - this.name = "CueLookupError"; - this.channelSlug = channelSlug; - this.videoId = videoId; - this.tried = tried; - } -} - -export function createCueSource({ - channelsDir = DEFAULT_CHANNELS_DIR, - siteOrigin = null, - cacheDir = DEFAULT_CACHE_DIR, - fetchImpl = globalThis.fetch, - log = () => {}, - // "auto" | "local" | "http" — see the note on divergence above. - prefer = "auto", - // Opt-in, because it is expensive: see resolveSiteId below. - resolveSiteIds = false, -} = {}) { - const mem = new Map(); - - async function getJson(url) { - if (mem.has(url)) return mem.get(url); - const disk = cacheDir ? path.join(cacheDir, cacheKey(url)) : null; - if (disk) { - try { - const cached = JSON.parse(await readFile(disk, "utf8")); - mem.set(url, cached); - return cached; - } catch { - /* cold cache */ - } - } - log(`fetch ${url}`); - const res = await fetchImpl(url); - if (!res.ok) throw new Error(`GET ${url} -> ${res.status}`); - const json = await res.json(); - mem.set(url, json); - if (disk) { - try { - await mkdir(path.dirname(disk), { recursive: true }); - await writeFile(disk, JSON.stringify(json)); - } catch { - // A cache we cannot write is a slow run, not a failed one. - } - } - return json; - } - - async function channelEntry(origin, channelSlug) { - const corpus = await getJson(`${origin}/corpus.json`); - const found = (corpus.channels ?? []).find((c) => c.slug === channelSlug); - if (!found) { - throw new CueLookupError( - `channel "${channelSlug}" is not in ${origin}/corpus.json`, - { channelSlug, videoId: null, tried: [`${origin}/corpus.json`] }, - ); - } - return found; - } - - // Map a LOCAL video id onto the id the site serves, by scanning the channel's - // pages for a record whose `webpageUrl` contains it. - // - // WHY THIS EXISTS: a Rumble video has two ids. The site (and the MCP) key it by - // the EMBED id; the local cue directory is named for the URL SLUG. A manifest - // hand-authored against local cue dirs therefore carries slugs that are absent - // from the published `slugToPage` — every Rumble clip misses. - // - // WHY IT IS OPT-IN: it downloads a channel's shards until it hits a match, and - // a shard is up to 8 MB. That is a reasonable price to pay knowingly and a - // terrible one to pay silently, so the direct lookup fails with instructions - // instead and this runs only when asked. - async function resolveSiteId(origin, entry, wanted) { - const manifest = await getJson(entry.manifests.transcripts); - log(`resolving "${wanted}" by scanning ${manifest.pageCount} shard(s) of ${entry.slug}`); - for (let n = 0; n < manifest.pageCount; n += 1) { - const page = await getJson(pageUrlFrom(entry.manifests.transcripts, n)); - const hit = page.find( - (r) => r.id === wanted || r.slug === wanted || String(r.webpageUrl ?? "").includes(wanted), - ); - if (hit) return hit; - } - return null; - } - - async function fromHttp(channelSlug, videoId, hints) { - const origin = hints.siteOrigin ?? siteOrigin; - if (!origin) { - throw new CueLookupError( - `no local cues for ${channelSlug}/${videoId} and no archive origin to fetch them from ` + - `(set provenance.siteOrigin in the manifest, or pass --site-origin / SITE_ORIGIN)`, - { channelSlug, videoId, tried: ["local"] }, - ); - } - const siteChannel = hints.siteChannel ?? channelSlug; - const siteVideo = hints.siteVideo ?? videoId; - const entry = await channelEntry(origin, siteChannel); - const manifest = await getJson(entry.manifests.transcripts); - const pageNumber = manifest.slugToPage?.[siteVideo]; - - if (pageNumber === undefined) { - if (resolveSiteIds) { - const hit = await resolveSiteId(origin, entry, siteVideo); - if (hit) return { ...hit, from: "http" }; - } - throw new CueLookupError( - `"${siteVideo}" is not in ${siteChannel}'s published slugToPage on ${origin}.\n` + - ` If this is a Rumble clip, the archive is keyed by the EMBED id while a local cue\n` + - ` directory is named for the URL SLUG — they differ. Either add "siteVideo" (and\n` + - ` "siteChannel" if it also differs) to this clip in the manifest, or re-run with\n` + - ` --resolve-site-ids to find it by scanning the channel's shards (slow: downloads\n` + - ` up to 8 MB per shard until it matches).\n` + - ` Note that a clip's citeUrl is NOT usable here — it may deliberately cite a\n` + - ` different recording (a mirror that reads better), whose clock is not the same.`, - { channelSlug, videoId, tried: [entry.manifests.transcripts] }, - ); - } - - const page = await getJson(pageUrlFrom(entry.manifests.transcripts, pageNumber)); - const record = page.find((r) => r.id === siteVideo || r.slug === siteVideo); - if (!record) { - throw new CueLookupError( - `${siteChannel}/${siteVideo} is on shard ${pageNumber} per the manifest, but no record ` + - `there has that id — the published archive is inconsistent`, - { channelSlug, videoId, tried: [pageUrlFrom(entry.manifests.transcripts, pageNumber)] }, - ); - } - return { ...record, from: "http" }; - } - - async function fromLocal(channelSlug, videoId) { - const p = path.join(channelsDir, channelSlug, "data", videoId, "transcript.cues.json"); - const parsed = JSON.parse(await readFile(p, "utf8")); - return { ...parsed, from: "local" }; - } - - return { - channelsDir, - /** - * The full record for one video: cues plus the metadata build-video needs. - * `hints` may carry `siteChannel` / `siteVideo` (when the published archive - * keys this recording differently) and `siteOrigin` (per-manifest override). - */ - prefer, - async load(channelSlug, videoId, hints = {}) { - if (prefer === "http") return await fromHttp(channelSlug, videoId, hints); - try { - return await fromLocal(channelSlug, videoId); - } catch (err) { - if (err?.code !== "ENOENT" && err?.code !== "ENOTDIR") throw err; - if (prefer === "local") { - throw new CueLookupError( - `no local cues for ${channelSlug}/${videoId} under ${channelsDir}, and ` + - `--cue-source local forbids falling back to the archive`, - { channelSlug, videoId, tried: [channelsDir] }, - ); - } - return await fromHttp(channelSlug, videoId, hints); - } - }, - }; -} diff --git a/scripts/report-to-video/package.json b/scripts/report-to-video/package.json @@ -1,25 +0,0 @@ -{ - "name": "report-to-video", - "version": "0.1.0", - "private": true, - "type": "module", - "description": "Turn a cited sweep report into a narrated-by-text video.", - "bin": { - "report-build-video": "./build-video.mjs", - "report-resolve-windows": "./resolve-windows.mjs", - "report-check-availability": "./check-availability.mjs", - "report-verify-build": "./verify-build.mjs", - "report-compose-chrome": "./compose-chrome.mjs" - }, - "exports": { - "./build-video": "./build-video.mjs", - "./check-availability": "./check-availability.mjs", - "./compose-chrome": "./compose-chrome.mjs", - "./cues": "./cues.mjs", - "./ledger-totals": "./ledger-totals.mjs", - "./package.json": "./package.json", - "./render-cards": "./render-cards.mjs", - "./resolve-windows": "./resolve-windows.mjs", - "./verify-build": "./verify-build.mjs" - } -} diff --git a/scripts/report-to-video/render-cards.mjs b/scripts/report-to-video/render-cards.mjs @@ -1,1530 +0,0 @@ -#!/usr/bin/env node -// render-cards.mjs — turn a video manifest's `card` entries into PNG stills. -// -// One PNG per card, written to <outDir>/cards/<id>.png at the manifest's render -// resolution. Text is laid out by ImageMagick's Pango delegate rather than -// ffmpeg's drawtext: Pango wraps, kerns and takes inline markup, so a card is a -// single markup string instead of a stack of hand-positioned drawtext filters. -// -// Card styles (manifest `style` field): -// title — the opening card: big heading, subtitle, provenance footer -// chapter — an act break: small amber kicker over a large heading -// bullets — heading plus a list of caveats -// sources — closing attribution -// -// `status` is RETIRED. It existed to quote a claim whose source had gone, and -// the `ledger` entry type does that better: it says the same words, alongside -// the arithmetic the claim moves, and it says WHY there is no footage from a -// probe rather than from a hand-written kicker that nothing re-checks. -// -// In the app: not used. On the CLI: -// node scripts/report-to-video/render-cards.mjs <manifest.json> [--out <dir>] -// -// Options: -// --out <dir> Output root (default: the manifest's directory + /out) -// --only <id> Render just one card, by manifest id -// -// Requires: ImageMagick built with Pango (magick -list format | grep PANGO). - -import { execFile } from "node:child_process"; -import { promisify } from "node:util"; -import { mkdir, writeFile, readFile } from "node:fs/promises"; -import path from "node:path"; -import { dateKey, ledgerTotals, rosterLine } from "./ledger-totals.mjs"; - -const execFileP = promisify(execFile); - -const RSVG = process.env.RSVG_BIN ?? "rsvg-convert"; -const QRENCODE = process.env.QRENCODE_BIN ?? "qrencode"; - -// Every card is drawn at the manifest's full `width`, but once a rail column is -// configured the RIGHT `rail.width` pixels of the frame belong to it. Text, -// rules and timeline nodes therefore lay out inside the CONTENT width, while the -// canvas stays full-width — the card ground is flat `pal.bg`, which is exactly -// what the rail wants behind it. -export const railWidth = (render) => render.rail?.width ?? 0; -export const contentWidth = (render) => render.width - railWidth(render); - -/** - * How wide a card's content may be. - * - * `hideRail` cards slide the rail off over their own dissolve, so they get the - * WHOLE frame. Everything else lays out inside the content width and leaves the - * rail column as ground. - */ -export const cardWidth = (card, render) => - card?.hideRail ? render.width : contentWidth(render); - -// Pango markup is XML-ish, so anything we interpolate has to be escaped first. -// Curly quotes and the ellipsis pass through fine; only these five matter. -function esc(s) { - return String(s) - .replace(/&/g, "&amp;") - .replace(/</g, "&lt;") - .replace(/>/g, "&gt;") - .replace(/"/g, "&quot;") - .replace(/'/g, "&apos;"); -} - -function span(text, { size, color, weight, family = "Fira Sans" }) { - const attrs = [`font_family="${family}"`, `size="${Math.round(size * 1024)}"`]; - if (color) attrs.push(`foreground="${color}"`); - if (weight) attrs.push(`weight="${weight}"`); - return `<span ${attrs.join(" ")}>${text}</span>`; -} - -// Each style returns Pango markup for the whole card body. Blank lines are real -// newlines in the markup — Pango honours them, which is how vertical rhythm is -// set without positioning each run separately. -function markupFor(card, pal) { - const H = (t, size = 62) => - span(esc(t), { size, color: pal.fg, weight: "bold" }); - const KICKER = (t) => - span(esc(t.toUpperCase()), { size: 24, color: pal.amber, weight: "bold" }); - const SUB = (t, size = 30) => span(esc(t), { size, color: pal.muted }); - - switch (card.style) { - case "title": - return [ - span(esc(card.heading), { size: 82, color: pal.fg, weight: "bold" }), - "", - span(esc(card.sub), { size: 38, color: pal.accent }), - "", - "", - SUB(card.foot, 24), - ].join("\n"); - - case "chapter": - return [ - card.kicker ? KICKER(card.kicker) : null, - card.kicker ? "" : null, - H(card.heading), - card.sub ? "" : null, - card.sub ? SUB(card.sub, 32) : null, - ] - .filter((l) => l !== null) - .join("\n"); - - case "bullets": { - const items = (card.bullets ?? []).flatMap((b) => [ - `${span("— ", { size: 30, color: pal.accent, weight: "bold" })}${span( - esc(b), - { size: 30, color: pal.fg }, - )}`, - "", - ]); - return [H(card.heading, 52), "", ...items].join("\n"); - } - - case "sources": - return [ - H(card.heading, 52), - "", - card.sub ? SUB(card.sub, 32) : null, - card.foot ? "" : null, - card.foot - ? card.foot - .split("\n") - .map((l) => SUB(l, 24)) - .join("\n") - : null, - ] - .filter((l) => l !== null) - .join("\n"); - - default: - return H(card.heading ?? card.id); - } -} - -// A chapter break that just states a heading is dead air — it stops the video to -// say something the next clip is about to say anyway. This draws the whole -// project arc instead, with the current step lit and everything before it -// filled, so the pause carries information: where we are and how far is left. -// -// Nodes come from the manifest's top-level `timelineNodes`; the card names its -// position with `step` (0-based). -async function renderTimelineCard(card, render, nodes, outDir) { - const pal = render.palette; - const { width, height } = render; - const outPath = path.join(outDir, "cards", `${card.id}.png`); - const dir = path.join(outDir, "cards"); - - const VW = cardWidth(card, render); - const x0 = 260; - const x1 = VW - 260; - const axisY = Math.round(height * 0.56); - const gap = (x1 - x0) / (nodes.length - 1); - const xs = nodes.map((_, i) => Math.round(x0 + i * gap)); - const cur = card.step; - - const args = ["-size", `${width}x${height}`, `xc:${pal.bg}`, "-strokewidth", "4"]; - - // Track: filled up to the current node, dim beyond it. - args.push( - "-stroke", pal.muted, "-fill", "none", - "-draw", `line ${xs[0]},${axisY} ${xs[xs.length - 1]},${axisY}`, - ); - if (cur > 0) { - args.push("-stroke", pal.accent, "-draw", `line ${xs[0]},${axisY} ${xs[cur]},${axisY}`); - } - - // Nodes: past and present filled, future hollow. The current one is larger and - // amber so the eye lands on it without needing a label to say "you are here". - nodes.forEach((_, i) => { - const r = i === cur ? 19 : 11; - const color = i === cur ? pal.amber : i < cur ? pal.accent : pal.bg; - args.push( - "-stroke", i <= cur ? (i === cur ? pal.amber : pal.accent) : pal.muted, - "-fill", color, - "-draw", `circle ${xs[i]},${axisY} ${xs[i] + r},${axisY}`, - ); - }); - - args.push("-stroke", "none"); - - // Heading, centred over the whole card. - const headMarkup = [ - span(esc((card.kicker ?? nodes[cur].label).toUpperCase()), { - size: 26, color: pal.amber, weight: "bold", - }), - "", - span(esc(card.heading ?? nodes[cur].title), { size: 62, color: pal.fg, weight: "bold" }), - card.sub ? "" : null, - card.sub ? span(esc(card.sub), { size: 30, color: pal.muted }) : null, - ] - .filter((l) => l !== null) - .join("\n"); - - // Heading, left-aligned on the same margin the other card styles use. - const headPath = path.join(dir, `${card.id}.head.pango`); - await writeFile(headPath, headMarkup, "utf8"); - args.push( - "(", "-size", `${VW - 460}x`, "-background", "none", - "-define", `pango:width=${VW - 460}`, - `pango:@${headPath}`, ")", - "-gravity", "NorthWest", - "-geometry", `+${Math.round(VW * 0.09) + 58}+${Math.round(height * 0.19)}`, - "-composite", - ); - - // Per-node date labels, centred under their dot. ImageMagick's - // `pango:alignment` define does not actually centre the text inside the box, - // so measure each rendered label and place it by hand instead of trusting it. - for (let i = 0; i < nodes.length; i += 1) { - const isCur = i === cur; - const labMarkup = span(esc(nodes[i].label), { - size: isCur ? 24 : 21, - color: isCur ? pal.fg : i < cur ? pal.muted : "#5c5570", - weight: isCur ? "bold" : "normal", - }); - const labPath = path.join(dir, `${card.id}.n${i}.pango`); - const labPng = path.join(dir, `${card.id}.n${i}.png`); - await writeFile(labPath, labMarkup, "utf8"); - await execFileP("magick", ["-background", "none", `pango:@${labPath}`, labPng]); - const { stdout } = await execFileP("magick", ["identify", "-format", "%w", labPng]); - const w = Number(stdout.trim()); - args.push(labPng, "-geometry", `+${xs[i] - Math.round(w / 2)}+${axisY + 44}`, "-composite"); - } - - args.push(outPath); - await execFileP("magick", args, { maxBuffer: 1 << 24 }); - return outPath; -} - -// Footer chrome, drawn once and overlaid on every clip: a track, a dot per -// section and its label. The *progress* along it is not baked in here — the fill -// bar and the amber marker are drawn by ffmpeg at encode time so they can slide -// between sections instead of cutting. See buildClipSegment in build-video.mjs. -// -// Returns the geometry the encoder needs to place those moving parts. -export async function renderFooterAssets(render, nodes, outDir) { - const pal = render.palette; - const { width } = render; - const FH = render.footerHeight ?? 92; - const dir = path.join(outDir, "cards"); - - // A cut whose clips are not a progression through time has nothing for a - // timeline to say, and a footer drawn anyway is chrome that has not earned its - // place. No nodes (or an explicit zero height) means no footer at all — the - // caller letterboxes against the header alone. - if (!nodes?.length || FH === 0) { - return { - footer: null, marker: null, bar: null, trackLen: 0, - footerHeight: 0, trackY: 0, xs: [], x0: 0, markerRadius: 0, - }; - } - - // The band itself stays full-frame so it reads as one strip running under the - // rail; only the TRACK is pulled in to the content width. - const x0 = 200; - const x1 = contentWidth(render) - 200; - const trackY = 26; - const gap = (x1 - x0) / (nodes.length - 1); - const xs = nodes.map((_, i) => Math.round(x0 + i * gap)); - - const footer = path.join(dir, "_footer.png"); - const args = [ - "-size", `${width}x${FH}`, `xc:${pal.bg}`, - "-strokewidth", "3", - "-stroke", "#3a3450", "-fill", "none", - "-draw", `line ${xs[0]},${trackY} ${xs[xs.length - 1]},${trackY}`, - "-stroke", "none", - ]; - for (const x of xs) { - args.push("-fill", "#4a4363", "-draw", `circle ${x},${trackY} ${x + 6},${trackY}`); - } - - // Two lines per node: what happened, then when. Both are measured and placed - // by hand — ImageMagick's `pango:alignment` define does not actually centre - // text inside its box. - for (let i = 0; i < nodes.length; i += 1) { - const lines = [ - { text: nodes[i].label, size: 18, color: pal.fg, dy: 18 }, - { text: nodes[i].date, size: 16, color: pal.muted, dy: 42 }, - ]; - for (const [k, ln] of lines.entries()) { - const pPath = path.join(dir, `_footer.n${i}.l${k}.pango`); - const pPng = path.join(dir, `_footer.n${i}.l${k}.png`); - await writeFile(pPath, span(esc(ln.text), { size: ln.size, color: ln.color }), "utf8"); - await execFileP("magick", ["-background", "none", `pango:@${pPath}`, pPng]); - const { stdout } = await execFileP("magick", ["identify", "-format", "%w", pPng]); - args.push( - pPng, - "-geometry", `+${xs[i] - Math.round(Number(stdout.trim()) / 2)}+${trackY + ln.dy}`, - "-composite", - ); - } - } - args.push(footer); - await execFileP("magick", args, { maxBuffer: 1 << 24 }); - - // The fill bar, as a strip to be TRANSLATED under a fixed crop rather than a - // drawbox whose width depends on `t`. drawbox has no time variable — its `t` - // is the box thickness — so the width expression the encoder used to build - // never evaluated and the bar was always full. Left half accent, right half - // transparent: sliding the crop window left across it grows the accent run. - const trackLen = xs[xs.length - 1] - x0; - const bar = path.join(dir, "_bar.png"); - await execFileP("magick", [ - "-size", `${trackLen * 2}x3`, "xc:none", - "-fill", pal.accent, "-stroke", "none", - "-draw", `rectangle 0,0 ${trackLen - 1},2`, - bar, - ]); - - // The travelling marker. - const marker = path.join(dir, "_marker.png"); - const r = 11; - await execFileP("magick", [ - "-size", `${r * 2 + 2}x${r * 2 + 2}`, "xc:none", - "-fill", pal.amber, "-stroke", "none", - "-draw", `circle ${r + 1},${r + 1} ${r * 2 + 1},${r + 1}`, - marker, - ]); - - return { footer, marker, bar, trackLen, footerHeight: FH, trackY, xs, x0, markerRadius: r }; -} - -// =========================================================================== -// The claim rail -// =========================================================================== -// A vertical ledger down the right edge that appends one row per claim as the -// video runs, with a live per-company tally beside it. The dates and the numbers -// are SPOKEN in the clips and shown only in the header citation line, so a viewer -// can hear "nearly ten" three times without ever seeing that the three refer to -// three different companies. The rail is what makes that visible. -// -// Every asset here is a STRIP, not a per-state still: one tall PNG whose window -// ffmpeg slides with an animated `crop`. That is the whole trick — swapping -// stills can only cut, but a crop can ease, and one input per moving part keeps -// the filtergraph small enough that ffmpeg does not deadlock on chained -// `-loop 1` inputs. See railFilterChain in build-video.mjs for the ramps. -// -// Assets are authored as SVG and rasterized with rsvg-convert rather than drawn -// with ImageMagick primitives: the strips need right-aligned columns, hairlines -// and ~500 individually-placed text runs, and one rsvg call beats four hundred -// `magick` invocations. Fonts inside the SVG resolve through fontconfig, so the -// family name has to MATCH the Pango cards ("Fira Sans"), not the font FILE that -// render.fontRegular points at. - -const RAIL_FONT = "Fira Sans"; - -// A tiny SVG text run. Everything is placed absolutely — no flow, no wrapping. -function svgText(x, y, text, o = {}) { - const a = [ - `x="${x}"`, `y="${y}"`, - `font-family="${o.family ?? RAIL_FONT}"`, - `font-size="${o.size ?? 15}"`, - `fill="${o.color}"`, - ]; - if (o.weight) a.push(`font-weight="${o.weight}"`); - if (o.anchor) a.push(`text-anchor="${o.anchor}"`); - if (o.ls) a.push(`letter-spacing="${o.ls}"`); - if (o.opacity != null) a.push(`opacity="${o.opacity}"`); - return `<text ${a.join(" ")}>${esc(text)}</text>`; -} - -const svgDoc = (w, h, body) => - `<svg xmlns="http://www.w3.org/2000/svg" width="${w}" height="${h}" ` + - `viewBox="0 0 ${w} ${h}">${body}</svg>`; - -// rsvg-convert is deterministic about output size in a way ImageMagick's RSVG -// delegate is not (its -density is ignored for sizing in some builds), so the -// pixel dimensions ffmpeg's crop arithmetic depends on are guaranteed here. -async function rasterize(svg, svgPath, pngPath, w, h) { - await writeFile(svgPath, svg, "utf8"); - await execFileP(RSVG, ["-w", String(w), "-h", String(h), "-o", pngPath, svgPath], { - maxBuffer: 1 << 26, - }); - return pngPath; -} - -// Truncate to a pixel budget. Fira Sans at these sizes averages ~0.50em per -// character; a hair conservative is right, because an overflowing row would run -// under the value column rather than wrap. -function fit(text, size, maxPx) { - const max = Math.max(4, Math.floor(maxPx / (size * 0.5))); - const t = String(text ?? ""); - return t.length <= max ? t : `${t.slice(0, max - 1).trimEnd()}…`; -} - -/** - * Every pixel measurement the rail needs, derived once so the asset builder and - * the filtergraph builder cannot disagree about a single one of them. - * - * The log window is deliberately sized to a WHOLE number of rows and butted - * against the bottom of the rail column: that is what lets the curtain (below) - * park exactly one window-height down and end up outside the rail entirely. - */ -export function railGeometry(render, nClaims) { - const rail = render.rail ?? {}; - const RW = rail.width ?? 500; - const HH = render.headerHeight ?? 56; - // reservedFooterHeight, NOT render.footerHeight. The two differ by 100px the - // moment the chart band is on (the band takes the footer's ground and 100 - // more), and deriving the rail's height from the smaller one ran the column - // a hundred pixels past the band's own top edge -- about two rows of the log - // window, drawn below the line everything else letterboxes to. - const FH = reservedFooterHeight(render); - const ROWH = rail.rowHeight ?? 38; - const PAD = rail.pad ?? 22; - const TALLYROWH = rail.tallyRowHeight ?? 40; - const nTracks = (rail.tracks ?? []).length; - - const RHGT = render.height - HH - FH; - // The tally's header belongs to the CHROME, not to the rolling block. Put it - // in the block and every roll scrolls a duplicate copy of it up through the - // window, which reads as noise rather than as a counter changing. - const TALLYHEAD_REL = rail.tallyTop ?? 76; - const TALLYTOP_REL = TALLYHEAD_REL + 26; - const TALLYH = nTracks * TALLYROWH; - - // The rolling CELL: the only part of a tally row that ever changes. The - // swatch and the company name sit left of it and belong to the chrome, so - // that a number moving does not drag its own label up the screen with it. - const CELLW = rail.cellWidth ?? 170; - const CELLX = RW - PAD - CELLW; - - // The roster he last enumerated. One line, and the point of it is that it - // stands still while the numbers above it move. - const ROSTERH = rail.rosterHeight ?? 26; - const ROSTERTOP_REL = TALLYTOP_REL + TALLYH + 14; - const ROSTERX = PAD + (rail.rosterLabelWidth ?? 62); - const ROSTERW = RW - PAD - ROSTERX; - - // The provenance tile, parked at the foot of the column: one QR per clip, - // bordered so it reads as a link rather than as decoration. - const QRSIZE = rail.qrSize ?? 132; - const TILEW = RW - 2 * PAD; - const TILEH = QRSIZE + 22; - const TILETOP_REL = RHGT - (rail.qrBottom ?? 12) - TILEH; - - const LOGTOP_REL = ROSTERTOP_REL + ROSTERH + 34; - // The log window is a WHOLE number of rows and stops short of the tile, so - // the curtain parks exactly one window-height down and the QR overlay (which - // comes after it in the chain) is never painted over. - const K = Math.max(1, Math.floor((TILETOP_REL - 12 - LOGTOP_REL) / ROWH)); - const LOGH = K * ROWH; - - return { - RW, PAD, ROWH, TALLYROWH, K, LOGH, TALLYH, RHGT, - CELLW, CELLX, - ROSTERH, ROSTERTOP_REL, ROSTERTOP: HH + ROSTERTOP_REL, ROSTERX, ROSTERW, - QRSIZE, TILEW, TILEH, TILETOP_REL, TILETOP: HH + TILETOP_REL, - VW: render.width - RW, - RX: render.width - RW, - RTOP: HH, - LOGTOP_REL, LOGTOP: HH + LOGTOP_REL, - TALLYTOP_REL, TALLYTOP: HH + TALLYTOP_REL, TALLYHEAD_REL, - nClaims, - logStripH: Math.max(LOGH, nClaims * ROWH), - }; -} - -/** - * How much of the frame the chrome band owns at the bottom. - * - * ONE definition, because three different renderers need it and they were - * already disagreeing: the closing chart drew its footnotes into the bottom - * 100px and the ledger scroll sized its window to `height - header - 100`, both - * of which are wrong the moment the band takes 200. The symptom is a card that - * looks finished in isolation and has its last two lines sitting under the - * chart in the cut. - */ -export function reservedFooterHeight(render) { - return render.chromeEngine === "hyperframes" - ? (render.chart?.height ?? 200) - : (render.footerHeight ?? 100); -} - -/** - * Every roll each tally cell will perform, as rows of one shared strip. - * - * --------------------------------------------------------------------------- - * Why the strip's LAYOUT carries the direction - * --------------------------------------------------------------------------- - * The whole block used to slide as one slab: when the coffee company's number - * changed, all four rows moved. Text that has not changed must not move, so - * each track now gets its own cell and its own y expression. - * - * A crop window can only walk a strip, and it walks in whichever direction its - * y expression takes it. So the DIRECTION of a roll is decided when the rows - * are laid out, not when the ramp is written: - * - * rise rows [old, new] crop walks DOWN the strip, content moves UP - * fall rows [new, old] crop walks UP the strip, content moves DOWN - * - * Between two transitions the crop steps instantly to the next pair's starting - * row. That step is invisible because the row it leaves and the row it arrives - * at hold IDENTICAL content — which is the reason every pair repeats the value - * it starts from rather than sharing a row with its neighbour. - * - * The delta chip rides along on both rows of the pair, and therefore stays on - * screen until the next change. That is deliberate: it reads as "how this - * number last moved", and blanking it at the step would make the invisible - * reposition visible. - * - * A repeated identical figure still rolls, upward. He said it again on a new - * date, and the `as of` line underneath is what changed. - * - * @returns {{lanes: Array<object>, rows: number}} one lane per track plus the - * roster lane, each with `rows` (the cells to draw) and `steps` (per claim, - * `null` or `{a, b}` — the row the roll starts on and the row it ends on). - */ -export function tallyTracks(ledger, tracks, opts = {}) { - const rosterLineOf = opts.rosterLine ?? (() => null); - const lanes = tracks.map((tr) => ({ - key: tr.key, track: tr, kind: "tally", - rows: [{ empty: true, track: tr }], - steps: new Array(ledger.length).fill(null), - cur: null, - })); - const byKey = Object.fromEntries(lanes.map((l) => [l.key, l])); - - const roster = { - key: "__roster", kind: "roster", - rows: [{ empty: true }], - steps: new Array(ledger.length).fill(null), - cur: null, - }; - - /** Lay a transition down as a pair of rows and record where it starts/ends. */ - const transition = (lane, from, to, rise) => { - const p = lane.rows.length; - if (rise) { - lane.rows.push(from, to); - return { a: p, b: p + 1 }; - } - lane.rows.push(to, from); - return { a: p + 1, b: p }; - }; - - ledger.forEach((c, i) => { - // ---- the four company cells ---- - const lane = byKey[c.scope ?? c.company]; - // UTTERED only. The tally says "the latest figure he has given", and a sum - // we performed is not one — showing 18 here while a card beside it says he - // never said eighteen makes the video contradict itself on screen. - if (lane && c.value != null && (!c.valueKind || c.valueKind === "uttered")) { - const prev = lane.cur; - const next = { - track: lane.track, - display: c.display ?? String(c.value), - value: c.value, - date: c.date, - population: c.population ?? null, - delta: prev ? Number((c.value - prev.value).toFixed(2)) : null, - }; - const from = prev ? { ...prev, delta: prev.delta } : { empty: true, track: lane.track }; - lane.steps[i] = transition(lane, from, next, !prev || c.value >= prev.value); - lane.cur = next; - } - - // ---- the roster ---- - // Only when the LINE changes. He enumerates the same two editors and one - // designer in October and again in December; rolling the line to arrive at - // the words it already said would animate the one thing that held still. - const line = Array.isArray(c.roles) && c.roles.length ? rosterLineOf(c.roles) : null; - if (line && line !== roster.cur?.line) { - const next = { line, date: c.date }; - const from = roster.cur ? { ...roster.cur } : { empty: true }; - roster.steps[i] = transition(roster, from, next, true); - roster.cur = next; - } - }); - - const all = [...lanes, roster]; - return { lanes: all, rows: Math.max(...all.map((l) => l.rows.length)) }; -} - -/** - * The colour of the population word under a figure. Never a new hue -- the - * chip is `pal.muted` text, so the dataviz gate does not have to be re-run. - */ -const POP_WORD = { - employees: "employees", - "full-time": "full time", - salaried: "salaried", - contractor: "contractors", - 1099: "1099", - people: "people", -}; - -/** A small solid triangle, because a font may not carry ▲ and tofu is worse. */ -function svgTri(x, y, up, color) { - const d = up - ? `M${x},${y + 7} L${x + 4.5},${y} L${x + 9},${y + 7} z` - : `M${x},${y} L${x + 4.5},${y + 7} L${x + 9},${y} z`; - return `<path d="${d}" fill="${color}"/>`; -} - -/** One rolling tally cell, drawn into a CELLW x TALLYROWH box at (x, y). */ -function tallyCellSvg(cell, x, y, g, pal, h) { - const right = x + g.CELLW; - const out = [`<rect x="${x}" y="${y}" width="${g.CELLW}" height="${h}" fill="${pal.bg}"/>`]; - if (cell.empty) { - out.push( - svgText(right, y + 24, "—", { size: 22, color: pal.muted, weight: "bold", anchor: "end", opacity: 0.5 }), - svgText(right, y + 37, "not yet stated", { size: 10.5, color: pal.muted, opacity: 0.6, anchor: "end" }), - ); - return out.join(""); - } - const colour = cell.track?.color ?? pal.fg; - out.push( - svgText(right, y + 24, cell.display, { size: 22, color: colour, weight: "bold", anchor: "end" }), - ); - if (cell.delta != null && cell.delta !== 0) { - const up = cell.delta > 0; - out.push( - svgTri(x, y + 11, up, colour), - svgText(x + 14, y + 22, `${up ? "+" : "−"}${Math.abs(cell.delta)}`, { - size: 13, color: colour, weight: "bold", - }), - ); - } - const chip = cell.population ? ` · ${POP_WORD[cell.population] ?? cell.population}` : ""; - out.push( - svgText(right, y + 37, `as of ${cell.date}${chip}`, { - size: 10.5, color: pal.muted, anchor: "end", opacity: 0.9, - }), - ); - return out.join(""); -} - -/** One roster line, drawn into a ROSTERW x ROSTERH box. */ -function rosterCellSvg(cell, x, y, g, pal) { - const out = [`<rect x="${x}" y="${y}" width="${g.ROSTERW}" height="${g.ROSTERH}" fill="${pal.bg}"/>`]; - out.push( - cell.empty - ? svgText(x, y + 18, "not yet enumerated", { size: 12.5, color: pal.muted, opacity: 0.55 }) - : svgText(x, y + 18, fit(cell.line, 13, g.ROSTERW), { size: 13, color: pal.muted }), - ); - return out.join(""); -} - -// --------------------------------------------------------------------------- -// The provenance tile -// --------------------------------------------------------------------------- -// A compilation asks the viewer to take the edit on trust. The QR is the -// antidote: it resolves to this clip's exact start in the archive's own viewer, -// so anyone can pull up the surrounding hour and check the cut is fair. -// -// It used to float over the bottom-right of the PICTURE, which is the one place -// in the frame the cut promises never to draw on. In the rail's foot it is a -// bordered tile that reads as a link, and it becomes one more strip walked by a -// crop -- one tile per segment, stepped instantaneously at the mid-dissolve, -// exactly like the other four. -// -// Two rules learned the hard way: it must be FULLY OPAQUE (a translucent QR -// will not scan) and it must keep its quiet zone (the white border is part of -// the symbol, not decoration). -export function qrUrlFor(entry, provenance) { - if (entry.type === "clip") { - // A mirror's LOCAL slug is not the id the site serves, so an explicit - // per-clip citeUrl always wins over the derived one. - return ( - entry.citeUrl ?? - `${provenance.siteOrigin}/?v=${encodeURIComponent( - `${entry.channel ?? provenance.channelSlug}/${entry.video}`, - )}&t=${Math.floor(entry.cite ?? entry.start)}` - ); - } - // A card is not a moment, so it gets the search rather than a timestamp. - // - // NOT `provenance.shareLink`. That link carries all 23 channel filters and is - // ~1.4k characters — a version-40 symbol, 177 modules inside a 132 px tile, - // which is roughly 0.7 px per module and unscannable. `qrLink` is the same - // query without the channel list (~200 chars, 63 modules, verified scannable - // at this size); the site origin is the fallback. A code nobody can scan is - // worse than a short one. - return provenance.qrLink ?? provenance.siteOrigin ?? ""; -} - -async function qrTileStrip(entries, provenance, render, g, outDir) { - const pal = render.palette; - const dir = path.join(outDir, "cards"); - const qrDir = path.join(outDir, "qr"); - const q = render.qr ?? {}; - const QR = g.QRSIZE; - const TH = g.TILEH; - - // One PNG per DISTINCT url, then a tile per segment referencing it. - const seen = new Map(); - const urls = entries.map((e) => qrUrlFor(e, provenance)); - for (const url of urls) { - if (seen.has(url)) continue; - const png = path.join(qrDir, `q${seen.size.toString().padStart(2, "0")}.png`); - await execFileP(QRENCODE, [ - "-o", png, "-s", String(q.scale ?? 4), "-m", String(q.quiet ?? 3), - "-l", q.ecc ?? "M", url, - ]); - // Nearest-neighbour to an exact box: a resampled QR blurs its module edges - // and stops scanning, and the geometry has to be known before this runs. - const sized = path.join(qrDir, `q${seen.size.toString().padStart(2, "0")}.${QR}.png`); - await execFileP("magick", [png, "-filter", "point", "-resize", `${QR}x${QR}!`, sized]); - seen.set(url, sized); - } - - const stripH = entries.length * TH; - const body = [`<rect x="0" y="0" width="${g.TILEW}" height="${stripH}" fill="${pal.bg}"/>`]; - const images = []; - entries.forEach((e, i) => { - const y = i * TH; - const isClip = e.type === "clip"; - body.push( - `<rect x="0.5" y="${y + 0.5}" width="${g.TILEW - 1}" height="${TH - 1}" fill="${pal.bg}" ` + - `stroke="${pal.accent}" stroke-width="1"/>`, - svgText(14, y + 32, "SCAN → JERALYZER", { - size: 12.5, color: pal.accent, weight: "bold", ls: 1.3, - }), - svgText(14, y + 58, isClip ? "this exact moment," : "the sweep this is cut from,", { - size: 13.5, color: pal.fg, - }), - svgText(14, y + 78, isClip ? "in the archive's own viewer" : "every claim, searchable", { - size: 13.5, color: pal.fg, - }), - svgText(14, y + 106, "the archive outlives the platform", { - size: 11.5, color: pal.muted, opacity: 0.8, - }), - ); - images.push({ png: seen.get(urls[i]), x: g.TILEW - QR - 11, y: y + 11 }); - }); - - const svgPath = path.join(dir, "_rail_qr.svg"); - const basePng = path.join(dir, "_rail_qr.base.png"); - await rasterize(svgDoc(g.TILEW, stripH, body.join("")), svgPath, basePng, g.TILEW, stripH); - - // The codes are composited rather than inlined: an <image href> in the SVG - // would be resampled by rsvg, and a resampled QR does not scan. - const out = path.join(dir, "_rail_qr.png"); - const args = [basePng]; - for (const im of images) args.push(im.png, "-geometry", `+${im.x}+${im.y}`, "-composite"); - args.push(out); - await execFileP("magick", args, { maxBuffer: 1 << 26 }); - return { path: out, tileH: TH, urls }; -} - -/** - * Build the rail strips. Returns their paths plus the geometry, so the caller - * never re-derives a size the pixels already committed to. - * - * `entries` and `provenance` are needed for the QR strip only; without them the - * tile is skipped and the rail is the four strips it always was. - */ -export async function renderRailAssets(render, ledger, outDir, entries = null, provenance = null) { - const pal = render.palette; - const rail = render.rail; - const tracks = rail.tracks ?? []; - const byKey = Object.fromEntries(tracks.map((t) => [t.key, t])); - const g = railGeometry(render, ledger.length); - const dir = path.join(outDir, "cards"); - const rule = rail.rule ?? "#2A322F"; - const P = (n) => path.join(dir, n); - - // ---- chrome: opaque, full rail column, never gated ------------------- - // It runs for the whole video rather than being switched on with enable=, - // which removes four expressions and four off-by-one opportunities. - // - // The tally SWATCH and LABEL live here. They never change, and while they - // rode in the rolling block a coffee figure changing dragged "The Quartering" - // up the screen with it — text moving for a reason that was not about it. - const chromeBody = [ - `<rect x="0" y="0" width="${g.RW}" height="${g.RHGT}" fill="${pal.bg}"/>`, - `<rect x="0" y="0" width="1" height="${g.RHGT}" fill="${rule}"/>`, - svgText(g.PAD, 32, "THE CLAIM LEDGER", { size: 15, color: pal.amber, weight: "bold", ls: 1.6 }), - svgText(g.PAD, 55, "every count he has given, as he gave it", { size: 14, color: pal.muted }), - `<rect x="${g.PAD}" y="70" width="${g.RW - 2 * g.PAD}" height="1" fill="${rule}"/>`, - svgText(g.PAD, g.TALLYHEAD_REL + 16, "LATEST FIGURE HE HAS GIVEN", { - size: 12, color: pal.muted, weight: "bold", ls: 1.4, - }), - ...tracks.map((tr, j) => { - const ry = g.TALLYTOP_REL + j * g.TALLYROWH; - return ( - `<rect x="${g.PAD}" y="${ry + 13}" width="10" height="10" fill="${tr.color}"/>` + - svgText(g.PAD + 20, ry + 22, fit(tr.label, 14, g.CELLX - g.PAD - 24), { - size: 14, color: pal.fg, - }) - ); - }), - `<rect x="${g.PAD}" y="${g.ROSTERTOP_REL - 8}" width="${g.RW - 2 * g.PAD}" height="1" fill="${rule}"/>`, - svgText(g.PAD, g.ROSTERTOP_REL + 19, "ROSTER", { - size: 11, color: pal.muted, weight: "bold", ls: 1.4, - }), - `<rect x="${g.PAD}" y="${g.LOGTOP_REL - 26}" width="${g.RW - 2 * g.PAD}" height="1" fill="${rule}"/>`, - svgText(g.PAD, g.LOGTOP_REL - 8, "IN THE ORDER STATED", { - size: 12, color: pal.muted, weight: "bold", ls: 1.4, - }), - ].join(""); - const chrome = await rasterize( - svgDoc(g.RW, g.RHGT, chromeBody), P("_rail_chrome.svg"), P("_rail_chrome.png"), g.RW, g.RHGT, - ); - - // ---- log strip: every claim, stacked, no padding --------------------- - const valX = g.RW - g.PAD; - const textX = g.PAD + 20; - const textBudget = valX - textX - 74; - const rows = ledger.map((c, i) => { - const y = i * g.ROWH; - // `scope` is the ADJUDICATED answer and `company` the undocumented guess it - // replaced. Falls back so a manifest with no adjudication yet still renders. - const tr = byKey[c.scope ?? c.company]; - // In `sourced` every row has a clip behind it, so `live` is always true - // there; `full` keeps the distinction because a stacked ledger card is a - // weaker citation than footage and must not look like one. - const live = !!c.entryId && !c.unsourced; - const ink = live ? pal.fg : pal.muted; - const dotFill = live ? (tr?.color ?? pal.accent) : "none"; - return [ - `<rect x="0" y="${y}" width="${g.RW}" height="${g.ROWH}" fill="${pal.bg}"/>`, - `<circle cx="${g.PAD + 5}" cy="${y + 19}" r="4.5" fill="${dotFill}" ` + - `stroke="${tr?.color ?? pal.muted}" stroke-width="1.5" opacity="${live ? 1 : 0.55}"/>`, - svgText(textX, y + 17, c.date, { size: 13.5, color: pal.muted, opacity: live ? 1 : 0.7 }), - svgText(valX, y + 19, c.display ?? "—", { - size: 18, color: live ? (tr?.color ?? pal.fg) : pal.muted, - weight: "bold", anchor: "end", opacity: live ? 1 : 0.65, - }), - svgText(textX, y + 33, fit(c.label ?? c.quote ?? "", 13, textBudget + 74), { - size: 13, color: ink, opacity: live ? 1 : 0.6, - }), - `<rect x="${g.PAD}" y="${y + g.ROWH - 1}" width="${g.RW - 2 * g.PAD}" height="1" fill="${rule}"/>`, - ].join(""); - }).join(""); - const log = await rasterize( - svgDoc(g.RW, g.logStripH, rows), P("_rail_log.svg"), P("_rail_log.png"), g.RW, g.logStripH, - ); - - // ---- curtain --------------------------------------------------------- - // With ONE log strip and the window parked at the top while the list is still - // filling, rows i+1…K-1 would show claims the video has not made yet. This - // opaque rectangle rides just below the last revealed row and, once the list - // is full, parks exactly one window-height down — outside the window. It is - // pal.bg precisely so that parking there is invisible; the QR tile overlays - // AFTER it, which is what stops the parked curtain covering the code. - const curtain = P("_rail_curtain.png"); - await execFileP("magick", ["-size", `${g.RW}x${g.LOGH}`, `xc:${pal.bg}`, curtain]); - - // ---- highlight ------------------------------------------------------- - const hl = await rasterize( - svgDoc(g.RW, g.ROWH, [ - `<rect x="0" y="0" width="${g.RW}" height="${g.ROWH}" fill="${pal.amber}" opacity="0.10"/>`, - `<rect x="${g.PAD - 12}" y="4" width="3" height="${g.ROWH - 8}" fill="${pal.amber}"/>`, - ].join("")), - P("_rail_hl.svg"), P("_rail_hl.png"), g.RW, g.ROWH, - ); - - // ---- tally strip: one COLUMN per lane, side by side ------------------- - // Four cells and the roster in one PNG: five crops at different x out of one - // input, rather than five inputs. Every column is as tall as the tallest, so - // a single strip height serves them all. - const { lanes, rows: nRows } = tallyTracks(ledger, tracks, { rosterLine }); - const cellH = g.TALLYROWH; - const laneX = []; - let sx = 0; - for (const lane of lanes) { - const w = lane.kind === "roster" ? g.ROSTERW : g.CELLW; - laneX.push({ x: sx, w }); - sx += w; - } - const stripW = sx; - const stripH = nRows * cellH; - const strip = [`<rect x="0" y="0" width="${stripW}" height="${stripH}" fill="${pal.bg}"/>`]; - lanes.forEach((lane, li) => { - const { x } = laneX[li]; - lane.rows.forEach((cell, ri) => { - strip.push( - lane.kind === "roster" - ? rosterCellSvg(cell, x, ri * cellH, g, pal) - : tallyCellSvg(cell, x, ri * cellH, g, pal, cellH), - ); - }); - }); - const tally = await rasterize( - svgDoc(stripW, stripH, strip.join("")), P("_rail_tally.svg"), P("_rail_tally.png"), stripW, stripH, - ); - - // ---- the provenance tile --------------------------------------------- - const qr = - entries && provenance && render.qr !== false - ? await qrTileStrip(entries, provenance, render, g, outDir) - : null; - - return { - chrome, log, curtain, hl, tally, qr, - geom: g, - lanes: lanes.map((lane, li) => ({ - key: lane.key, kind: lane.kind, steps: lane.steps, - x: laneX[li].x, w: laneX[li].w, cellH, - })), - }; -} - -// =========================================================================== -// Stacked ledger cards -// =========================================================================== -// The `full` cut's answer to the 28 claims the sweep found and no clip covers. -// -// Dimmed rail rows were the old answer, and they were confusing: a row with no -// audio behind it slid past with nothing to say for itself, and 28 of them read -// as padding rather than as evidence. So each one gets SCREEN TIME instead -- -// its date, its scope, its quote, his figure, and what that figure does to our -// running sum. Consecutive unclipped claims share a card and reveal in -// sequence, which is why 28 claims cost 14 cards and about 67 seconds. -// -// The reveal is the rail curtain's device: an opaque `pal.bg` rectangle walking -// down the card. Nothing fades, nothing moves; rows simply stop being covered. - -/** The reveal clock. One definition, because `scheduleClaims` pins to it. */ -export const LEDGER_LEAD = 0.9; -export const LEDGER_STEP = 1.3; -export const LEDGER_TAIL = 2.6; -export const ledgerRevealAt = (r) => LEDGER_LEAD + LEDGER_STEP * r; -export const ledgerSeconds = (n) => LEDGER_LEAD + LEDGER_STEP * n + LEDGER_TAIL - LEDGER_STEP; - -/** - * Why this claim is a line of text and not footage. - * - * "The upload is gone" and "we did not cut it" are different sentences, and - * saying the first about a live source is the kind of error that makes a whole - * compilation untrustworthy. So the words come from a `yt-dlp --simulate` - * probe, recorded in out/availability.json with the date it ran. - */ -export const SOURCE_TAG = { - ok: "not clipped", - deleted: "source deleted", - private: "source private", - "members-only": "members only", - restricted: "age-restricted", - "geo-blocked": "geo-blocked", - "no-cues": "no archived transcript", - maybe_missing: "source unreachable", -}; - -/** Greedy wrap to a pixel budget, at most `maxLines`, last line elided. */ -function wrapPx(text, size, maxPx, maxLines) { - const perChar = size * 0.5; - const cols = Math.max(8, Math.floor(maxPx / perChar)); - const words = String(text ?? "").split(/\s+/).filter(Boolean); - const lines = []; - let line = ""; - for (const w of words) { - if (line && (line + " " + w).length > cols) { - lines.push(line); - line = w; - if (lines.length === maxLines) break; - } else { - line = line ? line + " " + w : w; - } - } - if (lines.length < maxLines && line) lines.push(line); - if (lines.length === maxLines) { - const used = lines.join(" ").split(/\s+/).length; - if (used < words.length) lines[maxLines - 1] = fit(lines[maxLines - 1] + " …", size, maxPx); - } - return lines; -} - -/** - * One card carrying a run of consecutive unclipped claims. - * - * Returns the geometry the encoder needs to walk the curtain: where the rows - * start and how tall each one is. - */ -export async function renderLedgerCard(card, render, ledger, outDir, avail = null) { - const pal = render.palette; - const tracks = render.rail?.tracks ?? []; - const byKey = Object.fromEntries(tracks.map((t) => [t.key, t])); - const VW = cardWidth(card, render); - const H = render.height; - const dir = path.join(outDir, "cards"); - const rule = render.rail?.rule ?? "#2A322F"; - const RESERVED = reservedFooterHeight(render); - - const ids = card.claims ?? []; - const rows = ids.map((id) => ledger.find((c) => c.id === id)).filter(Boolean); - if (rows.length !== ids.length) { - const missing = ids.filter((id) => !ledger.some((c) => c.id === id)); - throw new Error(`ledger card ${card.id} names claims that are not in the ledger: ${missing.join(", ")}`); - } - - // The arithmetic is READ, never recomputed: one implementation of the walk, - // or the card and the chart band can disagree about the same sum. - const steps = new Map(ledgerTotals(ledger).steps.map((st) => [st.id, st])); - - const M = 96; - const ARITHW = 320; - const arithX = VW - M - ARITHW; - const quoteW = arithX - 70 - M; - - const body = [`<rect x="0" y="0" width="${VW}" height="${H}" fill="${pal.bg}"/>`]; - body.push( - svgText(M, 80, (card.kicker ?? "found in the sweep, not clipped here").toUpperCase(), { - size: 22, color: pal.amber, weight: "bold", ls: 1.4, - }), - svgText(M, 132, card.heading ?? "What the sweep found and this cut cannot show you", { - size: 40, color: pal.fg, weight: "bold", - }), - svgText(M, 168, card.sub ?? "his own words, and what they do to our running sum", { - size: 20, color: pal.muted, - }), - `<rect x="${M}" y="${192}" width="${VW - 2 * M}" height="1" fill="${rule}"/>`, - ); - - // Rows are a fixed height and the BLOCK is centred in what is left of the - // frame. Stretching two rows to fill 640px puts a hand's width of nothing - // between them; packing them at the top leaves the same gap in one lump at - // the bottom. Centring is the only arrangement that reads as deliberate. - const top = 214; - const available = H - RESERVED - top - 24; - const ROWH = Math.min(180, Math.floor(available / rows.length)); - const rowsTop = top + Math.floor((available - ROWH * rows.length) / 2); - - rows.forEach((c, r) => { - const y = rowsTop + r * ROWH; - const tr = byKey[c.scope ?? c.company]; - const st = steps.get(c.id); - const state = avail?.get(c.id) ?? null; - const tag = SOURCE_TAG[state] ?? "not clipped"; - - body.push( - svgText(M, y + 32, c.date, { size: 20, color: pal.muted }), - // The scope, as a bordered pill in its own colour. Which payroll a number - // is about is the whole argument, so it is never left to the ink alone. - `<rect x="${M + 148}" y="${y + 12}" width="${Math.max(120, (tr?.label?.length ?? 8) * 7.6 + 22)}" ` + - `height="26" rx="13" fill="none" stroke="${tr?.color ?? pal.muted}" stroke-width="1.2"/>`, - svgText(M + 159, y + 30, tr?.label ?? c.scope ?? "", { size: 14.5, color: tr?.color ?? pal.muted }), - svgText(M + 148 + Math.max(120, (tr?.label?.length ?? 8) * 7.6 + 22) + 16, y + 30, tag, { - size: 14.5, color: pal.muted, opacity: 0.85, - }), - svgText(arithX - 70, y + 40, c.display ?? "—", { - size: 34, color: tr?.color ?? pal.fg, weight: "bold", anchor: "end", - }), - ); - wrapPx(`“${c.quote ?? c.label ?? ""}”`, 25, quoteW, 2).forEach((line, li) => { - body.push(svgText(M, y + 76 + li * 33, line, { size: 25, color: pal.fg })); - }); - - // ---- the arithmetic column ---- - // Which layer this claim moved, lit; the others held, dimmed. The point is - // that the total on the right is OURS and is made of his own figures. - body.push( - svgText(arithX, y + 24, "OUR RUNNING SUM", { - size: 11, color: pal.muted, weight: "bold", ls: 1.3, - }), - ); - const basis = st?.impliedBasis ?? {}; - const companies = tracks.filter((t) => t.key !== "all"); - // Before any company has given a figure, our sum is not zero — it is - // undefined, and four dashes in a column say that far less clearly than - // one sentence does. - if (!companies.some((t) => basis[t.key])) { - body.push( - svgText(arithX, y + 52, "no company figure yet,", { size: 14, color: pal.muted }), - svgText(arithX, y + 74, "so our sum is not defined", { size: 14, color: pal.muted }), - ); - if (r < rows.length - 1) { - body.push( - `<rect x="${M}" y="${y + ROWH - 1}" width="${VW - 2 * M}" height="1" fill="${rule}" opacity="0.6"/>`, - ); - } - return; - } - companies.forEach((t, k) => { - const b = basis[t.key]; - const moved = (c.scope ?? c.company) === t.key; - const ty = y + 48 + k * 24; - body.push( - svgText(arithX, ty, fit(t.shortLabel ?? t.label, 13, ARITHW - 90), { - size: 13, color: moved ? t.color : pal.muted, opacity: moved ? 1 : 0.55, - }), - svgText(arithX + ARITHW, ty, b ? String(b.value) : "—", { - size: 17, color: moved ? t.color : pal.muted, weight: "bold", anchor: "end", - opacity: moved ? 1 : 0.55, - }), - ); - }); - const sy = y + 48 + companies.length * 24; - body.push( - `<rect x="${arithX}" y="${sy + 6}" width="${ARITHW}" height="1" fill="${rule}"/>`, - svgText(arithX, sy + 30, "IMPLIED", { size: 13, color: pal.fg, weight: "bold", ls: 1.2 }), - svgText(arithX + ARITHW, sy + 32, st?.implied == null ? "—" : String(st.implied), { - size: 22, color: pal.fg, weight: "bold", anchor: "end", - }), - ); - if (st?.impliedDelta) { - const up = st.impliedDelta > 0; - body.push( - svgTri(arithX + 84, sy + 22, up, pal.amber), - svgText(arithX + 98, sy + 30, `${up ? "+" : "−"}${Math.abs(st.impliedDelta)}`, { - size: 14, color: pal.amber, weight: "bold", - }), - ); - } - - if (r < rows.length - 1) { - body.push( - `<rect x="${M}" y="${y + ROWH - 1}" width="${VW - 2 * M}" height="1" fill="${rule}" opacity="0.6"/>`, - ); - } - }); - - const outPath = path.join(dir, `${card.id}.png`); - await rasterize(svgDoc(VW, H, body.join("")), path.join(dir, `${card.id}.svg`), outPath, VW, H); - return { path: outPath, width: VW, rowsTop, rowHeight: ROWH, rows: rows.length }; -} - -// =========================================================================== -// End sequence: the ledger scroll and the step chart -// =========================================================================== - -/** - * The whole ledger as one tall PNG for an animated crop to walk. - * - * --------------------------------------------------------------------------- - * One chronological line, a column per company - * --------------------------------------------------------------------------- - * It used to group by company: four blocks, each date-sorted inside itself. The - * cut plays in ONE chronology, and grouping at the end re-tells it in an order - * the viewer has not just watched — and it hides the only thing worth seeing - * here, which is that the four payrolls were being described in the same weeks. - * - * So: one date-ordered list, and the company is read from COLUMN POSITION. That - * makes colour the secondary encoding rather than the only one, which is the - * same rule the chart already runs under. - * - * The rail hides for this card (`hideRail`), so it is drawn at the FULL frame - * width rather than the content width. - * - * Returns the CONTENT HEIGHT because the scroll expression is written against - * it — crop clamps its own y, so an off-by-a-few degrades into a static last - * frame rather than an error, but only if the caller knows the real number. - */ -export async function renderScrollCard(card, render, ledger, outDir) { - const pal = render.palette; - const tracks = render.rail?.tracks ?? []; - const VW = cardWidth(card, render); - const dir = path.join(outDir, "cards"); - const rule = render.rail?.rule ?? "#2A322F"; - - const M = 96; - const ROWH = 42; - const body = []; - - // Columns. The value columns are right-aligned on their own gridline, so a - // number's horizontal position IS its company even before the colour reads. - const COLW = 152; - const dateX = M; - const colX = tracks.map((_, i) => M + 168 + i * COLW); - const popX = M + 168 + tracks.length * COLW + 24; - const labelX = popX + 132; - const labelW = VW - M - labelX; - - let y = 66; - body.push(svgText(M, y, card.heading ?? "THE COMPLETE LEDGER", { - size: 30, color: pal.fg, weight: "bold", ls: 1.5, - })); - y += 32; - body.push(svgText(M, y, card.sub ?? `${ledger.length} dated claims, in the order he made them`, { - size: 19, color: pal.muted, - })); - y += 44; - - // The column heads, which are the legend. No separate key: a company name - // over its own column of figures is the shortest legend there is. - body.push(`<rect x="${M}" y="${y - 4}" width="${VW - 2 * M}" height="1" fill="${rule}"/>`); - body.push(svgText(dateX, y + 26, "DATE", { size: 13, color: pal.muted, weight: "bold", ls: 1.3 })); - // No swatch beside the head: the head is already IN the track's colour, and - // the column position is the primary encoding either way. A swatch would only - // land on top of the words, since a right-anchored run cannot be measured - // here to leave room for one. - tracks.forEach((tr, i) => { - body.push( - svgText(colX[i], y + 26, fit(tr.shortLabel ?? tr.label, 13, COLW - 12), { - size: 13, color: tr.color, weight: "bold", anchor: "end", ls: 0.6, - }), - ); - }); - body.push( - svgText(popX, y + 26, "AS WHAT", { size: 13, color: pal.muted, weight: "bold", ls: 1.3 }), - svgText(labelX, y + 26, "WHAT HE SAID", { size: 13, color: pal.muted, weight: "bold", ls: 1.3 }), - ); - y += 40; - body.push(`<rect x="${M}" y="${y}" width="${VW - 2 * M}" height="1" fill="${rule}"/>`); - y += 8; - - const byKey = Object.fromEntries(tracks.map((t, i) => [t.key, i])); - const rows = [...ledger].sort((a, b) => dateKey(a.date).localeCompare(dateKey(b.date))); - for (const c of rows) { - const live = !!c.entryId && !c.unsourced; - const i = byKey[c.scope ?? c.company]; - const tr = tracks[i]; - body.push( - svgText(dateX, y + 26, c.date, { size: 18, color: pal.muted, opacity: live ? 1 : 0.7 }), - ); - if (tr) { - body.push( - svgText(colX[i], y + 26, c.display ?? "—", { - size: 21, color: live ? tr.color : pal.muted, weight: "bold", anchor: "end", - opacity: live ? 1 : 0.6, - }), - ); - } - body.push( - svgText(popX, y + 26, POP_WORD[c.population] ?? c.population ?? "", { - size: 15, color: pal.muted, opacity: live ? 0.9 : 0.6, - }), - svgText(labelX, y + 26, fit(c.label ?? "", 18, labelW), { - size: 18, color: live ? pal.fg : pal.muted, opacity: live ? 1 : 0.6, - }), - `<rect x="${M}" y="${y + ROWH - 1}" width="${VW - 2 * M}" height="1" fill="${rule}" opacity="0.5"/>`, - ); - y += ROWH; - } - y += 60; - - const contentHeight = y; - const outPath = path.join(dir, `${card.id}.png`); - await rasterize( - svgDoc(VW, contentHeight, `<rect x="0" y="0" width="${VW}" height="${contentHeight}" fill="${pal.bg}"/>${body.join("")}`), - path.join(dir, `${card.id}.svg`), outPath, VW, contentHeight, - ); - return { path: outPath, contentHeight, width: VW }; -} - -/** - * The four-series step chart, over the claims flagged `plotted`. - * - * COLOUR IS NOT THE ONLY ENCODING here, and that is a hard requirement rather - * than a flourish: no four-colour categorical palette clears the data-viz - * all-pairs CVD gate (three slots is the documented ceiling), so each series - * also carries a distinct dash pattern and a direct end-of-line label. The four - * hues themselves are the published artifact's, re-validated against this - * video's darker ground (#0F1312) on the adjacent pairlist — the pairlist for - * line charts — where all five checks pass. - */ -export async function renderChartCard(card, render, ledger, outDir) { - const pal = render.palette; - const tracks = render.rail?.tracks ?? []; - const VW = cardWidth(card, render); - const H = render.height; - const dir = path.join(outDir, "cards"); - const rule = render.rail?.rule ?? "#2A322F"; - - // The series come from ledger-totals, not from the legacy `plotted` flag. - // `plotted` was set under the OLD reading, in which a sum we performed sat in - // the same series as a figure he uttered. Drawing from it now would put 18 and - // 20 back on his line, after the whole point of the adjudication was to take - // them off it. - let totals = null; - try { - totals = ledgerTotals(ledger); - } catch { - // An unadjudicated ledger still renders -- as the three company series only, - // because the two totals are exactly what it cannot be trusted about. - totals = null; - } - const pts = ledger.filter((c) => c.value != null && (c.scope ?? c.company) !== "all"); - const yr = (d) => { - const [Y, M2, D2] = d.split("-").map(Number); - return Y + (M2 - 1) / 12 + (D2 - 1) / 365; - }; - const X0 = yr("2020-01-01"), X1 = yr("2026-12-31"); - const YMAX = - Math.max(21, ...pts.map((p) => p.value), ...(totals?.series.implied ?? []).map((p) => p.value)) + 1; - - const RESERVED = reservedFooterHeight(render); - const box = { l: 150, r: 300, t: 190, b: 130 + RESERVED }; - const plotW = VW - box.l - box.r; - const plotH = H - box.t - box.b; - const px = (v) => box.l + ((v - X0) / (X1 - X0)) * plotW; - const py = (v) => H - box.b - (v / YMAX) * plotH; - - const body = [`<rect x="0" y="0" width="${VW}" height="${H}" fill="${pal.bg}"/>`]; - body.push( - svgText(box.l, 78, "WHAT HE SAID, AND WHAT IT ADDS UP TO", { - size: 34, color: pal.fg, weight: "bold", ls: 1.5, - }), - svgText(box.l, 112, "every figure he utters, against the company he was talking about", { - size: 20, color: pal.muted, - }), - svgText(box.l, 146, "the heavy line is ours — his own per-company claims, added up", { - size: 18, color: pal.amber, - }), - ); - - // grid + axes - for (let gv = 0; gv <= YMAX - 1; gv += 5) { - body.push( - `<rect x="${box.l}" y="${py(gv)}" width="${plotW}" height="1" fill="${rule}"/>`, - svgText(box.l - 16, py(gv) + 6, String(gv), { size: 17, color: pal.muted, anchor: "end" }), - ); - } - body.push(svgText(box.l - 16, py(YMAX - 1) - 22, "PEOPLE", { - size: 13, color: pal.muted, weight: "bold", anchor: "end", ls: 1.2, - })); - for (let Y = 2020; Y <= 2026; Y += 1) { - const x = px(yr(`${Y}-01-01`)); - body.push( - `<rect x="${x}" y="${box.t}" width="1" height="${py(0) - box.t}" fill="${rule}" opacity="0.7"/>`, - svgText(x, py(0) + 30, String(Y), { size: 17, color: pal.muted, anchor: "middle" }), - ); - } - body.push(`<rect x="${box.l}" y="${py(0)}" width="${plotW}" height="2" fill="${pal.muted}"/>`); - - // One step path per series, plus a dot per claim and a direct end label. - // - // FIVE series, not four. The three companies are his, drawn as before. The - // fourth is what he says the WHOLE payroll is -- only ever a figure he utters - // as one number. The fifth is what his own per-company claims add up to, and - // it is ours: a heavy neutral step, because an aggregate is not a categorical - // peer of the things it aggregates and must not consume a palette slot. - const DASH = ["", "12 6", "3 7", "18 5 4 5"]; - const labels = []; - const drawn = []; - tracks.forEach((tr, ti) => { - if (tr.key === "all") return; - drawn.push({ - tr, dash: DASH[ti % 4], width: 3.5, - pts: pts.filter((c) => (c.scope ?? c.company) === tr.key) - .slice().sort((a, b) => a.date.localeCompare(b.date)) - .map((c) => ({ date: c.date, value: c.value, display: c.display, hedged: c.hedged })), - }); - }); - if (totals) { - const allTrack = tracks.find((t) => t.key === "all"); - drawn.push({ - tr: { key: "stated", color: allTrack?.color ?? pal.accent, label: "stated total" }, - dash: "18 5 4 5", width: 3.5, dots: true, - pts: totals.series.stated.map((p) => ({ date: p.date, value: p.value, display: String(p.value) })), - }); - drawn.push({ - tr: { key: "implied", color: pal.fg, label: "implied — our sum" }, - dash: "", width: 6, dots: false, - pts: totals.series.implied.map((p) => ({ date: p.date, value: p.value, display: String(p.value) })), - }); - } - - for (const sr of drawn) { - const mine = sr.pts; - if (!mine.length) continue; - // A step, not a line: the figure he gave holds until he gives another one, - // so the segment between two claims must be flat and the change vertical. - let d = ""; - let prevY = null; - for (const [i, c] of mine.entries()) { - const x = px(yr(c.date)); - const yv = py(c.value); - d += i === 0 - ? `M ${x.toFixed(1)} ${yv.toFixed(1)}` - : ` L ${x.toFixed(1)} ${prevY.toFixed(1)} L ${x.toFixed(1)} ${yv.toFixed(1)}`; - prevY = yv; - } - const last = mine[mine.length - 1]; - const lastY = py(last.value); - d += ` L ${(box.l + plotW).toFixed(1)} ${lastY.toFixed(1)}`; - body.push( - `<path d="${d}" fill="none" stroke="${sr.tr.color}" stroke-width="${sr.width}" ` + - `stroke-linejoin="round"${sr.dash ? ` stroke-dasharray="${sr.dash}"` : ""}/>`, - ); - if (sr.dots !== false) { - for (const c of mine) { - body.push( - `<circle cx="${px(yr(c.date)).toFixed(1)}" cy="${py(c.value).toFixed(1)}" r="${c.hedged ? 5 : 6}" ` + - `fill="${c.hedged ? pal.bg : sr.tr.color}" stroke="${sr.tr.color}" stroke-width="2.5"/>`, - ); - } - } - labels.push({ tr: sr.tr, last, lineY: lastY, y: lastY }); - } - - // The closing hold annotates the gap it has just finished drawing. - if (card.hold && totals && totals.final.stated != null && totals.final.implied != null) { - const xR = box.l + plotW; - const yS = py(totals.final.stated); - const yI = py(totals.final.implied); - body.push( - `<rect x="${(xR - 190).toFixed(1)}" y="${Math.min(yI, yS).toFixed(1)}" width="170" ` + - `height="${Math.abs(yS - yI).toFixed(1)}" fill="${tracks.find((t) => t.key === "all")?.color ?? pal.accent}" opacity="0.12"/>`, - `<path d="M ${(xR - 105).toFixed(1)} ${yI.toFixed(1)} L ${(xR - 105).toFixed(1)} ${yS.toFixed(1)}" ` + - `stroke="${pal.amber}" stroke-width="2"/>`, - svgText(xR - 96, (yI + yS) / 2 - 4, `gap ${Math.round(totals.final.implied - totals.final.stated)}`, { - size: 22, color: pal.amber, weight: "bold", - }), - svgText(xR - 96, (yI + yS) / 2 + 22, "between his last total and our sum", { - size: 14, color: pal.muted, - }), - ); - } - - // Three of the four series end within a couple of people of each other, so - // their direct labels land on top of one another. Push them apart and elbow a - // leader line back to the value each one actually belongs to — direct labels - // are the secondary encoding that lets a four-colour palette be legible at - // all, so an unreadable stack would defeat the point of having them. - const LBLH = 48; - labels.sort((a, b) => a.y - b.y); - for (let i = 1; i < labels.length; i += 1) { - labels[i].y = Math.max(labels[i].y, labels[i - 1].y + LBLH); - } - const overshoot = labels.length ? labels[labels.length - 1].y - (H - box.b - 10) : 0; - if (overshoot > 0) for (const l of labels) l.y -= overshoot; - for (const l of labels) { - const lx = box.l + plotW; - if (Math.abs(l.y - l.lineY) > 2) { - body.push( - `<path d="M ${lx} ${l.lineY.toFixed(1)} L ${lx + 9} ${l.lineY.toFixed(1)} ` + - `L ${lx + 9} ${l.y.toFixed(1)} L ${lx + 14} ${l.y.toFixed(1)}" fill="none" ` + - `stroke="${l.tr.color}" stroke-width="1.5" opacity="0.75"/>`, - ); - } - body.push( - svgText(lx + 20, l.y + 2, l.tr.label, { size: 18, color: l.tr.color, weight: "bold" }), - svgText(lx + 20, l.y + 22, `last stated ${l.last.display}`, { size: 14, color: pal.muted }), - ); - } - - body.push( - svgText(box.l, H - RESERVED - 56, "hollow dot = a hedge word (“nearly ten”, “a handful”), not a figure", { - size: 16, color: pal.muted, - }), - svgText(box.l, H - RESERVED - 30, "each series is dashed as well as coloured — the shapes carry the reading on their own; " + - "sums and midpoints are ours and are never drawn as his", { - size: 16, color: pal.muted, - }), - ); - - const outPath = path.join(dir, `${card.id}.png`); - await rasterize(svgDoc(VW, H, body.join("")), path.join(dir, `${card.id}.svg`), outPath, VW, H); - return { - path: outPath, - plotX: box.l, plotY: box.t, plotW, plotH: py(0) - box.t + 2, - }; -} - -export async function renderCard(card, render, outDir, nodes) { - if (card.style === "timeline") { - if (!nodes?.length) throw new Error(`card ${card.id} is style:timeline but no timelineNodes given`); - return renderTimelineCard(card, render, nodes, outDir); - } - return renderPlainCard(card, render, outDir); -} - -async function renderPlainCard(card, render, outDir) { - const pal = render.palette; - const { width, height } = render; - const VW = cardWidth(card, render); - const textWidth = Math.round(VW * 0.74); - const outPath = path.join(outDir, "cards", `${card.id}.png`); - - // Pango reads its markup from a file to keep it clear of shell/argv quoting. - const markupPath = path.join(outDir, "cards", `${card.id}.pango`); - await writeFile(markupPath, markupFor(card, pal), "utf8"); - - // One magick invocation: solid ground, an accent rule down the left margin, - // then the Pango block composited over it. The rule is what keeps the cards - // recognisably one family across styles. - const barX = Math.round(VW * 0.09); - const barTop = Math.round(height * 0.28); - const barBottom = Math.round(height * 0.72); - - const args = [ - "-size", `${width}x${height}`, - `xc:${pal.bg}`, - "-fill", pal.accent, - "-draw", `rectangle ${barX},${barTop} ${barX + 6},${barBottom}`, - "(", - // `-size` is still set to the full frame from the canvas above, and the - // pango delegate honours it — leaving it alone renders the text into a - // 1920x1080 box, which pins the block to the top and wraps at the frame - // edge instead of the margin. Reset it to the text column, height auto. - "-size", `${textWidth}x`, - "-background", "none", - "-define", `pango:width=${textWidth}`, - "-define", "pango:alignment=left", - "-define", "pango:wrap=word", - `pango:@${markupPath}`, - ")", - "-gravity", "West", - "-geometry", `+${barX + 58}+0`, - "-composite", - outPath, - ]; - - await execFileP("magick", args, { maxBuffer: 1 << 24 }); - return outPath; -} - -async function main() { - const argv = process.argv.slice(2); - const manifestPath = argv.find((a) => !a.startsWith("--")); - if (!manifestPath) { - console.error("usage: render-cards.mjs <manifest.json> [--out <dir>] [--only <id>]"); - process.exit(2); - } - const flag = (name) => { - const i = argv.indexOf(name); - return i >= 0 ? argv[i + 1] : undefined; - }; - - const manifest = JSON.parse(await readFile(manifestPath, "utf8")); - const outDir = flag("--out") ?? path.join(path.dirname(path.resolve(manifestPath)), "out"); - const only = flag("--only"); - - await mkdir(path.join(outDir, "cards"), { recursive: true }); - - const cards = manifest.timeline.filter( - (e) => e.type === "card" && (!only || e.id === only), - ); - for (const card of cards) { - const p = await renderCard(card, manifest.render, outDir, manifest.timelineNodes); - console.log(`card ${card.id} -> ${p}`); - } - console.log(`${cards.length} card(s) rendered`); -} - -if (import.meta.url === `file://${process.argv[1]}`) { - main().catch((err) => { - console.error(err); - process.exit(1); - }); -} diff --git a/scripts/report-to-video/resolve-windows.mjs b/scripts/report-to-video/resolve-windows.mjs @@ -1,241 +0,0 @@ -#!/usr/bin/env node -// resolve-windows.mjs — widen a manifest's clip windows to whole sentences. -// -// A manifest window starts life as the cue span covering a quote, and a cue -// boundary is a bad place to cut: ASR breaks cues where the caption line wrapped, -// which is routinely mid-sentence and often mid-word. Cutting there drops the -// lead-in that makes a quote make sense, and clips audibly start and stop in the -// middle of speech. -// -// This walks outward from the cue span to the nearest sentence boundary in the -// transcript — a cue whose text ends in . ? or ! — so the clip carries the whole -// thought. Word-level alignment is a separate, audio-side problem: build-video.mjs -// snaps the actual cut to a silence (see --fetch-pad / snapping there). -// -// Expansion is capped so a run-on passage can't drag a clip out to a minute. -// -// In the app: not used. On the CLI: -// node scripts/report-to-video/resolve-windows.mjs <manifest.json> [--write] -// -// Options: -// --write Rewrite the manifest in place (default: dry run, print a table) -// --max-lead <s> Max seconds to expand backwards (default 9) -// --max-tail <s> Max seconds to expand forwards (default 12) -// --site-origin <url> Archive to read cues from when there is no local corpus -// (defaults to the manifest's provenance.siteOrigin) -// --resolve-site-ids On a published-id miss, find the record by scanning the -// channel's shards. Slow; see cues.mjs. -// --cue-source <which> auto (default) | local | http. The two can disagree -// once a corpus moves past its last publish — see cues.mjs. -// -// A clip entry may set `lockStart` / `lockEnd` to pin that edge exactly. - -import { readFile, writeFile } from "node:fs/promises"; - -import { createCueSource, siteOriginFromManifest } from "./cues.mjs"; - - -const ENDS_SENTENCE = /[.!?]["'”’)\]]*\s*$/; - -// A cue that is only "[music]" or "[ __ ]" (the profanity bleep) carries no -// sentence signal; treat it as transparent so expansion walks past it. -const IS_FILLER = /^\s*(\[[^\]]*\]|>>|♪|—|-)*\s*$/; - -// The manifest stores times rounded to 2 dp, so a value read back from it can sit -// a hair BELOW the cue end it came from. Without a tolerance the end lookup then -// lands on the previous cue, the forward search runs on to the next sentence, and -// the clip grows a little every time this is run — it has to be a fixed point. -const EPS = 0.02; - - -function indexAt(cues, t, which) { - // First cue whose span contains t, else the nearest one on the right side. - let idx = cues.findIndex((c) => c.end > t); - if (idx < 0) idx = cues.length - 1; - if (which === "end") { - let j = cues.findIndex((c) => c.end >= t - EPS); - if (j < 0) j = cues.length - 1; - idx = j; - } - return idx; -} - -export function widen(cues, start, end, { maxLead = 8, maxTail = 12 } = {}) { - const isBoundary = (c) => ENDS_SENTENCE.test(c.text) && !IS_FILLER.test(c.text); - const i0 = indexAt(cues, start, "start"); - const i1 = indexAt(cues, end, "end"); - - // START: the latest cue that OPENS a sentence (i.e. its predecessor closes - // one) at or before the quote, within the lead budget. Finding no such cue - // means every candidate lead-in is a sentence fragment, so take none at all — - // a fragment is the irrelevant context we are trying to avoid, not context. - let si = null; - for (let i = i0; i > 0; i -= 1) { - if (start - cues[i].start > maxLead) break; - if (isBoundary(cues[i - 1])) { - si = i; - break; - } - } - if (si === null) si = i0; - - // END: the first cue that CLOSES a sentence at or after the quote. Never - // clamp to a budget here — stopping partway through a sentence is exactly the - // mid-thought ending this is meant to remove, so the budget only decides how - // far to look, and failing to find one falls back to the original cue end. - let ei = null; - for (let j = i1; j < cues.length; j += 1) { - if (cues[j].end - end > maxTail) break; - if (isBoundary(cues[j])) { - ei = j; - break; - } - } - if (ei === null) ei = i1; - - return { - start: cues[si].start, - end: cues[ei].end, - leadCues: i0 - si, - tailCues: ei - i1, - }; -} - -async function main() { - const argv = process.argv.slice(2); - const manifestPath = argv.find((a) => !a.startsWith("--")); - if (!manifestPath) { - console.error("usage: resolve-windows.mjs <manifest.json> [--write]"); - process.exit(2); - } - const num = (name, dflt) => { - const i = argv.indexOf(name); - return i >= 0 ? Number(argv[i + 1]) : dflt; - }; - // Lead is where the context lives — it is the run-up that makes a quote make - // sense. Tail only needs to finish the sentence, so it gets a smaller budget. - const opts = { maxLead: num("--max-lead", 8), maxTail: num("--max-tail", 12) }; - - const manifest = JSON.parse(await readFile(manifestPath, "utf8")); - const slug = manifest.provenance.channelSlug; - const cache = new Map(); - - // Cues come from a local corpus when there is one, and from the published - // archive the manifest was built against when there is not — so this runs in a - // clone with no `transcripts/` at all. See cues.mjs. - const flag = (name) => { - const i = argv.indexOf(name); - return i >= 0 ? argv[i + 1] : undefined; - }; - const cues = createCueSource({ - siteOrigin: flag("--site-origin") ?? process.env.SITE_ORIGIN ?? siteOriginFromManifest(manifest), - resolveSiteIds: argv.includes("--resolve-site-ids"), - prefer: flag("--cue-source") ?? "auto", - log: (m) => console.error(` · ${m}`), - }); - const loadCues = (videoId, channelSlug, hints) => - cues.load(channelSlug, videoId, hints).then((r) => r.cues); - - let changed = 0; - for (const e of manifest.timeline) { - if (e.type !== "clip") continue; - // A compilation can span several archived channels (the same streamer's VODs - // are mirrored across more than one), so a clip may name its own. Key the - // cache by channel too — the same id under a different slug is a different file. - // An author can trim a clip to land mid-cue on purpose — a cue often carries - // a whole paragraph, and cutting a quote short is an editorial decision. - // Widening would undo exactly that, so `lock` opts the clip out. - if (e.lock) { - console.log(`${e.id.padEnd(4)} ${e.video.padEnd(12)} locked, left at ${e.start.toFixed(1)}–${e.end.toFixed(1)}`); - continue; - } - const chan = e.channel ?? slug; - const key = `${chan}/${e.video}`; - if (!cache.has(key)) { - cache.set( - key, - await loadCues(e.video, chan, { siteChannel: e.siteChannel, siteVideo: e.siteVideo }), - ); - } - const cues = cache.get(key); - - const before = { start: e.start, end: e.end }; - const w = widen(cues, e.start, e.end, opts); - - // `lockStart` / `lockEnd` pin an edge to exactly what the author wrote. The - // escape hatch exists because sentence detection is only as good as the ASR's - // punctuation, and some uploads have none at all — and because an utterance's - // real trailing pause does not always line up with its last cue's end. - if (e.lockStart) w.start = before.start; - if (e.lockEnd) w.end = before.end; - const dLead = (before.start - w.start).toFixed(1); - const dTail = (w.end - before.end).toFixed(1); - const dur = (w.end - w.start).toFixed(1); - - // Ignore sub-frame drift so a re-run on an already-resolved manifest is a - // genuine no-op rather than a rewrite that nudges every window. - const moved = - Math.abs(w.start - before.start) > 0.05 || Math.abs(w.end - before.end) > 0.05; - if (moved) changed += 1; - console.log( - `${e.id.padEnd(4)} ${e.video.padEnd(12)} ` + - `${before.start.toFixed(1)}–${before.end.toFixed(1)} -> ` + - `${w.start.toFixed(1)}–${w.end.toFixed(1)} (+${dLead}s lead, +${dTail}s tail, ${dur}s)`, - ); - - if (moved) { - e.start = Number(w.start.toFixed(2)); - e.end = Number(w.end.toFixed(2)); - } - } - - // De-overlap clips that come from the SAME video. Widening is per-clip and - // blind to its neighbours, so a tail that finds no sentence boundary runs to - // the budget and can swallow the next clip's material — which plays as the - // same footage twice. (Real case: a 2024 upload whose ASR carries no - // punctuation at all in that stretch, so nothing stopped the search.) - // The later clip's start is the deliberate one, so trim the earlier clip's tail. - const byVideo = new Map(); - for (const e of manifest.timeline) { - if (e.type !== "clip") continue; - if (!byVideo.has(e.video)) byVideo.set(e.video, []); - byVideo.get(e.video).push(e); - } - for (const [video, list] of byVideo) { - if (list.length < 2) continue; - list.sort((a, b) => a.start - b.start); - for (let i = 0; i < list.length - 1; i += 1) { - const a = list[i]; - const b = list[i + 1]; - if (a.end <= b.start) continue; - const overlap = a.end - b.start; - if (a.lockEnd) { - console.log(` ⚠ ${a.id} overlaps ${b.id} by ${overlap.toFixed(1)}s but has lockEnd — not trimmed`); - continue; - } - a.end = Number(b.start.toFixed(2)); - changed += 1; - console.log( - ` de-overlap ${video}: ${a.id} trimmed ${overlap.toFixed(1)}s off its tail ` + - `(it ran into ${b.id})`, - ); - if (a.end - a.start < 3) { - console.log(` ⚠ ${a.id} is now only ${(a.end - a.start).toFixed(1)}s — check it`); - } - } - } - - if (argv.includes("--write")) { - await writeFile(manifestPath, JSON.stringify(manifest, null, 2) + "\n", "utf8"); - console.log(`\nwrote ${manifestPath} (${changed} window(s) changed)`); - } else { - console.log(`\ndry run — ${changed} window(s) would change; pass --write to apply`); - } -} - -if (import.meta.url === `file://${process.argv[1]}`) { - main().catch((err) => { - console.error(err); - process.exit(1); - }); -} diff --git a/scripts/report-to-video/verify-build.mjs b/scripts/report-to-video/verify-build.mjs @@ -1,110 +0,0 @@ -#!/usr/bin/env node -// verify-build.mjs — is the file that came out the file that was asked for? -// -// A build can exit 0 and still be wrong in ways nothing else notices: a concat -// that produced a zero-length file, a chapter pass that silently dropped -// markers, a timeline that lost a clip because --continue-on-error let it. Each -// of those looks like success at the terminal and like a finished video in a -// directory listing. -// -// So the last step of a build measures the deliverable and compares it to the -// manifest. Cheap (one ffprobe) and the only thing that closes the loop. -// -// node scripts/report-to-video/verify-build.mjs <manifest.json> [--out <dir>] -// [--variant sourced|full] [--json] - -import { execFile } from "node:child_process"; -import { promisify } from "node:util"; -import { readFile, stat } from "node:fs/promises"; -import path from "node:path"; - -import { selectVariant, variantPaths } from "./build-video.mjs"; - -const execFileP = promisify(execFile); -const FFPROBE = process.env.FFPROBE_BIN ?? "ffprobe"; - -export async function verifyBuild(manifestPath, { outDir, variant = "sourced" } = {}) { - // The SAME filter the build ran. Verifying the whole manifest against one - // variant's file would report a missing chapter for every entry the other cut - // carries -- i.e. it would be red exactly when the build was right. - const manifest = selectVariant( - JSON.parse(await readFile(manifestPath, "utf8")), - variant, - ); - const root = outDir ?? path.join(path.dirname(path.resolve(manifestPath)), "out"); - const file = variantPaths(root, manifest.slug, variant).final; - const problems = []; - - const st = await stat(file).catch(() => null); - if (!st) return { ok: false, file, problems: [`${file} does not exist`] }; - if (st.size < 1024) problems.push(`${file} is ${st.size} bytes`); - - const { stdout } = await execFileP(FFPROBE, [ - "-v", "error", - "-show_entries", "format=duration,size", - "-show_chapters", - "-of", "json", - file, - ], { maxBuffer: 1 << 24 }); - const probe = JSON.parse(stdout); - const duration = Number(probe.format?.duration ?? 0); - const chapters = (probe.chapters ?? []).length; - const entries = (manifest.timeline ?? []).length; - - if (!(duration > 0)) problems.push("duration is not greater than zero"); - - // Every timeline entry becomes a chapter, so a mismatch means the timeline and - // the file disagree about what is in it -- which is exactly the failure - // --continue-on-error is allowed to cause and must never cause silently. - if (chapters > 0 && chapters !== entries) { - problems.push(`${chapters} chapter(s) for ${entries} timeline entr(ies) — the cut is missing something`); - } - - // A rough floor: the sum of the windows, less the crossfades. Well under the - // real duration because snapping moves the cuts, but a file that came out at - // half the expected length did not build what was asked for. - const wanted = (manifest.timeline ?? []).reduce( - // `seconds` covers cards and the two end-sequence kinds (scroll, chart); - // only a clip's length has to be derived from its window. - (n, e) => n + (e.type === "clip" ? Math.max(0, (e.end ?? 0) - (e.start ?? 0)) : (e.seconds ?? 0)), - 0, - ); - if (wanted > 0 && duration < wanted * 0.5) { - problems.push(`${duration.toFixed(1)}s out of a timeline that asks for about ${wanted.toFixed(0)}s`); - } - - return { ok: problems.length === 0, variant, file, duration, chapters, entries, size: st.size, problems }; -} - -async function main() { - const argv = process.argv.slice(2); - const manifestPath = argv.find((a) => !a.startsWith("--")); - if (!manifestPath) { - console.error("usage: verify-build.mjs <manifest.json> [--out <dir>] [--variant sourced|full] [--json]"); - process.exit(2); - } - const flag = (n) => { const i = argv.indexOf(n); return i >= 0 ? argv[i + 1] : undefined; }; - const res = await verifyBuild(manifestPath, { - outDir: flag("--out"), - variant: flag("--variant") ?? "sourced", - }); - - if (argv.includes("--json")) { - console.log(JSON.stringify(res, null, 2)); - } else { - console.log( - `${res.file} (${res.variant})\n ${res.duration?.toFixed(1) ?? "?"}s · ${res.chapters ?? 0} chapter(s) for ` + - `${res.entries ?? 0} entr(ies) · ${((res.size ?? 0) / 1e6).toFixed(1)} MB`, - ); - for (const p of res.problems) console.log(` ** ${p}`); - if (res.ok) console.log(" ok"); - } - process.exit(res.ok ? 0 : 1); -} - -if (import.meta.url === `file://${process.argv[1]}`) { - main().catch((err) => { - console.error(err.message ?? err); - process.exit(1); - }); -} diff --git a/umtool/app/api/report/build/route.ts b/umtool/app/api/report/build/route.ts @@ -5,7 +5,7 @@ import { PRESETS, buildSteps } from "@/lib/report/driver.mjs"; import { clipsOf, readManifest } from "@/lib/projects/report.mjs"; import { projectRef } from "@/lib/projects"; import { recordBuildSnapshot } from "@/lib/report/snapshots.mjs"; -import { DEFAULT_VARIANT, VARIANTS, variantPaths } from "report-to-video/build-video"; +import { DEFAULT_VARIANT, VARIANTS, variantPaths } from "umtool-report-to-video/build-video"; export const dynamic = "force-dynamic"; diff --git a/umtool/bin/umtool.mjs b/umtool/bin/umtool.mjs @@ -513,7 +513,7 @@ async function cmdWindow() { const marks = ["lock", "lockStart", "lockEnd"].filter((k) => res.entry[k]); if (marks.length) console.log(` ${marks.join(", ")}`); console.log(`\nRun resolve-windows to see whether the widener agrees:`); - console.log(` node scripts/report-to-video/resolve-windows.mjs ${p.dir}/video.manifest.json`); + console.log(` node umtool/report-to-video/resolve-windows.mjs ${p.dir}/video.manifest.json`); } catch (e) { die(e?.message ?? String(e)); } diff --git a/umtool/components/projects/ClaimBench.tsx b/umtool/components/projects/ClaimBench.tsx @@ -8,7 +8,7 @@ import { SCOPE_CONFIDENCE, VALUE_KINDS, rosterLine, -} from "report-to-video/ledger-totals"; +} from "umtool-report-to-video/ledger-totals"; // Ruling on one claim. // diff --git a/umtool/components/projects/ClaimBenchPage.tsx b/umtool/components/projects/ClaimBenchPage.tsx @@ -3,7 +3,7 @@ import BrowseHeader from "@/components/BrowseHeader"; import ClaimBench, { type ClaimBenchData } from "./ClaimBench"; import { manifestToken } from "@/lib/report/manifest.mjs"; import { readClaimDetail, readManifest } from "@/lib/projects/report.mjs"; -import { ledgerTotals } from "report-to-video/ledger-totals"; +import { ledgerTotals } from "umtool-report-to-video/ledger-totals"; import type { ProjectRef } from "@/lib/project-types"; // The server half of the claim bench. diff --git a/umtool/docs/README.md b/umtool/docs/README.md @@ -40,8 +40,8 @@ umtool doctor # can this machine build at all? umtool check <slug> # 3. widen windows to whole sentences (dry first, then apply) -node scripts/report-to-video/resolve-windows.mjs ~/reports/<slug>/video.manifest.json -node scripts/report-to-video/resolve-windows.mjs ~/reports/<slug>/video.manifest.json --write +node umtool/report-to-video/resolve-windows.mjs ~/reports/<slug>/video.manifest.json +node umtool/report-to-video/resolve-windows.mjs ~/reports/<slug>/video.manifest.json --write # 4. bench any clip whose edges you are unsure of # /browse/<slug>/clip/<id> diff --git a/umtool/docs/authoring.md b/umtool/docs/authoring.md @@ -74,8 +74,8 @@ gone. ## 4. Widen to sentences ```sh -node scripts/report-to-video/resolve-windows.mjs <manifest> # dry -node scripts/report-to-video/resolve-windows.mjs <manifest> --write # apply +node umtool/report-to-video/resolve-windows.mjs <manifest> # dry +node umtool/report-to-video/resolve-windows.mjs <manifest> --write # apply ``` A cue boundary is a *line-wrap* boundary, so cutting there drops the lead-in that diff --git a/umtool/docs/report-video.md b/umtool/docs/report-video.md @@ -5,8 +5,8 @@ sources in their own voice, in order, each clip carrying a burned-in attribution and a QR that resolves to that exact moment in the archive's own viewer. The point is that a viewer does not have to take the edit on trust. -The pipeline lives at `scripts/report-to-video/` and its own -[README](../../scripts/report-to-video/README.md) is the source of truth for how it +The pipeline lives at `umtool/report-to-video/` and its own +[README](../report-to-video/README.md) is the source of truth for how it renders. This sheet covers what that README cannot know: what umtool reads, what umtool **writes**, and the rules a UI has to honour so the CLI and the app can never disagree. @@ -177,7 +177,7 @@ hazards were riding on it: as a single number for a named scope**. Sums and midpoints are ours and belong to the *implied* series, which says so on screen. -`scripts/report-to-video/ledger-totals.mjs` is the single implementation of that +`umtool/report-to-video/ledger-totals.mjs` is the single implementation of that arithmetic — the umtool reducer, the chart band and the closing card all import it, so they cannot disagree. It **refuses to run on an unadjudicated ledger** (`strict: true`); the inbox passes `strict: false` so it can show coherence rows diff --git a/umtool/e2e/clip-bench.spec.ts b/umtool/e2e/clip-bench.spec.ts @@ -143,7 +143,7 @@ test("the edit the bench predicts is a FIXED POINT for the widener", async ({ re const out = execFileSync( "node", - [path.join(UMTOOL, "..", "scripts", "report-to-video", "resolve-windows.mjs"), MANIFEST], + [path.join(UMTOOL, "report-to-video", "resolve-windows.mjs"), MANIFEST], { encoding: "utf8", env: { ...process.env, CHANNELS_DIR: path.join(FIXTURE, "channels") } }, ); // Setting an edge where a sentence actually ends means the next --write is a diff --git a/umtool/lib/decisions.ts b/umtool/lib/decisions.ts @@ -20,7 +20,7 @@ import { acceptedFor, readThumbAccepted, readThumbManifest, thumbNamesFor } from // of songs. // // `project` rather than `song`. It is the one neutral name that matters here: -// the second video pipeline in this repo (scripts/report-to-video) has open +// the second video pipeline in this repo (umtool/report-to-video) has open // decisions of exactly this shape -- a clip nobody judged, a manifest whose // sources have gone -- and when it arrives as a second project kind it should // land in this list rather than beside it. Nothing else in the type mentions a diff --git a/umtool/lib/projects/report.mjs b/umtool/lib/projects/report.mjs @@ -7,7 +7,7 @@ // and neither parses JSON. import { readdir, readFile, stat } from "node:fs/promises"; import path from "node:path"; -import { DEFAULT_VARIANT, cachedWindowsFor, findContainingWindow } from "report-to-video/build-video"; +import { DEFAULT_VARIANT, cachedWindowsFor, findContainingWindow } from "umtool-report-to-video/build-video"; /** * Where a build's per-entry segments live. @@ -22,8 +22,8 @@ const segmentDirs = (outDir) => [ path.join(outDir, DEFAULT_VARIANT, "segments"), path.join(outDir, "segments"), ]; -import { adjudicationGaps, ledgerTotals, unadjudicatedOf } from "report-to-video/ledger-totals"; -import { widen } from "report-to-video/resolve-windows"; +import { adjudicationGaps, ledgerTotals, unadjudicatedOf } from "umtool-report-to-video/ledger-totals"; +import { widen } from "umtool-report-to-video/resolve-windows"; // widen() and the cache's window naming are IMPORTED, never reimplemented. The // bench's "extend to sentence end" has to be the same function the CLI runs, or diff --git a/umtool/lib/report/driver.mjs b/umtool/lib/report/driver.mjs @@ -11,7 +11,7 @@ import path from "node:path"; /** Where the pipeline lives. One place, so a move is one edit. */ -export const PIPELINE_DIR = path.resolve(process.cwd(), "..", "scripts", "report-to-video"); +export const PIPELINE_DIR = path.resolve(process.cwd(), "report-to-video"); const script = (name) => path.join(PIPELINE_DIR, name); diff --git a/umtool/lib/report/export.mjs b/umtool/lib/report/export.mjs @@ -16,7 +16,7 @@ // Plain ESM, so `umtool export` runs from a terminal. import { readFile, readdir, stat } from "node:fs/promises"; import path from "node:path"; -import { DEFAULT_VARIANT, selectVariant, segmentOffsets } from "report-to-video/build-video"; +import { DEFAULT_VARIANT, selectVariant, segmentOffsets } from "umtool-report-to-video/build-video"; import { citeUrlFor, channelFor, readAvailability, readManifest } from "../projects/report.mjs"; export const EXPORT_FORMATS = ["toc-bbcode", "toc-markdown", "description", "chapters"]; diff --git a/umtool/lib/report/manifest.mjs b/umtool/lib/report/manifest.mjs @@ -28,7 +28,7 @@ import { SCOPE_CONFIDENCE, VALUE_KINDS, rolesGaps, -} from "report-to-video/ledger-totals"; +} from "umtool-report-to-video/ledger-totals"; // Its own write queue, not lib/state.ts's. // diff --git a/umtool/lib/report/serve.mjs b/umtool/lib/report/serve.mjs @@ -8,7 +8,7 @@ import path from "node:path"; import { REPORTS_ROOT, resolveInRoots } from "../paths.mjs"; import { walkProjects } from "../projects/walk.mjs"; -import { cachedWindowsFor } from "report-to-video/build-video"; +import { cachedWindowsFor } from "umtool-report-to-video/build-video"; import { clipsOf, readManifest } from "../projects/report.mjs"; export async function resolveClip(projectId, clipId) { diff --git a/umtool/package.json b/umtool/package.json @@ -17,7 +17,7 @@ "next": "16.2.3", "react": "19.2.4", "react-dom": "19.2.4", - "report-to-video": "workspace:*", + "umtool-report-to-video": "workspace:*", "tailwind-merge": "^3.6.0", "yt-dlp-transcript-common": "workspace:*" }, diff --git a/umtool/report-to-video/README.md b/umtool/report-to-video/README.md @@ -0,0 +1,786 @@ +# report-to-video + +Turns a cited sweep report into a video: the clips run in chronological order and +let the source speak for itself, with thin chrome carrying the citation and a +timeline of where you are. A video rendering of the reports we already write. + +Two scripts and a manifest: + +| file | lifetime | what it is | +|---|---|---| +| `build-video.mjs` | stable | manifest → mp4. Fetches clips, snaps cuts to silence, letterboxes them into the chrome, crossfades. | +| `render-cards.mjs` | stable | draws the timeline footer and marker, plus optional card stills. Imported by `build-video.mjs`. | +| `resolve-windows.mjs` | stable | widens clip windows from cue spans to whole sentences. Run once after authoring a manifest. | +| `<report>/video.manifest.json` | per report | the edit decision list. **This is the regeneration source of truth**, not the report. | + +``` +node umtool/report-to-video/resolve-windows.mjs ~/reports/<slug>/video.manifest.json --write +node umtool/report-to-video/build-video.mjs ~/reports/<slug>/video.manifest.json +``` + +Output lands in `<report dir>/out/`: `cards/`, `clips-raw/`, `segments/`, and the +finished `<slug>.mp4`. First worked example: `~/reports/ferret-rescue/`. + +## Why there is a manifest at all + +**A sweep report does not contain enough information to cut a video from.** Its +citations carry a single start second and nothing else — `momentUrl()` +(`common/lib/momentUrl.ts:106`) takes one `seconds` and floors it, and the MCP +`Snippet` type (`mcp/src/search.ts:41`) has no `end` field. There is no clip +length anywhere in a report. + +The end times do exist, they are just never rendered: every cue in +`transcripts/channels/<slug>/data/<id>/transcript.cues.json` is `{start, end, text}`. +So the manifest is built by matching each quote back to its covering cues and +recording the real window. That is also what handles ellipsis-joined citations — +a report quote like `"…" … "…"` is often two separate cue spans presented as one. + +The manifest additionally pins the provenance (share link, corpus handle, match +counts, the narrowing queries) so a rebuild months later is reproducible and the +video's own claims about its coverage can be checked. + +## Where a clip actually gets cut + +Three stages, because a cue span is the wrong answer twice over. + +1. **Cue span** — the raw window covering the quote, from `transcript.cues.json`. +2. **Sentence widening** (`resolve-windows.mjs`) — walk outward to the nearest cue + ending in `.`, `?` or `!`. A cue boundary is where the *caption line wrapped*, + so cutting there drops the run-up that makes a quote intelligible. + + Two asymmetries matter. The **start** takes a lead-in only if a real sentence + opening sits within `--max-lead` (8 s); otherwise it takes none, because a + half-sentence run-up is the irrelevant context you were trying to avoid, not + context. The **end** is never clamped to a budget — stopping partway through a + sentence is the exact mid-thought ending this removes — so `--max-tail` (12 s) + only bounds how far it looks before giving up and using the cue end. + + **Some uploads have no punctuation at all.** Older ASR in this corpus emits + unpunctuated cue text for whole videos, and sentence detection then has nothing + to find: widening degrades to the raw cue span at both ends. That is not a + silent failure you can ignore — it is what produced a clip opening mid-thought + on "higher than they can afford and because", and what let another clip's tail + run through its neighbour. For those videos, pick the window by reading the + cues and set it by hand; the de-overlap and silence passes still apply. + + **`lockStart` / `lockEnd` pin an edge** to exactly what the manifest says, and + neither widening nor de-overlap will move it. Reach for it when the utterance's + real trailing pause does not line up with its last cue's end — the snap picks + the *nearest* silence, and in speech over game audio the nearest one is often a + gap between syllables rather than the pause at the end of the thought. +3. **Silence snapping** (`build-video.mjs`) — a sentence boundary in the + *transcript* still isn't a boundary in the *audio*, so clips clip words in + half. Fetch `fetchPad` seconds wider than needed, run `silencedetect` over the + result, and move each cut to the nearest silence within `snapWindow`. Starts + land on a silence's END (just before speech resumes), ends on a silence's START + (just after speech stops). No silence close enough → keep the exact point; a + tight cut beats a cut in the wrong place. The build logs `start✓ end✓` per clip + so you can see which snapped. + + **The silence threshold is relative, and it has to be.** These are game + streams: the gaps between words are full of game audio and music — quiet, but + nowhere near silent. A fixed absolute threshold sits below the noise floor and + finds nothing. On one measured clip: mean volume −21 dB, **0** silences at + −32 dB, **25** at −26 dB. So each clip is measured with `volumedetect` first + and the threshold set `silenceRelDb` (default 6) below its own mean. If a rebuild + suddenly reports mostly `start– end–`, this is the knob. + +The trim happens during the burn-in encode, so snapping costs nothing extra. + +4. **De-overlap** (`resolve-windows.mjs`) — widening is per-clip and blind to its + neighbours, so two clips cut from the *same* video can end up overlapping, and + the overlap plays as the same footage twice. Any earlier clip whose tail runs + into a later clip's start is trimmed back to that start. This is not a rare + edge case: it fired on the first report, where a 2024 upload has **no + punctuation at all** in the relevant stretch, so sentence detection found + nothing and the tail ran the full budget straight through the next clip. + +## Regenerating and changing a video + +Everything is cached by content, so iteration is cheap: + +- **Reorder, drop or add clips** — edit `timeline`, re-run. Cached clips are not + refetched, so a re-cut costs an encode, not a download. +- **Change a clip's window** — edit `start`/`end`, re-run. A raw file's window is + in its name, so the cache is content-addressed; and since a request is satisfied + by any cached file that **contains** it, a nudge inside the existing pad costs + nothing at all. Only a window that escapes every cached file downloads again. +- **Preview one entry** — `--only <id>` builds a single segment and stops. +- **Work offline** — `--skip-fetch` fails loudly instead of downloading, so you + can confirm you are working entirely from cache. +- **Fetch one clip, wide** — `--fetch-only <id> --pad 20` puts a generous window + in the cache without building anything. This is what the umtool clip bench runs, + and containing-window reuse is what makes that fetch double as the build's cache. +- **Force a refetch** — delete `out/clips-raw/`, or pass `--no-reuse` to require an + exact-window file. +- **Iterate on the rail** — `--rail-only` re-runs only the rail chain over a cached + `out/<slug>.prerail.mp4`; `--preview <start> <dur>` does the same over a window. + `--no-rail` builds the cut without one. See + [The claim rail](#the-claim-rail-renderrail). + +Containing-window reuse was retrofitted, and the waste it removes is measurable: +`ferret-rescue/out/clips-raw` holds **31 files for 10 clips** because every window +edit before this downloaded the same material again — one source is there four +times over overlapping windows. The tightest containing file wins, not the widest, +because silence detection decodes the whole file and a 40 s file costs more than +the 24 s one that would also have done. + +After changing any window, re-run `resolve-windows.mjs --write` before building. +It is a **fixed point**: running it on an already-resolved manifest reports +`0 window(s) changed` and rewrites nothing. That property is load-bearing and was +not free — the manifest stores times rounded to 2 dp, so a value read back can sit +a hair below the cue end it came from, which lands the end lookup on the previous +cue and runs the search on to the *next* sentence. Left alone, every re-run grew +the same clip. Hence `EPS` in the end lookup and the 0.05 s deadband on applying a +change. + +## Driven from umtool + +The three CLIs are the source of truth and stay usable on their own; umtool drives +them rather than reimplementing them, so the UI and the terminal can never disagree +about a window, a format string or the Rumble retry. Three additions exist for that: + +- **`--progress ndjson`** — one JSON object per line instead of prose: + `start`, `card`, `clip`, `fetch`, `snap`, `segment`, `entry-failed`, `concat`, + `chapters`, `note`, `done`, `error`. The event set is exactly what was already + being printed; making it a format switch is what stops a wording change from + breaking the driver. +- **`--continue-on-error`** — record a failed entry and carry on. A dead source at + clip 14 of 19 otherwise throws away thirteen fetches already paid for. The run + still **refuses to concatenate** and exits non-zero: a finished file that quietly + lost a citation is worse than no file. +- **`buildVideo({manifestPath, opts, out, only, fetchOnly})`** is exported, and + `widen()` from `resolve-windows.mjs` already was. umtool imports `widen()` so the + bench's "extend to sentence end" is the CLI's own function, and **spawns** the + build — a 40-minute chain of yt-dlp and ffmpeg inside a request handler has no + cancellation story. + +`check-availability.mjs` is the fourth CLI and belongs at the *front* of a build: + +``` +node umtool/report-to-video/check-availability.mjs <manifest.json> +``` + +It runs `yt-dlp --simulate` once per distinct `(channel, video)` — no bytes +downloaded — classifies each failure (`deleted`, `private`, `restricted`, +`members-only`, `geo-blocked`, `no-cues`, `maybe_missing`), and writes +`out/availability.json` with a timestamp. This is the one fact about a manifest +that goes stale in *both* directions: a source can die after the manifest is +written, and a source annotated "gone" can come back. `no-cues` is called out +separately because it is a different bug — usually the Rumble two-ids trap, where +the manifest names the MCP video id while the cue file lives under the URL slug. + +## Manifest shape + +`timeline` is an ordered list; entries are `card` or `clip`. + +```jsonc +{ "type": "card", "id": "ch3", "style": "chapter", "seconds": 4.0, + "kicker": "March – November 2025", "heading": "Then: the county", + "sub": "Six months for the first approval" } + +{ "type": "clip", "id": "c04", "video": "uyz1_FIqIEk", + "start": 32989.56, "end": 32994.19, // cue-accurate, from transcript.cues.json + "cite": 32989, // the second shown in the attribution line + "quote": "The pre-application screening was approved by the county, dude." } +``` + +Two per-clip fields exist for compilations that span sources or need a hand-cut +window: + +- **`channel`** — the archived channel this clip's cue file lives under, overriding + `provenance.channelSlug`. A compilation about one person routinely spans several + mirror channels (`HasanAbiVODs` / `…VODs3` / `…VODsbackup`), and cue files are + keyed by channel, so a single manifest-wide slug cannot find them all. +- **`lock`** — exempt this clip from `resolve-windows`. Sentence-widening exists to + stop clips ending mid-thought, but that is exactly wrong when the author has + deliberately cut a quote short: a single cue often holds a whole paragraph, so + trimming to "I hate this country so much sometimes" and dropping the rest of the + sentence is an editorial decision that widening would silently undo. `lock` also + handles the reverse case — a clip whose lead-in would drag in seconds of some + *other* audio (a news package playing before the speaker starts). + +Clips also carry `section` and (auto-set) `sectionEnter`. Card styles — `title`, +`timeline`, `status`, `bullets`, `sources` — still work, but the ferret-rescue cut +uses none of them. `render` holds resolution, fps, fonts, palette and the knobs +(`fetchPad`, `snapWindow`, `silenceRelDb`, `transition`, `slideSeconds`, +`headerHeight`, `footerHeight`, and the optional `rail`); `provenance` holds the +sweep's scope and counts. Two further entry `type`s, `scroll` and `chart`, close a +cut off a top-level `ledger[]` — see [The claim rail](#the-claim-rail-renderrail). + +## Chrome, not cards + +**The ferret-rescue cut has no cards at all** — no title, no chapter breaks, no +closing slate. It is a cited timeline and nothing else: the clips run in +chronological order and the source material carries the argument. Cards remain +supported for reports that want them, but the default posture is that anything +drawn is an interruption which has to earn its place. + +Nothing is drawn *over* the picture either. The video is **letterboxed between** +thin chrome rather than overlaid by it: + +- **Header (`headerHeight`, 56 px).** The citation only — cleaned title · upload + date · timestamp. No quote: the clip is already saying it, and burning in a + transcription of speech you can hear is noise. +- **Footer (`footerHeight`, 100 px).** The timeline: one node per milestone, each + with a label and a month/year stamp beneath it. Drawn once by + `renderFooterAssets()`. +- **The marker slides.** On the first clip of each section (`sectionEnter`, set + automatically), the fill bar and the amber marker animate from the previous node + to the current one over `slideSeconds`. Everywhere else they hold position. The + motion is ffmpeg expressions on `crop`/`overlay`, so it costs nothing beyond the + encode that was happening anyway. + + > **This used to be half true.** Until 2026-08-19 only the marker moved. The + > fill bar was a `drawbox` whose width was `if(lt(t,0.9),…)` — but **`drawbox` + > has no time variable**: its `t` is the box *thickness*, and with `t=fill` + > that is effectively `INT_MAX`, so the comparison was always false and the + > width expression collapsed to its end value. (`drawbox=w='t*10':h=8:t=2` and + > `drawbox=w=20:h=8:t=2` produce an identical YAVG.) + > + > It was worse than "always full", because **`drawbox` reads `w=0` as *the + > input width***. Section 0's fill is 0 px, so every clip in the first section + > drew the bar across the **whole frame** — the progress track read 100 % + > complete on the opening clip of every cut that has a footer. Measured on + > `quartering-flagging-takedowns/n02`: 1708 accent px on the track row before, + > 0 after (the correct value), with the amber marker parked on node 0. + > + > It is now a `_bar.png` strip of `2·trackLen × 3` — accent on the left half, + > transparent on the right — translated under a fixed-width `crop`, which + > *does* evaluate `x` per frame. Verified on `ferret-rescue/c07`: 608 → 698 → + > 798 → 912 px across t = 0.0 … 0.9 s, then parked. + +Clips carry `section` (index into `timelineNodes`); nodes supply `label` and +`date`. Ordering clips chronologically is the author's job — the manifest plays in +the order it is written. + +A consequence worth knowing: 16:9 source into the reduced height leaves narrow +pillarbox bars. That is the price of never covering the picture, and it is why the +chrome is kept as thin as it is. + +**Both bands are optional, and turning them off is a real setting, not a hack.** +`footerHeight: 0` (or an empty `timelineNodes`) drops the timeline; `headerHeight: 0` +drops the citation line. With both at zero the clips fill the whole frame and +nothing is drawn at all — the filtergraph loses the overlays rather than compositing +invisible ones, and `renderFooterAssets` returns early instead of drawing PNGs +nobody uses. Two reasons this comes up: + +- A timeline footer only means something if the clips *are* a progression through + time. A cut ordered by argument rather than by date should not draw one. +- The header prints the **upload date of the archived copy**, which for a VOD-mirror + channel is often years after the stream (a Nov 2019 stream re-uploaded in Apr 2023 + reads `… November 6, 2019 … · 2023-04-06`). The title usually carries the true + date, so nothing is false, but on a cut spanning many re-uploads it reads badly. + +The `hasan-hate-america` cut runs with both off. Restoring them is a two-value edit. + +Note: commas inside an ffmpeg filter expression have to survive filtergraph +parsing — wrapping the expression in single quotes is what protects them. + +- **Segments crossfade** (`transition`, default 0.5 s). This forces a full + re-encode of the timeline via `xfade`/`acrossfade` — the concat demuxer can only + stream-copy hard cuts. Pass `--no-xfade` for a fast hard-cut build while + iterating; the last pass can add the transitions back. + +## Two cuts from one manifest (`--variant`) + +A sweep finds more claims than a cut can show footage for. There are two +defensible answers to that and they make different videos, so the manifest +describes both and one filter picks between them. + +| variant | output | ledger | what the viewer sees | +|---|---|---|---| +| `sourced` (default) | `out/<slug>.mp4` | claims with a clip | every row on screen has footage behind it | +| `full` | `out/<slug>-full.mp4` | every claim | the unclipped ones are stacked onto `ledger` cards | + +`selectVariant(manifest, variant)` runs **immediately after the manifest is read** +and is the whole mechanism. Three rules, in this order: + +1. a timeline entry tagged `"variant": "full"` survives only in that variant; +2. a claim survives only if the entry its `entryId` names survived — which is what + makes `sourced` a sourced-only ledger, because in `full` every claim is pinned, + to a clip or to a `ledger` card; +3. `card.variants[<name>]` field overrides are merged in and the key dropped. + +Nothing downstream learns about variants. `ledgerTotals`, the rail, the chart +band, `scheduleClaims`, the chapters and the closing cards already take the +ledger and the timeline as inputs. + +**Rule 3 exists because copy can be false in one cut.** A title card saying "48 +dated claims" is a lie in a cut that shows nineteen, and the closing sources card +says two claims stayed ambiguous — both of which happen to be unclipped, so in +`sourced` there are none. + +### Output layout + +``` +out/ + clips-raw/ SHARED — the only expensive thing in a build + availability.json SHARED — a fact about the manifest, not about a cut + <slug>.mp4 sourced + <slug>-full.mp4 full + sourced/{cards,segments,qr,chrome,schedule.json} + full/{cards,segments,qr,chrome,schedule.json} +``` + +`clips-raw` is shared deliberately: `sourced`'s clips are a subset of `full`'s, so +no clip is ever fetched twice. `sourced` writes `out/<slug>.mp4` because that is +the path umtool's build probe already looks for. + +`compose-chrome.mjs` and `verify-build.mjs` both take `--variant` for the same +reason: the band plots the ledger the cut carries, and verifying the whole +manifest against one variant's file would report a missing chapter for every +entry the other cut has. + +### The `ledger` entry type + +`full`'s answer to the claims no clip covers. One card per **run of consecutive +unclipped claims**, so 29 claims cost 13 cards and about 66 seconds. + +```jsonc +{ "type": "ledger", "id": "L10", "variant": "full", "seconds": 4.8, + "kicker": "November 2024", + "heading": "Ten and eight, named separately, in one breath", + "sub": "…", // optional + "claims": ["c12", "m14"] } // ledger ids, in ledger order +``` + +Each row draws `date · scope pill · why it is not footage · the quote · his +figure`, plus a right-hand **arithmetic column**: `media / coffee / publica` as +they stand, the implied total, the layer this claim just moved lit, and its delta. +The arithmetic is **read from `ledgerTotals`, never recomputed** — one walk, or +the card and the band disagree about the same sum. + +Rows **reveal in sequence** behind an opaque `pal.bg` rectangle walking down the +card: the rail curtain's device, exact because the card ground is flat. +`seconds` is derived (`2.2 + 1.3·rows`) rather than authored, because the rail +pins to the same clock — see below. + +**"Why it is not footage" comes from a probe, not from a hand-written kicker.** +`check-availability.mjs` now probes every **ledger** source as well as every clip +source, so a row says `source deleted`, `source unreachable` or `not clipped` +because `yt-dlp --simulate` said so on a recorded date. Several unclipped claims +are cut from videos this cut clips elsewhere, i.e. demonstrably live; saying "the +upload is gone" about one of those is the kind of error that discredits the whole +compilation. + +**Pins gain a within-segment offset.** A claim on a `ledger` card is pinned to its +own row's reveal (`ledgerRevealAt(r)`), not to the segment's mid-dissolve. +Otherwise four rail rows land on one frame, and the pin-order guard's strict +monotonicity breaks for no reason. In `full` this pins the rail almost exactly: +every claim has a segment, so `scheduleClaims` interpolates almost nothing. + +**`status` is retired.** It existed to quote a claim whose source had gone; a +`ledger` card does the same thing better, alongside the arithmetic the claim moves +and with the reason coming from the probe. + +## The claim rail (`render.rail`) + +A cut whose whole point is *which company a number was about* has a problem: the +dates and the figures are **spoken**, and shown only in the header's citation +line. A viewer can hear "nearly ten" three times without ever seeing that the +three refer to three different payrolls. + +`render.rail` adds a persistent **vertical ledger down the right edge**. It +appends one row per claim as the video runs, keeps a live per-company tally +beside it, and lists **every claim the sweep found** — not just the ones with a +clip behind them. Claims with no clip are dimmed (muted ink, hollow dot) and pass +with no audio; they are what stops the rail from implying the cut is the corpus. + +It is **entirely opt-in**. With no `render.rail` key the filtergraph is the one +that was there before, and output is byte-for-byte unchanged. + +```jsonc +"render": { + "rail": { + "width": 420, // picture shrinks to width - 420 + "rowHeight": 46, // one claim row; window height is a whole multiple + "tallyRowHeight": 44, + "tallyTop": 96, // y of the tally block inside the rail column + "pad": 22, + "slide": 0.55, // seconds per row-change ease + "rule": "#2A322F", + "tracks": [ // one per company; ORDER is the rail/legend order + { "key": "media", "label": "The Quartering · media", "color": "#22AB83" } + ] + } +} +``` + +and a top-level `ledger[]`, in **playback order** — each track's claims contiguous +and date-sorted within the track: + +```jsonc +{ "id": "m02", "date": "2022-09-17", "company": "media", + "value": 4, // null for a qualitative claim; the tally ignores those + "display": "4", // the badge + "label": "counts them out: one, two, three, four", + "quote": "…", "src": "…", + "hedged": false, // a hedge word, not a figure -> hollow dot on the chart + "plotted": true, // appears in the step chart + "entryId": "a01" } // pins the row to that timeline entry's segment +``` + +Pinned entries carry a `claim` back-reference so the link reads both ways. + +### How it is put together + +Everything that moves is **one tall strip walked by a fixed-size `crop`**, not a +per-state still, because swapping stills can only cut and a crop can ease. Five +strips, all bounded `-loop 1 -framerate <fps> -t <total+2>`: + +| strip | size | what it is | +|---|---|---| +| `_rail_chrome.png` | `RW × RHGT` | opaque panel, title, rules, **and the tally swatches and labels**. Runs the whole video — no `enable=` gates | +| `_rail_log.png` | `RW × N·rowHeight` | every claim, stacked, no padding | +| `_rail_curtain.png` | `RW × LOGH` | opaque `pal.bg` | +| `_rail_hl.png` | `RW × rowHeight` | the amber current-row marker | +| `_rail_tally.png` | `ΣlaneW × rows·cellH` | one COLUMN per lane — four rolling cells and the roster line | +| `_rail_qr.png` | `TILEW × segments·TILEH` | one provenance tile per segment | + +### The tally rolls one number at a time + +It used to be a column of four-row slabs walked by one `crop`: when the coffee +company's number changed, all four rows moved, and "The Quartering" slid up the +screen for a reason that had nothing to do with it. Text that has not changed +must not move. + +So the **swatch and the company label went into the static chrome**, and each +track got its own **rolling cell** — number, delta triangle, `as of <date>` and a +population chip, right-aligned in a ~170 px column. All the lanes live side by +side in **one** PNG, so it is still one input: five `crop`s at different `x`, five +overlays. + +**The direction of a roll is decided by the strip's LAYOUT, not by the ramp.** + +| | rows | the crop | what you see | +|---|---|---|---| +| rise | `[old, new]` | walks **down** | content moves **up** | +| fall | `[new, old]` | walks **up** | content moves **down** | + +Between transitions a one-frame `gte()` step repositions to the next pair's +starting row. That step is invisible **because both endpoint rows hold identical +content** — which is why every pair repeats the value it starts from instead of +sharing a row with its neighbour, and why the delta chip rides on both rows and +therefore stays on screen until the next change. + +``` +y_j(t) = r_j0 + Σ_k [ (a_k − b_{k−1})·gte(t,t_k) + (b_k − a_k)·ease(t_k) ] +``` + +Every term is cumulative and saturating — the rail's hard rule. + +A repeated identical figure still rolls, upward: he said it again on a new date, +and the `as of` line underneath is what changed. + +### The roster line + +A fifth lane under the tally, `2 editors · 1 designer`, in `pal.muted`. It moves +**only when the rendered line changes**, which in this corpus means it stands +still through October and December 2023 while the total above it goes from three +to four. That is the finding, drawn rather than asserted. + +It comes from an optional `roles` field on a ledger claim, and `rosterAt()` / +`rosterLine()` in `ledger-totals.mjs` are the one implementation, because a +chapter card states the same thing in words. + +```jsonc +"roles": [{ "role": "video editor", "count": 2, "verbatim": "two video editors" }] +``` + +`verbatim` is his words; `role` and `count` are our reading. `roles` is **not** one +of the six adjudication fields — most claims are a number and nothing else, and +gating the inbox on a field a handful of entries can carry would leave it +permanently red. + +### The QR moved into the rail's foot + +It used to float over the bottom-right of the **picture**, which is the one part +of the frame this cut promises never to draw on. It is now a bordered tile parked +at the foot of the rail column — `pal.accent` rule, `SCAN → JERALYZER` above the +code, what it opens below it — and one more strip: one tile per segment, +crop-walked with **instantaneous `gte()` steps** at segment mid-dissolves. A code +that eased into place would spend the ease unscannable. + +The tile overlays **after the curtain**, which is what stops the parked curtain +painting over it. + +> **A card cannot carry the report's share link.** That link carries all 23 +> channel filters and is ~1.4 k characters: a version-40 symbol, 177 modules in a +> 132 px tile, about 0.7 px per module. Cards get `provenance.qrLink` — the same +> query without the channel list, ~200 chars, 63 modules, verified scannable at +> this size — and fall back to `provenance.siteOrigin`. Manifests with **no** +> rail keep the old per-clip overlay in `buildClipSegment`, byte for byte. + +> **`railGeometry` used to derive its height from `render.footerHeight`.** With +> the chart band on, the band reserves 200 px and `footerHeight` says 100, so the +> rail column ran a hundred pixels — about two rows of its log window — past the +> line every other renderer letterboxes to. It reads `reservedFooterHeight(render)` +> now, the same fix the closing cards needed for the same reason. + +**The curtain is why one strip is enough.** With the window parked at the top +while the list is still filling, rows `i+1 … K-1` would show claims the video has +not made yet. The curtain is an opaque rectangle riding just below the last +revealed row; once the list is full it parks exactly one window-height down, +which is the bottom of the rail column — permanently outside the window. It is +`pal.bg` precisely so that parking there is invisible against the footer band. +Curtain and log **must share the same eased `P`**, or the curtain lags the rows +mid-slide and unrevealed claims flash into view. + +**Ramps are cumulative and saturating, never gated.** A piecewise sum of +`gte(t,sᵢ)·lt(t,sᵢ₊₁)·…` terms flashes to `y=0` for one frame at any boundary +gap, because every gate evaluates false at once and the sum collapses. Terms that +rise to their delta and stay cannot do that. + +**Scheduling.** Playback is ONE chronology across every company, and the ledger is +sorted the same way, so a claim's position in the rail *is* its position in time. +A claim with a clip behind it is pinned to that clip's segment; a claim on a +`ledger` card is pinned to its own row's reveal; the rest are spread evenly +between their neighbouring pins. A pin that runs backwards is refused — the +ledger and the timeline disagreeing about the order of events is a manifest bug, +and the whole cut rests on the two agreeing. + +Segment-level pins land at **`starts[i] + D/2`** — mid-dissolve, where the picture +is already crossfading and a ±3-frame error is invisible. + +The chain attaches **after the last `xfade` node**, inside the concat pass. `t` +there is absolute and continuous from 0, and nothing downstream of the last xfade +is dissolved — so it already has post-pass semantics without a second encode, +which would re-quantize crf-20 output at exactly the content that hurts most +(antialiased text on flat colour). + +### `--rail-only` and `--preview` + +`--rail-only` re-runs just the rail over a cached `out/<slug>.prerail.mp4`, +building that file from the existing segments the first time. Seconds instead of +the full concat. The hard-cut base is a separate file +(`<slug>.prerail-hardcut.mp4`), because the two timelines are different lengths +and a cached base from the wrong mode is a stale-cache trap the length assertion +would otherwise have to explain. + +**It is mandatory for `--no-xfade`**, not an optimisation: `concatHardCut` is +`-c copy` and a stream-copy mux cannot host a filtergraph at all. That path +concats to `.prerail.mp4`, asserts its length against `segmentOffsets().total`, +and then applies the rail. + +`--preview <start> <dur>` renders a window. `-ss` restarts `t` near zero, which +would put every absolute-time ramp in the wrong place — so the preview path +inserts `setpts=PTS+<start>/TB` before the rail chain and rebases afterwards. +Getting that wrong makes a working rail look broken. + +## The end sequence: `scroll` and `chart` + +Two timeline entry `type`s that exist to close a cut, both driven off the same +`ledger[]`: + +- **`scroll`** — the whole ledger as one tall PNG, walked by an animated `crop`. + **One chronological line, a column per company**: it used to group by company, + which re-told the cut's own order backwards and hid the only thing worth seeing + there — that the four payrolls were being described in the same weeks. Company + is read from COLUMN POSITION, so colour is the secondary encoding. + `hold` (default 2 s) buys a still moment at both ends; + `clip()` in the expression provides it for free, and crop's own clamping + degrades an off-by-a-few content height into a static last frame, not an error. +- **`chart`** — the four-series step chart over the `plotted` claims, authored as + SVG and rasterized with `rsvg-convert` (deterministic about output size in a + way ImageMagick's RSVG delegate is not). Fonts inside the SVG resolve through + **fontconfig, not `render.fontRegular`** — use the family name the Pango cards + use. + +The chart's wipe **cannot** be `crop=w='<ramp>'`: crop's `w` is config-time and +`t` is undefined there. It is a curtain instead — an opaque `pal.bg` rectangle +slid rightwards off the plot, which is exact because the card ground is flat. + +**Colour is not the only encoding on that chart, and that is a requirement.** No +four-colour categorical palette clears the data-viz all-pairs CVD gate (three +slots is the documented ceiling), so every series also carries a distinct dash +pattern and a direct end-of-line label with a leader elbow. The four hues are the +published artifact's, re-validated against this video's darker ground (`#0F1312`) +on the *adjacent* pairlist — the pairlist for line charts — where all five checks +pass (worst adjacent CVD ΔE 10.2 against a ≥8 target; normal-vision ΔE 17.4 +against a ≥15 floor). + +**`hideRail: true`** on a closing card slides the whole rail column off to the +right over that card's dissolve — one offset expression shared by every rail +overlay, so the column moves as one object — and lets the card render at the full +`width` instead of `contentWidth`. Not an `enable=` pop: a column that vanishes +between two frames reads as a dropped frame. The slide is cumulative and +saturating like every other ramp, so the rail does not come back; every card +after the first `hideRail` one should carry the flag too, or it lays out inside a +content width whose rail is no longer there. + +Both kinds go through the same `fps=,setsar=1` and the same `encodeArgs` as every +other segment. They have to: `xfade` rejects a mismatched link with *"First input +link parameters do not match"*, and that surfaces at concat time, after every +fetch has been paid for. + +## The ledger is adjudicated, and both totals depend on it + +`ledger[].company` used to be an **undocumented interpretation**, and four +different hazards were riding on it: + +| hazard | example | why it mattered | +|---|---|---| +| **scope ambiguity** | *"I have 10 employees, my coffee company employees… my editors"* | all-companies or coffee-only, depending on where the comma falls | +| **derived, not stated** | *"10 at coffee brand coffee, I've got eight staff for the live stream"* recorded as **18** | he never says 18 | +| **population drift** | 5 *"full-time salaried"*, 10 *"employees"*, 10 *"all basically contractors"* | different denominators, one series | +| **synthetic values** | 10.5 for *"about 10 people, 11 people"* | a midpoint we invented and attributed to him | + +So every ledger entry now carries six adjudicated fields — `scope`, +`scopeBasis`, `scopeConfidence`, `population`, `valueKind`, `flags` — settled +against **±90 s of surrounding context, never the quote alone**. A first-person +quote is routinely the host reading someone else's words or being sarcastic, and +neither is visible inside the quote. One claim in this corpus is a guest's +payroll rather than his, and it reads identically until you listen either side. + +`scopeConfidence: "unresolved"` is a legitimate outcome and **feeds neither +total**. + +**The rule that follows:** the *stated* series may contain only **a figure he +utters as a single number for a named scope**. Sums and midpoints are ours, and +live in the *implied* series, which says so on screen. + +`ledger-totals.mjs` is the one implementation of that arithmetic — the umtool +inbox, the chart band and the closing card all import it, so none of them can +disagree. It **refuses to run on an unadjudicated ledger**, because both totals +lie if you act on one. **Six** named predicates compute incoherence rather than +asserting it (`contradicts_component`, `same_day_conflict`, `self_negating`, +`population_mismatch`, `not_his_number`, `status_flip`). Deliberately **not** a +predicate: a large rise or fall between claims. Fluctuation is the subject, not a +defect. + +`status_flip` is the sixth: the same people described as staff and then as +contractors, or the reverse. Three details in it are load-bearing. + +- **`employees` and `people` are in NEITHER camp.** They are what he says when he + is not making a claim about status at all, and reading them as one side or the + other manufactures a reversal out of a change of vocabulary. +- **An `all` claim is comparable with any company; two companies are not + comparable with each other.** Without that asymmetry the corpus's clearest + reversal is invisible: December 2024's ten are the *channel's*, May 2025's ten + or eleven are *everything's*. +- **Only against the most recent comparable claim that carries a camp.** Fire on + every earlier pair and one 2022 "all 1099 and not full-time" flags each of the + next seven claims in turn — seven findings where there is one. Bounded this + way it fires at the TRANSITIONS, which is what a flip-flop is. + +It is not gated on the claim having a figure. *"That's why all my workers are +contract workers"* names no number and is the single clearest status claim here. + +Work the adjudication in umtool, at `/browse/<project>/claim/<id>`. Sign-off is +"the inbox is empty": `claim-unadjudicated` is **blocking**. + +## `render.chromeEngine: "hyperframes"` — the chart band + +Opt-in, and absent it the ffmpeg chrome path is byte-for-byte unchanged. + +The chart band **replaces** the footer node track, which only moved at section +handovers — precisely the fault it exists to fix. It takes the footer's ground +and 100 px more, and the picture loses that height (1500×924 → 1500×824). + +`compose-chrome.mjs` emits a HyperFrames project per region and renders it to a +**lossless RGBA PNG sequence**; `build-video.mjs` overlays the frames. Three +things about that are load-bearing: + +- **Footage never enters Chrome.** HyperFrames pre-extracts source video to JPEG + q95, which is unacceptable when the picture *is* the cited evidence. Only + chrome is composed there — about 37 % of full-frame pixels rather than 100 %. +- **The PNG regions overlay BEFORE the rail chain, not after.** The rail chain + ends in `format=yuv420p`, and overlaying an alpha sequence onto yuv420p is the + same alpha-subsampling trap the rail already documents, one layer later. +- **The playhead is driven by `out/schedule.json`**, which the build writes. + Recomputing claim times here would be a second implementation of + `segmentOffsets()` and would drift the first time `transition` changed. Same + rule as `widen()`: imported, never reimplemented. + +One clip-path sweeps the whole plot rather than a `stroke-dashoffset` per series. +The obvious build animates each path's dash offset, and it looks right for the +strokes and wrong for everything else: the gap band between the two totals is a +filled polygon with no stroke to offset, so it appears whole the moment it fades +in and the chart is showing an answer the playhead has not reached. + +**The flag colour is not the palette's amber.** `#E8A33F` sits at ΔE 12.4 from +the coffee series' `#D2732F` at *normal* vision — below the 15 floor — so a flag +badge beside a coffee mark was hard to tell from the coffee mark. `#E0E24A` +replaces it and adds **no new worst pair**: the worst CVD pair +(`#C55F9C↔#22AB83`, ΔE 5.4 deutan) and the worst normal-vision pair +(`#C55F9C↔#D2732F`, ΔE 16.3) are identical with and without it. The implied +total's `#EDF0EC` fails the categorical lightness and chroma checks *by design* — +it is an aggregate, not a categorical peer, so it is encoded by weight and +consumes no palette slot. + +## Things that cost time to find out + +**yt-dlp picks VP9 at `height<=720`, and that is a trap.** `--download-sections` +combined with `--force-keyframes-at-cuts` re-encodes, so a VP9 pick means +libvpx-vp9 — 27 seconds to cut a 5-second clip. It also writes a `.webm` and +appends that extension to whatever `-o` you gave, so the file never lands where +you asked and the run fails looking for it. Pin H.264/AAC in mp4 and the same cut +takes ~12 seconds. `build-video.mjs` does this and keeps a rename fallback for +the case where a fallback format still forces another container. + +**`--force-keyframes-at-cuts` is not optional here.** Without it the cut snaps to +the nearest preceding keyframe and can start seconds early. That is fine for a +human scrubbing a VOD; it is not fine when the clip *is* the citation. +`PlayerProvider.tsx:478` builds the copyable clip command without this flag — +correct for its purpose, wrong for ours. + +**`--ignore-config` is mandatory.** The operator's own yt-dlp config redirects +output to `~/Podcasts` and attaches thumbnail/metadata post-processors. Every +managed yt-dlp call in this repo passes `--ignore-config` for the same reason. + +**yt-dlp exit 101 is success**, not failure — it means a clean early stop. The +repo encodes this at `common/ytdlp/downloadOneManaged.ts:355`; a new caller has to +replicate it. + +**Local media will not help you.** Only 5 of 1,804 PirateSoftware video dirs hold +any media at all, and none are ones a report is likely to cite. Clips are a +network fetch. `metadata.info.json` and `transcript.cues.json` *are* present for +every video, so titles, dates, durations, webpage URLs and cue timings all come +from disk with no probe. + +**Upstream availability is load-bearing.** A clip can only be fetched while the +source is still up. Check `platform state` via `get_video_metadata` before +committing to a clip — a `deleted`/`maybe_missing` source needs a quote card +instead of footage. The archive outlives its sources, so a video built from an +old report will be *less* complete than the report unless this is handled +deliberately. + +**ImageMagick's `-size` leaks into the Pango group.** `-size 1920x1080 xc:BG` +followed by `( ... pango:@file )` renders the text into a full-frame box, which +pins it to the top and wraps at the frame edge instead of the text column. Reset +`-size` inside the parens. + +**ffmpeg `drawtext` does not wrap and hates punctuation.** Both are solved the +same way: wrap to a column count in JS, write to a file, and use +`textfile=`. Nothing then needs escaping. Stream titles also need emoji and +`!command` suffixes stripped or they render as tofu in the attribution line. + +**Segments are encoded to identical parameters on purpose** so the final +concatenation is a stream copy via the concat demuxer. Mismatched streams are the +usual reason a naive concat produces a broken or audio-desynced file. + +## Not done yet + +- **Narration is silent by design.** Cards carry the connective text; the only + audio is the clips'. A TTS layer would attach per card (`seconds` already gives + it a duration to fill) — deliberately deferred rather than designed out. +- **Snapping is silence-based, not word-based.** It finds gaps in the audio, which + is usually the same thing as a word boundary but is not guaranteed to be — + a speaker who does not pause gets the unsnapped cut. Forced alignment against + the transcript would be exact; `silencedetect` is a tenth of the work and + handles the cases that were actually audible. +- **Only the chart band is a HyperFrames region.** `chromeRegions()` returns one + entry. The rail and the header are still drawn by the ffmpeg chain, and porting + them is the rest of the job — the rail needs the ghost/hop convention (a row + with no clip fades in dimmed under a dashed rule and the amber highlight *hops + over* it to the next cited row) which the strip builders cannot express. Until + then `railFilterChain` and `renderFooterAssets` stay; they must not be retired + on the strength of the band alone. +- **`timelineNodes` / `section` / `sectionEnter` are still live.** They lose their + only consumer when the ffmpeg footer goes, not when the band arrives — so they + retire with `renderFooterAssets`, in that same commit. +- **Manifests are written by hand** from verified cue data. Deriving a first-draft + manifest automatically from a report's citations is the obvious next step; the + report parse is straightforward (`> "quote"` followed by + `— [title @ h:mm:ss](…?v=slug%2Fid&t=sec)`), the cue-matching is the real work. diff --git a/umtool/report-to-video/build-video.mjs b/umtool/report-to-video/build-video.mjs @@ -0,0 +1,1838 @@ +#!/usr/bin/env node +// build-video.mjs — render a cited sweep report into a narrated-by-text video. +// +// Takes a video manifest (see README.md next to this file) and produces one mp4: +// text cards state the findings, clips let the source say it in their own voice, +// and every clip carries a burned-in quote plus its attribution. +// +// Pipeline, per manifest entry: +// card -> still PNG (render-cards.mjs) -> N seconds of video + silent audio +// clip -> yt-dlp --download-sections (WIDE) -> silence-snap -> trim + burn +// then the segments are crossfaded together into the finished file. +// +// Three things worth knowing about how clips are cut: +// +// 1. Windows come from the manifest as absolute [start, end] seconds, derived +// from transcript.cues.json (which carries an END per cue). A sweep report +// only ever records a single start second, so windows cannot be recovered +// from the report alone. +// 2. Those windows are widened to sentence boundaries by resolve-windows.mjs, +// so a clip carries the run-up that makes the quote make sense. +// 3. A cue boundary is still not a *speech* boundary — cutting there clips +// words in half. So we fetch wider than needed and snap the real cut to a +// silence found in the audio. That is what makes clips start and end +// between words rather than through them. +// +// Fetched clips are cached by (video, start, end); re-running is cheap and only +// changed entries re-download. Delete out/clips-raw to force a refetch. +// +// In the app: not used. On the CLI: +// node umtool/report-to-video/build-video.mjs <manifest.json> [options] +// +// Options: +// --out <dir> Output root (default: manifest dir + /out) +// --variant <name> Which cut to build (sourced | full; default sourced) +// --skip-fetch Fail instead of downloading anything not already cached +// --only <id> Build a single entry's segment and stop (for iterating) +// --no-xfade Hard cuts instead of crossfades (much faster; concat copy) +// --progress ndjson One JSON event per line instead of prose (for umtool) +// --continue-on-error Record a failed entry and carry on, instead of aborting +// --fetch-only <id> Fetch one clip's window into clips-raw and stop +// --pad <s> Override render.fetchPad (the clip bench fetches wide) +// --site-origin <url> Archive to read cue windows from when there is no local +// corpus (defaults to the manifest's provenance.siteOrigin) +// --resolve-site-ids On a published-id miss, find the record by scanning the +// channel's shards. Slow; see cues.mjs. +// --cue-source <which> auto (default) | local | http. The two can disagree +// once a corpus moves past its last publish — see cues.mjs. +// --no-rail Skip the claim rail even when the manifest configures one +// --rail-only Re-run just the rail over out/<slug>.prerail.mp4 +// --preview <s> <d> Rail-only, over a <d>-second window starting at <s> +// +// Requires: yt-dlp, ffmpeg/ffprobe, ImageMagick with Pango. + +import { execFile } from "node:child_process"; +import { promisify } from "node:util"; +import { mkdir, writeFile, readFile, access, readdir, rename } from "node:fs/promises"; +import path from "node:path"; + +import { + renderCard, renderFooterAssets, renderRailAssets, renderScrollCard, renderChartCard, + renderLedgerCard, ledgerRevealAt, ledgerSeconds, + cardWidth, contentWidth, reservedFooterHeight, +} from "./render-cards.mjs"; +import { createCueSource, siteOriginFromManifest } from "./cues.mjs"; + +const execFileP = promisify(execFile); + +const YTDLP = process.env.YTDLP_BIN ?? "yt-dlp"; +const FFMPEG = process.env.FFMPEG_BIN ?? "ffmpeg"; +const FFPROBE = process.env.FFPROBE_BIN ?? "ffprobe"; +const QRENCODE = process.env.QRENCODE_BIN ?? "qrencode"; + +// Cue windows and per-video metadata come from a local corpus when there is one +// and from the published archive otherwise, so this runs in a clone with no +// `transcripts/` directory. Built once main() has the manifest (it carries the +// archive origin); see cues.mjs. +let CUES = null; + +const exists = (p) => access(p).then(() => true, () => false); + +// ---- variants ------------------------------------------------------------ +// ONE manifest, two cuts, one filter, applied once. +// +// The question the two variants answer differently is what to do with a claim +// the sweep found but no clip covers. `sourced` refuses to put it on screen at +// all -- every row the viewer sees has footage behind it. `full` gives each one +// a slot on a stacked ledger card, so nothing is dropped and the arithmetic of +// each layer is shown rather than asserted. +// +// Both end on the same three numbers. That is the point of shipping both: if +// the totals moved when the unsourced rows came off, the thesis would rest on +// rows nobody can check. +// +// The filter runs IMMEDIATELY after the manifest is read, and nothing +// downstream learns about variants. `ledgerTotals`, the rail, the chart band, +// `scheduleClaims`, the chapters and the scroll already take the ledger and the +// timeline as inputs, so selecting is the whole of the mechanism. +export const VARIANTS = ["sourced", "full"]; +/** The cut a caller means when it does not say. `out/<slug>.mp4`. */ +export const DEFAULT_VARIANT = "sourced"; + +/** + * The manifest as one variant sees it. + * + * Three things happen, in this order: + * + * 1. A timeline entry tagged `variant` survives only in that variant. (The + * stacked ledger cards are `variant: "full"`.) + * 2. A claim survives only if the entry it is pinned to survived. That single + * rule is what makes `sourced` a sourced-only ledger: in `full` every claim + * is pinned -- to a clip or to a ledger card -- so nothing is dropped. + * 3. `card.variants[<name>]` field overrides are merged in. The title and + * sources cards have to state their own scope honestly, and "50 dated + * claims" is simply false in `sourced`. + */ +export function selectVariant(manifest, variant = DEFAULT_VARIANT) { + if (!VARIANTS.includes(variant)) { + throw new Error(`unknown variant \`${variant}\` — one of ${VARIANTS.join(", ")}`); + } + const timeline = (manifest.timeline ?? []) + .filter((e) => !e.variant || e.variant === variant) + .map((e) => { + if (!e.variants) return e; + const { variants, ...rest } = e; + return { ...rest, ...(variants[variant] ?? {}) }; + }); + const kept = new Set(timeline.map((e) => e.id)); + const ledger = (manifest.ledger ?? []).filter((c) => c.entryId && kept.has(c.entryId)); + return { ...manifest, variant, timeline, ledger }; +} + +/** + * Where a variant's own working files live. + * + * `clips-raw` stays at the ROOT and is shared: it holds the only expensive + * thing in the build (network fetches), and `sourced`'s clips are a subset of + * `full`'s, so a shared cache means no clip is ever fetched twice. Everything + * else is per-variant, because every one of them differs between the two cuts. + * + * `sourced` writes `out/<slug>.mp4` -- the path umtool's build probe already + * looks for -- and `full` writes `out/<slug>-full.mp4` beside it. + */ +export function variantPaths(outRoot, slug, variant) { + return { + root: outRoot, + dir: path.join(outRoot, variant), + rawDir: path.join(outRoot, "clips-raw"), + final: path.join(outRoot, variant === "sourced" ? `${slug}.mp4` : `${slug}-${variant}.mp4`), + }; +} + +// ---- progress protocol --------------------------------------------------- +// This has two audiences: a human watching a terminal, and umtool's build driver +// reading the pipe. Rather than have the driver scrape prose (which would make +// every wording change a breaking change), `--progress ndjson` switches every +// line to one JSON object. The event set is exactly what was already being +// printed -- this is a formatting switch, not new instrumentation. +// +// Events: start, card, clip, fetch, snap, segment, entry-failed, concat, +// chapters, note, done. +const HUMAN = { + start: (e) => `${e.title} — ${e.entries} entr(ies)`, + card: (e) => `card ${e.id}`, + clip: (e) => `clip ${e.id} (${e.video}) §${e.section}${e.sectionEnter ? " ⟶" : ""}`, + fetch: (e) => + e.reuse + ? ` fetch ${e.id}: ${e.reuse} already covers ${hms(e.from)}–${hms(e.to)} — no download` + : e.cached + ? null + : ` fetch ${e.id}: ${e.video} ${hms(e.from)}–${hms(e.to)}`, + snap: (e) => + ` snap ${e.id}: ${e.start ? "start✓" : "start–"} ${e.end ? "end✓" : "end–"} ` + + `(${Number(e.seconds).toFixed(1)}s)`, + segment: () => null, + "entry-failed": (e) => ` ** ${e.id} failed: ${e.message}`, + concat: (e) => `${e.mode === "xfade" ? "crossfading" : "hard-cutting"} ${e.n} segments…`, + chapters: (e) => `chapters: ${e.n} marker(s) -> ${e.file}`, + note: (e) => e.message, + done: (e) => + e.duration === undefined + ? `built ${e.out}` + : `\n${e.out}\nduration=${e.duration}\nsize=${e.size}`, +}; + +let EMIT = (ev, fields = {}) => { + const line = HUMAN[ev]?.({ ev, ...fields }); + if (line) console.log(line); +}; + +export function setProgressMode(mode) { + EMIT = + mode === "ndjson" + ? (ev, fields = {}) => process.stdout.write(JSON.stringify({ ev, ...fields }) + "\n") + : (ev, fields = {}) => { + const line = HUMAN[ev]?.({ ev, ...fields }); + if (line) console.log(line); + }; +} + +function hms(total) { + const s = Math.floor(total); + const h = Math.floor(s / 3600); + const m = Math.floor((s % 3600) / 60); + const sec = s % 60; + return h > 0 + ? `${h}:${String(m).padStart(2, "0")}:${String(sec).padStart(2, "0")}` + : `${m}:${String(sec).padStart(2, "0")}`; +} + +// Stream titles here are full of emoji and !commands. drawtext renders them as +// tofu with a text font, and they add nothing to an attribution line. +function cleanTitle(title) { + return title + .replace(/[\u{1F000}-\u{1FFFF}\u{2600}-\u{27BF}\u{FE0F}]/gu, "") + .replace(/\s*[!@]\S+/g, "") + .replace(/\s{2,}/g, " ") + .replace(/[\s·|-]+$/, "") + .trim(); +} + +// drawtext does not wrap. Break to a character budget, write to a file, and use +// textfile= so nothing needs shell or filter escaping. +function wrap(text, cols) { + const words = text.split(/\s+/); + const lines = []; + let line = ""; + for (const w of words) { + if (line && (line + " " + w).length > cols) { + lines.push(line); + line = w; + } else { + line = line ? line + " " + w : w; + } + } + if (line) lines.push(line); + return lines.join("\n"); +} + +// The published shard record carries the same fields as a local cue file, so this +// reads identically whichever source answered. +async function videoMeta(videoId, channelSlug, hints = {}) { + const d = await CUES.load(channelSlug, videoId, hints); + return { title: d.title, uploadDate: d.uploadDate, webpageUrl: d.webpageUrl, duration: d.duration }; +} + +// The CONTAINER's duration is max(video, audio), and the audio is longer: the +// AAC encoder pads the front with ~21 ms of decoder delay, and a video duration +// is rarely an exact multiple of the frame interval. Either way the excess is +// small — and it ACCUMULATES through segmentOffsets, which subtracts one +// transition per segment and hands the result to xfade, the chapter marks and +// (now) the rail. A few hundred ms of drift by segment 20 is enough to land a +// rail row-change on the wrong side of a cut. +// +// The video stream's frame COUNT is the number the timeline actually runs on, +// so derive the duration from it. nb_frames is absent on some demuxers; fall +// back to the container rather than failing a build over a probe. +async function probeDuration(file, fps) { + if (fps) { + const { stdout } = await execFileP(FFPROBE, [ + "-v", "error", "-select_streams", "v:0", "-show_entries", "stream=nb_frames", + "-of", "default=nw=1:nk=1", file, + ]); + const n = Number(stdout.trim()); + if (Number.isFinite(n) && n > 0) return n / fps; + } + const { stdout } = await execFileP(FFPROBE, [ + "-v", "error", "-show_entries", "format=duration", + "-of", "default=nw=1:nk=1", file, + ]); + return Number(stdout.trim()); +} + +// yt-dlp exits 101 on a clean early stop (break-on-existing / max-downloads). +// The repo treats that as success everywhere else; do the same here. +const ytdlpOk = (err) => err?.code === 101; + +// ---- the clip cache ------------------------------------------------------ +// A raw clip's window is IN ITS NAME, which makes the file immutable and the +// cache content-addressed. The original lookup was for the exact name, so any +// change to a window -- a hand edit, a widen, a nudge in the clip bench -- was a +// fresh download of material already on disk. Measured on ferret-rescue: 31 +// files for 10 clips, one source fetched four times over overlapping windows. +// +// So: satisfy a request from ANY cached file that contains it. The TIGHTEST +// container wins, because detectSilence decodes the whole file and a 40s file +// costs more than the 14s one that would also have done. The clip bench fetches +// deliberately wide, and this is what makes that generous fetch become the +// build's cache rather than a second one. +const WINDOW_RE = /^(\d+(?:\.\d+)?)-(\d+(?:\.\d+)?)$/; + +// A window read back from a 2 dp manifest can sit a hair outside the file that +// produced it; the same tolerance resolve-windows.mjs uses for the same reason. +const WIN_EPS = 0.02; + +export async function cachedWindowsFor(rawDir, video) { + let names; + try { + names = await readdir(rawDir); + } catch { + return []; + } + const prefix = `${video}_`; + const out = []; + for (const name of names) { + if (!name.startsWith(prefix) || !name.endsWith(".mp4")) continue; + // The remainder must be exactly `a-b`, which is what stops a video id that + // is a prefix of another (or one containing `_`) from claiming its files. + const m = WINDOW_RE.exec(name.slice(prefix.length, -4)); + if (!m) continue; + out.push({ name, path: path.join(rawDir, name), from: Number(m[1]), to: Number(m[2]) }); + } + return out; +} + +/** The tightest cached file containing [from, to], or null. */ +export async function findContainingWindow(rawDir, video, from, to) { + const windows = await cachedWindowsFor(rawDir, video); + let best = null; + for (const w of windows) { + if (w.from > from + WIN_EPS || w.to < to - WIN_EPS) continue; + if (!best || w.to - w.from < best.to - best.from) best = w; + } + return best; +} + +async function fetchClip(entry, meta, render, rawDir, opts) { + // Deliberately over-fetch: the snapping pass below needs room on both sides to + // find a silence, and a clip that has no slack can only be cut where the cue + // happened to break — which is what put words in half in the first place. + const pad = opts.pad ?? render.fetchPad ?? 3.0; + const from = Math.max(0, entry.start - pad); + const to = entry.end + pad; + + // Shared across variants, and deliberately so: this is the only expensive + // thing in a build, and the two cuts overlap almost entirely. + const name = `${entry.video}_${from.toFixed(2)}-${to.toFixed(2)}.mp4`; + const dest = path.join(rawDir, name); + if (await exists(dest)) { + EMIT("fetch", { id: entry.id, video: entry.video, from, to, cached: true }); + return { path: dest, fetchStart: from, cached: true }; + } + if (!opts.noReuse) { + const hit = await findContainingWindow(rawDir, entry.video, from, to); + if (hit) { + EMIT("fetch", { + id: entry.id, video: entry.video, from, to, cached: true, reuse: hit.name, + }); + // fetchStart is the CACHED file's start, not the requested one -- every cut + // downstream is expressed relative to it, so reuse is transparent. + return { path: hit.path, fetchStart: hit.from, cached: true }; + } + } + if (opts.skipFetch) throw new Error(`--skip-fetch set and no cached window covers ${name}`); + + const maxH = render.maxHeightSource; + const fmt = [ + `bv*[vcodec^=avc1][height<=${maxH}]+ba[acodec^=mp4a]`, + `bv*[ext=mp4][height<=${maxH}]+ba[ext=m4a]`, + `b[ext=mp4][height<=${maxH}]`, + `b[height<=${maxH}]`, + ].join("/"); + + const argsWith = (extra) => [ + // The operator's own yt-dlp config redirects output and attaches thumbnail + // and metadata post-processors; without this the clips land elsewhere. + "--ignore-config", + "--no-playlist", + "--download-sections", `*${from.toFixed(2)}-${to.toFixed(2)}`, + // Without this the cut snaps to the nearest preceding keyframe, which can be + // seconds early — fine for scrubbing, not fine when the clip IS the citation. + "--force-keyframes-at-cuts", + ...extra, + // Pin H.264/AAC in mp4. Left alone yt-dlp picks VP9+Opus at these heights, + // and since --force-keyframes-at-cuts re-encodes, that means libvpx-vp9 — + // 27s to cut a 5s clip. It also writes .webm and appends that to -o. + "-f", fmt, + "--merge-output-format", "mp4", + "-o", dest, + "--", meta.webpageUrl, + ]; + + const attempt = async (extra) => { + try { + await execFileP(YTDLP, argsWith(extra), { maxBuffer: 1 << 26 }); + return null; + } catch (err) { + return ytdlpOk(err) ? null : err; + } + }; + + EMIT("fetch", { id: entry.id, video: entry.video, from, to, cached: false }); + let err = await attempt([]); + + // Rumble delivers HLS whose segments are named `.tar`, and ffmpeg 8's picky + // extension check rejects those outright — "URL ... is not in + // allowed_segment_extensions" — killing the fetch with exit 183. Rumble ships + // no progressive format to fall back to, so without this every Rumble-sourced + // clip is unbuildable. + // + // It has to be a RETRY, not a default: -extension_picky lives on the HLS + // demuxer, so passing it against a progressive URL (YouTube's googlevideo mp4) + // makes ffmpeg abort with "Option extension_picky not found" — i.e. adding it + // unconditionally trades a Rumble failure for a YouTube one. + if (err && /allowed_segment_extensions|allowed_extensions/.test(String(err.stderr ?? err.message ?? ""))) { + EMIT("note", { id: entry.id, message: ` ${entry.id}: HLS segment extension rejected, retrying with -extension_picky 0` }); + err = await attempt(["--downloader-args", "ffmpeg_i:-extension_picky 0"]); + } + if (err) { + throw new Error(`yt-dlp failed for ${entry.id} (${entry.video}): ${err.stderr ?? err.message}`); + } + if (!(await exists(dest))) { + // If a fallback format still forced another container, yt-dlp writes + // "<dest>.<realext>". Adopt it rather than failing the run. + const dir = path.dirname(dest); + const base = path.basename(dest); + const stray = (await readdir(dir)).find((f) => f.startsWith(base + ".")); + if (!stray) throw new Error(`yt-dlp reported success but produced no file for ${entry.id}`); + await rename(path.join(dir, stray), dest); + } + return { path: dest, fetchStart: from, cached: false }; +} + +// Parse ffmpeg's silencedetect output into [{s, e}] intervals, in seconds +// relative to the start of the given file. +async function detectSilence(file, render) { + const minDur = render.silenceMinDur ?? 0.09; + + // The threshold has to be RELATIVE to the clip, not absolute. These are game + // streams: the gaps between words are full of game audio and music, so they + // are quiet but nowhere near silent. A fixed -32 dB sits below the noise floor + // of a typical clip here and finds literally zero silences (measured: mean + // volume -21 dB, 0 hits at -32 dB, 25 hits at -26 dB). Measure the clip first + // and cut a few dB under its own mean instead. + const { stderr: volLog } = await execFileP( + FFMPEG, + ["-nostdin", "-i", file, "-af", "volumedetect", "-f", "null", "-"], + { maxBuffer: 1 << 26 }, + ).catch((e) => ({ stderr: e.stderr ?? "" })); + const meanMatch = (volLog ?? "").match(/mean_volume:\s*(-?[\d.]+) dB/); + const mean = meanMatch ? Number(meanMatch[1]) : -24; + const noise = Math.max(-45, Math.min(-18, mean - (render.silenceRelDb ?? 6))); + + // ffmpeg exits 0 here, so stderr comes back on the resolved result. + const { stderr } = await execFileP( + FFMPEG, + ["-nostdin", "-i", file, "-af", `silencedetect=noise=${noise.toFixed(1)}dB:d=${minDur}`, "-f", "null", "-"], + { maxBuffer: 1 << 26 }, + ).catch((e) => ({ stderr: e.stderr ?? "" })); + const log = stderr ?? ""; + + const out = []; + let open = null; + for (const line of log.split("\n")) { + const s = line.match(/silence_start:\s*(-?[\d.]+)/); + if (s) open = Number(s[1]); + const e = line.match(/silence_end:\s*(-?[\d.]+)/); + if (e && open !== null) { + out.push({ s: open, e: Number(e[1]) }); + open = null; + } + } + return out; +} + +// Snap a desired cut to the nearest silence, so the clip begins and ends between +// words instead of through one. Returns the desired point unchanged when no +// silence is close enough — better a tight cut than a cut in the wrong place. +function snap(desired, intervals, kind, window) { + let best = null; + for (const iv of intervals) { + // Starting: we want to resume just before speech does -> the silence's END. + // Ending: we want to stop just after speech does -> the silence's START. + const point = kind === "start" ? iv.e : iv.s; + const d = Math.abs(point - desired); + if (d > window) continue; + if (!best || d < best.d) best = { d, point }; + } + if (!best) return { at: desired, snapped: false }; + const lead = kind === "start" ? -0.10 : 0.18; + return { at: Math.max(0, best.point + lead), snapped: true }; +} + +// Encoder quality is manifest-driven so a cut can trade size for fidelity without +// editing this file. Defaults reproduce the original hardcoded settings exactly. +const encodeArgs = (render) => [ + "-c:v", "libx264", + "-preset", render.preset ?? "medium", + "-crf", String(render.crf ?? 20), + "-pix_fmt", "yuv420p", + "-r", String(render.fps), + "-c:a", "aac", + "-b:a", render.audioBitrate ?? "160k", + "-ar", String(render.audioRate), + "-ac", String(render.audioChannels), + "-movflags", "+faststart", +]; + +// The rail pass runs over an ALREADY ENCODED file, so its audio is already the +// finished AAC. Re-encoding it would cost a whole generation for nothing — and +// would make "the rail does not touch the audio" untrue. +const encodeArgsVideoOnly = (render) => [ + "-c:v", "libx264", + "-preset", render.preset ?? "medium", + "-crf", String(render.crf ?? 20), + "-pix_fmt", "yuv420p", + "-r", String(render.fps), + "-c:a", "copy", + "-movflags", "+faststart", +]; + +async function buildClipSegment(entry, meta, render, dirs, opts, chrome, nodes, provenance) { + const outDir = dirs.dir; + const { path: raw, fetchStart } = await fetchClip(entry, meta, render, dirs.rawDir, opts); + const seg = path.join(outDir, "segments", `${entry.id}.mp4`); + const pal = render.palette; + const { width, height } = render; + + // Desired cut points, expressed relative to the over-fetched file. + const wantA = entry.start - fetchStart; + const wantB = entry.end - fetchStart; + const win = render.snapWindow ?? 1.6; + + const sil = await detectSilence(raw, render); + const a = snap(wantA, sil, "start", win); + const b = snap(wantB, sil, "end", win); + // Never let snapping invert or collapse the window. + const cutA = Math.min(a.at, wantB - 1); + const cutB = Math.max(b.at, cutA + 1); + EMIT("snap", { id: entry.id, start: a.snapped, end: b.snapped, seconds: cutB - cutA }); + + const quotePath = path.join(outDir, "segments", `${entry.id}.quote.txt`); + const attribPath = path.join(outDir, "segments", `${entry.id}.attrib.txt`); + // Written for reference/diffing only — the quote is no longer drawn on screen. + await writeFile(quotePath, wrap(`“${entry.quote}”`, 92), "utf8"); + + const d = meta.uploadDate; + const date = `${d.slice(0, 4)}-${d.slice(4, 6)}-${d.slice(6, 8)}`; + await writeFile( + attribPath, + `${cleanTitle(meta.title)} · ${date} @ ${hms(entry.cite ?? entry.start)}`, + "utf8", + ); + + // The picture is the point. Nothing is drawn over it: the video is letterboxed + // between a thin citation header and a thin timeline footer, so the source + // material plays unobstructed and the additions stay subtle. + const HH = render.headerHeight ?? 56; + // The picture lives left of the rail column; the rail's own pixels are painted + // by the rail chain at concat time, over ground this pad leaves for it. + const VW = contentWidth(render); + const FH = chrome.footerHeight; + const hasFooter = FH > 0 && chrome.footer; + // headerHeight:0 drops the citation line too, leaving the clips alone on screen. + // Worth having: a cut whose sources are listed elsewhere does not need to carry + // its own attribution burnt into every frame. + const hasHeader = HH > 0; + const VH = height - HH - FH; + const trackAbsY = height - FH + chrome.trackY; + + // Where the progress marker travels this clip. Only the first clip of a + // section moves it; the rest hold it in place. + const T = render.slideSeconds ?? 0.9; + const xTo = chrome.xs[entry.section]; + const xFrom = entry.sectionEnter ? chrome.xs[Math.max(0, entry.section - 1)] : xTo; + // Commas inside a filter option have to survive filtergraph parsing; single + // quotes around the expression is what protects them. + const ramp = (a, b) => + a === b ? String(b) : `'if(lt(t,${T}),${a}+(${b}-${a})*t/${T},${b})'`; + const markX = ramp(xFrom - chrome.markerRadius, xTo - chrome.markerRadius); + + // The fill bar CANNOT be a drawbox with a `t`-dependent width. drawbox has no + // time variable at all: its `t` is the box THICKNESS, and with `t=fill` that + // is effectively INT_MAX, so the old `if(lt(t,0.9),…)` was always false and + // the bar was always drawn at its final width. (Proof: `drawbox=w='t*10'` and + // `drawbox=w=20` produce an identical YAVG.) Only the amber marker ever moved. + // + // So do it the way the rail does: a 2*LEN-wide strip, accent on the left half + // and transparent on the right, translated under a fixed-width crop. crop's + // x IS per-frame in `t`, and it clamps, so the ends are self-parking. + const fillA = xFrom - chrome.x0; + const fillB = xTo - chrome.x0; + const fillExpr = fillA === fillB + ? String(fillB) + : `${fillA}+(${fillB - fillA})*clip(t/${T},0,1)`; + + const base = [ + `scale=${VW}:${VH}:force_original_aspect_ratio=decrease`, + `pad=${VW}:${VH}:(ow-iw)/2:(oh-ih)/2:color=${pal.bg}`, + // Widen back to the full frame, leaving the rail column (if any) as ground. + `pad=${width}:${VH}:0:0:color=${pal.bg}`, + `pad=${width}:${height}:0:${HH}:color=${pal.bg}`, + "setsar=1", + `fps=${render.fps}`, + ...(hasHeader + ? [ + `drawbox=x=90:y=${Math.round((HH - 24) / 2)}:w=4:h=24:color=${pal.accent}:t=fill`, + [ + `drawtext=textfile='${attribPath}'`, + `fontfile='${render.fontRegular}'`, + "fontsize=22", + `fontcolor=${pal.muted}`, + "x=118", + `y=${Math.round((HH - 26) / 2)}`, + ].join(":"), + ] + : []), + ].join(","); + + // A manifest with a RAIL draws the code in the rail's foot instead, as one + // more strip: bottom-right of the frame, bordered, one per clip. It used to + // float over the bottom-right of the picture — the one part of the frame this + // cut promises never to draw on. Manifests with no rail keep the old overlay, + // byte for byte. + const qr = + render.qr === false || render.rail ? null : await qrForEntry(entry, provenance, render, outDir); + const qrM = render.qr?.margin ?? 28; + + // Bound the bar strip SHORTER than the clip. An overlay secondary that outruns + // the main extends the output, and the fix for that (shortest=1) would instead + // truncate the clip to the strip. Ending early is free: overlay's default + // eof_action=repeat holds the strip's last frame, which is the parked bar. + const barT = Math.max(0.2, cutB - cutA - 0.25); + + const inputs = ["-ss", cutA.toFixed(3), "-to", cutB.toFixed(3), "-i", raw]; + let nextIdx = 1; + let footerIdx, markerIdx, barIdx, qrIdx; + if (hasFooter) { + footerIdx = nextIdx++; inputs.push("-i", chrome.footer); + markerIdx = nextIdx++; inputs.push("-i", chrome.marker); + // A PNG fed with a plain -i through an ANIMATED crop is frozen: the crop + // sees one frame at t=0 and repeatlast repeats the already-cropped result. + // -loop 1 -framerate is what makes the strip a video the crop can walk. + barIdx = nextIdx++; + inputs.push( + "-loop", "1", "-framerate", String(render.fps), "-t", barT.toFixed(3), + "-i", chrome.bar, + ); + } + if (qr) { qrIdx = nextIdx++; inputs.push("-i", qr.png); } + + const parts = hasFooter + ? [ + `[0:v]${base}[b]`, + `[b][${footerIdx}:v]overlay=0:${height - FH}[f]`, + `[${barIdx}:v]crop=w=${chrome.trackLen}:h=3:x='${chrome.trackLen}-(${fillExpr})':y=0[bar]`, + `[f][bar]overlay=x=${chrome.x0}:y=${trackAbsY - 1}[g]`, + `[g][${markerIdx}:v]overlay=x=${markX}:y=${trackAbsY - chrome.markerRadius}[q]`, + ] + : [`[0:v]${base}[q]`]; + + // Sit above the footer when there is one, so the code never straddles the chrome. + parts.push( + qr + ? `[q][${qrIdx}:v]overlay=x=${VW}-w-${qrM}:y=H-h-${FH + qrM}[v]` + : `[q]null[v]`, + ); + + await execFileP( + FFMPEG, + [ + "-nostdin", "-v", "error", "-y", + ...inputs, + "-filter_complex", parts.join(";"), + "-map", "[v]", "-map", "0:a", + ...encodeArgs(render), + seg, + ], + { maxBuffer: 1 << 24 }, + ); + return seg; +} + +// ---- QR provenance code -------------------------------------------------- +// A compilation asks the viewer to take the edit on trust. The QR is the antidote: +// it resolves to this clip's exact START in the archive's own viewer, so anyone can +// pull up the surrounding hour and check that the cut is fair. Per clip, because a +// single code for the whole video would send everyone to the first citation. +// +// Two rules learned the hard way: it must be FULLY OPAQUE (a translucent QR will +// not scan) and it must keep its quiet zone (the white border is part of the +// symbol, not decoration). +async function qrForEntry(entry, provenance, render, outDir) { + const q = render.qr ?? {}; + // A mirror's LOCAL slug is not the id the site serves, and a clip taken from a + // copy whose archived transcript is broken should point at the copy that reads — + // so an explicit per-clip citeUrl always wins over the derived one. + const url = + entry.citeUrl ?? + `${provenance.siteOrigin}/?v=${encodeURIComponent( + `${entry.channel ?? provenance.channelSlug}/${entry.video}`, + )}&t=${Math.floor(entry.start)}`; + const png = path.join(outDir, "qr", `${entry.id}.png`); + await execFileP(QRENCODE, [ + "-o", png, + "-s", String(q.scale ?? 4), + "-m", String(q.quiet ?? 3), + "-l", q.ecc ?? "M", + url, + ]); + return { png, url }; +} + +async function buildCardSegment(card, render, outDir, nodes) { + const png = await renderCard(card, render, outDir, nodes); + const seg = path.join(outDir, "segments", `${card.id}.mp4`); + const dur = String(card.seconds); + + await execFileP( + FFMPEG, + [ + "-nostdin", "-v", "error", "-y", + // Without -framerate the image demuxer runs at its 25 fps default and the + // `-vf fps=30` below DUPLICATES a frame — at the segment's first frame, + // which is exactly where the next xfade seam lands. + "-loop", "1", "-framerate", String(render.fps), "-t", dur, "-i", png, + "-f", "lavfi", "-t", dur, + "-i", `anullsrc=channel_layout=stereo:sample_rate=${render.audioRate}`, + "-vf", `fps=${render.fps},setsar=1`, + ...encodeArgs(render), + "-shortest", + seg, + ], + { maxBuffer: 1 << 24 }, + ); + return seg; +} + +// =========================================================================== +// The claim rail +// =========================================================================== +// A persistent vertical ledger down the right edge, appending one row per claim +// as the video runs. It is folded into the concat pass rather than added as a +// second encode: the chain attaches AFTER the final xfade node, which already +// has post-pass semantics (nothing downstream of the last xfade is dissolved, +// and `t` there is absolute and continuous from 0). A separate pass would +// re-quantize crf-20 output, and antialiased text on flat colour is exactly the +// content that costs most. +// +// Five hard-won rules are load-bearing here; breaking any one produces a hang, +// a silently wrong-length file or a frozen overlay: +// +// 1. crop's w/h are CONFIG-TIME (`t` is undefined there) but x/y are +// per-frame. So every moving part is a fixed-size window walking a strip. +// 2. crop clamps x/y into range, so over-scroll is safe and self-parking. +// 3. A PNG on a plain -i through an animated crop is FROZEN. Every strip +// needs `-loop 1 -framerate <fps> -t <bound>`. +// 4. An UNBOUNDED `-loop 1` input deadlocks ffmpeg once several are chained. +// Hence `-t` on all five. +// 5. An overlay secondary longer than the main EXTENDS the output. The strips +// are deliberately longer (bound = total + 2), so `shortest=1` is required +// on EVERY rail overlay, not just the first. +// +// And two rendering ones: overlay's default `format=yuv420` subsamples alpha as +// well as chroma, which fringes 14–23 px rail text — so every rail overlay is +// `format=yuv444`, with a single `format=yuv420p` before the encoder. + +/** + * Fill an evenly-spaced schedule between known anchors. + * + * `known` holds the entries that are pinned to a segment; everything else is + * distributed linearly between its neighbouring pins, with `lo`/`hi` acting as + * virtual anchors just outside the run. + */ +function distribute(known, n, lo, hi) { + const at = new Array(n).fill(null); + for (const [i, t] of known) at[i] = t; + const pts = [[-1, lo], ...[...known].sort((a, b) => a[0] - b[0]), [n, hi]]; + for (let k = 0; k < pts.length - 1; k += 1) { + const [a, ta] = pts[k]; + const [b, tb] = pts[k + 1]; + for (let j = a + 1; j < b; j += 1) at[j] = ta + ((tb - ta) * (j - a)) / (b - a); + } + return at; +} + +/** + * When each ledger row appears, in finished-timeline seconds. + * + * ONE CHRONOLOGY. The cut plays in date order across every company, and the + * ledger is sorted the same way, so a claim's position in the rail IS its + * position in time. That collapses what used to live here: there is no longer a + * per-company span to bound a claim to, no contiguity rule to enforce, and no + * risk of scheduling a media claim over a coffee clip -- because "over a coffee + * clip" now means "later in the same chronology", which is exactly right. + * + * What remains is the part that was always doing the work: a claim WITH a clip + * behind it is pinned to that clip's segment, and the rest are spread evenly + * between their neighbouring pins. Monotonicity holds by construction, since + * both the pins and the rows are in date order. + * + * Every state change lands at `starts[i] + D/2` -- MID-DISSOLVE -- where the + * picture is already crossfading and a ±3-frame error is invisible. + */ +export function scheduleClaims(ledger, entries, starts, D, endBound) { + const segOf = new Map(entries.map((e, i) => [e.id, i])); + const mid = (seg) => starts[seg] + D / 2; + + // A stacked ledger card carries several claims, and each one has a MOMENT + // inside that card: the reveal of its own row. Pinning all of them to the + // segment's mid-dissolve would land four rail rows on one frame and, worse, + // break the pin-order guard's strict monotonicity for no reason. So a claim + // on such a card is pinned to its own row's reveal. + const within = new Map(); + for (const e of entries) { + if (e.type !== "ledger") continue; + (e.claims ?? []).forEach((cid, r) => within.set(`${e.id}|${cid}`, ledgerRevealAt(r))); + } + + const known = new Map(); + ledger.forEach((c, i) => { + const seg = c.entryId ? segOf.get(c.entryId) : undefined; + if (seg === undefined) return; + const off = within.get(`${c.entryId}|${c.id}`); + known.set(i, off === undefined ? mid(seg) : starts[seg] + off); + }); + if (!known.size) throw new Error("ledger: no claim is pinned to a clip, so nothing anchors the rail"); + + // A pin that runs backwards means the ledger and the timeline disagree about + // the order of events, which is a manifest bug rather than something to + // silently smooth over -- the whole cut rests on the two agreeing. + const pins = [...known].sort((a, b) => a[0] - b[0]); + for (let i = 1; i < pins.length; i += 1) { + if (pins[i][1] <= pins[i - 1][1]) { + throw new Error( + `ledger: ${ledger[pins[i][0]].id} is pinned to ${ledger[pins[i][0]].entryId}, which plays ` + + `before ${ledger[pins[i - 1][0]].id}'s clip — the ledger is not in the cut's order`, + ); + } + } + + // The first card is the title; the rail's own run opens just after it. + const times = distribute(known, ledger.length, mid(0), endBound); + + for (let i = 1; i < times.length; i += 1) { + if (times[i] <= times[i - 1]) times[i] = times[i - 1] + 1 / 30; + } + return times; +} + +/** + * The rail's filtergraph, as one builder with two call sites — the concat pass + * and `--rail-only` — so the two paths cannot drift. + * + * Ramps are CUMULATIVE AND SATURATING, never gated. A piecewise sum of + * `gte(t,s)*lt(t,s')*…` terms flashes to y=0 for one frame at any boundary gap, + * because every gate evaluates false at once and the sum collapses. Terms that + * rise to their delta and stay there cannot do that. + */ +export function railFilterChain(rail, assets, times, render, inLabel, firstInputIdx, bound, opts = {}) { + const g = assets.geom; + const SLIDE = rail.slide ?? 0.55; + const fps = render.fps; + + const P = (s) => `clip((t-${s.toFixed(3)})/${SLIDE},0,1)`; + // smoothstep() does not exist in ffmpeg's expression language. This is it. + const ease = (s) => { const p = P(s); return `${p}*${p}*(3-2*${p})`; }; + // Signed, and explicitly so. Joining terms with "+" was fine while every + // delta was a positive row height; a rolling cell FALLS as often as it rises, + // and `…+-40*x` is at best relying on ffmpeg's unary minus. + const sum = (y0, terms) => + terms.reduce( + (acc, t) => `${acc}${t.d < 0 ? "-" : "+"}${Math.abs(t.d)}*${t.f}`, + String(y0), + ); + const ramp = (y0, steps) => + sum(y0, steps.filter((st) => st.delta !== 0).map((st) => ({ d: st.delta, f: ease(st.at) }))); + + const { K, ROWH, RW, RX, RTOP, LOGH, LOGTOP, TALLYTOP } = g; + + const logY = ramp(0, times.map((t, i) => ({ at: t, delta: i + 1 > K ? ROWH : 0 }))); + // The curtain and the log MUST share the same eased P, or the curtain visibly + // lags the rows mid-slide and unrevealed claims flash into view. + const curtainY = ramp(LOGTOP, times.map((t, i) => ({ at: t, delta: i + 1 <= K ? ROWH : 0 }))); + const hlY = ramp(LOGTOP, times.map((t, i) => ({ at: t, delta: i > 0 && i < K ? ROWH : 0 }))); + + /** + * One lane's y, in the strip's own pixels. + * + * y(t) = r0 + Σ_k [ (a_k − b_{k−1})·gte(t,t_k) + (b_k − a_k)·ease(t_k) ] + * + * The first term is the instantaneous reposition to the next pair's starting + * row; the second is the roll itself. Both are CUMULATIVE AND SATURATING, + * which is the rail's hard rule: a gated piecewise sum flashes to y=0 for one + * frame at any boundary gap, because every gate goes false at once. + */ + const laneY = (lane) => { + const terms = []; + let prevB = 0; + lane.steps.forEach((st, i) => { + if (!st) return; + const at = times[i]; + const jump = (st.a - prevB) * lane.cellH; + const roll = (st.b - st.a) * lane.cellH; + if (jump !== 0) terms.push({ d: jump, f: `gte(t,${at.toFixed(3)})` }); + if (roll !== 0) terms.push({ d: roll, f: ease(at) }); + prevB = st.b; + }); + return sum(0, terms); + }; + + // The rail leaves by SLIDING OFF to the right, not by an enable= pop. One + // offset expression shared by every overlay, so the column moves as one + // object; `overlay`'s x is per-frame in `t`, which is what makes that + // possible at all. Cumulative and saturating, like everything else here. + const hideAt = opts.hideAt ?? null; + const OFF = hideAt == null ? "" : `+${RW + 8}*${ease(hideAt)}`; + const X = (x) => (OFF ? `'${x}${OFF}'` : String(x)); + + const i0 = firstInputIdx; + const files = [assets.chrome, assets.log, assets.curtain, assets.hl, assets.tally]; + if (assets.qr) files.push(assets.qr.path); + const inputs = files.flatMap((f) => [ + "-loop", "1", "-framerate", String(fps), "-t", bound.toFixed(3), "-i", f, + ]); + + const chain = [ + `[${i0 + 1}:v]crop=w=${RW}:h=${LOGH}:x=0:y='${logY}'[rlog]`, + `${inLabel}[${i0}:v]overlay=x=${X(RX)}:y=${RTOP}:format=yuv444:shortest=1[rr0]`, + `[rr0][rlog]overlay=x=${X(RX)}:y=${LOGTOP}:format=yuv444:shortest=1[rr1]`, + // The highlight goes UNDER the curtain: while the list is still filling, the + // row it marks has not been revealed yet, and the curtain is what hides it. + `[rr1][${i0 + 3}:v]overlay=x=${X(RX)}:y='${hlY}':format=yuv444:shortest=1[rr2]`, + `[rr2][${i0 + 2}:v]overlay=x=${X(RX)}:y='${curtainY}':format=yuv444:shortest=1[rr3]`, + ]; + + // One crop per lane out of the SINGLE tally strip. Four numbers that roll + // independently and a roster line that mostly does not, for one more input + // than the slab cost. + // + // `split` first, and it is NOT optional: a filtergraph link may be consumed + // exactly once, so five crops reading `[N:v]` is a parse error, not a + // shortcut. This is the whole reason the lanes share one PNG and still cost + // one input. + chain.push( + `[${i0 + 4}:v]split=${assets.lanes.length}${assets.lanes.map((_, j) => `[ts${j}]`).join("")}`, + ); + let lab = "[rr3]"; + assets.lanes.forEach((lane, j) => { + const isRoster = lane.kind === "roster"; + const h = isRoster ? g.ROSTERH : g.TALLYROWH; + const y = isRoster ? g.ROSTERTOP : TALLYTOP + j * g.TALLYROWH; + const x = isRoster ? RX + g.ROSTERX : RX + g.CELLX; + chain.push( + `[ts${j}]crop=w=${lane.w}:h=${h}:x=${lane.x}:y='${laneY(lane)}'[rc${j}]`, + `${lab}[rc${j}]overlay=x=${X(x)}:y=${y}:format=yuv444:shortest=1[rt${j}]`, + ); + lab = `[rt${j}]`; + }); + + // The provenance tile LAST, so the parked curtain cannot paint over it. + if (assets.qr) { + const qrY = sum(0, assets.qr.steps.map((st) => ({ d: st.delta, f: `gte(t,${st.at.toFixed(3)})` }))); + chain.push( + `[${i0 + 5}:v]crop=w=${g.TILEW}:h=${g.TILEH}:x=0:y='${qrY}'[rqr]`, + `${lab}[rqr]overlay=x=${X(RX + g.PAD)}:y=${g.TILETOP}:format=yuv444:shortest=1[rq]`, + ); + lab = "[rq]"; + } + + chain.push(`${lab}format=yuv420p[vout]`); + + return { inputs, chain: chain.join(";"), outLabel: "[vout]" }; +} + +/** + * The chrome as PNG-sequence overlays, for `render.chromeEngine: "hyperframes"`. + * + * OPT-IN, and absent it nothing below runs -- the ffmpeg chrome path is left + * byte-for-byte alone, which is the same bargain the rail was added under. + * + * The five ffmpeg traps the rail documents apply here unchanged, and two of them + * bite harder with an image sequence: + * + * * `format=yuv444` on EVERY overlay. overlay's default yuv420 subsamples + * ALPHA as well as chroma, which fringes small text -- and the band is + * nothing but small text. + * * `shortest=1` on EVERY overlay. A secondary longer than the main EXTENDS + * the output; the sequence is rendered to the same length as the concat, but + * a one-frame rounding difference either way must not change the duration. + * * One `format=yuv420p` before the encoder, once, at the end. + * + * A finite image sequence needs no `-t`: unlike `-loop 1` it ends by itself, so + * the deadlock the rail's five chained loops hit cannot happen here. + */ +export function chromeOverlayChain(render, regions, inLabel, firstInputIdx, opts = {}) { + const { outLabel = "[hfout]", final = true } = opts; + const inputs = []; + const parts = []; + let lab = inLabel; + regions.forEach((r, i) => { + inputs.push( + "-framerate", String(render.fps), + "-start_number", "1", + "-i", path.join(r.frames, "frame_%06d.png"), + ); + const idx = firstInputIdx + i; + const last = i === regions.length - 1; + const out = last && !final ? outLabel : `[hf${i}]`; + parts.push(`${lab}[${idx}:v]overlay=x=${r.x}:y=${r.y}:format=yuv444:shortest=1${out}`); + lab = out; + }); + if (final) parts.push(`${lab}format=yuv420p[vout]`); + return { + inputs, + chain: parts.join(";"), + outLabel: final ? "[vout]" : outLabel, + count: regions.length, + }; +} + +/** + * Where each rendered chrome region sits in the frame. + * + * The chart band REPLACES the footer node track rather than joining it: the + * track only moved at section handovers, which is precisely the fault the band + * exists to fix. So it takes the footer's ground and 100px more of it, and the + * picture loses that height. + */ +export function chromeRegions(render, outDir) { + const H = render.chart?.height ?? 200; + return [ + { + name: "chart", + frames: path.join(outDir, "chrome", "chart-frames"), + x: 0, + y: render.height - H, + width: contentWidth(render), + height: H, + }, + ]; +} + +/** + * The footer's stand-in when the chrome is drawn in a browser. + * + * It reserves the band's HEIGHT and draws nothing, so every segment letterboxes + * to the same picture box the overlay expects and the ground under the band is + * the palette background. `footer: null` is what switches the whole ffmpeg + * footer -- image, marker and fill bar -- off; the degenerate shape is the one + * renderFooterAssets already returns for a manifest with no nodes, so this path + * is not new. + */ +function reservedFooter(render) { + return { + footer: null, marker: null, bar: null, trackLen: 0, + footerHeight: render.chart?.height ?? 200, + trackY: 0, xs: [], x0: 0, markerRadius: 0, + }; +} + +/** + * Everything the rail chain needs that depends on the built segments. Returns + * null when the manifest does not ask for a rail — which is what keeps this + * whole feature opt-in and every existing report byte-for-byte unchanged. + */ +async function buildRailPlan(manifest, render, entries, segments, D, outDir) { + const rail = render.rail; + if (!rail || !manifest.ledger?.length) return null; + const { starts, total } = await segmentOffsets(segments, D, render.fps); + const assets = await renderRailAssets( + render, manifest.ledger, outDir, entries, manifest.provenance, + ); + // Every claim must be on the board before the closing ledger scroll reads it + // back, so the last section's spare rows are spread up to that segment. + const endIdx = entries.findIndex((e) => e.type === "scroll" || e.type === "chart"); + const endBound = endIdx > 0 ? starts[endIdx] : total; + const times = scheduleClaims(manifest.ledger, entries, starts, D, endBound); + + // The QR tile changes at the MID-DISSOLVE of every segment, instantaneously + // — a code that eased into place would spend the ease unscannable, and the + // picture is already crossfading there. + if (assets.qr) { + assets.qr.steps = entries.slice(1).map((_, i) => ({ + at: starts[i + 1] + D / 2, + delta: assets.geom.TILEH, + })); + } + + // Where the rail leaves. The closing ledger is a full-width card and the rail + // is the one thing on screen it would have to be read around, so the column + // slides off over that card's dissolve and does not come back. + const hideIdx = entries.findIndex((e) => e.hideRail); + const hideAt = hideIdx > 0 ? starts[hideIdx] : null; + + // The schedule, written down. + // + // The chart band has to sweep in step with the rail -- a playhead that tracks + // the current moment is the whole point of it -- and the only way it and the + // rail can be guaranteed to agree is for one of them to compute the schedule + // and the other to READ it. Recomputing from segment durations would be a + // second implementation of segmentOffsets(), and it would drift the first time + // the crossfade changed. Same rule as widen(): imported, never reimplemented. + await writeFile( + path.join(outDir, "schedule.json"), + JSON.stringify( + { + fps: render.fps, + transition: D, + total, + endBound, + segments: entries.map((e, i) => ({ id: e.id, type: e.type, start: starts[i] })), + claims: manifest.ledger.map((c, i) => ({ id: c.id, at: times[i], entryId: c.entryId ?? null })), + }, + null, + 2, + ) + "\n", + ); + + EMIT("note", { + message: `rail: ${manifest.ledger.length} claims, ${times.filter((_, i) => manifest.ledger[i].entryId).length} pinned, ` + + `window ${assets.geom.K} rows` + (hideAt == null ? "" : `, hides at ${hideAt.toFixed(1)}s`), + }); + return { assets, times, total, hideAt }; +} + +// ---- the end sequence ---------------------------------------------------- +// Two segment kinds that exist only to close the cut: the whole ledger read +// back in one scroll, then the same claims plotted. Both go through the SAME +// `fps=,setsar=1` and the SAME encodeArgs as every other segment — xfade +// rejects a mismatched link with "First input link parameters do not match", +// which would surface only at concat time, after every fetch has been paid for. + +async function buildScrollSegment(card, render, outDir, ledger) { + const { path: png, contentHeight, width: VW } = await renderScrollCard(card, render, ledger, outDir); + const seg = path.join(outDir, "segments", `${card.id}.mp4`); + const pal = render.palette; + const { width, height } = render; + const HH = render.headerHeight ?? 56; + // The band's height, not the manifest's footerHeight — otherwise the last + // 100px of the scroll play underneath the chart. + const FH = reservedFooterHeight(render); + const winH = height - HH - FH; + const dur = String(card.seconds); + + // clip() buys a free hold at BOTH ends, and crop's own clamping degrades an + // off-by-a-few contentHeight into a static last frame rather than an error. + const hold = card.hold ?? 2.0; + const travel = Math.max(0, contentHeight - winH); + const denom = Math.max(0.1, card.seconds - 2 * hold); + + await execFileP( + FFMPEG, + [ + "-nostdin", "-v", "error", "-y", + "-loop", "1", "-framerate", String(render.fps), "-t", dur, "-i", png, + "-f", "lavfi", "-t", dur, "-i", `color=c=${pal.bg}:s=${width}x${height}:r=${render.fps}`, + "-f", "lavfi", "-t", dur, + "-i", `anullsrc=channel_layout=stereo:sample_rate=${render.audioRate}`, + "-filter_complex", [ + `[0:v]crop=w=${VW}:h=${winH}:x=0:y='${travel}*clip((t-${hold})/${denom.toFixed(3)},0,1)'[win]`, + `[1:v][win]overlay=x=0:y=${HH}:shortest=1,fps=${render.fps},setsar=1[v]`, + ].join(";"), + "-map", "[v]", "-map", "2:a", + ...encodeArgs(render), + "-shortest", + seg, + ], + { maxBuffer: 1 << 24 }, + ); + return seg; +} + +/** + * A stacked ledger card: rows revealed in sequence by a walking curtain. + * + * The curtain is an opaque `pal.bg` rectangle that starts covering every row + * and steps down one row-height per reveal. Same device as the rail's, and for + * the same reason: the card ground is flat, so an opaque rectangle over it is + * an exact in-place wipe with no per-pixel filter. + * + * The ramp is CUMULATIVE AND SATURATING, like every other ramp here. + */ +async function buildLedgerSegment(card, render, outDir, ledger, avail) { + const geo = await renderLedgerCard(card, render, ledger, outDir, avail); + const seg = path.join(outDir, "segments", `${card.id}.mp4`); + const pal = render.palette; + const { width, height } = render; + // `seconds` is DERIVED, not authored: the pins that land claims on their own + // rows read the same clock, so a hand-set duration would silently move them. + const dur = String(card.seconds ?? ledgerSeconds(geo.rows)); + + const curtainH = height; + const curtain = path.join(outDir, "cards", `${card.id}.curtain.png`); + await execFileP("magick", [ + "-size", `${geo.width}x${curtainH}`, `xc:${pal.bg}`, curtain, + ]); + + const SLIDE = render.rail?.slide ?? 0.55; + const ease = (at) => { + const p = `clip((t-${at.toFixed(3)})/${SLIDE},0,1)`; + return `${p}*${p}*(3-2*${p})`; + }; + const y = [ + String(geo.rowsTop), + ...Array.from({ length: geo.rows }, (_, r) => `${geo.rowHeight}*${ease(ledgerRevealAt(r))}`), + ].join("+"); + + await execFileP( + FFMPEG, + [ + "-nostdin", "-v", "error", "-y", + "-f", "lavfi", "-t", dur, "-i", `color=c=${pal.bg}:s=${width}x${height}:r=${render.fps}`, + "-loop", "1", "-framerate", String(render.fps), "-t", dur, "-i", geo.path, + "-loop", "1", "-framerate", String(render.fps), "-t", dur, "-i", curtain, + "-f", "lavfi", "-t", dur, + "-i", `anullsrc=channel_layout=stereo:sample_rate=${render.audioRate}`, + "-filter_complex", [ + `[0:v][1:v]overlay=x=0:y=0:shortest=1[a]`, + `[a][2:v]overlay=x=0:y='${y}':shortest=1,fps=${render.fps},setsar=1[v]`, + ].join(";"), + "-map", "[v]", "-map", "3:a", + ...encodeArgs(render), + "-shortest", + seg, + ], + { maxBuffer: 1 << 24 }, + ); + return seg; +} + +async function buildChartSegment(card, render, outDir, ledger) { + const chart = await renderChartCard(card, render, ledger, outDir); + const seg = path.join(outDir, "segments", `${card.id}.mp4`); + const pal = render.palette; + const { width, height } = render; + const VW = cardWidth(card, render); + const dur = String(card.seconds); + + // The wipe CANNOT be `crop=w='<ramp>'` — crop's w is config-time and `t` is + // undefined there ("Error when evaluating the expression"). So: overlay the + // finished chart, then slide an opaque pal.bg rectangle rightwards off it. + // The card ground is flat pal.bg, so this is an exact in-place wipe with no + // per-pixel filter, and it draws the plot in like a plotter. + // `hold: true` -- the closing chart is a HOLD, not a reveal. + // + // The wipe existed because this card was the first and only time the viewer + // saw the numbers plotted. With the chart band drawing live under the whole + // cut, wiping it in again would re-tell a story the viewer has just watched + // happen. So the card opens on the finished plot and the seconds go to + // reading the final gap and its flags instead. + if (card.hold) { + await execFileP( + FFMPEG, + [ + "-nostdin", "-v", "error", "-y", + "-f", "lavfi", "-t", dur, "-i", `color=c=${pal.bg}:s=${width}x${height}:r=${render.fps}`, + "-loop", "1", "-framerate", String(render.fps), "-t", dur, "-i", chart.path, + "-f", "lavfi", "-t", dur, + "-i", `anullsrc=channel_layout=stereo:sample_rate=${render.audioRate}`, + "-filter_complex", + `[0:v][1:v]overlay=x=0:y=0:shortest=1,fps=${render.fps},setsar=1[v]`, + "-map", "[v]", "-map", "2:a", + ...encodeArgs(render), + "-shortest", + seg, + ], + { maxBuffer: 1 << 24 }, + ); + return seg; + } + + const wipeW = VW - chart.plotX; + const wipe = path.join(outDir, "cards", `${card.id}.wipe.png`); + await execFileP("magick", [ + "-size", `${wipeW}x${Math.round(chart.plotH)}`, `xc:${pal.bg}`, wipe, + ]); + + const wipeStart = card.wipeStart ?? 0.8; + const wipeDur = card.wipeSeconds ?? Math.max(1, card.seconds - wipeStart - 3.0); + + await execFileP( + FFMPEG, + [ + "-nostdin", "-v", "error", "-y", + "-f", "lavfi", "-t", dur, "-i", `color=c=${pal.bg}:s=${width}x${height}:r=${render.fps}`, + "-loop", "1", "-framerate", String(render.fps), "-t", dur, "-i", chart.path, + "-loop", "1", "-framerate", String(render.fps), "-t", dur, "-i", wipe, + "-f", "lavfi", "-t", dur, + "-i", `anullsrc=channel_layout=stereo:sample_rate=${render.audioRate}`, + "-filter_complex", [ + `[0:v][1:v]overlay=x=0:y=0:shortest=1[a]`, + `[a][2:v]overlay=x='${chart.plotX}+${wipeW}*clip((t-${wipeStart})/${wipeDur.toFixed(3)},0,1)'` + + `:y=${Math.round(chart.plotY)}:shortest=1,fps=${render.fps},setsar=1[v]`, + ].join(";"), + "-map", "[v]", "-map", "3:a", + ...encodeArgs(render), + "-shortest", + seg, + ], + { maxBuffer: 1 << 24 }, + ); + return seg; +} + +// Crossfade every segment into the next. This is a full re-encode of the +// timeline — the concat demuxer can only stream-copy hard cuts — so --no-xfade +// stays available for quick iteration. +async function concatWithXfade(segments, render, outPath, railPlan, chrome = null) { + const D = render.transition ?? 0.5; + const durs = []; + for (const s of segments) durs.push(await probeDuration(s, render.fps)); + + const inputs = segments.flatMap((s) => ["-i", s]); + const parts = []; + let vlab = "[0:v]"; + let alab = "[0:a]"; + let acc = durs[0]; + + for (let i = 1; i < segments.length; i += 1) { + const off = acc - D; + parts.push(`${vlab}[${i}:v]xfade=transition=fade:duration=${D}:offset=${off.toFixed(3)}[v${i}]`); + parts.push(`${alab}[${i}:a]acrossfade=d=${D}:c1=tri:c2=tri[a${i}]`); + vlab = `[v${i}]`; + alab = `[a${i}]`; + acc = acc + durs[i] - D; + } + + // The rail attaches to the LAST xfade node, so it runs after every dissolve + // and sees an absolute, continuous `t`. One encode, not two. + // + // When the chrome is rendered rather than drawn, the PNG regions go on FIRST + // and the rail chain reads their output. Not the other way round: the rail + // chain ends in `format=yuv420p`, and overlaying an alpha sequence onto + // yuv420p is the fringing trap the rail already documents, one layer later. + const chromeIn = chrome ?? null; + const railIn = chromeIn ? chromeIn.outLabel : vlab; + const rc = railPlan + ? railFilterChain( + render.rail, railPlan.assets, railPlan.times, render, + railIn, segments.length, railPlan.total + 2, + { hideAt: railPlan.hideAt }, + ) + : null; + const railInputs = rc ? rc.inputs.filter((a) => a === "-i").length : 0; + const hf = chromeIn + ? chromeOverlayChain(render, chromeIn.regions, vlab, segments.length + railInputs, { + outLabel: chromeIn.outLabel, + final: !rc, + }) + : null; + if (hf) parts.push(hf.chain); + if (rc) parts.push(rc.chain); + + const tail = rc ? rc.outLabel : hf ? hf.outLabel : vlab; + + await execFileP( + FFMPEG, + [ + "-nostdin", "-v", "error", "-y", + ...inputs, + ...(rc ? rc.inputs : []), + ...(hf ? hf.inputs : []), + "-filter_complex", parts.join(";"), + "-map", tail, "-map", alab, + ...encodeArgs(render), + outPath, + ], + { maxBuffer: 1 << 26 }, + ); +} + +/** + * Run the rail chain over an already-concatenated file. + * + * Two callers need this. `--rail-only` iterates on the rail in seconds instead + * of re-running the whole concat; and `--no-xfade` has no choice, because + * concatHardCut is `-c copy` and a stream-copy mux cannot host a filtergraph + * at all. + */ +async function applyRail(inPath, outPath, render, railPlan, preview) { + const rc = railFilterChain( + render.rail, railPlan.assets, railPlan.times, render, + preview ? "[base]" : "[0:v]", 1, railPlan.total + 2, + { hideAt: railPlan.hideAt }, + ); + const parts = []; + if (preview) { + // -ss restarts `t` near zero, which would put every absolute-time ramp in + // the wrong place — the rail would look broken while being correct. Shift + // the timestamps back to where the expressions think they are, then rebase + // them so the preview file still starts at 0. + parts.push(`[0:v]setpts=PTS+${preview.start.toFixed(3)}/TB[base]`); + } + parts.push(rc.chain); + const tail = preview ? "[vshift]" : rc.outLabel; + if (preview) parts.push(`${rc.outLabel}setpts=PTS-STARTPTS[vshift]`); + + await execFileP( + FFMPEG, + [ + "-nostdin", "-v", "error", "-y", + ...(preview ? ["-ss", String(preview.start), "-t", String(preview.dur)] : []), + "-i", inPath, + ...rc.inputs, + "-filter_complex", parts.join(";"), + "-map", tail, "-map", "0:a", + ...encodeArgsVideoOnly(render), + outPath, + ], + { maxBuffer: 1 << 26 }, + ); +} + +// ---- chapter markers ----------------------------------------------------- +// A compilation like this is a reference document as much as a video: the report +// cites moments, and a viewer wants to jump to them. Every clip therefore becomes +// a chapter. Offsets are derived exactly the way concatWithXfade derives its xfade +// offsets, so they stay correct for both crossfaded and hard-cut timelines. +// +// ffmetadata is a line-based format where =, ;, # and \ are structural, so a +// title carrying any of them has to be escaped or the file silently mis-parses. +const ffmetaEscape = (s) => String(s).replace(/([=;#\\])/g, "\\$1").replace(/\n/g, " "); + +export async function segmentOffsets(segments, D, fps) { + const durs = []; + for (const s of segments) durs.push(await probeDuration(s, fps)); + const starts = []; + let acc = 0; + for (let i = 0; i < durs.length; i += 1) { + starts.push(acc); + acc += durs[i] - (i < durs.length - 1 ? D : 0); + } + return { starts, total: acc }; +} + +async function chapterTitle(entry, index, provenance) { + if (entry.chapter) return entry.chapter; + if (entry.type !== "clip") return entry.title ?? entry.heading ?? `Card ${index + 1}`; + try { + const meta = await videoMeta(entry.video, entry.channel ?? provenance.channelSlug, { siteChannel: entry.siteChannel, siteVideo: entry.siteVideo }); + const d = String(meta.uploadDate ?? ""); + const date = /^\d{8}$/.test(d) ? `${d.slice(0, 4)}-${d.slice(4, 6)}-${d.slice(6, 8)}` : d; + const title = String(meta.title ?? entry.video); + return `${date} — ${title.length > 60 ? `${title.slice(0, 57)}…` : title}`.trim(); + } catch { + return `${index + 1}. ${entry.video}`; + } +} + +async function muxChapters(finalPath, entries, segments, D, outDir, provenance, fps) { + if (segments.length < 2) return; + const { starts, total } = await segmentOffsets(segments, D, fps); + const lines = [";FFMETADATA1", ""]; + for (let i = 0; i < entries.length; i += 1) { + // Land just PAST the crossfade, so the marker opens on the incoming clip + // rather than on the outgoing one mid-dissolve. + const start = i === 0 ? 0 : starts[i] + D; + const end = i === entries.length - 1 ? total : starts[i + 1] + D; + lines.push( + "[CHAPTER]", + "TIMEBASE=1/1000", + `START=${Math.round(start * 1000)}`, + `END=${Math.round(end * 1000)}`, + `title=${ffmetaEscape(await chapterTitle(entries[i], i, provenance))}`, + "", + ); + } + const metaPath = path.join(outDir, "chapters.ffmeta"); + await writeFile(metaPath, lines.join("\n"), "utf8"); + + // Stream copy — adding chapters must never re-encode the finished timeline. + const tmp = finalPath.replace(/\.mp4$/, ".chapters.mp4"); + await execFileP( + FFMPEG, + ["-nostdin", "-v", "error", "-y", "-i", finalPath, "-i", metaPath, + "-map", "0", "-map_metadata", "0", "-map_chapters", "1", "-c", "copy", tmp], + { maxBuffer: 1 << 24 }, + ); + await rename(tmp, finalPath); + EMIT("chapters", { n: entries.length, file: path.basename(metaPath) }); +} + +async function concatHardCut(segments, outDir, outPath) { + const listPath = path.join(outDir, "concat.txt"); + await writeFile(listPath, segments.map((s) => `file '${s}'`).join("\n") + "\n", "utf8"); + await execFileP( + FFMPEG, + ["-nostdin", "-v", "error", "-y", "-f", "concat", "-safe", "0", + // The concat demuxer stitches per-file timestamps; without generated PTS a + // stream copy can hand the next stage a discontinuous timeline, which the + // rail's absolute-time expressions would then read off by that much. + "-fflags", "+genpts", + "-i", listPath, "-c", "copy", outPath], + { maxBuffer: 1 << 24 }, + ); +} + +// A hard-cut concat and a crossfaded one are different lengths, so a cached +// prerail from one is a wrong base for the other. Keeping them in separate files +// means the mode can be switched without a stale-cache trap -- and without the +// length assertion below having to be the thing that explains it. +const prerailPath = (outDir, slug, D) => + path.join(outDir, `${slug}.prerail${D === 0 ? "-hardcut" : ""}.mp4`); + +// The finished timeline must be exactly as long as segmentOffsets says. Anything +// else means a filter changed the length behind our backs. +async function assertConcatLength(file, expected, fps, what) { + const got = await probeDuration(file, fps); + if (Math.abs(got - expected) > 1.5 / fps) { + throw new Error( + `${what}: duration ${got.toFixed(3)}s but the timeline is ${expected.toFixed(3)}s ` + + `(${((got - expected) * fps).toFixed(1)} frames out)` + + (/prerail/.test(what) ? " — delete it and let this rebuild it" : ""), + ); + } +} + +/** + * Build a manifest into a video. + * + * Exported so umtool's driver runs the SAME code the CLI does. It is still + * SPAWNED rather than imported by the app: a 40-minute chain of yt-dlp and + * ffmpeg inside a request handler has no cancellation story, and a runaway + * grandchild would outlive the request that started it. + */ +export async function buildVideo({ manifestPath, opts = {}, out, only, fetchOnly } = {}) { + const variant = opts.variant ?? "sourced"; + const whole = JSON.parse(await readFile(manifestPath, "utf8")); + const manifest = selectVariant(whole, variant); + const { render, provenance } = manifest; + + // The manifest already records which archive it was built against, so a clone + // with no corpus needs no extra configuration to read cue windows. + CUES = createCueSource({ + siteOrigin: opts.siteOrigin ?? process.env.SITE_ORIGIN ?? siteOriginFromManifest(whole), + resolveSiteIds: opts.resolveSiteIds === true, + prefer: opts.cueSource ?? "auto", + log: (m) => EMIT("log", { message: m }), + }); + const outRoot = out ?? path.join(path.dirname(path.resolve(manifestPath)), "out"); + const dirs = variantPaths(outRoot, manifest.slug, variant); + const outDir = dirs.dir; + + await mkdir(dirs.rawDir, { recursive: true }); + for (const d of ["cards", "segments", "qr"]) { + await mkdir(path.join(outDir, d), { recursive: true }); + } + + // Fetch one clip's window and stop. This is what the clip bench's "fetch 20s + // more" runs, so a bench fetch and a build fetch can never disagree about + // naming, format selection, the VP9 trap or the Rumble HLS retry. + // Reads the WHOLE manifest, not the variant's view of it: a clip bench fetch + // is about a moment in the corpus, and which cut happens to carry it is + // beside the point. + if (fetchOnly) { + let entry = whole.timeline.find((e) => e.id === fetchOnly); + // `!== "clip"`, not `=== "card"`. The timeline's vocabulary is OPEN -- one + // real manifest carries `scroll` and `chart` entries -- and the card-only + // check sent `undefined` into the fetcher for either of those. + if (entry && entry.type !== "clip") { + throw new Error(`${fetchOnly} is a ${entry.type ?? "non-clip"} entry, not a clip`); + } + if (!entry) { + // A LEDGER CLAIM. Adjudicating one means listening around the moment, and + // most of the ledger is cited by no clip at all -- so the claim page asks + // for a window the timeline has no entry for. It is fetched through this + // same path so the file lands in clips-raw under the build's own naming, + // inherits the format pin and the Rumble HLS retry, and is REUSED by a + // later build rather than fetched a second time. + const claim = (whole.ledger ?? []).find((e) => e.id === fetchOnly); + if (!claim) throw new Error(`no timeline entry or ledger claim with id ${fetchOnly}`); + if (!claim.video) throw new Error(`ledger claim ${fetchOnly} has no \`video\` to fetch`); + const at = Number(claim.cite); + if (!Number.isFinite(at)) throw new Error(`ledger claim ${fetchOnly} has no \`cite\` second`); + // A claim is a MOMENT, not a window: the pad is the whole point, so the + // entry is a hair either side of the cite and --pad does the rest. + entry = { + id: claim.id, + video: claim.video, + channel: claim.channel ?? null, + start: Math.max(0, at - 1), + end: at + 1, + }; + } + const meta = await videoMeta(entry.video, entry.channel ?? provenance.channelSlug, { siteChannel: entry.siteChannel, siteVideo: entry.siteVideo }); + const r = await fetchClip(entry, meta, render, dirs.rawDir, opts); + EMIT("done", { out: r.path, fetchStart: r.fetchStart, cached: r.cached }); + return { out: r.path, failures: [] }; + } + + // Footer chrome is shared by every clip, so build it once up front. + const hyper = render.chromeEngine === "hyperframes"; + const chrome = hyper + ? reservedFooter(render) + : await renderFooterAssets(render, manifest.timelineNodes, outDir); + + const entries = manifest.timeline.filter((e) => !only || e.id === only); + if (only && !entries.length) throw new Error(`no timeline entry with id ${only}`); + + // Read once, at the ROOT: a source's state is a fact about the manifest, not + // about a variant. A stacked ledger card says why each claim is text rather + // than footage, and this is where that answer comes from. + const availability = new Map( + ( + await readFile(path.join(dirs.root, "availability.json"), "utf8").then( + (j) => JSON.parse(j).sources ?? [], + () => [], + ) + ).flatMap((src) => (src.claims ?? []).map((id) => [id, src.state])), + ); + const segments = []; + const failures = []; + + const D = opts.noXfade || (render.transition ?? 0.5) === 0 ? 0 : render.transition ?? 0.5; + + // Retro-fit chapters onto an already-built file without re-encoding it. The + // per-clip segments are still on disk, which is all the offsets need. + if (opts.chaptersOnly) { + const finalPath = dirs.final; + const segs = entries.map((e) => path.join(outDir, "segments", `${e.id}.mp4`)); + for (const seg of segs) { + if (!(await exists(seg))) + throw new Error(`--chapters-only needs ${seg}, which is missing — run a full build first`); + } + await muxChapters(finalPath, entries, segs, D, outDir, provenance, render.fps); + return { out: finalPath, failures: [] }; + } + + // Re-run the rail over a cached concat instead of rebuilding the timeline. + // The rail is the part that gets iterated on; the 40-minute concat is not. + if (opts.railOnly) { + if (!render.rail) throw new Error("--rail-only needs render.rail in the manifest"); + const finalPath = dirs.final; + const prerail = prerailPath(outDir, manifest.slug, D); + const segs = entries.map((e) => path.join(outDir, "segments", `${e.id}.mp4`)); + for (const seg of segs) { + if (!(await exists(seg))) + throw new Error(`--rail-only needs ${seg}, which is missing — run a full build first`); + } + if (!(await exists(prerail))) { + EMIT("concat", { mode: D === 0 ? "hardcut" : "xfade", n: segs.length }); + if (D === 0) await concatHardCut(segs, outDir, prerail); + else await concatWithXfade(segs, render, prerail, null); + } + const railPlan = await buildRailPlan(manifest, render, entries, segs, D, outDir); + await assertConcatLength(prerail, railPlan.total, render.fps, + `cached ${path.basename(prerail)}`); + const out = opts.preview + ? path.join(outDir, `${manifest.slug}.preview.mp4`) + : finalPath; + await applyRail(prerail, out, render, railPlan, opts.preview ?? null); + if (!opts.preview) { + await assertConcatLength(out, railPlan.total, render.fps, "rail build"); + // applyRail re-encodes, so the chapters muxed onto the previous final are + // gone. Put them back, or --rail-only quietly ships a chapterless cut. + if (!opts.noChapters) { + await muxChapters(out, entries, segs, D, outDir, provenance, render.fps); + } + } + EMIT("done", { out, failures: [] }); + return { out, failures: [] }; + } + + EMIT("start", { title: manifest.title, entries: entries.length, out: outDir }); + for (let i = 0; i < entries.length; i += 1) { + const entry = entries[i]; + try { + if (entry.type === "card") { + EMIT("card", { id: entry.id, i, n: entries.length }); + segments.push(await buildCardSegment(entry, render, outDir, manifest.timelineNodes)); + } else if (entry.type === "scroll" || entry.type === "chart" || entry.type === "ledger") { + EMIT("card", { id: entry.id, i, n: entries.length }); + if (!manifest.ledger?.length) + throw new Error(`${entry.id} is type:${entry.type} but the manifest has no ledger[]`); + segments.push( + entry.type === "scroll" + ? await buildScrollSegment(entry, render, outDir, manifest.ledger) + : entry.type === "chart" + ? await buildChartSegment(entry, render, outDir, manifest.ledger) + : await buildLedgerSegment(entry, render, outDir, manifest.ledger, availability), + ); + } else { + const meta = await videoMeta(entry.video, entry.channel ?? provenance.channelSlug, { siteChannel: entry.siteChannel, siteVideo: entry.siteVideo }); + EMIT("clip", { + id: entry.id, i, n: entries.length, video: entry.video, + section: entry.section, sectionEnter: !!entry.sectionEnter, + }); + segments.push( + await buildClipSegment( + entry, meta, render, dirs, opts, chrome, manifest.timelineNodes, provenance, + ), + ); + } + EMIT("segment", { id: entry.id, path: segments[segments.length - 1] }); + } catch (err) { + // Without --continue-on-error a dead source at entry 14 of 19 throws away + // the thirteen fetches already paid for. With it, everything buildable is + // built and the run reports what was not. + if (!opts.continueOnError) throw err; + const message = err?.message ?? String(err); + failures.push({ id: entry.id, message }); + EMIT("entry-failed", { id: entry.id, message }); + } + } + + if (only) { + EMIT("done", { out: segments[0], failures }); + return { out: segments[0], failures }; + } + + // A timeline that silently lost a clip is a worse outcome than no file at all: + // the finished video would look complete and be missing a citation. So the + // segments are kept (they cost the fetches) and the concat is refused. + if (failures.length) { + EMIT("note", { + message: `refusing to concat: ${failures.length} of ${entries.length} entries failed ` + + `(${failures.map((f) => f.id).join(", ")})`, + }); + return { out: null, failures }; + } + + const final = dirs.final; + const railPlan = opts.noRail ? null : await buildRailPlan(manifest, render, entries, segments, D, outDir); + + // The rendered chrome, if this manifest asks for it. Absent, `chromePlan` is + // null and every line below behaves exactly as it did -- which is the claim + // the MD5 check tests. + let chromePlan = null; + if (hyper) { + const regions = chromeRegions(render, outDir); + for (const r of regions) { + if (!(await exists(path.join(r.frames, "frame_000001.png")))) { + throw new Error( + `render.chromeEngine is "hyperframes" but ${r.name} has no frames at ${r.frames}. ` + + `Run compose-chrome.mjs --region ${r.name} --render first.`, + ); + } + } + chromePlan = { regions, outLabel: "[hfout]" }; + EMIT("note", { message: `chrome: ${regions.map((r) => `${r.name} ${r.width}x${r.height}`).join(", ")} as png-sequence` }); + } + + // `transition: 0` is a real editorial choice, not just a speed knob: hard cuts + // hit harder on a compilation whose point is repetition. Honouring it here keeps + // the manifest the source of truth, so a rebuild does not silently re-add fades. + EMIT("concat", { mode: D === 0 ? "hardcut" : "xfade", n: segments.length }); + if (D === 0) { + if (chromePlan) { + throw new Error( + 'render.chromeEngine "hyperframes" needs a filtergraph, and `transition: 0` concatenates with ' + + "-c copy, which cannot host one. Give the manifest a transition, or drop chromeEngine.", + ); + } + // concatHardCut is `-c copy`, which cannot host a filtergraph, so the rail + // has to be a second pass here whether we like it or not. + const prerail = prerailPath(outDir, manifest.slug, D); + await concatHardCut(segments, outDir, railPlan ? prerail : final); + if (railPlan) { + await assertConcatLength(prerail, railPlan.total, render.fps, "hard-cut concat"); + await applyRail(prerail, final, render, railPlan, null); + } + } else { + await concatWithXfade(segments, render, final, railPlan, chromePlan); + } + + // Length is the canary for the two ways a rail input can go wrong: a file + // LONGER than the timeline means a strip outran the main (a missing + // shortest=1), and a hang means an unbounded -loop 1. + if (railPlan) await assertConcatLength(final, railPlan.total, render.fps, "rail build"); + + if (!opts.noChapters) await muxChapters(final, entries, segments, D, outDir, provenance, render.fps); + + const { stdout } = await execFileP(FFPROBE, [ + "-v", "error", "-show_entries", "format=duration,size", + "-of", "default=noprint_wrappers=1", final, + ]); + const probe = Object.fromEntries( + stdout.trim().split("\n").map((l) => l.split("=")), + ); + EMIT("done", { out: final, duration: Number(probe.duration), size: Number(probe.size) }); + return { out: final, failures }; +} + +async function main() { + const argv = process.argv.slice(2); + const manifestPath = argv.find((a) => !a.startsWith("--")); + if (!manifestPath) { + console.error( + "usage: build-video.mjs <manifest.json> [--out <dir>] [--variant sourced|full]\n" + + " [--only <id>] [--fetch-only <id>]\n" + + " [--pad <s>] [--skip-fetch] [--no-xfade] [--no-chapters] [--chapters-only]\n" + + " [--progress ndjson] [--continue-on-error] [--no-reuse]\n" + + " [--no-rail] [--rail-only] [--preview <start> <dur>]\n" + + " [--site-origin <url>] [--resolve-site-ids] [--cue-source auto|local|http]", + ); + process.exit(2); + } + const flag = (n) => { + const i = argv.indexOf(n); + return i >= 0 ? argv[i + 1] : undefined; + }; + setProgressMode(flag("--progress") ?? "human"); + + const padArg = flag("--pad"); + const opts = { + variant: flag("--variant") ?? "sourced", + skipFetch: argv.includes("--skip-fetch"), + continueOnError: argv.includes("--continue-on-error"), + noXfade: argv.includes("--no-xfade"), + noChapters: argv.includes("--no-chapters"), + chaptersOnly: argv.includes("--chapters-only"), + noReuse: argv.includes("--no-reuse"), + noRail: argv.includes("--no-rail"), + railOnly: argv.includes("--rail-only"), + pad: padArg === undefined ? undefined : Number(padArg), + siteOrigin: flag("--site-origin"), + resolveSiteIds: argv.includes("--resolve-site-ids"), + cueSource: flag("--cue-source"), + }; + const pv = argv.indexOf("--preview"); + if (pv >= 0) { + opts.preview = { start: Number(argv[pv + 1]), dur: Number(argv[pv + 2]) }; + opts.railOnly = true; + if (!Number.isFinite(opts.preview.start) || !Number.isFinite(opts.preview.dur)) { + console.error("--preview takes <start> <dur> in seconds"); + process.exit(2); + } + } + + const { failures } = await buildVideo({ + manifestPath, + opts, + out: flag("--out"), + only: flag("--only"), + fetchOnly: flag("--fetch-only"), + }); + // Non-zero on a partial run, so a caller that ignores the events still learns + // the build did not produce what was asked for. + if (failures.length) process.exit(1); +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main().catch((err) => { + EMIT("error", { message: err?.message ?? String(err) }); + console.error(err.message ?? err); + process.exit(1); + }); +} diff --git a/umtool/report-to-video/check-availability.mjs b/umtool/report-to-video/check-availability.mjs @@ -0,0 +1,182 @@ +#!/usr/bin/env node +// check-availability.mjs — is every source this manifest cites still fetchable? +// +// This is the one fact about a report video that goes stale in BOTH directions +// and that nothing on disk records. A source can be deleted between writing the +// manifest and building it (so a 40-minute build dies at clip 14 having paid for +// thirteen fetches), and a source can come back (so a manifest annotated "gone" +// stays wrong). Neither is visible until a build runs. +// +// So it runs first, it runs cheap, and it writes down when it ran. `--simulate` +// resolves formats without downloading a byte: a few seconds for a whole +// manifest against twenty-odd minutes for the build it protects. +// +// On the CLI: +// node umtool/report-to-video/check-availability.mjs <manifest.json> [--json] +// +// Options: +// --json Print the report as JSON instead of a table +// --out <dir> Output root (default: manifest dir + /out) +// --allow-missing Exit 0 even when a source is gone (report only) +// --max-age <days> Reuse a recorded verdict younger than this (default: 0) + +import { execFile } from "node:child_process"; +import { promisify } from "node:util"; +import { mkdir, readFile, writeFile } from "node:fs/promises"; +import path from "node:path"; + +import { DEFAULT_CHANNELS_DIR } from "./cues.mjs"; + +const execFileP = promisify(execFile); + +const YTDLP = process.env.YTDLP_BIN ?? "yt-dlp"; +// Resolved relative to the repo (see cues.mjs) rather than an absolute path in +// one machine's home directory, which every other clone would miss. +const CHANNELS_DIR = DEFAULT_CHANNELS_DIR; + +// yt-dlp says why in prose, and the distinction matters editorially: a private +// or removed video needs the clip converting to a quote card, while a network +// blip needs a retry. Anything unrecognised stays `maybe_missing` rather than +// being called deleted -- claiming a source is gone when it is not is the more +// expensive mistake, because it invites deleting a citation. +function classify(stderr) { + const t = String(stderr ?? ""); + if (/Private video|private/i.test(t)) return "private"; + if (/removed by the uploader|has been removed|no longer available|Video unavailable|does not exist|410/i.test(t)) + return "deleted"; + if (/age.?restrict|Sign in to confirm|confirm your age/i.test(t)) return "restricted"; + if (/members-only|join this channel/i.test(t)) return "members-only"; + if (/geo|not available in your country/i.test(t)) return "geo-blocked"; + return "maybe_missing"; +} + +async function cueMeta(videoId, channelSlug) { + const p = path.join(CHANNELS_DIR, channelSlug, "data", videoId, "transcript.cues.json"); + const d = JSON.parse(await readFile(p, "utf8")); + return { title: d.title, webpageUrl: d.webpageUrl, duration: d.duration }; +} + +export async function checkAvailability(manifestPath, { outDir, maxAgeDays = 0 } = {}) { + const manifest = JSON.parse(await readFile(manifestPath, "utf8")); + const slug = manifest.provenance?.channelSlug; + const dir = outDir ?? path.join(path.dirname(path.resolve(manifestPath)), "out"); + const file = path.join(dir, "availability.json"); + + // A clip may name its own channel: the same streamer's VODs are mirrored + // across more than one archive, and the same id under a different slug is a + // different file. So the unit of work is (channel, video), never video alone. + const wanted = new Map(); + const want = (channel, video) => { + const key = `${channel}/${video}`; + if (!wanted.has(key)) wanted.set(key, { key, channel, video, clips: [], claims: [] }); + return wanted.get(key); + }; + for (const e of manifest.timeline ?? []) { + if (e.type !== "clip") continue; + want(e.channel ?? slug, e.video).clips.push(e.id); + } + // LEDGER SOURCES TOO, not just the clipped ones. A cut that stacks its + // unclipped claims onto cards has to say WHY each one is a line of text + // rather than footage, and "the upload is gone" and "we did not cut it" are + // different sentences. Guessing which is which is how a live source ends up + // labelled deleted on screen. + for (const c of manifest.ledger ?? []) { + if (!c.video) continue; + want(c.channel ?? slug, c.video).claims.push(c.id); + } + + const prev = await readFile(file, "utf8").then( + (s) => JSON.parse(s), + () => ({ sources: [] }), + ); + const prevBy = new Map((prev.sources ?? []).map((s) => [s.key, s])); + const freshMs = maxAgeDays * 86400_000; + + const sources = []; + for (const w of wanted.values()) { + const was = prevBy.get(w.key); + if (freshMs > 0 && was?.checkedAt && Date.now() - Date.parse(was.checkedAt) < freshMs) { + sources.push({ ...was, clips: w.clips, claims: w.claims, reused: true }); + continue; + } + + let meta; + try { + meta = await cueMeta(w.video, w.channel); + } catch { + // No cue file is a DIFFERENT failure from a dead source, and it is the one + // the Rumble two-ids trap produces: the manifest names the MCP video id + // while the cues live under the URL slug. Build would die here too, so it + // is reported here rather than discovered twenty minutes in. + sources.push({ + ...w, ok: false, state: "no-cues", checkedAt: new Date().toISOString(), + error: `no transcript.cues.json under ${w.channel}/data/${w.video}`, + }); + continue; + } + + try { + await execFileP( + YTDLP, + ["--ignore-config", "--no-playlist", "--simulate", "--quiet", "--no-warnings", "--", meta.webpageUrl], + { maxBuffer: 1 << 24 }, + ); + sources.push({ + ...w, ok: true, state: "ok", title: meta.title, url: meta.webpageUrl, + checkedAt: new Date().toISOString(), error: null, + }); + } catch (err) { + const stderr = err?.stderr ?? err?.message ?? ""; + sources.push({ + ...w, ok: false, state: classify(stderr), title: meta.title, url: meta.webpageUrl, + checkedAt: new Date().toISOString(), error: String(stderr).trim().split("\n").slice(-3).join(" "), + }); + } + } + + const report = { manifest: path.resolve(manifestPath), checkedAt: new Date().toISOString(), sources }; + await mkdir(dir, { recursive: true }); + await writeFile(file, JSON.stringify(report, null, 2) + "\n", "utf8"); + return { ...report, file }; +} + +async function main() { + const argv = process.argv.slice(2); + const manifestPath = argv.find((a) => !a.startsWith("--")); + if (!manifestPath) { + console.error("usage: check-availability.mjs <manifest.json> [--json] [--out <dir>] [--allow-missing]"); + process.exit(2); + } + const flag = (n) => { + const i = argv.indexOf(n); + return i >= 0 ? argv[i + 1] : undefined; + }; + const report = await checkAvailability(manifestPath, { + outDir: flag("--out"), + maxAgeDays: Number(flag("--max-age") ?? 0), + }); + + if (argv.includes("--json")) { + console.log(JSON.stringify(report, null, 2)); + } else { + for (const s of report.sources) { + const mark = s.ok ? "ok " : "GONE"; + console.log( + `${mark} ${s.key.padEnd(40)} ${String(s.state).padEnd(14)} ` + + `${s.clips.length} clip(s)${s.reused ? " (cached)" : ""}`, + ); + if (!s.ok && s.error) console.log(` ${s.error}`); + } + const bad = report.sources.filter((s) => !s.ok).length; + console.log(`\n${report.sources.length} source(s), ${bad} unavailable -> ${report.file}`); + } + + if (report.sources.some((s) => !s.ok) && !argv.includes("--allow-missing")) process.exit(1); +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main().catch((err) => { + console.error(err.message ?? err); + process.exit(1); + }); +} diff --git a/scripts/report-to-video/compose-chrome.mjs b/umtool/report-to-video/compose-chrome.mjs diff --git a/umtool/report-to-video/cues.mjs b/umtool/report-to-video/cues.mjs @@ -0,0 +1,270 @@ +// Where a clip's caption cues come from. +// +// A clip window is widened from a cue span to a whole sentence, which needs cue +// END times. Nothing else in the pipeline carries them: a report citation is a +// single start second, and the MCP `Snippet` type has no `end` field. So this is +// the one place that answers "what are the real cue boundaries for this video". +// +// TWO SOURCES, SAME SHAPE. A local corpus stores each video as +// `<CHANNELS_DIR>/<slug>/data/<id>/transcript.cues.json`, and a *published* +// archive serves the same record inside a paginated shard. The two carry the +// same fields — `{ slug, id, channelSlug, title, uploadDate, duration, channel, +// description, platform, webpageUrl, cues: [{start, end, text}] }` — so one +// resolver serves both `loadCues` and `videoMeta`, and a caller cannot tell +// which it got beyond the `from` marker. +// +// That parity is what makes a corpus optional. Clone the repo, point a manifest +// at a public instance, and the video pipeline can cut clips without mirroring a +// single channel: the cue windows come over HTTP, and the media itself was +// always a network fetch (`yt-dlp --download-sections`). +// +// The shard walk is the contract published at `/corpus.json` under `shardScheme`: +// 1. GET <origin>/corpus.json -> channels[].manifests.transcripts +// 2. GET that manifest -> { pageCount, slugToPage: { <id>: N } } +// 3. GET page-<NNNN>.json (N zero-padded to 4) -> array of records +// 4. take the record whose `id` matches +// +// Local wins when present: it is faster, works offline, and is the operator's own +// data. HTTP is the fallback, not a preference. +// +// THE TWO SOURCES CAN DISAGREE, AND IT IS NOT ROUNDING. A published archive is a +// snapshot; a live corpus keeps moving. Re-synced platform captions, an +// auto-caption replacement or a re-transcription all rewrite a video's cues in +// place, and the archive keeps the text it was built from until it is rebuilt. +// Measured on this corpus (local 2026-08-13 against a 2026-08-07 publish): of +// four videos checked, three were byte-identical and one had 65 of its 84 cue +// texts changed with timings shifted by up to **2.24 s** — enough to cut a clip +// in the wrong place. +// +// So `prefer` is a real decision, not a micro-optimisation: +// "auto" (default) local when present, else HTTP. Right for an operator. +// "local" never fall back. Fail loudly instead of silently cutting from +// different cues than the ones a window was authored against. +// "http" always the archive. Right when you want the windows to match what a +// reader following the citation will actually see, and the only +// option that is reproducible on a machine with no corpus. +// Whatever answers, the returned record carries `from` so a caller can record it. + +import { readFile, writeFile, mkdir } from "node:fs/promises"; +import path from "node:path"; +import os from "node:os"; +import { createHash } from "node:crypto"; +import { fileURLToPath } from "node:url"; + +// This file lives at <repo>/umtool/report-to-video/, so the corpus a plain +// checkout would have is two levels up. Previously this defaulted to an absolute +// path inside the original author's home directory, which meant every other +// clone silently looked in a directory that does not exist. +const REPO_ROOT = path.resolve(path.dirname(fileURLToPath(import.meta.url)), "..", ".."); + +export const DEFAULT_CHANNELS_DIR = + process.env.CHANNELS_DIR ?? path.join(REPO_ROOT, "transcripts", "channels"); + +const DEFAULT_CACHE_DIR = + process.env.REPORT_CACHE_DIR ?? + path.join(os.homedir(), ".cache", "archilyzer-report-to-video"); + +// A shard page is capped at 8 MB and holds ~100 videos, so refetching one per +// clip — across two separate processes, resolve-windows then build-video — is +// the difference between usable and painful. Cached by URL on disk; archives are +// rebuilt rarely and a stale page only matters if the cues themselves changed. +function cacheKey(url) { + return createHash("sha1").update(url).digest("hex") + ".json"; +} + +export function pageFileName(pageNumber) { + return `page-${String(pageNumber).padStart(4, "0")}.json`; +} + +export function pageUrlFrom(manifestUrl, pageNumber) { + const u = new URL(manifestUrl); + u.pathname = u.pathname.replace(/[^/]+$/, pageFileName(pageNumber)); + return u.toString(); +} + +// The origin of the archive a manifest was built against. Every manifest already +// records this — `siteOrigin` explicitly, and `corpus` as `remote:<url>` or +// `local:<path>` — so the common case needs no configuration at all. +export function siteOriginFromManifest(manifest) { + const p = manifest?.provenance ?? {}; + if (typeof p.siteOrigin === "string" && p.siteOrigin.trim()) { + return p.siteOrigin.replace(/\/+$/, ""); + } + if (typeof p.corpus === "string" && p.corpus.startsWith("remote:")) { + return p.corpus.slice("remote:".length).replace(/\/+$/, ""); + } + if (typeof p.shareLink === "string" && /^https?:/.test(p.shareLink)) { + try { + return new URL(p.shareLink).origin; + } catch { + /* fall through */ + } + } + return null; +} + +export class CueLookupError extends Error { + constructor(message, { channelSlug, videoId, tried }) { + super(message); + this.name = "CueLookupError"; + this.channelSlug = channelSlug; + this.videoId = videoId; + this.tried = tried; + } +} + +export function createCueSource({ + channelsDir = DEFAULT_CHANNELS_DIR, + siteOrigin = null, + cacheDir = DEFAULT_CACHE_DIR, + fetchImpl = globalThis.fetch, + log = () => {}, + // "auto" | "local" | "http" — see the note on divergence above. + prefer = "auto", + // Opt-in, because it is expensive: see resolveSiteId below. + resolveSiteIds = false, +} = {}) { + const mem = new Map(); + + async function getJson(url) { + if (mem.has(url)) return mem.get(url); + const disk = cacheDir ? path.join(cacheDir, cacheKey(url)) : null; + if (disk) { + try { + const cached = JSON.parse(await readFile(disk, "utf8")); + mem.set(url, cached); + return cached; + } catch { + /* cold cache */ + } + } + log(`fetch ${url}`); + const res = await fetchImpl(url); + if (!res.ok) throw new Error(`GET ${url} -> ${res.status}`); + const json = await res.json(); + mem.set(url, json); + if (disk) { + try { + await mkdir(path.dirname(disk), { recursive: true }); + await writeFile(disk, JSON.stringify(json)); + } catch { + // A cache we cannot write is a slow run, not a failed one. + } + } + return json; + } + + async function channelEntry(origin, channelSlug) { + const corpus = await getJson(`${origin}/corpus.json`); + const found = (corpus.channels ?? []).find((c) => c.slug === channelSlug); + if (!found) { + throw new CueLookupError( + `channel "${channelSlug}" is not in ${origin}/corpus.json`, + { channelSlug, videoId: null, tried: [`${origin}/corpus.json`] }, + ); + } + return found; + } + + // Map a LOCAL video id onto the id the site serves, by scanning the channel's + // pages for a record whose `webpageUrl` contains it. + // + // WHY THIS EXISTS: a Rumble video has two ids. The site (and the MCP) key it by + // the EMBED id; the local cue directory is named for the URL SLUG. A manifest + // hand-authored against local cue dirs therefore carries slugs that are absent + // from the published `slugToPage` — every Rumble clip misses. + // + // WHY IT IS OPT-IN: it downloads a channel's shards until it hits a match, and + // a shard is up to 8 MB. That is a reasonable price to pay knowingly and a + // terrible one to pay silently, so the direct lookup fails with instructions + // instead and this runs only when asked. + async function resolveSiteId(origin, entry, wanted) { + const manifest = await getJson(entry.manifests.transcripts); + log(`resolving "${wanted}" by scanning ${manifest.pageCount} shard(s) of ${entry.slug}`); + for (let n = 0; n < manifest.pageCount; n += 1) { + const page = await getJson(pageUrlFrom(entry.manifests.transcripts, n)); + const hit = page.find( + (r) => r.id === wanted || r.slug === wanted || String(r.webpageUrl ?? "").includes(wanted), + ); + if (hit) return hit; + } + return null; + } + + async function fromHttp(channelSlug, videoId, hints) { + const origin = hints.siteOrigin ?? siteOrigin; + if (!origin) { + throw new CueLookupError( + `no local cues for ${channelSlug}/${videoId} and no archive origin to fetch them from ` + + `(set provenance.siteOrigin in the manifest, or pass --site-origin / SITE_ORIGIN)`, + { channelSlug, videoId, tried: ["local"] }, + ); + } + const siteChannel = hints.siteChannel ?? channelSlug; + const siteVideo = hints.siteVideo ?? videoId; + const entry = await channelEntry(origin, siteChannel); + const manifest = await getJson(entry.manifests.transcripts); + const pageNumber = manifest.slugToPage?.[siteVideo]; + + if (pageNumber === undefined) { + if (resolveSiteIds) { + const hit = await resolveSiteId(origin, entry, siteVideo); + if (hit) return { ...hit, from: "http" }; + } + throw new CueLookupError( + `"${siteVideo}" is not in ${siteChannel}'s published slugToPage on ${origin}.\n` + + ` If this is a Rumble clip, the archive is keyed by the EMBED id while a local cue\n` + + ` directory is named for the URL SLUG — they differ. Either add "siteVideo" (and\n` + + ` "siteChannel" if it also differs) to this clip in the manifest, or re-run with\n` + + ` --resolve-site-ids to find it by scanning the channel's shards (slow: downloads\n` + + ` up to 8 MB per shard until it matches).\n` + + ` Note that a clip's citeUrl is NOT usable here — it may deliberately cite a\n` + + ` different recording (a mirror that reads better), whose clock is not the same.`, + { channelSlug, videoId, tried: [entry.manifests.transcripts] }, + ); + } + + const page = await getJson(pageUrlFrom(entry.manifests.transcripts, pageNumber)); + const record = page.find((r) => r.id === siteVideo || r.slug === siteVideo); + if (!record) { + throw new CueLookupError( + `${siteChannel}/${siteVideo} is on shard ${pageNumber} per the manifest, but no record ` + + `there has that id — the published archive is inconsistent`, + { channelSlug, videoId, tried: [pageUrlFrom(entry.manifests.transcripts, pageNumber)] }, + ); + } + return { ...record, from: "http" }; + } + + async function fromLocal(channelSlug, videoId) { + const p = path.join(channelsDir, channelSlug, "data", videoId, "transcript.cues.json"); + const parsed = JSON.parse(await readFile(p, "utf8")); + return { ...parsed, from: "local" }; + } + + return { + channelsDir, + /** + * The full record for one video: cues plus the metadata build-video needs. + * `hints` may carry `siteChannel` / `siteVideo` (when the published archive + * keys this recording differently) and `siteOrigin` (per-manifest override). + */ + prefer, + async load(channelSlug, videoId, hints = {}) { + if (prefer === "http") return await fromHttp(channelSlug, videoId, hints); + try { + return await fromLocal(channelSlug, videoId); + } catch (err) { + if (err?.code !== "ENOENT" && err?.code !== "ENOTDIR") throw err; + if (prefer === "local") { + throw new CueLookupError( + `no local cues for ${channelSlug}/${videoId} under ${channelsDir}, and ` + + `--cue-source local forbids falling back to the archive`, + { channelSlug, videoId, tried: [channelsDir] }, + ); + } + return await fromHttp(channelSlug, videoId, hints); + } + }, + }; +} diff --git a/scripts/report-to-video/cues.test.mjs b/umtool/report-to-video/cues.test.mjs diff --git a/scripts/report-to-video/ledger-totals.mjs b/umtool/report-to-video/ledger-totals.mjs diff --git a/scripts/report-to-video/ledger-totals.test.mjs b/umtool/report-to-video/ledger-totals.test.mjs diff --git a/umtool/report-to-video/package.json b/umtool/report-to-video/package.json @@ -0,0 +1,25 @@ +{ + "name": "umtool-report-to-video", + "version": "0.1.0", + "private": true, + "type": "module", + "description": "Turn a cited sweep report into a narrated-by-text video.", + "bin": { + "report-build-video": "./build-video.mjs", + "report-resolve-windows": "./resolve-windows.mjs", + "report-check-availability": "./check-availability.mjs", + "report-verify-build": "./verify-build.mjs", + "report-compose-chrome": "./compose-chrome.mjs" + }, + "exports": { + "./build-video": "./build-video.mjs", + "./check-availability": "./check-availability.mjs", + "./compose-chrome": "./compose-chrome.mjs", + "./cues": "./cues.mjs", + "./ledger-totals": "./ledger-totals.mjs", + "./package.json": "./package.json", + "./render-cards": "./render-cards.mjs", + "./resolve-windows": "./resolve-windows.mjs", + "./verify-build": "./verify-build.mjs" + } +} diff --git a/umtool/report-to-video/render-cards.mjs b/umtool/report-to-video/render-cards.mjs @@ -0,0 +1,1530 @@ +#!/usr/bin/env node +// render-cards.mjs — turn a video manifest's `card` entries into PNG stills. +// +// One PNG per card, written to <outDir>/cards/<id>.png at the manifest's render +// resolution. Text is laid out by ImageMagick's Pango delegate rather than +// ffmpeg's drawtext: Pango wraps, kerns and takes inline markup, so a card is a +// single markup string instead of a stack of hand-positioned drawtext filters. +// +// Card styles (manifest `style` field): +// title — the opening card: big heading, subtitle, provenance footer +// chapter — an act break: small amber kicker over a large heading +// bullets — heading plus a list of caveats +// sources — closing attribution +// +// `status` is RETIRED. It existed to quote a claim whose source had gone, and +// the `ledger` entry type does that better: it says the same words, alongside +// the arithmetic the claim moves, and it says WHY there is no footage from a +// probe rather than from a hand-written kicker that nothing re-checks. +// +// In the app: not used. On the CLI: +// node umtool/report-to-video/render-cards.mjs <manifest.json> [--out <dir>] +// +// Options: +// --out <dir> Output root (default: the manifest's directory + /out) +// --only <id> Render just one card, by manifest id +// +// Requires: ImageMagick built with Pango (magick -list format | grep PANGO). + +import { execFile } from "node:child_process"; +import { promisify } from "node:util"; +import { mkdir, writeFile, readFile } from "node:fs/promises"; +import path from "node:path"; +import { dateKey, ledgerTotals, rosterLine } from "./ledger-totals.mjs"; + +const execFileP = promisify(execFile); + +const RSVG = process.env.RSVG_BIN ?? "rsvg-convert"; +const QRENCODE = process.env.QRENCODE_BIN ?? "qrencode"; + +// Every card is drawn at the manifest's full `width`, but once a rail column is +// configured the RIGHT `rail.width` pixels of the frame belong to it. Text, +// rules and timeline nodes therefore lay out inside the CONTENT width, while the +// canvas stays full-width — the card ground is flat `pal.bg`, which is exactly +// what the rail wants behind it. +export const railWidth = (render) => render.rail?.width ?? 0; +export const contentWidth = (render) => render.width - railWidth(render); + +/** + * How wide a card's content may be. + * + * `hideRail` cards slide the rail off over their own dissolve, so they get the + * WHOLE frame. Everything else lays out inside the content width and leaves the + * rail column as ground. + */ +export const cardWidth = (card, render) => + card?.hideRail ? render.width : contentWidth(render); + +// Pango markup is XML-ish, so anything we interpolate has to be escaped first. +// Curly quotes and the ellipsis pass through fine; only these five matter. +function esc(s) { + return String(s) + .replace(/&/g, "&amp;") + .replace(/</g, "&lt;") + .replace(/>/g, "&gt;") + .replace(/"/g, "&quot;") + .replace(/'/g, "&apos;"); +} + +function span(text, { size, color, weight, family = "Fira Sans" }) { + const attrs = [`font_family="${family}"`, `size="${Math.round(size * 1024)}"`]; + if (color) attrs.push(`foreground="${color}"`); + if (weight) attrs.push(`weight="${weight}"`); + return `<span ${attrs.join(" ")}>${text}</span>`; +} + +// Each style returns Pango markup for the whole card body. Blank lines are real +// newlines in the markup — Pango honours them, which is how vertical rhythm is +// set without positioning each run separately. +function markupFor(card, pal) { + const H = (t, size = 62) => + span(esc(t), { size, color: pal.fg, weight: "bold" }); + const KICKER = (t) => + span(esc(t.toUpperCase()), { size: 24, color: pal.amber, weight: "bold" }); + const SUB = (t, size = 30) => span(esc(t), { size, color: pal.muted }); + + switch (card.style) { + case "title": + return [ + span(esc(card.heading), { size: 82, color: pal.fg, weight: "bold" }), + "", + span(esc(card.sub), { size: 38, color: pal.accent }), + "", + "", + SUB(card.foot, 24), + ].join("\n"); + + case "chapter": + return [ + card.kicker ? KICKER(card.kicker) : null, + card.kicker ? "" : null, + H(card.heading), + card.sub ? "" : null, + card.sub ? SUB(card.sub, 32) : null, + ] + .filter((l) => l !== null) + .join("\n"); + + case "bullets": { + const items = (card.bullets ?? []).flatMap((b) => [ + `${span("— ", { size: 30, color: pal.accent, weight: "bold" })}${span( + esc(b), + { size: 30, color: pal.fg }, + )}`, + "", + ]); + return [H(card.heading, 52), "", ...items].join("\n"); + } + + case "sources": + return [ + H(card.heading, 52), + "", + card.sub ? SUB(card.sub, 32) : null, + card.foot ? "" : null, + card.foot + ? card.foot + .split("\n") + .map((l) => SUB(l, 24)) + .join("\n") + : null, + ] + .filter((l) => l !== null) + .join("\n"); + + default: + return H(card.heading ?? card.id); + } +} + +// A chapter break that just states a heading is dead air — it stops the video to +// say something the next clip is about to say anyway. This draws the whole +// project arc instead, with the current step lit and everything before it +// filled, so the pause carries information: where we are and how far is left. +// +// Nodes come from the manifest's top-level `timelineNodes`; the card names its +// position with `step` (0-based). +async function renderTimelineCard(card, render, nodes, outDir) { + const pal = render.palette; + const { width, height } = render; + const outPath = path.join(outDir, "cards", `${card.id}.png`); + const dir = path.join(outDir, "cards"); + + const VW = cardWidth(card, render); + const x0 = 260; + const x1 = VW - 260; + const axisY = Math.round(height * 0.56); + const gap = (x1 - x0) / (nodes.length - 1); + const xs = nodes.map((_, i) => Math.round(x0 + i * gap)); + const cur = card.step; + + const args = ["-size", `${width}x${height}`, `xc:${pal.bg}`, "-strokewidth", "4"]; + + // Track: filled up to the current node, dim beyond it. + args.push( + "-stroke", pal.muted, "-fill", "none", + "-draw", `line ${xs[0]},${axisY} ${xs[xs.length - 1]},${axisY}`, + ); + if (cur > 0) { + args.push("-stroke", pal.accent, "-draw", `line ${xs[0]},${axisY} ${xs[cur]},${axisY}`); + } + + // Nodes: past and present filled, future hollow. The current one is larger and + // amber so the eye lands on it without needing a label to say "you are here". + nodes.forEach((_, i) => { + const r = i === cur ? 19 : 11; + const color = i === cur ? pal.amber : i < cur ? pal.accent : pal.bg; + args.push( + "-stroke", i <= cur ? (i === cur ? pal.amber : pal.accent) : pal.muted, + "-fill", color, + "-draw", `circle ${xs[i]},${axisY} ${xs[i] + r},${axisY}`, + ); + }); + + args.push("-stroke", "none"); + + // Heading, centred over the whole card. + const headMarkup = [ + span(esc((card.kicker ?? nodes[cur].label).toUpperCase()), { + size: 26, color: pal.amber, weight: "bold", + }), + "", + span(esc(card.heading ?? nodes[cur].title), { size: 62, color: pal.fg, weight: "bold" }), + card.sub ? "" : null, + card.sub ? span(esc(card.sub), { size: 30, color: pal.muted }) : null, + ] + .filter((l) => l !== null) + .join("\n"); + + // Heading, left-aligned on the same margin the other card styles use. + const headPath = path.join(dir, `${card.id}.head.pango`); + await writeFile(headPath, headMarkup, "utf8"); + args.push( + "(", "-size", `${VW - 460}x`, "-background", "none", + "-define", `pango:width=${VW - 460}`, + `pango:@${headPath}`, ")", + "-gravity", "NorthWest", + "-geometry", `+${Math.round(VW * 0.09) + 58}+${Math.round(height * 0.19)}`, + "-composite", + ); + + // Per-node date labels, centred under their dot. ImageMagick's + // `pango:alignment` define does not actually centre the text inside the box, + // so measure each rendered label and place it by hand instead of trusting it. + for (let i = 0; i < nodes.length; i += 1) { + const isCur = i === cur; + const labMarkup = span(esc(nodes[i].label), { + size: isCur ? 24 : 21, + color: isCur ? pal.fg : i < cur ? pal.muted : "#5c5570", + weight: isCur ? "bold" : "normal", + }); + const labPath = path.join(dir, `${card.id}.n${i}.pango`); + const labPng = path.join(dir, `${card.id}.n${i}.png`); + await writeFile(labPath, labMarkup, "utf8"); + await execFileP("magick", ["-background", "none", `pango:@${labPath}`, labPng]); + const { stdout } = await execFileP("magick", ["identify", "-format", "%w", labPng]); + const w = Number(stdout.trim()); + args.push(labPng, "-geometry", `+${xs[i] - Math.round(w / 2)}+${axisY + 44}`, "-composite"); + } + + args.push(outPath); + await execFileP("magick", args, { maxBuffer: 1 << 24 }); + return outPath; +} + +// Footer chrome, drawn once and overlaid on every clip: a track, a dot per +// section and its label. The *progress* along it is not baked in here — the fill +// bar and the amber marker are drawn by ffmpeg at encode time so they can slide +// between sections instead of cutting. See buildClipSegment in build-video.mjs. +// +// Returns the geometry the encoder needs to place those moving parts. +export async function renderFooterAssets(render, nodes, outDir) { + const pal = render.palette; + const { width } = render; + const FH = render.footerHeight ?? 92; + const dir = path.join(outDir, "cards"); + + // A cut whose clips are not a progression through time has nothing for a + // timeline to say, and a footer drawn anyway is chrome that has not earned its + // place. No nodes (or an explicit zero height) means no footer at all — the + // caller letterboxes against the header alone. + if (!nodes?.length || FH === 0) { + return { + footer: null, marker: null, bar: null, trackLen: 0, + footerHeight: 0, trackY: 0, xs: [], x0: 0, markerRadius: 0, + }; + } + + // The band itself stays full-frame so it reads as one strip running under the + // rail; only the TRACK is pulled in to the content width. + const x0 = 200; + const x1 = contentWidth(render) - 200; + const trackY = 26; + const gap = (x1 - x0) / (nodes.length - 1); + const xs = nodes.map((_, i) => Math.round(x0 + i * gap)); + + const footer = path.join(dir, "_footer.png"); + const args = [ + "-size", `${width}x${FH}`, `xc:${pal.bg}`, + "-strokewidth", "3", + "-stroke", "#3a3450", "-fill", "none", + "-draw", `line ${xs[0]},${trackY} ${xs[xs.length - 1]},${trackY}`, + "-stroke", "none", + ]; + for (const x of xs) { + args.push("-fill", "#4a4363", "-draw", `circle ${x},${trackY} ${x + 6},${trackY}`); + } + + // Two lines per node: what happened, then when. Both are measured and placed + // by hand — ImageMagick's `pango:alignment` define does not actually centre + // text inside its box. + for (let i = 0; i < nodes.length; i += 1) { + const lines = [ + { text: nodes[i].label, size: 18, color: pal.fg, dy: 18 }, + { text: nodes[i].date, size: 16, color: pal.muted, dy: 42 }, + ]; + for (const [k, ln] of lines.entries()) { + const pPath = path.join(dir, `_footer.n${i}.l${k}.pango`); + const pPng = path.join(dir, `_footer.n${i}.l${k}.png`); + await writeFile(pPath, span(esc(ln.text), { size: ln.size, color: ln.color }), "utf8"); + await execFileP("magick", ["-background", "none", `pango:@${pPath}`, pPng]); + const { stdout } = await execFileP("magick", ["identify", "-format", "%w", pPng]); + args.push( + pPng, + "-geometry", `+${xs[i] - Math.round(Number(stdout.trim()) / 2)}+${trackY + ln.dy}`, + "-composite", + ); + } + } + args.push(footer); + await execFileP("magick", args, { maxBuffer: 1 << 24 }); + + // The fill bar, as a strip to be TRANSLATED under a fixed crop rather than a + // drawbox whose width depends on `t`. drawbox has no time variable — its `t` + // is the box thickness — so the width expression the encoder used to build + // never evaluated and the bar was always full. Left half accent, right half + // transparent: sliding the crop window left across it grows the accent run. + const trackLen = xs[xs.length - 1] - x0; + const bar = path.join(dir, "_bar.png"); + await execFileP("magick", [ + "-size", `${trackLen * 2}x3`, "xc:none", + "-fill", pal.accent, "-stroke", "none", + "-draw", `rectangle 0,0 ${trackLen - 1},2`, + bar, + ]); + + // The travelling marker. + const marker = path.join(dir, "_marker.png"); + const r = 11; + await execFileP("magick", [ + "-size", `${r * 2 + 2}x${r * 2 + 2}`, "xc:none", + "-fill", pal.amber, "-stroke", "none", + "-draw", `circle ${r + 1},${r + 1} ${r * 2 + 1},${r + 1}`, + marker, + ]); + + return { footer, marker, bar, trackLen, footerHeight: FH, trackY, xs, x0, markerRadius: r }; +} + +// =========================================================================== +// The claim rail +// =========================================================================== +// A vertical ledger down the right edge that appends one row per claim as the +// video runs, with a live per-company tally beside it. The dates and the numbers +// are SPOKEN in the clips and shown only in the header citation line, so a viewer +// can hear "nearly ten" three times without ever seeing that the three refer to +// three different companies. The rail is what makes that visible. +// +// Every asset here is a STRIP, not a per-state still: one tall PNG whose window +// ffmpeg slides with an animated `crop`. That is the whole trick — swapping +// stills can only cut, but a crop can ease, and one input per moving part keeps +// the filtergraph small enough that ffmpeg does not deadlock on chained +// `-loop 1` inputs. See railFilterChain in build-video.mjs for the ramps. +// +// Assets are authored as SVG and rasterized with rsvg-convert rather than drawn +// with ImageMagick primitives: the strips need right-aligned columns, hairlines +// and ~500 individually-placed text runs, and one rsvg call beats four hundred +// `magick` invocations. Fonts inside the SVG resolve through fontconfig, so the +// family name has to MATCH the Pango cards ("Fira Sans"), not the font FILE that +// render.fontRegular points at. + +const RAIL_FONT = "Fira Sans"; + +// A tiny SVG text run. Everything is placed absolutely — no flow, no wrapping. +function svgText(x, y, text, o = {}) { + const a = [ + `x="${x}"`, `y="${y}"`, + `font-family="${o.family ?? RAIL_FONT}"`, + `font-size="${o.size ?? 15}"`, + `fill="${o.color}"`, + ]; + if (o.weight) a.push(`font-weight="${o.weight}"`); + if (o.anchor) a.push(`text-anchor="${o.anchor}"`); + if (o.ls) a.push(`letter-spacing="${o.ls}"`); + if (o.opacity != null) a.push(`opacity="${o.opacity}"`); + return `<text ${a.join(" ")}>${esc(text)}</text>`; +} + +const svgDoc = (w, h, body) => + `<svg xmlns="http://www.w3.org/2000/svg" width="${w}" height="${h}" ` + + `viewBox="0 0 ${w} ${h}">${body}</svg>`; + +// rsvg-convert is deterministic about output size in a way ImageMagick's RSVG +// delegate is not (its -density is ignored for sizing in some builds), so the +// pixel dimensions ffmpeg's crop arithmetic depends on are guaranteed here. +async function rasterize(svg, svgPath, pngPath, w, h) { + await writeFile(svgPath, svg, "utf8"); + await execFileP(RSVG, ["-w", String(w), "-h", String(h), "-o", pngPath, svgPath], { + maxBuffer: 1 << 26, + }); + return pngPath; +} + +// Truncate to a pixel budget. Fira Sans at these sizes averages ~0.50em per +// character; a hair conservative is right, because an overflowing row would run +// under the value column rather than wrap. +function fit(text, size, maxPx) { + const max = Math.max(4, Math.floor(maxPx / (size * 0.5))); + const t = String(text ?? ""); + return t.length <= max ? t : `${t.slice(0, max - 1).trimEnd()}…`; +} + +/** + * Every pixel measurement the rail needs, derived once so the asset builder and + * the filtergraph builder cannot disagree about a single one of them. + * + * The log window is deliberately sized to a WHOLE number of rows and butted + * against the bottom of the rail column: that is what lets the curtain (below) + * park exactly one window-height down and end up outside the rail entirely. + */ +export function railGeometry(render, nClaims) { + const rail = render.rail ?? {}; + const RW = rail.width ?? 500; + const HH = render.headerHeight ?? 56; + // reservedFooterHeight, NOT render.footerHeight. The two differ by 100px the + // moment the chart band is on (the band takes the footer's ground and 100 + // more), and deriving the rail's height from the smaller one ran the column + // a hundred pixels past the band's own top edge -- about two rows of the log + // window, drawn below the line everything else letterboxes to. + const FH = reservedFooterHeight(render); + const ROWH = rail.rowHeight ?? 38; + const PAD = rail.pad ?? 22; + const TALLYROWH = rail.tallyRowHeight ?? 40; + const nTracks = (rail.tracks ?? []).length; + + const RHGT = render.height - HH - FH; + // The tally's header belongs to the CHROME, not to the rolling block. Put it + // in the block and every roll scrolls a duplicate copy of it up through the + // window, which reads as noise rather than as a counter changing. + const TALLYHEAD_REL = rail.tallyTop ?? 76; + const TALLYTOP_REL = TALLYHEAD_REL + 26; + const TALLYH = nTracks * TALLYROWH; + + // The rolling CELL: the only part of a tally row that ever changes. The + // swatch and the company name sit left of it and belong to the chrome, so + // that a number moving does not drag its own label up the screen with it. + const CELLW = rail.cellWidth ?? 170; + const CELLX = RW - PAD - CELLW; + + // The roster he last enumerated. One line, and the point of it is that it + // stands still while the numbers above it move. + const ROSTERH = rail.rosterHeight ?? 26; + const ROSTERTOP_REL = TALLYTOP_REL + TALLYH + 14; + const ROSTERX = PAD + (rail.rosterLabelWidth ?? 62); + const ROSTERW = RW - PAD - ROSTERX; + + // The provenance tile, parked at the foot of the column: one QR per clip, + // bordered so it reads as a link rather than as decoration. + const QRSIZE = rail.qrSize ?? 132; + const TILEW = RW - 2 * PAD; + const TILEH = QRSIZE + 22; + const TILETOP_REL = RHGT - (rail.qrBottom ?? 12) - TILEH; + + const LOGTOP_REL = ROSTERTOP_REL + ROSTERH + 34; + // The log window is a WHOLE number of rows and stops short of the tile, so + // the curtain parks exactly one window-height down and the QR overlay (which + // comes after it in the chain) is never painted over. + const K = Math.max(1, Math.floor((TILETOP_REL - 12 - LOGTOP_REL) / ROWH)); + const LOGH = K * ROWH; + + return { + RW, PAD, ROWH, TALLYROWH, K, LOGH, TALLYH, RHGT, + CELLW, CELLX, + ROSTERH, ROSTERTOP_REL, ROSTERTOP: HH + ROSTERTOP_REL, ROSTERX, ROSTERW, + QRSIZE, TILEW, TILEH, TILETOP_REL, TILETOP: HH + TILETOP_REL, + VW: render.width - RW, + RX: render.width - RW, + RTOP: HH, + LOGTOP_REL, LOGTOP: HH + LOGTOP_REL, + TALLYTOP_REL, TALLYTOP: HH + TALLYTOP_REL, TALLYHEAD_REL, + nClaims, + logStripH: Math.max(LOGH, nClaims * ROWH), + }; +} + +/** + * How much of the frame the chrome band owns at the bottom. + * + * ONE definition, because three different renderers need it and they were + * already disagreeing: the closing chart drew its footnotes into the bottom + * 100px and the ledger scroll sized its window to `height - header - 100`, both + * of which are wrong the moment the band takes 200. The symptom is a card that + * looks finished in isolation and has its last two lines sitting under the + * chart in the cut. + */ +export function reservedFooterHeight(render) { + return render.chromeEngine === "hyperframes" + ? (render.chart?.height ?? 200) + : (render.footerHeight ?? 100); +} + +/** + * Every roll each tally cell will perform, as rows of one shared strip. + * + * --------------------------------------------------------------------------- + * Why the strip's LAYOUT carries the direction + * --------------------------------------------------------------------------- + * The whole block used to slide as one slab: when the coffee company's number + * changed, all four rows moved. Text that has not changed must not move, so + * each track now gets its own cell and its own y expression. + * + * A crop window can only walk a strip, and it walks in whichever direction its + * y expression takes it. So the DIRECTION of a roll is decided when the rows + * are laid out, not when the ramp is written: + * + * rise rows [old, new] crop walks DOWN the strip, content moves UP + * fall rows [new, old] crop walks UP the strip, content moves DOWN + * + * Between two transitions the crop steps instantly to the next pair's starting + * row. That step is invisible because the row it leaves and the row it arrives + * at hold IDENTICAL content — which is the reason every pair repeats the value + * it starts from rather than sharing a row with its neighbour. + * + * The delta chip rides along on both rows of the pair, and therefore stays on + * screen until the next change. That is deliberate: it reads as "how this + * number last moved", and blanking it at the step would make the invisible + * reposition visible. + * + * A repeated identical figure still rolls, upward. He said it again on a new + * date, and the `as of` line underneath is what changed. + * + * @returns {{lanes: Array<object>, rows: number}} one lane per track plus the + * roster lane, each with `rows` (the cells to draw) and `steps` (per claim, + * `null` or `{a, b}` — the row the roll starts on and the row it ends on). + */ +export function tallyTracks(ledger, tracks, opts = {}) { + const rosterLineOf = opts.rosterLine ?? (() => null); + const lanes = tracks.map((tr) => ({ + key: tr.key, track: tr, kind: "tally", + rows: [{ empty: true, track: tr }], + steps: new Array(ledger.length).fill(null), + cur: null, + })); + const byKey = Object.fromEntries(lanes.map((l) => [l.key, l])); + + const roster = { + key: "__roster", kind: "roster", + rows: [{ empty: true }], + steps: new Array(ledger.length).fill(null), + cur: null, + }; + + /** Lay a transition down as a pair of rows and record where it starts/ends. */ + const transition = (lane, from, to, rise) => { + const p = lane.rows.length; + if (rise) { + lane.rows.push(from, to); + return { a: p, b: p + 1 }; + } + lane.rows.push(to, from); + return { a: p + 1, b: p }; + }; + + ledger.forEach((c, i) => { + // ---- the four company cells ---- + const lane = byKey[c.scope ?? c.company]; + // UTTERED only. The tally says "the latest figure he has given", and a sum + // we performed is not one — showing 18 here while a card beside it says he + // never said eighteen makes the video contradict itself on screen. + if (lane && c.value != null && (!c.valueKind || c.valueKind === "uttered")) { + const prev = lane.cur; + const next = { + track: lane.track, + display: c.display ?? String(c.value), + value: c.value, + date: c.date, + population: c.population ?? null, + delta: prev ? Number((c.value - prev.value).toFixed(2)) : null, + }; + const from = prev ? { ...prev, delta: prev.delta } : { empty: true, track: lane.track }; + lane.steps[i] = transition(lane, from, next, !prev || c.value >= prev.value); + lane.cur = next; + } + + // ---- the roster ---- + // Only when the LINE changes. He enumerates the same two editors and one + // designer in October and again in December; rolling the line to arrive at + // the words it already said would animate the one thing that held still. + const line = Array.isArray(c.roles) && c.roles.length ? rosterLineOf(c.roles) : null; + if (line && line !== roster.cur?.line) { + const next = { line, date: c.date }; + const from = roster.cur ? { ...roster.cur } : { empty: true }; + roster.steps[i] = transition(roster, from, next, true); + roster.cur = next; + } + }); + + const all = [...lanes, roster]; + return { lanes: all, rows: Math.max(...all.map((l) => l.rows.length)) }; +} + +/** + * The colour of the population word under a figure. Never a new hue -- the + * chip is `pal.muted` text, so the dataviz gate does not have to be re-run. + */ +const POP_WORD = { + employees: "employees", + "full-time": "full time", + salaried: "salaried", + contractor: "contractors", + 1099: "1099", + people: "people", +}; + +/** A small solid triangle, because a font may not carry ▲ and tofu is worse. */ +function svgTri(x, y, up, color) { + const d = up + ? `M${x},${y + 7} L${x + 4.5},${y} L${x + 9},${y + 7} z` + : `M${x},${y} L${x + 4.5},${y + 7} L${x + 9},${y} z`; + return `<path d="${d}" fill="${color}"/>`; +} + +/** One rolling tally cell, drawn into a CELLW x TALLYROWH box at (x, y). */ +function tallyCellSvg(cell, x, y, g, pal, h) { + const right = x + g.CELLW; + const out = [`<rect x="${x}" y="${y}" width="${g.CELLW}" height="${h}" fill="${pal.bg}"/>`]; + if (cell.empty) { + out.push( + svgText(right, y + 24, "—", { size: 22, color: pal.muted, weight: "bold", anchor: "end", opacity: 0.5 }), + svgText(right, y + 37, "not yet stated", { size: 10.5, color: pal.muted, opacity: 0.6, anchor: "end" }), + ); + return out.join(""); + } + const colour = cell.track?.color ?? pal.fg; + out.push( + svgText(right, y + 24, cell.display, { size: 22, color: colour, weight: "bold", anchor: "end" }), + ); + if (cell.delta != null && cell.delta !== 0) { + const up = cell.delta > 0; + out.push( + svgTri(x, y + 11, up, colour), + svgText(x + 14, y + 22, `${up ? "+" : "−"}${Math.abs(cell.delta)}`, { + size: 13, color: colour, weight: "bold", + }), + ); + } + const chip = cell.population ? ` · ${POP_WORD[cell.population] ?? cell.population}` : ""; + out.push( + svgText(right, y + 37, `as of ${cell.date}${chip}`, { + size: 10.5, color: pal.muted, anchor: "end", opacity: 0.9, + }), + ); + return out.join(""); +} + +/** One roster line, drawn into a ROSTERW x ROSTERH box. */ +function rosterCellSvg(cell, x, y, g, pal) { + const out = [`<rect x="${x}" y="${y}" width="${g.ROSTERW}" height="${g.ROSTERH}" fill="${pal.bg}"/>`]; + out.push( + cell.empty + ? svgText(x, y + 18, "not yet enumerated", { size: 12.5, color: pal.muted, opacity: 0.55 }) + : svgText(x, y + 18, fit(cell.line, 13, g.ROSTERW), { size: 13, color: pal.muted }), + ); + return out.join(""); +} + +// --------------------------------------------------------------------------- +// The provenance tile +// --------------------------------------------------------------------------- +// A compilation asks the viewer to take the edit on trust. The QR is the +// antidote: it resolves to this clip's exact start in the archive's own viewer, +// so anyone can pull up the surrounding hour and check the cut is fair. +// +// It used to float over the bottom-right of the PICTURE, which is the one place +// in the frame the cut promises never to draw on. In the rail's foot it is a +// bordered tile that reads as a link, and it becomes one more strip walked by a +// crop -- one tile per segment, stepped instantaneously at the mid-dissolve, +// exactly like the other four. +// +// Two rules learned the hard way: it must be FULLY OPAQUE (a translucent QR +// will not scan) and it must keep its quiet zone (the white border is part of +// the symbol, not decoration). +export function qrUrlFor(entry, provenance) { + if (entry.type === "clip") { + // A mirror's LOCAL slug is not the id the site serves, so an explicit + // per-clip citeUrl always wins over the derived one. + return ( + entry.citeUrl ?? + `${provenance.siteOrigin}/?v=${encodeURIComponent( + `${entry.channel ?? provenance.channelSlug}/${entry.video}`, + )}&t=${Math.floor(entry.cite ?? entry.start)}` + ); + } + // A card is not a moment, so it gets the search rather than a timestamp. + // + // NOT `provenance.shareLink`. That link carries all 23 channel filters and is + // ~1.4k characters — a version-40 symbol, 177 modules inside a 132 px tile, + // which is roughly 0.7 px per module and unscannable. `qrLink` is the same + // query without the channel list (~200 chars, 63 modules, verified scannable + // at this size); the site origin is the fallback. A code nobody can scan is + // worse than a short one. + return provenance.qrLink ?? provenance.siteOrigin ?? ""; +} + +async function qrTileStrip(entries, provenance, render, g, outDir) { + const pal = render.palette; + const dir = path.join(outDir, "cards"); + const qrDir = path.join(outDir, "qr"); + const q = render.qr ?? {}; + const QR = g.QRSIZE; + const TH = g.TILEH; + + // One PNG per DISTINCT url, then a tile per segment referencing it. + const seen = new Map(); + const urls = entries.map((e) => qrUrlFor(e, provenance)); + for (const url of urls) { + if (seen.has(url)) continue; + const png = path.join(qrDir, `q${seen.size.toString().padStart(2, "0")}.png`); + await execFileP(QRENCODE, [ + "-o", png, "-s", String(q.scale ?? 4), "-m", String(q.quiet ?? 3), + "-l", q.ecc ?? "M", url, + ]); + // Nearest-neighbour to an exact box: a resampled QR blurs its module edges + // and stops scanning, and the geometry has to be known before this runs. + const sized = path.join(qrDir, `q${seen.size.toString().padStart(2, "0")}.${QR}.png`); + await execFileP("magick", [png, "-filter", "point", "-resize", `${QR}x${QR}!`, sized]); + seen.set(url, sized); + } + + const stripH = entries.length * TH; + const body = [`<rect x="0" y="0" width="${g.TILEW}" height="${stripH}" fill="${pal.bg}"/>`]; + const images = []; + entries.forEach((e, i) => { + const y = i * TH; + const isClip = e.type === "clip"; + body.push( + `<rect x="0.5" y="${y + 0.5}" width="${g.TILEW - 1}" height="${TH - 1}" fill="${pal.bg}" ` + + `stroke="${pal.accent}" stroke-width="1"/>`, + svgText(14, y + 32, "SCAN → JERALYZER", { + size: 12.5, color: pal.accent, weight: "bold", ls: 1.3, + }), + svgText(14, y + 58, isClip ? "this exact moment," : "the sweep this is cut from,", { + size: 13.5, color: pal.fg, + }), + svgText(14, y + 78, isClip ? "in the archive's own viewer" : "every claim, searchable", { + size: 13.5, color: pal.fg, + }), + svgText(14, y + 106, "the archive outlives the platform", { + size: 11.5, color: pal.muted, opacity: 0.8, + }), + ); + images.push({ png: seen.get(urls[i]), x: g.TILEW - QR - 11, y: y + 11 }); + }); + + const svgPath = path.join(dir, "_rail_qr.svg"); + const basePng = path.join(dir, "_rail_qr.base.png"); + await rasterize(svgDoc(g.TILEW, stripH, body.join("")), svgPath, basePng, g.TILEW, stripH); + + // The codes are composited rather than inlined: an <image href> in the SVG + // would be resampled by rsvg, and a resampled QR does not scan. + const out = path.join(dir, "_rail_qr.png"); + const args = [basePng]; + for (const im of images) args.push(im.png, "-geometry", `+${im.x}+${im.y}`, "-composite"); + args.push(out); + await execFileP("magick", args, { maxBuffer: 1 << 26 }); + return { path: out, tileH: TH, urls }; +} + +/** + * Build the rail strips. Returns their paths plus the geometry, so the caller + * never re-derives a size the pixels already committed to. + * + * `entries` and `provenance` are needed for the QR strip only; without them the + * tile is skipped and the rail is the four strips it always was. + */ +export async function renderRailAssets(render, ledger, outDir, entries = null, provenance = null) { + const pal = render.palette; + const rail = render.rail; + const tracks = rail.tracks ?? []; + const byKey = Object.fromEntries(tracks.map((t) => [t.key, t])); + const g = railGeometry(render, ledger.length); + const dir = path.join(outDir, "cards"); + const rule = rail.rule ?? "#2A322F"; + const P = (n) => path.join(dir, n); + + // ---- chrome: opaque, full rail column, never gated ------------------- + // It runs for the whole video rather than being switched on with enable=, + // which removes four expressions and four off-by-one opportunities. + // + // The tally SWATCH and LABEL live here. They never change, and while they + // rode in the rolling block a coffee figure changing dragged "The Quartering" + // up the screen with it — text moving for a reason that was not about it. + const chromeBody = [ + `<rect x="0" y="0" width="${g.RW}" height="${g.RHGT}" fill="${pal.bg}"/>`, + `<rect x="0" y="0" width="1" height="${g.RHGT}" fill="${rule}"/>`, + svgText(g.PAD, 32, "THE CLAIM LEDGER", { size: 15, color: pal.amber, weight: "bold", ls: 1.6 }), + svgText(g.PAD, 55, "every count he has given, as he gave it", { size: 14, color: pal.muted }), + `<rect x="${g.PAD}" y="70" width="${g.RW - 2 * g.PAD}" height="1" fill="${rule}"/>`, + svgText(g.PAD, g.TALLYHEAD_REL + 16, "LATEST FIGURE HE HAS GIVEN", { + size: 12, color: pal.muted, weight: "bold", ls: 1.4, + }), + ...tracks.map((tr, j) => { + const ry = g.TALLYTOP_REL + j * g.TALLYROWH; + return ( + `<rect x="${g.PAD}" y="${ry + 13}" width="10" height="10" fill="${tr.color}"/>` + + svgText(g.PAD + 20, ry + 22, fit(tr.label, 14, g.CELLX - g.PAD - 24), { + size: 14, color: pal.fg, + }) + ); + }), + `<rect x="${g.PAD}" y="${g.ROSTERTOP_REL - 8}" width="${g.RW - 2 * g.PAD}" height="1" fill="${rule}"/>`, + svgText(g.PAD, g.ROSTERTOP_REL + 19, "ROSTER", { + size: 11, color: pal.muted, weight: "bold", ls: 1.4, + }), + `<rect x="${g.PAD}" y="${g.LOGTOP_REL - 26}" width="${g.RW - 2 * g.PAD}" height="1" fill="${rule}"/>`, + svgText(g.PAD, g.LOGTOP_REL - 8, "IN THE ORDER STATED", { + size: 12, color: pal.muted, weight: "bold", ls: 1.4, + }), + ].join(""); + const chrome = await rasterize( + svgDoc(g.RW, g.RHGT, chromeBody), P("_rail_chrome.svg"), P("_rail_chrome.png"), g.RW, g.RHGT, + ); + + // ---- log strip: every claim, stacked, no padding --------------------- + const valX = g.RW - g.PAD; + const textX = g.PAD + 20; + const textBudget = valX - textX - 74; + const rows = ledger.map((c, i) => { + const y = i * g.ROWH; + // `scope` is the ADJUDICATED answer and `company` the undocumented guess it + // replaced. Falls back so a manifest with no adjudication yet still renders. + const tr = byKey[c.scope ?? c.company]; + // In `sourced` every row has a clip behind it, so `live` is always true + // there; `full` keeps the distinction because a stacked ledger card is a + // weaker citation than footage and must not look like one. + const live = !!c.entryId && !c.unsourced; + const ink = live ? pal.fg : pal.muted; + const dotFill = live ? (tr?.color ?? pal.accent) : "none"; + return [ + `<rect x="0" y="${y}" width="${g.RW}" height="${g.ROWH}" fill="${pal.bg}"/>`, + `<circle cx="${g.PAD + 5}" cy="${y + 19}" r="4.5" fill="${dotFill}" ` + + `stroke="${tr?.color ?? pal.muted}" stroke-width="1.5" opacity="${live ? 1 : 0.55}"/>`, + svgText(textX, y + 17, c.date, { size: 13.5, color: pal.muted, opacity: live ? 1 : 0.7 }), + svgText(valX, y + 19, c.display ?? "—", { + size: 18, color: live ? (tr?.color ?? pal.fg) : pal.muted, + weight: "bold", anchor: "end", opacity: live ? 1 : 0.65, + }), + svgText(textX, y + 33, fit(c.label ?? c.quote ?? "", 13, textBudget + 74), { + size: 13, color: ink, opacity: live ? 1 : 0.6, + }), + `<rect x="${g.PAD}" y="${y + g.ROWH - 1}" width="${g.RW - 2 * g.PAD}" height="1" fill="${rule}"/>`, + ].join(""); + }).join(""); + const log = await rasterize( + svgDoc(g.RW, g.logStripH, rows), P("_rail_log.svg"), P("_rail_log.png"), g.RW, g.logStripH, + ); + + // ---- curtain --------------------------------------------------------- + // With ONE log strip and the window parked at the top while the list is still + // filling, rows i+1…K-1 would show claims the video has not made yet. This + // opaque rectangle rides just below the last revealed row and, once the list + // is full, parks exactly one window-height down — outside the window. It is + // pal.bg precisely so that parking there is invisible; the QR tile overlays + // AFTER it, which is what stops the parked curtain covering the code. + const curtain = P("_rail_curtain.png"); + await execFileP("magick", ["-size", `${g.RW}x${g.LOGH}`, `xc:${pal.bg}`, curtain]); + + // ---- highlight ------------------------------------------------------- + const hl = await rasterize( + svgDoc(g.RW, g.ROWH, [ + `<rect x="0" y="0" width="${g.RW}" height="${g.ROWH}" fill="${pal.amber}" opacity="0.10"/>`, + `<rect x="${g.PAD - 12}" y="4" width="3" height="${g.ROWH - 8}" fill="${pal.amber}"/>`, + ].join("")), + P("_rail_hl.svg"), P("_rail_hl.png"), g.RW, g.ROWH, + ); + + // ---- tally strip: one COLUMN per lane, side by side ------------------- + // Four cells and the roster in one PNG: five crops at different x out of one + // input, rather than five inputs. Every column is as tall as the tallest, so + // a single strip height serves them all. + const { lanes, rows: nRows } = tallyTracks(ledger, tracks, { rosterLine }); + const cellH = g.TALLYROWH; + const laneX = []; + let sx = 0; + for (const lane of lanes) { + const w = lane.kind === "roster" ? g.ROSTERW : g.CELLW; + laneX.push({ x: sx, w }); + sx += w; + } + const stripW = sx; + const stripH = nRows * cellH; + const strip = [`<rect x="0" y="0" width="${stripW}" height="${stripH}" fill="${pal.bg}"/>`]; + lanes.forEach((lane, li) => { + const { x } = laneX[li]; + lane.rows.forEach((cell, ri) => { + strip.push( + lane.kind === "roster" + ? rosterCellSvg(cell, x, ri * cellH, g, pal) + : tallyCellSvg(cell, x, ri * cellH, g, pal, cellH), + ); + }); + }); + const tally = await rasterize( + svgDoc(stripW, stripH, strip.join("")), P("_rail_tally.svg"), P("_rail_tally.png"), stripW, stripH, + ); + + // ---- the provenance tile --------------------------------------------- + const qr = + entries && provenance && render.qr !== false + ? await qrTileStrip(entries, provenance, render, g, outDir) + : null; + + return { + chrome, log, curtain, hl, tally, qr, + geom: g, + lanes: lanes.map((lane, li) => ({ + key: lane.key, kind: lane.kind, steps: lane.steps, + x: laneX[li].x, w: laneX[li].w, cellH, + })), + }; +} + +// =========================================================================== +// Stacked ledger cards +// =========================================================================== +// The `full` cut's answer to the 28 claims the sweep found and no clip covers. +// +// Dimmed rail rows were the old answer, and they were confusing: a row with no +// audio behind it slid past with nothing to say for itself, and 28 of them read +// as padding rather than as evidence. So each one gets SCREEN TIME instead -- +// its date, its scope, its quote, his figure, and what that figure does to our +// running sum. Consecutive unclipped claims share a card and reveal in +// sequence, which is why 28 claims cost 14 cards and about 67 seconds. +// +// The reveal is the rail curtain's device: an opaque `pal.bg` rectangle walking +// down the card. Nothing fades, nothing moves; rows simply stop being covered. + +/** The reveal clock. One definition, because `scheduleClaims` pins to it. */ +export const LEDGER_LEAD = 0.9; +export const LEDGER_STEP = 1.3; +export const LEDGER_TAIL = 2.6; +export const ledgerRevealAt = (r) => LEDGER_LEAD + LEDGER_STEP * r; +export const ledgerSeconds = (n) => LEDGER_LEAD + LEDGER_STEP * n + LEDGER_TAIL - LEDGER_STEP; + +/** + * Why this claim is a line of text and not footage. + * + * "The upload is gone" and "we did not cut it" are different sentences, and + * saying the first about a live source is the kind of error that makes a whole + * compilation untrustworthy. So the words come from a `yt-dlp --simulate` + * probe, recorded in out/availability.json with the date it ran. + */ +export const SOURCE_TAG = { + ok: "not clipped", + deleted: "source deleted", + private: "source private", + "members-only": "members only", + restricted: "age-restricted", + "geo-blocked": "geo-blocked", + "no-cues": "no archived transcript", + maybe_missing: "source unreachable", +}; + +/** Greedy wrap to a pixel budget, at most `maxLines`, last line elided. */ +function wrapPx(text, size, maxPx, maxLines) { + const perChar = size * 0.5; + const cols = Math.max(8, Math.floor(maxPx / perChar)); + const words = String(text ?? "").split(/\s+/).filter(Boolean); + const lines = []; + let line = ""; + for (const w of words) { + if (line && (line + " " + w).length > cols) { + lines.push(line); + line = w; + if (lines.length === maxLines) break; + } else { + line = line ? line + " " + w : w; + } + } + if (lines.length < maxLines && line) lines.push(line); + if (lines.length === maxLines) { + const used = lines.join(" ").split(/\s+/).length; + if (used < words.length) lines[maxLines - 1] = fit(lines[maxLines - 1] + " …", size, maxPx); + } + return lines; +} + +/** + * One card carrying a run of consecutive unclipped claims. + * + * Returns the geometry the encoder needs to walk the curtain: where the rows + * start and how tall each one is. + */ +export async function renderLedgerCard(card, render, ledger, outDir, avail = null) { + const pal = render.palette; + const tracks = render.rail?.tracks ?? []; + const byKey = Object.fromEntries(tracks.map((t) => [t.key, t])); + const VW = cardWidth(card, render); + const H = render.height; + const dir = path.join(outDir, "cards"); + const rule = render.rail?.rule ?? "#2A322F"; + const RESERVED = reservedFooterHeight(render); + + const ids = card.claims ?? []; + const rows = ids.map((id) => ledger.find((c) => c.id === id)).filter(Boolean); + if (rows.length !== ids.length) { + const missing = ids.filter((id) => !ledger.some((c) => c.id === id)); + throw new Error(`ledger card ${card.id} names claims that are not in the ledger: ${missing.join(", ")}`); + } + + // The arithmetic is READ, never recomputed: one implementation of the walk, + // or the card and the chart band can disagree about the same sum. + const steps = new Map(ledgerTotals(ledger).steps.map((st) => [st.id, st])); + + const M = 96; + const ARITHW = 320; + const arithX = VW - M - ARITHW; + const quoteW = arithX - 70 - M; + + const body = [`<rect x="0" y="0" width="${VW}" height="${H}" fill="${pal.bg}"/>`]; + body.push( + svgText(M, 80, (card.kicker ?? "found in the sweep, not clipped here").toUpperCase(), { + size: 22, color: pal.amber, weight: "bold", ls: 1.4, + }), + svgText(M, 132, card.heading ?? "What the sweep found and this cut cannot show you", { + size: 40, color: pal.fg, weight: "bold", + }), + svgText(M, 168, card.sub ?? "his own words, and what they do to our running sum", { + size: 20, color: pal.muted, + }), + `<rect x="${M}" y="${192}" width="${VW - 2 * M}" height="1" fill="${rule}"/>`, + ); + + // Rows are a fixed height and the BLOCK is centred in what is left of the + // frame. Stretching two rows to fill 640px puts a hand's width of nothing + // between them; packing them at the top leaves the same gap in one lump at + // the bottom. Centring is the only arrangement that reads as deliberate. + const top = 214; + const available = H - RESERVED - top - 24; + const ROWH = Math.min(180, Math.floor(available / rows.length)); + const rowsTop = top + Math.floor((available - ROWH * rows.length) / 2); + + rows.forEach((c, r) => { + const y = rowsTop + r * ROWH; + const tr = byKey[c.scope ?? c.company]; + const st = steps.get(c.id); + const state = avail?.get(c.id) ?? null; + const tag = SOURCE_TAG[state] ?? "not clipped"; + + body.push( + svgText(M, y + 32, c.date, { size: 20, color: pal.muted }), + // The scope, as a bordered pill in its own colour. Which payroll a number + // is about is the whole argument, so it is never left to the ink alone. + `<rect x="${M + 148}" y="${y + 12}" width="${Math.max(120, (tr?.label?.length ?? 8) * 7.6 + 22)}" ` + + `height="26" rx="13" fill="none" stroke="${tr?.color ?? pal.muted}" stroke-width="1.2"/>`, + svgText(M + 159, y + 30, tr?.label ?? c.scope ?? "", { size: 14.5, color: tr?.color ?? pal.muted }), + svgText(M + 148 + Math.max(120, (tr?.label?.length ?? 8) * 7.6 + 22) + 16, y + 30, tag, { + size: 14.5, color: pal.muted, opacity: 0.85, + }), + svgText(arithX - 70, y + 40, c.display ?? "—", { + size: 34, color: tr?.color ?? pal.fg, weight: "bold", anchor: "end", + }), + ); + wrapPx(`“${c.quote ?? c.label ?? ""}”`, 25, quoteW, 2).forEach((line, li) => { + body.push(svgText(M, y + 76 + li * 33, line, { size: 25, color: pal.fg })); + }); + + // ---- the arithmetic column ---- + // Which layer this claim moved, lit; the others held, dimmed. The point is + // that the total on the right is OURS and is made of his own figures. + body.push( + svgText(arithX, y + 24, "OUR RUNNING SUM", { + size: 11, color: pal.muted, weight: "bold", ls: 1.3, + }), + ); + const basis = st?.impliedBasis ?? {}; + const companies = tracks.filter((t) => t.key !== "all"); + // Before any company has given a figure, our sum is not zero — it is + // undefined, and four dashes in a column say that far less clearly than + // one sentence does. + if (!companies.some((t) => basis[t.key])) { + body.push( + svgText(arithX, y + 52, "no company figure yet,", { size: 14, color: pal.muted }), + svgText(arithX, y + 74, "so our sum is not defined", { size: 14, color: pal.muted }), + ); + if (r < rows.length - 1) { + body.push( + `<rect x="${M}" y="${y + ROWH - 1}" width="${VW - 2 * M}" height="1" fill="${rule}" opacity="0.6"/>`, + ); + } + return; + } + companies.forEach((t, k) => { + const b = basis[t.key]; + const moved = (c.scope ?? c.company) === t.key; + const ty = y + 48 + k * 24; + body.push( + svgText(arithX, ty, fit(t.shortLabel ?? t.label, 13, ARITHW - 90), { + size: 13, color: moved ? t.color : pal.muted, opacity: moved ? 1 : 0.55, + }), + svgText(arithX + ARITHW, ty, b ? String(b.value) : "—", { + size: 17, color: moved ? t.color : pal.muted, weight: "bold", anchor: "end", + opacity: moved ? 1 : 0.55, + }), + ); + }); + const sy = y + 48 + companies.length * 24; + body.push( + `<rect x="${arithX}" y="${sy + 6}" width="${ARITHW}" height="1" fill="${rule}"/>`, + svgText(arithX, sy + 30, "IMPLIED", { size: 13, color: pal.fg, weight: "bold", ls: 1.2 }), + svgText(arithX + ARITHW, sy + 32, st?.implied == null ? "—" : String(st.implied), { + size: 22, color: pal.fg, weight: "bold", anchor: "end", + }), + ); + if (st?.impliedDelta) { + const up = st.impliedDelta > 0; + body.push( + svgTri(arithX + 84, sy + 22, up, pal.amber), + svgText(arithX + 98, sy + 30, `${up ? "+" : "−"}${Math.abs(st.impliedDelta)}`, { + size: 14, color: pal.amber, weight: "bold", + }), + ); + } + + if (r < rows.length - 1) { + body.push( + `<rect x="${M}" y="${y + ROWH - 1}" width="${VW - 2 * M}" height="1" fill="${rule}" opacity="0.6"/>`, + ); + } + }); + + const outPath = path.join(dir, `${card.id}.png`); + await rasterize(svgDoc(VW, H, body.join("")), path.join(dir, `${card.id}.svg`), outPath, VW, H); + return { path: outPath, width: VW, rowsTop, rowHeight: ROWH, rows: rows.length }; +} + +// =========================================================================== +// End sequence: the ledger scroll and the step chart +// =========================================================================== + +/** + * The whole ledger as one tall PNG for an animated crop to walk. + * + * --------------------------------------------------------------------------- + * One chronological line, a column per company + * --------------------------------------------------------------------------- + * It used to group by company: four blocks, each date-sorted inside itself. The + * cut plays in ONE chronology, and grouping at the end re-tells it in an order + * the viewer has not just watched — and it hides the only thing worth seeing + * here, which is that the four payrolls were being described in the same weeks. + * + * So: one date-ordered list, and the company is read from COLUMN POSITION. That + * makes colour the secondary encoding rather than the only one, which is the + * same rule the chart already runs under. + * + * The rail hides for this card (`hideRail`), so it is drawn at the FULL frame + * width rather than the content width. + * + * Returns the CONTENT HEIGHT because the scroll expression is written against + * it — crop clamps its own y, so an off-by-a-few degrades into a static last + * frame rather than an error, but only if the caller knows the real number. + */ +export async function renderScrollCard(card, render, ledger, outDir) { + const pal = render.palette; + const tracks = render.rail?.tracks ?? []; + const VW = cardWidth(card, render); + const dir = path.join(outDir, "cards"); + const rule = render.rail?.rule ?? "#2A322F"; + + const M = 96; + const ROWH = 42; + const body = []; + + // Columns. The value columns are right-aligned on their own gridline, so a + // number's horizontal position IS its company even before the colour reads. + const COLW = 152; + const dateX = M; + const colX = tracks.map((_, i) => M + 168 + i * COLW); + const popX = M + 168 + tracks.length * COLW + 24; + const labelX = popX + 132; + const labelW = VW - M - labelX; + + let y = 66; + body.push(svgText(M, y, card.heading ?? "THE COMPLETE LEDGER", { + size: 30, color: pal.fg, weight: "bold", ls: 1.5, + })); + y += 32; + body.push(svgText(M, y, card.sub ?? `${ledger.length} dated claims, in the order he made them`, { + size: 19, color: pal.muted, + })); + y += 44; + + // The column heads, which are the legend. No separate key: a company name + // over its own column of figures is the shortest legend there is. + body.push(`<rect x="${M}" y="${y - 4}" width="${VW - 2 * M}" height="1" fill="${rule}"/>`); + body.push(svgText(dateX, y + 26, "DATE", { size: 13, color: pal.muted, weight: "bold", ls: 1.3 })); + // No swatch beside the head: the head is already IN the track's colour, and + // the column position is the primary encoding either way. A swatch would only + // land on top of the words, since a right-anchored run cannot be measured + // here to leave room for one. + tracks.forEach((tr, i) => { + body.push( + svgText(colX[i], y + 26, fit(tr.shortLabel ?? tr.label, 13, COLW - 12), { + size: 13, color: tr.color, weight: "bold", anchor: "end", ls: 0.6, + }), + ); + }); + body.push( + svgText(popX, y + 26, "AS WHAT", { size: 13, color: pal.muted, weight: "bold", ls: 1.3 }), + svgText(labelX, y + 26, "WHAT HE SAID", { size: 13, color: pal.muted, weight: "bold", ls: 1.3 }), + ); + y += 40; + body.push(`<rect x="${M}" y="${y}" width="${VW - 2 * M}" height="1" fill="${rule}"/>`); + y += 8; + + const byKey = Object.fromEntries(tracks.map((t, i) => [t.key, i])); + const rows = [...ledger].sort((a, b) => dateKey(a.date).localeCompare(dateKey(b.date))); + for (const c of rows) { + const live = !!c.entryId && !c.unsourced; + const i = byKey[c.scope ?? c.company]; + const tr = tracks[i]; + body.push( + svgText(dateX, y + 26, c.date, { size: 18, color: pal.muted, opacity: live ? 1 : 0.7 }), + ); + if (tr) { + body.push( + svgText(colX[i], y + 26, c.display ?? "—", { + size: 21, color: live ? tr.color : pal.muted, weight: "bold", anchor: "end", + opacity: live ? 1 : 0.6, + }), + ); + } + body.push( + svgText(popX, y + 26, POP_WORD[c.population] ?? c.population ?? "", { + size: 15, color: pal.muted, opacity: live ? 0.9 : 0.6, + }), + svgText(labelX, y + 26, fit(c.label ?? "", 18, labelW), { + size: 18, color: live ? pal.fg : pal.muted, opacity: live ? 1 : 0.6, + }), + `<rect x="${M}" y="${y + ROWH - 1}" width="${VW - 2 * M}" height="1" fill="${rule}" opacity="0.5"/>`, + ); + y += ROWH; + } + y += 60; + + const contentHeight = y; + const outPath = path.join(dir, `${card.id}.png`); + await rasterize( + svgDoc(VW, contentHeight, `<rect x="0" y="0" width="${VW}" height="${contentHeight}" fill="${pal.bg}"/>${body.join("")}`), + path.join(dir, `${card.id}.svg`), outPath, VW, contentHeight, + ); + return { path: outPath, contentHeight, width: VW }; +} + +/** + * The four-series step chart, over the claims flagged `plotted`. + * + * COLOUR IS NOT THE ONLY ENCODING here, and that is a hard requirement rather + * than a flourish: no four-colour categorical palette clears the data-viz + * all-pairs CVD gate (three slots is the documented ceiling), so each series + * also carries a distinct dash pattern and a direct end-of-line label. The four + * hues themselves are the published artifact's, re-validated against this + * video's darker ground (#0F1312) on the adjacent pairlist — the pairlist for + * line charts — where all five checks pass. + */ +export async function renderChartCard(card, render, ledger, outDir) { + const pal = render.palette; + const tracks = render.rail?.tracks ?? []; + const VW = cardWidth(card, render); + const H = render.height; + const dir = path.join(outDir, "cards"); + const rule = render.rail?.rule ?? "#2A322F"; + + // The series come from ledger-totals, not from the legacy `plotted` flag. + // `plotted` was set under the OLD reading, in which a sum we performed sat in + // the same series as a figure he uttered. Drawing from it now would put 18 and + // 20 back on his line, after the whole point of the adjudication was to take + // them off it. + let totals = null; + try { + totals = ledgerTotals(ledger); + } catch { + // An unadjudicated ledger still renders -- as the three company series only, + // because the two totals are exactly what it cannot be trusted about. + totals = null; + } + const pts = ledger.filter((c) => c.value != null && (c.scope ?? c.company) !== "all"); + const yr = (d) => { + const [Y, M2, D2] = d.split("-").map(Number); + return Y + (M2 - 1) / 12 + (D2 - 1) / 365; + }; + const X0 = yr("2020-01-01"), X1 = yr("2026-12-31"); + const YMAX = + Math.max(21, ...pts.map((p) => p.value), ...(totals?.series.implied ?? []).map((p) => p.value)) + 1; + + const RESERVED = reservedFooterHeight(render); + const box = { l: 150, r: 300, t: 190, b: 130 + RESERVED }; + const plotW = VW - box.l - box.r; + const plotH = H - box.t - box.b; + const px = (v) => box.l + ((v - X0) / (X1 - X0)) * plotW; + const py = (v) => H - box.b - (v / YMAX) * plotH; + + const body = [`<rect x="0" y="0" width="${VW}" height="${H}" fill="${pal.bg}"/>`]; + body.push( + svgText(box.l, 78, "WHAT HE SAID, AND WHAT IT ADDS UP TO", { + size: 34, color: pal.fg, weight: "bold", ls: 1.5, + }), + svgText(box.l, 112, "every figure he utters, against the company he was talking about", { + size: 20, color: pal.muted, + }), + svgText(box.l, 146, "the heavy line is ours — his own per-company claims, added up", { + size: 18, color: pal.amber, + }), + ); + + // grid + axes + for (let gv = 0; gv <= YMAX - 1; gv += 5) { + body.push( + `<rect x="${box.l}" y="${py(gv)}" width="${plotW}" height="1" fill="${rule}"/>`, + svgText(box.l - 16, py(gv) + 6, String(gv), { size: 17, color: pal.muted, anchor: "end" }), + ); + } + body.push(svgText(box.l - 16, py(YMAX - 1) - 22, "PEOPLE", { + size: 13, color: pal.muted, weight: "bold", anchor: "end", ls: 1.2, + })); + for (let Y = 2020; Y <= 2026; Y += 1) { + const x = px(yr(`${Y}-01-01`)); + body.push( + `<rect x="${x}" y="${box.t}" width="1" height="${py(0) - box.t}" fill="${rule}" opacity="0.7"/>`, + svgText(x, py(0) + 30, String(Y), { size: 17, color: pal.muted, anchor: "middle" }), + ); + } + body.push(`<rect x="${box.l}" y="${py(0)}" width="${plotW}" height="2" fill="${pal.muted}"/>`); + + // One step path per series, plus a dot per claim and a direct end label. + // + // FIVE series, not four. The three companies are his, drawn as before. The + // fourth is what he says the WHOLE payroll is -- only ever a figure he utters + // as one number. The fifth is what his own per-company claims add up to, and + // it is ours: a heavy neutral step, because an aggregate is not a categorical + // peer of the things it aggregates and must not consume a palette slot. + const DASH = ["", "12 6", "3 7", "18 5 4 5"]; + const labels = []; + const drawn = []; + tracks.forEach((tr, ti) => { + if (tr.key === "all") return; + drawn.push({ + tr, dash: DASH[ti % 4], width: 3.5, + pts: pts.filter((c) => (c.scope ?? c.company) === tr.key) + .slice().sort((a, b) => a.date.localeCompare(b.date)) + .map((c) => ({ date: c.date, value: c.value, display: c.display, hedged: c.hedged })), + }); + }); + if (totals) { + const allTrack = tracks.find((t) => t.key === "all"); + drawn.push({ + tr: { key: "stated", color: allTrack?.color ?? pal.accent, label: "stated total" }, + dash: "18 5 4 5", width: 3.5, dots: true, + pts: totals.series.stated.map((p) => ({ date: p.date, value: p.value, display: String(p.value) })), + }); + drawn.push({ + tr: { key: "implied", color: pal.fg, label: "implied — our sum" }, + dash: "", width: 6, dots: false, + pts: totals.series.implied.map((p) => ({ date: p.date, value: p.value, display: String(p.value) })), + }); + } + + for (const sr of drawn) { + const mine = sr.pts; + if (!mine.length) continue; + // A step, not a line: the figure he gave holds until he gives another one, + // so the segment between two claims must be flat and the change vertical. + let d = ""; + let prevY = null; + for (const [i, c] of mine.entries()) { + const x = px(yr(c.date)); + const yv = py(c.value); + d += i === 0 + ? `M ${x.toFixed(1)} ${yv.toFixed(1)}` + : ` L ${x.toFixed(1)} ${prevY.toFixed(1)} L ${x.toFixed(1)} ${yv.toFixed(1)}`; + prevY = yv; + } + const last = mine[mine.length - 1]; + const lastY = py(last.value); + d += ` L ${(box.l + plotW).toFixed(1)} ${lastY.toFixed(1)}`; + body.push( + `<path d="${d}" fill="none" stroke="${sr.tr.color}" stroke-width="${sr.width}" ` + + `stroke-linejoin="round"${sr.dash ? ` stroke-dasharray="${sr.dash}"` : ""}/>`, + ); + if (sr.dots !== false) { + for (const c of mine) { + body.push( + `<circle cx="${px(yr(c.date)).toFixed(1)}" cy="${py(c.value).toFixed(1)}" r="${c.hedged ? 5 : 6}" ` + + `fill="${c.hedged ? pal.bg : sr.tr.color}" stroke="${sr.tr.color}" stroke-width="2.5"/>`, + ); + } + } + labels.push({ tr: sr.tr, last, lineY: lastY, y: lastY }); + } + + // The closing hold annotates the gap it has just finished drawing. + if (card.hold && totals && totals.final.stated != null && totals.final.implied != null) { + const xR = box.l + plotW; + const yS = py(totals.final.stated); + const yI = py(totals.final.implied); + body.push( + `<rect x="${(xR - 190).toFixed(1)}" y="${Math.min(yI, yS).toFixed(1)}" width="170" ` + + `height="${Math.abs(yS - yI).toFixed(1)}" fill="${tracks.find((t) => t.key === "all")?.color ?? pal.accent}" opacity="0.12"/>`, + `<path d="M ${(xR - 105).toFixed(1)} ${yI.toFixed(1)} L ${(xR - 105).toFixed(1)} ${yS.toFixed(1)}" ` + + `stroke="${pal.amber}" stroke-width="2"/>`, + svgText(xR - 96, (yI + yS) / 2 - 4, `gap ${Math.round(totals.final.implied - totals.final.stated)}`, { + size: 22, color: pal.amber, weight: "bold", + }), + svgText(xR - 96, (yI + yS) / 2 + 22, "between his last total and our sum", { + size: 14, color: pal.muted, + }), + ); + } + + // Three of the four series end within a couple of people of each other, so + // their direct labels land on top of one another. Push them apart and elbow a + // leader line back to the value each one actually belongs to — direct labels + // are the secondary encoding that lets a four-colour palette be legible at + // all, so an unreadable stack would defeat the point of having them. + const LBLH = 48; + labels.sort((a, b) => a.y - b.y); + for (let i = 1; i < labels.length; i += 1) { + labels[i].y = Math.max(labels[i].y, labels[i - 1].y + LBLH); + } + const overshoot = labels.length ? labels[labels.length - 1].y - (H - box.b - 10) : 0; + if (overshoot > 0) for (const l of labels) l.y -= overshoot; + for (const l of labels) { + const lx = box.l + plotW; + if (Math.abs(l.y - l.lineY) > 2) { + body.push( + `<path d="M ${lx} ${l.lineY.toFixed(1)} L ${lx + 9} ${l.lineY.toFixed(1)} ` + + `L ${lx + 9} ${l.y.toFixed(1)} L ${lx + 14} ${l.y.toFixed(1)}" fill="none" ` + + `stroke="${l.tr.color}" stroke-width="1.5" opacity="0.75"/>`, + ); + } + body.push( + svgText(lx + 20, l.y + 2, l.tr.label, { size: 18, color: l.tr.color, weight: "bold" }), + svgText(lx + 20, l.y + 22, `last stated ${l.last.display}`, { size: 14, color: pal.muted }), + ); + } + + body.push( + svgText(box.l, H - RESERVED - 56, "hollow dot = a hedge word (“nearly ten”, “a handful”), not a figure", { + size: 16, color: pal.muted, + }), + svgText(box.l, H - RESERVED - 30, "each series is dashed as well as coloured — the shapes carry the reading on their own; " + + "sums and midpoints are ours and are never drawn as his", { + size: 16, color: pal.muted, + }), + ); + + const outPath = path.join(dir, `${card.id}.png`); + await rasterize(svgDoc(VW, H, body.join("")), path.join(dir, `${card.id}.svg`), outPath, VW, H); + return { + path: outPath, + plotX: box.l, plotY: box.t, plotW, plotH: py(0) - box.t + 2, + }; +} + +export async function renderCard(card, render, outDir, nodes) { + if (card.style === "timeline") { + if (!nodes?.length) throw new Error(`card ${card.id} is style:timeline but no timelineNodes given`); + return renderTimelineCard(card, render, nodes, outDir); + } + return renderPlainCard(card, render, outDir); +} + +async function renderPlainCard(card, render, outDir) { + const pal = render.palette; + const { width, height } = render; + const VW = cardWidth(card, render); + const textWidth = Math.round(VW * 0.74); + const outPath = path.join(outDir, "cards", `${card.id}.png`); + + // Pango reads its markup from a file to keep it clear of shell/argv quoting. + const markupPath = path.join(outDir, "cards", `${card.id}.pango`); + await writeFile(markupPath, markupFor(card, pal), "utf8"); + + // One magick invocation: solid ground, an accent rule down the left margin, + // then the Pango block composited over it. The rule is what keeps the cards + // recognisably one family across styles. + const barX = Math.round(VW * 0.09); + const barTop = Math.round(height * 0.28); + const barBottom = Math.round(height * 0.72); + + const args = [ + "-size", `${width}x${height}`, + `xc:${pal.bg}`, + "-fill", pal.accent, + "-draw", `rectangle ${barX},${barTop} ${barX + 6},${barBottom}`, + "(", + // `-size` is still set to the full frame from the canvas above, and the + // pango delegate honours it — leaving it alone renders the text into a + // 1920x1080 box, which pins the block to the top and wraps at the frame + // edge instead of the margin. Reset it to the text column, height auto. + "-size", `${textWidth}x`, + "-background", "none", + "-define", `pango:width=${textWidth}`, + "-define", "pango:alignment=left", + "-define", "pango:wrap=word", + `pango:@${markupPath}`, + ")", + "-gravity", "West", + "-geometry", `+${barX + 58}+0`, + "-composite", + outPath, + ]; + + await execFileP("magick", args, { maxBuffer: 1 << 24 }); + return outPath; +} + +async function main() { + const argv = process.argv.slice(2); + const manifestPath = argv.find((a) => !a.startsWith("--")); + if (!manifestPath) { + console.error("usage: render-cards.mjs <manifest.json> [--out <dir>] [--only <id>]"); + process.exit(2); + } + const flag = (name) => { + const i = argv.indexOf(name); + return i >= 0 ? argv[i + 1] : undefined; + }; + + const manifest = JSON.parse(await readFile(manifestPath, "utf8")); + const outDir = flag("--out") ?? path.join(path.dirname(path.resolve(manifestPath)), "out"); + const only = flag("--only"); + + await mkdir(path.join(outDir, "cards"), { recursive: true }); + + const cards = manifest.timeline.filter( + (e) => e.type === "card" && (!only || e.id === only), + ); + for (const card of cards) { + const p = await renderCard(card, manifest.render, outDir, manifest.timelineNodes); + console.log(`card ${card.id} -> ${p}`); + } + console.log(`${cards.length} card(s) rendered`); +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main().catch((err) => { + console.error(err); + process.exit(1); + }); +} diff --git a/umtool/report-to-video/resolve-windows.mjs b/umtool/report-to-video/resolve-windows.mjs @@ -0,0 +1,241 @@ +#!/usr/bin/env node +// resolve-windows.mjs — widen a manifest's clip windows to whole sentences. +// +// A manifest window starts life as the cue span covering a quote, and a cue +// boundary is a bad place to cut: ASR breaks cues where the caption line wrapped, +// which is routinely mid-sentence and often mid-word. Cutting there drops the +// lead-in that makes a quote make sense, and clips audibly start and stop in the +// middle of speech. +// +// This walks outward from the cue span to the nearest sentence boundary in the +// transcript — a cue whose text ends in . ? or ! — so the clip carries the whole +// thought. Word-level alignment is a separate, audio-side problem: build-video.mjs +// snaps the actual cut to a silence (see --fetch-pad / snapping there). +// +// Expansion is capped so a run-on passage can't drag a clip out to a minute. +// +// In the app: not used. On the CLI: +// node umtool/report-to-video/resolve-windows.mjs <manifest.json> [--write] +// +// Options: +// --write Rewrite the manifest in place (default: dry run, print a table) +// --max-lead <s> Max seconds to expand backwards (default 9) +// --max-tail <s> Max seconds to expand forwards (default 12) +// --site-origin <url> Archive to read cues from when there is no local corpus +// (defaults to the manifest's provenance.siteOrigin) +// --resolve-site-ids On a published-id miss, find the record by scanning the +// channel's shards. Slow; see cues.mjs. +// --cue-source <which> auto (default) | local | http. The two can disagree +// once a corpus moves past its last publish — see cues.mjs. +// +// A clip entry may set `lockStart` / `lockEnd` to pin that edge exactly. + +import { readFile, writeFile } from "node:fs/promises"; + +import { createCueSource, siteOriginFromManifest } from "./cues.mjs"; + + +const ENDS_SENTENCE = /[.!?]["'”’)\]]*\s*$/; + +// A cue that is only "[music]" or "[ __ ]" (the profanity bleep) carries no +// sentence signal; treat it as transparent so expansion walks past it. +const IS_FILLER = /^\s*(\[[^\]]*\]|>>|♪|—|-)*\s*$/; + +// The manifest stores times rounded to 2 dp, so a value read back from it can sit +// a hair BELOW the cue end it came from. Without a tolerance the end lookup then +// lands on the previous cue, the forward search runs on to the next sentence, and +// the clip grows a little every time this is run — it has to be a fixed point. +const EPS = 0.02; + + +function indexAt(cues, t, which) { + // First cue whose span contains t, else the nearest one on the right side. + let idx = cues.findIndex((c) => c.end > t); + if (idx < 0) idx = cues.length - 1; + if (which === "end") { + let j = cues.findIndex((c) => c.end >= t - EPS); + if (j < 0) j = cues.length - 1; + idx = j; + } + return idx; +} + +export function widen(cues, start, end, { maxLead = 8, maxTail = 12 } = {}) { + const isBoundary = (c) => ENDS_SENTENCE.test(c.text) && !IS_FILLER.test(c.text); + const i0 = indexAt(cues, start, "start"); + const i1 = indexAt(cues, end, "end"); + + // START: the latest cue that OPENS a sentence (i.e. its predecessor closes + // one) at or before the quote, within the lead budget. Finding no such cue + // means every candidate lead-in is a sentence fragment, so take none at all — + // a fragment is the irrelevant context we are trying to avoid, not context. + let si = null; + for (let i = i0; i > 0; i -= 1) { + if (start - cues[i].start > maxLead) break; + if (isBoundary(cues[i - 1])) { + si = i; + break; + } + } + if (si === null) si = i0; + + // END: the first cue that CLOSES a sentence at or after the quote. Never + // clamp to a budget here — stopping partway through a sentence is exactly the + // mid-thought ending this is meant to remove, so the budget only decides how + // far to look, and failing to find one falls back to the original cue end. + let ei = null; + for (let j = i1; j < cues.length; j += 1) { + if (cues[j].end - end > maxTail) break; + if (isBoundary(cues[j])) { + ei = j; + break; + } + } + if (ei === null) ei = i1; + + return { + start: cues[si].start, + end: cues[ei].end, + leadCues: i0 - si, + tailCues: ei - i1, + }; +} + +async function main() { + const argv = process.argv.slice(2); + const manifestPath = argv.find((a) => !a.startsWith("--")); + if (!manifestPath) { + console.error("usage: resolve-windows.mjs <manifest.json> [--write]"); + process.exit(2); + } + const num = (name, dflt) => { + const i = argv.indexOf(name); + return i >= 0 ? Number(argv[i + 1]) : dflt; + }; + // Lead is where the context lives — it is the run-up that makes a quote make + // sense. Tail only needs to finish the sentence, so it gets a smaller budget. + const opts = { maxLead: num("--max-lead", 8), maxTail: num("--max-tail", 12) }; + + const manifest = JSON.parse(await readFile(manifestPath, "utf8")); + const slug = manifest.provenance.channelSlug; + const cache = new Map(); + + // Cues come from a local corpus when there is one, and from the published + // archive the manifest was built against when there is not — so this runs in a + // clone with no `transcripts/` at all. See cues.mjs. + const flag = (name) => { + const i = argv.indexOf(name); + return i >= 0 ? argv[i + 1] : undefined; + }; + const cues = createCueSource({ + siteOrigin: flag("--site-origin") ?? process.env.SITE_ORIGIN ?? siteOriginFromManifest(manifest), + resolveSiteIds: argv.includes("--resolve-site-ids"), + prefer: flag("--cue-source") ?? "auto", + log: (m) => console.error(` · ${m}`), + }); + const loadCues = (videoId, channelSlug, hints) => + cues.load(channelSlug, videoId, hints).then((r) => r.cues); + + let changed = 0; + for (const e of manifest.timeline) { + if (e.type !== "clip") continue; + // A compilation can span several archived channels (the same streamer's VODs + // are mirrored across more than one), so a clip may name its own. Key the + // cache by channel too — the same id under a different slug is a different file. + // An author can trim a clip to land mid-cue on purpose — a cue often carries + // a whole paragraph, and cutting a quote short is an editorial decision. + // Widening would undo exactly that, so `lock` opts the clip out. + if (e.lock) { + console.log(`${e.id.padEnd(4)} ${e.video.padEnd(12)} locked, left at ${e.start.toFixed(1)}–${e.end.toFixed(1)}`); + continue; + } + const chan = e.channel ?? slug; + const key = `${chan}/${e.video}`; + if (!cache.has(key)) { + cache.set( + key, + await loadCues(e.video, chan, { siteChannel: e.siteChannel, siteVideo: e.siteVideo }), + ); + } + const cues = cache.get(key); + + const before = { start: e.start, end: e.end }; + const w = widen(cues, e.start, e.end, opts); + + // `lockStart` / `lockEnd` pin an edge to exactly what the author wrote. The + // escape hatch exists because sentence detection is only as good as the ASR's + // punctuation, and some uploads have none at all — and because an utterance's + // real trailing pause does not always line up with its last cue's end. + if (e.lockStart) w.start = before.start; + if (e.lockEnd) w.end = before.end; + const dLead = (before.start - w.start).toFixed(1); + const dTail = (w.end - before.end).toFixed(1); + const dur = (w.end - w.start).toFixed(1); + + // Ignore sub-frame drift so a re-run on an already-resolved manifest is a + // genuine no-op rather than a rewrite that nudges every window. + const moved = + Math.abs(w.start - before.start) > 0.05 || Math.abs(w.end - before.end) > 0.05; + if (moved) changed += 1; + console.log( + `${e.id.padEnd(4)} ${e.video.padEnd(12)} ` + + `${before.start.toFixed(1)}–${before.end.toFixed(1)} -> ` + + `${w.start.toFixed(1)}–${w.end.toFixed(1)} (+${dLead}s lead, +${dTail}s tail, ${dur}s)`, + ); + + if (moved) { + e.start = Number(w.start.toFixed(2)); + e.end = Number(w.end.toFixed(2)); + } + } + + // De-overlap clips that come from the SAME video. Widening is per-clip and + // blind to its neighbours, so a tail that finds no sentence boundary runs to + // the budget and can swallow the next clip's material — which plays as the + // same footage twice. (Real case: a 2024 upload whose ASR carries no + // punctuation at all in that stretch, so nothing stopped the search.) + // The later clip's start is the deliberate one, so trim the earlier clip's tail. + const byVideo = new Map(); + for (const e of manifest.timeline) { + if (e.type !== "clip") continue; + if (!byVideo.has(e.video)) byVideo.set(e.video, []); + byVideo.get(e.video).push(e); + } + for (const [video, list] of byVideo) { + if (list.length < 2) continue; + list.sort((a, b) => a.start - b.start); + for (let i = 0; i < list.length - 1; i += 1) { + const a = list[i]; + const b = list[i + 1]; + if (a.end <= b.start) continue; + const overlap = a.end - b.start; + if (a.lockEnd) { + console.log(` ⚠ ${a.id} overlaps ${b.id} by ${overlap.toFixed(1)}s but has lockEnd — not trimmed`); + continue; + } + a.end = Number(b.start.toFixed(2)); + changed += 1; + console.log( + ` de-overlap ${video}: ${a.id} trimmed ${overlap.toFixed(1)}s off its tail ` + + `(it ran into ${b.id})`, + ); + if (a.end - a.start < 3) { + console.log(` ⚠ ${a.id} is now only ${(a.end - a.start).toFixed(1)}s — check it`); + } + } + } + + if (argv.includes("--write")) { + await writeFile(manifestPath, JSON.stringify(manifest, null, 2) + "\n", "utf8"); + console.log(`\nwrote ${manifestPath} (${changed} window(s) changed)`); + } else { + console.log(`\ndry run — ${changed} window(s) would change; pass --write to apply`); + } +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main().catch((err) => { + console.error(err); + process.exit(1); + }); +} diff --git a/umtool/report-to-video/verify-build.mjs b/umtool/report-to-video/verify-build.mjs @@ -0,0 +1,110 @@ +#!/usr/bin/env node +// verify-build.mjs — is the file that came out the file that was asked for? +// +// A build can exit 0 and still be wrong in ways nothing else notices: a concat +// that produced a zero-length file, a chapter pass that silently dropped +// markers, a timeline that lost a clip because --continue-on-error let it. Each +// of those looks like success at the terminal and like a finished video in a +// directory listing. +// +// So the last step of a build measures the deliverable and compares it to the +// manifest. Cheap (one ffprobe) and the only thing that closes the loop. +// +// node umtool/report-to-video/verify-build.mjs <manifest.json> [--out <dir>] +// [--variant sourced|full] [--json] + +import { execFile } from "node:child_process"; +import { promisify } from "node:util"; +import { readFile, stat } from "node:fs/promises"; +import path from "node:path"; + +import { selectVariant, variantPaths } from "./build-video.mjs"; + +const execFileP = promisify(execFile); +const FFPROBE = process.env.FFPROBE_BIN ?? "ffprobe"; + +export async function verifyBuild(manifestPath, { outDir, variant = "sourced" } = {}) { + // The SAME filter the build ran. Verifying the whole manifest against one + // variant's file would report a missing chapter for every entry the other cut + // carries -- i.e. it would be red exactly when the build was right. + const manifest = selectVariant( + JSON.parse(await readFile(manifestPath, "utf8")), + variant, + ); + const root = outDir ?? path.join(path.dirname(path.resolve(manifestPath)), "out"); + const file = variantPaths(root, manifest.slug, variant).final; + const problems = []; + + const st = await stat(file).catch(() => null); + if (!st) return { ok: false, file, problems: [`${file} does not exist`] }; + if (st.size < 1024) problems.push(`${file} is ${st.size} bytes`); + + const { stdout } = await execFileP(FFPROBE, [ + "-v", "error", + "-show_entries", "format=duration,size", + "-show_chapters", + "-of", "json", + file, + ], { maxBuffer: 1 << 24 }); + const probe = JSON.parse(stdout); + const duration = Number(probe.format?.duration ?? 0); + const chapters = (probe.chapters ?? []).length; + const entries = (manifest.timeline ?? []).length; + + if (!(duration > 0)) problems.push("duration is not greater than zero"); + + // Every timeline entry becomes a chapter, so a mismatch means the timeline and + // the file disagree about what is in it -- which is exactly the failure + // --continue-on-error is allowed to cause and must never cause silently. + if (chapters > 0 && chapters !== entries) { + problems.push(`${chapters} chapter(s) for ${entries} timeline entr(ies) — the cut is missing something`); + } + + // A rough floor: the sum of the windows, less the crossfades. Well under the + // real duration because snapping moves the cuts, but a file that came out at + // half the expected length did not build what was asked for. + const wanted = (manifest.timeline ?? []).reduce( + // `seconds` covers cards and the two end-sequence kinds (scroll, chart); + // only a clip's length has to be derived from its window. + (n, e) => n + (e.type === "clip" ? Math.max(0, (e.end ?? 0) - (e.start ?? 0)) : (e.seconds ?? 0)), + 0, + ); + if (wanted > 0 && duration < wanted * 0.5) { + problems.push(`${duration.toFixed(1)}s out of a timeline that asks for about ${wanted.toFixed(0)}s`); + } + + return { ok: problems.length === 0, variant, file, duration, chapters, entries, size: st.size, problems }; +} + +async function main() { + const argv = process.argv.slice(2); + const manifestPath = argv.find((a) => !a.startsWith("--")); + if (!manifestPath) { + console.error("usage: verify-build.mjs <manifest.json> [--out <dir>] [--variant sourced|full] [--json]"); + process.exit(2); + } + const flag = (n) => { const i = argv.indexOf(n); return i >= 0 ? argv[i + 1] : undefined; }; + const res = await verifyBuild(manifestPath, { + outDir: flag("--out"), + variant: flag("--variant") ?? "sourced", + }); + + if (argv.includes("--json")) { + console.log(JSON.stringify(res, null, 2)); + } else { + console.log( + `${res.file} (${res.variant})\n ${res.duration?.toFixed(1) ?? "?"}s · ${res.chapters ?? 0} chapter(s) for ` + + `${res.entries ?? 0} entr(ies) · ${((res.size ?? 0) / 1e6).toFixed(1)} MB`, + ); + for (const p of res.problems) console.log(` ** ${p}`); + if (res.ok) console.log(" ok"); + } + process.exit(res.ok ? 0 : 1); +} + +if (import.meta.url === `file://${process.argv[1]}`) { + main().catch((err) => { + console.error(err.message ?? err); + process.exit(1); + }); +}