# The `report-video` kind A report video is a **cited timeline**: a sweep report's findings, said by the sources in their own voice, in order, each clip carrying a burned-in attribution and a QR that resolves to that exact moment in the archive's own viewer. The point is that a viewer does not have to take the edit on trust. The pipeline lives at `umtool/report-to-video/` and its own [README](../report-to-video/README.md) is the source of truth for how it renders. This sheet covers what that README cannot know: what umtool reads, what umtool **writes**, and the rules a UI has to honour so the CLI and the app can never disagree. ## The manifest is an EDL, and it is the project `/video.manifest.json` is both the marker file for the kind and the edit decision list. Nothing else needs to exist for a directory to be a report video. ```jsonc { "schemaVersion": 1, "slug": "quartering-gout", // out/.mp4 is the deliverable "title": "…", "subtitle": "…", "generatedOn": "2026-08-12", "provenance": { … }, // how the sweep was done, and its caveats "render": { … }, // resolution, fonts, palette, the cut knobs "timelineNodes": [ … ], // the footer's progress track, if any "timeline": [ … ] // the ordered cut. Array order IS the cut. } ``` **Ordering is array order.** There is no `index` field and no sort. A reorder is a move within the array, and it has to recompute `sectionEnter` — the flag that says "this clip is the first of its section", which is what makes the footer marker travel. umtool never reorders as a side effect of a window edit. ### A clip entry ```jsonc { "type": "clip", "id": "c04", "video": "uyz1_FIqIEk", "channel": "hasanabi-vods3", // optional: which archived channel's cues "channelTitle": "HasanAbi", // optional: what to CALL that channel on screen "start": 32980.24, "end": 32994.19, // the reviewed EXTENT, absolute source seconds, 2 dp "cutStart": 32986.1, "cutEnd": 32992.4, // the CUT that plays, inside the extent "lockCut": true, // the cut is deliberate; the resolver leaves it "cite": 32989, // the second shown in the attribution line "citeUrl": "https://…", // optional: overrides the derived QR target "quote": "…", // the words this clip exists for "title": "Monday Mail #11", // optional: overrides the record's title in the header "date": "2016-09-12", // optional: overrides the record's upload date "note": "…", // why it is in the cut (editorial, for humans) "correction": "…", // what the REPORT got wrong here; never rendered "verdict": "confirmed", // or "incorrect"; absent = nobody has looked yet "chapter": "…", // chapter title; falls back to date + title "section": 2, "sectionEnter": true, "lock": true, "lockStart": true, "lockEnd": true } ``` ### The attribution: `channelTitle`, `title`, `date`, `cite`, `citeUrl`, `quote` The header line is `${channel} · ${title ?? cleanTitle(record.title)} · ${date ?? record.uploadDate} @ ${hms(cite ?? start)}`, built by `report-to-video/attribution.mjs` — one module, imported by the renderer, by the clip bench's live preview and by the writer's validation, so a preview cannot promise a line the renderer would not draw. **The channel heads the line**, and it is not furniture. A compilation routinely spans a streamer's own channel, two VOD mirrors and a guest's show; before the channel was there every line looked like the same source, and the only way to tell whose room you were in was to recognise it. `channelName()` walks four answers, each because the next one is wrong somewhere: 1. **`channelTitle`** — the per-clip override. A record's uploader name can be a handle, a rebrand or an ALL-CAPS mirror name; this is the author saying what to call it in this cut. Hand-edited, deliberately: it is a fact about a *channel*, and a per-clip field in a bench invites a cut whose two clips from one channel disagree about its name. 2. **the cue record's `channel`** — the uploader's display name, carried by both a local `transcript.cues.json` and a published shard record. Free: that record is already loaded for the title and the upload date. 3. **`provenance.channel`** — the sweep's own display name, for a clip from the sweep's own channel. A hand-written manifest has it even when a cue record predates the field. 4. **the raw slug** — ugly on screen and meant to be. A visibly wrong name is a thing somebody fixes; a silently empty head just looks like a design. **`title` and `date` exist because the archive's answer is often the wrong one.** A VOD mirror's `uploadDate` is the date the *copy* was posted, routinely years after the stream; `destinys-child` overrides both on all 19 clips. `chapterTitle()` honours the same three so the mp4's chapter list and its burned-in header cannot disagree about who, what or when. > **The old "both absent reproduces the previous line byte for byte" property is > gone for clips, on purpose.** Adding the channel changes every clip header in > every cut, which is why it is a deliberate re-render rather than a quiet > upgrade: a cut built before this and a cut built after it differ on screen. It > still holds where no channel can be named at all, for card chapter titles, and > for `image` entries — a still's line is `title · date` and never carries a > channel, because there is no archived channel behind a screenshot. All five are **writable** — from the clip bench, from `PUT /api/report/window`, and from `umtool window --title/--date/--cite/--cite-url/--quote` — under the four rules below. An **empty value deletes the key**, the way the locks do: a manifest is read by humans and `"title": ""` is noise that reads like a decision. Two are checked rather than trusted: - **`date`** must be `YYYY-MM-DD` *and a day that exists*. The regex alone accepts `2025-02-31`, which reads as a date right up until somebody tries to check the clip against the stream it claims to come from. - **`citeUrl`** must be `http(s)`. A QR encoding anything else is the same class of defect as the 19 codes that shipped reading `undefined/?v=…`. ### `correction` — what the REPORT got wrong Free text, written while watching the clip, **never rendered and never read by the renderer**. It records a defect in the *report*, not in the cut: wrong speaker, wrong addressee, wrong date. `destinys-child`'s `c07` is the case it was built for — the report attributed a line to Destiny that Dan says, to Destiny, on a call. It is deliberately not a decision in the inbox: the fix belongs to the next sweep, not to this manifest, and nothing here can close it. `correctionsOf()` is the one definition; the project page lists it and `umtool corrections ` prints the same list as markdown with the QR's own moment link per bullet, ready to paste into the next prompt. ### Extent vs cut — `start`/`end` and `cutStart`/`cutEnd` **They answer different questions, and conflating them is why a reviewed clip used to have to be re-reviewed.** `start`/`end` is the **extent**: how much of this recording is worth having, decided once by somebody watching it — *"oh, I see how useful this clip could be"* — and not by any one paragraph's needs. `cutStart`/`cutEnd` is the **cut**: the sentences the quote is actually made of, which is what the video plays. The cut is **derived**, by `resolve-windows --cut-to-quote`, so the same reviewed extent can be re-cut for a different report without watching anything again. It must lie **inside** the extent; both fields or neither; at least half a second. The writer refuses anything else and `umtool check` reports it as `clip-cut-outside`, **blocking** — a cut outside its extent renders seconds nobody reviewed, and half a pair renders the whole extent while the manifest reads as though it were trimmed. - **The build** cuts to `[cutStart − render.leadIn, cutEnd]` (lead-in 0.4 s by default, clamped into the extent) when both are present, and to the extent when they are not — so a manifest written before any of this behaves exactly as it did. - **Fetching and the cache stay keyed on the EXTENT.** A cut is always inside it, so re-cutting never needs another download. - **`cite` is kept, never moved.** If it falls outside the cut the build says so and leaves it: a citation silently re-pointed is worse than one flagged. - **`lock` is about the extent; `lockCut` is about the cut.** The resolver honours both, and never overwrites existing cut fields without `--force-cut`. The matcher is deliberately loose, because a quote is prose a human wrote about speech a machine transcribed: case and punctuation are dropped, `[bracketed]` editorial words and `[Speaker]` markers are dropped, and an **ellipsis splits the quote into fragments** located independently — `"A … B"` is two places in the recording and the cut spans from the first to the last. A fragment counts as located at half its words; the whole quote must reach `--min-match` (0.6) or the clip is reported **UNMATCHED** with its best partial and nothing is written for it. That report is the useful output: a quote the transcript does not contain is a fact about the *report*, and a cut invented to hide it would be the worst of both. ### `verdict` — what the walk said `"confirmed"`, `"incorrect"`, or **absent**. Never rendered, like `correction`, and written by the same three writers (the clip bench's `y` and `x`, `PUT /api/report/window`, `umtool window --verdict confirmed|incorrect`). An empty value deletes the key. A clip is a *claim*: the report said somebody said this, here. Walking the cut is somebody checking that claim against the audio. | on disk | state | `correction` | |---|---|---| | `verdict: "confirmed"` | **confirmed** | optional — a *note* | | `verdict: "incorrect"` | **incorrect** | **required** | | neither | **not yet reviewed** | — | **The note is not always a complaint.** On a confirmed clip `correction` says why a good clip's window moved, or something a reader should know; on an incorrect one it says what the report got wrong. `umtool corrections` and the project page split on the *verdict*, not on "has text", so the next pass is never sent to fix a clip nobody complained about. **An incorrect verdict with no note is a complaint nobody can act on**, so the writer refuses it. The rule is a property of the entry *after* the patch, which makes both halves one check: setting `incorrect` with no note, and clearing the note off a clip that is already `incorrect`, are the same error and get the same sentence — *"an incorrect verdict needs its note: say what the report got wrong"*. The bench sends `{correction, verdict: "incorrect"}` as one patch for exactly that reason. **An absent key is the point.** A manifest nobody has walked says so by carrying nothing, rather than by carrying `"verdict": "unreviewed"` on sixty entries — which reads like a decision and would have to be written before the walk could start. **Legacy is read, not migrated.** A clip carrying a `correction` and no `verdict` was written before `incorrect` existed and means exactly that: `clipVerdict()` reads it as **incorrect**. The required-note rule is keyed on the explicit value, so those entries can still be cleared — a rule that refused to let somebody undo a note they wrote before the rule existed would be a trap. `reviewOf()` is the one definition of coverage — the bench header's `N reviewed`, the project page's "Walk the cut" button and `umtool corrections ` all read it, and `walkReadiness()` reads the same verdict rule to decide what the walk VISITS (unjudged and fetched — see [clip-bench.md](clip-bench.md#what-the-walk-skips)), and the last opens with `N clips: A incorrect, B confirmed (C with a note), D not yet reviewed (ids: …)` so the walk's coverage is visible beside the defects it found. A card entry is `{"type":"card", "id", "style", "seconds", …}` — see the pipeline README for the styles. Cards have no window and no source. **The entry vocabulary is OPEN, and code must treat it that way.** A CLIP is `type === "clip"`; everything else is a non-clip entry with no window and no source, and it is *not* necessarily a card. One real manifest (`quartering-employee-count`) carries `scroll` and `chart` entries beside its ten cards. Anything that branches on card-or-clip will send `undefined` into a path join the first time it meets a third type — which is exactly what 500'd the project page and crashed `umtool show`, on the one manifest of six that has any. The fixture now carries an entry of a type nothing in the code knows about, so this cannot come back. ### An image entry ```jsonc { "type": "image", "id": "i01a", "src": "nathan-pictures/IMG_7095.jpeg", // relative to the MANIFEST's directory "seconds": 6, // how long it is on screen; default 4 "redact": [[630, 925, 180, 145], // optional: solid boxes, in SOURCE pixels, [615, 1055, 185, 175]], // filled BEFORE the crop "crop": [0, 0, 1260, 1020], // optional [x, y, w, h], also source pixels "scroll": { "holdStart": 2, "holdEnd": 3 }, // optional: crawl it; `true` = 1.5 / 2.0 // …or, INSTEAD of `src`, a row of 2-4 captures, each with its own redact/crop: "panels": [{ "src": "post.png" }, { "src": "poster.jpg", "redact": [[630, 925, 180, 145]] }], "gap": 48, // px between panels; default 40 "title": "Bx (@bx_on_x) on X, September 2026 — 1/4", "date": "2026-08-19", // optional; appended as a clip's is "citeUrl": "https://…", // optional; the ONLY thing that draws a QR "note": "…" } ``` A still is the receipts a clip cannot say out loud — a post, a thread, a DM, a poster. Its header is `${title} · ${date}` and carries **no clock**: there is no moment inside a screenshot to cite, so a `@ 0:00` would be a number pointing at nothing. It has no window, no source and nothing to fetch, so `--fetch-only` on one is a no-op that says so rather than "not a clip", and `updateClip` refuses it like any other non-clip entry. **`crop` and `redact` are both expressed in the ORIGINAL file's pixels**, and redaction runs first so the numbers come straight off the file the author measured in an image viewer — not off a cropped intermediate whose origin moved. Both are validated against the source's real dimensions, because ffmpeg would clamp instead: a crop that silently became a different framing looks like a decision somebody made, and a redaction that silently became a smaller redaction is the exact failure the field exists to prevent. > **Redaction fills; it does not blur.** A screenshot can carry a live QR code, a > phone number or an address that must not ship in the video, and **a QR survives > mild blur** — a pixelated code still scans, which is a redaction that looks done > and is not. A solid box in `palette.bg` cannot be undone. `destinys-child`'s > `i00b` is a recruitment poster with two working codes, and it is checkable: > `zbarimg -q nathan-pictures/HScWQQWXEAE3Ie_.jpg` reads both off the source and > `zbarimg -q` on a frame grabbed from `out/sourced/segments/i00b.mp4` reads > none. **`scroll` is for a capture that does not fit.** A full tweet thread in a 1920x1024 hole leaves two bad options — shrink it until the text is unreadable, or slice it into four stills and make the viewer reassemble the thread — so the picture is scaled to the picture area's **width**, which is what makes a phone capture's text large, and the frame walks down it: hold the top for `holdStart`, travel linearly over `seconds - holdStart - holdEnd`, hold the bottom for `holdEnd`. Two caps: never wider than the area, and never more than 2x (a small capture blown up to 1920 wide is a wall of artefacts, so it is centred at 2x on `palette.bg` instead). A picture that fits after all is held still — a manifest saying `scroll` on a short capture is describing an intent, not demanding motion — and `seconds <= holdStart + holdEnd` is refused, because that is two holds with a jump cut between them rather than a slow crawl. > **The motion has to be `crop`'s `y`.** crop's `w`/`h` are config-time (`t` is > undefined there) but its `x`/`y` are evaluated per frame, which is exactly the > asymmetry that left the footer's fill bar drawn at its final width for weeks: > `drawbox` has no time variable at all. The input is `-loop 1 -framerate > -t `, because a PNG on a plain `-i` through an animated crop is > FROZEN — the crop sees one frame at `t=0` and repeatlast repeats the > already-cropped result. Both rules are the rail's, and a crawl is checkable > the same way: grab frames either side of each hold and compare them. **`panels` is the same still with several exhibits in it** — two phone captures that argue with each other, side by side. It replaces `src` (an entry with both is an error, and so is `panels` with `scroll`: those are two different answers to "it does not fit", and crawling a row would also have to decide which panel the crawl follows). Each panel goes through **its own** `redact` → `crop`, in its own source pixels, and is then scaled to the picture area's **height** — matching heights is what makes two differently-sized captures read as one exhibit rather than as two pictures that happen to be near each other. A row wider than the frame shrinks by **one factor applied to every panel**. A per-panel fit would also keep the row fitting, and would quietly restate the relative sizes of the exhibits — a claim about them nobody made. The gaps are ground rather than picture, so they keep their width while the pictures give way. Everything after the layout — header, `date`, `citeUrl`/QR, `seconds`, the chapter — is a single still's exactly. `destinys-child`'s `p05` is the real case: a 327x675 expulsion post beside an 853x1280 poster at `gap: 48` lays out as `496x1024 + 682x1024 = 1226px of 1920`, centred with 347px of background either side, and the poster's two recruitment codes filled black. A still **never derives** a QR target. A clip's code is derived from the archive that serves it; a screenshot has no such archive, and a code that resolves to the wrong place is worse than no code — so only an explicit `citeUrl` draws one. ## The on-screen deck (`render.chrome`) Opt-in, and it **replaces** the header, the footer's progress track and the corner QR described above — a deck manifest draws none of them, because the panel carries the citation itself. The full contract (every setting, the schedule document, the choreography) is the pipeline's own README ([`render.chrome` — the on-screen deck](../report-to-video/README.md#renderchrome--the-on-screen-deck)) and `plans/onscreen-deck.md`; this sheet covers what umtool reads, writes and serves. Any `clip`, `image` or `card` entry may carry `onscreen: { title?, subtitle? }` — nested, because `entry.title` already means the *stream's* title everywhere else on this page, the same reason `title`/`date` are their own fields above. `updateOnscreen(dir, { [id]: {title?, subtitle?} | null }, { token })` writes a whole batch through the same lock / backup / atomic-write / stale-token path as every other writer in this file; **one unknown id refuses the whole batch**, nothing partial is written. A value normalises to `null` — which DELETES the key — the same rule an empty attribution field follows. `updateClip` takes `onscreen` too, so a clip-bench save and an On-screen-table save are one rule, not two. A row may also carry the entry's fact-check claim beside its text, `{ title?, subtitle?, claim: { id, verdict } | null }` (`splitOnscreenRow`): `claim` is set on the entry (never inside `onscreen`), `null` removes it, and a row without the key leaves it as it is. After applying the batch the writer runs `validateClaims()` over the whole timeline — one verdict per claim id, none on a teaser — and a refusal writes nothing. The On-screen table shows a **claim** column (an id and a verdict) per row, `GET /api/report/onscreen` returns each row's `claim`, and PUT returns the `claims` it stored. The preview applies a draft's claims too: the build's stamps are placed again over its segments (`stampSchedule`), the deck's page draws the tally from them, and the stamps' own composition (`chrome/stamp-preview/`, served under `stamp-preview/`) is laid over the footage for the whole scrub. `updateChrome(dir, chrome | null, { token })` writes `render.chrome` itself: `validateChrome()` first, so a block umtool saves is one the build accepts, and `null` removes the key — turning the deck off without touching anything else under `render`. ### Preview, before and without a build `POST /api/report/chrome/preview` composes `out//chrome/deck-preview/` — never the build's own `chrome/deck/` or its render cache — from `{ project, variant?, draft? }`. `draft` is the On-screen table's unsaved rows, id → `{title?, subtitle?}` or `null`, applied on top of the saved manifest before a schedule is built. The schedule it draws from is the **build's own** (`out//schedule.json`, real probed durations) when one still lists this cut's entries in this order, else an **estimate** from the manifest alone (`estimated: true`) — so the section previews before anything has ever been built, and a finished build makes the preview's handovers land exactly where the render's will. `GET /api/report/chrome/files/[...path]` serves that preview directory, traversal-guarded — the composition's own HTML, loaded in an iframe that is **same-origin and unsandboxed** so its `@font-face` rules resolve (a sandboxed opaque origin fails every font fetch and falls back silently; see [quirks.md](quirks.md)). `GET /api/report/still?project&variant&(clip|at)` renders a **true** still, the same Chromium screenshot the pipeline's `--still` takes, of what is **saved** — not of the live preview's unsaved typing, which is the point of the two existing side by side. ### What a build leaves behind `GET /api/report/video?project&variant&kind=final|preview` range-serves either the deliverable or the short window `--chrome-preview` wrote, through `rangeResponse()` (`lib/report/serve.mjs`), shared with the segment route and the same `v`-matches-mtime caching rule. In the build chain (see [build.md](build.md)), the driver's `chromeOnly` option is `--chrome-only`, labelled **"re-render on-screen"**: it skips the availability preflight and the dry resolve, same as chapters-only and rail preview (none of the three touches a source), but keeps the build's own timeout rather than their five-minute cap, because a deck re-render is a render plus a concat of the whole cut, not a seconds-long retitle. The pipeline's other two deck flags, `--no-chrome` and `--chrome-preview`, are CLI-only — the driver does not wire them up. ### Naming The app calls this **"On-screen"** everywhere a person reads it. The manifest and the pipeline call it the deck (`render.chrome.layout: "deck"`), but "deck" is already the song kind's name in umtool, and one word meaning two things two pages apart is how somebody edits the wrong one. ## The three-stage window model A clip's window passes through three different notions of "where the cut is", and confusing them is the source of most window bugs. | Stage | Where it comes from | Whose problem | |---|---|---| | **1. Cue span** | `transcript.cues.json` — the cues covering the quote | The sweep. A report only records a single start second, so a window cannot be recovered from the report alone. | | **2. Sentence window** | `resolve-windows.mjs` walks outward to a cue ending in `.?!` | Meaning. A cue boundary is a *line-wrap* boundary; cutting there drops the lead-in that makes a quote make sense. | | **3. Audio cut** | `build-video.mjs` snaps to a silence found in the fetched audio | Sound. A sentence boundary is still not a *speech* boundary; snapping is what stops clips cutting through a word. | Only stage 2 is stored. Stage 1 is recoverable from the cue file, and stage 3 is recomputed on every build from the audio — which is why a build is reproducible from the manifest alone, and why the clip bench edits stage 2 and shows you stage 1. **Everything is in absolute source seconds**, from the manifest through the bench to the API. The fetched file's own start (`fetchStart`) is the only place a relative number appears, and it is always derived, never stored. ## Where the bytes come from: the editor, not yt-dlp A clip window is no longer downloaded by this pipeline. `POST /api/report/fetch` asks the **editor** (`POST /api/media/fetch-window`), which fetches it through its managed download path and writes it into the corpus: ``` channels//data//clips/-.mp4 channels//data//clips/-.json # who asked, and why ``` Four things follow, and each is the reason: - **Politeness.** The editor owns the channel's cookie policy and its `ytdlpExtraArgs`, the `--sleep-requests`, the per-platform 429 cooldown and the backoff that opens it. Running yt-dlp from here had none of them, and a burst is what trips a bot check. - **Reuse across reports.** `out/clips-raw` is one project's directory. The corpus is keyed by video, so a window one report paid for is a window the next one citing the same stream gets for free. Containment is the predicate, so a deliberately generous fetch covers its neighbours too. - **Provenance.** The sidecar records `requestedBy`, the manifest and clip it was wanted for, the reason in the author's own words, the pad, the size and the argv. A job log is pruned; the window is not, and a directory of windows nobody can account for is worse than no directory. - **It is not a download.** The fetch deliberately writes no `metadata.info.json`. The archive index keys a video's presence on that file, so a window must not make an undownloaded video look downloaded — the video's pipeline state and its absence from the index are unchanged. The editor's video page says so, in those words, under *Fetched windows*. `UMTOOL_LOCAL_FETCH=1` restores the old path (`build-video.mjs --fetch-only`) for a machine with no editor to ask. The step shape, the argv flags and the NDJSON events are identical either way, so the job runner, the bench's progress readout and the cache predicate cannot tell the two apart. The editor is named by `ARCHILYZER_EDITOR_URL` (default `http://localhost:3001`) and asked with the shared `WORKER_TOKEN` — the same secret `/api/worker/*` takes, because this endpoint spends a source's patience and must be no easier to reach. Both come from the environment, never from a request and never from a job step's `env` (which `jobView()` shows the browser). An unset token is an error that names the variable. ## Writing a window: the four rules umtool is a *second* writer of a file the CLI also writes. So: 1. **Round to 2 dp.** Non-negotiable. `resolve-windows.mjs` is a fixed point and both its `EPS` lookup tolerance and its 0.05 s deadband assume 2 dp storage. Writing 4 dp makes the widener grow the clip on every subsequent run. 2. **Re-read before writing, under `withStateLock`, and write atomically.** The file has two writers; a lost update here is lost human judgement. 3. **Preserve the CLI's formatting** — `JSON.stringify(m, null, 2) + "\n"`. The default `writeJsonAtomic` uses indent 1, which turns a two-number edit into a 600-line diff. 4. **Guard with an mtime token.** A `PUT` carrying a stale token is a 409, not a silent overwrite — an agent or a CLI run may have written in between. A `.bak` of the last hand-authored state is kept (rate-limited, not a rolling stack); `ferret-rescue/video.manifest.json.bak` already set that precedent. ## `lock`, `lockStart`, `lockEnd` These are **editorial acknowledgements**, not advanced options, and the data says so: across the six real manifests ferret-rescue locks 1 clip of 10 and *every other manifest locks 100%*. - **`lock`** — resolve-windows must not touch this clip at all. The two cases: the author deliberately cut a quote short (a single ASR cue often carries a whole paragraph, so trimming to the sentence that matters is a real decision that widening would undo), or the lead-in would drag in seconds of some *other* audio. A live case from `quartering-gout`: two clips whose source cue file has two cues spanning almost the whole runtime, where an unlocked widening pass would expand a 35-second window to a 2,488-second one. - **`lockStart` / `lockEnd`** — pin one edge exactly. Needed because sentence detection is only as good as the ASR's punctuation, and some uploads have none. `lockEnd` is also how the "ends mid-sentence" warning is *acknowledged*: setting it is the author saying "I meant to cut here", and it suppresses the warning. This matters — one 14-clip cut shipped with 8 clips ending mid-thought. **The rule that prevents silent loss:** if you move an edge in the bench to a value `widen()` would not produce, the bench offers to set the matching lock, defaulted on. Without it, the next `resolve-windows --write` reverts your edit. ## Provenance, and the two defects that shipped `provenance` is prose written by the sweep, and most of it is for humans. Three fields are load-bearing: - **`siteOrigin`** — the archive origin every QR is built from. **Nothing validated this**, and two real videos shipped broken: `quartering-employee-count` has no `siteOrigin` at all (19 QR codes encoding `undefined/?v=…`) and `ferret-rescue` has `http://localhost:3000` (codes that resolve to nothing on a phone). This is the single best reason to run `umtool check` before every build. - **`channelSlug`** — the default archived channel for cue lookups, overridable per clip. - **`availabilityCheckedOn`** — when sources were last confirmed fetchable. Stale in both directions; `check-availability.mjs` refreshes it. ## The Rumble two-ids rule A Rumble video has **two** ids: the site/MCP `video_id` is the *embed* id, while the local cue directory is named for the **URL slug**. The manifest's `video` field must be the local slug; the report cites the site id. A manifest that uses the site id finds no cue file — and finds that out twenty minutes into a build, unless `umtool check` ran first. `citeUrl` exists for the mirror case: a clip taken from a copy whose archived transcript is broken should point the QR at the copy that reads. ## Deleted sources A source can be deleted from YouTube after the manifest is written, and clips are a **live network fetch** — there is no local media. The archive's own availability data is a *snapshot*, so it goes stale in both directions. Options, in order of preference: cut the same moment from a live mirror (via a `CHANNELS_DIR` shadow directory — but check the archive's alignment statement first, because mirrors are not assumed to share a clock), or convert the clip to a quote card. ## The ledger, and why it must be adjudicated `ledger[]` is the claim set the rail and the closing cards read. Beside the descriptive fields (`date`, `value`, `display`, `label`, `quote`, `src`, `entryId`) every entry carries **six adjudication fields** — `scope`, `scopeBasis`, `scopeConfidence`, `population`, `valueKind`, `flags` — plus the `channel` / `video` / `cite` triple that lets a claim page fetch its own audio. They exist because `company` was an **undocumented interpretation** and four hazards were riding on it: | hazard | what it looked like | |---|---| | scope ambiguity | *"I have 10 employees, my coffee company employees…"* — all-companies or coffee-only, depending on where the comma falls | | derived, not stated | *"10 at coffee brand coffee, I've got eight staff for the live stream"* recorded as **18**, a figure nobody utters | | population drift | "full-time salaried", "employees" and "all basically contractors" compared as one series | | synthetic values | 10.5 for *"about 10 people, 11 people"* — a midpoint we invented and attributed to him | **The rule that follows:** a *stated* total may contain only **a figure he utters as a single number for a named scope**. Sums and midpoints are ours and belong to the *implied* series, which says so on screen. `umtool/report-to-video/ledger-totals.mjs` is the single implementation of that arithmetic — the umtool reducer, the chart band and the closing card all import it, so they cannot disagree. It **refuses to run on an unadjudicated ledger** (`strict: true`); the inbox passes `strict: false` so it can show coherence rows while the rest is still being worked. **Six** named predicates compute incoherence rather than asserting it: `contradicts_component`, `same_day_conflict`, `self_negating`, `population_mismatch`, `not_his_number`, `status_flip`. Deliberately **not** a predicate: a large rise or fall between claims — fluctuation is usually the *subject*, and flagging it would be putting a thumb on the scale. A seventh **optional** field, `roles`, records who he named when he enumerated rather than counted. It is not part of the gate — see [claim-bench.md](claim-bench.md). Rule on them at [`/browse//claim/`](claim-bench.md). ## One manifest, two cuts `timeline` entries may carry `"variant": "full"`, and cards may carry `variants: { sourced: { … } }` field overrides. `selectVariant()` in `build-video.mjs` applies both once, immediately after the manifest is read, and then filters `ledger` to the claims whose `entryId` survived. What umtool has to know about it: - **`out/.mp4` is still the `sourced` deliverable**, which is why `buildStateOf()` keeps working unchanged. `full` writes `out/-full.mp4` beside it, and registering that as a second output is a follow-up. - **`out/clips-raw` and `out/availability.json` stay at the root** and are shared; `cards/`, `segments/`, `qr/`, `chrome/` and `schedule.json` moved under `out//`. Anything reading a raw window (the clip bench, the claim bench) is unaffected; anything reading a segment needs a variant. - **The driver passes `--out /out`**, which is the ROOT, so it is correct as written and builds the default `sourced` cut. - A **`ledger`** timeline entry is a new type, and the vocabulary is open, so `clipsOf`/`cardsOf`/the "other entries" bucket already handle it. A window cannot be written to one — `updateClip()` refuses any non-clip. ## Deleted sources are a probe result, not an annotation `check-availability.mjs` probes every **ledger** source as well as every clip source, and `out/availability.json` records the verdict with the date it ran. The `full` cut prints that verdict on screen (`source deleted` / `source unreachable` / `not clipped`), so a hand-written "source since removed" in a card kicker is now a bug: several unclipped claims come from videos the cut clips elsewhere, i.e. demonstrably live. **Severity follows the CLIPS, not the source.** Widening the probe to ledger sources immediately put six permanent `blocking` rows in the inbox reading *"0 clip(s) cite it; the build dies here"* — a sentence that refutes itself. A dead source that no clip cites is exactly the case the `full` cut quotes on a card **because** it is gone, so it is `info`: named, dated, and not in anybody's way. It becomes blocking the moment a clip cites it. ## A field that changes meaning has to be chased through every renderer `valueKind: "derived"` was set on the 18, and the first build still shipped a rail tally reading **"Spans everything 18"** directly beside a card reading *"He never said eighteen"*. Three separate readers were still taking the old meaning: the tally took the latest value of any kind, the ledger's own `label` still said "the only explicit sum in the corpus", and the closing chart plotted off a legacy `plotted` flag. Nothing caught it. Not the tests, not `verify-build`, not `umtool check` — every one of them was true about a video that contradicted itself on screen. **Looking at a frame** was the only thing that found it. ## Discovered by getting it wrong once - **A build that loses a clip is worse than a build that fails.** `--continue-on-error` finishes everything buildable and then *refuses to concatenate*, because a finished file quietly missing a citation looks complete. - **The cache is content-addressed by window, and that was being wasted.** Before containing-window reuse, every window edit was a fresh download: `ferret-rescue/out/clips-raw` holds 31 files for 10 clips, one source four times over overlapping windows. - **Mirrors are not assumed to share a clock.** Timestamps mapped from one mirror to another have to be verified against the mirror's own cue text, not assumed. - **Every entry is a chapter, so counts must agree.** `verify-build.mjs` compares chapter count to *timeline* length, not clip count — a mismatch means the file and the timeline disagree about what is in it. - **`--chapters-only` is cheap and separate.** Retitling chapters does not need a re-encode; the per-clip segments on disk are all the offsets need. - **The channel's display name is on the CUE RECORD, not in `config.json`.** A corpus channel's `config.json` carries `name`, not `title`; the field the header wants is `channel` on the transcript record itself — which a published shard carries too, so naming the channel costs a build no lookup and works with no corpus at all. - **Blurring a QR code does not redact it.** Codes carry enough error correction to survive the sort of blur that reads as "censored" on screen, so a blurred recruitment link still scans off a paused video. `redact` fills solid, and the test for it is `zbarimg`, not an eyeball. - **A no-op still has to be a sentence.** `--fetch-only` on an image emits a `done` event with no file, and the human formatter used to print `built null` under it. An event a consumer needs is not an excuse for a line a reader does not. ## Sources are a view of the project `sourcesOf()` in `lib/projects/report.mjs` gives one row per distinct `(channel ?? provenance.channelSlug, video)`: the clips using it; the cue file's state (present / missing / no-punctuation); **cue coverage** — the last cue's end against the latest clip end, so a clip past the transcript's end reads as `cues end 880 s · c03 needs 897 s` instead of living in a `transcriptGapNote`; availability from `out/availability.json` with its age, or "never checked"; and the cite target — derived `${siteOrigin}/?v=…` or the per-clip `citeUrl` override, marked as such when it differs (every clip in one real manifest carries a `citeUrl`; only the Rumble ones point elsewhere). The project page, `umtool show` and the dashboard's Sources panel all read this one function. Nothing in it probes. ## Revisions, instead of `.bak` `umtool snapshot [--label ]` (or the button) copies the manifest byte-for-byte to `revisions/[-