commit 85d56ad907bdaed9c18234c84115b63184dfd593
parent c1da6127406dd069613e8d3e999f96f787cb7c42
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Fri, 9 Oct 2026 11:11:32 -0400
Merge main into r18/integration (umtool articles + notes, ops transcribe with word timings, fetch-windows, metadata refresh, channel create/rename/delete over ops); the changelogs keep both sides, r18 first
jobDetail reads both a publish stage's run and a fetch-windows batch;
GET_ARG_OPTIONAL keeps tags, channels and publish; PUBLISH.md keeps r18's
"three ways to drive it" table with main's fetch-windows row, and main's
two-step reports paragraph.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Diffstat:
204 files changed, 17456 insertions(+), 394 deletions(-)
diff --git a/AGENTS.md b/AGENTS.md
@@ -140,6 +140,21 @@ better), whose clock is not the same.
This tooling is on `main` as of 2026-08-20.
+# Operator notes on articles and videos
+
+```sh
+umtool notes --all --open # from the checkout: node umtool/bin/umtool.mjs notes --all --open
+```
+
+The operator leaves notes in umtool on articles (`/sites/<site>/<report>`) and on
+report-video projects (rows, takes, moments in a cut). They live in a `notes.json`
+beside the report's `report.json` (never published) or beside the project's
+`video.manifest.json`. `umtool notes <site>/<report>` prints each open note with its
+anchor resolved and the **source** file to edit — a draft, not the generated
+`report.json` or manifest. Act, regenerate, then `umtool notes reply <id> "<what
+changed>" --resolve`. Never hand-edit `notes.json`. See
+[umtool/docs/notes.md](umtool/docs/notes.md).
+
# The runtime container
`docker compose up -d` stands up a working archive: the editor plus Caddy, with
diff --git a/PUBLISH.md b/PUBLISH.md
@@ -57,17 +57,25 @@ for exactly the records that can hold one.
### Reports and cited sites
-A site with `reports` (site.json) has one more step before its build. It is not a
-stage: run it yourself (or from the site's **Reports** tab), or that site's build
-fails on the first citation with no prepared media. **reports prepare** cuts the
-evidence clip of every span its published reports cite — from the editor's clip
-windows, the saved-video store or a record's audio, fitted inside 1280×720 (H.264
-crf 23, AAC; an audio span is an `.m4a`) — and copies the screenshot and media of
-every cited post, only those, into `.export-index/sites/<id>/report-media/`, with a
-manifest (`index.json`) of each moment's file, size, hash and duration. Nothing is
-fetched: a citation whose media is not on disk, a clip over 24 MiB or an invalid
-report is listed and fails the run (exit 1, or a failed job) — fetch the window or
-persist the video, capture the post, and run it again; what is already cut is reused.
+A site with `reports` (site.json) has two more steps, before its build and on the
+host. First, **fetch the missing evidence**: *Fetch missing evidence* on the Reports
+tab, or `pnpm ops fetch-windows --json '{"siteId":"<id>"}' --wait`, fetches the
+window of every cited span whose media is not on disk, through the editor's managed
+path, as one paced job per platform (YouTube and Rumble side by side; at least 30 s
+between two windows on one platform, a 429 or two 403s in a row backing the
+platform off and stopping its job). `"dryRun": true` (*Preview missing evidence*)
+lists the windows, and the spans no window can fill — a video the source says is
+deleted or private, a channel off the site, a drive that is not mounted — without
+fetching. Running it again resumes: fetched windows are on disk. Then **reports
+prepare** cuts the evidence clip of every span its published reports
+cite — from the editor's clip windows, the saved-video store or a record's audio,
+fitted inside 1280×720 (H.264 crf 23, AAC; an audio span is an `.m4a`) — and copies
+the screenshot and media of every cited post, only those, into
+`.export-index/sites/<id>/report-media/`, with a manifest (`index.json`) of each
+moment's file, size, hash and duration. Prepare fetches nothing: a citation whose
+media is not on disk, a clip over 24 MiB or an invalid report is listed and fails
+the run (exit 1, or a failed job) — fetch the missing evidence or persist the video,
+capture the post, and run it again; what is already cut is reused.
When nothing is missing, prepare ends by **exporting** the reports (`reports export`
runs the same step alone): each published report is written, as compose would
@@ -273,6 +281,7 @@ the same lock.
| What the lane would run | /sites → **Publish now** | `"now"` | `publish now` |
| The status | the /sites Publish panel; /operations/publish | `pnpm ops get publish` | `publish status [--json]` |
| The source mirror alone | — (every homepage build runs it) | — | `source publish [--force] [--check] [--keep-scratch]`, `source audit [<git dir>]` |
+| Fetch the clip windows a site's reports cite and the disk lacks | its **Reports** tab → *Fetch missing evidence* (*Preview missing evidence* lists them) | `fetch-windows` (`{"siteId", "dryRun"?}`) | — |
| A site's report evidence media (then its exports) | its **Reports** tab → *Prepare evidence media* | `reports-prepare` (`{"siteId"}`) | `reports prepare <id>` |
| A site's report exports (HTML, PDF, Markdown, evidence pack) | its **Reports** tab → *Export reports* | `reports-export` (`{"siteId", "reportId"?, "formats"?}`) | `reports export <id> [--report <rid>] [--formats html,pdf,md,zip]` |
| A report from a /sweep report, an /ask answer or a report-to-video manifest, and a starter manifest from a report | — | — | `reports convert <sweep\|ask\|manifest> <in> --out <report.json> [--channels-dir <dir>]`, `reports to-manifest <report.json> --out <manifest.json>` |
diff --git a/REPORT.md b/REPORT.md
@@ -2,11 +2,11 @@
<!-- GENERATED by common/bin/file-schemas-docs.ts from the schemas and their *_FIELD_DOCS records — do not edit by hand. -->
-One cited report, format `"archilyzer-report"`, version 1, persisted to `transcripts/sites/<siteId>/reports/<reportId>/report.json` beside its `stills/` and `sources/<sourceId>/`; a relative path in it is relative to that directory. A site's `reports` list in `site.json` is the published, ordered list — see [SITE.md](SITE.md); a report directory it does not name is a draft. The schema is `common/lib/report/schema.ts`; its citations and sources are the citation model's — see [CITATIONS.md](CITATIONS.md).
+One cited report, format `"archilyzer-report"`, version 1, persisted to `transcripts/sites/<siteId>/reports/<reportId>/report.json` beside its `stills/` and `sources/<sourceId>/`; a relative path in it is relative to that directory. A site's `reports` list in `site.json` is the published, ordered list — see [SITE.md](SITE.md); a report directory it does not name is a draft. The schema is `common/lib/report/schema.ts`; its citations and sources are the citation model's — see [CITATIONS.md](CITATIONS.md). A `notes.json` beside it holds the operator's notes on the article (umtool, `umtool notes`; [umtool/docs/notes.md](umtool/docs/notes.md)) and is never published.
A **fact-check** (`"kind": "factcheck"`) is sections (chapters) of claims, each with a verdict and its findings. A **sweep** (`"kind": "sweep"`) is sections with no verdicts, or bodies that cite inline. Markdown fields (`summary`, a section's `body`, a claim's `findings`) cite with `[label](cite:<id>)`.
-`common/lib/report/validate.ts` reports every problem with its JSON path: an unknown key, a reference that names nothing (a listed citation, a claim's source sentence, a `cite:` link, the subject), a section or claim id used twice (they share one namespace: the report page's anchors), a sweep's claim with a verdict, `updated` before `published`, and every citation problem CITATIONS.md lists. Whether a still exists and whether a quote matches its cues are checked when the site is composed.
+`common/lib/report/validate.ts` reports every problem with its JSON path: an unknown key, a reference that names nothing (a listed citation, a claim's source sentence, a `cite:` link, the subject), a section or claim id used twice (they share one namespace: the report page's anchors), a sweep's claim with a verdict, `updated` before `published`, and every citation problem CITATIONS.md lists. Whether a still or the report's video exists, whether the video fits the publish limit, and whether a quote matches its cues are checked when the site is composed.
Regenerate this file with `pnpm --filter yt-dlp-transcript-common exec tsx bin/file-schemas-docs.ts`.
@@ -26,6 +26,7 @@ Regenerate this file with `pnpm --filter yt-dlp-transcript-common exec tsx bin/f
| `published` | no | When the report was published: `YYYY-MM-DD` or an ISO 8601 date-time with a zone. |
| `updated` | no | When it was last changed, in the same form; not before `published`. |
| `subject` | no | The document under review, when the report reviews one: `{ "source": "<id>" }`, an id in `sources`. |
+| `video` | no | A video of the report, shown at the head of its page under the title: `{ "src": "video.mp4", "poster": "poster.jpg", "caption": "…" }`. `src` is an mp4 and `poster` an image (png, jpg or webp), both relative to the report's directory; the caption is one line. Absent = none. |
| `verdicts` | no | Overrides of the shared verdict vocabulary's labels and colours, by verdict (`CORROBORATED`, `PARTLY`, `CONTRADICTED`, `NOT_FOUND`, `UNTESTABLE`): `{ "label": "…", "color": "#rrggbb" }`, each key optional. Absent = the shared defaults. |
| `sources` | no | The documents the report's `source` citations quote, by id — see [CITATIONS.md](CITATIONS.md). Absent = none. |
| `citations` | no | The report's citations, by id — see [CITATIONS.md](CITATIONS.md). A citation is cited from markdown with `[label](cite:<id>)` and listed under the claims that rest on it. Absent = none. |
diff --git a/RUNNING_IN_DOCKER.md b/RUNNING_IN_DOCKER.md
@@ -397,6 +397,10 @@ pnpm ops sync --json '{"slug":"the-quartering"}' --wait
pnpm ops metadata-scan --json '{"slug":"the-quartering"}'
pnpm ops refresh-metadata --json '{"slug":"the-quartering","id":"<videoId>"}' --wait
pnpm ops channel-config --json '{"slug":"the-quartering","patch":{"downloadFilterExclude":"rerun"}}'
+pnpm ops channel-config --json '{"slug":"example-x","sites":[],"excludeFromBuild":true}'
+pnpm ops create-channel --json '{"fields":{"name":"Example (X)","handling":"transcribe","url":"https://x.com/example"}}'
+pnpm ops rename-channel --json '{"slug":"exmaple-x","newSlug":"example-x"}'
+pnpm ops delete-channel --json '{"slug":"example-x","confirm":"example-x"}'
pnpm ops channel-priority --json '{"slugs":["the-quartering"],"operation":"download","tier":"paused"}'
pnpm ops lane --json '{"lane":"download","held":true}'
pnpm ops refresh-report --json '{"all":true}'
@@ -405,6 +409,7 @@ pnpm ops fetch-posts --json '{"slug":"example-x","older":true}' --wait
pnpm ops capture-posts --json '{"slug":"example-x","ids":["1234567890"]}' --wait
pnpm ops persist-videos --json '{"items":[{"slug":"example-channel","id":"abc123"}],"dryRun":true}'
pnpm ops get channel the-quartering
+pnpm ops get channels # every channel, its kind and its sites
pnpm ops list # every action name
```
@@ -420,7 +425,9 @@ Four things to know before you script against it:
`config.json`'s — `downloadFilterInclude` / `downloadFilterExclude` rather than
a `downloadFilter` object. That is what routes them through the form's own
validators, so a bad regex is refused here with the sentence the form shows.
- `""` clears a field, exactly as clearing the input does.
+ `""` clears a field, exactly as clearing the input does. `"sites"` is the
+ form's Sites section — the WHOLE membership set, `[]` for on no site — and
+ `create-channel`'s `"fields"` take the same names.
- **`keep-videos` sets the do-not-clean marker** — the video page's "Do not
clean" toggle, over every video of one channel whose title or description
matches `match`. `match` is matched exactly as a `downloadFilterInclude` is (a
diff --git a/common/bin/archilyzer.ts b/common/bin/archilyzer.ts
@@ -473,7 +473,7 @@ export const COMMANDS: Command[] = [
script(["duplicates"], "duplicate-shorts.ts",
"[--threshold N] [--all-durations] [--blocking title|duration|both] [--near F] [--tolerance N] … on-demand duplicate detection (after index + stats)", 8192),
script(["posts", "fetch"], "fetch-posts.ts",
- "--slug <channel> [--full | --older [--floor YYYY-MM-DD] [--force]] [--limit N] [--pages N] fetch a social channel's posts into its posts corpus (--older: walk back below the oldest archived post; --pages: a forum thread's latest N pages)"),
+ "--slug <channel> [--full | --older [--floor YYYY-MM-DD] [--from YYYY-MM-DD] [--force]] [--limit N] [--pages N] fetch a social channel's posts into its posts corpus (--older: walk back below the oldest archived post; --pages: a forum thread's latest N pages)"),
{
path: ["posts", "import-html"],
usage:
diff --git a/common/bin/fetch-posts.ts b/common/bin/fetch-posts.ts
@@ -2,7 +2,7 @@
// Fetch social posts for one channel into its on-disk posts corpus.
//
// pnpm --filter yt-dlp-transcript-common exec tsx bin/fetch-posts.ts \
-// --slug <channel-slug> [--full | --older [--floor YYYY-MM-DD] [--force]] [--limit N] [--pages N]
+// --slug <channel-slug> [--full | --older [--floor YYYY-MM-DD] [--from YYYY-MM-DD] [--force]] [--limit N] [--pages N]
//
// --pages caps how many pages one run reads, for a source read page by page (a
// forum thread: its latest N pages, newest first; the next run continues where
@@ -14,7 +14,9 @@
//
// --older walks the account's history backwards from the oldest archived post
// (X: search windows), below what the timeline reaches; --floor is the date it
-// stops at. Its position is saved apart from the timeline's, so a later run
+// stops at; --from starts it afresh at a date, replacing its saved position
+// (and a "complete") — for a gap above one surviving old post. Its position
+// is saved apart from the timeline's, so a later run
// continues it and a normal fetch is unaffected. An account that shows no
// posts (nothing archived, and the last timeline fetch read none) is refused
// a search walk unless --force.
@@ -28,7 +30,7 @@ const flags = parseFlags(process.argv.slice(2));
const slug = flags.slug;
if (!slug) {
console.error(
- "Usage: fetch-posts.ts --slug <channel-slug> [--full | --older [--floor YYYY-MM-DD] [--force]] [--limit N] [--pages N]",
+ "Usage: fetch-posts.ts --slug <channel-slug> [--full | --older [--floor YYYY-MM-DD] [--from YYYY-MM-DD] [--force]] [--limit N] [--pages N]",
);
process.exit(2);
}
@@ -52,6 +54,7 @@ fetchPosts({
full: flags.full === "true",
older: flags.older === "true",
floor: flags.floor,
+ from: flags.from,
force: flags.force === "true",
limit,
pages,
diff --git a/common/components/citations/citations.test.ts b/common/components/citations/citations.test.ts
@@ -194,6 +194,9 @@ test("cited markdown: cite links become inline cites, other links stay links, co
assert.match(html, /\[5\]/);
assert.match(html, /<a [^>]*href="https:\/\/example.org\/z" target="_blank" rel="noopener noreferrer"[^>]*>elsewhere<\/a>/);
assert.match(html, /<code[^>]*>\[x\]\(cite:c01\)<\/code>/);
+ // a label holding an editorial [insertion] is still a cite, with its number
+ const bracketed = render(h(CitedMarkdown, { citations: byId, children: "She said [“told [the mayor] so”](cite:c01)." }));
+ assert.match(bracketed, /data-inline-cite="c01"[^]*?>“told \[the mayor\] so”<\/a>[^]*?\[1\]/);
// a link to a place on the same page stays in the tab
const jump = render(h(CitedMarkdown, { citations: {}, children: "[the claims](#claims)" }));
assert.match(jump, /<a [^>]*href="#claims"[^>]*>the claims<\/a>/);
diff --git a/common/controller/fetchOlderPosts.test.ts b/common/controller/fetchOlderPosts.test.ts
@@ -592,3 +592,58 @@ test("a drain already set when a window ends stops before the pause, never waiti
assert.equal((await argvLines()).length - before, 1);
assert.equal((await readPostFetchState(channelRoot))?.older?.since, "2020-09-11");
});
+
+test("from starts a completed walk afresh above a sparse old post, and covers the gap", async () => {
+ const slug = "older-from";
+ const channelRoot = await makeChannel(slug);
+ // Only the OLDEST post survives in the archive (a timeline that returned
+ // one 2019 post), and the walk below it is already complete.
+ const oldest = POST_DATES[POST_DATES.length - 1];
+ const post = normalizeXTweet(
+ { tweet_id: idOf(oldest), date: oldest, content: "old", author: { name: HANDLE }, user: { name: HANDLE } },
+ slug,
+ );
+ assert.ok(post);
+ await writePosts(channelRoot, [post]);
+ await writePostFetchState(channelRoot, {
+ lastFetchedAt: "2026-01-01T00:00:00.000Z",
+ cursor: "1/TIMELINE-RESUME",
+ older: { since: "2019-08-20", until: "2019-11-21", emptyWindows: 4, complete: true, completeReason: "done" },
+ });
+
+ // Without a start date there is nothing to walk.
+ const idle = await fetchPosts({ paths: getPaths(), slug, settings: LOGIN, older: true, olderWindowPauseMs: 0 });
+ assert.equal(idle.written, 0);
+
+ const before = (await argvLines()).length;
+ const log: string[] = [];
+ const result = await fetchPosts({
+ paths: getPaths(),
+ slug,
+ settings: LOGIN,
+ older: true,
+ from: "2021-03-10",
+ olderWindowPauseMs: 0,
+ onLog: (l) => log.push(l),
+ });
+ assert.equal(result.ok, true, log.join("\n"));
+ assert.equal(result.written, 4);
+ assert.equal((await readSeenPostIds(channelRoot)).size, 5);
+ const queries = (await argvLines()).slice(before).map(queryOf);
+ assert.equal(queries[0], `from:${HANDLE} since:2020-12-11 until:2021-03-11 include:nativeretweets`);
+ assert.ok(log.some((l) => /starting afresh from 2021-03-10/.test(l)), log.join("\n"));
+ const state = await readPostFetchState(channelRoot);
+ assert.equal(state?.older?.complete, true);
+ assert.equal(state?.cursor, "1/TIMELINE-RESUME");
+});
+
+test("from is refused without older, and when it is not a date", async () => {
+ const slug = "older-from-refusals";
+ await makeChannel(slug);
+ const a = await fetchPosts({ paths: getPaths(), slug, settings: LOGIN, from: "2021-01-01" });
+ assert.equal(a.ok, false);
+ assert.match(a.error ?? "", /only to an older-posts fetch/);
+ const b = await fetchPosts({ paths: getPaths(), slug, settings: LOGIN, older: true, from: "2021-13-01" });
+ assert.equal(b.ok, false);
+ assert.match(b.error ?? "", /not a date/);
+});
diff --git a/common/controller/fetchPosts.ts b/common/controller/fetchPosts.ts
@@ -78,6 +78,13 @@ export type FetchPostsOptions = {
// later runs of the walk. Default: none (it stops after a run of empty
// windows, or at the account's creation date).
floor?: string;
+ // Start the older walk afresh, backwards from this date (YYYY-MM-DD), in
+ // place of its saved position — and of a "complete" it reached. The walk
+ // otherwise starts below the OLDEST archived post, so one surviving old post
+ // (a timeline that returns a handful of 2018 posts beside its 3,200 recent
+ // ones) puts years of history below a gap no window ever covers. Only with
+ // `older`.
+ from?: string;
// The older walk's pause between two windows. Default: the fetcher's own.
olderWindowPauseMs?: number;
// Walk older posts even on an account that shows none (see
@@ -96,6 +103,8 @@ export type FetchPostsOptions = {
// The refusal for both walks at once, shared with the server action so the
// button, the ops route and this controller say the same thing.
+export const OLDER_FROM_REFUSAL = "A start date (\"from\") applies only to an older-posts fetch.";
+
export const FULL_AND_OLDER_REFUSAL =
"A full re-fetch and an older-posts fetch are different walks — run one at a time.";
@@ -149,6 +158,14 @@ export async function fetchPosts(
if (opts.full && opts.older) {
return { ok: false, written: 0, skipped: 0, complete: false, error: FULL_AND_OLDER_REFUSAL };
}
+ for (const [name, day] of [["from", opts.from]] as const) {
+ if (day !== undefined && !isUtcDay(day)) {
+ return { ok: false, written: 0, skipped: 0, complete: false, error: `"${day}" is not a date (YYYY-MM-DD) for "${name}"` };
+ }
+ }
+ if (opts.from !== undefined && !opts.older) {
+ return { ok: false, written: 0, skipped: 0, complete: false, error: OLDER_FROM_REFUSAL };
+ }
if (opts.floor !== undefined && !isUtcDay(opts.floor)) {
return {
ok: false,
@@ -399,7 +416,8 @@ async function fetchOlderPosts(ctx: {
const { slug } = opts;
const log = (line: string) => opts.onLog?.(line);
const fetchOlder = fetcher.fetchOlder!;
- const prior = priorState?.older;
+ // A start date replaces the saved walk, complete or not.
+ const prior = opts.from === undefined ? priorState?.older : undefined;
if (prior?.complete) {
log(
@@ -418,8 +436,8 @@ async function fetchOlderPosts(ctx: {
}
if (empty) log("[warn] The account shows no posts; walking anyway, as forced.");
- const floor = opts.floor ?? prior?.floor;
- let accountCreatedAt = prior?.accountCreatedAt;
+ const floor = opts.floor ?? priorState?.older?.floor;
+ let accountCreatedAt = priorState?.older?.accountCreatedAt;
const startedAt = new Date().toISOString();
let position: OlderBackfillPosition;
if (prior) {
@@ -430,7 +448,7 @@ async function fetchOlderPosts(ctx: {
...(prior.maxId ? { maxId: prior.maxId } : {}),
};
} else {
- const oldest = await oldestPostCreatedAt(channelRoot);
+ const oldest = opts.from ?? (await oldestPostCreatedAt(channelRoot));
position = firstOlderWindow({
oldestArchivedAt: oldest ?? undefined,
floor: olderFloorDay(floor, accountCreatedAt),
@@ -439,7 +457,7 @@ async function fetchOlderPosts(ctx: {
log(
`Fetching older posts for ${slug} via ${fetcher.label} (@${handle}): ` +
- (prior ? "resuming at " : "starting at ") +
+ (prior ? "resuming at " : opts.from ? `starting afresh from ${opts.from} at ` : "starting at ") +
`${position.since} – ${position.until}` +
(floor ? `, down to ${floor}` : "") +
`; ${ctx.seenIds.size} already archived.`,
diff --git a/common/controller/fetchWindows.test.ts b/common/controller/fetchWindows.test.ts
@@ -0,0 +1,276 @@
+// fetchWindows (the batch) against a fake yt-dlp, through the real
+// fetchWindowManaged: what the platform answers is decided by a marker in the
+// item's URL, and every shared-state dependency — the cooldown, the hold, the
+// backoff, the clean record, the pause — is injected and recorded.
+//
+// Run with: pnpm --filter yt-dlp-transcript-common exec tsx --test controller/fetchWindows.test.ts
+
+import { test } from "node:test";
+import assert from "node:assert/strict";
+import { chmod, mkdir, mkdtemp, readFile, writeFile } from "node:fs/promises";
+import os from "node:os";
+import path from "node:path";
+import {
+ CLIP_WINDOW_MIN_GAP_SECONDS,
+ CLIP_WINDOW_PLATFORM_MIN_GAP_SECONDS,
+ fetchWindows,
+ type FetchWindowsDeps,
+ type FetchWindowsItem,
+} from "./fetchWindows";
+import type { Paths } from "../lib/paths";
+import type { ChannelConfig } from "../lib/channelConfig";
+import type { SiteSettings } from "../lib/settings";
+import type { JobProgress } from "../jobs/registry";
+
+const ROOT = await mkdtemp(path.join(os.tmpdir(), "fetchwindows-"));
+const BIN = path.join(ROOT, "fake-ytdlp.mjs");
+const ARGS_LOG = path.join(ROOT, "argv.jsonl");
+process.env.FAKE_ARGS_LOG = ARGS_LOG;
+
+// The URL decides: `/403` a Cloudflare refusal, `/429` a rate limit, `/gone` a
+// removed video; anything else writes the window.
+await writeFile(
+ BIN,
+ `#!/usr/bin/env node
+import { appendFileSync, writeFileSync } from "node:fs";
+const args = process.argv.slice(2);
+appendFileSync(process.env.FAKE_ARGS_LOG, JSON.stringify(args) + "\\n");
+const url = args[args.length - 1];
+const fail = (msg) => { process.stderr.write(msg + "\\n"); process.exit(1); };
+if (url.includes("/403")) fail("ERROR: [download] Got error: HTTP Error 403: Forbidden");
+if (url.includes("/429")) fail("ERROR: [youtube] x: HTTP Error 429: Too Many Requests");
+if (url.includes("/gone")) fail("ERROR: [youtube] x: Video unavailable. This video has been removed by the uploader");
+writeFileSync(args[args.indexOf("-o") + 1], "mp4");
+`,
+);
+await chmod(BIN, 0o755);
+
+let run = 0;
+// A fresh corpus per test, so one test's fetched windows are not another's cache.
+async function corpus(): Promise<Paths> {
+ const channelsDir = path.join(ROOT, `corpus-${++run}`, "channels");
+ await mkdir(channelsDir, { recursive: true });
+ return { ytdlpBin: BIN, channelsDir } as unknown as Paths;
+}
+
+async function spawns(): Promise<number> {
+ return (await readFile(ARGS_LOG, "utf8").catch(() => "")).trim().split("\n").filter(Boolean).length;
+}
+
+const CONFIG = { url: "https://www.youtube.com/@x", platform: "youtube" } as unknown as ChannelConfig;
+
+type Recorded = {
+ sleeps: number[];
+ backoffs: [string, string][];
+ cleans: string[];
+ progress: JobProgress[];
+};
+
+function harness(over: Partial<FetchWindowsDeps> = {}) {
+ const rec: Recorded = { sleeps: [], backoffs: [], cleans: [], progress: [] };
+ const deps: Partial<FetchWindowsDeps> = {
+ sleep: async (ms) => {
+ rec.sleeps.push(ms);
+ },
+ getSettings: () => ({ sleepBetweenDownloadsSeconds: 0 }) as unknown as SiteSettings,
+ readChannelConfig: async (_p, slug) => (slug === "c" ? CONFIG : null),
+ findVideoSourceUrl: async () => null,
+ assertTextReadable: async () => undefined,
+ cooldownRemainingMs: async () => 0,
+ heldRefusal: async () => null,
+ recordBackoff: async (platform, _p, cls) => {
+ rec.backoffs.push([platform, cls]);
+ },
+ recordClean: async (platform) => {
+ rec.cleans.push(platform);
+ return null;
+ },
+ gapRemainingMs: () => 0,
+ noteGap: () => {},
+ ...over,
+ };
+ return { rec, deps };
+}
+
+const item = (id: string, url: string, from = 10, to = 20): FetchWindowsItem => ({
+ slug: "c",
+ id,
+ from,
+ to,
+ webpageUrl: `https://www.youtube.com${url}`,
+});
+
+async function go(paths: Paths, items: FetchWindowsItem[], h: ReturnType<typeof harness>, extra: { gapMs?: number; signal?: AbortSignal; drainSignal?: AbortSignal } = {}) {
+ return fetchWindows({
+ paths,
+ items,
+ provenance: { requestedBy: "test" },
+ gapMs: extra.gapMs ?? 1000,
+ signal: extra.signal,
+ drainSignal: extra.drainSignal,
+ setProgress: (p) => h.rec.progress.push(p),
+ deps: h.deps,
+ });
+}
+
+test("a cached window costs no request and no pause", async () => {
+ const paths = await corpus();
+ const clips = path.join(paths.channelsDir, "c", "data", "v1", "clips");
+ await mkdir(clips, { recursive: true });
+ await writeFile(path.join(clips, "0.00-60.00.mp4"), "mp4");
+ const h = harness();
+ const before = await spawns();
+ const r = await go(paths, [item("v1", "/a", 5, 15), item("v1", "/b", 30, 40), item("v2", "/c")], h);
+ assert.equal(r.cached.length, 2);
+ assert.equal(r.fetched.length, 1);
+ assert.equal((await spawns()) - before, 1, "only the uncached window spawned");
+ assert.deepEqual(h.rec.sleeps, [], "one network fetch owes no pause");
+ assert.deepEqual(h.rec.cleans, ["youtube"]);
+ assert.deepEqual(h.rec.progress.at(-1), { metric: "clips", initial: 0, target: 3, current: 3 });
+});
+
+test("one 403 is an item failure; the next window is fetched after the gap", async () => {
+ const h = harness();
+ const r = await go(await corpus(), [item("v1", "/403"), item("v2", "/ok")], h);
+ assert.equal(r.stopped, undefined);
+ assert.equal(r.failed.length, 1);
+ assert.equal(r.failed[0].class, "network");
+ assert.equal(r.fetched.length, 1);
+ assert.deepEqual(h.rec.sleeps, [1000]);
+ assert.deepEqual(h.rec.backoffs, [], "one 403 does not back the platform off");
+});
+
+test("two 403s in a row back the platform off and stop the run", async () => {
+ const h = harness();
+ const r = await go(await corpus(), [item("v1", "/403"), item("v2", "/403"), item("v3", "/ok"), item("v4", "/ok")], h);
+ assert.equal(r.stopped, "network");
+ assert.equal(r.failed.length, 2);
+ assert.deepEqual(r.notAttempted.map((i) => i.id), ["v3", "v4"]);
+ assert.deepEqual(h.rec.backoffs, [["youtube", "network"]]);
+});
+
+test("a per-video failure between two 403s breaks the streak", async () => {
+ const h = harness();
+ const r = await go(await corpus(), [item("v1", "/403"), item("v2", "/gone"), item("v3", "/403"), item("v4", "/ok")], h);
+ assert.equal(r.stopped, undefined);
+ assert.deepEqual(r.failed.map((f) => f.class), ["network", "per_video", "network"]);
+ assert.equal(r.fetched.length, 1);
+ assert.deepEqual(h.rec.backoffs, []);
+});
+
+test("a 429 records the backoff and stops at once", async () => {
+ const h = harness();
+ const r = await go(await corpus(), [item("v1", "/ok"), item("v2", "/429"), item("v3", "/ok")], h);
+ assert.equal(r.stopped, "rate-limit");
+ assert.equal(r.fetched.length, 1);
+ assert.deepEqual(r.notAttempted.map((i) => i.id), ["v3"]);
+ assert.deepEqual(h.rec.backoffs, [["youtube", "rate_limit"]]);
+});
+
+test("a platform cooling down (or held) ends the run before its next fetch", async () => {
+ let calls = 0;
+ const h = harness({ cooldownRemainingMs: async () => (calls++ === 0 ? 0 : 60_000) });
+ const before = await spawns();
+ const r = await go(await corpus(), [item("v1", "/ok"), item("v2", "/ok"), item("v3", "/ok")], h);
+ assert.equal(r.stopped, "cooldown");
+ assert.equal(r.fetched.length, 1);
+ assert.deepEqual(r.notAttempted.map((i) => i.id), ["v2", "v3"]);
+ assert.equal((await spawns()) - before, 1);
+
+ const held = harness({ heldRefusal: async () => "youtube is held." });
+ const r2 = await go(await corpus(), [item("v1", "/ok")], held);
+ assert.equal(r2.stopped, "held");
+ assert.deepEqual(r2.notAttempted.map((i) => i.id), ["v1"]);
+});
+
+test("a drain stops between windows; the one in flight finishes", async () => {
+ const drain = new AbortController();
+ const h = harness({
+ recordClean: async () => {
+ drain.abort();
+ return null;
+ },
+ });
+ const r = await go(await corpus(), [item("v1", "/ok"), item("v2", "/ok")], h, { drainSignal: drain.signal });
+ assert.equal(r.stopped, "drain");
+ assert.equal(r.fetched.length, 1);
+ assert.deepEqual(r.notAttempted.map((i) => i.id), ["v2"]);
+});
+
+test("the default gap is the platform's, floored at the clip-window minimum and jittered", async () => {
+ const h = harness();
+ const r = await fetchWindows({
+ paths: await corpus(),
+ items: [item("v1", "/ok"), item("v2", "/ok"), item("v3", "/ok")],
+ provenance: { requestedBy: "test" },
+ deps: h.deps,
+ });
+ assert.equal(r.fetched.length, 3);
+ assert.equal(h.rec.sleeps.length, 2);
+ for (const ms of h.rec.sleeps) {
+ assert.ok(ms >= CLIP_WINDOW_MIN_GAP_SECONDS * 1000, `${ms} is at least the floor`);
+ assert.ok(ms <= CLIP_WINDOW_MIN_GAP_SECONDS * 1500 + 60_000, `${ms} is at most the floor plus half`);
+ }
+});
+
+test("an unknown channel, an unreadable one and a missing URL fail their items and the run goes on", async () => {
+ const h = harness({
+ readChannelConfig: async (_p, slug) => (slug === "nope" ? null : CONFIG),
+ assertTextReadable: async (_p, slug) => {
+ if (slug === "off") throw new Error("off/data is not a readable directory");
+ },
+ });
+ const r = await go(
+ await corpus(),
+ [
+ { ...item("v1", "/ok"), slug: "nope" },
+ { ...item("v1", "/ok"), slug: "off" },
+ { slug: "c", id: "v9", from: 1, to: 2 },
+ item("v2", "/ok"),
+ ],
+ h,
+ );
+ assert.deepEqual(r.failed.map((f) => f.class), ["unknown-channel", "unreachable", "no-url"]);
+ assert.equal(r.fetched.length, 1);
+ assert.deepEqual(h.rec.sleeps, [], "no network fetch preceded the one that ran");
+});
+
+test("a duplicated window is fetched once", async () => {
+ const h = harness();
+ const before = await spawns();
+ const r = await go(await corpus(), [item("v1", "/ok"), item("v1", "/ok")], h);
+ assert.equal(r.fetched.length, 1);
+ assert.equal((await spawns()) - before, 1);
+});
+
+test("rumble windows are further apart, whatever platform their channel names", async () => {
+ const h = harness();
+ const rumble = (id: string): FetchWindowsItem => ({ slug: "c", id, from: 10, to: 20, webpageUrl: `https://rumble.com/${id}-x.html` });
+ const r = await fetchWindows({
+ paths: await corpus(),
+ items: [rumble("v1"), rumble("v2")],
+ provenance: { requestedBy: "test" },
+ deps: h.deps,
+ });
+ assert.equal(r.fetched.length, 2);
+ assert.equal(h.rec.sleeps.length, 1);
+ const floor = CLIP_WINDOW_PLATFORM_MIN_GAP_SECONDS.rumble * 1000;
+ assert.ok(h.rec.sleeps[0] >= floor, `${h.rec.sleeps[0]} is at least rumble's ${floor}`);
+ assert.ok(h.rec.sleeps[0] <= floor * 1.5 + 60_000);
+});
+
+test("a second run on the same platform waits out the gap the first one set", async () => {
+ const next = new Map<string, number>();
+ const shared = {
+ gapRemainingMs: (key: string) => next.get(key) ?? 0,
+ noteGap: (key: string, ms: number) => void next.set(key, ms),
+ };
+ const first = harness(shared);
+ await go(await corpus(), [item("v1", "/ok"), item("v2", "/403")], first);
+ assert.deepEqual(first.rec.sleeps, [1000], "a run's own first fetch owes nothing");
+ assert.equal(next.get("clip-window:youtube"), 1000, "a refused fetch sets the gap too");
+
+ const second = harness(shared);
+ await go(await corpus(), [item("v3", "/ok"), item("v4", "/ok")], second);
+ assert.deepEqual(second.rec.sleeps, [1000, 1000], "the first fetch waits for the earlier run's gap");
+});
diff --git a/common/controller/fetchWindows.ts b/common/controller/fetchWindows.ts
@@ -0,0 +1,501 @@
+import path from "node:path";
+import type { Paths } from "../lib/paths";
+import type { ChannelConfig } from "../lib/channelConfig";
+import type { SiteSettings } from "../lib/settings";
+import { getSettings } from "../lib/settings";
+import type { DownloadFailureClass } from "../lib/availability";
+import { resolveCookiePolicy } from "../lib/cookiePolicy";
+import { assertChannelTextReadable } from "../lib/channelMedia";
+import { detectPlatform } from "../lib/platform";
+import { findContainingClipWindow } from "../lib/clipWindow-server";
+import type { JobProgress } from "../jobs/registry";
+import { downloadGapMs } from "../jobs/platformBackoff";
+import { notePlatformGap, platformGapRemainingMs } from "../jobs/platformGap";
+import {
+ heldPlatformRefusal,
+ platformCooldownRemainingMs,
+ recordDownloadBackoff,
+ recordPlatformClean,
+} from "../jobs/downloadBackoff";
+import {
+ FetchWindowError,
+ fetchWindowManaged,
+ type FetchWindowOpts,
+ type FetchWindowResult,
+} from "../ytdlp/fetchWindowManaged";
+import {
+ channelPaceSeconds,
+ channelPlatform,
+ configForVideoUrl,
+} from "../ytdlp/channelArgs";
+import {
+ platformMinGapSeconds,
+ staticSleepRequestsSeconds,
+} from "../ytdlp/platformArgs.mjs";
+import { readChannelConfig } from "./channels";
+import { findVideoSourceUrl } from "./undownloadedVideos";
+
+// FETCH A LIST OF CLIP WINDOWS, ONE PLATFORM'S, AS ONE PACED JOB.
+//
+// The batch form of the single `fetch-window` job. Evidence clips for a report
+// or a video used to be fetched one request at a time by hand-written shell
+// loops that owned the pacing and the "stop after two failures" rule; this is
+// that loop, inside the editor, where the platform's cooldown and pace live.
+//
+// EACH ITEM GOES THROUGH fetchWindowManaged, unchanged: the channel's cookie
+// policy and extra args, the auth retry, the HLS retry, the provenance sidecar.
+// What this adds is the space between items:
+//
+// - A CACHED WINDOW costs nothing: no request, no pause. The cache is asked
+// here first (the same containing-window rule fetchWindowManaged applies),
+// so a re-run of a half-done list walks straight to what is missing.
+// - BETWEEN TWO NETWORK FETCHES the batch downloads' own gap
+// (downloadGapMs), with a floor of CLIP_WINDOW_MIN_GAP_SECONDS — jittered
+// up to half again, so a platform is never asked on a fixed beat.
+// - BEFORE EACH NETWORK FETCH the platform's cooldown and hold. Either one
+// ends the run; the rest is left for the next.
+//
+// THE FAILURE RULE. The class comes from FetchWindowError, classified once in
+// fetchWindowManaged:
+// rate_limit (429, bot check, soft block) — the backoff is already
+// recorded; stop now.
+// network (a 403 above all: Rumble's Cloudflare, googlevideo refusing)
+// — one is an item failure, because a removed Rumble page
+// answers 403 too. A SECOND IN A ROW records a network backoff
+// and stops. Any other outcome between them breaks the streak.
+// anything else (unavailable, private, a cut that failed) — the item fails
+// and the run carries on.
+// A success records the platform clean (recordPlatformClean), which is how an
+// earlier backoff eases.
+//
+// RESUMABLE BY RE-RUNNING, as persistVideos: nothing is remembered between
+// runs, and a re-run's cache check skips every window the last one fetched.
+//
+// A JOB OF THIS KIND SPANS CHANNELS, so runManagedFunction's per-channel text
+// guard has no channel to ask. It is asked here instead, once per channel: an
+// item on a channel whose text is not readable fails as `unreachable` (its
+// clips/ directory lives in that text), and the rest carry on.
+
+// The least a batch waits between two clip-window fetches on one platform.
+// A window is a short request, which is exactly why a loop of them looks like
+// a scraper: 20 s apart was enough for YouTube to answer 403 (2026-10-07).
+export const CLIP_WINDOW_MIN_GAP_SECONDS = 30;
+
+// A platform whose windows must be further apart than that. Rumble's
+// Cloudflare puts the whole IP behind a JS challenge ("Just a moment…", 403 to
+// every rumble.com request, impersonated or not) after a handful of requests
+// in a few minutes, and a window costs three (page, embed JSON, HLS manifest):
+// on 2026-10-07 windows 30–45 s apart drew it within four, and it lifted again
+// in about five minutes of quiet.
+export const CLIP_WINDOW_PLATFORM_MIN_GAP_SECONDS: Readonly<Record<string, number>> =
+ Object.freeze({ rumble: 120 });
+
+export type FetchWindowsItem = {
+ slug: string;
+ id: string;
+ from: number;
+ to: number;
+ // `<reportId>#<citationId>`, or a manifest's clip id.
+ clipId?: string;
+ reason?: string;
+ pad?: number;
+ // The page URL when the caller has it; else resolved as the single fetch does.
+ webpageUrl?: string;
+};
+
+export type FetchWindowsFailureClass =
+ | DownloadFailureClass
+ // No channel config for the slug.
+ | "unknown-channel"
+ // The channel's text (and so its clips/) is not readable.
+ | "unreachable"
+ // No page URL to fetch from.
+ | "no-url";
+
+export type FetchWindowsFailure = {
+ item: FetchWindowsItem;
+ class: FetchWindowsFailureClass;
+ message: string;
+};
+
+export type FetchWindowsStop =
+ | "rate-limit"
+ | "network"
+ | "cooldown"
+ | "held"
+ | "drain"
+ | "cancel";
+
+export type FetchWindowsResult = {
+ fetched: FetchWindowsItem[];
+ cached: FetchWindowsItem[];
+ failed: FetchWindowsFailure[];
+ // Due a fetch but not attempted because the run stopped first.
+ notAttempted: FetchWindowsItem[];
+ stopped?: FetchWindowsStop;
+};
+
+export type FetchWindowsProvenance = {
+ requestedBy: string;
+ manifest?: string;
+ requestedAt?: string;
+};
+
+// Everything that touches the network, the disk's shared state or the clock.
+// Injectable so the tests decide what a platform answers and how long a pause
+// is.
+export type FetchWindowsDeps = {
+ fetchWindow: (opts: FetchWindowOpts) => Promise<FetchWindowResult>;
+ sleep: (ms: number, signal?: AbortSignal) => Promise<void>;
+ getSettings: () => SiteSettings;
+ readChannelConfig: (paths: Paths, slug: string) => Promise<ChannelConfig | null>;
+ findVideoSourceUrl: (
+ paths: Paths,
+ slug: string,
+ id: string,
+ config: ChannelConfig,
+ ) => Promise<string | null>;
+ assertTextReadable: (paths: Paths, slug: string) => Promise<unknown>;
+ cooldownRemainingMs: (platform: string, paths: Paths) => Promise<number>;
+ heldRefusal: (platform: string, paths: Paths) => Promise<string | null>;
+ recordBackoff: (
+ platform: string,
+ paths: Paths,
+ failureClass: "rate_limit" | "network",
+ ) => Promise<void>;
+ recordClean: (platform: string, paths: Paths) => Promise<string | null>;
+ // The gap ACROSS runs (jobs/platformGap.ts, keyed `clip-window:<platform>`):
+ // two batches back to back on one queue — a umtool manifest each — must not
+ // put the second one's first fetch right after the first one's last.
+ gapRemainingMs: (key: string) => number;
+ noteGap: (key: string, gapMs: number) => void;
+};
+
+function abortableSleep(ms: number, signal?: AbortSignal): Promise<void> {
+ if (ms <= 0 || signal?.aborted) return Promise.resolve();
+ return new Promise((resolve) => {
+ const onAbort = () => {
+ clearTimeout(t);
+ resolve();
+ };
+ const t = setTimeout(() => {
+ signal?.removeEventListener("abort", onAbort);
+ resolve();
+ }, ms);
+ signal?.addEventListener("abort", onAbort, { once: true });
+ });
+}
+
+const DEFAULT_DEPS: FetchWindowsDeps = {
+ fetchWindow: fetchWindowManaged,
+ sleep: abortableSleep,
+ getSettings,
+ readChannelConfig,
+ findVideoSourceUrl: (paths, slug, id, config) =>
+ findVideoSourceUrl(paths, slug, id, config),
+ assertTextReadable: (paths, slug) => assertChannelTextReadable(paths, slug),
+ cooldownRemainingMs: (platform, paths) =>
+ platformCooldownRemainingMs(platform, paths),
+ heldRefusal: (platform, paths) =>
+ heldPlatformRefusal(platform, "This batch", paths),
+ recordBackoff: (platform, paths, failureClass) =>
+ recordDownloadBackoff(platform, paths, failureClass),
+ recordClean: (platform, paths) => recordPlatformClean(platform, paths),
+ gapRemainingMs: (key) => platformGapRemainingMs(key),
+ noteGap: (key, gapMs) => notePlatformGap(key, gapMs),
+};
+
+export function fetchWindowsItemLabel(item: FetchWindowsItem): string {
+ const span = `${item.from.toFixed(2)}–${item.to.toFixed(2)}`;
+ return `${item.slug}/${item.id} ${span}${item.clipId ? ` (${item.clipId})` : ""}`;
+}
+
+// Duplicates collapse: one window asked for twice is fetched once, and the
+// first ask's clipId and reason are the ones recorded.
+export function dedupeFetchWindowsItems(
+ items: readonly FetchWindowsItem[],
+): FetchWindowsItem[] {
+ const seen = new Set<string>();
+ const out: FetchWindowsItem[] = [];
+ for (const item of items) {
+ const key = `${item.slug}\0${item.id}\0${item.from.toFixed(2)}\0${item.to.toFixed(2)}`;
+ if (seen.has(key)) continue;
+ seen.add(key);
+ out.push(item);
+ }
+ return out;
+}
+
+export async function fetchWindows({
+ paths,
+ items,
+ provenance,
+ maxHeight,
+ gapMs,
+ onLog,
+ signal,
+ drainSignal,
+ setProgress,
+ deps: depsOverride,
+}: {
+ paths: Paths;
+ items: FetchWindowsItem[];
+ provenance: FetchWindowsProvenance;
+ maxHeight?: number;
+ // The pause between two network fetches. Absent = the platform's batch gap,
+ // floored at CLIP_WINDOW_MIN_GAP_SECONDS.
+ gapMs?: number;
+ onLog?: (line: string) => void;
+ signal?: AbortSignal;
+ drainSignal?: AbortSignal;
+ setProgress?: (snap: JobProgress) => void;
+ deps?: Partial<FetchWindowsDeps>;
+}): Promise<FetchWindowsResult> {
+ const deps: FetchWindowsDeps = { ...DEFAULT_DEPS, ...depsOverride };
+ const log = (line: string) => onLog?.(line.endsWith("\n") ? line : `${line}\n`);
+ const fetchLog = (line: string) => onLog?.(line);
+ const fetchSignal = signal ?? new AbortController().signal;
+ const settings = deps.getSettings();
+ const requestedAt = provenance.requestedAt ?? new Date().toISOString();
+
+ const work = dedupeFetchWindowsItems(items);
+ const result: FetchWindowsResult = {
+ fetched: [],
+ cached: [],
+ failed: [],
+ notAttempted: [],
+ };
+ log(
+ `Fetch windows: ${work.length} window(s) for ${provenance.requestedBy}` +
+ `${provenance.manifest ? ` · ${provenance.manifest}` : ""}.`,
+ );
+
+ let done = 0;
+ const progress = () =>
+ setProgress?.({ metric: "clips", initial: 0, target: work.length, current: done });
+ progress();
+
+ const configs = new Map<string, ChannelConfig | null>();
+ const configOf = async (slug: string) => {
+ if (!configs.has(slug)) configs.set(slug, await deps.readChannelConfig(paths, slug));
+ return configs.get(slug) ?? null;
+ };
+ // A channel's text readability, asked once per channel per run.
+ const readable = new Map<string, string | null>();
+ const textProblem = async (slug: string): Promise<string | null> => {
+ if (!readable.has(slug)) {
+ readable.set(
+ slug,
+ await deps.assertTextReadable(paths, slug).then(
+ () => null,
+ (e: unknown) => (e as Error).message,
+ ),
+ );
+ }
+ return readable.get(slug) ?? null;
+ };
+
+ const fail = (
+ item: FetchWindowsItem,
+ cls: FetchWindowsFailureClass,
+ message: string,
+ ) => {
+ result.failed.push({ item, class: cls, message });
+ log(` ✗ ${fetchWindowsItemLabel(item)}: ${cls} — ${message}`);
+ };
+ const stop = (why: FetchWindowsStop, from: number) => {
+ result.stopped = why;
+ result.notAttempted.push(...work.slice(from));
+ };
+ const interrupted = (): FetchWindowsStop | null =>
+ signal?.aborted ? "cancel" : drainSignal?.aborted ? "drain" : null;
+
+ // Network attempts so far (the gap is owed before every one after the
+ // first), and the run of consecutive network-class failures.
+ let networkAttempts = 0;
+ let networkStreak = 0;
+
+ for (let i = 0; i < work.length; i++) {
+ const item = work[i];
+ const why = interrupted();
+ if (why) {
+ stop(why, i);
+ break;
+ }
+ const label = fetchWindowsItemLabel(item);
+
+ const config = await configOf(item.slug);
+ if (!config) {
+ fail(item, "unknown-channel", `no channel "${item.slug}"`);
+ done += 1;
+ progress();
+ continue;
+ }
+ const unreadable = await textProblem(item.slug);
+ if (unreadable) {
+ fail(item, "unreachable", unreadable);
+ done += 1;
+ progress();
+ continue;
+ }
+ const videoDir = path.join(paths.channelsDir, item.slug, "data", item.id);
+
+ // THE CACHE FIRST, here rather than only inside fetchWindowManaged, so a
+ // cached window owes no pause and is not counted as a network attempt.
+ const hit = await findContainingClipWindow(videoDir, item.from, item.to);
+ if (hit) {
+ result.cached.push(item);
+ log(` = ${label}: cached in clips/${hit.file}`);
+ done += 1;
+ progress();
+ continue;
+ }
+
+ const url =
+ item.webpageUrl?.trim() ||
+ (await deps.findVideoSourceUrl(paths, item.slug, item.id, config));
+ if (!url) {
+ fail(
+ item,
+ "no-url",
+ "no metadata.info.json and the playlist does not contain a matching entry",
+ );
+ done += 1;
+ progress();
+ continue;
+ }
+ // The cooldown's key, as the single fetch and every download path use it.
+ const platform = detectPlatform(url) ?? "unknown";
+
+ // Paced as the platform the URL points at, which a channel with no
+ // platform of its own (community-notes holds Rumble videos) does not name.
+ const paced = configForVideoUrl(config, url);
+ const gap =
+ gapMs ??
+ downloadGapMs(
+ config.sleepBetweenDownloadsSeconds ?? settings.sleepBetweenDownloadsSeconds,
+ channelPaceSeconds(paced),
+ staticSleepRequestsSeconds(channelPlatform(paced)),
+ {
+ minSeconds: Math.max(
+ CLIP_WINDOW_MIN_GAP_SECONDS,
+ platformMinGapSeconds(channelPlatform(paced)),
+ CLIP_WINDOW_PLATFORM_MIN_GAP_SECONDS[platform] ?? 0,
+ ),
+ },
+ );
+ const gapKey = `clip-window:${platform}`;
+ // Between this run's own fetches, the gap; before its first, whatever is
+ // left of the gap the last run on this platform set.
+ const wait = networkAttempts > 0 ? gap : deps.gapRemainingMs(gapKey);
+ if (wait > 0) {
+ log(
+ networkAttempts > 0
+ ? `Sleeping ${Math.round(wait / 1000)}s before the next fetch...`
+ : `${platform} was asked for a window by an earlier run: waiting ${Math.round(wait / 1000)}s.`,
+ );
+ await deps.sleep(wait, signal);
+ const after = interrupted();
+ if (after) {
+ stop(after, i);
+ break;
+ }
+ }
+
+ // ASKED BEFORE EVERY NETWORK FETCH: another job (or the lane) may have
+ // backed this platform off while this one slept.
+ const held = await deps.heldRefusal(platform, paths);
+ if (held) {
+ log(` ${label}: ${held} — stopping; the rest are left for a later run.`);
+ stop("held", i);
+ break;
+ }
+ const cooldownMs = await deps.cooldownRemainingMs(platform, paths);
+ if (cooldownMs > 0) {
+ log(
+ ` ${label}: ${platform} is in a rate-limit cooldown ` +
+ `(${Math.ceil(cooldownMs / 1000)}s remaining) — stopping; the rest are ` +
+ `left for a later run.`,
+ );
+ stop("cooldown", i);
+ break;
+ }
+
+ log(` ${label}: fetching…`);
+ networkAttempts += 1;
+ try {
+ const r = await deps.fetchWindow({
+ channelSlug: item.slug,
+ channelConfig: config,
+ paths,
+ videoDir,
+ videoId: item.id,
+ videoUrl: url,
+ cwd: path.join(paths.channelsDir, item.slug),
+ from: item.from,
+ to: item.to,
+ provenance: {
+ requestedBy: provenance.requestedBy,
+ manifest: provenance.manifest,
+ clipId: item.clipId,
+ reason: item.reason,
+ pad: item.pad,
+ requestedAt,
+ },
+ cookiePolicy: resolveCookiePolicy(settings, config),
+ maxHeight,
+ onLog: fetchLog,
+ signal: fetchSignal,
+ onPlatformBackoff: () => deps.recordBackoff(platform, paths, "rate_limit"),
+ });
+ networkStreak = 0;
+ (r.cached ? result.cached : result.fetched).push(item);
+ const clean = await deps.recordClean(platform, paths).catch(() => null);
+ if (clean) log(clean);
+ } catch (e) {
+ const message = (e as Error).message;
+ const cls: FetchWindowsFailureClass =
+ e instanceof FetchWindowError ? e.failureClass : "unknown";
+ fail(item, cls, message);
+ done += 1;
+ progress();
+ if (cls === "rate_limit") {
+ // fetchWindowManaged recorded the backoff before it threw.
+ log(` ${platform} rate-limited this run — stopping; the rest are left for a later run.`);
+ stop("rate-limit", i + 1);
+ break;
+ }
+ if (cls === "network") {
+ networkStreak += 1;
+ if (networkStreak >= 2) {
+ await deps.recordBackoff(platform, paths, "network").catch(() => {});
+ log(
+ ` ${platform} refused ${networkStreak} fetches in a row — backing it ` +
+ `off and stopping; the rest are left for a later run.`,
+ );
+ stop("network", i + 1);
+ break;
+ }
+ } else {
+ networkStreak = 0;
+ }
+ continue;
+ } finally {
+ // Settled, refused or not: the next fetch on this platform, in this run
+ // or the next one, starts no sooner than the gap from now.
+ deps.noteGap(gapKey, gap);
+ }
+ done += 1;
+ progress();
+ }
+
+ log(
+ `Fetch windows: ${result.fetched.length} fetched, ${result.cached.length} cached, ` +
+ `${result.failed.length} failed` +
+ (result.notAttempted.length
+ ? `, ${result.notAttempted.length} not attempted (stopped: ${result.stopped})`
+ : "") +
+ ".",
+ );
+ return result;
+}
diff --git a/common/controller/transcribeFile.test.ts b/common/controller/transcribeFile.test.ts
@@ -0,0 +1,463 @@
+import { test } from "node:test";
+import assert from "node:assert/strict";
+import {
+ chmod,
+ mkdir,
+ mkdtemp,
+ readFile,
+ readdir,
+ rm,
+ symlink,
+ writeFile,
+} from "node:fs/promises";
+import os from "node:os";
+import path from "node:path";
+import type { Worker } from "../lib/workers";
+
+// Run with:
+// pnpm --filter yt-dlp-transcript-common exec tsx --test controller/transcribeFile.test.ts
+//
+// The one-off file transcription (`pnpm ops transcribe`): the body's shape, the
+// refusals that come before any job, the window's ffmpeg cut and the cue
+// offset, which engine and model a worker resolves to — and one whole job, run
+// through the real registry and worker pool against a FAKE ffmpeg and a FAKE
+// engine (a parakeet worker whose `bin` is a script), in a temp corpus. Nothing
+// here spawns a real engine or touches a real transcripts dir.
+
+const ROOT = await mkdtemp(path.join(os.tmpdir(), "transcribe-file-"));
+const CORPUS = path.join(ROOT, "transcripts");
+const OUTSIDE = path.join(ROOT, "outside");
+const TMP = path.join(ROOT, "tmp");
+const BIN = path.join(ROOT, "bin");
+await mkdir(path.join(CORPUS, "channels", "chan", "data"), { recursive: true });
+await mkdir(OUTSIDE, { recursive: true });
+await mkdir(TMP, { recursive: true });
+await mkdir(BIN, { recursive: true });
+
+// A fake ffmpeg: records its argv, writes a WAV-sized file to its last arg —
+// an EMPTY one (a bare 44-byte header) for an input named *empty*.
+const FFMPEG_ARGS = path.join(ROOT, "ffmpeg-args.json");
+const ffmpeg = path.join(BIN, "ffmpeg");
+await writeFile(
+ ffmpeg,
+ `#!/usr/bin/env node
+const fs = require("node:fs");
+const args = process.argv.slice(2);
+fs.writeFileSync(${JSON.stringify(FFMPEG_ARGS)}, JSON.stringify(args));
+const input = args[args.indexOf("-i") + 1];
+const out = args[args.length - 1];
+fs.writeFileSync(out, Buffer.alloc(input.includes("empty") ? 44 : 4096));
+`,
+);
+await chmod(ffmpeg, 0o755);
+
+// A fake engine standing in for the parakeet wrapper: records its argv and
+// cwd, and writes chough-native JSON to its --output path (cue times from 0,
+// as a window's always are).
+const ENGINE_ARGS = path.join(ROOT, "engine-args.json");
+const engine = path.join(BIN, "fake-parakeet");
+await writeFile(
+ engine,
+ `#!/usr/bin/env node
+const fs = require("node:fs");
+const args = process.argv.slice(2);
+fs.writeFileSync(${JSON.stringify(ENGINE_ARGS)}, JSON.stringify({ args, cwd: process.cwd() }));
+const audio = args[args.length - 1];
+if (!fs.existsSync(audio)) { console.error("no audio " + audio); process.exit(3); }
+const out = args[args.indexOf("--output") + 1];
+fs.writeFileSync(out, JSON.stringify({ chunk_data: [
+ { start_time: 0.5, end_time: 1.25, text: " hello" },
+ { start_time: 2, end_time: 3.5, text: "world " },
+] }));
+`,
+);
+await chmod(engine, 0o755);
+
+const WORKERS: Worker[] = [
+ {
+ id: "gpu",
+ name: "GPU parakeet",
+ kind: "local",
+ enabled: true,
+ priority: 0,
+ appId: "parakeet",
+ config: { bin: engine, model: "/models/tdt.gguf" },
+ },
+ {
+ id: "cpu",
+ name: "CPU parakeet",
+ kind: "local",
+ enabled: true,
+ priority: 5,
+ appId: "parakeet",
+ config: { bin: engine, device: "cpu" },
+ },
+ {
+ id: "off",
+ name: "Switched off",
+ kind: "local",
+ enabled: false,
+ priority: 9,
+ appId: "parakeet",
+ config: { bin: engine },
+ },
+ {
+ id: "far",
+ name: "A remote",
+ kind: "remote",
+ enabled: true,
+ priority: 1,
+ remote: { baseUrl: "http://127.0.0.1:9", slots: 1 },
+ },
+];
+const SETTINGS_FILE = path.join(ROOT, "settings.json");
+const SETTINGS_TEXT = JSON.stringify({ workers: WORKERS }, null, 2);
+await writeFile(SETTINGS_FILE, SETTINGS_TEXT);
+
+// Set before getPaths (which caches) is first reached through the imports.
+process.env.TRANSCRIPTS_DIR = CORPUS;
+process.env.SETTINGS_FILE = SETTINGS_FILE;
+process.env.FFMPEG_BIN = ffmpeg;
+process.env.PARAKEET_MODEL = "/models/default.gguf";
+process.env.TMPDIR = TMP;
+
+const {
+ TRANSCRIBE_RESULT_MARKER,
+ checkTranscribeFileRequest,
+ corpusRootContaining,
+ corpusRoots,
+ describeTranscribeWorker,
+ enqueueTranscribeFile,
+ offsetCues,
+ parseTranscribeFileBody,
+ transcribeWorkerFilter,
+ windowOf,
+ windowWavArgs,
+ wordsFromTranscript,
+} = await import("./transcribeFile");
+const { getPaths } = await import("../lib/paths");
+const paths = getPaths();
+
+test.after(() => rm(ROOT, { recursive: true, force: true }));
+
+const MEDIA = path.join(OUTSIDE, "clip.mp4");
+await writeFile(MEDIA, "not really media");
+
+// --- the body ---------------------------------------------------------------
+
+test("a body needs an absolute path", () => {
+ assert.match(
+ (parseTranscribeFileBody({}) as { error: string }).error,
+ /"path" is required/,
+ );
+ assert.match(
+ (parseTranscribeFileBody({ path: "clip.mp4" }) as { error: string }).error,
+ /"path" must be absolute/,
+ );
+ assert.match(
+ (parseTranscribeFileBody({ path: 7 }) as { error: string }).error,
+ /"path" is required/,
+ );
+ const ok = parseTranscribeFileBody({ path: "/a/../b/clip.mp4" });
+ assert.deepEqual(ok, { ok: true, value: { path: "/b/clip.mp4" } });
+});
+
+test("a window must end after it starts, in non-negative seconds", () => {
+ const err = (b: Record<string, unknown>) =>
+ (parseTranscribeFileBody({ path: MEDIA, ...b }) as { error?: string }).error;
+ assert.match(err({ start: 10, end: 10 })!, /"end" \(10\) must be after "start" \(10\)/);
+ assert.match(err({ start: 10, end: 4 })!, /must be after/);
+ assert.match(err({ end: 0 })!, /"end" \(0\) must be after "start" \(0\)/);
+ assert.match(err({ start: -1 })!, /"start" must be a number of seconds/);
+ assert.match(err({ end: "30" })!, /"end" must be a number of seconds/);
+ assert.match(err({ start: Number.NaN })!, /"start" must be/);
+ assert.equal(err({ start: 10, end: 10.5 }), undefined);
+ assert.equal(err({ start: 10 }), undefined);
+ assert.equal(err({ end: 10 }), undefined);
+});
+
+test("workerId and out are checked for shape", () => {
+ const err = (b: Record<string, unknown>) =>
+ (parseTranscribeFileBody({ path: MEDIA, ...b }) as { error?: string }).error;
+ assert.match(err({ workerId: "" })!, /"workerId" must be a non-empty string/);
+ assert.match(err({ workerId: 3 })!, /"workerId"/);
+ assert.match(err({ out: "result.json" })!, /"out" must be an absolute path/);
+ assert.match(err({ out: 1 })!, /"out" must be a non-empty string/);
+});
+
+test("words is a boolean, kept only when true", () => {
+ const parsed = (b: Record<string, unknown>) => parseTranscribeFileBody({ path: MEDIA, ...b });
+ assert.match((parsed({ words: "yes" }) as { error: string }).error, /"words" must be true or false/);
+ const on = parsed({ words: true });
+ assert.ok(on.ok && on.value.words === true);
+ const off = parsed({ words: false });
+ assert.ok(off.ok && !("words" in off.value));
+});
+
+test("wordsFromTranscript shifts the engine's words onto the file's clock", () => {
+ const raw = JSON.stringify({
+ chunk_data: [{ start_time: 0, end_time: 1, text: "um so" }],
+ words: [
+ { w: " um", start: 0.12, end: 0.4, conf: 0.9 },
+ { w: "so", start: 0.5, end: 0.7 },
+ { w: " ", start: 0.8, end: 0.9 },
+ { w: "bad", start: "x", end: 1 },
+ ],
+ });
+ assert.deepEqual(wordsFromTranscript(raw, 120), [
+ { w: "um", start: 120.12, end: 120.4, conf: 0.9 },
+ { w: "so", start: 120.5, end: 120.7 },
+ ]);
+ // Another engine's document, or an older wrapper: no words, not a failure.
+ assert.deepEqual(wordsFromTranscript(JSON.stringify({ chunk_data: [] }), 0), []);
+ assert.deepEqual(wordsFromTranscript("not json", 0), []);
+});
+
+// --- the disk checks --------------------------------------------------------
+
+const ctx = { paths, workers: WORKERS };
+async function check(body: Record<string, unknown>, extra: object = {}) {
+ const parsed = parseTranscribeFileBody({ path: MEDIA, ...body });
+ assert.ok(parsed.ok, JSON.stringify(parsed));
+ const res = await checkTranscribeFileRequest(parsed.value, { ...ctx, ...extra });
+ return res.ok ? null : res.error;
+}
+
+test("a path that is not a readable file is refused", async () => {
+ assert.match((await check({ path: path.join(OUTSIDE, "nope.mp4") }))!, /does not exist/);
+ assert.match((await check({ path: OUTSIDE }))!, /is not a file/);
+ const locked = path.join(OUTSIDE, "locked.wav");
+ await writeFile(locked, "x");
+ await chmod(locked, 0o000);
+ // root reads anything; the refusal can only be seen as a user.
+ if (process.getuid?.() !== 0) {
+ assert.match((await check({ path: locked }))!, /is not readable/);
+ }
+ assert.equal(await check({}), null);
+});
+
+test("an out inside the corpus is refused, through a symlink too", async () => {
+ assert.match(
+ (await check({ out: path.join(CORPUS, "result.json") }))!,
+ /is inside the corpus/,
+ );
+ assert.match(
+ (await check({ out: path.join(CORPUS, "channels", "chan", "data", "x.json") }))!,
+ /is inside the corpus/,
+ );
+ // A link OUTSIDE the corpus that points INTO it.
+ const link = path.join(OUTSIDE, "into-corpus");
+ await symlink(path.join(CORPUS, "channels"), link);
+ assert.match(
+ (await check({ out: path.join(link, "r.json") }))!,
+ /is inside the corpus/,
+ );
+ // A channel's media linked off to another drive, reached through the corpus.
+ const drive = path.join(ROOT, "drive", "chan", "media");
+ await mkdir(drive, { recursive: true });
+ await symlink(drive, path.join(CORPUS, "channels", "chan", "media"));
+ assert.match(
+ (await check({ out: path.join(CORPUS, "channels", "chan", "media", "r.json") }))!,
+ /is inside the corpus/,
+ );
+ // ...and that drive written to directly is a storage location root's.
+ assert.equal(await check({ out: path.join(drive, "r.json") }), null);
+ assert.match(
+ (await check(
+ { out: path.join(drive, "r.json") },
+ { locations: [{ root: path.join(ROOT, "drive") }] },
+ ))!,
+ /is inside the corpus/,
+ );
+ assert.equal(await check({ out: path.join(OUTSIDE, "r.json") }), null);
+});
+
+test("an out that is the input, a directory, or in a missing directory is refused", async () => {
+ assert.match((await check({ out: MEDIA }))!, /is the input file itself/);
+ assert.match((await check({ out: OUTSIDE }))!, /is a directory/);
+ assert.match(
+ (await check({ out: path.join(OUTSIDE, "no", "such", "r.json") }))!,
+ /does not exist/,
+ );
+});
+
+test("corpusRoots and corpusRootContaining name the root", async () => {
+ const roots = corpusRoots(paths, [{ root: "/mnt/platter" }]);
+ assert.ok(roots.includes(path.resolve(CORPUS)));
+ assert.ok(roots.includes("/mnt/platter"));
+ assert.equal(await corpusRootContaining("/mnt/platter/x/media/a.json", roots), "/mnt/platter");
+ assert.equal(await corpusRootContaining("/mnt/platterx/a.json", roots), null);
+});
+
+test("an unknown workerId is refused, naming the known ones", async () => {
+ const err = await check({ workerId: "nope" });
+ assert.match(err!, /no worker "nope" — known: gpu, cpu, off, far/);
+});
+
+test("a remote workerId is refused, naming the local ones", async () => {
+ assert.match(
+ (await check({ workerId: "far" }))!,
+ /"far" is a remote worker; a file is transcribed on a local one \(gpu, cpu, off\)/,
+ );
+});
+
+test("a named worker switched off on the Workers page is refused, not waited for", async () => {
+ const workerStates = new Map([
+ ["off", { state: "disabled", degraded: false }],
+ ["cpu", { state: "enabled", degraded: true }],
+ ["gpu", { state: "enabled", degraded: false }],
+ ]);
+ assert.match((await check({ workerId: "off" }, { workerStates }))!, /"off" is disabled/);
+ assert.match((await check({ workerId: "cpu" }, { workerStates }))!, /"cpu" is degraded/);
+ assert.equal(await check({ workerId: "gpu" }, { workerStates }), null);
+});
+
+test("with no local worker configured, a default request is refused", async () => {
+ const err = await check({}, { workers: WORKERS.filter((w) => w.kind !== "local") });
+ assert.match(err!, /no local transcription worker is configured/);
+});
+
+// --- the pure pieces ----------------------------------------------------------
+
+test("the window is cut to 16 kHz mono WAV: input seek, then a duration", () => {
+ assert.deepEqual(windowWavArgs("/in.mp4", "/t/a.wav", { start: 120, end: 150.5 }), [
+ "-nostdin", "-hide_banner", "-v", "error", "-y",
+ "-ss", "120", "-i", "/in.mp4", "-t", "30.5",
+ "-vn", "-ac", "1", "-ar", "16000", "-c:a", "pcm_s16le", "-f", "wav", "/t/a.wav",
+ ]);
+ const whole = windowWavArgs("/in.mp4", "/t/a.wav");
+ assert.ok(!whole.includes("-ss") && !whole.includes("-t"));
+ const toEnd = windowWavArgs("/in.mp4", "/t/a.wav", { end: 40 });
+ assert.ok(!toEnd.includes("-ss"));
+ assert.deepEqual(toEnd.slice(toEnd.indexOf("-t"), toEnd.indexOf("-t") + 2), ["-t", "40"]);
+});
+
+test("cue times are shifted back onto the source's clock", () => {
+ assert.deepEqual(
+ offsetCues([{ start: 0.5, end: 1.25, text: "a" }, { start: 2.0004, end: 3, text: "b" }], 120),
+ [{ start: 120.5, end: 121.25, text: "a" }, { start: 122, end: 123, text: "b" }],
+ );
+ assert.deepEqual(offsetCues([{ start: 1, end: 2, text: "x" }], 0), [{ start: 1, end: 2, text: "x" }]);
+ assert.equal(windowOf({}), null);
+ assert.deepEqual(windowOf({ start: 5 }), { start: 5, end: null });
+ assert.deepEqual(windowOf({ end: 9 }), { start: 0, end: 9 });
+});
+
+test("a worker resolves to its engine and model, the app's default model when unset", () => {
+ assert.deepEqual(describeTranscribeWorker(WORKERS[0]), {
+ id: "gpu", name: "GPU parakeet", appId: "parakeet", model: "/models/tdt.gguf", device: null,
+ });
+ assert.deepEqual(describeTranscribeWorker(WORKERS[1]), {
+ id: "cpu", name: "CPU parakeet", appId: "parakeet", model: "/models/default.gguf", device: "cpu",
+ });
+ const chough = describeTranscribeWorker({
+ id: "c", name: "c", kind: "local", enabled: true, priority: 0, appId: "chough",
+ });
+ assert.equal(chough.appId, "chough");
+ assert.equal(chough.model, null);
+});
+
+test("the worker filter keeps to local workers, or to the one named", () => {
+ const any = transcribeWorkerFilter();
+ assert.deepEqual(WORKERS.filter(any).map((w) => w.id), ["gpu", "cpu", "off"]);
+ assert.deepEqual(WORKERS.filter(transcribeWorkerFilter("cpu")).map((w) => w.id), ["cpu"]);
+ assert.deepEqual(WORKERS.filter(transcribeWorkerFilter("far")).map((w) => w.id), []);
+});
+
+// --- one whole job ------------------------------------------------------------
+
+async function runJob(body: Record<string, unknown>) {
+ const res = await enqueueTranscribeFile(body, { paths });
+ if (!res.ok) return { error: res.error };
+ void res.stream.cancel();
+ const done = await res.done;
+ const log = await readFile(path.join(paths.jobsDir, `${res.jobId}.log`), "utf8");
+ const line = log.split("\n").find((l) => l.startsWith(TRANSCRIBE_RESULT_MARKER));
+ return {
+ status: done.status,
+ log,
+ result: line ? JSON.parse(line.slice(TRANSCRIBE_RESULT_MARKER.length)) : null,
+ };
+}
+
+test("a window is transcribed on the named worker, its cues on the source clock, out written", async () => {
+ const out = path.join(OUTSIDE, "result.json");
+ const run = await runJob({ path: MEDIA, start: 120, end: 150, workerId: "cpu", out });
+ assert.equal(run.status, "done", run.log);
+ const r = run.result;
+ assert.deepEqual(r.window, { start: 120, end: 150 });
+ assert.deepEqual(r.cues, [
+ { start: 120.5, end: 121.25, text: "hello" },
+ { start: 122, end: 123.5, text: "world" },
+ ]);
+ assert.equal(r.text, "hello world");
+ assert.deepEqual(r.worker, {
+ id: "cpu", name: "CPU parakeet", appId: "parakeet", model: "/models/default.gguf", device: "cpu",
+ });
+ assert.equal(r.transcriptFormat, "chough-json");
+ assert.equal(r.path, MEDIA);
+ // The same document in `out`.
+ assert.deepEqual(JSON.parse(await readFile(out, "utf8")), r);
+
+ // ffmpeg cut the window; the engine got the registry's command line for
+ // THAT worker (its device, the default model) and ran in the scratch dir.
+ const ff = JSON.parse(await readFile(FFMPEG_ARGS, "utf8")) as string[];
+ assert.deepEqual(ff.slice(ff.indexOf("-ss"), ff.indexOf("-ss") + 4), ["-ss", "120", "-i", MEDIA]);
+ assert.deepEqual(ff.slice(ff.indexOf("-t"), ff.indexOf("-t") + 2), ["-t", "30"]);
+ const eng = JSON.parse(await readFile(ENGINE_ARGS, "utf8")) as { args: string[]; cwd: string };
+ assert.deepEqual(eng.args.slice(0, 2), ["--model", "/models/default.gguf"]);
+ assert.ok(eng.args.includes("--device") && eng.args.includes("cpu"), eng.args.join(" "));
+ assert.equal(eng.args[eng.args.length - 1], "audio.wav");
+ assert.ok(eng.cwd.startsWith(TMP), `engine ran in ${eng.cwd}`);
+
+ // The scratch dir is gone, and nothing landed in the corpus but the job log.
+ assert.deepEqual(
+ (await readdir(TMP)).filter((n) => n.startsWith("archilyzer-transcribe-")),
+ [],
+ );
+ assert.deepEqual(
+ (await readdir(CORPUS)).filter((n) => n !== ".jobs" && n !== "channels"),
+ [],
+ );
+ assert.deepEqual(await readdir(path.join(CORPUS, "channels", "chan", "data")), []);
+ // settings.json is byte-for-byte what it was.
+ assert.equal(await readFile(SETTINGS_FILE, "utf8"), SETTINGS_TEXT);
+});
+
+test("with no workerId the pool's highest-priority local worker runs it, whole file, no offset", async () => {
+ const run = await runJob({ path: MEDIA });
+ assert.equal(run.status, "done", run.log);
+ assert.equal(run.result.worker.id, "gpu");
+ assert.equal(run.result.worker.model, "/models/tdt.gguf");
+ assert.equal(run.result.window, null);
+ assert.equal(run.result.cues[0].start, 0.5);
+ const ff = JSON.parse(await readFile(FFMPEG_ARGS, "utf8")) as string[];
+ assert.ok(!ff.includes("-ss") && !ff.includes("-t"));
+ assert.equal(await readFile(SETTINGS_FILE, "utf8"), SETTINGS_TEXT);
+});
+
+test("a window holding no audio fails the job with a sentence, and cleans up", async () => {
+ const empty = path.join(OUTSIDE, "empty.wav");
+ await writeFile(empty, "x");
+ const run = await runJob({ path: empty, start: 9999 });
+ assert.equal(run.status, "failed");
+ assert.match(run.log!, /no audio in .*start past the end/);
+ assert.equal(run.result, null);
+ assert.deepEqual(
+ (await readdir(TMP)).filter((n) => n.startsWith("archilyzer-transcribe-")),
+ [],
+ );
+});
+
+test("the guards answer before any job exists", async () => {
+ assert.match((await runJob({ path: "rel.mp4" })).error!, /must be absolute/);
+ assert.match((await runJob({ path: MEDIA, start: 5, end: 5 })).error!, /must be after/);
+ assert.match((await runJob({ path: MEDIA, workerId: "ghost" })).error!, /no worker "ghost"/);
+ assert.match(
+ (await runJob({ path: MEDIA, out: path.join(CORPUS, "x.json") })).error!,
+ /inside the corpus/,
+ );
+ // The pool seeded "off" from settings as disabled.
+ assert.match((await runJob({ path: MEDIA, workerId: "off" })).error!, /"off" is disabled/);
+ assert.equal(await readFile(SETTINGS_FILE, "utf8"), SETTINGS_TEXT);
+});
diff --git a/common/controller/transcribeFile.ts b/common/controller/transcribeFile.ts
@@ -0,0 +1,570 @@
+// ONE-OFF TRANSCRIPTION OF AN ARBITRARY FILE, AS AN EDITOR JOB — what
+// `pnpm ops transcribe` (`POST /api/ops/transcribe`) enqueues.
+//
+// A quote check or "what is audible in this clip" used to mean a hand-run
+// whisper-cli / parakeet-cli with a model path typed from memory. This runs the
+// SAME engine the corpus does, through the same path: the worker pool hands out
+// a lease (so a one-off never oversubscribes the GPU slot auto-transcribe is
+// using, and jumps ahead of queued background work as any manual transcribe
+// does), and `transcribeWithWorker` builds the command line from the worker's
+// config through the transcription-app registry. Nothing here knows an engine's
+// argv.
+//
+// WHAT IT TOUCHES. It reads `path` (anywhere, the corpus included) and writes
+// only to a scratch dir under the OS temp dir, removed afterwards, and to `out`
+// when given — which is refused inside the corpus (the transcripts dir, the
+// saved-video store, the sites dir, every storage location root; compared both
+// as written and resolved through symlinks, since a channel's `media` may be a
+// link to another drive). It never writes settings.json, a sidecar, or anything
+// under a channel. Its job log lands where every job's does.
+//
+// THE AUDIO IS ALWAYS A 16 kHz MONO WAV CUT BY ffmpeg, whole file or window.
+// Two reasons: the engines disagree about containers (whisper-cli wants audio,
+// parakeet's wrapper reads anything ffmpeg does), and parakeet's wrapper keeps
+// its resumable work dir BESIDE its input — given the source file directly, it
+// would create `.<name>.parakeet/` next to it, possibly inside the corpus.
+//
+// THE RESULT is the cue format the index uses ({start, end, text}, seconds),
+// shifted back into the SOURCE file's clock when a window was cut, plus which
+// worker, engine and model produced it. It is written to `out` when given and
+// always logged as ONE line starting with TRANSCRIBE_RESULT_MARKER, which is
+// how `pnpm ops transcribe --wait` prints it without reading any file on the
+// editor's disk.
+//
+// LOCAL WORKERS ONLY. A remote worker delegates by channel and video id (or by
+// an upload into ITS pool); a one-off file has neither identity, and the point
+// of the command is "this machine's engine". The default is the worker
+// auto-transcribe would get — the pool's highest-priority free one — among the
+// local workers.
+
+import os from "node:os";
+import path from "node:path";
+import { access, constants, mkdtemp, readFile, realpath, rm, stat } from "node:fs/promises";
+import { execa } from "execa";
+import type { Paths } from "../lib/paths";
+import type { Worker } from "../lib/workers";
+import type { Cue } from "../lib/vtt";
+import type { StorageLocation } from "../lib/storageLocations";
+import {
+ getTranscriptionApp,
+ type TranscriptOutputFormat,
+} from "../lib/transcriptionApps";
+import { detectTranscriptFormat, parseTranscriptJson } from "../lib/whisper";
+import { writeJsonAtomic } from "../lib/jsonFile-server";
+import { getSettings } from "../lib/settings";
+import { getWorkerPool, type WorkerFilter } from "../jobs/workerPool";
+import { makeTaskTracker } from "../jobs/taskHooks";
+import {
+ runManagedFunction,
+ type JobRunContext,
+ type StreamActionResult,
+} from "../jobs/streamCommand";
+import { transcribeWithWorker } from "./transcribeOne";
+
+export const TRANSCRIBE_FILE_JOB_KIND = "transcribe-file";
+
+// The log line carrying the result. One line, compact JSON after the marker.
+// scripts/archilyzer-ops.mjs matches the same string.
+export const TRANSCRIBE_RESULT_MARKER = "@@transcribe-result ";
+
+// The body `/api/ops/transcribe` accepts; anything else is a 400.
+export const TRANSCRIBE_FILE_BODY_KEYS = [
+ "path",
+ "start",
+ "end",
+ "workerId",
+ "out",
+ "words",
+] as const;
+
+const AUDIO_NAME = "audio.wav";
+// A WAV header with no samples after it: the window held no audio.
+const EMPTY_WAV_BYTES = 44;
+
+export type TranscribeFileRequest = {
+ path: string;
+ // Seconds into the source. Absent start = 0; absent end = to the end.
+ start?: number;
+ end?: number;
+ workerId?: string;
+ out?: string;
+ // True returns the engine's word timestamps as well as the cues. Only an
+ // engine that keeps them (parakeet) answers with any; the rest give none.
+ words?: boolean;
+};
+
+// One word as the engine timed it, on the source file's clock (seconds).
+export type TranscribedWord = { w: string; start: number; end: number; conf?: number };
+
+export type TranscribeWorkerInfo = {
+ id: string;
+ name: string;
+ appId: string;
+ model: string | null;
+ device: string | null;
+};
+
+export type TranscribeFileResult = {
+ version: 1;
+ path: string;
+ // Null when the whole file was transcribed. `end: null` = to the end.
+ window: { start: number; end: number | null } | null;
+ worker: TranscribeWorkerInfo;
+ transcriptFormat: TranscriptOutputFormat;
+ transcribedAt: string;
+ durationMs: number;
+ cues: Cue[];
+ text: string;
+ // Present only when the request asked for words: [] when the engine has none.
+ words?: TranscribedWord[];
+};
+
+type Check<T> = { ok: true; value: T } | { ok: false; error: string };
+
+// --- body ------------------------------------------------------------------
+
+function seconds(raw: unknown, key: string): Check<number | undefined> {
+ if (raw === undefined || raw === null) return { ok: true, value: undefined };
+ if (typeof raw !== "number" || !Number.isFinite(raw) || raw < 0) {
+ return { ok: false, error: `"${key}" must be a number of seconds, zero or more` };
+ }
+ return { ok: true, value: raw };
+}
+
+// The body's shape, judged without touching the disk. Every sentence here is
+// the refusal an ops caller reads.
+export function parseTranscribeFileBody(
+ body: Record<string, unknown>,
+): Check<TranscribeFileRequest> {
+ const p = body.path;
+ if (typeof p !== "string" || !p.trim()) {
+ return { ok: false, error: '"path" is required: an absolute path to an audio or video file' };
+ }
+ if (!path.isAbsolute(p)) {
+ return { ok: false, error: `"path" must be absolute (got "${p}")` };
+ }
+ const start = seconds(body.start, "start");
+ if (!start.ok) return start;
+ const end = seconds(body.end, "end");
+ if (!end.ok) return end;
+ if (end.value !== undefined && end.value <= (start.value ?? 0)) {
+ return {
+ ok: false,
+ error: `"end" (${end.value}) must be after "start" (${start.value ?? 0})`,
+ };
+ }
+ const workerId = body.workerId;
+ if (workerId !== undefined && (typeof workerId !== "string" || !workerId.trim())) {
+ return { ok: false, error: '"workerId" must be a non-empty string' };
+ }
+ const out = body.out;
+ if (out !== undefined) {
+ if (typeof out !== "string" || !out.trim()) {
+ return { ok: false, error: '"out" must be a non-empty string' };
+ }
+ if (!path.isAbsolute(out)) {
+ return { ok: false, error: `"out" must be an absolute path (got "${out}")` };
+ }
+ }
+ const words = body.words;
+ if (words !== undefined && typeof words !== "boolean") {
+ return { ok: false, error: '"words" must be true or false' };
+ }
+ return {
+ ok: true,
+ value: {
+ path: path.resolve(p),
+ ...(start.value !== undefined ? { start: start.value } : {}),
+ ...(end.value !== undefined ? { end: end.value } : {}),
+ ...(typeof workerId === "string" ? { workerId: workerId.trim() } : {}),
+ ...(typeof out === "string" ? { out: path.resolve(out) } : {}),
+ ...(words === true ? { words: true } : {}),
+ },
+ };
+}
+
+// --- the corpus fence --------------------------------------------------------
+
+// `p` with every symlink in its EXISTING prefix resolved; the part that does
+// not exist yet is appended as written. An `out` is usually a file that does
+// not exist, in a directory that does.
+export async function realpathDeep(p: string): Promise<string> {
+ const abs = path.resolve(p);
+ const rest: string[] = [];
+ let cur = abs;
+ for (;;) {
+ try {
+ const real = await realpath(cur);
+ return rest.length ? path.join(real, ...rest.reverse()) : real;
+ } catch {
+ const parent = path.dirname(cur);
+ if (parent === cur) return abs;
+ rest.push(path.basename(cur));
+ cur = parent;
+ }
+ }
+}
+
+function within(child: string, parent: string): boolean {
+ const rel = path.relative(parent, child);
+ return rel === "" || (!rel.startsWith("..") && !path.isAbsolute(rel));
+}
+
+// Every directory a write of ours must stay out of.
+export function corpusRoots(
+ paths: Pick<Paths, "transcriptsDir" | "channelsDir" | "savedVideosDir" | "sitesDir">,
+ locations: readonly Pick<StorageLocation, "root">[] = [],
+): string[] {
+ const roots = [
+ paths.transcriptsDir,
+ paths.channelsDir,
+ paths.savedVideosDir,
+ paths.sitesDir,
+ ...locations.map((l) => l.root),
+ ].filter((r): r is string => typeof r === "string" && r.trim() !== "");
+ return [...new Set(roots.map((r) => path.resolve(r)))];
+}
+
+// The root `target` lies under, or null. Both sides are compared as written
+// AND resolved: `channels/x/media` may be a link to another drive, and a
+// location root may itself be reached through a link.
+export async function corpusRootContaining(
+ target: string,
+ roots: readonly string[],
+): Promise<string | null> {
+ const forms = [path.resolve(target), await realpathDeep(target)];
+ for (const root of roots) {
+ const rootForms = [path.resolve(root), await realpathDeep(root)];
+ for (const t of forms) {
+ for (const r of rootForms) {
+ if (within(t, r)) return root;
+ }
+ }
+ }
+ return null;
+}
+
+// --- workers -----------------------------------------------------------------
+
+// Which workers may take the transcription: local ones, or the one named.
+export function transcribeWorkerFilter(workerId?: string): WorkerFilter {
+ return (w) => w.kind === "local" && (workerId === undefined || w.id === workerId);
+}
+
+// What a result records about the worker that produced it: the engine (app id)
+// and the model AFTER the app's own default, resolved by the same registry that
+// built the command line.
+export function describeTranscribeWorker(worker: Worker): TranscribeWorkerInfo {
+ const app = getTranscriptionApp(worker.appId);
+ const config = worker.config ?? {};
+ return {
+ id: worker.id,
+ name: worker.name,
+ appId: app.id,
+ model: app.resolveModel(config) ?? null,
+ device: config.device?.trim() || null,
+ };
+}
+
+// --- the disk checks ---------------------------------------------------------
+
+export type TranscribeFileContext = {
+ paths: Paths;
+ workers: readonly Worker[];
+ locations?: readonly Pick<StorageLocation, "root">[];
+ // The pool's runtime view (id → state). A named worker switched off on the
+ // Workers page is refused rather than waited for: the wait would be forever.
+ workerStates?: ReadonlyMap<string, { state: string; degraded: boolean }>;
+};
+
+export async function checkTranscribeFileRequest(
+ req: TranscribeFileRequest,
+ ctx: TranscribeFileContext,
+): Promise<Check<TranscribeFileRequest>> {
+ let st;
+ try {
+ st = await stat(req.path);
+ } catch {
+ return { ok: false, error: `"path" ${req.path} does not exist` };
+ }
+ if (!st.isFile()) {
+ return { ok: false, error: `"path" ${req.path} is not a file` };
+ }
+ try {
+ await access(req.path, constants.R_OK);
+ } catch {
+ return { ok: false, error: `"path" ${req.path} is not readable` };
+ }
+
+ if (req.out !== undefined) {
+ const root = await corpusRootContaining(
+ req.out,
+ corpusRoots(ctx.paths, ctx.locations ?? []),
+ );
+ if (root) {
+ return {
+ ok: false,
+ error: `"out" ${req.out} is inside the corpus (${root}) — this command never writes there; pick a path outside it`,
+ };
+ }
+ if ((await realpathDeep(req.out)) === (await realpathDeep(req.path))) {
+ return { ok: false, error: '"out" is the input file itself' };
+ }
+ const outStat = await stat(req.out).catch(() => null);
+ if (outStat?.isDirectory()) {
+ return { ok: false, error: `"out" ${req.out} is a directory; name a file` };
+ }
+ const dirStat = await stat(path.dirname(req.out)).catch(() => null);
+ if (!dirStat?.isDirectory()) {
+ return {
+ ok: false,
+ error: `"out": the directory ${path.dirname(req.out)} does not exist`,
+ };
+ }
+ }
+
+ const local = ctx.workers.filter((w) => w.kind === "local");
+ if (req.workerId !== undefined) {
+ const named = ctx.workers.find((w) => w.id === req.workerId);
+ if (!named) {
+ return {
+ ok: false,
+ error: `no worker "${req.workerId}" — known: ${
+ ctx.workers.map((w) => w.id).join(", ") || "none"
+ }`,
+ };
+ }
+ if (named.kind !== "local") {
+ return {
+ ok: false,
+ error: `worker "${req.workerId}" is a ${named.kind} worker; a file is transcribed on a local one (${
+ local.map((w) => w.id).join(", ") || "none configured"
+ })`,
+ };
+ }
+ const runtime = ctx.workerStates?.get(req.workerId);
+ if (runtime && (runtime.state !== "enabled" || runtime.degraded)) {
+ return {
+ ok: false,
+ error: `worker "${req.workerId}" is ${
+ runtime.degraded ? "degraded" : runtime.state
+ } on the Workers page — enable it there, or leave out "workerId"`,
+ };
+ }
+ } else if (local.length === 0) {
+ return { ok: false, error: "no local transcription worker is configured" };
+ }
+ return { ok: true, value: req };
+}
+
+// --- the work ----------------------------------------------------------------
+
+// ffmpeg's arguments for the 16 kHz mono WAV: input seeking for the start, a
+// duration for the end, no video.
+export function windowWavArgs(
+ src: string,
+ dst: string,
+ win: { start?: number; end?: number } = {},
+): string[] {
+ const start = win.start ?? 0;
+ const args = ["-nostdin", "-hide_banner", "-v", "error", "-y"];
+ if (start > 0) args.push("-ss", String(start));
+ args.push("-i", src);
+ if (win.end !== undefined) args.push("-t", String(win.end - start));
+ args.push("-vn", "-ac", "1", "-ar", "16000", "-c:a", "pcm_s16le", "-f", "wav", dst);
+ return args;
+}
+
+const ms = (n: number) => Math.round(n * 1000) / 1000;
+
+// Cues from a window start at zero; shift them back onto the source's clock.
+// The top-level `words` an engine asked with `words` wrote (parakeet's
+// wrapper, --words), shifted by `offset` onto the source file's clock. A
+// document without them -- another engine, or an older wrapper -- gives [].
+export function wordsFromTranscript(raw: string, offset: number): TranscribedWord[] {
+ let doc: unknown;
+ try {
+ doc = JSON.parse(raw);
+ } catch {
+ return [];
+ }
+ const list = (doc as { words?: unknown } | null)?.words;
+ if (!Array.isArray(list)) return [];
+ const out: TranscribedWord[] = [];
+ for (const w of list) {
+ if (!w || typeof w !== "object") continue;
+ const { w: text, start, end, conf } = w as Record<string, unknown>;
+ if (typeof text !== "string" || !text.trim()) continue;
+ if (typeof start !== "number" || typeof end !== "number") continue;
+ if (!Number.isFinite(start) || !Number.isFinite(end)) continue;
+ out.push({
+ w: text.trim(),
+ start: ms(start + offset),
+ end: ms(end + offset),
+ ...(typeof conf === "number" && Number.isFinite(conf) ? { conf } : {}),
+ });
+ }
+ return out;
+}
+
+export function offsetCues(cues: readonly Cue[], offset: number): Cue[] {
+ return cues.map((c) => ({
+ start: ms(c.start + offset),
+ end: ms(c.end + offset),
+ text: c.text,
+ }));
+}
+
+export function windowOf(
+ req: Pick<TranscribeFileRequest, "start" | "end">,
+): TranscribeFileResult["window"] {
+ if (req.start === undefined && req.end === undefined) return null;
+ return { start: req.start ?? 0, end: req.end ?? null };
+}
+
+function describeRequest(req: TranscribeFileRequest): string {
+ const win = windowOf(req);
+ return win
+ ? `${req.path} [${win.start}s – ${win.end === null ? "end" : `${win.end}s`}]`
+ : req.path;
+}
+
+export type RunTranscribeFileOpts = {
+ paths: Paths;
+ request: TranscribeFileRequest;
+ onLog: (line: string) => void;
+ signal?: AbortSignal;
+ ctx?: Pick<JobRunContext, "jobId" | "addTask" | "updateTask" | "removeTask" | "recordTaskDone">;
+};
+
+export async function runTranscribeFile(
+ opts: RunTranscribeFileOpts,
+): Promise<TranscribeFileResult> {
+ const { request: req, onLog, paths } = opts;
+ const started = Date.now();
+ const scratch = await mkdtemp(path.join(os.tmpdir(), "archilyzer-transcribe-"));
+ try {
+ const wav = path.join(scratch, AUDIO_NAME);
+ onLog(`Cutting ${describeRequest(req)} to 16 kHz mono WAV…`);
+ try {
+ await execa(paths.ffmpegBin, windowWavArgs(req.path, wav, req), {
+ cancelSignal: opts.signal,
+ });
+ } catch (err) {
+ if (opts.signal?.aborted) throw err;
+ const e = err as { stderr?: string; shortMessage?: string; message: string };
+ throw new Error(
+ `ffmpeg could not read ${req.path}: ${(e.stderr || e.shortMessage || e.message).trim()}`,
+ );
+ }
+ const wavStat = await stat(wav).catch(() => null);
+ if (!wavStat || wavStat.size <= EMPTY_WAV_BYTES) {
+ throw new Error(
+ `no audio in ${describeRequest(req)}${req.start ? " — does the window start past the end of the file?" : ""}`,
+ );
+ }
+
+ // Set by onWorker; a holder, so the closure's write is seen after the await.
+ const used: { worker?: Worker } = {};
+ onLog(
+ req.workerId
+ ? `Waiting for worker ${req.workerId}…`
+ : "Waiting for a free local worker…",
+ );
+ const label = `${path.basename(req.path)}${windowOf(req) ? " (window)" : ""}`;
+ const outcome = await transcribeWithWorker({
+ paths,
+ videoDir: scratch,
+ videoId: path.basename(req.path),
+ audioFilename: AUDIO_NAME,
+ strictAudio: true,
+ tracker: opts.ctx ? makeTaskTracker(opts.ctx, onLog) : undefined,
+ taskId: opts.ctx ? `file:${opts.ctx.jobId}` : undefined,
+ taskLabel: label,
+ onLog,
+ signal: opts.signal,
+ only: transcribeWorkerFilter(req.workerId),
+ onWorker: (w) => {
+ used.worker = w;
+ },
+ skipInlineDiarization: true,
+ ...(req.words ? { words: true } : {}),
+ });
+ const worker = used.worker;
+ if (outcome !== "transcribed" || !worker) {
+ throw new Error(
+ outcome === "paused"
+ ? "the transcription was stopped before it finished"
+ : `the transcription did not run (${outcome})`,
+ );
+ }
+ const raw = await readFile(path.join(scratch, "transcript.json"), "utf8");
+ const transcriptFormat = detectTranscriptFormat(raw) ?? "whisper-json";
+ const cues = offsetCues(parseTranscriptJson(raw, transcriptFormat), req.start ?? 0);
+ return {
+ version: 1,
+ path: req.path,
+ window: windowOf(req),
+ worker: describeTranscribeWorker(worker),
+ transcriptFormat,
+ transcribedAt: new Date().toISOString(),
+ durationMs: Date.now() - started,
+ cues,
+ text: cues.map((c) => c.text.trim()).filter(Boolean).join(" "),
+ ...(req.words ? { words: wordsFromTranscript(raw, req.start ?? 0) } : {}),
+ };
+ } finally {
+ await rm(scratch, { recursive: true, force: true }).catch(() => {});
+ }
+}
+
+// --- the job -----------------------------------------------------------------
+
+// Validate, then enqueue. Every refusal comes back as `{ ok: false, error }`
+// BEFORE a job exists, so the ops route answers it as a 400.
+export async function enqueueTranscribeFile(
+ body: Record<string, unknown>,
+ opts: { paths: Paths },
+): Promise<StreamActionResult> {
+ const parsed = parseTranscribeFileBody(body);
+ if (!parsed.ok) return parsed;
+ const settings = getSettings();
+ const pool = getWorkerPool();
+ const workerStates = new Map(
+ pool.summary().map((w) => [w.id, { state: w.state, degraded: w.degraded }]),
+ );
+ const checked = await checkTranscribeFileRequest(parsed.value, {
+ paths: opts.paths,
+ workers: settings.workers,
+ locations: settings.storage?.locations ?? [],
+ workerStates,
+ });
+ if (!checked.ok) return checked;
+ const req = checked.value;
+ return runManagedFunction({
+ kind: TRANSCRIBE_FILE_JOB_KIND,
+ // Parallel: the worker pool is what serialises transcriptions.
+ queueKey: "",
+ paths: opts.paths,
+ fn: async (onLog, signal, _setProgress, ctx) => {
+ onLog(`Transcribe ${describeRequest(req)}`);
+ const result = await runTranscribeFile({
+ paths: opts.paths,
+ request: req,
+ onLog,
+ signal,
+ ctx,
+ });
+ onLog(
+ `Transcribed with ${result.worker.appId} [${result.worker.id}]` +
+ `${result.worker.model ? ` model ${result.worker.model}` : ""}: ` +
+ `${result.cues.length} cue(s) in ${(result.durationMs / 1000).toFixed(1)}s`,
+ );
+ if (req.out) {
+ await writeJsonAtomic(req.out, result, { indent: 2 });
+ onLog(`Written to ${req.out}`);
+ }
+ onLog(`${TRANSCRIBE_RESULT_MARKER}${JSON.stringify(result)}`);
+ },
+ });
+}
diff --git a/common/controller/transcribeOne.ts b/common/controller/transcribeOne.ts
@@ -5,7 +5,11 @@ import type { Paths } from "../lib/paths";
import type { Worker } from "../lib/workers";
import { getTranscriptionApp } from "../lib/transcriptionApps";
import { detectTranscriptFormat } from "../lib/whisper";
-import { WORKER_DEGRADE_THRESHOLD, getWorkerPool } from "../jobs/workerPool";
+import {
+ WORKER_DEGRADE_THRESHOLD,
+ getWorkerPool,
+ type WorkerFilter,
+} from "../jobs/workerPool";
import type { TaskTracker } from "../jobs/taskHooks";
import { normalizeTranscript } from "./normalizeTranscript";
import { diarizeOneVideo } from "./diarizeOne";
@@ -91,6 +95,13 @@ export type TranscribeOneOptions = {
// resume. transcribeOneVideo returns "paused" so the batch leaves the video
// untranscribed (next run resumes). Ignored by engines without partial support.
partialSignal?: AbortSignal;
+ // True skips the inline diarization pass even when settings turn it on. A
+ // one-off file transcription (controller/transcribeFile.ts) runs in a scratch
+ // dir that is deleted afterwards, so a diarization there is work thrown away.
+ skipInlineDiarization?: boolean;
+ // True asks the engine to keep word timestamps in transcript.json (see
+ // TranscribeBuildInput.words). Only the one-off file transcription asks.
+ words?: boolean;
};
export type TranscribeOneOutcome = "transcribed" | "already-exists" | "paused";
@@ -189,6 +200,7 @@ export async function transcribeOneVideo(
audioFile: resolvedAudio,
outputBase: tmpBase,
config: appConfig,
+ ...(opts.words ? { words: true } : {}),
});
const child = execa(bin, build.argv, {
cwd: opts.videoDir,
@@ -302,7 +314,11 @@ export async function transcribeOneVideo(
// The "paused" early returns above correctly bypass this — there is no
// transcript yet, so there is nothing to diarize alongside.
const diarization = getSettings().diarization;
- if (diarization.enabled && diarization.inlineAfterTranscribe) {
+ if (
+ !opts.skipInlineDiarization &&
+ diarization.enabled &&
+ diarization.inlineAfterTranscribe
+ ) {
try {
const outcome = await diarizeOneVideo({
paths: opts.paths,
@@ -360,6 +376,16 @@ export type TranscribeWithWorkerOptions = {
// Auto-runner units pass true so they park BEHIND any manual (foreground)
// acquire in the worker pool — a manual transcribe preempts queued auto work.
background?: boolean;
+ // Narrows WHICH workers may take this video (the pool's `only`): a one-off
+ // file transcription keeps to local workers, or to the one it was told to use.
+ only?: WorkerFilter;
+ // Called with the worker each attempt runs on, before it starts — how a
+ // caller learns which engine and model produced the transcript.
+ onWorker?: (worker: Worker) => void;
+ // See TranscribeOneOptions.skipInlineDiarization.
+ skipInlineDiarization?: boolean;
+ // See TranscribeOneOptions.words.
+ words?: boolean;
};
// Acquire a worker from the global pool and transcribe one video through it,
@@ -383,11 +409,18 @@ export async function transcribeWithWorker(
if (opts.signal?.aborted || opts.drainSignal?.aborted) throw abortError();
let lease;
try {
- lease = await pool.acquire(acquireSignal, { background: opts.background });
- } catch {
+ lease = await pool.acquire(acquireSignal, {
+ background: opts.background,
+ ...(opts.only ? { only: opts.only } : {}),
+ });
+ } catch (err) {
+ // An `only` no configured worker passes is refused at once — that is
+ // not a cancel, and must not read as one.
+ if (opts.only && !acquireSignal?.aborted) throw err;
throw abortError(); // cancelled or drained while parked
}
const worker = lease.worker;
+ opts.onWorker?.(worker);
const task = opts.tracker?.start({
id: opts.taskId ?? opts.videoId,
label: opts.taskLabel ?? opts.videoId,
@@ -416,6 +449,8 @@ export async function transcribeWithWorker(
onProgress: task ? task.update : undefined,
signal: opts.signal,
partialSignal: partialController?.signal,
+ skipInlineDiarization: opts.skipInlineDiarization,
+ words: opts.words,
});
pool.markSuccess(worker.id);
return outcome;
diff --git a/common/jobs/jobDetail.test.ts b/common/jobs/jobDetail.test.ts
@@ -40,3 +40,27 @@ test("every other kind is unchanged: fetch-window keeps its phrase, the rest not
assert.equal(jobSpecDetail("sync", { kind: "sync", slug: "c", params: { runId: "zzzzzzzz" } }), undefined);
assert.equal(jobSpecDetail(undefined, undefined), undefined);
});
+
+const spec = (kind: string, params: Record<string, unknown>) => ({ kind, slug: "c", params });
+
+test("a single window names who asked and for which clip", () => {
+ assert.equal(
+ jobSpecDetail("fetch-window", spec("fetch-window", { requestedBy: "umtool", manifest: "m", clipId: "c01" })),
+ "umtool · m/c01",
+ );
+});
+
+test("a batch names who asked, for what, and how many windows", () => {
+ assert.equal(
+ jobSpecDetail("fetch-windows", spec("fetch-windows", { requestedBy: "reports", manifest: "demo-site", items: [{}, {}] })),
+ "reports · demo-site · 2 windows",
+ );
+ assert.equal(
+ jobSpecDetail("fetch-windows", spec("fetch-windows", { requestedBy: "umtool", items: [{}] })),
+ "umtool · 1 window",
+ );
+});
+
+test("other kinds add nothing", () => {
+ assert.equal(jobSpecDetail("sync", spec("sync", { requestedBy: "x" })), undefined);
+});
diff --git a/common/jobs/jobDetail.ts b/common/jobs/jobDetail.ts
@@ -27,9 +27,16 @@ export function jobSpecDetail(
if (!runId || !target) return undefined;
return `run ${runId.slice(-6)} · ${target}`;
}
- if (kind !== "fetch-window") return undefined;
+ if (kind !== "fetch-window" && kind !== "fetch-windows") return undefined;
const by = s(p.requestedBy);
if (!by) return undefined;
+ if (kind === "fetch-windows") {
+ // A batch: who asked, for what, and how many windows.
+ const n = Array.isArray(p.items) ? p.items.length : 0;
+ return [by, s(p.manifest), `${n} window${n === 1 ? "" : "s"}`]
+ .filter(Boolean)
+ .join(" · ");
+ }
const what = [s(p.manifest), s(p.clipId)].filter(Boolean).join("/");
return what ? `${by} · ${what}` : by;
}
diff --git a/common/jobs/jobKinds.test.ts b/common/jobs/jobKinds.test.ts
@@ -91,8 +91,12 @@ const ADDED_KINDS: Record<string, { label: string; drainable: boolean }> = {
"reports-prepare": { label: "Prepare report media", drainable: false },
// A report site's exports: one pass, cancelled rather than drained.
"reports-export": { label: "Export reports", drainable: false },
+ // A batch of clip windows: stops between windows, a re-run resumes.
+ "fetch-windows": { label: "Fetch windows", drainable: true },
// One video's metadata re-read: one spawn, cancelled rather than drained.
"refresh-metadata": { label: "Refresh metadata", drainable: false },
+ // One file through a local worker (pnpm ops transcribe): one engine run.
+ "transcribe-file": { label: "Transcribe file", drainable: false },
};
test("added kinds carry their pinned label and drainability", () => {
@@ -149,6 +153,7 @@ const TEXT_KINDS = [
"normalize-transcripts",
"purge-superseded-auto-subs",
"fetch-window",
+ "fetch-windows",
"evict-clips",
"metadata-scan",
"refresh-metadata",
diff --git a/common/jobs/jobKinds.ts b/common/jobs/jobKinds.ts
@@ -307,6 +307,23 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
needsMedia: false,
needsText: true,
},
+ // A LIST OF CLIP WINDOWS, one platform's, as one paced job
+ // (controller/fetchWindows.ts): a site's missing evidence, a umtool
+ // manifest's timeline, an explicit list. One job per platform queue, so
+ // YouTube and Rumble run side by side, each paced on its own. Drainable: it
+ // stops between windows and a re-run fetches only what is still missing,
+ // which is also why it is replayable. It spans channels and so starts with no
+ // channelSlug — the per-channel text check `needsText` stands for is made by
+ // the controller, once per channel, before that channel's first window.
+ "fetch-windows": {
+ kind: "fetch-windows",
+ label: "Fetch windows",
+ drainable: true,
+ replayable: true,
+ queueKeyStrategy: "platform",
+ needsMedia: false,
+ needsText: true,
+ },
"redownload-archive": {
kind: "redownload-archive",
label: "Archive source video",
@@ -863,6 +880,20 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
queueKeyStrategy: "custom",
needsMedia: true,
},
+ // ONE ARBITRARY FILE (or a window of it) through a local worker, for a quote
+ // check — `pnpm ops transcribe` (controller/transcribeFile.ts). No channel:
+ // it reads the file it was given and writes only to a scratch dir and the
+ // caller's `out`, never into the corpus, so neither guard applies. Parallel
+ // (queueKey ""): the worker pool serialises it against every other
+ // transcription. One engine run — cancelled, not drained; not replayable,
+ // since the file it names may be gone by the time anyone retries.
+ "transcribe-file": {
+ kind: "transcribe-file",
+ label: "Transcribe file",
+ drainable: false,
+ replayable: false,
+ queueKeyStrategy: "parallel",
+ },
"download-one-pipeline": {
kind: "download-one-pipeline",
drainable: false,
diff --git a/common/jobs/registry.ts b/common/jobs/registry.ts
@@ -22,7 +22,11 @@ export type JobProgressMetric =
// The metadata scan. Its progress CANNOT be re-counted from disk — the scan
// deliberately writes no video directory — so its runner always sets
// `current` itself. See ytdlp/metadataScan.ts.
- | "scans";
+ | "scans"
+ // Clip windows of a fetch-windows batch (controller/fetchWindows.ts). Spans
+ // channels, so there is no channel count to re-count; the runner always sets
+ // `current`.
+ | "clips";
export type JobProgress = {
metric: JobProgressMetric;
diff --git a/common/jobs/workerPool.test.ts b/common/jobs/workerPool.test.ts
@@ -460,3 +460,42 @@ test("legacy { background: true } maps to the background tier", async () => {
await b;
assert.deepEqual(order, ["F", "B"]);
});
+
+// opts.only (WorkerFilter): the one-off file transcription names a worker, or
+// keeps to local ones. The filter narrows the free-list and the parked grant;
+// a filter no configured worker passes is refused at once rather than parked.
+test("an acquire narrowed by `only` takes the named worker, even when a better one is free", async () => {
+ const pool = new WorkerPool();
+ pool.reconfigure([worker("gpu", 0), worker("cpu", 5)], { applyEnabled: true });
+ const lease = await pool.acquire(undefined, { only: (w) => w.id === "cpu" });
+ assert.equal(lease.worker.id, "cpu");
+ lease.release();
+});
+
+test("an acquire narrowed by `only` parks for its worker and skips others that free", async () => {
+ const pool = new WorkerPool();
+ pool.reconfigure([worker("gpu", 0), worker("cpu", 5)], { applyEnabled: true });
+ const gpu = await pool.acquire();
+ const cpu = await pool.acquire();
+ assert.equal(gpu.worker.id, "gpu");
+ let got: string | null = null;
+ const p = pool
+ .acquire(undefined, { only: (w) => w.id === "cpu" })
+ .then((l) => ((got = l.worker.id), l));
+ gpu.release(); // the wrong worker frees: the narrowed waiter stays parked
+ await new Promise((r) => setImmediate(r));
+ assert.equal(got, null);
+ cpu.release();
+ const l = await p;
+ assert.equal(l.worker.id, "cpu");
+ l.release();
+});
+
+test("an `only` filter no configured worker passes is refused, not parked", async () => {
+ const pool = new WorkerPool();
+ pool.reconfigure([worker("w1")], { applyEnabled: true });
+ await assert.rejects(
+ pool.acquire(undefined, { only: (w) => w.id === "nope" }),
+ /no configured worker can take this work/,
+ );
+});
diff --git a/common/jobs/workerPool.ts b/common/jobs/workerPool.ts
@@ -101,8 +101,16 @@ type Waiter = {
// FIRST waiter it matches, not blindly to the head — pump() scans past a
// waiter whose requirement this slot cannot satisfy.
requires?: readonly string[];
+ // A further restriction on WHICH worker may serve this waiter (acquire's
+ // opts.only). Matched alongside `requires`, never instead of it.
+ only?: WorkerFilter;
};
+// Narrows an acquire to particular workers — by id, by kind — on top of the
+// capability rule. The one-off file transcription uses it to stay on LOCAL
+// workers, or on the one worker the caller named.
+export type WorkerFilter = (worker: Worker) => boolean;
+
export class WorkerPool {
// Insertion order is the tiebreak for equal priority, so use a Map (ordered).
private entries = new Map<string, PoolEntry>();
@@ -316,21 +324,29 @@ export class WorkerPool {
return this.pausedSnapshot !== null;
}
- private eligible(entry: PoolEntry, requires?: readonly string[]): boolean {
+ private eligible(
+ entry: PoolEntry,
+ requires?: readonly string[],
+ only?: WorkerFilter,
+ ): boolean {
return (
entry.state === "enabled" &&
!entry.degraded &&
!entry.busy &&
- workerMatches(entry.config, requires)
+ workerMatches(entry.config, requires) &&
+ (!only || only(entry.config))
);
}
// Highest-priority free entry matching the requirement:
// (priority asc, insertion order asc).
- private pickFree(requires?: readonly string[]): PoolEntry | null {
+ private pickFree(
+ requires?: readonly string[],
+ only?: WorkerFilter,
+ ): PoolEntry | null {
let best: PoolEntry | null = null;
for (const entry of this.entries.values()) {
- if (!this.eligible(entry, requires)) continue;
+ if (!this.eligible(entry, requires, only)) continue;
if (!best || entry.config.priority < best.config.priority) best = entry;
}
return best;
@@ -365,7 +381,7 @@ export class WorkerPool {
let granted = false;
for (let i = 0; i < this.waiters.length; i++) {
const waiter = this.waiters[i];
- const entry = this.pickFree(waiter.requires);
+ const entry = this.pickFree(waiter.requires, waiter.only);
if (!entry) continue;
this.waiters.splice(i, 1);
waiter.onAbort?.();
@@ -405,12 +421,15 @@ export class WorkerPool {
// granted only on a worker that matches it. A requirement NO configured worker
// matches rejects immediately — parking on it would never resolve, and the
// caller should treat the work as unrunnable rather than wedge.
+ // opts.only narrows the lease further, to the workers the filter accepts (see
+ // WorkerFilter); a filter no configured worker passes rejects the same way.
acquire(
signal?: AbortSignal,
opts?: {
tier?: SchedulerTier;
background?: boolean;
requires?: readonly string[];
+ only?: WorkerFilter;
},
): Promise<Lease> {
this.ensureInit();
@@ -418,13 +437,14 @@ export class WorkerPool {
return Promise.reject(abortError());
}
const requires = opts?.requires;
- if (requires && requires.length > 0) {
+ const only = opts?.only;
+ if ((requires && requires.length > 0) || only) {
let satisfiable = false;
for (const e of this.entries.values()) {
// Config-level, ignoring busy/disabled/degraded on purpose: a matching
// worker that is merely busy or switched off is a reason to park, not
// to give up.
- if (workerMatches(e.config, requires)) {
+ if (workerMatches(e.config, requires) && (!only || only(e.config))) {
satisfiable = true;
break;
}
@@ -432,19 +452,21 @@ export class WorkerPool {
if (!satisfiable) {
return Promise.reject(
new Error(
- `no configured worker matches requirement [${requires.join(", ")}]`,
+ requires && requires.length > 0
+ ? `no configured worker matches requirement [${requires.join(", ")}]`
+ : "no configured worker can take this work",
),
);
}
}
- const entry = this.pickFree(requires);
+ const entry = this.pickFree(requires, only);
if (entry && this.waiters.length === 0) {
return Promise.resolve(this.grant(entry));
}
const tier: SchedulerTier =
opts?.tier ?? (opts?.background ? "background" : "foreground");
return new Promise<Lease>((resolve, reject) => {
- const waiter: Waiter = { resolve, reject, tier, requires };
+ const waiter: Waiter = { resolve, reject, tier, requires, only };
if (signal) {
const onAbort = () => {
const idx = this.waiters.indexOf(waiter);
@@ -493,7 +515,13 @@ export class WorkerPool {
}
if (!best) return null;
const claimed = best;
- if (this.waiters.some((w) => workerMatches(claimed.config, w.requires))) {
+ if (
+ this.waiters.some(
+ (w) =>
+ workerMatches(claimed.config, w.requires) &&
+ (!w.only || w.only(claimed.config)),
+ )
+ ) {
return null;
}
return this.grant(claimed);
diff --git a/common/lib/citations/citations.test.ts b/common/lib/citations/citations.test.ts
@@ -245,6 +245,16 @@ test("extractCiteRefs: every cite link in order, with labels and offsets; other
assert.deepEqual(extractCiteRefs("[empty](cite:)").map((r) => r.id), [""]);
});
+test("extractCiteRefs: a label may hold an editorial [insertion], never a lone ]", () => {
+ const md = "She said [“the plan is not what [the mayor] signed”](cite:p01), then [x ] y](cite:p02) and [a [b] [c] d](cite:p03).";
+ const refs = extractCiteRefs(md);
+ assert.deepEqual(refs.map((r) => [r.id, r.label]), [
+ ["p01", "“the plan is not what [the mayor] signed”"],
+ ["p03", "a [b] [c] d"],
+ ]);
+ assert.equal(md.slice(refs[0].offset, refs[0].offset + 2), "[“");
+});
+
test("extractCiteRefs: a cite link in code is text about the syntax", () => {
const md = [
"Write `[label](cite:id)` to cite, like [this](cite:c01).",
diff --git a/common/lib/citations/inline.ts b/common/lib/citations/inline.ts
@@ -28,9 +28,11 @@ export type CiteRef = {
offset: number;
};
-// `[label](cite:id)`. The label may not contain `]`; the id runs to the first
-// `)` or space (an empty id is still a ref, so validation can name it).
-const CITE_LINK_RE = /\[([^\]]*)\]\(\s*cite:([^)\s]*)\s*\)/g;
+// `[label](cite:id)`. The label may hold one level of balanced brackets — an
+// editorial insertion in a quote, `[“told [the mayor] so”](cite:id)`, which
+// markdown-to-jsx links too — but not a lone `]`; the id runs to the first `)`
+// or space (an empty id is still a ref, so validation can name it).
+const CITE_LINK_RE = /\[((?:[^[\]]|\[[^[\]]*\])*)\]\(\s*cite:([^)\s]*)\s*\)/g;
// Fenced blocks: a line opening with ``` or ~~~ (up to three spaces in) to
// the line closing it with at least as many of the same, or the end.
diff --git a/common/lib/report/docs.ts b/common/lib/report/docs.ts
@@ -32,7 +32,9 @@ export function renderReportMarkdown(): string {
"that directory. A site's `reports` list in `site.json` is the published, " +
"ordered list — see [SITE.md](SITE.md); a report directory it does not name is a " +
"draft. The schema is `common/lib/report/schema.ts`; its citations and sources " +
- "are the citation model's — see [CITATIONS.md](CITATIONS.md).",
+ "are the citation model's — see [CITATIONS.md](CITATIONS.md). A `notes.json` " +
+ "beside it holds the operator's notes on the article (umtool, `umtool notes`; " +
+ "[umtool/docs/notes.md](umtool/docs/notes.md)) and is never published.",
);
out.push("");
out.push(
@@ -48,8 +50,9 @@ export function renderReportMarkdown(): string {
"source sentence, a `cite:` link, the subject), a section or claim id used " +
"twice (they share one namespace: the report page's anchors), a sweep's claim " +
"with a verdict, `updated` before `published`, and every citation problem " +
- "CITATIONS.md lists. Whether a still exists and whether a quote matches its cues " +
- "are checked when the site is composed.",
+ "CITATIONS.md lists. Whether a still or the report's video exists, whether the " +
+ "video fits the publish limit, and whether a quote matches its cues are checked " +
+ "when the site is composed.",
);
out.push("");
out.push(REGENERATE_SCHEMA_DOC);
diff --git a/common/lib/report/report.test.ts b/common/lib/report/report.test.ts
@@ -94,7 +94,7 @@ test("every reference must resolve: listed citations, source sentences, cite lin
k1.citations = ["c01", "c99", "c01"];
k1.sourceQuote = { citation: "c01" }; // not a source citation
k1.findings = "See [gone](cite:c98) and [blank](cite:).";
- f.summary = "Also [missing](cite:zz).";
+ f.summary = "Also [missing](cite:zz) and [“told [the mayor] so”](cite:zy).";
const ps = validateReport(f);
assert.deepEqual(paths(ps).sort(), [
"sections[0].claims[0].citations[1]",
@@ -104,8 +104,10 @@ test("every reference must resolve: listed citations, source sentences, cite lin
"sections[0].claims[0].sourceQuote.citation",
"subject.source",
"summary",
+ "summary",
]);
assert.ok(ps.some((p) => /\[gone\]\(cite:c98\) names no citation/.test(p.message)));
+ assert.ok(ps.some((p) => /cite:zy\) names no citation/.test(p.message)));
assert.ok(ps.some((p) => /lists "c01" twice/.test(p.message)));
assert.ok(ps.some((p) => /names a video citation/.test(p.message)));
});
@@ -256,3 +258,15 @@ test("a claim's flag: one line of at most 60 characters", () => {
assert.deepEqual(paths(validateReport(withFlag(" "))), ["sections[0].claims[0].flag"]);
assert.deepEqual(paths(validateReport(withFlag("two\nlines"))), ["sections[0].claims[0].flag"]);
});
+
+test("a report's video: an mp4 and an image poster, relative to the report, and a one-line caption", () => {
+ const withVideo = (video: unknown) => ({ ...structuredClone(fixture()), video });
+ assert.deepEqual(validateReport(withVideo({ src: "video.mp4" })), []);
+ assert.deepEqual(validateReport(withVideo({ src: "media/cut.mp4", poster: "media/cut.jpg", caption: "The cut" })), []);
+ assert.deepEqual(paths(validateReport(withVideo({ src: "video.webm" }))), ["video.src"]);
+ assert.deepEqual(paths(validateReport(withVideo({ src: "../video.mp4" }))), ["video.src"]);
+ assert.deepEqual(paths(validateReport(withVideo({ src: "https://x.test/v.mp4" }))), ["video.src"]);
+ assert.deepEqual(paths(validateReport(withVideo({ src: "video.mp4", poster: "poster.gif" }))), ["video.poster"]);
+ assert.deepEqual(paths(validateReport(withVideo({ src: "video.mp4", caption: "two\nlines" }))), ["video.caption"]);
+ assert.equal(parseReport(withVideo({ src: "video.mp4", autoplay: true })).ok, false);
+});
diff --git a/common/lib/report/schema.ts b/common/lib/report/schema.ts
@@ -76,6 +76,7 @@ export const reportSchema = z.strictObject({
published: text.optional(),
updated: text.optional(),
subject: z.strictObject({ source: text }).optional(),
+ video: z.strictObject({ src: text, poster: text.optional(), caption: text.optional() }).optional(),
verdicts: z.partialRecord(verdict, verdictOverride).optional(),
sources: z.record(text, sourceSchema).optional(),
citations: z.record(text, citationSchema).optional(),
@@ -111,6 +112,8 @@ export const REPORT_FIELD_DOCS: FieldDocs<Report> = {
updated: "When it was last changed, in the same form; not before `published`.",
subject:
"The document under review, when the report reviews one: `{ \"source\": \"<id>\" }`, an id in `sources`.",
+ video:
+ "A video of the report, shown at the head of its page under the title: `{ \"src\": \"video.mp4\", \"poster\": \"poster.jpg\", \"caption\": \"…\" }`. `src` is an mp4 and `poster` an image (png, jpg or webp), both relative to the report's directory; the caption is one line. Absent = none.",
verdicts: `Overrides of the shared verdict vocabulary's labels and colours, by verdict (${VERDICTS.map((v) => `\`${v}\``).join(", ")}): \`{ "label": "…", "color": "#rrggbb" }\`, each key optional. Absent = the shared defaults.`,
sources: "The documents the report's `source` citations quote, by id — see [CITATIONS.md](CITATIONS.md). Absent = none.",
citations:
diff --git a/common/lib/report/validate.ts b/common/lib/report/validate.ts
@@ -18,7 +18,7 @@
// claim), and every `[label](cite:<id>)` link in the summary, the bodies
// and the findings.
//
-// What this cannot check is the disk and the corpus — that a still exists, that
+// What this cannot check is the disk and the corpus — that a still or the video exists, that
// a quote matches its cues; compose does those, where both are at hand.
import {
@@ -26,6 +26,7 @@ import {
isDateTime,
jsonPath,
problem,
+ relativePathProblem,
zodProblems,
type Parsed,
type PathSegment,
@@ -66,6 +67,21 @@ function reportProblems(report: Report, opts: ReportValidateOptions): Problem[]
if (report.subtitle !== undefined && /[\r\n]/.test(report.subtitle)) {
out.push(problem(["subtitle"], "must be one line"));
}
+ if (report.video) {
+ const { src, poster, caption } = report.video;
+ const srcProblem = relativePathProblem(src);
+ if (srcProblem) out.push(problem(["video", "src"], srcProblem));
+ else if (!/\.mp4$/i.test(src)) out.push(problem(["video", "src"], "must be an mp4"));
+ if (poster !== undefined) {
+ const posterProblem = relativePathProblem(poster);
+ if (posterProblem) out.push(problem(["video", "poster"], posterProblem));
+ else if (!/\.(png|jpe?g|webp)$/i.test(poster)) out.push(problem(["video", "poster"], "must be a png, jpg or webp"));
+ }
+ if (caption !== undefined) {
+ if (blank(caption)) out.push(problem(["video", "caption"], "must not be blank"));
+ else if (/[\r\n]/.test(caption)) out.push(problem(["video", "caption"], "must be one line"));
+ }
+ }
for (const key of ["published", "updated"] as const) {
const v = report[key];
if (v !== undefined && !isReportDate(v)) {
diff --git a/common/lib/report/views.test.ts b/common/lib/report/views.test.ts
@@ -78,7 +78,7 @@ const report: Report = {
{
id: "two",
title: "Two",
- body: "A body citing [a post](cite:p1).",
+ body: "A body citing [“a [quoted] post”](cite:p1).",
claims: [{ id: "c2", text: "Claim two.", verdict: "PARTLY", citations: ["a1"] }],
},
],
@@ -407,3 +407,11 @@ test("what the check found: ruled claims by verdict, changed first, confirmed la
long.sections[0].claims![0].gist = "x".repeat(241);
assert.ok(validateReport(long).some((p) => p.path === "sections[0].claims[0].gist"));
});
+
+test("a report's video rides its page view as site-root paths; a report without one carries none", () => {
+ const withVideo = { ...report, video: { src: "video.mp4", poster: "media/poster.jpg", caption: "The cut" } };
+ assert.deepEqual(validateReport(withVideo), []);
+ const v = buildReportPageView(withVideo, { record: (c) => record(c.channel, c.id) });
+ assert.deepEqual(v.video, { src: `/reports/${report.id}/video.mp4`, poster: `/reports/${report.id}/media/poster.jpg`, caption: "The cut" });
+ assert.equal("video" in view, false);
+});
diff --git a/common/lib/report/views.ts b/common/lib/report/views.ts
@@ -274,6 +274,8 @@ export type ReportPageView = {
series?: string;
title: string;
subtitle?: string;
+ // A video of the report, shown under the title: site-root paths.
+ video?: ReportVideoView;
summary?: string;
// How it was checked (markdown, cites nothing).
method?: string;
@@ -298,6 +300,8 @@ export type ReportPageView = {
history?: ReportHistoryRef;
};
+export type ReportVideoView = { src: string; poster?: string; caption?: string };
+
// What a report page offers to download: its exports (html, pdf, md, the
// evidence pack as zip) and its citations (json, csv). Each key is present
// only when the file is published.
@@ -528,6 +532,13 @@ export function buildReportPageView(report: Report, resolve: ReportViewResolver)
series: report.series,
title: report.title,
subtitle: report.subtitle,
+ video: report.video
+ ? defined({
+ src: reportAssetPath(report.id, report.video.src),
+ poster: report.video.poster ? reportAssetPath(report.id, report.video.poster) : undefined,
+ caption: report.video.caption,
+ })
+ : undefined,
summary: report.summary,
method: report.method,
published: report.published,
diff --git a/common/lib/transcriptionApps.test.ts b/common/lib/transcriptionApps.test.ts
@@ -0,0 +1,27 @@
+import { test } from "node:test";
+import assert from "node:assert/strict";
+import { getTranscriptionApp } from "./transcriptionApps";
+
+const input = { audioFile: "audio.wav", outputBase: "transcript.tmp", config: { model: "/m.gguf" } };
+
+test("parakeet keeps its words only when asked", () => {
+ const app = getTranscriptionApp("parakeet");
+ assert.ok(!app.build(input).argv.includes("--words"));
+ const asked = app.build({ ...input, words: true }).argv;
+ assert.ok(asked.includes("--words"));
+ // The audio file stays the last argument: the wrapper reads it positionally.
+ assert.equal(asked.at(-1), "audio.wav");
+});
+
+test("an engine with no words to give ignores the ask", () => {
+ for (const id of ["whisper.cpp", "chough"]) {
+ let app;
+ try {
+ app = getTranscriptionApp(id);
+ } catch {
+ continue;
+ }
+ if (app.id !== id) continue;
+ assert.deepEqual(app.build({ ...input, words: true }).argv, app.build(input).argv);
+ }
+});
diff --git a/common/lib/transcriptionApps.ts b/common/lib/transcriptionApps.ts
@@ -73,6 +73,11 @@ export type TranscribeBuildInput = {
audioFile: string; // basename relative to videoDir
outputBase: string; // e.g. "transcript.tmp-<pid>" (NO extension)
config: AppInstanceConfig;
+ // Ask the engine to keep its word timestamps in the output document (a
+ // top-level `words` array) as well as the grouped cues. Only parakeet has
+ // them to give; the others ignore it. Off for corpus transcriptions, whose
+ // transcript.json would otherwise carry every word twice.
+ words?: boolean;
};
export type TranscriptionApp = {
@@ -88,6 +93,10 @@ export type TranscriptionApp = {
};
// Default binary when the per-app `bin` override is empty.
defaultBin: () => string;
+ // The model this config runs with, after the app's own default — what a
+ // record of "which model produced this" should say. Undefined when the app
+ // picks one itself (chough with no model set).
+ resolveModel: (config: AppInstanceConfig) => string | undefined;
build: (input: TranscribeBuildInput) => TranscribeBuild;
makeProgressParser: () => { feed: (line: string) => ProgressUpdate | null };
// True when the engine can be stopped mid-run and still produce a usable
@@ -175,12 +184,13 @@ const whisperCpp: TranscriptionApp = {
label: "whisper.cpp (whisper-cli)",
fields: { model: true, customArgs: true },
defaultBin: () => getPaths().whisperBin,
+ resolveModel: (config) => config.model ?? getPaths().whisperModel,
build({ audioFile, outputBase, config }) {
const template =
config.customArgs && config.customArgs.length > 0
? config.customArgs
: DEFAULT_TRANSCRIBE_ARGS;
- const model = config.model ?? getPaths().whisperModel;
+ const model = whisperCpp.resolveModel(config) as string;
const argv = substitutePlaceholders(template, {
audioFile,
outputBase,
@@ -197,6 +207,7 @@ const chough: TranscriptionApp = {
label: "chough",
fields: { model: true, remoteUrl: true, chunkSize: true },
defaultBin: () => process.env.CHOUGH_BIN ?? "chough",
+ resolveModel: (config) => config.model?.trim() || undefined,
build({ audioFile, outputBase, config }) {
const argv = ["-f", "json", "-o", outputBase];
if (typeof config.chunkSize === "number" && config.chunkSize > 0) {
@@ -208,7 +219,7 @@ const chough: TranscriptionApp = {
argv.push("-r");
env.CHOUGH_URL = remoteUrl;
}
- const model = config.model?.trim();
+ const model = chough.resolveModel(config);
if (model) env.CHOUGH_MODEL = model;
argv.push(audioFile);
// chough writes EXACTLY the `-o` path — no ".json" is appended.
@@ -237,10 +248,12 @@ const parakeet: TranscriptionApp = {
// a partial transcript on SIGTERM — so a long run is interruptible.
supportsPartialStop: true,
defaultBin: () => getPaths().parakeetBin,
- build({ audioFile, outputBase, config }) {
+ resolveModel: (config) => config.model?.trim() || getPaths().parakeetModel,
+ build({ audioFile, outputBase, config, words }) {
const paths = getPaths();
- const model = config.model?.trim() || paths.parakeetModel;
+ const model = parakeet.resolveModel(config) as string;
const argv = ["--model", model, "--output", outputBase];
+ if (words) argv.push("--words");
if (typeof config.chunkSize === "number" && config.chunkSize > 0) {
argv.push("--segment", String(Math.floor(config.chunkSize)));
}
diff --git a/common/publish/composeReports.test.ts b/common/publish/composeReports.test.ts
@@ -54,6 +54,8 @@ const VIDEOS = "demo-channel";
const SOCIAL = "demo-social";
const REPORT = "demo-report";
const NOW_FLOOR = new Date().toISOString();
+// What a report's notes.json says: if it shows up in anything built, a note leaked.
+const NOTES_SENTINEL = "operator-note-never-published-7f3a";
const writeJson = (file: string, value: unknown) => {
mkdirSync(path.dirname(file), { recursive: true });
@@ -172,6 +174,8 @@ function seedSite(siteId: string, extra: Record<string, unknown> = {}, withRepor
writeJson(path.join(dir, "report.json"), report());
writeText(path.join(dir, "stills", "a01.png"), "png-bytes");
writeText(path.join(dir, "sources", "s0", "page.html"), "<p>saved copy, never published</p>");
+ // umtool's operator notes live beside report.json and are never published.
+ writeJson(path.join(dir, "notes.json"), { format: "umtool-notes", version: 1, notes: [{ text: NOTES_SENTINEL }] });
seedMedia(siteId);
}
@@ -507,6 +511,39 @@ test("a citation without prepared media fails compose with the list, unless --al
}
});
+test("a report's video: copied beside its page and named in the view; a missing or oversize one fails compose", async () => {
+ const dir = path.join(paths.sitesDir, "cited", "reports", REPORT);
+ const file = path.join(dir, "report.json");
+ writeJson(file, { ...report(), video: { src: "video.mp4", poster: "poster.jpg", caption: "The cut" } });
+ try {
+ await assert.rejects(compose("cited"), (e: unknown) => {
+ assert.ok(e instanceof ComposeReportsError);
+ assert.deepEqual(e.problems.map((p) => [p.kind, p.report]), [["report-video", REPORT], ["report-video", REPORT]]);
+ assert.match(e.message, /video\.mp4 does not exist/);
+ return true;
+ });
+ writeText(path.join(dir, "video.mp4"), "mp4-bytes");
+ writeText(path.join(dir, "poster.jpg"), "jpg-bytes");
+ await compose("cited");
+ assert.equal(readFileSync(pub("reports", REPORT, "video.mp4"), "utf8"), "mp4-bytes");
+ assert.equal(readFileSync(pub("reports", REPORT, "poster.jpg"), "utf8"), "jpg-bytes");
+ const view = readJson<{ video?: unknown }>(pub("reports", REPORT, "page.json"));
+ assert.deepEqual(view.video, { src: `/reports/${REPORT}/video.mp4`, poster: `/reports/${REPORT}/poster.jpg`, caption: "The cut" });
+ writeFileSync(path.join(dir, "video.mp4"), Buffer.alloc(25 * 1024 * 1024));
+ await assert.rejects(compose("cited"), (e: unknown) => {
+ assert.ok(e instanceof ComposeReportsError);
+ assert.deepEqual(e.problems.map((p) => p.kind), ["report-video"]);
+ assert.match(e.message, /over the publish limit/);
+ return true;
+ });
+ } finally {
+ writeJson(file, report());
+ rmSync(path.join(dir, "video.mp4"), { force: true });
+ rmSync(path.join(dir, "poster.jpg"), { force: true });
+ await compose("cited");
+ }
+});
+
test("a clip prepared for another span is stale: the reports changed since prepare", async () => {
const sidecar = path.join(reportMediaDir(paths, "cited"), `${CLIP}.json`);
const saved = readFileSync(sidecar, "utf8");
@@ -753,3 +790,27 @@ test("reports export: no browser skips the PDF with a note; no zip fails the pac
assert.equal(await exportMain({ siteId: "cited", formats: ["zip"], zipBin: path.join(ROOT, "no-such-zip") }, quiet), 1);
assert.equal(await exportMain({ siteId: "cited", formats: ["html", "md"] }, quiet), 0);
});
+
+test("a report's notes.json (umtool's operator notes) is never published, exported or committed to its history", async () => {
+ const { reportHistoryGitDir } = await import("./reportHistory");
+ const { execFileSync } = await import("node:child_process");
+ const notes = path.join(paths.sitesDir, "cited", "reports", REPORT, "notes.json");
+ assert.ok(existsSync(notes));
+ await exportSiteReports({ siteId: "cited", paths, now: () => new Date("2026-10-08T08:00:00Z"), openPdfPrinter: fakePrinter([]), onLog: () => {} });
+ await compose("cited");
+ const leaks = (root: string) =>
+ filesUnder(root).filter((f) => f.endsWith("notes.json") || readFileSync(path.join(root, f)).includes(NOTES_SENTINEL));
+ assert.deepEqual(leaks(paths.exportPublicDir), []);
+ assert.deepEqual(leaks(reportExportDir(paths, "cited", REPORT)), []);
+ const gitDir = reportHistoryGitDir(paths, "cited", REPORT);
+ const revs = execFileSync("git", ["--git-dir", gitDir, "rev-list", "--all"], { encoding: "utf8" }).trim().split("\n").filter(Boolean);
+ assert.ok(revs.length > 0);
+ for (const rev of revs) {
+ const names = execFileSync("git", ["--git-dir", gitDir, "ls-tree", "-r", "--name-only", rev], { encoding: "utf8" }).trim().split("\n");
+ assert.deepEqual(names.filter((n) => n.includes("notes")), [], rev);
+ for (const n of names) {
+ const blob = execFileSync("git", ["--git-dir", gitDir, "cat-file", "blob", `${rev}:${n}`]);
+ assert.ok(!blob.includes(NOTES_SENTINEL), `${rev}:${n}`);
+ }
+ }
+});
diff --git a/common/publish/composeReports.ts b/common/publish/composeReports.ts
@@ -57,6 +57,7 @@
import { copyFile, mkdir, readdir, readFile, rm, stat, writeFile } from "node:fs/promises";
import path from "node:path";
+import { publishFileSizeProblem } from "../lib/builtExport";
import type { Paths } from "../lib/paths";
import { siteChannelSlugs, type Site } from "../lib/site";
import { isCitedSite } from "../lib/siteSchema";
@@ -167,6 +168,7 @@ export type ComposeReportsProblemKind =
| "quote-drift"
| "missing-post"
| "missing-still"
+ | "report-video"
| "missing-media"
| "stale-media";
@@ -481,6 +483,16 @@ export async function resolveSiteReports(opts: ResolveSiteReportsOptions): Promi
// record is read.
const refusedChannel = new Set<string>();
for (const report of reports) {
+ // The report's video and its poster: on disk, and small enough to publish.
+ if (report.video) {
+ const dir = siteReportDir(paths, site.siteId, report.id);
+ for (const rel of [report.video.src, report.video.poster]) {
+ if (!rel) continue;
+ const st = await stat(path.join(dir, rel)).catch(() => null);
+ const why = !st?.isFile() ? `the video file ${rel} does not exist` : publishFileSizeProblem(rel, st.size);
+ if (why) problems.push({ kind: "report-video", report: report.id, message: why });
+ }
+ }
const all = report.citations ?? {};
// Only what the report cites: a citation it defines but never cites is
// not in its view, and is neither checked nor published.
@@ -836,6 +848,13 @@ export async function composeReports(opts: ComposeReportsOptions): Promise<Compo
await writeOut(publicDir, reportViewPath(report.id), json(view));
await writeOut(publicDir, reportCitationsDownloadPath(report.id, "json"), json(citationSet(report, view)));
await writeOut(publicDir, reportCitationsDownloadPath(report.id, "csv"), citationsCsv(view));
+ if (report.video && view.video) {
+ const dir = siteReportDir(paths, site.siteId, report.id);
+ await copyOut(publicDir, path.join(dir, report.video.src), view.video.src);
+ if (report.video.poster && view.video.poster) {
+ await copyOut(publicDir, path.join(dir, report.video.poster), view.video.poster);
+ }
+ }
for (const c of orderedCitations(view)) {
if (c.kind !== "source" || !c.image) continue;
const rel = (report.citations![c.id] as { image: string }).image;
diff --git a/common/publish/missingEvidenceWindows.test.ts b/common/publish/missingEvidenceWindows.test.ts
@@ -0,0 +1,131 @@
+// missingEvidenceWindows over a temp site: which cited spans are on disk, which
+// are windows to fetch, and which no window fetch can fill — through the real
+// report checker and the real tier lookup (ffprobe over a few frames of lavfi).
+//
+// Run with: pnpm --filter yt-dlp-transcript-common exec tsx --test publish/missingEvidenceWindows.test.ts
+
+import { after, test } from "node:test";
+import assert from "node:assert/strict";
+import { execFileSync } from "node:child_process";
+import { mkdirSync, mkdtempSync, rmSync, writeFileSync } from "node:fs";
+import { tmpdir } from "node:os";
+import path from "node:path";
+
+const ROOT = mkdtempSync(path.join(tmpdir(), "missing-evidence-"));
+Object.assign(process.env, {
+ TRANSCRIPTS_DIR: path.join(ROOT, "transcripts"),
+ SAVED_VIDEOS_DIR: path.join(ROOT, "saved-videos"),
+ SITES_DIR: path.join(ROOT, "transcripts", "sites"),
+ SETTINGS_FILE: path.join(ROOT, "settings.json"),
+ EXPORT_PUBLIC_DIR: path.join(ROOT, "public"),
+ EXPORT_INDEX_DIR: path.join(ROOT, ".export-index"),
+ EXPORT_BUILDS_DIR: path.join(ROOT, ".export-builds"),
+ ARCHILYZER_CONFIG_DIR: path.join(ROOT, "config"),
+});
+after(() => rmSync(ROOT, { recursive: true, force: true }));
+
+const { getPaths } = await import("../lib/paths");
+const { missingEvidenceWindows } = await import("./reportMedia");
+
+const paths = getPaths();
+const SITE = "demo-site";
+const CH = "demo-channel";
+const OFF = "off-site";
+
+const writeJson = (file: string, value: unknown) => {
+ mkdirSync(path.dirname(file), { recursive: true });
+ writeFileSync(file, JSON.stringify(value, null, 2));
+};
+const videoDir = (slug: string, id: string) => path.join(paths.channelsDir, slug, "data", id);
+
+for (const [slug, extra] of [
+ [CH, {}],
+ [OFF, {}],
+] as const) {
+ writeJson(path.join(paths.channelsDir, slug, "config.json"), {
+ handling: "transcribe",
+ name: slug,
+ url: "https://www.youtube.com/@demo",
+ ...extra,
+ });
+}
+// A fetched window holding 0–6 s of abc123.
+mkdirSync(path.join(videoDir(CH, "abc123"), "clips"), { recursive: true });
+execFileSync("ffmpeg", [
+ "-nostdin", "-v", "error", "-y",
+ "-f", "lavfi", "-i", "testsrc=size=160x90:rate=10:duration=6",
+ "-f", "lavfi", "-i", "sine=frequency=440:duration=6",
+ "-c:v", "libx264", "-preset", "ultrafast", "-pix_fmt", "yuv420p", "-c:a", "aac", "-shortest",
+ path.join(videoDir(CH, "abc123"), "clips", "0.00-6.00.mp4"),
+]);
+// A video the source says is gone.
+writeJson(path.join(videoDir(CH, "gone1"), "availability.json"), {
+ checkedAt: "2026-10-01T00:00:00.000Z",
+ availability: "deleted",
+});
+
+writeJson(path.join(paths.sitesDir, SITE, "site.json"), {
+ title: "Demo",
+ channels: [{ slug: CH }],
+ search: false,
+ reports: ["r1", "r2"],
+});
+const report = (id: string, citations: Record<string, unknown>) => ({
+ format: "archilyzer-report",
+ version: 1,
+ id,
+ kind: "sweep",
+ title: `Report ${id}`,
+ citations,
+ sections: [
+ { id: "s1", title: "One", body: Object.keys(citations).map((c) => `[${c}](cite:${c})`).join(" ") },
+ ],
+});
+writeJson(
+ path.join(paths.sitesDir, SITE, "reports", "r1", "report.json"),
+ report("r1", {
+ c01: { kind: "video", channel: CH, id: "abc123", start: 1, end: 2, quote: "on disk" },
+ c02: { kind: "video", channel: CH, id: "missing1", start: 10.004, end: 12.001, pad: { before: 1, after: 1 }, quote: "the one to fetch" },
+ c03: { kind: "video", channel: CH, id: "gone1", start: 5, end: 6, quote: "removed" },
+ c04: { kind: "video", channel: OFF, id: "elsewhere", start: 5, end: 6, quote: "not this site's" },
+ }),
+);
+writeJson(
+ path.join(paths.sitesDir, SITE, "reports", "r2", "report.json"),
+ report("r2", {
+ d01: { kind: "video", channel: CH, id: "missing1", start: 10.004, end: 12.001, pad: { after: 3 }, quote: "again" },
+ }),
+);
+
+test("on disk, to fetch, and what no window can fill", async () => {
+ const r = await missingEvidenceWindows(paths, SITE);
+ assert.equal(r.siteId, SITE);
+ assert.equal(r.onDisk, 1);
+ assert.deepEqual(r.problems, []);
+
+ // One window for the moment two reports cite, with the wider pad on each
+ // side (from the moment's rounded start and end), and the two decimals a
+ // window is named by.
+ assert.deepEqual(r.missing, [
+ {
+ slug: CH,
+ id: "missing1",
+ from: 9,
+ to: 15,
+ clipId: "r1#c02",
+ citations: ["r1#c02", "r2#d01"],
+ reason: "the one to fetch",
+ },
+ ]);
+
+ const why = Object.fromEntries(r.unfetchable.map((u) => [u.id, u.why]));
+ assert.deepEqual(why, {
+ gone1: "gone",
+ elsewhere: "not-in-site",
+ });
+ assert.match(r.unfetchable.find((u) => u.id === "gone1")!.message, /deleted/);
+});
+
+test("a site that does not exist is an error", async () => {
+ await assert.rejects(missingEvidenceWindows(paths, "no-such-site"), /no site "no-such-site"/);
+});
diff --git a/common/publish/reportMedia.ts b/common/publish/reportMedia.ts
@@ -34,13 +34,15 @@
//
// A RUN WITH PROBLEMS STILL WRITES THE MANIFEST (the editor lists the problems
// from it) and FAILS: a citation without media, a clip over the size limit, an
-// invalid or missing report. Nothing here fetches — the editor's fetch-window
-// and persist (for spans) and Capture posts (for posts) fill what is missing.
+// invalid or missing report. Nothing here fetches — the editor's fetch-windows
+// (for spans: `missingEvidenceWindows` below says which), persist, and Capture
+// posts (for posts) fill what is missing.
//
// AN UNMOUNTED DRIVE IS NOT A MISSING CLIP. A span whose media is not found on
// a channel whose media tier is not reachable (lib/channelMedia.ts) is
// reported as unreachable, with the reason, rather than as missing.
+import type { Dirent } from "node:fs";
import { readdir, rm } from "node:fs/promises";
import path from "node:path";
import { readJsonFile, writeJsonAtomic } from "../lib/jsonFile-server";
@@ -54,11 +56,15 @@ import { readChannelConfig } from "../controller/channels";
import { momentKey, momentOf, momentProblem, type Moment } from "../lib/citations/moments";
import type { CitationPad } from "../lib/citations/schema";
import { parseReport } from "../lib/report/validate";
-import type { Report } from "../lib/report/schema";
+import { isReportId, type Report } from "../lib/report/schema";
+import { isPermanentlyGone } from "../lib/availability";
+import { loadAvailability } from "../lib/availability-server";
+import { MAX_CLIP_WINDOW_SECONDS } from "../lib/clipWindow";
import {
citedEvidenceSpan,
isAudioOnlyPlatform,
prepareEvidenceClip,
+ resolveEvidenceSource,
widerPad,
type EvidenceKind,
type EvidenceMedia,
@@ -88,6 +94,23 @@ export function siteReportFile(paths: Paths, siteId: string, reportId: string):
return path.join(siteReportDir(paths, siteId, reportId), "report.json");
}
+// Every report directory of a site, published or draft: a directory under
+// `sites/<siteId>/reports/` whose name is a report id. Anything else there (a
+// stray file, a bad name) is not a report. Sorted. Whether one is PUBLISHED is
+// the site's `reports` list (site.json), not anything on disk.
+export async function listReportDirs(paths: Paths, siteId: string): Promise<string[]> {
+ let entries: Dirent[];
+ try {
+ entries = await readdir(path.join(siteDir(paths, siteId), "reports"), { withFileTypes: true });
+ } catch {
+ return [];
+ }
+ return entries
+ .filter((e) => e.isDirectory() && isReportId(e.name))
+ .map((e) => e.name)
+ .sort((a, b) => a.localeCompare(b));
+}
+
export type ReportMediaEntry =
| EvidenceMedia
| {
@@ -369,3 +392,144 @@ export async function readReportMediaIndex(paths: Paths, siteId: string): Promis
return v as ReportMediaIndex;
}
+
+// ---------------------------------------------------------------------------
+// WHICH WINDOWS A SITE STILL NEEDS — the list the editor's `fetch-windows` job
+// takes, so "fetch what the reports cite" is one request rather than a loop.
+//
+// The same enumeration prepare runs (loadSiteReports, citedMoments,
+// citedEvidenceSpan), and the same question prepare asks first
+// (resolveEvidenceSource) — WITHOUT cutting anything. A span prepare would
+// call `missing-media` comes back in `missing` as a window to fetch; a span
+// that cannot be fetched as a window comes back in `unfetchable` with the
+// reason, so a dry run tells the operator what no fetch will fix:
+//
+// not-in-site the channel is not one of the site's (prepare's not-in-site)
+// unreachable the channel's media is on a drive that is not answering —
+// not missing, and a fetch would write beside the wrong thing
+// gone availability.json says deleted, private or members-only
+// audio-only a feed record: its media is the episode's audio, which a
+// download fetches; a window selector has no picture to pick
+// too-long longer than one window may be — persist the video instead
+//
+// Spans already on disk count in `onDisk`. Posts are not windows and are left
+// to Capture posts.
+// ---------------------------------------------------------------------------
+
+export type EvidenceWindow = {
+ slug: string;
+ id: string;
+ from: number;
+ to: number;
+ // `<reportId>#<citationId>` of the first citation of the moment.
+ clipId: string;
+ // Every citation of the moment.
+ citations: string[];
+ // The first citation's quote, for the window's provenance.
+ reason?: string;
+};
+
+export type UnfetchableReason = "not-in-site" | "unreachable" | "gone" | "audio-only" | "too-long";
+
+export type UnfetchableEvidenceWindow = EvidenceWindow & {
+ why: UnfetchableReason;
+ message: string;
+};
+
+export type MissingEvidenceWindows = {
+ siteId: string;
+ missing: EvidenceWindow[];
+ unfetchable: UnfetchableEvidenceWindow[];
+ // Cited spans whose media is already on disk.
+ onDisk: number;
+ // The reports that could not be read (prepare's missing/invalid-report).
+ problems: ReportMediaProblem[];
+};
+
+// A window fetched for a span must hold it: its name is two decimals, so the
+// start rounds down and the end up.
+const floor2 = (n: number) => Math.floor(n * 100 + 1e-6) / 100;
+const ceil2 = (n: number) => Math.ceil(n * 100 - 1e-6) / 100;
+
+export async function missingEvidenceWindows(paths: Paths, siteId: string): Promise<MissingEvidenceWindows> {
+ if (!listSiteIds(paths).includes(siteId)) {
+ throw new Error(`no site "${siteId}" (no sites/${siteId}/site.json)`);
+ }
+ const site = getSite(siteId, paths);
+ const { reports, problems } = await loadSiteReports(paths, site);
+ const quotes = new Map<string, string>();
+ for (const report of reports) {
+ for (const [cid, c] of Object.entries(report.citations ?? {})) {
+ if (typeof c.quote === "string" && c.quote.trim()) quotes.set(`${report.id}#${cid}`, c.quote.trim());
+ }
+ }
+ const pool = siteChannelSlugs(site);
+ const configs = new Map<string, ChannelConfig | null>();
+ const configOf = async (slug: string) => {
+ if (!configs.has(slug)) configs.set(slug, await readChannelConfig(paths, slug));
+ return configs.get(slug) ?? null;
+ };
+ const out: MissingEvidenceWindows = { siteId, missing: [], unfetchable: [], onDisk: 0, problems };
+
+ for (const m of citedMoments(reports)) {
+ if (m.kind === "post" || m.moment.kind !== "span") continue;
+ const slug = m.moment.channel;
+ const id = m.moment.id;
+ const span = await citedEvidenceSpan(paths.channelsDir, slug, id, { start: m.moment.start, end: m.moment.end, pad: m.pad });
+ const clipId = m.citedBy[0];
+ const reason = quotes.get(clipId);
+ const window: EvidenceWindow = {
+ slug,
+ id,
+ from: floor2(span.from),
+ to: ceil2(span.to),
+ clipId,
+ citations: m.citedBy,
+ ...(reason ? { reason } : {}),
+ };
+ const unfetchable = (why: UnfetchableReason, message: string) =>
+ out.unfetchable.push({ ...window, why, message });
+
+ if (!pool.has(slug)) {
+ unfetchable("not-in-site", `channel "${slug}" is not one of this site's channels`);
+ continue;
+ }
+ const config = await configOf(slug);
+ const audioOnly = isAudioOnlyPlatform(config?.platform);
+ const source = await resolveEvidenceSource({
+ channelsDir: paths.channelsDir,
+ slug,
+ id,
+ span,
+ audio: m.kind === "audio" || audioOnly,
+ ffprobeBin: paths.ffprobeBin,
+ });
+ if (source) {
+ out.onDisk += 1;
+ continue;
+ }
+ const where = await inspectChannelMedia(paths, slug, config, { fresh: true }).catch(() => null);
+ if (where && where.status !== "ok" && where.status !== "in-place") {
+ unfetchable("unreachable", `the channel's media is ${where.status}${where.detail ? ` (${where.detail})` : ""}`);
+ continue;
+ }
+ const availability = (await loadAvailability(path.join(paths.channelsDir, slug, "data", id)))?.availability;
+ if (isPermanentlyGone(availability)) {
+ unfetchable("gone", `the source says the video is ${availability}`);
+ continue;
+ }
+ if (audioOnly) {
+ unfetchable("audio-only", "a feed record's media is its audio download, not a window");
+ continue;
+ }
+ if (window.to - window.from > MAX_CLIP_WINDOW_SECONDS) {
+ unfetchable(
+ "too-long",
+ `${Math.round(window.to - window.from)}s is longer than a window may be (${MAX_CLIP_WINDOW_SECONDS}s) — persist the video`,
+ );
+ continue;
+ }
+ out.missing.push(window);
+ }
+ return out;
+}
diff --git a/common/social/xenforoFetcher.test.ts b/common/social/xenforoFetcher.test.ts
@@ -544,3 +544,35 @@ test("downloadable media: files only", () => {
],
);
});
+
+test("capture: downloads what the post's page holds NOW, beside what the archive kept", async () => {
+ const outDir = await mkdtemp(path.join(os.tmpdir(), "forum-capture-live-"));
+ const video = "https://uploads.kiwifarms.st/data/video/9665/9665533-abc.mp4?hash=Zz";
+ const player =
+ `<div class="ephyra-player ephyra-player--video" data-media-type="video" data-duration-label="1:11"` +
+ ` data-source-fallback="${video.replace("https:", "")}" data-filename="clip.mp4"></div>`;
+ const postUrl = `${ORIGIN}/posts/9100/`;
+ const html = threadPage({
+ page: 1,
+ last: 1,
+ posts: [{ id: 9100, author: "Member1", userId: 51, ts: T0, position: 1, body: `for those that want proof:${player}` }],
+ });
+ const { loader, got } = captureLoader({ [postUrl]: html }, { x: 0, y: 0, width: 600, height: 200 });
+ const res = await captureForumPosts(
+ {
+ ids: ["9100"],
+ handle: "t",
+ accountUrl: THREAD_URL,
+ outDir,
+ // Archived by the parser of its day: the video was a bare "1:11", no media.
+ archived: new Map([["9100", { text: "for those that want proof:\n1:11", url: postUrl, media: [{ kind: "image", url: "https://images.example/b.jpg" }] }]]),
+ signal: new AbortController().signal,
+ },
+ { loader, pauseMs: () => 7_000, sleep: async () => {}, now: () => new Date(0) },
+ );
+ assert.deepEqual(res.outcomes.map((o) => [o.id, o.state, o.files]), [["9100", "captured", 3]]);
+ assert.deepEqual(got, [video, "https://images.example/b.jpg"], "the page's video first, then the archive's image");
+ const rec = JSON.parse(await readFile(path.join(outDir, "9100", "capture.json"), "utf8"));
+ assert.equal(rec.mediaState, "ok");
+ assert.deepEqual(rec.media.map((m: { url: string }) => m.url), [video, "https://images.example/b.jpg"]);
+});
diff --git a/common/social/xenforoFetcher.ts b/common/social/xenforoFetcher.ts
@@ -32,7 +32,7 @@
import { mkdir, writeFile } from "node:fs/promises";
import path from "node:path";
import { getPaths } from "../lib/paths";
-import type { Post } from "../lib/posts";
+import type { Post, PostMedia } from "../lib/posts";
import {
registerSocialFetcher,
type PostCaptureInput,
@@ -365,6 +365,10 @@ export type ForumShotResult = {
shot?: CapturedFile;
error?: string;
stop?: string;
+ // The post's media as its page reads NOW, parsed from the page just loaded.
+ // An archived record keeps what the parser of its day saw: before the
+ // player fix, a forum-hosted video was a bare duration and no media.
+ liveMedia?: PostMedia[];
};
export async function shootForumPost(
@@ -391,6 +395,14 @@ export async function shootForumPost(
stop: why.error,
};
}
+ let liveMedia: PostMedia[] | undefined;
+ try {
+ liveMedia = parseXenforoThreadPage(got.html, { channelSlug: "capture", pageUrl: postUrl }).posts.find(
+ (p) => p.id === id,
+ )?.media;
+ } catch {
+ liveMedia = undefined;
+ }
await page.waitForTimeout(1_000);
await page.evaluate(PREPARE_SCRIPT(id)).catch(() => {});
const snap = (await page.evaluate(SNAPSHOT_SCRIPT(id))) as ForumPostSnapshot;
@@ -413,7 +425,7 @@ export async function shootForumPost(
}
await mkdir(dir, { recursive: true });
await writeFile(path.join(dir, SHOT_FILENAME), png);
- return { state: "captured", shot: await describeCapturedFile(dir, SHOT_FILENAME, postUrl) };
+ return { state: "captured", shot: await describeCapturedFile(dir, SHOT_FILENAME, postUrl), ...(liveMedia ? { liveMedia } : {}) };
}
const EXT_BY_TYPE: Record<string, string> = {
@@ -438,12 +450,22 @@ function extFor(url: string, contentType: string | undefined): string {
// carries as files. Embeds and link cards are pages, not files.
export function downloadableMedia(
media: ReadonlyArray<{ kind: string; url: string }> | undefined,
+ ...more: ReadonlyArray<ReadonlyArray<{ kind: string; url: string }> | undefined>
): { kind: string; url: string }[] {
- return (media ?? []).filter(
- (m) => (m.kind === "image" || m.kind === "attachment" || m.kind === "video") && /^https?:\/\//.test(m.url),
- );
+ const out: { kind: string; url: string }[] = [];
+ for (const list of [media, ...more]) {
+ for (const m of list ?? []) {
+ if (!(m.kind === "image" || m.kind === "attachment" || m.kind === "video")) continue;
+ if (!/^https?:\/\//.test(m.url) || out.some((x) => x.url === m.url)) continue;
+ out.push(m);
+ }
+ }
+ return out;
}
+// A video can be hundreds of megabytes; an image is not.
+const MEDIA_TIMEOUT_MS = { video: 600_000, other: 60_000 };
+
export type ForumCaptureDeps = {
loader: ForumPageLoader & { page(): Promise<PageLike> };
pauseMs: () => number;
@@ -496,6 +518,7 @@ export async function captureForumPosts(
archived?.url && /^https?:\/\//.test(archived.url) ? archived.url : `${thread.origin}/posts/${id}/`;
let state: PostCaptureState | undefined = work.shot ? undefined : existing?.state;
+ let liveMedia: PostMedia[] | undefined;
let shot = work.shot ? undefined : existing?.shot;
let error: string | undefined;
let stop: string | undefined;
@@ -516,13 +539,15 @@ export async function captureForumPosts(
shot = res.shot;
error = res.error;
stop = res.stop;
+ liveMedia = res.liveMedia;
}
let mediaState: CaptureMediaState = work.media ? "skipped" : (existing?.mediaState ?? "skipped");
let media: CapturedFile[] = work.media ? [] : (existing?.media ?? []);
const postIsThere = state === undefined || state === "captured";
if (work.media && postIsThere && !stop) {
- const files = downloadableMedia(archived?.media);
+ // What the page holds now first, then anything only the archive kept.
+ const files = downloadableMedia(liveMedia, archived?.media);
if (files.length === 0) {
mediaState = "none";
state ??= "captured";
@@ -533,7 +558,8 @@ export async function captureForumPosts(
await contact();
if (signal.aborted) return stopped("Cancelled; the rest are left for a later run.");
try {
- const res = await page.request!.get(m.url, { timeout: 60_000, failOnStatusCode: false });
+ const timeout = m.kind === "video" ? MEDIA_TIMEOUT_MS.video : MEDIA_TIMEOUT_MS.other;
+ const res = await page.request!.get(m.url, { timeout, failOnStatusCode: false });
if (!res.ok()) {
failed++;
onLog?.(`${id}: media ${n + 1} answered HTTP ${res.status()}.`);
diff --git a/common/social/xenforoParse.test.ts b/common/social/xenforoParse.test.ts
@@ -243,3 +243,21 @@ test("a message with no id or no date is not a post", () => {
.replace(/datetime="[^"]*"/, "");
assert.equal(parseXenforoThreadPage(html, { channelSlug: "teapots" }).posts.length, 0);
});
+
+test("a Kiwi Farms player upload is video media named by its file, not a bare duration", () => {
+ const player =
+ `<div class="ephyra-player ephyra-player--video" id="media-9660049" data-engine="videojs"` +
+ ` data-manifest="/ephyra-stream/9619135/master.m3u8?attachment_id=9660049"` +
+ ` data-source-fallback="//uploads.kiwifarms.st/data/video/9619/9619135-10fa.mp4?hash=lZ6"` +
+ ` data-duration="209.207" data-duration-label="3:29" data-attachment-id="9660049"` +
+ ` data-media-type="video" data-filename="Teapot Intro [aLfKTn4x7q8].mp4">` +
+ `<button type="button" class="ephyra-lazy-poster"><img src="/ephyra-stream/9619135/poster.avif" alt=""></button>` +
+ `<span class="ephyra-duration">3:29</span></div>`;
+ const html = threadPage({ page: 1, last: 1, posts: [post(700, 7, `Before it goes:${player}Watch it.`)] });
+ const [p] = parseXenforoThreadPage(html, { channelSlug: "teapots" }).posts;
+ assert.deepEqual(p.media, [
+ { kind: "video", url: "https://uploads.kiwifarms.st/data/video/9619/9619135-10fa.mp4?hash=lZ6", name: "Teapot Intro [aLfKTn4x7q8].mp4" },
+ ]);
+ assert.ok(p.text.includes("[video: Teapot Intro [aLfKTn4x7q8].mp4 (3:29)]"), p.text);
+ assert.ok(!/^3:29$/m.test(p.text), "the duration label is not a line of its own");
+});
diff --git a/common/social/xenforoParse.ts b/common/social/xenforoParse.ts
@@ -410,6 +410,20 @@ function render(node: HtmlNode, ctx: BodyCtx, out: string[]): void {
return;
}
+ // Kiwi Farms' own player ("ephyra"): a div carrying the upload as data
+ // attributes, its <video> built by script. The file is the fallback source;
+ // without this branch only the player's duration label ("1:11") reached the
+ // text, and a forum-hosted clip — often the only copy — was not media.
+ if ((el.attrs["data-media-type"] === "video" || el.attrs["data-media-type"] === "audio") &&
+ el.attrs["data-source-fallback"]) {
+ const url = resolveForumUrl(el.attrs["data-source-fallback"], ctx.origin);
+ const name = el.attrs["data-filename"];
+ const label = el.attrs["data-duration-label"];
+ if (url) addMedia(ctx, { kind: "video", url, ...(name ? { name } : {}) });
+ out.push(SOFT, `[${el.attrs["data-media-type"]}${name ? `: ${name}` : ""}${label ? ` (${label})` : ""}]`, SOFT);
+ return;
+ }
+
if (el.tag === "video" || el.tag === "audio") {
const src =
el.attrs.src ??
diff --git a/common/views/jobRows.ts b/common/views/jobRows.ts
@@ -167,7 +167,9 @@ function computeJobProgressView(
now: number,
): JobRowView["progress"] {
const snap = job.progress;
- if (!snap || !stat) return undefined;
+ // A runner-reported `current` needs no channel stat: a batch spanning
+ // channels (fetch-windows) has none, and still knows its own count.
+ if (!snap || (!stat && snap.current === undefined)) return undefined;
// `current` is RE-COUNTED from disk (readChannelStat), never reported by the
// runner — which is why the digest metric needed its own on-disk counter
// (digestCount) rather than a number the batch could have just told us.
@@ -179,10 +181,10 @@ function computeJobProgressView(
const current =
snap.current ??
(snap.metric === "downloads"
- ? stat.downloadCount
+ ? stat!.downloadCount
: snap.metric === "digests"
- ? (stat.digestCount ?? 0)
- : stat.transcriptCount);
+ ? (stat!.digestCount ?? 0)
+ : stat!.transcriptCount);
const range = Math.max(0, snap.target - snap.initial);
const advance = Math.max(0, current - snap.initial);
const pct =
@@ -224,7 +226,7 @@ export function fromRecord(j: JobRecord, ctx: FromRecordContext): JobRowView {
channelSlug: j.channelSlug,
videoId: j.videoId,
progress:
- j.status === "running" && j.channelSlug
+ j.status === "running" && (j.channelSlug || j.progress?.current !== undefined)
? computeJobProgressView(j, ctx.stat, ctx.now)
: undefined,
tasks: j.tasks?.map((t) => ({
diff --git a/common/ytdlp/channelArgs.test.ts b/common/ytdlp/channelArgs.test.ts
@@ -3,6 +3,7 @@ import assert from "node:assert/strict";
import {
channelExtraArgs,
channelPaceSeconds,
+ configForVideoUrl,
pacedPlatformArgs,
platformArgs,
platformArgsForUrl,
@@ -198,3 +199,20 @@ test("one-off imports have their own floor: Odysee's 60 s, BitChute's batch floo
assert.equal(platformImportMinGapSeconds(p), 0, String(p));
}
});
+
+test("a video's own URL picks the platform args when the channel names none or another", () => {
+ const notes = cfg({ handling: "transcribe" } as Partial<ChannelConfig>);
+ assert.deepEqual(channelExtraArgs(notes), []);
+ assert.deepEqual(
+ channelExtraArgs(configForVideoUrl(notes, "https://rumble.com/v7em13s-x.html")),
+ RUMBLE,
+ );
+ const yt = cfg({ url: "https://www.youtube.com/@x", ytdlpExtraArgs: ["--limit-rate", "1M"] });
+ assert.deepEqual(
+ channelExtraArgs(configForVideoUrl(yt, "https://rumble.com/v7em13s-x.html")),
+ [...RUMBLE, "--limit-rate", "1M"],
+ "the channel's own args still apply",
+ );
+ assert.equal(configForVideoUrl(yt, "https://www.youtube.com/watch?v=abc"), yt);
+ assert.equal(configForVideoUrl(yt, "https://example.com/v.mp4"), yt);
+});
diff --git a/common/ytdlp/channelArgs.ts b/common/ytdlp/channelArgs.ts
@@ -53,6 +53,18 @@ export function channelPlatform(
return config.platform ?? detectPlatform(config.url);
}
+// The config to build a VIDEO's argv from: the channel's, with the platform
+// taken from the video's own URL when that names a different one. A channel
+// may hold another platform's videos — community-notes has no url or platform
+// of its own and holds Rumble videos — and a Rumble request without Rumble's
+// args (`--impersonate`) is refused at Cloudflare. The request goes where the
+// URL points, so the URL's platform decides the args and the pace.
+export function configForVideoUrl(config: ChannelConfig, videoUrl: string): ChannelConfig {
+ const platform = detectPlatform(videoUrl);
+ if (!platform || platform === channelPlatform(config)) return config;
+ return { ...config, platform, url: videoUrl };
+}
+
// The key the pacing state is kept under — the same one the rate-limit
// cooldown uses (`detectPlatform(url) ?? "unknown"`, autoRunner/runYtdlp), so a
// 429 recorded by any path paces every path.
diff --git a/common/ytdlp/fetchWindowManaged.test.ts b/common/ytdlp/fetchWindowManaged.test.ts
@@ -0,0 +1,124 @@
+// fetchWindowManaged against a fake yt-dlp: Rumble's `.tar` HLS segments are
+// refused by ffmpeg until the one retry adds -extension_picky 0, and the option
+// is never passed on a first try (against a progressive URL it is an error) —
+// except to a rumble.com URL, which gets it first and loses it on a retry.
+//
+// Run with: pnpm --filter yt-dlp-transcript-common exec tsx --test ytdlp/fetchWindowManaged.test.ts
+
+import { test } from "node:test";
+import assert from "node:assert/strict";
+import { chmod, mkdir, mkdtemp, readFile, writeFile } from "node:fs/promises";
+import os from "node:os";
+import path from "node:path";
+import { fetchWindowManaged, HLS_PICKY_RETRY_ARGS } from "./fetchWindowManaged";
+import type { Paths } from "../lib/paths";
+import type { ChannelConfig } from "../lib/channelConfig";
+
+const ROOT = await mkdtemp(path.join(os.tmpdir(), "fetchwindow-"));
+const BIN = path.join(ROOT, "fake-ytdlp.mjs");
+const ARGS_LOG = path.join(ROOT, "argv.jsonl");
+process.env.FAKE_ARGS_LOG = ARGS_LOG;
+
+// FAKE_MODE=rumble: refuses like ffmpeg 8 on Rumble's segments unless the
+// picky option is present. =progressive: fails if the option IS present, as
+// ffmpeg does against a non-HLS input. =broken: always fails, unrelated.
+await writeFile(
+ BIN,
+ `#!/usr/bin/env node
+import { appendFileSync, writeFileSync } from "node:fs";
+const args = process.argv.slice(2);
+appendFileSync(process.env.FAKE_ARGS_LOG, JSON.stringify(args) + "\\n");
+const picky = args.includes("ffmpeg_i:-extension_picky 0");
+const mode = process.env.FAKE_MODE;
+const fail = (msg) => { process.stderr.write(msg + "\\nERROR: ffmpeg exited with code 183\\n"); process.exit(1); };
+if (mode === "broken") fail("ERROR: something else entirely");
+if (mode === "rumble" && !picky) fail("[in#0] URL https://cdn.example/x.tar?r_file=media-0.ts is not in allowed_segment_extensions");
+if (mode === "progressive" && picky) fail("Option extension_picky not found.");
+writeFileSync(args[args.indexOf("-o") + 1], "mp4");
+`,
+);
+await chmod(BIN, 0o755);
+
+async function setup(name: string) {
+ const videoDir = path.join(ROOT, name, "data", "v1");
+ await mkdir(videoDir, { recursive: true });
+ return videoDir;
+}
+
+async function runs(): Promise<string[][]> {
+ return (await readFile(ARGS_LOG, "utf8").catch(() => ""))
+ .trim().split("\n").filter(Boolean).map((l) => JSON.parse(l) as string[]);
+}
+
+function opts(videoDir: string) {
+ return {
+ channelSlug: "c",
+ channelConfig: {} as ChannelConfig,
+ paths: { ytdlpBin: BIN } as unknown as Paths,
+ videoDir,
+ videoId: "v1",
+ videoUrl: "https://rumble.example/v1.html",
+ from: 10,
+ to: 20,
+ provenance: { requestedBy: "test" },
+ onLog: () => {},
+ signal: new AbortController().signal,
+ };
+}
+
+const hasPicky = (argv: string[]) => argv.join(" ").includes(HLS_PICKY_RETRY_ARGS.join(" "));
+
+test("a Rumble window refused for its segment extension is retried once with -extension_picky 0", async () => {
+ process.env.FAKE_MODE = "rumble";
+ const before = (await runs()).length;
+ const videoDir = await setup("rumble");
+ const res = await fetchWindowManaged(opts(videoDir));
+ const mine = (await runs()).slice(before);
+ assert.equal(mine.length, 2);
+ assert.equal(hasPicky(mine[0]), false, "never on the first try");
+ assert.equal(hasPicky(mine[1]), true);
+ assert.equal(res.cached, false);
+ assert.ok(res.provenance?.ytdlp === undefined);
+ assert.equal(await readFile(res.file, "utf8"), "mp4");
+});
+
+test("a progressive source fetches on the first try, without the HLS option", async () => {
+ process.env.FAKE_MODE = "progressive";
+ const before = (await runs()).length;
+ const res = await fetchWindowManaged(opts(await setup("progressive")));
+ const mine = (await runs()).slice(before);
+ assert.equal(mine.length, 1);
+ assert.equal(hasPicky(mine[0]), false);
+ assert.equal(res.cached, false);
+});
+
+test("any other failure is not retried with the HLS option", async () => {
+ process.env.FAKE_MODE = "broken";
+ const before = (await runs()).length;
+ await assert.rejects(fetchWindowManaged(opts(await setup("broken"))), /yt-dlp failed fetching/);
+ const mine = (await runs()).slice(before);
+ assert.equal(mine.length, 1);
+});
+
+const atRumble = (videoDir: string) => ({ ...opts(videoDir), videoUrl: "https://rumble.com/v1-x.html" });
+
+test("a rumble.com window sends -extension_picky 0 on the first try: one page load, not two", async () => {
+ process.env.FAKE_MODE = "rumble";
+ const before = (await runs()).length;
+ const res = await fetchWindowManaged(atRumble(await setup("rumble-first")));
+ const mine = (await runs()).slice(before);
+ assert.equal(mine.length, 1);
+ assert.equal(hasPicky(mine[0]), true);
+ assert.equal(await readFile(res.file, "utf8"), "mp4");
+});
+
+test("a rumble.com upload served progressive is retried once without the HLS option", async () => {
+ process.env.FAKE_MODE = "progressive";
+ const before = (await runs()).length;
+ const res = await fetchWindowManaged(atRumble(await setup("rumble-progressive")));
+ const mine = (await runs()).slice(before);
+ assert.equal(mine.length, 2);
+ assert.equal(hasPicky(mine[0]), true);
+ assert.equal(hasPicky(mine[1]), false);
+ assert.equal(await readFile(res.file, "utf8"), "mp4");
+});
diff --git a/common/ytdlp/fetchWindowManaged.ts b/common/ytdlp/fetchWindowManaged.ts
@@ -18,6 +18,8 @@ import {
AUTH_RETRY_CLASSES,
classifyDownloadFailure,
parseUnavailableFromStderr,
+ type Availability,
+ type DownloadFailureClass,
} from "../lib/availability";
import {
DEFAULT_COOKIE_MODE,
@@ -26,6 +28,7 @@ import {
type ResolvedCookiePolicy,
} from "../lib/cookiePolicy";
import type { ChannelConfig } from "../lib/channelConfig";
+import { detectPlatform } from "../lib/platform";
import type { Paths } from "../lib/paths";
import {
clipsDirFor,
@@ -39,7 +42,7 @@ import {
writeClipProvenance,
type ClipWindow,
} from "../lib/clipWindow-server";
-import { channelExtraArgs } from "./channelArgs";
+import { channelExtraArgs, configForVideoUrl } from "./channelArgs";
import { runOneYtdlp } from "./runOneYtdlp";
import { FULL_LOG_PROGRESS_ARGS } from "./downloadOneManaged";
import { clipFormatSelector } from "./downloadFormat";
@@ -49,10 +52,39 @@ import { clipFormatSelector } from "./downloadFormat";
// above 720 are thrown away after paying for them.
export const DEFAULT_CLIP_MAX_HEIGHT = 720;
+// Rumble serves HLS whose segments are named `.tar`, and ffmpeg 8 refuses them
+// ("URL … is not in allowed_segment_extensions", exit 183) — every Rumble window
+// failed. `-extension_picky 0` lets them through, but it is an option of the HLS
+// DEMUXER: against a progressive URL (YouTube's googlevideo mp4) ffmpeg aborts
+// with "Option extension_picky not found". So it is a RETRY on exactly that
+// refusal — umtool's build-video.mjs does the same (umtool/docs/quirks.md) —
+// EXCEPT on a rumble.com URL, where it goes on the first try: every attempt
+// loads the Rumble page again, and Cloudflare 403s some share of those loads,
+// so a retry every Rumble window needs doubled the refusals. An old Rumble
+// upload served progressive is the mirror case, retried once without it.
+export const HLS_EXTENSION_REFUSED = /allowed_segment_extensions|allowed_extensions/;
+export const HLS_PICKY_REFUSED = /Option extension_picky not found/;
+export const HLS_PICKY_RETRY_ARGS = ["--downloader-args", "ffmpeg_i:-extension_picky 0"];
+
// The clip format selector lives with the download presets (the "video_720"
// preset is built from it); re-exported so existing importers keep working.
export { clipFormatSelector };
+// A fetch that yt-dlp refused, with the refusal already classified. The
+// single-window job only shows the message; the batch (controller/fetchWindows)
+// reads `failureClass` to decide between carrying on, backing the platform off
+// and stopping — without parsing the message a second time.
+export class FetchWindowError extends Error {
+ constructor(
+ message: string,
+ readonly failureClass: DownloadFailureClass,
+ readonly availability: Availability,
+ ) {
+ super(message);
+ this.name = "FetchWindowError";
+ }
+}
+
export type FetchWindowProvenance = {
requestedBy: string;
manifest?: string;
@@ -162,7 +194,7 @@ export async function fetchWindowManaged(
await rm(part, { force: true });
const maxHeight = opts.maxHeight ?? DEFAULT_CLIP_MAX_HEIGHT;
- const argsWith = (cookies: string | undefined): string[] => [
+ const argsWith = (cookies: string | undefined, retryArgs: string[] = []): string[] => [
// The operator's own yt-dlp config redirects output and attaches thumbnail
// and metadata post-processors; without this the window lands elsewhere —
// and a metadata post-processor is exactly what must not run here.
@@ -183,7 +215,9 @@ export async function fetchWindowManaged(
"--merge-output-format",
"mp4",
...FULL_LOG_PROGRESS_ARGS,
- ...channelExtraArgs(opts.channelConfig, cookies),
+ // The video's URL picks the platform args: a channel may hold another
+ // platform's videos (configForVideoUrl).
+ ...channelExtraArgs(configForVideoUrl(opts.channelConfig, opts.videoUrl), cookies),
// THE NEGATIONS COME AFTER THE CHANNEL'S OWN ARGS, AND -o AFTER THOSE.
//
// `ytdlpExtraArgs` is free text an operator typed into a form; it is
@@ -205,13 +239,14 @@ export async function fetchWindowManaged(
"--no-write-subs",
"--no-write-auto-subs",
"--no-download-archive",
+ ...retryArgs,
"-o",
part,
"--",
opts.videoUrl,
];
- const run = async (cookies: string | undefined) =>
+ const run = async (cookies: string | undefined, retryArgs: string[] = []) =>
runOneYtdlp(
{
ytdlpBin: opts.paths.ytdlpBin,
@@ -219,11 +254,14 @@ export async function fetchWindowManaged(
signal: opts.signal,
},
opts.cwd ?? videoDir,
- argsWith(cookies),
+ argsWith(cookies, retryArgs),
);
- let args = argsWith(alwaysCookies(policy));
- let outcome = await run(alwaysCookies(policy));
+ let cookiesUsed = alwaysCookies(policy);
+ let picky = detectPlatform(opts.videoUrl) === "rumble" ? HLS_PICKY_RETRY_ARGS : [];
+ let args = argsWith(cookiesUsed, picky);
+ let outcome = await run(cookiesUsed, picky);
+ let backedOff = false;
if (outcome.exitCode !== 0) {
const availability = parseUnavailableFromStderr(outcome.stderrTail);
@@ -233,24 +271,53 @@ export async function fetchWindowManaged(
// Sync both honour. Recorded before the throw so the NEXT request is
// refused at the door rather than re-storming the source.
await opts.onPlatformBackoff?.("rate_limit");
+ backedOff = true;
} else if (AUTH_RETRY_CLASSES.has(availability)) {
const retryCookies = authRetryCookies(policy);
if (retryCookies) {
opts.onLog(
`Window fetch failed with ${availability}; retrying once with cookies.\n`,
);
- args = argsWith(retryCookies);
- outcome = await run(retryCookies);
+ cookiesUsed = retryCookies;
+ args = argsWith(retryCookies, picky);
+ outcome = await run(retryCookies, picky);
}
}
}
+ if (outcome.exitCode !== 0 && !picky.length && HLS_EXTENSION_REFUSED.test(outcome.stderrTail)) {
+ opts.onLog(
+ `Window fetch: ffmpeg refused the HLS segment extension; retrying once with -extension_picky 0.\n`,
+ );
+ await rm(part, { force: true });
+ picky = HLS_PICKY_RETRY_ARGS;
+ args = argsWith(cookiesUsed, picky);
+ outcome = await run(cookiesUsed, picky);
+ } else if (outcome.exitCode !== 0 && picky.length && HLS_PICKY_REFUSED.test(outcome.stderrTail)) {
+ opts.onLog(
+ `Window fetch: the source is not HLS after all; retrying once without -extension_picky.\n`,
+ );
+ await rm(part, { force: true });
+ picky = [];
+ args = argsWith(cookiesUsed, picky);
+ outcome = await run(cookiesUsed, picky);
+ }
+
if (outcome.exitCode !== 0) {
await rm(part, { force: true });
const tail = outcome.stderrTail.trim().split("\n").slice(-4).join(" / ");
- throw new Error(
+ // Classified from the LAST attempt, which is the one that decided.
+ const availability = parseUnavailableFromStderr(outcome.stderrTail);
+ const failureClass = classifyDownloadFailure(outcome.stderrTail, availability);
+ // A cookie retry that ran into a 429 is a 429 like any other.
+ if (failureClass === "rate_limit" && !backedOff) {
+ await opts.onPlatformBackoff?.("rate_limit");
+ }
+ throw new FetchWindowError(
`yt-dlp failed fetching ${from.toFixed(2)}–${to.toFixed(2)} of ` +
`${opts.videoId} (exit ${outcome.exitCode ?? "null"}): ${tail}`,
+ failureClass,
+ availability,
);
}
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -10,6 +10,12 @@
- **Substitute your own yt-dlp in Docker.** Point `YTDLP_BIN` at a zipapp you built, or set `YTDLP_SOURCE_HOST_DIR` to a yt-dlp checkout and start with `docker-compose.ytdlp.yml`: the image runs it with its own python, and nothing is rebuilt. Every editor boot logs `yt-dlp: <path> <version> (image|override)` (`MISSING` when it does not run; the editor still starts), and `YTDLP_AUTO_UPDATE` updates the image's yt-dlp only, warning instead of touching yours.
- **The Docker image can publish.** It carries python, `pipx` and a pinned `git-filter-repo`, so the homepage's `/source` mirror builds in the container; `docker-compose.source.yml` mounts your repository read-only for it, and the scrub rules and denylist live in the config volume (`/data/config/archilyzer`). Cloudflare and R2 credentials come from `.env`. Run publish commands with `docker compose exec editor pnpm archilyzer …`, not `run --rm`. The `homepage` service serves a local deploy from the builds volume once there is one. RUNNING_IN_DOCKER.md has a Windows checklist.
- **`archilyzer doctor` checks what a publish needs.** Which yt-dlp runs (the image's, the host's or an override, and whether it runs), whether the Cloudflare token and the R2 keys are set (never their values; R2 only when a bucket is configured) — judged exactly as a deploy judges them —, the wrangler a deploy runs (the pinned one or your `WRANGLER_BIN`, and that it starts and is the expected major), free space for the site bundles, the publish lock (free, held by a running stage, or left by one that is gone — with the command to clear it; never cleared for you), the index stamp's age and which sites were built from an older one, the repository the source mirror reads, the private config dir, and whether this Node is new enough for the pinned wrangler (deploys need 22).
+- **A file transcription can return word timings.** `pnpm ops transcribe` takes `"words": true` and adds `words: [{w, start, end, conf?}]` to its result, on the source file's clock like the cues. parakeet keeps the word timestamps it stitches from its windows (its wrapper's new `--words` writes them beside the cues); an engine without them answers `[]`. Corpus transcriptions do not ask, so their `transcript.json` is unchanged.
+- **A single file can be transcribed through the editor.** `pnpm ops transcribe` (`POST /api/ops/transcribe`) takes `{"path"}` — an absolute path to a local audio or video file — with `"start"`/`"end"` (seconds) for a window, `"workerId"` (a configured local worker; default: the one auto-transcribe would get) and `"out"` (an absolute path for the result). It runs as a "Transcribe file" job on `/jobs`, through the worker pool and the same engine, model and command line the corpus is transcribed with. A window is cut by ffmpeg to a temporary 16 kHz mono WAV, removed afterwards. The result is `{path, window, worker: {id, name, appId, model, device}, transcriptFormat, cues: [{start, end, text}], text, …}`, cue times on the source file's clock; it ends the job's log, is written to `out` when given, and `--wait` prints it on stdout. Refused before any job: a relative or unreadable path, an `end` not after `start`, an unknown or remote worker, a named worker switched off, an `out` inside the corpus (resolved through symlinks, storage locations included). Nothing is written to the corpus or to settings.
+- **Page titles no longer carry a paragraph of explanation.** /channels used to open with three lines of prose about where its numbers come from, and at a narrow window with the sidebar open its four buttons took the row and squeezed that prose to one word a line. It is now one status line — how many channels, how old the oldest report is, how many have none — and when the buttons do not fit beside the title they drop below it. Workers, Sites, Review and Operations lose their lead paragraph; Tags, Storage and Saved videos keep one line each, the instruction (rule edits apply at the next index build; media moves from a channel's Storage panel; the keep-latest window is in a channel's settings); the Monitor widget builder loses its own, which repeated the board's. Needs a restart of the editor.
+- **Clip windows can be fetched as a list, paced, in one job per platform.** `pnpm ops fetch-windows` (`POST /api/ops/fetch-windows`) takes `{"siteId"}` — every window a site's published reports cite and the disk does not hold — or `{"items": [{"slug", "id", "from", "to", "clipId"?, "reason"?}], "requestedBy", "manifest"?}`, with `"maxHeight"` and `"dryRun"`. A window already on disk is answered at once and joins no job; the rest are grouped by platform queue and fetched as one `fetch-windows` job per platform, so YouTube and Rumble run side by side, each window through the same managed fetch as a single one (cookie policy, auth and HLS retries, provenance). Between two fetches on a platform the job waits that platform's batch gap, at least 30 s and up to half again at random; before each it checks the platform's cooldown and hold and stops when either is set. A 429 backs the platform off and stops the job; one 403 is that window's failure, two in a row back the platform off and stop it. A platform cooling down or held is refused for its group before anything starts. The job is drainable, shows its progress as clip windows, and Retry or running the same body again fetches only what is still missing. A dry run lists the windows per platform, the ones on disk, and the spans no window can fill (a video the source says is deleted, private or members-only, a channel off the site, an unmounted drive). A site's **Reports** tab has **Fetch missing evidence** and **Preview missing evidence** beside **Prepare evidence media**, and umtool's `fetch-via-editor.mjs --all` sends a manifest's whole timeline as one request. A window fetch whose cookie retry runs into a 429 now records the cooldown too. Needs a restart of the editor.
+- **A social channel can be renamed and deleted from its page.** A posts channel's page (X, Bluesky, a forum thread) had no Danger zone, so it could not be renamed or deleted in the editor at all. It now has the one a video channel has, collapsed under the posts panel and opened by the same `?stage=danger` link: **Rename channel** and **Delete channel**, typing the slug to confirm, refused while the channel is busy, in the same sentences. Needs a restart of the editor.
+- **A channel can be created, renamed, deleted and put on sites over `pnpm ops`.** `create-channel` is the New channel form (`{"fields": {"name", "handling", "url", …}}`, the same field names as `channel-config`'s `patch`; `"slug"` and `"sites"` optional; the form's "Fetch playlist now", "Fetch posts now" and "Add to top of auto-queue" are off unless asked for, and a job they start is returned so `--wait` follows it). `rename-channel` (`{"slug", "newSlug"}`) and `delete-channel` (`{"slug", "confirm"}`, `confirm` repeating the slug) are the Danger zone's two forms. `channel-config` takes `"sites"` — the whole membership set, `[]` for on no site; an unknown site is refused rather than skipped — and `"excludeFromBuild"` / `"excludeFromCleanup"`, set to the value given rather than toggled, with or without a `patch`. `pnpm ops get channels` lists every channel with its kind, platform and the sites that carry it. Each runs the form's own action, so it refuses what the form refuses, in the same words. Needs a restart of the editor.
- **One video's metadata can be read again from its source.** A livestream that has just ended offers one fragmented audio format and no captions; hours later the same URL has plain formats and auto-captions, and the video's `metadata.info.json` still said what it said the first time. "Refresh metadata" under the video page's header, and `pnpm ops refresh-metadata --json '{"slug":"…","id":"…"}'`, re-read that one video with no subtitles and no media, on the platform's queue, with the channel's cookie policy and pace; a held platform or one in a rate-limit cooldown is refused, and a rate limit records the cooldown. The rewrite is recorded in the metadata history as `refresh`, and the job's log ends with what the source now says: `live_status`, how many formats, each audio-only format and its protocol, whether any is non-fragmented, the English subtitle and caption tracks, and the keys that changed. Nothing is downloaded or deleted, and no download attempt is recorded. A video not yet fetched into the archive is refused — a refresh never creates its directory — as are an archive.org record, a Wayback copy and a record completed from a podcast feed, whose metadata is not yt-dlp's. Needs a restart of the editor.
- **A renamed social channel's posts open again.** Each archived post carries the channel slug it was fetched under, and renaming the channel moves its directory without rewriting them, so every post of a renamed channel named the old slug: the index filed it under the new one, the post page and its thread could not find it there, and MCP links named a channel that no longer existed. A post's channel and slug are now read from the directory it is stored in, wherever it was fetched; the files are not rewritten. The next index build corrects the published records. Needs a restart of the editor.
- **A video's other English tracks are readable and searchable where their words differ.** Uploaded captions are not always a transcript of what was said, so the tracks beside the transcript stay: the served `en` beside `en-orig`, a regional or auto-translated track, and the captions a local transcription replaced. One is kept where its words differ from the transcript's and from every track kept before it; identical tracks, most of them, add nothing. The index keeps them in an `alts` sub-DB and writes `track` and `altTracks` onto the transcript record only then, so every other record's page is what it was. A search hit in a word only an alternate holds names the track; one every track says is found once, in the transcript. The video page's **Transcript** card reads the transcript and switches tracks ("Track: original audio captions ▾"); switching changes nothing on disk, and **Set as transcript** stays the way the transcript itself changes. English VTTs are no longer shipped as subtitle tracks. One notion of a track — ids, plain labels, which are kept, how a hit across them is found — lives in `common/lib/captionTracks.ts`.
diff --git a/editor/app/api/media/fetch-window/[jobId]/route.ts b/editor/app/api/media/fetch-window/[jobId]/route.ts
@@ -46,7 +46,9 @@ export async function GET(
return NextResponse.json({ error: "no such job" }, { status: 404 });
}
const kind = live?.kind ?? meta?.kind ?? entry?.kind;
- if (kind !== "fetch-window" && kind !== "redownload-archive") {
+ // A fetch-windows batch answers with its status and, on failure, its log's
+ // tail: it names no one file, so `file` stays absent.
+ if (kind !== "fetch-window" && kind !== "redownload-archive" && kind !== "fetch-windows") {
return NextResponse.json(
{ error: `job ${jobId} is a ${kind ?? "?"} job, not a media fetch` },
{ status: 400 },
@@ -59,7 +61,7 @@ export async function GET(
const out: Record<string, unknown> = { status, jobId };
if (exitCode !== undefined) out.exitCode = exitCode;
- if (status === "done") {
+ if (status === "done" && kind !== "fetch-windows") {
const slug = live?.channelSlug ?? meta?.channelSlug ?? spec?.slug;
const videoId = live?.videoId ?? meta?.videoId;
const p = spec?.params ?? {};
diff --git a/editor/app/api/ops/_lib.ts b/editor/app/api/ops/_lib.ts
@@ -2,7 +2,8 @@ import { NextResponse } from "next/server";
import { authorizeWorkerRequest } from "yt-dlp-transcript-common/lib/workerToken";
import { isValidChannelSlug } from "yt-dlp-transcript-common/controller/channels";
import { previewBranchProblem } from "yt-dlp-transcript-common/lib/pagesDeploy";
-import { isValidSiteId } from "yt-dlp-transcript-common/lib/site";
+import { isValidSiteId, listSiteIds } from "yt-dlp-transcript-common/lib/site";
+import { getPaths } from "yt-dlp-transcript-common/lib/paths";
import type { StreamActionResult } from "yt-dlp-transcript-common/jobs/streamCommand";
import type { QueueOutcome } from "../../channels/lib/queueForSlugs";
@@ -146,6 +147,77 @@ export function optBool(body: OpsBody, key: string): boolean | undefined {
return v;
}
+// A channel's site memberships, as the Configure form's Sites section posts
+// them: THE WHOLE SET — a site left out is a site the channel leaves, and `[]`
+// is "on no site". Undefined when absent (memberships untouched).
+//
+// STRICTER THAN THE FORM'S PARSER, ON PURPOSE. planSiteMembershipWrites skips a
+// siteId it does not know, because for the form that means the site was deleted
+// since the page loaded. Over HTTP it means a typo, and skipping it would answer
+// `{ ok: true }` about a membership that was never written — the silent success
+// this file's unknown-key rule exists to refuse. Group ids and new group names
+// are still the planner's to judge.
+export type OpsSiteMembership = {
+ siteId: string;
+ groupId?: string;
+ newGroupName?: string;
+};
+
+export function readSiteMemberships(
+ body: OpsBody,
+ key = "sites",
+): OpsSiteMembership[] | undefined {
+ const v = body[key];
+ if (v === undefined) return undefined;
+ if (!Array.isArray(v)) {
+ throw new OpsInputError(
+ `"${key}" must be an array of { siteId, groupId? | newGroupName? } ([] = on no site)`,
+ );
+ }
+ const known = new Set(listSiteIds(getPaths()));
+ const out: OpsSiteMembership[] = [];
+ for (const entry of v) {
+ if (typeof entry !== "object" || entry === null || Array.isArray(entry)) {
+ throw new OpsInputError(`each "${key}" entry must be an object with a "siteId"`);
+ }
+ const e = entry as Record<string, unknown>;
+ const stray = Object.keys(e).filter(
+ (k) => !["siteId", "groupId", "newGroupName"].includes(k),
+ );
+ if (stray.length) {
+ throw new OpsInputError(
+ `unknown key(s) in a "${key}" entry: ${stray.join(", ")} — accepted: siteId, groupId, newGroupName`,
+ );
+ }
+ if (typeof e.siteId !== "string" || !isValidSiteId(e.siteId)) {
+ throw new OpsInputError(`"${String(e.siteId)}" is not a valid site id`);
+ }
+ if (!known.has(e.siteId)) {
+ throw new OpsInputError(
+ `no site "${e.siteId}" — known: ${[...known].join(", ") || "none"}`,
+ );
+ }
+ for (const k of ["groupId", "newGroupName"] as const) {
+ if (e[k] !== undefined && typeof e[k] !== "string") {
+ throw new OpsInputError(`"${k}" in a "${key}" entry must be a string`);
+ }
+ }
+ if (e.groupId !== undefined && e.newGroupName !== undefined) {
+ throw new OpsInputError(
+ `a "${key}" entry takes "groupId" or "newGroupName", not both`,
+ );
+ }
+ out.push({
+ siteId: e.siteId,
+ ...(e.groupId !== undefined ? { groupId: e.groupId as string } : {}),
+ ...(e.newGroupName !== undefined
+ ? { newGroupName: e.newGroupName as string }
+ : {}),
+ });
+ }
+ return out;
+}
+
// A whole number above zero, or undefined when absent. A type check, not a
// rule: what the number means is the action's business.
export function optPositiveInt(body: OpsBody, key: string): number | undefined {
diff --git a/editor/app/api/ops/channel-config/route.ts b/editor/app/api/ops/channel-config/route.ts
@@ -5,12 +5,29 @@ import {
channelConfigToFormData,
validateChannelFormPatch,
} from "../../../channels/components/channelConfigToForm";
-import { updateChannelAction } from "../../../channels/actions";
-import { actionResponse, OpsInputError, ops, reqSlug } from "../_lib";
+import {
+ setChannelExclusionsAction,
+ updateChannelAction,
+} from "../../../channels/actions";
+import {
+ actionResponse,
+ OpsInputError,
+ ops,
+ optBool,
+ readSiteMemberships,
+ reqSlug,
+} from "../_lib";
export const dynamic = "force-dynamic";
-// POST { slug: string, patch: { <Configure-form field>: string|number|boolean|null } }
+// POST {
+// slug: string,
+// patch?: { <Configure-form field>: string|number|boolean|null },
+// sites?: [{ siteId, groupId? | newGroupName? }],
+// excludeFromBuild?: boolean,
+// excludeFromCleanup?: boolean,
+// }
+// — at least one of patch / sites / excludeFromBuild / excludeFromCleanup.
//
// The patch keys are the FORM's field names, not ChannelConfig's, because the
// form is what validates them: `downloadFilterInclude` / `downloadFilterExclude`
@@ -19,29 +36,68 @@ export const dynamic = "force-dynamic";
// channel's current form representation instead of being written directly.
//
// `""` (or null) clears a field, exactly as clearing the input does.
+//
+// `sites` is the Configure form's Sites section: the WHOLE membership set, so
+// a site left out is left, and `[]` takes the channel off every site. It rides
+// the same save as the patch, as it does in the form.
+//
+// The two exclusions are the /channels rack's build and cleanup toggles, set
+// to a value rather than flipped, since a caller cannot see the switch.
export async function POST(request: Request) {
- return ops(request, ["slug", "patch"], async (body) => {
- const slug = reqSlug(body, "slug");
- const patch = body.patch;
- if (
- typeof patch !== "object" ||
- patch === null ||
- Array.isArray(patch)
- ) {
- throw new OpsInputError('"patch" must be a JSON object');
- }
- // KEYS FIRST, CHANNEL SECOND. A misspelled field is a fact about the
- // request; reporting "channel not found" for it would hide the real error.
- try {
- validateChannelFormPatch(patch as Record<string, unknown>);
- } catch (e) {
- throw new OpsInputError((e as Error).message);
- }
- const paths = getPaths();
- const existing = await readChannelConfig(paths, slug);
- if (!existing) throw new OpsInputError(`Channel "${slug}" not found`);
- const fd = channelConfigToFormData(existing);
- applyChannelFormPatch(fd, patch as Record<string, unknown>);
- return actionResponse(await updateChannelAction(slug, undefined, fd));
- });
+ return ops(
+ request,
+ ["slug", "patch", "sites", "excludeFromBuild", "excludeFromCleanup"],
+ async (body) => {
+ const slug = reqSlug(body, "slug");
+ const patch = body.patch;
+ if (
+ patch !== undefined &&
+ (typeof patch !== "object" || patch === null || Array.isArray(patch))
+ ) {
+ throw new OpsInputError('"patch" must be a JSON object');
+ }
+ const sites = readSiteMemberships(body);
+ const excludeFromBuild = optBool(body, "excludeFromBuild");
+ const excludeFromCleanup = optBool(body, "excludeFromCleanup");
+ if (
+ patch === undefined &&
+ sites === undefined &&
+ excludeFromBuild === undefined &&
+ excludeFromCleanup === undefined
+ ) {
+ throw new OpsInputError(
+ 'nothing to change — give "patch", "sites", "excludeFromBuild" or "excludeFromCleanup"',
+ );
+ }
+ // KEYS FIRST, CHANNEL SECOND. A misspelled field is a fact about the
+ // request; reporting "channel not found" for it would hide the real error.
+ if (patch !== undefined) {
+ try {
+ validateChannelFormPatch(patch as Record<string, unknown>);
+ } catch (e) {
+ throw new OpsInputError((e as Error).message);
+ }
+ }
+ const paths = getPaths();
+ const existing = await readChannelConfig(paths, slug);
+ if (!existing) throw new OpsInputError(`Channel "${slug}" not found`);
+ if (patch !== undefined || sites !== undefined) {
+ const fd = channelConfigToFormData(existing);
+ if (patch !== undefined) {
+ applyChannelFormPatch(fd, patch as Record<string, unknown>);
+ }
+ if (sites !== undefined) {
+ fd.set("siteMembershipsJson", JSON.stringify(sites));
+ }
+ const saved = await updateChannelAction(slug, undefined, fd);
+ if (saved?.error) return actionResponse(saved);
+ }
+ return actionResponse(
+ await setChannelExclusionsAction(slug, {
+ excludeFromBuild,
+ excludeFromCleanup,
+ }),
+ );
+ },
+ );
}
diff --git a/editor/app/api/ops/channels/route.ts b/editor/app/api/ops/channels/route.ts
@@ -0,0 +1,44 @@
+import { NextResponse } from "next/server";
+import { getPaths } from "yt-dlp-transcript-common/lib/paths";
+import { listChannelConfigs } from "yt-dlp-transcript-common/controller/channels";
+import { isSocialChannel } from "yt-dlp-transcript-common/lib/channelConfig";
+import { siteChannelIndex } from "yt-dlp-transcript-common/lib/site";
+import { opsAuth } from "../_lib";
+
+export const dynamic = "force-dynamic";
+
+// GET /api/ops/channels
+//
+// Every channel, one line each: what a caller needs to pick a slug for the
+// other channel routes, and to see which sites carry it — a channel on no site
+// is private to this editor. Config reads only (no report, no media probe, no
+// directory walk); `get channel <slug>` is the deep read of one.
+export async function GET(request: Request) {
+ const denied = opsAuth(request);
+ if (denied) return denied;
+ const paths = getPaths();
+ const [rows, index] = await Promise.all([
+ listChannelConfigs(paths),
+ Promise.resolve(siteChannelIndex(paths)),
+ ]);
+ const sitesOf = new Map<string, string[]>();
+ for (const [siteId, slugs] of Object.entries(index)) {
+ for (const slug of slugs) {
+ sitesOf.set(slug, [...(sitesOf.get(slug) ?? []), siteId]);
+ }
+ }
+ return NextResponse.json({
+ ok: true,
+ channels: rows.map(({ slug, config }) => ({
+ slug,
+ name: config.name ?? slug,
+ kind: isSocialChannel(config) ? "social" : "video",
+ platform: config.platform ?? null,
+ handling: config.handling,
+ url: config.url ?? null,
+ sites: sitesOf.get(slug) ?? [],
+ ...(config.excludeFromBuild ? { excludeFromBuild: true } : {}),
+ ...(config.excludeFromCleanup ? { excludeFromCleanup: true } : {}),
+ })),
+ });
+}
diff --git a/editor/app/api/ops/create-channel/route.ts b/editor/app/api/ops/create-channel/route.ts
@@ -0,0 +1,74 @@
+import { NextResponse } from "next/server";
+import {
+ applyChannelFormPatch,
+ validateChannelFormPatch,
+} from "../../../channels/components/channelConfigToForm";
+import { createChannelFromForm } from "../../../channels/actions";
+import {
+ OpsInputError,
+ opsFail,
+ ops,
+ optBool,
+ readSiteMemberships,
+} from "../_lib";
+
+export const dynamic = "force-dynamic";
+
+// POST {
+// slug?: string, // else derived from fields.name, as the form does
+// fields: { name, handling, url?, platform?, sourceKind?, postFetcher?,
+// socialHandle?, … }, // the New channel form's field names
+// sites?: [{ siteId, groupId? | newGroupName? }], // absent = on no site
+// fetchPlaylist?: boolean, // the form's "Fetch playlist now"
+// fetchPostsNow?: boolean, // the form's "Fetch posts now"
+// prioritizeDownload?: boolean, // the form's "Add to top of auto-queue"
+// }
+//
+// The New channel form, over HTTP. `fields` takes exactly the keys
+// /api/ops/channel-config's `patch` takes (channelConfigToForm.ts), built into
+// the FormData the form would post, and the same parser refuses what the form
+// refuses: no name, a handling that is not "youtube" or "transcribe", a social
+// URL with no derivable handle, a bad filter regex.
+//
+// The three start-something options are OFF unless asked for. The form ticks
+// "Fetch playlist now" by default because a person creating a channel usually
+// wants it; a script creating one says so. Any job they start comes back in
+// `jobIds`, so `--wait` can follow it.
+export async function POST(request: Request) {
+ return ops(
+ request,
+ ["slug", "fields", "sites", "fetchPlaylist", "fetchPostsNow", "prioritizeDownload"],
+ async (body) => {
+ const fields = body.fields;
+ if (typeof fields !== "object" || fields === null || Array.isArray(fields)) {
+ throw new OpsInputError('"fields" must be a JSON object');
+ }
+ try {
+ validateChannelFormPatch(fields as Record<string, unknown>);
+ } catch (e) {
+ throw new OpsInputError((e as Error).message);
+ }
+ const fd = new FormData();
+ applyChannelFormPatch(fd, fields as Record<string, unknown>);
+ if (body.slug !== undefined) {
+ if (typeof body.slug !== "string") {
+ throw new OpsInputError('"slug" must be a string');
+ }
+ fd.set("slug", body.slug.trim());
+ }
+ const sites = readSiteMemberships(body);
+ if (sites) fd.set("siteMembershipsJson", JSON.stringify(sites));
+ for (const flag of ["fetchPlaylist", "fetchPostsNow", "prioritizeDownload"]) {
+ if (optBool(body, flag)) fd.set(flag, "on");
+ }
+ const result = await createChannelFromForm(fd);
+ if (!result.ok) return opsFail(result.error);
+ return NextResponse.json({
+ ok: true,
+ slug: result.slug,
+ jobIds: result.jobIds,
+ ...(result.jobIds.length === 1 ? { jobId: result.jobIds[0] } : {}),
+ });
+ },
+ );
+}
diff --git a/editor/app/api/ops/delete-channel/route.ts b/editor/app/api/ops/delete-channel/route.ts
@@ -0,0 +1,26 @@
+import { NextResponse } from "next/server";
+import { deleteChannelFromForm } from "../../../channels/actions";
+import { opsFail, ops, reqSlug, reqString } from "../_lib";
+
+export const dynamic = "force-dynamic";
+
+// POST { slug: string, confirm: string }
+//
+// The channel page's Danger → Delete, over HTTP: the whole channel directory
+// — every transcript, sidecar, post and media link — goes. `confirm` is the
+// form's "type the slug" box and must equal `slug` exactly; the action is what
+// compares them, and refuses in the form's own sentence. Every other refusal is
+// the form's too: a busy channel, media in transition.
+//
+// There is no undo, and transcripts/ is its own git repo — the only way back
+// is a commit there.
+export async function POST(request: Request) {
+ return ops(request, ["slug", "confirm"], async (body) => {
+ const slug = reqSlug(body, "slug");
+ const fd = new FormData();
+ fd.set("confirmSlug", reqString(body, "confirm"));
+ const result = await deleteChannelFromForm(slug, fd);
+ if (!result.ok) return opsFail(result.error);
+ return NextResponse.json({ ok: true, slug: result.slug, deleted: true });
+ });
+}
diff --git a/editor/app/api/ops/fetch-posts/route.ts b/editor/app/api/ops/fetch-posts/route.ts
@@ -10,14 +10,15 @@ import {
export const dynamic = "force-dynamic";
-// POST { slug, full?, older?, floor?, force?, limit?, pages?, queueKey? } -> { ok: true, jobId }
+// POST { slug, full?, older?, floor?, from?, force?, limit?, pages?, queueKey? } -> { ok: true, jobId }
//
// `pages` caps how many pages a run reads, for a channel read page by page (a
// forum thread: its latest N pages).
//
// A social channel's "Fetch posts" button, over HTTP — and its "Re-fetch full
// history" (`full`) and "Fetch older posts" (`older`, with an optional `floor`
-// date, YYYY-MM-DD, that the walk stops at). The job runs on the platform's
+// date, YYYY-MM-DD, that the walk stops at, and an optional `from` date it
+// starts afresh at, replacing its saved position). The job runs on the platform's
// queue, as the buttons' does.
//
// Every refusal is the action's own sentence: a channel that is not a social
@@ -28,7 +29,7 @@ export const dynamic = "force-dynamic";
export async function POST(request: Request) {
return ops(
request,
- ["slug", "full", "older", "floor", "force", "limit", "pages", "queueKey"],
+ ["slug", "full", "older", "floor", "from", "force", "limit", "pages", "queueKey"],
async (body) => {
const slug = reqSlug(body, "slug");
return jobResponse(
@@ -41,6 +42,7 @@ export async function POST(request: Request) {
optString(body, "floor"),
optBool(body, "force"),
optPositiveInt(body, "pages"),
+ optString(body, "from"),
),
);
},
diff --git a/editor/app/api/ops/fetch-windows/route.test.ts b/editor/app/api/ops/fetch-windows/route.test.ts
@@ -0,0 +1,134 @@
+import test from "node:test";
+import assert from "node:assert/strict";
+import { mkdir, mkdtemp, rm, writeFile } from "node:fs/promises";
+import os from "node:os";
+import path from "node:path";
+
+// Run with:
+// pnpm -C editor exec tsx --test "app/api/ops/fetch-windows/route.test.ts"
+//
+// The body's shape, the refusals that come before any job, and a dry run's
+// plan — answered from a temp corpus. Every request here is a dry run or a
+// refusal: no job is queued and nothing is fetched.
+
+const ROOT = await mkdtemp(path.join(os.tmpdir(), "fetch-windows-route-"));
+// Set before the route (and getPaths, which caches) is first imported.
+process.env.WORKER_TOKEN = "test-token";
+process.env.TRANSCRIPTS_DIR = ROOT;
+process.env.SETTINGS_FILE = path.join(ROOT, "settings.json");
+const { POST } = await import("./route");
+test.after(() => rm(ROOT, { recursive: true, force: true }));
+
+// Channels on two platforms, one video each, the YouTube one with a window
+// already on disk.
+const CH = path.join(ROOT, "channels");
+async function channel(slug: string, url: string, id: string, webpage: string) {
+ await mkdir(path.join(CH, slug, "data", id), { recursive: true });
+ await writeFile(
+ path.join(CH, slug, "config.json"),
+ JSON.stringify({ handling: "youtube", name: slug, url }),
+ );
+ await writeFile(
+ path.join(CH, slug, "data", id, "metadata.info.json"),
+ JSON.stringify({ id, webpage_url: webpage }),
+ );
+}
+await channel("yt-chan", "https://www.youtube.com/@yt", "vid1", "https://www.youtube.com/watch?v=vid1");
+await channel("rb-chan", "https://rumble.com/c/rb", "rb1", "https://rumble.com/rb1-x.html");
+// A channel whose own URL is no platform's, holding a Rumble video.
+await channel("mix-chan", "https://example.test/mix", "rb2", "https://rumble.com/rb2-y.html");
+await mkdir(path.join(CH, "yt-chan", "data", "vid1", "clips"), { recursive: true });
+await writeFile(path.join(CH, "yt-chan", "data", "vid1", "clips", "0.00-60.00.mp4"), "mp4");
+
+async function post(
+ body: Record<string, unknown>,
+): Promise<{ status: number; json: Record<string, unknown> & { error?: string } }> {
+ const res = await POST(
+ new Request("http://localhost/api/ops/fetch-windows", {
+ method: "POST",
+ headers: {
+ authorization: "Bearer test-token",
+ "content-type": "application/json",
+ },
+ body: JSON.stringify(body),
+ }),
+ );
+ return { status: res.status, json: (await res.json()) as Record<string, unknown> };
+}
+
+const item = (o: Record<string, unknown> = {}) => ({ slug: "yt-chan", id: "vid1", from: 100, to: 110, ...o });
+
+test("siteId or items, never both, never neither; unknown keys refused", async () => {
+ assert.match((await post({})).json.error!, /"siteId" .* or "items" .* is required/);
+ const both = await post({ siteId: "s", items: [item()], requestedBy: "t" });
+ assert.equal(both.status, 400);
+ assert.match(both.json.error!, /either "siteId" or "items"/);
+ const unknown = await post({ items: [item()], requestedBy: "t", gapMs: 1 });
+ assert.equal(unknown.status, 400);
+ assert.match(unknown.json.error!, /unknown key\(s\): gapMs/);
+});
+
+test("an item list is validated at the door", async () => {
+ const cases: [Record<string, unknown>, RegExp][] = [
+ [{ items: [], requestedBy: "t" }, /non-empty array/],
+ [{ items: [item({ slug: "../x" })], requestedBy: "t" }, /not a valid channel slug/],
+ [{ items: [item({ id: ".." })], requestedBy: "t" }, /not a video id/],
+ [{ items: [item({ from: 20, to: 10 })], requestedBy: "t" }, /must be less than to/],
+ [{ items: [item({ from: 0, to: 2000 })], requestedBy: "t" }, /at most 900s/],
+ [{ items: [item({ from: -1 })], requestedBy: "t" }, /non-negative seconds/],
+ [{ items: [item({ extra: 1 })], requestedBy: "t" }, /items\[0\]" has unknown key\(s\): extra/],
+ [{ items: [item()] }, /"requestedBy" is required/],
+ [{ items: [item()], requestedBy: "t", maxHeight: 50 }, /"maxHeight" must be a whole number of pixels/],
+ ];
+ for (const [body, re] of cases) {
+ const r = await post(body);
+ assert.equal(r.status, 400, JSON.stringify(body));
+ assert.match(r.json.error!, re);
+ }
+});
+
+test("a site's windows name the site; a missing site is refused", async () => {
+ const named = await post({ siteId: "demo-site", requestedBy: "me" });
+ assert.equal(named.status, 400);
+ assert.match(named.json.error!, /"requestedBy" goes with "items"/);
+ const missing = await post({ siteId: "demo-site" });
+ assert.equal(missing.status, 400);
+ assert.match(missing.json.error!, /No site "demo-site"/);
+});
+
+test("a dry run answers the cache, groups by platform queue, and names what it cannot resolve", async () => {
+ const r = await post({
+ requestedBy: "test",
+ dryRun: true,
+ items: [
+ item({ from: 10, to: 20 }), // inside the cached 0–60 window
+ item({ from: 100, to: 110 }),
+ item({ from: 100, to: 110 }), // a duplicate collapses
+ item({ slug: "rb-chan", id: "rb1", from: 5, to: 15 }),
+ item({ slug: "mix-chan", id: "rb2", from: 5, to: 15 }),
+ item({ slug: "no-such", id: "x", from: 5, to: 15 }),
+ ],
+ });
+ assert.equal(r.status, 200);
+ const j = r.json as {
+ dryRun: boolean;
+ groups: { platform: string; queueKey: string; items: { id: string; webpageUrl?: string }[] }[];
+ cached: { from: number }[];
+ unresolved: { error: string }[];
+ jobs?: unknown;
+ };
+ assert.equal(j.dryRun, true);
+ assert.equal(j.jobs, undefined, "a dry run starts nothing");
+ assert.deepEqual(j.cached.map((c) => c.from), [10]);
+ assert.deepEqual(
+ j.groups.map((g) => [g.platform, g.items.length]).sort(),
+ [["rumble", 2], ["youtube", 1]],
+ );
+ // The window's own URL picks the queue: the mix channel's Rumble video joins
+ // the Rumble job rather than starting a second one beside it.
+ assert.deepEqual(j.groups.map((g) => g.queueKey).sort(), ["platform:rumble", "platform:youtube"]);
+ const yt = j.groups.find((g) => g.platform === "youtube")!;
+ assert.equal(yt.items[0].webpageUrl, "https://www.youtube.com/watch?v=vid1");
+ assert.equal(j.unresolved.length, 1);
+ assert.match(j.unresolved[0].error, /Channel "no-such" not found/);
+});
diff --git a/editor/app/api/ops/fetch-windows/route.ts b/editor/app/api/ops/fetch-windows/route.ts
@@ -0,0 +1,179 @@
+import { NextResponse } from "next/server";
+import {
+ MAX_CLIP_WINDOW_SECONDS,
+ MAX_FETCH_MAX_HEIGHT,
+ MIN_FETCH_MAX_HEIGHT,
+ isFetchMaxHeight,
+} from "yt-dlp-transcript-common/lib/clipWindow";
+import { isValidSiteId } from "yt-dlp-transcript-common/lib/site";
+import type { FetchWindowsItem } from "yt-dlp-transcript-common/controller/fetchWindows";
+import {
+ fetchMissingEvidenceAction,
+ fetchWindowsAction,
+} from "../../../channels/[slug]/videos/fetchWindowsAction";
+import {
+ OpsInputError,
+ ops,
+ opsFail,
+ optBool,
+ optString,
+ reqSlug,
+ reqString,
+ reqVideoId,
+ type OpsBody,
+} from "../_lib";
+
+export const dynamic = "force-dynamic";
+
+// POST { siteId, maxHeight?, dryRun? }
+// | { items: [{ slug, id, from, to, clipId?, reason?, pad?, webpageUrl? }, …],
+// requestedBy, manifest?, maxHeight?, dryRun? }
+// dryRun -> { ok, dryRun: true, groups: [{ platform, queueKey, items }],
+// cached, unresolved[, onDisk, unfetchable] }
+// else -> { ok, dryRun: false, jobs: [{ platform, queueKey, jobId, items }],
+// jobIds, jobId (one job only), refused, cached, unresolved
+// [, onDisk, unfetchable] }
+//
+// Fetch clip windows through the managed path, ONE PACED JOB PER PLATFORM
+// QUEUE (controller/fetchWindows.ts): the batch form of
+// /api/media/fetch-window. `siteId` fetches every window the site's published
+// reports cite and the disk does not hold (`unfetchable` lists the spans no
+// window can fill, with why; `onDisk` counts the ones already there). `items`
+// fetches an explicit list, and must say who asked (`requestedBy`). A window
+// already on disk is answered in `cached` and joins no job; a platform cooling
+// down or held is `refused` for its group, and the other groups still start.
+// Re-running the same body is the resume: fetched windows are cached.
+//
+// A big list goes in a file: `pnpm ops fetch-windows --file list.json`.
+
+const ITEM_KEYS = ["slug", "id", "from", "to", "clipId", "reason", "pad", "webpageUrl"];
+
+function optItemString(e: OpsBody, key: string, where: string): string | undefined {
+ const v = e[key];
+ if (v === undefined) return undefined;
+ if (typeof v !== "string") throw new OpsInputError(`"${where}.${key}" must be a string`);
+ return v.trim() || undefined;
+}
+
+function reqSeconds(e: OpsBody, key: string, where: string): number {
+ const v = e[key];
+ if (typeof v !== "number" || !Number.isFinite(v) || v < 0) {
+ throw new OpsInputError(`"${where}.${key}" must be finite, non-negative seconds`);
+ }
+ return v;
+}
+
+function reqWindowItems(body: OpsBody): FetchWindowsItem[] {
+ const v = body.items;
+ if (!Array.isArray(v) || v.length === 0) {
+ throw new OpsInputError(
+ `"items" must be a non-empty array of { "slug", "id", "from", "to" }`,
+ );
+ }
+ return v.map((entry, i) => {
+ const where = `items[${i}]`;
+ if (typeof entry !== "object" || entry === null || Array.isArray(entry)) {
+ throw new OpsInputError(`"${where}" must be an object { "slug", "id", "from", "to" }`);
+ }
+ const e = entry as OpsBody;
+ const extra = Object.keys(e).filter((k) => !ITEM_KEYS.includes(k));
+ if (extra.length) {
+ throw new OpsInputError(
+ `"${where}" has unknown key(s): ${extra.join(", ")} — accepted: ${ITEM_KEYS.join(", ")}`,
+ );
+ }
+ const from = reqSeconds(e, "from", where);
+ const to = reqSeconds(e, "to", where);
+ if (from >= to) {
+ throw new OpsInputError(`"${where}": from (${from}) must be less than to (${to})`);
+ }
+ if (to - from > MAX_CLIP_WINDOW_SECONDS) {
+ throw new OpsInputError(
+ `"${where}": a window may be at most ${MAX_CLIP_WINDOW_SECONDS}s ` +
+ `(asked for ${Math.round(to - from)}s)`,
+ );
+ }
+ const pad = e.pad;
+ if (pad !== undefined && (typeof pad !== "number" || !Number.isFinite(pad))) {
+ throw new OpsInputError(`"${where}.pad" must be a number`);
+ }
+ return {
+ slug: reqSlug(e, "slug"),
+ id: reqVideoId(e, "id"),
+ from,
+ to,
+ clipId: optItemString(e, "clipId", where),
+ reason: optItemString(e, "reason", where)?.slice(0, 400),
+ ...(typeof pad === "number" ? { pad } : {}),
+ webpageUrl: optItemString(e, "webpageUrl", where),
+ };
+ });
+}
+
+function optMaxHeight(body: OpsBody): number | undefined {
+ const v = body.maxHeight;
+ if (v === undefined || v === null) return undefined;
+ if (!isFetchMaxHeight(v)) {
+ throw new OpsInputError(
+ `"maxHeight" must be a whole number of pixels from ` +
+ `${MIN_FETCH_MAX_HEIGHT} to ${MAX_FETCH_MAX_HEIGHT}`,
+ );
+ }
+ return v;
+}
+
+export async function POST(request: Request) {
+ return ops(
+ request,
+ ["siteId", "items", "requestedBy", "manifest", "maxHeight", "dryRun"],
+ async (body) => {
+ const maxHeight = optMaxHeight(body);
+ const dryRun = optBool(body, "dryRun");
+ if (body.siteId !== undefined && body.items !== undefined) {
+ throw new OpsInputError('send either "siteId" or "items", not both');
+ }
+ let result: Awaited<ReturnType<typeof fetchWindowsAction>>;
+ if (body.siteId !== undefined) {
+ // A site's windows say who asked themselves (the reports, by site).
+ for (const k of ["requestedBy", "manifest"]) {
+ if (body[k] !== undefined) {
+ throw new OpsInputError(`"${k}" goes with "items"; a site's windows name the site`);
+ }
+ }
+ const siteId = reqString(body, "siteId");
+ if (!isValidSiteId(siteId)) {
+ throw new OpsInputError(
+ `"${siteId}" is not a valid site id (lowercase letters, digits and "-"; must start with a letter or digit)`,
+ );
+ }
+ result = await fetchMissingEvidenceAction(siteId, { dryRun, maxHeight });
+ } else if (body.items !== undefined) {
+ const items = reqWindowItems(body);
+ // WHO ASKED IS NOT OPTIONAL, as on the single route: a window nobody
+ // can explain in six months is one nobody can clean up.
+ const requestedBy = reqString(body, "requestedBy");
+ result = await fetchWindowsAction({
+ items,
+ requestedBy,
+ manifest: optString(body, "manifest")?.trim() || undefined,
+ maxHeight,
+ dryRun,
+ });
+ } else {
+ throw new OpsInputError('"siteId" (a site\'s missing evidence) or "items" (a list of windows) is required');
+ }
+ if (!result.ok) return opsFail(result.error);
+ if (result.dryRun) {
+ const { jobs: _jobs, refused: _refused, ...plan } = result;
+ return NextResponse.json(plan);
+ }
+ const { groups: _groups, ...run } = result;
+ const jobIds = run.jobs.map((j) => j.jobId);
+ return NextResponse.json({
+ ...run,
+ jobIds,
+ ...(jobIds.length === 1 ? { jobId: jobIds[0] } : {}),
+ });
+ },
+ );
+}
diff --git a/editor/app/api/ops/rename-channel/route.test.ts b/editor/app/api/ops/rename-channel/route.test.ts
@@ -0,0 +1,153 @@
+import test from "node:test";
+import assert from "node:assert/strict";
+import { access, mkdir, mkdtemp, readFile, rm, writeFile } from "node:fs/promises";
+import os from "node:os";
+import path from "node:path";
+
+// Run with:
+// pnpm -C editor exec tsx --test "app/api/ops/rename-channel/route.test.ts"
+//
+// The channel lifecycle over ops — rename, delete, create, the membership and
+// exclusion half of channel-config, and the list — against a temp corpus with
+// two social channels and one site: every REFUSAL, which is answered before
+// anything is written. The successful writes end in revalidatePath, which needs
+// the request store a running Next server has and a bare tsx process has not
+// (safeRevalidate.ts); they are driven over HTTP in e2e/ops-api.spec.ts.
+
+const ROOT = await mkdtemp(path.join(os.tmpdir(), "channel-lifecycle-route-"));
+process.env.WORKER_TOKEN = "test-token";
+process.env.TRANSCRIPTS_DIR = ROOT;
+process.env.SETTINGS_FILE = path.join(ROOT, "settings.json");
+
+async function makeSocial(slug: string, handle: string): Promise<void> {
+ const dir = path.join(ROOT, "channels", slug);
+ await mkdir(path.join(dir, "posts"), { recursive: true });
+ await writeFile(
+ path.join(dir, "config.json"),
+ JSON.stringify({
+ handling: "transcribe",
+ sourceKind: "social",
+ platform: "twitter",
+ postFetcher: "x-gallery-dl",
+ socialHandle: handle,
+ name: `${handle} (X)`,
+ url: `https://x.com/${handle}`,
+ }),
+ );
+}
+await makeSocial("misspeled-x", "example_user");
+await makeSocial("other-x", "other_user");
+await mkdir(path.join(ROOT, "sites", "demo-site"), { recursive: true });
+await writeFile(
+ path.join(ROOT, "sites", "demo-site", "site.json"),
+ JSON.stringify({ siteId: "demo-site", siteTitle: "Demo", channels: [] }),
+);
+
+const rename = (await import("./route")).POST;
+const del = (await import("../delete-channel/route")).POST;
+const create = (await import("../create-channel/route")).POST;
+const config = (await import("../channel-config/route")).POST;
+const list = (await import("../channels/route")).GET;
+test.after(() => rm(ROOT, { recursive: true, force: true }));
+
+type Res = { status: number; json: Record<string, unknown> };
+async function call(
+ handler: (r: Request) => Promise<Response>,
+ body?: Record<string, unknown>,
+): Promise<Res> {
+ const res = await handler(
+ new Request("http://localhost/api/ops/x", {
+ method: body ? "POST" : "GET",
+ headers: {
+ authorization: "Bearer test-token",
+ "content-type": "application/json",
+ },
+ ...(body ? { body: JSON.stringify(body) } : {}),
+ }),
+ );
+ return { status: res.status, json: (await res.json()) as Record<string, unknown> };
+}
+const exists = (p: string) => access(p).then(() => true, () => false);
+const readSite = async () =>
+ JSON.parse(await readFile(path.join(ROOT, "sites", "demo-site", "site.json"), "utf8")) as {
+ channels: { slug: string }[];
+ };
+
+test("rename refuses a taken slug and an invalid one, in the form's sentences", async () => {
+ const taken = await call(rename, { slug: "misspeled-x", newSlug: "other-x" });
+ assert.equal(taken.status, 400);
+ assert.equal(taken.json.error, 'Channel "other-x" already exists');
+ const bad = await call(rename, { slug: "misspeled-x", newSlug: "../escape" });
+ assert.equal(bad.status, 400);
+ assert.match(String(bad.json.error), /not a valid channel slug/);
+ const missing = await call(rename, { slug: "nope", newSlug: "nope-2" });
+ assert.equal(missing.status, 400);
+ assert.equal(missing.json.error, 'Channel "nope" not found');
+});
+
+test("the list names every channel, its kind and its sites", async () => {
+ const listed = await call(list);
+ assert.equal(listed.status, 200);
+ const rows = listed.json.channels as { slug: string; sites: string[]; kind: string }[];
+ assert.deepEqual(rows.map((c) => c.slug).sort(), ["misspeled-x", "other-x"]);
+ assert.ok(rows.every((c) => c.kind === "social" && c.sites.length === 0));
+});
+
+test("channel-config refuses an unknown site and a malformed membership", async () => {
+ const unknown = await call(config, { slug: "misspeled-x", sites: [{ siteId: "no-such-site" }] });
+ assert.equal(unknown.status, 400);
+ assert.equal(unknown.json.error, 'no site "no-such-site" — known: demo-site');
+ const stray = await call(config, { slug: "misspeled-x", sites: [{ siteId: "demo-site", group: "a" }] });
+ assert.equal(stray.status, 400);
+ assert.match(String(stray.json.error), /unknown key\(s\) in a "sites" entry: group/);
+ const notArray = await call(config, { slug: "misspeled-x", sites: "demo-site" });
+ assert.equal(notArray.status, 400);
+ assert.match(String(notArray.json.error), /"sites" must be an array/);
+ // Nothing was written.
+ assert.deepEqual((await readSite()).channels, []);
+});
+
+test("channel-config refuses an empty body and a non-boolean exclusion", async () => {
+ const none = await call(config, { slug: "misspeled-x" });
+ assert.equal(none.status, 400);
+ assert.match(String(none.json.error), /nothing to change/);
+ const notBool = await call(config, { slug: "misspeled-x", excludeFromBuild: "yes" });
+ assert.equal(notBool.status, 400);
+ assert.equal(notBool.json.error, '"excludeFromBuild" must be a boolean');
+ const cfg = JSON.parse(
+ await readFile(path.join(ROOT, "channels", "misspeled-x", "config.json"), "utf8"),
+ );
+ assert.equal(cfg.excludeFromBuild, undefined);
+});
+
+test("create refuses what the New form refuses, before any channel exists", async () => {
+ const noName = await call(create, { fields: { handling: "transcribe" } });
+ assert.equal(noName.status, 400);
+ assert.equal(noName.json.error, "Name is required");
+ const badKey = await call(create, { fields: { name: "X", handlng: "youtube" } });
+ assert.equal(badKey.status, 400);
+ assert.match(String(badKey.json.error), /"handlng" is not a channel config field/);
+ const taken = await call(create, {
+ slug: "other-x",
+ fields: { name: "Other", handling: "transcribe" },
+ });
+ assert.equal(taken.status, 400);
+ assert.equal(taken.json.error, 'Channel "other-x" already exists');
+ const unknownSite = await call(create, {
+ slug: "made-x",
+ fields: { name: "Made", handling: "transcribe" },
+ sites: [{ siteId: "nope" }],
+ });
+ assert.equal(unknownSite.status, 400);
+ assert.equal(await exists(path.join(ROOT, "channels", "made-x")), false);
+});
+
+test("delete needs the slug repeated, and refuses before touching anything", async () => {
+ const wrong = await call(del, { slug: "other-x", confirm: "other" });
+ assert.equal(wrong.status, 400);
+ assert.equal(wrong.json.error, 'Type the channel slug "other-x" exactly to confirm deletion');
+ const missing = await call(del, { slug: "other-x" });
+ assert.equal(missing.status, 400);
+ assert.match(String(missing.json.error), /"confirm" is required/);
+ assert.equal(await exists(path.join(ROOT, "channels", "other-x", "config.json")), true);
+});
diff --git a/editor/app/api/ops/rename-channel/route.ts b/editor/app/api/ops/rename-channel/route.ts
@@ -0,0 +1,34 @@
+import { NextResponse } from "next/server";
+import { renameChannelFromForm } from "../../../channels/actions";
+import { opsFail, ops, reqSlug } from "../_lib";
+
+export const dynamic = "force-dynamic";
+
+// POST { slug: string, newSlug: string }
+//
+// The channel page's Danger → Rename, over HTTP. The form makes the operator
+// type the current slug to confirm; here the body naming `slug` IS that
+// confirmation, so it is passed through as the typed value and the action
+// applies every other rule unchanged — a busy channel (a running job, a lane
+// unit writing into data/, media in transition), a taken or invalid new slug.
+//
+// Answers { ok, slug: <new slug>, warnings }. A warning is a non-fatal
+// metadata-migration problem after the directory moved: the form can only log
+// it, and this can say it.
+export async function POST(request: Request) {
+ return ops(request, ["slug", "newSlug"], async (body) => {
+ const slug = reqSlug(body, "slug");
+ const newSlug = reqSlug(body, "newSlug");
+ const fd = new FormData();
+ fd.set("confirmSlug", slug);
+ fd.set("newSlug", newSlug);
+ const result = await renameChannelFromForm(slug, fd);
+ if (!result.ok) return opsFail(result.error);
+ return NextResponse.json({
+ ok: true,
+ slug: result.slug,
+ from: slug,
+ warnings: result.warnings,
+ });
+ });
+}
diff --git a/editor/app/api/ops/transcribe/route.test.ts b/editor/app/api/ops/transcribe/route.test.ts
@@ -0,0 +1,86 @@
+import test from "node:test";
+import assert from "node:assert/strict";
+import { mkdir, mkdtemp, readFile, rm, writeFile } from "node:fs/promises";
+import os from "node:os";
+import path from "node:path";
+
+// Run with:
+// pnpm -C editor exec tsx --test "app/api/ops/transcribe/route.test.ts"
+//
+// The adapter's surface: the token, the body's keys, and the controller's
+// refusals answered as 400s before any job exists. Every request here is
+// refused — no job is queued and no engine runs (the job itself is covered by
+// common/controller/transcribeFile.test.ts).
+
+const ROOT = await mkdtemp(path.join(os.tmpdir(), "transcribe-route-"));
+const CORPUS = path.join(ROOT, "transcripts");
+await mkdir(CORPUS, { recursive: true });
+const SETTINGS_FILE = path.join(ROOT, "settings.json");
+const SETTINGS_TEXT = JSON.stringify({
+ workers: [
+ {
+ id: "gpu",
+ name: "GPU",
+ kind: "local",
+ enabled: true,
+ priority: 0,
+ appId: "parakeet",
+ config: { model: "/models/x.gguf" },
+ },
+ ],
+});
+await writeFile(SETTINGS_FILE, SETTINGS_TEXT);
+const MEDIA = path.join(ROOT, "clip.wav");
+await writeFile(MEDIA, "x");
+// Set before the route (and getPaths, which caches) is first imported.
+process.env.WORKER_TOKEN = "test-token";
+process.env.TRANSCRIPTS_DIR = CORPUS;
+process.env.SETTINGS_FILE = SETTINGS_FILE;
+const { POST } = await import("./route");
+test.after(() => rm(ROOT, { recursive: true, force: true }));
+
+async function post(
+ body: Record<string, unknown>,
+ token = "test-token",
+): Promise<{ status: number; json: { ok?: boolean; error?: string; jobId?: string } }> {
+ const res = await POST(
+ new Request("http://localhost/api/ops/transcribe", {
+ method: "POST",
+ headers: {
+ authorization: `Bearer ${token}`,
+ "content-type": "application/json",
+ },
+ body: JSON.stringify(body),
+ }),
+ );
+ return { status: res.status, json: await res.json() };
+}
+
+test("a wrong token is a 401", async () => {
+ const { status } = await post({ path: MEDIA }, "wrong");
+ assert.equal(status, 401);
+});
+
+test("an unknown key is a 400 naming the accepted ones", async () => {
+ const { status, json } = await post({ path: MEDIA, model: "/x.gguf" });
+ assert.equal(status, 400);
+ assert.match(json.error!, /unknown key\(s\): model — this route accepts path, start, end, workerId, out/);
+});
+
+test("the controller's refusals come back as 400s", async () => {
+ const cases: [Record<string, unknown>, RegExp][] = [
+ [{ path: "clip.wav" }, /"path" must be absolute/],
+ [{ path: path.join(ROOT, "missing.wav") }, /does not exist/],
+ [{ path: MEDIA, start: 30, end: 10 }, /"end" \(10\) must be after "start" \(30\)/],
+ [{ path: MEDIA, workerId: "nope" }, /no worker "nope" — known: gpu/],
+ [{ path: MEDIA, out: path.join(CORPUS, "r.json") }, /inside the corpus/],
+ ];
+ for (const [body, re] of cases) {
+ const { status, json } = await post(body);
+ assert.equal(status, 400, JSON.stringify(body));
+ assert.equal(json.ok, false);
+ assert.match(json.error!, re);
+ assert.equal(json.jobId, undefined);
+ }
+ assert.equal(await readFile(SETTINGS_FILE, "utf8"), SETTINGS_TEXT);
+});
diff --git a/editor/app/api/ops/transcribe/route.ts b/editor/app/api/ops/transcribe/route.ts
@@ -0,0 +1,32 @@
+import { getPaths } from "yt-dlp-transcript-common/lib/paths";
+import {
+ TRANSCRIBE_FILE_BODY_KEYS,
+ enqueueTranscribeFile,
+} from "yt-dlp-transcript-common/controller/transcribeFile";
+import { jobResponse, ops } from "../_lib";
+
+export const dynamic = "force-dynamic";
+
+// POST { path, start?, end?, workerId?, out? } -> { ok: true, jobId }
+//
+// Transcribe ONE local file — or the window [start, end) seconds of it — on a
+// local transcription worker, as a job on /jobs: the engine, model and
+// command line the corpus uses, never a hand-run whisper or parakeet. No
+// `workerId` = the worker auto-transcribe would get (the pool's highest-priority
+// free local one).
+//
+// The job's log ends with the result as ONE line, `@@transcribe-result
+// {json}`: `{version, path, window, worker: {id, name, appId, model, device},
+// transcriptFormat, transcribedAt, durationMs, cues: [{start, end, text}],
+// text}`, cue times on the SOURCE file's clock. `pnpm ops transcribe --wait`
+// prints that JSON on stdout. `out` writes it to that file as well.
+//
+// AN ADAPTER: every refusal is the controller's (controller/transcribeFile.ts)
+// — a relative or unreadable `path`, `end` not after `start`, an unknown or
+// non-local `workerId`, a named worker switched off, an `out` inside the corpus
+// (resolved through symlinks) or in a directory that does not exist.
+export async function POST(request: Request) {
+ return ops(request, TRANSCRIBE_FILE_BODY_KEYS, async (body) =>
+ jobResponse(await enqueueTranscribeFile(body, { paths: getPaths() })),
+ );
+}
diff --git a/editor/app/channels/[slug]/page.tsx b/editor/app/channels/[slug]/page.tsx
@@ -204,6 +204,44 @@ export default async function ChannelDetailPage({
...(forumSession ? { forumSession } : {}),
}}
/>
+ {/* THE SAME DANGER ZONE A VIDEO CHANNEL HAS, which a social channel
+ used to lack entirely: this branch returns before the stage list,
+ so a misspelled X channel could not be renamed from the editor. The
+ forms, actions and busy guard are the video page's own; collapsed,
+ because it is a chore and not this page's subject, and opened by
+ the same `?stage=danger` link a video channel answers to. */}
+ <details
+ data-social-danger=""
+ open={sp.stage === "danger"}
+ className="rounded-lg border border-border bg-card p-4"
+ >
+ <summary className="cursor-pointer text-sm font-medium text-destructive">
+ Danger zone
+ </summary>
+ <div className="mt-3 flex flex-col gap-4">
+ <RenameChannelForm
+ slug={slug}
+ busyReason={channelMediaBusyReason(slug, "renaming it")}
+ action={
+ renameChannelAction.bind(null, slug) as (
+ prev: ActionResult,
+ formData: FormData,
+ ) => Promise<ActionResult>
+ }
+ />
+ <hr className="border-border" />
+ <DeleteChannelForm
+ slug={slug}
+ busyReason={channelMediaBusyReason(slug, "deleting it")}
+ action={
+ deleteChannelAction.bind(null, slug) as (
+ prev: ActionResult,
+ formData: FormData,
+ ) => Promise<ActionResult>
+ }
+ />
+ </div>
+ </details>
</div>
);
}
diff --git a/editor/app/channels/[slug]/socialActions.ts b/editor/app/channels/[slug]/socialActions.ts
@@ -23,7 +23,7 @@ import { isSocialChannel } from "yt-dlp-transcript-common/lib/channelConfig";
import {
emptyAccountOlderProblem,
fetchPosts,
- FULL_AND_OLDER_REFUSAL,
+ FULL_AND_OLDER_REFUSAL, OLDER_FROM_REFUSAL,
olderPostsProblem,
} from "yt-dlp-transcript-common/controller/fetchPosts";
import { isUtcDay } from "yt-dlp-transcript-common/social/olderBackfill";
@@ -166,6 +166,8 @@ export async function fetchPostsAction(
force?: boolean,
// A page walker's cap (a forum thread: its latest N pages this run).
pages?: number,
+ // The older walk's fresh start date (FetchPostsOptions.from).
+ from?: string,
): Promise<StreamActionResult> {
if (full && older) return { ok: false, error: FULL_AND_OLDER_REFUSAL };
if (pages !== undefined && !(Number.isInteger(pages) && pages > 0)) {
@@ -174,6 +176,10 @@ export async function fetchPostsAction(
if (floor !== undefined && !older) {
return { ok: false, error: "A floor date applies only to an older-posts fetch." };
}
+ if (from !== undefined && !older) return { ok: false, error: OLDER_FROM_REFUSAL };
+ if (from !== undefined && !isUtcDay(from)) {
+ return { ok: false, error: `"${from}" is not a date (YYYY-MM-DD).` };
+ }
if (force && !older) {
return { ok: false, error: "\"force\" applies only to an older-posts fetch." };
}
@@ -215,7 +221,7 @@ export async function fetchPostsAction(
spec: {
kind: "fetch-posts",
slug,
- params: { queueKey, full, limit, older, floor, force, ...(pages ? { pages } : {}) },
+ params: { queueKey, full, limit, older, floor, force, ...(pages ? { pages } : {}), ...(from ? { from } : {}) },
},
fn: async (onLog, signal, _progress, ctx) => {
const result = await fetchPosts({
@@ -225,6 +231,7 @@ export async function fetchPostsAction(
full,
older,
floor,
+ from,
force,
limit,
pages,
diff --git a/editor/app/channels/[slug]/videos/fetchWindowsAction.ts b/editor/app/channels/[slug]/videos/fetchWindowsAction.ts
@@ -0,0 +1,408 @@
+"use server";
+
+import path from "node:path";
+import { getPaths } from "yt-dlp-transcript-common/lib/paths";
+import { getSettings } from "yt-dlp-transcript-common/lib/settings";
+import { diskGate } from "yt-dlp-transcript-common/lib/diskSpace";
+import { formatBytes } from "yt-dlp-transcript-common/lib/format";
+import { resolveQueueKey } from "yt-dlp-transcript-common/lib/queueKeys";
+import { detectPlatform, queueKeyForUrl } from "yt-dlp-transcript-common/lib/platform";
+import {
+ MAX_CLIP_WINDOW_SECONDS,
+ isFetchMaxHeight,
+} from "yt-dlp-transcript-common/lib/clipWindow";
+import { findContainingClipWindow } from "yt-dlp-transcript-common/lib/clipWindow-server";
+import type { ChannelConfig } from "yt-dlp-transcript-common/lib/channelConfig";
+import {
+ isValidChannelSlug,
+ readChannelConfig,
+} from "yt-dlp-transcript-common/controller/channels";
+import { findVideoSourceUrl } from "yt-dlp-transcript-common/controller/undownloadedVideos";
+import {
+ fetchWindows,
+ fetchWindowsItemLabel,
+ dedupeFetchWindowsItems,
+ type FetchWindowsItem,
+} from "yt-dlp-transcript-common/controller/fetchWindows";
+import {
+ missingEvidenceWindows,
+ type UnfetchableEvidenceWindow,
+} from "yt-dlp-transcript-common/publish/reportMedia";
+import { isValidSiteId, listSiteIds } from "yt-dlp-transcript-common/lib/site";
+import {
+ heldPlatformRefusal,
+ platformCooldownRemainingMs,
+} from "yt-dlp-transcript-common/jobs/downloadBackoff";
+import {
+ runManagedFunction,
+ type StreamActionResult,
+} from "yt-dlp-transcript-common/jobs/streamCommand";
+import { safeRevalidate } from "../../../lib/safeRevalidate";
+
+// FETCH A LIST OF CLIP WINDOWS through the managed path, as ONE JOB PER
+// PLATFORM QUEUE (controller/fetchWindows.ts walks each one, paced).
+//
+// The batch form of fetchWindowAction (videos/[id]/videoActions.ts), with the
+// same rules at the door: a window of at most MAX_CLIP_WINDOW_SECONDS, a height
+// cap in range, a channel that exists, a URL that resolves. What it adds is the
+// fan-out: the list is grouped by the queue of each window's own URL
+// (`queueKeyForUrl`) and each group starts its
+// own job on that queue, so YouTube and Rumble run side by side, each behind
+// its own platform's other downloads — persistVideosAction's shape.
+//
+// A PLATFORM COOLING DOWN OR HELD is refused at the door for ITS group (the
+// sentence comes back in `refused`); the other groups still start. A window
+// already on disk is answered here (`cached`) and joins no job.
+//
+// Like the single fetch, NOT GATED BY THE DOWNLOAD PAUSE: an operator (or a
+// tool they are driving) asked for these seconds by hand.
+
+const ID_RE = /^[\w.-]+$/;
+const isVideoId = (v: string): boolean => ID_RE.test(v) && v !== "." && v !== "..";
+
+export type FetchWindowsRequest = {
+ items: FetchWindowsItem[];
+ // Who asked: the tool or surface. Recorded beside every window.
+ requestedBy: string;
+ manifest?: string;
+ maxHeight?: number;
+ dryRun?: boolean;
+ // A queue override for every group (resolveQueueKey's rules).
+ queueKey?: string;
+};
+
+export type FetchWindowsUnresolved = { item: FetchWindowsItem; error: string };
+
+export type FetchWindowsGroup = {
+ // The cooldown key (`detectPlatform(url)`), and the queue the job runs on.
+ platform: string;
+ queueKey: string;
+ items: FetchWindowsItem[];
+};
+
+export type FetchWindowsActionResult =
+ | { ok: false; error: string }
+ | {
+ ok: true;
+ dryRun: boolean;
+ // A real run: one per group that started.
+ jobs: { platform: string; queueKey: string; jobId: string; items: number }[];
+ // A dry run: what each job would be given.
+ groups: FetchWindowsGroup[];
+ // Groups refused at the door, with the refusal's own sentence.
+ refused: { platform: string; error: string; items: number }[];
+ cached: FetchWindowsItem[];
+ unresolved: FetchWindowsUnresolved[];
+ };
+
+// The window rules the HTTP door enforces, asked again here: a replayed spec is
+// a file on disk.
+function itemProblem(item: FetchWindowsItem): string | null {
+ if (!isValidChannelSlug(item.slug)) return `"${item.slug}" is not a channel slug`;
+ if (!isVideoId(item.id)) return `"${item.id}" is not a video id`;
+ const { from, to } = item;
+ if (
+ !Number.isFinite(from) ||
+ !Number.isFinite(to) ||
+ from < 0 ||
+ from >= to ||
+ to - from > MAX_CLIP_WINDOW_SECONDS
+ ) {
+ return (
+ `${from}–${to} is not a fetchable window ` +
+ `(at most ${MAX_CLIP_WINDOW_SECONDS}s, from < to, from >= 0)`
+ );
+ }
+ return null;
+}
+
+// Validate, answer the cache, resolve every URL and group by queue. Nothing is
+// started; nothing touches the network.
+async function planFetchWindows(
+ items: FetchWindowsItem[],
+ queueOverride: string | undefined,
+): Promise<{
+ groups: FetchWindowsGroup[];
+ cached: FetchWindowsItem[];
+ unresolved: FetchWindowsUnresolved[];
+}> {
+ const paths = getPaths();
+ const configs = new Map<string, ChannelConfig | null>();
+ const groups = new Map<string, FetchWindowsGroup>();
+ const cached: FetchWindowsItem[] = [];
+ const unresolved: FetchWindowsUnresolved[] = [];
+ for (const item of dedupeFetchWindowsItems(items)) {
+ const problem = itemProblem(item);
+ if (problem) {
+ unresolved.push({ item, error: problem });
+ continue;
+ }
+ if (!configs.has(item.slug)) {
+ configs.set(item.slug, await readChannelConfig(paths, item.slug));
+ }
+ const config = configs.get(item.slug);
+ if (!config) {
+ unresolved.push({ item, error: `Channel "${item.slug}" not found` });
+ continue;
+ }
+ const videoDir = path.join(paths.channelsDir, item.slug, "data", item.id);
+ if (await findContainingClipWindow(videoDir, item.from, item.to)) {
+ cached.push(item);
+ continue;
+ }
+ const url =
+ item.webpageUrl?.trim() ||
+ (await findVideoSourceUrl(paths, item.slug, item.id, config));
+ if (!url) {
+ unresolved.push({
+ item,
+ error:
+ "Could not determine the video URL: no metadata.info.json and the " +
+ "playlist does not contain a matching entry.",
+ });
+ continue;
+ }
+ // THE WINDOW'S OWN URL decides the queue, not the channel's: a channel
+ // whose URL is no platform's (a curated mix) can hold Rumble videos, and
+ // its queue would put a second Rumble job beside the first, each pacing
+ // only itself.
+ const queueKey = resolveQueueKey(queueKeyForUrl(url), queueOverride);
+ const platform = detectPlatform(url) ?? "unknown";
+ // One job per queue. Two platforms sharing a queue (an override) share a
+ // job too; the controller keys the cooldown per item, so each is honoured.
+ const group = groups.get(queueKey) ?? { platform, queueKey, items: [] };
+ group.items.push({ ...item, webpageUrl: url });
+ groups.set(queueKey, group);
+ }
+ return { groups: [...groups.values()], cached, unresolved };
+}
+
+// The door each group passes before its job is queued: the platform's hold,
+// then its cooldown — fetchWindowAction's checks and sentences.
+async function groupRefusal(platform: string): Promise<string | null> {
+ const paths = getPaths();
+ const held = await heldPlatformRefusal(platform, "This batch", paths);
+ if (held) return held;
+ const cooldownMs = await platformCooldownRemainingMs(platform, paths);
+ if (cooldownMs > 0) {
+ return (
+ `${platform} is in a rate-limit cooldown ` +
+ `(${Math.ceil(cooldownMs / 1000)}s remaining).`
+ );
+ }
+ return null;
+}
+
+// One platform's windows, as one job. Exported for Retry: the replay hands the
+// spec's items back here, and the controller's cache check skips whatever an
+// earlier run fetched.
+export async function fetchWindowsJobAction(req: {
+ items: FetchWindowsItem[];
+ requestedBy: string;
+ manifest?: string;
+ siteId?: string;
+ maxHeight?: number;
+ queueKey: string;
+}): Promise<StreamActionResult> {
+ if (req.items.length === 0) return { ok: false, error: "No windows to fetch." };
+ if (req.maxHeight !== undefined && !isFetchMaxHeight(req.maxHeight)) {
+ return {
+ ok: false,
+ error: `maxHeight ${req.maxHeight} is not a source height to cap a fetch at.`,
+ };
+ }
+ for (const item of req.items) {
+ const problem = itemProblem(item);
+ if (problem) return { ok: false, error: `${fetchWindowsItemLabel(item)}: ${problem}` };
+ }
+ const paths = getPaths();
+ const requestedAt = new Date().toISOString();
+ const slugs = [...new Set(req.items.map((i) => i.slug))];
+ return runManagedFunction({
+ kind: "fetch-windows",
+ queueKey: req.queueKey,
+ paths,
+ // A batch of ONE channel carries it, so /jobs and the media guard can name
+ // it; one spanning channels carries none, and the controller asks each
+ // channel's text guard itself.
+ ...(slugs.length === 1 ? { channelSlug: slugs[0] } : {}),
+ spec: {
+ kind: "fetch-windows",
+ // A spec needs a slug; the replay reads `params.items`.
+ slug: slugs[0],
+ params: {
+ items: req.items,
+ requestedBy: req.requestedBy,
+ queueKey: req.queueKey,
+ ...(req.manifest ? { manifest: req.manifest } : {}),
+ ...(req.siteId ? { siteId: req.siteId } : {}),
+ ...(req.maxHeight !== undefined ? { maxHeight: req.maxHeight } : {}),
+ },
+ },
+ fn: async (onLog, signal, setProgress, ctx) => {
+ const result = await fetchWindows({
+ paths,
+ items: req.items,
+ provenance: {
+ requestedBy: req.requestedBy,
+ manifest: req.manifest,
+ requestedAt,
+ },
+ maxHeight: req.maxHeight,
+ onLog,
+ signal,
+ drainSignal: ctx.drainSignal,
+ setProgress,
+ });
+ safeRevalidate([
+ ...new Set(req.items.map((i) => `/channels/${i.slug}/videos/${i.id}`)),
+ ]);
+ // A run that stopped short or lost a window did not do what it was
+ // asked: the job says so, and a re-run picks up the rest.
+ if (result.stopped && result.stopped !== "drain" && result.stopped !== "cancel") {
+ throw new Error(
+ `Stopped (${result.stopped}) with ${result.notAttempted.length} window(s) ` +
+ `not attempted — run it again later.`,
+ );
+ }
+ if (result.failed.length > 0) {
+ throw new Error(`${result.failed.length} window(s) failed to fetch.`);
+ }
+ },
+ });
+}
+
+// Fetch a list of windows ACROSS channels and platforms. A dry run answers with
+// the plan and starts nothing.
+export async function fetchWindowsAction(
+ req: FetchWindowsRequest & { siteId?: string },
+): Promise<FetchWindowsActionResult> {
+ const requestedBy = req.requestedBy.trim();
+ if (!requestedBy) {
+ return { ok: false, error: "requestedBy is required (who is asking for these bytes)." };
+ }
+ if (req.maxHeight !== undefined && !isFetchMaxHeight(req.maxHeight)) {
+ return {
+ ok: false,
+ error: `maxHeight ${req.maxHeight} is not a source height to cap a fetch at.`,
+ };
+ }
+ const { groups, cached, unresolved } = await planFetchWindows(req.items, req.queueKey);
+ const base = { groups, cached, unresolved, jobs: [], refused: [] };
+ if (req.dryRun) return { ok: true, dryRun: true, ...base };
+ if (groups.length === 0) return { ok: true, dryRun: false, ...base };
+
+ // A window is small, but it lands on the corpus disk: the floor an operator's
+ // click asks (manual mode), once, against the channels' text.
+ const paths = getPaths();
+ const disk = await diskGate(paths, getSettings(), {
+ mode: "manual",
+ dir: paths.channelsDir,
+ });
+ if (!disk.ok) {
+ return {
+ ok: false,
+ error:
+ `Low disk space: ${formatBytes(disk.freeBytes)} free, ` +
+ `${formatBytes(disk.thresholdBytes)} required. Free up space or ` +
+ `lower the floor in Settings.`,
+ };
+ }
+
+ const jobs: { platform: string; queueKey: string; jobId: string; items: number }[] = [];
+ const refused: { platform: string; error: string; items: number }[] = [];
+ for (const g of groups) {
+ const refusal = await groupRefusal(g.platform);
+ if (refusal) {
+ refused.push({ platform: g.platform, error: refusal, items: g.items.length });
+ continue;
+ }
+ let res: StreamActionResult;
+ try {
+ res = await fetchWindowsJobAction({
+ items: g.items,
+ requestedBy,
+ manifest: req.manifest,
+ siteId: req.siteId,
+ maxHeight: req.maxHeight,
+ queueKey: g.queueKey,
+ });
+ } catch (e) {
+ refused.push({ platform: g.platform, error: (e as Error).message, items: g.items.length });
+ continue;
+ }
+ if (!res.ok) {
+ refused.push({ platform: g.platform, error: res.error, items: g.items.length });
+ continue;
+ }
+ // Nobody reads the stream: the job's log is on disk.
+ void res.stream.cancel().catch(() => {});
+ jobs.push({
+ platform: g.platform,
+ queueKey: g.queueKey,
+ jobId: res.jobId,
+ items: g.items.length,
+ });
+ }
+ if (jobs.length === 0 && refused.length > 0) {
+ return {
+ ok: false,
+ error: refused.map((r) => `${r.platform}: ${r.error}`).join("; "),
+ };
+ }
+ return { ok: true, dryRun: false, groups, cached, unresolved, jobs, refused };
+}
+
+export type FetchMissingEvidenceResult =
+ | { ok: false; error: string }
+ | (Extract<FetchWindowsActionResult, { ok: true }> & {
+ siteId: string;
+ // Cited spans already on disk, and the ones no window fetch can fill.
+ onDisk: number;
+ unfetchable: UnfetchableEvidenceWindow[];
+ });
+
+// Every window a site's published reports cite and the disk does not hold
+// (publish/reportMedia.ts, missingEvidenceWindows), fetched as above. What the
+// Reports tab's "Fetch missing evidence" and `pnpm ops fetch-windows
+// {"siteId": …}` both run.
+export async function fetchMissingEvidenceAction(
+ siteId: string,
+ opts: { dryRun?: boolean; maxHeight?: number } = {},
+): Promise<FetchMissingEvidenceResult> {
+ const id = siteId.trim();
+ if (!isValidSiteId(id)) return { ok: false, error: `"${id}" is not a valid site id` };
+ const paths = getPaths();
+ if (!listSiteIds(paths).includes(id)) return { ok: false, error: `No site "${id}"` };
+ const need = await missingEvidenceWindows(paths, id);
+ const extra = { siteId: id, onDisk: need.onDisk, unfetchable: need.unfetchable };
+ if (need.missing.length === 0) {
+ return {
+ ok: true,
+ dryRun: opts.dryRun === true,
+ groups: [],
+ jobs: [],
+ refused: [],
+ cached: [],
+ unresolved: [],
+ ...extra,
+ };
+ }
+ const r = await fetchWindowsAction({
+ items: need.missing.map((w) => ({
+ slug: w.slug,
+ id: w.id,
+ from: w.from,
+ to: w.to,
+ clipId: w.clipId,
+ ...(w.reason ? { reason: w.reason.slice(0, 400) } : {}),
+ })),
+ requestedBy: "reports",
+ manifest: id,
+ siteId: id,
+ maxHeight: opts.maxHeight,
+ dryRun: opts.dryRun,
+ });
+ if (!r.ok) return r;
+ return { ...r, ...extra };
+}
diff --git a/editor/app/channels/actions.ts b/editor/app/channels/actions.ts
@@ -182,27 +182,53 @@ export async function probeChannelUrlAction(
}
}
+// What a create / rename / delete did, without the redirect. The form actions
+// below call these and redirect on `ok`; /api/ops calls them and answers with
+// JSON, because a route handler cannot follow a redirect() it did not ask for.
+// ONE body each, so the browser and an ops caller get the same refusal in the
+// same sentence (the ops layer's whole rule, api/ops/_lib.ts).
+export type ChannelLifecycleResult =
+ | {
+ ok: true;
+ slug: string;
+ // Jobs the step started on the way (create's first playlist or post
+ // fetch), so an ops caller can follow them.
+ jobIds: string[];
+ // Non-fatal problems after the move (rename's metadata migration).
+ warnings: string[];
+ }
+ | ({ ok: false } & NonNullable<FormErrorState>);
+
export async function createChannelAction(
_prev: ActionResult,
formData: FormData,
): Promise<ActionResult> {
+ const result = await createChannelFromForm(formData);
+ if (!result.ok) return { error: result.error, values: result.values };
+ redirect(`/channels/${result.slug}`);
+}
+
+export async function createChannelFromForm(
+ formData: FormData,
+): Promise<ChannelLifecycleResult> {
const values = formValues(formData);
+ const fail = (error: string) => ({ ok: false as const, error, values });
+ const jobIds: string[] = [];
let parsed;
try {
parsed = parseChannelForm(formData);
} catch (e) {
- return { error: (e as Error).message, values };
+ return fail((e as Error).message);
}
const { name, config } = parsed;
const slug = parsed.slug || slugify(name);
if (!slug) {
- return { error: "Could not derive a slug from the name", values };
+ return fail("Could not derive a slug from the name");
}
if (!isValidChannelSlug(slug)) {
- return {
- error: `"${slug}" is not a valid slug (letters, digits, ".", "_", "-"; must start with a letter or digit)`,
- values,
- };
+ return fail(
+ `"${slug}" is not a valid slug (letters, digits, ".", "_", "-"; must start with a letter or digit)`,
+ );
}
const paths = getPaths();
// Validate + plan the site-membership writes from the form's Sites section
@@ -218,10 +244,10 @@ export async function createChannelAction(
siteWrites = planSiteMembershipWrites(paths, slug, requests);
}
} catch (e) {
- return { error: (e as Error).message, values };
+ return fail((e as Error).message);
}
if (await channelExists(paths, slug)) {
- return { error: `Channel "${slug}" already exists`, values };
+ return fail(`Channel "${slug}" already exists`);
}
await createChannel(paths, slug, config);
// THE COMPILED TREES NAME THEIR CHANNELS. A channel created since the last
@@ -238,10 +264,9 @@ export async function createChannelAction(
await applySiteWrites(siteWrites, paths);
} catch (e) {
// The channel itself was created; don't redirect as if nothing happened.
- return {
- error: `Channel "${slug}" was created, but updating site memberships failed: ${(e as Error).message}. Open its Configure panel to retry.`,
- values,
- };
+ return fail(
+ `Channel "${slug}" was created, but updating site memberships failed: ${(e as Error).message}. Open its Configure panel to retry.`,
+ );
}
// URL-first onboarding: unless opted out ("Fetch playlist now", default on),
// immediately store the channel's playlist so undownloadedIds populates for
@@ -251,7 +276,10 @@ export async function createChannelAction(
if (config.url && formData.get("fetchPlaylist") != null) {
try {
const res = await storePlaylistAction(slug);
- if (res.ok) void res.stream.cancel();
+ if (res.ok) {
+ void res.stream.cancel();
+ jobIds.push(res.jobId);
+ }
} catch {
/* best-effort — the channel was created regardless */
}
@@ -262,7 +290,10 @@ export async function createChannelAction(
if (config.url && formData.get("fetchPostsNow") != null) {
try {
const res = await fetchPostsAction(slug);
- if (res.ok) void res.stream.cancel();
+ if (res.ok) {
+ void res.stream.cancel();
+ jobIds.push(res.jobId);
+ }
} catch {
/* best-effort — the channel was created regardless */
}
@@ -283,7 +314,7 @@ export async function createChannelAction(
revalidatePath("/channels");
revalidatePath("/sites");
revalidatePath("/");
- redirect(`/channels/${slug}`);
+ return { ok: true, slug, jobIds, warnings: [] };
}
export async function updateChannelAction(
@@ -453,6 +484,29 @@ export async function toggleChannelBuildInclusionAction(
return undefined;
}
+// SET, not toggle: the two exclusions above to a stated value, for a caller
+// that cannot see the current one before it clicks (/api/ops/channel-config).
+// `undefined` leaves a flag alone; `false` removes the key, as the toggles do.
+export async function setChannelExclusionsAction(
+ slug: string,
+ flags: { excludeFromBuild?: boolean; excludeFromCleanup?: boolean },
+): Promise<ActionResult> {
+ const paths = getPaths();
+ const existing = await readChannelConfig(paths, slug);
+ if (!existing) return { error: `Channel "${slug}" not found` };
+ const set: { excludeFromBuild?: true; excludeFromCleanup?: true } = {};
+ const unset: ("excludeFromBuild" | "excludeFromCleanup")[] = [];
+ for (const key of ["excludeFromBuild", "excludeFromCleanup"] as const) {
+ if (flags[key] === true) set[key] = true;
+ else if (flags[key] === false) unset.push(key);
+ }
+ if (Object.keys(set).length === 0 && unset.length === 0) return undefined;
+ await patchChannelConfig(paths, slug, set, { unset });
+ revalidatePath("/channels");
+ revalidatePath("/cleanup");
+ return undefined;
+}
+
// Toggle whether this channel's reclaimable bytes count toward the aggregate
// "cleanable data" total on the /cleanup page (and its sidebar badge). The
// cleanup sweeps themselves stay available regardless; this only flips the
@@ -565,13 +619,20 @@ export async function deleteChannelAction(
_prev: ActionResult,
formData: FormData,
): Promise<ActionResult> {
+ const result = await deleteChannelFromForm(slug, formData);
+ if (!result.ok) return { error: result.error, values: result.values };
+ redirect("/channels");
+}
+
+export async function deleteChannelFromForm(
+ slug: string,
+ formData: FormData,
+): Promise<ChannelLifecycleResult> {
const values = formValues(formData);
+ const fail = (error: string) => ({ ok: false as const, error, values });
const confirm = String(formData.get("confirmSlug") ?? "").trim();
if (confirm !== slug) {
- return {
- error: `Type the channel slug "${slug}" exactly to confirm deletion`,
- values,
- };
+ return fail(`Type the channel slug "${slug}" exactly to confirm deletion`);
}
// THE RENAME'S GUARD, AND DELETE NEEDED IT MORE. Renaming while a job runs
// orphans a registry entry keyed by the old slug; DELETING while one runs
@@ -581,18 +642,18 @@ export async function deleteChannelAction(
// races the writer for the tree and whichever loses reports an ENOENT nobody
// asked about.
const busy = channelMediaBusyReason(slug, "deleting it");
- if (busy) return { error: busy, values };
+ if (busy) return fail(busy);
// THE OTHER REFUSAL REACHES THE FORM THE SAME WAY. `deleteChannel` THROWS
// when `.relocating.json` is present — media in transition is not a channel
// anyone may delete — and an uncaught throw from a server action is a
// digest-shaped error page, not the sentence above it. Both refusals are
// refusals; they belong in the same place, on the same form, in the
- // operator's words. The `redirect` below stays OUTSIDE this: it throws
- // NEXT_REDIRECT as its control flow and a catch here would swallow it.
+ // operator's words. (The `redirect` is the caller's, OUTSIDE this: it
+ // throws NEXT_REDIRECT as its control flow and a catch here would swallow it.)
try {
await deleteChannel(getPaths(), slug);
} catch (e) {
- return { error: (e as Error).message, values };
+ return fail((e as Error).message);
}
// Same reason as createChannelAction: the deleted channel keeps a leaf in
// every compiled tree until something recompiles. A leaf matching nothing is
@@ -604,7 +665,7 @@ export async function deleteChannelAction(
}
revalidatePath("/channels");
revalidatePath("/");
- redirect("/channels");
+ return { ok: true, slug, jobIds: [], warnings: [] };
}
// Change a channel's slug (its on-disk directory name). High-friction: the
@@ -618,32 +679,38 @@ export async function renameChannelAction(
_prev: ActionResult,
formData: FormData,
): Promise<ActionResult> {
+ const result = await renameChannelFromForm(oldSlug, formData);
+ if (!result.ok) return { error: result.error, values: result.values };
+ redirect(`/channels/${result.slug}`);
+}
+
+export async function renameChannelFromForm(
+ oldSlug: string,
+ formData: FormData,
+): Promise<ChannelLifecycleResult> {
const values = formValues(formData);
+ const fail = (error: string) => ({ ok: false as const, error, values });
const confirm = String(formData.get("confirmSlug") ?? "").trim();
if (confirm !== oldSlug) {
- return {
- error: `Type the channel slug "${oldSlug}" exactly to confirm the rename`,
- values,
- };
+ return fail(`Type the channel slug "${oldSlug}" exactly to confirm the rename`);
}
const newSlug = String(formData.get("newSlug") ?? "").trim();
if (!newSlug) {
- return { error: "Enter a new slug", values };
+ return fail("Enter a new slug");
}
if (newSlug === oldSlug) {
- return { error: "The new slug is the same as the current one", values };
+ return fail("The new slug is the same as the current one");
}
if (!isValidChannelSlug(newSlug)) {
- return {
- error: `"${newSlug}" is not a valid slug (letters, digits, ".", "_", "-"; must start with a letter or digit)`,
- values,
- };
+ return fail(
+ `"${newSlug}" is not a valid slug (letters, digits, ".", "_", "-"; must start with a letter or digit)`,
+ );
}
const paths = getPaths();
const config = await readChannelConfig(paths, oldSlug);
- if (!config) return { error: `Channel "${oldSlug}" not found`, values };
+ if (!config) return fail(`Channel "${oldSlug}" not found`);
if (await channelExists(paths, newSlug)) {
- return { error: `Channel "${newSlug}" already exists`, values };
+ return fail(`Channel "${newSlug}" already exists`);
}
// THE REGISTRY IS HALF THE TRUTH, and this check used to be the other half's
@@ -654,13 +721,13 @@ export async function renameChannelAction(
// under it. One question, one answer, the same sentence the Storage panel
// and the bulk move say.
const busy = channelMediaBusyReason(oldSlug, "renaming it");
- if (busy) return { error: busy, values };
+ if (busy) return fail(busy);
let result;
try {
result = await renameChannel(paths, oldSlug, newSlug, config);
} catch (e) {
- return { error: (e as Error).message, values };
+ return fail((e as Error).message);
}
// THE PRIORITY DOCUMENT KEYS BY SLUG, so it has to follow the rename or the
// channel's tier, rank and per-operation overrides stay under a slug that no
@@ -685,7 +752,8 @@ export async function renameChannelAction(
);
}
// The directory move succeeded; any warnings are non-fatal metadata-migration
- // problems. Log them (we redirect on success, so there's no UI to show them).
+ // problems. Log them (the form redirects on success, so there's no UI to show
+ // them) and return them, for an ops caller, which can.
if (result.warnings.length > 0) {
console.warn(
`Channel rename ${oldSlug} -> ${newSlug} completed with warnings:`,
@@ -694,7 +762,7 @@ export async function renameChannelAction(
}
revalidatePath("/channels");
revalidatePath("/");
- redirect(`/channels/${newSlug}`);
+ return { ok: true, slug: newSlug, jobIds: [], warnings: result.warnings };
}
// ---------------------------------------------------------------------------
diff --git a/editor/app/channels/page.tsx b/editor/app/channels/page.tsx
@@ -357,55 +357,44 @@ export default async function ChannelsPage({
// fixed chrome and the rack between them is the only thing that moves.
// Below md there is no height and the document scrolls as it always has.
<div className="flex flex-col gap-3 md:h-[calc(100vh-3rem)] md:min-h-0">
- {/* NOWRAP ON md+, AND THE SUBTITLE IS WHAT GIVES. The freshness note is
- three lines of prose, not a tagline: at its max-w-3xl ceiling it
- claimed 768px of a 1216px content column, which left no room for the
- corpus-wide cluster and dropped it onto a band of its own — the same
- undifferentiated row of pills this change exists to remove. So the
- left block takes whatever is left and the paragraph rewraps into it
- (max-w-3xl stays a ceiling, never a floor), and the cluster keeps its
- intrinsic width at the top right. Below md the header still wraps and
- the cluster still stacks full-width. */}
- <header className="flex flex-wrap items-start justify-between gap-x-4 gap-y-2 shrink-0 md:flex-nowrap">
- <div className="min-w-0 md:flex-1">
+ {/* THE HEADER WRAPS; IT NEVER SQUEEZES. The title block has a 16rem
+ basis, so when the corpus-wide cluster cannot sit beside it at full
+ width the cluster drops onto its own line. It used to be nowrap with
+ the cluster shrink-0, and at a narrow md+ column (sidebar open) the
+ title block was left ~60px and its text ran one word per line. */}
+ <header className="flex flex-wrap items-start justify-between gap-x-4 gap-y-2 shrink-0">
+ <div className="min-w-0 flex-[1_1_16rem]">
<h1 className="text-2xl font-semibold">Channels</h1>
- {/* THE FRESHNESS NOTE IS THE PAGE'S CAVEAT, so it sits with the
- page's title. It used to be the last thing on the page, in 11px
- type under a floating bar, ~6,000px below the numbers it is a
- caveat ABOUT — which is to say it was written but not said. */}
+ {/* A STATUS LINE, NOT A LEAD PARAGRAPH. Every pipeline on the page is
+ read from each channel's last report, so the line says how old the
+ oldest one is and how many have none (the rack's Report column names
+ them). */}
{channels.length > 0 && (
<p
- className="max-w-3xl text-xs text-muted-foreground"
+ className="text-xs text-muted-foreground"
data-testid="channels-freshness"
>
<span className="tabular-nums">{channels.length}</span>{" "}
{channels.length === 1 ? "channel" : "channels"} ·{" "}
{freshness.oldest ? (
<>
- Every pipeline on this page — downloads, transcripts,
- digests and the speaker lanes — is read from each
- channel’s last report; the oldest on this page was
- generated{" "}
+ oldest report{" "}
<time dateTime={freshness.oldest}>
{new Date(freshness.oldest).toLocaleString()}
</time>
- . A just-finished job can take a moment to show up here.
</>
) : (
- <>
- No channel on this page has a generated report yet, so every
- pipeline reads empty. Run <em>Refresh report</em> from a row
- here, or <em>Update all reports</em> above, to populate them.
- </>
+ <>no reports yet — run <em>Update all reports</em></>
)}
{freshness.missing.length > 0 && freshness.oldest && (
<>
{" "}
- No report yet for {freshness.missing.join(", ")} — those rows
- read zero
- {sections
- ? ", and a group figure that counts one of them is marked with a trailing + to say it is a floor rather than a total."
- : "."}
+ ·{" "}
+ <span className="tabular-nums">
+ {freshness.missing.length}
+ </span>{" "}
+ without a report
+ {sections ? " (a group figure marked + is a floor)" : ""}
</>
)}
</p>
diff --git a/editor/app/jobs/components/JobProgressBars.tsx b/editor/app/jobs/components/JobProgressBars.tsx
@@ -158,6 +158,7 @@ const METRIC_LABELS: Record<JobProgressMetric, string> = {
digests: "Digests",
backfills: "Backfill",
scans: "Metadata scan",
+ clips: "Clip windows",
};
// Per-metric glyph for the compact line ("↓ 5/10").
@@ -167,6 +168,7 @@ const METRIC_PREFIX: Record<JobProgressMetric, string> = {
digests: "\u00b6 ",
backfills: "\u21ba ",
scans: "\u2315 ",
+ clips: "\u2702 ",
};
const METRIC_FILL: Record<JobProgressMetric, string> = {
@@ -175,6 +177,7 @@ const METRIC_FILL: Record<JobProgressMetric, string> = {
digests: "bg-info",
backfills: "bg-warning",
scans: "bg-info/60",
+ clips: "bg-success/60",
};
// One-line textual summary of a job's batch progress, e.g. "↓ 5/10 · ~2m left"
diff --git a/editor/app/jobs/jobReplayRegistry.ts b/editor/app/jobs/jobReplayRegistry.ts
@@ -57,6 +57,8 @@ import {
refreshVideoMetadataAction,
replayFetchWindowAction,
} from "../channels/[slug]/videos/[id]/videoActions";
+import { fetchWindowsJobAction } from "../channels/[slug]/videos/fetchWindowsAction";
+import type { FetchWindowsItem } from "yt-dlp-transcript-common/controller/fetchWindows";
import { reportsPrepareAction } from "../sites/lib/reportsPrepareAction";
import { reportsExportAction } from "../sites/lib/reportsExportAction";
import {
@@ -88,6 +90,33 @@ const num = (v: unknown): number | undefined =>
const strings = (v: unknown): string[] | undefined =>
Array.isArray(v) ? v.filter((k): k is string => typeof k === "string") : undefined;
+// A fetch-windows spec's items, keeping only the well-formed: the action
+// re-checks every window, so this only drops what is not even shaped like one.
+function windowItems(v: unknown): FetchWindowsItem[] {
+ if (!Array.isArray(v)) return [];
+ const out: FetchWindowsItem[] = [];
+ for (const raw of v) {
+ if (typeof raw !== "object" || raw === null) continue;
+ const r = raw as Record<string, unknown>;
+ const slug = str(r.slug);
+ const id = str(r.id);
+ const from = num(r.from);
+ const to = num(r.to);
+ if (!slug || !id || from === undefined || to === undefined) continue;
+ out.push({
+ slug,
+ id,
+ from,
+ to,
+ clipId: str(r.clipId),
+ reason: str(r.reason),
+ pad: num(r.pad),
+ webpageUrl: str(r.webpageUrl),
+ });
+ }
+ return out;
+}
+
// Flag-style params + the captured queueKey for a spec.
function params(spec: JobSpec): {
p: Record<string, unknown>;
@@ -170,6 +199,26 @@ export const JOB_REPLAY_HANDLERS: Record<string, ReplayHandler> = {
maxHeight: num(p.maxHeight),
});
},
+ // A batch of windows. Replay re-runs the same list on the same queue; the
+ // controller's cache check skips every window an earlier run fetched.
+ "fetch-windows": (spec) => {
+ const { p, queueKey } = params(spec);
+ const items = windowItems(p.items);
+ if (items.length === 0) {
+ return Promise.resolve({ ok: false, error: "Job spec has no windows." });
+ }
+ if (queueKey === undefined) {
+ return Promise.resolve({ ok: false, error: "Job spec is missing its queue." });
+ }
+ return fetchWindowsJobAction({
+ items,
+ requestedBy: str(p.requestedBy) ?? "unknown",
+ manifest: str(p.manifest),
+ siteId: str(p.siteId),
+ maxHeight: num(p.maxHeight),
+ queueKey,
+ });
+ },
// Both digest lanes replay through one action; the lane comes from params so a
// replayed metered run stays metered (and is refused if the lane has since
// been turned off, rather than quietly falling back to local).
@@ -318,6 +367,7 @@ export const JOB_REPLAY_HANDLERS: Record<string, ReplayHandler> = {
str(p.floor),
bool(p.force),
num(p.pages),
+ str(p.from),
);
},
// The ids are the spec's own (a capture is OF specific posts, unlike a
diff --git a/editor/app/operations/page.tsx b/editor/app/operations/page.tsx
@@ -27,11 +27,6 @@ export default async function OperationsPage() {
<h1 className="font-display text-2xl font-semibold tracking-tight">
Operations
</h1>
- <p className="text-sm text-muted-foreground">
- Every operation that turns this archive into a derived corpus, and what
- each one is doing right now. Open one to set its rules, arm its sweep or
- hold its lane.
- </p>
<OperationsBoard initial={initial} sync={sync} />
</div>
);
diff --git a/editor/app/review/page.tsx b/editor/app/review/page.tsx
@@ -34,9 +34,6 @@ export default async function ReviewPage() {
return (
<div className="flex flex-col gap-6">
<h1 className="text-2xl font-semibold">Review</h1>
- <p className="text-sm text-muted-foreground">
- Corpus review — findings a human decides, not work a lane runs.
- </p>
<AutoPausedSection rows={review.autoPaused} />
<DuplicatesSection
report={review.duplicates}
diff --git a/editor/app/saved-videos/page.tsx b/editor/app/saved-videos/page.tsx
@@ -57,10 +57,7 @@ export default async function SavedVideosPage() {
<header className="flex flex-col gap-1">
<h1 className="text-xl font-semibold">Saved videos</h1>
<p className="text-sm text-muted-foreground">
- Source video containers persisted by the keep-latest retention rule
- live in a separate store, leaving the main data volume holding only
- audio + transcripts. Configure a channel's keep-latest window and
- per-channel store dir under that channel's settings.
+ Set a channel's keep-latest window and store dir in its settings.
</p>
</header>
diff --git a/editor/app/sites/[siteId]/reports/page.tsx b/editor/app/sites/[siteId]/reports/page.tsx
@@ -6,6 +6,7 @@ import { formatBytes } from "yt-dlp-transcript-common/lib/format";
import { isCitedSite } from "yt-dlp-transcript-common/lib/site";
import { readReportMediaIndex } from "yt-dlp-transcript-common/publish/reportMedia";
import { PrepareReportMediaButton } from "../../components/PrepareReportMediaButton";
+import { FetchMissingEvidenceButton } from "../../components/FetchMissingEvidenceButton";
import { ExportReportsButton } from "../../components/ExportReportsButton";
import {
REPORT_EXPORT_FILENAMES,
@@ -125,13 +126,17 @@ export default async function SiteReportsPage({
<div>
<h2 className="text-lg font-semibold">Evidence media</h2>
<p className="text-sm text-muted-foreground">
- Cuts a clip of every span the published reports cite and copies every
- cited post capture, from the media already on disk, ready for the
- site's next build. Nothing is downloaded: a citation whose media
- is missing is listed as a problem — fetch its window or persist its
- video from the video page, then prepare again.
+ Fetch missing evidence downloads the window of every cited span
+ whose media is not on disk, one paced job per platform; Preview
+ lists them, and the spans no window can fill (a deleted video, a
+ channel off this site), without fetching. Prepare then cuts a clip
+ of every cited span and copies every cited post capture, from the
+ media on disk, ready for the site's next build — it downloads
+ nothing, and lists a citation whose media is still missing as a
+ problem.
</p>
</div>
+ <FetchMissingEvidenceButton siteId={siteId} />
<PrepareReportMediaButton siteId={siteId} />
<p className="text-sm text-muted-foreground">
Last prepare job:{" "}
diff --git a/editor/app/sites/components/FetchMissingEvidenceButton.tsx b/editor/app/sites/components/FetchMissingEvidenceButton.tsx
@@ -0,0 +1,91 @@
+"use client";
+
+import Link from "next/link";
+import { useState, useTransition } from "react";
+import { Button } from "yt-dlp-transcript-common/components/ui/button";
+import {
+ fetchMissingEvidenceAction,
+ type FetchMissingEvidenceResult,
+} from "../../channels/[slug]/videos/fetchWindowsAction";
+
+// "Fetch missing evidence": every window the site's published reports cite and
+// the disk does not hold, fetched as one paced job per platform
+// (fetch-windows). Preview answers the same question without starting
+// anything. It is not a streamed log because a run can start two jobs (YouTube
+// and Rumble side by side); each is linked, and its log is on /jobs.
+export function FetchMissingEvidenceButton({ siteId }: { siteId: string }) {
+ const [pending, start] = useTransition();
+ const [result, setResult] = useState<FetchMissingEvidenceResult | null>(null);
+ const run = (dryRun: boolean) => {
+ start(async () => {
+ setResult(await fetchMissingEvidenceAction(siteId, { dryRun }));
+ });
+ };
+ return (
+ <div className="flex flex-col gap-2">
+ <div className="flex flex-wrap items-center gap-2">
+ <Button type="button" variant="outline" disabled={pending} onClick={() => run(true)}>
+ Preview missing evidence
+ </Button>
+ <Button type="button" disabled={pending} onClick={() => run(false)}>
+ {pending ? "Working…" : "Fetch missing evidence"}
+ </Button>
+ </div>
+ {result && <FetchSummary result={result} />}
+ </div>
+ );
+}
+
+function FetchSummary({ result }: { result: FetchMissingEvidenceResult }) {
+ if (!result.ok) {
+ return (
+ <p role="alert" className="text-sm text-destructive">
+ {result.error}
+ </p>
+ );
+ }
+ const toFetch = result.groups.reduce((n, g) => n + g.items.length, 0);
+ return (
+ <div className="flex flex-col gap-1 text-sm" aria-label="Missing evidence">
+ <p>
+ {result.onDisk} cited span(s) on disk · {toFetch} window(s) to fetch
+ {result.cached.length > 0 && <> · {result.cached.length} already fetched</>}
+ {result.unfetchable.length > 0 && <> · {result.unfetchable.length} cannot be fetched</>}
+ </p>
+ {result.dryRun
+ ? result.groups.map((g) => (
+ <p key={g.queueKey} className="text-muted-foreground">
+ {g.platform}: {g.items.length} window(s) on {g.queueKey}
+ </p>
+ ))
+ : result.jobs.map((j) => (
+ <p key={j.jobId}>
+ {j.platform}: {j.items} window(s) —{" "}
+ <Link href={`/jobs/${j.jobId}`} className="font-mono underline">
+ {j.jobId}
+ </Link>
+ </p>
+ ))}
+ {!result.dryRun &&
+ result.refused.map((r) => (
+ <p key={r.platform} role="alert" className="text-destructive">
+ {r.platform}: {r.items} window(s) not started — {r.error}
+ </p>
+ ))}
+ {result.unresolved.map((u) => (
+ <p key={`${u.item.slug}/${u.item.id}/${u.item.from}`} className="text-destructive">
+ {u.item.clipId ?? `${u.item.slug}/${u.item.id}`}: {u.error}
+ </p>
+ ))}
+ {result.unfetchable.length > 0 && (
+ <ul className="list-disc pl-5 text-muted-foreground">
+ {result.unfetchable.map((u) => (
+ <li key={u.clipId}>
+ {u.clipId} ({u.slug}/{u.id}): {u.message}
+ </li>
+ ))}
+ </ul>
+ )}
+ </div>
+ );
+}
diff --git a/editor/app/sites/lib/reportListServer.ts b/editor/app/sites/lib/reportListServer.ts
@@ -1,12 +1,8 @@
import "server-only";
-import type { Dirent } from "node:fs";
-import { readdir } from "node:fs/promises";
-import path from "node:path";
import type { Paths } from "yt-dlp-transcript-common/lib/paths";
import { readJsonFile } from "yt-dlp-transcript-common/lib/jsonFile-server";
-import { isReportId } from "yt-dlp-transcript-common/lib/report/schema";
-import { siteDir, type Site } from "yt-dlp-transcript-common/lib/site";
-import { siteReportFile } from "yt-dlp-transcript-common/publish/reportMedia";
+import type { Site } from "yt-dlp-transcript-common/lib/site";
+import { listReportDirs, siteReportFile } from "yt-dlp-transcript-common/publish/reportMedia";
import {
publishableReportExports,
readReportExportManifest,
@@ -18,24 +14,11 @@ import { listAllJobs, type JobListEntry } from "yt-dlp-transcript-common/jobs/li
import { readJobMeta } from "yt-dlp-transcript-common/jobs/jobMeta";
import { reportRows, type ReportFileRead, type ReportRow } from "./reportList";
-// The report directories under `sites/<siteId>/reports/`: a directory whose
-// name is a report id. Anything else there (a stray file, a bad name) is not a
-// report and is not listed.
-async function reportDirIds(paths: Paths, siteId: string): Promise<string[]> {
- let entries: Dirent[];
- try {
- entries = await readdir(path.join(siteDir(paths, siteId), "reports"), { withFileTypes: true });
- } catch {
- return [];
- }
- return entries.filter((e) => e.isDirectory() && isReportId(e.name)).map((e) => e.name);
-}
-
// Every report of the site, published (in order) then drafts, each read and
// validated.
export async function readSiteReportRows(paths: Paths, site: Site): Promise<ReportRow[]> {
const published = site.reports ?? [];
- const dirIds = await reportDirIds(paths, site.siteId);
+ const dirIds = await listReportDirs(paths, site.siteId);
const ids = [...new Set([...published, ...dirIds])];
const reads = new Map<string, ReportFileRead>(
await Promise.all(
diff --git a/editor/app/sites/page.tsx b/editor/app/sites/page.tsx
@@ -84,11 +84,6 @@ export default async function SitesPage() {
+ New site
</Link>
</div>
- <p className="text-sm text-muted-foreground max-w-2xl">
- Each site is a selection + branding over the shared channel pool. The
- same channel can appear on several sites; its downloads are stored once
- and reused.
- </p>
{sites.length === 0 && (
<div className="flex flex-col gap-3 rounded border border-border p-4">
<p className="text-sm">
diff --git a/editor/app/storage/page.tsx b/editor/app/storage/page.tsx
@@ -22,17 +22,9 @@ export default async function StoragePage() {
<h1 className="text-2xl font-semibold">Storage</h1>
</div>
- <p className="text-sm text-muted-foreground max-w-3xl">
- A storage location is a named place a channel’s media — its big
- files, the audio and the raw live chat — may live, usually a second
- drive; a channel’s text never leaves the corpus volume. A channel
- is on a location when its <code>mediaDir</code> is under that
- location’s root; nothing is tagged, so moving a channel on or off
- one is a move, not a setting. When a drive comes back at a different
- mountpoint, <strong>re-point</strong> the location: it rewrites every
- channel’s symlink and <code>mediaDir</code> and moves no bytes. Move media onto a location from
- a channel’s <Link href="/channels" className="underline">Storage
- panel</Link>.
+ <p className="text-sm text-muted-foreground">
+ Move a channel’s media onto a location from its{" "}
+ <Link href="/channels" className="underline">Storage panel</Link>.
</p>
<StorageLocationsTable payload={payload} />
diff --git a/editor/app/tags/page.tsx b/editor/app/tags/page.tsx
@@ -50,13 +50,8 @@ export default async function TagsPage() {
<section className="flex flex-col gap-4">
<header className="flex flex-col gap-1">
<h1 className="text-2xl font-semibold">Tags</h1>
- <p className="max-w-3xl text-sm text-muted-foreground">
- A curated vocabulary that cuts across channels. A tag lands on a video
- three ways: a <strong>rule</strong> re-evaluated at every index build,
- a <strong>pin</strong> somebody made by hand, or an import from umtool
- — and a <strong>suppression</strong> rejects a rule's hit. Pins
- and suppressions are stored with their provenance; rule hits never
- are. Edit a rule, rebuild the index, done.
+ <p className="text-sm text-muted-foreground">
+ Rule edits apply at the next index build.
</p>
</header>
<EditorTagsClient
diff --git a/editor/app/widget/builder/page.tsx b/editor/app/widget/builder/page.tsx
@@ -27,16 +27,6 @@ export default async function WidgetBuilderPage() {
<div className="flex items-center justify-between">
<h1 className="text-2xl font-semibold">Monitor widget</h1>
</div>
- <p className="max-w-2xl text-sm text-muted-foreground">
- A read-only monitor with no sidebar, sized to sit in a small pinned
- window or an embedded <code className="font-mono"><iframe></code>.
- The board below is laid out the way the widget will be — drag the strips
- into the order you want to read them in, split them across rows and
- columns if you're giving it the width, and drop the ones you
- don't want into the tray. Give a row the leftover height with the
- rail on its left and the sections inside it scroll under their own
- headings. Then save it as a preset, or copy the link.
- </p>
<WidgetBuilder initialPresets={presets} />
</div>
);
diff --git a/editor/app/workers/page.tsx b/editor/app/workers/page.tsx
@@ -29,12 +29,6 @@ export default async function WorkersPage() {
return (
<div className="flex flex-col gap-4">
<h1 className="text-2xl font-semibold">Workers</h1>
- <p className="text-sm text-muted-foreground">
- Transcription workers — one slot each. Enable, disable, or drain a worker
- to free up a CPU/GPU for other programs, then turn it back on when done.
- Disabling every worker pauses running batches (they wait for a worker)
- instead of failing.
- </p>
<WorkersView initial={initial} />
<WorkersConfigForm
initial={settings.workers}
diff --git a/editor/e2e/channel-rename.spec.ts b/editor/e2e/channel-rename.spec.ts
@@ -1,3 +1,4 @@
+import { mkdir, writeFile } from "node:fs/promises";
import { test, expect, type Page } from "@playwright/test";
import { baseUrl } from "./baseUrl";
import {
@@ -5,6 +6,8 @@ import {
generateReport,
readJson,
resetData,
+ resolvePath,
+ writeChannelConfig,
writeSite,
} from "./helpers";
@@ -58,6 +61,60 @@ test("rename requires the exact current slug and then moves the channel", async
expect(site.channels.map((c) => c.slug)).toEqual([NEW]);
});
+// A SOCIAL CHANNEL HAS THE SAME DANGER ZONE. Its page returns before the stage
+// list a video channel gets, so it used to have no Rename at all — a misspelled
+// X channel could only be renamed by hand. The zone is collapsed under the posts
+// panel and opens on the same `?stage=danger` link.
+test("a social channel renames from its Danger zone and keeps its posts", async ({
+ page,
+}) => {
+ await resetData("empty");
+ const OLD_X = "misspeled-x";
+ const NEW_X = "misspelled-x";
+ await writeChannelConfig(OLD_X, {
+ handling: "transcribe",
+ sourceKind: "social",
+ platform: "twitter",
+ postFetcher: "x-gallery-dl",
+ socialHandle: "faketester",
+ name: "Fake Tester (X)",
+ url: "https://x.com/faketester",
+ });
+ const posts = resolvePath(`test-transcripts/channels/${OLD_X}/posts`);
+ await mkdir(posts, { recursive: true });
+ await writeFile(
+ `${posts}/2026-05.jsonl`,
+ JSON.stringify({
+ id: "1001",
+ slug: `${OLD_X}/1001`,
+ channelSlug: OLD_X,
+ author: "faketester",
+ createdAt: "2026-05-01T00:00:00.000Z",
+ uploadDate: "20260501",
+ text: "a post",
+ url: "https://x.com/faketester/status/1001",
+ platform: "x",
+ isReply: false,
+ isRepost: false,
+ links: [],
+ }) + "\n",
+ );
+
+ // Collapsed on a plain visit, open on the stage link.
+ await page.goto(`/channels/${OLD_X}`);
+ await expect(page.getByText("Danger zone")).toBeVisible();
+ await expect(page.getByRole("button", { name: "Rename channel" })).toBeHidden();
+ await page.goto(channelStage(OLD_X, "danger"));
+ await page.getByLabel("new slug").fill(NEW_X);
+ await page.getByLabel("confirm current slug").fill(OLD_X);
+ await page.getByRole("button", { name: "Rename channel" }).click();
+ await expect(page).toHaveURL(new RegExp(`/channels/${NEW_X}`));
+ // Still a social channel, with its one post.
+ await expect(page.locator("[data-social-panel]")).toContainText("Archived posts");
+ await expect(page.locator("[data-social-panel]")).toContainText("1");
+ expect((await page.request.get(`${baseUrl}/channels/${OLD_X}`)).status()).toBe(404);
+});
+
// RENAMING AND DELETING ARE REFUSED WHILE SOMETHING IS WRITING INTO data/.
//
// The rename always asked, but it asked the JOB REGISTRY alone — and the
diff --git a/editor/e2e/fetch-window.spec.ts b/editor/e2e/fetch-window.spec.ts
@@ -9,7 +9,7 @@
// Same token as /api/worker/* (the test server runs with
// WORKER_TOKEN=test-worker-token; see package.json dev:test).
-import { readdir, readFile, stat, writeFile } from "node:fs/promises";
+import { mkdir, readdir, readFile, stat, writeFile } from "node:fs/promises";
import { test, expect, type APIRequestContext } from "@playwright/test";
import {
generateReport,
@@ -51,6 +51,8 @@ async function invocations(): Promise<string> {
async function pollJob(
request: APIRequestContext,
jobId: string,
+ // A batch sleeps the clip-window gap (30–45 s) between two fetches.
+ timeout = 30_000,
): Promise<Record<string, unknown>> {
let last: Record<string, unknown> = {};
await expect
@@ -63,7 +65,7 @@ async function pollJob(
last = (await r.json()) as Record<string, unknown>;
return last.status as string;
},
- { timeout: 30_000 },
+ { timeout },
)
.not.toMatch(/^(queued|running)$/);
return last;
@@ -325,6 +327,84 @@ test("a 429 fails the job and puts the platform in cooldown", async ({
expect(body.cooldownMs).toBeGreaterThan(0);
});
+// THE BATCH: POST /api/ops/fetch-windows, one paced job per platform queue.
+// Three windows: one already on disk (no request, no pause), one the source
+// refuses with a 403, one that fetches. A single 403 is an item failure — a
+// removed Rumble page answers 403 too — so the run carries on past it and the
+// platform is NOT backed off; the job still ends failed, naming the window it
+// lost. Exactly one pause is owed (between the two network fetches), at the
+// clip-window floor of 30–45 s.
+test("a batch skips what is cached, survives one 403, and fetches the rest", async ({
+ request,
+}) => {
+ test.setTimeout(150_000);
+ await mkdir(resolvePath(rel(`data/${VIDEO}/clips`)), { recursive: true });
+ await writeFile(resolvePath(clipRel("0.00-30.00.mp4")), "already here");
+
+ const post = await request.post(`${baseUrl}/api/ops/fetch-windows`, {
+ headers: AUTH,
+ data: {
+ requestedBy: "umtool",
+ manifest: "demo-batch",
+ items: [
+ { slug: SLUG, id: VIDEO, from: 5, to: 10, clipId: "c01" },
+ {
+ slug: SLUG,
+ id: "win403vid1",
+ // The fake reads the sentinel out of the URL.
+ webpageUrl: "https://www.youtube.com/watch?v=win403vid1",
+ from: 5,
+ to: 10,
+ clipId: "c02",
+ },
+ { slug: SLUG, id: VIDEO, from: 50, to: 60, clipId: "c03" },
+ ],
+ },
+ });
+ expect(post.status()).toBe(200);
+ const body = (await post.json()) as {
+ cached: { clipId: string }[];
+ jobs: { platform: string; jobId: string; items: number }[];
+ jobId: string;
+ };
+ expect(body.cached.map((c) => c.clipId)).toEqual(["c01"]);
+ expect(body.jobs).toHaveLength(1);
+ expect(body.jobs[0]).toMatchObject({ platform: "youtube", items: 2 });
+
+ const finished = await pollJob(request, body.jobId, 90_000);
+ expect(finished.status).toBe("failed");
+ expect(String(finished.error)).toMatch(/1 window\(s\) failed to fetch/);
+ expect(String(finished.error)).toMatch(/1 fetched, 0 cached, 1 failed/);
+
+ expect(await exists(clipRel("50.00-60.00.mp4"))).toBe(true);
+ const sidecar = await readJson<{ requestedBy: string; manifest: string; clipId: string }>(
+ clipRel("50.00-60.00.json"),
+ );
+ expect(sidecar).toMatchObject({ requestedBy: "umtool", manifest: "demo-batch", clipId: "c03" });
+ expect(
+ await exists(`test-transcripts/channels/${SLUG}/data/win403vid1/clips/5.00-10.00.mp4`),
+ ).toBe(false);
+
+ // One 403 did not back the platform off: the next ask is not refused.
+ const next = await request.post(`${baseUrl}/api/ops/fetch-windows`, {
+ headers: AUTH,
+ data: {
+ requestedBy: "umtool",
+ dryRun: true,
+ items: [{ slug: SLUG, id: VIDEO, from: 100, to: 110 }],
+ },
+ });
+ expect(next.status()).toBe(200);
+ const plan = (await next.json()) as { groups: { platform: string }[] };
+ expect(plan.groups.map((g) => g.platform)).toEqual(["youtube"]);
+ const single = await request.post(`${baseUrl}/api/media/fetch-window`, {
+ headers: AUTH,
+ data: { channelSlug: SLUG, videoId: VIDEO, from: 100, to: 110, requestedBy: "umtool" },
+ });
+ expect(single.status()).toBe(202);
+ await pollJob(request, ((await single.json()) as { jobId: string }).jobId);
+});
+
// THE WHOLE RECORDING, when a window will not do — a tool that needs to re-cut
// freely, or a source whose windows would tile the entire runtime.
//
diff --git a/editor/e2e/fixtures/bin/fake-ytdlp.mjs b/editor/e2e/fixtures/bin/fake-ytdlp.mjs
@@ -573,6 +573,13 @@ async function main() {
);
process.exit(1);
}
+ // `win403`: Rumble's Cloudflare / googlevideo refusing the media request —
+ // a `network` failure, which the fetch-windows batch counts toward its
+ // two-in-a-row backoff.
+ if (url.toLowerCase().includes("win403")) {
+ process.stderr.write(`ERROR: [download] Got error: HTTP Error 403: Forbidden\n`);
+ process.exit(1);
+ }
await ensureDir(path.dirname(dest));
// Deterministic bytes, one chunk's worth, so a size assertion is stable.
await writeFile(dest, Buffer.alloc(CHUNK_BYTES, "w"));
diff --git a/editor/e2e/ops-api.spec.ts b/editor/e2e/ops-api.spec.ts
@@ -401,6 +401,92 @@ test("capture-posts refuses what it cannot capture, and an id not in the archive
expect(await listJobIds()).toEqual(before);
});
+// THE CHANNEL'S LIFECYCLE, OVER HTTP. create-channel is the New form,
+// rename-channel and delete-channel the Danger zone's two forms, channel-config
+// grows the Configure form's Sites section and the rack's two exclusion
+// switches, and `get channels` is the list a caller picks a slug from. Each is
+// the form's own action (actions.ts' *FromForm cores), so the files on disk are
+// what the browser would have left.
+test("a channel's lifecycle over ops: create on a site, list, re-site, rename, delete", async ({
+ request,
+}) => {
+ test.setTimeout(120_000);
+ await resetData("empty");
+ await settings();
+ await writeSite("mysite", { channels: [] });
+ const SITE = "test-transcripts/sites/mysite/site.json";
+ const list = async () => {
+ const res = await request.get(`${baseUrl}/api/ops/channels`, { headers: AUTH });
+ expect(res.status()).toBe(200);
+ return ((await res.json()) as {
+ channels: { slug: string; kind: string; sites: string[]; excludeFromBuild?: true }[];
+ }).channels;
+ };
+
+ const made = await ops(request, "create-channel", {
+ slug: "made-x",
+ fields: {
+ name: "Made (X)",
+ handling: "transcribe",
+ url: "https://x.com/made_user",
+ sourceKind: "social",
+ postFetcher: "x-gallery-dl",
+ },
+ sites: [{ siteId: "mysite" }],
+ });
+ expect(made.body).toEqual({ ok: true, slug: "made-x", jobIds: [] });
+ const cfg = await readJson<Record<string, unknown>>(
+ "test-transcripts/channels/made-x/config.json",
+ );
+ expect(cfg.sourceKind).toBe("social");
+ expect(cfg.socialHandle).toBe("made_user");
+ expect(
+ (await readJson<{ channels: { slug: string }[] }>(SITE)).channels.map((c) => c.slug),
+ ).toEqual(["made-x"]);
+ expect(await list()).toEqual([
+ { slug: "made-x", name: "Made (X)", kind: "social", platform: "twitter",
+ handling: "transcribe", url: "https://x.com/made_user", sites: ["mysite"] },
+ ]);
+
+ // Off every site, and out of the build — set, so a second call is a no-op.
+ for (let i = 0; i < 2; i++) {
+ const res = await ops(request, "channel-config", {
+ slug: "made-x",
+ sites: [],
+ excludeFromBuild: true,
+ });
+ expect(res.body).toEqual({ ok: true });
+ }
+ expect((await readJson<{ channels: unknown[] }>(SITE)).channels).toEqual([]);
+ const [row] = await list();
+ expect(row.sites).toEqual([]);
+ expect(row.excludeFromBuild).toBe(true);
+
+ const renamed = await ops(request, "rename-channel", {
+ slug: "made-x",
+ newSlug: "made-renamed-x",
+ });
+ expect(renamed.status).toBe(200);
+ expect((renamed.body as { slug?: string }).slug).toBe("made-renamed-x");
+ expect((await list()).map((c) => c.slug)).toEqual(["made-renamed-x"]);
+ expect(await pathExists("test-transcripts/channels/made-x")).toBe(false);
+
+ const refused = await ops(request, "delete-channel", {
+ slug: "made-renamed-x",
+ confirm: "made-x",
+ });
+ expect(refused.status).toBe(400);
+ expect(refused.body.error).toBe(
+ 'Type the channel slug "made-renamed-x" exactly to confirm deletion',
+ );
+ const gone = await ops(request, "delete-channel", {
+ slug: "made-renamed-x",
+ confirm: "made-renamed-x",
+ });
+ expect(gone.status).toBe(200);
+ expect(await list()).toEqual([]);
+});
+
test("channel-config round-trips a download filter and refuses a bad regex", async ({
page,
request,
diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md
@@ -2,6 +2,8 @@
## [Unreleased]
- **Posts withdrawn from a site leave a stand-in, not a stale copy.** When an X channel's posts stop being published on a site — while X posts are private, say — the site's build ships an empty posts manifest and an empty page at every address they were served from, sent with `Cache-Control: no-store`, so a reader (or the hub) asking for them gets "no posts" instead of the copy Cloudflare's edge kept for up to a week. The hub does the same for every X channel a public site carries.
+- **A quote with an editorial insertion is still a citation.** An inline citation whose quoted words hold brackets, `[“told [the mayor] so”](cite:id)`, is numbered, previews and links like any other, is listed in the reference list and its section's, and validation names its id when the report lacks it; it used to render as plain text. A label with a lone `]` is still not a link. The MCP's report reader counts it too. Needs `reports prepare` and a rebuild and deploy of each site with reports.
+- **A report can carry a video.** `report.json` `video` (`{ "src": "video.mp4", "poster": "poster.jpg", "caption": "…" }`, files in the report's directory: an mp4, a png/jpg/webp poster, a one-line caption) plays at the head of the report's page, under its header. Composing the site refuses a report whose video or poster is missing, or whose video is over the 24 MiB publish limit. Needs a rebuild and deploy of the site.
- **Search reads every English track of a video, and the transcript switches tracks.** Where a video has another English caption track whose words differ from its transcript — the uploaded captions beside the original audio's, a regional or auto-translated track — a query matches it too: a hit only that track holds says so ("in uploaded captions") and opens the transcript on that track at that moment, and a word both say is found once, in the transcript. The transcript reader shows a small "Track:" switcher beside the mode buttons on such a video; the transcript stays the default, and the choice rides on the share link (`vt`). Downloads and Copy MD take the track on show. Needs an index build and a rebuild and deploy of each site.
- **A citation of a Wayback Machine copy links its original and the copy.** A cited record downloaded from a Wayback capture shows "Original (may be gone)", the original at the cited second where its platform takes one, and "Wayback Machine copy, <capture date>", the capture page, which plays. Its moment link is the capture: a capture URL never takes a time param.
- **Transcripts read the original-audio captions.** Where a video has both, its transcript is YouTube's `en-orig` track (the captions of what was said) rather than the served `en`, which can reword it; a track with no text falls through to the next. Videos whose only captions are in cue blocks (some livestream recordings) have their text.
diff --git a/export/app/components/reports/ReportArticle.tsx b/export/app/components/reports/ReportArticle.tsx
@@ -269,6 +269,23 @@ export default function ReportArticle({ view }: { view: ReportPageView }) {
{view.subtitle && <p className="text-lg text-muted-foreground">{view.subtitle}</p>}
</header>
+ {/* The report as a video, when it has one. */}
+ {view.video && (
+ <figure data-report-video="" className="flex flex-col gap-2">
+ <video
+ controls
+ playsInline
+ preload="metadata"
+ poster={view.video.poster}
+ src={view.video.src}
+ className="aspect-video w-full rounded-lg border border-border bg-black"
+ />
+ {view.video.caption && (
+ <figcaption className="text-sm text-muted-foreground">{view.video.caption}</figcaption>
+ )}
+ </figure>
+ )}
+
{/* The quick take: the tally and the summary, then where to go next. */}
<section data-report-tier="1" aria-label="In brief" className="flex flex-col gap-4">
<TierMarker depth={1} />
diff --git a/package.json b/package.json
@@ -22,7 +22,7 @@
"e2e": "node scripts/worktree.mjs run -- pnpm --filter editor run e2e",
"wt": "node scripts/worktree.mjs",
"e2e:sharded": "node scripts/run-sharded-e2e.mjs",
- "test:scripts": "node --test scripts/*.test.mjs umtool/report-to-video/*.test.mjs umtool/lib/report/*.test.mjs",
+ "test:scripts": "node --test scripts/*.test.mjs umtool/report-to-video/*.test.mjs umtool/lib/report/*.test.mjs umtool/lib/annotations/*.test.mjs umtool/lib/articles/*.test.mjs",
"lint": "pnpm --filter export run lint",
"ops": "node scripts/archilyzer-ops.mjs"
},
diff --git a/scripts/archilyzer-ops.mjs b/scripts/archilyzer-ops.mjs
@@ -11,6 +11,7 @@
// pnpm ops <action> [--json '<body>' | --file <path>] [--wait]
// [--wait-timeout <seconds>] [--quiet]
// pnpm ops get channel <slug> [--counts]
+// pnpm ops get channels
// pnpm ops get tags [<tagId>]
// pnpm ops list
//
@@ -31,6 +32,11 @@
// pnpm ops import-video --json '{"slug":"demo-archive","url":"https://archive.org/details/example-item"}'
// pnpm ops import-archive-org --json '{"slug":"demo-archive","item":"example-item","match":"\\.mp4$"}' --wait
// pnpm ops channel-config --json '{"slug":"x","patch":{"downloadFilterExclude":"rerun"}}'
+// pnpm ops channel-config --json '{"slug":"x","sites":[{"siteId":"anilyzer"}]}'
+// pnpm ops channel-config --json '{"slug":"x","sites":[],"excludeFromBuild":true}'
+// pnpm ops create-channel --json '{"fields":{"name":"Example (X)","handling":"transcribe","url":"https://x.com/example"}}'
+// pnpm ops rename-channel --json '{"slug":"old-slug","newSlug":"new-slug"}'
+// pnpm ops delete-channel --json '{"slug":"x","confirm":"x"}'
// pnpm ops channel-priority --json '{"slugs":["x"],"operation":"download","tier":"paused"}'
// pnpm ops lane --json '{"lane":"download","held":true}'
// pnpm ops refresh-report --json '{"all":true}'
@@ -46,13 +52,18 @@
// pnpm ops reports-prepare --json '{"siteId":"demo-site"}' --wait
// pnpm ops reports-export --json '{"siteId":"demo-site","formats":["html","md"]}' --wait
// pnpm ops get channel the-quartering
+// pnpm ops get channels
// pnpm ops tags --json '{"op":"define","tag":{"id":"eva-collab","label":"Collab"}}'
// (a define is the WHOLE def: rules, and `sites` — the site ids the
// tag exists on, absent = every site — included)
// pnpm ops tag-videos --file ids.json
// pnpm ops persist-videos --file list.json --wait
+// pnpm ops fetch-windows --json '{"siteId":"demo-site","dryRun":true}'
+// pnpm ops fetch-windows --file windows.json --wait
// pnpm ops get tags eva-collab
// pnpm ops cut-release --json '{"workspace":"all","version":"next","commit":true}'
+// pnpm ops transcribe --json '{"path":"/abs/clip.mp4","start":120,"end":150}' --wait
+// pnpm ops transcribe --json '{"path":"/abs/a.wav","workerId":"parakeet-cpu","out":"/tmp/a.json"}' --wait
//
// --file reads the BODY from a JSON file, which is how a big one gets sent: a
// four-thousand-id tag-videos body is written by a script, not typed by a model
@@ -87,6 +98,9 @@ const GETTERS = {
// opt-in for the same reason the route makes it opt-in.
channel: (slug, counts) =>
`/api/ops/channel/${encodeURIComponent(slug)}${counts ? "?counts=1" : ""}`,
+ // Every channel, one line each: slug, name, kind, platform and the sites
+ // that carry it ([] = private to this editor).
+ channels: () => "/api/ops/channels",
// No argument: every definition with its pin/suppression counts. With one: a
// single tag's assignments, each carrying the provenance of the pin.
tags: (tag) =>
@@ -98,11 +112,16 @@ const GETTERS = {
// Nouns whose read takes no argument. `get channel` without a slug is a
// mistake; `get tags` without one is the whole vocabulary.
-const GET_ARG_OPTIONAL = new Set(["tags", "publish"]);
+const GET_ARG_OPTIONAL = new Set(["tags", "channels", "publish"]);
const ACTIONS = [
"channel-priority",
"channel-config",
+ // The channel's lifecycle: the New channel form, and the channel page's
+ // Danger → Rename and Danger → Delete. Synchronous; no job.
+ "create-channel",
+ "rename-channel",
+ "delete-channel",
"metadata-scan",
// ONE video's metadata.info.json re-read from its source ({slug, id}): no
// subtitles, no media; the rewrite lands in metadata.history.json.
@@ -166,13 +185,25 @@ const ACTIONS = [
// Persist specific videos, across channels, to the saved-video store
// ({items: [{slug, id}]}), paced and gated; a re-run resumes.
"persist-videos",
+ // Fetch clip windows through the managed path, one paced job per platform
+ // queue ({siteId} = a site's missing evidence, or {items, requestedBy}).
+ "fetch-windows",
// Cut a changelog's [Unreleased] into a dated release heading (release 10
// slice P). Synchronous. The same writer as `archilyzer release cut`, which
// needs no editor at all — this route exists only on an editor built from
// release 10 or later.
"cut-release",
+ // ONE local file, or a window of it, through a local transcription worker
+ // ({path, start?, end?, workerId?, out?}) — the corpus's own engine and
+ // model, as a job. With --wait the result JSON is what stdout carries.
+ "transcribe",
];
+// The log line a transcribe job ends with: this marker, then the result as
+// compact JSON. The same string is TRANSCRIBE_RESULT_MARKER in
+// common/controller/transcribeFile.ts.
+export const TRANSCRIBE_RESULT_MARKER = "@@transcribe-result ";
+
// The provenance a tag write from this CLI carries. Everything else ignores it.
function agentSource() {
return `agent:${process.env.ARCHILYZER_AGENT || "cli"}`;
@@ -296,6 +327,8 @@ export function parseArgs(argv) {
// a name are told apart; the route's own default (`agent:ops`) covers a
// caller that is neither.
...(action === "tag-videos" ? { defaultSource: agentSource() } : {}),
+ // A job whose log carries a RESULT, which --wait prints on stdout.
+ ...(action === "transcribe" ? { resultMarker: TRANSCRIBE_RESULT_MARKER } : {}),
wait,
quiet,
waitTimeout,
@@ -328,11 +361,45 @@ function printPreviewUrls(payload) {
for (const url of previewUrlsIn(payload)) console.error(`preview: ${url}`);
}
+// Pull a job's RESULT line out of its log as the log streams by. `feed` takes
+// each chunk and returns the text to echo — every complete line except the
+// result line; `finish` returns what is left and the parsed result (null when
+// the log never carried one). Line-buffered, because a poll may end mid-line.
+export function makeResultCapture(marker) {
+ let pending = "";
+ let result = null;
+ const take = (line) => {
+ if (line.startsWith(marker)) {
+ try {
+ result = JSON.parse(line.slice(marker.length));
+ return "";
+ } catch {
+ return `${line}\n`;
+ }
+ }
+ return `${line}\n`;
+ };
+ return {
+ feed(chunk) {
+ pending += chunk;
+ const lines = pending.split("\n");
+ pending = lines.pop() ?? "";
+ return lines.map(take).join("");
+ },
+ finish() {
+ const echo = pending ? take(pending).replace(/\n$/, "") : "";
+ pending = "";
+ return { echo, result };
+ },
+ };
+}
+
export function usage() {
return [
"Usage: pnpm ops <action> [--json '<body>' | --file <path>] [--wait]",
" [--wait-timeout <seconds>] [--quiet]",
" pnpm ops get channel <slug> [--counts]",
+ " pnpm ops get channels",
" pnpm ops get tags [<tagId>]",
" pnpm ops get publish",
" pnpm ops list",
@@ -373,6 +440,27 @@ export function usage() {
' report.md and evidence-pack.zip for the site\'s build to publish:',
' {"siteId", "reportId"?, "formats"?: ["html","pdf","md","zip"]}.',
"",
+ 'channel-config changes a channel as its Configure form does: {"slug"} and',
+ ' any of "patch" (form field names; "" clears one), "sites" (the WHOLE',
+ ' membership set: [{"siteId", "groupId"? | "newGroupName"?}], [] = on no',
+ ' site; an unknown site id is refused), "excludeFromBuild" and',
+ ' "excludeFromCleanup" (set to the value given, not toggled).',
+ "",
+ 'create-channel is the New channel form: {"fields": {"name", "handling":',
+ ' "youtube"|"transcribe", "url"?, "platform"?, "sourceKind"?, "postFetcher"?,',
+ ' "socialHandle"?, …}} with channel-config\'s patch keys; "slug"? (else',
+ ' derived from the name), "sites"? (absent = on no site). "fetchPlaylist",',
+ ' "fetchPostsNow" and "prioritizeDownload" are the form\'s checkboxes, OFF',
+ ' unless true; a job they start comes back as jobId(s), so --wait follows it.',
+ "",
+ 'rename-channel moves a channel to a new slug, as Danger → Rename does:',
+ ' {"slug", "newSlug"}. Refused while the channel is busy (a job, a lane unit,',
+ ' media in transition) or when the new slug is taken. Old links break.',
+ "",
+ 'delete-channel removes a channel\'s whole directory, as Danger → Delete does:',
+ ' {"slug", "confirm"} — "confirm" must repeat the slug. No undo outside the',
+ ' transcripts/ repo\'s own history.',
+ "",
'refresh-metadata re-reads ONE video\'s metadata.info.json from its source',
' (no subtitles, no media) on the platform\'s queue: {"slug", "id"}. The',
' job\'s log ends with what the source now says — live_status, formats,',
@@ -392,7 +480,9 @@ export function usage() {
'fetch-posts fetches a social channel\'s new posts: {"slug"}. "full": true',
' re-walks the whole timeline; "older": true walks back from the oldest',
" archived post through search (X; needs a login), saving its place for",
- ' the next run, down to "floor": "YYYY-MM-DD" when given. "limit": N caps',
+ ' the next run, down to "floor": "YYYY-MM-DD" when given; "from":',
+ ' "YYYY-MM-DD" starts the walk afresh there, replacing its saved place',
+ ' (and a "complete") — for a gap above one surviving old post. "limit": N caps',
' the posts one run reads; "pages": N caps the pages (a forum thread: its',
' latest N pages). "full" and "older" together are refused. An',
' older walk over an account that shows no posts (nothing archived, and',
@@ -421,6 +511,18 @@ export function usage() {
" queue; a low disk or a rate limit stops it, and running the same body",
" again resumes — saved videos are skipped.",
"",
+ 'fetch-windows fetches clip windows, one paced job per platform queue',
+ ' (YouTube and Rumble side by side): {"siteId"} fetches every window the',
+ ' site\'s published reports cite and the disk does not hold; {"items":',
+ ' [{"slug", "id", "from", "to", "clipId"?, "reason"?, "pad"?,',
+ ' "webpageUrl"?}, ...], "requestedBy", "manifest"?} fetches a list.',
+ ' "maxHeight" caps the source height (default 720). "dryRun": true lists',
+ ' the windows per platform, the ones already on disk ("cached") and the',
+ ' ones no fetch can fill ("unfetchable": deleted, off the site) and starts',
+ ' nothing. A platform cooling down or held is refused for its group; a',
+ ' 429, or two 403s in a row, backs the platform off and stops its job.',
+ ' Running the same body again resumes — fetched windows are cached.',
+ "",
'"preview": "<branch>" on deploy-site or build-deploy makes it a Cloudflare',
" Pages PREVIEW instead of production: the same bundle goes to a branch",
" alias, https://<branch>.<project>.pages.dev, and the live site is left",
@@ -447,6 +549,17 @@ export function usage() {
" (an older one answers 404); with no editor running, `archilyzer release",
" cut` does the same locally.",
"",
+ 'transcribe runs ONE local file through a local transcription worker, as a',
+ ' job: {"path": "/abs/file"} (audio or video), "start"/"end" (seconds) for a',
+ ' window, "workerId" (a settings worker id; default: the one auto-transcribe',
+ ' would get), "out" (an absolute path for the result JSON, never inside the',
+ " corpus). The result is {path, window, worker: {id, appId, model, device},",
+ " cues: [{start, end, text}], text, ...}, cue times on the file's own clock.",
+ ' "words": true adds words: [{w, start, end, conf?}] on the same clock -- from',
+ " parakeet, which keeps its word timestamps; [] from an engine that does not.",
+ " With --wait it is printed on stdout (the response and the log go to",
+ " stderr), so `pnpm ops transcribe ... --wait | jq -r .text` works.",
+ "",
"Env: ARCHILYZER_EDITOR_URL (default http://localhost:3001), WORKER_TOKEN,",
" ARCHILYZER_AGENT (provenance of a tag write; default \"cli\")",
].join("\n");
@@ -510,7 +623,13 @@ export async function followJob(jobId, quiet, opts = {}) {
const res = await doFetch(`${base}/api/jobs/${id}/log?from=${from}`);
if (!res.ok) return null;
const payload = await res.json();
- if (payload.content && !quiet) process.stderr.write(payload.content);
+ if (payload.content) {
+ // onContent sees every chunk, quiet or not, and says what to echo.
+ const echo = opts.onContent
+ ? opts.onContent(payload.content)
+ : payload.content;
+ if (echo && !quiet) process.stderr.write(echo);
+ }
// Only advance once the chunk is in hand: a poll that failed halfway
// must re-ask for the same offset.
from = payload.nextOffset ?? from;
@@ -653,7 +772,10 @@ async function main() {
console.error(`HTTP ${res.status}: ${text.slice(0, 500)}`);
return 1;
}
- console.log(JSON.stringify(payload, null, 2));
+ // A result-carrying action under --wait keeps stdout for the RESULT: the
+ // response goes to stderr with the log.
+ const resultMode = Boolean(parsed.wait && parsed.resultMarker);
+ (resultMode ? console.error : console.log)(JSON.stringify(payload, null, 2));
if (!res.ok || payload.ok === false) return 1;
if (!parsed.wait) {
printPreviewUrls(payload);
@@ -678,9 +800,21 @@ async function main() {
}
let worst = 0;
for (const jobId of jobIds) {
+ const capture = resultMode ? makeResultCapture(parsed.resultMarker) : null;
const status = await followJob(jobId, parsed.quiet, {
timeoutSeconds: parsed.waitTimeout ?? 0,
+ ...(capture ? { onContent: capture.feed } : {}),
});
+ if (capture) {
+ const { echo, result } = capture.finish();
+ if (echo && !parsed.quiet) process.stderr.write(echo);
+ if (result !== null) {
+ console.log(JSON.stringify(result, null, 2));
+ } else if (status === "done") {
+ console.error(`[${jobId}] finished but its log carries no result`);
+ worst = 1;
+ }
+ }
console.error(`[${jobId}] ${status}`);
if (status !== "done") worst = 1;
}
diff --git a/scripts/archilyzer-ops.test.mjs b/scripts/archilyzer-ops.test.mjs
@@ -6,7 +6,9 @@
import assert from "node:assert/strict";
import test from "node:test";
import {
+ TRANSCRIBE_RESULT_MARKER,
followJob,
+ makeResultCapture,
parseArgs,
previewUrlsIn,
usage,
@@ -462,6 +464,15 @@ test("persist-videos is a POST to its route, named in the usage", () => {
assert.match(usage(), /persist-videos/);
});
+test("fetch-windows is a POST to its route, named in the usage", () => {
+ const p = parseArgs(["fetch-windows", "--json", '{"siteId":"demo-site","dryRun":true}']);
+ assert.equal(p.method, "POST");
+ assert.equal(p.path, "/api/ops/fetch-windows");
+ assert.deepEqual(p.body, { siteId: "demo-site", dryRun: true });
+ assert.equal(parseArgs(["fetch-windows", "--file", "windows.json"]).bodyFile, "windows.json");
+ assert.match(usage(), /fetch-windows fetches clip windows/);
+});
+
test("feed-metadata posts {slug, dryRun} to /api/ops/feed-metadata", () => {
const p = parseArgs(["feed-metadata", "--json", '{"slug":"demo-channel","dryRun":true}']);
assert.equal(p.method, "POST");
@@ -485,3 +496,62 @@ test("refresh-metadata is a POST to its route, named in the usage", () => {
assert.match(usage(), /Actions:.*metadata-scan, refresh-metadata/);
assert.match(usage(), /refresh-metadata re-reads ONE video/);
});
+
+// --- transcribe: a job whose log carries a RESULT ---------------------------
+
+test("transcribe posts its body and carries the result marker", () => {
+ const p = parseArgs([
+ "transcribe",
+ "--json",
+ '{"path":"/abs/clip.mp4","start":120,"end":150}',
+ "--wait",
+ ]);
+ assert.equal(p.method, "POST");
+ assert.equal(p.path, "/api/ops/transcribe");
+ assert.deepEqual(p.body, { path: "/abs/clip.mp4", start: 120, end: 150 });
+ assert.equal(p.wait, true);
+ assert.equal(p.resultMarker, TRANSCRIBE_RESULT_MARKER);
+ // No other action carries one.
+ assert.equal(parseArgs(["sync", "--json", '{"slug":"x"}']).resultMarker, undefined);
+ assert.match(usage(), /transcribe runs ONE local file through a local transcription worker/);
+ assert.match(usage(), /it is printed on stdout/);
+});
+
+test("the result line is taken out of the log, across chunk boundaries", () => {
+ const result = { cues: [{ start: 120.5, end: 121, text: "hi" }], text: "hi" };
+ const line = `${TRANSCRIBE_RESULT_MARKER}${JSON.stringify(result)}`;
+ const cap = makeResultCapture(TRANSCRIBE_RESULT_MARKER);
+ let echoed = "";
+ // The log arrives in polls that split lines anywhere — the result line too.
+ const log = `Transcribe /abs/clip.mp4\nprogress 50%\n${line}\nafter\n`;
+ for (const chunk of [log.slice(0, 7), log.slice(7, 40), log.slice(40, 60), log.slice(60)]) {
+ echoed += cap.feed(chunk);
+ }
+ const { echo, result: got } = cap.finish();
+ echoed += echo;
+ assert.deepEqual(got, result);
+ assert.equal(echoed, "Transcribe /abs/clip.mp4\nprogress 50%\nafter\n");
+});
+
+test("a log with no result line captures null and echoes everything", () => {
+ const cap = makeResultCapture(TRANSCRIBE_RESULT_MARKER);
+ const echoed = cap.feed("one\ntwo");
+ const { echo, result } = cap.finish();
+ assert.equal(result, null);
+ assert.equal(echoed + echo, "one\ntwo");
+});
+
+test("followJob hands each chunk to onContent", async () => {
+ const seen = [];
+ const editor = fakeEditor({
+ log: [
+ { content: "a\n", nextOffset: 2, status: "running" },
+ { content: "b\n", nextOffset: 4, status: "done" },
+ ],
+ });
+ const status = await follow("j1", editor, {
+ onContent: (c) => (seen.push(c), ""),
+ });
+ assert.equal(status, "done");
+ assert.deepEqual(seen, ["a\n", "b\n"]);
+});
diff --git a/scripts/parakeet-stitch.mjs b/scripts/parakeet-stitch.mjs
@@ -48,6 +48,9 @@
// --max-cue <sec> cap a single cue's duration (default 8)
// -o, --output <f> output file (alternative to the positional arg)
// --keep-temp keep the temp working dir (for debugging)
+// --words also write the stitched words, as a top-level
+// `words: [{w, start, end, conf}]` beside chunk_data
+// (seconds; readers of chunk_data ignore it)
// -h, --help show this help
import { execFile } from "node:child_process";
@@ -92,6 +95,7 @@ function parseArgs(argv) {
gap: 0.8,
maxCue: 8,
keepTemp: false,
+ words: false,
output: undefined,
audio: undefined,
};
@@ -118,6 +122,7 @@ function parseArgs(argv) {
case "--max-cue": opts.maxCue = Number(next()); break;
case "-o": case "--output": opts.output = next(); break;
case "--keep-temp": opts.keepTemp = true; break;
+ case "--words": opts.words = true; break;
default:
if (a.startsWith("-")) fail(`unknown option ${a}`);
positional.push(a);
@@ -136,6 +141,19 @@ function numEnv(v, dflt) {
const round3 = (n) => Math.round(n * 1000) / 1000;
+// The stitched words as --words writes them: absolute seconds, empty words
+// dropped, conf kept only when the engine gave one.
+function wordsOut(stitched) {
+ return stitched
+ .filter((w) => String(w.w ?? "").trim() !== "")
+ .map((w) => ({
+ w: String(w.w).trim(),
+ start: round3(w.start),
+ end: round3(w.end),
+ ...(typeof w.conf === "number" ? { conf: w.conf } : {}),
+ }));
+}
+
// Seconds -> m:ss (or h:mm:ss). Used for the per-video ETA in progress lines.
function formatClock(totalSeconds) {
const s = Math.max(0, Math.floor(totalSeconds));
@@ -402,6 +420,7 @@ async function main() {
chunks: chunkData.length,
text,
chunk_data: chunkData,
+ ...(opts.words ? { words: wordsOut(stitched) } : {}),
};
const json = JSON.stringify(doc);
diff --git a/umtool/app/api/notes/context/route.ts b/umtool/app/api/notes/context/route.ts
@@ -0,0 +1,32 @@
+import { errorResponse, targetFrom } from "@/lib/annotations/server";
+import { digest } from "@/lib/annotations/digest.mjs";
+import { readNotes } from "@/lib/annotations/store.mjs";
+
+export const dynamic = "force-dynamic";
+
+// The agent brief: the same markdown `umtool notes <target>` prints, as
+// text/plain, so "Copy agent brief" and `curl` hand over the words an agent
+// would read on the command line.
+//
+// /api/notes/context?article=<site>/<report>[&status=open|resolved|all]
+// /api/notes/context?project=<id>[&status=…]
+export async function GET(request: Request) {
+ try {
+ const url = new URL(request.url);
+ const target = await targetFrom({ article: url.searchParams.get("article"), project: url.searchParams.get("project") });
+ const s = url.searchParams.get("status");
+ const status = s === "resolved" || s === "all" ? s : "open";
+ const read = await readNotes(target.file);
+ const cli = `umtool notes ${target.id}`;
+ const body =
+ !read.doc && !read.error
+ ? `No notes on ${target.id}.\n`
+ : `${await digest(
+ { kind: target.kind as "article" | "video-project", id: target.id, file: target.file, doc: read.doc, ...(read.error ? { error: read.error } : {}) },
+ { status, projectDir: "dir" in target ? target.dir : undefined },
+ )}\n\nRead these again with \`${cli}\` (from the repo checkout).\n`;
+ return new Response(body, { headers: { "content-type": "text/plain; charset=utf-8", "cache-control": "no-store" } });
+ } catch (err) {
+ return errorResponse(err);
+ }
+}
diff --git a/umtool/app/api/notes/route.ts b/umtool/app/api/notes/route.ts
@@ -0,0 +1,51 @@
+import { errorResponse, targetFrom } from "@/lib/annotations/server";
+import { readNotes } from "@/lib/annotations/store.mjs";
+import { writeNote } from "@/lib/annotations/targets.mjs";
+
+export const dynamic = "force-dynamic";
+
+// Notes on an article or a report-video project (lib/annotations/, docs/notes.md).
+//
+// GET ?article=<site>/<report> | ?project=<id>
+// { subject, file, token, doc, source, error? } -- doc null when
+// there are no notes; `source` is what a first write would record.
+// POST same params, body { token, op: { op: "add" | "edit" | "status" |
+// "reply" | "delete" | "delete-reply" | "source", ... } }
+// { doc, token, note }. 409 when `token` is stale or the file on
+// disk is not a notes doc; 400 for a refused op or target.
+//
+// Every write from here is the OPERATOR's. An agent writes through
+// `umtool notes`, which stamps "agent"; there is no way to claim to be one here.
+const NO_STORE = { "cache-control": "no-store" };
+
+function params(request: Request) {
+ const url = new URL(request.url);
+ return { article: url.searchParams.get("article"), project: url.searchParams.get("project") };
+}
+
+export async function GET(request: Request) {
+ try {
+ const target = await targetFrom(params(request));
+ const read = await readNotes(target.file);
+ const source = read.doc?.source ?? (await target.source());
+ return Response.json(
+ { subject: target.subject, file: target.file, token: read.token, doc: read.doc, source: source ?? null, ...(read.error ? { error: read.error } : {}) },
+ { headers: NO_STORE },
+ );
+ } catch (err) {
+ return errorResponse(err);
+ }
+}
+
+export async function POST(request: Request) {
+ try {
+ const target = await targetFrom(params(request));
+ const body = (await request.json().catch(() => null)) as { token?: unknown; op?: unknown } | null;
+ if (!body || typeof body.op !== "object" || body.op === null) return Response.json({ error: "body needs an op" }, { status: 400 });
+ const token = typeof body.token === "string" ? body.token : null;
+ const out = await writeNote(target, body.op as Record<string, unknown>, { by: "operator", token });
+ return Response.json(out, { headers: NO_STORE });
+ } catch (err) {
+ return errorResponse(err);
+ }
+}
diff --git a/umtool/app/api/report/chrome/route.ts b/umtool/app/api/report/chrome/route.ts
@@ -1,3 +1,4 @@
+import { withEditNotes } from "@/lib/report/guard";
import { ChromeRefused, StaleToken, manifestToken, updateChrome } from "@/lib/report/manifest.mjs";
import { resolveReport } from "@/lib/report/serve.mjs";
import { DECK_DEFAULTS, deckOn, resolveDeck, validateChrome } from "umtool-report-to-video/deck";
@@ -53,15 +54,18 @@ export async function PUT(request: Request) {
if ("error" in r) return Response.json({ error: r.error }, { status: r.status });
try {
- const res = await updateChrome(r.project.dir, (body.chrome ?? null) as Record<string, unknown> | null, {
- token: body.token === undefined ? null : String(body.token),
- });
+ const { result: res, editNotes } = await withEditNotes(r.project, () =>
+ updateChrome(r.project.dir, (body.chrome ?? null) as Record<string, unknown> | null, {
+ token: body.token === undefined ? null : String(body.token),
+ }),
+ );
return Response.json(
{
ok: true,
chrome: res.chrome,
deck: res.chrome ? resolveDeck({ ...(r.manifest.render ?? {}), chrome: res.chrome }) : null,
token: res.token,
+ editNotes,
},
{ headers: noStore },
);
diff --git a/umtool/app/api/report/claim/route.ts b/umtool/app/api/report/claim/route.ts
@@ -1,3 +1,4 @@
+import { withEditNotes } from "@/lib/report/guard";
import { StaleToken, manifestToken, updateClaim } from "@/lib/report/manifest.mjs";
import { resolveClaim } from "@/lib/report/serve.mjs";
import { readClaimDetail } from "@/lib/projects/report.mjs";
@@ -50,11 +51,16 @@ export async function PUT(request: Request) {
}
try {
- const { entry, token } = await updateClaim(r.project.dir, claimId, patch, {
- token: body.token === undefined ? null : String(body.token),
- });
+ const {
+ result: { entry, token },
+ editNotes,
+ } = await withEditNotes(r.project, () =>
+ updateClaim(r.project.dir, claimId, patch, {
+ token: body.token === undefined ? null : String(body.token),
+ }),
+ );
const detail = await readClaimDetail(r.project.dir, claimId);
- return Response.json({ claim: entry, gaps: detail?.gaps ?? [], token });
+ return Response.json({ claim: entry, gaps: detail?.gaps ?? [], token, editNotes });
} catch (err) {
// A stale token is a 409 and never a silent overwrite.
if (err instanceof StaleToken) {
diff --git a/umtool/app/api/report/cut/route.ts b/umtool/app/api/report/cut/route.ts
@@ -1,3 +1,4 @@
+import { withEditNotes } from "@/lib/report/guard";
import { StaleToken, updateClip } from "@/lib/report/manifest.mjs";
import { cuesInWindow } from "@/lib/projects/report.mjs";
import { resolveClip } from "@/lib/report/serve.mjs";
@@ -63,14 +64,16 @@ export async function POST(request: Request) {
}
try {
- const res = await updateClip(
- project.dir,
- clipId,
- { cutStart: hit.cutStart, cutEnd: hit.cutEnd },
- { token: body.token === undefined ? null : String(body.token) },
+ const { result: res, editNotes } = await withEditNotes(project, () =>
+ updateClip(
+ project.dir,
+ clipId,
+ { cutStart: hit.cutStart, cutEnd: hit.cutEnd },
+ { token: body.token === undefined ? null : String(body.token) },
+ ),
);
return Response.json(
- { ok: true, entry: res.entry, token: res.token, score: hit.score, matched: hit.matched },
+ { ok: true, entry: res.entry, token: res.token, score: hit.score, matched: hit.matched, editNotes },
{ headers: { "cache-control": "no-store" } },
);
} catch (e) {
diff --git a/umtool/app/api/report/moment/route.ts b/umtool/app/api/report/moment/route.ts
@@ -0,0 +1,26 @@
+import { projectRef } from "@/lib/projects";
+import { resolveMomentOnFile } from "@/lib/report/moments.mjs";
+
+export const dynamic = "force-dynamic";
+
+// GET ?project=<id>&file=<project-relative mp4>&t=<seconds>[&duration=<seconds>]
+//
+// What is on screen at `t` of a rendered cut (lib/report/moments.mjs): the
+// entry, its onscreen title and quote, the source second and its archive link,
+// and `approx` when the file and the build's schedule disagree. A timed note
+// stores this at write time. `{ entry: null }` when the file has no schedule
+// (a preview made without a build) -- the note then keeps `t` alone.
+export async function GET(request: Request) {
+ const url = new URL(request.url);
+ const project = await projectRef(url.searchParams.get("project") ?? "");
+ if (!project) return Response.json({ error: "no such project" }, { status: 404 });
+ const file = url.searchParams.get("file") ?? "";
+ if (!file || file.startsWith("/") || file.split("/").some((s) => s === ".." || s === "")) {
+ return Response.json({ error: "file must be relative to the project" }, { status: 400 });
+ }
+ const t = Number(url.searchParams.get("t"));
+ if (!Number.isFinite(t) || t < 0) return Response.json({ error: "t must be seconds ≥ 0" }, { status: 400 });
+ const d = Number(url.searchParams.get("duration"));
+ const r = await resolveMomentOnFile(project.dir, file, t, Number.isFinite(d) && d > 0 ? d : null);
+ return Response.json(r, { headers: { "cache-control": "no-store" } });
+}
diff --git a/umtool/app/api/report/onscreen/route.ts b/umtool/app/api/report/onscreen/route.ts
@@ -1,3 +1,4 @@
+import { withEditNotes } from "@/lib/report/guard";
import { StaleToken, manifestToken, updateOnscreen } from "@/lib/report/manifest.mjs";
import { postRows, scheduleForPreview } from "@/lib/report/onscreen.mjs";
import { resolveReport } from "@/lib/report/serve.mjs";
@@ -86,12 +87,14 @@ export async function PUT(request: Request) {
if ("error" in r) return Response.json({ error: r.error }, { status: r.status });
try {
- const res = await updateOnscreen(
- r.project.dir,
- body.onscreen as Record<string, { title?: string; subtitle?: string; claim?: Claim | null } | null>,
- { token: body.token === undefined ? null : String(body.token) },
+ const { result: res, editNotes } = await withEditNotes(r.project, () =>
+ updateOnscreen(
+ r.project.dir,
+ body.onscreen as Record<string, { title?: string; subtitle?: string; claim?: Claim | null } | null>,
+ { token: body.token === undefined ? null : String(body.token) },
+ ),
);
- return Response.json({ ok: true, onscreen: res.onscreen, claims: res.claims, token: res.token }, { headers: noStore });
+ return Response.json({ ok: true, onscreen: res.onscreen, claims: res.claims, token: res.token, editNotes }, { headers: noStore });
} catch (e) {
if (e instanceof StaleToken) {
return Response.json(
diff --git a/umtool/app/api/report/posts/route.ts b/umtool/app/api/report/posts/route.ts
@@ -1,3 +1,4 @@
+import { withEditNotes } from "@/lib/report/guard";
import { PostsRefused, StaleToken, updatePosts } from "@/lib/report/manifest.mjs";
import { resolveReport } from "@/lib/report/serve.mjs";
@@ -28,12 +29,14 @@ export async function PUT(request: Request) {
if ("error" in r) return Response.json({ error: r.error }, { status: r.status });
try {
- const res = await updatePosts(
- r.project.dir,
- body.posts as Record<string, { attachTo?: string | null; hide?: boolean }>,
- { token: body.token === undefined ? null : String(body.token) },
+ const { result: res, editNotes } = await withEditNotes(r.project, () =>
+ updatePosts(
+ r.project.dir,
+ body.posts as Record<string, { attachTo?: string | null; hide?: boolean }>,
+ { token: body.token === undefined ? null : String(body.token) },
+ ),
);
- return Response.json({ ok: true, posts: res.posts, token: res.token }, { headers: noStore });
+ return Response.json({ ok: true, posts: res.posts, token: res.token, editNotes }, { headers: noStore });
} catch (e) {
if (e instanceof StaleToken) {
return Response.json(
diff --git a/umtool/app/api/report/takes/preview/route.ts b/umtool/app/api/report/takes/preview/route.ts
@@ -0,0 +1,37 @@
+import { projectRef } from "@/lib/projects";
+import { rangeResponse } from "@/lib/report/serve.mjs";
+import { isTakeId, listTakes, previewFile } from "@/lib/report/takes.mjs";
+
+export const dynamic = "force-dynamic";
+
+// One take's preview mp4, to a <video> element. ?project&take&v.
+//
+// Names, never paths: the take must be one the listing returns -- so a take
+// whose take.json was skipped is not served either -- and its file is the one
+// its take.json names, confined to its own directory by previewFile.
+//
+// `v` is the video route's contract: the mtime the listing reported, and only
+// a matching one may be cached, because a re-render writes the same path.
+export async function GET(request: Request) {
+ const url = new URL(request.url);
+ const takeId = url.searchParams.get("take") ?? "";
+ if (!isTakeId(takeId)) return new Response("take must be a take id", { status: 400 });
+ const project = await projectRef(url.searchParams.get("project") ?? "");
+ if (!project) return new Response("no such project", { status: 404 });
+ const take = (await listTakes(project.dir)).takes.find((t) => t.id === takeId);
+ if (!take) return new Response("no such take", { status: 404 });
+ const file = await previewFile(project.dir, take);
+ if (!file) return new Response("this take has no preview yet", { status: 404 });
+
+ const fresh = url.searchParams.get("v") === String(file.mtimeMs);
+ return rangeResponse(request, {
+ abs: file.abs,
+ size: file.size,
+ headers: {
+ "content-type": "video/mp4",
+ "accept-ranges": "bytes",
+ "cache-control": fresh ? "private, max-age=3600, immutable" : "private, no-store",
+ "x-video-mtime": String(file.mtimeMs),
+ },
+ });
+}
diff --git a/umtool/app/api/report/takes/route.ts b/umtool/app/api/report/takes/route.ts
@@ -0,0 +1,15 @@
+import { projectRef } from "@/lib/projects";
+import { listTakes, readVerdicts } from "@/lib/report/takes.mjs";
+
+export const dynamic = "force-dynamic";
+
+// GET ?project=<id> every take, grouped, what was skipped and why, and the
+// verdicts so far. `exists: false` is a project with no
+// takes/ at all, which is the normal state, not an error.
+export async function GET(request: Request) {
+ const url = new URL(request.url);
+ const project = await projectRef(url.searchParams.get("project") ?? "");
+ if (!project) return Response.json({ error: "no such project" }, { status: 404 });
+ const [takes, verdicts] = await Promise.all([listTakes(project.dir), readVerdicts(project.dir)]);
+ return Response.json({ project: project.id, ...takes, verdicts }, { headers: { "cache-control": "no-store" } });
+}
diff --git a/umtool/app/api/report/takes/verdict/route.ts b/umtool/app/api/report/takes/verdict/route.ts
@@ -0,0 +1,39 @@
+import { projectRef } from "@/lib/projects";
+import { isTakeId, listTakes, readVerdicts, setTakeVerdict } from "@/lib/report/takes.mjs";
+
+export const dynamic = "force-dynamic";
+
+// GET ?project=<id> <project>/takes/verdicts.json, as read
+// POST { project, take, verdict?, note? } set one take's verdict ("like" |
+// "maybe" | "no" | null) and/or note;
+// an absent field keeps its value
+//
+// The take must be one the listing returns. A verdict on a take that is not
+// there would be written into a file another agent reads as a decision.
+export async function GET(request: Request) {
+ const url = new URL(request.url);
+ const project = await projectRef(url.searchParams.get("project") ?? "");
+ if (!project) return Response.json({ error: "no such project" }, { status: 404 });
+ return Response.json({ verdicts: await readVerdicts(project.dir) }, { headers: { "cache-control": "no-store" } });
+}
+
+export async function POST(request: Request) {
+ const body = (await request.json().catch(() => null)) as {
+ project?: string;
+ take?: string;
+ verdict?: string | null;
+ note?: string;
+ } | null;
+ if (!body || typeof body !== "object") return Response.json({ error: "bad json" }, { status: 400 });
+ if (!isTakeId(body.take)) return Response.json({ error: "take must be a take id" }, { status: 400 });
+ const project = await projectRef(String(body.project ?? ""));
+ if (!project) return Response.json({ error: "no such project" }, { status: 404 });
+ const { takes } = await listTakes(project.dir);
+ if (!takes.some((t) => t.id === body.take)) return Response.json({ error: "no such take" }, { status: 404 });
+ try {
+ const r = await setTakeVerdict(project.dir, body.take, { verdict: body.verdict, note: body.note });
+ return Response.json({ ok: true, ...r }, { headers: { "cache-control": "no-store" } });
+ } catch (e) {
+ return Response.json({ error: e instanceof Error ? e.message : String(e) }, { status: 400 });
+ }
+}
diff --git a/umtool/app/api/report/timeline/route.ts b/umtool/app/api/report/timeline/route.ts
@@ -0,0 +1,134 @@
+import { invalidateProjects } from "@/lib/projects";
+import { withEditNotes } from "@/lib/report/guard";
+import {
+ StaleToken,
+ StructureRefused,
+ duplicateEntry,
+ insertEntry,
+ manifestToken,
+ moveEntry,
+ removeEntry,
+ removePost,
+ undoStructural,
+ updateFactcheck,
+ updateTeaser,
+ upsertPost,
+} from "@/lib/report/manifest.mjs";
+import { resolveReport } from "@/lib/report/serve.mjs";
+import { deckOn } from "umtool-report-to-video/deck";
+import { VERDICTS, resolveFactcheck } from "umtool-report-to-video/factcheck";
+
+export const dynamic = "force-dynamic";
+
+// The STRUCTURE of the cut: what is in the timeline and in what order, a
+// teaser's lines, the posts, the fact-check's labels (lib/report/manifest.mjs,
+// "STRUCTURE"). The window route deliberately cannot move an entry; this is
+// the different button it points at.
+//
+// GET ?project=<id> the token, the timeline's ids in order, the
+// teasers, the posts and the fact-check block
+// POST { project, token, op, ... } one op:
+// move { id, at?, toIndex } toIndex is the index AFTER the move
+// remove { id, at? }
+// duplicate { id, at? }
+// insert { afterId: id | null, at?, entry: "<channel>/<video>@<start>-<end>" | { type, … } }
+// teaser { id, at?, patch: { lines?, beat?, dip?, tail?, tailWait? } }
+// post { post: { id, platform, date, text, url, … } } add or replace
+// post-remove { id }
+// factcheck { factcheck: { verdicts?, stamp?, tally? } | null }
+// undo {} restore the newest auto snapshot
+//
+// Every op snapshots the manifest first (`revisions/<stamp>-auto-before-<op>`),
+// is checked by the build's own validators, and on a GENERATED manifest leaves
+// an `edit` note for the agent (lib/report/guard.ts). A stale token is a 409.
+
+type Body = Record<string, unknown>;
+const noStore = { "cache-control": "no-store" };
+
+export async function GET(request: Request) {
+ const url = new URL(request.url);
+ const r = await resolveReport(url.searchParams.get("project") ?? "");
+ if ("error" in r) return Response.json({ error: r.error }, { status: r.status });
+ const entries = (r.manifest.timeline ?? []) as Record<string, unknown>[];
+ const timeline = entries.map((e) => ({ id: String(e.id), type: String(e.type ?? "entry") }));
+ // What the structure editors edit: each teaser as written, the posts, and
+ // the fact-check block with every default filled (the form's placeholders).
+ const teasers = entries
+ .map((e, at) => ({ e, at }))
+ .filter(({ e }) => e.type === "teaser")
+ .map(({ e, at }) => ({ at, id: String(e.id), lines: e.lines ?? [], beat: e.beat ?? null, dip: e.dip ?? null, tail: e.tail ?? null, tailWait: e.tailWait ?? null }));
+ const render = (r.manifest.render ?? {}) as Record<string, unknown>;
+ return Response.json(
+ {
+ timeline,
+ teasers,
+ posts: Array.isArray(r.manifest.posts) ? r.manifest.posts : [],
+ deckOn: deckOn(render),
+ factcheck: (render.chrome as { factcheck?: unknown } | undefined)?.factcheck ?? null,
+ factcheckResolved: resolveFactcheck(render),
+ verdicts: VERDICTS,
+ generatedBy: r.manifest.generatedBy ?? null,
+ token: await manifestToken(r.project.dir),
+ },
+ { headers: noStore },
+ );
+}
+
+const atOf = (v: unknown) => (v === undefined || v === null || v === "" ? null : Number(v));
+
+export async function POST(request: Request) {
+ let body: Body;
+ try {
+ body = await request.json();
+ } catch {
+ return Response.json({ error: "expected JSON" }, { status: 400 });
+ }
+ const r = await resolveReport(String(body.project ?? ""));
+ if ("error" in r) return Response.json({ error: r.error }, { status: r.status });
+ const dir = r.project.dir;
+ // A string, or no guard at all: `String(null)` would be the token "null",
+ // which matches no file and refuses every write.
+ const token = typeof body.token === "string" ? body.token : null;
+ const id = String(body.id ?? "");
+ const at = atOf(body.at);
+
+ const run = (): Promise<Record<string, unknown>> => {
+ switch (body.op) {
+ case "move":
+ return moveEntry(dir, id, Number(body.toIndex), { token, at });
+ case "remove":
+ return removeEntry(dir, id, { token, at });
+ case "duplicate":
+ return duplicateEntry(dir, id, { token, at });
+ case "insert":
+ return insertEntry(dir, body.afterId ? String(body.afterId) : null, body.entry, { token, at });
+ case "teaser":
+ return updateTeaser(dir, id, body.patch as Record<string, unknown>, { token, at });
+ case "post":
+ return upsertPost(dir, body.post as Record<string, unknown>, { token });
+ case "post-remove":
+ return removePost(dir, id, { token });
+ case "factcheck":
+ return updateFactcheck(dir, (body.factcheck ?? null) as Record<string, unknown> | null, { token });
+ case "undo":
+ return undoStructural(dir, { token });
+ default:
+ throw new Error("op must be move, remove, duplicate, insert, teaser, post, post-remove, factcheck or undo");
+ }
+ };
+
+ try {
+ const { result, editNotes } = await withEditNotes(r.project, run);
+ // revisions/ changed: the project page lists them.
+ invalidateProjects();
+ return Response.json({ ok: true, ...result, editNotes }, { headers: noStore });
+ } catch (e) {
+ if (e instanceof StaleToken) {
+ return Response.json({ error: e.message, expected: e.expected, got: e.got, stale: true }, { status: 409 });
+ }
+ if (e instanceof StructureRefused) {
+ return Response.json({ error: e.message, errors: e.errors }, { status: 400 });
+ }
+ return Response.json({ error: e instanceof Error ? e.message : String(e) }, { status: 400 });
+ }
+}
diff --git a/umtool/app/api/report/window/route.ts b/umtool/app/api/report/window/route.ts
@@ -1,3 +1,4 @@
+import { withEditNotes } from "@/lib/report/guard";
import { StaleToken, updateClip } from "@/lib/report/manifest.mjs";
import { resolveClip } from "@/lib/report/serve.mjs";
@@ -63,11 +64,13 @@ export async function PUT(request: Request) {
if (!Object.keys(patch).length) return Response.json({ error: "nothing to change" }, { status: 400 });
try {
- const res = await updateClip(r.project.dir, clipId, patch, {
- token: body.token === undefined ? null : String(body.token),
- });
+ const { result: res, editNotes } = await withEditNotes(r.project, () =>
+ updateClip(r.project.dir, clipId, patch, {
+ token: body.token === undefined ? null : String(body.token),
+ }),
+ );
return Response.json(
- { ok: true, entry: res.entry, before: res.before, token: res.token },
+ { ok: true, entry: res.entry, before: res.before, token: res.token, editNotes },
{ headers: { "cache-control": "no-store" } },
);
} catch (e) {
diff --git a/umtool/app/api/sites/evidence/route.ts b/umtool/app/api/sites/evidence/route.ts
@@ -0,0 +1,19 @@
+import { citationEvidence } from "@/lib/articles/evidence";
+import { readReportFile, siteById } from "@/lib/articles/sites";
+
+export const dynamic = "force-dynamic";
+
+// GET ?site&report&cite -- one citation's evidence (lib/articles/evidence.ts):
+// its quote and record, the transcript around it, and what can play it.
+export async function GET(request: Request) {
+ const url = new URL(request.url);
+ const site = url.searchParams.get("site") ?? "";
+ const reportId = url.searchParams.get("report") ?? "";
+ const cite = url.searchParams.get("cite") ?? "";
+ if (!siteById(site)) return Response.json({ error: `no site ${site}` }, { status: 404 });
+ const read = await readReportFile(site, reportId);
+ if (!read.report) return Response.json({ error: read.problems[0]?.message ?? "no report" }, { status: 404 });
+ const ev = await citationEvidence(site, read.report, cite);
+ if (!ev) return Response.json({ error: `no citation ${cite}` }, { status: 404 });
+ return Response.json(ev, { headers: { "cache-control": "no-store" } });
+}
diff --git a/umtool/app/api/sites/media/route.ts b/umtool/app/api/sites/media/route.ts
@@ -0,0 +1,94 @@
+import { readFile, realpath, stat } from "node:fs/promises";
+import path from "node:path";
+import { reportMediaDir, reportMediaIndexFile, siteReportDir } from "yt-dlp-transcript-common/publish/reportMedia";
+import { rangeResponse } from "@/lib/report/serve.mjs";
+import { CHANNELS_DIR, REPORTS_ROOT, SITES_DIR, inside } from "@/lib/paths";
+import { sitesPaths } from "@/lib/articles/sites";
+
+export const dynamic = "force-dynamic";
+
+// Article media, READ-ONLY, with byte ranges (lib/report/serve.mjs
+// rangeResponse, the one the cut player uses -- a <video> will not seek a
+// stream it was not given a 206 for).
+//
+// ?site&report&file a file in the report's directory: video.mp4,
+// poster.jpg, stills/… -- relative, no `..`, a media type,
+// and its real path under SITES_DIR (or REPORTS_ROOT: a
+// generator may link a take's preview in)
+// ?site&moment the site's PREPARED evidence clip for a moment, found
+// through report-media/index.json -- never a client path
+// ?corpus=<abs> a clip window, saved video, audio or post capture in
+// the corpus: lexically under CHANNELS_DIR (its real
+// path is on whatever drive the media tier links to)
+const TYPES: Record<string, string> = {
+ ".mp4": "video/mp4",
+ ".m4v": "video/mp4",
+ ".webm": "video/webm",
+ ".mkv": "video/x-matroska",
+ ".mov": "video/quicktime",
+ ".m4a": "audio/mp4",
+ ".mp3": "audio/mpeg",
+ ".opus": "audio/ogg",
+ ".ogg": "audio/ogg",
+ ".wav": "audio/wav",
+ ".jpg": "image/jpeg",
+ ".jpeg": "image/jpeg",
+ ".png": "image/png",
+ ".webp": "image/webp",
+ ".gif": "image/gif",
+};
+const SEGMENT = /^[a-z0-9][a-z0-9-]{0,63}$/;
+
+async function serve(request: Request, abs: string): Promise<Response> {
+ const type = TYPES[path.extname(abs).toLowerCase()];
+ if (!type) return new Response("not a media file", { status: 400 });
+ const st = await stat(/* turbopackIgnore: true */ abs).catch(() => null);
+ if (!st?.isFile()) return new Response("not found", { status: 404 });
+ return rangeResponse(request, {
+ abs,
+ size: st.size,
+ headers: { "content-type": type, "accept-ranges": "bytes", "cache-control": "private, no-store" },
+ });
+}
+
+async function reportFile(site: string, report: string, rel: string): Promise<string | null> {
+ if (!SEGMENT.test(site) || !SEGMENT.test(report)) return null;
+ if (!rel || rel.includes("\0") || rel.startsWith("/") || rel.split("/").some((s) => s === ".." || s === "" || s === ".")) return null;
+ const dir = siteReportDir(sitesPaths(), site, report);
+ const abs = path.join(/* turbopackIgnore: true */ dir, rel);
+ if (!inside(dir, abs)) return null;
+ const real = await realpath(/* turbopackIgnore: true */ abs).catch(() => null);
+ if (!real) return null;
+ const roots = await Promise.all([SITES_DIR, REPORTS_ROOT].map((r) => realpath(/* turbopackIgnore: true */ r).catch(() => r)));
+ return roots.some((r) => inside(r, real)) ? real : null;
+}
+
+async function preparedFile(site: string, moment: string): Promise<string | null> {
+ if (!SEGMENT.test(site)) return null;
+ try {
+ const index = JSON.parse(await readFile(/* turbopackIgnore: true */ reportMediaIndexFile(sitesPaths(), site), "utf8"));
+ const entry = index?.moments?.[moment];
+ if (!entry || typeof entry.file !== "string") return null;
+ const dir = reportMediaDir(sitesPaths(), site);
+ const abs = path.resolve(/* turbopackIgnore: true */ dir, entry.file);
+ return inside(dir, abs) ? abs : null;
+ } catch {
+ return null;
+ }
+}
+
+export async function GET(request: Request) {
+ const url = new URL(request.url);
+ const q = (k: string) => url.searchParams.get(k) ?? "";
+ if (url.searchParams.has("corpus")) {
+ const abs = path.resolve(/* turbopackIgnore: true */ q("corpus"));
+ if (abs !== q("corpus") || !inside(CHANNELS_DIR, abs)) return new Response("outside the corpus", { status: 400 });
+ return serve(request, abs);
+ }
+ if (url.searchParams.has("moment")) {
+ const abs = await preparedFile(q("site"), q("moment"));
+ return abs ? serve(request, abs) : new Response("no prepared clip for that moment", { status: 404 });
+ }
+ const abs = await reportFile(q("site"), q("report"), q("file"));
+ return abs ? serve(request, abs) : new Response("no such report file", { status: 404 });
+}
diff --git a/umtool/app/api/sites/workspace/route.ts b/umtool/app/api/sites/workspace/route.ts
@@ -0,0 +1,23 @@
+import { readFile } from "node:fs/promises";
+import { workspaceFile } from "@/lib/articles/workspace.mjs";
+
+export const dynamic = "force-dynamic";
+
+// GET ?ws=<name under REPORTS_ROOT>&rel=<listed file> -- one workspace file,
+// READ-ONLY, only one that lib/articles/workspace.mjs lists. HTML is served
+// under a CSP sandbox (no scripts, no same-origin) for the page's sandboxed
+// iframe; markdown and JSON as plain text.
+export async function GET(request: Request) {
+ const url = new URL(request.url);
+ const f = await workspaceFile(url.searchParams.get("ws") ?? "", url.searchParams.get("rel") ?? "");
+ if (!f) return new Response("not a workspace file", { status: 404 });
+ const body = await readFile(/* turbopackIgnore: true */ f.real);
+ const html = f.rel.endsWith(".html");
+ return new Response(body, {
+ headers: {
+ "content-type": html ? "text/html; charset=utf-8" : f.rel.endsWith(".json") ? "application/json; charset=utf-8" : "text/plain; charset=utf-8",
+ "cache-control": "no-store",
+ ...(html ? { "content-security-policy": "sandbox; default-src 'none'; img-src data:; style-src 'unsafe-inline'" } : {}),
+ },
+ });
+}
diff --git a/umtool/app/browse/decisions/page.tsx b/umtool/app/browse/decisions/page.tsx
@@ -122,7 +122,9 @@ export default async function DecisionsPage({
<section key={id} data-project={id}>
<h2 className="mb-1.5 flex items-baseline gap-2">
<Link
- href={`/browse/${id}`}
+ // An article's notes are listed under its page's path
+ // (lib/decisions.ts noteDecisions), not a project's.
+ href={id.startsWith("sites/") ? `/${id}` : `/browse/${id}`}
className="font-mono text-[13px] text-[var(--color-text)] hover:text-[var(--color-sel)]"
>
{id}
diff --git a/umtool/app/sites/[site]/[report]/evidence/page.tsx b/umtool/app/sites/[site]/[report]/evidence/page.tsx
@@ -0,0 +1,60 @@
+import { notFound } from "next/navigation";
+import { reportCitationNumbers } from "yt-dlp-transcript-common/lib/report/uses";
+import BrowseHeader from "@/components/BrowseHeader";
+import EvidenceWalk, { type WalkItem } from "@/components/articles/EvidenceWalk";
+import { readArticleNotes, readReportFile, siteById } from "@/lib/articles/sites";
+import { corpusNotesFile } from "@/lib/paths";
+import { linkedProjects } from "@/lib/articles/links.mjs";
+import { sourceFor } from "@/lib/articles/sources.mjs";
+import type { NotesRead } from "@/lib/annotations/types";
+
+export const dynamic = "force-dynamic";
+
+// Every citation of one article, one per screen, in the article's own
+// numbering. `?c=<citationId>` opens on that citation.
+export default async function EvidenceWalkPage({
+ params,
+ searchParams,
+}: {
+ params: Promise<{ site: string; report: string }>;
+ searchParams: Promise<{ c?: string }>;
+}) {
+ const { site: siteId, report: reportId } = await params;
+ const { c } = await searchParams;
+ const site = siteById(siteId);
+ if (!site) notFound();
+ const read = await readReportFile(site.siteId, reportId);
+ if (!read.report) notFound();
+ const r = read.report;
+ const numbers = reportCitationNumbers(r);
+ const items: WalkItem[] = [...numbers.entries()]
+ .sort((a, b) => a[1] - b[1])
+ .map(([id, number]) => ({ id, number, kind: r.citations?.[id]?.kind ?? "?", quote: r.citations?.[id]?.quote ?? "" }));
+ const at = Math.max(0, items.findIndex((i) => i.id === c));
+ const [notes, source] = await Promise.all([readArticleNotes(site.siteId, reportId), sourceFor(site.siteId, reportId)]);
+ const links = await linkedProjects(site.siteId, reportId, { workspace: source?.workspace ?? null });
+ const initialNotes: NotesRead = {
+ subject: { kind: "article", site: site.siteId, report: reportId },
+ file: corpusNotesFile(site.siteId, reportId) ?? "",
+ token: notes.token,
+ doc: notes.doc,
+ source: notes.doc?.source ?? null,
+ };
+ return (
+ <div className="flex h-full flex-col">
+ <BrowseHeader
+ active="sites"
+ crumbs={[
+ { href: "/sites", label: "sites" },
+ { href: `/sites/${site.siteId}`, label: site.siteId },
+ { href: `/sites/${site.siteId}/${reportId}`, label: reportId },
+ { label: "evidence" },
+ ]}
+ note={`${items.length} citations`}
+ />
+ <EvidenceWalk site={site.siteId} report={reportId} items={items} initial={at} initialNotes={initialNotes}
+ videoProjects={links.linked.map((p: { id: string; title: string; generatedBy: string | null }) => ({ id: p.id, title: p.title, generatedBy: p.generatedBy }))}
+ />
+ </div>
+ );
+}
diff --git a/umtool/app/sites/[site]/[report]/page.tsx b/umtool/app/sites/[site]/[report]/page.tsx
@@ -0,0 +1,180 @@
+import path from "node:path";
+import Link from "next/link";
+import { notFound } from "next/navigation";
+import { readRevisionHead, REPORT_HISTORY_GIT_DIRNAME } from "yt-dlp-transcript-common/publish/reportHistory";
+import { siteReportDir } from "yt-dlp-transcript-common/publish/reportMedia";
+import BrowseHeader from "@/components/BrowseHeader";
+import ArticleReader from "@/components/articles/ArticleReader";
+import WorkspacePanel from "@/components/articles/WorkspacePanel";
+import { badgeVariants } from "@/components/ui/badge";
+import { articleView } from "@/lib/articles/article";
+import { articleWorkspaceListing, openWorkspaceFile } from "@/lib/articles/files";
+import { linkedProjects } from "@/lib/articles/links.mjs";
+import { readArticleNotes, readReportFile, siteById, sitesPaths } from "@/lib/articles/sites";
+import { sourceFor } from "@/lib/articles/sources.mjs";
+import { corpusNotesFile } from "@/lib/paths";
+import type { NotesRead } from "@/lib/annotations/types";
+
+export const dynamic = "force-dynamic";
+
+// One article: the reader with its notes (default), or `?tab=source` -- the
+// workspace files it was written from, its draft marked.
+//
+// ?status=open|resolved|all the notes filter it opens with
+// ?note=<id> the note it opens on (the decisions inbox links here)
+// ?ws=&rel= a workspace file open on the source tab
+
+type Search = { tab?: string; status?: string; note?: string; ws?: string; rel?: string };
+
+export default async function ArticlePage({
+ params,
+ searchParams,
+}: {
+ params: Promise<{ site: string; report: string }>;
+ searchParams: Promise<Search>;
+}) {
+ const { site: siteId, report: reportId } = await params;
+ const sp = await searchParams;
+ const site = siteById(siteId);
+ if (!site) notFound();
+ const read = await readReportFile(site.siteId, reportId);
+ if (!read.report && read.problems[0]?.message === "no report.json") notFound();
+
+ const [notes, source] = await Promise.all([readArticleNotes(site.siteId, reportId), sourceFor(site.siteId, reportId)]);
+ const links = await linkedProjects(site.siteId, reportId, { workspace: source?.workspace ?? null });
+ const gitDir = path.join(/* turbopackIgnore: true */ siteReportDir(sitesPaths(), site.siteId, reportId), REPORT_HISTORY_GIT_DIRNAME);
+ const head = await readRevisionHead(gitDir).catch(() => null);
+ const published = (site.reports ?? []).includes(reportId);
+ const r = read.report;
+ const tab = sp.tab === "source" ? "source" : "article";
+ const base = `/sites/${site.siteId}/${reportId}`;
+ const media = (file: string) =>
+ `/api/sites/media?site=${encodeURIComponent(site.siteId)}&report=${encodeURIComponent(reportId)}&file=${encodeURIComponent(file)}`;
+
+ const meta = (
+ <div data-article-meta className="flex flex-wrap items-center gap-x-3 gap-y-1 text-[12px] text-[var(--color-dim)]">
+ <Link href={`/sites/${site.siteId}`} className="hover:text-[var(--color-text)]">
+ {site.siteTitle || site.siteId}
+ </Link>
+ <span data-status={published ? "published" : "draft"} className={badgeVariants({ variant: published ? "on" : "neutral", size: "sm" })}>
+ {published ? "published" : "draft"}
+ </span>
+ {r?.published && <span>published {r.published}</span>}
+ {r?.updated && <span>updated {r.updated}</span>}
+ {head && <span data-revisions={head.revision}>{head.revision} revision{head.revision === 1 ? "" : "s"}</span>}
+ {links.linked.map((p: { id: string }) => (
+ <Link key={p.id} href={`/browse/${p.id}`} data-project-link={p.id} className="font-mono text-[var(--color-sel)] hover:underline">
+ {p.id}
+ </Link>
+ ))}
+ {links.possible.map((p: { id: string }) => (
+ <Link key={p.id} href={`/browse/${p.id}`} className="font-mono hover:underline" title="shares the slug; not linked">
+ {p.id} (possible)
+ </Link>
+ ))}
+ <nav className="ml-auto flex gap-2" aria-label="article views">
+ <Link href={base} aria-current={tab === "article" ? "page" : undefined} className={tab === "article" ? "text-[var(--color-sel)]" : "hover:text-[var(--color-text)]"}>
+ article
+ </Link>
+ <Link href={`${base}?tab=source`} aria-current={tab === "source" ? "page" : undefined} className={tab === "source" ? "text-[var(--color-sel)]" : "hover:text-[var(--color-text)]"}>
+ source
+ </Link>
+ <Link href={`${base}/evidence`} className="hover:text-[var(--color-text)]">
+ evidence
+ </Link>
+ </nav>
+ {source && (
+ <div className="w-full font-mono text-[11px]" title={source.how}>
+ {source.draft && <span>draft {source.draft}</span>}
+ {source.generator && <span className="ml-3">generator {source.generator}</span>}
+ </div>
+ )}
+ </div>
+ );
+
+ const header = (
+ <BrowseHeader
+ active="sites"
+ crumbs={[{ href: "/sites", label: "sites" }, { href: `/sites/${site.siteId}`, label: site.siteId }, { label: reportId }]}
+ note={`${notes.doc?.notes.filter((n) => n.status === "open").length ?? 0} open notes`}
+ />
+ );
+
+ if (!r) {
+ return (
+ <div className="flex h-full flex-col">
+ {header}
+ <main className="deck-main flex-1 p-4">
+ <div className="mx-auto max-w-[72ch] space-y-3">
+ {meta}
+ <p className="text-[13px] text-[var(--color-bad)]">report.json does not read as a report:</p>
+ <ul className="list-disc pl-5 text-[12px] text-[var(--color-dim)]">
+ {read.problems.slice(0, 20).map((p, i) => (
+ <li key={i}>{p.message}</li>
+ ))}
+ </ul>
+ </div>
+ </main>
+ </div>
+ );
+ }
+
+ if (tab === "source") {
+ const ws = await articleWorkspaceListing(source?.workspace);
+ const opened =
+ sp.ws && sp.rel
+ ? await openWorkspaceFile(sp.ws, sp.rel)
+ : ws && source?.draft
+ ? await openWorkspaceFile(ws.name, path.relative(ws.dir, source.draft))
+ : null;
+ const draftRel = ws && source?.draft ? `${ws.name}/${path.relative(ws.dir, source.draft)}` : null;
+ return (
+ <div className="flex h-full flex-col">
+ {header}
+ <main className="deck-main flex-1 p-4">
+ <div className="space-y-4">
+ {meta}
+ <WorkspacePanel
+ workspaces={ws ? [ws] : []}
+ opened={opened}
+ highlight={draftRel}
+ hrefFor={(w, rel) => `${base}?${new URLSearchParams({ tab: "source", ws: w, rel })}`}
+ />
+ </div>
+ </main>
+ </div>
+ );
+ }
+
+ const { view, error } = await articleView(r);
+ const initialNotes: NotesRead = {
+ subject: { kind: "article", site: site.siteId, report: reportId },
+ file: corpusNotesFile(site.siteId, reportId) ?? "",
+ token: notes.token,
+ doc: notes.doc,
+ source: notes.doc?.source ?? source ?? null,
+ ...(notes.error ? { error: notes.error } : {}),
+ };
+ const status = sp.status === "resolved" || sp.status === "all" ? sp.status : "open";
+ const video = r.video
+ ? { file: r.video.src, src: media(r.video.src), poster: r.video.poster ? media(r.video.poster) : undefined, caption: r.video.caption }
+ : null;
+
+ return (
+ <div className="flex h-full flex-col">
+ {header}
+ <ArticleReader
+ site={site.siteId}
+ report={reportId}
+ view={view}
+ viewError={error}
+ initialNotes={initialNotes}
+ initialFilter={status}
+ initialNote={sp.note ?? null}
+ meta={meta}
+ video={video}
+ videoProjects={links.linked.map((p: { id: string; title: string; generatedBy: string | null }) => ({ id: p.id, title: p.title, generatedBy: p.generatedBy }))}
+ />
+ </div>
+ );
+}
diff --git a/umtool/app/sites/[site]/page.tsx b/umtool/app/sites/[site]/page.tsx
@@ -0,0 +1,123 @@
+import Link from "next/link";
+import { notFound } from "next/navigation";
+import BrowseHeader from "@/components/BrowseHeader";
+import ArticleTable, { articleHref } from "@/components/articles/ArticleTable";
+import SiteChips from "@/components/articles/SiteChips";
+import WorkspacePanel from "@/components/articles/WorkspacePanel";
+import { badgeVariants } from "@/components/ui/badge";
+import { readSiteRow, siteById } from "@/lib/articles/sites";
+import { openWorkspaceFile, siteWorkspaceListings, takeTally } from "@/lib/articles/files";
+import { videoProjects } from "@/lib/articles/links.mjs";
+
+export const dynamic = "force-dynamic";
+
+// One site: its articles, the videos they play, the umtool projects those
+// videos were cut in (with their takes), and the workspace files the articles
+// were written from. `?ws=&rel=` opens a workspace file below.
+
+export default async function SitePage({
+ params,
+ searchParams,
+}: {
+ params: Promise<{ site: string }>;
+ searchParams: Promise<{ ws?: string; rel?: string }>;
+}) {
+ const { site: siteId } = await params;
+ const sp = await searchParams;
+ const site = siteById(siteId);
+ if (!site) notFound();
+ const row = await readSiteRow(site);
+
+ const projects = await videoProjects();
+ const linkedIds = [...new Set(row.articles.flatMap((a) => a.projects.linked.map((p) => p.id)))];
+ const linked = await Promise.all(
+ linkedIds.map(async (id) => {
+ const p = projects.find((x: { id: string }) => x.id === id) as { id: string; dir: string };
+ return { id: p.id, tally: await takeTally(p.dir), articles: row.articles.filter((a) => a.projects.linked.some((l) => l.id === id)) };
+ }),
+ );
+ const videos = row.articles.filter((a) => a.hasVideo);
+ const workspaces = await siteWorkspaceListings(site.siteId, row.articles.map((a) => a.id));
+ const opened = sp.ws && sp.rel ? await openWorkspaceFile(sp.ws, sp.rel) : null;
+ const hrefFor = (ws: string, rel: string) =>
+ `/sites/${site.siteId}?${new URLSearchParams({ ws, rel }).toString()}#files`;
+ const media = (id: string, file: string) =>
+ `/api/sites/media?site=${encodeURIComponent(site.siteId)}&report=${encodeURIComponent(id)}&file=${file}`;
+
+ return (
+ <div className="flex h-full flex-col">
+ <BrowseHeader
+ active="sites"
+ crumbs={[{ href: "/sites", label: "sites" }, { label: row.title }]}
+ note={`${row.published} published · ${row.drafts} drafts · ${row.openNotes} open notes`}
+ />
+ <main className="deck-main flex-1 space-y-6 p-4">
+ <div className="flex flex-wrap items-baseline gap-2">
+ <h1 className="text-[16px] font-semibold text-[var(--color-text)]">{row.title}</h1>
+ <span className="font-mono text-[11px] text-[var(--color-dim)]">{row.siteId}</span>
+ <SiteChips site={row} />
+ </div>
+
+ <section data-section="articles">
+ <h2 className="micro mb-1.5">articles</h2>
+ <ArticleTable articles={row.articles} />
+ </section>
+
+ <section data-section="videos">
+ <h2 className="micro mb-1.5">report videos {videos.length}</h2>
+ {videos.length === 0 ? (
+ <p className="text-[12px] text-[var(--color-dim)]">none</p>
+ ) : (
+ <div className="grid grid-cols-[repeat(auto-fill,minmax(280px,1fr))] gap-3">
+ {videos.map((a) => (
+ <figure key={a.id} data-report-video={a.id} className="rounded border border-[var(--color-line)] bg-[var(--color-panel)] p-2">
+ <video
+ controls
+ preload="none"
+ src={media(a.id, "video.mp4")}
+ poster={a.hasPoster ? media(a.id, "poster.jpg") : undefined}
+ className="aspect-video w-full rounded bg-black"
+ />
+ <figcaption className="mt-1 text-[12px]">
+ <Link href={articleHref(a)} className="text-[var(--color-text)] hover:text-[var(--color-sel)]">
+ {a.title}
+ </Link>
+ </figcaption>
+ </figure>
+ ))}
+ </div>
+ )}
+ </section>
+
+ <section data-section="projects">
+ <h2 className="micro mb-1.5">video projects {linked.length}</h2>
+ {linked.length === 0 ? (
+ <p className="text-[12px] text-[var(--color-dim)]">none linked</p>
+ ) : (
+ <ul className="space-y-1">
+ {linked.map((p) => (
+ <li key={p.id} data-video-project={p.id} className="flex flex-wrap items-baseline gap-2 rounded border border-[var(--color-line)] bg-[var(--color-panel)] px-3 py-1.5 text-[12px]">
+ <Link href={`/browse/${p.id}`} className="font-mono text-[var(--color-sel)] hover:underline">
+ {p.id}
+ </Link>
+ <span className="text-[var(--color-dim)]">{p.articles.map((a) => a.title).join(", ")}</span>
+ <Link href={`/browse/${p.id}/takes`} className="ml-auto text-[var(--color-dim)] hover:text-[var(--color-text)]" data-takes={p.tally.takes}>
+ {p.tally.takes} takes
+ </Link>
+ <span className={badgeVariants({ variant: "neutral", size: "sm" })}>like {p.tally.like}</span>
+ <span className={badgeVariants({ variant: "neutral", size: "sm" })}>maybe {p.tally.maybe}</span>
+ <span className={badgeVariants({ variant: "neutral", size: "sm" })}>no {p.tally.no}</span>
+ </li>
+ ))}
+ </ul>
+ )}
+ </section>
+
+ <section data-section="files" id="files">
+ <h2 className="micro mb-1.5">workspace files</h2>
+ <WorkspacePanel workspaces={workspaces} opened={opened} hrefFor={hrefFor} />
+ </section>
+ </main>
+ </div>
+ );
+}
diff --git a/umtool/app/sites/page.tsx b/umtool/app/sites/page.tsx
@@ -0,0 +1,87 @@
+import Link from "next/link";
+import BrowseHeader from "@/components/BrowseHeader";
+import ArticleTable from "@/components/articles/ArticleTable";
+import SiteChips from "@/components/articles/SiteChips";
+import { badgeVariants } from "@/components/ui/badge";
+import { listSiteRows, type ArticleRow } from "@/lib/articles/sites";
+
+export const dynamic = "force-dynamic";
+
+// Every site's articles, private sites first. Zero client JS: the filters are
+// links that change searchParams (/browse/decisions' idiom), so a filtered
+// view is one pasteable URL.
+
+type Search = { site?: string; status?: string; notes?: string };
+
+export default async function SitesPage({ searchParams }: { searchParams: Promise<Search> }) {
+ const sp = await searchParams;
+ const all = await listSiteRows();
+
+ const keep = (a: ArticleRow) =>
+ (!sp.status || a.status === sp.status) && (sp.notes !== "open" || a.openNotes > 0);
+ const sites = all
+ .filter((s) => !sp.site || s.siteId === sp.site)
+ .map((s) => ({ ...s, shown: s.articles.filter(keep) }))
+ .filter((s) => s.shown.length > 0 || (!sp.status && !sp.notes));
+
+ const articles = all.flatMap((s) => s.articles);
+ const open = articles.reduce((n, a) => n + a.openNotes, 0);
+ const qs = (next: Partial<Search>) => {
+ const p = new URLSearchParams();
+ for (const [k, v] of Object.entries({ ...sp, ...next })) if (v) p.set(k, String(v));
+ const s = p.toString();
+ return `/sites${s ? `?${s}` : ""}`;
+ };
+
+ return (
+ <div className="flex h-full flex-col">
+ <BrowseHeader active="sites" crumbs={[{ label: "sites" }]} note={`${all.length} sites · ${articles.length} articles · ${open} open notes`} />
+ <main className="deck-main flex-1 p-4">
+ <div className="mb-3 flex flex-wrap items-center gap-1.5">
+ <span className="micro">site</span>
+ <Chip href={qs({ site: "" })} on={!sp.site} label={`all ${all.length}`} />
+ {all.map((s) => (
+ <Chip key={s.siteId} href={qs({ site: s.siteId })} on={sp.site === s.siteId} label={s.siteId} />
+ ))}
+ <span className="micro ml-3">status</span>
+ <Chip href={qs({ status: "" })} on={!sp.status} label="all" />
+ <Chip href={qs({ status: "published" })} on={sp.status === "published"} label={`published ${articles.filter((a) => a.status === "published").length}`} />
+ <Chip href={qs({ status: "draft" })} on={sp.status === "draft"} label={`draft ${articles.filter((a) => a.status === "draft").length}`} />
+ <span className="micro ml-3">notes</span>
+ <Chip href={qs({ notes: "" })} on={sp.notes !== "open"} label="all" />
+ <Chip href={qs({ notes: "open" })} on={sp.notes === "open"} label={`open ${articles.filter((a) => a.openNotes > 0).length}`} />
+ </div>
+
+ {sites.length === 0 ? (
+ <p className="text-[12px] text-[var(--color-dim)]">{all.length === 0 ? "no sites" : "nothing matches that filter"}</p>
+ ) : (
+ <div className="space-y-6">
+ {sites.map((s) => (
+ <section key={s.siteId} data-site={s.siteId}>
+ <h2 className="mb-1.5 flex flex-wrap items-baseline gap-2">
+ <Link href={`/sites/${s.siteId}`} className="text-[14px] font-semibold text-[var(--color-text)] hover:text-[var(--color-sel)]">
+ {s.title}
+ </Link>
+ <span className="font-mono text-[11px] text-[var(--color-dim)]">{s.siteId}</span>
+ <SiteChips site={s} />
+ <span className="micro" data-counts={`${s.published}/${s.drafts}`}>
+ {s.published} published · {s.drafts} draft{s.drafts === 1 ? "" : "s"}
+ </span>
+ </h2>
+ <ArticleTable articles={s.shown} />
+ </section>
+ ))}
+ </div>
+ )}
+ </main>
+ </div>
+ );
+}
+
+function Chip({ href, on, label }: { href: string; on: boolean; label: string }) {
+ return (
+ <Link href={href} aria-current={on ? "true" : undefined} className={badgeVariants({ variant: on ? "on" : "neutral" })}>
+ {label}
+ </Link>
+ );
+}
diff --git a/umtool/bin/umtool.mjs b/umtool/bin/umtool.mjs
@@ -36,6 +36,9 @@
// umtool diff <project> <snapshot> what changed since that snapshot
// umtool export <project> --format toc-bbcode|toc-markdown|description|chapters [--variant V]
// umtool check-sources [<project>…] prints the re-check chain
+// umtool notes [<site>/<report> | <project> | --all] [--open|--resolved|--all-status] [--json]
+// umtool notes reply <id> "<text>" [--resolve] | resolve | wontfix | reopen <id>
+// the operator's notes, and the agent's answers
import process from "node:process";
import {
PROJECT_KINDS,
@@ -65,6 +68,7 @@ import { buildSteps, checkSourcesSteps, PRESETS } from "../lib/report/driver.mjs
import { openIndex, signRecord } from "../lib/projects/index-db.mjs";
import { pipelineProcessesFor } from "../lib/report/busy.mjs";
import { probeTools } from "../lib/tools.mjs";
+import { notesCommand } from "../lib/annotations/cli.mjs";
import { scaffoldReportVideo } from "../lib/projects/scaffold.mjs";
import { CACHE_DIR, INDEX_DIR, MEDIA_ROOT, MEDIA_TIERED, OLD_CACHE_DIR } from "../lib/paths.mjs";
import {
@@ -698,6 +702,8 @@ function usage() {
" umtool diff <project> <snapshot> what changed since that snapshot",
" umtool export <project> --format toc-bbcode|toc-markdown|description|chapters [--variant V]",
" umtool check-sources [<project>…] prints the re-check chain (never-checked/old when no args)",
+ " umtool notes [<site>/<report>|<project>|--all] the operator's notes, as markdown (open ones)",
+ " umtool notes reply <id> \"<text>\" [--resolve] answer one; resolve | wontfix | reopen <id>",
"",
`reading ${REPORTS_ROOT} (set REPORTS_DIR to move it)`,
"",
@@ -725,6 +731,7 @@ const COMMANDS = {
diff: cmdDiff,
export: cmdExport,
"check-sources": cmdCheckSources,
+ notes: async () => process.exit(await notesCommand(argv.slice(argv.indexOf("notes") + 1))),
help: usage,
};
diff --git a/umtool/components/AppNav.tsx b/umtool/components/AppNav.tsx
@@ -6,7 +6,8 @@ import NavGroup from "./NavGroup";
// is waiting, the two benches that are not a project (mix, find), and the song
// piles folded under one entry.
//
-// SEVEN visible entries, and the cap is still NINE. A tenth wraps the header on
+// SEVEN visible entries (home, browse, decisions, sites, mix, find, song ▸),
+// and the cap is still NINE. A tenth wraps the header on
// a laptop, and a nav that wraps stops reading as one row of places and starts
// reading as a list. The next tool goes UNDER one of these, not beside them --
// which is exactly what happened to the four judging piles and the sources
@@ -30,6 +31,9 @@ export default function AppNav({ active }: { active: string }) {
// The worklist across every project, not a sixth pile. It sits beside
// browse because that is where every decision it names gets settled.
{ href: "/browse/decisions", label: "decisions" },
+ // Every site's articles -- published and drafts -- with their notes, their
+ // evidence and the workspace they were written in.
+ { href: "/sites", label: "sites" },
{ href: "/mix", label: "mix" },
// Every occurrence of a word across the corpus. It sits with browse because
// what it retrieves is raw material for a build, not a pile to judge.
diff --git a/umtool/components/articles/AddToVideo.tsx b/umtool/components/articles/AddToVideo.tsx
@@ -0,0 +1,78 @@
+"use client";
+
+import Link from "next/link";
+import { useState } from "react";
+import type { Evidence } from "@/lib/articles/evidence";
+
+// "Add to video": the cited span as a new clip at the END of a linked
+// report-video project's timeline, through the structure route
+// (/api/report/timeline, op insert). That route snapshots the manifest first
+// (Undo on the project page restores it) and, on a GENERATED manifest, leaves
+// an `edit` note so the agent ports the clip into the generator's inputs.
+
+export type VideoProjectLink = { id: string; title: string; generatedBy: string | null };
+
+export default function AddToVideo({ ev, projects }: { ev: Evidence; projects: VideoProjectLink[] }) {
+ const [project, setProject] = useState(projects[0]?.id ?? "");
+ const [state, setState] = useState<{ busy: boolean; done?: { id: string; project: string }; error?: string }>({ busy: false });
+ if (!projects.length || !ev.record || ev.start === undefined || ev.end === undefined) return null;
+ const chosen = projects.find((p) => p.id === project) ?? projects[0];
+
+ const add = async () => {
+ setState({ busy: true });
+ try {
+ const g = await fetch(`/api/report/timeline?project=${encodeURIComponent(chosen.id)}`, { cache: "no-store" });
+ const gj = await g.json();
+ if (!g.ok) throw new Error(gj.error ?? g.statusText);
+ const last = gj.timeline.length ? gj.timeline[gj.timeline.length - 1] : null;
+ const entry: Record<string, unknown> = {
+ type: "clip",
+ channel: ev.record!.channel,
+ video: ev.record!.id,
+ start: ev.start,
+ end: ev.end,
+ quote: ev.quote,
+ };
+ if (/^[A-Za-z0-9_-]{1,60}$/.test(ev.cite)) entry.id = `cite-${ev.cite}`;
+ const r = await fetch("/api/report/timeline", {
+ method: "POST",
+ headers: { "content-type": "application/json" },
+ body: JSON.stringify({ project: chosen.id, token: gj.token, op: "insert", afterId: last?.id ?? null, at: last ? gj.timeline.length - 1 : null, entry }),
+ });
+ const j = await r.json();
+ if (!r.ok) throw new Error(j.error ?? r.statusText);
+ setState({ busy: false, done: { id: j.id ?? j.result?.id ?? "the clip", project: chosen.id } });
+ } catch (err) {
+ setState({ busy: false, error: err instanceof Error ? err.message : String(err) });
+ }
+ };
+
+ return (
+ <div data-add-to-video className="flex flex-wrap items-center gap-2 border-t border-[var(--color-line)] pt-2">
+ {projects.length > 1 ? (
+ <select value={chosen.id} onChange={(e) => setProject(e.target.value)} aria-label="video project" className="rounded border border-[var(--color-line)] bg-transparent px-1 py-0.5 font-mono text-[11px]">
+ {projects.map((p) => (
+ <option key={p.id} value={p.id}>
+ {p.id}
+ </option>
+ ))}
+ </select>
+ ) : (
+ <span className="font-mono text-[11px] text-[var(--color-dim)]">{chosen.id}</span>
+ )}
+ <button type="button" onClick={add} disabled={state.busy} className="rounded border border-[var(--color-line)] px-2 py-0.5 hover:border-[var(--color-sel)] disabled:opacity-50">
+ Add to video
+ </button>
+ {chosen.generatedBy && <span className="text-[11px] text-[var(--color-dim)]">generated; leaves an edit note</span>}
+ {state.done && (
+ <span role="status" className="text-[11px]">
+ added <span className="font-mono">{state.done.id}</span> —{" "}
+ <Link href={`/browse/${state.done.project}`} className="underline">
+ open
+ </Link>
+ </span>
+ )}
+ {state.error && <span role="alert" className="text-[11px] text-[var(--color-bad)]">{state.error}</span>}
+ </div>
+ );
+}
diff --git a/umtool/components/articles/ArticleBody.tsx b/umtool/components/articles/ArticleBody.tsx
@@ -0,0 +1,178 @@
+"use client";
+
+import { createContext, memo, useContext, type AnchorHTMLAttributes, type ReactNode } from "react";
+import { Markdown } from "yt-dlp-transcript-common/components/Markdown";
+import type { CitationView, ReportPageView } from "yt-dlp-transcript-common/lib/report/views";
+
+// The article's text, as blocks the notes anchor to (`data-block`: title,
+// subtitle, summary, method, each section by id). MEMOISED and never
+// re-rendered once mounted: the reader wraps notes' quotes in <mark>s by
+// editing this DOM directly (components/articles/anchorDom.ts), which is safe
+// only because React has no reason to touch it again. Everything that changes
+// -- which citation is open, which note is selected -- goes through the
+// context below, whose value never changes.
+//
+// A citation is a button (its label, which is part of the text) and its
+// number (which is not: `data-anchor-skip`). Clicking it opens the evidence
+// panel; nothing here links out of umtool.
+
+export type ArticleActions = {
+ openCite: (id: string) => void;
+ noteSection: (section: string, title: string) => void;
+};
+
+export const ArticleActionsContext = createContext<ArticleActions | null>(null);
+
+const CitationsContext = createContext<Readonly<Record<string, CitationView>>>({});
+
+const CITE_SCHEME = "cite:";
+
+function CiteLink({ href, children }: AnchorHTMLAttributes<HTMLAnchorElement>) {
+ const actions = useContext(ArticleActionsContext);
+ const citations = useContext(CitationsContext);
+ const id = href?.startsWith(CITE_SCHEME) ? href.slice(CITE_SCHEME.length).trim() : null;
+ if (id === null) {
+ return (
+ <a href={href} target="_blank" rel="noopener noreferrer" className="text-[var(--color-sel)] underline decoration-[var(--color-sel)]/40">
+ {children}
+ </a>
+ );
+ }
+ const c = citations[id];
+ return (
+ <>
+ <button
+ type="button"
+ data-cite={id}
+ onClick={() => actions?.openCite(id)}
+ title={c ? `${c.quote}${c.speaker ? ` — ${c.speaker}` : ""}` : `citation ${id}`}
+ className="cursor-pointer rounded-sm text-left text-[var(--color-text)] underline decoration-[var(--color-sel)] decoration-dotted underline-offset-2 hover:bg-[var(--color-panel-2)] data-[has-note=true]:bg-[color-mix(in_srgb,var(--color-dirty)_22%,transparent)]"
+ >
+ {children}
+ </button>
+ <sup data-anchor-skip className="ml-0.5 font-mono text-[10px] text-[var(--color-sel)]">
+ {c?.number ?? "?"}
+ </sup>
+ </>
+ );
+}
+
+// markdown-to-jsx's elements carry common's class names, which umtool's
+// stylesheet does not have; these descendant rules are umtool's.
+const MD =
+ "text-[14px] leading-relaxed text-[var(--color-text)] [&_p]:my-2.5 [&_ul]:my-2 [&_ul]:list-disc [&_ul]:pl-5 [&_ol]:my-2 [&_ol]:list-decimal [&_ol]:pl-5 [&_li]:my-1 [&_blockquote]:my-2 [&_blockquote]:border-l-2 [&_blockquote]:border-[var(--color-line)] [&_blockquote]:pl-3 [&_blockquote]:text-[var(--color-dim)] [&_strong]:font-semibold [&_em]:italic [&_h3]:mt-4 [&_h3]:font-semibold [&_h4]:mt-3 [&_h4]:font-semibold [&_code]:font-mono [&_code]:text-[12px] [&_hr]:my-4 [&_hr]:border-[var(--color-line)]";
+
+function Md({ text }: { text: string }) {
+ return (
+ <Markdown className={MD} linkComponent={CiteLink}>
+ {text}
+ </Markdown>
+ );
+}
+
+function NoteButton({ section, title }: { section: string; title: string }) {
+ const actions = useContext(ArticleActionsContext);
+ return (
+ <button
+ type="button"
+ data-anchor-skip
+ aria-label={`note on ${title}`}
+ onClick={() => actions?.noteSection(section, title)}
+ className="ml-2 align-middle text-[11px] font-normal text-[var(--color-dim)] opacity-60 hover:text-[var(--color-sel)] hover:opacity-100"
+ >
+ + note
+ </button>
+ );
+}
+
+function ArticleBodyInner({ view, video }: { view: ReportPageView; video?: ReactNode }) {
+ return (
+ <CitationsContext.Provider value={view.citations}>
+ <div data-article-body className="space-y-5">
+ <header className="space-y-1">
+ {view.series && <div className="micro">{view.series}</div>}
+ <h1 data-block="title" className="text-[22px] font-semibold leading-tight text-[var(--color-text)]">
+ {view.title}
+ </h1>
+ {view.subtitle && (
+ <p data-block="subtitle" className="text-[15px] text-[var(--color-dim)]">
+ {view.subtitle}
+ </p>
+ )}
+ </header>
+ {/* The report's own video, under its title as the published page has
+ it. Outside every data-block, so it is never part of a quote. */}
+ {video}
+ {view.summary && (
+ <section data-block="summary" className="rounded border border-[var(--color-line)] bg-[var(--color-panel)] px-4 py-2">
+ <div className="micro pt-1" data-anchor-skip>
+ summary
+ <NoteButton section="summary" title="summary" />
+ </div>
+ <Md text={view.summary} />
+ </section>
+ )}
+ {view.sections.map((s) => (
+ <section key={s.id} id={`s-${s.id}`} data-block={s.id} className="scroll-mt-4">
+ <h2 className="mt-2 text-[17px] font-semibold text-[var(--color-text)]">
+ {s.title}
+ <NoteButton section={s.id} title={s.title} />
+ </h2>
+ {s.body && <Md text={s.body} />}
+ {s.claims.map((cl) => (
+ <div key={cl.id} data-claim={cl.id} className="my-3 rounded border border-[var(--color-line)] bg-[var(--color-panel)] px-3 py-2">
+ {cl.title && <div className="text-[13px] font-semibold text-[var(--color-text)]">{cl.title}</div>}
+ <div className="flex items-start gap-2">
+ <div className="min-w-0 flex-1">
+ <Md text={cl.text} />
+ </div>
+ {cl.verdict && (
+ <span data-anchor-skip className="mt-2 shrink-0 rounded border border-[var(--color-line)] px-1.5 py-0.5 font-mono text-[10px] uppercase text-[var(--color-dim)]">
+ {view.verdicts?.[cl.verdict]?.label ?? cl.verdict}
+ </span>
+ )}
+ </div>
+ {cl.findings && <Md text={cl.findings} />}
+ {cl.citations.length > 0 && (
+ <div data-anchor-skip className="mt-1 flex flex-wrap gap-1">
+ {cl.citations.map((id) => (
+ <CiteChip key={id} id={id} />
+ ))}
+ </div>
+ )}
+ </div>
+ ))}
+ </section>
+ ))}
+ {view.method && (
+ <section data-block="method" className="border-t border-[var(--color-line)] pt-3">
+ <div className="micro" data-anchor-skip>
+ method
+ <NoteButton section="method" title="method" />
+ </div>
+ <Md text={view.method} />
+ </section>
+ )}
+ </div>
+ </CitationsContext.Provider>
+ );
+}
+
+function CiteChip({ id }: { id: string }) {
+ const actions = useContext(ArticleActionsContext);
+ const c = useContext(CitationsContext)[id];
+ return (
+ <button
+ type="button"
+ data-cite={id}
+ onClick={() => actions?.openCite(id)}
+ title={c?.quote ?? id}
+ className="rounded border border-[var(--color-line)] px-1.5 py-0.5 font-mono text-[10px] text-[var(--color-sel)] hover:border-[var(--color-sel)] data-[has-note=true]:border-[var(--color-dirty)]"
+ >
+ {c?.number ?? "?"} {c?.label ?? c?.speaker ?? id}
+ </button>
+ );
+}
+
+const ArticleBody = memo(ArticleBodyInner, (a, b) => a.view === b.view && a.video === b.video);
+export default ArticleBody;
diff --git a/umtool/components/articles/ArticleReader.tsx b/umtool/components/articles/ArticleReader.tsx
@@ -0,0 +1,413 @@
+"use client";
+
+import { useCallback, useEffect, useLayoutEffect, useMemo, useRef, useState, type ReactNode } from "react";
+import type { ReportPageView } from "yt-dlp-transcript-common/lib/report/views";
+import CopyButton from "@/components/CopyButton";
+import { locateQuote, quoteAnchor } from "@/lib/annotations/anchor.mjs";
+import type { Anchor, Note, NotesRead } from "@/lib/annotations/types";
+import { useNotes } from "@/lib/annotations/useNotes";
+import { ShareNotes } from "@/components/notes/NotesProvider";
+import ArticleBody, { ArticleActionsContext, type ArticleActions } from "./ArticleBody";
+import ArticleVideo from "./ArticleVideo";
+import type { VideoProjectLink } from "./AddToVideo";
+import { blockText, selectionIn, unwrapMarks, wrapRange } from "./anchorDom";
+import EvidencePanel, { useEvidence } from "./EvidencePanel";
+import { Composer, NoteCard, anchorLabel } from "./NoteCards";
+
+// The article, its notes, and its evidence, in one screen: the text in a
+// ~70ch column, the notes in a rail beside it (a drawer below 1100px).
+//
+// select text → "Note" (or `n`) a note on that quote
+// "+ note" on a heading a note on the section
+// a citation its evidence in the rail; a note on it there
+// "+ whole article" a note on all of it
+//
+// Notes on a quote are found again in the text as it is NOW (lib/annotations/
+// anchor.mjs: the quote, then its context) and marked; one whose quote is gone
+// is still listed, flagged "orphaned", under its section.
+//
+// Keys (not while typing): j/k move between notes, r resolves, e edits, n
+// notes the selection (or the whole article), Esc closes whatever is open.
+
+type Filter = "open" | "resolved" | "all";
+
+const MARK_CLASS =
+ "rounded-sm bg-[color-mix(in_srgb,var(--color-sel)_22%,transparent)] text-inherit data-[active=true]:bg-[color-mix(in_srgb,var(--color-sel)_48%,transparent)] data-[status=resolved]:bg-transparent data-[status=resolved]:underline data-[status=resolved]:decoration-[var(--color-good)] data-[status=wontfix]:bg-transparent";
+
+type ComposerState = { anchor: Anchor; label: string } | null;
+
+export default function ArticleReader({
+ site,
+ report,
+ view,
+ initialNotes,
+ initialFilter = "open",
+ initialNote = null,
+ meta,
+ video,
+ videoProjects = [],
+ viewError,
+}: {
+ site: string;
+ report: string;
+ view: ReportPageView;
+ initialNotes: NotesRead;
+ initialFilter?: Filter;
+ initialNote?: string | null;
+ meta: ReactNode;
+ /** The report's own video (report.json `video`), played with timed notes. */
+ video?: { file: string; src: string; poster?: string; caption?: string } | null;
+ viewError?: string | null;
+ /** The article's LINKED video projects: where "Add to video" puts a cited span. */
+ videoProjects?: VideoProjectLink[];
+}) {
+ const notesApi = useNotes({ article: `${site}/${report}` }, { initial: initialNotes });
+ const { notes, write, busy, error } = notesApi;
+ // Stable across note writes: ArticleBody is memoised on it (its marks are
+ // laid over the rendered DOM). The player reads live notes through
+ // <ShareNotes>, not through this element.
+ const videoSlot = useMemo(
+ () =>
+ video ? (
+ <ArticleVideo site={site} report={report} file={video.file} src={video.src} poster={video.poster} caption={video.caption} />
+ ) : null,
+ [site, report, video?.file, video?.src, video?.poster, video?.caption],
+ );
+ const [filter, setFilter] = useState<Filter>(initialFilter);
+ const [selected, setSelected] = useState<string | null>(initialNote);
+ const [editing, setEditing] = useState<string | null>(null);
+ const [composer, setComposer] = useState<ComposerState>(null);
+ const [pending, setPending] = useState<{ anchor: Anchor; label: string; x: number; y: number } | null>(null);
+ const [cite, setCite] = useState<string | null>(null);
+ const [orphans, setOrphans] = useState<Set<string>>(new Set());
+ const [hover, setHover] = useState<string | null>(null);
+ const [drawer, setDrawer] = useState(false);
+ const bodyRef = useRef<HTMLDivElement>(null);
+ const evidence = useEvidence(site, report, cite);
+
+ const titles = useMemo(() => {
+ const t: Record<string, string> = { title: "title", subtitle: "subtitle", summary: "summary", method: "method" };
+ for (const s of view.sections) t[s.id] = s.title;
+ return t;
+ }, [view]);
+ const numbers = useMemo(() => Object.fromEntries(Object.entries(view.citations).map(([id, c]) => [id, c.number])), [view]);
+ const blockOrder = useMemo(() => ["title", "subtitle", "summary", ...view.sections.map((s) => s.id), "method"], [view]);
+
+ const visible = useMemo(() => {
+ const keep = (n: Note) => filter === "all" || (filter === "open" ? n.status === "open" : n.status !== "open");
+ const rank = (n: Note) => {
+ const a = n.anchor;
+ if (a.kind === "whole") return -1;
+ if (a.kind === "text" || a.kind === "section") return blockOrder.indexOf(a.section) + 0.5;
+ if (a.kind === "cite") return 1000 + (numbers[a.cite] ?? 0);
+ return 2000;
+ };
+ return notes.filter(keep).sort((a, b) => rank(a) - rank(b) || a.at.localeCompare(b.at));
+ }, [notes, filter, blockOrder, numbers]);
+
+ // ---- marks ---------------------------------------------------------------
+ useLayoutEffect(() => {
+ const root = bodyRef.current;
+ if (!root) return;
+ unwrapMarks(root);
+ const lost = new Set<string>();
+ for (const n of visible) {
+ const a = n.anchor;
+ if (a.kind !== "text") continue;
+ const block = root.querySelector(`[data-block="${CSS.escape(a.section)}"]`);
+ if (!block) {
+ lost.add(n.id);
+ continue;
+ }
+ const at = locateQuote(blockText(block).text, a);
+ if (!at.found) {
+ lost.add(n.id);
+ continue;
+ }
+ wrapRange(block, at.start, at.end, { "data-note": n.id, "data-status": n.status }, MARK_CLASS);
+ }
+ // Citation notes: mark the citation itself.
+ for (const el of Array.from(root.querySelectorAll("[data-cite][data-has-note]"))) el.removeAttribute("data-has-note");
+ for (const n of visible) {
+ if (n.anchor.kind === "cite") {
+ for (const el of Array.from(root.querySelectorAll(`[data-cite="${CSS.escape(n.anchor.cite)}"]`))) el.setAttribute("data-has-note", "true");
+ }
+ }
+ setOrphans((prev) => (prev.size === lost.size && [...lost].every((id) => prev.has(id)) ? prev : lost));
+ }, [visible]);
+
+ // The selected / hovered note's marks light up.
+ useEffect(() => {
+ const root = bodyRef.current;
+ if (!root) return;
+ for (const m of Array.from(root.querySelectorAll("mark[data-note]"))) {
+ const id = m.getAttribute("data-note");
+ m.setAttribute("data-active", String(id === selected || id === hover));
+ }
+ }, [selected, hover, visible]);
+
+ const focusNote = useCallback((id: string, scroll = true) => {
+ setSelected(id);
+ if (!scroll) return;
+ requestAnimationFrame(() => {
+ const mark = bodyRef.current?.querySelector(`mark[data-note="${CSS.escape(id)}"]`);
+ mark?.scrollIntoView({ block: "center", behavior: "smooth" });
+ document.querySelector(`[data-note-card="${CSS.escape(id)}"]`)?.scrollIntoView({ block: "nearest" });
+ });
+ }, []);
+
+ // A deep link (`?note=`) lands on its note, whatever the filter says.
+ useEffect(() => {
+ if (!initialNote) return;
+ const n = notes.find((x) => x.id === initialNote);
+ if (n && filter !== "all" && (filter === "open") !== (n.status === "open")) setFilter("all");
+ focusNote(initialNote);
+ // once, on mount
+ // eslint-disable-next-line react-hooks/exhaustive-deps
+ }, []);
+
+ // ---- selection → "Note" ---------------------------------------------------
+ const readSelection = useCallback(() => {
+ const root = bodyRef.current;
+ if (!root) return;
+ const sel = selectionIn(root);
+ if (!sel) {
+ setPending(null);
+ return;
+ }
+ const section = sel.block.getAttribute("data-block") ?? "";
+ const q = quoteAnchor(blockText(sel.block).text, sel.start, sel.end);
+ if (!q.quote) {
+ setPending(null);
+ return;
+ }
+ const anchor: Anchor = { kind: "text", section, ...q };
+ setPending({ anchor, label: anchorLabel(anchor, titles, numbers), x: sel.rect.right, y: sel.rect.top });
+ }, [titles, numbers]);
+
+ const openComposer = useCallback((anchor: Anchor) => {
+ setComposer({ anchor, label: anchorLabel(anchor, titles, numbers) });
+ setPending(null);
+ setDrawer(true);
+ window.getSelection()?.removeAllRanges();
+ }, [titles, numbers]);
+
+ // The body's buttons talk to the reader through a context whose value never
+ // changes (so the memoised body never re-renders): it reads the latest
+ // callbacks from a ref.
+ const latest = useRef({ setCite, openComposer, setDrawer });
+ latest.current = { setCite, openComposer, setDrawer };
+ const actions = useMemo<ArticleActions>(
+ () => ({
+ openCite: (id) => {
+ latest.current.setCite(id);
+ latest.current.setDrawer(true);
+ },
+ noteSection: (section) => latest.current.openComposer({ kind: "section", section }),
+ }),
+ [],
+ );
+
+ const add = async (anchor: Anchor, text: string) => {
+ const note = await write({ op: "add", text, anchor });
+ if (note) {
+ setComposer(null);
+ if (filter === "resolved") setFilter("open");
+ focusNote(note.id, false);
+ }
+ };
+
+ // ---- keys -----------------------------------------------------------------
+ useEffect(() => {
+ const onKey = (e: KeyboardEvent) => {
+ const t = e.target as HTMLElement | null;
+ if (t && (t.closest("input, textarea, select, [contenteditable=true]") || e.metaKey || e.ctrlKey || e.altKey)) return;
+ const idx = visible.findIndex((n) => n.id === selected);
+ if (e.key === "j" || e.key === "k") {
+ if (!visible.length) return;
+ e.preventDefault();
+ const next = e.key === "j" ? (idx < 0 ? 0 : Math.min(visible.length - 1, idx + 1)) : idx < 0 ? 0 : Math.max(0, idx - 1);
+ focusNote(visible[next].id);
+ } else if (e.key === "r" && idx >= 0) {
+ e.preventDefault();
+ void write({ op: "status", id: visible[idx].id, status: "resolved" });
+ } else if (e.key === "e" && idx >= 0 && visible[idx].author === "operator") {
+ e.preventDefault();
+ setEditing(visible[idx].id);
+ } else if (e.key === "n") {
+ e.preventDefault();
+ if (pending) openComposer(pending.anchor);
+ else if (cite) openComposer({ kind: "cite", cite });
+ else openComposer({ kind: "whole" });
+ } else if (e.key === "Escape") {
+ setComposer(null);
+ setEditing(null);
+ setPending(null);
+ setCite(null);
+ setDrawer(false);
+ }
+ };
+ window.addEventListener("keydown", onKey);
+ return () => window.removeEventListener("keydown", onKey);
+ }, [visible, selected, pending, cite, write, focusNote, openComposer]);
+
+ const counts = {
+ open: notes.filter((n) => n.status === "open").length,
+ resolved: notes.filter((n) => n.status !== "open").length,
+ all: notes.length,
+ };
+ const citeNotes = cite ? notes.filter((n) => n.anchor.kind === "cite" && n.anchor.cite === cite) : [];
+ const card = (n: Note) => (
+ <NoteCard
+ key={n.id}
+ note={n}
+ label={anchorLabel(n.anchor, titles, numbers)}
+ orphaned={orphans.has(n.id)}
+ selected={selected === n.id}
+ editing={editing === n.id}
+ busy={busy}
+ onSelect={() => focusNote(n.id)}
+ onHover={(on) => setHover(on ? n.id : null)}
+ onEdit={() => setEditing(n.id)}
+ onCancelEdit={() => setEditing(null)}
+ write={write}
+ />
+ );
+
+ return (
+ <div className="flex min-h-0 flex-1">
+ <main className="deck-main min-w-0 flex-1 p-4">
+ <div className="mx-auto max-w-[72ch] space-y-4">
+ {meta}
+ {viewError && <p className="text-[12px] text-[var(--color-bad)]">citations not shown: {viewError}</p>}
+ <ArticleActionsContext.Provider value={actions}>
+ <div
+ ref={bodyRef}
+ onMouseUp={() => setTimeout(readSelection, 0)}
+ onKeyUp={(e) => e.shiftKey && readSelection()}
+ onMouseOver={(e) => {
+ const m = (e.target as HTMLElement).closest?.("mark[data-note]");
+ setHover(m ? m.getAttribute("data-note") : null);
+ }}
+ onClick={(e) => {
+ const m = (e.target as HTMLElement).closest?.("mark[data-note]");
+ const id = m?.getAttribute("data-note");
+ if (id && window.getSelection()?.isCollapsed) {
+ setDrawer(true);
+ focusNote(id, false);
+ document.querySelector(`[data-note-card="${CSS.escape(id)}"]`)?.scrollIntoView({ block: "nearest" });
+ }
+ }}
+ >
+ <ShareNotes target={{ article: `${site}/${report}` }} notes={notesApi}>
+ <ArticleBody view={view} video={videoSlot} />
+ </ShareNotes>
+ </div>
+ </ArticleActionsContext.Provider>
+ </div>
+ </main>
+
+ {pending && (
+ <button
+ type="button"
+ onMouseDown={(e) => e.preventDefault()}
+ onClick={() => openComposer(pending.anchor)}
+ style={{ left: Math.min(pending.x + 6, window.innerWidth - 70), top: Math.min(window.innerHeight - 32, Math.max(8, pending.y - 30)) }}
+ className="fixed z-50 rounded bg-[var(--color-sel)] px-2 py-0.5 text-[12px] font-medium text-[var(--color-ink)] shadow"
+ >
+ Note
+ </button>
+ )}
+
+ <button
+ type="button"
+ onClick={() => setDrawer((d) => !d)}
+ className="fixed bottom-3 right-3 z-30 rounded border border-[var(--color-line)] bg-[var(--color-panel)] px-2 py-1 text-[12px] min-[1100px]:hidden"
+ >
+ notes {counts.open}
+ </button>
+
+ <aside
+ aria-label="notes"
+ data-drawer={drawer ? "open" : "closed"}
+ className={`fixed inset-y-0 right-0 z-40 flex w-[min(380px,92vw)] flex-col border-l border-[var(--color-line)] bg-[var(--color-ink)] transition-transform min-[1100px]:static min-[1100px]:z-auto min-[1100px]:w-[360px] min-[1100px]:shrink-0 min-[1100px]:translate-x-0 ${drawer ? "translate-x-0" : "translate-x-full"}`}
+ >
+ <div className="flex flex-wrap items-center gap-1 border-b border-[var(--color-line)] p-2">
+ {(["open", "resolved", "all"] as Filter[]).map((f) => (
+ <button
+ key={f}
+ type="button"
+ aria-pressed={filter === f}
+ onClick={() => setFilter(f)}
+ className={`rounded border px-1.5 py-0.5 font-mono text-[11px] ${filter === f ? "border-[var(--color-sel)] text-[var(--color-sel)]" : "border-[var(--color-line)] text-[var(--color-dim)]"}`}
+ >
+ {f} {counts[f]}
+ </button>
+ ))}
+ <button
+ type="button"
+ onClick={() => openComposer({ kind: "whole" })}
+ className="ml-auto text-[11px] text-[var(--color-dim)] hover:text-[var(--color-sel)]"
+ >
+ + whole article
+ </button>
+ <button type="button" aria-label="close notes" onClick={() => setDrawer(false)} className="text-[12px] text-[var(--color-dim)] min-[1100px]:hidden">
+ ✕
+ </button>
+ </div>
+ <div className="flex items-center gap-2 border-b border-[var(--color-line)] px-2 py-1">
+ <CopyButton
+ url={`/api/notes/context?${new URLSearchParams({ article: `${site}/${report}` })}`}
+ label="copy agent brief"
+ title={`umtool notes ${site}/${report}`}
+ className="shrink-0 whitespace-nowrap"
+ />
+ {notesApi.source?.draft && (
+ <span className="truncate font-mono text-[10px] text-[var(--color-dim)]" title={notesApi.source.how ?? ""}>
+ edit {notesApi.source.draft}
+ </span>
+ )}
+ </div>
+ <div className="min-h-0 flex-1 space-y-2 overflow-y-auto p-2">
+ {error && <p className="text-[12px] text-[var(--color-bad)]">{error}</p>}
+
+ {cite && (
+ <section data-evidence-rail className="space-y-2 rounded border border-[var(--color-line)] p-2">
+ <div className="flex items-center">
+ <span className="micro">citation [{numbers[cite] ?? cite}]</span>
+ <a href={`/sites/${site}/${report}/evidence?c=${encodeURIComponent(cite)}`} className="ml-auto text-[11px] text-[var(--color-dim)] hover:text-[var(--color-sel)]">
+ walk →
+ </a>
+ <button type="button" aria-label="close evidence" onClick={() => setCite(null)} className="ml-2 text-[12px] text-[var(--color-dim)] hover:text-[var(--color-text)]">
+ ✕
+ </button>
+ </div>
+ <EvidencePanel ev={evidence.ev} error={evidence.error} number={numbers[cite]} projects={videoProjects} />
+ {citeNotes.length > 0 && <ul className="space-y-2">{citeNotes.map(card)}</ul>}
+ {composer?.anchor.kind === "cite" && composer.anchor.cite === cite ? (
+ <Composer label={composer.label} busy={busy} onSave={(text) => add(composer.anchor, text)} onCancel={() => setComposer(null)} />
+ ) : (
+ <button type="button" onClick={() => openComposer({ kind: "cite", cite })} className="text-[11px] text-[var(--color-dim)] hover:text-[var(--color-sel)]">
+ + note on this citation
+ </button>
+ )}
+ </section>
+ )}
+
+ {composer && !(composer.anchor.kind === "cite" && composer.anchor.cite === cite) && (
+ <Composer label={composer.label} busy={busy} onSave={(text) => add(composer.anchor, text)} onCancel={() => setComposer(null)} />
+ )}
+
+ {visible.length === 0 ? (
+ <p className="text-[12px] text-[var(--color-dim)]">{notes.length === 0 ? "no notes" : `no ${filter} notes`}</p>
+ ) : (
+ <ul data-notes-list className="space-y-2">
+ {/* A note on the open citation is listed with its evidence above. */}
+ {visible.filter((n) => !(cite && n.anchor.kind === "cite" && n.anchor.cite === cite)).map(card)}
+ </ul>
+ )}
+ </div>
+ </aside>
+ </div>
+ );
+}
diff --git a/umtool/components/articles/ArticleTable.tsx b/umtool/components/articles/ArticleTable.tsx
@@ -0,0 +1,103 @@
+import Link from "next/link";
+import { badgeVariants } from "@/components/ui/badge";
+import { fmtAgo } from "@/lib/format";
+import type { ArticleRow } from "@/lib/articles/sites";
+
+// One row per article: what it is, where it stands, what is waiting on it,
+// and where it came from. Server-rendered, zero JS -- the /browse idiom.
+
+const posterUrl = (a: ArticleRow) =>
+ `/api/sites/media?site=${encodeURIComponent(a.site)}&report=${encodeURIComponent(a.id)}&file=poster.jpg`;
+
+const baseName = (p: string) => p.split("/").pop() ?? p;
+
+export function articleHref(a: { site: string; id: string }) {
+ return `/sites/${a.site}/${a.id}`;
+}
+
+export default function ArticleTable({ articles }: { articles: ArticleRow[] }) {
+ if (articles.length === 0) return <p className="text-[12px] text-[var(--color-dim)]">no articles</p>;
+ return (
+ <table className="w-full border-collapse text-[12px]" data-testid="article-table">
+ <thead>
+ <tr className="text-left text-[11px] text-[var(--color-dim)]">
+ <th className="w-[72px] py-1 font-normal" />
+ <th className="py-1 font-normal">article</th>
+ <th className="py-1 font-normal">status</th>
+ <th className="py-1 font-normal">updated</th>
+ <th className="py-1 text-right font-normal">cites</th>
+ <th className="py-1 text-right font-normal">notes</th>
+ <th className="py-1 pl-3 font-normal">video project</th>
+ <th className="py-1 font-normal">source</th>
+ </tr>
+ </thead>
+ <tbody>
+ {articles.map((a) => (
+ <tr
+ key={`${a.site}/${a.id}`}
+ data-article={`${a.site}/${a.id}`}
+ data-status={a.status}
+ className="border-t border-[var(--color-line)] align-top"
+ >
+ <td className="py-1.5 pr-2">
+ {a.hasPoster ? (
+ // eslint-disable-next-line @next/next/no-img-element
+ <img src={posterUrl(a)} alt="" loading="lazy" className="h-9 w-16 rounded object-cover" />
+ ) : (
+ <div className="h-9 w-16 rounded bg-[var(--color-panel-2)]" />
+ )}
+ </td>
+ <td className="py-1.5 pr-3">
+ <Link href={articleHref(a)} className="text-[var(--color-text)] hover:text-[var(--color-sel)]">
+ {a.title}
+ </Link>
+ <div className="font-mono text-[11px] text-[var(--color-dim)]">
+ {a.id}
+ {a.series ? ` · ${a.series}` : ""}
+ {a.problem ? <span className="text-[var(--color-bad)]"> · {a.problems} problem{a.problems === 1 ? "" : "s"}</span> : null}
+ </div>
+ </td>
+ <td className="py-1.5 pr-3">
+ <span className={badgeVariants({ variant: a.status === "published" ? "on" : "neutral", size: "sm" })}>
+ {a.status}
+ </span>
+ </td>
+ <td className="num py-1.5 pr-3 text-[var(--color-dim)]">
+ {a.updated ?? (a.mtimeMs ? fmtAgo(a.mtimeMs) : "—")}
+ </td>
+ <td className="num py-1.5 pr-3 text-right">{a.citations}</td>
+ <td className="num py-1.5 text-right">
+ {a.openNotes > 0 ? (
+ <Link
+ href={`${articleHref(a)}?status=open`}
+ data-open-notes={a.openNotes}
+ className={badgeVariants({ variant: "open", size: "sm" })}
+ >
+ {a.openNotes} open
+ </Link>
+ ) : (
+ <span className="text-[var(--color-dim)]">{a.notes || "—"}</span>
+ )}
+ </td>
+ <td className="py-1.5 pl-3 pr-3">
+ {a.projects.linked.map((p) => (
+ <Link key={p.id} href={`/browse/${p.id}`} data-project-link={p.id} className="block font-mono text-[11px] text-[var(--color-sel)] hover:underline">
+ {p.id}
+ </Link>
+ ))}
+ {a.projects.possible.map((p) => (
+ <Link key={p.id} href={`/browse/${p.id}`} title="shares the slug; not linked" className="block font-mono text-[11px] text-[var(--color-dim)] hover:underline">
+ {p.id} (possible)
+ </Link>
+ ))}
+ {a.projects.linked.length + a.projects.possible.length === 0 && <span className="text-[var(--color-dim)]">—</span>}
+ </td>
+ <td className="py-1.5 font-mono text-[11px] text-[var(--color-dim)]" title={a.source?.how ?? ""}>
+ {a.source?.draft ? baseName(a.source.draft) : a.source?.generator ? baseName(a.source.generator) : "—"}
+ </td>
+ </tr>
+ ))}
+ </tbody>
+ </table>
+ );
+}
diff --git a/umtool/components/articles/ArticleVideo.tsx b/umtool/components/articles/ArticleVideo.tsx
@@ -0,0 +1,34 @@
+"use client";
+
+import { TimedVideo } from "@/components/notes/TimedNotes";
+
+// THE ARTICLE'S OWN VIDEO -- the report's video (report.json `video.src`), as
+// the published page plays it, with timed notes under it: `n` or Mark at the
+// playhead writes a moment anchor `{ kind: "moment", file: <video.src>, t }`
+// to the ARTICLE's notes. The handle is the reader's own, shared through
+// <ShareNotes> (components/notes/NotesProvider.tsx), so a mark and a text note
+// never race each other's token. No schedule: a report video is a published
+// file, so a mark keeps its time only. Keep this the only place the article
+// page renders its video.
+export default function ArticleVideo({
+ site,
+ report,
+ file,
+ src,
+ poster,
+ caption,
+}: {
+ site: string;
+ report: string;
+ file: string;
+ src: string;
+ poster?: string;
+ caption?: string;
+}) {
+ return (
+ <figure data-article-video className="space-y-1">
+ <TimedVideo target={{ article: `${site}/${report}` }} file={file} src={src} poster={poster} testId="article-video" showErrors={false} />
+ {caption && <figcaption className="text-[12px] text-[var(--color-dim)]">{caption}</figcaption>}
+ </figure>
+ );
+}
diff --git a/umtool/components/articles/EvidencePanel.tsx b/umtool/components/articles/EvidencePanel.tsx
@@ -0,0 +1,167 @@
+"use client";
+
+import { forwardRef, useEffect, useImperativeHandle, useRef, useState } from "react";
+import CopyButton from "@/components/CopyButton";
+import type { Evidence } from "@/lib/articles/evidence";
+import AddToVideo, { type VideoProjectLink } from "./AddToVideo";
+
+// One citation's evidence: what it quotes, who and when, the transcript around
+// it with the cited cues marked, and the media that plays it -- the prepared
+// clip, a fetched window or a saved file, seeked to the cited second -- or,
+// when there is none on disk, the MCP line that fetches it through the editor.
+
+export type EvidenceHandle = { togglePlay: () => void };
+
+const fmt = (s: number) => {
+ const m = Math.floor(s / 60);
+ const r = Math.floor(s % 60);
+ return `${m}:${String(r).padStart(2, "0")}`;
+};
+
+export function useEvidence(site: string, report: string, cite: string | null) {
+ const [ev, setEv] = useState<Evidence | null>(null);
+ const [error, setError] = useState<string | null>(null);
+ useEffect(() => {
+ if (!cite) return;
+ let live = true;
+ setEv(null);
+ setError(null);
+ fetch(`/api/sites/evidence?${new URLSearchParams({ site, report, cite })}`, { cache: "no-store" })
+ .then(async (r) => {
+ const j = await r.json();
+ if (!r.ok) throw new Error(j.error ?? r.statusText);
+ if (live) setEv(j);
+ })
+ .catch((err) => live && setError(err instanceof Error ? err.message : String(err)));
+ return () => {
+ live = false;
+ };
+ }, [site, report, cite]);
+ return { ev, error };
+}
+
+const EvidencePanel = forwardRef<
+ EvidenceHandle,
+ { ev: Evidence | null; error: string | null; number?: number; autoPlay?: boolean; projects?: VideoProjectLink[] }
+>(
+ function EvidencePanel({ ev, error, number, autoPlay = false, projects = [] }, ref) {
+ const media = useRef<HTMLMediaElement | null>(null);
+ useImperativeHandle(ref, () => ({
+ togglePlay: () => {
+ const m = media.current;
+ if (!m) return;
+ if (m.paused) void m.play().catch(() => {});
+ else m.pause();
+ },
+ }));
+
+ if (error) return <p className="text-[12px] text-[var(--color-bad)]">{error}</p>;
+ if (!ev) return <p className="text-[12px] text-[var(--color-dim)]">loading…</p>;
+
+ const play = ev.play;
+ // File time of a record second: the cited span starts `offset` into the file.
+ const fileT = (t: number) => (play && ev.start !== undefined ? Math.max(0, t - ev.start + play.offset) : 0);
+ const startAt = play ? (play.kind === "prepared" ? 0 : Math.max(0, play.offset - 3)) : 0;
+ const seek = (t: number) => {
+ if (media.current) {
+ media.current.currentTime = fileT(t);
+ void media.current.play().catch(() => {});
+ }
+ };
+
+ return (
+ <div data-evidence={ev.cite} className="space-y-2 text-[12px]">
+ <blockquote className="border-l-2 border-[var(--color-sel)] pl-2 text-[13px] text-[var(--color-text)]">
+ {number !== undefined && <span className="mr-1 font-mono text-[11px] text-[var(--color-sel)]">[{number}]</span>}“{ev.quote}”
+ </blockquote>
+ <div className="text-[var(--color-dim)]">
+ {[ev.speaker ?? ev.record?.channelTitle, ev.date, ev.label].filter(Boolean).join(" · ")}
+ {ev.start !== undefined && ev.end !== undefined && (
+ <span className="num ml-1">
+ @ {fmt(ev.start)}–{fmt(ev.end)}
+ </span>
+ )}
+ </div>
+ {ev.record?.title && <div className="text-[var(--color-dim)]">{ev.record.title}</div>}
+ {ev.originalUrl && (
+ <a href={ev.originalUrl} target="_blank" rel="noopener noreferrer" className="text-[var(--color-sel)] hover:underline">
+ original ↗
+ </a>
+ )}
+
+ {play ? (
+ <div data-play={play.kind}>
+ {play.audio ? (
+ <audio
+ ref={(el) => {
+ media.current = el;
+ }}
+ controls
+ src={play.url}
+ autoPlay={autoPlay}
+ onLoadedMetadata={(e) => {
+ e.currentTarget.currentTime = startAt;
+ }}
+ className="w-full"
+ />
+ ) : (
+ <video
+ ref={(el) => {
+ media.current = el;
+ }}
+ controls
+ src={play.url}
+ autoPlay={autoPlay}
+ onLoadedMetadata={(e) => {
+ e.currentTarget.currentTime = startAt;
+ }}
+ className="aspect-video w-full rounded bg-black"
+ />
+ )}
+ <div className="micro mt-0.5">{play.label}</div>
+ </div>
+ ) : ev.fetchLine ? (
+ <div data-play="none" className="space-y-1 rounded border border-[var(--color-line)] bg-[var(--color-panel-2)] p-2">
+ <div className="text-[var(--color-dim)]">not on disk — fetch it through the editor:</div>
+ <code className="block break-all font-mono text-[11px] text-[var(--color-text)]">{ev.fetchLine}</code>
+ <CopyButton text={ev.fetchLine} label="copy fetch_clip" />
+ </div>
+ ) : null}
+
+ {ev.post && (
+ <div data-post className="space-y-1 rounded border border-[var(--color-line)] p-2">
+ {ev.post.author && <div className="text-[var(--color-dim)]">{ev.post.author}</div>}
+ {ev.post.text && <p className="whitespace-pre-wrap text-[var(--color-text)]">{ev.post.text}</p>}
+ {ev.post.shot && (
+ // eslint-disable-next-line @next/next/no-img-element
+ <img src={ev.post.shot} alt="capture of the post" className="w-full rounded border border-[var(--color-line)]" />
+ )}
+ </div>
+ )}
+
+ {ev.cues.length > 0 ? (
+ <ol data-cues className="space-y-0.5">
+ {ev.cues.map((c) => (
+ <li key={c.start} data-cited={c.cited ? "true" : undefined}>
+ <button
+ type="button"
+ onClick={() => seek(c.start)}
+ disabled={!play}
+ className={`flex w-full gap-2 rounded px-1 text-left ${c.cited ? "bg-[var(--color-panel-2)] text-[var(--color-text)]" : "text-[var(--color-dim)]"} enabled:hover:text-[var(--color-text)]`}
+ >
+ <span className="num shrink-0 font-mono text-[10px] leading-5">{fmt(c.start)}</span>
+ <span>{c.text}</span>
+ </button>
+ </li>
+ ))}
+ </ol>
+ ) : ev.cuesNote ? (
+ <div className="micro">{ev.cuesNote}</div>
+ ) : null}
+ <AddToVideo ev={ev} projects={projects} />
+ </div>
+ );
+ },
+);
+
+export default EvidencePanel;
diff --git a/umtool/components/articles/EvidenceWalk.tsx b/umtool/components/articles/EvidenceWalk.tsx
@@ -0,0 +1,129 @@
+"use client";
+
+import type { VideoProjectLink } from "./AddToVideo";
+import Link from "next/link";
+import { useCallback, useEffect, useRef, useState } from "react";
+import type { NotesRead } from "@/lib/annotations/types";
+import { useNotes } from "@/lib/annotations/useNotes";
+import EvidencePanel, { useEvidence, type EvidenceHandle } from "./EvidencePanel";
+import { Composer, NoteCard } from "./NoteCards";
+
+// Every citation of one article, one per screen: the evidence on the left, its
+// notes on the right. j/k (or ←/→) move, space plays, n notes it, Esc cancels.
+// The citation in view is in the URL (`?c=`), so a walk can be resumed.
+
+export type WalkItem = { id: string; number: number; kind: string; quote: string };
+
+export default function EvidenceWalk({
+ site,
+ report,
+ items,
+ initial,
+ initialNotes,
+ videoProjects = [],
+}: {
+ site: string;
+ report: string;
+ items: WalkItem[];
+ initial: number;
+ initialNotes: NotesRead;
+ videoProjects?: VideoProjectLink[];
+}) {
+ const [idx, setIdx] = useState(initial);
+ const [composing, setComposing] = useState(false);
+ const item = items[idx];
+ const { ev, error } = useEvidence(site, report, item?.id ?? null);
+ const { notes, write, busy, error: notesError } = useNotes({ article: `${site}/${report}` }, { initial: initialNotes });
+ const panel = useRef<EvidenceHandle>(null);
+
+ const go = useCallback(
+ (to: number) => {
+ const next = Math.max(0, Math.min(items.length - 1, to));
+ setIdx(next);
+ setComposing(false);
+ const url = new URL(window.location.href);
+ url.searchParams.set("c", items[next].id);
+ window.history.replaceState(null, "", url);
+ },
+ [items],
+ );
+
+ useEffect(() => {
+ const onKey = (e: KeyboardEvent) => {
+ const t = e.target as HTMLElement | null;
+ if (t?.closest("input, textarea, select, [contenteditable=true]") || e.metaKey || e.ctrlKey || e.altKey) return;
+ if (e.key === "j" || e.key === "ArrowRight") {
+ e.preventDefault();
+ go(idx + 1);
+ } else if (e.key === "k" || e.key === "ArrowLeft") {
+ e.preventDefault();
+ go(idx - 1);
+ } else if (e.key === " ") {
+ e.preventDefault();
+ panel.current?.togglePlay();
+ } else if (e.key === "n") {
+ e.preventDefault();
+ setComposing(true);
+ } else if (e.key === "Escape") {
+ setComposing(false);
+ }
+ };
+ window.addEventListener("keydown", onKey);
+ return () => window.removeEventListener("keydown", onKey);
+ }, [idx, go]);
+
+ if (!item) return <p className="p-4 text-[12px] text-[var(--color-dim)]">no citations</p>;
+ const mine = notes.filter((n) => n.anchor.kind === "cite" && n.anchor.cite === item.id);
+
+ return (
+ <div data-walk={item.id} className="grid min-h-0 flex-1 grid-cols-1 gap-4 overflow-auto p-4 min-[1100px]:grid-cols-[minmax(0,1.3fr)_minmax(320px,1fr)]">
+ <div className="min-w-0 space-y-2">
+ <div className="flex items-center gap-2 text-[12px]">
+ <button type="button" onClick={() => go(idx - 1)} disabled={idx === 0} className="rounded border border-[var(--color-line)] px-2 disabled:opacity-30">
+ ← k
+ </button>
+ <span className="num" data-walk-position>
+ {idx + 1} / {items.length}
+ </span>
+ <button type="button" onClick={() => go(idx + 1)} disabled={idx === items.length - 1} className="rounded border border-[var(--color-line)] px-2 disabled:opacity-30">
+ j →
+ </button>
+ <span className="micro ml-2">{item.kind}</span>
+ <span className="micro ml-auto">space plays · n notes</span>
+ </div>
+ <EvidencePanel key={item.id} ref={panel} ev={ev} error={error} number={item.number} projects={videoProjects} />
+ </div>
+ <div className="min-w-0 space-y-2">
+ <div className="micro">notes on [{item.number}]</div>
+ {notesError && <p className="text-[12px] text-[var(--color-bad)]">{notesError}</p>}
+ {mine.length > 0 && (
+ <ul className="space-y-2">
+ {mine.map((n) => (
+ <NoteCard key={n.id} note={n} label={`citation [${item.number}]`} busy={busy} write={write} />
+ ))}
+ </ul>
+ )}
+ {composing ? (
+ <Composer
+ label={`citation [${item.number}]`}
+ busy={busy}
+ onSave={async (text) => {
+ const note = await write({ op: "add", text, anchor: { kind: "cite", cite: item.id } });
+ if (note) setComposing(false);
+ }}
+ onCancel={() => setComposing(false)}
+ />
+ ) : (
+ <button type="button" onClick={() => setComposing(true)} className="text-[12px] text-[var(--color-dim)] hover:text-[var(--color-sel)]">
+ + note on this citation
+ </button>
+ )}
+ <div className="pt-2">
+ <Link href={`/sites/${site}/${report}`} className="text-[12px] text-[var(--color-dim)] hover:text-[var(--color-text)]">
+ ← article
+ </Link>
+ </div>
+ </div>
+ </div>
+ );
+}
diff --git a/umtool/components/articles/NoteCards.tsx b/umtool/components/articles/NoteCards.tsx
@@ -0,0 +1,240 @@
+"use client";
+
+import { useEffect, useRef, useState } from "react";
+import { NOTE_TEXT_LIMIT } from "@/lib/annotations/types";
+import type { Anchor, Note, NoteOp } from "@/lib/annotations/types";
+import { fmtAgo } from "@/lib/format";
+
+// A note in the rail, and the box that writes one. Shared by the article
+// reader and the evidence walk.
+
+export function anchorLabel(a: Anchor, titles: Record<string, string>, cites: Record<string, number | undefined>): string {
+ const sec = (id: string) => titles[id] ?? id;
+ switch (a.kind) {
+ case "text":
+ return `${sec(a.section)} · “${a.quote.length > 60 ? `${a.quote.slice(0, 59)}…` : a.quote}”`;
+ case "section":
+ return `section · ${sec(a.section)}`;
+ case "cite":
+ return `citation [${cites[a.cite] ?? a.cite}]`;
+ case "whole":
+ return "whole article";
+ case "moment":
+ return `${a.file} @ ${a.t.toFixed(1)}s`;
+ default:
+ return a.kind;
+ }
+}
+
+export function Composer({
+ label,
+ initial = "",
+ busy,
+ onSave,
+ onCancel,
+ saveLabel = "save note",
+}: {
+ label: string;
+ initial?: string;
+ busy?: boolean;
+ onSave: (text: string) => void | Promise<void>;
+ onCancel: () => void;
+ saveLabel?: string;
+}) {
+ const [text, setText] = useState(initial);
+ const ref = useRef<HTMLTextAreaElement>(null);
+ useEffect(() => ref.current?.focus(), []);
+ const save = () => {
+ if (text.trim()) void onSave(text);
+ };
+ return (
+ <div data-composer className="space-y-1 rounded border border-[var(--color-sel)] bg-[var(--color-panel)] p-2">
+ <div className="micro truncate" title={label}>
+ {label}
+ </div>
+ <textarea
+ ref={ref}
+ aria-label="note text"
+ value={text}
+ maxLength={NOTE_TEXT_LIMIT}
+ onChange={(e) => setText(e.target.value)}
+ onKeyDown={(e) => {
+ if (e.key === "Enter" && (e.metaKey || e.ctrlKey)) {
+ e.preventDefault();
+ save();
+ } else if (e.key === "Escape") {
+ e.preventDefault();
+ onCancel();
+ }
+ }}
+ rows={4}
+ className="w-full resize-y rounded border border-[var(--color-line)] bg-[var(--color-ink)] p-1.5 text-[12px] text-[var(--color-text)] outline-none focus:border-[var(--color-sel)]"
+ />
+ <div className="flex items-center gap-2">
+ <button
+ type="button"
+ onClick={save}
+ disabled={busy || !text.trim()}
+ className="rounded bg-[var(--color-sel)] px-2 py-0.5 text-[12px] text-[var(--color-ink)] disabled:opacity-40"
+ >
+ {saveLabel}
+ </button>
+ <button type="button" onClick={onCancel} className="text-[12px] text-[var(--color-dim)] hover:text-[var(--color-text)]">
+ cancel
+ </button>
+ <span className="micro ml-auto">ctrl+enter</span>
+ </div>
+ </div>
+ );
+}
+
+const STATUS_TONE: Record<Note["status"], string> = {
+ open: "text-[var(--color-dirty)] border-[var(--color-dirty)]",
+ resolved: "text-[var(--color-good)] border-[var(--color-good)]",
+ wontfix: "text-[var(--color-dim)] border-[var(--color-line)]",
+};
+
+export function NoteCard({
+ note,
+ label,
+ orphaned,
+ selected,
+ editing,
+ busy,
+ onSelect,
+ onHover,
+ onEdit,
+ onCancelEdit,
+ write,
+}: {
+ note: Note;
+ label: string;
+ orphaned?: boolean;
+ selected?: boolean;
+ editing?: boolean;
+ busy?: boolean;
+ onSelect?: () => void;
+ onHover?: (on: boolean) => void;
+ onEdit?: () => void;
+ onCancelEdit?: () => void;
+ write: (op: NoteOp) => Promise<unknown>;
+}) {
+ const [replying, setReplying] = useState(false);
+ const mine = note.author === "operator";
+ return (
+ <li
+ data-note-card={note.id}
+ data-status={note.status}
+ data-orphaned={orphaned ? "true" : undefined}
+ aria-current={selected ? "true" : undefined}
+ onMouseEnter={() => onHover?.(true)}
+ onMouseLeave={() => onHover?.(false)}
+ onClick={onSelect}
+ className={`cursor-default space-y-1 rounded border bg-[var(--color-panel)] p-2 text-[12px] ${selected ? "border-[var(--color-sel)]" : "border-[var(--color-line)]"}`}
+ >
+ <div className="flex items-center gap-1.5">
+ <span className={`rounded border px-1 font-mono text-[10px] uppercase ${STATUS_TONE[note.status]}`}>{note.status}</span>
+ {note.author === "agent" && <span className="rounded border border-[var(--color-meter)] px-1 font-mono text-[10px] text-[var(--color-meter)]">agent</span>}
+ {orphaned && (
+ <span title="the quoted text is no longer in the article" className="rounded border border-[var(--color-bad)] px-1 font-mono text-[10px] text-[var(--color-bad)]">
+ orphaned
+ </span>
+ )}
+ <span className="num ml-auto text-[10px] text-[var(--color-dim)]" title={note.updatedAt}>
+ {fmtAgo(Date.parse(note.updatedAt))}
+ </span>
+ </div>
+ <div className="truncate text-[11px] text-[var(--color-dim)]" title={label}>
+ {label}
+ </div>
+ {editing ? (
+ <Composer
+ label="edit note"
+ initial={note.text}
+ busy={busy}
+ saveLabel="save"
+ onSave={async (text) => {
+ await write({ op: "edit", id: note.id, text });
+ onCancelEdit?.();
+ }}
+ onCancel={() => onCancelEdit?.()}
+ />
+ ) : (
+ <p data-note-text className="whitespace-pre-wrap text-[var(--color-text)]">
+ {note.text}
+ </p>
+ )}
+ {note.replies.length > 0 && (
+ <ul className="space-y-1 border-l border-[var(--color-line)] pl-2">
+ {note.replies.map((r, i) => (
+ <li key={`${r.at}-${i}`} data-reply={r.author} className="text-[12px]">
+ <span className={`mr-1 font-mono text-[10px] ${r.author === "agent" ? "text-[var(--color-meter)]" : "text-[var(--color-dim)]"}`}>{r.author}</span>
+ <span className="whitespace-pre-wrap text-[var(--color-text)]">{r.text}</span>
+ </li>
+ ))}
+ </ul>
+ )}
+ {replying ? (
+ <Composer
+ label="reply"
+ busy={busy}
+ saveLabel="reply"
+ onSave={async (text) => {
+ await write({ op: "reply", id: note.id, text });
+ setReplying(false);
+ }}
+ onCancel={() => setReplying(false)}
+ />
+ ) : (
+ <div className="flex flex-wrap gap-2 pt-0.5 text-[11px]">
+ {note.status === "open" ? (
+ <>
+ <Act onClick={() => write({ op: "status", id: note.id, status: "resolved" })} disabled={busy}>
+ resolve
+ </Act>
+ <Act onClick={() => write({ op: "status", id: note.id, status: "wontfix" })} disabled={busy}>
+ won’t fix
+ </Act>
+ </>
+ ) : (
+ <Act onClick={() => write({ op: "status", id: note.id, status: "open" })} disabled={busy}>
+ reopen
+ </Act>
+ )}
+ <Act onClick={() => setReplying(true)} disabled={busy}>
+ reply
+ </Act>
+ {mine && onEdit && (
+ <Act onClick={onEdit} disabled={busy}>
+ edit
+ </Act>
+ )}
+ <Act
+ onClick={() => {
+ if (window.confirm("Delete this note?")) void write({ op: "delete", id: note.id });
+ }}
+ disabled={busy}
+ >
+ delete
+ </Act>
+ </div>
+ )}
+ </li>
+ );
+}
+
+function Act({ onClick, disabled, children }: { onClick: () => void; disabled?: boolean; children: React.ReactNode }) {
+ return (
+ <button
+ type="button"
+ onClick={(e) => {
+ e.stopPropagation();
+ onClick();
+ }}
+ disabled={disabled}
+ className="text-[var(--color-dim)] hover:text-[var(--color-sel)] disabled:opacity-40"
+ >
+ {children}
+ </button>
+ );
+}
diff --git a/umtool/components/articles/SiteChips.tsx b/umtool/components/articles/SiteChips.tsx
@@ -0,0 +1,14 @@
+import { badgeVariants } from "@/components/ui/badge";
+
+// A site's audience, listing and search, as three short chips.
+export default function SiteChips({ site }: { site: { private: boolean; listed: boolean; search: boolean } }) {
+ return (
+ <span className="inline-flex gap-1">
+ <span data-audience={site.private ? "private" : "public"} className={badgeVariants({ variant: site.private ? "meter" : "neutral", size: "sm" })}>
+ {site.private ? "private" : "public"}
+ </span>
+ <span className={badgeVariants({ variant: "info", size: "sm" })}>{site.listed ? "listed" : "unlisted"}</span>
+ <span className={badgeVariants({ variant: "info", size: "sm" })}>{site.search ? "search" : "cited only"}</span>
+ </span>
+ );
+}
diff --git a/umtool/components/articles/WorkspacePanel.tsx b/umtool/components/articles/WorkspacePanel.tsx
@@ -0,0 +1,110 @@
+import Link from "next/link";
+import { Markdown } from "@/lib/markdown";
+import { fmtAgo, fmtBytes } from "@/lib/format";
+import type { OpenedFile, WorkspaceListing } from "@/lib/articles/files";
+
+// The files an article was written from, and the one that is open. Zero JS:
+// each file is a link that sets `?ws=&rel=` on the page it sits on; markdown
+// renders, a draft's JSON is pretty-printed with each top-level key folded,
+// and HTML opens in a sandboxed iframe (no scripts, no same origin).
+
+export default function WorkspacePanel({
+ workspaces,
+ opened,
+ hrefFor,
+ highlight,
+}: {
+ workspaces: WorkspaceListing[];
+ opened: OpenedFile | null;
+ hrefFor: (ws: string, rel: string) => string;
+ /** A file to mark (the article's own draft), as `<ws>/<rel>`. */
+ highlight?: string | null;
+}) {
+ if (workspaces.length === 0) return <p className="text-[12px] text-[var(--color-dim)]">no workspace found</p>;
+ return (
+ <div className="grid gap-4 min-[1100px]:grid-cols-[minmax(240px,320px)_1fr]" data-testid="workspace-panel">
+ <div className="space-y-3">
+ {workspaces.map((w) => (
+ <div key={w.name} data-workspace={w.name}>
+ <div className="mb-1 font-mono text-[11px] text-[var(--color-dim)]">{w.dir}</div>
+ <ul className="space-y-0.5">
+ {w.files.map((f) => {
+ const on = opened?.ws === w.name && opened.rel === f.rel;
+ const mine = highlight === `${w.name}/${f.rel}`;
+ return (
+ <li key={f.rel} className="flex items-baseline gap-2 text-[12px]">
+ <Link
+ href={hrefFor(w.name, f.rel)}
+ aria-current={on ? "true" : undefined}
+ data-file={f.rel}
+ className={`truncate font-mono ${on ? "text-[var(--color-sel)]" : mine ? "text-[var(--color-text)]" : "text-[var(--color-dim)] hover:text-[var(--color-text)]"}`}
+ >
+ {f.rel}
+ </Link>
+ {mine && <span className="micro">this article</span>}
+ <span className="num ml-auto shrink-0 text-[10px] text-[var(--color-dim)]">
+ {fmtBytes(f.bytes)} · {fmtAgo(f.mtimeMs)}
+ </span>
+ </li>
+ );
+ })}
+ </ul>
+ </div>
+ ))}
+ </div>
+ <div className="min-w-0">{opened ? <Opened file={opened} /> : <p className="text-[12px] text-[var(--color-dim)]">pick a file</p>}</div>
+ </div>
+ );
+}
+
+function Opened({ file }: { file: OpenedFile }) {
+ const head = <div className="mb-2 font-mono text-[11px] text-[var(--color-dim)]">{file.rel}</div>;
+ if (file.kind === "error") {
+ return (
+ <div>
+ {head}
+ <p className="text-[12px] text-[var(--color-bad)]">{file.message}</p>
+ </div>
+ );
+ }
+ if (file.kind === "html") {
+ return (
+ <div>
+ {head}
+ <iframe title={file.rel} src={file.url} sandbox="" className="h-[70vh] w-full rounded border border-[var(--color-line)] bg-white" />
+ </div>
+ );
+ }
+ if (file.kind === "json") {
+ const v = file.value;
+ const entries = v && typeof v === "object" && !Array.isArray(v) ? Object.entries(v as Record<string, unknown>) : null;
+ return (
+ <div data-opened="json">
+ {head}
+ {entries ? (
+ <div className="space-y-1">
+ {entries.map(([k, val]) => (
+ <details key={k} open={typeof val !== "object" || val === null} className="rounded border border-[var(--color-line)] bg-[var(--color-panel)]">
+ <summary className="cursor-pointer px-2 py-1 font-mono text-[12px] text-[var(--color-text)]">
+ {k}
+ {Array.isArray(val) ? <span className="micro ml-2">{val.length} items</span> : null}
+ </summary>
+ <pre className="max-h-[50vh] overflow-auto whitespace-pre-wrap px-2 pb-2 font-mono text-[11px] text-[var(--color-dim)]">
+ {JSON.stringify(val, null, 2)}
+ </pre>
+ </details>
+ ))}
+ </div>
+ ) : (
+ <pre className="overflow-auto whitespace-pre-wrap font-mono text-[11px]">{JSON.stringify(v, null, 2)}</pre>
+ )}
+ </div>
+ );
+ }
+ return (
+ <div data-opened="md">
+ {head}
+ <Markdown text={file.text} />
+ </div>
+ );
+}
diff --git a/umtool/components/articles/anchorDom.ts b/umtool/components/articles/anchorDom.ts
@@ -0,0 +1,98 @@
+// The article's TEXT, as the notes see it, and the <mark>s that show them.
+//
+// A block (`[data-block]`: title, subtitle, summary, method, or a section) has
+// one plain text: its text nodes in document order, MINUS anything inside
+// `[data-anchor-skip]` -- a citation's superscript number, a "+ note" button --
+// so a quote never carries a "3" from a cite marker and re-anchors the same
+// way `umtool notes` reads the section from report.json. lib/annotations/
+// anchor.mjs locates a quote in that text; these turn offsets into DOM and back.
+//
+// The marks are DOM the app adds AFTER React has rendered the (memoised,
+// never re-rendered) article body, and removes before adding them again.
+
+export const SKIP = "[data-anchor-skip]";
+
+export type TextModel = { text: string; nodes: { node: Text; start: number }[] };
+
+export function blockText(root: Element): TextModel {
+ const walker = document.createTreeWalker(root, NodeFilter.SHOW_TEXT, {
+ acceptNode: (n) => {
+ const skip = n.parentElement?.closest(SKIP);
+ return skip && root.contains(skip) ? NodeFilter.FILTER_REJECT : NodeFilter.FILTER_ACCEPT;
+ },
+ });
+ const nodes: TextModel["nodes"] = [];
+ let text = "";
+ for (let n = walker.nextNode(); n; n = walker.nextNode()) {
+ const t = n as Text;
+ nodes.push({ node: t, start: text.length });
+ text += t.data;
+ }
+ return { text, nodes };
+}
+
+/** The text offset of a DOM point inside `root`. */
+export function offsetOf(root: Element, container: Node, offset: number): number {
+ const { nodes, text } = blockText(root);
+ const point = document.createRange();
+ point.setStart(container, offset);
+ for (const { node, start } of nodes) {
+ if (node === container) return start + Math.min(offset, node.data.length);
+ const r = document.createRange();
+ r.selectNodeContents(node);
+ if (r.compareBoundaryPoints(Range.START_TO_START, point) >= 0) return start;
+ }
+ return text.length;
+}
+
+export function unwrapMarks(root: Element) {
+ const parents = new Set<Node>();
+ for (const m of Array.from(root.querySelectorAll("mark[data-note]"))) {
+ const p = m.parentNode;
+ if (!p) continue;
+ while (m.firstChild) p.insertBefore(m.firstChild, m);
+ p.removeChild(m);
+ parents.add(p);
+ }
+ for (const p of parents) p.normalize();
+}
+
+/** Wrap [start, end) of a block's text in marks carrying `attrs`. Returns the marks. */
+export function wrapRange(root: Element, start: number, end: number, attrs: Record<string, string>, className: string): HTMLElement[] {
+ const { nodes } = blockText(root);
+ const out: HTMLElement[] = [];
+ for (const { node, start: ns } of nodes) {
+ const ne = ns + node.data.length;
+ if (ne <= start || ns >= end) continue;
+ const a = Math.max(start, ns) - ns;
+ const b = Math.min(end, ne) - ns;
+ if (b <= a) continue;
+ let t: Text = node;
+ if (a > 0) t = t.splitText(a);
+ if (b - a < t.data.length) t.splitText(b - a);
+ // Whitespace between two block elements is not worth a mark.
+ if (!t.data.trim()) continue;
+ const mark = document.createElement("mark");
+ for (const [k, v] of Object.entries(attrs)) mark.setAttribute(k, v);
+ mark.className = className;
+ t.parentNode!.insertBefore(mark, t);
+ mark.appendChild(t);
+ out.push(mark);
+ }
+ return out;
+}
+
+/** The block a selection lies in, and its offsets, or null (collapsed, or across blocks). */
+export function selectionIn(container: Element): { block: Element; start: number; end: number; rect: DOMRect } | null {
+ const sel = window.getSelection();
+ if (!sel || sel.rangeCount === 0 || sel.isCollapsed) return null;
+ const range = sel.getRangeAt(0);
+ const el = (n: Node) => (n.nodeType === Node.ELEMENT_NODE ? (n as Element) : n.parentElement);
+ const a = el(range.startContainer)?.closest("[data-block]");
+ const b = el(range.endContainer)?.closest("[data-block]");
+ if (!a || a !== b || !container.contains(a)) return null;
+ const start = offsetOf(a, range.startContainer, range.startOffset);
+ const end = offsetOf(a, range.endContainer, range.endOffset);
+ if (end <= start) return null;
+ return { block: a, start, end, rect: range.getBoundingClientRect() };
+}
diff --git a/umtool/components/notes/AnchoredNotes.tsx b/umtool/components/notes/AnchoredNotes.tsx
@@ -0,0 +1,74 @@
+"use client";
+
+import { useState } from "react";
+import { badgeVariants } from "@/components/ui/badge";
+import type { Anchor, Note } from "@/lib/annotations/types";
+import type { UseNotes } from "@/lib/annotations/useNotes";
+import { NoteComposer, NoteThread } from "./NoteThread";
+
+// The notes on one THING: a timeline row (`entry`), a take (`take`). A count
+// of the open ones, folded; unfolded, every note on it and a box for another.
+
+type Pin = { kind: "entry"; entry: string } | { kind: "take"; take: string };
+
+export const notesOn = (notes: Note[], pin: Pin) =>
+ notes.filter((n) => {
+ const a = n.anchor as Anchor;
+ if (pin.kind === "entry") {
+ return (a.kind === "entry" && a.entry === pin.entry) || (a.kind === "edit" && a.entry === pin.entry);
+ }
+ return (a.kind === "take" && a.take === pin.take) || (a.kind === "moment" && a.take === pin.take);
+ });
+
+export default function AnchoredNotes({
+ notes,
+ pin,
+ startOpen = false,
+ label = "notes",
+}: {
+ notes: UseNotes;
+ pin: Pin;
+ startOpen?: boolean;
+ label?: string;
+}) {
+ const [open, setOpen] = useState(startOpen);
+ const mine = notesOn(notes.notes, pin).filter((n) => n.anchor.kind !== "moment");
+ const openCount = mine.filter((n) => n.status === "open").length;
+ const key = pin.kind === "entry" ? pin.entry : pin.take;
+ return (
+ <div data-anchored-notes={key} data-open-notes={openCount} className="text-[11px]">
+ <button
+ type="button"
+ data-action="toggle-notes"
+ onClick={() => setOpen((o) => !o)}
+ className={openCount ? badgeVariants({ variant: "open", size: "sm" }) : "text-[var(--color-sel)] hover:underline"}
+ >
+ {openCount ? `${openCount} open ${openCount === 1 ? "note" : "notes"}` : mine.length ? `${label} (${mine.length})` : `+ ${label.replace(/s$/, "")}`}
+ </button>
+ {open && (
+ <div className="mt-1 space-y-1 rounded bg-[var(--color-panel-2)] p-2">
+ {mine.map((n) => (
+ <NoteThread
+ key={n.id}
+ note={n}
+ write={notes.write}
+ busy={notes.busy}
+ head={
+ n.anchor.kind === "edit" ? (
+ <span className={badgeVariants({ variant: "info", size: "sm" })}>edit · {n.anchor.field}</span>
+ ) : null
+ }
+ />
+ ))}
+ <NoteComposer
+ busy={notes.busy}
+ placeholder={pin.kind === "entry" ? `a note on ${key}` : `a note on this take`}
+ testId={`note-input-${key}`}
+ onAdd={(text) => void notes.write({ op: "add", text, anchor: pin })}
+ />
+ {notes.error && <div className="text-[var(--color-bad)]">{notes.error}</div>}
+ </div>
+ )}
+ </div>
+ );
+}
diff --git a/umtool/components/notes/GeneratedBanner.tsx b/umtool/components/notes/GeneratedBanner.tsx
@@ -0,0 +1,15 @@
+// One line, wherever a generated manifest can be edited: the project page,
+// the clip bench, the On-screen section. Edits are still allowed; each one
+// leaves an `edit` note for the agent that runs the generator
+// (lib/report/guard.ts).
+export default function GeneratedBanner({ generatedBy }: { generatedBy: string | null | undefined }) {
+ if (!generatedBy) return null;
+ return (
+ <p
+ data-testid="generated-banner"
+ className="rounded border border-[var(--color-dirty)] px-2 py-1 text-[11px] text-[var(--color-dirty)]"
+ >
+ Generated by <code className="font-mono">{generatedBy}</code>; a rebuild of manifests overwrites edits made here.
+ </p>
+ );
+}
diff --git a/umtool/components/notes/NoteThread.tsx b/umtool/components/notes/NoteThread.tsx
@@ -0,0 +1,136 @@
+"use client";
+
+import { useState } from "react";
+import { badgeVariants } from "@/components/ui/badge";
+import type { Note, NoteOp } from "@/lib/annotations/types";
+
+// One note: its text, who wrote it, its status, the replies, and what can be
+// done to it. Shared by the timed notes, the row notes and the take notes.
+
+const STATUS_TONE = { open: "open", resolved: "info", wontfix: "info" } as const;
+
+export function NoteThread({
+ note,
+ write,
+ busy,
+ head,
+}: {
+ note: Note;
+ write: (op: NoteOp) => Promise<unknown>;
+ busy: boolean;
+ /** What the note is on, drawn before its text (a timestamp, an entry id). */
+ head?: React.ReactNode;
+}) {
+ const [reply, setReply] = useState<string | null>(null);
+ const open = note.status === "open";
+ return (
+ <div
+ data-note-id={note.id}
+ data-note-status={note.status}
+ className={`space-y-0.5 text-[12px] ${open ? "" : "opacity-60"}`}
+ >
+ <div className="flex flex-wrap items-baseline gap-1.5">
+ {head}
+ {note.author === "agent" && <span className={badgeVariants({ variant: "meter", size: "sm" })}>agent</span>}
+ {!open && <span className={badgeVariants({ variant: STATUS_TONE[note.status], size: "sm" })}>{note.status}</span>}
+ <span className="flex-1 whitespace-pre-wrap text-[var(--color-text)]">{note.text}</span>
+ <span className="flex gap-1.5 text-[10px]">
+ {open ? (
+ <>
+ <button type="button" data-note-action="resolve" disabled={busy} onClick={() => void write({ op: "status", id: note.id, status: "resolved" })} className="text-[var(--color-good)] hover:underline disabled:opacity-40">
+ resolve
+ </button>
+ <button type="button" data-note-action="wontfix" disabled={busy} onClick={() => void write({ op: "status", id: note.id, status: "wontfix" })} className="text-[var(--color-dim)] hover:underline disabled:opacity-40">
+ won’t fix
+ </button>
+ </>
+ ) : (
+ <button type="button" data-note-action="reopen" disabled={busy} onClick={() => void write({ op: "status", id: note.id, status: "open" })} className="text-[var(--color-sel)] hover:underline disabled:opacity-40">
+ reopen
+ </button>
+ )}
+ <button type="button" data-note-action="reply" disabled={busy} onClick={() => setReply((r) => (r === null ? "" : null))} className="text-[var(--color-sel)] hover:underline disabled:opacity-40">
+ reply
+ </button>
+ {note.author === "operator" && (
+ <button type="button" data-note-action="delete" disabled={busy} onClick={() => void write({ op: "delete", id: note.id })} className="text-[var(--color-dim)] hover:text-[var(--color-bad)] disabled:opacity-40">
+ delete
+ </button>
+ )}
+ </span>
+ </div>
+ {note.replies.map((r, i) => (
+ <div key={i} data-note-reply={i} className="ml-4 flex flex-wrap items-baseline gap-1.5 text-[11px]">
+ {r.author === "agent" && <span className={badgeVariants({ variant: "meter", size: "sm" })}>agent</span>}
+ <span className="whitespace-pre-wrap text-[var(--color-dim)]">{r.text}</span>
+ </div>
+ ))}
+ {reply !== null && (
+ <input
+ autoFocus
+ value={reply}
+ data-note-reply-input=""
+ onChange={(e) => setReply(e.target.value)}
+ onKeyDown={(e) => {
+ if (e.key === "Enter" && reply.trim()) {
+ void write({ op: "reply", id: note.id, text: reply });
+ setReply(null);
+ }
+ if (e.key === "Escape") setReply(null);
+ }}
+ placeholder="reply — Enter to send"
+ className="ml-4 w-[calc(100%-1rem)] rounded border border-[var(--color-line)] bg-[var(--color-ink)] px-1.5 py-0.5 text-[12px] text-[var(--color-text)] outline-none focus:border-[var(--color-sel)]"
+ />
+ )}
+ </div>
+ );
+}
+
+/** A one-line composer: Enter adds, Escape cancels. */
+export function NoteComposer({
+ onAdd,
+ onCancel,
+ busy,
+ placeholder,
+ head,
+ testId,
+}: {
+ onAdd: (text: string) => void;
+ onCancel?: () => void;
+ busy: boolean;
+ placeholder: string;
+ head?: React.ReactNode;
+ testId?: string;
+}) {
+ const [text, setText] = useState("");
+ const add = () => {
+ if (!text.trim()) return;
+ onAdd(text);
+ setText("");
+ };
+ return (
+ <div className="flex flex-wrap items-center gap-2">
+ {head}
+ <input
+ autoFocus
+ value={text}
+ data-testid={testId}
+ onChange={(e) => setText(e.target.value)}
+ onKeyDown={(e) => {
+ if (e.key === "Enter") add();
+ if (e.key === "Escape") onCancel?.();
+ }}
+ placeholder={placeholder}
+ className="min-w-[14rem] flex-1 rounded border border-[var(--color-line)] bg-[var(--color-ink)] px-1.5 py-0.5 text-[12px] text-[var(--color-text)] outline-none focus:border-[var(--color-sel)]"
+ />
+ <button
+ type="button"
+ disabled={busy || !text.trim()}
+ onClick={add}
+ className="rounded border border-[var(--color-good)] px-2 py-0.5 text-[11px] text-[var(--color-good)] disabled:opacity-40"
+ >
+ add
+ </button>
+ </div>
+ );
+}
diff --git a/umtool/components/notes/NotesProvider.tsx b/umtool/components/notes/NotesProvider.tsx
@@ -0,0 +1,65 @@
+"use client";
+
+import { createContext, useContext, useEffect } from "react";
+import { notesQuery, useNotes, type NotesTarget, type UseNotes } from "@/lib/annotations/useNotes";
+
+// ONE handle on a notes.json per page.
+//
+// Every write carries the token its handle last read, so two handles on the
+// same file in one page -- the timeline's row notes and the final video's
+// timed notes, say -- would 409 each other on every other write. A page wraps
+// itself in <NotesProvider target>, and every notes component under it shares
+// that handle; a component outside any provider (or under one for another
+// file) opens its own.
+
+const Ctx = createContext<{ key: string; notes: UseNotes } | null>(null);
+
+/** Something on the page wrote to a project's notes server-side (an edit note): re-read them. */
+export const NOTES_CHANGED = "umtool:notes-changed";
+export function announceNotesChanged(project: string) {
+ window.dispatchEvent(new CustomEvent(NOTES_CHANGED, { detail: { project } }));
+}
+
+/** Re-read `notes` when the page says its project's notes changed under it. */
+export function useNotesRefresh(target: NotesTarget | null, notes: UseNotes) {
+ const project = target && "project" in target ? target.project : null;
+ const { reload } = notes;
+ useEffect(() => {
+ if (!project) return;
+ const on = (e: Event) => {
+ if ((e as CustomEvent).detail?.project === project) void reload();
+ };
+ window.addEventListener(NOTES_CHANGED, on);
+ return () => window.removeEventListener(NOTES_CHANGED, on);
+ }, [project, reload]);
+}
+
+export function NotesProvider({ target, children }: { target: NotesTarget; children: React.ReactNode }) {
+ const notes = useNotes(target);
+ useNotesRefresh(target, notes);
+ return <Ctx.Provider value={{ key: notesQuery(target), notes }}>{children}</Ctx.Provider>;
+}
+
+/**
+ * Share a handle the page already holds (the article reader's) with the notes
+ * components under it, so they write with ITS token rather than opening a
+ * second one on the same file.
+ */
+export function ShareNotes({ target, notes, children }: { target: NotesTarget; notes: UseNotes; children: React.ReactNode }) {
+ return <Ctx.Provider value={{ key: notesQuery(target), notes }}>{children}</Ctx.Provider>;
+}
+
+/** The page's shared handle for `target`, or a handle of this component's own. */
+export function useSharedNotes(target: NotesTarget | null): UseNotes {
+ const ctx = useContext(Ctx);
+ const shared = !!(ctx && target && ctx.key === notesQuery(target));
+ const own = useNotes(shared ? null : target);
+ useNotesRefresh(shared ? null : target, own);
+ return shared ? ctx!.notes : own;
+}
+
+/** True when a write's response says the server wrote notes too (an edit on a generated manifest). */
+export const wroteNotes = (j: unknown): boolean => {
+ const e = (j as { editNotes?: { added?: number; updated?: number; deleted?: number } | null })?.editNotes;
+ return !!e && (e.added ?? 0) + (e.updated ?? 0) + (e.deleted ?? 0) > 0;
+};
diff --git a/umtool/components/notes/TimedNotes.tsx b/umtool/components/notes/TimedNotes.tsx
@@ -0,0 +1,239 @@
+"use client";
+
+import { useCallback, useEffect, useState } from "react";
+import { badgeVariants } from "@/components/ui/badge";
+import type { Anchor, MomentResolved, Note } from "@/lib/annotations/types";
+import type { NotesTarget, UseNotes } from "@/lib/annotations/useNotes";
+import { useSharedNotes } from "./NotesProvider";
+import { NoteComposer, NoteThread } from "./NoteThread";
+
+// Notes at a SECOND of a rendered video, in the notes.json of whatever the
+// video belongs to (a report-video project, an article). The song tool's
+// MomentMarks keeps its own store; this is the same idea bound to
+// lib/annotations, so an agent reads it back with everything else.
+//
+// `n` (with the video focused) or "Mark" opens a note at the playhead. When
+// the video is a report project's build, the second is RESOLVED at write time
+// (/api/report/moment: the entry on screen, its title and quote, the source
+// second and its archive link) and stored in the note -- a rebuild moves
+// entries, and the agent must see what was on screen when the key was pressed.
+// A file with no build schedule keeps `t` alone.
+
+const ts = (t: number) => {
+ const m = Math.floor(t / 60);
+ const s = t - m * 60;
+ return `${m}:${s.toFixed(1).padStart(4, "0")}`;
+};
+
+type MomentAnchor = Extract<Anchor, { kind: "moment" }>;
+const isMomentOn = (file: string) => (n: Note): n is Note & { anchor: MomentAnchor } =>
+ n.anchor.kind === "moment" && n.anchor.file === file;
+
+export default function TimedNotes({
+ notes,
+ file,
+ video,
+ take,
+ resolveProject,
+ showErrors = true,
+}: {
+ notes: UseNotes;
+ /** The file's path as the anchor stores it: project-relative (`takes/<id>/preview.mp4`), or the article's `video.mp4`. */
+ file: string;
+ video: HTMLVideoElement | null;
+ /** The take this file is a preview of, stored on the anchor. */
+ take?: string;
+ /** Resolve each mark against this project's build schedule. */
+ resolveProject?: string | null;
+ /** Show the handle's errors here. Off when the handle is shared with a page that shows them itself (the article reader). */
+ showErrors?: boolean;
+}) {
+ const [duration, setDuration] = useState<number | null>(null);
+ const [pending, setPending] = useState<number | null>(null);
+ const marks = notes.notes.filter(isMomentOn(file)).sort((a, b) => a.anchor.t - b.anchor.t);
+
+ const mark = useCallback(() => {
+ if (!video) return;
+ setPending(Number(video.currentTime.toFixed(2)));
+ }, [video]);
+
+ useEffect(() => {
+ if (!video) return;
+ const dur = () => setDuration(Number.isFinite(video.duration) ? video.duration : null);
+ const key = (e: KeyboardEvent) => {
+ if ((e.key === "n" || e.key === "N") && !e.ctrlKey && !e.metaKey && !e.altKey) {
+ e.preventDefault();
+ e.stopPropagation();
+ mark();
+ }
+ };
+ dur();
+ video.addEventListener("loadedmetadata", dur);
+ video.addEventListener("durationchange", dur);
+ video.addEventListener("keydown", key);
+ return () => {
+ video.removeEventListener("loadedmetadata", dur);
+ video.removeEventListener("durationchange", dur);
+ video.removeEventListener("keydown", key);
+ };
+ }, [video, mark]);
+
+ const seek = (t: number) => {
+ if (!video) return;
+ video.currentTime = t;
+ video.focus();
+ };
+
+ const add = async (t: number, text: string) => {
+ let resolved: (MomentResolved & { entry?: string | null }) | null = null;
+ if (resolveProject) {
+ const q = new URLSearchParams({ project: resolveProject, file, t: String(t) });
+ if (duration) q.set("duration", String(duration));
+ resolved = await fetch(`/api/report/moment?${q}`, { cache: "no-store" })
+ .then((r) => (r.ok ? r.json() : null))
+ .catch(() => null);
+ }
+ const anchor: MomentAnchor = { kind: "moment", file, t };
+ if (take) anchor.take = take;
+ if (resolved?.entry) {
+ anchor.entry = resolved.entry;
+ const { entry: _e, ...rest } = resolved as Record<string, unknown>;
+ delete rest.schedule;
+ anchor.resolved = rest as MomentResolved;
+ }
+ await notes.write({ op: "add", text, anchor });
+ };
+
+ return (
+ <div data-timed-notes={file} data-marks={marks.length} className="space-y-1">
+ {/* the tick strip: where the marks are, under the player */}
+ {duration ? (
+ <div className="relative h-2 rounded bg-[var(--color-panel-2)]" data-tick-strip="">
+ {marks.map((n) => (
+ <button
+ key={n.id}
+ type="button"
+ title={`${ts(n.anchor.t)} — ${n.text}`}
+ onClick={() => seek(n.anchor.t)}
+ data-tick={n.anchor.t}
+ className={`absolute top-0 h-2 w-1 -translate-x-1/2 rounded ${n.status === "open" ? "bg-[var(--color-dirty)]" : "bg-[var(--color-dim)]"}`}
+ style={{ left: `${Math.min(100, (n.anchor.t / duration) * 100)}%` }}
+ />
+ ))}
+ </div>
+ ) : null}
+ <div className="flex flex-wrap items-center gap-2 text-[11px]">
+ <button
+ type="button"
+ data-action="mark"
+ disabled={!video}
+ onClick={mark}
+ className="rounded border border-[var(--color-sel)] px-2 py-0.5 text-[11px] text-[var(--color-sel)] disabled:opacity-40"
+ >
+ Mark <kbd>n</kbd>
+ </button>
+ {marks.length > 0 && (
+ <span className="text-[var(--color-dim)]">
+ {marks.filter((m) => m.status === "open").length} open of {marks.length}
+ </span>
+ )}
+ {showErrors && notes.error && <span className="text-[var(--color-bad)]">{notes.error}</span>}
+ </div>
+ {pending !== null && (
+ <div className="rounded bg-[var(--color-panel-2)] p-2">
+ <NoteComposer
+ busy={notes.busy}
+ placeholder="what is wrong here"
+ testId="timed-note-input"
+ head={<span className="num text-[11px] text-[var(--color-meter)]">{ts(pending)}</span>}
+ onCancel={() => setPending(null)}
+ onAdd={(text) => {
+ const t = pending;
+ setPending(null);
+ void add(t, text);
+ }}
+ />
+ </div>
+ )}
+ {marks.length > 0 && (
+ <ul className="space-y-1">
+ {marks.map((n) => {
+ const r = n.anchor.resolved;
+ return (
+ <li key={n.id} data-mark-at={n.anchor.t} data-mark-entry={n.anchor.entry ?? ""}>
+ <NoteThread
+ note={n}
+ write={notes.write}
+ busy={notes.busy}
+ head={
+ <>
+ <button
+ type="button"
+ onClick={() => seek(n.anchor.t)}
+ className="num text-[11px] text-[var(--color-sel)] underline"
+ title="seek here"
+ >
+ {ts(n.anchor.t)}
+ </button>
+ {n.anchor.entry && (
+ <span className="font-mono text-[11px] text-[var(--color-dim)]" title={r?.quote ?? undefined}>
+ {n.anchor.entry}
+ {r?.title ? ` · ${r.title}` : ""}
+ </span>
+ )}
+ {r?.approx && <span className={badgeVariants({ variant: "info", size: "sm" })}>approx</span>}
+ {r?.url && (
+ <a href={r.url} target="_blank" rel="noreferrer" className="text-[11px] text-[var(--color-sel)] hover:underline">
+ source
+ </a>
+ )}
+ </>
+ }
+ />
+ </li>
+ );
+ })}
+ </ul>
+ )}
+ </div>
+ );
+}
+
+/**
+ * A video with its timed notes under it: what the project's final cut, a
+ * take's preview or an article's report video mounts. `notes` defaults to the
+ * page's shared handle for `target` (NotesProvider), else one of its own.
+ */
+export function TimedVideo({
+ target,
+ file,
+ src,
+ take,
+ resolveProject,
+ notes: given,
+ poster,
+ testId,
+ showErrors,
+ className = "aspect-video w-full rounded border border-[var(--color-line)] bg-black",
+}: {
+ target: NotesTarget;
+ file: string;
+ src: string;
+ take?: string;
+ resolveProject?: string | null;
+ notes?: UseNotes;
+ poster?: string;
+ testId?: string;
+ showErrors?: boolean;
+ className?: string;
+}) {
+ const own = useSharedNotes(given ? null : target);
+ const notes = given ?? own;
+ const [el, setEl] = useState<HTMLVideoElement | null>(null);
+ return (
+ <div className="space-y-1">
+ <video ref={setEl} data-testid={testId} src={src} poster={poster} controls preload="metadata" playsInline className={className} />
+ <TimedNotes notes={notes} file={file} video={el} take={take} resolveProject={resolveProject} showErrors={showErrors} />
+ </div>
+ );
+}
diff --git a/umtool/components/projects/ClipBench.tsx b/umtool/components/projects/ClipBench.tsx
@@ -1,5 +1,8 @@
"use client";
+import AnchoredNotes from "@/components/notes/AnchoredNotes";
+import GeneratedBanner from "@/components/notes/GeneratedBanner";
+import { useSharedNotes, wroteNotes } from "@/components/notes/NotesProvider";
import { useCallback, useEffect, useRef, useState } from "react";
import Link from "next/link";
import { useRouter } from "next/navigation";
@@ -256,6 +259,8 @@ export type ClipBenchData = {
siblings: Sibling[];
/** Is the cut built with the on-screen panel (`render.chrome`)? */
deckOn: boolean;
+ /** The manifest's `generatedBy`: edits here are noted for the agent that runs it. */
+ generatedBy?: string | null;
};
const hms = (t: number) => {
@@ -377,6 +382,9 @@ type Session = {
};
export default function ClipBench({ data }: { data: ClipBenchData }) {
+ // The project's notes: this clip's row notes, and the edit notes a save on a
+ // generated manifest leaves (the save says so, and they are re-read).
+ const notes = useSharedNotes({ project: data.project });
const [clip, setClip] = useState<Clip>(data.clip);
const [windows, setWindows] = useState<Win[]>(data.windows);
const [cues, setCues] = useState<Cue[]>(data.cues);
@@ -1084,6 +1092,7 @@ export default function ClipBench({ data }: { data: ClipBenchData }) {
});
const j = (await r.json()) as Record<string, unknown>;
setBusy(null);
+ if (wroteNotes(j)) void notes.reload();
if (!r.ok) {
setNote(
j.stale
@@ -1134,7 +1143,7 @@ export default function ClipBench({ data }: { data: ClipBenchData }) {
void refresh();
return true;
},
- [data.project, clip],
+ [data.project, clip, notes.reload],
);
/**
@@ -1484,7 +1493,7 @@ export default function ClipBench({ data }: { data: ClipBenchData }) {
// Returned, not just stored: "did anything actually arrive" is a question
// the caller has to answer before it claims a fetch worked.
return j;
- }, [data.project, clip.id]);
+ }, [data.project, clip.id, notes.reload]);
// ---- fetching more --------------------------------------------------------
//
@@ -1684,6 +1693,7 @@ export default function ClipBench({ data }: { data: ClipBenchData }) {
}
setClip((prev) => fromEntry(prev, j.entry ?? {}));
token.current = String(j.token ?? "");
+ if (wroteNotes(j)) void notes.reload();
setCutScore(j.score ?? null);
setNote(`cut to the quote (match ${(j.score ?? 0).toFixed(2)})`);
}, [data.project, clip.id]);
@@ -1849,6 +1859,7 @@ export default function ClipBench({ data }: { data: ClipBenchData }) {
data-verdict={verdict}
className="flex flex-col gap-2 lg:h-full lg:min-h-0 lg:overflow-hidden"
>
+ <GeneratedBanner generatedBy={data.generatedBy} />
{/* ---- where you are, where you can go, and what will be burned in ---- */}
<div className="flex flex-wrap items-center gap-x-3 gap-y-1 text-[11px]">
<span className="num font-mono text-[var(--color-text)]">
@@ -2942,6 +2953,9 @@ export default function ClipBench({ data }: { data: ClipBenchData }) {
})}
</div>
+ {/* ---- notes on this clip, for whoever acts on them ---- */}
+ <AnchoredNotes notes={notes} pin={{ kind: "entry", entry: clip.id }} startOpen={false} />
+
{/* ---- on-screen: what the panel under the footage says ----
Beside the header's fields because it is the same sitting: the
words you hear are the words a title should summarise. Both
diff --git a/umtool/components/projects/ClipBenchPage.tsx b/umtool/components/projects/ClipBenchPage.tsx
@@ -105,6 +105,7 @@ export default async function ClipBenchPage({
// Whether the cut is built with the on-screen panel. Decides whether the
// bench composes a preview of it; the fields are there either way.
deckOn: deckOn(manifest.render ?? {}),
+ generatedBy: typeof manifest.generatedBy === "string" ? manifest.generatedBy : null,
view,
windows: windows.map((w: { name: string; from: number; to: number }) => ({
name: w.name,
diff --git a/umtool/components/projects/OnscreenSection.tsx b/umtool/components/projects/OnscreenSection.tsx
@@ -1,5 +1,10 @@
"use client";
+import GeneratedBanner from "@/components/notes/GeneratedBanner";
+import { TimedVideo } from "@/components/notes/TimedNotes";
+import StructureEditors from "./StructureEditors";
+import { MANIFEST_CHANGED } from "./timelineApi";
+import { announceNotesChanged, wroteNotes } from "@/components/notes/NotesProvider";
import { useCallback, useEffect, useMemo, useRef, useState } from "react";
import { badgeVariants } from "@/components/ui/badge";
import { buttonVariants } from "@/components/ui/button";
@@ -772,8 +777,11 @@ export default function OnscreenSection({
project,
entries,
built,
+ generatedBy = null,
}: {
project: string;
+ /** The manifest's `generatedBy`: its edits are noted for the agent, and the section says so. */
+ generatedBy?: string | null;
/**
* Is the default cut's deliverable on disk? The server already knows, and
* asking the video route about a file that is not there is a 404 in the
@@ -821,6 +829,9 @@ export default function OnscreenSection({
const [job, setJob] = useState<JobView | null>(null);
const [jobError, setJobError] = useState<string | null>(null);
const [finalV, setFinalV] = useState<number | null | "none">(null);
+ // The built file's project-relative path (the video route's x-video): what a
+ // timed note on it is anchored to.
+ const [finalRel, setFinalRel] = useState<string | null>(null);
// The switch as pressed, while its write is in flight: the box follows the
// hand at once and falls back to the manifest's answer if the write fails.
const [switching, setSwitching] = useState<boolean | null>(null);
@@ -933,6 +944,7 @@ export default function OnscreenSection({
const loadFinal = useCallback(async () => {
const r = await fetch(`/api/report/video?${q}&kind=final`, { method: "HEAD", cache: "no-store" }).catch(() => null);
const m = r?.ok ? r.headers.get("x-video-mtime") : null;
+ setFinalRel(r?.ok ? r.headers.get("x-video") : null);
setFinalV(m ? Number(m) : "none");
}, [q]);
@@ -959,6 +971,28 @@ export default function OnscreenSection({
// `built` and `variant` are read once per load; loadFinal already follows the variant.
}, [loadChrome, loadRows, loadFinal, recompose]);
+ // A structural write elsewhere on the page (the timeline, the teaser and
+ // posts editors) changed the manifest: re-read the tokens and the rows,
+ // keeping every unsaved edit here, or the next save would 409.
+ const formDirtyRef = useRef(formDirty);
+ formDirtyRef.current = formDirty;
+ useEffect(() => {
+ const on = (e: Event) => {
+ if ((e as CustomEvent).detail?.project !== project) return;
+ void (async () => {
+ const r = await fetch(`/api/report/chrome?project=${encodeURIComponent(project)}`, { cache: "no-store" });
+ const j = (await r.json().catch(() => null)) as (ChromeDoc & { error?: string }) | null;
+ if (r.ok && j) {
+ token.current = j.token;
+ if (!formDirtyRef.current) setDoc(j);
+ }
+ await loadRows(true);
+ })();
+ };
+ window.addEventListener(MANIFEST_CHANGED, on);
+ return () => window.removeEventListener(MANIFEST_CHANGED, on);
+ }, [project, loadRows]);
+
// ---- writing ------------------------------------------------------------
const queued = useCallback(<T,>(fn: () => Promise<T>): Promise<T> => {
const run = saving.current.catch(() => null).then(fn);
@@ -978,6 +1012,7 @@ export default function OnscreenSection({
body: JSON.stringify({ project, chrome, token: token.current }),
});
const j = (await r.json()) as Record<string, unknown>;
+ if (wroteNotes(j)) announceNotesChanged(project);
setBusy(null);
if (!r.ok) {
if (j.stale) {
@@ -1072,6 +1107,7 @@ export default function OnscreenSection({
body: JSON.stringify({ project, onscreen: draftMap(), token: token.current }),
});
const j = (await r.json()) as Record<string, unknown>;
+ if (wroteNotes(j)) announceNotesChanged(project);
setBusy(null);
if (!r.ok) {
// The drafts stay exactly as typed. A refused batch is a typo or a
@@ -1114,6 +1150,7 @@ export default function OnscreenSection({
body: JSON.stringify({ project, posts: postsDraftMap(), token: token.current }),
});
const j = (await r.json()) as Record<string, unknown>;
+ if (wroteNotes(j)) announceNotesChanged(project);
setBusy(null);
if (!r.ok) {
if (j.stale) {
@@ -1626,6 +1663,7 @@ export default function OnscreenSection({
data-onscreen={on ? "on" : "off"}
className="space-y-3 rounded border border-[var(--color-line)] bg-[var(--color-panel)] p-3"
>
+ <GeneratedBanner generatedBy={generatedBy} />
{/* ---- the switch ---- */}
<div className="flex flex-wrap items-center gap-x-3 gap-y-1">
<h2 className="micro">on-screen</h2>
@@ -2050,12 +2088,12 @@ export default function OnscreenSection({
<div className="space-y-1">
<span className="micro">the built video</span>
{typeof finalV === "number" ? (
- <video
- data-testid="onscreen-final-video"
+ <TimedVideo
+ target={{ project }}
+ testId="onscreen-final-video"
+ file={finalRel ?? `out/${variant || "sourced"}/final.mp4`}
src={`/api/report/video?${q}&kind=final&v=${finalV}`}
- controls
- preload="metadata"
- className="aspect-video w-full rounded border border-[var(--color-line)] bg-black"
+ resolveProject={project}
/>
) : (
<p data-testid="onscreen-no-final" className="text-[11px] text-[var(--color-dim)]">
@@ -2093,6 +2131,14 @@ export default function OnscreenSection({
<div className="mt-1.5">{postsBlock}</div>
</details>
)}
+
+ {/* ---- the lists: teasers, posts, fact-check labels ---- */}
+ <details data-testid="structure-folded">
+ <summary className="cursor-pointer text-[11px] text-[var(--color-dim)]">teasers, posts and fact-check labels</summary>
+ <div className="mt-1.5">
+ <StructureEditors project={project} />
+ </div>
+ </details>
</section>
);
}
diff --git a/umtool/components/projects/ProjectView.tsx b/umtool/components/projects/ProjectView.tsx
@@ -6,6 +6,7 @@ import ReportProject from "./ReportProject";
import ClipBenchPage from "./ClipBenchPage";
import ClaimBenchPage from "./ClaimBenchPage";
import SweepProject from "./SweepProject";
+import TakesPage from "./TakesPage";
// ---------------------------------------------------------------------------
// The one place a kind id is matched against a component.
@@ -49,6 +50,9 @@ export default async function ProjectView({
if (rest.length === 2 && rest[0] === "claim") {
return <ClaimBenchPage project={project} claimId={rest[1]} />;
}
+ // `/browse/<project>/takes` -- alternative renders of one part of the
+ // cut, side by side, each judged like / maybe / no.
+ if (rest.length === 1 && rest[0] === "takes") return <TakesPage project={project} />;
return notFound();
}
case "sweep-report": {
diff --git a/umtool/components/projects/ReportProject.tsx b/umtool/components/projects/ReportProject.tsx
@@ -17,12 +17,16 @@ import {
import { EXPORT_FORMATS, exportableVariants } from "@/lib/report/export.mjs";
import { diffManifests, formatChange } from "@/lib/report/manifest-diff.mjs";
import { listSnapshots, readSnapshot } from "@/lib/report/snapshots.mjs";
+import { listTakes } from "@/lib/report/takes.mjs";
import DeliverSection from "./DeliverSection";
import FetchUnfetchedButton from "./FetchUnfetchedButton";
import OnscreenSection from "./OnscreenSection";
import ReportBuildChain from "./ReportBuildChain";
import SnapshotButton from "./SnapshotButton";
import TagCitedButton from "./TagCitedButton";
+import { RowControls, RowNotes, TimelineList } from "./TimelineEditor";
+import GeneratedBanner from "@/components/notes/GeneratedBanner";
+import { NotesProvider } from "@/components/notes/NotesProvider";
import { decisionsForProject } from "@/lib/projects";
import { badgeVariants, type BadgeVariants } from "@/components/ui/badge";
import { buttonVariants } from "@/components/ui/button";
@@ -127,6 +131,7 @@ export default async function ReportProject({
const noteFiles = names.filter((n) => /\.md$/i.test(n) && !/^readme\.md$/i.test(n) && n !== "sweep-report.md").sort();
const noteTexts = await Promise.all(noteFiles.map((n) => readFile(path.join(project.dir, n), "utf8").catch(() => "")));
const variants = await exportableVariants(project.dir);
+ const takes = await listTakes(project.dir);
// Provenance splits by SHAPE: a scalar or a URL stays in the table; a long
// string, or one with a newline, is prose the author wrote and reads as a
@@ -226,6 +231,17 @@ export default async function ReportProject({
</span>
</div>
)}
+ {/* Alternative renders of part of the cut, made elsewhere and judged
+ on their own page. Only when there is a takes/ to look at. */}
+ {takes.exists && (
+ <Link
+ href={`/browse/${project.id}/takes`}
+ data-takes-link={takes.takes.length}
+ className={buttonVariants({ size: "lg" })}
+ >
+ Takes — {takes.takes.length}
+ </Link>
+ )}
{build.built && (
<div className="text-right text-[11px] text-[var(--color-dim)]">
<div className="font-mono text-[var(--color-text)]">{build.slug}.mp4</div>
@@ -244,7 +260,9 @@ export default async function ReportProject({
)}
</div>
+ <NotesProvider target={{ project: project.id }}>
<main className="deck-main flex-1 space-y-4 p-4">
+ <GeneratedBanner generatedBy={typeof m.generatedBy === "string" ? m.generatedBy : null} />
{/* --- the author's own account of the cut, first ------------------ */}
{readme && (
<section data-readme="" className="rounded border border-[var(--color-line)] bg-[var(--color-panel)] px-3 py-2">
@@ -417,6 +435,7 @@ export default async function ReportProject({
because the panel is part of the video that gets delivered. */}
<OnscreenSection
project={project.id}
+ generatedBy={typeof m.generatedBy === "string" ? m.generatedBy : null}
entries={entries.map((e) => ({ id: e.id, kind: e.kind, segment: !!e.segment }))}
built={!!build.built}
/>
@@ -446,8 +465,8 @@ export default async function ReportProject({
/>
</div>
)}
- <ul className="space-y-1">
- {entries.map((e) => {
+ <TimelineList project={project.id} count={entries.length}>
+ {entries.map((e, at) => {
// Anything that is not a CLIP renders generically. The timeline's
// vocabulary is open -- one real manifest carries `scroll` and
// `chart` beside its cards -- and a page that only knows two words
@@ -455,17 +474,23 @@ export default async function ReportProject({
if (e.kind !== "clip") {
return (
<li
- key={e.id}
+ key={`${e.id}@${at}`}
data-entry={e.id}
data-kind={e.kind}
- className="flex flex-wrap items-baseline gap-2 rounded border border-dashed border-[var(--color-line)] px-3 py-1.5 text-[12px]"
+ data-at={at}
+ tabIndex={0}
+ className="flex flex-wrap items-baseline gap-2 rounded border border-dashed border-[var(--color-line)] px-3 py-1.5 text-[12px] outline-none focus:border-[var(--color-sel)]"
>
+ <RowControls id={e.id} at={at} />
<span className="font-mono text-[var(--color-dim)]">{e.id}</span>
<Pill>{e.style ? `${e.kind} · ${e.style}` : e.kind}</Pill>
<span className="text-[var(--color-text)]">
{e.heading ?? e.title ?? e.label ?? ""}
</span>
{e.seconds != null && <span className="num micro ml-auto">{e.seconds}s</span>}
+ <div className="basis-full">
+ <RowNotes id={e.id} project={project.id} />
+ </div>
</li>
);
}
@@ -474,15 +499,18 @@ export default async function ReportProject({
const midSentence = e.endsSentence === false && !e.lockEnd && !e.lock;
return (
<li
- key={e.id}
+ key={`${e.id}@${at}`}
data-entry={e.id}
data-kind="clip"
+ data-at={at}
+ tabIndex={0}
data-cached={e.cached ? "1" : "0"}
data-fetched={e.fetched ? "1" : "0"}
data-mid-sentence={midSentence ? "1" : "0"}
- className="rounded border border-[var(--color-line)] bg-[var(--color-panel)] px-3 py-1.5"
+ className="rounded border border-[var(--color-line)] bg-[var(--color-panel)] px-3 py-1.5 outline-none focus:border-[var(--color-sel)]"
>
<div className="flex flex-wrap items-baseline gap-2 text-[12px]">
+ <RowControls id={e.id} at={at} />
<span className="font-mono text-[var(--color-sel)]">{e.id}</span>
<span className="num font-mono text-[11px] text-[var(--color-dim)]">
{e.video} {hms(e.start)}–{hms(e.end)} ({(e.end - e.start).toFixed(1)}s)
@@ -555,10 +583,11 @@ export default async function ReportProject({
set <code className="font-mono">lock</code> if this window is deliberate
</p>
)}
+ <RowNotes id={e.id} project={project.id} />
</li>
);
})}
- </ul>
+ </TimelineList>
<div className="mt-2">
<Link
href={`/browse/${project.id}${showAll ? "" : "?all=1"}`}
@@ -758,6 +787,7 @@ export default async function ReportProject({
</section>
)}
</main>
+ </NotesProvider>
</div>
);
}
diff --git a/umtool/components/projects/StructureEditors.tsx b/umtool/components/projects/StructureEditors.tsx
@@ -0,0 +1,286 @@
+"use client";
+
+import { useRouter } from "next/navigation";
+import { useCallback, useEffect, useRef, useState } from "react";
+import { announceNotesChanged, wroteNotes } from "@/components/notes/NotesProvider";
+import { buttonVariants } from "@/components/ui/button";
+import { MANIFEST_CHANGED, announceManifestChanged, refusalOf, timelineOp } from "./timelineApi";
+
+// The parts of the cut that are lists, edited in place: each teaser's lines
+// and timing, the posts, and the fact-check's labels and colours. Every save
+// is one structural write (POST /api/report/timeline): checked by the build's
+// own validators, snapshotted first, and -- on a generated manifest -- noted
+// for the agent.
+
+type Teaser = { at: number; id: string; lines: unknown[]; beat: number | null; dip: { fade: number; black: number } | null; tail: string | null; tailWait: number | null };
+type Post = Record<string, unknown> & { id: string };
+type Verdict = { label: string; color: string };
+type Doc = {
+ teasers: Teaser[];
+ posts: Post[];
+ deckOn: boolean;
+ factcheck: { verdicts?: Record<string, Partial<Verdict>>; stamp?: Record<string, unknown>; tally?: Record<string, unknown> } | null;
+ factcheckResolved: { verdicts: Record<string, Verdict>; stamp: { seconds: number; position: string }; tally: { show: boolean; position: string } };
+ verdicts: string[];
+ token: string | null;
+};
+
+const input =
+ "rounded border border-[var(--color-line)] bg-[var(--color-ink)] px-1.5 py-0.5 text-[12px] text-[var(--color-text)] outline-none focus:border-[var(--color-sel)]";
+const num = (v: string) => (v.trim() === "" ? null : Number(v));
+
+export default function StructureEditors({ project }: { project: string }) {
+ const router = useRouter();
+ const [doc, setDoc] = useState<Doc | null>(null);
+ const [error, setError] = useState<string | null>(null);
+ const [busy, setBusy] = useState(false);
+
+ const load = useCallback(async () => {
+ const r = await fetch(`/api/report/timeline?project=${encodeURIComponent(project)}`, { cache: "no-store" });
+ const j = await r.json();
+ if (r.ok) setDoc(j as Doc);
+ else setError(String(j.error ?? r.status));
+ }, [project]);
+
+ useEffect(() => {
+ void load();
+ const on = (e: Event) => {
+ if ((e as CustomEvent).detail?.project === project) void load();
+ };
+ window.addEventListener(MANIFEST_CHANGED, on);
+ return () => window.removeEventListener(MANIFEST_CHANGED, on);
+ }, [load, project]);
+
+ const run = async (op: string, args: Record<string, unknown>) => {
+ setBusy(true);
+ setError(null);
+ const res = await timelineOp(project, doc?.token ?? null, op, args);
+ setBusy(false);
+ if (!res.ok) {
+ setError(refusalOf(res));
+ if (res.json.stale) await load();
+ return false;
+ }
+ // Adopt the new token now: the reload the announcement starts may land
+ // after the next save is pressed.
+ if (typeof res.json.token === "string") setDoc((d) => (d ? { ...d, token: res.json.token as string } : d));
+ announceManifestChanged(project);
+ if (wroteNotes(res.json)) announceNotesChanged(project);
+ router.refresh();
+ return true;
+ };
+
+ if (!doc) return error ? <p className="text-[11px] text-[var(--color-bad)]">{error}</p> : null;
+ return (
+ <div data-testid="structure-editors" className="space-y-3">
+ {error && (
+ <p data-testid="structure-error" className="text-[11px] text-[var(--color-bad)]">
+ {error}
+ </p>
+ )}
+ {doc.teasers.map((t) => (
+ <TeaserEditor key={`${t.id}@${t.at}`} teaser={t} busy={busy} save={(patch) => run("teaser", { id: t.id, at: t.at, patch })} />
+ ))}
+ <PostsEditor posts={doc.posts} busy={busy} save={(post) => run("post", { post })} remove={(id) => run("post-remove", { id })} />
+ {doc.deckOn && <FactcheckEditor doc={doc} busy={busy} save={(factcheck) => run("factcheck", { factcheck })} />}
+ </div>
+ );
+}
+
+function TeaserEditor({ teaser: t, busy, save }: { teaser: Teaser; busy: boolean; save: (patch: Record<string, unknown>) => Promise<boolean> }) {
+ // Plain lines edit as one per row; a teaser whose lines carry settings
+ // (break, role, replace…) edits as the JSON it is, so nothing is dropped.
+ const plain = t.lines.every((l) => typeof l === "string");
+ const linesText = () => (plain ? (t.lines as string[]).join("\n") : JSON.stringify(t.lines, null, 2));
+ const [lines, setLines] = useState(linesText);
+ const [beat, setBeat] = useState(t.beat == null ? "" : String(t.beat));
+ const [tail, setTail] = useState(t.tail ?? "");
+ const [tailWait, setTailWait] = useState(t.tailWait == null ? "" : String(t.tailWait));
+ const [fade, setFade] = useState(t.dip ? String(t.dip.fade) : "");
+ const [black, setBlack] = useState(t.dip ? String(t.dip.black) : "");
+ const [local, setLocal] = useState<string | null>(null);
+ // Unsaved edits win over a reload: the re-read every structural write sets
+ // off can land after somebody has started typing again.
+ const dirty = useRef(false);
+ const edit = <T,>(set: (v: T) => void) => (v: T) => {
+ dirty.current = true;
+ set(v);
+ };
+ useEffect(() => {
+ if (dirty.current) return;
+ setLines(linesText());
+ setBeat(t.beat == null ? "" : String(t.beat));
+ setTail(t.tail ?? "");
+ setTailWait(t.tailWait == null ? "" : String(t.tailWait));
+ setFade(t.dip ? String(t.dip.fade) : "");
+ setBlack(t.dip ? String(t.dip.black) : "");
+ // eslint-disable-next-line react-hooks/exhaustive-deps
+ }, [t]);
+
+ const submit = () => {
+ setLocal(null);
+ let parsed: unknown[];
+ try {
+ parsed = plain ? lines.split("\n").map((l) => l.trim()).filter(Boolean) : JSON.parse(lines);
+ } catch {
+ setLocal("the lines are not valid JSON");
+ return;
+ }
+ const dip = fade.trim() || black.trim() ? { fade: num(fade), black: num(black) } : null;
+ void save({ lines: parsed, beat: num(beat), tail: tail.trim() || null, tailWait: num(tailWait), dip }).then((ok) => {
+ if (ok) dirty.current = false;
+ });
+ };
+
+ return (
+ <div data-teaser-editor={t.id} className="space-y-1 rounded border border-[var(--color-line)] p-2">
+ <div className="micro">teaser · {t.id}</div>
+ <textarea data-testid={`teaser-lines-${t.id}`} rows={Math.min(8, Math.max(2, lines.split("\n").length))} value={lines} onChange={(e) => edit(setLines)(e.target.value)} className={`${input} w-full font-mono`} />
+ <div className="flex flex-wrap items-center gap-2 text-[11px] text-[var(--color-dim)]">
+ <label>beat <input data-testid={`teaser-beat-${t.id}`} value={beat} onChange={(e) => edit(setBeat)(e.target.value)} className={`${input} w-14`} /></label>
+ <label>tail <input value={tail} onChange={(e) => edit(setTail)(e.target.value)} className={`${input} w-16`} /></label>
+ <label>tail wait <input value={tailWait} onChange={(e) => edit(setTailWait)(e.target.value)} className={`${input} w-14`} /></label>
+ <label>dip fade <input value={fade} onChange={(e) => edit(setFade)(e.target.value)} className={`${input} w-14`} /></label>
+ <label>black <input value={black} onChange={(e) => edit(setBlack)(e.target.value)} className={`${input} w-14`} /></label>
+ <button type="button" data-testid={`teaser-save-${t.id}`} disabled={busy} onClick={submit} className={buttonVariants({ variant: "primary", size: "sm" })}>
+ Save teaser
+ </button>
+ {local && <span className="text-[var(--color-bad)]">{local}</span>}
+ </div>
+ </div>
+ );
+}
+
+const POST_FIELDS = ["platform", "author", "handle", "date", "url", "shot", "flag"] as const;
+
+function PostsEditor({
+ posts,
+ busy,
+ save,
+ remove,
+}: {
+ posts: Post[];
+ busy: boolean;
+ save: (post: Post) => Promise<boolean>;
+ remove: (id: string) => Promise<boolean>;
+}) {
+ const [adding, setAdding] = useState(false);
+ return (
+ <div data-testid="posts-editor" className="space-y-1">
+ <div className="flex items-center gap-2">
+ <span className="micro">posts — {posts.length}</span>
+ <button type="button" data-testid="post-add" onClick={() => setAdding((a) => !a)} className="text-[11px] text-[var(--color-sel)] hover:underline">
+ + post
+ </button>
+ </div>
+ {adding && <PostRow post={{ id: "", platform: "x" }} fresh busy={busy} save={async (p) => (await save(p)) && (setAdding(false), true)} />}
+ {posts.map((p) => (
+ <PostRow key={p.id} post={p} busy={busy} save={save} remove={() => void remove(p.id)} />
+ ))}
+ </div>
+ );
+}
+
+function PostRow({ post, fresh = false, busy, save, remove }: { post: Post; fresh?: boolean; busy: boolean; save: (p: Post) => Promise<boolean>; remove?: () => void }) {
+ const [d, setD] = useState<Record<string, string>>(() => {
+ const out: Record<string, string> = { id: post.id, text: String(post.text ?? "") };
+ for (const k of POST_FIELDS) out[k] = post[k] == null ? "" : String(post[k]);
+ return out;
+ });
+ const set = (k: string, v: string) => setD((x) => ({ ...x, [k]: v }));
+ return (
+ <div data-post-row={post.id || "new"} className="space-y-1 rounded border border-[var(--color-line)] p-2 text-[11px]">
+ <div className="flex flex-wrap items-center gap-1.5 text-[var(--color-dim)]">
+ {fresh ? (
+ <label>id <input data-testid="post-id" value={d.id} onChange={(e) => set("id", e.target.value)} className={`${input} w-28 font-mono`} /></label>
+ ) : (
+ <span className="font-mono text-[var(--color-text)]">{post.id}</span>
+ )}
+ <select value={d.platform} onChange={(e) => set("platform", e.target.value)} className={`${input} text-[11px]`}>
+ {["x", "bluesky", "web"].map((p) => (
+ <option key={p}>{p}</option>
+ ))}
+ </select>
+ {(["author", "handle", "date", "url", "shot", "flag"] as const).map((k) => (
+ <label key={k}>
+ {k} <input data-testid={`post-${k}`} value={d[k]} onChange={(e) => set(k, e.target.value)} className={`${input} ${k === "url" ? "w-56" : "w-28"}`} />
+ </label>
+ ))}
+ </div>
+ <textarea data-testid="post-text" rows={2} value={d.text} onChange={(e) => set("text", e.target.value)} className={`${input} w-full`} />
+ <div className="flex gap-2">
+ <button
+ type="button"
+ data-testid="post-save"
+ disabled={busy}
+ onClick={() => {
+ const p: Post = { id: d.id.trim(), text: d.text };
+ for (const k of POST_FIELDS) p[k] = d[k];
+ void save(p);
+ }}
+ className={buttonVariants({ variant: "primary", size: "sm" })}
+ >
+ {fresh ? "Add post" : "Save post"}
+ </button>
+ {remove && (
+ <button type="button" data-testid="post-remove" disabled={busy} onClick={remove} className={buttonVariants({ variant: "destructive", size: "sm" })}>
+ Remove
+ </button>
+ )}
+ </div>
+ </div>
+ );
+}
+
+function FactcheckEditor({ doc, busy, save }: { doc: Doc; busy: boolean; save: (f: Record<string, unknown> | null) => Promise<boolean> }) {
+ const given = doc.factcheck ?? {};
+ const [v, setV] = useState<Record<string, { label: string; color: string }>>(() =>
+ Object.fromEntries(doc.verdicts.map((k) => [k, { label: String(given.verdicts?.[k]?.label ?? ""), color: String(given.verdicts?.[k]?.color ?? "") }])),
+ );
+ const [seconds, setSeconds] = useState(given.stamp?.seconds == null ? "" : String(given.stamp.seconds));
+ const submit = () => {
+ const verdicts: Record<string, Record<string, string>> = {};
+ for (const [k, o] of Object.entries(v)) {
+ const one: Record<string, string> = {};
+ if (o.label.trim()) one.label = o.label.trim();
+ if (o.color.trim()) one.color = o.color.trim();
+ if (Object.keys(one).length) verdicts[k] = one;
+ }
+ // Only what differs from the defaults is written, as the deck's settings are.
+ const out: Record<string, unknown> = { ...given };
+ if (Object.keys(verdicts).length) out.verdicts = verdicts;
+ else delete out.verdicts;
+ const stamp = { ...(given.stamp ?? {}) } as Record<string, unknown>;
+ if (seconds.trim()) stamp.seconds = Number(seconds);
+ else delete stamp.seconds;
+ if (Object.keys(stamp).length) out.stamp = stamp;
+ else delete out.stamp;
+ void save(Object.keys(out).length ? out : null);
+ };
+ return (
+ <div data-testid="factcheck-editor" className="space-y-1 rounded border border-[var(--color-line)] p-2 text-[11px]">
+ <div className="micro">fact-check labels</div>
+ <div className="grid grid-cols-[max-content_1fr_max-content] items-center gap-x-2 gap-y-1">
+ {doc.verdicts.map((k) => {
+ const def = doc.factcheckResolved.verdicts[k];
+ return (
+ <div key={k} className="contents">
+ <span className="font-mono text-[var(--color-dim)]">{k}</span>
+ <input data-testid={`fc-label-${k}`} placeholder={def?.label} value={v[k]?.label ?? ""} onChange={(e) => setV((x) => ({ ...x, [k]: { ...x[k], label: e.target.value } }))} className={input} />
+ <span className="flex items-center gap-1">
+ <input data-testid={`fc-color-${k}`} placeholder={def?.color} value={v[k]?.color ?? ""} onChange={(e) => setV((x) => ({ ...x, [k]: { ...x[k], color: e.target.value } }))} className={`${input} w-20 font-mono`} />
+ <span className="inline-block h-3 w-3 rounded" style={{ background: v[k]?.color || def?.color }} />
+ </span>
+ </div>
+ );
+ })}
+ </div>
+ <div className="flex items-center gap-2 text-[var(--color-dim)]">
+ <label>stamp seconds <input value={seconds} placeholder={String(doc.factcheckResolved.stamp.seconds)} onChange={(e) => setSeconds(e.target.value)} className={`${input} w-14`} /></label>
+ <button type="button" data-testid="fc-save" disabled={busy} onClick={submit} className={buttonVariants({ variant: "primary", size: "sm" })}>
+ Save labels
+ </button>
+ </div>
+ </div>
+ );
+}
diff --git a/umtool/components/projects/TakesBench.tsx b/umtool/components/projects/TakesBench.tsx
@@ -0,0 +1,359 @@
+"use client";
+
+import { useCallback, useEffect, useRef, useState } from "react";
+import AnchoredNotes from "@/components/notes/AnchoredNotes";
+import { NotesProvider, useSharedNotes } from "@/components/notes/NotesProvider";
+import TimedNotes from "@/components/notes/TimedNotes";
+import { badgeVariants, type BadgeVariants } from "@/components/ui/badge";
+import { buttonVariants } from "@/components/ui/button";
+import { fmtAgo } from "@/lib/format";
+
+// The takes of one project, a group at a time, every preview on screen at once.
+//
+// ONE AUDIO SOURCE. Every preview is muted except the one being listened to:
+// the "listen" toggle picks it, and so does pressing play (or unmuting) on a
+// single preview -- the last thing you reached for is what you hear. "Play all
+// from start" keeps whichever is being listened to, so a group can be watched
+// in sync against one soundtrack, or in silence.
+//
+// Buttons stay disabled until mount, as VerdictChip's do: a click before
+// hydration sends nothing, and a verdict lost that way looks like one never
+// given.
+
+export type Take = {
+ id: string;
+ group: string;
+ order: number;
+ label: string;
+ kind: "reference" | "similar" | "different";
+ summary: string;
+ changes: string[];
+ preview: string;
+ seconds: number | null;
+ builtAt: string | null;
+ previewSize: number | null;
+ previewMtimeMs: number | null;
+};
+
+export type Verdict = "like" | "maybe" | "no";
+export type TakeVerdict = { verdict: Verdict | null; note: string; at: string };
+
+const KIND_TONE: Record<Take["kind"], BadgeVariants["variant"]> = {
+ reference: "on",
+ similar: "info",
+ different: "meter",
+};
+
+// Literal class strings, as VerdictChip's are: Tailwind finds classes by
+// reading the source, so a class assembled from a variable is never generated.
+const TONES = {
+ good: {
+ on: "border-[var(--color-good)] bg-[color-mix(in_srgb,var(--color-good)_18%,transparent)] text-[var(--color-good)]",
+ off: "border-[var(--color-line)] text-[var(--color-dim)] hover:border-[var(--color-good)] hover:text-[var(--color-good)]",
+ },
+ dirty: {
+ on: "border-[var(--color-dirty)] bg-[color-mix(in_srgb,var(--color-dirty)_18%,transparent)] text-[var(--color-dirty)]",
+ off: "border-[var(--color-line)] text-[var(--color-dim)] hover:border-[var(--color-dirty)] hover:text-[var(--color-dirty)]",
+ },
+ bad: {
+ on: "border-[var(--color-bad)] bg-[color-mix(in_srgb,var(--color-bad)_18%,transparent)] text-[var(--color-bad)]",
+ off: "border-[var(--color-line)] text-[var(--color-dim)] hover:border-[var(--color-bad)] hover:text-[var(--color-bad)]",
+ },
+ sel: {
+ on: "border-[var(--color-sel)] bg-[color-mix(in_srgb,var(--color-sel)_18%,transparent)] text-[var(--color-sel)]",
+ off: "border-[var(--color-line)] text-[var(--color-dim)] hover:border-[var(--color-sel)] hover:text-[var(--color-sel)]",
+ },
+} as const;
+type Tone = keyof typeof TONES;
+const tone = (t: Tone, on: boolean) => (on ? TONES[t].on : TONES[t].off);
+
+const VERDICTS: { id: Verdict; label: string; tone: Tone }[] = [
+ { id: "like", label: "Like", tone: "good" },
+ { id: "maybe", label: "Maybe", tone: "dirty" },
+ { id: "no", label: "No", tone: "bad" },
+];
+
+const previewSrc = (project: string, t: Take) =>
+ `/api/report/takes/preview?${new URLSearchParams({
+ project,
+ take: t.id,
+ ...(t.previewMtimeMs != null ? { v: String(t.previewMtimeMs) } : {}),
+ })}`;
+
+export default function TakesBench({
+ project,
+ groups,
+ initial,
+}: {
+ project: string;
+ groups: { group: string; takes: Take[] }[];
+ initial: Record<string, TakeVerdict>;
+}) {
+ const [verdicts, setVerdicts] = useState(initial);
+ const [listen, setListen] = useState<string | null>(null);
+ const [ready, setReady] = useState(false);
+ const videos = useRef(new Map<string, HTMLVideoElement>());
+ const listenRef = useRef<string | null>(null);
+ // True while "play all" is starting its videos, so their play events do not
+ // each claim the audio.
+ const starting = useRef(false);
+
+ useEffect(() => setReady(true), []);
+ useEffect(() => {
+ listenRef.current = listen;
+ for (const [id, v] of videos.current) v.muted = id !== listen;
+ }, [listen]);
+
+ const els = (g: { takes: Take[] }) =>
+ g.takes.map((t) => videos.current.get(t.id)).filter((v): v is HTMLVideoElement => !!v);
+
+ const playAll = async (g: { takes: Take[] }) => {
+ starting.current = true;
+ const list = els(g);
+ for (const v of list) {
+ v.pause();
+ v.currentTime = 0;
+ }
+ await Promise.allSettled(list.map((v) => v.play()));
+ starting.current = false;
+ };
+
+ const save = async (take: string, patch: { verdict?: Verdict | null; note?: string }) => {
+ const r = await fetch("/api/report/takes/verdict", {
+ method: "POST",
+ headers: { "content-type": "application/json" },
+ body: JSON.stringify({ project, take, ...patch }),
+ cache: "no-store",
+ });
+ const j = (await r.json()) as { error?: string; verdicts?: Record<string, TakeVerdict> };
+ if (!r.ok || !j.verdicts) throw new Error(j.error ?? `HTTP ${r.status}`);
+ setVerdicts(j.verdicts);
+ };
+
+ return (
+ <NotesProvider target={{ project }}>
+ <div className="space-y-6">
+ {groups.map((g) => {
+ const tally = VERDICTS.map((v) => [v.id, g.takes.filter((t) => verdicts[t.id]?.verdict === v.id).length] as const)
+ .filter(([, n]) => n > 0)
+ .map(([id, n]) => `${n} ${id}`)
+ .join(" · ");
+ return (
+ <section key={g.group} data-group={g.group}>
+ <div className="mb-2 flex flex-wrap items-center gap-2">
+ <h2 className="micro">
+ {g.group} — {g.takes.length}
+ </h2>
+ {tally && <span className="num text-[11px] text-[var(--color-dim)]">{tally}</span>}
+ <span className="ml-auto flex gap-1.5">
+ <button
+ type="button"
+ data-action="play-all"
+ disabled={!ready}
+ className={buttonVariants({ variant: "primary", size: "sm" })}
+ onClick={() => void playAll(g)}
+ >
+ Play all from start
+ </button>
+ <button
+ type="button"
+ data-action="pause-all"
+ disabled={!ready}
+ className={buttonVariants({ size: "sm" })}
+ onClick={() => els(g).forEach((v) => v.pause())}
+ >
+ Pause all
+ </button>
+ </span>
+ </div>
+ <div className="grid grid-cols-1 gap-3 md:grid-cols-2 xl:grid-cols-3">
+ {g.takes.map((t) => (
+ <TakeCard
+ key={t.id}
+ project={project}
+ take={t}
+ src={previewSrc(project, t)}
+ verdict={verdicts[t.id] ?? null}
+ listening={listen === t.id}
+ ready={ready}
+ onListen={() => setListen((cur) => (cur === t.id ? null : t.id))}
+ videoRef={(el) => {
+ if (el) {
+ el.muted = t.id !== listenRef.current;
+ videos.current.set(t.id, el);
+ } else {
+ videos.current.delete(t.id);
+ }
+ }}
+ onPlay={() => {
+ if (!starting.current && listenRef.current !== t.id) setListen(t.id);
+ }}
+ onUnmute={() => {
+ if (listenRef.current !== t.id) setListen(t.id);
+ }}
+ save={(patch) => save(t.id, patch)}
+ />
+ ))}
+ </div>
+ </section>
+ );
+ })}
+ </div>
+ </NotesProvider>
+ );
+}
+
+function TakeCard({
+ project,
+ take: t,
+ src,
+ verdict,
+ listening,
+ ready,
+ onListen,
+ videoRef,
+ onPlay,
+ onUnmute,
+ save,
+}: {
+ project: string;
+ take: Take;
+ src: string;
+ verdict: TakeVerdict | null;
+ listening: boolean;
+ ready: boolean;
+ onListen: () => void;
+ videoRef: (el: HTMLVideoElement | null) => void;
+ onPlay: () => void;
+ onUnmute: () => void;
+ save: (patch: { verdict?: Verdict | null; note?: string }) => Promise<void>;
+}) {
+ const notes = useSharedNotes({ project });
+ // The element, for the timed notes; the parent's callback (audio routing)
+ // still gets it. Stable, so it runs once per mount, not once per render.
+ const [el, setEl] = useState<HTMLVideoElement | null>(null);
+ const parentRef = useRef(videoRef);
+ parentRef.current = videoRef;
+ const bindVideo = useCallback((node: HTMLVideoElement | null) => {
+ parentRef.current(node);
+ setEl(node);
+ }, []);
+ const [note, setNote] = useState(verdict?.note ?? "");
+ const [busy, setBusy] = useState(false);
+ const [error, setError] = useState<string | null>(null);
+ const saved = verdict?.note ?? "";
+ // Adopt a server value that moved underneath (another tab, a refresh).
+ useEffect(() => setNote(verdict?.note ?? ""), [verdict?.note]);
+
+ const run = async (patch: { verdict?: Verdict | null; note?: string }) => {
+ setBusy(true);
+ setError(null);
+ try {
+ await save(patch);
+ } catch (e) {
+ setError(e instanceof Error ? e.message : String(e));
+ } finally {
+ setBusy(false);
+ }
+ };
+ const commitNote = () => {
+ if (note.replace(/\s+$/, "") !== saved) void run({ note });
+ };
+ const current = verdict?.verdict ?? null;
+ const hue = VERDICTS.find((v) => v.id === current)?.tone;
+
+ return (
+ <article
+ data-take={t.id}
+ data-kind={t.kind}
+ data-verdict={current ?? ""}
+ className="flex flex-col overflow-hidden rounded border bg-[var(--color-panel)]"
+ style={{ borderColor: `var(--color-${hue ?? "line"})` }}
+ >
+ {t.previewSize != null ? (
+ <video
+ ref={bindVideo}
+ src={src}
+ controls
+ preload="metadata"
+ playsInline
+ onPlay={onPlay}
+ onVolumeChange={(e) => {
+ if (!e.currentTarget.muted) onUnmute();
+ }}
+ className="aspect-video w-full bg-black"
+ />
+ ) : (
+ <div data-no-preview="" className="flex aspect-video w-full items-center justify-center bg-black text-[11px] text-[var(--color-dim)]">
+ no preview yet
+ </div>
+ )}
+ <div className="flex flex-1 flex-col gap-1.5 px-3 py-2">
+ <div className="flex flex-wrap items-center gap-2">
+ <h3 className="text-[13px] text-[var(--color-text)]">{t.label}</h3>
+ <span className={badgeVariants({ variant: KIND_TONE[t.kind], size: "sm" })}>{t.kind}</span>
+ <span className="num text-[11px] text-[var(--color-dim)]" suppressHydrationWarning>
+ {t.seconds != null && `${t.seconds.toFixed(1)} s`}
+ {t.builtAt && !Number.isNaN(Date.parse(t.builtAt)) && ` · ${fmtAgo(Date.parse(t.builtAt))}`}
+ </span>
+ <button
+ type="button"
+ data-action="listen"
+ aria-pressed={listening}
+ disabled={!ready || t.previewSize == null}
+ className={`ml-auto rounded border px-2 py-0.5 text-[11px] ${tone("sel", listening)}`}
+ onClick={onListen}
+ >
+ {listening ? "listening" : "listen"}
+ </button>
+ </div>
+ {t.summary && <p className="text-[12px] text-[var(--color-text)]">{t.summary}</p>}
+ {t.changes.length > 0 && (
+ <ul className="list-disc pl-4 text-[11px] text-[var(--color-dim)]">
+ {t.changes.map((c, i) => (
+ <li key={i}>{c}</li>
+ ))}
+ </ul>
+ )}
+ <div className="mt-auto flex flex-wrap items-center gap-1.5 pt-1">
+ {VERDICTS.map((v) => (
+ <button
+ key={v.id}
+ type="button"
+ data-verdict-button={v.id}
+ aria-pressed={current === v.id}
+ disabled={!ready || busy}
+ className={`rounded border px-2.5 py-1 text-[12px] ${tone(v.tone, current === v.id)}`}
+ onClick={() => void run({ verdict: current === v.id ? null : v.id })}
+ >
+ {v.label}
+ </button>
+ ))}
+ <input
+ value={note}
+ onChange={(e) => setNote(e.target.value)}
+ onBlur={commitNote}
+ onKeyDown={(e) => {
+ if (e.key === "Enter") commitNote();
+ }}
+ disabled={!ready}
+ placeholder="note"
+ aria-label={`note on ${t.label}`}
+ data-take-note=""
+ maxLength={2000}
+ className="min-w-0 flex-1 rounded border border-[var(--color-line)] bg-[var(--color-ink)] px-1.5 py-1 text-[12px] text-[var(--color-text)] outline-none focus:border-[var(--color-sel)]"
+ />
+ </div>
+ {error && <span className="text-[11px] text-[var(--color-bad)]">{error}</span>}
+ {/* The conversation about this take: notes on the take as a whole, and
+ notes at a second of its preview, resolved to the entry on screen.
+ Both land in the project's notes.json, which the agent that made
+ the takes reads back (`umtool notes`) beside verdicts.json. */}
+ <AnchoredNotes notes={notes} pin={{ kind: "take", take: t.id }} label="take notes" />
+ {t.previewSize != null && (
+ <TimedNotes notes={notes} file={`takes/${t.id}/${t.preview}`} video={el} take={t.id} resolveProject={project} />
+ )}
+ </div>
+ </article>
+ );
+}
diff --git a/umtool/components/projects/TakesPage.tsx b/umtool/components/projects/TakesPage.tsx
@@ -0,0 +1,58 @@
+import Link from "next/link";
+import BrowseHeader from "@/components/BrowseHeader";
+import { listTakes, readVerdicts } from "@/lib/report/takes.mjs";
+import type { ProjectRef } from "@/lib/project-types";
+import TakesBench, { type Take, type TakeVerdict } from "./TakesBench";
+
+// `/browse/<project>/takes` -- alternative renders of one part of the cut,
+// watched side by side and judged like / maybe / no.
+//
+// The takes are made elsewhere (lib/report/takes.mjs says by whom and in what
+// shape); this page only reads them, and writes takes/verdicts.json.
+
+export default async function TakesPage({ project }: { project: ProjectRef }) {
+ const [listing, verdicts] = await Promise.all([listTakes(project.dir), readVerdicts(project.dir)]);
+ const takes = listing.takes as Take[];
+ const judged = takes.filter((t) => verdicts[t.id]?.verdict).length;
+
+ return (
+ <div className="flex h-full flex-col">
+ <BrowseHeader
+ crumbs={[
+ { href: "/browse", label: "projects" },
+ { href: `/browse/${project.id}`, label: project.id },
+ { label: "takes" },
+ ]}
+ note={takes.length ? `${takes.length} takes · ${judged} judged` : undefined}
+ />
+ <main className="deck-main flex-1 space-y-4 p-4">
+ {takes.length === 0 ? (
+ <p data-takes-empty="" className="text-[12px] text-[var(--color-dim)]">
+ No takes yet.{" "}
+ <Link href={`/browse/${project.id}`} className="text-[var(--color-sel)] hover:underline">
+ back to {project.id}
+ </Link>
+ </p>
+ ) : (
+ <TakesBench
+ project={project.id}
+ groups={listing.groups.map((g) => ({ group: g.group, takes: g.takes as Take[] }))}
+ initial={verdicts as Record<string, TakeVerdict>}
+ />
+ )}
+ {listing.skipped.length > 0 && (
+ <section data-takes-skipped="" className="text-[11px] text-[var(--color-dim)]">
+ <h2 className="micro mb-1">skipped — {listing.skipped.length}</h2>
+ <ul className="space-y-0.5">
+ {listing.skipped.map((s) => (
+ <li key={s.dir} data-skipped={s.dir}>
+ <code className="font-mono text-[var(--color-text)]">takes/{s.dir}</code> — {s.why}
+ </li>
+ ))}
+ </ul>
+ </section>
+ )}
+ </main>
+ </div>
+ );
+}
diff --git a/umtool/components/projects/TimelineEditor.tsx b/umtool/components/projects/TimelineEditor.tsx
@@ -0,0 +1,205 @@
+"use client";
+
+import { useRouter } from "next/navigation";
+import { createContext, useCallback, useContext, useEffect, useRef, useState } from "react";
+import AnchoredNotes from "@/components/notes/AnchoredNotes";
+import { announceNotesChanged, useSharedNotes, wroteNotes } from "@/components/notes/NotesProvider";
+import { buttonVariants } from "@/components/ui/button";
+import { MANIFEST_CHANGED, announceManifestChanged, refusalOf, timelineOp } from "./timelineApi";
+
+// Re-ordering the cut, on the project page.
+//
+// The rows stay the server's (each `<li data-entry data-at>` the page always
+// drew); this list adds what moves them. Drag a row by its handle, or focus it
+// and press alt+↑ / alt+↓; each row's menu duplicates it, removes it, or
+// inserts a clip after it (`<channel>/<video>@<start>-<end>`). Every one is a
+// structural write (POST /api/report/timeline): snapshotted first, so "Undo"
+// restores the cut as it was before the last burst of edits.
+
+type Ctx = {
+ project: string;
+ busy: boolean;
+ run: (op: string, args: Record<string, unknown>) => Promise<boolean>;
+ dragging: React.MutableRefObject<{ id: string; at: number } | null>;
+};
+const TimelineCtx = createContext<Ctx | null>(null);
+
+const rowOf = (el: EventTarget | null) => (el instanceof Element ? (el.closest("li[data-entry][data-at]") as HTMLLIElement | null) : null);
+const rowKey = (li: HTMLLIElement) => ({ id: li.dataset.entry ?? "", at: Number(li.dataset.at) });
+
+export function TimelineList({ project, count, children }: { project: string; count: number; children: React.ReactNode }) {
+ const router = useRouter();
+ // The manifest's write token, in a ref as well as state: a key pressed before
+ // the first read came back must wait for it, not send no token at all.
+ const tokenRef = useRef<string | null>(null);
+ const [busy, setBusy] = useState(false);
+ const [error, setError] = useState<string | null>(null);
+ const [over, setOver] = useState<number | null>(null);
+ const dragging = useRef<{ id: string; at: number } | null>(null);
+
+ const loadToken = useCallback(async () => {
+ const r = await fetch(`/api/report/timeline?project=${encodeURIComponent(project)}`, { cache: "no-store" });
+ const j = await r.json().catch(() => ({}));
+ if (r.ok) tokenRef.current = typeof j.token === "string" ? j.token : null;
+ return tokenRef.current;
+ }, [project]);
+ useEffect(() => {
+ void loadToken();
+ const on = (e: Event) => {
+ if ((e as CustomEvent).detail?.project === project) void loadToken();
+ };
+ window.addEventListener(MANIFEST_CHANGED, on);
+ return () => window.removeEventListener(MANIFEST_CHANGED, on);
+ }, [loadToken, project]);
+
+ const run = useCallback(
+ async (op: string, args: Record<string, unknown>) => {
+ setBusy(true);
+ setError(null);
+ const res = await timelineOp(project, tokenRef.current ?? (await loadToken()), op, args);
+ setBusy(false);
+ if (!res.ok) {
+ setError(refusalOf(res));
+ if (res.json.stale) await loadToken();
+ return false;
+ }
+ tokenRef.current = typeof res.json.token === "string" ? res.json.token : null;
+ announceManifestChanged(project);
+ if (wroteNotes(res.json)) announceNotesChanged(project);
+ router.refresh();
+ return true;
+ },
+ [project, loadToken, router],
+ );
+
+ const onKeyDown = (e: React.KeyboardEvent) => {
+ if (!e.altKey || (e.key !== "ArrowUp" && e.key !== "ArrowDown")) return;
+ if (e.target instanceof HTMLElement && /^(INPUT|TEXTAREA|SELECT)$/.test(e.target.tagName)) return;
+ const li = rowOf(e.target);
+ if (!li || busy) return;
+ const { id, at } = rowKey(li);
+ const to = at + (e.key === "ArrowUp" ? -1 : 1);
+ if (to < 0 || to >= count) return;
+ e.preventDefault();
+ void run("move", { id, at, toIndex: to }).then((ok) => {
+ // Keep the moved row focused after the page redraws it.
+ if (ok) setTimeout(() => (document.querySelector(`li[data-at="${to}"]`) as HTMLElement | null)?.focus(), 400);
+ });
+ };
+
+ return (
+ <TimelineCtx.Provider value={{ project, busy, run, dragging }}>
+ <div className="mb-1.5 flex flex-wrap items-center gap-2 text-[11px]">
+ <button type="button" data-testid="timeline-undo" disabled={busy} onClick={() => void run("undo", {})} className={buttonVariants({ size: "sm" })}>
+ Undo last edit
+ </button>
+ <span className="text-[var(--color-dim)]">
+ drag ⠿ or <kbd>alt</kbd>+<kbd>↑</kbd>/<kbd>↓</kbd> to move
+ </span>
+ {busy && <span className="text-[var(--color-dim)]">saving…</span>}
+ {error && (
+ <span data-testid="timeline-error" className="text-[var(--color-bad)]">
+ {error}
+ </span>
+ )}
+ </div>
+ <ul
+ className="space-y-1"
+ data-testid="timeline-list"
+ data-drop-at={over ?? ""}
+ onKeyDown={onKeyDown}
+ onDragOver={(e) => {
+ if (!dragging.current) return;
+ const li = rowOf(e.target);
+ if (!li) return;
+ e.preventDefault();
+ setOver(rowKey(li).at);
+ }}
+ onDragLeave={() => setOver(null)}
+ onDrop={(e) => {
+ const from = dragging.current;
+ const li = rowOf(e.target);
+ dragging.current = null;
+ setOver(null);
+ if (!from || !li) return;
+ e.preventDefault();
+ const to = rowKey(li).at;
+ if (to !== from.at) void run("move", { id: from.id, at: from.at, toIndex: to });
+ }}
+ >
+ {children}
+ </ul>
+ </TimelineCtx.Provider>
+ );
+}
+
+/** Inside one row: the drag handle and the row's menu. */
+export function RowControls({ id, at }: { id: string; at: number }) {
+ const ctx = useContext(TimelineCtx);
+ const [menu, setMenu] = useState(false);
+ const [insert, setInsert] = useState<string | null>(null);
+ if (!ctx) return null;
+ const { busy, run, dragging } = ctx;
+ return (
+ <>
+ <span
+ draggable={!busy}
+ data-drag-handle={id}
+ title="drag to move"
+ onDragStart={(e) => {
+ dragging.current = { id, at };
+ e.dataTransfer.effectAllowed = "move";
+ e.dataTransfer.setData("text/plain", id);
+ }}
+ onDragEnd={() => (dragging.current = null)}
+ className="cursor-grab select-none text-[13px] leading-none text-[var(--color-dim)]"
+ >
+ ⠿
+ </span>
+ <span className="relative">
+ <button type="button" data-row-menu={id} onClick={() => setMenu((m) => !m)} className="px-1 text-[12px] text-[var(--color-dim)] hover:text-[var(--color-text)]" aria-label={`actions for ${id}`}>
+ ⋯
+ </button>
+ {menu && (
+ <span className="absolute right-0 z-10 mt-1 flex w-36 flex-col rounded border border-[var(--color-line)] bg-[var(--color-panel)] p-1 text-[11px] shadow">
+ <button type="button" data-row-action="duplicate" disabled={busy} onClick={() => (setMenu(false), void run("duplicate", { id, at }))} className="px-1 py-0.5 text-left hover:bg-[var(--color-panel-2)]">
+ duplicate
+ </button>
+ <button type="button" data-row-action="insert" disabled={busy} onClick={() => (setMenu(false), setInsert(""))} className="px-1 py-0.5 text-left hover:bg-[var(--color-panel-2)]">
+ insert clip after
+ </button>
+ <button type="button" data-row-action="remove" disabled={busy} onClick={() => (setMenu(false), void run("remove", { id, at }))} className="px-1 py-0.5 text-left text-[var(--color-bad)] hover:bg-[var(--color-panel-2)]">
+ remove
+ </button>
+ </span>
+ )}
+ </span>
+ {insert !== null && (
+ <input
+ autoFocus
+ data-testid={`insert-after-${id}`}
+ value={insert}
+ placeholder="<channel>/<video>@<start>-<end>"
+ onChange={(e) => setInsert(e.target.value)}
+ onKeyDown={(e) => {
+ if (e.key === "Escape") setInsert(null);
+ if (e.key === "Enter" && insert.trim()) {
+ void run("insert", { afterId: id, at, entry: insert.trim() }).then((ok) => ok && setInsert(null));
+ }
+ }}
+ className="w-64 rounded border border-[var(--color-line)] bg-[var(--color-ink)] px-1.5 py-0.5 font-mono text-[11px] text-[var(--color-text)] outline-none focus:border-[var(--color-sel)]"
+ />
+ )}
+ </>
+ );
+}
+
+/** Under one row: its notes (and the edit notes a generated manifest left on it). */
+export function RowNotes({ id, project }: { id: string; project: string }) {
+ const notes = useSharedNotes({ project });
+ return (
+ <div className="mt-1">
+ <AnchoredNotes notes={notes} pin={{ kind: "entry", entry: id }} />
+ </div>
+ );
+}
diff --git a/umtool/components/projects/timelineApi.ts b/umtool/components/projects/timelineApi.ts
@@ -0,0 +1,39 @@
+"use client";
+
+// The client's side of POST /api/report/timeline, and the event every part of
+// a project page listens to after the manifest changed under it.
+//
+// The timeline editor, the structure editors and the On-screen section each
+// hold a manifest TOKEN; a structural write changes the file, so after one
+// every other holder must re-read before its next save would 409.
+
+export const MANIFEST_CHANGED = "umtool:manifest-changed";
+
+export function announceManifestChanged(project: string) {
+ window.dispatchEvent(new CustomEvent(MANIFEST_CHANGED, { detail: { project } }));
+}
+
+export type TimelineResult = {
+ ok: boolean;
+ status: number;
+ json: Record<string, unknown> & { error?: string; errors?: string[]; stale?: boolean; token?: string };
+};
+
+export async function timelineOp(project: string, token: string | null, op: string, args: Record<string, unknown> = {}): Promise<TimelineResult> {
+ const r = await fetch("/api/report/timeline", {
+ method: "POST",
+ headers: { "content-type": "application/json" },
+ body: JSON.stringify({ project, token, op, ...args }),
+ cache: "no-store",
+ });
+ const json = (await r.json().catch(() => ({ error: `HTTP ${r.status}` }))) as TimelineResult["json"];
+ return { ok: r.ok, status: r.status, json };
+}
+
+/** A refusal in words: the build's sentences when there are several. */
+export const refusalOf = (res: TimelineResult) =>
+ res.json.stale
+ ? "the manifest changed since this page read it — reloaded; try again"
+ : res.json.errors?.length
+ ? res.json.errors.join("; ")
+ : String(res.json.error ?? `HTTP ${res.status}`);
diff --git a/umtool/docs/README.md b/umtool/docs/README.md
@@ -93,6 +93,8 @@ defects that have already shipped in real videos: a manifest with no `siteOrigin
| [mix-from-a-project.md](mix-from-a-project.md) | Deep-linking a clip into `/mix` |
| [index.md](index.md) | The LMDB index, and why the filesystem stays the model |
| [cli.md](cli.md) | `umtool ls / show / check / build / window / …` |
+| [notes.md](notes.md) | Operator notes on articles and videos; `umtool notes`, the agent loop |
+| [sites.md](sites.md) | `/sites`: every site's articles, their media, workspaces and notes |
| [authoring.md](authoring.md) | Writing a manifest from a sweep report |
| [e2e.md](e2e.md) | The fixture, the stubs, the global queue |
| [quirks.md](quirks.md) | Everything that cost time to find out |
diff --git a/umtool/docs/e2e.md b/umtool/docs/e2e.md
@@ -37,6 +37,7 @@ rendered over a deliverable would be indistinguishable from a person doing it.
| `onscreen-build-fixture` | built with the deck, then re-rendered on-screen (`--chrome-only`) |
| `gone-fixture` | its source is gone — the preflight must block it |
| `no-origin-fixture` / `localhost-fixture` | the two defects that shipped |
+| `takes-fixture` | takes **write** `takes/verdicts.json` — two groups, a take with no preview, one skipped take.json, a work dir that is not a take |
| `bike-fixture` | the third kind |
| `find/` | shadowed by a tool page |
| `deep/nested/solo-fixture` | a pass-through chain, for the collapse |
diff --git a/umtool/docs/notes.md b/umtool/docs/notes.md
@@ -0,0 +1,112 @@
+# Notes: the operator writes them in umtool, an agent acts on them
+
+A note is a sentence or two the operator leaves on an **article** (a report on a
+site, `/sites/<site>/<report>`) or on a **report-video project** (a timeline row, a
+take, a moment in a rendered cut). It is saved where the agent that made the thing
+can read it back, act on the right SOURCE file, and reply or resolve.
+
+```sh
+umtool notes --all --open # every notes file with an open note
+umtool notes candalyzer/polemic-israel # one article's open notes, as markdown
+umtool notes candace/polemic-israel # one video project's (plus every take's verdict)
+umtool notes reply n_mgk3x0a1b2 "Changed 'said' to 'claimed' in drafts/israel.json" --resolve
+umtool notes resolve | wontfix | reopen n_mgk3x0a1b2
+```
+
+Run it from the repo checkout (or with `TRANSCRIPTS_DIR`/`SITES_DIR` and
+`REPORTS_DIR` set): like `CHANNELS_DIR`, the sites directory is found by walking up
+from the working directory to the checkout.
+
+## The agent loop
+
+1. `umtool notes --all --open` — what is waiting.
+2. `umtool notes <target>` — each open note with its anchor resolved against what
+ is on disk now, and the **Source** line naming the file to edit.
+3. Edit the SOURCE, not the output. An article's `report.json` is regenerated from
+ a draft (`<workspace>/polemics/drafts/<slug>.json`) by a generator
+ (`make-site.py`, `polemics.py`); a generated video manifest (`generatedBy`) is
+ rewritten by its `make-videos.py`. An edit to the output is lost on the next run.
+4. Regenerate.
+5. `umtool notes reply <id> "<what changed>" --resolve` — or reply without
+ `--resolve` to ask a question. The operator sees agent replies badged "agent".
+
+**Never hand-edit `notes.json`.** The CLI and the app write through one store that
+validates, takes a lock, and refuses a stale write; a hand edit races the page and
+can lose a reply. If the recorded source is wrong, correct it:
+`umtool notes source <target> --draft <path> [--generator <path>] [--how <why>]`.
+
+The same digest is served as text at `GET /api/notes/context?article=<site>/<report>`
+(or `?project=<id>`); the page's "Copy agent brief" copies it.
+
+## Where notes live
+
+| Subject | File |
+|---|---|
+| An article | `transcripts/sites/<site>/reports/<report>/notes.json`, beside `report.json` |
+| A report-video project | `<project>/notes.json`, beside `video.manifest.json` |
+
+The article file is the one corpus file umtool writes, and only through
+`isCorpusNotesFile` (`lib/paths.mjs`): exactly `sites/<site>/reports/<id>/notes.json`,
+in an existing report directory that is not a symlink out of its site. The
+generators overwrite only `report.json`, `video.mp4` and `poster.jpg` in a report
+directory, so notes survive a regenerate; compose and the report history never read
+it (`common/publish/composeReports.test.ts` holds that), so a note is never published.
+
+## The file
+
+```jsonc
+{ "format": "umtool-notes", "version": 1,
+ "subject": { "kind": "article", "site": "candalyzer", "report": "polemic-israel" },
+ // | { "kind": "video-project", "project": "candace/polemic-israel" }
+ "source": { "draft": "~/reports/candace/polemics/drafts/israel.json",
+ "generator": "~/reports/candace/site/polemics.py",
+ "how": "draft matched by id polemic-israel; generator names candalyzer" },
+ "notes": [ { "id": "n_…", "status": "open", // open | resolved | wontfix
+ "author": "operator", "text": "…", // operator | agent
+ "at": "…", "updatedAt": "…",
+ "anchor": { … },
+ "replies": [ { "author": "agent", "text": "…", "at": "…" } ],
+ "resolvedAt": "…", "resolvedBy": "agent" } ] }
+```
+
+`source` is filled on the first write. For an article it comes from
+`lib/articles/sources.mjs`: a draft whose `id` or file name is the report id (allowing
+a `polemic-` prefix either way; the workspace whose generator names the site wins a
+tie), and the generator under `<workspace>/polemics` or `<workspace>/site` that names
+the site or the report. For a video project it is the manifest and, when it has
+`generatedBy`, the generator.
+
+The last note deleted deletes the file. A `notes.json` that does not parse is never
+overwritten — fix it by hand first.
+
+## Anchors
+
+| `kind` | Fields | Resolved for the agent as |
+|---|---|---|
+| `text` | `section` (a section id, or `title`/`subtitle`/`summary`/`method`), `quote`, `prefix`, `suffix` | the section title and the sentence holding the quote |
+| `cite` | `cite` (a citation id) | the citation's quote, speaker, date, `<channel>/<id>@start-end` |
+| `section` | `section` | the section title |
+| `whole` | — | the whole article or project |
+| `moment` | `file` (relative mp4), `t`, `take?`, `entry?`, `resolved?` | the time, and the entry, quote and source second it resolved to when written |
+| `entry` | `entry` (a timeline entry id) | the entry's title, quote and source span |
+| `take` | `take` (a take id) | the take's label and summary, and its verdict |
+| `edit` | `entry?`, `field`, `from`, `to` | an edit made in umtool to a GENERATED manifest — port it into the generator's inputs |
+
+A text anchor is a quote with 32 characters of context either side (the W3C
+TextQuoteSelector), never an offset, so it survives a regenerated `report.json`:
+`lib/annotations/anchor.mjs` finds the quote exactly, then with whitespace, quotes
+and case normalised, then — when the quote itself was rewritten — between its
+surviving context. A note it cannot place is **orphaned**: still shown, pinned to its
+section, with its original quote.
+
+## Code
+
+| Module | What |
+|---|---|
+| `lib/annotations/shape.mjs` | the contract: constants, validation, the ops (pure, client-safe) |
+| `lib/annotations/anchor.mjs` | re-anchoring (pure, client-safe) |
+| `lib/annotations/store.mjs` | read, lockfile (`notes.json.lock`, stale after 30 s), token, tmp+rename |
+| `lib/annotations/targets.mjs` | article / project → file, subject, source; `listNotesFiles` |
+| `lib/annotations/digest.mjs`, `cli.mjs` | the markdown digest and `umtool notes` |
+| `app/api/notes/route.ts` | GET / POST (operator-stamped; 409 on a stale token) |
+| `lib/annotations/useNotes.ts` | the page's hook |
diff --git a/umtool/docs/quirks.md b/umtool/docs/quirks.md
@@ -26,7 +26,21 @@ and Rumble ships no progressive fallback, so every Rumble clip is unbuildable
without it. But the option lives on the **HLS demuxer**: pass it against a
progressive URL (YouTube's googlevideo mp4) and ffmpeg aborts with "Option
extension_picky not found". Adding it unconditionally trades a Rumble failure for
-a YouTube one.
+a YouTube one. The editor's window fetch keys it on the URL instead: a rumble.com
+URL gets it on the FIRST try (a retry loads the Rumble page again, and Cloudflare
+403s a share of page loads), any other URL only as the retry.
+
+**A Rumble video gets Rumble's args from its URL, not its channel.** A channel
+with no platform of its own (community-notes) holds Rumble videos; fetched with
+that channel's args they went without `--impersonate` and Cloudflare 403'd them.
+
+**A 403 backs the platform off in a batch.** `fetch-via-editor.mjs --all` (the
+editor's `fetch-windows` job) treats one 403 as that clip's failure — a removed
+Rumble page answers 403 too — but two in a row back the platform off and stop the
+job, as a 429 does at once. The windows are 30–45 s apart; 20 s drew YouTube 403s.
+Rumble's are 120–180 s apart, across jobs too: its Cloudflare puts the whole IP
+behind a JS challenge ("Just a moment…", 403 to everything) after a handful of
+requests in a few minutes, and it lifts after a few quiet minutes.
**`--force-keyframes-at-cuts` matters because the clip IS the citation.** Without
it the cut snaps to the nearest preceding keyframe, which can be seconds early.
diff --git a/umtool/docs/report-video.md b/umtool/docs/report-video.md
@@ -692,6 +692,85 @@ containing a newline (`coverage`, `transcriptGapNote`, `qrNote`, `revisionNote`,
beside the manifest (`build-notes.md`, `ADJUDICATION-DIVERGENCE.md`) are listed
there too. No field is moved or rewritten.
+## Takes: alternative renders, judged side by side
+
+`/browse/<project>/takes`, linked from the project page when `takes/` exists.
+Something else renders the takes; umtool reads them and records the verdicts.
+
+```
+takes/<id>/take.json id (== dir, [a-z0-9-]), group, order, label,
+ kind reference|similar|different, preview;
+ optional summary, changes[], seconds, builtAt
+takes/<id>/preview.mp4 whatever `preview` names, inside the take dir
+takes/verdicts.json { "<id>": { verdict: like|maybe|no|null, note, at } }
+```
+
+A take.json that fails validation is listed under **skipped** with the reason;
+a directory with no take.json is not a take. A row whose verdict is null and
+note empty is removed. A `verdicts.json` that does not parse is never
+overwritten. The rules are `lib/report/takes.mjs`; the routes are
+`/api/report/takes` (list), `/takes/preview` (ranges) and `/takes/verdict`.
+
+Each take also takes notes (a `take` anchor) and timed notes on its preview;
+the verdicts and those notes reach the agent through `umtool notes <project>`,
+which lists every take with its verdict and note ([notes.md](notes.md)).
+
+## Operator notes, timed notes and the generated-manifest guard
+
+Notes on a project live in `<project>/notes.json` (shape and CLI:
+[notes.md](notes.md)): **row notes** on a timeline entry (`entry` anchor, an open
+count per row on the project page and in the clip bench), **take notes**, **timed
+notes** and **edit notes**.
+
+**Timed notes.** On the built cut (On-screen → the built video), on every take
+preview and on an article's report video: `n` with the video focused, or Mark,
+writes a note at the playhead, and the tick strip under the player seeks. On a
+project the second is resolved against the build's `schedule.json` (the
+deliverable `out/<slug>.mp4` against `out/sourced/`, a variant's file against its
+own; a take's preview against `takes/<id>/out/<variant>/`) to the entry on screen:
+its title, quote, source second and archive link, stored in the note as
+`resolved`. A take's `preview.mp4` is that build's output copied, and its length
+equals the schedule's `total` exactly (all 9 takes of candace/polemic-israel), so
+a mark is exact when the lengths match within 0.25 s and `approx` when they
+differ, when the schedule is estimated, or when the second is a held frame past
+the clip's source. `lib/report/moments.mjs`, `GET /api/report/moment`.
+
+**Generated manifests.** A manifest with `generatedBy` shows one line on the
+project page, the clip bench and the On-screen section: "Generated by
+`<generatedBy>`; a rebuild of manifests overwrites edits made here." Edits are
+still allowed. Every manifest writer ROUTE goes through `withEditNotes`
+(`lib/report/guard.ts`), which diffs the manifest before and after and records
+each change as an `edit` note `{ entry?, field, from, to }`: repeated saves of one
+field coalesce into one note, and putting a field back deletes its note unless it
+has replies. The agent ports each change into the generator's inputs (BEATS,
+drafts), regenerates and resolves the note. `umtool window` is the agent's own
+tool and is not wrapped.
+
+## Structural edits
+
+The project page's timeline re-orders by drag or alt+↑/↓; each row's ⋯ menu
+duplicates it, removes it, or inserts a clip after it
+(`<channel>/<video>@<start>-<end>`). The On-screen section edits a teaser's lines,
+beat, tail, tail wait and dip; the posts (add, edit, remove); and each verdict's
+label and colour and the stamp seconds. An article's citation can be added to the
+end of a linked project's timeline ("Add to video", [sites.md](sites.md)).
+
+All of it is `POST /api/report/timeline` (`move`, `remove`, `duplicate`, `insert`,
+`teaser`, `post`, `post-remove`, `factcheck`, `undo`; `GET` gives the token), with
+the writers in `lib/report/manifest.mjs` beside the window writers: the mtime
+token, tmp + rename, 2 dp. Each op is checked by the build's own validators
+(`deck.mjs`, `factcheck.mjs`) and refused only for what it breaks, and snapshots
+the manifest first (`revisions/<stamp>-auto-before-<op>.manifest.json`, at most
+one per op every two minutes). **Undo** restores the newest such snapshot byte for
+byte (the current state is kept as `undo-saved`).
+
+A move recomputes `sectionEnter` — a clip enters when its `section` differs from
+the previous clip's, the first clip included (`lib/report/sections.mjs`) — and only
+in a manifest that already carries the flag; `section` itself never changes.
+report-to-video only READS the flag. Checked read-only on all 44 manifests under
+~/reports: no mismatch with the stored flags, and 3,490 move-and-back round trips
+byte-identical.
+
## Exports
`umtool export <project> --format toc-bbcode|toc-markdown|description|chapters`
diff --git a/umtool/docs/sites.md b/umtool/docs/sites.md
@@ -0,0 +1,92 @@
+# /sites — every site's articles
+
+Every report on every site under `transcripts/sites/` — published and drafts —
+read in place: no private site built, no port served. umtool **reads** the sites;
+the one file it writes there is a report's `notes.json` (docs/notes.md).
+
+| page | what |
+|---|---|
+| `/sites` | every site (private first), every article: status, updated, citations, open notes, poster, video project, source draft. Filters are links: `?site=`, `?status=published\|draft`, `?notes=open` |
+| `/sites/<site>` | its articles, its report videos (playable), the umtool projects they were cut in (takes, like/maybe/no), the workspace files they were written from (`?ws=&rel=` opens one) |
+| `/sites/<site>/<report>` | the article with its notes (below) |
+| `/sites/<site>/<report>?tab=source` | the article's workspace files, its own draft opened |
+| `/sites/<site>/<report>/evidence` | every citation, one per screen (`?c=<id>`) |
+
+## Where an article comes from
+
+- **Site and reports**: common's `listSites`/`getSite` (site.json, the editor's
+ defaults), `listReportDirs` (the enumerator the editor's Reports tab uses), each
+ `report.json` through the report validator. **Published** = named in site.json
+ `reports`, in that order; every other report directory is a **draft**.
+- **Source** (`lib/articles/sources.mjs`): the draft under
+ `~/reports/<ws>/polemics/drafts/` whose `id` or file name is the report id,
+ `polemic-` allowed on either side, preferring a workspace whose generator names
+ the site; the generator is the `*.py`/`*.mts` under `<ws>/polemics` or
+ `<ws>/site` that names the site or the report. Shown with its reason; written
+ into `notes.json` `source` on the first note. **Edit the draft, not
+ report.json** — the generator overwrites report.json.
+- **Video project** (`lib/articles/links.mjs`): a report-video manifest with a
+ top-level `"article": "<site>/<report>"` is linked to that article and nothing
+ else (build-video.mjs ignores the key). Otherwise a manifest `slug` equal to the
+ report id (± `polemic-`) in the draft's workspace, when exactly one matches;
+ several are listed as "possible".
+
+## Reading and noting
+
+The article renders in a ~70ch column with its notes in a rail (a drawer below
+1100px). Its video sits under the title (`components/articles/ArticleVideo.tsx`)
+and takes timed notes: `n` with the video focused, or **Mark**, notes the
+playhead as a `moment` anchor on the article's notes ([report-video.md](report-video.md)).
+
+| to note | do |
+|---|---|
+| a passage | select it → **Note** (or `n`) |
+| a section | **+ note** on its heading |
+| a citation | click it → the evidence panel → **+ note on this citation** |
+| the article | **+ whole article** |
+
+A passage note keeps its quote and 32 characters either side; it is found again
+in the text as it is now (`lib/annotations/anchor.mjs`: exact, then
+whitespace/quote/case-normalised, then by its surviving context) and marked.
+One whose quote is gone is listed **orphaned**. Keys, when not typing: `j`/`k`
+move, `r` resolves, `e` edits, `n` notes, `Esc` closes. `?status=` and `?note=`
+open the rail on a filter or a note (the decisions inbox links to the note).
+**Copy agent brief** copies `/api/notes/context` — the `umtool notes` digest.
+
+Open notes are `open-note` rows in `/browse/decisions` and on the dashboard.
+
+## Evidence
+
+A citation's panel: its quote, speaker, date, label and original link; ±6 cues
+of its record's `transcript.cues.json`, the cited ones marked (click one to seek);
+and media, best first:
+
+1. the site's **prepared** clip (`.export-index/sites/<site>/report-media/`);
+2. a clip **window** the editor fetched (`data/<id>/clips/`);
+3. a **saved** video or audio file in `data/<id>/` (through its media-tier link);
+4. otherwise the `fetch_clip` MCP line to fetch it through the editor. umtool
+ never runs yt-dlp, and its own fetch client needs a project manifest.
+
+A post shows its text and its capture screenshot. The walk
+(`/evidence`) is the same panel one citation at a time: `j`/`k` (or ←/→),
+`space` plays, `n` notes.
+
+**Add to video** (in the panel, when the article has a linked video project)
+puts the cited span at the end of that project's timeline as a clip
+(`/api/report/timeline` insert): the manifest is snapshotted first, so the
+project page's **Undo** takes it back, and on a generated manifest an `edit` note
+tells the agent to port it into the generator's inputs.
+
+## Routes (read-only)
+
+| route | serves |
+|---|---|
+| `/api/sites/media?site&report&file` | a file in the report dir (video.mp4, poster.jpg, stills/…), byte ranges |
+| `/api/sites/media?site&moment` | the prepared clip of a moment, via report-media/index.json |
+| `/api/sites/media?corpus=<abs>` | a media file lexically under CHANNELS_DIR |
+| `/api/sites/evidence?site&report&cite` | one citation's evidence (JSON) |
+| `/api/sites/workspace?ws&rel` | a listed workspace file; HTML under a CSP sandbox |
+
+`SITES_DIR` (else `TRANSCRIPTS_DIR/sites`, else the checkout's
+`transcripts/sites`) is where all of it is read; the e2e suite points it at its
+fixture (`e2e/fixtures/sites-fixture.mjs`).
diff --git a/umtool/e2e/article-evidence.spec.ts b/umtool/e2e/article-evidence.spec.ts
@@ -0,0 +1,71 @@
+import { test, expect } from "@playwright/test";
+import { rmSync } from "node:fs";
+import path from "node:path";
+import { fileURLToPath } from "node:url";
+import { WARM_TIMEOUT, warm } from "./warm";
+
+// ---------------------------------------------------------------------------
+// A citation's evidence: the transcript around it, the window that plays it,
+// or -- with nothing on disk -- the fetch_clip line; and the one-per-screen
+// walk.
+//
+// priv/polemic-alpha c1 → sitechan/sv1, a clip window on disk
+// c2 → sitechan/sv2, cues only
+// ---------------------------------------------------------------------------
+
+const HERE = path.dirname(fileURLToPath(import.meta.url));
+const NOTES = path.join(HERE, "..", ".e2e-song", "sites", "priv", "reports", "polemic-alpha", "notes.json");
+
+
+test.beforeAll(async ({ playwright }) => {
+ test.setTimeout(WARM_TIMEOUT);
+ await warm(playwright, ["/sites/priv/polemic-alpha", "/sites/priv/polemic-alpha/evidence", "/api/sites/evidence?site=priv&report=polemic-alpha&cite=c1", "/api/notes?article=priv/polemic-alpha"]);
+});
+
+test.beforeEach(() => rmSync(NOTES, { force: true }));
+test.afterAll(() => rmSync(NOTES, { force: true }));
+
+test("a citation opens its evidence: the cited cue marked, the window playing from the span", async ({ page }) => {
+ await page.goto("/sites/priv/polemic-alpha");
+ await page.locator("button[data-cite='c1']").first().click();
+ const ev = page.locator("[data-evidence='c1']");
+ await expect(ev).toContainText("The vote was rigged and everybody knew.");
+ await expect(ev.locator("[data-play]")).toHaveAttribute("data-play", "window");
+ await expect(ev.locator("[data-cited='true']")).toHaveCount(1);
+ await expect(ev.locator("[data-cited='true']")).toContainText("The vote was rigged");
+ await expect(ev.locator("[data-cues] li")).toHaveCount(6);
+ const video = ev.locator("video");
+ await expect(video).toHaveAttribute("src", /corpus=.*sv1.*clips/);
+ await expect.poll(() => video.evaluate((v: HTMLVideoElement) => v.currentTime)).toBeGreaterThanOrEqual(0);
+});
+
+test("with nothing on disk it says so and offers the fetch_clip line, never yt-dlp", async ({ page }) => {
+ await page.goto("/sites/priv/polemic-alpha");
+ await page.locator("button[data-cite='c2']").first().click();
+ const ev = page.locator("[data-evidence='c2']");
+ await expect(ev.locator("[data-play]")).toHaveAttribute("data-play", "none");
+ await expect(ev).toContainText('fetch_clip {"channel":"sitechan","video":"sv2","start":8,"end":12');
+ await expect(ev).not.toContainText("yt-dlp");
+});
+
+test("the walk: one citation per screen, j/k, n notes it, the URL follows", async ({ page }) => {
+ await page.setViewportSize({ width: 1366, height: 768 });
+ await page.goto("/sites/priv/polemic-alpha/evidence");
+ await expect(page.locator("[data-walk='c1']")).toBeVisible();
+ await expect(page.locator("[data-walk-position]")).toHaveText("1 / 2");
+ await page.keyboard.press("j");
+ await expect(page.locator("[data-walk='c2']")).toBeVisible();
+ await expect(page).toHaveURL(/\?c=c2$/);
+ await page.keyboard.press("n");
+ await page.getByLabel("note text").fill("Needs a clip.");
+ await page.keyboard.press("Control+Enter");
+ await expect(page.locator("[data-walk='c2'] [data-note-card]")).toHaveCount(1);
+ await page.keyboard.press("k");
+ await expect(page.locator("[data-walk='c1'] [data-note-card]")).toHaveCount(0);
+
+ // It fits the laptop screen: no horizontal scroll.
+ expect(await page.evaluate(() => document.documentElement.scrollWidth <= window.innerWidth)).toBe(true);
+
+ await page.goto("/sites/priv/polemic-alpha/evidence?c=c2");
+ await expect(page.locator("[data-walk='c2'] [data-note-card]")).toContainText("Needs a clip.");
+});
diff --git a/umtool/e2e/article-integration.spec.ts b/umtool/e2e/article-integration.spec.ts
@@ -0,0 +1,106 @@
+import { test, expect, type Locator } from "@playwright/test";
+import { copyFileSync, existsSync, readFileSync, readdirSync, rmSync } from "node:fs";
+import path from "node:path";
+import { fileURLToPath } from "node:url";
+import { WARM_TIMEOUT, warm } from "./warm";
+
+// ---------------------------------------------------------------------------
+// Where the two tracks meet on the article page:
+//
+// * the report's own video carries timed notes, written to the ARTICLE's
+// notes through the reader's one handle (a mark and a text note do not
+// 409 each other);
+// * "Add to video" puts a cited span at the end of the linked project's
+// timeline; the project is GENERATED, so an `edit` note is left for the
+// agent, and the change is undoable from the auto snapshot.
+//
+// priv/polemic-alpha video.mp4 (4 s); c1 → sitechan/sv1@3-6
+// sitews/polemic-alpha the linked project (generatedBy set), one clip a1
+// ---------------------------------------------------------------------------
+
+const HERE = path.dirname(fileURLToPath(import.meta.url));
+const FIX = path.join(HERE, "..", ".e2e-song");
+const NOTES = path.join(FIX, "sites", "priv", "reports", "polemic-alpha", "notes.json");
+const PROJ = path.join(FIX, "sitews", "polemic-alpha");
+const MANIFEST = path.join(PROJ, "video.manifest.json");
+const SAVED = path.join(PROJ, ".manifest.e2e-integration");
+const PROJ_NOTES = path.join(PROJ, "notes.json");
+
+const readJson = (f: string) => JSON.parse(readFileSync(f, "utf8"));
+
+async function seek(video: Locator, t: number) {
+ await video.evaluate(async (el: HTMLVideoElement, at: number) => {
+ if (el.readyState < 1) await new Promise((r) => el.addEventListener("loadedmetadata", r, { once: true }));
+ el.currentTime = at;
+ await new Promise((r) => el.addEventListener("seeked", r, { once: true }));
+ }, t);
+}
+
+function restoreProject() {
+ if (existsSync(SAVED)) {
+ copyFileSync(SAVED, MANIFEST);
+ rmSync(SAVED, { force: true });
+ }
+ rmSync(PROJ_NOTES, { force: true });
+ rmSync(`${MANIFEST}.bak`, { force: true });
+ const rev = path.join(PROJ, "revisions");
+ if (existsSync(rev)) for (const f of readdirSync(rev)) if (f.includes("auto-before")) rmSync(path.join(rev, f), { recursive: true, force: true });
+}
+
+test.beforeAll(async ({ playwright }) => {
+ test.setTimeout(WARM_TIMEOUT);
+ await warm(playwright, ["/sites/priv/polemic-alpha", "/api/notes?article=priv/polemic-alpha", "/api/report/timeline?project=sitews/polemic-alpha"]);
+});
+test.beforeEach(() => {
+ rmSync(NOTES, { force: true });
+ copyFileSync(MANIFEST, SAVED);
+});
+test.afterEach(() => {
+ rmSync(NOTES, { force: true });
+ restoreProject();
+});
+
+test("the article's video takes timed notes, on the article's notes, beside a text note", async ({ page }) => {
+ await page.goto("/sites/priv/polemic-alpha");
+ const video = page.getByTestId("article-video");
+ await expect(video).toBeVisible();
+ const scope = page.locator("[data-timed-notes='video.mp4']");
+ await seek(video, 1);
+ await scope.locator("[data-action='mark']").click();
+ const input = scope.getByTestId("timed-note-input");
+ await input.fill("the cut is late here");
+ await input.press("Enter");
+ await expect(scope.locator("[data-mark-at]")).toHaveCount(1);
+
+ // The rail's handle is the same one: a whole-article note right after does not 409.
+ await page.getByRole("button", { name: "+ whole article" }).click();
+ await page.getByLabel("note text").fill("Retitle it.");
+ await page.getByRole("button", { name: "save note" }).click();
+ await expect.poll(() => (existsSync(NOTES) ? readJson(NOTES).notes.length : 0)).toBe(2);
+ const doc = readJson(NOTES);
+ expect(doc.subject).toEqual({ kind: "article", site: "priv", report: "polemic-alpha" });
+ const moment = doc.notes.find((n: { anchor: { kind: string } }) => n.anchor.kind === "moment");
+ expect(moment.anchor).toMatchObject({ kind: "moment", file: "video.mp4", t: 1 });
+ expect(moment.author).toBe("operator");
+});
+
+test("Add to video: the cited span lands at the end of the linked project, with an edit note; undo restores it", async ({ page, request }) => {
+ const before = readJson(MANIFEST).timeline.length;
+ await page.goto("/sites/priv/polemic-alpha");
+ await page.locator("button[data-cite='c1']").first().click();
+ const add = page.locator("[data-evidence='c1'] [data-add-to-video]");
+ await expect(add).toContainText("sitews/polemic-alpha");
+ await add.getByRole("button", { name: "Add to video" }).click();
+ await expect(add.getByRole("status")).toContainText("cite-c1");
+
+ const m = readJson(MANIFEST);
+ expect(m.timeline.length).toBe(before + 1);
+ expect(m.timeline.at(-1)).toMatchObject({ type: "clip", id: "cite-c1", channel: "sitechan", video: "sv1", start: 3, end: 6 });
+ const notes = readJson(PROJ_NOTES).notes;
+ expect(notes.some((n: { anchor: { kind: string } }) => n.anchor.kind === "edit")).toBe(true);
+
+ const g = await (await request.get("/api/report/timeline?project=sitews/polemic-alpha")).json();
+ const undo = await request.post("/api/report/timeline", { data: { project: "sitews/polemic-alpha", token: g.token, op: "undo" } });
+ expect(undo.ok()).toBe(true);
+ expect(readJson(MANIFEST).timeline.length).toBe(before);
+});
diff --git a/umtool/e2e/article-notes.spec.ts b/umtool/e2e/article-notes.spec.ts
@@ -0,0 +1,186 @@
+import { test, expect, type Page } from "@playwright/test";
+import { existsSync, readFileSync, rmSync } from "node:fs";
+import path from "node:path";
+import { fileURLToPath } from "node:url";
+import { WARM_TIMEOUT, warm } from "./warm";
+
+// ---------------------------------------------------------------------------
+// Notes on an article: written beside report.json (the one corpus file umtool
+// writes), anchored to a quote and found again, resolved, deleted -- and the
+// last delete removes the file.
+//
+// priv/polemic-alpha WRITES sites/priv/reports/polemic-alpha/notes.json
+// ---------------------------------------------------------------------------
+
+const HERE = path.dirname(fileURLToPath(import.meta.url));
+const NOTES = path.join(HERE, "..", ".e2e-song", "sites", "priv", "reports", "polemic-alpha", "notes.json");
+const PAGE = "/sites/priv/polemic-alpha";
+
+
+test.beforeAll(async ({ playwright }) => {
+ test.setTimeout(WARM_TIMEOUT);
+ await warm(playwright, ["/sites/priv/polemic-alpha", "/api/notes?article=priv/polemic-alpha", "/sites", "/browse/decisions?kind=open-note", "/api/sites/evidence?site=priv&report=polemic-alpha&cite=c1"]);
+});
+
+test.beforeEach(() => rmSync(NOTES, { force: true }));
+test.afterAll(() => rmSync(NOTES, { force: true }));
+
+async function select(page: Page, block: string, text: string) {
+ await page.evaluate(
+ ([block, text]) => {
+ const root = document.querySelector(`[data-block="${block}"]`)!;
+ const w = document.createTreeWalker(root, NodeFilter.SHOW_TEXT);
+ for (let n = w.nextNode() as Text | null; n; n = w.nextNode() as Text | null) {
+ const i = n.data.indexOf(text);
+ if (i < 0) continue;
+ const r = document.createRange();
+ r.setStart(n, i);
+ r.setEnd(n, i + text.length);
+ const s = getSelection()!;
+ s.removeAllRanges();
+ s.addRange(r);
+ return;
+ }
+ throw new Error(`no "${text}" in ${block}`);
+ },
+ [block, text],
+ );
+ await page.locator(`[data-block="${block}"]`).dispatchEvent("mouseup");
+}
+
+test("select text, Note, save: a mark on the quote, a note beside report.json", async ({ page }) => {
+ await page.goto(PAGE);
+ await select(page, "first", "Nobody checked the claim");
+ await page.getByRole("button", { name: "Note", exact: true }).click();
+ await page.getByLabel("note text").fill("Source this sentence.");
+ await page.getByRole("button", { name: "save note" }).click();
+
+ await expect(page.locator("mark[data-note]")).toHaveText("Nobody checked the claim");
+ const doc = JSON.parse(readFileSync(NOTES, "utf8"));
+ expect(doc.subject).toEqual({ kind: "article", site: "priv", report: "polemic-alpha" });
+ expect(doc.source.draft).toMatch(/sitews\/polemics\/drafts\/alpha\.json$/);
+ expect(doc.notes[0]).toMatchObject({ status: "open", author: "operator", text: "Source this sentence." });
+ expect(doc.notes[0].anchor).toMatchObject({ kind: "text", section: "first", quote: "Nobody checked the claim" });
+ expect(doc.notes[0].anchor.prefix).toContain("first stream.");
+
+ // Found again after a reload, and hovering the card lights the mark.
+ await page.reload();
+ const card = page.locator("[data-note-card]");
+ await expect(card).toHaveCount(1);
+ await card.hover();
+ await expect(page.locator("mark[data-note]")).toHaveAttribute("data-active", "true");
+});
+
+test("section, whole-article and citation notes; resolve, reopen, reply, filters", async ({ page }) => {
+ await page.goto(PAGE);
+ await page.getByRole("button", { name: "note on Later" }).click();
+ await page.getByLabel("note text").fill("Section note.");
+ await page.keyboard.press("Control+Enter");
+ await expect(page.locator("[data-note-card]")).toHaveCount(1);
+
+ await page.getByRole("button", { name: "+ whole article" }).click();
+ await page.getByLabel("note text").fill("Whole note.");
+ await page.getByRole("button", { name: "save note" }).click();
+ await expect(page.locator("[data-note-card]")).toHaveCount(2);
+
+ await page.locator("button[data-cite='c1']").first().click();
+ await expect(page.locator("[data-evidence='c1']")).toBeVisible();
+ await page.getByRole("button", { name: "+ note on this citation" }).click();
+ await page.getByLabel("note text").fill("Right second?");
+ await page.getByRole("button", { name: "save note" }).click();
+ await expect(page.locator("[data-evidence-rail] [data-note-card]")).toHaveCount(1);
+ await expect(page.locator("button[data-cite='c1'][data-has-note='true']").first()).toBeVisible();
+
+ const notes = JSON.parse(readFileSync(NOTES, "utf8")).notes;
+ expect(notes.map((n: { anchor: { kind: string } }) => n.anchor.kind).sort()).toEqual(["cite", "section", "whole"]);
+
+ // resolve the whole-article note; the open filter drops it
+ const whole = page.locator("[data-note-card]", { hasText: "Whole note." });
+ await whole.getByRole("button", { name: "resolve" }).click();
+ await expect(page.locator("[data-note-card]", { hasText: "Whole note." })).toHaveCount(0);
+ await page.getByRole("button", { name: /^resolved 1$/ }).click();
+ const resolved = page.locator("[data-note-card]", { hasText: "Whole note." });
+ await expect(resolved).toHaveAttribute("data-status", "resolved");
+ await resolved.getByRole("button", { name: "reopen" }).click();
+ await page.getByRole("button", { name: /^all 3$/ }).click();
+ const reopened = page.locator("[data-note-card]", { hasText: "Whole note." });
+ await expect(reopened).toHaveAttribute("data-status", "open");
+ await reopened.getByRole("button", { name: "reply" }).click();
+ await page.getByLabel("note text").fill("A reply.");
+ await reopened.getByRole("button", { name: "reply", exact: true }).click();
+ await expect(reopened.locator("[data-reply='operator']")).toContainText("A reply.");
+});
+
+test("keys: j selects, r resolves, n notes the selection", async ({ page }) => {
+ await page.goto(PAGE);
+ await select(page, "later", "on a different show");
+ await expect(page.getByRole("button", { name: "Note", exact: true })).toBeVisible();
+ await page.keyboard.press("n");
+ await page.getByLabel("note text").fill("Which show?");
+ await page.keyboard.press("Control+Enter");
+ await expect(page.locator("mark[data-note]")).toHaveText("on a different show");
+ await page.locator("main").click({ position: { x: 5, y: 5 } });
+ await page.keyboard.press("j");
+ await expect(page.locator("[data-note-card]")).toHaveAttribute("aria-current", "true");
+ await page.keyboard.press("r");
+ await expect(page.locator("[data-note-card]")).toHaveCount(0);
+ expect(JSON.parse(readFileSync(NOTES, "utf8")).notes[0].status).toBe("resolved");
+});
+
+test("a quote that is gone is still listed, flagged orphaned", async ({ page, request }) => {
+ const res = await request.post("/api/notes?article=priv/polemic-alpha", {
+ data: { op: { op: "add", text: "About a sentence the agent rewrote.", anchor: { kind: "text", section: "first", quote: "a sentence that is not in the article", prefix: "", suffix: "" } } },
+ });
+ expect(res.status()).toBe(200);
+ await page.goto(PAGE);
+ await expect(page.locator("[data-note-card][data-orphaned='true']")).toHaveCount(1);
+ await expect(page.locator("mark[data-note]")).toHaveCount(0);
+});
+
+test("a stale token is a 409 and the page re-reads; the last delete removes notes.json", async ({ page, request }) => {
+ await page.goto(PAGE);
+ await page.getByRole("button", { name: "+ whole article" }).click();
+ await page.getByLabel("note text").fill("Mine.");
+ await page.getByRole("button", { name: "save note" }).click();
+ await expect(page.locator("[data-note-card]")).toHaveCount(1);
+
+ // An agent writes in between (no token: the CLI's way).
+ const id = JSON.parse(readFileSync(NOTES, "utf8")).notes[0].id;
+ const stale = await request.post("/api/notes?article=priv/polemic-alpha", { data: { token: "absent", op: { op: "status", id, status: "resolved" } } });
+ expect(stale.status()).toBe(409);
+ await request.post("/api/notes?article=priv/polemic-alpha", { data: { op: { op: "reply", id, text: "seen" } } });
+
+ await page.locator("[data-note-card]").getByRole("button", { name: "resolve" }).click();
+ await expect(page.getByText(/reloaded; check and try again/)).toBeVisible();
+ await expect(page.locator("[data-note-card] [data-reply]")).toContainText("seen");
+
+ page.on("dialog", (d) => d.accept());
+ await page.locator("[data-note-card]").getByRole("button", { name: "delete" }).click();
+ await expect(page.locator("[data-note-card]")).toHaveCount(0);
+ expect(existsSync(NOTES)).toBe(false);
+});
+
+test("open notes are open decisions, linked to the note; ?note= opens on it", async ({ page, request }) => {
+ await request.post("/api/notes?article=priv/polemic-alpha", { data: { op: { op: "add", text: "For the inbox.", anchor: { kind: "section", section: "later" } } } });
+ const id = JSON.parse(readFileSync(NOTES, "utf8")).notes[0].id;
+
+ await page.goto("/sites?notes=open");
+ await expect(page.locator("[data-article='priv/polemic-alpha'] [data-open-notes]")).toHaveAttribute("data-open-notes", "1");
+
+ await page.goto("/browse/decisions?kind=open-note");
+ const row = page.locator("[data-project='sites/priv/polemic-alpha'] [data-decision='open-note']");
+ await expect(row).toContainText("For the inbox.");
+ await row.getByRole("link").click();
+ await expect(page).toHaveURL(new RegExp(`/sites/priv/polemic-alpha\\?note=${id}`));
+ await expect(page.locator(`[data-note-card='${id}']`)).toHaveAttribute("aria-current", "true");
+});
+
+test("writes outside a report are refused", async ({ request }) => {
+ for (const q of ["article=priv/../pub", "article=priv/nope", "article=priv", "article=../x/y"]) {
+ const r = await request.post(`/api/notes?${q}`, { data: { op: { op: "add", text: "x", anchor: { kind: "whole" } } } });
+ expect([400, 403, 404]).toContain(r.status());
+ }
+ const bad = await request.post("/api/notes?article=priv/polemic-alpha", { data: { op: { op: "add", text: "x", anchor: { kind: "moment", file: "../../x.mp4", t: 1 } } } });
+ expect(bad.status()).toBe(400);
+ expect(existsSync(NOTES)).toBe(false);
+});
diff --git a/umtool/e2e/fixtures/make-fixture.mjs b/umtool/e2e/fixtures/make-fixture.mjs
@@ -22,6 +22,7 @@ import { spawnSync } from "node:child_process";
import path from "node:path";
import { SONG_DATA } from "../../song/paths.mjs";
import { songCapabilities } from "./song-capabilities.mjs";
+import { makeSitesFixture } from "./sites-fixture.mjs";
const CODE = path.resolve(path.dirname(new URL(import.meta.url).pathname), "..", "..", "song");
const dest = path.resolve(process.argv[2] ?? path.join(process.cwd(), ".e2e-song"));
@@ -1774,7 +1775,118 @@ writeFileSync(
"# deliverables-fixture — share batch `first`\n\n- **d03_2025-03-03_A-Fixture-Stream.mp4**\n",
);
+// -- THE TAKES FIXTURE --------------------------------------------------------
+//
+// Alternative renders of one part of a cut, as another agent leaves them in
+// takes/<id>/. Its own project because the spec WRITES takes/verdicts.json.
+//
+// finale ref (reference, order 1), slow (similar, 2), hard-cut (different,
+// 3, its preview not rendered yet -- listed, with no player)
+// opening intro-a (similar, 5)
+// bad-kind a take.json whose kind is not one of the three -> skipped, with
+// the reason on the page
+// current/ a work directory with no take.json -> not a take, not listed
+//
+// report-fixture has no takes/ at all, which is the empty state.
+const TAKES = writeProject(
+ "takes-fixture",
+ manifest("takes-fixture", "The Takes Fixture", { siteOrigin: "https://archive.example" }, [
+ { type: "clip", id: "c01", video: "vid1", start: 0, end: 3, cite: 0, section: 0, lock: true, quote: "q" },
+ ]),
+);
+const take = (id, doc, { hz = null } = {}) => {
+ const dir = path.join(TAKES, "takes", id);
+ mkdirSync(dir, { recursive: true });
+ writeFileSync(path.join(dir, "take.json"), JSON.stringify({ id, preview: "preview.mp4", ...doc }, null, 2));
+ if (hz) {
+ ff([
+ "-f", "lavfi", "-i", "testsrc=size=320x180:rate=15:duration=2",
+ "-f", "lavfi", "-i", `sine=frequency=${hz}:duration=2`,
+ "-c:v", "libx264", "-pix_fmt", "yuv420p", "-c:a", "aac", "-shortest",
+ "-movflags", "+faststart",
+ path.join(dir, "preview.mp4"),
+ ]);
+ }
+};
+take("ref", { group: "finale", order: 1, label: "As built", kind: "reference", summary: "The current ending.", changes: [], seconds: 2 }, { hz: 330 });
+take("slow", { group: "finale", order: 2, label: "Slow burn", kind: "similar", summary: "Same words; slower beat.", changes: ["beat 1.35 s (was 1.05)", "dip 1.4 s (was 0.6)"], seconds: 2, builtAt: "2026-10-06T23:10:00Z" }, { hz: 440 });
+take("hard-cut", { group: "finale", order: 3, label: "Hard cut", kind: "different", summary: "No dip at all.", changes: ["no fade"] });
+take("intro-a", { group: "opening", order: 5, label: "Cold open", kind: "similar", summary: "Starts on the quote.", seconds: 2 }, { hz: 550 });
+take("bad-kind", { group: "finale", order: 4, label: "Bad", kind: "maybe" });
+mkdirSync(path.join(TAKES, "takes", "current", "out"), { recursive: true });
+
+// -- THE VIDEO-NOTES AND TIMELINE FIXTURES ------------------------------------
+//
+// video-notes-fixture is a GENERATED manifest (`generatedBy`) with a built cut
+// and its schedule, and a take with its own: timed notes resolve against them,
+// and every edit made to it leaves an `edit` note (video-notes.spec.ts).
+// timeline-fixture is hand-written, with a teaser, three clips, a post riding
+// on the second and the deck on: the structural edits' subject
+// (timeline-edit.spec.ts). Both are written by specs; nothing else reads them.
+const twoSeconds = (file, hz) =>
+ ff([
+ "-f", "lavfi", "-i", "testsrc=size=320x180:rate=15:duration=2",
+ "-f", "lavfi", "-i", `sine=frequency=${hz}:duration=2`,
+ "-c:v", "libx264", "-pix_fmt", "yuv420p", "-c:a", "aac", "-shortest",
+ "-movflags", "+faststart",
+ file,
+ ]);
+const NOTES_SCHEDULE = {
+ version: 1,
+ kind: "deck",
+ estimated: false,
+ fps: 15,
+ transition: 0,
+ total: 2,
+ segments: [
+ { id: "k1", type: "card", start: 0, duration: 0.5, end: 0.5, title: "Opening" },
+ { id: "n01", type: "clip", start: 0.5, duration: 0.7, end: 1.2, title: "The first claim" },
+ { id: "n02", type: "clip", start: 1.2, duration: 0.8, end: 2, title: "The second claim" },
+ ],
+};
+const VNOTES = writeProject("video-notes-fixture", {
+ ...manifest("video-notes-fixture", "The Video Notes Fixture", { siteOrigin: "https://archive.example" }, [
+ { type: "card", id: "k1", heading: "Opening" },
+ { type: "clip", id: "n01", video: "vid1", start: 0, end: 3, cite: 0, section: 0, lock: true, quote: "This is a complete sentence." },
+ { type: "clip", id: "n02", video: "vid1", start: 9, end: 12, cite: 9, section: 0, lock: true, quote: "Another whole sentence entirely." },
+ ]),
+ generatedBy: "polemics/video/make-videos.py",
+});
+{
+ // The deck on: the On-screen section shows the built cut (and its timed
+ // notes) only under it.
+ const m = JSON.parse(readFileSync(path.join(VNOTES, "video.manifest.json"), "utf8"));
+ m.render.chrome = { engine: "hyperframes", layout: "deck" };
+ writeFileSync(path.join(VNOTES, "video.manifest.json"), JSON.stringify(m, null, 2) + "\n");
+}
+mkdirSync(path.join(VNOTES, "out", "sourced"), { recursive: true });
+twoSeconds(path.join(VNOTES, "out", "video-notes-fixture.mp4"), 300);
+writeFileSync(path.join(VNOTES, "out", "sourced", "schedule.json"), JSON.stringify(NOTES_SCHEDULE, null, 2));
+{
+ const dir = path.join(VNOTES, "takes", "alt");
+ mkdirSync(path.join(dir, "out", "sourced"), { recursive: true });
+ writeFileSync(
+ path.join(dir, "take.json"),
+ JSON.stringify({ id: "alt", group: "cut", order: 1, label: "Alternate", kind: "similar", summary: "Tighter.", preview: "preview.mp4", seconds: 2 }, null, 2),
+ );
+ twoSeconds(path.join(dir, "preview.mp4"), 360);
+ writeFileSync(path.join(dir, "out", "sourced", "schedule.json"), JSON.stringify(NOTES_SCHEDULE, null, 2));
+}
+writeProject("timeline-fixture", {
+ ...manifest("timeline-fixture", "The Timeline Fixture", { siteOrigin: "https://archive.example" }, [
+ { type: "teaser", id: "t1", lines: ["THE PROMISE"] },
+ { type: "clip", id: "a01", video: "vid1", start: 0, end: 3, cite: 0, section: 1, sectionEnter: true, lock: true, quote: "one" },
+ { type: "clip", id: "a02", video: "vid1", start: 9, end: 12, cite: 9, section: 1, lock: true, quote: "two" },
+ { type: "clip", id: "a03", video: "vid1", start: 12, end: 15, cite: 12, section: 2, sectionEnter: true, lock: true, quote: "three" },
+ ]),
+ posts: [{ id: "p1", platform: "x", author: "Someone", handle: "@someone", date: "2024-01-02", text: "A post.", url: "https://x.com/someone/status/1", attachTo: "a02" }],
+});
+
+const { sites: SITES } = makeSitesFixture({ dest, reports, channels: CHANNELS });
+
console.log(`fixture at ${dest}`);
+console.log(` SITES_DIR=${SITES}`);
+console.log(` video-notes-fixture (generated, built, schedule + take alt), timeline-fixture (teaser, a01-a03, post p1)`);
if (planned) console.log(` planned clip (used in a build): ${planned}`);
console.log(` videos/: alpha (4 cuts, 3 variants), beta (2 cuts), deck (1 cut, 2 variants)`);
console.log(` deck: 1 spec error, 1 stale recipe, 1 unjudged variant, 1 judged one`);
@@ -1804,4 +1916,5 @@ console.log(` onscreen-fixture (writable, deck on, unbuilt), onscreen
console.log(` onscreen-posts-fixture (writable, deck on, three posts, unbuilt)`);
console.log(` onscreen-feed-fixture (writable, posts feed, built by the spec)`);
console.log(` deliver-stop-fixture (writable: six confirmed clips to cut, for Stop and resume)`);
+console.log(` takes-fixture (writable: takes/ with 4 takes over 2 groups, 1 skipped)`);
console.log(` ${taken} candidate files copied, 2 mix tracks synthesised`);
diff --git a/umtool/e2e/fixtures/sites-fixture.mjs b/umtool/e2e/fixtures/sites-fixture.mjs
@@ -0,0 +1,171 @@
+// The fixture's SITES_DIR (`<dest>/sites`), which the e2e server reads instead
+// of the real transcripts/sites (playwright.config.ts). The suite WRITES notes
+// there -- a report's notes.json is the one corpus file umtool writes -- so it
+// must never be the real one.
+//
+// Called by make-fixture.mjs after the projects and the channels exist, so a
+// report here can cite the fixture's channels and link its projects.
+//
+// sites/priv private, cited-only. Reports:
+// polemic-alpha PUBLISHED: summary, two sections, cites sitechan/sv1
+// (a clip window on disk: plays) and sitechan/sv2 (cues,
+// no media: the fetch_clip line); video.mp4 + poster.jpg
+// polemic-beta a DRAFT (not in site.json's reports)
+// sites/pub public. Reports:
+// gamma PUBLISHED, one section, no citations
+// <dest>/sitews/ the WORKSPACE (a direct child of REPORTS_ROOT, which the
+// e2e server takes as <dest>): polemics/drafts/{alpha,beta}.json,
+// polemics/make-site.py naming `priv`, polemics/out/alpha.{md,html},
+// NOTES.md
+// <dest>/sitews/polemic-alpha/
+// the article's report-video project: a manifest whose slug
+// is polemic-alpha, two takes, one verdict
+// channels/sitechan/data/{sv1,sv2} the cited records
+import { spawnSync } from "node:child_process";
+import { mkdirSync, writeFileSync } from "node:fs";
+import path from "node:path";
+
+const put = (file, value) => {
+ mkdirSync(path.dirname(file), { recursive: true });
+ writeFileSync(file, typeof value === "string" ? value : JSON.stringify(value, null, 2) + "\n");
+};
+
+const ff = (args) => {
+ const r = spawnSync("ffmpeg", ["-nostdin", "-v", "error", "-y", ...args], { encoding: "utf8" });
+ if (r.status !== 0) throw new Error(`ffmpeg failed: ${r.stderr || r.status}`);
+};
+
+const clip = (out, seconds, hz) => {
+ mkdirSync(path.dirname(out), { recursive: true });
+ ff([
+ "-f", "lavfi", "-i", `testsrc=size=320x180:rate=15:duration=${seconds}`,
+ "-f", "lavfi", "-i", `sine=frequency=${hz}:duration=${seconds}`,
+ "-c:v", "libx264", "-pix_fmt", "yuv420p", "-c:a", "aac", "-shortest", "-movflags", "+faststart", out,
+ ]);
+};
+
+export const SITE_CUES = {
+ sv1: [
+ [0, 3, "Opening line of the first source."],
+ [3, 6, "The vote was rigged and everybody knew."],
+ [6, 9, "Nobody checked it at the time."],
+ [9, 12, "Later the story changed again."],
+ [12, 15, "A fifth line for context."],
+ [15, 18, "And a sixth to close."],
+ ],
+ sv2: [
+ [0, 4, "The second source starts here."],
+ [4, 8, "She said it on a different show."],
+ [8, 12, "Then she said the opposite."],
+ ],
+};
+
+const ALPHA_BODY_1 =
+ "She said [the vote was rigged](cite:c1) in the first stream. Nobody checked the claim at the time, and the clip shows it.\n\nThe second paragraph repeats the claim for the record.";
+const ALPHA_BODY_2 = "Two years later [she said the opposite](cite:c2), on a different show.";
+
+const siteJson = (siteId, title, extra) => ({
+ siteId,
+ siteTitle: title,
+ groups: [{ id: "default", name: "All channels", selectedByDefault: true, order: 0, inline: true }],
+ defaultGroupId: "default",
+ channels: [{ slug: "sitechan", order: 0 }],
+ ...extra,
+});
+
+/**
+ * @param {{ dest: string, reports: string, channels: string }} at
+ */
+export function makeSitesFixture({ dest, channels }) {
+ const sites = path.join(dest, "sites");
+ mkdirSync(sites, { recursive: true });
+
+ // ---- the cited records ----------------------------------------------------
+ for (const [vid, rows] of Object.entries(SITE_CUES)) {
+ put(path.join(channels, "sitechan", "data", vid, "transcript.cues.json"), {
+ title: `Site source ${vid}`,
+ uploadDate: "20240315",
+ channel: "Site Channel",
+ webpageUrl: `https://www.youtube.com/watch?v=${vid}`,
+ duration: rows[rows.length - 1][1],
+ cues: rows.map(([start, end, text]) => ({ start, end, text })),
+ });
+ }
+ // sv1 has a clip window the editor fetched: [0, 12] holds the cited [3, 6].
+ clip(path.join(channels, "sitechan", "data", "sv1", "clips", "0.00-12.00.mp4"), 12, 523);
+
+ // ---- the sites ------------------------------------------------------------
+ put(path.join(sites, "priv", "site.json"), siteJson("priv", "Private Fixture", { audience: "private", search: false, reports: ["polemic-alpha"] }));
+ put(path.join(sites, "pub", "site.json"), siteJson("pub", "Public Fixture", { reports: ["gamma"] }));
+
+ const alphaDir = path.join(sites, "priv", "reports", "polemic-alpha");
+ put(path.join(alphaDir, "report.json"), {
+ format: "archilyzer-report",
+ version: 1,
+ id: "polemic-alpha",
+ kind: "sweep",
+ series: "Polemics",
+ title: "Alpha: the rigged vote",
+ subtitle: "What she said, and when",
+ summary: "She said one thing in 2020 and the opposite later.",
+ published: "2026-10-01",
+ updated: "2026-10-07",
+ video: { src: "video.mp4", poster: "poster.jpg" },
+ citations: {
+ c1: { kind: "video", channel: "sitechan", id: "sv1", start: 3, end: 6, quote: "The vote was rigged and everybody knew.", speaker: "Site Channel", date: "2024-03-15" },
+ c2: { kind: "video", channel: "sitechan", id: "sv2", start: 8, end: 12, quote: "Then she said the opposite.", speaker: "Site Channel", date: "2024-03-15" },
+ },
+ sections: [
+ { id: "first", title: "The first claim", body: ALPHA_BODY_1 },
+ { id: "later", title: "Later", body: ALPHA_BODY_2 },
+ ],
+ });
+ clip(path.join(alphaDir, "video.mp4"), 4, 440);
+ ff(["-f", "lavfi", "-i", "testsrc=size=320x180:rate=1:duration=1", "-frames:v", "1", path.join(alphaDir, "poster.jpg")]);
+
+ put(path.join(sites, "priv", "reports", "polemic-beta", "report.json"), {
+ format: "archilyzer-report",
+ version: 1,
+ id: "polemic-beta",
+ kind: "sweep",
+ title: "Beta: a draft",
+ summary: "A draft nobody has published.",
+ sections: [{ id: "only", title: "Only section", body: "The draft's only paragraph." }],
+ });
+
+ put(path.join(sites, "pub", "reports", "gamma", "report.json"), {
+ format: "archilyzer-report",
+ version: 1,
+ id: "gamma",
+ kind: "sweep",
+ title: "Gamma on the public site",
+ published: "2026-09-01",
+ sections: [{ id: "g", title: "Gamma", body: "A public article with nothing cited." }],
+ });
+
+ // ---- the workspace they were written in ------------------------------------
+ const ws = path.join(dest, "sitews");
+ put(path.join(ws, "polemics", "drafts", "alpha.json"), { id: "polemic-alpha", title: "Alpha: the rigged vote", sections: [] });
+ put(path.join(ws, "polemics", "drafts", "beta.json"), { id: "polemic-beta", title: "Beta: a draft", sections: [] });
+ put(path.join(ws, "polemics", "make-site.py"), 'SITE = "priv"\nfor f in DRAFTS.glob("drafts/*.json"):\n rid = f"polemic-{f.stem}"\n');
+ put(path.join(ws, "polemics", "out", "alpha.md"), "# Alpha\n\nThe rendered draft of **alpha**.\n");
+ put(path.join(ws, "polemics", "out", "alpha.html"), "<!doctype html><h1>Alpha html</h1><script>document.title='ran'</script>\n");
+ put(path.join(ws, "NOTES.md"), "# Workspace notes\n\n- one\n- two\n");
+
+ // ---- the article's video project -------------------------------------------
+ const proj = path.join(ws, "polemic-alpha");
+ put(path.join(proj, "video.manifest.json"), {
+ schemaVersion: 1,
+ slug: "polemic-alpha",
+ title: "Alpha: the rigged vote",
+ generatedBy: "polemics/video/make-videos.py",
+ provenance: { channelSlug: "sitechan", siteOrigin: "https://priv.example" },
+ timeline: [{ type: "clip", id: "a1", video: "sv1", channel: "sitechan", start: 3, end: 6, quote: "The vote was rigged" }],
+ });
+ for (const [id, order] of [["deck", 1], ["tight", 2]]) {
+ put(path.join(proj, "takes", id, "take.json"), { id, group: "cut", order, label: id, kind: order === 1 ? "reference" : "similar", preview: "preview.mp4" });
+ }
+ put(path.join(proj, "takes", "verdicts.json"), { deck: { verdict: "like", note: "", at: "2026-10-07T00:00:00Z" } });
+
+ return { sites, workspace: ws, project: proj };
+}
diff --git a/umtool/e2e/sites.spec.ts b/umtool/e2e/sites.spec.ts
@@ -0,0 +1,88 @@
+import { test, expect } from "@playwright/test";
+import { WARM_TIMEOUT, warm } from "./warm";
+
+// ---------------------------------------------------------------------------
+// /sites: every site's articles, a site's page, its media and workspace files.
+//
+// priv/polemic-alpha published, a video + poster, linked to the project
+// sitews/polemic-alpha (slug in the draft's workspace)
+// priv/polemic-beta a draft
+// pub/gamma published on the public site
+// (e2e/fixtures/sites-fixture.mjs)
+// ---------------------------------------------------------------------------
+
+
+test.beforeAll(async ({ playwright }) => {
+ test.setTimeout(WARM_TIMEOUT);
+ await warm(playwright, ["/sites", "/sites/priv", "/sites/priv/polemic-alpha", "/api/sites/media?site=priv&report=polemic-alpha&file=poster.jpg", "/api/sites/workspace?ws=sitews&rel=NOTES.md"]);
+});
+
+test("/sites lists every site, private first, with published and draft articles", async ({ page }) => {
+ const res = await page.goto("/sites");
+ expect(res?.status()).toBe(200);
+ await expect(page.getByRole("link", { name: "sites", exact: true }).first()).toHaveAttribute("aria-current", "page");
+
+ const sites = page.locator("[data-site]");
+ await expect(sites).toHaveCount(2);
+ await expect(sites.nth(0)).toHaveAttribute("data-site", "priv");
+ await expect(sites.nth(1)).toHaveAttribute("data-site", "pub");
+ await expect(page.locator("[data-site='priv'] [data-audience]")).toHaveAttribute("data-audience", "private");
+ await expect(page.locator("[data-site='priv'] [data-counts]")).toHaveAttribute("data-counts", "1/1");
+
+ const alpha = page.locator("[data-article='priv/polemic-alpha']");
+ await expect(alpha).toHaveAttribute("data-status", "published");
+ await expect(alpha).toContainText("Alpha: the rigged vote");
+ await expect(alpha.locator("[data-project-link='sitews/polemic-alpha']")).toBeVisible();
+ await expect(alpha).toContainText("alpha.json");
+ await expect(alpha.locator("img")).toHaveCount(1);
+ await expect(page.locator("[data-article='priv/polemic-beta']")).toHaveAttribute("data-status", "draft");
+});
+
+test("the filters are links, and compose", async ({ page }) => {
+ await page.goto("/sites");
+ await page.getByRole("link", { name: /^draft \d+$/ }).click();
+ await expect(page).toHaveURL(/status=draft/);
+ await expect(page.locator("[data-article]")).toHaveCount(1);
+ await expect(page.locator("[data-article='priv/polemic-beta']")).toBeVisible();
+
+ await page.goto("/sites?site=pub");
+ await expect(page.locator("[data-site]")).toHaveCount(1);
+ await expect(page.locator("[data-article='pub/gamma']")).toBeVisible();
+
+ await page.goto("/sites?notes=open");
+ await expect(page.locator("[data-article]")).toHaveCount(0);
+});
+
+test("a site's page plays its report videos, lists its project's takes and its workspace", async ({ page, request }) => {
+ await page.goto("/sites/priv");
+ await expect(page.locator("[data-report-video='polemic-alpha'] video")).toHaveAttribute("src", /file=video\.mp4/);
+ const project = page.locator("[data-video-project='sitews/polemic-alpha']");
+ await expect(project.locator("[data-takes]")).toHaveAttribute("data-takes", "2");
+ await expect(project).toContainText("like 1");
+
+ // media: ranged, read-only, and nothing outside the report dir
+ const v = await request.get("/api/sites/media?site=priv&report=polemic-alpha&file=video.mp4", { headers: { range: "bytes=0-99" } });
+ expect(v.status()).toBe(206);
+ expect(v.headers()["content-type"]).toBe("video/mp4");
+ expect((await request.get("/api/sites/media?site=priv&report=polemic-alpha&file=../polemic-beta/report.json")).status()).toBe(404);
+ expect((await request.get("/api/sites/media?site=priv&report=polemic-alpha&file=report.json")).status()).toBe(400);
+ expect((await request.get("/api/sites/media?corpus=/etc/passwd")).status()).toBe(400);
+
+ // workspace files: markdown renders, a draft folds, HTML is sandboxed
+ const files = page.locator("[data-section='files']");
+ await files.locator("[data-file='NOTES.md']").click();
+ await expect(page.locator("[data-opened='md']")).toContainText("Workspace notes");
+ await page.locator("[data-file='polemics/drafts/alpha.json']").click();
+ await expect(page.locator("[data-opened='json']")).toContainText("polemic-alpha");
+ await page.locator("[data-file='polemics/out/alpha.html']").click();
+ await expect(page.locator("iframe[sandbox='']")).toHaveCount(1);
+ const html = await request.get("/api/sites/workspace?ws=sitews&rel=polemics/out/alpha.html");
+ expect(html.headers()["content-security-policy"]).toContain("sandbox");
+ expect((await request.get("/api/sites/workspace?ws=sitews&rel=../sites/priv/site.json")).status()).toBe(404);
+ expect((await request.get("/api/sites/workspace?ws=..&rel=NOTES.md")).status()).toBe(404);
+});
+
+test("an unknown site or article is a 404", async ({ page }) => {
+ expect((await page.goto("/sites/nope"))?.status()).toBe(404);
+ expect((await page.goto("/sites/priv/nope"))?.status()).toBe(404);
+});
diff --git a/umtool/e2e/takes.spec.ts b/umtool/e2e/takes.spec.ts
@@ -0,0 +1,143 @@
+import { test, expect } from "@playwright/test";
+import { readFileSync } from "node:fs";
+import path from "node:path";
+import { fileURLToPath } from "node:url";
+
+// ---------------------------------------------------------------------------
+// Takes: alternative renders of part of a cut, side by side, judged.
+//
+// takes-fixture WRITES takes/verdicts.json. Two groups, a take with no
+// preview yet, one skipped take.json, a work dir that is
+// not a take (make-fixture.mjs says which is which)
+// report-fixture no takes/ at all -- the empty state
+// ---------------------------------------------------------------------------
+
+const HERE = path.dirname(fileURLToPath(import.meta.url));
+const PROJECT_DIR = path.join(HERE, "..", ".e2e-song", "reports", "takes-fixture");
+const VERDICTS = path.join(PROJECT_DIR, "takes", "verdicts.json");
+const PAGE = "/browse/reports/takes-fixture/takes";
+
+test("a project with no takes/ says so, and its page has no takes link", async ({ page }) => {
+ const res = await page.goto("/browse/reports/report-fixture/takes");
+ expect(res?.status()).toBe(200);
+ await expect(page.locator("[data-takes-empty]")).toContainText("No takes yet");
+
+ await page.goto("/browse/reports/report-fixture");
+ await expect(page.locator("[data-takes-link]")).toHaveCount(0);
+});
+
+test("takes are grouped and ordered; the skipped one says why; a work dir is not a take", async ({ page }) => {
+ await page.goto("/browse/reports/takes-fixture");
+ const link = page.locator("[data-takes-link]");
+ await expect(link).toHaveAttribute("data-takes-link", "4");
+ await link.click();
+ await expect(page).toHaveURL(new RegExp(`${PAGE}$`));
+
+ const groups = page.locator("[data-group]");
+ await expect(groups).toHaveCount(2);
+ await expect(groups.nth(0)).toHaveAttribute("data-group", "finale");
+ await expect(groups.nth(1)).toHaveAttribute("data-group", "opening");
+
+ const finale = page.locator("[data-group='finale'] [data-take]");
+ expect(await finale.evaluateAll((els) => els.map((e) => e.getAttribute("data-take")))).toEqual([
+ "ref",
+ "slow",
+ "hard-cut",
+ ]);
+ await expect(page.locator("[data-take='ref']")).toHaveAttribute("data-kind", "reference");
+ await expect(page.locator("[data-take='slow']")).toContainText("beat 1.35 s (was 1.05)");
+ await expect(page.locator("[data-take='hard-cut'] [data-no-preview]")).toBeVisible();
+ await expect(page.locator("[data-take='hard-cut'] video")).toHaveCount(0);
+
+ await expect(page.locator("[data-skipped='bad-kind']")).toContainText("kind must be one of");
+ await expect(page.locator("[data-skipped='current']")).toHaveCount(0);
+ await expect(page.locator("[data-take='current']")).toHaveCount(0);
+});
+
+test("a preview is served with byte ranges, and only for a listed take", async ({ request }) => {
+ const q = (take: string) => `/api/report/takes/preview?project=reports/takes-fixture&take=${take}`;
+ const whole = await request.get(q("ref"));
+ expect(whole.status()).toBe(200);
+ expect(whole.headers()["content-type"]).toBe("video/mp4");
+ const size = Number(whole.headers()["content-length"]);
+ expect(size).toBeGreaterThan(100);
+
+ const part = await request.get(q("ref"), { headers: { range: "bytes=0-99" } });
+ expect(part.status()).toBe(206);
+ expect(part.headers()["content-range"]).toBe(`bytes 0-99/${size}`);
+ expect((await part.body()).length).toBe(100);
+
+ expect((await request.get(q("bad-kind"))).status()).toBe(404);
+ expect((await request.get(q("current"))).status()).toBe(404);
+ expect((await request.get(q("hard-cut"))).status()).toBe(404);
+ expect((await request.get(q("..%2Fref"))).status()).toBe(400);
+});
+
+test("verdicts and notes land in takes/verdicts.json in the agreed shape", async ({ page, request }) => {
+ await page.goto(PAGE);
+ const slow = page.locator("[data-take='slow']");
+ const like = slow.locator("[data-verdict-button='like']");
+ await expect(like).toBeEnabled();
+ await like.click();
+ await expect(slow).toHaveAttribute("data-verdict", "like");
+
+ const note = slow.locator("[data-take-note]");
+ await expect(async () => {
+ await note.fill("the wait is right");
+ await expect(note).toHaveValue("the wait is right");
+ }).toPass();
+ await note.press("Enter");
+ await expect
+ .poll(() => {
+ try {
+ return JSON.parse(readFileSync(VERDICTS, "utf8")).slow?.note;
+ } catch {
+ return null;
+ }
+ })
+ .toBe("the wait is right");
+
+ const onDisk = JSON.parse(readFileSync(VERDICTS, "utf8"));
+ expect(Object.keys(onDisk.slow).sort()).toEqual(["at", "note", "verdict"]);
+ expect(onDisk.slow.verdict).toBe("like");
+
+ // Survives a reload; pressing the lit verdict again clears it but keeps the note.
+ await page.reload();
+ await expect(slow).toHaveAttribute("data-verdict", "like");
+ await expect(note).toHaveValue("the wait is right");
+ await slow.locator("[data-verdict-button='like']").click();
+ await expect(slow).toHaveAttribute("data-verdict", "");
+ await expect.poll(() => JSON.parse(readFileSync(VERDICTS, "utf8")).slow?.verdict).toBeNull();
+
+ // A verdict on a take that is not listed is refused, not written.
+ const bad = await request.post("/api/report/takes/verdict", {
+ data: { project: "reports/takes-fixture", take: "bad-kind", verdict: "like" },
+ });
+ expect(bad.status()).toBe(404);
+ expect(JSON.parse(readFileSync(VERDICTS, "utf8"))["bad-kind"]).toBeUndefined();
+});
+
+test("listen picks the one audio source; play all starts the group from zero", async ({ page }) => {
+ await page.goto(PAGE);
+ const muted = (id: string) => page.locator(`[data-take='${id}'] video`).evaluate((v: HTMLVideoElement) => v.muted);
+
+ await page.locator("[data-take='slow'] [data-action='listen']").click();
+ await expect(page.locator("[data-take='slow'] [data-action='listen']")).toHaveAttribute("aria-pressed", "true");
+ await expect.poll(() => muted("slow")).toBe(false);
+ expect(await muted("ref")).toBe(true);
+ expect(await muted("intro-a")).toBe(true);
+
+ await page.locator("[data-take='ref'] [data-action='listen']").click();
+ await expect.poll(() => muted("ref")).toBe(false);
+ expect(await muted("slow")).toBe(true);
+
+ await page.locator("[data-group='finale'] [data-action='play-all']").click();
+ const playing = (id: string) =>
+ page.locator(`[data-take='${id}'] video`).evaluate((v: HTMLVideoElement) => !v.paused && v.currentTime > 0);
+ await expect.poll(() => playing("ref")).toBe(true);
+ await expect.poll(() => playing("slow")).toBe(true);
+ // The other group was not started, and play all did not move the audio.
+ expect(await playing("intro-a")).toBe(false);
+ expect(await muted("ref")).toBe(false);
+ expect(await muted("slow")).toBe(true);
+});
diff --git a/umtool/e2e/timeline-edit.spec.ts b/umtool/e2e/timeline-edit.spec.ts
@@ -0,0 +1,136 @@
+import { test, expect, type Page } from "@playwright/test";
+import { copyFileSync, existsSync, readFileSync, readdirSync, rmSync, writeFileSync } from "node:fs";
+import path from "node:path";
+import { fileURLToPath } from "node:url";
+
+// ---------------------------------------------------------------------------
+// Structural edits to a report video (lib/report/manifest.mjs "STRUCTURE",
+// POST /api/report/timeline):
+//
+// timeline-fixture hand-written: teaser t1, clips a01 a02 a03, post p1 on
+// a02. WRITES its manifest and revisions/; every test
+// starts from the fixture's manifest.
+//
+// Re-order with alt+↓ and by dragging, undo, the row menu (duplicate, remove,
+// insert after), a refusal in the build's words, and the teaser, posts and
+// fact-check editors.
+// ---------------------------------------------------------------------------
+
+const HERE = path.dirname(fileURLToPath(import.meta.url));
+const DIR = path.join(HERE, "..", ".e2e-song", "reports", "timeline-fixture");
+const MANIFEST = path.join(DIR, "video.manifest.json");
+const PRISTINE = path.join(DIR, "video.manifest.pristine");
+const PAGE = "/browse/reports/timeline-fixture";
+
+type Entry = Record<string, unknown> & { id: string };
+const manifest = () => JSON.parse(readFileSync(MANIFEST, "utf8")) as { timeline: Entry[]; posts?: Entry[]; render: Record<string, unknown> };
+const order = () => manifest().timeline.map((e) => e.id);
+const rows = (page: Page) => page.locator("li[data-entry][data-kind]");
+const shownOrder = (page: Page) => rows(page).evaluateAll((els) => els.map((e) => e.getAttribute("data-entry")));
+
+test.beforeAll(() => {
+ if (!existsSync(PRISTINE)) copyFileSync(MANIFEST, PRISTINE);
+});
+test.beforeEach(() => {
+ copyFileSync(PRISTINE, MANIFEST);
+ rmSync(path.join(DIR, "revisions"), { recursive: true, force: true });
+ rmSync(path.join(DIR, "notes.json"), { force: true });
+});
+
+async function open(page: Page) {
+ await page.goto(PAGE);
+ await expect(page.getByTestId("timeline-undo")).toBeEnabled();
+ // The list has read its token once the page is hydrated.
+ await page.waitForLoadState("networkidle");
+}
+
+test("alt+↓ moves a row, recomputes sectionEnter, and Undo puts it back byte for byte", async ({ page }) => {
+ const before = readFileSync(MANIFEST, "utf8");
+ await open(page);
+ await page.locator("li[data-entry='a01'][data-kind]").focus();
+ await page.keyboard.press("Alt+ArrowDown");
+ await expect.poll(order).toEqual(["t1", "a02", "a01", "a03"]);
+ await expect.poll(() => shownOrder(page)).toEqual(["t1", "a02", "a01", "a03"]);
+ const m = manifest();
+ expect(m.timeline[1].sectionEnter).toBe(true);
+ expect("sectionEnter" in m.timeline[2]).toBe(false);
+ expect(readdirSync(path.join(DIR, "revisions")).some((n) => n.includes("auto-before-move"))).toBe(true);
+
+ await page.getByTestId("timeline-undo").click();
+ await expect.poll(() => readFileSync(MANIFEST, "utf8")).toBe(before);
+ await expect.poll(() => shownOrder(page)).toEqual(["t1", "a01", "a02", "a03"]);
+});
+
+test("a row dragged by its handle lands where it is dropped", async ({ page }) => {
+ await open(page);
+ await page.locator("[data-drag-handle='a03']").dragTo(page.locator("li[data-entry='a01'][data-kind]"));
+ await expect.poll(order).toEqual(["t1", "a03", "a01", "a02"]);
+});
+
+test("the row menu duplicates, removes, inserts — and a refusal is in the build's words", async ({ page }) => {
+ await open(page);
+ await page.locator("[data-row-menu='a03']").click();
+ await page.locator("[data-row-action='duplicate']").click();
+ await expect.poll(order).toEqual(["t1", "a01", "a02", "a03", "a03-copy"]);
+
+ await expect(page.locator("[data-row-menu='a03-copy']")).toBeVisible();
+ await page.locator("[data-row-menu='a03-copy']").click();
+ await page.locator("[data-row-action='remove']").click();
+ await expect.poll(order).toEqual(["t1", "a01", "a02", "a03"]);
+
+ await page.locator("[data-row-menu='a03']").click();
+ await page.locator("[data-row-action='insert']").click();
+ await page.getByTestId("insert-after-a03").fill("testchan/vid1@3-6");
+ await page.getByTestId("insert-after-a03").press("Enter");
+ await expect.poll(order).toEqual(["t1", "a01", "a02", "a03", "vid1-3"]);
+ expect(manifest().timeline[4]).toEqual({ type: "clip", id: "vid1-3", channel: "testchan", video: "vid1", start: 3, end: 6 });
+
+ // p1 rides on a02: removing it is refused, and nothing is written.
+ const was = readFileSync(MANIFEST, "utf8");
+ await page.locator("[data-row-menu='a02']").click();
+ await page.locator("[data-row-action='remove']").click();
+ await expect(page.getByTestId("timeline-error")).toContainText("attachTo");
+ expect(readFileSync(MANIFEST, "utf8")).toBe(was);
+});
+
+test("the teaser's lines and beat, edited in place; an empty teaser is refused", async ({ page }) => {
+ await open(page);
+ await page.getByTestId("structure-folded").locator("summary").click();
+ await page.getByTestId("teaser-lines-t1").fill("THE PROMISE\nAND WHAT HAPPENED");
+ await page.getByTestId("teaser-beat-t1").fill("1.2");
+ await page.getByTestId("teaser-save-t1").click();
+ await expect.poll(() => manifest().timeline[0].lines).toEqual(["THE PROMISE", "AND WHAT HAPPENED"]);
+ expect(manifest().timeline[0].beat).toBe(1.2);
+
+ await page.getByTestId("teaser-lines-t1").fill("");
+ await page.getByTestId("teaser-save-t1").click();
+ await expect(page.getByTestId("structure-error")).toContainText("lines must be a list");
+});
+
+test("a post added, then removed; the fact-check's labels once the deck is on", async ({ page }) => {
+ // The fact-check is drawn by the deck: turn it on in the fixture first.
+ const m = manifest();
+ m.render.chrome = { engine: "hyperframes", layout: "deck" };
+ writeFileSync(MANIFEST, JSON.stringify(m, null, 2) + "\n");
+
+ await open(page);
+ await page.getByTestId("structure-folded").locator("summary").click();
+ await page.getByTestId("post-add").click();
+ const fresh = page.locator("[data-post-row='new']");
+ await fresh.getByTestId("post-id").fill("p2");
+ await fresh.getByTestId("post-date").fill("2024-02-03");
+ await fresh.getByTestId("post-url").fill("https://x.com/someone/status/2");
+ await fresh.getByTestId("post-text").fill("Another post.");
+ await fresh.getByTestId("post-save").click();
+ await expect.poll(() => (manifest().posts ?? []).map((p) => p.id)).toEqual(["p1", "p2"]);
+
+ await page.locator("[data-post-row='p2']").getByTestId("post-remove").click();
+ await expect.poll(() => (manifest().posts ?? []).map((p) => p.id)).toEqual(["p1"]);
+
+ await page.getByTestId("fc-label-CONTRADICTED").fill("NOPE");
+ await page.getByTestId("fc-color-CONTRADICTED").fill("#ff0000");
+ await page.getByTestId("fc-save").click();
+ await expect
+ .poll(() => (manifest().render.chrome as { factcheck?: unknown }).factcheck)
+ .toEqual({ verdicts: { CONTRADICTED: { label: "NOPE", color: "#ff0000" } } });
+});
diff --git a/umtool/e2e/video-notes.spec.ts b/umtool/e2e/video-notes.spec.ts
@@ -0,0 +1,151 @@
+import { test, expect, type Locator, type Page } from "@playwright/test";
+import { copyFileSync, existsSync, readFileSync, rmSync } from "node:fs";
+import path from "node:path";
+import { fileURLToPath } from "node:url";
+
+// ---------------------------------------------------------------------------
+// Notes on a report video (lib/annotations, components/notes):
+//
+// video-notes-fixture a GENERATED manifest (`generatedBy`) with a built cut,
+// its schedule, and one take (`alt`) with its own. WRITES
+// its notes.json; every test starts from the fixture's
+// manifest and no notes.
+//
+// What is proved: a timed note at a second of the built cut and of a take's
+// preview, resolved to the entry on screen; take notes and row notes; and that
+// an edit made here to a generated manifest leaves an `edit` note, coalesced,
+// and gone again when the edit is put back.
+// ---------------------------------------------------------------------------
+
+const HERE = path.dirname(fileURLToPath(import.meta.url));
+const DIR = path.join(HERE, "..", ".e2e-song", "reports", "video-notes-fixture");
+const MANIFEST = path.join(DIR, "video.manifest.json");
+const PRISTINE = path.join(DIR, "video.manifest.pristine");
+const NOTES = path.join(DIR, "notes.json");
+const PROJECT = "reports/video-notes-fixture";
+const PAGE = `/browse/${PROJECT}`;
+
+type Note = { id: string; status: string; author: string; text: string; anchor: Record<string, unknown>; replies: unknown[] };
+const notesOnDisk = (): Note[] => (existsSync(NOTES) ? JSON.parse(readFileSync(NOTES, "utf8")).notes : []);
+
+test.beforeAll(() => {
+ if (!existsSync(PRISTINE)) copyFileSync(MANIFEST, PRISTINE);
+});
+test.beforeEach(() => {
+ copyFileSync(PRISTINE, MANIFEST);
+ rmSync(NOTES, { force: true });
+ rmSync(path.join(DIR, "revisions"), { recursive: true, force: true });
+});
+
+/** Seek a <video> to `t` and wait for it to land. */
+async function seek(video: Locator, t: number) {
+ await video.evaluate(async (el: HTMLVideoElement, at: number) => {
+ if (el.readyState < 1) await new Promise((r) => el.addEventListener("loadedmetadata", r, { once: true }));
+ el.currentTime = at;
+ await new Promise((r) => el.addEventListener("seeked", r, { once: true }));
+ }, t);
+}
+
+async function addTimedNote(page: Page, scope: Locator, text: string, via: "button" | "key") {
+ if (via === "button") await scope.locator("[data-action='mark']").click();
+ else {
+ await scope.locator("video").focus();
+ await page.keyboard.press("n");
+ }
+ const input = scope.getByTestId("timed-note-input");
+ await input.fill(text);
+ await input.press("Enter");
+}
+
+test("the generated banner is on the project page and the clip bench", async ({ page }) => {
+ await page.goto(PAGE);
+ await expect(page.getByTestId("generated-banner").first()).toContainText(
+ "Generated by polemics/video/make-videos.py; a rebuild of manifests overwrites edits made here.",
+ );
+ await page.goto(`${PAGE}/clip/n01`);
+ await expect(page.getByTestId("generated-banner")).toContainText("polemics/video/make-videos.py");
+});
+
+test("a timed note on the built cut resolves to the entry on screen and its source", async ({ page }) => {
+ await page.goto(PAGE);
+ const video = page.getByTestId("onscreen-final-video");
+ await expect(video).toBeVisible();
+ const scope = page.locator("[data-timed-notes='out/video-notes-fixture.mp4']");
+ await seek(video, 0.8);
+ await addTimedNote(page, scope, "the claim card is late", "button");
+
+ await expect(scope.locator("[data-mark-entry='n01']")).toContainText("the claim card is late");
+ await expect(scope.locator("[data-mark-entry='n01']")).toContainText("The first claim");
+ await expect(scope.locator("[data-tick]")).toHaveCount(1);
+
+ const [n] = notesOnDisk();
+ expect(n.author).toBe("operator");
+ expect(n.anchor).toMatchObject({ kind: "moment", file: "out/video-notes-fixture.mp4", t: 0.8, entry: "n01" });
+ expect(n.anchor.resolved).toMatchObject({ title: "The first claim", channel: "testchan", video: "vid1", sourceT: 0.3 });
+ expect(String((n.anchor.resolved as { url: string }).url)).toBe("https://archive.example/?v=testchan%2Fvid1&t=0");
+ expect((n.anchor.resolved as { approx?: boolean }).approx).toBeUndefined();
+
+ // Delete it: the last note takes the file with it.
+ await scope.locator(`[data-note-id='${n.id}'] [data-note-action='delete']`).click();
+ await expect(scope.locator("[data-mark-at]")).toHaveCount(0);
+ await expect.poll(() => existsSync(NOTES)).toBe(false);
+});
+
+test("a take: notes on the take, and a timed note on its preview with `n`", async ({ page }) => {
+ await page.goto(`${PAGE}/takes`);
+ const card = page.locator("[data-take='alt']");
+ await seek(card.locator("video"), 1.5);
+ await addTimedNote(page, card, "second claim runs long", "key");
+ await expect(card.locator("[data-mark-entry='n02']")).toContainText("second claim runs long");
+
+ await card.locator("[data-anchored-notes='alt'] [data-action='toggle-notes']").click();
+ await card.getByTestId("note-input-alt").fill("prefer this one, but tighter");
+ await card.getByTestId("note-input-alt").press("Enter");
+ await expect(card.locator("[data-anchored-notes='alt']")).toHaveAttribute("data-open-notes", "1");
+
+ const notes = notesOnDisk();
+ expect(notes.map((n) => n.anchor.kind).sort()).toEqual(["moment", "take"]);
+ expect(notes.find((n) => n.anchor.kind === "moment")!.anchor).toMatchObject({ file: "takes/alt/preview.mp4", take: "alt", entry: "n02" });
+ expect(notes.find((n) => n.anchor.kind === "take")!.anchor).toEqual({ kind: "take", take: "alt" });
+
+ // Resolve the take note; it stays, shown as resolved.
+ const takeNote = notes.find((n) => n.anchor.kind === "take")!;
+ await card.locator(`[data-note-id='${takeNote.id}'] [data-note-action='resolve']`).click();
+ await expect(card.locator(`[data-note-id='${takeNote.id}']`)).toHaveAttribute("data-note-status", "resolved");
+ await expect(card.locator("[data-anchored-notes='alt']")).toHaveAttribute("data-open-notes", "0");
+ expect(notesOnDisk().find((n) => n.id === takeNote.id)!.status).toBe("resolved");
+});
+
+test("an edit to a generated manifest leaves one edit note, coalesced, gone when put back", async ({ request }) => {
+ const token = async () => (await (await request.get(`/api/report/clip?project=${PROJECT}&clip=n01`)).json()).token as string;
+ const put = async (title: string) =>
+ request.put("/api/report/window", { data: { project: PROJECT, clip: "n01", token: await token(), title } });
+
+ let r = await (await put("A new title")).json();
+ expect(r.editNotes).toMatchObject({ generatedBy: "polemics/video/make-videos.py", added: 1 });
+ r = await (await put("A newer title")).json();
+ expect(r.editNotes).toMatchObject({ added: 0, updated: 1 });
+ const [n] = notesOnDisk();
+ expect(notesOnDisk()).toHaveLength(1);
+ expect(n.anchor).toEqual({ kind: "edit", entry: "n01", field: "title", from: null, to: "A newer title" });
+ expect(n.text).toContain("polemics/video/make-videos.py");
+
+ const doc = JSON.parse(readFileSync(NOTES, "utf8"));
+ expect(doc.source).toMatchObject({ manifest: expect.stringContaining("video-notes-fixture/video.manifest.json") });
+
+ r = await (await put("")).json();
+ expect(r.editNotes).toMatchObject({ deleted: 1 });
+ expect(existsSync(NOTES)).toBe(false);
+});
+
+test("a row note on the project page, counted on the row and kept across a reload", async ({ page }) => {
+ await page.goto(PAGE);
+ const row = page.locator("[data-anchored-notes='n02']");
+ await row.locator("[data-action='toggle-notes']").click();
+ await page.getByTestId("note-input-n02").fill("check the date on this one");
+ await page.getByTestId("note-input-n02").press("Enter");
+ await expect(row).toHaveAttribute("data-open-notes", "1");
+ await page.reload();
+ await expect(page.locator("[data-anchored-notes='n02']")).toHaveAttribute("data-open-notes", "1");
+ expect(notesOnDisk()[0].anchor).toEqual({ kind: "entry", entry: "n02" });
+});
diff --git a/umtool/e2e/warm.ts b/umtool/e2e/warm.ts
@@ -0,0 +1,17 @@
+import { test, type PlaywrightWorkerArgs } from "@playwright/test";
+
+// The e2e server is `next dev`: the first request to a route COMPILES it, and
+// the article page pulls in a large graph (common's report views, markdown,
+// the evidence resolver). On a loaded machine that first compile alone has
+// taken over 30 s -- a test's whole budget -- so each spec file that visits
+// these routes compiles them first, in a beforeAll with its own timeout.
+export const WARM_TIMEOUT = 180_000;
+
+export async function warm(playwright: PlaywrightWorkerArgs["playwright"], urls: string[]) {
+ const ctx = await playwright.request.newContext({ baseURL: test.info().project.use.baseURL });
+ try {
+ for (const u of urls) await ctx.get(u, { timeout: 170_000 });
+ } finally {
+ await ctx.dispose();
+ }
+}
diff --git a/umtool/lib/annotations/anchor.mjs b/umtool/lib/annotations/anchor.mjs
@@ -0,0 +1,176 @@
+// Finding a text anchor again in text that may have changed.
+//
+// A note on an article is anchored by its QUOTE plus 32 characters of context
+// either side (the W3C TextQuoteSelector), never by an offset: report.json is
+// regenerated from a draft, and an offset into the old text points at nothing
+// after the agent edits the paragraph above it.
+//
+// PURE and client-safe: no node imports. The article page runs it against the
+// rendered DOM text of a section; `umtool notes` runs it against the section's
+// plain text; the unit test runs it against both.
+
+export const CONTEXT = 32;
+
+/** Common-suffix length of `a` and `b` (how much of the prefix still precedes). */
+function suffixMatch(a, b) {
+ let n = 0;
+ while (n < a.length && n < b.length && a[a.length - 1 - n] === b[b.length - 1 - n]) n += 1;
+ return n;
+}
+/** Common-prefix length of `a` and `b` (how much of the suffix still follows). */
+function prefixMatch(a, b) {
+ let n = 0;
+ while (n < a.length && n < b.length && a[n] === b[n]) n += 1;
+ return n;
+}
+
+function allIndexes(hay, needle) {
+ const out = [];
+ if (!needle) return out;
+ for (let i = hay.indexOf(needle); i !== -1; i = hay.indexOf(needle, i + 1)) out.push(i);
+ return out;
+}
+
+/** Of several hits, the one whose surroundings best match prefix/suffix. Ties: the first. */
+function best(hay, hits, len, prefix, suffix) {
+ let top = hits[0];
+ let topScore = -1;
+ for (const i of hits) {
+ const score =
+ suffixMatch(hay.slice(Math.max(0, i - prefix.length), i), prefix) +
+ prefixMatch(hay.slice(i + len, i + len + suffix.length), suffix);
+ if (score > topScore) {
+ top = i;
+ topScore = score;
+ }
+ }
+ return top;
+}
+
+/**
+ * Whitespace collapsed (and typographic quotes/dashes folded), with a map from
+ * each normalised index back to the original one. `lower` also lowercases.
+ */
+export function normalise(text, { lower = false } = {}) {
+ const chars = [];
+ const map = [];
+ let space = false;
+ for (let i = 0; i < text.length; i += 1) {
+ let c = text[i];
+ if (/\s/.test(c)) {
+ if (space || chars.length === 0) continue;
+ space = true;
+ chars.push(" ");
+ map.push(i);
+ continue;
+ }
+ space = false;
+ if (c === "‘" || c === "’") c = "'";
+ else if (c === "“" || c === "”") c = '"';
+ else if (c === "–" || c === "—") c = "-";
+ else if (c === "…") c = ".";
+ if (lower) c = c.toLowerCase();
+ chars.push(c);
+ map.push(i);
+ }
+ if (chars[chars.length - 1] === " ") {
+ chars.pop();
+ map.pop();
+ }
+ map.push(text.length);
+ return { text: chars.join(""), map };
+}
+
+const norm = (s, lower) => normalise(s ?? "", { lower }).text;
+
+/**
+ * Locate a text anchor. `{ found: true, start, end, how }` with offsets into
+ * `text`, or `{ found: false }` -- an ORPHANED note, still shown, pinned to its
+ * section. `how`: "exact", "normalised" (whitespace/quotes/case differ), or
+ * "context" (the quote itself was edited, but what came before and after it is
+ * still there, close together).
+ *
+ * @param {string} text
+ * @param {{ quote: string, prefix?: string, suffix?: string }} anchor
+ */
+export function locateQuote(text, anchor) {
+ const quote = anchor?.quote ?? "";
+ const prefix = anchor?.prefix ?? "";
+ const suffix = anchor?.suffix ?? "";
+ if (!text || !quote) return { found: false };
+
+ const exact = allIndexes(text, quote);
+ if (exact.length) {
+ const i = best(text, exact, quote.length, prefix, suffix);
+ return { found: true, start: i, end: i + quote.length, how: "exact" };
+ }
+
+ for (const lower of [false, true]) {
+ const n = normalise(text, { lower });
+ const q = norm(quote, lower);
+ if (!q) continue;
+ const hits = allIndexes(n.text, q);
+ if (hits.length) {
+ const i = best(n.text, hits, q.length, norm(prefix, lower), norm(suffix, lower));
+ return { found: true, start: n.map[i], end: n.map[i + q.length - 1] + 1, how: "normalised" };
+ }
+ }
+
+ // The quote was rewritten. If the context on BOTH sides survives, close
+ // together, the span between them is where it was.
+ const n = normalise(text, { lower: true });
+ const p = norm(prefix, true);
+ const s = norm(suffix, true);
+ if (p.length >= 8 && s.length >= 8) {
+ const q = norm(quote, true);
+ for (const pi of allIndexes(n.text, p)) {
+ const from = pi + p.length;
+ const si = n.text.indexOf(s, from);
+ if (si === -1) continue;
+ const span = si - from;
+ if (span <= 0 || span > Math.max(q.length * 2, q.length + 80)) continue;
+ // Trim the separator spaces the normalised text keeps around the span.
+ let a = from;
+ let b = si;
+ while (a < b && n.text[a] === " ") a += 1;
+ while (b > a && n.text[b - 1] === " ") b -= 1;
+ if (a >= b) continue;
+ return { found: true, start: n.map[a], end: n.map[b - 1] + 1, how: "context" };
+ }
+ }
+ return { found: false };
+}
+
+/**
+ * The anchor for a selection `[start, end)` of `text`: the quote (trimmed of
+ * surrounding whitespace) and up to CONTEXT characters either side.
+ *
+ * @param {string} text
+ * @param {number} start
+ * @param {number} end
+ * @param {number} [context]
+ */
+export function quoteAnchor(text, start, end, context = CONTEXT) {
+ let a = Math.max(0, Math.min(start, end));
+ let b = Math.min(text.length, Math.max(start, end));
+ while (a < b && /\s/.test(text[a])) a += 1;
+ while (b > a && /\s/.test(text[b - 1])) b -= 1;
+ return {
+ quote: text.slice(a, b),
+ prefix: text.slice(Math.max(0, a - context), a),
+ suffix: text.slice(b, b + context),
+ };
+}
+
+/** The sentence of `text` around `[start, end)`, for a reader with no page open. */
+export function sentenceAround(text, start, end, max = 400) {
+ const before = text.slice(0, start);
+ const after = text.slice(end);
+ const boundary = /[.?!]\s+|\n/g;
+ let from = 0;
+ for (let m = boundary.exec(before); m; m = boundary.exec(before)) from = m.index + m[0].length;
+ const m = after.search(/[.?!](\s|$)|\n/);
+ const to = m === -1 ? text.length : end + m + (after[m] === "\n" ? 0 : 1);
+ const out = text.slice(from, to).replace(/\s+/g, " ").trim();
+ return out.length > max ? `${out.slice(0, max - 1)}…` : out;
+}
diff --git a/umtool/lib/annotations/anchor.test.mjs b/umtool/lib/annotations/anchor.test.mjs
@@ -0,0 +1,65 @@
+// Re-anchoring a quote in text that changed.
+//
+// Run with: pnpm test:scripts
+import assert from "node:assert/strict";
+import test from "node:test";
+import { locateQuote, normalise, quoteAnchor, sentenceAround } from "./anchor.mjs";
+
+const P1 = "She said the vote was rigged in 2020. Nobody checked the claim at the time.";
+const P2 = "Two years later she said the vote was rigged again, on a different show.";
+const TEXT = `${P1}\n\n${P2}`;
+
+test("an exact quote is found, and the context picks between two copies", () => {
+ const second = TEXT.indexOf("the vote was rigged", P1.length);
+ const a = quoteAnchor(TEXT, second, second + "the vote was rigged".length);
+ assert.equal(a.quote, "the vote was rigged");
+ const r = locateQuote(TEXT, a);
+ assert.deepEqual([r.found, r.start, r.how], [true, second, "exact"]);
+ const first = locateQuote(TEXT, quoteAnchor(TEXT, TEXT.indexOf("the vote"), TEXT.indexOf("the vote") + 19));
+ assert.equal(first.start, TEXT.indexOf("the vote"));
+});
+
+test("a paragraph edited ABOVE the quote does not move the note off it", () => {
+ const start = TEXT.indexOf("on a different show");
+ const a = quoteAnchor(TEXT, start, start + "on a different show".length);
+ const edited = `A new opening paragraph the agent added.\n\n${P1.replace("Nobody checked", "No outlet checked")}\n\n${P2}`;
+ const r = locateQuote(edited, a);
+ assert.equal(r.found, true);
+ assert.equal(edited.slice(r.start, r.end), "on a different show");
+});
+
+test("whitespace, typographic quotes and case still match (normalised)", () => {
+ const a = { quote: "it's the\nclaim", prefix: "", suffix: "" };
+ const t = "And then: It’s the claim, again.";
+ const r = locateQuote(t, a);
+ assert.equal(r.found, true);
+ assert.equal(r.how, "normalised");
+ assert.equal(t.slice(r.start, r.end), "It’s the claim");
+});
+
+test("a rewritten quote is found by its surviving context; a vanished one is orphaned", () => {
+ const start = TEXT.indexOf("Nobody checked the claim");
+ const a = quoteAnchor(TEXT, start, start + "Nobody checked the claim".length);
+ const rewritten = TEXT.replace("Nobody checked the claim", "No outlet verified it");
+ const r = locateQuote(rewritten, a);
+ assert.equal(r.found, true);
+ assert.equal(r.how, "context");
+ assert.equal(rewritten.slice(r.start, r.end), "No outlet verified it");
+
+ const gone = locateQuote("A completely different article.", a);
+ assert.deepEqual(gone, { found: false });
+ assert.deepEqual(locateQuote("", a), { found: false });
+});
+
+test("normalise maps back to the original offsets", () => {
+ const n = normalise(" a \n\n b—c ");
+ assert.equal(n.text, "a b-c");
+ assert.deepEqual(n.map.slice(0, 5), [2, 3, 7, 8, 9]);
+});
+
+test("sentenceAround gives the sentence holding the quote", () => {
+ const s = TEXT.indexOf("Nobody checked");
+ assert.equal(sentenceAround(TEXT, s, s + 6), "Nobody checked the claim at the time.");
+ const f = TEXT.indexOf("Two years");
+ assert.equal(sentenceAround(TEXT, f, f + 3), P2);
+});
diff --git a/umtool/lib/annotations/cli.mjs b/umtool/lib/annotations/cli.mjs
@@ -0,0 +1,140 @@
+// `umtool notes` -- the agent's side of the notes the operator writes in the
+// app. It reads and writes through the same store (./store.mjs) and targets
+// (./targets.mjs) the app does, and stamps every write `author: agent`. An
+// agent never hand-edits notes.json: this validates, locks, and keeps the
+// operator's page from losing a reply (docs/notes.md).
+//
+// umtool notes [--all] every notes file, with open counts
+// umtool notes <site>/<report> | <project> the digest (open notes)
+// [--open | --resolved | --all-status] [--json]
+// umtool notes reply <id> "<text>" [--resolve] [--in <target>]
+// umtool notes resolve | wontfix | reopen <id> [--in <target>]
+// umtool notes source <target> [--draft P] [--generator P] [--how T]
+import { REPORTS_ROOT, SITES_DIR } from "../paths.mjs";
+import { digest } from "./digest.mjs";
+import { readNotes } from "./store.mjs";
+import { articleTarget, listNotesFiles, projectTarget, resolveTarget, writeNote } from "./targets.mjs";
+
+const USAGE = [
+ "usage: umtool notes [--all] every notes file and its open count",
+ " umtool notes <site>/<report> | <project> the notes, as markdown (open ones)",
+ " [--open | --resolved | --all-status] [--json]",
+ " umtool notes reply <id> \"<text>\" [--resolve] answer a note (and close it)",
+ " umtool notes resolve | wontfix | reopen <id>",
+ " umtool notes source <target> [--draft P] [--generator P] [--how T]",
+ "",
+ "Notes are the operator's, written in umtool (/sites, a video project). Act on one",
+ "by editing its SOURCE file and regenerating; never edit notes.json by hand.",
+].join("\n");
+
+class CliError extends Error {}
+
+/**
+ * @param {string[]} args everything after `notes`
+ * @param {{ sitesDir?: string, reportsRoot?: string, log?: (s: string) => void }} [opts]
+ * @returns {Promise<number>} exit code
+ */
+export async function notesCommand(args, { sitesDir = SITES_DIR, reportsRoot = REPORTS_ROOT, log = console.log } = {}) {
+ const flags = new Set(args.filter((a) => a.startsWith("--")));
+ const val = (n) => {
+ const i = args.indexOf(n);
+ return i >= 0 && i + 1 < args.length ? args[i + 1] : undefined;
+ };
+ const VALUED = new Set(["--in", "--draft", "--generator", "--how"]);
+ const pos = args.filter((a, i) => !a.startsWith("--") && !(i > 0 && VALUED.has(args[i - 1])));
+ const json = flags.has("--json");
+ const opts = { sitesDir, reportsRoot };
+ const print = (v) => log(json ? JSON.stringify(v, null, 2) : v);
+
+ try {
+ const verb = pos[0];
+ if (flags.has("--help") || verb === "help") {
+ log(USAGE);
+ return 0;
+ }
+ if (verb === "reply" || verb === "resolve" || verb === "wontfix" || verb === "reopen") {
+ const id = pos[1];
+ if (!id) throw new CliError(`which note? \`umtool notes ${verb} <id>\``);
+ const where = await findNote(id, val("--in"), opts);
+ let op;
+ if (verb === "reply") {
+ const text = pos[2];
+ if (!text) throw new CliError('what reply? `umtool notes reply <id> "<text>" [--resolve]`');
+ op = { op: "reply", id, text, resolve: flags.has("--resolve") };
+ } else {
+ op = { op: "status", id, status: verb === "reopen" ? "open" : verb === "wontfix" ? "wontfix" : "resolved" };
+ }
+ const r = await writeNote(where.target, op, { by: "agent" });
+ if (json) print({ ok: true, target: where.target.id, file: where.target.file, note: r.note });
+ else log(`${id} in ${where.target.id}: ${r.note?.status}${verb === "reply" ? `, ${r.note?.replies.length} repl${r.note?.replies.length === 1 ? "y" : "ies"}` : ""}`);
+ return 0;
+ }
+ if (verb === "source") {
+ const spec = pos[1];
+ if (!spec) throw new CliError("which target? `umtool notes source <site>/<report> --draft P`");
+ const target = await resolveTarget(spec, opts);
+ const cur = await readNotes(target.file);
+ if (!cur.doc) throw new CliError(`${target.id} has no notes; the source is recorded with the first note`);
+ const source = { ...(cur.doc.source ?? {}) };
+ for (const k of ["draft", "generator", "how"]) if (val(`--${k}`) !== undefined) source[k] = val(`--${k}`);
+ const r = await writeNote(target, { op: "source", source }, { by: "agent" });
+ print(json ? { ok: true, source: r.doc?.source ?? null } : `${target.id}: source ${JSON.stringify(r.doc?.source ?? {})}`);
+ return 0;
+ }
+
+ const status = flags.has("--all-status") ? "all" : flags.has("--resolved") ? "resolved" : "open";
+ if (!verb || flags.has("--all")) {
+ const files = await listNotesFiles(opts);
+ if (json) {
+ print(files.map((f) => ({ kind: f.kind, id: f.id, file: f.file, open: f.doc ? f.doc.notes.filter((n) => n.status === "open").length : null, total: f.doc?.notes.length ?? null, error: f.error })));
+ return 0;
+ }
+ if (!files.length) {
+ log(`no notes under ${sitesDir} or ${reportsRoot}`);
+ return 0;
+ }
+ const w = Math.max(...files.map((f) => f.id.length));
+ let open = 0;
+ for (const f of files) {
+ const o = f.doc ? f.doc.notes.filter((n) => n.status === "open").length : 0;
+ open += o;
+ if (status === "open" && !o && !f.error) continue;
+ log(`${f.id.padEnd(w)} ${f.kind === "article" ? "article" : "video "} ${f.error ? `UNREADABLE: ${f.error}` : `${o} open / ${f.doc.notes.length}`}`);
+ }
+ log(`\n${open} open note(s) in ${files.length} file(s). \`umtool notes <id>\` for one.`);
+ return 0;
+ }
+
+ const target = await resolveTarget(verb, opts);
+ const read = await readNotes(target.file);
+ const entry = { kind: target.kind, id: target.id, file: target.file, doc: read.doc, ...(read.error ? { error: read.error } : {}) };
+ if (json) {
+ print({ ...entry, source: read.doc?.source ?? (await target.source()) ?? null });
+ return 0;
+ }
+ if (!read.doc && !read.error) {
+ log(`no notes on ${target.id} (${target.file})`);
+ return 0;
+ }
+ log(await digest(entry, { status, sitesDir, reportsRoot, projectDir: target.dir }));
+ return 0;
+ } catch (err) {
+ console.error(err instanceof Error ? err.message : String(err));
+ return err instanceof CliError || err?.name === "NoteError" || err?.name === "TargetError" ? 2 : 1;
+ }
+}
+
+/** The notes file holding note `id`: the one named by --in, else a search of every file. */
+async function findNote(id, inSpec, opts) {
+ if (inSpec) {
+ const target = await resolveTarget(inSpec, opts);
+ const r = await readNotes(target.file);
+ if (!r.doc?.notes.some((n) => n.id === id)) throw new CliError(`no note ${id} in ${target.id}`);
+ return { target };
+ }
+ const hits = (await listNotesFiles(opts)).filter((f) => f.doc?.notes.some((n) => n.id === id));
+ if (!hits.length) throw new CliError(`no note ${id} (see \`umtool notes --all\`)`);
+ if (hits.length > 1) throw new CliError(`${id} is in ${hits.length} files; say which with --in: ${hits.map((h) => h.id).join(", ")}`);
+ const hit = hits[0];
+ return { target: hit.kind === "article" ? await articleTarget(hit.id, opts) : await projectTarget(hit.id, opts) };
+}
diff --git a/umtool/lib/annotations/cli.test.mjs b/umtool/lib/annotations/cli.test.mjs
@@ -0,0 +1,103 @@
+// `umtool notes`: the agent's loop, end to end, against a temp SITES_DIR and
+// REPORTS_DIR -- read the digest, reply and resolve, reopen, list.
+//
+// Run with: pnpm test:scripts
+import assert from "node:assert/strict";
+import { mkdir, mkdtemp, readFile, rm, writeFile } from "node:fs/promises";
+import { tmpdir } from "node:os";
+import path from "node:path";
+import test from "node:test";
+import { notesCommand } from "./cli.mjs";
+import { articleTarget, projectTarget, writeNote } from "./targets.mjs";
+import { clearSourcesCache } from "../articles/sources.mjs";
+
+const REPORT = {
+ format: "archilyzer-report",
+ version: 1,
+ id: "polemic-x",
+ kind: "sweep",
+ title: "On X",
+ summary: "She said **X** twice. Then she [denied it](cite:c1).",
+ citations: { c1: { kind: "video", channel: "ch", id: "v1", start: 10, end: 20, quote: "I never said X", speaker: "Her" } },
+ sections: [{ id: "s1", title: "The first time", body: "In 2019 she said X on a podcast. Nobody noticed." }],
+};
+
+async function fixture() {
+ const root = await mkdtemp(path.join(tmpdir(), "umtool-notes-cli-"));
+ const sites = path.join(root, "sites");
+ const reports = path.join(root, "reports");
+ await mkdir(path.join(sites, "priv", "reports", "polemic-x"), { recursive: true });
+ await writeFile(path.join(sites, "priv", "site.json"), "{}");
+ await writeFile(path.join(sites, "priv", "reports", "polemic-x", "report.json"), JSON.stringify(REPORT));
+ await mkdir(path.join(reports, "ws", "polemics", "drafts"), { recursive: true });
+ await writeFile(path.join(reports, "ws", "polemics", "drafts", "x.json"), JSON.stringify({ id: "polemic-x" }));
+ await writeFile(path.join(reports, "ws", "polemics", "make-site.py"), 'SITE = "priv" # drafts\n');
+ const proj = path.join(reports, "ws", "polemic-x");
+ await mkdir(proj, { recursive: true });
+ await writeFile(
+ path.join(proj, "video.manifest.json"),
+ JSON.stringify({ schemaVersion: 1, slug: "polemic-x", generatedBy: "polemics/make-site.py", timeline: [{ type: "clip", id: "e1", channel: "ch", video: "v1", start: 10, end: 20, quote: "I never said X", onscreen: { title: "Denial" } }] }),
+ );
+ clearSourcesCache();
+ return { root, sites, reports, opts: { sitesDir: sites, reportsRoot: reports } };
+}
+
+async function run(args, opts) {
+ const out = [];
+ const code = await notesCommand(args, { ...opts, log: (s) => out.push(String(s)) });
+ return { code, text: out.join("\n") };
+}
+
+test("the agent loop: digest, reply --resolve, reopen; every write is the agent's", async () => {
+ const { root, sites, opts } = await fixture();
+ const t = await articleTarget("priv/polemic-x", opts);
+ const a = await writeNote(t, { op: "add", text: "Too strong; say 'claimed'.", anchor: { kind: "text", section: "s1", quote: "she said X", prefix: "In 2019 ", suffix: " on a podcast" } }, { by: "operator" });
+ await writeNote(t, { op: "add", text: "Is this the right clip?", anchor: { kind: "cite", cite: "c1" } }, { by: "operator" });
+
+ const d = await run(["priv/polemic-x"], opts);
+ assert.equal(d.code, 0);
+ assert.match(d.text, /# Notes on priv\/polemic-x — On X/);
+ assert.match(d.text, /edit `.*ws\/polemics\/drafts\/x\.json`; `report\.json` is regenerated by `.*make-site\.py`/);
+ assert.match(d.text, /“she said X” in “The first time”/);
+ assert.match(d.text, /> In 2019 she said X on a podcast\./);
+ assert.match(d.text, /citation `c1` — ch\/v1@10-20 \(Her\)/);
+ assert.match(d.text, /2 open, 0 closed/);
+
+ const r = await run(["reply", a.note.id, "Changed to 'claimed' in drafts/x.json", "--resolve"], opts);
+ assert.equal(r.code, 0, r.text);
+ const doc = JSON.parse(await readFile(path.join(sites, "priv", "reports", "polemic-x", "notes.json"), "utf8"));
+ const n = doc.notes.find((x) => x.id === a.note.id);
+ assert.equal(n.status, "resolved");
+ assert.equal(n.resolvedBy, "agent");
+ assert.deepEqual(n.replies.map((x) => x.author), ["agent"]);
+
+ assert.match((await run(["priv/polemic-x"], opts)).text, /1 open, 1 closed/);
+ assert.match((await run(["priv/polemic-x", "--resolved"], opts)).text, /Changed to 'claimed'/);
+ assert.equal((await run(["reopen", a.note.id], opts)).code, 0);
+ const list = await run(["--all"], opts);
+ assert.match(list.text, /priv\/polemic-x\s+article\s+2 open \/ 2/);
+
+ // refusals exit 2 and write nothing
+ assert.equal((await run(["reply", "n_nosuchnote"], opts)).code, 2);
+ assert.equal((await run(["reply", a.note.id], opts)).code, 2);
+ assert.equal((await run(["priv/../etc"], opts)).code, 2);
+ await rm(root, { recursive: true });
+});
+
+test("a video project's digest names the generator and lists the takes' verdicts", async () => {
+ const { root, reports, opts } = await fixture();
+ const proj = path.join(reports, "ws", "polemic-x");
+ await mkdir(path.join(proj, "takes", "deck"), { recursive: true });
+ await writeFile(path.join(proj, "takes", "deck", "take.json"), JSON.stringify({ id: "deck", group: "open", order: 1, label: "Deck first", kind: "similar", preview: "preview.mp4" }));
+ await writeFile(path.join(proj, "takes", "verdicts.json"), JSON.stringify({ deck: { verdict: "like", note: "keep the beat", at: "x" } }));
+ const t = await projectTarget("ws/polemic-x", opts);
+ await writeNote(t, { op: "add", text: "Cut this.", anchor: { kind: "entry", entry: "e1" } }, { by: "operator" });
+ await writeNote(t, { op: "add", text: "quote changed", anchor: { kind: "edit", entry: "e1", field: "quote", from: "a", to: "b" } }, { by: "operator" });
+ const d = await run(["ws/polemic-x"], opts);
+ assert.equal(d.code, 0, d.text);
+ assert.match(d.text, /video\.manifest\.json` is regenerated by `.*ws\/polemics\/make-site\.py`/);
+ assert.match(d.text, /timeline entry `e1` \(clip\) “Denial”/);
+ assert.match(d.text, /`e1\.quote`: "a" → "b"/);
+ assert.match(d.text, /- `deck` “Deck first” \(open, similar\): like — keep the beat/);
+ await rm(root, { recursive: true });
+});
diff --git a/umtool/lib/annotations/digest.mjs b/umtool/lib/annotations/digest.mjs
@@ -0,0 +1,207 @@
+// Notes as an AGENT reads them: markdown, every anchor resolved to something a
+// reader with no page open can act on, and the file to edit named first.
+//
+// `umtool notes <target>` prints this; GET /api/notes/context serves the same
+// text, so "Copy agent brief" on a page and an agent's CLI hand over the same
+// words. An anchor is resolved against what is on disk NOW:
+//
+// text the section's title and the sentence holding the quote, found
+// again with the same re-anchoring the page uses (ORPHANED when the
+// quote is gone -- the note still prints, with its quote)
+// cite the citation's quote, speaker, date and `<channel>/<id>@start-end`
+// moment the time, and the entry/source it resolved to when it was written
+// entry the timeline entry's title and quote
+// take the take's label and summary, and the operator's verdict on it
+// edit what changed, from → to, to port into the generator's inputs
+//
+// A video project's digest also lists every take with its verdict and note
+// (takes/verdicts.json): the agent that rendered the takes reads them here.
+import { readFile } from "node:fs/promises";
+import path from "node:path";
+import { REPORTS_ROOT, SITES_DIR } from "../paths.mjs";
+import { listTakes, readVerdicts } from "../report/takes.mjs";
+import { locateQuote, sentenceAround } from "./anchor.mjs";
+
+const readJson = (file) => readFile(/* turbopackIgnore: true */ file, "utf8").then(JSON.parse, () => null);
+
+/** Markdown to the plain text a reader sees: links to their labels, emphasis and code marks dropped. */
+export function plainText(md) {
+ return String(md ?? "")
+ .replace(/!\[([^\]]*)\]\([^)]*\)/g, "$1")
+ .replace(/\[([^\]]*)\]\([^)]*\)/g, "$1")
+ .replace(/^#{1,6}\s+/gm, "")
+ .replace(/^\s*>\s?/gm, "")
+ .replace(/^\s*[-*+]\s+/gm, "")
+ .replace(/(\*\*|__|\*|_|`)/g, "");
+}
+
+/**
+ * The plain text of one block of a report, as the page renders it: a section's
+ * title, body and claims; or the title/subtitle/summary/method.
+ */
+export function blockText(report, block) {
+ if (!report) return { title: block, text: "" };
+ if (["title", "subtitle", "summary", "method"].includes(block)) {
+ return { title: block, text: plainText(report[block] ?? "") };
+ }
+ const s = (report.sections ?? []).find((x) => x.id === block);
+ if (!s) return null;
+ const parts = [s.title, plainText(s.body ?? "")];
+ for (const c of s.claims ?? []) parts.push(c.title ?? "", plainText(c.text), plainText(c.findings ?? ""));
+ return { title: s.title, text: parts.filter(Boolean).join("\n") };
+}
+
+const clip = (s, n = 300) => {
+ const t = String(s ?? "").replace(/\s+/g, " ").trim();
+ return t.length > n ? `${t.slice(0, n - 1)}…` : t;
+};
+const quoteLine = (s) => `> ${clip(s, 400)}`;
+const hms = (t) => {
+ const s = Math.max(0, Math.round(Number(t) || 0));
+ const h = Math.floor(s / 3600);
+ const m = Math.floor((s % 3600) / 60);
+ const ss = String(s % 60).padStart(2, "0");
+ return h ? `${h}:${String(m).padStart(2, "0")}:${ss}` : `${m}:${ss}`;
+};
+
+/** One anchor, resolved, as markdown lines. */
+function anchorLines(a, ctx) {
+ const { report, manifest, takes, verdicts } = ctx;
+ switch (a.kind) {
+ case "whole":
+ return [ctx.kind === "article" ? "**On:** the whole article" : "**On:** the whole project"];
+ case "section": {
+ const b = blockText(report, a.section);
+ return [`**On:** section “${b?.title ?? a.section}”${b ? "" : " (section no longer exists)"}`];
+ }
+ case "text": {
+ const b = blockText(report, a.section);
+ if (!b) return [`**On:** text in section \`${a.section}\` — ORPHANED (section no longer exists)`, quoteLine(a.quote)];
+ const hit = locateQuote(b.text, a);
+ if (!hit.found) return [`**On:** text in “${b.title}” — ORPHANED (quote no longer in the section)`, quoteLine(a.quote)];
+ const exact = b.text.slice(hit.start, hit.end);
+ return [
+ `**On:** “${clip(exact, 200)}” in “${b.title}”${hit.how === "exact" ? "" : ` (found ${hit.how})`}`,
+ quoteLine(sentenceAround(b.text, hit.start, hit.end)),
+ ];
+ }
+ case "cite": {
+ const c = report?.citations?.[a.cite];
+ if (!c) return [`**On:** citation \`${a.cite}\` (no longer in the report)`];
+ const who = [c.speaker, c.date].filter(Boolean).join(", ");
+ const where =
+ c.kind === "video" || c.kind === "audio"
+ ? `${c.channel}/${c.id}@${c.start}-${c.end}`
+ : c.kind === "post"
+ ? `post ${c.channel}/${c.id}`
+ : c.kind === "page"
+ ? c.url
+ : `source ${c.source}`;
+ return [`**On:** citation \`${a.cite}\`${c.label ? ` “${c.label}”` : ""} — ${where}${who ? ` (${who})` : ""}`, quoteLine(c.quote)];
+ }
+ case "moment": {
+ const r = a.resolved ?? {};
+ const take = a.take ? ` of take \`${a.take}\`` : "";
+ const lines = [`**On:** ${hms(a.t)} in \`${a.file}\`${take}${r.approx ? " (approximate)" : ""}`];
+ const entry = a.entry ?? r.entry;
+ if (entry || r.title) lines.push(`entry \`${entry ?? "?"}\`${r.title ? ` “${clip(r.title, 160)}”` : ""}`);
+ if (r.quote) lines.push(quoteLine(r.quote));
+ if (r.channel && r.video) lines.push(`source ${r.channel}/${r.video}${r.sourceT !== undefined ? ` @ ${hms(r.sourceT)}` : ""}${r.url ? ` — ${r.url}` : ""}`);
+ return lines;
+ }
+ case "entry": {
+ const e = (manifest?.timeline ?? []).find((x) => x.id === a.entry);
+ if (!e) return [`**On:** timeline entry \`${a.entry}\` (no longer in the manifest)`];
+ const title = e.title ?? e.onscreen?.title;
+ const lines = [`**On:** timeline entry \`${a.entry}\` (${e.type ?? "entry"})${title ? ` “${clip(title, 160)}”` : ""}`];
+ if (e.quote) lines.push(quoteLine(e.quote));
+ if (e.type === "clip" && e.channel && e.video) lines.push(`source ${e.channel}/${e.video}@${e.start}-${e.end}`);
+ return lines;
+ }
+ case "take": {
+ const t = takes?.find((x) => x.id === a.take);
+ const v = verdicts?.[a.take];
+ const lines = [`**On:** take \`${a.take}\`${t ? ` “${t.label}” (${t.group})` : " (no longer in takes/)"}${v?.verdict ? ` — verdict: ${v.verdict}` : ""}`];
+ if (t?.summary) lines.push(`summary: ${clip(t.summary, 300)}`);
+ return lines;
+ }
+ case "edit":
+ return [
+ `**On:** an edit made in umtool to a GENERATED manifest — port it into the generator's inputs`,
+ `\`${a.entry ? `${a.entry}.` : ""}${a.field}\`: ${clip(JSON.stringify(a.from), 300)} → ${clip(JSON.stringify(a.to), 300)}`,
+ ];
+ default:
+ return [`**On:** ${JSON.stringify(a)}`];
+ }
+}
+
+/** "Edit X; Y is regenerated by Z" -- the line an agent most needs. */
+export function sourceLine(kind, source) {
+ if (!source) return kind === "article" ? "**Source:** unknown — find the draft before editing report.json, which a generator may overwrite." : "**Source:** the manifest.";
+ if (kind === "article") {
+ if (source.draft && source.generator) return `**Source:** edit \`${source.draft}\`; \`report.json\` is regenerated by \`${source.generator}\`.`;
+ if (source.draft) return `**Source:** edit \`${source.draft}\`.`;
+ if (source.generator) return `**Source:** \`report.json\` is written by \`${source.generator}\`; edit its inputs.`;
+ } else {
+ if (source.generator) return `**Source:** \`${source.manifest}\` is regenerated by \`${source.generator}\`; edit its inputs (BEATS, drafts), not the manifest.`;
+ if (source.manifest) return `**Source:** edit \`${source.manifest}\`.`;
+ }
+ return `**Source:** ${source.how ?? "unknown"}`;
+}
+
+const STATUS = { open: (n) => n.status === "open", resolved: (n) => n.status !== "open", all: () => true };
+
+/**
+ * The digest of one notes file.
+ *
+ * @param {{ kind: "article" | "video-project", id: string, file: string, doc: any, error?: string }} entry
+ * @param {{ status?: "open" | "resolved" | "all", sitesDir?: string, reportsRoot?: string, projectDir?: string }} [opts]
+ */
+export async function digest(entry, { status = "open", sitesDir = SITES_DIR, reportsRoot = REPORTS_ROOT, projectDir } = {}) {
+ const lines = [];
+ const doc = entry.doc;
+ const ctx = { kind: entry.kind, report: null, manifest: null, takes: null, verdicts: null };
+ let heading;
+ if (entry.kind === "article") {
+ const [site, report] = entry.id.split("/");
+ ctx.report = await readJson(path.join(/* turbopackIgnore: true */ sitesDir, site, "reports", report, "report.json"));
+ heading = `# Notes on ${entry.id}${ctx.report?.title ? ` — ${ctx.report.title}` : ""}`;
+ } else {
+ const dir = projectDir ?? path.dirname(/* turbopackIgnore: true */ entry.file);
+ ctx.manifest = await readJson(path.join(/* turbopackIgnore: true */ dir, "video.manifest.json"));
+ const t = await listTakes(dir);
+ ctx.takes = t.takes;
+ ctx.verdicts = await readVerdicts(dir);
+ heading = `# Notes on ${entry.id}${ctx.manifest?.title ? ` — ${ctx.manifest.title}` : ""}`;
+ }
+ lines.push(heading, "");
+ lines.push(`file: \`${entry.file}\``);
+ if (entry.error) {
+ lines.push("", `**notes.json does not parse:** ${entry.error}. Nothing here may write it until it is fixed by hand.`);
+ return lines.join("\n");
+ }
+ lines.push(sourceLine(entry.kind, doc?.source));
+ if (doc?.source?.how) lines.push(`(${doc.source.how})`);
+ const all = doc?.notes ?? [];
+ const shown = all.filter(STATUS[status] ?? STATUS.open);
+ const open = all.filter((n) => n.status === "open").length;
+ lines.push("", `${open} open, ${all.length - open} closed${status === "all" ? "" : `; showing ${status}`}.`);
+ for (const n of shown) {
+ lines.push("", `## ${n.id} — ${n.status}${n.status !== "open" && n.resolvedBy ? ` by ${n.resolvedBy}` : ""} (${n.author}, ${n.at.slice(0, 16).replace("T", " ")})`);
+ lines.push(...anchorLines(n.anchor, ctx));
+ lines.push("", n.text);
+ for (const r of n.replies) lines.push("", `- **${r.author}** (${r.at.slice(0, 16).replace("T", " ")}): ${r.text.replace(/\n/g, "\n ")}`);
+ }
+ if (entry.kind === "video-project" && ctx.takes?.length) {
+ lines.push("", "## Takes (takes/verdicts.json)");
+ for (const t of ctx.takes) {
+ const v = ctx.verdicts?.[t.id];
+ lines.push(`- \`${t.id}\` “${t.label}” (${t.group}, ${t.kind}): ${v?.verdict ?? "no verdict"}${v?.note ? ` — ${clip(v.note, 400)}` : ""}`);
+ }
+ }
+ if (shown.length) {
+ lines.push("", "Act on a note by editing the SOURCE above and regenerating, then:");
+ lines.push(`\`umtool notes reply <id> "what you changed" --resolve\` (or \`umtool notes reply <id> "question"\` to ask).`);
+ }
+ return lines.join("\n");
+}
diff --git a/umtool/lib/annotations/server.ts b/umtool/lib/annotations/server.ts
@@ -0,0 +1,27 @@
+import { projectRef } from "@/lib/projects";
+import { articleTarget, projectTargetFor, TargetError } from "./targets.mjs";
+
+// The app's side of lib/annotations/targets.mjs: a request's `article` or
+// `project` parameter to a target, using the app's memoised project walk.
+// Everything about WHERE a note may be written is decided in targets.mjs.
+
+export type Target = Awaited<ReturnType<typeof articleTarget>> | Awaited<ReturnType<typeof projectTargetFor>>;
+
+export async function targetFrom(params: { article?: string | null; project?: string | null }): Promise<Target> {
+ if (params.article && params.project) throw new TargetError("pass article or project, not both");
+ if (params.article) return articleTarget(params.article);
+ if (params.project) {
+ const p = await projectRef(params.project);
+ if (!p) throw new TargetError(`no project ${params.project}`, 404);
+ return projectTargetFor(p);
+ }
+ throw new TargetError("pass ?article=<site>/<report> or ?project=<id>");
+}
+
+/** A thrown error as a response: typed refusals keep their status, the rest are 500s. */
+export function errorResponse(err: unknown): Response {
+ const status = typeof (err as { status?: unknown })?.status === "number" ? (err as { status: number }).status : null;
+ const name = (err as Error)?.name;
+ const code = status ?? (name === "NoteError" ? 400 : 500);
+ return Response.json({ error: err instanceof Error ? err.message : String(err) }, { status: code });
+}
diff --git a/umtool/lib/annotations/shape.mjs b/umtool/lib/annotations/shape.mjs
@@ -0,0 +1,302 @@
+// notes.json: its constants, its validation and the edits made to it -- PURE
+// (no node imports), so the store (./store.mjs, node), the CLI and a client
+// component all hold the same rules. ./types.ts gives them types.
+//
+// { "format": "umtool-notes", "version": 1,
+// "subject": { kind: "article", site, report } | { kind: "video-project", project },
+// "source": { draft?, generator?, manifest?, how? }, which file to EDIT
+// "notes": [ { id, status, author, text, at, updatedAt, anchor, replies,
+// resolvedAt?, resolvedBy? } ] }
+//
+// docs/notes.md is the prose.
+
+export const NOTES_FORMAT = "umtool-notes";
+export const NOTES_VERSION = 1;
+export const NOTE_TEXT_LIMIT = 8000;
+export const NOTE_STATUSES = ["open", "resolved", "wontfix"];
+export const NOTE_AUTHORS = ["operator", "agent"];
+export const ANCHOR_KINDS = ["text", "cite", "section", "whole", "moment", "entry", "take", "edit"];
+export const REPORT_BLOCKS = ["title", "subtitle", "summary", "method"];
+
+const ID = /^[A-Za-z0-9][A-Za-z0-9._:-]{0,127}$/;
+const TAKE_ID = /^[a-z0-9][a-z0-9-]{0,63}$/;
+const NOTE_ID = /^n_[a-z0-9]{4,32}$/;
+const isStr = (v) => typeof v === "string";
+const short = (v, max) => isStr(v) && v.length <= max;
+
+export const isNoteId = (v) => isStr(v) && NOTE_ID.test(v);
+
+/** A fresh note id: time then randomness, base36. */
+export function newNoteId(now = Date.now()) {
+ return `n_${now.toString(36)}${Math.random().toString(36).slice(2, 6).padEnd(4, "0")}`;
+}
+
+/** A moment's rel path: relative, no `..`, no empty segment, no backslash. */
+function relFile(v) {
+ return short(v, 512) && v.length > 0 && !v.startsWith("/") && !/[\\\0]/.test(v) && !v.split("/").some((s) => s === ".." || s === "" || s === ".");
+}
+
+/** A JSON value small enough to keep: an edit's from/to. */
+function smallJson(v) {
+ if (v === undefined) return true;
+ try {
+ return JSON.stringify(v).length <= 8000;
+ } catch {
+ return false;
+ }
+}
+
+const RESOLVED_STR = ["entry", "title", "quote", "channel", "video", "url"];
+
+/**
+ * An anchor as stored, or `{ error }`. Extra keys are dropped; a text anchor's
+ * prefix/suffix default to "".
+ *
+ * @param {unknown} raw
+ * @returns {{ anchor: Record<string, unknown> } | { error: string }}
+ */
+export function validateAnchor(raw) {
+ if (!raw || typeof raw !== "object" || Array.isArray(raw)) return { error: "anchor is not an object" };
+ const a = /** @type {Record<string, unknown>} */ (raw);
+ switch (a.kind) {
+ case "whole":
+ return { anchor: { kind: "whole" } };
+ case "section":
+ if (!short(a.section, 128) || !ID.test(a.section)) return { error: "section anchor needs a section id" };
+ return { anchor: { kind: "section", section: a.section } };
+ case "text": {
+ if (!short(a.section, 128) || !ID.test(a.section)) return { error: "text anchor needs a section id" };
+ if (!short(a.quote, 4000) || !a.quote.trim()) return { error: "text anchor needs a quote" };
+ if (a.prefix !== undefined && !short(a.prefix, 256)) return { error: "prefix is not a short string" };
+ if (a.suffix !== undefined && !short(a.suffix, 256)) return { error: "suffix is not a short string" };
+ return { anchor: { kind: "text", section: a.section, quote: a.quote, prefix: a.prefix ?? "", suffix: a.suffix ?? "" } };
+ }
+ case "cite":
+ if (!short(a.cite, 128) || !ID.test(a.cite)) return { error: "cite anchor needs a citation id" };
+ return { anchor: { kind: "cite", cite: a.cite } };
+ case "entry":
+ if (!short(a.entry, 128) || !ID.test(a.entry)) return { error: "entry anchor needs an entry id" };
+ return { anchor: { kind: "entry", entry: a.entry } };
+ case "take":
+ if (!isStr(a.take) || !TAKE_ID.test(a.take)) return { error: "take anchor needs a take id" };
+ return { anchor: { kind: "take", take: a.take } };
+ case "moment": {
+ if (!relFile(a.file)) return { error: "moment anchor needs a relative file" };
+ const t = Number(a.t);
+ if (!Number.isFinite(t) || t < 0) return { error: "moment anchor needs t ≥ 0" };
+ const out = { kind: "moment", file: a.file, t: Number(t.toFixed(2)) };
+ if (a.take !== undefined) {
+ if (!isStr(a.take) || !TAKE_ID.test(a.take)) return { error: "moment take is not a take id" };
+ out.take = a.take;
+ }
+ if (a.entry !== undefined) {
+ if (!short(a.entry, 128) || !ID.test(a.entry)) return { error: "moment entry is not an entry id" };
+ out.entry = a.entry;
+ }
+ if (a.resolved !== undefined) {
+ if (!a.resolved || typeof a.resolved !== "object" || Array.isArray(a.resolved)) return { error: "resolved is not an object" };
+ const r = /** @type {Record<string, unknown>} */ (a.resolved);
+ const res = {};
+ for (const k of RESOLVED_STR) if (short(r[k], 2000)) res[k] = r[k];
+ if (typeof r.sourceT === "number" && Number.isFinite(r.sourceT)) res.sourceT = Number(r.sourceT.toFixed(2));
+ if (r.approx === true) res.approx = true;
+ out.resolved = res;
+ }
+ return { anchor: out };
+ }
+ case "edit": {
+ if (!short(a.field, 128) || !a.field) return { error: "edit anchor needs a field" };
+ if (a.entry !== undefined && (!short(a.entry, 128) || !ID.test(a.entry))) return { error: "edit entry is not an entry id" };
+ if (!smallJson(a.from) || !smallJson(a.to)) return { error: "edit from/to too large" };
+ const out = { kind: "edit", field: a.field, from: a.from ?? null, to: a.to ?? null };
+ if (a.entry !== undefined) out.entry = a.entry;
+ return { anchor: out };
+ }
+ default:
+ return { error: `anchor kind must be one of ${ANCHOR_KINDS.join(", ")}` };
+ }
+}
+
+/** @returns {{ subject: Record<string, string> } | { error: string }} */
+export function validateSubject(raw) {
+ if (!raw || typeof raw !== "object") return { error: "subject is not an object" };
+ const s = /** @type {Record<string, unknown>} */ (raw);
+ if (s.kind === "article" && isStr(s.site) && TAKE_ID.test(s.site) && isStr(s.report) && TAKE_ID.test(s.report)) {
+ return { subject: { kind: "article", site: s.site, report: s.report } };
+ }
+ if (s.kind === "video-project" && short(s.project, 512) && s.project.length > 0) {
+ return { subject: { kind: "video-project", project: s.project } };
+ }
+ return { error: "subject must be an article {site, report} or a video-project {project}" };
+}
+
+function cleanSource(raw) {
+ if (!raw || typeof raw !== "object" || Array.isArray(raw)) return undefined;
+ const out = {};
+ for (const k of ["draft", "generator", "manifest", "how"]) if (short(raw[k], 1000) && raw[k]) out[k] = raw[k];
+ return Object.keys(out).length ? out : undefined;
+}
+
+function cleanText(v) {
+ if (!isStr(v)) throw new NoteError("text must be a string");
+ const t = v.replace(/\s+$/, "");
+ if (!t.trim()) throw new NoteError("text is empty");
+ if (t.length > NOTE_TEXT_LIMIT) throw new NoteError(`text is over ${NOTE_TEXT_LIMIT} characters`);
+ return t;
+}
+
+/** A refused edit: the caller's fault, a 400. */
+export class NoteError extends Error {
+ constructor(message) {
+ super(message);
+ this.name = "NoteError";
+ }
+}
+
+/**
+ * A parsed notes.json, checked. `{ doc }` or `{ error }`. A note that is not
+ * the contract's shape is an error for the whole file -- the file is never
+ * "repaired" by dropping it, because the next write would erase it.
+ */
+export function parseNotesDoc(raw) {
+ if (!raw || typeof raw !== "object" || Array.isArray(raw)) return { error: "not an object" };
+ if (raw.format !== NOTES_FORMAT) return { error: `format is not ${NOTES_FORMAT}` };
+ if (raw.version !== NOTES_VERSION) return { error: `version ${raw.version} is not ${NOTES_VERSION}` };
+ const subject = validateSubject(raw.subject);
+ if ("error" in subject) return subject;
+ if (!Array.isArray(raw.notes)) return { error: "notes is not a list" };
+ const notes = [];
+ for (const [i, n] of raw.notes.entries()) {
+ const where = `notes[${i}]`;
+ if (!n || typeof n !== "object") return { error: `${where} is not an object` };
+ if (!isNoteId(n.id)) return { error: `${where}.id is not a note id` };
+ if (!NOTE_STATUSES.includes(n.status)) return { error: `${where}.status` };
+ if (!NOTE_AUTHORS.includes(n.author)) return { error: `${where}.author` };
+ if (!isStr(n.text) || !isStr(n.at) || !isStr(n.updatedAt)) return { error: `${where} text/at/updatedAt` };
+ const a = validateAnchor(n.anchor);
+ if ("error" in a) return { error: `${where}.anchor: ${a.error}` };
+ if (!Array.isArray(n.replies)) return { error: `${where}.replies is not a list` };
+ const replies = [];
+ for (const r of n.replies) {
+ if (!r || !NOTE_AUTHORS.includes(r.author) || !isStr(r.text) || !isStr(r.at)) return { error: `${where}.replies` };
+ replies.push({ author: r.author, text: r.text, at: r.at });
+ }
+ const note = { id: n.id, status: n.status, author: n.author, text: n.text, at: n.at, updatedAt: n.updatedAt, anchor: a.anchor, replies };
+ if (isStr(n.resolvedAt)) note.resolvedAt = n.resolvedAt;
+ if (NOTE_AUTHORS.includes(n.resolvedBy)) note.resolvedBy = n.resolvedBy;
+ notes.push(note);
+ }
+ const doc = { format: NOTES_FORMAT, version: NOTES_VERSION, subject: subject.subject };
+ const source = cleanSource(raw.source);
+ if (source) doc.source = source;
+ doc.notes = notes;
+ return { doc };
+}
+
+export function emptyDoc(subject, source) {
+ const doc = { format: NOTES_FORMAT, version: NOTES_VERSION, subject };
+ const s = cleanSource(source);
+ if (s) doc.source = s;
+ doc.notes = [];
+ return doc;
+}
+
+function find(doc, id) {
+ const note = doc.notes.find((n) => n.id === id);
+ if (!note) throw new NoteError(`no note ${id}`);
+ return note;
+}
+
+function setStatus(note, status, by, now) {
+ if (!NOTE_STATUSES.includes(status)) throw new NoteError(`status must be one of ${NOTE_STATUSES.join(", ")}`);
+ note.status = status;
+ note.updatedAt = now;
+ if (status === "open") {
+ delete note.resolvedAt;
+ delete note.resolvedBy;
+ } else {
+ note.resolvedAt = now;
+ note.resolvedBy = by;
+ }
+}
+
+/**
+ * Apply one op to a doc, in place. Returns the note touched (null on delete).
+ * `by` is who is writing: the app stamps "operator", the CLI "agent".
+ *
+ * { op: "add", text, anchor } a new open note
+ * { op: "edit", id, text?, anchor? } rewrite it (its author only)
+ * { op: "status", id, status } open | resolved | wontfix
+ * { op: "reply", id, text, resolve? } a threaded reply, optionally resolving
+ * { op: "delete", id } remove the note
+ * { op: "delete-reply", id, index } remove one reply
+ * { op: "source", source } correct which file to edit
+ *
+ * @param {Record<string, any>} doc
+ * @param {Record<string, any>} op
+ * @param {"operator" | "agent"} by
+ * @param {string} [now]
+ */
+export function applyOp(doc, op, by, now = new Date().toISOString()) {
+ if (!NOTE_AUTHORS.includes(by)) throw new NoteError("author must be operator or agent");
+ switch (op?.op) {
+ case "add": {
+ const a = validateAnchor(op.anchor);
+ if ("error" in a) throw new NoteError(a.error);
+ let id = newNoteId();
+ while (doc.notes.some((n) => n.id === id)) id = newNoteId();
+ const note = { id, status: "open", author: by, text: cleanText(op.text), at: now, updatedAt: now, anchor: a.anchor, replies: [] };
+ doc.notes.push(note);
+ return note;
+ }
+ case "edit": {
+ const note = find(doc, op.id);
+ if (note.author !== by) throw new NoteError(`only the ${note.author} edits this note; reply instead`);
+ if (op.text !== undefined) note.text = cleanText(op.text);
+ if (op.anchor !== undefined) {
+ const a = validateAnchor(op.anchor);
+ if ("error" in a) throw new NoteError(a.error);
+ note.anchor = a.anchor;
+ }
+ note.updatedAt = now;
+ return note;
+ }
+ case "status": {
+ const note = find(doc, op.id);
+ setStatus(note, op.status, by, now);
+ return note;
+ }
+ case "reply": {
+ const note = find(doc, op.id);
+ note.replies.push({ author: by, text: cleanText(op.text), at: now });
+ note.updatedAt = now;
+ if (op.resolve) setStatus(note, "resolved", by, now);
+ return note;
+ }
+ case "delete": {
+ const i = doc.notes.findIndex((n) => n.id === op.id);
+ if (i === -1) throw new NoteError(`no note ${op.id}`);
+ doc.notes.splice(i, 1);
+ return null;
+ }
+ case "delete-reply": {
+ const note = find(doc, op.id);
+ const i = Number(op.index);
+ if (!Number.isInteger(i) || i < 0 || i >= note.replies.length) throw new NoteError("no such reply");
+ if (note.replies[i].author !== by) throw new NoteError(`only the ${note.replies[i].author} deletes that reply`);
+ note.replies.splice(i, 1);
+ note.updatedAt = now;
+ return note;
+ }
+ case "source": {
+ const s = cleanSource(op.source);
+ if (s) doc.source = s;
+ else delete doc.source;
+ return null;
+ }
+ default:
+ throw new NoteError("op must be add, edit, status, reply, delete, delete-reply or source");
+ }
+}
+
+export const openNotes = (doc) => (doc ? doc.notes.filter((n) => n.status === "open") : []);
diff --git a/umtool/lib/annotations/store.mjs b/umtool/lib/annotations/store.mjs
@@ -0,0 +1,157 @@
+// notes.json on disk: read, lock, apply one op, write. Shared by the app's
+// /api/notes and `umtool notes`, so the operator's page and an agent's CLI go
+// through ONE writer with one set of rules (./shape.mjs).
+//
+// Three writers can race on one file -- the page, an agent's `umtool notes
+// reply`, a second agent -- and they are different PROCESSES, so the in-process
+// queues lib/state.ts and lib/report/manifest.mjs use are not enough here:
+//
+// * a LOCKFILE (`notes.json.lock`, created O_EXCL) serialises
+// read-modify-write across processes; one left by a dead writer is stale
+// after 30 s and is taken over;
+// * the write is tmp + rename, so a reader never sees half a file;
+// * every write from the page carries the TOKEN it read (the file's mtime in
+// ns, or "absent"); a stale one is a 409, never a silent overwrite of a
+// reply an agent wrote in between;
+// * a notes.json that exists and does not parse is NEVER overwritten (the
+// takes.mjs rule) -- writing over it would erase every note in it;
+// * the last note deleted deletes the file: an empty notes.json says nothing.
+//
+// WHERE a write may land is the caller's to decide BEFORE it gets here
+// (lib/annotations/targets.mjs: isCorpusNotesFile for an article, a project
+// directory for a video) -- `writeOp` takes the file it is given.
+import { open, readFile, rename, stat, unlink, writeFile } from "node:fs/promises";
+import path from "node:path";
+import { NoteError, applyOp, emptyDoc, parseNotesDoc } from "./shape.mjs";
+
+export { NoteError };
+export const LOCK_STALE_MS = 30_000;
+const LOCK_WAIT_MS = 10_000;
+
+/** A write whose token no longer matches the file: someone wrote in between. */
+export class StaleNotes extends Error {
+ constructor(expected, got) {
+ super(`the notes changed since you read them (${got} vs ${expected})`);
+ this.name = "StaleNotes";
+ this.status = 409;
+ }
+}
+
+/** A notes.json that exists and is not one: refused, never overwritten. */
+export class NotesUnreadable extends Error {
+ constructor(file, why) {
+ super(`${path.basename(file)} does not parse (${why}); not overwriting it`);
+ this.name = "NotesUnreadable";
+ this.status = 409;
+ }
+}
+
+/** The file's identity for a write guard: mtime in ns plus size, or "absent". */
+export async function notesToken(file) {
+ const st = await stat(/* turbopackIgnore: true */ file, { bigint: true }).catch(() => null);
+ return st ? `${st.mtimeNs}-${st.size}` : "absent";
+}
+
+/**
+ * `{ doc, token }` -- doc null when there is no file. `{ error }` too when the
+ * file is there and is not a notes doc (the page shows it; writes refuse).
+ *
+ * @param {string} file
+ * @returns {Promise<{ doc: import("./types").NotesDoc | null, token: string, error?: string }>}
+ */
+export async function readNotes(file) {
+ const token = await notesToken(file);
+ let text;
+ try {
+ text = await readFile(/* turbopackIgnore: true */ file, "utf8");
+ } catch (err) {
+ if (/** @type {NodeJS.ErrnoException} */ (err).code === "ENOENT") return { doc: null, token: "absent" };
+ return { doc: null, token, error: String(/** @type {Error} */ (err).message ?? err) };
+ }
+ let raw;
+ try {
+ raw = JSON.parse(text);
+ } catch (err) {
+ return { doc: null, token, error: `not JSON: ${/** @type {Error} */ (err).message}` };
+ }
+ const r = parseNotesDoc(raw);
+ if ("error" in r) return { doc: null, token, error: r.error };
+ return { doc: r.doc, token };
+}
+
+const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
+
+/**
+ * Run `fn` holding `<file>.lock`. A lock older than LOCK_STALE_MS is a dead
+ * writer's and is removed; otherwise wait (up to 10 s) and retry.
+ *
+ * @template T
+ * @param {string} file
+ * @param {() => Promise<T>} fn
+ * @param {{ waitMs?: number, staleMs?: number }} [opts]
+ * @returns {Promise<T>}
+ */
+export async function withNotesLock(file, fn, { waitMs = LOCK_WAIT_MS, staleMs = LOCK_STALE_MS } = {}) {
+ const lock = `${file}.lock`;
+ const deadline = Date.now() + waitMs;
+ let delay = 15;
+ for (;;) {
+ try {
+ const h = await open(/* turbopackIgnore: true */ lock, "wx");
+ await h.writeFile(`${process.pid} ${new Date().toISOString()}\n`).catch(() => {});
+ await h.close();
+ break;
+ } catch (err) {
+ if (/** @type {NodeJS.ErrnoException} */ (err).code !== "EEXIST") throw err;
+ const st = await stat(/* turbopackIgnore: true */ lock).catch(() => null);
+ if (st && Date.now() - st.mtimeMs > staleMs) {
+ await unlink(/* turbopackIgnore: true */ lock).catch(() => {});
+ continue;
+ }
+ if (Date.now() > deadline) throw new Error(`${path.basename(lock)} is held; try again`);
+ await sleep(delay);
+ delay = Math.min(delay * 2, 250);
+ }
+ }
+ try {
+ return await fn();
+ } finally {
+ await unlink(/* turbopackIgnore: true */ lock).catch(() => {});
+ }
+}
+
+/**
+ * Apply one op to the notes at `file` and write it back. `init` is the
+ * subject (and source) a NEW file starts with; an existing file keeps its own
+ * subject, and its `source` unless the op is `source`. `token` (when given)
+ * must match the file as it is now, or StaleNotes. `by` is "operator" (the
+ * app) or "agent" (the CLI).
+ *
+ * Returns the doc as written (null when the file was deleted), the new token,
+ * and the note the op touched.
+ *
+ * @param {string} file
+ * @param {{ subject: Record<string, string>, source?: Record<string, string> | null }} init
+ * @param {Record<string, any>} op
+ * @param {{ by: "operator" | "agent", token?: string | null }} opts
+ * @returns {Promise<{ doc: import("./types").NotesDoc | null, token: string, note: import("./types").Note | null }>}
+ */
+export async function writeOp(file, init, op, { by, token = null }) {
+ return withNotesLock(file, async () => {
+ const current = await readNotes(file);
+ if (current.error) throw new NotesUnreadable(file, current.error);
+ if (token !== null && token !== undefined && token !== current.token) throw new StaleNotes(token, current.token);
+ const doc = /** @type {any} */ (current.doc ?? emptyDoc(init.subject, init.source ?? undefined));
+ const note = applyOp(doc, op, by);
+ if (doc.notes.length === 0) {
+ await unlink(/* turbopackIgnore: true */ file).catch((err) => {
+ if (err.code !== "ENOENT") throw err;
+ });
+ return { doc: null, token: "absent", note };
+ }
+ const tmp = `${file}.tmp-${process.pid}-${Math.random().toString(36).slice(2, 8)}`;
+ await writeFile(/* turbopackIgnore: true */ tmp, JSON.stringify(doc, null, 2) + "\n", "utf8");
+ await rename(/* turbopackIgnore: true */ tmp, file);
+ return { doc, token: await notesToken(file), note };
+ });
+}
diff --git a/umtool/lib/annotations/store.test.mjs b/umtool/lib/annotations/store.test.mjs
@@ -0,0 +1,206 @@
+// notes.json: the store, the corpus write predicate and the targets.
+//
+// Run with: pnpm test:scripts
+import assert from "node:assert/strict";
+import { mkdir, mkdtemp, readFile, rm, stat, symlink, utimes, writeFile } from "node:fs/promises";
+import { tmpdir } from "node:os";
+import path from "node:path";
+import test from "node:test";
+import { corpusNotesFile, isCorpusNotesFile } from "../paths.mjs";
+import { NoteError, applyOp, emptyDoc, parseNotesDoc, validateAnchor } from "./shape.mjs";
+import { NotesUnreadable, StaleNotes, readNotes, withNotesLock, writeOp } from "./store.mjs";
+import { articleTarget, listNotesFiles, writeNote } from "./targets.mjs";
+
+const SUBJECT = { kind: "article", site: "s1", report: "r1" };
+
+async function tmp() {
+ return mkdtemp(path.join(tmpdir(), "umtool-notes-"));
+}
+
+test("add, reply, resolve, reopen, delete: round-trip, and the last delete removes the file", async () => {
+ const dir = await tmp();
+ const file = path.join(dir, "notes.json");
+ const a = await writeOp(file, { subject: SUBJECT, source: { draft: "~/d.json" } }, { op: "add", text: "fix this", anchor: { kind: "whole" } }, { by: "operator" });
+ assert.equal(a.doc.notes.length, 1);
+ assert.equal(a.doc.source.draft, "~/d.json");
+ const id = a.note.id;
+ assert.match(id, /^n_[a-z0-9]+$/);
+
+ const back = await readNotes(file);
+ assert.deepEqual(back.doc, a.doc);
+ assert.equal(back.token, a.token);
+
+ const r = await writeOp(file, { subject: SUBJECT }, { op: "reply", id, text: "done in drafts/x.json", resolve: true }, { by: "agent", token: a.token });
+ assert.equal(r.note.status, "resolved");
+ assert.equal(r.note.resolvedBy, "agent");
+ assert.equal(r.note.replies[0].author, "agent");
+
+ const o = await writeOp(file, { subject: SUBJECT }, { op: "status", id, status: "open" }, { by: "operator" });
+ assert.equal(o.note.status, "open");
+ assert.equal(o.note.resolvedAt, undefined);
+
+ const d = await writeOp(file, { subject: SUBJECT }, { op: "delete", id }, { by: "operator" });
+ assert.equal(d.doc, null);
+ await assert.rejects(stat(file), /ENOENT/);
+ await rm(dir, { recursive: true });
+});
+
+test("a stale token is a 409 and the file is untouched", async () => {
+ const dir = await tmp();
+ const file = path.join(dir, "notes.json");
+ const a = await writeOp(file, { subject: SUBJECT }, { op: "add", text: "one", anchor: { kind: "whole" } }, { by: "operator", token: "absent" });
+ await writeOp(file, { subject: SUBJECT }, { op: "add", text: "two (agent)", anchor: { kind: "whole" } }, { by: "agent" });
+ const before = await readFile(file, "utf8");
+ await assert.rejects(
+ writeOp(file, { subject: SUBJECT }, { op: "add", text: "three", anchor: { kind: "whole" } }, { by: "operator", token: a.token }),
+ (err) => err instanceof StaleNotes && err.status === 409,
+ );
+ assert.equal(await readFile(file, "utf8"), before);
+ // "absent" against a file that exists is stale too
+ await assert.rejects(
+ writeOp(file, { subject: SUBJECT }, { op: "add", text: "x", anchor: { kind: "whole" } }, { by: "operator", token: "absent" }),
+ StaleNotes,
+ );
+ await rm(dir, { recursive: true });
+});
+
+test("an unparseable notes.json is never overwritten", async () => {
+ const dir = await tmp();
+ const file = path.join(dir, "notes.json");
+ await writeFile(file, "{ not json");
+ assert.match((await readNotes(file)).error, /not JSON/);
+ await assert.rejects(writeOp(file, { subject: SUBJECT }, { op: "add", text: "x", anchor: { kind: "whole" } }, { by: "operator" }), NotesUnreadable);
+ assert.equal(await readFile(file, "utf8"), "{ not json");
+ // a parseable file with a bad note is refused the same way
+ await writeFile(file, JSON.stringify({ format: "umtool-notes", version: 1, subject: SUBJECT, notes: [{ id: "bad" }] }));
+ await assert.rejects(writeOp(file, { subject: SUBJECT }, { op: "add", text: "x", anchor: { kind: "whole" } }, { by: "operator" }), NotesUnreadable);
+ await rm(dir, { recursive: true });
+});
+
+test("lock contention: parallel writers all land; a held lock waits; a stale lock is taken over", async () => {
+ const dir = await tmp();
+ const file = path.join(dir, "notes.json");
+ await Promise.all(
+ Array.from({ length: 12 }, (_, i) =>
+ writeOp(file, { subject: SUBJECT }, { op: "add", text: `n${i}`, anchor: { kind: "whole" } }, { by: i % 2 ? "agent" : "operator" }),
+ ),
+ );
+ assert.equal((await readNotes(file)).doc.notes.length, 12);
+
+ // Another process holds it (a fresh lock file): we wait, then give up.
+ await writeFile(`${file}.lock`, "999999 now\n");
+ await assert.rejects(withNotesLock(file, async () => 1, { waitMs: 120 }), /held/);
+ // The same lock, 31 s old: a dead writer's, taken over.
+ const old = new Date(Date.now() - 31_000);
+ await utimes(`${file}.lock`, old, old);
+ assert.equal(await withNotesLock(file, async () => 2, { waitMs: 120 }), 2);
+ await assert.rejects(stat(`${file}.lock`), /ENOENT/);
+ await rm(dir, { recursive: true });
+});
+
+test("ops refuse what the contract does not allow", () => {
+ const doc = emptyDoc(SUBJECT);
+ assert.throws(() => applyOp(doc, { op: "add", text: " ", anchor: { kind: "whole" } }, "operator"), NoteError);
+ assert.throws(() => applyOp(doc, { op: "add", text: "x", anchor: { kind: "nope" } }, "operator"), NoteError);
+ assert.throws(() => applyOp(doc, { op: "add", text: "x".repeat(8001), anchor: { kind: "whole" } }, "operator"), NoteError);
+ const n = applyOp(doc, { op: "add", text: "x", anchor: { kind: "whole" } }, "operator");
+ assert.throws(() => applyOp(doc, { op: "edit", id: n.id, text: "agent rewrites it" }, "agent"), /reply instead/);
+ assert.throws(() => applyOp(doc, { op: "status", id: n.id, status: "done" }, "agent"), NoteError);
+ assert.throws(() => applyOp(doc, { op: "reply", id: "n_missing0", text: "x" }, "agent"), /no note/);
+ assert.throws(() => applyOp(doc, { op: "frobnicate" }, "agent"), NoteError);
+ assert.throws(() => applyOp(doc, { op: "add", text: "x", anchor: { kind: "whole" } }, "someone"), NoteError);
+ // source: set, then cleared
+ applyOp(doc, { op: "source", source: { draft: "~/x.json", junk: 1 } }, "agent");
+ assert.deepEqual(doc.source, { draft: "~/x.json" });
+ assert.ok(parseNotesDoc(JSON.parse(JSON.stringify(doc))).doc);
+});
+
+test("anchors: every kind validates; bad ones refuse", () => {
+ const ok = [
+ { kind: "whole" },
+ { kind: "section", section: "s-2" },
+ { kind: "text", section: "summary", quote: "the claim", prefix: "before ", suffix: " after" },
+ { kind: "cite", cite: "c12" },
+ { kind: "moment", file: "takes/deck/preview.mp4", t: 12.345, take: "deck", entry: "e3", resolved: { title: "T", sourceT: 81.234, approx: true, bogus: 1 } },
+ { kind: "entry", entry: "clip-4" },
+ { kind: "take", take: "cold-open" },
+ { kind: "edit", entry: "e3", field: "quote", from: "a", to: "b" },
+ ];
+ for (const a of ok) assert.ok("anchor" in validateAnchor(a), JSON.stringify(a));
+ assert.equal(validateAnchor(ok[4]).anchor.t, 12.35);
+ assert.deepEqual(validateAnchor(ok[4]).anchor.resolved, { title: "T", sourceT: 81.23, approx: true });
+ assert.equal(validateAnchor({ kind: "text", section: "s", quote: "q" }).anchor.prefix, "");
+ const bad = [
+ null,
+ { kind: "text", section: "s" },
+ { kind: "moment", file: "../x.mp4", t: 1 },
+ { kind: "moment", file: "/abs.mp4", t: 1 },
+ { kind: "moment", file: "a.mp4", t: -1 },
+ { kind: "take", take: "Bad Id" },
+ { kind: "section", section: "has space" },
+ { kind: "edit", field: "" },
+ ];
+ for (const a of bad) assert.ok("error" in validateAnchor(a), JSON.stringify(a));
+});
+
+async function sitesFixture() {
+ const root = await tmp();
+ const sites = path.join(root, "sites");
+ await mkdir(path.join(sites, "s1", "reports", "r1"), { recursive: true });
+ await writeFile(path.join(sites, "s1", "site.json"), "{}");
+ await writeFile(path.join(sites, "s1", "reports", "r1", "report.json"), "{}");
+ const outside = path.join(root, "outside", "r2");
+ await mkdir(outside, { recursive: true });
+ await symlink(outside, path.join(sites, "s1", "reports", "r2"));
+ return { root, sites };
+}
+
+test("isCorpusNotesFile: exactly sites/<site>/reports/<id>/notes.json, and nothing else", async () => {
+ const { root, sites } = await sitesFixture();
+ const opt = { sitesDir: sites };
+ const good = path.join(sites, "s1", "reports", "r1", "notes.json");
+ assert.equal(corpusNotesFile("s1", "r1", opt), good);
+ assert.equal(await isCorpusNotesFile(good, opt), true);
+ // traversal, spelled several ways
+ assert.equal(await isCorpusNotesFile(path.join(sites, "s1", "reports", "r1", "..", "r1", "notes.json").replace(/\/r1\/notes/, "/../r1/r1/notes"), opt), false);
+ assert.equal(await isCorpusNotesFile(`${sites}/s1/reports/../reports/r1/notes.json`, opt), false);
+ assert.equal(await isCorpusNotesFile(`${sites}/s1/reports//r1/notes.json`, opt), false);
+ assert.equal(await isCorpusNotesFile("s1/reports/r1/notes.json", opt), false);
+ // wrong name, wrong depth, wrong middle segment, bad ids
+ assert.equal(await isCorpusNotesFile(path.join(sites, "s1", "reports", "r1", "report.json"), opt), false);
+ assert.equal(await isCorpusNotesFile(path.join(sites, "s1", "reports", "r1", "x", "notes.json"), opt), false);
+ assert.equal(await isCorpusNotesFile(path.join(sites, "s1", "stills", "r1", "notes.json"), opt), false);
+ assert.equal(await isCorpusNotesFile(path.join(sites, "S1", "reports", "r1", "notes.json"), opt), false);
+ assert.equal(corpusNotesFile("s1", "../r1", opt), null);
+ // a report dir that is a symlink out of the site
+ assert.equal(await isCorpusNotesFile(path.join(sites, "s1", "reports", "r2", "notes.json"), opt), false);
+ // a report that does not exist: a note never creates its directory
+ assert.equal(await isCorpusNotesFile(path.join(sites, "s1", "reports", "r9", "notes.json"), opt), false);
+ // a notes.json that is itself a symlink
+ await writeFile(path.join(root, "elsewhere.json"), "{}");
+ await symlink(path.join(root, "elsewhere.json"), good);
+ assert.equal(await isCorpusNotesFile(good, opt), false);
+ await rm(root, { recursive: true });
+});
+
+test("articleTarget: writes land beside report.json with the discovered source; refusals are typed", async () => {
+ const { root, sites } = await sitesFixture();
+ const reports = path.join(root, "reports");
+ await mkdir(path.join(reports, "ws", "polemics", "drafts"), { recursive: true });
+ await writeFile(path.join(reports, "ws", "polemics", "drafts", "r1.json"), JSON.stringify({ id: "r1" }));
+ await writeFile(path.join(reports, "ws", "polemics", "make-site.py"), 'OUT = "sites/s1/reports"\nfor d in drafts: pass\n');
+ const opt = { sitesDir: sites, reportsRoot: reports };
+
+ const t = await articleTarget("s1/r1", opt);
+ const w = await writeNote(t, { op: "add", text: "x", anchor: { kind: "whole" } }, { by: "operator" });
+ assert.equal(w.doc.source.draft.endsWith(path.join("ws", "polemics", "drafts", "r1.json")), true);
+ assert.equal(w.doc.source.generator.endsWith("make-site.py"), true);
+ const list = await listNotesFiles(opt);
+ assert.deepEqual(list.map((l) => [l.kind, l.id, l.doc.notes.length]), [["article", "s1/r1", 1]]);
+
+ await assert.rejects(articleTarget("s1/r9", opt), (e) => e.status === 404);
+ await assert.rejects(articleTarget("s1/r2", opt), (e) => e.status === 403);
+ await assert.rejects(articleTarget("s1/../r1", opt), (e) => e.status === 400);
+ await assert.rejects(articleTarget("s1", opt), (e) => e.status === 400);
+ await rm(root, { recursive: true });
+});
diff --git a/umtool/lib/annotations/targets.mjs b/umtool/lib/annotations/targets.mjs
@@ -0,0 +1,185 @@
+// What a note is ON, and therefore which notes.json it lives in -- decided
+// here, once, for the app's /api/notes and for `umtool notes`.
+//
+// article `SITES_DIR/<site>/reports/<report>/notes.json`, beside
+// report.json. The generators that write a report dir
+// overwrite only report.json, video.mp4 and poster.jpg, so it
+// survives a regenerate; the compose stage and the report
+// history never read it (common/publish tests hold that), so
+// it is never published. The ONE corpus file umtool writes,
+// and only through isCorpusNotesFile.
+// video-project `<project>/notes.json`, beside video.manifest.json. A
+// report-video project under REPORTS_ROOT, found by the same
+// walk `umtool ls` uses.
+import { readdir, readFile, realpath, stat } from "node:fs/promises";
+import path from "node:path";
+import { NOTES_FILENAME, REPORTS_ROOT, SEGMENT_RE, SITES_DIR, corpusNotesFile, inside, isCorpusNotesFile } from "../paths.mjs";
+import { projectRefs, resolveProject } from "../projects/core.mjs";
+import { kindTakesNotes } from "../projects/kinds.mjs";
+import { sourceFor, tildify } from "../articles/sources.mjs";
+import { readNotes, writeOp } from "./store.mjs";
+
+const exists = (p) => stat(/* turbopackIgnore: true */ p).then(() => true, () => false);
+
+/** A refused target: the caller's fault (400/404). */
+export class TargetError extends Error {
+ constructor(message, status = 400) {
+ super(message);
+ this.name = "TargetError";
+ this.status = status;
+ }
+}
+
+/**
+ * An article target, checked: the site and report ids, the report directory
+ * on disk, and the notes path through isCorpusNotesFile.
+ *
+ * @param {string} spec `<site>/<report>`
+ * @param {{ sitesDir?: string, reportsRoot?: string }} [opts]
+ */
+export async function articleTarget(spec, { sitesDir = SITES_DIR, reportsRoot = REPORTS_ROOT } = {}) {
+ const [site, report, ...rest] = String(spec ?? "").split("/");
+ if (rest.length || !SEGMENT_RE.test(site ?? "") || !SEGMENT_RE.test(report ?? "")) {
+ throw new TargetError(`not an article: ${JSON.stringify(spec)} (want <site>/<report>)`);
+ }
+ const file = corpusNotesFile(site, report, { sitesDir });
+ if (!file || !(await exists(path.dirname(/* turbopackIgnore: true */ file)))) throw new TargetError(`no report ${site}/${report}`, 404);
+ if (!(await isCorpusNotesFile(file, { sitesDir }))) throw new TargetError(`refusing to write ${file}`, 403);
+ return {
+ kind: "article",
+ id: `${site}/${report}`,
+ file,
+ subject: { kind: "article", site, report },
+ // Lazy: the scan reads every workspace's generators, and a read of an
+ // existing file never needs it.
+ source: async () => sourceFor(site, report, { reportsRoot }),
+ };
+}
+
+/**
+ * A video-project target: a report-video project under REPORTS_ROOT, by id,
+ * unique name or directory.
+ *
+ * @param {string} spec
+ * @param {{ reportsRoot?: string }} [opts]
+ */
+export async function projectTarget(spec, { reportsRoot = REPORTS_ROOT } = {}) {
+ const r = await resolveProject(String(spec ?? ""), reportsRoot);
+ if (r.ambiguous) throw new TargetError(`${spec} names ${r.ambiguous.length} projects: ${r.ambiguous.map((p) => p.id).join(", ")}`);
+ if (!r.project) throw new TargetError(`no project ${spec}`, 404);
+ return projectTargetFor(r.project, { reportsRoot });
+}
+
+/**
+ * The target for a project ref already in hand (the app's memoised walk).
+ *
+ * @param {{ id: string, dir: string, kind: string }} p
+ * @param {{ reportsRoot?: string }} [opts]
+ */
+export async function projectTargetFor(p, { reportsRoot = REPORTS_ROOT } = {}) {
+ if (!kindTakesNotes(p.kind)) throw new TargetError(`${p.id} is a ${p.kind} project; notes are for report videos`);
+ const [realRoot, realDir] = await Promise.all([
+ realpath(/* turbopackIgnore: true */ reportsRoot).catch(() => null),
+ realpath(/* turbopackIgnore: true */ p.dir).catch(() => null),
+ ]);
+ if (!realRoot || !realDir || !inside(realRoot, realDir) || realRoot === realDir) {
+ throw new TargetError(`${p.id} is not under the reports root`, 403);
+ }
+ const file = path.join(/* turbopackIgnore: true */ p.dir, NOTES_FILENAME);
+ return {
+ kind: "video-project",
+ id: p.id,
+ dir: p.dir,
+ file,
+ subject: { kind: "video-project", project: p.id },
+ source: async () => projectSource(p.dir, reportsRoot),
+ };
+}
+
+/**
+ * A report-video project's source: its manifest, and -- when the manifest is
+ * generated -- the generator, resolved against the project's workspace (the
+ * first directory under REPORTS_ROOT) when that file exists.
+ */
+export async function projectSource(dir, reportsRoot = REPORTS_ROOT) {
+ const manifest = path.join(/* turbopackIgnore: true */ dir, "video.manifest.json");
+ const out = { manifest: tildify(manifest) };
+ let generatedBy = null;
+ try {
+ const m = JSON.parse(await readFile(/* turbopackIgnore: true */ manifest, "utf8"));
+ if (typeof m.generatedBy === "string" && m.generatedBy.trim()) generatedBy = m.generatedBy.trim();
+ } catch {
+ // no manifest, or not JSON: the manifest path is still the place to look
+ }
+ if (!generatedBy) {
+ out.how = "hand-edited manifest";
+ return out;
+ }
+ const rel = path.relative(/* turbopackIgnore: true */ reportsRoot, dir);
+ const ws = rel && !rel.startsWith("..") ? path.join(/* turbopackIgnore: true */ reportsRoot, rel.split(path.sep)[0]) : null;
+ const candidates = [ws && path.join(/* turbopackIgnore: true */ ws, generatedBy), path.join(/* turbopackIgnore: true */ dir, generatedBy)].filter(Boolean);
+ let gen = null;
+ for (const c of candidates) {
+ if (await exists(c)) {
+ gen = c;
+ break;
+ }
+ }
+ out.generator = gen ? tildify(gen) : generatedBy;
+ out.how = `manifest is generated by ${generatedBy}; edit its inputs, then regenerate`;
+ return out;
+}
+
+/**
+ * Resolve what the CLI was handed: `<site>/<report>` when that report exists,
+ * else a project.
+ */
+export async function resolveTarget(spec, opts = {}) {
+ const parts = String(spec ?? "").split("/");
+ if (parts.length === 2 && parts.every((s) => SEGMENT_RE.test(s))) {
+ const dir = path.join(/* turbopackIgnore: true */ opts.sitesDir ?? SITES_DIR, parts[0], "reports", parts[1]);
+ if (await exists(dir)) return articleTarget(spec, opts);
+ }
+ return projectTarget(spec, opts);
+}
+
+/**
+ * Every notes.json there is: articles under SITES_DIR, projects under
+ * REPORTS_ROOT. `{ target, file, doc, error? }` each; sorted by id.
+ *
+ * @param {{ sitesDir?: string, reportsRoot?: string }} [opts]
+ */
+export async function listNotesFiles({ sitesDir = SITES_DIR, reportsRoot = REPORTS_ROOT } = {}) {
+ const out = [];
+ for (const site of await readdir(/* turbopackIgnore: true */ sitesDir).catch(() => [])) {
+ if (!SEGMENT_RE.test(site)) continue;
+ for (const report of await readdir(/* turbopackIgnore: true */ path.join(/* turbopackIgnore: true */ sitesDir, site, "reports")).catch(() => [])) {
+ if (!SEGMENT_RE.test(report)) continue;
+ const file = path.join(/* turbopackIgnore: true */ sitesDir, site, "reports", report, NOTES_FILENAME);
+ if (!(await exists(file))) continue;
+ out.push({ kind: "article", id: `${site}/${report}`, file, ...(await readNotes(file)) });
+ }
+ }
+ for (const p of await projectRefs(reportsRoot)) {
+ if (!kindTakesNotes(p.kind)) continue;
+ const file = path.join(/* turbopackIgnore: true */ p.dir, NOTES_FILENAME);
+ if (!(await exists(file))) continue;
+ out.push({ kind: "video-project", id: p.id, projectKind: p.kind, file, ...(await readNotes(file)) });
+ }
+ return out.sort((a, b) => a.kind.localeCompare(b.kind) || a.id.localeCompare(b.id));
+}
+
+/**
+ * One write to a target's notes. A file that does not exist yet starts with
+ * the target's subject and its discovered source (which the agent may correct
+ * later with a `source` op); an existing file keeps both.
+ *
+ * @param {{ file: string, subject: Record<string, string>, source: () => Promise<Record<string, string> | null> }} target
+ * @param {Record<string, any>} op
+ * @param {{ by: "operator" | "agent", token?: string | null }} opts
+ */
+export async function writeNote(target, op, opts) {
+ const now = await readNotes(target.file);
+ const source = now.doc || now.error ? undefined : ((await target.source()) ?? undefined);
+ return writeOp(target.file, { subject: target.subject, source }, op, opts);
+}
diff --git a/umtool/lib/annotations/types.ts b/umtool/lib/annotations/types.ts
@@ -0,0 +1,98 @@
+// The notes shapes, with NO server imports (the lib/note-types.ts rule): a
+// client component may take a value from here without dragging node:fs into
+// the browser bundle. The store that reads and writes them is ./store.mjs,
+// shared by the app and `umtool notes`; docs/notes.md is the prose.
+
+// The values are ./shape.mjs's (pure, shared with the store and the CLI);
+// this file adds the types.
+export {
+ ANCHOR_KINDS,
+ NOTE_AUTHORS,
+ NOTE_STATUSES,
+ NOTE_TEXT_LIMIT,
+ NOTES_FORMAT,
+ NOTES_VERSION,
+ REPORT_BLOCKS,
+} from "./shape.mjs";
+export { CONTEXT as ANCHOR_CONTEXT } from "./anchor.mjs";
+
+export type NoteStatus = "open" | "resolved" | "wontfix";
+export type NoteAuthor = "operator" | "agent";
+
+/** What a moment resolved to when it was written, for the agent reading it back. */
+export type MomentResolved = {
+ entry?: string;
+ title?: string;
+ quote?: string;
+ channel?: string;
+ video?: string;
+ sourceT?: number;
+ url?: string;
+ /** The schedule did not match this file exactly (a preview, not out/). */
+ approx?: boolean;
+};
+
+export type Anchor =
+ | { kind: "text"; section: string; quote: string; prefix: string; suffix: string }
+ | { kind: "cite"; cite: string }
+ | { kind: "section"; section: string }
+ | { kind: "whole" }
+ | { kind: "moment"; file: string; t: number; take?: string; entry?: string; resolved?: MomentResolved }
+ | { kind: "entry"; entry: string }
+ | { kind: "take"; take: string }
+ | { kind: "edit"; entry?: string; field: string; from: unknown; to: unknown };
+
+export type AnchorKind = Anchor["kind"];
+
+export type NoteReply = { author: NoteAuthor; text: string; at: string };
+
+export type Note = {
+ id: string;
+ status: NoteStatus;
+ author: NoteAuthor;
+ text: string;
+ at: string;
+ updatedAt: string;
+ anchor: Anchor;
+ replies: NoteReply[];
+ resolvedAt?: string;
+ resolvedBy?: NoteAuthor;
+};
+
+export type NotesSubject =
+ | { kind: "article"; site: string; report: string }
+ | { kind: "video-project"; project: string };
+
+/** Which file an agent should edit to act on a note. Filled by umtool, correctable. */
+export type NotesSource = { draft?: string; generator?: string; manifest?: string; how?: string };
+
+export type NotesDoc = {
+ format: "umtool-notes";
+ version: 1;
+ subject: NotesSubject;
+ source?: NotesSource;
+ notes: Note[];
+};
+
+/** What GET /api/notes returns. `token` goes back on every write (409 when stale). */
+export type NotesRead = {
+ subject: NotesSubject;
+ file: string;
+ token: string;
+ doc: NotesDoc | null;
+ source: NotesSource | null;
+ error?: string;
+};
+
+/** One write. The server stamps author "operator" on everything the UI sends. */
+export type NoteOp =
+ | { op: "add"; text: string; anchor: Anchor }
+ | { op: "edit"; id: string; text?: string; anchor?: Anchor }
+ | { op: "status"; id: string; status: NoteStatus }
+ | { op: "reply"; id: string; text: string; resolve?: boolean }
+ | { op: "delete"; id: string }
+ | { op: "delete-reply"; id: string; index: number }
+ | { op: "source"; source: NotesSource };
+
+export const openCount = (doc: NotesDoc | null | undefined): number =>
+ doc ? doc.notes.filter((n) => n.status === "open").length : 0;
diff --git a/umtool/lib/annotations/useNotes.ts b/umtool/lib/annotations/useNotes.ts
@@ -0,0 +1,92 @@
+"use client";
+
+import { useCallback, useEffect, useRef, useState } from "react";
+import type { NoteOp, NotesRead, Note, NotesDoc } from "./types";
+
+// The page's handle on one notes.json: read it, write one op at a time with the
+// token it read, and on a 409 re-read and say so rather than retrying blind --
+// an agent may have replied in between, and the operator should see that reply
+// before their edit lands on top of it.
+
+export type NotesTarget = { article: string } | { project: string };
+
+export function notesQuery(target: NotesTarget): string {
+ return "article" in target ? `article=${encodeURIComponent(target.article)}` : `project=${encodeURIComponent(target.project)}`;
+}
+
+export type UseNotes = {
+ doc: NotesDoc | null;
+ notes: Note[];
+ source: NotesRead["source"];
+ error: string | null;
+ loading: boolean;
+ busy: boolean;
+ /** Apply one op. Resolves to the note touched (null on delete), or null after an error (shown in `error`). */
+ write: (op: NoteOp) => Promise<Note | null>;
+ reload: () => Promise<void>;
+};
+
+export function useNotes(target: NotesTarget | null, { initial }: { initial?: NotesRead | null } = {}): UseNotes {
+ const [read, setRead] = useState<NotesRead | null>(initial ?? null);
+ const [error, setError] = useState<string | null>(initial?.error ?? null);
+ const [loading, setLoading] = useState(!initial && !!target);
+ const [busy, setBusy] = useState(false);
+ const token = useRef<string | null>(initial?.token ?? null);
+ const query = target ? notesQuery(target) : null;
+
+ const reload = useCallback(async () => {
+ if (!query) return;
+ setLoading(true);
+ try {
+ const res = await fetch(`/api/notes?${query}`, { cache: "no-store" });
+ const j = await res.json();
+ if (!res.ok) throw new Error(j.error ?? res.statusText);
+ token.current = j.token;
+ setRead(j);
+ setError(j.error ?? null);
+ } catch (err) {
+ setError(err instanceof Error ? err.message : String(err));
+ } finally {
+ setLoading(false);
+ }
+ }, [query]);
+
+ useEffect(() => {
+ if (!initial) void reload();
+ // `initial` is the server render's read; only a changed target re-reads.
+ // eslint-disable-next-line react-hooks/exhaustive-deps
+ }, [reload]);
+
+ const write = useCallback(
+ async (op: NoteOp): Promise<Note | null> => {
+ if (!query) return null;
+ setBusy(true);
+ try {
+ const res = await fetch(`/api/notes?${query}`, {
+ method: "POST",
+ headers: { "content-type": "application/json" },
+ body: JSON.stringify({ token: token.current, op }),
+ });
+ const j = await res.json();
+ if (res.status === 409) {
+ await reload();
+ setError(`${j.error ?? "the notes changed"} -- reloaded; check and try again`);
+ return null;
+ }
+ if (!res.ok) throw new Error(j.error ?? res.statusText);
+ token.current = j.token;
+ setRead((r) => (r ? { ...r, doc: j.doc, token: j.token, source: j.doc?.source ?? r.source } : r));
+ setError(null);
+ return j.note ?? null;
+ } catch (err) {
+ setError(err instanceof Error ? err.message : String(err));
+ return null;
+ } finally {
+ setBusy(false);
+ }
+ },
+ [query, reload],
+ );
+
+ return { doc: read?.doc ?? null, notes: read?.doc?.notes ?? [], source: read?.source ?? null, error, loading, busy, write, reload };
+}
diff --git a/umtool/lib/articles/article.ts b/umtool/lib/articles/article.ts
@@ -0,0 +1,191 @@
+import { readFile, stat } from "node:fs/promises";
+import path from "node:path";
+import {
+ buildReportPageView,
+ REPORT_PAGE_FORMAT,
+ REPORT_VIEWS_VERSION,
+ type RecordView,
+ type ReportPageView,
+} from "yt-dlp-transcript-common/lib/report/views";
+import type { Report } from "yt-dlp-transcript-common/lib/report/schema";
+import { platformMomentUrl } from "yt-dlp-transcript-common/lib/momentUrl";
+import { readAllPosts } from "yt-dlp-transcript-common/lib/posts-server";
+import type { Post } from "yt-dlp-transcript-common/lib/posts";
+import { readCues } from "@/lib/projects/report.mjs";
+import { CHANNELS_DIR } from "@/lib/paths";
+
+// A report's PAGE VIEW, built the way the export site builds it
+// (common/lib/report/views.ts buildReportPageView) but resolved against what
+// umtool can read without the LMDB index or a compose run: each cited record's
+// own files on disk. compose's resolveSiteReports is not reused: it verifies,
+// prepares and THROWS on any problem, and half of what this page is for is
+// reading drafts that have problems.
+//
+// A record resolves from its transcript.cues.json (title, date, the uploader's
+// display name, webpageUrl -- through lib/projects/report.mjs readCues, which is
+// memoised on the file's mtime), else its metadata.info.json, else the
+// citation's own label, speaker and date. A post resolves from the channel's
+// posts. Nothing here fails the page: a record that cannot be read is a card
+// with less on it.
+
+const isoDay = (d: unknown): string | undefined => {
+ const s = typeof d === "string" ? d : "";
+ if (/^\d{8}$/.test(s)) return `${s.slice(0, 4)}-${s.slice(4, 6)}-${s.slice(6, 8)}`;
+ if (/^\d{4}-\d{2}-\d{2}/.test(s)) return s.slice(0, 10);
+ return undefined;
+};
+
+export const recordDir = (channel: string, id: string) =>
+ path.join(/* turbopackIgnore: true */ CHANNELS_DIR, channel, "data", id);
+
+type RecordMeta = { title?: string; date?: string; channelTitle?: string; webpageUrl?: string };
+
+const metaMemo = new Map<string, { key: string; value: RecordMeta | null }>();
+
+/** What a record says about itself; null when its directory holds neither file. */
+export async function recordMeta(channel: string, id: string): Promise<RecordMeta | null> {
+ if (!/^[A-Za-z0-9_.@-]+$/.test(channel) || !/^[A-Za-z0-9_.@-]+$/.test(id)) return null;
+ const dir = recordDir(channel, id);
+ const cues = (await readCues(path.join(/* turbopackIgnore: true */ dir, "transcript.cues.json"))) as
+ | { title?: string; uploadDate?: string; webpageUrl?: string; channel?: string }
+ | null;
+ if (cues && (cues.title || cues.webpageUrl)) {
+ return { title: cues.title, date: isoDay(cues.uploadDate), channelTitle: cues.channel, webpageUrl: cues.webpageUrl };
+ }
+ const file = path.join(/* turbopackIgnore: true */ dir, "metadata.info.json");
+ const st = await stat(/* turbopackIgnore: true */ file).catch(() => null);
+ if (!st) return null;
+ const key = `${Math.round(st.mtimeMs)}-${st.size}`;
+ const hit = metaMemo.get(file);
+ if (hit?.key === key) return hit.value;
+ let value: RecordMeta | null = null;
+ try {
+ const m = JSON.parse(await readFile(/* turbopackIgnore: true */ file, "utf8"));
+ value = {
+ title: typeof m.title === "string" ? m.title : undefined,
+ date: isoDay(m.upload_date),
+ channelTitle: typeof m.uploader === "string" ? m.uploader : typeof m.channel === "string" ? m.channel : undefined,
+ webpageUrl: typeof m.webpage_url === "string" ? m.webpage_url : undefined,
+ };
+ } catch {
+ value = null;
+ }
+ metaMemo.set(file, { key, value });
+ return value;
+}
+
+// A channel's posts, by id. A big X archive is thousands of posts, so one read
+// per channel per minute.
+const postsMemo = new Map<string, { at: number; value: Promise<Map<string, Post>> }>();
+export function channelPosts(channel: string): Promise<Map<string, Post>> {
+ const hit = postsMemo.get(channel);
+ if (hit && Date.now() - hit.at < 60_000) return hit.value;
+ const value = readAllPosts(path.join(/* turbopackIgnore: true */ CHANNELS_DIR, channel))
+ .then((list) => new Map(list.map((p) => [p.id, p])))
+ .catch(() => new Map<string, Post>());
+ postsMemo.set(channel, { at: Date.now(), value });
+ return value;
+}
+
+/** The capture screenshot of a post, if the editor took one. */
+export async function postShotFile(channel: string, id: string): Promise<string | null> {
+ if (!/^[A-Za-z0-9_.@-]+$/.test(channel) || !/^[A-Za-z0-9_.@-]+$/.test(id)) return null;
+ const file = path.join(/* turbopackIgnore: true */ CHANNELS_DIR, channel, "posts-media", id, "shot.png");
+ return (await stat(/* turbopackIgnore: true */ file).catch(() => null))?.isFile() ? file : null;
+}
+
+export const corpusMediaUrl = (abs: string) => `/api/sites/media?corpus=${encodeURIComponent(abs)}`;
+
+type Cite = NonNullable<Report["citations"]>[string];
+
+async function recordViewOf(
+ c: Cite,
+ metaOf: (channel: string, id: string) => Promise<RecordMeta | null> = recordMeta,
+): Promise<RecordView | undefined> {
+ if (c.kind === "video" || c.kind === "audio") {
+ const m = await metaOf(c.channel, c.id);
+ return {
+ channel: c.channel,
+ id: c.id,
+ ...(m?.channelTitle || c.speaker ? { channelTitle: m?.channelTitle ?? c.speaker } : {}),
+ ...(m?.title || c.label ? { title: m?.title ?? c.label } : {}),
+ ...(m?.date || c.date ? { date: m?.date ?? c.date } : {}),
+ ...(m?.webpageUrl ? { originalUrl: platformMomentUrl(m.webpageUrl, null, c.start) ?? m.webpageUrl } : {}),
+ };
+ }
+ if (c.kind === "post") {
+ const post = (await channelPosts(c.channel)).get(c.id);
+ return {
+ channel: c.channel,
+ id: c.id,
+ ...(post?.authorName || post?.author ? { channelTitle: post.authorName ?? post.author } : {}),
+ ...(post?.createdAt ? { date: post.createdAt.slice(0, 10) } : c.date ? { date: c.date } : {}),
+ ...(post?.platform ? { platform: post.platform } : {}),
+ ...(post?.url ? { originalUrl: post.url } : {}),
+ };
+ }
+ return undefined;
+}
+
+/**
+ * The page view of a report, or -- when the report's citations cannot be built
+ * into one (a draft naming a source it does not define) -- the same view with
+ * no citations, and the reason.
+ */
+export async function articleView(report: Report): Promise<{ view: ReportPageView; error: string | null }> {
+ const records = new Map<Cite, RecordView>();
+ const posts = new Map<Cite, { author?: string; text?: string; shot?: string }>();
+ // A fact-check cites the same few records hundreds of times: read each once,
+ // all at the same time (recordMeta memoises on the file, not on a promise).
+ const metas = new Map<string, Promise<RecordMeta | null>>();
+ const metaOf = (channel: string, id: string) => {
+ const k = `${channel}/${id}`;
+ if (!metas.has(k)) metas.set(k, recordMeta(channel, id));
+ return metas.get(k)!;
+ };
+ await Promise.all(
+ Object.values(report.citations ?? {}).map(async (c) => {
+ const r = await recordViewOf(c, metaOf);
+ if (r) records.set(c, r);
+ if (c.kind === "post") {
+ const post = (await channelPosts(c.channel)).get(c.id);
+ const shot = await postShotFile(c.channel, c.id);
+ posts.set(c, {
+ ...(post ? { author: post.authorName ?? post.author, text: post.text } : {}),
+ ...(shot ? { shot: corpusMediaUrl(shot) } : {}),
+ });
+ }
+ }),
+ );
+ try {
+ const view = buildReportPageView(report, {
+ record: (c) => records.get(c as Cite) ?? { channel: c.channel, id: c.id },
+ post: (c) => posts.get(c as Cite),
+ });
+ return { view, error: null };
+ } catch (err) {
+ const view = {
+ format: REPORT_PAGE_FORMAT,
+ version: REPORT_VIEWS_VERSION,
+ id: report.id,
+ kind: report.kind,
+ ...(report.series ? { series: report.series } : {}),
+ title: report.title,
+ ...(report.subtitle ? { subtitle: report.subtitle } : {}),
+ ...(report.summary ? { summary: report.summary } : {}),
+ ...(report.method ? { method: report.method } : {}),
+ ...(report.published ? { published: report.published } : {}),
+ ...(report.updated ? { updated: report.updated } : {}),
+ sources: {},
+ verdicts: {},
+ citations: {},
+ sections: report.sections.map((s) => ({
+ id: s.id,
+ title: s.title,
+ ...(s.body ? { body: s.body } : {}),
+ claims: (s.claims ?? []).map((cl) => ({ ...cl, citations: cl.citations ?? [] })),
+ })),
+ } as unknown as ReportPageView;
+ return { view, error: err instanceof Error ? err.message : String(err) };
+ }
+}
diff --git a/umtool/lib/articles/evidence.ts b/umtool/lib/articles/evidence.ts
@@ -0,0 +1,182 @@
+import { readFile } from "node:fs/promises";
+import path from "node:path";
+import type { Report } from "yt-dlp-transcript-common/lib/report/schema";
+import { momentKeyOf } from "yt-dlp-transcript-common/lib/citations/moments";
+import { evidenceSpan, resolveEvidenceSource } from "yt-dlp-transcript-common/lib/evidenceClip-server";
+import { reportMediaDir, reportMediaIndexFile } from "yt-dlp-transcript-common/publish/reportMedia";
+import { readCues } from "@/lib/projects/report.mjs";
+import { CHANNELS_DIR } from "@/lib/paths";
+import { channelPosts, corpusMediaUrl, postShotFile, recordDir, recordMeta } from "./article";
+import { sitesPaths } from "./sites";
+
+// What a citation's EVIDENCE panel shows: the cited seconds with the transcript
+// around them, and something to play -- found, never fetched.
+//
+// Playback, best first:
+// 1. prepared the site's own evidence clip (.export-index/sites/<site>/
+// report-media/, what the build publishes), when prepare has run;
+// 2. window a clip window the editor fetched into data/<id>/clips/;
+// 3. saved a saved video or audio file in data/<id>/ (through its media
+// tier link -- read, never written);
+// 4. otherwise nothing to play, and the line that fetches it through the
+// editor: the MCP `fetch_clip` tool. umtool's own fetch client
+// (api/report/fetch) is a manifest's, so it cannot ask for a
+// window no project names. Never yt-dlp.
+
+export const CONTEXT_CUES = 6;
+
+export type EvidenceCue = { start: number; end: number; text: string; cited: boolean };
+
+export type EvidencePlay = {
+ kind: "prepared" | "window" | "saved" | "audio";
+ url: string;
+ /** Seconds into the file where the cited span starts. */
+ offset: number;
+ audio: boolean;
+ label: string;
+};
+
+export type Evidence = {
+ cite: string;
+ kind: string;
+ quote: string;
+ speaker?: string;
+ date?: string;
+ label?: string;
+ originalUrl?: string;
+ record?: { channel: string; id: string; title?: string; channelTitle?: string };
+ start?: number;
+ end?: number;
+ cues: EvidenceCue[];
+ cuesNote?: string;
+ play: EvidencePlay | null;
+ fetchLine?: string;
+ post?: { author?: string; text?: string; url?: string; shot?: string };
+};
+
+type Cite = NonNullable<Report["citations"]>[string];
+
+/** The ±CONTEXT_CUES cues around [start, end], the overlapping ones marked. */
+export async function cueContext(channel: string, id: string, start: number, end: number) {
+ const file = path.join(/* turbopackIgnore: true */ recordDir(channel, id), "transcript.cues.json");
+ const doc = (await readCues(file)) as { cues?: { start: number; end: number; text?: string }[] } | null;
+ const cues = doc?.cues ?? [];
+ if (!cues.length) return { cues: [] as EvidenceCue[], note: doc ? "no cues" : "no transcript.cues.json" };
+ const EPS = 0.05;
+ let first = cues.findIndex((c) => c.end > start + EPS);
+ if (first < 0) first = cues.length - 1;
+ let last = first;
+ while (last + 1 < cues.length && cues[last + 1].start < end - EPS) last += 1;
+ const from = Math.max(0, first - CONTEXT_CUES);
+ const to = Math.min(cues.length - 1, last + CONTEXT_CUES);
+ return {
+ cues: cues.slice(from, to + 1).map((c, i) => ({
+ start: c.start,
+ end: c.end,
+ text: String(c.text ?? "").replace(/\s+/g, " ").trim(),
+ cited: from + i >= first && from + i <= last,
+ })),
+ };
+}
+
+async function preparedClip(siteId: string, c: Cite): Promise<EvidencePlay | null> {
+ if (c.kind !== "video" && c.kind !== "audio") return null;
+ const key = momentKeyOf(c);
+ if (!key) return null;
+ try {
+ const index = JSON.parse(await readFile(/* turbopackIgnore: true */ reportMediaIndexFile(sitesPaths(), siteId), "utf8"));
+ const entry = index?.moments?.[key];
+ if (!entry || (entry.kind !== "video" && entry.kind !== "audio") || typeof entry.file !== "string") return null;
+ const abs = path.join(/* turbopackIgnore: true */ reportMediaDir(sitesPaths(), siteId), entry.file);
+ let from = evidenceSpan(c).from;
+ try {
+ const side = JSON.parse(await readFile(/* turbopackIgnore: true */ abs.replace(/\.(mp4|m4a)$/, ".json"), "utf8"));
+ if (typeof side?.span?.from === "number") from = side.span.from;
+ } catch {
+ // no sidecar: the citation's own pad is the best guess
+ }
+ return {
+ kind: "prepared",
+ url: `/api/sites/media?site=${encodeURIComponent(siteId)}&moment=${encodeURIComponent(key)}`,
+ offset: Math.max(0, c.start - from),
+ audio: entry.kind === "audio",
+ label: "prepared evidence clip",
+ };
+ } catch {
+ return null;
+ }
+}
+
+async function corpusClip(c: Cite): Promise<EvidencePlay | null> {
+ if (c.kind !== "video" && c.kind !== "audio") return null;
+ const span = { from: c.start, to: c.end };
+ for (const audio of c.kind === "audio" ? [true] : [false, true]) {
+ const hit = await resolveEvidenceSource({ channelsDir: CHANNELS_DIR, slug: c.channel, id: c.id, span, audio }).catch(() => null);
+ if (!hit) continue;
+ const kind = hit.kind === "corpus-window" ? "window" : hit.kind === "saved-video" ? "saved" : "audio";
+ return {
+ kind,
+ url: corpusMediaUrl(hit.path),
+ offset: Math.max(0, c.start - hit.windowStart),
+ audio: hit.kind === "audio",
+ label: kind === "window" ? `clip window ${hit.name}` : kind === "saved" ? `saved ${hit.name}` : `audio ${hit.name}`,
+ };
+ }
+ return null;
+}
+
+/** The MCP line that fetches this span through the editor. */
+export function fetchClipLine(c: { channel: string; id: string; start: number; end: number }, reason: string): string {
+ const s = (n: number) => Number(n.toFixed(2));
+ return `fetch_clip ${JSON.stringify({ channel: c.channel, video: c.id, start: s(c.start), end: s(c.end), reason })}`;
+}
+
+export async function citationEvidence(siteId: string, report: Report, citeId: string): Promise<Evidence | null> {
+ const c = report.citations?.[citeId];
+ if (!c) return null;
+ const base: Evidence = {
+ cite: citeId,
+ kind: c.kind,
+ quote: c.quote,
+ ...(c.speaker ? { speaker: c.speaker } : {}),
+ ...(c.date ? { date: c.date } : {}),
+ ...(c.label ? { label: c.label } : {}),
+ cues: [],
+ play: null,
+ };
+ if (c.kind === "video" || c.kind === "audio") {
+ const meta = await recordMeta(c.channel, c.id);
+ const ctx = await cueContext(c.channel, c.id, c.start, c.end);
+ const play = (await preparedClip(siteId, c)) ?? (await corpusClip(c));
+ return {
+ ...base,
+ start: c.start,
+ end: c.end,
+ record: { channel: c.channel, id: c.id, ...(meta?.title ? { title: meta.title } : {}), ...(meta?.channelTitle ? { channelTitle: meta.channelTitle } : {}) },
+ ...(meta?.webpageUrl ? { originalUrl: meta.webpageUrl } : {}),
+ cues: ctx.cues,
+ ...(ctx.note ? { cuesNote: ctx.note } : {}),
+ play,
+ ...(play ? {} : { fetchLine: fetchClipLine(c, `${siteId}/${report.id} ${citeId}`) }),
+ };
+ }
+ if (c.kind === "post") {
+ const post = (await channelPosts(c.channel)).get(c.id);
+ const shot = await postShotFile(c.channel, c.id);
+ return {
+ ...base,
+ record: { channel: c.channel, id: c.id },
+ ...(post?.url ? { originalUrl: post.url } : {}),
+ post: {
+ ...(post ? { author: post.authorName ?? post.author, text: post.text, url: post.url } : {}),
+ ...(shot ? { shot: corpusMediaUrl(shot) } : {}),
+ },
+ };
+ }
+ if (c.kind === "page") return { ...base, originalUrl: c.url };
+ if (c.kind === "source") {
+ const s = report.sources?.[c.source];
+ return { ...base, ...(s?.url ? { originalUrl: s.url } : {}), ...(s?.title ? { label: c.label ?? s.title } : {}) };
+ }
+ return base;
+}
diff --git a/umtool/lib/articles/files.ts b/umtool/lib/articles/files.ts
@@ -0,0 +1,77 @@
+import { readFile } from "node:fs/promises";
+import path from "node:path";
+import { REPORTS_ROOT } from "@/lib/paths";
+import { listTakes, readVerdicts } from "@/lib/report/takes.mjs";
+import { siteWorkspaces, tildify } from "./sources.mjs";
+import { untildify } from "./links.mjs";
+import { workspaceFile, workspaceFiles } from "./workspace.mjs";
+
+// The page side of lib/articles/workspace.mjs and the video projects' takes.
+
+export type WorkspaceListing = {
+ /** The workspace's name under REPORTS_ROOT (what /api/sites/workspace takes). */
+ name: string;
+ dir: string;
+ files: { rel: string; kind: "draft" | "out" | "doc"; bytes: number; mtimeMs: number }[];
+};
+
+export async function listingOf(dir: string): Promise<WorkspaceListing> {
+ const files = (await workspaceFiles(dir)).map(({ rel, kind, bytes, mtimeMs }: { rel: string; kind: WorkspaceListing["files"][number]["kind"]; bytes: number; mtimeMs: number }) => ({ rel, kind, bytes, mtimeMs }));
+ return { name: path.basename(dir), dir: tildify(dir), files };
+}
+
+/** Every workspace a site's articles were written in. */
+export async function siteWorkspaceListings(siteId: string, reportIds: string[]): Promise<WorkspaceListing[]> {
+ const dirs: string[] = await siteWorkspaces(siteId, reportIds, { reportsRoot: REPORTS_ROOT });
+ return Promise.all(dirs.map(listingOf));
+}
+
+/** One article's workspace (its source's), or null. */
+export async function articleWorkspaceListing(workspace: string | undefined | null): Promise<WorkspaceListing | null> {
+ if (!workspace) return null;
+ const abs = untildify(workspace);
+ if (path.dirname(abs) !== path.resolve(REPORTS_ROOT)) return null;
+ return listingOf(abs);
+}
+
+export type OpenedFile =
+ | { ws: string; rel: string; kind: "md"; text: string }
+ | { ws: string; rel: string; kind: "json"; value: unknown; text: string }
+ | { ws: string; rel: string; kind: "html"; url: string }
+ | { ws: string; rel: string; kind: "error"; message: string };
+
+const MAX = 2 * 1024 * 1024;
+
+/** A workspace file opened for the page, or an error saying why not. */
+export async function openWorkspaceFile(ws: string, rel: string): Promise<OpenedFile> {
+ const f = await workspaceFile(ws, rel);
+ if (!f) return { ws, rel, kind: "error", message: "not a listed workspace file" };
+ if (rel.endsWith(".html")) {
+ return { ws, rel, kind: "html", url: `/api/sites/workspace?ws=${encodeURIComponent(ws)}&rel=${encodeURIComponent(rel)}` };
+ }
+ if (f.bytes > MAX) return { ws, rel, kind: "error", message: `${f.bytes} bytes; too big to show` };
+ const text = await readFile(/* turbopackIgnore: true */ f.real, "utf8");
+ if (rel.endsWith(".json")) {
+ try {
+ return { ws, rel, kind: "json", value: JSON.parse(text), text };
+ } catch (err) {
+ return { ws, rel, kind: "error", message: `not JSON: ${(err as Error).message}` };
+ }
+ }
+ return { ws, rel, kind: "md", text };
+}
+
+export type TakeTally = { takes: number; like: number; maybe: number; no: number; skipped: number };
+
+/** How many takes a video project has, and how they were judged. */
+export async function takeTally(dir: string): Promise<TakeTally> {
+ const [t, v] = await Promise.all([listTakes(dir), readVerdicts(dir)]);
+ const rows = Object.values(v) as { verdict: string | null }[];
+ return {
+ takes: t.takes.length,
+ like: rows.filter((r) => r.verdict === "like").length,
+ maybe: rows.filter((r) => r.verdict === "maybe").length,
+ no: rows.filter((r) => r.verdict === "no").length,
+ skipped: t.skipped.length,
+ };
+}
diff --git a/umtool/lib/articles/links.mjs b/umtool/lib/articles/links.mjs
@@ -0,0 +1,100 @@
+// Which umtool report-video project is an article's video.
+//
+// Two ways, in order:
+//
+// 1. The manifest says so: a top-level `"article": "<site>/<report>"` in
+// video.manifest.json. build-video.mjs never reads top-level keys it does
+// not know (it reads `generatedBy` no more than this), so the key costs
+// the render nothing. A manifest that names an article is linked to it
+// and to nothing else.
+// 2. The slug matches: the manifest's `slug` (else the project directory's
+// name) is the report id, `polemic-<id>`, or the id without `polemic-` --
+// and the project lives in the same workspace as the article's draft
+// (lib/articles/sources.mjs). A UNIQUE match is linked; two or more are
+// only "possible", and the page says so rather than picking one.
+//
+// Plain ESM, so `umtool notes` can name an article's video too.
+import { readFile } from "node:fs/promises";
+import os from "node:os";
+import path from "node:path";
+import { REPORTS_ROOT } from "../paths.mjs";
+import { projectRefs } from "../projects/core.mjs";
+import { kindLinksArticles } from "../projects/kinds.mjs";
+import { idKeys } from "./sources.mjs";
+
+const CACHE_MS = 30_000;
+/** @type {Map<string, { at: number, value: Promise<any[]> }>} */
+const cache = new Map();
+
+/** `~/x` back to an absolute path. */
+export function untildify(p) {
+ if (typeof p !== "string") return p;
+ return p === "~" || p.startsWith("~/") ? path.join(/* turbopackIgnore: true */ os.homedir(), p.slice(2)) : p;
+}
+
+/**
+ * Every report-video project with what linking needs from its manifest:
+ * `{ id, dir, name, slug, article, generatedBy, title }`.
+ *
+ * @param {string} [reportsRoot]
+ */
+export function videoProjects(reportsRoot = REPORTS_ROOT) {
+ const hit = cache.get(reportsRoot);
+ if (hit && Date.now() - hit.at < CACHE_MS) return hit.value;
+ const value = (async () => {
+ const out = [];
+ for (const p of await projectRefs(reportsRoot)) {
+ if (!kindLinksArticles(p.kind)) continue;
+ let m = {};
+ try {
+ m = JSON.parse(await readFile(/* turbopackIgnore: true */ path.join(/* turbopackIgnore: true */ p.dir, "video.manifest.json"), "utf8"));
+ } catch {
+ // a project with no readable manifest still links by its directory name
+ }
+ out.push({
+ id: p.id,
+ dir: p.dir,
+ name: p.name,
+ slug: typeof m.slug === "string" && m.slug ? m.slug : p.name,
+ article: typeof m.article === "string" ? m.article : null,
+ generatedBy: typeof m.generatedBy === "string" ? m.generatedBy : null,
+ title: typeof m.title === "string" ? m.title : p.name,
+ });
+ }
+ return out.sort((a, b) => a.id.localeCompare(b.id));
+ })();
+ cache.set(reportsRoot, { at: Date.now(), value });
+ value.catch(() => cache.delete(reportsRoot));
+ return value;
+}
+
+export function clearLinksCache() {
+ cache.clear();
+}
+
+const inside = (root, p) => p === root || p.startsWith(root + path.sep);
+
+/**
+ * The projects linked to one article: `{ linked, possible, how }`.
+ *
+ * @param {string} siteId
+ * @param {string} reportId
+ * @param {{ reportsRoot?: string, workspace?: string | null, projects?: any[] }} [opts]
+ * `workspace` is the article's workspace dir (sourceFor's, `~` allowed); without
+ * one, slug matches are only ever "possible".
+ */
+export async function linkedProjects(siteId, reportId, { reportsRoot = REPORTS_ROOT, workspace = null, projects } = {}) {
+ const all = projects ?? (await videoProjects(reportsRoot));
+ const key = `${siteId}/${reportId}`;
+ const declared = all.filter((p) => p.article === key);
+ if (declared.length) return { linked: declared, possible: [], how: "manifest names the article" };
+
+ const keys = idKeys(reportId);
+ // A manifest that names a DIFFERENT article is never a slug match.
+ const bySlug = all.filter((p) => !p.article && keys.has(p.slug));
+ const ws = workspace ? untildify(workspace) : null;
+ const inWs = ws ? bySlug.filter((p) => inside(ws, p.dir)) : [];
+ if (inWs.length === 1) return { linked: inWs, possible: [], how: `slug ${inWs[0].slug} in the article's workspace` };
+ const possible = inWs.length > 1 ? inWs : bySlug;
+ return { linked: [], possible, how: possible.length ? `${possible.length} project(s) share the slug` : "no project" };
+}
diff --git a/umtool/lib/articles/sites.ts b/umtool/lib/articles/sites.ts
@@ -0,0 +1,177 @@
+import { readFile, stat } from "node:fs/promises";
+import path from "node:path";
+import { getPaths, type Paths } from "yt-dlp-transcript-common/lib/paths";
+import { getSite, isListedSite, isPrivateSite, listSites, type Site } from "yt-dlp-transcript-common/lib/site";
+import { listReportDirs, siteReportDir } from "yt-dlp-transcript-common/publish/reportMedia";
+import { parseReport } from "yt-dlp-transcript-common/lib/report/validate";
+import type { Report } from "yt-dlp-transcript-common/lib/report/schema";
+import { CHANNELS_DIR, REPORTS_ROOT, SITES_DIR } from "@/lib/paths";
+import { readNotes } from "@/lib/annotations/store.mjs";
+import { corpusNotesFile } from "@/lib/paths";
+import type { NotesDoc } from "@/lib/annotations/types";
+import { sourceFor } from "./sources.mjs";
+import { linkedProjects, videoProjects } from "./links.mjs";
+
+// Every site's articles, as umtool reads them: the site list from common
+// (site.json through getSite, so a default is the editor's default), each
+// report directory through common's listReportDirs (the editor's report tab
+// uses the same enumerator), each report.json through the report document's
+// own validator, and published-or-draft from the site's `reports` order. Plus
+// what only umtool knows: its notes, its source draft, its video project.
+//
+// READ-ONLY. The one thing umtool writes under SITES_DIR is a notes.json, and
+// that goes through lib/annotations, never here.
+
+/** common's Paths, with the two roots umtool resolves itself (and e2e confines). */
+export function sitesPaths(): Paths {
+ return { ...getPaths(), sitesDir: SITES_DIR, channelsDir: CHANNELS_DIR };
+}
+
+export type ArticleStatus = "published" | "draft";
+
+export type ProjectLinkRow = { id: string; title: string; slug: string };
+
+export type ArticleRow = {
+ site: string;
+ id: string;
+ title: string;
+ series: string | null;
+ kind: string | null;
+ status: ArticleStatus;
+ published: string | null;
+ updated: string | null;
+ /** report.json's mtime, for "updated" when the report names no date. */
+ mtimeMs: number | null;
+ citations: number;
+ notes: number;
+ openNotes: number;
+ hasVideo: boolean;
+ hasPoster: boolean;
+ /** report.json missing, unparseable, or invalid: the first problem, else null. */
+ problem: string | null;
+ problems: number;
+ source: { draft?: string; generator?: string; how?: string; workspace?: string } | null;
+ projects: { linked: ProjectLinkRow[]; possible: ProjectLinkRow[] };
+};
+
+export type SiteRow = {
+ siteId: string;
+ title: string;
+ private: boolean;
+ listed: boolean;
+ search: boolean;
+ published: number;
+ drafts: number;
+ openNotes: number;
+ articles: ArticleRow[];
+};
+
+const exists = (p: string) => stat(/* turbopackIgnore: true */ p).then((s) => s.isFile(), () => false);
+
+export type ArticleRead = {
+ report: Report | null;
+ problems: { path?: string; message: string }[];
+ mtimeMs: number | null;
+};
+
+/** One report.json, read and validated; never throws. */
+export async function readReportFile(siteId: string, reportId: string): Promise<ArticleRead> {
+ const file = path.join(/* turbopackIgnore: true */ siteReportDir(sitesPaths(), siteId, reportId), "report.json");
+ let text: string;
+ let mtimeMs: number | null = null;
+ try {
+ const [t, st] = await Promise.all([readFile(/* turbopackIgnore: true */ file, "utf8"), stat(/* turbopackIgnore: true */ file)]);
+ text = t;
+ mtimeMs = Math.round(st.mtimeMs);
+ } catch {
+ return { report: null, problems: [{ message: "no report.json" }], mtimeMs: null };
+ }
+ let raw: unknown;
+ try {
+ raw = JSON.parse(text);
+ } catch (err) {
+ return { report: null, problems: [{ message: `report.json is not JSON: ${(err as Error).message}` }], mtimeMs };
+ }
+ const parsed = parseReport(raw, { id: reportId });
+ if (!parsed.ok) return { report: null, problems: parsed.problems, mtimeMs };
+ return { report: parsed.value, problems: parsed.problems, mtimeMs };
+}
+
+export async function readArticleNotes(siteId: string, reportId: string): Promise<{ doc: NotesDoc | null; token: string; error?: string }> {
+ const file = corpusNotesFile(siteId, reportId);
+ if (!file) return { doc: null, token: "absent" };
+ return readNotes(file);
+}
+
+async function articleRow(site: Site, id: string, published: Set<string>, projects: Awaited<ReturnType<typeof videoProjects>>): Promise<ArticleRow> {
+ const [read, notes, source] = await Promise.all([
+ readReportFile(site.siteId, id),
+ readArticleNotes(site.siteId, id),
+ sourceFor(site.siteId, id, { reportsRoot: REPORTS_ROOT }),
+ ]);
+ const dir = siteReportDir(sitesPaths(), site.siteId, id);
+ const r = read.report;
+ const links = await linkedProjects(site.siteId, id, { workspace: source?.workspace ?? null, projects });
+ const row = (p: { id: string; title: string; slug: string }) => ({ id: p.id, title: p.title, slug: p.slug });
+ return {
+ site: site.siteId,
+ id,
+ title: r?.title ?? id,
+ series: r?.series ?? null,
+ kind: r?.kind ?? null,
+ status: published.has(id) ? "published" : "draft",
+ published: r?.published ?? null,
+ updated: r?.updated ?? r?.published ?? null,
+ mtimeMs: read.mtimeMs,
+ citations: Object.keys(r?.citations ?? {}).length,
+ notes: notes.doc?.notes.length ?? 0,
+ openNotes: notes.doc?.notes.filter((n) => n.status === "open").length ?? 0,
+ hasVideo: !!r?.video?.src && (await exists(path.join(/* turbopackIgnore: true */ dir, r.video.src))),
+ hasPoster: !!r?.video?.poster && (await exists(path.join(/* turbopackIgnore: true */ dir, r.video.poster))),
+ problem: read.problems[0]?.message ?? null,
+ problems: read.problems.length,
+ source,
+ projects: { linked: links.linked.map(row), possible: links.possible.map(row) },
+ };
+}
+
+function siteFlags(site: Site) {
+ return { private: isPrivateSite(site), listed: isListedSite(site), search: site.search !== false };
+}
+
+/** One site with every article, published (in the site's order) then drafts (by id). */
+export async function readSiteRow(site: Site, projects?: Awaited<ReturnType<typeof videoProjects>>): Promise<SiteRow> {
+ const all = projects ?? (await videoProjects(REPORTS_ROOT));
+ const published = new Set(site.reports ?? []);
+ const dirs = await listReportDirs(sitesPaths(), site.siteId);
+ const ids = [...(site.reports ?? []), ...dirs.filter((d) => !published.has(d))];
+ const articles = await Promise.all(ids.map((id) => articleRow(site, id, published, all)));
+ return {
+ siteId: site.siteId,
+ title: site.siteTitle || site.siteId,
+ ...siteFlags(site),
+ published: articles.filter((a) => a.status === "published").length,
+ drafts: articles.filter((a) => a.status === "draft").length,
+ openNotes: articles.reduce((n, a) => n + a.openNotes, 0),
+ articles,
+ };
+}
+
+/** Every site, private first, then by id. */
+export async function listSiteRows(): Promise<SiteRow[]> {
+ const projects = await videoProjects(REPORTS_ROOT);
+ const sites = listSites(sitesPaths());
+ const rows = await Promise.all(sites.map((s) => readSiteRow(s, projects)));
+ return rows.sort((a, b) => Number(b.private) - Number(a.private) || a.siteId.localeCompare(b.siteId));
+}
+
+/** A site by id, or null (a bad id or no site.json). */
+export function siteById(siteId: string): Site | null {
+ try {
+ const paths = sitesPaths();
+ if (!listSites(paths).some((s) => s.siteId === siteId)) return null;
+ return getSite(siteId, paths);
+ } catch {
+ return null;
+ }
+}
diff --git a/umtool/lib/articles/sources.mjs b/umtool/lib/articles/sources.mjs
@@ -0,0 +1,219 @@
+// Which file an agent should EDIT to change an article.
+//
+// A report.json under transcripts/sites/ is generated: a workspace under
+// ~/reports keeps the draft (`<ws>/polemics/drafts/<slug>.json`, the source of
+// truth) and a generator script that writes report.json from it
+// (`<ws>/polemics/make-site.py`, `<ws>/site/polemics.py`, ...). A note that
+// says "fix this sentence" is useless to an agent that edits report.json -- the
+// next generator run puts the old sentence back. So every notes.json carries a
+// `source` block naming the draft and the generator, found here.
+//
+// The match is a heuristic, written down with its reason (`how`), and the
+// agent may correct it (`umtool notes source`):
+//
+// draft a drafts/*.json whose `id`, or file name, is the report id --
+// allowing for a `polemic-` prefix on either side (candalyzer's
+// polemic-israel is drafts/israel.json with id polemic-israel;
+// jeralyzer-private's `blame` is drafts/blame.json with id
+// polemic-blame). Several matches: the one whose workspace has a
+// generator naming the site wins; still several, none is chosen.
+// generator a *.py / *.mts under <ws>/polemics or <ws>/site that names the
+// site (or the report id), preferring one that names the report
+// id itself, then one that reads the drafts. Backups
+// (`make-report.pre-2026-10-05.py`: a second dot) are skipped.
+//
+// Cheap: one readdir per workspace and one read per generator, cached for 30 s
+// like the project walk.
+import { readdir, readFile, stat } from "node:fs/promises";
+import os from "node:os";
+import path from "node:path";
+import { REPORTS_ROOT } from "../paths.mjs";
+
+const CACHE_MS = 30_000;
+const GEN_DIRS = ["polemics", "site"];
+const GEN_EXT = /^[^.]+\.(py|mts|mjs|sh)$/;
+const MAX_GEN_BYTES = 2 * 1024 * 1024;
+
+/** `~/…` for a path under the home directory: what a note shows an agent. */
+export function tildify(abs) {
+ const home = os.homedir();
+ return abs === home || abs.startsWith(home + path.sep) ? `~${abs.slice(home.length)}` : abs;
+}
+
+/** The ids a report or draft may go by: itself, without `polemic-`, with it. */
+export function idKeys(id) {
+ const bare = id.replace(/^polemic-/, "");
+ return new Set([id, bare, `polemic-${bare}`]);
+}
+
+/** Does `text` name `id` as a whole token (so `jasolyzer` is not `jasolyzer-private`)? */
+export function names(text, id) {
+ const esc = id.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
+ return new RegExp(`(?<![A-Za-z0-9_-])${esc}(?![A-Za-z0-9_-])`).test(text);
+}
+
+const isDir = (p) => stat(/* turbopackIgnore: true */ p).then((s) => s.isDirectory(), () => false);
+
+/** @type {Map<string, { at: number, value: Promise<any[]> }>} */
+const cache = new Map();
+
+/**
+ * Every workspace under `reportsRoot` that has drafts or a generator dir:
+ * `{ dir, name, drafts: [{ file, slug, id }], generators: [{ file, text }] }`.
+ *
+ * @param {string} [reportsRoot]
+ */
+export function scanWorkspaces(reportsRoot = REPORTS_ROOT) {
+ const hit = cache.get(reportsRoot);
+ if (hit && Date.now() - hit.at < CACHE_MS) return hit.value;
+ const value = scan(reportsRoot);
+ cache.set(reportsRoot, { at: Date.now(), value });
+ value.catch(() => cache.delete(reportsRoot));
+ return value;
+}
+
+/** Forget the scan (a test that writes a fixture, then reads it). */
+export function clearSourcesCache() {
+ cache.clear();
+}
+
+async function scan(reportsRoot) {
+ const entries = await readdir(/* turbopackIgnore: true */ reportsRoot, { withFileTypes: true }).catch(() => []);
+ const out = [];
+ for (const e of entries) {
+ if (e.name.startsWith(".") || e.name === "data") continue;
+ const dir = path.join(/* turbopackIgnore: true */ reportsRoot, e.name);
+ if (!(e.isDirectory() || (e.isSymbolicLink() && (await isDir(dir))))) continue;
+ const drafts = [];
+ const draftsDir = path.join(/* turbopackIgnore: true */ dir, "polemics", "drafts");
+ for (const f of await readdir(/* turbopackIgnore: true */ draftsDir).catch(() => [])) {
+ if (!f.endsWith(".json")) continue;
+ const file = path.join(/* turbopackIgnore: true */ draftsDir, f);
+ let id = null;
+ try {
+ const j = JSON.parse(await readFile(/* turbopackIgnore: true */ file, "utf8"));
+ if (j && typeof j.id === "string") id = j.id;
+ } catch {
+ // an unreadable draft still matches by its file name
+ }
+ drafts.push({ file, slug: f.slice(0, -5), id });
+ }
+ const generators = [];
+ for (const g of GEN_DIRS) {
+ const gdir = path.join(/* turbopackIgnore: true */ dir, g);
+ for (const f of await readdir(/* turbopackIgnore: true */ gdir).catch(() => [])) {
+ if (!GEN_EXT.test(f)) continue;
+ const file = path.join(/* turbopackIgnore: true */ gdir, f);
+ const st = await stat(/* turbopackIgnore: true */ file).catch(() => null);
+ if (!st?.isFile() || st.size > MAX_GEN_BYTES) continue;
+ generators.push({ file, text: await readFile(/* turbopackIgnore: true */ file, "utf8").catch(() => "") });
+ }
+ }
+ if (drafts.length || generators.length) {
+ drafts.sort((a, b) => a.file.localeCompare(b.file));
+ generators.sort((a, b) => a.file.localeCompare(b.file));
+ out.push({ dir, name: e.name, drafts, generators });
+ }
+ }
+ return out.sort((a, b) => a.name.localeCompare(b.name));
+}
+
+/** Is `draft` this report's? */
+function draftMatches(draft, keys) {
+ return (draft.id !== null && keys.has(draft.id)) || keys.has(draft.slug) || keys.has(`polemic-${draft.slug}`);
+}
+
+/**
+ * The source of truth for one report, or null when no workspace claims it.
+ *
+ * @param {string} siteId
+ * @param {string} reportId
+ * @param {{ reportsRoot?: string }} [opts]
+ * @returns {Promise<{ draft?: string, generator?: string, how: string, workspace?: string } | null>}
+ */
+export async function sourceFor(siteId, reportId, { reportsRoot = REPORTS_ROOT } = {}) {
+ const workspaces = await scanWorkspaces(reportsRoot);
+ const keys = idKeys(reportId);
+ const namesSite = (ws) => ws.generators.some((g) => names(g.text, siteId));
+
+ const cands = [];
+ for (const ws of workspaces) {
+ for (const d of ws.drafts) {
+ if (!draftMatches(d, keys)) continue;
+ const score = (namesSite(ws) ? 4 : 0) + (d.id === reportId ? 2 : 0) + (d.slug === reportId.replace(/^polemic-/, "") ? 1 : 0);
+ cands.push({ ws, d, score });
+ }
+ }
+ cands.sort((a, b) => b.score - a.score);
+ const top = cands[0];
+ const unique = top && (cands.length === 1 || cands[1].score < top.score);
+ const draft = unique ? top : null;
+
+ const pool = draft ? [draft.ws] : workspaces;
+ let gen = null;
+ let genScore = 0;
+ for (const ws of pool) {
+ for (const g of ws.generators) {
+ const site = names(g.text, siteId);
+ const report = names(g.text, reportId);
+ if (!site && !report) continue;
+ const base = path.basename(g.file);
+ const score =
+ (report ? 3 : 0) + (site ? 2 : 0) + (draft && /\bdrafts\b/.test(g.text) ? 1 : 0) + (/^(make-site|polemics)\./.test(base) ? 0.5 : 0);
+ if (score > genScore) {
+ gen = { ws, g };
+ genScore = score;
+ }
+ }
+ }
+ // Without a draft, a generator that only names the SITE is every report's
+ // generator and says nothing about this one; keep it only if it names the id.
+ // A bare id (`deleted`, `poker`) is also an English word, so naming it is
+ // only evidence when the generator names the site too.
+ if (!draft && gen && !(names(gen.g.text, reportId) && (reportId.includes("-") || names(gen.g.text, siteId)))) gen = null;
+ if (!draft && !gen) {
+ if (cands.length > 1) {
+ return { how: `several drafts match ${reportId}: ${cands.map((c) => tildify(c.d.file)).join(", ")}; none chosen` };
+ }
+ return null;
+ }
+
+ const how = [];
+ if (draft) {
+ const by = draft.d.id === reportId ? `id ${reportId}` : draft.d.id && keys.has(draft.d.id) ? `id ${draft.d.id}` : `file name ${draft.d.slug}`;
+ how.push(`draft matched by ${by}`);
+ if (cands.length > 1) how.push(`preferred over ${cands.length - 1} other`);
+ }
+ if (gen) {
+ const named = [siteId, reportId].filter((id) => names(gen.g.text, id));
+ how.push(`generator names ${named.join(" and ")}`);
+ }
+ const out = { how: how.join("; ") };
+ if (draft) out.draft = tildify(draft.d.file);
+ if (gen) out.generator = tildify(gen.g.file);
+ out.workspace = tildify((draft?.ws ?? gen?.ws).dir);
+ return out;
+}
+
+/**
+ * The workspace directories a site's articles come from: every workspace that
+ * holds a matched draft for one of `reportIds`, or a generator that names the
+ * site. Absolute paths, sorted.
+ *
+ * @param {string} siteId
+ * @param {string[]} reportIds
+ * @param {{ reportsRoot?: string }} [opts]
+ */
+export async function siteWorkspaces(siteId, reportIds, { reportsRoot = REPORTS_ROOT } = {}) {
+ const workspaces = await scanWorkspaces(reportsRoot);
+ const dirs = new Set();
+ for (const ws of workspaces) {
+ if (ws.generators.some((g) => names(g.text, siteId))) dirs.add(ws.dir);
+ }
+ for (const id of reportIds) {
+ const s = await sourceFor(siteId, id, { reportsRoot });
+ const ws = s?.draft ? workspaces.find((w) => w.drafts.some((d) => tildify(d.file) === s.draft)) : null;
+ if (ws) dirs.add(ws.dir);
+ }
+ return [...dirs].sort();
+}
diff --git a/umtool/lib/articles/sources.test.mjs b/umtool/lib/articles/sources.test.mjs
@@ -0,0 +1,69 @@
+// Finding an article's draft and generator, on the three workspace layouts
+// the live private sites use.
+//
+// Run with: pnpm test:scripts
+import assert from "node:assert/strict";
+import { mkdir, mkdtemp, rm, writeFile } from "node:fs/promises";
+import { tmpdir } from "node:os";
+import path from "node:path";
+import test from "node:test";
+import { clearSourcesCache, names, siteWorkspaces, sourceFor } from "./sources.mjs";
+
+async function put(file, text) {
+ await mkdir(path.dirname(file), { recursive: true });
+ await writeFile(file, text);
+}
+
+async function fixture() {
+ const root = await mkdtemp(path.join(tmpdir(), "umtool-sources-"));
+ // candalyzer: drafts/<bare>.json with id polemic-<bare>; generator site/polemics.py;
+ // a hand-written fact-check whose generator names it; a backup generator.
+ await put(path.join(root, "candace/polemics/drafts/israel.json"), JSON.stringify({ id: "polemic-israel" }));
+ await put(path.join(root, "candace/site/polemics.py"), 'OUT = ".../sites/candalyzer/reports"\nfor f in DRAFTS.glob("drafts/*.json"): rid = f"polemic-{slug}"\n');
+ await put(path.join(root, "candace/site/make-report.py"), 'SITE = "candalyzer"\nREPORT = "deconstruction-fact-check"\n');
+ await put(path.join(root, "candace/site/make-report.pre-2026-10-05.py"), 'REPORT = "polemic-israel" # candalyzer\n');
+ // jasolyzer-private: drafts/<bare>.json, generator polemics/make-site.py
+ await put(path.join(root, "pirate/polemics/drafts/skg.json"), JSON.stringify({ id: "polemic-skg" }));
+ await put(path.join(root, "pirate/polemics/make-site.py"), 'SITE = "jasolyzer-private"\nrid = f"polemic-{slug}" # from drafts\n');
+ // jeralyzer-private: the report id is BARE (`blame`), the draft id is polemic-blame
+ await put(path.join(root, "quartering/polemics/drafts/blame.json"), JSON.stringify({ id: "polemic-blame" }));
+ await put(path.join(root, "quartering/polemics/make-site.py"), 'SITE = "jeralyzer-private"\nrid = slug # drafts\n');
+ // a decoy: a public site's workspace with a same-named draft
+ await put(path.join(root, "decoy/polemics/drafts/blame.json"), JSON.stringify({ id: "blame" }));
+ await put(path.join(root, "decoy/polemics/make-site.py"), 'SITE = "jeralyzer"\n');
+ clearSourcesCache();
+ return root;
+}
+
+test("each layout finds its draft and generator", async () => {
+ const root = await fixture();
+ const o = { reportsRoot: root };
+ const isr = await sourceFor("candalyzer", "polemic-israel", o);
+ assert.equal(isr.draft, path.join(root, "candace/polemics/drafts/israel.json"));
+ assert.equal(isr.generator, path.join(root, "candace/site/polemics.py"));
+ assert.match(isr.how, /draft matched by id polemic-israel/);
+
+ const fc = await sourceFor("candalyzer", "deconstruction-fact-check", o);
+ assert.equal(fc.draft, undefined);
+ assert.equal(fc.generator, path.join(root, "candace/site/make-report.py"));
+
+ const skg = await sourceFor("jasolyzer-private", "polemic-skg", o);
+ assert.equal(skg.draft, path.join(root, "pirate/polemics/drafts/skg.json"));
+ assert.equal(skg.generator, path.join(root, "pirate/polemics/make-site.py"));
+
+ // The decoy's draft id is exactly `blame`, but its workspace never names the site.
+ const blame = await sourceFor("jeralyzer-private", "blame", o);
+ assert.equal(blame.draft, path.join(root, "quartering/polemics/drafts/blame.json"));
+ assert.equal(blame.generator, path.join(root, "quartering/polemics/make-site.py"));
+ assert.match(blame.how, /preferred over 1 other/);
+
+ assert.equal(await sourceFor("candalyzer", "no-such-report", o), null);
+ assert.deepEqual(await siteWorkspaces("jasolyzer-private", ["polemic-skg"], o), [path.join(root, "pirate")]);
+ await rm(root, { recursive: true });
+});
+
+test("names() matches whole ids only", () => {
+ assert.equal(names('"jasolyzer-private"', "jasolyzer"), false);
+ assert.equal(names("sites/jasolyzer/reports", "jasolyzer"), true);
+ assert.equal(names("polemic-blame", "blame"), false);
+});
diff --git a/umtool/lib/articles/workspace.mjs b/umtool/lib/articles/workspace.mjs
@@ -0,0 +1,65 @@
+// The files an article was WRITTEN from, for reading beside it: a workspace's
+// drafts, the rendered drafts and briefs under polemics/, and the notes a
+// workspace keeps at its top level. Read-only -- this lists and serves, it
+// never writes a workspace.
+//
+// polemics/drafts/*.json the drafts (source of truth)
+// polemics/out/*.{md,html} what the drafts render to
+// polemics/*.md BRIEF.md, NOTES.md, PRIVACY-SWEEP.md, …
+// <ws>/{NOTES,BRIEF,LEADS,PLAN,PRIVACY-SWEEP}.md
+// <ws>/site/PLAN.md, <ws>/site/MERGE-PLAN.md
+import { readdir, realpath, stat } from "node:fs/promises";
+import path from "node:path";
+import { REPORTS_ROOT } from "../paths.mjs";
+
+export const TOP_LEVEL = ["NOTES.md", "BRIEF.md", "LEADS.md", "PLAN.md", "PRIVACY-SWEEP.md"];
+const SITE_LEVEL = ["PLAN.md", "MERGE-PLAN.md"];
+export const WORKSPACE_EXT = /\.(md|html|json)$/;
+
+const statFile = (p) => stat(/* turbopackIgnore: true */ p).then((s) => (s.isFile() ? s : null), () => null);
+
+/**
+ * Every listed file of one workspace: `{ rel, abs, kind, bytes, mtimeMs }`,
+ * `kind` one of "draft" | "out" | "doc". Sorted: drafts, docs, outputs; by name.
+ *
+ * @param {string} wsDir
+ */
+export async function workspaceFiles(wsDir) {
+ const out = [];
+ const add = async (rel, kind) => {
+ const abs = path.join(/* turbopackIgnore: true */ wsDir, rel);
+ const st = await statFile(abs);
+ if (st) out.push({ rel, abs, kind, bytes: st.size, mtimeMs: Math.round(st.mtimeMs) });
+ };
+ const ls = (rel) => readdir(/* turbopackIgnore: true */ path.join(/* turbopackIgnore: true */ wsDir, rel)).catch(() => []);
+ for (const f of await ls("polemics/drafts")) if (f.endsWith(".json")) await add(`polemics/drafts/${f}`, "draft");
+ for (const f of await ls("polemics/out")) if (/\.(md|html)$/.test(f)) await add(`polemics/out/${f}`, "out");
+ for (const f of await ls("polemics")) if (f.endsWith(".md")) await add(`polemics/${f}`, "doc");
+ for (const f of TOP_LEVEL) await add(f, "doc");
+ for (const f of SITE_LEVEL) await add(`site/${f}`, "doc");
+ const rank = { draft: 0, doc: 1, out: 2 };
+ return out.sort((a, b) => rank[a.kind] - rank[b.kind] || a.rel.localeCompare(b.rel));
+}
+
+/**
+ * A workspace file a client named, or null: `ws` must be a directory directly
+ * under REPORTS_ROOT and `rel` one of the files workspaceFiles lists for it,
+ * and its REAL path must stay inside the workspace.
+ *
+ * @param {string} ws the workspace's name under REPORTS_ROOT
+ * @param {string} rel
+ * @param {{ reportsRoot?: string }} [opts]
+ */
+export async function workspaceFile(ws, rel, { reportsRoot = REPORTS_ROOT } = {}) {
+ if (typeof ws !== "string" || !/^[A-Za-z0-9][A-Za-z0-9._-]*$/.test(ws) || ws === "..") return null;
+ if (typeof rel !== "string" || rel.includes("\0") || rel.split("/").some((s) => s === ".." || s === "" || s === ".")) return null;
+ const dir = path.join(/* turbopackIgnore: true */ reportsRoot, ws);
+ const listed = (await workspaceFiles(dir)).find((f) => f.rel === rel);
+ if (!listed) return null;
+ const [realDir, realAbs] = await Promise.all([
+ realpath(/* turbopackIgnore: true */ dir).catch(() => null),
+ realpath(/* turbopackIgnore: true */ listed.abs).catch(() => null),
+ ]);
+ if (!realDir || !realAbs || !realAbs.startsWith(realDir + path.sep)) return null;
+ return { ...listed, real: realAbs };
+}
diff --git a/umtool/lib/decisions.ts b/umtool/lib/decisions.ts
@@ -4,6 +4,7 @@ import { DEFAULT_TARGET, loudnessVerdict } from "./loudness-types";
import { buildStatus, readManifest } from "./manifest";
import { readSpec, validateSpec } from "./spec";
import { readNotes, type NoteMap } from "./notes";
+import { listNotesFiles } from "./annotations/targets.mjs";
import { acceptedFor, readThumbAccepted, readThumbManifest, thumbNamesFor } from "./thumbs";
// ---------------------------------------------------------------------------
@@ -298,3 +299,79 @@ export function decisionsMarkdown(items: Decision[]): string[] {
}
return out;
}
+
+// ---------------------------------------------------------------------------
+// Open NOTES -- on an article (sites/<site>/reports/<id>/notes.json) or on a
+// report-video project (<project>/notes.json) -- are open decisions: somebody
+// asked for a change and nobody has answered it. One row per open note, kind
+// `open-note`, linked to the note on its page. A notes.json that does not
+// parse is blocking: nothing can write to it until somebody fixes it by hand.
+//
+// An ARTICLE is not a project, so its row's `project` is its page's path
+// (`sites/<site>/<report>`), which is also where its href points.
+// ---------------------------------------------------------------------------
+
+const firstLine = (s: string, max = 140) => {
+ const line = s.split("\n").find((l) => l.trim()) ?? "";
+ return line.length > max ? `${line.slice(0, max - 1)}…` : line;
+};
+
+function anchorLabel(a: { kind: string; [k: string]: unknown }): string {
+ switch (a.kind) {
+ case "text":
+ return `“${firstLine(String(a.quote ?? ""), 48)}”`;
+ case "cite":
+ return `cite ${a.cite}`;
+ case "section":
+ return `section ${a.section}`;
+ case "moment":
+ return `${a.file} @ ${Number(a.t).toFixed(1)}s`;
+ case "entry":
+ return `entry ${a.entry}`;
+ case "take":
+ return `take ${a.take}`;
+ case "edit":
+ return `edit ${a.field}${a.entry ? ` on ${a.entry}` : ""}`;
+ default:
+ return "whole";
+ }
+}
+
+export async function noteDecisions(): Promise<Decision[]> {
+ const files = await listNotesFiles().catch(() => []);
+ const out: Decision[] = [];
+ for (const f of files) {
+ const article = f.kind === "article";
+ const project = article ? `sites/${f.id}` : f.id;
+ const page = article ? `/sites/${f.id}` : `/browse/${f.id}`;
+ if (f.error || !f.doc) {
+ out.push({
+ kind: "unreadable-notes",
+ project,
+ projectKind: article ? "article" : ((f as { projectKind?: string }).projectKind ?? "project"),
+ target: "notes.json",
+ why: `notes.json does not parse (${f.error ?? "unknown"}); nothing will write to it until it is fixed`,
+ href: page,
+ severity: "blocking",
+ at: Date.now(),
+ });
+ continue;
+ }
+ for (const n of f.doc.notes) {
+ if (n.status !== "open") continue;
+ const replies = n.replies.length ? ` · ${n.replies.length} repl${n.replies.length === 1 ? "y" : "ies"}` : "";
+ const at = Date.parse(n.updatedAt);
+ out.push({
+ kind: "open-note",
+ project,
+ projectKind: article ? "article" : ((f as { projectKind?: string }).projectKind ?? "project"),
+ target: anchorLabel(n.anchor as { kind: string }),
+ why: `${firstLine(n.text)}${replies}`,
+ href: `${page}?note=${encodeURIComponent(n.id)}`,
+ severity: "open",
+ at: Number.isFinite(at) ? at : 0,
+ });
+ }
+ }
+ return out;
+}
diff --git a/umtool/lib/paths.mjs b/umtool/lib/paths.mjs
@@ -5,6 +5,7 @@
// `umtool ls` and the page it is supposed to describe. lib/paths.ts re-exports
// everything here with types; nothing computes a root twice.
import { existsSync } from "node:fs";
+import { lstat, realpath } from "node:fs/promises";
import os from "node:os";
import path from "node:path";
import { SONG_DATA, SONG_REPORTS } from "../song/paths.mjs";
@@ -208,12 +209,74 @@ export const CHANNELS_DIR = path.resolve(
: path.join(/* turbopackIgnore: true */ REPO_ROOT, "transcripts", "channels")),
);
+// The SITES -- `transcripts/sites/<site>/`, each a site.json and its reports
+// (`reports/<id>/report.json`, `video.mp4`, `poster.jpg`). Resolved the way
+// CHANNELS_DIR is, and the way common/lib/paths.ts resolves its sitesDir, so
+// one `SITES_DIR` confines both. Readable only; the ONE file in it umtool may
+// write is a report's notes.json, and only through isCorpusNotesFile below.
+export const SITES_DIR = path.resolve(
+ /* turbopackIgnore: true */
+ process.env.SITES_DIR ??
+ (process.env.TRANSCRIPTS_DIR
+ ? path.join(/* turbopackIgnore: true */ process.env.TRANSCRIPTS_DIR, "sites")
+ : path.join(/* turbopackIgnore: true */ REPO_ROOT, "transcripts", "sites")),
+);
+
export const READ_ROOTS = dedupe(
process.env.MIX_ROOTS
? process.env.MIX_ROOTS.split(":").filter(Boolean)
- : [SONG_REPORTS, REPORTS_ROOT, SONG_DATA, SONG_SCRATCH, CHANNELS_DIR, MEDIA_ROOT],
+ : [SONG_REPORTS, REPORTS_ROOT, SONG_DATA, SONG_SCRATCH, CHANNELS_DIR, MEDIA_ROOT, SITES_DIR],
);
+/** A site id or a report id: one lowercase url-safe segment (common/lib/report/schema.ts REPORT_ID_RE). */
+export const SEGMENT_RE = /^[a-z0-9][a-z0-9-]{0,63}$/;
+export const NOTES_FILENAME = "notes.json";
+
+/**
+ * The ONE corpus path umtool may write: `SITES_DIR/<site>/reports/<id>/notes.json`.
+ *
+ * Lexically: exactly four segments under SITES_DIR, the second `reports`, the
+ * site and report ids in the report-id grammar, the file named notes.json, and
+ * no `..` or doubled separator anywhere. Then on disk: the report directory's
+ * REAL path must be the lexical one under the real SITES_DIR -- a report dir
+ * that is a symlink out of its site (or a site dir that is one) is refused --
+ * and it must already exist: a note never creates a report directory. A
+ * notes.json that exists and is not a plain file (a symlink) is refused too.
+ *
+ * WRITE_ROOTS is untouched: nothing else in the corpus becomes writable.
+ *
+ * @param {string} abs
+ * @param {{ sitesDir?: string }} [opts]
+ * @returns {Promise<boolean>}
+ */
+export async function isCorpusNotesFile(abs, { sitesDir = SITES_DIR } = {}) {
+ if (typeof abs !== "string" || !path.isAbsolute(abs) || abs.includes("\0")) return false;
+ if (path.resolve(/* turbopackIgnore: true */ abs) !== abs) return false;
+ const root = path.resolve(/* turbopackIgnore: true */ sitesDir);
+ const rel = path.relative(/* turbopackIgnore: true */ root, abs);
+ if (!rel || rel.startsWith("..") || path.isAbsolute(rel)) return false;
+ const parts = rel.split(path.sep);
+ if (parts.length !== 4) return false;
+ const [site, reports, report, name] = parts;
+ if (!SEGMENT_RE.test(site) || reports !== "reports" || !SEGMENT_RE.test(report) || name !== NOTES_FILENAME) {
+ return false;
+ }
+ const [realRoot, realDir] = await Promise.all([
+ realpath(/* turbopackIgnore: true */ root).catch(() => null),
+ realpath(/* turbopackIgnore: true */ path.dirname(/* turbopackIgnore: true */ abs)).catch(() => null),
+ ]);
+ if (!realRoot || !realDir) return false;
+ if (realDir !== path.join(/* turbopackIgnore: true */ realRoot, site, "reports", report)) return false;
+ const st = await lstat(/* turbopackIgnore: true */ abs).catch(() => null);
+ return !st || st.isFile();
+}
+
+/** `SITES_DIR/<site>/reports/<report>/notes.json`, or null for a bad id. */
+export function corpusNotesFile(site, report, { sitesDir = SITES_DIR } = {}) {
+ if (!SEGMENT_RE.test(String(site)) || !SEGMENT_RE.test(String(report))) return null;
+ return path.join(/* turbopackIgnore: true */ path.resolve(/* turbopackIgnore: true */ sitesDir), site, "reports", report, NOTES_FILENAME);
+}
+
export const WRITE_ROOTS = dedupe(
process.env.MIX_WRITE_ROOTS
? process.env.MIX_WRITE_ROOTS.split(":").filter(Boolean)
diff --git a/umtool/lib/paths.ts b/umtool/lib/paths.ts
@@ -20,7 +20,9 @@ import path from "node:path";
// ---------------------------------------------------------------------------
export {
CACHE_DIR,
+ CHANNELS_DIR,
cacheFile,
+ corpusNotesFile,
INDEX_DIR,
MEDIA_ROOT,
MEDIA_ROOTS,
@@ -31,8 +33,10 @@ export {
SONG_DATA,
SONG_REPORTS,
SONG_SCRATCH,
+ SITES_DIR,
WRITE_ROOTS,
inside,
+ isCorpusNotesFile,
labelFor,
mediaMirror,
resolveInRoots,
diff --git a/umtool/lib/projects.ts b/umtool/lib/projects.ts
@@ -5,7 +5,7 @@ import { collapseFolders, foldersFor, walkProjects } from "./projects/walk.mjs";
import { SONG_KIND } from "./projects/song.mjs";
import { openIndex, signRecord } from "./projects/index-db.mjs";
import { BROWSE_ROOT } from "./browse";
-import { decisionsForSong } from "./decisions";
+import { decisionsForSong, noteDecisions } from "./decisions";
import { listMedia, listMediaUnder, type MediaRow } from "./media";
import type { Decision } from "./decisions";
import type {
@@ -302,8 +302,10 @@ const RANK: Record<string, number> = { blocking: 0, open: 1, info: 2 };
export async function openDecisions(): Promise<Decision[]> {
const refs = await projectRefs();
const per = await Promise.all(refs.map(decisionsForProject));
- return per
- .flat()
+ // Open notes on articles and report videos (lib/decisions.ts noteDecisions):
+ // an article is not a project, so they are added here, not per project.
+ const notes = await noteDecisions();
+ return [...per.flat(), ...notes]
.sort((a, b) => RANK[a.severity] - RANK[b.severity] || b.at - a.at);
}
diff --git a/umtool/lib/projects/kinds.mjs b/umtool/lib/projects/kinds.mjs
@@ -104,6 +104,12 @@ export const PROJECT_KINDS = [
// `brand` offers report-to-video's presets (render.brand); none is the
// default and writes the manifest it always did.
scaffold: { fields: ["from", "siteOrigin", "seed", "brand"], brands: BRAND_CHOICES },
+ // Takes a notes.json beside its manifest (lib/annotations/targets.mjs):
+ // timed notes on its cuts, row notes, take notes, edit notes.
+ notes: true,
+ // Its manifest can be an ARTICLE's video (lib/articles/links.mjs): a
+ // top-level `article`, else its slug against the report id.
+ linksArticles: true,
},
{
id: "song",
@@ -182,6 +188,12 @@ if (process.env.E2E_UMTOOL_EXTRA_KINDS) {
export const kindById = (id) => PROJECT_KINDS.find((k) => k.id === id) ?? null;
+/** Does a project of this kind keep a notes.json (lib/annotations)? Declared on the kind, never branched on its id. */
+export const kindTakesNotes = (id) => kindById(id)?.notes === true;
+
+/** Can a project of this kind be an article's video (lib/articles/links.mjs)? Declared on the kind. */
+export const kindLinksArticles = (id) => kindById(id)?.linksArticles === true;
+
/** What a client component needs, with none of what it must not have. */
export const kindMeta = (k) => ({
id: k.id,
diff --git a/umtool/lib/report/edit-notes.mjs b/umtool/lib/report/edit-notes.mjs
@@ -0,0 +1,174 @@
+// An edit made in umtool to a GENERATED manifest, written down for the agent
+// that generates it.
+//
+// A manifest with `generatedBy` (polemics/video/make-videos.py, …) is rebuilt
+// from the generator's inputs, and the rebuild overwrites whatever was edited
+// here. So edits are still allowed -- the operator is watching the cut and the
+// fix belongs there -- and each one becomes an `edit` note in the project's
+// notes.json (lib/annotations/): which entry, which field, from what, to what.
+// The agent ports it into the generator's inputs (BEATS, drafts) and resolves
+// the note, and the next rebuild keeps it.
+//
+// ONE wrapper does this for every manifest writer (lib/report/guard.ts), by
+// diffing the manifest before and after the write -- so a writer added later
+// is covered without knowing this exists.
+//
+// Repeated saves of one field COALESCE: an open edit note on the same entry and
+// field keeps its original `from` and takes the new `to`, and an edit that
+// returns the field to its `from` deletes the note. Dragging a window five
+// times is one note, and dragging it back is none.
+import { readNotes } from "../annotations/store.mjs";
+import { writeNote } from "../annotations/targets.mjs";
+
+const MAX_VALUE = 3000;
+
+/** A value small enough to keep in a note; a large one is summarised. */
+function keep(v) {
+ if (v === undefined) return null;
+ const s = JSON.stringify(v);
+ if (s.length <= MAX_VALUE) return v;
+ return `(${Array.isArray(v) ? `${v.length} items` : "object"}, ${s.length} characters)`;
+}
+
+const ID_LISTS = ["posts", "ledger"];
+
+const same = (a, b) => JSON.stringify(a) === JSON.stringify(b);
+
+/** Key the timeline by `id` (and its variant, since twins share an id). */
+function entryKey(e) {
+ return e?.variant ? `${e.id}@${e.variant}` : String(e?.id);
+}
+
+/**
+ * Every change between two manifests, as `{ entry?, field, from, to }`.
+ *
+ * timeline entry field { entry: id, field: "<key>" }
+ * an entry added { entry: id, field: "timeline+", from: null, to: <entry> }
+ * an entry removed { entry: id, field: "timeline-", from: <entry>, to: null }
+ * the order { field: "timeline.order", from: [ids], to: [ids] } (survivors only)
+ * a post, a ledger claim { entry: id, field: "posts.<key>" | "posts+" | "posts-" } (ledger alike)
+ * anything else { field: "<top>.<key>" } one level down (render.chrome, provenance.siteOrigin)
+ *
+ * @param {Record<string, any>} before
+ * @param {Record<string, any>} after
+ */
+export function editsBetween(before, after) {
+ const out = [];
+ const a = (before?.timeline ?? []).filter((e) => e && e.id != null);
+ const b = (after?.timeline ?? []).filter((e) => e && e.id != null);
+ const byA = new Map(a.map((e) => [entryKey(e), e]));
+ const byB = new Map(b.map((e) => [entryKey(e), e]));
+ for (const [k, e] of byA) if (!byB.has(k)) out.push({ entry: String(e.id), field: "timeline-", from: keep(e), to: null });
+ for (const [k, e] of byB) if (!byA.has(k)) out.push({ entry: String(e.id), field: "timeline+", from: null, to: keep(e) });
+ const sa = a.map(entryKey).filter((k) => byB.has(k));
+ const sb = b.map(entryKey).filter((k) => byA.has(k));
+ if (!same(sa, sb)) out.push({ field: "timeline.order", from: keep(sa), to: keep(sb) });
+ for (const [k, x] of byA) {
+ const y = byB.get(k);
+ if (!y) continue;
+ for (const f of new Set([...Object.keys(x), ...Object.keys(y)])) {
+ // sectionEnter follows the order; the order change already says it.
+ if (f === "sectionEnter" || same(x[f], y[f])) continue;
+ out.push({ entry: String(x.id), field: f, from: keep(x[f]), to: keep(y[f]) });
+ }
+ }
+
+ // Lists of things with ids -- the posts, the ledger's claims -- by id.
+ for (const list of ID_LISTS) {
+ const pa = new Map((Array.isArray(before?.[list]) ? before[list] : []).map((p) => [String(p?.id), p]));
+ const pb = new Map((Array.isArray(after?.[list]) ? after[list] : []).map((p) => [String(p?.id), p]));
+ for (const [id, p] of pa) if (!pb.has(id)) out.push({ entry: id, field: `${list}-`, from: keep(p), to: null });
+ for (const [id, p] of pb) {
+ const q = pa.get(id);
+ if (!q) {
+ out.push({ entry: id, field: `${list}+`, from: null, to: keep(p) });
+ continue;
+ }
+ for (const f of new Set([...Object.keys(q), ...Object.keys(p)])) {
+ if (!same(q[f], p[f])) out.push({ entry: id, field: `${list}.${f}`, from: keep(q[f]), to: keep(p[f]) });
+ }
+ }
+ }
+
+ for (const top of new Set([...Object.keys(before ?? {}), ...Object.keys(after ?? {})])) {
+ if (top === "timeline" || ID_LISTS.includes(top)) continue;
+ const x = before?.[top];
+ const y = after?.[top];
+ if (same(x, y)) continue;
+ const isObj = (v) => v && typeof v === "object" && !Array.isArray(v);
+ if (isObj(x) && isObj(y)) {
+ for (const f of new Set([...Object.keys(x), ...Object.keys(y)])) {
+ if (!same(x[f], y[f])) out.push({ field: `${top}.${f}`, from: keep(x[f]), to: keep(y[f]) });
+ }
+ } else {
+ out.push({ field: top, from: keep(x), to: keep(y) });
+ }
+ }
+ return out;
+}
+
+/** The sentence an edit note carries; the anchor carries the values. */
+export function editText(edit, generatedBy) {
+ const where = edit.entry ? `${edit.entry} ` : "";
+ const what =
+ edit.field === "timeline+"
+ ? "added to the timeline"
+ : edit.field === "timeline-"
+ ? "removed from the timeline"
+ : edit.field === "timeline.order"
+ ? "the timeline was re-ordered"
+ : edit.field.endsWith("+")
+ ? `${edit.field.slice(0, -1)} entry added`
+ : edit.field.endsWith("-")
+ ? `${edit.field.slice(0, -1)} entry removed`
+ : `${edit.field} changed`;
+ return `${where}${what} in umtool. Port it into the inputs of ${generatedBy}; a rebuild of manifests overwrites it.`;
+}
+
+/**
+ * Write `edits` into a project's notes as `edit` notes, coalescing with open
+ * ones on the same entry and field (see the top of this file). Returns how many
+ * notes were added, updated and deleted. Errors are returned, not thrown: the
+ * manifest write already happened, and a notes file that will not take a note
+ * must not turn a saved edit into a reported failure.
+ *
+ * @param {{ file: string, subject: Record<string, string>, source: () => Promise<any> }} target
+ * @param {Array<{ entry?: string, field: string, from: unknown, to: unknown }>} edits
+ * @param {string} generatedBy
+ */
+export async function recordEdits(target, edits, generatedBy) {
+ const counts = { added: 0, updated: 0, deleted: 0, errors: /** @type {string[]} */ ([]) };
+ for (const edit of edits) {
+ try {
+ const { doc } = await readNotes(target.file);
+ const open = (doc?.notes ?? []).find(
+ (n) =>
+ n.status === "open" &&
+ n.author === "operator" &&
+ n.anchor.kind === "edit" &&
+ n.anchor.field === edit.field &&
+ (n.anchor.entry ?? null) === (edit.entry ?? null),
+ );
+ if (open) {
+ const from = /** @type {any} */ (open.anchor).from;
+ // Put back as it was: the note says nothing -- unless somebody has
+ // already replied to it, and then it stays for them to resolve.
+ if (same(from, edit.to) && !open.replies.length) {
+ await writeNote(target, { op: "delete", id: open.id }, { by: "operator" });
+ counts.deleted += 1;
+ } else {
+ const anchor = { kind: "edit", field: edit.field, from, to: edit.to, ...(edit.entry ? { entry: edit.entry } : {}) };
+ await writeNote(target, { op: "edit", id: open.id, anchor }, { by: "operator" });
+ counts.updated += 1;
+ }
+ continue;
+ }
+ const anchor = { kind: "edit", field: edit.field, from: edit.from, to: edit.to, ...(edit.entry ? { entry: edit.entry } : {}) };
+ await writeNote(target, { op: "add", text: editText(edit, generatedBy), anchor }, { by: "operator" });
+ counts.added += 1;
+ } catch (e) {
+ counts.errors.push(e instanceof Error ? e.message : String(e));
+ }
+ }
+ return counts;
+}
diff --git a/umtool/lib/report/edit-notes.test.mjs b/umtool/lib/report/edit-notes.test.mjs
@@ -0,0 +1,67 @@
+// Edits to a generated manifest, as notes: the diff, and the coalescing.
+//
+// Run with: pnpm test:scripts
+import assert from "node:assert/strict";
+import { mkdtemp, rm } from "node:fs/promises";
+import { tmpdir } from "node:os";
+import path from "node:path";
+import test from "node:test";
+import { readNotes } from "../annotations/store.mjs";
+import { editsBetween, editText, recordEdits } from "./edit-notes.mjs";
+
+const m = () => ({
+ generatedBy: "polemics/video/make-videos.py",
+ render: { fps: 30, chrome: { engine: "hyperframes" } },
+ timeline: [
+ { type: "clip", id: "a", start: 1, end: 5, quote: "q" },
+ { type: "clip", id: "b", start: 10, end: 15 },
+ { type: "card", id: "k" },
+ ],
+ posts: [{ id: "p1", text: "t", attachTo: "a" }],
+ ledger: [{ id: "c1", scope: "x" }],
+});
+
+test("editsBetween names each change by entry and field", () => {
+ const before = m();
+ const after = m();
+ after.timeline[0].quote = "new";
+ after.timeline[0].start = 2;
+ after.timeline = [after.timeline[1], after.timeline[0], { type: "card", id: "k2" }];
+ after.posts[0].attachTo = "b";
+ after.ledger[0].scope = "y";
+ after.render.chrome.layout = "deck";
+ const e = editsBetween(before, after);
+ const keyOf = (x) => `${x.entry ?? ""}|${x.field}`;
+ assert.deepEqual(
+ e.map(keyOf).sort(),
+ ["a|quote", "a|start", "c1|ledger.scope", "k2|timeline+", "k|timeline-", "p1|posts.attachTo", "|render.chrome", "|timeline.order"].sort(),
+ );
+ const quote = e.find((x) => x.field === "quote");
+ assert.deepEqual([quote.from, quote.to], ["q", "new"]);
+ assert.deepEqual(editsBetween(before, m()), []);
+ assert.match(editText({ entry: "a", field: "quote" }, "gen.py"), /^a quote changed in umtool\. Port it into the inputs of gen\.py/);
+});
+
+test("recordEdits adds, coalesces, and deletes a note an edit put back", async () => {
+ const dir = await mkdtemp(path.join(tmpdir(), "umtool-editnotes-"));
+ const target = {
+ file: path.join(dir, "notes.json"),
+ subject: { kind: "video-project", project: "x/y" },
+ source: async () => ({ manifest: "~/x/y/video.manifest.json" }),
+ };
+ try {
+ let c = await recordEdits(target, [{ entry: "a", field: "start", from: 1, to: 2 }], "gen.py");
+ assert.deepEqual([c.added, c.updated, c.deleted], [1, 0, 0]);
+ c = await recordEdits(target, [{ entry: "a", field: "start", from: 2, to: 3 }], "gen.py");
+ assert.deepEqual([c.added, c.updated], [0, 1]);
+ let doc = (await readNotes(target.file)).doc;
+ assert.equal(doc.notes.length, 1);
+ assert.deepEqual([doc.notes[0].anchor.from, doc.notes[0].anchor.to], [1, 3]);
+ assert.equal(doc.source.manifest, "~/x/y/video.manifest.json");
+ c = await recordEdits(target, [{ entry: "a", field: "start", from: 3, to: 1 }], "gen.py");
+ assert.equal(c.deleted, 1);
+ assert.equal((await readNotes(target.file)).doc, null);
+ } finally {
+ await rm(dir, { recursive: true });
+ }
+});
diff --git a/umtool/lib/report/guard.ts b/umtool/lib/report/guard.ts
@@ -0,0 +1,38 @@
+import { projectTargetFor } from "@/lib/annotations/targets.mjs";
+import { readManifest } from "@/lib/projects/report.mjs";
+import { editsBetween, recordEdits } from "./edit-notes.mjs";
+
+// THE one wrapper every manifest writer's route goes through.
+//
+// A manifest with `generatedBy` is rebuilt by its generator, and the rebuild
+// overwrites edits made here (the banner on the project page, the bench and
+// the On-screen section says so). The edit is still made; what this adds is a
+// record of it: the manifest is read before and after the write, and every
+// change becomes an `edit` note in the project's notes.json for the agent to
+// port into the generator's inputs (lib/report/edit-notes.mjs). A hand-edited
+// manifest (no `generatedBy`) is written exactly as before.
+//
+// The notes are written AFTER the manifest and never fail the request: the
+// edit is saved either way, and `editNotes.errors` says when its note is not.
+
+export type EditNotes = { generatedBy: string; added: number; updated: number; deleted: number; errors: string[] };
+
+export async function withEditNotes<T>(
+ project: { id: string; dir: string; kind: string },
+ write: () => Promise<T>,
+): Promise<{ result: T; editNotes: EditNotes | null }> {
+ const before = await readManifest(project.dir);
+ const result = await write();
+ const generatedBy = typeof before?.generatedBy === "string" ? before.generatedBy.trim() : "";
+ if (!generatedBy) return { result, editNotes: null };
+ const after = await readManifest(project.dir);
+ const edits = editsBetween(before, after);
+ if (!edits.length) return { result, editNotes: { generatedBy, added: 0, updated: 0, deleted: 0, errors: [] } };
+ try {
+ const target = await projectTargetFor(project);
+ const counts = await recordEdits(target, edits, generatedBy);
+ return { result, editNotes: { generatedBy, ...counts } };
+ } catch (e) {
+ return { result, editNotes: { generatedBy, added: 0, updated: 0, deleted: 0, errors: [e instanceof Error ? e.message : String(e)] } };
+ }
+}
diff --git a/umtool/lib/report/manifest.mjs b/umtool/lib/report/manifest.mjs
@@ -30,10 +30,12 @@ import {
rolesGaps,
} from "umtool-report-to-video/ledger-totals";
import { isCalendarDate } from "umtool-report-to-video/attribution";
-import { normalizeOnscreen, validateChrome, validatePosts } from "umtool-report-to-video/deck";
+import { normalizeOnscreen, validateChrome, validatePosts, validateTeaser, validateTeasers } from "umtool-report-to-video/deck";
+import { applySectionEnter } from "./sections.mjs";
import { normalizeClaim, validateClaims } from "umtool-report-to-video/factcheck";
import { parseMuteFrom } from "./playback.mjs";
import { DELIVERABLES_MODES } from "./storage.mjs";
+import { createSnapshot, listSnapshots } from "./snapshots.mjs";
// Its own write queue, not lib/state.ts's.
//
@@ -787,3 +789,381 @@ export async function updateStorage(dir, { deliverables } = {}, { token = null }
return { storage: manifest.storage, token: nextToken, changed: true };
});
}
+
+
+// ---------------------------------------------------------------------------
+// STRUCTURE: re-ordering the cut, and the edits that add or remove a thing.
+//
+// Every writer above changes the fields of something already in the manifest.
+// These change WHAT IS IN IT -- the order of the timeline, which entries it
+// holds, a teaser's lines, the posts, the fact-check's labels -- so they are
+// the edits a person wants to take back. Each one snapshots the manifest into
+// revisions/ first (`auto-before-<op>`, at most one per op every two minutes,
+// so a burst of drags is one step back), and `undoStructural` restores the
+// newest of those.
+//
+// Same four rules as the rest of the file: the mtime token, tmp + rename under
+// the lock, 2 dp, and validation by the BUILD's own checks (deck.mjs,
+// factcheck.mjs) before anything is written, so a manifest these accept is one
+// the build accepts.
+//
+// An entry is named by its id. A timeline may repeat an id across variants
+// (`variant: "sourced"` / `"full"` twins); then the caller passes `at`, the
+// index it means, and a stale `at` (the id is not there any more) refuses.
+// ---------------------------------------------------------------------------
+
+export const AUTO_SNAPSHOT_PREFIX = "auto-before-";
+const AUTO_SNAPSHOT_EVERY_MS = 2 * 60 * 1000;
+const ENTRY_ID_RE = /^[A-Za-z0-9_-]{1,64}$/;
+
+/**
+ * Copy the manifest into revisions/ as `auto-before-<op>`, unless one for the
+ * same op was taken in the last two minutes. Never fatal: an undo point that
+ * cannot be taken is reported, and the edit still lands.
+ */
+export async function autoSnapshot(dir, op, { now = Date.now() } = {}) {
+ const label = `${AUTO_SNAPSHOT_PREFIX}${op}`;
+ try {
+ const recent = (await listSnapshots(dir)).find((s) => !s.legacy && s.label === label);
+ if (recent && now - recent.mtimeMs < AUTO_SNAPSHOT_EVERY_MS) return { skipped: true, rel: recent.rel };
+ return { skipped: false, ...(await createSnapshot(dir, { label })) };
+ } catch (e) {
+ return { skipped: true, error: e instanceof Error ? e.message : String(e) };
+ }
+}
+
+/** The index of entry `id`: `at` when it names it, else the only entry with that id. */
+export function entryIndex(timeline, id, at = null) {
+ if (at !== null && at !== undefined) {
+ const i = Number(at);
+ if (!Number.isInteger(i) || timeline[i]?.id !== id) {
+ throw new Error(`timeline[${at}] is not ${id} any more — reload`);
+ }
+ return i;
+ }
+ const hits = [];
+ timeline.forEach((e, i) => {
+ if (e?.id === id) hits.push(i);
+ });
+ if (!hits.length) throw new Error(`no timeline entry with id ${id}`);
+ if (hits.length > 1) throw new Error(`${id} is in the timeline ${hits.length} times — say which (at)`);
+ return hits[0];
+}
+
+/** An id not yet in the timeline: `base`, else `base-2`, `base-3`, … */
+export function freshEntryId(timeline, base) {
+ const clean = String(base).replace(/[^A-Za-z0-9_-]+/g, "-").replace(/^-+|-+$/g, "").slice(0, 56) || "entry";
+ const ids = new Set(timeline.map((e) => e?.id));
+ if (!ids.has(clean)) return clean;
+ for (let n = 2; ; n += 1) if (!ids.has(`${clean}-${n}`)) return `${clean}-${n}`;
+}
+
+/** Every reason the build would refuse the manifest's structure, after an edit. */
+function structureErrors(manifest) {
+ return [
+ ...validatePosts(manifest.posts, manifest.timeline ?? [], manifest.render),
+ ...validateClaims(manifest),
+ ...validateTeasers(manifest),
+ ];
+}
+
+/** The build's sentences, thrown whole so a route can return each one. */
+export class StructureRefused extends Error {
+ /** @param {string[]} errors */
+ constructor(errors) {
+ super(errors.join("; "));
+ this.name = "StructureRefused";
+ this.errors = errors;
+ }
+}
+
+/**
+ * One structural write: token, read, `mutate` (which throws to refuse), the
+ * build's checks, the auto snapshot of the file as it still is, then the write.
+ */
+async function structural(dir, op, token, mutate) {
+ return withManifestLock(async () => {
+ const current = await manifestToken(dir);
+ if (token !== null && current !== token) throw new StaleToken(token, current);
+ const manifest = JSON.parse(await readFile(manifestFile(dir), "utf8"));
+ if (!Array.isArray(manifest.timeline)) manifest.timeline = [];
+ // Only what THIS edit breaks refuses it: a manifest that already carries a
+ // problem the build would name (somebody's hand edit) can still be
+ // re-ordered, and the problem is still the build's to report.
+ const already = new Set(structureErrors(manifest));
+ const result = mutate(manifest);
+ const errors = structureErrors(manifest).filter((e) => !already.has(e));
+ if (errors.length) throw new StructureRefused(errors);
+ applySectionEnter(manifest.timeline);
+ const snapshot = await autoSnapshot(dir, op);
+ const nextToken = await writeManifestAtomic(dir, manifest);
+ return { ...result, snapshot, token: nextToken };
+ });
+}
+
+/**
+ * Move one entry to `toIndex` (its index in the timeline AFTER the move).
+ * `sectionEnter` is recomputed by report-to-video's own rule (sections.mjs),
+ * and only on a manifest that already uses it.
+ *
+ * @param {string} dir
+ * @param {string} id
+ * @param {number} toIndex
+ * @param {{ token?: string | null, at?: number | null }} [opts]
+ */
+export async function moveEntry(dir, id, toIndex, { token = null, at = null } = {}) {
+ return structural(dir, "move", token, (m) => {
+ const from = entryIndex(m.timeline, id, at);
+ const to = Number(toIndex);
+ if (!Number.isInteger(to) || to < 0 || to >= m.timeline.length) {
+ throw new Error(`toIndex must be 0–${m.timeline.length - 1}`);
+ }
+ if (to === from) throw new Error(`${id} is already at ${to}`);
+ const [e] = m.timeline.splice(from, 1);
+ m.timeline.splice(to, 0, e);
+ return { id, from, to };
+ });
+}
+
+/**
+ * Remove one entry. Refused when something still points at it -- a post
+ * attached to a clip, a claim -- in the build's own words.
+ *
+ * @param {string} dir
+ * @param {string} id
+ * @param {{ token?: string | null, at?: number | null }} [opts]
+ */
+export async function removeEntry(dir, id, { token = null, at = null } = {}) {
+ return structural(dir, "remove", token, (m) => {
+ const i = entryIndex(m.timeline, id, at);
+ const [removed] = m.timeline.splice(i, 1);
+ return { id, at: i, removed };
+ });
+}
+
+/** Copy one entry to just after itself, under a fresh id. A copied claim is dropped (one claim, one entry). *
+ * @param {string} dir
+ * @param {string} id
+ * @param {{ token?: string | null, at?: number | null }} [opts]
+ */
+export async function duplicateEntry(dir, id, { token = null, at = null } = {}) {
+ return structural(dir, "duplicate", token, (m) => {
+ const i = entryIndex(m.timeline, id, at);
+ const copy = JSON.parse(JSON.stringify(m.timeline[i]));
+ copy.id = freshEntryId(m.timeline, `${id}-copy`);
+ delete copy.claim;
+ m.timeline.splice(i + 1, 0, copy);
+ return { id: copy.id, at: i + 1, entry: copy };
+ });
+}
+
+/** `<channel>/<video>@<start>-<end>`, in seconds. */
+const CLIP_SPEC_RE = /^([A-Za-z0-9._-]+)\/([^@\s/]+)@(\d+(?:\.\d+)?)-(\d+(?:\.\d+)?)$/;
+
+/**
+ * What `insertEntry` will put in the timeline, checked: a clip spec string
+ * (`<channel>/<video>@<start>-<end>`) becomes a bare clip; an object is an
+ * entry as written, given a fresh id when it has none (or one already taken).
+ *
+ * @param {Array<Record<string, unknown>>} timeline
+ * @param {unknown} spec
+ */
+export function entryFromSpec(timeline, spec) {
+ if (typeof spec === "string") {
+ const mm = spec.trim().match(CLIP_SPEC_RE);
+ if (!mm) throw new Error("a clip is <channel>/<video>@<start>-<end> in seconds");
+ const [, channel, video, s, e] = mm;
+ return clipEntry(timeline, { type: "clip", channel, video, start: Number(s), end: Number(e) });
+ }
+ if (!spec || typeof spec !== "object" || Array.isArray(spec)) throw new Error("an entry is an object, or a clip spec string");
+ const entry = JSON.parse(JSON.stringify(spec));
+ if (JSON.stringify(entry).length > 20000) throw new Error("that entry is too large");
+ if (typeof entry.type !== "string" || !entry.type) throw new Error("an entry needs a type");
+ if (entry.type === "clip") return clipEntry(timeline, entry);
+ const base = typeof entry.id === "string" && ENTRY_ID_RE.test(entry.id) ? entry.id : entry.type;
+ entry.id = freshEntryId(timeline, base);
+ return entry;
+}
+
+function clipEntry(timeline, e) {
+ if (typeof e.video !== "string" || !e.video.trim()) throw new Error("a clip needs its video id");
+ for (const k of ["start", "end"]) {
+ const v = Number(e[k]);
+ if (!Number.isFinite(v) || v < 0) throw new Error(`a clip's ${k} must be a number ≥ 0`);
+ e[k] = round2(v);
+ }
+ if (e.end - e.start < 0.5) throw new Error(`a clip must be at least half a second (${e.start}–${e.end})`);
+ if (e.channel !== undefined && (typeof e.channel !== "string" || !e.channel)) delete e.channel;
+ const base = typeof e.id === "string" && ENTRY_ID_RE.test(e.id) ? e.id : `${e.video}-${Math.floor(e.start)}`;
+ // `type` and `id` first: the manifests are read by humans.
+ const { type: _t, id: _i, ...rest } = e;
+ return { type: "clip", id: freshEntryId(timeline, base), ...rest };
+}
+
+/**
+ * Insert an entry after `afterId` (null or "": at the start). `spec` is a clip
+ * spec string or an entry object (entryFromSpec).
+ *
+ * @param {string} dir
+ * @param {string | null} afterId
+ * @param {unknown} spec
+ * @param {{ token?: string | null, at?: number | null }} [opts] `at` is afterId's index
+ */
+export async function insertEntry(dir, afterId, spec, { token = null, at = null } = {}) {
+ return structural(dir, "insert", token, (m) => {
+ const i = afterId ? entryIndex(m.timeline, afterId, at) + 1 : 0;
+ const entry = entryFromSpec(m.timeline, spec);
+ m.timeline.splice(i, 0, entry);
+ return { id: entry.id, at: i, entry };
+ });
+}
+
+const TEASER_PATCH_KEYS = ["lines", "beat", "dip", "tail", "tailWait"];
+
+/**
+ * Patch one teaser: its lines, beat, dip, tail and tail wait. A key given as
+ * null (or "") is removed; a key not given is kept. Checked with the build's
+ * validateTeaser (which checks the dip too) against the patched entry.
+ *
+ * @param {string} dir
+ * @param {string} id
+ * @param {Record<string, unknown>} patch
+ * @param {{ token?: string | null, at?: number | null }} [opts]
+ */
+export async function updateTeaser(dir, id, patch, { token = null, at = null } = {}) {
+ if (!patch || typeof patch !== "object" || Array.isArray(patch)) throw new Error("a teaser patch is an object");
+ const keys = Object.keys(patch);
+ const bad = keys.filter((k) => !TEASER_PATCH_KEYS.includes(k));
+ if (bad.length) throw new Error(`${bad.join(", ")}: not something this writer changes (${TEASER_PATCH_KEYS.join(", ")})`);
+ if (!keys.length) throw new Error("nothing to change");
+ return structural(dir, "teaser", token, (m) => {
+ const e = m.timeline[entryIndex(m.timeline, id, at)];
+ if (e.type !== "teaser") throw new Error(`${id} is a ${e.type ?? "non-teaser"} entry, not a teaser`);
+ for (const k of keys) {
+ const v = patch[k];
+ if (v === null || v === "" || v === undefined) delete e[k];
+ else if (k === "beat" || k === "tailWait") e[k] = round2(Number(v));
+ else if (k === "dip") {
+ e.dip = typeof v === "object" && v ? { fade: round2(Number(v.fade)), black: round2(Number(v.black)) } : v;
+ } else if (k === "tail") e.tail = String(v);
+ else e[k] = v;
+ }
+ const errors = validateTeaser(e);
+ if (errors.length) throw new StructureRefused(errors);
+ return { id, entry: e };
+ });
+}
+
+const POST_TEXT_KEYS = ["platform", "author", "handle", "date", "text", "url", "shot", "flag", "accent", "logo", "siteChannel", "siteUrl", "postId", "variant"];
+
+/**
+ * Add a post, or replace the one with its id. The value is the post as it
+ * should be stored: an empty optional string is dropped rather than written.
+ * `attachTo` and `hide` are kept from the stored post unless the value names
+ * them. Checked with the build's validatePosts.
+ *
+ * @param {string} dir
+ * @param {Record<string, any>} post
+ * @param {{ token?: string | null }} [opts]
+ */
+export async function upsertPost(dir, post, { token = null } = {}) {
+ if (!post || typeof post !== "object" || Array.isArray(post)) throw new Error("a post is an object");
+ if (typeof post.id !== "string" || !ENTRY_ID_RE.test(post.id)) throw new Error("a post needs an id: letters, digits, dashes, underscores");
+ return structural(dir, "post", token, (m) => {
+ if (!Array.isArray(m.posts)) m.posts = [];
+ const i = m.posts.findIndex((p) => p?.id === post.id);
+ const prev = i >= 0 ? m.posts[i] : {};
+ const next = { id: post.id };
+ for (const k of POST_TEXT_KEYS) {
+ const v = k in post ? post[k] : prev[k];
+ if (v === undefined || v === null || (typeof v === "string" && !v.trim())) continue;
+ next[k] = typeof v === "string" ? v.trim() : v;
+ }
+ for (const k of ["attachTo", "hide"]) {
+ const v = k in post ? post[k] : prev[k];
+ if (v === undefined || v === null || v === "" || v === false) continue;
+ next[k] = v;
+ }
+ if (i >= 0) m.posts[i] = next;
+ else m.posts.push(next);
+ return { post: next, created: i < 0 };
+ });
+}
+
+/** Remove one post by id. *
+ * @param {string} dir
+ * @param {string} id
+ * @param {{ token?: string | null }} [opts]
+ */
+export async function removePost(dir, id, { token = null } = {}) {
+ return structural(dir, "post", token, (m) => {
+ const i = (m.posts ?? []).findIndex((p) => p?.id === id);
+ if (i < 0) throw new Error(`no post with id ${id}`);
+ const [removed] = m.posts.splice(i, 1);
+ if (!m.posts.length) delete m.posts;
+ return { removed };
+ });
+}
+
+/**
+ * Set or remove `render.chrome.factcheck`: the stamp, the tally, and each
+ * verdict's label and colour. Only on a manifest whose deck is on (the
+ * fact-check is drawn by the deck); null removes it. Checked by the build's
+ * validateChrome against the rest of the render block.
+ *
+ * @param {string} dir
+ * @param {Record<string, unknown> | null} factcheck
+ * @param {{ token?: string | null }} [opts]
+ */
+export async function updateFactcheck(dir, factcheck, { token = null } = {}) {
+ if (factcheck === undefined) throw new Error("factcheck must be an object, or null to remove it");
+ return structural(dir, "factcheck", token, (m) => {
+ const render = m.render ?? {};
+ if (!render.chrome || typeof render.chrome !== "object") {
+ throw new Error("the fact-check is drawn by the on-screen deck, and this manifest has none — turn the deck on first");
+ }
+ const chrome = { ...render.chrome };
+ if (factcheck === null) delete chrome.factcheck;
+ else chrome.factcheck = factcheck;
+ const { chrome: _old, ...rest } = render;
+ const already = new Set(validateChrome(render.chrome, rest));
+ const errors = validateChrome(chrome, rest).filter((e) => !already.has(e));
+ if (errors.length) throw new StructureRefused(errors);
+ m.render = { ...render, chrome };
+ return { factcheck: chrome.factcheck ?? null };
+ });
+}
+
+/**
+ * Take back the newest structural edit: restore the newest `auto-before-<op>`
+ * snapshot, byte for byte. The manifest as it is now is kept first
+ * (`undo-saved`), and the restored snapshot is renamed `undone-<op>` so the
+ * next undo goes one step further back rather than round in a circle.
+ *
+ * @param {string} dir
+ * @param {{ token?: string | null }} [opts]
+ */
+export async function undoStructural(dir, { token = null } = {}) {
+ return withManifestLock(async () => {
+ const current = await manifestToken(dir);
+ if (token !== null && current !== token) throw new StaleToken(token, current);
+ const target = (await listSnapshots(dir)).find((s) => !s.legacy && s.label?.startsWith(AUTO_SNAPSHOT_PREFIX));
+ if (!target) throw new Error("nothing to undo: no automatic snapshot in revisions/");
+ const op = target.label.slice(AUTO_SNAPSHOT_PREFIX.length);
+ const text = await readFile(path.join(dir, target.rel), "utf8");
+ JSON.parse(text); // a snapshot that does not parse is not restored over a manifest that does
+ let saved = null;
+ try {
+ saved = (await createSnapshot(dir, { label: "undo-saved" })).rel;
+ } catch {
+ // within the same second as another snapshot: the state is already kept
+ }
+ const file = manifestFile(dir);
+ const tmp = `${file}.tmp-${process.pid}-${Math.random().toString(36).slice(2, 8)}`;
+ await writeFile(tmp, text, "utf8");
+ await rename(tmp, file);
+ const undone = target.rel.replace(`-${target.label}.manifest.json`, `-undone-${op}.manifest.json`);
+ await rename(path.join(dir, target.rel), path.join(dir, undone)).catch(() => {});
+ return { restored: target.rel, op, saved, token: await manifestToken(dir) };
+ });
+}
diff --git a/umtool/lib/report/moments.mjs b/umtool/lib/report/moments.mjs
@@ -0,0 +1,155 @@
+// A moment on a rendered cut -> the entry playing there.
+//
+// A timed note is a second on a FILE (`out/sourced/<slug>.mp4`, a take's
+// `takes/<id>/preview.mp4`). What makes it worth an agent's time is what that
+// second resolves to: the entry on screen, its onscreen title and quote, and
+// the source second with a link into the archive. That join is the build's
+// `schedule.json` (build-video.mjs writeChromeSchedule: `out/<variant>/`, or a
+// take's own `takes/<id>/out/<variant>/`), which says where each entry starts
+// in the cut.
+//
+// THE SCHEDULE MATCHES THE BUILD'S OUTPUT, and a take's preview.mp4 is that
+// output copied: measured on every take of candace/polemic-israel, the
+// preview's duration equals the schedule's `total` to the millisecond. So a
+// mark is exact when the file it was made on is as long as the schedule says,
+// and APPROXIMATE (`approx: true`) when it is not -- a different preset, a
+// trimmed preview -- or when the schedule is an estimate, or the second falls
+// in a held frame past the clip's own source.
+//
+// The resolution is written INTO the note at write time (`anchor.resolved`):
+// a later rebuild moves entries around, and the agent reading the note must
+// see what was on screen when the operator pressed the key.
+import { readdir, readFile, stat } from "node:fs/promises";
+import path from "node:path";
+import { DEFAULT_VARIANT } from "umtool-report-to-video/build-video";
+import { channelFor } from "../projects/report.mjs";
+
+const TAKE_RE = /^[a-z0-9][a-z0-9-]{0,63}$/;
+const APPROX_TOLERANCE = 0.25;
+
+/**
+ * Where a moment's file sits: in a take (`takes/<id>/…`) and/or a variant's
+ * output dir (`…out/<variant>/…`). Pure.
+ *
+ * @param {string} rel project-relative, `/`-separated
+ */
+export function momentFileInfo(rel) {
+ const parts = String(rel ?? "").split("/");
+ let take = null;
+ let i = 0;
+ if (parts[0] === "takes" && TAKE_RE.test(parts[1] ?? "")) {
+ take = parts[1];
+ i = 2;
+ }
+ const variant = parts[i] === "out" && parts.length > i + 2 && TAKE_RE.test(parts[i + 1]) ? parts[i + 1] : null;
+ return { take, variant };
+}
+
+async function readJson(file) {
+ try {
+ return JSON.parse(await readFile(/* turbopackIgnore: true */ file, "utf8"));
+ } catch {
+ return null;
+ }
+}
+
+/**
+ * The schedule and manifest a moment on `rel` resolves against, or null when
+ * there is no schedule. A take's own manifest wins over the project's.
+ *
+ * @param {string} projectDir
+ * @param {string} rel
+ */
+export async function scheduleForFile(projectDir, rel) {
+ const { take, variant } = momentFileInfo(rel);
+ const base = take ? path.join(/* turbopackIgnore: true */ projectDir, "takes", take) : projectDir;
+ let variants = [];
+ if (variant) variants = [variant];
+ else {
+ const dirs = await readdir(/* turbopackIgnore: true */ path.join(/* turbopackIgnore: true */ base, "out"), { withFileTypes: true }).catch(() => []);
+ variants = dirs.filter((d) => d.isDirectory() || d.isSymbolicLink()).map((d) => d.name);
+ // A deliverable names its cut (`<slug>-full.mp4`; the default cut is
+ // `<slug>.mp4`): that variant's schedule first, then the default's.
+ const named = (v) => path.basename(String(rel)).endsWith(`-${v}.mp4`);
+ const rank = (v) => (named(v) ? 0 : v === DEFAULT_VARIANT ? 1 : 2);
+ variants.sort((a, b) => rank(a) - rank(b) || a.localeCompare(b));
+ }
+ for (const v of variants) {
+ const file = path.join(/* turbopackIgnore: true */ base, "out", v, "schedule.json");
+ const schedule = await readJson(file);
+ if (!schedule || !Array.isArray(schedule.segments)) continue;
+ const manifest =
+ (take ? await readJson(path.join(/* turbopackIgnore: true */ base, "video.manifest.json")) : null) ??
+ (await readJson(path.join(/* turbopackIgnore: true */ projectDir, "video.manifest.json")));
+ const st = await stat(/* turbopackIgnore: true */ file).catch(() => null);
+ return {
+ schedule,
+ manifest,
+ variant: v,
+ take,
+ scheduleRel: path.relative(/* turbopackIgnore: true */ projectDir, file).split(path.sep).join("/"),
+ scheduleMtimeMs: st ? Math.round(st.mtimeMs) : null,
+ };
+ }
+ return null;
+}
+
+/**
+ * The entry playing at `t` seconds of a cut. Pure.
+ *
+ * During a crossfade both segments are on screen; the incoming one is taken
+ * from half-way through it. `duration` is the file's own (the player knows it):
+ * when it differs from the schedule's total the result is `approx`.
+ *
+ * @param {{ schedule: any, manifest: any, variant?: string | null, t: number, duration?: number | null }} args
+ * @returns {{ entry: string | null, title?: string, quote?: string, channel?: string, video?: string, sourceT?: number, url?: string, approx?: boolean }}
+ */
+export function resolveMoment({ schedule, manifest, variant = null, t, duration = null }) {
+ const segs = Array.isArray(schedule?.segments) ? schedule.segments : [];
+ if (!segs.length || !Number.isFinite(t)) return { entry: null };
+ const D = Number(schedule.transition) || 0;
+ let i = 0;
+ for (let k = 0; k < segs.length; k += 1) if (Number(segs[k].start) <= t - D / 2) i = k;
+ const seg = segs[i];
+ const out = { entry: String(seg.id) };
+ let approx = schedule.estimated === true;
+ if (Number.isFinite(duration) && Number.isFinite(Number(schedule.total)) && Math.abs(duration - Number(schedule.total)) > APPROX_TOLERANCE) {
+ approx = true;
+ }
+ const e = (manifest?.timeline ?? []).find((x) => x?.id === seg.id && (!x.variant || !variant || x.variant === variant)) ?? null;
+ const title = seg.title ?? e?.onscreen?.title ?? e?.title ?? e?.heading ?? null;
+ if (title) out.title = String(title);
+ if (e?.quote) out.quote = String(e.quote);
+ if (e?.type === "clip") {
+ const from = Number(e.cutStart ?? e.start);
+ const to = Number(e.cutEnd ?? e.end);
+ const into = Math.max(0, t - Number(seg.start));
+ if (Number.isFinite(from) && Number.isFinite(to)) {
+ if (from + into > to + 0.05) approx = true; // a held frame past the clip's own source
+ out.sourceT = Number(Math.min(to, from + into).toFixed(2));
+ const channel = channelFor(manifest, e);
+ if (channel) out.channel = channel;
+ out.video = String(e.video);
+ const origin = manifest?.provenance?.siteOrigin;
+ if (origin && channel) {
+ out.url = `${origin}/?v=${encodeURIComponent(`${channel}/${e.video}`)}&t=${Math.floor(out.sourceT)}`;
+ }
+ }
+ }
+ if (approx) out.approx = true;
+ return out;
+}
+
+/**
+ * What a moment on `rel` at `t` resolves to, or `{ entry: null }` with no schedule.
+ *
+ * @param {string} projectDir
+ * @param {string} rel
+ * @param {number} t
+ * @param {number | null} [duration]
+ */
+export async function resolveMomentOnFile(projectDir, rel, t, duration = null) {
+ const s = await scheduleForFile(projectDir, rel);
+ if (!s) return { entry: null, schedule: null };
+ return { ...resolveMoment({ ...s, t, duration }), schedule: s.scheduleRel };
+}
diff --git a/umtool/lib/report/moments.test.mjs b/umtool/lib/report/moments.test.mjs
@@ -0,0 +1,81 @@
+// A second on a rendered cut -> the entry, title, quote and source second.
+//
+// Run with: pnpm test:scripts
+import assert from "node:assert/strict";
+import { mkdir, mkdtemp, rm, writeFile } from "node:fs/promises";
+import { tmpdir } from "node:os";
+import path from "node:path";
+import test from "node:test";
+import { momentFileInfo, resolveMoment, resolveMomentOnFile } from "./moments.mjs";
+
+const manifest = {
+ provenance: { siteOrigin: "https://arch.test", channelSlug: "chan" },
+ timeline: [
+ { type: "teaser", id: "tz", lines: ["X"] },
+ { type: "clip", id: "c1", video: "v1", start: 100, end: 110, cutStart: 102, cutEnd: 108, quote: "the words", onscreen: { title: "Own title" } },
+ { type: "clip", id: "c2", video: "v2", channel: "other", start: 50, end: 60 },
+ ],
+};
+const schedule = {
+ transition: 0.5,
+ total: 20.5,
+ segments: [
+ { id: "tz", start: 0, end: 4.5, title: "Teaser" },
+ { id: "c1", start: 4, end: 10.5 },
+ { id: "c2", start: 10, end: 20.5, title: "Schedule title" },
+ ],
+};
+
+test("momentFileInfo reads the take and the variant from the path", () => {
+ assert.deepEqual(momentFileInfo("takes/deck/preview.mp4"), { take: "deck", variant: null });
+ assert.deepEqual(momentFileInfo("takes/deck/out/full/x.mp4"), { take: "deck", variant: "full" });
+ assert.deepEqual(momentFileInfo("out/sourced/slug.mp4"), { take: null, variant: "sourced" });
+ assert.deepEqual(momentFileInfo("video.mp4"), { take: null, variant: null });
+});
+
+test("resolveMoment: the entry on screen, the incoming one past half a crossfade", () => {
+ assert.equal(resolveMoment({ schedule, manifest, t: 1 }).entry, "tz");
+ assert.equal(resolveMoment({ schedule, manifest, t: 4.1 }).entry, "tz", "still the outgoing one");
+ const c1 = resolveMoment({ schedule, manifest, t: 6, duration: 20.5 });
+ assert.deepEqual(c1, {
+ entry: "c1",
+ title: "Own title",
+ quote: "the words",
+ sourceT: 104,
+ channel: "chan",
+ video: "v1",
+ url: "https://arch.test/?v=chan%2Fv1&t=104",
+ });
+ const c2 = resolveMoment({ schedule, manifest, t: 12 });
+ assert.equal(c2.title, "Schedule title");
+ assert.equal(c2.channel, "other");
+ assert.equal(c2.sourceT, 52);
+});
+
+test("approx: a file of another length, an estimated schedule, a held frame", () => {
+ assert.equal(resolveMoment({ schedule, manifest, t: 6, duration: 30 }).approx, true);
+ assert.equal(resolveMoment({ schedule: { ...schedule, estimated: true }, manifest, t: 6 }).approx, true);
+ // c1's cut is 6 s; 7.5 s in is a held frame
+ const held = resolveMoment({ schedule, manifest, t: 4 + 7.5 - 0.01 + 0, duration: 20.5 });
+ assert.equal(held.entry, "c2");
+ const c1held = resolveMoment({ schedule: { ...schedule, segments: [schedule.segments[0], { id: "c1", start: 4 }] }, manifest, t: 12, duration: 20.5 });
+ assert.equal(c1held.sourceT, 108);
+ assert.equal(c1held.approx, true);
+});
+
+test("resolveMomentOnFile finds a take's schedule, preferring the default variant", async () => {
+ const dir = await mkdtemp(path.join(tmpdir(), "umtool-moments-"));
+ try {
+ await writeFile(path.join(dir, "video.manifest.json"), JSON.stringify(manifest));
+ await mkdir(path.join(dir, "takes", "t1", "out", "full"), { recursive: true });
+ await mkdir(path.join(dir, "takes", "t1", "out", "sourced"), { recursive: true });
+ await writeFile(path.join(dir, "takes", "t1", "out", "sourced", "schedule.json"), JSON.stringify(schedule));
+ await writeFile(path.join(dir, "takes", "t1", "out", "full", "schedule.json"), JSON.stringify({ ...schedule, segments: [{ id: "c2", start: 0 }] }));
+ const r = await resolveMomentOnFile(dir, "takes/t1/preview.mp4", 6, 20.5);
+ assert.equal(r.entry, "c1");
+ assert.equal(r.schedule, "takes/t1/out/sourced/schedule.json");
+ assert.deepEqual(await resolveMomentOnFile(dir, "takes/none/preview.mp4", 6), { entry: null, schedule: null });
+ } finally {
+ await rm(dir, { recursive: true });
+ }
+});
diff --git a/umtool/lib/report/sections.mjs b/umtool/lib/report/sections.mjs
@@ -0,0 +1,64 @@
+// `sectionEnter`: the flag on the first clip of each section, which is what
+// makes the legacy footer's marker slide from one node to the next
+// (report-to-video/build-video.mjs reads it; README "The marker slides").
+//
+// report-to-video only ever READS it; the generators author it. This writes
+// the rule down for umtool's one writer that re-orders a timeline
+// (lib/report/manifest.mjs moveEntry), derived from what the builder reads and
+// checked against every real manifest: in the one that carries the flag
+// (quartering-employee-count, seven of nineteen clips) a clip carries it
+// exactly when its `section` differs from the previous CLIP's -- the first
+// clip counts, a card between two clips does not break a section, and a clip
+// with no `section` never enters one. Re-applying it to every real manifest
+// under ~/reports changes nothing (lib/report/sections.test.mjs has the shape).
+//
+// `section` itself is the author's: it says which chapter an entry belongs
+// to, and a move does not change that. Only the flag follows the order.
+//
+// A manifest that carries no `sectionEnter` key at all (every deck-era cut:
+// the deck draws no footer) is left exactly as it is: `applySectionEnter`
+// changes nothing unless some entry already carries the key.
+
+/**
+ * The flag each entry should carry, by index: true on a clip whose `section`
+ * is set and differs from the previous clip's; false everywhere else.
+ *
+ * @param {Array<Record<string, unknown>>} timeline
+ * @returns {boolean[]}
+ */
+export function sectionEnterFlags(timeline) {
+ let prev;
+ return (timeline ?? []).map((e) => {
+ if (e?.type !== "clip") return false;
+ const s = e.section;
+ const enters = s !== undefined && s !== null && s !== prev;
+ prev = s;
+ return enters;
+ });
+}
+
+/** Does this timeline use the flag at all? */
+export const usesSectionEnter = (timeline) => (timeline ?? []).some((e) => e && Object.hasOwn(e, "sectionEnter"));
+
+/**
+ * Recompute `sectionEnter` in place after a re-order. Only on a timeline that
+ * already uses it; written as `true` or removed (an absent key reads as false,
+ * and `"sectionEnter": false` is noise a human reads as a decision). Returns
+ * the ids whose flag changed.
+ *
+ * @param {Array<Record<string, unknown>>} timeline
+ * @returns {string[]}
+ */
+export function applySectionEnter(timeline) {
+ if (!usesSectionEnter(timeline)) return [];
+ const flags = sectionEnterFlags(timeline);
+ const changed = [];
+ timeline.forEach((e, i) => {
+ const was = e.sectionEnter === true;
+ if (flags[i] === was) return;
+ if (flags[i]) e.sectionEnter = true;
+ else delete e.sectionEnter;
+ changed.push(String(e.id));
+ });
+ return changed;
+}
diff --git a/umtool/lib/report/sections.test.mjs b/umtool/lib/report/sections.test.mjs
@@ -0,0 +1,42 @@
+// The sectionEnter rule, against the shape of the one manifest that uses it.
+import assert from "node:assert/strict";
+import test from "node:test";
+import { applySectionEnter, sectionEnterFlags, usesSectionEnter } from "./sections.mjs";
+
+const clip = (id, section, extra = {}) => ({ type: "clip", id, section, ...extra });
+
+test("a clip enters a section when its section differs from the previous clip's", () => {
+ const tl = [
+ { type: "card", id: "t0" },
+ clip("a", 1),
+ clip("b", 1),
+ { type: "card", id: "mid" },
+ clip("c", 1),
+ clip("d", 2),
+ clip("e"),
+ clip("f", 2),
+ { type: "scroll", id: "s" },
+ ];
+ assert.deepEqual(sectionEnterFlags(tl), [false, true, false, false, false, true, false, true, false]);
+});
+
+test("applySectionEnter leaves a timeline that never used the flag alone", () => {
+ const tl = [clip("a", 1), clip("b", 2)];
+ assert.deepEqual(applySectionEnter(tl), []);
+ assert.equal(usesSectionEnter(tl), false);
+ assert.equal("sectionEnter" in tl[0], false);
+});
+
+test("applySectionEnter is the identity on a timeline already in order, and fixes a move", () => {
+ const tl = [clip("a", 1, { sectionEnter: true }), clip("b", 1), clip("c", 2, { sectionEnter: true }), clip("d", 2)];
+ assert.deepEqual(applySectionEnter(tl), []);
+ // d moved to the front: d enters 2, a enters 1, c no longer enters (b was 1, c is 2 -> still enters)
+ const moved = [tl[3], tl[0], tl[1], tl[2]];
+ assert.deepEqual(applySectionEnter(moved).sort(), ["d"]);
+ assert.equal(moved[0].sectionEnter, true);
+ assert.equal(moved[3].sectionEnter, true);
+ // and moving it back removes the flag again (deleted, never written false)
+ const back = [moved[1], moved[2], moved[3], moved[0]];
+ assert.deepEqual(applySectionEnter(back), ["d"]);
+ assert.equal("sectionEnter" in back[3], false);
+});
diff --git a/umtool/lib/report/structure.test.mjs b/umtool/lib/report/structure.test.mjs
@@ -0,0 +1,235 @@
+// The structural writers: move, remove, duplicate, insert, the teaser, the
+// posts, the fact-check, and undo. Each goes through the token, the lock, the
+// build's own checks and an automatic snapshot, and each refusal leaves the
+// file byte-for-byte as it was.
+//
+// Run with: pnpm test:scripts
+import assert from "node:assert/strict";
+import { mkdtemp, readFile, readdir, rm, utimes, writeFile } from "node:fs/promises";
+import { tmpdir } from "node:os";
+import path from "node:path";
+import test from "node:test";
+
+import {
+ MANIFEST_NAME,
+ StaleToken,
+ StructureRefused,
+ duplicateEntry,
+ entryFromSpec,
+ insertEntry,
+ manifestToken,
+ moveEntry,
+ removeEntry,
+ removePost,
+ undoStructural,
+ updateFactcheck,
+ updateTeaser,
+ upsertPost,
+} from "./manifest.mjs";
+
+const base = () => ({
+ slug: "t",
+ provenance: { siteOrigin: "https://example.test", channelSlug: "chan" },
+ render: { width: 1920, height: 1080, fps: 30, transition: 0.5 },
+ timeline: [
+ { type: "teaser", id: "tz", lines: ["THE PROMISE"] },
+ { type: "clip", id: "c01", video: "v1", start: 10, end: 20, section: 1, sectionEnter: true },
+ { type: "clip", id: "c02", video: "v2", start: 30, end: 41, section: 1 },
+ { type: "clip", id: "c03", video: "v3", start: 50, end: 55, section: 2, sectionEnter: true },
+ ],
+ posts: [
+ { id: "p1", platform: "x", date: "2024-01-02", text: "hello", url: "https://x.com/a/status/1", attachTo: "c02" },
+ ],
+});
+
+async function project(manifest = base()) {
+ const dir = await mkdtemp(path.join(tmpdir(), "umtool-structure-"));
+ await writeFile(path.join(dir, MANIFEST_NAME), JSON.stringify(manifest, null, 2) + "\n");
+ return dir;
+}
+const readRaw = (dir) => readFile(path.join(dir, MANIFEST_NAME), "utf8");
+const read = async (dir) => JSON.parse(await readRaw(dir));
+const ids = (m) => m.timeline.map((e) => e.id);
+const revisions = async (dir) => (await readdir(path.join(dir, "revisions")).catch(() => [])).sort();
+
+test("moveEntry: re-orders, recomputes sectionEnter, snapshots once per burst", async () => {
+ const dir = await project();
+ try {
+ const r = await moveEntry(dir, "c03", 1, { token: await manifestToken(dir) });
+ assert.deepEqual([r.from, r.to], [3, 1]);
+ const m = await read(dir);
+ assert.deepEqual(ids(m), ["tz", "c03", "c01", "c02"]);
+ // c03 (section 2) now first: it enters; c01 (section 1, after a 2) enters; c02 does not
+ assert.equal(m.timeline[1].sectionEnter, true);
+ assert.equal(m.timeline[2].sectionEnter, true);
+ assert.equal("sectionEnter" in m.timeline[3], false);
+ // The file is still the CLI's formatting.
+ assert.equal(await readRaw(dir), JSON.stringify(m, null, 2) + "\n");
+ assert.equal((await revisions(dir)).filter((n) => n.includes("auto-before-move")).length, 1);
+ await moveEntry(dir, "c03", 3);
+ assert.equal((await revisions(dir)).filter((n) => n.includes("auto-before-move")).length, 1, "throttled");
+ await assert.rejects(moveEntry(dir, "c03", 3), /already at 3/);
+ await assert.rejects(moveEntry(dir, "c03", 9), /toIndex/);
+ await assert.rejects(moveEntry(dir, "nope", 0), /no timeline entry/);
+ } finally {
+ await rm(dir, { recursive: true });
+ }
+});
+
+test("a stale token refuses every structural write and leaves the file alone", async () => {
+ const dir = await project();
+ try {
+ const before = await readRaw(dir);
+ for (const call of [
+ () => moveEntry(dir, "c03", 1, { token: "1" }),
+ () => removeEntry(dir, "c03", { token: "1" }),
+ () => duplicateEntry(dir, "c03", { token: "1" }),
+ () => insertEntry(dir, "c03", "chan/vx@1-5", { token: "1" }),
+ () => updateTeaser(dir, "tz", { beat: 1 }, { token: "1" }),
+ () => upsertPost(dir, { id: "p2" }, { token: "1" }),
+ () => removePost(dir, "p1", { token: "1" }),
+ () => undoStructural(dir, { token: "1" }),
+ ]) {
+ await assert.rejects(call(), StaleToken);
+ }
+ assert.equal(await readRaw(dir), before);
+ } finally {
+ await rm(dir, { recursive: true });
+ }
+});
+
+test("removeEntry: refused while a post rides on the clip; fine once it does not", async () => {
+ const dir = await project();
+ try {
+ const before = await readRaw(dir);
+ await assert.rejects(removeEntry(dir, "c02"), (e) => e instanceof StructureRefused && /attachTo/.test(e.message));
+ assert.equal(await readRaw(dir), before);
+ const r = await removeEntry(dir, "c03");
+ assert.equal(r.removed.id, "c03");
+ assert.deepEqual(ids(await read(dir)), ["tz", "c01", "c02"]);
+ } finally {
+ await rm(dir, { recursive: true });
+ }
+});
+
+test("duplicateEntry and insertEntry: fresh ids, the clip spec, 2 dp, at the right place", async () => {
+ const dir = await project();
+ try {
+ const d = await duplicateEntry(dir, "c01");
+ assert.equal(d.id, "c01-copy");
+ const again = await duplicateEntry(dir, "c01");
+ assert.equal(again.id, "c01-copy-2");
+ const ins = await insertEntry(dir, "c02", "mychan/abc123@12.3456-20.1");
+ assert.deepEqual(ins.entry, { type: "clip", id: "abc123-12", channel: "mychan", video: "abc123", start: 12.35, end: 20.1 });
+ const first = await insertEntry(dir, null, { type: "card", heading: "Start" });
+ assert.equal(first.at, 0);
+ assert.equal(first.id, "card");
+ const m = await read(dir);
+ assert.deepEqual(ids(m), ["card", "tz", "c01", "c01-copy-2", "c01-copy", "c02", "abc123-12", "c03"]);
+ await assert.rejects(insertEntry(dir, "c02", "not a spec"), /channel/);
+ await assert.rejects(insertEntry(dir, "c02", "chan/v@5-5.2"), /half a second/);
+ // A teaser that the build would refuse is refused here.
+ await assert.rejects(insertEntry(dir, "c02", { type: "teaser", lines: [] }), StructureRefused);
+ } finally {
+ await rm(dir, { recursive: true });
+ }
+});
+
+test("entryFromSpec never reuses an id", () => {
+ const tl = [{ id: "a" }, { id: "v-1" }];
+ assert.equal(entryFromSpec(tl, { type: "card", id: "a" }).id, "a-2");
+ assert.equal(entryFromSpec(tl, "c/v@1-3").id, "v-1-2");
+});
+
+test("updateTeaser: lines, beat, tail, dip — validated by the build's own check", async () => {
+ const dir = await project();
+ try {
+ const r = await updateTeaser(dir, "tz", { lines: ["THE PROMISE", "AND WHAT HAPPENED"], beat: 1.2345, tail: "?", tailWait: 1 });
+ assert.equal(r.entry.beat, 1.23);
+ await assert.rejects(updateTeaser(dir, "tz", { lines: [] }), StructureRefused);
+ await assert.rejects(updateTeaser(dir, "tz", { tail: null }), /tailWait/);
+ await updateTeaser(dir, "tz", { tail: null, tailWait: null });
+ const m = await read(dir);
+ assert.equal("tail" in m.timeline[0], false);
+ await assert.rejects(updateTeaser(dir, "c01", { beat: 1 }), /not a teaser/);
+ await assert.rejects(updateTeaser(dir, "tz", { seconds: 4 }), /not something/);
+ } finally {
+ await rm(dir, { recursive: true });
+ }
+});
+
+test("upsertPost / removePost: add, replace keeping attachTo, refuse what validatePosts refuses", async () => {
+ const dir = await project();
+ try {
+ const add = await upsertPost(dir, { id: "p2", platform: "bluesky", date: "2024-03-04", text: " words ", url: "https://bsky.app/x", author: "" });
+ assert.equal(add.created, true);
+ assert.deepEqual(add.post, { id: "p2", platform: "bluesky", date: "2024-03-04", text: "words", url: "https://bsky.app/x" });
+ const rep = await upsertPost(dir, { id: "p1", platform: "x", date: "2024-01-02", text: "edited", url: "https://x.com/a/status/1" });
+ assert.equal(rep.post.attachTo, "c02");
+ await assert.rejects(upsertPost(dir, { id: "p3", platform: "myspace", date: "2024", text: "x", url: "http://no" }), StructureRefused);
+ await removePost(dir, "p2");
+ await removePost(dir, "p1");
+ assert.equal("posts" in (await read(dir)), false);
+ await assert.rejects(removePost(dir, "p1"), /no post/);
+ } finally {
+ await rm(dir, { recursive: true });
+ }
+});
+
+test("updateFactcheck: needs the deck; labels and colours checked by validateChrome", async () => {
+ const dir = await project();
+ try {
+ await assert.rejects(updateFactcheck(dir, { stamp: { seconds: 3 } }), /deck on first/);
+ const m = await read(dir);
+ m.render.chrome = { engine: "hyperframes", layout: "deck" };
+ await writeFile(path.join(dir, MANIFEST_NAME), JSON.stringify(m, null, 2) + "\n");
+ const r = await updateFactcheck(dir, { verdicts: { CONTRADICTED: { label: "NOPE", color: "#ff0000" } }, stamp: { seconds: 4 } });
+ assert.equal(r.factcheck.stamp.seconds, 4);
+ await assert.rejects(updateFactcheck(dir, { stamp: { seconds: 99 } }), StructureRefused);
+ await updateFactcheck(dir, null);
+ assert.equal("factcheck" in (await read(dir)).render.chrome, false);
+ } finally {
+ await rm(dir, { recursive: true });
+ }
+});
+
+test("undoStructural restores the newest automatic snapshot, keeps the current state, and walks back", async () => {
+ const dir = await project();
+ try {
+ const original = await readRaw(dir);
+ await assert.rejects(undoStructural(dir), /nothing to undo/);
+ await moveEntry(dir, "c03", 1);
+ // Age the move's snapshot so the next op is not throttled into it.
+ const rev = path.join(dir, "revisions");
+ for (const n of await readdir(rev)) await utimes(path.join(rev, n), new Date(Date.now() - 600_000), new Date(Date.now() - 600_000));
+ await new Promise((r) => setTimeout(r, 1100));
+ await removeEntry(dir, "c01");
+ const afterRemove = ids(await read(dir));
+ assert.deepEqual(afterRemove, ["tz", "c03", "c02"]);
+
+ const u1 = await undoStructural(dir);
+ assert.equal(u1.op, "remove");
+ assert.deepEqual(ids(await read(dir)), ["tz", "c03", "c01", "c02"]);
+ await new Promise((r) => setTimeout(r, 1100));
+ const u2 = await undoStructural(dir);
+ assert.equal(u2.op, "move");
+ assert.equal(await readRaw(dir), original, "byte for byte");
+ const names = await revisions(dir);
+ assert.ok(names.some((n) => n.includes("undone-move")));
+ assert.ok(names.some((n) => n.includes("undo-saved")));
+ } finally {
+ await rm(dir, { recursive: true });
+ }
+});
+
+test("a manifest that already has a build problem can still be re-ordered", async () => {
+ const m = base();
+ m.posts[0].attachTo = "gone"; // already broken before the edit
+ const dir = await project(m);
+ try {
+ await moveEntry(dir, "c03", 1);
+ assert.deepEqual(ids(await read(dir)), ["tz", "c03", "c01", "c02"]);
+ } finally {
+ await rm(dir, { recursive: true });
+ }
+});
diff --git a/umtool/lib/report/takes.mjs b/umtool/lib/report/takes.mjs
@@ -0,0 +1,322 @@
+// Takes: alternative renders of one part of a report video, side by side.
+//
+// Something ELSE renders them -- an agent trying five endings -- into
+//
+// <project>/takes/<takeId>/take.json
+// <project>/takes/<takeId>/preview.mp4 (whatever take.json names)
+// <project>/takes/<takeId>/video.manifest.json (what it was built from)
+//
+// and this reads them back, serves their previews and records which ones the
+// operator liked in <project>/takes/verdicts.json, which that agent reads.
+// So take.json is a contract written by somebody else: every field is checked,
+// and a take that fails is SKIPPED WITH A REASON rather than dropped or half
+// shown. A directory with no take.json is not a take (a work dir, a build in
+// progress) and is not listed at all.
+//
+// Plain ESM for the same reason manifest.mjs is: `node --test` runs it with no
+// TypeScript, and the rules about which file a request may open live here,
+// tested, rather than in a route.
+import { readdir, readFile, realpath, rename, stat, writeFile } from "node:fs/promises";
+import path from "node:path";
+import { inside } from "../paths.mjs";
+
+export const TAKES_DIR = "takes";
+export const TAKE_FILE = "take.json";
+export const VERDICTS_FILE = "verdicts.json";
+
+/** A take id is its directory's name, and one url-safe segment. */
+export const TAKE_ID = /^[a-z0-9][a-z0-9-]{0,63}$/;
+/**
+ * @param {unknown} v
+ * @returns {v is string}
+ */
+export const isTakeId = (v) => typeof v === "string" && TAKE_ID.test(v);
+
+export const TAKE_KINDS = ["reference", "similar", "different"];
+export const TAKE_VERDICTS = ["like", "maybe", "no"];
+
+/** A note is a sentence or two about one take, not an essay. */
+export const TAKE_NOTE_LIMIT = 2000;
+
+export const takesDirOf = (projectDir) => path.join(projectDir, TAKES_DIR);
+export const verdictsFileOf = (projectDir) => path.join(takesDirOf(projectDir), VERDICTS_FILE);
+
+const str = (v) => (typeof v === "string" ? v.trim() : "");
+
+/**
+ * One take.json, checked. `{ take }` or `{ error }` -- never a partial take.
+ *
+ * Required: id (== the directory), group, order, label, kind, preview.
+ * Optional: summary (""), changes ([]), seconds (null), builtAt (null). An
+ * optional field of the wrong type is an error too: a `changes` that is a
+ * string would otherwise render as one bullet per character.
+ *
+ * `preview` is RELATIVE to the take's directory and may not climb out of it;
+ * that is checked again against the real path when the file is served.
+ *
+ * @param {string} dirName
+ * @param {unknown} raw
+ */
+export function parseTake(dirName, raw) {
+ if (!isTakeId(dirName)) return { error: `directory name is not a take id (${TAKE_ID})` };
+ if (!raw || typeof raw !== "object" || Array.isArray(raw)) return { error: "take.json is not an object" };
+ const j = /** @type {Record<string, unknown>} */ (raw);
+ if (j.id !== dirName) return { error: `id ${JSON.stringify(j.id ?? null)} is not the directory name` };
+ const group = str(j.group);
+ if (!group) return { error: "group is missing" };
+ if (typeof j.order !== "number" || !Number.isFinite(j.order)) return { error: "order is not a number" };
+ const label = str(j.label);
+ if (!label) return { error: "label is missing" };
+ if (!TAKE_KINDS.includes(/** @type {string} */ (j.kind))) {
+ return { error: `kind must be one of ${TAKE_KINDS.join(", ")}` };
+ }
+ const preview = str(j.preview);
+ if (!preview) return { error: "preview is missing" };
+ if (path.isAbsolute(preview) || /[\\\0]/.test(preview) || preview.split("/").some((s) => s === ".." || s === "")) {
+ return { error: "preview must be a relative path inside the take" };
+ }
+ if (j.summary !== undefined && typeof j.summary !== "string") return { error: "summary is not a string" };
+ if (j.changes !== undefined && !(Array.isArray(j.changes) && j.changes.every((c) => typeof c === "string"))) {
+ return { error: "changes is not a list of strings" };
+ }
+ if (j.seconds !== undefined && j.seconds !== null && (typeof j.seconds !== "number" || !Number.isFinite(j.seconds))) {
+ return { error: "seconds is not a number" };
+ }
+ if (j.builtAt !== undefined && j.builtAt !== null && typeof j.builtAt !== "string") {
+ return { error: "builtAt is not a string" };
+ }
+ return {
+ take: {
+ id: dirName,
+ group,
+ order: j.order,
+ label,
+ kind: /** @type {"reference" | "similar" | "different"} */ (j.kind),
+ summary: str(j.summary),
+ changes: /** @type {string[]} */ (j.changes ?? []).map((c) => c.trim()).filter(Boolean),
+ preview,
+ seconds: typeof j.seconds === "number" ? j.seconds : null,
+ builtAt: typeof j.builtAt === "string" ? j.builtAt : null,
+ },
+ };
+}
+
+/**
+ * Every take of a project, grouped, plus what was skipped and why.
+ *
+ * Groups are in order of their lowest `order`, then by name; takes within a
+ * group by `order`, then id. Each take carries its preview's size and mtime
+ * (null when the file is not there yet -- an agent may still be rendering it),
+ * and the mtime is what the page passes as `v`, because a re-render writes the
+ * same path.
+ *
+ * @param {string} projectDir
+ */
+export async function listTakes(projectDir) {
+ const root = takesDirOf(projectDir);
+ const entries = await readdir(root, { withFileTypes: true }).catch(() => null);
+ if (!entries) return { exists: false, takes: [], groups: [], skipped: [] };
+
+ const takes = [];
+ const skipped = [];
+ for (const e of entries) {
+ if (e.name.startsWith(".")) continue;
+ const dir = path.join(root, e.name);
+ const isDir = e.isDirectory() || (e.isSymbolicLink() && (await stat(dir).then((s) => s.isDirectory(), () => false)));
+ if (!isDir) continue;
+ let text;
+ try {
+ text = await readFile(path.join(dir, TAKE_FILE), "utf8");
+ } catch {
+ continue; // not a take
+ }
+ let raw;
+ try {
+ raw = JSON.parse(text);
+ } catch (err) {
+ skipped.push({ dir: e.name, why: `take.json does not parse: ${err instanceof Error ? err.message : err}` });
+ continue;
+ }
+ const r = parseTake(e.name, raw);
+ if ("error" in r) {
+ skipped.push({ dir: e.name, why: r.error });
+ continue;
+ }
+ const file = await previewFile(projectDir, r.take);
+ takes.push({ ...r.take, previewSize: file?.size ?? null, previewMtimeMs: file?.mtimeMs ?? null });
+ }
+
+ takes.sort((a, b) => a.order - b.order || a.id.localeCompare(b.id));
+ const byGroup = new Map();
+ for (const t of takes) {
+ if (!byGroup.has(t.group)) byGroup.set(t.group, []);
+ byGroup.get(t.group).push(t);
+ }
+ const groups = [...byGroup.entries()]
+ .map(([group, list]) => ({ group, takes: list }))
+ .sort((a, b) => a.takes[0].order - b.takes[0].order || a.group.localeCompare(b.group));
+ skipped.sort((a, b) => a.dir.localeCompare(b.dir));
+ return { exists: true, takes, groups, skipped };
+}
+
+/**
+ * A take's preview, if it is a file whose REAL path is inside that take's own
+ * directory. A symlink out of it, or a take directory that is itself a link
+ * out of `takes/`, is refused, as deckPreviewFile refuses one.
+ *
+ * @param {string} projectDir
+ * @param {{ id: string, preview: string }} take an already-parsed take
+ * @returns {Promise<{ abs: string, size: number, mtimeMs: number } | null>}
+ */
+export async function previewFile(projectDir, take) {
+ if (!isTakeId(take?.id)) return null;
+ const root = takesDirOf(projectDir);
+ const dir = path.join(root, take.id);
+ const abs = path.resolve(dir, take.preview);
+ if (!inside(dir, abs) || abs === dir) return null;
+ const [realRoot, realDir, realAbs] = await Promise.all(
+ [root, dir, abs].map((p) => realpath(p).catch(() => null)),
+ );
+ if (!realRoot || !realDir || !realAbs) return null;
+ if (!inside(realRoot, realDir) || realDir === realRoot) return null;
+ if (!inside(realDir, realAbs) || realAbs === realDir) return null;
+ const st = await stat(realAbs).catch(() => null);
+ if (!st?.isFile()) return null;
+ return { abs: realAbs, size: st.size, mtimeMs: Math.round(st.mtimeMs) };
+}
+
+// ---------------------------------------------------------------------------
+// Verdicts.
+//
+// { "<takeId>": { "verdict": "like" | "maybe" | "no" | null, "note": string, "at": ISO } }
+//
+// Read back by the agent that made the takes, so the shape is fixed and every
+// row carries all three keys. A row whose verdict is null and note empty is
+// REMOVED rather than stored -- it says nothing.
+//
+// Its own write queue, as manifest.mjs has its own: lib/state.ts is TypeScript
+// (and serialises the song state files, an unrelated set). Within a process the
+// queue serialises read-modify-write; across processes the tmp + rename keeps
+// a reader from ever seeing half a file.
+// ---------------------------------------------------------------------------
+
+/** @type {Promise<unknown>} */
+let queue = Promise.resolve();
+/**
+ * @template T
+ * @param {() => Promise<T>} fn
+ * @returns {Promise<T>}
+ */
+function withTakesLock(fn) {
+ const run = queue.then(fn, fn);
+ queue = run.then(
+ () => undefined,
+ () => undefined,
+ );
+ return run;
+}
+
+/** Rows that are not the contract's shape are dropped on read, never repaired on disk. */
+function cleanVerdicts(raw) {
+ const out = {};
+ if (!raw || typeof raw !== "object" || Array.isArray(raw)) return out;
+ for (const [id, row] of Object.entries(raw)) {
+ if (!isTakeId(id) || !row || typeof row !== "object") continue;
+ const verdict = TAKE_VERDICTS.includes(row.verdict) ? row.verdict : null;
+ const note = typeof row.note === "string" ? row.note : "";
+ const at = typeof row.at === "string" ? row.at : "";
+ out[id] = { verdict, note, at };
+ }
+ return out;
+}
+
+/**
+ * @param {string} projectDir
+ * @returns {Promise<Record<string, { verdict: "like" | "maybe" | "no" | null, note: string, at: string }>>}
+ */
+export async function readVerdicts(projectDir) {
+ try {
+ return cleanVerdicts(JSON.parse(await readFile(verdictsFileOf(projectDir), "utf8")));
+ } catch {
+ return {};
+ }
+}
+
+/**
+ * Set one take's verdict and/or note. A field left undefined keeps its value;
+ * `verdict: null` clears it. Returns the row as stored (null when removed) and
+ * the whole file.
+ *
+ * A verdicts.json that exists and does not parse is an ERROR, not an empty
+ * map: writing over it would erase every judgement in it.
+ *
+ * @param {string} projectDir
+ * @param {string} takeId
+ * @param {{ verdict?: string | null, note?: string }} patch
+ */
+export async function setTakeVerdict(projectDir, takeId, patch) {
+ if (!isTakeId(takeId)) throw new Error("not a take id");
+ if (patch.verdict !== undefined && patch.verdict !== null && !TAKE_VERDICTS.includes(patch.verdict)) {
+ throw new Error(`verdict must be one of ${TAKE_VERDICTS.join(", ")} or null`);
+ }
+ if (patch.note !== undefined && typeof patch.note !== "string") throw new Error("note must be a string");
+ const file = verdictsFileOf(projectDir);
+ return withTakesLock(async () => {
+ let text = null;
+ try {
+ text = await readFile(file, "utf8");
+ } catch (err) {
+ if (/** @type {NodeJS.ErrnoException} */ (err).code !== "ENOENT") throw err;
+ }
+ let map = {};
+ if (text !== null) {
+ try {
+ map = cleanVerdicts(JSON.parse(text));
+ } catch {
+ throw new Error(`${VERDICTS_FILE} does not parse; not overwriting it`);
+ }
+ }
+ const prev = map[takeId] ?? { verdict: null, note: "", at: "" };
+ const row = {
+ verdict: patch.verdict === undefined ? prev.verdict : patch.verdict,
+ note: patch.note === undefined ? prev.note : patch.note.slice(0, TAKE_NOTE_LIMIT).replace(/\s+$/, ""),
+ at: new Date().toISOString(),
+ };
+ if (row.verdict === null && !row.note) delete map[takeId];
+ else map[takeId] = row;
+ const tmp = `${file}.tmp-${process.pid}-${Math.random().toString(36).slice(2, 8)}`;
+ await writeFile(tmp, JSON.stringify(map, null, 2) + "\n", "utf8");
+ await rename(tmp, file);
+ return { entry: map[takeId] ?? null, verdicts: map };
+ });
+}
+
+/**
+ * Every take with its verdict, for an agent reading the conversation back
+ * (`umtool notes`, the notes digest): label, group, summary and changes from
+ * take.json, and the operator's verdict and note from verdicts.json. A verdict
+ * on a take that no longer exists is kept, with `missing: true`, rather than
+ * dropped -- the agent may have deleted the take it was about.
+ *
+ * @param {string} projectDir
+ * @returns {Promise<Array<{ id: string, group: string | null, label: string | null, summary: string, changes: string[], verdict: "like" | "maybe" | "no" | null, note: string, at: string, missing?: true }>>}
+ */
+export async function takesWithVerdicts(projectDir) {
+ const [listing, verdicts] = await Promise.all([listTakes(projectDir), readVerdicts(projectDir)]);
+ const out = listing.takes.map((t) => ({
+ id: t.id,
+ group: t.group,
+ label: t.label,
+ summary: t.summary,
+ changes: t.changes,
+ verdict: verdicts[t.id]?.verdict ?? null,
+ note: verdicts[t.id]?.note ?? "",
+ at: verdicts[t.id]?.at ?? "",
+ }));
+ const known = new Set(out.map((t) => t.id));
+ for (const [id, v] of Object.entries(verdicts)) {
+ if (known.has(id)) continue;
+ out.push({ id, group: null, label: null, summary: "", changes: [], verdict: v.verdict, note: v.note, at: v.at, missing: true });
+ }
+ return out;
+}
diff --git a/umtool/lib/report/takes.test.mjs b/umtool/lib/report/takes.test.mjs
@@ -0,0 +1,262 @@
+// Takes: what take.json may say, which preview a request may open, and the
+// verdicts file another agent reads back.
+//
+// Run with: pnpm test:scripts
+import assert from "node:assert/strict";
+import { mkdir, mkdtemp, readFile, rm, symlink, writeFile } from "node:fs/promises";
+import { tmpdir } from "node:os";
+import path from "node:path";
+import test from "node:test";
+
+import {
+ isTakeId,
+ listTakes,
+ parseTake,
+ previewFile,
+ readVerdicts,
+ setTakeVerdict,
+ verdictsFileOf,
+} from "./takes.mjs";
+
+const good = (over = {}) => ({
+ id: "slow-burn",
+ group: "finale",
+ order: 3,
+ label: "Slow burn",
+ kind: "similar",
+ summary: "Same words; slower beat",
+ changes: ["beat 1.35 s (was 1.05)"],
+ preview: "preview.mp4",
+ seconds: 38.2,
+ builtAt: "2026-10-06T23:10:00Z",
+ ...over,
+});
+
+async function project() {
+ const dir = await mkdtemp(path.join(tmpdir(), "takes-"));
+ return dir;
+}
+
+async function writeTake(projectDir, id, doc, { preview = true } = {}) {
+ const dir = path.join(projectDir, "takes", id);
+ await mkdir(dir, { recursive: true });
+ await writeFile(path.join(dir, "take.json"), typeof doc === "string" ? doc : JSON.stringify(doc));
+ if (preview) await writeFile(path.join(dir, "preview.mp4"), "not really an mp4");
+ return dir;
+}
+
+test("take ids are one lowercase url segment", () => {
+ for (const ok of ["a", "slow-burn", "take-01", "0", "a".repeat(64)]) assert.equal(isTakeId(ok), true, ok);
+ for (const bad of ["", "-a", "Slow", "a_b", "a.b", "a/b", "..", "a".repeat(65), null, 3]) {
+ assert.equal(isTakeId(bad), false, String(bad));
+ }
+});
+
+test("parseTake accepts the contract and fills the optional fields", () => {
+ const r = parseTake("slow-burn", good());
+ assert.ok("take" in r);
+ assert.deepEqual(r.take, {
+ id: "slow-burn",
+ group: "finale",
+ order: 3,
+ label: "Slow burn",
+ kind: "similar",
+ summary: "Same words; slower beat",
+ changes: ["beat 1.35 s (was 1.05)"],
+ preview: "preview.mp4",
+ seconds: 38.2,
+ builtAt: "2026-10-06T23:10:00Z",
+ });
+ const min = parseTake("x", { id: "x", group: "g", order: 0, label: "X", kind: "reference", preview: "p.mp4" });
+ assert.ok("take" in min);
+ assert.equal(min.take.summary, "");
+ assert.deepEqual(min.take.changes, []);
+ assert.equal(min.take.seconds, null);
+ assert.equal(min.take.builtAt, null);
+});
+
+test("parseTake refuses a take it cannot show truthfully", () => {
+ const cases = [
+ ["slow-burn", [], "not an object"],
+ ["slow-burn", good({ id: "other" }), "not the directory"],
+ ["Bad_Dir", good({ id: "Bad_Dir" }), "directory name"],
+ ["slow-burn", good({ group: " " }), "group"],
+ ["slow-burn", good({ order: "3" }), "order"],
+ ["slow-burn", good({ order: Number.NaN }), "order"],
+ ["slow-burn", good({ label: undefined }), "label"],
+ ["slow-burn", good({ kind: "other" }), "kind"],
+ ["slow-burn", good({ preview: "" }), "preview"],
+ ["slow-burn", good({ preview: "/etc/passwd" }), "relative"],
+ ["slow-burn", good({ preview: "../other/preview.mp4" }), "relative"],
+ ["slow-burn", good({ preview: "a//b.mp4" }), "relative"],
+ ["slow-burn", good({ changes: "beat 1.35" }), "changes"],
+ ["slow-burn", good({ changes: [1] }), "changes"],
+ ["slow-burn", good({ summary: 4 }), "summary"],
+ ["slow-burn", good({ seconds: "38" }), "seconds"],
+ ["slow-burn", good({ builtAt: 1 }), "builtAt"],
+ ];
+ for (const [dir, doc, want] of cases) {
+ const r = parseTake(dir, doc);
+ assert.ok("error" in r, `${JSON.stringify(doc)} should fail`);
+ assert.match(r.error, new RegExp(want), r.error);
+ }
+});
+
+test("listTakes: no takes/ is a state of its own, not an error", async () => {
+ const dir = await project();
+ try {
+ assert.deepEqual(await listTakes(dir), { exists: false, takes: [], groups: [], skipped: [] });
+ } finally {
+ await rm(dir, { recursive: true, force: true });
+ }
+});
+
+test("listTakes groups, sorts, skips with a reason, and ignores what is not a take", async () => {
+ const dir = await project();
+ try {
+ await writeTake(dir, "slow-burn", good());
+ await writeTake(dir, "ref", good({ id: "ref", label: "Reference", kind: "reference", order: 1 }));
+ await writeTake(dir, "hard-cut", good({ id: "hard-cut", label: "Hard cut", kind: "different", order: 2 }), { preview: false });
+ await writeTake(dir, "intro-a", good({ id: "intro-a", group: "opening", order: 5 }));
+ await writeTake(dir, "broken", "{ not json");
+ await writeTake(dir, "wrong-kind", good({ id: "wrong-kind", kind: "maybe" }));
+ // A work directory and loose files: not takes, not listed, not "skipped".
+ await mkdir(path.join(dir, "takes", "current", "out"), { recursive: true });
+ await writeFile(path.join(dir, "takes", "make-takes.py"), "");
+ await writeFile(path.join(dir, "takes", "verdicts.json"), "{}");
+
+ const r = await listTakes(dir);
+ assert.equal(r.exists, true);
+ assert.deepEqual(r.groups.map((g) => [g.group, g.takes.map((t) => t.id)]), [
+ ["finale", ["ref", "hard-cut", "slow-burn"]],
+ ["opening", ["intro-a"]],
+ ]);
+ assert.deepEqual(r.skipped.map((s) => s.dir), ["broken", "wrong-kind"]);
+ assert.match(r.skipped[0].why, /does not parse/);
+ assert.match(r.skipped[1].why, /kind/);
+ const hard = r.takes.find((t) => t.id === "hard-cut");
+ assert.equal(hard.previewSize, null, "a preview not rendered yet is listed, without a size");
+ const ref = r.takes.find((t) => t.id === "ref");
+ assert.equal(ref.previewSize, "not really an mp4".length);
+ assert.equal(typeof ref.previewMtimeMs, "number");
+ } finally {
+ await rm(dir, { recursive: true, force: true });
+ }
+});
+
+test("previewFile never leaves the take's own directory", async () => {
+ const dir = await project();
+ const outside = await mkdtemp(path.join(tmpdir(), "takes-outside-"));
+ try {
+ const takeDir = await writeTake(dir, "ok", good({ id: "ok" }));
+ await writeFile(path.join(outside, "secret.mp4"), "secret");
+ assert.ok(await previewFile(dir, { id: "ok", preview: "preview.mp4" }));
+
+ // Lexical escapes, even if parseTake were bypassed.
+ assert.equal(await previewFile(dir, { id: "ok", preview: "../../video.manifest.json" }), null);
+ assert.equal(await previewFile(dir, { id: "../ok", preview: "preview.mp4" }), null);
+ assert.equal(await previewFile(dir, { id: "ok", preview: "." }), null);
+
+ // A symlinked preview pointing out of the take.
+ await symlink(path.join(outside, "secret.mp4"), path.join(takeDir, "leak.mp4"));
+ assert.equal(await previewFile(dir, { id: "ok", preview: "leak.mp4" }), null);
+
+ // A take directory that is itself a link out of takes/.
+ await mkdir(path.join(outside, "linked"));
+ await writeFile(path.join(outside, "linked", "preview.mp4"), "x");
+ await symlink(path.join(outside, "linked"), path.join(dir, "takes", "linked"));
+ assert.equal(await previewFile(dir, { id: "linked", preview: "preview.mp4" }), null);
+
+ // A directory is not a file; a missing file is null.
+ await mkdir(path.join(takeDir, "sub"));
+ assert.equal(await previewFile(dir, { id: "ok", preview: "sub" }), null);
+ assert.equal(await previewFile(dir, { id: "ok", preview: "nope.mp4" }), null);
+ } finally {
+ await rm(dir, { recursive: true, force: true });
+ await rm(outside, { recursive: true, force: true });
+ }
+});
+
+test("setTakeVerdict writes the contract's shape, merges, and removes an empty row", async () => {
+ const dir = await project();
+ try {
+ await mkdir(path.join(dir, "takes"), { recursive: true });
+ assert.deepEqual(await readVerdicts(dir), {});
+
+ const a = await setTakeVerdict(dir, "slow-burn", { verdict: "like" });
+ assert.equal(a.entry.verdict, "like");
+ assert.equal(a.entry.note, "");
+ assert.ok(!Number.isNaN(Date.parse(a.entry.at)));
+
+ // A note alone keeps the verdict; trailing whitespace is not kept.
+ await setTakeVerdict(dir, "slow-burn", { note: "the wait is right \n" });
+ await setTakeVerdict(dir, "ref", { verdict: "no", note: "too fast" });
+ const onDisk = JSON.parse(await readFile(verdictsFileOf(dir), "utf8"));
+ assert.deepEqual(Object.keys(onDisk).sort(), ["ref", "slow-burn"]);
+ assert.deepEqual(
+ { verdict: onDisk["slow-burn"].verdict, note: onDisk["slow-burn"].note },
+ { verdict: "like", note: "the wait is right" },
+ );
+ assert.deepEqual(Object.keys(onDisk["ref"]).sort(), ["at", "note", "verdict"]);
+
+ // Clearing the verdict keeps a row with a note; clearing both removes it.
+ await setTakeVerdict(dir, "ref", { verdict: null });
+ assert.deepEqual((await readVerdicts(dir)).ref.verdict, null);
+ const gone = await setTakeVerdict(dir, "ref", { note: "" });
+ assert.equal(gone.entry, null);
+ assert.deepEqual(Object.keys(await readVerdicts(dir)), ["slow-burn"]);
+ } finally {
+ await rm(dir, { recursive: true, force: true });
+ }
+});
+
+test("setTakeVerdict refuses bad input and never overwrites a file it cannot read", async () => {
+ const dir = await project();
+ try {
+ await mkdir(path.join(dir, "takes"), { recursive: true });
+ await assert.rejects(setTakeVerdict(dir, "../x", { verdict: "like" }), /take id/);
+ await assert.rejects(setTakeVerdict(dir, "a", { verdict: "keep" }), /verdict must be/);
+ await assert.rejects(setTakeVerdict(dir, "a", { note: 4 }), /note must be/);
+
+ await writeFile(verdictsFileOf(dir), "{ half a file");
+ await assert.rejects(setTakeVerdict(dir, "a", { verdict: "like" }), /does not parse/);
+ assert.equal(await readFile(verdictsFileOf(dir), "utf8"), "{ half a file");
+ } finally {
+ await rm(dir, { recursive: true, force: true });
+ }
+});
+
+test("concurrent verdicts on different takes are all kept", async () => {
+ const dir = await project();
+ try {
+ await mkdir(path.join(dir, "takes"), { recursive: true });
+ const ids = Array.from({ length: 12 }, (_, i) => `t${i}`);
+ await Promise.all(ids.map((id, i) => setTakeVerdict(dir, id, { verdict: ["like", "maybe", "no"][i % 3] })));
+ assert.deepEqual(Object.keys(await readVerdicts(dir)).sort(), [...ids].sort());
+ } finally {
+ await rm(dir, { recursive: true, force: true });
+ }
+});
+
+test("takesWithVerdicts joins take.json and verdicts.json, keeping a verdict on a vanished take", async () => {
+ const { takesWithVerdicts } = await import("./takes.mjs");
+ const dir = await mkdtemp(path.join(tmpdir(), "umtool-takes-digest-"));
+ try {
+ await mkdir(path.join(dir, "takes", "a"), { recursive: true });
+ await writeFile(
+ path.join(dir, "takes", "a", "take.json"),
+ JSON.stringify({ id: "a", group: "g", order: 1, label: "A", kind: "reference", preview: "p.mp4", summary: "s" }),
+ );
+ await writeFile(
+ path.join(dir, "takes", "verdicts.json"),
+ JSON.stringify({ a: { verdict: "like", note: "yes", at: "x" }, gone: { verdict: "no", note: "", at: "y" } }),
+ );
+ const rows = await takesWithVerdicts(dir);
+ assert.deepEqual(rows.map((r) => [r.id, r.label, r.verdict, r.note, r.missing ?? false]), [
+ ["a", "A", "like", "yes", false],
+ ["gone", null, "no", "", true],
+ ]);
+ } finally {
+ await rm(dir, { recursive: true });
+ }
+});
diff --git a/umtool/playwright.config.ts b/umtool/playwright.config.ts
@@ -76,6 +76,10 @@ export default defineConfig({
// var. CHANNELS_DIR has to be said explicitly: it is where a report
// video's cue files live, and its default is the real 3 GB corpus.
`CHANNELS_DIR=${FIXTURE}/channels ` +
+ // The sites (/sites, article notes): the fixture's own, never the real
+ // transcripts/sites -- the one corpus file umtool writes is a report's
+ // notes.json, and the suite writes them.
+ `SITES_DIR=${FIXTURE}/sites ` +
// The cache (the project index, posters, analyses) is no longer under
// SONG_DIR (release 17): its default is the user's ~/.cache, which a
// suite must never write. The fixture's own, rebuilt with it every run.
diff --git a/umtool/report-to-video/README.md b/umtool/report-to-video/README.md
@@ -411,7 +411,9 @@ characters, trimmed; an empty value deletes the key, the way the attribution
fields do. Without it the deck's title is empty (a card falls back to
`heading`) and the subtitle is auto-built: a clip's channel · title · date (the
channel only when the cut spans more than one), an image's `title · date`, a
-card's `sub`. The QR follows a clip's corner-QR rule unchanged (`citeUrl`, else
+card's `sub`. A line too long for its column gives way in its middle parts
+(the title ends in "…"): its first part and its last (the date) always show
+whole. The QR follows a clip's corner-QR rule unchanged (`citeUrl`, else
the site link at the clip's start); an image draws one only with an explicit
`citeUrl`; a card never does.
@@ -481,8 +483,10 @@ footage, so it rides on a clip.
its day), `@handle · Bluesky` (or X), the words — paragraphs kept, clamped to
`maxLines` with an ellipsis — and a QR of the post's page on the archive
(below), or of its own `url`. A post with a `shot` draws the screenshot in
- place of the words, as wide as the card's text and no taller than a full card
- of words, or than `shotMaxHeight` px when that is set.
+ place of the words, always as wide as the card's text (so it reads), in a
+ viewport no taller than a full card of words, or than `shotMaxHeight` px when
+ that is set; a taller screenshot holds on its top for 1.2 s once the card has
+ landed, pans to its bottom, and holds there 1.2 s before the card leaves.
- **Marks.** A post's own `accent`, `logo` and `flag` set one kind of card apart
from another at a glance — a source document's sentence beside a platform post,
say: the rail and rim in the accent, the logo above the QR in the top corner,
@@ -597,6 +601,62 @@ a verdict named without a colour keeps the default's), the stamp's seconds and
corner, and whether and where the tally is drawn; `validateChrome()` refuses an
unknown key there as everywhere in the block.
+### `thread` and `render.chrome.threads` — the thread rail
+
+A cut whose clips make a few lines of argument can draw them as a rail of cards
+down the frame's left side. Each clip names its thread; the list names the
+threads in the rail's order, each with an optional outcome.
+
+```jsonc
+"render": { "chrome": { …, "threads": { "list": [
+ { "id": "bet", "label": "The bet: her career", // ≤ 32 characters, one line
+ "outcome": { "verdict": "CONTRADICTED", "label": "Walked back" } }, // label optional (≤ 24): else the verdict's
+ { "id": "aside", "label": "An aside" } ] } } } // no outcome: never stamped
+{ "type": "clip", "id": "c07", …, "thread": "bet" }
+```
+
+- **The layout.** With the rail on, the footage box moves to the frame's right
+ edge (24 px in) and the rail takes the left, as tall as the footage: a card
+ per thread (1–8), each its number, a dot per clip and its label.
+- **The motion.** The card of the thread on screen is lit and a string draws
+ from it to the picture; a clip's dot fills as the clip comes in; when the
+ thread's last clip ends (1.8 s before it hands over), its outcome is stamped
+ on its card in the verdict's colour. Before its thread plays a card is dim,
+ after it rests quieter, so by the last clip the rail is the whole argument.
+ The rail steps aside while a popup post has moved the footage over it.
+- **Checks.** `validateThreads()` (via `validateChrome()`) refuses a bad list
+ and the feed layout beside it; `validateThreadEntries()` refuses an entry
+ naming an unlisted thread, a teaser in a thread and a listed thread with no
+ clip. `schedule.json` gains `threads: { threads, runs, asides }` only when
+ the rail is on. `chrome-threads.mjs` is the page; `compose-chrome.mjs
+ --region threads` renders it; the build lays it last, like the stamps.
+
+### `render.chrome.flips` — THEN and NOW, back to back
+
+A cut built of pairs: a line from THEN and the opposite line from NOW. A panel
+down the frame's left (the rail's region: a cut has the rail or the panel)
+holds one pair while it plays.
+
+```jsonc
+"render": { "transition": 0.12, "chrome": { …, "deck": { "footageScale": 0.78 }, "flips": { "pairs": [
+ { "id": "vax", "topic": "Vaccines", // ≤ 32 characters
+ "then": { "entry": "c01", "when": "2019", "words": "…" }, // when ≤ 18, words ≤ 110: verbatim
+ "now": { "entry": "c02", "when": "2025", "words": "…" } } ] } } }
+```
+
+- **The motion.** The pair rises in with its THEN clip: its place (`03 / 11`),
+ the topic, the THEN card (tag, when, words) lit. As the NOW clip starts its
+ card slams in under it with a flash, the THEN card dims, and when both
+ `when`s carry a year the years between roll up on an odometer
+ ("+7 years later"). The pair lifts away as its NOW clip ends.
+- **Checks.** `validateFlips()` (via `validateChrome()`) refuses a bad pair
+ and the rail beside it; `validateFlipEntries()` refuses a side naming no
+ timeline entry, a teaser, a THEN after its NOW and an entry in two pairs.
+ `schedule.json` gains `flips: { pairs, asides }`. `chrome-flips.mjs` is the
+ page; `compose-chrome.mjs --region flips` renders it.
+- **Punch.** The panel is built for short clips — the line and nothing else —
+ and near-hard cuts (`transition` 0.12); a smaller `footageScale` gives it room.
+
### The `image` entry type
A still: the receipts a clip cannot say out loud — a post, a thread, a DM, a
@@ -776,8 +836,8 @@ falls on the teaser.
"dip": { "fade": 1.2, "black": 0.6 } } // optional: go to black before it (below)
```
-- **`lines`** — 1 to 5, each one line of at most 80 characters: a string, or
- `{ text, break }`, where `break` is the END of `text` drawn as a smaller,
+- **`lines`** — 1 to 8 in at most 5 rows (below), each one line of at most 80 characters: a string, or
+ `{ text, break, replace?, role?, hold?, together?, lead? }`, where `break` is the END of `text` drawn as a smaller,
wide-tracked second tier under the rest that pops 3/7 of a beat after it
(0.3 s at the default `beat`).
The words are data: they are drawn uppercase, and kept as written everywhere
@@ -785,6 +845,39 @@ falls on the teaser.
small wide-tracked overline between two accent rules, the last a mid-size
kicker (a date), and everything between a big heavy title; two lines are an
overline and a title; one is a title.
+- **Rows that say several things in turn: `replace`.** A line written
+ `{ "text": …, "replace": true }` takes the ROW of the line before it: it
+ slams in where that one was, and the one it replaces rises 64 px, shrinks
+ to 0.9×, blurs and fades over 0.3 s from its pop (`TEASER_MOTION.out*`),
+ gone before the new line's impact. So one row can step through a sequence
+ — his estimates, each pushing out the last — and a last line with
+ `replace` can resolve it (`"2027"` pushing out the last estimate). The
+ frame holds at most **5 rows** (`TEASER_LIMITS.rows`) of at most **8 lines**;
+ roles follow the ROW's position as they followed the line's, and a
+ replacing line takes the role of the line it replaces. The lines of a row
+ are stacked in one grid cell (`.slot`), so the row is as tall as its
+ tallest line and the layout never moves as they change. The first line
+ cannot replace.
+- **A line's own size: `role`.** `"overline"`, `"title"` or `"kicker"` on a
+ line overrides the role its row's position gives it — its type, its fit
+ and its hit. A row of estimates at `"kicker"` keeps the title above them
+ the biggest words on the card; `"2027"` at `"title"` makes the answer
+ outweigh them. A kicker-sized row above a title gets 44 px of air.
+- **`hold`** — 0 to 4 seconds on a line: held after its pop (after its
+ second tier's) before the next line pops, so a line about to be pushed out
+ — or the last before the kicker — can be read. On the last line it holds
+ the tail back as well.
+- **`together`** — `true` on a line with a `break`: its second tier pops
+ WITH it, on the line's one hit, instead of 3/7 of a beat later on a
+ lighter hit of its own. For a row that steps through quote-and-date pairs,
+ so each step is one hit, not two.
+- **`lead`** — a few words drawn ABOVE a line as a small wide-tracked tier,
+ the way a `break` is drawn below it: a label for the line ("Breaking
+ Ground" over "2027"). It drops in with its line's pop, on that line's hit,
+ and fits as a second tier does (at most 66 characters).
+- A teaser that uses none of these composes the page it always did,
+ byte for byte (`chrome-teaser.test.mjs` pins it), so no existing cut
+ re-renders.
- **What fits the frame** (`TEASER_LIMITS.fit`): a row wider than 80 % of the
frame shrinks to its role's floor and no further, so each row is also held to
what fits at that floor — a **title** at most **34** characters, a
diff --git a/umtool/report-to-video/build-video.mjs b/umtool/report-to-video/build-video.mjs
@@ -109,10 +109,12 @@ import { ensureWriteDir } from "../lib/report/storage.mjs";
import {
assertChrome, deckGeometry, deckOn, deckSchedule, endFadeOf, feedGeometry, feedOn, frameCount, hidesDeck, MUTE_FADE,
dipOf, muteSegmentSeconds, playWindow, postsGeometry, postWindows, resolveDeck, scheduleFrom, snapWindow, stampGeometry,
- teaserHits, teaserSeconds, teaserTitle,
+ teaserHits, teaserSeconds, teaserTitle, threadsGeometry,
validateCutEdits, validatePosts, validateTeasers,
} from "./deck.mjs";
import { validateClaims } from "./factcheck.mjs";
+import { validateThreadEntries } from "./threads.mjs";
+import { validateFlipEntries } from "./flips.mjs";
// The per-platform yt-dlp args (Rumble's `--impersonate chrome`): the ONE table,
// in common, plain JS so bare `node` can load it.
import { platformArgsForUrl } from "yt-dlp-transcript-common/ytdlp/platformArgs.mjs";
@@ -2298,7 +2300,7 @@ export function chromeOverlayChain(render, regions, inLabel, firstInputIdx, opts
// The posts feed (`name: "feed"`) and the fact-check stamps (`name:
// "stamp"`) are whole-cut sequences like the deck's, laid exactly as the
// deck's is.
- const deck = r.name === "deck" || r.name === "feed" || r.name === "stamp";
+ const deck = r.name === "deck" || r.name === "feed" || r.name === "stamp" || r.name === "threads" || r.name === "flips";
const posts = r.name === "posts";
inputs.push(
...(deck || posts ? ["-reinit_filter", "0"] : []),
@@ -2363,6 +2365,12 @@ export function feedRegion(render, frames) {
return { name: "feed", frames, ...feedGeometry(render).column };
}
+/** The thread rail as an overlay region: its frames at threadsGeometry's box. */
+export function threadsRegion(render, frames, name = "threads") {
+ const { cards: _c, ...box } = threadsGeometry(render);
+ return { name, frames, ...box };
+}
+
/** The fact-check stamps as an overlay region: their frames at stampGeometry's box. */
export function stampRegion(render, frames, schedule) {
return { name: "stamp", frames, ...stampGeometry(render, { feed: schedule.layout === "feed" }) };
@@ -3584,6 +3592,47 @@ async function renderDeck({ manifestPath, render, outDir, variant, schedule, fro
});
regions.push(stampRegion(render, st.frames, schedule));
}
+
+ // The thread rail (a schedule with `threads`): one sequence for the whole
+ // cut -- or the deck's window -- beside the footage, laid like the deck's.
+ if (schedule.threads?.threads?.length) {
+ EMIT("chrome", { phase: "compose", region: "threads", ...(duration != null ? { from, duration } : {}) });
+ const t1 = Date.now();
+ const th = await composeChrome({
+ manifestPath, outDir, variant, region: "threads", doRender: true,
+ fps: render.fps, workers: 2, quality: "high", format: "png-sequence",
+ ...(duration != null ? { from, duration } : {}),
+ });
+ if (th.frameCount !== want) {
+ throw new Error(`the thread rail's sequence is ${th.frameCount} frames but the deck's is ${want}`);
+ }
+ EMIT("chrome", {
+ phase: th.cached ? "cached" : "render", region: "threads",
+ frames: th.frameCount, key: th.key, dir: th.frames,
+ seconds: Number(((Date.now() - t1) / 1000).toFixed(1)),
+ });
+ regions.push(threadsRegion(render, th.frames));
+ }
+
+ // The flips panel (a schedule with `flips`): the same region as the rail.
+ if (schedule.flips?.pairs?.length) {
+ EMIT("chrome", { phase: "compose", region: "flips", ...(duration != null ? { from, duration } : {}) });
+ const t1 = Date.now();
+ const fl = await composeChrome({
+ manifestPath, outDir, variant, region: "flips", doRender: true,
+ fps: render.fps, workers: 2, quality: "high", format: "png-sequence",
+ ...(duration != null ? { from, duration } : {}),
+ });
+ if (fl.frameCount !== want) {
+ throw new Error(`the flips panel's sequence is ${fl.frameCount} frames but the deck's is ${want}`);
+ }
+ EMIT("chrome", {
+ phase: fl.cached ? "cached" : "render", region: "flips",
+ frames: fl.frameCount, key: fl.key, dir: fl.frames,
+ seconds: Number(((Date.now() - t1) / 1000).toFixed(1)),
+ });
+ regions.push(threadsRegion(render, fl.frames, "flips"));
+ }
return { regions, outLabel: "[hfout]" };
}
@@ -3882,7 +3931,10 @@ export async function buildVideo({ manifestPath, opts = {}, out, only, fetchOnly
// A clip's `muteFrom` and `render.endFade`, checked against the WHOLE
// manifest, deck or not: both are made where the cut is joined.
{
- const errors = [...validateCutEdits(whole), ...validateTeasers(whole), ...validateClaims(whole)];
+ const errors = [
+ ...validateCutEdits(whole), ...validateTeasers(whole), ...validateClaims(whole), ...validateThreadEntries(whole),
+ ...validateFlipEntries(whole),
+ ];
if (errors.length) throw new Error(`manifest: ${errors.join("; ")}`);
}
const deck = deckOn(render);
diff --git a/umtool/report-to-video/chrome-deck.mjs b/umtool/report-to-video/chrome-deck.mjs
@@ -485,9 +485,15 @@ export function deckHtml(schedule, render, opts = {}) {
text-overflow: ellipsis; font-family: 'DeckSansBold', sans-serif;
font-size: ${t.titleSize}px; line-height: ${t.titleBox}px; letter-spacing: -0.012em;
color: ${pal.fg}; }
- .deck-sub { display: block; max-width: ${t.width}px; white-space: nowrap; overflow: hidden;
- text-overflow: ellipsis; font-size: ${t.subtitleSize}px; line-height: ${t.subtitleBox}px;
+ /* The source line is a row of parts: when it runs long, the middle parts
+ (the title) give way and the first (who) and last (the date) always
+ show whole; a line of one or two parts shrinks its first. */
+ .deck-sub { display: flex; max-width: ${t.width}px; white-space: nowrap; overflow: hidden;
+ font-size: ${t.subtitleSize}px; line-height: ${t.subtitleBox}px;
letter-spacing: 0.005em; color: ${pal.muted}; }
+ .deck-sub .part, .deck-sub .sep { flex: none; }
+ .deck-sub .part:not(:first-child):not(:last-child), .deck-sub .part:first-child:nth-last-child(-n+3) {
+ flex: 0 1 auto; min-width: 0; overflow: hidden; text-overflow: ellipsis; }
.deck-sub .sep { color: ${pal.accent}; padding: 0 0.42em; font-family: 'DeckSansBold', sans-serif; }
.blade { position: absolute; right: 0; top: ${Math.round(t.titleBox * 0.14)}px; width: 4px;
height: ${Math.round(t.titleBox * 0.72)}px; border-radius: 2px; background: ${pal.accent};
diff --git a/umtool/report-to-video/chrome-flips.mjs b/umtool/report-to-video/chrome-flips.mjs
@@ -0,0 +1,190 @@
+// The FLIPS panel's composition: one HyperFrames page for a whole cut, the
+// frame's left side beside the footage (deck.mjs threadsGeometry -- the region
+// the thread rail would take), one pair on it at a time (flips.mjs): its
+// topic, the THEN card, the years between, the NOW card slamming in.
+//
+// PURE, like chrome-threads.mjs: a schedule and a render block in, an HTML
+// string out; compose-chrome.mjs renders it (`region: "flips"`).
+import { pageDuration, threadsGeometry } from "./deck.mjs";
+import { mix, rgba } from "./chrome-deck.mjs";
+import { flipCues } from "./flips.mjs";
+
+const esc = (s) =>
+ String(s ?? "")
+ .replace(/&/g, "&")
+ .replace(/</g, "<")
+ .replace(/>/g, ">")
+ .replace(/"/g, """);
+
+const r4 = (v) => Math.round(v * 10000) / 10000;
+
+/** THEN's colour when the palette names none (`palette.then`): a cool blue against the accent's NOW. */
+export const THEN_COLOR = "#6fa8dc";
+
+/**
+ * The panel's HTML, for the whole cut -- or a window of it (`from`/`duration`).
+ * `fonts` = `{ regular, bold }`; `gsap` the vendored script. `?still=<t>` and
+ * the preview's `deck:seek` take CUT seconds.
+ */
+export function flipsHtml(schedule, render, opts = {}) {
+ const sched = schedule.flips;
+ if (!sched?.pairs?.length) throw new Error("flips: the schedule has no pairs");
+ const geo = threadsGeometry(render);
+ const pal = render.palette;
+ const W = geo.width, H = geo.height;
+ const C = geo.cards;
+ const fonts = opts.fonts ?? {};
+ const gsapSrc = opts.gsap ?? "assets/gsap.min.js";
+ const total = schedule.total;
+ const from = Number(opts.from ?? 0);
+ const dur = opts.duration != null ? r4(Number(opts.duration)) : r4(total - from);
+ if (!(dur > 0)) throw new Error(`flips: nothing to render from ${from}s of a ${total}s cut`);
+ const windowed = from > 0 || Math.abs(dur - total) > 1e-6;
+ const { init, cues } = flipCues(sched);
+ const thenC = /^#[0-9a-fA-F]{6}$/.test(pal.then ?? "") ? pal.then : THEN_COLOR;
+ const nowC = pal.accent;
+ // Type from the panel's width: 298 px at the deck's default footage, wider with less footage.
+ const k = Math.max(0.8, Math.min(1.4, C.width / 300));
+ const S = (v) => Math.round(v * k);
+ const topicSize = S(30), idxSize = S(15), whenSize = S(50), wordsSize = S(27), tagSize = S(14), gapSize = S(22);
+ const wordsLine = Math.round(wordsSize * 1.22);
+ const gapLine = Math.round(gapSize * 1.2);
+
+ const odo = (years) =>
+ Array.from({ length: years + 1 }, (_, i) => `<span class="digit">+${i}</span>`).join("");
+ const sideHtml = (cls, tag, s, color, key) =>
+ `<div class="card ${cls}" data-k="${key}" style="--c:${color}; --cg:${rgba(color, 0.55)}; --cb:${rgba(mix(pal.bg, color, 0.16), 0.95)}">` +
+ `<div class="row"><span class="tag">${tag}</span><span class="when">${esc(s.when)}</span></div>` +
+ `<div class="words">“${esc(s.words)}”</div>`;
+
+ const pairsHtml = sched.pairs
+ .map((p, j) =>
+ `<div class="pair" data-pair="${esc(p.id)}" data-k="p${j}">` +
+ `<div class="top"><span class="idx">${String(p.index).padStart(2, "0")} / ${String(p.of).padStart(2, "0")}</span>` +
+ `<div class="topic">${esc(p.topic)}</div></div>` +
+ sideHtml("then", "THEN", p.then, thenC, `t${j}`) + `</div>` +
+ (p.years != null
+ ? `<div class="gap" data-k="g${j}"><span class="odo"><span class="strip" data-k="gc${j}">${odo(p.years)}</span></span>` +
+ `<span class="unit">${p.years === 1 ? "year" : "years"} later</span></div>`
+ : `<div class="gap blank"></div>`) +
+ sideHtml("now", "NOW", p.now, nowC, `n${j}`) + `<div class="flash" data-k="nf${j}"></div></div>` +
+ `</div>`)
+ .join("\n ");
+
+ const data = {
+ total: r4(total),
+ window: windowed ? { from: r4(from), dur } : null,
+ wordsSize,
+ ids: sched.pairs.map((p) => p.id),
+ init,
+ cues: cues.map(({ why, ...c }) => c),
+ };
+ const json = JSON.stringify(data).replace(/</g, "\\u003c");
+ const fps = schedule.fps ?? render.fps ?? 30;
+
+ return `<!doctype html>
+<html lang="en">
+ <head>
+ <meta charset="UTF-8" />
+ <meta name="viewport" content="width=${W}, height=${H}" />
+ <script src="${esc(gsapSrc)}"></script>
+ <style>
+ @font-face { font-family: 'DeckSans'; font-weight: 400; font-style: normal; src: url('${esc(fonts.regular ?? "")}'); }
+ @font-face { font-family: 'DeckSansBold'; font-weight: 400; font-style: normal; src: url('${esc(fonts.bold ?? "")}'); }
+ * { margin: 0; padding: 0; box-sizing: border-box; }
+ html, body { width: ${W}px; height: ${H}px; overflow: hidden; background: transparent; }
+ body { font-family: 'DeckSans', sans-serif; font-synthesis: none; color: ${pal.fg};
+ -webkit-font-smoothing: antialiased; text-rendering: geometricPrecision; }
+ #root { position: relative; width: ${W}px; height: ${H}px; overflow: hidden; }
+ #flips-clip, .panel { position: absolute; left: 0; top: 0; width: ${W}px; height: ${H}px; }
+ /* One pair at a time, centred in the panel's height. */
+ .pair { position: absolute; left: ${C.x}px; top: 0; width: ${C.width}px; height: ${H}px; visibility: hidden; opacity: 0;
+ display: flex; flex-direction: column; justify-content: center; gap: ${S(10)}px; }
+ .top { margin-bottom: ${S(6)}px; }
+ .idx { font-family: 'DeckSansBold', sans-serif; font-size: ${idxSize}px; letter-spacing: 0.14em; color: ${pal.muted};
+ font-variant-numeric: tabular-nums; }
+ .topic { margin-top: ${S(4)}px; font-family: 'DeckSansBold', sans-serif; font-size: ${topicSize}px;
+ line-height: ${Math.round(topicSize * 1.1)}px; letter-spacing: 0.04em; text-transform: uppercase; color: ${pal.fg};
+ display: -webkit-box; -webkit-box-orient: vertical; -webkit-line-clamp: 2; overflow: hidden; }
+ /* A side: its tag and when on a row, then the words. */
+ .card { position: relative; border-radius: 10px; padding: ${S(12)}px ${S(14)}px ${S(14)}px;
+ background: var(--cb); border-left: 6px solid var(--c);
+ box-shadow: inset 0 0 0 1px ${rgba(pal.fg, 0.08)}, 0 10px 26px rgba(0, 0, 0, 0.35); transform-origin: 20% 50%; }
+ .now { visibility: hidden; opacity: 0; }
+ .row { display: flex; align-items: baseline; gap: ${S(10)}px; }
+ .tag { font-family: 'DeckSansBold', sans-serif; font-size: ${tagSize}px; line-height: ${tagSize + 8}px; letter-spacing: 0.14em;
+ padding: 0 ${S(8)}px; border-radius: 4px; color: ${pal.bg}; background: var(--c); }
+ .when { font-family: 'DeckSansBold', sans-serif; font-size: ${whenSize}px; line-height: ${Math.round(whenSize * 1.05)}px;
+ color: var(--c); font-variant-numeric: tabular-nums; letter-spacing: -0.01em; }
+ .words { margin-top: ${S(6)}px; font-family: 'DeckSansBold', sans-serif; font-size: ${wordsSize}px; line-height: ${wordsLine}px;
+ color: ${pal.fg}; display: -webkit-box; -webkit-box-orient: vertical; -webkit-line-clamp: 5; overflow: hidden; }
+ .flash { position: absolute; inset: -4px; border-radius: 12px; opacity: 0; pointer-events: none;
+ box-shadow: 0 0 0 3px ${nowC}, 0 0 34px 10px ${rgba(nowC, 0.6)}; }
+ /* The years between: an odometer, then the unit. */
+ .gap { display: flex; align-items: center; justify-content: center; gap: ${S(8)}px; height: ${gapLine + S(6)}px;
+ visibility: hidden; opacity: 0; }
+ .gap.blank { visibility: hidden; }
+ .odo { display: inline-block; height: ${gapLine}px; overflow: hidden; }
+ .strip { display: flex; flex-direction: column; }
+ .digit { display: block; height: ${gapLine}px; font-family: 'DeckSansBold', sans-serif; font-size: ${gapSize}px;
+ line-height: ${gapLine}px; color: ${nowC}; font-variant-numeric: tabular-nums; text-align: right; }
+ .unit { font-size: ${gapSize}px; line-height: ${gapLine}px; color: ${pal.muted}; letter-spacing: 0.02em; }
+ </style>
+ </head>
+ <body>
+ <div id="root" data-composition-id="flips" data-start="0" data-duration="${pageDuration(dur, fps)}"
+ data-width="${W}" data-height="${H}">
+ <div id="flips-clip" class="clip" data-start="0" data-duration="${pageDuration(dur, fps)}" data-track-index="1">
+ <div class="panel" data-k="panel">
+ ${pairsHtml}
+ </div>
+ </div>
+ </div>
+
+ <script id="flips-data" type="application/json">${json}</script>
+ <script>
+ const S = JSON.parse(document.getElementById("flips-data").textContent);
+ const byK = {};
+ for (const el of document.querySelectorAll("[data-k]")) byK[el.dataset.k] = el;
+ for (const k of Object.keys(S.init)) if (byK[k]) gsap.set(byK[k], S.init[k]);
+ const inner = gsap.timeline({ paused: true });
+ for (const c of S.cues) {
+ const el = byK[c.k];
+ if (!el) continue;
+ inner.fromTo(el, c.from, { ...c.to, duration: c.dur, ease: c.ease, immediateRender: false }, c.at);
+ }
+ inner.set({}, {}, S.total);
+ let tl = inner;
+ if (S.window) {
+ tl = gsap.timeline({ paused: true });
+ tl.add(inner.tweenFromTo(S.window.from, S.window.from + S.window.dur, { duration: S.window.dur, ease: "none" }), 0);
+ }
+ window.__timelines = window.__timelines || {};
+ window.__timelines["flips"] = tl;
+ const ready = document.fonts.load(S.wordsSize + "px DeckSansBold").catch(() => {}).then(() => {
+ document.documentElement.dataset.fit = "1";
+ });
+ const params = new URLSearchParams(location.search);
+ const local = (t) => {
+ const v = Number(t) || 0;
+ return S.window ? Math.max(0, Math.min(S.window.dur, v - S.window.from)) : Math.max(0, v);
+ };
+ const still = params.get("still");
+ if (still !== null) {
+ tl.seek(local(still), false);
+ ready.then(() => tl.seek(local(still), false));
+ }
+ if (params.get("preview") === "1") {
+ window.addEventListener("message", (e) => {
+ const m = e.data || {};
+ if (m.type === "deck:seek") tl.seek(local(m.t), false);
+ });
+ ready.then(() => {
+ if (window.parent !== window) window.parent.postMessage({ type: "flips:ready", total: S.total, ids: S.ids }, "*");
+ });
+ }
+ </script>
+ </body>
+</html>
+`;
+}
diff --git a/umtool/report-to-video/chrome-posts.mjs b/umtool/report-to-video/chrome-posts.mjs
@@ -93,12 +93,19 @@ export function embedFn(name, fn) {
* every card still up leaves over its `out` (a front-loaded fade and a slight
* shrink).
*
+ * `pans` are, per card, how many px its screenshot runs past its viewport (0:
+ * it fits). A tall screenshot is drawn at the card's full width -- readable --
+ * in a viewport no taller than the column allows, and pans up through the
+ * rest (`s<j>`, its y): it holds `panHold` s on its top once the card has
+ * landed, glides to its bottom, and holds there `panHold` s before it leaves.
+ * With too little time for the holds, the glide takes all of it.
+ *
* @returns {{ tops: number[], init: Record<string, object>,
* cues: Array<{ k: string, at: number, dur: number, from: object, to: object, ease: string, why: string }> }}
*/
export function postsCues({
posts, heights, column, gap = 14, enter = 0.55, slide = 0.35, enterX = 624,
- glowAt = 0.3, glowUp = 0.18, glowDown = 1.1, glowRest = 0.3,
+ glowAt = 0.3, glowUp = 0.18, glowDown = 1.1, glowRest = 0.3, pans = [], panHold = 1.2,
}) {
const R = (v) => Math.round(v * 10000) / 10000;
const MIN = 0.001;
@@ -112,6 +119,7 @@ export function postsCues({
for (let j = 0; j < posts.length; j += 1) {
init[`c${j}`] = { autoAlpha: 0, x: enterX, scale: 1 };
init[`g${j}`] = { opacity: 0 };
+ if (pans[j] > 0) init[`s${j}`] = { y: 0 };
}
const ev = [];
@@ -135,6 +143,14 @@ export function postsCues({
// The flare, as the card lands; then it settles to a quiet rim.
add(`g${j}`, t + glowAt, glowUp, { opacity: 1 }, "power2.out", `glow ${p.id}`);
add(`g${j}`, t + glowAt + glowUp, glowDown, { opacity: glowRest }, "power2.inOut", `settle ${p.id}`);
+ if (pans[j] > 0) {
+ const landed = t + enter;
+ const leave = p.out[0];
+ let a = landed + panHold;
+ let b = leave - panHold;
+ if (b - a < 1) { a = landed; b = Math.max(landed + MIN, leave); }
+ add(`s${j}`, a, b - a, { y: -pans[j] }, "sine.inOut", `pan ${p.id}`);
+ }
visible.push(j);
}
for (const j of visible) {
@@ -252,7 +268,7 @@ export function postsHtml(schedule, render, window, opts = {}) {
return (
`<article class="post${shot ? " has-shot" : ""}${logo ? " has-logo" : ""}" data-post="${esc(p.id)}" data-k="c${j}"${style}>` +
(shot
- ? `<div class="body shot-body">${flag}<img class="shot" src="${esc(shot)}" alt=""></div>`
+ ? `<div class="body shot-body">${flag}<div class="shot-view"><img class="shot" data-k="s${j}" src="${esc(shot)}" alt=""></div></div>`
: `<div class="body">${flag}` +
`<div class="meta">` +
(platform ? `<span class="platform">${esc(platform)}</span>` : "") +
@@ -355,11 +371,14 @@ export function postsHtml(schedule, render, window, opts = {}) {
display: -webkit-box; -webkit-box-orient: vertical; -webkit-line-clamp: ${set.maxLines}; }
.para + .para { margin-top: ${Math.round(lineH * 0.36)}px; }
.para.gone { display: none; }
- /* A post's screenshot in place of its words: as wide as the words would
- be, no taller than a full card of them -- or than \`shotMaxHeight\`. */
+ /* A post's screenshot in place of its words: always as wide as the words
+ would be, so it reads; its viewport is no taller than a full card of
+ them -- or than \`shotMaxHeight\` -- and a taller one pans up through
+ it (postsCues \`pans\`). */
.shot-body { padding: ${pad - 6}px; }
- .shot { display: block; width: 100%; height: auto; max-height: ${set.shotMaxHeight ?? Math.round(metaSize * 1.3) + 8 + set.maxLines * lineH + 2 * pad}px;
- object-fit: contain; object-position: left top; border-radius: 6px; }
+ .shot-view { position: relative; overflow: hidden; border-radius: 6px;
+ max-height: ${set.shotMaxHeight ?? Math.round(metaSize * 1.3) + 8 + set.maxLines * lineH + 2 * pad}px; }
+ .shot { display: block; width: 100%; height: auto; }
/* The source cell: the QR in a cell a shade down, as on the deck. */
.plate { position: absolute; right: 0; top: 0; bottom: 0; width: ${plateW}px;
display: flex; flex-direction: column; align-items: center; justify-content: center; gap: 10px;
@@ -428,7 +447,14 @@ export function postsHtml(schedule, render, window, opts = {}) {
const cards = P.ids.map((_, j) => byK["c" + j]);
cards.forEach(clampText);
const heights = cards.map((c) => c.offsetHeight);
- const plan = postsCues({ posts: P.posts, heights, column: P.column, gap: P.gap,
+ // How far each screenshot runs past its viewport: that much to pan.
+ const pans = cards.map((c) => {
+ const view = c.querySelector(".shot-view");
+ const img = view && view.querySelector("img.shot");
+ return img ? Math.max(0, Math.round(img.offsetHeight - view.clientHeight)) : 0;
+ });
+ document.documentElement.dataset.pans = pans.join(",");
+ const plan = postsCues({ posts: P.posts, heights, pans, column: P.column, gap: P.gap,
enter: P.enter, slide: P.slide, enterX: P.enterX, glowAt: P.glowAt,
glowUp: P.glowUp, glowDown: P.glowDown, glowRest: P.glowRest });
cards.forEach((c, j) => { c.style.top = plan.tops[j] + "px"; });
diff --git a/umtool/report-to-video/chrome-posts.test.mjs b/umtool/report-to-video/chrome-posts.test.mjs
@@ -467,7 +467,7 @@ test("a post with a screenshot draws it in place of its text card, its QR cell k
const open = html.indexOf('data-post="a"');
const card = html.slice(open, html.indexOf("<article", open + 1));
assert.match(html, /<article class="post has-shot" data-post="a"/);
- assert.ok(card.includes('<img class="shot" src="assets/shot00.png" alt="">'));
+ assert.match(card, /<div class="shot-view"><img class="shot" data-k="s\d+" src="assets\/shot00\.png" alt=""><\/div>/);
assert.doesNotMatch(card, /class="para"|class="meta"/, "no words beside the picture");
assert.ok(card.includes('<img src="assets/qr00.png"'), "the QR stays");
// The post without one is the text card it always was.
@@ -495,7 +495,7 @@ test("a post's accent, logo and flag mark its card; a post without them is drawn
const card = html.slice(open - 60, html.indexOf("<article", open + 1));
assert.match(card, /class="post has-shot has-logo"/);
assert.match(card, /style="--acc: #c98fd6; --acc-rim: rgba\(201, ?143, ?214, ?0\.95\)/);
- assert.ok(card.includes('<div class="flag">No source in the article</div><img class="shot"'), "the flag heads the picture");
+ assert.ok(card.includes('<div class="flag">No source in the article</div><div class="shot-view"><img class="shot"'), "the flag heads the picture");
assert.ok(card.includes('<img class="logo" src="assets/logo00.png" alt=""><img src="assets/qr00.png"'), "logo above the QR");
const b = html.slice(html.indexOf('data-post="b"') - 60);
assert.doesNotMatch(b.slice(0, b.indexOf("</article>")), /--acc|class="flag"|class="logo"/);
@@ -508,11 +508,43 @@ test("a post's accent, logo and flag mark its card; a post without them is drawn
assert.match(bad, /flag must be a short label/);
});
-test("shotMaxHeight caps a screenshot in px; unset, a full card of words does", () => {
+test("shotMaxHeight caps a screenshot's viewport in px; unset, a full card of words does", () => {
const sched = schedule([POST("a", "2024-10-19T17:01:17.640Z", { shot: "shots/a.png" })]);
const win = snapWindow(postWindows(sched)[0], { fps: 30, total: sched.total });
- const cap = (render) => /\.shot \{[^}]*max-height: (\d+)px/.exec(postsHtml(sched, render, win, { fonts: FONTS, shotSrcs: { a: "assets/shot00.png" } }))[1];
+ const cap = (render) => /\.shot-view \{[^}]*max-height: (\d+)px/.exec(postsHtml(sched, render, win, { fonts: FONTS, shotSrcs: { a: "assets/shot00.png" } }))[1];
const tall = { ...RENDER, chrome: { ...RENDER.chrome, deck: { ...RENDER.chrome.deck, posts: { ...RENDER.chrome.deck?.posts, shotMaxHeight: 820 } } } };
assert.equal(cap(tall), "820");
assert.notEqual(cap(RENDER), "820");
});
+
+test("postsCues: a tall screenshot holds on its top, pans to its bottom, and holds before it leaves", () => {
+ const posts = [{ id: "p0", appear: 10, out: [30, 30.5] }, { id: "p1", appear: 12, out: [30, 30.5] }];
+ const plan = postsCues({ posts, heights: [800, 200], pans: [600, 0], column: 2000, enter: 0.5, panHold: 1.2 });
+ assert.deepEqual(plan.init.s0, { y: 0 });
+ assert.equal(plan.init.s1, undefined, "a screenshot that fits does not pan");
+ const pan = plan.cues.filter((c) => c.k === "s0");
+ assert.equal(pan.length, 1);
+ near(pan[0].at, 10 + 0.5 + 1.2, "after the card lands and a hold on its top");
+ near(pan[0].at + pan[0].dur, 30 - 1.2, "a hold on its bottom before it leaves");
+ assert.deepEqual([pan[0].from, pan[0].to], [{ y: 0 }, { y: -600 }]);
+ // Too little time for the holds: the glide takes all of it.
+ const tight = postsCues({ posts: [{ id: "q", appear: 0, out: [2, 2.5] }], heights: [800], pans: [300], column: 2000, enter: 0.5 });
+ const g = tight.cues.find((c) => c.k === "s0");
+ near(g.at, 0.5, "from the landing");
+ near(g.at + g.dur, 2, "to the leave");
+});
+
+test("postsHtml: a screenshot sits in a viewport at the card's width, and the page measures its pan", () => {
+ const html = postsHtml(
+ { fps: 30, total: 40, segments: [{ id: "c1", start: 0, duration: 40 }], posts: [
+ { id: "x1", segment: "c1", slot: 0, appear: 2, out: [30, 30.5], platform: "x", handle: "a", date: "2026-01-01", text: "t", shot: "s.png" },
+ ] },
+ { ...RENDER, chrome: { ...RENDER.chrome, deck: { posts: { shotMaxHeight: 800 } } } },
+ { segment: "c1", from: 0, to: 31 },
+ { shotSrcs: { x1: "assets/shot0.png" } },
+ );
+ assert.match(html, /<div class="shot-view"><img class="shot" data-k="s0" src="assets\/shot0.png"/);
+ assert.match(html, /\.shot \{ display: block; width: 100%; height: auto; \}/);
+ assert.doesNotMatch(html, /\.shot \{[^}]*object-fit/, "never shrunk to fit");
+ assert.match(html, /const pans = cards\.map/);
+});
diff --git a/umtool/report-to-video/chrome-teaser.mjs b/umtool/report-to-video/chrome-teaser.mjs
@@ -164,6 +164,22 @@ export function teaserCues({ lines, tail = "", seconds, motion = TEASER_MOTION,
put(`l${i}.rules`, { scaleX: 0 });
add(`l${i}.rules`, hit - T(0.04), T(0.6), { scaleX: 1 }, "expo.out", `${why} rules`);
}
+ // Pushed out: the line after it takes its row (`replace`), and as that
+ // one slams in this one rises, shrinks a little, blurs and is gone.
+ if (b.outAt != null) {
+ put(`l${i}.o`, { y: 0 });
+ add(`l${i}.o`, b.outAt - T(0.02), T(m.out), { autoAlpha: 0, y: -m.outRise }, "power2.in", `${why} pushed out`);
+ add(`l${i}`, b.outAt - T(0.02), T(m.out), { scale: m.outScale }, "power2.in", `${why} pushed out (scale)`);
+ add(`l${i}.t`, b.outAt - T(0.02), T(m.out), { "--blur": `${m.outBlur}px` }, "power2.in", `${why} pushed out (blur)`);
+ }
+ // A lead -- the small tier above the line -- drops in with the line's pop,
+ // as a second tier rises in under one.
+ if (l.lead) {
+ put(`l${i}.lead`, { autoAlpha: 0, y: -16, scale: 1.12 });
+ put(`l${i}.leadt`, { "--blur": "10px" });
+ add(`l${i}.lead`, b.at, T(0.5), { autoAlpha: 1, y: 0, scale: 1 }, "expo.out", `${why} lead`);
+ add(`l${i}.leadt`, b.at, T(0.32), { "--blur": "0px" }, "power2.out", `${why} lead focus`);
+ }
if (l.sub && b.subAt != null) {
put(`l${i}.sub`, { autoAlpha: 0, y: 16, scale: 1.12 });
put(`l${i}.subt`, { "--blur": "10px" });
@@ -252,7 +268,7 @@ export function teaserHtml(entry, render, opts = {}) {
const maxW = Math.round(W * 0.8);
const ty = TEASER_TYPE;
- const lineHtml = lines
+ const lineHtmls = lines
.map((l, i) => {
const k = `l${i}`;
const isLast = i === lines.length - 1;
@@ -270,6 +286,9 @@ export function teaserHtml(entry, render, opts = {}) {
`<div class="flash" data-k="${k}.flash"></div>` +
`<div class="streak" data-k="${k}.streak"></div>` +
rules +
+ (l.lead
+ ? `<div class="sub lead" data-k="${k}.lead"><div class="row"><span class="subt" data-k="${k}.leadt">${esc(l.lead)}</span></div></div>`
+ : "") +
`<div class="pop" data-k="${k}"><div class="row">` +
`<span class="txt" data-k="${k}.t">${esc(l.head)}</span>${l.sub ? "" : tailHtml}</div></div>` +
(l.sub
@@ -277,8 +296,18 @@ export function teaserHtml(entry, render, opts = {}) {
: "") +
`</div>`
);
- })
+ });
+ // A row that several lines take in turn (`replace`) stacks them in one
+ // cell, so the row is as tall as its tallest and each pushes the last out
+ // in place. A row of one line is the line itself, as it always was.
+ const rowsOf = [];
+ lines.forEach((l, i) => { (rowsOf[l.row] ??= []).push(i); });
+ const lineHtml = rowsOf
+ .map((members) => members.length === 1
+ ? lineHtmls[members[0]]
+ : `<div class="slot r-${lines[members[0]].role}">\n ${members.map((i) => lineHtmls[i]).join("\n ")}\n </div>`)
.join("\n ");
+ const slotted = rowsOf.some((m) => m.length > 1);
const data = {
seconds,
@@ -338,8 +367,19 @@ export function teaserHtml(entry, render, opts = {}) {
.kicker .txt { color: ${pal.fg}; }
.overline { margin-bottom: 38px; }
.title + .title { margin-top: 10px; }
- .title .sub { margin-top: 14px; }
- .kicker { margin-top: 70px; }
+ .title .sub { margin-top: 14px; }${lines.some((l) => l.lead) ? `
+ /* A lead: the small tier above its line, as a break is the one below. */
+ .line .sub.lead { margin-top: 0; margin-bottom: 10px; }` : ""}
+ .kicker { margin-top: 70px; }${slotted || lines.some((l, i) => l.role === "kicker" && i < lines.length - 1) ? `
+ /* Rows several lines take in turn: one grid cell, the lines stacked in
+ it; the row carries the margins its lines would have. */
+ .slot { position: relative; display: grid; justify-items: center; align-items: center; }
+ .slot > .line { grid-area: 1 / 1; margin: 0; }
+ .line.title + .slot.r-title, .slot.r-title + .line.title, .slot + .slot { margin-top: 10px; }
+ .slot.r-kicker { margin-top: 70px; }
+ .slot.r-overline { margin-bottom: 38px; }
+ /* A kicker-sized row above a title (a line's own role) keeps air between them. */
+ .kicker + .title, .slot.r-kicker + .title { margin-top: 44px; }` : ""}
/* An overline between two hairline rules in the accent. */
.rules { position: absolute; left: -132px; right: -132px; top: 50%; height: 2px; transform-origin: 50% 50%; }
.rule { position: absolute; top: 0; width: 96px; height: 2px; border-radius: 1px; }
diff --git a/umtool/report-to-video/chrome-teaser.test.mjs b/umtool/report-to-video/chrome-teaser.test.mjs
@@ -33,15 +33,26 @@ test("a valid teaser has nothing to say; every bad shape is a sentence", () => {
assert.deepEqual(validateTeaser(FERRET), []);
assert.deepEqual(validateTeaser({ ...FERRET, tail: undefined, hits: false }), []);
const bad = (patch) => validateTeaser({ ...FERRET, ...patch }).join(" | ");
- assert.match(bad({ lines: [] }), /lines must be a list of 1 to 5/);
- assert.match(bad({ lines: ["a", "b", "c", "d", "e", "f"] }), /1 to 5/);
+ assert.match(bad({ lines: [] }), /lines must be a list of 1 to 8/);
+ assert.match(bad({ lines: ["a", "b", "c", "d", "e", "f", "g", "h", "i"] }), /1 to 8/);
+ // Six ROWS do not fit; six lines in five rows do.
+ assert.match(bad({ lines: ["a", "b", "c", "d", "e", "f"], seconds: undefined }), /make 6 rows, and at most 5 fit/);
+ assert.deepEqual(validateTeaser({ ...FERRET, seconds: undefined, lines: ["a", "b", "c", { text: "c2", replace: true }, "d", "e"] }), []);
assert.match(bad({ lines: ["ok", ""] }), /lines\[1\] must be words/);
assert.match(bad({ lines: ["two\nlines"] }), /one line/);
assert.match(bad({ lines: ["x".repeat(81)] }), /81 characters/);
assert.match(bad({ lines: [{ text: "Abc", brk: "c" }] }), /lines\[0\]\.brk is not a teaser line field/);
assert.match(bad({ lines: [{ text: "The big one", break: "small" }] }), /must be the end of its text/);
assert.match(bad({ lines: [{ text: "whole", break: "whole" }] }), /leaves nothing for the first tier/);
- assert.match(bad({ lines: [42] }), /string or \{ text, break \}/);
+ assert.match(bad({ lines: [42] }), /string or \{ text, break, replace, role, hold, together, lead \}/);
+ assert.match(bad({ lines: ["a", { text: "b", lead: "x".repeat(67) }] }), /lines\[1\]\.lead is 67 characters, and at most 66 fit/);
+ assert.match(bad({ lines: ["a", { text: "b", lead: "" }] }), /lines\[1\]\.lead must be words/);
+ assert.match(bad({ lines: ["a", { text: "b", together: true }] }), /lines\[1\]\.together pops the second tier with its line, and it has no break/);
+ assert.match(bad({ lines: ["a", { text: "b c", break: "c", together: 1 }] }), /lines\[1\]\.together must be true or false/);
+ assert.match(bad({ lines: [{ text: "First", replace: true }] }), /lines\[0\]\.replace: the first line has no line before it/);
+ assert.match(bad({ lines: ["a", { text: "b", replace: "yes" }] }), /lines\[1\]\.replace must be true or false/);
+ assert.match(bad({ lines: ["a", { text: "b", role: "huge" }] }), /lines\[1\]\.role must be one of overline, title, kicker/);
+ assert.match(bad({ lines: ["a", { text: "b", hold: 5 }] }), /lines\[1\]\.hold must be from 0 to 4 seconds/);
assert.match(bad({ seconds: 2 }), /seconds must be from 3 to 20/);
assert.match(bad({ seconds: "7" }), /seconds must be from 3 to 20, or absent/);
assert.match(bad({ beat: 0.3 }), /beat must be from 0\.4 to 2\.5 seconds/);
@@ -400,3 +411,115 @@ test("tailWait: the tail enters that long after the last hit; the swell and the
assert.match(teaserHtml(FERRET, RENDER), /<span class="tail-hang"><span class="tail" data-k="tail">/);
assert.match(teaserHtml(FERRET, RENDER), /\.tail-hang \{ display: inline-block; width: 0; white-space: nowrap; \}/);
});
+
+// ---- rows several lines take in turn (`replace`), a line's own `role`, `hold` ----
+
+const ESTIMATES = Object.freeze({
+ type: "teaser", id: "fin", beat: 1,
+ lines: [
+ "Pirate Software",
+ { text: "The Largest Ferret Rescue in the United States", break: "in the United States" },
+ { text: "“Months” March 2024", break: "March 2024" },
+ { text: "“This year” January 2025", break: "January 2025", replace: true },
+ { text: "“Within the next six months” August 2026", break: "August 2026", replace: true, hold: 0.8 },
+ "2027",
+ ],
+ tail: "?",
+});
+
+test("replace: a line takes the row before it; roles follow rows; a line's own role wins", () => {
+ assert.deepEqual(validateTeaser(ESTIMATES), []);
+ const lines = teaserLines(ESTIMATES);
+ assert.deepEqual(lines.map((l) => l.row), [0, 1, 2, 2, 2, 3]);
+ // Four rows: overline, two titles, the kicker -- the estimates inherit their row's.
+ assert.deepEqual(lines.map((l) => l.role), ["overline", "title", "title", "title", "title", "kicker"]);
+ assert.deepEqual(lines.map((l) => l.replace ?? false), [false, false, false, true, true, false]);
+ // A replacing line keeps a role the line it replaces was given.
+ const small = { ...ESTIMATES, lines: ESTIMATES.lines.map((l, i) => (i === 2 ? { ...l, role: "kicker" } : l)) };
+ assert.deepEqual(teaserLines(small).map((l) => l.role), ["overline", "title", "kicker", "kicker", "kicker", "kicker"]);
+ // An entry that uses none of it normalises exactly as before: no new keys.
+ for (const l of teaserLines(FERRET)) assert.deepEqual(Object.keys(l).sort(), ["head", "role", "row", "sub", "text"]);
+});
+
+test("replace: the pushed line leaves as the next pops; hold waits before the next line", () => {
+ const lines = teaserLines(ESTIMATES);
+ const m = teaserMotion(1);
+ const { beats, cues, init } = teaserCues({ lines, tail: "?", seconds: 14, motion: m });
+ const b = beats.lines;
+ // Only a replaced line has an outAt, at its replacer's pop.
+ assert.deepEqual(b.map((x) => x.outAt ?? null), [null, null, b[3].at, b[4].at, null, null]);
+ // The hold: the kicker pops a beat AND the hold after the last estimate's second tier.
+ assert.equal(Math.round((b[5].at - b[4].subAt) * 1e4) / 1e4, Math.round((m.gap + 0.8) * 1e4) / 1e4);
+ // Pushed out: gone (autoAlpha 0) and risen, by the replacer's impact.
+ for (const i of [2, 3]) {
+ const out = cues.find((c) => c.why === `line ${i} pushed out`);
+ assert.deepEqual(out.to, { autoAlpha: 0, y: -m.outRise });
+ assert.ok(out.at + out.dur <= b[i + 1].impact + 0.12, `line ${i} is gone as the next lands`);
+ }
+ // Seek-safe and every from stated, as for any teaser.
+ const state = JSON.parse(JSON.stringify(init));
+ for (const c of cues) {
+ for (const [p, v] of Object.entries(c.from)) assert.deepEqual(v, state[c.k][p], `${c.k}.${p} at ${c.at}`);
+ Object.assign(state[c.k], c.to);
+ }
+ // The end: the replaced lines gone, the rest shown.
+ assert.deepEqual([0, 1, 2, 3, 4, 5].map((i) => state[`l${i}.o`].autoAlpha), [1, 1, 0, 0, 1, 1]);
+});
+
+test("replace: one stacked row in the page, a hit per line, the tail on the last", () => {
+ const html = teaserHtml(ESTIMATES, RENDER);
+ assert.equal((html.match(/<div class="slot r-title">/g) ?? []).length, 1);
+ // The three estimates are in that row, in order.
+ const slot = html.slice(html.indexOf('<div class="slot'), html.indexOf('data-line="5"'));
+ assert.ok(slot.indexOf('data-line="2"') < slot.indexOf('data-line="3"') && slot.indexOf('data-line="3"') < slot.indexOf('data-line="4"'));
+ assert.ok(html.indexOf('data-line="5"') < html.indexOf('data-k="tail"'));
+ assert.match(html, /\.slot > \.line \{ grid-area: 1 \/ 1; margin: 0; \}/);
+ // A teaser without a stacked row carries none of the slot's style.
+ assert.ok(!teaserHtml(FERRET, RENDER).includes(".slot"));
+ // Every line sounds its own hit, the estimates as titles.
+ const hits = teaserHits(ESTIMATES).filter((h) => h.kind === "hit" && h.role !== "sub");
+ assert.deepEqual(hits.map((h) => h.role), ["overline", "title", "title", "title", "title", "kicker"]);
+});
+
+test("together: a line's second tier pops with it, on its one hit", () => {
+ const together = {
+ ...ESTIMATES,
+ lines: ESTIMATES.lines.map((l, i) => (i >= 2 && i <= 4 ? { ...l, together: true } : l)),
+ };
+ assert.deepEqual(validateTeaser(together), []);
+ const lines = teaserLines(together);
+ assert.deepEqual(lines.map((l) => l.together ?? false), [false, false, true, true, true, false]);
+ const m = teaserMotion(1);
+ const plain = teaserCues({ lines: teaserLines(ESTIMATES), tail: "?", seconds: 14, motion: m }).beats.lines;
+ const { beats } = teaserCues({ lines, tail: "?", seconds: 14, motion: m });
+ // The date pops at its line's pop, and the next line counts from there.
+ for (const i of [2, 3, 4]) assert.equal(beats.lines[i].subAt, beats.lines[i].at);
+ assert.equal(Math.round((beats.lines[3].at - beats.lines[2].at) * 1e4) / 1e4, m.gap);
+ // The title keeps its two-step: its tier a beat-fraction after it, as without.
+ assert.equal(beats.lines[1].subAt, plain[1].subAt);
+ // One hit per estimate: no lighter second-tier hit under them; the title keeps its.
+ const subs = teaserHits(together).filter((h) => h.role === "sub").map((h) => h.at);
+ assert.deepEqual(subs, [beats.lines[1].subAt]);
+});
+
+test("lead: a small tier above its line, popping with it on its one hit", () => {
+ const entry = {
+ ...ESTIMATES,
+ lines: [...ESTIMATES.lines.slice(0, 5), { text: "2027", replace: true, role: "title", lead: "Breaking Ground" }],
+ };
+ assert.deepEqual(validateTeaser(entry), []);
+ const lines = teaserLines(entry);
+ assert.equal(lines[5].lead, "Breaking Ground");
+ const html = teaserHtml(entry, RENDER);
+ // Above the year, in its line, before its pop; the tail still hangs off the year.
+ const at = html.indexOf('data-line="5"');
+ assert.ok(at < html.indexOf('data-k="l5.lead"') && html.indexOf('data-k="l5.lead"') < html.indexOf('data-k="l5"'));
+ assert.ok(html.indexOf('data-k="l5"') < html.indexOf('data-k="tail"'));
+ assert.ok(html.includes("Breaking Ground"));
+ // It drops in at the line's pop; no hit of its own.
+ const { cues, beats } = teaserCues({ lines, tail: "?", seconds: 16, motion: teaserMotion(1) });
+ assert.equal(cues.find((c) => c.why === "line 5 lead").at, beats.lines[5].at);
+ assert.equal(teaserHits(entry).filter((h) => h.kind === "hit").length, teaserHits(ESTIMATES).filter((h) => h.kind === "hit").length);
+ // A teaser without a lead carries none of its style.
+ assert.ok(!teaserHtml(ESTIMATES, RENDER).includes(".sub.lead"));
+});
diff --git a/umtool/report-to-video/chrome-threads.mjs b/umtool/report-to-video/chrome-threads.mjs
@@ -0,0 +1,218 @@
+// The THREAD RAIL's composition: one HyperFrames page for a whole cut, the
+// frame's left side beside the footage (deck.mjs threadsGeometry), a card per
+// thread (threads.mjs). The card of the thread on screen is lit and strung
+// across to the picture; its dots fill clip by clip; its outcome is stamped on
+// it as its last clip ends. The rail steps aside while a popup post has moved
+// the footage over it.
+//
+// PURE, like chrome-stamp.mjs: a schedule and a render block in, an HTML
+// string out. compose-chrome.mjs copies the assets in beside it, writes it and
+// renders it (`region: "threads"`); the build lays the frames over the cut as
+// it lays the deck's.
+//
+// The timeline is threads.mjs `threadCues`, each cue stating its from -- a
+// render is a seek per frame, from parallel workers, in any order.
+import { pageDuration, threadsGeometry } from "./deck.mjs";
+import { mix, rgba } from "./chrome-deck.mjs";
+import { threadCues } from "./threads.mjs";
+
+const esc = (s) =>
+ String(s ?? "")
+ .replace(/&/g, "&")
+ .replace(/</g, "<")
+ .replace(/>/g, ">")
+ .replace(/"/g, """);
+
+const r4 = (v) => Math.round(v * 10000) / 10000;
+
+/**
+ * The cards' sizes in the rail's height for `n` threads: each card `h` px tall
+ * (at most `max`), `gap` apart, the stack centred (`top`).
+ */
+export function railLayout(height, n, { gap = 14, max = 190 } = {}) {
+ const h = Math.min(max, Math.floor((height - (n - 1) * gap) / n));
+ const total = n * h + (n - 1) * gap;
+ return { h, gap, top: Math.max(0, Math.floor((height - total) / 2)) };
+}
+
+/**
+ * The rail's HTML, for the whole cut -- or a window of it (`from`/`duration`,
+ * as the deck's). `fonts` = `{ regular, bold }` asset-relative paths
+ * (DeckSans / DeckSansBold); `gsap` the vendored script. `?still=<t>` and the
+ * preview's `deck:seek` take CUT seconds.
+ */
+export function threadsHtml(schedule, render, opts = {}) {
+ const sched = schedule.threads;
+ if (!sched?.threads?.length) throw new Error("threads: the schedule has no thread rail");
+ const geo = threadsGeometry(render);
+ const pal = render.palette;
+ const W = geo.width, H = geo.height;
+ const C = geo.cards;
+ const fonts = opts.fonts ?? {};
+ const gsapSrc = opts.gsap ?? "assets/gsap.min.js";
+ const total = schedule.total;
+ const from = Number(opts.from ?? 0);
+ const dur = opts.duration != null ? r4(Number(opts.duration)) : r4(total - from);
+ if (!(dur > 0)) throw new Error(`threads: nothing to render from ${from}s of a ${total}s cut`);
+ const windowed = from > 0 || Math.abs(dur - total) > 1e-6;
+ const { init, cues } = threadCues(sched);
+ const n = sched.threads.length;
+ const lay = railLayout(C.height, n);
+ // Type scales with the card: a rail of eight is denser than a rail of three.
+ const k = Math.min(1, lay.h / 170);
+ const idxSize = Math.round(17 * Math.max(0.8, k));
+ const labelSize = Math.round(27 * Math.max(0.72, k));
+ const labelLine = Math.round(labelSize * 1.12);
+ const outSize = Math.round(16 * Math.max(0.8, k));
+ const dot = Math.round(11 * Math.max(0.8, k));
+ const pad = Math.round(16 * Math.max(0.75, k));
+ const rail = 5;
+
+ const cardsHtml = sched.threads
+ .map((t, j) => {
+ const top = lay.top + j * (lay.h + lay.gap);
+ const oc = t.outcome?.color ?? pal.accent;
+ const dots = t.clips.map((_, d) => `<i class="dot" data-k="d${j}-${d}"></i>`).join("");
+ return (
+ `<div class="card" data-thread="${esc(t.id)}" data-k="c${j}" style="top:${top}px; --oc:${esc(oc)}; ` +
+ `--og:${rgba(oc, 0.6)}; --ob:${rgba(mix(pal.bg, oc, 0.14), 0.92)}">` +
+ `<div class="wash" data-k="w${j}"></div>` +
+ `<div class="head"><span class="idx">${String(j + 1).padStart(2, "0")}</span><span class="dots">${dots}</span></div>` +
+ `<div class="label">${esc(t.label)}</div>` +
+ (t.outcome
+ ? `<div class="outcome" data-k="o${j}"><span class="word">${esc(t.outcome.label)}</span>` +
+ `<span class="oflash" data-k="f${j}"></span></div>`
+ : "") +
+ `</div>` +
+ `<div class="string" data-k="s${j}" style="top:${top + Math.round(lay.h / 2) - 1}px"></div>`
+ );
+ })
+ .join("\n ");
+
+ const data = {
+ total: r4(total),
+ window: windowed ? { from: r4(from), dur } : null,
+ labelSize,
+ ids: sched.threads.map((t) => t.id),
+ init,
+ cues: cues.map(({ why, ...c }) => c),
+ };
+ const json = JSON.stringify(data).replace(/</g, "\\u003c");
+ const fps = schedule.fps ?? render.fps ?? 30;
+ const cardBg = mix(mix(pal.bg, pal.fg, 0.07), pal.accent, 0.05);
+ const litBg = mix(mix(pal.bg, pal.fg, 0.12), pal.accent, 0.14);
+
+ return `<!doctype html>
+<html lang="en">
+ <head>
+ <meta charset="UTF-8" />
+ <meta name="viewport" content="width=${W}, height=${H}" />
+ <script src="${esc(gsapSrc)}"></script>
+ <style>
+ /* The deck's faces, under the deck's private names -- see docs/quirks.md. */
+ @font-face { font-family: 'DeckSans'; font-weight: 400; font-style: normal;
+ src: url('${esc(fonts.regular ?? "")}'); }
+ @font-face { font-family: 'DeckSansBold'; font-weight: 400; font-style: normal;
+ src: url('${esc(fonts.bold ?? "")}'); }
+ * { margin: 0; padding: 0; box-sizing: border-box; }
+ html, body { width: ${W}px; height: ${H}px; overflow: hidden; background: transparent; }
+ body { font-family: 'DeckSans', sans-serif; font-synthesis: none; color: ${pal.fg};
+ -webkit-font-smoothing: antialiased; text-rendering: geometricPrecision; }
+ #root { position: relative; width: ${W}px; height: ${H}px; overflow: hidden; }
+ #threads-clip { position: absolute; left: 0; top: 0; width: ${W}px; height: ${H}px; }
+ .rail { position: absolute; left: 0; top: 0; width: ${W}px; height: ${H}px; }
+ /* A card: its number, a dot per clip, its label; the outcome stamped at
+ its foot. Ahead of its thread it is dim; lit while it plays; after,
+ it keeps its outcome at a quieter rest. */
+ .card { position: absolute; left: ${C.x}px; width: ${C.width}px; height: ${lay.h}px; opacity: 0;
+ border-radius: 10px; overflow: hidden; background: ${cardBg};
+ border-left: ${rail}px solid ${pal.accent};
+ box-shadow: inset 0 0 0 1px ${rgba(pal.fg, 0.1)}; padding: ${pad - 2}px ${pad}px ${pad}px ${pad}px; }
+ .wash { position: absolute; inset: 0; opacity: 0; pointer-events: none; background: ${litBg};
+ box-shadow: inset 0 0 0 2px ${rgba(pal.accent, 0.9)}, inset 0 0 26px ${rgba(pal.accent, 0.35)}; }
+ .head { position: relative; display: flex; align-items: center; justify-content: space-between; gap: 8px;
+ height: ${idxSize + 8}px; }
+ .idx { font-family: 'DeckSansBold', sans-serif; font-size: ${idxSize}px; line-height: ${idxSize + 8}px;
+ letter-spacing: 0.12em; color: ${pal.accent}; font-variant-numeric: tabular-nums; }
+ .dots { display: flex; flex-wrap: wrap; justify-content: flex-end; gap: ${Math.round(dot * 0.55)}px; }
+ .dot { display: block; width: ${dot}px; height: ${dot}px; border-radius: 50%; background: ${pal.fg};
+ box-shadow: 0 0 8px ${rgba(pal.fg, 0.45)}; }
+ .label { position: relative; margin-top: ${Math.round(pad * 0.45)}px; font-family: 'DeckSansBold', sans-serif;
+ font-size: ${labelSize}px; line-height: ${labelLine}px; letter-spacing: -0.005em; color: ${pal.fg};
+ display: -webkit-box; -webkit-box-orient: vertical; -webkit-line-clamp: 2; overflow: hidden; }
+ /* The outcome: a small rubber stamp in its verdict's colour. */
+ .outcome { position: absolute; left: ${pad}px; bottom: ${pad - 4}px; visibility: hidden; opacity: 0;
+ max-width: ${C.width - 2 * pad - rail}px; padding: 3px 10px 4px; border-radius: 5px;
+ background: var(--ob); border: 2px solid var(--oc); transform-origin: 30% 50%; rotate: -4deg; }
+ .outcome .word { display: block; white-space: nowrap; overflow: hidden; text-overflow: ellipsis;
+ font-family: 'DeckSansBold', sans-serif; font-size: ${outSize}px; line-height: ${outSize + 6}px;
+ letter-spacing: 0.08em; text-transform: uppercase; color: var(--oc); }
+ .oflash { position: absolute; inset: -3px; border-radius: 6px; opacity: 0; pointer-events: none;
+ box-shadow: 0 0 0 2px var(--oc), 0 0 22px 6px var(--og); }
+ /* The string: from the lit card across to the picture's edge. */
+ .string { position: absolute; left: ${C.x + C.width}px; width: ${W - C.x - C.width}px; height: 3px;
+ border-radius: 2px; transform-origin: 0% 50%; transform: scaleX(0);
+ background: linear-gradient(90deg, ${pal.accent} 0%, ${rgba(pal.accent, 0.35)} 100%);
+ box-shadow: 0 0 10px ${rgba(pal.accent, 0.7)}; }
+ </style>
+ </head>
+ <body>
+ <div id="root" data-composition-id="threads" data-start="0" data-duration="${pageDuration(dur, fps)}"
+ data-width="${W}" data-height="${H}">
+ <div id="threads-clip" class="clip" data-start="0" data-duration="${pageDuration(dur, fps)}" data-track-index="1">
+ <div class="rail" data-k="rail">
+ ${cardsHtml}
+ </div>
+ </div>
+ </div>
+
+ <script id="threads-data" type="application/json">${json}</script>
+ <script>
+ const S = JSON.parse(document.getElementById("threads-data").textContent);
+ const byK = {};
+ for (const el of document.querySelectorAll("[data-k]")) byK[el.dataset.k] = el;
+
+ for (const k of Object.keys(S.init)) if (byK[k]) gsap.set(byK[k], S.init[k]);
+ const inner = gsap.timeline({ paused: true });
+ for (const c of S.cues) {
+ const el = byK[c.k];
+ if (!el) continue;
+ inner.fromTo(el, c.from, { ...c.to, duration: c.dur, ease: c.ease, immediateRender: false }, c.at);
+ }
+ inner.set({}, {}, S.total);
+ let tl = inner;
+ if (S.window) {
+ tl = gsap.timeline({ paused: true });
+ tl.add(inner.tweenFromTo(S.window.from, S.window.from + S.window.dur, { duration: S.window.dur, ease: "none" }), 0);
+ }
+ window.__timelines = window.__timelines || {};
+ window.__timelines["threads"] = tl;
+
+ const ready = document.fonts.load(S.labelSize + "px DeckSansBold").catch(() => {}).then(() => {
+ document.documentElement.dataset.fit = "1";
+ });
+
+ const params = new URLSearchParams(location.search);
+ const local = (t) => {
+ const v = Number(t) || 0;
+ return S.window ? Math.max(0, Math.min(S.window.dur, v - S.window.from)) : Math.max(0, v);
+ };
+ const still = params.get("still");
+ if (still !== null) {
+ tl.seek(local(still), false);
+ ready.then(() => tl.seek(local(still), false));
+ }
+ if (params.get("preview") === "1") {
+ window.addEventListener("message", (e) => {
+ const m = e.data || {};
+ if (m.type === "deck:seek") tl.seek(local(m.t), false);
+ });
+ ready.then(() => {
+ if (window.parent !== window) window.parent.postMessage({ type: "threads:ready", total: S.total, ids: S.ids }, "*");
+ });
+ }
+ </script>
+ </body>
+</html>
+`;
+}
diff --git a/umtool/report-to-video/compose-chrome.mjs b/umtool/report-to-video/compose-chrome.mjs
@@ -47,6 +47,8 @@ import { deckHtml, GSAP_FILE } from "./chrome-deck.mjs";
import { postsHtml, snapWindow, windowPosts } from "./chrome-posts.mjs";
import { feedHtml } from "./chrome-feed.mjs";
import { stampHtml } from "./chrome-stamp.mjs";
+import { threadsHtml } from "./chrome-threads.mjs";
+import { flipsHtml } from "./chrome-flips.mjs";
const run = promisify(execFile);
@@ -692,6 +694,16 @@ async function regionHtml(region, { manifest, manifestDir, base, projDir, assets
const fonts = await copyFonts(manifest.render, assetsDir, { regular: "DeckSans", bold: "DeckSansBold" }, { strict: true });
return stampHtml(schedule, manifest.render, { fonts, from, duration });
}
+ if (region === "threads") {
+ // The thread rail (chrome-threads.mjs): the deck's faces, no QR.
+ const fonts = await copyFonts(manifest.render, assetsDir, { regular: "DeckSans", bold: "DeckSansBold" }, { strict: true });
+ return threadsHtml(schedule, manifest.render, { fonts, from, duration });
+ }
+ if (region === "flips") {
+ // The flips panel (chrome-flips.mjs): the deck's faces, no QR.
+ const fonts = await copyFonts(manifest.render, assetsDir, { regular: "DeckSans", bold: "DeckSansBold" }, { strict: true });
+ return flipsHtml(schedule, manifest.render, { fonts, from, duration });
+ }
if (region === "teaser") {
// A page module reached by a dynamic import, so nothing that imports this
// file -- umtool's preview helper, the build -- loads its face's URL
@@ -722,9 +734,21 @@ async function framesOnDisk(dir) {
}
/** Run the renderer, its chatter to stderr (stdout may be a build's NDJSON). */
+/**
+ * The renderer's environment: ours without DISPLAY and WAYLAND_DISPLAY. It is
+ * headless and needs no display, and a stale one breaks it: with DISPLAY naming
+ * an X server that has gone (Xwayland killed by the OOM killer, say), ANGLE's
+ * SwiftShader tries to connect to it, fails, and every render dies with
+ * "assertSwiftShader ... vendor=''".
+ */
+export function rendererEnv(env = process.env) {
+ const { DISPLAY: _d, WAYLAND_DISPLAY: _w, ...rest } = env;
+ return rest;
+}
+
function runRenderer(cmd, args) {
return new Promise((resolve, reject) => {
- const child = spawn(cmd, args, { stdio: ["ignore", "pipe", "pipe"] });
+ const child = spawn(cmd, args, { stdio: ["ignore", "pipe", "pipe"], env: rendererEnv() });
let tail = "";
const keep = (b) => {
process.stderr.write(b);
@@ -772,6 +796,10 @@ function runRenderer(cmd, args) {
* `chrome/stamp-frames/` and their `.key`, a window `stamp-from<s>[-frames]`,
* cached as the deck's are.
*
+ * Threads (`region: "threads"`, a schedule with `threads`): the thread rail
+ * for the whole cut, exactly as the stamps -- project `chrome/threads/`,
+ * frames `chrome/threads-frames/`, windowed and cached alike.
+ *
* Teaser (`region: "teaser"`, `segment` = the entry's id): the whole frame,
* `seconds` long, drawn from the entry alone (no schedule):
* - project `chrome/teaser-<id>/` (`teaser-preview-<id>/` when `preview`),
@@ -809,9 +837,10 @@ export async function composeChrome({
// The regions keyed by the render cache: the two drawn from the deck's
// schedule, and a teaser, drawn from its own timeline entry.
- const keyed = region === "deck" || region === "feed" || region === "posts" || region === "teaser" || region === "stamp";
+ const keyed = region === "deck" || region === "feed" || region === "posts" || region === "teaser" || region === "stamp" ||
+ region === "threads" || region === "flips";
// The regions drawn over the whole cut from its schedule, windowable alike.
- const wholeCut = region === "deck" || region === "feed" || region === "stamp";
+ const wholeCut = region === "deck" || region === "feed" || region === "stamp" || region === "threads" || region === "flips";
let teaser = null;
if (region === "teaser") {
teaser = (manifest.timeline ?? []).find((e) => e.id === segment && e.type === "teaser") ?? null;
@@ -832,6 +861,12 @@ export async function composeChrome({
if (region === "feed" && sched.layout !== "feed") {
throw new Error("the feed region needs a feed's schedule (layout \"feed\": posts.layout \"feed\" and posts to draw)");
}
+ if (region === "flips" && !sched.flips?.pairs?.length) {
+ throw new Error("the flips region needs a schedule with flip pairs (render.chrome.flips)");
+ }
+ if (region === "threads" && !sched.threads?.threads?.length) {
+ throw new Error("the threads region needs a schedule with a thread rail (render.chrome.threads)");
+ }
if (region === "stamp" && !sched.factcheck?.stamps?.length) {
throw new Error("the stamp region needs a schedule that stamps a claim (an entry with a `claim`)");
}
@@ -905,7 +940,7 @@ export async function composeChrome({
"--virtual-time-budget=6000",
`--screenshot=${out}`,
`file://${path.join(projDir, "index.html")}?still=${Number(still)}`,
- ], { maxBuffer: 1 << 26 });
+ ], { maxBuffer: 1 << 26, env: rendererEnv() });
return { ...result, still: out };
}
@@ -960,7 +995,7 @@ if (import.meta.url === `file://${process.argv[1]}`) {
const manifestPath = argv.find((a, i) => !a.startsWith("--") && !VALUED.has(argv[i - 1]));
if (!manifestPath) {
console.error(
- "usage: compose-chrome.mjs <manifest.json> [--region chart|deck|feed|stamp|posts|teaser] [--variant sourced|full]\n" +
+ "usage: compose-chrome.mjs <manifest.json> [--region chart|deck|feed|stamp|threads|flips|posts|teaser] [--variant sourced|full]\n" +
" [--segment <id>] (posts: the clip whose window to compose; teaser: its entry)\n" +
" [--from <s>] [--duration <s>] [--out <dir>] [--preview]\n" +
" [--still <s> --png <path>]\n" +
@@ -983,7 +1018,7 @@ if (import.meta.url === `file://${process.argv[1]}`) {
still: num("--still"),
png: flag("--png"),
// The deck renders four-wide by default; the band keeps the renderer's own default.
- workers: num("--workers") ?? (region === "deck" || region === "feed" || region === "teaser" ? 4 : region === "posts" || region === "stamp" ? 2 : null),
+ workers: num("--workers") ?? (region === "deck" || region === "feed" || region === "teaser" ? 4 : region === "posts" || region === "stamp" || region === "threads" || region === "flips" ? 2 : null),
quality: flag("--quality") ?? "high",
format: flag("--format") ?? "png-sequence",
fps: num("--fps"),
diff --git a/umtool/report-to-video/deck.mjs b/umtool/report-to-video/deck.mjs
@@ -20,6 +20,11 @@ import { attributionParts, deckSubtitle } from "./attribution.mjs";
import {
claimOf, originalUrlAt, resolveFactcheck, roundStamps, stampSchedule, validateFactcheck,
} from "./factcheck.mjs";
+import { threadOf, threadSchedule, threadsOn, validateThreads } from "./threads.mjs";
+import { flipSchedule, flipsOn, validateFlips } from "./flips.mjs";
+
+/** Is the left panel on: the thread rail (threads.mjs) or the flips panel (flips.mjs) -- one region, one or the other. */
+export const railOn = (render) => threadsOn(render) || flipsOn(render);
/** The renderer version, pinned. It is part of the cache key: a new renderer is new frames. */
export const HYPERFRAMES_PKG_DEFAULT = "hyperframes@0.8.24";
@@ -132,7 +137,7 @@ export function validateChrome(chrome, render = {}) {
const errors = [];
if (chrome === undefined || chrome === null) return errors;
if (!isObj(chrome)) return ["render.chrome must be an object"];
- unknownKeys(chrome, ["engine", "layout", "deck", "factcheck"], "render.chrome", errors);
+ unknownKeys(chrome, ["engine", "layout", "deck", "factcheck", "threads", "flips"], "render.chrome", errors);
if (chrome.engine !== "hyperframes") errors.push('render.chrome.engine must be "hyperframes"');
if (chrome.layout !== "deck") errors.push('render.chrome.layout must be "deck"');
if (render.rail) errors.push("render.chrome (the deck) and render.rail cannot both be set");
@@ -144,6 +149,11 @@ export function validateChrome(chrome, render = {}) {
errors.push(...validateEndFade(render));
// The fact-check stamps and tally (factcheck.mjs): drawn only by the deck.
errors.push(...validateFactcheck(chrome.factcheck));
+ // The thread rail (threads.mjs): the deck's layout only, beside the footage.
+ errors.push(...validateThreads(chrome.threads));
+ // The flips panel (flips.mjs): the same region as the rail, so not both.
+ errors.push(...validateFlips(chrome.flips));
+ if (threadsOn({ chrome }) && flipsOn({ chrome })) errors.push("render.chrome.threads and render.chrome.flips share the left panel: set one");
const d = chrome.deck ?? {};
if (!isObj(d)) return [...errors, "render.chrome.deck must be an object"];
const w = "render.chrome.deck";
@@ -266,6 +276,14 @@ export function validateChrome(chrome, render = {}) {
// sees the posts.
if (d.posts !== undefined && resolveDeck({ chrome }).posts.show) errors.push(...postsFitErrors({ ...render, chrome }));
const room = g.H - g.deck.height;
+ if (railOn({ chrome })) {
+ const which = threadsOn({ chrome }) ? "render.chrome.threads" : "render.chrome.flips";
+ if (resolveDeck({ chrome }).posts.layout === "feed") errors.push(`${which} needs the popup posts, not the feed`);
+ const rail = threadsGeometry({ ...render, chrome });
+ if (rail.cards.width < 220) {
+ errors.push(`${which} leaves the rail ${rail.cards.width}px wide (at least 220: a smaller footageScale)`);
+ }
+ }
if (g.footage.height > room) {
const max = Math.floor((room / g.H) * 1000) / 1000;
errors.push(
@@ -321,7 +339,9 @@ export const even = (v) => Math.round(v / 2) * 2;
*
* The footage box keeps the FRAME's aspect, is `footageScale` of its width
* (evened), and is centred in the area above the deck. For 1920×1080 at 0.82
- * with a 190 px deck that is 1574×886 at (173, 2).
+ * with a 190 px deck that is 1574×886 at (173, 2). With the thread rail on
+ * (threads.mjs) the box moves to the frame's right edge, `railGap` in, and the
+ * rail takes the left: 1574×886 at (322, 2).
*
* @returns {{ W:number, H:number,
* footage:{x:number,y:number,width:number,height:number},
@@ -334,13 +354,33 @@ export function deckGeometry(render) {
const dh = deck.height;
const fw = even(W * deck.footageScale);
const fh = even((fw * H) / W);
+ const fx = railOn(render) ? W - fw - railGap(render) : Math.floor((W - fw) / 2);
return {
W, H,
- footage: { x: Math.floor((W - fw) / 2), y: Math.floor((H - dh - fh) / 2), width: fw, height: fh },
+ footage: { x: fx, y: Math.floor((H - dh - fh) / 2), width: fw, height: fh },
deck: { x: 0, y: H - dh, width: W, height: dh },
};
}
+/** The air between the frame's edges, the thread rail and the footage. */
+export const railGap = (render) => Math.round(24 * ((render?.width ?? 1920) / 1920));
+
+/**
+ * The thread rail's region (threads.mjs, chrome-threads.mjs): the frame's left
+ * side beside the footage, as tall as the footage box, from the frame's edge
+ * to the footage's -- so a lit card's string can reach the picture. `cards`
+ * is where the cards sit inside it (region-local), `railGap` in from its left
+ * and short of the footage.
+ */
+export function threadsGeometry(render) {
+ const { footage } = deckGeometry(render);
+ const gap = railGap(render);
+ return {
+ x: 0, y: footage.y, width: footage.x, height: footage.height,
+ cards: { x: gap, y: 0, width: footage.x - 2 * gap, height: footage.height },
+ };
+}
+
/**
* Where things sit INSIDE the deck region (region-local pixels). The
* composition draws from this; umtool's preview frames the same rect. A
@@ -607,6 +647,10 @@ export function deckSchedule({
const feed = placed.length > 0 && deck.posts.layout === "feed";
const moves = placed.length && !feed ? footageMoves({ posts: placed, segments: segs, render }) : [];
const stamps = stampSchedule({ segments: segs.map((s, i) => ({ ...s, claim: claimOf(entries[i]) })), D, total, render });
+ const threads = threadsOn(render)
+ ? threadSchedule({ segments: segs.map((s, i) => ({ ...s, thread: threadOf(entries[i]) })), D, total, render, moves })
+ : null;
+ const flips = flipsOn(render) ? flipSchedule({ segments: segs, total, render, moves }) : null;
return {
version: 1,
kind: "deck",
@@ -638,6 +682,10 @@ export function deckSchedule({
}),
// The fact-check stamps: present only when a claim is stamped.
...(stamps.length ? { factcheck: { stamps: roundStamps(stamps) } } : {}),
+ // The thread rail: present only when it is on.
+ ...(threads ? { threads } : {}),
+ // The flips panel: present only when it is on.
+ ...(flips ? { flips } : {}),
// Present only when there are posts to draw, so a cut without them writes
// the schedule it always did.
...(placed.length ? { posts: roundPosts(placed) } : {}),
@@ -1394,32 +1442,66 @@ export function muteSegmentSeconds({ entry, record = null, render = {}, seconds
* spill at these counts.
*/
export const TEASER_LIMITS = Object.freeze({
- lines: [1, 5], seconds: [3, 20], chars: 80, tail: 8,
+ lines: [1, 8], rows: 5, seconds: [3, 20], chars: 80, tail: 8, hold: 4,
fit: Object.freeze({ overline: 64, title: 34, kicker: 56, sub: 66 }),
});
-const LINE_KEYS = ["text", "break"];
+const LINE_KEYS = ["text", "break", "replace", "role", "hold", "together", "lead"];
+
+/** The roles a line may be drawn in; a line's `role` names one of them. */
+export const TEASER_ROLES = Object.freeze(["overline", "title", "kicker"]);
/**
- * A teaser's lines, normalised: `{ text, head, sub, role }` each, `head` the
- * part drawn on the first tier and `sub` the second tier (`break`) or null.
+ * A teaser's lines, normalised: `{ text, head, sub, role, row }` each, `head`
+ * the part drawn on the first tier and `sub` the second tier (`break`) or null.
* Trims; assumes `validateTeaser` passed.
*
- * @returns {Array<{ text: string, head: string, sub: string|null, role: "overline"|"title"|"kicker" }>}
+ * ROWS, NOT LINES, HAVE POSITIONS. A line with `replace: true` takes the row
+ * of the line before it -- it slams in where that one was and pushes it out
+ * (`teaserTimes` gives the pushed line its `outAt`) -- so a row can say
+ * several things in turn: an estimate, then the next one. `row` is the line's
+ * row, and roles follow the ROW's position as they always followed the line's
+ * (with three or more rows the first is the overline, the last the kicker,
+ * the rest titles). A line's own `role` overrides that; a replacing line
+ * without one takes the role of the line it replaces. `replace` and `hold`
+ * (seconds held after the line before the next one pops) are present only
+ * when set, so an entry that uses neither normalises exactly as it did. So is
+ * `together`: a line with a `break` whose second tier pops WITH it, on its one
+ * hit, rather than a beat-fraction later on a lighter hit of its own. And
+ * `lead`: a small wide-tracked tier ABOVE the line (a label: "Breaking
+ * ground" over "2027"), drawn as a second tier is and popping with its line.
+ *
+ * @returns {Array<{ text: string, head: string, sub: string|null, role: "overline"|"title"|"kicker",
+ * row: number, replace?: true, hold?: number, together?: true, lead?: string }>}
*/
export function teaserLines(entry) {
const lines = Array.isArray(entry?.lines) ? entry.lines : [];
- const n = lines.length;
- return lines.map((l, i) => {
+ let row = -1;
+ const rowOf = lines.map((l, i) => (i > 0 && isObj(l) && l.replace === true ? row : (row += 1)));
+ const n = row + 1;
+ const out = [];
+ lines.forEach((l, i) => {
const text = String(isObj(l) ? l.text ?? "" : l ?? "").trim();
const brk = isObj(l) && typeof l.break === "string" ? l.break.trim() : "";
const sub = brk && text.endsWith(brk) && text.length > brk.length ? brk : null;
const head = sub ? text.slice(0, text.length - sub.length).trim() : text;
- const role = n >= 3 ? (i === 0 ? "overline" : i === n - 1 ? "kicker" : "title")
- : n === 2 ? (i === 0 ? "overline" : "title")
+ const r = rowOf[i];
+ const replace = i > 0 && isObj(l) && l.replace === true;
+ const byRow = n >= 3 ? (r === 0 ? "overline" : r === n - 1 ? "kicker" : "title")
+ : n === 2 ? (r === 0 ? "overline" : "title")
: "title";
- return { text, head, sub, role };
+ const own = isObj(l) && TEASER_ROLES.includes(l.role) ? l.role : null;
+ const role = own ?? (replace ? out[i - 1].role : byRow);
+ const hold = isObj(l) && typeof l.hold === "number" && l.hold > 0 ? l.hold : null;
+ const together = !!sub && isObj(l) && l.together === true;
+ const lead = isObj(l) && typeof l.lead === "string" && l.lead.trim() ? l.lead.trim() : null;
+ out.push({
+ text, head, sub, role, row: r,
+ ...(replace ? { replace: true } : {}), ...(hold ? { hold } : {}), ...(together ? { together: true } : {}),
+ ...(lead ? { lead } : {}),
+ });
});
+ return out;
}
/** The teaser's tail, trimmed, or "" for none. */
@@ -1451,6 +1533,10 @@ export function teaserTitle(entry) {
export const TEASER_MOTION = Object.freeze({
first: 0.55, gap: 0.7, sub: 0.3, slam: 1.42, under: 0.968, hit: 0.2, settle: 0.5,
blur: 18, tailAfter: 0.8, tailDur: 1.7, endRoom: 1.2, push: 1.065, grainHz: 12,
+ // A line pushed out of its row by the next (`replace`): over `out` seconds
+ // from that one's pop it rises `outRise` px, shrinks to `outScale` and blurs
+ // to `outBlur` px as it fades -- gone before the new line's slam lands.
+ out: 0.3, outRise: 64, outScale: 0.9, outBlur: 12,
});
/** The beats an entry's `beat` may be: from 0.4 s (packed) to 2.5 s (a pause between each). */
@@ -1471,7 +1557,8 @@ export function teaserTailWait(entry) {
if (entry?.tailWait !== undefined && entry?.tailWait !== null) return Number(entry.tailWait);
const m = teaserMotion(entry?.beat);
const lines = teaserLines(entry);
- const lastSub = !!lines[lines.length - 1]?.sub;
+ const last = lines[lines.length - 1];
+ const lastSub = !!last?.sub && !last.together;
return Math.round((m.tailAfter - (lastSub ? 0 : m.hit)) * 10000) / 10000;
}
@@ -1514,18 +1601,26 @@ export function teaserTimes(lines, tail, m = TEASER_MOTION) {
let t = m.first;
const raw = [];
lines.forEach((l, i) => {
- if (i > 0) t += m.gap;
+ if (i > 0) t += m.gap + (lines[i - 1].hold ?? 0);
const at = t;
- const subAt = l.sub ? at + m.sub : null;
+ // `together`: the second tier pops with its line (one pop, one hit).
+ const subAt = l.sub ? at + (l.together ? 0 : m.sub) : null;
if (subAt != null) t = subAt;
raw.push({ at, subAt });
});
+ // A line held after its pop holds the tail back too.
+ t += lines[lines.length - 1]?.hold ?? 0;
const tailRaw = tail ? t + m.tailAfter : null;
const endRaw = tailRaw != null ? tailRaw + m.tailDur : t + m.hit + m.settle;
const r = (v) => Math.round(v * 10000) / 10000;
const T = r;
return {
- lines: raw.map((b) => ({ at: T(b.at), impact: r(T(b.at) + T(m.hit)), subAt: b.subAt == null ? null : T(b.subAt) })),
+ // `outAt`: a line the next one replaces is pushed out as that one pops.
+ // Only on such a line, so a teaser without `replace` times as it did.
+ lines: raw.map((b, i) => ({
+ at: T(b.at), impact: r(T(b.at) + T(m.hit)), subAt: b.subAt == null ? null : T(b.subAt),
+ ...(lines[i + 1]?.replace ? { outAt: T(raw[i + 1].at) } : {}),
+ })),
tailAt: tailRaw == null ? null : T(tailRaw),
tailDur: T(m.tailDur),
end: T(endRaw),
@@ -1649,7 +1744,8 @@ function dippedMotion(entry, lead) {
const r = (v) => Math.round(v * 10000) / 10000;
if (entry?.tailWait !== undefined && entry?.tailWait !== null && teaserTail(entry)) {
const lines = teaserLines(entry);
- const lastSub = !!lines[lines.length - 1]?.sub;
+ const last = lines[lines.length - 1];
+ const lastSub = !!last?.sub && !last.together;
m = Object.freeze({ ...m, tailAfter: r(Number(entry.tailWait) + (lastSub ? 0 : m.hit)) });
}
if (!dipOf(entry)) return m;
@@ -1717,7 +1813,9 @@ export function teaserHits(entry, D = 0.5, fps = 30) {
};
lines.forEach((l, i) => {
out.push({ kind: "hit", at: times.lines[i].impact, role: l.role, ...HIT[l.role] });
- if (l.sub && times.lines[i].subAt != null) out.push({ kind: "hit", at: times.lines[i].subAt, role: "sub", ...HIT.sub });
+ if (l.sub && !l.together && times.lines[i].subAt != null) {
+ out.push({ kind: "hit", at: times.lines[i].subAt, role: "sub", ...HIT.sub });
+ }
});
if (tail && times.tailAt != null) {
out.push({ kind: "swell", at: times.tailAt, role: "tail", gain: 0.34, decay: 0.7, f0: 46, f1: 62, dur: times.tailDur });
@@ -1763,11 +1861,39 @@ export function validateTeaser(entry, where = `timeline entry ${entry?.id ?? "?"
if (!Array.isArray(lines) || lines.length < lo || lines.length > hi) {
errors.push(`${where}.lines must be a list of ${lo} to ${hi} lines`);
} else {
+ // What the frame holds is ROWS: a replacing line shares the row before it.
+ const rows = lines.filter((l, i) => !(i > 0 && isObj(l) && l.replace === true)).length;
+ if (rows > TEASER_LIMITS.rows) {
+ errors.push(
+ `${where}.lines make ${rows} rows, and at most ${TEASER_LIMITS.rows} fit the frame ` +
+ `-- set "replace": true on a line to have it take the row of the line before it`,
+ );
+ }
lines.forEach((l, i) => {
const w = `${where}.lines[${i}]`;
if (typeof l === "string") { oneLine(l, w); return; }
- if (!isObj(l)) { errors.push(`${w} must be a string or { text, break }`); return; }
+ if (!isObj(l)) { errors.push(`${w} must be a string or { text, break, replace, role, hold, together, lead }`); return; }
for (const k of Object.keys(l)) if (!LINE_KEYS.includes(k)) errors.push(`${w}.${k} is not a teaser line field`);
+ if (l.replace !== undefined && l.replace !== null) {
+ if (typeof l.replace !== "boolean") errors.push(`${w}.replace must be true or false`);
+ else if (l.replace && i === 0) errors.push(`${w}.replace: the first line has no line before it to replace`);
+ }
+ if (l.role !== undefined && l.role !== null && !TEASER_ROLES.includes(l.role)) {
+ errors.push(`${w}.role must be one of ${TEASER_ROLES.join(", ")}`);
+ }
+ if (l.together !== undefined && l.together !== null) {
+ if (typeof l.together !== "boolean") errors.push(`${w}.together must be true or false`);
+ else if (l.together && !(typeof l.break === "string" && l.break.trim())) {
+ errors.push(`${w}.together pops the second tier with its line, and it has no break`);
+ }
+ }
+ if (l.lead !== undefined && l.lead !== null && oneLine(l.lead, `${w}.lead`)
+ && l.lead.trim().length > TEASER_LIMITS.fit.sub) {
+ errors.push(`${w}.lead is ${l.lead.trim().length} characters, and at most ${TEASER_LIMITS.fit.sub} fit the frame as a small tier`);
+ }
+ if (l.hold !== undefined && l.hold !== null && !numIn(l.hold, 0, TEASER_LIMITS.hold)) {
+ errors.push(`${w}.hold must be from 0 to ${TEASER_LIMITS.hold} seconds`);
+ }
if (!oneLine(l.text, `${w}.text`)) return;
if (l.break === undefined || l.break === null) return;
if (!oneLine(l.break, `${w}.break`)) return;
diff --git a/umtool/report-to-video/fetch-via-editor.mjs b/umtool/report-to-video/fetch-via-editor.mjs
@@ -56,8 +56,13 @@ function die(message) {
const manifestPath = argv.find((a) => !a.startsWith("-") && a.endsWith(".json"));
const clipId = flag("--fetch-only") ?? flag("--clip");
-if (!manifestPath || !clipId) {
- die("usage: fetch-via-editor.mjs <manifest.json> --fetch-only <clipId> [--full] [--max-height N] [--pad-before N] [--pad-after N] [--progress ndjson]");
+// THE WHOLE TIMELINE AS ONE ASK: every clip entry's window, posted as one
+// `fetch-windows` call. The editor answers the windows already on disk at once
+// and fetches the rest as one paced job per platform — the pacing, the 403
+// streak and the 429 stop live there, not in a shell loop here.
+const wantAll = argv.includes("--all");
+if (!manifestPath || (!clipId && !wantAll) || (clipId && wantAll)) {
+ die("usage: fetch-via-editor.mjs <manifest.json> (--fetch-only <clipId> [--full] | --all) [--max-height N] [--pad N] [--pad-before N] [--pad-after N] [--progress ndjson]");
}
// THE WHOLE RECORDING INSTEAD OF A WINDOW. For a clip whose windows would tile
@@ -102,6 +107,111 @@ if (!token) {
const whole = JSON.parse(await readFile(manifestPath, "utf8"));
const provenance = whole.provenance ?? {};
+const pad = Number(flag("--pad") ?? 3);
+const padBefore = Number(flag("--pad-before") ?? pad);
+const padAfter = Number(flag("--pad-after") ?? pad);
+// TWO DECIMALS, matching the editor's own naming (common/lib/clipWindow.ts) and
+// the build's. The name IS the window, so a request that rounds differently
+// addresses a different file and the cache misses forever.
+const windowOf = (e) => ({
+ from: Number(Math.max(0, Number(e.start) - padBefore).toFixed(2)),
+ to: Number((Number(e.end) + padAfter).toFixed(2)),
+});
+const manifestId = provenance.manifestId ?? path.basename(path.dirname(path.resolve(manifestPath)));
+
+const headers = {
+ authorization: `Bearer ${token}`,
+ "content-type": "application/json",
+};
+
+async function ask(url, init) {
+ try {
+ return await fetch(url, init);
+ } catch (err) {
+ die(`could not reach the editor at ${editorUrl}: ${err.message}`);
+ }
+}
+
+if (wantAll) {
+ if (argv.includes("--full")) die("--all fetches windows; --full is one clip's whole recording");
+ const clips = (whole.timeline ?? []).filter((e) => e.type === "clip");
+ const items = [];
+ for (const e of clips) {
+ const slug = e.channel ?? provenance.channelSlug;
+ if (!slug || !e.video) die(`${e.id} has no channel or video to fetch`);
+ items.push({
+ slug,
+ id: e.video,
+ ...windowOf(e),
+ clipId: e.id,
+ pad: Math.max(padBefore, padAfter),
+ ...(e.webpageUrl ? { webpageUrl: e.webpageUrl } : {}),
+ reason: String(e.note ?? e.quote ?? `clip window with ${padBefore}s before / ${padAfter}s after`).slice(0, 400),
+ });
+ }
+ if (items.length === 0) {
+ EMIT("note", { message: "no clip entries in the timeline — nothing to fetch" });
+ EMIT("done", { out: null, nothingToFetch: true });
+ process.exit(0);
+ }
+ const res = await ask(`${editorUrl}/api/ops/fetch-windows`, {
+ method: "POST",
+ headers,
+ body: JSON.stringify({
+ items,
+ requestedBy: "umtool",
+ manifest: manifestId,
+ ...(maxHeight !== undefined ? { maxHeight } : {}),
+ }),
+ });
+ const body = await res.json().catch(() => ({}));
+ if (!res.ok || !body.ok) {
+ die(`the editor refused (HTTP ${res.status}): ${body.error ?? "no reason given"}`);
+ }
+ EMIT("note", {
+ message:
+ `${items.length} clip(s): ${body.cached.length} already on disk, ` +
+ `${body.jobs.reduce((n, j) => n + j.items, 0)} queued in ${body.jobs.length} job(s)`,
+ });
+ for (const u of body.unresolved ?? []) {
+ EMIT("note", { message: ` ${u.item.clipId ?? u.item.id}: ${u.error}` });
+ }
+ for (const r of body.refused ?? []) {
+ EMIT("note", { message: ` ${r.platform}: ${r.items} window(s) not started — ${r.error}` });
+ }
+ // Every job, polled to its end. A job that stopped short (a rate limit, two
+ // 403s) fails with the reason; asking again later resumes it, the fetched
+ // windows answering from the cache.
+ const pending = new Map(body.jobs.map((j) => [j.jobId, j]));
+ const failed = [];
+ const deadline = Date.now() + 6 * 60 * 60_000;
+ const last = new Map();
+ while (pending.size > 0) {
+ if (Date.now() > deadline) die(`gave up waiting for editor job(s) ${[...pending.keys()].join(", ")}`);
+ await new Promise((r) => setTimeout(r, POLL_MS * 5));
+ for (const [jobId, j] of pending) {
+ const poll = await ask(`${editorUrl}/api/media/fetch-window/${jobId}`, { headers });
+ const p = await poll.json().catch(() => ({}));
+ if (!poll.ok) die(`polling editor job ${jobId} failed (HTTP ${poll.status}): ${p.error ?? ""}`);
+ if (p.status !== last.get(jobId)) {
+ last.set(jobId, p.status);
+ EMIT("note", { message: ` editor job ${jobId} (${j.platform}, ${j.items} window(s)): ${p.status}` });
+ }
+ if (p.status === "done") pending.delete(jobId);
+ else if (p.status === "failed" || p.status === "cancelled") {
+ pending.delete(jobId);
+ failed.push(`${jobId} ${p.status}: ${String(p.error ?? "").split("\n").slice(-3).join(" / ")}`);
+ }
+ }
+ }
+ const unfinished = failed.length + (body.refused?.length ?? 0) + (body.unresolved?.length ?? 0);
+ if (unfinished > 0) {
+ die(`not every window was fetched:\n ${[...failed, ...(body.refused ?? []).map((r) => `${r.platform}: ${r.error}`)].join("\n ") || "see the notes above"}`);
+ }
+ EMIT("done", { out: null, all: true, clips: items.length });
+ process.exit(0);
+}
+
// Same resolution build-video.mjs's --fetch-only does, and for the same
// reasons: a still has nothing to fetch, a non-clip entry is an error, and a
// LEDGER CLAIM is a moment rather than a window (most of a ledger is cited by
@@ -136,37 +246,16 @@ if (!channelSlug) {
die(`${clipId} has no channel, and the manifest's provenance names none`);
}
-const pad = Number(flag("--pad") ?? 3);
-const padBefore = Number(flag("--pad-before") ?? pad);
-const padAfter = Number(flag("--pad-after") ?? pad);
-// TWO DECIMALS, matching the editor's own naming (common/lib/clipWindow.ts) and
-// the build's. The name IS the window, so a request that rounds differently
-// addresses a different file and the cache misses forever.
-const from = Number(Math.max(0, Number(entry.start) - padBefore).toFixed(2));
-const to = Number((Number(entry.end) + padAfter).toFixed(2));
+const { from, to } = windowOf(entry);
// WHY THESE SECONDS. Stored beside the file so a directory of windows can be
// read back months later. The clip's own note is the closest thing the manifest
// has to a reason; the manifest id and the clip id say the rest.
-const manifestId = provenance.manifestId ?? path.basename(path.dirname(path.resolve(manifestPath)));
const reason =
entry.note ??
entry.quote ??
`clip window with ${padBefore}s before / ${padAfter}s after`;
-const headers = {
- authorization: `Bearer ${token}`,
- "content-type": "application/json",
-};
-
-async function ask(url, init) {
- try {
- return await fetch(url, init);
- } catch (err) {
- die(`could not reach the editor at ${editorUrl}: ${err.message}`);
- }
-}
-
EMIT("fetch", {
id: entry.id,
video: entry.video,
diff --git a/umtool/report-to-video/flips.mjs b/umtool/report-to-video/flips.mjs
@@ -0,0 +1,188 @@
+// FLIPS: a cut built of pairs -- a line from THEN and the opposite line from
+// NOW, back to back -- with a panel down the frame's left that holds the pair
+// while it plays: the topic, the THEN card (when, and the words) lit as the
+// first clip plays, then the NOW card slamming in under it as the second
+// starts, the years between them counting up. When the pair ends it lifts
+// away and the next one comes in. Nothing is said for her: the panel's words
+// are hers, verbatim, a few of them.
+//
+// render.chrome: "flips": { "pairs": [ { "id": "vax", "topic": "Vaccines",
+// "then": { "entry": "f-vax-then", "when": "2019", "words": "…" },
+// "now": { "entry": "f-vax-now", "when": "2025", "words": "…" } } ] }
+//
+// With flips on, the footage box moves to the frame's right edge and the panel
+// takes the left, as the thread rail's does (deck.mjs deckGeometry,
+// threadsGeometry); the two are one region and a cut has one or the other.
+//
+// PURE, like threads.mjs. chrome-flips.mjs draws the panel.
+import { planCues } from "./threads.mjs";
+
+/** The limits: pairs, a topic's, a `when`'s and a side's words' characters. */
+export const FLIPS_LIMITS = Object.freeze({ pairs: Object.freeze([1, 24]), topic: 32, when: 18, words: 110 });
+
+/**
+ * The panel's motion, in seconds: a pair rises in over `enter`, its NOW card
+ * slams over `slam` (from `slamScale`) with a flash, the gap counts over
+ * `count`, the THEN card dims over `dim`; a pair lifts away over `leave`.
+ */
+export const FLIP_MOTION = Object.freeze({
+ enter: 0.35, slam: 0.22, slamScale: 1.35, flashUp: 0.05, flashDown: 0.5, count: 0.6, dim: 0.3, leave: 0.3, aside: 0.4,
+});
+
+/** How bright the THEN card rests once NOW has landed. */
+export const THEN_DIM = 0.5;
+
+const ID_RE = /^[A-Za-z0-9][A-Za-z0-9_-]{0,47}$/;
+const isObj = (v) => v !== null && typeof v === "object" && !Array.isArray(v);
+const oneLine = (v, max) => typeof v === "string" && v.trim() !== "" && !/[\n\r]/.test(v) && v.length <= max;
+
+/** Is the flips panel on: a `render.chrome.flips` with at least one pair. */
+export function flipsOn(render) {
+ const f = render?.chrome?.flips;
+ return isObj(f) && Array.isArray(f.pairs) && f.pairs.length > 0;
+}
+
+/**
+ * Every reason `render.chrome.flips` cannot be built, as sentences. Shape
+ * only; `validateFlipEntries` checks the pairs against the timeline.
+ *
+ * @returns {string[]}
+ */
+export function validateFlips(f, where = "render.chrome.flips") {
+ if (f === undefined) return [];
+ if (!isObj(f)) return [`${where} must be an object`];
+ const errors = [];
+ for (const k of Object.keys(f)) if (k !== "pairs") errors.push(`${where}.${k} is not a flips setting`);
+ const [lo, hi] = FLIPS_LIMITS.pairs;
+ if (!Array.isArray(f.pairs) || f.pairs.length < lo || f.pairs.length > hi) return [...errors, `${where}.pairs must be ${lo} to ${hi} pairs`];
+ const ids = new Set();
+ f.pairs.forEach((p, i) => {
+ const w = `${where}.pairs[${i}]`;
+ if (!isObj(p)) { errors.push(`${w} must be an object`); return; }
+ for (const k of Object.keys(p)) if (!["id", "topic", "then", "now"].includes(k)) errors.push(`${w}.${k} is not a pair key`);
+ if (typeof p.id !== "string" || !ID_RE.test(p.id)) errors.push(`${w}.id must be letters, digits, _ or -`);
+ else if (ids.has(p.id)) errors.push(`${w}.id ${p.id} is listed twice`);
+ else ids.add(p.id);
+ if (!oneLine(p.topic, FLIPS_LIMITS.topic)) errors.push(`${w}.topic must be one line of at most ${FLIPS_LIMITS.topic} characters`);
+ for (const side of ["then", "now"]) {
+ const s = p[side];
+ const ws = `${w}.${side}`;
+ if (!isObj(s)) { errors.push(`${ws} must be an object`); continue; }
+ for (const k of Object.keys(s)) if (!["entry", "when", "words"].includes(k)) errors.push(`${ws}.${k} is not a side key`);
+ if (typeof s.entry !== "string" || !s.entry) errors.push(`${ws}.entry must name a timeline entry`);
+ if (!oneLine(s.when, FLIPS_LIMITS.when)) errors.push(`${ws}.when must be one line of at most ${FLIPS_LIMITS.when} characters`);
+ if (!oneLine(s.words, FLIPS_LIMITS.words)) errors.push(`${ws}.words must be one line of at most ${FLIPS_LIMITS.words} characters`);
+ }
+ });
+ return errors;
+}
+
+/**
+ * The pairs against the timeline: each side names an entry of the cut, a clip
+ * (not a teaser), the THEN side before the NOW side, and no entry in two pairs.
+ *
+ * @returns {string[]}
+ */
+export function validateFlipEntries(manifest) {
+ const render = manifest?.render ?? {};
+ if (!flipsOn(render)) return [];
+ const errors = [];
+ const at = new Map((manifest.timeline ?? []).map((e, i) => [e?.id, { e, i }]));
+ const used = new Map();
+ render.chrome.flips.pairs.forEach((p, i) => {
+ const w = `render.chrome.flips.pairs[${i}] (${p?.id ?? "?"})`;
+ const idx = {};
+ for (const side of ["then", "now"]) {
+ const id = p?.[side]?.entry;
+ const hit = at.get(id);
+ if (!hit) { errors.push(`${w}.${side}.entry ${id} is not in the timeline`); continue; }
+ if (hit.e.type === "teaser") errors.push(`${w}.${side}.entry ${id} is a teaser, not a clip`);
+ if (used.has(id)) errors.push(`${w}.${side}.entry ${id} is already in pair ${used.get(id)}`);
+ used.set(id, p.id);
+ idx[side] = hit.i;
+ }
+ if (idx.then !== undefined && idx.now !== undefined && idx.then >= idx.now) errors.push(`${w}: its THEN entry must come before its NOW entry`);
+ });
+ return errors;
+}
+
+/** The years between two `when`s, when both carry one (`2019`, `May 2019`): else null. */
+export function yearsBetween(a, b) {
+ const y = (s) => {
+ const m = /(19|20)\d\d/.exec(String(s ?? ""));
+ return m ? Number(m[0]) : null;
+ };
+ const ya = y(a), yb = y(b);
+ return ya != null && yb != null && yb > ya ? yb - ya : null;
+}
+
+/**
+ * The panel's schedule, in the cut's clock: each pair with its two sides'
+ * spans (`from`: the clip's start; `to`: the next clip's start, or the cut's
+ * end) and the years between; `asides` as the thread rail's (a popup post's
+ * move to the end of its clip).
+ *
+ * @returns {{ pairs: Array<{ id: string, topic: string, index: number, of: number, years: number|null,
+ * then: { segment: string, from: number, to: number, when: string, words: string },
+ * now: { segment: string, from: number, to: number, when: string, words: string } }>,
+ * asides: Array<{ from: number, to: number }> }}
+ */
+export function flipSchedule({ segments, total, render, moves = [] }) {
+ const R = (v) => Math.round(v * 1000) / 1000;
+ const endOf = (i) => (i + 1 < segments.length ? segments[i + 1].start : total);
+ const idx = new Map(segments.map((s, i) => [s.id, i]));
+ const list = flipsOn(render) ? render.chrome.flips.pairs : [];
+ const side = (s) => {
+ const i = idx.get(s.entry);
+ return i === undefined ? null : { segment: s.entry, from: R(segments[i].start), to: R(endOf(i)), when: s.when, words: s.words };
+ };
+ const pairs = list
+ .map((p) => ({ id: p.id, topic: p.topic, then: side(p.then), now: side(p.now), years: yearsBetween(p.then.when, p.now.when) }))
+ .filter((p) => p.then && p.now)
+ .sort((a, b) => a.then.from - b.then.from)
+ .map((p, i, all) => ({ ...p, index: i + 1, of: all.length }));
+ const asides = moves.map((m) => {
+ const i = idx.get(m.segment);
+ return { from: R(m.at), to: R(i === undefined ? total : endOf(i)) };
+ });
+ return { pairs, asides };
+}
+
+/**
+ * Everything the panel's timeline does, as data. Pair j is `p<j>` (autoAlpha,
+ * y); its THEN card `t<j>` (opacity), NOW card `n<j>` (autoAlpha, scale, x)
+ * with its flash `nf<j>`, the gap `g<j>` (autoAlpha) and its count `gc<j>`
+ * (an odometer strip of 0..years, rolled by yPercent -- a transform, so a
+ * seek from anywhere lands on the same digit); the whole panel is `panel`.
+ *
+ * @returns {{ init: Record<string, object>, cues: Array<object> }}
+ */
+export function flipCues(sched) {
+ const m = FLIP_MOTION;
+ const init = { panel: { autoAlpha: 1, x: 0 } };
+ const ev = [];
+ const add = (k, at, dur, to, ease, why) => ev.push({ k, at, dur, to, ease, why });
+ sched.pairs.forEach((p, j) => {
+ init[`p${j}`] = { autoAlpha: 0, y: 60 };
+ init[`t${j}`] = { opacity: 1 };
+ init[`n${j}`] = { autoAlpha: 0, scale: m.slamScale, x: 40 };
+ init[`nf${j}`] = { opacity: 0 };
+ init[`g${j}`] = { autoAlpha: 0 };
+ init[`gc${j}`] = { yPercent: 0 };
+ add(`p${j}`, p.then.from, m.enter, { autoAlpha: 1, y: 0 }, "power3.out", `${p.id} in`);
+ add(`n${j}`, p.now.from, m.slam, { autoAlpha: 1, scale: 1, x: 0 }, "power4.in", `${p.id} now`);
+ add(`nf${j}`, p.now.from + m.slam, m.flashUp, { opacity: 1 }, "none", `${p.id} flash`);
+ add(`nf${j}`, p.now.from + m.slam + m.flashUp, m.flashDown, { opacity: 0 }, "power2.out", `${p.id} flash`);
+ add(`t${j}`, p.now.from, m.dim, { opacity: THEN_DIM }, "power2.out", `${p.id} then dims`);
+ if (p.years != null) {
+ add(`g${j}`, p.now.from, m.slam, { autoAlpha: 1 }, "power2.out", `${p.id} gap`);
+ add(`gc${j}`, p.now.from + 0.05, m.count, { yPercent: -100 * (p.years / (p.years + 1)) }, "power2.out", `${p.id} count`);
+ }
+ add(`p${j}`, p.now.to - m.leave, m.leave, { autoAlpha: 0, y: -50 }, "power2.in", `${p.id} out`);
+ });
+ for (const a of sched.asides) {
+ add("panel", a.from, m.aside, { autoAlpha: 0, x: -40 }, "power2.in", "aside for a post");
+ add("panel", a.to, m.aside, { autoAlpha: 1, x: 0 }, "power2.out", "back after a post");
+ }
+ return { init, cues: planCues(init, ev) };
+}
diff --git a/umtool/report-to-video/flips.test.mjs b/umtool/report-to-video/flips.test.mjs
@@ -0,0 +1,110 @@
+// Tests for the flips panel: its settings and validation (flips.mjs, through
+// validateChrome), the pairs against the timeline, the schedule, the cues and
+// the page (chrome-flips.mjs).
+//
+// Run with: pnpm test:scripts
+import assert from "node:assert/strict";
+import test from "node:test";
+
+import { flipCues, flipSchedule, flipsOn, FLIPS_LIMITS, THEN_DIM, validateFlipEntries, validateFlips, yearsBetween } from "./flips.mjs";
+import { deckGeometry, deckSchedule, railGap, validateChrome } from "./deck.mjs";
+import { flipsHtml } from "./chrome-flips.mjs";
+
+const PALETTE = { bg: "#12101a", fg: "#f4f1ea", muted: "#9a93ad", accent: "#e5534b", amber: "#ffc860" };
+const PAIRS = [
+ { id: "vax", topic: "Vaccines", then: { entry: "a", when: "2019", words: "I trust doctors" }, now: { entry: "b", when: "2025", words: "Never again" } },
+ { id: "dw", topic: "The Daily Wire", then: { entry: "c", when: "May 2021", words: "Best place" }, now: { entry: "d", when: "Ep 300", words: "Worst place" } },
+];
+const CHROME = { engine: "hyperframes", layout: "deck", deck: {}, flips: { pairs: PAIRS } };
+const RENDER = { width: 1920, height: 1080, fps: 30, transition: 0.12, palette: PALETTE, chrome: CHROME };
+const TIMELINE = ["a", "b", "c", "d"].map((id) => ({ id, type: "clip" }));
+
+test("validation: the pairs' shape, and not beside the thread rail", () => {
+ assert.equal(flipsOn(RENDER), true);
+ assert.deepEqual(validateFlips({ pairs: PAIRS }), []);
+ assert.deepEqual(validateChrome(CHROME, RENDER), []);
+ assert.match(validateFlips({ pairs: [] }).join(";"), /1 to 24 pairs/);
+ const bad = validateFlips({
+ pairs: [
+ { id: "x", topic: "t", then: { entry: "a", when: "2019", words: "w".repeat(FLIPS_LIMITS.words + 1) }, now: { entry: "b", when: "x", words: "y" }, extra: 1 },
+ { id: "x", topic: "", then: null, now: { entry: "", when: "", words: "y" } },
+ ],
+ }).join(";");
+ assert.match(bad, /pairs\[0\]\.extra is not a pair key/);
+ assert.match(bad, /then\.words must be one line/);
+ assert.match(bad, /x is listed twice/);
+ assert.match(bad, /pairs\[1\]\.topic must be/);
+ assert.match(bad, /pairs\[1\]\.then must be an object/);
+ assert.match(bad, /now\.entry must name a timeline entry/);
+ const both = { ...CHROME, threads: { list: [{ id: "t", label: "T" }] } };
+ assert.match(validateChrome(both, RENDER).join(";"), /share the left panel/);
+});
+
+test("the pairs against the timeline: entries exist, THEN first, one pair each", () => {
+ assert.deepEqual(validateFlipEntries({ render: RENDER, timeline: TIMELINE }), []);
+ const errs = validateFlipEntries({ render: RENDER, timeline: [{ id: "b", type: "clip" }, { id: "a", type: "teaser" }] }).join(";");
+ assert.match(errs, /vax\)?: its THEN entry must come before its NOW entry/);
+ assert.match(errs, /then\.entry a is a teaser/);
+ assert.match(errs, /then\.entry c is not in the timeline/);
+ const twice = { ...RENDER, chrome: { ...CHROME, flips: { pairs: [PAIRS[0], { ...PAIRS[1], then: { ...PAIRS[1].then, entry: "b" } }] } } };
+ assert.match(validateFlipEntries({ render: twice, timeline: TIMELINE }).join(";"), /b is already in pair vax/);
+});
+
+test("the footage moves right for the panel, as for the rail", () => {
+ const g = deckGeometry(RENDER).footage;
+ assert.equal(g.x, 1920 - g.width - railGap(RENDER));
+});
+
+test("years between two whens, when both carry a year", () => {
+ assert.equal(yearsBetween("2019", "2025"), 6);
+ assert.equal(yearsBetween("May 2021", "October 2026"), 5);
+ assert.equal(yearsBetween("Ep 12", "2025"), null);
+ assert.equal(yearsBetween("2025", "2019"), null);
+});
+
+const SEGS = [
+ { id: "a", start: 0, duration: 6 },
+ { id: "b", start: 5.88, duration: 5 },
+ { id: "c", start: 10.76, duration: 7 },
+ { id: "d", start: 17.64, duration: 4 },
+];
+
+test("the schedule: each pair's spans, its years, its place", () => {
+ const s = flipSchedule({ segments: SEGS, total: 21.64, render: RENDER, moves: [{ segment: "c", at: 11 }] });
+ assert.deepEqual(s.pairs.map((p) => [p.id, p.index, p.of, p.years]), [["vax", 1, 2, 6], ["dw", 2, 2, null]]);
+ assert.deepEqual(s.pairs[0].then, { segment: "a", from: 0, to: 5.88, when: "2019", words: "I trust doctors" });
+ assert.deepEqual([s.pairs[1].now.from, s.pairs[1].now.to], [17.64, 21.64]);
+ assert.deepEqual(s.asides, [{ from: 11, to: 17.64 }]);
+});
+
+test("the cues: a pair rises with THEN, NOW slams and THEN dims, the count rolls to its year, the pair leaves", () => {
+ const s = flipSchedule({ segments: SEGS, total: 21.64, render: RENDER });
+ const { init, cues } = flipCues(s);
+ for (const c of cues) for (const k of Object.keys(c.to)) assert.notEqual(c.from[k], undefined, `${c.k} ${k} has a from`);
+ assert.equal(cues.find((c) => c.k === "p0").at, 0);
+ const slam = cues.find((c) => c.k === "n0");
+ assert.equal(slam.at, 5.88);
+ assert.deepEqual(slam.to, { autoAlpha: 1, scale: 1, x: 0 });
+ assert.equal(cues.find((c) => c.k === "t0").to.opacity, THEN_DIM);
+ const count = cues.find((c) => c.k === "gc0");
+ assert.ok(Math.abs(count.to.yPercent - -100 * (6 / 7)) < 1e-9);
+ assert.equal(cues.find((c) => c.k === "gc1"), undefined, "no years, no count");
+ assert.deepEqual(init.gc0, { yPercent: 0 });
+ const leaves = cues.filter((c) => c.k === "p0" && c.to.autoAlpha === 0);
+ assert.equal(leaves.length, 1);
+ assert.ok(leaves[0].at + leaves[0].dur <= 10.76 + 1e-9, "gone by the next pair");
+});
+
+test("the page: one pair block per pair, both sides' words, an odometer of the years", () => {
+ const entries = SEGS.map((s) => ({ id: s.id, type: "clip" }));
+ const schedule = deckSchedule({ entries, durs: SEGS.map((s) => s.duration), D: 0.12, render: RENDER });
+ assert.ok(schedule.flips, "the deck's schedule carries the pairs");
+ const html = flipsHtml(schedule, RENDER, { fonts: { regular: "r.ttf", bold: "b.ttf" } });
+ assert.equal((html.match(/class="pair"/g) ?? []).length, 2);
+ assert.match(html, /“I trust doctors”/);
+ assert.match(html, /“Never again”/);
+ assert.equal((html.match(/class="digit"/g) ?? []).length, 7);
+ assert.match(html, /years later/);
+ assert.match(html, /01 \/ 02/);
+ assert.throws(() => flipsHtml({ ...schedule, flips: undefined }, RENDER), /no pairs/);
+});
diff --git a/umtool/report-to-video/threads.mjs b/umtool/report-to-video/threads.mjs
@@ -0,0 +1,283 @@
+// THREADS: a cut's clips grouped into the few lines of argument they make,
+// drawn as a rail of cards down the frame's left side. The card of the thread
+// on screen is lit and strung to the footage; a dot fills for each of its
+// clips as it plays; when the thread's last clip ends, its outcome is stamped
+// on its card. By the last clip the rail is the whole argument at a glance.
+//
+// render.chrome: "threads": { "list": [ { "id": "bet", "label": "The bet",
+// "outcome": { "verdict": "CONTRADICTED" } } ] }
+// timeline entry: "thread": "bet"
+//
+// An outcome names a verdict of the shared vocabulary; its label and colour
+// are the cut's (`factcheck.verdicts`), or the outcome's own `label`. A clip
+// with no `thread` belongs to none: nothing is lit while it plays.
+//
+// With threads on, the footage box moves to the frame's right edge and the
+// rail takes the left (deck.mjs deckGeometry, threadsGeometry). A popup post
+// moves the footage over the rail, so the rail steps aside while one is up.
+//
+// PURE, like factcheck.mjs: the settings, their validation, the schedule and
+// the cues. No fs. deck.mjs imports this file -- never the other way round.
+// chrome-threads.mjs draws the rail.
+import { resolveFactcheck, VERDICTS } from "./factcheck.mjs";
+
+/** The limits: how many threads, a label's and an outcome label's characters. */
+export const THREADS_LIMITS = Object.freeze({ threads: Object.freeze([1, 8]), label: 32, outcome: 24 });
+
+/**
+ * The rail's motion, in seconds: a card lights over `light` and its string
+ * draws over `string`; a clip's dot fills over `dot`; an outcome slams in over
+ * `slam` and flashes; the rail steps aside (and back) over `aside`. The cards
+ * come up one after another at the start, `stagger` apart.
+ */
+export const THREAD_MOTION = Object.freeze({
+ light: 0.45, string: 0.5, dot: 0.3, slam: 0.3, flashUp: 0.06, flashDown: 0.6, aside: 0.4, intro: 0.5, stagger: 0.08,
+});
+
+/** How long before its thread's last clip ends the outcome lands, at most. */
+export const OUTCOME_LEAD = 1.8;
+
+/** A thread id: letters, digits and `_ -`. */
+const THREAD_ID_RE = /^[A-Za-z0-9][A-Za-z0-9_-]{0,47}$/;
+
+const isObj = (v) => v !== null && typeof v === "object" && !Array.isArray(v);
+const oneLine = (v) => typeof v === "string" && v.trim() !== "" && !/[\n\r]/.test(v);
+
+/** Is the rail on: a `render.chrome.threads` with at least one thread. */
+export function threadsOn(render) {
+ const t = render?.chrome?.threads;
+ return isObj(t) && Array.isArray(t.list) && t.list.length > 0;
+}
+
+/**
+ * The threads with their outcomes resolved: `[{ id, label, outcome: { verdict,
+ * label, color } | null }]`, in the rail's order (the list's).
+ */
+export function resolveThreads(render) {
+ if (!threadsOn(render)) return [];
+ const verdicts = resolveFactcheck(render).verdicts;
+ return render.chrome.threads.list.map((t) => {
+ const o = isObj(t.outcome) ? t.outcome : null;
+ const v = o ? verdicts[o.verdict] : null;
+ return {
+ id: t.id,
+ label: t.label,
+ outcome: o && v ? { verdict: o.verdict, label: o.label ?? v.label, color: v.color } : null,
+ };
+ });
+}
+
+/**
+ * Every reason `render.chrome.threads` cannot be built, as sentences (empty:
+ * it can). validateChrome calls this, so umtool's writer and the build refuse
+ * the same things.
+ *
+ * @returns {string[]}
+ */
+export function validateThreads(t, where = "render.chrome.threads") {
+ if (t === undefined) return [];
+ if (!isObj(t)) return [`${where} must be an object`];
+ const errors = [];
+ for (const k of Object.keys(t)) if (k !== "list") errors.push(`${where}.${k} is not a threads setting`);
+ const [lo, hi] = THREADS_LIMITS.threads;
+ if (!Array.isArray(t.list) || t.list.length < lo || t.list.length > hi) {
+ return [...errors, `${where}.list must be ${lo} to ${hi} threads`];
+ }
+ const seen = new Set();
+ t.list.forEach((th, i) => {
+ const w = `${where}.list[${i}]`;
+ if (!isObj(th)) { errors.push(`${w} must be an object`); return; }
+ for (const k of Object.keys(th)) if (!["id", "label", "outcome"].includes(k)) errors.push(`${w}.${k} is not a thread key`);
+ if (typeof th.id !== "string" || !THREAD_ID_RE.test(th.id)) errors.push(`${w}.id must be letters, digits, _ or -`);
+ else if (seen.has(th.id)) errors.push(`${w}.id ${th.id} is listed twice`);
+ else seen.add(th.id);
+ if (!oneLine(th.label) || th.label.length > THREADS_LIMITS.label) {
+ errors.push(`${w}.label must be one line of at most ${THREADS_LIMITS.label} characters`);
+ }
+ if (th.outcome !== undefined) {
+ const o = th.outcome;
+ if (!isObj(o)) { errors.push(`${w}.outcome must be an object`); return; }
+ for (const k of Object.keys(o)) if (!["verdict", "label"].includes(k)) errors.push(`${w}.outcome.${k} is not an outcome key`);
+ if (!VERDICTS.includes(o.verdict)) errors.push(`${w}.outcome.verdict must be one of ${VERDICTS.join(", ")}`);
+ if (o.label !== undefined && (!oneLine(o.label) || o.label.length > THREADS_LIMITS.outcome)) {
+ errors.push(`${w}.outcome.label must be one line of at most ${THREADS_LIMITS.outcome} characters`);
+ }
+ }
+ });
+ return errors;
+}
+
+/** The thread an entry belongs to, or null. */
+export function threadOf(entry) {
+ return typeof entry?.thread === "string" && entry.thread ? entry.thread : null;
+}
+
+/**
+ * Every `thread` in the timeline, checked against the list: it names a listed
+ * thread, it is not on a teaser, and every listed thread has a clip. The
+ * build refuses with these before it fetches.
+ *
+ * @returns {string[]}
+ */
+export function validateThreadEntries(manifest) {
+ const errors = [];
+ const render = manifest?.render ?? {};
+ const listed = threadsOn(render) ? new Set(render.chrome.threads.list.map((t) => t?.id)) : new Set();
+ const used = new Set();
+ (manifest?.timeline ?? []).forEach((e, i) => {
+ if (e?.thread === undefined || e?.thread === null) return;
+ const where = `timeline[${i}] (${e.id ?? "?"}).thread`;
+ if (typeof e.thread !== "string" || !e.thread) { errors.push(`${where} must be a thread id`); return; }
+ if (e.type === "teaser") { errors.push(`${where}: a teaser belongs to no thread`); return; }
+ if (!listed.has(e.thread)) { errors.push(`${where} ${e.thread} is not in render.chrome.threads.list`); return; }
+ used.add(e.thread);
+ });
+ for (const id of listed) if (!used.has(id)) errors.push(`render.chrome.threads: thread ${id} has no clip`);
+ return errors;
+}
+
+/**
+ * The rail's schedule, in the cut's clock. `segments` are the schedule's, in
+ * order, each `{ id, start, duration, thread }`; `moves` are the footage's
+ * moves for popup posts (deck.mjs footageMoves).
+ *
+ * - `threads`: each listed thread with its `clips` (`{ segment, at }`: when
+ * its dot fills -- halfway into the dissolve that brings the clip in) and,
+ * with an outcome, `closeAt`: when it is stamped -- `OUTCOME_LEAD` before
+ * its last clip ends, but never in that clip's first half second, and never
+ * while the rail is aside (then as it is back);
+ * - `runs`: when each thread is the one on screen (`{ thread, from, to }`:
+ * consecutive clips of one thread make one run, ending as the next clip
+ * starts to come in);
+ * - `asides`: when the rail steps aside (`{ from, to }`: a popup post's move
+ * to the end of its clip).
+ *
+ * @returns {{ threads: Array<{ id: string, label: string, outcome: object|null,
+ * clips: Array<{ segment: string, at: number }>, closeAt?: number }>,
+ * runs: Array<{ thread: string, from: number, to: number }>,
+ * asides: Array<{ from: number, to: number }> }}
+ */
+export function threadSchedule({ segments, D, total, render, moves = [] }) {
+ const R = (v) => Math.round(v * 1000) / 1000;
+ const endOf = (i) => (i + 1 < segments.length ? segments[i + 1].start : total);
+ const threads = resolveThreads(render).map((t) => ({ ...t, clips: [] }));
+ const byId = new Map(threads.map((t) => [t.id, t]));
+ const last = new Map();
+ segments.forEach((s, i) => {
+ const t = s.thread ? byId.get(s.thread) : null;
+ if (!t) return;
+ t.clips.push({ segment: s.id, at: R(i === 0 ? s.start : s.start + D / 2) });
+ last.set(t.id, i);
+ });
+ const asides = moves.map((m) => {
+ const i = segments.findIndex((s) => s.id === m.segment);
+ return { from: R(m.at), to: R(i < 0 ? total : endOf(i)) };
+ });
+ for (const t of threads) {
+ if (!t.outcome || !last.has(t.id)) continue;
+ const i = last.get(t.id);
+ const s = segments[i];
+ let at = Math.max(s.start + 0.5, endOf(i) - D - OUTCOME_LEAD);
+ // Stamped where it can be seen: a rail stepped aside for a post gets it as it comes back.
+ const hidden = asides.find((a) => at >= a.from - 0.2 && at < a.to + THREAD_MOTION.aside);
+ if (hidden) at = hidden.to + THREAD_MOTION.aside + 0.1;
+ t.closeAt = R(at);
+ }
+ const runs = [];
+ segments.forEach((s, i) => {
+ if (!s.thread || !byId.has(s.thread)) return;
+ const prev = runs[runs.length - 1];
+ if (prev && prev.thread === s.thread && prev.lastIdx === i - 1) {
+ prev.to = R(endOf(i));
+ prev.lastIdx = i;
+ } else runs.push({ thread: s.thread, from: R(s.start), to: R(endOf(i)), lastIdx: i });
+ });
+ return { threads, runs: runs.map(({ lastIdx, ...r }) => r), asides };
+}
+
+/**
+ * Order a page's cue events, clamp each so it never starts before the last on
+ * its own element ends, and state every from (the last `to` on its element,
+ * or its `init`): a render is a seek per frame, in any order, so a cue must
+ * never depend on what played before it. The same bookkeeping as the
+ * stamps' and the posts'.
+ */
+export function planCues(init, events, instant = 0.001) {
+ const r4 = (v) => Math.round(v * 10000) / 10000;
+ const ev = events.map((e, n) => ({ ...e, at: r4(e.at), dur: r4(Math.max(instant, e.dur)), n }));
+ ev.sort((a, b) => a.at - b.at || a.n - b.n);
+ const state = Object.fromEntries(Object.entries(init).map(([k, v]) => [k, { ...v }]));
+ const freeAt = new Map();
+ const cues = [];
+ for (const e of ev) {
+ let { at, dur } = e;
+ const free = freeAt.get(e.k) ?? 0;
+ if (at < free) {
+ const end = at + dur;
+ at = r4(free);
+ dur = r4(Math.max(instant, end - at));
+ }
+ const cur = state[e.k] ?? (state[e.k] = {});
+ const from = {};
+ for (const p of Object.keys(e.to)) from[p] = cur[p];
+ Object.assign(cur, e.to);
+ freeAt.set(e.k, r4(at + dur));
+ cues.push({ k: e.k, at, dur, from, to: e.to, ease: e.ease, why: e.why });
+ }
+ return cues;
+}
+
+/** A card's resting opacity before its thread plays, and after. */
+export const CARD_REST = Object.freeze({ ahead: 0.42, done: 0.8 });
+
+/** A dot before its clip plays: there, hollow-looking and small, so a card shows how many clips its thread has. */
+export const DOT_AHEAD = Object.freeze({ opacity: 0.22, scale: 0.7 });
+
+/**
+ * Everything the rail's timeline does, as data. Card j is `c<j>` (opacity), its
+ * lit wash `w<j>` and string `s<j>` (scaleX from its left), its dots
+ * `d<j>-<n>`, its outcome `o<j>` (autoAlpha, scale) with its flash `f<j>`; the
+ * whole rail is `rail` (autoAlpha, x).
+ *
+ * @returns {{ init: Record<string, object>, cues: Array<object> }}
+ */
+export function threadCues(sched) {
+ const m = THREAD_MOTION;
+ const init = { rail: { autoAlpha: 1, x: 0 } };
+ const ev = [];
+ const add = (k, at, dur, to, ease, why) => ev.push({ k, at, dur, to, ease, why });
+ const firstRun = new Map();
+ for (const r of sched.runs) if (!firstRun.has(r.thread)) firstRun.set(r.thread, r.from);
+ sched.threads.forEach((t, j) => {
+ init[`c${j}`] = { opacity: 0 };
+ init[`w${j}`] = { opacity: 0 };
+ init[`s${j}`] = { scaleX: 0 };
+ t.clips.forEach((_, n) => { init[`d${j}-${n}`] = { ...DOT_AHEAD }; });
+ // A card whose thread is already on screen by the end of its intro comes up lit.
+ const lit = (firstRun.get(t.id) ?? Infinity) <= m.stagger * j + m.intro;
+ add(`c${j}`, m.stagger * j, m.intro, { opacity: lit ? 1 : CARD_REST.ahead }, "power2.out", `intro ${t.id}`);
+ t.clips.forEach((c, n) => add(`d${j}-${n}`, c.at, m.dot, { opacity: 1, scale: 1 }, "back.out(2)", `dot ${t.id} ${c.segment}`));
+ if (t.outcome && t.closeAt != null) {
+ init[`o${j}`] = { autoAlpha: 0, scale: 1.6 };
+ init[`f${j}`] = { opacity: 0 };
+ add(`o${j}`, t.closeAt, m.slam, { autoAlpha: 1, scale: 1 }, "power4.in", `outcome ${t.id}`);
+ add(`f${j}`, t.closeAt + m.slam, m.flashUp, { opacity: 1 }, "none", `outcome ${t.id} flash`);
+ add(`f${j}`, t.closeAt + m.slam + m.flashUp, m.flashDown, { opacity: 0 }, "power2.out", `outcome ${t.id} flash`);
+ }
+ });
+ const idx = new Map(sched.threads.map((t, j) => [t.id, j]));
+ for (const r of sched.runs) {
+ const j = idx.get(r.thread);
+ add(`c${j}`, r.from, m.light, { opacity: 1 }, "power2.out", `light ${r.thread}`);
+ add(`w${j}`, r.from, m.light, { opacity: 1 }, "power2.out", `light ${r.thread}`);
+ add(`s${j}`, r.from + 0.1, m.string, { scaleX: 1 }, "power3.out", `string ${r.thread}`);
+ add(`s${j}`, r.to - m.string * 0.6, m.string * 0.6, { scaleX: 0 }, "power2.in", `unstring ${r.thread}`);
+ add(`w${j}`, r.to, m.light, { opacity: 0 }, "power2.inOut", `dim ${r.thread}`);
+ add(`c${j}`, r.to, m.light, { opacity: CARD_REST.done }, "power2.inOut", `dim ${r.thread}`);
+ }
+ for (const a of sched.asides) {
+ add("rail", a.from, m.aside, { autoAlpha: 0, x: -40 }, "power2.in", "aside for a post");
+ add("rail", a.to, m.aside, { autoAlpha: 1, x: 0 }, "power2.out", "back after a post");
+ }
+ return { init, cues: planCues(init, ev) };
+}
diff --git a/umtool/report-to-video/threads.test.mjs b/umtool/report-to-video/threads.test.mjs
@@ -0,0 +1,201 @@
+// Tests for the thread rail: its settings and their validation (threads.mjs,
+// through validateChrome), the footage box moving aside for it, its schedule
+// (dots, runs, outcomes, asides), its cues, its page (chrome-threads.mjs) and
+// its overlay region.
+//
+// Run with: pnpm test:scripts
+import assert from "node:assert/strict";
+import test from "node:test";
+
+import {
+ CARD_REST, OUTCOME_LEAD, planCues, resolveThreads, threadCues, threadOf, threadSchedule, threadsOn, THREADS_LIMITS,
+ validateThreadEntries, validateThreads,
+} from "./threads.mjs";
+import { deckGeometry, deckSchedule, railGap, threadsGeometry, validateChrome } from "./deck.mjs";
+import { railLayout, threadsHtml } from "./chrome-threads.mjs";
+import { threadsRegion } from "./build-video.mjs";
+
+const PALETTE = { bg: "#12101a", fg: "#f4f1ea", muted: "#9a93ad", accent: "#a97bff", amber: "#ffc860" };
+const LIST = [
+ { id: "bet", label: "The bet", outcome: { verdict: "CONTRADICTED" } },
+ { id: "proof", label: "Definitive proof", outcome: { verdict: "NOT_FOUND", label: "Still coming" } },
+ { id: "aside", label: "An aside" },
+];
+const CHROME = { engine: "hyperframes", layout: "deck", deck: {}, threads: { list: LIST } };
+const RENDER = { width: 1920, height: 1080, fps: 30, transition: 0.5, palette: PALETTE, chrome: CHROME };
+const PLAIN = { ...RENDER, chrome: { engine: "hyperframes", layout: "deck", deck: {} } };
+
+test("the rail is on only with a list of threads", () => {
+ assert.equal(threadsOn(RENDER), true);
+ assert.equal(threadsOn(PLAIN), false);
+ assert.equal(threadsOn({ chrome: { threads: { list: [] } } }), false);
+});
+
+test("an outcome takes its verdict's colour and the cut's label, or its own label", () => {
+ const t = resolveThreads({
+ ...RENDER, chrome: { ...CHROME, factcheck: { verdicts: { CONTRADICTED: { label: "Walked back" } } } },
+ });
+ assert.equal(t[0].outcome.label, "Walked back");
+ assert.match(t[0].outcome.color, /^#[0-9a-f]{6}$/i);
+ assert.equal(t[1].outcome.label, "Still coming");
+ assert.equal(t[2].outcome, null);
+});
+
+test("validation: the list, ids, labels and outcomes", () => {
+ assert.deepEqual(validateThreads(undefined), []);
+ assert.deepEqual(validateThreads({ list: LIST }), []);
+ assert.match(validateThreads({ list: [] }).join(";"), /1 to 8 threads/);
+ assert.match(validateThreads({ list: LIST, side: "left" }).join(";"), /side is not a threads setting/);
+ const bad = validateThreads({
+ list: [
+ { id: "a b", label: "x" },
+ { id: "dup", label: "x" },
+ { id: "dup", label: "y".repeat(THREADS_LIMITS.label + 1) },
+ { id: "o", label: "z", outcome: { verdict: "MAYBE", label: "two\nlines" } },
+ ],
+ }).join(";");
+ assert.match(bad, /list\[0\]\.id must be/);
+ assert.match(bad, /dup is listed twice/);
+ assert.match(bad, /list\[2\]\.label must be one line/);
+ assert.match(bad, /outcome\.verdict must be one of/);
+ assert.match(bad, /outcome\.label must be one line/);
+ // validateChrome reads it, and refuses the feed beside it.
+ assert.deepEqual(validateChrome(CHROME, RENDER), []);
+ assert.match(
+ validateChrome({ ...CHROME, deck: { posts: { layout: "feed" } } }, RENDER).join(";"),
+ /threads needs the popup posts/,
+ );
+ assert.match(validateChrome({ ...CHROME, deck: { footageScale: 0.95 } }, RENDER).join(";"), /rail \d+px wide/);
+});
+
+test("entries: a listed thread, not on a teaser, and every thread used", () => {
+ const timeline = [
+ { id: "a", type: "clip", thread: "bet" },
+ { id: "b", type: "clip", thread: "proof" },
+ { id: "c", type: "clip", thread: "aside" },
+ ];
+ assert.deepEqual(validateThreadEntries({ render: RENDER, timeline }), []);
+ const errs = validateThreadEntries({
+ render: RENDER,
+ timeline: [{ id: "a", type: "clip", thread: "nope" }, { id: "t", type: "teaser", thread: "bet" }],
+ }).join(";");
+ assert.match(errs, /nope is not in render\.chrome\.threads\.list/);
+ assert.match(errs, /a teaser belongs to no thread/);
+ assert.match(errs, /thread proof has no clip/);
+ assert.equal(threadOf({ thread: "bet" }), "bet");
+ assert.equal(threadOf({}), null);
+});
+
+test("the footage moves to the right edge and the rail takes the left", () => {
+ const plain = deckGeometry(PLAIN).footage;
+ const g = deckGeometry(RENDER).footage;
+ assert.equal(g.width, plain.width);
+ assert.equal(g.x, 1920 - g.width - railGap(RENDER));
+ const rail = threadsGeometry(RENDER);
+ assert.deepEqual([rail.x, rail.y, rail.width, rail.height], [0, g.y, g.x, g.height]);
+ assert.equal(rail.cards.x, railGap(RENDER));
+ assert.equal(rail.cards.x + rail.cards.width + railGap(RENDER), g.x);
+ const { cards: _c, ...box } = rail;
+ assert.deepEqual(threadsRegion(RENDER, "/f"), { name: "threads", frames: "/f", ...box });
+});
+
+const SEGS = [
+ { id: "s0", start: 0, duration: 10, thread: "bet" },
+ { id: "s1", start: 9.5, duration: 10, thread: "bet" },
+ { id: "s2", start: 19, duration: 10, thread: null },
+ { id: "s3", start: 28.5, duration: 10, thread: "proof" },
+ { id: "s4", start: 38, duration: 10, thread: "bet" },
+ { id: "s5", start: 47.5, duration: 6, thread: "aside" },
+];
+
+test("the schedule: dots, runs, outcomes on the last clip, asides for posts", () => {
+ const moves = [{ segment: "s3", at: 30, segmentAt: 1.5, seconds: 0.6 }];
+ const s = threadSchedule({ segments: SEGS, D: 0.5, total: 53.5, render: RENDER, moves });
+ const bet = s.threads[0];
+ assert.deepEqual(bet.clips.map((c) => [c.segment, c.at]), [["s0", 0], ["s1", 9.75], ["s4", 38.25]]);
+ // Stamped OUTCOME_LEAD (and the dissolve) before its last clip hands over.
+ assert.equal(bet.closeAt, 47.5 - 0.5 - OUTCOME_LEAD);
+ // A thread without an outcome is never closed.
+ assert.equal(s.threads[2].closeAt, undefined);
+ assert.deepEqual(s.runs, [
+ { thread: "bet", from: 0, to: 19 },
+ { thread: "proof", from: 28.5, to: 38 },
+ { thread: "bet", from: 38, to: 47.5 },
+ { thread: "aside", from: 47.5, to: 53.5 },
+ ]);
+ assert.deepEqual(s.asides, [{ from: 30, to: 38 }]);
+ // An outcome that would land while the rail is aside lands as it comes back.
+ const hid = threadSchedule({
+ segments: SEGS, D: 0.5, total: 53.5, render: RENDER, moves: [{ segment: "s4", at: 40 }],
+ });
+ assert.equal(hid.threads[0].closeAt, 47.5 + 0.4 + 0.1);
+ // A clip too short for the lead is stamped half a second in, never before.
+ const short = threadSchedule({
+ segments: [{ id: "x", start: 0, duration: 1.2, thread: "proof" }], D: 0.5, total: 1.2, render: RENDER,
+ });
+ assert.equal(short.threads[1].closeAt, 0.5);
+});
+
+test("the cues state every from, light and dim each run, and stamp each outcome once", () => {
+ const sched = threadSchedule({ segments: SEGS, D: 0.5, total: 53.5, render: RENDER, moves: [{ segment: "s3", at: 30 }] });
+ const { init, cues } = threadCues(sched);
+ for (const c of cues) {
+ for (const k of Object.keys(c.to)) assert.notEqual(c.from[k], undefined, `${c.k} ${k} has a from`);
+ }
+ // Its thread opens the cut, so its card comes up lit; it lights again for its second run.
+ const lit = cues.filter((c) => c.k === "c0" && c.to.opacity === 1).map((c) => c.at);
+ assert.equal(lit[0], 0);
+ assert.equal(lit[lit.length - 1], 38);
+ assert.equal(cues.find((c) => c.k === "c1").to.opacity, CARD_REST.ahead);
+ const rests = cues.filter((c) => c.k === "c0" && c.to.opacity === CARD_REST.done);
+ assert.equal(rests.length, 2);
+ assert.equal(cues.filter((c) => c.k === "o0" && c.to.autoAlpha === 1).length, 1);
+ assert.equal(init.o2, undefined, "no outcome, no stamp");
+ assert.deepEqual(init["d0-0"], { opacity: 0.22, scale: 0.7 }, "a dot is there, small and dim, before its clip");
+ assert.deepEqual(cues.filter((c) => c.k === "rail").map((c) => c.to.autoAlpha), [0, 1]);
+ // One element's cues never overlap.
+ const byK = {};
+ for (const c of cues) {
+ if (byK[c.k] !== undefined) assert.ok(c.at >= byK[c.k] - 1e-9, `${c.k} overlaps at ${c.at}`);
+ byK[c.k] = c.at + c.dur;
+ }
+});
+
+test("planCues clamps a cue behind its element's last and carries the state", () => {
+ const cues = planCues({ a: { x: 0 } }, [
+ { k: "a", at: 0, dur: 2, to: { x: 1 }, ease: "none" },
+ { k: "a", at: 1, dur: 2, to: { x: 2 }, ease: "none" },
+ ]);
+ assert.deepEqual(cues.map((c) => [c.at, c.dur, c.from.x, c.to.x]), [[0, 2, 0, 1], [2, 1, 1, 2]]);
+});
+
+test("the page: a card per thread, a dot per clip, an outcome where one is set", () => {
+ const segments = SEGS.map(({ thread, ...s }) => s);
+ const entries = SEGS.map((s) => ({ id: s.id, type: "clip", ...(s.thread ? { thread: s.thread } : {}) }));
+ const schedule = deckSchedule({ entries, durs: SEGS.map((s) => s.duration), D: 0.5, render: RENDER });
+ assert.ok(schedule.threads, "the deck's schedule carries the rail");
+ assert.equal(segments.length, schedule.segments.length);
+ const html = threadsHtml(schedule, RENDER, { fonts: { regular: "assets/r.ttf", bold: "assets/b.ttf" } });
+ assert.equal((html.match(/class="card"/g) ?? []).length, 3);
+ assert.equal((html.match(/class="dot"/g) ?? []).length, 5);
+ assert.equal((html.match(/class="outcome"/g) ?? []).length, 2);
+ assert.match(html, /data-composition-id="threads"/);
+ assert.match(html, />Still coming</);
+ assert.throws(() => threadsHtml({ ...schedule, threads: undefined }, RENDER), /no thread rail/);
+ // No rail, no `threads` in the schedule: a cut without one writes what it always did.
+ const plain = deckSchedule({ entries, durs: SEGS.map((s) => s.duration), D: 0.5, render: PLAIN });
+ assert.equal("threads" in plain, false);
+});
+
+test("railLayout: cards share the height, capped, the stack centred", () => {
+ assert.deepEqual(railLayout(886, 5), { h: 166, gap: 14, top: 0 });
+ const three = railLayout(886, 3);
+ assert.equal(three.h, 190);
+ assert.equal(three.top, Math.floor((886 - 3 * 190 - 2 * 14) / 2));
+});
+
+test("the renderer runs without a display: a stale DISPLAY would break SwiftShader", async () => {
+ const { rendererEnv } = await import("./compose-chrome.mjs");
+ const env = rendererEnv({ DISPLAY: ":0.0", WAYLAND_DISPLAY: "wayland-0", PATH: "/usr/bin", HOME: "/h" });
+ assert.deepEqual(env, { PATH: "/usr/bin", HOME: "/h" });
+});