commit a3c3f4d6f1751f701350fd6154881e0afab45785
parent 6d34aa9d6ea97406c4e6f27af121a7a982315ea9
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Mon, 5 Oct 2026 03:38:11 -0400
records: REPORT.md on making a report from what exists; PUBLISH.md row; report-to-video README on converting and an image's quote; an [Unreleased] bullet
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Diffstat:
5 files changed, 35 insertions(+), 0 deletions(-)
diff --git a/PUBLISH.md b/PUBLISH.md
@@ -54,6 +54,7 @@ entry points in `common/publish/build.ts`.
| The homepage | /sites → Homepage → **Build homepage** (tick *Deploy after build*) / **Deploy homepage**, with an optional preview branch | `build-homepage` (`{"deploy":true}` to deploy after), `deploy-homepage` (`{"preview":"<branch>"}`) | `build homepage [--no-source]`, `deploy homepage [--preview <branch>]` |
| The source mirror alone | — (every homepage build runs it) | — | `source publish [--force] [--check] [--keep-scratch]`, `source audit [<git dir>]` |
| A site's report evidence media | — (a Reports tab is to come) | `reports-prepare` (`{"siteId"}`) | `reports prepare <id>` |
+| A report from a /sweep report, an /ask answer or a report-to-video manifest, and a starter manifest from a report | — | — | `reports convert <sweep\|ask\|manifest> <in> --out <report.json> [--channels-dir <dir>]`, `reports to-manifest <report.json> --out <manifest.json>` |
`pnpm archilyzer <command>` is the short form of
`pnpm --filter yt-dlp-transcript-common exec tsx bin/archilyzer.ts <command>`;
diff --git a/REPORT.md b/REPORT.md
@@ -61,3 +61,7 @@ The shared vocabulary (`common/lib/report/verdicts.mjs` — the one copy; report
| `CONTRADICTED` | Contradicted | `#e5534b` |
| `NOT_FOUND` | Not found | `#8b93a7` |
| `UNTESTABLE` | Untestable | `#7d8fd6` |
+
+## Making a report from what exists
+
+`archilyzer reports convert <sweep|ask|manifest> <in> --out <report.json>` makes a report of a /sweep report (markdown: its `#` heading the title, each `##` section a section, each list item or paragraph that cites a claim), an /ask answer (its `[n @ mm:ss]` markers resolved against its sources) or a report-to-video manifest (chapter cards the sections, `claim` entries the claims, clips, stills and posts the citations), and writes it only when it validates. A /sweep or /ask citation carries one second: `--channels-dir <transcripts/channels>` widens it through the record's cues to whole sentences (else it is the second plus 10 s) and finds the channel that keeps a post cited by its platform link. `archilyzer reports to-manifest <report.json> --out <manifest.json>` goes the other way: a starter manifest — a title card, a chapter card per section, each claim's still and clips stamped with its verdict, its posts. The converters are `common/lib/report/convert*.ts`; the warnings say what a conversion left out.
diff --git a/common/lib/report/docs.ts b/common/lib/report/docs.ts
@@ -74,5 +74,23 @@ export function renderReportMarkdown(): string {
out.push(`| \`${v}\` | ${cell(d.label)} | \`${d.color}\` |`);
}
out.push("");
+ out.push("## Making a report from what exists");
+ out.push("");
+ out.push(
+ "`archilyzer reports convert <sweep|ask|manifest> <in> --out <report.json>` makes a " +
+ "report of a /sweep report (markdown: its `#` heading the title, each `##` section a " +
+ "section, each list item or paragraph that cites a claim), an /ask answer (its " +
+ "`[n @ mm:ss]` markers resolved against its sources) or a report-to-video manifest " +
+ "(chapter cards the sections, `claim` entries the claims, clips, stills and posts the " +
+ "citations), and writes it only when it validates. A /sweep or /ask citation carries " +
+ "one second: `--channels-dir <transcripts/channels>` widens it through the record's " +
+ "cues to whole sentences (else it is the second plus 10 s) and finds the channel that " +
+ "keeps a post cited by its platform link. `archilyzer reports to-manifest <report.json> " +
+ "--out <manifest.json>` goes the other way: a starter manifest — a title card, a " +
+ "chapter card per section, each claim's still and clips stamped with its verdict, its " +
+ "posts. The converters are `common/lib/report/convert*.ts`; the warnings say what a " +
+ "conversion left out.",
+ );
+ out.push("");
return out.join("\n");
}
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -4,6 +4,7 @@
- **A site can say what it publishes, and which reports.** `site.json` takes `publish` — `"full"`, the searchable corpus every site has been (the default, never written), or `"cited"`, only the site's reports and the moments they cite — and `reports`, the ordered ids of its published reports (each a slug; invalid and repeated ids are dropped). The site form has a Publish control and lists the site's reports read-only; saving the form keeps the stored list. A cited site still builds as a full one until the reports pipeline applies the scope. SITE.md documents both keys.
- **A report and its citations now have one written format, checked before anything is built from them.** A cited report is a `report.json` (`archilyzer-report`, version 1): a summary, then sections of claims, each claim with an optional verdict, the reviewed document's own sentence, findings in markdown and the citations it rests on. A citation is one of five kinds — a span of a video, a span of an audio record, a post, a sentence of a source document, or a web page — with a verbatim quote, and is cited from any markdown in the report as `[label](cite:<id>)`. The checker lists every problem at once with where it is: a citation, a source or a `cite:` link that names nothing, a span that ends before it starts or runs past 120 seconds with its context, a still that points outside the report's folder, an id used twice. Each cited span and post has one page address, `/m/<channel>/<id>/<start>-<end>/` or `/m/<channel>/<id>/`. Nothing builds or shows reports yet. The fact-check verdicts (Corroborated, Partly true, Contradicted, Not found, Untestable) and their colours are now kept in one place, which the report video's stamps and tally read too. `REPORT.md` and `CITATIONS.md` list every key.
- **A site's reports can have their evidence media prepared: every cited span cut to a clip, every cited post's capture copied.** `archilyzer reports prepare <site>`, the `reports-prepare` job (`POST /api/ops/reports-prepare`, `pnpm ops reports-prepare`) reads the site's published reports, checks them, and for each cited moment cuts the span (with its context) out of the media already on disk — a fetched clip window, the saved video, or the recording's audio — fitted inside 1280×720 with H.264 and AAC, or as an `.m4a` for an audio span; and copies the screenshot and attached media of each cited post, and of no other post, beside them. Everything lands in the site's build staging (`.export-index/sites/<site>/report-media/`) with an `index.json` naming each moment's file, size, checksum, size in pixels and duration. A clip is cut once and reused while its source file and span are unchanged; a clip or capture no longer cited is removed. Nothing is fetched: a citation whose media is not on disk, a post without a screenshot, a clip over 24 MiB, a citation of a channel outside the site, a post the site may not show, or an invalid or missing report is listed with the citations it affects, and the run fails (exit 1, or a failed job) — a span on a drive that is not mounted is reported as such rather than as missing. The lookup of a span's media on disk is now shared with report-to-video, which finds the same files it did.
+- **A report can be made from a /sweep report, an /ask answer or a report video's manifest, and a report video's manifest from a report.** `archilyzer reports convert sweep|ask|manifest <in> --out <report.json>` writes a `report.json`. From a /sweep report (markdown): its first `#` heading is the title, the text before the first section the summary, each `##` section a section, and each list item or paragraph with a citing link a claim — a line that is only a citation under a quote cites that quote — with every archive moment link, archive post link, moment page link and X or Bluesky post link turned into a citation and the link into `[label](cite:<id>)`; a citation's quote is the quoted words nearest before its link. From an /ask answer (`{ "answer", "sources" }`, a saved chat or an assistant message, as JSON): each `[n]` or `[n @ mm:ss]` marker becomes a citation of source `n` at the excerpt line nearest the marked time, quoting that line, or of the post. From a manifest: chapter cards are the sections, entries carrying `claim` the claims with their verdicts (the first still a claim's source sentence), clips the video citations, stills the source citations and `posts` the post citations; ledger rows are claims too. A /sweep or /ask citation names one second: with `--channels-dir <transcripts/channels>` its span is the cited cue run on to the quote's length and widened to whole sentences, as `resolve-windows.mjs` widens a clip (the widening now lives in common and is shared), and a post cited by its X or Bluesky link is found in the channel that archives it; without it the span is the second plus 10 seconds and such a post stays a plain link. `archilyzer reports to-manifest <report.json> --out <manifest.json>` writes a starter manifest: a title card, a chapter card per section, the section's cited clips, and per claim its source sentence's still and its clips, each with the claim's verdict, and its posts attached to its last clip (a post's platform, link and date come from the channels tree, or for an X post from its id). Both check what they write against the report format and write nothing when it has problems (exit 1); everything a conversion had to leave out or guess is printed as a warning. A manifest image may carry `quote`, the words the still shows, which the converters keep.
- **A long report video no longer runs out of memory while its clips are crossfaded.** `build-video.mjs` used to join every segment of a cut in one ffmpeg command, which grows with the number of segments: a cut of a few hundred clips could use more memory than the machine had and be stopped. Past 24 segments the build now crossfades them in batches of consecutive segments, each into a file under `out/<variant>/xfade-batches/`, then crossfades those files together with the same transition and lays the on-screen deck, the rail and the dips over them. Every transition, the deck's schedule and the chapters land on the same frames as before, at the cost of one more video encode on such a cut. The batch size is `render.xfadeBatch` in the manifest or `REPORT_VIDEO_XFADE_BATCH` in the environment (which wins); `0` never batches. A batch file is reused while its segments, their holds and moves and the encode settings are unchanged, so a `--chrome-only` run that moves no footage redoes only the last pass.
- **Auto-download no longer tries a video the metadata scan already found members-only or private.** The scan records why it could not read a video, but only a failed download used to take a video out of the auto-download queue, so each members-only video the scan had found was still downloaded once: four yt-dlp requests, two with browser cookies, about 30 seconds each. A members-only or private answer from the scan now keeps the video out of the queue and counts it under the channel's members-only or private exclusions, and it stays in **Needs cookies** for a manual cookie run. A scan that reads the video later lifts this. A video the scan saw only as "Video unavailable" is still tried, since YouTube gives that answer when it is throttling too.
- **X post fetches stop on Drain, and an account with no posts is not searched.** Draining a `fetch-posts` job used to do nothing until gallery-dl finished its whole run. Now the timeline fetch stops at the next page boundary (at once when gallery-dl is between pages or waiting out a rate limit), the older-posts walk stops its current window's search at once and never starts the 45–120 second pause between windows, and both keep their resume point: the job ends done, not failed, and the log says "Drained; the next run resumes …". An older-posts walk is refused when nothing is archived and the last timeline fetch finished having read no posts, since it would only repeat empty searches; `"force": true` (`--force` on `archilyzer posts fetch`) walks anyway. A walk with nothing archived that finds nothing ends after two empty three-month windows instead of four, and records why; a walk that has posts keeps the year-of-empty-windows rule. Capture-posts already stopped between posts on Drain.
diff --git a/umtool/report-to-video/README.md b/umtool/report-to-video/README.md
@@ -256,6 +256,16 @@ the manifest names the MCP video id while the cue file lives under the URL slug.
`timeline` is an ordered list; entries are `card`, `clip` or `image` (plus
`scroll`, `chart`, `ledger` and `teaser` — the vocabulary is open).
+A manifest and a report site's `report.json` convert into each other:
+`archilyzer reports to-manifest <report.json> --out video.manifest.json` writes a
+starter cut of a report (a title card, a chapter card per section, each claim's
+still and clips carrying its `claim`, its posts), and `archilyzer reports convert
+manifest video.manifest.json --out report.json` makes a report of a manifest — see
+REPORT.md, "Making a report from what exists". An image's `quote` is what the
+converters carry as the still's words; nothing in a build reads it. The sentence
+widening `resolve-windows.mjs` does is `common/lib/cueWiden.mjs`, which the
+converters share.
+
```jsonc
{ "type": "card", "id": "ch3", "style": "chapter", "seconds": 4.0,
"kicker": "March – November 2025", "heading": "Then: the county",
@@ -582,6 +592,7 @@ poster — shown for `seconds` and then gone.
"title": "Bx (@bx_on_x) on X, Sept 2026 — 1/4",
"date": "2026-08-19", // optional; appended like a clip's
"citeUrl": "https://…", // optional; the ONLY thing that draws a QR
+ "quote": "…", // optional: the words the still shows (not drawn)
"note": "…" }
```