commit a39ef7882425b75632ce785d1276666db3f8bd55
parent 5dd228561903b51bf4d51b0e1b3caf29f29ed90f
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Mon, 5 Oct 2026 03:11:17 -0400
records: PUBLISH.md on reports prepare; an [Unreleased] bullet
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Diffstat:
2 files changed, 13 insertions(+), 0 deletions(-)
diff --git a/PUBLISH.md b/PUBLISH.md
@@ -53,6 +53,7 @@ entry points in `common/publish/build.ts`.
| The hub | /sites → Hub → **Build hub** / **Deploy hub** | `build-hub`, `deploy-hub` | `build hub`, `deploy hub [--preview <branch>]` |
| The homepage | /sites → Homepage → **Build homepage** (tick *Deploy after build*) / **Deploy homepage**, with an optional preview branch | `build-homepage` (`{"deploy":true}` to deploy after), `deploy-homepage` (`{"preview":"<branch>"}`) | `build homepage [--no-source]`, `deploy homepage [--preview <branch>]` |
| The source mirror alone | — (every homepage build runs it) | — | `source publish [--force] [--check] [--keep-scratch]`, `source audit [<git dir>]` |
+| A site's report evidence media | — (a Reports tab is to come) | `reports-prepare` (`{"siteId"}`) | `reports prepare <id>` |
`pnpm archilyzer <command>` is the short form of
`pnpm --filter yt-dlp-transcript-common exec tsx bin/archilyzer.ts <command>`;
@@ -61,6 +62,17 @@ this machine has what a build needs. From `export/`, `pnpm run build` is
`archilyzer build site` and `pnpm run deploy` is `archilyzer deploy site`; both take
the site from `SITE_ID` when no id is given.
+A site with `reports` (site.json) has one more step, before its build and on the
+host: **reports prepare** cuts the evidence clip of every span its published reports
+cite — from the editor's clip windows, the saved-video store or a record's audio,
+fitted inside 1280×720 (H.264 crf 23, AAC; an audio span is an `.m4a`) — and copies
+the screenshot and media of every cited post, only those, into
+`.export-index/sites/<id>/report-media/`, with a manifest (`index.json`) of each
+moment's file, size, hash and duration. Nothing is fetched: a citation whose media
+is not on disk, a clip over 24 MiB or an invalid report is listed and fails the run
+(exit 1, or a failed job) — fetch the window or persist the video, capture the
+post, and run it again; what is already cut is reused.
+
Deploy-only ships whatever is in `export/out`, which the basic build composes one
site at a time into a single shared directory — so it **refuses, before starting a
job, if `export/out` holds a build of another site** (or no build at all), naming the
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -3,6 +3,7 @@
## [Unreleased]
- **A site can say what it publishes, and which reports.** `site.json` takes `publish` — `"full"`, the searchable corpus every site has been (the default, never written), or `"cited"`, only the site's reports and the moments they cite — and `reports`, the ordered ids of its published reports (each a slug; invalid and repeated ids are dropped). The site form has a Publish control and lists the site's reports read-only; saving the form keeps the stored list. A cited site still builds as a full one until the reports pipeline applies the scope. SITE.md documents both keys.
- **A report and its citations now have one written format, checked before anything is built from them.** A cited report is a `report.json` (`archilyzer-report`, version 1): a summary, then sections of claims, each claim with an optional verdict, the reviewed document's own sentence, findings in markdown and the citations it rests on. A citation is one of five kinds — a span of a video, a span of an audio record, a post, a sentence of a source document, or a web page — with a verbatim quote, and is cited from any markdown in the report as `[label](cite:<id>)`. The checker lists every problem at once with where it is: a citation, a source or a `cite:` link that names nothing, a span that ends before it starts or runs past 120 seconds with its context, a still that points outside the report's folder, an id used twice. Each cited span and post has one page address, `/m/<channel>/<id>/<start>-<end>/` or `/m/<channel>/<id>/`. Nothing builds or shows reports yet. The fact-check verdicts (Corroborated, Partly true, Contradicted, Not found, Untestable) and their colours are now kept in one place, which the report video's stamps and tally read too. `REPORT.md` and `CITATIONS.md` list every key.
+- **A site's reports can have their evidence media prepared: every cited span cut to a clip, every cited post's capture copied.** `archilyzer reports prepare <site>`, the `reports-prepare` job (`POST /api/ops/reports-prepare`, `pnpm ops reports-prepare`) reads the site's published reports, checks them, and for each cited moment cuts the span (with its context) out of the media already on disk — a fetched clip window, the saved video, or the recording's audio — fitted inside 1280×720 with H.264 and AAC, or as an `.m4a` for an audio span; and copies the screenshot and attached media of each cited post, and of no other post, beside them. Everything lands in the site's build staging (`.export-index/sites/<site>/report-media/`) with an `index.json` naming each moment's file, size, checksum, size in pixels and duration. A clip is cut once and reused while its source file and span are unchanged; a clip or capture no longer cited is removed. Nothing is fetched: a citation whose media is not on disk, a post without a screenshot, a clip over 24 MiB, a citation of a channel outside the site, a post the site may not show, or an invalid or missing report is listed with the citations it affects, and the run fails (exit 1, or a failed job) — a span on a drive that is not mounted is reported as such rather than as missing. The lookup of a span's media on disk is now shared with report-to-video, which finds the same files it did.
- **A long report video no longer runs out of memory while its clips are crossfaded.** `build-video.mjs` used to join every segment of a cut in one ffmpeg command, which grows with the number of segments: a cut of a few hundred clips could use more memory than the machine had and be stopped. Past 24 segments the build now crossfades them in batches of consecutive segments, each into a file under `out/<variant>/xfade-batches/`, then crossfades those files together with the same transition and lays the on-screen deck, the rail and the dips over them. Every transition, the deck's schedule and the chapters land on the same frames as before, at the cost of one more video encode on such a cut. The batch size is `render.xfadeBatch` in the manifest or `REPORT_VIDEO_XFADE_BATCH` in the environment (which wins); `0` never batches. A batch file is reused while its segments, their holds and moves and the encode settings are unchanged, so a `--chrome-only` run that moves no footage redoes only the last pass.
- **Auto-download no longer tries a video the metadata scan already found members-only or private.** The scan records why it could not read a video, but only a failed download used to take a video out of the auto-download queue, so each members-only video the scan had found was still downloaded once: four yt-dlp requests, two with browser cookies, about 30 seconds each. A members-only or private answer from the scan now keeps the video out of the queue and counts it under the channel's members-only or private exclusions, and it stays in **Needs cookies** for a manual cookie run. A scan that reads the video later lifts this. A video the scan saw only as "Video unavailable" is still tried, since YouTube gives that answer when it is throttling too.
- **X post fetches stop on Drain, and an account with no posts is not searched.** Draining a `fetch-posts` job used to do nothing until gallery-dl finished its whole run. Now the timeline fetch stops at the next page boundary (at once when gallery-dl is between pages or waiting out a rate limit), the older-posts walk stops its current window's search at once and never starts the 45–120 second pause between windows, and both keep their resume point: the job ends done, not failed, and the log says "Drained; the next run resumes …". An older-posts walk is refused when nothing is archived and the last timeline fetch finished having read no posts, since it would only repeat empty searches; `"force": true` (`--force` on `archilyzer posts fetch`) walks anyway. A walk with nothing archived that finds nothing ends after two empty three-month windows instead of four, and records why; a walk that has posts keeps the year-of-empty-windows rule. Capture-posts already stopped between posts on Drain.