# Report sites — a site built around cited reports Status: BUILT — R0–R8 merged 2026-10-05 (plus R4b page size); first private report site built locally. RH (per-report revision history) built 2026-10-05, see "History". Report-only is derived from a per-site `search` switch (search off); export's `start` serves moment dirs (`scripts/serve-out.mjs`). Follow-ups: cited corpus.json still describes shards, voice checks recorded in citations, saveSiteAction onto patchSite. ## What it is Two independent axes on a site: | Axis | Values | Today | |---|---|---| | `reports` | 0..n published reports, each a cited document rendered as pages | none | | `search` | on — the searchable corpus (today); off — ONLY the moments the site's reports cite | always on | - A **report site** is search off (`search: false`) + ≥1 report: no search, no browse, no full transcripts, no archives. - A **full site with reports** keeps everything it has and gains a Reports section; its moment pages also link into the corpus (`/?v=…&t=`). - Every citation in a report opens a **moment page** on the site: the quote, a self-hosted evidence clip of the cited span (720p, ± a little context), the surrounding transcript lines, record metadata (title, channel, date), and a link to the original at its time. A post citation shows the captured screenshot and media (capture-posts) and the post text instead of a clip. - A report claim can carry the reviewed source's own sentence as an image (shoot-page) and the source's archive links. First user: one private, local-only report site. The machinery is generic so later sites (and later reports on existing full sites) need configuration, not code. ## The report document — `report.json` (format `archilyzer-report`, version 1) ```jsonc { "format": "archilyzer-report", "version": 1, "id", "kind": "factcheck" | "sweep", "title", "subtitle", "summary" /* md */, "method" /* md, optional: "How it was checked" */, "published", "updated", "subject": { "source": "s0" }, // the document under review (optional) "verdicts": { "PARTLY": { "label": "…" } }, // optional label/colour overrides of the shared vocab "sources": { "s0": { "kind": "article", "title", "url", "publisher", "author", "date", "archives": [{ "label", "url", "context" }], // links in context, as the source had them "saved": "sources/s0/page.html" } }, // input for stills only — NEVER published "citations": { "c01": { "kind": "video", "channel", "id", "start", "end", "quote", "pad": { "before": 5, "after": 5 } }, "p01": { "kind": "post", "channel", "id", "quote" }, "a01": { "kind": "source", "source": "s0", "quote", "image": "stills/a01.png" } }, "sections": [ { "id", "title", "body" /* md */, "claims": [ { "id", "text", "verdict" /* optional */, "gist" /* optional, ≤240: its line in "What the check found" */, "flag" /* optional, ≤60 chars: a pill on the card */, "sourceQuote": { "citation": "a01" }, "findings": "md with [label](cite:c01) refs", "citations": ["c01", "p01"] } ] } ] } ``` - A fact-check is sections (chapters) → claims with a verdict. A sweep report is sections with no verdicts, or bodies with inline `cite:` links. - The verdict vocabulary (CORROBORATED, PARTLY, CONTRADICTED, NOT_FOUND, UNTESTABLE + labels/colours) moves to common; `umtool/report-to-video/factcheck.mjs` re-exports it — one copy. - Files per site: `transcripts/sites//reports//{report.json, stills/, sources//…}`. `site.json` `reports: [ids]` is the published, ordered list; an unlisted directory is a draft. - Validation (zod, `common/lib/report/`): ids resolve, every `cite:` ref exists, spans are sane (cap ~120 s), stills exist; at compose time every video quote is checked against its cue window (fail loudly on drift; a platform with two ids must say which). ## The citation system (operator steer, 2026-10-05: "a well-integrated citation system that can be useful on any type of report") Citations are the core; reports are one consumer. - **One model** — `common/lib/citations/` (reports import it): kinds `video` (span), `audio` (span; podcasts), `post` (optional thread), `source` (a document's sentence: still + the source's archive links), `page` (url + archived url); a versioned discriminated union, extensible. Every citation: verbatim `quote`, optional `speaker`, `date`, `label`, `note`, and a COMPUTED `verification` block (quote score against the cues, voice check, when, method) filled at compose, never typed by hand. - **One inline syntax** in any report markdown: `[label](cite:id)`; per-report numbering by first appearance; stable anchors `#c-`. - **Shared components** (`common/components/citations/`), used by the export site, editor previews and umtool's report preview: inline marker with a hover/tap preview card (quote, speaker, date, thumbnail or post shot, play for spans), citation card, numbered reference list per section/report, verification badge. - **Moment pages show back-links** — "Cited in": every report/claim on the site that cites the moment (a pure index built at compose). - **Round trip** — converters bring `/sweep` markdown, `/ask` answers and umtool manifests into the model; a report can emit a report-to-video manifest (the video and the page share one source); the report page offers its citations as JSON and CSV; MCP may later serve `get_report` / citations. ## New `site.json` keys | Key | Default | Notes | |---|---|---| | `search` | `true` | `false` publishes only reports + their moments (report-only). Only `false` is written. | | `reports` | `[]` | Ordered report ids under `reports/`. | The editor form has a Search checkbox. A `publish: "cited"` read from a site.json reads as `search: false` and is written as such on save. `corpus.json` says `site.scope: "cited"` for a report-only site. `channels` keeps its meaning (the pool a site's citations may resolve against). Docs in `SITE_FIELD_DOCS`, `SITE.md` regenerated; the editor's site form gets the Search checkbox; a new **Reports** tab under `/sites//` lists reports, their validation problems, and "Prepare evidence media". ## Routes (export app) | Route | What | |---|---| | `/reports/` | Report index; the HOME page of a `cited` site | | `/reports//` | The report: summary, tally (fact-check), sections → claims (verdict chip, source sentence still, findings with citation links, archive links in context) | | `/m///-/` | Video moment page | | `/m///` | Post moment page | - Server-rendered at build from JSON that compose writes under `public/reports/**` and `public/m/**`, read like `export/app/lib/archives.ts`. - `generateStaticParams` must never be empty (Next 16 static export error E87 on an empty dynamic route): a site with zero reports emits one placeholder param that renders "no reports". - A `cited` site's shell has no workspace/search, no `/ask`, no downloads/duplicates; the home page must not call `countTranscripts()` (it throws without a summaries manifest). - A full site shows a **Reports** header link when it has reports. ## Evidence media - Clips are cut ON THE HOST, before the per-site build, by a prepare step (`archilyzer reports prepare ` and an editor job kind with `needsMedia`): the container comes from the corpus clip window (`data//clips/-.*`) or the saved-video store pointer; cut accurately, 1280×720 H.264 crf 23, AAC. Cached at `.export-index/sites//report-media/.mp4` (hash = source + span + profile), so the docker build (read-only mounts, no ffmpeg) only copies. - The cutter's tier lookup and cut args move from `umtool/report-to-video/sources.mjs` / `umtool/lib/report/` into `common/lib/evidenceClip-server.ts`; umtool re-exports (common may not import umtool). - A citation whose media is missing fails prepare with a list (the editor's fetch-window / persist fill it). - Published layout: `out/media/clips///-.mp4`, `out/media/posts///…` (only cited captures; the post-visibility rule still applies), `out/reports//stills/*.png`. - Limits: refuse a file over 24 MiB (Pages: 25 MiB), keep the site under the 20k-file cap. - The existing guard test that forbids publish code from naming `posts-media` is amended to allow exactly the one module that copies CITED captures, with a test that it copies nothing else. ## Exports A report must survive a takedown as files anyone can save and host again (slice RX). - `archilyzer reports export [--report ] [--formats html,pdf,md,zip] [--allow-missing-media]` and the editor's `reports-export` job (prepare's queue; Reports tab → **Export reports**) run `common/publish/reportExports.ts`; `reports prepare` runs it at its end when nothing is missing. It resolves the reports exactly as compose does (`resolveSiteReports`) and writes, per report, to `.export-index/sites//report-exports//`: - `report.html` — ONE file from `common/lib/report/exportHtml.ts` (pure view → HTML): inline CSS, no script; stills and post screenshots recompressed on the host (WebP, else JPEG, ≤ 1200 px wide) into data URIs; clips linked on the site, never inlined. The heading mirrors the site: the series and the title on two lines (or the title alone), one line of dates and the export's revision, the document under review as a cited line on its rail, the subtitle. Then the three tiers, each opened by a hairline with depth dots: the tally, sections → claims (verdict chip, the document's sentence as its still — else its words, else the claim's text, plain — findings with `[n]` markers, evidence with the added mark beside its number), the documents quoted with their archive links, and the numbered references (quote, speaker, date, record, the original at its time, archive links, the moment page and clip when the site has a `siteUrl`). - `report.pdf` — that HTML printed by headless Chromium (`importPlaywright`, A4); skipped with a note where Playwright or its browser is missing, never a failure. - `report.md` — plain Markdown (`exportMarkdown.ts`), `[label](cite:id)` → `label [n]`, numbered references. - `evidence-pack.zip` — `/report.html` playing its own `media/` (clips, stills, screenshots), `report.md`, `citations.json`/`.csv`; packed by the system `zip` with sorted names, fixed times and modes (the same report checked at the same time packs to the same bytes). No `zip` fails that format, naming it. - `export.json` — each file's size and sha256, the sha256 of the report.json it was made from, the footer, notes. Each run replaces the directory: nothing is left from another version. - Every export ends with the footer `Revision N · · report sha256 · commit `; an unknown part is left out. Today: the report's date and hash. The revision history (slice RH) fills `revision` and `commit` in ONE place, `exportFooterFor` (`reportExports.ts`). - Compose (`publishableReportExports`, `common/publish/reportExportFiles.ts`) copies `report.html`, `report.pdf`, `report.md` and the pack into `public/reports//` only when made from the report.json as it is now and each is within the shared publish limit (`PUBLISH_MAX_FILE_BYTES`, 24 MiB, `lib/builtExport.ts`); a pack over it stays local. The view's `downloads` lists what was published and the report page's download line shows HTML · PDF · Markdown · Evidence pack · Citations JSON · CSV. `reports/` is allowed wholesale by the cited audit. ## History Readers can check that a report is the one published and see every edit, with hashes (slice RH). History is PER REPORT, never site-wide. - **Store** — a bare git repo per report beside its report.json: `sites//reports//history-git/` (`common/publish/reportHistory.ts`; never a path segment named `.git`). It is inside the corpus repo's tree; nothing writes a `.gitignore` for it — the operator adds `history-git/` there. One branch, `main`. - **Revisions** — `reports export` (`reportExports.ts` `writeReportExports`, given the store and the report.json's bytes) commits one when the sha256 of report.json differs from the newest revision's; an unchanged report commits nothing. A commit holds `report.json` (byte for byte), `report.md` (the Markdown export) and `exports.json` (format `archilyzer-report-revision-exports`: the sha256 and size of every export file made and of the citations JSON and CSV). Message: `Revision N`, a blank line, `- ` lines of `reportChangeSummary` (`common/lib/report/revisions.ts`, pure): title/series/subtitle changes, claims added/removed, verdict changes, claims edited (title, text, findings), citations added/removed, quotes edited, summary/method edited; the first revision counts what it holds; anything else is one "Other edits" line. Written with git plumbing (`hash-object`, `mktree`, `commit-tree --no-gpg-sign`, `update-ref` with the old value). - **Identity** — author and committer = the site's title, `noreply@.invalid`; `GIT_AUTHOR_DATE` / `GIT_COMMITTER_DATE` as `@ +0000`. git runs with `extendEnv: false` and a minimal env through `cleanGitEnv`: PATH, HOME/XDG_CONFIG_HOME at a path that does not exist, `GIT_CONFIG_NOSYSTEM=1`, `GIT_CONFIG_GLOBAL=/dev/null`, `TZ=UTC`, plus `-c commit.gpgSign=false -c core.hooksPath=/dev/null`. - **Footer and commit** — an export's footer is `Revision N · · report sha256 <12>`: N is the revision the export belongs to (the next one when report.json changed), computed before the files are written. The footer never names the commit: the commit holds report.md, whose footer would have to name the commit holding it. The history page and `history.json` map each revision's report sha256 to its commit; `export.json` records the revision, commit, date and summary once committed. - **Publish** — compose reads every revision into `reports//history/history.json` (`ReportHistoryView`, format `archilyzer-report-history`: per revision N, date, commit, report sha256, summary, and `claimChanges` — each changed claim's title/text/verdict/findings as a word diff of `eq`/`ins`/`del` runs) and stages a dumb-HTTP clone at `reports//history/repo/` as the source mirror does: a fresh `clone --bare --no-local --single-branch`, `repack -a -d --max-pack-size=20m`, `pack-refs --all`, `update-server-info`, then an allowlisted copy (`HEAD`, `packed-refs`, `info/refs`, `objects/info/packs`, `objects/pack/pack-*.{pack,idx}`, `refs/heads/main`). The page view carries `history: { revision, date, href, current }` (`current` false when report.json changed after the newest revision and was not exported again — the line then reads "Edited since revision N"). The sitemap lists the history page. - **Export site** — `/reports//history/` (static params as the report page's): `git clone /reports//history/repo` (the site-root path when `siteUrl` is unset), then each revision newest first with its hashes, summary and inline ``/`` diff. The report header's date line ends with "revision N", linked to the history ("edited since revision N" when the report changed after it). - **Audit** — `reportHistoryProblem` (`lib/builtExport.ts`, asked by `builtSiteProblem` and `builtBundleProblem` for every build) refuses any file in a `reports//history/repo/` outside the allowlist or over `PUBLISH_MAX_FILE_BYTES`; the cited audit counts the clone's files toward Pages' 20,000. - **Editor** — the Reports tab's Exports list shows each report's revision, commit and last change (from `export.json`). - **Tests** — `lib/report/revisions.test.ts` (summary, word diff, claim changes), `publish/reportHistory.test.ts` (commit only on change; site identity and UTC dates with an operator's GIT_*/EMAIL/TZ set; no operator string in any object; the staged clone served by `export/scripts/serve-out.mjs` and cloned over HTTP), compose and audit tests; the report-site e2e fixture has two revisions (`source/demo-factcheck/report.v1.json`, then `report.json`) and checks the page line, the history page's hashes, summary and diff, and a `git clone` over the e2e server. ## Compose and the contract - `compose-site` gains a reports stage: resolve citations against the shared transcript trees, verify quotes, write report + moment JSON, copy media and stills. - Search off (`search: false`) prunes everything corpus-shaped from `public/` (summaries, stats, transcripts, subs, posts, digests, archives, duplicates, tags, aliases, chart templates, service worker) — `public/` is shared by every site's build in turn, so stale data from the previous site must not survive — and an **out/ allowlist audit** refuses a cited build that contains anything outside reports/moments/media/stills/shell assets. - Always emit `site.json` (`channels: []` for a cited site) and `corpus.json` with `site.scope: "cited"`, `audience`, and `reports: { index: "/reports/index.json" }` — the private-deploy guard and compose's cross-site check read them. `CONTRACT.corpusSpec` → 5; readers tolerate a cited scope (an old reader sees an empty corpus, which is safe). `llms.txt` and the sitemap get a cited variant listing report routes. - Hub and homepage exclude `cited` sites (`isListedSite` folds them out, as a private site); MCP serves `list_reports` / `get_report` and says "cited-only site: N report(s)" (R5). - umtool's cue walk (`cues.mjs`) cannot resolve cues off a cited site — it refuses one by name (R5). ## Converters (to make reports from what exists) - umtool video manifest / ledger → report.json. - Sweep markdown (`/sweep` output: `[title @ mm:ss](/?v=…&t=…)`, `[post by …](url)`) → report.json, ends widened with `resolve-windows widen`. - `shoot-page` batch items from each claim's `sourceQuote`. - A bespoke report shape (any other generator) converts outside the repo and validates against the schema. ## Slices | Slice | What | Size | Needs | |---|---|---|---| | R0 | `common/lib/citations/` + `common/lib/report/`: schemas, validation, inline-cite extraction + numbering, moment keys, back-link index, shared verdict vocab; `CITATIONS.md` + `REPORT.md` generated | M | — | | R1 | site keys `search` / `reports` + site form control | S | — | | R2 | evidence media: cutter port to common, `reports prepare` CLI + job kind, cited-capture copier + guard amendment | L | R0 | | R3 | compose: reports stage, cited prune, allowlist audit, always-emit site/corpus, llms/sitemap variants | L | R0, R2 interface | | R4 | export app: routes, shared citation components (inline marker + preview, cards, reference list, verification badge), Reports header link, cited shell, placeholder params | L | R0 (fixture JSON) | | R5 | consumers: corpusSpec 5, reader tolerance, hub/homepage exclusion, builtExport audit hook | M | R0 | | R6 | converters: manifest ↔ report (both ways), sweep md → report, /ask answer → report, shoot-page items; citations JSON/CSV export | M | R0 | | R7 | editor Reports tab: list, problems, Prepare media, private build | M | R1, R2 | | R8 | e2e: a cited fixture site (own playwright config), compose + audit unit tests | M | R3, R4 | ``` wave 1: R0, R1 wave 2: R2, R4, R5, R6 (after R0) wave 3: R3 (after R2), R7 (after R1, R2) wave 4: R8 (after R3, R4) ``` ## Risks 1. Stale `public/` leaking corpus data into a cited build → explicit prune + out/ allowlist audit before anything is served or deployed. 2. A cited build without `site.json` / `corpus.json` silently disables the private-deploy guard → always emit. 3. Next 16 empty-params build failure → placeholder param. 4. Post privacy: only cited captures, the visibility rule applies, saved source pages are never published. 5. Moment JSON carries only cited fields and a bounded cue context. 6. Size: span cap, 24 MiB per-file refusal, 20k files. 7. Cue drift between a corpus and an earlier publish; platforms with two ids → validated at compose, loud failure. 8. Contract change for third-party readers → spec bump, safe degradation.