commit 171f0e934bc129c41b5546376bbfbaa83638ecf8
parent 57bd41f23d9f1394b0495d6278592fcf47240c0b
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Mon, 5 Oct 2026 03:39:18 -0400
Merge report-r3-compose (compose writes reports + moments with computed quote verification; a cited site prunes the corpus and always emits site.json/corpus.json (spec 5); the cited out/ allowlist audit guards build and deploy)
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Diffstat:
22 files changed, 2095 insertions(+), 37 deletions(-)
diff --git a/.gitignore b/.gitignore
@@ -77,6 +77,10 @@ yarn-error.log*
# per-site bulk-download archive zips + manifest (regenerate with `pnpm build`)
# (no trailing slash — may be a worktree symlink, as above)
/export/public/archives
+# a report site's reports, moment pages and cited media (compose-site's reports stage)
+/export/public/reports
+/export/public/m
+/export/public/media
# oversize archives staged for the deploy-time R2 upload
/export/.r2-staging/
/export/.export-index/
diff --git a/ENVIRONMENT.md b/ENVIRONMENT.md
@@ -132,6 +132,7 @@ The publish pipeline sets these for a process it spawns. Listed so a reader know
| `SITE_ID` | — | Which site a compose or an export build is for. `archilyzer build site <id>` sets it; `compose site` and `build site` fall back to it when no id is given. | common/bin/compose-site.ts, export/app/lib/site.ts |
| `INSTANCE_MODE` | a site | `hub` makes the export build the hub. Set by `archilyzer build hub`. | export/app/lib/mode.ts, common/lib/archive/contract.ts |
| `BUILD_ARCHIVES` | on | `0` skips archive-zip generation for one build (`--skip-archives`). | common/bin/compose-site.ts, common/bin/build-archives.ts |
+| `REPORTS_ALLOW_MISSING_MEDIA` | off | `1` lets a report citation whose evidence media was not prepared through compose (`--allow-missing-media`): its moment page renders without a clip. Off, compose fails with the list. | common/bin/compose-site.ts |
| `ARCHIVES_READONLY` | off | `1` inside a docker-mode build container: materialize archives, never write the shared cache. | common/bin/compose-site.ts |
| `HOMEPAGE_PUBLIC_DIR` | `<repo>/homepage/public` | Where `compose homepage` and `source publish` write. | common/bin/compose-homepage.ts, common/publish/source.ts |
diff --git a/PUBLISH.md b/PUBLISH.md
@@ -46,7 +46,7 @@ entry points in `common/publish/build.ts`.
| To | Editor | `pnpm ops` | `pnpm archilyzer …` |
|---|---|---|---|
-| Build one site | a site's **Publish** tab → *Build static export* | `build-site` | `build site <id> [--nodata] [--skip-archives]` |
+| Build one site | a site's **Publish** tab → *Build static export* | `build-site` | `build site <id> [--nodata] [--skip-archives] [--allow-missing-media]` |
| Deploy the built site | *Deploy to production* / *Deploy preview* | `deploy-site` | `deploy site <id> [--preview <branch>]` |
| Build, then deploy | *Build & deploy* | `build-deploy` | `build site <id>` then `deploy site <id>` |
| Build every site | /sites → **Build all sites** | — | `build all [--skip-archives]` |
@@ -74,6 +74,32 @@ is not on disk, a clip over 24 MiB or an invalid report is listed and fails the
(exit 1, or a failed job) — fetch the window or persist the video, capture the
post, and run it again; what is already cut is reused.
+The build's compose then writes the reports from what prepare left
+(`common/publish/composeReports.ts`): each report's page, its citations as
+`citations.json` and `citations.csv`, its cited stills (never a saved source copy),
+one page per cited moment with the transcript lines around it, and the prepared clips
+and captures — only those the published reports cite. Every quote is checked as it is
+composed: a span's against its cues within 5 s either side (a record whose `en` track
+has no cues is read from `en-orig`), a post's against its text, scored as the share of
+the quote's words found there; the score, the time and the method are written into the
+citation, replacing any typed by hand. Compose fails, before it writes any of it, with
+the list of everything wrong: an invalid report, a citation of a channel outside the
+site or a post the site may not carry, a missing record, still or post, a quote below
+60 %, and a citation whose media was not prepared, or was cut for a span the report no
+longer cites. `--allow-missing-media` (on `compose site` and `build site`) lets the
+last two through; their pages render without a clip.
+
+A **cited** site (`publish: "cited"`) publishes its reports and nothing else. Its
+compose removes every corpus-shaped file the shared `export/public` holds — summaries,
+stats, transcripts, subs, posts, digests, archives, the duplicates, tags, aliases and
+chart files, the service worker — writes `site.json` with no channels and `corpus.json`
+with `site.scope: "cited"` (spec 5), and an `llms.txt` and sitemap listing the reports
+and moment pages. After `next build`, the built `out/` is audited: a cited build that
+holds anything but `_next/`, the reports, the moment pages, the cited media, the
+shell's own pages and files, or a file over 25 MiB, or more than 20,000 files, fails
+the build, and every deploy path refuses it. A site switched to cited is refused at
+deploy until it is built again.
+
Deploy-only ships whatever is in `export/out`, which the basic build composes one
site at a time into a single shared directory — so it **refuses, before starting a
job, if `export/out` holds a build of another site** (or no build at all), naming the
diff --git a/common/bin/archilyzer.ts b/common/bin/archilyzer.ts
@@ -46,12 +46,17 @@ export const COMMANDS: Command[] = [
},
{
path: ["compose", "site"],
- usage: "<id> compose one site's export/public (default: SITE_ID)",
+ usage:
+ "<id> [--allow-missing-media] compose one site's export/public (default: SITE_ID); --allow-missing-media lets a report citation with no prepared media through",
+ flags: { "allow-missing-media": "boolean" },
maxPositionals: 1,
- run: async ({ positionals, env }) => {
+ run: async ({ positionals, flags, env }) => {
const siteId = siteIdFrom(positionals, env, "compose site");
if (!siteId) return 2;
- await (await import("./compose-site")).main({ siteId });
+ await (await import("./compose-site")).main({
+ siteId,
+ allowMissingMedia: flags["allow-missing-media"] === true,
+ });
return 0;
},
},
@@ -74,8 +79,8 @@ export const COMMANDS: Command[] = [
{
path: ["build", "site"],
usage:
- "<id> [--nodata] [--skip-archives] data phase + compose + next build into export/out (default id: SITE_ID)",
- flags: { nodata: "boolean", "skip-archives": "boolean" },
+ "<id> [--nodata] [--skip-archives] [--allow-missing-media] data phase + compose + next build into export/out (default id: SITE_ID)",
+ flags: { nodata: "boolean", "skip-archives": "boolean", "allow-missing-media": "boolean" },
maxPositionals: 1,
run: async ({ positionals, flags, env }) => {
const siteId = siteIdFrom(positionals, env, "build site");
@@ -95,6 +100,7 @@ export const COMMANDS: Command[] = [
signal: interrupted(),
skipData: flags.nodata === true,
skipArchives: flags["skip-archives"] === true,
+ allowMissingMedia: flags["allow-missing-media"] === true,
});
if (code !== 0) console.error(`build site ${siteId}: failed (exit ${code})`);
return code;
diff --git a/common/bin/compose-hub.ts b/common/bin/compose-hub.ts
@@ -97,6 +97,10 @@ export const SITE_ONLY_PUBLIC_ENTRIES: readonly string[] = [
"digests",
"stats",
"archives",
+ // A report site's reports, moments and cited media (publish/composeReports.ts).
+ "reports",
+ "m",
+ "media",
"site.json",
"tags.json",
"duplicates.json",
diff --git a/common/bin/compose-site.ts b/common/bin/compose-site.ts
@@ -9,9 +9,17 @@
// public/transcripts/<slug>/ <- shared/transcripts/<slug>/ (site's members)
// public/stats/* <- sites/<id>/stats/ (per-site)
// public/chart-templates.json <- sites/<id>/chart-templates.json (per-site)
+// public/reports/, m/, media/ <- the site's reports, resolved against the
+// corpus (publish/composeReports.ts)
//
// Only the site's member channels are composed, so each deployed bundle holds
// only that site's data. Checked-in static assets in public/ are left intact.
+//
+// A CITED site (site.json `publish: "cited"`) publishes its reports and the
+// moments they cite, and nothing else: its compose removes every corpus-shaped
+// entry (CORPUS_PUBLIC_ENTRIES) instead of composing it, then writes the
+// reports, site.json (no channels) and corpus.json (`site.scope: "cited"`) —
+// see composeCitedSite.
import path from "node:path";
import { cp, link, mkdir, rm, readdir, access, readFile, writeFile, stat, rename } from "node:fs/promises";
@@ -63,7 +71,13 @@ import {
import { readChannelConfig } from "../controller/channels";
import { builtSiteIdIn } from "../lib/builtExport";
import { publishedMemberSlugs } from "../lib/postsVisibility";
-import { isPrivateSite } from "../lib/siteSchema";
+import { isCitedSite, isPrivateSite } from "../lib/siteSchema";
+import { MANIFEST_VERSION, SUMMARIES_PAGE_SIZE } from "../lib/manifest";
+import {
+ composeReports,
+ reportRoutes,
+ type ComposedReports,
+} from "../publish/composeReports";
import { runIfEntryPoint } from "./_cli";
import { copyPublicFile, ownDir, writePublicFile } from "./_publicFile";
@@ -84,11 +98,17 @@ async function emitFederationFiles(
renderHeadersFile("compose-site.ts"),
);
+ // A cited site has no summaries: its descriptor names no channel, and it is
+ // ALWAYS written — the deploy guards read the bundle's identity from it.
+ const cited = isCitedSite(site);
const manifestPath = path.join(paths.exportSummariesDir, "manifest.json");
- if (!(await exists(manifestPath))) return; // no composed data → no descriptor
- const manifest = JSON.parse(await readFile(manifestPath, "utf8")) as Manifest;
+ if (!cited && !(await exists(manifestPath))) return; // no composed data → no descriptor
+ const manifest = cited
+ ? emptyManifest(site)
+ : (JSON.parse(await readFile(manifestPath, "utf8")) as Manifest);
const descriptor = buildSiteDescriptor(site, manifest, resolveSocialLinks(site), {
- pwa: shipsPwa(site),
+ // A cited site is a handful of pages: nothing to install or cache offline.
+ pwa: !cited && shipsPwa(site),
hubUrl: resolveHubUrl(site),
});
await writePublicFile(
@@ -97,6 +117,89 @@ async function emitFederationFiles(
);
}
+// A cited site's stand-in for the summaries manifest its descriptor would be
+// built from: no channels, fresh as of this compose.
+function emptyManifest(site: Site): Manifest {
+ return {
+ version: MANIFEST_VERSION,
+ totalCount: 0,
+ pageSize: SUMMARIES_PAGE_SIZE,
+ pageCount: 0,
+ generatedAt: new Date().toISOString(),
+ channels: [],
+ groups: site.groups,
+ defaultGroupId: site.defaultGroupId,
+ siteId: site.siteId,
+ };
+}
+
+// Everything a full compose writes into public/ that is the CORPUS — what a
+// cited site must not ship, whichever site's compose left it there. Removed
+// by composeCitedSite before it writes anything; lib/builtExport.ts
+// citedBuildProblem audits the built out/ against the same promise. A new
+// corpus-shaped entry in this file's full compose belongs here too.
+export const CORPUS_PUBLIC_ENTRIES: readonly string[] = [
+ "summaries",
+ "stats",
+ "transcripts",
+ "subs",
+ "posts",
+ "digests",
+ "archives",
+ "chart-templates.json",
+ "search-aliases.json",
+ TAGS_FILENAME,
+ DUPLICATES_FILENAME,
+ "sw.js",
+ // The hub's, which a site never ships.
+ "hub-sites.json",
+ "hub-summary.json",
+ // Rewritten for the cited site below.
+ "corpus.json",
+ "llms.txt",
+ "robots.txt",
+ "sitemap.xml",
+];
+
+// Compose a CITED site: prune the corpus, then its reports and its contract.
+async function composeCitedSite(
+ site: Site,
+ paths: ReturnType<typeof getPaths>,
+ cachePath: string,
+ opts: { allowMissingMedia: boolean },
+): Promise<void> {
+ for (const entry of CORPUS_PUBLIC_ENTRIES) {
+ await rm(path.join(paths.exportPublicDir, entry), { recursive: true, force: true });
+ }
+ // …and an earlier full compose's staged oversize archives, which the deploy
+ // would otherwise upload under this site's name.
+ await rm(path.join(path.dirname(paths.exportPublicDir), ".r2-staging", site.siteId), {
+ recursive: true,
+ force: true,
+ });
+ console.log(`[compose] ${site.siteId} publishes only its reports: the corpus is not composed.`);
+
+ const composed = await composeReports({
+ paths,
+ site,
+ allowMissingMedia: opts.allowMissingMedia,
+ log: console.log,
+ });
+ await emitFederationFiles(site, paths);
+ await emitAiFiles(site, paths, composed);
+ // Nothing of the corpus is in public/ now: the next full compose of this
+ // site composes every stage afresh.
+ await writeComposeCache(cachePath, { transcripts: {}, subs: {}, posts: {}, digests: {} });
+ console.log(
+ `Composed cited site "${site.siteId}" into ${paths.exportPublicDir} ` +
+ `(${composed.reports.length} report(s), ${composed.moments.length} moment(s)).`,
+ );
+}
+
+// `--allow-missing-media` (archilyzer compose site / build site), or the
+// environment variable the build passes it to compose as.
+export const ALLOW_MISSING_MEDIA_ENV = "REPORTS_ALLOW_MISSING_MEDIA";
+
// Emit the AI-discovery surface — a small FIXED set of site-root files
// (llms.txt, corpus.json, robots.txt, sitemap.xml). These document how to
// navigate the already-served paginated shards; they never enumerate per-video
@@ -105,6 +208,7 @@ async function emitFederationFiles(
async function emitAiFiles(
site: Site,
paths: ReturnType<typeof getPaths>,
+ composed: ComposedReports,
): Promise<void> {
const sitePath = path.join(paths.exportPublicDir, "site.json");
if (!(await exists(sitePath))) return; // no composed data → nothing to describe
@@ -171,6 +275,8 @@ async function emitAiFiles(
// A private site says so in its own corpus.json — the bundle's word that
// the deploy guard (lib/builtExport.ts builtAudienceProblem) reads.
private: isPrivateSite(site),
+ reportCount: composed.reports.length,
+ cited: isCitedSite(site),
});
await writePublicFile(
path.join(paths.exportPublicDir, "corpus.json"),
@@ -178,7 +284,7 @@ async function emitAiFiles(
);
await writePublicFile(
path.join(paths.exportPublicDir, "llms.txt"),
- renderSiteLlmsTxt(corpus),
+ renderSiteLlmsTxt(corpus, { reports: composed.reports }),
);
await writePublicFile(
path.join(paths.exportPublicDir, "robots.txt"),
@@ -189,11 +295,14 @@ async function emitAiFiles(
// siteUrl; otherwise clear any stale copy from a previous build.
const sitemapPath = path.join(paths.exportPublicDir, "sitemap.xml");
if (descriptor.siteUrl) {
- const routes = ["/", "/changelog"];
+ // A cited site's home IS its report index; a full site's reports follow
+ // its own routes.
+ const routes = isCitedSite(site) ? ["/"] : ["/", "/changelog"];
if (hasArchives) routes.push("/downloads");
if (await exists(path.join(paths.exportPublicDir, DUPLICATES_FILENAME))) {
routes.push("/duplicates");
}
+ routes.push(...reportRoutes(composed).filter((r) => !routes.includes(r)));
await writePublicFile(
sitemapPath,
renderSitemapXml({ siteUrl: descriptor.siteUrl, routes }),
@@ -695,7 +804,7 @@ export async function reconcileChannelTree(
// `archilyzer compose site <id>` passes the id; the bare bin (export's
// compose:site script) reads SITE_ID, as it always has.
export async function main(
- opts: { siteId?: string; paths?: Paths } = {},
+ opts: { siteId?: string; paths?: Paths; allowMissingMedia?: boolean } = {},
): Promise<void> {
const siteId = opts.siteId?.trim() || process.env.SITE_ID;
if (!siteId) {
@@ -756,6 +865,13 @@ export async function main(
// short part-way leaves the next one nothing to trust.
await rm(path.join(paths.exportPublicDir, "site.json"), { force: true });
+ const allowMissingMedia =
+ opts.allowMissingMedia === true || process.env[ALLOW_MISSING_MEDIA_ENV] === "1";
+ if (isCitedSite(site)) {
+ await composeCitedSite(site, paths, cachePath, { allowMissingMedia });
+ return;
+ }
+
// --- per-site aggregates (whole-dir swaps), gated on the source signature ---
const summariesSrc = path.join(paths.exportSitesIndexDir, siteId, "summaries");
const summariesSig = await dirSignature(summariesSrc);
@@ -1037,6 +1153,11 @@ export async function main(
cache.duplicates = `${wrote ? "written" : "empty"}|${dupKey}`;
}
+ // --- reports: the site's published reports and the moments they cite ---
+ // (a site with none: whatever an earlier compose left is removed). Before
+ // site.json: a stage that fails leaves public/ naming no site.
+ const composed = await composeReports({ paths, site, allowMissingMedia, log: console.log });
+
// --- federation contract: /site.json descriptor + CORS _headers ---
await emitFederationFiles(site, paths);
// A site's bundle is not a hub's. export/public is shared with the hub build,
@@ -1054,8 +1175,8 @@ export async function main(
await composeArchives(site, memberSlugs, paths);
// --- AI discovery: llms.txt / corpus.json / robots.txt / sitemap.xml ---
- // (after site.json + archives — both feed into these fixed-count files)
- await emitAiFiles(site, paths);
+ // (after site.json + archives + reports — all feed into these fixed-count files)
+ await emitAiFiles(site, paths, composed);
// Persist the incremental-compose signatures for the next build.
await writeComposeCache(cachePath, cache);
diff --git a/common/controller/curatedTagsBuild.test.ts b/common/controller/curatedTagsBuild.test.ts
@@ -217,7 +217,7 @@ test("compose 1: the site publishes /tags.json and corpus.json points at it", as
const corpus = JSON.parse(
readFileSync(path.join(paths.exportPublicDir, "corpus.json"), "utf8"),
);
- assert.equal(corpus.spec, 4);
+ assert.equal(corpus.spec, 5);
assert.equal(corpus.tags.videoField, "curatedTags");
assert.match(corpus.tags.url, /\/tags\.json$/);
assert.match(
diff --git a/common/lib/archive/contract.test.ts b/common/lib/archive/contract.test.ts
@@ -164,7 +164,7 @@ test("shipsPwa: the site flag, or hub mode", () => {
test("CONTRACT is frozen where it is published", () => {
// These are on the wire. A change here is a change to every deployed archive's
// machine contract, so it belongs in a slice that says so, not in a refactor.
- assert.equal(CONTRACT.corpusSpec, 4);
+ assert.equal(CONTRACT.corpusSpec, 5);
assert.equal(CONTRACT.siteDescriptor, 1);
assert.equal(CONTRACT.manifest, 3);
assert.equal(CONTRACT.transcriptsManifest, 1);
diff --git a/common/lib/archive/contract.ts b/common/lib/archive/contract.ts
@@ -31,8 +31,10 @@ import { TAGS_FILENAME } from "../curatedTags";
// per-field notes in lib/corpus.ts for what a bump means to a client.
export const CONTRACT = {
// /corpus.json's own `spec`. `generator` is deliberately unversioned; see the
- // note in lib/corpus.ts for why a credit line did not bump the spec.
- corpusSpec: 4,
+ // note in lib/corpus.ts for why a credit line did not bump the spec. 5: a
+ // site may publish reports (`reports`), and a CITED site (`site.scope:
+ // "cited"`) publishes only them — no channels, no shards.
+ corpusSpec: 5,
// /site.json's `contract` (siteDescriptor.ts).
siteDescriptor: 1,
// The four manifest versions (manifest.ts). Each is the version field of one
diff --git a/common/lib/builtExport.test.ts b/common/lib/builtExport.test.ts
@@ -8,8 +8,12 @@ import {
builtHomepageAt,
builtHomepageProblem,
builtHubProblem,
+ builtScopeProblem,
builtSiteIdIn,
builtSiteProblem,
+ citedBuildProblem,
+ deployAudienceProblem,
+ PAGES_MAX_FILES,
} from "./builtExport";
function tempOut(siteJson?: string): { dir: string; cleanup: () => void } {
@@ -256,3 +260,133 @@ test("builtSiteProblem refuses exactly what builtBundleProblem refuses", () => {
torn.cleanup();
}
});
+
+// ─── The cited out/ audit ───
+
+// A cited build as next build leaves it: compose's files, Next's shell, the
+// reports, moments and cited media.
+function citedOut(extra: string[] = []): { dir: string; cleanup: () => void } {
+ const t = bundle({
+ site: { siteId: "reports-site", channels: [] },
+ corpus: { spec: 5, kind: "site", site: { id: "reports-site", scope: "cited", audience: "public" } },
+ });
+ const files = [
+ "_next/static/chunks/app.js",
+ "_not-found/index.html",
+ "404/index.html",
+ "404.html",
+ "index.html",
+ "index.txt",
+ "__next._tree.txt",
+ "__next.!KHdvcmtzcGFjZSk.__PAGE__.txt",
+ "ask/index.html",
+ "changelog/index.html",
+ "downloads/index.html",
+ "duplicates/index.html",
+ "offline/index.html",
+ "reports/index.json",
+ "reports/index.html",
+ "reports/_none/index.html",
+ "reports/demo-report/index.html",
+ "reports/demo-report/page.json",
+ "reports/demo-report/stills/a01.png",
+ "m/index.json",
+ "m/_none/index.html",
+ "m/demo-channel/abc123/10.00-20.00/index.html",
+ "media/clips/demo-channel/abc123/10.00-20.00.mp4",
+ "media/posts/demo-social/3kabc/shot.png",
+ "icons/icon-192.png",
+ "favicon.ico",
+ "manifest.webmanifest",
+ "llms.txt",
+ "robots.txt",
+ "sitemap.xml",
+ "_headers",
+ "next.svg",
+ ...extra,
+ ];
+ for (const f of files) {
+ mkdirSync(path.dirname(path.join(t.dir, f)), { recursive: true });
+ writeFileSync(path.join(t.dir, f), "x");
+ }
+ return t;
+}
+
+test("a cited build holding only reports, moments, cited media and the shell passes, and deploys", () => {
+ const t = citedOut();
+ try {
+ assert.equal(citedBuildProblem(t.dir), null);
+ assert.equal(builtBundleProblem(t.dir, "reports-site"), null);
+ assert.equal(builtSiteProblem(t.dir, "reports-site"), null);
+ } finally {
+ t.cleanup();
+ }
+});
+
+test("a cited build holding anything corpus-shaped is refused — by the audit and by both deploy checks", () => {
+ for (const extra of [
+ "summaries/manifest.json",
+ "transcripts/demo-channel/page-0000.json",
+ "posts/demo-social/manifest.json",
+ "archives/manifest.json",
+ "search-aliases.json",
+ "tags.json",
+ "sw.js",
+ "media/thumbnails/a.jpg",
+ "media/stray.mp4",
+ ]) {
+ const t = citedOut([extra]);
+ try {
+ const problem = citedBuildProblem(t.dir);
+ assert.ok(problem, extra);
+ assert.match(problem, /is a cited build, which publishes only reports and their moments, but it also holds/);
+ assert.ok(problem.includes(extra.startsWith("media/") ? extra.split("/").slice(0, 2).join("/") : extra.split("/")[0]), `${extra}: ${problem}`);
+ assert.equal(builtBundleProblem(t.dir, "reports-site"), problem, extra);
+ assert.equal(builtSiteProblem(t.dir, "reports-site"), problem, extra);
+ } finally {
+ t.cleanup();
+ }
+ }
+});
+
+test("the audit leaves a full build alone, whatever it holds", () => {
+ const t = bundle({ site: { siteId: "anilyzer" }, corpus: { spec: 5, site: { id: "anilyzer" } } });
+ mkdirSync(path.join(t.dir, "transcripts", "x"), { recursive: true });
+ try {
+ assert.equal(citedBuildProblem(t.dir), null);
+ assert.equal(builtBundleProblem(t.dir, "anilyzer"), null);
+ } finally {
+ t.cleanup();
+ }
+});
+
+test("a cited build over the Pages file limit is refused", () => {
+ const t = citedOut();
+ try {
+ const dir = path.join(t.dir, "m", "many");
+ mkdirSync(dir, { recursive: true });
+ for (let i = 0; i <= PAGES_MAX_FILES; i++) writeFileSync(path.join(dir, String(i)), "");
+ assert.match(citedBuildProblem(t.dir)!, /over Pages' limit of 20000/);
+ } finally {
+ t.cleanup();
+ }
+});
+
+test("a site configured cited with a full build is refused at deploy; a cited build of it is not", () => {
+ const full = bundle({ site: { siteId: "reports-site" }, corpus: { spec: 5, site: { id: "reports-site" } } });
+ const cited = citedOut();
+ try {
+ const site = { siteId: "reports-site", publish: "cited" };
+ assert.match(builtScopeProblem(site, full.dir)!, /publishes only its reports \(publish: cited\), but .* holds a full build/);
+ assert.equal(deployAudienceProblem(site, full.dir), builtScopeProblem(site, full.dir));
+ assert.equal(builtScopeProblem(site, cited.dir), null);
+ assert.equal(builtScopeProblem({ siteId: "reports-site" }, full.dir), null);
+ // Nothing built: the identity checks answer that, not this one.
+ const empty = tempOut();
+ assert.equal(builtScopeProblem(site, empty.dir), null);
+ empty.cleanup();
+ } finally {
+ full.cleanup();
+ cited.cleanup();
+ }
+});
diff --git a/common/lib/builtExport.ts b/common/lib/builtExport.ts
@@ -13,7 +13,7 @@
// public/ into out/. So the check is a file read, and it is cheap enough to do
// before every deploy.
-import { existsSync, readFileSync, statSync } from "node:fs";
+import { existsSync, readdirSync, readFileSync, statSync, type Dirent } from "node:fs";
import path from "node:path";
/**
@@ -64,7 +64,7 @@ export function builtSiteProblem(outDir: string, siteId: string): string | null
if (corpusSiteIdIn(outDir) !== asked) {
return `export/out holds an incomplete build of "${asked}" (its corpus.json does not name it) — build ${asked} first`;
}
- return null;
+ return citedBuildProblem(outDir);
}
/**
@@ -95,7 +95,7 @@ export function builtBundleProblem(outDir: string, siteId: string): string | nul
if (described !== asked) {
return `${outDir} describes "${described}", not "${asked}" (corpus.json)`;
}
- return null;
+ return citedBuildProblem(outDir);
}
/**
@@ -149,10 +149,172 @@ export function builtAudienceProblem(outDir: string): string | null {
* logs or throws, or null.
*/
export function deployAudienceProblem(
- site: { siteId: string; audience?: string },
+ site: { siteId: string; audience?: string; publish?: string },
+ outDir: string,
+): string | null {
+ return siteDeployProblem(site) ?? builtAudienceProblem(outDir) ?? builtScopeProblem(site, outDir);
+}
+
+// ─── A CITED build (site.json `publish: "cited"`) ───
+//
+// A cited site publishes its reports and the moments they cite, and nothing
+// else (plans/report-sites.md). export/public is shared by every site's build
+// in turn, so compose prunes everything corpus-shaped before it writes a cited
+// site — and this audit is the second line: after `next build`, a cited out/
+// may hold only what the list below names. Anything else (a stale summaries
+// tree, a transcripts shard, an archive, a service worker) refuses the build
+// and every deploy of it: builtSiteProblem and builtBundleProblem ask it, so
+// runDeployIntoLog, the container deploy phase, deploySite, the editor's
+// deploy actions and docker/build-site.sh all refuse what it refuses, and
+// `archilyzer build site` fails on it (publish/build.ts runBuildPhase).
+//
+// The bundle names its own scope — corpus.json's `site.scope: "cited"`, which
+// compose always writes for a cited site — so the audit needs no site config.
+// A site configured cited whose out/ is a FULL build is builtScopeProblem's.
+
+// What a cited out/ may hold at its top level.
+export const CITED_OUT_ALLOWED_DIRS: readonly string[] = [
+ // Next's assets, and the routes a cited build still renders (the shell's
+ // pages that say "not on this site", the not-found page).
+ "_next",
+ "_not-found",
+ "404",
+ "ask",
+ "changelog",
+ "downloads",
+ "duplicates",
+ "offline",
+ // What compose's reports stage writes: the reports and their stills, the
+ // moment pages, the cited media (with each route's placeholder, `_none`).
+ "reports",
+ "m",
+ "media",
+ "icons",
+];
+
+export const CITED_OUT_ALLOWED_FILES: readonly string[] = [
+ "site.json",
+ "corpus.json",
+ "llms.txt",
+ "robots.txt",
+ "sitemap.xml",
+ "_headers",
+ "index.html",
+ "index.txt",
+ "404.html",
+ "favicon.ico",
+ "manifest.webmanifest",
+ // export/public's checked-in assets.
+ "file.svg",
+ "globe.svg",
+ "next.svg",
+ "vercel.svg",
+ "window.svg",
+];
+
+// Next's per-segment payloads for the root route (`__next._tree.txt`, …).
+const NEXT_SEGMENT_FILE_RE = /^__next\..+\.txt$/;
+
+// What `media/` may hold: the cited clips and post captures, nothing else.
+export const CITED_MEDIA_ALLOWED_DIRS: readonly string[] = ["clips", "posts"];
+
+// Cloudflare Pages' own limits: a file of at most 25 MiB, at most 20,000 files.
+export const PAGES_MAX_FILE_BYTES = 25 * 1024 * 1024;
+export const PAGES_MAX_FILES = 20_000;
+
+// corpus.json's `site.scope`, or null.
+function builtScopeIn(outDir: string): string | null {
+ try {
+ const parsed: unknown = JSON.parse(readFileSync(path.join(outDir, "corpus.json"), "utf8"));
+ const scope = (parsed as { site?: { scope?: unknown } } | null)?.site?.scope;
+ return typeof scope === "string" ? scope : null;
+ } catch {
+ return null;
+ }
+}
+
+// Whether the build in `outDir` is a cited site's (its corpus.json says so).
+export function isCitedBuild(outDir: string): boolean {
+ return builtScopeIn(outDir) === "cited";
+}
+
+/**
+ * Why the CITED build in `outDir` may not ship, as one sentence naming what
+ * it holds that a cited site may not — or null when it may, or when the build
+ * is not a cited one.
+ */
+export function citedBuildProblem(outDir: string): string | null {
+ if (!isCitedBuild(outDir)) return null;
+ const extra: string[] = [];
+ let entries: Dirent[];
+ try {
+ entries = readdirSync(outDir, { withFileTypes: true });
+ } catch {
+ return `${outDir} is a cited build that cannot be read`;
+ }
+ for (const e of entries) {
+ if (e.isDirectory()) {
+ if (!CITED_OUT_ALLOWED_DIRS.includes(e.name)) extra.push(`${e.name}/`);
+ } else if (!CITED_OUT_ALLOWED_FILES.includes(e.name) && !NEXT_SEGMENT_FILE_RE.test(e.name)) {
+ extra.push(e.name);
+ }
+ }
+ const mediaDir = path.join(outDir, "media");
+ if (existsSync(mediaDir)) {
+ for (const e of readdirSync(mediaDir, { withFileTypes: true })) {
+ if (!e.isDirectory() || !CITED_MEDIA_ALLOWED_DIRS.includes(e.name)) {
+ extra.push(`media/${e.name}${e.isDirectory() ? "/" : ""}`);
+ }
+ }
+ }
+ if (extra.length > 0) {
+ extra.sort();
+ const shown = extra.slice(0, 12).join(", ") + (extra.length > 12 ? `, … (${extra.length} in all)` : "");
+ return (
+ `${outDir} is a cited build, which publishes only reports and their moments, ` +
+ `but it also holds ${shown} — compose the site again`
+ );
+ }
+ // The Pages limits, which a cited build is small enough to walk for.
+ let files = 0;
+ const oversize: string[] = [];
+ const walk = (dir: string, rel: string): void => {
+ for (const e of readdirSync(dir, { withFileTypes: true })) {
+ const p = path.join(dir, e.name);
+ const r = rel ? `${rel}/${e.name}` : e.name;
+ if (e.isDirectory()) walk(p, r);
+ else {
+ files++;
+ if (statSync(p).size > PAGES_MAX_FILE_BYTES) oversize.push(r);
+ }
+ }
+ };
+ walk(outDir, "");
+ if (oversize.length > 0) {
+ return `${outDir} holds ${oversize.length} file(s) over Pages' 25 MiB limit: ${oversize.slice(0, 5).join(", ")}`;
+ }
+ if (files > PAGES_MAX_FILES) {
+ return `${outDir} holds ${files} files, over Pages' limit of ${PAGES_MAX_FILES}`;
+ }
+ return null;
+}
+
+/**
+ * Why `outDir` may not be deployed as `site` because of what it PUBLISHES, as
+ * one sentence, or null: a site configured cited (`publish: "cited"`) whose
+ * build is not a cited one was built before the switch, and would ship the
+ * whole corpus. Asked beside the audience refusals (deployAudienceProblem).
+ */
+export function builtScopeProblem(
+ site: { siteId: string; publish?: string },
outDir: string,
): string | null {
- return siteDeployProblem(site) ?? builtAudienceProblem(outDir);
+ if (site.publish !== "cited" || !existsSync(path.join(outDir, "corpus.json"))) return null;
+ if (isCitedBuild(outDir)) return null;
+ return (
+ `Site "${site.siteId}" publishes only its reports (publish: cited), but ${outDir} ` +
+ `holds a full build — build ${site.siteId} again, then deploy`
+ );
}
// corpus.json's `site.id`, or null when there is no readable one.
@@ -200,7 +362,18 @@ export function builtHubProblem(outDir: string): string | null {
// The per-site data trees a hub bundle must never carry (the trees of
// compose-hub's SITE_ONLY_PUBLIC_ENTRIES).
-const HUB_FORBIDDEN_TREES = ["summaries", "transcripts", "subs", "posts", "digests", "stats", "archives"];
+const HUB_FORBIDDEN_TREES = [
+ "summaries",
+ "transcripts",
+ "subs",
+ "posts",
+ "digests",
+ "stats",
+ "archives",
+ "reports",
+ "m",
+ "media",
+];
/**
* Why `outDir` — the homepage package's `homepage/out` — may not be deployed as
diff --git a/common/lib/citations/verify.test.ts b/common/lib/citations/verify.test.ts
@@ -0,0 +1,49 @@
+import { test } from "node:test";
+import assert from "node:assert/strict";
+import {
+ QUOTE_CHECK_METHOD,
+ QUOTE_DRIFT_THRESHOLD,
+ cueWindowText,
+ quoteDrifted,
+ quoteTokens,
+ quoteVerification,
+ tokenRecall,
+} from "./verify";
+
+test("tokens: lowercased, punctuation stripped, split on whitespace", () => {
+ assert.deepEqual(quoteTokens("Don't stop — the U.S. bridge, OK?"), ["dont", "stop", "the", "us", "bridge", "ok"]);
+ assert.deepEqual(quoteTokens("Café 2019\nnew"), ["café", "2019", "new"]);
+ assert.deepEqual(quoteTokens("…"), []);
+});
+
+test("recall: a verbatim quote scores 1 however much more the window says", () => {
+ assert.equal(tokenRecall("I was there for it.", "The bridge opened in spring, I was there for it. Anyway."), 1);
+});
+
+test("recall: each window token is used once; missing tokens lower the score", () => {
+ assert.equal(tokenRecall("the the the", "the"), 1 / 3);
+ assert.equal(tokenRecall("he opened the bridge himself", "the bridge opened"), 3 / 5);
+ assert.equal(tokenRecall("", "anything"), 0);
+ assert.equal(tokenRecall("words", ""), 0);
+});
+
+test("the cue window takes every cue overlapping the span ± 5 s", () => {
+ const cues = [
+ { start: 0, end: 4, text: "too early" },
+ { start: 4, end: 6, text: "edge before" },
+ { start: 10, end: 20, text: "inside" },
+ { start: 24, end: 26, text: "edge after" },
+ { start: 26, end: 30, text: "too late" },
+ ];
+ assert.equal(cueWindowText(cues, 10, 20), "edge before inside edge after");
+ assert.equal(cueWindowText(cues, 10, 20, 0), "inside");
+});
+
+test("the block: rounded score, the time, the method; drift below the threshold", () => {
+ const v = quoteVerification("he opened the bridge himself", "the bridge opened", "2026-10-05T12:00:00.000Z");
+ assert.deepEqual(v, { quoteScore: 0.6, quoteCheckedAt: "2026-10-05T12:00:00.000Z", method: QUOTE_CHECK_METHOD });
+ assert.equal(quoteDrifted(v), false);
+ assert.equal(quoteDrifted({ quoteScore: QUOTE_DRIFT_THRESHOLD - 0.01 }), true);
+ assert.equal(quoteDrifted({}), true);
+ assert.ok(!("voiceChecked" in v), "a voice is never vouched for");
+});
diff --git a/common/lib/citations/verify.ts b/common/lib/citations/verify.ts
@@ -0,0 +1,97 @@
+// QUOTE VERIFICATION — how compose checks that a citation's `quote` is what
+// the record says at the cited place, and the block it writes.
+//
+// THE METHOD (QUOTE_CHECK_METHOD, written into every block so a reader knows
+// what the number means):
+//
+// 1. The record's text at the cited place: for a span, the text of every cue
+// overlapping [start − 5 s, end + 5 s] (QUOTE_WINDOW_SLACK_SECONDS) — a cue
+// boundary is where a caption line wrapped, so a quote may run a little
+// past the span; for a post, the post's text.
+// 2. Both are normalised the same way: lowercased (Unicode-aware), every
+// character that is not a letter, a digit or whitespace removed (so
+// "don't" is "dont" and "U.S." is "us"), split on whitespace.
+// 3. The score is TOKEN RECALL: the share of the quote's tokens found in the
+// record's tokens, each record token used at most once (a multiset). A
+// verbatim quote scores 1 however much more the window says; a quote
+// with no tokens scores 0.
+//
+// The score is rounded to two decimals. Below QUOTE_DRIFT_THRESHOLD the quote
+// has DRIFTED from the record (a paraphrase, the wrong span, cues that moved
+// since the quote was taken) and compose fails on it — CITATIONS.md promises
+// a quote is verbatim, and this is the check that holds it to that.
+//
+// The block is COMPUTED: compose overwrites whatever a document carried
+// (lib/citations/schema.ts verificationSchema), and nothing here can vouch for
+// a voice, so `voiceChecked` is never written.
+//
+// Pure, no imports but types: the export site and the browser can use it.
+
+import type { CitationVerification } from "./schema";
+
+export const QUOTE_WINDOW_SLACK_SECONDS = 5;
+
+export const QUOTE_DRIFT_THRESHOLD = 0.6;
+
+export const QUOTE_CHECK_METHOD =
+ "token recall v1: quote vs the cues within ±5 s of the span (a post: its text); lowercase, punctuation stripped";
+
+type TimedText = { start: number; end: number; text: string };
+
+// A text's comparison tokens (step 2 above).
+export function quoteTokens(text: string): string[] {
+ return text
+ .toLowerCase()
+ .replace(/[^\p{L}\p{N}\s]/gu, "")
+ .split(/\s+/)
+ .filter(Boolean);
+}
+
+// The share of `quote`'s tokens found in `text` (step 3), unrounded.
+export function tokenRecall(quote: string, text: string): number {
+ const want = quoteTokens(quote);
+ if (want.length === 0) return 0;
+ const have = new Map<string, number>();
+ for (const t of quoteTokens(text)) have.set(t, (have.get(t) ?? 0) + 1);
+ let found = 0;
+ for (const t of want) {
+ const n = have.get(t) ?? 0;
+ if (n > 0) {
+ found++;
+ have.set(t, n - 1);
+ }
+ }
+ return found / want.length;
+}
+
+// The text of every cue overlapping [start − slack, end + slack], in order.
+export function cueWindowText(
+ cues: readonly TimedText[],
+ start: number,
+ end: number,
+ slack = QUOTE_WINDOW_SLACK_SECONDS,
+): string {
+ const from = start - slack;
+ const to = end + slack;
+ return cues
+ .filter((c) => c.end > from && c.start < to)
+ .map((c) => c.text)
+ .join(" ");
+}
+
+export function roundScore(score: number): number {
+ return Math.round(score * 100) / 100;
+}
+
+// The verification block for a quote checked against `text` at `checkedAt`.
+export function quoteVerification(quote: string, text: string, checkedAt: string): CitationVerification {
+ return {
+ quoteScore: roundScore(tokenRecall(quote, text)),
+ quoteCheckedAt: checkedAt,
+ method: QUOTE_CHECK_METHOD,
+ };
+}
+
+export function quoteDrifted(v: Pick<CitationVerification, "quoteScore">): boolean {
+ return (v.quoteScore ?? 0) < QUOTE_DRIFT_THRESHOLD;
+}
diff --git a/common/lib/corpus.test.ts b/common/lib/corpus.test.ts
@@ -190,14 +190,15 @@ test("Use with AI is the homepage's AI and MCP doc; the chat is the instance's /
}
});
-test("the spec is 4, and `generator` is still not why", () => {
+test("the spec is 5, and `generator` is still not why", () => {
// Guard on the reasoning, not just the number: a bump announces a new
- // FETCHABLE document. spec 4 is /tags.json. The informational credit string
+ // FETCHABLE document. spec 4 is /tags.json, spec 5 the reports index (and a
+ // cited site that publishes only reports). The informational credit string
// added at spec 3's time breaks no reader and did not bump anything — if
// either assertion is ever updated, the version-history comment in corpus.ts
// must justify why.
- assert.equal(CORPUS_SPEC_VERSION, 4);
- assert.equal(buildSiteCorpus(descriptor(), { hasArchives: false }).spec, 4);
+ assert.equal(CORPUS_SPEC_VERSION, 5);
+ assert.equal(buildSiteCorpus(descriptor(), { hasArchives: false }).spec, 5);
});
test("the tags pointer is present only when the site published one", () => {
diff --git a/common/lib/corpus.ts b/common/lib/corpus.ts
@@ -1,6 +1,7 @@
import type { PublicSiteDescriptor } from "./siteDescriptor";
import { AI_DOC_URL, PROJECT_GENERATOR } from "./project";
import { TAGS_FILENAME } from "./curatedTags";
+import { REPORTS_INDEX_PATH } from "./report/views";
import {
CONTRACT,
archiveUrl,
@@ -34,6 +35,13 @@ import {
// where X collabs". The per-record key itself is additive and bumps no manifest
// version — an older reader ignores an unknown field, as it always has.
//
+// v5: reports — a site may publish cited reports, announced as `reports`
+// ({ index: "/reports/index.json" }, the report index the export's Reports
+// pages read); and a CITED site (site.json `publish: "cited"`) publishes ONLY
+// them, saying so as `site.scope: "cited"` with no channels and zero totals.
+// An older reader sees an empty corpus there, which is the safe reading: there
+// is no shard to fetch.
+//
// NOT a v4: the `generator` field added below is deliberately unversioned. Every
// prior bump announced a new FETCHABLE LAYER — a reader that ignored it would
// miss data it could otherwise have retrieved. `generator` is an informational
@@ -177,7 +185,12 @@ export type SiteCorpus = {
hubUrl?: string;
// Present only on a PRIVATE site's build (site.json `audience`, release 17
// slice XP): the operator's own reading copy, which no deploy path ships.
- audience?: "private";
+ // A CITED site's build always says who it is for, "public" included.
+ audience?: "private" | "public";
+ // Present only on a CITED site's build: it publishes its reports and the
+ // moments they cite, and nothing else — `channels` is empty, there are no
+ // shards (spec 5).
+ scope?: "cited";
};
totals: { channels: number; videos: number };
channels: CorpusChannel[];
@@ -193,6 +206,10 @@ export type SiteCorpus = {
tags?: { url: string; videoField: "curatedTags"; description: string };
// Present when this build ships bulk-download archives (whole-channel zips).
bulkArchives?: { manifest: string; note: string };
+ // Present when this site publishes at least one report (spec 5): the report
+ // index, each entry linking its page; every report's page view and
+ // citations sit beside it under /reports/<id>/.
+ reports?: { index: string; count: number; description: string };
// Pointer to the human page on using the archive with AI: the homepage's AI
// and MCP doc (AI_DOC_URL; the site's own /use-with-ai page until release
// 16). The BYO-key chat is the site's /ask/.
@@ -235,6 +252,12 @@ export function buildSiteCorpus(
// A private site's build (site.json `audience: "private"`): corpus.json's
// `site.audience` says so. Absent/false leaves corpus.json as before.
private?: boolean;
+ // How many reports this build published (compose's reports stage). Absent
+ // or 0 leaves corpus.json without a `reports` pointer.
+ reportCount?: number;
+ // A CITED site's build: `site.scope: "cited"` and `site.audience` always.
+ // Its descriptor carries no channels, so the corpus has none.
+ cited?: boolean;
},
): SiteCorpus {
const base = descriptor.siteUrl;
@@ -274,7 +297,8 @@ export function buildSiteCorpus(
description: descriptor.siteDescription,
...(descriptor.siteUrl ? { url: descriptor.siteUrl } : {}),
...(descriptor.hubUrl ? { hubUrl: descriptor.hubUrl } : {}),
- ...(opts.private ? { audience: "private" as const } : {}),
+ ...(opts.private ? { audience: "private" as const } : opts.cited ? { audience: "public" as const } : {}),
+ ...(opts.cited ? { scope: "cited" as const } : {}),
},
totals: { channels: channels.length, videos },
channels,
@@ -304,6 +328,19 @@ export function buildSiteCorpus(
"the platform's own keywords. Absent on archives built before spec 4.",
};
}
+ if (opts.reportCount) {
+ corpus.reports = {
+ index: join(base, REPORTS_INDEX_PATH),
+ count: opts.reportCount,
+ description:
+ "Cited reports. The index lists each report (title, kind, counts, " +
+ "`href` of its page); /reports/<id>/page.json is one report with every " +
+ "citation resolved, and /reports/<id>/citations.json (an " +
+ "archilyzer-citations set) and citations.csv are its citations as " +
+ "data. A video or audio citation's moment page is /m/<channel>/<id>/" +
+ "<start>-<end>/ (moment.json beside it), a post's /m/<channel>/<id>/.",
+ };
+ }
if (opts.hasArchives) {
corpus.bulkArchives = {
manifest: join(base, "/archives/manifest.json"),
@@ -355,9 +392,27 @@ export function buildHubCorpus(
// index. (Bounded either way — this is per-channel, never per-video.)
const LLMS_INLINE_LIMIT = 100;
+// A published report as llms.txt and the sitemap list it: its title and page.
+export type LlmsReport = { title: string; href: string; subtitle?: string };
+
+function pushReports(out: string[], base: string | undefined, reports: readonly LlmsReport[]): void {
+ for (const r of reports.slice(0, LLMS_INLINE_LIMIT)) {
+ out.push(`- [${r.title}](${join(base, r.href)})${r.subtitle ? `: ${r.subtitle}` : ""}`);
+ }
+ if (reports.length > LLMS_INLINE_LIMIT) {
+ out.push(`- …and ${reports.length - LLMS_INLINE_LIMIT} more — see the report index.`);
+ }
+}
+
// Render the per-site llms.txt (llmstxt.org convention: H1 + blockquote summary
-// + linked sections).
-export function renderSiteLlmsTxt(corpus: SiteCorpus): string {
+// + linked sections). A CITED site gets its own variant: its reports, and no
+// corpus layer, since it publishes none. `reports` lists the site's published
+// reports, in order (compose's reports stage).
+export function renderSiteLlmsTxt(
+ corpus: SiteCorpus,
+ opts: { reports?: readonly LlmsReport[] } = {},
+): string {
+ if (corpus.site.scope === "cited") return renderCitedLlmsTxt(corpus, opts.reports ?? []);
const base = corpus.site.url;
const out: string[] = [];
out.push(`# ${corpus.site.title}`);
@@ -408,6 +463,15 @@ export function renderSiteLlmsTxt(corpus: SiteCorpus): string {
`and live-chat zips for offline ingestion.`,
);
}
+ if (corpus.reports && opts.reports?.length) {
+ out.push("");
+ out.push("## Reports");
+ out.push(
+ `- [Report index](${corpus.reports.index}): cited reports — each citation ` +
+ `opens a moment page with the quote, its evidence and a link into the corpus.`,
+ );
+ pushReports(out, base, opts.reports);
+ }
out.push("");
out.push("## Channels");
const shown = corpus.channels.slice(0, LLMS_INLINE_LIMIT);
@@ -424,6 +488,48 @@ export function renderSiteLlmsTxt(corpus: SiteCorpus): string {
return out.join("\n") + "\n";
}
+// The cited variant: a site that publishes reports and the moments they cite,
+// and nothing else — no search, no transcripts, no shards to describe.
+function renderCitedLlmsTxt(corpus: SiteCorpus, reports: readonly LlmsReport[]): string {
+ const base = corpus.site.url;
+ const out: string[] = [];
+ out.push(`# ${corpus.site.title}`);
+ out.push("");
+ out.push(
+ `> ${corpus.site.description ? corpus.site.description.trim() + " " : ""}` +
+ `${reports.length} cited report(s). This site publishes only its reports and ` +
+ `the moments they cite — there is no searchable corpus here.`,
+ );
+ out.push("");
+ out.push("## Reports");
+ if (corpus.reports) {
+ out.push(
+ `- [Report index](${corpus.reports.index}): every report, as JSON; ` +
+ `/reports/<id>/page.json is one report with its citations resolved.`,
+ );
+ }
+ pushReports(out, base, reports);
+ out.push("");
+ out.push("## Citations");
+ out.push(
+ `- Each report's citations as data: /reports/<id>/citations.json (an ` +
+ `archilyzer-citations set) and /reports/<id>/citations.csv.`,
+ );
+ out.push(
+ `- A cited video or audio span has a moment page, /m/<channel>/<id>/<start>-<end>/ ` +
+ `(moment.json beside it): the quote, the evidence clip, the transcript lines ` +
+ `around it, the record, a link to the original and every report citing it. ` +
+ `A cited post's is /m/<channel>/<id>/.`,
+ );
+ out.push(
+ `- [corpus.json](${corpusUrl(base)}): this site's contract — \`site.scope\` is ` +
+ `"cited", and there are no channels or shards.`,
+ );
+ out.push("");
+ out.push(`Generated by ${corpus.generator}`);
+ return out.join("\n") + "\n";
+}
+
// Render the aggregate hub llms.txt.
export function renderHubLlmsTxt(corpus: HubCorpus): string {
const base = corpus.hub.url;
diff --git a/common/lib/envVars.ts b/common/lib/envVars.ts
@@ -141,6 +141,7 @@ const DECLARED: EnvVarDecl[] = [
{ name: "SITE_ID", audience: "internal", default: "—", readBy: "common/bin/compose-site.ts, export/app/lib/site.ts", doc: "Which site a compose or an export build is for. `archilyzer build site <id>` sets it; `compose site` and `build site` fall back to it when no id is given." },
{ name: "INSTANCE_MODE", audience: "internal", default: "a site", readBy: "export/app/lib/mode.ts, common/lib/archive/contract.ts", doc: "`hub` makes the export build the hub. Set by `archilyzer build hub`." },
{ name: "BUILD_ARCHIVES", audience: "internal", default: "on", readBy: "common/bin/compose-site.ts, common/bin/build-archives.ts", doc: "`0` skips archive-zip generation for one build (`--skip-archives`)." },
+ { name: "REPORTS_ALLOW_MISSING_MEDIA", audience: "internal", default: "off", readBy: "common/bin/compose-site.ts", doc: "`1` lets a report citation whose evidence media was not prepared through compose (`--allow-missing-media`): its moment page renders without a clip. Off, compose fails with the list." },
{ name: "ARCHIVES_READONLY", audience: "internal", default: "off", readBy: "common/bin/compose-site.ts", doc: "`1` inside a docker-mode build container: materialize archives, never write the shared cache." },
{ name: "HOMEPAGE_PUBLIC_DIR", audience: "internal", default: "`<repo>/homepage/public`", readBy: "common/bin/compose-homepage.ts, common/publish/source.ts", doc: "Where `compose homepage` and `source publish` write." },
diff --git a/common/publish/build.test.ts b/common/publish/build.test.ts
@@ -114,6 +114,16 @@ test("buildSiteSteps: skipData drops the data phase; skipArchives sets BUILD_ARC
);
});
+test("buildSiteSteps: allowMissingMedia lets compose through a report citation with no prepared media", () => {
+ const [compose] = buildSiteSteps({ siteId: "a", paths, skipData: true, allowMissingMedia: true, baseEnv: {} });
+ assert.equal(compose.args.join(" "), "run compose:site");
+ assert.equal(compose.env.REPORTS_ALLOW_MISSING_MEDIA, "1");
+ assert.equal(
+ buildSiteSteps({ siteId: "a", paths, baseEnv: {} })[0].env.REPORTS_ALLOW_MISSING_MEDIA,
+ undefined,
+ );
+});
+
test("buildHubSteps: compose:hub, then next build with INSTANCE_MODE=hub, in export/", () => {
const steps = buildHubSteps({ paths, baseEnv: { PATH: "/bin" } });
const env = {
diff --git a/common/publish/build.ts b/common/publish/build.ts
@@ -18,7 +18,9 @@ import {
builtAudienceProblem,
builtBundleProblem,
builtHubProblem,
+ builtScopeProblem,
builtSiteProblem,
+ citedBuildProblem,
deployAudienceProblem,
siteDeployProblem,
} from "../lib/builtExport";
@@ -81,6 +83,8 @@ export function buildSiteSteps(opts: {
paths: Paths;
skipData?: boolean;
skipArchives?: boolean;
+ // Let a report citation with no prepared media through compose.
+ allowMissingMedia?: boolean;
baseEnv?: NodeJS.ProcessEnv;
}): BuildStep[] {
const { paths } = opts;
@@ -94,6 +98,8 @@ export function buildSiteSteps(opts: {
// makes compose-site skip generation this build regardless of the
// global/site flags.
...(opts.skipArchives ? { BUILD_ARCHIVES: "0" } : {}),
+ // compose-site.ts ALLOW_MISSING_MEDIA_ENV (`--allow-missing-media`).
+ ...(opts.allowMissingMedia ? { REPORTS_ALLOW_MISSING_MEDIA: "1" } : {}),
};
const step = (args: string[]): BuildStep => ({
command: "pnpm",
@@ -134,7 +140,7 @@ export async function runBuildPhase(
signal: AbortSignal,
siteId: string,
paths: Paths,
- opts?: { skipData?: boolean; skipArchives?: boolean },
+ opts?: { skipData?: boolean; skipArchives?: boolean; allowMissingMedia?: boolean },
): Promise<number> {
// Skipping the data rebuild composes from the existing .export-index staging
// (see buildSiteSteps).
@@ -149,11 +155,29 @@ export async function runBuildPhase(
if (skipArchives) {
onLog("[notice] Skipping archive-zip generation for this build.\n");
}
- return runSteps(
+ const code = await runSteps(
onLog,
signal,
- buildSiteSteps({ siteId, paths, skipData, skipArchives }),
+ buildSiteSteps({
+ siteId,
+ paths,
+ skipData,
+ skipArchives,
+ allowMissingMedia: opts?.allowMissingMedia === true,
+ }),
);
+ if (code !== 0) return code;
+ // A CITED site's build must hold nothing but its reports, their moments and
+ // the shell (lib/builtExport.ts citedBuildProblem) — and must BE a cited
+ // build. Failing here is loud and early; every deploy path asks again.
+ const outDir = resolveOutDir(siteId, paths);
+ const scopeProblem =
+ citedBuildProblem(outDir) ?? builtScopeProblem(getSite(siteId, paths), outDir);
+ if (scopeProblem) {
+ onLog(`[build] REFUSED — ${scopeProblem}.\n`);
+ return 1;
+ }
+ return 0;
}
// Where the basic (host) compose staged this site's oversize archives for R2
@@ -750,12 +774,13 @@ function resolved(opts: PublishOpts): {
/** Build one site into export/out (basic/host mode). Returns the exit code. */
export async function buildSite(
siteId: string,
- opts: PublishOpts & { skipData?: boolean; skipArchives?: boolean } = {},
+ opts: PublishOpts & { skipData?: boolean; skipArchives?: boolean; allowMissingMedia?: boolean } = {},
): Promise<number> {
const { paths, onLog, signal } = resolved(opts);
return runBuildPhase(onLog, signal, siteId.trim(), paths, {
skipData: opts.skipData,
skipArchives: opts.skipArchives,
+ allowMissingMedia: opts.allowMissingMedia,
});
}
@@ -795,6 +820,9 @@ export async function deploySite(
// Before the R2 upload below: a bundle built private is never deployed.
const builtPrivate = builtAudienceProblem(outDir);
if (builtPrivate) throw new Error(`${builtPrivate}. Build ${site.siteId} again, then deploy.`);
+ // …nor a full build of a site that now publishes only its reports.
+ const builtScope = builtScopeProblem(site, outDir);
+ if (builtScope) throw new Error(`${builtScope}.`);
// The production path logs no banner and gains none here: its log has
// always opened on wrangler's own first line.
if (branch) {
diff --git a/common/publish/composeReports.test.ts b/common/publish/composeReports.test.ts
@@ -0,0 +1,551 @@
+// Integration: the reports stage of compose, through the REAL site compose
+// (bin/compose-site.ts main) over a temp corpus.
+//
+// One video channel (a record whose cues are fresh enough to read from its
+// VTT, and one whose `en` track parses to no cues so `en-orig` is read), one
+// Bluesky channel with a posts archive, a fact-check citing all five kinds
+// (and defining one citation it never cites), stills, a saved source copy, and
+// a prepared media manifest with fake clips and a capture — the cache
+// `archilyzer reports prepare` would have written. Three sites over it: a
+// CITED one, a FULL one with the same report, and a full one with none.
+//
+// Run with: node_modules/.bin/tsx --test publish/composeReports.test.ts
+
+import { after, test } from "node:test";
+import assert from "node:assert/strict";
+import { existsSync, mkdirSync, mkdtempSync, readdirSync, readFileSync, rmSync, writeFileSync } from "node:fs";
+import { tmpdir } from "node:os";
+import path from "node:path";
+
+const ROOT = mkdtempSync(path.join(tmpdir(), "compose-reports-"));
+const PINNED: Record<string, string> = {
+ TRANSCRIPTS_DIR: path.join(ROOT, "transcripts"),
+ SAVED_VIDEOS_DIR: path.join(ROOT, "saved-videos"),
+ SITES_DIR: path.join(ROOT, "transcripts", "sites"),
+ SETTINGS_FILE: path.join(ROOT, "settings.json"),
+ EXPORT_PUBLIC_DIR: path.join(ROOT, "public"),
+ EXPORT_INDEX_DIR: path.join(ROOT, ".export-index"),
+ EXPORT_BUILDS_DIR: path.join(ROOT, ".export-builds"),
+ EDITOR_CHANGELOG_FILE: path.join(ROOT, "editor-CHANGELOG.md"),
+ EXPORT_CHANGELOG_FILE: path.join(ROOT, "export-CHANGELOG.md"),
+ CHARTS_CONFIG_FILE: path.join(ROOT, "chart-templates.json"),
+ SEARCH_ALIASES_FILE: path.join(ROOT, "transcripts", "search-aliases.json"),
+ CURATED_TAGS_FILE: path.join(ROOT, "transcripts", "tags.json"),
+ ARCHILYZER_CONFIG_DIR: path.join(ROOT, "config"),
+ ARCHILYZER_SOURCE_SCRATCH: path.join(ROOT, "source-scratch"),
+};
+Object.assign(process.env, PINNED);
+delete process.env.REPORTS_ALLOW_MISSING_MEDIA;
+after(() => rmSync(ROOT, { recursive: true, force: true }));
+
+const { getPaths } = await import("../lib/paths");
+const { buildIndex } = await import("../controller/buildIndex");
+const { writePosts } = await import("../lib/posts-server");
+const { main: composeSite } = await import("../bin/compose-site");
+const { ComposeReportsError, CITATIONS_CSV_COLUMNS } = await import("./composeReports");
+const { REPORT_MEDIA_FORMAT, REPORT_MEDIA_VERSION, reportMediaDir, reportMediaIndexFile } = await import("./reportMedia");
+const { QUOTE_CHECK_METHOD } = await import("../lib/citations/verify");
+const { parseCitationSet } = await import("../lib/citations/validate");
+const { CONTRACT } = await import("../lib/archive/contract");
+const { citedBuildProblem, builtBundleProblem } = await import("../lib/builtExport");
+
+const paths = getPaths();
+const VIDEOS = "demo-channel";
+const SOCIAL = "demo-social";
+const REPORT = "demo-report";
+const NOW_FLOOR = new Date().toISOString();
+
+const writeJson = (file: string, value: unknown) => {
+ mkdirSync(path.dirname(file), { recursive: true });
+ writeFileSync(file, JSON.stringify(value, null, 2));
+};
+const writeText = (file: string, text: string) => {
+ mkdirSync(path.dirname(file), { recursive: true });
+ writeFileSync(file, text);
+};
+const readJson = <T = Record<string, unknown>>(file: string): T => JSON.parse(readFileSync(file, "utf8")) as T;
+const pub = (...p: string[]) => path.join(paths.exportPublicDir, ...p);
+
+// Every file under a directory, relative, sorted.
+function filesUnder(dir: string): string[] {
+ const out: string[] = [];
+ const walk = (d: string, rel: string) => {
+ for (const e of readdirSync(d, { withFileTypes: true })) {
+ const r = rel ? `${rel}/${e.name}` : e.name;
+ if (e.isDirectory()) walk(path.join(d, e.name), r);
+ else out.push(r);
+ }
+ };
+ if (existsSync(dir)) walk(dir, "");
+ return out.sort();
+}
+
+// A YouTube-shaped VTT: parseVtt keeps only lines carrying inline timing.
+const ts = (s: number) => new Date(s * 1000).toISOString().slice(11, 23);
+const vtt = (cues: [number, number, string][], tagged = true) =>
+ "WEBVTT\nKind: captions\nLanguage: en\n\n" +
+ cues
+ .map(([a, b, text]) => `${ts(a)} --> ${ts(b)} align:start position:0%\n${text}${tagged ? `<${ts(a)}><c></c>` : ""}\n`)
+ .join("\n");
+
+const CLIP = "0".repeat(31) + "1";
+const AUDIO_CLIP = "0".repeat(31) + "2";
+const UNCITED_CLIP = "0".repeat(31) + "3";
+
+function report(over: { c01Quote?: string } = {}) {
+ return {
+ format: "archilyzer-report",
+ version: 1,
+ id: REPORT,
+ kind: "factcheck",
+ title: "Checking a demo article",
+ subtitle: "Four claims, one stream",
+ published: "2026-10-01",
+ subject: { source: "s0" },
+ sources: {
+ s0: {
+ kind: "article",
+ title: "A demo article",
+ url: "https://example.org/article",
+ saved: "sources/s0/page.html",
+ },
+ },
+ citations: {
+ c01: {
+ kind: "video",
+ channel: VIDEOS,
+ id: "abc123",
+ start: 10,
+ end: 20,
+ pad: { before: 2, after: 3 },
+ quote: over.c01Quote ?? "The bridge opened in the spring, I was there for it.",
+ // Hand-typed: compose overwrites it.
+ verification: { quoteScore: 1, quoteCheckedAt: "2020-01-01T00:00:00Z", voiceChecked: true, method: "by hand" },
+ },
+ c02: { kind: "audio", channel: VIDEOS, id: "def456", start: 5, end: 9, quote: "words only the original track has" },
+ p01: { kind: "post", channel: SOCIAL, id: "3kabc", quote: "Posted to settle it", date: "2025-03-14" },
+ a01: { kind: "source", source: "s0", quote: "He opened the bridge himself.", image: "stills/a01.png" },
+ w01: {
+ kind: "page",
+ url: "https://example.org/page",
+ quote: "a page says so",
+ verification: { quoteScore: 0.5, quoteCheckedAt: "2020-01-01T00:00:00Z" },
+ },
+ u01: { kind: "video", channel: VIDEOS, id: "abc123", start: 40, end: 45, quote: "never cited" },
+ },
+ sections: [
+ {
+ id: "bridge",
+ title: "The bridge",
+ claims: [
+ {
+ id: "claim-1",
+ text: "He opened the bridge himself.",
+ verdict: "CONTRADICTED",
+ sourceQuote: { citation: "a01" },
+ findings: "He says [it opened without him](cite:c01), and [posted so](cite:p01).",
+ citations: ["c01", "c02", "w01"],
+ },
+ ],
+ },
+ ],
+ };
+}
+
+function seedSite(siteId: string, extra: Record<string, unknown> = {}, withReport = true) {
+ writeJson(path.join(paths.sitesDir, siteId, "site.json"), {
+ siteId,
+ siteTitle: `Site ${siteId}`,
+ siteDescription: "fixture",
+ headerTitle: siteId,
+ homeTagline: "",
+ socialLinks: [],
+ groups: [{ id: "default", name: "All channels", selectedByDefault: true }],
+ defaultGroupId: "default",
+ channels: [VIDEOS, SOCIAL].map((slug) => ({ slug, groupId: "default" })),
+ siteUrl: `https://${siteId}.example.test`,
+ archives: false,
+ ...extra,
+ });
+ if (!withReport) return;
+ const dir = path.join(paths.sitesDir, siteId, "reports", REPORT);
+ writeJson(path.join(dir, "report.json"), report());
+ writeText(path.join(dir, "stills", "a01.png"), "png-bytes");
+ writeText(path.join(dir, "sources", "s0", "page.html"), "<p>saved copy, never published</p>");
+ seedMedia(siteId);
+}
+
+// What `archilyzer reports prepare` leaves: the manifest, the clips with
+// their sidecars, the cited capture.
+function seedMedia(siteId: string, drop: string[] = []) {
+ const dir = reportMediaDir(paths, siteId);
+ const clip = (hash: string, kind: "video" | "audio", span: { from: number; to: number }) => {
+ const file = `${hash}${kind === "video" ? ".mp4" : ".m4a"}`;
+ writeText(path.join(dir, file), `${kind}-bytes-${hash}`);
+ const media = { kind, file, bytes: 20, sha256: "f".repeat(64), width: null, height: null, durationSec: span.to - span.from };
+ writeJson(path.join(dir, `${hash}.json`), { ...media, profile: "evidence-v1", span, source: { kind: "corpus-window", name: "x" } });
+ return media;
+ };
+ writeText(path.join(dir, "posts", SOCIAL, "3kabc", "shot.png"), "shot");
+ writeText(path.join(dir, "posts", SOCIAL, "3kabc", "photo.jpg"), "photo");
+ const moments: Record<string, unknown> = {
+ [`${VIDEOS}/abc123/10.00-20.00`]: clip(CLIP, "video", { from: 8, to: 23 }),
+ [`${VIDEOS}/def456/5.00-9.00`]: clip(AUDIO_CLIP, "audio", { from: 5, to: 9 }),
+ [`${VIDEOS}/abc123/40.00-45.00`]: clip(UNCITED_CLIP, "video", { from: 40, to: 45 }),
+ [`${SOCIAL}/3kabc`]: {
+ kind: "post",
+ file: `posts/${SOCIAL}/3kabc/shot.png`,
+ bytes: 4,
+ sha256: "e".repeat(64),
+ width: null,
+ height: null,
+ durationSec: null,
+ media: [{ file: `posts/${SOCIAL}/3kabc/photo.jpg`, bytes: 5, sha256: "d".repeat(64) }],
+ },
+ };
+ for (const k of drop) delete moments[k];
+ writeJson(reportMediaIndexFile(paths, siteId), {
+ format: REPORT_MEDIA_FORMAT,
+ version: REPORT_MEDIA_VERSION,
+ siteId,
+ preparedAt: "2026-10-05T00:00:00.000Z",
+ moments,
+ problems: [],
+ });
+}
+
+async function seedCorpus() {
+ writeJson(paths.settingsFile, {});
+ writeJson(path.join(paths.channelsDir, VIDEOS, "config.json"), {
+ handling: "youtube",
+ name: "Demo Channel",
+ url: "https://www.youtube.com/@demo/videos",
+ });
+ const meta = (id: string) => ({
+ id,
+ title: `Demo stream ${id}`,
+ channel: VIDEOS,
+ upload_date: "20260110",
+ duration: 600,
+ webpage_url: `https://www.youtube.com/watch?v=${id}`,
+ extractor_key: "Youtube",
+ });
+ const abc = path.join(paths.channelsDir, VIDEOS, "data", "abc123");
+ writeJson(path.join(abc, "metadata.info.json"), meta("abc123"));
+ writeText(
+ path.join(abc, "transcript.en.vtt"),
+ vtt([
+ [0, 6, "Okay so somebody asked about the bridge."],
+ [6, 10, "Let me be clear about this one."],
+ [10, 15, "The bridge opened in the spring,"],
+ [15, 20, "I was there for it."],
+ [20, 30, "Anyway, back to the mail."],
+ [30, 40, "Next letter."],
+ [40, 45, "never cited"],
+ [60, 70, "Far outside any window."],
+ ]),
+ );
+ // The served `en` track parses to no cues; the original track has them.
+ const def = path.join(paths.channelsDir, VIDEOS, "data", "def456");
+ writeJson(path.join(def, "metadata.info.json"), meta("def456"));
+ writeText(path.join(def, "transcript.en.vtt"), vtt([[5, 9, "lost words"]], false));
+ writeText(path.join(def, "transcript.en-orig.vtt"), vtt([[5, 9, "Words only the original track has."]]));
+
+ writeJson(path.join(paths.channelsDir, SOCIAL, "config.json"), {
+ handling: "youtube",
+ name: "Demo Social",
+ url: "https://bsky.app/profile/demo.example",
+ sourceKind: "social",
+ platform: "bluesky",
+ socialHandle: "demo.example",
+ });
+ await writePosts(path.join(paths.channelsDir, SOCIAL), [
+ {
+ id: "3kabc",
+ slug: `${SOCIAL}/3kabc`,
+ channelSlug: SOCIAL,
+ author: "demo.example",
+ authorName: "Demo",
+ createdAt: "2025-03-14T12:00:00.000Z",
+ uploadDate: "20250314",
+ text: "Posted to settle it. The bridge was open before I got there.",
+ url: "https://bsky.app/profile/demo.example/post/3kabc",
+ platform: "bluesky",
+ isReply: false,
+ isRepost: false,
+ links: [],
+ },
+ ]);
+
+ seedSite("cited", { publish: "cited", reports: [REPORT] });
+ seedSite("full", { reports: [REPORT] });
+ seedSite("plain", {}, false);
+ await buildIndex({ paths, onLog: () => {} });
+}
+
+async function compose(siteId: string, opts: { allowMissingMedia?: boolean } = {}) {
+ const log = console.log;
+ console.log = () => {};
+ try {
+ await composeSite({ siteId, paths, ...opts });
+ } finally {
+ console.log = log;
+ }
+}
+
+await seedCorpus();
+
+test("a cited site: exactly its reports, moments, cited media and contract — the corpus pruned", async () => {
+ // A full compose leaves the corpus in public/, as the shared public dir
+ // would hold it after another site's build; and some stray files besides.
+ await compose("plain");
+ assert.ok(existsSync(pub("transcripts", VIDEOS)));
+ for (const f of ["sw.js", "tags.json", "duplicates.json", "hub-sites.json", "chart-templates.json"]) writeText(pub(f), "stale");
+ writeText(pub("archives", "manifest.json"), "{}");
+ writeText(pub("digests", VIDEOS, "manifest.json"), "{}");
+
+ await compose("cited");
+ const abc = `${VIDEOS}/abc123/10.00-20.00`;
+ const def = `${VIDEOS}/def456/5.00-9.00`;
+ assert.deepEqual(filesUnder(paths.exportPublicDir), [
+ "_headers",
+ "corpus.json",
+ "llms.txt",
+ `m/${VIDEOS}/abc123/10.00-20.00/moment.json`,
+ `m/${VIDEOS}/def456/5.00-9.00/moment.json`,
+ `m/${SOCIAL}/3kabc/moment.json`,
+ "m/index.json",
+ `media/clips/${abc}.mp4`,
+ `media/clips/${def}.m4a`,
+ `media/posts/${SOCIAL}/3kabc/photo.jpg`,
+ `media/posts/${SOCIAL}/3kabc/shot.png`,
+ `reports/${REPORT}/citations.csv`,
+ `reports/${REPORT}/citations.json`,
+ `reports/${REPORT}/page.json`,
+ `reports/${REPORT}/stills/a01.png`,
+ "reports/index.json",
+ "robots.txt",
+ "site.json",
+ "sitemap.xml",
+ ]);
+ assert.equal(readFileSync(pub("media", "clips", `${abc}.mp4`), "utf8"), `video-bytes-${CLIP}`);
+ assert.equal(readFileSync(pub("media", "clips", `${def}.m4a`), "utf8"), `audio-bytes-${AUDIO_CLIP}`);
+ // The uncited citation's clip, prepared, is not published; nor its page.
+ assert.ok(!JSON.stringify(readJson(pub("m", "index.json"))).includes("40.00"));
+});
+
+test("a cited site's contract: site.json without channels, corpus.json scope cited at spec 5, the cited llms.txt and sitemap", () => {
+ const site = readJson<{ siteId: string; channels: unknown[]; pwa: boolean }>(pub("site.json"));
+ assert.equal(site.siteId, "cited");
+ assert.deepEqual(site.channels, []);
+ assert.equal(site.pwa, false);
+ const corpus = readJson<{
+ spec: number;
+ site: { id: string; scope?: string; audience?: string };
+ channels: unknown[];
+ totals: { channels: number; videos: number };
+ reports?: { index: string; count: number };
+ }>(pub("corpus.json"));
+ assert.equal(corpus.spec, CONTRACT.corpusSpec);
+ assert.equal(corpus.spec, 5);
+ assert.equal(corpus.site.id, "cited");
+ assert.equal(corpus.site.scope, "cited");
+ assert.equal(corpus.site.audience, "public");
+ assert.deepEqual(corpus.channels, []);
+ assert.deepEqual(corpus.totals, { channels: 0, videos: 0 });
+ assert.equal(corpus.reports?.index, "https://cited.example.test/reports/index.json");
+ assert.equal(corpus.reports?.count, 1);
+
+ const llms = readFileSync(pub("llms.txt"), "utf8");
+ assert.match(llms, /^# Site cited/);
+ assert.match(llms, /\[Checking a demo article\]\(https:\/\/cited\.example\.test\/reports\/demo-report\/\): Four claims, one stream/);
+ assert.match(llms, /there is no searchable corpus here/);
+ assert.doesNotMatch(llms, /## Channels|shardScheme|page-<NNNN>|transcripts\//);
+
+ const sitemap = readFileSync(pub("sitemap.xml"), "utf8");
+ const locs = [...sitemap.matchAll(/<loc>([^<]+)<\/loc>/g)].map((m) => m[1].replace("https://cited.example.test", ""));
+ assert.deepEqual(locs, [
+ "/",
+ "/reports/",
+ "/reports/demo-report/",
+ `/m/${VIDEOS}/abc123/10.00-20.00/`,
+ `/m/${VIDEOS}/def456/5.00-9.00/`,
+ `/m/${SOCIAL}/3kabc/`,
+ ]);
+ // The bundle names itself and its scope: the deploy guards take it.
+ assert.equal(builtBundleProblem(paths.exportPublicDir, "cited"), null);
+ assert.equal(citedBuildProblem(paths.exportPublicDir), null);
+});
+
+type View = { citations: Record<string, { verification?: Record<string, unknown>; record?: Record<string, unknown>; shot?: string; text?: string; author?: string }> };
+
+test("verification is computed at compose: a hand-typed block is overwritten, a page's dropped, en-orig read when en has no cues", () => {
+ const view = readJson<View>(pub("reports", REPORT, "page.json"));
+ const v = view.citations.c01.verification!;
+ assert.equal(v.quoteScore, 1);
+ assert.equal(v.method, QUOTE_CHECK_METHOD);
+ assert.ok((v.quoteCheckedAt as string) >= NOW_FLOOR, "checked now, not when the document says");
+ assert.ok(!("voiceChecked" in v), "a voice is never vouched for by compose");
+ assert.equal(view.citations.c02.verification?.quoteScore, 1, "checked against the en-orig track");
+ assert.equal(view.citations.p01.verification?.quoteScore, 1, "a post's quote against its text");
+ assert.equal(view.citations.w01.verification, undefined);
+ assert.equal(view.citations.a01.verification, undefined);
+ assert.equal(view.citations.u01, undefined, "a citation never cited is not in the view");
+ // The record, resolved from the corpus; a cited site links no corpus.
+ assert.deepEqual(view.citations.c01.record, {
+ channel: VIDEOS,
+ channelTitle: "Demo Channel",
+ id: "abc123",
+ title: "Demo stream abc123",
+ date: "2026-01-10",
+ platform: "youtube",
+ originalUrl: "https://www.youtube.com/watch?v=abc123&t=10s",
+ });
+ assert.equal(view.citations.p01.shot, `/media/posts/${SOCIAL}/3kabc/shot.png`);
+ assert.equal(view.citations.p01.author, "Demo (@demo.example)");
+});
+
+test("moment pages: the clip with its pad, the bounded cue context, the post's capture, cited in", () => {
+ const span = readJson<Record<string, unknown> & { cues: { start: number; text: string; inSpan: boolean }[]; citedIn: { href: string }[] }>(
+ pub("m", VIDEOS, "abc123", "10.00-20.00", "moment.json"),
+ );
+ assert.equal(span.kind, "video");
+ assert.deepEqual(span.clip, { src: `/media/clips/${VIDEOS}/abc123/10.00-20.00.mp4`, start: 8, end: 23 });
+ assert.equal(span.start, 10);
+ assert.equal(span.end, 20);
+ // ±15 s of context, never the whole record.
+ assert.deepEqual(
+ span.cues.map((q) => [q.start, q.inSpan]),
+ [
+ [0, false],
+ [6, false],
+ [10, true],
+ [15, true],
+ [20, false],
+ [30, false],
+ ],
+ );
+ assert.deepEqual(
+ span.citedIn.map((e) => e.href),
+ [`/reports/${REPORT}/#claim-1`],
+ );
+ const audio = readJson<Record<string, unknown>>(pub("m", VIDEOS, "def456", "5.00-9.00", "moment.json"));
+ assert.equal(audio.kind, "audio");
+ assert.deepEqual(audio.clip, { src: `/media/clips/${VIDEOS}/def456/5.00-9.00.m4a`, start: 5, end: 9 });
+
+ const post = readJson<Record<string, unknown> & { post: unknown; record: Record<string, unknown> }>(
+ pub("m", SOCIAL, "3kabc", "moment.json"),
+ );
+ assert.equal(post.kind, "post");
+ assert.equal(post.date, "2025-03-14");
+ assert.deepEqual(post.post, {
+ author: "Demo (@demo.example)",
+ text: "Posted to settle it. The bridge was open before I got there.",
+ shot: `/media/posts/${SOCIAL}/3kabc/shot.png`,
+ media: [{ src: `/media/posts/${SOCIAL}/3kabc/photo.jpg`, kind: "image" }],
+ });
+ assert.equal(post.record.originalUrl, "https://bsky.app/profile/demo.example/post/3kabc");
+});
+
+test("the citations as files: a valid citation set without the saved copy, and one CSV row per citation", () => {
+ const set = readJson(pub("reports", REPORT, "citations.json"));
+ const parsed = parseCitationSet(set);
+ assert.ok(parsed.ok, JSON.stringify(parsed.problems));
+ assert.deepEqual(parsed.problems, []);
+ assert.deepEqual(Object.keys((set as { citations: object }).citations), ["a01", "c01", "p01", "c02", "w01"]);
+ assert.ok(!JSON.stringify(set).includes("saved"), "a source's saved copy is never published");
+
+ const csv = readFileSync(pub("reports", REPORT, "citations.csv"), "utf8").trimEnd().split("\r\n");
+ assert.equal(csv[0], CITATIONS_CSV_COLUMNS.join(","));
+ assert.equal(csv.length, 6);
+ assert.equal(
+ csv[2],
+ `c01,2,video,${VIDEOS},abc123,10,20,"The bridge opened in the spring, I was there for it.",,2026-01-10,https://www.youtube.com/watch?v=abc123&t=10s,/m/${VIDEOS}/abc123/10.00-20.00/,1`,
+ );
+ assert.match(csv[1], /^a01,1,source,,,,,He opened the bridge himself\.,,,https:\/\/example\.org\/article,,$/);
+});
+
+test("a quote that drifted from its cues fails compose, before anything is written", async () => {
+ const file = path.join(paths.sitesDir, "cited", "reports", REPORT, "report.json");
+ writeJson(file, report({ c01Quote: "He said he cut the ribbon himself that morning." }));
+ try {
+ await assert.rejects(compose("cited"), (e: unknown) => {
+ assert.ok(e instanceof ComposeReportsError);
+ assert.deepEqual(
+ e.problems.map((p) => [p.kind, p.citation]),
+ [["quote-drift", `${REPORT}#c01`]],
+ );
+ return true;
+ });
+ assert.ok(!existsSync(pub("reports")));
+ assert.ok(!existsSync(pub("site.json")), "a failed compose leaves public/ naming no site");
+ } finally {
+ writeJson(file, report());
+ }
+});
+
+test("a citation without prepared media fails compose with the list, unless --allow-missing-media", async () => {
+ seedMedia("cited", [`${VIDEOS}/abc123/10.00-20.00`]);
+ try {
+ await assert.rejects(compose("cited"), (e: unknown) => {
+ assert.ok(e instanceof ComposeReportsError);
+ assert.deepEqual(
+ e.problems.map((p) => [p.kind, p.moment]),
+ [["missing-media", `${VIDEOS}/abc123/10.00-20.00`]],
+ );
+ assert.match(e.message, /cited by demo-report#c01/);
+ return true;
+ });
+ await compose("cited", { allowMissingMedia: true });
+ const m = readJson<Record<string, unknown>>(pub("m", VIDEOS, "abc123", "10.00-20.00", "moment.json"));
+ assert.equal(m.clip, undefined);
+ assert.ok(!existsSync(pub("media", "clips", VIDEOS, "abc123")));
+ } finally {
+ seedMedia("cited");
+ }
+});
+
+test("a clip prepared for another span is stale: the reports changed since prepare", async () => {
+ const sidecar = path.join(reportMediaDir(paths, "cited"), `${CLIP}.json`);
+ const saved = readFileSync(sidecar, "utf8");
+ writeJson(sidecar, { ...JSON.parse(saved), span: { from: 10, to: 20 } });
+ try {
+ await assert.rejects(compose("cited"), (e: unknown) => {
+ assert.ok(e instanceof ComposeReportsError);
+ assert.deepEqual(e.problems.map((p) => p.kind), ["stale-media"]);
+ return true;
+ });
+ } finally {
+ writeFileSync(sidecar, saved);
+ }
+});
+
+test("a full site with reports keeps its corpus and links each moment into it", async () => {
+ await compose("full");
+ assert.ok(existsSync(pub("transcripts", VIDEOS)), "the corpus is composed as ever");
+ assert.ok(existsSync(pub("summaries", "manifest.json")));
+ const corpus = readJson<{ spec: number; site: { scope?: string; audience?: string }; channels: unknown[]; reports?: { count: number } }>(
+ pub("corpus.json"),
+ );
+ assert.equal(corpus.spec, 5);
+ assert.equal(corpus.site.scope, undefined);
+ assert.equal(corpus.site.audience, undefined);
+ assert.equal(corpus.channels.length, 2);
+ assert.equal(corpus.reports?.count, 1);
+ const m = readJson<{ record: { corpusUrl?: string } }>(pub("m", VIDEOS, "abc123", "10.00-20.00", "moment.json"));
+ assert.equal(m.record.corpusUrl, `/?v=${VIDEOS}%2Fabc123&t=10`);
+ const post = readJson<{ record: { corpusUrl?: string } }>(pub("m", SOCIAL, "3kabc", "moment.json"));
+ assert.equal(post.record.corpusUrl, `/?v=${SOCIAL}%2F3kabc&vm=post`);
+ const llms = readFileSync(pub("llms.txt"), "utf8");
+ assert.match(llms, /## Reports[^]*Checking a demo article[^]*## Channels/);
+ assert.match(readFileSync(pub("sitemap.xml"), "utf8"), /\/reports\/demo-report\//);
+ assert.equal(citedBuildProblem(paths.exportPublicDir), null, "the audit leaves a full build alone");
+});
+
+test("a site with no reports ships none of the last site's", async () => {
+ await compose("cited");
+ await compose("plain");
+ for (const entry of ["reports", "m", "media"]) assert.ok(!existsSync(pub(entry)), entry);
+ const corpus = readJson<{ reports?: unknown }>(pub("corpus.json"));
+ assert.equal(corpus.reports, undefined);
+});
diff --git a/common/publish/composeReports.ts b/common/publish/composeReports.ts
@@ -0,0 +1,742 @@
+// THE REPORTS STAGE OF COMPOSE — a site's published reports, resolved against
+// the corpus and written as the views the export's Reports pages read
+// (lib/report/views.ts names every file), with the media and stills they cite
+// (plans/report-sites.md, "Compose and the contract").
+//
+// Runs for EVERY site's compose, full or cited: it first removes what an
+// earlier compose — of this site or another, public/ is shared — left under
+// `reports/`, `m/` and `media/`, so a site with no reports ships none.
+//
+// What it reads:
+// - the site's `reports` (site.json, in order), each report.json parsed and
+// validated by the document's own checker (./reportMedia.ts
+// loadSiteReports); any problem fails the stage;
+// - per cited video/audio record: transcript.cues.json when it is fresh,
+// else metadata.info.json and the raw transcript; when those cues are
+// empty, the English VTT tracks in turn, `en-orig` first (a served `en`
+// track can parse to no cues);
+// - per cited post: the channel's posts archive (lib/posts-server.ts), and
+// the post visibility rule (lib/postsVisibility.ts);
+// - the media `archilyzer reports prepare` cut and copied for the site
+// (./reportMedia.ts — the manifest and the cache beside it). This stage
+// never cuts: the build has no ffmpeg.
+//
+// What it computes: every citation's VERIFICATION, overwriting whatever the
+// document carried (lib/citations/verify.ts): a span's quote against its cue
+// window, a post's against its text; a `source` or `page` citation keeps none.
+// A quote that drifted fails the stage.
+//
+// What it writes, under the public dir:
+// reports/index.json the report index
+// reports/<id>/page.json each report's page view
+// reports/<id>/citations.{json,csv} its citations as data (the JSON is
+// an `archilyzer-citations` set;
+// never a source's `saved` copy)
+// reports/<id>/<still> each cited source still
+// m/index.json, m/<key>/moment.json one moment view per cited moment,
+// "cited in" across every report
+// media/clips/<channel>/<id>/<s>-<e>.mp4 a span's prepared clip (.m4a for
+// a clip cut as audio)
+// media/posts/<channel>/<id>/<file> a cited post's prepared capture
+// — ONLY what the published reports cite. A citation whose media was not
+// prepared fails the stage with the list, unless `allowMissingMedia`: its page
+// then renders without a clip.
+//
+// THE STAGE FAILS BEFORE IT WRITES: every problem is collected first and
+// thrown together (ComposeReportsError), so a failed compose leaves no half
+// set of reports behind.
+
+import { copyFile, mkdir, readdir, readFile, rm, stat, writeFile } from "node:fs/promises";
+import path from "node:path";
+import type { Paths } from "../lib/paths";
+import { siteChannelSlugs, type Site } from "../lib/site";
+import { isCitedSite } from "../lib/siteSchema";
+import { getSettings } from "../lib/settings";
+import { postsVisibleTo } from "../lib/postsVisibility";
+import { assertChannelTextReadable } from "../lib/channelMedia";
+import type { ChannelConfig } from "../lib/channelConfig";
+import { readChannelConfig } from "../controller/channels";
+import { isCuesJsonFresh, readNormalizedTranscript } from "../controller/normalizeTranscript";
+import { loadRawMetadataFromDir, summarize } from "../lib/transcripts-server";
+import type { TranscriptSummary } from "../lib/transcripts";
+import { parseVtt, type Cue } from "../lib/vtt";
+import { parseTranscriptJson } from "../lib/whisper";
+import { WHISPER_FILENAME, isEnglishVtt, resolvePrimaryVtt } from "../lib/videoStatus";
+import { platformMomentUrl } from "../lib/momentUrl";
+import type { Platform } from "../lib/platform";
+import { readAllPosts } from "../lib/posts-server";
+import type { Post } from "../lib/posts";
+import { readJsonFile } from "../lib/jsonFile-server";
+import { momentPath, parseMomentKey, type SpanMoment } from "../lib/citations/moments";
+import { CITATIONS_VERSION, type Citation, type PostCitation, type SpanCitation } from "../lib/citations/schema";
+import { cueWindowText, quoteDrifted, quoteVerification, QUOTE_DRIFT_THRESHOLD } from "../lib/citations/verify";
+import { buildCitedIn } from "../lib/report/citedIn";
+import type { Report } from "../lib/report/schema";
+import { reportCitationNumbers } from "../lib/report/uses";
+import {
+ MOMENT_INDEX_FORMAT,
+ MOMENT_PAGE_FORMAT,
+ MOMENTS_INDEX_PATH,
+ REPORT_INDEX_FORMAT,
+ REPORT_VIEWS_VERSION,
+ REPORTS_INDEX_PATH,
+ buildReportPageView,
+ citedInViews,
+ evidenceClipPath,
+ momentViewPath,
+ orderedCitations,
+ reportCitationsDownloadPath,
+ reportIndexEntry,
+ reportViewPath,
+ type CitationView,
+ type CueLineView,
+ type MomentPageView,
+ type MomentPostView,
+ type RecordView,
+ type ReportIndexEntry,
+ type ReportIndexView,
+ type ReportPageView,
+} from "../lib/report/views";
+import { evidenceSpan, isAudioOnlyPlatform, type EvidenceSpan } from "../lib/evidenceClip-server";
+import {
+ citedMoments,
+ loadSiteReports,
+ readReportMediaIndex,
+ reportMediaDir,
+ siteReportDir,
+ type ReportMediaEntry,
+} from "./reportMedia";
+
+// The public dir's entries this stage owns. Every compose removes them first.
+export const REPORT_PUBLIC_ENTRIES: readonly string[] = ["reports", "m", "media"];
+
+// The transcript lines a moment page shows either side of its span, and at
+// most how many: bounded context, never the record (plans/report-sites.md,
+// Risks 5).
+export const MOMENT_CUE_CONTEXT_SECONDS = 15;
+export const MOMENT_CUE_LINES_MAX = 80;
+
+export const CITATIONS_CSV_COLUMNS = [
+ "citation",
+ "number",
+ "kind",
+ "channel",
+ "id",
+ "start",
+ "end",
+ "quote",
+ "speaker",
+ "date",
+ "originalUrl",
+ "momentPath",
+ "quoteScore",
+] as const;
+
+export type ComposeReportsProblemKind =
+ | "missing-report"
+ | "invalid-report"
+ | "not-in-site"
+ | "not-visible"
+ | "unreadable"
+ | "missing-record"
+ | "no-cues"
+ | "quote-drift"
+ | "missing-post"
+ | "missing-still"
+ | "missing-media"
+ | "stale-media";
+
+export type ComposeReportsProblem = {
+ kind: ComposeReportsProblemKind;
+ message: string;
+ report?: string;
+ // `<reportId>#<citationId>`.
+ citation?: string;
+ moment?: string;
+ // A JSON path in the report (an invalid report's problems).
+ path?: string;
+};
+
+// The kinds `allowMissingMedia` lets through: the page renders without a clip.
+const MEDIA_PROBLEMS: ReadonlySet<ComposeReportsProblemKind> = new Set(["missing-media", "stale-media"]);
+
+export class ComposeReportsError extends Error {
+ constructor(readonly problems: ComposeReportsProblem[]) {
+ super(
+ `the site's reports cannot be composed (${problems.length} problem(s)):\n` +
+ formatComposeReportsProblems(problems).map((l) => ` ${l}`).join("\n"),
+ );
+ this.name = "ComposeReportsError";
+ }
+}
+
+export function formatComposeReportsProblems(problems: readonly ComposeReportsProblem[]): string[] {
+ return problems.map((p) => {
+ const where = p.citation ?? p.moment ?? `${p.report ?? "?"}${p.path ? ` at ${p.path}` : ""}`;
+ return `${p.kind}: ${where}: ${p.message}`;
+ });
+}
+
+export type ComposeReportsOptions = {
+ paths: Paths;
+ site: Site;
+ // Default: paths.exportPublicDir.
+ publicDir?: string;
+ // Let a citation without prepared media through (`--allow-missing-media`).
+ allowMissingMedia?: boolean;
+ // `social.x.visibility` and friends; default the live settings.
+ settings?: { social?: { x?: { visibility?: unknown } } };
+ now?: () => Date;
+ log?: (line: string) => void;
+};
+
+export type ComposedReports = {
+ // The index's entries, in the site's order.
+ reports: ReportIndexEntry[];
+ // Every moment page written.
+ moments: string[];
+ // Media problems let through by `allowMissingMedia`.
+ allowed: ComposeReportsProblem[];
+};
+
+// ─── Reading the corpus ───
+
+type CitedRecord = {
+ summary: Pick<TranscriptSummary, "id" | "slug" | "title" | "uploadDate" | "platform" | "webpageUrl" | "channel">;
+ cues: Cue[];
+};
+
+// The English VTT tracks of a video dir, `en-orig` first, then the order
+// resolvePrimaryVtt prefers.
+function englishVttsByPreference(entries: readonly string[]): string[] {
+ const vtts = entries.filter(isEnglishVtt);
+ const ordered: string[] = [];
+ const orig = vtts.find((n) => n === "transcript.en-orig.vtt");
+ if (orig) ordered.push(orig);
+ const rest = vtts.filter((n) => n !== orig);
+ while (rest.length > 0) {
+ const best = resolvePrimaryVtt(rest)!;
+ ordered.push(best);
+ rest.splice(rest.indexOf(best), 1);
+ }
+ return ordered;
+}
+
+async function readCues(file: string, kind: "vtt" | "whisper"): Promise<Cue[]> {
+ try {
+ const raw = await readFile(file, "utf8");
+ return kind === "vtt" ? parseVtt(raw) : parseTranscriptJson(raw);
+ } catch {
+ return [];
+ }
+}
+
+// A cited record: its summary and its cues, or null when the data dir holds
+// neither a normalized transcript nor metadata.
+export async function readCitedRecord(
+ channelsDir: string,
+ slug: string,
+ id: string,
+ channelName?: string,
+): Promise<CitedRecord | null> {
+ const dir = path.join(channelsDir, slug, "data", id);
+ const entries = await readdir(dir).catch(() => [] as string[]);
+ if (entries.length === 0) return null;
+ let summary: CitedRecord["summary"] | null = null;
+ let cues: Cue[] = [];
+ const fresh = await isCuesJsonFresh(dir);
+ if (fresh.fresh) {
+ const n = await readNormalizedTranscript(fresh.cuesPath);
+ if (n) {
+ summary = n;
+ cues = n.cues ?? [];
+ }
+ }
+ if (!summary) {
+ const meta = await loadRawMetadataFromDir(dir);
+ if (!meta) return null;
+ summary = summarize(slug, id, meta, channelName);
+ if (entries.includes(WHISPER_FILENAME)) cues = await readCues(path.join(dir, WHISPER_FILENAME), "whisper");
+ if (cues.length === 0) {
+ const primary = resolvePrimaryVtt(entries);
+ if (primary) cues = await readCues(path.join(dir, primary), "vtt");
+ }
+ }
+ if (cues.length === 0) {
+ for (const name of englishVttsByPreference(entries)) {
+ cues = await readCues(path.join(dir, name), "vtt");
+ if (cues.length > 0) break;
+ }
+ }
+ return { summary, cues };
+}
+
+const isoDay = (uploadDate: string | undefined): string | undefined =>
+ uploadDate && /^\d{8}$/.test(uploadDate)
+ ? `${uploadDate.slice(0, 4)}-${uploadDate.slice(4, 6)}-${uploadDate.slice(6, 8)}`
+ : undefined;
+
+// A record in this site's corpus — the viewer's `?v=` link (a FULL site only).
+function corpusLink(slug: string, params: Record<string, string>): string {
+ return `/?${new URLSearchParams({ v: slug, ...params }).toString()}`;
+}
+
+const cueLine = (text: string) => text.replace(/\s+/g, " ").trim();
+
+const IMAGE_EXTS = new Set([".png", ".jpg", ".jpeg", ".gif", ".webp", ".avif"]);
+
+function csvCell(v: unknown): string {
+ if (v === undefined || v === null) return "";
+ const s = String(v);
+ return /[",\r\n]/.test(s) ? `"${s.replace(/"/g, '""')}"` : s;
+}
+
+// A report's citations as CSV: one row per citation it cites, in number order.
+export function citationsCsv(view: ReportPageView): string {
+ const rows: string[] = [CITATIONS_CSV_COLUMNS.join(",")];
+ for (const c of orderedCitations(view)) {
+ const span = c.kind === "video" || c.kind === "audio" ? c : null;
+ const recorded = c.kind === "video" || c.kind === "audio" || c.kind === "post" ? c : null;
+ const row: Record<(typeof CITATIONS_CSV_COLUMNS)[number], unknown> = {
+ citation: c.id,
+ number: c.number,
+ kind: c.kind,
+ channel: recorded?.record.channel,
+ id: recorded?.record.id,
+ start: span?.start,
+ end: span?.end,
+ quote: c.quote,
+ speaker: c.speaker,
+ date: c.date ?? recorded?.record.date,
+ originalUrl: recorded ? recorded.record.originalUrl : c.href,
+ momentPath: recorded?.href,
+ quoteScore: c.verification?.quoteScore,
+ };
+ rows.push(CITATIONS_CSV_COLUMNS.map((k) => csvCell(row[k])).join(","));
+ }
+ return rows.join("\r\n") + "\r\n";
+}
+
+// A report's citations as an `archilyzer-citations` set (CITATIONS.md): the
+// citations it cites, verification computed, and the sources they quote —
+// never a source's `saved` copy.
+export function citationSet(report: Report, view: ReportPageView): unknown {
+ const cited = orderedCitations(view).map((c) => c.id);
+ const all = report.citations ?? {};
+ const citations = Object.fromEntries(cited.map((id) => [id, all[id]]));
+ const sourceIds = new Set<string>();
+ for (const id of cited) if (all[id].kind === "source") sourceIds.add((all[id] as { source: string }).source);
+ if (report.subject) sourceIds.add(report.subject.source);
+ const sources = Object.fromEntries(
+ [...sourceIds]
+ .filter((id) => report.sources?.[id])
+ .map((id) => {
+ const { saved: _saved, ...rest } = report.sources![id];
+ void _saved;
+ return [id, rest];
+ }),
+ );
+ return {
+ format: "archilyzer-citations",
+ version: CITATIONS_VERSION,
+ ...(Object.keys(sources).length > 0 ? { sources } : {}),
+ citations,
+ };
+}
+
+const json = (value: unknown) => `${JSON.stringify(value, null, 2)}\n`;
+
+async function writeOut(publicDir: string, urlPath: string, data: string): Promise<void> {
+ const file = path.join(publicDir, ...urlPath.split("/").filter(Boolean));
+ await mkdir(path.dirname(file), { recursive: true });
+ await writeFile(file, data);
+}
+
+async function copyOut(publicDir: string, src: string, urlPath: string): Promise<void> {
+ const file = path.join(publicDir, ...urlPath.split("/").filter(Boolean));
+ await mkdir(path.dirname(file), { recursive: true });
+ await copyFile(src, file);
+}
+
+const isFile = async (p: string) => (await stat(p).catch(() => null))?.isFile() === true;
+
+const sameSpan = (a: EvidenceSpan, b: EvidenceSpan) =>
+ Math.abs(a.from - b.from) < 0.001 && Math.abs(a.to - b.to) < 0.001;
+
+// ─── The stage ───
+
+export async function composeReports(opts: ComposeReportsOptions): Promise<ComposedReports> {
+ const { paths, site } = opts;
+ const publicDir = opts.publicDir ?? paths.exportPublicDir;
+ const log = opts.log ?? (() => {});
+ const now = (opts.now?.() ?? new Date()).toISOString();
+ const cited = isCitedSite(site);
+
+ // Whatever an earlier compose left. rm removes a link, never its target (a
+ // worktree's public/ entries may be links into the primary checkout).
+ for (const entry of REPORT_PUBLIC_ENTRIES) {
+ await rm(path.join(publicDir, entry), { recursive: true, force: true });
+ }
+ if ((site.reports ?? []).length === 0) {
+ log("[reports] none published.");
+ return { reports: [], moments: [], allowed: [] };
+ }
+
+ const problems: ComposeReportsProblem[] = [];
+ const loaded = await loadSiteReports(paths, site);
+ for (const p of loaded.problems) {
+ problems.push({
+ kind: p.kind === "missing-report" ? "missing-report" : "invalid-report",
+ message: p.message,
+ report: p.report,
+ path: p.path,
+ });
+ }
+ // Compose works on its own copy: verification is overwritten below.
+ const reports: Report[] = loaded.problems.length > 0 ? [] : structuredClone(loaded.reports);
+
+ const settings = opts.settings ?? getSettings();
+ const pool = siteChannelSlugs(site);
+ const configs = new Map<string, ChannelConfig | null>();
+ const configOf = async (slug: string) => {
+ if (!configs.has(slug)) configs.set(slug, await readChannelConfig(paths, slug).catch(() => null));
+ return configs.get(slug) ?? null;
+ };
+ // Per channel, whether its text can be read (a legacy or migrating channel
+ // cannot), as the problem's sentence or null.
+ const unreadable = new Map<string, string | null>();
+ const textProblem = async (slug: string): Promise<string | null> => {
+ if (!unreadable.has(slug)) {
+ try {
+ await assertChannelTextReadable(paths, slug, await configOf(slug));
+ unreadable.set(slug, null);
+ } catch (e) {
+ unreadable.set(slug, (e as Error).message);
+ }
+ }
+ return unreadable.get(slug) ?? null;
+ };
+ const records = new Map<string, CitedRecord | null>();
+ const recordOf = async (slug: string, id: string) => {
+ const k = `${slug}/${id}`;
+ if (!records.has(k)) records.set(k, await readCitedRecord(paths.channelsDir, slug, id, (await configOf(slug))?.name));
+ return records.get(k) ?? null;
+ };
+ const postsByChannel = new Map<string, Map<string, Post>>();
+ const postOf = async (slug: string, id: string) => {
+ if (!postsByChannel.has(slug)) {
+ const posts = await readAllPosts(path.join(paths.channelsDir, slug)).catch(() => [] as Post[]);
+ postsByChannel.set(slug, new Map(posts.map((p) => [p.id, p])));
+ }
+ return postsByChannel.get(slug)!.get(id) ?? null;
+ };
+
+ // Resolve and verify every citation the reports cite. A channel outside the
+ // site's pool, or a post the site may not carry, is refused before its
+ // record is read.
+ const refusedChannel = new Set<string>();
+ for (const report of reports) {
+ const all = report.citations ?? {};
+ // Only what the report cites: a citation it defines but never cites is
+ // not in its view, and is neither checked nor published.
+ const used = new Set(reportCitationNumbers(report).keys());
+ for (const [cid, c] of Object.entries(all)) {
+ const ref = `${report.id}#${cid}`;
+ if (!used.has(cid)) continue;
+ if (c.kind === "source" || c.kind === "page") {
+ delete c.verification;
+ if (c.kind === "source" && c.image) {
+ const src = path.join(siteReportDir(paths, site.siteId, report.id), c.image);
+ if (!(await isFile(src))) {
+ problems.push({ kind: "missing-still", citation: ref, report: report.id, message: `the still ${c.image} does not exist` });
+ }
+ }
+ continue;
+ }
+ if (!pool.has(c.channel)) {
+ problems.push({ kind: "not-in-site", citation: ref, report: report.id, message: `channel "${c.channel}" is not one of this site's channels` });
+ refusedChannel.add(c.channel);
+ continue;
+ }
+ const text = await textProblem(c.channel);
+ if (text) {
+ problems.push({ kind: "unreadable", citation: ref, report: report.id, message: text });
+ continue;
+ }
+ if (c.kind === "post") {
+ if (!postsVisibleTo(site, await configOf(c.channel), settings)) {
+ problems.push({ kind: "not-visible", citation: ref, report: report.id, message: `this site may not carry posts of "${c.channel}" (the post visibility rule)` });
+ continue;
+ }
+ const post = await postOf(c.channel, c.id);
+ if (!post) {
+ problems.push({ kind: "missing-post", citation: ref, report: report.id, message: `no post ${c.id} in the posts archive of "${c.channel}"` });
+ continue;
+ }
+ c.verification = quoteVerification(c.quote, post.text, now);
+ } else {
+ const record = await recordOf(c.channel, c.id);
+ if (!record) {
+ problems.push({ kind: "missing-record", citation: ref, report: report.id, message: `no record ${c.channel}/${c.id} (no metadata or transcript in its data dir)` });
+ continue;
+ }
+ if (record.cues.length === 0) {
+ problems.push({ kind: "no-cues", citation: ref, report: report.id, message: `${c.channel}/${c.id} has no transcript cues to check the quote against` });
+ continue;
+ }
+ c.verification = quoteVerification(c.quote, cueWindowText(record.cues, c.start, c.end), now);
+ }
+ if (quoteDrifted(c.verification)) {
+ problems.push({
+ kind: "quote-drift",
+ citation: ref,
+ report: report.id,
+ message:
+ `the quote matches ${Math.round((c.verification.quoteScore ?? 0) * 100)}% of what the record says there ` +
+ `(at least ${Math.round(QUOTE_DRIFT_THRESHOLD * 100)}% is required): quote it verbatim, or fix the span`,
+ });
+ }
+ }
+ }
+
+ // The prepared media, per moment.
+ const media = await readReportMediaIndex(paths, site.siteId);
+ const cacheDir = reportMediaDir(paths, site.siteId);
+ const citedIn = buildCitedIn(reports);
+ const momentInfo = new Map(citedMoments(reports).map((m) => [m.key, m]));
+ const mediaOf = new Map<string, ReportMediaEntry>();
+ for (const key of Object.keys(citedIn)) {
+ const m = momentInfo.get(key);
+ if (!m || refusedChannel.has(m.moment.channel)) continue;
+ const entry = media?.siteId === site.siteId ? media.moments[key] : undefined;
+ const citedBy = citedIn[key].map((e) => `${e.reportId}#${e.citationId}`);
+ const missing = (kind: ComposeReportsProblemKind, message: string) =>
+ problems.push({ kind, moment: key, message: `${message} (cited by ${[...new Set(citedBy)].join(", ")})` });
+ if (!entry) {
+ missing("missing-media", "no prepared media — run Prepare evidence media (archilyzer reports prepare)");
+ continue;
+ }
+ const files = entry.kind === "post" ? [entry.file, ...entry.media.map((f) => f.file)] : [entry.file];
+ const absent = [];
+ for (const f of files) if (!(await isFile(path.join(cacheDir, f)))) absent.push(f);
+ if (absent.length > 0) {
+ missing("missing-media", `the prepared media is gone from the cache (${absent.join(", ")}) — prepare again`);
+ continue;
+ }
+ if (entry.kind !== "post" && m.moment.kind === "span") {
+ // The clip was cut for a span and pad; a report changed since must not
+ // ship the old cut under the new moment's page.
+ const sidecar = await readJsonFile(path.join(cacheDir, entry.file.replace(/\.[^.]+$/, ".json")));
+ const cut = sidecar.ok ? (sidecar.value as { span?: EvidenceSpan }).span : undefined;
+ const want = evidenceSpan({ start: m.moment.start, end: m.moment.end, pad: m.pad });
+ if (cut && !sameSpan(cut, want)) {
+ missing("stale-media", `the clip was cut for ${cut.from}–${cut.to} s, the reports now cite ${want.from}–${want.to} s — prepare again`);
+ continue;
+ }
+ }
+ mediaOf.set(key, entry);
+ }
+
+ const allowed = opts.allowMissingMedia ? problems.filter((p) => MEDIA_PROBLEMS.has(p.kind)) : [];
+ const fatal = problems.filter((p) => !allowed.includes(p));
+ if (fatal.length > 0) throw new ComposeReportsError(fatal);
+ for (const line of formatComposeReportsProblems(allowed)) log(`[reports] allowed (--allow-missing-media): ${line}`);
+
+ // ─── The views ───
+
+ const recordView = async (c: SpanCitation | PostCitation): Promise<RecordView> => {
+ const config = await configOf(c.channel);
+ if (c.kind === "post") {
+ const post = (await postOf(c.channel, c.id))!;
+ return defined({
+ channel: c.channel,
+ channelTitle: config?.name ?? post.authorName,
+ id: c.id,
+ date: post.createdAt.slice(0, 10),
+ platform: post.platform,
+ originalUrl: post.url,
+ corpusUrl: cited ? undefined : corpusLink(`${c.channel}/${c.id}`, { vm: "post" }),
+ });
+ }
+ const { summary } = (await recordOf(c.channel, c.id))!;
+ const audioOnly = isAudioOnlyPlatform(config?.platform);
+ const seconds = Math.max(0, Math.floor(c.start));
+ return defined({
+ channel: c.channel,
+ channelTitle: config?.name ?? (summary.channel || undefined),
+ id: c.id,
+ title: summary.title,
+ date: isoDay(summary.uploadDate),
+ platform: summary.platform,
+ originalUrl: audioOnly
+ ? summary.webpageUrl || undefined
+ : (platformMomentUrl(summary.webpageUrl, summary.platform as Platform, c.start) ?? undefined),
+ // The viewer keys a record by its PUBLISHED slug, which on a platform
+ // with two ids is not the data dir's name.
+ corpusUrl: cited ? undefined : corpusLink(summary.slug ?? `${c.channel}/${summary.id}`, seconds > 0 ? { t: String(seconds) } : {}),
+ });
+ };
+ // buildReportPageView's resolver is synchronous: resolve every cited
+ // record first, by the citation it is resolved for.
+ const recordViews = new WeakMap<Citation, RecordView>();
+ for (const report of reports) {
+ const all = report.citations ?? {};
+ for (const id of reportCitationNumbers(report).keys()) {
+ const c = all[id];
+ if (c.kind === "video" || c.kind === "audio" || c.kind === "post") recordViews.set(c, await recordView(c));
+ }
+ }
+ const postShot = (c: { channel: string; id: string }): string | undefined => {
+ const entry = mediaOf.get(`${c.channel}/${c.id}`);
+ return entry?.kind === "post" ? `/media/${entry.file}` : undefined;
+ };
+ const views: ReportPageView[] = reports.map((report) =>
+ buildReportPageView(report, {
+ record: (c) => recordViews.get(c)!,
+ post: (c) => {
+ const post = postsByChannel.get(c.channel)?.get(c.id);
+ return post ? { author: postAuthor(post), text: post.text, shot: postShot(c) } : undefined;
+ },
+ downloads: {
+ json: reportCitationsDownloadPath(report.id, "json"),
+ csv: reportCitationsDownloadPath(report.id, "csv"),
+ },
+ }),
+ );
+
+ const index: ReportIndexView = {
+ format: REPORT_INDEX_FORMAT,
+ version: REPORT_VIEWS_VERSION,
+ reports: views.map(reportIndexEntry),
+ };
+
+ const moments: MomentPageView[] = [];
+ for (const [key, entries] of Object.entries(citedIn)) {
+ const first = entries[0];
+ const report = reports.find((r) => r.id === first.reportId)!;
+ const c = report.citations![first.citationId] as Citation;
+ if (c.kind !== "video" && c.kind !== "audio" && c.kind !== "post") continue;
+ const view = views.find((v) => v.id === report.id)!.citations[first.citationId] as Extract<
+ CitationView,
+ { kind: "video" | "audio" | "post" }
+ >;
+ const entry = mediaOf.get(key);
+ const common = {
+ format: MOMENT_PAGE_FORMAT,
+ version: REPORT_VIEWS_VERSION,
+ key,
+ record: view.record,
+ quote: c.quote,
+ ...(c.speaker ? { speaker: c.speaker } : {}),
+ ...(c.date ? { date: c.date } : {}),
+ ...(c.verification ? { verification: c.verification } : {}),
+ citedIn: citedInViews(entries, views),
+ } as const;
+ if (c.kind === "post") {
+ const post = postsByChannel.get(c.channel)!.get(c.id)!;
+ const postView: MomentPostView = {
+ author: postAuthor(post),
+ text: post.text,
+ ...(entry?.kind === "post"
+ ? {
+ shot: `/media/${entry.file}`,
+ ...(entry.media.length > 0
+ ? {
+ media: entry.media.map((f) => ({
+ src: `/media/${f.file}`,
+ kind: IMAGE_EXTS.has(path.extname(f.file).toLowerCase()) ? ("image" as const) : ("video" as const),
+ })),
+ }
+ : {}),
+ }
+ : {}),
+ };
+ moments.push({ ...common, kind: "post", post: postView });
+ continue;
+ }
+ const m = parseMomentKey(key) as SpanMoment;
+ const info = momentInfo.get(key)!;
+ const span = evidenceSpan({ start: m.start, end: m.end, pad: info.pad });
+ const { cues } = records.get(`${c.channel}/${c.id}`)!;
+ const from = m.start - MOMENT_CUE_CONTEXT_SECONDS;
+ const to = m.end + MOMENT_CUE_CONTEXT_SECONDS;
+ const lines: CueLineView[] = cues
+ .filter((q) => q.end > from && q.start < to)
+ .slice(0, MOMENT_CUE_LINES_MAX)
+ .map((q) => ({ start: q.start, end: q.end, text: cueLine(q.text), inSpan: q.end > m.start && q.start < m.end }));
+ const audio = entry?.kind === "audio";
+ moments.push({
+ ...common,
+ // A span whose clip had to be cut as sound is heard, not watched.
+ kind: audio ? "audio" : info.kind === "audio" ? "audio" : "video",
+ start: m.start,
+ end: m.end,
+ ...(entry && entry.kind !== "post" ? { clip: { src: clipPath(m, entry.kind), start: span.from, end: span.to } } : {}),
+ cues: lines,
+ });
+ }
+ moments.sort((a, b) => (a.key < b.key ? -1 : a.key > b.key ? 1 : 0));
+
+ // ─── Writing ───
+
+ await writeOut(publicDir, REPORTS_INDEX_PATH, json(index));
+ for (let i = 0; i < reports.length; i++) {
+ const report = reports[i];
+ const view = views[i];
+ await writeOut(publicDir, reportViewPath(report.id), json(view));
+ await writeOut(publicDir, reportCitationsDownloadPath(report.id, "json"), json(citationSet(report, view)));
+ await writeOut(publicDir, reportCitationsDownloadPath(report.id, "csv"), citationsCsv(view));
+ for (const c of orderedCitations(view)) {
+ if (c.kind !== "source" || !c.image) continue;
+ const rel = (report.citations![c.id] as { image: string }).image;
+ await copyOut(publicDir, path.join(siteReportDir(paths, site.siteId, report.id), rel), c.image);
+ }
+ }
+ await writeOut(
+ publicDir,
+ MOMENTS_INDEX_PATH,
+ json({ format: MOMENT_INDEX_FORMAT, version: REPORT_VIEWS_VERSION, moments: moments.map((m) => m.key) }),
+ );
+ let copied = 0;
+ for (const m of moments) {
+ await writeOut(publicDir, momentViewPath(m.key), json(m));
+ const entry = mediaOf.get(m.key);
+ if (!entry) continue;
+ if (entry.kind === "post") {
+ for (const f of [entry.file, ...entry.media.map((x) => x.file)]) {
+ await copyOut(publicDir, path.join(cacheDir, f), `/media/${f}`);
+ copied++;
+ }
+ } else {
+ await copyOut(publicDir, path.join(cacheDir, entry.file), clipPath(parseMomentKey(m.key) as SpanMoment, entry.kind));
+ copied++;
+ }
+ }
+ log(
+ `[reports] ${reports.length} report(s), ${moments.length} moment page(s), ${copied} media file(s)` +
+ `${allowed.length ? `, ${allowed.length} without media` : ""}.`,
+ );
+ return { reports: index.reports, moments: moments.map((m) => m.key), allowed };
+}
+
+// A clip's published path: the moment's (lib/report/views.ts), `.m4a` for a
+// clip cut as sound.
+function clipPath(m: SpanMoment, kind: "video" | "audio"): string {
+ const p = evidenceClipPath(m);
+ return kind === "audio" ? p.replace(/\.mp4$/, ".m4a") : p;
+}
+
+function postAuthor(post: Post): string {
+ const handle = post.author.startsWith("@") ? post.author : `@${post.author}`;
+ return post.authorName ? `${post.authorName} (${handle})` : handle;
+}
+
+const defined = <T extends object>(o: T): T =>
+ Object.fromEntries(Object.entries(o).filter(([, v]) => v !== undefined)) as T;
+
+// The site-root routes the reports add to a sitemap: the index, each report,
+// each moment page.
+export function reportRoutes(composed: Pick<ComposedReports, "reports" | "moments">): string[] {
+ if (composed.reports.length === 0) return [];
+ return ["/reports/", ...composed.reports.map((r) => r.href), ...composed.moments.map((k) => momentPath(k))];
+}
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -1,6 +1,7 @@
# Changelog
## [Unreleased]
+- **A site's build composes its reports, and a site that publishes only its reports ships nothing else.** Every site's compose now writes the reports its `site.json` publishes: each report's page and its citations as `citations.json` and `citations.csv` under `/reports/<id>/`, its cited stills, a page per cited moment with the record, the transcript lines around the span and every report that cites it, and the clips and post captures `archilyzer reports prepare` made for it, only the cited ones. Each quote is checked against the record as it is composed (a span's against its cues within 5 s either side, read from `en-orig` when the `en` track has no cues; a post's against its text) and the score, time and method are written into the citation, replacing any typed by hand. The build stops with the list of every problem before anything is written: an invalid report, a citation of a channel outside the site or of a post the site may not carry, a missing record, still or post, a quote that matches less than 60 % of what the record says, and a citation without prepared media or with media cut for another span (`--allow-missing-media` on `archilyzer compose site` and `build site` lets those two through, without a clip). A site with `publish: "cited"` removes everything corpus-shaped from `export/public` before it writes its reports, and its built `out/` is checked against what a cited site may hold: anything else, a file over 25 MiB or more than 20,000 files fails the build, and every deploy path (the Publish tab, `deploy site`, Build & deploy, Build & deploy all, the container build) refuses it, as it refuses a site set to cited whose last build was a full one. The hub's compose removes a report site's files too.
- **A site has a Reports tab.** `/sites/<site>/reports` lists every report under the site's `reports/` directory — the published ones in their order, then the drafts — with its kind, dates, sections, claims, citations by kind and, for a fact-check, how many claims carry each verdict. Each report's problems, from the same checker the prepare step and the build use, open under it. A draft with no problems can be published, and a published report moved up or down or unpublished; each writes only the site's `reports` list, applied to the list as it is on disk at that moment, so it never overwrites another change to the site. "Prepare evidence media" queues the `reports-prepare` job, and beside it the tab shows the last prepared media (moments by kind, total size, problems by kind) and links the last prepare job. What the site publishes (full or cited) is shown with a link to Settings, where it is changed.
- **A site can say what it publishes, and which reports.** `site.json` takes `publish` — `"full"`, the searchable corpus every site has been (the default, never written), or `"cited"`, only the site's reports and the moments they cite — and `reports`, the ordered ids of its published reports (each a slug; invalid and repeated ids are dropped). The site form has a Publish control and lists the site's reports read-only; saving the form keeps the stored list. A cited site still builds as a full one until the reports pipeline applies the scope. SITE.md documents both keys.
- **A report and its citations now have one written format, checked before anything is built from them.** A cited report is a `report.json` (`archilyzer-report`, version 1): a summary, then sections of claims, each claim with an optional verdict, the reviewed document's own sentence, findings in markdown and the citations it rests on. A citation is one of five kinds — a span of a video, a span of an audio record, a post, a sentence of a source document, or a web page — with a verbatim quote, and is cited from any markdown in the report as `[label](cite:<id>)`. The checker lists every problem at once with where it is: a citation, a source or a `cite:` link that names nothing, a span that ends before it starts or runs past 120 seconds with its context, a still that points outside the report's folder, an id used twice. Each cited span and post has one page address, `/m/<channel>/<id>/<start>-<end>/` or `/m/<channel>/<id>/`. Nothing builds or shows reports yet. The fact-check verdicts (Corroborated, Partly true, Contradicted, Not found, Untestable) and their colours are now kept in one place, which the report video's stamps and tally read too. `REPORT.md` and `CITATIONS.md` list every key.
diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md
@@ -1,6 +1,7 @@
# Changelog
## [Unreleased]
+- **`corpus.json` is spec 5: it names a site's reports, and a site that publishes only reports says so.** A site with reports adds `reports` to its `corpus.json` (`index`: `/reports/index.json`, the count, and how to read a report's page, its citations and its moment pages) and a Reports section to `llms.txt`; its sitemap lists the report and moment pages. A site that publishes only its reports has `"scope": "cited"` and its audience under `site`, no channels and zero totals, an `llms.txt` that lists its reports and how their citations and moment pages are read, and a `site.json` with no channels. A reader that does not know spec 5 sees an empty corpus there. Needs a rebuild and deploy of each site.
- **A site can show cited reports, and every citation opens on a page of its own.** A site built with reports has a **Reports** link in its header and a page at `/reports/` listing them. A report's page has its title, subtitle, dates and the document under review with its archive links; a fact-check's tally of verdicts; the summary; the sections and their claims, each with its verdict, the document's own sentence as an image, the findings and the evidence cards; a numbered reference list; and links to download its citations as JSON and CSV. A citation in the text shows as its words plus a number: hovering it, focusing the number or tapping it once shows a card of the citation (the quote, who said it and when, a picture or the post's screenshot, and how closely the quote matched the transcript when it was checked); the words open what it cites and the number jumps to its reference. A cited span of a video or audio record opens at `/m/<channel>/<id>/<start>-<end>/`: a short clip of the span with a little context either side, the quote, the transcript lines around it, the record's title, channel and date, a link to the original at that time, and every report on the site that cites it. A cited post opens at `/m/<channel>/<id>/` with its screenshot and text. A site that publishes only its reports (`site.json` `publish: "cited"`) opens on the report index and has no search, Ask AI, downloads or duplicates. A site with no reports is unchanged. Needs a rebuild and deploy of each site.
## [0.11.1] - 2026-10-01