commit a4468f46b7adee23e49ad825f217d5a2ac8d676b parent 77099b41c695b06247ecb60c23e2fdfcd8c1a0b9 Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st> Date: Mon, 5 Oct 2026 12:32:50 -0400 Merge site/search-toggle (site.json search switch replaces publish; search off = report-only; export start serves moment dirs) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Diffstat:
36 files changed, 326 insertions(+), 273 deletions(-)
diff --git a/ENVIRONMENT.md b/ENVIRONMENT.md @@ -110,7 +110,7 @@ Every local server's default port, from `common/lib/ports.mjs`. The primary chec | `EDITOR_PORT` | `3001` | Editor real dev/start (`pnpm dev:editor`). A worktree adds its offset (`pnpm wt list`). | common/lib/ports.mjs | | `PORT` | `3011` | Editor test server + Playwright editor baseURL. A worktree adds its offset (`pnpm wt list`). | common/lib/ports.mjs | | `EXPORT_PORT` | `3010` | Export server launched by the editor e2e. A worktree adds its offset (`pnpm wt list`). | common/lib/ports.mjs | -| `EXPORT_DEV_PORT` | `3000` | Export real dev (`pnpm dev:export`). A worktree adds its offset (`pnpm wt list`). | common/lib/ports.mjs | +| `EXPORT_DEV_PORT` | `3000` | Export real dev (`pnpm dev:export`) and its built `out/` (`pnpm start:export`). A worktree adds its offset (`pnpm wt list`). | common/lib/ports.mjs | | `EXPORT_E2E_PORT` | `3020` | Export's own Playwright suite. A worktree adds its offset (`pnpm wt list`). | common/lib/ports.mjs | | `OLLAMA_STUB_PORT` | `11435` | Digest-lane stub server in the editor e2e suite. A worktree adds its offset (`pnpm wt list`). | common/lib/ports.mjs | | `HOMEPAGE_DEV_PORT` | `3030` | Homepage real dev (`pnpm dev:homepage`). A worktree adds its offset (`pnpm wt list`). | common/lib/ports.mjs | diff --git a/PUBLISH.md b/PUBLISH.md @@ -89,7 +89,8 @@ site or a post the site may not carry, a missing record, still or post, a quote longer cites. `--allow-missing-media` (on `compose site` and `build site`) lets the last two through; their pages render without a clip. -A **cited** site (`publish: "cited"`) publishes its reports and nothing else. Its +A **cited** (report-only) site is one with search off (`site.json` `search: false`; +a legacy `publish: "cited"` reads the same): it publishes its reports and nothing else. Its compose removes every corpus-shaped file the shared `export/public` holds — summaries, stats, transcripts, subs, posts, digests, archives, the duplicates, tags, aliases and chart files, the service worker — writes `site.json` with no channels and `corpus.json` @@ -97,8 +98,8 @@ with `site.scope: "cited"` (spec 5), and an `llms.txt` and sitemap listing the r and moment pages. After `next build`, the built `out/` is audited: a cited build that holds anything but `_next/`, the reports, the moment pages, the cited media, the shell's own pages and files, or a file over 25 MiB, or more than 20,000 files, fails -the build, and every deploy path refuses it. A site switched to cited is refused at -deploy until it is built again. +the build, and every deploy path refuses it. A site switched to search off is refused at +deploy until it is built again. With no published reports it is an empty report index. A cited site is not a searchable archive, so the family never lists it, whatever `listed` says (`isListedSite`): it is no hub member, has no homepage card, counts in @@ -113,10 +114,11 @@ fact-check citing every kind, its moments and their media) in a stage project at repo root, `.e2e-report-site/` — never in `export/`, whose `public/` and `out/` are the checkout's — audits its `out/` as above and reads it in a browser. `pnpm --filter export run build:report-fixture` builds and audits it without the -browser. The suite serves the build with `export/e2e-report/serve.mjs`, not `serve`: -`serve` (export's `start`, the container's `site` service) takes a video or audio -moment's directory name, `<start>-<end>` in seconds with decimals, for a file with an -extension, and lists the directory instead of serving its page. +browser. The suite, and export's `start` (`pnpm start:export`), serve the build with +`export/scripts/serve-out.mjs`, not `serve`: `serve` (still the container's `site` +service) takes a video or audio moment's directory name, `<start>-<end>` in seconds +with decimals, for a file with an extension, and lists the directory instead of +serving its page. Deploy-only ships whatever is in `export/out`, which the basic build composes one site at a time into a single shared directory — so it **refuses, before starting a diff --git a/SITE.md b/SITE.md @@ -25,9 +25,9 @@ Regenerate this file with `pnpm --filter yt-dlp-transcript-common exec tsx bin/f | [`siteUrl`](#siteurl) | absent | | [`listed`](#listed) | `true` | | [`audience`](#audience) | absent | -| [`publish`](#publish) | `"full"` | | [`reports`](#reports) | `[]` | | [`relatedSites`](#relatedsites) | `[]` | +| [`search`](#search) | `true` | | [`pwa`](#pwa) | `false` | | [`archives`](#archives) | `true` | | [`duplicates`](#duplicates) | `true` | @@ -163,7 +163,7 @@ Default: absent ## `listed` -Whether the family lists this site. Opt-OUT: absent/true = listed, only an explicit `false` is written. An unlisted site still builds and deploys as before, and its own pages are unchanged; it is left out of the homepage (cards, chart, `/stats`), the hub (members, federated search, `/corpus.json`, `/llms.txt`), every other site's footer, and the published `channel-sites.json` and pooled `stats/`. A channel only unlisted sites expose is in none of the family's public totals; a channel a listed site also exposes is credited to the listed one. A private site and a cited site (`publish: "cited"`) are never listed, whatever `listed` says. +Whether the family lists this site. Opt-OUT: absent/true = listed, only an explicit `false` is written. An unlisted site still builds and deploys as before, and its own pages are unchanged; it is left out of the homepage (cards, chart, `/stats`), the hub (members, federated search, `/corpus.json`, `/llms.txt`), every other site's footer, and the published `channel-sites.json` and pooled `stats/`. A channel only unlisted sites expose is in none of the family's public totals; a channel a listed site also exposes is credited to the listed one. A private site and a report-only site (`search: false`) are never listed, whatever `listed` says. Default: `true` @@ -173,12 +173,6 @@ Who this site is built for. `"public"` (the default; absent) or `"private"`: the Default: absent -## `publish` - -What this site publishes. `"full"` (the default; absent): the searchable corpus of its `channels`. `"cited"`: only the site's `reports` and the moments they cite — no search, no browse, no full transcripts, no archives; `channels` is then the pool its citations may resolve against. The cited scope is applied by the reports pipeline at build time. A cited site is not a searchable archive, so the family does not list it (as `listed: false`, whatever `listed` says): no hub membership, no homepage card or totals, no footer link from other sites. Only `"cited"` is written; any other value reads as `"full"`. - -Default: `"full"` - ## `reports` The site's published reports, in display order: report ids (lowercase slugs, `[a-z0-9][a-z0-9-]*`), each a directory under `sites/<siteId>/reports/`. A report directory not named here is a draft and is not published. Invalid and repeated ids are dropped. Absent/empty = no reports. @@ -208,6 +202,12 @@ Default: [] ``` +## `search` + +Whether this site publishes its searchable corpus. Opt-OUT: absent/true = on, only an explicit `false` is written. Off = report-only: the site publishes only its `reports` and the moments they cite — no search, browse, transcripts or archives; `channels` is the pool its citations resolve against, and its `/corpus.json` says `"scope": "cited"`. A report-only site is never listed (as `listed: false`, whatever `listed` says). Legacy `publish: "cited"` reads as `false`. + +Default: `true` + ## `pwa` Whether this site ships an installable PWA (service worker + web manifest). Default false: a "dumb instance" that serves the CORS-enabled JSON federation contract but is not independently installable, so a visitor trusts only the hub PWA. Stored only when true. diff --git a/common/bin/compose-hub.test.ts b/common/bin/compose-hub.test.ts @@ -172,7 +172,7 @@ test("an unlisted site is in none of the hub's files; a listed one is in each", } }); -// Report sites: a cited site (`publish: "cited"`) publishes reports, not a +// Report sites: a report-only site (`search: false`) publishes reports, not a // searchable archive — no hub member, whatever `listed` says. A full site with // reports stays one. test("a cited site is no hub member; a full site with reports is", async () => { @@ -192,7 +192,7 @@ test("a cited site is no hub member; a full site with reports is", async () => { write("fixture-cited", { siteTitle: "Cited Fixture", siteUrl: "https://fixture-cited.example", - publish: "cited", + search: false, listed: true, }); console.log = () => {}; diff --git a/common/bin/compose-site.ts b/common/bin/compose-site.ts @@ -15,7 +15,7 @@ // Only the site's member channels are composed, so each deployed bundle holds // only that site's data. Checked-in static assets in public/ are left intact. // -// A CITED site (site.json `publish: "cited"`) publishes its reports and the +// A CITED site (site.json `search: false`) publishes its reports and the // moments they cite, and nothing else: its compose removes every corpus-shaped // entry (CORPUS_PUBLIC_ENTRIES) instead of composing it, then writes the // reports, site.json (no channels) and corpus.json (`site.scope: "cited"`) — diff --git a/common/controller/poolSummary.test.ts b/common/controller/poolSummary.test.ts @@ -32,7 +32,7 @@ test("channel-sites.json names listed sites only; a channel only an unlisted sit test("channel-sites.json leaves out a cited site", () => { const sites = [ parseSite("fixture-a", { channels: [{ slug: "shared" }] }), - parseSite("fixture-cited", { publish: "cited", channels: [{ slug: "shared" }, { slug: "own" }] }), + parseSite("fixture-cited", { search: false, channels: [{ slug: "shared" }, { slug: "own" }] }), ]; assert.deepEqual(channelSitesOf(sites), { shared: ["fixture-a"] }); }); diff --git a/common/lib/builtExport.test.ts b/common/lib/builtExport.test.ts @@ -376,8 +376,8 @@ test("a site configured cited with a full build is refused at deploy; a cited bu const full = bundle({ site: { siteId: "reports-site" }, corpus: { spec: 5, site: { id: "reports-site" } } }); const cited = citedOut(); try { - const site = { siteId: "reports-site", publish: "cited" }; - assert.match(builtScopeProblem(site, full.dir)!, /publishes only its reports \(publish: cited\), but .* holds a full build/); + const site = { siteId: "reports-site", search: false }; + assert.match(builtScopeProblem(site, full.dir)!, /publishes only its reports \(search: false\), but .* holds a full build/); assert.equal(deployAudienceProblem(site, full.dir), builtScopeProblem(site, full.dir)); assert.equal(builtScopeProblem(site, cited.dir), null); assert.equal(builtScopeProblem({ siteId: "reports-site" }, full.dir), null); diff --git a/common/lib/builtExport.ts b/common/lib/builtExport.ts @@ -15,6 +15,7 @@ import { existsSync, readdirSync, readFileSync, statSync, type Dirent } from "node:fs"; import path from "node:path"; +import { isCitedSite } from "./siteSchema"; /** * The site id of the build sitting in `outDir`, or null when there is no @@ -149,13 +150,13 @@ export function builtAudienceProblem(outDir: string): string | null { * logs or throws, or null. */ export function deployAudienceProblem( - site: { siteId: string; audience?: string; publish?: string }, + site: { siteId: string; audience?: string; search?: boolean }, outDir: string, ): string | null { return siteDeployProblem(site) ?? builtAudienceProblem(outDir) ?? builtScopeProblem(site, outDir); } -// ─── A CITED build (site.json `publish: "cited"`) ─── +// ─── A CITED build (a report-only site: site.json `search: false`) ─── // // A cited site publishes its reports and the moments they cite, and nothing // else (plans/report-sites.md). export/public is shared by every site's build @@ -301,18 +302,18 @@ export function citedBuildProblem(outDir: string): string | null { /** * Why `outDir` may not be deployed as `site` because of what it PUBLISHES, as - * one sentence, or null: a site configured cited (`publish: "cited"`) whose + * one sentence, or null: a site configured report-only (`search: false`) whose * build is not a cited one was built before the switch, and would ship the * whole corpus. Asked beside the audience refusals (deployAudienceProblem). */ export function builtScopeProblem( - site: { siteId: string; publish?: string }, + site: { siteId: string; search?: boolean }, outDir: string, ): string | null { - if (site.publish !== "cited" || !existsSync(path.join(outDir, "corpus.json"))) return null; + if (!isCitedSite(site) || !existsSync(path.join(outDir, "corpus.json"))) return null; if (isCitedBuild(outDir)) return null; return ( - `Site "${site.siteId}" publishes only its reports (publish: cited), but ${outDir} ` + + `Site "${site.siteId}" publishes only its reports (search: false), but ${outDir} ` + `holds a full build — build ${site.siteId} again, then deploy` ); } diff --git a/common/lib/corpus.ts b/common/lib/corpus.ts @@ -37,7 +37,7 @@ import { // // v5: reports — a site may publish cited reports, announced as `reports` // ({ index: "/reports/index.json" }, the report index the export's Reports -// pages read); and a CITED site (site.json `publish: "cited"`) publishes ONLY +// pages read); and a CITED site (site.json `search: false`) publishes ONLY // them, saying so as `site.scope: "cited"` with no channels and zero totals. // An older reader sees an empty corpus there, which is the safe reading: there // is no shard to fetch. diff --git a/common/lib/homepageSummary.test.ts b/common/lib/homepageSummary.test.ts @@ -305,10 +305,10 @@ test("an unlisted site is in no array and no total; a channel it shares is the l assert.equal(s.version, 6); }); -// Report sites: a cited site (`publish: "cited"`) is not a searchable archive, +// Report sites: a report-only site (`search: false`) is not a searchable archive, // so the homepage lists it no more than an unlisted one. test("a cited site is in no array and no total, exactly as an unlisted one", () => { - const cited = { ...site("zeta", ["q1", "a1"], "https://zeta.example"), publish: "cited" } as Site; + const cited = { ...site("zeta", ["q1", "a1"], "https://zeta.example"), search: false } as Site; const channelSites = { ...CHANNEL_SITES, a1: ["alpha", "zeta"], q1: ["zeta"] }; const own = [stat({ channelSlug: "q1", id: "q-1", uploadDate: "20251101", duration: 7200 })]; const s = buildHomepageSummary([...STATS, ...own], channelSites, [...SITES, cited], NOW); diff --git a/common/lib/homepageSummary.ts b/common/lib/homepageSummary.ts @@ -45,7 +45,7 @@ import { VIDEO_STATES, type VideoState } from "./availability"; // `availability` included. No field was added or removed; the number says the // totals' scope moved. // -// Still v6: a cited report site (site.json `publish: "cited"`) is left out +// Still v6: a cited report site (site.json `search: false`) is left out // exactly as an unlisted one (lib/siteSchema.ts isListedSite) — the same scope // rule, applied to a kind of site no summary had yet counted. export const HOMEPAGE_SUMMARY_VERSION = 6; diff --git a/common/lib/ports.mjs b/common/lib/ports.mjs @@ -35,7 +35,7 @@ export const PORTS = Object.freeze({ EDITOR_PORT: { base: 3001, what: "editor real dev/start (`pnpm dev:editor`)" }, PORT: { base: 3011, what: "editor test server + Playwright editor baseURL" }, EXPORT_PORT: { base: 3010, what: "export server launched by the editor e2e" }, - EXPORT_DEV_PORT: { base: 3000, what: "export real dev (`pnpm dev:export`)" }, + EXPORT_DEV_PORT: { base: 3000, what: "export real dev (`pnpm dev:export`) and its built `out/` (`pnpm start:export`)" }, EXPORT_E2E_PORT: { base: 3020, what: "export's own Playwright suite" }, OLLAMA_STUB_PORT: { base: 11435, what: "digest-lane stub server in the editor e2e suite" }, HOMEPAGE_DEV_PORT: { base: 3030, what: "homepage real dev (`pnpm dev:homepage`)" }, diff --git a/common/lib/siteSchema.test.ts b/common/lib/siteSchema.test.ts @@ -66,7 +66,7 @@ test("empty, null, [] and a number all read as the defaults, every key emitted", assert.equal(want.pwa, false); assert.equal(want.listed, true); assert.deepEqual(want.relatedSites, []); - assert.equal(want.publish, "full"); + assert.equal(want.search, true); assert.deepEqual(want.reports, []); assert.equal(want.socialLinks, undefined); assert.ok("socialLinks" in want); @@ -239,7 +239,7 @@ function fixtures(): Array<[string, unknown]> { cloudflareProject: "p", siteUrl: "https://s.example//", listed: false, - publish: "cited", + search: false, reports: ["r-one", "r-one", "Bad", "", 7, "r-two"], relatedSites: [{ siteIds: ["x", "x", "BAD"] }, { label: " ", siteIds: [] }], pwa: true, @@ -316,18 +316,35 @@ test("listed: absent reads listed, only an explicit false unlists, and only fals assert.equal(getSite("s", paths).listed, true); }); -test("publish: absent reads full, only cited is kept, and only cited is written", () => { - for (const v of [undefined, "full", "FULL", "Cited", "corpus", true, 1, null]) { - const site = parseSite("s", v === undefined ? {} : { publish: v }); - assert.equal(site.publish, "full", String(v)); +test("search: absent/true is on, only false is report-only, and only false is written", () => { + for (const v of [undefined, true, "false", 0, null, "off"]) { + const site = parseSite("s", v === undefined ? {} : { search: v }); + assert.equal(site.search, true, String(v)); assert.equal(isCitedSite(site), false, String(v)); - assert.equal("publish" in siteToDisk(site), false, String(v)); + assert.equal("search" in siteToDisk(site), false, String(v)); } - const cited = parseSite("s", { publish: "cited" }); - assert.equal(cited.publish, "cited"); - assert.equal(isCitedSite(cited), true); - assert.equal(siteToDisk(cited).publish, "cited"); + const off = parseSite("s", { search: false }); + assert.equal(off.search, false); + assert.equal(isCitedSite(off), true); + assert.equal(siteToDisk(off).search, false); assert.equal(isCitedSite({}), false); + // A search-off site with no reports is still report-only (an empty index). + assert.equal(isCitedSite(parseSite("s", { search: false, reports: [] })), true); +}); + +test("legacy publish: \"cited\" reads as search off and is rewritten as search: false", () => { + const legacy = parseSite("s", { publish: "cited" }); + assert.equal(legacy.search, false); + assert.equal(isCitedSite(legacy), true); + assert.equal("publish" in legacy, false); + const disk = siteToDisk(legacy); + assert.equal(disk.search, false); + assert.equal("publish" in disk, false); + // Any other publish value is nothing; an explicit search wins over it. + for (const v of ["full", "Cited", true, 1]) { + assert.equal(parseSite("s", { publish: v }).search, true, String(v)); + } + assert.equal(parseSite("s", { publish: "cited", search: true }).search, true); }); test("reports: ordered slugs, invalid and repeated ids dropped, written only when non-empty", () => { @@ -347,19 +364,20 @@ test("reports: ordered slugs, invalid and repeated ids dropped, written only whe assert.deepEqual(siteToDisk({ ...site, reports: ["ok", "NOPE", "ok"] }).reports, ["ok"]); }); -test("writeSite → getSite round-trips a cited site and its reports; a full site writes neither key", async () => { +test("writeSite → getSite round-trips a report-only site and its reports; a full site writes neither key", async () => { const paths = scratchPaths(await mkdtemp(path.join(os.tmpdir(), "site-"))); - const cited = parseSite("s", { publish: "cited", reports: ["first", "second"] }); + const cited = parseSite("s", { search: false, reports: ["first", "second"] }); await writeSite(cited, paths); assert.deepEqual(getSite("s", paths), cited); const disk = JSON.parse(await readFile(siteConfigFile(paths, "s"), "utf8")); - assert.equal(disk.publish, "cited"); + assert.equal(disk.search, false); + assert.equal("publish" in disk, false); assert.deepEqual(disk.reports, ["first", "second"]); - await writeSite({ ...cited, publish: "full", reports: [] }, paths); + await writeSite({ ...cited, search: true, reports: [] }, paths); const plain = JSON.parse(await readFile(siteConfigFile(paths, "s"), "utf8")); - assert.equal("publish" in plain, false); + assert.equal("search" in plain, false); assert.equal("reports" in plain, false); - assert.equal(getSite("s", paths).publish, "full"); + assert.equal(getSite("s", paths).search, true); assert.deepEqual(getSite("s", paths).reports, []); }); @@ -376,14 +394,14 @@ test("isListedSite is the key's default; channelsOnlyOnUnlistedSites keeps a sha assert.deepEqual([...channelsOnlyOnUnlistedSites([sites[0]])], []); }); -test("a cited site is never listed, whatever `listed` says; a full site with reports is listed as before", () => { - assert.equal(isListedSite({ publish: "cited" }), false); - assert.equal(isListedSite({ publish: "cited", listed: true }), false); - assert.equal(isListedSite({ publish: "full" }), true); +test("a report-only site is never listed, whatever `listed` says; a full site with reports is listed as before", () => { + assert.equal(isListedSite({ search: false }), false); + assert.equal(isListedSite({ search: false, listed: true }), false); + assert.equal(isListedSite({ search: true }), true); assert.equal(isListedSite(parseSite("full-with-reports", { reports: ["demo-report"] })), true); const sites = [ parseSite("shown", { channels: [{ slug: "shared" }] }), - parseSite("reports", { publish: "cited", channels: [{ slug: "shared" }, { slug: "pool-only" }] }), + parseSite("reports", { search: false, channels: [{ slug: "shared" }, { slug: "pool-only" }] }), ]; // A cited site's channels are its citations' pool, not a published archive. assert.deepEqual([...channelsOnlyOnUnlistedSites(sites)], ["pool-only"]); @@ -419,6 +437,7 @@ test("writeSite → getSite round-trips, and the file holds only non-defaults", assert.equal("transcriptDownloads" in disk, false); assert.equal("pwa" in disk, false); assert.equal("listed" in disk, false); + assert.equal("search" in disk, false); assert.equal("publish" in disk, false); assert.equal("reports" in disk, false); }); @@ -431,7 +450,7 @@ test("patchSite applies a patch to the site on disk now, keeps every other key, const site = parseSite("s", { siteTitle: "Mine", archives: false, - publish: "cited", + search: false, reports: ["one", "two"], }); await writeSite(site, paths); diff --git a/common/lib/siteSchema.ts b/common/lib/siteSchema.ts @@ -90,22 +90,21 @@ export function isPrivateSite(site: Pick<Site, "audience">): boolean { return site.audience === "private"; } -// What a site publishes (report sites). "full" (the default, never written) is -// the searchable corpus every site has always been. "cited" publishes only the -// site's reports and the moments they cite; the reports pipeline applies that -// scope, and until it does a cited site builds as a full one. -export type SitePublish = "full" | "cited"; - -export const SITE_PUBLISH_SCOPES: readonly SitePublish[] = ["full", "cited"]; - -export function isSitePublish(v: unknown): v is SitePublish { - return v === "full" || v === "cited"; +// THE ONE PREDICATE for a report-only ("cited") site: search is off. Such a +// site publishes only its reports and the moments they cite (corpus.json +// `site.scope: "cited"`). Absent or true is a full, searchable site. +export function isCitedSite(site: Pick<Site, "search">): boolean { + return site.search === false; } -// THE ONE PREDICATE for a cited-only site. Absent or anything but "cited" reads -// as full. -export function isCitedSite(site: Pick<Site, "publish">): boolean { - return site.publish === "cited"; +// LEGACY: `publish: "cited"` (the pre-`search` report-only switch) reads as +// `search: false` when the file sets no `search`. siteToDisk never writes +// `publish`, so the next save migrates the file. +function migrateLegacyPublish(raw: unknown): unknown { + if (!raw || typeof raw !== "object" || Array.isArray(raw)) return raw; + const r = raw as Record<string, unknown>; + if (r.publish !== "cited" || r.search !== undefined) return raw; + return { ...r, search: false }; } // A report id: a lowercase slug, the directory name under @@ -149,9 +148,9 @@ export type Site = { siteUrl?: string; listed?: boolean; audience?: SiteAudience; - publish?: SitePublish; reports?: string[]; relatedSites?: RelatedSiteGroup[]; + search?: boolean; pwa?: boolean; archives?: boolean; duplicates?: boolean; @@ -184,15 +183,15 @@ export const SITE_FIELD_DOCS: FieldDocs<Site> = { siteUrl: "Absolute public URL of this site's deployment, e.g. `https://jeralyzer.pages.dev` (trimmed, trailing slashes removed; anything not absolute http(s) is dropped). Drives the cross-site footer: a site with no siteUrl is omitted from every other site's list.", listed: - "Whether the family lists this site. Opt-OUT: absent/true = listed, only an explicit `false` is written. An unlisted site still builds and deploys as before, and its own pages are unchanged; it is left out of the homepage (cards, chart, `/stats`), the hub (members, federated search, `/corpus.json`, `/llms.txt`), every other site's footer, and the published `channel-sites.json` and pooled `stats/`. A channel only unlisted sites expose is in none of the family's public totals; a channel a listed site also exposes is credited to the listed one. A private site and a cited site (`publish: \"cited\"`) are never listed, whatever `listed` says.", + "Whether the family lists this site. Opt-OUT: absent/true = listed, only an explicit `false` is written. An unlisted site still builds and deploys as before, and its own pages are unchanged; it is left out of the homepage (cards, chart, `/stats`), the hub (members, federated search, `/corpus.json`, `/llms.txt`), every other site's footer, and the published `channel-sites.json` and pooled `stats/`. A channel only unlisted sites expose is in none of the family's public totals; a channel a listed site also exposes is credited to the listed one. A private site and a report-only site (`search: false`) are never listed, whatever `listed` says.", audience: 'Who this site is built for. `"public"` (the default; absent) or `"private"`: the operator\'s own reading copy, built on this machine and never deployed — every deploy path (Build & deploy, Deploy, `archilyzer deploy site`, Build & deploy all, docker/publish-site.sh) refuses it before any upload, while a build without a deploy still works. A private site is never listed (as `listed: false`, whatever `listed` says), publishes no `hubUrl`, and its `/corpus.json` says `"audience": "private"`. Content kept from the public — X posts while `social.x.visibility` is `"private"` — is built only into private sites. Only `"private"` is written.', - publish: - 'What this site publishes. `"full"` (the default; absent): the searchable corpus of its `channels`. `"cited"`: only the site\'s `reports` and the moments they cite — no search, no browse, no full transcripts, no archives; `channels` is then the pool its citations may resolve against. The cited scope is applied by the reports pipeline at build time. A cited site is not a searchable archive, so the family does not list it (as `listed: false`, whatever `listed` says): no hub membership, no homepage card or totals, no footer link from other sites. Only `"cited"` is written; any other value reads as `"full"`.', reports: "The site's published reports, in display order: report ids (lowercase slugs, `[a-z0-9][a-z0-9-]*`), each a directory under `sites/<siteId>/reports/`. A report directory not named here is a draft and is not published. Invalid and repeated ids are dropped. Absent/empty = no reports.", relatedSites: "Pulls specific siblings to the front of the footer's cross-site list, in named groups. Siblings not named here fall into a trailing \"Other sites\" group. Absent/empty = one flat list of every sibling.", + search: + "Whether this site publishes its searchable corpus. Opt-OUT: absent/true = on, only an explicit `false` is written. Off = report-only: the site publishes only its `reports` and the moments they cite — no search, browse, transcripts or archives; `channels` is the pool its citations resolve against, and its `/corpus.json` says `\"scope\": \"cited\"`. A report-only site is never listed (as `listed: false`, whatever `listed` says). Legacy `publish: \"cited\"` reads as `false`.", pwa: "Whether this site ships an installable PWA (service worker + web manifest). Default false: a \"dumb instance\" that serves the CORS-enabled JSON federation contract but is not independently installable, so a visitor trusts only the hub PWA. Stored only when true.", archives: @@ -225,11 +224,11 @@ export function isValidSiteId(id: unknown): id is string { // // A PRIVATE site (`audience: "private"`) is never listed, whatever `listed` // says: it is never deployed, so there is nothing at its URL to list. Nor is a -// CITED site (`publish: "cited"`): it publishes reports and the moments they +// report-only site (`search: false`): it publishes reports and the moments they // cite, not a searchable archive, so it is no hub member (federated search // would find no channel there), no homepage card and in no family total. A // full site with reports is listed as before. -export function isListedSite(site: Pick<Site, "listed" | "audience" | "publish">): boolean { +export function isListedSite(site: Pick<Site, "listed" | "audience" | "search">): boolean { return site.listed !== false && !isPrivateSite(site) && !isCitedSite(site); } @@ -239,7 +238,7 @@ export function isListedSite(site: Pick<Site, "listed" | "audience" | "publish"> // and a channel no site exposes (pool-only) is not here either — the family's // instance-wide totals have always counted it. export function channelsOnlyOnUnlistedSites( - sites: readonly Pick<Site, "listed" | "audience" | "publish" | "channels">[], + sites: readonly Pick<Site, "listed" | "audience" | "search" | "channels">[], ): Set<string> { const onListed = new Set<string>(); const onUnlisted = new Set<string>(); @@ -364,12 +363,10 @@ export const siteFieldsSchema = z.object({ audience: settingsField((v): SiteAudience | undefined => v === "private" ? "private" : undefined, ).describe(d.audience), - // Only "cited" is kept; absent (and anything else) is the full default. - publish: settingsField((v): SitePublish => - v === "cited" ? "cited" : "full", - ).describe(d.publish), reports: settingsField(parseSiteReports).describe(d.reports), relatedSites: settingsField(parseRelatedSites).describe(d.relatedSites), + // Opt-out: only an explicit false is report-only. Absent/true stays on. + search: settingsField((v): boolean => v !== false).describe(d.search), pwa: settingsField((v): boolean => v === true).describe(d.pwa), // Opt-out: only an explicit false disables. Absent/true stays on. archives: settingsField((v): boolean => v !== false).describe(d.archives), @@ -389,7 +386,7 @@ export const siteFieldsSchema = z.object({ // output lacks `socialLinks`, `accent`, … on a file that does not spell them — // where parseSite has always emitted them, present and `undefined`. Rebuilding // from SITE_KEYS keeps that shape (and the key order) exactly. -export const siteSchema = siteFieldsSchema.transform((s): Site => { +export const siteSchema = z.preprocess(migrateLegacyPublish, siteFieldsSchema).transform((s): Site => { const groups = s.groups; const resolved: Partial<Record<keyof Site, unknown>> = { ...s, @@ -463,10 +460,10 @@ export function siteToDisk(site: Site): Site { ...(site.listed === false ? { listed: false } : {}), // Public is the default: only the private audience is persisted. ...(isPrivateSite(site) ? { audience: "private" as const } : {}), - // Full is the default: only the cited scope is persisted. - ...(isCitedSite(site) ? { publish: "cited" as const } : {}), ...(reports.length > 0 ? { reports } : {}), ...(relatedSites.length > 0 ? { relatedSites } : {}), + // Search is the default: only report-only is persisted (never `publish`). + ...(isCitedSite(site) ? { search: false } : {}), ...(site.pwa ? { pwa: true } : {}), // Persist only the non-default: archives is on unless explicitly disabled. ...(site.archives === false ? { archives: false } : {}), diff --git a/common/publish/composeReports.test.ts b/common/publish/composeReports.test.ts @@ -278,7 +278,7 @@ async function seedCorpus() { }, ]); - seedSite("cited", { publish: "cited", reports: [REPORT] }); + seedSite("cited", { search: false, reports: [REPORT] }); seedSite("full", { reports: [REPORT] }); seedSite("plain", {}, false); await buildIndex({ paths, onLog: () => {} }); diff --git a/common/publish/reportMedia.test.ts b/common/publish/reportMedia.test.ts @@ -108,7 +108,7 @@ for (const id of ["111", "222"]) { writeJson(path.join(paths.sitesDir, SITE, "site.json"), { title: "Demo", channels: [{ slug: CH }, { slug: X }], - publish: "cited", + search: false, reports: ["r1", "r2", "r-gone"], }); writeJson( diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -4,7 +4,7 @@ - **A report video's cue lookup names a site that publishes only its reports.** Pointing a report-to-video manifest at such a site (`corpus.json` `site.scope: "cited"`) used to fail with "channel … is not in corpus.json"; it now says the site publishes no transcripts and to use a full archive or a local corpus. The site form's **Publish** hint says a cited-only site is never listed on the homepage or the hub. - **A site's build composes its reports, and a site that publishes only its reports ships nothing else.** Every site's compose now writes the reports its `site.json` publishes: each report's page and its citations as `citations.json` and `citations.csv` under `/reports/<id>/`, its cited stills, a page per cited moment with the record, the transcript lines around the span and every report that cites it, and the clips and post captures `archilyzer reports prepare` made for it, only the cited ones. Each quote is checked against the record as it is composed (a span's against its cues within 5 s either side, read from `en-orig` when the `en` track has no cues; a post's against its text) and the score, time and method are written into the citation, replacing any typed by hand. The build stops with the list of every problem before anything is written: an invalid report, a citation of a channel outside the site or of a post the site may not carry, a missing record, still or post, a quote that matches less than 60 % of what the record says, and a citation without prepared media or with media cut for another span (`--allow-missing-media` on `archilyzer compose site` and `build site` lets those two through, without a clip). A site with `publish: "cited"` removes everything corpus-shaped from `export/public` before it writes its reports, and its built `out/` is checked against what a cited site may hold: anything else, a file over 25 MiB or more than 20,000 files fails the build, and every deploy path (the Publish tab, `deploy site`, Build & deploy, Build & deploy all, the container build) refuses it, as it refuses a site set to cited whose last build was a full one. The hub's compose removes a report site's files too. - **A site has a Reports tab.** `/sites/<site>/reports` lists every report under the site's `reports/` directory — the published ones in their order, then the drafts — with its kind, dates, sections, claims, citations by kind and, for a fact-check, how many claims carry each verdict. Each report's problems, from the same checker the prepare step and the build use, open under it. A draft with no problems can be published, and a published report moved up or down or unpublished; each writes only the site's `reports` list, applied to the list as it is on disk at that moment, so it never overwrites another change to the site. "Prepare evidence media" queues the `reports-prepare` job, and beside it the tab shows the last prepared media (moments by kind, total size, problems by kind) and links the last prepare job. What the site publishes (full or cited) is shown with a link to Settings, where it is changed. -- **A site can say what it publishes, and which reports.** `site.json` takes `publish` — `"full"`, the searchable corpus every site has been (the default, never written), or `"cited"`, only the site's reports and the moments they cite — and `reports`, the ordered ids of its published reports (each a slug; invalid and repeated ids are dropped). The site form has a Publish control and lists the site's reports read-only; saving the form keeps the stored list. A cited site still builds as a full one until the reports pipeline applies the scope. SITE.md documents both keys. +- **A site can publish reports, and with search off it is only its reports.** `site.json` takes `reports`, the ordered ids of its published reports (each a slug; invalid and repeated ids are dropped), and `search` (default on; only `false` is written): a site with search off publishes only its reports and the moments they cite. The site form has a **Search** checkbox and lists the site's reports read-only (saving the form keeps the stored list), and the Reports tab says which the site is. SITE.md documents both keys. - **A report and its citations now have one written format, checked before anything is built from them.** A cited report is a `report.json` (`archilyzer-report`, version 1): a summary, then sections of claims, each claim with an optional verdict, the reviewed document's own sentence, findings in markdown and the citations it rests on. A citation is one of five kinds — a span of a video, a span of an audio record, a post, a sentence of a source document, or a web page — with a verbatim quote, and is cited from any markdown in the report as `[label](cite:<id>)`. The checker lists every problem at once with where it is: a citation, a source or a `cite:` link that names nothing, a span that ends before it starts or runs past 120 seconds with its context, a still that points outside the report's folder, an id used twice. Each cited span and post has one page address, `/m/<channel>/<id>/<start>-<end>/` or `/m/<channel>/<id>/`. Nothing builds or shows reports yet. The fact-check verdicts (Corroborated, Partly true, Contradicted, Not found, Untestable) and their colours are now kept in one place, which the report video's stamps and tally read too. `REPORT.md` and `CITATIONS.md` list every key. - **A site's reports can have their evidence media prepared: every cited span cut to a clip, every cited post's capture copied.** `archilyzer reports prepare <site>`, the `reports-prepare` job (`POST /api/ops/reports-prepare`, `pnpm ops reports-prepare`) reads the site's published reports, checks them, and for each cited moment cuts the span (with its context) out of the media already on disk — a fetched clip window, the saved video, or the recording's audio — fitted inside 1280×720 with H.264 and AAC, or as an `.m4a` for an audio span; and copies the screenshot and attached media of each cited post, and of no other post, beside them. Everything lands in the site's build staging (`.export-index/sites/<site>/report-media/`) with an `index.json` naming each moment's file, size, checksum, size in pixels and duration. A clip is cut once and reused while its source file and span are unchanged; a clip or capture no longer cited is removed. Nothing is fetched: a citation whose media is not on disk, a post without a screenshot, a clip over 24 MiB, a citation of a channel outside the site, a post the site may not show, or an invalid or missing report is listed with the citations it affects, and the run fails (exit 1, or a failed job) — a span on a drive that is not mounted is reported as such rather than as missing. The lookup of a span's media on disk is now shared with report-to-video, which finds the same files it did. - **A report can be made from a /sweep report, an /ask answer or a report video's manifest, and a report video's manifest from a report.** `archilyzer reports convert sweep|ask|manifest <in> --out <report.json>` writes a `report.json`. From a /sweep report (markdown): its first `#` heading is the title, the text before the first section the summary, each `##` section a section, and each list item or paragraph with a citing link a claim — a line that is only a citation under a quote cites that quote — with every archive moment link, archive post link, moment page link and X or Bluesky post link turned into a citation and the link into `[label](cite:<id>)`; a citation's quote is the quoted words nearest before its link. From an /ask answer (`{ "answer", "sources" }`, a saved chat or an assistant message, as JSON): each `[n]` or `[n @ mm:ss]` marker becomes a citation of source `n` at the excerpt line nearest the marked time, quoting that line, or of the post. From a manifest: chapter cards are the sections, entries carrying `claim` the claims with their verdicts (the first still a claim's source sentence), clips the video citations, stills the source citations and `posts` the post citations; ledger rows are claims too. A /sweep or /ask citation names one second: with `--channels-dir <transcripts/channels>` its span is the cited cue run on to the quote's length and widened to whole sentences, as `resolve-windows.mjs` widens a clip (the widening now lives in common and is shared), and a post cited by its X or Bluesky link is found in the channel that archives it; without it the span is the second plus 10 seconds and such a post stays a plain link. `archilyzer reports to-manifest <report.json> --out <manifest.json>` writes a starter manifest: a title card, a chapter card per section, the section's cited clips, and per claim its source sentence's still and its clips, each with the claim's verdict, and its posts attached to its last clip (a post's platform, link and date come from the channels tree, or for an X post from its id). Both check what they write against the report format and write nothing when it has problems (exit 1); everything a conversion had to leave out or guess is printed as a warning. A manifest image may carry `quote`, the words the still shows, which the converters keep. diff --git a/editor/app/sites/[siteId]/reports/page.tsx b/editor/app/sites/[siteId]/reports/page.tsx @@ -58,8 +58,8 @@ export default async function SiteReportsPage({ <p className="text-sm"> <span className="font-medium">Publishes:</span>{" "} {cited - ? "Cited moments only — its reports and the moments they cite" - : "Full archive — the searchable corpus of its channels"}{" "} + ? "Reports only — search is off" + : "Full archive — search is on"}{" "} <span className="text-muted-foreground"> (changed on{" "} <Link href={`/sites/${siteId}`} className="underline"> @@ -70,8 +70,8 @@ export default async function SiteReportsPage({ </p> {cited && publishedCount === 0 && ( <p role="status" className="text-sm text-warning"> - This site publishes only cited moments and has no published report: - its build would have nothing to show. + Search is off and no report is published: the site is an empty + report index. </p> )} </section> diff --git a/editor/app/sites/actions.ts b/editor/app/sites/actions.ts @@ -118,8 +118,8 @@ export async function saveSiteAction( return { ok: false, error: `Unknown audience "${audienceRaw}".`, values }; } const isPrivate = audienceRaw === "private"; - // What the site publishes (report sites): only "cited" is kept. The report - // list is not on this form; the stored one is carried over. + // Search (opt-out; off = report-only) and the report list, which is not on + // this form: the stored one is carried over. let storedReports: Pick<Site, "reports"> | undefined; try { storedReports = getSite(siteId, getPaths()); @@ -127,7 +127,6 @@ export async function saveSiteAction( storedReports = undefined; } const reportKeys = reportSiteKeysFromForm(formData, storedReports); - if (!reportKeys.ok) return { ok: false, error: reportKeys.error, values }; // Hub parent (per-site override of the family default) + PWA opt-in. const hubUrlRaw = String(formData.get("hubUrl") ?? "").trim(); @@ -266,7 +265,7 @@ export async function saveSiteAction( // The Site is rebuilt from the form: a key missing here is dropped on save. ...(listed ? {} : { listed: false }), ...(isPrivate ? { audience: "private" as const } : {}), - ...reportKeys.keys, + ...reportKeys, ...(hubUrl ? { hubUrl } : {}), ...(pwa ? { pwa: true } : {}), ...(archives ? {} : { archives: false }), diff --git a/editor/app/sites/components/SiteForm.tsx b/editor/app/sites/components/SiteForm.tsx @@ -361,26 +361,18 @@ export function SiteForm({ initial, channels, allSites, isNew }: Props) { into while Settings keeps X posts private. </span> </label> - <label className="flex flex-col gap-1 text-sm"> - <span className="font-medium">Publish</span> - <SeededSelect - state={state} - name="publish" - aria-label="Publish" - initial={initial.publish === "cited" ? "cited" : "full"} - className="w-fit rounded border border-border bg-card px-2 py-1 text-sm" - > - <option value="full">Full archive — the searchable corpus of its channels</option> - <option value="cited">Cited moments only — its reports and the moments they cite</option> - </SeededSelect> - <span className="text-xs text-muted-foreground"> - A cited-only site publishes no search, browse, full transcripts or - archives; its channels are the pool its reports' citations resolve - against. The reports pipeline applies this scope when the site is built. - It is never listed on the homepage or the hub: it is not a searchable - archive. - </span> + <label className="flex items-center gap-2 text-sm"> + <input + type="checkbox" + name="search" + defaultChecked={seedChecked(state, "search", initial.search !== false)} + className="accent-brand" + /> + Search </label> + <p className="-mt-2 text-xs text-muted-foreground"> + Off: the site is its reports only. + </p> <div className="flex flex-col gap-1 text-sm"> <span className="font-medium">Reports</span> {initial.reports && initial.reports.length > 0 ? ( diff --git a/editor/app/sites/lib/reportSiteKeys.test.ts b/editor/app/sites/lib/reportSiteKeys.test.ts @@ -10,36 +10,22 @@ function form(entries: Record<string, string>): FormData { return fd; } -test("full (or no field) writes no publish key; cited is kept", () => { - assert.deepEqual(reportSiteKeysFromForm(form({ publish: "full" }), undefined), { - ok: true, - keys: {}, - }); - assert.deepEqual(reportSiteKeysFromForm(form({}), undefined), { ok: true, keys: {} }); - assert.deepEqual(reportSiteKeysFromForm(form({ publish: "cited" }), undefined), { - ok: true, - keys: { publish: "cited" }, - }); -}); - -test("an unknown publish scope is refused, not read as full", () => { - for (const v of ["Cited", "corpus", ""]) { - const r = reportSiteKeysFromForm(form({ publish: v }), undefined); - assert.equal(r.ok, false, v); - if (!r.ok) assert.match(r.error, /Unknown publish scope/); - } +test("search checked writes no key; unchecked (no field) writes search: false", () => { + assert.deepEqual(reportSiteKeysFromForm(form({ search: "on" }), undefined), {}); + assert.deepEqual(reportSiteKeysFromForm(form({}), undefined), { search: false }); + // The old select's field is not read: only the checkbox decides. + assert.deepEqual(reportSiteKeysFromForm(form({ search: "on", publish: "cited" }), undefined), {}); }); test("the stored report list is carried over in order; the form cannot set it", () => { const stored = { reports: ["second", "first"] }; assert.deepEqual( - reportSiteKeysFromForm(form({ publish: "full", reports: "injected" }), stored), - { ok: true, keys: { reports: ["second", "first"] } }, + reportSiteKeysFromForm(form({ search: "on", reports: "injected" }), stored), + { reports: ["second", "first"] }, ); // An empty or missing stored list writes no key; a bad stored id is dropped. - assert.deepEqual(reportSiteKeysFromForm(form({}), { reports: [] }), { ok: true, keys: {} }); - assert.deepEqual(reportSiteKeysFromForm(form({}), { reports: ["ok", "NOT OK"] }), { - ok: true, - keys: { reports: ["ok"] }, + assert.deepEqual(reportSiteKeysFromForm(form({ search: "on" }), { reports: [] }), {}); + assert.deepEqual(reportSiteKeysFromForm(form({ search: "on" }), { reports: ["ok", "NOT OK"] }), { + reports: ["ok"], }); }); diff --git a/editor/app/sites/lib/reportSiteKeys.ts b/editor/app/sites/lib/reportSiteKeys.ts @@ -1,35 +1,23 @@ -// The site form's two report-site keys (site.json `publish` and `reports`), +// The site form's two report-site keys (site.json `search` and `reports`), // pure so the parse is tested without a server (saveSiteAction calls it). // -// `publish` is the form's select; only "cited" is kept, as the file writes it. +// `search` is the form's checkbox, opt-out like `archives`: an unchecked box +// sends no field and is written as `search: false` (report-only). A legacy +// `publish: "cited"` read from the file is never written back. // `reports` is NOT edited by this form (the site's Reports tab manages it): the // Site is rebuilt from the form on save, so the list already stored is carried // over, and saving any other field never drops a published report. -import { - isSitePublish, - parseSiteReports, - type Site, -} from "yt-dlp-transcript-common/lib/site"; - -export type ReportSiteKeys = - | { ok: true; keys: Pick<Site, "publish" | "reports"> } - | { ok: false; error: string }; +import { parseSiteReports, type Site } from "yt-dlp-transcript-common/lib/site"; export function reportSiteKeysFromForm( formData: FormData, stored: Pick<Site, "reports"> | undefined, -): ReportSiteKeys { - const raw = String(formData.get("publish") ?? "full"); - if (!isSitePublish(raw)) { - return { ok: false, error: `Unknown publish scope "${raw}".` }; - } +): Pick<Site, "search" | "reports"> { + const search = formData.get("search") === "on"; const reports = parseSiteReports(stored?.reports); return { - ok: true, - keys: { - ...(raw === "cited" ? { publish: "cited" as const } : {}), - ...(reports.length > 0 ? { reports } : {}), - }, + ...(search ? {} : { search: false }), + ...(reports.length > 0 ? { reports } : {}), }; } diff --git a/editor/e2e/helpers.ts b/editor/e2e/helpers.ts @@ -153,6 +153,9 @@ export async function writeSite( : {}), ...(site.siteUrl ? { siteUrl: site.siteUrl } : {}), ...(site.listed === false ? { listed: false } : {}), + ...(site.search === false ? { search: false } : {}), + // The legacy report-only key, written as given (siteSchema migrates it). + ...(site.publish !== undefined ? { publish: site.publish } : {}), ...(site.relatedSites ? { relatedSites: site.relatedSites } : {}), }; await writeFile( diff --git a/editor/e2e/sites-crud.spec.ts b/editor/e2e/sites-crud.spec.ts @@ -345,6 +345,50 @@ test("archives + per-video transcript downloads opt-outs round-trip", async ({ await expect(downloads).toBeChecked(); }); +test("the search switch round-trips; a legacy publish: cited reads as off and saves as search: false", async ({ + page, +}) => { + await resetData("empty"); + await writeSite("reportonly", { siteTitle: "Report Only", publish: "cited" }); + + type SearchSiteFile = { search?: boolean; publish?: string }; + const file = "test-transcripts/sites/reportonly/site.json"; + const search = page.getByRole("checkbox", { name: "Search", exact: true }); + const save = async () => { + await page.getByRole("button", { name: /save site/i }).click(); + await expect( + page.getByRole("status").filter({ hasText: "Saved" }), + ).toBeVisible(); + }; + + // The legacy key reads as search off; saving rewrites it as search: false. + await page.goto("/sites/reportonly"); + await expect(search).not.toBeChecked(); + await save(); + await expect(async () => { + const site = await readJson<SearchSiteFile>(file); + expect(site.search).toBe(false); + expect("publish" in site).toBe(false); + }).toPass({ timeout: 10_000 }); + + await page.goto("/sites/reportonly/reports"); + await expect(page.getByText("Reports only — search is off")).toBeVisible(); + + // Ticking it back on removes the key — only the non-default is persisted. + await page.goto("/sites/reportonly"); + await search.check(); + await save(); + await expect(async () => { + const site = await readJson<SearchSiteFile>(file); + expect("search" in site).toBe(false); + }).toPass({ timeout: 10_000 }); + + await page.goto("/sites/reportonly"); + await expect(search).toBeChecked(); + await page.goto("/sites/reportonly/reports"); + await expect(page.getByText("Full archive — search is on")).toBeVisible(); +}); + test("the listed opt-out round-trips, and a save of another field keeps it", async ({ page, }) => { diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md @@ -1,8 +1,9 @@ # Changelog ## [Unreleased] +- **`pnpm start:export` serves a built site's moment pages.** It used `serve`, which listed a video or audio moment's directory (`3126.00-3151.00`) instead of serving its page; it now runs `export/scripts/serve-out.mjs`, which serves directories as Cloudflare Pages does, on `EXPORT_DEV_PORT` (3000). - **MCP: a site's reports can be read, and a site that publishes only reports says so instead of looking empty.** Two new tools: `list_reports` lists the reports a site publishes (id, title, kind, claim and citation counts, a fact-check's verdict tally, its page), and `get_report` reads one — its tally, then each section's claims with their verdicts and findings and, for every citation, the verbatim quote, the original (the platform at the cited second, the post, the document) and the site's moment page; `section` reads one section. On a site that publishes only its reports, `list_channels`, `list_sources` and `resolve_source` say "cited-only site: N report(s)" where they said "No channels found"; on a site with reports as well, `list_sources` and `resolve_source` say how many. The archive readers (local, remote) read `corpus.json` spec 5: a cited site is an empty corpus without an error, and a local copy of one is never read from a stale `transcripts/` folder. A hub has no reports of its own; `list_reports` says to name a member site. -- **The hub leaves out a site that publishes only its reports.** A site with `publish: "cited"` is not a searchable archive, so it is no hub member (not in federated search, the hub's `corpus.json` or `llms.txt`) and no other site's footer links to it, whatever its **List on the Archilyzer homepage and hub** setting says. A site with reports that publishes its full corpus is a member as before. Needs a hub rebuild and deploy once such a site exists. +- **The hub leaves out a site that publishes only its reports.** A site with search off (`search: false`) is not a searchable archive, so it is no hub member (not in federated search, the hub's `corpus.json` or `llms.txt`) and no other site's footer links to it, whatever its **List on the Archilyzer homepage and hub** setting says. A site with reports that publishes its full corpus is a member as before. Needs a hub rebuild and deploy once such a site exists. - **`corpus.json` is spec 5: it names a site's reports, and a site that publishes only reports says so.** A site with reports adds `reports` to its `corpus.json` (`index`: `/reports/index.json`, the count, and how to read a report's page, its citations and its moment pages) and a Reports section to `llms.txt`; its sitemap lists the report and moment pages. A site that publishes only its reports has `"scope": "cited"` and its audience under `site`, no channels and zero totals, an `llms.txt` that lists its reports and how their citations and moment pages are read, and a `site.json` with no channels. A reader that does not know spec 5 sees an empty corpus there. Needs a rebuild and deploy of each site. - **A site can show cited reports, and every citation opens on a page of its own.** A site built with reports has a **Reports** link in its header and a page at `/reports/` listing them. A report's page has its title, subtitle, dates and the document under review, its archive links listed once under it and folded away ("N archive links in context"); a fact-check's tally of verdicts; the summary; the sections and their claims, each with its verdict, the document's own sentence as an image with a link back to the document and only the archive links that sit in that sentence (at most five), the findings and the evidence cards; a numbered reference list; and links to download its citations as JSON and CSV. A citation in the text shows as its words plus a number: hovering it, focusing the number or tapping it once shows a card of the citation (the quote, who said it and when, a picture or the post's screenshot, and how closely the quote matched the transcript when it was checked); the words open what it cites and the number jumps to its reference. A cited span of a video or audio record opens at `/m/<channel>/<id>/<start>-<end>/`: a short clip of the span with a little context either side, the quote, the transcript lines around it, the record's title, channel and date, a link to the original at that time, and every report on the site that cites it. A cited post opens at `/m/<channel>/<id>/` with its screenshot and text. A site that publishes only its reports (`site.json` `publish: "cited"`) opens on the report index and has no search, Ask AI, downloads or duplicates. A site with no reports is unchanged. Needs a rebuild and deploy of each site. diff --git a/export/app/(workspace)/layout.tsx b/export/app/(workspace)/layout.tsx @@ -10,7 +10,7 @@ import SiteWorkspace from "./SiteWorkspace"; // // Hub mode keeps its own self-contained pages (<HubHome/>, <AskHub/>), so here // it's a straight pass-through — the fork is preserved exactly as before. So -// is a CITED site (site.json `publish: "cited"`): it has no corpus to search, +// is a CITED site (site.json `search: false`): it has no corpus to search, // and its home is the report index. export default function WorkspaceLayout({ children, diff --git a/export/app/components/Footer.tsx b/export/app/components/Footer.tsx @@ -19,7 +19,7 @@ import { isCitedSite } from "yt-dlp-transcript-common/lib/siteSchema"; export default function Footer() { const site = currentSite(); - // A cited site (site.json `publish: "cited"`) has no archives to download + // A cited site (site.json `search: false`) has no archives to download // and nothing for an AI to read; its footer keeps the rest. const cited = instanceMode() === "site" && isCitedSite(site); const showArchives = !cited && hasArchives(); diff --git a/export/app/lib/nav.ts b/export/app/lib/nav.ts @@ -9,7 +9,7 @@ export type NavLink = { href: string; label: string; external?: boolean }; // site, so a plain <a> in the same tab, not a client navigation. Reports // shows when this build has a report (public/reports/index.json). // -// A CITED site (site.json `publish: "cited"`) publishes only its reports and +// A CITED site (site.json `search: false`) publishes only its reports and // their moments: no search, no chat, no downloads or duplicates, and nothing // for an AI to read — its nav is Reports alone. export function headerNavLinks(has: { diff --git a/export/app/lib/reports.test.ts b/export/app/lib/reports.test.ts @@ -22,7 +22,7 @@ fs.mkdirSync(publicDir, { recursive: true }); fs.mkdirSync(path.join(sitesDir, "demo-cited"), { recursive: true }); fs.writeFileSync( path.join(sitesDir, "demo-cited", "site.json"), - JSON.stringify({ siteTitle: "Demo Reports", publish: "cited", reports: ["demo-factcheck"] }), + JSON.stringify({ siteTitle: "Demo Reports", search: false, reports: ["demo-factcheck"] }), ); process.env.EXPORT_PUBLIC_DIR = publicDir; process.env.TRANSCRIPTS_DIR = path.join(root, "transcripts"); diff --git a/export/e2e-report/fixtures/sites/reportsite/site.json b/export/e2e-report/fixtures/sites/reportsite/site.json @@ -4,6 +4,6 @@ "siteDescription": "A cited fixture site: one fact-check and the moments it cites.", "headerTitle": "Demo Reports", "siteUrl": "https://reports.example.org", - "publish": "cited", + "search": false, "reports": ["demo-factcheck"] } diff --git a/export/e2e-report/serve.mjs b/export/e2e-report/serve.mjs @@ -1,103 +0,0 @@ -// A static server for the staged cited build (playwright.report.config.ts), -// serving a directory the way Cloudflare Pages and the container's Caddy do: -// `/dir/` is `dir/index.html`, `/dir` redirects to `/dir/`. -// -// Not `serve` (export's `start`): serve-handler stats a path whose last segment -// looks like it has an extension as a file, and lists the directory instead of -// serving its index.html — and a span moment's directory is named -// `<start>-<end>` in seconds with two decimals (`3126.00-3151.00`), so every -// video and audio moment page came back as a directory listing. -// -// Byte ranges are honoured (a media element asks for them). -// -// node e2e-report/serve.mjs <dir> <port> - -import fs from "node:fs"; -import http from "node:http"; -import path from "node:path"; - -const [rootArg, portArg] = process.argv.slice(2); -if (!rootArg || !portArg) { - console.error("usage: serve.mjs <dir> <port>"); - process.exit(2); -} -const ROOT = path.resolve(rootArg); -const PORT = Number(portArg); - -const TYPES = { - ".html": "text/html; charset=utf-8", - ".txt": "text/plain; charset=utf-8", - ".json": "application/json; charset=utf-8", - ".js": "text/javascript; charset=utf-8", - ".css": "text/css; charset=utf-8", - ".csv": "text/csv; charset=utf-8", - ".xml": "application/xml; charset=utf-8", - ".webmanifest": "application/manifest+json", - ".svg": "image/svg+xml", - ".png": "image/png", - ".ico": "image/x-icon", - ".woff2": "font/woff2", - ".mp4": "video/mp4", - ".m4a": "audio/mp4", -}; - -function statOrNull(p) { - try { - return fs.statSync(p); - } catch { - return null; - } -} - -function sendFile(req, res, file, status = 200) { - const size = fs.statSync(file).size; - const headers = { - "Content-Type": TYPES[path.extname(file).toLowerCase()] ?? "application/octet-stream", - "Accept-Ranges": "bytes", - }; - const range = status === 200 ? /^bytes=(\d*)-(\d*)$/.exec(req.headers.range ?? "") : null; - if (range && (range[1] || range[2])) { - let start = range[1] ? Number(range[1]) : size - Number(range[2]); - let end = range[1] && range[2] ? Number(range[2]) : size - 1; - start = Math.max(0, start); - end = Math.min(end, size - 1); - if (start > end) { - res.writeHead(416, { "Content-Range": `bytes */${size}` }).end(); - return; - } - res.writeHead(206, { ...headers, "Content-Range": `bytes ${start}-${end}/${size}`, "Content-Length": end - start + 1 }); - if (req.method === "HEAD") return void res.end(); - fs.createReadStream(file, { start, end }).pipe(res); - return; - } - res.writeHead(status, { ...headers, "Content-Length": size }); - if (req.method === "HEAD") return void res.end(); - fs.createReadStream(file).pipe(res); -} - -http - .createServer((req, res) => { - const url = new URL(req.url ?? "/", "http://localhost"); - let pathname; - try { - pathname = decodeURIComponent(url.pathname); - } catch { - return void res.writeHead(400).end(); - } - const file = path.join(ROOT, pathname); - if (file !== ROOT && !file.startsWith(ROOT + path.sep)) return void res.writeHead(403).end(); - const st = statOrNull(file); - if (st?.isDirectory()) { - if (!pathname.endsWith("/")) { - return void res.writeHead(308, { Location: `${url.pathname}/${url.search}` }).end(); - } - const index = path.join(file, "index.html"); - if (statOrNull(index)?.isFile()) return sendFile(req, res, index); - } else if (st?.isFile()) { - return sendFile(req, res, file); - } - const notFound = path.join(ROOT, "404.html"); - if (statOrNull(notFound)?.isFile()) return sendFile(req, res, notFound, 404); - res.writeHead(404).end(); - }) - .listen(PORT, () => console.log(`[report-site] serving ${ROOT} on http://localhost:${PORT}`)); diff --git a/export/package.json b/export/package.json @@ -14,7 +14,7 @@ "compose:hub": "tsx ../common/bin/archilyzer.ts compose hub", "build": "tsx ../common/bin/archilyzer.ts build site", "build:hub": "tsx ../common/bin/archilyzer.ts build hub", - "start": "serve out", + "start": "node scripts/serve-out.mjs out ${EXPORT_DEV_PORT:-3000}", "lint": "eslint", "test": "tsx --test \"app/**/*.test.ts\"", "e2e": "node ../scripts/queue-lock.mjs --ports EXPORT_E2E_PORT:3020 -- playwright test", diff --git a/export/playwright.report.config.ts b/export/playwright.report.config.ts @@ -9,7 +9,7 @@ import { REPORT_SITE_OUT, stageReportSite } from "./e2e-report/stage"; // The CITED report-site suite: a real static export of a cited fixture site // (one fact-check citing every kind, its moments, their media), staged and // built OUTSIDE export/ — e2e-report/stage.ts says why — and served as Pages -// serves a deployed site (e2e-report/serve.mjs; not `serve`, which lists a +// serves a deployed site (scripts/serve-out.mjs; not `serve`, which lists a // span moment's directory instead of serving its page). What the specs see is // what a cited build ships, and audit.spec.ts holds the same out/ to the cited // allowlist audit. @@ -34,7 +34,7 @@ export default defineConfig({ fullyParallel: false, workers: 1, webServer: { - command: `node e2e-report/serve.mjs ${REPORT_SITE_OUT} ${PORT}`, + command: `node scripts/serve-out.mjs ${REPORT_SITE_OUT} ${PORT}`, url: `${baseURL}/site.json`, timeout: 60_000, reuseExistingServer: !process.env.CI, diff --git a/export/scripts/serve-out.mjs b/export/scripts/serve-out.mjs @@ -0,0 +1,118 @@ +// A static server for a built site (export's `start`, and the report-site e2e +// in playwright.report.config.ts), serving a directory the way Cloudflare +// Pages does: `/dir/` is `dir/index.html`, `/dir` redirects to `/dir/`, `/x` +// is `x.html` when there is one. +// +// Not `serve`: serve-handler stats a path whose last segment looks like it has +// an extension as a file, and lists the directory instead of serving its +// index.html — and a span moment's directory is named `<start>-<end>` in +// seconds with two decimals (`3126.00-3151.00`), so every video and audio +// moment page came back as a directory listing. +// +// Byte ranges are honoured (a media element asks for them). +// +// node scripts/serve-out.mjs <dir> <port> + +import fs from "node:fs"; +import http from "node:http"; +import path from "node:path"; + +const [rootArg, portArg] = process.argv.slice(2); +if (!rootArg || !portArg) { + console.error("usage: serve-out.mjs <dir> <port>"); + process.exit(2); +} +const ROOT = path.resolve(rootArg); +const PORT = Number(portArg); + +const TYPES = { + ".html": "text/html; charset=utf-8", + ".txt": "text/plain; charset=utf-8", + ".json": "application/json; charset=utf-8", + ".js": "text/javascript; charset=utf-8", + ".css": "text/css; charset=utf-8", + ".csv": "text/csv; charset=utf-8", + ".xml": "application/xml; charset=utf-8", + ".webmanifest": "application/manifest+json", + ".svg": "image/svg+xml", + ".png": "image/png", + ".ico": "image/x-icon", + ".woff2": "font/woff2", + ".woff": "font/woff", + ".jpg": "image/jpeg", + ".jpeg": "image/jpeg", + ".webp": "image/webp", + ".gif": "image/gif", + ".mp4": "video/mp4", + ".webm": "video/webm", + ".m4a": "audio/mp4", + ".mp3": "audio/mpeg", + ".vtt": "text/vtt; charset=utf-8", + ".srt": "text/plain; charset=utf-8", + ".md": "text/markdown; charset=utf-8", + ".zip": "application/zip", + ".wasm": "application/wasm", +}; + +function statOrNull(p) { + try { + return fs.statSync(p); + } catch { + return null; + } +} + +function sendFile(req, res, file, status = 200) { + const size = fs.statSync(file).size; + const headers = { + "Content-Type": TYPES[path.extname(file).toLowerCase()] ?? "application/octet-stream", + "Accept-Ranges": "bytes", + }; + const range = status === 200 ? /^bytes=(\d*)-(\d*)$/.exec(req.headers.range ?? "") : null; + if (range && (range[1] || range[2])) { + let start = range[1] ? Number(range[1]) : size - Number(range[2]); + let end = range[1] && range[2] ? Number(range[2]) : size - 1; + start = Math.max(0, start); + end = Math.min(end, size - 1); + if (start > end) { + res.writeHead(416, { "Content-Range": `bytes */${size}` }).end(); + return; + } + res.writeHead(206, { ...headers, "Content-Range": `bytes ${start}-${end}/${size}`, "Content-Length": end - start + 1 }); + if (req.method === "HEAD") return void res.end(); + fs.createReadStream(file, { start, end }).pipe(res); + return; + } + res.writeHead(status, { ...headers, "Content-Length": size }); + if (req.method === "HEAD") return void res.end(); + fs.createReadStream(file).pipe(res); +} + +http + .createServer((req, res) => { + const url = new URL(req.url ?? "/", "http://localhost"); + let pathname; + try { + pathname = decodeURIComponent(url.pathname); + } catch { + return void res.writeHead(400).end(); + } + const file = path.join(ROOT, pathname); + if (file !== ROOT && !file.startsWith(ROOT + path.sep)) return void res.writeHead(403).end(); + const st = statOrNull(file); + if (st?.isDirectory()) { + if (!pathname.endsWith("/")) { + return void res.writeHead(308, { Location: `${url.pathname}/${url.search}` }).end(); + } + const index = path.join(file, "index.html"); + if (statOrNull(index)?.isFile()) return sendFile(req, res, index); + } else if (st?.isFile()) { + return sendFile(req, res, file); + } else if (!pathname.endsWith("/") && statOrNull(`${file}.html`)?.isFile()) { + return sendFile(req, res, `${file}.html`); + } + const notFound = path.join(ROOT, "404.html"); + if (statOrNull(notFound)?.isFile()) return sendFile(req, res, notFound, 404); + res.writeHead(404).end(); + }) + .listen(PORT, () => console.log(`[serve-out] serving ${ROOT} on http://localhost:${PORT}`)); diff --git a/mcp/README.md b/mcp/README.md @@ -331,7 +331,8 @@ verdict tally, then each claim with its verdict, its findings and, per citation, the quote, the original (the platform at the cited second, the post, the document) and the site's moment page (`/m/<channel>/<id>/<start>-<end>/`). -A **cited-only** site (`corpus.json` `site.scope: "cited"`) publishes its reports +A **cited-only** site (`corpus.json` `site.scope: "cited"`; built from a site with +search off, `site.json` `search: false`) publishes its reports and the moments they cite and nothing else: no channels, no transcripts, nothing to search. `list_channels`, `list_sources` and `resolve_source` say "cited-only site: N report(s)" there rather than reporting an empty corpus. Reports are per diff --git a/plans/report-sites.md b/plans/report-sites.md @@ -1,6 +1,6 @@ # Report sites — a site built around cited reports -Status: BUILT — R0–R8 merged 2026-10-05 (plus R4b page size); first private report site built locally. Follow-ups: a static server for moment dirs (serve lists them), cited corpus.json still describes shards, voice checks recorded in citations, saveSiteAction onto patchSite. +Status: BUILT — R0–R8 merged 2026-10-05 (plus R4b page size); first private report site built locally. 2026-10-05: `publish` replaced by a per-site `search` switch; report-only is derived (search off); export's `start` serves moment dirs (`scripts/serve-out.mjs`). Follow-ups: cited corpus.json still describes shards, voice checks recorded in citations, saveSiteAction onto patchSite. ## What it is @@ -82,6 +82,11 @@ Citations are the core; reports are one consumer. | `publish` | `"full"` | `"cited"` publishes only reports + their moments. Only the non-default is written. | | `reports` | `[]` | Ordered report ids under `reports/`. | +Ruling 2026-10-05: `publish` is replaced by `search` (boolean, default `true`; only `false` is written). A +site with search off is report-only — what `publish: "cited"` was. A legacy `publish: "cited"` reads as +`search: false` and is rewritten as such on save. The editor form's Publish select is a Search checkbox. +`corpus.json` keeps `site.scope: "cited"`. + `channels` keeps its meaning (the pool a site's citations may resolve against). Docs in `SITE_FIELD_DOCS`, `SITE.md` regenerated; the editor's site form gets the publish control; a new **Reports** tab under `/sites/<id>/` lists reports, their validation problems, and "Prepare evidence media".