Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit bcd1ef21054fec0f9f13600514e38c15ac7fb286
parent 521a327e919f623be547ce529c4aa894f2a6436d
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon,  5 Oct 2026 02:36:37 -0400

Merge report-r0-schema (the citation model in common/lib/citations and the report document in common/lib/report: schemas, validation, moment keys, inline cites, cited-in index, one verdict vocabulary; CITATIONS.md + REPORT.md)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

# Conflicts:
#	editor/CHANGELOG.md

Diffstat:
ACITATIONS.md | 134+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
AREPORT.md | 62++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/bin/archilyzer.ts | 2+-
Mcommon/bin/file-schemas-docs.ts | 15++++++++++-----
Acommon/lib/citations/citations.test.ts | 266+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/citations/docs.ts | 167+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/citations/inline.ts | 92+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/citations/moments.ts | 151++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/citations/schema.ts | 268+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/citations/validate.ts | 295+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/report/citedIn.ts | 50++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/report/docs.test.ts | 46++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/report/docs.ts | 78++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/report/report.test.ts | 241+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/report/schema.ts | 123+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/report/uses.ts | 85+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/report/validate.ts | 158+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/report/verdicts.mjs | 81+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/report/verdicts.ts | 40++++++++++++++++++++++++++++++++++++++++
Mcommon/package.json | 1+
Meditor/CHANGELOG.md | 1+
Mumtool/report-to-video/factcheck.mjs | 45+++++++++++++++++----------------------------
Mumtool/report-to-video/factcheck.test.mjs | 14++++++++++++--
23 files changed, 2379 insertions(+), 36 deletions(-)

diff --git a/CITATIONS.md b/CITATIONS.md @@ -0,0 +1,134 @@ +# Citations + +<!-- GENERATED by common/bin/file-schemas-docs.ts from the schemas and their *_FIELD_DOCS records — do not edit by hand. --> + +The one citation model, for every document that cites the corpus: a report (`report.json` — see [REPORT.md](REPORT.md)), a standalone citation set (`archilyzer-citations`, below), and whatever comes next. The schema is `common/lib/citations/schema.ts`; the value rules are `common/lib/citations/validate.ts`, which reports every problem with its JSON path instead of stopping at the first. + +A citation is a discriminated union on `kind`: `video` and `audio` (a span of a record's media), `post` (a post record), `source` (a sentence of a document under review) and `page` (a web page outside the corpus). Every kind carries the common fields. A container names its citations by id in a `citations` map, and its documents in a `sources` map; an id is letters, digits and `_ . : -`, starting with a letter or digit, at most 64, and no id names both a source and a citation. Unknown keys are refused, and so is an unknown kind: a new kind is a new version. + +**Citing inline.** In any markdown a document carries, `[label](cite:<id>)` cites the citation `<id>` (a link inside a code span or a fenced block is text, not a citation). Citations are numbered per document by first appearance in reading order — a citation cited again keeps its number — and each has the stable anchor `#c-<id>`. + +**Moments.** A `video` or `audio` citation opens the page `/m/<channel>/<id>/<start>-<end>/`, a `post` citation `/m/<channel>/<id>/`; `source` and `page` citations have no page of their own. `<start>` and `<end>` are written with two decimals, always (`12.50`, never `12.5`), so a moment has one URL; the pad is not part of it. `<channel>` and `<id>` are single safe path segments (letters, digits, `.`, `_`, `-`; a channel starts with a letter or digit, an id may start with `-`). Moment keys are `common/lib/citations/moments.ts`. + +**Spans.** `start` is at least 0, `end` is after it (by at least 0.01 s), each `pad` is at least 0, and the span with its pad is at most 120 s. + +**Paths.** A still (`image`) or a saved document (`saved`) is relative to the container's directory: `/`-separated, never absolute, with no empty, `.` or `..` segment. + +Regenerate this file with `pnpm --filter yt-dlp-transcript-common exec tsx bin/file-schemas-docs.ts`. + +## Every kind + +#### Common fields + +| Key | Required | Description | +|---|---|---| +| `quote` | yes | The cited words, VERBATIM — as the record says them (a span's cues, a post's text, the source's sentence, the page's text). Never a paraphrase: compose checks a span's quote against its cues and fails on drift. | +| `speaker` | no | Who says the quote, when that is not the record's own channel (a guest, a co-host, a caller). | +| `date` | no | When the quote was said or written: `YYYY`, `YYYY-MM`, `YYYY-MM-DD` or an ISO 8601 date-time. Absent = the record's own date. | +| `label` | no | A short display name for the citation (one line), used where its number alone would be too little. | +| `note` | no | An editorial note shown with the citation (plain text): context the quote needs. | +| `verification` | no | COMPUTED, not authored: what checking this citation found, written by compose. A hand-typed block that could not have come from a check (a score outside 0–1, a score without its time) is a validation problem. | + +#### `verification` + +| Key | Required | Description | +|---|---|---| +| `quoteScore` | no | How closely `quote` matches what the record says at the cited place, from 0 (nothing alike) to 1 (verbatim). Written by the check, with `quoteCheckedAt`. | +| `quoteCheckedAt` | no | When the quote was checked: an ISO 8601 date-time. Written with `quoteScore`. | +| `voiceChecked` | no | True when the speaker's voice in the cited span was checked against `speaker`. Only a span (video, audio) has a voice to check. | +| `method` | no | What did the checking, e.g. the cue-window comparison and its version. Free text, one line. | + +## The kinds + +#### `"video"` + +| Key | Required | Description | +|---|---|---| +| `kind` | yes | `"video"`. | +| `channel` | yes | The record's channel slug (its directory under `transcripts/channels/`). A path segment of the moment page. | +| `id` | yes | The record's video id (its directory under `data/`; may start with `-`). A path segment of the moment page. | +| `start` | yes | Where the cited span starts, in seconds from the start of the record. | +| `end` | yes | Where the cited span ends, in seconds; after `start`. | +| `pad` | no | Context around the span in the evidence clip, in seconds. Absent = none. The span plus its pad is at most 120 s. | + +#### `"audio"` + +| Key | Required | Description | +|---|---|---| +| `kind` | yes | `"audio"`: a span of a record with no picture worth showing (a podcast). Rendered with a poster. | +| `channel` | yes | The record's channel slug (its directory under `transcripts/channels/`). A path segment of the moment page. | +| `id` | yes | The record's video id (its directory under `data/`; may start with `-`). A path segment of the moment page. | +| `start` | yes | Where the cited span starts, in seconds from the start of the record. | +| `end` | yes | Where the cited span ends, in seconds; after `start`. | +| `pad` | no | Context around the span in the evidence clip, in seconds. Absent = none. The span plus its pad is at most 120 s. | + +#### `pad` (video, audio) + +| Key | Required | Description | +|---|---|---| +| `before` | no | Seconds of context before `start`; ≥ 0. Absent = 0. | +| `after` | no | Seconds of context after `end`; ≥ 0. Absent = 0. | + +#### `"post"` + +| Key | Required | Description | +|---|---|---| +| `kind` | yes | `"post"`. | +| `channel` | yes | The post's channel slug. A path segment of the moment page. | +| `id` | yes | The post's id. A path segment of the moment page. | +| `thread` | no | True to show the post with the thread it belongs to. Absent = the post alone. | + +#### `"source"` + +| Key | Required | Description | +|---|---|---| +| `kind` | yes | `"source"`: a sentence of a document under review. | +| `source` | yes | The id of the document in the container's `sources`. Its archive links are shown with the citation. | +| `image` | no | A still of the sentence as the document shows it: a path relative to the container's directory (e.g. `stills/a01.png`), never absolute, never leaving it. | + +#### `"page"` + +| Key | Required | Description | +|---|---|---| +| `kind` | yes | `"page"`: a web page outside the corpus. | +| `url` | yes | The page's address: an http(s) URL. | +| `title` | no | The page's title. | +| `archiveUrl` | no | An archived copy of the page (an http(s) URL), shown beside the live link. | + +## Sources + +The documents a `source` citation quotes, in a container's `sources` map by id. A saved copy (`saved`) is the input stills are shot from and is never published. + +#### `sources.<id>` + +| Key | Required | Description | +|---|---|---| +| `kind` | yes | What the document is: `article`, `page`, `video`, `post`, `document`, `other`. | +| `title` | yes | The document's title. | +| `url` | no | Where the document lives: an http(s) URL. | +| `publisher` | no | Who published it (the outlet, the site). | +| `author` | no | Who wrote it. | +| `date` | no | When it was published: `YYYY`, `YYYY-MM`, `YYYY-MM-DD` or an ISO 8601 date-time. | +| `archives` | no | The document's archive links, in context — as the document had them. Shown with every citation of it. | +| `saved` | no | A saved copy of the document, relative to the container's directory (e.g. `sources/s0/page.html`): the input the stills are shot from. NEVER published. | + +#### `sources.<id>.archives[]` + +| Key | Required | Description | +|---|---|---| +| `label` | yes | The link's text. | +| `url` | yes | The archived copy: an http(s) URL. | +| `context` | no | The words around the link in the document, so a reader sees what it was offered as. | + +## A citation set + +Citations outside a report — what a converter of a sweep, an answer or a video manifest emits — are one JSON document, format `"archilyzer-citations"`, version 1. + +#### The document + +| Key | Required | Description | +|---|---|---| +| `format` | yes | `"archilyzer-citations"`. | +| `version` | yes | `1`. | +| `sources` | no | The documents the `source` citations quote, by id. Absent = none. | +| `citations` | yes | The citations, by id. An id is what a `[label](cite:<id>)` link names. | diff --git a/REPORT.md b/REPORT.md @@ -0,0 +1,62 @@ +# report.json keys + +<!-- GENERATED by common/bin/file-schemas-docs.ts from the schemas and their *_FIELD_DOCS records — do not edit by hand. --> + +One cited report, format `"archilyzer-report"`, version 1, persisted to `transcripts/sites/<siteId>/reports/<reportId>/report.json` beside its `stills/` and `sources/<sourceId>/`; a relative path in it is relative to that directory. A site's `reports` list in `site.json` is the published, ordered list — see [SITE.md](SITE.md); a report directory it does not name is a draft. The schema is `common/lib/report/schema.ts`; its citations and sources are the citation model's — see [CITATIONS.md](CITATIONS.md). + +A **fact-check** (`"kind": "factcheck"`) is sections (chapters) of claims, each with a verdict and its findings. A **sweep** (`"kind": "sweep"`) is sections with no verdicts, or bodies that cite inline. Markdown fields (`summary`, a section's `body`, a claim's `findings`) cite with `[label](cite:<id>)`. + +`common/lib/report/validate.ts` reports every problem with its JSON path: an unknown key, a reference that names nothing (a listed citation, a claim's source sentence, a `cite:` link, the subject), a section or claim id used twice (they share one namespace: the report page's anchors), a sweep's claim with a verdict, `updated` before `published`, and every citation problem CITATIONS.md lists. Whether a still exists and whether a quote matches its cues are checked when the site is composed. + +Regenerate this file with `pnpm --filter yt-dlp-transcript-common exec tsx bin/file-schemas-docs.ts`. + +## The document + +| Key | Required | Description | +|---|---|---| +| `format` | yes | `"archilyzer-report"`. | +| `version` | yes | `1`. | +| `id` | yes | The report's id: a lowercase slug (`[a-z0-9][a-z0-9-]*`, at most 64), its directory name under `reports/` and the last segment of its page, `/reports/<id>/`. Must match the directory. | +| `kind` | yes | `"factcheck"` — sections of claims, each with a verdict — or `"sweep"` — sections with no verdicts, or bodies with inline citations. | +| `title` | yes | The report's title. | +| `subtitle` | no | A line under the title. | +| `summary` | no | The report's summary, in markdown, shown before the sections. May cite inline: `[label](cite:<id>)`. | +| `published` | no | When the report was published: `YYYY-MM-DD` or an ISO 8601 date-time with a zone. | +| `updated` | no | When it was last changed, in the same form; not before `published`. | +| `subject` | no | The document under review, when the report reviews one: `{ "source": "<id>" }`, an id in `sources`. | +| `verdicts` | no | Overrides of the shared verdict vocabulary's labels and colours, by verdict (`CORROBORATED`, `PARTLY`, `CONTRADICTED`, `NOT_FOUND`, `UNTESTABLE`): `{ "label": "…", "color": "#rrggbb" }`, each key optional. Absent = the shared defaults. | +| `sources` | no | The documents the report's `source` citations quote, by id — see [CITATIONS.md](CITATIONS.md). Absent = none. | +| `citations` | no | The report's citations, by id — see [CITATIONS.md](CITATIONS.md). A citation is cited from markdown with `[label](cite:<id>)` and listed under the claims that rest on it. Absent = none. | +| `sections` | yes | The report's sections, in order. | + +#### `sections[]` + +| Key | Required | Description | +|---|---|---| +| `id` | yes | The section's id (letters, digits, `_ . : -`; at most 64): its anchor on the report page. Unique among the report's section and claim ids. | +| `title` | yes | The section's heading. | +| `body` | no | Markdown under the heading. May cite inline. | +| `claims` | no | The section's claims, in order. Absent = none. | + +#### `sections[].claims[]` + +| Key | Required | Description | +|---|---|---| +| `id` | yes | The claim's id (letters, digits, `_ . : -`; at most 64): its anchor on the report page. Unique among the report's section and claim ids. | +| `text` | yes | The claim, as stated by the document under review (plain text). | +| `verdict` | no | The ruling on the claim: `CORROBORATED`, `PARTLY`, `CONTRADICTED`, `NOT_FOUND`, `UNTESTABLE`. A fact-check's claim may leave it out (not yet ruled); a sweep's carries none. | +| `sourceQuote` | no | The document's own sentence making the claim: `{ "citation": "<id>" }`, naming a `source` citation (its still is shown with the claim). | +| `findings` | no | What the evidence shows, in markdown, citing inline: `[label](cite:<id>)`. | +| `citations` | no | The citations the claim rests on, in the order they are listed under it. Each must exist; none twice. | + +## The verdicts + +The shared vocabulary (`common/lib/report/verdicts.mjs` — the one copy; report-to-video's fact-check stamps read it too), in the order a tally lists them. A report's `verdicts` overrides a label (one line, at most 24 characters) or a colour (`#rgb` or `#rrggbb`), each on its own. + +| Verdict | Default label | Default colour | +|---|---|---| +| `CORROBORATED` | Corroborated | `#3fbf7f` | +| `PARTLY` | Partly true | `#e3b23c` | +| `CONTRADICTED` | Contradicted | `#e5534b` | +| `NOT_FOUND` | Not found | `#8b93a7` | +| `UNTESTABLE` | Untestable | `#7d8fd6` | diff --git a/common/bin/archilyzer.ts b/common/bin/archilyzer.ts @@ -296,7 +296,7 @@ export const COMMANDS: Command[] = [ }, { path: ["docs", "files"], - usage: "[--check] write SITE.md + CHANNEL.md from the file schemas", + usage: "[--check] write SITE.md + CHANNEL.md + REPORT.md + CITATIONS.md from the file schemas", flags: { check: "boolean" }, run: async ({ flags }) => (await import("./file-schemas-docs")).main({ check: flags.check === true }), diff --git a/common/bin/file-schemas-docs.ts b/common/bin/file-schemas-docs.ts @@ -1,16 +1,17 @@ #!/usr/bin/env tsx -// WRITE SITE.md AND CHANNEL.md FROM THE FILE SCHEMAS. +// WRITE SITE.md, CHANNEL.md, REPORT.md AND CITATIONS.md FROM THE FILE SCHEMAS. // // Usage (from the repo root): // pnpm --filter yt-dlp-transcript-common exec tsx bin/file-schemas-docs.ts // pnpm --filter yt-dlp-transcript-common exec tsx bin/file-schemas-docs.ts --check // -// `--check` writes nothing and exits 1 if either committed file differs from +// `--check` writes nothing and exits 1 if any committed file differs from // what the schemas generate (the same claim common/lib/fileSchemaDocs.test.ts -// makes). The sibling of settings-example.ts (SETTINGS.md). +// and lib/report/docs.test.ts make). The sibling of settings-example.ts +// (SETTINGS.md). // -// Reads no site.json and no config.json: both outputs are functions of the -// schemas' docs records alone. +// Reads no site.json, config.json or report.json: every output is a function of +// the schemas' docs records alone. import { readFile, writeFile } from "node:fs/promises"; import path from "node:path"; @@ -19,6 +20,8 @@ import { renderChannelMarkdown, renderSiteMarkdown, } from "../lib/fileSchemaDocs"; +import { renderCitationsMarkdown } from "../lib/citations/docs"; +import { renderReportMarkdown } from "../lib/report/docs"; import { parseFlags } from "./_parseFlags"; import { runIfEntryPoint } from "./_cli"; @@ -27,6 +30,8 @@ const REPO = path.resolve(path.dirname(fileURLToPath(import.meta.url)), "..", ". const FILE_SCHEMA_OUTPUTS: ReadonlyArray<[string, () => string]> = [ ["SITE.md", renderSiteMarkdown], ["CHANNEL.md", renderChannelMarkdown], + ["REPORT.md", renderReportMarkdown], + ["CITATIONS.md", renderCitationsMarkdown], ]; // `check` writes nothing and returns 1 when a committed file is stale. diff --git a/common/lib/citations/citations.test.ts b/common/lib/citations/citations.test.ts @@ -0,0 +1,266 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + CITATION_KINDS, + citationSchema, + MAX_CITATION_SPAN_SECONDS, + type Citation, +} from "./schema"; +import { + citationProblems, + isPartialDate, + jsonPath, + parseCitationSet, + relativePathProblem, + validateCitationSet, + type Problem, +} from "./validate"; +import { + momentKey, + momentKeyOf, + momentOf, + momentPath, + parseMomentKey, + parseMomentPath, + isSafeIdSegment, + isSafeChannelSegment, +} from "./moments"; +import { citationAnchor, citeHref, extractCiteRefs, numberCitations } from "./inline"; + +const video = (over: Partial<Extract<Citation, { kind: "video" }>> = {}): Citation => ({ + kind: "video", + channel: "demo-channel", + id: "abc123", + start: 10, + end: 20, + quote: "the words as said", + ...over, +}); + +const paths = (ps: Problem[]) => ps.map((p) => p.path); + +const SET = { + format: "archilyzer-citations", + version: 1, + sources: { + s0: { + kind: "article", + title: "A demo article", + url: "https://example.test/article", + archives: [{ label: "archived", url: "https://archive.example.test/a", context: "as linked here" }], + saved: "sources/s0/page.html", + }, + }, + citations: { + c01: { kind: "video", channel: "demo-channel", id: "-abc123", start: 1.5, end: 9.25, quote: "a span", pad: { before: 5, after: 5 } }, + c02: { kind: "audio", channel: "demo-podcast", id: "ep42", start: 60, end: 75, quote: "an episode", speaker: "A guest" }, + p01: { kind: "post", channel: "demo-channel", id: "1234567890", quote: "a post", thread: true }, + a01: { kind: "source", source: "s0", quote: "the sentence", image: "stills/a01.png" }, + w01: { kind: "page", url: "https://example.test/page", title: "A page", archiveUrl: "https://archive.example.test/p", quote: "on the page" }, + }, +}; + +// ─── schema ─── + +test("a citation set with every kind parses with no problems", () => { + const r = parseCitationSet(SET); + assert.ok(r.ok); + assert.deepEqual(r.problems, []); +}); + +test("every kind is a member of the union, and the union is closed", () => { + assert.deepEqual( + CITATION_KINDS.map((s) => s.shape.kind.value), + ["video", "audio", "post", "source", "page"], + ); + assert.equal(citationSchema.safeParse({ kind: "tweet", quote: "x" }).success, false); +}); + +test("shape problems: unknown keys, wrong types, a missing quote, the wrong format or version", () => { + const bad = { + ...SET, + version: 2, + citations: { + c01: { ...SET.citations.c01, colour: "red" }, + c02: { ...SET.citations.c02, start: "60" }, + p01: { kind: "post", channel: "demo-channel", id: "1" }, + }, + }; + const r = parseCitationSet(bad); + assert.equal(r.ok, false); + const ps = paths(r.problems); + assert.ok(ps.includes("version"), ps.join()); + assert.ok(ps.includes("citations.c01"), ps.join()); // the unknown key + assert.ok(ps.includes("citations.c02.start"), ps.join()); + assert.ok(ps.includes("citations.p01.quote"), ps.join()); + assert.equal(parseCitationSet({ ...SET, format: "other" }).ok, false); +}); + +// ─── value rules ─── + +test("span rules: end after start, pad at least 0, at most the cap with the pad", () => { + assert.deepEqual(citationProblems(video(), ["c"]), []); + assert.deepEqual(paths(citationProblems(video({ end: 10 }), ["c"])), ["c.end"]); + assert.deepEqual(paths(citationProblems(video({ end: 5 }), ["c"])), ["c.end"]); + assert.deepEqual(paths(citationProblems(video({ start: -1 }), ["c"])), ["c.start"]); + assert.deepEqual(paths(citationProblems(video({ pad: { before: -1 } }), ["c"])), ["c.pad.before"]); + // At the cap exactly is fine; past it, counting the pad, is not. + assert.deepEqual(citationProblems(video({ start: 0, end: MAX_CITATION_SPAN_SECONDS }), ["c"]), []); + assert.deepEqual(citationProblems(video({ start: 0, end: 110, pad: { before: 5, after: 5 } }), ["c"]), []); + const over = citationProblems(video({ start: 0, end: 110, pad: { before: 5, after: 6 } }), ["c"]); + assert.deepEqual(paths(over), ["c"]); + assert.match(over[0].message, /121 s \(at most 120 s\)/); + // A span that vanishes at the key's precision. + assert.deepEqual(paths(citationProblems(video({ start: 1.001, end: 1.004 }), ["c"])), ["c.end"]); +}); + +test("channel and id must be safe path segments; an id may start with -", () => { + assert.deepEqual(citationProblems(video({ id: "-dQw4w9WgXcQ" }), ["c"]), []); + assert.deepEqual(citationProblems(video({ id: "_x" }), ["c"]), []); + for (const id of ["..", ".", "a/b", "a\\b", "", "a b", "a?b"]) { + assert.deepEqual(paths(citationProblems(video({ id }), ["c"])), ["c.id"], id); + } + for (const channel of ["-lead", "..", "a/b", ""]) { + assert.deepEqual(paths(citationProblems(video({ channel }), ["c"])), ["c.channel"], channel); + } +}); + +test("common fields: a blank quote, a two-line label, a bad date", () => { + const ps = citationProblems(video({ quote: " ", label: "a\nb", date: "2024-02-30", speaker: "" }), ["c"]); + assert.deepEqual(paths(ps).sort(), ["c.date", "c.label", "c.quote", "c.speaker"]); +}); + +test("verification is computed: values that no check writes are flagged", () => { + const ok = video({ verification: { quoteScore: 0.93, quoteCheckedAt: "2026-10-01T12:00:00Z", voiceChecked: true, method: "cue-window" } }); + assert.deepEqual(citationProblems(ok, ["c"]), []); + const ps = citationProblems(video({ verification: { quoteScore: 1.4, quoteCheckedAt: "yesterday" } }), ["c"]); + assert.deepEqual(paths(ps).sort(), ["c.verification.quoteCheckedAt", "c.verification.quoteScore"]); + assert.deepEqual(paths(citationProblems(video({ verification: { quoteScore: 0.5 } }), ["c"])), ["c.verification"]); + const post: Citation = { kind: "post", channel: "demo-channel", id: "1", quote: "q", verification: { voiceChecked: true } }; + assert.deepEqual(paths(citationProblems(post, ["p"])), ["p.verification.voiceChecked"]); +}); + +test("source and page citations: the source must exist, paths may not escape, URLs are http(s)", () => { + const set = structuredClone(SET) as typeof SET & { citations: Record<string, unknown> }; + set.citations.a02 = { kind: "source", source: "s9", quote: "q", image: "../outside.png" }; + set.citations.w02 = { kind: "page", url: "javascript:alert(1)", quote: "q", archiveUrl: "ftp://x" }; + (set.sources.s0 as { saved: string }).saved = "/abs/page.html"; + const ps = paths(validateCitationSet(set)); + assert.deepEqual(ps.sort(), [ + "citations.a02.image", + "citations.a02.source", + "citations.w02.archiveUrl", + "citations.w02.url", + "sources.s0.saved", + ]); +}); + +test("ids: reference ids only, and no id names both a source and a citation", () => { + const set = structuredClone(SET) as unknown as { citations: Record<string, unknown>; sources: Record<string, unknown> }; + set.citations["bad id"] = { ...SET.citations.p01 }; + set.citations.s0 = { ...SET.citations.p01 }; + const ps = validateCitationSet(set); + assert.deepEqual(paths(ps).sort(), ['citations["bad id"]', "sources.s0"]); +}); + +test("relative paths: relative, forward slashes, inside the directory", () => { + assert.equal(relativePathProblem("stills/a01.png"), null); + for (const p of ["/etc/x", "C:/x", "a\\b", "../x", "a/../../x", "a//b", "./a", "https://x/y", "", "a/"]) { + assert.notEqual(relativePathProblem(p), null, p); + } +}); + +test("dates: partial dates and zoned date-times, real calendar days only", () => { + for (const d of ["2024", "2024-02", "2024-02-29", "2026-10-01T12:00Z", "2026-10-01T12:00:00.5+02:00"]) { + assert.ok(isPartialDate(d), d); + } + for (const d of ["2023-02-29", "2024-13", "24-01-01", "2026-10-01T12:00", "October 2026"]) { + assert.ok(!isPartialDate(d), d); + } +}); + +test("jsonPath: plain keys dotted, other keys quoted, indexes bracketed", () => { + assert.equal(jsonPath(["sections", 0, "claims", 2, "findings"]), "sections[0].claims[2].findings"); + assert.equal(jsonPath(["citations", "c.1", "end"]), 'citations["c.1"].end'); + assert.equal(jsonPath([]), ""); +}); + +// ─── moments ─── + +test("moment keys: video and audio by span, post by id, none for source and page", () => { + const set = parseCitationSet(SET); + assert.ok(set.ok); + const c = set.value.citations; + assert.equal(momentKeyOf(c.c01), "demo-channel/-abc123/1.50-9.25"); + assert.equal(momentKeyOf(c.c02), "demo-podcast/ep42/60.00-75.00"); + assert.equal(momentKeyOf(c.p01), "demo-channel/1234567890"); + assert.equal(momentKeyOf(c.a01), null); + assert.equal(momentKeyOf(c.w01), null); + assert.equal(momentPath(momentOf(c.c01)!), "/m/demo-channel/-abc123/1.50-9.25/"); + assert.equal(momentPath("demo-channel/1234567890"), "/m/demo-channel/1234567890/"); +}); + +test("moment keys round-trip, an id starting with - included; the pad is not part of the key", () => { + for (const key of ["demo-channel/-abc123/0.00-12.34", "demo-channel/abc_123/3600.10-3610.00", "demo-channel/-1"]) { + const m = parseMomentKey(key); + assert.ok(m, key); + assert.equal(momentKey(m!), key); + assert.deepEqual(parseMomentPath(momentPath(m!)), m); + assert.deepEqual(parseMomentPath(`/m/${key}`), m); + } + assert.equal(momentKeyOf(video({ pad: { before: 3 } })), momentKeyOf(video())); + // Rounded to the key's precision, and the moment carries the rounded numbers. + const m = momentOf(video({ start: 1.234, end: 5.678 })); + assert.deepEqual(m, { kind: "span", channel: "demo-channel", id: "abc123", start: 1.23, end: 5.68 }); +}); + +test("moment keys: only the canonical spelling parses; unsafe segments never key", () => { + for (const key of [ + "demo-channel/abc/1.5-2.00", + "demo-channel/abc/01.00-2.00", + "demo-channel/abc/2.00-1.00", + "demo-channel/abc/1.00-1.00", + "demo-channel/abc/-1.00-2.00", + "demo-channel/../1.00-2.00", + "demo-channel", + "demo-channel/abc/1.00-2.00/x", + "-x/abc", + ]) { + assert.equal(parseMomentKey(key), null, key); + } + assert.equal(parseMomentPath("/reports/x/"), null); + assert.throws(() => momentKey({ kind: "post", channel: "demo-channel", id: "a/b" }), /safe path segment/); + assert.equal(momentKeyOf(video({ id: ".." })), null); + assert.ok(isSafeIdSegment("-abc") && !isSafeChannelSegment("-abc")); +}); + +// ─── inline citations ─── + +test("extractCiteRefs: every cite link in order, with labels and offsets; other links ignored", () => { + const md = "He said so [here](cite:c01) and [again](cite:c02), see [the site](https://example.test) and [x]( cite:c01 )."; + const refs = extractCiteRefs(md); + assert.deepEqual(refs.map((r) => [r.id, r.label]), [["c01", "here"], ["c02", "again"], ["c01", "x"]]); + assert.equal(md.slice(refs[0].offset, refs[0].offset + 6), "[here]"); + assert.deepEqual(extractCiteRefs(undefined), []); + assert.deepEqual(extractCiteRefs("[empty](cite:)").map((r) => r.id), [""]); +}); + +test("extractCiteRefs: a cite link in code is text about the syntax", () => { + const md = [ + "Write `[label](cite:id)` to cite, like [this](cite:c01).", + "```md", + "[in a fence](cite:c09)", + "```", + "~~~~", + "[tilde fence](cite:c08)", + "~~~~", + "And ``[double](cite:c07)`` too, then [after](cite:c02).", + ].join("\n"); + assert.deepEqual(extractCiteRefs(md).map((r) => r.id), ["c01", "c02"]); +}); + +test("numberCitations: by first appearance, a repeat keeps its number", () => { + assert.deepEqual([...numberCitations(["c03", "c01", "c03", "c02", "c01"])], [["c03", 1], ["c01", 2], ["c02", 3]]); + assert.equal(citeHref("c01"), "cite:c01"); + assert.equal(citationAnchor("c01"), "c-c01"); +}); diff --git a/common/lib/citations/docs.ts b/common/lib/citations/docs.ts @@ -0,0 +1,167 @@ +// CITATIONS.md — the citation model's key tables, GENERATED from the schemas +// (./schema.ts): the keys and their order from each member's zod shape, +// required-or-not from the same shape, the prose from the *_FIELD_DOCS records. +// +// A pure renderer, the sibling of lib/fileSchemaDocs.ts (SITE.md, CHANNEL.md): +// common/bin/file-schemas-docs.ts writes the file and citations/docs.test.ts +// asserts the committed bytes are what this returns, so it cannot be edited by +// hand. + +import type { z } from "zod"; +import { cell } from "../settingsDocs"; +import { + AUDIO_CITATION_FIELD_DOCS, + CITATION_COMMON_FIELD_DOCS, + CITATION_PAD_FIELD_DOCS, + CITATION_SET_FIELD_DOCS, + CITATION_SET_FORMAT, + CITATION_VERIFICATION_FIELD_DOCS, + CITATIONS_VERSION, + MAX_CITATION_SPAN_SECONDS, + PAGE_CITATION_FIELD_DOCS, + POST_CITATION_FIELD_DOCS, + SOURCE_ARCHIVE_FIELD_DOCS, + SOURCE_CITATION_FIELD_DOCS, + SOURCE_FIELD_DOCS, + VIDEO_CITATION_FIELD_DOCS, + audioCitationSchema, + citationSetSchema, + pageCitationSchema, + postCitationSchema, + sourceArchiveSchema, + sourceCitationSchema, + sourceSchema, + verificationSchema, + videoCitationSchema, +} from "./schema"; + +export const GENERATED_SCHEMA_DOC = + "<!-- GENERATED by common/bin/file-schemas-docs.ts from the schemas and their *_FIELD_DOCS records — do not edit by hand. -->"; + +export const REGENERATE_SCHEMA_DOC = + "Regenerate this file with " + + "`pnpm --filter yt-dlp-transcript-common exec tsx bin/file-schemas-docs.ts`."; + +type Shape = Record<string, z.ZodType>; + +// Whether a key of a zod object shape may be left out. +export function isOptionalKey(shape: Shape, key: string): boolean { + return shape[key].safeParse(undefined).success; +} + +// One key table, `Key | Required | Description`, in `docs` order; only the +// keys `docs` names (a member's own keys, when the common ones are shared). +export function renderKeyTable( + out: string[], + heading: string, + shape: Shape, + docs: Readonly<Record<string, string>>, +): void { + out.push(heading); + out.push(""); + out.push("| Key | Required | Description |"); + out.push("|---|---|---|"); + for (const [key, text] of Object.entries(docs)) { + if (!(key in shape)) throw new Error(`${heading}: ${key} is documented but not in the schema`); + out.push(`| \`${key}\` | ${isOptionalKey(shape, key) ? "no" : "yes"} | ${cell(text)} |`); + } + out.push(""); +} + +export function renderCitationsMarkdown(): string { + const out: string[] = []; + out.push("# Citations"); + out.push(""); + out.push(GENERATED_SCHEMA_DOC); + out.push(""); + out.push( + "The one citation model, for every document that cites the corpus: a report " + + "(`report.json` — see [REPORT.md](REPORT.md)), a standalone citation set " + + `(\`${CITATION_SET_FORMAT}\`, below), and whatever comes next. The schema is ` + + "`common/lib/citations/schema.ts`; the value rules are " + + "`common/lib/citations/validate.ts`, which reports every problem with its JSON " + + "path instead of stopping at the first.", + ); + out.push(""); + out.push( + "A citation is a discriminated union on `kind`: `video` and `audio` (a span of a " + + "record's media), `post` (a post record), `source` (a sentence of a document under " + + "review) and `page` (a web page outside the corpus). Every kind carries the common " + + "fields. A container names its citations by id in a `citations` map, and its " + + "documents in a `sources` map; an id is letters, digits and `_ . : -`, starting " + + "with a letter or digit, at most 64, and no id names both a source and a citation. " + + "Unknown keys are refused, and so is an unknown kind: a new kind is a new version.", + ); + out.push(""); + out.push( + "**Citing inline.** In any markdown a document carries, `[label](cite:<id>)` cites " + + "the citation `<id>` (a link inside a code span or a fenced block is text, not a " + + "citation). Citations are numbered per document by first appearance in reading " + + "order — a citation cited again keeps its number — and each has the stable anchor " + + "`#c-<id>`.", + ); + out.push(""); + out.push( + "**Moments.** A `video` or `audio` citation opens the page " + + "`/m/<channel>/<id>/<start>-<end>/`, a `post` citation `/m/<channel>/<id>/`; " + + "`source` and `page` citations have no page of their own. `<start>` and `<end>` " + + "are written with two decimals, always (`12.50`, never `12.5`), so a moment has " + + "one URL; the pad is not part of it. `<channel>` and `<id>` are single safe path " + + "segments (letters, digits, `.`, `_`, `-`; a channel starts with a letter or " + + "digit, an id may start with `-`). Moment keys are `common/lib/citations/moments.ts`.", + ); + out.push(""); + out.push( + `**Spans.** \`start\` is at least 0, \`end\` is after it (by at least 0.01 s), each \`pad\` ` + + `is at least 0, and the span with its pad is at most ${MAX_CITATION_SPAN_SECONDS} s.`, + ); + out.push(""); + out.push( + "**Paths.** A still (`image`) or a saved document (`saved`) is relative to the " + + "container's directory: `/`-separated, never absolute, with no empty, `.` or `..` " + + "segment.", + ); + out.push(""); + out.push(REGENERATE_SCHEMA_DOC); + out.push(""); + + out.push("## Every kind"); + out.push(""); + renderKeyTable(out, "#### Common fields", videoCitationSchema.shape, CITATION_COMMON_FIELD_DOCS); + renderKeyTable(out, "#### `verification`", verificationSchema.shape, CITATION_VERIFICATION_FIELD_DOCS); + + out.push("## The kinds"); + out.push(""); + renderKeyTable(out, '#### `"video"`', videoCitationSchema.shape, VIDEO_CITATION_FIELD_DOCS); + renderKeyTable(out, '#### `"audio"`', audioCitationSchema.shape, AUDIO_CITATION_FIELD_DOCS); + renderKeyTable( + out, + "#### `pad` (video, audio)", + videoCitationSchema.shape.pad.unwrap().shape, + CITATION_PAD_FIELD_DOCS, + ); + renderKeyTable(out, '#### `"post"`', postCitationSchema.shape, POST_CITATION_FIELD_DOCS); + renderKeyTable(out, '#### `"source"`', sourceCitationSchema.shape, SOURCE_CITATION_FIELD_DOCS); + renderKeyTable(out, '#### `"page"`', pageCitationSchema.shape, PAGE_CITATION_FIELD_DOCS); + + out.push("## Sources"); + out.push(""); + out.push( + "The documents a `source` citation quotes, in a container's `sources` map by id. " + + "A saved copy (`saved`) is the input stills are shot from and is never published.", + ); + out.push(""); + renderKeyTable(out, "#### `sources.<id>`", sourceSchema.shape, SOURCE_FIELD_DOCS); + renderKeyTable(out, "#### `sources.<id>.archives[]`", sourceArchiveSchema.shape, SOURCE_ARCHIVE_FIELD_DOCS); + + out.push("## A citation set"); + out.push(""); + out.push( + `Citations outside a report — what a converter of a sweep, an answer or a video ` + + `manifest emits — are one JSON document, format \`"${CITATION_SET_FORMAT}"\`, ` + + `version ${CITATIONS_VERSION}.`, + ); + out.push(""); + renderKeyTable(out, "#### The document", citationSetSchema.shape, CITATION_SET_FIELD_DOCS); + return out.join("\n"); +} diff --git a/common/lib/citations/inline.ts b/common/lib/citations/inline.ts @@ -0,0 +1,92 @@ +// INLINE CITATIONS — the one syntax for citing in markdown, and the numbers a +// reader sees. +// +// [label](cite:<id>) +// +// anywhere a document's markdown is (a report's summary, a section's body, a +// claim's findings). `<id>` names a citation in the same document. A renderer +// replaces the link with a numbered marker that previews the citation and +// jumps to it; the citation's anchor is `#c-<id>` (`citationAnchor`), stable +// across edits because it is the id, not the number. +// +// NUMBERS ARE PER DOCUMENT, BY FIRST APPEARANCE: the first citation cited is +// [1], the next new one [2], and a citation cited again keeps its number. The +// document decides the order it is read in (lib/report/uses.ts walks a +// report); `numberCitations` only counts. +// +// Code is not citing: a `cite:` link inside an inline code span or a fenced +// block is text about the syntax, and is skipped. +// +// Pure, no imports: the export site's pages and the browser can use it. + +export const CITE_SCHEME = "cite:"; + +export type CiteRef = { + id: string; + label: string; + // Where the link starts in the markdown (a UTF-16 offset). + offset: number; +}; + +// `[label](cite:id)`. The label may not contain `]`; the id runs to the first +// `)` or space (an empty id is still a ref, so validation can name it). +const CITE_LINK_RE = /\[([^\]]*)\]\(\s*cite:([^)\s]*)\s*\)/g; + +// Fenced blocks: a line opening with ``` or ~~~ (up to three spaces in) to +// the line closing it with at least as many of the same, or the end. +const FENCE_OPEN_RE = /^ {0,3}(`{3,}|~{3,})/; + +// An inline code span: a run of backticks, to the next run of the same length. +const CODE_SPAN_RE = /(?<!`)(`+)(?!`)[\s\S]*?(?<!`)\1(?!`)/g; + +const blank = (s: string) => s.replace(/[^\n]/g, " "); + +// The markdown with every fenced block and code span blanked to spaces, so +// offsets still point into the original. +function blankCode(md: string): string { + const lines = md.split("\n"); + let fence: string | null = null; + for (let i = 0; i < lines.length; i++) { + if (fence) { + const close = new RegExp(`^ {0,3}${fence[0] === "`" ? "`" : "~"}{${fence.length},}\\s*$`); + if (close.test(lines[i])) fence = null; + lines[i] = blank(lines[i]); + continue; + } + const open = FENCE_OPEN_RE.exec(lines[i]); + if (open) { + fence = open[1]; + lines[i] = blank(lines[i]); + } + } + return lines.join("\n").replace(CODE_SPAN_RE, blank); +} + +// Every `cite:` link in the markdown, in order. +export function extractCiteRefs(md: string | null | undefined): CiteRef[] { + if (!md) return []; + const scan = blankCode(md); + const out: CiteRef[] = []; + for (const m of scan.matchAll(CITE_LINK_RE)) { + out.push({ id: m[2], label: md.slice(m.index + 1, m.index + 1 + m[1].length), offset: m.index }); + } + return out; +} + +// Numbers by first appearance: each id's number, from 1, in the order the ids +// are given; a repeated id keeps its first number. +export function numberCitations(ids: Iterable<string>): Map<string, number> { + const out = new Map<string, number>(); + for (const id of ids) if (!out.has(id)) out.set(id, out.size + 1); + return out; +} + +// The href an inline citation is written with. +export function citeHref(id: string): string { + return `${CITE_SCHEME}${id}`; +} + +// The anchor (fragment id, without `#`) of a citation's entry in a document. +export function citationAnchor(id: string): string { + return `c-${id}`; +} diff --git a/common/lib/citations/moments.ts b/common/lib/citations/moments.ts @@ -0,0 +1,151 @@ +// MOMENT KEYS — the one name of a cited moment, and the URL of its page. +// +// span (video, audio) <channel>/<id>/<start>-<end> page /m/<channel>/<id>/<start>-<end>/ +// post <channel>/<id> page /m/<channel>/<id>/ +// +// `source` and `page` citations have no moment: they are shown where they are +// cited, never on a page of their own. +// +// THE KEY IS THE PAGE, so it must be a function of the cited numbers alone: +// `start` and `end` are written with TWO DECIMALS, always (`Number#toFixed(2)`, +// the same precision as a clip window's name, lib/clipWindow.ts), and a parse +// accepts only that spelling — `12.5` is not a key, `12.50` is — so one moment +// has exactly one URL. Two citations of the same span share a page; the pad is +// not part of the key (it widens the clip, not the moment). Consumers that cut +// or show the span use the ROUNDED numbers (`momentOf`), so the page, the clip +// file and the key agree. +// +// SLUG SAFETY: the channel and the id are path segments of a page that a static +// build writes to disk, so each must be one safe segment — no `/`, no `\`, not +// `.` or `..`, nothing outside `[A-Za-z0-9._-]`. A video id MAY START WITH `-` +// (YouTube ids do); that is a safe path segment, but such an id must never be +// handed to a command line as a bare argument. `momentKey` throws on an unsafe +// segment rather than build a path from it; `momentKeyOf` answers null. +// +// Pure, no imports but types: the export site's pages and the browser can use it. + +import type { Citation } from "./schema"; + +export const MOMENT_SECONDS_DECIMALS = 2; + +// The route prefix of every moment page. +export const MOMENT_ROUTE_PREFIX = "/m/"; + +// A channel slug: the corpus's own rule (controller/channels.ts +// CHANNEL_SLUG_RE, which lib may not import), capped in length. +const CHANNEL_SEGMENT_RE = /^[A-Za-z0-9][A-Za-z0-9._-]{0,127}$/; + +// A record id: like a slug, but it may also start with `-` or `_`. +const ID_SEGMENT_RE = /^[A-Za-z0-9_-][A-Za-z0-9._-]{0,127}$/; + +// A span segment as a key spells it: two decimals each, no sign, no leading +// zeros but the one before the point. +const SPAN_SEGMENT_RE = /^(0|[1-9]\d*)\.(\d{2})-(0|[1-9]\d*)\.(\d{2})$/; + +export function isSafeChannelSegment(v: unknown): v is string { + return typeof v === "string" && v !== ".." && CHANNEL_SEGMENT_RE.test(v); +} + +export function isSafeIdSegment(v: unknown): v is string { + return typeof v === "string" && v !== "." && v !== ".." && ID_SEGMENT_RE.test(v); +} + +export type SpanMoment = { kind: "span"; channel: string; id: string; start: number; end: number }; +export type PostMoment = { kind: "post"; channel: string; id: string }; +export type Moment = SpanMoment | PostMoment; + +// Seconds as a key spells them. +export function formatMomentSeconds(s: number): string { + return s.toFixed(MOMENT_SECONDS_DECIMALS); +} + +// Seconds rounded to the key's precision, as a number. +export function roundMomentSeconds(s: number): number { + return Number(formatMomentSeconds(s)); +} + +// The moment a citation opens, its span rounded to the key's precision; null +// for a kind without a page (source, page). +export function momentOf(c: Citation): Moment | null { + switch (c.kind) { + case "video": + case "audio": + return { + kind: "span", + channel: c.channel, + id: c.id, + start: roundMomentSeconds(c.start), + end: roundMomentSeconds(c.end), + }; + case "post": + return { kind: "post", channel: c.channel, id: c.id }; + default: + return null; + } +} + +// Why a moment cannot be keyed, as a sentence, or null when it can. +export function momentProblem(m: Moment): string | null { + if (!isSafeChannelSegment(m.channel)) return `channel ${JSON.stringify(m.channel)} is not a safe path segment`; + if (!isSafeIdSegment(m.id)) return `id ${JSON.stringify(m.id)} is not a safe path segment`; + if (m.kind === "span") { + if (!Number.isFinite(m.start) || !Number.isFinite(m.end) || m.start < 0) { + return "a span's start and end must be numbers, the start at least 0"; + } + if (!(roundMomentSeconds(m.end) > roundMomentSeconds(m.start))) { + return `the span ends (${formatMomentSeconds(m.end)}) at or before it starts (${formatMomentSeconds(m.start)}) at the key's precision`; + } + } + return null; +} + +// The moment's key. Throws on a moment `momentProblem` refuses. +export function momentKey(m: Moment): string { + const problem = momentProblem(m); + if (problem) throw new Error(`moment key: ${problem}`); + const base = `${m.channel}/${m.id}`; + return m.kind === "span" ? `${base}/${formatMomentSeconds(m.start)}-${formatMomentSeconds(m.end)}` : base; +} + +// The citation's moment key, or null: a kind without a page, or a citation +// whose moment cannot be keyed (validation names why). +export function momentKeyOf(c: Citation): string | null { + const m = momentOf(c); + if (!m || momentProblem(m)) return null; + return momentKey(m); +} + +// The page of a moment (or of a key): `/m/<key>/`. +export function momentPath(m: Moment | string): string { + return `${MOMENT_ROUTE_PREFIX}${typeof m === "string" ? m : momentKey(m)}/`; +} + +// A key back into its moment, or null for anything that is not a key exactly +// as `momentKey` spells it. +export function parseMomentKey(key: string): Moment | null { + const parts = key.split("/"); + if (parts.length === 2) { + const [channel, id] = parts; + if (!isSafeChannelSegment(channel) || !isSafeIdSegment(id)) return null; + return { kind: "post", channel, id }; + } + if (parts.length === 3) { + const [channel, id, spanPart] = parts; + if (!isSafeChannelSegment(channel) || !isSafeIdSegment(id)) return null; + const m = SPAN_SEGMENT_RE.exec(spanPart); + if (!m) return null; + const start = Number(`${m[1]}.${m[2]}`); + const end = Number(`${m[3]}.${m[4]}`); + if (!(end > start)) return null; + return { kind: "span", channel, id, start, end }; + } + return null; +} + +// A moment page's path (`/m/<key>/`, the trailing slash optional) back into +// its moment, or null. +export function parseMomentPath(p: string): Moment | null { + if (!p.startsWith(MOMENT_ROUTE_PREFIX)) return null; + const key = p.slice(MOMENT_ROUTE_PREFIX.length).replace(/\/$/, ""); + return parseMomentKey(key); +} diff --git a/common/lib/citations/schema.ts b/common/lib/citations/schema.ts @@ -0,0 +1,268 @@ +// THE CITATION MODEL — one definition of what a citation is, for every document +// that cites the corpus: a report (`report.json`, lib/report/), a standalone +// citation set (`archilyzer-citations`, below — what a converter of a sweep, an +// /ask answer or a video manifest emits), and whatever comes next. +// +// A citation is a discriminated union on `kind`: +// +// video a span of a video record: channel, id, start, end, pad +// audio a span of an audio record (a podcast): the same fields as video; +// rendered with a poster instead of a picture +// post a post record: channel, id, thread +// source a sentence of a document under review: the source's id in +// `sources`, a still of the sentence (`image`); rendered with that +// source's archive links +// page a web page that is not in the corpus: url, title, archiveUrl +// +// and every kind carries the common fields: a VERBATIM `quote`, and optional +// `speaker`, `date`, `label`, `note`, and `verification` — the one block that is +// COMPUTED (filled by compose when it checks the quote against the cues, the +// voice against the speaker), never typed by hand; the validator flags a block +// whose values could not have come from a check. +// +// EXTENDING THE UNION: a new kind is one more member schema in CITATION_KINDS, +// its *_FIELD_DOCS record (CITATIONS.md is generated from them), its moment +// shape in ./moments.ts when it has a page of its own, and its rules in +// ./validate.ts. A reader of an older version refuses a kind it does not know +// (the union is closed), which is why the containers carry a `version`. +// +// THE SPLIT, and why: zod here checks SHAPE only — types, required keys, +// literal kinds, no unknown keys. Every rule about VALUES (ids, spans, paths, +// dates, URLs, references between citations and sources) is ./validate.ts's, +// which runs after a successful parse and reports EVERY problem with its JSON +// path, instead of the first shape error hiding the rest. +// +// SERVER-ONLY: zod. A client importer takes the types with `import type`; the +// pure helpers (./moments.ts, ./inline.ts) import nothing from here but types. + +import { z } from "zod"; +import type { FieldDocs } from "../fieldDocs"; + +// The version of the citation model a container declares. Bumped when a reader +// of the old version would misread a document of the new one. +export const CITATIONS_VERSION = 1; + +// The longest span a video or audio citation may cut, pad included, in seconds. +// A moment page carries one self-hosted evidence clip of it; this keeps every +// clip a citation (not a re-upload) and well under a static host's file limit. +export const MAX_CITATION_SPAN_SECONDS = 120; + +const text = z.string(); + +export const verificationSchema = z.strictObject({ + quoteScore: z.number().optional(), + quoteCheckedAt: text.optional(), + voiceChecked: z.boolean().optional(), + method: text.optional(), +}); + +export type CitationVerification = z.infer<typeof verificationSchema>; + +export const CITATION_VERIFICATION_FIELD_DOCS: FieldDocs<CitationVerification> = { + quoteScore: + "How closely `quote` matches what the record says at the cited place, from 0 (nothing alike) to 1 (verbatim). Written by the check, with `quoteCheckedAt`.", + quoteCheckedAt: "When the quote was checked: an ISO 8601 date-time. Written with `quoteScore`.", + voiceChecked: + "True when the speaker's voice in the cited span was checked against `speaker`. Only a span (video, audio) has a voice to check.", + method: "What did the checking, e.g. the cue-window comparison and its version. Free text, one line.", +}; + +const common = { + quote: text, + speaker: text.optional(), + date: text.optional(), + label: text.optional(), + note: text.optional(), + verification: verificationSchema.optional(), +}; + +const pad = z.strictObject({ + before: z.number().optional(), + after: z.number().optional(), +}); + +const span = { + channel: text, + id: text, + start: z.number(), + end: z.number(), + pad: pad.optional(), +}; + +export const videoCitationSchema = z.strictObject({ kind: z.literal("video"), ...span, ...common }); +export const audioCitationSchema = z.strictObject({ kind: z.literal("audio"), ...span, ...common }); +export const postCitationSchema = z.strictObject({ + kind: z.literal("post"), + channel: text, + id: text, + thread: z.boolean().optional(), + ...common, +}); +export const sourceCitationSchema = z.strictObject({ + kind: z.literal("source"), + source: text, + image: text.optional(), + ...common, +}); +export const pageCitationSchema = z.strictObject({ + kind: z.literal("page"), + url: text, + title: text.optional(), + archiveUrl: text.optional(), + ...common, +}); + +// Every member, in the order CITATIONS.md documents them. +export const CITATION_KINDS = [ + videoCitationSchema, + audioCitationSchema, + postCitationSchema, + sourceCitationSchema, + pageCitationSchema, +] as const; + +export const citationSchema = z.discriminatedUnion("kind", [...CITATION_KINDS]); + +export type VideoCitation = z.infer<typeof videoCitationSchema>; +export type AudioCitation = z.infer<typeof audioCitationSchema>; +export type PostCitation = z.infer<typeof postCitationSchema>; +export type SourceCitation = z.infer<typeof sourceCitationSchema>; +export type PageCitation = z.infer<typeof pageCitationSchema>; +export type Citation = z.infer<typeof citationSchema>; +export type CitationKind = Citation["kind"]; +// The kinds that cite a span of a record's media: a clip, a start and an end. +export type SpanCitation = VideoCitation | AudioCitation; +export type CitationPad = z.infer<typeof pad>; + +export const SPAN_KINDS: readonly CitationKind[] = ["video", "audio"]; + +export function isSpanCitation(c: Citation): c is SpanCitation { + return c.kind === "video" || c.kind === "audio"; +} + +export type CitationCommon = Pick<VideoCitation, keyof typeof common>; + +export const CITATION_COMMON_FIELD_DOCS: FieldDocs<CitationCommon> = { + quote: + "The cited words, VERBATIM — as the record says them (a span's cues, a post's text, the source's sentence, the page's text). Never a paraphrase: compose checks a span's quote against its cues and fails on drift.", + speaker: "Who says the quote, when that is not the record's own channel (a guest, a co-host, a caller).", + date: "When the quote was said or written: `YYYY`, `YYYY-MM`, `YYYY-MM-DD` or an ISO 8601 date-time. Absent = the record's own date.", + label: "A short display name for the citation (one line), used where its number alone would be too little.", + note: "An editorial note shown with the citation (plain text): context the quote needs.", + verification: + "COMPUTED, not authored: what checking this citation found, written by compose. A hand-typed block that could not have come from a check (a score outside 0–1, a score without its time) is a validation problem.", +}; + +type Own<T> = Omit<T, keyof CitationCommon>; + +export const VIDEO_CITATION_FIELD_DOCS: FieldDocs<Own<VideoCitation>> = { + kind: '`"video"`.', + channel: "The record's channel slug (its directory under `transcripts/channels/`). A path segment of the moment page.", + id: "The record's video id (its directory under `data/`; may start with `-`). A path segment of the moment page.", + start: "Where the cited span starts, in seconds from the start of the record.", + end: "Where the cited span ends, in seconds; after `start`.", + pad: "Context around the span in the evidence clip, in seconds. Absent = none. The span plus its pad is at most 120 s.", +}; + +export const AUDIO_CITATION_FIELD_DOCS: FieldDocs<Own<AudioCitation>> = { + ...VIDEO_CITATION_FIELD_DOCS, + kind: '`"audio"`: a span of a record with no picture worth showing (a podcast). Rendered with a poster.', +}; + +export const CITATION_PAD_FIELD_DOCS: FieldDocs<CitationPad> = { + before: "Seconds of context before `start`; ≥ 0. Absent = 0.", + after: "Seconds of context after `end`; ≥ 0. Absent = 0.", +}; + +export const POST_CITATION_FIELD_DOCS: FieldDocs<Own<PostCitation>> = { + kind: '`"post"`.', + channel: "The post's channel slug. A path segment of the moment page.", + id: "The post's id. A path segment of the moment page.", + thread: "True to show the post with the thread it belongs to. Absent = the post alone.", +}; + +export const SOURCE_CITATION_FIELD_DOCS: FieldDocs<Own<SourceCitation>> = { + kind: '`"source"`: a sentence of a document under review.', + source: "The id of the document in the container's `sources`. Its archive links are shown with the citation.", + image: + "A still of the sentence as the document shows it: a path relative to the container's directory (e.g. `stills/a01.png`), never absolute, never leaving it.", +}; + +export const PAGE_CITATION_FIELD_DOCS: FieldDocs<Own<PageCitation>> = { + kind: '`"page"`: a web page outside the corpus.', + url: "The page's address: an http(s) URL.", + title: "The page's title.", + archiveUrl: "An archived copy of the page (an http(s) URL), shown beside the live link.", +}; + +// ─── Sources: the documents a `source` citation quotes ─── + +export const SOURCE_KINDS = ["article", "page", "video", "post", "document", "other"] as const; + +export const sourceArchiveSchema = z.strictObject({ + label: text, + url: text, + context: text.optional(), +}); + +export const sourceSchema = z.strictObject({ + kind: z.enum(SOURCE_KINDS), + title: text, + url: text.optional(), + publisher: text.optional(), + author: text.optional(), + date: text.optional(), + archives: z.array(sourceArchiveSchema).optional(), + saved: text.optional(), +}); + +export type Source = z.infer<typeof sourceSchema>; +export type SourceArchive = z.infer<typeof sourceArchiveSchema>; + +export const SOURCE_FIELD_DOCS: FieldDocs<Source> = { + kind: `What the document is: ${SOURCE_KINDS.map((k) => `\`${k}\``).join(", ")}.`, + title: "The document's title.", + url: "Where the document lives: an http(s) URL.", + publisher: "Who published it (the outlet, the site).", + author: "Who wrote it.", + date: "When it was published: `YYYY`, `YYYY-MM`, `YYYY-MM-DD` or an ISO 8601 date-time.", + archives: "The document's archive links, in context — as the document had them. Shown with every citation of it.", + saved: + "A saved copy of the document, relative to the container's directory (e.g. `sources/s0/page.html`): the input the stills are shot from. NEVER published.", +}; + +export const SOURCE_ARCHIVE_FIELD_DOCS: FieldDocs<SourceArchive> = { + label: "The link's text.", + url: "The archived copy: an http(s) URL.", + context: "The words around the link in the document, so a reader sees what it was offered as.", +}; + +// ─── A standalone citation set ─── + +export const CITATION_SET_FORMAT = "archilyzer-citations"; + +export const citationSetSchema = z.strictObject({ + format: z.literal(CITATION_SET_FORMAT), + version: z.literal(CITATIONS_VERSION), + sources: z.record(z.string(), sourceSchema).optional(), + citations: z.record(z.string(), citationSchema), +}); + +export type CitationSet = z.infer<typeof citationSetSchema>; + +export const CITATION_SET_FIELD_DOCS: FieldDocs<CitationSet> = { + format: `\`"${CITATION_SET_FORMAT}"\`.`, + version: `\`${CITATIONS_VERSION}\`.`, + sources: "The documents the `source` citations quote, by id. Absent = none.", + citations: "The citations, by id. An id is what a `[label](cite:<id>)` link names.", +}; + +// A reference id: a citation's, a source's, a section's or a claim's. Letters, +// digits and `_ . : -`, starting with a letter or digit, at most 64 — the same +// rule umtool's fact-check claim ids follow. An id is used in a URL fragment +// (`#c-<id>`), never as a path segment. +export const REF_ID_RE = /^[A-Za-z0-9][A-Za-z0-9_.:-]{0,63}$/; + +export function isRefId(v: unknown): v is string { + return typeof v === "string" && REF_ID_RE.test(v); +} diff --git a/common/lib/citations/validate.ts b/common/lib/citations/validate.ts @@ -0,0 +1,295 @@ +// CITATION VALIDATION — every rule about a citation's VALUES, as a list of +// problems with JSON paths. Never throws. +// +// ./schema.ts checks shape (zod); this checks what shape cannot: ids, spans, +// path safety, dates and URLs, references from a citation to its source, and a +// computed `verification` block that could not have come from a check. A +// document validator (lib/report/validate.ts, `validateCitationSet` below) +// runs zod first, maps its issues to the same problem list, and runs these only +// on a document that parsed — so a reader gets every value problem at once. + +import type { z } from "zod"; +import { + MAX_CITATION_SPAN_SECONDS, + citationSetSchema, + isRefId, + isSpanCitation, + type Citation, + type CitationSet, + type CitationVerification, + type Source, +} from "./schema"; +import { formatMomentSeconds, isSafeChannelSegment, isSafeIdSegment, roundMomentSeconds } from "./moments"; + +export type Problem = { + // Where, as a JSON path from the document's root: `citations.c01.end`, + // `sections[0].claims[2].findings`; `""` for the root. + path: string; + message: string; +}; + +export type PathSegment = string | number; + +const IDENT_RE = /^[A-Za-z_$][A-Za-z0-9_$]*$/; + +// A JSON path: `.key` for a plain key, `["k.y"]` for any other, `[i]` for an +// index. +export function jsonPath(segs: readonly PathSegment[]): string { + let out = ""; + for (const s of segs) { + if (typeof s === "number") out += `[${s}]`; + else if (IDENT_RE.test(s)) out += out ? `.${s}` : s; + else out += `[${JSON.stringify(s)}]`; + } + return out; +} + +export function problem(at: readonly PathSegment[], message: string): Problem { + return { path: jsonPath(at), message }; +} + +// zod's issues as problems, under `at`. +export function zodProblems(error: z.ZodError, at: readonly PathSegment[] = []): Problem[] { + return error.issues.map((i) => + problem([...at, ...i.path.map((p) => (typeof p === "number" ? p : String(p)))], i.message), + ); +} + +const blank = (s: string) => !/\S/.test(s); + +// `YYYY`, `YYYY-MM`, `YYYY-MM-DD`, or a date-time with a zone. +const PARTIAL_DATE_RE = /^(\d{4})(?:-(\d{2})(?:-(\d{2}))?)?$/; +const DATE_TIME_RE = /^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}(?::\d{2}(?:\.\d+)?)?(?:Z|[+-]\d{2}:\d{2})$/; + +export function isDateTime(v: string): boolean { + return DATE_TIME_RE.test(v) && Number.isFinite(Date.parse(v)) && isPartialDate(v.slice(0, 10)); +} + +export function isPartialDate(v: string): boolean { + if (DATE_TIME_RE.test(v)) return isDateTime(v); + const m = PARTIAL_DATE_RE.exec(v); + if (!m) return false; + if (m[2] === undefined) return true; + const month = Number(m[2]); + if (month < 1 || month > 12) return false; + if (m[3] === undefined) return true; + const day = Number(m[3]); + const days = new Date(Date.UTC(Number(m[1]), month, 0)).getUTCDate(); + return day >= 1 && day <= days; +} + +export function isHttpUrl(v: string): boolean { + try { + const u = new URL(v); + return (u.protocol === "http:" || u.protocol === "https:") && !!u.hostname; + } catch { + return false; + } +} + +// Why `p` is not a safe path relative to a document's directory, or null: it +// must be relative, `/`-separated, with no empty, `.` or `..` segment — so it +// can neither start elsewhere nor climb out. +export function relativePathProblem(p: string): string | null { + if (blank(p)) return "is empty"; + if (p.includes("\\")) return "must use `/`, not `\\`"; + if (p.startsWith("/") || /^[A-Za-z]:/.test(p)) return "must be relative, not absolute"; + if (/^[A-Za-z][A-Za-z0-9+.-]*:/.test(p)) return "must be a path, not a URL"; + if (p.split("/").some((s) => s === "" || s === "." || s === "..")) { + return "must not have empty, `.` or `..` segments (it may not leave the document's directory)"; + } + if (/[\x00-\x1f]/.test(p)) return "must not contain control characters"; + return null; +} + +function textProblems( + out: Problem[], + at: readonly PathSegment[], + value: string | undefined, + { required = false, oneLine = false }: { required?: boolean; oneLine?: boolean } = {}, +): void { + if (value === undefined) return; + if (blank(value)) { + out.push(problem(at, required ? "must not be blank" : "must not be blank (leave it out instead)")); + return; + } + if (oneLine && /[\r\n]/.test(value)) out.push(problem(at, "must be one line")); +} + +function dateProblems(out: Problem[], at: readonly PathSegment[], value: string | undefined): void { + if (value === undefined) return; + if (!isPartialDate(value)) { + out.push(problem(at, "must be a date: YYYY, YYYY-MM, YYYY-MM-DD or an ISO 8601 date-time with a zone")); + } +} + +function urlProblems(out: Problem[], at: readonly PathSegment[], value: string | undefined): void { + if (value === undefined) return; + if (!isHttpUrl(value)) out.push(problem(at, "must be an http(s) URL")); +} + +function pathProblems(out: Problem[], at: readonly PathSegment[], value: string | undefined): void { + if (value === undefined) return; + const why = relativePathProblem(value); + if (why) out.push(problem(at, why)); +} + +function verificationProblems( + out: Problem[], + at: readonly PathSegment[], + v: CitationVerification, + c: Citation, +): void { + if (v.quoteScore !== undefined && !(v.quoteScore >= 0 && v.quoteScore <= 1)) { + out.push(problem([...at, "quoteScore"], "must be from 0 to 1 — a check writes it; it is not typed by hand")); + } + if (v.quoteCheckedAt !== undefined && !isDateTime(v.quoteCheckedAt)) { + out.push(problem([...at, "quoteCheckedAt"], "must be an ISO 8601 date-time with a zone")); + } + if ((v.quoteScore === undefined) !== (v.quoteCheckedAt === undefined)) { + out.push( + problem( + at, + "quoteScore and quoteCheckedAt are written together by the check — one without the other was not", + ), + ); + } + if (v.voiceChecked !== undefined && !isSpanCitation(c)) { + out.push(problem([...at, "voiceChecked"], `a ${c.kind} citation has no voice to check`)); + } + textProblems(out, [...at, "method"], v.method, { oneLine: true }); +} + +export type CitationContext = { + // The container's sources, for a `source` citation's reference. + sources?: Readonly<Record<string, Source>>; +}; + +// Every problem with one citation, under `at` (its path in the document). +export function citationProblems( + c: Citation, + at: readonly PathSegment[], + ctx: CitationContext = {}, +): Problem[] { + const out: Problem[] = []; + textProblems(out, [...at, "quote"], c.quote, { required: true }); + textProblems(out, [...at, "speaker"], c.speaker, { oneLine: true }); + textProblems(out, [...at, "label"], c.label, { oneLine: true }); + textProblems(out, [...at, "note"], c.note); + dateProblems(out, [...at, "date"], c.date); + if (c.verification) verificationProblems(out, [...at, "verification"], c.verification, c); + + switch (c.kind) { + case "video": + case "audio": { + if (!isSafeChannelSegment(c.channel)) { + out.push(problem([...at, "channel"], "must be a channel slug (one safe path segment)")); + } + if (!isSafeIdSegment(c.id)) { + out.push(problem([...at, "id"], "must be a record id (one safe path segment: letters, digits, `.`, `_`, `-`)")); + } + const before = c.pad?.before ?? 0; + const after = c.pad?.after ?? 0; + if (c.pad?.before !== undefined && c.pad.before < 0) out.push(problem([...at, "pad", "before"], "must be at least 0")); + if (c.pad?.after !== undefined && c.pad.after < 0) out.push(problem([...at, "pad", "after"], "must be at least 0")); + if (c.start < 0) out.push(problem([...at, "start"], "must be at least 0")); + if (!(c.end > c.start)) { + out.push(problem([...at, "end"], `must be after start (${c.start})`)); + } else if (!(roundMomentSeconds(c.end) > roundMomentSeconds(c.start))) { + out.push( + problem( + [...at, "end"], + `rounds to the start (${formatMomentSeconds(c.start)}): a span is at least 0.01 s at the moment key's precision`, + ), + ); + } else { + const total = c.end - c.start + Math.max(0, before) + Math.max(0, after); + if (total > MAX_CITATION_SPAN_SECONDS) { + out.push( + problem( + at, + `the span with its pad is ${Math.round(total * 100) / 100} s (at most ${MAX_CITATION_SPAN_SECONDS} s)`, + ), + ); + } + } + break; + } + case "post": + if (!isSafeChannelSegment(c.channel)) { + out.push(problem([...at, "channel"], "must be a channel slug (one safe path segment)")); + } + if (!isSafeIdSegment(c.id)) { + out.push(problem([...at, "id"], "must be a record id (one safe path segment: letters, digits, `.`, `_`, `-`)")); + } + break; + case "source": + if (!ctx.sources || !Object.hasOwn(ctx.sources, c.source)) { + out.push(problem([...at, "source"], `names no source (${JSON.stringify(c.source)} is not in sources)`)); + } + pathProblems(out, [...at, "image"], c.image); + break; + case "page": + urlProblems(out, [...at, "url"], c.url); + urlProblems(out, [...at, "archiveUrl"], c.archiveUrl); + textProblems(out, [...at, "title"], c.title, { oneLine: true }); + break; + } + return out; +} + +// Every problem with one source, under `at`. +export function sourceProblems(s: Source, at: readonly PathSegment[]): Problem[] { + const out: Problem[] = []; + textProblems(out, [...at, "title"], s.title, { required: true, oneLine: true }); + urlProblems(out, [...at, "url"], s.url); + textProblems(out, [...at, "publisher"], s.publisher, { oneLine: true }); + textProblems(out, [...at, "author"], s.author, { oneLine: true }); + dateProblems(out, [...at, "date"], s.date); + (s.archives ?? []).forEach((a, i) => { + textProblems(out, [...at, "archives", i, "label"], a.label, { required: true, oneLine: true }); + urlProblems(out, [...at, "archives", i, "url"], a.url); + textProblems(out, [...at, "archives", i, "context"], a.context); + }); + pathProblems(out, [...at, "saved"], s.saved); + return out; +} + +// Every problem with a container's `sources` and `citations` maps: each id is a +// reference id, no id names both a source and a citation (a `cite:` link names +// citations only — one id meaning two things is a trap), and every entry's own +// problems. +export function citationMapProblems( + citations: Readonly<Record<string, Citation>>, + sources: Readonly<Record<string, Source>> | undefined, + at: readonly PathSegment[] = [], +): Problem[] { + const out: Problem[] = []; + for (const [id, s] of Object.entries(sources ?? {})) { + const where = [...at, "sources", id]; + if (!isRefId(id)) out.push(problem(where, "is not a reference id (letters, digits, `_ . : -`; at most 64)")); + if (Object.hasOwn(citations, id)) out.push(problem(where, "is also a citation's id — ids name one thing")); + out.push(...sourceProblems(s, where)); + } + for (const [id, c] of Object.entries(citations)) { + const where = [...at, "citations", id]; + if (!isRefId(id)) out.push(problem(where, "is not a reference id (letters, digits, `_ . : -`; at most 64)")); + out.push(...citationProblems(c, where, { sources })); + } + return out; +} + +export type Parsed<T> = { ok: true; value: T; problems: Problem[] } | { ok: false; problems: Problem[] }; + +// A standalone citation set (`archilyzer-citations`): parsed, and every problem. +// `ok` means it parsed; `problems` may still be non-empty. +export function parseCitationSet(raw: unknown): Parsed<CitationSet> { + const r = citationSetSchema.safeParse(raw); + if (!r.success) return { ok: false, problems: zodProblems(r.error) }; + return { ok: true, value: r.data, problems: citationMapProblems(r.data.citations, r.data.sources) }; +} + +// Every problem with a citation set; empty when it is sound. +export function validateCitationSet(raw: unknown): Problem[] { + return parseCitationSet(raw).problems; +} diff --git a/common/lib/report/citedIn.ts b/common/lib/report/citedIn.ts @@ -0,0 +1,50 @@ +// "CITED IN" — the back-link index a moment page reads: for every moment the +// given reports cite, each place that cites it. +// +// moment key → [{ reportId, sectionId, claimId, citationId }] +// +// Built at compose from the site's published reports, in the order given, each +// report in reading order (./uses.ts). A place is listed once however many +// times it cites the moment there (a claim that cites a span in its findings +// and lists it under itself is one back-link); two citations of the same span +// in one claim are two. Citations without a moment (`source`, `page`), and +// references to citations the report does not define, are not indexed. +// +// Pure: the input is parsed reports, the output plain JSON. + +import { momentKeyOf } from "../citations/moments"; +import type { Report } from "./schema"; +import { reportCitationUses } from "./uses"; + +export type CitedIn = { + reportId: string; + // null when cited in the report's summary. + sectionId: string | null; + // null when cited in the summary or a section's body. + claimId: string | null; + citationId: string; +}; + +export function buildCitedIn(reports: readonly Report[]): Record<string, CitedIn[]> { + const out: Record<string, CitedIn[]> = {}; + const seen = new Set<string>(); + for (const report of reports) { + const citations = report.citations ?? {}; + for (const use of reportCitationUses(report)) { + if (!Object.hasOwn(citations, use.citationId)) continue; + const key = momentKeyOf(citations[use.citationId]); + if (!key) continue; + const entry: CitedIn = { + reportId: report.id, + sectionId: use.sectionId, + claimId: use.claimId, + citationId: use.citationId, + }; + const dedup = JSON.stringify([key, entry.reportId, entry.sectionId, entry.claimId, entry.citationId]); + if (seen.has(dedup)) continue; + seen.add(dedup); + (out[key] ??= []).push(entry); + } + } + return out; +} diff --git a/common/lib/report/docs.test.ts b/common/lib/report/docs.test.ts @@ -0,0 +1,46 @@ +import { readFileSync } from "node:fs"; +import path from "node:path"; +import { fileURLToPath } from "node:url"; +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { renderReportMarkdown } from "./docs"; +import { renderCitationsMarkdown } from "../citations/docs"; +import { claimSchema, reportSchema, sectionSchema } from "./schema"; +import { CITATION_KINDS } from "../citations/schema"; +import { VERDICTS } from "./verdicts"; + +// REPORT.md and CITATIONS.md are GENERATED from the schemas +// (common/bin/file-schemas-docs.ts). A hand edit to either, or a schema change +// without a regenerate, fails here. + +const REPO = path.resolve(path.dirname(fileURLToPath(import.meta.url)), "..", "..", ".."); + +for (const [name, render] of [ + ["REPORT.md", renderReportMarkdown], + ["CITATIONS.md", renderCitationsMarkdown], +] as const) { + test(`${name} is what the schema generates`, () => { + const committed = readFileSync(path.join(REPO, name), "utf8"); + assert.equal( + committed, + render(), + `${name} is stale: run pnpm --filter yt-dlp-transcript-common exec tsx bin/file-schemas-docs.ts`, + ); + }); +} + +test("REPORT.md has a row for every key of the document, a section and a claim, and every verdict", () => { + const md = renderReportMarkdown(); + for (const shape of [reportSchema.shape, sectionSchema.shape, claimSchema.shape]) { + for (const key of Object.keys(shape)) assert.ok(md.includes(`| \`${key}\` |`), key); + } + for (const v of VERDICTS) assert.ok(md.includes(`| \`${v}\` |`), v); +}); + +test("CITATIONS.md documents every key of every kind", () => { + const md = renderCitationsMarkdown(); + for (const member of CITATION_KINDS) { + assert.ok(md.includes(`#### \`"${member.shape.kind.value}"\``), member.shape.kind.value); + for (const key of Object.keys(member.shape)) assert.ok(md.includes(`| \`${key}\` |`), key); + } +}); diff --git a/common/lib/report/docs.ts b/common/lib/report/docs.ts @@ -0,0 +1,78 @@ +// REPORT.md — `report.json`'s key tables, GENERATED from the schema +// (./schema.ts) and the verdict vocabulary (./verdicts.mjs). The citation keys +// are CITATIONS.md's (lib/citations/docs.ts), linked rather than repeated. +// +// A pure renderer: common/bin/file-schemas-docs.ts writes the file and +// report/docs.test.ts asserts the committed bytes are what this returns. + +import { cell } from "../settingsDocs"; +import { GENERATED_SCHEMA_DOC, REGENERATE_SCHEMA_DOC, renderKeyTable } from "../citations/docs"; +import { + CLAIM_FIELD_DOCS, + REPORT_FIELD_DOCS, + REPORT_FORMAT, + REPORT_VERSION, + SECTION_FIELD_DOCS, + claimSchema, + reportSchema, + sectionSchema, +} from "./schema"; +import { VERDICT_DEFAULTS, VERDICT_LABEL_MAX, VERDICTS } from "./verdicts"; + +export function renderReportMarkdown(): string { + const out: string[] = []; + out.push("# report.json keys"); + out.push(""); + out.push(GENERATED_SCHEMA_DOC); + out.push(""); + out.push( + `One cited report, format \`"${REPORT_FORMAT}"\`, version ${REPORT_VERSION}, ` + + "persisted to `transcripts/sites/<siteId>/reports/<reportId>/report.json` beside " + + "its `stills/` and `sources/<sourceId>/`; a relative path in it is relative to " + + "that directory. A site's `reports` list in `site.json` is the published, " + + "ordered list — see [SITE.md](SITE.md); a report directory it does not name is a " + + "draft. The schema is `common/lib/report/schema.ts`; its citations and sources " + + "are the citation model's — see [CITATIONS.md](CITATIONS.md).", + ); + out.push(""); + out.push( + 'A **fact-check** (`"kind": "factcheck"`) is sections (chapters) of claims, each ' + + 'with a verdict and its findings. A **sweep** (`"kind": "sweep"`) is sections ' + + "with no verdicts, or bodies that cite inline. Markdown fields (`summary`, a " + + "section's `body`, a claim's `findings`) cite with `[label](cite:<id>)`.", + ); + out.push(""); + out.push( + "`common/lib/report/validate.ts` reports every problem with its JSON path: an " + + "unknown key, a reference that names nothing (a listed citation, a claim's " + + "source sentence, a `cite:` link, the subject), a section or claim id used " + + "twice (they share one namespace: the report page's anchors), a sweep's claim " + + "with a verdict, `updated` before `published`, and every citation problem " + + "CITATIONS.md lists. Whether a still exists and whether a quote matches its cues " + + "are checked when the site is composed.", + ); + out.push(""); + out.push(REGENERATE_SCHEMA_DOC); + out.push(""); + renderKeyTable(out, "## The document", reportSchema.shape, REPORT_FIELD_DOCS); + renderKeyTable(out, "#### `sections[]`", sectionSchema.shape, SECTION_FIELD_DOCS); + renderKeyTable(out, "#### `sections[].claims[]`", claimSchema.shape, CLAIM_FIELD_DOCS); + + out.push("## The verdicts"); + out.push(""); + out.push( + "The shared vocabulary (`common/lib/report/verdicts.mjs` — the one copy; " + + "report-to-video's fact-check stamps read it too), in the order a tally lists " + + `them. A report's \`verdicts\` overrides a label (one line, at most ${VERDICT_LABEL_MAX} ` + + "characters) or a colour (`#rgb` or `#rrggbb`), each on its own.", + ); + out.push(""); + out.push("| Verdict | Default label | Default colour |"); + out.push("|---|---|---|"); + for (const v of VERDICTS) { + const d = VERDICT_DEFAULTS[v]; + out.push(`| \`${v}\` | ${cell(d.label)} | \`${d.color}\` |`); + } + out.push(""); + return out.join("\n"); +} diff --git a/common/lib/report/report.test.ts b/common/lib/report/report.test.ts @@ -0,0 +1,241 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { parseReport, validateReport } from "./validate"; +import { reportCitationNumbers, reportCitationUses } from "./uses"; +import { buildCitedIn } from "./citedIn"; +import { VERDICTS, VERDICT_DEFAULTS, isVerdict, resolveVerdicts } from "./verdicts"; +import type { Report } from "./schema"; +import type { Problem } from "../citations/validate"; + +const paths = (ps: Problem[]) => ps.map((p) => p.path); + +function fixture(): Record<string, unknown> & Report { + return { + format: "archilyzer-report", + version: 1, + id: "demo-report", + kind: "factcheck", + title: "A demo fact-check", + subtitle: "What the article says, against the record", + summary: "The article gets one thing [right](cite:c02) and one [wrong](cite:c01).", + published: "2026-10-01", + updated: "2026-10-04T09:30:00Z", + subject: { source: "s0" }, + verdicts: { PARTLY: { label: "Half true" } }, + sources: { + s0: { + kind: "article", + title: "A demo article", + url: "https://example.test/article", + archives: [{ label: "archived", url: "https://archive.example.test/a" }], + saved: "sources/s0/page.html", + }, + }, + citations: { + c01: { kind: "video", channel: "demo-channel", id: "-abc123", start: 10, end: 20, quote: "q1", pad: { before: 5, after: 5 } }, + c02: { kind: "video", channel: "demo-channel", id: "def456", start: 30, end: 42.5, quote: "q2" }, + p01: { kind: "post", channel: "demo-channel", id: "1234567890", quote: "q3" }, + a01: { kind: "source", source: "s0", quote: "the claim", image: "stills/a01.png" }, + a02: { kind: "source", source: "s0", quote: "another claim" }, + w01: { kind: "page", url: "https://example.test/page", quote: "q4" }, + }, + sections: [ + { + id: "ch1", + title: "Chapter one", + body: "Background, see [this](cite:w01).", + claims: [ + { + id: "k1", + text: "The first claim.", + verdict: "CONTRADICTED", + sourceQuote: { citation: "a01" }, + findings: "The record says otherwise: [at 0:10](cite:c01), and [the post](cite:p01).", + citations: ["c01", "p01"], + }, + { + id: "k2", + text: "The second claim.", + verdict: "PARTLY", + sourceQuote: { citation: "a02" }, + findings: "Partly: [here](cite:c02) and [here again](cite:c01).", + citations: ["c02"], + }, + ], + }, + { id: "ch2", title: "Chapter two", claims: [{ id: "k3", text: "Not yet ruled." }] }, + ], + }; +} + +test("a sound fact-check parses with no problems", () => { + const r = parseReport(fixture(), { id: "demo-report" }); + assert.ok(r.ok); + assert.deepEqual(r.problems, []); +}); + +test("schema rejects: wrong format/version, unknown keys, an unknown verdict, a missing title", () => { + const f = fixture() as Record<string, unknown>; + const bad = { ...f, format: "other", version: 2, extra: true, title: undefined, verdicts: { MAYBE: {} } }; + const r = parseReport(bad); + assert.equal(r.ok, false); + const ps = paths(r.problems); + for (const want of ["format", "version", "title", "verdicts.MAYBE"]) assert.ok(ps.includes(want), `${want} in ${ps}`); + assert.ok(ps.includes(""), `the unknown key at the root in ${ps}`); + const claim = structuredClone(fixture()); + (claim.sections[0].claims![0] as Record<string, unknown>).verdict = "MOSTLY"; + assert.deepEqual(paths(validateReport(claim)), ["sections[0].claims[0].verdict"]); +}); + +test("every reference must resolve: listed citations, source sentences, cite links, the subject", () => { + const f = fixture(); + f.subject = { source: "s9" }; + const k1 = f.sections[0].claims![0]; + k1.citations = ["c01", "c99", "c01"]; + k1.sourceQuote = { citation: "c01" }; // not a source citation + k1.findings = "See [gone](cite:c98) and [blank](cite:)."; + f.summary = "Also [missing](cite:zz)."; + const ps = validateReport(f); + assert.deepEqual(paths(ps).sort(), [ + "sections[0].claims[0].citations[1]", + "sections[0].claims[0].citations[2]", + "sections[0].claims[0].findings", + "sections[0].claims[0].findings", + "sections[0].claims[0].sourceQuote.citation", + "subject.source", + "summary", + ]); + assert.ok(ps.some((p) => /\[gone\]\(cite:c98\) names no citation/.test(p.message))); + assert.ok(ps.some((p) => /lists "c01" twice/.test(p.message))); + assert.ok(ps.some((p) => /names a video citation/.test(p.message))); +}); + +test("citation problems surface under the report's paths", () => { + const f = fixture(); + f.citations!.c01 = { ...(f.citations!.c01 as object), end: 200 } as never; + f.citations!.a01 = { ...(f.citations!.a01 as object), image: "../../etc/passwd" } as never; + assert.deepEqual(paths(validateReport(f)).sort(), ["citations.a01.image", "citations.c01"]); +}); + +test("ids: report id is a slug and matches its directory; section and claim ids unique together", () => { + const f = fixture(); + f.sections[1].id = "k1"; + f.sections[1].claims![0].id = "bad id"; + const ps = validateReport(f, { id: "other-report" }); + assert.deepEqual(paths(ps).sort(), ["id", "sections[1].claims[0].id", "sections[1].id"]); + assert.ok(ps.some((p) => /already the id of sections\[0\]\.claims\[0\]/.test(p.message))); + assert.deepEqual(paths(validateReport({ ...fixture(), id: "Demo_Report" })), ["id"]); +}); + +test("dates, verdict overrides, a sweep with a verdict", () => { + const f = fixture(); + f.published = "2026-10-05"; + f.updated = "2026-10-04"; + f.verdicts = { PARTLY: { label: "x".repeat(25), color: "orange" } }; + const ps = validateReport(f); + assert.deepEqual(paths(ps).sort(), ["updated", "verdicts.PARTLY.color", "verdicts.PARTLY.label"]); + assert.deepEqual(paths(validateReport({ ...fixture(), published: "1 Oct 2026" })), ["published"]); + assert.deepEqual(paths(validateReport({ ...fixture(), kind: "sweep" })), [ + "sections[0].claims[0].verdict", + "sections[0].claims[1].verdict", + ]); +}); + +test("a sweep with inline citations in its bodies and no claims is a report", () => { + const sweep = { + format: "archilyzer-report", + version: 1, + id: "demo-sweep", + kind: "sweep", + title: "A demo sweep", + citations: { c01: { kind: "video", channel: "demo-channel", id: "abc123", start: 1, end: 2, quote: "q" } }, + sections: [{ id: "s1", title: "Found", body: "It came up [once](cite:c01)." }], + }; + assert.deepEqual(validateReport(sweep), []); +}); + +test("uses: reading order — summary, body, then each claim's source sentence, findings, list", () => { + const r = parseReport(fixture()); + assert.ok(r.ok); + const uses = reportCitationUses(r.value).map((u) => `${u.field}:${u.citationId}`); + assert.deepEqual(uses, [ + "summary:c02", + "summary:c01", + "body:w01", + "sourceQuote:a01", + "findings:c01", + "findings:p01", + "citations:c01", + "citations:p01", + "sourceQuote:a02", + "findings:c02", + "findings:c01", + "citations:c02", + ]); + assert.deepEqual([...reportCitationNumbers(r.value)], [ + ["c02", 1], + ["c01", 2], + ["w01", 3], + ["a01", 4], + ["p01", 5], + ["a02", 6], + ]); +}); + +test("numbers skip dangling references", () => { + const f = fixture(); + f.summary = "[x](cite:nope) then [y](cite:c01)"; + const r = parseReport(f); + assert.ok(r.ok); + assert.equal(reportCitationNumbers(r.value).get("c01"), 1); + assert.equal(reportCitationNumbers(r.value).has("nope"), false); +}); + +test("cited-in: moment key → each place that cites it, across reports, deduplicated", () => { + const a = parseReport(fixture()); + const second = fixture(); + second.id = "second-report"; + second.summary = undefined; + second.sections = [{ id: "only", title: "Only", body: "Again [here](cite:c01)." }]; + const b = parseReport(second); + assert.ok(a.ok && b.ok); + const index = buildCitedIn([a.value, b.value]); + assert.deepEqual(Object.keys(index), [ + "demo-channel/def456/30.00-42.50", + "demo-channel/-abc123/10.00-20.00", + "demo-channel/1234567890", + ]); + assert.deepEqual(index["demo-channel/-abc123/10.00-20.00"], [ + { reportId: "demo-report", sectionId: null, claimId: null, citationId: "c01" }, + { reportId: "demo-report", sectionId: "ch1", claimId: "k1", citationId: "c01" }, + { reportId: "demo-report", sectionId: "ch1", claimId: "k2", citationId: "c01" }, + { reportId: "second-report", sectionId: "only", claimId: null, citationId: "c01" }, + ]); + assert.deepEqual(index["demo-channel/1234567890"], [ + { reportId: "demo-report", sectionId: "ch1", claimId: "k1", citationId: "p01" }, + ]); +}); + +test("cited-in: two citations of one span in one place are two back-links; source and page citations have none", () => { + const f = fixture(); + f.citations!.c03 = { ...(f.citations!.c01 as object), quote: "the same span, quoted again" } as never; + f.sections[0].claims![0].citations = ["c01", "c03"]; + const r = parseReport(f); + assert.ok(r.ok); + const entries = buildCitedIn([r.value])["demo-channel/-abc123/10.00-20.00"]; + assert.deepEqual( + entries.filter((e) => e.claimId === "k1").map((e) => e.citationId), + ["c01", "c03"], + ); + const keys = Object.keys(buildCitedIn([r.value])); + assert.ok(!keys.some((k) => k.includes("example.test"))); +}); + +test("the verdict vocabulary: five verdicts, a default for each, overrides laid over one key at a time", () => { + assert.deepEqual([...VERDICTS], ["CORROBORATED", "PARTLY", "CONTRADICTED", "NOT_FOUND", "UNTESTABLE"]); + assert.deepEqual(Object.keys(VERDICT_DEFAULTS), [...VERDICTS]); + assert.ok(isVerdict("PARTLY") && !isVerdict("partly")); + const v = resolveVerdicts({ PARTLY: { label: "Half true" }, MAYBE: { label: "x" } }); + assert.deepEqual(v.PARTLY, { label: "Half true", color: VERDICT_DEFAULTS.PARTLY.color }); + assert.deepEqual(Object.keys(v), [...VERDICTS]); +}); diff --git a/common/lib/report/schema.ts b/common/lib/report/schema.ts @@ -0,0 +1,123 @@ +// THE REPORT DOCUMENT — one definition of `report.json` (format +// `archilyzer-report`), used by the validator, the compose stage, the export +// site's report pages and REPORT.md. +// +// A report is a cited document: a summary, then sections, each with an +// optional markdown body and claims. A FACT-CHECK (`kind: "factcheck"`) is +// sections (chapters) of claims, each with a verdict from the shared +// vocabulary (./verdicts.mjs) and its findings; a SWEEP (`kind: "sweep"`) is +// sections with no verdicts, or bodies with inline citations. +// +// ITS CITATIONS ARE THE CITATION MODEL'S (lib/citations/): the `sources` and +// `citations` maps are that model's schemas, cited inline with +// `[label](cite:<id>)` and listed per claim. Nothing about a citation is +// defined here. +// +// Lives at `transcripts/sites/<siteId>/reports/<reportId>/report.json`, beside +// its `stills/` and `sources/<sourceId>/`; a relative path in it (a still, a +// saved source) is relative to that directory. +// +// As in lib/citations/schema.ts, zod checks SHAPE only; every rule about +// values — ids, references, `cite:` links, the verdict overrides, dates — is +// ./validate.ts's, which reports every problem with its JSON path. +// +// SERVER-ONLY: zod. A client importer takes the types with `import type`. + +import { z } from "zod"; +import type { FieldDocs } from "../fieldDocs"; +import { citationSchema, sourceSchema } from "../citations/schema"; +import { VERDICTS, type Verdict } from "./verdicts"; + +export const REPORT_FORMAT = "archilyzer-report"; +export const REPORT_VERSION = 1; + +export const REPORT_KINDS = ["factcheck", "sweep"] as const; +export type ReportKind = (typeof REPORT_KINDS)[number]; + +const text = z.string(); +const verdict = z.enum(VERDICTS as unknown as [Verdict, ...Verdict[]]); + +export const claimSchema = z.strictObject({ + id: text, + text, + verdict: verdict.optional(), + sourceQuote: z.strictObject({ citation: text }).optional(), + findings: text.optional(), + citations: z.array(text).optional(), +}); + +export const sectionSchema = z.strictObject({ + id: text, + title: text, + body: text.optional(), + claims: z.array(claimSchema).optional(), +}); + +const verdictOverride = z.strictObject({ label: text.optional(), color: text.optional() }); + +export const reportSchema = z.strictObject({ + format: z.literal(REPORT_FORMAT), + version: z.literal(REPORT_VERSION), + id: text, + kind: z.enum(REPORT_KINDS), + title: text, + subtitle: text.optional(), + summary: text.optional(), + published: text.optional(), + updated: text.optional(), + subject: z.strictObject({ source: text }).optional(), + verdicts: z.partialRecord(verdict, verdictOverride).optional(), + sources: z.record(text, sourceSchema).optional(), + citations: z.record(text, citationSchema).optional(), + sections: z.array(sectionSchema), +}); + +export type Claim = z.infer<typeof claimSchema>; +export type Section = z.infer<typeof sectionSchema>; +export type Report = z.infer<typeof reportSchema>; +export type ReportSubject = NonNullable<Report["subject"]>; +export type ClaimSourceQuote = NonNullable<Claim["sourceQuote"]>; + +// A report id is its directory name under `reports/` and a path segment of its +// page (`/reports/<id>/`): a lowercase slug, like a site id. +export const REPORT_ID_RE = /^[a-z0-9][a-z0-9-]{0,63}$/; + +export function isReportId(v: unknown): v is string { + return typeof v === "string" && REPORT_ID_RE.test(v); +} + +export const REPORT_FIELD_DOCS: FieldDocs<Report> = { + format: `\`"${REPORT_FORMAT}"\`.`, + version: `\`${REPORT_VERSION}\`.`, + id: "The report's id: a lowercase slug (`[a-z0-9][a-z0-9-]*`, at most 64), its directory name under `reports/` and the last segment of its page, `/reports/<id>/`. Must match the directory.", + kind: '`"factcheck"` — sections of claims, each with a verdict — or `"sweep"` — sections with no verdicts, or bodies with inline citations.', + title: "The report's title.", + subtitle: "A line under the title.", + summary: "The report's summary, in markdown, shown before the sections. May cite inline: `[label](cite:<id>)`.", + published: "When the report was published: `YYYY-MM-DD` or an ISO 8601 date-time with a zone.", + updated: "When it was last changed, in the same form; not before `published`.", + subject: + "The document under review, when the report reviews one: `{ \"source\": \"<id>\" }`, an id in `sources`.", + verdicts: `Overrides of the shared verdict vocabulary's labels and colours, by verdict (${VERDICTS.map((v) => `\`${v}\``).join(", ")}): \`{ "label": "…", "color": "#rrggbb" }\`, each key optional. Absent = the shared defaults.`, + sources: "The documents the report's `source` citations quote, by id — see [CITATIONS.md](CITATIONS.md). Absent = none.", + citations: + "The report's citations, by id — see [CITATIONS.md](CITATIONS.md). A citation is cited from markdown with `[label](cite:<id>)` and listed under the claims that rest on it. Absent = none.", + sections: "The report's sections, in order.", +}; + +export const SECTION_FIELD_DOCS: FieldDocs<Section> = { + id: "The section's id (letters, digits, `_ . : -`; at most 64): its anchor on the report page. Unique among the report's section and claim ids.", + title: "The section's heading.", + body: "Markdown under the heading. May cite inline.", + claims: "The section's claims, in order. Absent = none.", +}; + +export const CLAIM_FIELD_DOCS: FieldDocs<Claim> = { + id: "The claim's id (letters, digits, `_ . : -`; at most 64): its anchor on the report page. Unique among the report's section and claim ids.", + text: "The claim, as stated by the document under review (plain text).", + verdict: `The ruling on the claim: ${VERDICTS.map((v) => `\`${v}\``).join(", ")}. A fact-check's claim may leave it out (not yet ruled); a sweep's carries none.`, + sourceQuote: + "The document's own sentence making the claim: `{ \"citation\": \"<id>\" }`, naming a `source` citation (its still is shown with the claim).", + findings: "What the evidence shows, in markdown, citing inline: `[label](cite:<id>)`.", + citations: "The citations the claim rests on, in the order they are listed under it. Each must exist; none twice.", +}; diff --git a/common/lib/report/uses.ts b/common/lib/report/uses.ts @@ -0,0 +1,85 @@ +// EVERY PLACE A REPORT CITES, in reading order — the one walk the validator, +// the numbering and the back-link index share, so they cannot disagree about +// what a report cites or in what order. +// +// Reading order: the summary; then each section's body, and each of its +// claims in turn — the claim's source sentence, its findings, then the +// citations listed under it. Within a markdown field, `cite:` links in the +// order they are written (lib/citations/inline.ts). +// +// Pure, no imports but types and the inline parser: the export site can use it. + +import { extractCiteRefs, numberCitations } from "../citations/inline"; +import type { PathSegment } from "../citations/validate"; +import type { Report } from "./schema"; + +export type CitationUseField = "summary" | "body" | "sourceQuote" | "findings" | "citations"; + +export type CitationUse = { + citationId: string; + // The section and claim the use is in; null in the summary (both) or a + // section's body (the claim). + sectionId: string | null; + claimId: string | null; + field: CitationUseField; + // The JSON path of the field (with the list index for `citations`). + path: PathSegment[]; + // An inline link's label; absent for a listed citation. + label?: string; +}; + +export function reportCitationUses(report: Report): CitationUse[] { + const out: CitationUse[] = []; + const inline = ( + md: string | undefined, + field: CitationUseField, + path: PathSegment[], + sectionId: string | null, + claimId: string | null, + ) => { + for (const ref of extractCiteRefs(md)) { + out.push({ citationId: ref.id, sectionId, claimId, field, path, label: ref.label }); + } + }; + inline(report.summary, "summary", ["summary"], null, null); + report.sections.forEach((section, si) => { + const sp: PathSegment[] = ["sections", si]; + inline(section.body, "body", [...sp, "body"], section.id, null); + (section.claims ?? []).forEach((claim, ci) => { + const cp: PathSegment[] = [...sp, "claims", ci]; + if (claim.sourceQuote) { + out.push({ + citationId: claim.sourceQuote.citation, + sectionId: section.id, + claimId: claim.id, + field: "sourceQuote", + path: [...cp, "sourceQuote", "citation"], + }); + } + inline(claim.findings, "findings", [...cp, "findings"], section.id, claim.id); + (claim.citations ?? []).forEach((id, i) => { + out.push({ + citationId: id, + sectionId: section.id, + claimId: claim.id, + field: "citations", + path: [...cp, "citations", i], + }); + }); + }); + }); + return out; +} + +// Each citation's number in the report, from 1, by first appearance in reading +// order. Only citations the report defines are numbered (a dangling reference +// is a validation problem, not a number); a citation the report defines but +// never cites has none. +export function reportCitationNumbers(report: Report): Map<string, number> { + const defined = report.citations ?? {}; + return numberCitations( + reportCitationUses(report) + .map((u) => u.citationId) + .filter((id) => Object.hasOwn(defined, id)), + ); +} diff --git a/common/lib/report/validate.ts b/common/lib/report/validate.ts @@ -0,0 +1,158 @@ +// REPORT VALIDATION — every problem with a `report.json`, as a list with JSON +// paths. Never throws. +// +// Shape first (./schema.ts, zod): a report that does not parse gets zod's +// problems and nothing else, since the rest cannot be walked. A report that +// parses gets every value problem at once: +// +// - its id is a report id (and, when asked, its directory's name); +// - the dates are dates, `updated` not before `published`; +// - the verdict overrides follow the shared rule (./verdicts.mjs); +// - section and claim ids are reference ids, unique together (they are the +// report page's anchors); +// - a sweep's claims carry no verdict; +// - every citation and source is sound (lib/citations/validate.ts: ids, +// spans, safe paths, URLs, verification); +// - every reference resolves: `subject.source`, each claim's `sourceQuote` +// (to a `source` citation), each listed citation (none twice in one +// claim), and every `[label](cite:<id>)` link in the summary, the bodies +// and the findings. +// +// What this cannot check is the disk and the corpus — that a still exists, that +// a quote matches its cues; compose does those, where both are at hand. + +import { + citationMapProblems, + isDateTime, + jsonPath, + problem, + zodProblems, + type Parsed, + type PathSegment, + type Problem, +} from "../citations/validate"; +import { isRefId } from "../citations/schema"; +import { reportSchema, isReportId, type Report } from "./schema"; +import { reportCitationUses } from "./uses"; +import { verdictOverrideProblems } from "./verdicts"; + +const DATE_RE = /^\d{4}-\d{2}-\d{2}$/; + +function isReportDate(v: string): boolean { + if (DATE_RE.test(v)) return isDateTime(`${v}T00:00Z`); + return isDateTime(v); +} + +const blank = (s: string) => !/\S/.test(s); + +export type ReportValidateOptions = { + // The report's directory name: its `id` must be this. + id?: string; +}; + +function reportProblems(report: Report, opts: ReportValidateOptions): Problem[] { + const out: Problem[] = []; + const citations = report.citations ?? {}; + const sources = report.sources ?? {}; + + if (!isReportId(report.id)) { + out.push(problem(["id"], "must be a lowercase slug (`[a-z0-9][a-z0-9-]*`, at most 64)")); + } else if (opts.id !== undefined && opts.id !== report.id) { + out.push(problem(["id"], `is ${JSON.stringify(report.id)} but the report's directory is ${JSON.stringify(opts.id)}`)); + } + if (blank(report.title)) out.push(problem(["title"], "must not be blank")); + else if (/[\r\n]/.test(report.title)) out.push(problem(["title"], "must be one line")); + if (report.subtitle !== undefined && /[\r\n]/.test(report.subtitle)) { + out.push(problem(["subtitle"], "must be one line")); + } + for (const key of ["published", "updated"] as const) { + const v = report[key]; + if (v !== undefined && !isReportDate(v)) { + out.push(problem([key], "must be YYYY-MM-DD or an ISO 8601 date-time with a zone")); + } + } + if ( + report.published !== undefined && + report.updated !== undefined && + isReportDate(report.published) && + isReportDate(report.updated) + ) { + const p = Date.parse(DATE_RE.test(report.published) ? `${report.published}T00:00Z` : report.published); + const u = Date.parse(DATE_RE.test(report.updated) ? `${report.updated}T00:00Z` : report.updated); + if (u < p) out.push(problem(["updated"], "is before published")); + } + + for (const [v, override] of Object.entries(report.verdicts ?? {})) { + for (const message of verdictOverrideProblems(override, "")) { + // The shared rule names its own sub-path (`.label`); split it back off. + const m = /^\.(\w+) (.*)$/s.exec(message); + out.push(m ? problem(["verdicts", v, m[1]], m[2]) : problem(["verdicts", v], message.trim())); + } + } + + if (report.subject && !Object.hasOwn(sources, report.subject.source)) { + out.push(problem(["subject", "source"], `names no source (${JSON.stringify(report.subject.source)} is not in sources)`)); + } + + out.push(...citationMapProblems(citations, report.sources)); + + const anchors = new Map<string, string>(); + const anchor = (id: string, at: PathSegment[]) => { + if (!isRefId(id)) { + out.push(problem([...at, "id"], "is not a reference id (letters, digits, `_ . : -`; at most 64)")); + return; + } + const first = anchors.get(id); + if (first) out.push(problem([...at, "id"], `${JSON.stringify(id)} is already the id of ${first}`)); + else anchors.set(id, jsonPath(at)); + }; + report.sections.forEach((section, si) => { + const sp: PathSegment[] = ["sections", si]; + anchor(section.id, sp); + if (blank(section.title)) out.push(problem([...sp, "title"], "must not be blank")); + (section.claims ?? []).forEach((claim, ci) => { + const cp: PathSegment[] = [...sp, "claims", ci]; + anchor(claim.id, cp); + if (blank(claim.text)) out.push(problem([...cp, "text"], "must not be blank")); + if (report.kind === "sweep" && claim.verdict !== undefined) { + out.push(problem([...cp, "verdict"], "a sweep's claims carry no verdict (make the report a factcheck)")); + } + const listed = new Set<string>(); + (claim.citations ?? []).forEach((id, i) => { + if (listed.has(id)) out.push(problem([...cp, "citations", i], `lists ${JSON.stringify(id)} twice`)); + listed.add(id); + }); + }); + }); + + for (const use of reportCitationUses(report)) { + const id = use.citationId; + if (use.field === "citations" || use.field === "sourceQuote") { + if (!Object.hasOwn(citations, id)) { + out.push(problem(use.path, `names no citation (${JSON.stringify(id)} is not in citations)`)); + } else if (use.field === "sourceQuote" && citations[id].kind !== "source") { + out.push(problem(use.path, `names a ${citations[id].kind} citation; a claim's source sentence is a \`source\` citation`)); + } + continue; + } + if (id === "") out.push(problem(use.path, `the link [${use.label}](cite:) names no citation`)); + else if (!Object.hasOwn(citations, id)) { + out.push(problem(use.path, `the link [${use.label}](cite:${id}) names no citation (${JSON.stringify(id)} is not in citations)`)); + } + } + return out; +} + +// A report: parsed, and every problem. `ok` means it parsed (its shape is a +// report); `problems` may still be non-empty, and a report with problems must +// not be published. +export function parseReport(raw: unknown, opts: ReportValidateOptions = {}): Parsed<Report> { + const r = reportSchema.safeParse(raw); + if (!r.success) return { ok: false, problems: zodProblems(r.error) }; + return { ok: true, value: r.data, problems: reportProblems(r.data, opts) }; +} + +// Every problem with a report; empty when it is sound. +export function validateReport(raw: unknown, opts: ReportValidateOptions = {}): Problem[] { + return parseReport(raw, opts).problems; +} diff --git a/common/lib/report/verdicts.mjs b/common/lib/report/verdicts.mjs @@ -0,0 +1,81 @@ +// THE VERDICT VOCABULARY — THE one copy: the rulings a fact-check claim may +// carry, their default labels and colours, and the limits on an override. +// +// Plain JS (with JSDoc types) for the reason ytdlp/platformArgs.mjs is: umtool's +// report-to-video scripts are `.mjs` run by bare `node` and cannot import +// TypeScript, and factcheck.mjs (the video's stamps and tally) takes its +// defaults from here. `./verdicts.ts` re-exports it with types for every TS +// caller (report.json's schema and validator, the export site's report pages). +// Never copy these values into another file. + +/** The verdicts a claim may carry, in the order a tally lists them. */ +export const VERDICTS = Object.freeze(["CORROBORATED", "PARTLY", "CONTRADICTED", "NOT_FOUND", "UNTESTABLE"]); + +/** Each verdict's default label and colour. An override names only what it changes. */ +export const VERDICT_DEFAULTS = Object.freeze({ + CORROBORATED: Object.freeze({ label: "Corroborated", color: "#3fbf7f" }), + PARTLY: Object.freeze({ label: "Partly true", color: "#e3b23c" }), + CONTRADICTED: Object.freeze({ label: "Contradicted", color: "#e5534b" }), + NOT_FOUND: Object.freeze({ label: "Not found", color: "#8b93a7" }), + UNTESTABLE: Object.freeze({ label: "Untestable", color: "#7d8fd6" }), +}); + +/** The longest label an override may set, in characters (trimmed). */ +export const VERDICT_LABEL_MAX = 24; + +/** An override's colour: `#rgb` or `#rrggbb`. */ +export const VERDICT_COLOR_RE = /^#(?:[0-9a-fA-F]{3}|[0-9a-fA-F]{6})$/; + +/** @param {unknown} v */ +export function isVerdict(v) { + return typeof v === "string" && VERDICTS.includes(v); +} + +const isObj = (v) => v !== null && typeof v === "object" && !Array.isArray(v); + +/** + * Every verdict's label and colour with the overrides laid over the defaults, + * one verdict at a time and one key at a time. A non-object override, or one + * for a verdict that is not in the vocabulary, is ignored: validation is the + * caller's (factcheck.mjs `validateFactcheck`, report/validate.ts). + * + * @param {unknown} overrides + * @returns {Record<string, { label: string, color: string }>} + */ +export function resolveVerdicts(overrides) { + const given = isObj(overrides) ? overrides : {}; + /** @type {Record<string, { label: string, color: string }>} */ + const out = {}; + for (const v of VERDICTS) { + out[v] = { ...VERDICT_DEFAULTS[v], ...(isObj(given[v]) ? given[v] : {}) }; + } + return out; +} + +/** + * Every reason a `{ label?, color? }` override cannot be used, as sentences + * naming `where` (empty: it can). Unknown keys are refused. The one rule for + * both a report's `verdicts` and a video manifest's `render.chrome.factcheck.verdicts`. + * + * @param {unknown} v + * @param {string} where + * @returns {string[]} + */ +export function verdictOverrideProblems(v, where) { + if (!isObj(v)) return [`${where} must be { label, color }`]; + const errors = []; + for (const k of Object.keys(v)) { + if (k !== "label" && k !== "color") errors.push(`${where}.${k} is not a verdict setting (label, color)`); + } + if (v.label !== undefined) { + if (typeof v.label !== "string" || !v.label.trim()) errors.push(`${where}.label must be words, not empty`); + else if (/[\r\n]/.test(v.label)) errors.push(`${where}.label must be one line`); + else if (v.label.trim().length > VERDICT_LABEL_MAX) { + errors.push(`${where}.label is ${v.label.trim().length} characters (at most ${VERDICT_LABEL_MAX})`); + } + } + if (v.color !== undefined && !(typeof v.color === "string" && VERDICT_COLOR_RE.test(v.color))) { + errors.push(`${where}.color must be a hex colour, #rgb or #rrggbb`); + } + return errors; +} diff --git a/common/lib/report/verdicts.ts b/common/lib/report/verdicts.ts @@ -0,0 +1,40 @@ +// The verdict vocabulary, typed. The values live in ./verdicts.mjs — plain JS so +// umtool's report-to-video scripts can import the same copy — and this module +// is how every TS caller reads them: with the verdict as a union, not a string. + +import { + VERDICT_COLOR_RE as VERDICT_COLOR_RE_JS, + VERDICT_DEFAULTS as VERDICT_DEFAULTS_JS, + VERDICT_LABEL_MAX as VERDICT_LABEL_MAX_JS, + VERDICTS as VERDICTS_JS, + resolveVerdicts as resolveVerdictsJs, + verdictOverrideProblems as verdictOverrideProblemsJs, +} from "./verdicts.mjs"; + +export type Verdict = "CORROBORATED" | "PARTLY" | "CONTRADICTED" | "NOT_FOUND" | "UNTESTABLE"; + +export type VerdictStyle = { label: string; color: string }; + +export type VerdictOverride = { label?: string; color?: string }; + +export const VERDICTS = VERDICTS_JS as readonly Verdict[]; + +export const VERDICT_DEFAULTS = VERDICT_DEFAULTS_JS as Readonly<Record<Verdict, Readonly<VerdictStyle>>>; + +export const VERDICT_LABEL_MAX: number = VERDICT_LABEL_MAX_JS; + +export const VERDICT_COLOR_RE: RegExp = VERDICT_COLOR_RE_JS; + +export function isVerdict(v: unknown): v is Verdict { + return typeof v === "string" && (VERDICTS as readonly string[]).includes(v); +} + +// Defaults with the overrides laid over them, one key at a time. +export function resolveVerdicts(overrides: unknown): Record<Verdict, VerdictStyle> { + return resolveVerdictsJs(overrides) as Record<Verdict, VerdictStyle>; +} + +// Every reason an override cannot be used, as sentences naming `where`. +export function verdictOverrideProblems(v: unknown, where: string): string[] { + return verdictOverrideProblemsJs(v, where); +} diff --git a/common/package.json b/common/package.json @@ -33,6 +33,7 @@ "./components/*": "./components/*.tsx", "./lib/detectPlatform.mjs": "./lib/detectPlatform.mjs", "./lib/ports.mjs": "./lib/ports.mjs", + "./lib/report/verdicts.mjs": "./lib/report/verdicts.mjs", "./lib/toolProbe.mjs": "./lib/toolProbe.mjs", "./ytdlp/platformArgs.mjs": "./ytdlp/platformArgs.mjs", "./lib/*": "./lib/*.ts", diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -2,6 +2,7 @@ ## [Unreleased] - **A site can say what it publishes, and which reports.** `site.json` takes `publish` — `"full"`, the searchable corpus every site has been (the default, never written), or `"cited"`, only the site's reports and the moments they cite — and `reports`, the ordered ids of its published reports (each a slug; invalid and repeated ids are dropped). The site form has a Publish control and lists the site's reports read-only; saving the form keeps the stored list. A cited site still builds as a full one until the reports pipeline applies the scope. SITE.md documents both keys. +- **A report and its citations now have one written format, checked before anything is built from them.** A cited report is a `report.json` (`archilyzer-report`, version 1): a summary, then sections of claims, each claim with an optional verdict, the reviewed document's own sentence, findings in markdown and the citations it rests on. A citation is one of five kinds — a span of a video, a span of an audio record, a post, a sentence of a source document, or a web page — with a verbatim quote, and is cited from any markdown in the report as `[label](cite:<id>)`. The checker lists every problem at once with where it is: a citation, a source or a `cite:` link that names nothing, a span that ends before it starts or runs past 120 seconds with its context, a still that points outside the report's folder, an id used twice. Each cited span and post has one page address, `/m/<channel>/<id>/<start>-<end>/` or `/m/<channel>/<id>/`. Nothing builds or shows reports yet. The fact-check verdicts (Corroborated, Partly true, Contradicted, Not found, Untestable) and their colours are now kept in one place, which the report video's stamps and tally read too. `REPORT.md` and `CITATIONS.md` list every key. - **A long report video no longer runs out of memory while its clips are crossfaded.** `build-video.mjs` used to join every segment of a cut in one ffmpeg command, which grows with the number of segments: a cut of a few hundred clips could use more memory than the machine had and be stopped. Past 24 segments the build now crossfades them in batches of consecutive segments, each into a file under `out/<variant>/xfade-batches/`, then crossfades those files together with the same transition and lays the on-screen deck, the rail and the dips over them. Every transition, the deck's schedule and the chapters land on the same frames as before, at the cost of one more video encode on such a cut. The batch size is `render.xfadeBatch` in the manifest or `REPORT_VIDEO_XFADE_BATCH` in the environment (which wins); `0` never batches. A batch file is reused while its segments, their holds and moves and the encode settings are unchanged, so a `--chrome-only` run that moves no footage redoes only the last pass. - **Auto-download no longer tries a video the metadata scan already found members-only or private.** The scan records why it could not read a video, but only a failed download used to take a video out of the auto-download queue, so each members-only video the scan had found was still downloaded once: four yt-dlp requests, two with browser cookies, about 30 seconds each. A members-only or private answer from the scan now keeps the video out of the queue and counts it under the channel's members-only or private exclusions, and it stays in **Needs cookies** for a manual cookie run. A scan that reads the video later lifts this. A video the scan saw only as "Video unavailable" is still tried, since YouTube gives that answer when it is throttling too. - **X post fetches stop on Drain, and an account with no posts is not searched.** Draining a `fetch-posts` job used to do nothing until gallery-dl finished its whole run. Now the timeline fetch stops at the next page boundary (at once when gallery-dl is between pages or waiting out a rate limit), the older-posts walk stops its current window's search at once and never starts the 45–120 second pause between windows, and both keep their resume point: the job ends done, not failed, and the log says "Drained; the next run resumes …". An older-posts walk is refused when nothing is archived and the last timeline fetch finished having read no posts, since it would only repeat empty searches; `"force": true` (`--force` on `archilyzer posts fetch`) walks anyway. A walk with nothing archived that finds nothing ends after two empty three-month windows instead of four, and records why; a walk that has posts keeps the year-of-empty-windows rule. Capture-posts already stopped between posts on Drain. diff --git a/umtool/report-to-video/factcheck.mjs b/umtool/report-to-video/factcheck.mjs @@ -17,8 +17,20 @@ // round -- so the build, the composition and umtool all read one copy. // chrome-stamp.mjs draws the stamps; chrome-deck.mjs draws the tally. +// THE VERDICT VOCABULARY IS COMMON'S (common/lib/report/verdicts.mjs): the +// verdicts, their default labels and colours and the override rules are one +// copy shared with report.json and the export site's report pages. This file +// re-exports the list and lays the stamp and tally settings beside the labels. +import { + VERDICT_DEFAULTS, + VERDICT_LABEL_MAX, + VERDICTS, + resolveVerdicts, + verdictOverrideProblems, +} from "yt-dlp-transcript-common/lib/report/verdicts.mjs"; + /** The verdicts a claim may carry, in the order the tally lists them. */ -export const VERDICTS = Object.freeze(["CORROBORATED", "PARTLY", "CONTRADICTED", "NOT_FOUND", "UNTESTABLE"]); +export { VERDICTS }; /** Where the stamp sits in the footage box. */ export const STAMP_POSITIONS = Object.freeze(["top-left", "top-right", "bottom-left", "bottom-right", "center"]); @@ -33,19 +45,13 @@ export const TALLY_POSITIONS = Object.freeze(["right", "left"]); * (top-right by default) do not use. */ export const FACTCHECK_DEFAULTS = Object.freeze({ - verdicts: Object.freeze({ - CORROBORATED: Object.freeze({ label: "Corroborated", color: "#3fbf7f" }), - PARTLY: Object.freeze({ label: "Partly true", color: "#e3b23c" }), - CONTRADICTED: Object.freeze({ label: "Contradicted", color: "#e5534b" }), - NOT_FOUND: Object.freeze({ label: "Not found", color: "#8b93a7" }), - UNTESTABLE: Object.freeze({ label: "Untestable", color: "#7d8fd6" }), - }), + verdicts: VERDICT_DEFAULTS, stamp: Object.freeze({ seconds: 3, position: "top-left" }), tally: Object.freeze({ show: true, position: "right" }), }); /** The limits: a stamp's seconds on screen, a label's characters. */ -export const FACTCHECK_LIMITS = Object.freeze({ seconds: Object.freeze([1, 10]), label: 24 }); +export const FACTCHECK_LIMITS = Object.freeze({ seconds: Object.freeze([1, 10]), label: VERDICT_LABEL_MAX }); /** * The stamp's motion, in seconds: it slams in over `slam` (landing is when the @@ -60,7 +66,6 @@ const CLAIM_ID_RE = /^[A-Za-z0-9][A-Za-z0-9_.:-]{0,63}$/; const isObj = (v) => v !== null && typeof v === "object" && !Array.isArray(v); const numIn = (v, lo, hi) => typeof v === "number" && Number.isFinite(v) && v >= lo && v <= hi; -const HEX_RE = /^#(?:[0-9a-fA-F]{3}|[0-9a-fA-F]{6})$/; function unknownKeys(obj, allowed, where, errors) { for (const k of Object.keys(obj)) { @@ -71,13 +76,8 @@ function unknownKeys(obj, allowed, where, errors) { /** The fact-check settings with every default filled in. */ export function resolveFactcheck(render) { const f = isObj(render?.chrome?.factcheck) ? render.chrome.factcheck : {}; - const given = isObj(f.verdicts) ? f.verdicts : {}; - const verdicts = {}; - for (const v of VERDICTS) { - verdicts[v] = { ...FACTCHECK_DEFAULTS.verdicts[v], ...(isObj(given[v]) ? given[v] : {}) }; - } return { - verdicts, + verdicts: resolveVerdicts(f.verdicts), stamp: { ...FACTCHECK_DEFAULTS.stamp, ...(isObj(f.stamp) ? f.stamp : {}) }, tally: { ...FACTCHECK_DEFAULTS.tally, ...(isObj(f.tally) ? f.tally : {}) }, }; @@ -101,18 +101,7 @@ export function validateFactcheck(f, where = "render.chrome.factcheck") { for (const [k, v] of Object.entries(f.verdicts)) { const w = `${where}.verdicts.${k}`; if (!VERDICTS.includes(k)) { errors.push(`${w} is not a verdict (${VERDICTS.join(", ")})`); continue; } - if (!isObj(v)) { errors.push(`${w} must be { label, color }`); continue; } - unknownKeys(v, ["label", "color"], w, errors); - if (v.label !== undefined) { - if (typeof v.label !== "string" || !v.label.trim()) errors.push(`${w}.label must be words, not empty`); - else if (/[\r\n]/.test(v.label)) errors.push(`${w}.label must be one line`); - else if (v.label.trim().length > FACTCHECK_LIMITS.label) { - errors.push(`${w}.label is ${v.label.trim().length} characters (at most ${FACTCHECK_LIMITS.label})`); - } - } - if (v.color !== undefined && !(typeof v.color === "string" && HEX_RE.test(v.color))) { - errors.push(`${w}.color must be a hex colour, #rgb or #rrggbb`); - } + errors.push(...verdictOverrideProblems(v, w)); } } } diff --git a/umtool/report-to-video/factcheck.test.mjs b/umtool/report-to-video/factcheck.test.mjs @@ -11,10 +11,13 @@ import path from "node:path"; import test from "node:test"; import { - claimOf, FACTCHECK_DEFAULTS, normalizeClaim, originalUrlAt, resolveFactcheck, roundStamps, STAMP_MOTION, stampSchedule, + claimOf, FACTCHECK_DEFAULTS, FACTCHECK_LIMITS, normalizeClaim, originalUrlAt, resolveFactcheck, roundStamps, STAMP_MOTION, stampSchedule, tallyCues, tallyOf, tallySteps, tallyVerdicts, validateClaims, validateFactcheck, VERDICTS, } from "./factcheck.mjs"; import { + VERDICT_DEFAULTS, VERDICT_LABEL_MAX, VERDICTS as COMMON_VERDICTS, +} from "yt-dlp-transcript-common/lib/report/verdicts.mjs"; +import { deckGeometry, deckLayout, deckQrUrl, deckSchedule, DECK_DEFAULTS, estimateSchedule, feedGeometry, stampGeometry, validateChrome, } from "./deck.mjs"; @@ -42,6 +45,13 @@ const sched = (render = RENDER, entries = ENTRIES, D = 0.5, metas = []) => // ---- the claim ---------------------------------------------------------------- +test("the verdict vocabulary is common's, one copy: the same objects, not equal copies", () => { + assert.equal(VERDICTS, COMMON_VERDICTS); + assert.equal(FACTCHECK_DEFAULTS.verdicts, VERDICT_DEFAULTS); + assert.deepEqual(Object.keys(FACTCHECK_DEFAULTS.verdicts), [...VERDICTS]); + assert.equal(FACTCHECK_LIMITS.label, VERDICT_LABEL_MAX); +}); + test("normalizeClaim: trims the id, refuses a bad shape, an unknown field, a verdict off the list", () => { assert.equal(normalizeClaim(undefined), null); assert.equal(normalizeClaim(null), null); @@ -114,7 +124,7 @@ test("validateChrome refuses a bad factcheck block in sentences, unknown keys in for (const want of [ /render\.chrome\.factcheck\.colour is not a factcheck setting/, /render\.chrome\.factcheck\.verdicts\.MAYBE is not a verdict/, - /render\.chrome\.factcheck\.verdicts\.PARTLY\.size is not a factcheck setting/, + /render\.chrome\.factcheck\.verdicts\.PARTLY\.size is not a verdict setting/, /render\.chrome\.factcheck\.verdicts\.PARTLY\.label is 25 characters \(at most 24\)/, /render\.chrome\.factcheck\.verdicts\.PARTLY\.color must be a hex colour/, /render\.chrome\.factcheck\.verdicts\.NOT_FOUND must be \{ label, color \}/,