Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 38bc728ec0cb330a8d70705264977b06f0e7af85
parent b230836b0653a8cbc9d122d8cd2ce0240f7a32f7
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon,  5 Oct 2026 17:43:41 -0400

Merge report/header-card

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
MPUBLISH.md | 2+-
MREPORT.md | 2+-
Mcommon/components/citations/CitationCard.tsx | 42++++++++++++++++++++++++++----------------
Mcommon/lib/report/export.test.ts | 35++++++++++++++++++++++++-----------
Mcommon/lib/report/exportHtml.ts | 72+++++++++++++++++++++++++++++++++++++-----------------------------------
Mcommon/lib/report/exportMarkdown.ts | 48+++++++++++++++++++++---------------------------
Mcommon/lib/report/schema.ts | 2+-
Mcommon/lib/report/views.test.ts | 22----------------------
Mcommon/lib/report/views.ts | 58++++------------------------------------------------------
Meditor/CHANGELOG.md | 34+++++++++++++++++-----------------
Mexport/CHANGELOG.md | 18+++++++++---------
Mexport/app/components/reports/ReportArticle.tsx | 77+++++++++++++++++++++++++++++++++++++----------------------------------------
Mexport/app/components/reports/parts.tsx | 25++++++++++++++++++-------
Mexport/app/lib/reports.test.ts | 71++++++++++++++++++++++++++++++++++++++++++++++++-----------------------
Mexport/e2e-report/report-site.spec.ts | 29+++++++++++++++++++----------
Mexport/fixtures/report-site/public/m/demo-channel/abc123/3126.00-3151.00/moment.json | 6+++---
Mexport/fixtures/report-site/public/m/demo-podcast/ep-042/610.50-628.00/moment.json | 6+++---
Mexport/fixtures/report-site/public/m/demo-social/1234567890/moment.json | 2+-
Mexport/fixtures/report-site/public/reports/demo-factcheck/page.json | 1+
Mexport/fixtures/report-site/public/reports/index.json | 1+
Mexport/fixtures/report-site/source/demo-factcheck/report.json | 1+
Mexport/fixtures/report-site/source/demo-factcheck/report.v1.json | 1+
Mmcp/README.md | 2+-
Mplans/report-sites.md | 29++++++++++++++---------------
24 files changed, 289 insertions(+), 297 deletions(-)

diff --git a/PUBLISH.md b/PUBLISH.md @@ -141,7 +141,7 @@ deploy until it is built again. With no published reports it is an empty report A cited site is not a searchable archive, so the family never lists it, whatever `listed` says (`isListedSite`): it is no hub member, has no homepage card, counts in no family total and no other site's footer links to it. A full site with reports is -listed as before. Readers of the contract see a cited site as an empty corpus with +listed as `listed` says. Readers of the contract see a cited site as an empty corpus with reports: the MCP server's `list_reports` / `get_report` read them (and its `list_channels` says "cited-only site: N report(s)"), and umtool's cue walk refuses one by name — it has no transcripts to cut a report video from. diff --git a/REPORT.md b/REPORT.md @@ -18,7 +18,7 @@ Regenerate this file with `pnpm --filter yt-dlp-transcript-common exec tsx bin/f | `version` | yes | `1`. | | `id` | yes | The report's id: a lowercase slug (`[a-z0-9][a-z0-9-]*`, at most 64), its directory name under `reports/` and the last segment of its page, `/reports/<id>/`. Must match the directory. | | `kind` | yes | `"factcheck"` — sections of claims, each with a verdict — or `"sweep"` — sections with no verdicts, or bodies with inline citations. | -| `series` | no | The series the report belongs to: a recurring name for its kind of report. It heads the report's page in place of its title, is shown on its own line above the title, in the accent colour, in place of the kind's label, in the report list, and is joined to it as `<series>: <title>` where one line names the report (a page title, a link, a cited-in entry). | +| `series` | no | The series the report belongs to: a recurring name for its kind of report. It is shown on its own line above the title, in the accent colour, at the head of the report's page, its exports and the report list (where a report with no series shows its kind instead), and is joined to the title as `<series>: <title>` where one line names the report (a page title, a link, a cited-in entry). | | `title` | yes | The report's title. | | `subtitle` | no | A line under the title. | | `summary` | no | The report's summary, in markdown, shown before the sections. May cite inline: `[label](cite:<id>)`. | diff --git a/common/components/citations/CitationCard.tsx b/common/components/citations/CitationCard.tsx @@ -198,9 +198,11 @@ function PageLinks({ c }: { c: PageCitationView }) { } // Evidence the report's author found — the document under review did not give -// it (`origin: "added"`): the project's mark and the words, e.g. "Not in the -// article". A card shows both; a reference entry or a preview, the mark alone -// (the words for a screen reader). The claim's flag pill wears the same mark. +// it (`origin: "added"`): the project's mark, never words on screen (the +// words, e.g. "Not in the article", are its tooltip and what a screen reader +// says). A card wears it under its number box, as wide as the box (with no +// number, as wide as a one-digit box); a reference entry or a preview, inline +// after the kind. The claim's flag pill wears the same mark. export function AddedMark({ className }: { className?: string }) { return <BrandMark palette={ICON_PALETTES.archilyzer} className={cn("size-3.5 shrink-0", className)} />; } @@ -221,6 +223,7 @@ export function CitationCard({ const Icon = KIND_ICON[c.kind]; const { title, meta } = heading(c); const showNumber = variant !== "reference" && c.number !== undefined; + const markUnder = added && variant === "card"; const postText = c.kind === "post" && c.text && c.text.trim() !== c.quote.trim() ? c.text : null; return ( <div @@ -233,26 +236,33 @@ export function CitationCard({ className, )} > - {added && variant === "card" && ( - <p - data-citation-added="" - className="flex items-center gap-1.5 font-mono text-[10px] uppercase tracking-[0.14em] text-muted-foreground" - > - <AddedMark /> - {addedLabel} - </p> - )} <div className="flex min-w-0 items-start gap-2"> - {showNumber && ( - <span className="mt-px shrink-0 rounded bg-brand-soft px-1.5 font-mono text-xs font-semibold leading-5 text-brand-strong"> - {c.number} + {(showNumber || markUnder) && ( + <span className="mt-px inline-flex shrink-0 flex-col items-stretch gap-1"> + {showNumber && ( + <span className="rounded bg-brand-soft px-1.5 font-mono text-xs font-semibold leading-5 text-brand-strong"> + {c.number} + </span> + )} + {markUnder && ( + <span + data-citation-added="" + title={addedLabel} + // Sized by the number box, never sizing it: no width of its own + // (the svg would bring 300px) and at least the column's. + className={cn("block opacity-80", showNumber ? "w-0 min-w-full" : "w-[calc(1ch+0.75rem)] font-mono text-xs")} + > + <BrandMark palette={ICON_PALETTES.archilyzer} className="block h-auto w-full" /> + <span className="sr-only">{addedLabel}</span> + </span> + )} </span> )} <div className="min-w-0 flex-1"> <div className="flex items-center gap-1.5 font-mono text-[10px] uppercase tracking-[0.14em] text-muted-foreground"> <Icon className="size-3 shrink-0" aria-hidden /> {CITATION_KIND_LABELS[c.kind]} - {added && variant !== "card" && ( + {added && !markUnder && ( <span data-citation-added="" title={addedLabel} className="inline-flex items-center"> <AddedMark className="size-3" /> <span className="sr-only">{addedLabel}</span> diff --git a/common/lib/report/export.test.ts b/common/lib/report/export.test.ts @@ -102,15 +102,15 @@ test("the HTML export is deterministic, holds no script, and every image is a da assert.doesNotMatch(html, /<link |@import|url\(/); }); -test("the heading: the series (else the title), the dates and revision, the document as a cited line", () => { +test("the heading: the series and the title (or the title alone), the dates and revision, the document as a cited line", () => { const html = reportExportHtml(view, opts); assert.match( html, - /<header class="report-head"><h1>Demo Checks<\/h1>\n<p class="dates" data-report-dates="">2026-10-01 · updated 2026-10-04<\/p>\n<p class="subject" data-subject-source="s0" style="border-left-color:#aa3300"><a href="https:\/\/example\.org\/a">An article<\/a> <span class="meta">— A\. Writer · Example Gazette<\/span><\/p>\n<p class="subtitle">Two claims, tested<\/p><\/header>/, + /<header class="report-head"><h1><span class="series" data-report-series="">Demo Checks<\/span> <span class="title">A demo fact-check<\/span><\/h1>\n<p class="dates" data-report-dates="">2026-10-01 · updated 2026-10-04<\/p>\n<p class="subject" data-subject-source="s0" style="border-left-color:#aa3300"><a href="https:\/\/example\.org\/a">An article<\/a> <span class="meta">— <span style="white-space:nowrap">A\. Writer<\/span> · <span style="white-space:nowrap">Example Gazette<\/span><\/span><\/p>\n<p class="subtitle">Two claims, tested<\/p><\/header>/, ); assert.match(html, /<title>Demo Checks: A demo fact-check<\/title>/, "the title stays the document's name"); assert.doesNotMatch(html, /data-report-attribution|data-report-byline|Under review|Fact-check by|class="eyebrow"/); - // No series: the title is the heading. + // No series: the title alone. assert.match(reportExportHtml({ ...view, series: undefined }, opts), /<h1>A demo fact-check<\/h1>/); // No subject: no cited line. assert.doesNotMatch(reportExportHtml({ ...view, subject: undefined }, opts), /class="subject"/); @@ -133,7 +133,8 @@ test("the three tiers: the quick take, what the check found, every claim from ho const at = order.map((m) => html.indexOf(m)); assert.ok(at.every((i) => i >= 0), JSON.stringify(at)); assert.deepEqual([...at].sort((a, b) => a - b), at, "in the page's order"); - assert.match(html, /<div class="tier" data-tier-marker="1"><span class="dots" aria-hidden="true"><i class="on"><\/i><i class=""><\/i><i class=""><\/i><\/span><span class="min">1 min<\/span>/); + assert.match(html, /<div class="tier" data-tier-marker="1" aria-hidden="true"><span class="dots"><i class="on"><\/i><i class=""><\/i><i class=""><\/i><\/span><span class="rule"><\/span><\/div>/); + assert.doesNotMatch(html, /\d+ min\b/); assert.match(html, /<a href="#found">What the check found<\/a> · <a href="#claims">Every claim<\/a> · <a href="#references">References<\/a>/); // Grouped in FOUND_VERDICT_ORDER, each row its title, gist and flag, to the claim. const groups = [...html.matchAll(/data-found-group="([A-Z_]+)"/g)].map((m) => m[1]); @@ -152,13 +153,22 @@ test("a claim: its flag, the subject's sentence on the accent rail with no link, assert.match(claim, /data-verdict="CONTRADICTED"[^]*<span class="flag" data-claim-flag=""><svg class="mark"[^>]*>[^]*<\/svg>No source given<\/span><h3>One<\/h3>/); assert.match(claim, /<figure class="sentence" data-source-sentence="s1" style="border-left-color:#aa3300">/); assert.doesNotMatch(claim, /from <a href="#source-s0">/, "the rail says whose sentence it is"); + // The document's sentence stands for the claim: no paraphrase beside it. + assert.doesNotMatch(claim, /Claim one\.|data-claim-text|says:/); + // Without a sentence, a titled claim's text follows plain, unquoted. + const noSentence = structuredClone(view); + delete noSentence.sections[0].claims[0].sourceQuote; + const plainHtml = reportExportHtml(noSentence, opts); + assert.match(plainHtml, /<h3>One<\/h3><\/div>\n<p class="claim-text" data-claim-text="">Claim one\.<\/p>/); + assert.match(reportExportMarkdown(noSentence, opts), /^### One\n\n(?:.*\n\n)?Claim one\.\n/m); + assert.doesNotMatch(reportExportMarkdown(view, opts), /Claim one\./); assert.match(html, /<p class="subject" data-subject-source="s0" style="border-left-color:#aa3300">/, "the heading's cited line wears the same rail"); // Added first, marked; the subject's own folded under "In the article". - assert.match(claim, /<li data-evidence="v1"><a href="#c-v1">\[1\]<\/a> <span class="added" data-citation-added=""><svg class="mark"[^]*?<\/svg>Not in the article<\/span> Video/); + assert.match(claim, /<li data-evidence="v1"><a href="#c-v1">\[1\]<\/a> <span class="added" data-citation-added="" title="Not in the article"><svg class="mark"[^]*?<\/svg><span class="sr">Not in the article<\/span><\/span> Video/); assert.match(claim, /<details class="given" data-subject-evidence=""><summary>In the article \(1\)<\/summary><ul class="evidence"><li data-evidence="a1">/); assert.match(reportExportHtml(view, { ...opts, print: true }), /<details class="given" data-subject-evidence="" open>/, "printed open"); // The reference list marks added evidence too. - assert.match(html, /<li id="c-v1"[^>]*><p class="quote"><q>one \* two<\/q><\/p><p class="meta"><span class="added"/); + assert.match(html, /<li id="c-v1"[^>]*><p class="quote"><span class="added" data-citation-added=""[^]*?<\/span><\/span> <q>one \* two<\/q><\/p>/); }); test("verdicts, inline markers [n] to the reference list, and references with every link", () => { @@ -212,23 +222,26 @@ test("the Markdown export: the heading lines, `label [n]` citations, numbered re assert.equal(reportExportMarkdown(view, opts), md); assert.ok( md.startsWith( - "# Demo Checks\n\n2026-10-01 · updated 2026-10-04\n\n" + + "**Demo Checks**\n\n# A demo fact-check\n\n2026-10-01 · updated 2026-10-04\n\n" + "[An article](https://example.org/a) — A. Writer · Example Gazette\n\n*Two claims, tested*\n", ), md.slice(0, 300), ); + assert.ok(reportExportMarkdown({ ...view, series: undefined }, opts).startsWith("# A demo fact-check\n\n2026-10-01"), "no series: the title alone"); assert.match(md, /It starts here \[1\]\./); assert.match(md, /See this \[3\], \*\*bold\*\*/); assert.match(md, /^### One$/m); assert.match(md, /^\*\*Verdict: Contradicted\*\* \*\[No source given\]\*$/m); - // The tiers as plain lines; what the check found by verdict; how it was checked. - assert.deepEqual([...md.matchAll(/^— (\d+) min —$/gm)].map((m) => m[1]), ["1", "1", "1"]); + // Each tier after a rule (and the footer after one); what the check found by verdict; how it was checked. + assert.equal([...md.matchAll(/^---$/gm)].length, 4); + assert.doesNotMatch(md, /\d+ min\b/); assert.match(md, /## What the check found\n\n### Contradicted \(1\)\n\n- One — The recording says otherwise\. \*\[No source given\]\*\n/); assert.match(md, /## Every claim, with its evidence\n\n\*\*How it was checked\*\*\n\nEach quote was checked against the recording\.\n/); // Evidence by origin: the added marked in words, the subject's under its own line. - assert.match(md, /Evidence:\n\n- \[1\] Not in the article: Video, /); + assert.match(md, /Evidence:\n\n- \[1\] Video, /); + assert.doesNotMatch(md, /Not in the/, "the Markdown carries no added marker"); assert.match(md, /In the article:\n\n- \[4\] Audio, /); - assert.match(md, /^1\. “one \\\* two” \n Not in the article \n/m); + assert.match(md, /^1\. “one \\\* two” \n Video · /m); assert.doesNotMatch(reportExportMarkdown({ ...view, kind: "sweep" }, opts), /What the check found/); const refs = [...md.matchAll(/^(\d+)\. “(.*)”/gm)].map((m) => `${m[1]}:${m[2]}`); assert.deepEqual(refs, ["1:one \\* two", "2:four", "3:five", "4:two", "5:three"]); diff --git a/common/lib/report/exportHtml.ts b/common/lib/report/exportHtml.ts @@ -10,16 +10,17 @@ // answers — a data: URI for the one-file export, a relative path in the pack. // Clips are LINKED, never inlined (the pack plays them from its media/). // -// What it holds, as the site's report page has it: the heading (the series, -// else the title; one line of dates and the revision; the document under -// review as a cited line on its colour's rail; the subtitle), then the three tiers, each -// opened by a marker (depth dots, minutes to read): the quick take (the +// What it holds, as the site's report page has it: the heading (the series +// on one line and the title on the next, or the title alone; one line of +// dates and the revision; the document under review as a cited line on its +// colour's rail; the subtitle), then the three tiers, each opened by a +// marker (a hairline and depth dots): the quick take (the // tally, the summary, jump links); what the check found (a fact-check's // claims by verdict, a line each with its gist and flag); and every claim // with its evidence, opening with how it was checked — each claim its // verdict and flag pill, the document's sentence on its rail, the findings // with numbered markers, and the evidence: what the report added first -// (the project's mark, "Not in the article"), then the rest, then what the +// (the project's mark beside its number), then the rest, then what the // document gave itself under "In the article (n)" (a <details>, printed // open). Then the documents quoted with their archive links, the numbered // references (quote, speaker, date, record, the original at its time, the @@ -39,7 +40,6 @@ import { reportFullTitle, reportPagePath, reportRevisionLabel, - reportTierMinutes, sourceAnchor, spanLabel, subjectMetaParts, @@ -278,6 +278,8 @@ main{max-width:46rem;margin:0 auto;padding:2rem 1rem 3rem} a{color:var(--accent);text-decoration:underline;text-decoration-thickness:1px;text-underline-offset:2px;overflow-wrap:anywhere} h1,h2,h3,h4{line-height:1.25;margin:0} h1{font-size:2rem;font-weight:600;letter-spacing:-.01em;overflow-wrap:anywhere} +h1 .series{display:block;color:var(--accent)} +h1 .title{display:block;font-weight:400} h2{font-size:1.5rem;font-weight:600;margin-top:2.5rem;padding-top:1rem;border-top:1px solid var(--border)} h3{font-size:1.15rem;font-weight:600} h4,h5,h6{font-size:1rem;font-weight:600;margin:1rem 0 .25rem} @@ -300,7 +302,6 @@ header.report-head{padding-bottom:1.25rem;border-bottom:1px solid var(--border)} .tier .dots{display:inline-flex;gap:.25rem} .tier .dots i{display:inline-block;width:.4rem;height:.4rem;border-radius:50%;border:1px solid var(--muted)} .tier .dots i.on{background:var(--accent);border-color:var(--accent)} -.tier .min{font:.75rem/1 ui-monospace,SFMono-Regular,Menlo,monospace;color:var(--muted)} .tier .rule{flex:1;height:1px;background:var(--border)} .tier+h2{margin-top:.5rem;padding-top:0;border-top:0} h2.section{font-size:1.3rem} @@ -313,7 +314,8 @@ h2.section{font-size:1.3rem} .method{margin:.75rem 0} .flag,.added{display:inline-flex;align-items:center;gap:.35rem;white-space:nowrap} .flag{border:1px solid color-mix(in srgb,var(--accent) 40%,#fff);background:color-mix(in srgb,var(--accent) 8%,#fff);color:var(--accent);border-radius:999px;padding:.05rem .6rem;font-size:.78rem;font-weight:500} -.added{font:.7rem/1.2 ui-monospace,SFMono-Regular,Menlo,monospace;text-transform:uppercase;letter-spacing:.12em;color:var(--muted)} +.added{opacity:.8;vertical-align:-.1em} +.sr{position:absolute;width:1px;height:1px;overflow:hidden;clip:rect(0 0 0 0);white-space:nowrap} svg.mark{flex:none;border-radius:22%} details.given{margin:.4rem 0 0} details.given summary{cursor:pointer;font:.8rem/1.4 ui-monospace,SFMono-Regular,Menlo,monospace;color:var(--muted)} @@ -326,8 +328,7 @@ nav.toc ol{margin:.5rem 0;padding-left:1.4rem} .claim{margin:1.25rem 0;padding:1rem 1.1rem;border:1px solid var(--border);border-radius:8px} .claim-head{display:flex;flex-wrap:wrap;align-items:baseline;gap:.4rem .7rem} .claim-head h3{flex:1 1 15rem} -.says{color:var(--muted);font-size:.95rem} -.says q{color:var(--fg)} +.claim-text{color:var(--muted);font-size:.95rem} figure.sentence{margin:.75rem 0;padding-left:.75rem;border-left:4px solid var(--border)} figure.sentence img{display:block;border:1px solid var(--border);border-radius:4px;background:#fff;break-inside:avoid} figure.sentence figcaption{font-size:.85rem;color:var(--muted);margin-top:.3rem} @@ -415,14 +416,15 @@ function link(href: string, text?: string): string { } // Evidence the report's author found (`origin: "added"`): the project's mark -// (the site's AddedMark, as a static SVG) and the words. +// alone (the site's AddedMark, as a static SVG), beside the citation's number; +// the words are its tooltip and what a screen reader says. const ADDED_MARK = markSvg(ICON_PALETTES.archilyzer).replace( "<svg ", '<svg class="mark" width="12" height="12" aria-hidden="true" focusable="false" ', ); function addedTag(label: string): string { - return `<span class="added" data-citation-added="">${ADDED_MARK}${escapeHtml(label)}</span>`; + return `<span class="added" data-citation-added="" title="${escapeHtml(label)}">${ADDED_MARK}<span class="sr">${escapeHtml(label)}</span></span>`; } function flagPill(flag: string): string { @@ -476,8 +478,7 @@ function reference(c: CitationView, opts: ReportExportHtmlOptions, addedLabel: s const extra = [c.label, c.note].filter(Boolean).map((x) => `<p class="meta">${escapeHtml(x!)}</p>`).join(""); return ( `<li id="${escapeHtml(citationAnchor(c.id))}" value="${c.number ?? ""}" data-reference="${escapeHtml(c.id)}">` + - `<p class="quote"><q>${escapeHtml(c.quote)}</q></p>` + - (c.origin === "added" ? `<p class="meta">${addedTag(addedLabel)}</p>` : "") + + `<p class="quote">${c.origin === "added" ? `${addedTag(addedLabel)} ` : ""}<q>${escapeHtml(c.quote)}</q></p>` + `<p class="meta">${meta.join(" · ")}</p>` + extra + (links.length > 0 ? `<p class="links">${links.join("<br>")}</p>` : "") + @@ -497,7 +498,6 @@ function evidenceItem(c: CitationView, addedLabel: string): string { type ClaimContext = { view: ReportPageView; - subjectLabel: string; subjectNoun: string; cite: (id: string) => CiteRef | undefined; opts: ReportExportHtmlOptions; @@ -511,9 +511,12 @@ function claimHtml(claim: ClaimView, x: ClaimContext): string { `<div class="claim-head">${claim.verdict ? verdictChip(view, claim.verdict) : ""}` + `${claim.flag ? flagPill(claim.flag) : ""}<h3>${escapeHtml(claim.title ?? claim.text)}</h3></div>`, ); - if (claim.title) parts.push(`<p class="says">${escapeHtml(x.subjectLabel)}: <q>${escapeHtml(claim.text)}</q></p>`); + // The document's own sentence (its still, else its verbatim words) when the + // claim cites one; else the claim as the report states it, plain. const sentence = claim.sourceQuote ? view.citations[claim.sourceQuote] : undefined; - if (sentence?.kind === "source") { + if (sentence?.kind !== "source") { + if (claim.title) parts.push(`<p class="claim-text" data-claim-text="">${escapeHtml(claim.text)}</p>`); + } else { // On a rail in the document's colour; a sentence of the document under // review needs no link to it (the rail says whose it is). const isSubject = sentence.sourceId === view.subject; @@ -563,13 +566,10 @@ function railStyle(accent: string | undefined): string { } // Where a tier of the page begins: a hairline with, at its left, how deep it -// goes (one, two or three dots) and how long it takes to read. -function tierMarker(depth: 1 | 2 | 3, minutes: number): string { +// goes (one, two or three dots). Decoration only. +function tierMarker(depth: 1 | 2 | 3): string { const dots = [1, 2, 3].map((i) => `<i class="${i <= depth ? "on" : ""}"></i>`).join(""); - return ( - `<div class="tier" data-tier-marker="${depth}"><span class="dots" aria-hidden="true">${dots}</span>` + - `<span class="min">${Math.max(1, minutes)} min</span><span class="rule" aria-hidden="true"></span></div>` - ); + return `<div class="tier" data-tier-marker="${depth}" aria-hidden="true"><span class="dots">${dots}</span><span class="rule"></span></div>`; } // ─── The document ─── @@ -583,18 +583,20 @@ export function reportExportHtml(view: ReportPageView, opts: ReportExportHtmlOpt const md = (text: string, headingBase = 3) => markdownToHtml(text, cite, { siteUrl: opts.siteUrl, headingBase }); const subject = view.subject ? view.sources[view.subject] : undefined; const subjectNoun = subject?.kind === "article" ? "article" : "source"; - const subjectLabel = `The ${subjectNoun} says`; const tally = isFactcheck ? verdictTally(view) : []; const groups = isFactcheck ? foundGroups(view) : []; - const minutes = reportTierMinutes(view); const claimCount = view.sections.reduce((n, s) => n + s.claims.length, 0); const pageUrl = siteLink(opts.siteUrl, reportPagePath(view.id)); - const ctx: ClaimContext = { view, subjectLabel, subjectNoun, cite, opts }; - - // The heading, in the site's order: the series (else the title), one line - // of dates and this export's revision, the document under review as a - // cited line on its colour's rail, the subtitle. - const head: string[] = [`<h1>${escapeHtml(view.series ?? view.title)}</h1>`]; + const ctx: ClaimContext = { view, subjectNoun, cite, opts }; + + // The heading, in the site's order: the series on one line and the title + // on the next (or the title alone), one line of dates and this export's + // revision, the document under review as a cited line on its colour's + // rail, the subtitle. + const name = view.series + ? `<span class="series" data-report-series="">${escapeHtml(view.series)}</span> <span class="title">${escapeHtml(view.title)}</span>` + : escapeHtml(view.title); + const head: string[] = [`<h1>${name}</h1>`]; const historyUrl = view.history ? siteLink(opts.siteUrl, view.history.href) : undefined; const revision = opts.footer.revision !== undefined @@ -605,7 +607,7 @@ export function reportExportHtml(view: ReportPageView, opts: ReportExportHtmlOpt const dateLine = [...reportDateParts(view).map(escapeHtml), ...(revision ? [revision] : [])]; if (dateLine.length > 0) head.push(`<p class="dates" data-report-dates="">${dateLine.join(" · ")}</p>`); if (subject) { - const meta = subjectMetaParts(subject).map(escapeHtml).join(" · "); + const meta = subjectMetaParts(subject).map((x) => `<span style="white-space:nowrap">${escapeHtml(x)}</span>`).join(" · "); head.push( `<p class="subject" data-subject-source="${escapeHtml(subject.id)}"${railStyle(subject.accent)}>${sourceTitleHtml(subject)}` + `${meta ? ` <span class="meta">— ${meta}</span>` : ""}</p>`, @@ -615,7 +617,7 @@ export function reportExportHtml(view: ReportPageView, opts: ReportExportHtmlOpt const body: string[] = [`<header class="report-head">${head.join("\n")}</header>`]; // Tier 1, the quick take: the tally, the summary, where to go next. - const quick: string[] = [tierMarker(1, minutes.quick)]; + const quick: string[] = [tierMarker(1)]; if (tally.length > 0) { quick.push( `<p class="label">${plural(claimCount, "claim")} checked</p>` + @@ -634,7 +636,7 @@ export function reportExportHtml(view: ReportPageView, opts: ReportExportHtmlOpt // Tier 2, what the check found: every ruled claim by verdict, one line each. if (groups.length > 0) { body.push( - `<section id="found" data-report-tier="2">${tierMarker(2, minutes.found ?? 0)}<h2>What the check found</h2>` + + `<section id="found" data-report-tier="2">${tierMarker(2)}<h2>What the check found</h2>` + groups .map( (g) => @@ -654,7 +656,7 @@ export function reportExportHtml(view: ReportPageView, opts: ReportExportHtmlOpt } // Tier 3, every claim with its evidence, opening with how it was checked. - const full: string[] = [tierMarker(3, minutes.claims), `<h2>Every claim, with its evidence</h2>`]; + const full: string[] = [tierMarker(3), `<h2>Every claim, with its evidence</h2>`]; if (view.method) full.push(`<div class="method" data-report-method=""><h3 class="label">How it was checked</h3>${md(view.method)}</div>`); if (view.sections.length > 1) { full.push( diff --git a/common/lib/report/exportMarkdown.ts b/common/lib/report/exportMarkdown.ts @@ -22,7 +22,6 @@ import { reportDateParts, reportPagePath, reportRevisionLabel, - reportTierMinutes, spanLabel, subjectMetaParts, verdictTally, @@ -87,7 +86,7 @@ function citationDate(c: CitationView): string | undefined { return dateLabel(c.date ?? (c.kind === "video" || c.kind === "audio" || c.kind === "post" ? c.record.date : undefined)); } -function referenceLines(c: CitationView, siteUrl: string | undefined, addedLabel: string): string[] { +function referenceLines(c: CitationView, siteUrl: string | undefined): string[] { const meta = [ CITATION_KIND_LABELS[c.kind], c.speaker ? mdText(c.speaker) : null, @@ -96,7 +95,6 @@ function referenceLines(c: CitationView, siteUrl: string | undefined, addedLabel c.verification?.quoteScore !== undefined ? `quote match ${Math.round(c.verification.quoteScore * 100)}%` : null, ].filter(Boolean); const out = [`${c.number ?? "-"}. “${mdText(c.quote)}”`]; - if (c.origin === "added") out.push(` ${addedLabel}`); out.push(` ${meta.join(" · ")}`); for (const x of [c.label, c.note]) if (x) out.push(` ${mdText(x)}`); const link = (label: string, url: string) => out.push(` ${label}: <${url}>`); @@ -127,32 +125,28 @@ function referenceLines(c: CitationView, siteUrl: string | undefined, addedLabel return out.map((l, i, a) => (i < a.length - 1 ? `${l} ` : l)); } -// A tier's start, as a plain line: `— 6 min —`. -function tierLine(minutes: number): string { - return `— ${Math.max(1, minutes)} min —`; -} - -function evidenceLine(c: CitationView, addedLabel: string): string { - return `- [${c.number ?? "?"}] ${c.origin === "added" ? `${addedLabel}: ` : ""}${CITATION_KIND_LABELS[c.kind]}, ${citationWhere(c)}: “${mdText(c.quote)}”`; +function evidenceLine(c: CitationView): string { + return `- [${c.number ?? "?"}] ${CITATION_KIND_LABELS[c.kind]}, ${citationWhere(c)}: “${mdText(c.quote)}”`; } function claimLines( claim: ClaimView, view: ReportPageView, - subjectLabel: string, subjectNoun: string, siteUrl: string | undefined, ): string[] { - const addedLabel = `Not in the ${subjectNoun}`; const out: string[] = [`### ${mdText(claim.title ?? claim.text)}`, ""]; const marks = [ claim.verdict ? `**Verdict: ${mdText(view.verdicts[claim.verdict].label)}**` : null, claim.flag ? `*[${mdText(claim.flag)}]*` : null, ].filter(Boolean); if (marks.length > 0) out.push(marks.join(" "), ""); - if (claim.title) out.push(`${subjectLabel}: “${mdText(claim.text)}”`, ""); + // The document's own sentence when the claim cites one; else the claim as + // the report states it, plain. const sentence = claim.sourceQuote ? view.citations[claim.sourceQuote] : undefined; - if (sentence?.kind === "source") { + if (sentence?.kind !== "source") { + if (claim.title) out.push(mdText(claim.text), ""); + } else { out.push(`> “${mdText(sentence.quote)}” [${sentence.number ?? "?"}]`); if (sentence.sourceId !== view.subject) out.push(">", `> — from *${mdText(sentence.sourceTitle)}*`); out.push(""); @@ -167,8 +161,8 @@ function claimLines( // the document under review gave itself. const shown = [...evidence.filter((c) => c.origin === "added"), ...evidence.filter((c) => c.origin === undefined)]; const given = evidence.filter((c) => c.origin === "subject"); - if (shown.length > 0) out.push("Evidence:", "", ...shown.map((c) => evidenceLine(c, addedLabel)), ""); - if (given.length > 0) out.push(`In the ${subjectNoun}:`, "", ...given.map((c) => evidenceLine(c, addedLabel)), ""); + if (shown.length > 0) out.push("Evidence:", "", ...shown.map((c) => evidenceLine(c)), ""); + if (given.length > 0) out.push(`In the ${subjectNoun}:`, "", ...given.map((c) => evidenceLine(c)), ""); return out; } @@ -176,14 +170,14 @@ export function reportExportMarkdown(view: ReportPageView, opts: ReportExportOpt const out: string[] = []; const subject = view.subject ? view.sources[view.subject] : undefined; const subjectNoun = subject?.kind === "article" ? "article" : "source"; - const subjectLabel = `The ${subjectNoun} says`; const pageUrl = siteLink(opts.siteUrl, reportPagePath(view.id)); const isFactcheck = view.kind === "factcheck"; - const minutes = reportTierMinutes(view); - // The series (else the title), one line of dates and this export's - // revision, the document under review as a cited line, the subtitle. - out.push(`# ${mdText(view.series ?? view.title)}`, ""); + // The series on its own line and the title as the heading (or the title + // alone), one line of dates and this export's revision, the document under + // review as a cited line, the subtitle. + if (view.series) out.push(`**${mdText(view.series)}**`, ""); + out.push(`# ${mdText(view.title)}`, ""); const historyUrl = view.history ? siteLink(opts.siteUrl, view.history.href) : undefined; const revision = opts.footer.revision !== undefined @@ -199,8 +193,8 @@ export function reportExportMarkdown(view: ReportPageView, opts: ReportExportOpt } if (view.subtitle) out.push(`*${mdText(view.subtitle)}*`, ""); - // Tier 1: the quick take. - out.push(tierLine(minutes.quick), ""); + // Tier 1: the quick take, after a rule (each tier opens with one). + out.push("---", ""); const tally = isFactcheck ? verdictTally(view) : []; if (tally.length > 0) { const claims = view.sections.reduce((n, s) => n + s.claims.length, 0); @@ -214,7 +208,7 @@ export function reportExportMarkdown(view: ReportPageView, opts: ReportExportOpt // Tier 2: what the check found (a fact-check with verdicts). const groups = isFactcheck ? foundGroups(view) : []; if (groups.length > 0) { - out.push(tierLine(minutes.found ?? 0), "", "## What the check found", ""); + out.push("---", "", "## What the check found", ""); for (const g of groups) { out.push(`### ${mdText(view.verdicts[g.verdict].label)} (${g.claims.length})`, ""); for (const c of g.claims) { @@ -225,12 +219,12 @@ export function reportExportMarkdown(view: ReportPageView, opts: ReportExportOpt } // Tier 3: every claim, with its evidence. - out.push(tierLine(minutes.claims), "", "## Every claim, with its evidence", ""); + out.push("---", "", "## Every claim, with its evidence", ""); if (view.method) out.push("**How it was checked**", "", citedMarkdownToPlain(view.method, view, opts.siteUrl).trim(), ""); for (const s of view.sections) { out.push(`## ${mdText(s.title)}`, ""); if (s.body) out.push(citedMarkdownToPlain(s.body, view, opts.siteUrl).trim(), ""); - for (const c of s.claims) out.push(...claimLines(c, view, subjectLabel, subjectNoun, opts.siteUrl)); + for (const c of s.claims) out.push(...claimLines(c, view, subjectNoun, opts.siteUrl)); } const sources = Object.values(view.sources); @@ -249,7 +243,7 @@ export function reportExportMarkdown(view: ReportPageView, opts: ReportExportOpt const refs = orderedCitations(view); if (refs.length > 0) { out.push("## References", ""); - for (const c of refs) out.push(...referenceLines(c, opts.siteUrl, `Not in the ${subjectNoun}`), ""); + for (const c of refs) out.push(...referenceLines(c, opts.siteUrl), ""); } out.push("---", "", reportExportFooterLine(opts.footer)); diff --git a/common/lib/report/schema.ts b/common/lib/report/schema.ts @@ -101,7 +101,7 @@ export const REPORT_FIELD_DOCS: FieldDocs<Report> = { version: `\`${REPORT_VERSION}\`.`, id: "The report's id: a lowercase slug (`[a-z0-9][a-z0-9-]*`, at most 64), its directory name under `reports/` and the last segment of its page, `/reports/<id>/`. Must match the directory.", kind: '`"factcheck"` — sections of claims, each with a verdict — or `"sweep"` — sections with no verdicts, or bodies with inline citations.', - series: "The series the report belongs to: a recurring name for its kind of report. It heads the report's page in place of its title, is shown on its own line above the title, in the accent colour, in place of the kind's label, in the report list, and is joined to it as `<series>: <title>` where one line names the report (a page title, a link, a cited-in entry).", + series: "The series the report belongs to: a recurring name for its kind of report. It is shown on its own line above the title, in the accent colour, at the head of the report's page, its exports and the report list (where a report with no series shows its kind instead), and is joined to the title as `<series>: <title>` where one line names the report (a page title, a link, a cited-in entry).", title: "The report's title.", subtitle: "A line under the title.", summary: "The report's summary, in markdown, shown before the sections. May cite inline: `[label](cite:<id>)`.", diff --git a/common/lib/report/views.test.ts b/common/lib/report/views.test.ts @@ -16,9 +16,6 @@ import { citedInViews, evidenceClipPath, foundGroups, - readingMinutes, - reportTierMinutes, - wordCount, momentViewPath, orderedCitations, reportAssetPath, @@ -371,25 +368,6 @@ test("a citation's origin rides its view; a report with no subject carries none" assert.ok(validateReport(bad).length > 0); }); -test("reading time: words (a cite link is its label) at 230 a minute, rounded up", () => { - assert.equal(wordCount("He says [so on stream](cite:c01), twice."), 6); - assert.equal(wordCount("**Bold** and `code` — it's 2019."), 5); - assert.equal(wordCount(undefined), 0); - assert.equal(readingMinutes(0), 0); - assert.equal(readingMinutes(1), 1); - assert.equal(readingMinutes(230), 1); - assert.equal(readingMinutes(231), 2); - const long = structuredClone(report); - long.method = "word ".repeat(460); - const m = reportTierMinutes(buildReportPageView(long, { record: (c) => record(c.channel, c.id) })); - assert.ok(m.claims >= 3, `the method is read with the claims: ${m.claims}`); - assert.equal(m.quick, reportTierMinutes(view).quick, "the method is not in the quick take"); - const sweep = structuredClone(report); - sweep.kind = "sweep"; - for (const s of sweep.sections) for (const c of s.claims ?? []) delete c.verdict; - assert.equal(reportTierMinutes(buildReportPageView(sweep, { record: (c) => record(c.channel, c.id) })).found, undefined); -}); - test("what the check found: ruled claims by verdict, changed first, confirmed last; gist and method ride the view", () => { const r = structuredClone(report); r.method = "Searched the transcripts."; diff --git a/common/lib/report/views.ts b/common/lib/report/views.ts @@ -267,8 +267,8 @@ export type ReportPageView = { version: typeof REPORT_VIEWS_VERSION; id: string; kind: ReportKind; - // Its own line above the title, in place of the kind's label; `<series>: <title>` - // where one line names the report (reportFullTitle). + // Its own line above the title; `<series>: <title>` where one line names + // the report (reportFullTitle). series?: string; title: string; subtitle?: string; @@ -306,8 +306,8 @@ export type VerdictCount = { verdict: Verdict; count: number }; export type ReportIndexEntry = { id: string; kind: ReportKind; - // Its own line above the title, in place of the kind's label; `<series>: <title>` - // where one line names the report (reportFullTitle). + // Its own line above the title; `<series>: <title>` where one line names + // the report (reportFullTitle). series?: string; title: string; subtitle?: string; @@ -686,56 +686,6 @@ export function foundGroups(view: Pick<ReportPageView, "sections">): FoundGroup[ ); } -export const READING_WPM = 230; - -// The words a reader reads in a text: a cite link counts as its label, and -// markdown's punctuation as nothing. -export function wordCount(text: string | undefined): number { - if (!text) return 0; - return (text.replace(/\]\([^)]*\)/g, "]").match(/[\p{L}\p{N}][\p{L}\p{N}'’.-]*/gu) ?? []).length; -} - -// Minutes to read so many words, rounded up; nothing to read is no minutes. -export function readingMinutes(words: number): number { - return words > 0 ? Math.ceil(words / READING_WPM) : 0; -} - -export type ReportTierMinutes = { quick: number; found?: number; claims: number }; - -// Each tier's reading time: the quick take (the tally and the summary), what -// the check found (each row's title and gist; a fact-check with verdicts -// only), and every claim in full (the method, section bodies, each claim's -// text and findings, and the quote of every citation they show, once). -export function reportTierMinutes(view: ReportPageView): ReportTierMinutes { - const sum = (xs: (string | undefined)[]) => xs.reduce((n, x) => n + wordCount(x), 0); - const tally = - view.kind === "factcheck" - ? verdictTally(view).map((t) => `${view.verdicts[t.verdict]?.label ?? t.verdict} ${t.count}`) - : []; - const groups = view.kind === "factcheck" ? foundGroups(view) : []; - const claims = view.sections.flatMap((s) => s.claims); - const quoted = new Set<string>(); - for (const c of claims) { - if (c.sourceQuote) quoted.add(c.sourceQuote); - for (const id of c.citations) quoted.add(id); - } - return defined({ - quick: readingMinutes(sum([view.summary, ...tally])), - found: - groups.length > 0 - ? readingMinutes(sum(groups.flatMap((g) => g.claims.flatMap((c) => [c.title ?? c.text, c.gist])))) - : undefined, - claims: readingMinutes( - sum([ - view.method, - ...view.sections.map((s) => s.body), - ...claims.flatMap((c) => [c.title, c.text, c.findings]), - ...[...quoted].map((id) => view.citations[id]?.quote), - ]), - ), - }); -} - // The report's entry in the index. export function reportIndexEntry(view: ReportPageView): ReportIndexEntry { const tally = view.kind === "factcheck" ? verdictTally(view) : undefined; diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,23 +1,23 @@ # Changelog ## [Unreleased] -- **Exporting a changed report records a new revision of it.** `reports export` (and **Export reports** on a site's Reports tab, and the end of a prepare) now commits a revision to the report's own git history, `sites/<site>/reports/<id>/history-git/`, whenever its `report.json` changed since the last one: the `report.json`, its Markdown export and the checksums of every export file, with a message of `Revision N` and a summary of the change. A re-export of an unchanged report records nothing. The commits carry the site's name and a `noreply@<site>.invalid` address with dates in UTC, never your git name, email or time zone. The Reports tab shows each report's revision, its commit and the last change under **Exports**, and the site's next build publishes the history. Add `history-git/` to the corpus repository's `.gitignore`. -- **archive.org files come over BitTorrent when possible, else straight from archive.org — never through yt-dlp.** An archive.org import fetched the record with yt-dlp, whose archive.org extractor fails on some items (an audio item: "opening play-av tag not found"). Now the chosen file is fetched from the item's own torrent (`<identifier>_archive.torrent`, which lists archive.org as a web seed, so other peers take load off archive.org) with aria2c, only that file of the item, and seeded afterwards for 10 minutes or to a ratio of 1, whichever comes first; the log shows "torrent: <file> (n of m pieces, peers p, web seed yes)" and "seeding 10 min…". With no aria2c, a torrent that does not carry the file, or no progress for 5 minutes, it is downloaded directly from `archive.org/download/…` instead (resumable, backing off on 429/503), and the log says "fell back to direct download: <reason>". Every file is checked against archive.org's sha1/md5: a mismatch is downloaded once more directly, a second one fails the record. The record is written from the item's metadata: `metadata.info.json` with the file's page, the canonical id, the duration ffprobe measures and archive.org's playable copies of the file, the `archiveorg.json` provenance as before (a mirror's original title, date and uploader), and `audio.<fmt>` — an audio file already in the channel's format is used as is, anything else goes through the app's audio extraction, a video kept in the saved-video store when the channel keeps sources. An .avi/.mpeg/.flac/.wav original is fetched as archive.org's mp4 or mp3 of it. aria2c runs in its own process group: cancelling the job stops it and everything it started, and it stops itself if the editor exits. New settings block `archiveOrg` (`torrent`, `seedMinutes`, `seedRatio`, `stallMinutes`, `maxPeers`, `maxDownloadKiBps`, `maxUploadKiBps`), `ARIA2C_BIN`, an aria2c row in `archilyzer doctor`, and `aria2` in the runtime Docker images. +- **Exporting a changed report records a new revision of it.** `reports export` (and **Export reports** on a site's Reports tab, and the end of a prepare) commits a revision to the report's own git history, `sites/<site>/reports/<id>/history-git/`, whenever its `report.json` changed since the last one: the `report.json`, its Markdown export and the checksums of every export file, with a message of `Revision N` and a summary of the change. A re-export of an unchanged report records nothing. The commits carry the site's name and a `noreply@<site>.invalid` address with dates in UTC, never your git name, email or time zone. The Reports tab shows each report's revision, its commit and the last change under **Exports**, and the site's next build publishes the history. Add `history-git/` to the corpus repository's `.gitignore`. +- **archive.org files come over BitTorrent when possible, else straight from archive.org — never through yt-dlp.** The chosen file of an archive.org import is fetched from the item's own torrent (`<identifier>_archive.torrent`, which lists archive.org as a web seed, so other peers take load off archive.org) with aria2c, only that file of the item, and seeded afterwards for 10 minutes or to a ratio of 1, whichever comes first; the log shows "torrent: <file> (n of m pieces, peers p, web seed yes)" and "seeding 10 min…". With no aria2c, a torrent that does not carry the file, or no progress for 5 minutes, it is downloaded directly from `archive.org/download/…` instead (resumable, backing off on 429/503), and the log says "fell back to direct download: <reason>". Every file is checked against archive.org's sha1/md5: a mismatch is downloaded once more directly, a second one fails the record. The record is written from the item's metadata: `metadata.info.json` with the file's page, the canonical id, the duration ffprobe measures and archive.org's playable copies of the file, the `archiveorg.json` provenance (a mirror's original title, date and uploader), and `audio.<fmt>` — an audio file already in the channel's format is used as is, anything else goes through the app's audio extraction, a video kept in the saved-video store when the channel keeps sources. An .avi/.mpeg/.flac/.wav original is fetched as archive.org's mp4 or mp3 of it. aria2c runs in its own process group: cancelling the job stops it and everything it started, and it stops itself if the editor exits. New settings block `archiveOrg` (`torrent`, `seedMinutes`, `seedRatio`, `stallMinutes`, `maxPeers`, `maxDownloadKiBps`, `maxUploadKiBps`), `ARIA2C_BIN`, an aria2c row in `archilyzer doctor`, and `aria2` in the runtime Docker images. - **A forum thread can be archived as a posts source.** A XenForo thread URL (`…/threads/<title>.<id>/`; Kiwi Farms is recognised by host) makes a forum-thread channel — platform "xenforo", one channel per thread, each forum post a post — searchable and readable like X and Bluesky posts, in the editor, the export and the MCP (`get_thread` gives a forum post's conversation: the posts it quotes and the posts quoting it). **Fetch posts** reads the thread in a headless browser, newest page first, one page at a time with a 10–20 s pause (the channel key `postPagePauseSeconds` sets it), and stops at already-archived posts; a **Latest N pages** box (`archilyzer posts fetch --pages N`) caps a run, and the next run continues where it stopped. The browser keeps one profile per forum host, so a browser check it clears once (KiwiFlare's proof of work, say) stays cleared; a check that does not clear within a minute, a captcha, a login wall or a refusal stops the run with the reason and keeps its place — never retried at once. **Connect forum session** on the channel page opens that profile in a window on the editor's machine, at the thread, for the operator to clear it or log in. **Import saved pages** (`archilyzer posts import-html <slug> <file-or-dir>…`) reads thread pages saved from a browser ("Save page as", complete or HTML only) through the same parser: new posts are added and a post saved again after an edit is updated. **Capture posts** works on forum posts: a screenshot of the post and its attached files, through the same profile. A forum post keeps its thread title, page, position, author id, last-edit time, quoted posts and its media links; quoted text is marked with "> " lines. -- **A site's reports can be exported as files a reader saves and hosts again.** `archilyzer reports export <site> [--report <id>] [--formats html,pdf,md,zip]`, the `reports-export` job (`POST /api/ops/reports-export`, `pnpm ops reports-export`, and **Export reports** on a site's Reports tab) write each published report, checked as the build checks it, into `.export-index/sites/<site>/report-exports/<report>/`: `report.html`, one self-contained page (its own style, no script, stills and post screenshots inlined and recompressed, clips linked on the site); `report.pdf`, that page printed by headless Chromium, skipped with a note where there is none; `report.md`, plain Markdown with numbered references; and `evidence-pack.zip`, the page with its clips, stills and screenshots as files plus the Markdown and the citations, packed by the system `zip` (a host without it fails that format, naming it). An `export.json` names each file's size and checksum and the checksum of the report.json it was made from; every export ends with the report's date and the start of that checksum. Preparing the evidence media now exports at its end when nothing is missing, on the same queue. The build publishes an export beside the report only when it was made from the report as it is now and is at most 24 MiB — a larger evidence pack stays local — and the Reports tab lists each report's exports, their sizes and which the next build publishes. The 24 MiB limit the source mirror and the evidence clips already kept is now one shared number. +- **A site's reports can be exported as files a reader saves and hosts again.** `archilyzer reports export <site> [--report <id>] [--formats html,pdf,md,zip]`, the `reports-export` job (`POST /api/ops/reports-export`, `pnpm ops reports-export`, and **Export reports** on a site's Reports tab) write each published report, checked as the build checks it, into `.export-index/sites/<site>/report-exports/<report>/`: `report.html`, one self-contained page (its own style, no script, stills and post screenshots inlined and recompressed, clips linked on the site); `report.pdf`, that page printed by headless Chromium, skipped with a note where there is none; `report.md`, plain Markdown with numbered references; and `evidence-pack.zip`, the page with its clips, stills and screenshots as files plus the Markdown and the citations, packed by the system `zip` (a host without it fails that format, naming it). An `export.json` names each file's size and checksum and the checksum of the report.json it was made from; every export ends with the report's date and the start of that checksum. Preparing the evidence media exports at its end when nothing is missing, on the same queue. The build publishes an export beside the report only when it was made from the report as it is now and is at most 24 MiB — a larger evidence pack stays local — and the Reports tab lists each report's exports, their sizes and which the next build publishes. The 24 MiB limit is one number, shared with the source mirror and the evidence clips. - **archive.org items are a source (`platform: "archiveorg"`).** A channel can hold recordings imported from archive.org and transcribe them like any transcribe channel. **Import video** takes an item page (`https://archive.org/details/<identifier>`) when the item holds one media file, or ONE file of a multi-file item (`…/details/<identifier>/<file>`); an item with several media files is refused with the way to choose files. `pnpm ops import-archive-org --json '{"slug":…,"item":…,"files":[…]}'` (or `"match": "<regex>"`, `"dryRun": true`) imports chosen files of one item as one drainable job. A whole item's id is its identifier; a file's is `<identifier>__<slug>-<hash>`, stable and unique per file. Each record keeps an `archiveorg.json` sidecar — the item's title, date, creator and collections, its torrent, and for a mirror of a YouTube upload the original's id, URL, title and upload date read from the info.json uploaded beside it — and its metadata takes the file's own page and title (and a mirror's original title and date), recorded in the metadata history as `archiveorg-provenance`. The video page says "Archived on archive.org: <item> · torrent" and, for a mirror, "Originally on YouTube: <url> (uploaded <date>)". The channel form offers archive.org in both platform lists. -- **Polite to archive.org.** archive.org runs on its own queue (`platform:archiveorg`), one download at a time, with a jittered pause of at least 8 s between files (the channel's or the global `sleepBetweenDownloadsSeconds` when longer). Its metadata API is asked once per item (cached for 6 h), with an identifying User-Agent, at most one request at a time and 2 s apart, honouring `Retry-After` and backing off exponentially on 429/503, stopping after four attempts. yt-dlp's archive.org spawns carry `--sleep-requests 2` and an exponential `--retry-sleep`; nothing downloads in parallel ranges. A file already downloaded is never fetched again, and a bulk import stops on a rate limit or after three failures in a row — re-running it resumes. The "auto" download format on archive.org takes the uploader's original mp4/mkv/webm, or an audio item's MP3. -- **A report video's cue lookup names a site that publishes only its reports.** Pointing a report-to-video manifest at such a site (`corpus.json` `site.scope: "cited"`) used to fail with "channel … is not in corpus.json"; it now says the site publishes no transcripts and to use a full archive or a local corpus. The site form's **Publish** hint says a cited-only site is never listed on the homepage or the hub. -- **A site's build composes its reports, and a site that publishes only its reports ships nothing else.** Every site's compose now writes the reports its `site.json` publishes: each report's page and its citations as `citations.json` and `citations.csv` under `/reports/<id>/`, its cited stills, a page per cited moment with the record, the transcript lines around the span and every report that cites it, and the clips and post captures `archilyzer reports prepare` made for it, only the cited ones. Each quote is checked against the record as it is composed (a span's against its cues within 5 s either side, read from `en-orig` when the `en` track has no cues; a post's against its text) and the score, time and method are written into the citation, replacing any typed by hand. The build stops with the list of every problem before anything is written: an invalid report, a citation of a channel outside the site or of a post the site may not carry, a missing record, still or post, a quote that matches less than 60 % of what the record says, and a citation without prepared media or with media cut for another span (`--allow-missing-media` on `archilyzer compose site` and `build site` lets those two through, without a clip). A site with `publish: "cited"` removes everything corpus-shaped from `export/public` before it writes its reports, and its built `out/` is checked against what a cited site may hold: anything else, a file over 25 MiB or more than 20,000 files fails the build, and every deploy path (the Publish tab, `deploy site`, Build & deploy, Build & deploy all, the container build) refuses it, as it refuses a site set to cited whose last build was a full one. The hub's compose removes a report site's files too. +- **Polite to archive.org.** archive.org runs on its own queue (`platform:archiveorg`), one download at a time, with a jittered pause of at least 8 s between files (the channel's or the global `sleepBetweenDownloadsSeconds` when longer). Its metadata API is asked once per item (cached for 6 h), with an identifying User-Agent, at most one request at a time and 2 s apart, honouring `Retry-After` and backing off exponentially on 429/503, stopping after four attempts. A file already downloaded is never fetched again, and a bulk import stops on a rate limit or after three failures in a row — re-running it resumes. +- **A report video's cue lookup names a site that publishes only its reports.** Pointed at such a site (`corpus.json` `site.scope: "cited"`), a report-to-video manifest's cue lookup says the site publishes no transcripts and to use a full archive or a local corpus. +- **A site's build composes its reports, and a site that publishes only its reports ships nothing else.** Every site's compose writes the reports its `site.json` publishes: each report's page and its citations as `citations.json` and `citations.csv` under `/reports/<id>/`, its cited stills, a page per cited moment with the record, the transcript lines around the span and every report that cites it, and the clips and post captures `archilyzer reports prepare` made for it, only the cited ones. Each quote is checked against the record as it is composed (a span's against its cues within 5 s either side, read from `en-orig` when the `en` track has no cues; a post's against its text) and the score, time and method are written into the citation, replacing any typed by hand. The build stops with the list of every problem before anything is written: an invalid report, a citation of a channel outside the site or of a post the site may not carry, a missing record, still or post, a quote that matches less than 60 % of what the record says, and a citation without prepared media or with media cut for another span (`--allow-missing-media` on `archilyzer compose site` and `build site` lets those two through, without a clip). A site with search off (`search: false`) removes everything corpus-shaped from `export/public` before it writes its reports, and its built `out/` is checked against what a cited site may hold: anything else, a file over 25 MiB or more than 20,000 files fails the build, and every deploy path (the Publish tab, `deploy site`, Build & deploy, Build & deploy all, the container build) refuses it, as it refuses a site with search off whose last build was a full one. The hub's compose removes a report site's files too. - **A site has a Reports tab.** `/sites/<site>/reports` lists every report under the site's `reports/` directory — the published ones in their order, then the drafts — with its kind, dates, sections, claims, citations by kind and, for a fact-check, how many claims carry each verdict. Each report's problems, from the same checker the prepare step and the build use, open under it. A draft with no problems can be published, and a published report moved up or down or unpublished; each writes only the site's `reports` list, applied to the list as it is on disk at that moment, so it never overwrites another change to the site. "Prepare evidence media" queues the `reports-prepare` job, and beside it the tab shows the last prepared media (moments by kind, total size, problems by kind) and links the last prepare job. What the site publishes (full or cited) is shown with a link to Settings, where it is changed. - **A site can publish reports, and with search off it is only its reports.** `site.json` takes `reports`, the ordered ids of its published reports (each a slug; invalid and repeated ids are dropped), and `search` (default on; only `false` is written): a site with search off publishes only its reports and the moments they cite. The site form has a **Search** checkbox and lists the site's reports read-only (saving the form keeps the stored list), and the Reports tab says which the site is. SITE.md documents both keys. -- **A report and its citations now have one written format, checked before anything is built from them.** A cited report is a `report.json` (`archilyzer-report`, version 1): a summary, then sections of claims, each claim with an optional verdict, the reviewed document's own sentence, findings in markdown and the citations it rests on. A citation is one of five kinds — a span of a video, a span of an audio record, a post, a sentence of a source document, or a web page — with a verbatim quote, and is cited from any markdown in the report as `[label](cite:<id>)`. The checker lists every problem at once with where it is: a citation, a source or a `cite:` link that names nothing, a span that ends before it starts or runs past 120 seconds with its context, a still that points outside the report's folder, an id used twice. Each cited span and post has one page address, `/m/<channel>/<id>/<start>-<end>/` or `/m/<channel>/<id>/`. Nothing builds or shows reports yet. The fact-check verdicts (Corroborated, Partly true, Contradicted, Not found, Untestable) and their colours are now kept in one place, which the report video's stamps and tally read too. `REPORT.md` and `CITATIONS.md` list every key. -- **A site's reports can have their evidence media prepared: every cited span cut to a clip, every cited post's capture copied.** `archilyzer reports prepare <site>`, the `reports-prepare` job (`POST /api/ops/reports-prepare`, `pnpm ops reports-prepare`) reads the site's published reports, checks them, and for each cited moment cuts the span (with its context) out of the media already on disk — a fetched clip window, the saved video, or the recording's audio — fitted inside 1280×720 with H.264 and AAC, or as an `.m4a` for an audio span; and copies the screenshot and attached media of each cited post, and of no other post, beside them. Everything lands in the site's build staging (`.export-index/sites/<site>/report-media/`) with an `index.json` naming each moment's file, size, checksum, size in pixels and duration. A clip is cut once and reused while its source file and span are unchanged; a clip or capture no longer cited is removed. Nothing is fetched: a citation whose media is not on disk, a post without a screenshot, a clip over 24 MiB, a citation of a channel outside the site, a post the site may not show, or an invalid or missing report is listed with the citations it affects, and the run fails (exit 1, or a failed job) — a span on a drive that is not mounted is reported as such rather than as missing. The lookup of a span's media on disk is now shared with report-to-video, which finds the same files it did. +- **A report and its citations have one written format, checked before anything is built from them.** A cited report is a `report.json` (`archilyzer-report`, version 1): a summary, then sections of claims, each claim with an optional verdict, the reviewed document's own sentence, findings in markdown and the citations it rests on. A citation is one of five kinds — a span of a video, a span of an audio record, a post, a sentence of a source document, or a web page — with a verbatim quote, and is cited from any markdown in the report as `[label](cite:<id>)`. The checker lists every problem at once with where it is: a citation, a source or a `cite:` link that names nothing, a span that ends before it starts or runs past 120 seconds with its context, a still that points outside the report's folder, an id used twice. Each cited span and post has one page address, `/m/<channel>/<id>/<start>-<end>/` or `/m/<channel>/<id>/`. The fact-check verdicts (Corroborated, Partly true, Contradicted, Not found, Untestable) and their colours are kept in one place, which the report video's stamps and tally read too. `REPORT.md` and `CITATIONS.md` list every key. +- **A site's reports can have their evidence media prepared: every cited span cut to a clip, every cited post's capture copied.** `archilyzer reports prepare <site>`, the `reports-prepare` job (`POST /api/ops/reports-prepare`, `pnpm ops reports-prepare`) reads the site's published reports, checks them, and for each cited moment cuts the span (with its context) out of the media already on disk — a fetched clip window, the saved video, or the recording's audio — fitted inside 1280×720 with H.264 and AAC, or as an `.m4a` for an audio span; and copies the screenshot and attached media of each cited post, and of no other post, beside them. Everything lands in the site's build staging (`.export-index/sites/<site>/report-media/`) with an `index.json` naming each moment's file, size, checksum, size in pixels and duration. A clip is cut once and reused while its source file and span are unchanged; a clip or capture no longer cited is removed. Nothing is fetched: a citation whose media is not on disk, a post without a screenshot, a clip over 24 MiB, a citation of a channel outside the site, a post the site may not show, or an invalid or missing report is listed with the citations it affects, and the run fails (exit 1, or a failed job) — a span on a drive that is not mounted is reported as such rather than as missing. The lookup of a span's media on disk is the one report-to-video uses. - **A report can be made from a /sweep report, an /ask answer or a report video's manifest, and a report video's manifest from a report.** `archilyzer reports convert sweep|ask|manifest <in> --out <report.json>` writes a `report.json`. From a /sweep report (markdown): its first `#` heading is the title, the text before the first section the summary, each `##` section a section, and each list item or paragraph with a citing link a claim — a line that is only a citation under a quote cites that quote — with every archive moment link, archive post link, moment page link and X or Bluesky post link turned into a citation and the link into `[label](cite:<id>)`; a citation's quote is the quoted words nearest before its link. From an /ask answer (`{ "answer", "sources" }`, a saved chat or an assistant message, as JSON): each `[n]` or `[n @ mm:ss]` marker becomes a citation of source `n` at the excerpt line nearest the marked time, quoting that line, or of the post. From a manifest: chapter cards are the sections, entries carrying `claim` the claims with their verdicts (the first still a claim's source sentence), clips the video citations, stills the source citations and `posts` the post citations; ledger rows are claims too. A /sweep or /ask citation names one second: with `--channels-dir <transcripts/channels>` its span is the cited cue run on to the quote's length and widened to whole sentences, as `resolve-windows.mjs` widens a clip (the widening now lives in common and is shared), and a post cited by its X or Bluesky link is found in the channel that archives it; without it the span is the second plus 10 seconds and such a post stays a plain link. `archilyzer reports to-manifest <report.json> --out <manifest.json>` writes a starter manifest: a title card, a chapter card per section, the section's cited clips, and per claim its source sentence's still and its clips, each with the claim's verdict, and its posts attached to its last clip (a post's platform, link and date come from the channels tree, or for an X post from its id). Both check what they write against the report format and write nothing when it has problems (exit 1); everything a conversion had to leave out or guess is printed as a warning. A manifest image may carry `quote`, the words the still shows, which the converters keep. - **A long report video no longer runs out of memory while its clips are crossfaded.** `build-video.mjs` used to join every segment of a cut in one ffmpeg command, which grows with the number of segments: a cut of a few hundred clips could use more memory than the machine had and be stopped. Past 24 segments the build now crossfades them in batches of consecutive segments, each into a file under `out/<variant>/xfade-batches/`, then crossfades those files together with the same transition and lays the on-screen deck, the rail and the dips over them. Every transition, the deck's schedule and the chapters land on the same frames as before, at the cost of one more video encode on such a cut. The batch size is `render.xfadeBatch` in the manifest or `REPORT_VIDEO_XFADE_BATCH` in the environment (which wins); `0` never batches. A batch file is reused while its segments, their holds and moves and the encode settings are unchanged, so a `--chrome-only` run that moves no footage redoes only the last pass. - **Auto-download no longer tries a video the metadata scan already found members-only or private.** The scan records why it could not read a video, but only a failed download used to take a video out of the auto-download queue, so each members-only video the scan had found was still downloaded once: four yt-dlp requests, two with browser cookies, about 30 seconds each. A members-only or private answer from the scan now keeps the video out of the queue and counts it under the channel's members-only or private exclusions, and it stays in **Needs cookies** for a manual cookie run. A scan that reads the video later lifts this. A video the scan saw only as "Video unavailable" is still tried, since YouTube gives that answer when it is throttling too. -- **X post fetches stop on Drain, and an account with no posts is not searched.** Draining a `fetch-posts` job used to do nothing until gallery-dl finished its whole run. Now the timeline fetch stops at the next page boundary (at once when gallery-dl is between pages or waiting out a rate limit), the older-posts walk stops its current window's search at once and never starts the 45–120 second pause between windows, and both keep their resume point: the job ends done, not failed, and the log says "Drained; the next run resumes …". An older-posts walk is refused when nothing is archived and the last timeline fetch finished having read no posts, since it would only repeat empty searches; `"force": true` (`--force` on `archilyzer posts fetch`) walks anyway. A walk with nothing archived that finds nothing ends after two empty three-month windows instead of four, and records why; a walk that has posts keeps the year-of-empty-windows rule. Capture-posts already stopped between posts on Drain. -- **Capturing an X post that is an Article also saves the article.** An X Article (a long-form post) is archived as nothing but its link, and gallery-dl cannot read its body. When `pnpm ops capture-posts` meets a post whose archived text, or whose card on the page, links to an article, it now opens the article in the same X profile, after the same 4–10 second pause, and saves beside the post's capture: `article.json` (the title, author, date and every heading, paragraph, quote, list item, image, link and embedded post in reading order, an embedded post by its URL), `article.md` (the same as readable text), `article.png` (the whole article as shown, cut off at 16,000 pixels tall and marked `trimmed` when longer), `article.html` (the article as X served it, so it can be read again without going back to X) and the article's pictures as `article-img-1.jpg`, `article-img-2.png`, … at full size, fetched through the same browser session. `capture.json` records the article's state, title, block count and every file's size and SHA-256. `"articles": false` leaves articles alone; `"shots": false, "media": false, "articles": true` reads only the articles. An article already captured, deleted or unavailable is not opened again unless `"force": true`; one that failed is tried on the next run. If X asks to log in on the article page, the job stops there as it does for a post. The post viewer's capture panel in the editor shows the article's title with a link to `article.md`. +- **X post fetches stop on Drain, and an account with no posts is not searched.** Draining a `fetch-posts` job used to do nothing until gallery-dl finished its whole run. Now the timeline fetch stops at the next page boundary (at once when gallery-dl is between pages or waiting out a rate limit), the older-posts walk stops its current window's search at once and never starts the 45–120 second pause between windows, and both keep their resume point: the job ends done, not failed, and the log says "Drained; the next run resumes …". An older-posts walk is refused when nothing is archived and the last timeline fetch finished having read no posts, since it would only repeat empty searches; `"force": true` (`--force` on `archilyzer posts fetch`) walks anyway. A walk with nothing archived that finds nothing ends after two empty three-month windows, and records why; a walk that has posts keeps the year-of-empty-windows rule. Capture-posts already stopped between posts on Drain. +- **Capturing an X post that is an Article also saves the article.** An X Article (a long-form post) is archived as nothing but its link, and gallery-dl cannot read its body. When `pnpm ops capture-posts` meets a post whose archived text, or whose card on the page, links to an article, it opens the article in the same X profile, after the same 4–10 second pause, and saves beside the post's capture: `article.json` (the title, author, date and every heading, paragraph, quote, list item, image, link and embedded post in reading order, an embedded post by its URL), `article.md` (the same as readable text), `article.png` (the whole article as shown, cut off at 16,000 pixels tall and marked `trimmed` when longer), `article.html` (the article as X served it, so it can be read again without going back to X) and the article's pictures as `article-img-1.jpg`, `article-img-2.png`, … at full size, fetched through the same browser session. `capture.json` records the article's state, title, block count and every file's size and SHA-256. `"articles": false` leaves articles alone; `"shots": false, "media": false, "articles": true` reads only the articles. An article already captured, deleted or unavailable is not opened again unless `"force": true`; one that failed is tried on the next run. If X asks to log in on the article page, the job stops there as it does for a post. The post viewer's capture panel in the editor shows the article's title with a link to `article.md`. - **A report video can play a clip that has only sound, and a clip can be a file beside the manifest.** When a clip's source has no picture, `build-video.mjs` plays it under a poster: a card with the clip's channel, title and date, the size of the picture area, with the sound's waveform moving along its foot (`render.audioPoster.waveform: false` keeps it still). The segment matches every other one in size, frame rate and sound, and the header, footer and on-screen deck are drawn over it as over footage. A video's saved sound (`audio.mp3` and the like in its folder) is now a source the build can cut from, after every saved picture: before any download when the clip has no picture to fetch (`"audioOnly": true` on the clip, `"preferLocalAudio": true` in `render`, a podcast or feed record, or a record with no page). A clip that should have a picture is not quietly played from its sound: with `--no-network` or `--skip-fetch`, one whose picture is missing still stops the build, listed as needing a download with a note that its sound is on disk, so `--no-network` still proves every picture is there. Add `--audio-fallback` to play such clips from their sound under the poster instead; they are listed apart (and logged as `audio-fallback`), and clips that have no picture to fetch are listed as playing from audio only rather than refused. A clip may also give `"src"` (a video or audio file) and `"cues"` (its transcript, either a `transcript.cues.json` or a `parakeet-stitch` transcript), both relative to the manifest, instead of a channel and video: it plays the whole file unless `start`/`end` cut inside it, `resolve-windows.mjs` widens it with those cues, it gets no QR unless it has a `citeUrl`, and a path that leaves the manifest's folder (or an absolute one, without `"allowAbsoluteSrc": true` in `render`), a missing file or an unreadable transcript stops the build before anything runs, naming the clip. - **A report build cuts from media already on disk before it downloads anything, and `--no-network` makes sure it never does.** For each clip, `build-video.mjs` now looks, in order, in the project's own `out/clips-raw`, in the clip windows the editor fetched into the channel (`channels/<slug>/data/<id>/clips/`), and in a saved whole source video (through the saved-video store's pointer, or a `source-media` file still in the video's folder), and cuts from the first that holds the clip plus its fetch pad; only when none does is the window downloaded. A file that is a link to a drive that is not mounted counts as not there, and the next place is tried. The build prints one line per clip naming where its source came from (`raw-cache`, `corpus-window`, `saved-video`, or a network fetch). With `--no-network`, every clip's source is found before anything is rendered, and if any clip would need a download the build stops at once and lists each one (its position in the timeline, channel, video and the span it needs). umtool's clip bench reads the same three places, so a clip it shows as fetched is one the build cuts from without downloading. - **umtool's report videos can show a highlighted sentence from a saved article.** `node umtool/report-to-video/shoot-page.mjs --page <saved page.html> --quote "<sentence>" --out <shot.png>` opens a web page saved to disk, finds the sentence in its text, highlights it and saves a PNG of the paragraph that holds it, ready to be a report manifest's `image` entry. `--batch <items.json> --out <dir>` does a list of `{ id, page, quote, context? }` at once and writes `<id>.png` for each plus a `results.json` recording each shot's crop, the matched text and the block it shot. The page is opened offline: nothing is fetched except files saved beside it, and its own scripts do not run unless `--js` is given. The sentence is found whether its quotes and apostrophes are curly or straight, across links and emphasis, and through non-breaking spaces, soft hyphens and line breaks in the page's source. A sentence that is not on the page is listed in `results.json` and on the terminal, and the run ends with an error rather than leaving it out. `--color` sets the highlight; `context` picks one occurrence of a sentence that appears more than once. On a page where a whole post is one block of paragraphs separated by line breaks, `--crop mark` (or an item's `"crop": "mark"`) shoots only the sentence's own lines and one whole line above and below (`--context-lines` sets how many) instead of the whole post; `results.json` records which crop each shot used. @@ -25,8 +25,8 @@ - **Persist a list of videos, across channels, to the saved-video store.** `pnpm ops persist-videos --json '{"items":[{"slug":"<channel>","id":"<video id>"}, …]}'` (or `--file list.json` for a long list) re-fetches the source container of each video that is not saved yet, the way **Persist source video** does on a video's page. `"format": "original"` or `"video_720"` picks the quality (default: each channel's own), and `"replace": "above-height"` also re-fetches a saved video whose recorded height is unknown or above that quality — the old file is removed only after the new one is saved, and the saved video keeps its retention class. Downloads run one at a time, one job per channel on that channel's download queue, with the usual gap between them (`"gapMs"` overrides it); `"minFreeMemMb"` holds each download until that much memory is free. A low disk or a rate limit stops the job, and running the same list again picks up where it left off: saved videos are skipped. `"dryRun": true` answers with what a run would do — saved, saved above the height, to fetch, no source URL, unknown — and starts nothing. Jobs show as **Persist videos**, can be drained, and can be retried from /jobs. - **Capture specific X posts: a screenshot of each, and its attached media.** `pnpm ops capture-posts --json '{"slug":"<channel>","ids":["<post id>", …]}'` shoots each post as X shows it, through the connected X profile, and downloads its pictures and videos with gallery-dl, into the channel's `posts-media/<post id>/` beside a `capture.json` that records when, from which URLs, and each file's size and SHA-256. Every id must already be in the channel's posts archive; one that is not is refused by name and nothing runs. `"shots": false` or `"media": false` skips that half, and posts already captured are skipped unless `"force": true`. The job runs on the X queue with a post fetch, so the two never run at once, and waits a random 4–10 seconds before each request to X, as fetches do. A deleted post, or one behind its account's wall (protected, suspended, gone), is recorded as such in the channel's deleted-post record; a post behind a sensitive-media warning is opened and shot. If X asks to log in, or answers "Something went wrong", the job stops at that post and leaves the rest for a later run. Captures are never published: the export does not read them. - **A source video can be saved at 720p for clip and editing work.** **Persist source video** on a video's page has a **Quality** select: **Original** (the best video and audio, as every persist has been) or **Video 720p (H.264, for clips/editing)**, which saves an H.264 mp4 at most 720p tall — smaller, and quick to cut. When a source has nothing at or under 720p in H.264 it takes 480p, and when it has neither it takes whatever is best and says so in the job's log, with the height it got. The default is the new **Source video quality** under **Settings**, which a channel can override in its Advanced settings; the whole-recording fetch (`full: true`) and **Persist kept now** follow the channel's, else the global, choice. A persisted video's **Source video** card now shows the format that was saved (height and codec) and the quality asked for; videos persisted before this show nothing new. "Video 720p" is also offered as a download format. -- **A report cut can be a fact-check: each claim gets a verdict stamp over the footage, and a tally in the on-screen deck counts them.** Any clip, still or card entry of a report manifest can say which claim it is evidence for and the verdict, `"claim": { "id": "k3", "verdict": "CONTRADICTED" }` (one of CORROBORATED, PARTLY, CONTRADICTED, NOT_FOUND, UNTESTABLE). At the end of the last entry carrying each claim, the verdict slams in over the picture as a stamp in its colour and leaves with the transition, and a row of counts in the deck, one per verdict the cut uses, steps up as each stamp lands. `render.chrome.factcheck` sets each verdict's label and colour, how long the stamp is up (1–10 seconds, 3 by default) and which corner of the footage it sits in, and whether the tally is drawn beside the QR or before the title. A claim is a new key, separate from the clip review's `verdict`; one claim id given two verdicts, a claim on a teaser, or an unknown setting is refused with a sentence before a build fetches anything. umtool's On-screen table has a **claim** column (an id and a verdict per row), saved with the titles, and its preview shows the stamps and the tally before a build. The deck's QR can link each clip's original instead of the archive with `render.chrome.deck.qr.links: "original"` — a YouTube video at the clip's second, a Rumble page or an X post as they are; a clip's `citeUrl` still wins. A manifest post can carry `shot`, a screenshot (PNG, JPEG or WebP, beside the manifest), which its card draws in place of the post's text, its QR kept. A cut with none of these builds exactly as before. -- **X posts are fetched more slowly, with random gaps.** Every read of X now waits a random 4 to 10 seconds before each request to X, where it used to page as fast as X answered, and always waits out a rate limit rather than pushing through. When fetching older posts, the pause between one three-month window and the next is a random 45 to 120 seconds instead of a fixed 15. A deep walk of an account's history takes longer; a routine fetch of new posts takes a few seconds more. +- **A report cut can be a fact-check: each claim gets a verdict stamp over the footage, and a tally in the on-screen deck counts them.** Any clip, still or card entry of a report manifest can say which claim it is evidence for and the verdict, `"claim": { "id": "k3", "verdict": "CONTRADICTED" }` (one of CORROBORATED, PARTLY, CONTRADICTED, NOT_FOUND, UNTESTABLE). At the end of the last entry carrying each claim, the verdict slams in over the picture as a stamp in its colour and leaves with the transition, and a row of counts in the deck, one per verdict the cut uses, steps up as each stamp lands. `render.chrome.factcheck` sets each verdict's label and colour, how long the stamp is up (1–10 seconds, 3 by default) and which corner of the footage it sits in, and whether the tally is drawn beside the QR or before the title. A claim is a new key, separate from the clip review's `verdict`; one claim id given two verdicts, a claim on a teaser, or an unknown setting is refused with a sentence before a build fetches anything. umtool's On-screen table has a **claim** column (an id and a verdict per row), saved with the titles, and its preview shows the stamps and the tally before a build. The deck's QR can link each clip's original instead of the archive with `render.chrome.deck.qr.links: "original"` — a YouTube video at the clip's second, a Rumble page or an X post as they are; a clip's `citeUrl` still wins. A manifest post can carry `shot`, a screenshot (PNG, JPEG or WebP, beside the manifest), which its card draws in place of the post's text, its QR kept. None of these changes a cut that does not set it. +- **X posts are fetched more slowly, with random gaps.** Every read of X now waits a random 4 to 10 seconds before each request to X, where it used to page as fast as X answered, and always waits out a rate limit rather than pushing through. When fetching older posts, the pause between one three-month window and the next is a random 45 to 120 seconds. A deep walk of an account's history takes longer; a routine fetch of new posts takes a few seconds more. - **The MCP's search tools take `date_from` and `date_to` as `2024-10-26` as well as `20241026`, and refuse a date they cannot read.** `search_transcripts` and `enumerate_matches` used to accept only `YYYYMMDD`: any other spelling was dropped with a footer warning and the search ran with no date bound, so a whole-corpus count could be read as the bounded one. Dashed, slashed and dotted dates and ISO timestamps are now normalised, and anything else is an error and nothing is searched. - **An X channel can fetch posts older than its timeline reaches.** X's timeline only pages back so far, so a fetch could end, and call the history done, well short of an account's first post. The new **Fetch older posts** button on an X channel's page (or `pnpm ops fetch-posts --json '{"slug":"<channel>","older":true}'`) walks back from the oldest archived post through X search, three months at a time, and saves posts the same way a normal fetch does; posts already archived are skipped. It needs a login, as search does: without one it stops at once and the channel shows **Needs credentials**. A run saves its place as it goes and stops after three hours; the next run continues from there. The walk ends at the account's creation date, after a year of windows with no posts, or at a date you give as `"floor": "YYYY-MM-DD"`, and the page's **Older posts** line then says it is complete; running it again says so and fetches nothing. A normal **Fetch posts** is unaffected and still fetches new posts from the top. Bluesky channels have no such button: their fetch already reads the whole history. - **`pnpm ops transcribe-bucket` transcribes a channel's downloaded-but-untranscribed videos**, as the channel page's **Transcribe N downloaded** button does, on the transcription queue. `"ids"` runs only some of them; each must be in the bucket, and a stray id is refused by name. `retry-bucket` is not the way to do this: it retries downloads, and counts a video whose audio is on disk as complete. @@ -39,13 +39,13 @@ - **umtool's report videos keep every clip's sound on its picture.** In a crossfaded cut each clip's audio was placed by the audio's own length and its picture by the picture's, and an encoded clip's audio is routinely a few to twenty milliseconds shorter or longer than its video, so the sound drifted further ahead clip by clip: by the end of a seventeen-clip cut it was a third of a second early, and two seconds on one with title and sources cards. Each clip's sound is now padded or trimmed to exactly its picture's length before the crossfade. Every crossfaded report video changes when it is rebuilt, and is in sync; a hard-cut video was not affected. - **umtool's report videos can wear an on-screen deck: one panel under the footage for the whole cut, with a pip timeline, a title per clip, its source and date, and its QR.** A report manifest whose `render` says `"chrome": { "engine": "hyperframes", "layout": "deck" }` scales the footage into a box above a 190 px panel (both sizes are settings) and draws, over the whole cut, one unlabelled pip per clip on a track that fills as the cut plays, the clip's own title from `onscreen.title`, a subtitle naming the recording and its date (the channel too when the cut spans more than one; `onscreen.subtitle` replaces it), and the clip's QR. At each clip change the marker travels to the next pip and the title, subtitle and QR hand over; over a card the panel slides away and comes back after. The citation header, the corner QR and the section footer are not drawn on such a cut, and chapters take the clip's on-screen title. Every setting (sizes, spacing, date format, what the subtitle names, whether cards keep the panel, the motion's timings) is in `render.chrome.deck` and checked when it is saved; an unknown or out-of-range one is refused with a sentence saying why. The panel is rendered once per cut by HyperFrames (pinned to 0.8.24; `HYPERFRAMES_PKG` or `HYPERFRAMES_BIN` override it) and reused until its text or settings change. `build-video.mjs --chrome-only` redraws it over the built segments without rebuilding or fetching anything, `--no-chrome` builds the framed cut without it, and `--chrome-preview <at> <dur>` renders a short window. In umtool, the report page has an **On-screen** section — a switch, the settings, a table of every entry's title and subtitle with the automatic subtitle as its placeholder and a character counter, a live preview with a scrubber, a true still, **Re-render on-screen** and the built video — and the clip bench has on-screen title and subtitle fields with the panel previewed over the clip. The deck changes nothing, byte for byte, in a cut whose manifest has no `render.chrome`. - **A report cut that wears the on-screen deck can show posts — Bluesky or X statements — as cards over the footage.** A report manifest's `posts` list (each with its platform, handle, date, words and link) is drawn near the end of the clip each post belongs with: the clip whose recording most closely precedes it by date, unless the post names one with `attachTo`; `hide` leaves one out. A clip's posts appear four seconds apart and stack down a column at the frame's top right; as the first appears, the footage eases aside (to 86 % of its box, at the far side) to make room, and the clip's last frame is held, in silence, for 2.5 seconds so the last post can be read; then they all leave together in the change to the next clip, which comes in at the normal size. When the column is full the oldest slide up and out. Each card slides in from the edge of the frame and flares in the deck's accent as it lands; it has an accent rail down its edge and shows the post's date, a platform label ("Bluesky" or "X") beside `@handle`, its words in paragraphs up to seven lines with an ellipsis, and a QR of the post's link, in the deck's colours and faces. The hold and the move are made where the cut is joined, not in a clip, so `--chrome-only` changes them without rebuilding one; chapters and the deck's timing count the hold. The timing, the hold (`hold`, 0 turns it off), the move (`shift`: its scale and seconds, or `false`), the column's side, width and inset, the QR size and the line limit are settings under `render.chrome.deck.posts`, and a bad post or setting is refused with a sentence before a build fetches anything. Only the seconds the cards are up are rendered, one short sequence per clip, cached like the deck; `--chrome-only`, `--chrome-preview` and a hard-cut cut lay them as they lay the deck, and `--no-chrome` draws neither — though it still holds and moves the footage, which are part of the cut rather than the chrome. A first post that appears inside the hold still moves the footage, and a hold is a whole number of frames. `posts` changes nothing in a cut that has none, and without the deck it is not drawn at all. -- **A report cut's posts can be a feed: a column beside the footage for the whole cut, each post ticking in as its clip starts, with no pause.** `render.chrome.deck.posts.layout: "feed"` (the default, `"popup"`, is the cards described above) puts every post in one column on the right, standing on the deck so the two read as one L-shaped panel around the picture. The footage of every clip and still is framed, for the whole cut, into the box left beside the column (1272×716 at 1920×1080 with the default 600 px column, against 1574×886 under the deck alone); nothing moves and nothing is held, so the cut is as long as its clips. Before the first post the column is just its header — the platform and the handle, or "Posts" when there are several authors — over its ground. Each post ticks in at the start of the clip it belongs with, a crossfade's length in (just after the crossfade into it, and as far into the first clip, which has none), a clip's next ones `posts.seconds` apart (closer on a short clip): it lands at the top with an accent flare and keeps a lit rail while it is the newest, and the posts already up slide down to make room; when the column is full the oldest fade out at the bottom. Each card shows the post's date, its words up to `maxLines`, and its QR. Over a card or the teaser the column slides out of the frame with the deck and comes back after. The column is drawn by one composition for the whole cut, cached like the deck's. Switching the layout reframes every segment, so it takes a normal build (with `--skip-fetch` it re-cuts from the cached windows); `--chrome-only` over segments framed for the other layout is refused with a sentence naming them, because each segment's `<id>.cut.json` now records the box it was framed into. In umtool, the posts settings have a **layout** switch, and the live preview shows the column for the whole scrub with the footage in its box. A cut in the popup layout, or without posts, builds exactly as before. -- **A report clip can go silent partway through, a report cut can fade out at its end, and the deck's QR names its site in larger type.** A clip's `muteFrom` (in the recording's own seconds, inside the clip) silences it from that second to its end while the picture plays on, after a 40 ms fade that ends there, so nothing clicks and no next word leaks in; a hold on that clip stays silent. `render.endFade` (seconds; 0, the default, is off) fades the cut's last segment, whatever it is — a clip with its hold, a closing card or a teaser — to the background colour and to silence over its final seconds, all of it when the segment is shorter, and the deck stays drawn over it. Both are applied where the cut is joined, so `--chrome-only` changes them without rebuilding a clip, and a value out of range is refused with a sentence before a build fetches anything. Each clip build now writes `<id>.cut.json` beside its segment, saying where in the recording the segment really starts after its cut was snapped to a silence; `muteFrom` is measured from it, and a segment built before this measures from the clip's unsnapped start and says so. The site's name beside the deck's QR is now exactly as long as the code is tall, for any site. Neither key changes a cut that does not set it. +- **A report cut's posts can be a feed: a column beside the footage for the whole cut, each post ticking in as its clip starts, with no pause.** `render.chrome.deck.posts.layout: "feed"` (the default, `"popup"`, is the cards described above) puts every post in one column on the right, standing on the deck so the two read as one L-shaped panel around the picture. The footage of every clip and still is framed, for the whole cut, into the box left beside the column (1272×716 at 1920×1080 with the default 600 px column, against 1574×886 under the deck alone); nothing moves and nothing is held, so the cut is as long as its clips. Before the first post the column is just its header — the platform and the handle, or "Posts" when there are several authors — over its ground. Each post ticks in at the start of the clip it belongs with, a crossfade's length in (just after the crossfade into it, and as far into the first clip, which has none), a clip's next ones `posts.seconds` apart (closer on a short clip): it lands at the top with an accent flare and keeps a lit rail while it is the newest, and the posts already up slide down to make room; when the column is full the oldest fade out at the bottom. Each card shows the post's date, its words up to `maxLines`, and its QR. Over a card or the teaser the column slides out of the frame with the deck and comes back after. The column is drawn by one composition for the whole cut, cached like the deck's. Switching the layout reframes every segment, so it takes a normal build (with `--skip-fetch` it re-cuts from the cached windows); `--chrome-only` over segments framed for the other layout is refused with a sentence naming them, because each segment's `<id>.cut.json` records the box it was framed into. In umtool, the posts settings have a **layout** switch, and the live preview shows the column for the whole scrub with the footage in its box. The feed changes nothing in a cut in the popup layout or without posts. +- **A report clip can go silent partway through, a report cut can fade out at its end, and the deck's QR names its site in larger type.** A clip's `muteFrom` (in the recording's own seconds, inside the clip) silences it from that second to its end while the picture plays on, after a 40 ms fade that ends there, so nothing clicks and no next word leaks in; a hold on that clip stays silent. `render.endFade` (seconds; 0, the default, is off) fades the cut's last segment, whatever it is — a clip with its hold, a closing card or a teaser — to the background colour and to silence over its final seconds, all of it when the segment is shorter, and the deck stays drawn over it. Both are applied where the cut is joined, so `--chrome-only` changes them without rebuilding a clip, and a value out of range is refused with a sentence before a build fetches anything. Each clip build writes `<id>.cut.json` beside its segment, saying where in the recording the segment really starts after its cut was snapped to a silence; `muteFrom` is measured from it, and a segment built before this measures from the clip's unsnapped start and says so. The site's name beside the deck's QR is exactly as long as the code is tall, for any site. Neither key changes a cut that does not set it. - **umtool's clip bench stops exactly where a range ends, and sets a clip's mute mark.** The bench's **play selection**, the edge auditions, the auto-audition and a click on a transcript line now play the window's sound through the browser's Web Audio, from a decode made on the server by ffmpeg — the same timeline the build cuts on — and each stops on the audio clock where its range ends, at every speed. They used to play on the video element and were stopped when it next reported its time, which overran the end by up to a quarter of a second, by a different amount each time. The picture follows, muted. If the sound cannot be decoded, the video element plays as before and the bench says the playback is approximate and why. The mute mark sets the clip's `muteFrom`: `m` puts it at the playhead, **pick on waveform** puts it where you click, `;` and `'` nudge it (with shift, by half a second), and `M` or **clear mute** removes it. It is saved with the window like the edges, every playback goes silent at it with the build's own 40 ms fade, and a window save that would leave it outside the clip is refused unless the same save moves or clears it — or, when it is within 0.02 s of the new edge, moves it onto that edge. The decoded sound is served by a new `GET /api/report/audio`, at most 120 seconds of a cached window at a time, as WAV. - **A report cut can end on a teaser card: a few lines popping in over a dark cinematic ground, with a trailer hit under each.** A report manifest's `teaser` entry (`lines`, `seconds`, an optional `tail`) is a full-frame card drawn from its own words, one to five lines each popping in top to bottom with a scale overshoot, a blur that sharpens, and a flash of the accent; with three or more lines the first is a small overline, the last a mid-size date, and the ones between a big title. A line written as `{ "text": …, "break": … }` draws its ending as a smaller second tier a beat later, and the tail fades in on its own off the right of the last line, which stays centred by its own words; `"tailWait"` (0.3–6 seconds) sets how long after the last hit it comes. Under each pop is a synthesised boom, the title's the biggest, and under the tail a low swell; `"hits": false` makes the card silent. `"beat"` (0.4–2.5 seconds, default 0.7) sets the time from one pop — and its hit — to the next, the second tier and the tail's wait slowing with it; `seconds` may be left out for exactly the length the beats need, and a `seconds` too short for them is refused with that length rather than played faster. Put after the last clip, it joins with the ordinary crossfade and takes the cut's end fade. A line too long to fit the frame at its smallest size is refused with a sentence saying how many characters fit (a title holds 34). It is rendered once and re-rendered when its words change, `--chrome-only` included, and its chapter is its lines. umtool shows it as a card row named by its lines; its words are edited in the manifest. -- **A report cut can go to black before its teaser, and the teaser rises out of the black.** A teaser entry's `"dip": { "fade": …, "black": … }` fades the whole frame before it — the footage, the on-screen deck, the posts feed and anything else drawn over the cut — to black over the previous segment's last `fade` seconds (0.3–4), its sound to silence with it, then holds `black` seconds (0–3) of black, taken to the nearest whole frame at the cut's frame rate. The teaser then opens out of it: the letterbox is already closed, the ground and its light stay dark until the first line slams in and come up with its hit, and a synthesised riser swells under the black into that first hit. The deck and the feed leave under the black instead of sliding away over the crossfade. The black is the start of the teaser's own segment, so the teaser is that much longer and nothing else moves; with a dip, the teaser's `seconds` counts from where the light comes up. The fade is made where the cut is joined, so `--chrome-only` changes it without rebuilding a clip. A dip anywhere but on a teaser, on the first entry, or out of range is refused with a sentence before a build fetches anything. A cut without a dip builds exactly as before. +- **A report cut can go to black before its teaser, and the teaser rises out of the black.** A teaser entry's `"dip": { "fade": …, "black": … }` fades the whole frame before it — the footage, the on-screen deck, the posts feed and anything else drawn over the cut — to black over the previous segment's last `fade` seconds (0.3–4), its sound to silence with it, then holds `black` seconds (0–3) of black, taken to the nearest whole frame at the cut's frame rate. The teaser then opens out of it: the letterbox is already closed, the ground and its light stay dark until the first line slams in and come up with its hit, and a synthesised riser swells under the black into that first hit. The deck and the feed leave under the black instead of sliding away over the crossfade. The black is the start of the teaser's own segment, so the teaser is that much longer and nothing else moves; with a dip, the teaser's `seconds` counts from where the light comes up. The fade is made where the cut is joined, so `--chrome-only` changes it without rebuilding a clip. A dip anywhere but on a teaser, on the first entry, or out of range is refused with a sentence before a build fetches anything. A dip changes nothing in a cut without one. - **A report cut with `transition: 0` builds when its output folder was given as a relative path.** The hard-cut concat listed its segments relative to the working directory, and ffmpeg reads that list relative to the list file's own folder, so every hard-cut build with a relative `--out` failed at the concat. The list now names each segment by its full path. -- **A post's QR in a report video opens the post's page on the archive, not on Bluesky or X.** A post drawn by a report cut's on-screen deck (the popup cards or the feed) used to carry a QR of its bsky.app or x.com link; it now opens the post on the archive the report was made from (`provenance.siteOrigin`), the site's post view at `?v=<channel>%2F<id>&vm=post`, which shows the post and links on to the original, so the QR keeps working if the post or the platform goes away. The build finds the archive channel that keeps each post (a post channel such as `piratesoftware-bsky` is its own channel on the archive) from the archive's `corpus.json` and its posts manifests, cached on disk like the cue lookups, and writes the link into `schedule.json`, so `--chrome-only` and umtool's On-screen preview show the same QR; the manifest is not rewritten. Each post gets a line in the build's log saying which channel it was found in, or that the archive does not have it, or did not answer, and that its QR then links the original — never a failed build. A post can pin its channel with `siteChannel`, or a page of its own with `siteUrl` (and give `postId` when its link does not carry its id), and `render.chrome.deck.posts.links: "original"` (the **qr links** setting in umtool) keeps the old QRs. Rebuilding a cut with posts, or `--chrome-only`, re-renders its posts with the new QRs. The archive MCP's `get_post` also names the post's archive page, as `- archive: <url>`, when the source is a site. +- **A post's QR in a report video opens the post's page on the archive, not on Bluesky or X.** A post drawn by a report cut's on-screen deck (the popup cards or the feed) carries a QR that opens the post on the archive the report was made from (`provenance.siteOrigin`), the site's post view at `?v=<channel>%2F<id>&vm=post`, which shows the post and links on to the original, so the QR keeps working if the post or the platform goes away. The build finds the archive channel that keeps each post (a post channel such as `piratesoftware-bsky` is its own channel on the archive) from the archive's `corpus.json` and its posts manifests, cached on disk like the cue lookups, and writes the link into `schedule.json`, so `--chrome-only` and umtool's On-screen preview show the same QR; the manifest is not rewritten. Each post gets a line in the build's log saying which channel it was found in, or that the archive does not have it, or did not answer, and that its QR then links the original — never a failed build. A post can pin its channel with `siteChannel`, or a page of its own with `siteUrl` (and give `postId` when its link does not carry its id), and `render.chrome.deck.posts.links: "original"` (the **qr links** setting in umtool) links the bsky.app or x.com post instead. The archive MCP's `get_post` also names the post's archive page, as `- archive: <url>`, when the source is a site. - **`archilyzer doctor` checks the image Build all builds sites in.** When a container engine answers, a new **build image** section says whether the image named under **Settings → Build pipeline** is there, when it was built and how big it is. It warns when the image is missing, or older than the last change to its Dockerfile, and prints the one command that rebuilds it. Build all still builds or refreshes the image itself before it builds any site; the warning tells you ahead of time that the next Build all will spend that time. With no container engine the check is skipped in one line, and with no corpus it is only a note. It never fails the doctor. - **The site build image runs Node 22 and pnpm 11**, the versions the rest of the workspace runs on, instead of Node 20 and pnpm 9, which did not read the workspace's install rules. The next Build all rebuilds the image from its first step, reinstalling every dependency, before it builds any site. - **Build all sites works in containers again.** Every site's container build had been failing while it prerendered `/favicon.ico`. Each site now builds from the data composed for it, never from files baked into the build image. A bundle whose `site.json` and `corpus.json` do not both name its site is refused before it is handed back or deployed. The image carries no corpus data, and its build context is about 7 MB from any checkout. @@ -58,7 +58,7 @@ - **A form whose save is refused keeps what you typed.** Every editor form put its plain fields back to the stored values when its save was refused — a site's ID rejected, a page size out of range, a slug already taken — so everything typed had to be typed again. A refused save now leaves every field as you left it, beside the reason: **Settings**; a site's form (new and existing); the hub's config on `/sites`; **Cut release**; a channel's form (new and **Configure**), **Rename** and **Delete**; a video's **Delete directory**; **Drive health timing** on `/storage`; the backup config on `/saved-videos`; the sync operation's controls; the **Digest**, **Diarization**, **Speaker attribution** and **Speaker work lane** settings; and the worker list on `/workers`. A save that succeeds behaves as before, with one difference you may notice: a drop-down, and a checkbox or choice that the page tracks as you change it (a cadence, a worker's **Enabled**, a social link's **Keep in header**, a site membership, a site's accent), now shows what was saved. A form's own drop-downs used to go back to what the page had loaded with until a reload, and a second save from the same page sent that old choice again; the others went back until the page next refreshed itself (every 5 seconds by default). - **A media move no longer starts over a job that is writing into the channel, holds the channel's writers while it runs, and makes its copy match the source before it verifies — so a transcription or a download during a move cannot fail it.** A move that has waited its turn behind other moves now checks again when it starts: if a job is running on the channel, or an auto-queue lane is working on one of its videos, it stops at once and says which ("a transcription of abc123 is running (Transcribe all, job …) — wait for it or cancel it"), with nothing copied — and a job you have just cancelled counts until it has actually stopped ("is stopping … — wait for it to stop"); **Preview** says the same, and the Storage panel's blocked message now names the job too. While a move's marker stands, the channel is held: every lane skips it, and every job that reads or writes its media (single-video transcriptions, downloads and transcodes and the availability checks now included) refuses to start, including one that was already queued when the move began. The rack shows a **media held** chip in the channel's Tier cell and the Storage panel says "Held: its media is moving"; both go when the move finishes or its marker is cleared. The copy is now followed by a pass that makes the destination copy match the source — files the source no longer has are removed from the copy, never from the source — so a file written or deleted during the copy (a transcriber's scratch folder, say) no longer fails the check, and **Resume move** finishes a move whose copy holds such leftovers. Every file removed from a copy is listed in the move's log, and **Preview** says so when a copy from an earlier attempt is already there. If the source keeps changing, the move stops and lists what differs: extra on the destination, missing there, or changed. A new **Reconcile and resume** button beside **Resume move** lists those differences, makes the copy match and finishes the move, so no file has to be deleted by hand. The saved-video store's move does the same matching and the same check before it starts. Needs a rebuild and restart of the editor. - **Connecting an X account opens your own browser, and the X fetchers can use your everyday browser's X login instead.** **Settings → X account session → Connect X account** used to open Playwright's bundled Chromium with its automation signals on (the "controlled by automated test software" bar, `navigator.webdriver`): Google's sign-in refused it and X's own login form stalled in it. It now opens your Chromium or Chrome when one is installed (`ARCHILYZER_X_BROWSER` names another; Playwright's bundled Chromium otherwise), without those signals. Google's sign-in may still refuse an embedded browser; X's password login is the reliable path. A new **Login source** choice (`social.x.cookieSource` in `settings.json`) says where the X fetchers' login comes from: **Browser login** hands gallery-dl `--cookies-from-browser` with your `cookiesFromBrowser` on every fetch, so the login lasts as long as you stay logged in to x.com in that browser and no window is needed; **Connected profile** is the session broker, as before. Left on **Automatic**, it is the browser login when `cookiesFromBrowser` is set and no profile is connected, and the profile otherwise. **Check** says which source is in use, whether an X login is visible in it and when it was last used (the browser's cookies are read from a private copy, never written; this reads Firefox's, and gallery-dl reads Chromium's itself). Needs a rebuild and restart of the editor. -- **X posts can be kept off every public site.** **Settings → X account session** has a new choice, **Where X posts appear** (`social.x.visibility` in `settings.json`): **Public**, the default, builds an X channel's posts into every site that has the channel, as before; **Private** leaves every X channel out of every public site's build — its posts, its posts manifest entry and its place in the channel list, `site.json` and `corpus.json` — and builds it only into private sites (below). Nothing on disk changes and fetching goes on. A site already published changes on its next build and deploy, and a public site whose only posts were X posts loses its **Posts** box under **Search in**. Choosing the X login source no longer forgets this choice, and choosing this one keeps the login source. Needs a rebuild and restart of the editor, then a rebuild and deploy of every site and the hub. +- **X posts can be kept off every public site.** **Settings → X account session** has a new choice, **Where X posts appear** (`social.x.visibility` in `settings.json`): **Public**, the default, builds an X channel's posts into every site that has the channel, as before; **Private** leaves every X channel out of every public site's build — its posts, its posts manifest entry and its place in the channel list, `site.json` and `corpus.json` — and builds it only into private sites (below). Nothing on disk changes and fetching goes on. A site already published changes on its next build and deploy, and a public site whose only posts were X posts loses its **Posts** box under **Search in**. Choosing the X login source keeps this choice, and choosing this one keeps the login source. Needs a rebuild and restart of the editor, then a rebuild and deploy of every site and the hub. - **A site can be private: built for reading on this machine, never deployed and never listed.** A site's settings have a new **Audience** choice (`audience` in `site.json`; only `"private"` is written). A private site is refused by every deploy — **Build & deploy**, **Deploy**, `archilyzer deploy site`, `pnpm ops build-deploy` and `deploy-site`, **Build & deploy all** (which builds it and skips its deploy) and `docker/publish-site.sh` — before anything is uploaded, in a sentence naming the audience; **Build** still builds it. It is left off the homepage, the hub and every other site's footer whatever **List on the Archilyzer homepage and hub** says, it publishes no hub URL, and its `corpus.json` says `"audience": "private"`, so a build of it is refused too if the site is switched back to public before it is rebuilt. Point the MCP at a private site's build to ask about what only it holds. Needs a rebuild and restart of the editor. - **A site built on the host right after another site no longer ships that site's channel list.** A site's **Build** (and Build all without containers) composes every site into the same folder, and the step that copies a site's summaries, stats and duplicates was skipped when that site's own data had not changed, even though another site had composed there since. The site then went out with the other site's channel list and stats. Those steps now run again whenever another site composed last. Needs a rebuild and restart of the editor. - **The hub no longer ships the data of the last site built before it.** **Build hub** built from the folder a site's build had just filled, so the hub carried that site's summaries, transcripts, posts and other data, and served them. The hub's build now clears every site's data first and uses the global search aliases, and **Deploy hub** refuses a hub build that still carries a site's data. Needs a rebuild and restart of the editor, then a rebuild and deploy of the hub. diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md @@ -1,18 +1,18 @@ # Changelog ## [Unreleased] -- **A report shows its revision, and every edit to it can be checked.** A report's date line now ends with "revision N", linking to its history ("edited since revision N" when the report has changed since). The history page, `/reports/<id>/history/`, lists every revision, newest first: its number, date (UTC), commit hash and the sha256 of its `report.json`, what changed (claims added or removed, verdicts changed, claims edited, citations added or removed, quotes edited, title, series or subtitle changed), and each changed claim's title, text, verdict and findings with the words removed struck through and the words added marked. The same data is in `history.json` beside the page. Each report's history is its own git repository, published for cloning: `git clone <site>/reports/<id>/history/repo`. Each commit names the site as its author, with its date in UTC. The footer of the report's HTML, PDF and Markdown downloads now begins with the revision number, and its sha256 can be looked up on the history page. Needs `reports export` and a rebuild and deploy of each site with reports. -- **A report can be saved whole: as one HTML page, a PDF, Markdown, or an evidence pack.** A report page's download line now reads HTML · PDF · Markdown · Evidence pack · Citations JSON · CSV, each listed only when the site publishes it. The HTML is one file that opens with no network: the report with its verdicts, the document's sentences and the post screenshots inside it, numbered citations, and a reference list giving each quote's speaker, date, record, the original at its time and the moment page on the site. The PDF is that page printed. The Markdown is the same report as plain text with numbered references. The evidence pack is a zip of the page with its clips, stills and screenshots beside it, so the clips play offline. Each ends with a line naming the report's date and the start of its checksum. Needs `reports export` (or prepare) and a rebuild and deploy of each site with reports. -- **A report's claim can carry a flag, its header names the document under review, and a site with one report names it in the browser tab.** `report.json` claim `flag` (one line, at most 60 characters) shows as a small pill in the accent colour beside the claim's verdict, e.g. "No source given". On a report-only site with one report, the home page's tab title is the report's, as on the report's own page, where it was the site's title alone. A report's page header is its series (or, with none, its title), then one small line of dates and the revision ("2026-10-04 · updated 2026-10-05 · revision 1"; "updated" only when it differs), then a card for the document under review: its title linking to the document, "<author> · <publisher> · <date>", and its archive links folded away, on a left rail in the document's colour (a source's `accent`, `"#rrggbb"`; without one, the border colour), then the subtitle. The kind's label no longer sits above the title on the report's page, and the page names no byline or site of its own. A claim's sentence from the document under review no longer links up to the document's card: it sits on the same rail. A sentence of another document keeps its "from <title>" link. A report's citation can say where its evidence came from (`origin`: `"subject"`, the document under review gave it; `"added"`, the report's author found it). A claim lists what the report added first, each card marked with the Archilyzer mark and "Not in the article" ("Not in the source"), then evidence of unknown origin, then what the document gave itself folded under "In the article (n)"; the reference list marks an added citation with the mark alone, and a claim's flag pill wears the same mark. A fact-check's page reads in three tiers, each opened by a hairline with one, two or three dots and its reading time (at 230 words a minute): the quick take (the tally, the summary, and links to what the check found, every claim and the downloads); **What the check found**, every ruled claim grouped by verdict (contradicted, not found, partly, untestable, corroborated), one line each linking to the claim, with its `gist` (a new optional claim field, one line, at most 240 characters) and its flag; and **Every claim, with its evidence**, which opens with **How it was checked** (`method`, a new optional report field in markdown). A report of kind `sweep` has the first and last tiers only. Needs a rebuild and deploy of the site. -- **Forum posts read like the other posts.** A post from a forum-thread channel shows its place in the thread (#N), an "edited" mark, the thread's title and its media as links, and opening its thread shows its conversation — the posts it quotes and the posts quoting it — rather than the whole forum thread. +- **A report shows its revision, and every edit to it can be checked.** A report's date line ends with "revision N", linking to its history ("edited since revision N" when the report has changed since). The history page, `/reports/<id>/history/`, lists every revision, newest first: its number, date (UTC), commit hash and the sha256 of its `report.json`, what changed (claims added or removed, verdicts changed, claims edited, citations added or removed, quotes edited, title, series or subtitle changed), and each changed claim's title, text, verdict and findings with the words removed struck through and the words added marked. The same data is in `history.json` beside the page. Each report's history is its own git repository, published for cloning: `git clone <site>/reports/<id>/history/repo`. Each commit names the site as its author, with its date in UTC. The footer of the report's HTML, PDF and Markdown downloads begins with the revision number, and its sha256 can be looked up on the history page. Needs `reports export` and a rebuild and deploy of each site with reports. +- **A report can be saved whole: as one HTML page, a PDF, Markdown, or an evidence pack.** A report page's download line reads HTML · PDF · Markdown · Evidence pack · Citations JSON · CSV, each listed only when the site publishes it. The HTML is one file that opens with no network: the report with its verdicts, the document's sentences and the post screenshots inside it, numbered citations, and a reference list giving each quote's speaker, date, record, the original at its time and the moment page on the site. The PDF is that page printed. The Markdown is the same report as plain text with numbered references. The evidence pack is a zip of the page with its clips, stills and screenshots beside it, so the clips play offline. Each ends with a line naming the report's revision, its date and the start of its checksum. Needs `reports export` (or prepare) and a rebuild and deploy of each site with reports. +- **A report's claim can carry a flag, its header names the document under review, and a site with one report names it in the browser tab.** `report.json` claim `flag` (one line, at most 60 characters) shows as a small pill in the accent colour beside the claim's verdict, e.g. "No source given". On a report-only site with one report, the home page's tab title is the report's, as on the report's own page. A report's page header is its name — with a `series`, the series on one line in the accent colour and the title on the line below; without one, the title — then one small line of dates and the revision ("2026-10-04 · updated 2026-10-05 · revision 1"; "updated" only when it differs), then a card for the document under review: its title linking to the document, "<author> · <publisher> · <date>", and its archive links folded away, on a left rail in the document's colour (a source's `accent`, `"#rrggbb"`; without one, the border colour), then the subtitle. The page names no byline or site of its own. A claim that cites the document's own sentence shows that sentence — its still, else its words — on the same rail, with no link up to the card and no paraphrase beside it; a sentence of another document links "from <title>" to that document's box; a titled claim with no such sentence shows its text under the title, plain. A report's citation can say where its evidence came from (`origin`: `"subject"`, the document under review gave it; `"added"`, the report's author found it). A claim lists what the report added first, each card marked with the Archilyzer mark under its number ("Not in the article", or "Not in the source", is the mark's tooltip and what a screen reader says), then evidence of unknown origin, then what the document gave itself folded under "In the article (n)"; the reference list marks an added citation with the mark too, and a claim's flag pill wears the same mark. A fact-check's page reads in three tiers, each opened by a hairline with one, two or three dots: the quick take (the tally, the summary, and links to what the check found, every claim and the downloads); **What the check found**, every ruled claim grouped by verdict (contradicted, not found, partly, untestable, corroborated), one line each linking to the claim, with its `gist` (a new optional claim field, one line, at most 240 characters) and its flag; and **Every claim, with its evidence**, which opens with **How it was checked** (`method`, a new optional report field in markdown). A report of kind `sweep` has the first and last tiers only. Needs a rebuild and deploy of the site. +- **Forum posts read like the other posts.** A post from a forum-thread channel shows its place in the thread (#N), an "edited" mark, the thread's title and its media as links, and opening its thread shows its conversation: the posts it quotes and the posts quoting it. - **archive.org records play and are cited with their downloads.** A record imported from archive.org plays its file in the page's own player, which seeks to a cited second and follows the transcript. A citation of one links "archive.org" (the file's page) and its "torrent"; a citation of an archive.org mirror of a YouTube upload links the original on YouTube at the cited second, then "archive.org" and "torrent", on the citation cards and the moment pages. -- **A report-only site with one report opens on that report.** Its home page is the report itself, its header links nothing, and `/reports/` forwards home: there is no index of one. With more reports the home page is the list, without repeating the site's title under the header; a list entry is the report's name, subtitle and dates (its counts and tally are on its page). Pages a report-only site does not have link home. -- **A report can belong to a series.** `report.json` `series` heads the report's page in place of its title (the header's card names the document), and in the report list is shown on its own line above the title, in the accent colour, in place of the kind's label ("Fact-check"); a page title, a cited-in link, `llms.txt` and the MCP name it `<series>: <title>`. +- **A report-only site with one report opens on that report.** Its home page is the report itself, its header links nothing, and `/reports/` forwards home: there is no index of one. With more reports the home page is the list, with no heading of its own; a list entry is the report's name, subtitle and dates (its counts and tally are on its page). Pages a report-only site does not have link home. +- **A report can belong to a series.** `report.json` `series` is shown on its own line above the report's title, in the accent colour, at the head of the report's page and its downloads and in the report list, where a report without one shows its kind ("Fact-check"); a page title, a cited-in link, `llms.txt` and the MCP name it `<series>: <title>`. - **`pnpm start:export` serves a built site's moment pages.** It used `serve`, which listed a video or audio moment's directory (`3126.00-3151.00`) instead of serving its page; it now runs `export/scripts/serve-out.mjs`, which serves directories as Cloudflare Pages does, on `EXPORT_DEV_PORT` (3000). -- **MCP: a site's reports can be read, and a site that publishes only reports says so instead of looking empty.** Two new tools: `list_reports` lists the reports a site publishes (id, title, kind, claim and citation counts, a fact-check's verdict tally, its page), and `get_report` reads one — its tally, then each section's claims with their verdicts and findings and, for every citation, the verbatim quote, the original (the platform at the cited second, the post, the document) and the site's moment page; `section` reads one section. On a site that publishes only its reports, `list_channels`, `list_sources` and `resolve_source` say "cited-only site: N report(s)" where they said "No channels found"; on a site with reports as well, `list_sources` and `resolve_source` say how many. The archive readers (local, remote) read `corpus.json` spec 5: a cited site is an empty corpus without an error, and a local copy of one is never read from a stale `transcripts/` folder. A hub has no reports of its own; `list_reports` says to name a member site. -- **The hub leaves out a site that publishes only its reports.** A site with search off (`search: false`) is not a searchable archive, so it is no hub member (not in federated search, the hub's `corpus.json` or `llms.txt`) and no other site's footer links to it, whatever its **List on the Archilyzer homepage and hub** setting says. A site with reports that publishes its full corpus is a member as before. Needs a hub rebuild and deploy once such a site exists. +- **MCP: a site's reports can be read, and a site that publishes only reports says so.** Two new tools: `list_reports` lists the reports a site publishes (id, title, kind, claim and citation counts, a fact-check's verdict tally, its page), and `get_report` reads one — its tally, then each section's claims with their verdicts and findings and, for every citation, the verbatim quote, the original (the platform at the cited second, the post, the document) and the site's moment page; `section` reads one section. On a site that publishes only its reports, `list_channels`, `list_sources` and `resolve_source` say "cited-only site: N report(s)"; on a site with reports as well, `list_sources` and `resolve_source` say how many. The archive readers (local, remote) read `corpus.json` spec 5: a cited site is an empty corpus without an error, and a local copy of one is never read from a stale `transcripts/` folder. A hub has no reports of its own; `list_reports` says to name a member site. +- **The hub leaves out a site that publishes only its reports.** A site with search off (`search: false`) is not a searchable archive, so it is no hub member (not in federated search, the hub's `corpus.json` or `llms.txt`) and no other site's footer links to it, whatever its **List on the Archilyzer homepage and hub** setting says. A site with reports that publishes its full corpus is a member. Needs a hub rebuild and deploy once such a site exists. - **`corpus.json` is spec 5: it names a site's reports, and a site that publishes only reports says so.** A site with reports adds `reports` to its `corpus.json` (`index`: `/reports/index.json`, the count, and how to read a report's page, its citations and its moment pages) and a Reports section to `llms.txt`; its sitemap lists the report and moment pages. A site that publishes only its reports has `"scope": "cited"` and its audience under `site`, no channels and zero totals, an `llms.txt` that lists its reports and how their citations and moment pages are read, and a `site.json` with no channels. A reader that does not know spec 5 sees an empty corpus there. Needs a rebuild and deploy of each site. -- **A site can show cited reports, and every citation opens on a page of its own.** A site built with reports has a **Reports** link in its header and a page at `/reports/` listing them. A report's page has its title, subtitle, dates and the document under review, its archive links listed once under it and folded away ("N archive links"); a fact-check's tally of verdicts; the summary; the sections and their claims, each with its verdict, the document's own sentence as an image with a link back to the document and only the archive links that sit in that sentence (at most five), the findings and the evidence cards; a numbered reference list; and links to download its citations as JSON and CSV. A citation in the text shows as its words plus a number: hovering it, focusing the number or tapping it once shows a card of the citation (the quote, who said it and when, a picture or the post's screenshot, and how closely the quote matched the transcript when it was checked); the words open what it cites and the number jumps to its reference. A cited span of a video or audio record opens at `/m/<channel>/<id>/<start>-<end>/`: a short clip of the span with a little context either side, the quote, the transcript lines around it, the record's title, channel and date, a link to the original at that time, and every report on the site that cites it. A cited post opens at `/m/<channel>/<id>/` with its screenshot and text. A site that publishes only its reports (`site.json` `publish: "cited"`) opens on the report index and has no search, Ask AI, downloads or duplicates. A site with no reports is unchanged. Needs a rebuild and deploy of each site. +- **A site can show cited reports, and every citation opens on a page of its own.** A site built with reports has a **Reports** link in its header and a page at `/reports/` listing them. A report's page has its name, dates and the document under review, with the document's archive links listed once and folded away; a fact-check's tally of verdicts; the summary; the sections and their claims, each with its verdict, the document's own sentence as an image with only the archive links that sit in that sentence (at most five), the findings and the evidence cards; a numbered reference list; and its downloads. A citation in the text shows as its words plus a number: hovering it, focusing the number or tapping it once shows a card of the citation (the quote, who said it and when, a picture or the post's screenshot, and how closely the quote matched the transcript when it was checked); the words open what it cites and the number jumps to its reference. A cited span of a video or audio record opens at `/m/<channel>/<id>/<start>-<end>/`: a short clip of the span with a little context either side, the quote, the transcript lines around it, the record's title, channel and date, a link to the original at that time, and every report on the site that cites it. A cited post opens at `/m/<channel>/<id>/` with its screenshot and text. A site that publishes only its reports (`site.json` `search: false`) has no search, Ask AI, transcript downloads or duplicates. A site with no reports is unchanged. Needs a rebuild and deploy of each site. ## [0.11.1] - 2026-10-01 - **Use with AI goes to the Archilyzer site's AI and MCP doc; the page on each site is gone.** The header's, the slide-out menu's, the footer's and Ask AI's **Use with AI** keep their label and open https://archilyzer.pages.dev/docs/ai-and-mcp/ in the same tab, on every site and the hub, where one block says how to run Claude Code against any archive (the source, `pnpm install`, `claude mcp add archilyzer`, `/ask`). `/use-with-ai/` is no longer built. `corpus.json`'s `useWithAi` names the doc; `llms.txt`'s Ask AI section lists the site's `/ask/` chat and the doc; the sitemap drops `/use-with-ai`. Needs a rebuild and deploy of each site and the hub. diff --git a/export/app/components/reports/ReportArticle.tsx b/export/app/components/reports/ReportArticle.tsx @@ -10,7 +10,6 @@ import { orderedCitations, reportDateParts, reportRevisionLabel, - reportTierMinutes, sourceAnchor, verdictTally, type CitationView, @@ -18,17 +17,20 @@ import { type ReportPageView, type SourceCitationView, } from "yt-dlp-transcript-common/lib/report/views"; -import { ArchiveList, SourceBlock, SubjectCard, textLink } from "./parts"; +import { ArchiveList, ReportName, SourceBlock, SubjectCard, textLink } from "./parts"; // ONE REPORT, from its view (common/lib/report/views.ts): the header (its -// series, else its title; one line of dates and the revision, linked to its -// history; the document under review as a card on its colour's rail, with its -// archive links; the subtitle), a fact-check's tally, the summary, the sections and their claims — each claim -// its verdict and flag, the document's own sentence (its still), the findings with -// their inline citations, and the evidence cards (what the report added first, -// marked; what the document gave itself folded away last) — then the numbered -// reference list the inline markers jump to, and the downloads: the report as -// files (HTML, PDF, Markdown, the evidence pack) and its citations as data. +// series and its title, or the title alone; one line of dates and the +// revision, linked to its history; the document under review as a card on its +// colour's rail, with its archive links; the subtitle), the three tiers each +// opened by a depth marker — the quick take (a fact-check's tally, the +// summary), what the check found, and every claim with its evidence: each +// claim its verdict and flag, the document's own sentence (its still), the +// findings with their inline citations, and the evidence cards (what the +// report added first, marked; what the document gave itself folded away +// last) — then the numbered reference list the inline markers jump to, and +// the downloads: the report as files (HTML, PDF, Markdown, the evidence pack) +// and its citations as data. const proseClass = "text-[0.95rem] leading-relaxed text-foreground"; @@ -106,14 +108,12 @@ function FlagPill({ flag, className }: { flag: string; className?: string }) { } // Where a tier of the page begins: a hairline with, at its left, how deep it -// goes (one, two or three dots filled in the accent) and how long it takes to -// read. Read aloud as one plain line ("Next: more detail, about 6 minutes"). -function TierMarker({ depth, minutes, spoken }: { depth: 1 | 2 | 3; minutes: number; spoken: string }) { - const time = `${Math.max(1, minutes)} min`; - const said = `${spoken}, about ${Math.max(1, minutes)} minute${Math.max(1, minutes) === 1 ? "" : "s"}`; +// goes (one, two or three dots filled in the accent). Decoration only: the +// tier's own label or heading names it. +function TierMarker({ depth }: { depth: 1 | 2 | 3 }) { return ( - <div data-tier-marker={depth} className="flex items-center gap-2.5"> - <span aria-hidden="true" className="flex items-center gap-1"> + <div data-tier-marker={depth} aria-hidden="true" className="flex items-center gap-2.5"> + <span className="flex items-center gap-1"> {[1, 2, 3].map((i) => ( <span key={i} @@ -126,28 +126,26 @@ function TierMarker({ depth, minutes, spoken }: { depth: 1 | 2 | 3; minutes: num /> ))} </span> - <span aria-hidden="true" className="font-mono text-xs text-muted-foreground"> - {time} - </span> - <span aria-hidden="true" className="h-px flex-1 bg-border" /> - <span className="sr-only">{said}</span> + <span className="h-px flex-1 bg-border" /> </div> ); } +// A claim: its verdict, flag and title; the document's own sentence (its +// still, else its verbatim words) when the claim cites one, else the claim as +// the report states it, plain; the findings; the evidence. function Claim({ claim, view, - subjectLabel, subjectNoun, }: { claim: ClaimView; view: ReportPageView; - subjectLabel: string; // What the document under review is called: "article" or "source". subjectNoun: string; }) { - const sentence = claim.sourceQuote ? view.citations[claim.sourceQuote] : undefined; + const cited = claim.sourceQuote ? view.citations[claim.sourceQuote] : undefined; + const sentence = cited?.kind === "source" ? cited : undefined; const evidence = claim.citations .filter((id) => id !== claim.sourceQuote) .map((id) => view.citations[id]) @@ -174,13 +172,14 @@ function Claim({ </a> </h3> </div> - {claim.title && ( - <p className="text-sm text-muted-foreground"> - {subjectLabel}: <q className="text-foreground">{claim.text}</q> - </p> - )} - {sentence?.kind === "source" && ( + {sentence ? ( <SourceSentence c={sentence} archives={claim.archives} isSubject={sentence.sourceId === view.subject} /> + ) : ( + claim.title && ( + <p data-claim-text="" className="text-sm text-muted-foreground"> + {claim.text} + </p> + ) )} {claim.findings && ( <CitedMarkdown className={proseClass}> @@ -221,7 +220,6 @@ export default function ReportArticle({ view }: { view: ReportPageView }) { const references = orderedCitations(view); const subject = view.subject ? view.sources[view.subject] : undefined; const subjectNoun = subject?.kind === "article" ? "article" : "source"; - const subjectLabel = `The ${subjectNoun} says`; const dates = reportDateParts(view); const claimCount = view.sections.reduce((n, s) => n + s.claims.length, 0); // The report as files (its exports), then its citations as data: only what @@ -231,7 +229,6 @@ export default function ReportArticle({ view }: { view: ReportPageView }) { return href ? [[key, label, href] as const] : []; }); const groups = isFactcheck ? foundGroups(view) : []; - const minutes = reportTierMinutes(view); const jumps = [ ...(groups.length > 0 ? [{ href: "#found", label: "What the check found" }] : []), { href: "#claims", label: "Every claim" }, @@ -246,12 +243,12 @@ export default function ReportArticle({ view }: { view: ReportPageView }) { // serialized once, not once per marker. <CitationsProvider citations={view.citations}> <article data-report={view.id} className="mx-auto flex w-full max-w-3xl flex-col gap-8"> - {/* The report's name (its series, else its title: with a series, the - card below names the document), one line of dates and the + {/* The report's name (its series on one line, its title on the next; + without a series, the title alone), one line of dates and the revision, the document under review, the subtitle. */} <header className="flex flex-col gap-3"> <h1 className="font-display text-3xl font-semibold leading-tight tracking-tight text-foreground [overflow-wrap:anywhere] sm:text-4xl"> - {view.series ?? view.title} + <ReportName series={view.series} title={view.title} /> </h1> {(dates.length > 0 || view.history) && ( <p data-report-dates="" className="font-mono text-xs text-muted-foreground"> @@ -274,7 +271,7 @@ export default function ReportArticle({ view }: { view: ReportPageView }) { {/* The quick take: the tally and the summary, then where to go next. */} <section data-report-tier="1" aria-label="In brief" className="flex flex-col gap-4"> - <TierMarker depth={1} minutes={minutes.quick} spoken="In brief" /> + <TierMarker depth={1} /> {tally.length > 0 && ( <div className="flex flex-col gap-2"> <p className="font-mono text-[10px] uppercase tracking-[0.16em] text-muted-foreground"> @@ -303,7 +300,7 @@ export default function ReportArticle({ view }: { view: ReportPageView }) { {/* What the check found: every ruled claim by verdict, one line each. */} {groups.length > 0 && ( <section id="found" data-report-tier="2" aria-labelledby="found-heading" className="flex scroll-mt-20 flex-col gap-4"> - <TierMarker depth={2} minutes={minutes.found ?? 0} spoken="Next: more detail" /> + <TierMarker depth={2} /> <h2 id="found-heading" className="font-display text-2xl font-semibold tracking-tight text-foreground"> What the check found </h2> @@ -336,7 +333,7 @@ export default function ReportArticle({ view }: { view: ReportPageView }) { {/* Every claim, with its evidence. */} <div id="claims" data-report-tier="3" className="flex scroll-mt-20 flex-col gap-8"> <div className="flex flex-col gap-4"> - <TierMarker depth={3} minutes={minutes.claims} spoken="Next: every claim in full" /> + <TierMarker depth={3} /> <h2 className="font-display text-2xl font-semibold tracking-tight text-foreground"> Every claim, with its evidence </h2> @@ -378,7 +375,7 @@ export default function ReportArticle({ view }: { view: ReportPageView }) { </CitedMarkdown> )} {s.claims.map((c) => ( - <Claim key={c.id} claim={c} view={view} subjectLabel={subjectLabel} subjectNoun={subjectNoun} /> + <Claim key={c.id} claim={c} view={view} subjectNoun={subjectNoun} /> ))} </section> ))} diff --git a/export/app/components/reports/parts.tsx b/export/app/components/reports/parts.tsx @@ -1,3 +1,4 @@ +import { Fragment } from "react"; import Link from "next/link"; import { ExternalLink } from "lucide-react"; import { sourceAnchor, subjectMetaParts, type SourceView } from "yt-dlp-transcript-common/lib/report/views"; @@ -16,11 +17,11 @@ export function Eyebrow({ children }: { children: React.ReactNode }) { return <p className="font-mono text-xs uppercase tracking-[0.18em] text-brand">{children}</p>; } -// A report's name as a heading, in the index and on the history page. With a -// series, the series is its own line in the accent (the tag, like a -// wordmark's lead) and the title the next, in regular weight — both inside -// the heading, so its name reads "<series> <title>"; without one, the title -// alone (the kind's label is then the eyebrow above). +// A report's name as a heading, on its page, in the index and on the history +// page. With a series, the series is its own line in the accent (the tag, +// like a wordmark's lead) and the title the next, in regular weight — both +// inside the heading, so its name reads "<series> <title>"; without one, the +// title alone. export function ReportName({ series, title }: { series?: string; title: string }) { if (!series) return <>{title}</>; return ( @@ -85,7 +86,8 @@ export function SourceBlock({ source, label }: { source: SourceView; label: stri // whose it is. It is the document's one place on the page, so it carries the // anchor a citation's "from <title>" links to. export function SubjectCard({ source }: { source: SourceView }) { - const meta = subjectMetaParts(source).join(" · "); + // Each part stays whole on a narrow screen (a date never splits at its dashes). + const meta = subjectMetaParts(source); const n = source.archives.length; return ( <div @@ -108,7 +110,16 @@ export function SubjectCard({ source }: { source: SourceView }) { source.title )} </p> - {meta && <p className="text-xs text-muted-foreground">{meta}</p>} + {meta.length > 0 && ( + <p className="text-xs text-muted-foreground"> + {meta.map((part, i) => ( + <Fragment key={i}> + {i > 0 && " · "} + <span className="whitespace-nowrap">{part}</span> + </Fragment> + ))} + </p> + )} {n > 0 && ( <details data-source-archives={n} className="text-xs"> <summary className="cursor-pointer text-muted-foreground hover:text-foreground"> diff --git a/export/app/lib/reports.test.ts b/export/app/lib/reports.test.ts @@ -82,10 +82,14 @@ test("params: every report and every moment, as the index files list them", () = test("the report page: header, tally, sections and claims, inline cites, references, downloads", async () => { const params = Promise.resolve({ reportId: "demo-factcheck" }); const html = render(await reportPage.default({ params })); - // the header: the title (no series), one muted line of dates and the - // revision linked to its history, the document under review as a card, the subtitle + // the header: the series on one line and the title on the next, one muted + // line of dates and the revision linked to its history, the document under + // review as a card, the subtitle const header = html.slice(html.indexOf("<header"), html.indexOf("</header>")); - assert.match(header, /<h1[^>]*>Checking an example article<\/h1>/); + assert.match( + header, + /<h1[^>]*><span data-report-series="" class="block font-semibold text-brand">Demo Checks<\/span> <span class="block font-normal">Checking an example article<\/span><\/h1>/, + ); assert.match( header, /<\/h1><p data-report-dates=""[^>]*>2026-10-01 · updated 2026-10-04 · <a href="\/reports\/demo-factcheck\/history\/" data-report-revision="2"[^>]*>revision 2<\/a><\/p><div id="source-s0" data-subject-source="s0"/, @@ -98,7 +102,7 @@ test("the report page: header, tally, sections and claims, inline cites, referen const card = header.slice(header.indexOf('<div id="source-s0"')); assert.match(card, /^<div id="source-s0" data-subject-source="s0" class="[^"]*border-l-4[^"]*" style="border-left-color:#8a6fb0">/); assert.match(card, /<a href="https:\/\/example.org\/articles\/demo" target="_blank" rel="noopener noreferrer"[^>]*>An example article about a demo channel<\/a>/); - assert.match(card, /<p[^>]*>A. Writer · Example Gazette · 2026-09-20<\/p>/); + assert.match(card, /<p[^>]*><span class="whitespace-nowrap">A. Writer<\/span> · <span class="whitespace-nowrap">Example Gazette<\/span> · <span class="whitespace-nowrap">2026-09-20<\/span><\/p>/); assert.match(card, /<details data-source-archives="1"[^>]*><summary[^>]*>1 archive link<\/summary>/); assert.ok(card.includes("as published on the day")); assert.ok(!html.includes("Read in the edition published on the day"), "the subject's note is not in the card"); @@ -112,8 +116,11 @@ test("the report page: header, tally, sections and claims, inline cites, referen assert.match(html, /data-verdict-tally=""/); for (const v of ["CORROBORATED", "PARTLY", "CONTRADICTED", "UNTESTABLE"]) assert.match(html, new RegExp(`data-verdict="${v}"`)); assert.match(html, /<section id="bridge"/); - // a claim with a title: the title heads it, the claim's text follows as what the article says - assert.match(html, /<article id="claim-1"[^]*?<h3[^>]*><a href="#claim-1"[^>]*>“He opened it himself”<\/a><\/h3>[^]*?The article says: <q[^>]*>He opened the bridge himself in 2018.<\/q>/); + // a claim with a title and the article's sentence: the title heads it, the + // sentence's still follows (its words its alt), with no paraphrase beside it + const titled = html.slice(html.indexOf('<article id="claim-1"'), html.indexOf("</article>", html.indexOf('<article id="claim-1"'))); + assert.match(titled, /<h3[^>]*><a href="#claim-1"[^>]*>“He opened it himself”<\/a><\/h3><\/div><figure data-source-sentence="a01"[^>]*><img src="[^"]*a01.png" alt="He opened the bridge himself in 2018."/); + assert.doesNotMatch(titled, /says:|<q[^>]*>He opened the bridge|data-claim-text/); // a claim without one is headed by its text assert.match(html, /<h3[^>]*><a href="#claim-2"[^>]*>He went back to the bridge every year.<\/a>/); assert.match(html, /<article id="claim-1"/); @@ -128,7 +135,7 @@ test("the report page: header, tally, sections and claims, inline cites, referen assert.match(html, /href="\/reports\/demo-factcheck\/citations.json" download=""/); assert.match(html, /href="\/reports\/demo-factcheck\/citations.csv" download=""/); const meta = await reportPage.generateMetadata({ params }); - assert.equal(meta.title, "Checking an example article"); + assert.equal(meta.title, "Demo Checks: Checking an example article"); // a claim's flag: a pill on its card; an unflagged claim carries none assert.match(html, /<article id="claim-4"[^]*?<span data-claim-flag=""[^>]*><svg[^>]*data-brand-mark=""[^]*?<\/svg>No source given<\/span>/); // a claim's evidence: what the report added first, marked; then evidence of @@ -136,7 +143,11 @@ test("the report page: header, tally, sections and claims, inline cites, referen const claim1 = html.slice(html.indexOf('<article id="claim-1"'), html.indexOf("</article>", html.indexOf('<article id="claim-1"'))); const order = [...claim1.matchAll(/data-citation="([^"]+)"/g)].map((m) => m[1]); assert.deepEqual(order, ["c01", "au1", "w01"]); - assert.match(claim1, /<div data-citation="c01"[^>]*><p data-citation-added=""[^>]*><svg[^>]*data-brand-mark=""[^]*?<\/svg>Not in the article<\/p>/); + // the added card: the mark alone, under its number box, the words its tooltip + assert.match( + claim1, + /<div data-citation="c01"[^>]*><div[^>]*><span class="[^"]*flex-col[^"]*"><span[^>]*>1<\/span><span data-citation-added="" title="Not in the article"[^>]*><svg[^>]*data-brand-mark=""[^]*?<\/svg><span class="sr-only">Not in the article<\/span><\/span><\/span>/, + ); assert.equal([...claim1.matchAll(/data-citation-added=/g)].length, 1, "only the added card is marked"); assert.match(claim1, /<details data-subject-evidence=""[^>]*><summary[^>]*>In the article \(1\)<\/summary>[^]*?data-citation="w01"/); assert.doesNotMatch(claim1, /<details[^>]* open/, "folded by default"); @@ -153,9 +164,10 @@ test("the report page: header, tally, sections and claims, inline cites, referen const tier2 = html.slice(html.indexOf('data-report-tier="2"'), html.indexOf('data-report-tier="3"')); const tier3 = html.slice(html.indexOf('data-report-tier="3"')); assert.doesNotMatch(tier1 + tier2, /<details/); - assert.match(tier1, /data-tier-marker="1"[^]*?<span class="sr-only">In brief, about 1 minute<\/span>/); - assert.match(tier2, /data-tier-marker="2"[^]*?<span class="sr-only">Next: more detail, about 1 minute<\/span>/); - assert.match(tier3, /data-tier-marker="3"[^]*?<span class="sr-only">Next: every claim in full, about 1 minute<\/span>/); + assert.match(tier1, /<div data-tier-marker="1" aria-hidden="true"/); + assert.match(tier2, /<div data-tier-marker="2" aria-hidden="true"/); + assert.match(tier3, /<div data-tier-marker="3" aria-hidden="true"/); + assert.doesNotMatch(html, />\s*\d+ min\s*<|\d+ minutes?\b/, "no reading time"); assert.equal([...tier2.matchAll(/data-depth-dot="filled"/g)].length, 2); assert.match(tier1, /data-verdict-tally=""[^]*?The article makes four claims/); const jumps = [...tier1.slice(tier1.indexOf("data-report-jumps")).matchAll(/<a href="([^"]+)"[^>]*>([^<]+)<\/a>/g)].map((m) => `${m[1]} ${m[2]}`); @@ -174,14 +186,14 @@ test("the report page: header, tally, sections and claims, inline cites, referen assert.match(tier3, /<h2[^>]*>Every claim, with its evidence<\/h2><div data-report-method=""[^>]*><h3[^>]*>How it was checked<\/h3>[^]*?searched for in the channel/); }); -test("the report header: a series heads it in place of the title; an edited report says so; no accent, the border's rail", () => { +test("the report header: no series, the title alone; an edited report says so; no accent, the border's rail", () => { const view = fixture.buildFixtureReportView(); const s0 = view.sources.s0; const html = render( React.createElement(ReportArticle, { view: { ...view, - series: "On the Record", + series: undefined, updated: view.published, history: { ...view.history!, current: false }, sources: { ...view.sources, s0: { ...s0, accent: undefined, author: undefined, date: undefined } }, @@ -189,14 +201,26 @@ test("the report header: a series heads it in place of the title; an edited repo }), ); const header = html.slice(html.indexOf("<header"), html.indexOf("</header>")); - // the series is the heading; the title is the card's document, not repeated - assert.match(header, /<h1[^>]*>On the Record<\/h1>/); - assert.ok(!header.includes(view.title), "with a series the report's title is not in the header"); + // no series: the title alone is the heading + assert.match(header, /<h1[^>]*>Checking an example article<\/h1>/); + assert.doesNotMatch(header, /data-report-series/); // no 'updated' when it is the published date; the revision, edited since assert.match(header, /<p data-report-dates=""[^>]*>2026-10-01 · <a [^>]*data-report-revision="2"[^>]*>edited since revision 2<\/a><\/p>/); // without an accent the card's rail is the border colour; only the parts it has assert.match(header, /<div id="source-s0" data-subject-source="s0" class="[^"]*border-l-4[^"]*border-border"(?! style)/); - assert.match(header, /<p[^>]*>Example Gazette<\/p>/); + assert.match(header, /<p[^>]*><span class="whitespace-nowrap">Example Gazette<\/span><\/p>/); +}); + +test("a titled claim with no sentence of the document: its text follows plain, unquoted and unlabelled", () => { + const view = fixture.buildFixtureReportView(); + const sections = view.sections.map((s) => ({ + ...s, + claims: s.claims.map((c) => (c.id === "claim-1" ? { ...c, sourceQuote: undefined } : c)), + })); + const html = render(React.createElement(ReportArticle, { view: { ...view, sections } })); + const claim = html.slice(html.indexOf('<article id="claim-1"'), html.indexOf("</article>", html.indexOf('<article id="claim-1"'))); + assert.match(claim, /<\/h3><\/div><p data-claim-text=""[^>]*>He opened the bridge himself in 2018.<\/p>/); + assert.doesNotMatch(claim, /data-source-sentence|says:|<q[^>]*>He opened/); }); test("a video moment: the clip, quote, verification, cue lines, original link, cited in", async () => { @@ -210,7 +234,7 @@ test("a video moment: the clip, quote, verification, cue lines, original link, c assert.equal([...html.matchAll(/data-in-span=""/g)].length, 2); assert.match(html, /href="https:\/\/media.example.org\/demo-channel\/abc123\?t=3126"[^>]*>Original at 52:06/); assert.match(html, /data-cited-in=""/); - assert.match(html, /href="\/reports\/demo-factcheck\/"[^>]*>Checking an example article/); + assert.match(html, /href="\/reports\/demo-factcheck\/"[^>]*>Demo Checks: Checking an example article/); assert.match(html, /href="\/reports\/demo-factcheck\/#claim-1"/); assert.ok(html.includes("— “He opened it himself”"), "a cited-in entry names the claim by its title"); assert.match(html, /<span>Summary<\/span>/, "a summary citation is named as the summary"); @@ -245,7 +269,7 @@ test("one report: /reports/ is no index, it forwards to the home page (the repor test("a cited site with one report: its home IS the report, and never counts transcripts", async () => { // There is no summaries manifest here: countTranscripts() would throw. const html = render(await home.default()); - assert.match(html, /<h1[^>]*>Checking an example article<\/h1>/); + assert.match(html, /<h1[^>]*><span data-report-series=""[^>]*>Demo Checks<\/span> <span[^>]*>Checking an example article<\/span><\/h1>/); assert.match(html, /data-claim=/); assert.doesNotMatch(html, /data-report-index/); // the site's title is the site header's alone: the report names no attribution @@ -254,7 +278,7 @@ test("a cited site with one report: its home IS the report, and never counts tra // its tab names the report as the report's own page does const reportMeta = await reportPage.generateMetadata({ params: Promise.resolve({ reportId: "demo-factcheck" }) }); assert.deepEqual(home.generateMetadata(), reportMeta); - assert.equal(home.generateMetadata().title, "Checking an example article"); + assert.equal(home.generateMetadata().title, "Demo Checks: Checking an example article"); }); test("a cited site with more than one report: its home is the index, with no heading of its own", async () => { @@ -262,7 +286,8 @@ test("a cited site with more than one report: its home is the index, with no hea const original = fs.readFileSync(file, "utf8"); try { const index = JSON.parse(original); - index.reports = [...index.reports, { ...index.reports[0], id: "demo-second", series: "On the Record", title: "A second report", href: "/reports/demo-second/" }]; + const { series: _series, ...unseried } = index.reports[0]; + index.reports = [unseried, { ...index.reports[0], id: "demo-second", series: "On the Record", title: "A second report", href: "/reports/demo-second/" }]; fs.writeFileSync(file, JSON.stringify(index)); const html = render(await home.default()); assert.match(html, /data-report-index/); @@ -275,12 +300,12 @@ test("a cited site with more than one report: its home is the index, with no hea assert.match(idx, /data-report-entry="demo-factcheck"/); assert.doesNotMatch(idx, /claims ·|data-verdict-tally/, "counts and tally are the report page's"); assert.doesNotMatch(idx, /http-equiv="refresh"/); - // A series is its own line above the title, in the accent, in place of the kind's eyebrow. + // A series is its own line above the title, in the accent, with no kind eyebrow. const second = html.slice(html.indexOf('data-report-entry="demo-second"')); assert.match(second, /<span data-report-series="" class="block font-semibold text-brand">On the Record<\/span> <span class="block font-normal">A second report<\/span>/); assert.doesNotMatch(second.slice(0, second.indexOf("</li>")), /Fact-check/); const first = html.slice(html.indexOf('data-report-entry="demo-factcheck"')); - assert.match(first.slice(0, first.indexOf("</li>")), /Fact-check/, "no series: the kind's eyebrow as before"); + assert.match(first.slice(0, first.indexOf("</li>")), /Fact-check/, "no series: the kind's eyebrow"); } finally { fs.writeFileSync(file, original); } diff --git a/export/e2e-report/report-site.spec.ts b/export/e2e-report/report-site.spec.ts @@ -37,7 +37,9 @@ async function expectImageLoaded(img: Locator): Promise<void> { test("one report: the home page IS the report, and the header links nothing", async ({ page }) => { await page.goto("/"); - await expect(page.locator("h1")).toHaveText("Checking an example article"); + // The series on one line, the title on the next, both in the heading. + await expect(page.locator("h1")).toHaveText("Demo Checks Checking an example article"); + await expect(page.locator("h1 [data-report-series]")).toHaveText("Demo Checks"); // One line of dates and the revision; then the document under review as a // card on its rail: its title linked to it, author · publisher · date, no label. await expect(page.locator("[data-report-dates]")).toHaveText("2026-10-01 · updated 2026-10-04 · revision 2"); @@ -52,7 +54,7 @@ test("one report: the home page IS the report, and the header links nothing", as await expect(page.locator("[data-report-attribution], [data-report-byline]")).toHaveCount(0); await expect(page.locator("article[data-claim]")).toHaveCount(5); // The tab names the report, as the report's own page does. - await expect(page).toHaveTitle("Checking an example article — Demo Reports"); + await expect(page).toHaveTitle("Demo Checks: Checking an example article — Demo Reports"); await expect(page.locator("[data-report-index]")).toHaveCount(0); await expect(page.locator("header nav a")).toHaveCount(0); }); @@ -94,7 +96,7 @@ test("the report: verdict chips, the tally, the source's sentence as a still, nu // A claim's flag is a pill on its card; an unflagged claim has none. await expect(page.locator("article#claim-4 [data-claim-flag]")).toHaveText("No source given"); await expect(page.locator("article[data-claim] [data-claim-flag]")).toHaveCount(1); - await expect(page).toHaveTitle("Checking an example article — Demo Reports"); + await expect(page).toHaveTitle("Demo Checks: Checking an example article — Demo Reports"); // The claim's sentence in the article, as the reviewed page showed it. await expectImageLoaded(page.locator('[data-source-sentence="a01"] img[src$="/stills/a01.png"]')); @@ -117,7 +119,7 @@ test("the report: verdict chips, the tally, the source's sentence as a still, nu } }); -test("three tiers, each marked by its depth and reading time; nothing above the claims is folded away", async ({ page }) => { +test("three tiers, each marked by its depth; nothing above the claims is folded away", async ({ page }) => { await page.goto(REPORT); const tiers = page.locator("[data-report-tier]"); expect(await tiers.evaluateAll((els) => els.map((e) => e.getAttribute("data-report-tier")))).toEqual(["1", "2", "3"]); @@ -125,7 +127,7 @@ test("three tiers, each marked by its depth and reading time; nothing above the const marker = page.locator(`[data-report-tier="${depth}"] [data-tier-marker="${depth}"]`); await expect(marker).toBeVisible(); await expect(marker.locator('[data-depth-dot="filled"]')).toHaveCount(depth); - await expect(marker).toContainText(/\d+ min/); + await expect(marker).toHaveText(""); } await expect(page.locator('[data-report-tier="1"] details, [data-report-tier="2"] details')).toHaveCount(0); @@ -162,10 +164,17 @@ test("a claim's evidence: what the report added first and marked, the article's const cards = claim.locator("[data-citation]"); await expect(cards).toHaveCount(3); expect(await cards.evaluateAll((els) => els.map((e) => e.getAttribute("data-citation")))).toEqual(["c01", "au1", "w01"]); - // Added: the project's mark and the words; nothing else in the claim is marked. + // Added: the project's mark alone, under the card's number box and as wide + // (the words are its tooltip); nothing else in the claim is marked. const added = claim.locator('[data-citation="c01"] [data-citation-added]'); - await expect(added).toHaveText("Not in the article"); - await expect(added.locator("svg[data-brand-mark]")).toBeVisible(); + await expect(added).toHaveAttribute("title", "Not in the article"); + const mark = added.locator("svg[data-brand-mark]"); + await expect(mark).toBeVisible(); + const numberBox = (await added.locator("xpath=preceding-sibling::span[1]").boundingBox())!; + const markBox = (await mark.boundingBox())!; + expect(Math.abs(markBox.width - numberBox.width)).toBeLessThan(1); + expect(Math.abs(markBox.x - numberBox.x)).toBeLessThan(1); + expect(markBox.y).toBeGreaterThanOrEqual(numberBox.y + numberBox.height); await expect(claim.locator("[data-citation-added]")).toHaveCount(1); // The flag pill wears the same mark. await expect(page.locator("article#claim-4 [data-claim-flag] svg[data-brand-mark]")).toBeVisible(); @@ -258,9 +267,9 @@ test("the report downloads as HTML, PDF, Markdown and an evidence pack, its cita const body = await res.body(); if (href.endsWith("citations.json")) expect(JSON.parse(body.toString()).format).toBe("archilyzer-citations"); else if (href.endsWith(".csv")) expect(body.toString()).toContain("c01"); - else if (href.endsWith(".html")) expect(body.toString()).toContain("<title>Checking an example article</title>"); + else if (href.endsWith(".html")) expect(body.toString()).toContain("<title>Demo Checks: Checking an example article</title>"); else if (href.endsWith(".pdf")) expect(body.subarray(0, 5).toString()).toBe("%PDF-"); - else if (href.endsWith(".md")) expect(body.toString()).toContain("# Checking an example article"); + else if (href.endsWith(".md")) expect(body.toString()).toContain("**Demo Checks**\n\n# Checking an example article\n"); else expect(body.subarray(0, 4).toString("hex")).toBe("504b0304"); } }); diff --git a/export/fixtures/report-site/public/m/demo-channel/abc123/3126.00-3151.00/moment.json b/export/fixtures/report-site/public/m/demo-channel/abc123/3126.00-3151.00/moment.json @@ -72,7 +72,7 @@ "claimId": null, "citationId": "c01", "href": "/reports/demo-factcheck/", - "reportTitle": "Checking an example article", + "reportTitle": "Demo Checks: Checking an example article", "number": 1 }, { @@ -81,7 +81,7 @@ "claimId": "claim-1", "citationId": "c01", "href": "/reports/demo-factcheck/#claim-1", - "reportTitle": "Checking an example article", + "reportTitle": "Demo Checks: Checking an example article", "sectionTitle": "The bridge", "claimTitle": "“He opened it himself”", "claimText": "He opened the bridge himself in 2018.", @@ -98,7 +98,7 @@ "claimId": "claim-5", "citationId": "c01", "href": "/reports/demo-factcheck/#claim-5", - "reportTitle": "Checking an example article", + "reportTitle": "Demo Checks: Checking an example article", "sectionTitle": "Moving away", "claimText": "He was at the opening.", "verdict": "CORROBORATED", diff --git a/export/fixtures/report-site/public/m/demo-podcast/ep-042/610.50-628.00/moment.json b/export/fixtures/report-site/public/m/demo-podcast/ep-042/610.50-628.00/moment.json @@ -51,7 +51,7 @@ "claimId": null, "citationId": "au1", "href": "/reports/demo-factcheck/", - "reportTitle": "Checking an example article", + "reportTitle": "Demo Checks: Checking an example article", "number": 2 }, { @@ -60,7 +60,7 @@ "claimId": "claim-1", "citationId": "au1", "href": "/reports/demo-factcheck/#claim-1", - "reportTitle": "Checking an example article", + "reportTitle": "Demo Checks: Checking an example article", "sectionTitle": "The bridge", "claimTitle": "“He opened it himself”", "claimText": "He opened the bridge himself in 2018.", @@ -77,7 +77,7 @@ "claimId": "claim-2", "citationId": "au1", "href": "/reports/demo-factcheck/#claim-2", - "reportTitle": "Checking an example article", + "reportTitle": "Demo Checks: Checking an example article", "sectionTitle": "The bridge", "claimText": "He went back to the bridge every year.", "verdict": "PARTLY", diff --git a/export/fixtures/report-site/public/m/demo-social/1234567890/moment.json b/export/fixtures/report-site/public/m/demo-social/1234567890/moment.json @@ -25,7 +25,7 @@ "claimId": "claim-3", "citationId": "p01", "href": "/reports/demo-factcheck/#claim-3", - "reportTitle": "Checking an example article", + "reportTitle": "Demo Checks: Checking an example article", "sectionTitle": "Moving away", "claimText": "He has said many times that he would move away.", "verdict": "CONTRADICTED", diff --git a/export/fixtures/report-site/public/reports/demo-factcheck/page.json b/export/fixtures/report-site/public/reports/demo-factcheck/page.json @@ -3,6 +3,7 @@ "version": 1, "id": "demo-factcheck", "kind": "factcheck", + "series": "Demo Checks", "title": "Checking an example article", "subtitle": "Four claims about a demo channel, tested against what was said on air", "summary": "The article makes four claims. The recordings support one, partly support another, and contradict a third; the fourth cannot be tested. The host said the opening date out loud [on stream](cite:c01), and later [repeated it](cite:au1).", diff --git a/export/fixtures/report-site/public/reports/index.json b/export/fixtures/report-site/public/reports/index.json @@ -5,6 +5,7 @@ { "id": "demo-factcheck", "kind": "factcheck", + "series": "Demo Checks", "title": "Checking an example article", "subtitle": "Four claims about a demo channel, tested against what was said on air", "published": "2026-10-01", diff --git a/export/fixtures/report-site/source/demo-factcheck/report.json b/export/fixtures/report-site/source/demo-factcheck/report.json @@ -3,6 +3,7 @@ "version": 1, "id": "demo-factcheck", "kind": "factcheck", + "series": "Demo Checks", "title": "Checking an example article", "subtitle": "Four claims about a demo channel, tested against what was said on air", "summary": "The article makes four claims. The recordings support one, partly support another, and contradict a third; the fourth cannot be tested. The host said the opening date out loud [on stream](cite:c01), and later [repeated it](cite:au1).", diff --git a/export/fixtures/report-site/source/demo-factcheck/report.v1.json b/export/fixtures/report-site/source/demo-factcheck/report.v1.json @@ -3,6 +3,7 @@ "version": 1, "id": "demo-factcheck", "kind": "factcheck", + "series": "Demo Checks", "title": "Checking an example article", "subtitle": "Three claims about a demo channel, tested against what was said on air", "summary": "The article makes four claims. The recordings partly support two and contradict a third; the fourth cannot be tested. The host said the opening date out loud [on stream](cite:c01), and later [repeated it](cite:au1).", diff --git a/mcp/README.md b/mcp/README.md @@ -335,7 +335,7 @@ A **cited-only** site (`corpus.json` `site.scope: "cited"`; built from a site wi search off, `site.json` `search: false`) publishes its reports and the moments they cite and nothing else: no channels, no transcripts, nothing to search. `list_channels`, `list_sources` and `resolve_source` say "cited-only -site: N report(s)" there rather than reporting an empty corpus. Reports are per +site: N report(s)" there. Reports are per site: a hub has none of its own, so name a member with `source:"remote:<url>"`. ## How it reads the corpus diff --git a/plans/report-sites.md b/plans/report-sites.md @@ -1,7 +1,7 @@ # Report sites — a site built around cited reports Status: BUILT — R0–R8 merged 2026-10-05 (plus R4b page size); first private report site built locally. RH -(per-report revision history) built 2026-10-05, see "History". 2026-10-05: `publish` replaced by a per-site `search` switch; report-only is derived (search off); export's `start` serves moment dirs (`scripts/serve-out.mjs`). Follow-ups: cited corpus.json still describes shards, voice checks recorded in citations, saveSiteAction onto patchSite. +(per-report revision history) built 2026-10-05, see "History". Report-only is derived from a per-site `search` switch (search off); export's `start` serves moment dirs (`scripts/serve-out.mjs`). Follow-ups: cited corpus.json still describes shards, voice checks recorded in citations, saveSiteAction onto patchSite. ## What it is @@ -10,9 +10,9 @@ Two independent axes on a site: | Axis | Values | Today | |---|---|---| | `reports` | 0..n published reports, each a cited document rendered as pages | none | -| `publish` | `"full"` — the searchable corpus (today); `"cited"` — ONLY the moments the site's reports cite | always full | +| `search` | on — the searchable corpus (today); off — ONLY the moments the site's reports cite | always on | -- A **report site** is `publish: "cited"` + ≥1 report: no search, no browse, no full transcripts, no archives. +- A **report site** is search off (`search: false`) + ≥1 report: no search, no browse, no full transcripts, no archives. - A **full site with reports** keeps everything it has and gains a Reports section; its moment pages also link into the corpus (`/?v=…&t=`). - Every citation in a report opens a **moment page** on the site: the quote, a self-hosted evidence clip of the @@ -80,16 +80,14 @@ Citations are the core; reports are one consumer. | Key | Default | Notes | |---|---|---| -| `publish` | `"full"` | `"cited"` publishes only reports + their moments. Only the non-default is written. | +| `search` | `true` | `false` publishes only reports + their moments (report-only). Only `false` is written. | | `reports` | `[]` | Ordered report ids under `reports/`. | -Ruling 2026-10-05: `publish` is replaced by `search` (boolean, default `true`; only `false` is written). A -site with search off is report-only — what `publish: "cited"` was. A legacy `publish: "cited"` reads as -`search: false` and is rewritten as such on save. The editor form's Publish select is a Search checkbox. -`corpus.json` keeps `site.scope: "cited"`. +The editor form has a Search checkbox. A `publish: "cited"` read from a site.json reads as `search: false` and +is written as such on save. `corpus.json` says `site.scope: "cited"` for a report-only site. `channels` keeps its meaning (the pool a site's citations may resolve against). Docs in `SITE_FIELD_DOCS`, -`SITE.md` regenerated; the editor's site form gets the publish control; a new **Reports** tab under +`SITE.md` regenerated; the editor's site form gets the Search checkbox; a new **Reports** tab under `/sites/<id>/` lists reports, their validation problems, and "Prepare evidence media". ## Routes (export app) @@ -136,10 +134,11 @@ A report must survive a takedown as files anyone can save and host again (slice `.export-index/sites/<id>/report-exports/<reportId>/`: - `report.html` — ONE file from `common/lib/report/exportHtml.ts` (pure view → HTML): inline CSS, no script; stills and post screenshots recompressed on the host (WebP, else JPEG, ≤ 1200 px wide) into data URIs; clips - linked on the site, never inlined. The heading mirrors the site: the series (else the title), one line of - dates and the export's revision, the document under review as a cited line on its rail, the subtitle. - Then the tally, sections → claims (verdict chip, the document's sentence as its still, findings with `[n]` - markers), the documents quoted with their archive links, and the numbered references (quote, speaker, date, + linked on the site, never inlined. The heading mirrors the site: the series and the title on two lines (or the + title alone), one line of dates and the export's revision, the document under review as a cited line on its rail, the subtitle. + Then the three tiers, each opened by a hairline with depth dots: the tally, sections → claims (verdict + chip, the document's sentence as its still — else its words, else the claim's text, plain — findings with + `[n]` markers, evidence with the added mark beside its number), the documents quoted with their archive links, and the numbered references (quote, speaker, date, record, the original at its time, archive links, the moment page and clip when the site has a `siteUrl`). - `report.pdf` — that HTML printed by headless Chromium (`importPlaywright`, A4); skipped with a note where Playwright or its browser is missing, never a failure. @@ -213,7 +212,7 @@ PER REPORT, never site-wide. - `compose-site` gains a reports stage: resolve citations against the shared transcript trees, verify quotes, write report + moment JSON, copy media and stills. -- `publish: "cited"` prunes everything corpus-shaped from `public/` (summaries, stats, transcripts, subs, posts, +- Search off (`search: false`) prunes everything corpus-shaped from `public/` (summaries, stats, transcripts, subs, posts, digests, archives, duplicates, tags, aliases, chart templates, service worker) — `public/` is shared by every site's build in turn, so stale data from the previous site must not survive — and an **out/ allowlist audit** refuses a cited build that contains anything outside reports/moments/media/stills/shell assets. @@ -238,7 +237,7 @@ PER REPORT, never site-wide. | Slice | What | Size | Needs | |---|---|---|---| | R0 | `common/lib/citations/` + `common/lib/report/`: schemas, validation, inline-cite extraction + numbering, moment keys, back-link index, shared verdict vocab; `CITATIONS.md` + `REPORT.md` generated | M | — | -| R1 | site keys `publish` / `reports` + site form control | S | — | +| R1 | site keys `search` / `reports` + site form control | S | — | | R2 | evidence media: cutter port to common, `reports prepare` CLI + job kind, cited-capture copier + guard amendment | L | R0 | | R3 | compose: reports stage, cited prune, allowlist audit, always-emit site/corpus, llms/sitemap variants | L | R0, R2 interface | | R4 | export app: routes, shared citation components (inline marker + preview, cards, reference list, verification badge), Reports header link, cited shell, placeholder params | L | R0 (fixture JSON) |