Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 57bd41f23d9f1394b0495d6278592fcf47240c0b
parent c5a01bff2c05b43577bd4f6abd94ea7b95db68d1
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon,  5 Oct 2026 03:38:48 -0400

Merge report-r6-converters (sweep markdown, /ask answers and report-to-video manifests convert into the report/citation model; reports convert back to a starter manifest; archilyzer reports convert / to-manifest)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
MPUBLISH.md | 1+
MREPORT.md | 4++++
Mcommon/bin/archilyzer.ts | 44++++++++++++++++++++++++++++++++++++++++++++
Acommon/bin/reports-convert.test.ts | 212+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/bin/reports-convert.ts | 140+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/cueWiden.mjs | 99+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/report/convert-server.ts | 123+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/report/convert.test.ts | 618+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/report/convertAsk.ts | 172+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/report/convertManifest.ts | 552+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/report/convertShared.ts | 628+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/report/convertSweep.ts | 52++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/lib/report/docs.ts | 18++++++++++++++++++
Mcommon/package.json | 1+
Meditor/CHANGELOG.md | 1+
Mumtool/report-to-video/README.md | 11+++++++++++
Mumtool/report-to-video/resolve-windows.mjs | 70+++++-----------------------------------------------------------------
17 files changed, 2681 insertions(+), 65 deletions(-)

diff --git a/PUBLISH.md b/PUBLISH.md @@ -54,6 +54,7 @@ entry points in `common/publish/build.ts`. | The homepage | /sites → Homepage → **Build homepage** (tick *Deploy after build*) / **Deploy homepage**, with an optional preview branch | `build-homepage` (`{"deploy":true}` to deploy after), `deploy-homepage` (`{"preview":"<branch>"}`) | `build homepage [--no-source]`, `deploy homepage [--preview <branch>]` | | The source mirror alone | — (every homepage build runs it) | — | `source publish [--force] [--check] [--keep-scratch]`, `source audit [<git dir>]` | | A site's report evidence media | — (a Reports tab is to come) | `reports-prepare` (`{"siteId"}`) | `reports prepare <id>` | +| A report from a /sweep report, an /ask answer or a report-to-video manifest, and a starter manifest from a report | — | — | `reports convert <sweep\|ask\|manifest> <in> --out <report.json> [--channels-dir <dir>]`, `reports to-manifest <report.json> --out <manifest.json>` | `pnpm archilyzer <command>` is the short form of `pnpm --filter yt-dlp-transcript-common exec tsx bin/archilyzer.ts <command>`; diff --git a/REPORT.md b/REPORT.md @@ -61,3 +61,7 @@ The shared vocabulary (`common/lib/report/verdicts.mjs` — the one copy; report | `CONTRADICTED` | Contradicted | `#e5534b` | | `NOT_FOUND` | Not found | `#8b93a7` | | `UNTESTABLE` | Untestable | `#7d8fd6` | + +## Making a report from what exists + +`archilyzer reports convert <sweep|ask|manifest> <in> --out <report.json>` makes a report of a /sweep report (markdown: its `#` heading the title, each `##` section a section, each list item or paragraph that cites a claim), an /ask answer (its `[n @ mm:ss]` markers resolved against its sources) or a report-to-video manifest (chapter cards the sections, `claim` entries the claims, clips, stills and posts the citations), and writes it only when it validates. A /sweep or /ask citation carries one second: `--channels-dir <transcripts/channels>` widens it through the record's cues to whole sentences (else it is the second plus 10 s) and finds the channel that keeps a post cited by its platform link. `archilyzer reports to-manifest <report.json> --out <manifest.json>` goes the other way: a starter manifest — a title card, a chapter card per section, each claim's still and clips stamped with its verdict, its posts. The converters are `common/lib/report/convert*.ts`; the warnings say what a conversion left out. diff --git a/common/bin/archilyzer.ts b/common/bin/archilyzer.ts @@ -162,6 +162,50 @@ export const COMMANDS: Command[] = [ }, }, { + path: ["reports", "convert"], + usage: + "<sweep|ask|manifest> <in> --out <report.json> [--channels-dir <dir>] [--id <id>] [--title <title>] a /sweep report (markdown), an /ask answer or a report-to-video manifest as a report.json, written only when it validates (--channels-dir: widen spans from the cues, find posts' channels)", + flags: { out: "string", "channels-dir": "string", id: "string", title: "string" }, + maxPositionals: 2, + run: async ({ positionals, flags }) => { + const [from, input] = positionals; + const { convertMain, isConvertFrom, CONVERT_FROM } = await import("./reports-convert"); + if (!isConvertFrom(from) || !input || typeof flags.out !== "string") { + console.error(`reports convert: give <${CONVERT_FROM.join("|")}> <in> --out <report.json>`); + return 2; + } + return convertMain({ + from, + input, + out: flags.out, + ...(typeof flags["channels-dir"] === "string" ? { channelsDir: flags["channels-dir"] } : {}), + ...(typeof flags.id === "string" ? { id: flags.id } : {}), + ...(typeof flags.title === "string" ? { title: flags.title } : {}), + }); + }, + }, + { + path: ["reports", "to-manifest"], + usage: + "<report.json> --out <manifest.json> [--channels-dir <dir>] [--site-origin <url>] a starter report-to-video manifest from a report (a chapter card per section, a claim's still and clips stamped with its verdict, its posts)", + flags: { out: "string", "channels-dir": "string", "site-origin": "string" }, + maxPositionals: 1, + run: async ({ positionals, flags }) => { + const [input] = positionals; + if (!input || typeof flags.out !== "string") { + console.error("reports to-manifest: give <report.json> --out <manifest.json>"); + return 2; + } + const { toManifestMain } = await import("./reports-convert"); + return toManifestMain({ + input, + out: flags.out, + ...(typeof flags["channels-dir"] === "string" ? { channelsDir: flags["channels-dir"] } : {}), + ...(typeof flags["site-origin"] === "string" ? { siteOrigin: flags["site-origin"] } : {}), + }); + }, + }, + { path: ["source", "publish"], usage: "[--force] [--check] [--keep-scratch] the scrubbed git mirror, raw tree, history pages (stagit, when installed) and tarball into homepage/public, behind the denied-literal gate (--check: audit and count, write nothing)", diff --git a/common/bin/reports-convert.test.ts b/common/bin/reports-convert.test.ts @@ -0,0 +1,212 @@ +// The converters over files: the disk callbacks (cues behind the text guard, +// a post's channel and record off the posts archive) and the two CLI mains, +// on a temp channels tree. + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { existsSync, mkdirSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import path from "node:path"; +import { convertMain, toManifestMain } from "./reports-convert"; +import { diskCuesOf, diskPostChannelOf, diskPostOf } from "../lib/report/convert-server"; +import { resolveCommand } from "./_cli"; +import { COMMANDS } from "./archilyzer"; + +const ORIGIN = "https://archive.example"; + +function tree() { + const root = mkdtempSync(path.join(tmpdir(), "reports-convert-")); + const channels = path.join(root, "channels"); + const rec = path.join(channels, "demo-channel", "data", "abc123"); + mkdirSync(rec, { recursive: true }); + writeFileSync( + path.join(rec, "transcript.cues.json"), + JSON.stringify({ + cues: [ + { start: 60.1, end: 62.0, text: "and then the host says" }, + { start: 58.2, end: 60.1, text: "This is where it starts." }, + { start: 62.4, end: 64.9, text: "we never said that." }, + ], + }), + ); + const social = path.join(channels, "demo-social"); + mkdirSync(path.join(social, "posts"), { recursive: true }); + writeFileSync(path.join(social, "posts-archive"), "twitter 987654321\n"); + writeFileSync( + path.join(social, "posts", "2024-05.jsonl"), + JSON.stringify({ + id: "987654321", + channelSlug: "demo-social", + author: "demo", + authorName: "Demo Account", + createdAt: "2024-05-01T12:00:00.000Z", + text: "a platform post", + url: "https://x.com/demo/status/987654321", + platform: "twitter", + }) + "\n", + ); + // A second channel holding the same post id: the handle picks. + const mirror = path.join(channels, "other-mirror"); + mkdirSync(mirror, { recursive: true }); + writeFileSync(path.join(mirror, "posts-archive"), "twitter 987654321\n"); + return { root, channels, done: () => rmSync(root, { recursive: true, force: true }) }; +} + +const quiet = () => { + const lines = { log: [] as string[], error: [] as string[] }; + return { lines, out: { log: (s: string) => lines.log.push(s), error: (s: string) => lines.error.push(s) } }; +}; + +test("diskCuesOf reads a record's cues sorted; a missing record or an unsafe id is null", async () => { + const t = tree(); + try { + const cuesOf = diskCuesOf(t.channels); + const cues = await cuesOf("demo-channel", "abc123"); + assert.deepEqual(cues?.map((c) => c.start), [58.2, 60.1, 62.4]); + assert.equal(await cuesOf("demo-channel", "nope"), null); + assert.equal(await cuesOf("demo-channel", ".."), null); + } finally { + t.done(); + } +}); + +test("diskCuesOf: a channel whose text is not readable is a warning, not a read", async () => { + const t = tree(); + try { + const legacy = path.join(t.channels, "legacy-channel"); + mkdirSync(legacy, { recursive: true }); + writeFileSync(path.join(legacy, "config.json"), JSON.stringify({ dataDir: "/elsewhere/legacy-channel/data" })); + const warnings: string[] = []; + assert.equal(await diskCuesOf(t.channels, warnings)("legacy-channel", "abc123"), null); + assert.equal(warnings.length, 1); + assert.match(warnings[0], /^legacy-channel: /); + } finally { + t.done(); + } +}); + +test("diskPostChannelOf finds the channel that keeps a post, the handle breaking a tie; diskPostOf reads its record", async () => { + const t = tree(); + try { + const channelOf = diskPostChannelOf(t.channels); + assert.equal(await channelOf({ platform: "x", id: "987654321", handle: "demo" }), "demo-social"); + assert.equal(await channelOf({ platform: "x", id: "987654321", handle: "other" }), "other-mirror"); + assert.equal(await channelOf({ platform: "x", id: "1" }), null); + const postOf = diskPostOf(t.channels); + assert.deepEqual(await postOf("demo-social", "987654321"), { + platform: "x", + url: "https://x.com/demo/status/987654321", + createdAt: "2024-05-01T12:00:00.000Z", + author: "demo", + authorName: "Demo Account", + }); + assert.equal(await postOf("demo-social", "2"), null); + } finally { + t.done(); + } +}); + +test("reports convert sweep: widened from the cues, the platform post found, written", async () => { + const t = tree(); + try { + const input = path.join(t.root, "sweep.md"); + writeFileSync( + input, + [ + "# A demo sweep", + "", + "## Findings", + "", + `- "we never said that" — [Demo @ 1:02](${ORIGIN}/?v=demo-channel%2Fabc123&t=62)`, + `- "a platform post" [post by demo](https://x.com/demo/status/987654321)`, + ].join("\n"), + ); + const out = path.join(t.root, "reports", "demo-sweep", "report.json"); + const { lines, out: o } = quiet(); + assert.equal(await convertMain({ from: "sweep", input, out, channelsDir: t.channels }, o), 0); + const report = JSON.parse(readFileSync(out, "utf8")); + assert.equal(report.id, "a-demo-sweep"); + assert.deepEqual([report.citations.c01.start, report.citations.c01.end], [60.1, 64.9]); + assert.equal(report.citations.p01.channel, "demo-social"); + assert.deepEqual(lines.error, []); + assert.match(lines.log[0], /report a-demo-sweep \(sweep\) — 1 section\(s\), 2 claim\(s\), 2 citation\(s\)/); + } finally { + t.done(); + } +}); + +test("reports convert refuses to write a report with problems; an unreadable input is exit 2", async () => { + const t = tree(); + try { + const input = path.join(t.root, "answer.json"); + writeFileSync(input, JSON.stringify({ answer: "Said [1 @ 0:05].", sources: [{ key: "demo-channel/abc123", snippets: [] }] })); + const out = path.join(t.root, "out.json"); + const { lines, out: o } = quiet(); + assert.equal(await convertMain({ from: "ask", input, out, id: "Not A Slug" }, o), 1); + assert.equal(existsSync(out), false); + assert.ok(lines.error.some((l) => l.startsWith(" id: must be a lowercase slug"))); + assert.equal(await convertMain({ from: "ask", input: path.join(t.root, "missing.json"), out }, quiet().out), 2); + writeFileSync(input, JSON.stringify({ unrelated: true })); + assert.equal(await convertMain({ from: "ask", input, out }, quiet().out), 2); + } finally { + t.done(); + } +}); + +test("reports to-manifest: a starter manifest, its posts read off the channels tree", async () => { + const t = tree(); + try { + const input = path.join(t.root, "report.json"); + writeFileSync( + input, + JSON.stringify({ + format: "archilyzer-report", + version: 1, + id: "demo-report", + kind: "sweep", + title: "A demo report", + citations: { + c01: { kind: "video", channel: "demo-channel", id: "abc123", start: 60.1, end: 64.9, quote: "we never said that." }, + p01: { kind: "post", channel: "demo-social", id: "987654321", quote: "a platform post" }, + }, + sections: [{ id: "s1", title: "Findings", claims: [{ id: "k1", text: "It was said.", citations: ["c01", "p01"] }] }], + }), + ); + const out = path.join(t.root, "video.manifest.json"); + const { lines, out: o } = quiet(); + assert.equal(await toManifestMain({ input, out, channelsDir: t.channels, siteOrigin: ORIGIN }, o), 0); + const m = JSON.parse(readFileSync(out, "utf8")); + assert.deepEqual( + m.timeline.map((e: { type: string; id: string }) => `${e.type}:${e.id}`), + ["card:title", "card:s1", "clip:c01"], + ); + assert.deepEqual(m.posts, [ + { + id: "p01", + platform: "x", + author: "Demo Account", + handle: "demo", + date: "2024-05-01T12:00:00.000Z", + text: "a platform post", + url: "https://x.com/demo/status/987654321", + attachTo: "c01", + siteChannel: "demo-social", + postId: "987654321", + }, + ]); + assert.equal(m.provenance.siteOrigin, ORIGIN); + assert.deepEqual(lines.error, []); + // An invalid report is refused, and nothing is written. + writeFileSync(input, JSON.stringify({ format: "archilyzer-report" })); + rmSync(out); + assert.equal(await toManifestMain({ input, out }, quiet().out), 1); + assert.equal(existsSync(out), false); + } finally { + t.done(); + } +}); + +test("the archilyzer table routes reports convert and reports to-manifest", () => { + assert.deepEqual(resolveCommand(COMMANDS, ["reports", "convert", "sweep", "in.md"])?.command.path, ["reports", "convert"]); + assert.deepEqual(resolveCommand(COMMANDS, ["reports", "to-manifest", "r.json"])?.command.path, ["reports", "to-manifest"]); +}); diff --git a/common/bin/reports-convert.ts b/common/bin/reports-convert.ts @@ -0,0 +1,140 @@ +// `archilyzer reports convert <sweep|ask|manifest> <in> --out <report.json>` +// and `archilyzer reports to-manifest <report.json> --out <manifest.json>` — +// the converters (lib/report/convert*.ts) over files. +// +// `--channels-dir <dir>` (a `transcripts/channels` tree) lets a conversion +// read what the documents do not carry: a cited record's cues, so a /sweep or +// /ask citation's one second widens to whole sentences; the channel that keeps +// a post cited by its platform link; a post's platform, link and date for a +// manifest's `posts`. Without it a span is the second plus 10 s and those +// posts are left out, each with a warning. Nothing is fetched. +// +// The output is checked before it is written: `convert` writes only a report +// the report validator passes, `to-manifest` reads only one. Exit 0 written +// (warnings on stderr), 1 refused with problems (nothing written), 2 an input +// that cannot be read. + +import { readFile } from "node:fs/promises"; +import path from "node:path"; +import { writeJsonAtomic } from "../lib/jsonFile-server"; +import type { Problem } from "../lib/citations/validate"; +import { askToReport, readAskAnswer } from "../lib/report/convertAsk"; +import { manifestToReport, reportToManifest } from "../lib/report/convertManifest"; +import { sweepToReport } from "../lib/report/convertSweep"; +import type { ConvertContext, Converted } from "../lib/report/convertShared"; +import { diskCuesOf, diskPostChannelOf, diskPostOf } from "../lib/report/convert-server"; + +type Out = { log: (s: string) => void; error: (s: string) => void }; + +export const CONVERT_FROM = ["sweep", "ask", "manifest"] as const; +export type ConvertFrom = (typeof CONVERT_FROM)[number]; + +export function isConvertFrom(v: unknown): v is ConvertFrom { + return typeof v === "string" && (CONVERT_FROM as readonly string[]).includes(v); +} + +function printProblems(out: Out, what: string, problems: Problem[]): void { + out.error(`${what}: ${problems.length} problem(s):`); + for (const p of problems) out.error(` ${p.path || "(root)"}: ${p.message}`); +} + +async function readInput(file: string, out: Out, what: string): Promise<string | null> { + try { + return await readFile(file, "utf8"); + } catch (err) { + out.error(`${what}: cannot read ${file}: ${(err as Error).message}`); + return null; + } +} + +function parseJson(text: string, file: string, out: Out, what: string): unknown { + try { + return JSON.parse(text); + } catch (err) { + out.error(`${what}: ${file} is not JSON: ${(err as Error).message}`); + return undefined; + } +} + +export async function convertMain( + opts: { from: ConvertFrom; input: string; out: string; channelsDir?: string; id?: string; title?: string }, + out: Out = console, +): Promise<number> { + const what = `reports convert ${opts.from}`; + const text = await readInput(opts.input, out, what); + if (text === null) return 2; + const warnings: string[] = []; + const ctx: ConvertContext = opts.channelsDir + ? { cuesOf: diskCuesOf(opts.channelsDir, warnings), postChannelOf: diskPostChannelOf(opts.channelsDir) } + : {}; + const named = { ...ctx, ...(opts.id ? { id: opts.id } : {}), ...(opts.title ? { title: opts.title } : {}) }; + + let converted: Converted; + if (opts.from === "sweep") { + converted = await sweepToReport(text, named); + } else { + const raw = parseJson(text, opts.input, out, what); + if (raw === undefined) return 2; + if (opts.from === "ask") { + const answer = readAskAnswer(raw); + if (!answer) { + out.error(`${what}: ${opts.input} is not an /ask answer ({ answer, sources }, a saved chat or an assistant message)`); + return 2; + } + converted = await askToReport(answer, named); + } else { + converted = await manifestToReport(raw, named); + const hasStills = Object.values(converted.report.citations ?? {}).some((c) => c.kind === "source" && c.image); + if (hasStills && path.resolve(path.dirname(opts.input)) !== path.resolve(path.dirname(opts.out))) { + warnings.push( + "the report's stills are paths relative to the manifest's directory — copy them beside the report (or write it there)", + ); + } + } + } + + for (const w of [...warnings, ...converted.warnings]) out.error(`warning: ${w}`); + if (converted.problems.length) { + printProblems(out, `${what}: not written`, converted.problems); + return 1; + } + await writeJsonAtomic(opts.out, converted.report, { mkdir: true }); + const r = converted.report; + const claims = r.sections.reduce((n, s) => n + (s.claims?.length ?? 0), 0); + out.log( + `${opts.out}: report ${r.id} (${r.kind}) — ${r.sections.length} section(s), ${claims} claim(s), ${Object.keys(r.citations ?? {}).length} citation(s)`, + ); + return 0; +} + +export async function toManifestMain( + opts: { input: string; out: string; channelsDir?: string; siteOrigin?: string }, + out: Out = console, +): Promise<number> { + const what = "reports to-manifest"; + const text = await readInput(opts.input, out, what); + if (text === null) return 2; + const raw = parseJson(text, opts.input, out, what); + if (raw === undefined) return 2; + const result = await reportToManifest(raw, { + ...(opts.siteOrigin ? { siteOrigin: opts.siteOrigin } : {}), + ...(opts.channelsDir ? { postOf: diskPostOf(opts.channelsDir) } : {}), + }); + if (!result.ok) { + printProblems(out, `${what}: ${opts.input} is not a sound report`, result.problems); + return 1; + } + const warnings = [...result.warnings]; + const timeline = result.manifest.timeline as { type: string }[]; + if ( + timeline.some((e) => e.type === "image") && + path.resolve(path.dirname(opts.input)) !== path.resolve(path.dirname(opts.out)) + ) { + warnings.push("the images' src are paths relative to the report's directory — copy the stills beside the manifest"); + } + for (const w of warnings) out.error(`warning: ${w}`); + await writeJsonAtomic(opts.out, result.manifest, { mkdir: true }); + const posts = (result.manifest.posts as unknown[] | undefined)?.length ?? 0; + out.log(`${opts.out}: ${timeline.length} timeline entr${timeline.length === 1 ? "y" : "ies"}, ${posts} post(s)`); + return 0; +} diff --git a/common/lib/cueWiden.mjs b/common/lib/cueWiden.mjs @@ -0,0 +1,99 @@ +// cueWiden.mjs — widen a span of a record from cue edges to whole sentences. +// +// A span starts life as the cue span covering a quote, and a cue boundary is a +// bad place to cut: ASR breaks cues where the caption line wrapped, which is +// routinely mid-sentence and often mid-word. `widen` walks outward from the +// span to the nearest sentence boundary in the cues — a cue whose text ends in +// . ? or ! — so the span carries the whole thought. +// +// Two consumers ask it and must get the same answer: +// - umtool's report-to-video (`resolve-windows.mjs`, which re-exports it and +// widens a manifest's clip windows, and `cutToQuote`); +// - the report converters (lib/report/convert*.ts), which widen a citation +// that carries one second into a span. +// So it lives HERE, once. Plain ESM with no imports: umtool's scripts run +// under bare node, which cannot load a `.ts` file. + +/** @typedef {{ start: number, end: number, text: string }} Cue */ + +const ENDS_SENTENCE = /[.!?]["'”’)\]]*\s*$/; + +// A cue that is only "[music]" or "[ __ ]" (the profanity bleep) carries no +// sentence signal; treat it as transparent so expansion walks past it. +const IS_FILLER = /^\s*(\[[^\]]*\]|>>|♪|—|-)*\s*$/; + +// A manifest stores times rounded to 2 dp, so a value read back from it can sit +// a hair BELOW the cue end it came from. Without a tolerance the end lookup then +// lands on the previous cue, the forward search runs on to the next sentence, and +// the span grows a little every time this is run — it has to be a fixed point. +export const CUE_EPS = 0.02; + +/** + * First cue whose span contains t, else the nearest one on the right side. + * @param {readonly Cue[]} cues + * @param {number} t + * @param {"start" | "end"} which + */ +function indexAt(cues, t, which) { + let idx = cues.findIndex((c) => c.end > t); + if (idx < 0) idx = cues.length - 1; + if (which === "end") { + let j = cues.findIndex((c) => c.end >= t - CUE_EPS); + if (j < 0) j = cues.length - 1; + idx = j; + } + return idx; +} + +/** + * The span widened to sentence edges, within the lead budget at the start; the + * end looks at most `maxTail` past the span and is never clamped short of a + * sentence it found. `cues` is sorted by start and not empty. + * + * @param {readonly Cue[]} cues + * @param {number} start + * @param {number} end + * @param {{ maxLead?: number, maxTail?: number }} [opts] + * @returns {{ start: number, end: number, leadCues: number, tailCues: number }} + */ +export function widen(cues, start, end, { maxLead = 8, maxTail = 12 } = {}) { + /** @param {Cue} c */ + const isBoundary = (c) => ENDS_SENTENCE.test(c.text) && !IS_FILLER.test(c.text); + const i0 = indexAt(cues, start, "start"); + const i1 = indexAt(cues, end, "end"); + + // START: the latest cue that OPENS a sentence (i.e. its predecessor closes + // one) at or before the quote, within the lead budget. Finding no such cue + // means every candidate lead-in is a sentence fragment, so take none at all — + // a fragment is the irrelevant context we are trying to avoid, not context. + let si = null; + for (let i = i0; i > 0; i -= 1) { + if (start - cues[i].start > maxLead) break; + if (isBoundary(cues[i - 1])) { + si = i; + break; + } + } + if (si === null) si = i0; + + // END: the first cue that CLOSES a sentence at or after the quote. Never + // clamp to a budget here — stopping partway through a sentence is exactly the + // mid-thought ending this is meant to remove, so the budget only decides how + // far to look, and failing to find one falls back to the original cue end. + let ei = null; + for (let j = i1; j < cues.length; j += 1) { + if (cues[j].end - end > maxTail) break; + if (isBoundary(cues[j])) { + ei = j; + break; + } + } + if (ei === null) ei = i1; + + return { + start: cues[si].start, + end: cues[ei].end, + leadCues: i0 - si, + tailCues: ei - i1, + }; +} diff --git a/common/lib/report/convert-server.ts b/common/lib/report/convert-server.ts @@ -0,0 +1,123 @@ +// The converters' disk half: what a channels tree can tell a converter, as the +// callbacks the converters take — a record's cues, the archive channel that +// keeps a post, a post's record. The CLI is bin/reports-convert.ts. +// +// Reads only TEXT: `channels/<slug>/data/<id>/transcript.cues.json` (behind +// the text guard, lib/channelMedia.ts — a legacy channel or one mid-migration +// is "no cues", with a warning, never a read of the wrong tree), and each +// channel's `posts-archive` and `posts/*.jsonl`. Writes nothing. +// +// SERVER-ONLY (node:fs). + +import { readdir, readFile } from "node:fs/promises"; +import path from "node:path"; +import { assertChannelTextReadable } from "../channelMedia"; +import { isSafeChannelSegment, isSafeIdSegment } from "../citations/moments"; +import { readAllPosts, readSeenPostIds } from "../posts-server"; +import type { Post } from "../posts"; +import type { Cue, CuesOf, PostChannelOf } from "./convertShared"; +import type { PostOf, PostRecord } from "./convertManifest"; + +function isCue(v: unknown): v is Cue { + if (!v || typeof v !== "object") return false; + const c = v as Record<string, unknown>; + return typeof c.start === "number" && typeof c.end === "number" && typeof c.text === "string" && c.end >= c.start; +} + +// A record's cues off the channels tree, cached per record. A channel whose +// text is not readable is one warning and no cues. +export function diskCuesOf(channelsDir: string, warnings: string[] = []): CuesOf { + const guarded = new Map<string, Promise<boolean>>(); + const cache = new Map<string, Promise<Cue[] | null>>(); + const readable = (slug: string) => { + let p = guarded.get(slug); + if (!p) { + p = assertChannelTextReadable({ channelsDir }, slug).then( + () => true, + (err: Error) => { + warnings.push(`${slug}: ${err.message} — no cues read from it`); + return false; + }, + ); + guarded.set(slug, p); + } + return p; + }; + return (channel, id) => { + if (!isSafeChannelSegment(channel) || !isSafeIdSegment(id)) return Promise.resolve(null); + const key = `${channel}/${id}`; + let p = cache.get(key); + if (!p) { + p = (async () => { + if (!(await readable(channel))) return null; + let doc: unknown; + try { + doc = JSON.parse(await readFile(path.join(channelsDir, channel, "data", id, "transcript.cues.json"), "utf8")); + } catch { + return null; + } + const cues = (doc as { cues?: unknown })?.cues; + if (!Array.isArray(cues)) return null; + return cues.filter(isCue).sort((a, b) => a.start - b.start); + })(); + cache.set(key, p); + } + return p; + }; +} + +// The channel directories under the tree, by name. +async function channelSlugs(channelsDir: string): Promise<string[]> { + const entries = await readdir(channelsDir, { withFileTypes: true }).catch(() => []); + return entries.filter((e) => e.isDirectory() && isSafeChannelSegment(e.name)).map((e) => e.name).sort(); +} + +const squash = (v: string) => v.toLowerCase().replace(/^@/, "").replace(/[^a-z0-9]+/g, ""); + +// The archive channel that keeps a post: every channel's posts archive, read +// once; among several, the one whose slug carries the post's handle. +export function diskPostChannelOf(channelsDir: string): PostChannelOf { + let index: Promise<Map<string, string[]>> | null = null; + const build = async () => { + const byId = new Map<string, string[]>(); + for (const slug of await channelSlugs(channelsDir)) { + for (const id of await readSeenPostIds(path.join(channelsDir, slug))) { + byId.set(id, [...(byId.get(id) ?? []), slug]); + } + } + return byId; + }; + return async ({ id, handle }) => { + index ??= build(); + const slugs = (await index).get(id) ?? []; + if (slugs.length <= 1) return slugs[0] ?? null; + const h = handle ? squash(handle.split(".")[0]) : ""; + return (h && slugs.find((s) => squash(s).includes(h))) || slugs[0]; + }; +} + +// A post's record off its channel's archive, each channel read once. +export function diskPostOf(channelsDir: string): PostOf { + const channels = new Map<string, Promise<Map<string, Post>>>(); + return async (channel, id) => { + if (!isSafeChannelSegment(channel)) return null; + let p = channels.get(channel); + if (!p) { + p = readAllPosts(path.join(channelsDir, channel)).then( + (posts) => new Map(posts.map((post) => [post.id, post])), + () => new Map(), + ); + channels.set(channel, p); + } + const post = (await p).get(id); + if (!post) return null; + const rec: PostRecord = { + platform: post.platform === "bluesky" ? "bluesky" : "x", + url: post.url, + createdAt: post.createdAt, + author: post.author, + }; + if (post.authorName) rec.authorName = post.authorName; + return rec; + }; +} diff --git a/common/lib/report/convert.test.ts b/common/lib/report/convert.test.ts @@ -0,0 +1,618 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + DEFAULT_SPAN_SECONDS, + IdAllocator, + parseCitationHref, + spanAt, + structureMarkdown, + type Cue, + type CuesOf, +} from "./convertShared"; +import { sweepToReport } from "./convertSweep"; +import { askToReport, parseClock, readAskAnswer } from "./convertAsk"; +import { manifestToReport, reportToManifest, STARTER_RENDER } from "./convertManifest"; +import { validateReport } from "./validate"; +import type { Report } from "./schema"; + +// ─── Fixtures: one channel of two records, one post channel ─── + +const CUES: Record<string, Cue[]> = { + "demo-channel/abc123": [ + { start: 58.2, end: 60.1, text: "This is where it starts." }, + { start: 60.1, end: 62.0, text: "and then the host says" }, + { start: 62.4, end: 64.9, text: "we never said that" }, + { start: 64.9, end: 67.3, text: "not once, not ever." }, + { start: 67.3, end: 70.0, text: "Next topic." }, + ], +}; + +const cuesOf: CuesOf = async (channel, id) => CUES[`${channel}/${id}`] ?? null; + +const ORIGIN = "https://archive.example"; +const moment = (channel: string, id: string, t: number, encode = true) => + `${ORIGIN}/?v=${encode ? encodeURIComponent(`${channel}/${id}`) : `${channel}/${id}`}&t=${t}`; + +// ─── Links ─── + +test("parseCitationHref reads viewer moments, archive posts, moment pages and platform posts", () => { + assert.deepEqual(parseCitationHref(moment("demo-channel", "abc123", 62)), { + kind: "span", + channel: "demo-channel", + id: "abc123", + seconds: 62, + }); + assert.deepEqual(parseCitationHref(moment("demo-channel", "-def456", 5, false)), { + kind: "span", + channel: "demo-channel", + id: "-def456", + seconds: 5, + }); + assert.deepEqual(parseCitationHref(`${ORIGIN}/?v=demo-social%2F1234567890&vm=post`), { + kind: "post", + channel: "demo-social", + id: "1234567890", + }); + assert.deepEqual(parseCitationHref(`${ORIGIN}/m/demo-channel/abc123/12.50-20.00/`), { + kind: "span", + channel: "demo-channel", + id: "abc123", + seconds: 12.5, + end: 20, + }); + assert.deepEqual(parseCitationHref("https://x.com/demo/status/987654321"), { + kind: "post-original", + platform: "x", + id: "987654321", + handle: "demo", + }); + assert.deepEqual(parseCitationHref("https://bsky.app/profile/demo.example/post/3abcxyz"), { + kind: "post-original", + platform: "bluesky", + id: "3abcxyz", + handle: "demo.example", + }); + assert.equal(parseCitationHref("https://www.youtube.com/watch?v=abc123&t=5s"), null); + assert.equal(parseCitationHref("https://example.test/article"), null); + assert.equal(parseCitationHref("cite:c01"), null); +}); + +// ─── Spans ─── + +test("spanAt: no cues is the second plus the default span", async () => { + assert.deepEqual(await spanAt("demo-channel", "nocues", 30, "anything", {}), { + start: 30, + end: 30 + DEFAULT_SPAN_SECONDS, + from: "default", + }); + assert.deepEqual(await spanAt("demo-channel", "nocues", 30, "x", { spanSeconds: 4 }), { start: 30, end: 34, from: "default" }); +}); + +test("spanAt: the cited cue, run on to the quote's length, widened to whole sentences", async () => { + // t=62 cites the cue starting at 62.4; the quote is two cues long; the + // sentence opens after "This is where it starts." and closes at "not ever.". + const span = await spanAt("demo-channel", "abc123", 62, "we never said that, not once, not ever", { cuesOf }); + assert.deepEqual(span, { start: 60.1, end: 67.3, from: "cues" }); +}); + +test("spanAt: an end the link gives is kept", async () => { + assert.deepEqual(await spanAt("demo-channel", "abc123", 12.5, "q", { cuesOf }, 20), { start: 12.5, end: 20, from: "link" }); +}); + +test("IdAllocator keeps a sound free id, repairs and suffixes the rest", () => { + const ids = new IdAllocator(); + assert.equal(ids.claim("c01"), "c01"); + assert.equal(ids.claim("c01"), "c01-2"); + assert.equal(ids.claim("-x y"), "x-y"); + assert.equal(ids.next("c"), "c02"); + assert.equal(ids.next("c"), "c03"); +}); + +// ─── The markdown engine ─── + +test("structureMarkdown: title, summary, sections, claims and bodies; code is not citing", async () => { + const md = [ + "# The title", + "", + "A summary [citing](x:1).", + "", + "## Section one", + "", + "- A claim [cite](x:2)", + "- Not a claim.", + "", + "```", + "[not](x:3)", + "```", + "", + "### A sub-heading stays in the body", + "", + "## Section one", + "", + "Plain paragraph with `[code](x:4)`.", + ].join("\n"); + const seen: string[] = []; + const doc = await structureMarkdown( + md, + async ({ href }) => { + seen.push(href); + return `c${href.slice(2)}`; + }, + { defaultSection: "Default" }, + ); + assert.deepEqual(seen, ["x:1", "x:2"]); + assert.equal(doc.title, "The title"); + assert.equal(doc.summary, "A summary [citing](cite:c1)."); + assert.deepEqual( + doc.sections.map((s) => s.id), + ["section-one", "section-one-2"], + ); + assert.deepEqual(doc.sections[0].claims, [{ id: "k01", text: "A claim", citations: ["c2"] }]); + assert.equal(doc.sections[0].body, "- Not a claim.\n\n```\n[not](x:3)\n```\n\n### A sub-heading stays in the body"); + assert.equal(doc.sections[1].body, "Plain paragraph with `[code](x:4)`."); +}); + +test("structureMarkdown: no section headings is one section of the default title", async () => { + const doc = await structureMarkdown("One [a](x:1).\n\nTwo.", async () => "c1", { defaultSection: "Answer" }); + assert.equal(doc.title, null); + assert.equal(doc.summary, ""); + assert.equal(doc.sections.length, 1); + assert.equal(doc.sections[0].title, "Answer"); + assert.equal(doc.sections[0].claims?.[0].text, "One."); + assert.equal(doc.sections[0].body, "Two."); +}); + +// ─── /sweep ─── + +const SWEEP_MD = [ + "# Demo sweep", + "", + `What the sweep looked for; [the method](https://example.test/method).`, + "", + "## First topic", + "", + `- "We never said that." — [Demo episode @ 1:02](${moment("demo-channel", "abc123", 62)})`, + "- Background with no citation.", + "", + `> "Another line here."`, + "", + `— [Demo episode @ 2:00](${moment("demo-channel", "abc123", 120, false)})`, + "", + "## Posts", + "", + `- The account posted it: "the post words" [post by demo, 2024-05-01](${ORIGIN}/?v=demo-social%2F1234567890&vm=post)`, + `- On X: "a platform post" [post by demo](https://x.com/demo/status/987654321)`, + `- "first" [A @ 0:05](${moment("demo-channel", "abc123", 5)}) and later "second" [B @ 0:30](${moment("demo-channel", "-def456", 30)}).`, +].join("\n"); + +test("sweepToReport: sections, claims, citations, quotes and a sound report", async () => { + const { report, warnings, problems } = await sweepToReport(SWEEP_MD, { + postChannelOf: async ({ id }) => (id === "987654321" ? "demo-social" : null), + }); + assert.deepEqual(problems, []); + assert.deepEqual(warnings, []); + assert.equal(report.id, "demo-sweep"); + assert.equal(report.kind, "sweep"); + assert.equal(report.title, "Demo sweep"); + assert.equal(report.summary, "What the sweep looked for; [the method](https://example.test/method)."); + assert.deepEqual( + report.sections.map((s) => [s.id, s.title]), + [ + ["first-topic", "First topic"], + ["posts", "Posts"], + ], + ); + const [first, posts] = report.sections; + assert.equal(first.body, "- Background with no citation."); + assert.deepEqual(first.claims, [ + { id: "k01", text: '"We never said that."', citations: ["c01"] }, + { id: "k02", text: '"Another line here."', citations: ["c02"] }, + ]); + assert.deepEqual(report.citations?.c01, { + kind: "video", + channel: "demo-channel", + id: "abc123", + start: 62, + end: 62 + DEFAULT_SPAN_SECONDS, + quote: "We never said that.", + label: "Demo episode", + }); + assert.equal(report.citations?.c02.quote, "Another line here."); + assert.deepEqual(report.citations?.p01, { + kind: "post", + channel: "demo-social", + id: "1234567890", + quote: "the post words", + label: "post by demo, 2024-05-01", + date: "2024-05-01", + }); + assert.deepEqual(report.citations?.p02, { + kind: "post", + channel: "demo-social", + id: "987654321", + quote: "a platform post", + label: "post by demo", + }); + // Two citations with words between them keep their places in findings. + const both = posts.claims?.[2]; + assert.deepEqual(both?.citations, ["c03", "c04"]); + assert.equal(both?.findings, '"first" [A @ 0:05](cite:c03) and later "second" [B @ 0:30](cite:c04).'); + assert.equal(report.citations?.c03.quote, "first"); + assert.equal(report.citations?.c04.quote, "second"); + assert.deepEqual(validateReport(report), []); +}); + +test("sweepToReport: cues widen the spans; a platform post with no channel stays a plain link", async () => { + const { report, warnings, problems } = await sweepToReport(SWEEP_MD, { cuesOf }); + assert.deepEqual(problems, []); + const c01 = report.citations?.c01; + assert.ok(c01 && c01.kind === "video"); + assert.equal(c01.start, 60.1); + assert.equal(c01.end, 67.3); + assert.ok(warnings.some((w) => w.startsWith("https://x.com/demo/status/987654321: no archive channel"))); + assert.ok(warnings.some((w) => w.startsWith("demo-channel/-def456: no cues"))); + assert.match(JSON.stringify(report.sections), /\[post by demo\]\(https:\/\/x\.com\/demo\/status\/987654321\)/); +}); + +// ─── /ask ─── + +const ASK = { + question: "What did the host say?", + answer: "The host said it [1 @ 1:01], then posted about it [2].\n\nA stray marker [9].\n\n`[1]` in code.", + sources: [ + { + key: "demo-channel/abc123", + title: "Demo episode", + snippets: [ + { clock: "1:02", seconds: 62, text: "we never said that" }, + { clock: "5:00", seconds: 300, text: "something else" }, + ], + }, + { key: "demo-social/1234567890", title: "a post", isPost: true, uploadDate: "20240501", snippets: [{ clock: "", seconds: 0, text: "the post words" }] }, + ], +}; + +test("parseClock reads mm:ss and h:mm:ss", () => { + assert.equal(parseClock("1:02"), 62); + assert.equal(parseClock("1:00:05"), 3605); + assert.equal(parseClock("x:10"), null); +}); + +test("readAskAnswer takes an answer, a saved chat and an assistant message", () => { + assert.deepEqual(readAskAnswer(ASK), ASK); + const chat = { + messages: [ + { role: "user", content: "the question" }, + { role: "assistant", content: "the answer [1]", sources: ASK.sources }, + ], + }; + assert.deepEqual(readAskAnswer(chat), { answer: "the answer [1]", sources: ASK.sources, question: "the question" }); + assert.deepEqual(readAskAnswer({ ...chat, report: "the report [1]", reportSources: ASK.sources }), { + answer: "the report [1]", + sources: ASK.sources, + question: undefined, + }); + assert.equal(readAskAnswer({ role: "assistant", content: "x" })?.answer, "x"); + assert.equal(readAskAnswer({ nothing: true }), null); + assert.equal(readAskAnswer([]), null); +}); + +test("askToReport: markers become citations snapped to excerpt lines; a stray marker is a warning", async () => { + const { report, warnings, problems } = await askToReport(ASK, { cuesOf }); + assert.deepEqual(problems, []); + assert.equal(report.title, "What did the host say?"); + assert.equal(report.id, "what-did-the-host-say"); + assert.equal(report.subtitle, undefined); + assert.equal(report.sections.length, 1); + assert.equal(report.sections[0].title, "Answer"); + const claim = report.sections[0].claims?.[0]; + assert.deepEqual(claim?.citations, ["c01", "p01"]); + assert.equal( + claim?.findings, + "The host said it [1 @ 1:01](cite:c01), then posted about it [2](cite:p01).", + ); + // 1:01 snaps to the excerpt line at 62; the cues widen it to the sentence. + assert.deepEqual(report.citations?.c01, { + kind: "video", + channel: "demo-channel", + id: "abc123", + start: 60.1, + end: 67.3, + quote: "we never said that", + label: "Demo episode", + }); + assert.deepEqual(report.citations?.p01, { + kind: "post", + channel: "demo-social", + id: "1234567890", + quote: "the post words", + date: "2024-05-01", + }); + assert.equal(report.sections[0].body, "A stray marker [9].\n\n`[1]` in code."); + assert.deepEqual(warnings, ["[9] names no source (the answer has 2)"]); +}); + +// ─── report-to-video manifests ─── + +function demoManifest() { + return { + schemaVersion: 1, + slug: "demo-factcheck", + title: "A demo fact-check", + subtitle: "What was said, against the record", + provenance: { siteOrigin: ORIGIN, channelSlug: "demo-channel" }, + render: { width: 1920 }, + timeline: [ + { type: "card", id: "t0", style: "title", heading: "A demo fact-check" }, + { type: "card", id: "ch1", style: "chapter", heading: "Chapter one", sub: "The first part" }, + { + type: "image", + id: "a01", + src: "stills/a01.png", + title: "A demo article", + quote: "The article's claim.", + date: "2026-01-02", + citeUrl: "https://example.test/article", + claim: { id: "k1", verdict: "CONTRADICTED" }, + }, + { type: "clip", id: "c01", video: "abc123", start: 60.1, end: 67.3, cite: 60, quote: "we never said that", claim: { id: "k1", verdict: "CONTRADICTED" } }, + { type: "clip", id: "c02", channel: "demo-other", video: "-def456", start: 10, end: 20, cite: 10, quote: "a second record", claim: { id: "k1", verdict: "CONTRADICTED" } }, + { type: "card", id: "ch2", style: "chapter", heading: "Chapter two" }, + { type: "clip", id: "c03", video: "abc123", start: 100, end: 110, cite: 100, quote: "context, no claim" }, + { type: "clip", id: "c04", video: "abc123", start: 200, end: 210, cite: 200, quote: "pinned by the ledger" }, + { type: "clip", id: "x01", src: "media/episode.mp3", start: 1, end: 5 }, + ], + posts: [ + { + id: "pst1", + platform: "x", + date: "2026-01-03", + text: "the post words", + url: "https://x.com/i/status/1234567890", + attachTo: "c01", + siteChannel: "demo-social", + }, + ], + ledger: [ + { id: "L1", date: "2026-01-04", label: "a ledger label", quote: "the ledger's words", channel: "demo-channel", video: "abc123", cite: 200, entryId: "c04" }, + { id: "L2", date: "2026-01-05", quote: "never clipped", channel: "demo-channel", video: "abc123", cite: 62 }, + ], + }; +} + +test("manifestToReport: chapters, claims with verdicts, the source sentence, posts, the ledger", async () => { + const { report, warnings, problems } = await manifestToReport(demoManifest(), { cuesOf }); + assert.deepEqual(problems, []); + assert.equal(report.id, "demo-factcheck"); + assert.equal(report.kind, "factcheck"); + assert.equal(report.subtitle, "What was said, against the record"); + assert.deepEqual( + report.sections.map((s) => [s.id, s.title, s.body]), + [ + ["ch1", "Chapter one", "The first part"], + ["ch2", "Chapter two", "- “context, no claim” [abc123](cite:c03)"], + ["ledger", "Not clipped", undefined], + ], + ); + assert.deepEqual(report.sections[0].claims, [ + { + id: "k1", + text: "The article's claim.", + verdict: "CONTRADICTED", + sourceQuote: { citation: "a01" }, + citations: ["c01", "pst1", "c02"], + }, + ]); + assert.deepEqual(report.sections[1].claims, [ + { id: "L1", title: "a ledger label", text: "the ledger's words", citations: ["c04"] }, + ]); + assert.deepEqual(report.sections[2].claims, [{ id: "L2", text: "never clipped", citations: ["l-L2"] }]); + assert.deepEqual(report.citations?.["l-L2"], { + kind: "video", + channel: "demo-channel", + id: "abc123", + start: 60.1, + end: 67.3, + quote: "never clipped", + }); + assert.deepEqual(report.citations?.c02, { + kind: "video", + channel: "demo-other", + id: "-def456", + start: 10, + end: 20, + quote: "a second record", + }); + assert.deepEqual(report.citations?.a01, { + kind: "source", + source: "src-a01", + quote: "The article's claim.", + image: "stills/a01.png", + }); + assert.deepEqual(report.sources?.["src-a01"], { + kind: "other", + title: "A demo article", + url: "https://example.test/article", + date: "2026-01-02", + }); + assert.deepEqual(report.citations?.pst1, { + kind: "post", + channel: "demo-social", + id: "1234567890", + quote: "the post words", + date: "2026-01-03", + }); + assert.deepEqual(warnings, ["timeline[8] x01: a clip of its own media (src) cites no record — left out"]); +}); + +test("reportToManifest: a starter cut — title card, a chapter per section, the claim's still and clips with its verdict, its posts", async () => { + const { report } = await manifestToReport(demoManifest()); + const result = await reportToManifest(report, { siteOrigin: ORIGIN, today: "2026-10-05" }); + assert.ok(result.ok); + const m = result.manifest; + assert.equal(m.slug, "demo-factcheck"); + assert.equal(m.generatedOn, "2026-10-05"); + assert.deepEqual(m.render, STARTER_RENDER); + assert.deepEqual(m.provenance, { siteOrigin: ORIGIN, channelSlug: "demo-channel", channel: "" }); + const timeline = m.timeline as Record<string, unknown>[]; + assert.deepEqual( + timeline.map((e) => `${e.type}:${e.id}`), + ["card:title", "card:ch1", "image:a01", "clip:c01", "clip:c02", "card:ch2", "clip:c03", "clip:c04", "card:ledger", "clip:l-L2"], + ); + assert.deepEqual(timeline[3], { + type: "clip", + id: "c01", + channel: "demo-channel", + video: "abc123", + start: 60.1, + end: 67.3, + cite: 60, + quote: "we never said that", + claim: { id: "k1", verdict: "CONTRADICTED" }, + }); + // A claim without a verdict is not stamped. + assert.equal(timeline[7].claim, undefined); + assert.deepEqual(m.posts, [ + { + id: "pst1", + platform: "x", + date: "2026-01-03", + text: "the post words", + url: "https://x.com/i/status/1234567890", + attachTo: "c01", + siteChannel: "demo-social", + postId: "1234567890", + }, + ]); +}); + +test("round trip: manifest → report → manifest keeps every clip, still, chapter, claim and post", async () => { + const before = demoManifest(); + const { report, problems } = await manifestToReport(before); + assert.deepEqual(problems, []); + const result = await reportToManifest(report); + assert.ok(result.ok); + const after = result.manifest; + const project = (timeline: Record<string, unknown>[]) => + timeline + .filter((e) => (e.type === "clip" && !e.src) || e.type === "image" || (e.type === "card" && e.style === "chapter")) + .map((e) => + e.type === "card" + ? { type: "card", heading: e.heading } + : e.type === "image" + ? { type: "image", id: e.id, src: e.src, quote: e.quote, title: e.title, claim: e.claim } + : { + type: "clip", + id: e.id, + channel: e.channel ?? "demo-channel", + video: e.video, + start: e.start, + end: e.end, + quote: e.quote, + claim: e.claim, + }, + ); + assert.deepEqual(project(after.timeline as Record<string, unknown>[]), [ + ...project(before.timeline), + // The ledger's claim that no clip pinned gains a clip of its own moment. + { type: "card", heading: "Not clipped" }, + { type: "clip", id: "l-L2", channel: "demo-channel", video: "abc123", start: 62, end: 72, quote: "never clipped", claim: undefined }, + ]); + const posts = (after.posts as Record<string, unknown>[]).map(({ id, platform, date, text, url, attachTo, siteChannel }) => ({ + id, + platform, + date, + text, + url, + attachTo, + siteChannel, + })); + assert.deepEqual(posts, before.posts); +}); + +function demoReport(): Report { + return { + format: "archilyzer-report", + version: 1, + id: "demo-report", + kind: "factcheck", + title: "A demo report", + sources: { s0: { kind: "article", title: "A demo article", url: "https://example.test/a", date: "2026-01-02" } }, + citations: { + a01: { kind: "source", source: "s0", quote: "The first claim.", image: "stills/a01.png" }, + c01: { kind: "video", channel: "demo-channel", id: "abc123", start: 10, end: 20, quote: "q1", pad: { before: 5, after: 5 } }, + c02: { kind: "audio", channel: "demo-podcast", id: "ep-042", start: 30.5, end: 42, quote: "q2", date: "2025-03-04" }, + p01: { kind: "post", channel: "demo-social", id: "1234567890", quote: "the post words", date: "2026-01-03T10:00:00Z" }, + w01: { kind: "page", url: "https://example.test/page", quote: "q4" }, + }, + sections: [ + { + id: "ch1", + title: "Chapter one", + body: "Context: [here](cite:c02).", + claims: [ + { + id: "k1", + text: "The first claim.", + verdict: "PARTLY", + sourceQuote: { citation: "a01" }, + findings: "See [this](cite:c01) and [the page](cite:w01).", + citations: ["c01", "p01"], + }, + ], + }, + ], + }; +} + +test("round trip: report → manifest → report keeps sections, claims, verdicts, spans and quotes", async () => { + const before = demoReport(); + const result = await reportToManifest(before); + assert.ok(result.ok); + assert.deepEqual(result.warnings, ["w01: a page citation has no manifest entry — left out"]); + const { report: after, problems } = await manifestToReport(result.manifest); + assert.deepEqual(problems, []); + assert.equal(after.kind, "factcheck"); + assert.deepEqual(after.sections[0].claims, [ + { id: "k1", text: "The first claim.", verdict: "PARTLY", sourceQuote: { citation: "a01" }, citations: ["c01", "p01"] }, + ]); + assert.equal(after.sections[0].body, "- “q2” [ep-042](cite:c02)"); + const c = after.citations ?? {}; + // A pad is the site's, not the cut's: it does not make the trip. + assert.deepEqual(c.c01, { kind: "video", channel: "demo-channel", id: "abc123", start: 10, end: 20, quote: "q1" }); + assert.deepEqual(c.c02, { kind: "audio", channel: "demo-podcast", id: "ep-042", start: 30.5, end: 42, quote: "q2", date: "2025-03-04" }); + assert.deepEqual(c.p01, before.citations?.p01); + assert.equal(c.a01.kind === "source" && c.a01.image, "stills/a01.png"); + assert.equal(c.a01.quote, "The first claim."); +}); + +test("reportToManifest refuses a report that does not validate", async () => { + const bad = { ...demoReport(), citations: {} }; + const result = await reportToManifest(bad); + assert.equal(result.ok, false); + assert.ok(!result.ok && result.problems.some((p) => p.path === "sections[0].claims[0].citations[0]")); +}); + +test("reportToManifest: a post with no known platform or link is left out with a warning", async () => { + const r = demoReport(); + r.citations!.p01 = { kind: "post", channel: "demo-social", id: "3abcxyz", quote: "words", date: "2026-01-03" }; + const result = await reportToManifest(r); + assert.ok(result.ok); + assert.equal(result.manifest.posts, undefined); + assert.ok(result.warnings.some((w) => w.startsWith("p01: post demo-social/3abcxyz — its platform and link are unknown"))); + const withRecord = await reportToManifest(r, { + postOf: async () => ({ platform: "bluesky", url: "https://bsky.app/profile/demo.example/post/3abcxyz", author: "demo.example" }), + }); + assert.ok(withRecord.ok); + assert.deepEqual((withRecord.manifest.posts as Record<string, unknown>[])[0], { + id: "p01", + platform: "bluesky", + handle: "demo.example", + date: "2026-01-03", + text: "words", + url: "https://bsky.app/profile/demo.example/post/3abcxyz", + attachTo: "c01", + siteChannel: "demo-social", + postId: "3abcxyz", + }); +}); diff --git a/common/lib/report/convertAsk.ts b/common/lib/report/convertAsk.ts @@ -0,0 +1,172 @@ +// An /ask ANSWER → `report.json` (kind "sweep"). +// +// The export's /ask writes an answer (or a running report) in markdown that +// cites its sources with numbered markers — `[n]`, or `[n @ mm:ss]` for a +// moment (also `h:mm:ss`, stray spaces tolerated) — where `n` indexes the +// answer's source list, the retrieved records (export/app/ask/citations.tsx +// parses the same markers; export/app/lib/askRetrieval.ts `RetrievedVideo` is +// a source). Accepted as input, all JSON: +// +// { "answer": "<md>", "sources": [<source>…], "question"?: "…" } +// a saved chat ({ "messages": […], "report"?, "reportSources"? }): its +// report and report sources when it has them, else its last answer and +// that answer's sources, with the question before it +// one assistant message ({ "role": "assistant", "content", "sources" }) +// +// Each marker becomes `[label](cite:<id>)`: a post source a post citation +// (its words the quote), any other source a span at the marked second — +// snapped to the nearest second the source's excerpt lines carry, as the +// /ask page does, so a slightly-off model time lands on a real line — or at +// its first excerpt line when the marker has no time; its quote is that +// line's text. Then the answer is structured like a sweep (./convertShared.ts +// `structureMarkdown`): paragraphs and list items that cite are claims. + +import { + CitationRegistry, + emptyReport, + finish, + reportIdOf, + structureMarkdown, + type ConvertContext, + type Converted, +} from "./convertShared"; + +export type AskSource = { + // `<channel>/<id>`: the record (or post) the source is. + key: string; + title?: string; + uploadDate?: string; + snippets?: { clock?: string; seconds: number; text: string }[]; + isPost?: boolean; +}; + +export type AskAnswer = { answer: string; sources: AskSource[]; question?: string }; + +export type AskConvertOptions = ConvertContext & { + id?: string; + // Default: the question, else "Answer". + title?: string; + sectionTitle?: string; +}; + +const isObj = (v: unknown): v is Record<string, unknown> => !!v && typeof v === "object" && !Array.isArray(v); + +function sourcesOf(v: unknown): AskSource[] | null { + if (!Array.isArray(v)) return null; + return v.filter((s): s is AskSource => isObj(s) && typeof s.key === "string"); +} + +// The answer, its sources and its question out of any accepted shape; null for +// JSON that is none of them. +export function readAskAnswer(raw: unknown): AskAnswer | null { + if (!isObj(raw)) return null; + const question = typeof raw.question === "string" ? raw.question : undefined; + for (const [text, list] of [ + ["answer", "sources"], + ["report", "reportSources"], + ] as const) { + const sources = sourcesOf(raw[list]); + if (typeof raw[text] === "string" && sources) return { answer: raw[text] as string, sources, question }; + } + if (raw.role === "assistant" && typeof raw.content === "string") { + return { answer: raw.content, sources: sourcesOf(raw.sources) ?? [], question }; + } + if (Array.isArray(raw.messages)) { + const messages = raw.messages.filter(isObj); + for (let i = messages.length - 1; i >= 0; i--) { + const m = messages[i]; + if (m.role !== "assistant" || typeof m.content !== "string" || !m.content.trim()) continue; + const asked = messages.slice(0, i).reverse().find((u) => u.role === "user" && typeof u.content === "string"); + return { answer: m.content, sources: sourcesOf(m.sources) ?? [], question: question ?? (asked?.content as string | undefined) }; + } + } + return null; +} + +// `mm:ss` or `h:mm:ss` as whole seconds; null when malformed. +export function parseClock(s: string): number | null { + const parts = s.split(":"); + if (parts.length < 2 || parts.length > 3) return null; + let total = 0; + for (const p of parts) { + if (!/^\d+$/.test(p)) return null; + total = total * 60 + Number(p); + } + return total; +} + +// One citation marker: a source number, an optional `@ mm:ss`, stray spaces +// tolerated — not one already followed by a link target. +const MARKER_RE = /\[\s*(\d+)\s*(?:@\s*(\d{1,2}(?::\d{2}){1,2})\s*)?\](?!\()/g; + +// The excerpt line a marker cites: the one nearest the marked second, else +// the first. +function snippetAt(src: AskSource, seconds: number | null): { seconds: number; text: string } | null { + const lines = src.snippets ?? []; + if (!lines.length) return null; + if (seconds === null) return lines[0]; + return lines.reduce((best, s) => (Math.abs(s.seconds - seconds) < Math.abs(best.seconds - seconds) ? s : best)); +} + +const dateOfUpload = (d: string | undefined) => + d && /^\d{8}$/.test(d) ? `${d.slice(0, 4)}-${d.slice(4, 6)}-${d.slice(6, 8)}` : undefined; + +export async function askToReport(input: AskAnswer, opts: AskConvertOptions = {}): Promise<Converted> { + const warnings: string[] = []; + const registry = new CitationRegistry(opts, warnings); + + // The markers, outside code, as `[label](cite:<id>)`. + const segments = input.answer.split(/(```[\s\S]*?```|`[^`]*`)/g); + for (let i = 0; i < segments.length; i += 2) { + let out = ""; + let last = 0; + for (const m of segments[i].matchAll(MARKER_RE)) { + out += segments[i].slice(last, m.index); + last = m.index + m[0].length; + const n = Number(m[1]); + const src = input.sources[n - 1]; + if (!src) { + warnings.push(`[${m[1]}${m[2] ? ` @ ${m[2]}` : ""}] names no source (the answer has ${input.sources.length})`); + out += m[0]; + continue; + } + const slash = src.key.indexOf("/"); + if (slash <= 0) { + warnings.push(`source ${n} (${src.key}) is not a <channel>/<id> key — left as it is`); + out += m[0]; + continue; + } + const channel = src.key.slice(0, slash); + const id = src.key.slice(slash + 1); + const marked = m[2] ? parseClock(m[2]) : null; + let cid: string; + if (src.isPost) { + const words = (src.snippets ?? []).map((s) => s.text).join("\n").trim() || src.title || id; + cid = registry.post(channel, id, words, { date: dateOfUpload(src.uploadDate) }); + } else { + const line = snippetAt(src, marked); + if (!line) warnings.push(`source ${n} (${src.key}) has no excerpt lines — its quote is its title`); + cid = await registry.span(channel, id, line?.seconds ?? marked ?? 0, line?.text.trim() || src.title || id, { + label: src.title, + }); + } + out += `[${m[0].slice(1, -1).trim()}](cite:${cid})`; + } + segments[i] = out + segments[i].slice(last); + } + + const doc = await structureMarkdown( + segments.join(""), + async ({ href }) => (href.startsWith("cite:") ? href.slice("cite:".length) : null), + { defaultSection: opts.sectionTitle ?? "Answer" }, + ); + const question = input.question?.replace(/\s+/g, " ").trim(); + const title = opts.title ?? doc.title ?? question ?? "Answer"; + const report = emptyReport(reportIdOf(opts.id, title), "sweep", title); + if (question && question !== title) report.subtitle = question; + if (doc.summary) report.summary = doc.summary; + report.citations = registry.citations; + report.sections = doc.sections; + if (!Object.keys(registry.citations).length) warnings.push("the answer cites no source"); + return finish(report, warnings); +} diff --git a/common/lib/report/convertManifest.ts b/common/lib/report/convertManifest.ts @@ -0,0 +1,552 @@ +// A report-to-video MANIFEST ↔ `report.json`, both ways. The video and the +// report page can be cut from one source: a manifest becomes a report, and a +// report becomes a starter manifest to refine in umtool. +// +// The manifest is umtool/report-to-video's (its README, "Manifest shape"). +// What maps to what: +// +// manifest report +// ──────────────────────────────────────── ────────────────────────────────────── +// slug, title, subtitle id, title, subtitle +// card (style "chapter", or none) a section: heading → title, sub → body +// clip (channel ?? provenance.channelSlug, a video citation, id = the clip's id; +// video, cutStart ?? start, an `audioOnly` clip an audio one +// cutEnd ?? end, quote, date, note) +// image (src, or its first panel; quote, a source citation (its still = src) +// title, date, citeUrl, note) and a source of its own +// posts[] (siteChannel, its id, text, date) a post citation, id = the post's id +// `claim: { id, verdict }` on entries a claim with that verdict; its first +// image is its source sentence +// ledger[] (id, quote, label, date, a claim with no verdict (text = the +// entryId, channel, video, cite) quote, title = the label) +// +// An entry that carries no claim is cited in its section's body, one line +// each. A post goes with the clip it is attached to (`attachTo`, else the +// clip it follows by date, else the first) — under that clip's claim, or in +// its section's body. A claim's text is its ledger entry's quote, else the +// heading of a card that carries it, else its source sentence, else its first +// clip's quote. A report with any verdict is a fact-check, else a sweep. +// +// Left out, with a warning: a clip with its own media (`src`: no record to +// cite), an entry without a quote, a post no archive channel is known to keep. +// A manifest has more than a report (render settings, holds, the deck, cut +// edits, redactions, the ledger's adjudication) and a report more than a +// manifest (findings, a claim's own title, citation pads, labels, speakers); +// neither survives the trip. Image paths are written as the manifest has +// them: relative to ITS directory (the CLI says so when the report is written +// elsewhere). +// +// A report → a STARTER manifest: no cold open — a title card, then per section +// a chapter card, its body's citations in order, and per claim its source +// sentence's still (an `image`) and its video and audio citations (`clip`s), +// each carrying `claim: { id, verdict }` when the claim has a verdict, and its +// posts in `posts[]` attached to the claim's last clip. Spans are written as +// the report has them, so `resolve-windows.mjs` and the clip bench start from +// the cited sentences. A post needs its platform, its link and its date: from +// `postOf` (the CLI reads the channel's archive) or, for a numeric X id, its +// `x.com/i/status/` link and the citation's date; one that has none of them is +// left out with a warning. + +import type { Citation, Source, SourceCitation } from "../citations/schema"; +import { isHttpUrl, isPartialDate } from "../citations/validate"; +import { + IdAllocator, + emptyReport, + finish, + citedIds, + reportIdOf, + spanAt, + type CitationMap, + type ConvertContext, + type Converted, +} from "./convertShared"; +import type { Claim, Report, Section } from "./schema"; +import { isVerdict, type Verdict } from "./verdicts"; +import { parseReport } from "./validate"; +import type { Problem } from "../citations/validate"; + +type Obj = Record<string, unknown>; + +const isObj = (v: unknown): v is Obj => !!v && typeof v === "object" && !Array.isArray(v); +const str = (v: unknown): string | undefined => (typeof v === "string" && v.trim() ? v : undefined); +const num = (v: unknown): number | undefined => (typeof v === "number" && Number.isFinite(v) ? v : undefined); + +export type ManifestConvertOptions = ConvertContext & { id?: string; title?: string }; + +// A manifest entry's claim, when it carries a sound one. +function claimOfEntry(e: Obj): { id: string; verdict: Verdict } | null { + if (!isObj(e.claim) || e.type === "teaser") return null; + const id = str(e.claim.id)?.trim(); + return id && isVerdict(e.claim.verdict) ? { id, verdict: e.claim.verdict } : null; +} + +// The id a post has on its platform: `postId`, else the link's. +export function postNativeId(post: Obj): string | null { + const given = str(post.postId); + if (given && /^[A-Za-z0-9_-]{1,128}$/.test(given)) return given; + try { + const u = new URL(String(post.url ?? "")); + const m = /\/post\/([A-Za-z0-9]+)\/?$/.exec(u.pathname) ?? /\/status(?:es)?\/(\d+)(?:\/|$)/.exec(u.pathname); + return m ? m[1] : null; + } catch { + return null; + } +} + +const dayOf = (d: unknown) => (typeof d === "string" && /^\d{4}-\d{2}-\d{2}/.test(d) ? d.slice(0, 10) : null); + +// The clip a post goes with, as the build attaches it: `attachTo`, else the +// clip whose date most closely precedes the post's (ties to the later in the +// cut), else the first. +function clipForPost(post: Obj, clips: Obj[]): Obj | null { + const named = str(post.attachTo); + if (named) { + const c = clips.find((e) => e.id === named); + if (c) return c; + } + const day = dayOf(post.date); + let best: Obj | null = null; + if (day) { + for (const c of clips) { + const cd = dayOf(c.date); + if (cd && cd <= day && (!best || cd >= (dayOf(best.date) as string))) best = c; + } + } + return best ?? clips[0] ?? null; +} + +export async function manifestToReport(manifest: unknown, opts: ManifestConvertOptions = {}): Promise<Converted> { + const warnings: string[] = []; + const m: Obj = isObj(manifest) ? manifest : {}; + const timeline = (Array.isArray(m.timeline) ? m.timeline : []).filter(isObj); + const ledger = (Array.isArray(m.ledger) ? m.ledger : []).filter(isObj); + const posts = (Array.isArray(m.posts) ? m.posts : []).filter(isObj); + const prov = isObj(m.provenance) ? m.provenance : {}; + const title = opts.title ?? str(m.title) ?? str(m.slug) ?? "Report"; + const report = emptyReport(reportIdOf(opts.id ?? str(m.slug), title), "sweep", title); + if (str(m.subtitle)) report.subtitle = m.subtitle as string; + + const citations: CitationMap = {}; + const sources: Record<string, Source> = {}; + const citeIds = new IdAllocator(); + const anchors = new IdAllocator(); + const where = (e: Obj, i: number) => `timeline[${i}] ${str(e.id) ?? "?"}`; + + // ── The citations every entry is ── + const citeOf = new Map<Obj, string>(); + for (const [i, e] of timeline.entries()) { + if (e.type === "clip") { + if (str(e.src)) { + warnings.push(`${where(e, i)}: a clip of its own media (src) cites no record — left out`); + continue; + } + const channel = str(e.channel) ?? str(prov.channelSlug); + const video = str(e.video); + const start = num(e.cutStart) ?? num(e.start); + const end = num(e.cutEnd) ?? num(e.end); + const quote = str(e.quote); + if (!channel || !video || start === undefined || end === undefined) { + warnings.push(`${where(e, i)}: a clip needs a channel, a video, a start and an end — left out`); + continue; + } + if (!quote) { + warnings.push(`${where(e, i)}: a clip without a quote — left out (a citation's quote is verbatim)`); + continue; + } + const cid = citeIds.claim(str(e.id) ?? "c", "c"); + const c: Citation = { kind: e.audioOnly === true ? "audio" : "video", channel, id: video, start, end, quote }; + if (str(e.date) && isPartialDate(e.date as string)) c.date = e.date as string; + if (str(e.note)) c.note = e.note as string; + citations[cid] = c; + citeOf.set(e, cid); + } else if (e.type === "image") { + const panels = Array.isArray(e.panels) ? e.panels.filter(isObj) : []; + const src = str(e.src) ?? str(panels[0]?.src); + if (!src) { + warnings.push(`${where(e, i)}: an image with no src — left out`); + continue; + } + if (panels.length > 1) warnings.push(`${where(e, i)}: ${panels.length} panels — the citation's still is the first`); + let quote = str(e.quote); + if (!quote) { + quote = str(e.title) ?? str(e.id) ?? src; + warnings.push(`${where(e, i)}: an image without a quote — its citation quotes its title; give it the still's words`); + } + const cid = citeIds.claim(str(e.id) ?? "a", "a"); + const sid = citeIds.claim(`src-${cid}`); + const source: Source = { kind: "other", title: str(e.title) ?? cid }; + if (str(e.citeUrl) && isHttpUrl(e.citeUrl as string)) source.url = e.citeUrl as string; + if (str(e.date) && isPartialDate(e.date as string)) source.date = e.date as string; + sources[sid] = source; + const c: SourceCitation = { kind: "source", source: sid, quote, image: src }; + if (str(e.note)) c.note = e.note as string; + citations[cid] = c; + citeOf.set(e, cid); + } + } + + // ── Posts: each a citation, and the clip it goes with ── + const clips = timeline.filter((e) => e.type === "clip" && citeOf.has(e)); + const postsOfClip = new Map<Obj, string[]>(); + const loosePosts: string[] = []; + for (const [i, p] of posts.entries()) { + if (p.hide === true) continue; + const native = postNativeId(p); + const platform = p.platform === "bluesky" ? "bluesky" : "x"; + const channel = + str(p.siteChannel) ?? + (native && opts.postChannelOf ? await opts.postChannelOf({ platform, id: native, handle: str(p.handle) }) : null); + const text = str(p.text); + if (!native || !channel || !text) { + warnings.push( + `posts[${i}] ${str(p.id) ?? "?"}: ${!native ? "no id on its platform" : !text ? "no text" : "no archive channel known to keep it (siteChannel)"} — left out`, + ); + continue; + } + const cid = citeIds.claim(str(p.id) ?? "p", "p"); + citations[cid] = { + kind: "post", + channel, + id: native, + quote: text, + ...(str(p.date) && isPartialDate(p.date as string) ? { date: p.date as string } : {}), + }; + const clip = clipForPost(p, clips); + if (clip) postsOfClip.set(clip, [...(postsOfClip.get(clip) ?? []), cid]); + else loosePosts.push(cid); + } + + // ── Ledger claims: what the ledger says a claim is, and the clip it pins ── + const ledgerById = new Map<string, Obj>(); + for (const l of ledger) if (str(l.id)) ledgerById.set(l.id as string, l); + const timelineClaims = new Set(timeline.map(claimOfEntry).filter((c) => c).map((c) => c!.id)); + const ledgerClaimOfClip = new Map<string, string>(); + for (const l of ledger) { + const id = str(l.id); + const pinned = str(l.entryId); + if (id && pinned && !timelineClaims.has(id)) ledgerClaimOfClip.set(pinned, id); + } + + // ── The walk: sections, claims, bodies ── + const sections: { section: Section; body: string[] }[] = []; + const claims = new Map<string, { claim: Claim; texts: { from: string; text: string }[] }>(); + const current = () => { + if (!sections.length) { + sections.push({ section: { id: anchors.claim("opening"), title: "Opening", claims: [] }, body: [] }); + } + return sections[sections.length - 1]; + }; + const claimFor = (id: string, verdict: Verdict | null) => { + let k = claims.get(id); + if (!k) { + const claim: Claim = { id: anchors.claim(id, "k"), text: "", citations: [] }; + if (verdict) claim.verdict = verdict; + k = { claim, texts: [] }; + claims.set(id, k); + current().section.claims!.push(claim); + } + return k; + }; + const list = (claim: Claim, cid: string) => { + if (!claim.citations!.includes(cid)) claim.citations!.push(cid); + }; + const bodyLine = (cid: string) => { + const c = citations[cid]; + const quote = c.quote.replace(/\s+/g, " ").trim(); + const label = + c.kind === "post" ? "post" : c.kind === "source" ? (sources[c.source]?.title ?? cid) : c.kind === "page" ? cid : `${c.id}`; + return `- “${quote}” [${label.replace(/[[\]]/g, "")}](cite:${cid})`; + }; + + for (const e of timeline) { + const tagged = claimOfEntry(e); + if (e.type === "card") { + if (tagged) { + const k = claimFor(tagged.id, tagged.verdict); + if (str(e.heading)) k.texts.push({ from: "card", text: e.heading as string }); + continue; + } + const style = e.style; + if ((style === undefined || style === "chapter") && str(e.heading)) { + const section: Section = { id: anchors.claim(str(e.id) ?? "section", "section"), title: e.heading as string, claims: [] }; + sections.push({ section, body: str(e.sub) ? [e.sub as string] : [] }); + } + continue; + } + const cid = citeOf.get(e); + if (!cid) continue; + const ledgerClaim = !tagged && str(e.id) ? ledgerClaimOfClip.get(e.id as string) : undefined; + const claimId = tagged?.id ?? ledgerClaim; + if (claimId) { + const k = claimFor(claimId, tagged?.verdict ?? null); + const c = citations[cid]; + if (c.kind === "source" && !k.claim.sourceQuote) { + k.claim.sourceQuote = { citation: cid }; + k.texts.push({ from: "source", text: c.quote }); + } else { + list(k.claim, cid); + if (c.kind !== "source") k.texts.push({ from: "clip", text: c.quote }); + } + for (const pid of postsOfClip.get(e) ?? []) list(k.claim, pid); + } else { + const body = current().body; + body.push(bodyLine(cid)); + for (const pid of postsOfClip.get(e) ?? []) body.push(bodyLine(pid)); + } + } + + // Ledger claims no clip pins: their own moment, in a section of their own. + const unpinned = ledger.filter((l) => { + const id = str(l.id); + return id && !claims.has(id) && !timelineClaims.has(id); + }); + if (unpinned.length) { + sections.push({ section: { id: anchors.claim("ledger"), title: "Not clipped", claims: [] }, body: [] }); + for (const l of unpinned) { + const k = claimFor(l.id as string, null); + const channel = str(l.channel) ?? str(prov.channelSlug); + const video = str(l.video); + const second = num(l.cite) ?? num(l.start); + const quote = str(l.quote); + if (channel && video && second !== undefined && quote) { + const span = await spanAt(channel, video, second, quote, opts, num(l.end)); + const cid = citeIds.claim(`l-${l.id as string}`); + citations[cid] = { kind: "video", channel, id: video, start: span.start, end: span.end, quote }; + list(k.claim, cid); + } else { + warnings.push(`ledger ${l.id as string}: no channel, video, second and quote to cite — the claim has no citation`); + } + } + } + if (loosePosts.length) { + const body = current().body; + for (const pid of loosePosts) body.push(bodyLine(pid)); + } + + // ── Claim texts and titles ── + for (const [id, { claim, texts }] of claims) { + const l = ledgerById.get(id); + const pick = (from: string) => texts.find((t) => t.from === from)?.text; + const text = str(l?.quote) ?? pick("card") ?? pick("source") ?? pick("clip"); + if (str(l?.label) && str(l?.quote)) claim.title = l!.label as string; + if (text) claim.text = text.replace(/\s+/g, " ").trim(); + else { + claim.text = id; + warnings.push(`claim ${id}: nothing in the manifest states it — its text is its id`); + } + if (!claim.citations!.length) delete claim.citations; + } + + report.kind = [...claims.values()].some((k) => k.claim.verdict) ? "factcheck" : "sweep"; + report.sources = sources; + report.citations = citations; + report.sections = sections.map(({ section, body }) => { + const out: Section = { id: section.id, title: section.title }; + if (body.length) out.body = body.join("\n"); + if (section.claims!.length) out.claims = section.claims; + return out; + }); + return finish(report, warnings); +} + +// ─── report → manifest ─── + +// What a post record says that a post citation does not (the CLI reads it off +// the channel's archive). +export type PostRecord = { + platform: "x" | "bluesky"; + url: string; + createdAt?: string; + author?: string; + authorName?: string; +}; + +export type PostOf = (channel: string, id: string) => Promise<PostRecord | null>; + +export type ManifestOptions = { + // The archive the cut's QR codes link (provenance.siteOrigin). Default "" + // — which `umtool check` blocks on until it is set, as for `umtool new`. + siteOrigin?: string; + postOf?: PostOf; + // Replaces the starter `render` block. + render?: Record<string, unknown>; + // `generatedOn`; default today. + today?: string; +}; + +// The render block `umtool new` writes (umtool/lib/projects/scaffold.mjs +// `skeleton`), so a converted report renders as a new project does. Keep the +// two alike. +export const STARTER_RENDER: Readonly<Record<string, unknown>> = Object.freeze({ + width: 1920, + height: 1080, + fps: 30, + audioRate: 48000, + audioChannels: 2, + maxHeightSource: 1080, + fontRegular: "/usr/share/fonts/TTF/FiraSans-Regular.ttf", + fontBold: "/usr/share/fonts/TTF/FiraSans-Bold.ttf", + palette: { bg: "#12100c", fg: "#f6f1e6", muted: "#a2957f", accent: "#c8752a", amber: "#ffc860" }, + transition: 0.4, + fetchPad: 3, + snapWindow: 1.6, + silenceMinDur: 0.09, + silenceRelDb: 6, + headerHeight: 56, + footerHeight: 0, + crf: 21, + preset: "slow", + qr: { scale: 4, quiet: 3, ecc: "M", margin: 28 }, +}); + +// A manifest entry's `date` is a calendar day. +const DAY_RE = /^\d{4}-\d{2}-\d{2}$/; + +export const CARD_SECONDS = { title: 4, chapter: 3.5 } as const; +export const IMAGE_SECONDS = 6; + +export type ManifestResult = + | { ok: true; manifest: Record<string, unknown>; warnings: string[] } + | { ok: false; problems: Problem[] }; + +export async function reportToManifest(raw: unknown, opts: ManifestOptions = {}): Promise<ManifestResult> { + const parsed = parseReport(raw); + if (!parsed.ok || parsed.problems.length) return { ok: false, problems: parsed.problems }; + const report: Report = parsed.value; + const warnings: string[] = []; + const citations = report.citations ?? {}; + const sources = report.sources ?? {}; + // Timeline ids are file names in a build: letters, digits, `_` and `-`. + const entryIds = new IdAllocator(/[^A-Za-z0-9_-]+/g); + const postIds = new IdAllocator(/[^A-Za-z0-9_-]+/g); + const timeline: Record<string, unknown>[] = []; + const posts: Record<string, unknown>[] = []; + const postsDone = new Set<string>(); + const channels = new Map<string, number>(); + let lastClip: string | null = null; + + timeline.push({ + type: "card", + id: entryIds.claim("title"), + style: "title", + seconds: CARD_SECONDS.title, + heading: report.title, + sub: report.subtitle ?? "", + }); + + const claimTag = (claim: Claim | null) => (claim?.verdict ? { claim: { id: claim.id, verdict: claim.verdict } } : {}); + + const emit = async (cid: string, claim: Claim | null, sourceSentence = false) => { + const c = citations[cid]; + if (!c) return; + switch (c.kind) { + case "video": + case "audio": { + const id = entryIds.claim(cid, "c"); + timeline.push({ + type: "clip", + id, + channel: c.channel, + video: c.id, + start: c.start, + end: c.end, + cite: Math.floor(c.start), + quote: c.quote, + ...(c.kind === "audio" ? { audioOnly: true } : {}), + ...(c.date && DAY_RE.test(c.date) ? { date: c.date } : {}), + ...(c.note ? { note: c.note } : {}), + ...claimTag(claim), + }); + channels.set(c.channel, (channels.get(c.channel) ?? 0) + 1); + lastClip = id; + return; + } + case "source": { + if (!c.image) { + warnings.push( + `${cid}: ${sourceSentence ? "a claim's source sentence" : "a source citation"} with no still — no image entry (shoot it, then set its image)`, + ); + return; + } + const source = sources[c.source]; + const dayOfSource = [c.date, source?.date].find((d) => d && DAY_RE.test(d)); + timeline.push({ + type: "image", + id: entryIds.claim(cid, "a"), + src: c.image, + seconds: IMAGE_SECONDS, + title: source?.title ?? cid, + quote: c.quote, + ...(dayOfSource ? { date: dayOfSource } : {}), + ...(source?.url ? { citeUrl: source.url } : {}), + ...(c.note ? { note: c.note } : {}), + ...claimTag(claim), + }); + return; + } + case "post": { + if (postsDone.has(cid)) return; + postsDone.add(cid); + const rec = opts.postOf ? await opts.postOf(c.channel, c.id) : null; + const platform = rec?.platform ?? (/^\d+$/.test(c.id) ? "x" : null); + const url = rec?.url ?? (platform === "x" ? `https://x.com/i/status/${c.id}` : null); + const date = c.date ?? rec?.createdAt; + if (!platform || !url || !date || !/^\d{4}-\d{2}-\d{2}/.test(date)) { + warnings.push( + `${cid}: post ${c.channel}/${c.id} — ${!platform || !url ? "its platform and link are unknown (give a channels dir)" : "no date"} — left out of posts`, + ); + return; + } + posts.push({ + id: postIds.claim(cid, "p"), + platform, + ...(rec?.authorName ? { author: rec.authorName } : {}), + ...(rec?.author ? { handle: rec.author } : {}), + date, + text: c.quote.slice(0, 3000), + url, + attachTo: lastClip, + siteChannel: c.channel, + ...(/^[A-Za-z0-9_-]{1,128}$/.test(c.id) ? { postId: c.id } : {}), + }); + return; + } + case "page": + warnings.push(`${cid}: a page citation has no manifest entry — left out`); + return; + } + }; + + for (const section of report.sections) { + timeline.push({ + type: "card", + id: entryIds.claim(section.id, "section"), + style: "chapter", + seconds: CARD_SECONDS.chapter, + heading: section.title, + }); + for (const cid of citedIds(section.body)) await emit(cid, null); + for (const claim of section.claims ?? []) { + if (claim.sourceQuote) await emit(claim.sourceQuote.citation, claim, true); + const order = [...(claim.citations ?? []), ...citedIds(claim.findings)]; + for (const cid of [...new Set(order)]) { + if (cid !== claim.sourceQuote?.citation) await emit(cid, claim); + } + } + } + + const channelSlug = [...channels.entries()].sort((a, b) => b[1] - a[1])[0]?.[0] ?? ""; + const manifest: Record<string, unknown> = { + schemaVersion: 1, + slug: report.id, + title: report.title, + subtitle: report.subtitle ?? "", + generatedOn: opts.today ?? new Date().toISOString().slice(0, 10), + provenance: { siteOrigin: opts.siteOrigin ?? "", channelSlug, channel: "" }, + render: { ...(opts.render ?? STARTER_RENDER) }, + timelineNodes: [], + timeline, + ...(posts.length ? { posts } : {}), + }; + return { ok: true, manifest, warnings }; +} diff --git a/common/lib/report/convertShared.ts b/common/lib/report/convertShared.ts @@ -0,0 +1,628 @@ +// THE CONVERTERS' SHARED HALF — what bringing a cited document from elsewhere +// into the citation model takes, whatever the document was: +// +// - reading a link a report cites with (`parseCitationHref`): an archive +// viewer moment (`<origin>/?v=<channel>%2F<id>&t=<s>`, what /sweep and +// /ask write), an archive post (`…&vm=post`), a moment page +// (`/m/<key>/`), or a post on its own platform (x.com, bsky.app); +// - turning ONE cited second into a span (`spanAt`): a record's cues, when +// the caller can read them, widened to whole sentences by the one widening +// (lib/cueWiden.mjs); else the second plus a default span; +// - the markdown engine (`structureMarkdown`): a document's headings into +// sections, its list items and paragraphs that carry a citation into +// claims, everything else into the sections' bodies, and every citing link +// rewritten to `[label](cite:<id>)`; +// - ids, and the finished report checked by the report validator. +// +// The converters themselves are ./convertSweep.ts, ./convertAsk.ts and +// ./convertManifest.ts. Nothing here reads the disk: the cues and the channel +// that keeps a post arrive as callbacks (./convert-server.ts has the disk +// ones), so every converter runs in a test on literals. + +import { widen } from "../cueWiden.mjs"; +import { MAX_CITATION_SPAN_SECONDS, REF_ID_RE, type Citation } from "../citations/schema"; +import { citeHref, extractCiteRefs } from "../citations/inline"; +import { parseMomentPath, roundMomentSeconds } from "../citations/moments"; +import { isPartialDate, type Problem } from "../citations/validate"; +import { REPORT_FORMAT, REPORT_ID_RE, REPORT_VERSION, type Claim, type Report, type Section } from "./schema"; +import { parseReport } from "./validate"; + +export type Cue = { start: number; end: number; text: string }; + +// A record's cues, sorted by start, or null when they cannot be read. +export type CuesOf = (channel: string, id: string) => Promise<readonly Cue[] | null>; + +export type PostPlatformName = "x" | "bluesky"; + +// The archive channel that keeps a post, or null. +export type PostChannelOf = (post: { platform: PostPlatformName; id: string; handle?: string }) => Promise<string | null>; + +export type ConvertContext = { + cuesOf?: CuesOf; + postChannelOf?: PostChannelOf; + // A cited second with no cues to read becomes [second, second + this]. + spanSeconds?: number; +}; + +// What a converter hands back: the report, what it had to guess or leave out +// (warnings), and the report validator's problems — a report with problems is +// not to be written. +export type Converted = { report: Report; warnings: string[]; problems: Problem[] }; + +// A cited second with no cues is this long. A /sweep or /ask citation carries +// one second, never an end; ten seconds is about a spoken sentence. +export const DEFAULT_SPAN_SECONDS = 10; + +// ─── Links ─── + +export type ParsedCitationHref = + // An archive viewer moment, or a moment page: `seconds` is where it starts; + // a moment page also carries its end. + | { kind: "span"; channel: string; id: string; seconds: number; end?: number } + // An archived post on the archive. + | { kind: "post"; channel: string; id: string } + // A post on its own platform: the archive channel is not in the link. + | { kind: "post-original"; platform: PostPlatformName; id: string; handle?: string }; + +const X_HOSTS = new Set(["x.com", "twitter.com", "mobile.twitter.com", "www.x.com", "www.twitter.com", "mobile.x.com"]); +const BSKY_HOSTS = new Set(["bsky.app", "www.bsky.app"]); + +// What a link cites, or null for a link that cites nothing the model knows (a +// platform's video page, an article, a relative link). +export function parseCitationHref(href: string): ParsedCitationHref | null { + let u: URL; + try { + u = new URL(href.trim()); + } catch { + return null; + } + if (u.protocol !== "http:" && u.protocol !== "https:") return null; + const host = u.hostname.toLowerCase(); + + if (X_HOSTS.has(host)) { + const m = /^\/([^/]+)\/status(?:es)?\/(\d+)(?:\/|$)/.exec(u.pathname); + if (!m) return null; + return { kind: "post-original", platform: "x", id: m[2], ...(m[1] !== "i" ? { handle: m[1] } : {}) }; + } + if (BSKY_HOSTS.has(host)) { + const m = /^\/profile\/([^/]+)\/post\/([A-Za-z0-9]+)\/?$/.exec(u.pathname); + if (!m) return null; + return { kind: "post-original", platform: "bluesky", id: m[2], handle: m[1] }; + } + + const moment = parseMomentPath(u.pathname); + if (moment) { + return moment.kind === "span" + ? { kind: "span", channel: moment.channel, id: moment.id, seconds: moment.start, end: moment.end } + : { kind: "post", channel: moment.channel, id: moment.id }; + } + + // The viewer: `v` is `<channel>/<id>` (URL-encoded or not), `t` whole seconds. + const v = u.searchParams.get("v"); + if (!v) return null; + const slash = v.indexOf("/"); + if (slash <= 0 || slash === v.length - 1) return null; + const channel = v.slice(0, slash); + const id = v.slice(slash + 1); + if (u.searchParams.get("vm") === "post") return { kind: "post", channel, id }; + const t = u.searchParams.get("t"); + const seconds = t && /^\d+(\.\d+)?$/.test(t) ? Number(t) : 0; + return { kind: "span", channel, id, seconds }; +} + +// ─── Spans ─── + +const EPS = 0.02; + +const wordCount = (s: string) => s.split(/\s+/).filter((w) => /[\p{L}\p{N}]/u.test(w)).length; + +const round2 = (n: number) => roundMomentSeconds(n); + +// The cue a cited second points at: the viewer's `t` is a cue's start, +// floored, so the cue that starts in [t, t + 1) is the one cited; else the cue +// that holds t; else the nearest one after it. +function cueIndexAt(cues: readonly Cue[], t: number): number { + const starting = cues.findIndex((c) => c.start >= t - EPS && c.start < t + 1); + if (starting >= 0) return starting; + const holding = cues.findIndex((c) => c.start <= t + EPS && c.end > t); + if (holding >= 0) return holding; + const after = cues.findIndex((c) => c.start > t); + return after >= 0 ? after : cues.length - 1; +} + +export type Span = { start: number; end: number; from: "cues" | "default" | "link" }; + +// One cited second as a span. With the record's cues: the cue it cites, run +// on cue by cue until it holds as many words as the quote, then widened to +// whole sentences — and if the widened span is longer than a citation may be, +// the unwidened cue run, and failing that the default. Without cues (or past +// their end): [second, second + spanSeconds]. An `end` the link already gave +// (a moment page) is kept as it is. +export async function spanAt( + channel: string, + id: string, + seconds: number, + quote: string, + ctx: ConvertContext, + end?: number, +): Promise<Span> { + if (end !== undefined && end > seconds) return { start: round2(seconds), end: round2(end), from: "link" }; + const fallback: Span = { + start: round2(seconds), + end: round2(seconds + (ctx.spanSeconds ?? DEFAULT_SPAN_SECONDS)), + from: "default", + }; + const cues = ctx.cuesOf ? await ctx.cuesOf(channel, id) : null; + if (!cues || cues.length === 0 || seconds > cues[cues.length - 1].end) return fallback; + const i = cueIndexAt(cues, seconds); + const want = Math.max(1, wordCount(quote)); + let j = i; + let have = wordCount(cues[i].text); + while (have < want && j + 1 < cues.length && cues[j + 1].end - cues[i].start <= MAX_CITATION_SPAN_SECONDS) { + j += 1; + have += wordCount(cues[j].text); + } + const fits = (s: number, e: number) => e > s && e - s <= MAX_CITATION_SPAN_SECONDS && round2(e) > round2(s); + const w = widen(cues, cues[i].start, cues[j].end); + if (fits(w.start, w.end)) return { start: round2(w.start), end: round2(w.end), from: "cues" }; + if (fits(cues[i].start, cues[j].end)) return { start: round2(cues[i].start), end: round2(cues[j].end), from: "cues" }; + return fallback; +} + +// ─── Ids ─── + +// Unique ids in one namespace, each a reference id (lib/citations/schema.ts +// REF_ID_RE): a wanted id is kept when it is free and sound, else made sound +// and suffixed `-2`, `-3`, …. +export class IdAllocator { + private readonly taken = new Set<string>(); + private readonly counters = new Map<string, number>(); + + constructor(private readonly charset: RegExp = /[^A-Za-z0-9_.:-]+/g) {} + + has(id: string): boolean { + return this.taken.has(id); + } + + claim(wanted: string, fallback = "x"): string { + let base = wanted.replace(this.charset, "-").replace(/^[^A-Za-z0-9]+/, "").slice(0, 56); + if (!base) base = fallback; + let id = base; + for (let n = 2; this.taken.has(id) || !REF_ID_RE.test(id); n++) id = `${base}-${n}`; + this.taken.add(id); + return id; + } + + // The next free `<prefix>01`, `<prefix>02`, …. + next(prefix: string): string { + let n = this.counters.get(prefix) ?? 0; + let id: string; + do { + n += 1; + id = `${prefix}${String(n).padStart(2, "0")}`; + } while (this.taken.has(id)); + this.counters.set(prefix, n); + this.taken.add(id); + return id; + } +} + +// A heading or a title as an id: lowercase words joined by `-`. +export function slugOf(text: string, max = 48): string { + return text + .normalize("NFKD") + .replace(/[̀-ͯ]/g, "") + .toLowerCase() + .replace(/[^a-z0-9]+/g, "-") + .replace(/^-+|-+$/g, "") + .slice(0, max) + .replace(/-+$/, ""); +} + +// A report id (a lowercase slug, REPORT_ID_RE) from what the caller asked for, +// else from the title. +export function reportIdOf(wanted: string | undefined, title: string): string { + if (wanted !== undefined) return wanted; + const slug = slugOf(title, 64); + return REPORT_ID_RE.test(slug) ? slug : "report"; +} + +// ─── Markdown ─── + +// `[label](href)`: a label may hold one level of brackets (a title with +// `[live]` in it); the href runs to the first space or `)`, optionally in +// `<…>`, optionally followed by a quoted title. +const LINK_RE = /\[((?:[^[\]]|\[[^[\]]*\])*)\]\(\s*<?([^)\s>]+)>?(?:\s+"[^"]*")?\s*\)/g; +const CODE_SPAN_RE = /(?<!`)(`+)(?!`)[\s\S]*?(?<!`)\1(?!`)/g; +const HEADING_RE = /^ {0,3}(#{1,6})\s+(.*?)\s*#*\s*$/; +const FENCE_RE = /^ {0,3}(`{3,}|~{3,})/; +const LIST_ITEM_RE = /^ ?([-*+]|\d{1,9}[.)])\s+/; +const QUOTE_LINE_RE = /^ {0,3}>/; +const RULE_RE = /^ {0,3}([-*_])(\s*\1){2,}\s*$/; + +type MdHeading = { kind: "heading"; level: number; text: string }; +type MdBlock = { kind: "block"; md: string; code: boolean }; +type MdPart = MdHeading | MdBlock; + +// The markdown as headings and blocks: a fenced block, a top-level list item +// (with its continuation and nested lines), a run of quote lines (and the +// lines that lazily continue it), or a paragraph. Blank lines and rules +// separate them. +function partsOf(md: string): MdPart[] { + const out: MdPart[] = []; + const lines = md.replace(/\r\n?/g, "\n").split("\n"); + let cur: string[] = []; + let curQuote = false; + const flush = () => { + if (cur.length) out.push({ kind: "block", md: cur.join("\n"), code: false }); + cur = []; + curQuote = false; + }; + for (let i = 0; i < lines.length; i++) { + const line = lines[i]; + const fence = FENCE_RE.exec(line); + if (fence) { + flush(); + const close = new RegExp(`^ {0,3}${fence[1][0] === "`" ? "`" : "~"}{${fence[1].length},}\\s*$`); + const body = [line]; + while (++i < lines.length) { + body.push(lines[i]); + if (close.test(lines[i])) break; + } + out.push({ kind: "block", md: body.join("\n"), code: true }); + continue; + } + const heading = HEADING_RE.exec(line); + if (heading) { + flush(); + out.push({ kind: "heading", level: heading[1].length, text: heading[2] }); + continue; + } + if (!/\S/.test(line) || RULE_RE.test(line)) { + flush(); + continue; + } + if (LIST_ITEM_RE.test(line)) { + flush(); + cur.push(line); + continue; + } + const quote = QUOTE_LINE_RE.test(line); + if (quote && cur.length && !curQuote) flush(); + if (quote) curQuote = true; + cur.push(line); + } + flush(); + return out; +} + +// Markdown as the plain text a claim is: no list or quote markers, no +// emphasis or code ticks, a link as its label, one line. +export function plainText(md: string): string { + return md + .split("\n") + .map((l) => l.replace(LIST_ITEM_RE, "").replace(/^ {0,3}(>\s?)+/, "")) + .join(" ") + .replace(LINK_RE, (_m, label: string) => label) + .replace(/(\*\*|__|\*|_|`)(?=\S)([\s\S]*?\S)\1/g, "$2") + .replace(/\s+/g, " ") + .trim(); +} + +// The quoted words in a stretch of text: every “…” or "…" run, joined by an +// ellipsis (a report quotes one passage as `"A" … "B"`). +export function quotedIn(text: string): string | null { + const runs = [...text.matchAll(/“([^”]+)”|"([^"]+)"/g)] + .map((m) => (m[1] ?? m[2]).trim()) + .filter((q) => wordCount(q) > 0); + return runs.length ? runs.join(" … ") : null; +} + +const TRIM_SEP_RE = /^[\s—–\-:;,|]+|[\s—–\-:;,|(]+$/g; + +// What a citing link resolves to: the id of the citation it now names, or null +// to leave the link as it is. +export type CiteResolver = (link: { + href: string; + label: string; + // The quoted words nearest before the link in its block, else in the block + // before it (a quote line followed by its citation line), else the block's + // plain text. + quote: string; +}) => Promise<string | null>; + +export type StructuredMarkdown = { + title: string | null; + summary: string; + sections: Section[]; +}; + +type Blk = { md: string; text: string; ids: string[]; interleaved: boolean }; + +function stripMarkers(md: string): string { + const lines = md.split("\n"); + const allQuoted = lines.every((l) => QUOTE_LINE_RE.test(l) || !/\S/.test(l)); + return lines + .map((l, i) => { + let s = i === 0 ? l.replace(LIST_ITEM_RE, "") : l.replace(/^ {2,4}/, ""); + if (allQuoted) s = s.replace(/^ {0,3}>\s?/, ""); + return s; + }) + .join("\n") + .trim(); +} + +// One block with its citing links rewritten. A link inside a code span is +// text, as it is to lib/citations/inline.ts. +async function rewriteBlock(md: string, previousText: string, resolve: CiteResolver): Promise<Blk> { + const scan = md.replace(CODE_SPAN_RE, (s) => s.replace(/[^\n]/g, " ")); + const links = [...scan.matchAll(LINK_RE)]; + const ownText = plainText(md.replace(LINK_RE, "")).replace(TRIM_SEP_RE, ""); + const ids: string[] = []; + let out = ""; + let last = 0; + const withoutLinks: string[] = []; + let between = 0; + for (const m of links) { + const start = m.index; + const end = start + m[0].length; + const label = md.slice(start + 1, start + 1 + m[1].length); + const href = m[2]; + const before = plainText(md.slice(last, start)); + const quote = + quotedIn(md.slice(last, start)) ?? + quotedIn(md.slice(0, start)) ?? + quotedIn(md.slice(end)) ?? + (ownText || + quotedIn(previousText) || + previousText.replace(TRIM_SEP_RE, "") || + label); + const id = await resolve({ href, label, quote }); + out += md.slice(last, start); + withoutLinks.push(md.slice(last, start)); + if (id) { + if (ids.length && before.replace(TRIM_SEP_RE, "")) between += 1; + ids.push(id); + out += `[${label}](${citeHref(id)})`; + } else { + out += md.slice(start, end); + withoutLinks.push(label); + } + last = end; + } + out += md.slice(last); + withoutLinks.push(md.slice(last)); + const text = plainText(withoutLinks.join("")).replace(/\s+([.,;:!?])/g, "$1").replace(TRIM_SEP_RE, ""); + return { md: out, text, ids: [...new Set(ids)], interleaved: between > 0 }; +} + +// The engine. `sectionLevel` headings (the shallowest below the title) open +// sections; the first `#` heading, when it comes before any section, is the +// title; deeper headings stay in the body. In a section, a block that cites +// becomes a claim — a block that is ONLY citations (a `— [title @ 1:02](…)` +// line under a quote) lends them to the block before it, which becomes the +// claim — and every other block joins the section's body. Before the first +// section, everything is the summary (its links rewritten, no claims). With no +// section headings at all, everything after the title is one section, +// `defaultSection`. +export async function structureMarkdown( + md: string, + resolve: CiteResolver, + { defaultSection }: { defaultSection: string }, +): Promise<StructuredMarkdown> { + const parts = partsOf(md); + let title: string | null = null; + const firstHeading = parts.findIndex((p) => p.kind === "heading"); + const firstBlock = parts.findIndex((p) => p.kind === "block"); + if (firstHeading >= 0 && (parts[firstHeading] as MdHeading).level === 1 && (firstBlock < 0 || firstHeading < firstBlock)) { + title = plainText((parts[firstHeading] as MdHeading).text); + parts.splice(firstHeading, 1); + } + const levels = parts.filter((p): p is MdHeading => p.kind === "heading").map((p) => p.level); + const sectionLevel = levels.length ? Math.min(...levels) : null; + + const anchors = new IdAllocator(/[^a-z0-9-]+/g); + const summary: string[] = []; + const sections: { section: Section; body: string[] }[] = []; + const open = (heading: string) => { + const id = anchors.claim(slugOf(heading) || "section", "section"); + sections.push({ section: { id, title: heading, claims: [] }, body: [] }); + }; + if (sectionLevel === null) open(defaultSection); + + // The block before, while it is still a candidate to take a citation-only + // block's citations: its rewritten markdown, its text, and where it went. + let prev: { blk: Blk; into: "body" | "claim"; at: number } | null = null; + let prevText = ""; + + for (const part of parts) { + if (part.kind === "heading") { + prev = null; + prevText = ""; + if (part.level === sectionLevel) { + open(plainText(part.text)); + continue; + } + const line = `${"#".repeat(part.level)} ${part.text}`; + if (sections.length) sections[sections.length - 1].body.push(line); + else summary.push(line); + continue; + } + if (part.code) { + if (sections.length) sections[sections.length - 1].body.push(part.md); + else summary.push(part.md); + prev = null; + prevText = ""; + continue; + } + const blk = await rewriteBlock(part.md, prevText, resolve); + prevText = plainText(part.md.replace(LINK_RE, "")); + if (!sections.length) { + summary.push(blk.md); + continue; + } + const cur = sections[sections.length - 1]; + const claims = cur.section.claims!; + if (!blk.ids.length) { + cur.body.push(blk.md); + prev = { blk, into: "body", at: cur.body.length - 1 }; + continue; + } + if (!blk.text && prev) { + // Citations alone: they cite the block before. + if (prev.into === "body") { + cur.body.splice(prev.at, 1); + claims.push(claimOf(anchors, { ...prev.blk, md: `${prev.blk.md}\n${blk.md}`, ids: blk.ids, interleaved: false })); + } else { + const claim = claims[prev.at]; + claim.citations = [...new Set([...(claim.citations ?? []), ...blk.ids])]; + } + prev = null; + continue; + } + claims.push(claimOf(anchors, blk)); + prev = { blk, into: "claim", at: claims.length - 1 }; + } + + return { + title, + summary: summary.join("\n\n").trim(), + sections: sections.map(({ section, body }) => { + const out: Section = { id: section.id, title: section.title }; + const text = body.join("\n\n").trim(); + if (text) out.body = text; + if (section.claims!.length) out.claims = section.claims; + return out; + }), + }; +} + +function claimOf(anchors: IdAllocator, blk: Blk): Claim { + const claim: Claim = { id: anchors.next("k"), text: blk.text || plainText(blk.md), citations: blk.ids }; + if (blk.interleaved) claim.findings = stripMarkers(blk.md); + return claim; +} + +// ─── The finished report ─── + +export function emptyReport(id: string, kind: Report["kind"], title: string): Report { + return { format: REPORT_FORMAT, version: REPORT_VERSION, id, kind, title, sections: [] }; +} + +// The report with empty optional maps left out, checked. +export function finish(report: Report, warnings: string[]): Converted { + const out: Report = { ...report }; + if (out.citations && Object.keys(out.citations).length === 0) delete out.citations; + if (out.sources && Object.keys(out.sources).length === 0) delete out.sources; + if (out.summary !== undefined && !/\S/.test(out.summary)) delete out.summary; + const parsed = parseReport(out); + return { report: out, warnings, problems: parsed.problems }; +} + +// Every `cite:` id a markdown names, in order. +export function citedIds(md: string | undefined): string[] { + return extractCiteRefs(md).map((r) => r.id); +} + +// A citation map entry, typed for the converters that build one. +export type CitationMap = Record<string, Citation>; + +// ─── The citations a converter collects ─── + +const CLOCK_SUFFIX_RE = /\s*@\s*\d{1,2}(?::\d{2}){1,2}\s*$/; +const POST_LABEL_DATE_RE = /,\s*(\d{4}-\d{2}-\d{2}(?:T[^\s,]*)?)\s*$/; + +// A citing link's label as the citation's display name: one line, without the +// `@ mm:ss` the moment already carries. +export function labelOf(label: string): string | undefined { + const s = plainText(label).replace(CLOCK_SUFFIX_RE, "").trim(); + return s || undefined; +} + +// The citations one document cites, each once: a span by its record and +// second, a post by its channel and id. Ids are `c01`, `c02`, … for spans and +// `p01`, … for posts, in order of first citing. +export class CitationRegistry { + readonly citations: CitationMap = {}; + private readonly ids = new IdAllocator(); + private readonly byKey = new Map<string, string>(); + private readonly warnedCues = new Set<string>(); + + constructor( + private readonly ctx: ConvertContext, + private readonly warnings: string[], + ) {} + + async span( + channel: string, + id: string, + seconds: number, + quote: string, + opts: { label?: string; end?: number; kind?: "video" | "audio"; date?: string } = {}, + ): Promise<string> { + const key = `span ${channel}/${id} ${seconds} ${opts.end ?? ""}`; + const known = this.byKey.get(key); + if (known) return known; + const span = await spanAt(channel, id, seconds, quote, this.ctx, opts.end); + if (span.from === "default" && this.ctx.cuesOf && !this.warnedCues.has(`${channel}/${id}`)) { + this.warnedCues.add(`${channel}/${id}`); + this.warnings.push( + `${channel}/${id}: no cues to widen from — its citations end ${this.ctx.spanSeconds ?? DEFAULT_SPAN_SECONDS} s after the cited second`, + ); + } + const cid = this.ids.next("c"); + this.citations[cid] = { + kind: opts.kind ?? "video", + channel, + id, + start: span.start, + end: span.end, + quote, + ...(opts.label ? { label: opts.label } : {}), + ...(opts.date ? { date: opts.date } : {}), + }; + this.byKey.set(key, cid); + return cid; + } + + post(channel: string, id: string, quote: string, opts: { label?: string; date?: string } = {}): string { + const key = `post ${channel}/${id}`; + const known = this.byKey.get(key); + if (known) return known; + const cid = this.ids.next("p"); + this.citations[cid] = { + kind: "post", + channel, + id, + quote, + ...(opts.label ? { label: opts.label } : {}), + ...(opts.date ? { date: opts.date } : {}), + }; + this.byKey.set(key, cid); + return cid; + } + + // A citing link (an archive moment or post, a moment page, a post on its + // platform) as a citation's id; null, with a warning when it looked like a + // citation, for a link to leave as it is. + async link(href: string, label: string, quote: string): Promise<string | null> { + const parsed = parseCitationHref(href); + if (!parsed) return null; + if (parsed.kind === "span") { + return this.span(parsed.channel, parsed.id, parsed.seconds, quote, { label: labelOf(label), end: parsed.end }); + } + const date = POST_LABEL_DATE_RE.exec(label)?.[1]; + const postOpts = { label: labelOf(label), ...(date && isPartialDate(date) ? { date } : {}) }; + if (parsed.kind === "post") return this.post(parsed.channel, parsed.id, quote, postOpts); + const channel = this.ctx.postChannelOf + ? await this.ctx.postChannelOf({ platform: parsed.platform, id: parsed.id, handle: parsed.handle }) + : null; + if (!channel) { + this.warnings.push( + `${href}: no archive channel keeps this post${this.ctx.postChannelOf ? "" : " (give a channels dir to look it up)"} — left as a plain link`, + ); + return null; + } + return this.post(channel, parsed.id, quote, postOpts); + } +} diff --git a/common/lib/report/convertSweep.ts b/common/lib/report/convertSweep.ts @@ -0,0 +1,52 @@ +// A /sweep REPORT (markdown) → `report.json` (kind "sweep"). +// +// What /sweep writes (mcp/src/instructions.ts): `## sections` of findings, +// each cited as `[title @ mm:ss](<origin>/?v=<channel>%2F<id>&t=<seconds>)` +// for a video and `[post by <author>, <date>](<url>)` for a post — the url an +// archive post (`…&vm=post`) or the post on its platform. The engine +// (./convertShared.ts `structureMarkdown`) makes the document's first `#` +// heading the title, the text before the first section the summary, each +// section a section, and each list item or paragraph that cites a claim; every +// citing link becomes `[label](cite:<id>)`. +// +// A citation's quote is the quoted words nearest before its link (`"…"` or +// `“…”`, several joined by an ellipsis), else the claim's text — VERBATIM is +// the model's rule, and compose checks it against the cues. Its span is the +// cited second widened through the record's cues when the caller can read +// them (`ctx.cuesOf`), else the second plus `ctx.spanSeconds`. A post on its +// platform is cited only when `ctx.postChannelOf` finds the archive channel +// that keeps it; otherwise its link stays a plain link, with a warning. + +import { + CitationRegistry, + emptyReport, + finish, + reportIdOf, + structureMarkdown, + type ConvertContext, + type Converted, +} from "./convertShared"; + +export type SweepConvertOptions = ConvertContext & { + // The report's id; default: the title's slug. + id?: string; + // The report's title; default: the markdown's first `#` heading. + title?: string; + // The one section of a sweep with no section headings. + sectionTitle?: string; +}; + +export async function sweepToReport(md: string, opts: SweepConvertOptions = {}): Promise<Converted> { + const warnings: string[] = []; + const registry = new CitationRegistry(opts, warnings); + const doc = await structureMarkdown(md, (link) => registry.link(link.href, link.label, link.quote), { + defaultSection: opts.sectionTitle ?? "Findings", + }); + const title = opts.title ?? doc.title ?? "Sweep"; + const report = emptyReport(reportIdOf(opts.id, title), "sweep", title); + if (doc.summary) report.summary = doc.summary; + report.citations = registry.citations; + report.sections = doc.sections; + if (!Object.keys(registry.citations).length) warnings.push("the markdown cites nothing the citation model knows"); + return finish(report, warnings); +} diff --git a/common/lib/report/docs.ts b/common/lib/report/docs.ts @@ -74,5 +74,23 @@ export function renderReportMarkdown(): string { out.push(`| \`${v}\` | ${cell(d.label)} | \`${d.color}\` |`); } out.push(""); + out.push("## Making a report from what exists"); + out.push(""); + out.push( + "`archilyzer reports convert <sweep|ask|manifest> <in> --out <report.json>` makes a " + + "report of a /sweep report (markdown: its `#` heading the title, each `##` section a " + + "section, each list item or paragraph that cites a claim), an /ask answer (its " + + "`[n @ mm:ss]` markers resolved against its sources) or a report-to-video manifest " + + "(chapter cards the sections, `claim` entries the claims, clips, stills and posts the " + + "citations), and writes it only when it validates. A /sweep or /ask citation carries " + + "one second: `--channels-dir <transcripts/channels>` widens it through the record's " + + "cues to whole sentences (else it is the second plus 10 s) and finds the channel that " + + "keeps a post cited by its platform link. `archilyzer reports to-manifest <report.json> " + + "--out <manifest.json>` goes the other way: a starter manifest — a title card, a " + + "chapter card per section, each claim's still and clips stamped with its verdict, its " + + "posts. The converters are `common/lib/report/convert*.ts`; the warnings say what a " + + "conversion left out.", + ); + out.push(""); return out.join("\n"); } diff --git a/common/package.json b/common/package.json @@ -31,6 +31,7 @@ "./components/urlState": "./components/urlState.ts", "./components/virtualizer": "./components/virtualizer.ts", "./components/*": "./components/*.tsx", + "./lib/cueWiden.mjs": "./lib/cueWiden.mjs", "./lib/detectPlatform.mjs": "./lib/detectPlatform.mjs", "./lib/evidenceClip.mjs": "./lib/evidenceClip.mjs", "./lib/ports.mjs": "./lib/ports.mjs", diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -5,6 +5,7 @@ - **A site can say what it publishes, and which reports.** `site.json` takes `publish` — `"full"`, the searchable corpus every site has been (the default, never written), or `"cited"`, only the site's reports and the moments they cite — and `reports`, the ordered ids of its published reports (each a slug; invalid and repeated ids are dropped). The site form has a Publish control and lists the site's reports read-only; saving the form keeps the stored list. A cited site still builds as a full one until the reports pipeline applies the scope. SITE.md documents both keys. - **A report and its citations now have one written format, checked before anything is built from them.** A cited report is a `report.json` (`archilyzer-report`, version 1): a summary, then sections of claims, each claim with an optional verdict, the reviewed document's own sentence, findings in markdown and the citations it rests on. A citation is one of five kinds — a span of a video, a span of an audio record, a post, a sentence of a source document, or a web page — with a verbatim quote, and is cited from any markdown in the report as `[label](cite:<id>)`. The checker lists every problem at once with where it is: a citation, a source or a `cite:` link that names nothing, a span that ends before it starts or runs past 120 seconds with its context, a still that points outside the report's folder, an id used twice. Each cited span and post has one page address, `/m/<channel>/<id>/<start>-<end>/` or `/m/<channel>/<id>/`. Nothing builds or shows reports yet. The fact-check verdicts (Corroborated, Partly true, Contradicted, Not found, Untestable) and their colours are now kept in one place, which the report video's stamps and tally read too. `REPORT.md` and `CITATIONS.md` list every key. - **A site's reports can have their evidence media prepared: every cited span cut to a clip, every cited post's capture copied.** `archilyzer reports prepare <site>`, the `reports-prepare` job (`POST /api/ops/reports-prepare`, `pnpm ops reports-prepare`) reads the site's published reports, checks them, and for each cited moment cuts the span (with its context) out of the media already on disk — a fetched clip window, the saved video, or the recording's audio — fitted inside 1280×720 with H.264 and AAC, or as an `.m4a` for an audio span; and copies the screenshot and attached media of each cited post, and of no other post, beside them. Everything lands in the site's build staging (`.export-index/sites/<site>/report-media/`) with an `index.json` naming each moment's file, size, checksum, size in pixels and duration. A clip is cut once and reused while its source file and span are unchanged; a clip or capture no longer cited is removed. Nothing is fetched: a citation whose media is not on disk, a post without a screenshot, a clip over 24 MiB, a citation of a channel outside the site, a post the site may not show, or an invalid or missing report is listed with the citations it affects, and the run fails (exit 1, or a failed job) — a span on a drive that is not mounted is reported as such rather than as missing. The lookup of a span's media on disk is now shared with report-to-video, which finds the same files it did. +- **A report can be made from a /sweep report, an /ask answer or a report video's manifest, and a report video's manifest from a report.** `archilyzer reports convert sweep|ask|manifest <in> --out <report.json>` writes a `report.json`. From a /sweep report (markdown): its first `#` heading is the title, the text before the first section the summary, each `##` section a section, and each list item or paragraph with a citing link a claim — a line that is only a citation under a quote cites that quote — with every archive moment link, archive post link, moment page link and X or Bluesky post link turned into a citation and the link into `[label](cite:<id>)`; a citation's quote is the quoted words nearest before its link. From an /ask answer (`{ "answer", "sources" }`, a saved chat or an assistant message, as JSON): each `[n]` or `[n @ mm:ss]` marker becomes a citation of source `n` at the excerpt line nearest the marked time, quoting that line, or of the post. From a manifest: chapter cards are the sections, entries carrying `claim` the claims with their verdicts (the first still a claim's source sentence), clips the video citations, stills the source citations and `posts` the post citations; ledger rows are claims too. A /sweep or /ask citation names one second: with `--channels-dir <transcripts/channels>` its span is the cited cue run on to the quote's length and widened to whole sentences, as `resolve-windows.mjs` widens a clip (the widening now lives in common and is shared), and a post cited by its X or Bluesky link is found in the channel that archives it; without it the span is the second plus 10 seconds and such a post stays a plain link. `archilyzer reports to-manifest <report.json> --out <manifest.json>` writes a starter manifest: a title card, a chapter card per section, the section's cited clips, and per claim its source sentence's still and its clips, each with the claim's verdict, and its posts attached to its last clip (a post's platform, link and date come from the channels tree, or for an X post from its id). Both check what they write against the report format and write nothing when it has problems (exit 1); everything a conversion had to leave out or guess is printed as a warning. A manifest image may carry `quote`, the words the still shows, which the converters keep. - **A long report video no longer runs out of memory while its clips are crossfaded.** `build-video.mjs` used to join every segment of a cut in one ffmpeg command, which grows with the number of segments: a cut of a few hundred clips could use more memory than the machine had and be stopped. Past 24 segments the build now crossfades them in batches of consecutive segments, each into a file under `out/<variant>/xfade-batches/`, then crossfades those files together with the same transition and lays the on-screen deck, the rail and the dips over them. Every transition, the deck's schedule and the chapters land on the same frames as before, at the cost of one more video encode on such a cut. The batch size is `render.xfadeBatch` in the manifest or `REPORT_VIDEO_XFADE_BATCH` in the environment (which wins); `0` never batches. A batch file is reused while its segments, their holds and moves and the encode settings are unchanged, so a `--chrome-only` run that moves no footage redoes only the last pass. - **Auto-download no longer tries a video the metadata scan already found members-only or private.** The scan records why it could not read a video, but only a failed download used to take a video out of the auto-download queue, so each members-only video the scan had found was still downloaded once: four yt-dlp requests, two with browser cookies, about 30 seconds each. A members-only or private answer from the scan now keeps the video out of the queue and counts it under the channel's members-only or private exclusions, and it stays in **Needs cookies** for a manual cookie run. A scan that reads the video later lifts this. A video the scan saw only as "Video unavailable" is still tried, since YouTube gives that answer when it is throttling too. - **X post fetches stop on Drain, and an account with no posts is not searched.** Draining a `fetch-posts` job used to do nothing until gallery-dl finished its whole run. Now the timeline fetch stops at the next page boundary (at once when gallery-dl is between pages or waiting out a rate limit), the older-posts walk stops its current window's search at once and never starts the 45–120 second pause between windows, and both keep their resume point: the job ends done, not failed, and the log says "Drained; the next run resumes …". An older-posts walk is refused when nothing is archived and the last timeline fetch finished having read no posts, since it would only repeat empty searches; `"force": true` (`--force` on `archilyzer posts fetch`) walks anyway. A walk with nothing archived that finds nothing ends after two empty three-month windows instead of four, and records why; a walk that has posts keeps the year-of-empty-windows rule. Capture-posts already stopped between posts on Drain. diff --git a/umtool/report-to-video/README.md b/umtool/report-to-video/README.md @@ -256,6 +256,16 @@ the manifest names the MCP video id while the cue file lives under the URL slug. `timeline` is an ordered list; entries are `card`, `clip` or `image` (plus `scroll`, `chart`, `ledger` and `teaser` — the vocabulary is open). +A manifest and a report site's `report.json` convert into each other: +`archilyzer reports to-manifest <report.json> --out video.manifest.json` writes a +starter cut of a report (a title card, a chapter card per section, each claim's +still and clips carrying its `claim`, its posts), and `archilyzer reports convert +manifest video.manifest.json --out report.json` makes a report of a manifest — see +REPORT.md, "Making a report from what exists". An image's `quote` is what the +converters carry as the still's words; nothing in a build reads it. The sentence +widening `resolve-windows.mjs` does is `common/lib/cueWiden.mjs`, which the +converters share. + ```jsonc { "type": "card", "id": "ch3", "style": "chapter", "seconds": 4.0, "kicker": "March – November 2025", "heading": "Then: the county", @@ -582,6 +592,7 @@ poster — shown for `seconds` and then gone. "title": "Bx (@bx_on_x) on X, Sept 2026 — 1/4", "date": "2026-08-19", // optional; appended like a clip's "citeUrl": "https://…", // optional; the ONLY thing that draws a QR + "quote": "…", // optional: the words the still shows (not drawn) "note": "…" } ``` diff --git a/umtool/report-to-video/resolve-windows.mjs b/umtool/report-to-video/resolve-windows.mjs @@ -51,72 +51,12 @@ import path from "node:path"; import { createCueSource, siteOriginFromManifest } from "./cues.mjs"; import { clipLabel, createLocalMedia, hasLocalMedia } from "./local-media.mjs"; +// THE SENTENCE WIDENING IS COMMON'S (common/lib/cueWiden.mjs): one copy, shared +// with the report converters, which widen a citation's one second into a span. +// This file re-exports it. +import { CUE_EPS as EPS, widen } from "yt-dlp-transcript-common/lib/cueWiden.mjs"; -const ENDS_SENTENCE = /[.!?]["'”’)\]]*\s*$/; - -// A cue that is only "[music]" or "[ __ ]" (the profanity bleep) carries no -// sentence signal; treat it as transparent so expansion walks past it. -const IS_FILLER = /^\s*(\[[^\]]*\]|>>|♪|—|-)*\s*$/; - -// The manifest stores times rounded to 2 dp, so a value read back from it can sit -// a hair BELOW the cue end it came from. Without a tolerance the end lookup then -// lands on the previous cue, the forward search runs on to the next sentence, and -// the clip grows a little every time this is run — it has to be a fixed point. -const EPS = 0.02; - - -function indexAt(cues, t, which) { - // First cue whose span contains t, else the nearest one on the right side. - let idx = cues.findIndex((c) => c.end > t); - if (idx < 0) idx = cues.length - 1; - if (which === "end") { - let j = cues.findIndex((c) => c.end >= t - EPS); - if (j < 0) j = cues.length - 1; - idx = j; - } - return idx; -} - -export function widen(cues, start, end, { maxLead = 8, maxTail = 12 } = {}) { - const isBoundary = (c) => ENDS_SENTENCE.test(c.text) && !IS_FILLER.test(c.text); - const i0 = indexAt(cues, start, "start"); - const i1 = indexAt(cues, end, "end"); - - // START: the latest cue that OPENS a sentence (i.e. its predecessor closes - // one) at or before the quote, within the lead budget. Finding no such cue - // means every candidate lead-in is a sentence fragment, so take none at all — - // a fragment is the irrelevant context we are trying to avoid, not context. - let si = null; - for (let i = i0; i > 0; i -= 1) { - if (start - cues[i].start > maxLead) break; - if (isBoundary(cues[i - 1])) { - si = i; - break; - } - } - if (si === null) si = i0; - - // END: the first cue that CLOSES a sentence at or after the quote. Never - // clamp to a budget here — stopping partway through a sentence is exactly the - // mid-thought ending this is meant to remove, so the budget only decides how - // far to look, and failing to find one falls back to the original cue end. - let ei = null; - for (let j = i1; j < cues.length; j += 1) { - if (cues[j].end - end > maxTail) break; - if (isBoundary(cues[j])) { - ei = j; - break; - } - } - if (ei === null) ei = i1; - - return { - start: cues[si].start, - end: cues[ei].end, - leadCues: i0 - si, - tailCues: ei - i1, - }; -} +export { widen }; // --------------------------------------------------------------------------- // THE CUT INSIDE THE EXTENT.