Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit f83ae7767edf136806650b74cb9bbbe3432ec1a6
parent ddedd92b1d30571a3117d9c4cd126e9b4bfde0db
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon,  5 Oct 2026 12:32:50 -0400

changelogs: the unreleased report-site bullets describe the search switch, not publish

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 3+--
Mexport/CHANGELOG.md | 2+-
2 files changed, 2 insertions(+), 3 deletions(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,11 +1,10 @@ # Changelog ## [Unreleased] -- **A site with search off is a report-only site.** `site.json` `search` (default on; only `false` is written) replaces `publish`: a site with search off publishes only its reports and the moments they cite. The site form's **Publish** select is a **Search** checkbox, and the Reports tab says which the site is. A `publish: "cited"` already on disk reads as search off and is saved as `search: false`. - **A report video's cue lookup names a site that publishes only its reports.** Pointing a report-to-video manifest at such a site (`corpus.json` `site.scope: "cited"`) used to fail with "channel … is not in corpus.json"; it now says the site publishes no transcripts and to use a full archive or a local corpus. The site form's **Publish** hint says a cited-only site is never listed on the homepage or the hub. - **A site's build composes its reports, and a site that publishes only its reports ships nothing else.** Every site's compose now writes the reports its `site.json` publishes: each report's page and its citations as `citations.json` and `citations.csv` under `/reports/<id>/`, its cited stills, a page per cited moment with the record, the transcript lines around the span and every report that cites it, and the clips and post captures `archilyzer reports prepare` made for it, only the cited ones. Each quote is checked against the record as it is composed (a span's against its cues within 5 s either side, read from `en-orig` when the `en` track has no cues; a post's against its text) and the score, time and method are written into the citation, replacing any typed by hand. The build stops with the list of every problem before anything is written: an invalid report, a citation of a channel outside the site or of a post the site may not carry, a missing record, still or post, a quote that matches less than 60 % of what the record says, and a citation without prepared media or with media cut for another span (`--allow-missing-media` on `archilyzer compose site` and `build site` lets those two through, without a clip). A site with `publish: "cited"` removes everything corpus-shaped from `export/public` before it writes its reports, and its built `out/` is checked against what a cited site may hold: anything else, a file over 25 MiB or more than 20,000 files fails the build, and every deploy path (the Publish tab, `deploy site`, Build & deploy, Build & deploy all, the container build) refuses it, as it refuses a site set to cited whose last build was a full one. The hub's compose removes a report site's files too. - **A site has a Reports tab.** `/sites/<site>/reports` lists every report under the site's `reports/` directory — the published ones in their order, then the drafts — with its kind, dates, sections, claims, citations by kind and, for a fact-check, how many claims carry each verdict. Each report's problems, from the same checker the prepare step and the build use, open under it. A draft with no problems can be published, and a published report moved up or down or unpublished; each writes only the site's `reports` list, applied to the list as it is on disk at that moment, so it never overwrites another change to the site. "Prepare evidence media" queues the `reports-prepare` job, and beside it the tab shows the last prepared media (moments by kind, total size, problems by kind) and links the last prepare job. What the site publishes (full or cited) is shown with a link to Settings, where it is changed. -- **A site can say what it publishes, and which reports.** `site.json` takes `publish` — `"full"`, the searchable corpus every site has been (the default, never written), or `"cited"`, only the site's reports and the moments they cite — and `reports`, the ordered ids of its published reports (each a slug; invalid and repeated ids are dropped). The site form has a Publish control and lists the site's reports read-only; saving the form keeps the stored list. A cited site still builds as a full one until the reports pipeline applies the scope. SITE.md documents both keys. +- **A site can publish reports, and with search off it is only its reports.** `site.json` takes `reports`, the ordered ids of its published reports (each a slug; invalid and repeated ids are dropped), and `search` (default on; only `false` is written): a site with search off publishes only its reports and the moments they cite. The site form has a **Search** checkbox and lists the site's reports read-only (saving the form keeps the stored list), and the Reports tab says which the site is. SITE.md documents both keys. - **A report and its citations now have one written format, checked before anything is built from them.** A cited report is a `report.json` (`archilyzer-report`, version 1): a summary, then sections of claims, each claim with an optional verdict, the reviewed document's own sentence, findings in markdown and the citations it rests on. A citation is one of five kinds — a span of a video, a span of an audio record, a post, a sentence of a source document, or a web page — with a verbatim quote, and is cited from any markdown in the report as `[label](cite:<id>)`. The checker lists every problem at once with where it is: a citation, a source or a `cite:` link that names nothing, a span that ends before it starts or runs past 120 seconds with its context, a still that points outside the report's folder, an id used twice. Each cited span and post has one page address, `/m/<channel>/<id>/<start>-<end>/` or `/m/<channel>/<id>/`. Nothing builds or shows reports yet. The fact-check verdicts (Corroborated, Partly true, Contradicted, Not found, Untestable) and their colours are now kept in one place, which the report video's stamps and tally read too. `REPORT.md` and `CITATIONS.md` list every key. - **A site's reports can have their evidence media prepared: every cited span cut to a clip, every cited post's capture copied.** `archilyzer reports prepare <site>`, the `reports-prepare` job (`POST /api/ops/reports-prepare`, `pnpm ops reports-prepare`) reads the site's published reports, checks them, and for each cited moment cuts the span (with its context) out of the media already on disk — a fetched clip window, the saved video, or the recording's audio — fitted inside 1280×720 with H.264 and AAC, or as an `.m4a` for an audio span; and copies the screenshot and attached media of each cited post, and of no other post, beside them. Everything lands in the site's build staging (`.export-index/sites/<site>/report-media/`) with an `index.json` naming each moment's file, size, checksum, size in pixels and duration. A clip is cut once and reused while its source file and span are unchanged; a clip or capture no longer cited is removed. Nothing is fetched: a citation whose media is not on disk, a post without a screenshot, a clip over 24 MiB, a citation of a channel outside the site, a post the site may not show, or an invalid or missing report is listed with the citations it affects, and the run fails (exit 1, or a failed job) — a span on a drive that is not mounted is reported as such rather than as missing. The lookup of a span's media on disk is now shared with report-to-video, which finds the same files it did. - **A report can be made from a /sweep report, an /ask answer or a report video's manifest, and a report video's manifest from a report.** `archilyzer reports convert sweep|ask|manifest <in> --out <report.json>` writes a `report.json`. From a /sweep report (markdown): its first `#` heading is the title, the text before the first section the summary, each `##` section a section, and each list item or paragraph with a citing link a claim — a line that is only a citation under a quote cites that quote — with every archive moment link, archive post link, moment page link and X or Bluesky post link turned into a citation and the link into `[label](cite:<id>)`; a citation's quote is the quoted words nearest before its link. From an /ask answer (`{ "answer", "sources" }`, a saved chat or an assistant message, as JSON): each `[n]` or `[n @ mm:ss]` marker becomes a citation of source `n` at the excerpt line nearest the marked time, quoting that line, or of the post. From a manifest: chapter cards are the sections, entries carrying `claim` the claims with their verdicts (the first still a claim's source sentence), clips the video citations, stills the source citations and `posts` the post citations; ledger rows are claims too. A /sweep or /ask citation names one second: with `--channels-dir <transcripts/channels>` its span is the cited cue run on to the quote's length and widened to whole sentences, as `resolve-windows.mjs` widens a clip (the widening now lives in common and is shared), and a post cited by its X or Bluesky link is found in the channel that archives it; without it the span is the second plus 10 seconds and such a post stays a plain link. `archilyzer reports to-manifest <report.json> --out <manifest.json>` writes a starter manifest: a title card, a chapter card per section, the section's cited clips, and per claim its source sentence's still and its clips, each with the claim's verdict, and its posts attached to its last clip (a post's platform, link and date come from the channels tree, or for an X post from its id). Both check what they write against the report format and write nothing when it has problems (exit 1); everything a conversion had to leave out or guess is printed as a warning. A manifest image may carry `quote`, the words the still shows, which the converters keep. diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md @@ -3,7 +3,7 @@ ## [Unreleased] - **`pnpm start:export` serves a built site's moment pages.** It used `serve`, which listed a video or audio moment's directory (`3126.00-3151.00`) instead of serving its page; it now runs `export/scripts/serve-out.mjs`, which serves directories as Cloudflare Pages does, on `EXPORT_DEV_PORT` (3000). - **MCP: a site's reports can be read, and a site that publishes only reports says so instead of looking empty.** Two new tools: `list_reports` lists the reports a site publishes (id, title, kind, claim and citation counts, a fact-check's verdict tally, its page), and `get_report` reads one — its tally, then each section's claims with their verdicts and findings and, for every citation, the verbatim quote, the original (the platform at the cited second, the post, the document) and the site's moment page; `section` reads one section. On a site that publishes only its reports, `list_channels`, `list_sources` and `resolve_source` say "cited-only site: N report(s)" where they said "No channels found"; on a site with reports as well, `list_sources` and `resolve_source` say how many. The archive readers (local, remote) read `corpus.json` spec 5: a cited site is an empty corpus without an error, and a local copy of one is never read from a stale `transcripts/` folder. A hub has no reports of its own; `list_reports` says to name a member site. -- **The hub leaves out a site that publishes only its reports.** A site with `publish: "cited"` is not a searchable archive, so it is no hub member (not in federated search, the hub's `corpus.json` or `llms.txt`) and no other site's footer links to it, whatever its **List on the Archilyzer homepage and hub** setting says. A site with reports that publishes its full corpus is a member as before. Needs a hub rebuild and deploy once such a site exists. +- **The hub leaves out a site that publishes only its reports.** A site with search off (`search: false`) is not a searchable archive, so it is no hub member (not in federated search, the hub's `corpus.json` or `llms.txt`) and no other site's footer links to it, whatever its **List on the Archilyzer homepage and hub** setting says. A site with reports that publishes its full corpus is a member as before. Needs a hub rebuild and deploy once such a site exists. - **`corpus.json` is spec 5: it names a site's reports, and a site that publishes only reports says so.** A site with reports adds `reports` to its `corpus.json` (`index`: `/reports/index.json`, the count, and how to read a report's page, its citations and its moment pages) and a Reports section to `llms.txt`; its sitemap lists the report and moment pages. A site that publishes only its reports has `"scope": "cited"` and its audience under `site`, no channels and zero totals, an `llms.txt` that lists its reports and how their citations and moment pages are read, and a `site.json` with no channels. A reader that does not know spec 5 sees an empty corpus there. Needs a rebuild and deploy of each site. - **A site can show cited reports, and every citation opens on a page of its own.** A site built with reports has a **Reports** link in its header and a page at `/reports/` listing them. A report's page has its title, subtitle, dates and the document under review, its archive links listed once under it and folded away ("N archive links in context"); a fact-check's tally of verdicts; the summary; the sections and their claims, each with its verdict, the document's own sentence as an image with a link back to the document and only the archive links that sit in that sentence (at most five), the findings and the evidence cards; a numbered reference list; and links to download its citations as JSON and CSV. A citation in the text shows as its words plus a number: hovering it, focusing the number or tapping it once shows a card of the citation (the quote, who said it and when, a picture or the post's screenshot, and how closely the quote matched the transcript when it was checked); the words open what it cites and the number jumps to its reference. A cited span of a video or audio record opens at `/m/<channel>/<id>/<start>-<end>/`: a short clip of the span with a little context either side, the quote, the transcript lines around it, the record's title, channel and date, a link to the original at that time, and every report on the site that cites it. A cited post opens at `/m/<channel>/<id>/` with its screenshot and text. A site that publishes only its reports (`site.json` `publish: "cited"`) opens on the report index and has no search, Ask AI, downloads or duplicates. A site with no reports is unchanged. Needs a rebuild and deploy of each site.