Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit d37027f8fd23096f14090d57a94520fff253efa9
parent b9e2849d2c90bbcc4e0544c471ec6318af974c2f
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon,  5 Oct 2026 05:23:12 -0400

records: [Unreleased] bullets (export, homepage, editor), PUBLISH.md on how the family and the contract's readers treat a cited site, the plan's compose notes

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
MPUBLISH.md | 8++++++++
Meditor/CHANGELOG.md | 1+
Mexport/CHANGELOG.md | 2++
Mhomepage/CHANGELOG.md | 1+
Mplans/report-sites.md | 5+++--
5 files changed, 15 insertions(+), 2 deletions(-)

diff --git a/PUBLISH.md b/PUBLISH.md @@ -100,6 +100,14 @@ shell's own pages and files, or a file over 25 MiB, or more than 20,000 files, f the build, and every deploy path refuses it. A site switched to cited is refused at deploy until it is built again. +A cited site is not a searchable archive, so the family never lists it, whatever +`listed` says (`isListedSite`): it is no hub member, has no homepage card, counts in +no family total and no other site's footer links to it. A full site with reports is +listed as before. Readers of the contract see a cited site as an empty corpus with +reports: the MCP server's `list_reports` / `get_report` read them (and its +`list_channels` says "cited-only site: N report(s)"), and umtool's cue walk refuses +one by name — it has no transcripts to cut a report video from. + Deploy-only ships whatever is in `export/out`, which the basic build composes one site at a time into a single shared directory — so it **refuses, before starting a job, if `export/out` holds a build of another site** (or no build at all), naming the diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,7 @@ # Changelog ## [Unreleased] +- **A report video's cue lookup names a site that publishes only its reports.** Pointing a report-to-video manifest at such a site (`corpus.json` `site.scope: "cited"`) used to fail with "channel … is not in corpus.json"; it now says the site publishes no transcripts and to use a full archive or a local corpus. The site form's **Publish** hint says a cited-only site is never listed on the homepage or the hub. - **A site's build composes its reports, and a site that publishes only its reports ships nothing else.** Every site's compose now writes the reports its `site.json` publishes: each report's page and its citations as `citations.json` and `citations.csv` under `/reports/<id>/`, its cited stills, a page per cited moment with the record, the transcript lines around the span and every report that cites it, and the clips and post captures `archilyzer reports prepare` made for it, only the cited ones. Each quote is checked against the record as it is composed (a span's against its cues within 5 s either side, read from `en-orig` when the `en` track has no cues; a post's against its text) and the score, time and method are written into the citation, replacing any typed by hand. The build stops with the list of every problem before anything is written: an invalid report, a citation of a channel outside the site or of a post the site may not carry, a missing record, still or post, a quote that matches less than 60 % of what the record says, and a citation without prepared media or with media cut for another span (`--allow-missing-media` on `archilyzer compose site` and `build site` lets those two through, without a clip). A site with `publish: "cited"` removes everything corpus-shaped from `export/public` before it writes its reports, and its built `out/` is checked against what a cited site may hold: anything else, a file over 25 MiB or more than 20,000 files fails the build, and every deploy path (the Publish tab, `deploy site`, Build & deploy, Build & deploy all, the container build) refuses it, as it refuses a site set to cited whose last build was a full one. The hub's compose removes a report site's files too. - **A site has a Reports tab.** `/sites/<site>/reports` lists every report under the site's `reports/` directory — the published ones in their order, then the drafts — with its kind, dates, sections, claims, citations by kind and, for a fact-check, how many claims carry each verdict. Each report's problems, from the same checker the prepare step and the build use, open under it. A draft with no problems can be published, and a published report moved up or down or unpublished; each writes only the site's `reports` list, applied to the list as it is on disk at that moment, so it never overwrites another change to the site. "Prepare evidence media" queues the `reports-prepare` job, and beside it the tab shows the last prepared media (moments by kind, total size, problems by kind) and links the last prepare job. What the site publishes (full or cited) is shown with a link to Settings, where it is changed. - **A site can say what it publishes, and which reports.** `site.json` takes `publish` — `"full"`, the searchable corpus every site has been (the default, never written), or `"cited"`, only the site's reports and the moments they cite — and `reports`, the ordered ids of its published reports (each a slug; invalid and repeated ids are dropped). The site form has a Publish control and lists the site's reports read-only; saving the form keeps the stored list. A cited site still builds as a full one until the reports pipeline applies the scope. SITE.md documents both keys. diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md @@ -1,6 +1,8 @@ # Changelog ## [Unreleased] +- **MCP: a site's reports can be read, and a site that publishes only reports says so instead of looking empty.** Two new tools: `list_reports` lists the reports a site publishes (id, title, kind, claim and citation counts, a fact-check's verdict tally, its page), and `get_report` reads one — its tally, then each section's claims with their verdicts and findings and, for every citation, the verbatim quote, the original (the platform at the cited second, the post, the document) and the site's moment page; `section` reads one section. On a site that publishes only its reports, `list_channels`, `list_sources` and `resolve_source` say "cited-only site: N report(s)" where they said "No channels found"; on a site with reports as well, `list_sources` and `resolve_source` say how many. The archive readers (local, remote) read `corpus.json` spec 5: a cited site is an empty corpus without an error, and a local copy of one is never read from a stale `transcripts/` folder. A hub has no reports of its own; `list_reports` says to name a member site. +- **The hub leaves out a site that publishes only its reports.** A site with `publish: "cited"` is not a searchable archive, so it is no hub member (not in federated search, the hub's `corpus.json` or `llms.txt`) and no other site's footer links to it, whatever its **List on the Archilyzer homepage and hub** setting says. A site with reports that publishes its full corpus is a member as before. Needs a hub rebuild and deploy once such a site exists. - **`corpus.json` is spec 5: it names a site's reports, and a site that publishes only reports says so.** A site with reports adds `reports` to its `corpus.json` (`index`: `/reports/index.json`, the count, and how to read a report's page, its citations and its moment pages) and a Reports section to `llms.txt`; its sitemap lists the report and moment pages. A site that publishes only its reports has `"scope": "cited"` and its audience under `site`, no channels and zero totals, an `llms.txt` that lists its reports and how their citations and moment pages are read, and a `site.json` with no channels. A reader that does not know spec 5 sees an empty corpus there. Needs a rebuild and deploy of each site. - **A site can show cited reports, and every citation opens on a page of its own.** A site built with reports has a **Reports** link in its header and a page at `/reports/` listing them. A report's page has its title, subtitle, dates and the document under review, its archive links listed once under it and folded away ("N archive links in context"); a fact-check's tally of verdicts; the summary; the sections and their claims, each with its verdict, the document's own sentence as an image with a link back to the document and only the archive links that sit in that sentence (at most five), the findings and the evidence cards; a numbered reference list; and links to download its citations as JSON and CSV. A citation in the text shows as its words plus a number: hovering it, focusing the number or tapping it once shows a card of the citation (the quote, who said it and when, a picture or the post's screenshot, and how closely the quote matched the transcript when it was checked); the words open what it cites and the number jumps to its reference. A cited span of a video or audio record opens at `/m/<channel>/<id>/<start>-<end>/`: a short clip of the span with a little context either side, the quote, the transcript lines around it, the record's title, channel and date, a link to the original at that time, and every report on the site that cites it. A cited post opens at `/m/<channel>/<id>/` with its screenshot and text. A site that publishes only its reports (`site.json` `publish: "cited"`) opens on the report index and has no search, Ask AI, downloads or duplicates. A site with no reports is unchanged. Needs a rebuild and deploy of each site. diff --git a/homepage/CHANGELOG.md b/homepage/CHANGELOG.md @@ -1,6 +1,7 @@ # Homepage Changelog ## [Unreleased] +- **A site that publishes only its reports is not on the homepage.** A site with `publish: "cited"` has no Official Instances card, chart series, `/stats` entry or recent item, is not in `channel-sites.json`, and the channels only it carries count in no total — as an unlisted site, whatever its listing setting says. A site with reports that publishes its full corpus is listed as before. - **The AI and MCP doc has a Ten-minute setup.** Right after the MCP server's introduction, one block runs Claude Code against a published archive, the Jeralyzer as the example: clone the source (or unpack the tarball on Downloads), `pnpm install`, `claude mcp add archilyzer`, start `claude` and try `/ask`; then what it needs, why the server must be registered as `archilyzer` (the shipped `/ask` and `/sweep` call `mcp__archilyzer__…`), the two optional editor lines for `fetch_clip`, `TRANSCRIPT_HUB_URL`, where the `mcp.json` form for other clients is, and WSL2 on Windows. "What it can do" is a heading of its own after it. Every archive's **Use with AI** link now lands on this page. - **The Ten-minute setup starts the server with the source's own command, `pnpm --silent -C "$PWD" archilyzer mcp`**, as the README and the MCP server's README do. `--silent` keeps pnpm's own lines off the output the client reads the server's replies on, and a note says so. The note on other clients gives the `mcp.json` entry's arguments in the same form. - **`/source/` links the source's history.** A History block — how many commits (past 10,000, "the latest 10,000 of N"), the newest one (linking to its page), and links to the Log, the Refs and the Atom feed — shows when the build published the history pages (`/source/git/`, rendered by stagit); without them there is no History block. The history pages open on the homepage's ground (the reader's stored choice, else Dark; without JavaScript, the system's), start with one line back to `/source/`, and their Files page is an index into the raw tree. The e2e shows the page with and without a history from a fixture publish (`E2E_SOURCE_PUBLIC_DIR`, never read by a production build), and walks the real pages when the checkout has published them. diff --git a/plans/report-sites.md b/plans/report-sites.md @@ -131,8 +131,9 @@ Citations are the core; reports are one consumer. `audience`, and `reports: { index: "/reports/index.json" }` — the private-deploy guard and compose's cross-site check read them. `CONTRACT.corpusSpec` → 5; readers tolerate a cited scope (an old reader sees an empty corpus, which is safe). `llms.txt` and the sitemap get a cited variant listing report routes. -- Hub and homepage exclude `cited` sites for now; MCP may later gain `list_reports` / `get_report`. -- umtool's cue walk (`cues.mjs`) cannot resolve cues off a cited site — documented. +- Hub and homepage exclude `cited` sites (`isListedSite` folds them out, as a private site); MCP serves + `list_reports` / `get_report` and says "cited-only site: N report(s)" (R5). +- umtool's cue walk (`cues.mjs`) cannot resolve cues off a cited site — it refuses one by name (R5). ## Converters (to make reports from what exists)