Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit c65c3df271a76212daac1cba4d9c735c66bb6213
parent adecfb2b8992a406b1c8ce650bd75afb70ff06b2
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon,  5 Oct 2026 15:15:11 -0400

docs: archive.org items — README section with operator steps, CHANNEL.md platform list, changelogs

The README section names the WARC/Wayback records as out of scope and where
they would go.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
MCHANNEL.md | 2+-
MREADME.md | 36++++++++++++++++++++++++++++++++++++
Mcommon/lib/channelConfig.ts | 2+-
Meditor/CHANGELOG.md | 2++
Mexport/CHANGELOG.md | 1+
Mmcp/README.md | 2+-
6 files changed, 42 insertions(+), 3 deletions(-)

diff --git a/CHANNEL.md b/CHANNEL.md @@ -18,7 +18,7 @@ Regenerate this file with `pnpm --filter yt-dlp-transcript-common exec tsx bin/f | `sourceKind` | config | What KIND of source this is: `"video"` (default — yt-dlp + transcription) or `"social"` (an account fetched into the posts corpus, skipped by the video scan). A separate axis from `handling`, so every binary handling branch stays binary. | | `postFetcher` | config | Social channels only: which social fetcher drives ingest (e.g. `"bluesky-atproto"`, `"x-gallery-dl"`). Absent = resolve by URL detection. Trimmed. | | `socialHandle` | config | Social channels only: the bare account handle (a leading "@" is stripped). Derived from `url` at creation but stored, so a later URL-format change upstream cannot silently re-point ingest at a different account. | -| `platform` | config | The source platform (youtube, rumble, …). An unknown value is dropped. | +| `platform` | config | The source platform: youtube, rumble, odysee, twitch, kick, archiveorg (archive.org items — imported, never listed; see README.md, "archive.org items"), twitter or bluesky. An unknown value is dropped. | | `name` | config | Display name. | | `url` | config | The channel / playlist / account URL syncs enumerate. Absent = the channel is never auto-synced. | | `audioFormat` | config | `"m4a"`, `"mp3"` or `"opus"`: the audio a transcribe-handling download keeps. | diff --git a/README.md b/README.md @@ -130,6 +130,42 @@ pnpm start:export # serve it at http://localhost:3000 **[SETUP.md](SETUP.md)** has the full per-OS instructions. Below is the short version. +### archive.org items + +A channel can hold recordings from archive.org — a whole item +(`https://archive.org/details/<identifier>`, when it holds one media file) or single +files of a multi-file item (`…/details/<identifier>/<file>`, e.g. one video of a +channel archive). They are imported, never listed, and transcribed like any +transcribe channel: + +1. **Create the channel**: Platform *archive.org*, handling *Transcribe*, and no URL + (an archive.org channel is never auto-synced). +2. **Import one recording**: the channel's *Import video* with the item or file URL, + or `pnpm ops import-video --json '{"slug":"<channel>","url":"https://archive.org/details/<identifier>"}'`. + An item with several media files is refused and its files are named. +3. **Import chosen files of one item**: + `pnpm ops import-archive-org --json '{"slug":"<channel>","item":"<identifier>","files":["<file>",…]}'` + — or `"match": "<regex>"` over the item's file names; add `"dryRun": true` to see the + list first. One job, one file at a time. + +A whole item's video id is its identifier; a file's is `<identifier>__<slug>-<hash>`. +Each record keeps `archiveorg.json` beside its metadata: the item's title, date, +creator and collections, the item's torrent, and — for a mirror of a YouTube upload — +the original's id, URL, title and upload date, read from the `.info.json` uploaded +with it. The video page shows both; a citation of the record links **archive.org** +and the **torrent** (and, for a mirror, the original **YouTube** upload at the cited +second), so a reader can fetch the file and check it. + +**It is polite to archive.org**: its own job queue, one download at a time with a +jittered pause of at least 8 s between files, the item's metadata asked once (cached), +an identifying User-Agent, `Retry-After` and exponential backoff honoured, a stop +after repeated failures, and a file on disk never fetched again +(`common/lib/archiveOrgClient.ts`, `common/controller/archiveOrgImport.ts`). + +Wayback Machine captures (web.archive.org pages, WARC records) are a different kind of +record and are not handled by this; they would be a source of their own beside +`common/lib/archiveOrg.ts`. + ### Requirements Always needed, to install and run the apps: diff --git a/common/lib/channelConfig.ts b/common/lib/channelConfig.ts @@ -157,7 +157,7 @@ export const CHANNEL_CONFIG_FIELD_DOCS: FieldDocs<ChannelConfig> = { 'Social channels only: which social fetcher drives ingest (e.g. `"bluesky-atproto"`, `"x-gallery-dl"`). Absent = resolve by URL detection. Trimmed.', socialHandle: 'Social channels only: the bare account handle (a leading "@" is stripped). Derived from `url` at creation but stored, so a later URL-format change upstream cannot silently re-point ingest at a different account.', - platform: "The source platform (youtube, rumble, …). An unknown value is dropped.", + platform: "The source platform: youtube, rumble, odysee, twitch, kick, archiveorg (archive.org items — imported, never listed; see README.md, \"archive.org items\"), twitter or bluesky. An unknown value is dropped.", name: "Display name.", url: "The channel / playlist / account URL syncs enumerate. Absent = the channel is never auto-synced.", audioFormat: '`"m4a"`, `"mp3"` or `"opus"`: the audio a transcribe-handling download keeps.', diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,8 @@ # Changelog ## [Unreleased] +- **archive.org items are a source (`platform: "archiveorg"`).** A channel can hold recordings imported from archive.org and transcribe them like any transcribe channel. **Import video** takes an item page (`https://archive.org/details/<identifier>`) when the item holds one media file, or ONE file of a multi-file item (`…/details/<identifier>/<file>`); an item with several media files is refused with the way to choose files. `pnpm ops import-archive-org --json '{"slug":…,"item":…,"files":[…]}'` (or `"match": "<regex>"`, `"dryRun": true`) imports chosen files of one item as one drainable job. A whole item's id is its identifier; a file's is `<identifier>__<slug>-<hash>`, stable and unique per file. Each record keeps an `archiveorg.json` sidecar — the item's title, date, creator and collections, its torrent, and for a mirror of a YouTube upload the original's id, URL, title and upload date read from the info.json uploaded beside it — and its metadata takes the file's own page and title (and a mirror's original title and date), recorded in the metadata history as `archiveorg-provenance`. The video page says "Archived on archive.org: <item> · torrent" and, for a mirror, "Originally on YouTube: <url> (uploaded <date>)". The channel form offers archive.org in both platform lists. +- **Polite to archive.org.** archive.org runs on its own queue (`platform:archiveorg`), one download at a time, with a jittered pause of at least 8 s between files (the channel's or the global `sleepBetweenDownloadsSeconds` when longer). Its metadata API is asked once per item (cached for 6 h), with an identifying User-Agent, at most one request at a time and 2 s apart, honouring `Retry-After` and backing off exponentially on 429/503, stopping after four attempts. yt-dlp's archive.org spawns carry `--sleep-requests 2` and an exponential `--retry-sleep`; nothing downloads in parallel ranges. A file already downloaded is never fetched again, and a bulk import stops on a rate limit or after three failures in a row — re-running it resumes. The "auto" download format on archive.org takes the uploader's original mp4/mkv/webm, or an audio item's MP3. - **A report video's cue lookup names a site that publishes only its reports.** Pointing a report-to-video manifest at such a site (`corpus.json` `site.scope: "cited"`) used to fail with "channel … is not in corpus.json"; it now says the site publishes no transcripts and to use a full archive or a local corpus. The site form's **Publish** hint says a cited-only site is never listed on the homepage or the hub. - **A site's build composes its reports, and a site that publishes only its reports ships nothing else.** Every site's compose now writes the reports its `site.json` publishes: each report's page and its citations as `citations.json` and `citations.csv` under `/reports/<id>/`, its cited stills, a page per cited moment with the record, the transcript lines around the span and every report that cites it, and the clips and post captures `archilyzer reports prepare` made for it, only the cited ones. Each quote is checked against the record as it is composed (a span's against its cues within 5 s either side, read from `en-orig` when the `en` track has no cues; a post's against its text) and the score, time and method are written into the citation, replacing any typed by hand. The build stops with the list of every problem before anything is written: an invalid report, a citation of a channel outside the site or of a post the site may not carry, a missing record, still or post, a quote that matches less than 60 % of what the record says, and a citation without prepared media or with media cut for another span (`--allow-missing-media` on `archilyzer compose site` and `build site` lets those two through, without a clip). A site with `publish: "cited"` removes everything corpus-shaped from `export/public` before it writes its reports, and its built `out/` is checked against what a cited site may hold: anything else, a file over 25 MiB or more than 20,000 files fails the build, and every deploy path (the Publish tab, `deploy site`, Build & deploy, Build & deploy all, the container build) refuses it, as it refuses a site set to cited whose last build was a full one. The hub's compose removes a report site's files too. - **A site has a Reports tab.** `/sites/<site>/reports` lists every report under the site's `reports/` directory — the published ones in their order, then the drafts — with its kind, dates, sections, claims, citations by kind and, for a fact-check, how many claims carry each verdict. Each report's problems, from the same checker the prepare step and the build use, open under it. A draft with no problems can be published, and a published report moved up or down or unpublished; each writes only the site's `reports` list, applied to the list as it is on disk at that moment, so it never overwrites another change to the site. "Prepare evidence media" queues the `reports-prepare` job, and beside it the tab shows the last prepared media (moments by kind, total size, problems by kind) and links the last prepare job. What the site publishes (full or cited) is shown with a link to Settings, where it is changed. diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md @@ -1,6 +1,7 @@ # Changelog ## [Unreleased] +- **archive.org records play and are cited with their downloads.** A record imported from archive.org plays its file in the page's own player, which seeks to a cited second and follows the transcript. A citation of one links "archive.org" (the file's page) and its "torrent"; a citation of an archive.org mirror of a YouTube upload links the original on YouTube at the cited second, then "archive.org" and "torrent", on the citation cards and the moment pages. - **A report-only site with one report opens on that report.** Its home page is the report itself, its header links nothing, and `/reports/` forwards home: there is no index of one. With more reports the home page is the list, without repeating the site's title under the header; a list entry is the report's name, subtitle and dates (its counts and tally are on its page). Pages a report-only site does not have link home. - **A report can belong to a series.** `report.json` `series` leads the report's title in the accent colour, in place of the kind's label ("Fact-check"); a page title, a cited-in link, `llms.txt` and the MCP name it `<series>: <title>`. - **`pnpm start:export` serves a built site's moment pages.** It used `serve`, which listed a video or audio moment's directory (`3126.00-3151.00`) instead of serving its page; it now runs `export/scripts/serve-out.mjs`, which serves directories as Cloudflare Pages does, on `EXPORT_DEV_PORT` (3000). diff --git a/mcp/README.md b/mcp/README.md @@ -53,7 +53,7 @@ compact `[mm:ss|<seconds>]`. The expansion rule: **full moment link = `[title @ 2:36](<moment_base>156)`. The seconds are floored exactly like the inline links', so both styles cite the identical second. A video whose base can't be built (no viewer origin and a platform whose time param doesn't take -raw seconds — Twitch — or doesn't exist — Rumble/Kick) omits the line: cite its +raw seconds — Twitch — or doesn't exist — Rumble/Kick/archive.org) omits the line: cite its `- source:` URL plain instead. `open_link` results and `get_transcript` stay inline-linked (future work).