commit 0ff9ad4f43e42d1b06b0838532ebbd65aafaffed
parent 8716c71e27bd55393ce4dd72e3d558bd98277eaa
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Mon, 5 Oct 2026 16:02:42 -0400
docs: archive.org over BitTorrent — README (aria2c, seeding defaults, the fallback), changelog
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Diffstat:
2 files changed, 25 insertions(+), 4 deletions(-)
diff --git a/README.md b/README.md
@@ -156,11 +156,31 @@ with it. The video page shows both; a citation of the record links **archive.org
and the **torrent** (and, for a mirror, the original **YouTube** upload at the cited
second), so a reader can fetch the file and check it.
+**Files come over BitTorrent when possible** — to be extra polite to archive.org. Every
+item has a torrent (`<identifier>_archive.torrent`) that lists archive.org itself as a
+web seed, so a torrent client takes what other peers have from them and only the rest
+from archive.org. With [aria2](https://aria2.github.io/) installed (`aria2c` on PATH, or
+`ARIA2C_BIN`; the Docker images include it), an import fetches just the chosen file
+from the item's torrent, then **seeds it for 10 minutes or to a ratio of 1, whichever
+comes first** — the job holds archive.org's queue while it seeds. Without aria2c, when
+the torrent does not carry the file, or when the torrent makes no progress for 5
+minutes, the file is downloaded directly from `archive.org/download/…` (resumed with a
+Range request, backing off on 429/503) and the log says why ("fell back to direct
+download: …"). Either way the file is checked against the item's sha1/md5; a mismatch
+is downloaded once more directly, and a second mismatch fails the record. No yt-dlp is
+involved: the record's metadata is written from the item's metadata API, and an audio
+item that yt-dlp could not read imports like any other. The defaults are
+settings.json's `archiveOrg` block (`torrent`, `seedMinutes`, `seedRatio`,
+`stallMinutes`, `maxPeers`, `maxDownloadKiBps`, `maxUploadKiBps` —
+[SETTINGS.md](SETTINGS.md#archiveorg)); `"torrent": false` always downloads directly.
+`archilyzer doctor` reports whether aria2c is there.
+
**It is polite to archive.org**: its own job queue, one download at a time with a
-jittered pause of at least 8 s between files, the item's metadata asked once (cached),
-an identifying User-Agent, `Retry-After` and exponential backoff honoured, a stop
-after repeated failures, and a file on disk never fetched again
-(`common/lib/archiveOrgClient.ts`, `common/controller/archiveOrgImport.ts`).
+jittered pause of at least 8 s between files, the item's metadata and torrent asked
+once (cached), an identifying User-Agent, `Retry-After` and exponential backoff
+honoured, a stop after repeated failures, and a file on disk never fetched again
+(`common/lib/archiveOrgClient.ts`, `common/controller/archiveOrgImport.ts`,
+`common/controller/archiveOrgDownload.ts`).
Wayback Machine captures (web.archive.org pages, WARC records) are a different kind of
record and are not handled by this; they would be a source of their own beside
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -1,6 +1,7 @@
# Changelog
## [Unreleased]
+- **archive.org files come over BitTorrent when possible, else straight from archive.org — never through yt-dlp.** An archive.org import fetched the record with yt-dlp, whose archive.org extractor fails on some items (an audio item: "opening play-av tag not found"). Now the chosen file is fetched from the item's own torrent (`<identifier>_archive.torrent`, which lists archive.org as a web seed, so other peers take load off archive.org) with aria2c, only that file of the item, and seeded afterwards for 10 minutes or to a ratio of 1, whichever comes first; the log shows "torrent: <file> (n of m pieces, peers p, web seed yes)" and "seeding 10 min…". With no aria2c, a torrent that does not carry the file, or no progress for 5 minutes, it is downloaded directly from `archive.org/download/…` instead (resumable, backing off on 429/503), and the log says "fell back to direct download: <reason>". Every file is checked against archive.org's sha1/md5: a mismatch is downloaded once more directly, a second one fails the record. The record is written from the item's metadata: `metadata.info.json` with the file's page, the canonical id, the duration ffprobe measures and archive.org's playable copies of the file, the `archiveorg.json` provenance as before (a mirror's original title, date and uploader), and `audio.<fmt>` — an audio file already in the channel's format is used as is, anything else goes through the app's audio extraction, a video kept in the saved-video store when the channel keeps sources. An .avi/.mpeg/.flac/.wav original is fetched as archive.org's mp4 or mp3 of it. aria2c runs in its own process group: cancelling the job stops it and everything it started, and it stops itself if the editor exits. New settings block `archiveOrg` (`torrent`, `seedMinutes`, `seedRatio`, `stallMinutes`, `maxPeers`, `maxDownloadKiBps`, `maxUploadKiBps`), `ARIA2C_BIN`, an aria2c row in `archilyzer doctor`, and `aria2` in the runtime Docker images.
- **A forum thread can be archived as a posts source.** A XenForo thread URL (`…/threads/<title>.<id>/`; Kiwi Farms is recognised by host) makes a forum-thread channel — platform "xenforo", one channel per thread, each forum post a post — searchable and readable like X and Bluesky posts, in the editor, the export and the MCP (`get_thread` gives a forum post's conversation: the posts it quotes and the posts quoting it). **Fetch posts** reads the thread in a headless browser, newest page first, one page at a time with a 10–20 s pause (the channel key `postPagePauseSeconds` sets it), and stops at already-archived posts; a **Latest N pages** box (`archilyzer posts fetch --pages N`) caps a run, and the next run continues where it stopped. The browser keeps one profile per forum host, so a browser check it clears once (KiwiFlare's proof of work, say) stays cleared; a check that does not clear within a minute, a captcha, a login wall or a refusal stops the run with the reason and keeps its place — never retried at once. **Connect forum session** on the channel page opens that profile in a window on the editor's machine, at the thread, for the operator to clear it or log in. **Import saved pages** (`archilyzer posts import-html <slug> <file-or-dir>…`) reads thread pages saved from a browser ("Save page as", complete or HTML only) through the same parser: new posts are added and a post saved again after an edit is updated. **Capture posts** works on forum posts: a screenshot of the post and its attached files, through the same profile. A forum post keeps its thread title, page, position, author id, last-edit time, quoted posts and its media links; quoted text is marked with "> " lines.
- **A site's reports can be exported as files a reader saves and hosts again.** `archilyzer reports export <site> [--report <id>] [--formats html,pdf,md,zip]`, the `reports-export` job (`POST /api/ops/reports-export`, `pnpm ops reports-export`, and **Export reports** on a site's Reports tab) write each published report, checked as the build checks it, into `.export-index/sites/<site>/report-exports/<report>/`: `report.html`, one self-contained page (its own style, no script, stills and post screenshots inlined and recompressed, clips linked on the site); `report.pdf`, that page printed by headless Chromium, skipped with a note where there is none; `report.md`, plain Markdown with numbered references; and `evidence-pack.zip`, the page with its clips, stills and screenshots as files plus the Markdown and the citations, packed by the system `zip` (a host without it fails that format, naming it). An `export.json` names each file's size and checksum and the checksum of the report.json it was made from; every export ends with the report's date and the start of that checksum. Preparing the evidence media now exports at its end when nothing is missing, on the same queue. The build publishes an export beside the report only when it was made from the report as it is now and is at most 24 MiB — a larger evidence pack stays local — and the Reports tab lists each report's exports, their sizes and which the next build publishes. The 24 MiB limit the source mirror and the evidence clips already kept is now one shared number.
- **archive.org items are a source (`platform: "archiveorg"`).** A channel can hold recordings imported from archive.org and transcribe them like any transcribe channel. **Import video** takes an item page (`https://archive.org/details/<identifier>`) when the item holds one media file, or ONE file of a multi-file item (`…/details/<identifier>/<file>`); an item with several media files is refused with the way to choose files. `pnpm ops import-archive-org --json '{"slug":…,"item":…,"files":[…]}'` (or `"match": "<regex>"`, `"dryRun": true`) imports chosen files of one item as one drainable job. A whole item's id is its identifier; a file's is `<identifier>__<slug>-<hash>`, stable and unique per file. Each record keeps an `archiveorg.json` sidecar — the item's title, date, creator and collections, its torrent, and for a mirror of a YouTube upload the original's id, URL, title and upload date read from the info.json uploaded beside it — and its metadata takes the file's own page and title (and a mirror's original title and date), recorded in the metadata history as `archiveorg-provenance`. The video page says "Archived on archive.org: <item> · torrent" and, for a mirror, "Originally on YouTube: <url> (uploaded <date>)". The channel form offers archive.org in both platform lists.