Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit ccc2a271b480439c2f460ac9d38a122cb3b9bad8
parent 63fb2c0083af66b91d6e667e88e98aac51472e6b
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Tue,  6 Oct 2026 07:29:05 -0400

docs: BitChute — README section, CHANNEL.md platform list, MCP README, changelogs

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
MCHANNEL.md | 2+-
MREADME.md | 21+++++++++++++++++++++
Mcommon/lib/channelConfig.ts | 2+-
Meditor/CHANGELOG.md | 3+++
Mexport/CHANGELOG.md | 1+
Mmcp/README.md | 2+-
6 files changed, 28 insertions(+), 3 deletions(-)

diff --git a/CHANNEL.md b/CHANNEL.md @@ -19,7 +19,7 @@ Regenerate this file with `pnpm --filter yt-dlp-transcript-common exec tsx bin/f | `postFetcher` | config | Social channels only: which social fetcher drives ingest (e.g. `"bluesky-atproto"`, `"x-gallery-dl"`). Absent = resolve by URL detection. Trimmed. | | `socialHandle` | config | Social channels only: the bare account handle (a leading "@" is stripped). Derived from `url` at creation but stored, so a later URL-format change upstream cannot silently re-point ingest at a different account. | | `postPagePauseSeconds` | config | Social channels that are read page by page (a forum thread) only: the pause between two page loads, in seconds; each pause is jittered to 0.85–1.65× of it. Absent = the fetcher's own (12 s, so 10–20 s); floored at 5, capped at 600. | -| `platform` | config | The source platform: youtube, rumble, odysee, twitch, kick, archiveorg (archive.org items — imported, never listed; see README.md, "archive.org items"), twitter, bluesky or xenforo (a forum thread). An unknown value is dropped. | +| `platform` | config | The source platform: youtube, rumble, odysee, twitch, kick, archiveorg (archive.org items — imported, never listed; see README.md, "archive.org items"), bitchute (BitChute videos and channels; see README.md, "BitChute"), twitter, bluesky or xenforo (a forum thread). An unknown value is dropped. | | `name` | config | Display name. | | `url` | config | The channel / playlist / account URL syncs enumerate. Absent = the channel is never auto-synced. | | `audioFormat` | config | `"m4a"`, `"mp3"` or `"opus"`: the audio a transcribe-handling download keeps. | diff --git a/README.md b/README.md @@ -191,6 +191,27 @@ Wayback Machine captures (web.archive.org pages, WARC records) are a different k record and are not handled by this; they would be a source of their own beside `common/lib/archiveOrg.ts`. +### BitChute + +BitChute is a platform like YouTube, Rumble and Odysee (`platform: "bitchute"`): + +- **A channel**: make a channel whose URL is the BitChute channel + (`https://www.bitchute.com/channel/<name>/`) and sync it — yt-dlp lists its videos. +- **One video**: **Import video** on any channel with its page + (`https://www.bitchute.com/video/<id>/`), or + `pnpm ops import-video --json '{"slug":"<channel>","url":"https://www.bitchute.com/video/<id>/"}'`. + +A record plays its mp4 in the archive's own player, which seeks to a cited second +(BitChute's embed does not); a citation links the BitChute page, which takes no start +time. + +**It is polite to BitChute**, which rate-limits: its own job queue, one transfer at a +time; `--sleep-requests 3` and an exponential `--retry-sleep` on every yt-dlp spawn; at +least 60 s, jittered, between two videos (`sleepBetweenDownloadsSeconds` when longer); +nothing asked while BitChute is in a rate-limit cooldown, which a 429 starts; and a video +on disk never fetched again (`common/ytdlp/platformArgs.mjs`, +`common/controller/bitchuteImport.ts`). + ### Requirements Always needed, to install and run the apps: diff --git a/common/lib/channelConfig.ts b/common/lib/channelConfig.ts @@ -160,7 +160,7 @@ export const CHANNEL_CONFIG_FIELD_DOCS: FieldDocs<ChannelConfig> = { 'Social channels only: the bare account handle (a leading "@" is stripped). Derived from `url` at creation but stored, so a later URL-format change upstream cannot silently re-point ingest at a different account.', postPagePauseSeconds: "Social channels that are read page by page (a forum thread) only: the pause between two page loads, in seconds; each pause is jittered to 0.85–1.65× of it. Absent = the fetcher's own (12 s, so 10–20 s); floored at 5, capped at 600.", - platform: "The source platform: youtube, rumble, odysee, twitch, kick, archiveorg (archive.org items — imported, never listed; see README.md, \"archive.org items\"), twitter, bluesky or xenforo (a forum thread). An unknown value is dropped.", + platform: "The source platform: youtube, rumble, odysee, twitch, kick, archiveorg (archive.org items — imported, never listed; see README.md, \"archive.org items\"), bitchute (BitChute videos and channels; see README.md, \"BitChute\"), twitter, bluesky or xenforo (a forum thread). An unknown value is dropped.", name: "Display name.", url: "The channel / playlist / account URL syncs enumerate. Absent = the channel is never auto-synced.", audioFormat: '`"m4a"`, `"mp3"` or `"opus"`: the audio a transcribe-handling download keeps.', diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -8,6 +8,9 @@ - **A site's reports can be exported as files a reader saves and hosts again.** `archilyzer reports export <site> [--report <id>] [--formats html,pdf,md,zip]`, the `reports-export` job (`POST /api/ops/reports-export`, `pnpm ops reports-export`, and **Export reports** on a site's Reports tab) write each published report, checked as the build checks it, into `.export-index/sites/<site>/report-exports/<report>/`: `report.html`, one self-contained page (its own style, no script, stills and post screenshots inlined and recompressed, clips linked on the site); `report.pdf`, that page printed by headless Chromium, skipped with a note where there is none; `report.md`, plain Markdown with numbered references; and `evidence-pack.zip`, the page with its clips, stills and screenshots as files plus the Markdown and the citations, packed by the system `zip` (a host without it fails that format, naming it). An `export.json` names each file's size and checksum and the checksum of the report.json it was made from; every export ends with the report's date and the start of that checksum. Preparing the evidence media exports at its end when nothing is missing, on the same queue. The build publishes an export beside the report only when it was made from the report as it is now and is at most 24 MiB — a larger evidence pack stays local — and the Reports tab lists each report's exports, their sizes and which the next build publishes. The 24 MiB limit is one number, shared with the source mirror and the evidence clips. - **archive.org items are a source (`platform: "archiveorg"`).** A channel can hold recordings imported from archive.org and transcribe them like any transcribe channel. **Import video** takes an item page (`https://archive.org/details/<identifier>`) when the item holds one media file, or ONE file of a multi-file item (`…/details/<identifier>/<file>`); an item with several media files is refused with the way to choose files. `pnpm ops import-archive-org --json '{"slug":…,"item":…,"files":[…]}'` (or `"match": "<regex>"`, `"dryRun": true`) imports chosen files of one item as one drainable job. A whole item's id is its identifier; a file's is `<identifier>__<slug>-<hash>`, stable and unique per file. Each record keeps an `archiveorg.json` sidecar — the item's title, date, creator and collections, its torrent, and for a mirror of a YouTube upload the original's id, URL, title and upload date read from the info.json uploaded beside it — and its metadata takes the file's own page and title (and a mirror's original title and date), recorded in the metadata history as `archiveorg-provenance`. The video page says "Archived on archive.org: <item> · torrent" and, for a mirror, "Originally on YouTube: <url> (uploaded <date>)". The channel form offers archive.org in both platform lists. One file of a multi-file item with no uploaded info.json is dated by the `YYYYMMDD` its file name starts with (after any `<word>_`, a real calendar day only) and, when the item gives it no title of its own, titled from that name with the date, the `[<n> views]` count, the YouTube id and the extension taken off and ` _ ` read as ` | ` — the raw name stays in `archiveorg.json` as `file` — and `archilyzer archive-org refresh <slug> [--dry-run]` brings a channel's existing file records to the same title and date, offline, as `archiveorg-provenance` entries in the metadata history, printing old → new. A file with no date in its own name takes the one its folder starts with (`<YYYYMMDD>_<title>/…`), and any file of an item, not only a mirror, is dated that way. - **Polite to archive.org.** archive.org runs on its own queue (`platform:archiveorg`), one download at a time, with a jittered pause of at least 8 s between files (the channel's or the global `sleepBetweenDownloadsSeconds` when longer). Its metadata API is asked once per item (cached for 6 h), with an identifying User-Agent, at most one request at a time and 2 s apart, honouring `Retry-After` and backing off exponentially on 429/503, stopping after four attempts. A file already downloaded is never fetched again, and a bulk import stops on a rate limit or after three failures in a row — re-running it resumes. +- **BitChute videos and channels are a source (`platform: "bitchute"`).** A BitChute URL (`bitchute.com`, `www.`, `old.`; `/video/<id>/` or `/embed/<id>/`) is detected as BitChute and its id is the video id; a record yt-dlp's BitChute extractor wrote is labelled "bitchute", plays its mp4 in the page's own player (seeking to a cited second, a link to the page when the file will not load), and is cited with the bare page (BitChute's watch page takes no start time). **Import video** takes a BitChute video page into any channel; a channel or playlist page is refused with the way to take one. A channel whose URL is a BitChute channel (`https://www.bitchute.com/channel/<name>/`) syncs like any other, listed through yt-dlp's BitChute channel extractor. The channel form offers BitChute in both platform lists, and the duplicates page names it. +- **Polite to BitChute.** BitChute runs on its own queue (`platform:bitchute`), one transfer at a time. Its yt-dlp spawns carry `--sleep-requests 3` and an exponential `--retry-sleep` (2 s doubling to 120 s); a video is one plain HTTP stream, never parallel ranges. Between two BitChute videos a batch download, a sync, a persist and the auto-download lane wait at least 60 s, plus up to half again at random (the channel's or the global `sleepBetweenDownloadsSeconds` when longer, plus any pace a rate limit added) — the per-platform floor is `PLATFORM_MIN_GAP_SECONDS` in `common/ytdlp/platformArgs.mjs`. A BitChute import asks nothing while BitChute is held or in a rate-limit cooldown, never fetches a video already on disk, and records what BitChute answered on its pacing state: a 429 backs every BitChute path off, a clean download settles it. +- **A record indexed before its platform was known is relabelled by the next index build.** BitChute records downloaded before the bitchute platform existed carry "youtube" in the index and in their `transcript.cues.json`; the first index build after an update that teaches the app such a platform reads every stored summary once (from LMDB, no disk), re-derives the ones whose own page says otherwise from their metadata, and logs "Platform labels v1: N record(s) … re-derived." A normalized transcript whose summary is labelled so is re-derived the same way by the index and by report composition. The pass is recorded and does not run again, except while a channel is held. - **A report video's cue lookup names a site that publishes only its reports.** Pointed at such a site (`corpus.json` `site.scope: "cited"`), a report-to-video manifest's cue lookup says the site publishes no transcripts and to use a full archive or a local corpus. - **A site's build composes its reports, and a site that publishes only its reports ships nothing else.** Every site's compose writes the reports its `site.json` publishes: each report's page and its citations as `citations.json` and `citations.csv` under `/reports/<id>/`, its cited stills, a page per cited moment with the record, the transcript lines around the span and every report that cites it, and the clips and post captures `archilyzer reports prepare` made for it, only the cited ones. Each quote is checked against the record as it is composed (a span's against its cues within 5 s either side, read from `en-orig` when the `en` track has no cues; a post's against its text) and the score, time and method are written into the citation, replacing any typed by hand. The build stops with the list of every problem before anything is written: an invalid report, a citation of a channel outside the site or of a post the site may not carry, a missing record, still or post, a quote that matches less than 60 % of what the record says, and a citation without prepared media or with media cut for another span (`--allow-missing-media` on `archilyzer compose site` and `build site` lets those two through, without a clip). A site with search off (`search: false`) removes everything corpus-shaped from `export/public` before it writes its reports, and its built `out/` is checked against what a cited site may hold: anything else, a file over 25 MiB or more than 20,000 files fails the build, and every deploy path (the Publish tab, `deploy site`, Build & deploy, Build & deploy all, the container build) refuses it, as it refuses a site with search off whose last build was a full one. The hub's compose removes a report site's files too. - **A site has a Reports tab.** `/sites/<site>/reports` lists every report under the site's `reports/` directory — the published ones in their order, then the drafts — with its kind, dates, sections, claims, citations by kind and, for a fact-check, how many claims carry each verdict. Each report's problems, from the same checker the prepare step and the build use, open under it. A draft with no problems can be published, and a published report moved up or down or unpublished; each writes only the site's `reports` list, applied to the list as it is on disk at that moment, so it never overwrites another change to the site. "Prepare evidence media" queues the `reports-prepare` job, and beside it the tab shows the last prepared media (moments by kind, total size, problems by kind) and links the last prepare job. What the site publishes (full or cited) is shown with a link to Settings, where it is changed. diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md @@ -6,6 +6,7 @@ - **A report's claim can carry a flag, its header names the document under review, and a site with one report names it in the browser tab.** `report.json` claim `flag` (one line, at most 60 characters) shows as a small pill in the accent colour beside the claim's verdict, e.g. "No source given". On a report-only site with one report, the home page's tab title is the report's, as on the report's own page. A report's page header is its name — with a `series`, the series on one line in the accent colour and the title on the line below; without one, the title — then one small line of dates and the revision ("2026-10-04 · updated 2026-10-05 · revision 1"; "updated" only when it differs), then a card for the document under review: its title linking to the document, "<author> · <publisher> · <date>", and its archive links folded away, on a left rail in the document's colour (a source's `accent`, `"#rrggbb"`; without one, the border colour), then the subtitle. The page names no byline or site of its own. A claim that cites the document's own sentence shows that sentence — its still, else its words — on the same rail, with no link up to the card and no paraphrase beside it; a sentence of another document links "from <title>" to that document's box; a titled claim with no such sentence shows its text under the title, plain. A report's citation can say where its evidence came from (`origin`: `"subject"`, the document under review gave it; `"added"`, the report's author found it). A claim lists what the report added first, each card marked with the Archilyzer mark under its number ("Not in the article", or "Not in the source", is the mark's tooltip and what a screen reader says), then evidence of unknown origin, then what the document gave itself folded under "In the article (n)"; the reference list marks an added citation with the mark too, and a claim's flag pill wears the same mark. A fact-check's page reads in three tiers, each opened by a hairline with one, two or three dots: the quick take (the tally, the summary, and links to what the check found, every claim and the downloads); **What the check found**, every ruled claim grouped by verdict (contradicted, not found, partly, untestable, corroborated), one line each linking to the claim, with its `gist` (a new optional claim field, one line, at most 240 characters) and its flag; and **Every claim, with its evidence**, which opens with **How it was checked** (`method`, a new optional report field in markdown). A report of kind `sweep` has the first and last tiers only. Needs a rebuild and deploy of the site. - **Forum posts read like the other posts.** A post from a forum-thread channel shows its place in the thread (#N), an "edited" mark, the thread's title and its media as links, and opening its thread shows its conversation: the posts it quotes and the posts quoting it. - **archive.org records play and are cited with their downloads.** A record imported from archive.org plays its file in the page's own player, which seeks to a cited second and follows the transcript. A citation of one links "archive.org" (the file's page) and its "torrent"; a citation of an archive.org mirror of a YouTube upload links the original on YouTube at the cited second, then "archive.org" and "torrent", on the citation cards and the moment pages. +- **BitChute records play, and are labelled BitChute.** A record from BitChute plays its mp4 in the page's own player, which seeks to a cited second and follows the transcript, and opens the BitChute page when the file will not load. A citation of one links the BitChute page, with no start time. The duplicates page names the platform "BitChute". - **A report-only site with one report opens on that report.** Its home page is the report itself, its header links nothing, and `/reports/` forwards home: there is no index of one. With more reports the home page is the list, with no heading of its own; a list entry is the report's name, subtitle and dates (its counts and tally are on its page). Pages a report-only site does not have link home. - **A report can belong to a series.** `report.json` `series` is shown on its own line above the report's title, in the accent colour, at the head of the report's page and its downloads and in the report list, where a report without one shows its kind ("Fact-check"); a page title, a cited-in link, `llms.txt` and the MCP name it `<series>: <title>`. - **`pnpm start:export` serves a built site's moment pages.** It used `serve`, which listed a video or audio moment's directory (`3126.00-3151.00`) instead of serving its page; it now runs `export/scripts/serve-out.mjs`, which serves directories as Cloudflare Pages does, on `EXPORT_DEV_PORT` (3000). diff --git a/mcp/README.md b/mcp/README.md @@ -53,7 +53,7 @@ compact `[mm:ss|<seconds>]`. The expansion rule: **full moment link = `[title @ 2:36](<moment_base>156)`. The seconds are floored exactly like the inline links', so both styles cite the identical second. A video whose base can't be built (no viewer origin and a platform whose time param doesn't take -raw seconds — Twitch — or doesn't exist — Rumble/Kick/archive.org) omits the line: cite its +raw seconds — Twitch — or doesn't exist — Rumble/Kick/archive.org/BitChute) omits the line: cite its `- source:` URL plain instead. `open_link` results and `get_transcript` stay inline-linked (future work).