Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit ad31a10124df4828c6a7a73e325382e49f660238
parent ccc2a271b480439c2f460ac9d38a122cb3b9bad8
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Tue,  6 Oct 2026 07:37:03 -0400

import: a BitChute or archive.org URL runs on its platform's queue from the channel page too; e2e for a BitChute import

The channel page's queue control starts at the channel's own queue and always
sends it, so an import from the page ran on the channel's queue, not the
platform's. That value now reads as "no choice made"; any other queue chosen
still wins.

e2e (import-video.spec.ts): a BitChute import into an Odysee transcribe
channel takes the canonical page, carries --sleep-requests 3, is labelled
from the BitChute extractor, runs on platform:bitchute, and a second import
is refused as already downloaded. fake-ytdlp writes BitChute's metadata for a
bitchute.com URL.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 2+-
Meditor/app/channels/[slug]/pipelineActions.ts | 16++++++++++------
Meditor/e2e/fixtures/bin/fake-ytdlp.mjs | 12++++++++++++
Meditor/e2e/import-video.spec.ts | 54+++++++++++++++++++++++++++++++++++++++++++++++++++++-
4 files changed, 76 insertions(+), 8 deletions(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -9,7 +9,7 @@ - **archive.org items are a source (`platform: "archiveorg"`).** A channel can hold recordings imported from archive.org and transcribe them like any transcribe channel. **Import video** takes an item page (`https://archive.org/details/<identifier>`) when the item holds one media file, or ONE file of a multi-file item (`…/details/<identifier>/<file>`); an item with several media files is refused with the way to choose files. `pnpm ops import-archive-org --json '{"slug":…,"item":…,"files":[…]}'` (or `"match": "<regex>"`, `"dryRun": true`) imports chosen files of one item as one drainable job. A whole item's id is its identifier; a file's is `<identifier>__<slug>-<hash>`, stable and unique per file. Each record keeps an `archiveorg.json` sidecar — the item's title, date, creator and collections, its torrent, and for a mirror of a YouTube upload the original's id, URL, title and upload date read from the info.json uploaded beside it — and its metadata takes the file's own page and title (and a mirror's original title and date), recorded in the metadata history as `archiveorg-provenance`. The video page says "Archived on archive.org: <item> · torrent" and, for a mirror, "Originally on YouTube: <url> (uploaded <date>)". The channel form offers archive.org in both platform lists. One file of a multi-file item with no uploaded info.json is dated by the `YYYYMMDD` its file name starts with (after any `<word>_`, a real calendar day only) and, when the item gives it no title of its own, titled from that name with the date, the `[<n> views]` count, the YouTube id and the extension taken off and ` _ ` read as ` | ` — the raw name stays in `archiveorg.json` as `file` — and `archilyzer archive-org refresh <slug> [--dry-run]` brings a channel's existing file records to the same title and date, offline, as `archiveorg-provenance` entries in the metadata history, printing old → new. A file with no date in its own name takes the one its folder starts with (`<YYYYMMDD>_<title>/…`), and any file of an item, not only a mirror, is dated that way. - **Polite to archive.org.** archive.org runs on its own queue (`platform:archiveorg`), one download at a time, with a jittered pause of at least 8 s between files (the channel's or the global `sleepBetweenDownloadsSeconds` when longer). Its metadata API is asked once per item (cached for 6 h), with an identifying User-Agent, at most one request at a time and 2 s apart, honouring `Retry-After` and backing off exponentially on 429/503, stopping after four attempts. A file already downloaded is never fetched again, and a bulk import stops on a rate limit or after three failures in a row — re-running it resumes. - **BitChute videos and channels are a source (`platform: "bitchute"`).** A BitChute URL (`bitchute.com`, `www.`, `old.`; `/video/<id>/` or `/embed/<id>/`) is detected as BitChute and its id is the video id; a record yt-dlp's BitChute extractor wrote is labelled "bitchute", plays its mp4 in the page's own player (seeking to a cited second, a link to the page when the file will not load), and is cited with the bare page (BitChute's watch page takes no start time). **Import video** takes a BitChute video page into any channel; a channel or playlist page is refused with the way to take one. A channel whose URL is a BitChute channel (`https://www.bitchute.com/channel/<name>/`) syncs like any other, listed through yt-dlp's BitChute channel extractor. The channel form offers BitChute in both platform lists, and the duplicates page names it. -- **Polite to BitChute.** BitChute runs on its own queue (`platform:bitchute`), one transfer at a time. Its yt-dlp spawns carry `--sleep-requests 3` and an exponential `--retry-sleep` (2 s doubling to 120 s); a video is one plain HTTP stream, never parallel ranges. Between two BitChute videos a batch download, a sync, a persist and the auto-download lane wait at least 60 s, plus up to half again at random (the channel's or the global `sleepBetweenDownloadsSeconds` when longer, plus any pace a rate limit added) — the per-platform floor is `PLATFORM_MIN_GAP_SECONDS` in `common/ytdlp/platformArgs.mjs`. A BitChute import asks nothing while BitChute is held or in a rate-limit cooldown, never fetches a video already on disk, and records what BitChute answered on its pacing state: a 429 backs every BitChute path off, a clean download settles it. +- **Polite to BitChute.** BitChute runs on its own queue (`platform:bitchute`), one transfer at a time. Its yt-dlp spawns carry `--sleep-requests 3` and an exponential `--retry-sleep` (2 s doubling to 120 s); a video is one plain HTTP stream, never parallel ranges. Between two BitChute videos a batch download, a sync, a persist and the auto-download lane wait at least 60 s, plus up to half again at random (the channel's or the global `sleepBetweenDownloadsSeconds` when longer, plus any pace a rate limit added) — the per-platform floor is `PLATFORM_MIN_GAP_SECONDS` in `common/ytdlp/platformArgs.mjs`. An import of a BitChute or archive.org URL runs on that platform's queue unless another queue than the channel's own is chosen in the queue control. A BitChute import asks nothing while BitChute is held or in a rate-limit cooldown, never fetches a video already on disk, and records what BitChute answered on its pacing state: a 429 backs every BitChute path off, a clean download settles it. - **A record indexed before its platform was known is relabelled by the next index build.** BitChute records downloaded before the bitchute platform existed carry "youtube" in the index and in their `transcript.cues.json`; the first index build after an update that teaches the app such a platform reads every stored summary once (from LMDB, no disk), re-derives the ones whose own page says otherwise from their metadata, and logs "Platform labels v1: N record(s) … re-derived." A normalized transcript whose summary is labelled so is re-derived the same way by the index and by report composition. The pass is recorded and does not run again, except while a channel is held. - **A report video's cue lookup names a site that publishes only its reports.** Pointed at such a site (`corpus.json` `site.scope: "cited"`), a report-to-video manifest's cue lookup says the site publishes no transcripts and to use a full archive or a local corpus. - **A site's build composes its reports, and a site that publishes only its reports ships nothing else.** Every site's compose writes the reports its `site.json` publishes: each report's page and its citations as `citations.json` and `citations.csv` under `/reports/<id>/`, its cited stills, a page per cited moment with the record, the transcript lines around the span and every report that cites it, and the clips and post captures `archilyzer reports prepare` made for it, only the cited ones. Each quote is checked against the record as it is composed (a span's against its cues within 5 s either side, read from `en-orig` when the `en` track has no cues; a post's against its text) and the score, time and method are written into the citation, replacing any typed by hand. The build stops with the list of every problem before anything is written: an invalid report, a citation of a channel outside the site or of a post the site may not carry, a missing record, still or post, a quote that matches less than 60 % of what the record says, and a citation without prepared media or with media cut for another span (`--allow-missing-media` on `archilyzer compose site` and `build site` lets those two through, without a clip). A site with search off (`search: false`) removes everything corpus-shaped from `export/public` before it writes its reports, and its built `out/` is checked against what a cited site may hold: anything else, a file over 25 MiB or more than 20,000 files fails the build, and every deploy path (the Publish tab, `deploy site`, Build & deploy, Build & deploy all, the container build) refuses it, as it refuses a site with search off whose last build was a full one. The hub's compose removes a report site's files too. diff --git a/editor/app/channels/[slug]/pipelineActions.ts b/editor/app/channels/[slug]/pipelineActions.ts @@ -585,14 +585,18 @@ export async function importVideoAction( const err = await lowDiskError(paths, slug); if (err) return err; const settings = getSettings(); + // A URL on a platform with its own queue runs there. The channel page's + // queue control starts at the CHANNEL's queue, which is therefore read as + // "no choice made"; any other queue the operator picks still wins. + const channelQueue = downloadQueueKey(channelConfig); + const ownQueue = archiveOrg || bitchute ? platformQueueKey(urlPlatform) : null; + const override = + ownQueue && queueKey !== undefined && queueKey.trim() === channelQueue + ? undefined + : queueKey; return runManagedFunction({ kind: "import-one", - queueKey: resolveQueueKey( - archiveOrg || bitchute - ? platformQueueKey(urlPlatform) - : downloadQueueKey(channelConfig), - queueKey, - ), + queueKey: resolveQueueKey(ownQueue ?? channelQueue, override), paths, channelSlug: slug, videoId, diff --git a/editor/e2e/fixtures/bin/fake-ytdlp.mjs b/editor/e2e/fixtures/bin/fake-ytdlp.mjs @@ -111,6 +111,17 @@ async function writeMetadata(videoDir, id, opts = {}) { ? `https://odysee.com/${id}` : `https://www.youtube.com/watch?v=${id}`, }; + // A bitchute.com URL: what yt-dlp's BitChute extractor writes — its key, the + // video page, and the one mp4 on a media host. + if (opts.bitchute) { + const file = `https://seed901.bitchute.com/FakeHash0001/${id}.mp4`; + meta.extractor = "BitChute"; + meta.extractor_key = "BitChute"; + meta.webpage_url = `https://www.bitchute.com/video/${id}/`; + meta.channel_url = "https://www.bitchute.com/channel/fakechannel/"; + meta.formats = [{ format_id: "0", ext: "mp4", url: file }]; + meta.url = file; + } // For the no-subs-fallback tests we need yt-dlp's metadata to reflect // whether the video advertises caption tracks. The default (no opts) // omits both keys so the rest of the existing suite keeps its behavior. @@ -171,6 +182,7 @@ function urlSentinels(url) { wasLive: lower.includes("waslive") || lower.includes("livevid"), longDuration: lower.includes("longvideo"), odysee: lower.includes("odyseevid"), + bitchute: /^https?:\/\/([a-z0-9-]+\.)?bitchute\.com\//.test(lower), }; } diff --git a/editor/e2e/import-video.spec.ts b/editor/e2e/import-video.spec.ts @@ -1,4 +1,4 @@ -import { readFile } from "node:fs/promises"; +import { readdir, readFile } from "node:fs/promises"; import { test, expect } from "@playwright/test"; import { channelStage, @@ -69,3 +69,55 @@ test("import rejects a non-URL input", async ({ page }) => { { timeout: 10_000 }, ); }); + +// A BitChute video imported into an Odysee (transcribe) channel runs the way BitChute's own +// downloads do: its canonical page, BitChute's queue (one transfer at a time) +// and yt-dlp args — the queue control's default (the channel's queue) is no +// choice — the record labelled from yt-dlp's BitChute extractor, and a second +// import of it refused without asking BitChute again. +test("a BitChute import runs on BitChute's queue and pace, and is never fetched twice", async ({ + page, +}) => { + test.setTimeout(90_000); + await resetData("one-transcribe-channel"); + await generateReport(page, "test-transcribe"); + await page.goto(channelStage("test-transcribe", "playlist")); + + const id = "bcOneoff0001"; + await page + .getByLabel("Video URL to import") + .fill(`https://old.bitchute.com/video/${id}/`); + const importBtn = page.getByRole("button", { name: "Import video" }); + await importBtn.click(); + + const log = page.getByLabel("Import video output"); + await expect(log).toContainText(`https://www.bitchute.com/video/${id}/`, { + timeout: 30_000, + }); + await expect(log).toContainText("--sleep-requests 3"); + + const metaPath = `test-transcripts/channels/test-transcribe/data/${id}/metadata.info.json`; + await expect.poll(() => pathExists(metaPath), { timeout: 30_000 }).toBe(true); + const meta = JSON.parse(await readFile(resolvePath(metaPath), "utf8")); + expect(meta.extractor_key).toBe("BitChute"); + + // The job ran on BitChute's queue, not the channel's. + await expect(async () => { + const dir = resolvePath("test-transcripts/.jobs"); + const queues: string[] = []; + for (const name of await readdir(dir)) { + if (!name.endsWith(".meta.json")) continue; + const m = JSON.parse(await readFile(`${dir}/${name}`, "utf8")); + if (m.kind === "import-one" && m.videoId === id) queues.push(m.queueKey); + } + expect(queues).toEqual(["platform:bitchute"]); + }).toPass({ timeout: 15_000 }); + + // Imported again: refused as already downloaded, nothing fetched. + await expect(importBtn).toBeEnabled({ timeout: 30_000 }); + await importBtn.click(); + await expect(page.getByLabel("Import video notice")).toContainText( + "Already downloaded", + { timeout: 10_000 }, + ); +});