Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 9b649912101756b6a746db374b3b40af5966564d
parent 56ff1ce13a13dbaee0cae4b996415e97899762c2
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Tue,  6 Oct 2026 07:37:35 -0400

Merge sources/bitchute (BitChute as a platform: detection, polite import and pacing, native player, relabel pass)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
MCHANNEL.md | 2+-
MREADME.md | 21+++++++++++++++++++++
Mcommon/components/PlayerProvider.tsx | 14++++++++++----
Mcommon/controller/autoRunner.ts | 6+++++-
Acommon/controller/bitchuteImport.test.ts | 77+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/controller/bitchuteImport.ts | 52++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/controller/buildIndex.ts | 47++++++++++++++++++++++++++++++++++++++++++++++-
Acommon/controller/buildIndexPlatformLabels.test.ts | 201+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/controller/persistVideos.ts | 6+++++-
Mcommon/jobs/platformBackoff.test.ts | 17+++++++++++++++++
Mcommon/jobs/platformBackoff.ts | 14+++++++++++++-
Acommon/lib/bitchute.test.ts | 117+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/bitchute.ts | 40++++++++++++++++++++++++++++++++++++++++
Mcommon/lib/channelConfig.ts | 2+-
Mcommon/lib/detectPlatform.mjs | 3+++
Mcommon/lib/duplicates.ts | 1+
Mcommon/lib/momentUrl.ts | 14++++++++------
Mcommon/lib/platform.ts | 5+++++
Mcommon/lib/transcripts-server.ts | 41+++++++++++++++++++++++++++++++++++------
Mcommon/lib/videoId.ts | 6++++++
Mcommon/publish/composeReports.ts | 9+++++++--
Mcommon/ytdlp/channelArgs.test.ts | 21+++++++++++++++++++++
Mcommon/ytdlp/channelArgs.ts | 2++
Mcommon/ytdlp/downloadFormat.ts | 3++-
Mcommon/ytdlp/platformArgs.mjs | 41+++++++++++++++++++++++++++++++++++++++++
Mcommon/ytdlp/runYtdlp.ts | 10+++++++++-
Meditor/CHANGELOG.md | 3+++
Meditor/app/channels/[slug]/pipelineActions.ts | 89++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-----------
Meditor/app/channels/components/ChannelForm.tsx | 2++
Meditor/e2e/fixtures/bin/fake-ytdlp.mjs | 12++++++++++++
Meditor/e2e/import-video.spec.ts | 54+++++++++++++++++++++++++++++++++++++++++++++++++++++-
Mexport/CHANGELOG.md | 1+
Mexport/app/duplicates/DuplicatesClient.tsx | 1+
Mmcp/README.md | 2+-
34 files changed, 896 insertions(+), 40 deletions(-)

diff --git a/CHANNEL.md b/CHANNEL.md @@ -19,7 +19,7 @@ Regenerate this file with `pnpm --filter yt-dlp-transcript-common exec tsx bin/f | `postFetcher` | config | Social channels only: which social fetcher drives ingest (e.g. `"bluesky-atproto"`, `"x-gallery-dl"`). Absent = resolve by URL detection. Trimmed. | | `socialHandle` | config | Social channels only: the bare account handle (a leading "@" is stripped). Derived from `url` at creation but stored, so a later URL-format change upstream cannot silently re-point ingest at a different account. | | `postPagePauseSeconds` | config | Social channels that are read page by page (a forum thread) only: the pause between two page loads, in seconds; each pause is jittered to 0.85–1.65× of it. Absent = the fetcher's own (12 s, so 10–20 s); floored at 5, capped at 600. | -| `platform` | config | The source platform: youtube, rumble, odysee, twitch, kick, archiveorg (archive.org items — imported, never listed; see README.md, "archive.org items"), twitter, bluesky or xenforo (a forum thread). An unknown value is dropped. | +| `platform` | config | The source platform: youtube, rumble, odysee, twitch, kick, archiveorg (archive.org items — imported, never listed; see README.md, "archive.org items"), bitchute (BitChute videos and channels; see README.md, "BitChute"), twitter, bluesky or xenforo (a forum thread). An unknown value is dropped. | | `name` | config | Display name. | | `url` | config | The channel / playlist / account URL syncs enumerate. Absent = the channel is never auto-synced. | | `audioFormat` | config | `"m4a"`, `"mp3"` or `"opus"`: the audio a transcribe-handling download keeps. | diff --git a/README.md b/README.md @@ -191,6 +191,27 @@ Wayback Machine captures (web.archive.org pages, WARC records) are a different k record and are not handled by this; they would be a source of their own beside `common/lib/archiveOrg.ts`. +### BitChute + +BitChute is a platform like YouTube, Rumble and Odysee (`platform: "bitchute"`): + +- **A channel**: make a channel whose URL is the BitChute channel + (`https://www.bitchute.com/channel/<name>/`) and sync it — yt-dlp lists its videos. +- **One video**: **Import video** on any channel with its page + (`https://www.bitchute.com/video/<id>/`), or + `pnpm ops import-video --json '{"slug":"<channel>","url":"https://www.bitchute.com/video/<id>/"}'`. + +A record plays its mp4 in the archive's own player, which seeks to a cited second +(BitChute's embed does not); a citation links the BitChute page, which takes no start +time. + +**It is polite to BitChute**, which rate-limits: its own job queue, one transfer at a +time; `--sleep-requests 3` and an exponential `--retry-sleep` on every yt-dlp spawn; at +least 60 s, jittered, between two videos (`sleepBetweenDownloadsSeconds` when longer); +nothing asked while BitChute is in a rate-limit cooldown, which a 429 starts; and a video +on disk never fetched again (`common/ytdlp/platformArgs.mjs`, +`common/controller/bitchuteImport.ts`). + ### Requirements Always needed, to install and run the apps: diff --git a/common/components/PlayerProvider.tsx b/common/components/PlayerProvider.tsx @@ -55,12 +55,18 @@ const KickPlayer = dynamic(() => import("./KickPlayer"), { ssr: false, }); -// archive.org records play their file in a native <video> (FilePlayer): the -// embed takes no start time we can rely on, the file seeks. +// archive.org and BitChute records play their file in a native <video> +// (FilePlayer): neither embed takes a start time we can rely on, the file +// seeks. A file that will not load degrades to a link to the source page. const FilePlayer = dynamic(() => import("./FilePlayer"), { ssr: false, }); +// The platforms whose records play their own file (summary `mediaUrl`). +function playsFile(platform: Platform): boolean { + return platform === "archiveorg" || platform === "bitchute"; +} + type PlayerHandle = | Pick<ReactPlayerType, "seekTo"> | RumblePlayerHandle @@ -526,7 +532,7 @@ export function PlayerProvider({ const t = data.platform === "youtube" || data.platform === "kick" || - (data.platform === "archiveorg" && data.mediaUrl) + (playsFile(data.platform) && data.mediaUrl) ? currentTime : (urlTime ?? 0); const secs = Math.max(0, Math.floor(t)); @@ -1037,7 +1043,7 @@ export function PlayerProvider({ } onError={() => setKickError(true)} /> - ) : data.platform === "archiveorg" && data.mediaUrl && !fileError ? ( + ) : playsFile(data.platform) && data.mediaUrl && !fileError ? ( <FilePlayer key={data.id} ref={(p: FilePlayerHandle | null) => { diff --git a/common/controller/autoRunner.ts b/common/controller/autoRunner.ts @@ -80,7 +80,10 @@ import { prunePlatformPacing, pruneSubtitleDeferrals, } from "../jobs/platformBackoff"; -import { staticSleepRequestsSeconds } from "../ytdlp/platformArgs.mjs"; +import { + platformMinGapSeconds, + staticSleepRequestsSeconds, +} from "../ytdlp/platformArgs.mjs"; import { applyUnitOutcome } from "../jobs/unitOutcome"; import { type DownloadFailureClass } from "../lib/availability"; import { resolveCookiePolicy } from "../lib/cookiePolicy"; @@ -2048,6 +2051,7 @@ async function runLoop( unitSettings.sleepBetweenDownloadsSeconds, currentPaceSeconds(kindState.platformPace, unitPlatform, base), base, + { minSeconds: platformMinGapSeconds(unitPlatform) }, ); if (gap > 0) platformNextStartAt.set(unitPlatform, settledAt + gap); else platformNextStartAt.delete(unitPlatform); diff --git a/common/controller/bitchuteImport.test.ts b/common/controller/bitchuteImport.test.ts @@ -0,0 +1,77 @@ +// Run with: pnpm --filter yt-dlp-transcript-common exec tsx --test controller/bitchuteImport.test.ts + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { importPlatformSignal, resolveBitchuteImportUrl } from "./bitchuteImport"; +import type { DownloadOutcomeRecord } from "../lib/downloadOutcome"; + +const ID = "Zq3xVb7Kp2Lm"; +const PAGE = `https://www.bitchute.com/video/${ID}/`; + +test("every URL of one video resolves to its canonical page", () => { + for (const url of [ + PAGE, + `https://www.bitchute.com/video/${ID}`, + `https://old.bitchute.com/video/${ID}/`, + `https://bitchute.com/embed/${ID}/`, + ]) { + assert.deepEqual(resolveBitchuteImportUrl(url), { ok: true, id: ID, url: PAGE }, url); + } +}); + +test("a channel or playlist page is refused with the way to take one", () => { + for (const url of [ + "https://www.bitchute.com/channel/examplechannel/", + "https://www.bitchute.com/playlist/Pl4yL1st0001/", + ]) { + const r = resolveBitchuteImportUrl(url); + assert.equal(r.ok, false, url); + assert.match((r as { error: string }).error, /make it a channel with that URL and sync it/); + } +}); + +function outcome(over: Partial<DownloadOutcomeRecord>): DownloadOutcomeRecord { + return { + videoId: ID, + status: "ok", + startedAt: "2026-01-01T00:00:00.000Z", + finishedAt: "2026-01-01T00:01:00.000Z", + attempts: [], + ...over, + }; +} + +test("a 429 backs BitChute off; a download that came down settles it; a per-video failure says nothing", () => { + assert.equal(importPlatformSignal(outcome({ status: "ok" })), "clean"); + assert.equal(importPlatformSignal(outcome({ status: "ok-auto-transcribed" })), "clean"); + assert.equal( + importPlatformSignal(outcome({ status: "failed", failureClass: "rate_limit" })), + "rate_limit", + ); + assert.equal( + importPlatformSignal(outcome({ status: "failed", failureClass: "network" })), + "network", + ); + // Classified from the last attempt when the record carries no class. + assert.equal( + importPlatformSignal( + outcome({ + status: "failed", + attempts: [ + { mode: "default", startedAt: "", finishedAt: "", error: "ERROR: [BitChute] x: HTTP Error 429: Too Many Requests" } as never, + ], + }), + ), + "rate_limit", + ); + assert.equal( + importPlatformSignal(outcome({ status: "failed", failureClass: "per_video" })), + null, + ); + assert.equal(importPlatformSignal(outcome({ status: "skipped-filtered" })), null); + // The media came down; only a subtitle fetch was refused. + assert.equal( + importPlatformSignal(outcome({ status: "ok", failureClass: "subs_rate_limit" })), + null, + ); +}); diff --git a/common/controller/bitchuteImport.ts b/common/controller/bitchuteImport.ts @@ -0,0 +1,52 @@ +// A one-off BitChute import (import-video with a bitchute.com URL), made polite. +// +// BitChute rate-limits, so an import of one of its videos — into whatever +// channel — runs the way a BitChute channel's own downloads do: on BitChute's +// queue (`platform:bitchute`, one transfer at a time), with BitChute's yt-dlp +// args (`PLATFORM_ARGS.bitchute`), never while BitChute is held or in a +// rate-limit cooldown, never for a video already on disk, and what BitChute +// answered is recorded on its shared pacing state, so a 429 backs every +// BitChute path off and a clean download settles it. The editor's +// importVideoAction does the wiring; these are its pure rules. + +import { classifyDownloadFailure } from "../lib/availability"; +import { defaultWebpageUrl } from "../lib/platform"; +import { extractVideoId } from "../lib/videoId"; +import type { DownloadOutcomeRecord } from "../lib/downloadOutcome"; + +export type BitchuteImportUrl = + | { ok: true; id: string; url: string } + | { ok: false; error: string }; + +// The canonical page of the video a BitChute URL names (www., old. and embed +// URLs all become https://www.bitchute.com/video/<id>/). A channel, playlist or +// profile page names no one video: it is refused with the way to list one. +export function resolveBitchuteImportUrl(url: string): BitchuteImportUrl { + const id = extractVideoId(url); + if (!id || !/^[\w-]+$/.test(id)) { + return { + ok: false, + error: + "A BitChute import takes one video's page (https://www.bitchute.com/video/<id>/). " + + "To take a whole BitChute channel, make it a channel with that URL and sync it.", + }; + } + return { ok: true, id, url: defaultWebpageUrl("bitchute", id) }; +} + +// What one import's outcome tells BitChute's pacing state: a rate limit or a +// network failure backs the platform off, a download that came down settles +// it, and anything else (a per-video failure, a filter skip) tells it nothing. +export function importPlatformSignal( + outcome: Pick<DownloadOutcomeRecord, "status" | "failureClass" | "attempts">, +): "rate_limit" | "network" | "clean" | null { + if (outcome.status.startsWith("ok")) { + return outcome.failureClass === "subs_rate_limit" ? null : "clean"; + } + if (outcome.status !== "failed") return null; + const last = outcome.attempts[outcome.attempts.length - 1]; + const cls = + outcome.failureClass ?? + classifyDownloadFailure(last?.error ?? "", last?.availabilityClass); + return cls === "rate_limit" || cls === "network" ? cls : null; +} diff --git a/common/controller/buildIndex.ts b/common/controller/buildIndex.ts @@ -46,6 +46,7 @@ import { parseVtt, type Cue } from "../lib/vtt"; import { parseTranscriptJson } from "../lib/whisper"; import { parseLiveChat } from "../lib/liveChat"; import { + platformLabelStale, summarize, toDisplaySummary, type RawMetadata, @@ -192,6 +193,18 @@ import { // so the clearAsync() enumeration above is unchanged too. const SCHEMA_VERSION = 13; +// PLATFORM LABELS — also NOT a schema bump. When the app learns a platform +// whose records may already be indexed under another label (BitChute's were +// summarized "youtube", with no file to play), a bump would wipe the cache and +// re-read every video directory to fix a handful. Instead, the first build +// that sees a new PLATFORM_LABELS_VERSION reads every stored summary out of +// LMDB once (no disk), queues each one platformLabelStale() flags as changed, +// and records the version — only when no channel is held, so a held channel's +// records are still looked at once its text is back. Bump this whenever +// platformFromMetadata learns such a platform. +const PLATFORM_LABELS_VERSION = 1; +const PLATFORM_LABELS_KEY = "platformLabels"; + // Per-channel post stats, persisted so per-site aggregates survive a no-op // rebuild that doesn't re-encode the post pages. Mirrors ChannelSubsStat. type ChannelPostsStat = { @@ -764,6 +777,32 @@ export async function buildIndex({ ); } + // Summaries labelled before their platform was known (PLATFORM_LABELS_ + // VERSION): re-derived like a changed record. A schema bump has cleared + // every summary, so it has none to look at. + const relabelDue = + meta.get(PLATFORM_LABELS_KEY) !== PLATFORM_LABELS_VERSION; + if (relabelDue && !schemaBumped) { + const queued = new Set( + [...added, ...changed].map((s) => pathKeyId([s.channelSlug, s.videoDir])), + ); + let relabelled = 0; + for (const s of live) { + const pk: PathKey = [s.channelSlug, s.videoDir]; + if (queued.has(pathKeyId(pk))) continue; + const prev = mtimes.get(pk); + if (!prev) continue; + const sum = sums.get(prev.indexKey); + if (sum && platformLabelStale(sum)) { + changed.push(s); + relabelled++; + } + } + log( + `Platform labels v${PLATFORM_LABELS_VERSION}: ${relabelled} record(s) labelled before their platform was known, re-derived.`, + ); + } + const anyMutations = added.length > 0 || changed.length > 0 || removed.length > 0; @@ -846,6 +885,9 @@ export async function buildIndex({ channel: s.configName ?? rest.channel, }; cueList = cuesField; + // Normalized before its platform was known: the summary is + // re-derived from the metadata below, the cues are kept. + if (platformLabelStale(summary)) summary = undefined; } } if (!summary) { @@ -857,7 +899,7 @@ export async function buildIndex({ parsedMeta, s.configName, ); - if (s.transcriptMs !== null && s.transcriptKind) { + if (cueList === undefined && s.transcriptMs !== null && s.transcriptKind) { try { const raw = await readFile(s.transcriptPath, "utf8"); cueList = @@ -2287,6 +2329,9 @@ export async function buildIndex({ } for (const k of staleFpKeys) meta.remove(k); await meta.put(INDEX_SCANNED_AT_KEY, scanStartedAt); + if (relabelDue && held.size === 0) { + await meta.put(PLATFORM_LABELS_KEY, PLATFORM_LABELS_VERSION); + } await meta.flushed; await root.close(); diff --git a/common/controller/buildIndexPlatformLabels.test.ts b/common/controller/buildIndexPlatformLabels.test.ts @@ -0,0 +1,201 @@ +// Integration: a record labelled before its platform was known is relabelled +// by the index — through the REAL buildIndex, over a temp corpus. +// +// A BitChute record downloaded before the bitchute platform existed was +// summarized "youtube" (and with no file to play), and a transcribed one +// carries that summary frozen in its transcript.cues.json. Its files never +// change again, so the mtime diff alone would never look at it: the build's +// one-off platform-labels pass (PLATFORM_LABELS_VERSION) queues it, and the +// per-video step re-derives a stale normalized summary from the metadata. +// +// Run with: node_modules/.bin/tsx --test common/controller/buildIndexPlatformLabels.test.ts + +import { after, test } from "node:test"; +import assert from "node:assert/strict"; +import { mkdirSync, mkdtempSync, rmSync, utimesSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import path from "node:path"; + +const ROOT = mkdtempSync(path.join(tmpdir(), "build-index-labels-")); +const PINNED: Record<string, string> = { + TRANSCRIPTS_DIR: path.join(ROOT, "transcripts"), + SAVED_VIDEOS_DIR: path.join(ROOT, "saved-videos"), + SITES_DIR: path.join(ROOT, "transcripts", "sites"), + SETTINGS_FILE: path.join(ROOT, "settings.json"), + EXPORT_PUBLIC_DIR: path.join(ROOT, "public"), + EXPORT_INDEX_DIR: path.join(ROOT, ".export-index"), + EXPORT_BUILDS_DIR: path.join(ROOT, ".export-builds"), + EDITOR_CHANGELOG_FILE: path.join(ROOT, "editor-CHANGELOG.md"), + EXPORT_CHANGELOG_FILE: path.join(ROOT, "export-CHANGELOG.md"), + CHARTS_CONFIG_FILE: path.join(ROOT, "chart-templates.json"), + SEARCH_ALIASES_FILE: path.join(ROOT, "transcripts", "search-aliases.json"), + CURATED_TAGS_FILE: path.join(ROOT, "transcripts", "tags.json"), + ARCHILYZER_CONFIG_DIR: path.join(ROOT, "config"), + ARCHILYZER_SOURCE_SCRATCH: path.join(ROOT, "source-scratch"), +}; +Object.assign(process.env, PINNED); +delete process.env.ARCHILYZER_INDEX_ALLOW_HELD; +after(() => rmSync(ROOT, { recursive: true, force: true })); + +const { getPaths } = await import("../lib/paths"); +const { buildIndex } = await import("./buildIndex"); +const { open } = await import("lmdb"); + +const paths = getPaths(); +const CHANNEL = "example-channel"; +const SITE = "testsite"; +const BC = "Zq3xVb7Kp2Lm"; +const YT = "AbC123xyz_9"; +const BC_PAGE = `https://www.bitchute.com/video/${BC}/`; +const BC_FILE = `https://seed901.bitchute.com/AbCdEfGhIjKl/${BC}.mp4`; + +const writeJson = (file: string, value: unknown) => { + mkdirSync(path.dirname(file), { recursive: true }); + writeFileSync(file, JSON.stringify(value, null, 2)); +}; +const dirOf = (id: string) => path.join(paths.channelsDir, CHANNEL, "data", id); + +const VTT = + "WEBVTT\nKind: captions\nLanguage: en\n\n" + + "00:00:00.000 --> 00:00:05.000 align:start position:0%\n" + + "First<00:00:01.000><c> caption</c><00:00:02.000><c> line.</c>\n"; + +const CUES = [{ start: 0, end: 5, text: "Normalized line." }]; + +function seed(): void { + rmSync(paths.transcriptsDir, { recursive: true, force: true }); + rmSync(PINNED.EXPORT_INDEX_DIR, { recursive: true, force: true }); + writeFileSync(paths.settingsFile, "{}"); + writeJson(path.join(paths.channelsDir, CHANNEL, "config.json"), { + handling: "transcribe", + name: CHANNEL, + }); + writeJson(path.join(paths.sitesDir, SITE, "site.json"), { + siteId: SITE, + siteTitle: "Test Site", + siteDescription: "fixture", + headerTitle: "Test Site", + homeTagline: "", + socialLinks: [], + groups: [{ id: "default", name: "All channels", selectedByDefault: true }], + defaultGroupId: "default", + channels: [{ slug: CHANNEL, groupId: "default" }], + }); + writeJson(path.join(dirOf(YT), "metadata.info.json"), { + id: YT, + title: "A YouTube video", + upload_date: "20260601", + duration: 120, + webpage_url: `https://www.youtube.com/watch?v=${YT}`, + extractor_key: "Youtube", + }); + writeFileSync(path.join(dirOf(YT), "transcript.en.vtt"), VTT); + writeJson(path.join(dirOf(BC), "metadata.info.json"), { + id: BC, + title: "A BitChute video", + upload_date: "20260602", + duration: 61, + webpage_url: BC_PAGE, + extractor: "BitChute", + extractor_key: "BitChute", + formats: [{ format_id: "0", ext: "mp4", url: BC_FILE }], + url: BC_FILE, + }); + writeFileSync(path.join(dirOf(BC), "transcript.en.vtt"), VTT); + // The normalized transcript an earlier build of the app wrote: its summary + // says "youtube" and carries no file. Newer than everything, so fresh. + const cuesPath = path.join(dirOf(BC), "transcript.cues.json"); + writeJson(cuesPath, { + version: 2, + source: "vtt", + slug: `${CHANNEL}/${BC}`, + id: BC, + channelSlug: CHANNEL, + title: "A BitChute video", + uploadDate: "20260602", + duration: 61, + channel: CHANNEL, + description: "", + tags: [], + isLivestream: false, + ageRestricted: false, + platform: "youtube", + webpageUrl: BC_PAGE, + cues: CUES, + }); + const later = new Date(Date.now() + 10_000); + utimesSync(cuesPath, later, later); +} + +type Summary = { id: string; platform: string; mediaUrl?: string }; + +function withIndex<T>(fn: (db: (name: string) => ReturnType<ReturnType<typeof open>["openDB"]>) => T): T { + const root = open({ path: paths.lmdbPath, maxDbs: 18, compression: true }); + try { + return fn((name) => root.openDB({ name, encoding: "msgpack" })); + } finally { + root.close(); + } +} +const summaryOf = (id: string): Summary => + withIndex((db) => { + for (const { value } of db("sums").getRange()) { + if ((value as Summary).id === id) return value as Summary; + } + throw new Error(`${id} is not indexed`); + }); +const cuesOf = (id: string) => + withIndex((db) => { + for (const { key, value } of db("sums").getRange()) { + if ((value as Summary).id === id) return db("cues").get(key) as typeof CUES; + } + return undefined; + }); + +async function runIndex(): Promise<string[]> { + const log: string[] = []; + await buildIndex({ paths, onLog: (s) => log.push(s) }); + return log; +} + +test("a normalized summary frozen as youtube is re-derived as bitchute, with its file; its cues are kept", async () => { + seed(); + await runIndex(); + const bc = summaryOf(BC); + assert.equal(bc.platform, "bitchute"); + assert.equal(bc.mediaUrl, BC_FILE); + assert.deepEqual(cuesOf(BC), CUES); + assert.equal(summaryOf(YT).platform, "youtube"); +}); + +test("a summary already indexed under the old label is relabelled once, by the platform-labels pass", async () => { + // The index as an older build left it: the record stored as "youtube", and + // no platform-labels version recorded. + withIndex((db) => { + const sums = db("sums"); + for (const { key, value } of sums.getRange()) { + const v = value as Summary; + if (v.id === BC) { + const stale = { ...v, platform: "youtube" } as Summary; + delete stale.mediaUrl; + sums.putSync(key, stale); + } + } + db("meta").removeSync("platformLabels"); + }); + assert.equal(summaryOf(BC).platform, "youtube"); + + const log = await runIndex(); + assert.ok( + log.some((l) => /Platform labels v\d+: 1 record\(s\)/.test(l)), + log.join("\n"), + ); + const bc = summaryOf(BC); + assert.equal(bc.platform, "bitchute"); + assert.equal(bc.mediaUrl, BC_FILE); + assert.equal(summaryOf(YT).platform, "youtube"); + + // Recorded: the next build does not look again. + const again = await runIndex(); + assert.equal(again.some((l) => /Platform labels/.test(l)), false, again.join("\n")); +}); diff --git a/common/controller/persistVideos.ts b/common/controller/persistVideos.ts @@ -31,7 +31,10 @@ import { channelPlatform, pacingPlatformKey, } from "../ytdlp/channelArgs"; -import { staticSleepRequestsSeconds } from "../ytdlp/platformArgs.mjs"; +import { + platformMinGapSeconds, + staticSleepRequestsSeconds, +} from "../ytdlp/platformArgs.mjs"; import { downloadGapMs } from "../jobs/platformBackoff"; import { recordDownloadBackoff } from "../jobs/downloadBackoff"; @@ -408,6 +411,7 @@ export async function persistVideos({ config.sleepBetweenDownloadsSeconds ?? settings.sleepBetweenDownloadsSeconds, channelPaceSeconds(config), staticSleepRequestsSeconds(channelPlatform(config)), + { minSeconds: platformMinGapSeconds(channelPlatform(config)) }, ); if (gap > 0) { log(`Sleeping ${gap / 1000}s before the next download...`); diff --git a/common/jobs/platformBackoff.test.ts b/common/jobs/platformBackoff.test.ts @@ -198,6 +198,23 @@ test("the lane gap is the operator's sleep plus the pace above its base", () => assert.equal(downloadGapMs(-5, 0.5, 1), 0); }); +test("a platform floor raises the operator's sleep to it and jitters it by up to half again", () => { + // Below the floor: the floor, stretched by random * 50 %. + assert.equal(downloadGapMs(0, 3, 3, { minSeconds: 60, random: 0 }), 60_000); + assert.equal(downloadGapMs(30, 3, 3, { minSeconds: 60, random: 0.5 }), 75_000); + assert.equal(downloadGapMs(30, 3, 3, { minSeconds: 60, random: 1 }), 90_000); + // Above the floor the operator's sleep is the base; a raised pace still adds. + assert.equal(downloadGapMs(120, 3, 3, { minSeconds: 60, random: 0 }), 120_000); + assert.equal(downloadGapMs(0, 6, 3, { minSeconds: 60, random: 0 }), 63_000); + // Out-of-range randoms are clamped; an absent random still lands in range. + assert.equal(downloadGapMs(0, 3, 3, { minSeconds: 60, random: 7 }), 90_000); + const g = downloadGapMs(0, 3, 3, { minSeconds: 60 }); + assert.ok(g >= 60_000 && g <= 90_000, String(g)); + // No floor: exactly the old gap, never jittered. + assert.equal(downloadGapMs(30, 1, 1, { minSeconds: 0, random: 1 }), 30_000); + assert.equal(downloadGapMs(30, 1, 1, {}), 30_000); +}); + test("a held platform's backoff survives the prune; a hold with no backoff is dropped", () => { const now = 10 * BACKOFF_MAX_MS; const state = { diff --git a/common/jobs/platformBackoff.ts b/common/jobs/platformBackoff.ts @@ -303,12 +303,24 @@ export function currentPaceSeconds( // sleeps between two videos: the operator's `sleepBetweenDownloadsSeconds` // plus the pace ABOVE ITS BASE. The base pace is already paid inside every // spawn (`--sleep-requests`); what the gap adds is the part a rate limit added. +// +// A PLATFORM FLOOR (`floor.minSeconds`, the platform's PLATFORM_MIN_GAP_SECONDS +// in ytdlp/platformArgs.mjs — BitChute's 60 s): the operator's sleep is raised +// to at least the floor and then stretched by up to half again at random +// (`floor.random`, in [0, 1)), so a platform that rate-limits is never asked on +// a fixed beat. With no floor the gap is exactly what it always was. export function downloadGapMs( sleepBetweenDownloadsSeconds: number, paceSeconds: number, baseSeconds: number, + floor: { minSeconds?: number; random?: number } = {}, ): number { - const sleep = Math.max(0, sleepBetweenDownloadsSeconds); + let sleep = Math.max(0, sleepBetweenDownloadsSeconds); + const min = floor.minSeconds ?? 0; + if (min > 0) { + const r = Math.min(1, Math.max(0, floor.random ?? Math.random())); + sleep = Math.max(min, sleep) * (1 + 0.5 * r); + } const extra = Math.max(0, paceSeconds - baseSeconds); return Math.round((sleep + extra) * 1000); } diff --git a/common/lib/bitchute.test.ts b/common/lib/bitchute.test.ts @@ -0,0 +1,117 @@ +// BitChute as a platform: detection, ids, the extractor's label, the playable +// file, the queue, moment links and the "auto" format. Synthetic ids only. +// +// Run with: pnpm --filter yt-dlp-transcript-common exec tsx --test lib/bitchute.test.ts + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + PLATFORM_VALUES, + defaultWebpageUrl, + detectPlatform, + isSocialPlatform, + queueKeyForUrl, +} from "./platform"; +import { extractVideoId } from "./videoId"; +import { bitchutePlayableUrl } from "./bitchute"; +import { platformFromMetadata, summarize } from "./transcripts-server"; +import { platformMomentBaseUrl, platformMomentUrl } from "./momentUrl"; +import { resolveDownloadFormatSelector } from "../ytdlp/downloadFormat"; +import { dataDirIdForUrl } from "../ytdlp/runYtdlp"; + +const ID = "Zq3xVb7Kp2Lm"; +const PAGE = `https://www.bitchute.com/video/${ID}/`; +const FILE = `https://seed901.bitchute.com/AbCdEfGhIjKl/${ID}.mp4`; + +test("bitchute is a video platform", () => { + assert.ok(PLATFORM_VALUES.includes("bitchute")); + assert.equal(isSocialPlatform("bitchute"), false); +}); + +test("every BitChute host is detected as bitchute; look-alikes are not", () => { + for (const url of [ + PAGE, + `https://bitchute.com/video/${ID}/`, + `https://old.bitchute.com/video/${ID}/`, + `https://www.bitchute.com/embed/${ID}/`, + "https://www.bitchute.com/channel/examplechannel/", + "https://api.bitchute.com/api/beta/video", + FILE, + ]) { + assert.equal(detectPlatform(url), "bitchute", url); + } + assert.equal(detectPlatform("https://notbitchute.com/video/x/"), null); + assert.equal(detectPlatform("https://www.youtube.com/watch?v=AbC123xyz_9"), "youtube"); +}); + +test("a video or embed URL gives the video id; a channel, playlist or profile gives none", () => { + assert.equal(extractVideoId(PAGE), ID); + assert.equal(extractVideoId(`https://www.bitchute.com/video/${ID}`), ID); + assert.equal(extractVideoId(`https://old.bitchute.com/video/${ID}/?list=x`), ID); + assert.equal(extractVideoId(`https://www.bitchute.com/embed/${ID}/`), ID); + assert.equal(extractVideoId("https://www.bitchute.com/channel/examplechannel/"), null); + assert.equal(extractVideoId("https://www.bitchute.com/playlist/Pl4yL1st0001/"), null); + assert.equal(extractVideoId("https://www.bitchute.com/profile/Pr0f1le00001/"), null); + // A single-video spawn pins its output to data/<id>/. + assert.equal(dataDirIdForUrl(PAGE), ID); +}); + +test("the canonical page is /video/<id>/, and BitChute has its own queue", () => { + assert.equal(defaultWebpageUrl("bitchute", ID), PAGE); + assert.equal(queueKeyForUrl(PAGE), "platform:bitchute"); + assert.equal(queueKeyForUrl("https://www.bitchute.com/channel/examplechannel/"), "platform:bitchute"); +}); + +test("yt-dlp's BitChute extractors label a record bitchute, as does a BitChute page", () => { + assert.equal(platformFromMetadata({ extractor_key: "BitChute" }), "bitchute"); + assert.equal(platformFromMetadata({ extractor_key: "BitChuteChannel" }), "bitchute"); + assert.equal(platformFromMetadata({ extractor: "BitChute" }), "bitchute"); + assert.equal(platformFromMetadata({ extractor_key: "Generic", webpage_url: PAGE }), "bitchute"); + // An unknown host still lands where it always has. + assert.equal(platformFromMetadata({ extractor_key: "Generic", webpage_url: "https://example.com/a.mp4" }), "youtube"); +}); + +test("the playable file is the offered mp4 on a bitchute.com host", () => { + assert.equal(bitchutePlayableUrl({ formats: [{ format_id: "0", ext: "mp4", url: FILE }] } as never), FILE); + // The top-level url is the fallback. + assert.equal(bitchutePlayableUrl({ url: FILE, ext: "mp4" }), FILE); + // A live stream's manifest, or a file on another host, is not played. + assert.equal(bitchutePlayableUrl({ formats: [{ ext: "mp4", url: "https://seed901.bitchute.com/x/live.m3u8" }] }), undefined); + assert.equal(bitchutePlayableUrl({ formats: [{ ext: "mp4", url: "https://cdn.example.com/a.mp4" }] }), undefined); + assert.equal(bitchutePlayableUrl({}), undefined); +}); + +test("summarize: a BitChute record keeps its id and page and carries the file to play", () => { + const s = summarize("example-channel", ID, { + id: ID, + title: "A synthetic upload", + upload_date: "20240102", + duration: 61, + extractor: "BitChute", + extractor_key: "BitChute", + webpage_url: PAGE, + formats: [{ format_id: "0", ext: "mp4", url: FILE }], + url: FILE, + }); + assert.equal(s.platform, "bitchute"); + assert.equal(s.id, ID); + assert.equal(s.slug, `example-channel/${ID}`); + assert.equal(s.webpageUrl, PAGE); + assert.equal(s.mediaUrl, FILE); + // No page URL on record: the canonical one. + const bare = summarize("example-channel", ID, { id: ID, upload_date: "20240102", extractor_key: "BitChute" }); + assert.equal(bare.webpageUrl, PAGE); + assert.equal(bare.mediaUrl, undefined); +}); + +test("a moment on BitChute links the bare page: its watch page takes no start time", () => { + assert.equal(platformMomentUrl(PAGE, "bitchute", 125), PAGE); + assert.equal(platformMomentUrl(PAGE, null, 125), PAGE); + assert.equal(platformMomentBaseUrl(PAGE, "bitchute"), null); +}); + +test('"auto" takes BitChute\'s one file', () => { + // BitChute offers one format (an mp4 with no codec fields): `bestaudio` + // matches nothing and `worst` takes that file. + assert.equal(resolveDownloadFormatSelector("auto", "bitchute"), "bestaudio/worst"); +}); diff --git a/common/lib/bitchute.ts b/common/lib/bitchute.ts @@ -0,0 +1,40 @@ +// BitChute records — the pure rules that are BitChute's own. +// +// A BitChute video is fetched by yt-dlp's BitChute extractor, which offers ONE +// format: the uploaded mp4, served as a plain file from a `seedNNN.bitchute.com` +// media host (`https://seed123.bitchute.com/<channel hash>/<id>.mp4`). That file +// is what the viewer plays: BitChute's embed (bitchute.com/embed/<id>/) takes +// no start time it honours, and a native <video> over the file seeks +// (common/components/FilePlayer.tsx), as archive.org records do. + +type FormatLike = { url?: unknown; ext?: unknown; protocol?: unknown }; + +const MEDIA_HOST_RE = /^https:\/\/[a-z0-9-]+\.bitchute\.com\//i; +const PLAYABLE_EXT_RE = /\.(mp4|webm|m4v)(?:[?#]|$)/i; + +function playable(url: unknown, ext: unknown, protocol?: unknown): url is string { + if (typeof url !== "string" || !MEDIA_HOST_RE.test(url)) return false; + // An HLS manifest carries ext "mp4" too; it is a playlist, not a file. + if (/\.m3u8(?:[?#]|$)/i.test(url)) return false; + if (typeof protocol === "string" && protocol.startsWith("m3u8")) return false; + if (typeof ext === "string" && /^(mp4|webm|m4v)$/i.test(ext)) return true; + return PLAYABLE_EXT_RE.test(url); +} + +// The file a BitChute record's player plays: the first offered format that is +// a plain mp4/webm on a bitchute.com host, else the record's top-level `url` +// (yt-dlp writes the chosen format's there). An HLS manifest (a live stream) +// is not a file; undefined when there is none. +export function bitchutePlayableUrl(meta: { + formats?: unknown; + url?: unknown; + ext?: unknown; + protocol?: unknown; +}): string | undefined { + const formats = Array.isArray(meta.formats) ? (meta.formats as FormatLike[]) : []; + for (const f of formats) { + if (f && playable(f.url, f.ext, f.protocol)) return f.url as string; + } + if (playable(meta.url, meta.ext, meta.protocol)) return meta.url as string; + return undefined; +} diff --git a/common/lib/channelConfig.ts b/common/lib/channelConfig.ts @@ -160,7 +160,7 @@ export const CHANNEL_CONFIG_FIELD_DOCS: FieldDocs<ChannelConfig> = { 'Social channels only: the bare account handle (a leading "@" is stripped). Derived from `url` at creation but stored, so a later URL-format change upstream cannot silently re-point ingest at a different account.', postPagePauseSeconds: "Social channels that are read page by page (a forum thread) only: the pause between two page loads, in seconds; each pause is jittered to 0.85–1.65× of it. Absent = the fetcher's own (12 s, so 10–20 s); floored at 5, capped at 600.", - platform: "The source platform: youtube, rumble, odysee, twitch, kick, archiveorg (archive.org items — imported, never listed; see README.md, \"archive.org items\"), twitter, bluesky or xenforo (a forum thread). An unknown value is dropped.", + platform: "The source platform: youtube, rumble, odysee, twitch, kick, archiveorg (archive.org items — imported, never listed; see README.md, \"archive.org items\"), bitchute (BitChute videos and channels; see README.md, \"BitChute\"), twitter, bluesky or xenforo (a forum thread). An unknown value is dropped.", name: "Display name.", url: "The channel / playlist / account URL syncs enumerate. Absent = the channel is never auto-synced.", audioFormat: '`"m4a"`, `"mp3"` or `"opus"`: the audio a transcribe-handling download keeps.', diff --git a/common/lib/detectPlatform.mjs b/common/lib/detectPlatform.mjs @@ -54,6 +54,9 @@ export function detectPlatform(url) { // archive.org ITEMS only. Not web.archive.org (the Wayback Machine's page // captures are a different kind of record) — see lib/archiveOrgId.ts. if (host === "archive.org" || host === "www.archive.org") return "archiveorg"; + // bitchute.com, www. and old. (the pages), api. and the seedNNN. media + // hosts — every host BitChute serves from is the one platform's. + if (host === "bitchute.com" || host.endsWith(".bitchute.com")) return "bitchute"; if (XENFORO_HOSTS.some((h) => host === h || host.endsWith(`.${h}`))) { return "xenforo"; } diff --git a/common/lib/duplicates.ts b/common/lib/duplicates.ts @@ -263,6 +263,7 @@ const PLATFORM_PREFERENCE: ReadonlyArray<string> = [ "odysee", "kick", "twitch", + "bitchute", ]; function platformRank(platform: string): number { diff --git a/common/lib/momentUrl.ts b/common/lib/momentUrl.ts @@ -44,7 +44,8 @@ function twitchTime(totalSeconds: number): string { // Best-effort deep link into a platform's own watch page at `seconds`. Mirrors // the per-platform time params the in-app players build. Platforms whose watch -// page has no reliable start param (Rumble, Kick) get the bare `webpageUrl`. +// page has no reliable start param (Rumble, Kick, archive.org, BitChute) get +// the bare `webpageUrl`. // Returns null only when there is no `webpageUrl` to work from. export function platformMomentUrl( webpageUrl: string | null | undefined, @@ -71,10 +72,11 @@ export function platformMomentUrl( case "twitch": u.searchParams.set("t", twitchTime(secs)); return u.toString(); - // Rumble / Kick / archive.org / unknown: the watch page has no dependable - // start param — return the plain webpage URL rather than an invalid seek. - // (archive.org's own player takes none we can rely on; the archive's - // viewer plays the file itself and seeks it — PlayerProvider.) + // Rumble / Kick / archive.org / BitChute / unknown: the watch page has no + // dependable start param — return the plain webpage URL rather than an + // invalid seek. (archive.org's and BitChute's own players take none we can + // rely on; the archive's viewer plays the file itself and seeks it — + // PlayerProvider.) default: return webpageUrl; } @@ -166,7 +168,7 @@ export function viewerMomentBaseUrl( // webpage URL), a base MUST be appendable — so only platforms whose time param // takes raw seconds qualify. Twitch is excluded (its `t` takes an `XhYmZs` // token, so appending an integer would be an invalid seek); Rumble/Kick/ -// archive.org/unknown have no dependable start param at all. Null in every +// archive.org/BitChute/unknown have no dependable start param at all. Null in every // non-appendable case. export function platformMomentBaseUrl( webpageUrl: string | null | undefined, diff --git a/common/lib/platform.ts b/common/lib/platform.ts @@ -11,6 +11,8 @@ export type Platform = // archive.org items (lib/archiveOrgId.ts): a whole item, or one file inside // a multi-file item. | "archiveorg" + // BitChute videos (bitchute.com/video/<id>/) and channels. + | "bitchute" | "twitter" | "bluesky" | "xenforo"; @@ -22,6 +24,7 @@ export const PLATFORM_VALUES: ReadonlyArray<Platform> = [ "twitch", "kick", "archiveorg", + "bitchute", "twitter", "bluesky", "xenforo", @@ -63,6 +66,8 @@ export function defaultWebpageUrl(platform: Platform, id: string): string { const file = /^(.+?)__.*-[0-9a-f]{8}$/.exec(id); return `https://archive.org/details/${file ? file[1] : id}`; } + // BitChute's canonical id is its video id, which is also yt-dlp's. + if (platform === "bitchute") return `https://www.bitchute.com/video/${id}/`; // Social posts: /i/status/<id> resolves without knowing the handle. Bluesky // has no handle-free permalink, so this is only a last-resort fallback — // every archived post carries its own canonical `url` (see postPermalink). diff --git a/common/lib/transcripts-server.ts b/common/lib/transcripts-server.ts @@ -3,6 +3,7 @@ import path from "node:path"; import { formatDate, formatDuration } from "./format"; import { defaultWebpageUrl, detectPlatform } from "./platform"; import { archiveOrgPlayableUrl } from "./archiveOrg"; +import { bitchutePlayableUrl } from "./bitchute"; import { archiveOrgVideoIdFromNativeId } from "./archiveOrgId"; import type { DisplaySummary, Platform, TranscriptSummary } from "./transcripts"; import type { MediaType, VideoStat, VideoStatus } from "./stats"; @@ -41,10 +42,14 @@ export type RawMetadata = { // yt-dlp's coarse kind: "video" | "livestream" | "short" (YouTube). Absent on // platforms that don't distinguish, where we fall back to is/was_live. media_type?: string; - // Every format the extractor offered. Read only for archive.org, where each - // is a plain download URL and one of them is what the player plays - // (lib/archiveOrg.ts archiveOrgPlayableUrl). + // Every format the extractor offered. Read only for archive.org and + // BitChute, where each is a plain download URL and one of them is what the + // player plays (lib/archiveOrg.ts archiveOrgPlayableUrl, lib/bitchute.ts). formats?: unknown; + // The chosen format's URL, which yt-dlp writes at the top level. Read only + // for BitChute, as the playable file's fallback. + url?: unknown; + ext?: unknown; }; // Read and parse a video's metadata.info.json into the typed RawMetadata @@ -86,12 +91,33 @@ export function platformFromMetadata(meta: RawMetadata): Platform { if (/^twitch/i.test(key)) return "twitch"; if (/^kick/i.test(key)) return "kick"; if (/^archive\.?org$/i.test(key)) return "archiveorg"; + // `BitChute` (a video) and `BitChuteChannel` (a listing). + if (/^bitchute/i.test(key)) return "bitchute"; if (/^youtube/i.test(key)) return "youtube"; const fromPage = detectPlatform(meta.webpage_url); - if (fromPage === "archiveorg") return fromPage; + if (fromPage && PAGE_DECIDES.includes(fromPage)) return fromPage; return "youtube"; } +// The platforms whose page decides a record's label even under an extractor +// the app does not know (above). +const PAGE_DECIDES: ReadonlyArray<Platform> = ["archiveorg", "bitchute"]; + +// A SUMMARY LABELLED BEFORE ITS PLATFORM WAS KNOWN: its own page is on a +// platform whose page decides the label, and the label says otherwise — what a +// BitChute record summarized before the bitchute platform existed carries +// ("youtube", and no file to play). The index re-derives such a summary from +// the record's metadata (controller/buildIndex.ts), and so does every reader +// of a normalized transcript.cues.json, whose summary was frozen when it was +// written. Pure: the summary alone decides. +export function platformLabelStale(summary: { + platform?: Platform; + webpageUrl?: string; +}): boolean { + const fromPage = detectPlatform(summary.webpageUrl); + return fromPage !== null && PAGE_DECIDES.includes(fromPage) && fromPage !== summary.platform; +} + // The broad "is this a livestream (or stream VOD/upcoming)" notion used by the // coverage detector and the download-time duration guard, so both skip the same // content (stream captures have unreliable metadata durations). Distinct from @@ -141,10 +167,13 @@ export function summarize( webpageUrl: meta.webpage_url ?? defaultWebpageUrl(platform, id), // Kick VODs play from a persisted HLS manifest (no iframe embed exists). hlsUrl: platform === "kick" ? meta.manifest_url : undefined, - // archive.org plays the file itself in a native <video>, which seeks. + // archive.org and BitChute play the file itself in a native <video>, + // which seeks. ...(platform === "archiveorg" ? { mediaUrl: archiveOrgPlayableUrl(meta) } - : {}), + : platform === "bitchute" + ? { mediaUrl: bitchutePlayableUrl(meta) } + : {}), }; } diff --git a/common/lib/videoId.ts b/common/lib/videoId.ts @@ -44,6 +44,12 @@ export function extractVideoId(url: string): string | null { if (lastColon > 0) return decoded.slice(lastColon + 1); return null; } + if (host === "bitchute.com" || host.endsWith(".bitchute.com")) { + // A video page or its embed: /video/<id>/, /embed/<id>/ (www. or old.). + // A channel, playlist or profile page names no video. + const segs = u.pathname.split("/").filter(Boolean); + return (segs[0] === "video" || segs[0] === "embed") && segs[1] ? segs[1] : null; + } if (host.endsWith("twitch.tv")) { // VODs: /videos/<id> or legacy /<channel>/v/<id>; clips: // clips.twitch.tv/<slug> or /<channel>/clip/<slug>. Grab the segment diff --git a/common/publish/composeReports.ts b/common/publish/composeReports.ts @@ -66,7 +66,11 @@ import { assertChannelTextReadable } from "../lib/channelMedia"; import type { ChannelConfig } from "../lib/channelConfig"; import { readChannelConfig } from "../controller/channels"; import { isCuesJsonFresh, readNormalizedTranscript } from "../controller/normalizeTranscript"; -import { loadRawMetadataFromDir, summarize } from "../lib/transcripts-server"; +import { + loadRawMetadataFromDir, + platformLabelStale, + summarize, +} from "../lib/transcripts-server"; import type { TranscriptSummary } from "../lib/transcripts"; import { parseVtt, type Cue } from "../lib/vtt"; import { parseTranscriptJson } from "../lib/whisper"; @@ -275,7 +279,8 @@ export async function readCitedRecord( const fresh = await isCuesJsonFresh(dir); if (fresh.fresh) { const n = await readNormalizedTranscript(fresh.cuesPath); - if (n) { + // A summary frozen before its platform was known is re-derived below. + if (n && !platformLabelStale(n)) { summary = n; cues = n.cues ?? []; } diff --git a/common/ytdlp/channelArgs.test.ts b/common/ytdlp/channelArgs.test.ts @@ -5,7 +5,10 @@ import { channelPaceSeconds, pacedPlatformArgs, platformArgs, + platformArgsForUrl, PLATFORM_ARGS, + PLATFORM_MIN_GAP_SECONDS, + platformMinGapSeconds, staticSleepRequestsSeconds, withSleepRequests, } from "./channelArgs"; @@ -164,3 +167,21 @@ test("archive.org is paced and backs off, and keeps the fetched page as webpage_ // No parallel transfer is ever asked for. assert.equal(args.some((a) => /concurrent|downloader|^-N$/.test(a)), false); }); + +test("BitChute is paced hardest: 3 s between requests, exponential retry sleeps, a 60 s floor between videos", () => { + const args = platformArgs("bitchute"); + assert.equal(staticSleepRequestsSeconds("bitchute"), 3); + assert.deepEqual(args.slice(0, 2), ["--sleep-requests", "3"]); + assert.ok(args.includes("http:exp=2:120")); + assert.ok(args.includes("extractor:exp=2:120")); + assert.deepEqual(platformArgsForUrl("https://www.bitchute.com/video/Zq3xVb7Kp2Lm/"), args); + // No impersonation, and no parallel transfer is ever asked for. + assert.equal(args.includes("--impersonate"), false); + assert.equal(args.some((a) => /concurrent|downloader|^-N$/.test(a)), false); + assert.equal(platformMinGapSeconds("bitchute"), 60); + assert.equal(PLATFORM_MIN_GAP_SECONDS.bitchute, 60); + // Every other platform keeps the operator's gap. + for (const p of ["youtube", "rumble", "odysee", "archiveorg", "unknown", null]) { + assert.equal(platformMinGapSeconds(p), 0, String(p)); + } +}); diff --git a/common/ytdlp/channelArgs.ts b/common/ytdlp/channelArgs.ts @@ -39,6 +39,8 @@ export { PLATFORM_ARGS, platformArgs, platformArgsForUrl, + PLATFORM_MIN_GAP_SECONDS, + platformMinGapSeconds, staticSleepRequestsSeconds, withSleepRequests, } from "./platformArgs.mjs"; diff --git a/common/ytdlp/downloadFormat.ts b/common/ytdlp/downloadFormat.ts @@ -6,7 +6,8 @@ import type { Platform } from "../lib/platform"; // `original` format (every HLS rung is CDN-truncated to a few minutes), so auto // prefers `original` there, archive.org gets the uploader's original file // (ARCHIVE_ORG_AUTO_FORMAT_SELECTOR), and the historical `bestaudio/worst` -// everywhere else. +// everywhere else — BitChute included, whose one format (an mp4 with no codec +// fields) `bestaudio` never matches and `worst` takes. export type DownloadFormatPreset = | "auto" | "original" diff --git a/common/ytdlp/platformArgs.mjs b/common/ytdlp/platformArgs.mjs @@ -48,6 +48,16 @@ import { detectPlatform } from "../lib/detectPlatform.mjs"; // (reconcileVideoDirs.ts) — the file's record would be merged into the item's. // The import always fetches by the canonical file page, so the two agree; the // regex only takes an archive.org URL, and leaves any other untouched. +// +// bitchute: BitChute rate-limits (its API answered HTTP 429 on 2026-10-06), so +// it is paced harder than any other platform: `--sleep-requests 3` between +// the extractor's requests (two API calls per video, then a HEAD per media +// host it tries), and the same exponential `--retry-sleep` as archive.org so a +// refused or dropped request waits 2 s doubling to 120 s instead of retrying +// at once. A video is one plain mp4 over one HTTP stream — no fragments, no +// parallel ranges, and nothing here passes `-N` or an external downloader. No +// `--impersonate`: the extractor's API answers without a browser fingerprint. +// Between two VIDEOS the floor is PLATFORM_MIN_GAP_SECONDS below. /** @type {Readonly<Partial<Record<Platform, readonly string[]>>>} */ export const PLATFORM_ARGS = Object.freeze({ rumble: Object.freeze(["--impersonate", "chrome", "--sleep-requests", "1"]), @@ -62,9 +72,40 @@ export const PLATFORM_ARGS = Object.freeze({ "--parse-metadata", "original_url:(?P<webpage_url>https://archive\\.org/(?:details|embed|download)/.+)", ]), + bitchute: Object.freeze([ + "--sleep-requests", + "3", + "--retry-sleep", + "http:exp=2:120", + "--retry-sleep", + "extractor:exp=2:120", + ]), +}); + +// THE FLOOR UNDER THE GAP BETWEEN TWO VIDEOS on a platform, in seconds — for a +// platform that must be asked less often than the operator's +// `sleepBetweenDownloadsSeconds` would ask it. Every gap a batch download, a +// sync, a persist or the auto-download lane waits on that platform is at least +// this, plus up to half again at random so a run never settles into a fixed +// beat (jobs/platformBackoff.ts, downloadGapMs). A platform absent here keeps +// the operator's gap exactly, unjittered. +/** @type {Readonly<Partial<Record<Platform, number>>>} */ +export const PLATFORM_MIN_GAP_SECONDS = Object.freeze({ + bitchute: 60, }); /** + * @param {string | null | undefined} platform + * @returns {number} + */ +export function platformMinGapSeconds(platform) { + if (!platform) return 0; + /** @type {Record<string, number | undefined>} */ + const table = PLATFORM_MIN_GAP_SECONDS; + return table[platform] ?? 0; +} + +/** * @param {Platform | null | undefined} platform * @returns {string[]} */ diff --git a/common/ytdlp/runYtdlp.ts b/common/ytdlp/runYtdlp.ts @@ -19,6 +19,7 @@ import { channelPaceSeconds, channelPlatform, pacedPlatformArgs, + platformMinGapSeconds, staticSleepRequestsSeconds, } from "./channelArgs"; import { isRealAudioFile } from "../lib/videoStatus"; @@ -1077,6 +1078,9 @@ export async function runManagedDownloads( // rate limit doubled the platform's pace, a batch spaces its videos further // apart too. Read per gap, so a 429 in this batch slows the rest of it. const basePace = staticSleepRequestsSeconds(channelPlatform(effectiveChannelConfig)); + // A platform's floor under the gap (BitChute's 60 s, jittered) — see + // PLATFORM_MIN_GAP_SECONDS in platformArgs.mjs. + const minGap = platformMinGapSeconds(channelPlatform(effectiveChannelConfig)); const paceNow = deps.paceSeconds ?? (() => channelPaceSeconds(effectiveChannelConfig)); const recordSubs = deps.recordSubtitleDeferral ?? @@ -1223,7 +1227,9 @@ export async function runManagedDownloads( // declined slept 30 s after nothing but its metadata prefetch. Every // other outcome still sleeps — a real fetch, success or failure, and // every failure, per-video ones included (see declinedWithoutMediaFetch). - const gapMs = downloadGapMs(sleepSeconds, paceNow(), basePace); + const gapMs = downloadGapMs(sleepSeconds, paceNow(), basePace, { + minSeconds: minGap, + }); if ( gapMs > 0 && !isLast && @@ -1465,6 +1471,7 @@ async function downloadMissingSubs(opts: RunYtdlpOpts): Promise<void> { opts.channelConfig.sleepBetweenDownloadsSeconds ?? getSettings().sleepBetweenDownloadsSeconds; const subsBasePace = staticSleepRequestsSeconds(channelPlatform(opts.channelConfig)); + const subsMinGap = platformMinGapSeconds(channelPlatform(opts.channelConfig)); await Promise.all( tofetch.map((url, index) => limit(async () => { @@ -1475,6 +1482,7 @@ async function downloadMissingSubs(opts: RunYtdlpOpts): Promise<void> { subsSleepSeconds, channelPaceSeconds(opts.channelConfig), subsBasePace, + { minSeconds: subsMinGap }, ); if (gapMs > 0) { opts.onLog(`Sleeping ${gapMs / 1000}s before the next video...\n`); diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -8,6 +8,9 @@ - **A site's reports can be exported as files a reader saves and hosts again.** `archilyzer reports export <site> [--report <id>] [--formats html,pdf,md,zip]`, the `reports-export` job (`POST /api/ops/reports-export`, `pnpm ops reports-export`, and **Export reports** on a site's Reports tab) write each published report, checked as the build checks it, into `.export-index/sites/<site>/report-exports/<report>/`: `report.html`, one self-contained page (its own style, no script, stills and post screenshots inlined and recompressed, clips linked on the site); `report.pdf`, that page printed by headless Chromium, skipped with a note where there is none; `report.md`, plain Markdown with numbered references; and `evidence-pack.zip`, the page with its clips, stills and screenshots as files plus the Markdown and the citations, packed by the system `zip` (a host without it fails that format, naming it). An `export.json` names each file's size and checksum and the checksum of the report.json it was made from; every export ends with the report's date and the start of that checksum. Preparing the evidence media exports at its end when nothing is missing, on the same queue. The build publishes an export beside the report only when it was made from the report as it is now and is at most 24 MiB — a larger evidence pack stays local — and the Reports tab lists each report's exports, their sizes and which the next build publishes. The 24 MiB limit is one number, shared with the source mirror and the evidence clips. - **archive.org items are a source (`platform: "archiveorg"`).** A channel can hold recordings imported from archive.org and transcribe them like any transcribe channel. **Import video** takes an item page (`https://archive.org/details/<identifier>`) when the item holds one media file, or ONE file of a multi-file item (`…/details/<identifier>/<file>`); an item with several media files is refused with the way to choose files. `pnpm ops import-archive-org --json '{"slug":…,"item":…,"files":[…]}'` (or `"match": "<regex>"`, `"dryRun": true`) imports chosen files of one item as one drainable job. A whole item's id is its identifier; a file's is `<identifier>__<slug>-<hash>`, stable and unique per file. Each record keeps an `archiveorg.json` sidecar — the item's title, date, creator and collections, its torrent, and for a mirror of a YouTube upload the original's id, URL, title and upload date read from the info.json uploaded beside it — and its metadata takes the file's own page and title (and a mirror's original title and date), recorded in the metadata history as `archiveorg-provenance`. The video page says "Archived on archive.org: <item> · torrent" and, for a mirror, "Originally on YouTube: <url> (uploaded <date>)". The channel form offers archive.org in both platform lists. One file of a multi-file item with no uploaded info.json is dated by the `YYYYMMDD` its file name starts with (after any `<word>_`, a real calendar day only) and, when the item gives it no title of its own, titled from that name with the date, the `[<n> views]` count, the YouTube id and the extension taken off and ` _ ` read as ` | ` — the raw name stays in `archiveorg.json` as `file` — and `archilyzer archive-org refresh <slug> [--dry-run]` brings a channel's existing file records to the same title and date, offline, as `archiveorg-provenance` entries in the metadata history, printing old → new. A file with no date in its own name takes the one its folder starts with (`<YYYYMMDD>_<title>/…`), and any file of an item, not only a mirror, is dated that way. - **Polite to archive.org.** archive.org runs on its own queue (`platform:archiveorg`), one download at a time, with a jittered pause of at least 8 s between files (the channel's or the global `sleepBetweenDownloadsSeconds` when longer). Its metadata API is asked once per item (cached for 6 h), with an identifying User-Agent, at most one request at a time and 2 s apart, honouring `Retry-After` and backing off exponentially on 429/503, stopping after four attempts. A file already downloaded is never fetched again, and a bulk import stops on a rate limit or after three failures in a row — re-running it resumes. +- **BitChute videos and channels are a source (`platform: "bitchute"`).** A BitChute URL (`bitchute.com`, `www.`, `old.`; `/video/<id>/` or `/embed/<id>/`) is detected as BitChute and its id is the video id; a record yt-dlp's BitChute extractor wrote is labelled "bitchute", plays its mp4 in the page's own player (seeking to a cited second, a link to the page when the file will not load), and is cited with the bare page (BitChute's watch page takes no start time). **Import video** takes a BitChute video page into any channel; a channel or playlist page is refused with the way to take one. A channel whose URL is a BitChute channel (`https://www.bitchute.com/channel/<name>/`) syncs like any other, listed through yt-dlp's BitChute channel extractor. The channel form offers BitChute in both platform lists, and the duplicates page names it. +- **Polite to BitChute.** BitChute runs on its own queue (`platform:bitchute`), one transfer at a time. Its yt-dlp spawns carry `--sleep-requests 3` and an exponential `--retry-sleep` (2 s doubling to 120 s); a video is one plain HTTP stream, never parallel ranges. Between two BitChute videos a batch download, a sync, a persist and the auto-download lane wait at least 60 s, plus up to half again at random (the channel's or the global `sleepBetweenDownloadsSeconds` when longer, plus any pace a rate limit added) — the per-platform floor is `PLATFORM_MIN_GAP_SECONDS` in `common/ytdlp/platformArgs.mjs`. An import of a BitChute or archive.org URL runs on that platform's queue unless another queue than the channel's own is chosen in the queue control. A BitChute import asks nothing while BitChute is held or in a rate-limit cooldown, never fetches a video already on disk, and records what BitChute answered on its pacing state: a 429 backs every BitChute path off, a clean download settles it. +- **A record indexed before its platform was known is relabelled by the next index build.** BitChute records downloaded before the bitchute platform existed carry "youtube" in the index and in their `transcript.cues.json`; the first index build after an update that teaches the app such a platform reads every stored summary once (from LMDB, no disk), re-derives the ones whose own page says otherwise from their metadata, and logs "Platform labels v1: N record(s) … re-derived." A normalized transcript whose summary is labelled so is re-derived the same way by the index and by report composition. The pass is recorded and does not run again, except while a channel is held. - **A report video's cue lookup names a site that publishes only its reports.** Pointed at such a site (`corpus.json` `site.scope: "cited"`), a report-to-video manifest's cue lookup says the site publishes no transcripts and to use a full archive or a local corpus. - **A site's build composes its reports, and a site that publishes only its reports ships nothing else.** Every site's compose writes the reports its `site.json` publishes: each report's page and its citations as `citations.json` and `citations.csv` under `/reports/<id>/`, its cited stills, a page per cited moment with the record, the transcript lines around the span and every report that cites it, and the clips and post captures `archilyzer reports prepare` made for it, only the cited ones. Each quote is checked against the record as it is composed (a span's against its cues within 5 s either side, read from `en-orig` when the `en` track has no cues; a post's against its text) and the score, time and method are written into the citation, replacing any typed by hand. The build stops with the list of every problem before anything is written: an invalid report, a citation of a channel outside the site or of a post the site may not carry, a missing record, still or post, a quote that matches less than 60 % of what the record says, and a citation without prepared media or with media cut for another span (`--allow-missing-media` on `archilyzer compose site` and `build site` lets those two through, without a clip). A site with search off (`search: false`) removes everything corpus-shaped from `export/public` before it writes its reports, and its built `out/` is checked against what a cited site may hold: anything else, a file over 25 MiB or more than 20,000 files fails the build, and every deploy path (the Publish tab, `deploy site`, Build & deploy, Build & deploy all, the container build) refuses it, as it refuses a site with search off whose last build was a full one. The hub's compose removes a report site's files too. - **A site has a Reports tab.** `/sites/<site>/reports` lists every report under the site's `reports/` directory — the published ones in their order, then the drafts — with its kind, dates, sections, claims, citations by kind and, for a fact-check, how many claims carry each verdict. Each report's problems, from the same checker the prepare step and the build use, open under it. A draft with no problems can be published, and a published report moved up or down or unpublished; each writes only the site's `reports` list, applied to the list as it is on disk at that moment, so it never overwrites another change to the site. "Prepare evidence media" queues the `reports-prepare` job, and beside it the tab shows the last prepared media (moments by kind, total size, problems by kind) and links the last prepare job. What the site publishes (full or cited) is shown with a link to Settings, where it is changed. diff --git a/editor/app/channels/[slug]/pipelineActions.ts b/editor/app/channels/[slug]/pipelineActions.ts @@ -21,6 +21,10 @@ import { runArchiveOrgImport, } from "yt-dlp-transcript-common/controller/archiveOrgImport"; import { + importPlatformSignal, + resolveBitchuteImportUrl, +} from "yt-dlp-transcript-common/controller/bitchuteImport"; +import { heldPlatformRefusal, platformCooldownRemainingMs, recordDownloadBackoff, @@ -519,7 +523,14 @@ export async function importVideoAction( // BitTorrent when it can, else straight from archive.org, never yt-dlp // (controller/archiveOrgDownload.ts) — and a file already on disk is not // fetched again. - const archiveOrg = detectPlatform(videoUrl) === "archiveorg"; + const urlPlatform = detectPlatform(videoUrl); + const archiveOrg = urlPlatform === "archiveorg"; + // A BitChute URL likewise (controller/bitchuteImport.ts): canonicalized, run + // on BitChute's own queue with BitChute's args whatever the channel's + // platform, refused while BitChute is held or cooling down, never fetched + // again when on disk — and what BitChute answers is recorded on its pacing + // state below. + const bitchute = urlPlatform === "bitchute"; let downloadConfig = channelConfig; if (archiveOrg) { const refused = await archiveOrgRefusal(paths, "The import"); @@ -544,6 +555,29 @@ export async function importVideoAction( downloadConfig = { ...channelConfig, platform: "archiveorg" }; } } + if (bitchute) { + const refused = await platformRefusal(paths, "bitchute", "BitChute", "The import"); + if (refused) return { ok: false, info: true, error: refused }; + const resolved = resolveBitchuteImportUrl(videoUrl); + if (!resolved.ok) return { ok: false, error: resolved.error }; + videoUrl = resolved.url; + if ( + await destinationExists( + path.join(paths.channelsDir, slug, "data"), + resolved.id, + channelConfig.handling, + ) + ) { + return { + ok: false, + info: true, + error: `Already downloaded: data/${resolved.id}/ — BitChute is not asked for it again.`, + }; + } + if (channelConfig.platform !== "bitchute") { + downloadConfig = { ...channelConfig, platform: "bitchute" }; + } + } // Best-effort canonical id: used only for revalidation/labels. When null, // downloadOneManaged falls back to %(id)s and the reconcile pass repairs the // dir, so we don't hard-fail here. @@ -551,12 +585,18 @@ export async function importVideoAction( const err = await lowDiskError(paths, slug); if (err) return err; const settings = getSettings(); + // A URL on a platform with its own queue runs there. The channel page's + // queue control starts at the CHANNEL's queue, which is therefore read as + // "no choice made"; any other queue the operator picks still wins. + const channelQueue = downloadQueueKey(channelConfig); + const ownQueue = archiveOrg || bitchute ? platformQueueKey(urlPlatform) : null; + const override = + ownQueue && queueKey !== undefined && queueKey.trim() === channelQueue + ? undefined + : queueKey; return runManagedFunction({ kind: "import-one", - queueKey: resolveQueueKey( - archiveOrg ? platformQueueKey("archiveorg") : downloadQueueKey(channelConfig), - queueKey, - ), + queueKey: resolveQueueKey(ownQueue ?? channelQueue, override), paths, channelSlug: slug, videoId, @@ -567,7 +607,7 @@ export async function importVideoAction( kind: "download", }); try { - await downloadOneManaged({ + const outcome = await downloadOneManaged({ channelSlug: slug, channelConfig: downloadConfig, paths, @@ -579,6 +619,22 @@ export async function importVideoAction( globalSkipLiveDownloads: settings.skipLiveDownloads, appendArchive: true, }); + if (bitchute) { + // BitChute's shared pacing state learns what it said: a 429 backs + // every BitChute path off (the lane, a sync, the next import). + // Best-effort, as the batch downloads' bookkeeping is. + const answer = importPlatformSignal(outcome); + try { + if (answer === "rate_limit" || answer === "network") { + await recordDownloadBackoff("bitchute", paths, answer); + } else if (answer === "clean") { + const line = await recordPlatformClean("bitchute", paths); + if (line) task.onLog(line); + } + } catch { + /* shared-state write is best-effort */ + } + } // ADD-ONLY: record the imported video in the roster so a one-off import // is a known member of the channel rather than an orphan dir, and so a // later retry of the same URL is possible even if it never appears in a @@ -604,18 +660,27 @@ export async function importVideoAction( }); } -// archive.org held, or in a rate-limit cooldown: a sentence, else null. An -// import asks archive.org nothing while it has asked us to wait. -async function archiveOrgRefusal(paths: Paths, what: string): Promise<string | null> { - const held = await heldPlatformRefusal("archiveorg", what, paths); +// A platform held, or in a rate-limit cooldown: a sentence, else null. An +// import asks a platform nothing while it has asked us to wait. +async function platformRefusal( + paths: Paths, + platform: string, + label: string, + what: string, +): Promise<string | null> { + const held = await heldPlatformRefusal(platform, what, paths); if (held) return held; - const remainingMs = await platformCooldownRemainingMs("archiveorg", paths); + const remainingMs = await platformCooldownRemainingMs(platform, paths); if (remainingMs > 0) { - return `archive.org is in a rate-limit cooldown (${Math.ceil(remainingMs / 1000)}s remaining). ${what} can run once it lapses.`; + return `${label} is in a rate-limit cooldown (${Math.ceil(remainingMs / 1000)}s remaining). ${what} can run once it lapses.`; } return null; } +function archiveOrgRefusal(paths: Paths, what: string): Promise<string | null> { + return platformRefusal(paths, "archiveorg", "archive.org", what); +} + // IMPORT CHOSEN FILES OF ONE archive.org ITEM (`pnpm ops import-archive-org`): // one job on archive.org's own queue that imports the files one at a time, // with a jittered pause between them, skipping any already downloaded, and diff --git a/editor/app/channels/components/ChannelForm.tsx b/editor/app/channels/components/ChannelForm.tsx @@ -385,6 +385,7 @@ export function ChannelForm({ <option value="twitch">Twitch</option> <option value="kick">Kick</option> <option value="archiveorg">archive.org</option> + <option value="bitchute">BitChute</option> <option value="twitter">X / Twitter (posts)</option> <option value="bluesky">Bluesky (posts)</option> <option value="xenforo">Forum thread — XenForo (posts)</option> @@ -506,6 +507,7 @@ export function ChannelForm({ <option value="twitch">Twitch</option> <option value="kick">Kick</option> <option value="archiveorg">archive.org</option> + <option value="bitchute">BitChute</option> <option value="twitter">X / Twitter (posts)</option> <option value="bluesky">Bluesky (posts)</option> <option value="xenforo">Forum thread — XenForo (posts)</option> diff --git a/editor/e2e/fixtures/bin/fake-ytdlp.mjs b/editor/e2e/fixtures/bin/fake-ytdlp.mjs @@ -111,6 +111,17 @@ async function writeMetadata(videoDir, id, opts = {}) { ? `https://odysee.com/${id}` : `https://www.youtube.com/watch?v=${id}`, }; + // A bitchute.com URL: what yt-dlp's BitChute extractor writes — its key, the + // video page, and the one mp4 on a media host. + if (opts.bitchute) { + const file = `https://seed901.bitchute.com/FakeHash0001/${id}.mp4`; + meta.extractor = "BitChute"; + meta.extractor_key = "BitChute"; + meta.webpage_url = `https://www.bitchute.com/video/${id}/`; + meta.channel_url = "https://www.bitchute.com/channel/fakechannel/"; + meta.formats = [{ format_id: "0", ext: "mp4", url: file }]; + meta.url = file; + } // For the no-subs-fallback tests we need yt-dlp's metadata to reflect // whether the video advertises caption tracks. The default (no opts) // omits both keys so the rest of the existing suite keeps its behavior. @@ -171,6 +182,7 @@ function urlSentinels(url) { wasLive: lower.includes("waslive") || lower.includes("livevid"), longDuration: lower.includes("longvideo"), odysee: lower.includes("odyseevid"), + bitchute: /^https?:\/\/([a-z0-9-]+\.)?bitchute\.com\//.test(lower), }; } diff --git a/editor/e2e/import-video.spec.ts b/editor/e2e/import-video.spec.ts @@ -1,4 +1,4 @@ -import { readFile } from "node:fs/promises"; +import { readdir, readFile } from "node:fs/promises"; import { test, expect } from "@playwright/test"; import { channelStage, @@ -69,3 +69,55 @@ test("import rejects a non-URL input", async ({ page }) => { { timeout: 10_000 }, ); }); + +// A BitChute video imported into an Odysee (transcribe) channel runs the way BitChute's own +// downloads do: its canonical page, BitChute's queue (one transfer at a time) +// and yt-dlp args — the queue control's default (the channel's queue) is no +// choice — the record labelled from yt-dlp's BitChute extractor, and a second +// import of it refused without asking BitChute again. +test("a BitChute import runs on BitChute's queue and pace, and is never fetched twice", async ({ + page, +}) => { + test.setTimeout(90_000); + await resetData("one-transcribe-channel"); + await generateReport(page, "test-transcribe"); + await page.goto(channelStage("test-transcribe", "playlist")); + + const id = "bcOneoff0001"; + await page + .getByLabel("Video URL to import") + .fill(`https://old.bitchute.com/video/${id}/`); + const importBtn = page.getByRole("button", { name: "Import video" }); + await importBtn.click(); + + const log = page.getByLabel("Import video output"); + await expect(log).toContainText(`https://www.bitchute.com/video/${id}/`, { + timeout: 30_000, + }); + await expect(log).toContainText("--sleep-requests 3"); + + const metaPath = `test-transcripts/channels/test-transcribe/data/${id}/metadata.info.json`; + await expect.poll(() => pathExists(metaPath), { timeout: 30_000 }).toBe(true); + const meta = JSON.parse(await readFile(resolvePath(metaPath), "utf8")); + expect(meta.extractor_key).toBe("BitChute"); + + // The job ran on BitChute's queue, not the channel's. + await expect(async () => { + const dir = resolvePath("test-transcripts/.jobs"); + const queues: string[] = []; + for (const name of await readdir(dir)) { + if (!name.endsWith(".meta.json")) continue; + const m = JSON.parse(await readFile(`${dir}/${name}`, "utf8")); + if (m.kind === "import-one" && m.videoId === id) queues.push(m.queueKey); + } + expect(queues).toEqual(["platform:bitchute"]); + }).toPass({ timeout: 15_000 }); + + // Imported again: refused as already downloaded, nothing fetched. + await expect(importBtn).toBeEnabled({ timeout: 30_000 }); + await importBtn.click(); + await expect(page.getByLabel("Import video notice")).toContainText( + "Already downloaded", + { timeout: 10_000 }, + ); +}); diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md @@ -6,6 +6,7 @@ - **A report's claim can carry a flag, its header names the document under review, and a site with one report names it in the browser tab.** `report.json` claim `flag` (one line, at most 60 characters) shows as a small pill in the accent colour beside the claim's verdict, e.g. "No source given". On a report-only site with one report, the home page's tab title is the report's, as on the report's own page. A report's page header is its name — with a `series`, the series on one line in the accent colour and the title on the line below; without one, the title — then one small line of dates and the revision ("2026-10-04 · updated 2026-10-05 · revision 1"; "updated" only when it differs), then a card for the document under review: its title linking to the document, "<author> · <publisher> · <date>", and its archive links folded away, on a left rail in the document's colour (a source's `accent`, `"#rrggbb"`; without one, the border colour), then the subtitle. The page names no byline or site of its own. A claim that cites the document's own sentence shows that sentence — its still, else its words — on the same rail, with no link up to the card and no paraphrase beside it; a sentence of another document links "from <title>" to that document's box; a titled claim with no such sentence shows its text under the title, plain. A report's citation can say where its evidence came from (`origin`: `"subject"`, the document under review gave it; `"added"`, the report's author found it). A claim lists what the report added first, each card marked with the Archilyzer mark under its number ("Not in the article", or "Not in the source", is the mark's tooltip and what a screen reader says), then evidence of unknown origin, then what the document gave itself folded under "In the article (n)"; the reference list marks an added citation with the mark too, and a claim's flag pill wears the same mark. A fact-check's page reads in three tiers, each opened by a hairline with one, two or three dots: the quick take (the tally, the summary, and links to what the check found, every claim and the downloads); **What the check found**, every ruled claim grouped by verdict (contradicted, not found, partly, untestable, corroborated), one line each linking to the claim, with its `gist` (a new optional claim field, one line, at most 240 characters) and its flag; and **Every claim, with its evidence**, which opens with **How it was checked** (`method`, a new optional report field in markdown). A report of kind `sweep` has the first and last tiers only. Needs a rebuild and deploy of the site. - **Forum posts read like the other posts.** A post from a forum-thread channel shows its place in the thread (#N), an "edited" mark, the thread's title and its media as links, and opening its thread shows its conversation: the posts it quotes and the posts quoting it. - **archive.org records play and are cited with their downloads.** A record imported from archive.org plays its file in the page's own player, which seeks to a cited second and follows the transcript. A citation of one links "archive.org" (the file's page) and its "torrent"; a citation of an archive.org mirror of a YouTube upload links the original on YouTube at the cited second, then "archive.org" and "torrent", on the citation cards and the moment pages. +- **BitChute records play, and are labelled BitChute.** A record from BitChute plays its mp4 in the page's own player, which seeks to a cited second and follows the transcript, and opens the BitChute page when the file will not load. A citation of one links the BitChute page, with no start time. The duplicates page names the platform "BitChute". - **A report-only site with one report opens on that report.** Its home page is the report itself, its header links nothing, and `/reports/` forwards home: there is no index of one. With more reports the home page is the list, with no heading of its own; a list entry is the report's name, subtitle and dates (its counts and tally are on its page). Pages a report-only site does not have link home. - **A report can belong to a series.** `report.json` `series` is shown on its own line above the report's title, in the accent colour, at the head of the report's page and its downloads and in the report list, where a report without one shows its kind ("Fact-check"); a page title, a cited-in link, `llms.txt` and the MCP name it `<series>: <title>`. - **`pnpm start:export` serves a built site's moment pages.** It used `serve`, which listed a video or audio moment's directory (`3126.00-3151.00`) instead of serving its page; it now runs `export/scripts/serve-out.mjs`, which serves directories as Cloudflare Pages does, on `EXPORT_DEV_PORT` (3000). diff --git a/export/app/duplicates/DuplicatesClient.tsx b/export/app/duplicates/DuplicatesClient.tsx @@ -32,6 +32,7 @@ const PLATFORM_LABEL: Record<string, string> = { twitch: "Twitch", kick: "Kick", archiveorg: "archive.org", + bitchute: "BitChute", }; const MATCH_LABEL: Record<DuplicateCluster["matchKind"], string> = { diff --git a/mcp/README.md b/mcp/README.md @@ -53,7 +53,7 @@ compact `[mm:ss|<seconds>]`. The expansion rule: **full moment link = `[title @ 2:36](<moment_base>156)`. The seconds are floored exactly like the inline links', so both styles cite the identical second. A video whose base can't be built (no viewer origin and a platform whose time param doesn't take -raw seconds — Twitch — or doesn't exist — Rumble/Kick/archive.org) omits the line: cite its +raw seconds — Twitch — or doesn't exist — Rumble/Kick/archive.org/BitChute) omits the line: cite its `- source:` URL plain instead. `open_link` results and `get_transcript` stay inline-linked (future work).