Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit c4eea25ab4f57dde7730ab376313bea0bbb6e395
parent 9b649912101756b6a746db374b3b40af5966564d
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Tue,  6 Oct 2026 09:11:52 -0400

Merge sources/wayback (Wayback captures named by what they copy; wayback.json provenance; wayback refresh; archived-copy links)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
MREADME.md | 23+++++++++++++++++++++++
Mcommon/bin/archilyzer.ts | 19+++++++++++++++++++
Acommon/bin/wayback-refresh.ts | 70++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/controller/reconcileVideoDirs.ts | 9++++++++-
Mcommon/controller/rosterStore.ts | 27+++++++++++++++++++++++++++
Acommon/controller/waybackRefresh.test.ts | 242+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/controller/waybackRefresh.ts | 342+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/lib/metadataHistory.ts | 4++++
Mcommon/lib/momentUrl.ts | 4++++
Mcommon/lib/report/views.ts | 6++++--
Mcommon/lib/sidecar-server.test.ts | 4+++-
Mcommon/lib/transcripts-server.ts | 14+++++++++++---
Mcommon/lib/videoId.ts | 20++++++++++++++++++++
Acommon/lib/wayback-server.ts | 43+++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/wayback.test.ts | 162+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/wayback.ts | 211+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/publish/composeReports.ts | 29+++++++++++++++++++++++++++--
Acommon/publish/composeReportsWayback.test.ts | 54++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/ytdlp/downloadOneManaged.ts | 31+++++++++++++++++++++++++++++--
Meditor/CHANGELOG.md | 2++
Meditor/app/channels/[slug]/videos/[id]/page.tsx | 24++++++++++++++++++++++++
Mexport/CHANGELOG.md | 1+
22 files changed, 1330 insertions(+), 11 deletions(-)

diff --git a/README.md b/README.md @@ -212,6 +212,29 @@ nothing asked while BitChute is in a rate-limit cooldown, which a 429 starts; an on disk never fetched again (`common/ytdlp/platformArgs.mjs`, `common/controller/bitchuteImport.ts`). +### Wayback Machine captures + +**Import video** takes a Wayback Machine capture +(`https://web.archive.org/web/<timestamp>/<original>`, with or without a replay +modifier such as `id_`) and yt-dlp downloads it: an archived YouTube page through its +Wayback extractor, a raw media file as the file. The record is named by what the capture +is OF: an archived YouTube page by its YouTube id, a JW Player file +(`…/videos/<id>-<rendition>.mp4` on `cdn.jwplayer.com`, `content.jwplatform.com`, +`videos-fms.jwpsrv.com`) by its media id. Every download of a capture writes +`wayback.json` beside its metadata: the original URL (an archived YouTube page's watch +URL), the capture's timestamp, the capture as a page that plays and as its raw bytes. + +The video page says "Archived copy (Wayback Machine, <capture date>) of <original>". A +citation of one links the original, marked as the original and as possibly gone, and the +Wayback copy; its moment link is the capture, which plays. + +`pnpm archilyzer wayback refresh <channel> [--titles <file>] [--dry-run]` brings records +imported before this up to it, offline: `wayback.json`, the dir renamed to its id through +the snapshot's own reconcile pass (the roster entry moves with it), and with `--titles` (a +JSON file of `id → {title, upload_date}`) the title and date of a raw file that has none. +A record a running job holds is skipped and named. It prints old → new; a second run +changes nothing. + ### Requirements Always needed, to install and run the apps: diff --git a/common/bin/archilyzer.ts b/common/bin/archilyzer.ts @@ -351,6 +351,25 @@ export const COMMANDS: Command[] = [ }, }, { + path: ["wayback", "refresh"], + usage: + "<slug> [--titles <json>] [--dry-run] bring a channel's Wayback Machine copies up to the Wayback rules: wayback.json (original URL, capture time), the dir renamed to its canonical id through the snapshot's reconcile (roster moved with it), and with --titles (a file of id → {title, upload_date}) a raw file's title and date; offline, skips a record a live job holds; prints old → new", + flags: { titles: "string", "dry-run": "boolean" }, + maxPositionals: 1, + run: async ({ positionals, flags }) => { + const [slug] = positionals; + if (!slug) { + console.error("wayback refresh: which channel? Pass its slug."); + return 2; + } + return (await import("./wayback-refresh")).main({ + slug, + dryRun: flags["dry-run"] === true, + ...(typeof flags.titles === "string" ? { titlesFile: flags.titles } : {}), + }); + }, + }, + { path: ["feeds", "backfill-metadata"], usage: "<slug> [--feed <url>] [--dry-run] complete a podcast channel's records (title, date, description, duration) from its RSS feed: one fetch of the feed (default: the channel's url), no media; --dry-run counts matched / unmatched / already complete and writes nothing", diff --git a/common/bin/wayback-refresh.ts b/common/bin/wayback-refresh.ts @@ -0,0 +1,70 @@ +// `archilyzer wayback refresh <slug> [--titles <json>] [--dry-run]` — every +// Wayback Machine copy of a channel brought up to the Wayback rules +// (controller/waybackRefresh.ts): its `wayback.json`, its dir under its +// canonical id (through the snapshot's own reconcile pass, roster moved with +// it), and with `--titles` a file mapping id → {title, upload_date} for a raw +// file that has neither. Offline. Prints old → new per record; a second run +// changes nothing. + +import { readFile } from "node:fs/promises"; +import { isValidChannelSlug } from "../controller/channels"; +import { refreshWaybackRecords, type WaybackTitle } from "../controller/waybackRefresh"; +import { getPaths, type Paths } from "../lib/paths"; + +async function readTitles(file: string): Promise<Record<string, WaybackTitle>> { + const v = JSON.parse(await readFile(file, "utf8")) as unknown; + if (!v || typeof v !== "object" || Array.isArray(v)) { + throw new Error(`${file}: expected an object of id → {title, upload_date}`); + } + const out: Record<string, WaybackTitle> = {}; + for (const [id, raw] of Object.entries(v as Record<string, unknown>)) { + if (!raw || typeof raw !== "object" || Array.isArray(raw)) throw new Error(`${file}: "${id}" is not an object`); + const r = raw as Record<string, unknown>; + out[id] = { + ...(typeof r.title === "string" ? { title: r.title } : {}), + ...(typeof r.upload_date === "string" ? { upload_date: r.upload_date } : {}), + }; + } + return out; +} + +export async function main(opts: { + slug: string; + dryRun: boolean; + titlesFile?: string; + paths?: Paths; +}): Promise<number> { + if (!isValidChannelSlug(opts.slug)) { + console.error(`wayback refresh: "${opts.slug}" is not a channel slug`); + return 2; + } + try { + const titles = opts.titlesFile ? await readTitles(opts.titlesFile) : undefined; + const r = await refreshWaybackRecords({ + slug: opts.slug, + paths: opts.paths ?? getPaths(), + dryRun: opts.dryRun, + titles, + onLog: (line) => console.log(line.replace(/\n$/, "")), + }); + const would = opts.dryRun ? "would be " : ""; + const renamed = r.records.filter((x) => x.rename === "renamed" || x.rename === "merged").length; + const held = r.records.filter((x) => x.rename === "held" || x.rename === "conflict").length; + const sidecars = r.records.filter((x) => x.sidecar).length; + const patched = r.records.filter((x) => Object.keys(x.changes).length > 0).length; + for (const id of r.unmatchedTitles) console.log(`--titles: no Wayback record "${id}"`); + for (const f of r.failed) console.error(`${f.id}: failed — ${f.error}`); + console.log( + `${opts.dryRun ? "dry run: " : ""}${r.records.length} Wayback records, ` + + `${renamed} ${would}renamed, ${sidecars} wayback.json ${would}written, ` + + `${patched} ${would}retitled` + + (held > 0 ? `, ${held} not renamed` : "") + + (r.failed.length > 0 ? `, ${r.failed.length} failed` : "") + + ".", + ); + return r.failed.length > 0 ? 1 : 0; + } catch (err) { + console.error(`wayback refresh: ${(err as Error).message}`); + return 1; + } +} diff --git a/common/controller/reconcileVideoDirs.ts b/common/controller/reconcileVideoDirs.ts @@ -28,6 +28,10 @@ export type ReconcileOpts = { channelDir: string; dryRun?: boolean; onLog?: (s: string) => void; + // Reconcile only the dirs this accepts (default: every dir). A caller that + // knows which records it means to move (`archilyzer wayback refresh`, + // controller/waybackRefresh.ts) leaves the rest to the snapshot's pass. + only?: (dirName: string) => boolean; }; // Files the canonical dir's copy should win on a name collision: these are @@ -136,7 +140,10 @@ export async function reconcileVideoDirs( const entries = await readdir(dataDir, { withFileTypes: true }).catch( () => [], ); - const dirNames = entries.filter((e) => e.isDirectory()).map((e) => e.name); + const dirNames = entries + .filter((e) => e.isDirectory()) + .map((e) => e.name) + .filter((name) => !opts.only || opts.only(name)); for (const name of dirNames) { try { diff --git a/common/controller/rosterStore.ts b/common/controller/rosterStore.ts @@ -141,6 +141,33 @@ export function mergeRoster( return { ...roster, version: ROSTER_VERSION, updatedAt: now, entries }; } +// A RENAMED RECORD keeps its roster entry under its new id. Not a removal: +// the entry moves (url, firstSeenAt, source and all), which is what a dir +// renamed to its canonical id (reconcileVideoDirs.ts) needs — left under the +// old id it would read as a video the channel has and nobody downloaded. +// When both ids have an entry the new one's stays and the earlier +// firstSeenAt wins. Returns the SAME object when nothing moved. +export function renameRosterEntries( + roster: Roster, + renames: ReadonlyArray<{ from: string; to: string }>, + now: string, +): Roster { + const entries: Record<string, RosterEntry> = { ...roster.entries }; + let changed = false; + for (const { from, to } of renames) { + const prev = entries[from]; + if (!prev || from === to) continue; + const there = entries[to]; + entries[to] = there + ? { ...there, firstSeenAt: prev.firstSeenAt && prev.firstSeenAt < there.firstSeenAt ? prev.firstSeenAt : there.firstSeenAt, url: there.url || prev.url } + : prev; + delete entries[from]; + changed = true; + } + if (!changed) return roster; + return { ...roster, version: ROSTER_VERSION, updatedAt: now, entries }; +} + // Stamp the outcome of an enumeration. Kept separate from mergeRoster because a // REJECTED enumeration still merges (additively, losing nothing) while recording // that its listing was not trusted — that record is what the next enumeration diff --git a/common/controller/waybackRefresh.test.ts b/common/controller/waybackRefresh.test.ts @@ -0,0 +1,242 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { lstat, mkdir, mkdtemp, readFile, readdir, readlink, stat, symlink, writeFile } from "node:fs/promises"; +import { tmpdir } from "node:os"; +import path from "node:path"; +import type { Paths } from "../lib/paths"; +import { buildWaybackProvenance } from "../lib/wayback"; +import { loadMetadataHistory } from "../lib/metadataHistory-server"; +import { refreshWaybackRecords } from "./waybackRefresh"; +import { renameRosterEntries, type Roster } from "./rosterStore"; +import { main as refreshCli } from "../bin/wayback-refresh"; + +// Run with: +// pnpm --filter yt-dlp-transcript-common exec tsx --test controller/waybackRefresh.test.ts +// +// A temp corpus of Wayback copies imported before the app knew what one was: +// an archived YouTube page in `watch/`, two raw JW Player files in +// `<jwId>-<rendition>.mp4/`, and a plain record. Every id here is invented. + +const SLUG = "demo-wayback"; +const YT = "Abc123def45"; +const PAGE = `https://web.archive.org/web/20220102030405/https://www.youtube.com/watch?v=${YT}`; +const JW_A = "Qw3rTy12"; +const JW_B = "Zx9vBn34"; +const fileUrl = (jw: string, ts: string) => + `https://web.archive.org/web/${ts}id_/https://videos-fms.jwpsrv.com/content/conversions/AcCt1234/videos/${jw}-12345678.mp4?token=0_abc_0xdef`; +const FILE_A = fileUrl(JW_A, "20200102030405"); +const FILE_B = fileUrl(JW_B, "20200103030405"); + +async function withCorpus( + fn: (paths: Paths, dataDir: string, channelDir: string) => Promise<void>, +): Promise<void> { + const dir = await mkdtemp(path.join(tmpdir(), "ttb-wayback-")); + const transcriptsDir = path.join(dir, "corpus"); + const paths = { + transcriptsDir, + channelsDir: path.join(transcriptsDir, "channels"), + jobsDir: path.join(transcriptsDir, ".jobs"), + } as Paths; + const channelDir = path.join(paths.channelsDir, SLUG); + const dataDir = path.join(channelDir, "data"); + await mkdir(dataDir, { recursive: true }); + await mkdir(paths.jobsDir, { recursive: true }); + await writeFile( + path.join(channelDir, "config.json"), + JSON.stringify({ name: "Demo", platform: "archiveorg", handling: "transcribe" }), + ); + const record = async (name: string, info: Record<string, unknown>) => { + await mkdir(path.join(dataDir, name), { recursive: true }); + await writeFile(path.join(dataDir, name, "metadata.info.json"), JSON.stringify(info)); + }; + await record("watch", { + id: YT, + title: "An archived upload", + upload_date: "20090102", + extractor_key: "YoutubeWebArchive", + webpage_url: PAGE, + webpage_url_basename: "watch", + }); + for (const [jw, url] of [[JW_A, FILE_A], [JW_B, FILE_B]] as const) { + await record(`${jw}-12345678.mp4`, { + id: `${jw}-12345678`, + title: `${jw}-12345678`, + extractor_key: "Generic", + webpage_url: url, + webpage_url_basename: `${jw}-12345678.mp4`, + }); + } + await record("Plain12345a", { id: "Plain12345a", title: "Plain", upload_date: "20200101", extractor_key: "Youtube", webpage_url: "https://www.youtube.com/watch?v=Plain12345a" }); + // A tiered file: a relative link into media/<old id>/ (release 17). + await mkdir(path.join(channelDir, "media", `${JW_A}-12345678.mp4`), { recursive: true }); + await writeFile(path.join(channelDir, "media", `${JW_A}-12345678.mp4`, "audio.mp3"), "audio"); + await symlink( + path.join("..", "..", "media", `${JW_A}-12345678.mp4`, "audio.mp3"), + path.join(dataDir, `${JW_A}-12345678.mp4`, "audio.mp3"), + ); + const entry = (url: string) => ({ url, firstSeenAt: "2026-01-01T00:00:00.000Z", lastListedAt: "2026-01-01T00:00:00.000Z", source: "import" }); + await writeFile( + path.join(channelDir, "roster.json"), + JSON.stringify({ + version: 1, + updatedAt: "2026-01-01T00:00:00.000Z", + lastSweep: null, + entries: { + watch: entry(PAGE), + [`${JW_A}-12345678.mp4`]: entry(FILE_A), + [`${JW_B}-12345678.mp4`]: entry(FILE_B), + Plain12345a: entry("https://www.youtube.com/watch?v=Plain12345a"), + }, + }), + ); + await fn(paths, dataDir, channelDir); +} + +const titles = { + [JW_A]: { title: "Show: First Guest", upload_date: "2019-03-03" }, + [`${JW_B}-12345678.mp4`]: { title: "Show: Second Guest", upload_date: "20190310" }, + [YT]: { title: "Not applied: the page has a title", upload_date: "20000101" }, + nobody: { title: "No such record" }, +}; + +test("a dry run reports every rename and title and writes nothing", async () => { + await withCorpus(async (paths, dataDir, channelDir) => { + const rosterBefore = await readFile(path.join(channelDir, "roster.json"), "utf8"); + const r = await refreshWaybackRecords({ slug: SLUG, paths, dryRun: true, titles }); + assert.deepEqual( + r.records.map((x) => [x.from, x.to, x.rename, x.sidecar]), + [ + [`${JW_A}-12345678.mp4`, JW_A, "renamed", true], + [`${JW_B}-12345678.mp4`, JW_B, "renamed", true], + ["watch", YT, "renamed", true], + ], + ); + assert.deepEqual(r.records[0].changes, { + title: { from: `${JW_A}-12345678`, to: "Show: First Guest" }, + upload_date: { from: null, to: "20190303" }, + }); + assert.deepEqual(r.records[2].changes, {}); + assert.deepEqual(r.unmatchedTitles, ["nobody"]); + assert.deepEqual((await readdir(dataDir)).sort(), [`${JW_A}-12345678.mp4`, `${JW_B}-12345678.mp4`, "Plain12345a", "watch"].sort()); + assert.equal(await readFile(path.join(channelDir, "roster.json"), "utf8"), rosterBefore); + await assert.rejects(stat(path.join(dataDir, "watch", "wayback.json"))); + }); +}); + +test("a real run renames through the reconcile pass, moves the roster, writes sidecars and titles; a second run changes nothing", async () => { + await withCorpus(async (paths, dataDir, channelDir) => { + const r = await refreshWaybackRecords({ slug: SLUG, paths, titles, now: () => new Date("2026-02-02T00:00:00.000Z") }); + assert.equal(r.failed.length, 0); + assert.deepEqual((await readdir(dataDir)).sort(), [JW_A, JW_B, "Plain12345a", YT].sort()); + + // The tier link moved with its dir and still resolves. + const link = path.join(dataDir, JW_A, "audio.mp3"); + assert.ok((await lstat(link)).isSymbolicLink()); + assert.equal(await readlink(link), path.join("..", "..", "media", `${JW_A}-12345678.mp4`, "audio.mp3")); + assert.equal(await readFile(link, "utf8"), "audio"); + + const roster = JSON.parse(await readFile(path.join(channelDir, "roster.json"), "utf8")); + assert.deepEqual(Object.keys(roster.entries).sort(), [JW_A, JW_B, "Plain12345a", YT].sort()); + assert.equal(roster.entries[YT].url, PAGE); + assert.equal(roster.entries[YT].firstSeenAt, "2026-01-01T00:00:00.000Z"); + + assert.deepEqual(JSON.parse(await readFile(path.join(dataDir, YT, "wayback.json"), "utf8")), buildWaybackProvenance(PAGE)); + assert.equal(JSON.parse(await readFile(path.join(dataDir, JW_A, "wayback.json"), "utf8")).originalUrl, FILE_A.replace(/^.*?id_\//, "")); + await assert.rejects(stat(path.join(dataDir, "Plain12345a", "wayback.json"))); + + const info = JSON.parse(await readFile(path.join(dataDir, JW_B, "metadata.info.json"), "utf8")); + assert.equal(info.title, "Show: Second Guest"); + assert.equal(info.upload_date, "20190310"); + const page = JSON.parse(await readFile(path.join(dataDir, YT, "metadata.info.json"), "utf8")); + assert.equal(page.title, "An archived upload"); + assert.equal(page.upload_date, "20090102"); + const history = await loadMetadataHistory(path.join(dataDir, JW_A)); + assert.equal(history?.entries.at(-1)?.by, "wayback-provenance"); + + const again = await refreshWaybackRecords({ slug: SLUG, paths, titles }); + assert.deepEqual( + again.records.map((x) => [x.from, x.to, x.rename ?? null, x.sidecar, Object.keys(x.changes).length]), + [ + [YT, YT, null, false, 0], + [JW_A, JW_A, null, false, 0], + [JW_B, JW_B, null, false, 0], + ], + ); + }); +}); + +test("a record a live job names is held; a dead writer's job is not", async () => { + await withCorpus(async (paths, dataDir) => { + const meta = (id: string, videoId: string, pid: number) => + writeFile( + path.join(paths.jobsDir, `${id}.meta.json`), + JSON.stringify({ id, kind: "transcribe-one", queueKey: "transcription", channelSlug: SLUG, videoId, status: "running", queuedAt: 1, pid }), + ); + await meta("01AAAAAAAAAAAAAAAAAAAAAAAA", `${JW_A}-12345678.mp4`, process.ppid); + await meta("01BBBBBBBBBBBBBBBBBBBBBBBB", "watch", 2 ** 22 + 12345); + // An auto-queue lane's newest pick is in flight; an old one is not. + paths.autoQueueStateFile = path.join(paths.transcriptsDir, ".auto-queue", "state.json"); + await mkdir(path.dirname(paths.autoQueueStateFile), { recursive: true }); + const at = Date.parse("2026-02-02T00:00:00.000Z"); + await writeFile( + paths.autoQueueStateFile, + JSON.stringify({ + transcription: { + picks: [ + { at, leafId: "x", videoId: `${JW_B}-12345678.mp4`, channelSlug: SLUG }, + // Finished: its outcome sidecar is newer than the pick (below). + { at: at - 1000, leafId: "x", videoId: "watch", channelSlug: SLUG }, + ], + }, + // Too old to be in flight. + download: { picks: [{ at: at - 7 * 3600 * 1000, leafId: "x", videoId: "watch", channelSlug: SLUG }] }, + }), + ); + await writeFile(path.join(dataDir, "watch", "transcribe-outcome.json"), "{}"); + const r = await refreshWaybackRecords({ slug: SLUG, paths, titles, now: () => new Date(at + 60_000) }); + assert.equal(r.records.find((x) => x.from === `${JW_B}-12345678.mp4`)!.rename, "held"); + const a = r.records.find((x) => x.from === `${JW_A}-12345678.mp4`)!; + assert.equal(a.rename, "held"); + assert.equal(a.to, a.from); + assert.deepEqual(a.changes, {}); + assert.equal(r.records.find((x) => x.from === "watch")!.rename, "renamed"); + assert.deepEqual((await readdir(dataDir)).sort(), [`${JW_A}-12345678.mp4`, `${JW_B}-12345678.mp4`, "Plain12345a", YT].sort()); + }); +}); + +test("renameRosterEntries moves an entry and keeps the earlier sighting", () => { + const e = (url: string, firstSeenAt: string) => ({ url, firstSeenAt, lastListedAt: firstSeenAt, source: "import" as const }); + const roster: Roster = { + version: 1, + updatedAt: "", + lastSweep: null, + entries: { old: e("u-old", "2026-01-01"), both: e("u-both", "2026-01-01"), keep: e("u-keep", "2026-03-03") }, + }; + const next = renameRosterEntries(roster, [{ from: "old", to: "new" }, { from: "both", to: "keep" }], "now"); + assert.deepEqual(Object.keys(next.entries).sort(), ["keep", "new"]); + assert.equal(next.entries.new.url, "u-old"); + assert.equal(next.entries.keep.firstSeenAt, "2026-01-01"); + assert.equal(next.entries.keep.url, "u-keep"); + assert.equal(renameRosterEntries(next, [{ from: "gone", to: "x" }], "now"), next); +}); + +test("the CLI: --titles from a file, old → new printed, exit 0", async () => { + await withCorpus(async (paths, dataDir) => { + const file = path.join(paths.transcriptsDir, "titles.json"); + await writeFile(file, JSON.stringify(titles)); + const lines: string[] = []; + const orig = console.log; + console.log = (...a: unknown[]) => void lines.push(a.join(" ")); + try { + assert.equal(await refreshCli({ slug: SLUG, dryRun: false, titlesFile: file, paths }), 0); + } finally { + console.log = orig; + } + const out = lines.join("\n"); + assert.match(out, new RegExp(`watch → ${YT}`)); + assert.match(out, new RegExp(`${JW_A}-12345678\\.mp4 → ${JW_A}`)); + assert.match(out, /3 Wayback records, 3 renamed, 3 wayback\.json written, 2 retitled\./); + assert.ok((await readdir(dataDir)).includes(YT)); + assert.equal(await refreshCli({ slug: "Not A Slug", dryRun: true, paths }), 2); + }); +}); diff --git a/common/controller/waybackRefresh.ts b/common/controller/waybackRefresh.ts @@ -0,0 +1,342 @@ +// WAYBACK MACHINE COPIES, BROUGHT UP TO THE WAYBACK RULES — offline. +// +// A record imported from a Wayback capture (lib/wayback.ts) before the app +// knew what one was carries no `wayback.json`, and its dir is named by the last +// segment of the capture URL: `watch` for an archived YouTube page (a name +// every such capture shares), `<jwId>-<rendition>.mp4` for a raw JW Player +// file. This brings every such record of a channel up to the rules: +// +// wayback.json written from the capture URL (lib/wayback-server.ts) +// data/<id>/ renamed to its canonical id (lib/videoId.ts) by the +// snapshot's own pass, reconcileVideoDirs — the media +// tier's relative links move with the dir, and a dir +// already there is merged, never overwritten +// roster.json the entry moved to the new id (renameRosterEntries) +// metadata.info.json with `titles`: the title and upload_date the operator +// found for a record that has none (a raw file's title +// is its file name), through `patchMetadataInfo`, so the +// change is in metadata.history.json as +// `wayback-provenance` +// +// A record a live job names (a `.jobs/` meta, queued or running, whose writer +// is alive, or one of an auto-queue lane's newest picks) is skipped and +// reported; a live job over the whole channel holds every record. Only what differs is written, so a second run writes nothing. +// No network. + +import path from "node:path"; +import { readdir, readFile, stat } from "node:fs/promises"; +import { getPaths, type Paths } from "../lib/paths"; +import { readJsonFile } from "../lib/jsonFile-server"; +import { assertChannelTextReadable, readRelocationMarker } from "../lib/channelMedia"; +import { parseWaybackUrl } from "../lib/wayback"; +import { ensureWaybackProvenance } from "../lib/wayback-server"; +import { extractVideoId } from "../lib/videoId"; +import { patchMetadataInfo } from "../lib/metadataHistory-server"; +import { readChannelConfig } from "./channels"; +import { listChannelVideoIds } from "./keptVideos"; +import { reconcileVideoDirs } from "./reconcileVideoDirs"; +import { loadRoster, renameRosterEntries, writeRoster } from "./rosterStore"; +import { writerIsGone } from "../jobs/bootQueuedJobs"; +import type { JobMeta } from "../jobs/jobMeta"; +import type { AutoQueuePick } from "../jobs/autoQueueState"; + +// What the operator found for a record: its title and the day it is of. +export type WaybackTitle = { title?: string; upload_date?: string }; + +export type WaybackRecordResult = { + // The dir's name before, and after (the same when it was not renamed). + from: string; + to: string; + // The capture the record was fetched from. + captureUrl: string; + rename?: "renamed" | "merged" | "conflict" | "held"; + // Why a rename did not happen (a live job, a conflict, a capture that is + // not the record's webpage_url). + note?: string; + // wayback.json was (or would be) written. + sidecar: boolean; + // Per key: the value before and after. + changes: Partial<Record<"title" | "upload_date", { from: unknown; to: unknown }>>; +}; + +export type WaybackRefreshResult = { + records: WaybackRecordResult[]; + // `titles` entries no Wayback record of the channel matched. + unmatchedTitles: string[]; + failed: { id: string; error: string }[]; +}; + +type Info = Record<string, unknown>; + +async function readInfo(videoDir: string): Promise<Info | null> { + const read = await readJsonFile(path.join(videoDir, "metadata.info.json")); + return read.ok && read.value && typeof read.value === "object" && !Array.isArray(read.value) + ? (read.value as Info) + : null; +} + +// The URL the managed download ran yt-dlp on: the last argument of the +// command line it logged (`$ yt-dlp … -- <url>`). +async function urlFromDownloadLog(videoDir: string): Promise<string | null> { + let head: string; + try { + head = (await readFile(path.join(videoDir, "download.log"), "utf8")).slice(0, 16 * 1024); + } catch { + return null; + } + for (const line of head.split("\n")) { + const m = /^\$ yt-dlp .* -- (\S+)\s*$/.exec(line); + if (m && parseWaybackUrl(m[1])) return m[1]; + } + return null; +} + +// The capture a record was fetched from: its webpage_url, its original_url, +// or the URL its download log ran on. +async function captureUrlOf(videoDir: string, info: Info | null): Promise<string | null> { + for (const key of ["webpage_url", "original_url"]) { + const v = info?.[key]; + if (typeof v === "string" && parseWaybackUrl(v)) return v; + } + return urlFromDownloadLog(videoDir); +} + +// The live jobs on a channel: the record ids they name, and whether one of +// them covers the whole channel. +async function liveJobs( + paths: Paths, + slug: string, + now: number, +): Promise<{ ids: Set<string>; channelWide: string[] }> { + const ids = new Set<string>(); + const channelWide: string[] = []; + const names = paths.jobsDir ? await readdir(paths.jobsDir).catch(() => [] as string[]) : []; + for (const name of names) { + if (!name.endsWith(".meta.json")) continue; + const read = await readJsonFile(path.join(paths.jobsDir, name)); + if (!read.ok || !read.value || typeof read.value !== "object") continue; + const meta = read.value as JobMeta; + if (meta.channelSlug !== slug) continue; + if (meta.status !== "queued" && meta.status !== "running") continue; + if (writerIsGone(meta)) continue; + if (meta.videoId) ids.add(meta.videoId); + else channelWide.push(`${meta.kind} ${meta.id}`); + } + // The auto-queue lanes are jobs of no channel, and the record each is on + // is its newest pick (jobs/autoQueueState.ts). A lane runs a few workers, so + // the newest few recent picks of this channel are held. + if (paths.autoQueueStateFile) { + const read = await readJsonFile(paths.autoQueueStateFile); + const state = read.ok && read.value && typeof read.value === "object" ? (read.value as Record<string, unknown>) : {}; + const since = now - LANE_PICK_HOLD_MS; + for (const [kind, kindState] of Object.entries(state)) { + const picks = (kindState as { picks?: unknown } | null)?.picks; + if (!Array.isArray(picks)) continue; + for (const pick of picks.slice(0, LANE_PICKS_HELD) as Partial<AutoQueuePick>[]) { + if (pick?.channelSlug !== slug || typeof pick.videoId !== "string") continue; + const at = pick.at ?? 0; + if (at < since) continue; + // A pick whose outcome sidecar was written since is finished. + const outcome = LANE_OUTCOME[kind]; + if (outcome) { + const done = await stat(path.join(paths.channelsDir, slug, "data", pick.videoId, outcome)).catch(() => null); + if (done && done.mtimeMs >= at) continue; + } + ids.add(pick.videoId); + } + } + } + return { ids, channelWide }; +} + +// How many of a lane's newest picks may still be in flight, and for how long. +const LANE_PICKS_HELD = 4; +const LANE_PICK_HOLD_MS = 6 * 60 * 60 * 1000; +// The sidecar a lane's run writes when it ends, by lane kind. +const LANE_OUTCOME: Record<string, string> = { + transcription: "transcribe-outcome.json", + download: "download-outcome.json", +}; + +// A title that is not one: absent, or the file's own name (what yt-dlp's +// generic extractor titles a raw file with). +function hasRealTitle(info: Info, names: string[]): boolean { + const t = typeof info.title === "string" ? info.title.trim() : ""; + if (!t) return false; + const own = new Set<string>(names); + for (const key of ["id", "display_id", "webpage_url_basename"]) { + const v = info[key]; + if (typeof v === "string") { + own.add(v); + own.add(v.replace(/\.[A-Za-z0-9]{2,4}$/, "")); + } + } + return !own.has(t); +} + +// `YYYYMMDD` from `YYYYMMDD` or `YYYY-MM-DD`; null otherwise. +export function normalizeUploadDate(v: unknown): string | null { + if (typeof v !== "string") return null; + const s = v.trim(); + if (/^\d{8}$/.test(s)) return s; + const m = /^(\d{4})-(\d{2})-(\d{2})$/.exec(s); + return m ? `${m[1]}${m[2]}${m[3]}` : null; +} + +export async function refreshWaybackRecords(opts: { + slug: string; + paths?: Paths; + dryRun?: boolean; + // Record id (its new id, or its old dir name) → its title and date. + titles?: Record<string, WaybackTitle>; + onLog?: (line: string) => void; + now?: () => Date; +}): Promise<WaybackRefreshResult> { + const paths = opts.paths ?? getPaths(); + const log = opts.onLog ?? (() => {}); + const dryRun = opts.dryRun === true; + const config = await readChannelConfig(paths, opts.slug); + if (!config) throw new Error(`Channel "${opts.slug}" not found`); + // An unreadable text tier is not an empty channel (AGENTS.md). + await assertChannelTextReadable(paths, opts.slug, config); + if (await readRelocationMarker(paths, opts.slug)) { + throw new Error(`Channel "${opts.slug}" is relocating (.relocating.json) — run again when the move is done`); + } + + const channelDir = path.join(paths.channelsDir, opts.slug); + const dataDir = path.join(channelDir, "data"); + const result: WaybackRefreshResult = { records: [], unmatchedTitles: [], failed: [] }; + const jobs = await liveJobs(paths, opts.slug, (opts.now?.() ?? new Date()).getTime()); + + // ─── Find the Wayback records ─── + type Found = { rec: WaybackRecordResult; info: Info | null; canonical: string | null }; + const found: Found[] = []; + for (const name of (await listChannelVideoIds(paths, opts.slug)).sort()) { + const videoDir = path.join(dataDir, name); + try { + const info = await readInfo(videoDir); + const captureUrl = await captureUrlOf(videoDir, info); + if (!captureUrl) continue; + // The snapshot names a dir by its webpage_url (reconcileVideoDirs.ts), + // so that is the only name a rename can give it that lasts. + const webpageUrl = typeof info?.webpage_url === "string" ? info.webpage_url : null; + const canonical = webpageUrl && parseWaybackUrl(webpageUrl) ? extractVideoId(webpageUrl) : null; + const rec: WaybackRecordResult = { from: name, to: name, captureUrl, sidecar: false, changes: {} }; + if (canonical && canonical !== name) rec.to = canonical; + else if (!canonical && webpageUrl !== captureUrl) { + rec.note = "webpage_url is not the capture; the dir keeps its name"; + } + found.push({ rec, info, canonical }); + } catch (err) { + result.failed.push({ id: name, error: (err as Error).message }); + } + } + + const held = (rec: WaybackRecordResult): string | null => { + if (jobs.channelWide.length > 0) return `a live job holds the channel (${jobs.channelWide.join(", ")})`; + if (jobs.ids.has(rec.from) || jobs.ids.has(rec.to)) return "a live job names this record"; + return null; + }; + + // ─── Rename, through the snapshot's own pass ─── + const toRename = new Set<string>(); + for (const { rec } of found) { + if (rec.to === rec.from) continue; + const why = held(rec); + if (why) { + rec.rename = "held"; + rec.note = why; + rec.to = rec.from; + continue; + } + toRename.add(rec.from); + } + if (toRename.size > 0) { + const r = await reconcileVideoDirs({ channelDir, dryRun, only: (name) => toRename.has(name) }); + const byFrom = new Map(found.map((f) => [f.rec.from, f.rec])); + for (const x of r.renamed) { + const rec = byFrom.get(x.from); + if (rec) rec.rename = "renamed"; + } + for (const x of r.merged) { + const rec = byFrom.get(x.from); + if (rec) { + rec.rename = "merged"; + rec.note = `merged into the existing ${x.to}/ (${x.movedFiles.length} files)`; + } + } + for (const x of r.conflicts) { + const rec = byFrom.get(x.from); + if (rec) { + rec.rename = "conflict"; + rec.note = x.reason; + rec.to = rec.from; + } + } + const moved = [...r.renamed, ...r.merged].map((x) => ({ from: x.from, to: x.to })); + if (!dryRun && moved.length > 0) { + const now = (opts.now?.() ?? new Date()).toISOString(); + const before = await loadRoster(paths, opts.slug); + const after = renameRosterEntries(before, moved, now); + if (after !== before) await writeRoster(paths, opts.slug, after); + } + } + + // ─── The sidecar, and the titles ─── + const titles = opts.titles ?? {}; + const usedTitles = new Set<string>(); + for (const { rec, info } of found) { + // A dry run renamed nothing: the record is still under its old name. + const videoDir = path.join(dataDir, dryRun ? rec.from : rec.to); + try { + const s = await ensureWaybackProvenance(videoDir, rec.captureUrl, { dryRun }); + rec.sidecar = s.written; + const key = rec.to in titles ? rec.to : rec.from in titles ? rec.from : null; + if (key === null || !info) continue; + usedTitles.add(key); + if (held(rec)) { + rec.note = rec.note ?? held(rec)!; + continue; + } + const want = titles[key]; + const patch: Record<string, unknown> = {}; + const title = typeof want.title === "string" ? want.title.trim() : ""; + if (title && !hasRealTitle(info, [rec.from, rec.to]) && info.title !== title) { + patch.title = title; + rec.changes.title = { from: info.title ?? null, to: title }; + } + const date = normalizeUploadDate(want.upload_date); + if (want.upload_date !== undefined && !date) { + rec.note = `upload_date "${String(want.upload_date)}" is not YYYYMMDD or YYYY-MM-DD`; + } else if (date && normalizeUploadDate(info.upload_date) === null) { + patch.upload_date = date; + rec.changes.upload_date = { from: info.upload_date ?? null, to: date }; + } + if (Object.keys(patch).length > 0 && !dryRun) { + await patchMetadataInfo(videoDir, patch, { by: "wayback-provenance", requestedBy: "cli", onLog: log }); + } + } catch (err) { + result.failed.push({ id: rec.from, error: (err as Error).message }); + } + } + result.unmatchedTitles = Object.keys(titles).filter((k) => !usedTitles.has(k)).sort(); + result.records = found.map((f) => f.rec); + for (const rec of result.records) log(formatWaybackRecord(rec, dryRun)); + return result; +} + +// One record as the CLI prints it: `old → new`, then what was written. +export function formatWaybackRecord(rec: WaybackRecordResult, dryRun: boolean): string { + const would = dryRun ? "would be " : ""; + const head = + rec.from === rec.to + ? `${rec.from} (name kept)` + : `${rec.from} → ${rec.to}${rec.rename === "merged" ? " (merged)" : ""}`; + const lines = [head]; + if (rec.note) lines.push(` ${rec.rename === "held" || rec.rename === "conflict" ? "NOT RENAMED: " : ""}${rec.note}`); + if (rec.sidecar) lines.push(` wayback.json ${would}written`); + for (const [k, c] of Object.entries(rec.changes)) { + lines.push(` ${k}: ${JSON.stringify(c!.from)} → ${JSON.stringify(c!.to)}`); + } + return lines.join("\n"); +} diff --git a/common/lib/metadataHistory.ts b/common/lib/metadataHistory.ts @@ -68,6 +68,10 @@ export const METADATA_HISTORY_WRITERS = [ // metadata API (controller/archiveOrgDownload.ts) — archive.org records are // not fetched by yt-dlp. "archiveorg-import", + // A Wayback Machine copy's title and date, set from what the operator found + // for it (`archilyzer wayback refresh --titles`, controller/waybackRefresh.ts) + // — a raw media file captured by the Wayback Machine carries no title. + "wayback-provenance", ] as const; export type MetadataHistoryWriter = (typeof METADATA_HISTORY_WRITERS)[number]; diff --git a/common/lib/momentUrl.ts b/common/lib/momentUrl.ts @@ -19,6 +19,7 @@ // report-citation UI, and build tools alike. import { detectPlatform, type Platform } from "./platform"; +import { parseWaybackUrl } from "./wayback"; export type MomentUrlInput = { // Public origin of the archilyzer viewer that owns this video (a RemoteSource @@ -55,6 +56,9 @@ export function platformMomentUrl( if (!webpageUrl) return null; const secs = Math.max(0, Math.floor(seconds || 0)); if (secs <= 0) return webpageUrl; + // A Wayback capture (lib/wayback.ts) is the page that plays, and a time + // param would name a different URL — one the Wayback Machine never captured. + if (parseWaybackUrl(webpageUrl)) return webpageUrl; const plat = platform ?? detectPlatform(webpageUrl); let u: URL; try { diff --git a/common/lib/report/views.ts b/common/lib/report/views.ts @@ -140,11 +140,13 @@ export type RecordView = { originalUrl?: string; // What `originalUrl` is, when "Original" would not say: "archive.org" for an // archive.org record, "YouTube" for an archive.org mirror of a YouTube - // upload (whose originalUrl is the upload at the cited second). + // upload (whose originalUrl is the upload at the cited second), "Original + // (may be gone)" for a Wayback Machine copy (lib/wayback.ts). originalLabel?: string; // Where a reader can fetch the recording itself to check it, derived from // the record's provenance (lib/archiveOrg.ts archiveOrgCitationLinks): the - // archive.org page and the item's torrent. Absent for a record with none. + // archive.org page and the item's torrent; a Wayback copy's capture page + // (lib/wayback.ts waybackCitationLinks). Absent for a record with none. downloads?: { label: string; url: string }[]; // The record in this site's corpus (`/?v=<channel>/<id>&t=<s>`): a FULL site // only — a cited site has no corpus to open. diff --git a/common/lib/sidecar-server.test.ts b/common/lib/sidecar-server.test.ts @@ -11,6 +11,7 @@ import "./attribution-server"; import "./diarization-server"; import "./digest-server"; import "./metadataHistory-server"; +import "./wayback-server"; import { availabilitySidecar, loadAvailability, @@ -42,7 +43,7 @@ async function scratch(): Promise<string> { return mkdtemp(path.join(os.tmpdir(), "sidecar-")); } -test("every declared sidecar filename escapes SUB_FILE_RE, and all eleven are declared", () => { +test("every declared sidecar filename escapes SUB_FILE_RE, and all twelve are declared", () => { assert.deepEqual([...SIDECAR_FILENAMES].sort(), [ "ai-digest.json", "ai-digest.overrides.json", @@ -55,6 +56,7 @@ test("every declared sidecar filename escapes SUB_FILE_RE, and all eleven are de "exclude-truncated-check.json", "metadata.history.json", "transcribe-outcome.json", + "wayback.json", ]); for (const name of SIDECAR_FILENAMES) { assert.ok(!SUB_FILE_RE.test(name), name); diff --git a/common/lib/transcripts-server.ts b/common/lib/transcripts-server.ts @@ -5,6 +5,8 @@ import { defaultWebpageUrl, detectPlatform } from "./platform"; import { archiveOrgPlayableUrl } from "./archiveOrg"; import { bitchutePlayableUrl } from "./bitchute"; import { archiveOrgVideoIdFromNativeId } from "./archiveOrgId"; +import { parseWaybackUrl } from "./wayback"; +import { extractVideoId } from "./videoId"; import type { DisplaySummary, Platform, TranscriptSummary } from "./transcripts"; import type { MediaType, VideoStat, VideoStatus } from "./stats"; import type { VideoState } from "./availability"; @@ -77,7 +79,8 @@ export function loadRawMetadataFromDir( // The platform a record is from, by yt-dlp's extractor. archive.org's // extractor is `ArchiveOrg` (key) / `archive.org` (name) — NOT web.archive.org's -// `YoutubeWebArchive`, which is a YouTube video. +// `YoutubeWebArchive`, which is a YouTube video: the original's platform (the +// record's `wayback.json` says it is a copy, lib/wayback-server.ts). // // AN UNKNOWN EXTRACTOR: the record's own page decides when its host is one the // app knows (lib/detectPlatform.mjs); otherwise "youtube", as it always has @@ -144,12 +147,17 @@ export function summarize( const platform = platformFromMetadata(meta); // archive.org: yt-dlp's id for one file of an item is `<identifier>/<path>`, // which is not a slug; the canonical id (lib/archiveOrgId.ts) is. + // A Wayback capture (lib/wayback.ts): the id its dir is named by — what + // the capture is of (lib/videoId.ts) — not yt-dlp's, which for a raw media + // file is the file's name (`<jwId>-<rendition>`). + const waybackId = parseWaybackUrl(meta.webpage_url) ? extractVideoId(meta.webpage_url!) : null; const id = - platform === "odysee" + waybackId ?? + (platform === "odysee" ? (meta.webpage_url_basename ?? meta.id ?? videoDir) : platform === "archiveorg" ? (archiveOrgVideoIdFromNativeId(meta.id) ?? videoDir) - : (meta.id ?? videoDir); + : (meta.id ?? videoDir)); const dateFromDir = videoDir.match(/^(\d{8})(?:_|$)/)?.[1]; return { slug: `${channelSlug}/${id}`, diff --git a/common/lib/videoId.ts b/common/lib/videoId.ts @@ -10,11 +10,31 @@ // for the native-id resolution that reads metadata.info.json. import { isArchiveOrgItemHost, parseArchiveOrgUrl, archiveOrgVideoId } from "./archiveOrgId"; +import { isJwPlayerHost, isWaybackHost, jwPlayerMediaId, parseWaybackUrl } from "./wayback"; export function extractVideoId(url: string): string | null { + return extractVideoIdAt(url, 0); +} + +function extractVideoIdAt(url: string, depth: number): string | null { try { const u = new URL(url); const host = u.hostname.toLowerCase(); + if (isWaybackHost(host)) { + // A Wayback capture is named by what it is a capture OF + // (lib/wayback.ts): an archived YouTube page by its YouTube id, a JW + // Player file by its media id. Its own path's last segment is the + // original's (`watch`, `<id>-<rendition>.mp4`) — a name two captures + // share. A capture of a capture is not unwrapped twice. + const ref = depth === 0 ? parseWaybackUrl(url) : null; + return ref ? extractVideoIdAt(ref.originalUrl, depth + 1) : null; + } + if (isJwPlayerHost(host)) { + // One media id across every rendition and host; a JW URL that names + // none falls through to the last segment below. + const jw = jwPlayerMediaId(u); + if (jw) return jw; + } if (isArchiveOrgItemHost(host)) { // A whole item → its identifier; one file inside an item → a stable // `<identifier>__<slug>-<hash>` (lib/archiveOrgId.ts). A URL that names diff --git a/common/lib/wayback-server.ts b/common/lib/wayback-server.ts @@ -0,0 +1,43 @@ +// WAYBACK PROVENANCE ON DISK — the `wayback.json` sidecar. +// +// A record downloaded from a Wayback Machine capture (lib/wayback.ts) carries +// what it is a copy of: the original URL, the capture's timestamp, the capture +// as a page that plays and as its raw bytes. Built from the capture URL alone +// — no request — so it is written on every download of one +// (ytdlp/downloadOneManaged.ts) and by `archilyzer wayback refresh` for a +// record imported before this existed (controller/waybackRefresh.ts). + +import { + WAYBACK_PROVENANCE_FILENAME, + buildWaybackProvenance, + coerceWaybackProvenance, + sameWaybackProvenance, + type WaybackProvenance, +} from "./wayback"; +import { sidecar, sidecarField } from "./sidecar-server"; + +export const waybackProvenanceSidecar = sidecar( + WAYBACK_PROVENANCE_FILENAME, + sidecarField(coerceWaybackProvenance), +); + +export const { load: loadWaybackProvenance, write: writeWaybackProvenance } = + waybackProvenanceSidecar; + +// The sidecar for a record fetched by `url`: written when the URL is a +// capture and the one on disk is absent or says otherwise. Returns what the +// record now carries (null for a URL that is not a capture) and whether it was +// (or, with `dryRun`, would be) written. +export async function ensureWaybackProvenance( + videoDir: string, + url: string, + opts: { dryRun?: boolean; onLog?: (line: string) => void } = {}, +): Promise<{ provenance: WaybackProvenance | null; written: boolean }> { + const next = buildWaybackProvenance(url); + if (!next) return { provenance: null, written: false }; + const prev = await loadWaybackProvenance(videoDir); + if (prev && sameWaybackProvenance(prev, next)) return { provenance: prev, written: false }; + if (!opts.dryRun) await writeWaybackProvenance(videoDir, next); + opts.onLog?.(`Wayback provenance: a capture of ${next.originalUrl} (${next.captureTs}).\n`); + return { provenance: next, written: true }; +} diff --git a/common/lib/wayback.test.ts b/common/lib/wayback.test.ts @@ -0,0 +1,162 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { mkdtemp, readFile, writeFile } from "node:fs/promises"; +import { tmpdir } from "node:os"; +import path from "node:path"; +import { + buildWaybackProvenance, + coerceWaybackProvenance, + jwPlayerMediaId, + parseWaybackUrl, + waybackCaptureDate, + waybackCitationLinks, +} from "./wayback"; +import { ensureWaybackProvenance, loadWaybackProvenance } from "./wayback-server"; +import { extractVideoId } from "./videoId"; +import { platformMomentUrl } from "./momentUrl"; +import { platformFromMetadata, summarize } from "./transcripts-server"; + +// Run with: +// pnpm --filter yt-dlp-transcript-common exec tsx --test lib/wayback.test.ts +// +// Every id, account and token here is invented. + +const YT = "Abc123def45"; +const JW = "Qw3rTy12"; +const ARCHIVED_PAGE = `https://web.archive.org/web/20210102030405/https://www.youtube.com/watch?v=${YT}`; +const JW_FILE = `https://videos-fms.jwpsrv.com/content/conversions/AcCt1234/videos/${JW}-12345678.mp4?token=0_abc_0xdef`; +const ARCHIVED_FILE = `https://web.archive.org/web/20200102030405id_/${JW_FILE}`; + +test("parseWaybackUrl: timestamp, modifier and the original with its own query", () => { + assert.deepEqual(parseWaybackUrl(ARCHIVED_PAGE), { + captureTs: "20210102030405", + modifier: "", + originalUrl: `https://www.youtube.com/watch?v=${YT}`, + }); + assert.deepEqual(parseWaybackUrl(ARCHIVED_FILE), { + captureTs: "20200102030405", + modifier: "id_", + originalUrl: JW_FILE, + }); + // A short timestamp, another modifier, no scheme, a collapsed scheme, an + // encoded original, the older path without /web/. + assert.equal(parseWaybackUrl("https://web.archive.org/web/2019im_/example.com/a.png")?.originalUrl, "http://example.com/a.png"); + assert.equal(parseWaybackUrl("https://web.archive.org/web/2019/https:/example.com/x")?.originalUrl, "https://example.com/x"); + assert.equal( + parseWaybackUrl("https://web.archive.org/web/2019/https%3A%2F%2Fexample.com%2Fx")?.originalUrl, + "https://example.com/x", + ); + assert.equal(parseWaybackUrl("http://wayback.archive.org/20190101000000/http://example.com/")?.captureTs, "20190101000000"); + // Not captures. + assert.equal(parseWaybackUrl("https://web.archive.org/web/*/example.com"), null); + assert.equal(parseWaybackUrl("https://archive.org/details/some-item"), null); + assert.equal(parseWaybackUrl(`https://www.youtube.com/watch?v=${YT}`), null); + assert.equal(parseWaybackUrl("not a url"), null); +}); + +test("jwPlayerMediaId: the media id across hosts and renditions", () => { + assert.equal(jwPlayerMediaId(JW_FILE), JW); + assert.equal(jwPlayerMediaId(`https://cdn.jwplayer.com/videos/${JW}-AbCdEf12.mp4`), JW); + assert.equal(jwPlayerMediaId(`https://content.jwplatform.com/videos/${JW}.mp4`), JW); + assert.equal(jwPlayerMediaId(`https://cdn.jwplayer.com/manifests/${JW}.m3u8`), JW); + assert.equal(jwPlayerMediaId(`https://cdn.jwplayer.com/v2/media/${JW}`), JW); + assert.equal(jwPlayerMediaId("https://cdn.jwplayer.com/libraries/AbCd1234.js"), null); + assert.equal(jwPlayerMediaId(`https://example.com/videos/${JW}-1.mp4`), null); +}); + +test("extractVideoId unwraps a capture into what it is a capture of", () => { + assert.equal(extractVideoId(ARCHIVED_PAGE), YT); + assert.equal(extractVideoId(`https://web.archive.org/web/2021if_/https://youtu.be/${YT}`), YT); + assert.equal(extractVideoId(ARCHIVED_FILE), JW); + assert.equal(extractVideoId(JW_FILE), JW); + // Unwrapped once: a capture of a capture names no id. + assert.equal(extractVideoId(`https://web.archive.org/web/2021/${ARCHIVED_PAGE}`), null); + // A Wayback page that is not a capture names no record. + assert.equal(extractVideoId("https://web.archive.org/web/*/example.com"), null); + // A capture of a page the app does not know: the original's last segment. + assert.equal(extractVideoId("https://web.archive.org/web/2021/https://example.com/media/clip-7.mp4"), "clip-7.mp4"); + // Nothing else moved. + assert.equal(extractVideoId(`https://www.youtube.com/watch?v=${YT}`), YT); + assert.equal(extractVideoId("https://example.com/feed/episode-1.mp3"), "episode-1.mp3"); +}); + +test("buildWaybackProvenance: an archived YouTube page's original is its watch URL", () => { + assert.deepEqual(buildWaybackProvenance(`https://web.archive.org/web/20210102id_/https://m.youtube.com/watch?v=${YT}&feature=share`), { + originalUrl: `https://www.youtube.com/watch?v=${YT}`, + captureTs: "20210102", + waybackUrl: `https://web.archive.org/web/20210102/https://m.youtube.com/watch?v=${YT}&feature=share`, + rawUrl: `https://web.archive.org/web/20210102id_/https://m.youtube.com/watch?v=${YT}&feature=share`, + }); + const file = buildWaybackProvenance(ARCHIVED_FILE)!; + assert.equal(file.originalUrl, JW_FILE); + assert.equal(file.waybackUrl, `https://web.archive.org/web/20200102030405/${JW_FILE}`); + assert.equal(file.rawUrl, ARCHIVED_FILE); + assert.equal(buildWaybackProvenance(JW_FILE), null); + assert.deepEqual(coerceWaybackProvenance(JSON.parse(JSON.stringify(file))), file); + assert.equal(coerceWaybackProvenance({ ...file, captureTs: "yesterday" }), null); + assert.equal(coerceWaybackProvenance([]), null); +}); + +test("waybackCaptureDate and the citation links", () => { + assert.equal(waybackCaptureDate("20210102030405"), "2021-01-02"); + assert.equal(waybackCaptureDate("2021"), "2021"); + assert.equal(waybackCaptureDate("202101"), "2021-01"); + const prov = buildWaybackProvenance(ARCHIVED_PAGE)!; + const links = waybackCitationLinks(prov, { + originalMomentUrl: platformMomentUrl(prov.originalUrl, null, 90), + }); + assert.deepEqual(links.original, { + label: "Original (may be gone)", + url: `https://www.youtube.com/watch?v=${YT}&t=90s`, + }); + assert.deepEqual(links.copy, { + label: "Wayback Machine copy, 2021-01-02", + url: `https://web.archive.org/web/20210102030405/https://www.youtube.com/watch?v=${YT}`, + }); +}); + +test("platformMomentUrl: a capture is the page that plays, never given a time param", () => { + assert.equal(platformMomentUrl(ARCHIVED_PAGE, "youtube", 125), ARCHIVED_PAGE); + assert.equal(platformMomentUrl(`https://www.youtube.com/watch?v=${YT}`, "youtube", 125), `https://www.youtube.com/watch?v=${YT}&t=125s`); +}); + +test("platformFromMetadata: YoutubeWebArchive is the original's platform, YouTube", () => { + assert.equal(platformFromMetadata({ extractor_key: "YoutubeWebArchive", extractor: "web.archive:youtube", webpage_url: ARCHIVED_PAGE }), "youtube"); + assert.equal(platformFromMetadata({ extractor: "web.archive:youtube", webpage_url: ARCHIVED_PAGE }), "youtube"); +}); + +test("summarize: a capture's id is its dir's name, not yt-dlp's file-name id", () => { + const s = summarize("demo", `${JW}`, { + id: `${JW}-12345678`, + title: `${JW}-12345678`, + extractor_key: "Generic", + webpage_url: ARCHIVED_FILE, + }); + assert.equal(s.id, JW); + assert.equal(s.slug, `demo/${JW}`); + const page = summarize("demo", YT, { id: YT, title: "A title", extractor_key: "YoutubeWebArchive", webpage_url: ARCHIVED_PAGE }); + assert.equal(page.id, YT); + assert.equal(page.platform, "youtube"); + assert.equal(page.webpageUrl, ARCHIVED_PAGE); +}); + +test("ensureWaybackProvenance writes the sidecar once, and nothing for a non-capture", async () => { + const dir = await mkdtemp(path.join(tmpdir(), "wayback-sidecar-")); + assert.deepEqual(await ensureWaybackProvenance(dir, JW_FILE), { provenance: null, written: false }); + assert.equal(await loadWaybackProvenance(dir), null); + + const dry = await ensureWaybackProvenance(dir, ARCHIVED_PAGE, { dryRun: true }); + assert.equal(dry.written, true); + assert.equal(await loadWaybackProvenance(dir), null); + + const first = await ensureWaybackProvenance(dir, ARCHIVED_PAGE); + assert.equal(first.written, true); + const onDisk = JSON.parse(await readFile(path.join(dir, "wayback.json"), "utf8")); + assert.deepEqual(onDisk, buildWaybackProvenance(ARCHIVED_PAGE)); + assert.equal((await ensureWaybackProvenance(dir, ARCHIVED_PAGE)).written, false); + + // A sidecar for another capture is rewritten. + await writeFile(path.join(dir, "wayback.json"), JSON.stringify(buildWaybackProvenance(ARCHIVED_FILE))); + assert.equal((await ensureWaybackProvenance(dir, ARCHIVED_PAGE)).written, true); + assert.deepEqual(await loadWaybackProvenance(dir), buildWaybackProvenance(ARCHIVED_PAGE)); +}); diff --git a/common/lib/wayback.ts b/common/lib/wayback.ts @@ -0,0 +1,211 @@ +// THE WAYBACK MACHINE — a capture of some other URL, and what it was a copy of. +// +// A Wayback capture URL wraps the original: `https://web.archive.org/web/ +// <timestamp>[<modifier>]/<original>`, where the timestamp is 4–14 digits +// (yyyy[MM[dd[hh[mm[ss]]]]]) and the modifier is the replay mode — none (the +// page in the Wayback frame, which is what plays), `id_` (the raw bytes as +// captured), `im_`, `if_`, `js_`, `cs_`, `oe_`, … The original keeps its own +// query string, so it is the rest of the path PLUS the capture URL's search. +// +// yt-dlp downloads either kind through the editor's import: an archived +// YouTube page through its `YoutubeWebArchive` extractor (the record's id is +// the YouTube id), a raw media file through `generic` (an id that is the file +// name). Neither knows it is a copy; the `wayback.json` sidecar +// (lib/wayback-server.ts) records that, and lib/videoId.ts names the record by +// what the capture is OF. +// +// A LEAF: no imports, so lib/videoId.ts (itself a leaf apart from +// archiveOrgId.ts) can unwrap a capture without a cycle. + +export const WAYBACK_PROVENANCE_FILENAME = "wayback.json"; + +// Hosts that serve Wayback captures. archive.org's own hosts serve items +// (lib/archiveOrgId.ts), never captures. +const WAYBACK_HOSTS = new Set(["web.archive.org", "wayback.archive.org"]); + +export function isWaybackHost(host: string): boolean { + return WAYBACK_HOSTS.has(host.toLowerCase()); +} + +export type WaybackRef = { + // The capture's timestamp as the URL gives it (4–14 digits). + captureTs: string; + // The replay modifier (`id_`, `im_`, …), "" for the framed page. + modifier: string; + // The URL the capture is of, as the capture URL names it (scheme added + // when the capture URL left it off). + originalUrl: string; +}; + +// `/web/<ts><mod>/<original>`, or the older `/<ts><mod>/<original>`. +const CAPTURE_PATH_RE = /^\/(?:web\/)?(\d{4,14})([a-z]{2}_)?\/(.+)$/s; + +// A Wayback capture URL's parts, or null for anything else (a calendar page +// `/web/*/<url>`, a search, another host). +export function parseWaybackUrl(url: string | null | undefined): WaybackRef | null { + if (!url) return null; + let u: URL; + try { + u = new URL(url); + } catch { + return null; + } + if (!isWaybackHost(u.hostname)) return null; + const m = CAPTURE_PATH_RE.exec(u.pathname); + if (!m) return null; + let rest = m[3]; + // A percent-encoded original (`https%3A%2F%2F…`) is the same URL. + if (/^https?%3a/i.test(rest)) { + try { + rest = decodeURIComponent(rest); + } catch { + return null; + } + } + // The WHATWG parser keeps `https://` inside a path as written, but a capture + // URL that went through a path normaliser arrives as `https:/host`. + rest = rest.replace(/^(https?):\/(?!\/)/i, "$1://"); + if (!/^https?:\/\//i.test(rest)) rest = `http://${rest}`; + const original = `${rest}${u.search}${u.hash}`; + try { + new URL(original); + } catch { + return null; + } + return { captureTs: m[1], modifier: m[2] ?? "", originalUrl: original }; +} + +// The capture as a page that plays (the Wayback frame, no modifier). +export function waybackPageUrl(ref: Pick<WaybackRef, "captureTs" | "originalUrl">): string { + return `https://web.archive.org/web/${ref.captureTs}/${ref.originalUrl}`; +} + +// The capture's raw bytes (`id_`), the form yt-dlp fetches a media file by. +export function waybackRawUrl(ref: Pick<WaybackRef, "captureTs" | "originalUrl">): string { + return `https://web.archive.org/web/${ref.captureTs}id_/${ref.originalUrl}`; +} + +// `YYYY-MM-DD` of a capture timestamp, or the year (/month) when that is all +// it names. +export function waybackCaptureDate(captureTs: string): string { + const y = captureTs.slice(0, 4); + const mo = captureTs.slice(4, 6); + const d = captureTs.slice(6, 8); + return [y, mo, d].filter((p) => p.length === 2 || p.length === 4).join("-"); +} + +// ─── JW Player ─── + +// The hosts JW Player serves a media file from. A file there is +// `…/videos/<mediaId>-<rendition>.<ext>` (cdn.jwplayer.com/videos/…, +// content.jwplatform.com/videos/…, videos-fms.jwpsrv.com/content/conversions/ +// <account>/videos/…); the media id is eight alphanumerics and names the video +// across every rendition. +const JW_HOSTS = ["jwplayer.com", "jwplatform.com", "jwpsrv.com"]; + +export function isJwPlayerHost(host: string): boolean { + const h = host.toLowerCase(); + return JW_HOSTS.some((d) => h === d || h.endsWith(`.${d}`)); +} + +const JW_FILE_RE = /\/videos\/([A-Za-z0-9]{8})(?:-[A-Za-z0-9]+)?\.[A-Za-z0-9]+$/; +const JW_MEDIA_RE = /\/(?:manifests|v2\/media|previews)\/([A-Za-z0-9]{8})(?:[-./]|$)/; + +// The JW media id a JW Player file or manifest URL names, or null. +export function jwPlayerMediaId(url: string | URL): string | null { + let u: URL; + try { + u = typeof url === "string" ? new URL(url) : url; + } catch { + return null; + } + if (!isJwPlayerHost(u.hostname)) return null; + const m = JW_FILE_RE.exec(u.pathname) ?? JW_MEDIA_RE.exec(u.pathname); + return m ? m[1] : null; +} + +// ─── The sidecar's record ─── + +export type WaybackProvenance = { + // What the capture is a copy of. For an archived YouTube page, its watch + // URL (`https://www.youtube.com/watch?v=<id>`), whatever form was captured. + originalUrl: string; + // The capture's timestamp (4–14 digits). + captureTs: string; + // The capture as a page that plays. + waybackUrl: string; + // The capture's raw bytes. + rawUrl: string; +}; + +// A YouTube URL as its watch page, or null for any other URL. +function youtubeWatchUrl(url: string): string | null { + let u: URL; + try { + u = new URL(url); + } catch { + return null; + } + const host = u.hostname.toLowerCase(); + let id: string | null = null; + if (host === "youtu.be") id = u.pathname.split("/").filter(Boolean)[0] ?? null; + else if (host === "youtube.com" || host.endsWith(".youtube.com")) { + id = u.searchParams.get("v"); + if (!id) { + const segs = u.pathname.split("/").filter(Boolean); + if ((segs[0] === "embed" || segs[0] === "shorts" || segs[0] === "v" || segs[0] === "live") && segs[1]) id = segs[1]; + } + } else return null; + return id ? `https://www.youtube.com/watch?v=${id}` : null; +} + +// The sidecar for a capture URL, or null when the URL is not one. +export function buildWaybackProvenance(url: string): WaybackProvenance | null { + const ref = parseWaybackUrl(url); + if (!ref) return null; + return { + originalUrl: youtubeWatchUrl(ref.originalUrl) ?? ref.originalUrl, + captureTs: ref.captureTs, + waybackUrl: waybackPageUrl(ref), + rawUrl: waybackRawUrl(ref), + }; +} + +export function coerceWaybackProvenance(value: unknown): WaybackProvenance | null { + if (!value || typeof value !== "object" || Array.isArray(value)) return null; + const v = value as Record<string, unknown>; + const str = (k: string) => (typeof v[k] === "string" && v[k] ? (v[k] as string) : null); + const originalUrl = str("originalUrl"); + const captureTs = str("captureTs"); + const waybackUrl = str("waybackUrl"); + const rawUrl = str("rawUrl"); + if (!originalUrl || !captureTs || !/^\d{4,14}$/.test(captureTs) || !waybackUrl || !rawUrl) return null; + return { originalUrl, captureTs, waybackUrl, rawUrl }; +} + +export function sameWaybackProvenance(a: WaybackProvenance, b: WaybackProvenance): boolean { + return ( + a.originalUrl === b.originalUrl && + a.captureTs === b.captureTs && + a.waybackUrl === b.waybackUrl && + a.rawUrl === b.rawUrl + ); +} + +// ─── Citing it ─── + +export type WaybackLink = { label: string; url: string }; + +// What a citation of an archived copy links: the original, named as the +// original and as possibly gone (a capture exists because it may be), and the +// Wayback copy, which plays. `originalMomentUrl` is the original at the cited +// second when its platform takes one (lib/momentUrl.ts). +export function waybackCitationLinks( + prov: WaybackProvenance, + opts: { originalMomentUrl?: string | null } = {}, +): { original: WaybackLink; copy: WaybackLink } { + return { + original: { label: "Original (may be gone)", url: opts.originalMomentUrl || prov.originalUrl }, + copy: { label: `Wayback Machine copy, ${waybackCaptureDate(prov.captureTs)}`, url: prov.waybackUrl }, + }; +} diff --git a/common/publish/composeReports.ts b/common/publish/composeReports.ts @@ -78,6 +78,8 @@ import { WHISPER_FILENAME, isEnglishVtt, resolvePrimaryVtt } from "../lib/videoS import { platformMomentUrl } from "../lib/momentUrl"; import { archiveOrgCitationLinks, type ArchiveOrgProvenance } from "../lib/archiveOrg"; import { loadArchiveOrgProvenance } from "../lib/archiveOrg-server"; +import { WAYBACK_PROVENANCE_FILENAME, waybackCitationLinks, type WaybackProvenance } from "../lib/wayback"; +import { loadWaybackProvenance } from "../lib/wayback-server"; import type { Platform } from "../lib/platform"; import { readAllPosts } from "../lib/posts-server"; import type { Post } from "../lib/posts"; @@ -236,6 +238,9 @@ type CitedRecord = { // An archive.org record's provenance (its torrent, a mirror's original); // null for every other record. archiveOrg: ArchiveOrgProvenance | null; + // A Wayback Machine capture's provenance (lib/wayback.ts): what the record + // is an archived copy of. Null for every other record. + wayback: WaybackProvenance | null; }; // The English VTT tracks of a video dir, `en-orig` first, then the order @@ -316,7 +321,8 @@ export async function readCitedRecord( } if (tracks.length === 0 && cues.length > 0) tracks.push({ name: "cues", cues }); const archiveOrg = summary.platform === "archiveorg" ? await loadArchiveOrgProvenance(dir) : null; - return { summary, cues, tracks, archiveOrg }; + const wayback = entries.includes(WAYBACK_PROVENANCE_FILENAME) ? await loadWaybackProvenance(dir) : null; + return { summary, cues, tracks, archiveOrg, wayback }; } const isoDay = (uploadDate: string | undefined): string | undefined => @@ -629,9 +635,28 @@ export async function resolveSiteReports(opts: ResolveSiteReportsOptions): Promi corpusUrl: cited ? undefined : corpusLink(`${c.channel}/${c.id}`, { vm: "post" }), }); } - const { summary, archiveOrg } = (await recordOf(c.channel, c.id))!; + const { summary, archiveOrg, wayback } = (await recordOf(c.channel, c.id))!; const audioOnly = isAudioOnlyPlatform(config?.platform); const seconds = Math.max(0, Math.floor(c.start)); + if (wayback) { + // An archived copy (Wayback Machine): the original, named as the + // original and as possibly gone, then the copy, which plays. + const links = waybackCitationLinks(wayback, { + originalMomentUrl: platformMomentUrl(wayback.originalUrl, null, c.start), + }); + return defined({ + channel: c.channel, + channelTitle: config?.name ?? (summary.channel || undefined), + id: c.id, + title: summary.title, + date: isoDay(summary.uploadDate), + platform: summary.platform, + originalUrl: links.original.url, + originalLabel: links.original.label, + downloads: [links.copy], + corpusUrl: cited ? undefined : corpusLink(summary.slug ?? `${c.channel}/${summary.id}`, seconds > 0 ? { t: String(seconds) } : {}), + }); + } if (summary.platform === "archiveorg") { // archive.org: the original (YouTube at the second, for a mirror; else // the archive.org page) plus the downloads a reader can check it from. diff --git a/common/publish/composeReportsWayback.test.ts b/common/publish/composeReportsWayback.test.ts @@ -0,0 +1,54 @@ +// A cited Wayback Machine copy carries its provenance into compose, and the +// links a citation of it shows are derived from that (lib/wayback.ts +// waybackCitationLinks): the original at the second, named as the original and +// as possibly gone, then the Wayback copy. Every id here is invented. +// +// Run with: node_modules/.bin/tsx --test publish/composeReportsWayback.test.ts + +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { mkdirSync, mkdtempSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import path from "node:path"; +import { readCitedRecord } from "./composeReports"; +import { buildWaybackProvenance, waybackCitationLinks } from "../lib/wayback"; +import { platformMomentUrl } from "../lib/momentUrl"; + +const YT = "Xyz987abc65"; +const CAPTURE = `https://web.archive.org/web/20190807060504/https://www.youtube.com/watch?v=${YT}`; + +test("readCitedRecord loads a Wayback copy's provenance; others get null", async () => { + const channels = mkdtempSync(path.join(tmpdir(), "compose-wayback-")); + const dir = path.join(channels, "demo-wayback", "data", YT); + mkdirSync(dir, { recursive: true }); + writeFileSync( + path.join(dir, "metadata.info.json"), + JSON.stringify({ + id: YT, + extractor_key: "YoutubeWebArchive", + title: "An archived upload", + upload_date: "20090102", + webpage_url: CAPTURE, + }), + ); + const prov = buildWaybackProvenance(CAPTURE)!; + writeFileSync(path.join(dir, "wayback.json"), JSON.stringify(prov)); + + const rec = await readCitedRecord(channels, "demo-wayback", YT); + assert.ok(rec); + assert.equal(rec.summary.platform, "youtube"); + assert.equal(rec.summary.id, YT); + assert.deepEqual(rec.wayback, prov); + // The record's own moment link is the capture, which plays. + assert.equal(platformMomentUrl(rec.summary.webpageUrl, "youtube", 42), CAPTURE); + const links = waybackCitationLinks(rec.wayback!, { + originalMomentUrl: platformMomentUrl(rec.wayback!.originalUrl, null, 42), + }); + assert.deepEqual(links.original, { label: "Original (may be gone)", url: `https://www.youtube.com/watch?v=${YT}&t=42s` }); + assert.deepEqual(links.copy, { label: "Wayback Machine copy, 2019-08-07", url: CAPTURE }); + + const plain = path.join(channels, "demo-wayback", "data", "Plain12345a"); + mkdirSync(plain, { recursive: true }); + writeFileSync(path.join(plain, "metadata.info.json"), JSON.stringify({ id: "Plain12345a", extractor_key: "Youtube" })); + assert.equal((await readCitedRecord(channels, "demo-wayback", "Plain12345a"))?.wayback, null); +}); diff --git a/common/ytdlp/downloadOneManaged.ts b/common/ytdlp/downloadOneManaged.ts @@ -1,7 +1,7 @@ import { removeMediaFile } from "../lib/mediaTier-server"; import { tierVideoDir } from "../lib/mediaTier-server"; import path from "node:path"; -import { appendFile, mkdir, readdir, readFile, rm, stat } from "node:fs/promises"; +import { access, appendFile, mkdir, readdir, readFile, rm, stat } from "node:fs/promises"; import { createWriteStream, type Dirent, type WriteStream } from "node:fs"; import { execa } from "execa"; import { @@ -50,6 +50,8 @@ import { writeDownloadOutcome } from "../lib/downloadOutcome-server"; import { formatBytes } from "../lib/format"; import { recordAvailability } from "../lib/availability-server"; import { withMetadataHistory } from "../lib/metadataHistory-server"; +import { ensureWaybackProvenance } from "../lib/wayback-server"; +import { parseWaybackUrl } from "../lib/wayback"; import { loadRawMetadata, loadRawMetadataFromDir, @@ -656,12 +658,37 @@ export async function downloadOneManaged( if (detectPlatform(opts.videoUrl) === "archiveorg") { return await downloadArchiveOrgManaged(opts, opts.archiveOrgDeps); } - return await runManagedDownload(opts, channelDir, startedAt, canonicalId); + const outcome = await runManagedDownload(opts, channelDir, startedAt, canonicalId); + await recordWaybackProvenance(opts, channelDir, canonicalId); + return outcome; } finally { logStream?.end(); } } +// A WAYBACK CAPTURE IS A COPY (lib/wayback.ts): once the record exists — its +// metadata.info.json written by the prefetch or the download — the +// `wayback.json` sidecar says of what. Built from the URL alone, so it costs no +// request; never fails the download. +async function recordWaybackProvenance( + opts: ManagedDownloadOpts, + channelDir: string, + canonicalId: string | null, +): Promise<void> { + if (!canonicalId || !parseWaybackUrl(opts.videoUrl)) return; + const videoDir = path.join(channelDir, "data", canonicalId); + try { + await access(path.join(videoDir, "metadata.info.json")); + } catch { + return; + } + try { + await ensureWaybackProvenance(videoDir, opts.videoUrl, { onLog: opts.onLog }); + } catch (err) { + opts.onLog(`Wayback provenance not written: ${(err as Error).message}\n`); + } +} + async function runManagedDownload( opts: ManagedDownloadOpts, channelDir: string, diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,8 @@ # Changelog ## [Unreleased] +- **A Wayback Machine capture is a copy, and says of what.** A capture URL (`web.archive.org/web/<timestamp>[id_|im_|…]/<original>`) names its record by what it is a capture of: an archived YouTube page by its YouTube id (no longer `watch`), a JW Player file by its media id (no longer `<id>-<rendition>.mp4`). Every download of a capture writes `wayback.json` (the original URL, the capture's timestamp, the capture page and its raw bytes); the video page says "Archived copy (Wayback Machine, <date>) of <original>"; a citation links the original, marked as possibly gone, and the Wayback copy, and its moment link is the capture, which plays (a capture URL never takes a time param). An existing record whose page is a capture is renamed to its id by the next snapshot. +- **`archilyzer wayback refresh <slug> [--titles <file>] [--dry-run]`** brings a channel's Wayback copies up to that offline: `wayback.json`, the dir renamed through the snapshot's own reconcile pass with its roster entry moved, and with `--titles` (`id → {title, upload_date}`) the title and date of a raw file that has none, recorded in the metadata history as `wayback-provenance`. A record a live job holds is skipped and named; a second run changes nothing. - **A cited moment at the very end of a recording prepares.** Prepare evidence media cuts a clip whose padding runs past the recording's end at the end (the recording's duration from its metadata), where it found no media for the padded span; a span that starts past the end is still refused. report-to-video keeps its strict rule. - **Exporting a changed report records a new revision of it.** `reports export` (and **Export reports** on a site's Reports tab, and the end of a prepare) commits a revision to the report's own git history, `sites/<site>/reports/<id>/history-git/`, whenever its `report.json` changed since the last one: the `report.json`, its Markdown export and the checksums of every export file, with a message of `Revision N` and a summary of the change. A re-export of an unchanged report records nothing. The commits carry the site's name and a `noreply@<site>.invalid` address with dates in UTC, never your git name, email or time zone. The Reports tab shows each report's revision, its commit and the last change under **Exports**, and the site's next build publishes the history. Add `history-git/` to the corpus repository's `.gitignore`. - **archive.org files come over BitTorrent when possible, else straight from archive.org — never through yt-dlp.** The chosen file of an archive.org import is fetched from the item's own torrent (`<identifier>_archive.torrent`, which lists archive.org as a web seed, so other peers take load off archive.org) with aria2c, only that file of the item, and seeded afterwards for 10 minutes or to a ratio of 1, whichever comes first; the log shows "torrent: <file> (n of m pieces, peers p, web seed yes)" and "seeding 10 min…". With no aria2c, a torrent that does not carry the file, or no progress for 5 minutes, it is downloaded directly from `archive.org/download/…` instead (resumable, backing off on 429/503), and the log says "fell back to direct download: <reason>". Every file is checked against archive.org's sha1/md5: a mismatch is downloaded once more directly, a second one fails the record. The record is written from the item's metadata: `metadata.info.json` with the file's page, the canonical id, the duration ffprobe measures and archive.org's playable copies of the file, the `archiveorg.json` provenance (a mirror's original title, date and uploader), and `audio.<fmt>` — an audio file already in the channel's format is used as is, anything else goes through the app's audio extraction, a video kept in the saved-video store when the channel keeps sources. An .avi/.mpeg/.flac/.wav original is fetched as archive.org's mp4 or mp3 of it. aria2c runs in its own process group: cancelling the job stops it and everything it started, and it stops itself if the editor exits. New settings block `archiveOrg` (`torrent`, `seedMinutes`, `seedRatio`, `stallMinutes`, `maxPeers`, `maxDownloadKiBps`, `maxUploadKiBps`), `ARIA2C_BIN`, an aria2c row in `archilyzer doctor`, and `aria2` in the runtime Docker images. diff --git a/editor/app/channels/[slug]/videos/[id]/page.tsx b/editor/app/channels/[slug]/videos/[id]/page.tsx @@ -13,6 +13,8 @@ import { isExcludedFromTruncatedCheck } from "yt-dlp-transcript-common/lib/exclu import { loadSavedVideo } from "yt-dlp-transcript-common/lib/savedVideo-server"; import { loadArchiveOrgProvenance } from "yt-dlp-transcript-common/lib/archiveOrg-server"; import type { ArchiveOrgProvenance } from "yt-dlp-transcript-common/lib/archiveOrg"; +import { loadWaybackProvenance } from "yt-dlp-transcript-common/lib/wayback-server"; +import { waybackCaptureDate, type WaybackProvenance } from "yt-dlp-transcript-common/lib/wayback"; import { getPaths } from "yt-dlp-transcript-common/lib/paths"; import { readVideoMetadataForDisplay, @@ -139,6 +141,8 @@ export default async function VideoDetailPage({ // Where an archive.org record came from (lib/archiveOrg-server.ts): the // item and its torrent, and a mirror's original. Absent everywhere else. const archiveOrg = await loadArchiveOrgProvenance(videoDir); + // What a Wayback Machine copy is a copy of (lib/wayback-server.ts). + const wayback = await loadWaybackProvenance(videoDir); // The windows another tool asked this editor to fetch. One readdir of // data/<id>/clips/ plus a stat per file — and no per-CHANNEL count anywhere, // because that would be a walk of every video dir to draw one number. @@ -191,6 +195,7 @@ export default async function VideoDetailPage({ excludedFromTruncatedCheck, savedVideo, archiveOrg, + wayback, clipWindows, vttProvenance, coverage, @@ -219,6 +224,7 @@ export default async function VideoDetailPage({ excludedFromTruncatedCheck, savedVideo, archiveOrg, + wayback, clipWindows, vttProvenance, coverage, @@ -285,6 +291,7 @@ export default async function VideoDetailPage({ )} </div> {archiveOrg && <ArchiveOrgProvenanceLine prov={archiveOrg} />} + {wayback && <WaybackProvenanceLine prov={wayback} />} {meta.description && ( <details className="text-sm"> <summary className="cursor-pointer text-muted-foreground hover:text-foreground"> @@ -385,6 +392,23 @@ function ArchiveOrgProvenanceLine({ prov }: { prov: ArchiveOrgProvenance }) { ); } +// "Archived copy (Wayback Machine, <capture date>) of <original>". +function WaybackProvenanceLine({ prov }: { prov: WaybackProvenance }) { + const link = "underline hover:text-foreground"; + return ( + <div aria-label="Wayback provenance" className="text-sm text-muted-foreground"> + Archived copy ( + <a href={prov.waybackUrl} target="_blank" rel="noreferrer" className={link}> + Wayback Machine, {waybackCaptureDate(prov.captureTs)} + </a> + ) of{" "} + <a href={prov.originalUrl} target="_blank" rel="noreferrer" className={`${link} break-all`}> + {prov.originalUrl} + </a> + </div> + ); +} + function formatUploadDate(s: string): string { // yt-dlp emits YYYYMMDD. Render as YYYY-MM-DD; pass through anything else. if (/^\d{8}$/.test(s)) { diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md @@ -1,6 +1,7 @@ # Changelog ## [Unreleased] +- **A citation of a Wayback Machine copy links its original and the copy.** A cited record downloaded from a Wayback capture shows "Original (may be gone)", the original at the cited second where its platform takes one, and "Wayback Machine copy, <capture date>", the capture page, which plays. Its moment link is the capture: a capture URL never takes a time param. - **A report shows its revision, and every edit to it can be checked.** A report's date line ends with "revision N", linking to its history ("edited since revision N" when the report has changed since). The history page, `/reports/<id>/history/`, lists every revision, newest first: its number, date (UTC), commit hash and the sha256 of its `report.json`, what changed (claims added or removed, verdicts changed, claims edited, citations added or removed, quotes edited, title, series or subtitle changed), and each changed claim's title, text, verdict and findings with the words removed struck through and the words added marked. The same data is in `history.json` beside the page. Each report's history is its own git repository, published for cloning: `git clone <site>/reports/<id>/history/repo`. Each commit names the site as its author, with its date in UTC. The footer of the report's HTML, PDF and Markdown downloads begins with the revision number, and its sha256 can be looked up on the history page. Needs `reports export` and a rebuild and deploy of each site with reports. - **A report can be saved whole: as one HTML page, a PDF, Markdown, or an evidence pack.** A report page's download line reads HTML · PDF · Markdown · Evidence pack · Citations JSON · CSV, each listed only when the site publishes it. The HTML is one file that opens with no network: the report with its verdicts, the document's sentences and the post screenshots inside it, numbered citations, and a reference list giving each quote's speaker, date, record, the original at its time and the moment page on the site. The PDF is that page printed. The Markdown is the same report as plain text with numbered references. The evidence pack is a zip of the page with its clips, stills and screenshots beside it, so the clips play offline. Each ends with a line naming the report's revision, its date and the start of its checksum. Needs `reports export` (or prepare) and a rebuild and deploy of each site with reports. - **A report's claim can carry a flag, its header names the document under review, and a site with one report names it in the browser tab.** `report.json` claim `flag` (one line, at most 60 characters) shows as a small pill in the accent colour beside the claim's verdict, e.g. "No source given". On a report-only site with one report, the home page's tab title is the report's, as on the report's own page. A report's page header is its name — with a `series`, the series on one line in the accent colour and the title on the line below; without one, the title — then one small line of dates and the revision ("2026-10-04 · updated 2026-10-05 · revision 1"; "updated" only when it differs), then a card for the document under review: its title linking to the document, "<author> · <publisher> · <date>", and its archive links folded away, on a left rail in the document's colour (a source's `accent`, `"#rrggbb"`; without one, the border colour), then the subtitle. The page names no byline or site of its own. A claim that cites the document's own sentence shows that sentence — its still, else its words — on the same rail, with no link up to the card and no paraphrase beside it; a sentence of another document links "from <title>" to that document's box; a titled claim with no such sentence shows its text under the title, plain. A report's citation can say where its evidence came from (`origin`: `"subject"`, the document under review gave it; `"added"`, the report's author found it). A claim lists what the report added first, each card marked with the Archilyzer mark under its number ("Not in the article", or "Not in the source", is the mark's tooltip and what a screen reader says), then evidence of unknown origin, then what the document gave itself folded under "In the article (n)"; the reference list marks an added citation with the mark too, and a claim's flag pill wears the same mark. A fact-check's page reads in three tiers, each opened by a hairline with one, two or three dots: the quick take (the tally, the summary, and links to what the check found, every claim and the downloads); **What the check found**, every ruled claim grouped by verdict (contradicted, not found, partly, untestable, corroborated), one line each linking to the claim, with its `gist` (a new optional claim field, one line, at most 240 characters) and its flag; and **Every claim, with its evidence**, which opens with **How it was checked** (`method`, a new optional report field in markdown). A report of kind `sweep` has the first and last tiers only. Needs a rebuild and deploy of the site.