Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit b60db16f328ad7639678e03415549ed5a677baf7
parent f0db76164ba11789ac3a589232803a83576b1bfa
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon,  5 Oct 2026 19:19:54 -0400

Merge sources/archive-org-titles (archive.org file records: clean titles and upload dates from the file name; archive-org refresh)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
MREADME.md | 7++++++-
Mcommon/bin/archilyzer.ts | 18++++++++++++++++++
Acommon/bin/archive-org-refresh.ts | 37+++++++++++++++++++++++++++++++++++++
Acommon/controller/archiveOrgRefresh.test.ts | 167+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/controller/archiveOrgRefresh.ts | 106+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/lib/archiveOrg.test.ts | 61+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/lib/archiveOrg.ts | 45+++++++++++++++++++++++++++++++++++++++++----
Mcommon/lib/archiveOrgId.test.ts | 55+++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/lib/archiveOrgId.ts | 59+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/CHANGELOG.md | 2+-
10 files changed, 551 insertions(+), 6 deletions(-)

diff --git a/README.md b/README.md @@ -154,7 +154,12 @@ creator and collections, the item's torrent, and — for a mirror of a YouTube u the original's id, URL, title and upload date, read from the `.info.json` uploaded with it. The video page shows both; a citation of the record links **archive.org** and the **torrent** (and, for a mirror, the original **YouTube** upload at the cited -second), so a reader can fetch the file and check it. +second), so a reader can fetch the file and check it. A file with no uploaded +`.info.json` is dated by the `YYYYMMDD` its name starts with and, when the item gives it +no title of its own, titled from its name — the leading date, the `[<n> views]` count, +the YouTube id and the extension off — and +`pnpm archilyzer archive-org refresh <channel> [--dry-run]` brings a channel's existing +records to that title and date, offline, printing old → new. **Files come over BitTorrent when possible** — to be extra polite to archive.org. Every item has a torrent (`<identifier>_archive.torrent`) that lists archive.org itself as a diff --git a/common/bin/archilyzer.ts b/common/bin/archilyzer.ts @@ -333,6 +333,24 @@ export const COMMANDS: Command[] = [ }, }, { + path: ["archive-org", "refresh"], + usage: + "<slug> [--dry-run] bring a channel's archive.org file records up to their provenance: a name-only mirror's title (where the item gives the file none) and upload date from its file name; offline, through the metadata history; prints old → new", + flags: { "dry-run": "boolean" }, + maxPositionals: 1, + run: async ({ positionals, flags }) => { + const [slug] = positionals; + if (!slug) { + console.error("archive-org refresh: which channel? Pass its slug."); + return 2; + } + return (await import("./archive-org-refresh")).main({ + slug, + dryRun: flags["dry-run"] === true, + }); + }, + }, + { path: ["feeds", "backfill-metadata"], usage: "<slug> [--feed <url>] [--dry-run] complete a podcast channel's records (title, date, description, duration) from its RSS feed: one fetch of the feed (default: the channel's url), no media; --dry-run counts matched / unmatched / already complete and writes nothing", diff --git a/common/bin/archive-org-refresh.ts b/common/bin/archive-org-refresh.ts @@ -0,0 +1,37 @@ +// `archilyzer archive-org refresh <slug> [--dry-run]` — every archive.org file +// record of a channel brought up to its provenance's rules +// (controller/archiveOrgRefresh.ts): the title a file's name carries (where the +// item gives the file none) and the day it starts with, for a mirror with no +// uploaded info.json. Offline: nothing is fetched. Prints old → new per key; a +// second run changes nothing. + +import { isValidChannelSlug } from "../controller/channels"; +import { refreshArchiveOrgRecords } from "../controller/archiveOrgRefresh"; +import { getPaths, type Paths } from "../lib/paths"; + +export async function main(opts: { slug: string; dryRun: boolean; paths?: Paths }): Promise<number> { + if (!isValidChannelSlug(opts.slug)) { + console.error(`archive-org refresh: "${opts.slug}" is not a channel slug`); + return 2; + } + try { + const r = await refreshArchiveOrgRecords({ + slug: opts.slug, + paths: opts.paths ?? getPaths(), + dryRun: opts.dryRun, + onLog: (line) => console.log(line.replace(/\n$/, "")), + }); + const would = opts.dryRun ? "would be " : ""; + console.log( + `${opts.dryRun ? "dry run: " : ""}${r.fileRecords} archive.org file records, ` + + `${r.changed.length} ${would}refreshed, ` + + `${r.sidecars} provenance sidecars ${would}updated` + + (r.failed.length > 0 ? `, ${r.failed.length} failed` : "") + + ".", + ); + return r.failed.length > 0 ? 1 : 0; + } catch (err) { + console.error(`archive-org refresh: ${(err as Error).message}`); + return 1; + } +} diff --git a/common/controller/archiveOrgRefresh.test.ts b/common/controller/archiveOrgRefresh.test.ts @@ -0,0 +1,167 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { mkdir, mkdtemp, readFile, rm, writeFile } from "node:fs/promises"; +import { tmpdir } from "node:os"; +import path from "node:path"; +import type { Paths } from "../lib/paths"; +import { buildArchiveOrgProvenance, parseArchiveOrgItemMetadata } from "../lib/archiveOrg"; +import { archiveOrgVideoId } from "../lib/archiveOrgId"; +import { loadMetadataHistory } from "../lib/metadataHistory-server"; +import { refreshArchiveOrgRecords } from "./archiveOrgRefresh"; +import { main as refreshCli } from "../bin/archive-org-refresh"; + +// Run with: +// pnpm --filter yt-dlp-transcript-common exec tsx --test controller/archiveOrgRefresh.test.ts +// +// A temp corpus of archive.org records written before file-name titles: the +// raw name as the title, the item's date, and no mirror title or date. Every name here is invented. + +const SLUG = "demo-archive"; +const ITEM = "example-channel-archive"; +const RAW = "extras_20210102 A Talk _ Some Show [4321 views]-Pqr678stu_4.mp4"; +const TITLED = "Second Upload-Def456uvw_8.mp4"; +const MIRRORED = "Third Upload-Ghi789rst_7.mp4"; + +const item = parseArchiveOrgItemMetadata({ + metadata: { identifier: ITEM, title: "Example Channel Archive", date: "2024-01-02" }, + files: [ + { name: RAW, source: "original" }, + { name: TITLED, source: "original", title: "Second, as titled on archive.org" }, + { name: MIRRORED, source: "original" }, + ], +})!; + +const stem = (f: string) => f.replace(/\.mp4$/, ""); + +// The sidecar as an older import wrote it: no title or date for a name-only mirror. +function oldProvenance(file: string, infoJson?: unknown) { + const prov = buildArchiveOrgProvenance({ + ref: { identifier: ITEM, file }, + item, + infoJson, + fetchedAt: "2026-01-01T00:00:00.000Z", + }); + if (prov.mirror?.from === "file-name") { + delete prov.mirror.title; + delete prov.mirror.uploadDate; + } + return prov; +} + +async function withCorpus(fn: (paths: Paths, dataDir: string) => Promise<void>): Promise<void> { + const dir = await mkdtemp(path.join(tmpdir(), "ttb-ia-titles-")); + const transcriptsDir = path.join(dir, "corpus"); + const paths = { transcriptsDir, channelsDir: path.join(transcriptsDir, "channels") } as Paths; + const channelDir = path.join(paths.channelsDir, SLUG); + const dataDir = path.join(channelDir, "data"); + await mkdir(dataDir, { recursive: true }); + await writeFile( + path.join(channelDir, "config.json"), + JSON.stringify({ name: "Demo", platform: "archiveorg", handling: "transcribe" }), + ); + const seed: [string, unknown, Record<string, unknown>][] = [ + [archiveOrgVideoId({ identifier: ITEM, file: RAW }), oldProvenance(RAW), { title: stem(RAW), timestamp: 1704153600 }], + [archiveOrgVideoId({ identifier: ITEM, file: TITLED }), oldProvenance(TITLED), { title: "Second, as titled on archive.org" }], + [ + archiveOrgVideoId({ identifier: ITEM, file: MIRRORED }), + oldProvenance(MIRRORED, { id: "Ghi789rst_7", extractor_key: "Youtube", title: "The Original Title" }), + { title: "The Original Title" }, + ], + // A whole item: not one file of many, never re-titled. + [ + "example-single", + buildArchiveOrgProvenance({ + ref: { identifier: "example-single" }, + item: { ...item, metadata: { identifier: "example-single", title: "Single" } }, + fetchedAt: "2026-01-01T00:00:00.000Z", + }), + { title: "Single" }, + ], + ]; + for (const [id, prov, info] of seed) { + await mkdir(path.join(dataDir, id), { recursive: true }); + await writeFile(path.join(dataDir, id, "archiveorg.json"), JSON.stringify(prov)); + await writeFile( + path.join(dataDir, id, "metadata.info.json"), + JSON.stringify({ id, extractor_key: "ArchiveOrg", ...info, upload_date: "20240102" }), + ); + } + try { + await fn(paths, dataDir); + } finally { + await rm(dir, { recursive: true, force: true }); + } +} + +const RAW_ID = archiveOrgVideoId({ identifier: ITEM, file: RAW }); + +test("refresh: a dry run lists old → new and writes nothing", async () => { + await withCorpus(async (paths, dataDir) => { + const before = await readFile(path.join(dataDir, RAW_ID, "metadata.info.json"), "utf8"); + const beforeProv = await readFile(path.join(dataDir, RAW_ID, "archiveorg.json"), "utf8"); + const r = await refreshArchiveOrgRecords({ paths, slug: SLUG, dryRun: true }); + assert.equal(r.fileRecords, 3); + assert.deepEqual(r.changed, [ + { + id: RAW_ID, + changes: { + title: { from: stem(RAW), to: "A Talk | Some Show" }, + upload_date: { from: "20240102", to: "20210102" }, + timestamp: { from: 1704153600, to: null }, + }, + }, + ]); + assert.equal(r.sidecars, 1); + assert.equal(await readFile(path.join(dataDir, RAW_ID, "metadata.info.json"), "utf8"), before); + assert.equal(await readFile(path.join(dataDir, RAW_ID, "archiveorg.json"), "utf8"), beforeProv); + assert.equal(await loadMetadataHistory(path.join(dataDir, RAW_ID)), null); + + // The CLI over the same corpus: prints the change, exits 0, writes nothing. + const lines: string[] = []; + const orig = console.log; + console.log = (...a: unknown[]) => void lines.push(a.join(" ")); + try { + assert.equal(await refreshCli({ slug: SLUG, dryRun: true, paths }), 0); + } finally { + console.log = orig; + } + const out = lines.join("\n"); + assert.match(out, /→ "A Talk \| Some Show"/); + assert.match(out, /upload_date: "20240102"\n {4}→ "20210102"/); + assert.match(out, /dry run: 3 archive\.org file records, 1 would be refreshed/); + assert.equal(await readFile(path.join(dataDir, RAW_ID, "metadata.info.json"), "utf8"), before); + }); +}); + +test("refresh: title and date go through the metadata history, the raw name stays as `file`; idempotent", async () => { + await withCorpus(async (paths, dataDir) => { + const r = await refreshArchiveOrgRecords({ paths, slug: SLUG }); + assert.equal(r.changed.length, 1); + assert.deepEqual(r.failed, []); + const dir = path.join(dataDir, RAW_ID); + const info = JSON.parse(await readFile(path.join(dir, "metadata.info.json"), "utf8")); + assert.equal(info.title, "A Talk | Some Show"); + assert.equal(info.upload_date, "20210102"); + assert.equal(info.timestamp ?? null, null); + assert.equal(Object.keys(info).at(-1), "upload_date"); + const prov = JSON.parse(await readFile(path.join(dir, "archiveorg.json"), "utf8")); + assert.equal(prov.file, RAW); + assert.equal(prov.mirror.title, "A Talk | Some Show"); + assert.equal(prov.mirror.uploadDate, "20210102"); + const entry = (await loadMetadataHistory(dir))?.entries.at(-1); + assert.equal(entry?.by, "archiveorg-provenance"); + assert.deepEqual(entry?.changed.title, { from: stem(RAW), to: "A Talk | Some Show" }); + assert.deepEqual(entry?.changed.upload_date, { from: "20240102", to: "20210102" }); + + const again = await refreshArchiveOrgRecords({ paths, slug: SLUG }); + assert.deepEqual(again.changed, []); + assert.equal(again.sidecars, 0); + assert.equal((await loadMetadataHistory(dir))?.entries.length, 1); + }); +}); + +test("refresh: an unknown channel is refused", async () => { + await withCorpus(async (paths) => { + await assert.rejects(refreshArchiveOrgRecords({ paths, slug: "no-such-channel" }), /not found/); + }); +}); diff --git a/common/controller/archiveOrgRefresh.ts b/common/controller/archiveOrgRefresh.ts @@ -0,0 +1,106 @@ +// ARCHIVE.ORG FILE RECORDS, BROUGHT UP TO THEIR PROVENANCE'S RULES — offline. +// +// A record imported as ONE FILE of a multi-file item whose mirror has no +// uploaded info.json was titled with the raw file name and dated with the +// item's date. The rule is lib/archiveOrg.ts (withFileNameFields → +// archiveOrgMetadataPatch): the title its name carries (where the item gives +// the file none) and the day its name starts with. This brings every such +// record of a channel up to it: +// +// archiveorg.json the mirror's `title` and `uploadDate`, when the rule +// gives different ones (from a name only; an +// info.json's are never touched); `file` keeps the raw +// name +// metadata.info.json `title`, `upload_date` and `timestamp` (dropped +// beside a mirror's date, as at import), through +// `patchMetadataInfo`, so the change is in +// metadata.history.json as `archiveorg-provenance` +// +// Only those keys — every other key the provenance decides was written at +// import. Only when they differ, so a second run writes nothing. No network. + +import path from "node:path"; +import { getPaths, type Paths } from "../lib/paths"; +import { readJsonFile } from "../lib/jsonFile-server"; +import { assertChannelTextReadable } from "../lib/channelMedia"; +import { archiveOrgMetadataPatch, withFileNameFields } from "../lib/archiveOrg"; +import { loadArchiveOrgProvenance, writeArchiveOrgProvenance } from "../lib/archiveOrg-server"; +import { patchMetadataInfo } from "../lib/metadataHistory-server"; +import { readChannelConfig } from "./channels"; +import { listChannelVideoIds } from "./keptVideos"; + +const REFRESHED_KEYS = ["title", "upload_date", "timestamp"] as const; + +export type ArchiveOrgRecordChange = { + id: string; + // Per key: the value before and after (null: absent / removed). + changes: Partial<Record<(typeof REFRESHED_KEYS)[number], { from: unknown; to: unknown }>>; +}; + +export type ArchiveOrgRefreshResult = { + // archive.org records of one file of an item. + fileRecords: number; + changed: ArchiveOrgRecordChange[]; + // Sidecars whose mirror title or date was (or would be) rewritten. + sidecars: number; + failed: { id: string; error: string }[]; +}; + +export async function refreshArchiveOrgRecords(opts: { + slug: string; + paths?: Paths; + dryRun?: boolean; + onLog?: (line: string) => void; +}): Promise<ArchiveOrgRefreshResult> { + const paths = opts.paths ?? getPaths(); + const log = opts.onLog ?? (() => {}); + const config = await readChannelConfig(paths, opts.slug); + if (!config) throw new Error(`Channel "${opts.slug}" not found`); + // An unreadable text tier is not an empty channel (AGENTS.md). + await assertChannelTextReadable(paths, opts.slug, config); + + const dataDir = path.join(paths.channelsDir, opts.slug, "data"); + const result: ArchiveOrgRefreshResult = { fileRecords: 0, changed: [], sidecars: 0, failed: [] }; + for (const id of (await listChannelVideoIds(paths, opts.slug)).sort()) { + const videoDir = path.join(dataDir, id); + const prov = await loadArchiveOrgProvenance(videoDir); + if (!prov?.file) continue; + result.fileRecords++; + try { + const next = withFileNameFields(prov); + if (next !== prov) { + result.sidecars++; + if (!opts.dryRun) await writeArchiveOrgProvenance(videoDir, next); + } + const read = await readJsonFile(path.join(videoDir, "metadata.info.json")); + const info = + read.ok && read.value && typeof read.value === "object" && !Array.isArray(read.value) + ? (read.value as Record<string, unknown>) + : null; + if (!info) continue; + const full = archiveOrgMetadataPatch(next, info); + const patch: Record<string, unknown> = {}; + const change: ArchiveOrgRecordChange = { id, changes: {} }; + for (const key of REFRESHED_KEYS) { + if (!(key in full)) continue; + patch[key] = full[key]; + change.changes[key] = { from: info[key] ?? null, to: full[key] }; + } + if (Object.keys(patch).length === 0) continue; + result.changed.push(change); + log( + `${id}\n` + + Object.entries(change.changes) + .map(([k, c]) => ` ${k}: ${JSON.stringify(c!.from)}\n → ${JSON.stringify(c!.to)}`) + .join("\n"), + ); + if (!opts.dryRun) { + await patchMetadataInfo(videoDir, patch, { by: "archiveorg-provenance", requestedBy: "cli", onLog: log }); + } + } catch (err) { + result.failed.push({ id, error: (err as Error).message }); + log(`${id}: failed — ${(err as Error).message}`); + } + } + return result; +} diff --git a/common/lib/archiveOrg.test.ts b/common/lib/archiveOrg.test.ts @@ -6,6 +6,7 @@ import { archiveOrgPlayableUrl, buildArchiveOrgProvenance, coerceArchiveOrgProvenance, + withFileNameFields, findArchiveOrgInfoJson, listArchiveOrgMediaFiles, parseArchiveOrgItemMetadata, @@ -143,6 +144,7 @@ test("without an info.json the mirror is known by name only; a plain item is no platform: "youtube", id: "Ghi789rst_7", url: "https://www.youtube.com/watch?v=Ghi789rst_7", + title: "Third Upload", from: "file-name", }); const byIdent = buildArchiveOrgProvenance({ @@ -367,3 +369,62 @@ test("a mirror's record is the original's: title, date, uploader from its upload assert.equal(info.duration, 10); assert.equal(info.webpage_url, "https://archive.org/details/youtube-AbC123xyz_9"); }); + +test("one file of many with no title of its own is titled and dated from its name, the name kept as `file`", () => { + const NAMED = "extras_20210102 A Talk _ Some Show [4321 views]-Pqr678stu_4.mp4"; + const many = parseArchiveOrgItemMetadata({ + metadata: { identifier: ITEM, title: "Example Channel Archive", date: "2024-01-02" }, + files: [ + { name: NAMED, source: "original", format: "MPEG4" }, + { name: V2, source: "original", format: "MPEG4", title: "Second, as titled on archive.org" }, + ], + })!; + const prov = buildArchiveOrgProvenance({ + ref: { identifier: ITEM, file: NAMED }, + item: many, + fetchedAt: "2026-01-01T00:00:00.000Z", + }); + assert.equal(prov.file, NAMED); + assert.equal(prov.mirror?.from, "file-name"); + assert.equal(prov.mirror?.title, "A Talk | Some Show"); + // The name's leading date is the original's day, not the item's. + assert.equal(prov.mirror?.uploadDate, "20210102"); + const rec = archiveOrgInfoJson({ prov, item: many, file: NAMED, fetched: many.files[0] }); + assert.equal(rec.title, "A Talk | Some Show"); + assert.equal(rec.upload_date, "20210102"); + assert.equal(Object.keys(rec).at(-1), "upload_date"); + // A record with archive.org's date and a timestamp: the name's day replaces + // both, as a mirror's date always has. + const patch = archiveOrgMetadataPatch(prov, { title: "x", upload_date: "20240102", timestamp: 1704153600 }); + assert.equal(patch.upload_date, "20210102"); + assert.equal(patch.timestamp, null); + + // A file with its own title in the item keeps it; the mirror is not titled from the name. + const titled = buildArchiveOrgProvenance({ + ref: { identifier: ITEM, file: V2 }, + item: many, + fetchedAt: "2026-01-01T00:00:00.000Z", + }); + assert.equal(titled.mirror?.title, undefined); + assert.equal(archiveOrgMetadataPatch(titled, {}).title, "Second, as titled on archive.org"); + + // An older sidecar (no mirror title) is brought up to the rule, once. + const old = { ...prov, mirror: { ...prov.mirror! } }; + delete old.mirror.title; + delete old.mirror.uploadDate; + const fixed = withFileNameFields(old); + assert.equal(fixed.mirror?.title, "A Talk | Some Show"); + assert.equal(fixed.mirror?.uploadDate, "20210102"); + assert.equal(withFileNameFields(fixed), fixed); + // An info.json's title and date are the original's own and are never replaced. + const fromInfo = { + ...prov, + mirror: { ...prov.mirror!, from: "info-json" as const, title: "Original", uploadDate: "20200101" }, + }; + assert.equal(withFileNameFields(fromInfo), fromInfo); + // An info.json without a date takes the name's. + const infoNoDate = { ...fromInfo, mirror: { ...fromInfo.mirror } }; + delete (infoNoDate.mirror as { uploadDate?: string }).uploadDate; + assert.equal(withFileNameFields(infoNoDate).mirror?.uploadDate, "20210102"); + assert.equal(withFileNameFields(infoNoDate).mirror?.title, "Original"); +}); diff --git a/common/lib/archiveOrg.ts b/common/lib/archiveOrg.ts @@ -24,6 +24,8 @@ import { archiveOrgTorrentUrl, archiveOrgVideoId, parseArchiveOrgUrl, + dateFromMirrorFileName, + titleFromMirrorFileName, youtubeIdFromFileName, youtubeIdFromIdentifier, type ArchiveOrgRef, @@ -292,7 +294,37 @@ export function buildArchiveOrgProvenance(opts: { ...(mirror ? { mirror } : {}), fetchedAt: opts.fetchedAt, }; - return stripUndefinedDeep(prov); + return stripUndefinedDeep(withFileNameFields(prov)); +} + +// A mirror's original known only by name (no info.json, or one without these +// fields) for one file of a multi-file item takes them from the file's name +// (archiveOrgId.ts): +// +// title titleFromMirrorFileName, unless the item gives the file a +// title of its own (`fileTitle`) +// uploadDate dateFromMirrorFileName, the leading `[<word>_]YYYYMMDD` +// +// What a name gives is re-derived every time, so a refresh follows the rule; +// an info.json's title and date are the original's own and always win. +// Returns `prov` itself when nothing changes. +export function withFileNameFields(prov: ArchiveOrgProvenance): ArchiveOrgProvenance { + const mirror = prov.mirror; + if (!prov.file || !mirror) return prov; + const byName = mirror.from !== "info-json"; + const next: ArchiveOrgMirror = { ...mirror }; + if (!prov.fileTitle && (byName || !mirror.title)) { + const title = titleFromMirrorFileName(prov.file); + if (title) next.title = title; + else delete next.title; + } + if (byName || !mirror.uploadDate) { + const date = dateFromMirrorFileName(prov.file); + if (date) next.uploadDate = date; + else delete next.uploadDate; + } + if (JSON.stringify(next) === JSON.stringify(mirror)) return prov; + return { ...prov, mirror: next }; } function stripUndefinedDeep<T>(v: T): T { @@ -340,8 +372,11 @@ export function coerceArchiveOrgProvenance(value: unknown): ArchiveOrgProvenance // `extractVideoId(webpage_url)` (reconcileVideoDirs.ts), so an // entry left with the item's page would be merged into the item. // title the original's title (a mirror), else the file's own title, -// else its file name — never the item's for one file of many. -// upload_date the original's date (a mirror), with its timestamp dropped +// else the title its file name carries (date, view count, id +// and extension off; titleFromMirrorFileName), else the name — +// never the item's for one file of many. +// upload_date the original's date (a mirror: its info.json's, else the +// date its file name starts with), with its timestamp dropped // description the original's (a mirror), when it had one // uploader the original's uploader, else the item's public credit — never // the uploading account's e-mail address @@ -355,7 +390,9 @@ export function archiveOrgMetadataPatch( want.webpage_url = prov.fileUrl ?? prov.itemUrl; const mirror = prov.mirror; const fileName = prov.file ? (prov.file.split("/").pop() ?? prov.file).replace(/\.[A-Za-z0-9]{1,8}$/, "") : undefined; - const title = mirror?.title ?? (prov.file ? (prov.fileTitle ?? fileName) : undefined); + const title = + mirror?.title ?? + (prov.file ? (prov.fileTitle ?? titleFromMirrorFileName(prov.file) ?? fileName) : undefined); if (title) want.title = title; if (mirror?.uploadDate) want.upload_date = mirror.uploadDate; if (mirror?.description) want.description = mirror.description; diff --git a/common/lib/archiveOrgId.test.ts b/common/lib/archiveOrgId.test.ts @@ -7,6 +7,8 @@ import { archiveOrgVideoId, archiveOrgVideoIdFromNativeId, parseArchiveOrgUrl, + dateFromMirrorFileName, + titleFromMirrorFileName, youtubeIdFromFileName, youtubeIdFromIdentifier, } from "./archiveOrgId"; @@ -93,3 +95,56 @@ test("YouTube ids read from mirror names", () => { assert.equal(youtubeIdFromFileName("interview-performance.mp4"), null); assert.equal(youtubeIdFromFileName("plain.mp4"), null); }); + +test("a mirrored file's name gives its title: date, view count, id and extension off", () => { + const cases: [string, string | null][] = [ + // `YYYYMMDD <title> [<n> views]-<id>.<ext>` + ["20210102 Example Talk at the Hall [1234 views]-AbC123xyz_9.mp4", "Example Talk at the Hall"], + // A `<word>_` before the date, ` _ ` for ` | `. + ["extras_20200304 A Long Chat _ Guest Name _ TOPIC _ Some Show [9876543 views]-Def456uvw_8.mp4", + "A Long Chat | Guest Name | TOPIC | Some Show"], + // No date; a separator inside the title; a dashed id that starts with `-`. + ["Speaker - 'A Quoted Line!' & Why! _ Daily Clip--bC123xyz_9.mp4", "Speaker - 'A Quoted Line!' & Why! | Daily Clip"], + // The bracketed id, a date then ` - `, commas in the count. + ["20191231 - New Year Stream [1,234 views] [Ghi789rst_7].mkv", "New Year Stream"], + // A sanitised colon: an underscore glued to a word, a space after it. + ["20220505 Part One_ The Beginning-Jkl012mno_6.webm", "Part One: The Beginning"], + // An underscore anywhere else stays. + ["snake_case and_more [12 views]-Mno345pqr_5.mp4", "snake_case and_more"], + // Not a calendar day (no 31st of April): kept. + ["20210431 Not a Day-Abc234def_5.mp4", "20210431 Not a Day"], + // Not a calendar date: kept. Brackets that are not a count: kept. + ["20221340 Numbers First [NEW] Thing-Pqr678stu_4.mp4", "20221340 Numbers First [NEW] Thing"], + // A dot inside the title survives; only the one extension goes. + ["20180101 Talk with A.B. Someone-Stu901vwx_3.mp4", "Talk with A.B. Someone"], + // A path inside the item: only its last part. + ["sub/dir/20180101 Inside a Folder-Vwx234yza_2.mp4", "Inside a Folder"], + // A name with no id at all is cleaned the same way. + ["20180101 Plain Recording.mp3", "Plain Recording"], + // Nothing sensible left. + ["20180101 [55 views]-Yza567bcd_1.mp4", null], + ["___.mp4", null], + ]; + for (const [name, want] of cases) assert.equal(titleFromMirrorFileName(name), want, name); +}); + +test("a mirrored file's name gives its upload day: the leading [word_]YYYYMMDD, a real calendar day only", () => { + const cases: [string, string | null][] = [ + ["20210102 Example Talk [1234 views]-AbC123xyz_9.mp4", "20210102"], + ["extras_20200304 A Long Chat _ Some Show-Def456uvw_8.mp4", "20200304"], + ["20191231 - New Year Stream [Ghi789rst_7].mkv", "20191231"], + ["sub/dir/20180101 Inside a Folder-Vwx234yza_2.mp4", "20180101"], + ["20180101.mp3", "20180101"], + // Leap days: only in a leap year. + ["20200229 Leap Day-Jkl012mno_6.mp4", "20200229"], + ["20190229 No Leap Day-Jkl012mno_6.mp4", null], + ["20210431 Not a Day-Abc234def_5.mp4", null], + ["20221340 Numbers First-Pqr678stu_4.mp4", null], + // Not at the start, or glued to the title: no date. + ["Talk from 20210102-Stu901vwx_3.mp4", null], + ["20210102Talk-Stu901vwx_3.mp4", null], + ["two_words_20210102 Talk-Stu901vwx_3.mp4", null], + ["Plain Recording.mp3", null], + ]; + for (const [name, want] of cases) assert.equal(dateFromMirrorFileName(name), want, name); +}); diff --git a/common/lib/archiveOrgId.ts b/common/lib/archiveOrgId.ts @@ -169,3 +169,62 @@ export function youtubeIdFromFileName(file: string): string | null { if (dashed && /[A-Z0-9_-]/.test(dashed[1])) return dashed[1]; return null; } + +// THE TITLE A MIRRORED FILE'S NAME CARRIES, for one file of a multi-file item +// that has no title of its own in the item. Archiving tools name a file +// `[<collection>_]YYYYMMDD <title> [<n> views]-<id>.<ext>`, and a file-name +// sanitiser writes a title's ` | ` as ` _ `. What is taken off: +// +// - the extension (one), and the YouTube id youtubeIdFromFileName finds, +// with its `-`/`_`/space or brackets; +// - a trailing ` [<n> views]` count; +// - a leading `YYYYMMDD` date, alone or after one `<word>_` prefix, when it +// is a calendar date (dateFromMirrorFileName reads it); the raw name stays +// in the provenance (`file`); +// +// and what is put back: ` _ ` → ` | `, and `word_ next` → `word: next` (an +// underscore glued to a word and followed by a space is a sanitised colon; an +// underscore anywhere else is left as it is). Spaces collapse. Null when no +// letter or digit remains. +const MIRROR_DATE_PREFIX = /^(?:[A-Za-z0-9]+_)?((?:19|20)\d{2})(\d{2})(\d{2})(?:\s+-\s+|\s+|_|$)/; + +// The leading `[<word>_]YYYYMMDD` of a file name's stem, when it is a real +// calendar day (no 31st of April, no 29th of February outside a leap year): +// the date as YYYYMMDD and the stem after it. Null otherwise. +function mirrorDatePrefix(stem: string): { date: string; rest: string } | null { + const m = MIRROR_DATE_PREFIX.exec(stem); + if (!m) return null; + const [y, mo, d] = [Number(m[1]), Number(m[2]), Number(m[3])]; + const day = new Date(Date.UTC(y, mo - 1, d)); + if (day.getUTCFullYear() !== y || day.getUTCMonth() !== mo - 1 || day.getUTCDate() !== d) return null; + return { date: `${m[1]}${m[2]}${m[3]}`, rest: stem.slice(m[0].length) }; +} + +function mirrorFileStem(file: string): string { + return (file.split("/").pop() ?? file).replace(/\.[A-Za-z0-9]{1,8}$/, ""); +} + +// THE DATE A MIRRORED FILE'S NAME CARRIES: the leading `[<word>_]YYYYMMDD` +// archiving tools write (the original's upload day), as YYYYMMDD, only for a +// real calendar day. Null when the name starts with no such date. +export function dateFromMirrorFileName(file: string): string | null { + return mirrorDatePrefix(mirrorFileStem(file))?.date ?? null; +} + +export function titleFromMirrorFileName(file: string): string | null { + const name = file.split("/").pop() ?? file; + let t = mirrorFileStem(name); + const id = youtubeIdFromFileName(name); + if (id) { + if (t.endsWith(`[${id}]`)) t = t.slice(0, -(id.length + 2)); + else if (t.endsWith(id)) t = t.slice(0, -id.length).replace(/[-_ ]$/, ""); + } + t = t.replace(/\s*\[\d[\d,]*\s+views?\]\s*$/i, ""); + t = mirrorDatePrefix(t)?.rest ?? t; + t = t + .replace(/\s+_\s+/g, " | ") + .replace(/(?<=[\p{L}\p{N})\]!?'’"])_(?=\s+\S)/gu, ":") + .replace(/\s+/g, " ") + .trim(); + return /[\p{L}\p{N}]/u.test(t) ? t : null; +} diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -6,7 +6,7 @@ - **archive.org files come over BitTorrent when possible, else straight from archive.org — never through yt-dlp.** The chosen file of an archive.org import is fetched from the item's own torrent (`<identifier>_archive.torrent`, which lists archive.org as a web seed, so other peers take load off archive.org) with aria2c, only that file of the item, and seeded afterwards for 10 minutes or to a ratio of 1, whichever comes first; the log shows "torrent: <file> (n of m pieces, peers p, web seed yes)" and "seeding 10 min…". With no aria2c, a torrent that does not carry the file, or no progress for 5 minutes, it is downloaded directly from `archive.org/download/…` instead (resumable, backing off on 429/503), and the log says "fell back to direct download: <reason>". Every file is checked against archive.org's sha1/md5: a mismatch is downloaded once more directly, a second one fails the record. The record is written from the item's metadata: `metadata.info.json` with the file's page, the canonical id, the duration ffprobe measures and archive.org's playable copies of the file, the `archiveorg.json` provenance (a mirror's original title, date and uploader), and `audio.<fmt>` — an audio file already in the channel's format is used as is, anything else goes through the app's audio extraction, a video kept in the saved-video store when the channel keeps sources. An .avi/.mpeg/.flac/.wav original is fetched as archive.org's mp4 or mp3 of it. aria2c runs in its own process group: cancelling the job stops it and everything it started, and it stops itself if the editor exits. New settings block `archiveOrg` (`torrent`, `seedMinutes`, `seedRatio`, `stallMinutes`, `maxPeers`, `maxDownloadKiBps`, `maxUploadKiBps`), `ARIA2C_BIN`, an aria2c row in `archilyzer doctor`, and `aria2` in the runtime Docker images. - **A forum thread can be archived as a posts source.** A XenForo thread URL (`…/threads/<title>.<id>/`; Kiwi Farms is recognised by host) makes a forum-thread channel — platform "xenforo", one channel per thread, each forum post a post — searchable and readable like X and Bluesky posts, in the editor, the export and the MCP (`get_thread` gives a forum post's conversation: the posts it quotes and the posts quoting it). **Fetch posts** reads the thread in a headless browser, newest page first, one page at a time with a 10–20 s pause (the channel key `postPagePauseSeconds` sets it), and stops at already-archived posts; a **Latest N pages** box (`archilyzer posts fetch --pages N`) caps a run, and the next run continues where it stopped. The browser keeps one profile per forum host, so a browser check it clears once (KiwiFlare's proof of work, say) stays cleared; a check that does not clear within a minute, a captcha, a login wall or a refusal stops the run with the reason and keeps its place — never retried at once. **Connect forum session** on the channel page opens that profile in a window on the editor's machine, at the thread, for the operator to clear it or log in. **Import saved pages** (`archilyzer posts import-html <slug> <file-or-dir>…`) reads thread pages saved from a browser ("Save page as", complete or HTML only) through the same parser: new posts are added and a post saved again after an edit is updated; a page saved from one of a forum's mirror domains (kiwifarms.net for a kiwifarms.st channel) is a page of the same thread. **Capture posts** works on forum posts: a screenshot of the post and its attached files, through the same profile. A forum post keeps its thread title, page, position, author id, last-edit time, quoted posts and its media links; quoted text is marked with "> " lines. - **A site's reports can be exported as files a reader saves and hosts again.** `archilyzer reports export <site> [--report <id>] [--formats html,pdf,md,zip]`, the `reports-export` job (`POST /api/ops/reports-export`, `pnpm ops reports-export`, and **Export reports** on a site's Reports tab) write each published report, checked as the build checks it, into `.export-index/sites/<site>/report-exports/<report>/`: `report.html`, one self-contained page (its own style, no script, stills and post screenshots inlined and recompressed, clips linked on the site); `report.pdf`, that page printed by headless Chromium, skipped with a note where there is none; `report.md`, plain Markdown with numbered references; and `evidence-pack.zip`, the page with its clips, stills and screenshots as files plus the Markdown and the citations, packed by the system `zip` (a host without it fails that format, naming it). An `export.json` names each file's size and checksum and the checksum of the report.json it was made from; every export ends with the report's date and the start of that checksum. Preparing the evidence media exports at its end when nothing is missing, on the same queue. The build publishes an export beside the report only when it was made from the report as it is now and is at most 24 MiB — a larger evidence pack stays local — and the Reports tab lists each report's exports, their sizes and which the next build publishes. The 24 MiB limit is one number, shared with the source mirror and the evidence clips. -- **archive.org items are a source (`platform: "archiveorg"`).** A channel can hold recordings imported from archive.org and transcribe them like any transcribe channel. **Import video** takes an item page (`https://archive.org/details/<identifier>`) when the item holds one media file, or ONE file of a multi-file item (`…/details/<identifier>/<file>`); an item with several media files is refused with the way to choose files. `pnpm ops import-archive-org --json '{"slug":…,"item":…,"files":[…]}'` (or `"match": "<regex>"`, `"dryRun": true`) imports chosen files of one item as one drainable job. A whole item's id is its identifier; a file's is `<identifier>__<slug>-<hash>`, stable and unique per file. Each record keeps an `archiveorg.json` sidecar — the item's title, date, creator and collections, its torrent, and for a mirror of a YouTube upload the original's id, URL, title and upload date read from the info.json uploaded beside it — and its metadata takes the file's own page and title (and a mirror's original title and date), recorded in the metadata history as `archiveorg-provenance`. The video page says "Archived on archive.org: <item> · torrent" and, for a mirror, "Originally on YouTube: <url> (uploaded <date>)". The channel form offers archive.org in both platform lists. +- **archive.org items are a source (`platform: "archiveorg"`).** A channel can hold recordings imported from archive.org and transcribe them like any transcribe channel. **Import video** takes an item page (`https://archive.org/details/<identifier>`) when the item holds one media file, or ONE file of a multi-file item (`…/details/<identifier>/<file>`); an item with several media files is refused with the way to choose files. `pnpm ops import-archive-org --json '{"slug":…,"item":…,"files":[…]}'` (or `"match": "<regex>"`, `"dryRun": true`) imports chosen files of one item as one drainable job. A whole item's id is its identifier; a file's is `<identifier>__<slug>-<hash>`, stable and unique per file. Each record keeps an `archiveorg.json` sidecar — the item's title, date, creator and collections, its torrent, and for a mirror of a YouTube upload the original's id, URL, title and upload date read from the info.json uploaded beside it — and its metadata takes the file's own page and title (and a mirror's original title and date), recorded in the metadata history as `archiveorg-provenance`. The video page says "Archived on archive.org: <item> · torrent" and, for a mirror, "Originally on YouTube: <url> (uploaded <date>)". The channel form offers archive.org in both platform lists. One file of a multi-file item with no uploaded info.json is dated by the `YYYYMMDD` its file name starts with (after any `<word>_`, a real calendar day only) and, when the item gives it no title of its own, titled from that name with the date, the `[<n> views]` count, the YouTube id and the extension taken off and ` _ ` read as ` | ` — the raw name stays in `archiveorg.json` as `file` — and `archilyzer archive-org refresh <slug> [--dry-run]` brings a channel's existing file records to the same title and date, offline, as `archiveorg-provenance` entries in the metadata history, printing old → new. - **Polite to archive.org.** archive.org runs on its own queue (`platform:archiveorg`), one download at a time, with a jittered pause of at least 8 s between files (the channel's or the global `sleepBetweenDownloadsSeconds` when longer). Its metadata API is asked once per item (cached for 6 h), with an identifying User-Agent, at most one request at a time and 2 s apart, honouring `Retry-After` and backing off exponentially on 429/503, stopping after four attempts. A file already downloaded is never fetched again, and a bulk import stops on a rate limit or after three failures in a row — re-running it resumes. - **A report video's cue lookup names a site that publishes only its reports.** Pointed at such a site (`corpus.json` `site.scope: "cited"`), a report-to-video manifest's cue lookup says the site publishes no transcripts and to use a full archive or a local corpus. - **A site's build composes its reports, and a site that publishes only its reports ships nothing else.** Every site's compose writes the reports its `site.json` publishes: each report's page and its citations as `citations.json` and `citations.csv` under `/reports/<id>/`, its cited stills, a page per cited moment with the record, the transcript lines around the span and every report that cites it, and the clips and post captures `archilyzer reports prepare` made for it, only the cited ones. Each quote is checked against the record as it is composed (a span's against its cues within 5 s either side, read from `en-orig` when the `en` track has no cues; a post's against its text) and the score, time and method are written into the citation, replacing any typed by hand. The build stops with the list of every problem before anything is written: an invalid report, a citation of a channel outside the site or of a post the site may not carry, a missing record, still or post, a quote that matches less than 60 % of what the record says, and a citation without prepared media or with media cut for another span (`--allow-missing-media` on `archilyzer compose site` and `build site` lets those two through, without a clip). A site with search off (`search: false`) removes everything corpus-shaped from `export/public` before it writes its reports, and its built `out/` is checked against what a cited site may hold: anything else, a file over 25 MiB or more than 20,000 files fails the build, and every deploy path (the Publish tab, `deploy site`, Build & deploy, Build & deploy all, the container build) refuses it, as it refuses a site with search off whose last build was a full one. The hub's compose removes a report site's files too.