Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 12353c3cc6b1c40d6bace8529aa5ec0e88d144a6
parent c4eea25ab4f57dde7730ab376313bea0bbb6e395
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Tue,  6 Oct 2026 09:29:51 -0400

Merge transcripts/en-track-fallback (one caption-track rule: en-orig first, empty tracks fall through; cue-block VTTs parse; one-shot index re-read)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

# Conflicts:
#	common/lib/sidecar-server.test.ts
#	editor/CHANGELOG.md
#	export/CHANGELOG.md

Diffstat:
MPUBLISH.md | 5+++--
MREADME.md | 5++++-
Mcommon/controller/buildIndex.ts | 117++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-----
Acommon/controller/buildIndexCaptionTrack.test.ts | 156+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/controller/channelSnapshot.ts | 13+++++++++----
Acommon/controller/normalizeCaptionTrack.test.ts | 160+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/controller/normalizeTranscript.ts | 143++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++---------------
Acommon/lib/__fixtures__/vtt-cue-blocks.vtt | 29+++++++++++++++++++++++++++++
Acommon/lib/__fixtures__/vtt-rolling.vtt | 31+++++++++++++++++++++++++++++++
Acommon/lib/captionTrack.test.ts | 117+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/lib/sidecar-server.test.ts | 2++
Mcommon/lib/subtitleProvenance.ts | 25+++++++++++++++++++++++--
Acommon/lib/transcriptPin-server.ts | 39+++++++++++++++++++++++++++++++++++++++
Mcommon/lib/videoStatus.ts | 133++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++---------------
Mcommon/lib/vtt.test.ts | 60+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-
Mcommon/lib/vtt.ts | 39+++++++++++++++++++++++++++++++++++----
Mcommon/publish/composeReports.ts | 31++++---------------------------
Meditor/CHANGELOG.md | 3+++
Meditor/app/channels/[slug]/videos/[id]/components/cards/TranscriptSourceSection.tsx | 11++++++-----
Meditor/app/channels/[slug]/videos/[id]/components/cards/videoFiles.ts | 2++
Meditor/app/channels/[slug]/videos/[id]/videoActions.ts | 9++++++---
Mexport/CHANGELOG.md | 1+
Mumtool/report-to-video/cues.mjs | 44+++++++++++++++++++++++++++++++++++++++++++-
Mumtool/report-to-video/cues.test.mjs | 53+++++++++++++++++++++++++++++++++++++++++++++++++++++
24 files changed, 1121 insertions(+), 107 deletions(-)

diff --git a/PUBLISH.md b/PUBLISH.md @@ -116,8 +116,9 @@ mirror is (a fresh bare clone, `repack -a -d`, `update-server-info`, then only ` `packed-refs`, `info/refs`, `objects/info/packs`, the packs and `refs/heads/main`), so `git clone <siteUrl>/reports/<id>/history/repo` works. The deploy audit refuses any other file in a clone or one over 24 MiB, and counts the clone's files toward Pages' 20,000. Every quote is checked as it is -composed: a span's against its cues within 5 s either side (a record whose `en` track -has no cues is read from `en-orig`), a post's against its text, scored as the share of +composed: a span's against its cues within 5 s either side (read under the caption-track +rule — `en-orig` first, then the first English track with cues — and against every +other English track, the best match shown), a post's against its text, scored as the share of the quote's words found there; the score, the time and the method are written into the citation, replacing any typed by hand. Compose fails, before it writes any of it, with the list of everything wrong: an invalid report, a citation of a channel outside the diff --git a/README.md b/README.md @@ -382,7 +382,10 @@ differs is emphasis: - **Captions you already publish are reused.** Channels whose platform publishes captions need no transcription backend at all; those are taken directly. There is also an opt-in lane that *replaces* platform auto-captions with your own transcripts - where you would rather have the better text. + where you would rather have the better text. Where YouTube serves both, the + original-audio captions (`en-orig`) are read before the served `en` track, which can + reword what was said; a track with no text falls through to the next, and a video's + page can pin another (`common/lib/videoStatus.ts`, "the caption-track rule"). - **The output is yours to brand.** Site title, header, description, tagline, social links and channel grouping are all per-site configuration. - **One corpus can publish several sites.** Channels are grouped into sites, so a diff --git a/common/controller/buildIndex.ts b/common/controller/buildIndex.ts @@ -107,7 +107,11 @@ import { import { resolveChannelGroupId } from "../lib/channelGroups"; import type { Paths } from "../lib/paths"; import { + CAPTION_TRACK_RULE_VERSION, + captionInputs, + englishVttsByPreference, pickIndexTranscript, + readEnglishVttCues, readSubTracks, readVideoFiles, type IndexTranscript, @@ -205,6 +209,39 @@ const SCHEMA_VERSION = 13; const PLATFORM_LABELS_VERSION = 1; const PLATFORM_LABELS_KEY = "platformLabels"; +// CAPTION TRACK — the same one-shot shape, for the caption-track rule +// (CAPTION_TRACK_RULE_VERSION, lib/videoStatus.ts). When the rule changes, the +// first build that sees the new version re-reads, from disk, every caption +// record the change can reach — one with more than one English VTT (the rule +// may now pick another), or whose stored cues are empty (the cue-block shape +// used to parse to nothing) — and nothing else; then records the version, +// again only when no channel is held. Each re-read is reported per channel: +// how many now read different text, how many had none and now do. +const CAPTION_TRACK_KEY = "captionTrackRule"; + +// A cue list's text, for "did the words change" — timing alone is not a +// different transcript. +function cueText(list: readonly Cue[]): string { + return list.map((c) => c.text).join("\n"); +} + +// Whether the cue list stored under a key is empty or absent, WITHOUT decoding +// it: a non-empty list of cues is far longer than the few bytes an empty +// msgpack array takes, and decoding every caption record of the corpus to ask +// this would cost the full read the pass exists to avoid. +function storedCuesEmpty(db: { getBinaryFast(key: IndexKey): Buffer | undefined }, key: IndexKey): boolean { + const raw = db.getBinaryFast(key); + return raw === undefined || raw.length <= 4; +} + +export type CaptionTrackChannelReport = { + reread: number; + // Re-read records whose caption text is not what the index held. + changed: number; + // Of those, the ones the index held no text for. + zeroToText: number; +}; + // Per-channel post stats, persisted so per-site aggregates survive a no-op // rebuild that doesn't re-encode the post pages. Mirrors ChannelSubsStat. type ChannelPostsStat = { @@ -299,6 +336,8 @@ type LiveEntry = { transcriptPath: string; transcriptMs: number | null; transcriptKind: IndexTranscript["kind"] | null; + // How many English VTTs the dir holds (the caption-track pass's question). + englishVttCount: number; subTracks: SubTrack[]; subsMs: number | null; availabilityMs: number | null; @@ -429,10 +468,17 @@ async function scanSource( : path.join(fullVideoDir, "transcript.json"); let transcriptMs: number | null = null; if (picked) { - try { - transcriptMs = (await stat(transcriptPath)).mtimeMs; - } catch { - transcriptMs = null; + // Captions: the newest of every caption input (each English VTT and + // the operator's pin), since the cues may come from any of them. + const inputs = + picked.kind === "vtt" ? captionInputs(files.entries) : [picked.filename]; + for (const name of inputs) { + try { + const ms = (await stat(path.join(fullVideoDir, name))).mtimeMs; + if (transcriptMs === null || ms > transcriptMs) transcriptMs = ms; + } catch { + // ignore + } } } const subTracks = await readSubTracks(fullVideoDir); @@ -482,6 +528,7 @@ async function scanSource( transcriptPath, transcriptMs, transcriptKind: picked?.kind ?? null, + englishVttCount: englishVttsByPreference(files.entries).length, subTracks, subsMs, availabilityMs, @@ -542,6 +589,9 @@ export type BuildIndexResult = { // Channels whose media could not be read, so they were not rescanned: their // index records and shared pages were kept as they were (see the header). heldChannels: string[]; + // Per channel, what a caption-track pass re-read and changed; empty when no + // pass ran (CAPTION_TRACK_KEY). + captionTrack: Record<string, CaptionTrackChannelReport>; }; export type BuildIndexOptions = { @@ -803,6 +853,31 @@ export async function buildIndex({ ); } + // Caption records the caption-track rule may now read differently + // (CAPTION_TRACK_KEY). A schema bump re-reads everything anyway. + const captionTrackDue = + meta.get(CAPTION_TRACK_KEY) !== CAPTION_TRACK_RULE_VERSION; + // pathKeyIds of the records this pass re-reads. + const captionReread = new Set<string>(); + if (captionTrackDue && !schemaBumped) { + const queued = new Set( + [...added, ...changed].map((s) => pathKeyId([s.channelSlug, s.videoDir])), + ); + for (const s of live) { + if (s.transcriptKind !== "vtt") continue; + const pk: PathKey = [s.channelSlug, s.videoDir]; + const prev = mtimes.get(pk); + if (!prev) continue; + if (s.englishVttCount < 2 && !storedCuesEmpty(cues, prev.indexKey)) continue; + captionReread.add(pathKeyId(pk)); + if (!queued.has(pathKeyId(pk))) changed.push(s); + } + log( + `Caption track v${CAPTION_TRACK_RULE_VERSION}: ${captionReread.size} record(s) re-read.`, + ); + } + const captionReport = new Map<string, CaptionTrackChannelReport>(); + const anyMutations = added.length > 0 || changed.length > 0 || removed.length > 0; @@ -901,11 +976,11 @@ export async function buildIndex({ ); if (cueList === undefined && s.transcriptMs !== null && s.transcriptKind) { try { - const raw = await readFile(s.transcriptPath, "utf8"); cueList = s.transcriptKind === "vtt" - ? parseVtt(raw) - : parseTranscriptJson(raw); + ? // The caption-track rule (lib/videoStatus.ts). + (await readEnglishVttCues(videoFullDir))?.cues + : parseTranscriptJson(await readFile(s.transcriptPath, "utf8")); } catch { cueList = undefined; } @@ -923,6 +998,12 @@ export async function buildIndex({ const pk: PathKey = [s.channelSlug, s.videoDir]; const prev = mtimes.get(pk); + // What the last build held for this video's captions, for the + // caption-track report — read before the put below replaces it. + const heldBefore = + prev && captionReread.has(pathKeyId(pk)) + ? cueText(cues.get(prev.indexKey) ?? []) + : undefined; // What the last build held for this video's sub tracks, read before // a re-keyed record is removed: a tiered live chat whose raw cannot // be read now keeps the cues it had (below). @@ -938,6 +1019,17 @@ export async function buildIndex({ if (cueList) cues.put(indexKey, cueList); else cues.remove(indexKey); + if (heldBefore !== undefined) { + const r = captionReport.get(s.channelSlug) ?? { reread: 0, changed: 0, zeroToText: 0 }; + r.reread++; + const now = cueText(cueList ?? []); + if (heldBefore !== now) { + r.changed++; + if (heldBefore === "" && now !== "") r.zeroToText++; + } + captionReport.set(s.channelSlug, r); + } + const parsedSubs: StoredSubs = []; // A tiered track that could be neither read nor kept: `subsMs` is // stored null so the next build's scan sees a change and retries. @@ -1112,6 +1204,13 @@ export async function buildIndex({ } } + // What the caption-track pass changed, per channel (CAPTION_TRACK_KEY). + for (const [slug, r] of [...captionReport].sort(([a], [b]) => (a < b ? -1 : a > b ? 1 : 0))) { + log( + `Caption track v${CAPTION_TRACK_RULE_VERSION}: ${slug}: ${r.reread} re-read, ${r.changed} now read different text, ${r.zeroToText} had no text and now do.`, + ); + } + for (const { pathKey, indexKey } of removed) { sums.remove(indexKey); cues.remove(indexKey); @@ -2332,6 +2431,9 @@ export async function buildIndex({ if (relabelDue && held.size === 0) { await meta.put(PLATFORM_LABELS_KEY, PLATFORM_LABELS_VERSION); } + if (captionTrackDue && held.size === 0) { + await meta.put(CAPTION_TRACK_KEY, CAPTION_TRACK_RULE_VERSION); + } await meta.flushed; await root.close(); @@ -2356,5 +2458,6 @@ export async function buildIndex({ removed: removed.length, shortCircuited: !sharedNeedsBuild && sitesBuilt === 0, heldChannels: [...held.keys()], + captionTrack: Object.fromEntries(captionReport), }; } diff --git a/common/controller/buildIndexCaptionTrack.test.ts b/common/controller/buildIndexCaptionTrack.test.ts @@ -0,0 +1,156 @@ +// Integration: the caption-track pass (CAPTION_TRACK_RULE_VERSION) re-reads, +// through the REAL buildIndex over a temp corpus, exactly the records an older +// caption-track rule may have read differently — one with two English VTTs, +// one whose stored cues are empty — and leaves every other record alone. +// +// Run with: node_modules/.bin/tsx --test common/controller/buildIndexCaptionTrack.test.ts + +import { after, test } from "node:test"; +import assert from "node:assert/strict"; +import { mkdirSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import path from "node:path"; + +const ROOT = mkdtempSync(path.join(tmpdir(), "build-index-captions-")); +const PINNED: Record<string, string> = { + TRANSCRIPTS_DIR: path.join(ROOT, "transcripts"), + SAVED_VIDEOS_DIR: path.join(ROOT, "saved-videos"), + SITES_DIR: path.join(ROOT, "transcripts", "sites"), + SETTINGS_FILE: path.join(ROOT, "settings.json"), + EXPORT_PUBLIC_DIR: path.join(ROOT, "public"), + EXPORT_INDEX_DIR: path.join(ROOT, ".export-index"), + EXPORT_BUILDS_DIR: path.join(ROOT, ".export-builds"), + EDITOR_CHANGELOG_FILE: path.join(ROOT, "editor-CHANGELOG.md"), + EXPORT_CHANGELOG_FILE: path.join(ROOT, "export-CHANGELOG.md"), + CHARTS_CONFIG_FILE: path.join(ROOT, "chart-templates.json"), + SEARCH_ALIASES_FILE: path.join(ROOT, "transcripts", "search-aliases.json"), + CURATED_TAGS_FILE: path.join(ROOT, "transcripts", "tags.json"), + ARCHILYZER_CONFIG_DIR: path.join(ROOT, "config"), + ARCHILYZER_SOURCE_SCRATCH: path.join(ROOT, "source-scratch"), +}; +Object.assign(process.env, PINNED); +delete process.env.ARCHILYZER_INDEX_ALLOW_HELD; +after(() => rmSync(ROOT, { recursive: true, force: true })); + +const { getPaths } = await import("../lib/paths"); +const { buildIndex } = await import("./buildIndex"); +const { open } = await import("lmdb"); + +const paths = getPaths(); +const CHANNEL = "example-channel"; +const SITE = "testsite"; +const BOTH = "BothTracks01"; // served en (cue blocks) + en-orig (rolling) +const BLOCKS = "CueBlocks001"; // a lone cue-block en +const ROLL = "RollingOnly1"; // a lone rolling en — nothing for the pass to do + +const fixture = (name: string) => + readFileSync(path.join(import.meta.dirname, "..", "lib", "__fixtures__", name), "utf8"); +const ROLLING = fixture("vtt-rolling.vtt"); +const CUE_BLOCKS = fixture("vtt-cue-blocks.vtt"); + +const writeJson = (file: string, value: unknown) => { + mkdirSync(path.dirname(file), { recursive: true }); + writeFileSync(file, JSON.stringify(value, null, 2)); +}; +const dirOf = (id: string) => path.join(paths.channelsDir, CHANNEL, "data", id); + +function seed(): void { + rmSync(paths.transcriptsDir, { recursive: true, force: true }); + rmSync(PINNED.EXPORT_INDEX_DIR, { recursive: true, force: true }); + writeFileSync(paths.settingsFile, "{}"); + writeJson(path.join(paths.channelsDir, CHANNEL, "config.json"), { + handling: "youtube", + name: CHANNEL, + }); + writeJson(path.join(paths.sitesDir, SITE, "site.json"), { + siteId: SITE, + siteTitle: "Test Site", + siteDescription: "fixture", + headerTitle: "Test Site", + homeTagline: "", + socialLinks: [], + groups: [{ id: "default", name: "All channels", selectedByDefault: true }], + defaultGroupId: "default", + channels: [{ slug: CHANNEL, groupId: "default" }], + }); + for (const [i, id] of [BOTH, BLOCKS, ROLL].entries()) { + writeJson(path.join(dirOf(id), "metadata.info.json"), { + id, + title: `Video ${id}`, + upload_date: `2026060${i + 1}`, + duration: 30, + webpage_url: `https://www.youtube.com/watch?v=${id}`, + extractor_key: "Youtube", + }); + } + writeFileSync(path.join(dirOf(BOTH), "transcript.en.vtt"), CUE_BLOCKS); + writeFileSync(path.join(dirOf(BOTH), "transcript.en-orig.vtt"), ROLLING); + writeFileSync(path.join(dirOf(BLOCKS), "transcript.en.vtt"), CUE_BLOCKS); + writeFileSync(path.join(dirOf(ROLL), "transcript.en.vtt"), ROLLING); +} + +type Summary = { id: string }; +type Cue = { start: number; end: number; text: string }; + +function withIndex<T>(fn: (db: (name: string) => ReturnType<ReturnType<typeof open>["openDB"]>) => T): T { + const root = open({ path: paths.lmdbPath, maxDbs: 18, compression: true }); + try { + return fn((name) => root.openDB({ name, encoding: "msgpack" })); + } finally { + root.close(); + } +} +const keyOf = (db: (name: string) => ReturnType<ReturnType<typeof open>["openDB"]>, id: string) => { + for (const { key, value } of db("sums").getRange()) { + if ((value as Summary).id === id) return key; + } + throw new Error(`${id} is not indexed`); +}; +const cuesOf = (id: string) => withIndex((db) => db("cues").get(keyOf(db, id)) as Cue[] | undefined); + +async function runIndex(): Promise<string[]> { + const log: string[] = []; + await buildIndex({ paths, onLog: (s) => log.push(s) }); + return log; +} + +test("a fresh index reads en-orig beside a served en, and a lone cue-block en as text", async () => { + seed(); + const log = await runIndex(); + assert.equal(cuesOf(BOTH)?.[0].text, "are talking about the harbor"); + assert.equal(cuesOf(BLOCKS)?.length, 7); + assert.equal(cuesOf(ROLL)?.length, 3); + // A first build has nothing indexed under an older rule: no pass. + assert.equal(log.some((l) => /Caption track/.test(l)), false, log.join("\n")); +}); + +test("an index built under the old rule is re-read once, for exactly the records the rule reaches", async () => { + // The index as an older build left it: the served en's text for BOTH, no + // cues for BLOCKS, and no caption-track version recorded. + const heldRoll = cuesOf(ROLL); + withIndex((db) => { + const cues = db("cues"); + cues.putSync(keyOf(db, BOTH), [{ start: 0, end: 6, text: "served words" }]); + cues.putSync(keyOf(db, BLOCKS), []); + // ROLL's record is marked so a re-read would show. + cues.putSync(keyOf(db, ROLL), [{ start: 0, end: 1, text: "untouched" }]); + db("meta").removeSync("captionTrackRule"); + }); + + const log = await runIndex(); + assert.ok(log.includes("Caption track v1: 2 record(s) re-read."), log.join("\n")); + assert.ok( + log.includes( + `Caption track v1: ${CHANNEL}: 2 re-read, 2 now read different text, 1 had no text and now do.`, + ), + log.join("\n"), + ); + assert.equal(cuesOf(BOTH)?.[0].text, "are talking about the harbor"); + assert.equal(cuesOf(BLOCKS)?.length, 7); + assert.deepEqual(cuesOf(ROLL), [{ start: 0, end: 1, text: "untouched" }]); + assert.notDeepEqual(heldRoll, cuesOf(ROLL)); + + // Recorded: the next build does not look again. + const again = await runIndex(); + assert.equal(again.some((l) => /Caption track/.test(l)), false, again.join("\n")); +}); diff --git a/common/controller/channelSnapshot.ts b/common/controller/channelSnapshot.ts @@ -10,6 +10,7 @@ import { isVideoTranscribed, readVideoFiles, LIVE_CHAT_FILENAME, + ORIG_VTT_FILENAME, VTT_FILENAME, type VideoFiles, } from "../lib/videoStatus"; @@ -74,7 +75,7 @@ import { loadMaybeMissing } from "./quickAvailabilityCheck"; import { loadRoster } from "./rosterStore"; import { deriveChannelSets } from "./channelSets"; import { readTranscriptCoverage } from "./normalizeTranscript"; -import { resolveVttProvenance } from "../lib/subtitleProvenance"; +import { resolveCaptionsProvenance } from "../lib/subtitleProvenance"; import { isIncompleteTranscript } from "../lib/transcriptCoverage"; // THE PER-OPERATION WORK LISTS, keyed by operation id. @@ -1007,8 +1008,10 @@ export async function generateChannelSnapshot( // both the auto-subs work lane (no whisper yet) and the superseded // backup inventory (whisper already won) need it. Same conditional // per-video sidecar read pattern as the cues.json coverage read above. + // Over every English VTT (resolveCaptionsProvenance): the rule reads + // en-orig even beside a human `en`, which is not ASR-only. const vttProvenance = files.ytVttFile - ? await resolveVttProvenance(dir, files.ytVttFile) + ? await resolveCaptionsProvenance(dir, files.entries) : null; // Only transcribed videos can carry a digest, so everything else skips // the sidecar read entirely — the same conditional per-video @@ -1414,11 +1417,13 @@ export async function generateChannelSnapshot( // A transcript that exists only under a non-standard VTT name — either a // regional/auto English track (transcript.en-US.vtt) now picked up by the // fallback, or a foreign-only transcript.<lang>.vtt that isn't recognized as - // English. Whisper transcripts and the canonical transcript.en.vtt are fine. + // English. Whisper transcripts, the canonical transcript.en.vtt and the + // original-audio transcript.en-orig.vtt (the rule's first pick) are fine. if ( !files.hasWhisper && files.hasNonCanonicalVtt && - files.ytVttFile !== VTT_FILENAME + files.ytVttFile !== VTT_FILENAME && + files.ytVttFile !== ORIG_VTT_FILENAME ) { nonStandardVtt.push(id); } diff --git a/common/controller/normalizeCaptionTrack.test.ts b/common/controller/normalizeCaptionTrack.test.ts @@ -0,0 +1,160 @@ +// transcript.cues.json under the caption-track rule: normalize records the +// track it read and the rule that chose it, and a cues.json made under an older +// rule is stale exactly where the rule could now read something else. +// +// Run with: node_modules/.bin/tsx --test common/controller/normalizeCaptionTrack.test.ts + +import { after, test } from "node:test"; +import assert from "node:assert/strict"; +import { mkdirSync, mkdtempSync, readFileSync, rmSync, utimesSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import path from "node:path"; +import { isCuesJsonFresh, normalizeTranscript, readNormalizedTranscript } from "./normalizeTranscript"; +import { CAPTION_TRACK_RULE_VERSION, TRANSCRIPT_PIN_FILENAME } from "../lib/videoStatus"; + +const ROOT = mkdtempSync(path.join(tmpdir(), "normalize-caption-")); +after(() => rmSync(ROOT, { recursive: true, force: true })); + +const fixture = (name: string) => + readFileSync(path.join(import.meta.dirname, "..", "lib", "__fixtures__", name), "utf8"); +const ROLLING = fixture("vtt-rolling.vtt"); +const CUE_BLOCKS = fixture("vtt-cue-blocks.vtt"); + +const META = { + id: "vid", + title: "A video", + upload_date: "20260601", + duration: 30, + webpage_url: "https://www.youtube.com/watch?v=vid", + extractor_key: "Youtube", +}; + +let n = 0; +function videoDir(files: Record<string, string>): string { + const dir = path.join(ROOT, `c${n}`, "data", `vid${n++}`); + mkdirSync(dir, { recursive: true }); + writeFileSync(path.join(dir, "metadata.info.json"), JSON.stringify(META)); + for (const [name, body] of Object.entries(files)) writeFileSync(path.join(dir, name), body); + return dir; +} + +// A cues.json as the app wrote it before the rule had a version: no track +// recorded, newer than every input. +function oldCuesJson(dir: string, cues: { start: number; end: number; text: string }[]): void { + const p = path.join(dir, "transcript.cues.json"); + writeFileSync( + p, + JSON.stringify({ version: 2, source: "vtt", transcriptFormat: "vtt", id: "vid", slug: "c/vid", cues }), + ); + const later = new Date(Date.now() + 10_000); + utimesSync(p, later, later); +} + +test("normalize reads en-orig, and records the track and the rule in the file's first bytes", async () => { + const dir = videoDir({ "transcript.en.vtt": CUE_BLOCKS, "transcript.en-orig.vtt": ROLLING }); + const out = await normalizeTranscript({ videoDir: dir, channelSlug: "c" }); + assert.equal(out.status, "wrote"); + const raw = readFileSync(path.join(dir, "transcript.cues.json"), "utf8"); + assert.match(raw.slice(0, 200), /"vttFile":"transcript\.en-orig\.vtt","captionTrackRule":\d+/); + const n = await readNormalizedTranscript(path.join(dir, "transcript.cues.json")); + assert.equal(n?.vttFile, "transcript.en-orig.vtt"); + assert.equal(n?.captionTrackRule, CAPTION_TRACK_RULE_VERSION); + assert.equal(n?.cues?.[0].text, "are talking about the harbor"); + assert.equal((await isCuesJsonFresh(dir)).fresh, true); + assert.equal((await normalizeTranscript({ videoDir: dir, channelSlug: "c" })).status, "fresh"); +}); + +test("a cues.json from before the rule is fresh beside a byte-identical en/en-orig pair: it reads the same", async () => { + const dir = videoDir({ "transcript.en.vtt": ROLLING, "transcript.en-orig.vtt": ROLLING }); + oldCuesJson(dir, [{ start: 0, end: 1, text: "same words" }]); + assert.equal((await isCuesJsonFresh(dir)).fresh, true); + assert.equal((await normalizeTranscript({ videoDir: dir, channelSlug: "c" })).status, "fresh"); +}); + +test("a cues.json from before the rule is fresh beside two tracks the older rule ranked the same way", async () => { + const dir = videoDir({ "transcript.en.vtt": CUE_BLOCKS, "transcript.en-US.vtt": ROLLING }); + oldCuesJson(dir, [{ start: 0, end: 1, text: "kept" }]); + assert.equal((await isCuesJsonFresh(dir)).fresh, true); +}); + +test("a cues.json from before the rule with no cues is stale beside two tracks: the fallback may find text", async () => { + const dir = videoDir({ "transcript.en.vtt": ROLLING, "transcript.en-orig.vtt": ROLLING }); + oldCuesJson(dir, []); + assert.equal((await isCuesJsonFresh(dir)).fresh, false); +}); + +test("a cues.json from before the rule is stale beside an en and en-orig that differ, and normalize rewrites it", async () => { + const dir = videoDir({ "transcript.en.vtt": CUE_BLOCKS, "transcript.en-orig.vtt": ROLLING }); + oldCuesJson(dir, [{ start: 0, end: 1, text: "served words" }]); + assert.deepEqual(await isCuesJsonFresh(dir), { + fresh: false, + reason: "stale", + cuesPath: path.join(dir, "transcript.cues.json"), + }); + assert.equal((await normalizeTranscript({ videoDir: dir, channelSlug: "c" })).status, "wrote"); + assert.equal((await isCuesJsonFresh(dir)).fresh, true); +}); + +test("a cues.json from before the rule is stale over a lone cue-block track (it parsed to nothing)", async () => { + const dir = videoDir({ "transcript.en.vtt": CUE_BLOCKS }); + oldCuesJson(dir, []); + assert.equal((await isCuesJsonFresh(dir)).fresh, false); + await normalizeTranscript({ videoDir: dir, channelSlug: "c" }); + const n = await readNormalizedTranscript(path.join(dir, "transcript.cues.json")); + assert.equal(n?.cues?.length, 7); + assert.equal((await isCuesJsonFresh(dir)).fresh, true); +}); + +test("a cues.json from before the rule that holds cues stays fresh over a lone cue-block track: the old parse made none", async () => { + const dir = videoDir({ "transcript.en.vtt": CUE_BLOCKS }); + oldCuesJson(dir, [{ start: 0, end: 1, text: "written by hand" }]); + assert.equal((await isCuesJsonFresh(dir)).fresh, true); +}); + +test("a cues.json from before the rule stays fresh over a lone rolling track: nothing to choose, same parse", async () => { + const dir = videoDir({ "transcript.en.vtt": ROLLING }); + oldCuesJson(dir, [{ start: 0, end: 1, text: "kept" }]); + assert.equal((await isCuesJsonFresh(dir)).fresh, true); + assert.equal((await normalizeTranscript({ videoDir: dir, channelSlug: "c" })).status, "fresh"); +}); + +test("a newer secondary track, or a new pin, makes cues.json stale", async () => { + const dir = videoDir({ "transcript.en.vtt": CUE_BLOCKS, "transcript.en-orig.vtt": ROLLING }); + const at = (name: string, s: number) => { + const t = new Date(Date.now() + s * 1000); + utimesSync(path.join(dir, name), t, t); + }; + await normalizeTranscript({ videoDir: dir, channelSlug: "c" }); + at("transcript.cues.json", 10); + assert.equal((await isCuesJsonFresh(dir)).fresh, true); + // Not the track the rule reads — still an input (the fallback may read it). + at("transcript.en.vtt", 20); + assert.equal((await isCuesJsonFresh(dir)).reason, "stale"); + await normalizeTranscript({ videoDir: dir, channelSlug: "c" }); + at("transcript.cues.json", 30); + assert.equal((await isCuesJsonFresh(dir)).fresh, true); + + // The operator pins transcript.en.vtt: stale, and the rewrite reads en. + writeFileSync( + path.join(dir, TRANSCRIPT_PIN_FILENAME), + JSON.stringify({ from: "transcript.en.vtt", pinnedAt: "" }), + ); + at(TRANSCRIPT_PIN_FILENAME, 40); + assert.equal((await isCuesJsonFresh(dir)).fresh, false); + await normalizeTranscript({ videoDir: dir, channelSlug: "c" }); + const n = await readNormalizedTranscript(path.join(dir, "transcript.cues.json")); + assert.equal(n?.vttFile, "transcript.en.vtt"); +}); + +test("whisper wins: a whisper cues.json is never asked about caption tracks", async () => { + const dir = videoDir({ + "transcript.en.vtt": CUE_BLOCKS, + "transcript.en-orig.vtt": ROLLING, + "transcript.json": JSON.stringify({ transcription: [{ offsets: { from: 0, to: 1000 }, text: "spoken" }] }), + }); + const p = path.join(dir, "transcript.cues.json"); + writeFileSync(p, JSON.stringify({ version: 2, source: "whisper", cues: [{ start: 0, end: 1, text: "spoken" }] })); + const later = new Date(Date.now() + 10_000); + utimesSync(p, later, later); + assert.equal((await isCuesJsonFresh(dir)).fresh, true); +}); diff --git a/common/controller/normalizeTranscript.ts b/common/controller/normalizeTranscript.ts @@ -4,9 +4,9 @@ // shape that buildIndex would emit, plus a `source` marker. import path from "node:path"; -import { readdir, readFile, stat } from "node:fs/promises"; +import { open, readdir, readFile, stat } from "node:fs/promises"; import { writeJsonAtomic } from "../lib/jsonFile-server"; -import { parseVtt, type Cue } from "../lib/vtt"; +import { hasWordTiming, type Cue } from "../lib/vtt"; import { detectTranscriptFormat, parseTranscriptJson, @@ -19,13 +19,17 @@ import { type TranscriptCoverage, } from "../lib/transcriptCoverage"; import { + CAPTION_TRACK_RULE_VERSION, CUES_JSON_FILENAME, META_FILENAME, + ORIG_VTT_FILENAME, VTT_FILENAME, WHISPER_FILENAME, + captionInputs, + englishVttsByPreference, pickIndexTranscript, + readEnglishVttCues, readVideoFiles, - resolvePrimaryVtt, type IndexTranscript, } from "../lib/videoStatus"; @@ -36,6 +40,12 @@ export const CUES_FILE_VERSION = 2; export type NormalizedTranscript = TranscriptDetail & { version: number; source: IndexTranscript["kind"] | "live_chat"; + // source "vtt" only: the English VTT the cues were read from, and the + // caption-track rule that chose it (CAPTION_TRACK_RULE_VERSION). Written + // right after `source`, so they sit in the file's first bytes and + // isCuesJsonFresh can read them without parsing the cues. + vttFile?: string; + captionTrackRule?: number; // The raw transcript format this was parsed from. Durable per-video record so // a later re-normalize knows how to read the raw file without re-sniffing. transcriptFormat?: TranscriptOutputFormat; @@ -76,24 +86,14 @@ export async function normalizeTranscript( const picked = pickIndexTranscript(files); if (!picked) return { status: "skipped", reason: "no-raw-transcript" }; - const cuesPath = path.join(opts.videoDir, CUES_JSON_FILENAME); const metaPath = path.join(opts.videoDir, META_FILENAME); const transcriptPath = path.join(opts.videoDir, picked.filename); - const [cuesMs, metaStatMs, transcriptStatMs] = await Promise.all([ - mtimeMs(cuesPath), - mtimeMs(metaPath), - mtimeMs(transcriptPath), - ]); - - if ( - !opts.force && - cuesMs !== null && - metaStatMs !== null && - transcriptStatMs !== null && - cuesMs >= metaStatMs && - cuesMs >= transcriptStatMs - ) { + // The same question every reader of cues.json asks (isCuesJsonFresh), so a + // normalize pass rewrites exactly the files they refuse. + const freshness = await isCuesJsonFresh(opts.videoDir); + const cuesPath = freshness.cuesPath; + if (!opts.force && freshness.fresh) { return { status: "fresh", cuesPath }; } @@ -106,14 +106,20 @@ export async function normalizeTranscript( opts.configName, ); - const rawTranscript = await readFile(transcriptPath, "utf8"); let cues: Cue[]; let transcriptFormat: TranscriptOutputFormat; + let vttFile: string | undefined; try { if (picked.kind === "vtt") { - cues = parseVtt(rawTranscript); + // The caption-track rule: the first English VTT, in preference order, + // with cues (lib/videoStatus.ts). + const read = await readEnglishVttCues(opts.videoDir, files.entries); + if (!read) throw new Error("no English VTT could be read"); + cues = read.cues; + vttFile = read.filename; transcriptFormat = "vtt"; } else { + const rawTranscript = await readFile(transcriptPath, "utf8"); // Resolve the JSON format: authoritative hint -> recorded per-video tag // -> content sniff -> whisper fallback (the only app before chough). transcriptFormat = @@ -132,6 +138,9 @@ export async function normalizeTranscript( const out: NormalizedTranscript = { version: CUES_FILE_VERSION, source: picked.kind, + ...(vttFile !== undefined + ? { vttFile, captionTrackRule: CAPTION_TRACK_RULE_VERSION } + : {}), transcriptFormat, ...summary, cues, @@ -140,7 +149,7 @@ export async function normalizeTranscript( // Compact, no trailing newline: transcript.cues.json's historical bytes. await writeJsonAtomic(cuesPath, out, { indent: 0, newline: false }); opts.log?.( - `Normalized ${opts.channelSlug}/${path.basename(opts.videoDir)} (${transcriptFormat}, ${cues.length} cues)`, + `Normalized ${opts.channelSlug}/${path.basename(opts.videoDir)} (${vttFile ?? transcriptFormat}, ${cues.length} cues)`, ); return { status: "wrote", cuesPath }; } @@ -214,22 +223,33 @@ export type CuesFreshReason = // Helper: given a video dir, decide whether transcript.cues.json (if present) // is at least as new as metadata.info.json and the raw transcript file. Used // by buildIndex to know whether it can trust cues.json without re-parsing. +// +// CAPTIONS: the raw transcript is EVERY caption input (captionInputs — each +// English VTT, since the content fallback may read any of them, and the +// operator's pin), and a cues.json that is new enough must also have been made +// under the current caption-track rule. Made under an older one, it is `stale` +// where the rule could read it differently: its cues may come from a track the +// rule no longer picks (a served `en` that differs from the `en-orig` beside +// it), or be the zero cues a cue-block VTT used to parse to +// (followsCaptionTrackRule). Telling costs a few small reads, never a parse. export async function isCuesJsonFresh( videoDir: string, ): Promise<{ fresh: boolean; reason: CuesFreshReason; cuesPath: string }> { const cuesPath = path.join(videoDir, CUES_JSON_FILENAME); const metaPath = path.join(videoDir, META_FILENAME); - // Resolve the actual primary VTT (may be a regional/auto English track like - // transcript.en-US.vtt) rather than assuming the literal transcript.en.vtt. const entries = await readdir(videoDir).catch(() => [] as string[]); - const vttPath = path.join(videoDir, resolvePrimaryVtt(entries) ?? VTT_FILENAME); const whisperPath = path.join(videoDir, WHISPER_FILENAME); - const [cuesMs, metaMs, vttMs, whisperMs] = await Promise.all([ + const inputs = captionInputs(entries); + const [cuesMs, metaMs, whisperMs, ...inputMs] = await Promise.all([ mtimeMs(cuesPath), mtimeMs(metaPath), - mtimeMs(vttPath), mtimeMs(whisperPath), + ...inputs.map((n) => mtimeMs(path.join(videoDir, n))), ]); + const vttMs = inputMs.reduce<number | null>( + (max, ms) => (ms !== null && (max === null || ms > max) ? ms : max), + null, + ); // Prefer whisper if present (matches pickIndexTranscript priority). const rawMs = whisperMs ?? vttMs; // Ordered to match normalizeTranscript's OWN precedence (metadata, then raw, @@ -243,5 +263,76 @@ export async function isCuesJsonFresh( if (cuesMs < metaMs || cuesMs < rawMs) { return { fresh: false, reason: "stale", cuesPath }; } + if (whisperMs === null && !(await followsCaptionTrackRule(videoDir, cuesPath, entries))) { + return { fresh: false, reason: "stale", cuesPath }; + } return { fresh: true, reason: "fresh", cuesPath }; } + +const HEAD_BYTES = 2048; + +async function readHead(file: string, bytes = HEAD_BYTES): Promise<string> { + const fh = await open(file, "r"); + try { + const buf = Buffer.alloc(bytes); + const { bytesRead } = await fh.read(buf, 0, bytes, 0); + return buf.subarray(0, bytesRead).toString("utf8"); + } finally { + await fh.close(); + } +} + +async function readTail(file: string, bytes = 64): Promise<string> { + const fh = await open(file, "r"); + try { + const { size } = await fh.stat(); + const start = Math.max(0, size - bytes); + const buf = Buffer.alloc(size - start); + const { bytesRead } = await fh.read(buf, 0, buf.length, start); + return buf.subarray(0, bytesRead).toString("utf8"); + } finally { + await fh.close(); + } +} + +// Whether a caption cues.json, already new enough by mtime, was made under the +// current caption-track rule — normalize records it as `captionTrackRule` in +// the file's first bytes — or, made before the rule had a version, holds what +// the rule would read anyway. The older rule ranked transcript.en.vtt first +// and parsed a cue-block VTT to no cues, so an unversioned file is stale only: +// +// when it holds NO cues (the cue-block parse, or the fallback to the next +// track, may find text now) — the cues are the last key normalize writes, +// so that is the file's tail; unless its one English VTT has word timing, +// whose parse did not change and which had nothing to choose from; +// when transcript.en.vtt and transcript.en-orig.vtt are both present and +// differ — it was read from en, and en-orig is read now. A byte-identical +// pair (equal size: on this corpus every one of 50,461 equal-size pairs was +// byte-identical) reads the same either way. +// +// One to three small reads (a head, a tail, two stats), never a parse. A +// mis-read only costs a rewrite, after which the record is there. +async function followsCaptionTrackRule( + videoDir: string, + cuesPath: string, + entries: readonly string[], +): Promise<boolean> { + const vtts = englishVttsByPreference(entries); + if (vtts.length === 0) return true; + try { + if (vtts.length === 1 && hasWordTiming(await readHead(path.join(videoDir, vtts[0])))) { + return true; + } + const m = (await readHead(cuesPath)).match(/"captionTrackRule"\s*:\s*(\d+)/); + if (m !== null && Number(m[1]) === CAPTION_TRACK_RULE_VERSION) return true; + if (/"cues"\s*:\s*\[\s*\]\s*\}\s*$/.test(await readTail(cuesPath))) return false; + if (!entries.includes(VTT_FILENAME) || !entries.includes(ORIG_VTT_FILENAME)) return true; + const [en, orig] = await Promise.all([ + stat(path.join(videoDir, VTT_FILENAME)), + stat(path.join(videoDir, ORIG_VTT_FILENAME)), + ]); + return en.size === orig.size; + } catch { + return false; + } +} diff --git a/common/lib/__fixtures__/vtt-cue-blocks.vtt b/common/lib/__fixtures__/vtt-cue-blocks.vtt @@ -0,0 +1,29 @@ +WEBVTT +Kind: captions +Language: en + +00:00:00.960 --> 00:00:06.000 +welcome back everyone today we are talking&nbsp; +about the harbor bridge finally finally&nbsp;&nbsp; + +00:00:06.000 --> 00:00:09.800 +I have been teasing that all week and let me&nbsp; +just say to every reader out there that&nbsp;&nbsp; + +00:00:09.800 --> 00:00:16.200 +no matter what you build in life you never&nbsp; +ever skip the load test plus later on&nbsp;&nbsp; + +00:00:16.200 --> 00:00:20.383 +we will look at the tide tables for the&nbsp; +north pier &amp; the &lt;old&gt; lighthouse + +00:00:20.383 --> 00:00:20.393 +[Applause] + +00:00:20.393 --> 00:00:20.400 +[Music] + +00:00:20.400 --> 00:00:26.640 +<i>right</i> to begin let us start with the&nbsp; +end the conclusion here I have said this&nbsp;&nbsp; diff --git a/common/lib/__fixtures__/vtt-rolling.vtt b/common/lib/__fixtures__/vtt-rolling.vtt @@ -0,0 +1,31 @@ +WEBVTT +Kind: captions +Language: en + +00:00:00.960 --> 00:00:03.070 align:start position:0% + +welcome<00:00:01.160><c> back</c><00:00:01.319><c> everyone</c><00:00:01.520><c> today</c><00:00:01.760><c> we</c> + +00:00:03.070 --> 00:00:03.080 align:start position:0% +welcome back everyone today we + + +00:00:03.080 --> 00:00:05.670 align:start position:0% +welcome back everyone today we +are<00:00:03.240><c> talking</c><00:00:03.639><c> about</c><00:00:04.319><c> the</c><00:00:04.759><c> harbor</c> + +00:00:05.670 --> 00:00:05.680 align:start position:0% +are talking about the harbor + + +00:00:05.680 --> 00:00:06.869 align:start position:0% +are talking about the harbor +bridge<00:00:06.000><c> finally</c><00:00:06.120><c> finally</c> + +00:00:06.869 --> 00:00:06.879 align:start position:0% +bridge finally finally + + +00:00:06.879 --> 00:00:09.310 align:start position:0% +bridge finally finally +I<00:00:07.000><c> have</c><00:00:07.120><c> been</c><00:00:07.279><c> teasing</c><00:00:07.560><c> that</c> diff --git a/common/lib/captionTrack.test.ts b/common/lib/captionTrack.test.ts @@ -0,0 +1,117 @@ +// The caption-track rule (lib/videoStatus.ts): which English VTT a video's +// transcript is read from, by name and then by content. +// +// Run with: node_modules/.bin/tsx --test common/lib/captionTrack.test.ts + +import { after, test } from "node:test"; +import assert from "node:assert/strict"; +import { mkdirSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs"; +import { tmpdir } from "node:os"; +import path from "node:path"; +import { + CAPTION_TRACK_RULE_VERSION, + TRANSCRIPT_PIN_FILENAME, + captionInputs, + englishVttsByPreference, + readEnglishVttCues, + readSubTracks, + resolvePrimaryVtt, +} from "./videoStatus"; + +const ROOT = mkdtempSync(path.join(tmpdir(), "caption-track-")); +after(() => rmSync(ROOT, { recursive: true, force: true })); + +const fixture = (name: string) => + readFileSync(path.join(import.meta.dirname, "__fixtures__", name), "utf8"); + +const ROLLING = fixture("vtt-rolling.vtt"); +const CUE_BLOCKS = fixture("vtt-cue-blocks.vtt"); +const EMPTY = "WEBVTT\nKind: captions\nLanguage: en\n\n"; + +let n = 0; +function videoDir(files: Record<string, string>): string { + const dir = path.join(ROOT, `v${n++}`); + mkdirSync(dir, { recursive: true }); + for (const [name, body] of Object.entries(files)) writeFileSync(path.join(dir, name), body); + return dir; +} + +test("en-orig ranks above en; regional above auto-translated; translations are not English", () => { + const entries = [ + "transcript.en-en-US.vtt", + "transcript.en.vtt", + "transcript.es-en-US.vtt", + "transcript.en-GB.vtt", + "transcript.en-orig.vtt", + "transcript.json", + ]; + assert.deepEqual(englishVttsByPreference(entries), [ + "transcript.en-orig.vtt", + "transcript.en.vtt", + "transcript.en-GB.vtt", + "transcript.en-en-US.vtt", + ]); + assert.equal(resolvePrimaryVtt(entries), "transcript.en-orig.vtt"); + assert.equal(resolvePrimaryVtt(["transcript.en.vtt", "transcript.en-US.vtt"]), "transcript.en.vtt"); + assert.equal(resolvePrimaryVtt(["transcript.es.vtt"]), null); +}); + +test("the operator's pin puts transcript.en.vtt first, and is a caption input", () => { + const entries = ["transcript.en-orig.vtt", "transcript.en.vtt", TRANSCRIPT_PIN_FILENAME]; + assert.equal(resolvePrimaryVtt(entries), "transcript.en.vtt"); + assert.deepEqual(captionInputs(entries), [ + "transcript.en.vtt", + "transcript.en-orig.vtt", + TRANSCRIPT_PIN_FILENAME, + ]); + // A pin with no English VTT is no input. + assert.deepEqual(captionInputs([TRANSCRIPT_PIN_FILENAME]), []); +}); + +test("readEnglishVttCues reads en-orig when both tracks have text", async () => { + const dir = videoDir({ + "transcript.en.vtt": CUE_BLOCKS, + "transcript.en-orig.vtt": ROLLING, + }); + const got = await readEnglishVttCues(dir); + assert.equal(got?.filename, "transcript.en-orig.vtt"); + assert.equal(got?.cues[0].text, "are talking about the harbor"); +}); + +test("readEnglishVttCues falls back past a track with no cues", async () => { + const dir = videoDir({ + "transcript.en-orig.vtt": EMPTY, + "transcript.en.vtt": CUE_BLOCKS, + }); + const got = await readEnglishVttCues(dir); + assert.equal(got?.filename, "transcript.en.vtt"); + assert.equal(got?.cues.length, 7); +}); + +test("readEnglishVttCues: every track empty is the first track with no cues; none is null", async () => { + const dir = videoDir({ + "transcript.en-orig.vtt": EMPTY, + "transcript.en.vtt": EMPTY, + }); + assert.deepEqual(await readEnglishVttCues(dir), { filename: "transcript.en-orig.vtt", cues: [] }); + assert.equal(await readEnglishVttCues(videoDir({ "transcript.es.vtt": ROLLING })), null); +}); + +test("readSubTracks lists the served en as an alternate beside an en-orig primary", async () => { + const dir = videoDir({ + "transcript.en.vtt": CUE_BLOCKS, + "transcript.en-orig.vtt": ROLLING, + "transcript.es.vtt": ROLLING, + }); + const tracks = (await readSubTracks(dir)).map((t) => t.track).sort(); + assert.deepEqual(tracks, ["en", "es"]); +}); + +test("umtool's copy of the caption-track rule version matches", () => { + const src = readFileSync( + path.join(import.meta.dirname, "..", "..", "umtool", "report-to-video", "cues.mjs"), + "utf8", + ); + const m = src.match(/export const CAPTION_TRACK_RULE_VERSION = (\d+);/); + assert.equal(Number(m?.[1]), CAPTION_TRACK_RULE_VERSION); +}); diff --git a/common/lib/sidecar-server.test.ts b/common/lib/sidecar-server.test.ts @@ -11,6 +11,7 @@ import "./attribution-server"; import "./diarization-server"; import "./digest-server"; import "./metadataHistory-server"; +import "./transcriptPin-server"; import "./wayback-server"; import { availabilitySidecar, @@ -56,6 +57,7 @@ test("every declared sidecar filename escapes SUB_FILE_RE, and all twelve are de "exclude-truncated-check.json", "metadata.history.json", "transcribe-outcome.json", + "transcript-pin.json", "wayback.json", ]); for (const name of SIDECAR_FILENAMES) { diff --git a/common/lib/subtitleProvenance.ts b/common/lib/subtitleProvenance.ts @@ -1,7 +1,7 @@ import path from "node:path"; import { createReadStream } from "node:fs"; import { readFile } from "node:fs/promises"; -import type { VideoFiles } from "./videoStatus"; +import { englishVttsByPreference, type VideoFiles } from "./videoStatus"; // Where a VTT transcript came from: YouTube's speech recognition ("asr", what // yt-dlp downloads under --write-auto-subs) or a human-authored/uploaded track @@ -140,6 +140,27 @@ export async function resolveVttProvenance( } } +// Where a video's English CAPTIONS came from, taken over every English VTT and +// not just the one the caption-track rule reads: "asr" only when each of them +// is, "manual" when any is, else "unknown". The rule prefers the original-audio +// ASR track (en-orig) even beside a human `en`, so asking only the track it +// reads would call a video with human captions ASR-only and schedule them for +// replacement. Single-track dirs pay the one sniff they always did. +export async function resolveCaptionsProvenance( + videoDir: string, + entries: readonly string[], +): Promise<SubtitleProvenance | null> { + const vtts = englishVttsByPreference(entries); + if (vtts.length === 0) return null; + let unknown = false; + for (const name of vtts) { + const p = await resolveVttProvenance(videoDir, name); + if (p === "manual") return "manual"; + if (p === "unknown") unknown = true; + } + return unknown ? "unknown" : "asr"; +} + // True when this video's ONLY transcript is YouTube ASR — the work-lane // candidate rule, shared by the snapshot buckets, the whisper gate // (transcribeOneFromQueue) and the auto-runner's download override so all four @@ -152,5 +173,5 @@ export async function isAutoSubsOnly( if (!files.ytVttFile || files.hasWhisper || files.isUntranscribable) { return false; } - return (await resolveVttProvenance(videoDir, files.ytVttFile)) === "asr"; + return (await resolveCaptionsProvenance(videoDir, files.entries)) === "asr"; } diff --git a/common/lib/transcriptPin-server.ts b/common/lib/transcriptPin-server.ts @@ -0,0 +1,39 @@ +// THE OPERATOR'S CAPTION-TRACK PICK. The editor's "Set as transcript" copies +// the chosen track to transcript.en.vtt and writes this sidecar beside it; +// while it is present, the caption-track rule ranks transcript.en.vtt first +// (lib/videoStatus.ts, englishVttsByPreference) — above en-orig, which the rule +// otherwise prefers. Only the file's PRESENCE is read by the rule; the record +// says which track was copied and when, for a person reading the dir. +// +// SERVER-ONLY (node:fs). + +import { TRANSCRIPT_PIN_FILENAME } from "./videoStatus"; +import { sidecar, sidecarField } from "./sidecar-server"; + +export type TranscriptPinRecord = { + // The track the operator picked (copied to transcript.en.vtt). + from: string; + pinnedAt: string; +}; + +// Any parseable record still pins: the rule reads presence, so the record's +// shape must never make a pin read as absent here and present there. +export function coerceTranscriptPin(value: unknown): TranscriptPinRecord { + const v = value as Partial<TranscriptPinRecord> | null; + return { + from: typeof v?.from === "string" ? v.from : "", + pinnedAt: typeof v?.pinnedAt === "string" ? v.pinnedAt : "", + }; +} + +export const transcriptPinSidecar = sidecar( + TRANSCRIPT_PIN_FILENAME, + sidecarField(coerceTranscriptPin), +); + +export async function pinTranscript(videoDir: string, from: string): Promise<void> { + await transcriptPinSidecar.write(videoDir, { + from, + pinnedAt: new Date().toISOString(), + }); +} diff --git a/common/lib/videoStatus.ts b/common/lib/videoStatus.ts @@ -1,12 +1,14 @@ import path from "node:path"; import { readdir, readFile, stat } from "node:fs/promises"; import { isPartAudioFile, isRealAudioFile } from "./mediaFiles"; +import { parseVtt, type Cue } from "./vtt"; export type VideoFiles = { hasMeta: boolean; hasYtVtt: boolean; - // The resolved primary English VTT filename (transcript.en.vtt when present, - // otherwise the best regional/auto English track — see resolvePrimaryVtt). + // The resolved primary English VTT filename by name (transcript.en-orig.vtt, + // else transcript.en.vtt, else the best regional/auto English track — see + // resolvePrimaryVtt; the cues may come from a later one, readEnglishVttCues). // Null when no English VTT exists. hasYtVtt === (ytVttFile !== null). ytVttFile: string | null; // True when a transcript.<lang>.vtt exists under a non-canonical name (i.e. @@ -40,8 +42,9 @@ export type VideoFiles = { export type IndexTranscript = | { kind: "whisper"; filename: "transcript.json" } - // filename is usually "transcript.en.vtt" but may be a regional/auto English - // track (e.g. "transcript.en-US.vtt") when YouTube served no plain `en` track. + // filename is the most preferred English track by name (resolvePrimaryVtt): + // transcript.en-orig.vtt, transcript.en.vtt, or a regional/auto one such as + // transcript.en-US.vtt. The cues are read with readEnglishVttCues. | { kind: "vtt"; filename: string }; export const VTT_FILENAME = "transcript.en.vtt"; @@ -92,18 +95,50 @@ function isSubExt(value: string): value is SubExt { return (SUB_EXT_VALUES as readonly string[]).includes(value); } -// Resolve the primary English transcript VTT in a video dir. Normally this is -// the canonical transcript.en.vtt, but YouTube sometimes serves a video's -// English captions only under regional/auto codes (transcript.en-US.vtt, -// transcript.en-en-US.vtt, transcript.en-orig.vtt) with no plain `en` track. A -// file counts as English iff the FIRST segment of its language code is `en`, so -// translations like transcript.ab-en-US.vtt / transcript.es-en-US.vtt are -// excluded. Preference: en (canonical) > en-orig (original audio) > -// regional/manual en-US,en-GB,… > auto-translated en-en-* variants. +// THE CAPTION-TRACK RULE — which English VTT a video's transcript is read from. +// It lives here and nowhere else; every reader that turns captions into cues +// (the index, normalize, report compose) goes through englishVttsByPreference / +// readEnglishVttCues, and the MCP, the export and report-to-video read what +// those wrote. +// +// A file counts as English iff the FIRST segment of its language code is `en`, +// so translations like transcript.ab-en-US.vtt / transcript.es-en-US.vtt are +// excluded. Preference: +// +// en-orig the captions of the ORIGINAL audio — the speaker's words. +// en canonical. Usually the same text as en-orig, but for some videos +// YouTube serves a rewritten/translated `en` that changes facts +// (a date, a word), and for some livestream VODs one in a cue-block +// shape with no word timing. +// en-US, en-GB, … regional/manual. +// en-en-* auto-translated en→en variants. +// +// AN OPERATOR'S PICK BEATS ALL OF IT. The editor's "Set as transcript" copies +// the chosen track to transcript.en.vtt and writes TRANSCRIPT_PIN_FILENAME +// beside it (lib/transcriptPin-server.ts); while that file is present, +// transcript.en.vtt ranks first. Only its presence is read — the rule stays a +// function of the listing. +// +// Name order is half the rule. The other half is CONTENT: the transcript is the +// first track in this order that parses to at least one cue (readEnglishVttCues), +// so a track that is present but empty never hides one that has text. +// CAPTION_TRACK_RULE_VERSION names this rule; bump it when the order or the +// fallback changes, and the next index build re-reads the records the change +// can reach (buildIndex.ts, "Caption track"), and a cues.json normalized under +// an older one reads as stale where it could differ (normalizeTranscript.ts). +// umtool/report-to-video/cues.mjs carries a copy (plain node, no tsx); change +// one, change the other — captionTrack.test.ts holds them equal. +export const CAPTION_TRACK_RULE_VERSION = 1; + +export const ORIG_VTT_FILENAME = "transcript.en-orig.vtt"; +// Not transcript.<x>.<y>: SUB_FILE_RE would read it as a subtitle track. +export const TRANSCRIPT_PIN_FILENAME = "transcript-pin.json"; + const EN_VTT_RE = /^transcript\.(en(?:-[^.]+)?)\.vtt$/; -function englishVttRank(track: string): number { - if (track === "en") return 0; - if (track === "en-orig") return 1; +function englishVttRank(track: string, pinned: boolean): number { + if (track === "en" && pinned) return -1; + if (track === "en-orig") return 0; + if (track === "en") return 1; if (/^en-en(?:-|$)/.test(track)) return 3; // auto-translated en→en variants return 2; // regional/manual en-US, en-GB, … } @@ -115,15 +150,63 @@ export function isEnglishVtt(name: string): boolean { return EN_VTT_RE.test(name); } -export function resolvePrimaryVtt(entries: string[]): string | null { - let best: { name: string; rank: number } | null = null; +// Every English VTT in a listing, most preferred first (ties by name, so the +// order never depends on readdir's). +export function englishVttsByPreference(entries: readonly string[]): string[] { + const pinned = entries.includes(TRANSCRIPT_PIN_FILENAME); + const ranked: { name: string; rank: number }[] = []; for (const e of entries) { const m = e.match(EN_VTT_RE); - if (!m) continue; - const rank = englishVttRank(m[1]); - if (!best || rank < best.rank) best = { name: e, rank }; + if (m) ranked.push({ name: e, rank: englishVttRank(m[1], pinned) }); + } + ranked.sort((a, b) => a.rank - b.rank || (a.name < b.name ? -1 : a.name > b.name ? 1 : 0)); + return ranked.map((r) => r.name); +} + +// The most preferred English VTT BY NAME — no file is read. This is the +// track the index stats for change detection and the one the editor labels the +// primary; the cues themselves come from readEnglishVttCues, which moves past +// it when it parses to nothing. +export function resolvePrimaryVtt(entries: readonly string[]): string | null { + return englishVttsByPreference(entries)[0] ?? null; +} + +// The files a caption transcript is derived from: every English VTT (the +// content fallback can reach any of them) and the operator's pin. Their newest +// mtime is what a cues.json, or an index record, is compared against. +export function captionInputs(entries: readonly string[]): string[] { + const out = englishVttsByPreference(entries); + if (out.length > 0 && entries.includes(TRANSCRIPT_PIN_FILENAME)) { + out.push(TRANSCRIPT_PIN_FILENAME); + } + return out; +} + +// The cues of a video's caption transcript: the first English VTT, in +// preference order, that parses to at least one cue. When every track parses +// to nothing, the most preferred one is returned with its empty list (the +// video has captions, they say nothing). Null when there is no English VTT, or +// none can be read. `entries` is the caller's readdir of `videoDir`, when it +// has one. +export async function readEnglishVttCues( + videoDir: string, + entries?: readonly string[], +): Promise<{ filename: string; cues: Cue[] } | null> { + const names = englishVttsByPreference( + entries ?? (await readdir(videoDir).catch(() => [] as string[])), + ); + let first: { filename: string; cues: Cue[] } | null = null; + for (const filename of names) { + let cues: Cue[]; + try { + cues = parseVtt(await readFile(path.join(videoDir, filename), "utf8")); + } catch { + continue; + } + if (cues.length > 0) return { filename, cues }; + first ??= { filename, cues }; } - return best?.name ?? null; + return first; } // Any transcript.<lang>.vtt file (any language code). These are the candidate @@ -194,9 +277,11 @@ export async function readSubTracks(videoDir: string): Promise<SubTrack[]> { const primaryVtt = resolvePrimaryVtt(entries); const tracks: SubTrack[] = []; for (const entry of entries) { - if (entry === VTT_FILENAME || entry === WHISPER_FILENAME) continue; - // The resolved primary English VTT (e.g. transcript.en-US.vtt when there's - // no transcript.en.vtt) is the main transcript, not an alternate sub-track. + if (entry === WHISPER_FILENAME) continue; + // The resolved primary English VTT (transcript.en-orig.vtt, or e.g. + // transcript.en-US.vtt when there is nothing better) is the main + // transcript, not an alternate sub-track. Any other English track — the + // served transcript.en.vtt beside an en-orig — is an alternate. if (entry === primaryVtt) continue; if (entry === CUES_JSON_FILENAME) continue; if (entry === LIVE_CHAT_CUES_FILENAME) continue; diff --git a/common/lib/vtt.test.ts b/common/lib/vtt.test.ts @@ -1,6 +1,14 @@ import { test } from "node:test"; import assert from "node:assert/strict"; -import { cuesToText, cuesToSrt, type Cue } from "./vtt"; +import { readFileSync } from "node:fs"; +import path from "node:path"; +import { cuesToText, cuesToSrt, hasWordTiming, parseVtt, type Cue } from "./vtt"; + +// Both fixtures keep the structure of real YouTube downloads (one rolling +// auto-caption track, one served `en` track of a livestream VOD) with the +// words replaced. +const fixture = (name: string) => + readFileSync(path.join(import.meta.dirname, "__fixtures__", name), "utf8"); // Run with: pnpm --filter yt-dlp-transcript-common exec tsx --test common/lib/vtt.test.ts @@ -37,3 +45,53 @@ test("cuesToSrt clamps negative times to zero", () => { const neg = [{ start: -3, end: -1, text: "neg" }]; assert.equal(cuesToSrt(neg), "1\n00:00:00,000 --> 00:00:00,000\nneg\n"); }); + +test("parseVtt reads a rolling auto-caption track one new line per cue", () => { + const cues = parseVtt(fixture("vtt-rolling.vtt")); + // Not asserted: the FIRST cue, whose carried-over line is a lone space — + // the body read stops at a whitespace-only line, so that cue reads as empty. + // That is the rolling parse as it stands, unchanged here. + assert.deepEqual( + cues.slice(-3).map((c) => c.text), + [ + "are talking about the harbor", + "bridge finally finally", + "I have been teasing that", + ], + ); + assert.equal(cues.at(-3)!.start, 3.08); +}); + +test("parseVtt reads a cue-block track (no word timing, &nbsp; line ends) — every line of every cue", () => { + const src = fixture("vtt-cue-blocks.vtt"); + assert.equal(hasWordTiming(src), false); + const cues = parseVtt(src); + assert.equal(cues.length, 7); + assert.deepEqual(cues[0], { + start: 0.96, + end: 6, + text: "welcome back everyone today we are talking about the harbor bridge finally finally", + }); + // Entities decoded after the tags are stripped: an escaped "<old>" is text. + assert.equal(cues[3].text, "we will look at the tide tables for the north pier & the <old> lighthouse"); + assert.equal(cues[4].text, "[Applause]"); + // A real tag is stripped. + assert.equal(cues[6].text, "right to begin let us start with the end the conclusion here I have said this"); +}); + +test("parseVtt: the shape is decided per document — a rolling track's untagged repeat cues are not read as text", () => { + const src = fixture("vtt-rolling.vtt"); + assert.equal(hasWordTiming(src), true); + // Seven cues in the file; the repeats add nothing. + assert.equal(parseVtt(src).length, 3); +}); + +test("parseVtt reads cue identifiers and settings as structure, not text", () => { + const src = + "WEBVTT\n\nNOTE a comment\n\n1\n00:00:01.000 --> 00:00:02.500 line:90%\n<v Host>First line\n\n" + + "2\n00:00:03.000 --> 00:00:04.000\nSecond line\n"; + assert.deepEqual(parseVtt(src), [ + { start: 1, end: 2.5, text: "First line" }, + { start: 3, end: 4, text: "Second line" }, + ]); +}); diff --git a/common/lib/vtt.ts b/common/lib/vtt.ts @@ -2,6 +2,13 @@ export type Cue = { start: number; end: number; text: string }; const TIMING_TAG_RE = /<\d{2}:\d{2}:\d{2}\.\d{3}>/; +// Whether VTT text carries inline word timing (<hh:mm:ss.mmm>) — the mark of +// YouTube's rolling auto-caption shape, which parseVtt reads differently from +// plain cue blocks. +export function hasWordTiming(src: string): boolean { + return TIMING_TAG_RE.test(src); +} + function parseTimestamp(ts: string): number { const m = ts.match(/(\d+):(\d+):(\d+)\.(\d+)/); if (!m) return 0; @@ -27,9 +34,32 @@ function stripTags(s: string): string { .trim(); } +// Strip EVERY tag from a plain cue-block line (<i>, <b>, <v Speaker>, a stray +// <c>), then decode entities — in that order, so an escaped "&lt;3" survives as +// text instead of being read as the start of a tag. +function stripPlainLine(s: string): string { + return decodeEntities(s.replace(/<[^>]*>/g, "")) + .replace(/\s+/g, " ") + .trim(); +} + +// Two VTT shapes, decided once per DOCUMENT: +// +// ROLLING (YouTube auto-captions, word timing). Each cue repeats the line +// before it and adds one new line carrying inline <hh:mm:ss.mmm><c>…</c> word +// timing; between them sits a ~10 ms cue that repeats the text with no tags. +// Only the tagged line is new, so only it is kept — reading every line would +// say each sentence two or three times. +// +// CUE BLOCKS (uploaded captions, and the `en` track YouTube serves for some +// livestream VODs: two lines a cue, `&nbsp;` at each line end, no word +// timing). Every line of a cue is its text. A document of this shape has no +// timing tag anywhere, which is what tells the two apart — read as ROLLING it +// came out as zero cues, and the video as textless. export function parseVtt(src: string): Cue[] { const lines = src.replace(/\r\n/g, "\n").split("\n"); const cues: Cue[] = []; + const rolling = hasWordTiming(src); let i = 0; while (i < lines.length) { @@ -50,12 +80,13 @@ export function parseVtt(src: string): Cue[] { i++; } - const taggedLines = body.filter((l) => TIMING_TAG_RE.test(l)); let text: string; - if (taggedLines.length > 0) { - text = stripTags(taggedLines[taggedLines.length - 1]); + if (!rolling) { + text = body.map(stripPlainLine).filter(Boolean).join(" "); } else { - continue; + const taggedLines = body.filter((l) => TIMING_TAG_RE.test(l)); + if (taggedLines.length === 0) continue; + text = stripTags(taggedLines[taggedLines.length - 1]); } if (text) cues.push({ start, end, text }); diff --git a/common/publish/composeReports.ts b/common/publish/composeReports.ts @@ -74,7 +74,7 @@ import { import type { TranscriptSummary } from "../lib/transcripts"; import { parseVtt, type Cue } from "../lib/vtt"; import { parseTranscriptJson } from "../lib/whisper"; -import { WHISPER_FILENAME, isEnglishVtt, resolvePrimaryVtt } from "../lib/videoStatus"; +import { WHISPER_FILENAME, englishVttsByPreference, readEnglishVttCues } from "../lib/videoStatus"; import { platformMomentUrl } from "../lib/momentUrl"; import { archiveOrgCitationLinks, type ArchiveOrgProvenance } from "../lib/archiveOrg"; import { loadArchiveOrgProvenance } from "../lib/archiveOrg-server"; @@ -243,22 +243,6 @@ type CitedRecord = { wayback: WaybackProvenance | null; }; -// The English VTT tracks of a video dir, `en-orig` first, then the order -// resolvePrimaryVtt prefers. -function englishVttsByPreference(entries: readonly string[]): string[] { - const vtts = entries.filter(isEnglishVtt); - const ordered: string[] = []; - const orig = vtts.find((n) => n === "transcript.en-orig.vtt"); - if (orig) ordered.push(orig); - const rest = vtts.filter((n) => n !== orig); - while (rest.length > 0) { - const best = resolvePrimaryVtt(rest)!; - ordered.push(best); - rest.splice(rest.indexOf(best), 1); - } - return ordered; -} - async function readCues(file: string, kind: "vtt" | "whisper"): Promise<Cue[]> { try { const raw = await readFile(file, "utf8"); @@ -295,17 +279,10 @@ export async function readCitedRecord( if (!meta) return null; summary = summarize(slug, id, meta, channelName); if (entries.includes(WHISPER_FILENAME)) cues = await readCues(path.join(dir, WHISPER_FILENAME), "whisper"); - if (cues.length === 0) { - const primary = resolvePrimaryVtt(entries); - if (primary) cues = await readCues(path.join(dir, primary), "vtt"); - } - } - if (cues.length === 0) { - for (const name of englishVttsByPreference(entries)) { - cues = await readCues(path.join(dir, name), "vtt"); - if (cues.length > 0) break; - } } + // The caption-track rule (videoStatus.ts): the first English VTT, `en-orig` + // first, that has cues. + if (cues.length === 0) cues = (await readEnglishVttCues(dir, entries))?.cues ?? []; const tracks: CitedRecord["tracks"] = []; if (fresh.fresh) { const n = await readNormalizedTranscript(fresh.cuesPath); diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -3,6 +3,9 @@ ## [Unreleased] - **A Wayback Machine capture is a copy, and says of what.** A capture URL (`web.archive.org/web/<timestamp>[id_|im_|…]/<original>`) names its record by what it is a capture of: an archived YouTube page by its YouTube id (no longer `watch`), a JW Player file by its media id (no longer `<id>-<rendition>.mp4`). Every download of a capture writes `wayback.json` (the original URL, the capture's timestamp, the capture page and its raw bytes); the video page says "Archived copy (Wayback Machine, <date>) of <original>"; a citation links the original, marked as possibly gone, and the Wayback copy, and its moment link is the capture, which plays (a capture URL never takes a time param). An existing record whose page is a capture is renamed to its id by the next snapshot. - **`archilyzer wayback refresh <slug> [--titles <file>] [--dry-run]`** brings a channel's Wayback copies up to that offline: `wayback.json`, the dir renamed through the snapshot's own reconcile pass with its roster entry moved, and with `--titles` (`id → {title, upload_date}`) the title and date of a raw file that has none, recorded in the metadata history as `wayback-provenance`. A record a live job holds is skipped and named; a second run changes nothing. +- **A video's captions are read from its original-audio track first, and a track with no text never hides one that has it.** Where YouTube serves both, `transcript.en-orig.vtt` (the captions of the original audio) is read before `transcript.en.vtt`, whose text can be a rewrite of what was said; then regional tracks (`en-US`, `en-GB`, …), then auto-translated `en-en-*` ones. The transcript is the first track in that order that has cues. One rule (`englishVttsByPreference` / `readEnglishVttCues` in `common/lib/videoStatus.ts`) serves the index, normalize and report compose, so search, the export, the MCP and report videos read the same words. **Set as transcript** on a video's page copies the chosen track to `transcript.en.vtt` and pins it there with `transcript-pin.json`, which ranks it first; deleting `transcript-pin.json` returns the video to the automatic pick. The served `en` track is listed as an alternate subtitle track where `en-orig` is the transcript. Videos with a human-made `en` track are not counted as auto-captions-only, so the replace-auto-captions lane still leaves them alone. +- **Captions in cue blocks are read.** A VTT with no inline word timing (uploaded captions, and the `en` track YouTube serves for some livestream recordings: two lines a cue, `&nbsp;` at each line end) is read cue by cue; it used to parse to no cues, which left those videos with no text in the index. +- **The next index build re-reads the caption records the new rule reaches, once.** A record with more than one English track, or with no cues stored, is re-read from disk; nothing else is. The build log says `Caption track v1: N record(s) re-read.` and, per channel, how many now read different text and how many had none and now do. The version is recorded only when no channel is held. A `transcript.cues.json` written before the rule is stale only where the rule reads something else — its `transcript.en.vtt` and `transcript.en-orig.vtt` differ, or it holds no cues and a cue-block parse or the next track may have them — so a **Normalize** run rewrites exactly those, and the digest and attribution lanes hold them until it does; a cues file now records the track it was read from (`vttFile`) and the rule (`captionTrackRule`). umtool's report-to-video refuses such a stale local record by name rather than cut from it. - **A cited moment at the very end of a recording prepares.** Prepare evidence media cuts a clip whose padding runs past the recording's end at the end (the recording's duration from its metadata), where it found no media for the padded span; a span that starts past the end is still refused. report-to-video keeps its strict rule. - **Exporting a changed report records a new revision of it.** `reports export` (and **Export reports** on a site's Reports tab, and the end of a prepare) commits a revision to the report's own git history, `sites/<site>/reports/<id>/history-git/`, whenever its `report.json` changed since the last one: the `report.json`, its Markdown export and the checksums of every export file, with a message of `Revision N` and a summary of the change. A re-export of an unchanged report records nothing. The commits carry the site's name and a `noreply@<site>.invalid` address with dates in UTC, never your git name, email or time zone. The Reports tab shows each report's revision, its commit and the last change under **Exports**, and the site's next build publishes the history. Add `history-git/` to the corpus repository's `.gitignore`. - **archive.org files come over BitTorrent when possible, else straight from archive.org — never through yt-dlp.** The chosen file of an archive.org import is fetched from the item's own torrent (`<identifier>_archive.torrent`, which lists archive.org as a web seed, so other peers take load off archive.org) with aria2c, only that file of the item, and seeded afterwards for 10 minutes or to a ratio of 1, whichever comes first; the log shows "torrent: <file> (n of m pieces, peers p, web seed yes)" and "seeding 10 min…". With no aria2c, a torrent that does not carry the file, or no progress for 5 minutes, it is downloaded directly from `archive.org/download/…` instead (resumable, backing off on 429/503), and the log says "fell back to direct download: <reason>". Every file is checked against archive.org's sha1/md5: a mismatch is downloaded once more directly, a second one fails the record. The record is written from the item's metadata: `metadata.info.json` with the file's page, the canonical id, the duration ffprobe measures and archive.org's playable copies of the file, the `archiveorg.json` provenance (a mirror's original title, date and uploader), and `audio.<fmt>` — an audio file already in the channel's format is used as is, anything else goes through the app's audio extraction, a video kept in the saved-video store when the channel keeps sources. An .avi/.mpeg/.flac/.wav original is fetched as archive.org's mp4 or mp3 of it. aria2c runs in its own process group: cancelling the job stops it and everything it started, and it stops itself if the editor exits. New settings block `archiveOrg` (`torrent`, `seedMinutes`, `seedRatio`, `stallMinutes`, `maxPeers`, `maxDownloadKiBps`, `maxUploadKiBps`), `ARIA2C_BIN`, an aria2c row in `archilyzer doctor`, and `aria2` in the runtime Docker images. diff --git a/editor/app/channels/[slug]/videos/[id]/components/cards/TranscriptSourceSection.tsx b/editor/app/channels/[slug]/videos/[id]/components/cards/TranscriptSourceSection.tsx @@ -5,7 +5,7 @@ import type { SubtitleProvenance } from "yt-dlp-transcript-common/lib/subtitlePr import { setPrimaryTranscriptAction, } from "../../videoActions"; -import { CANONICAL_VTT, WHISPER_FILENAME } from "./videoFiles"; +import { CANONICAL_VTT, TRANSCRIPT_PIN_FILENAME, WHISPER_FILENAME } from "./videoFiles"; const PROVENANCE_LABEL: Record<SubtitleProvenance, string> = { asr: "YouTube auto-captions", @@ -46,10 +46,11 @@ export function TranscriptSourceSection({ <div className="flex flex-col gap-3"> <p className="text-sm text-muted-foreground"> Pick which subtitle track is this video&apos;s primary transcript. The - chosen track is copied to <code>{CANONICAL_VTT}</code> — the canonical - name the index and viewer read. Reversible: delete{" "} - <code>{CANONICAL_VTT}</code> in Files to fall back to the automatic - English pick, or choose another track to switch. + chosen track is copied to <code>{CANONICAL_VTT}</code> and pinned there + with <code>{TRANSCRIPT_PIN_FILENAME}</code>. Unpinned, the transcript + is the first English track with text, <code>transcript.en-orig.vtt</code>{" "} + first. Reversible: delete <code>{TRANSCRIPT_PIN_FILENAME}</code> in + Files to fall back to that pick, or choose another track to switch. </p> {hasWhisper && ( <p diff --git a/editor/app/channels/[slug]/videos/[id]/components/cards/videoFiles.ts b/editor/app/channels/[slug]/videos/[id]/components/cards/videoFiles.ts @@ -14,6 +14,8 @@ export type VideoFile = { export const WHISPER_FILENAME = "transcript.json"; export const CANONICAL_VTT = "transcript.en.vtt"; +// lib/videoStatus.ts TRANSCRIPT_PIN_FILENAME (this module is client-side). +export const TRANSCRIPT_PIN_FILENAME = "transcript-pin.json"; export function isTranscriptVttName(name: string): boolean { return /^transcript\.[^.]+\.vtt$/.test(name); diff --git a/editor/app/channels/[slug]/videos/[id]/videoActions.ts b/editor/app/channels/[slug]/videos/[id]/videoActions.ts @@ -46,6 +46,7 @@ import { import { onDrive } from "yt-dlp-transcript-common/lib/storageHealth"; import { isTierable } from "yt-dlp-transcript-common/lib/mediaTier"; import { setExcludedFromTruncatedCheck } from "yt-dlp-transcript-common/lib/excludeTruncatedCheck-server"; +import { pinTranscript } from "yt-dlp-transcript-common/lib/transcriptPin-server"; import { pruneFailedTranscriptions } from "yt-dlp-transcript-common/controller/failedTranscriptions"; import { transcodeAudio } from "yt-dlp-transcript-common/controller/transcode"; import { @@ -583,9 +584,10 @@ export async function deleteVideoFileAction( // Promote a transcript.<lang>.vtt track to the canonical transcript.en.vtt so // the index, snapshot, and viewer all treat it as the primary transcript. The // chosen file is copied (not moved) so the original language-coded track is kept -// and the choice stays reversible/repeatable — delete transcript.en.vtt to fall -// back to the automatic regional-English pick, or pick a different track to -// switch again. +// and the choice stays reversible/repeatable — delete transcript-pin.json to +// fall back to the automatic pick, or pick a different track to switch again. +// The pin is what makes the copy win: the caption-track rule otherwise ranks +// transcript.en-orig.vtt above transcript.en.vtt (lib/videoStatus.ts). export async function setPrimaryTranscriptAction( slug: string, videoId: string, @@ -611,6 +613,7 @@ export async function setPrimaryTranscriptAction( if (filename !== VTT_FILENAME) { await writeFileAtomic(path.join(videoDir, VTT_FILENAME), raw); } + await pinTranscript(videoDir, filename); // The video page reads the dir directly, so it reflects the new primary right // away. The channel list + diagnostics bucket read the cached snapshot and // refresh on the next snapshot regeneration (same as the other video actions). diff --git a/export/CHANGELOG.md b/export/CHANGELOG.md @@ -2,6 +2,7 @@ ## [Unreleased] - **A citation of a Wayback Machine copy links its original and the copy.** A cited record downloaded from a Wayback capture shows "Original (may be gone)", the original at the cited second where its platform takes one, and "Wayback Machine copy, <capture date>", the capture page, which plays. Its moment link is the capture: a capture URL never takes a time param. +- **Transcripts read the original-audio captions.** Where a video has both, its transcript is YouTube's `en-orig` track (the captions of what was said) rather than the served `en`, which can reword it; a track with no text falls through to the next. Videos whose only captions are in cue blocks (some livestream recordings) have their text. - **A report shows its revision, and every edit to it can be checked.** A report's date line ends with "revision N", linking to its history ("edited since revision N" when the report has changed since). The history page, `/reports/<id>/history/`, lists every revision, newest first: its number, date (UTC), commit hash and the sha256 of its `report.json`, what changed (claims added or removed, verdicts changed, claims edited, citations added or removed, quotes edited, title, series or subtitle changed), and each changed claim's title, text, verdict and findings with the words removed struck through and the words added marked. The same data is in `history.json` beside the page. Each report's history is its own git repository, published for cloning: `git clone <site>/reports/<id>/history/repo`. Each commit names the site as its author, with its date in UTC. The footer of the report's HTML, PDF and Markdown downloads begins with the revision number, and its sha256 can be looked up on the history page. Needs `reports export` and a rebuild and deploy of each site with reports. - **A report can be saved whole: as one HTML page, a PDF, Markdown, or an evidence pack.** A report page's download line reads HTML · PDF · Markdown · Evidence pack · Citations JSON · CSV, each listed only when the site publishes it. The HTML is one file that opens with no network: the report with its verdicts, the document's sentences and the post screenshots inside it, numbered citations, and a reference list giving each quote's speaker, date, record, the original at its time and the moment page on the site. The PDF is that page printed. The Markdown is the same report as plain text with numbered references. The evidence pack is a zip of the page with its clips, stills and screenshots beside it, so the clips play offline. Each ends with a line naming the report's revision, its date and the start of its checksum. Needs `reports export` (or prepare) and a rebuild and deploy of each site with reports. - **A report's claim can carry a flag, its header names the document under review, and a site with one report names it in the browser tab.** `report.json` claim `flag` (one line, at most 60 characters) shows as a small pill in the accent colour beside the claim's verdict, e.g. "No source given". On a report-only site with one report, the home page's tab title is the report's, as on the report's own page. A report's page header is its name — with a `series`, the series on one line in the accent colour and the title on the line below; without one, the title — then one small line of dates and the revision ("2026-10-04 · updated 2026-10-05 · revision 1"; "updated" only when it differs), then a card for the document under review: its title linking to the document, "<author> · <publisher> · <date>", and its archive links folded away, on a left rail in the document's colour (a source's `accent`, `"#rrggbb"`; without one, the border colour), then the subtitle. The page names no byline or site of its own. A claim that cites the document's own sentence shows that sentence — its still, else its words — on the same rail, with no link up to the card and no paraphrase beside it; a sentence of another document links "from <title>" to that document's box; a titled claim with no such sentence shows its text under the title, plain. A report's citation can say where its evidence came from (`origin`: `"subject"`, the document under review gave it; `"added"`, the report's author found it). A claim lists what the report added first, each card marked with the Archilyzer mark under its number ("Not in the article", or "Not in the source", is the mark's tooltip and what a screen reader says), then evidence of unknown origin, then what the document gave itself folded under "In the article (n)"; the reference list marks an added citation with the mark too, and a claim's flag pill wears the same mark. A fact-check's page reads in three tiers, each opened by a hairline with one, two or three dots: the quick take (the tally, the summary, and links to what the check found, every claim and the downloads); **What the check found**, every ruled claim grouped by verdict (contradicted, not found, partly, untestable, corroborated), one line each linking to the claim, with its `gist` (a new optional claim field, one line, at most 240 characters) and its flag; and **Every claim, with its evidence**, which opens with **How it was checked** (`method`, a new optional report field in markdown). A report of kind `sweep` has the first and last tiers only. Needs a rebuild and deploy of the site. diff --git a/umtool/report-to-video/cues.mjs b/umtool/report-to-video/cues.mjs @@ -56,7 +56,7 @@ // a channel the stale copy lacked is found. Manifest and shard URLs carry no // version, so a cue window does not move because of it. -import { readFile, writeFile, mkdir, lstat, readlink } from "node:fs/promises"; +import { readFile, writeFile, mkdir, lstat, readlink, readdir, stat } from "node:fs/promises"; import path from "node:path"; import os from "node:os"; import { createHash } from "node:crypto"; @@ -114,6 +114,40 @@ export function siteOriginFromManifest(manifest) { return null; } +// A TWIN OF `CAPTION_TRACK_RULE_VERSION` (`common/lib/videoStatus.ts`) — the +// caption-track rule a `transcript.cues.json` records as `captionTrackRule`. +// Copied for the same reason as the text guard below: umtool's bins run under +// plain node. Change one, change the other. +export const CAPTION_TRACK_RULE_VERSION = 1; +const ENGLISH_VTT_RE = /^transcript\.en(?:-[^.]+)?\.vtt$/; + +// A local caption record normalized under an older caption-track rule, where +// the rule now reads other words — the same test as `isCuesJsonFresh` +// (`common/controller/normalizeTranscript.ts`, followsCaptionTrackRule): no +// cues at all (a cue-block VTT used to parse to nothing; a lone track with word +// timing is genuinely empty and is let through), or a transcript.en.vtt and +// transcript.en-orig.vtt that differ (the older rule read en; en-orig is read +// now). Cutting from it would widen clips on text the corpus no longer +// publishes, so it is refused with the fix, not used. +async function staleCaptionRecord(videoDir, record) { + if (record?.source !== "vtt" || record.captionTrackRule === CAPTION_TRACK_RULE_VERSION) { + return false; + } + const entries = await readdir(videoDir).catch(() => []); + const english = entries.filter((e) => ENGLISH_VTT_RE.test(e)); + if (!Array.isArray(record.cues) || record.cues.length === 0) { + if (english.length !== 1) return true; + const head = (await readFile(path.join(videoDir, english[0]), "utf8").catch(() => "")).slice(0, 2048); + return !/<\d{2}:\d{2}:\d{2}\.\d{3}>/.test(head); + } + if (!entries.includes("transcript.en.vtt") || !entries.includes("transcript.en-orig.vtt")) return false; + const [en, orig] = await Promise.all([ + stat(path.join(videoDir, "transcript.en.vtt")), + stat(path.join(videoDir, "transcript.en-orig.vtt")), + ]); + return en.size !== orig.size; +} + export class CueLookupError extends Error { constructor(message, { channelSlug, videoId, tried }) { super(message); @@ -423,6 +457,14 @@ export function createCueSource({ await assertChannelReachable(channelSlug); const p = path.join(channelsDir, channelSlug, "data", videoId, "transcript.cues.json"); const parsed = JSON.parse(await readFile(p, "utf8")); + if (await staleCaptionRecord(path.dirname(p), parsed)) { + throw new CueLookupError( + `${channelSlug}/${videoId}: transcript.cues.json was normalized under an older caption-track rule ` + + `(it holds the served en track's words where en-orig is read now, or no words at all). ` + + `Run Normalize for channel ${channelSlug} in the editor, or pass --cue-source http.`, + { channelSlug, videoId, tried: [p] }, + ); + } return { ...parsed, from: "local" }; } diff --git a/umtool/report-to-video/cues.test.mjs b/umtool/report-to-video/cues.test.mjs @@ -12,6 +12,7 @@ import { tmpdir } from "node:os"; import path from "node:path"; import { + CAPTION_TRACK_RULE_VERSION, createCueSource, pageFileName, pageUrlFrom, @@ -148,6 +149,58 @@ test("a local corpus is preferred over the network", async () => { } }); +// --- a caption record normalized under an older caption-track rule --------- + +test("a local caption record from before the caption-track rule is refused where the rule could read other words", async () => { + const dir = await mkdtemp(path.join(tmpdir(), "cues-rule-")); + try { + const vdir = path.join(dir, "chan", "data", "vid1"); + await mkdir(vdir, { recursive: true }); + await writeFile(path.join(vdir, "transcript.en.vtt"), "WEBVTT\n\nserved words\n"); + await writeFile(path.join(vdir, "transcript.en-orig.vtt"), "WEBVTT\n\nwhat was said\n"); + await writeFile(path.join(vdir, "transcript.cues.json"), JSON.stringify({ ...RECORD, source: "vtt" })); + const seen = []; + const src = createCueSource({ channelsDir: dir, siteOrigin: ORIGIN, cacheDir: null, fetchImpl: stubFetch(ROUTES, seen) }); + await assert.rejects(src.load("chan", "vid1"), (err) => { + assert.equal(err.name, "CueLookupError"); + assert.match(err.message, /older caption-track rule.*Run Normalize for channel chan/s); + return true; + }); + assert.deepEqual(seen, [], "never answered from the archive instead"); + + // Normalized under the current rule: read as usual. + await writeFile( + path.join(vdir, "transcript.cues.json"), + JSON.stringify({ ...RECORD, source: "vtt", captionTrackRule: CAPTION_TRACK_RULE_VERSION }), + ); + assert.equal((await src.load("chan", "vid1")).from, "local"); + + // An unversioned record beside a byte-identical pair reads the same either way. + await writeFile(path.join(vdir, "transcript.en.vtt"), "WEBVTT\n\nwhat was said\n"); + await writeFile(path.join(vdir, "transcript.cues.json"), JSON.stringify({ ...RECORD, source: "vtt" })); + assert.equal((await src.load("chan", "vid1")).from, "local"); + } finally { + await rm(dir, { recursive: true, force: true }); + } +}); + +test("a lone-track caption record with cues needs no rule to be read", async () => { + const dir = await mkdtemp(path.join(tmpdir(), "cues-rule-")); + try { + const vdir = path.join(dir, "chan", "data", "vid1"); + await mkdir(vdir, { recursive: true }); + await writeFile(path.join(vdir, "transcript.en.vtt"), "WEBVTT\n"); + await writeFile(path.join(vdir, "transcript.cues.json"), JSON.stringify({ ...RECORD, source: "vtt" })); + const src = createCueSource({ channelsDir: dir, siteOrigin: ORIGIN, cacheDir: null, fetchImpl: stubFetch(ROUTES) }); + assert.equal((await src.load("chan", "vid1")).from, "local"); + // …but one with no cues is refused: a cue-block track used to parse to none. + await writeFile(path.join(vdir, "transcript.cues.json"), JSON.stringify({ ...RECORD, source: "vtt", cues: [] })); + await assert.rejects(src.load("chan", "vid1"), /older caption-track rule/); + } finally { + await rm(dir, { recursive: true, force: true }); + } +}); + // --- a channel the corpus holds but whose text it cannot read --------------- // // The bug these cover: on the RETIRED layout `data/` is a symlink to another