Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 76b5500e2da4f9469d31cf1c6d340307f840f7a4
parent 797c449c2d534a2a780d93cd71fb1520ed28bae1
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sat, 26 Sep 2026 18:32:14 -0400

Merge editor/full-fetch-media — release 10 slice N: the whole-recording fetch downloads anyway (forceMedia through the no-captions fallback: Persist source video, fetch_clip full: true and Persist kept now fetch and persist the source on a captions-only channel even with a transcript on disk; a Whisper transcript untouched, no audio extracted beside one), and every metadata.info.json rewrite appends a normalized diff to metadata.history.json (volatile keys fingerprinted, counters apart, content from → to, cap 200) with a read-only panel on the video page

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
MAGENTS.md | 13+++++++++++++
Mcommon/controller/persistKept.ts | 9+++++++--
Acommon/controller/persistKeptForceMedia.test.ts | 137+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/metadataHistory-server.test.ts | 128+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/metadataHistory-server.ts | 131+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/metadataHistory.test.ts | 215+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/metadataHistory.ts | 372+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/lib/sidecar-server.test.ts | 4+++-
Acommon/views/metadataHistoryView.test.ts | 129+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/views/metadataHistoryView.ts | 155+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/ytdlp/downloadOneManaged.ts | 195+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++----------------
Acommon/ytdlp/forceMedia.test.ts | 302++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/ytdlp/runYtdlp.ts | 28+++++++++++++++++++++++++---
Meditor/CHANGELOG.md | 2++
Aeditor/app/channels/[slug]/videos/[id]/components/MetadataHistoryDetails.tsx | 81+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/app/channels/[slug]/videos/[id]/components/cards/SourceVideoSection.tsx | 20++++++++++++--------
Meditor/app/channels/[slug]/videos/[id]/lib/videoChoreCards.ts | 2+-
Meditor/app/channels/[slug]/videos/[id]/page.tsx | 10++++++++++
Meditor/app/channels/[slug]/videos/[id]/videoActions.ts | 14+++++++++++---
Meditor/e2e/fetch-window.spec.ts | 62+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-
Meditor/e2e/fixtures/bin/fake-ytdlp.mjs | 18++++++++++++++++++
Aeditor/e2e/persist-youtube-handling.spec.ts | 175+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mplans/FACTS.md | 21++++++++++++++++-----
Mplans/release-10.md | 236+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
24 files changed, 2397 insertions(+), 62 deletions(-)

diff --git a/AGENTS.md b/AGENTS.md @@ -89,6 +89,19 @@ machine with no editor. The editor fetches only for a channel it already archive `transcripts/channels/<slug>/`), so an MCP pointed at a public site with a fresh editor gets a 404 `Channel "<slug>" not found` on every clip. +**`full: true` and "Persist source video" download the source even when a transcript or +captions exist** (`forceMedia` on `downloadOneManaged`, release 10 slice N; "Persist kept +now" too). On a youtube-handling channel that pass used to be skipped for any video with a +transcript or captions — the job ended `done` with no file. With a transcript on disk the +forced pass moves the container into the saved-video store and does nothing else: no +subtitles, no `audio.<fmt>`, no transcription. YouTube's subtitles are fetched again first, +as on any re-download; a Whisper transcript is not touched, and no audio is extracted beside +a transcript. **Every rewrite of `metadata.info.json` appends to +`metadata.history.json`** beside it (`common/lib/metadataHistory.ts`): the old and new values +of changed content keys (title, description, …), the counters that moved, and +formats/thumbnails/caption URLs only by fingerprint; 200 entries, newest last. The video +page shows it under the description. + `umtool/report-to-video/` renders a cited sweep report to an mp4. What it needs: - **yt-dlp** — clip media is fetched over the network per clip. Local media does not diff --git a/common/controller/persistKept.ts b/common/controller/persistKept.ts @@ -15,8 +15,9 @@ import { downloadOneManaged } from "../ytdlp/downloadOneManaged"; // after those videos had already been downloaded audio-only. // // For each kept video whose source isn't saved yet, it re-fetches the container -// via downloadOneManaged with keepSourceVideoOverride (which app-extracts audio -// and moves the container into the saved store), without disturbing the existing +// via downloadOneManaged with keepSourceVideoOverride + forceMedia (which moves +// the container into the saved store; a youtube-handling video that already has +// a transcript gets no audio extracted), without disturbing the existing // transcript. Mirrors redownloadToArchiveAction (videoActions.ts) but loops over // the whole window. Videos already saved, or whose URL can't be resolved, are // skipped rather than failing the pass. @@ -130,6 +131,10 @@ export async function persistKept({ globalSkipLiveDownloads: settings.skipLiveDownloads, appendArchive: true, keepSourceVideoOverride: true, + // "Persist kept" means persist: on a youtube-handling channel the media + // pass is otherwise skipped for any video with a transcript or + // captions, which is every kept one (release 10 slice N). + forceMedia: true, }); result.persisted += 1; } catch (e) { diff --git a/common/controller/persistKeptForceMedia.test.ts b/common/controller/persistKeptForceMedia.test.ts @@ -0,0 +1,137 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { chmod, mkdir, mkdtemp, readFile, readdir, rm, writeFile } from "node:fs/promises"; +import { tmpdir } from "node:os"; +import path from "node:path"; +import type { Paths } from "../lib/paths"; +import type { ChannelConfig } from "../lib/channelConfig"; +import { persistKept } from "./persistKept"; + +// Run with: +// pnpm --filter yt-dlp-transcript-common exec tsx --test controller/persistKeptForceMedia.test.ts +// +// "PERSIST KEPT" MEANS PERSIST (release 10 slice N, review L1). On a +// youtube-handling channel every kept video has a transcript or captions, so +// without `forceMedia` the bulk pass ran two subtitle-shaped yt-dlp passes per +// video and persisted nothing — the same silent no-op "Persist source video" +// had. `persistKept.test.ts` never reaches downloadOneManaged, and the e2e +// "Persist kept now" check uses a transcribe fixture, so this is the one test +// that fails when persistKept stops passing the flag. A separate file because +// persistKept reads settings through `getPaths()`, which caches the first +// SETTINGS_FILE it sees: this process sets it before anything else can. + +const ID = "H64QQZuw-aA"; +const VIDEO = `https://www.youtube.com/watch?v=${ID}`; + +// The same fake as ytdlp/forceMedia.test.ts: the prefetch copies meta.json in +// as the info json, the subtitle pass writes nothing, and only a media pass +// (the `source-media` output template) writes a container. +const FAKE_YTDLP = `#!/usr/bin/env node +const fs = require("node:fs"); +const path = require("node:path"); +const argv = process.argv.slice(2); +const root = process.env.FAKE_ROOT; +fs.appendFileSync(path.join(root, "argv.log"), JSON.stringify(argv) + "\\n"); +const has = (f) => argv.includes(f); +const arg = (f) => { const i = argv.indexOf(f); return i < 0 ? undefined : argv[i + 1]; }; +const info = arg("--load-info-json"); +const id = info ? path.basename(path.dirname(info)) : "${ID}"; +const dir = path.join("data", id); +fs.mkdirSync(dir, { recursive: true }); +if (has("--skip-download") && has("--write-info-json") && has("--no-write-subs")) { + fs.copyFileSync(path.join(root, "meta.json"), path.join(dir, "metadata.info.json")); + process.exit(0); +} +if (has("--skip-download") && has("--write-auto-subs")) { + console.log("DLOM_ARCHIVE youtube " + id); + process.exit(0); +} +if ((arg("-o") || "").includes("source-media")) { + fs.writeFileSync(path.join(dir, "source-media.mp4"), "container bytes for " + id); + console.log("DLOM_ARCHIVE youtube " + id); + process.exit(0); +} +console.error("fake: unknown invocation " + argv.join(" ")); +process.exit(2); +`; + +test("persistKept forces the media pass: a kept youtube-handling video with a transcript is persisted, with no audio", async () => { + const root = await mkdtemp(path.join(tmpdir(), "persist-kept-force-")); + try { + // Disk floor off, so the per-item gate cannot turn this into a skip on a + // full disk. Set before persistKept's first getSettings() → getPaths(). + const settingsFile = path.join(root, "settings.json"); + await writeFile(settingsFile, JSON.stringify({ minFreeDiskGB: 0 })); + process.env.SETTINGS_FILE = settingsFile; + process.env.TRANSCRIPTS_DIR = root; + process.env.FAKE_ROOT = root; + + const ytdlp = path.join(root, "fake-ytdlp.cjs"); + await writeFile(ytdlp, FAKE_YTDLP); + await chmod(ytdlp, 0o755); + const meta = { + id: ID, + title: "A kept video", + upload_date: "20260101", + duration: 60, + webpage_url: VIDEO, + subtitles: {}, + automatic_captions: { en: [{ url: "https://example/cap" }] }, + }; + await writeFile(path.join(root, "meta.json"), JSON.stringify(meta)); + + const paths = { + transcriptsDir: root, + channelsDir: path.join(root, "channels"), + savedVideosDir: path.join(root, "saved-videos"), + ytdlpBin: ytdlp, + ffmpegBin: path.join(root, "no-ffmpeg"), + ffprobeBin: path.join(root, "no-ffprobe"), + } as Paths; + const videoDir = path.join(paths.channelsDir, "c", "data", ID); + await mkdir(videoDir, { recursive: true }); + await writeFile(path.join(videoDir, "metadata.info.json"), JSON.stringify(meta)); + const transcript = '{"transcription":[{"text":"kept"}]}'; + await writeFile(path.join(videoDir, "transcript.json"), transcript); + + let log = ""; + const result = await persistKept({ + paths, + channelSlug: "c", + channelConfig: { + handling: "youtube", + url: "https://www.youtube.com/@c/videos", + keepLatest: 1, + } as ChannelConfig, + onLog: (s) => { + log += s; + }, + }); + + assert.equal(result.kept, 1); + assert.equal(result.persisted, 1); + assert.match(log, /forceMedia: downloading the source although a transcript is on disk/); + // The media pass ran: three spawns, the last without --skip-download. + const argvs = (await readFile(path.join(root, "argv.log"), "utf8")) + .split("\n") + .filter(Boolean) + .map((l) => JSON.parse(l) as string[]); + assert.equal(argvs.length, 3); + assert.ok(!argvs[2].includes("--skip-download")); + // Persisted, and nothing else: the pointer is there, the container is in + // the store, no audio beside the transcript, the transcript untouched. + const pointer = JSON.parse( + await readFile(path.join(videoDir, "saved-video.json"), "utf8"), + ) as { dir: string; file: string; keepReason?: string }; + assert.equal(pointer.keepReason, "override"); + assert.equal( + await readFile(path.join(pointer.dir, pointer.file), "utf8"), + `container bytes for ${ID}`, + ); + const files = await readdir(videoDir); + assert.deepEqual(files.filter((f) => /^(audio|source-media)\./.test(f)), []); + assert.equal(await readFile(path.join(videoDir, "transcript.json"), "utf8"), transcript); + } finally { + await rm(root, { recursive: true, force: true }); + } +}); diff --git a/common/lib/metadataHistory-server.test.ts b/common/lib/metadataHistory-server.test.ts @@ -0,0 +1,128 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { mkdtemp, readFile, readdir, rm, writeFile } from "node:fs/promises"; +import os from "node:os"; +import path from "node:path"; +import { + loadMetadataHistory, + metadataHistoryPath, + withMetadataHistory, +} from "./metadataHistory-server"; + +// Run with: +// pnpm --filter yt-dlp-transcript-common exec tsx --test lib/metadataHistory-server.test.ts + +async function withDir(fn: (dir: string) => Promise<void>): Promise<void> { + const dir = await mkdtemp(path.join(os.tmpdir(), "metadata-history-")); + try { + await fn(dir); + } finally { + await rm(dir, { recursive: true, force: true }); + } +} + +const INFO = "metadata.info.json"; +const meta = (title: string, view_count = 1) => + JSON.stringify({ id: "v1", title, view_count, formats: [{ url: `u-${title}` }] }); + +test("a rewrite that changed the file appends one entry", async () => { + await withDir(async (dir) => { + await writeFile(path.join(dir, INFO), meta("old", 10)); + const out = await withMetadataHistory( + dir, + { by: "prefetch", requestedBy: "umtool" }, + async () => { + await writeFile(path.join(dir, INFO), meta("new", 12)); + return 42; + }, + ); + assert.equal(out, 42); + const h = await loadMetadataHistory(dir); + assert.equal(h?.entries.length, 1); + const e = h!.entries[0]; + assert.equal(e.by, "prefetch"); + assert.equal(e.requestedBy, "umtool"); + assert.deepEqual(e.changed.title, { from: "old", to: "new" }); + assert.deepEqual(e.counters.view_count, [10, 12]); + assert.deepEqual(e.volatile, ["formats"]); + assert.ok(e.from.at, "the old file's mtime is recorded"); + + // A second rewrite appends, newest last. + await withMetadataHistory(dir, { by: "audio-check" }, async () => { + await writeFile(path.join(dir, INFO), meta("newer", 12)); + }); + const h2 = await loadMetadataHistory(dir); + assert.deepEqual( + h2?.entries.map((x) => x.by), + ["prefetch", "audio-check"], + ); + assert.equal("requestedBy" in h2!.entries[1], false); + // Written through the sidecar: indented JSON with a trailing newline, and + // no temp file left behind. + const raw = await readFile(metadataHistoryPath(dir), "utf8"); + assert.ok(raw.startsWith("{\n") && raw.endsWith("}\n")); + assert.deepEqual((await readdir(dir)).sort(), [ + "metadata.history.json", + "metadata.info.json", + ]); + }); +}); + +test("an identical rewrite records nothing", async () => { + await withDir(async (dir) => { + await writeFile(path.join(dir, INFO), meta("same")); + await withMetadataHistory(dir, { by: "prefetch" }, async () => { + await writeFile(path.join(dir, INFO), meta("same")); + }); + assert.equal(await loadMetadataHistory(dir), null); + assert.deepEqual(await readdir(dir), [INFO]); + }); +}); + +test("no prior file records nothing: a first write is not a rewrite", async () => { + await withDir(async (dir) => { + await withMetadataHistory(dir, { by: "prefetch" }, async () => { + await writeFile(path.join(dir, INFO), meta("first")); + }); + assert.equal(await loadMetadataHistory(dir), null); + assert.deepEqual(await readdir(dir), [INFO]); + }); +}); + +test("a run that throws still records the rewrite, and its error comes through", async () => { + await withDir(async (dir) => { + await writeFile(path.join(dir, INFO), meta("before")); + await assert.rejects( + withMetadataHistory(dir, { by: "download-one" }, async () => { + await writeFile(path.join(dir, INFO), meta("after")); + throw new Error("yt-dlp exited 1"); + }), + /yt-dlp exited 1/, + ); + const h = await loadMetadataHistory(dir); + assert.equal(h?.entries.length, 1); + assert.equal(h!.entries[0].by, "download-one"); + assert.deepEqual(h!.entries[0].changed.title, { from: "before", to: "after" }); + }); +}); + +test("a failure to record is logged, never thrown", async () => { + await withDir(async (dir) => { + await writeFile(path.join(dir, INFO), meta("a")); + // A DIRECTORY where the history file goes: the atomic rename onto it fails. + await import("node:fs/promises").then((fs) => + fs.mkdir(metadataHistoryPath(dir)), + ); + let log = ""; + const out = await withMetadataHistory( + dir, + { by: "prefetch", onLog: (s) => (log += s) }, + async () => { + await writeFile(path.join(dir, INFO), meta("b")); + return "ran"; + }, + ); + assert.equal(out, "ran"); + assert.match(log, /Could not record the metadata history/); + }); +}); diff --git a/common/lib/metadataHistory-server.ts b/common/lib/metadataHistory-server.ts @@ -0,0 +1,131 @@ +// THE METADATA HISTORY SIDECAR, and the wrap every metadata.info.json writer +// runs inside. Release 10 slice N; the rules (what is stored, what is only +// fingerprinted, when an entry is written at all) are in metadataHistory.ts. +// +// ONE FILE, `metadata.history.json`, beside the file it describes — not a +// directory. A new directory under data/<id>/ would need the `clips/` treatment +// in every video-dir enumerator (the snapshot, the storage view, the +// reconciler); a bounded file needs none. +// +// IT IS NOT IN discardPrefetchDir's ALLOW-LIST, deliberately +// (downloadOneManaged.ts, PREFETCH_OWN_FILES). A history exists only because a +// metadata.info.json was already in the directory when some pass rewrote it — +// the directory predates the pass now deciding whether to delete it, and the +// history is the one record of what the source used to say. So a title-filter +// rejection of such a directory keeps it and writes the outcome sidecar, which +// is the cheap direction of being wrong (one metadata stub) that the allow-list +// is built on. The case is narrow: a rejection discards its own prefetch dir, +// so this needs a directory left by an earlier pass that was NOT rejected (a +// failed download, say) and a filter that rejects it now. +// +// THE FILE yt-dlp WRITES IS STILL THE FILE. Nothing here changes what lands in +// metadata.info.json; it reads the old bytes before the spawn and the new ones +// after it, and appends what moved. +// +// SERVER-ONLY (node:fs). + +import path from "node:path"; +import { readFile, stat } from "node:fs/promises"; +import { withJsonFileLock } from "./jsonFile-server"; +import { sidecar, sidecarField } from "./sidecar-server"; +import { + METADATA_HISTORY_FILENAME, + appendMetadataHistoryEntry, + buildMetadataHistoryEntry, + coerceMetadataHistory, + type MetadataHistoryEntry, + type MetadataHistoryWriter, +} from "./metadataHistory"; + +const INFO_JSON = "metadata.info.json"; + +export const metadataHistorySidecar = sidecar( + METADATA_HISTORY_FILENAME, + sidecarField(coerceMetadataHistory), +); + +export const { path: metadataHistoryPath, load: loadMetadataHistory } = + metadataHistorySidecar; + +export type MetadataSnapshot = { text: Buffer; mtime: Date }; + +// The info json as it is on disk right now, or null when there is none (or it +// cannot be read — the history is best-effort and never the reason a download +// fails). +export async function snapshotMetadata( + videoDir: string, +): Promise<MetadataSnapshot | null> { + const file = path.join(videoDir, INFO_JSON); + try { + const [text, st] = await Promise.all([readFile(file), stat(file)]); + return { text, mtime: st.mtime }; + } catch { + return null; + } +} + +export type MetadataRewriteContext = { + by: MetadataHistoryWriter; + requestedBy?: string; +}; + +// Compare the file now against `before` and append an entry when it was +// rewritten. Returns the entry, or null when there was nothing to record: no +// old file (a first write is not a rewrite), no new file, or the same bytes. +// +// A READ-MODIFY-WRITE, so it holds the path's lock: two writers finishing at +// once (a prefetch and an audio-check pass on one video are sequential, but a +// second job on the same video is not) must not each append to the file they +// read and drop the other's entry. +export async function recordMetadataRewrite( + videoDir: string, + before: MetadataSnapshot | null, + ctx: MetadataRewriteContext, +): Promise<MetadataHistoryEntry | null> { + if (!before) return null; + const after = await snapshotMetadata(videoDir); + if (!after) return null; + const entry = buildMetadataHistoryEntry({ + before, + after, + at: new Date().toISOString(), + by: ctx.by, + ...(ctx.requestedBy ? { requestedBy: ctx.requestedBy } : {}), + }); + if (!entry) return null; + await withJsonFileLock(metadataHistorySidecar.path(videoDir), async () => { + const existing = await loadMetadataHistory(videoDir); + await metadataHistorySidecar.write( + videoDir, + appendMetadataHistoryEntry(existing, entry), + ); + }); + return entry; +} + +// Run one metadata writer with the history kept: snapshot before `run()`, +// compare and append after it — whether it succeeded OR threw, because a +// yt-dlp that fails late (a prefetch that wrote the file and then hit a +// subtitle error) has still rewritten it. `run`'s own result or error is +// returned unchanged; a failure to record is logged through `onLog` and +// swallowed. +export async function withMetadataHistory<T>( + videoDir: string, + ctx: MetadataRewriteContext & { onLog?: (line: string) => void }, + run: () => Promise<T>, +): Promise<T> { + const before = await snapshotMetadata(videoDir); + try { + return await run(); + } finally { + if (before) { + try { + await recordMetadataRewrite(videoDir, before, ctx); + } catch (err) { + ctx.onLog?.( + `Could not record the metadata history: ${(err as Error).message}\n`, + ); + } + } + } +} diff --git a/common/lib/metadataHistory.test.ts b/common/lib/metadataHistory.test.ts @@ -0,0 +1,215 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + METADATA_HISTORY_CAP, + METADATA_HISTORY_MAX_VALUE_BYTES, + appendMetadataHistoryEntry, + buildMetadataHistoryEntry, + coerceMetadataHistory, + diffMetadata, + isValueDigest, + normalizeForDiff, + type MetadataHistoryEntry, +} from "./metadataHistory"; + +// Run with: +// pnpm --filter yt-dlp-transcript-common exec tsx --test lib/metadataHistory.test.ts + +// A small but real-shaped info json: content, counters, and the volatile keys +// that move on every fetch (signed format URLs, epoch). +function info(over: Record<string, unknown> = {}): Record<string, unknown> { + return { + id: "dbnS-cBgStY", + title: "The original title", + description: "First description.", + duration: 3600, + view_count: 100, + like_count: 7, + formats: [{ format_id: "18", url: "https://example/sig=aaa" }], + thumbnails: [{ url: "https://i.ytimg.com/a.jpg" }], + subtitles: {}, + automatic_captions: { en: [{ url: "https://example/cap=1" }] }, + epoch: 1_700_000_000, + _version: { version: "2026.09.01" }, + ...over, + }; +} + +const text = (o: unknown) => JSON.stringify(o); + +function entry( + before: Record<string, unknown>, + after: Record<string, unknown>, +): MetadataHistoryEntry | null { + return buildMetadataHistoryEntry({ + before: { text: text(before), mtime: new Date("2026-09-01T00:00:00.000Z") }, + after: { text: text(after) }, + at: "2026-09-26T12:00:00.000Z", + by: "prefetch", + }); +} + +test("a formats-only rewrite is an all-volatile entry: nothing changed, nothing stored", () => { + const e = entry( + info(), + info({ + formats: [{ format_id: "18", url: "https://example/sig=bbb" }], + epoch: 1_700_000_999, + }), + ); + assert.ok(e); + assert.deepEqual(e.changed, {}); + assert.deepEqual(e.added, {}); + assert.deepEqual(e.removed, {}); + assert.deepEqual(e.counters, {}); + assert.deepEqual(e.volatile, ["epoch", "formats"]); + // The volatile VALUES are never stored — only that they moved. + assert.ok(!JSON.stringify(e).includes("sig=bbb")); + assert.equal(e.by, "prefetch"); + assert.equal(e.from.at, "2026-09-01T00:00:00.000Z"); + assert.notEqual(e.from.sha256, e.to.sha256); + assert.equal(e.to.bytes, Buffer.byteLength(text(info({ + formats: [{ format_id: "18", url: "https://example/sig=bbb" }], + epoch: 1_700_000_999, + })))); + // Small: the operator asked to be told every rewrite, and this is what it costs. + assert.ok(JSON.stringify(e).length < 400, String(JSON.stringify(e).length)); +}); + +test("a title change is changed.title, both sides whole", () => { + const e = entry(info(), info({ title: "A new title" })); + assert.ok(e); + assert.deepEqual(e.changed, { + title: { from: "The original title", to: "A new title" }, + }); + assert.deepEqual(e.volatile, []); +}); + +test("a new caption language is changed.subtitles_langs, though subtitles itself is volatile", () => { + const e = entry( + info(), + info({ subtitles: { en: [{ url: "https://example/sub=1" }] } }), + ); + assert.ok(e); + assert.deepEqual(e.changed, { + subtitles_langs: { from: [], to: ["en"] }, + }); + assert.deepEqual(e.volatile, ["subtitles"]); +}); + +test("a caption map appearing where there was none is added.<x>_langs", () => { + const before = info(); + delete before.subtitles; + const e = entry(before, info({ subtitles: { de: [] } })); + assert.ok(e); + assert.deepEqual(e.added, { subtitles_langs: ["de"] }); +}); + +test("counters that moved land in counters only, [from, to]", () => { + const e = entry(info(), info({ view_count: 150, like_count: 7 })); + assert.ok(e); + assert.deepEqual(e.counters, { view_count: [100, 150] }); + assert.deepEqual(e.changed, {}); + // A counter the old file lacked reads from null. + const e2 = entry(info(), info({ comment_count: 3 })); + assert.deepEqual(e2?.counters, { comment_count: [null, 3] }); +}); + +test("added and removed content keys", () => { + const before = info({ availability: "public" }); + const after = info({ chapters: [{ title: "Intro", start_time: 0 }] }); + const e = entry(before, after); + assert.ok(e); + assert.deepEqual(e.added, { chapters: [{ title: "Intro", start_time: 0 }] }); + assert.deepEqual(e.removed, { availability: "public" }); +}); + +test("key order alone is not a change, at any depth", () => { + const d = diffMetadata( + { title: "t", chapters: [{ a: 1, b: 2 }], formats: [{ x: 1, y: 2 }] }, + { chapters: [{ b: 2, a: 1 }], formats: [{ y: 2, x: 1 }], title: "t" }, + ); + assert.deepEqual(d, { + changed: {}, + added: {}, + removed: {}, + counters: {}, + volatile: [], + }); +}); + +test("a stored value over 16 KB is replaced by its { sha256, bytes }", () => { + const long = "x".repeat(METADATA_HISTORY_MAX_VALUE_BYTES + 10); + const e = entry(info(), info({ description: long })); + assert.ok(e); + const { from, to } = e.changed.description as { from: unknown; to: unknown }; + assert.equal(from, "First description."); + assert.ok(isValueDigest(to), JSON.stringify(to).slice(0, 80)); + assert.equal((to as { bytes: number }).bytes, long.length + 2); // its JSON: quotes + // And a value at the cap is kept whole. + const atCap = "y".repeat(METADATA_HISTORY_MAX_VALUE_BYTES - 2); + const e2 = entry(info(), info({ description: atCap })); + assert.equal((e2?.changed.description as { to: unknown }).to, atCap); +}); + +test("byte-identical → no entry", () => { + assert.equal(entry(info(), info()), null); +}); + +test("an unparseable side records the rewrite with no diff", () => { + const e = buildMetadataHistoryEntry({ + before: { text: "{ torn" }, + after: { text: text(info()) }, + at: "2026-09-26T12:00:00.000Z", + by: "audio-check", + requestedBy: "mcp", + }); + assert.ok(e); + assert.deepEqual(e.unparseable, ["from"]); + assert.deepEqual(e.changed, {}); + assert.deepEqual(e.added, {}); + assert.equal(e.requestedBy, "mcp"); + assert.equal("at" in e.from, false); +}); + +test("the cap drops the oldest, newest stays last", () => { + let h = null as ReturnType<typeof appendMetadataHistoryEntry> | null; + const mk = (i: number) => + ({ + ...entry(info(), info({ title: `t${i}` }))!, + at: `2026-09-26T12:00:${String(i % 60).padStart(2, "0")}.000Z`, + }) as MetadataHistoryEntry; + for (let i = 0; i < METADATA_HISTORY_CAP + 5; i++) { + h = appendMetadataHistoryEntry(h, mk(i)); + } + assert.equal(h!.entries.length, METADATA_HISTORY_CAP); + assert.deepEqual( + (h!.entries[0].changed.title as { to: string }).to, + "t5", + ); + assert.deepEqual( + (h!.entries.at(-1)!.changed.title as { to: string }).to, + `t${METADATA_HISTORY_CAP + 4}`, + ); + // A smaller cap, as passed. + const small = appendMetadataHistoryEntry(h, mk(999), 3); + assert.equal(small.entries.length, 3); +}); + +test("normalizeForDiff splits content, counters and volatile fingerprints", () => { + const n = normalizeForDiff(info()); + assert.deepEqual(Object.keys(n.counters).sort(), ["like_count", "view_count"]); + assert.deepEqual(n.content.automatic_captions_langs, ["en"]); + assert.deepEqual(n.content.subtitles_langs, []); + assert.equal("formats" in n.content, false); + assert.match(n.volatile.formats, /^[0-9a-f]{64}$/); +}); + +test("coerceMetadataHistory keeps good entries and drops torn ones", () => { + const good = entry(info(), info({ title: "x" }))!; + assert.equal(coerceMetadataHistory(null), null); + assert.equal(coerceMetadataHistory({ entries: "no" }), null); + assert.deepEqual(coerceMetadataHistory({ entries: [good, { at: 1 }, null] }), { + entries: [good], + }); +}); diff --git a/common/lib/metadataHistory.ts b/common/lib/metadataHistory.ts @@ -0,0 +1,372 @@ +// WHAT CHANGED EACH TIME yt-dlp REWROTE A VIDEO'S metadata.info.json. +// +// Release 10 slice N. Every managed download rewrites `data/<id>/metadata.info.json` +// (the metadata prefetch has no existence check, by design: it is what the +// filters decide on), so the file only ever says what the source says TODAY. A +// title that changed upstream, a description that was edited, a caption track +// that appeared — all of it was overwritten without a trace. The history keeps +// one entry per rewrite, beside the video, in `metadata.history.json`. +// +// PURE: no fs. The server half (metadataHistory-server.ts) snapshots the file +// around a yt-dlp spawn and hands both texts here, so every rule below is +// testable without a directory. +// +// ── WHAT IS STORED, AND WHAT IS ONLY FINGERPRINTED ─────────────────────────── +// +// A real info json is ~50 KB and ~83 % of it is `formats`: signed URLs that +// differ on EVERY fetch. Storing old versions whole would be 50 KB of noise per +// rewrite. So the keys fall into three sets: +// +// VOLATILE compared by fingerprint (sha256 of their canonical JSON) and never +// stored — an entry lists which of them moved, nothing more. The +// two caption maps ALSO contribute their language-key lists as +// content keys (`subtitles_langs`, `automatic_captions_langs`), so a +// caption track appearing IS a meaningful change. +// COUNTERS view/like/comment counts: they drift every fetch, and an entry +// records only the ones that moved, as [from, to]. +// CONTENT everything else (title, description, duration, availability, +// chapters, uploader, …): deep-compared on the raw value and stored +// WHOLE on both sides, because a description's from/to is exactly +// what the operator wants to read. A single value over 16 KB is +// replaced by its { sha256, bytes }. +// +// AN ENTRY IS APPENDED ON EVERY REWRITE whose bytes differ — even one where +// only the volatile keys moved. The operator wants to know it was rewritten; +// such an entry is ~250 bytes. Byte-identical → no entry. No old file → no +// entry (a first write is not a rewrite). + +import { createHash } from "node:crypto"; + +export const METADATA_HISTORY_FILENAME = "metadata.history.json"; + +// Newest last; the oldest are dropped past this. Bounded so the file lives with +// the video with no eviction pass of its own. +export const METADATA_HISTORY_CAP = 200; + +// A stored value whose JSON is larger than this is replaced by its digest. +export const METADATA_HISTORY_MAX_VALUE_BYTES = 16 * 1024; + +// Who rewrote the file: the three yt-dlp spawns that write it. +export const METADATA_HISTORY_WRITERS = [ + // downloadOneManaged's attempt 0, on every managed download. + "prefetch", + // The audio-checked primary, which re-extracts on purpose (its restarts + // would outlive the prefetch's pinned format URLs). + "audio-check", + // runYtdlp's legacy single-video `download-one-audio` job. + "download-one", +] as const; +export type MetadataHistoryWriter = (typeof METADATA_HISTORY_WRITERS)[number]; + +export const VOLATILE_KEYS: ReadonlySet<string> = new Set([ + "formats", + "requested_formats", + "requested_downloads", + "thumbnails", + "thumbnail", + "heatmap", + "automatic_captions", + "subtitles", + "epoch", + "_version", + "_format_sort_fields", + "url", + "http_headers", + "format", + "format_id", + "format_note", + "filesize", + "filesize_approx", + "tbr", + "abr", + "vbr", + "asr", + "acodec", + "vcodec", + "fps", + "width", + "height", + "resolution", + "aspect_ratio", + "dynamic_range", + "protocol", + "ext", + "audio_channels", + "container", + "downloader_options", + "quality", + "source_preference", + "has_drm", + "language_preference", +]); + +export const COUNTER_KEYS: ReadonlySet<string> = new Set([ + "view_count", + "like_count", + "dislike_count", + "comment_count", + "channel_follower_count", + "average_rating", + "repost_count", +]); + +// The volatile caption maps whose LANGUAGE LIST is content: the URLs inside +// them are signed and move every fetch, but a language appearing or vanishing +// is a real change to what the video offers. +const LANG_LIST_KEYS: Readonly<Record<string, string>> = { + subtitles: "subtitles_langs", + automatic_captions: "automatic_captions_langs", +}; + +// A file, identified. +export type FileFingerprint = { sha256: string; bytes: number }; + +// What stands in for a stored value over METADATA_HISTORY_MAX_VALUE_BYTES. +export type ValueDigest = { sha256: string; bytes: number }; + +export type MetadataHistoryEntry = { + // When the rewrite was recorded (after the spawn returned). + at: string; + by: MetadataHistoryWriter; + // Who asked for the download, when the caller knew: the saved-video origin's + // requester on a whole-recording fetch ("umtool", "mcp", …). + requestedBy?: string; + // The file before; `at` is its mtime, i.e. when THAT version was written. + from: FileFingerprint & { at?: string }; + to: FileFingerprint; + // Content keys present on both sides whose values differ. + changed: Record<string, { from: unknown; to: unknown }>; + // Content keys only the new file has / only the old one had. + added: Record<string, unknown>; + removed: Record<string, unknown>; + // Counters that moved, [from, to]; null for a side that lacked the key. + counters: Record<string, [unknown, unknown]>; + // Volatile keys whose fingerprint differs (present on one side only counts). + volatile: string[]; + // A side that was not a JSON object (a torn or foreign file). The diff is + // then left empty: comparing against `{}` would list every key as added or + // removed, which is a lie about the source. The fingerprints still say the + // file changed. + unparseable?: Array<"from" | "to">; +}; + +export type MetadataHistory = { entries: MetadataHistoryEntry[] }; + +type Json = unknown; +type JsonObject = Record<string, Json>; + +function isPlainObject(v: unknown): v is JsonObject { + return typeof v === "object" && v !== null && !Array.isArray(v); +} + +function sha256(text: string | Buffer): string { + return createHash("sha256").update(text).digest("hex"); +} + +// JSON with object keys sorted at every depth, so two files that hold the same +// value in a different key order compare equal. `undefined` (an absent key) +// has no JSON and stays undefined. +export function canonicalJson(v: Json): string | undefined { + if (v === undefined) return undefined; + return JSON.stringify(sortKeys(v)); +} + +function sortKeys(v: Json): Json { + if (Array.isArray(v)) return v.map(sortKeys); + if (isPlainObject(v)) { + const out: JsonObject = {}; + for (const k of Object.keys(v).sort()) out[k] = sortKeys(v[k]); + return out; + } + return v; +} + +function jsonEqual(a: Json, b: Json): boolean { + return canonicalJson(a) === canonicalJson(b); +} + +function fingerprint(v: Json): string { + return sha256(canonicalJson(v) ?? ""); +} + +// A value as it is STORED in an entry: itself, or its digest when its JSON is +// over the cap. Measured in UTF-8 bytes, the unit the file is. +export function guardStoredValue(v: Json): Json { + const text = JSON.stringify(v); + if (text === undefined) return null; + const bytes = Buffer.byteLength(text, "utf8"); + if (bytes <= METADATA_HISTORY_MAX_VALUE_BYTES) return v; + const digest: ValueDigest = { sha256: sha256(text), bytes }; + return digest; +} + +// Is this stored value a digest standing in for an oversized one? (A real +// metadata value with exactly these two keys and a 64-hex sha is not a +// plausible yt-dlp field.) +export function isValueDigest(v: unknown): v is ValueDigest { + if (!isPlainObject(v)) return false; + const keys = Object.keys(v); + return ( + keys.length === 2 && + typeof v.sha256 === "string" && + /^[0-9a-f]{64}$/.test(v.sha256) && + typeof v.bytes === "number" + ); +} + +export type NormalizedMetadata = { + content: JsonObject; + counters: JsonObject; + // Volatile key -> fingerprint of its value. + volatile: Record<string, string>; +}; + +// Split one parsed info json into the three sets above. +export function normalizeForDiff(meta: JsonObject): NormalizedMetadata { + const content: JsonObject = {}; + const counters: JsonObject = {}; + const volatile: Record<string, string> = {}; + for (const [key, value] of Object.entries(meta)) { + if (VOLATILE_KEYS.has(key)) { + volatile[key] = fingerprint(value); + const langKey = LANG_LIST_KEYS[key]; + if (langKey && isPlainObject(value)) { + content[langKey] = Object.keys(value).sort(); + } + } else if (COUNTER_KEYS.has(key)) { + counters[key] = value; + } else { + content[key] = value; + } + } + return { content, counters, volatile }; +} + +export type MetadataDiff = Pick< + MetadataHistoryEntry, + "changed" | "added" | "removed" | "counters" | "volatile" +>; + +// Union of both sides' keys: the new file's order first, then keys only the +// old one had — so an entry reads in the order the source writes its fields. +function unionKeys(a: object, b: object): string[] { + const out = Object.keys(b); + const seen = new Set(out); + for (const k of Object.keys(a)) if (!seen.has(k)) out.push(k); + return out; +} + +export function diffMetadata(oldMeta: JsonObject, newMeta: JsonObject): MetadataDiff { + const a = normalizeForDiff(oldMeta); + const b = normalizeForDiff(newMeta); + const changed: MetadataDiff["changed"] = {}; + const added: MetadataDiff["added"] = {}; + const removed: MetadataDiff["removed"] = {}; + for (const key of unionKeys(a.content, b.content)) { + const inOld = Object.hasOwn(a.content, key); + const inNew = Object.hasOwn(b.content, key); + if (inOld && inNew) { + if (!jsonEqual(a.content[key], b.content[key])) { + changed[key] = { + from: guardStoredValue(a.content[key]), + to: guardStoredValue(b.content[key]), + }; + } + } else if (inNew) { + added[key] = guardStoredValue(b.content[key]); + } else { + removed[key] = guardStoredValue(a.content[key]); + } + } + const counters: MetadataDiff["counters"] = {}; + for (const key of unionKeys(a.counters, b.counters)) { + const from = Object.hasOwn(a.counters, key) ? a.counters[key] : null; + const to = Object.hasOwn(b.counters, key) ? b.counters[key] : null; + if (!jsonEqual(from, to)) counters[key] = [from, to]; + } + const volatile = unionKeys(a.volatile, b.volatile) + .filter((k) => a.volatile[k] !== b.volatile[k]) + .sort(); + return { changed, added, removed, counters, volatile }; +} + +function parseObject(text: string | Buffer): JsonObject | null { + try { + const v: unknown = JSON.parse(typeof text === "string" ? text : text.toString("utf8")); + return isPlainObject(v) ? v : null; + } catch { + return null; + } +} + +// One rewrite, as an entry — or null when there is nothing to record (the new +// bytes are the old bytes). +export function buildMetadataHistoryEntry(input: { + before: { text: string | Buffer; mtime?: Date | string }; + after: { text: string | Buffer }; + at: string; + by: MetadataHistoryWriter; + requestedBy?: string; +}): MetadataHistoryEntry | null { + const fromSha = sha256(input.before.text); + const toSha = sha256(input.after.text); + if (fromSha === toSha) return null; + const mtime = input.before.mtime; + const from: MetadataHistoryEntry["from"] = { + sha256: fromSha, + bytes: Buffer.byteLength(input.before.text), + ...(mtime !== undefined + ? { at: typeof mtime === "string" ? mtime : mtime.toISOString() } + : {}), + }; + const to = { sha256: toSha, bytes: Buffer.byteLength(input.after.text) }; + const oldMeta = parseObject(input.before.text); + const newMeta = parseObject(input.after.text); + const unparseable: Array<"from" | "to"> = []; + if (!oldMeta) unparseable.push("from"); + if (!newMeta) unparseable.push("to"); + const diff: MetadataDiff = + oldMeta && newMeta + ? diffMetadata(oldMeta, newMeta) + : { changed: {}, added: {}, removed: {}, counters: {}, volatile: [] }; + return { + at: input.at, + by: input.by, + ...(input.requestedBy ? { requestedBy: input.requestedBy } : {}), + from, + to, + ...diff, + ...(unparseable.length > 0 ? { unparseable } : {}), + }; +} + +// Append, newest last, keeping at most `cap` (the oldest go first). +export function appendMetadataHistoryEntry( + history: MetadataHistory | null, + entry: MetadataHistoryEntry, + cap: number = METADATA_HISTORY_CAP, +): MetadataHistory { + const entries = [...(history?.entries ?? []), entry]; + return { entries: entries.slice(-Math.max(1, cap)) }; +} + +// The sidecar's shape check. An entry that is not recognisably one is dropped +// rather than failing the whole file: one torn entry must not hide the other +// 199 from the page. +export function coerceMetadataHistory(value: unknown): MetadataHistory | null { + if (!isPlainObject(value) || !Array.isArray(value.entries)) return null; + const entries = value.entries.filter( + (e): e is MetadataHistoryEntry => + isPlainObject(e) && + typeof e.at === "string" && + typeof e.by === "string" && + isPlainObject(e.from) && + isPlainObject(e.to) && + isPlainObject(e.changed) && + isPlainObject(e.added) && + isPlainObject(e.removed) && + isPlainObject(e.counters) && + Array.isArray(e.volatile), + ); + return { entries }; +} diff --git a/common/lib/sidecar-server.test.ts b/common/lib/sidecar-server.test.ts @@ -9,6 +9,7 @@ import { SUB_FILE_RE } from "./videoStatus"; import "./attribution-server"; import "./diarization-server"; import "./digest-server"; +import "./metadataHistory-server"; import { availabilitySidecar, loadAvailability, @@ -40,7 +41,7 @@ async function scratch(): Promise<string> { return mkdtemp(path.join(os.tmpdir(), "sidecar-")); } -test("every declared sidecar filename escapes SUB_FILE_RE, and all nine are declared", () => { +test("every declared sidecar filename escapes SUB_FILE_RE, and all ten are declared", () => { assert.deepEqual([...SIDECAR_FILENAMES].sort(), [ "ai-digest.json", "ai-digest.overrides.json", @@ -50,6 +51,7 @@ test("every declared sidecar filename escapes SUB_FILE_RE, and all nine are decl "do-not-clean.json", "download-outcome.json", "exclude-truncated-check.json", + "metadata.history.json", "transcribe-outcome.json", ]); for (const name of SIDECAR_FILENAMES) { diff --git a/common/views/metadataHistoryView.test.ts b/common/views/metadataHistoryView.test.ts @@ -0,0 +1,129 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + METADATA_VALUE_DISPLAY_CHARS, + NOTHING_MEANINGFUL, + metadataHistoryView, + relativeAgo, + valueView, +} from "./metadataHistoryView"; +import type { MetadataHistoryEntry } from "../lib/metadataHistory"; + +function e(over: Partial<MetadataHistoryEntry>): MetadataHistoryEntry { + return { + at: "2026-09-26T12:00:00.000Z", + by: "prefetch", + from: { sha256: "a".repeat(64), bytes: 10 }, + to: { sha256: "b".repeat(64), bytes: 11 }, + changed: {}, + added: {}, + removed: {}, + counters: {}, + volatile: ["epoch", "formats"], + ...over, + }; +} + +const NOW = Date.parse("2026-09-26T12:05:00.000Z"); + +test("no history, or an empty one, draws nothing", () => { + assert.equal(metadataHistoryView(null, NOW), null); + assert.equal(metadataHistoryView({ entries: [] }, NOW), null); +}); + +test("the line: count, how lately, by whom, and what moved", () => { + const v = metadataHistoryView( + { + entries: [ + e({ at: "2026-09-25T12:00:00.000Z" }), + e({ + at: "2026-09-26T12:00:00.000Z", + changed: { title: { from: "Old", to: "New" } }, + counters: { view_count: [10, 12] }, + }), + ], + }, + NOW, + ); + assert.ok(v); + assert.equal(v.count, 2); + assert.equal(v.line, "Metadata rewritten 2× · last 5m ago by prefetch: title"); + // Newest first. + assert.deepEqual( + v.entries.map((x) => x.at), + ["2026-09-26T12:00:00.000Z", "2026-09-25T12:00:00.000Z"], + ); + assert.equal(v.entries[0].atLabel, "2026-09-26 12:00 UTC"); + assert.deepEqual( + v.entries[0].changes.map((c) => [c.key, c.kind, c.from?.text, c.to?.text]), + [ + ["title", "changed", "Old", "New"], + ["view_count", "counter", "10", "12"], + ], + ); + assert.equal(v.entries[1].summary, NOTHING_MEANINGFUL); +}); + +test("counters alone are named; an all-volatile rewrite says nothing meaningful", () => { + const counters = metadataHistoryView( + { entries: [e({ counters: { view_count: [1, 2], like_count: [null, 3] } })] }, + NOW, + )!; + assert.equal( + counters.line, + "Metadata rewritten 1× · last 5m ago by prefetch: only counters (view_count, like_count)", + ); + assert.equal(counters.entries[0].changes[1].from, null); + const none = metadataHistoryView({ entries: [e({})] }, NOW)!; + assert.match(none.line, /: nothing meaningful \(formats only\)$/); + assert.deepEqual(none.entries[0].volatile, ["epoch", "formats"]); +}); + +test("added and removed keys, requestedBy, and the unreadable case", () => { + const v = metadataHistoryView( + { + entries: [ + e({ + by: "audio-check", + requestedBy: "umtool", + added: { chapters: [{ title: "Intro" }] }, + removed: { availability: "public" }, + }), + e({ unparseable: ["from"], volatile: [] }), + ], + }, + NOW, + )!; + assert.equal(v.entries[1].summary, "chapters, availability"); + assert.equal(v.entries[1].requestedBy, "umtool"); + assert.deepEqual( + v.entries[1].changes.map((c) => [c.kind, c.from?.text ?? null, c.to?.text ?? null]), + [ + ["added", null, '[{"title":"Intro"}]'], + ["removed", "public", null], + ], + ); + assert.equal(v.entries[0].summary, "unreadable metadata (no diff)"); + assert.equal(v.entries[0].unparseable, true); +}); + +test("long values are cut at 300 with the whole text kept; digests read as such", () => { + const long = "d".repeat(METADATA_VALUE_DISPLAY_CHARS + 50); + const cut = valueView(long); + assert.equal(cut.truncated, true); + assert.equal(cut.text.length, METADATA_VALUE_DISPLAY_CHARS + 1); + assert.ok(cut.text.endsWith("…")); + assert.equal(cut.full, long); + const short = valueView("short"); + assert.deepEqual(short, { text: "short", full: "short", truncated: false }); + const digest = valueView({ sha256: "c".repeat(64), bytes: 20000 }); + assert.match(digest.text, /^\(20000 bytes, too long to keep; sha256 cccccccccccc…\)$/); +}); + +test("relativeAgo steps s → m → h → d, and a future stamp is just now", () => { + assert.equal(relativeAgo(4_000), "4s ago"); + assert.equal(relativeAgo(5 * 60_000), "5m ago"); + assert.equal(relativeAgo(3 * 3_600_000), "3h ago"); + assert.equal(relativeAgo(2 * 86_400_000), "2d ago"); + assert.equal(relativeAgo(-10), "just now"); +}); diff --git a/common/views/metadataHistoryView.ts b/common/views/metadataHistoryView.ts @@ -0,0 +1,155 @@ +import { + isValueDigest, + type MetadataHistory, + type MetadataHistoryEntry, +} from "../lib/metadataHistory"; + +// THE VIDEO PAGE'S METADATA HISTORY: one line, and the entries behind it. +// +// Release 10 slice N. `metadata.history.json` records every rewrite of the +// video's metadata.info.json (lib/metadataHistory.ts); this turns the loaded +// sidecar into what the page draws — a line saying how often and how lately it +// was rewritten and what moved, and the entries newest first, each key with its +// from → to. Read-only: nothing here writes, and there is no control. +// +// PURE, like everything in this directory: the history and the clock arrive +// as arguments, so "5m ago" is testable. +// +// ONE SENTENCE FOR "NOTHING MOVED". Almost every rewrite changes only the +// signed format URLs and the counters; the line must say so plainly rather +// than list `formats`, which means nothing to an operator. Content keys are +// the news; counters are named only when nothing else moved. + +// A value longer than this is cut on the page, with the whole text in `full` +// (the renderer's `title=`). +export const METADATA_VALUE_DISPLAY_CHARS = 300; + +export type MetadataValueView = { + text: string; + full: string; + truncated: boolean; +}; + +export type MetadataKeyChangeView = { + key: string; + kind: "changed" | "added" | "removed" | "counter"; + // null for the side that lacked the key (an added key has no `from`). + from: MetadataValueView | null; + to: MetadataValueView | null; +}; + +export type MetadataHistoryEntryView = { + // ISO, as stored. + at: string; + // "2026-09-26 19:40 UTC" — fixed, so a server render never depends on the + // host's locale. + atLabel: string; + by: string; + requestedBy: string | null; + // The entry's changed keys, or the sentence for none. + summary: string; + changes: MetadataKeyChangeView[]; + // Volatile keys that moved (by fingerprint only; never shown as values). + volatile: string[]; + // A side was not JSON, so the entry carries no diff. + unparseable: boolean; +}; + +export type MetadataHistoryView = { + count: number; + // `Metadata rewritten N× · last <relative time> by <by>: <summary>`. + line: string; + // Newest first. + entries: MetadataHistoryEntryView[]; +}; + +export const NOTHING_MEANINGFUL = "nothing meaningful (formats only)"; + +// Compact relative time, the steps the editor's widget uses (s → m → h → d). +export function relativeAgo(deltaMs: number): string { + if (!Number.isFinite(deltaMs) || deltaMs < 0) return "just now"; + const s = Math.round(deltaMs / 1000); + if (s < 60) return `${s}s ago`; + if (s < 3600) return `${Math.round(s / 60)}m ago`; + if (s < 86400) return `${Math.round(s / 3600)}h ago`; + return `${Math.round(s / 86400)}d ago`; +} + +function atLabel(iso: string): string { + const d = new Date(iso); + if (Number.isNaN(d.getTime())) return iso; + return `${d.toISOString().slice(0, 16).replace("T", " ")} UTC`; +} + +export function valueView(v: unknown): MetadataValueView { + const full = isValueDigest(v) + ? `(${v.bytes} bytes, too long to keep; sha256 ${v.sha256.slice(0, 12)}…)` + : typeof v === "string" + ? v + : (JSON.stringify(v) ?? "null"); + const truncated = full.length > METADATA_VALUE_DISPLAY_CHARS; + return { + text: truncated ? `${full.slice(0, METADATA_VALUE_DISPLAY_CHARS)}…` : full, + full, + truncated, + }; +} + +function changesOf(e: MetadataHistoryEntry): MetadataKeyChangeView[] { + const out: MetadataKeyChangeView[] = []; + for (const [key, c] of Object.entries(e.changed)) { + out.push({ key, kind: "changed", from: valueView(c.from), to: valueView(c.to) }); + } + for (const [key, v] of Object.entries(e.added)) { + out.push({ key, kind: "added", from: null, to: valueView(v) }); + } + for (const [key, v] of Object.entries(e.removed)) { + out.push({ key, kind: "removed", from: valueView(v), to: null }); + } + for (const [key, pair] of Object.entries(e.counters)) { + const [from, to] = Array.isArray(pair) ? pair : [null, null]; + out.push({ + key, + kind: "counter", + from: from === null ? null : valueView(from), + to: to === null ? null : valueView(to), + }); + } + return out; +} + +export function entrySummary(e: MetadataHistoryEntry): string { + const content = [ + ...Object.keys(e.changed), + ...Object.keys(e.added), + ...Object.keys(e.removed), + ]; + if (content.length > 0) return content.join(", "); + const counters = Object.keys(e.counters); + if (counters.length > 0) return `only counters (${counters.join(", ")})`; + if (e.unparseable?.length) return "unreadable metadata (no diff)"; + return NOTHING_MEANINGFUL; +} + +export function metadataHistoryView( + history: MetadataHistory | null, + nowMs: number, +): MetadataHistoryView | null { + const stored = history?.entries ?? []; + if (stored.length === 0) return null; + const entries: MetadataHistoryEntryView[] = [...stored].reverse().map((e) => ({ + at: e.at, + atLabel: atLabel(e.at), + by: e.by, + requestedBy: e.requestedBy ?? null, + summary: entrySummary(e), + changes: changesOf(e), + volatile: [...e.volatile], + unparseable: Boolean(e.unparseable?.length), + })); + const last = entries[0]; + const line = + `Metadata rewritten ${stored.length}× · last ` + + `${relativeAgo(nowMs - Date.parse(last.at))} by ${last.by}: ${last.summary}`; + return { count: stored.length, line, entries }; +} diff --git a/common/ytdlp/downloadOneManaged.ts b/common/ytdlp/downloadOneManaged.ts @@ -39,6 +39,7 @@ import { } from "../lib/downloadOutcome"; import { writeDownloadOutcome } from "../lib/downloadOutcome-server"; import { recordAvailability } from "../lib/availability-server"; +import { withMetadataHistory } from "../lib/metadataHistory-server"; import { loadRawMetadata, loadRawMetadataFromDir, @@ -141,6 +142,16 @@ export type ManagedDownloadOpts = { // Per-run override: extract audio now and discard the container even for a // video the keep-latest rule would otherwise persist (the save-disk backfill). extractImmediately?: boolean; + // THE CALLER WANTS THE MEDIA (release 10 slice N): a transcript or captions + // on disk are not a reason to skip the download. Only youtube handling ever + // skipped it — its media pass is the no-subs fallback (attempt 3), gated on + // "no transcript and no captions" — so this opens that gate. With a + // transcript on disk the forced pass touches no transcript and extracts no + // audio: it persists the container and nothing else (see attempt 3). Set by + // "Persist source video", the whole-recording fetch and "Persist kept now"; + // never by the download lane, re-acquire or import. Transcribe handling + // already downloads, so it is unaffected. + forceMedia?: boolean; // Per-run override of channelConfig.audioFormat for the extracted audio. audioFormatOverride?: AudioFormat; // Resolved yt-dlp `-f` download-format preset (override > channel > global), @@ -235,6 +246,11 @@ async function finalizeAppExtraction(opts: { category: PersistenceDecision["category"]; // Threaded straight onto the pointer; see ManagedDownloadOpts.persistOrigin. origin?: SavedVideoOrigin; + // PERSIST ONLY when false: move the container, extract nothing. A forced + // media download of a video that already has a transcript (attempt 3, + // `keepTranscript`) wants the source kept, and an audio.<fmt> beside an + // existing transcript is bytes nothing will read. Default true. + extractAudio?: boolean; onLog: (s: string) => void; signal: AbortSignal; }): Promise<void> { @@ -246,25 +262,37 @@ async function finalizeAppExtraction(opts: { ); return; } - try { - await transcodeAudio({ - paths: opts.paths, - videoDir: opts.videoDir, - sourceFilename: source, - targetFormat: opts.fmt, - onLog: opts.onLog, - signal: opts.signal, - }); - } catch (err) { + if (opts.extractAudio === false) { opts.onLog( - `App extraction failed (${source} -> audio.${opts.fmt}): ${(err as Error).message}. Keeping the source container.\n`, + `Persist only: a transcript is on disk, so no audio.${opts.fmt} is extracted from ${source}.\n`, ); - return; - } - if (!opts.persist) { - await rm(path.join(opts.videoDir, source), { force: true }); - opts.onLog(`Discarded source container ${source} (audio-only).\n`); - return; + if (!opts.persist) { + // Unreachable from attempt 3 today (it does not download when the plan + // persists nothing), kept so the option means one thing on its own. + opts.onLog(`Not persisting ${source}; it stays in the data dir.\n`); + return; + } + } else { + try { + await transcodeAudio({ + paths: opts.paths, + videoDir: opts.videoDir, + sourceFilename: source, + targetFormat: opts.fmt, + onLog: opts.onLog, + signal: opts.signal, + }); + } catch (err) { + opts.onLog( + `App extraction failed (${source} -> audio.${opts.fmt}): ${(err as Error).message}. Keeping the source container.\n`, + ); + return; + } + if (!opts.persist) { + await rm(path.join(opts.videoDir, source), { force: true }); + opts.onLog(`Discarded source container ${source} (audio-only).\n`); + return; + } } // Persist: move the container into the saved-video store + write a pointer. const storeDir = savedVideoDir( @@ -376,6 +404,9 @@ const PREFETCH_OWN_FILES: ReadonlySet<string> = new Set([ // one a previous rejection wrote before this rule existed. It records what // happened, never what is on disk, so it is ours to drop with the rest. "download-outcome.json", + // NOT metadata.history.json, deliberately: a history means an earlier + // metadata.info.json was here before this pass, and it is the one record of + // what the source used to say (lib/metadataHistory-server.ts). ]); async function discardPrefetchDir( @@ -655,10 +686,21 @@ async function runManagedDownload( "--", opts.videoUrl, ]; - const prefetchRes = await runOneYtdlp( - opts, - channelDir, - buildPrefetchArgs(prefetchCookies), + // EVERY PREFETCH REWRITES metadata.info.json (no existence check — the + // filters decide on today's metadata), so each spawn runs inside the + // history wrap: what moved since the last version is appended to + // metadata.history.json. The file yt-dlp writes is still the file. + const prefetchHistory = { + by: "prefetch" as const, + ...(opts.persistOrigin?.requestedBy + ? { requestedBy: opts.persistOrigin.requestedBy } + : {}), + onLog: opts.onLog, + }; + const prefetchRes = await withMetadataHistory( + videoDir, + prefetchHistory, + () => runOneYtdlp(opts, channelDir, buildPrefetchArgs(prefetchCookies)), ); lastFullTail = prefetchRes.stderrTail; const prefetchAvail = attemptSucceeded(prefetchRes.exitCode) @@ -694,10 +736,11 @@ async function runManagedDownload( opts.onLog( `Metadata prefetch auth-required (${prefetchAvail}); retrying with --cookies-from-browser ${prefetchRetryCookies}\n`, ); - const retryRes = await runOneYtdlp( - opts, - channelDir, - buildPrefetchArgs(prefetchRetryCookies), + const retryRes = await withMetadataHistory( + videoDir, + prefetchHistory, + () => + runOneYtdlp(opts, channelDir, buildPrefetchArgs(prefetchRetryCookies)), ); lastFullTail = retryRes.stderrTail; const retryAvail = attemptSucceeded(retryRes.exitCode) @@ -1027,16 +1070,34 @@ async function runManagedDownload( // output to data/<canonicalId>/ for every platform, so the canonical id is // the dir. const expectedVideoIdHint = canonicalId; - const audioOutcome = await runAudioCheckedYtdlp({ - paths: opts.paths, - channelDir, - channelConfig: opts.channelConfig, - ytdlpArgs: primaryArgs, - onLog: opts.onLog, - signal: opts.signal, - archiveMarker: ARCHIVE_MARKER, - expectedVideoIdHint, - }); + const runAudioCheck = () => + runAudioCheckedYtdlp({ + paths: opts.paths, + channelDir, + channelConfig: opts.channelConfig, + ytdlpArgs: primaryArgs, + onLog: opts.onLog, + signal: opts.signal, + archiveMarker: ARCHIVE_MARKER, + expectedVideoIdHint, + }); + // The audio-checked primary writes the info json itself (it re-extracts on + // purpose), so it is the second rewrite of the file in one download and + // gets its own history entry. Only when the dir is known up front — an + // unidentifiable URL lands in data/%(id)s/, found only afterwards. + const audioOutcome = canonicalId + ? await withMetadataHistory( + path.join(channelDir, "data", canonicalId), + { + by: "audio-check", + ...(opts.persistOrigin?.requestedBy + ? { requestedBy: opts.persistOrigin.requestedBy } + : {}), + onLog: opts.onLog, + }, + runAudioCheck, + ) + : await runAudioCheck(); primaryRes = { exitCode: audioOutcome.ytdlpExitCode, stderrTail: audioOutcome.stderrTail, @@ -1225,6 +1286,11 @@ async function runManagedDownload( } // ---------- Attempt 3: no-subs fallback (youtube handling only) ---------- + // Also the youtube-handling MEDIA download a `forceMedia` caller asked for: + // this is the only pass that fetches media for a youtube-handling channel, + // and its own gate ("no transcript and no captions") is exactly the reason + // "Persist source video" and the whole-recording fetch used to do nothing on + // a video that already had a transcript — two subtitle passes, no file. if ( lastSucceeded && opts.channelConfig.handling === "youtube" && @@ -1232,10 +1298,40 @@ async function runManagedDownload( ) { const hasTranscript = await hasAnyTranscriptOnDisk(videoDir); const noCaptions = await metadataReportsNoCaptions(videoDir); - if (!hasTranscript && noCaptions) { + const noSubsFallback = !hasTranscript && noCaptions; + // Forced only when the fallback would NOT have run on its own: a video with + // no transcript and no captions takes today's path whoever asked. + const forced = !noSubsFallback && opts.forceMedia === true; + // THE TRANSCRIPT ON DISK IS NEVER TOUCHED. A forced pass over a video that + // has one fetches the source and persists it, and does nothing else: no + // subtitles (refused on the command line, after the channel's own args, so + // a channel carrying --write-auto-subs cannot overwrite the transcript), no + // audio extraction, no transcription, no short-audio verdict on audio it + // never made. With NO transcript on disk (captions listed, none fetched — + // a language the channel's sub-langs does not match) the forced pass is + // today's fallback in full: download, extract, transcribe inline if + // configured, persist. + const keepTranscript = forced && hasTranscript; + if (keepTranscript && !plan.persist) { + // PERSIST-ONLY WITH NOTHING TO PERSIST. Every forceMedia caller sets the + // keep-source override, so the plan persists; a caller that did not would + // be asking for a download with no destination — the transcript rules + // out audio, and the plan rules out the container. opts.onLog( - `No subs available for ${videoId}; falling back to audio download + whisper${plan.persist ? " (keeping source video)" : ""}.\n`, + `forceMedia: a transcript is on disk and this download keeps no source video (${plan.reason}); nothing to fetch.\n`, ); + } else if (noSubsFallback || forced) { + if (forced) { + opts.onLog( + `forceMedia: downloading the source although ${ + hasTranscript ? "a transcript is on disk" : "captions exist" + }\n`, + ); + } else { + opts.onLog( + `No subs available for ${videoId}; falling back to audio download + whisper${plan.persist ? " (keeping source video)" : ""}.\n`, + ); + } const fallbackConfig: ChannelConfig = { ...opts.channelConfig, handling: "transcribe", @@ -1267,6 +1363,9 @@ async function runManagedDownload( "--print", `after_video:${ARCHIVE_MARKER} %(extractor)s %(id)s`, ...channelConfigArgs(fallbackConfig, fallbackCookieOverride), + // yt-dlp takes the LAST occurrence of an option, so the refusals go + // after the channel's own args (the chat-only pass's rule). + ...(keepTranscript ? ["--no-write-subs", "--no-write-auto-subs"] : []), ...sourceArgs( opts.videoUrl, path.join(videoDir, "metadata.info.json"), @@ -1290,7 +1389,27 @@ async function runManagedDownload( }); if (fallbackRes.archiveLine) lastArchiveLine = fallbackRes.archiveLine; - if (attemptSucceeded(fallbackRes.exitCode)) { + if (attemptSucceeded(fallbackRes.exitCode) && keepTranscript) { + // Persist only. Not a fallback to transcription — the transcript is + // the one already on disk — so `fellBackToTranscribe` stays unset and + // the status stays what the subtitle pass made it. + if (plan.extractionMode === "app") { + await finalizeAppExtraction({ + paths: opts.paths, + channelSlug: opts.channelSlug, + channelConfig: opts.channelConfig, + videoDir, + videoId, + fmt, + persist: plan.persist, + category: plan.category, + origin: opts.persistOrigin, + extractAudio: false, + onLog: opts.onLog, + signal: opts.signal, + }); + } + } else if (attemptSucceeded(fallbackRes.exitCode)) { fellBackToTranscribe = true; // In app mode, extract audio.<fmt> from the downloaded source container // (and keep/discard it) before any inline whisper can read the audio. diff --git a/common/ytdlp/forceMedia.test.ts b/common/ytdlp/forceMedia.test.ts @@ -0,0 +1,302 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + chmod, + mkdir, + mkdtemp, + readFile, + readdir, + rm, + writeFile, +} from "node:fs/promises"; +import { tmpdir } from "node:os"; +import path from "node:path"; +import type { Paths } from "../lib/paths"; +import type { ChannelConfig } from "../lib/channelConfig"; +import { downloadOneManaged, type ManagedDownloadOpts } from "./downloadOneManaged"; + +// Run with: +// pnpm --filter yt-dlp-transcript-common exec tsx --test ytdlp/forceMedia.test.ts +// +// THE WHOLE-RECORDING FETCH DOWNLOADS ANYWAY (release 10 slice N). On a +// youtube-handling channel the only pass that fetches media is attempt 3, the +// no-subs fallback, gated on "no transcript and no captions" — so "Persist +// source video" and `fetch_clip` with `full: true` were a silent no-op on any +// video that already had a transcript: two subtitle passes, no file. +// `forceMedia` opens the gate; with a transcript on disk the forced pass +// persists the container and does nothing else. The yt-dlp here is a temp +// node script that logs every argv and writes what the real one would; no +// network. + +const ID = "H64QQZuw-aA"; +const VIDEO = `https://www.youtube.com/watch?v=${ID}`; + +// The fake: the metadata prefetch writes meta.json's content as the info json; +// the subtitle pass writes nothing (the transcript is seeded); the media pass +// (the `source-media` output template) writes the container. Every spawn +// appends its argv, one JSON line, to `argv.log`. +const FAKE_YTDLP = `#!/usr/bin/env node +const fs = require("node:fs"); +const path = require("node:path"); +const argv = process.argv.slice(2); +const root = process.env.FAKE_ROOT; +fs.appendFileSync(path.join(root, "argv.log"), JSON.stringify(argv) + "\\n"); +const has = (f) => argv.includes(f); +const arg = (f) => { const i = argv.indexOf(f); return i < 0 ? undefined : argv[i + 1]; }; +const info = arg("--load-info-json"); +const id = info ? path.basename(path.dirname(info)) : "${ID}"; +const dir = path.join("data", id); +fs.mkdirSync(dir, { recursive: true }); +if (has("--skip-download") && has("--write-info-json") && has("--no-write-subs")) { + fs.copyFileSync(path.join(root, "meta.json"), path.join(dir, "metadata.info.json")); + process.exit(0); +} +if (has("--skip-download") && has("--write-auto-subs")) { + console.log("DLOM_ARCHIVE youtube " + id); + process.exit(0); +} +if ((arg("-o") || "").includes("source-media")) { + fs.writeFileSync(path.join(dir, "source-media.mp4"), "container bytes for " + id); + console.log("DLOM_ARCHIVE youtube " + id); + process.exit(0); +} +console.error("fake: unknown invocation " + argv.join(" ")); +process.exit(2); +`; + +// ffmpeg for the app extraction: writes its last argument (the temp output). +const FAKE_FFMPEG = `#!/bin/sh +for last; do :; done +echo "extracted audio" > "$last" +`; + +type Run = { + argvs: string[][]; + record: Awaited<ReturnType<typeof downloadOneManaged>>; + log: string; + videoDir: string; + storeDir: string; + files: string[]; +}; + +async function runWith(opts: { + // The info json the prefetch writes. + meta: Record<string, unknown>; + // Seed a transcript before the run. + transcript?: { name: string; body: string }; + // Seed an info json before the run (so the prefetch is a REWRITE). + priorMeta?: Record<string, unknown>; + config?: Partial<ChannelConfig>; + managed?: Partial<ManagedDownloadOpts>; +}): Promise<Run & { cleanup: () => Promise<void> }> { + const root = await mkdtemp(path.join(tmpdir(), "force-media-")); + const ytdlp = path.join(root, "fake-ytdlp.cjs"); + const ffmpeg = path.join(root, "fake-ffmpeg.sh"); + await writeFile(ytdlp, FAKE_YTDLP); + await writeFile(ffmpeg, FAKE_FFMPEG); + await chmod(ytdlp, 0o755); + await chmod(ffmpeg, 0o755); + await writeFile(path.join(root, "meta.json"), JSON.stringify(opts.meta)); + process.env.FAKE_ROOT = root; + const paths = { + transcriptsDir: root, + channelsDir: path.join(root, "channels"), + savedVideosDir: path.join(root, "saved-videos"), + ytdlpBin: ytdlp, + ffmpegBin: ffmpeg, + ffprobeBin: path.join(root, "no-ffprobe"), + } as Paths; + const videoDir = path.join(paths.channelsDir, "c", "data", ID); + await mkdir(videoDir, { recursive: true }); + if (opts.transcript) { + await writeFile(path.join(videoDir, opts.transcript.name), opts.transcript.body); + } + if (opts.priorMeta) { + await writeFile( + path.join(videoDir, "metadata.info.json"), + JSON.stringify(opts.priorMeta), + ); + } + let log = ""; + const record = await downloadOneManaged({ + channelSlug: "c", + channelConfig: { + handling: "youtube", + url: "https://www.youtube.com/@c/videos", + ...opts.config, + } as ChannelConfig, + paths, + videoUrl: VIDEO, + onLog: (s) => { + log += s; + }, + signal: new AbortController().signal, + ...opts.managed, + }); + const argvs = (await readFile(path.join(root, "argv.log"), "utf8")) + .split("\n") + .filter(Boolean) + .map((l) => JSON.parse(l) as string[]); + return { + argvs, + record, + log, + videoDir, + storeDir: path.join(paths.savedVideosDir, "c", ID), + files: (await readdir(videoDir)).sort(), + cleanup: () => rm(root, { recursive: true, force: true }), + }; +} + +const META = { + id: ID, + title: "A video", + duration: 60, + webpage_url: VIDEO, + subtitles: {}, + automatic_captions: { en: [{ url: "https://example/cap" }] }, +}; +const TRANSCRIPT = { name: "transcript.json", body: '{"segments":[{"text":"kept"}]}' }; +const FORCED: Partial<ManagedDownloadOpts> = { + keepSourceVideoOverride: true, + forceMedia: true, + persistOrigin: { requestedBy: "mcp", requestedAt: "2026-09-26T00:00:00.000Z" }, +}; + +test("forceMedia with a transcript on disk: the source is fetched and persisted, nothing else", async () => { + const r = await runWith({ meta: META, transcript: TRANSCRIPT, managed: FORCED }); + try { + // Prefetch, the subtitle pass as every re-download runs it, then the media. + assert.equal(r.argvs.length, 3); + assert.ok(r.argvs[1].includes("--skip-download")); + const media = r.argvs[2]; + assert.ok(!media.includes("--skip-download"), media.join(" ")); + assert.ok(media.includes("--load-info-json")); + assert.deepEqual( + media.slice(media.indexOf("-f"), media.indexOf("-f") + 2), + ["-f", "bestvideo*+bestaudio/best"], + ); + // The transcript cannot be rewritten, even by a channel's own args: the + // refusals come after everything the channel passes. + assert.ok(media.indexOf("--no-write-auto-subs") > media.indexOf("--print")); + assert.ok(media.includes("--no-write-subs")); + assert.match(r.log, /forceMedia: downloading the source although a transcript is on disk\n/); + + // The pointer landed and the container is in the store. + const pointer = JSON.parse( + await readFile(path.join(r.videoDir, "saved-video.json"), "utf8"), + ) as Record<string, unknown>; + assert.equal(pointer.keepReason, "override"); + assert.equal((pointer.origin as { requestedBy: string }).requestedBy, "mcp"); + assert.equal( + await readFile(path.join(r.storeDir, "source-media.mp4"), "utf8"), + `container bytes for ${ID}`, + ); + // The transcript is byte-for-byte what it was, and no audio was made. + assert.equal( + await readFile(path.join(r.videoDir, TRANSCRIPT.name), "utf8"), + TRANSCRIPT.body, + ); + assert.deepEqual( + r.files.filter((f) => f.startsWith("audio.") || f.startsWith("source-media")), + [], + ); + assert.match(r.log, /Persist only: a transcript is on disk, so no audio\.mp3 is extracted/); + // Not a fallback to transcription, and a clean status. + assert.equal(r.record.status, "ok"); + assert.equal(r.record.fellBackToTranscribe, undefined); + assert.deepEqual( + r.record.attempts.map((a) => [a.kind, a.ytdlpExitCode]), + [ + ["metadata-prefetch", 0], + ["primary", 0], + ["no-subs-fallback", 0], + ], + ); + } finally { + await r.cleanup(); + } +}); + +test("without forceMedia the same video is untouched: two subtitle-shaped passes, no file (today's behaviour)", async () => { + const r = await runWith({ + meta: META, + transcript: TRANSCRIPT, + managed: { keepSourceVideoOverride: true }, + }); + try { + assert.equal(r.argvs.length, 2); + assert.ok(r.argvs.every((a) => a.includes("--skip-download"))); + assert.ok(!r.files.includes("saved-video.json")); + assert.doesNotMatch(r.log, /forceMedia/); + } finally { + await r.cleanup(); + } +}); + +test("forceMedia with captions listed but no transcript: today's fallback in full (download, extract, persist)", async () => { + const r = await runWith({ meta: META, managed: FORCED }); + try { + assert.equal(r.argvs.length, 3); + const media = r.argvs[2]; + assert.ok(!media.includes("--skip-download")); + // No transcript to protect, so no refusals — the args are the fallback's. + assert.ok(!media.includes("--no-write-auto-subs")); + assert.match(r.log, /forceMedia: downloading the source although captions exist\n/); + assert.ok(r.files.includes("saved-video.json")); + assert.ok(r.files.includes("audio.mp3"), r.files.join(",")); + assert.equal(r.record.fellBackToTranscribe, true); + } finally { + await r.cleanup(); + } +}); + +test("no transcript and no captions: the fallback runs as it always has, forced or not", async () => { + const noCaps = { ...META, automatic_captions: {} }; + const r = await runWith({ meta: noCaps, managed: FORCED }); + try { + assert.equal(r.argvs.length, 3); + assert.match(r.log, /No subs available for .*falling back to audio download \+ whisper \(keeping source video\)/); + assert.doesNotMatch(r.log, /forceMedia/); + assert.ok(r.files.includes("audio.mp3")); + } finally { + await r.cleanup(); + } +}); + +test("forceMedia with a transcript but a plan that persists nothing fetches nothing", async () => { + const r = await runWith({ + meta: META, + transcript: TRANSCRIPT, + managed: { forceMedia: true, keepSourceVideoOverride: false }, + }); + try { + assert.equal(r.argvs.length, 2); + assert.match(r.log, /forceMedia: a transcript is on disk and this download keeps no source video/); + assert.equal(r.record.status, "ok"); + } finally { + await r.cleanup(); + } +}); + +test("the prefetch's rewrite of metadata.info.json lands in metadata.history.json, with who asked", async () => { + const r = await runWith({ + meta: { ...META, title: "Renamed upstream", view_count: 12 }, + priorMeta: { ...META, view_count: 10 }, + transcript: TRANSCRIPT, + managed: FORCED, + }); + try { + const h = JSON.parse( + await readFile(path.join(r.videoDir, "metadata.history.json"), "utf8"), + ) as { entries: Array<Record<string, unknown>> }; + assert.equal(h.entries.length, 1); + const e = h.entries[0]; + assert.equal(e.by, "prefetch"); + assert.equal(e.requestedBy, "mcp"); + assert.deepEqual(e.changed, { title: { from: "A video", to: "Renamed upstream" } }); + assert.deepEqual(e.counters, { view_count: [10, 12] }); + } finally { + await r.cleanup(); + } +}); diff --git a/common/ytdlp/runYtdlp.ts b/common/ytdlp/runYtdlp.ts @@ -30,6 +30,7 @@ import { type ResolvedCookiePolicy, } from "../lib/cookiePolicy"; import { resolveEffectiveAvailability } from "../lib/availability-server"; +import { withMetadataHistory } from "../lib/metadataHistory-server"; import { settledByTitleFilterIds, upsertMetadataScan, @@ -288,8 +289,8 @@ export function outputArgsForUrl( opts: { mediaName?: string } = {}, ): string[] { const mediaName = opts.mediaName ?? "audio"; - const id = extractVideoId(url); - if (id && /^[\w.-]+$/.test(id) && id !== "." && id !== "..") { + const id = dataDirIdForUrl(url); + if (id) { return [ "-o", `data/${id}/${mediaName}.%(ext)s`, @@ -314,6 +315,14 @@ export function outputArgsForUrl( return OUTPUT_ARGS; } +// The data/<id>/ directory outputArgsForUrl pins a URL's output to, or null +// when it falls back to yt-dlp's %(id)s (no canonical id, or one that is not a +// safe directory name) — i.e. the directory is only known afterwards. +export function dataDirIdForUrl(url: string): string | null { + const id = extractVideoId(url); + return id && /^[\w.-]+$/.test(id) && id !== "." && id !== ".." ? id : null; +} + // Thrown by enumeratePlaylistUrls when a listing is rate-limited part-way. // // A full enumeration that the platform rate-limited part-way through. What it @@ -1377,7 +1386,20 @@ async function downloadOneAudio(opts: RunYtdlpOpts): Promise<void> { "--", opts.singleVideoUrl, ]; - await runChildAndStream(opts, root, args); + // --write-info-json rewrites the video's metadata.info.json; the history + // keeps what moved (lib/metadataHistory-server.ts). Only when the output dir + // is pinned — a %(id)s dir is known only after the run. + const dirId = dataDirIdForUrl(opts.singleVideoUrl); + const run = () => runChildAndStream(opts, root, args); + if (dirId) { + await withMetadataHistory( + path.join(root, "data", dirId), + { by: "download-one", onLog: opts.onLog }, + run, + ); + } else { + await run(); + } await safeBackfillAvailability(opts); } diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -7,6 +7,8 @@ - **`/jobs` says why a job was cancelled at boot.** A job the boot settled shows its reason under its status on `/jobs` and as *Cancelled because* on its own page: for example "server restarted; the scheduler re-derives syncs" or "superseded by a newer queued job (…)". The reason used to be only in the job's log. - **The server log says how often a queued job skips its page refresh.** When a queued job finishes outside any request, the editor skips its page refresh and notes it in the log. The note used to appear once and never again. Now the first one after a quiet spell is logged at once, any more in the next 10 minutes are counted, and one line at the end gives the count, with a running total. - **An agent working through the MCP server asks the editor for a clip instead of running yt-dlp.** The MCP server has a new tool, `fetch_clip`. Given a citation's channel, video id, start and end and a one-line reason, it asks the local editor for that window through `POST /api/media/fetch-window`: the same paced, cookie-aware job umtool uses, which records who asked and why beside the file. It answers with the file's path in the corpus (`channels/<slug>/data/<id>/clips/`). The window is the cited span with 3 seconds either side, at most 15 minutes. `full: true` asks for the whole recording instead, which lands in the saved-video store and needs a video the editor already knows. A Rumble citation's id (the embed id the archive publishes) is mapped to the id the editor names the video's folder by, through the archive record's link. The editor must already archive the channel: pointed at a public site with a fresh editor, every clip gets a 404 `Channel "<slug>" not found`. The tool waits up to 90 seconds by default (at most 300) and otherwise returns the job's id, to wait on with `job`; the fetch carries on in the editor either way. While it waits it sends a progress notification per poll to a client that asks for progress. A client whose requests time out at 60 seconds (the MCP SDK's default) must raise that or pass `wait_seconds` of 50 or less. No request to the editor waits more than 15 seconds. If the editor stops answering mid-fetch, the answer gives the job's id and says not to ask again from scratch. The `/ask` and `/sweep` plans now tell the agent to use it and never to run yt-dlp itself. The MCP needs `ARCHILYZER_EDITOR_URL` and `WORKER_TOKEN` (the editor's own) in its environment, so re-register it with the two `--env` lines in the README; without them the tool says so and fetches nothing. The MCP server itself still writes nothing. The README's `yt-dlp --download-sections` command is now only the fallback for a machine with no editor. +- **"Persist source video" and a whole-recording fetch download the video even when it already has a transcript.** On a channel that takes YouTube's subtitles, the button — and a whole-recording request from umtool or the MCP server's `fetch_clip` with `full: true`, which run the same job — fetched only the subtitles again when the video already had a transcript or captions, and finished with no file. It now downloads the source and moves it into the saved-video store. YouTube's subtitles are fetched again first, as on any re-download; a Whisper transcript is not touched, and no audio is extracted beside a transcript. **Persist kept now** on a channel's Cleanup stage does the same for every kept video, so on such a channel it now downloads each kept video's source. The video page's Source video card offers the button on these channels too; it used to say persistence was for transcribe-handling channels only. +- **Each video keeps a history of how its metadata changed at the source.** Every download that rewrites a video's `metadata.info.json` and changes anything in it adds one entry to `metadata.history.json` beside it: the old and new value of each field that changed (title, description, duration, availability, chapters and the rest), the view, like and comment counts that moved, and which of the fields that change on every fetch (format URLs, thumbnails, caption URLs) differed, compared by fingerprint only. A caption language appearing or disappearing counts as a change. The newest 200 entries are kept. The video page shows the history under the description: "Metadata rewritten N× · last … by …: <what changed>", with each entry's old → new values when opened. The history starts with the first rewrite after this update. ## [0.9.0] - 2026-09-26 - **Every page now has a ground and an accent to choose, and the five theme families are gone.** The theme menu (the palette button beside the quick toggle, in the editor's sidebar and in the header of every published site, the hub and the homepage) has two groups. **Base** is System, Light, Sepia or Dark; Sepia is new, a warm paper ground for long reading. **Accent** is Signal, Brass, Vermilion, Violet, Sakura, Blue or Green, with the site's own tagged *default*; a site with a custom hex offers it first as *Site colour*. The quick toggle cycles System → Light → Sepia → Dark. A published site opens on the reader's system setting, in the accent its site form sets. The hub and the homepage open on Dark, in Signal, even with JavaScript off, and the editor follows the system, in Signal. Each accent has a value for each ground that reads at 4.5:1, and a custom hex is darkened or lightened per ground to match. A reader's accent is remembered only while it differs from the site's: picking the site's own again forgets it, so the reader follows the site if its accent changes later. Base, Archive, Selenized, Swiss and Archilyzer are gone. A choice made before this update carries over once: light stays light (Archive light becomes Sepia), dark stays dark and system stays system; the family itself is dropped. Headings are Archivo, text is IBM Plex Sans and figures are IBM Plex Mono everywhere, with one corner radius. Success, warning and other status text reads at 4.5:1 on its own tinted fill on every ground; on Light, success and warning are a shade deeper than before for it. Chart colours are fixed per ground and never follow the accent; the third is a violet, well clear of the red that marks a recording as gone. The phone's browser bar takes the page's ground, not the accent. Needs a rebuild and deploy of every site, the hub and the homepage. diff --git a/editor/app/channels/[slug]/videos/[id]/components/MetadataHistoryDetails.tsx b/editor/app/channels/[slug]/videos/[id]/components/MetadataHistoryDetails.tsx @@ -0,0 +1,81 @@ +import { Fragment } from "react"; +import type { + MetadataHistoryView, + MetadataKeyChangeView, + MetadataValueView, +} from "yt-dlp-transcript-common/views/metadataHistoryView"; + +// WHAT CHANGED EACH TIME THE METADATA WAS REWRITTEN (release 10 slice N). +// +// Drawn in the page header, beside the Description, because it is about the +// same thing: the video's metadata as the source reports it. One line — how +// often, how lately, by what, and what moved — and the entries behind it, +// newest first. Read-only and server-rendered: the view +// (views/metadataHistoryView.ts) is already folded, so there is no client +// state and no control. + +const KIND_NOTE: Record<MetadataKeyChangeView["kind"], string> = { + changed: "", + added: " (new)", + removed: " (gone)", + counter: "", +}; + +function Value({ v }: { v: MetadataValueView | null }) { + if (!v) return <span className="text-muted-foreground">—</span>; + return ( + <span + className="whitespace-pre-wrap break-words" + title={v.truncated ? v.full : undefined} + > + {v.text} + </span> + ); +} + +export function MetadataHistoryDetails({ view }: { view: MetadataHistoryView }) { + return ( + <details className="text-sm" aria-label="metadata history"> + <summary className="cursor-pointer text-muted-foreground hover:text-foreground"> + {view.line} + </summary> + <ol className="mt-2 flex flex-col gap-3" aria-label="metadata history entries"> + {view.entries.map((e, i) => ( + <li + key={`${e.at}-${i}`} + className="flex flex-col gap-1 rounded border border-border px-3 py-2" + > + <div className="text-muted-foreground"> + <span className="tabular-nums">{e.atLabel}</span> + {` · by ${e.by}`} + {e.requestedBy ? ` for ${e.requestedBy}` : ""} + {` · ${e.summary}`} + </div> + {e.changes.length > 0 && ( + <dl className="grid grid-cols-[auto_1fr] gap-x-3 gap-y-1"> + {e.changes.map((c) => ( + <Fragment key={`${c.kind}-${c.key}`}> + <dt className="font-mono text-xs text-muted-foreground pt-0.5"> + {c.key} + {KIND_NOTE[c.kind]} + </dt> + <dd className="min-w-0"> + <Value v={c.from} /> + <span className="text-muted-foreground"> → </span> + <Value v={c.to} /> + </dd> + </Fragment> + ))} + </dl> + )} + {e.volatile.length > 0 && ( + <div className="text-xs text-muted-foreground"> + Also moved (compared by fingerprint only): {e.volatile.join(", ")} + </div> + )} + </li> + ))} + </ol> + </details> + ); +} diff --git a/editor/app/channels/[slug]/videos/[id]/components/cards/SourceVideoSection.tsx b/editor/app/channels/[slug]/videos/[id]/components/cards/SourceVideoSection.tsx @@ -37,14 +37,11 @@ export function SourceVideoSection({ // video is already persisted, so the DOM is unchanged. const [ranHere, setRanHere] = useState(false); - if (handling !== "transcribe") { - return ( - <p className="text-sm text-muted-foreground"> - Source-video persistence applies to transcribe-handling channels only. - </p> - ); - } - + // EVERY CHANNEL GETS THE CONTROLS (release 10 slice N). A youtube-handling + // channel used to be told persistence was for transcribe handling only — + // which was true in effect, because its one media pass was skipped whenever a + // transcript or captions existed. The run now forces that pass (forceMedia), + // so the button does what it says on every channel. return ( <div className="flex flex-col gap-3"> {savedVideo ? ( @@ -53,6 +50,13 @@ export function SourceVideoSection({ videoId={videoId} savedVideo={savedVideo} /> + ) : handling === "youtube" ? ( + <p className="text-sm text-muted-foreground"> + Download this video&apos;s full source container and move it into + the saved-video store. YouTube&apos;s subtitles are fetched again + first, as on any re-download; a Whisper transcript is not touched, + and no audio is extracted beside a transcript. + </p> ) : ( <p className="text-sm text-muted-foreground"> Re-fetch this video&apos;s full source container and move it into the diff --git a/editor/app/channels/[slug]/videos/[id]/lib/videoChoreCards.ts b/editor/app/channels/[slug]/videos/[id]/lib/videoChoreCards.ts @@ -113,7 +113,7 @@ export const VIDEO_CHORE_CARDS: readonly VideoChoreCard[] = [ { id: "source-video-persistence", component: SourceVideoSection, - shown: "always (transcribe-handling channels get the controls)", + shown: "always, with its controls on every channel (youtube handling too since release 10 slice N)", why: "a storage move into the saved-video store; not derived from the transcript", }, { diff --git a/editor/app/channels/[slug]/videos/[id]/page.tsx b/editor/app/channels/[slug]/videos/[id]/page.tsx @@ -7,6 +7,8 @@ import type { Dirent } from "node:fs"; import { readChannelConfig } from "yt-dlp-transcript-common/controller/channels"; import { loadDownloadOutcome } from "yt-dlp-transcript-common/lib/downloadOutcome-server"; import { loadAvailability } from "yt-dlp-transcript-common/lib/availability-server"; +import { loadMetadataHistory } from "yt-dlp-transcript-common/lib/metadataHistory-server"; +import { metadataHistoryView } from "yt-dlp-transcript-common/views/metadataHistoryView"; import { listClipWindows } from "yt-dlp-transcript-common/lib/clipWindow-server"; import { isDoNotClean } from "yt-dlp-transcript-common/lib/doNotClean-server"; import { isExcludedFromTruncatedCheck } from "yt-dlp-transcript-common/lib/excludeTruncatedCheck-server"; @@ -37,6 +39,7 @@ import { RunningJobsList } from "../../../../jobs/components/RunningJobsList"; import { VideoPanel, type VideoFile } from "./components/VideoPanel"; import { OperationPanel } from "./components/OperationPanel"; import { TagsPanel } from "./components/TagsPanel"; +import { MetadataHistoryDetails } from "./components/MetadataHistoryDetails"; import { loadVideoTags } from "./lib/videoTags"; import { loadVideoOperationPanels } from "./lib/videoOperationPanels"; @@ -96,6 +99,12 @@ export default async function VideoDetailPage({ const downloadOutcome = await loadDownloadOutcome(videoDir); const availabilityRecord = await loadAvailability(videoDir); const availabilityHistory = availabilityRecord?.history ?? []; + // Every rewrite of metadata.info.json that changed its bytes (release 10 + // slice N). One small bounded file, read once; null → nothing drawn. + const metadataHistory = metadataHistoryView( + await loadMetadataHistory(videoDir), + Date.now(), + ); const doNotClean = await isDoNotClean(videoDir); const excludedFromTruncatedCheck = await isExcludedFromTruncatedCheck(videoDir); @@ -210,6 +219,7 @@ export default async function VideoDetailPage({ </p> </details> )} + {metadataHistory && <MetadataHistoryDetails view={metadataHistory} />} </header> <RunningJobsList jobs={activeJobs} hideChannelSlug hideVideoId /> diff --git a/editor/app/channels/[slug]/videos/[id]/videoActions.ts b/editor/app/channels/[slug]/videos/[id]/videoActions.ts @@ -234,9 +234,12 @@ export async function downloadVideoPipelineAction( // Re-fetch an already-downloaded video purely to archive its SOURCE container, // without disturbing the existing transcript. Forces the persistence rule on -// (keepSourceVideoOverride=true) so downloadOneManaged downloads the full video, -// app-extracts audio, and keeps the container (Phase 3 moves it to the saved -// store). Works on a video already in the archive — downloadOneManaged has no +// (keepSourceVideoOverride=true) so downloadOneManaged downloads the full video +// and keeps the container (Phase 3 moves it to the saved store), and forces the +// media download itself (forceMedia) on a youtube-handling channel, whose only +// media pass is otherwise skipped when a transcript or captions exist; there, +// with a transcript on disk, no audio is extracted — the container is the +// point. Works on a video already in the archive — downloadOneManaged has no // archive prefilter, so it always re-downloads. export async function redownloadToArchiveAction( slug: string, @@ -300,6 +303,11 @@ async function archiveSourceVideo( globalSkipLiveDownloads: settings.skipLiveDownloads, appendArchive: true, keepSourceVideoOverride: true, + // The operator (or the tool behind a whole-recording fetch) asked for + // the SOURCE: a transcript or captions on disk are not a reason to + // skip it. Without this a youtube-handling video with a transcript + // got two subtitle passes and no file (release 10 slice N). + forceMedia: true, persistOrigin, }); safeRevalidate([ diff --git a/editor/e2e/fetch-window.spec.ts b/editor/e2e/fetch-window.spec.ts @@ -9,10 +9,11 @@ // Same token as /api/worker/* (the test server runs with // WORKER_TOKEN=test-worker-token; see package.json dev:test). -import { readFile, stat } from "node:fs/promises"; +import { readdir, readFile, stat } from "node:fs/promises"; import { test, expect, type APIRequestContext } from "@playwright/test"; import { generateReport, + readJson, resetData, resolvePath, writeChannelConfig, @@ -395,3 +396,62 @@ test("a full-source fetch lands in the saved-video store and is cached on a seco expect(Number(cached.bytes)).toBe(Number(pointer.bytes)); expect(await invocations()).toBe(invBefore); }); + +// THE SAME ASK ON THE CHANNEL AS IT IS (release 10 slice N). The test above +// switches the channel to transcribe handling, because that used to be the only +// way the container reached the store: on this subtitles-first channel the one +// media pass (the no-subs fallback) was skipped for any video with a +// transcript, so the job ended `done` with no file — the silent no-op slice M's +// live proof found. `forceMedia` opens that pass; with a transcript on disk it +// persists the container and extracts nothing. +test("a full-source fetch on a youtube-handling video that already has a transcript downloads anyway", async ({ + request, +}) => { + test.setTimeout(90_000); + // The fixture's downloaded video: metadata.info.json + transcript.en.vtt. + const HAS_TRANSCRIPT = "fake00000001"; + const post = await request.post(`${baseUrl}/api/media/fetch-window`, { + headers: AUTH, + data: { + channelSlug: SLUG, + videoId: HAS_TRANSCRIPT, + full: true, + requestedBy: "mcp", + manifest: "demo-report", + reason: "the whole recording, captions or not", + }, + }); + expect(post.status(), await post.text()).toBe(202); + const { jobId } = (await post.json()) as { jobId: string }; + + const finished = await pollJob(request, jobId); + expect(finished.status, JSON.stringify(finished)).toBe("done"); + // A FILE, which is the whole difference. + expect(String(finished.file)).toMatch(/source-media\./); + expect(Number(finished.bytes)).toBeGreaterThan(0); + + const dir = resolvePath(rel(`data/${HAS_TRANSCRIPT}`)); + const files = await readdir(dir); + expect(files).toContain("transcript.en.vtt"); + expect(files).toContain("saved-video.json"); + // Nothing extracted beside the transcript. + expect(files.filter((f) => /^(audio|source-media)\./.test(f))).toEqual([]); + + // AND THE METADATA REWRITE IS ON RECORD, with who asked. The fixture's own + // info json differs from what the fake's prefetch writes (another title), so + // this fetch is a rewrite that changed content. + const history = await readJson<{ + entries: Array<{ + by: string; + requestedBy?: string; + changed: Record<string, { from: unknown; to: unknown }>; + }>; + }>(rel(`data/${HAS_TRANSCRIPT}/metadata.history.json`)); + expect(history.entries).toHaveLength(1); + expect(history.entries[0].by).toBe("prefetch"); + expect(history.entries[0].requestedBy).toBe("mcp"); + expect(history.entries[0].changed.title).toEqual({ + from: "Synthetic Test Video 1", + to: `Synthetic ${HAS_TRANSCRIPT}`, + }); +}); diff --git a/editor/e2e/fixtures/bin/fake-ytdlp.mjs b/editor/e2e/fixtures/bin/fake-ytdlp.mjs @@ -132,12 +132,30 @@ async function writeMetadata(videoDir, id, opts = {}) { meta.was_live = true; meta.live_status = "was_live"; } + Object.assign(meta, await readMetadataOverrides()); await writeFile( path.join(videoDir, "metadata.info.json"), JSON.stringify(meta), ); } +// THE SOURCE CHANGED ITS METADATA. A spec that needs the second fetch of a +// video to differ from the first (the metadata history, release 10 slice N) +// writes `.fake-ytdlp-metadata.json` in the channel root (the fake's cwd): a +// JSON object whose keys are laid over every info json this fake writes from +// then on — `{ "title": "Renamed", "view_count": 150 }`. Absent (every other +// spec), nothing changes and the output stays byte-stable. +async function readMetadataOverrides() { + try { + const parsed = JSON.parse(await readFile(".fake-ytdlp-metadata.json", "utf8")); + return parsed && typeof parsed === "object" && !Array.isArray(parsed) + ? parsed + : {}; + } catch { + return {}; + } +} + // Sentinels encoded in a video id/URL select metadata variants so a fixture // can deterministically exercise each filter/fallback branch. function urlSentinels(url) { diff --git a/editor/e2e/persist-youtube-handling.spec.ts b/editor/e2e/persist-youtube-handling.spec.ts @@ -0,0 +1,175 @@ +// THE WHOLE-RECORDING FETCH DOWNLOADS ANYWAY, AND EVERY METADATA REWRITE KEEPS +// THE OLD VERSION (release 10 slice N). +// +// On a youtube-handling channel the only pass that fetches media is the no-subs +// fallback, and its gate is "no transcript and no captions". So "Persist source +// video" (and `fetch_clip` with `full: true`, the same function) on a video that +// already had a transcript ran two subtitle passes and produced no file — the +// live proof of slice M found it on teamrcn. No spec covered it: every persist +// spec used a transcribe-handling fixture. This one uses the youtube fixture's +// already-downloaded video, transcript and all. +// +// The second test is the history: two downloads of one video where the source's +// metadata changed in between (the fake's `.fake-ytdlp-metadata.json` knob), +// and the video page saying so. + +import { readFile, readdir, stat, writeFile } from "node:fs/promises"; +import { test, expect } from "@playwright/test"; +import { readJson, resetData, resolvePath, writeSettings } from "./helpers"; +import { baseUrl } from "./baseUrl"; + +const SLUG = "test-youtube"; +const rel = (p: string) => `test-transcripts/channels/${SLUG}/${p}`; +// The fixture's downloaded video: metadata.info.json + transcript.en.vtt. +const HAS_TRANSCRIPT = "fake00000001"; +// In the playlist, never downloaded. +const UNDOWNLOADED = "fake00000002"; + +type Outcome = { + status: string; + startedAt: string; + fellBackToTranscribe?: boolean; + attempts: Array<{ kind: string; ytdlpExitCode: number | null }>; +}; + +async function outcomeOf(id: string): Promise<Outcome | null> { + try { + return await readJson<Outcome>(rel(`data/${id}/download-outcome.json`)); + } catch { + return null; + } +} + +test("Persist source video on a youtube-handling video with a transcript downloads the source and leaves the transcript alone", async ({ + page, +}) => { + test.setTimeout(120_000); + await resetData("youtube-with-playlist"); + // Inline whisper ON, so the claim "nothing but the container" bites: had the + // forced pass extracted audio and fallen through to transcription, the + // whisper transcript below would have been rewritten. + await writeSettings({ inlineTranscribeOnFallback: true }); + const whisper = JSON.stringify({ + transcription: [{ text: "the transcript already on disk" }], + }); + const whisperPath = resolvePath(rel(`data/${HAS_TRANSCRIPT}/transcript.json`)); + await writeFile(whisperPath, whisper); + await fetch(`${baseUrl}/api/test/invalidate-cache`).catch(() => {}); + + await page.goto(`/channels/${SLUG}/videos/${HAS_TRANSCRIPT}`); + await page.getByLabel("Source video stage summary").click(); + // The card used to say persistence was for transcribe-handling channels only + // and offered no button here at all. + const run = page.getByRole("button", { + name: "Persist source video", + exact: true, + }); + await expect(run).toBeEnabled(); + await run.click(); + + const log = page.getByLabel(`Persist source video for ${HAS_TRANSCRIPT} output`); + await expect(log).toContainText( + "forceMedia: downloading the source although a transcript is on disk", + { timeout: 30_000 }, + ); + await expect(page.getByLabel(`unpersist source video ${HAS_TRANSCRIPT}`)).toBeVisible({ + timeout: 30_000, + }); + await expect(log).toContainText("Persisted source video to"); + + // yt-dlp was asked for MEDIA: the fake reaches its source-media branch only + // for an invocation without --skip-download. + const inv = await readFile(resolvePath(rel("fake-ytdlp.invocations")), "utf8"); + expect(inv).toContain( + `download-source-media:https://www.youtube.com/watch?v=${HAS_TRANSCRIPT}`, + ); + + // The pointer landed, naming a real file in the store. + const pointer = await readJson<{ + dir: string; + file: string; + bytes: number; + keepReason?: string; + }>(rel(`data/${HAS_TRANSCRIPT}/saved-video.json`)); + expect(pointer.keepReason).toBe("override"); + expect(pointer.file).toMatch(/^source-media\./); + expect((await stat(`${pointer.dir}/${pointer.file}`)).size).toBe(pointer.bytes); + + // The transcript is byte-for-byte what it was, and nothing was extracted + // beside it: no audio, no container left in the data dir. + expect(await readFile(whisperPath, "utf8")).toBe(whisper); + const files = await readdir(resolvePath(rel(`data/${HAS_TRANSCRIPT}`))); + expect(files.filter((f) => /^(audio|source-media)\./.test(f))).toEqual([]); + + // A clean download, and not a fallback to transcription. + const outcome = await outcomeOf(HAS_TRANSCRIPT); + expect(outcome?.status).toBe("ok"); + expect(outcome?.fellBackToTranscribe).toBeUndefined(); + expect(outcome?.attempts.map((a) => a.kind)).toEqual([ + "metadata-prefetch", + "primary", + "no-subs-fallback", + ]); +}); + +test("a download whose metadata changed upstream appends to metadata.history.json, and the video page says so", async ({ + page, +}) => { + test.setTimeout(120_000); + await resetData("youtube-with-playlist"); + const knob = resolvePath(rel(".fake-ytdlp-metadata.json")); + const historyRel = rel(`data/${UNDOWNLOADED}/metadata.history.json`); + + // First download: a first write of metadata.info.json is not a rewrite. + await writeFile(knob, JSON.stringify({ view_count: 100 })); + await page.goto(`/channels/${SLUG}/videos/${UNDOWNLOADED}`); + await page.getByRole("button", { name: /^Run download pipeline$/ }).click(); + await expect.poll(() => outcomeOf(UNDOWNLOADED), { timeout: 30_000 }).not.toBeNull(); + const first = (await outcomeOf(UNDOWNLOADED))!; + expect(first.status).toBe("ok"); + await expect(readJson(historyRel)).rejects.toThrow(); + + // The source renames the video and it gains views; download again. + await writeFile( + knob, + JSON.stringify({ + title: `Synthetic ${UNDOWNLOADED} (renamed upstream)`, + view_count: 150, + }), + ); + await page.reload(); + await page.getByRole("button", { name: /^Run download pipeline$/ }).click(); + await expect + .poll(async () => (await outcomeOf(UNDOWNLOADED))?.startedAt, { + timeout: 30_000, + }) + .not.toBe(first.startedAt); + + type History = { + entries: Array<{ + by: string; + changed: Record<string, { from: unknown; to: unknown }>; + counters: Record<string, [unknown, unknown]>; + }>; + }; + const history = await readJson<History>(historyRel); + expect(history.entries).toHaveLength(1); + expect(history.entries[0].by).toBe("prefetch"); + expect(history.entries[0].changed).toEqual({ + title: { + from: `Synthetic ${UNDOWNLOADED}`, + to: `Synthetic ${UNDOWNLOADED} (renamed upstream)`, + }, + }); + expect(history.entries[0].counters).toEqual({ view_count: [100, 150] }); + + // The page: one line, and the entry behind it with its from → to. + await page.goto(`/channels/${SLUG}/videos/${UNDOWNLOADED}`); + const line = page.getByText(/^Metadata rewritten 1× · last .* by prefetch: title$/); + await expect(line).toBeVisible(); + await line.click(); + const entries = page.getByLabel("metadata history entries"); + await expect(entries).toContainText(`Synthetic ${UNDOWNLOADED} (renamed upstream)`); + await expect(entries).toContainText("view_count"); + await expect(entries).toContainText("100 → 150"); +}); diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -6757,11 +6757,22 @@ S4, as shipped"; `release-10.md` "Slice L2 / L1, as shipped". Every anchor below the editor's `fetchFullSourceAction` uses the saved-video store (`transcripts/saved-videos/<slug>/ <id>/source-media.<ext>`, pointer `data/<id>/saved-video.json` with `keepReason: "override"` and `origin`), job kind `redownload-archive`, and takes NO URL (404 when the video has neither - metadata nor a playlist entry). It is a silent no-op on a youtube-handling video that already has - a transcript (`downloadOneManaged.ts:1233-1235`: the keep-source override reaches only transcribe - handling or the no-captions fallback) and it rewrites `metadata.info.json`; the window path never - writes metadata. Neither path dedupes a running job: a repeated request while one runs queues a - second fetch, which is why every text says "resume with job". + metadata nor a playlist entry). It WAS a silent no-op on a youtube-handling video that already had + a transcript (the keep-source override reached only transcribe handling or the no-captions + fallback) — **fixed by release 10 slice N**: `archiveSourceVideo` and `persistKept` pass + `forceMedia: true`, so the no-subs fallback (attempt 3, `downloadOneManaged.ts` `forced` / + `keepTranscript`) runs whenever the subtitle pass succeeded; with a transcript on disk it adds + `--no-write-subs --no-write-auto-subs` after the channel's args and finalizes persist-only + (`finalizeAppExtraction` `extractAudio: false`: no `audio.<fmt>`, no whisper, no short-audio + probe, `fellBackToTranscribe` unset). The primary subtitle pass still runs as on any re-download, + and yt-dlp re-writes an existing subtitle file by default (`YoutubeDL.existing_file`, + `default_overwrite=True`, with `overwrites` popped when unset — verified in the installed + `yt-dlp-patched` 2026.08.19), so a `transcript.<lang>.vtt` gets a new mtime; a whisper + `transcript.json` is untouched. It rewrites `metadata.info.json` (every managed download's + prefetch does); since slice N each rewrite that changes the bytes appends to + `metadata.history.json` (`common/lib/metadataHistory.ts`). The window path never writes metadata. + Neither path dedupes a running job: a repeated request while one runs queues a second fetch, which + is why every text says "resume with job". - The editor must already HAVE the channel (404 `Channel "<slug>" not found` otherwise); a public-only setup gets the "no editor configured" error naming both env vars, and the README's `yt-dlp --download-sections` is the no-editor fallback only. diff --git a/plans/release-10.md b/plans/release-10.md @@ -1166,6 +1166,242 @@ nit N1 taken). **Gates** (`m-gate4.log`): tsc clean (whole workspace); **mcp 269/269**, 16 s. +### Slice N, as shipped — the whole-recording fetch downloads anyway; metadata history (2026-09-26) + +Branch `editor/full-fetch-media` off `main` `555bc454`, worktree +`/home/user/Projects/editor-full-fetch-media`, one Opus implementer. The plan is +[`full-fetch-forces-media.md`](full-fetch-forces-media.md) (Decisions 1–6, Implementation 1–8). It +answers the operator's ask of 2026-09-26: the whole-recording fetch downloads whatever is on disk, +and every rewrite of the metadata is kept so changes can be detected. No settings, site or channel +key; no new job kind; no new directory. + +**`forceMedia`** (`common/ytdlp/downloadOneManaged.ts`). +- One option on `ManagedDownloadOpts` (:154): the caller wants the media; a transcript or captions + on disk are not a reason to skip. `archiveSourceVideo` (`videoActions.ts:310`, so both "Persist + source video" and `fetchFullSourceAction`) and `persistKept` (`persistKept.ts:137`) set it. The + download lane, re-acquire and import do not. +- **Attempt 3** (:1285-): `forced = !noSubsFallback && forceMedia` (:1304) — a video with no + transcript and no captions takes today's path whoever asked. The pass then runs when the primary + succeeded, and logs one line first: `forceMedia: downloading the source although a transcript is + on disk` (or `… although captions exist`). +- **`keepTranscript = forced && hasTranscript`** (:1314) makes it persist-only: + - `--no-write-subs --no-write-auto-subs` after the channel's own args (:1368), so a channel whose + `ytdlpExtraArgs` carry `--write-auto-subs` cannot overwrite the transcript (the chat-only + pass's last-occurrence rule); + - `finalizeAppExtraction` gains `extractAudio?: boolean` (:253); `false` moves the container into + the store and extracts nothing, logging `Persist only: a transcript is on disk, so no + audio.<fmt> is extracted from <source>.`; + - no inline whisper, no short-audio probe (it would probe audio the pass never made); + `fellBackToTranscribe` stays unset and the status stays the subtitle pass's (`ok`); the attempt + is still recorded as `no-subs-fallback`, n: 3. +- With `forceMedia` and NO transcript (captions listed, none fetched), the pass is today's fallback + in full: download, extract, transcribe inline if configured, persist. +- A caller that does not set `forceMedia` gets byte-for-byte today's behaviour: the same gate, log, + args and finalize. Every existing e2e arg-ordering assertion passed unchanged. + +**Metadata history.** +- `common/lib/metadataHistory.ts` (pure; `node:crypto` only): the key sets exactly as the plan + lists them; `normalizeForDiff`, `diffMetadata`, `buildMetadataHistoryEntry` (null when the sha is + unchanged), `appendMetadataHistoryEntry` (cap 200, oldest dropped), the 16 KB guard + (`guardStoredValue` → `{ sha256, bytes }`, `isValueDigest`), `coerceMetadataHistory` (drops a torn + entry instead of failing the file). Volatile values are compared on canonical (key-sorted) JSON. +- `common/lib/metadataHistory-server.ts`: `sidecar("metadata.history.json", …)` (indented, trailing + newline); `snapshotMetadata`; `recordMetadataRewrite`, a read-modify-write under + `withJsonFileLock`; and `withMetadataHistory(videoDir, {by, requestedBy, onLog}, run)`. That + snapshots before `run()` and records in a `finally`, so a failed run is recorded too. It returns + or rethrows `run`'s own result, and logs and swallows a failure to record. +- Wrapped, with `requestedBy` = `persistOrigin.requestedBy` where there is one: + - the prefetch spawn and its auth retry (`by: "prefetch"`, :700); + - the audio-check primary when the canonical id is known (`by: "audio-check"`, :1088); + - `downloadOneAudio` when its output dir is pinned (`by: "download-one"`, + `runYtdlp.ts:1395`). + `dataDirIdForUrl` (`runYtdlp.ts:321`) is the id test `outputArgsForUrl` already made, now named. +- **Not on `discardPrefetchDir`'s allow-list**, said at `PREFETCH_OWN_FILES` (:407) and in the + server module's head. +- `SIDECAR_FILENAMES`' enumeration test: ten names. + +**The surface.** +- `common/views/metadataHistoryView.ts` (pure, the clock as `nowMs`): `metadataHistoryView(history, + nowMs)` → `{count, line, entries}` newest first, or null. + - The line is `Metadata rewritten N× · last <5m ago> by <by>: <summary>`. + - Values are cut at 300 characters, with the full text kept for `title=`. + - `atLabel` is fixed UTC. +- `page.tsx` reads the sidecar once beside `availability.json`. The new server component + `components/MetadataHistoryDetails.tsx` draws it in the page header, under the Description: + a `<details>` whose summary is the line, with each entry's keys as from → to and the volatile + keys named. No client state and no control. + +**Where the code departs from, or fills in, the plan** (for the reviewer): +- **The "Persist source video" button had to be un-gated.** `SourceVideoSection` returned + "Source-video persistence applies to transcribe-handling channels only." for any other handling, + so the plan's e2e 6(a) — click the button on a youtube-handling video — had no button to click. + - The gate is gone. + - A youtube channel gets its own description line: "…YouTube's subtitles are fetched again + first, as on any re-download; a Whisper transcript is not touched, and no audio is extracted + beside a transcript." (review M1: the first wording said the transcript was "kept", which a + re-fetched VTT is not). + - Labels, aria-labels and test ids are unchanged. + - `videoChoreCards.ts`'s `shown` text is updated. +- **`keepTranscript` with a plan that persists nothing fetches nothing.** The plan said "do not + assume `plan.persist`". With a transcript the pass may extract no audio, so a plan that keeps no + container leaves the download with no destination. It logs `forceMedia: a transcript is on disk + and this download keeps no source video (<reason>); nothing to fetch.` and skips. Unreachable + today: every `forceMedia` caller sets `keepSourceVideoOverride: true`. +- **The summary names counters when nothing else moved:** `only counters (view_count, …)`. + - "nothing meaningful (formats only)" is kept for an all-volatile entry; + `unreadable metadata (no diff)` for a torn side. + - Real rewrites nearly always move a counter, so without this the line would read "formats only" + over a view count that changed. +- **An unparseable side** records the rewrite (both fingerprints) with an empty diff and + `unparseable: ["from"|"to"]`, rather than diffing against `{}` (which would list every key). +- **A counter present on one side only** is `[null, x]` in `counters`, not in added/removed. +- **The surface is in the header, not in a card.** There is no metadata card. The header is where + the page shows the metadata, and a header `<details>` is visible without opening a collapsed + card. +- **README: nothing.** `grep -n "Persist source video" README.md` is empty. Its `full: true` line + ("it needs a video the editor already knows") is still true. + +| sha | what | +|---|---| +| `92ae844c` | `common:` `metadataHistory.ts` + `metadataHistory-server.ts`; `metadataHistory.test.ts` (13), `metadataHistory-server.test.ts` (5); the sidecar enumeration test names ten | +| `51166736` | `common:` `forceMedia` (attempt-3 gate, the log line, the subs refusals, persist-only finalize); the three history wraps + `dataDirIdForUrl`; `archiveSourceVideo` and `persistKept` pass it; `forceMedia.test.ts` (6) | +| `1877d3a5` | `editor:` `views/metadataHistoryView.ts` (+ test, 6), `MetadataHistoryDetails.tsx`, the page loader; the Source video card un-gated | +| `92c26c73` | `e2e:` `persist-youtube-handling.spec.ts` (2), `fetch-window.spec.ts` +1, the fake's `.fake-ytdlp-metadata.json` knob | +| `6adbd455` | `docs:` AGENTS.md, "Clips and report-to-video" | +| `31eeece4` | `common:` the `PREFETCH_OWN_FILES` comment (comment only) | +| _this_ | `plans:` this record; two `[Unreleased]` bullets; the FACTS amendment (slice M's "silent no-op" is fixed, with the VTT-overwrite fact) | + +**Gates**, all from the worktree root: +- **tsc** (`pnpm -r --no-bail --workspace-concurrency=1 exec tsc --noEmit`) clean before every + commit (`n-tsc1.log` … `n-tsc5.log`). +- **common 1,984/1,984** (on `6adbd455`). That is 1,954 + 30: `metadataHistory` 13, `-server` 5, `forceMedia` 6, + the view 6. +- **editor unit 85/85.** +- **`test:scripts` 173 + 1 skip of 174.** +- **mcp 269/269**, unchanged (`n-gates1.log`). +- **`pnpm --filter editor exec next build`** ok: compiled in 17.5 s, 50 s total (`n-build1.log`). +- **e2e** (queued, detached, `export/public` links in place, no dangling links): **78 passed, + 0 failed, 8.7 min** (`n-e2e1.log`), on `92c26c73` (the two commits after it are AGENTS.md and + a comment). The specs: + - the plan's: `saved-videos`, `fetch-window` and the new `persist-youtube-handling`; + - the grep for `downloadOneManaged|Persist source video|persist-kept|redownload`: `import-video` + and `incomplete-transcript` (plus `saved-videos`); + - the paths this slice touched: `no-subs-fallback` (attempt 3), `video-page` (the header), + `title-filter` and `chat-only` (the prefetch wrap beside `discardPrefetchDir`), `skip-live` + (the prefetch), `audio-check-scenarios` (the audio-check wrap). +- **The new tests bite.** + - e2e: with `forceMedia: true` removed from `archiveSourceVideo` and `withMetadataHistory` made a + pass-through (uncommitted), `persist-youtube-handling.spec.ts fetch-window.spec.ts` gave + **6 passed, 3 failed, 1.3 min** (`n-bite1.log`), and the three failures are exactly the new + tests: + - the full fetch ends `done` with `file` undefined (slice M's live no-op, reproduced); + - no `forceMedia:` line; + - no `metadata.history.json`. + - unit: with the attempt-3 gate forced off, 3 of the 6 `forceMedia.test.ts` fail; with + `extractAudio: true`, 1 fails. +- **Numbers tool: none**, as the plan says. Nothing was run against the live :3001 editor or the + real corpus. + +**Found and left.** +- **The primary subtitle pass rewrites a `transcript.<lang>.vtt`.** + - This is not new: the decision keeps that pass "as today", and every re-download already does + it. yt-dlp deletes and re-fetches an existing subtitle file unless `--no-overwrites` is passed + (`YoutubeDL.existing_file`, `default_overwrite=True`; verified in the installed + `yt-dlp-patched` 2026.08.19). + - So the rollout's check "the transcript's mtime unchanged" on `teamrcn/dbnS-cBgStY` holds for a + whisper `transcript.json`, not for a VTT. A VTT's bytes are what YouTube serves now. + - What slice N guarantees is that the MEDIA pass writes no transcript. The new e2e asserts it on a + `transcript.json`, with inline whisper on. +- **A title-filter rejection now keeps a directory that has a history** (the plan's allow-list + rule). This takes a dir left by an earlier pass that was not rejected (a failed download) and a + filter that rejects it now. Real rewrites always differ (`epoch`), so such a dir gets a history + and keeps its metadata stub. +- **Two history entries per audio-checked download** (the prefetch, then the audio-check primary's + own re-extraction), as the plan's wraps imply. +- **No history for an unidentifiable URL** (the `%(id)s` dir is known only afterwards), nor for a + rewrite by a pass the plan did not wrap: the subtitle primary and the fallback load the + prefetched info json and write none, unless a channel's `ytdlpExtraArgs` carry + `--write-info-json`. +- **"Persist kept now" on a youtube-handling channel now downloads every kept video's source** + (the plan's known limitation). `persistKept` still counts a returned-but-failed download as + `persisted` (pre-existing). +- **A forced download that fails on a video with a transcript marks the whole download failed** + (review L2; a low for a later slice, not fixed here). + - The `else` branch after the attempt-3 run (`downloadOneManaged.ts:1470-1472`) sets `status = + "failed"` and `lastSucceeded = false` for a `keepTranscript` pass too. + - So the video page shows "Download failed" on a video whose transcript is fine. + - The whole-recording fetch job still ends `done` with no file, and the MCP says "finished but + named no file". + - yt-dlp's `source-media.*.part` (from `bestvideo*+bestaudio`, possibly several GB) can stay in + `data/<id>/` until a later persist resumes it. + - These are the same mechanics as today's no-subs fallback, now reachable on youtube channels. + - The review's option: on a `keepTranscript` failure, keep the subtitle pass's status and record + only the failed attempt. + +**Review fixes** (review SHIP AFTER FIXES, `n-review.md`: one medium, three lows; the three +questions ruled as asked — the card un-gated, no `--no-overwrites`, the history file kept off +`PREFETCH_OWN_FILES`). + +1. **M1: nothing the operator reads says the transcript is "kept"** (`3067f8af`). On a + youtube-handling channel the subtitle pass deletes and re-fetches `transcript.<lang>.vtt` (and + the live chat) before the media pass runs, so "kept" was true only of a Whisper + `transcript.json`. + - The Source video card, the `[Unreleased]` bullet and the AGENTS.md paragraph now say: + "YouTube's subtitles are fetched again first, as on any re-download; a Whisper transcript is + not touched, and no audio is extracted beside a transcript." + - The AGENTS.md sentence also says the pass was skipped for a transcript *or captions* (nit N5). +2. **L1: a test that `persistKept` passes `forceMedia`** (`42f836cd`, + `controller/persistKeptForceMedia.test.ts`, 1 test). + - Setup: youtube handling, `keepLatest: 1`, a kept video with a `transcript.json`, + `SETTINGS_FILE` set to a temp file with the disk floor off. + - Asserts: three spawns, the last without `--skip-download`; `persisted: 1`; the pointer + (`override`) and the container in the store; no `audio.*` or `source-media.*` in the data dir; + the transcript's bytes unchanged. + - It bites: with `persistKept.ts:137` removed it fails. +3. **L2:** recorded above, under "Found and left". +4. **L3: the full editor suite** (below). +- **Nits not taken:** + - N1 is covered by M1's "beside a transcript". + - N2, N3 (a 300-unit cut can split a surrogate pair), N4 and N6 are cosmetic. + +| sha | what | +|---|---| +| `3067f8af` | M1: the card text, the changelog bullet, the AGENTS.md clause (+ N5) | +| `42f836cd` | L1: `persistKeptForceMedia.test.ts` | +| _this_ | `plans:` these fixes, L2, the full-suite gate | + +**Gates on `42f836cd`:** +- **tsc:** clean before each commit (`n-tsc6.log`, `n-tsc7.log`). +- **common 1,985/1,985** (+1, the L1 test). +- **editor unit 85/85.** +- **`test:scripts` 173 + 1 skip.** +- **mcp 269/269.** +- **Editor build:** `pnpm --filter editor exec next build` ok, compiled in 19.9 s, 54 s total + (`n-gates2.log`). +- **The FULL editor e2e suite** (`pnpm e2e`, no spec list; queued, detached, `export/public` links + in place with none dangling): **641 passed, 2 failed, 12 skipped of 655, 48.6 min** + (`n-e2e-full1.log`). Each failure was re-run alone ×3 + (`pnpm e2e cookies-mode.spec.ts:241 tags.spec.ts:220 --repeat-each=3`): **5 passed, 1 failed, + 3.0 min** (`n-rerun1.log`). + - `tags.spec.ts:220` ("a video whose description and transcript are large previews and + renders"): **3/3 alone — a flake.** + - The failed assertion is the `addTag` helper's `tags-saved` wait (5 s), under load. + - The test downloads nothing, so no `metadata.history.json` exists and the new header block + renders nothing. + - `cookies-mode.spec.ts:241` ("defer: … downloads via the bucket button"): **2/3 alone — an + intermittent that predates this slice.** + - Both failures are the same: `getByLabel('Retry needs cookies output')` → "element(s) not + found" while waiting for "Managed download complete". The captured log shows the download + itself finished (`single-url fetched vidcookiegated1`). + - The cause: `NeedsCookiesList` (`editor/app/channels/[slug]/components/stages/DownloadStage.tsx`) + returns `null` when the bucket empties, so the refresh after the successful download + unmounts the card and its run log before the last line arrives. That is the FACTS + "a run log lives in the panel's React state" hazard, in a file this slice does not touch. + - Nothing of this slice runs on that path. The gated video's first prefetch fails before + writing metadata, so there is no history. The retry is not a forced download, and its + transcript comes from the subtitle pass, so there is no attempt 3. + - Left for a later slice: render the card while its run lives, as `SourceVideoSection` does. + ## Rollout Nothing is rolled out, except that **Jeralyzer is already on the brand, in Signal** (a build-deploy