import { test } from "node:test"; import assert from "node:assert/strict"; import { canonicalScanId, resolveScanErrorId } from "./metadataScan"; // Run with: // pnpm --filter yt-dlp-transcript-common exec tsx --test common/ytdlp/metadataScan.test.ts // yt-dlp's `id` is the NATIVE extractor id. Every consumer of the scan store // keys by the CANONICAL id (extractVideoId of the URL — the name data// // carries). They coincide on YouTube and diverge everywhere else, and keyed // natively nothing on those platforms would ever settle: the settled set would // be full of ids no playlist entry and no directory has ever been called. test("YouTube: native and canonical are the same id", () => { assert.equal( canonicalScanId({ id: "dQw4w9WgXcQ", webpage_url: "https://www.youtube.com/watch?v=dQw4w9WgXcQ", }), "dQw4w9WgXcQ", ); }); test("Rumble: the URL slug wins over the embed id", () => { // The case that motivated this: the archive keys by the embed id, a local cue // dir is named for the URL slug (see AGENTS.md). assert.equal( canonicalScanId({ id: "v2apmfn", webpage_url: "https://rumble.com/v2846lb-some-title.html", }), "v2846lb", ); }); test("Twitch: the canonical id keeps the leading v the native id drops", () => { assert.equal( canonicalScanId({ id: "2735747903", webpage_url: "https://www.twitch.tv/videos/v2735747903", }), "v2735747903", ); }); test("no usable URL falls back to the native id rather than dropping the record", () => { assert.equal(canonicalScanId({ id: "abc123" }), "abc123"); assert.equal(canonicalScanId({ id: "abc123", webpage_url: "not a url" }), "abc123"); assert.equal(canonicalScanId({}), ""); }); // --- error lines ------------------------------------------------------------- // yt-dlp names the NATIVE id on an ERROR: line and gives no URL to canonicalize. const URLS = new Map([ ["v2846lb", "https://rumble.com/v2846lb-some-title.html"], ["dQw4w9WgXcQ", "https://www.youtube.com/watch?v=dQw4w9WgXcQ"], ]); test("an error id resolves through a record already read this run", () => { assert.equal( resolveScanErrorId("v2apmfn", new Map([["v2apmfn", "v2846lb"]]), URLS), "v2846lb", ); }); test("an error id that IS a canonical id is taken as-is (the YouTube case)", () => { assert.equal(resolveScanErrorId("dQw4w9WgXcQ", new Map(), URLS), "dQw4w9WgXcQ"); }); test("an error id found inside a target URL resolves to that target", () => { // Nothing was read for it this run — the very first video failing is exactly // when this matters — so the only link left is the URL itself. assert.equal(resolveScanErrorId("v2846lb-some", new Map(), URLS), "v2846lb"); }); test("an unrecognizable error id is kept, not dropped", () => { // A wrong key for an ERROR row costs one re-read next run. Dropping it would // make the backlog never reach zero. assert.equal(resolveScanErrorId("whoknows", new Map(), URLS), "whoknows"); });