Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 39a36becdc4f81628c2fa78b59cf0ff0536e0320
parent 62cb88b946f1fa076ca6b5e60fe17a0a16c34468
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sat, 26 Sep 2026 12:52:38 -0400

Merge r10/runner-lows — release 10 slice L2: YouTube's "try again later" soft block backs off (download batch, runner, scan, a prefetch early exit, the availability check stops); the pre-clean gate never judges an unprobed suspect; the boot pass waits at most 60 s for the storage pass; /jobs shows cancelReason; the safeRevalidate warning is counted

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
Acommon/controller/checkAvailability.test.ts | 132+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/controller/checkAvailability.ts | 72+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++---
Mcommon/controller/verifyBeforeClean.test.ts | 181++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-
Mcommon/controller/verifyBeforeClean.ts | 28+++++++++++++++++++++++++++-
Mcommon/jobs/bootQueuedJobs.test.ts | 117++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-
Mcommon/jobs/bootQueuedJobs.ts | 96+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++--
Mcommon/jobs/listJobs.test.ts | 37+++++++++++++++++++++++++++++++++++++
Mcommon/jobs/listJobs.ts | 10++++++++++
Mcommon/lib/availability.test.ts | 93++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-
Mcommon/lib/availability.ts | 34++++++++++++++++++++++++++++++++++
Mcommon/views/jobRowView.ts | 4++++
Mcommon/views/jobRows.test.ts | 22++++++++++++++++++++++
Mcommon/views/jobRows.ts | 1+
Mcommon/ytdlp/downloadOneManaged.ts | 68+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-----
Mcommon/ytdlp/managedDownloadsSleep.test.ts | 89++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++---
Mcommon/ytdlp/metadataScan.ts | 12++++++++++--
Acommon/ytdlp/prefetchRateLimit.test.ts | 115+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/ytdlp/runYtdlp.ts | 9++++++---
Meditor/CHANGELOG.md | 4++++
Meditor/app/jobs/[id]/page.tsx | 27+++++++++++++++++++++++++--
Meditor/app/jobs/components/JobRow.tsx | 30++++++++++++++++++++++++++++++
Meditor/app/lib/safeRevalidate.test.ts | 139++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-------
Meditor/app/lib/safeRevalidate.ts | 174++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-------------
Meditor/e2e/jobs-filters.spec.ts | 49+++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/instrumentation.ts | 49++++++++++++++++++++++---------------------------
Mplans/FACTS.md | 26+++++++++++++++++++++++---
Mplans/release-10.md | 248+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
27 files changed, 1774 insertions(+), 92 deletions(-)

diff --git a/common/controller/checkAvailability.test.ts b/common/controller/checkAvailability.test.ts @@ -0,0 +1,132 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { chmod, mkdir, mkdtemp, readFile, rm, writeFile } from "node:fs/promises"; +import { tmpdir } from "node:os"; +import path from "node:path"; +import type { Paths } from "../lib/paths"; +import { runAvailabilityCheck } from "./checkAvailability"; + +// Run with: +// pnpm --filter yt-dlp-transcript-common exec tsx --test controller/checkAvailability.test.ts +// +// THE AVAILABILITY CHECK STOPS ON A RATE LIMIT (release 10, L2 review). Its +// probes are one `--dump-json` each; before, a rate-limited probe was recorded +// and the next one started anyway, into the same refusal. Now the first +// `rate_limit` probe records its own `error`, no further probe starts, and the +// platform cooldown is recorded once. A temp yt-dlp script answers every probe +// with the given line and counts the spawns; the cooldown is injected. + +process.env.SETTINGS_FILE = path.join(tmpdir(), "ttb-avail-no-such-settings.json"); + +const IDS = ["vid00000001", "vid00000002", "vid00000003", "vid00000004"]; + +async function run(stderrLine: string): Promise<{ + spawns: number; + backoffs: string[]; + result: Awaited<ReturnType<typeof runAvailabilityCheck>>; + recorded: Record<string, string | null>; + log: string[]; +}> { + const dir = await mkdtemp(path.join(tmpdir(), "ttb-avail-")); + try { + const counter = path.join(dir, "spawns"); + const bin = path.join(dir, "fake-ytdlp.sh"); + await writeFile( + bin, + `#!/bin/sh\necho x >> '${counter}'\necho ${JSON.stringify(stderrLine)} >&2\nexit 1\n`, + ); + await chmod(bin, 0o755); + const paths = { + channelsDir: path.join(dir, "channels"), + ytdlpBin: bin, + } as Paths; + const ch = path.join(paths.channelsDir, "ch"); + await mkdir(ch, { recursive: true }); + await writeFile( + path.join(ch, "config.json"), + JSON.stringify({ handling: "youtube", url: "https://www.youtube.com/@ch/videos" }), + ); + for (const id of IDS) { + const vd = path.join(ch, "data", id); + await mkdir(vd, { recursive: true }); + await writeFile( + path.join(vd, "metadata.info.json"), + JSON.stringify({ id, webpage_url: `https://www.youtube.com/watch?v=${id}` }), + ); + } + const backoffs: string[] = []; + const log: string[] = []; + const result = await runAvailabilityCheck({ + channelSlug: "ch", + paths, + mode: "recheck-all", + ignoreShard: true, + concurrency: 1, + onLog: (m) => log.push(m), + onPlatformBackoff: (pf) => { + backoffs.push(pf); + }, + }); + const spawns = (await readFile(counter, "utf8").catch(() => "")) + .split("\n") + .filter(Boolean).length; + const recorded: Record<string, string | null> = {}; + for (const id of IDS) { + const raw = await readFile( + path.join(ch, "data", id, "availability.json"), + "utf8", + ).catch(() => null); + recorded[id] = raw ? (JSON.parse(raw) as { availability: string }).availability : null; + } + return { spawns, backoffs, result, recorded, log }; + } finally { + await rm(dir, { recursive: true, force: true }); + } +} + +test("a soft-blocked probe stops the check: one spawn, one cooldown, the rest skipped", async () => { + const r = await run( + "ERROR: [youtube] vid00000001: This content isn't available, try again later. " + + "The current session has been rate-limited by YouTube for up to an hour.", + ); + assert.equal(r.spawns, 1); + assert.equal(r.result.blocked, true); + assert.match(r.result.blockMessage ?? "", /try again later/); + assert.equal(r.result.probedIds.length, 1); + assert.equal(r.result.attempted, 1); + assert.equal(r.result.skipped, IDS.length - 1); + // The one probe records its observation — `error`, never `deleted` — and + // nothing else is written. + const written = Object.values(r.recorded).filter((v) => v !== null); + assert.deepEqual(written, ["error"]); + assert.equal(r.recorded[r.result.probedIds[0]], "error"); + // The cooldown, once, under the channel's platform key. + assert.deepEqual(r.backoffs, ["youtube"]); + assert.ok(r.log.some((l) => /^STOPPED: the source is rate-limiting availability probes/.test(l))); +}); + +test("a 429 and the bot check stop it the same way", async () => { + for (const line of [ + "ERROR: [youtube] vid00000001: Unable to download webpage: HTTP Error 429: Too Many Requests", + "ERROR: [youtube] vid00000001: Sign in to confirm you’re not a bot. Use --cookies-from-browser or --cookies for the authentication.", + ]) { + const r = await run(line); + assert.equal(r.spawns, 1, line); + assert.equal(r.result.blocked, true, line); + assert.deepEqual(r.backoffs, ["youtube"], line); + } +}); + +test("any other failure does not stop it: every video is probed, no cooldown", async () => { + for (const line of [ + "ERROR: [youtube] vid00000001: Unable to download webpage: HTTP Error 403: Forbidden", + "ERROR: [youtube] vid00000001: Video unavailable", + ]) { + const r = await run(line); + assert.equal(r.spawns, IDS.length, line); + assert.equal(r.result.blocked, false, line); + assert.equal(r.result.blockMessage, undefined, line); + assert.deepEqual([...r.result.probedIds].sort(), IDS, line); + assert.deepEqual(r.backoffs, [], line); + } +}); diff --git a/common/controller/checkAvailability.ts b/common/controller/checkAvailability.ts @@ -6,9 +6,11 @@ import { AVAILABILITY_VALUES, EXCLUDED_FROM_DOWNLOAD, availabilityFromJsonField, + classifyDownloadFailure, parseUnavailableFromStderr, type Availability, } from "../lib/availability"; +import { recordDownloadBackoff } from "../jobs/downloadBackoff"; import { loadAvailability, recordAvailability, @@ -65,6 +67,11 @@ export type CheckAvailabilityOpts = { ignoreShard?: boolean; onLog?: (msg: string) => void; signal?: AbortSignal; + // Records the shared per-platform cooldown when a probe is rate-limited + // (see `blocked` below). Defaults to recordDownloadBackoff — the helper and + // key a download's 429 and a manual Sync use. A test seam; production + // callers pass nothing. + onPlatformBackoff?: (platform: string) => void | Promise<void>; }; export type CheckAvailabilityResult = { @@ -72,6 +79,17 @@ export type CheckAvailabilityResult = { checked: number; skipped: number; byStatus: Record<Availability, number>; + // THE CHECK STOPS ON A RATE LIMIT (release 10, L2 review). A probe that + // classifies `rate_limit` — HTTP 429, the bot check, YouTube's soft block + // "…isn't available, try again later" — records its own `error` observation + // and then no further probe is started: every later id is `skipped`, and + // the platform cooldown is recorded once. The ids that WERE probed are in + // `probedIds`; anything else a caller reads for this run is an OLD record, + // and a caller that acts on the result must not judge it (verifyBeforeClean + // leaves it unverified). + blocked: boolean; + blockMessage?: string; + probedIds: string[]; }; function emptyByStatus(): Record<Availability, number> { @@ -114,6 +132,7 @@ export async function runAvailabilityCheck({ ignoreShard = false, onLog, signal, + onPlatformBackoff, }: CheckAvailabilityOpts): Promise<CheckAvailabilityResult> { const log = onLog ?? ((m: string) => console.log(m)); const channelDir = path.join(paths.channelsDir, channelSlug); @@ -151,13 +170,23 @@ export async function runAvailabilityCheck({ log( `Shard availability: save-only — slice ${(shardResult.config?.shardIndex ?? shardIndex ?? 0) + 1}/${shardResult.config?.totalShards ?? shardTotal} (${videoDirs.length} item(s)) persisted; not checking.`, ); - return { attempted: 0, checked: 0, skipped: 0, byStatus: emptyByStatus() }; + return { + attempted: 0, + checked: 0, + skipped: 0, + byStatus: emptyByStatus(), + blocked: false, + probedIds: [], + }; } let attempted = 0; let checked = 0; let skipped = 0; const byStatus = emptyByStatus(); + // Set by the first rate-limited probe; see CheckAvailabilityResult.blocked. + let blocked: { message: string; url: string } | null = null; + const probedIds: string[] = []; // Pre-scan: count how many already have a parsed availability.json so the // "skipped" total in the summary is unambiguous (no, we are not skipping @@ -177,7 +206,9 @@ export async function runAvailabilityCheck({ await Promise.all( videoDirs.map((id) => limit(async () => { - if (signal?.aborted) { + // A rate-limited source is not asked again in this run: each further + // probe would be another request into the same refusal. + if (signal?.aborted || blocked) { skipped++; return; } @@ -264,6 +295,15 @@ export async function runAvailabilityCheck({ const trimmed = stderr.trim().split("\n").slice(-3).join("\n"); if (trimmed) error = trimmed; } + if ( + !blocked && + classifyDownloadFailure(stderr, availability) === "rate_limit" + ) { + blocked = { + message: (error ?? stderr.trim()).slice(0, 300), + url, + }; + } } await recordAvailability(videoDir, { @@ -279,11 +319,29 @@ export async function runAvailabilityCheck({ }); checked++; byStatus[availability]++; + probedIds.push(id); log(`Check ${id}: ${availability}`); }), ), ); + if (blocked) { + const b = blocked as { message: string; url: string }; + const platform = detectPlatform(config?.url ?? b.url) ?? "unknown"; + const unprobed = videoDirs.length - probedIds.length; + log( + `STOPPED: the source is rate-limiting availability probes (${b.message}). ` + + `${unprobed} video(s) were not probed this run; a ${platform} cooldown has been recorded.`, + ); + try { + await (onPlatformBackoff ?? ((pf) => recordDownloadBackoff(pf, paths)))( + platform, + ); + } catch { + /* the cooldown is best-effort; the stop above already happened */ + } + } + // Type guard: every key in AVAILABILITY_VALUES has been incremented above. void AVAILABILITY_VALUES; @@ -294,5 +352,13 @@ export async function runAvailabilityCheck({ `Availability check done. attempted=${attempted} checked=${checked} skipped=${skipped} (${summary}).`, ); - return { attempted, checked, skipped, byStatus }; + return { + attempted, + checked, + skipped, + byStatus, + blocked: blocked !== null, + ...(blocked ? { blockMessage: (blocked as { message: string }).message } : {}), + probedIds, + }; } diff --git a/common/controller/verifyBeforeClean.test.ts b/common/controller/verifyBeforeClean.test.ts @@ -1,6 +1,6 @@ import { test } from "node:test"; import assert from "node:assert/strict"; -import { mkdir, mkdtemp, readFile, rm, writeFile } from "node:fs/promises"; +import { chmod, mkdir, mkdtemp, readFile, rm, writeFile } from "node:fs/promises"; import { tmpdir } from "node:os"; import path from "node:path"; import type { Paths } from "../lib/paths"; @@ -175,3 +175,182 @@ test("an unreachable source fails CLOSED: nothing is cleared to delete", async ( assert.equal(excludedIds(verdicts).size, 2); }); }); + +// A RATE-LIMITED CONFIRMATION (release 10, L2 review). The tier C check now +// stops at its first rate-limited probe. The suspects it never reached still +// carry whatever availability.json they had — here a months-old `public` — +// and judged on that they would be CLEARED FOR DELETE. They must come back +// unverified instead. A temp yt-dlp script answers the flat listing (one +// listed video, so the rest are suspects) and each probe per `probe`. +async function withFakeSource( + probe: string, + fn: (paths: Paths, probes: () => Promise<string[]>) => Promise<void>, +): Promise<void> { + const dir = await mkdtemp(path.join(tmpdir(), "ttb-vbc-src-")); + const probesFile = path.join(dir, "probes"); + const bin = path.join(dir, "fake-ytdlp.sh"); + await writeFile( + bin, + [ + "#!/bin/sh", + 'for a in "$@"; do last="$a"; done', + 'case " $* " in', + ' *" --flat-playlist "*) echo "https://www.youtube.com/watch?v=listed00001"; exit 0;;', + "esac", + `echo "$last" >> '${probesFile}'`, + probe, + "", + ].join("\n"), + ); + await chmod(bin, 0o755); + const paths = { + channelsDir: path.join(dir, "channels"), + ytdlpBin: bin, + autoQueueStateFile: path.join(dir, ".auto-queue", "state.json"), + } as Paths; + try { + await fn(paths, async () => + (await readFile(probesFile, "utf8").catch(() => "")) + .split("\n") + .filter(Boolean), + ); + } finally { + await rm(dir, { recursive: true, force: true }); + } +} + +async function seedProbeable( + paths: Paths, + id: string, + opts: { noUrl?: boolean } = {}, +): Promise<void> { + const dir = path.join(paths.channelsDir, "ch", "data", id); + await mkdir(dir, { recursive: true }); + await writeFile( + path.join(dir, "metadata.info.json"), + JSON.stringify( + opts.noUrl + ? { id } + : { id, webpage_url: `https://www.youtube.com/watch?v=${id}` }, + ), + ); + // A STALE verdict from long ago: exactly what must not clear a delete. + await writeFile( + path.join(dir, "availability.json"), + JSON.stringify({ checkedAt: "2026-01-01T00:00:00.000Z", availability: "public" }), + ); +} + +const SUSPECTS = ["suspect0001", "suspect0002", "suspect0003"]; + +test("a rate-limited confirmation leaves every suspect it did not probe unverified: nothing is cleaned on an old record", async () => { + await withFakeSource( + `echo "ERROR: [youtube] x: This content isn't available, try again later." >&2; exit 1`, + async (paths, probes) => { + await seedConfig(paths, "ch", { + handling: "youtube", + url: "https://www.youtube.com/@ch/videos", + }); + await seedProbeable(paths, "listed00001"); + for (const id of SUSPECTS) await seedProbeable(paths, id); + + const verdicts = await verifyBeforeClean({ + channelSlug: "ch", + paths, + candidateIds: ["listed00001", ...SUSPECTS], + onLog: () => {}, + }); + + // ONE probe: the check stopped at the soft block. + assert.equal((await probes()).length, 1); + // Every suspect is protected: the probed one read `error`, the other two + // were never asked and are NOT judged on their stale `public`. + assert.deepEqual([...verdicts.unverified].sort(), SUSPECTS); + assert.equal(verdicts.pinned.size, 0); + const excluded = excludedIds(verdicts); + for (const id of SUSPECTS) assert.ok(excluded.has(id), id); + // The listed video is still listed: cleanable as before. + assert.equal(excluded.has("listed00001"), false); + // …and the platform cooldown was recorded, once, under youtube. + const state = JSON.parse( + await readFile(paths.autoQueueStateFile, "utf8"), + ) as { download: { platformBackoff: Record<string, { until: number }> } }; + assert.ok(state.download.platformBackoff.youtube.until > Date.now()); + }, + ); +}); + +test("an unblocked confirmation is judged exactly as before", async () => { + await withFakeSource( + [ + 'case "$last" in', + ` *suspect0001*) echo '{"id":"suspect0001","availability":"public"}';;`, + ` *suspect0002*) echo "ERROR: [youtube] suspect0002: Video unavailable" >&2; exit 1;;`, + ` *) echo "ERROR: [youtube] x: Unable to download webpage: HTTP Error 403: Forbidden" >&2; exit 1;;`, + "esac", + ].join("\n"), + async (paths, probes) => { + await seedConfig(paths, "ch", { + handling: "youtube", + url: "https://www.youtube.com/@ch/videos", + }); + await seedProbeable(paths, "listed00001"); + for (const id of SUSPECTS) await seedProbeable(paths, id); + + const verdicts = await verifyBeforeClean({ + channelSlug: "ch", + paths, + candidateIds: ["listed00001", ...SUSPECTS], + onLog: () => {}, + }); + + // Every suspect probed (a 403 is not a rate limit, so nothing stops). + assert.equal((await probes()).length, 3); + // public → cleanable; deleted → pinned; error → unverified. + assert.equal(excludedIds(verdicts).has("suspect0001"), false); + assert.equal(verdicts.pinned.get("suspect0002"), "deleted"); + assert.deepEqual([...verdicts.unverified], ["suspect0003"]); + // No cooldown for a run nothing rate-limited. + await assert.rejects(readFile(paths.autoQueueStateFile, "utf8")); + }, + ); +}); + +test("an unblocked confirmation that could not probe a suspect (no webpage_url) leaves it unverified, not judged on its old record", async () => { + await withFakeSource( + [ + 'case "$last" in', + ` *suspect0001*) echo '{"id":"suspect0001","availability":"public"}';;`, + ` *) echo "ERROR: [youtube] x: Video unavailable" >&2; exit 1;;`, + "esac", + ].join("\n"), + async (paths, probes) => { + await seedConfig(paths, "ch", { + handling: "youtube", + url: "https://www.youtube.com/@ch/videos", + }); + await seedProbeable(paths, "listed00001"); + await seedProbeable(paths, "suspect0001"); + // Not listed, and nothing to probe it by — but its old record says + // `public`, which would clear it for delete if it were read. + await seedProbeable(paths, "nourl000001", { noUrl: true }); + + const verdicts = await verifyBeforeClean({ + channelSlug: "ch", + paths, + candidateIds: ["listed00001", "suspect0001", "nourl000001"], + onLog: () => {}, + }); + + // The check was not blocked: it probed the one suspect it could. + assert.deepEqual(await probes(), ["https://www.youtube.com/watch?v=suspect0001"]); + assert.deepEqual([...verdicts.unverified], ["nourl000001"]); + const excluded = excludedIds(verdicts); + assert.ok(excluded.has("nourl000001")); + // The probed public suspect and the listed video stay cleanable. + assert.equal(excluded.has("suspect0001"), false); + assert.equal(excluded.has("listed00001"), false); + await assert.rejects(readFile(paths.autoQueueStateFile, "utf8")); + }, + ); +}); diff --git a/common/controller/verifyBeforeClean.ts b/common/controller/verifyBeforeClean.ts @@ -153,8 +153,9 @@ export async function verifyBeforeClean({ log( `Verify before clean: ${suspects.length} candidate(s) missing from the fresh listing — confirming individually…`, ); + let check: Awaited<ReturnType<typeof runAvailabilityCheck>>; try { - await runAvailabilityCheck({ + check = await runAvailabilityCheck({ channelSlug, paths, mode: "recheck-all", @@ -182,8 +183,33 @@ export async function verifyBeforeClean({ return verdicts; } + // NOTHING THE CHECK DID NOT PROBE IS JUDGED (release 10, L2 review). A + // suspect the check never asked about still has whatever `availability.json` + // it already had — which can be months old and say `public`. Read below, + // that stale record would clear the video for an irreversible delete. So + // every suspect missing from `probedIds` is `unverified`: skipped, no + // marker, retried on the next sweep. Two ways to be missing: the check + // stopped at a rate-limited probe before reaching it (`blocked`), or it was + // skipped — no `webpage_url` in its metadata, so there was nothing to probe. + // The second was judged on its old record before this; no video on the live + // corpus was in that state (2026-09-26 review scan), and none can be now. + const probed = new Set(check.probedIds); + const unprobed = new Set(suspects.filter((id) => !probed.has(id))); + if (unprobed.size > 0) { + log( + (check.blocked + ? `Verify before clean: the source rate-limited the confirmation — ` + : `Verify before clean: the confirmation could not probe every suspect — `) + + `${unprobed.size} suspect(s) it did not reach are left unverified, not judged on an old record.`, + ); + } + for (const id of suspects) { if (signal?.aborted) break; + if (unprobed.has(id)) { + verdicts.unverified.add(id); + continue; + } const availability = await resolveEffectiveAvailability( path.join(dataDir, id), ); diff --git a/common/jobs/bootQueuedJobs.test.ts b/common/jobs/bootQueuedJobs.test.ts @@ -6,7 +6,13 @@ import path from "node:path"; import type { Paths } from "../lib/paths"; import type { JobMeta } from "./jobMeta"; import type { JobSpec } from "./jobSpec"; -import { settleQueuedJobMetas, type RequeueFn } from "./bootQueuedJobs"; +import { + STORAGE_PASS_WAIT_MS, + settleAfterStoragePass, + settleQueuedJobMetas, + waitForStoragePass, + type RequeueFn, +} from "./bootQueuedJobs"; // Run with: // pnpm --filter yt-dlp-transcript-common exec tsx --test jobs/bootQueuedJobs.test.ts @@ -336,3 +342,112 @@ test("a missing .jobs dir is nothing to do", async () => { }); assert.deepEqual(res, { requeued: [], cancelled: [] }); }); + +// THE WAIT ON THE STORAGE PASS IS BOUNDED (release 10, L2). A probe stuck on a +// hung mount (a `stat` that never returns) used to hold the settle forever, so +// every stale meta stayed `queued`. Timeouts are shortened here; production +// waits STORAGE_PASS_WAIT_MS. + +// A pass we finish by hand, so a test can prove the timeout did not. +function deferred(): { + promise: Promise<string>; + resolve: (v: string) => void; +} { + let resolve!: (v: string) => void; + const promise = new Promise<string>((r) => { + resolve = r; + }); + return { promise, resolve }; +} + +test("a storage pass that never finishes: the settle runs after the timeout, and says so", async () => { + const f = await fixture([{ id: "A1", spec: SPEC, channelSlug: "teamrcn" }]); + try { + const rq = recordingRequeue(); + const lines: string[] = []; + const never = new Promise<void>(() => {}); + const started = Date.now(); + const res = await settleAfterStoragePass(never, { + paths: f.paths, + requeue: rq.fn, + bootedAt: BOOT, + log: (l) => lines.push(l), + waitMs: 50, + }); + assert.ok(Date.now() - started >= 45, "it did wait, for the bound"); + assert.deepEqual(res.requeued, [{ id: "A1", newId: "NEW1" }]); + assert.equal((await f.read("A1")).status, "cancelled"); + // The timeout line comes first, then the settle's own lines. + assert.match(lines[0], /^\[boot\] storage pass still running after 0 s; settling queued jobs without it/); + assert.match(lines.at(-1) ?? "", /re-queued 1, cancelled 0/); + } finally { + await rm(f.root, { recursive: true, force: true }); + } +}); + +test("the timeout does not cancel the storage pass: it finishes later, with its own result, and is logged", async () => { + const pass = deferred(); + const lines: string[] = []; + let t = 1_000; + const outcome = await waitForStoragePass(pass.promise, { + timeoutMs: 20, + log: (l) => lines.push(l), + now: () => t, + }); + assert.equal(outcome, "timed-out"); + assert.equal(lines.length, 1); + // Another reader of the pass (the storage pass's own caller) still gets its + // value when it ends — the race was over a derived promise. + t = 1_000 + 95_000; + pass.resolve("probed 2"); + assert.equal(await pass.promise, "probed 2"); + await new Promise((r) => setImmediate(r)); + assert.equal(lines.length, 2); + assert.match( + lines[1], + /^\[boot\] storage pass finished 95 s after the queued-job pass began waiting \(it stopped waiting at 0 s\)$/, + ); +}); + +test("a storage pass that finishes in time is waited for, and nothing is logged about it", async () => { + const f = await fixture([{ id: "A1", spec: SPEC, channelSlug: "teamrcn" }]); + try { + let passDone = false; + const pass = new Promise<void>((r) => + setTimeout(() => { + passDone = true; + r(); + }, 30), + ); + const lines: string[] = []; + let requeuedAfterPass: boolean | undefined; + await settleAfterStoragePass(pass, { + paths: f.paths, + requeue: async () => { + requeuedAfterPass = passDone; + return { ok: true, jobId: "NEW1" }; + }, + bootedAt: BOOT, + log: (l) => lines.push(l), + waitMs: 5_000, + }); + assert.equal(requeuedAfterPass, true, "the settle waited for the pass"); + assert.equal(lines.filter((l) => /storage pass/.test(l)).length, 0); + } finally { + await rm(f.root, { recursive: true, force: true }); + } +}); + +test("a storage pass that throws counts as finished: no timeout, the settle runs", async () => { + const lines: string[] = []; + const outcome = await waitForStoragePass( + Promise.reject(new Error("findmnt exploded")), + { timeoutMs: 5_000, log: (l) => lines.push(l) }, + ); + assert.equal(outcome, "done"); + assert.deepEqual(lines, []); +}); + +test("the production bound is 60 s", () => { + assert.equal(STORAGE_PASS_WAIT_MS, 60_000); +}); diff --git a/common/jobs/bootQueuedJobs.ts b/common/jobs/bootQueuedJobs.ts @@ -80,7 +80,7 @@ export function specKey(spec: JobSpec): string { }); } -export async function settleQueuedJobMetas(opts: { +export type SettleQueuedOpts = { paths: Paths; // null: cancel, never re-queue (an idle boot, or the e2e test server). requeue: RequeueFn | null; @@ -91,7 +91,99 @@ export async function settleQueuedJobMetas(opts: { log?: (line: string) => void; idleReason?: string; maxAgeMs?: number; -}): Promise<BootQueuedResult> { +}; + +// THE WAIT ON THE STORAGE BOOT PASS, BOUNDED (release 10, L2). +// +// The settle waits for the storage pass (release 9 review, LOW-1) so a channel +// whose location is mid-autoRepoint is reachable when its job is re-queued. It +// waited with no bound, and one probe in that pass is not bounded either: every +// findmnt has FINDMNT_TIMEOUT_MS (3 s), but the `stat` of a location's root and +// the statfs for its free space are plain syscalls, and on a hung network mount +// they never return. The boot pass then never ran, and every `queued` meta +// stayed `queued` on /jobs for the life of the process. +// +// 60 s. A healthy pass takes milliseconds (findmnt answers in under 10 ms here). +// Its bounded worst case is two or three 3 s findmnts per location — identity, +// fstab, or where the uuid is mounted — plus, for a location being re-pointed, +// the preflight's second probe and a stat per channel on it: ~12 s a location. +// 60 s covers several locations at that worst case; past it the pass is stuck on +// a syscall that will not answer, and waiting longer buys nothing. +// +// On a timeout the settle runs anyway. What a job re-queued for a channel the +// pass has not reached then meets depends on why the pass is slow: +// - an UNMOUNTED drive (the root is simply absent): the media guard +// (jobs/jobKinds.ts `needsMedia`, lib/channelMedia.ts) gets ENOENT at +// once, so the job is refused — at submission, which closes the old meta +// `cancelled` with the error, or when it starts, with the reason in its +// log. /jobs says so, and Retry is one click. +// - a HUNG mount, the case this bound exists for: the guard's own `stat` +// (lib/channelMedia.ts) hangs on the same syscall, so that job waits with +// it. Re-queues run one at a time (below), so every re-queue after it stays +// `queued` until the mount answers. Every cancel has run by then — cancels +// come first — and re-queues are few (one at the first live boot), so this +// is left, not fixed: see the release 10 record, L2 "found and left". +// The storage pass is NOT cancelled (it has no signal, and a re-point it +// already enqueued must finish): it runs on, and a line says when it ends. +export const STORAGE_PASS_WAIT_MS = 60_000; + +export async function waitForStoragePass( + pass: Promise<unknown>, + opts: { + timeoutMs?: number; + log?: (line: string) => void; + now?: () => number; + } = {}, +): Promise<"done" | "timed-out"> { + const timeoutMs = opts.timeoutMs ?? STORAGE_PASS_WAIT_MS; + const log = opts.log ?? (() => {}); + const now = opts.now ?? Date.now; + const startedAt = now(); + // A pass that throws has still finished; the settle does not care how. This + // is a DERIVED promise — racing it never touches the pass itself. + const finished = pass.then( + () => "done" as const, + () => "done" as const, + ); + let timer: ReturnType<typeof setTimeout> | undefined; + const timedOut = new Promise<"timed-out">((resolve) => { + timer = setTimeout(() => resolve("timed-out"), timeoutMs); + }); + const outcome = await Promise.race([finished, timedOut]); + clearTimeout(timer); + if (outcome === "timed-out") { + const secs = (ms: number) => `${Math.round(ms / 1000)} s`; + log( + `[boot] storage pass still running after ${secs(timeoutMs)}; settling queued jobs without it ` + + `(a hung mount? a job re-queued for a channel on it will wait on the same mount, ` + + `and the re-queues after it with it)`, + ); + void finished.then(() => + log( + `[boot] storage pass finished ${secs(now() - startedAt)} after the queued-job pass began waiting ` + + `(it stopped waiting at ${secs(timeoutMs)})`, + ), + ); + } + return outcome; +} + +// What instrumentation.ts runs: the settle, after the storage pass or after +// STORAGE_PASS_WAIT_MS, whichever comes first. +export async function settleAfterStoragePass( + storagePass: Promise<unknown>, + opts: SettleQueuedOpts & { waitMs?: number }, +): Promise<BootQueuedResult> { + await waitForStoragePass(storagePass, { + timeoutMs: opts.waitMs, + log: opts.log, + }); + return settleQueuedJobMetas(opts); +} + +export async function settleQueuedJobMetas( + opts: SettleQueuedOpts, +): Promise<BootQueuedResult> { const log = opts.log ?? (() => {}); const maxAgeMs = opts.maxAgeMs ?? REQUEUE_MAX_AGE_MS; const result: BootQueuedResult = { requeued: [], cancelled: [] }; diff --git a/common/jobs/listJobs.test.ts b/common/jobs/listJobs.test.ts @@ -90,6 +90,43 @@ test("listAllJobs before-cursor skips newer entries", async () => { }); }); +// The boot pass (bootQueuedJobs.ts) closes a meta a restart left `queued` as +// `cancelled` with a `cancelReason`; /jobs draws it from the entry. +test("a cancelled meta's cancelReason reaches the entry; any other status drops it", async () => { + await withJobs(async (paths) => { + const write = async (id: string, status: string) => { + await writeFile(path.join(paths.jobsDir, `${id}.log`), "log\n"); + await writeFile( + path.join(paths.jobsDir, `${id}.meta.json`), + JSON.stringify({ + id, + kind: "whisper-all", + queueKey: "q", + status, + queuedAt: BASE, + endedAt: BASE + 1, + cancelReason: "queued before the last restart, stale", + }), + ); + }; + const cancelled = ulid(BASE); + const done = ulid(BASE + 1); + await write(cancelled, "cancelled"); + await write(done, "done"); + const one = await getJobEntry(paths, cancelled); + assert.equal(one?.status, "cancelled"); + assert.equal(one?.cancelReason, "queued before the last restart, stale"); + const page = await listAllJobs(paths); + const byId = new Map(page.entries.map((e) => [e.id, e])); + assert.equal( + byId.get(cancelled)?.cancelReason, + "queued before the last restart, stale", + ); + // A reason beside a status it does not explain is not carried. + assert.equal(byId.get(done)?.cancelReason, undefined); + }); +}); + test("getJobEntry resolves one job and null for unknown ids", async () => { await withJobs(async (paths, seed) => { const id = await seed(0, BASE); diff --git a/common/jobs/listJobs.ts b/common/jobs/listJobs.ts @@ -30,6 +30,11 @@ export type JobListEntry = { // A short phrase naming what this particular job is for, when its kind alone // does not say (a fetch-window job's requester and clip). See jobSpecDetail. detail?: string; + // Why a `cancelled` job was cancelled, when something other than a person + // pressing Cancel decided it — today only the boot pass (bootQueuedJobs.ts), + // e.g. "server restarted; the scheduler re-derives syncs". From the sidecar; + // absent on every other job. + cancelReason?: string; }; export type JobsPage = { @@ -135,6 +140,11 @@ async function buildEntry( inRegistry: false, replayable: Boolean(meta.spec), detail: jobSpecDetail(meta.kind, meta.spec), + // Only on a job that did end `cancelled`: a reason beside any other + // status would explain something that did not happen. + ...(meta.status === "cancelled" && meta.cancelReason + ? { cancelReason: meta.cancelReason } + : {}), logPath, logSize, }; diff --git a/common/lib/availability.test.ts b/common/lib/availability.test.ts @@ -1,6 +1,10 @@ import { test } from "node:test"; import assert from "node:assert/strict"; -import { classifyDownloadFailure, parseUnavailableFromStderr } from "./availability"; +import { + classifyDownloadFailure, + isSoftBlock, + parseUnavailableFromStderr, +} from "./availability"; // Run with: // pnpm --filter yt-dlp-transcript-common exec tsx --test lib/availability.test.ts @@ -35,3 +39,90 @@ test("HTTP 410 Gone is a removed video, not an error", () => { assert.equal(parseUnavailableFromStderr("HTTP Error 403: Forbidden"), "error"); assert.equal(parseUnavailableFromStderr("HTTP Error 500: Internal Server Error"), "error"); }); + +// YOUTUBE'S SOFT BLOCK (release 10, L2). The strings below are real: +// - the bare line is the one reported in yt-dlp issue #11426 ("[YouTube] This +// content isn't available, try again later.", after ~280 videos of a +// playlist), which is what a yt-dlp older than #12958 prints; +// - the long ones are what `yt_dlp/extractor/youtube/_video.py` builds from that +// reason since #12958 (the build this editor runs, 2026.08.19) — "The current +// session" without cookies, "Your account" with them. +// yt-dlp's wiki (Extractors, "This content isn't available, try again later") +// puts the limit at ~300 videos/hour for a guest session, ~2000 signed in. +const SOFT_BLOCK_SESSION = + "ERROR: [youtube] dQw4w9WgXcQ: This content isn't available, try again later. " + + "The current session has been rate-limited by YouTube for up to an hour. " + + "It is recommended to use `-t sleep` to add a delay between video requests to avoid " + + "exceeding the rate limit. For more information, refer to " + + "https://github.com/yt-dlp/yt-dlp/wiki/Extractors#this-content-isnt-available-try-again-later"; +const SOFT_BLOCK_ACCOUNT = SOFT_BLOCK_SESSION.replace( + "The current session", + "Your account", +); +const SOFT_BLOCK_BARE = + "ERROR: [youtube] H64QQZuw-aA: This content isn't available, try again later."; +// The same reason with a right single quote, in case YouTube sends one. +const SOFT_BLOCK_CURLY = SOFT_BLOCK_BARE.replace("isn't", "isn’t"); + +test("the soft block ('try again later') is a rate limit, not a deleted video", () => { + for (const stderr of [ + SOFT_BLOCK_SESSION, + SOFT_BLOCK_ACCOUNT, + SOFT_BLOCK_BARE, + SOFT_BLOCK_CURLY, + ]) { + assert.equal(isSoftBlock(stderr), true, stderr); + // Not a property of the video: transient, so never excluded as gone. + assert.equal(parseUnavailableFromStderr(stderr), "error", stderr); + // …and a batch-level signal that backs the platform off. + assert.equal( + classifyDownloadFailure(stderr, parseUnavailableFromStderr(stderr)), + "rate_limit", + stderr, + ); + } + // Inside a multi-line tail (a WARNING first, the ERROR last) it still counts. + const tail = `WARNING: [youtube] dQw4w9WgXcQ: nsig extraction slow\n${SOFT_BLOCK_SESSION}\n`; + assert.equal(classifyDownloadFailure(tail, parseUnavailableFromStderr(tail)), "rate_limit"); +}); + +test("a soft block wins over a per-video class a caller already holds", () => { + // A `deleted` read before release 10, or from another line of the same tail, + // must not turn the soft block back into a per-video skip. + assert.equal(classifyDownloadFailure(SOFT_BLOCK_BARE, "deleted"), "rate_limit"); + assert.equal(classifyDownloadFailure(SOFT_BLOCK_SESSION, "members_only"), "rate_limit"); +}); + +test("a genuinely removed or unavailable video is still deleted / per_video", () => { + // yt-dlp's wording for videos that are really gone, soft-block-free. + const gone = [ + "ERROR: [youtube] dQw4w9WgXcQ: Video unavailable", + "ERROR: [youtube] dQw4w9WgXcQ: Video unavailable. This video has been removed by the uploader", + "ERROR: [youtube] dQw4w9WgXcQ: Video unavailable. This video is no longer available because the YouTube account associated with this video has been terminated.", + "ERROR: [youtube] dQw4w9WgXcQ: This video has been removed for violating YouTube's Terms of Service", + "ERROR: [Rumble] v7e07us: Unable to download webpage: HTTP Error 410: Gone (caused by <HTTPError 410: Gone>)", + ]; + for (const stderr of gone) { + assert.equal(isSoftBlock(stderr), false, stderr); + assert.equal(parseUnavailableFromStderr(stderr), "deleted", stderr); + assert.equal( + classifyDownloadFailure(stderr, parseUnavailableFromStderr(stderr)), + "per_video", + stderr, + ); + } + // The pattern's boundary: "content isn't available" with no retry advice + // keeps the `deleted` reading it has had since the availability check began. + const noAdvice = "ERROR: [youtube] dQw4w9WgXcQ: This content isn't available."; + assert.equal(isSoftBlock(noAdvice), false); + assert.equal(parseUnavailableFromStderr(noAdvice), "deleted"); + // The other per-video classes are untouched by the soft-block check. + const priv = + "ERROR: [youtube] dQw4w9WgXcQ: Private video. Sign in if you've been granted access to this video"; + assert.equal(parseUnavailableFromStderr(priv), "private"); + assert.equal(classifyDownloadFailure(priv, "private"), "per_video"); + const members = + "ERROR: [youtube] dQw4w9WgXcQ: Join this channel to get access to members-only content like this video, and other exclusive perks."; + assert.equal(parseUnavailableFromStderr(members), "members_only"); + assert.equal(classifyDownloadFailure(members, "members_only"), "per_video"); +}); diff --git a/common/lib/availability.ts b/common/lib/availability.ts @@ -154,10 +154,39 @@ export type AvailabilityRecord = { export const AVAILABILITY_FILENAME = "availability.json"; +// YouTube's SOFT BLOCK. YouTube answers a session it is rate-limiting with the +// playability reason "This content isn't available, try again later." — worded +// like a removed video, and it is not one: the same video plays for anyone else +// and for this session an hour later. yt-dlp (since #12958, 2025-04) re-words +// it to say so: +// +// ERROR: [youtube] <id>: This content isn't available, try again later. The +// current session has been rate-limited by YouTube for up to an hour. It is +// recommended to use `-t sleep` to add a delay between video requests to +// avoid exceeding the rate limit. For more information, refer to +// https://github.com/yt-dlp/yt-dlp/wiki/Extractors#this-content-isnt-available-try-again-later +// +// ("Your account has been rate-limited …" when cookies were passed). Until +// release 10 its "content isn't available" matched the `deleted` group below, so +// the soft block read as a removed video: per_video, no platform cooldown, the +// batch kept going, and the video was excluded from download as gone. It is a +// rate limit, and it is classified as one: `parseUnavailableFromStderr` says +// "error" (transient) and `classifyDownloadFailure` says `rate_limit`. +// +// "try again later" is what marks it: a bare "Video unavailable" or "This +// content isn't available." with no retry advice is still read as removed. The +// apostrophe may be a right single quote, so match either. +export function isSoftBlock(stderr: string): boolean { + return /isn['’]?t available,? try again later/i.test(stderr); +} + // Map yt-dlp's stderr text to an Availability. yt-dlp does not expose a // stable machine-readable reason on failure, so we pattern-match the human // messages it emits across YouTube, Rumble, and a few other extractors. export function parseUnavailableFromStderr(stderr: string): Availability { + // Checked FIRST: the soft block's wording would otherwise match `deleted` + // below. It says nothing about the video, so it is an error, not a class. + if (isSoftBlock(stderr)) return "error"; const s = stderr.toLowerCase(); if ( /private video/.test(s) || @@ -235,6 +264,11 @@ export function classifyDownloadFailure( stderrTail: string, availabilityClass: Availability | undefined, ): DownloadFailureClass { + // The soft block wins over a per-video class, whichever parser produced it: + // a caller holding a `deleted` read from before release 10 (or from a tail + // that also names a removed video) must still back off. Erring this way costs + // one cooldown; erring the other way keeps requesting into the block. + if (isSoftBlock(stderrTail)) return "rate_limit"; if ( availabilityClass !== undefined && PER_VIDEO_CLASSES.includes(availabilityClass) diff --git a/common/views/jobRowView.ts b/common/views/jobRowView.ts @@ -94,6 +94,10 @@ export type JobRowView = { // jobs/jobDetail.ts from the job's own replay spec; absent for every kind // that has nothing to add, so no existing row changes. detail?: string; + // Why a `cancelled` row was cancelled when no person pressed Cancel — the + // boot pass's reason for a job a restart left queued ("server restarted; + // the scheduler re-derives syncs"). History rows only; see listJobs.ts. + cancelReason?: string; // Which adapter built it. Never rendered; tests and the merge read it — // except "runner": an auto-queue lane's in-flight unit (fromInFlight), which // links to a job page only when `inRegistry` says it has one. diff --git a/common/views/jobRows.test.ts b/common/views/jobRows.test.ts @@ -146,6 +146,28 @@ test("fromEntry keeps the history fields, archived included", () => { assert.equal(row.source, "archive"); // An old log with no sidecar has no kind at all; the cell renders "—". assert.equal(fromEntry({ ...e, kind: undefined }).kind, ""); + // No reason on the entry, no reason on the row (not even the key). + assert.equal("cancelReason" in row, false); +}); + +test("fromEntry carries a cancelled job's cancelReason to the row", () => { + const e: JobListEntry = { + id: newJobId(), + kind: "sync", + channelSlug: "teamrcn", + status: "cancelled", + queuedAt: 1, + endedAt: 9, + inRegistry: false, + replayable: true, + logPath: "/x.log", + logSize: 64, + cancelReason: "server restarted; the scheduler re-derives syncs", + }; + assert.equal( + fromEntry(e).cancelReason, + "server restarted; the scheduler re-derives syncs", + ); }); test("orderLiveRows: running, then queued IN QUEUE ORDER, then what just ended", () => { diff --git a/common/views/jobRows.ts b/common/views/jobRows.ts @@ -282,6 +282,7 @@ export function fromEntry(e: JobListEntry): JobRowView { inRegistry: e.inRegistry, replayable: e.replayable, detail: e.detail, + ...(e.cancelReason ? { cancelReason: e.cancelReason } : {}), source: "archive", }; } diff --git a/common/ytdlp/downloadOneManaged.ts b/common/ytdlp/downloadOneManaged.ts @@ -719,6 +719,35 @@ async function runManagedDownload( } } + // A RATE-LIMITED PREFETCH ENDS THE VIDEO HERE (release 10, L2 review). + // Every failed prefetch used to fall through to the real download, so a + // 429, a bot check or YouTube's soft block ("…isn't available, try again + // later") cost one more request into the same refusal before the batch + // could back off. That request is only skipped for a `rate_limit` class: + // any other failure (a 403, a removed video, an auth gate) keeps today's + // flow, where attempt 1 and its own auth retry still get their chance. The + // record carries `failureClass: "rate_limit"`, so the batch's cooldown and + // abort and the runner's per-video deferral follow exactly as before. + const lastPrefetch = attempts.at(-1); + if ( + lastPrefetch && + !attemptSucceeded(lastPrefetch.ytdlpExitCode) && + classifyDownloadFailure(lastFullTail, lastPrefetch.availabilityClass) === + "rate_limit" + ) { + opts.onLog( + `Metadata prefetch for ${canonicalId} was rate-limited by the source; ` + + `not attempting the download (the platform backs off instead).\n`, + ); + return writeOutcome(opts, videoDir, { + videoId: canonicalId, + status: "failed", + startedAt, + attempts, + lastFullTail, + }); + } + const metaPath = path.join(videoDir, "metadata.info.json"); const metadata = await loadRawMetadata(metaPath); // Only wire --load-info-json into the real attempts when we actually have @@ -1364,24 +1393,53 @@ async function runManagedDownload( } // ---------- Sidecar ---------- + return writeOutcome(opts, videoDir, { + videoId, + status, + startedAt, + attempts, + lastFullTail, + fellBackToTranscribe, + shortAudio: shortAudioInfo, + }); +} + +// THE SIDECAR AND THE AVAILABILITY HISTORY, written one way by every exit that +// ends a download's attempts: the end of the main path, and a prefetch the +// source rate-limited (release 10, L2 review). A filter-skip writes its own +// record and has no failure class; it does not come through here. +async function writeOutcome( + opts: ManagedDownloadOpts, + videoDir: string, + o: { + videoId: string; + status: DownloadOutcomeStatus; + startedAt: string; + attempts: DownloadAttempt[]; + lastFullTail: string; + fellBackToTranscribe?: boolean; + shortAudio?: DownloadOutcomeRecord["shortAudio"]; + }, +): Promise<DownloadOutcomeRecord> { + const { status, attempts } = o; const finishedAt = new Date().toISOString(); // Classify a failed download against the FULL stderr tail of the last attempt // so a rate-limit logged as a WARNING (then masked by a different final error) // is still caught. Skipped for successes and filter-skips. const isFailure = status === "failed" || status === "failed-corrupt-source"; const failureClass = isFailure - ? classifyDownloadFailure(lastFullTail, attempts.at(-1)?.availabilityClass) + ? classifyDownloadFailure(o.lastFullTail, attempts.at(-1)?.availabilityClass) : undefined; const record: DownloadOutcomeRecord = { - videoId, + videoId: o.videoId, webpageUrl: opts.videoUrl, status, - startedAt, + startedAt: o.startedAt, finishedAt, attempts, ...(failureClass ? { failureClass } : {}), - ...(fellBackToTranscribe ? { fellBackToTranscribe: true } : {}), - ...(shortAudioInfo ? { shortAudio: shortAudioInfo } : {}), + ...(o.fellBackToTranscribe ? { fellBackToTranscribe: true } : {}), + ...(o.shortAudio ? { shortAudio: o.shortAudio } : {}), }; // Only write the sidecar if we know which dir to put it in. If the very first // attempt failed before metadata could be written, the data/<id> dir may not diff --git a/common/ytdlp/managedDownloadsSleep.test.ts b/common/ytdlp/managedDownloadsSleep.test.ts @@ -115,13 +115,96 @@ test("a video the download filter declined does not sleep", async () => { }); // Release 9 review: a per-video failure KEEPS the pace, even one that never -// got past the prefetch. YouTube's soft block ("This content isn't available, -// try again later") classifies as deleted → per_video; skipping the sleep there -// would fire prefetches back to back into the block. +// got past the prefetch — a prefetch is still a request. (Release 9's example, +// YouTube's soft block, is no longer per-video: see the release 10 case below.) test("a per-video failure at the prefetch still sleeps", async () => { assert.equal(await sleepsFor([MEMBERS_ONLY, FETCHED]), 1); }); +// THE SOFT BLOCK BACKS OFF (release 10, L2). The outcome is built the way +// downloadOneManaged builds it — availabilityClass from parseUnavailableFromStderr, +// failureClass from classifyDownloadFailure over the same tail — from yt-dlp's +// real line, so this runs the classifier and the batch loop together. A soft +// block must record the platform cooldown and stop the batch; a video that is +// really gone must do neither (it is that video's property, the batch goes on). +const { classifyDownloadFailure, parseUnavailableFromStderr } = await import( + "../lib/availability" +); + +function failedPrefetch(stderr: string): DownloadOutcomeRecord { + const availabilityClass = parseUnavailableFromStderr(stderr); + return outcome( + "failed", + [{ ...PREFETCH, ytdlpExitCode: 1, availabilityClass, error: stderr }], + classifyDownloadFailure(stderr, availabilityClass), + ); +} + +async function batch(outcomes: DownloadOutcomeRecord[]): Promise<{ + downloads: number; + backoffs: string[]; + aborted: boolean; +}> { + let i = 0; + const backoffs: string[] = []; + const urls = outcomes.map( + (_, n) => + `https://www.youtube.com/watch?v=vid${String(n).padStart(8, "0")}`, + ); + const channelConfig = { + handling: "youtube", + url: "https://www.youtube.com/@c/videos", + } as never; + const res = await runManagedDownloads( + { + channelSlug: "c", + mode: "download-missing" as never, + channelConfig, + paths: getPaths(), + onLog: () => {}, + signal: new AbortController().signal, + // The production default: a batch-level failure stops the batch. + onPlatformBackoff: (cls) => { + backoffs.push(cls); + }, + }, + urls, + channelConfig, + undefined, + { + downloadOne: async () => outcomes[i++], + sleep: async () => {}, + }, + ); + return { downloads: i, backoffs, aborted: res.firstFailure !== null }; +} + +test("a soft block backs the platform off and stops the batch; a deleted video does not", async () => { + const softBlock = failedPrefetch( + "ERROR: [youtube] H64QQZuw-aA: This content isn't available, try again later. " + + "The current session has been rate-limited by YouTube for up to an hour. " + + "It is recommended to use `-t sleep` to add a delay between video requests to avoid " + + "exceeding the rate limit. For more information, refer to " + + "https://github.com/yt-dlp/yt-dlp/wiki/Extractors#this-content-isnt-available-try-again-later", + ); + assert.equal(softBlock.failureClass, "rate_limit"); + assert.equal(softBlock.attempts[0].availabilityClass, "error"); + const blocked = await batch([softBlock, FETCHED, FETCHED]); + assert.deepEqual(blocked.backoffs, ["rate_limit"]); + assert.equal(blocked.aborted, true); + assert.equal(blocked.downloads, 1, "nothing after the soft block is requested"); + + const gone = failedPrefetch( + "ERROR: [youtube] dQw4w9WgXcQ: Video unavailable. This video has been removed by the uploader", + ); + assert.equal(gone.failureClass, "per_video"); + assert.equal(gone.attempts[0].availabilityClass, "deleted"); + const kept = await batch([gone, FETCHED, FETCHED]); + assert.deepEqual(kept.backoffs, []); + assert.equal(kept.aborted, false); + assert.equal(kept.downloads, 3, "a removed video does not stop the batch"); +}); + test("the last video never sleeps, whatever it was", async () => { assert.equal(await sleepsFor([FETCHED]), 0); assert.equal(await sleepsFor([FILTERED, FETCHED]), 0); diff --git a/common/ytdlp/metadataScan.ts b/common/ytdlp/metadataScan.ts @@ -21,6 +21,7 @@ import { execa } from "execa"; import { classifyDownloadFailure, isBotCheck, + isSoftBlock, parseUnavailableFromStderr, } from "../lib/availability"; import type { ChannelConfig } from "../lib/channelConfig"; @@ -484,10 +485,17 @@ export async function runMetadataScan( const failure = classifyDownloadFailure(trimmed, cls); // A batch-level signal: continuing would just hammer the source, and on a // bot check every subsequent entry fails the same way. Stop the pass and - // let the caller decide whether cookies are the answer. + // let the caller decide whether cookies are the answer. YouTube's + // explicit soft block ("…isn't available, try again later") lands here + // too since release 10 — before, it was recorded as a `deleted` error on + // its id and the pass kept going into it. if (failure === "rate_limit") { block = { - kind: isBotCheck(trimmed) ? "bot-check" : "rate-limit", + kind: isBotCheck(trimmed) + ? "bot-check" + : isSoftBlock(trimmed) + ? "soft-block" + : "rate-limit", message: trimmed.slice(0, 300), }; child.kill("SIGTERM"); diff --git a/common/ytdlp/prefetchRateLimit.test.ts b/common/ytdlp/prefetchRateLimit.test.ts @@ -0,0 +1,115 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { chmod, mkdtemp, readFile, rm, writeFile } from "node:fs/promises"; +import { tmpdir } from "node:os"; +import path from "node:path"; +import type { Paths } from "../lib/paths"; +import type { ChannelConfig } from "../lib/channelConfig"; +import { downloadOneManaged } from "./downloadOneManaged"; + +// Run with: +// pnpm --filter yt-dlp-transcript-common exec tsx --test ytdlp/prefetchRateLimit.test.ts +// +// A RATE-LIMITED PREFETCH ENDS THE VIDEO (release 10, L2 review). The per-video +// download runs a metadata prefetch first; a failed prefetch used to fall +// through to the real download whatever the failure, so a soft block cost one +// more request into it. Now a `rate_limit` prefetch stops there, and every +// other failure keeps today's flow. The yt-dlp here is a temp script that +// counts its spawns and always fails with the given line — no network, no +// real yt-dlp. + +const VIDEO = "https://www.youtube.com/watch?v=H64QQZuw-aA"; + +async function runWith(stderrLine: string): Promise<{ + spawns: number; + record: Awaited<ReturnType<typeof downloadOneManaged>>; + outcomeOnDisk: { status: string; failureClass?: string }; + log: string; +}> { + const root = await mkdtemp(path.join(tmpdir(), "prefetch-rl-")); + try { + const counter = path.join(root, "spawns"); + const bin = path.join(root, "fake-ytdlp.sh"); + await writeFile( + bin, + `#!/bin/sh\necho x >> '${counter}'\necho ${JSON.stringify(stderrLine)} >&2\nexit 1\n`, + ); + await chmod(bin, 0o755); + const paths = { + channelsDir: path.join(root, "channels"), + ytdlpBin: bin, + } as Paths; + let log = ""; + const record = await downloadOneManaged({ + channelSlug: "c", + channelConfig: { + handling: "youtube", + url: "https://www.youtube.com/@c/videos", + } as ChannelConfig, + paths, + videoUrl: VIDEO, + onLog: (s) => { + log += s; + }, + signal: new AbortController().signal, + }); + const spawns = (await readFile(counter, "utf8").catch(() => "")) + .split("\n") + .filter(Boolean).length; + const outcomeOnDisk = JSON.parse( + await readFile( + path.join(paths.channelsDir, "c", "data", "H64QQZuw-aA", "download-outcome.json"), + "utf8", + ), + ); + return { spawns, record, outcomeOnDisk, log }; + } finally { + await rm(root, { recursive: true, force: true }); + } +} + +test("a soft-blocked prefetch is the only request: no download attempt follows", async () => { + const r = await runWith( + "ERROR: [youtube] H64QQZuw-aA: This content isn't available, try again later. " + + "The current session has been rate-limited by YouTube for up to an hour.", + ); + assert.equal(r.spawns, 1); + assert.equal(r.record.attempts.length, 1); + assert.equal(r.record.attempts[0].kind, "metadata-prefetch"); + assert.equal(r.record.status, "failed"); + assert.equal(r.record.failureClass, "rate_limit"); + // The sidecar is written the same way as on the main path. + assert.equal(r.outcomeOnDisk.status, "failed"); + assert.equal(r.outcomeOnDisk.failureClass, "rate_limit"); + assert.match(r.log, /rate-limited by the source; not attempting the download/); +}); + +test("an HTTP 429 at the prefetch stops the same way", async () => { + const r = await runWith( + "ERROR: [youtube] H64QQZuw-aA: Unable to download webpage: HTTP Error 429: Too Many Requests", + ); + assert.equal(r.spawns, 1); + assert.equal(r.record.failureClass, "rate_limit"); +}); + +test("any other prefetch failure still reaches the download attempt (a 403: two requests)", async () => { + const r = await runWith( + "ERROR: [youtube] H64QQZuw-aA: Unable to download webpage: HTTP Error 403: Forbidden", + ); + assert.equal(r.spawns, 2); + assert.deepEqual( + r.record.attempts.map((a) => a.kind), + ["metadata-prefetch", "primary"], + ); + assert.equal(r.record.status, "failed"); + assert.equal(r.record.failureClass, "network"); + assert.doesNotMatch(r.log, /not attempting the download/); +}); + +test("a removed video still reaches the download attempt too (per-video flow unchanged)", async () => { + const r = await runWith( + "ERROR: [youtube] H64QQZuw-aA: Video unavailable. This video has been removed by the uploader", + ); + assert.equal(r.spawns, 2); + assert.equal(r.record.failureClass, "per_video"); +}); diff --git a/common/ytdlp/runYtdlp.ts b/common/ytdlp/runYtdlp.ts @@ -901,9 +901,12 @@ export type ManagedDownloadsDeps = { // status skipped-filtered, and still asked). // // A per_video FAILURE is deliberately NOT here, even one that never got past -// the prefetch (release 9 review): YouTube's soft block ("This content isn't -// available, try again later") classifies as `deleted` → per_video, so -// skipping the sleep there would fire prefetches back to back into the block. +// the prefetch (release 9 review): a prefetch is still a request, and a run of +// "unavailable" answers is how a throttled source looks (see SOFT_BLOCK_STREAK +// in metadataScan.ts). Release 9's own reason — YouTube's explicit soft block +// ("This content isn't available, try again later") reading as `deleted` → +// per_video — is gone since release 10: it is `rate_limit` now (isSoftBlock), +// backs the platform off and stops the batch. The pace stays anyway. export function declinedWithoutMediaFetch( outcome: DownloadOutcomeRecord, ): boolean { diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -2,6 +2,10 @@ ## [Unreleased] - **umtool's report videos can wear the Archilyzer Media brand, and the channel's YouTube picture, watermark and banner are generated.** A report video whose manifest says `"render": { "brand": "archilyzer-media" }` (umtool's **new project** form now has a brand choice, and `umtool new` takes `--brand archilyzer-media`) is drawn in the channel's slate-and-teal palette and its three faces, with a title card that carries the Archilyzer Media lockup and the found line, the mark at the head of every clip's citation line, a 20-second end card whose right half is left empty for YouTube's end-screen videos (`render.endCard` changes its length or, with `false`, drops it), and a 1280 × 720 thumbnail from a clip still and a headline (`thumbnail` in the manifest; `build-video.mjs --thumbnail` makes it alone). A manifest without the key renders exactly as before, byte for byte. `pnpm --filter yt-dlp-transcript-common exec tsx bin/brand-media.ts` writes the channel picture, the video watermark, the banner and a preview of its phone crop to `~/reports/archilyzer-media/brand/`, with an `INDEX.html` that walks through the YouTube Studio upload. The fonts (Archivo, IBM Plex Sans, IBM Plex Mono) ship in `umtool/report-to-video/fonts/` under the SIL Open Font License; nothing is installed system-wide. +- **YouTube's "try again later" block now backs off instead of reading as a deleted video.** When YouTube rate-limits a session it answers "This content isn't available, try again later." The editor read that as a removed video: the download moved straight on to the next video (into the same block), no cooldown was recorded, and the video was set aside as deleted, so later download runs skipped it. It is now handled like an HTTP 429. A download stops its batch and records the platform cooldown that the auto-download runner and Sync honour, and the runner defers the video for 6 hours. A metadata scan stops (retrying once with cookies when the channel has them) instead of recording the video as gone. The availability check records a temporary error instead of "deleted", and stops probing for the rest of that run. Every rate limit also ends things sooner now: a download whose metadata request is refused no longer goes on to try the download itself, and the availability check stops at the first refused probe instead of asking about every remaining video. When the check that runs before cleaning audio is cut short this way, or has no link to check a video by, the cleanup keeps the audio of every video it did not reach, rather than trusting what an older check said about it. Videos that really are gone ("Video unavailable", removed by the uploader, a terminated account, Rumble's 410) are still recorded as deleted. +- **Jobs a restart left queued are settled even when a drive hangs.** At boot the editor settles those jobs after it has checked where its storage locations are. A hung network mount could stall that check forever, and the jobs then stayed "queued" on `/jobs`. The settling now waits at most 60 seconds, logs `[boot] storage pass still running after 60 s …` and carries on. The storage check keeps running and logs when it ends. A job re-queued for a channel on the hung drive itself still waits for that drive, and any re-queued after it wait too. +- **`/jobs` says why a job was cancelled at boot.** A job the boot settled shows its reason under its status on `/jobs` and as *Cancelled because* on its own page: for example "server restarted; the scheduler re-derives syncs" or "superseded by a newer queued job (…)". The reason used to be only in the job's log. +- **The server log says how often a queued job skips its page refresh.** When a queued job finishes outside any request, the editor skips its page refresh and notes it in the log. The note used to appear once and never again. Now the first one after a quiet spell is logged at once, any more in the next 10 minutes are counted, and one line at the end gives the count, with a running total. ## [0.9.0] - 2026-09-26 - **Every page now has a ground and an accent to choose, and the five theme families are gone.** The theme menu (the palette button beside the quick toggle, in the editor's sidebar and in the header of every published site, the hub and the homepage) has two groups. **Base** is System, Light, Sepia or Dark; Sepia is new, a warm paper ground for long reading. **Accent** is Signal, Brass, Vermilion, Violet, Sakura, Blue or Green, with the site's own tagged *default*; a site with a custom hex offers it first as *Site colour*. The quick toggle cycles System → Light → Sepia → Dark. A published site opens on the reader's system setting, in the accent its site form sets. The hub and the homepage open on Dark, in Signal, even with JavaScript off, and the editor follows the system, in Signal. Each accent has a value for each ground that reads at 4.5:1, and a custom hex is darkened or lightened per ground to match. A reader's accent is remembered only while it differs from the site's: picking the site's own again forgets it, so the reader follows the site if its accent changes later. Base, Archive, Selenized, Swiss and Archilyzer are gone. A choice made before this update carries over once: light stays light (Archive light becomes Sepia), dark stays dark and system stays system; the family itself is dropped. Headings are Archivo, text is IBM Plex Sans and figures are IBM Plex Mono everywhere, with one corner radius. Success, warning and other status text reads at 4.5:1 on its own tinted fill on every ground; on Light, success and warning are a shade deeper than before for it. Chart colours are fixed per ground and never follow the accent; the third is a violet, well clear of the red that marks a recording as gone. The phone's browser bar takes the page's ground, not the accent. Needs a rebuild and deploy of every site, the hub and the homepage. diff --git a/editor/app/jobs/[id]/page.tsx b/editor/app/jobs/[id]/page.tsx @@ -92,6 +92,16 @@ export default async function JobDetailPage({ label="Exit code" value={typeof job.exitCode === "number" ? String(job.exitCode) : "—"} /> + {/* Why, when no person pressed Cancel: the boot pass's reason for a + job a restart left queued (JobMeta.cancelReason). */} + {job.status === "cancelled" && job.cancelReason && ( + <Cell + label="Cancelled because" + value={job.cancelReason} + className="col-span-2 sm:col-span-4" + testId="cancel-reason" + /> + )} </dl> <p className="text-xs text-muted-foreground font-mono"> log: {path.relative(getPaths().monorepoRoot, job.logPath)} @@ -106,9 +116,22 @@ export default async function JobDetailPage({ ); } -function Cell({ label, value }: { label: string; value: string }) { +function Cell({ + label, + value, + className = "", + testId, +}: { + label: string; + value: string; + className?: string; + testId?: string; +}) { return ( - <div className="rounded border border-border px-3 py-2"> + <div + className={`rounded border border-border px-3 py-2 ${className}`} + data-testid={testId} + > <div className="text-xs uppercase tracking-wide text-muted-foreground"> {label} </div> diff --git a/editor/app/jobs/components/JobRow.tsx b/editor/app/jobs/components/JobRow.tsx @@ -122,6 +122,31 @@ export function statusColor(status: string): string { } } +// WHY A JOB WAS CANCELLED, when no person pressed Cancel: the boot pass's +// reason for a job a restart left queued (JobRowView.cancelReason, from the +// meta). Drawn wherever the `cancelled` pill is, as plain text under or beside +// it; absent on every other row. +function cancelReasonOf(job: JobRowView): string | undefined { + return job.status === "cancelled" ? job.cancelReason : undefined; +} + +function CancelReason({ job, inline }: { job: JobRowView; inline?: boolean }) { + const reason = cancelReasonOf(job); + if (!reason) return null; + return ( + <span + data-testid="cancel-reason" + className={ + inline + ? "text-xs text-muted-foreground" + : "self-start max-w-72 text-xs text-muted-foreground break-words" + } + > + {inline ? `· ${reason}` : reason} + </span> + ); +} + // A row with a page at /jobs/<id>. A runner's in-flight unit that is a task // on its runner's job (a transcription, a digest) is not a registry job, so it // does not link to one; a download unit is, and carries its id. @@ -221,6 +246,7 @@ function JobRowHeading({ ? `uppercase tracking-wide px-1.5 py-0.5 rounded text-[10px] ${statusColor(job.status)}` : `text-xs uppercase tracking-wide px-2 py-0.5 rounded ${statusColor(job.status)}` } + title={cancelReasonOf(job)} > {job.status} </span> @@ -277,6 +303,9 @@ function JobRowHeading({ {job.background && " · waiting behind a manual job"} </span> )} + {/* Its counterpart for a cancelled job: why. Compact rows carry it as + the pill's title instead — they are one line, always. */} + {!compact && <CancelReason job={job} inline />} </> ); } @@ -512,6 +541,7 @@ function TableRow({ > {j.status} </span> + <CancelReason job={j} /> {j.stuck && ( <span data-stuck={j.stuck.reason} diff --git a/editor/app/lib/safeRevalidate.test.ts b/editor/app/lib/safeRevalidate.test.ts @@ -1,6 +1,8 @@ import { test } from "node:test"; import assert from "node:assert/strict"; import { + SKIP_REPORT_WINDOW_MS, + SkipReporter, resetSafeRevalidateWarning, runGuarded, safeRevalidate, @@ -10,7 +12,9 @@ import { // // Outside a Next request (which is where a QUEUED job's body runs, and where // this test runs) revalidatePath throws the missing-store invariant. The helper -// must swallow exactly that, warn once per process, and rethrow anything else. +// must swallow exactly that, rethrow anything else, and — since release 10 — +// COUNT every skip: the first in a quiet period is logged at once, the rest of +// its window are reported together when the window ends. function captureWarn<T>(fn: () => T): { result: T; warnings: string[] } { const warnings: string[] = []; @@ -25,16 +29,130 @@ function captureWarn<T>(fn: () => T): { result: T; warnings: string[] } { } } -test("outside a request, safeRevalidate does not throw and warns once", () => { +// A reporter on a hand-driven clock. `advance` moves time and, unless told +// not to, fires the end-of-window timer if it came due — like the event loop. +function harness(windowMs = 60_000) { + let t = 1_000_000; + const lines: string[] = []; + type Pending = { at: number; fn: () => void; live: boolean }; + let pending: Pending | null = null; + const reporter = new SkipReporter({ + windowMs, + now: () => t, + warn: (l) => lines.push(l), + schedule: (fn, ms) => { + const entry: Pending = { at: t + ms, fn, live: true }; + pending = entry; + return { + cancel: () => { + entry.live = false; + }, + }; + }, + }); + const current = (): Pending | null => pending; + return { + reporter, + lines, + advance(ms: number, opts: { fire?: boolean } = {}) { + t += ms; + const p = current(); + if (opts.fire !== false && p && p.live && p.at <= t) { + pending = null; + p.fn(); + } + }, + armed: () => Boolean(current()?.live), + }; +} + +test("outside a request, safeRevalidate does not throw and logs the first skip with its paths", () => { resetSafeRevalidateWarning(); const { warnings } = captureWarn(() => { safeRevalidate(["/channels/x", "/channels", ["/operations/[id]", "page"]]); safeRevalidate(["/"]); }); + resetSafeRevalidateWarning(); // drop the pending end-of-window report const ours = warnings.filter((w) => w.includes("[safeRevalidate]")); + // ONE line for two calls inside one window: the second is counted. assert.equal(ours.length, 1); assert.match(ours[0], /\/channels\/x/); assert.match(ours[0], /\/operations\/\[id\] \(page\)/); + assert.match(ours[0], /1 skip\(s\) since this process started/); + assert.match(ours[0], /next 10 min are counted and reported together/); +}); + +test("one call is one skip, however many targets it names", () => { + const h = harness(); + h.reporter.record(["/a", "/b", "tag:c"]); + assert.equal(h.lines.length, 1); + assert.match(h.lines[0], /skipped revalidating \/a, \/b, tag:c\./); +}); + +test("skips inside the window are counted, and reported once when it ends", () => { + const h = harness(60_000); + h.reporter.record(["/channels/one"]); + h.advance(1_000); + for (let i = 0; i < 5; i++) { + h.reporter.record([`/channels/n${i}`]); + h.advance(1_000); + } + // Nothing more logged while the window is open: no flooding. + assert.equal(h.lines.length, 1); + assert.equal(h.armed(), true); + h.advance(60_000); + assert.equal(h.lines.length, 2); + assert.equal( + h.lines[1], + "[safeRevalidate] 5 more skip(s) with no request store in the last 1 min " + + "(latest: /channels/n4); 6 since this process started.", + ); +}); + +test("a quiet window reports nothing at its end, and the next skip is logged at once", () => { + const h = harness(60_000); + h.reporter.record(["/a"]); + // No second skip: no timer armed, no end-of-window line. + assert.equal(h.armed(), false); + h.advance(10 * 60_000); + assert.equal(h.lines.length, 1); + h.reporter.record(["/b"]); + assert.equal(h.lines.length, 2); + assert.match( + h.lines[1], + /skipped revalidating \/b\. .*2 skip\(s\) since this process started/, + ); +}); + +test("a skip after the window, before its timer ran, reports the old window first", () => { + const h = harness(60_000); + h.reporter.record(["/a"]); + h.advance(30_000); + h.reporter.record(["/b"]); // counted; the report is due at +60 s + // The clock passes the window but the timer has not run yet (a busy loop). + h.advance(40_000, { fire: false }); + h.reporter.record(["/c"]); // a new window: old report first, then this one + assert.equal(h.lines.length, 3); + assert.equal( + h.lines[1], + "[safeRevalidate] 1 more skip(s) with no request store in the last 1 min " + + "(latest: /b); 2 since this process started.", + ); + assert.match(h.lines[2], /skipped revalidating \/c\. .*3 skip\(s\) since/); + // The old timer was cancelled: it reports nothing when it would have fired. + assert.equal(h.armed(), false); +}); + +test("the process-wide reporter is on globalThis, so every module instance shares one", () => { + resetSafeRevalidateWarning(); + captureWarn(() => safeRevalidate(["/x"])); + assert.ok(globalThis.__yttSafeRevalidateSkips__ instanceof SkipReporter); + resetSafeRevalidateWarning(); + assert.equal(globalThis.__yttSafeRevalidateSkips__, undefined); +}); + +test("the production window is 10 minutes", () => { + assert.equal(SKIP_REPORT_WINDOW_MS, 10 * 60 * 1000); }); test("the real revalidatePath does throw here (the premise of the helper)", async () => { @@ -45,17 +163,16 @@ test("the real revalidatePath does throw here (the premise of the helper)", asyn ); }); -test("any other error is rethrown", () => { - resetSafeRevalidateWarning(); +test("any other error is rethrown; a revalidation that runs is not a skip", () => { assert.throws( () => - runGuarded( - () => { - throw new Error("disk on fire"); - }, - "/x", - ["/x"], - ), + runGuarded(() => { + throw new Error("disk on fire"); + }), /disk on fire/, ); + assert.equal( + runGuarded(() => {}), + false, + ); }); diff --git a/editor/app/lib/safeRevalidate.ts b/editor/app/lib/safeRevalidate.ts @@ -13,8 +13,8 @@ import { revalidatePath, revalidateTag } from "next/cache"; // Nothing is lost by skipping the revalidation there: every page these paths // name is dynamic and re-reads disk on the next request, and the snapshot // scheduler revalidates the channel pages itself when it regenerates. So the -// missing-store invariant is swallowed (logged once per process, naming the -// paths), and ANY OTHER error is rethrown — a real failure stays a failure. +// missing-store invariant is swallowed (and reported — see SkipReporter), and +// ANY OTHER error is rethrown — a real failure stays a failure. // // Use it ONLY inside job bodies and job hooks. A plain server action (a form // handler that revalidates and returns) runs inside its request and keeps @@ -24,8 +24,6 @@ export type RevalidateTarget = string | [string, "page" | "layout"]; const MISSING_STORE = /static generation store missing/; -let warned = false; - function isMissingStore(err: unknown): boolean { return err instanceof Error && MISSING_STORE.test(err.message); } @@ -34,25 +32,139 @@ function describe(target: RevalidateTarget): string { return typeof target === "string" ? target : `${target[0]} (${target[1]})`; } +// HOW OFTEN IT HAPPENS, WITHOUT A LINE EACH TIME (release 10, L2). +// +// Release 9 warned once per module instance and then went silent, so the log +// could not say whether this was one job a day or every job in the queue — and +// Next can load a module more than once, so "once" was not even once. Now each +// skip (one `safeRevalidate` call, however many targets it names) is COUNTED, +// on a process-wide holder: +// - the first skip in a quiet period is logged at once, with its paths, and +// opens a SKIP_REPORT_WINDOW_MS window; +// - further skips inside the window are counted, not logged; +// - when the window ends, one line reports how many there were and the latest +// paths (nothing, if there were none). +// So the log carries at most two lines per window, and every skip is in a +// number. The running total since the process started is on every line. +export const SKIP_REPORT_WINDOW_MS = 10 * 60 * 1000; + +type Timer = { cancel: () => void }; + +export type SkipReporterDeps = { + windowMs?: number; + now?: () => number; + // Arms the end-of-window report. The default is an unref'd setTimeout: a log + // line must never hold the process open. + schedule?: (fn: () => void, ms: number) => Timer; + warn?: (line: string) => void; +}; + +function defaultSchedule(fn: () => void, ms: number): Timer { + const t = setTimeout(fn, ms); + t.unref?.(); + return { cancel: () => clearTimeout(t) }; +} + +function minutes(ms: number): string { + const m = Math.round(ms / 60_000); + return m >= 1 ? `${m} min` : `${Math.round(ms / 1000)} s`; +} + +export class SkipReporter { + private readonly windowMs: number; + private readonly now: () => number; + private readonly schedule: (fn: () => void, ms: number) => Timer; + private readonly warn: (line: string) => void; + private total = 0; + // End of the open window; at or before `now()` means no window is open. + private windowEndsAt = 0; + private suppressed = 0; + private latest = ""; + private timer: Timer | null = null; + + constructor(deps: SkipReporterDeps = {}) { + this.windowMs = deps.windowMs ?? SKIP_REPORT_WINDOW_MS; + this.now = deps.now ?? Date.now; + this.schedule = deps.schedule ?? defaultSchedule; + this.warn = deps.warn ?? ((line) => console.warn(line)); + } + + // One skip: one safeRevalidate call that had no request store. + record(labels: readonly string[]): void { + const now = this.now(); + const paths = labels.join(", "); + if (now >= this.windowEndsAt) { + // A quiet period ended (or this is the first skip): report what the last + // window held, if its timer has not already, then this one, at once. + this.flush(); + this.total++; + this.windowEndsAt = now + this.windowMs; + this.warn( + `[safeRevalidate] no request store (a queued job ran outside a request); ` + + `skipped revalidating ${paths}. The pages re-read disk on their next request. ` + + `${this.total} skip(s) since this process started; more in the next ` + + `${minutes(this.windowMs)} are counted and reported together.`, + ); + return; + } + this.total++; + this.suppressed++; + this.latest = paths; + if (!this.timer) { + this.timer = this.schedule(() => { + this.timer = null; + this.flush(); + }, this.windowEndsAt - now); + } + } + + // Report the counted skips, if any. Called by the window's timer, and by the + // first skip of the next window in case the timer has not fired yet. + flush(): void { + if (this.timer) { + this.timer.cancel(); + this.timer = null; + } + if (this.suppressed === 0) return; + this.warn( + `[safeRevalidate] ${this.suppressed} more skip(s) with no request store in the last ` + + `${minutes(this.windowMs)} (latest: ${this.latest}); ` + + `${this.total} since this process started.`, + ); + this.suppressed = 0; + this.latest = ""; + } + + // Test seam: drop the pending report without writing it. + dispose(): void { + this.timer?.cancel(); + this.timer = null; + } +} + +// ONE PER PROCESS, not per module instance: on globalThis, like the job +// registry. Not reset by /api/test/invalidate-cache — it decides only how many +// log lines a skip costs, and nothing reads that log. +declare global { + // eslint-disable-next-line no-var + var __yttSafeRevalidateSkips__: SkipReporter | undefined; +} + +function skipReporter(): SkipReporter { + globalThis.__yttSafeRevalidateSkips__ ??= new SkipReporter(); + return globalThis.__yttSafeRevalidateSkips__; +} + // The guard itself, exported for the unit test so the rethrow path can be -// exercised without a Next runtime. `call` is one revalidation. -export function runGuarded( - call: () => void, - label: string, - allLabels: readonly string[], -): void { +// exercised without a Next runtime. `call` is one revalidation. True when it +// was skipped for want of a request store. +export function runGuarded(call: () => void): boolean { try { call(); + return false; } catch (err) { if (!isMissingStore(err)) throw err; - if (!warned) { - warned = true; - console.warn( - `[safeRevalidate] no request store (a queued job ran outside a request); ` + - `skipped revalidating ${allLabels.join(", ")} — first skip was ${label}. ` + - `Logged once per process; the pages re-read disk on their next request.`, - ); - } + return true; } } @@ -60,23 +172,31 @@ export function safeRevalidate( paths: RevalidateTarget[], tags: string[] = [], ): void { - const labels = [...paths.map(describe), ...tags.map((t) => `tag:${t}`)]; + let skipped = false; for (const target of paths) { - runGuarded( - () => + if ( + runGuarded(() => typeof target === "string" ? revalidatePath(target) : revalidatePath(target[0], target[1]), - describe(target), - labels, - ); + ) + ) { + skipped = true; + } } for (const tag of tags) { - runGuarded(() => revalidateTag(tag, "max"), `tag:${tag}`, labels); + if (runGuarded(() => revalidateTag(tag, "max"))) skipped = true; + } + if (skipped) { + skipReporter().record([ + ...paths.map(describe), + ...tags.map((t) => `tag:${t}`), + ]); } } -// Test seam: the once-per-process latch. +// Test seam: a fresh process-wide reporter (any pending report dropped). export function resetSafeRevalidateWarning(): void { - warned = false; + globalThis.__yttSafeRevalidateSkips__?.dispose(); + globalThis.__yttSafeRevalidateSkips__ = undefined; } diff --git a/editor/e2e/jobs-filters.spec.ts b/editor/e2e/jobs-filters.spec.ts @@ -21,6 +21,9 @@ type SeedJob = { startedAt?: number; endedAt?: number; exitCode?: number; + // Written by the boot pass (common/jobs/bootQueuedJobs.ts) on a job a + // restart left queued. + cancelReason?: string; }; async function seedJob(job: SeedJob): Promise<void> { @@ -65,6 +68,52 @@ test("archived jobs keep their metadata from the sidecar", async ({ page }) => { ); }); +// Release 10 (L2): a job the boot pass closed says WHY, on its row and on its +// page — the reason used to be only in the meta and the log's last line. +test("a job the boot pass cancelled shows its reason on /jobs and on its page", async ({ + page, +}) => { + await resetData(); + const reason = "server restarted; the scheduler re-derives syncs"; + await seedJob({ + id: "boot-cancelled-1", + kind: "sync", + status: "cancelled", + channelSlug: "alpha", + queueKey: "platform:youtube", + queuedAt: 1000, + endedAt: 2000, + cancelReason: reason, + }); + // A cancel with no recorded reason (a person pressed Cancel) draws none. + await seedJob({ + id: "plain-cancelled-1", + kind: "sync", + status: "cancelled", + channelSlug: "alpha", + queueKey: "platform:youtube", + queuedAt: 1000, + startedAt: 1000, + endedAt: 2000, + }); + + await page.goto("/jobs"); + const row = rowById(page, "boot-cancelled-1"); + await expect(row).toContainText("cancelled"); + await expect(row.getByTestId("cancel-reason")).toHaveText(reason); + await expect( + rowById(page, "plain-cancelled-1").getByTestId("cancel-reason"), + ).toHaveCount(0); + + await page.goto("/jobs/boot-cancelled-1"); + const cell = page.getByTestId("cancel-reason"); + await expect(cell).toContainText("Cancelled because"); + await expect(cell).toContainText(reason); + await page.goto("/jobs/plain-cancelled-1"); + await expect(page.getByText("Exit code")).toBeVisible(); + await expect(page.getByTestId("cancel-reason")).toHaveCount(0); +}); + test("filters hide refresh-report by default, persist, and reset", async ({ page, }) => { diff --git a/editor/instrumentation.ts b/editor/instrumentation.ts @@ -119,11 +119,14 @@ export async function register() { // too, but there it ONLY cancels — an idle boot must not resume work — and so // does the e2e test server, whose leftover metas belong to a previous run's // fixture. It waits for the storage pass above, so a channel being - // re-pointed is reachable when its job is re-queued. See + // re-pointed is reachable when its job is re-queued — for at most + // STORAGE_PASS_WAIT_MS (60 s, release 10): a probe stuck on a hung mount + // must not keep every stale meta `queued` for the life of the process. The + // timeout is logged, and the storage pass itself runs on untouched. See // common/jobs/bootQueuedJobs.ts. Lazy, voided, best-effort: never blocks // readiness. try { - const { settleQueuedJobMetas } = await import( + const { settleAfterStoragePass } = await import( "yt-dlp-transcript-common/jobs/bootQueuedJobs" ); const { getPaths } = await import("yt-dlp-transcript-common/lib/paths"); @@ -132,31 +135,23 @@ export async function register() { ); const testServer = process.env.EDITOR_TEST_ROUTES === "1"; const cancelOnly = idle || testServer; - void storagePass - .then(() => - settleQueuedJobMetas({ - paths: getPaths(), - bootedAt, - isLive: (id) => getRegistry().get(id) !== undefined, - log: (line) => console.log(line), - idleReason: idle - ? "idle boot" - : testServer - ? "test server" - : undefined, - requeue: cancelOnly - ? null - : async (spec) => { - const { runJobSpec } = await import("./app/jobs/runJobSpec"); - const res = await runJobSpec(spec); - if (!res.ok) return { ok: false, error: res.error }; - // Nobody reads this stream; release it as the ops routes do. - void res.stream.cancel(); - return { ok: true, jobId: res.jobId }; - }, - }), - ) - .catch(() => {}); + void settleAfterStoragePass(storagePass, { + paths: getPaths(), + bootedAt, + isLive: (id) => getRegistry().get(id) !== undefined, + log: (line) => console.log(line), + idleReason: idle ? "idle boot" : testServer ? "test server" : undefined, + requeue: cancelOnly + ? null + : async (spec) => { + const { runJobSpec } = await import("./app/jobs/runJobSpec"); + const res = await runJobSpec(spec); + if (!res.ok) return { ok: false, error: res.error }; + // Nobody reads this stream; release it as the ops routes do. + void res.stream.cancel(); + return { ok: true, jobId: res.jobId }; + }, + }).catch(() => {}); } catch { /* a boot pass that fails to start must not block server readiness */ } diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -6417,13 +6417,29 @@ Line numbers are `plans/FACTS.md` lines at `e172749b`, before this record's in-p request store and Next throws `Invariant: static generation store missing`, which marked the job `failed` after its work was done. Job bodies and `onDone`/`afterRun` hooks call `safeRevalidate(paths, tags?)` (`editor/app/lib/safeRevalidate.ts`): it swallows exactly that - invariant (one warning per bundle), rethrows anything else. Request-context server actions keep - the plain call. + invariant, rethrows anything else. Request-context server actions keep the plain call. + **Release 10 (L2):** every skip is counted by one `SkipReporter` per process + (`globalThis.__yttSafeRevalidateSkips__`): the first after a quiet spell is logged at once and + opens a 10 min window (`SKIP_REPORT_WINDOW_MS`), the rest are reported as one + `[safeRevalidate] N more skip(s) …; T since this process started.` line when it ends. It was one + warning per module instance. - **`--sleep-requests 1` is a YouTube platform arg** (`common/ytdlp/platformArgs.mjs`), like Rumble's; a channel's `ytdlpExtraArgs` comes after and wins. - **`runManagedDownloads` skips the inter-download sleep only for `skipped-filtered`** (no media request was made). Every failure class still sleeps — YouTube's soft block "try again later" classifies as `deleted`/per_video, so skipping there would hammer a soft block. + **Release 10 (L2): no longer.** `isSoftBlock` (`common/lib/availability.ts`, "isn't available, + try again later", either apostrophe) is checked first: `parseUnavailableFromStderr` → `error`, + `classifyDownloadFailure` → `rate_limit` even over a per-video class. The batch records the + platform cooldown and aborts; the metadata scan stops as a `soft-block` block. Every failure + still sleeps. **L2 review fixes:** a metadata prefetch that classifies `rate_limit` ends + `downloadOneManaged` there (no attempt 1; `writeOutcome` writes the sidecar for both exits), and + `runAvailabilityCheck` stops at its first `rate_limit` probe (`blocked`, `probedIds`, one + `recordDownloadBackoff`). **`verifyBeforeClean` leaves every tier C suspect the check did not probe + `unverified`** (not in `probedIds`: the check was blocked before it, or it had no `webpage_url`) — + never judged on its old `availability.json`, which is what would clear an irreversible delete. No live sidecar held the string (2026-09-26 read-only grep of 136,401 + `availability.json`/`download-outcome.json` and every `metadata-scan.json`), but an availability + check stores no text for `deleted`, so an earlier misread there cannot be found after the fact. - **A job record carries `progressAt`** (live-only, stamped by the one progress writer and by task completion). `/jobs` marks `possibly-stalled` only when tasks are idle AND progress has not moved for 10 min (`common/views/jobRows.ts`). @@ -6432,7 +6448,11 @@ Line numbers are `plans/FACTS.md` lines at `e172749b`, before this record's in-p metas queued > 24 h before the boot are cancelled as stale; of the rest only the newest per kind + channel + bucket + params is re-queued via the retry path; the others are cancelled as superseded; `ARCHILYZER_IDLE_BOOT` cancels everything. A `cancelReason` field on the meta says why. First live - boot: re-queued 1, cancelled 390. + boot: re-queued 1, cancelled 390. **Release 10 (L2):** the wait on the storage pass is bounded + (`settleAfterStoragePass`, `STORAGE_PASS_WAIT_MS` 60 s; a `stat`/statfs on a hung mount has no + timeout of its own) and logged when it fires; the storage pass is not cancelled. `cancelReason` + is drawn on `/jobs` (`JobListEntry` → `JobRowView.cancelReason`, only on a `cancelled` meta; + `data-testid="cancel-reason"` on the row and the job page's "Cancelled because" cell). - **The hub embeds `public/hub-summary.json`** at `compose:hub` (`common/controller/poolSummary.ts` shared with `compose-homepage`; `common/lib/hubSummary.ts` projects `official` + per-site figures from the same `buildHomepageSummary`). Optional end to end: missing/404/malformed → cards without diff --git a/plans/release-10.md b/plans/release-10.md @@ -224,6 +224,254 @@ safeRevalidate helper. **S4 — Archilyzer Media** (the YouTube channel's assets and umtool's opt-in `render.brand` preset) merges after the release-10 cut, from `brand/media`, and is **not part of the site rollout**: it touches only `common/lib` + `common/bin` and umtool; see `brand-and-themes.md`, "Slice S4, as shipped". +### Slice L2, as shipped — runner lows (2026-09-26) + +Branch `r10/runner-lows` off `main` `5dfc9c3a`, one Opus implementer, beside L1 (hub lows) and brand +S4. Items 6–9 of "Planned: the lows". Nothing on disk moves, and no settings, site or channel key +changes; `JobListEntry.cancelReason` and `JobRowView.cancelReason` are additive and optional. + +**6 — YouTube's soft block backs off.** YouTube answers a session it is rate-limiting with the +playability reason "This content isn't available, try again later." Its "content isn't available" +matched `parseUnavailableFromStderr`'s `deleted` group, so it was `per_video`: no platform cooldown, +the batch kept requesting into the block, and the video joined `EXCLUDED_FROM_DOWNLOAD` as gone. +`isSoftBlock` (`common/lib/availability.ts`, `/isn['’]?t available,? try again later/i`) is now +checked first in both functions: `parseUnavailableFromStderr` → `error` (transient: not excluded, not +pinned as gone), `classifyDownloadFailure` → `rate_limit`, even over a per-video class a caller +already holds. What that drives is the existing path, unchanged: `runManagedDownloads` records the +platform cooldown and aborts the batch; the auto runner's `applyUnitOutcome` backs the platform off +and defers the video 6 h; `fetchWindowManaged` records the cooldown; the metadata scan stops the +pass (its block `kind` is `"soft-block"` for this line, which only changes the log wording; the one +cookie retry every block gets is unchanged). "Try again later" is what marks it. A bare "Video +unavailable", "…removed by the uploader", a terminated account, a ToS removal and Rumble's `HTTP Error +410: Gone` stay `deleted` / `per_video`, and so does "This content isn't available." with no retry +advice. Every failure still sleeps between downloads; the release 9 comments that gave the soft block +as the reason now say why the pace stays anyway. +- **The strings, and where they came from.** No sidecar on the live corpus holds one (read-only grep, + 2026-09-26: 136,401 `availability.json` / `download-outcome.json` files and every + `metadata-scan.json`, zero hits for "try again later", "content isn't available" or "rate-limited by + YouTube"). The tests use: + - `ERROR: [youtube] H64QQZuw-aA: This content isn't available, try again later.` — the line + reported in yt-dlp issue #11426 (what a yt-dlp before #12958 prints); + - `ERROR: [youtube] <id>: This content isn't available, try again later. The current session has + been rate-limited by YouTube for up to an hour. It is recommended to use `-t sleep` to add a delay + between video requests to avoid exceeding the rate limit. For more information, refer to + https://github.com/yt-dlp/yt-dlp/wiki/Extractors#this-content-isnt-available-try-again-later` — + built by `yt_dlp/extractor/youtube/_video.py` (yt-dlp #12958, commit `26feac3dd`) in the build the + editor runs (`~/Projects/yt-dlp-patched`, 2026.08.19); "Your account" instead of "The current + session" when cookies were passed; + - the bare line with a right single quote. + yt-dlp's wiki puts the limit at ~300 videos/hour for a guest session, ~2,000 signed in. + +**7 — the boot pass's wait on the storage pass is bounded.** `settleAfterStoragePass` +(`common/jobs/bootQueuedJobs.ts`) waits for the storage pass for at most `STORAGE_PASS_WAIT_MS`, +then settles anyway; `instrumentation.ts` calls it in place of `storagePass.then(settle…)`. +- **60 s, and why.** A healthy pass takes milliseconds: `findmnt -T` answers in under 10 ms on this + machine. Its bounded worst case is two or three `findmnt`s per location at `FINDMNT_TIMEOUT_MS` + (3 s) — identity and fstab, or where the uuid is mounted — plus, for a location being re-pointed, + the preflight's second probe and a stat per channel on it: about 12 s a location. Nothing bounds + the `stat` of a location's root or its statfs, and on a hung network mount they never return. 60 s + covers several locations at their worst; past it the pass is stuck on a syscall, and waiting buys + nothing. +- **When it fires:** `[boot] storage pass still running after 60 s; settling queued jobs without it + (a hung mount? a job re-queued for a channel on it will wait on the same mount, and the re-queues + after it with it)`, and, when the pass finally ends, `[boot] storage pass finished N s after the + queued-job pass began waiting (it stopped waiting at 60 s)`. The race is over a derived promise, + so the storage pass is never cancelled and a re-point it enqueued runs to the end. A pass that + throws counts as finished. +- **What a re-queue then meets** (corrected in review; as first shipped this said every such job is + refused at once). For an UNMOUNTED drive, the media guard's `stat` gets ENOENT at once, so the job + is refused — at submission, closing the old meta `cancelled` with the error, or when it starts — + and `/jobs` says so. For a HUNG mount, the case the bound exists for, the guard's own `stat` + (`lib/channelMedia.ts`) hangs on the same syscall; re-queues run one at a time, so the ones after + it stay `queued` until the mount answers. Every cancel has run by then. Left: see below. + +**8 — `/jobs` shows `cancelReason`.** `JobListEntry.cancelReason` is read from the meta only when its +status is `cancelled`; `fromEntry` copies it to `JobRowView.cancelReason`. `JobRow.tsx` draws it +wherever the cancelled pill is: under the pill in the table's Status cell, at the end of a card row's +heading, and as the pill's `title` on a compact row. The job page (`/jobs/[id]`) adds a full-width +**Cancelled because** cell. All carry `data-testid="cancel-reason"`. Everything is added; no existing +label, text or test id changed (`editor/e2e` had no `cancel-reason`, `Cancelled because` or detail-page +cell assertions to collide with). A live row never has one: only the boot pass writes it, to metas +from before the boot. + +**9 — the safeRevalidate warning is counted.** `SkipReporter` (`editor/app/lib/safeRevalidate.ts`), +one per process on `globalThis.__yttSafeRevalidateSkips__` (LOW-4: one per module instance was not +even once). A skip is one `safeRevalidate` call with no request store, however many targets it names. +The first after a quiet spell is logged at once with its paths, under the old prefix, and opens a +`SKIP_REPORT_WINDOW_MS` (10 min) window. The rest of the window are counted, and one line at its end +reports them: `[safeRevalidate] N more skip(s) with no request store in the last 10 min (latest: +…); T since this process started.` So at most two lines per window. The end-of-window timer is +unref'd. It is not reset by `/api/test/invalidate-cache`: it decides only how many log lines a skip +costs, and nothing reads that log. + +| sha | what | +|---|---| +| `deaff1d9` | 6: `isSoftBlock`; `parseUnavailableFromStderr` → `error`, `classifyDownloadFailure` → `rate_limit`; the scan's `soft-block` kind; stale comments; `availability.test.ts` +3, `managedDownloadsSleep.test.ts` +1 | +| `d2ff1411` | 7: `STORAGE_PASS_WAIT_MS`, `waitForStoragePass`, `settleAfterStoragePass`, `instrumentation.ts`; `bootQueuedJobs.test.ts` +5 | +| `904ffc4e` | 8: `cancelReason` through `listJobs` → `fromEntry` → `JobRow` and the job page; `listJobs.test.ts` +1, `jobRows.test.ts` +1, `jobs-filters.spec.ts` +1 | +| `11446fba` | 9: `SkipReporter`, one per process; `safeRevalidate.test.ts` 3 → 9 | +| `fc2ce63c` | 8: a card row's reason at the end of its heading, not beside the pill | +| `5b657aeb` | `plans:` this record, FACTS (three release 9 bullets amended), the editor `[Unreleased]` bullets | + +**Gates**, all from the worktree root. +- **tsc** (`pnpm -r --no-bail --workspace-concurrency=1 exec tsc --noEmit`) clean before every code + commit (`l2-tsc-1..5.log`). +- **common 1,924/1,924** (1,913 + 11: availability +3, managedDownloadsSleep +1, bootQueuedJobs +5, + listJobs +1, jobRows +1); **editor unit 85/85** (79 + 6, `safeRevalidate.test.ts`); + **`test:scripts` 162 + 1 skip of 163**; **mcp 219/219** (`l2-units.log`, on `11446fba`). +- **Editor build** `pnpm --filter editor exec next build` ok: 89 s on `11446fba` (`l2-build.log`) and + 52 s on the code tip `fc2ce63c` (`l2-build-2.log`). + `export/public` in the worktree: the generated paths linked from the primary, `sw.js` a plain copy. +- **EDITOR e2e.** Spec list `l2-specs.txt`, 37 specs: every spec that opens `/jobs` or its jobs + (`jobs jobs-filters jobs-retry jobs-channel jobs-active-order jobs-reorder jobs-batch-tasks-drain + job-stream-cancel cancel queue queues dashboard widget channel-work`), the boot and runner specs + (`auto-queue lane-runner ops-api backfill`), downloads and their classification (`availability + availability-backfill metadata-scan-softblock metadata-scan-botcheck pacing retry-bucket + rumble-sweep fetch-window undownloaded download-part-files partial-downloads-bucket title-filter + pre-clean-availability maybe-missing`), sync (`sync-deep sync-break-on-existing scheduler`), a + revalidating job (`truncated-check`) and `storage-locations`, all `.spec.ts`. A content grep for + the prompt's five words matches 122 of 123 specs ("sync" is in "async"), so the list is by name and + by what each spec drives. + - Run 1 (`l2-e2e-1.log`, on `fc2ce63c`, 7 min in the queue behind L1): **213 passed, 2 failed, + 13.9 min** (215 tests). `ops-api.spec.ts:447` (B1) printed the new first line live: `[safeRevalidate] + no request store …; skipped revalidating /channels/slow-b, /channels, /operations/[id] (page), + /cleanup, /. … 1 skip(s) since this process started; more in the next 10 min are counted and + reported together.` + - `channel-work.spec.ts:208`, 233 ms: `EEXIST: file already exists, mkdir + '…/editor/test-transcripts/channels'` in `resetData`'s fixture copy (`helpers.ts:69`), before the + test touched the page — the pre-existing fixture-reset race S3 met in `pipeline.spec.ts:164`. + - `lane-runner.spec.ts:203`, 7.9 s: `apiRequestContext.get: read ECONNRESET` on a + `/api/auto-queue/status` poll; the next four lane-runner tests passed. + - Rerun alone, `channel-work.spec.ts:208 lane-runner.spec.ts:203 --repeat-each 3` + (`l2-e2e-rerun.log`, 4m55s in the queue behind S4): **6 passed, 0 failed, 1.0 min**. Both are + **pre-existing flakes**, not this slice: the EEXIST is the fixture-reset race in the helper, and the + reset was a socket dropped mid-poll by the dev server. The slice's diff touches neither + (`git diff --stat 5dfc9c3a HEAD -- editor/e2e/helpers.ts editor/app/api/auto-queue + common/controller/autoRunner.ts common/jobs/autoQueueState.ts editor/app/channels` is empty). + - The primary checkout's `export/public/sw.js` was not written (10,027 B, md5 `55cbf381…`, mtime + 2026-09-25 20:47:37, before and after). +- **Numbers tools: none.** + +**Found and left.** +- ~~A soft-blocked prefetch still makes the primary attempt~~ and ~~the availability check has no + rate-limit stop~~: both done in the review fixes, below. +- **A hung mount can hold the boot pass's re-queues.** After the 60 s timeout, a job re-queued for a + channel on a hung mount waits on the media guard's `stat`, and the re-queues after it wait with it + (above). Cancels are unaffected, and re-queues are few (one at the first live boot). A bound on the + guard's `stat`, or re-queues that do not wait on each other, would fix it. +- **An old yt-dlp's soft-block line on a channel listing.** A flat-playlist listing that fails with + the bare pre-#12958 line now classifies `rate_limit`, so `enumeratePlaylistUrls` throws + `EnumerationIncompleteError` (`runYtdlp.ts:420`) and a full sweep falls back to its paged walk (one + more page request). The current yt-dlp's longer line took that path before this slice (it matches + `/rate-limit/`), and a flat listing never reaches the playability check that prints it. +- **Manual batches have no per-video deferral.** A video that answers "try again later" every time + stops every manual download-missing at itself, on every run; the auto runner defers it 6 h + instead. It errs toward backing off. +- **Earlier misreads cannot be found.** An availability check stores no stderr for a `deleted` + result, so a soft block it read before this release is indistinguishable from a real removal. None + shows in the download outcomes or scan stores (above). +- **No e2e drives the soft block through a download.** The fake yt-dlp has no sentinel for the + "try again later" line; the chain is pinned by the classifier tests and by + `managedDownloadsSleep.test.ts`, which feeds the real line through the classifier into the batch + loop. +- `plans/STATE.md` still lists items 6–9 as open; it is shared with L1, so the merge updates it. +- ~~An unblocked tier C check judges a suspect with no `webpage_url` on its old record~~ + (pre-existing, seen while fixing): fixed in the re-read, below. +- **New low: `resolveMaybeMissingState` reads an `error` probe newer than the scan as `available`** + (`stateFromAvailability`'s default; pre-existing, seen while fixing). A maybe-missing video whose + confirm probe failed (a 403, a network error, or the one rate-limited probe of a blocked run) + stops being flagged. It is display-only: it feeds the maybe-missing state in `buildIndex.ts` + (~:1180), i.e. the published site's presence badge. No deletion path reads it; the clean gate uses + `resolveEffectiveAvailability`. The fix is one line (`error` → `maybe_missing`), but it changes + published presence semantics and the maybe-missing count, so it needs its own export check: not + in this slice. Before this slice the soft block published a false "deleted", which was worse. + +**Review fixes (review verdict SHIP AFTER FIXES, `l2-review.md`).** + +- **Item 7's wording (should-fix).** The comment, the timeout line and this record promised that a job + re-queued for a channel the storage pass had not reached is refused at once. True for an unmounted + drive; not for a hung mount, where the media guard's own `stat` hangs too and the sequential + re-queues behind it stay `queued`. Corrected in all three; listed under found and left. No + behaviour change. +- **A rate-limited prefetch ends the download (question a).** In `runManagedDownload`, after the + prefetch's auth retry: when the last prefetch attempt classifies `rate_limit`, one log line + (`Metadata prefetch for <id> was rate-limited by the source; not attempting the download …`) and + the record `failed` / `rate_limit` with only the prefetch attempt(s). The sidecar and + availability-history tail is one helper, `writeOutcome`, which both exits call. Every other failure + class keeps today's flow. The batch's cooldown and abort, and the runner's deferral, follow from the + record unchanged. +- **The availability check stops on a rate limit (question b), with the data-loss guard in the same + commit.** + - `runAvailabilityCheck`: the first probe that classifies `rate_limit` records its own `error` + and sets `blocked`. No further probe starts; the rest count as `skipped`, and a `STOPPED: …` line + says how many. After the pass the platform cooldown is recorded once, best-effort, with + `recordDownloadBackoff(detectPlatform(config?.url ?? url) ?? "unknown", paths)`, the key the batch + and the Sync gate use (`onPlatformBackoff` is the test seam). The result gains `blocked`, + `blockMessage` and `probedIds`. + - **`verifyBeforeClean`: every tier C suspect the check did not probe is `unverified`.** Judged on + its old `availability.json` (a months-old `public`), it would have been cleared for an + irreversible delete. As first fixed (`b8053655`) this applied only to a BLOCKED check. The re-read + (`52bb00ef`) dropped that condition, so a suspect an unblocked check skipped for having no + `webpage_url` is unverified too. That closes a pre-existing gap. The review's read-only scan found + no such video on the live corpus (79,700 dirs, 1,186 clean candidates), so nothing changes today; + such a video is simply never cleaned until it has a `webpage_url`. A suspect every probe reached + is judged exactly as before. + - The other callers are unaffected, as the review said. `checkKeptDeleted` only pins, and an unprobed + id keeps its old verdict (pinning is the safe direction). The full-sweep confirm leaves an + unprobed id `maybe_missing`, because its probe time is older than the scan's. +- **Two found-and-left lines** (above): an old yt-dlp's soft-block line on a flat listing, and manual + batches stopping at a persistently soft-blocked video. + +| sha | what | +|---|---| +| `61c03ac1` | (a) the rate-limited prefetch exit, `writeOutcome`; `prefetchRateLimit.test.ts` (4) | +| `24e0f7d5` | item 7's comment and timeout line say what a hung mount does | +| `b8053655` | (b) the availability check's stop + cooldown + `probedIds`; `verifyBeforeClean` leaves unprobed suspects unverified; `checkAvailability.test.ts` (3), `verifyBeforeClean.test.ts` +2 | +| `89218c38` | `plans:` this paragraph, the item 7 record text, found and left, FACTS, the `[Unreleased]` bullets | +| `52bb00ef` | (re-read) every unprobed tier C suspect is unverified, blocked or not; `verifyBeforeClean.test.ts` +1 (a no-URL suspect in an unblocked check) | +| _this_ | `plans:` the re-read fix and the `resolveMaybeMissingState` low in this record, FACTS | + +**Tests, and the proof they bite.** +- `prefetchRateLimit.test.ts` drives the real `downloadOneManaged` against a temp yt-dlp script that + counts its spawns. A soft block gives 1 spawn, 1 attempt and a `rate_limit` sidecar, and a 429 is + the same. A 403 still reaches the primary (2 spawns, `network`), and so does a removed video (2, + `per_video`). With the exit disabled, the soft-block and 429 cases fail. +- `checkAvailability.test.ts` covers four probeable videos: + - a soft block gives 1 spawn, 3 skipped, one `error` written and one youtube cooldown; + - a 429 and the bot check behave the same; + - a 403 and a removed video probe all four, with no cooldown. +- `verifyBeforeClean.test.ts` runs through the real quick check and tier C: + - a blocked confirmation leaves all 3 suspects `unverified`, although two carry a stale `public`; + the listed video stays cleanable, and the cooldown lands in the state file; + - an unblocked one probes all 3 and judges `public` / `deleted` / 403 as before; + - with the guard disabled, the blocked case fails. + +**Re-gate on `b8053655`:** +- tsc clean before each fix commit (`l2-tsc-6..8.log`). +- Unit tests (`l2-units-2.log`): common **1,933/1,933** (1,924 + 9), editor unit **85/85**, `test:scripts` + **162 + 1 skip of 163**, mcp **219/219**. +- Editor `next build` ok, 43 s (`l2-build-3.log`). +- EDITOR e2e (`l2-e2e-2.log`): the 37 specs plus 9 cleanup / availability specs (`cleanup-actionable + cleanup-holds cleanup-page do-not-clean shard cookies-mode whisper channels-actions + auto-subs-replace`; `pre-clean-availability`, `availability`, `availability-backfill` and + `maybe-missing` were already in the list). **256 passed, 0 failed, 16.2 min**, no queue wait. +- The primary's `export/public/sw.js` was untouched again. + +**Re-read fix (`l2-review.md`, "Re-read of fixes": no must-fix, one should-fix), on `52bb00ef`.** +- The fix is `52bb00ef`, above. +- The new `verifyBeforeClean.test.ts` case puts a no-URL suspect with a stale `public` into an unblocked + check: + - only the other suspect is probed; + - the no-URL one comes back `unverified` and excluded; + - the probed `public` suspect and the listed video stay cleanable; + - no cooldown is recorded. + With the old blocked-only condition it fails; the blocked and unblocked cases still pass. +- tsc clean (`l2-tsc-9.log`). +- common **1,934/1,934** (+1, `l2-units-3.log`). +- EDITOR e2e `cleanup-actionable cleanup-holds cleanup-page do-not-clean pre-clean-availability` + (`l2-e2e-3.log`): **12 passed, 0 failed, 1.1 min** (9 s in the queue). `pre-clean-availability` + was added to the four named specs because it drives all three tiers of the gate end to end. + ## Rollout Nothing is rolled out. The live :3001 editor still runs `0213f6c8` (the pre-brand build); the five