import { link, mkdir, readFile, writeFile } from "node:fs/promises"; import { join } from "node:path"; import { test, expect, type APIRequestContext, type Page } from "@playwright/test"; import { ulid } from "yt-dlp-transcript-common/jobs/ulid"; import { baseUrl } from "./baseUrl"; import { channelStage, generateReport, resetData, resolvePath, } from "./helpers"; // THE DASHBOARD ANSWERS WHILE A REPORT REGENERATES (release 17 slice D0). // // On 2026-10-01 the live editor's `/`, `/channels` and `/jobs` gave no // response for over an hour while two `refresh-report` jobs walked 2,000- and // 3,260-video channels in its own process, side by side, and three more // `refresh-report` metas still read `running` hours after the process that ran // them was gone. The first case below is the first half: two channels whose // reports take several seconds each, regenerated by "Update all reports" — one // after the other, never side by side — with `/` and `/jobs` polled every 2 s // throughout. The second is the ghost: a `running` meta a dead process left // behind is closed as interrupted by the boot pass, one a live process owns is // not, and neither stops a media move. // // THE BUDGET IS 5 s, OR THREE TIMES WHAT THE SAME PAGE TOOK JUST BEFORE THE // REGENERATION, WHICHEVER IS LONGER. The suite runs under `next dev` on a // machine other suites and builds share; at a load average of 30 a page that // renders in 0.3 s on a quiet machine takes 2–4 s with nothing regenerating at // all. What this pins is that the regeneration does not starve the pages, not // how fast the machine is — so the idle measurement taken a moment before sets // the floor, and on a quiet machine the budget is simply 5 s. CAPPED AT 15 s: // on a machine so loaded that idle pages take over 5 s, the floor stops // growing, and the case fails rather than stretching to hide starvation. // // What the case pins, honestly: the serial queue (read from the jobs' own // records) and starvation at the scale of seconds. The yield between chunks is // pinned by a unit test (common/controller/snapshotYield.test.ts): on a dev // server under load `main`'s walk answered inside such a budget too. const OPS_AUTH = { authorization: "Bearer test-worker-token" }; // THE BIG CHANNELS. Every video dir hardlinks ONE ~400 KB metadata.info.json // — the size of a long VOD's, the file the walk parses — so 600 of them cost // one file's bytes and a second to make. No `webpage_url`: the reconcile pass // at the walk's start parses each file for it and, finding none, renames // nothing (one shared id would merge every dir into one). The archive names a // non-YouTube extractor, so the walk parses each file a second time for its // native id — the Rumble-channel path. 600 is several seconds of walk under // `next start` on a quiet machine and well over a minute under `next dev` at a // load average of 30; 2,000 did not finish inside five minutes there. const BIG = ["big-channel-a", "big-channel-b"]; const BIG_VIDEOS = 600; async function seedBigChannel(slug: string): Promise { const channelDir = resolvePath(`test-transcripts/channels/${slug}`); const dataDir = join(channelDir, "data"); await mkdir(dataDir, { recursive: true }); await writeFile( join(channelDir, "config.json"), JSON.stringify({ handling: "youtube", name: "Big Channel" }), ); const formats = Array.from({ length: 400 }, (_, i) => ({ format_id: `hls-${i}`, url: `https://example.invalid/${"x".repeat(600)}${i}`, ext: "mp4", protocol: "m3u8_native", width: 1280, height: 720, tbr: 1234.5 + i, http_headers: { "User-Agent": `Mozilla/5.0 ${"y".repeat(80)}`, Accept: "*/*" }, fragments: [ { url: "seg0", duration: 6 }, { url: "seg1", duration: 6 }, ], })); const template = resolvePath(`test-transcripts/.${slug}-meta.json`); await writeFile( template, JSON.stringify({ id: "native", title: "A long VOD", duration: 3600, formats }), ); const ids = Array.from( { length: BIG_VIDEOS }, (_, i) => `${slug.slice(-1)}big${String(i).padStart(8, "0")}`, ); for (let i = 0; i < ids.length; i += 100) { await Promise.all( ids.slice(i, i + 100).map(async (id) => { const dir = join(dataDir, id); await mkdir(dir); await link(template, join(dir, "metadata.info.json")); }), ); } await writeFile( join(channelDir, "archive"), ids.map((id) => `rumble ${id}\n`).join(""), ); await writeFile(join(channelDir, "playlist"), ""); } type Sample = { path: string; ms: number; status: number }; async function timedGet( request: APIRequestContext, path: string, ): Promise { const t = Date.now(); const res = await request.get(`${baseUrl}${path}`, { timeout: 60_000 }); // The body too: a page that sends its head and stalls is not an answer. await res.body(); return { path, ms: Date.now() - t, status: res.status() }; } type JobMetaOnDisk = { id: string; kind: string; channelSlug?: string; status: string; startedAt?: number; endedAt?: number; }; test("/ and /jobs answer while two large reports regenerate, one after the other", async ({ request, }) => { test.setTimeout(420_000); await resetData("empty"); for (const slug of BIG) await seedBigChannel(slug); // Warm both routes first: under `next dev` the first request compiles the // page, which is the dev server's cost and not what this measures. Then the // idle floor: the slowest of three more rounds, with nothing regenerating. for (const path of ["/", "/jobs"]) { expect((await timedGet(request, path)).status).toBe(200); } // And the ops route that starts the regeneration: its first request compiles // it, which under load held `/` for 15 s in one full-suite run — a dev // server's compile, not a walk. A body naming neither form is refused (400) // after the route has loaded, and starts nothing. const warm = await request.post(`${baseUrl}/api/ops/refresh-report`, { headers: OPS_AUTH, data: {}, timeout: 120_000, }); expect(warm.status()).toBe(400); const idle: Sample[] = []; for (let i = 0; i < 3; i++) { for (const path of ["/", "/jobs"]) idle.push(await timedGet(request, path)); } const budgetMs = Math.min( Math.max(5_000, 3 * Math.max(...idle.map((s) => s.ms))), 15_000, ); // "Update all reports", through the ops API: a refresh-report job per // channel on the serial queue. The call answers once they are QUEUED, with // their ids; "while they regenerate" lasts until both jobs' records say // they have ended. const startedAt = Date.now(); const res = await request.post(`${baseUrl}/api/ops/refresh-report`, { headers: OPS_AUTH, data: { all: true }, timeout: 120_000, }); expect(res.status()).toBe(200); const done = { body: (await res.json()) as { queued?: string[]; jobIds?: string[] } }; expect([...(done.body.queued ?? [])].sort()).toEqual(BIG); expect(done.body.jobIds).toHaveLength(2); const ended = async (): Promise => { for (const id of done.body.jobIds ?? []) { const m = await readFile(resolvePath(`test-transcripts/.jobs/${id}.meta.json`), "utf8") .then((raw) => JSON.parse(raw) as JobMetaOnDisk) .catch(() => null); if (!m || m.status === "queued" || m.status === "running") return false; } return true; }; let finishedAt: number | null = null; const during: Sample[] = []; while (finishedAt === null) { if (Date.now() - startedAt > 360_000) throw new Error("the regenerations did not end in 6 min"); const tick = Date.now(); for (const path of ["/", "/jobs"]) { const s = await timedGet(request, path); during.push(s); console.log(`[dashboard-answers] +${tick - startedAt} ms ${path} ${s.ms} ms`); } if (await ended()) { finishedAt = Date.now(); break; } const wait = 2_000 - (Date.now() - tick); if (wait > 0) await new Promise((r) => setTimeout(r, wait)); } for (const slug of BIG) { const snapshot = JSON.parse( await readFile( resolvePath(`test-transcripts/channels/${slug}/snapshot.json`), "utf8", ), ) as { totals: { videos: number } }; expect(snapshot.totals.videos).toBe(BIG_VIDEOS); } // ONE AFTER THE OTHER: the second started when the first had ended. Read // off the jobs' records on disk; the terminal write lands just after the // job's stream closes, so it is waited for. const readMetas = () => Promise.all( (done.body.jobIds ?? []).map( async (id) => JSON.parse( await readFile(resolvePath(`test-transcripts/.jobs/${id}.meta.json`), "utf8"), ) as JobMetaOnDisk, ), ); await expect .poll(async () => (await readMetas()).map((m) => `${m.kind} ${m.status}`), { timeout: 10_000, }) .toEqual(["refresh-report done", "refresh-report done"]); const metas = await readMetas(); const [first, second] = [...metas].sort( (a, b) => (a.startedAt ?? 0) - (b.startedAt ?? 0), ); expect(second.startedAt ?? 0).toBeGreaterThanOrEqual(first.endedAt ?? Infinity); // The fixture has to make the walk long enough to be polled through, or the // budget below proves nothing. const regenMs = (finishedAt ?? Date.now()) - startedAt; test.info().annotations.push({ type: "timings", description: `idle max ${Math.max(...idle.map((s) => s.ms))} ms, budget ${budgetMs} ms; ` + `regeneration ${regenMs} ms; ${during.map((s) => `${s.path} ${s.ms}`).join(", ")}`, }); expect(regenMs, "the regenerations took several seconds").toBeGreaterThan(4_000); for (const path of ["/", "/jobs"]) { expect( during.filter((s) => s.path === path).length, `${path} was polled during the regeneration`, ).toBeGreaterThanOrEqual(2); } for (const s of during) { expect(s.status, s.path).toBe(200); expect(s.ms, `${s.path} answered in ${s.ms} ms (budget ${budgetMs} ms)`).toBeLessThan(budgetMs); } }); // --------------------------------------------------------------------------- // The ghost. const SLUG = "test-youtube"; async function quiet(page: Page): Promise { await expect .poll( async () => { const res = await page.request.get(`${baseUrl}/api/jobs/active`); const body = await res.json(); const jobs: { channelSlug?: string; status: string }[] = Array.isArray( body, ) ? body : (body.jobs ?? []); return jobs.filter( (j) => j.channelSlug === SLUG && (j.status === "running" || j.status === "queued"), ).length; }, { timeout: 30_000 }, ) .toBe(0); } // A `running` refresh-report meta for SLUG, as a process that died mid-walk // leaves it: queued and started an hour ago, its log last written then. async function plantGhost(pid: number): Promise { const id = ulid(Date.now() - 60 * 60 * 1000); const jobsDir = resolvePath("test-transcripts/.jobs"); await mkdir(jobsDir, { recursive: true }); const at = Date.now() - 60 * 60 * 1000; await writeFile( join(jobsDir, `${id}.meta.json`), JSON.stringify({ id, kind: "refresh-report", queueKey: "", channelSlug: SLUG, status: "running", queuedAt: at, startedAt: at, pid, }), ); await writeFile( join(jobsDir, `${id}.log`), `Regenerating report for ${SLUG}…\n`, ); return id; } async function metaOf(id: string): Promise<{ status: string; cancelReason?: string }> { return JSON.parse( await readFile(resolvePath(`test-transcripts/.jobs/${id}.meta.json`), "utf8"), ); } test("a ghost running meta from a dead process is closed as interrupted, and does not block a move", async ({ page, }, testInfo) => { test.setTimeout(180_000); await resetData("one-youtube-channel-with-data"); await generateReport(page, SLUG); await quiet(page); // Past pid_max: no process has it. And this test runner's own pid: a live // process that is not the editor — an `archilyzer run` beside it. const ghost = await plantGhost(2 ** 22 + 1); const live = await plantGhost(process.pid); const res = await page.request.post(`${baseUrl}/api/test/settle-running-metas`); expect(res.ok()).toBe(true); const { interrupted } = (await res.json()) as { interrupted: { id: string }[] }; expect(interrupted.map((j) => j.id)).toContain(ghost); expect(interrupted.map((j) => j.id)).not.toContain(live); const closed = await metaOf(ghost); expect(closed.status).toBe("cancelled"); expect(closed.cancelReason).toMatch(/^interrupted: /); expect((await metaOf(live)).status).toBe("running"); // /jobs says why, on the job's own page. await page.goto(`/jobs/${ghost}`); await expect(page.getByTestId("cancel-reason")).toContainText("interrupted"); // Neither the closed ghost nor the live one is a writer on the channel: the // Storage panel offers the move and the move completes. const root = testInfo.outputPath("ghost-root"); await mkdir(root, { recursive: true }); await page.goto(channelStage(SLUG, "storage")); await page.getByLabel("destination root").fill(root); await page.getByRole("button", { name: "Preview", exact: true }).click(); await expect(page.getByLabel("relocation preview")).toBeVisible({ timeout: 15_000, }); const moveButton = page.getByRole("button", { name: "Move media" }); await expect(moveButton).toBeEnabled(); await moveButton.click(); await expect(page.getByLabel("Move media output")).toContainText("Moved", { timeout: 60_000, }); });