Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 8b39d63fb75f67c5ee216df409ea3b08faec7e0a
parent 348813a3afa0d7df4a1a2b75c8417d16939ed5e9
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Wed, 19 Aug 2026 23:41:04 -0400

auto-queue: the newest upload first, and a page that says what it's doing

Both runners always took pending[leaf][0], and that order came from
snapshot.json's bucket arrays — sorted alphabetically by video id, which is
arbitrary for YouTube ids and oldest-first for YYYYMMDD_ dir names. Nothing
made today's upload jump a 9,000-video backlog.

A per-runner `order` (listed | newest | oldest) sorts videos WITHIN a rule,
across every channel and bucket it claims; the tree still decides which rule
goes first. Claiming is untouched — buildPendingByLeaf takes an optional
comparator and sorts only the finished per-leaf list, so an absent comparator
reproduces today's order byte for byte.

The recency key is layered, and the layer that matters is not the obvious one.
transcripts/index.mdb's byChannel sub-DB is keyed [slug, uploadDate, id], so a
key-only scan dates 78,583 videos in ~300ms — but it is the TRANSCRIPT index,
holding only videos that already have a transcript, while auto-transcribe's
whole candidate set is downloadedNoTranscript. Measured on the real corpus it
dated 122 of the 870 genuinely pending. The gap is closed by an 8KB tail read
of metadata.info.json (868/870 at 0.19ms each, memoized — upload dates never
change). Undownloaded videos have no date anywhere on disk and are estimated
from playlist position, marked `estimated` and rendered with ≈. Anything still
undatable sorts oldest; a missing or locked index degrades to today's ordering
rather than failing a dispatch.

The page was two identical panels of counts and a pick log, and the payload it
already received held the answers it never showed. It is now a dispatcher
board: a deck for both runners, then per lane — what it would dispatch next and
why that rule, what is in flight, and a claim ladder that folds the tree, the
per-rule counts and the pick log into one numbered object, each rung's rail
carrying both depth and live occupancy, each count opening onto the ordered
pending ids. Plus the answer to the question operators actually arrive with:
every idleReason the loop already computed and threw away.

Also: snooze (idle, not stopped — persisted, so it survives a restart and
resumes itself), an unsaved-changes bar with Discard and a beforeunload guard
where edits used to vanish silently, and names for three controls a screen
reader could only announce as glyphs.

Structural notes for anyone editing the page: each kind stays a literal
<section> with its <h2>, nothing inside it may be a <section>, the deck names
runners in plain text not headings, and the policy controls stay native
select/checkbox — ~1,200 lines of e2e scope themselves on exactly that. Lanes
now stamp data-hydrated, because a state update landing before hydration is
discarded silently and that cost this change one red test.

43 unit tests, 26 e2e across auto-queue / auto-subs-replace /
new-channel-onboarding, tsc and build clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Diffstat:
Mcommon/controller/autoRunner.ts | 292+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++----
Acommon/controller/recencyIndex.test.ts | 243+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/controller/recencyIndex.ts | 487+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/jobs/autoQueuePolicy.test.ts | 174+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/jobs/autoQueuePolicy.ts | 55+++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/CHANGELOG.md | 7+++++++
Meditor/app/auto-queue/actions.ts | 37+++++++++++++++++++++++++++++++++++++
Meditor/app/auto-queue/components/AutoQueueView.tsx | 253+++++++++++++++++++++++++++++++------------------------------------------------
Aeditor/app/auto-queue/components/ClaimLadder.tsx | 48++++++++++++++++++++++++++++++++++++++++++++++++
Aeditor/app/auto-queue/components/DispatchDeck.tsx | 96+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aeditor/app/auto-queue/components/HowPriorityWorks.tsx | 51+++++++++++++++++++++++++++++++++++++++++++++++++++
Aeditor/app/auto-queue/components/InFlightList.tsx | 63+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aeditor/app/auto-queue/components/LadderRung.tsx | 501+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aeditor/app/auto-queue/components/LaneHeader.tsx | 141+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aeditor/app/auto-queue/components/NextUp.tsx | 102+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/app/auto-queue/components/PolicyTreeEditor.tsx | 557+++++++++++++++++++++++++------------------------------------------------------
Aeditor/app/auto-queue/components/SaveBar.tsx | 81+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aeditor/app/auto-queue/components/SnoozeControl.tsx | 92+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aeditor/app/auto-queue/components/dispatch.ts | 194+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/app/auto-queue/page.tsx | 20+++++++-------------
Meditor/app/auto-queue/status.ts | 24+++++++++++++++++++++---
Meditor/e2e/auto-queue.spec.ts | 277++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-
22 files changed, 3224 insertions(+), 571 deletions(-)

diff --git a/common/controller/autoRunner.ts b/common/controller/autoRunner.ts @@ -15,15 +15,26 @@ import { makeTaskTracker } from "../jobs/taskHooks"; import type { JobRunContext } from "../jobs/streamCommand"; import { type ActiveCounts, + type AutoQueueOrder, type AutoQueuePolicy, type ChannelWork, type WorkPick, buildPendingByLeaf, + flattenLeaves, selectNextWork, defaultBucketsForPolicy, selectableBucketsForKind, } from "../jobs/autoQueuePolicy"; import { + type RecencyKey, + buildRecencyKeys, + makeRecencyComparator, +} from "./recencyIndex"; + +// Re-exported so the editor's status payload can name the type without reaching +// past the runner into the index module it is an implementation detail of. +export type RecencyKeyView = RecencyKey; +import { type AutoQueueKind, readAutoQueueState, recordPick, @@ -95,11 +106,39 @@ export type AutoRunnerInFlight = { startedAt: number; }; +// Why the runner is up but dispatching nothing. Every one of these was already +// computed inside next() (or limit()) and thrown away, so the operator's most +// common question — "it's running, why isn't it doing anything?" — had no answer +// on the page. Recorded on the live record at each early return and cleared on a +// successful pick. +export type AutoRunnerIdleReason = + // The job record is gone or no longer running. + | "stopped" + // policy.enabled went false; the runner is shutting itself down. + | "disabled" + // policy.snoozeUntil is in the future — idle on purpose, not stopped. + | "snoozed" + // The global downloads pause (settings.downloadsPaused). Download only. + | "downloads-paused" + // diskGate refused to start more work. Download only. + | "disk-gate" + // Every pending video belongs to a platform inside a rate-limit/network + // cooldown window. Download only. + | "cooldown" + // No enabled, non-degraded worker slot exists. Transcription only. + | "no-workers" + // There is genuinely nothing to do. + | "no-pending" + // Work exists but every path to it is at a maxWorkers ceiling (a node cap, or + // the per-platform one-download-at-a-time gate). + | "capped"; + type RunnerLive = { jobId: string; inFlight: Map<string, AutoRunnerInFlight>; active: ActiveCounts; startedAt: number; + idleReason: AutoRunnerIdleReason | null; }; type AutoRunnerSingleton = { runners: Map<AutoQueueKind, RunnerLive> }; @@ -123,6 +162,9 @@ export type AutoRunnerStatus = { startedAt: number | null; inFlight: AutoRunnerInFlight[]; activeByNode: ActiveCounts; + // Why nothing is being dispatched right now, or null when the last scheduling + // decision was a grant. A stopped runner always reads "stopped". + idleReason: AutoRunnerIdleReason | null; }; export function getAutoRunnerStatus(kind: AutoQueueKind): AutoRunnerStatus { @@ -135,6 +177,7 @@ export function getAutoRunnerStatus(kind: AutoQueueKind): AutoRunnerStatus { startedAt: live?.startedAt ?? null, inFlight: live ? [...live.inFlight.values()] : [], activeByNode: live ? { ...live.active } : {}, + idleReason: running ? (live?.idleReason ?? null) : "stopped", }; } @@ -194,24 +237,191 @@ async function buildChannelWork( return { channels, owner }; } -// Per-leaf count of pending (snapshot-derived) work for the status panel — what -// the operator sees as "cornbreadman: 12 pending". Uses the same matching as the -// runner so the numbers line up with what would actually be picked. -export async function computeLeafPendingCounts( +// The policy's `order`, defaulted for settings files written before the field +// existed. One place, so the runner and the status page can never disagree about +// what an absent field means. +export function orderOf(policy: Pick<AutoQueuePolicy, "order">): AutoQueueOrder { + return policy.order ?? "listed"; +} + +// Build the comparator buildPendingByLeaf sorts each leaf with, plus the keys it +// was built from (the UI renders them next to the drill-down ids, which is how an +// ordering change gets verified by eye). Returns a null comparator for +// order:"listed" so the historical path is "don't sort", not "sort by identity". +// +// Never throws: a recency lookup failing is not a reason to stop dispatching +// work, so a broken index degrades to today's ordering. +async function recencyOrdering( + kind: AutoQueueKind, + paths: Paths, + policy: AutoQueuePolicy, + meta: ReadonlyArray<ChannelMeta>, + channels: ReadonlyArray<ChannelWork>, +): Promise<{ + compare: ((a: string, b: string) => number) | null; + keys: Map<string, RecencyKey>; +}> { + const order = orderOf(policy); + if (order === "listed") return { compare: null, keys: new Map() }; + // Candidates AND their owning channel, derived in one pass from the same + // projection buildPendingByLeaf claims from — first-writer-wins, exactly as + // buildChannelWork builds its own owner map, so the two cannot disagree. + const candidateIds = new Set<string>(); + const owner = new Map<string, string>(); + for (const ch of channels) { + for (const ids of Object.values(ch.buckets)) { + for (const id of ids) { + candidateIds.add(id); + if (!owner.has(id)) owner.set(id, ch.slug); + } + } + } + try { + const keys = await buildRecencyKeys({ + paths, + meta, + candidateIds, + owner, + // Only auto-download faces videos with NOTHING on disk to date them by. + // Every transcription candidate is already downloaded, so its + // metadata.info.json is there for the tail read and the playlist reads + // would be pure cost. + interpolate: kind === "download", + }); + return { compare: makeRecencyComparator(keys, order), keys }; + } catch { + return { compare: null, keys: new Map() }; + } +} + +// How many drill-down ids the status payload carries per rule. Enough to see +// whether an ordering change did what you asked; small enough that a 9,000-video +// catch-all doesn't ship 9,000 strings on a 3-second poll. +const PENDING_HEAD = 20; + +// The video the policy would hand out next, and enough context to say WHY that +// one. `skippedLeafIds` are the rules ahead of the chosen one in priority order +// that had nothing pending — which is the whole answer to "why is it working on +// rule 2?". +export type NextUpView = { + videoId: string; + channelSlug: string | null; + leafId: string; + path: string[]; + recency: RecencyKey | null; + skippedLeafIds: string[]; +}; + +export type LeafPending = { + counts: Record<string, number>; + // The first PENDING_HEAD ids of each leaf, in the exact order the runner would + // hand them out. + head: Record<string, string[]>; + // videoId -> owning channel slug, for linking a drill-down id to its page. + owner: Record<string, string>; + // videoId -> recency sort key, only for the ids in `head`, and only when the + // policy actually orders by recency. + recency: Record<string, RecencyKey>; + // What the policy would pick right now, or null when it would pick nothing. + nextUp: NextUpView | null; +}; + +// Per-leaf pending (snapshot-derived) work for the status panel — what the +// operator sees as "cornbreadman: 12 pending". Uses the same matching AND the +// same ordering as the runner, so the numbers and the drill-down line up with +// what would actually be picked. +export async function computeLeafPending( kind: AutoQueueKind, paths: Paths = getPaths(), -): Promise<Record<string, number>> { +): Promise<LeafPending> { const policy = getSettings().autoQueue[kind]; const meta = await listChannelMeta(paths); - const { channels } = await buildChannelWork(paths, kind, meta); + const { channels, owner } = await buildChannelWork(paths, kind, meta); + const { compare, keys } = await recencyOrdering( + kind, + paths, + policy, + meta, + channels, + ); const pending = buildPendingByLeaf( policy.root, channels, defaultBucketsForPolicy(kind, policy), + compare ? { compare } : undefined, ); + // A video already in flight is not "next up" — drop the live runner's set + // before asking the policy, exactly as next() does. + const live = getSingleton().runners.get(kind); + if (live && live.inFlight.size > 0) { + removeIds(pending, new Set(live.inFlight.keys())); + } + const counts: Record<string, number> = {}; - for (const [leafId, ids] of Object.entries(pending)) counts[leafId] = ids.length; - return counts; + const head: Record<string, string[]> = {}; + const ownerOut: Record<string, string> = {}; + const recency: Record<string, RecencyKey> = {}; + for (const [leafId, ids] of Object.entries(pending)) { + counts[leafId] = ids.length; + const slice = ids.slice(0, PENDING_HEAD); + head[leafId] = slice; + for (const id of slice) { + const slug = owner.get(id); + if (slug) ownerOut[id] = slug; + const key = keys.get(id); + if (key) recency[id] = key; + } + } + + // THE TRAP: selectNextWork advances SWRR fairness by mutating + // runtime.currentWeights in place. This runs on a 3-second status poll, so + // asking the live runtime would let merely HAVING the page open skew a + // round-robin group's rotation. Deep-clone first; the clone is discarded. + const state = await readAutoQueueState(paths); + const runtime = { + currentWeights: { ...state[kind].runtime.currentWeights }, + }; + const pick = selectNextWork( + policy.root, + pending, + runtime, + live ? { ...live.active } : {}, + ); + let nextUp: NextUpView | null = null; + if (pick) { + const order = flattenLeaves(policy.root); + const at = order.findIndex((l) => l.id === pick.leafId); + nextUp = { + videoId: pick.videoId, + channelSlug: owner.get(pick.videoId) ?? null, + leafId: pick.leafId, + path: pick.path, + recency: keys.get(pick.videoId) ?? null, + skippedLeafIds: + at <= 0 + ? [] + : order + .slice(0, at) + .filter((l) => (counts[l.id] ?? 0) === 0) + .map((l) => l.id), + }; + } + + return { counts, head, owner: ownerOut, recency, nextUp }; +} + +// Back-compat shim for callers that only ever wanted the counts. +export async function computeLeafPendingCounts( + kind: AutoQueueKind, + paths: Paths = getPaths(), +): Promise<Record<string, number>> { + return (await computeLeafPending(kind, paths)).counts; +} + +function countPending(pending: Record<string, string[]>): number { + let total = 0; + for (const ids of Object.values(pending)) total += ids.length; + return total; } function removeIds( @@ -330,9 +540,15 @@ async function runLoop( // itself while draining, so this never needs to check the drain signal. const limit = (): number => { const policy = getSettings().autoQueue[kind]; - return kind === "transcription" - ? Math.min(eligibleSlots(), policy.maxWorkers ?? Number.POSITIVE_INFINITY) - : (policy.maxWorkers ?? Number.POSITIVE_INFINITY); + if (kind !== "transcription") { + return policy.maxWorkers ?? Number.POSITIVE_INFINITY; + } + const slots = eligibleSlots(); + // A zero limit means runPool never calls next(), so this is the only place + // "every worker is off or degraded" can be observed — without it the page + // would show a running runner with no explanation at all. + if (slots === 0 && live.inFlight.size === 0) live.idleReason = "no-workers"; + return Math.min(slots, policy.maxWorkers ?? Number.POSITIVE_INFINITY); }; // Select the next unit and RESERVE its slot (so the next selection sees it), @@ -343,20 +559,38 @@ async function runLoop( // (a hard cancel flips it to "cancelled" and also fires `signal`). const rec = getRegistry().get(live.jobId); if (!rec || rec.status !== "running") { + live.idleReason = "stopped"; stopController.abort(); return null; } const settings = getSettings(); const policy: AutoQueuePolicy = settings.autoQueue[kind]; if (!policy.enabled) { + live.idleReason = "disabled"; onLog(`Auto-${kind} disabled — stopping runner.`); stopController.abort(); return null; } + // Snooze: idle, do NOT stop. Same shape as the downloads-pause branch below + // — settings are re-read every iteration, so the runner wakes by itself the + // moment the deadline passes, and because the deadline lives in settings.json + // it survives a server restart. getSettings() already normalizes a lapsed + // snooze to null; the clock check is belt-and-braces for a clock that moved. + const snoozeUntil = policy.snoozeUntil ?? null; + if (snoozeUntil !== null && snoozeUntil > Date.now()) { + if (live.idleReason !== "snoozed") { + onLog( + `Auto-${kind} snoozed until ${new Date(snoozeUntil).toLocaleString()}.`, + ); + } + live.idleReason = "snoozed"; + return null; + } // Global downloads pause: like the `enabled` flag, this is re-read each // iteration. Rather than stop the runner, idle it (return null) so it // resumes dispatching on the next tick once unpaused — no restart needed. if (kind === "download" && settings.downloadsPaused) { + live.idleReason = "downloads-paused"; return null; } // THE UNATTENDED DISK GATE. This runner is the one path that dispatches @@ -375,6 +609,7 @@ async function runLoop( diskIdle = true; onLog(`Auto-download idle: ${gate.message}.`); } + live.idleReason = "disk-gate"; return null; } if (diskIdle) { @@ -397,17 +632,29 @@ async function runLoop( ); const { channels, owner } = await buildChannelWork(paths, kind, metaCache); // Re-derived each iteration from the freshly-read policy, like `enabled`, so - // toggling the replace-auto-captions lane takes effect without a restart. + // toggling the replace-auto-captions lane (or the ordering) takes effect + // without a restart. The recency index behind the comparator has its own + // 30s TTL, so this is a map lookup per candidate on all but one tick in ten. + const { compare } = await recencyOrdering( + kind, + paths, + policy, + metaCache, + channels, + ); const pending = buildPendingByLeaf( policy.root, channels, defaultBucketsForPolicy(kind, policy), + compare ? { compare } : undefined, ); const exclude = new Set<string>([...live.inFlight.keys(), ...completed]); removeIds(pending, exclude); + const pendingBeforeGates = countPending(pending); // Download only: drop videos whose platform already has an in-flight // download (busy) OR is in a rate-limit/network backoff window (cooling // down), so a busy/throttled platform yields to the next-priority free one. + let anyCooling = false; if (kind === "download") { // Merge in any cooldown a manual sync/import wrote to the shared state // (read-modify-write from outside the runner) since our last persist. @@ -428,7 +675,13 @@ async function runLoop( if (n >= PER_PLATFORM_CAP) skip.add(pf); } for (const pf of Object.keys(kindState.platformBackoff)) { - if (isCoolingDown(kindState.platformBackoff, pf, now)) skip.add(pf); + if (isCoolingDown(kindState.platformBackoff, pf, now)) { + skip.add(pf); + // A platform skipped for a cooldown is a different answer to "why is + // it idle?" than one skipped for being busy, so the two are tracked + // apart rather than both reading as "capped". + anyCooling = true; + } } if (skip.size > 0) { for (const leafId of Object.keys(pending)) { @@ -441,7 +694,16 @@ async function runLoop( } const pick = selectNextWork(policy.root, pending, runtime, live.active); - if (!pick) return null; + if (!pick) { + // Attribute the idleness. Nothing pending at all is a different situation + // from work that exists but is unreachable, and both differ from work the + // platform gates just took away — the page says which. + if (pendingBeforeGates === 0) live.idleReason = "no-pending"; + else if (countPending(pending) === 0) { + live.idleReason = anyCooling ? "cooldown" : "capped"; + } else live.idleReason = "capped"; + return null; + } const channelSlug = owner.get(pick.videoId); if (!channelSlug) { @@ -449,6 +711,7 @@ async function runLoop( markCompleted(pick.videoId); return null; } + live.idleReason = null; const unitPlatform = kind === "download" ? platformKey(channelSlug, slugToPlatform) : null; @@ -802,6 +1065,7 @@ export async function startAutoRunner( inFlight: new Map(), active: {}, startedAt: Date.now(), + idleReason: null, }; const result = await runManagedFunction({ diff --git a/common/controller/recencyIndex.test.ts b/common/controller/recencyIndex.test.ts @@ -0,0 +1,243 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { mkdir, mkdtemp, rm, writeFile } from "node:fs/promises"; +import { tmpdir } from "node:os"; +import path from "node:path"; +import { + type RecencyKey, + buildRecencyKeys, + clearRecencyCache, + interpolateFromPlaylist, + makeRecencyComparator, +} from "./recencyIndex"; +import type { Paths } from "../lib/paths"; + +// Run with: node_modules/.bin/tsx --test common/controller/recencyIndex.test.ts + +function run( + playlist: string[], + dates: Record<string, string>, + wanted: string[], +): Record<string, RecencyKey> { + const out = new Map<string, RecencyKey>(); + interpolateFromPlaylist( + playlist, + new Map(Object.entries(dates)), + new Set(wanted), + out, + ); + return Object.fromEntries(out); +} + +// --- Layer 2: playlist-neighbour interpolation ------------------------------- + +test("interpolation: an undated id takes the nearest PRECEDING date", () => { + // The playlist is newest-first, so the entry above an undownloaded video is + // an upper bound on its age — the best estimate available, since an + // undownloaded video has no metadata.info.json to read a real date from. + const keys = run( + ["a", "gap1", "gap2", "b", "gap3"], + { a: "20260801", b: "20260101" }, + ["gap1", "gap2", "gap3"], + ); + assert.deepEqual(keys.gap1, { key: "20260801", estimated: true }); + assert.deepEqual(keys.gap2, { key: "20260801", estimated: true }); + assert.deepEqual(keys.gap3, { key: "20260101", estimated: true }); +}); + +test("interpolation: an id ahead of every date sorts above everything", () => { + // The point of the whole feature: a video uploaded today, not yet downloaded, + // sits at the head of a newest-first playlist with nothing dated above it. It + // must jump the backlog rather than land next to the channel's oldest work. + const keys = run(["brand-new", "a"], { a: "20260801" }, ["brand-new"]); + assert.equal(keys["brand-new"].estimated, true); + const cmp = makeRecencyComparator( + new Map(Object.entries(keys) as [string, RecencyKey][]).set("old", { + key: "20261231", + estimated: false, + }), + "newest", + )!; + assert.deepEqual(["old", "brand-new"].sort(cmp), ["brand-new", "old"]); +}); + +test("interpolation: an anchorless playlist is left alone", () => { + // With no dated entry anywhere we know nothing about the channel's timeline. + // Handing every id the top sentinel would let one freshly-added channel with + // 9,000 undownloaded videos monopolize the head of the queue, so this bails + // and lets the dir-prefix fallback (layer 3) decide instead. + assert.deepEqual(run(["x", "y"], {}, ["x", "y"]), {}); +}); + +test("interpolation: only ids the caller asked for are keyed, once each", () => { + const wanted = new Set(["gap1"]); + const out = new Map<string, RecencyKey>(); + interpolateFromPlaylist( + ["a", "gap1", "gap2"], + new Map([["a", "20260801"]]), + wanted, + out, + ); + assert.deepEqual([...out.keys()], ["gap1"]); + // The wanted set is drained as it is satisfied, so the caller can stop early. + assert.equal(wanted.size, 0); +}); + +// --- Layer 2: the metadata.info.json tail read ------------------------------- + +// Write a metadata.info.json whose upload_date sits near the END of the file, +// behind `padBytes` of other JSON — which is how yt-dlp actually writes them, +// and the reason a tail read works at all. +async function writeMeta( + channelsDir: string, + slug: string, + id: string, + uploadDate: string | null, + padBytes = 0, +): Promise<void> { + const dir = path.join(channelsDir, slug, "data", id); + await mkdir(dir, { recursive: true }); + const body: Record<string, unknown> = { id, description: "x".repeat(padBytes) }; + if (uploadDate) body.upload_date = uploadDate; + await writeFile(path.join(dir, "metadata.info.json"), JSON.stringify(body)); +} + +test("tail read dates a downloaded-but-untranscribed video", async () => { + // The case the LMDB layer structurally cannot serve: auto-transcribe's whole + // candidate set is videos with no transcript, so none of them are in the + // transcript index. Measured on the live corpus, layer 1 covered 122 of 870 + // and this layer covered 868. + const dir = await mkdtemp(path.join(tmpdir(), "recency-tail-")); + try { + clearRecencyCache(); + const channelsDir = path.join(dir, "channels"); + await writeMeta(channelsDir, "ch", "vid-new", "20260812", 400); + await writeMeta(channelsDir, "ch", "vid-old", "20240101", 400); + // No upload_date at all, and one with no metadata file whatsoever. + await writeMeta(channelsDir, "ch", "vid-undated", null, 100); + const paths = { + lmdbPath: path.join(dir, "none.mdb"), + channelsDir, + } as Paths; + const keys = await buildRecencyKeys({ + paths, + meta: [{ slug: "ch" }], + candidateIds: new Set(["vid-new", "vid-old", "vid-undated", "vid-absent"]), + owner: new Map([ + ["vid-new", "ch"], + ["vid-old", "ch"], + ["vid-undated", "ch"], + ["vid-absent", "ch"], + ]), + fresh: true, + }); + assert.deepEqual(keys.get("vid-new"), { key: "20260812", estimated: false }); + assert.deepEqual(keys.get("vid-old"), { key: "20240101", estimated: false }); + // Both misses fall through to layer 4 rather than vanishing from the queue. + assert.deepEqual(keys.get("vid-undated"), { key: "", estimated: false }); + assert.deepEqual(keys.get("vid-absent"), { key: "", estimated: false }); + assert.deepEqual( + ["vid-old", "vid-new"].sort(makeRecencyComparator(keys, "newest")!), + ["vid-new", "vid-old"], + ); + } finally { + await rm(dir, { recursive: true, force: true }); + } +}); + +test("tail read is skipped for ids with no known owner", async () => { + // Without an owner there is no channel dir to look in; those ids must fall + // through rather than probe every channel. + const dir = await mkdtemp(path.join(tmpdir(), "recency-noowner-")); + try { + clearRecencyCache(); + const channelsDir = path.join(dir, "channels"); + await writeMeta(channelsDir, "ch", "vid", "20260812"); + const keys = await buildRecencyKeys({ + paths: { lmdbPath: path.join(dir, "none.mdb"), channelsDir } as Paths, + meta: [{ slug: "ch" }], + candidateIds: new Set(["vid"]), + fresh: true, + }); + assert.deepEqual(keys.get("vid"), { key: "", estimated: false }); + } finally { + await rm(dir, { recursive: true, force: true }); + } +}); + +// --- Layer 4 + end-to-end degradation --------------------------------------- + +test("buildRecencyKeys: no index and no metadata still keys every candidate", async () => { + // A missing/locked index must never fail a dispatch — it just means fewer + // known dates. The YYYYMMDD_ dir-name prefix is the last real signal. + const dir = await mkdtemp(path.join(tmpdir(), "recency-")); + try { + clearRecencyCache(); + const paths = { + lmdbPath: path.join(dir, "does-not-exist.mdb"), + channelsDir: path.join(dir, "channels"), + } as Paths; + const keys = await buildRecencyKeys({ + paths, + meta: [{ slug: "ch" }], + candidateIds: new Set(["20260812_talk", "20240101_old", "dQw4w9WgXcQ"]), + fresh: true, + }); + assert.deepEqual(keys.get("20260812_talk"), { + key: "20260812", + estimated: false, + }); + assert.deepEqual(keys.get("20240101_old"), { + key: "20240101", + estimated: false, + }); + // An undatable id sorts oldest rather than being dropped from the queue. + assert.deepEqual(keys.get("dQw4w9WgXcQ"), { key: "", estimated: false }); + } finally { + await rm(dir, { recursive: true, force: true }); + } +}); + +test("buildRecencyKeys: an empty candidate set does no work", async () => { + const keys = await buildRecencyKeys({ + paths: { lmdbPath: "/nope", channelsDir: "/nope" } as Paths, + meta: [{ slug: "ch" }], + candidateIds: new Set(), + fresh: true, + }); + assert.equal(keys.size, 0); +}); + +// --- The comparator ---------------------------------------------------------- + +test("makeRecencyComparator: newest is descending, oldest ascending", () => { + const keys = new Map<string, RecencyKey>([ + ["old", { key: "20240101", estimated: false }], + ["mid", { key: "20250601", estimated: false }], + ["new", { key: "20260812", estimated: false }], + ]); + assert.deepEqual( + ["old", "new", "mid"].sort(makeRecencyComparator(keys, "newest")!), + ["new", "mid", "old"], + ); + assert.deepEqual( + ["new", "old", "mid"].sort(makeRecencyComparator(keys, "oldest")!), + ["old", "mid", "new"], + ); +}); + +test("makeRecencyComparator: listed returns null, so nothing is sorted", () => { + // Not an identity comparator — null, so buildPendingByLeaf skips the sort + // entirely and today's order is reproduced by not touching it. + assert.equal(makeRecencyComparator(new Map(), "listed"), null); +}); + +test("makeRecencyComparator: an unknown id sorts oldest, never crashes", () => { + const keys = new Map<string, RecencyKey>([ + ["known", { key: "20260101", estimated: false }], + ]); + assert.deepEqual( + ["ghost", "known"].sort(makeRecencyComparator(keys, "newest")!), + ["known", "ghost"], + ); +}); diff --git a/common/controller/recencyIndex.ts b/common/controller/recencyIndex.ts @@ -0,0 +1,487 @@ +import { existsSync } from "node:fs"; +import { open as openFile, readFile } from "node:fs/promises"; +import path from "node:path"; +import { open } from "lmdb"; +import type { Paths } from "../lib/paths"; +import { mapConcurrent } from "../lib/concurrency"; +import { extractVideoId } from "../lib/videoId"; +import { uploadKeyFor } from "./keptVideos"; +import type { AutoQueueOrder } from "../jobs/autoQueuePolicy"; + +// Upload-date lookup for the auto-queue's "newest first" ordering. +// +// The problem: the runners pick pending[leaf][0], and that order comes from +// snapshot.json's bucket arrays, which channelSnapshot sorts lexicographically +// by video id. For YouTube ids that is arbitrary; for YYYYMMDD_-prefixed dir +// names it is oldest-first. Neither lets a today's-upload jump a 9,000-video +// backlog. To sort by recency we need an upload date per candidate video, for +// tens of thousands of candidates, on every scheduling tick — so it has to be +// nearly free. +// +// It nearly is, because the transcript index already stores it. buildIndex's +// `byChannel` sub-DB is keyed [channelSlug, uploadDate, id] with a constant +// value, so a KEY-ONLY range scan reads every video's date without decoding a +// single value. What that scan cannot supply is covered by a bounded tail read. +// +// Four layers, in order: +// 1. the LMDB `byChannel` scan — ~76 ms warm for all 78,583 indexed videos, +// 99% of the corpus. BUT it is the TRANSCRIPT index, so it holds only +// videos that already have a transcript. Measured against the live corpus, +// that covers just 122 of the 870 videos auto-transcribe actually has +// pending — the bucket it orders is `downloadedNoTranscript`, which is by +// definition the set this layer cannot see. It is still layer 1 because it +// dates the whole corpus for one cheap scan, and it supplies the anchors +// layer 3 interpolates between; +// 2. an 8 KB TAIL read of each remaining video's metadata.info.json, regexed +// for upload_date. Measured on the real backlog: 868/870 hits (99.8%) at +// 0.19 ms/file, 164 ms for the lot. This is the layer that actually orders +// auto-transcribe. Memoized process-wide — an upload date never changes — +// so the cost is paid once per video, not once per refresh; +// 3. playlist-neighbour interpolation, for auto-download — an UNdownloaded +// video has no metadata.info.json at all, so it has no date anywhere on +// disk. The stored playlist is newest-first, so the nearest PRECEDING +// dated entry is an upper bound on its age; +// 4. uploadKeyFor's YYYYMMDD_ dir-name prefix, then "" (sorts oldest). +// +// Reading each metadata.info.json in FULL was measured and rejected: ~6.5 min +// and ~41 GiB of I/O. So was roster.json's firstSeenAt (99.7% of entries +// collapse onto one seed timestamp) and the served stats pages (per-site +// filtered, capped at 30,990 rows). +// +// Everything here degrades rather than throws: a missing, locked or corrupt +// index just means fewer known dates, never a runner that fails to dispatch. + +// A video's recency sort key. `key` is a YYYYMMDD string (or one of the two +// sentinels below); `estimated` marks a key that was interpolated rather than +// read, so the UI can render it as "≈2026-08-12" instead of claiming precision. +export type RecencyKey = { key: string; estimated: boolean }; + +// Sorts above every real date. Given to an undownloaded video that sits ahead of +// every dated entry in its channel's newest-first playlist — i.e. it was +// uploaded after everything we have. That is precisely the video this feature +// exists to promote, so it goes to the very top of the archive, not merely to +// the top of its own channel. Guarded by requiring at least one dated anchor in +// the playlist (see interpolateFromPlaylist), so a brand-new channel with +// nothing downloaded cannot flood the head of the queue. +const NEWER_THAN_ANYTHING = "￿"; + +// Sorts below every real date — what uploadKeyFor already returns for an +// undatable id. Named only for readability at the call sites. +const UNKNOWN = ""; + +// --- Layer 1: the LMDB byChannel key scan ----------------------------------- + +// [channelSlug, uploadDate, id] — buildIndex.ts's ChannelKey. +type ChannelKey = [string, string, string]; + +export type UploadDateIndex = { + // slug -> (videoId -> YYYYMMDD). Empty when the index is unavailable. + datesBySlug: Map<string, Map<string, string>>; + close(): Promise<void>; +}; + +const EMPTY_INDEX: UploadDateIndex = { + datesBySlug: new Map(), + close: async () => {}, +}; + +// Read-only view of the index's upload dates, scanned per channel. Modeled on +// openChannelSigner (common/lib/channelSignature.ts): guard on existsSync, wrap +// the open in try/catch, and degrade to an empty stub rather than fail — the +// runner must keep dispatching even when the index is mid-rebuild or absent. +export function openUploadDateIndex( + paths: Paths, + slugs: ReadonlyArray<string>, +): UploadDateIndex { + if (!existsSync(paths.lmdbPath)) return EMPTY_INDEX; + let root: ReturnType<typeof open>; + try { + root = open({ path: paths.lmdbPath, readOnly: true, maxDbs: 14 }); + } catch { + return EMPTY_INDEX; + } + try { + const byChannel = root.openDB<number, ChannelKey>({ + name: "byChannel", + encoding: "msgpack", + }); + const datesBySlug = new Map<string, Map<string, string>>(); + for (const slug of slugs) { + const dates = new Map<string, string>(); + // Key-only iteration: getRange yields {key, value} but we never touch + // .value, and lmdb-js decodes values lazily, so no msgpack decode happens. + for (const { key } of byChannel.getRange({ + start: [slug], + end: [slug, "￿"], + })) { + const k = key as ChannelKey; + if (k[0] !== slug) break; + // First writer wins, matching the archive's own dedup: an id can only + // appear once per channel, but a mid-rebuild index may briefly hold a + // stale row alongside a fresh one. + if (!dates.has(k[2])) dates.set(k[2], k[1]); + } + datesBySlug.set(slug, dates); + } + return { + datesBySlug, + close: async () => { + await root.close(); + }, + }; + } catch { + void root.close().catch(() => {}); + return EMPTY_INDEX; + } +} + +// --- Layer 2: the metadata.info.json tail read ------------------------------- + +// yt-dlp writes upload_date near the END of metadata.info.json, so an 8 KB tail +// finds it without paying for a multi-hundred-KB file (some carry every +// subtitle track and comment). Measured 868/870 on the live corpus. +const TAIL_BYTES = 8192; +const UPLOAD_DATE_RE = /"upload_date":\s*"(\d{8})"/; + +// Upload dates are immutable, so a date read once is a date forever. Memoizing +// process-wide turns this layer from a per-refresh cost into a one-off: after +// the first fill the whole layer is map lookups. `null` memoizes a genuine miss +// (no metadata, or upload_date outside the tail) so it is not retried every +// refresh either. +const tailMemo = new Map<string, string | null>(); +// Backstop against unbounded growth on a corpus far larger than this one. +const TAIL_MEMO_CAP = 200_000; + +// How many tail reads one build may perform. At 0.19 ms each this is ~1.5 s +// worst case, and only on the FIRST refresh after a huge backlog appears — +// subsequent refreshes hit the memo. Anything past the cap falls through to +// layer 4 for now and is picked up by a later refresh, which is the right +// failure mode: an unrankable video sorts oldest, i.e. to the back of a +// newest-first queue, rather than blocking a scheduling tick. +const TAIL_READS_PER_BUILD = 8000; +const TAIL_READ_CONCURRENCY = 32; + +async function readTailUploadDate(file: string): Promise<string | null> { + let fh; + try { + fh = await openFile(file, "r"); + const { size } = await fh.stat(); + if (size === 0) return null; + const len = Math.min(TAIL_BYTES, size); + const buf = Buffer.allocUnsafe(len); + await fh.read(buf, 0, len, size - len); + // latin1 never throws on a multi-byte sequence split by the tail boundary, + // and the pattern we want is pure ASCII. + return buf.toString("latin1").match(UPLOAD_DATE_RE)?.[1] ?? null; + } catch { + return null; + } finally { + await fh?.close().catch(() => {}); + } +} + +// Date the ids in `wanted` from their on-disk metadata. Removes each id it +// keys from `wanted`, like interpolateFromPlaylist. +async function datesFromMetadata( + paths: Paths, + owner: ReadonlyMap<string, string>, + wanted: Set<string>, + out: Map<string, RecencyKey>, +): Promise<void> { + const todo: string[] = []; + for (const id of wanted) { + const memo = tailMemo.get(id); + if (memo !== undefined) { + if (memo) { + out.set(id, { key: memo, estimated: false }); + wanted.delete(id); + } + continue; + } + if (!owner.has(id)) continue; + todo.push(id); + if (todo.length >= TAIL_READS_PER_BUILD) break; + } + if (todo.length === 0) return; + const dates = await mapConcurrent(todo, TAIL_READ_CONCURRENCY, (id) => + readTailUploadDate( + path.join( + paths.channelsDir, + owner.get(id) as string, + "data", + id, + "metadata.info.json", + ), + ), + ); + for (const [i, id] of todo.entries()) { + const date = dates[i]; + if (tailMemo.size < TAIL_MEMO_CAP) tailMemo.set(id, date); + if (!date) continue; + out.set(id, { key: date, estimated: false }); + wanted.delete(id); + } +} + +// --- Layer 3: playlist-neighbour interpolation ------------------------------ + +async function readPlaylistIds( + paths: Paths, + slug: string, +): Promise<string[]> { + try { + const raw = await readFile( + path.join(paths.channelsDir, slug, "playlist"), + "utf8", + ); + const ids: string[] = []; + for (const line of raw.split("\n")) { + const trimmed = line.trim(); + if (!trimmed) continue; + const id = extractVideoId(trimmed); + if (id) ids.push(id); + } + return ids; + } catch { + return []; + } +} + +// Estimate a date for each id in `wanted` from its position in the channel's +// newest-first playlist: take the nearest PRECEDING entry whose date we know. +// An id ahead of every known date is newer than everything we have and gets +// NEWER_THAN_ANYTHING. +// +// Requires at least one dated anchor: with no anchor we know nothing about the +// channel's timeline, and handing every id the top sentinel would let a freshly +// added 9,000-video channel monopolize the head of the queue. Anchorless +// channels fall through to layer 4 instead. +// +// Exported for the unit test: its two inputs (a newest-first playlist and a +// partial date map) are exactly what a fixture can supply, whereas driving it +// through buildRecencyKeys would require standing up an LMDB. +export function interpolateFromPlaylist( + playlistIds: ReadonlyArray<string>, + dates: ReadonlyMap<string, string>, + // Mutated: an id keyed here is removed, so the caller can stop early once + // every missing id has been estimated. + wanted: Set<string>, + out: Map<string, RecencyKey>, +): void { + let anchored = false; + for (const id of playlistIds) { + if (dates.has(id)) { + anchored = true; + break; + } + } + if (!anchored) return; + let last: string | null = null; + for (const id of playlistIds) { + const known = dates.get(id); + if (known) { + last = known; + continue; + } + if (!wanted.delete(id)) continue; + out.set(id, { + key: last ?? NEWER_THAN_ANYTHING, + estimated: true, + }); + } +} + +// --- The public builder ------------------------------------------------------ + +export type BuildRecencyKeysArgs = { + paths: Paths; + // The channels whose work is in play — the runner's own channel-meta list. + meta: ReadonlyArray<{ slug: string }>; + // Every video id the caller might sort. Ids outside this set are not keyed. + candidateIds: ReadonlySet<string>; + // videoId -> owning channel slug. Required for the tail-read layer, which has + // to know which channel dir a video lives in. Ids absent from this map skip + // layer 2. The runner derives it from the same projection it builds the + // candidate set from, so the two can never disagree. + owner?: ReadonlyMap<string, string>; + // Read playlists for layer 2. Auto-transcribe only ever considers already- + // downloaded videos, which the index already dates, so it passes false and + // skips the reads entirely. Default true. + interpolate?: boolean; + // Bypass the TTL cache (tests, and the offline sanity script). + fresh?: boolean; +}; + +// What the expensive half of a build produces: everything derived from disk, +// independent of which ids the caller happens to be asking about. Cached, so a +// runner ticking every few seconds pays for the index scan and the playlist +// reads at most once per TTL. +type RecencySources = { + at: number; + // Every indexed video's date, merged across channels (an id belongs to one + // channel, so the merge is lossless in practice; first writer wins). + dates: Map<string, string>; + // Per channel, kept separately because interpolation needs a channel's own + // anchors, not the corpus's. + datesBySlug: Map<string, Map<string, string>>; + // slug -> playlist ids in listing (newest-first) order. Only populated when + // the caller asked to interpolate. + playlists: Map<string, string[]>; + interpolated: boolean; +}; + +// Matches the runner's own CHANNEL_LIST_TTL_MS: the two caches expire together, +// so a channel added mid-run becomes visible to the meta list and to its dates +// on the same tick rather than one lagging the other. +const RECENCY_TTL_MS = 30_000; + +let cache: RecencySources | null = null; +let cacheKey = ""; + +// Exported for tests, which need a clean slate between fixtures. +export function clearRecencyCache(): void { + cache = null; + cacheKey = ""; + tailMemo.clear(); +} + +async function loadSources( + paths: Paths, + slugs: ReadonlyArray<string>, + interpolate: boolean, + fresh: boolean, +): Promise<RecencySources> { + const key = `${paths.lmdbPath}\u0000${slugs.join(",")}`; + const now = Date.now(); + if ( + !fresh && + cache && + cacheKey === key && + now - cache.at < RECENCY_TTL_MS && + // A cached scan that skipped playlists cannot serve a request that needs + // them; the reverse is fine. + (!interpolate || cache.interpolated) + ) { + return cache; + } + + const index = openUploadDateIndex(paths, slugs); + const dates = new Map<string, string>(); + const datesBySlug = index.datesBySlug; + try { + for (const perChannel of datesBySlug.values()) { + for (const [id, date] of perChannel) if (!dates.has(id)) dates.set(id, date); + } + } finally { + await index.close().catch(() => {}); + } + + const playlists = new Map<string, string[]>(); + if (interpolate) { + for (const slug of slugs) { + // Require at least one indexed video. Layer 2 can add anchors later, but a + // channel with no transcript at all is one whose whole playlist would fall + // to layer 4 anyway — so skip the read rather than parse it and discard it. + if (!datesBySlug.get(slug)?.size) continue; + const ids = await readPlaylistIds(paths, slug); + if (ids.length > 0) playlists.set(slug, ids); + } + } + + const sources: RecencySources = { + at: now, + dates, + datesBySlug, + playlists, + interpolated: interpolate, + }; + if (!fresh) { + cache = sources; + cacheKey = key; + } + return sources; +} + +export async function buildRecencyKeys({ + paths, + meta, + candidateIds, + owner, + interpolate = true, + fresh = false, +}: BuildRecencyKeysArgs): Promise<Map<string, RecencyKey>> { + const out = new Map<string, RecencyKey>(); + if (candidateIds.size === 0) return out; + const slugs = meta.map((m) => m.slug); + const sources = await loadSources(paths, slugs, interpolate, fresh); + + // Layer 1: the index scan. + const missing = new Set<string>(); + for (const id of candidateIds) { + const date = sources.dates.get(id); + if (date) out.set(id, { key: date, estimated: false }); + else missing.add(id); + } + + // Layer 2: on-disk metadata. This is the one that actually orders + // auto-transcribe, whose entire candidate set is by definition absent from the + // transcript index. + if (owner && missing.size > 0) { + await datesFromMetadata(paths, owner, missing, out); + } + + // Layer 3: playlist interpolation, for videos with nothing on disk at all. + // Anchors are the index dates PLUS anything layer 2 just read — a downloaded + // -but-untranscribed video is invisible to the index yet makes a perfectly + // good anchor, and folding it in tightens every estimate around it. + if (interpolate && missing.size > 0) { + for (const [slug, playlistIds] of sources.playlists) { + if (missing.size === 0) break; + const anchors = new Map(sources.datesBySlug.get(slug) ?? []); + if (owner) { + for (const [id, key] of out) { + if (!key.estimated && key.key && owner.get(id) === slug) { + anchors.set(id, key.key); + } + } + } + interpolateFromPlaylist(playlistIds, anchors, missing, out); + } + } + + // Layer 4: the YYYYMMDD_ dir-name prefix, else "" (oldest). Reuses the same + // helper the on-disk keep-window uses, so an id keys identically here and + // there rather than growing a second set of fallback rules. + for (const id of candidateIds) { + if (out.has(id)) continue; + out.set(id, { key: uploadKeyFor(undefined, id), estimated: false }); + } + return out; +} + +// --- Comparator --------------------------------------------------------------- + +// Total order over video ids for a given policy `order`. Returns null for +// "listed", which the caller passes straight through to buildPendingByLeaf as an +// absent comparator — so the historical order is reproduced by not sorting at +// all, not by sorting with an identity comparator. +// +// Equal keys compare 0 deliberately: Array#sort is stable, so videos sharing an +// upload date keep their "listed" order (playlist order for undownloadedIds, +// id order elsewhere) instead of being shuffled by an arbitrary tiebreak. +export function makeRecencyComparator( + keys: ReadonlyMap<string, RecencyKey>, + order: AutoQueueOrder, +): ((a: string, b: string) => number) | null { + if (order === "listed") return null; + // Keys are YYYYMMDD strings, so lexicographic order IS chronological order. + // "newest" therefore sorts DESCENDING: a smaller (older) key must come later, + // which is a positive comparator result. + const olderFirst = order === "newest" ? 1 : -1; + return (a, b) => { + const ka = keys.get(a)?.key ?? UNKNOWN; + const kb = keys.get(b)?.key ?? UNKNOWN; + if (ka === kb) return 0; + return ka < kb ? olderFirst : -olderFirst; + }; +} diff --git a/common/jobs/autoQueuePolicy.test.ts b/common/jobs/autoQueuePolicy.test.ts @@ -537,3 +537,177 @@ test("sanitizeAutoQueue defaults replaceAutoSubs to false", () => { true, ); }); + +// --- Ordering within a rule (policy.order) ---------------------------------- + +// The comparator the runner supplies is built from real upload dates; here a +// literal date map stands in, so these tests exercise the ENGINE's contract — +// sort each leaf's finished list, change nothing else — without any disk. +function byDate( + dates: Record<string, string>, + dir: "newest" | "oldest", +): (a: string, b: string) => number { + const olderFirst = dir === "newest" ? 1 : -1; + return (a, b) => { + const ka = dates[a] ?? ""; + const kb = dates[b] ?? ""; + if (ka === kb) return 0; + return ka < kb ? olderFirst : -olderFirst; + }; +} + +const ORDER_CHANNELS: ChannelWork[] = [ + { + slug: "cornbreadman", + platform: "odysee", + buckets: { downloadedNoTranscript: ["c_old", "c_new"], failedListed: [] }, + }, + { + slug: "destiny", + platform: "youtube", + buckets: { downloadedNoTranscript: ["d_mid"], failedListed: [] }, + }, +]; +const ORDER_DATES = { + c_old: "20240101", + c_new: "20260812", + d_mid: "20250601", +}; +const ORDER_ROOT: AutoQueueGroup = { + id: "root", + mode: "strict", + children: [{ id: "all", match: { type: "all" } }], +}; + +test("buildPendingByLeaf: no comparator reproduces today's order exactly", () => { + // The whole compatibility claim for this feature in one assertion: an absent + // `compare` must not merely produce a "similar" list, it must produce the + // identical one, so every existing settings.json keeps its behaviour. + const before = buildPendingByLeaf(ORDER_ROOT, ORDER_CHANNELS, [ + "downloadedNoTranscript", + "failedListed", + ]); + const after = buildPendingByLeaf( + ORDER_ROOT, + ORDER_CHANNELS, + ["downloadedNoTranscript", "failedListed"], + {}, + ); + assert.deepEqual(before.all, ["c_old", "c_new", "d_mid"]); + assert.deepEqual(after, before); +}); + +test("buildPendingByLeaf: a comparator reorders within a leaf, across channels", () => { + const newest = buildPendingByLeaf( + ORDER_ROOT, + ORDER_CHANNELS, + ["downloadedNoTranscript", "failedListed"], + { compare: byDate(ORDER_DATES, "newest") }, + ); + assert.deepEqual(newest.all, ["c_new", "d_mid", "c_old"]); + + const oldest = buildPendingByLeaf( + ORDER_ROOT, + ORDER_CHANNELS, + ["downloadedNoTranscript", "failedListed"], + { compare: byDate(ORDER_DATES, "oldest") }, + ); + assert.deepEqual(oldest.all, ["c_old", "d_mid", "c_new"]); +}); + +test("ordering sorts INSIDE a rule; rule order still wins", () => { + // The documented semantic. `corn` is declared first, so its videos are served + // first even though destiny's is newer than one of them — ordering never + // promotes a video past a higher-priority RULE. + const root: AutoQueueGroup = { + id: "root", + mode: "strict", + children: [ + { id: "corn", match: { type: "channel", value: "cornbreadman" } }, + { id: "rest", match: { type: "all" } }, + ], + }; + const pending = buildPendingByLeaf( + root, + ORDER_CHANNELS, + ["downloadedNoTranscript", "failedListed"], + { compare: byDate(ORDER_DATES, "newest") }, + ); + assert.deepEqual(pending.corn, ["c_new", "c_old"]); + assert.deepEqual(pending.rest, ["d_mid"]); + assert.deepEqual(drain(root, pending), ["corn", "corn", "rest"]); +}); + +test("ordering does not change CLAIMING (bucket priority still wins)", () => { + // A video in two buckets is still attributed once, to the higher-priority + // bucket — sorting happens after claiming, never instead of it. + const channels: ChannelWork[] = [ + { + slug: "cornbreadman", + platform: "odysee", + buckets: { + partialDownloads: ["shared"], + undownloadedIds: ["shared", "fresh"], + }, + }, + ]; + const root: AutoQueueGroup = { + id: "root", + mode: "strict", + children: [ + { id: "partials", match: { type: "all", bucket: "partialDownloads" } }, + { id: "rest", match: { type: "all" } }, + ], + }; + const pending = buildPendingByLeaf( + root, + channels, + ["partialDownloads", "undownloadedIds"], + { compare: byDate({ shared: "20200101", fresh: "20260101" }, "newest") }, + ); + assert.deepEqual(pending.partials, ["shared"]); + assert.deepEqual(pending.rest, ["fresh"]); +}); + +test("equal keys keep listed order (stable sort, no arbitrary tiebreak)", () => { + const channels: ChannelWork[] = [ + { + slug: "cornbreadman", + platform: "odysee", + // Same upload date: whatever order they arrived in must survive, which is + // what preserves playlist order for undownloadedIds. + buckets: { undownloadedIds: ["zzz", "aaa", "mmm"] }, + }, + ]; + const pending = buildPendingByLeaf( + { id: "root", mode: "strict", children: [{ id: "all", match: { type: "all" } }] }, + channels, + ["undownloadedIds"], + { compare: byDate({ zzz: "20260101", aaa: "20260101", mmm: "20260101" }, "newest") }, + ); + assert.deepEqual(pending.all, ["zzz", "aaa", "mmm"]); +}); + +test("sanitize: order defaults to listed and rejects junk", () => { + const s = sanitizeAutoQueue({ + transcription: { enabled: true, order: "newest", root: {} }, + download: { enabled: true, order: "sideways", root: {} }, + }); + assert.equal(s.transcription.order, "newest"); + assert.equal(s.download.order, "listed"); + // A settings.json written before the field existed reads as today's behaviour. + assert.equal(sanitizeAutoQueue({}).transcription.order, "listed"); + assert.equal(defaultAutoQueue().download.order, "listed"); +}); + +test("sanitize: a lapsed snooze normalizes to null, a future one survives", () => { + const future = Date.now() + 60_000; + const s = sanitizeAutoQueue({ + transcription: { snoozeUntil: future, root: {} }, + download: { snoozeUntil: Date.now() - 60_000, root: {} }, + }); + assert.equal(s.transcription.snoozeUntil, future); + assert.equal(s.download.snoozeUntil, null); + assert.equal(sanitizeAutoQueue({}).transcription.snoozeUntil, null); + assert.equal(sanitizeAutoQueue({ download: { snoozeUntil: "soon" } }).download.snoozeUntil, null); +}); diff --git a/common/jobs/autoQueuePolicy.ts b/common/jobs/autoQueuePolicy.ts @@ -26,6 +26,24 @@ export const AUTO_QUEUE_MODES: ReadonlyArray<AutoQueueMode> = [ export type AutoQueueMatchType = "channel" | "platform" | "all"; +// How videos are ordered WITHIN a rule, across every channel and bucket that +// rule claims. "listed" is the historical behaviour: whatever order +// buildPendingByLeaf accumulated, which is bucket order over channel order over +// each bucket array's own order (channelSnapshot sorts most buckets +// lexicographically by video id, and leaves undownloadedIds in playlist order). +// "newest"/"oldest" sort each leaf's claimed list by an upload-date key supplied +// by the caller as a comparator — the engine stays pure and never reads disk. +// +// This deliberately does NOT reorder RULES: the tree is what expresses +// priority. A newest-first archive is one catch-all rule with order "newest". +export type AutoQueueOrder = "listed" | "newest" | "oldest"; + +export const AUTO_QUEUE_ORDERS: ReadonlyArray<AutoQueueOrder> = [ + "listed", + "newest", + "oldest", +]; + export type AutoQueueMatch = { type: AutoQueueMatchType; // Channel slug (type=channel) or platform name (type=platform). Ignored for @@ -77,6 +95,15 @@ export type AutoQueuePolicy = { // per-channel opt-in without flipping this switch. Optional: settings written // before this field existed lack it; the sanitizer defaults it to false. replaceAutoSubs?: boolean; + // Ordering within each rule (see AutoQueueOrder). Optional exactly like + // replaceAutoSubs: settings files written before this field existed lack it, + // and the sanitizer defaults them to "listed" (today's behaviour). + order?: AutoQueueOrder; + // Epoch ms until which this runner idles WITHOUT stopping: next() returns null + // so the loop stays up, re-reads settings each iteration, and resumes by + // itself when the moment passes. null/absent/past = not snoozed. Survives a + // restart because it lives in settings.json, not in runner memory. + snoozeUntil?: number | null; root: AutoQueueGroup; }; @@ -204,6 +231,15 @@ export function buildPendingByLeaf( root: AutoQueueNode, channels: ReadonlyArray<ChannelWork>, defaultBuckets: ReadonlyArray<string>, + opts?: { + // Optional total order applied to each leaf's FINISHED list, after claiming. + // Claiming itself is untouched — a video in two buckets is still attributed + // once, to the higher-priority bucket — so this only decides which of a + // leaf's own videos goes first. Supplied by the runner from the recency + // index (common/controller/recencyIndex.ts) when policy.order != "listed"; + // omitting it reproduces the historical order byte-for-byte. + compare?: (a: string, b: string) => number; + }, ): Record<string, string[]> { const leaves = flattenLeaves(root); const pending: Record<string, string[]> = {}; @@ -224,6 +260,10 @@ export function buildPendingByLeaf( } } } + if (opts?.compare) { + const compare = opts.compare; + for (const ids of Object.values(pending)) ids.sort(compare); + } return pending; } @@ -399,6 +439,8 @@ export function defaultAutoQueuePolicy(): AutoQueuePolicy { enabled: false, maxWorkers: null, replaceAutoSubs: false, + order: "listed", + snoozeUntil: null, root: { id: "root", mode: "strict", weight: 1, maxWorkers: null, children: [] }, }; } @@ -410,6 +452,16 @@ export function defaultAutoQueue(): AutoQueueSettings { }; } +// A snooze that has already lapsed is not a snooze: normalizing it to null here +// means every reader (runner, status payload, UI) can treat "non-null" as "still +// snoozed" without repeating the clock comparison. Re-sanitized on every read of +// settings.json, so a stale value self-clears without anyone writing. +function sanitizeSnooze(value: unknown): number | null { + if (typeof value !== "number" || !Number.isFinite(value)) return null; + const at = Math.floor(value); + return at > Date.now() ? at : null; +} + function sanitizePolicy(value: unknown): AutoQueuePolicy { const r = (value ?? {}) as Record<string, unknown>; const seen = new Set<string>(); @@ -419,6 +471,9 @@ function sanitizePolicy(value: unknown): AutoQueuePolicy { // Opt-in only: anything but an explicit `true` (including a missing field on // a pre-existing settings.json) leaves the lane off. replaceAutoSubs: r.replaceAutoSubs === true, + // Anything unrecognised (including a missing field) means today's behaviour. + order: r.order === "newest" || r.order === "oldest" ? r.order : "listed", + snoozeUntil: sanitizeSnooze(r.snoozeUntil), root: sanitizeRoot(r.root, seen), }; } diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,13 @@ # Changelog ## [Unreleased] +- **The auto-queue can be told to do the newest uploads first, and it now genuinely does.** Both runners always took the first video off a rule's pile, and that pile's order came straight from the channel snapshot, which sorts most buckets **alphabetically by video id** — arbitrary for YouTube ids, and oldest-first for the date-prefixed folder names some sites use. So when a channel uploaded today, nothing made that video jump the nine-thousand-video backlog in front of it; the only reason auto-download roughly worked was that its one bucket happens to be left in playlist order. There is now an **Order** setting per runner — *Listed order* (what you have today, and still the default), *Newest first*, *Oldest first*. It sorts the videos **inside** each rule, across every channel and bucket that rule claims; the rule list still decides which rule goes first, because that is what the rule list is for. For a straight newest-first archive, use one catch-all rule. Worth knowing before you switch it on: under *Newest first* the retry and partial-download buckets lose their head start, so a half-finished download can end up waiting behind fresh work. The page says so next to the setting. +- **Working out how recent 79,000 videos are turned out to be nearly free, once we stopped guessing where the dates were.** The obvious source — reading each video's metadata file — is about six and a half minutes and 41 GB of reading, on every scheduling decision, which is a non-starter. The transcript index already holds a date per video in a form that can be scanned without decoding anything: **78,583 videos in well under a second**. That covers the corpus, but it turned out **not** to cover the videos auto-transcribe actually queues, because the index only holds videos that already *have* a transcript and auto-transcribe's whole job is the ones that don't — of the 870 videos genuinely pending here, it knew the date of **122**. The gap is closed by reading the last 8 KB of each remaining video's metadata file, where the upload date happens to sit: **868 of the 870, at a fifth of a millisecond each**, and remembered afterwards so it is paid once rather than every few seconds. Videos not downloaded yet have no date anywhere on disk at all, so auto-download estimates one from the video's position in the channel's newest-first listing; those show with a `≈`, and a brand-new upload with nothing dated above it goes to the front, which is the entire point. Anything still undatable sorts to the back rather than disappearing, and a missing or busy index degrades to the old ordering instead of stopping the runner. +- **The auto-queue page now answers "what is it doing, and why not?".** It was two identical panels of pending counts and a pick log — and the payload it was already receiving contained the answers to both questions, thrown away on arrival. A **dispatch deck** across the top gives both runners at a glance, so you never scroll to find out about the other one. Each lane then reads top to bottom as the questions you actually arrive with. **Next up** names the exact video the policy would hand out next, with its upload date, the rule that claimed it, and *why that rule* — "rule 1 has no pending work" — which is also the fastest way to see that an ordering change did what you asked. **In flight** lists what is running right now with how long each has been going, which the page received and rendered as a bare count. The policy tree became a **claim ladder**: one rung per rule, numbered by its real priority, saying in plain language what it matches, with a hairline down its left edge whose fill shows how much of the runner that rung is currently holding. The pending count on each rung opens to show the actual videos, in the exact order they will be handed out. This folds three previously separate sections — the tree, "Pending per rule" and "Recent picks", which you had to cross-reference by eye — into one object. The lane header gains how long the runner has been up and a link to its job log, both of which were on the wire and discarded. +- **A running auto-queue runner that is doing nothing now says why.** Every reason it stops short — nothing pending, every route at a worker cap, no enabled worker, the disk gate, a rate-limit cooldown, the global downloads pause — was already worked out inside the loop and then discarded, so a runner with nothing to do looked exactly like a wedged one. The lane now states the reason in words. +- **Edits to an auto-queue policy no longer vanish when you navigate away.** Reordering rules, changing a match, adding a group — none of it was saved until you pressed the button, and nothing on the page said so, so leaving the page silently discarded the lot. There is now an **unsaved changes** marker with **Discard changes** beside Save, and the browser asks before you leave. The page also picks up policy changes made elsewhere — "Add to top of auto-queue" on the channels table rewrites this very policy — instead of showing a stale copy, and it will not do that while you have edits in progress. +- **An auto-queue runner can be snoozed instead of stopped.** Stopping it means remembering to start it again; a snooze — 1 hour, 4 hours, or until tomorrow morning — leaves the runner up and idle, and it resumes by itself when the time passes. Because the deadline is stored in settings rather than held in the running process, it survives a restart. **Wake now** ends it early. +- **Three auto-queue controls that a screen reader could not name now have names.** The group-mode dropdown and the rule match-type dropdown had no label at all, and the move-up / move-down buttons were announced as their arrow glyphs. - **The Homepage page is about the project's site now, not "the hub".** The `homepage` package stopped being *our shelf of archives* and became **Archilyzer's own site** — marketing home, documentation, source download — with the cross-site dashboard moved to `/stats/`. The editor page says so, and names the split it now sits on: the wordmark and page title are the **product's** identity and are no longer editable here, while the social links and the deploy target remain **yours**. The practical consequence to know about: the deployed site's meta description no longer comes from `homepage.json`, so the corpus-flavoured blurb that used to describe it (*"Search the transcripts of your favorite and least favorite creators!"*) has been replaced by a description of the software. - **Placeholder URLs stopped advertising a domain that isn't ours.** Four form hints and two library comments offered `https://archilyzer.com` as the example hub URL. That domain **belongs to someone else** — it redirects to an unrelated page — so anyone who typed the example verbatim would have pointed their sites at a stranger. Every occurrence now reads `https://archilyzer.pages.dev`, which is the real one. - **The backfill card now breaks its figure down by kind, and every lane card gained a second line.** Backfill read "77,952 reachable · 77,134 need media", which looks inverted and is not: three kinds are running, and summing them produces a number in no unit at all. 99.5% of that "reachable" is attribution-from-text, which costs about one model call per transcript *chunk*; the whole of "need media" is diarization, which is a few hundred audio passes. Worse, the two overlap — most of the videos needing media for diarization are also inside the text-attribution total — so the corpus read as simultaneously all-actionable and all-blocked, and the biggest single fact was invisible: **77,463 videos are blocked on diarization's output**, which is precisely why attribution can only run text-only. The card now prints a line per kind (`Speaker diarization 329 · 77,134 need media`), each clause dropped at zero, and never adds them together. The other three lanes stopped being a status word over a sentence: transcription and downloads name their backlog and how many channels it spans, transcription adds worker occupancy, downloads adds free disk against the floor, and digest names how many channels the layer has reached plus the videos waiting on a transcript or held for want of a normalized one. diff --git a/editor/app/auto-queue/actions.ts b/editor/app/auto-queue/actions.ts @@ -15,6 +15,7 @@ import { isGroup, type AutoQueueGroup, type AutoQueueNode, + type AutoQueueOrder, } from "yt-dlp-transcript-common/jobs/autoQueuePolicy"; export type SaveResult = { ok: true } | { ok: false; error: string }; @@ -31,10 +32,17 @@ export async function saveAutoQueueAction( maxWorkers: number | null; // Opt in to the lowest-priority replace-auto-captions lane (default false). replaceAutoSubs: boolean; + // Ordering WITHIN each rule (newest / oldest upload first, or listed order). + order: AutoQueueOrder; root: AutoQueueGroup; }, ): Promise<SaveResult> { const current = getSettings(); + // NOTE: this object lists every persisted policy field EXPLICITLY, so a field + // added to AutoQueuePolicy and forgotten here is silently dropped on every + // save rather than failing loudly. `snoozeUntil` is deliberately carried over + // from `current` instead of taken from the form: the snooze is set by its own + // action, and a policy save (e.g. reordering rules) must not cancel it. const next: SiteSettings = { ...current, autoQueue: { @@ -43,6 +51,8 @@ export async function saveAutoQueueAction( enabled: input.enabled, maxWorkers: input.maxWorkers, replaceAutoSubs: input.replaceAutoSubs === true, + order: input.order, + snoozeUntil: current.autoQueue[kind].snoozeUntil ?? null, root: input.root, }, }, @@ -80,6 +90,33 @@ export async function stopAutoQueueAction( return { ok: true }; } +// Idle a runner until `untilMs` (epoch ms) without stopping it, or wake it now +// with null. Deliberately NOT part of saveAutoQueueAction: snoozing is a +// one-click operational act, and routing it through the policy form would mean a +// pending tree edit had to be saved (or discarded) to snooze. The runner re-reads +// settings every iteration, so this takes effect on the next tick — and because +// it lives in settings.json rather than runner memory, it survives a restart. +export async function snoozeAutoQueueAction( + kind: AutoQueueKind, + untilMs: number | null, +): Promise<SaveResult> { + const current = getSettings(); + const next: SiteSettings = { + ...current, + autoQueue: { + ...current.autoQueue, + [kind]: { ...current.autoQueue[kind], snoozeUntil: untilMs }, + }, + }; + try { + await writeSettings(next); + } catch (e) { + return { ok: false, error: (e as Error).message }; + } + revalidatePath("/auto-queue"); + return { ok: true }; +} + // Recursively drop every leaf that matches this channel, so re-prioritizing the // same channel doesn't accumulate duplicate leaves (a group emptied of children // is kept — sanitizeAutoQueue tolerates it, and removing it could orphan a diff --git a/editor/app/auto-queue/components/AutoQueueView.tsx b/editor/app/auto-queue/components/AutoQueueView.tsx @@ -7,14 +7,27 @@ import type { AutoQueueKindStatus, PlatformCooldownView, } from "../status"; +import { DispatchDeck } from "./DispatchDeck"; +import { InFlightList } from "./InFlightList"; +import { LaneHeader } from "./LaneHeader"; +import { NextUp } from "./NextUp"; import { PolicyTreeEditor } from "./PolicyTreeEditor"; +import { SnoozeControl } from "./SnoozeControl"; +import { type Channel, formatClock, formatCooldown, leafOrder } from "./dispatch"; -// Live panel for the auto-queue runners: one section per kind (auto-transcribe / -// auto-download), each with a status header + Start/Stop, the policy-tree editor, -// snapshot-derived pending counts per rule, and the recent pick log. Polls -// /api/auto-queue/status like the scheduler panel. - -type Channel = { slug: string; name: string | null }; +// The dispatcher board. A deck showing both runners, then one lane per runner +// answering, top to bottom: what is it about to do, what is it doing, by what +// rules, and what did it just do. +// +// STRUCTURAL CONTRACT, load-bearing for ~1,200 lines of e2e: +// * each kind is a literal <section> containing an <h2> named exactly +// "Auto-transcribe" / "Auto-download", with the policy editor and the +// Start/Drain/Stop buttons inside it; +// * NOTHING inside a kind section is itself a <section> — the suite scopes with +// locator("section", { has: heading }), and a nested one would match two +// ancestors, failing every scoped lookup on strict mode; +// * the deck names runners in plain text, never as headings, for the same +// reason. export function AutoQueueView({ initial, @@ -29,6 +42,7 @@ export function AutoQueueView({ }) { const [data, setData] = useState<AutoQueueStatusPayload>(initial); const mounted = useRef(true); + const now = useNow(); const refresh = useCallback(async () => { try { @@ -52,35 +66,39 @@ export function AutoQueueView({ return ( <div className="flex flex-col gap-8"> - <KindPanel + <DispatchDeck data={data} now={now} /> + <KindLane kind="transcription" title="Auto-transcribe" status={data.transcription} channels={channels} platforms={platforms} buckets={bucketsByKind.transcription} + now={now} onRefresh={refresh} /> - <KindPanel + <KindLane kind="download" title="Auto-download" status={data.download} channels={channels} platforms={platforms} buckets={bucketsByKind.download} + now={now} onRefresh={refresh} /> </div> ); } -function KindPanel({ +function KindLane({ kind, title, status, channels, platforms, buckets, + now, onRefresh, }: { kind: AutoQueueKind; @@ -89,10 +107,10 @@ function KindPanel({ channels: Channel[]; platforms: string[]; buckets: string[]; + now: number | null; onRefresh: () => Promise<void>; }) { const [busy, setBusy] = useState(false); - const running = status.runner.running; const control = useCallback( async (action: "start" | "stop" | "drain") => { @@ -111,60 +129,21 @@ function KindPanel({ [kind, onRefresh], ); - const leafName = (leafId: string) => leafLabel(status.policy.root, leafId); - const totalPending = Object.values(status.pendingByLeaf).reduce( - (a, b) => a + b, - 0, - ); - return ( - <section className="flex flex-col gap-4"> - <div className="flex flex-wrap items-center gap-3"> - <h2 className="text-lg font-semibold">{title}</h2> - <span - className={`text-sm rounded-full px-3 py-1 border ${ - running - ? "border-success/30 bg-success-soft text-success" - : "border-border bg-card text-muted-foreground" - }`} - > - {running ? "Runner running" : "Runner stopped"} - </span> - <span className="text-sm text-muted-foreground"> - {status.runner.inFlight.length} in flight · {totalPending} pending - </span> - <span className="ml-auto flex gap-2"> - <button - type="button" - aria-label={`Start ${title}`} - onClick={() => control("start")} - disabled={busy || running} - className="px-3 py-1.5 rounded-md bg-primary text-primary-foreground text-sm font-medium hover:opacity-90 disabled:opacity-50" - > - Start - </button> - <button - type="button" - aria-label={`Drain ${title}`} - title="Finish the in-flight items, start no new ones, then stop" - onClick={() => control("drain")} - disabled={busy || !running} - className="px-3 py-1.5 rounded-md border border-warning/30 text-sm font-medium text-warning hover:bg-warning-soft disabled:opacity-50" - > - Drain - </button> - <button - type="button" - aria-label={`Stop ${title}`} - title="Stop now, aborting any in-flight item" - onClick={() => control("stop")} - disabled={busy || !running} - className="px-3 py-1.5 rounded-md border border-border text-sm font-medium hover:bg-muted disabled:opacity-50" - > - Stop - </button> - </span> - </div> + // data-hydrated is a TESTING AFFORDANCE, and a deliberate one. A React state + // update that lands before hydration is silently discarded — no error, no + // request, nothing — and this repo has lost days to that failure mode more + // than once. `now` is null until the client mount effect runs, so it is an + // exact "React is live on this subtree" signal that already exists; exposing + // it lets a test wait for the page rather than race it. + <section className="flex flex-col gap-4" data-hydrated={now !== null ? "true" : undefined}> + <LaneHeader + title={title} + status={status} + now={now} + busy={busy} + onControl={control} + /> {kind === "transcription" && ( <p className="text-xs text-muted-foreground"> @@ -176,102 +155,68 @@ function KindPanel({ )} {status.cooldowns.length > 0 && ( - <CooldownBanner cooldowns={status.cooldowns} /> + <CooldownStrip cooldowns={status.cooldowns} now={now} /> )} + <SnoozeControl + kind={kind} + title={title} + snoozeUntil={status.policy.snoozeUntil ?? null} + now={now} + onChanged={onRefresh} + /> + + <NextUp status={status} channels={channels} /> + <InFlightList status={status} channels={channels} now={now} /> + <PolicyTreeEditor kind={kind} - initialEnabled={status.policy.enabled} - initialMaxWorkers={status.policy.maxWorkers} - initialReplaceAutoSubs={status.policy.replaceAutoSubs === true} - initialRoot={status.policy.root} + status={status} channels={channels} platforms={platforms} buckets={buckets} /> - <div className="grid gap-5 md:grid-cols-2"> - <div className="flex flex-col gap-2"> - <h3 className="text-sm font-semibold">Pending per rule</h3> - {Object.keys(status.pendingByLeaf).length === 0 ? ( - <p className="text-sm text-muted-foreground">No rules configured.</p> - ) : ( - <ul className="flex flex-col gap-1 text-sm"> - {Object.entries(status.pendingByLeaf).map(([leafId, count]) => ( - <li - key={leafId} - className="flex justify-between gap-3 text-muted-foreground" - > - <span>{leafName(leafId)}</span> - <span className="tabular-nums">{count}</span> - </li> - ))} - </ul> - )} - </div> - - <div className="flex flex-col gap-2"> - <h3 className="text-sm font-semibold">Recent picks</h3> - {status.picks.length === 0 ? ( - <p className="text-sm text-muted-foreground">Nothing picked yet.</p> - ) : ( - <ul className="flex flex-col gap-1 text-sm"> - {status.picks.slice(0, 15).map((p, i) => ( - <li - key={`${p.at}-${i}`} - className="flex flex-wrap gap-x-3 text-muted-foreground" - > - <span className="tabular-nums text-muted-foreground"> - {formatClock(p.at)} - </span> - <span> - {p.channelSlug}/{p.videoId} - </span> - <span className="text-muted-foreground">{leafName(p.leafId)}</span> - </li> - ))} - </ul> - )} - </div> - </div> + <RecentPicks status={status} /> </section> ); } -// Human label for a leaf id from the current policy tree (its match), falling -// back to the raw id. -function leafLabel( - node: AutoQueueKindStatus["policy"]["root"], - leafId: string, -): string { - type N = - | typeof node - | { id: string; match?: { type: string; value?: string; bucket?: string } }; - const walk = (n: N): string | null => { - if (n.id === leafId && "match" in n && n.match) { - const m = n.match; - const base = - m.type === "all" ? "all channels" : `${m.type}: ${m.value ?? "?"}`; - return m.bucket ? `${base} [${m.bucket}]` : base; - } - const children = (n as { children?: N[] }).children; - if (children) { - for (const c of children) { - const found = walk(c); - if (found) return found; - } - } - return null; - }; - return walk(node) ?? leafId; -} - -function formatClock(ms: number): string { - if (!ms) return "—"; - return new Date(ms).toLocaleTimeString([], { - hour: "2-digit", - minute: "2-digit", - }); +// The pick log, now cross-referenced against the ladder by ORDINAL rather than +// by the leafLabel() id lookup this page used to carry — the rung numbers are +// right there, so "rule 2" is a pointer you can follow with your eye. +function RecentPicks({ status }: { status: AutoQueueKindStatus }) { + const leaves = leafOrder(status.policy.root); + return ( + <div className="flex flex-col gap-1"> + <p className="font-mono text-xs uppercase tracking-[0.14em] text-muted-foreground"> + Recent picks + </p> + {status.picks.length === 0 ? ( + <p className="text-sm text-muted-foreground">Nothing picked yet.</p> + ) : ( + <ul className="flex flex-col gap-0.5 text-sm"> + {status.picks.slice(0, 15).map((p, i) => { + const at = leaves.findIndex((l) => l.id === p.leafId); + return ( + <li + key={`${p.at}-${i}`} + className="flex flex-wrap gap-x-3 text-muted-foreground" + > + <span className="tabular-nums">{formatClock(p.at)}</span> + <span className="font-mono text-foreground"> + {p.channelSlug}/{p.videoId} + </span> + <span className="text-xs"> + {at >= 0 ? `rule ${at + 1}` : p.leafId} + </span> + </li> + ); + })} + </ul> + )} + </div> + ); } // Wall-clock that re-renders each second, starting null so SSR and the first @@ -288,8 +233,13 @@ function useNow(): number | null { // Platforms paused by a rate-limit/network backoff. A manual Sync on one of // these is refused until it lapses; the runner skips it meanwhile. -function CooldownBanner({ cooldowns }: { cooldowns: PlatformCooldownView[] }) { - const now = useNow(); +function CooldownStrip({ + cooldowns, + now, +}: { + cooldowns: PlatformCooldownView[]; + now: number | null; +}) { return ( <div className="flex flex-col gap-1 rounded-md border border-warning/30 bg-warning-soft px-3 py-2 text-sm"> <span className="font-medium text-warning"> @@ -314,10 +264,3 @@ function CooldownBanner({ cooldowns }: { cooldowns: PlatformCooldownView[] }) { </div> ); } - -function formatCooldown(secs: number): string { - if (secs < 60) return `${secs}s`; - const m = Math.floor(secs / 60); - const s = secs % 60; - return s === 0 ? `${m}m` : `${m}m ${s}s`; -} diff --git a/editor/app/auto-queue/components/ClaimLadder.tsx b/editor/app/auto-queue/components/ClaimLadder.tsx @@ -0,0 +1,48 @@ +"use client"; + +import type { AutoQueueGroup } from "yt-dlp-transcript-common/jobs/autoQueuePolicy"; +import { LadderRung, type RungData, type RungOps } from "./LadderRung"; +import type { Channel } from "./dispatch"; + +// The policy tree, rendered as the dispatch mechanism it actually is: one rung +// per node, in priority order, each carrying its own live occupancy and backlog. +// +// A plain <div>, never a <section> — see NextUp for why a nested section breaks +// every scoped lookup in the e2e suite. + +export function ClaimLadder({ + root, + channels, + platforms, + buckets, + data, + ops, +}: { + root: AutoQueueGroup; + channels: Channel[]; + platforms: string[]; + buckets: string[]; + data: RungData; + ops: RungOps; +}) { + return ( + <div className="flex flex-col gap-2 rounded-md border border-border bg-card px-3 py-2"> + <p className="font-mono text-xs uppercase tracking-[0.14em] text-muted-foreground"> + Policy + </p> + <LadderRung + node={root} + depth={0} + parentId={null} + parentMode={null} + index={0} + siblingCount={1} + channels={channels} + platforms={platforms} + buckets={buckets} + data={data} + ops={ops} + /> + </div> + ); +} diff --git a/editor/app/auto-queue/components/DispatchDeck.tsx b/editor/app/auto-queue/components/DispatchDeck.tsx @@ -0,0 +1,96 @@ +"use client"; + +import type { AutoQueueStatusPayload, AutoQueueKindStatus } from "../status"; +import { formatElapsed, idleReasonText } from "./dispatch"; + +// Both runners at a glance, above the fold, so you never scroll to learn the +// state of the other one. State ONLY — every control lives in the lane below. +// That is not just tidiness: two Start buttons would give the page two elements +// with the same accessible name, and the e2e suite addresses them by name. +// +// The runner names here are deliberately PLAIN TEXT, not headings. The whole +// suite scopes itself with locator("section", { has: heading "Auto-transcribe" }), +// and a second heading with that name would make every one of those lookups +// ambiguous. + +export function DispatchDeck({ + data, + now, +}: { + data: AutoQueueStatusPayload; + now: number | null; +}) { + return ( + <div className="rounded-lg border border-border bg-card"> + <p className="border-b border-border px-4 py-2 font-mono text-xs uppercase tracking-[0.14em] text-muted-foreground"> + Dispatch + </p> + <div className="divide-y divide-border"> + <DeckRow label="Auto-transcribe" status={data.transcription} now={now} /> + <DeckRow label="Auto-download" status={data.download} now={now} /> + </div> + </div> + ); +} + +function DeckRow({ + label, + status, + now, +}: { + label: string; + status: AutoQueueKindStatus; + now: number | null; +}) { + const running = status.runner.running; + const inFlight = status.runner.inFlight.length; + const pending = Object.values(status.pendingByLeaf).reduce((a, b) => a + b, 0); + const idle = running ? idleReasonText(status.runner.idleReason, status.kind) : null; + // Working > held > stopped. A running runner with nothing in flight is not the + // same as a stopped one, and the dot is the only thing that says so at a + // glance — so it takes the "attention" tone rather than the neutral one. + const tone = !running + ? "bg-muted-foreground/40" + : inFlight > 0 + ? "bg-info animate-pulse motion-reduce:animate-none" + : "bg-warning"; + + return ( + <div className="flex flex-wrap items-baseline gap-x-3 gap-y-1 px-4 py-2.5 text-sm"> + <span + aria-hidden="true" + className={`size-2 shrink-0 self-center rounded-full ${tone}`} + /> + <span className="min-w-40 font-medium text-foreground">{label}</span> + <span className="text-muted-foreground"> + {!running ? ( + "not running" + ) : ( + <> + up{" "} + <span className="tabular-nums"> + {now !== null && status.runner.startedAt + ? formatElapsed(now - status.runner.startedAt) + : "—"} + </span> + {idle ? ` · idle: ${idle}` : ""} + </> + )} + </span> + <span className="ml-auto flex items-baseline gap-4 text-muted-foreground"> + <span> + <span className="font-display tabular-nums text-foreground"> + {inFlight} + </span>{" "} + in flight + </span> + <span> + <span className="font-display tabular-nums text-foreground"> + {pending.toLocaleString()} + </span>{" "} + pending + </span> + </span> + </div> + ); +} diff --git a/editor/app/auto-queue/components/HowPriorityWorks.tsx b/editor/app/auto-queue/components/HowPriorityWorks.tsx @@ -0,0 +1,51 @@ +"use client"; + +import { useState } from "react"; +import Link from "next/link"; +import { + Collapsible, + CollapsibleContent, + CollapsibleTrigger, +} from "yt-dlp-transcript-common/components/ui/collapsible"; + +// The five-line paragraph that used to sit under the page title, folded away. +// It is reference material — true, worth having, and read once — so it was +// costing every subsequent visit the height of the answer to "what is it doing +// right now", which is the question people actually arrive with. + +export function HowPriorityWorks() { + const [open, setOpen] = useState(false); + return ( + <Collapsible open={open} onOpenChange={setOpen}> + <CollapsibleTrigger className="text-sm text-muted-foreground underline underline-offset-2 hover:text-foreground"> + How priority works {open ? "⌃" : "⌄"} + </CollapsibleTrigger> + <CollapsibleContent> + <div className="mt-2 flex flex-col gap-2 border-l-2 border-border pl-3 text-sm text-muted-foreground"> + <p> + Rules are read top to bottom: the highest one with available work + claims the next freed worker slot, and when its work runs out the + runner falls through to the next rule automatically. Wrap rules in a + group set to <em>round-robin</em> or <em>weighted-fair</em> to + alternate between them instead. + </p> + <p> + <strong className="text-foreground">Order</strong> sorts videos{" "} + <em>inside</em> each rule, across every channel and bucket that rule + claims. Rule order still wins — for a pure newest-first archive, use + one catch-all rule. + </p> + <p> + A video is claimed by exactly one rule (the first that matches), so + overlapping rules never double-process it. This is independent of the{" "} + <Link href="/scheduler" className="underline"> + sync schedule + </Link> + , which only decides when to re-fetch each channel, and manual + batches keep working alongside it. + </p> + </div> + </CollapsibleContent> + </Collapsible> + ); +} diff --git a/editor/app/auto-queue/components/InFlightList.tsx b/editor/app/auto-queue/components/InFlightList.tsx @@ -0,0 +1,63 @@ +"use client"; + +import Link from "next/link"; +import type { AutoQueueKindStatus } from "../status"; +import { type Channel, formatElapsed, leafOrder, leafSentence } from "./dispatch"; + +// What the runner is doing RIGHT NOW. `runner.inFlight` has always been in the +// status payload — video id, owning channel, the leaf that claimed it and when +// it started — and the old page rendered a count and threw the rest away. + +export function InFlightList({ + status, + channels, + now, +}: { + status: AutoQueueKindStatus; + channels: Channel[]; + now: number | null; +}) { + const items = status.runner.inFlight; + if (items.length === 0) return null; + const leaves = leafOrder(status.policy.root); + + return ( + <div className="flex flex-col gap-1 rounded-md border border-border bg-card px-3 py-2"> + <p className="font-mono text-xs uppercase tracking-[0.14em] text-muted-foreground"> + In flight + </p> + <ul className="flex flex-col gap-1"> + {[...items] + .sort((a, b) => a.startedAt - b.startedAt) + .map((item) => { + const at = leaves.findIndex((l) => l.id === item.leafId); + const leaf = at >= 0 ? leaves[at] : null; + return ( + <li + key={item.videoId} + className="flex flex-wrap items-baseline gap-x-2 gap-y-0.5 text-sm" + > + <span + aria-hidden="true" + className="size-1.5 shrink-0 self-center rounded-full bg-info animate-pulse motion-reduce:animate-none" + /> + <Link + href={`/channels/${item.channelSlug}/videos/${item.videoId}`} + className="font-mono text-foreground underline underline-offset-2 hover:text-brand" + > + {item.channelSlug}/{item.videoId} + </Link> + <span className="text-xs text-muted-foreground"> + {at >= 0 ? `rule ${at + 1}` : item.leafId} + {leaf ? ` · ${leafSentence(leaf, channels)}` : ""} + </span> + <span className="ml-auto tabular-nums text-xs text-muted-foreground"> + {now !== null ? formatElapsed(now - item.startedAt) : "—"} + </span> + </li> + ); + })} + </ul> + </div> + ); +} diff --git a/editor/app/auto-queue/components/LadderRung.tsx b/editor/app/auto-queue/components/LadderRung.tsx @@ -0,0 +1,501 @@ +"use client"; + +import Link from "next/link"; +import { useState } from "react"; +import { + type AutoQueueGroup, + type AutoQueueLeaf, + type AutoQueueMode, + type AutoQueueNode, + AUTO_QUEUE_MODES, + isGroup, +} from "yt-dlp-transcript-common/jobs/autoQueuePolicy"; +import { Button } from "yt-dlp-transcript-common/components/ui/button"; +import { + Collapsible, + CollapsibleContent, + CollapsibleTrigger, +} from "yt-dlp-transcript-common/components/ui/collapsible"; +import { + type Channel, + MODE_GLOSS, + NUM_CLASS, + SELECT_CLASS, + formatRecency, + leafSentence, +} from "./dispatch"; + +// One rung of the claim ladder: a rule (or a group of them) as a single line +// that carries its ordinal, what it matches, how much of the runner it is +// currently holding, and how much work it owns — and is simultaneously the +// control that edits it. +// +// This replaces a tree editor that sat above a separate "Pending per rule" list +// and a separate "Recent picks" list, all three cross-referenced by a leafLabel() +// id lookup. Putting them on one object deletes that lookup and answers +// "which rule is doing the work" by looking at it. +// +// DEPTH IS THE RAIL, NOT AN INDENT. Each level draws its own hairline on the +// left, and that same hairline carries the claim fill — so nesting and +// occupancy are one mark instead of a margin plus a chip. + +export type RungData = { + pendingByLeaf: Record<string, number>; + pendingHeadByLeaf: Record<string, string[]>; + ownerByVideo: Record<string, string>; + recencyByVideo: Record<string, { key: string; estimated: boolean }>; + activeByNode: Record<string, number>; + // Denominator for the claim fill: the runner's live capacity. + capacity: number; + nextUpLeafId: string | null; + // Pre-order leaf ids, so a rung knows its own ordinal without re-walking. + leafIds: string[]; +}; + +export type RungOps = { + update: (id: string, fn: (n: AutoQueueNode) => AutoQueueNode) => void; + remove: (id: string) => void; + addChild: (parentId: string, child: AutoQueueNode) => void; + move: (parentId: string, index: number, dir: -1 | 1) => void; + moveToTop: (parentId: string, index: number) => void; + makeLeaf: (type: "channel" | "platform" | "all") => AutoQueueLeaf; + makeGroup: () => AutoQueueGroup; +}; + +export function LadderRung(props: { + node: AutoQueueNode; + depth: number; + parentId: string | null; + parentMode: AutoQueueMode | null; + index: number; + siblingCount: number; + channels: Channel[]; + platforms: string[]; + buckets: string[]; + data: RungData; + ops: RungOps; +}) { + const { node, depth, parentId, parentMode, index, siblingCount, data, ops } = + props; + const group = isGroup(node); + const active = data.activeByNode[node.id] ?? 0; + const isNext = !group && data.nextUpLeafId === node.id; + + return ( + <div className="flex gap-2"> + <ClaimRail active={active} capacity={data.capacity} /> + <div className="flex min-w-0 flex-1 flex-col gap-2 pb-1"> + <div className="flex flex-wrap items-center gap-2"> + {group ? ( + <GroupControls node={node as AutoQueueGroup} update={ops.update} /> + ) : ( + <LeafControls {...props} leaf={node as AutoQueueLeaf} /> + )} + + {parentMode === "weighted-fair" && ( + <label className="flex items-center gap-1 text-xs text-muted-foreground"> + weight + <input + type="number" + min={1} + value={node.weight ?? 1} + onChange={(e) => + ops.update(node.id, (n) => ({ + ...n, + weight: Math.max(1, Number.parseInt(e.target.value, 10) || 1), + })) + } + className={`${NUM_CLASS} w-16`} + /> + </label> + )} + + <label className="flex items-center gap-1 text-xs text-muted-foreground"> + max + <input + type="number" + min={1} + placeholder="∞" + value={node.maxWorkers ?? ""} + onChange={(e) => + ops.update(node.id, (n) => ({ + ...n, + maxWorkers: + e.target.value.trim() === "" + ? null + : Math.max(1, Number.parseInt(e.target.value, 10) || 1), + })) + } + className={`${NUM_CLASS} w-16`} + /> + </label> + + {active > 0 && ( + <span className="tabular-nums text-xs text-info"> + {active} working + </span> + )} + + {!group && ( + <PendingDrilldown + leafId={node.id} + ordinal={data.leafIds.indexOf(node.id) + 1} + data={data} + /> + )} + + {isNext && ( + <span className="rounded-full border border-info/30 bg-info-soft px-2 py-0.5 text-xs text-info"> + next up + </span> + )} + + <span className="ml-auto flex items-center gap-1"> + {parentId && siblingCount > 1 && ( + <> + <Button + type="button" + size="xs" + variant="outline" + disabled={index === 0} + aria-label="Move to top" + title="Move to top" + onClick={() => ops.moveToTop(parentId, index)} + > + <span aria-hidden="true">⤒</span> + </Button> + {/* Named. These two were addressed by a bare glyph, so a screen + reader announced the button as "up arrow, button". */} + <Button + type="button" + size="xs" + variant="outline" + disabled={index === 0} + aria-label="Move up" + title="Move up" + onClick={() => ops.move(parentId, index, -1)} + > + <span aria-hidden="true">↑</span> + </Button> + <Button + type="button" + size="xs" + variant="outline" + disabled={index === siblingCount - 1} + aria-label="Move down" + title="Move down" + onClick={() => ops.move(parentId, index, 1)} + > + <span aria-hidden="true">↓</span> + </Button> + </> + )} + {parentId && ( + <Button + type="button" + size="xs" + variant="outline" + className="text-destructive hover:text-destructive" + onClick={() => ops.remove(node.id)} + > + Remove + </Button> + )} + </span> + </div> + + {group && ( + <> + <div className="flex flex-col gap-2"> + {(node as AutoQueueGroup).children.map((child, i) => ( + <LadderRung + key={child.id} + {...props} + node={child} + depth={depth + 1} + parentId={node.id} + parentMode={(node as AutoQueueGroup).mode} + index={i} + siblingCount={(node as AutoQueueGroup).children.length} + /> + ))} + {(node as AutoQueueGroup).children.length === 0 && ( + <p className="text-xs text-muted-foreground"> + Empty group — add a rule or nested group below. + </p> + )} + </div> + <div className="flex flex-wrap gap-2"> + <AddButton onClick={() => ops.addChild(node.id, ops.makeLeaf("channel"))}> + + Channel rule + </AddButton> + <AddButton onClick={() => ops.addChild(node.id, ops.makeLeaf("platform"))}> + + Platform rule + </AddButton> + <AddButton onClick={() => ops.addChild(node.id, ops.makeLeaf("all"))}> + + Catch-all rule + </AddButton> + <AddButton onClick={() => ops.addChild(node.id, ops.makeGroup())}> + + Group + </AddButton> + </div> + </> + )} + </div> + </div> + ); +} + +function AddButton({ + onClick, + children, +}: { + onClick: () => void; + children: React.ReactNode; +}) { + return ( + <Button type="button" size="xs" variant="outline" onClick={onClick}> + {children} + </Button> + ); +} + +// The signature mark. A hairline that says depth by existing at all, and says +// occupancy by how much of it is lit — so a rung holding two of the runner's +// three slots reads without a number, and the number beside it is confirmation +// rather than the only channel. +function ClaimRail({ active, capacity }: { active: number; capacity: number }) { + const pct = + capacity > 0 ? Math.min(100, Math.round((active / capacity) * 100)) : 0; + return ( + <div + aria-hidden="true" + className="relative w-0.5 shrink-0 self-stretch rounded-full bg-border" + > + {pct > 0 && ( + <span + className="absolute inset-x-0 bottom-0 rounded-full bg-info" + style={{ height: `${pct}%` }} + /> + )} + </div> + ); +} + +// The count, and behind it the actual videos in the exact order they would be +// handed out. buildPendingByLeaf already returned these arrays and +// computeLeafPendingCounts threw everything but .length away — which is also why +// there was previously no way to check an ordering change by eye. +function PendingDrilldown({ + leafId, + ordinal, + data, +}: { + leafId: string; + ordinal: number; + data: RungData; +}) { + const [open, setOpen] = useState(false); + const count = data.pendingByLeaf[leafId] ?? 0; + const head = data.pendingHeadByLeaf[leafId] ?? []; + + if (count === 0) { + return <span className="tabular-nums text-xs text-muted-foreground">0</span>; + } + + return ( + <Collapsible open={open} onOpenChange={setOpen} className="contents"> + <CollapsibleTrigger asChild> + <button + type="button" + aria-label={`Show pending videos for rule ${ordinal}`} + className="rounded px-1 tabular-nums text-xs text-foreground underline underline-offset-2 hover:text-brand" + > + {count.toLocaleString()} + <span aria-hidden="true"> {open ? "⌃" : "⌄"}</span> + </button> + </CollapsibleTrigger> + <CollapsibleContent className="w-full"> + <ul className="mt-1 flex flex-col gap-0.5 border-l border-border pl-3"> + {head.map((id) => { + const slug = data.ownerByVideo[id]; + const recency = formatRecency(data.recencyByVideo[id]); + return ( + <li + key={id} + className="flex flex-wrap items-baseline gap-x-2 text-xs" + > + {slug ? ( + <Link + href={`/channels/${slug}/videos/${id}`} + className="font-mono text-muted-foreground underline underline-offset-2 hover:text-brand" + > + {slug}/{id} + </Link> + ) : ( + <span className="font-mono text-muted-foreground">{id}</span> + )} + {recency && ( + <span className="tabular-nums text-muted-foreground"> + {recency} + </span> + )} + </li> + ); + })} + {count > head.length && ( + <li className="text-xs text-muted-foreground"> + …and {(count - head.length).toLocaleString()} more + </li> + )} + </ul> + </CollapsibleContent> + </Collapsible> + ); +} + +function GroupControls({ + node, + update, +}: { + node: AutoQueueGroup; + update: RungOps["update"]; +}) { + return ( + <> + <span className="font-mono text-xs uppercase tracking-[0.14em] text-muted-foreground"> + Group + </span> + {/* Named: this select had no label at all. */} + <select + aria-label="group mode" + value={node.mode} + onChange={(e) => + update(node.id, (n) => ({ ...n, mode: e.target.value as AutoQueueMode })) + } + className={SELECT_CLASS} + > + {AUTO_QUEUE_MODES.map((m) => ( + <option key={m} value={m}> + {m} + </option> + ))} + </select> + <span className="text-xs text-muted-foreground"> + {MODE_GLOSS[node.mode]} + </span> + </> + ); +} + +function LeafControls(props: { + leaf: AutoQueueLeaf; + channels: Channel[]; + platforms: string[]; + buckets: string[]; + data: RungData; + ops: RungOps; +}) { + const { leaf, channels, platforms, buckets, data, ops } = props; + const type = leaf.match.type; + const ordinal = data.leafIds.indexOf(leaf.id) + 1; + return ( + <> + {/* Legitimate numbering: this ordinal IS the priority the engine reads, + not decoration on an unordered list. */} + <span className="w-5 shrink-0 font-display text-sm font-semibold tabular-nums text-muted-foreground"> + {ordinal > 0 ? ordinal : "—"} + </span> + {/* Named: the match-type select had no label. */} + <select + aria-label="rule match type" + value={type} + onChange={(e) => + ops.update(leaf.id, (n) => ({ + ...n, + match: { type: e.target.value as AutoQueueLeaf["match"]["type"] }, + })) + } + className={SELECT_CLASS} + > + <option value="channel">channel</option> + <option value="platform">platform</option> + <option value="all">all</option> + </select> + + {type === "channel" && ( + <select + aria-label="channel rule value" + value={leaf.match.value ?? ""} + onChange={(e) => + ops.update(leaf.id, (n) => ({ + ...n, + match: { ...(n as AutoQueueLeaf).match, value: e.target.value }, + })) + } + className={SELECT_CLASS} + > + <option value="">— pick channel —</option> + {channels.map((c) => ( + <option key={c.slug} value={c.slug}> + {c.name ?? c.slug} + </option> + ))} + </select> + )} + + {type === "platform" && ( + <select + aria-label="platform rule value" + value={leaf.match.value ?? ""} + onChange={(e) => + ops.update(leaf.id, (n) => ({ + ...n, + match: { ...(n as AutoQueueLeaf).match, value: e.target.value }, + })) + } + className={SELECT_CLASS} + > + <option value="">— pick platform —</option> + {platforms.map((p) => ( + <option key={p} value={p}> + {p} + </option> + ))} + </select> + )} + + <label className="flex items-center gap-1 text-xs text-muted-foreground"> + bucket + {/* The option TEXT stays the raw bucket name: it is how an operator + names a bucket to the runner, and one spec asserts exactly one + option named `downloadedAutoSubsOnly` exists page-wide. The + plain-language gloss goes in the sentence below instead. */} + <select + value={leaf.match.bucket ?? ""} + aria-label="rule bucket" + onChange={(e) => + ops.update(leaf.id, (n) => { + const match = { ...(n as AutoQueueLeaf).match }; + if (e.target.value) match.bucket = e.target.value; + else delete match.bucket; + return { ...n, match }; + }) + } + className={SELECT_CLASS} + > + <option value="">all buckets (default)</option> + {buckets.map((b) => ( + <option key={b} value={b}> + {b} + </option> + ))} + </select> + </label> + + <span className="text-xs text-muted-foreground"> + {leafSentence(leaf, channels)} + </span> + </> + ); +} diff --git a/editor/app/auto-queue/components/LaneHeader.tsx b/editor/app/auto-queue/components/LaneHeader.tsx @@ -0,0 +1,141 @@ +"use client"; + +import Link from "next/link"; +import { Button } from "yt-dlp-transcript-common/components/ui/button"; +import type { AutoQueueKindStatus } from "../status"; +import { formatElapsed, idleReasonText } from "./dispatch"; + +// The lane's identity line: the <h2> every e2e lookup keys off, the state pill, +// the figures, and the three controls. +// +// PINNED, do not rename: the heading text, the pill's "Runner running" / +// "Runner stopped" wording, and the Start/Drain/Stop aria-labels. The suite +// scopes whole tests on them. What is NEW is that the payload's startedAt and +// jobId are finally rendered — both have always been on the wire and both were +// dropped on the floor, so "how long has this been up" and "show me its log" +// had no answer on the page. + +export function LaneHeader({ + title, + status, + now, + busy, + onControl, +}: { + title: string; + status: AutoQueueKindStatus; + now: number | null; + busy: boolean; + onControl: (action: "start" | "stop" | "drain") => void; +}) { + const running = status.runner.running; + const inFlight = status.runner.inFlight.length; + const pending = Object.values(status.pendingByLeaf).reduce((a, b) => a + b, 0); + const idle = running ? idleReasonText(status.runner.idleReason, status.kind) : null; + + return ( + <div className="flex flex-col gap-2"> + <div className="flex flex-wrap items-center gap-x-3 gap-y-2"> + <h2 className="font-display text-lg font-semibold tracking-tight"> + {title} + </h2> + <span + className={`rounded-full border px-3 py-1 text-sm ${ + running + ? "border-success/30 bg-success-soft text-success" + : "border-border bg-card text-muted-foreground" + }`} + > + {running ? "Runner running" : "Runner stopped"} + </span> + <span className="ml-auto flex gap-2"> + <Button + type="button" + size="sm" + aria-label={`Start ${title}`} + onClick={() => onControl("start")} + disabled={busy || running} + > + Start + </Button> + <Button + type="button" + size="sm" + variant="outline" + aria-label={`Drain ${title}`} + title="Finish the in-flight items, start no new ones, then stop" + onClick={() => onControl("drain")} + disabled={busy || !running} + className="border-warning/30 text-warning hover:bg-warning-soft hover:text-warning" + > + Drain + </Button> + <Button + type="button" + size="sm" + variant="outline" + aria-label={`Stop ${title}`} + title="Stop now, aborting any in-flight item" + onClick={() => onControl("stop")} + disabled={busy || !running} + > + Stop + </Button> + </span> + </div> + + <p className="flex flex-wrap items-center gap-x-2 gap-y-1 text-sm text-muted-foreground"> + <span className="tabular-nums text-foreground">{inFlight}</span> + <span>in flight</span> + <Sep /> + <span className="tabular-nums text-foreground"> + {pending.toLocaleString()} + </span> + <span>pending</span> + {running && status.runner.startedAt !== null && ( + <> + <Sep /> + <span> + up{" "} + <span className="tabular-nums"> + {now !== null ? formatElapsed(now - status.runner.startedAt) : "—"} + </span> + </span> + </> + )} + {status.runner.jobId && ( + <> + <Sep /> + <Link + href={`/jobs/${status.runner.jobId}`} + className="underline underline-offset-2 hover:text-foreground" + > + job log + </Link> + </> + )} + </p> + + {/* The page's answer to the most common question an operator arrives + with. Only rendered while the runner is UP: for a stopped runner the + pill above already says everything, and a second line repeating it + would collide with the suite's getByText("Runner stopped"). */} + {idle && ( + <p className="text-sm text-warning"> + <span className="font-mono text-xs uppercase tracking-[0.14em]"> + Idle + </span>{" "} + — {idle}. + </p> + )} + </div> + ); +} + +function Sep() { + return ( + <span aria-hidden="true" className="text-border"> + · + </span> + ); +} diff --git a/editor/app/auto-queue/components/NextUp.tsx b/editor/app/auto-queue/components/NextUp.tsx @@ -0,0 +1,102 @@ +"use client"; + +import Link from "next/link"; +import type { AutoQueueKindStatus } from "../status"; +import { + type Channel, + ORDER_LABEL, + formatRecency, + leafOrder, + leafSentence, +} from "./dispatch"; + +// What the runner would dispatch next, and WHY that one. +// +// Nothing here needed new machinery: the pick comes from the same +// buildChannelWork → buildPendingByLeaf → selectNextWork chain the runner uses, +// the rule path is WorkPick.path, and "why not the rule above it" is just the +// higher-priority leaves whose pending count is zero. The page simply never +// asked. +// +// A plain <div>, never a <section>: a nested section inside a kind section makes +// locator("section", { has }) match two ancestors and every scoped lookup in the +// suite fails on strict mode. + +export function NextUp({ + status, + channels, +}: { + status: AutoQueueKindStatus; + channels: Channel[]; +}) { + const next = status.nextUp; + const order = status.policy.order ?? "listed"; + + if (!next) { + return ( + <Panel> + <p className="text-sm text-muted-foreground"> + Nothing to dispatch — every rule is empty or capped. + </p> + </Panel> + ); + } + + const leaves = leafOrder(status.policy.root); + const at = leaves.findIndex((l) => l.id === next.leafId); + const leaf = at >= 0 ? leaves[at] : null; + const recency = formatRecency(next.recency ?? undefined); + const skipped = next.skippedLeafIds + .map((id) => leaves.findIndex((l) => l.id === id)) + .filter((i) => i >= 0); + + return ( + <Panel> + <p className="flex flex-wrap items-baseline gap-x-2 gap-y-1"> + {next.channelSlug ? ( + <Link + href={`/channels/${next.channelSlug}/videos/${next.videoId}`} + className="font-mono text-sm text-foreground underline underline-offset-2 hover:text-brand" + > + {next.channelSlug}/{next.videoId} + </Link> + ) : ( + <span className="font-mono text-sm text-foreground"> + {next.videoId} + </span> + )} + {recency && ( + <span className="tabular-nums text-sm text-muted-foreground"> + {recency} + </span> + )} + {at >= 0 && ( + <span className="text-sm text-muted-foreground"> + <span aria-hidden="true">← </span> + rule {at + 1} + {leaf ? ` · ${leafSentence(leaf, channels)}` : ""} + </span> + )} + </p> + <p className="text-xs text-muted-foreground"> + {skipped.length > 0 + ? `${skipped + .map((i) => `rule ${i + 1}`) + .join(", ")} ${skipped.length === 1 ? "has" : "have"} no pending work` + : "highest-priority rule with work"} + {order !== "listed" ? ` · ${ORDER_LABEL[order].toLowerCase()}` : ""} + </p> + </Panel> + ); +} + +function Panel({ children }: { children: React.ReactNode }) { + return ( + <div className="flex flex-col gap-1 rounded-md border border-border bg-card px-3 py-2"> + <p className="font-mono text-xs uppercase tracking-[0.14em] text-muted-foreground"> + Next up + </p> + {children} + </div> + ); +} diff --git a/editor/app/auto-queue/components/PolicyTreeEditor.tsx b/editor/app/auto-queue/components/PolicyTreeEditor.tsx @@ -1,27 +1,34 @@ "use client"; -import { useState } from "react"; +import { useEffect, useMemo, useRef, useState } from "react"; import { type AutoQueueGroup, type AutoQueueLeaf, - type AutoQueueMode, type AutoQueueNode, - AUTO_QUEUE_MODES, + type AutoQueueOrder, + AUTO_QUEUE_ORDERS, isGroup, } from "yt-dlp-transcript-common/jobs/autoQueuePolicy"; import type { AutoQueueKind } from "yt-dlp-transcript-common/jobs/autoQueueState"; import { saveAutoQueueAction, type SaveResult } from "../actions"; +import type { AutoQueueKindStatus } from "../status"; +import { ClaimLadder } from "./ClaimLadder"; +import type { RungOps } from "./LadderRung"; +import { SaveBar } from "./SaveBar"; +import { + type Channel, + ORDER_HINT, + ORDER_LABEL, + SELECT_CLASS, + NUM_CLASS, + leafOrder, + orderTradeoff, +} from "./dispatch"; -// A nested, controlled editor for one runner's policy tree. Groups carry a mode -// (strict / round-robin / weighted-fair) and optional worker cap; leaves match a -// channel / platform / everything, optionally narrowed to a snapshot bucket. The -// whole tree is held in React state and saved through a typed server action -// (writeSettings re-sanitizes it server-side). - -const inputClass = - "rounded border border-border bg-card px-2 py-1 text-sm"; -const btnClass = - "px-2 py-0.5 rounded border border-border text-xs hover:bg-muted"; +// Owns one runner's editable policy: the tree, the switches, and the save. The +// DRAWING of the tree moved to ClaimLadder/LadderRung, which also render the +// live counts — so what used to be a form sitting above a separate counts list +// is now one object. let idSeq = 0; function newId(): string { @@ -121,58 +128,95 @@ function makeGroup(): AutoQueueGroup { }; } -// --- Component -------------------------------------------------------------- +// --- Form state -------------------------------------------------------------- + +type Form = { + enabled: boolean; + maxWorkers: number | null; + replaceAutoSubs: boolean; + order: AutoQueueOrder; + root: AutoQueueGroup; +}; + +function formOf(status: AutoQueueKindStatus): Form { + return { + enabled: status.policy.enabled, + maxWorkers: status.policy.maxWorkers, + replaceAutoSubs: status.policy.replaceAutoSubs === true, + order: status.policy.order ?? "listed", + root: status.policy.root, + }; +} export function PolicyTreeEditor({ kind, - initialEnabled, - initialMaxWorkers, - initialReplaceAutoSubs, - initialRoot, + status, channels, platforms, buckets, }: { kind: AutoQueueKind; - initialEnabled: boolean; - initialMaxWorkers: number | null; - initialReplaceAutoSubs: boolean; - initialRoot: AutoQueueGroup; - channels: { slug: string; name: string | null }[]; + status: AutoQueueKindStatus; + channels: Channel[]; platforms: string[]; buckets: string[]; }) { - const [root, setRoot] = useState<AutoQueueGroup>(initialRoot); - const [enabled, setEnabled] = useState(initialEnabled); - const [maxWorkers, setMaxWorkers] = useState<number | null>(initialMaxWorkers); - const [replaceAutoSubs, setReplaceAutoSubs] = useState( - initialReplaceAutoSubs, - ); + const [form, setForm] = useState<Form>(() => formOf(status)); + // The last value we know is on disk. Everything dirty-related is a comparison + // against this, so "unsaved" means genuinely unsaved rather than "different + // from what the page was first rendered with". + const [baseline, setBaseline] = useState<Form>(() => formOf(status)); const [saving, setSaving] = useState(false); const [result, setResult] = useState<SaveResult | null>(null); - const update = (id: string, fn: (n: AutoQueueNode) => AutoQueueNode) => - setRoot((r) => mapNode(r, id, fn) as AutoQueueGroup); - const remove = (id: string) => setRoot((r) => removeFrom(r, id)); - const addChild = (parentId: string, child: AutoQueueNode) => - setRoot((r) => addChildTo(r, parentId, child) as AutoQueueGroup); - const move = (parentId: string, index: number, dir: -1 | 1) => - setRoot((r) => moveChildIn(r, parentId, index, dir) as AutoQueueGroup); - const moveToTop = (parentId: string, index: number) => - setRoot((r) => moveChildToTop(r, parentId, index) as AutoQueueGroup); + const dirty = useMemo( + () => JSON.stringify(form) !== JSON.stringify(baseline), + [form, baseline], + ); + // A ref so the poll effect can read dirtiness without re-subscribing to it. + const dirtyRef = useRef(dirty); + dirtyRef.current = dirty; + + // Adopt the server's copy when it changes and we have nothing to lose. This + // is what makes an off-page write visible — "Top of queue" on the channels + // table rewrites this very policy — and it also picks up the ids the server + // sanitizer assigned after our own save. Never runs while dirty, so an edit + // in progress is never clobbered by a poll. + const incoming = JSON.stringify(formOf(status)); + useEffect(() => { + if (dirtyRef.current) return; + if (incoming === JSON.stringify(baseline)) return; + const next = JSON.parse(incoming) as Form; + setBaseline(next); + setForm(next); + }, [incoming, baseline]); + + const setRoot = (fn: (r: AutoQueueGroup) => AutoQueueGroup) => + setForm((f) => ({ ...f, root: fn(f.root) })); + + const ops: RungOps = { + update: (id, fn) => setRoot((r) => mapNode(r, id, fn) as AutoQueueGroup), + remove: (id) => setRoot((r) => removeFrom(r, id)), + addChild: (parentId, child) => + setRoot((r) => addChildTo(r, parentId, child) as AutoQueueGroup), + move: (parentId, index, dir) => + setRoot((r) => moveChildIn(r, parentId, index, dir) as AutoQueueGroup), + moveToTop: (parentId, index) => + setRoot((r) => moveChildToTop(r, parentId, index) as AutoQueueGroup), + makeLeaf, + makeGroup, + }; const save = async () => { setSaving(true); setResult(null); + const attempt = form; try { - setResult( - await saveAutoQueueAction(kind, { - enabled, - maxWorkers, - replaceAutoSubs, - root, - }), - ); + const res = await saveAutoQueueAction(kind, attempt); + setResult(res); + // Only the values we actually persisted become the new baseline; a failed + // save must stay dirty so the work is still recoverable. + if (res.ok) setBaseline(attempt); } catch (e) { setResult({ ok: false, error: (e as Error).message }); } finally { @@ -180,44 +224,101 @@ export function PolicyTreeEditor({ } }; + const kindWord = kind === "transcription" ? "transcribe" : "download"; + const tradeoff = orderTradeoff(kind, form.order); + // What the claim rail's fill is a fraction OF: the runner's own ceiling when + // it has one, else however much it is currently carrying (so a rung holding + // everything in flight reads as full). + const capacity = + status.policy.maxWorkers ?? Math.max(status.runner.inFlight.length, 1); + return ( - <div className="flex flex-col gap-3 border border-border rounded p-3"> - <div className="flex flex-wrap items-center gap-4"> + <div className="flex flex-col gap-3"> + <div className="flex flex-wrap items-center gap-x-5 gap-y-2"> <label className="flex items-center gap-2 text-sm font-medium"> <input type="checkbox" - checked={enabled} - onChange={(e) => setEnabled(e.target.checked)} + checked={form.enabled} + onChange={(e) => + setForm((f) => ({ ...f, enabled: e.target.checked })) + } /> - Enable auto-{kind === "transcription" ? "transcribe" : "download"} + Enable auto-{kindWord} + </label> + <label className="flex items-center gap-2 text-sm"> + <span className="text-muted-foreground">Order</span> + <select + aria-label={`video order for auto-${kindWord}`} + value={form.order} + onChange={(e) => + setForm((f) => ({ + ...f, + order: e.target.value as AutoQueueOrder, + })) + } + className={SELECT_CLASS} + > + {AUTO_QUEUE_ORDERS.map((o) => ( + <option key={o} value={o}> + {ORDER_LABEL[o]} + </option> + ))} + </select> </label> <label className="flex items-center gap-2 text-sm"> - <span>Runner max workers</span> + <span className="text-muted-foreground">Runner max workers</span> <input type="number" min={1} - value={maxWorkers ?? ""} + value={form.maxWorkers ?? ""} placeholder="∞" onChange={(e) => - setMaxWorkers( - e.target.value.trim() === "" - ? null - : Math.max(1, Number.parseInt(e.target.value, 10) || 1), - ) + setForm((f) => ({ + ...f, + maxWorkers: + e.target.value.trim() === "" + ? null + : Math.max(1, Number.parseInt(e.target.value, 10) || 1), + })) } - className={`${inputClass} w-20`} + className={`${NUM_CLASS} w-20`} /> </label> </div> + <p className="text-xs text-muted-foreground"> + {ORDER_HINT} + {tradeoff ? ` ${tradeoff}` : ""} + </p> + + <ClaimLadder + root={form.root} + channels={channels} + platforms={platforms} + buckets={buckets} + ops={ops} + data={{ + pendingByLeaf: status.pendingByLeaf, + pendingHeadByLeaf: status.pendingHeadByLeaf, + ownerByVideo: status.ownerByVideo, + recencyByVideo: status.recencyByVideo, + activeByNode: status.runner.activeByNode, + capacity, + nextUpLeafId: status.nextUp?.leafId ?? null, + leafIds: leafOrder(form.root).map((l) => l.id), + }} + /> + {/* Lowest-priority lane, off by default. Appending the opt-in bucket to the TAIL of the default union is what makes it lowest priority: pending work is claimed bucket-by-bucket in list order. */} <label className="flex items-start gap-2 text-sm"> <input type="checkbox" - checked={replaceAutoSubs} - onChange={(e) => setReplaceAutoSubs(e.target.checked)} + checked={form.replaceAutoSubs} + onChange={(e) => + setForm((f) => ({ ...f, replaceAutoSubs: e.target.checked })) + } aria-label={`replace YouTube auto-captions for auto-${kind}`} className="mt-1" /> @@ -229,336 +330,22 @@ export function PolicyTreeEditor({ ? "transcribe videos whose only transcript is YouTube's speech recognition (and whose audio is already downloaded)" : "download audio for videos whose only transcript is YouTube's speech recognition"} . Strictly lowest priority, and off by default — every candidate - costs a download plus a transcription. Rules below can also target + costs a download plus a transcription. Rules above can also target the bucket directly for per-channel opt-in. </span> </span> </label> - <NodeEditor - node={root} - depth={0} - parentId={null} - parentMode={null} - index={0} - siblingCount={1} - channels={channels} - platforms={platforms} - buckets={buckets} - update={update} - remove={remove} - addChild={addChild} - move={move} - moveToTop={moveToTop} + <SaveBar + dirty={dirty} + saving={saving} + result={result} + onSave={save} + onDiscard={() => { + setForm(baseline); + setResult(null); + }} /> - - <div className="flex items-center gap-3"> - <button - type="button" - onClick={save} - disabled={saving} - className="px-3 py-1.5 rounded-md bg-primary text-primary-foreground text-sm font-medium hover:opacity-90 disabled:opacity-50" - > - {saving ? "Saving…" : "Save policy"} - </button> - {result?.ok === true && ( - <span role="status" className="text-sm text-success"> - Saved. - </span> - )} - {result?.ok === false && ( - <span role="alert" className="text-sm text-destructive"> - {result.error} - </span> - )} - </div> - </div> - ); -} - -type EditorProps = { - node: AutoQueueNode; - depth: number; - parentId: string | null; - parentMode: AutoQueueMode | null; - index: number; - siblingCount: number; - channels: { slug: string; name: string | null }[]; - platforms: string[]; - buckets: string[]; - update: (id: string, fn: (n: AutoQueueNode) => AutoQueueNode) => void; - remove: (id: string) => void; - addChild: (parentId: string, child: AutoQueueNode) => void; - move: (parentId: string, index: number, dir: -1 | 1) => void; - moveToTop: (parentId: string, index: number) => void; -}; - -function NodeEditor(props: EditorProps) { - const { node, depth, parentId, parentMode, index, siblingCount } = props; - const group = isGroup(node); - return ( - <div - className="flex flex-col gap-2 rounded border border-border p-2" - style={{ marginLeft: depth > 0 ? 12 : 0 }} - > - <div className="flex flex-wrap items-center gap-2"> - <span className="text-xs font-semibold text-muted-foreground"> - {group ? "Group" : "Rule"} - </span> - - {group ? ( - <ModeSelect node={node} update={props.update} /> - ) : ( - <LeafControls {...props} leaf={node} /> - )} - - {/* weight, shown when the PARENT is weighted-fair */} - {parentMode === "weighted-fair" && ( - <label className="flex items-center gap-1 text-xs text-muted-foreground"> - weight - <input - type="number" - min={1} - value={node.weight ?? 1} - onChange={(e) => - props.update(node.id, (n) => ({ - ...n, - weight: Math.max(1, Number.parseInt(e.target.value, 10) || 1), - })) - } - className={`${inputClass} w-16`} - /> - </label> - )} - - <label className="flex items-center gap-1 text-xs text-muted-foreground"> - max - <input - type="number" - min={1} - placeholder="∞" - value={node.maxWorkers ?? ""} - onChange={(e) => - props.update(node.id, (n) => ({ - ...n, - maxWorkers: - e.target.value.trim() === "" - ? null - : Math.max(1, Number.parseInt(e.target.value, 10) || 1), - })) - } - className={`${inputClass} w-16`} - /> - </label> - - <span className="ml-auto flex items-center gap-1"> - {parentId && siblingCount > 1 && ( - <> - <button - type="button" - className={btnClass} - disabled={index === 0} - aria-label="Move to top" - title="Move to top" - onClick={() => props.moveToTop(parentId, index)} - > - ⤒ - </button> - <button - type="button" - className={btnClass} - disabled={index === 0} - onClick={() => props.move(parentId, index, -1)} - > - ↑ - </button> - <button - type="button" - className={btnClass} - disabled={index === siblingCount - 1} - onClick={() => props.move(parentId, index, 1)} - > - ↓ - </button> - </> - )} - {parentId && ( - <button - type="button" - className={`${btnClass} text-destructive`} - onClick={() => props.remove(node.id)} - > - Remove - </button> - )} - </span> - </div> - - {group && ( - <> - <div className="flex flex-col gap-2"> - {(node as AutoQueueGroup).children.map((child, i) => ( - <NodeEditor - key={child.id} - {...props} - node={child} - depth={depth + 1} - parentId={node.id} - parentMode={(node as AutoQueueGroup).mode} - index={i} - siblingCount={(node as AutoQueueGroup).children.length} - /> - ))} - {(node as AutoQueueGroup).children.length === 0 && ( - <p className="text-xs text-muted-foreground pl-2"> - Empty group — add a rule or nested group below. - </p> - )} - </div> - <div className="flex flex-wrap gap-2"> - <button - type="button" - className={btnClass} - onClick={() => props.addChild(node.id, makeLeaf("channel"))} - > - + Channel rule - </button> - <button - type="button" - className={btnClass} - onClick={() => props.addChild(node.id, makeLeaf("platform"))} - > - + Platform rule - </button> - <button - type="button" - className={btnClass} - onClick={() => props.addChild(node.id, makeLeaf("all"))} - > - + Catch-all rule - </button> - <button - type="button" - className={btnClass} - onClick={() => props.addChild(node.id, makeGroup())} - > - + Group - </button> - </div> - </> - )} </div> ); } - -function ModeSelect({ - node, - update, -}: { - node: AutoQueueNode; - update: (id: string, fn: (n: AutoQueueNode) => AutoQueueNode) => void; -}) { - return ( - <select - value={(node as AutoQueueGroup).mode} - onChange={(e) => - update(node.id, (n) => ({ ...n, mode: e.target.value as AutoQueueMode })) - } - className={inputClass} - > - {AUTO_QUEUE_MODES.map((m) => ( - <option key={m} value={m}> - {m} - </option> - ))} - </select> - ); -} - -function LeafControls(props: EditorProps & { leaf: AutoQueueLeaf }) { - const { leaf, channels, platforms, buckets, update } = props; - const type = leaf.match.type; - return ( - <> - <select - value={type} - onChange={(e) => - update(leaf.id, (n) => ({ - ...n, - match: { type: e.target.value as AutoQueueLeaf["match"]["type"] }, - })) - } - className={inputClass} - > - <option value="channel">channel</option> - <option value="platform">platform</option> - <option value="all">all</option> - </select> - - {type === "channel" && ( - <select - aria-label="channel rule value" - value={leaf.match.value ?? ""} - onChange={(e) => - update(leaf.id, (n) => ({ - ...n, - match: { ...(n as AutoQueueLeaf).match, value: e.target.value }, - })) - } - className={inputClass} - > - <option value="">— pick channel —</option> - {channels.map((c) => ( - <option key={c.slug} value={c.slug}> - {c.name ?? c.slug} - </option> - ))} - </select> - )} - - {type === "platform" && ( - <select - aria-label="platform rule value" - value={leaf.match.value ?? ""} - onChange={(e) => - update(leaf.id, (n) => ({ - ...n, - match: { ...(n as AutoQueueLeaf).match, value: e.target.value }, - })) - } - className={inputClass} - > - <option value="">— pick platform —</option> - {platforms.map((p) => ( - <option key={p} value={p}> - {p} - </option> - ))} - </select> - )} - - <label className="flex items-center gap-1 text-xs text-muted-foreground"> - bucket - <select - value={leaf.match.bucket ?? ""} - onChange={(e) => - update(leaf.id, (n) => { - const match = { ...(n as AutoQueueLeaf).match }; - if (e.target.value) match.bucket = e.target.value; - else delete match.bucket; - return { ...n, match }; - }) - } - className={inputClass} - > - <option value="">all buckets (default)</option> - {buckets.map((b) => ( - <option key={b} value={b}> - {b} - </option> - ))} - </select> - </label> - </> - ); -} diff --git a/editor/app/auto-queue/components/SaveBar.tsx b/editor/app/auto-queue/components/SaveBar.tsx @@ -0,0 +1,81 @@ +"use client"; + +import { useEffect } from "react"; +import { Button } from "yt-dlp-transcript-common/components/ui/button"; +import type { SaveResult } from "../actions"; + +// Save, and the thing this page never had: a signal that there is something to +// save. Editing the tree used to change nothing visible, so navigating away +// silently discarded the work — the classic "I clicked Move to top and it did +// nothing" report. +// +// The Save button itself is ALWAYS rendered, never gated on dirtiness. That is +// deliberate: three e2e specs click "Save policy" by name, and a button that +// appears only under a condition is a button that can vanish under a race. The +// dirty state is carried by the rail, the wording and the Discard button beside +// it, which is where it belongs anyway. + +export function SaveBar({ + dirty, + saving, + result, + onSave, + onDiscard, +}: { + dirty: boolean; + saving: boolean; + result: SaveResult | null; + onSave: () => void; + onDiscard: () => void; +}) { + // The browser-level guard. Only armed while there is something to lose. + useEffect(() => { + if (!dirty) return; + const onBeforeUnload = (e: BeforeUnloadEvent) => e.preventDefault(); + window.addEventListener("beforeunload", onBeforeUnload); + return () => window.removeEventListener("beforeunload", onBeforeUnload); + }, [dirty]); + + return ( + <div + className={`flex flex-wrap items-center gap-3 rounded-md border-l-2 px-3 py-2 ${ + dirty + ? "border-l-warning bg-warning-soft/40" + : "border-l-transparent bg-transparent" + }`} + > + {dirty && ( + <span className="font-mono text-xs uppercase tracking-[0.14em] text-warning"> + Unsaved changes + </span> + )} + <Button type="button" size="sm" onClick={onSave} disabled={saving}> + {saving ? "Saving…" : "Save policy"} + </Button> + {dirty && ( + <Button + type="button" + size="sm" + variant="outline" + onClick={onDiscard} + disabled={saving} + > + Discard changes + </Button> + )} + {/* The ONLY role="status" on this page, twice over (one per lane). A spec + asserts getByRole("status").first() reads exactly "Saved.", so any + other live region here would break it. */} + {result?.ok === true && ( + <span role="status" className="text-sm text-success"> + Saved. + </span> + )} + {result?.ok === false && ( + <span role="alert" className="text-sm text-destructive"> + {result.error} + </span> + )} + </div> + ); +} diff --git a/editor/app/auto-queue/components/SnoozeControl.tsx b/editor/app/auto-queue/components/SnoozeControl.tsx @@ -0,0 +1,92 @@ +"use client"; + +import { useState } from "react"; +import { Button } from "yt-dlp-transcript-common/components/ui/button"; +import type { AutoQueueKind } from "yt-dlp-transcript-common/jobs/autoQueueState"; +import { snoozeAutoQueueAction } from "../actions"; +import { snoozePresets } from "./dispatch"; + +// Idle a runner for a while without stopping it. +// +// The distinction matters operationally: stopping loses the in-flight work and +// needs someone to remember to start it again, whereas a snooze leaves the loop +// up, re-reading settings every iteration, so it resumes by itself. And because +// the deadline is persisted in settings.json rather than held in runner memory, +// it survives a server restart — which a "pause" flag in the process would not. + +export function SnoozeControl({ + kind, + title, + snoozeUntil, + now, + onChanged, +}: { + kind: AutoQueueKind; + title: string; + snoozeUntil: number | null; + now: number | null; + onChanged: () => Promise<void>; +}) { + const [busy, setBusy] = useState(false); + const snoozed = snoozeUntil !== null && (now === null || snoozeUntil > now); + + const set = async (until: number | null) => { + setBusy(true); + try { + await snoozeAutoQueueAction(kind, until); + await onChanged(); + } finally { + setBusy(false); + } + }; + + if (snoozed) { + return ( + <div className="flex flex-wrap items-center gap-3 rounded-md border border-warning/30 bg-warning-soft px-3 py-2 text-sm"> + <span className="text-warning"> + Snoozed until{" "} + <span className="tabular-nums"> + {new Date(snoozeUntil).toLocaleString([], { + hour: "2-digit", + minute: "2-digit", + month: "short", + day: "numeric", + })} + </span> + {" — the runner stays up and resumes by itself."} + </span> + <Button + type="button" + size="xs" + variant="outline" + aria-label={`Wake ${title} now`} + disabled={busy} + onClick={() => set(null)} + > + Wake now + </Button> + </div> + ); + } + + return ( + <div className="flex flex-wrap items-center gap-2 text-sm text-muted-foreground"> + <span className="font-mono text-xs uppercase tracking-[0.14em]"> + Snooze + </span> + {snoozePresets(now ?? Date.now()).map((p) => ( + <Button + key={p.label} + type="button" + size="xs" + variant="outline" + aria-label={`Snooze ${title} for ${p.label}`} + disabled={busy || now === null} + onClick={() => set(p.until)} + > + {p.label} + </Button> + ))} + </div> + ); +} diff --git a/editor/app/auto-queue/components/dispatch.ts b/editor/app/auto-queue/components/dispatch.ts @@ -0,0 +1,194 @@ +import type { + AutoQueueGroup, + AutoQueueLeaf, + AutoQueueNode, + AutoQueueOrder, +} from "yt-dlp-transcript-common/jobs/autoQueuePolicy"; +import { isGroup } from "yt-dlp-transcript-common/jobs/autoQueuePolicy"; +import type { AutoRunnerIdleReason } from "yt-dlp-transcript-common/controller/autoRunner"; +import type { AutoQueueKind } from "yt-dlp-transcript-common/jobs/autoQueueState"; + +// Shared vocabulary for the dispatcher board. Everything here turns a wire value +// into the words an operator would use, in one place, so the deck, the ladder +// and the next-up line cannot describe the same state three different ways. + +export type Channel = { slug: string; name: string | null }; + +// --- Rules as sentences ------------------------------------------------------ + +// What a bucket actually means, in words. Deliberately NOT used for the bucket +// <select>'s option labels: an e2e spec asserts exactly one option named +// `downloadedAutoSubsOnly` page-wide, and the picker is also how an operator +// names a bucket to the runner — so the raw name stays there and the gloss lives +// in the rung's sentence. +const BUCKET_GLOSS: Record<string, string> = { + downloadedNoTranscript: "downloaded, not transcribed", + failedListed: "retries only", + partialDownloads: "partial downloads", + undownloadedIds: "not yet downloaded", + downloadedAutoSubsOnly: "YouTube captions, audio in hand", + autoSubsOnly: "YouTube captions", +}; + +export function bucketGloss(bucket: string | undefined): string { + if (!bucket) return "all buckets"; + return BUCKET_GLOSS[bucket] ?? bucket; +} + +// A leaf, read as a sentence rather than as a record: "cornbreadman · retries +// only" instead of "channel: cornbreadman [failedListed]". +export function leafSentence(leaf: AutoQueueLeaf, channels: Channel[]): string { + const m = leaf.match; + let who: string; + if (m.type === "all") who = "any channel"; + else if (m.type === "channel") { + const found = channels.find((c) => c.slug === m.value); + who = m.value ? (found?.name ?? m.value) : "no channel picked"; + } else who = m.value ? `${m.value} channels` : "no platform picked"; + return `${who} · ${bucketGloss(m.bucket)}`; +} + +// Plain-language gloss for a group's competition mode. The mode word alone +// ("strict") tells an operator nothing about what it does to their queue. +export const MODE_GLOSS: Record<string, string> = { + strict: "first rule with work wins", + "round-robin": "take turns between rules", + "weighted-fair": "share by weight", +}; + +// Pre-order walk, matching flattenLeaves — the ORDINAL a rung shows is its +// position in this list, which is genuinely the priority order the engine reads. +export function leafOrder(root: AutoQueueNode): AutoQueueLeaf[] { + if (!isGroup(root)) return [root]; + const out: AutoQueueLeaf[] = []; + for (const child of (root as AutoQueueGroup).children) { + out.push(...leafOrder(child)); + } + return out; +} + +// --- Ordering ---------------------------------------------------------------- + +export const ORDER_LABEL: Record<AutoQueueOrder, string> = { + listed: "Listed order", + newest: "Newest first", + oldest: "Oldest first", +}; + +export const ORDER_HINT = + "Newest first orders the videos inside each rule. Rule order still wins; " + + "for a pure newest-first archive, use one catch-all rule."; + +export function orderTradeoff( + kind: AutoQueueKind, + order: AutoQueueOrder, +): string | null { + if (order === "listed") return null; + return kind === "download" + ? "Partial downloads lose their head start, so a half-finished download can wait behind fresh work." + : "Retries lose their head start, so a failed video can wait behind fresh work."; +} + +// --- Idle reasons ------------------------------------------------------------ + +// The answer to "it's running, why isn't it doing anything?". Kept out of the +// components so the deck and the lane say the same sentence. +export function idleReasonText( + reason: AutoRunnerIdleReason | null, + kind: AutoQueueKind, +): string | null { + const what = kind === "transcription" ? "transcribe" : "download"; + switch (reason) { + case "no-pending": + return "nothing pending — every rule is empty"; + case "capped": + return "work exists, but every route to it is at a worker cap"; + case "cooldown": + return "every pending platform is in a rate-limit cooldown"; + case "no-workers": + return "no enabled worker to run it"; + case "disk-gate": + return "disk gate closed — not enough free space"; + case "downloads-paused": + return "downloads are paused globally"; + case "snoozed": + return "snoozed"; + case "disabled": + return `auto-${what} was switched off`; + case "stopped": + case null: + return null; + } +} + +// --- Formatting -------------------------------------------------------------- + +export function formatClock(ms: number): string { + if (!ms) return "—"; + return new Date(ms).toLocaleTimeString([], { + hour: "2-digit", + minute: "2-digit", + }); +} + +// Compact elapsed time: 42s, 3m21s, 14m, 2h06m. Used for uptime and per-unit +// age, where a full duration string would swamp the line it sits on. +export function formatElapsed(ms: number): string { + const s = Math.max(0, Math.floor(ms / 1000)); + if (s < 60) return `${s}s`; + const m = Math.floor(s / 60); + if (m < 60) { + const rem = s % 60; + return rem === 0 ? `${m}m` : `${m}m${String(rem).padStart(2, "0")}s`; + } + const h = Math.floor(m / 60); + return `${h}h${String(m % 60).padStart(2, "0")}m`; +} + +// A recency key as a date, marked when it was estimated rather than read. +export function formatRecency( + key: { key: string; estimated: boolean } | undefined, +): string | null { + if (!key || !key.key) return null; + // The "newer than anything we know" sentinel — there is no date to show. + if (!/^\d{8}$/.test(key.key)) return key.estimated ? "≈ newest" : null; + const iso = `${key.key.slice(0, 4)}-${key.key.slice(4, 6)}-${key.key.slice(6)}`; + return key.estimated ? `≈${iso}` : iso; +} + +export function formatCooldown(secs: number): string { + if (secs < 60) return `${secs}s`; + const m = Math.floor(secs / 60); + const s = secs % 60; + return s === 0 ? `${m}m` : `${m}m ${s}s`; +} + +// Snooze presets, resolved at click time so "until tomorrow" means the next +// local 09:00 rather than a fixed offset. +export function snoozePresets(now: number): { label: string; until: number }[] { + const tomorrow = new Date(now); + tomorrow.setDate(tomorrow.getDate() + 1); + tomorrow.setHours(9, 0, 0, 0); + return [ + { label: "1 hour", until: now + 3_600_000 }, + { label: "4 hours", until: now + 4 * 3_600_000 }, + { label: "Until tomorrow", until: tomorrow.getTime() }, + ]; +} + +// --- Control classes --------------------------------------------------------- + +// The policy controls stay NATIVE <select> / <input type=checkbox> rather than +// the kit's Radix wrappers, and that is a deliberate, load-bearing choice: +// * `selectOption("alpha")` in the suite requires a real <select>; +// * `getByRole("option", { name: "downloadedAutoSubsOnly" })` must find exactly +// one option page-wide — Radix's items are portalled and unmounted while the +// menu is closed, so that count would be zero; +// * `getByLabel(/Enable auto-transcribe/).check()` is unambiguous on a real +// checkbox input. +// They are token-styled to sit with the rest of the kit, so only the mechanism +// differs, not the look. +export const SELECT_CLASS = + "rounded border border-border bg-background px-2 py-1 text-sm text-foreground"; +export const NUM_CLASS = + "rounded border border-border bg-background px-2 py-1 text-sm text-foreground tabular-nums"; diff --git a/editor/app/auto-queue/page.tsx b/editor/app/auto-queue/page.tsx @@ -6,6 +6,7 @@ import { PLATFORM_VALUES } from "yt-dlp-transcript-common/lib/platform"; import { selectableBucketsForKind } from "yt-dlp-transcript-common/jobs/autoQueuePolicy"; import { buildAutoQueueStatusPayload } from "./status"; import { AutoQueueView } from "./components/AutoQueueView"; +import { HowPriorityWorks } from "./components/HowPriorityWorks"; export const dynamic = "force-dynamic"; @@ -33,25 +34,18 @@ export default async function AutoQueuePage() { return ( <div className="flex flex-col gap-4"> <div className="flex items-center justify-between"> - <h1 className="text-2xl font-semibold">Auto-queue</h1> + <h1 className="font-display text-2xl font-semibold tracking-tight"> + Auto-queue + </h1> <Link href="/scheduler" className="text-sm underline"> Sync schedule </Link> </div> <p className="text-sm text-muted-foreground"> - Automatically pick the next transcription across all channels by a - priority policy, instead of running one channel batch at a time. Order the - rules top-to-bottom for strict priority, or wrap rules in a group set to{" "} - <em>round-robin</em> / <em>weighted-fair</em> to alternate between them. - The highest-priority channel with available work claims the next freed - worker slot; when its work runs out the runner falls back to the next - rule automatically. This is independent of the{" "} - <Link href="/scheduler" className="underline"> - sync schedule - </Link>{" "} - (which only decides when to re-fetch each channel) and manual batches keep - working alongside it. + Decide, continuously and unattended, which video across every channel + gets worked on next. </p> + <HowPriorityWorks /> <AutoQueueView initial={initial} channels={channelOptions} diff --git a/editor/app/auto-queue/status.ts b/editor/app/auto-queue/status.ts @@ -2,7 +2,9 @@ import { getPaths } from "yt-dlp-transcript-common/lib/paths"; import { getSettings } from "yt-dlp-transcript-common/lib/settings"; import { type AutoRunnerStatus, - computeLeafPendingCounts, + type NextUpView, + type RecencyKeyView, + computeLeafPending, getAutoRunnerStatus, } from "yt-dlp-transcript-common/controller/autoRunner"; import { @@ -32,6 +34,18 @@ export type AutoQueueKindStatus = { policy: AutoQueuePolicy; runner: AutoRunnerStatus; pendingByLeaf: Record<string, number>; + // The first few pending ids of each rule, in the order the runner would hand + // them out — the drill-down behind each rung's count, and the only way to + // check an ordering change by eye. + pendingHeadByLeaf: Record<string, string[]>; + // videoId -> owning channel slug, for the drill-down's links. Covers exactly + // the ids in pendingHeadByLeaf (plus nextUp's). + ownerByVideo: Record<string, string>; + // videoId -> recency sort key, for the same ids, when the policy orders by + // recency. Absent entries just render without a date. + recencyByVideo: Record<string, RecencyKeyView>; + // What the policy would dispatch next, and which rules it passed over. + nextUp: NextUpView | null; picks: AutoQueuePick[]; // Platforms paused by a rate-limit/network backoff (download kind only). cooldowns: PlatformCooldownView[]; @@ -47,7 +61,7 @@ async function buildKind(kind: AutoQueueKind): Promise<AutoQueueKindStatus> { const policy = getSettings().autoQueue[kind]; const runner = getAutoRunnerStatus(kind); const state = await readAutoQueueState(paths); - const pendingByLeaf = await computeLeafPendingCounts(kind, paths); + const pending = await computeLeafPending(kind, paths); const now = Date.now(); // Surface platforms still inside their cooldown window so the operator can see // why an otherwise-pending platform isn't being serviced (and a manual sync on @@ -62,7 +76,11 @@ async function buildKind(kind: AutoQueueKind): Promise<AutoQueueKindStatus> { kind, policy, runner, - pendingByLeaf, + pendingByLeaf: pending.counts, + pendingHeadByLeaf: pending.head, + ownerByVideo: pending.owner, + recencyByVideo: pending.recency, + nextUp: pending.nextUp, picks: state[kind].picks, cooldowns, }; diff --git a/editor/e2e/auto-queue.spec.ts b/editor/e2e/auto-queue.spec.ts @@ -169,8 +169,15 @@ async function stopRunner( } type KindStatus = { - runner: { running: boolean; jobId: string | null }; + runner: { + running: boolean; + jobId: string | null; + idleReason: string | null; + }; + policy: { order?: string; snoozeUntil?: number | null }; pendingByLeaf: Record<string, number>; + pendingHeadByLeaf: Record<string, string[]>; + nextUp: { videoId: string; leafId: string } | null; picks: { leafId: string; videoId: string; channelSlug: string }[]; }; type Status = { transcription: KindStatus; download: KindStatus }; @@ -356,6 +363,7 @@ test("UI: build a policy in the editor, save it, and start the runner", async ({ has: page.getByRole("heading", { name: "Auto-transcribe" }), }); await expect(section.getByText("Runner stopped")).toBeVisible(); + await awaitHydration(section); // Add a channel rule and point it at alpha. await section.getByRole("button", { name: "+ Channel rule" }).click(); @@ -411,6 +419,8 @@ test("UI: Move to top jumps a rule to the front of its siblings", async ({ has: page.getByRole("heading", { name: "Auto-transcribe" }), }); + await awaitHydration(section); + // Add three channel rules pointing at alpha, beta, gamma (in that order). for (const slug of ["alpha", "beta", "gamma"]) { await section.getByRole("button", { name: "+ Channel rule" }).click(); @@ -752,3 +762,268 @@ test("download: prioritizes channels across the per-platform queue", async ({ "b2", ]); }); + +// A React state update that lands before hydration is discarded silently, so a +// select or a click made too early leaves the page looking changed while the +// component still holds the old value — and the next save writes the old value. +// The lane stamps data-hydrated once its client mount effect has run. +async function awaitHydration( + section: import("@playwright/test").Locator, +): Promise<void> { + await expect(section).toHaveAttribute("data-hydrated", "true", { + timeout: 20_000, + }); +} + +// --- Ordering within a rule (policy.order) ---------------------------------- +// +// The fixtures here write snapshot.json directly and stand up no LMDB index and +// no metadata.info.json, so the recency key comes from the LAST layer of +// common/controller/recencyIndex.ts: the YYYYMMDD_ prefix on the video dir name. +// That is deliberate — it exercises the ordering end to end (policy → runner → +// pick order) without needing a real index in the test corpus. + +test("UI: the order select round-trips through save", async ({ page }) => { + await resetData(null); + await makeChannel("alpha", []); + await writeSettings({ + adminTitle: "Test Admin", + maxTranscriptPageBytes: 8388608, + sleepBetweenDownloadsSeconds: 0, + minFreeDiskGB: 0, + workers: ONE_WORKER, + autoQueue: { + transcription: { + enabled: false, + maxWorkers: 1, + root: { id: "root", mode: "strict", children: [] }, + }, + download: {}, + }, + }); + + await page.goto("/auto-queue"); + const section = page.locator("section", { + has: page.getByRole("heading", { name: "Auto-transcribe" }), + }); + + await awaitHydration(section); + const order = section.getByLabel("video order for auto-transcribe"); + await expect(order).toHaveValue("listed"); + await order.selectOption("newest"); + // Proves the change reached React state, not just the DOM node — and covers + // the new unsaved-changes marker on the way past. + await expect(section.getByText("Unsaved changes")).toBeVisible(); + await section.getByRole("button", { name: "Save policy" }).click(); + await expect(section.getByText("Saved.")).toBeVisible(); + + await expect + .poll( + async () => { + const s = await readJson<{ + autoQueue?: { transcription?: { order?: string } }; + }>("test-settings.json").catch(() => null); + return s?.autoQueue?.transcription?.order ?? null; + }, + { timeout: 10_000 }, + ) + .toBe("newest"); + + // And it comes back that way, rather than being silently dropped by the + // action's explicit field list on the way through. + await page.reload(); + await expect( + page + .locator("section", { + has: page.getByRole("heading", { name: "Auto-transcribe" }), + }) + .getByLabel("video order for auto-transcribe"), + ).toHaveValue("newest"); +}); + +test("newest first: the freshest upload jumps the older backlog", async ({ + request, +}) => { + await resetData(null); + // Deliberately created oldest-first, which is also how they sort + // lexicographically — so "listed" order would serve them in exactly the wrong + // order and any pass here has to come from the recency key. + await makeChannel("alpha", [ + "20240101_old", + "20250601_mid", + "20260812_new", + ]); + + const root: Group = { + id: "root", + mode: "strict", + children: [{ id: "leaf-alpha", match: { type: "channel", value: "alpha" } }], + }; + await writeSettings({ + adminTitle: "Test Admin", + maxTranscriptPageBytes: 8388608, + sleepBetweenDownloadsSeconds: 0, + minFreeDiskGB: 0, + workers: ONE_WORKER, + autoQueue: { + transcription: { enabled: true, maxWorkers: 1, order: "newest", root }, + download: {}, + }, + }); + + await startRunner(request); + await expect + .poll(async () => (await getStatus(request)).transcription.picks.length, { + timeout: 60_000, + }) + .toBe(3); + + expect(pickOrder(await getStatus(request))).toEqual([ + "20260812_new", + "20250601_mid", + "20240101_old", + ]); +}); + +test("listed order is unchanged — the same fixture serves oldest first", async ({ + request, +}) => { + // The compatibility half of the claim above: with no `order` field at all (an + // existing settings.json), the identical fixture must still drain in the + // order it always did. + await resetData(null); + await makeChannel("alpha", [ + "20240101_old", + "20250601_mid", + "20260812_new", + ]); + + const root: Group = { + id: "root", + mode: "strict", + children: [{ id: "leaf-alpha", match: { type: "channel", value: "alpha" } }], + }; + await writeSettings({ + adminTitle: "Test Admin", + maxTranscriptPageBytes: 8388608, + sleepBetweenDownloadsSeconds: 0, + minFreeDiskGB: 0, + workers: ONE_WORKER, + autoQueue: transcriptionAutoQueue(root, 1), + }); + + await startRunner(request); + await expect + .poll(async () => (await getStatus(request)).transcription.picks.length, { + timeout: 60_000, + }) + .toBe(3); + + expect(pickOrder(await getStatus(request))).toEqual([ + "20240101_old", + "20250601_mid", + "20260812_new", + ]); +}); + +// --- "It's running — why isn't it doing anything?" --------------------------- + +test("a running runner says why it is idle", async ({ page, request }) => { + await resetData(null); + // A channel with nothing pending: the runner comes up, finds no work, and + // used to sit there looking identical to one that was wedged. + await makeChannel("alpha", []); + + const root: Group = { + id: "root", + mode: "strict", + children: [{ id: "leaf-alpha", match: { type: "channel", value: "alpha" } }], + }; + await writeSettings({ + adminTitle: "Test Admin", + maxTranscriptPageBytes: 8388608, + sleepBetweenDownloadsSeconds: 0, + minFreeDiskGB: 0, + workers: ONE_WORKER, + autoQueue: transcriptionAutoQueue(root, 1), + }); + + await startRunner(request); + await expect + .poll( + async () => (await getStatus(request)).transcription.runner.idleReason, + { timeout: 20_000 }, + ) + .toBe("no-pending"); + + await page.goto("/auto-queue"); + const section = page.locator("section", { + has: page.getByRole("heading", { name: "Auto-transcribe" }), + }); + await expect(section.getByText(/nothing pending/)).toBeVisible({ + timeout: 20_000, + }); +}); + +// --- Snooze ----------------------------------------------------------------- + +test("snooze idles the runner without stopping it, and Wake now resumes", async ({ + page, + request, +}) => { + await resetData(null); + await makeChannel("alpha", ["a1"]); + + const root: Group = { + id: "root", + mode: "strict", + children: [{ id: "leaf-alpha", match: { type: "channel", value: "alpha" } }], + }; + await writeSettings({ + adminTitle: "Test Admin", + maxTranscriptPageBytes: 8388608, + sleepBetweenDownloadsSeconds: 0, + minFreeDiskGB: 0, + workers: ONE_WORKER, + autoQueue: { + transcription: { + enabled: true, + maxWorkers: 1, + snoozeUntil: Date.now() + 3_600_000, + root, + }, + download: {}, + }, + }); + + await startRunner(request); + + // IDLE, not stopped: the loop stays up so it can notice the deadline pass. + await expect + .poll( + async () => { + const s = (await getStatus(request)).transcription; + return { running: s.runner.running, idle: s.runner.idleReason }; + }, + { timeout: 20_000 }, + ) + .toEqual({ running: true, idle: "snoozed" }); + expect(await allTranscribed("alpha", ["a1"])).toBe(false); + + // Waking is a live settings change, not a restart: the same runner picks the + // work up on its next iteration. + await page.goto("/auto-queue"); + const section = page.locator("section", { + has: page.getByRole("heading", { name: "Auto-transcribe" }), + }); + await expect(section.getByText(/Snoozed until/)).toBeVisible({ + timeout: 20_000, + }); + await awaitHydration(section); + await section.getByRole("button", { name: "Wake Auto-transcribe now" }).click(); + + await expect + .poll(async () => allTranscribed("alpha", ["a1"]), { timeout: 60_000 }) + .toBe(true); + expect((await getStatus(request)).transcription.runner.running).toBe(true); +});