commit 0a37f7cacabb46b59f92b900fbe44883d4d0d058
parent 7cfcab884b11e6c0696d4f03d1caaa3254914a36
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Fri, 26 Jun 2026 15:59:23 -0400
Phase 8: one-click retry + retry-all for failed jobs
Granular control (priority feature #2): re-run a failed job from its stored
spec, jumping ahead of other queued work.
- retryJobAction(id): resolve the spec from the live registry or the on-disk
meta sidecar (so an evicted/archived failure is still retryable), re-run via
runJobSpec (bucket jobs re-derive from current state), then promote the new
job to the front of its queue (reusing Phase 7's promote — same observable
"runs next" as an urgent tier, without threading tier through every action).
- retryAllFailedAction(): re-run every listed failed job that has a spec.
- RetryJobButton on each failed+bookmarkable row in the Jobs table; a
RetryAllFailedButton in the page header when any such job exists. Kinds with
no replay descriptor offer no Retry.
Also adds CHANGELOG entries for retry, reorder/promote (Phase 7), and the
foundational scheduler/runPool/jobKinds refactor.
Verification: jobs-retry e2e (retry re-runs from the sidecar-staged failure;
spec-less failures show no Retry) + jobs-filters/jobs/bookmarks green; typecheck
(common+editor) clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Diffstat:
7 files changed, 234 insertions(+), 1 deletion(-)
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -1,6 +1,9 @@
# Changelog
## [Unreleased]
+- **One-click Retry for failed jobs (plus "Retry all failed").** A failed job that carries a replay descriptor (any bookmarkable kind — sync, download-missing, transcribe-all, retry-bucket, …) now shows a **Retry** button on the Jobs history table, and the page header gains a **Retry all failed** button whenever at least one such job is listed. Retry re-runs the job from its stored spec exactly like a bookmark re-run (so bucket jobs re-derive from the channel's *current* state), and the re-run **jumps ahead of other queued work** (it's promoted to the front of its queue, reusing the new reorder machinery) so a fix-and-retry runs next rather than at the back of the line. The spec is resolved from the live registry or, for an evicted/archived job, from its on-disk `<id>.meta.json` sidecar — so even a failure the 100-job cap has dropped is still retryable. Kinds with no replay descriptor (e.g. `import-one`) intentionally offer no Retry. See `editor/app/jobs/actions.ts` (`retryJobAction` / `retryAllFailedAction`), the new `RetryJobButton` / `RetryAllFailedButton`, and `editor/e2e/jobs-retry.spec.ts`.
+- **Reorder and promote queued jobs from Active Jobs.** A queued job's row now carries **Promote / ↑ / ↓** controls (mirroring the auto-queue policy editor's move buttons) to change its order within its queue — Promote sends it to the front so it runs next, ↑/↓ nudge it one slot. Only actionable moves render (the first-queued job shows no up/promote, the last no down), and a running job is never displaced. See `editor/app/jobs/components/ReorderJobButtons.tsx`, the `reorderJobAction` / `promoteJobAction` server actions, and `editor/e2e/jobs-reorder.spec.ts`.
+- **Job-queue internals unified onto one scheduler + one concurrency primitive (foundational refactor; no behavior change beyond the two features above).** The "foreground preempts background" priority was previously implemented twice (once for job queue ordering, once for worker-slot waiters); both now share a single `compareTier` comparator in a new `common/jobs/scheduler.ts` (priority tiers urgent/foreground/background, per-queue concurrency, and the reorder/promote operations the UI uses), which the registry delegates its queue ordering to. The auto-transcribe/-download runner's hand-written fill-to-capacity loop **and** the whisper batch's `Promise.all` are both replaced by one audited `runPool()` primitive (`common/jobs/concurrentRunner.ts`) that encodes the no-event-loop-spin wait once — structurally eliminating the "Drain all hangs" bug class rather than patching each loop. Per-kind metadata (label, drainability, bookmarkability) is consolidated into one `common/jobs/jobKinds.ts` table, and the bookmark/replay dispatch is now a data-driven lookup, so adding a job kind touches ~1 file instead of ~5. Covered by new unit tests (`scheduler`, `concurrentRunner`, `jobKinds`) and the existing queue/drain e2e suites.
- **"Drain all" no longer hangs the server when an auto-transcribe/-download unit is actively running.** A second, distinct drain hang remained after the earlier parked-unit fix: the runner's internal `waitNext()` helper short-circuited to an *immediately-resolved* promise whenever a signal was *already* aborted — so once a soft drain fired (and `drainSignal` stays aborted for the rest of the run), every wait while in-flight units were still finishing returned with no delay. That turned the runner's poll loops into a timer-less **microtask spin** that starved the Node event loop (the spin was reached first in the main fill loop's at-capacity wait once the drain target dropped to 0, before the terminal drain-wait was ever hit). A transcription that drain deliberately lets finish completes via a child-process `exit` event — a *macrotask* — which the spin never let run, so the in-flight count never reached zero, a CPU core pegged, and the whole app appeared frozen. (The earlier fix only covered *parked* units, which settle via microtasks and so cleared even under the spin; a genuinely *running* unit depends on a macrotask and didn't.) `waitNext()` now only fast-paths a real pending `wake()`; an already-aborted signal falls through to a real timer, so all three wait sites pace instead of spinning while still being woken promptly by a finishing unit or by an abort firing mid-wait. The `whisper-all` batch was never affected (it `await Promise.all(...)` with no manual poll loop). See `common/controller/autoRunner.ts` (`waitNext`) and the new running-unit drain regression test in `editor/e2e/auto-queue.spec.ts`.
- **First-class video-persistence UI (phase 5, the final phase): a Saved Videos area, per-channel retention controls, and per-video persist/unpersist.** The video-persistence subsystem built up over phases 1–4 is now driveable end to end from the editor. A new top-level **Saved videos** page (`/saved-videos`, in the Pool nav) summarizes the whole saved-video store — total count and size, per-channel breakdown (count, size, how many carry a backup checksum), the default store location, and the last backup time — and hosts the **backup configuration** (destination, scheduled on/off, interval) plus **Back up now** / **Verify backup** buttons. Each channel's **Cleanup stage** gains a **Retention & persistence** section (shown whenever keep-latest is on or the channel has saved videos) with live counts and three buttons: **Check kept videos** (re-probe the window for deleted-from-source videos and pin them), **Persist kept now** (a new bulk catch-up pass that re-fetches the source container for any in-window video whose source isn't saved yet — `persistKeptAction` / `common/controller/persistKept.ts`), and **Back up saved videos**. The **channel settings form** adds a Retention & persistence section: **keep latest** (window size), **extraction mode** (yt-dlp vs app-side ffmpeg), and a per-channel **saved-video store dir** override. Each **video page** gains a **Source video** card showing persisted status (file, size, stored time, keep reason, sha256, location) with an **Unpersist** control that moves the container back into the data dir, or a **Persist source video** button (re-fetch + archive) when it isn't saved yet. New job-kind label for `persist-kept`; `persist-kept` is re-runnable from bookmarks. The saved-video actions moved from `editor/app/savedVideos/` to `editor/app/saved-videos/` to match the route. Covered by `editor/e2e/saved-videos.spec.ts` and `common/controller/persistKept.test.ts`. (Deferred: surfacing kept-check/persist as Actionable-page rows, and streaming the player directly from the store — unpersist brings the container back to the data dir to play it.)
- **Backups for the saved-video store: rsync mirror + per-backup manifest + drift verification (phase 4 of the video-persistence subsystem; backend + scheduler, UI lands later).** The (large, often irreplaceable) saved source videos can now be **backed up to a configured destination**. A backup walks every saved-video pointer across all channels (so per-channel store overrides are covered automatically) and **`rsync`-mirrors each container** into `<dest>/<slug>/<videoId>/` — incremental and resumable (`-a --partial`), additive (no deletes), so re-running only transfers changed or new files. It writes a **`backup-manifest.json`** at the destination root recording each container's canonical location, byte size, and a **streamed sha256**, and caches that hash back onto the live pointer. A **verify** step reads the manifest back and reports drift in four buckets — `missing`, `sizeMismatch`, `checksumMismatch` (re-hashing each present file), and `extra` (containers at the destination the manifest doesn't know about). New global settings block **`savedVideoBackup`** (`{ enabled, dest, intervalMinutes }`; a blank `dest` forces `enabled` off) plus a **`RSYNC_BIN`** env override. When enabled with a destination, the **sync scheduler** runs the backup automatically on its own cadence (a global, not per-channel, job — suppressed during quiet hours, tracked via `lastSavedVideoBackupAt`). Backups can also be run/verified manually via `backupSavedVideosAction` / `verifySavedVideoBackupAction` (managed jobs on a dedicated `saved-videos` queue). The destination is treated as a local filesystem path (a mounted backup disk). See the new `common/lib/savedVideoBackup.ts` (manifest types/parse), `common/controller/{backupSavedVideos,savedVideoInventory}.ts` (+ tests), `common/lib/paths.ts` (`rsyncBin`), `common/lib/settings.ts` (`savedVideoBackup`), `common/jobs/syncSchedulerState.ts`, `editor/app/savedVideos/backupActions.ts`, and `editor/app/scheduler/runTick.ts`.
diff --git a/editor/app/jobs/actions.ts b/editor/app/jobs/actions.ts
@@ -4,6 +4,10 @@ import { revalidatePath } from "next/cache";
import { getPaths } from "yt-dlp-transcript-common/lib/paths";
import { clearArchivedLogs } from "yt-dlp-transcript-common/jobs/listJobs";
import { getRegistry } from "yt-dlp-transcript-common/jobs/registry";
+import { readJobMeta } from "yt-dlp-transcript-common/jobs/jobMeta";
+import type { JobSpec } from "yt-dlp-transcript-common/jobs/jobSpec";
+import type { StreamActionResult } from "yt-dlp-transcript-common/jobs/streamCommand";
+import { runJobSpec } from "./runJobSpec";
export async function cancelJobAction(id: string): Promise<{ ok: boolean }> {
const ok = getRegistry().cancel(id);
@@ -55,6 +59,56 @@ export async function promoteJobAction(id: string): Promise<{ ok: boolean }> {
return { ok };
}
+// Resolve a job's replay descriptor: prefer the live registry record, fall back
+// to the on-disk meta sidecar so a failed job the 100-job cap evicted (or one
+// from before a restart) can still be retried. Null if the kind isn't
+// replayable (no spec was ever attached).
+async function resolveSpec(id: string): Promise<JobSpec | null> {
+ const live = getRegistry().get(id)?.spec;
+ if (live) return live;
+ const meta = await readJobMeta(getPaths(), id);
+ return meta?.spec ?? null;
+}
+
+// One-click retry: re-run a (typically failed) job from its stored spec. Like a
+// bookmark re-run, bucket jobs re-derive from the channel's CURRENT state. The
+// re-run jumps ahead of other queued work (run next) by reusing the same
+// promote the reorder buttons use — a no-op if it starts immediately.
+export async function retryJobAction(id: string): Promise<StreamActionResult> {
+ const spec = await resolveSpec(id);
+ if (!spec) {
+ return {
+ ok: false,
+ error: "This job can't be retried — it has no replay descriptor.",
+ };
+ }
+ const res = await runJobSpec(spec);
+ if (res.ok) getRegistry().promote(res.jobId);
+ revalidatePath("/jobs");
+ return res;
+}
+
+// Retry every currently-listed failed job that has a spec. Returns how many were
+// re-launched (kinds without a replay descriptor are skipped).
+export async function retryAllFailedAction(): Promise<{ count: number }> {
+ const registry = getRegistry();
+ const failed = registry
+ .list()
+ .filter((j): j is typeof j & { spec: JobSpec } =>
+ Boolean(j.status === "failed" && j.spec),
+ );
+ let count = 0;
+ for (const j of failed) {
+ const res = await runJobSpec(j.spec);
+ if (res.ok) {
+ registry.promote(res.jobId);
+ count++;
+ }
+ }
+ revalidatePath("/jobs");
+ return { count };
+}
+
export async function clearArchivedAction(): Promise<{ deleted: number }> {
const deleted = await clearArchivedLogs(getPaths());
revalidatePath("/jobs");
diff --git a/editor/app/jobs/components/JobsTable.tsx b/editor/app/jobs/components/JobsTable.tsx
@@ -5,6 +5,7 @@ import Link from "next/link";
import type { JobListEntry } from "yt-dlp-transcript-common/jobs/listJobs";
import { CancelJobButton } from "./CancelJobButton";
import { BookmarkJobButton } from "./BookmarkJobButton";
+import { RetryJobButton } from "./RetryJobButton";
import { jobKindLabel } from "../jobKindLabels";
import {
clearJobsFilters,
@@ -280,6 +281,9 @@ export function JobsTable({ jobs }: { jobs: JobListEntry[] }) {
</td>
<td className="px-3 py-2 text-right">
<div className="flex items-center justify-end gap-2">
+ {j.status === "failed" && j.bookmarkable && (
+ <RetryJobButton jobId={j.id} />
+ )}
{j.bookmarkable && <BookmarkJobButton jobId={j.id} />}
{(j.status === "running" || j.status === "queued") && (
<CancelJobButton jobId={j.id} />
diff --git a/editor/app/jobs/components/RetryAllFailedButton.tsx b/editor/app/jobs/components/RetryAllFailedButton.tsx
@@ -0,0 +1,30 @@
+"use client";
+
+import { useState } from "react";
+import { useRouter } from "next/navigation";
+import { retryAllFailedAction } from "../actions";
+
+// Re-run every currently-listed failed job that has a replay descriptor. Shown
+// only when there's at least one such job (the page decides).
+export function RetryAllFailedButton() {
+ const [busy, setBusy] = useState(false);
+ const router = useRouter();
+ return (
+ <button
+ type="button"
+ disabled={busy}
+ onClick={async () => {
+ setBusy(true);
+ try {
+ await retryAllFailedAction();
+ router.refresh();
+ } finally {
+ setBusy(false);
+ }
+ }}
+ className="px-3 py-1.5 rounded border border-blue-300 dark:border-blue-800 text-sm font-medium text-blue-700 dark:text-blue-300 hover:bg-blue-50 dark:hover:bg-blue-950 disabled:opacity-50"
+ >
+ {busy ? "Retrying…" : "Retry all failed"}
+ </button>
+ );
+}
diff --git a/editor/app/jobs/components/RetryJobButton.tsx b/editor/app/jobs/components/RetryJobButton.tsx
@@ -0,0 +1,45 @@
+"use client";
+
+import { useState } from "react";
+import { useRouter } from "next/navigation";
+import { retryJobAction } from "../actions";
+
+type Props = {
+ jobId: string;
+};
+
+// Re-run a failed job from its stored spec (bucket jobs re-derive from the
+// channel's current state). The re-run jumps ahead of other queued work.
+export function RetryJobButton({ jobId }: Props) {
+ const [busy, setBusy] = useState(false);
+ const [note, setNote] = useState<string | null>(null);
+ const router = useRouter();
+ return (
+ <span className="inline-flex items-center gap-1">
+ <button
+ type="button"
+ disabled={busy}
+ onClick={async () => {
+ setBusy(true);
+ setNote(null);
+ try {
+ const res = await retryJobAction(jobId);
+ if (!res.ok) setNote(res.error);
+ router.refresh();
+ } finally {
+ setBusy(false);
+ }
+ }}
+ aria-label={`retry job ${jobId}`}
+ className="px-2 py-1 rounded border border-blue-300 dark:border-blue-800 text-xs font-medium text-blue-700 dark:text-blue-300 hover:bg-blue-50 dark:hover:bg-blue-950 disabled:opacity-50"
+ >
+ {busy ? "Retrying…" : "Retry"}
+ </button>
+ {note && (
+ <span className="text-xs text-zinc-500" title={note}>
+ {note}
+ </span>
+ )}
+ </span>
+ );
+}
diff --git a/editor/app/jobs/page.tsx b/editor/app/jobs/page.tsx
@@ -3,6 +3,7 @@ import { listAllJobs } from "yt-dlp-transcript-common/jobs/listJobs";
import { getPaths } from "yt-dlp-transcript-common/lib/paths";
import { ClearArchivedButton } from "./components/ClearArchivedButton";
import { JobsTable } from "./components/JobsTable";
+import { RetryAllFailedButton } from "./components/RetryAllFailedButton";
import { BookmarksMenu } from "./components/BookmarksMenu";
import { loadBookmarksView } from "./loadBookmarks";
@@ -15,11 +16,17 @@ export default async function JobsPage() {
listAllJobs(getPaths()),
loadBookmarksView(),
]);
+ const hasRetryableFailed = jobs.some(
+ (j) => j.status === "failed" && j.bookmarkable,
+ );
return (
<div className="flex flex-col gap-4">
<div className="flex items-center justify-between">
<h1 className="text-2xl font-semibold">Jobs</h1>
- <ClearArchivedButton />
+ <div className="flex items-center gap-2">
+ {hasRetryableFailed && <RetryAllFailedButton />}
+ <ClearArchivedButton />
+ </div>
</div>
<BookmarksMenu bookmarks={bookmarks} missingSlugs={missingSlugs} />
{jobs.length === 0 ? (
diff --git a/editor/e2e/jobs-retry.spec.ts b/editor/e2e/jobs-retry.spec.ts
@@ -0,0 +1,90 @@
+// One-click retry: a failed job that carries a replay descriptor shows a Retry
+// button that re-runs it (resolving the spec from the on-disk meta sidecar, so
+// even an evicted/archived job is retryable). A failed job WITHOUT a spec offers
+// no Retry. Failed jobs are staged as sidecar + log files so the scenario is
+// deterministic and exercises the sidecar-fallback path directly.
+
+import { mkdir, writeFile } from "node:fs/promises";
+import { test, expect } from "@playwright/test";
+import { resetData, resolvePath } from "./helpers";
+import { baseUrl } from "./baseUrl";
+
+async function invalidateCache() {
+ await fetch(`${baseUrl}/api/test/invalidate-cache`).catch(() => {});
+}
+
+async function makeSlowChannel(slug: string, name: string) {
+ const root = resolvePath(`test-transcripts/channels/${slug}`);
+ await mkdir(root, { recursive: true });
+ await writeFile(
+ `${root}/config.json`,
+ JSON.stringify({
+ handling: "youtube",
+ name,
+ url: `https://www.youtube.com/@${slug}/videos`,
+ ytdlpExtraArgs: ["--test-slow"],
+ }),
+ );
+}
+
+// Stage an archived job on disk: listAllJobs iterates `.log` files and merges
+// the `.meta.json` sidecar, so both are needed for the row to appear.
+async function writeArchivedJob(
+ id: string,
+ meta: Record<string, unknown>,
+): Promise<void> {
+ const dir = resolvePath("test-transcripts/.jobs");
+ await mkdir(dir, { recursive: true });
+ await writeFile(`${dir}/${id}.log`, `archived log for ${id}\n`);
+ await writeFile(`${dir}/${id}.meta.json`, JSON.stringify({ id, ...meta }));
+}
+
+test("Retry re-runs a failed job from its spec; spec-less failures offer none", async ({
+ page,
+}) => {
+ test.setTimeout(60_000);
+ await resetData(null);
+ await makeSlowChannel("retry-ch", "Retry Ch");
+
+ const ts = 1_700_000_000_000;
+ // A failed SYNC carrying a replay spec → retryable.
+ await writeArchivedJob("failedsync1", {
+ kind: "sync",
+ channelSlug: "retry-ch",
+ queueKey: "platform:youtube",
+ status: "failed",
+ queuedAt: ts,
+ startedAt: ts,
+ endedAt: ts + 1000,
+ exitCode: 1,
+ spec: { kind: "sync", slug: "retry-ch" },
+ });
+ // A failed job with NO spec → not retryable.
+ await writeArchivedJob("failednospec1", {
+ kind: "import-one",
+ channelSlug: "retry-ch",
+ status: "failed",
+ queuedAt: ts,
+ startedAt: ts,
+ endedAt: ts + 1000,
+ exitCode: 1,
+ });
+ await invalidateCache();
+
+ await page.goto("/jobs");
+
+ const failedRow = page.getByRole("row").filter({ hasText: "failedsync1" });
+ await expect(failedRow.getByRole("button", { name: "Retry" })).toBeVisible();
+ // The spec-less failed job offers no Retry.
+ const noSpecRow = page.getByRole("row").filter({ hasText: "failednospec1" });
+ await expect(noSpecRow.getByRole("button", { name: "Retry" })).toHaveCount(0);
+
+ // Retry re-runs the sync (slow channel → it shows up running).
+ await failedRow.getByRole("button", { name: "Retry" }).click();
+
+ const runningSync = page
+ .getByRole("row")
+ .filter({ hasText: "retry-ch" })
+ .filter({ hasText: "running" });
+ await expect(runningSync).toBeVisible({ timeout: 15_000 });
+});