Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 0a37f7cacabb46b59f92b900fbe44883d4d0d058
parent 7cfcab884b11e6c0696d4f03d1caaa3254914a36
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Fri, 26 Jun 2026 15:59:23 -0400

Phase 8: one-click retry + retry-all for failed jobs

Granular control (priority feature #2): re-run a failed job from its stored
spec, jumping ahead of other queued work.

- retryJobAction(id): resolve the spec from the live registry or the on-disk
  meta sidecar (so an evicted/archived failure is still retryable), re-run via
  runJobSpec (bucket jobs re-derive from current state), then promote the new
  job to the front of its queue (reusing Phase 7's promote — same observable
  "runs next" as an urgent tier, without threading tier through every action).
- retryAllFailedAction(): re-run every listed failed job that has a spec.
- RetryJobButton on each failed+bookmarkable row in the Jobs table; a
  RetryAllFailedButton in the page header when any such job exists. Kinds with
  no replay descriptor offer no Retry.

Also adds CHANGELOG entries for retry, reorder/promote (Phase 7), and the
foundational scheduler/runPool/jobKinds refactor.

Verification: jobs-retry e2e (retry re-runs from the sidecar-staged failure;
spec-less failures show no Retry) + jobs-filters/jobs/bookmarks green; typecheck
(common+editor) clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 3+++
Meditor/app/jobs/actions.ts | 54++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/app/jobs/components/JobsTable.tsx | 4++++
Aeditor/app/jobs/components/RetryAllFailedButton.tsx | 30++++++++++++++++++++++++++++++
Aeditor/app/jobs/components/RetryJobButton.tsx | 45+++++++++++++++++++++++++++++++++++++++++++++
Meditor/app/jobs/page.tsx | 9++++++++-
Aeditor/e2e/jobs-retry.spec.ts | 90+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
7 files changed, 234 insertions(+), 1 deletion(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,9 @@ # Changelog ## [Unreleased] +- **One-click Retry for failed jobs (plus "Retry all failed").** A failed job that carries a replay descriptor (any bookmarkable kind — sync, download-missing, transcribe-all, retry-bucket, …) now shows a **Retry** button on the Jobs history table, and the page header gains a **Retry all failed** button whenever at least one such job is listed. Retry re-runs the job from its stored spec exactly like a bookmark re-run (so bucket jobs re-derive from the channel's *current* state), and the re-run **jumps ahead of other queued work** (it's promoted to the front of its queue, reusing the new reorder machinery) so a fix-and-retry runs next rather than at the back of the line. The spec is resolved from the live registry or, for an evicted/archived job, from its on-disk `<id>.meta.json` sidecar — so even a failure the 100-job cap has dropped is still retryable. Kinds with no replay descriptor (e.g. `import-one`) intentionally offer no Retry. See `editor/app/jobs/actions.ts` (`retryJobAction` / `retryAllFailedAction`), the new `RetryJobButton` / `RetryAllFailedButton`, and `editor/e2e/jobs-retry.spec.ts`. +- **Reorder and promote queued jobs from Active Jobs.** A queued job's row now carries **Promote / ↑ / ↓** controls (mirroring the auto-queue policy editor's move buttons) to change its order within its queue — Promote sends it to the front so it runs next, ↑/↓ nudge it one slot. Only actionable moves render (the first-queued job shows no up/promote, the last no down), and a running job is never displaced. See `editor/app/jobs/components/ReorderJobButtons.tsx`, the `reorderJobAction` / `promoteJobAction` server actions, and `editor/e2e/jobs-reorder.spec.ts`. +- **Job-queue internals unified onto one scheduler + one concurrency primitive (foundational refactor; no behavior change beyond the two features above).** The "foreground preempts background" priority was previously implemented twice (once for job queue ordering, once for worker-slot waiters); both now share a single `compareTier` comparator in a new `common/jobs/scheduler.ts` (priority tiers urgent/foreground/background, per-queue concurrency, and the reorder/promote operations the UI uses), which the registry delegates its queue ordering to. The auto-transcribe/-download runner's hand-written fill-to-capacity loop **and** the whisper batch's `Promise.all` are both replaced by one audited `runPool()` primitive (`common/jobs/concurrentRunner.ts`) that encodes the no-event-loop-spin wait once — structurally eliminating the "Drain all hangs" bug class rather than patching each loop. Per-kind metadata (label, drainability, bookmarkability) is consolidated into one `common/jobs/jobKinds.ts` table, and the bookmark/replay dispatch is now a data-driven lookup, so adding a job kind touches ~1 file instead of ~5. Covered by new unit tests (`scheduler`, `concurrentRunner`, `jobKinds`) and the existing queue/drain e2e suites. - **"Drain all" no longer hangs the server when an auto-transcribe/-download unit is actively running.** A second, distinct drain hang remained after the earlier parked-unit fix: the runner's internal `waitNext()` helper short-circuited to an *immediately-resolved* promise whenever a signal was *already* aborted — so once a soft drain fired (and `drainSignal` stays aborted for the rest of the run), every wait while in-flight units were still finishing returned with no delay. That turned the runner's poll loops into a timer-less **microtask spin** that starved the Node event loop (the spin was reached first in the main fill loop's at-capacity wait once the drain target dropped to 0, before the terminal drain-wait was ever hit). A transcription that drain deliberately lets finish completes via a child-process `exit` event — a *macrotask* — which the spin never let run, so the in-flight count never reached zero, a CPU core pegged, and the whole app appeared frozen. (The earlier fix only covered *parked* units, which settle via microtasks and so cleared even under the spin; a genuinely *running* unit depends on a macrotask and didn't.) `waitNext()` now only fast-paths a real pending `wake()`; an already-aborted signal falls through to a real timer, so all three wait sites pace instead of spinning while still being woken promptly by a finishing unit or by an abort firing mid-wait. The `whisper-all` batch was never affected (it `await Promise.all(...)` with no manual poll loop). See `common/controller/autoRunner.ts` (`waitNext`) and the new running-unit drain regression test in `editor/e2e/auto-queue.spec.ts`. - **First-class video-persistence UI (phase 5, the final phase): a Saved Videos area, per-channel retention controls, and per-video persist/unpersist.** The video-persistence subsystem built up over phases 1–4 is now driveable end to end from the editor. A new top-level **Saved videos** page (`/saved-videos`, in the Pool nav) summarizes the whole saved-video store — total count and size, per-channel breakdown (count, size, how many carry a backup checksum), the default store location, and the last backup time — and hosts the **backup configuration** (destination, scheduled on/off, interval) plus **Back up now** / **Verify backup** buttons. Each channel's **Cleanup stage** gains a **Retention & persistence** section (shown whenever keep-latest is on or the channel has saved videos) with live counts and three buttons: **Check kept videos** (re-probe the window for deleted-from-source videos and pin them), **Persist kept now** (a new bulk catch-up pass that re-fetches the source container for any in-window video whose source isn't saved yet — `persistKeptAction` / `common/controller/persistKept.ts`), and **Back up saved videos**. The **channel settings form** adds a Retention & persistence section: **keep latest** (window size), **extraction mode** (yt-dlp vs app-side ffmpeg), and a per-channel **saved-video store dir** override. Each **video page** gains a **Source video** card showing persisted status (file, size, stored time, keep reason, sha256, location) with an **Unpersist** control that moves the container back into the data dir, or a **Persist source video** button (re-fetch + archive) when it isn't saved yet. New job-kind label for `persist-kept`; `persist-kept` is re-runnable from bookmarks. The saved-video actions moved from `editor/app/savedVideos/` to `editor/app/saved-videos/` to match the route. Covered by `editor/e2e/saved-videos.spec.ts` and `common/controller/persistKept.test.ts`. (Deferred: surfacing kept-check/persist as Actionable-page rows, and streaming the player directly from the store — unpersist brings the container back to the data dir to play it.) - **Backups for the saved-video store: rsync mirror + per-backup manifest + drift verification (phase 4 of the video-persistence subsystem; backend + scheduler, UI lands later).** The (large, often irreplaceable) saved source videos can now be **backed up to a configured destination**. A backup walks every saved-video pointer across all channels (so per-channel store overrides are covered automatically) and **`rsync`-mirrors each container** into `<dest>/<slug>/<videoId>/` — incremental and resumable (`-a --partial`), additive (no deletes), so re-running only transfers changed or new files. It writes a **`backup-manifest.json`** at the destination root recording each container's canonical location, byte size, and a **streamed sha256**, and caches that hash back onto the live pointer. A **verify** step reads the manifest back and reports drift in four buckets — `missing`, `sizeMismatch`, `checksumMismatch` (re-hashing each present file), and `extra` (containers at the destination the manifest doesn't know about). New global settings block **`savedVideoBackup`** (`{ enabled, dest, intervalMinutes }`; a blank `dest` forces `enabled` off) plus a **`RSYNC_BIN`** env override. When enabled with a destination, the **sync scheduler** runs the backup automatically on its own cadence (a global, not per-channel, job — suppressed during quiet hours, tracked via `lastSavedVideoBackupAt`). Backups can also be run/verified manually via `backupSavedVideosAction` / `verifySavedVideoBackupAction` (managed jobs on a dedicated `saved-videos` queue). The destination is treated as a local filesystem path (a mounted backup disk). See the new `common/lib/savedVideoBackup.ts` (manifest types/parse), `common/controller/{backupSavedVideos,savedVideoInventory}.ts` (+ tests), `common/lib/paths.ts` (`rsyncBin`), `common/lib/settings.ts` (`savedVideoBackup`), `common/jobs/syncSchedulerState.ts`, `editor/app/savedVideos/backupActions.ts`, and `editor/app/scheduler/runTick.ts`. diff --git a/editor/app/jobs/actions.ts b/editor/app/jobs/actions.ts @@ -4,6 +4,10 @@ import { revalidatePath } from "next/cache"; import { getPaths } from "yt-dlp-transcript-common/lib/paths"; import { clearArchivedLogs } from "yt-dlp-transcript-common/jobs/listJobs"; import { getRegistry } from "yt-dlp-transcript-common/jobs/registry"; +import { readJobMeta } from "yt-dlp-transcript-common/jobs/jobMeta"; +import type { JobSpec } from "yt-dlp-transcript-common/jobs/jobSpec"; +import type { StreamActionResult } from "yt-dlp-transcript-common/jobs/streamCommand"; +import { runJobSpec } from "./runJobSpec"; export async function cancelJobAction(id: string): Promise<{ ok: boolean }> { const ok = getRegistry().cancel(id); @@ -55,6 +59,56 @@ export async function promoteJobAction(id: string): Promise<{ ok: boolean }> { return { ok }; } +// Resolve a job's replay descriptor: prefer the live registry record, fall back +// to the on-disk meta sidecar so a failed job the 100-job cap evicted (or one +// from before a restart) can still be retried. Null if the kind isn't +// replayable (no spec was ever attached). +async function resolveSpec(id: string): Promise<JobSpec | null> { + const live = getRegistry().get(id)?.spec; + if (live) return live; + const meta = await readJobMeta(getPaths(), id); + return meta?.spec ?? null; +} + +// One-click retry: re-run a (typically failed) job from its stored spec. Like a +// bookmark re-run, bucket jobs re-derive from the channel's CURRENT state. The +// re-run jumps ahead of other queued work (run next) by reusing the same +// promote the reorder buttons use — a no-op if it starts immediately. +export async function retryJobAction(id: string): Promise<StreamActionResult> { + const spec = await resolveSpec(id); + if (!spec) { + return { + ok: false, + error: "This job can't be retried — it has no replay descriptor.", + }; + } + const res = await runJobSpec(spec); + if (res.ok) getRegistry().promote(res.jobId); + revalidatePath("/jobs"); + return res; +} + +// Retry every currently-listed failed job that has a spec. Returns how many were +// re-launched (kinds without a replay descriptor are skipped). +export async function retryAllFailedAction(): Promise<{ count: number }> { + const registry = getRegistry(); + const failed = registry + .list() + .filter((j): j is typeof j & { spec: JobSpec } => + Boolean(j.status === "failed" && j.spec), + ); + let count = 0; + for (const j of failed) { + const res = await runJobSpec(j.spec); + if (res.ok) { + registry.promote(res.jobId); + count++; + } + } + revalidatePath("/jobs"); + return { count }; +} + export async function clearArchivedAction(): Promise<{ deleted: number }> { const deleted = await clearArchivedLogs(getPaths()); revalidatePath("/jobs"); diff --git a/editor/app/jobs/components/JobsTable.tsx b/editor/app/jobs/components/JobsTable.tsx @@ -5,6 +5,7 @@ import Link from "next/link"; import type { JobListEntry } from "yt-dlp-transcript-common/jobs/listJobs"; import { CancelJobButton } from "./CancelJobButton"; import { BookmarkJobButton } from "./BookmarkJobButton"; +import { RetryJobButton } from "./RetryJobButton"; import { jobKindLabel } from "../jobKindLabels"; import { clearJobsFilters, @@ -280,6 +281,9 @@ export function JobsTable({ jobs }: { jobs: JobListEntry[] }) { </td> <td className="px-3 py-2 text-right"> <div className="flex items-center justify-end gap-2"> + {j.status === "failed" && j.bookmarkable && ( + <RetryJobButton jobId={j.id} /> + )} {j.bookmarkable && <BookmarkJobButton jobId={j.id} />} {(j.status === "running" || j.status === "queued") && ( <CancelJobButton jobId={j.id} /> diff --git a/editor/app/jobs/components/RetryAllFailedButton.tsx b/editor/app/jobs/components/RetryAllFailedButton.tsx @@ -0,0 +1,30 @@ +"use client"; + +import { useState } from "react"; +import { useRouter } from "next/navigation"; +import { retryAllFailedAction } from "../actions"; + +// Re-run every currently-listed failed job that has a replay descriptor. Shown +// only when there's at least one such job (the page decides). +export function RetryAllFailedButton() { + const [busy, setBusy] = useState(false); + const router = useRouter(); + return ( + <button + type="button" + disabled={busy} + onClick={async () => { + setBusy(true); + try { + await retryAllFailedAction(); + router.refresh(); + } finally { + setBusy(false); + } + }} + className="px-3 py-1.5 rounded border border-blue-300 dark:border-blue-800 text-sm font-medium text-blue-700 dark:text-blue-300 hover:bg-blue-50 dark:hover:bg-blue-950 disabled:opacity-50" + > + {busy ? "Retrying…" : "Retry all failed"} + </button> + ); +} diff --git a/editor/app/jobs/components/RetryJobButton.tsx b/editor/app/jobs/components/RetryJobButton.tsx @@ -0,0 +1,45 @@ +"use client"; + +import { useState } from "react"; +import { useRouter } from "next/navigation"; +import { retryJobAction } from "../actions"; + +type Props = { + jobId: string; +}; + +// Re-run a failed job from its stored spec (bucket jobs re-derive from the +// channel's current state). The re-run jumps ahead of other queued work. +export function RetryJobButton({ jobId }: Props) { + const [busy, setBusy] = useState(false); + const [note, setNote] = useState<string | null>(null); + const router = useRouter(); + return ( + <span className="inline-flex items-center gap-1"> + <button + type="button" + disabled={busy} + onClick={async () => { + setBusy(true); + setNote(null); + try { + const res = await retryJobAction(jobId); + if (!res.ok) setNote(res.error); + router.refresh(); + } finally { + setBusy(false); + } + }} + aria-label={`retry job ${jobId}`} + className="px-2 py-1 rounded border border-blue-300 dark:border-blue-800 text-xs font-medium text-blue-700 dark:text-blue-300 hover:bg-blue-50 dark:hover:bg-blue-950 disabled:opacity-50" + > + {busy ? "Retrying…" : "Retry"} + </button> + {note && ( + <span className="text-xs text-zinc-500" title={note}> + {note} + </span> + )} + </span> + ); +} diff --git a/editor/app/jobs/page.tsx b/editor/app/jobs/page.tsx @@ -3,6 +3,7 @@ import { listAllJobs } from "yt-dlp-transcript-common/jobs/listJobs"; import { getPaths } from "yt-dlp-transcript-common/lib/paths"; import { ClearArchivedButton } from "./components/ClearArchivedButton"; import { JobsTable } from "./components/JobsTable"; +import { RetryAllFailedButton } from "./components/RetryAllFailedButton"; import { BookmarksMenu } from "./components/BookmarksMenu"; import { loadBookmarksView } from "./loadBookmarks"; @@ -15,11 +16,17 @@ export default async function JobsPage() { listAllJobs(getPaths()), loadBookmarksView(), ]); + const hasRetryableFailed = jobs.some( + (j) => j.status === "failed" && j.bookmarkable, + ); return ( <div className="flex flex-col gap-4"> <div className="flex items-center justify-between"> <h1 className="text-2xl font-semibold">Jobs</h1> - <ClearArchivedButton /> + <div className="flex items-center gap-2"> + {hasRetryableFailed && <RetryAllFailedButton />} + <ClearArchivedButton /> + </div> </div> <BookmarksMenu bookmarks={bookmarks} missingSlugs={missingSlugs} /> {jobs.length === 0 ? ( diff --git a/editor/e2e/jobs-retry.spec.ts b/editor/e2e/jobs-retry.spec.ts @@ -0,0 +1,90 @@ +// One-click retry: a failed job that carries a replay descriptor shows a Retry +// button that re-runs it (resolving the spec from the on-disk meta sidecar, so +// even an evicted/archived job is retryable). A failed job WITHOUT a spec offers +// no Retry. Failed jobs are staged as sidecar + log files so the scenario is +// deterministic and exercises the sidecar-fallback path directly. + +import { mkdir, writeFile } from "node:fs/promises"; +import { test, expect } from "@playwright/test"; +import { resetData, resolvePath } from "./helpers"; +import { baseUrl } from "./baseUrl"; + +async function invalidateCache() { + await fetch(`${baseUrl}/api/test/invalidate-cache`).catch(() => {}); +} + +async function makeSlowChannel(slug: string, name: string) { + const root = resolvePath(`test-transcripts/channels/${slug}`); + await mkdir(root, { recursive: true }); + await writeFile( + `${root}/config.json`, + JSON.stringify({ + handling: "youtube", + name, + url: `https://www.youtube.com/@${slug}/videos`, + ytdlpExtraArgs: ["--test-slow"], + }), + ); +} + +// Stage an archived job on disk: listAllJobs iterates `.log` files and merges +// the `.meta.json` sidecar, so both are needed for the row to appear. +async function writeArchivedJob( + id: string, + meta: Record<string, unknown>, +): Promise<void> { + const dir = resolvePath("test-transcripts/.jobs"); + await mkdir(dir, { recursive: true }); + await writeFile(`${dir}/${id}.log`, `archived log for ${id}\n`); + await writeFile(`${dir}/${id}.meta.json`, JSON.stringify({ id, ...meta })); +} + +test("Retry re-runs a failed job from its spec; spec-less failures offer none", async ({ + page, +}) => { + test.setTimeout(60_000); + await resetData(null); + await makeSlowChannel("retry-ch", "Retry Ch"); + + const ts = 1_700_000_000_000; + // A failed SYNC carrying a replay spec → retryable. + await writeArchivedJob("failedsync1", { + kind: "sync", + channelSlug: "retry-ch", + queueKey: "platform:youtube", + status: "failed", + queuedAt: ts, + startedAt: ts, + endedAt: ts + 1000, + exitCode: 1, + spec: { kind: "sync", slug: "retry-ch" }, + }); + // A failed job with NO spec → not retryable. + await writeArchivedJob("failednospec1", { + kind: "import-one", + channelSlug: "retry-ch", + status: "failed", + queuedAt: ts, + startedAt: ts, + endedAt: ts + 1000, + exitCode: 1, + }); + await invalidateCache(); + + await page.goto("/jobs"); + + const failedRow = page.getByRole("row").filter({ hasText: "failedsync1" }); + await expect(failedRow.getByRole("button", { name: "Retry" })).toBeVisible(); + // The spec-less failed job offers no Retry. + const noSpecRow = page.getByRole("row").filter({ hasText: "failednospec1" }); + await expect(noSpecRow.getByRole("button", { name: "Retry" })).toHaveCount(0); + + // Retry re-runs the sync (slow channel → it shows up running). + await failedRow.getByRole("button", { name: "Retry" }).click(); + + const runningSync = page + .getByRole("row") + .filter({ hasText: "retry-ch" }) + .filter({ hasText: "running" }); + await expect(runningSync).toBeVisible({ timeout: 15_000 }); +});