Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 49814eee56824921830960c7b0ca813b08e301d5
parent 051f308bcaea03c739f95c77d5118715c9c47b2d
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Fri, 19 Jun 2026 19:04:00 -0400

Fix retry-bucket prefilter skipping partials with an audio-check snapshot

destinationExists() decided a transcribe channel's video was "already
complete" via a hand-rolled `startsWith("audio.") && !endsWith(".part")`
check, which wrongly matched audio-integrity snapshots
(audio.<ext>.part.good/.part.testing), sidecars (audio.info.json,
audio.live_chat.json), and temp files. A genuine audio.<ext>.part sitting
next to a .part.good snapshot was prefiltered out ("Nothing to fetch")
even though the partial-downloads bucket — which uses isRealAudioFile —
correctly excludes those same files.

Reuse isRealAudioFile in destinationExists so the prefilter and the bucket
agree. Add a regression e2e test that seeds a .part.good snapshot beside
the .part and asserts the partial is kept (verified it fails without the
fix).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

Diffstat:
Mcommon/ytdlp/runYtdlp.ts | 11+++++++----
Meditor/CHANGELOG.md | 1+
Meditor/e2e/partial-downloads-bucket.spec.ts | 37+++++++++++++++++++++++++++++++++++++
3 files changed, 45 insertions(+), 4 deletions(-)

diff --git a/common/ytdlp/runYtdlp.ts b/common/ytdlp/runYtdlp.ts @@ -13,6 +13,7 @@ import { getSettings } from "../lib/settings"; import { checkDiskSpace } from "../lib/diskSpace"; import { formatBytes } from "../lib/format"; import { detectPlatform } from "../lib/platform"; +import { isRealAudioFile } from "../lib/videoStatus"; import type { Paths } from "../lib/paths"; import { EXCLUDED_FROM_DOWNLOAD, @@ -878,11 +879,13 @@ export async function destinationExists( } // Transcribe channels treat raw audio as a download in progress so we don't // re-fetch it before whisper runs. YouTube channels expect a .vtt; an audio - // file alone shouldn't suppress the next sync. + // file alone shouldn't suppress the next sync. Use isRealAudioFile (the same + // predicate the partial-downloads bucket uses) so audio-check snapshots + // (audio.<ext>.part.good/.part.testing), sidecars (audio.info.json, + // audio.live_chat.json), and temp files don't get mistaken for finished audio + // — otherwise a genuine partial carrying a .part.good snapshot is skipped. if (handling === "transcribe") { - return entries.some( - (e) => e.startsWith("audio.") && !e.endsWith(".part"), - ); + return entries.some(isRealAudioFile); } return false; } diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,7 @@ # Changelog ## [Unreleased] +- **"Retry partial downloads" no longer skips partials that carry an audio-check snapshot.** A transcribe channel's retry-bucket prefilter decided a video was "already complete" by looking for any `audio.*` file not ending in `.part` — which wrongly matched the audio-integrity snapshots (`audio.<ext>.part.good` / `.part.testing`) and sidecars (`audio.info.json`, `audio.live_chat.json`, `audio.*.tmp-*`) left in a partial video's dir. So a genuine `audio.<ext>.part` that happened to sit next to a `.part.good` snapshot got prefiltered out (`Prefilter: 0 missing destination files, N already complete` → `Nothing to fetch`), even though that same snapshot is excluded when the video is placed in the **Partial downloads** bucket. The prefilter (`destinationExists`) now reuses the same `isRealAudioFile` predicate the bucket uses, so the two agree and genuine partials resume. See `common/ytdlp/runYtdlp.ts` and `common/lib/videoStatus.ts`. - **One-click "Drain all" to spin work down before a server restart.** The `/jobs/active` header gained a **Drain all** button that, in a single confirmed action, **drains every running job** (lets in-flight sub-operations finish, starts no new ones) and **cancels every queued job** — so you can wind the queue down gracefully before restarting the server instead of draining/cancelling each job row by hand. It reuses the existing per-job drain path (`requestDrain` drains running jobs and cancels queued ones), so semantics match the per-row **Drain**/**Cancel** buttons exactly; non-drainable running kinds are marked draining and finish on their own. A confirm prompt guards the bulk action. See `editor/app/jobs/components/DrainAllButton.tsx` and `drainAllAction` in `editor/app/jobs/actions.ts`. - **Queued jobs now have a Cancel button right on the row.** On `/jobs/active`, a queued (not-yet-started) job previously offered no inline way to cancel it — you had to expand its log to find a control. Each active row now shows a **Cancel** button directly (queued *and* running), so you can drop a queued job without opening anything. The running-only **Drain** button is unchanged. See `editor/app/jobs/components/RunningJobsList.tsx`. - **The Bookmarks panel is now a compact one-click run menu, with management moved to its own page.** The bookmarks block atop `/jobs` and `/jobs/active` used to render a heavy multi-row card per bookmark — inline rename, tags, created/last-run metadata, a confirm-gated delete, and a full streaming **Run again** button that popped a tall inline log box on every run — which buried the live job list below it. It's now a tight **wrap of quick-access buttons**, one per bookmark, labeled with the bookmark's name (the full kind/bucket/channel shows on hover). Clicking a button **launches the job fire-and-forget** — no inline shell, no expanding log — and the relaunched job simply appears in the live list below (it runs to completion server-side whether or not the page watches its stream); the button briefly reads *Launching… → Launched ✓*, an empty re-derived bucket still reads as a neutral *"Nothing to retry right now"* line, and a bookmark whose channel was deleted is disabled with a *channel missing* hint. The heavier controls — rename, delete, created/last-run metadata, and the full streaming **Run again** with its inline log — move to a new **`/jobs/bookmarks`** management page, reachable via the **Manage** link in the menu header. See `editor/app/jobs/components/BookmarksMenu.tsx`, `editor/app/jobs/components/BookmarkRunButton.tsx`, and `editor/app/jobs/bookmarks/page.tsx`. diff --git a/editor/e2e/partial-downloads-bucket.spec.ts b/editor/e2e/partial-downloads-bucket.spec.ts @@ -64,3 +64,40 @@ test("resume action re-downloads the partial video", async ({ page }) => { await pathExists(`${CHANNEL_ROOT}/data/vidA/audio.m4a`), ).toBe(true); }); + +test("resume retries a partial that carries an audio-check snapshot", async ({ + page, +}) => { + // Regression: a genuine audio.<ext>.part alongside an audio-check snapshot + // (audio.<ext>.part.good) used to be prefiltered out as "already complete" + // because destinationExists() matched any audio.* not ending in .part. The + // prefilter must agree with the bucket and keep the partial. + test.setTimeout(60_000); + await resetData("one-transcribe-channel-with-audio"); + await rename( + resolvePath(`${CHANNEL_ROOT}/data/vidA/audio.m4a`), + resolvePath(`${CHANNEL_ROOT}/data/vidA/audio.m4a.part`), + ); + // The leftover audio-check snapshot that previously fooled the prefilter. + await writeFile( + resolvePath(`${CHANNEL_ROOT}/data/vidA/audio.m4a.part.good`), + "snapshot\n", + ); + await writeFile( + resolvePath(`${CHANNEL_ROOT}/playlist`), + "https://www.youtube.com/watch?v=vidA\n", + ); + + await page.goto(`/channels/${CHANNEL}`); + + await page + .getByLabel("retry resume partial downloads bucket") + .getByRole("button", { name: /^Retry \(1\)$/ }) + .click(); + + const log = page.getByLabel("Retry resume partial downloads output"); + await expect(log).toContainText("Prefilter: 1 missing destination files", { + timeout: 30_000, + }); + await expect(log).not.toContainText("Nothing to fetch."); +});