commit 14f94531e4b479ea2b35cd3e53a982f95e15986a
parent 0ddd9434e22300fb54d9edda93105e19ce9f7ee5
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Sun, 14 Jun 2026 11:18:24 -0400
Merge branch 'worktree-fix-nft-warning'
Diffstat:
4 files changed, 13 insertions(+), 7 deletions(-)
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -8,7 +8,7 @@
- **Remote workers: offload transcription to another instance of this app on your LAN.** Add a **remote** worker in the Settings list with the base URL of another instance (e.g. `http://gpu-box.lan:3001`) and a shared token. When a video is dispatched to it, this instance uploads the audio over HTTP, the remote transcribes it through *its own* worker pool (picking among its local engines), streams progress and log back, and this instance pulls the finished `transcript.json` and normalizes it locally — so the remote needs no knowledge of your channels, just CPU/GPU. The protocol lives under `/api/worker/*` and is **disabled unless `WORKER_TOKEN` is set** in the environment, so an instance is never an open transcription server by accident; every request carries `Authorization: Bearer <token>`, validated with a constant-time compare against the accepting instance's own `WORKER_TOKEN` (never against settings). Uploaded audio and the produced transcript live in a scratch dir that's cleaned up once the result is pulled (or the job is cancelled). If a remote returns a transport error mid-job, the video is automatically retried on another worker; a genuine transcription failure on the remote is not retried. On a transport failure the remote's `GET /api/worker/health` is probed, and a remote confirmed **down** is auto-disabled (shown "degraded" on the Workers page, with **Enable** to retry once it's back) so neither the current video nor later ones keep burning attempts on it — they fail over to a healthy worker. A worker that racks up repeated failures while still reachable is auto-disabled after a few strikes.
- **New Workers page (`/workers`) with live status and runtime controls.** Lists every worker with its state (idle / busy / draining / disabled / degraded) and the video it's currently transcribing with per-task progress. Each worker can be **disabled** (stop taking new work immediately; in-flight transcriptions keep running), **drained** (stop taking new work but let the current video finish — the graceful "free up the GPU when it's done" path), or **enabled** again — without editing settings, so you can hand a CPU/GPU back to other programs and reclaim it later. A **Pause all** button disables every worker at once and remembers each one's state; **Resume all** restores them exactly. These runtime controls are transient (a restart returns workers to their configured enabled state); the Settings list is where the persisted defaults live.
- **parakeet.cpp transcriptions are resumable, and can be paused mid-run.** The overlapping-segment wrapper now writes each window's raw parakeet-cli JSON to a per-audio work dir (`.<audio>.parakeet/`) as it finishes, and stitches the final transcript only once *all* windows are done (then removes the work dir). Re-running the same transcription picks up the cached windows and only does what's missing — yt-dlp-style resume, so a crash, cancel, or pause never loses completed windows. A busy parakeet worker on the Workers page shows a **Stop & keep progress** button: it finishes the in-flight window, stops and frees the worker (no transcript written yet), and the next "Transcribe missing" resumes from the cached windows and completes. Handy to reclaim a GPU mid-run. (whisper.cpp/chough run as a single pass and don't offer this.)
-- **Selectable compute device for parakeet.cpp.** A parakeet worker gained a **Device** field (e.g. `cuda:0`, `cpu`) passed through to `parakeet-cli` (`--device`, also honored as the `PARAKEET_DEVICE` env). Combined with one-worker-per-slot and Copy, you can pin different parakeet workers to different GPUs.
+- **Selectable compute device for parakeet.cpp.** A parakeet worker gained a **Device** field that forces the compute device — `cpu` to run on CPU, or a specific GPU like `CUDA0` / `Vulkan1`. parakeet.cpp's `parakeet-cli` has no `--device` flag and otherwise auto-grabs the first GPU the ggml registry reports, so the wrapper now exports the choice as the **`PARAKEET_DEVICE`** environment variable to the CLI (previously it was passed as a non-existent `--device` flag, which the CLI ignored — so a worker set to `cpu` still ran on the GPU). Combined with one-worker-per-slot and Copy, you can pin different parakeet workers to different devices.
- **The Workers page and Active Jobs page cross-reference each other.** Each busy worker on `/workers` now shows what it's transcribing right now — the video (linked), the channel it's in, a live elapsed timer, percent, and the engine's progress detail — not just a bare bar. Conversely, every in-flight transcription on `/jobs/active` now says which worker it's running **on** (e.g. "Transcribing <id> on GPU"), so you can see how a batch is spread across your workers at a glance.
- **Batches pause instead of failing when no worker is available.** If every worker is disabled (or you hit **Pause all**) while a transcription batch is running, the batch parks — it keeps its in-flight video to completion, starts no new ones, and stays **running** on `/jobs/active` rather than failing the remaining videos. Re-enabling any worker (or **Resume all**) immediately resumes it where it left off. A video whose worker fails for a transport reason (e.g. a remote worker that went away) is automatically retried on another worker before being recorded as failed.
- **Git worktrees can run in parallel on non-colliding ports (dev tooling).** Two checkouts of the repo (via `git worktree`) can now run their dev servers and e2e suites at the same time without port clashes. A new helper, `scripts/worktree.mjs` (exposed as `pnpm wt`), assigns each worktree a port block offset by `index * 100` based on its position in `git worktree list` — the main worktree keeps the original defaults (editor 3001, test 3011, export 3010/3000/3020), worktree #1 gets 31xx, and so on. `pnpm dev:editor`, `pnpm dev:export`, `pnpm start:export`, and `pnpm e2e` route through `wt run`, which injects the assigned ports, so they "just work" per worktree; the editor/export `package.json` port flags and the Playwright configs now honor these env vars (previously `pnpm dev:test` hardcoded 3011, so a custom `PORT` only moved the URL Playwright waited on, not the server). E2E specs that hit the editor's test API now derive the base URL from `PLAYWRIGHT_BASE_URL` (centralized in `editor/e2e/baseUrl.ts`) instead of hardcoding `localhost:3011`. `pnpm wt add <branch>` creates a sibling worktree pre-seeded with `settings.json` and prints its ports; `--share-data` links it to the main worktree's downloaded `transcripts/` for read-mostly reuse (with an LMDB concurrent-write caveat). See `WORKTREES.md`.
diff --git a/editor/app/settings/components/WorkersField.tsx b/editor/app/settings/components/WorkersField.tsx
@@ -305,7 +305,7 @@ function WorkerCard({
label="Device"
value={cfg.device ?? ""}
onChange={(v) => onPatchConfig({ device: v })}
- hint="parakeet compute device passed to parakeet-cli (--device), e.g. cuda:0 or cpu. Blank = parakeet-cli's default."
+ hint="parakeet compute device (sets PARAKEET_DEVICE). e.g. cpu to force CPU, or CUDA0 / Vulkan1 for a specific GPU. Blank = parakeet-cli's default (first GPU)."
id={`${uid}-device`}
/>
)}
diff --git a/editor/e2e/parakeet-partial.spec.ts b/editor/e2e/parakeet-partial.spec.ts
@@ -73,7 +73,7 @@ test("a parakeet worker's device is configurable and persists", async ({
await page.goto("/settings");
// Migration makes worker 1 whisper-cpp; switch it to parakeet to reveal Device.
await page.getByLabel("worker 1 engine").selectOption("parakeet");
- await page.getByLabel(/^Device/).fill("cuda:0");
+ await page.getByLabel(/^Device/).fill("CUDA0");
await page.getByRole("button", { name: /save settings/i }).click();
await expect(
page.getByRole("status").filter({ hasText: "Saved" }),
@@ -83,5 +83,5 @@ test("a parakeet worker's device is configurable and persists", async ({
workers?: { appId?: string; config?: { device?: string } }[];
}>("test-settings.json");
expect(saved.workers?.[0].appId).toBe("parakeet");
- expect(saved.workers?.[0].config?.device).toBe("cuda:0");
+ expect(saved.workers?.[0].config?.device).toBe("CUDA0");
});
diff --git a/scripts/parakeet-stitch.mjs b/scripts/parakeet-stitch.mjs
@@ -33,7 +33,9 @@
// --overlap <sec> window overlap (PARAKEET_OVERLAP_SEC, default 6)
// --decoder <ctc|tdt> passed through to parakeet-cli (PARAKEET_DECODER)
// --lang <locale> passed through to parakeet-cli (PARAKEET_LANG)
-// --device <dev> compute device passed to parakeet-cli (PARAKEET_DEVICE)
+// --device <dev> compute device, exported to parakeet-cli as the
+// PARAKEET_DEVICE env var (parakeet.cpp has no --device
+// flag): "cpu", "CUDA0", "Vulkan1", ... (PARAKEET_DEVICE)
//
// Resumable, yt-dlp style: each window's raw parakeet-cli JSON is written to a
// per-audio work dir (".<audio>.parakeet/" beside the audio) as it completes,
@@ -193,8 +195,12 @@ async function transcribeWindow(opts, wav) {
const args = ["transcribe", "--model", opts.model, "--timestamps", "--json", "--input", wav];
if (opts.decoder) args.push("--decoder", opts.decoder);
if (opts.lang) args.push("--lang", opts.lang);
- if (opts.device) args.push("--device", opts.device);
- const { stdout } = await execFileP(opts.cli, args, { maxBuffer: 256 * 1024 * 1024 });
+ // parakeet.cpp's parakeet-cli has NO --device flag; the compute device is
+ // chosen solely via the PARAKEET_DEVICE env var (e.g. "cpu", "CUDA0",
+ // "Vulkan1"). Its default is the first GPU the ggml registry reports, so we
+ // MUST set the env var to honor a "cpu" (or specific-GPU) request.
+ const env = opts.device ? { ...process.env, PARAKEET_DEVICE: opts.device } : process.env;
+ const { stdout } = await execFileP(opts.cli, args, { env, maxBuffer: 256 * 1024 * 1024 });
return parseCliJson(stdout);
}