commit 6af4a7ffa092a38f1f29e4bee4de91b493fa85cc
parent 112c773e1c772408b0d5974c18381b19aa9dc760
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Sat, 8 Aug 2026 21:45:48 -0400
Surface the backfill pause where the work is watched
The hold already existed: backfillBatch's limit() re-reads
settings.backfill.enabled at DISPATCH time and returns 0, so the pool
idle-waits, the job stays alive and keeps its place, and resuming re-derives
nothing. It survives a restart because it is persisted rather than applied to a
live pool. What was missing was a reachable control -- the only switch was on
the Settings page, which is a strange place to look for a control over a job you
are watching run on the dashboard.
THIS REVISES A DOCUMENTED DECISION. BackfillSweepControls carried a note
explaining why it had one button and not two: at the default weight the lane is
already idle-only, so a manual pause looked redundant. That covered the wrong
hazard. Idle-only yielding handles 'get out of the transcription lane's way'; it
does nothing for 'this is a desktop someone is sitting at, and diarization pins
four cores for hours'. The revised rationale is in that file's header.
The button writes THE SAME FIELD as the Settings checkbox rather than a new
backfillPaused flag, so the two cannot drift -- one hold with two writers that
could disagree is worse than no button. The test asserts the field, not just the
label, for that reason, and asserts the button renders without a sweep armed
(a hand-clicked backfill-channel job deserves a pause too) and that pausing does
NOT disarm the sweep -- otherwise 'pause' quietly becomes 'stop'.
Verified: editor e2e backfill.spec 12/12 (11 pre-existing + 1 new), tsc clean in
common and editor, pnpm build clean -- the last because this is a client
component and the suite runs in dev mode, which never prerenders.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Diffstat:
4 files changed, 144 insertions(+), 11 deletions(-)
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -1,6 +1,7 @@
# Changelog
## [Unreleased]
+- **The backfill lane can now be paused from the dashboard, next to the work it pauses.** The hold itself is not new — the lane has always checked, every time it goes to pick up the next video, whether it is still switched on, and parked itself without ending if it wasn't. What was missing was anywhere sensible to flip it: the only switch lived on the Settings page, which is a strange place to look for a control over a job you are watching run. **Pause now sits beside Start Backfill Sweep, and it is a hold rather than a stop** — the job stays alive, keeps its place, and picks up within a few seconds when you resume, with nothing recomputed and nothing lost. It also survives a restart, because it is a remembered preference rather than something applied to a running pool. This matters most for speaker diarization, which is the one backfill that is genuinely heavy: it runs on the processor rather than the graphics card, pins several cores, and takes hours per long video — so on a desktop you are actually sitting at, "not right now" needs to be one click, for reasons the scheduler cannot possibly know about. **Pause and Stop are deliberately different.** Pause keeps the sweep armed and holds it; Stop ends the sweep, letting the video in flight finish first, and you re-arm it later. The button and the Settings checkbox are the same switch under the hood, so they can never disagree about whether the lane is running.
- **The archive can now put names to the speakers — two ways, both switched off until you decide they are worth the time.** Diarization already captures *that* the speaker changed; it records anonymous clusters, not people. Naming them is a separate pass, and there are two honest ways to do it that are not interchangeable. **From the audio:** the diarizer has already grouped every voice, so the model is only asked to put a name to each group given samples of what it said — **about one model call for a whole video**, and the speaker boundaries come from the audio rather than from a guess. **From the transcript alone:** for videos with no diarization, which today is nearly all of them, the model has to read the transcript to find the changes in the first place — roughly **one call per chunk**, about 194,000 across this archive, which is the same order as the AI digest sweep and would compete with it for the same card. So both are built and **neither is armed**: they are off, under a master switch that is also off, and the honest next step is a small measured pilot rather than committing 25–55 days of GPU time to the worse of the two lanes. **Quality is not oversold anywhere in the UI.** A name from the audio carries a confidence the model was actually asked for; a name from the transcript carries none, because the model was asked where the speaker changes, not how sure it is who anyone is — and inventing a number there would be exactly the wrong kind of confident. The transcript-only lane misfires on rapid back-and-forth, and on auto-caption channels there are no speaker turns to find at all. **The two lanes share one file, and the rule about that is the important part.** Naming from the audio may replace a transcript-only record — that is an upgrade, and it is queued for you automatically the moment a video gains diarization. The reverse can never happen: a transcript-only pass will not overwrite a record made from audio, even when forced, even when the audio and its diarization have since been deleted and that record is the only thing left that knows who was speaking. **Both lanes plug into the Backfill lane rather than being new machinery**, so they inherit the resource share, the corpus-wide sweep, the restart-survival and every indicator — the channel's Backfill card, `/actionable`, the dashboard and the widget all report them beside diarization, with reachable work and needs-its-media-back still counted separately. That separation matters more here than it did for diarization: only one video in this archive currently *has* the diarization the good lane needs, so a single “remaining” figure would be 73,000 videos of work no button can start. **Staleness is provenance, not age**, as everywhere else: a name is stale when the engine, the model or the prompt generation differs from what would be produced now — and, uniquely here, when the video has been **re-diarized underneath it**, because a cluster number means nothing except relative to the run that produced it. There is a Prompt generation number in Settings you can raise to force a corpus-wide redo; it cannot be set lower, because pinning it to a superseded prompt would freeze that prompt's output into the archive looking current.
- **Backfill is now a first-class thing the system knows about, with a lane of its own and a remainder you can see.** Every derived-data feature that lands runs into the same wall: the corpus that already exists does not have what it needs. Diarization hit it first — the input it wants is audio, which Clean-audio deletes once a video is transcribed — and the only catch-up was a per-channel button written for that one feature, with no way to ask how much of the corpus was missing it and no way to run it alongside new-video work without one starving the other. So a feature now **declares** what it needs and how to tell whether a video has it, and gets the lane, the resource share and the indicator for free. **The number this reports is split in two, and that is the whole design.** On this corpus 836 videos still have their media and about 76,270 do not — 91× more — so a single "remaining" figure would be dominated by work no button can start, and every channel would sit at the top of every list forever. Reachable work and needs-re-acquiring are therefore separate numbers on all four surfaces: the channel's new **Backfill** stage card, a new section on `/actionable` (with the second figure in its own column, exactly as "Est. reclaim" has), a dashboard instrument, and an optional widget strip. **The lane runs on its own queue**, so it is genuinely concurrent with transcription rather than sitting behind it, and how much of the machine it may take is one setting: **0 (the default) means idle-only** — full speed while transcription is quiet, standing aside the instant it isn't — and any value above 0 is a guaranteed share, floored at 1 so a small number is a slow lane and not a stopped one. Catch-up on a corpus that already exists must never slow down new arrivals, and the default enforces that rather than trusting it. There is a corpus-wide **Start Backfill Sweep** on the dashboard that persists its scope along with the flag and resumes itself after a restart. **Staleness now means provenance, not age.** A diarization sidecar records the models and clustering threshold that produced it, and the lane compares that against what the current settings *would* produce — so changing the threshold, which is the single knob most likely to make you want a re-run, finally shows as work instead of leaving the corpus looking finished. A field the record predates compares equal to today's default, so adding one does not invalidate everything on disk. **Re-downloading deleted media is off by default and is the part to read carefully.** With it on, each file is fetched, used, and deleted again immediately in a `finally` — whether the backfill succeeded, failed, or crashed — unless the video is marked "do not clean", and nothing starts at all when free disk is under the configured floor. On a disk at 97% a leak here fills it, which is why those four behaviours have e2e tests of their own.
- **Speaker diarization can now be captured while the audio still exists, and Clean audio will no longer delete audio out from under it.** Audio is the one input in this pipeline that goes away: the Clean-audio sweep removes it as soon as a video has a transcript, and nothing downstream can reconstruct it. Everything built on top of speaker turns — attribution, naming the speakers, badges, quote filtering — can be redone from a saved file at any time, but the turns themselves can only be extracted from audio. So this ships the perishable half only: a new **Diarization** section in Settings, off by default, that captures speaker-labelled time ranges to `diarization.json` beside each transcript, and a **Diarize speakers** button in every channel's Cleanup stage that backfills over whatever audio is still on disk. Turning capture on also arms a guard: the sweep now refuses to delete audio for a transcribed video that has no sidecar yet, reports it as `awaiting diarization`, and drops that video out of the channel's "Est. reclaim" so the estimate matches what the sweep will actually take. That hold is the point — it is what keeps the audio alive long enough to be captured — but it does mean enabling this stops reclaiming disk until diarization catches up, and the setting says so. **Inline-after-transcribe is a separate switch, and it is off by default for a measured reason.** Transcription on this box runs at 221 s per audio-hour on the GPU — 16.3× real time, averaged over 3,602 real videos from their own `transcribe-outcome.json` records — while diarization runs on the CPU at roughly 500–680, because no diarization model has been ported to ggml and neither ONNX nor PyTorch has a Vulkan compute path on Linux. Diarization is therefore 2–3× slower than the transcription it would follow, so running it inline makes the whole pipeline 3–4× slower and idles the GPU. The intended shape for a large batch is the opposite: turn capture on so the audio is held, leave inline off so transcription runs at full speed, then backfill. The clustering threshold defaults to **0.9**, not sherpa-onnx's own 0.5, and that is also measured: on a six-minute excerpt of a known two-person interview, 0.5 produced 22 speaker clusters and 0.9 produced 6, with the top two at 40%/40% of talk time — recognizably the two hosts. It still over-splits, which is why every sidecar records the engine, both model names, the library version and the threshold that produced it: over-splitting is the recoverable direction, and a later pass can re-cluster or re-run selectively without ever needing the audio back. The engine sits behind a path (`DIARIZE_BIN` → `scripts/diarize.mjs`, driving sherpa-onnx through `scripts/diarize-sherpa.py`) exactly as the parakeet wrapper does, so swapping in pyannote later is a settings change rather than a code change. A failing diarizer never fails a transcription, and never counts as "done" — the sweep keeps that video's audio. Full spike results, including the disk arithmetic and what is still unmeasured, are in `plans/diarization-spike-results.md`.
diff --git a/editor/app/jobs/actions.ts b/editor/app/jobs/actions.ts
@@ -274,6 +274,56 @@ export async function stopBackfillSweepAction(): Promise<BackfillSweepResult> {
}
}
+// PAUSE / RESUME the backfill lane.
+//
+// WHY THIS TOGGLES `backfill.enabled` RATHER THAN A NEW `backfillPaused` FIELD.
+// The digest lane has both a sweep flag and a separate `digestsPaused`; the
+// backfill lane has only the one switch, and that switch already behaves
+// exactly like a pause — `backfillBatch`'s limit() re-reads it at DISPATCH time
+// and returns 0, which makes the pool idle-WAIT rather than finish. The job
+// stays alive, logs "Backfill lane disabled in settings — holding", and resumes
+// within one 3s poll with nothing re-derived. Adding a second field would give
+// the same hold two writers that could disagree, and would silently divorce
+// this button from the "Run the backfill lane" checkbox in Settings.
+//
+// So: one field, two places to set it, and they cannot drift.
+//
+// This REVISES the "one button, not two" note in BackfillSweepControls.tsx.
+// That reasoning was that idle-only weighting made a manual hold unnecessary —
+// true for standing aside from transcription, but it gives an operator no way
+// to stop a long CPU-bound backfill for reasons of their own (the machine is a
+// desktop someone is using). See that file's header for the revised rationale.
+//
+// Needs no boot hook: the flag is consulted at dispatch, not applied to a live
+// pool, and it is already persisted in settings.json.
+export type BackfillPauseResult = { ok: boolean; error?: string };
+
+async function setBackfillLaneEnabled(
+ enabled: boolean,
+): Promise<BackfillPauseResult> {
+ try {
+ const current = getSettings();
+ if (current.backfill.enabled !== enabled) {
+ await writeSettings({
+ ...current,
+ backfill: { ...current.backfill, enabled },
+ });
+ }
+ } catch (e) {
+ return { ok: false, error: (e as Error).message };
+ }
+ revalidatePath("/jobs");
+ return { ok: true };
+}
+
+export async function pauseBackfillAction(): Promise<BackfillPauseResult> {
+ return setBackfillLaneEnabled(false);
+}
+
+export async function resumeBackfillAction(): Promise<BackfillPauseResult> {
+ return setBackfillLaneEnabled(true);
+}
+
// Retention scopes offered by the ClearLogsMenu. "all" clears every finished
// job's log; the day-scopes clear anything older than that. Running/queued jobs
// are never deleted (see pruneJobLogs).
diff --git a/editor/app/jobs/components/BackfillSweepControls.tsx b/editor/app/jobs/components/BackfillSweepControls.tsx
@@ -3,22 +3,42 @@
import { useEffect, useState, useTransition } from "react";
import { useRouter } from "next/navigation";
import {
+ pauseBackfillAction,
+ resumeBackfillAction,
startBackfillSweepAction,
stopBackfillSweepAction,
} from "../actions";
-// Arm / disarm the corpus-wide backfill sweep.
+// Arm / disarm the corpus-wide backfill sweep, and hold it without ending it.
//
-// One button, not two — unlike DigestSweepControls, which pairs a sweep switch
-// with a pause. The backfill lane has no separate pause because it does not need
-// one: at the default weight it is ALREADY idle-only, standing aside whenever
-// transcription works, and the lane switch in Settings is the hold. Adding a
-// second control that looks like the digest pause would invite the confusion
-// that pair exists to prevent.
+// TWO BUTTONS, AND THIS REVISES AN EARLIER DECISION. This component used to
+// carry one button and a note explaining why: at the default weight the lane is
+// ALREADY idle-only, standing aside whenever transcription works, so a manual
+// pause looked redundant, and a second control was thought to invite exactly the
+// confusion DigestSweepControls' pair exists to prevent.
//
-// Stopping drains rather than cancels: the video in flight finishes instead of
-// being thrown away, and the stop reaches the per-channel job the sweep is
-// waiting on rather than meaning "after this channel".
+// That reasoning covered the wrong hazard. Idle-only yielding handles "get out
+// of the transcription lane's way"; it does nothing for "this is a desktop
+// someone is sitting at, and diarization pins four cores for hours." Backfill
+// work is CPU-bound and long — a diarization pass over the corpus runs for
+// days — so the operator needs a hold for reasons the scheduler cannot see. The
+// hold already existed; it was only reachable from the Settings page, which is
+// a strange place to look for a control over a job you are watching run.
+//
+// So the pause is surfaced here, and it writes THE SAME FIELD the Settings
+// checkbox writes (settings.backfill.enabled) rather than a new `backfillPaused`
+// flag. One field, two places to set it, and they cannot drift. See
+// pauseBackfillAction in ../actions for why that is the right field.
+//
+// It is a REAL pause, not a stop: backfillBatch's limit() re-reads the flag at
+// dispatch and returns 0, so the pool idle-waits, the job stays alive, and
+// resuming re-derives nothing. It also survives a restart, because it is
+// persisted rather than applied to a live pool.
+//
+// Stopping the SWEEP, by contrast, drains rather than cancels: the video in
+// flight finishes instead of being thrown away, and the stop reaches the
+// per-channel job the sweep is waiting on rather than meaning "after this
+// channel". Pause when you want it back; stop when you don't.
export function BackfillSweepControls({
sweeping,
laneEnabled,
@@ -78,12 +98,32 @@ export function BackfillSweepControls({
>
{sweeping ? "Stop Backfill Sweep" : "Start Backfill Sweep"}
</button>
+ <button
+ type="button"
+ disabled={disabled}
+ aria-label={laneEnabled ? "pause backfill" : "resume backfill"}
+ onClick={() =>
+ run(laneEnabled ? pauseBackfillAction : resumeBackfillAction)
+ }
+ title={
+ laneEnabled
+ ? "Hold the backfill lane without ending anything. A running job idles at zero and keeps its place; nothing is re-derived on resume. Survives a restart."
+ : "Resume the backfill lane. A held job picks up within a few seconds — it was holding, not stopped."
+ }
+ className={
+ laneEnabled
+ ? "px-3 py-1.5 rounded-md border border-border text-sm hover:bg-muted disabled:opacity-50"
+ : "px-3 py-1.5 rounded-md bg-warning text-warning-foreground text-sm font-medium hover:bg-warning/90 disabled:opacity-50"
+ }
+ >
+ {laneEnabled ? "Pause Backfill" : "Resume Backfill"}
+ </button>
{sweeping && !laneEnabled && (
<span
aria-label="backfill lane off"
className="text-xs text-muted-foreground"
>
- lane off in Settings — the sweep is holding
+ the sweep is holding
</span>
)}
{error && (
diff --git a/editor/e2e/backfill.spec.ts b/editor/e2e/backfill.spec.ts
@@ -584,3 +584,45 @@ test("actionable shows the backfill section with both numbers", async ({
// does — never added to the count beside it.
await expect(section).toContainText("Needs media");
});
+
+// PAUSE IS A HOLD, NOT A STOP, and it is reachable from where the work is
+// watched. The hold has always existed — backfillBatch's limit() re-reads
+// settings.backfill.enabled at dispatch and returns 0, so the pool idle-waits
+// and the job keeps its place — but the only way to set it was the Settings
+// page, which is a strange place to look for a control over a job you are
+// watching run on the dashboard.
+//
+// The button writes THE SAME FIELD as that checkbox rather than a second
+// `backfillPaused` flag, so the two cannot drift; this asserts the field,
+// not just the label, for exactly that reason. It also asserts the button
+// renders WITHOUT a sweep armed — an operator pauses a hand-clicked
+// backfill-channel job too, not only a sweep.
+test("the dashboard pauses and resumes the backfill lane", async ({ page }) => {
+ test.setTimeout(SLOW);
+ await resetData("one-transcribe-channel-with-audio");
+ await writeSettings(backfillSettings());
+
+ const laneEnabled = async () =>
+ (
+ await readJson<{ backfill?: { enabled?: boolean } }>(
+ "test-settings.json",
+ ).catch(() => ({}) as { backfill?: { enabled?: boolean } })
+ ).backfill?.enabled ?? false;
+
+ await page.goto("/");
+ const pause = page.getByRole("button", { name: "pause backfill" });
+ await expect(pause).toBeEnabled({ timeout: 30_000 });
+ await pause.click();
+
+ await expect.poll(laneEnabled, { timeout: 30_000 }).toBe(false);
+
+ // The sweep flag is untouched: pausing the lane must not disarm a sweep, or
+ // "pause" would quietly become "stop" and days of queued work would need
+ // re-arming.
+ expect(await sweepFlag()).toBe(false);
+
+ const resume = page.getByRole("button", { name: "resume backfill" });
+ await expect(resume).toBeEnabled({ timeout: 30_000 });
+ await resume.click();
+ await expect.poll(laneEnabled, { timeout: 30_000 }).toBe(true);
+});