commit 8aa041621e9e4aca45b3ec85311ba994105dcc76
parent df58ee9d3a6f511549b086ee0bd76378b0538b37
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Sat, 22 Aug 2026 00:44:56 -0400
digest: rehome the lane's guards onto declared rules
laneBackfillKinds' own header names the blocker: digest carries
digestsPaused, the yield with its CPU-worker carve-out, spendCapUsd,
the remoteEnabled fail-fast, the engine probe() fail-fast and
duplicate-cluster sharing — "not one of them is expressible as
backfillLimit()'s single scalar". All six lived in digestBatch's
closure, so an arbiter wanting to dispatch digest would have had to
re-implement them from memory. That is why there are two schedulers.
common/controller/laneGuards.ts declares them in the two shapes a
dispatcher can act on, and the split is forced by WHEN they can be
asked:
* a PREFLIGHT before any pool exists — remoteEnabled and probe()
both throw today and must, because a configuration problem does
not resolve itself and parking a pool on one idles forever while
looking busy. It refuses a disabled metered lane WITHOUT probing.
* a GATE per pull, returning a HOLD and never a stop. runPool
idle-waits at a zero limit; next() returning null ends the job,
and every pause in this repo depends on that distinction.
The spend cap reads a per-RUN accumulator, so it is an argument rather
than a module-level reach — which is also what makes it testable.
Edge-logging stays in digestBatch (the rule is pure) and now tracks
the REASON rather than a boolean, so a lane going from yielding to
paused says so instead of staying silent.
laneSharesDuplicates() declares cluster sharing: a chapter list is
about what was SAID, so an aligned mirror gets one free; diarization
and attribution are grounded in one audio track and cannot share.
One ordering change, deliberate and documented: the remoteEnabled
refusal now happens after resolveDigestTarget rather than before, so
both fail-fasts are one call. Same error, same throw, a few sidecar
reads earlier — and the probe is still not reached, so a disabled lane
still costs no network call.
752 common tests (8 new), 31/31 digest+backfill e2e, tsc clean.
plans/unified-operations-model.md steps 2 and 4 marked done.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Diffstat:
5 files changed, 425 insertions(+), 76 deletions(-)
diff --git a/common/controller/digestBatch.ts b/common/controller/digestBatch.ts
@@ -27,7 +27,6 @@ import type { JobProgress } from "../jobs/registry";
import { isSectionFresh, type DigestSectionKind } from "../lib/digest";
import {
digestLaneFor,
- laneYieldsToTranscription,
} from "../lib/backfillKinds";
import { loadDigest } from "../lib/digest-server";
import {
@@ -36,7 +35,11 @@ import {
} from "./digestTarget";
import { CUES_JSON_FILENAME } from "../lib/videoStatus";
import { digestVideo } from "./digestVideo";
-import { transcriptionActivity } from "./digestYield";
+import {
+ digestGate,
+ digestPreflight,
+ laneSharesDuplicates,
+} from "./laneGuards";
import type { AutoQueueOrder } from "../jobs/autoQueuePolicy";
import { batchRecencyComparator } from "./batchRecency";
import {
@@ -151,11 +154,6 @@ export async function runDigestBatch(
const settings = getSettings();
const digestSettings = settings.digest;
const lane = opts.lane ?? "local";
- if (lane === "remote" && !digestSettings.remoteEnabled) {
- throw new Error(
- "The metered digest lane is disabled (settings.digest.remoteEnabled). Enable it in Settings before running it.",
- );
- }
// ONE derivation, shared with countMissingDigests and the channel snapshot's
// noDigest bucket — see digestTarget.ts. It must match what digestVideo will
// actually chunk with, or the batch's freshness check and the writer would
@@ -171,16 +169,21 @@ export async function runDigestBatch(
const dataDir = path.join(opts.paths.channelsDir, opts.channelSlug, "data");
- // Fail fast and loudly rather than 1,100 times in a row: an unreachable engine
- // is a configuration problem, and discovering it per-item wastes the log.
- if (!(await app.probe(config))) {
- throw new Error(
- `Digest engine ${app.id} is not reachable. ` +
- (app.lane === "local-gpu"
- ? "Is the ollama service running (systemctl status ollama)?"
- : "Is the claude CLI installed and on PATH (set CLAUDE_BIN)?"),
- );
- }
+ // BOTH fail-fasts, asked as ONE declared rule. Fail fast and loudly rather
+ // than 1,100 times in a row: an unreachable engine is a configuration
+ // problem, and discovering it per-item wastes the log. It still THROWS here,
+ // byte-identically — the message is the rule's, and the rule is what a future
+ // dispatcher asks instead of re-implementing this from memory.
+ //
+ // Deliberately after resolveDigestTarget, not before: `app` and `config` come
+ // out of it, and probing an engine the run was never going to use would be a
+ // network call for nothing.
+ const preflight = await digestPreflight({
+ settings,
+ lane,
+ app: { id: app.id, lane: app.lane, probe: () => app.probe(config) },
+ });
+ if (!preflight.ok) throw new Error(preflight.error);
const allDirs = await readdir(dataDir).catch(() => [] as string[]);
const onDisk = new Set(allDirs);
@@ -188,8 +191,16 @@ export async function runDigestBatch(
? opts.ids.filter((id) => onDisk.has(id))
: allDirs;
+ // WHETHER A LANE CAN SHARE A DUPLICATE'S OUTPUT IS DECLARED, not assumed from
+ // the fact that this file is digestBatch. A chapter list is about what was
+ // SAID, so an aligned mirror gets the canonical member's digest for free
+ // (~11% of the sweep); diarization and attribution are grounded in one audio
+ // track and one set of cue timings and can share nothing. Asking the rule is
+ // what lets a dispatcher that does not know which operation it is holding get
+ // this right. An explicit `useClusters: false` still wins — this is the
+ // default, not a lock.
const clusterPlan =
- opts.useClusters === false
+ opts.useClusters === false || !laneSharesDuplicates(digestLaneFor(app.lane))
? null
: (opts.clusterPlan ?? (await buildDigestClusterPlan(opts.paths)));
@@ -461,9 +472,11 @@ export async function runDigestBatch(
}
};
- // Edge-triggered, so the yield logs twice per contention window rather than
- // once per poll over a multi-week sweep.
- let yielding = false;
+ // Edge-triggered, so a hold logs twice per contention window rather than once
+ // per poll over a multi-week sweep. It tracks the REASON rather than a
+ // boolean: a lane that goes from yielding to paused has changed state and
+ // should say so, where a boolean would stay `true` and stay silent.
+ let heldReason: "paused" | "yield" | "spend-cap" | null = null;
const NEVER = new AbortController().signal;
// Local lane: 1. The GPU is the bottleneck and a second concurrent generation
@@ -480,56 +493,38 @@ export async function runDigestBatch(
next,
run: runOne,
limit: () => {
- // Re-read at DISPATCH time so a pause takes effect within one poll and
- // survives a restart with no boot hook. Returning 0 makes runPool
- // idle-wait, which is a pause; returning null from next() would END the
- // batch, which is not.
- if (getSettings().digest.digestsPaused) return 0;
- // Step aside for whisper. Same mechanism as the pause and for the same
- // reason it works: a zero limit HOLDS the pool instead of ending the
- // batch, so the sweep resumes the moment the card is free without
- // re-deriving anything. Only the local (GPU) lane yields — the metered
- // lane is network-bound and competes for nothing here.
- // WHICH LANE YIELDS IS NOW DECLARED, not re-tested here. digestLaneFor
- // maps the engine to the digest operation's registered lane, and
- // laneYieldsToTranscription reads `contendsFor: "gpu"` off it. Same
- // decision, same behaviour — but stated once, next to the queue key it
- // belongs with, so a future GPU operation cannot answer it differently.
- if (
- laneYieldsToTranscription(digestLaneFor(app.lane)) &&
- getSettings().digest.yieldToTranscription
- ) {
- const activity = transcriptionActivity();
- if (activity.busy) {
- // Logged on the EDGE only. An operator watching a sweep sit at zero
- // throughput has to be able to tell yielding from wedged, but a line
- // per poll would bury the job log over a multi-week run.
- if (!yielding) {
- yielding = true;
- log(
- `Yielding the GPU to transcription (${activity.reason}); the digest lane will resume when it is free.`,
- );
- }
- return 0;
- }
- if (yielding) {
- yielding = false;
- log("Transcription finished; resuming the digest lane.");
- }
- }
- if (
- app.metered &&
- digestSettings.spendCapUsd > 0 &&
- result.costUsd >= digestSettings.spendCapUsd
- ) {
- if (!result.spendCapped) {
- result.spendCapped = true;
- log(
- `Spend cap reached ($${result.costUsd.toFixed(2)} of $${digestSettings.spendCapUsd.toFixed(2)}) — parking the metered lane.`,
- );
+ // EVERY GUARD, ASKED AS ONE DECLARED RULE. Settings are re-read at
+ // DISPATCH time so a pause takes effect within one poll and survives a
+ // restart with no boot hook (the downloadsPaused pattern), and the answer
+ // is a HOLD: returning 0 makes runPool idle-wait, where returning null
+ // from next() would END the batch. Every pause in this repo depends on
+ // that distinction.
+ const gate = digestGate({
+ settings: getSettings(),
+ appLane: app.lane,
+ metered: app.metered,
+ // Per-RUN, not per-settings: the spend cap is a per-job ceiling, so it
+ // is read off the live accumulator rather than from disk.
+ costUsd: result.costUsd,
+ });
+ if (gate.hold) {
+ // LOGGED ON THE EDGE ONLY, and the edge is tracked here rather than in
+ // the rule: the rule is pure and a line per poll would bury a job log
+ // over a multi-week run. An operator watching a lane sit at zero
+ // throughput still has to be able to tell yielding from wedged.
+ if (heldReason !== gate.reason) {
+ heldReason = gate.reason;
+ log(gate.message);
}
+ // The one piece of state a caller keeps: the result flag surfaces a
+ // metered run that stopped early for money rather than for work.
+ if (gate.reason === "spend-cap") result.spendCapped = true;
return 0;
}
+ if (heldReason === "yield") {
+ log("Transcription finished; resuming the digest lane.");
+ }
+ heldReason = null;
return concurrency;
},
signal: opts.signal ?? NEVER,
diff --git a/common/controller/laneGuards.test.ts b/common/controller/laneGuards.test.ts
@@ -0,0 +1,152 @@
+import { test } from "node:test";
+import assert from "node:assert/strict";
+import { defaultDigest, type SiteSettings } from "../lib/settings";
+import { digestLaneFor } from "../lib/backfillKinds";
+import { BACKFILL_QUEUE } from "../lib/queueKeys";
+import {
+ digestGate,
+ digestPreflight,
+ laneSharesDuplicates,
+} from "./laneGuards";
+
+// Run with: node_modules/.bin/tsx --test common/controller/laneGuards.test.ts
+//
+// The yield branch is deliberately NOT unit-tested here: it consults the live
+// worker pool and job registry through transcriptionActivity(), which fails
+// open to "nothing running" in a bare process. Stubbing those would test the
+// stub. The yield is covered end to end by digest.spec.ts and backfill.spec.ts,
+// and the DECISION it turns on — that only a gpu-contending lane yields — is
+// asserted structurally below.
+
+function settingsWith(patch: Partial<SiteSettings["digest"]>): SiteSettings {
+ return { digest: { ...defaultDigest(), ...patch } } as SiteSettings;
+}
+
+test("a pause holds the lane, and beats every other reason", () => {
+ const gate = digestGate({
+ // Paused AND over a spend cap: an operator who paused the lane must be told
+ // about the pause, not sent to look at a billing setting.
+ settings: settingsWith({ digestsPaused: true, spendCapUsd: 1 }),
+ appLane: "remote-api",
+ metered: true,
+ costUsd: 99,
+ });
+ assert.equal(gate.hold, true);
+ assert.equal(gate.hold && gate.reason, "paused");
+});
+
+test("the spend cap parks the METERED lane only", () => {
+ const settings = settingsWith({ spendCapUsd: 5 });
+ const over = digestGate({
+ settings,
+ appLane: "remote-api",
+ metered: true,
+ costUsd: 5,
+ });
+ assert.equal(over.hold, true);
+ assert.equal(over.hold && over.reason, "spend-cap");
+ assert.match(over.hold ? over.message : "", /\$5\.00 of \$5\.00/);
+
+ // A local run costs nothing per call, so the cap has nothing to cap.
+ assert.equal(
+ digestGate({ settings, appLane: "local-gpu", metered: false, costUsd: 500 })
+ .hold,
+ false,
+ );
+ // And the cap is opt-in: 0 means no cap, not "cap at zero".
+ assert.equal(
+ digestGate({
+ settings: settingsWith({ spendCapUsd: 0 }),
+ appLane: "remote-api",
+ metered: true,
+ costUsd: 500,
+ }).hold,
+ false,
+ );
+});
+
+test("under the cap the lane goes, at the cap it holds", () => {
+ const settings = settingsWith({ spendCapUsd: 5 });
+ const input = { settings, appLane: "remote-api" as const, metered: true };
+ assert.equal(digestGate({ ...input, costUsd: 4.99 }).hold, false);
+ // >=, not >: the cap is a ceiling on cumulative spend, so reaching it stops.
+ assert.equal(digestGate({ ...input, costUsd: 5 }).hold, true);
+});
+
+test("only a gpu-contending lane can yield", () => {
+ // The decision the yield turns on, asserted where it is DECLARED rather than
+ // where it is consumed. A network lane competes for nothing local, so parking
+ // it would idle a lane that costs nothing to keep running.
+ assert.equal(digestLaneFor("local-gpu").contendsFor, "gpu");
+ assert.equal(digestLaneFor("remote-api").contendsFor, "network");
+ // With yielding switched off entirely, even the GPU lane goes.
+ assert.equal(
+ digestGate({
+ settings: settingsWith({ yieldToTranscription: false }),
+ appLane: "local-gpu",
+ metered: false,
+ costUsd: 0,
+ }).hold,
+ false,
+ );
+});
+
+test("preflight refuses the metered lane while it is disabled, without probing", async () => {
+ let probed = false;
+ const verdict = await digestPreflight({
+ settings: settingsWith({ remoteEnabled: false }),
+ lane: "remote",
+ app: {
+ id: "claude",
+ lane: "remote-api",
+ probe: async () => {
+ probed = true;
+ return true;
+ },
+ },
+ });
+ assert.equal(verdict.ok, false);
+ assert.match(verdict.ok ? "" : verdict.error, /remoteEnabled/);
+ // A disabled lane must not cost a network call to find out it is disabled.
+ assert.equal(probed, false);
+});
+
+test("preflight refuses an unreachable engine, and names the fix", async () => {
+ const local = await digestPreflight({
+ settings: settingsWith({}),
+ lane: "local",
+ app: { id: "ollama", lane: "local-gpu", probe: async () => false },
+ });
+ assert.equal(local.ok, false);
+ assert.match(local.ok ? "" : local.error, /ollama service/);
+
+ const remote = await digestPreflight({
+ settings: settingsWith({ remoteEnabled: true }),
+ lane: "remote",
+ app: { id: "claude", lane: "remote-api", probe: async () => false },
+ });
+ assert.equal(remote.ok, false);
+ assert.match(remote.ok ? "" : remote.error, /CLAUDE_BIN/);
+});
+
+test("preflight passes a reachable, enabled lane", async () => {
+ const verdict = await digestPreflight({
+ settings: settingsWith({ remoteEnabled: true }),
+ lane: "remote",
+ app: { id: "claude", lane: "remote-api", probe: async () => true },
+ });
+ assert.equal(verdict.ok, true);
+});
+
+test("only the digest lanes share a duplicate's output", () => {
+ // A chapter list is about what was SAID, so an aligned mirror gets the
+ // canonical member's digest for free. Diarization and attribution are
+ // grounded in one audio track and one set of cue timings and can share
+ // nothing — a mirror's clock is not the same clock.
+ assert.equal(laneSharesDuplicates(digestLaneFor("local-gpu")), true);
+ assert.equal(laneSharesDuplicates(digestLaneFor("remote-api")), true);
+ assert.equal(
+ laneSharesDuplicates({ queueKey: BACKFILL_QUEUE, contendsFor: "cpu" }),
+ false,
+ );
+});
diff --git a/common/controller/laneGuards.ts b/common/controller/laneGuards.ts
@@ -0,0 +1,196 @@
+import type { SiteSettings } from "../lib/settings";
+import {
+ digestLaneFor,
+ laneYieldsToTranscription,
+ type BackfillLane,
+} from "../lib/backfillKinds";
+import type { DigestLane } from "../lib/digest";
+import { transcriptionActivity } from "./digestYield";
+
+// THE LANE'S RULES, declared once, in the two shapes a dispatcher can act on.
+//
+// This exists because of a sentence in laneBackfillKinds' own header:
+//
+// "digest carries digestsPaused, the yield-to-transcription carve-out (with
+// its CPU-worker exemption), spendCapUsd on the metered lane, the
+// remoteEnabled fail-fast, the engine probe() fail-fast … and not one of
+// them is expressible as backfillLimit()'s single scalar."
+//
+// That is the whole reason there are two schedulers for one kind of work. Every
+// one of those guards lived inside digestBatch's own closure, where nothing
+// else could ask about it — so an arbiter that wanted to dispatch digest would
+// have had to re-implement all six, correctly, from memory. They live here now,
+// and digestBatch calls them, so there is exactly one definition and no second
+// opinion to drift.
+//
+// TWO SHAPES, AND THE SPLIT IS FORCED BY WHEN THEY CAN BE ASKED:
+//
+// * a PREFLIGHT is answered before any pool exists. `remoteEnabled` and the
+// engine `probe()` both THROW today, before runPool is constructed, and
+// they have to: an unreachable engine is a configuration problem, and
+// discovering it per-item wastes 1,100 log lines. A dispatcher needs to ask
+// that question before it commits a job, not after.
+// * a GATE is answered at DISPATCH time, on every pull. It returns a HOLD,
+// never a stop — runPool idle-waits at a zero limit, whereas next()
+// returning null ends the job. Every pause in this repo depends on that
+// distinction, and stating it here is what keeps the next dispatcher from
+// re-deciding it.
+//
+// Nothing here reads disk or mutates settings. The spend cap is the one rule
+// that needs per-RUN state, and it takes it as an argument rather than reaching
+// for a module-level accumulator — see DigestGateInput.costUsd.
+
+// A gate's answer. `hold` carries the reason in words, because an operator
+// watching a lane sit at zero throughput has to be able to tell YIELDING from
+// WEDGED, and those look identical from outside.
+export type LaneGate =
+ | { hold: false }
+ | {
+ hold: true;
+ // Machine-readable, for a surface that wants to branch.
+ reason: "paused" | "yield" | "spend-cap";
+ // Human-readable, logged on the EDGE only — a line per poll would bury a
+ // job log over a multi-week run.
+ message: string;
+ };
+
+export const GO: LaneGate = { hold: false };
+
+export type PreflightVerdict = { ok: true } | { ok: false; error: string };
+
+// --- Digest ------------------------------------------------------------------
+
+export type DigestPreflightInput = {
+ settings: SiteSettings;
+ // Which lane was ASKED for, not which one is configured.
+ lane: "local" | "remote";
+ app: {
+ id: string;
+ lane: DigestLane;
+ // The engine's own reachability check. Awaited here so the caller has one
+ // place to ask "can this lane run at all".
+ probe: () => Promise<boolean>;
+ };
+};
+
+// Can this digest lane run at all?
+//
+// Both answers are refusals rather than holds, and deliberately: a disabled
+// metered lane and an unreachable engine are configuration problems that will
+// not resolve on their own, so parking the pool on them would idle forever
+// while looking busy.
+export async function digestPreflight(
+ input: DigestPreflightInput,
+): Promise<PreflightVerdict> {
+ if (input.lane === "remote" && !input.settings.digest.remoteEnabled) {
+ return {
+ ok: false,
+ error:
+ "The metered digest lane is disabled (settings.digest.remoteEnabled). Enable it in Settings before running it.",
+ };
+ }
+ if (!(await input.app.probe())) {
+ return {
+ ok: false,
+ error:
+ `Digest engine ${input.app.id} is not reachable. ` +
+ (input.app.lane === "local-gpu"
+ ? "Is the ollama service running (systemctl status ollama)?"
+ : "Is the claude CLI installed and on PATH (set CLAUDE_BIN)?"),
+ };
+ }
+ return { ok: true };
+}
+
+export type DigestGateInput = {
+ settings: SiteSettings;
+ // The engine's lane, which decides both the queue key and whether this run
+ // contends for the GPU. Passed rather than re-derived so a caller that has
+ // already resolved its app cannot disagree with one that has not.
+ appLane: DigestLane;
+ metered: boolean;
+ // Cumulative metered spend SO FAR IN THIS RUN. The cap is per-job, so it
+ // cannot be read from settings alone — and passing it in is what keeps this
+ // function pure enough to test.
+ costUsd: number;
+};
+
+// Should the digest lane hold right now?
+//
+// Order matters and is the order it was in: an explicit pause beats a yield
+// beats a spend cap. A paused lane that reported "yielding" would send an
+// operator to look at the transcription queue for a switch they turned off
+// themselves.
+export function digestGate(input: DigestGateInput): LaneGate {
+ const digest = input.settings.digest;
+
+ // Re-read at DISPATCH time so a pause takes effect within one poll and
+ // survives a restart with no boot hook — the downloadsPaused pattern.
+ if (digest.digestsPaused) {
+ return {
+ hold: true,
+ reason: "paused",
+ message: "Digests are paused; the lane is holding.",
+ };
+ }
+
+ // Step aside for whisper. WHICH LANE YIELDS IS DECLARED, not re-tested here:
+ // digestLaneFor maps the engine to the digest operation's registered lane and
+ // laneYieldsToTranscription reads `contendsFor: "gpu"` off it, so a future GPU
+ // operation cannot answer this differently.
+ //
+ // The CPU carve-out lives one level down, in transcriptionActivity(): a
+ // worker pinned to `device: "cpu"` is not GPU contention, and treating it as
+ // such once stalled the digest lane for transcription competing for zero
+ // shaders. A worker with NO device set still triggers the yield — the unknown
+ // case fails safe.
+ if (
+ laneYieldsToTranscription(digestLaneFor(input.appLane)) &&
+ digest.yieldToTranscription
+ ) {
+ const activity = transcriptionActivity();
+ if (activity.busy) {
+ return {
+ hold: true,
+ reason: "yield",
+ message: `Yielding the GPU to transcription (${activity.reason}); the digest lane will resume when it is free.`,
+ };
+ }
+ }
+
+ if (
+ input.metered &&
+ digest.spendCapUsd > 0 &&
+ input.costUsd >= digest.spendCapUsd
+ ) {
+ return {
+ hold: true,
+ reason: "spend-cap",
+ message:
+ `Spend cap reached ($${input.costUsd.toFixed(2)} of ` +
+ `$${digest.spendCapUsd.toFixed(2)}) — parking the metered lane.`,
+ };
+ }
+
+ return GO;
+}
+
+// --- What a lane does with a duplicate cluster -------------------------------
+
+// Whether an operation's output can be SHARED between two recordings of the
+// same event rather than generated twice.
+//
+// Digest can: a chapter list is about what was said, so an aligned mirror gets
+// the canonical member's digest for free (worth ~11% of the sweep). Diarization
+// and attribution cannot — they are grounded in a specific audio track and a
+// specific set of cue timings, and a mirror's clock is not the same one.
+//
+// Declared here rather than inferred from an id, so a dispatcher can ask the
+// question without knowing which operation it is holding. `useClusters: false`
+// on a run still overrides it; this is the DEFAULT, not a lock.
+export function laneSharesDuplicates(lane: BackfillLane): boolean {
+ return (
+ lane.queueKey === digestLaneFor("local-gpu").queueKey ||
+ lane.queueKey === digestLaneFor("remote-api").queueKey
+ );
+}
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -10,6 +10,7 @@
- **The auto-queue page is a console for all four pipelines, not two stacked copies of one.** It rendered auto-transcribe and auto-download as two identical full-height panels and had no room for the other two pipelines at all — so "the digest lane is idle because transcription has the GPU" was a fact you could only assemble by visiting three pages. The page now opens on a **rail**: one line per pipeline, always present, whichever lane you are looking at. Each line carries a band showing what that pipeline has done, what it can do **right now**, what is waiting on an operation upstream, what is held by a gate, and what has lost its media — five separate figures that are never added together, because on this archive they are ninety-one times apart and one "remaining" number would say the same thing about a finished lane and a lane that cannot start. Exactly one saturated colour appears on the page, and it marks work that can happen right now; everything else is texture, distinguished by **fill pattern before hue** — solid, hatched, dotted, hollow — because a four-colour scheme was measured here and none of them survives colour-blindness on every pair. Each band is drawn against **its own** population and says so, since a digest's eligible videos are genuinely a different set from a diarization's. Where an archive cannot yet say how many videos a pipeline is done with, the band draws as an empty outline and says *coverage unknown* rather than showing 0%. Below the rail, a **dropdown** puts one lane in full view: the rules ladder for the two runners, and for digest and backfill the state, both switches, and what is running. The rail is what makes the switch safe — you never have to change lanes to learn that this one is idle because another one is, which on this archive is the normal case rather than the exception.
- **Order and Reach sit at the foot of the lane, and now cover every pipeline.** *Newest first* used to be a control on two of the four lanes, halfway up a form. It is now the same block on all four, with the trade-off spelled out beneath it, and it is joined by **Reach** — *Within each channel* or *Across all channels* — which is disabled while the order is left at *Listed*, because a reach with no order to apply is not a setting. The digest and backfill lanes keep **both** of their switches, under the names they have everywhere else: the **sweep** decides whether there is a corpus-wide pass at all, and the **pause** decides whether the lane consumes it. They are deliberately not merged into one button: one is cheap to undo and the other costs a week of GPU time.
- **A runner is a lane, not a job, and the active-jobs screen finally says so.** Each always-on runner took a full bordered card — the same card a channel with real work in it gets — to say "running", which pushed the jobs that actually have progress bars below the fold. The runners are now a **strip** at the top: one line each, with the lane's state and, when it is not working, *why* — "nothing pending", "sweep armed, lane paused", "no enabled worker". Digest and backfill are on the strip too, so all four pipelines are visible at once from a screen that previously knew about two. Every control a runner had is still there, on its line. A screen with no jobs at all now shows the strip rather than only the words "No active jobs" — "every lane is idle and here is why" is the answer that screen was previously unable to give.
+- **The digest lane's six safety rules are now written down where anything can read them.** They lived inside one function's closure — the global pause, standing aside for transcription (and the carve-out that stops a CPU-pinned worker counting as GPU contention), the metered spend cap, the two refusals for a disabled lane and an unreachable engine, and the rule that lets an aligned re-upload borrow a digest instead of generating one. Nothing outside that function could ask about any of them, which is exactly why this archive ended up with two separate schedulers for one kind of work: anything else wanting to run a digest would have had to reimplement all six from memory, correctly. They are one module now, and the digest runner calls it, so there is one definition and no second opinion to drift from it. Nothing changes about what happens — a paused lane still holds rather than stopping, a disabled metered lane still refuses without making a network call to find out, and the log still says which of the reasons is holding you up, printed once when the state changes rather than once per poll for a week.
- **Newest-first could quietly hand a video another channel's upload date.** The date lookup's first and cheapest layer scans the transcript index, which is keyed by the id the *platform* reports; every video the auto-queue actually asks about is named by its **folder**, and on this archive those two disagree for about one video in seven — a Rumble folder is named for the URL, while its recorded id is the embed id. Matched across the whole corpus, one channel's folder name could collide with a different channel's recorded id and inherit its date, which is a *wrong* answer rather than a missing one, and wrong dates are exactly what an ordering setting cannot survive. The lookup is now scoped to the channel the video belongs to, so a folder name that is not an id in its own channel falls through to reading the date out of the folder it actually names. Separately, the first pass over a very large backlog can only date so many videos at once; it now says so once in the log, because the remainder sorting to the back and being picked up on the next pass is the design, not a fault.
- **The auto-queue can be told to do the newest uploads first, and it now genuinely does.** Both runners always took the first video off a rule's pile, and that pile's order came straight from the channel snapshot, which sorts most buckets **alphabetically by video id** — arbitrary for YouTube ids, and oldest-first for the date-prefixed folder names some sites use. So when a channel uploaded today, nothing made that video jump the nine-thousand-video backlog in front of it; the only reason auto-download roughly worked was that its one bucket happens to be left in playlist order. There is now an **Order** setting per runner — *Listed order* (what you have today, and still the default), *Newest first*, *Oldest first*. It sorts the videos **inside** each rule, across every channel and bucket that rule claims; the rule list still decides which rule goes first, because that is what the rule list is for. For a straight newest-first archive, use one catch-all rule. Worth knowing before you switch it on: under *Newest first* the retry and partial-download buckets lose their head start, so a half-finished download can end up waiting behind fresh work. The page says so next to the setting.
- **Working out how recent 79,000 videos are turned out to be nearly free, once we stopped guessing where the dates were.** The obvious source — reading each video's metadata file — is about six and a half minutes and 41 GB of reading, on every scheduling decision, which is a non-starter. The transcript index already holds a date per video in a form that can be scanned without decoding anything: **78,583 videos in well under a second**. That covers the corpus, but it turned out **not** to cover the videos auto-transcribe actually queues, because the index only holds videos that already *have* a transcript and auto-transcribe's whole job is the ones that don't — of the 870 videos genuinely pending here, it knew the date of **122**. The gap is closed by reading the last 8 KB of each remaining video's metadata file, where the upload date happens to sit: **868 of the 870, at a fifth of a millisecond each**, and remembered afterwards so it is paid once rather than every few seconds. Videos not downloaded yet have no date anywhere on disk at all, so auto-download estimates one from the video's position in the channel's newest-first listing; those show with a `≈`, and a brand-new upload with nothing dated above it goes to the front, which is the entire point. Anything still undatable sorts to the back rather than disappearing, and a missing or busy index degrades to the old ordering instead of stopping the runner.
diff --git a/plans/unified-operations-model.md b/plans/unified-operations-model.md
@@ -92,18 +92,23 @@ change** — every one of them moves live numbers on a 78,000-video corpus.
it is 99.87% of the corpus — the replacement must keep that refusal, and note that
`blocked` now removes untranscribed videos from the count, which will change the
published coverage percentage.
-2. **A leaf can name an operation.** `buildPendingByLeaf` reads `snapshot.backfill[op].ids`.
- At this point the tree can *express* backfill work without anything dispatching it.
+2. **A leaf can name an operation.** — **DONE.** `AutoQueueMatch.operation` beside `bucket`,
+ drawing from a separate `ChannelWork.operations` id space, with the claim key widened to
+ `${operation}\0${id}` so digest and diarization cannot steal each other's work. The tree
+ can *express* operation work; nothing dispatches it yet.
3. **One arbiter per lane, not per subsystem.** A single runner reads the tree, groups
selected work by `lane.queueKey`, and dispatches one job per key. This is where
`digestSweep` and `backfillSweep` finally merge — they are the same loop with different
constants.
-4. **Move the guards onto declared lane rules.** The blocker is that `digestBatch.limit()`
- carries guards no scalar share can express: `digestsPaused`, `yieldToTranscription` plus
- its CPU-worker carve-out, `spendCapUsd` on the metered lane, the `remoteEnabled`
- fail-fast, the engine `probe()` fail-fast, and duplicate-cluster sharing. Each needs a
- home in the lane rule *before* step 3 can carry digest. `contendsFor` is the first of
- these and already exists.
+4. **Move the guards onto declared lane rules.** — **DONE.** `common/controller/laneGuards.ts`
+ holds all six, in the two shapes a dispatcher can act on: a **preflight** answered before
+ any pool exists (`remoteEnabled`, the engine `probe()` — both still throw, because a
+ configuration problem will not resolve itself and parking on it would idle forever while
+ looking busy) and a **gate** answered per pull (`digestsPaused`, the yield with its
+ CPU-worker carve-out, `spendCapUsd`), which returns a HOLD and never a stop.
+ `laneSharesDuplicates()` covers cluster sharing. `digestBatch` calls all of them, so
+ behaviour is unchanged and there is one definition rather than two. 8 unit tests; the
+ yield branch stays covered end to end because it reads live pool state.
5. **One pause model.** `settings.backfill.enabled`, `digest.digestsPaused`,
`transcriptionsPaused` and `downloadsPaused` become node state on the tree. Keep the
property all four already have and that makes them safe: a pause returns `limit() === 0`,