commit 5d1a28630eb51cf216d81fa4f104e58dc9f64349
parent 1476904975d1d186034e0f0c6a84bce9279d3adb
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Mon, 27 Jul 2026 00:22:37 -0400
Digest: ship the measured-best defaults, and make the counters honest
Three things that each independently made a validation run misleading.
1. The shipped defaults were the configuration the bake-off measured as
WORST. round2.md scored qwen2.5:7b@8192/chunk-local best (11.8%
zero-yield, 13.37 chapters/h, 24:13 worst gap) against @16384/absolute
(33.3%, 8.19, 1:27:48) — but DEFAULT_DIGEST_TIMESTAMP_MODE was
"absolute" and DEFAULT_DIGEST_NUM_CTX 16384, so a fresh install ran the
loser. Now chunk-local / 8192 / 600 cues.
PAIRED WITH A PROMPT_VERSION BUMP (1 -> 2), which is the whole point.
digestPromptVariant derives its string RELATIVE to the defaults, so a
value equal to the default yields `undefined`. Flipping the defaults
alone would have made the 39 pilot digests — written under
absolute/1200 with promptVariant: undefined — compare FRESH under the
new defaults, freezing bake-off-losing output into the corpus forever.
isSectionFresh compares promptVersion, so the bump invalidates them
explicitly. Cost: 39 regenerations. Tests pin the pairing.
2. The digest pending count did not mean what it said. buckets.noDigest
asked only `engine && hasItems`, while its label and stageStatus.ts
claimed a full identity check, so after any config change the stage
read "All digested" while countMissingDigests reported the whole
channel stale — a disagreement of an entire channel. New
resolveDigestTarget() is the single derivation; batch runner, counter
and snapshot all use it. Per-video cost is zero extra I/O. Coverage by
engine stays identity-blind on purpose.
3. countMissingDigests always resolved the LOCAL app, so a metered-lane
job's progress bar was sized against the wrong engine's identity. It
takes the lane now; both call sites already knew it.
Also §1 of the plan: measured --near sensitivity for duplicate detection.
0.6 was rejecting real cross-platform mirrors in bulk — at 0.35 confirmed
pairs go 2,717 -> 7,329 and aligned mirrors 1,599 -> 3,952, a strict
superset, and all 4,456 new clusters inspect as genuine (99.3%
cross-platform AND cross-channel, 96% byte-identical titles, scores in a
tight 0.45-0.60 band just under the old cutoff). Cause: the two sides are
different ASR engines. Recorded in FACTS.md/STATE.md; the default is NOT
changed here — it alters what every built site asserts, so it belongs in
its own commit. transcripts/duplicates.json restored byte-identical.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Diffstat:
12 files changed, 456 insertions(+), 136 deletions(-)
diff --git a/common/controller/channelSnapshot.ts b/common/controller/channelSnapshot.ts
@@ -24,6 +24,8 @@ import {
} from "../lib/availability-server";
import { isDoNotClean } from "../lib/doNotClean-server";
import { loadDigest } from "../lib/digest-server";
+import { isSectionFresh } from "../lib/digest";
+import { resolveDigestTarget } from "./digestTarget";
import { isExcludedFromTruncatedCheck } from "../lib/excludeTruncatedCheck-server";
import { loadDownloadOutcome } from "../lib/downloadOutcome-server";
import type { Paths } from "../lib/paths";
@@ -148,11 +150,17 @@ export type ChannelSnapshot = {
// (retry-bucket with forceCookies). Optional: older snapshots lack it;
// readers must default to [].
needsCookies: string[];
- // Transcribed videos with no (or an outdated-shape) ai-digest.json — the AI
- // digest layer's work list, and the denominator for corpus coverage during
- // the backfill. Only TRANSCRIBED videos are listed: a video without a
- // transcript is a transcription problem, not a digest one. Optional: older
- // snapshots lack it; readers must default to [].
+ // Transcribed videos whose digest is missing or STALE against the local
+ // lane's current freshness target (engine, requested model, prompt version,
+ // prompt shape, context hash) — the AI digest layer's work list, and the
+ // denominator for corpus coverage during the backfill. This is the same
+ // target countMissingDigests and runDigestBatch compute, deliberately: a
+ // bucket that silently meant something weaker than its name is how a sweep
+ // reports "nothing to do" on a corpus that needs redoing.
+ //
+ // Only TRANSCRIBED videos are listed: a video without a transcript is a
+ // transcription problem, not a digest one. Optional: older snapshots lack
+ // it; readers must default to [].
noDigest: string[];
};
undownloadedIds: string[];
@@ -458,6 +466,24 @@ export async function generateChannelSnapshot(
const supersededAutoSubs: string[] = [];
const noDigest: string[] = [];
const digestEngines: Record<string, number> = {};
+ // The SAME freshness target countMissingDigests and runDigestBatch use, so
+ // the Digest stage's count and the batch runner's progress target cannot
+ // disagree. This bucket used to ask only "is there an ai-digest.json with
+ // items?", which meant that after any prompt/model/context change the stage
+ // read "All digested." while the batch reported the whole channel as stale.
+ //
+ // Affordable in this hot path because the per-video cost is ZERO extra I/O:
+ // the digest sidecar is already loaded above, and isSectionFresh is a pure
+ // comparison. Only the target itself is new work, and it is resolved ONCE per
+ // channel (a settings read, a registry lookup, and one small context file).
+ //
+ // The LOCAL lane is the target on purpose: it is the lane that carries the
+ // corpus, and the metered lane exists only for the long tail.
+ const digestTarget = await resolveDigestTarget({
+ paths,
+ channelSlug: slug,
+ lane: "local",
+ });
let transcribedWithAudioBytes = 0;
let multipleAudioFormatsBytes = 0;
let foreignAudioBytes = 0;
@@ -610,11 +636,24 @@ export async function generateChannelSnapshot(
const engine = chapters?.provenance.appId ?? tags?.provenance.appId;
const hasItems =
(chapters?.items.length ?? 0) > 0 || (tags?.items.length ?? 0) > 0;
+ // Coverage-by-engine answers "what produced what is on disk?" and stays
+ // identity-BLIND — a section generated by an older prompt version was
+ // still generated by that engine, and hiding it would make the coverage
+ // split lie during exactly the config change it exists to survey.
if (engine && hasItems) {
digestEngines[engine] = (digestEngines[engine] ?? 0) + 1;
- } else {
- noDigest.push(id);
}
+ // The WORK LIST is identity-aware: a video whose digest predates the
+ // current identity is work, not coverage. A digest shared from a
+ // duplicate cluster's canonical member counts as done — the canonical
+ // member's own freshness is what drives regeneration, and the share is
+ // re-applied from it (isSharedFrom's contract).
+ const fresh =
+ digest?.derivedFrom != null ||
+ digestTarget.sections.every((section) =>
+ isSectionFresh(digest, section, digestTarget.target),
+ );
+ if (!fresh) noDigest.push(id);
}
if (files.isUntranscribable) {
untranscribable.push(id);
diff --git a/common/controller/digestBatch.ts b/common/controller/digestBatch.ts
@@ -24,18 +24,12 @@ import { getSettings } from "../lib/settings";
import { runPool } from "../jobs/concurrentRunner";
import type { TaskTracker } from "../jobs/taskHooks";
import type { JobProgress } from "../jobs/registry";
-import {
- getDigestApp,
- type DigestAppConfig,
-} from "../lib/digestApps";
-import {
- digestPromptVariant,
- isSectionFresh,
- type DigestSectionKind,
-} from "../lib/digest";
+import { isSectionFresh, type DigestSectionKind } from "../lib/digest";
import { loadDigest } from "../lib/digest-server";
-import { PROMPT_VERSION, maxCuesForContext } from "../lib/digestPrompt";
-import { readDigestContext } from "../lib/digestContext-server";
+import {
+ resolveDigestTarget,
+ type DigestLaneChoice,
+} from "./digestTarget";
import { CUES_JSON_FILENAME } from "../lib/videoStatus";
import { digestVideo } from "./digestVideo";
import {
@@ -142,23 +136,20 @@ export async function runDigestBatch(
"The metered digest lane is disabled (settings.digest.remoteEnabled). Enable it in Settings before running it.",
);
}
- const appId =
- lane === "remote" ? digestSettings.remoteAppId : digestSettings.localAppId;
- const app = getDigestApp(appId);
- const config: DigestAppConfig = digestSettings.apps[app.id] ?? {};
- const sections = opts.sections ?? digestSettings.sections;
- const modelRequested = config.model?.trim() || app.defaultModel();
- const timestampMode = digestSettings.timestampMode;
- // Must match what digestVideo will actually chunk with, or the batch's
- // freshness check and the writer would disagree on the identity and every
- // video would look stale forever.
- const promptVariant = digestPromptVariant({
- ...digestSettings,
- maxCues: maxCuesForContext(config.numCtx),
+ // ONE derivation, shared with countMissingDigests and the channel snapshot's
+ // noDigest bucket — see digestTarget.ts. It must match what digestVideo will
+ // actually chunk with, or the batch's freshness check and the writer would
+ // disagree on the identity and every video would look stale forever.
+ const resolved = await resolveDigestTarget({
+ paths: opts.paths,
+ channelSlug: opts.channelSlug,
+ lane,
+ sections: opts.sections,
});
+ const { app, config, sections, modelRequested, timestampMode, context } =
+ resolved;
const dataDir = path.join(opts.paths.channelsDir, opts.channelSlug, "data");
- const context = await readDigestContext(opts.paths, opts.channelSlug);
// Fail fast and loudly rather than 1,100 times in a row: an unreachable engine
// is a configuration problem, and discovering it per-item wastes the log.
@@ -243,13 +234,7 @@ export async function runDigestBatch(
spendCapped: false,
};
- const freshnessTarget = {
- appId: app.id,
- model: modelRequested,
- promptVersion: PROMPT_VERSION,
- contextHash: context.hash,
- ...(promptVariant ? { promptVariant } : {}),
- };
+ const freshnessTarget = resolved.target;
// Ids already handed out this run. Combined with the cursor below, this is what
// makes the disk re-derivation O(n) overall rather than O(n²): the cursor only
@@ -429,26 +414,24 @@ export async function runDigestBatch(
// How many of a channel's videos still need a digest at the CURRENT prompt/model
// identity. Used to scope a job's progress bar, and cheap enough to call before
// starting one (a small sidecar read per video, no transcript parsing).
+//
+// `lane` is REQUIRED-in-spirit: it used to be absent and the local app was
+// resolved unconditionally, so a metered-lane job's progress bar was sized
+// against the LOCAL engine's identity — and, once numCtx joined the identity,
+// against the local app's context size too. Both call sites already know their
+// lane. It still defaults to "local" so the meaning of an unqualified call is
+// the same as it always was, rather than silently changing under old callers.
export async function countMissingDigests(
paths: Paths,
channelSlug: string,
ids?: ReadonlyArray<string>,
+ lane: DigestLaneChoice = "local",
): Promise<number> {
- const settings = getSettings().digest;
- const app = getDigestApp(settings.localAppId);
- const config = settings.apps[app.id] ?? {};
- const context = await readDigestContext(paths, channelSlug);
- const promptVariant = digestPromptVariant({
- ...settings,
- maxCues: maxCuesForContext(config.numCtx),
+ const { target, sections } = await resolveDigestTarget({
+ paths,
+ channelSlug,
+ lane,
});
- const target = {
- appId: app.id,
- model: config.model?.trim() || app.defaultModel(),
- promptVersion: PROMPT_VERSION,
- contextHash: context.hash,
- ...(promptVariant ? { promptVariant } : {}),
- };
const dataDir = path.join(paths.channelsDir, channelSlug, "data");
const dirs = ids
? [...ids]
@@ -456,7 +439,7 @@ export async function countMissingDigests(
let missing = 0;
for (const id of dirs) {
const record = await loadDigest(path.join(dataDir, id));
- const fresh = settings.sections.every((section) =>
+ const fresh = sections.every((section) =>
isSectionFresh(record, section, target),
);
if (!fresh) missing++;
diff --git a/common/controller/digestTarget.ts b/common/controller/digestTarget.ts
@@ -0,0 +1,95 @@
+// The ONE place a digest freshness target is derived.
+//
+// Three callers need to answer "would we regenerate this section?" — the batch
+// runner (to skip), countMissingDigests (to size a progress bar), and the
+// channel snapshot (to fill the noDigest bucket). They MUST agree, and before
+// this module they did not: the snapshot asked only "does an ai-digest.json with
+// items exist?", so after any config change the Digest stage read "All digested"
+// while the batch reported the whole channel as stale. Two counters that
+// disagree by an entire channel is how a sweep reports "nothing to do" on a
+// corpus that needs redoing.
+//
+// The lane is a PARAMETER, never assumed. countMissingDigests used to resolve
+// settings.localAppId unconditionally, so a metered-lane job's progress target
+// was computed against the local engine's identity — and, once numCtx moved into
+// the recorded identity, mis-derived promptVariant from the local app's context
+// size as well.
+
+import type { Paths } from "../lib/paths";
+import { getSettings } from "../lib/settings";
+import { getDigestApp, type DigestApp, type DigestAppConfig } from "../lib/digestApps";
+import {
+ digestPromptVariant,
+ type DigestFreshnessTarget,
+ type DigestSectionKind,
+ type DigestTimestampMode,
+} from "../lib/digest";
+import { PROMPT_VERSION, maxCuesForContext } from "../lib/digestPrompt";
+import {
+ readDigestContext,
+ type DigestContext,
+} from "../lib/digestContext-server";
+
+// Which lane to resolve against. The string form the editor actions already
+// speak, rather than DigestLane, so call sites pass what they already hold.
+export type DigestLaneChoice = "local" | "remote";
+
+export type ResolvedDigestTarget = {
+ target: DigestFreshnessTarget;
+ // Which sections must ALL be fresh for a video to count as digested.
+ sections: DigestSectionKind[];
+ app: DigestApp;
+ config: DigestAppConfig;
+ // What the config asked for. Distinct from what the engine reports having run
+ // — freshness compares this one (see DigestProvenance.modelRequested).
+ modelRequested: string;
+ timestampMode: DigestTimestampMode;
+ // Cues per chunk at this app's context size. The batch passes it on to the
+ // chunker, so the identity and the actual chunking cannot drift apart.
+ maxCues: number;
+ // The channel-context note + its hash. Returned rather than re-read because
+ // the hash is already IN the target above: reading it twice invites the two
+ // copies to disagree.
+ context: DigestContext;
+};
+
+export async function resolveDigestTarget(opts: {
+ paths: Paths;
+ channelSlug: string;
+ lane?: DigestLaneChoice;
+ // Overrides the lane's configured app. The bake-off harness drives this.
+ appId?: string;
+ sections?: DigestSectionKind[];
+}): Promise<ResolvedDigestTarget> {
+ const digestSettings = getSettings().digest;
+ const appId =
+ opts.appId ??
+ (opts.lane === "remote"
+ ? digestSettings.remoteAppId
+ : digestSettings.localAppId);
+ const app = getDigestApp(appId);
+ const config: DigestAppConfig = digestSettings.apps[app.id] ?? {};
+ const modelRequested = config.model?.trim() || app.defaultModel();
+ const maxCues = maxCuesForContext(config.numCtx);
+ const promptVariant = digestPromptVariant({
+ ...digestSettings,
+ maxCues,
+ });
+ const context = await readDigestContext(opts.paths, opts.channelSlug);
+ return {
+ target: {
+ appId: app.id,
+ model: modelRequested,
+ promptVersion: PROMPT_VERSION,
+ contextHash: context.hash,
+ ...(promptVariant ? { promptVariant } : {}),
+ },
+ sections: opts.sections ?? digestSettings.sections,
+ app,
+ config,
+ modelRequested,
+ timestampMode: digestSettings.timestampMode,
+ maxCues,
+ context,
+ };
+}
diff --git a/common/lib/digest.test.ts b/common/lib/digest.test.ts
@@ -260,22 +260,30 @@ test("isSharedFrom recognizes a digest copied from a cluster's canonical member"
// ---------------------------------------------------------------------------
test("digestPromptVariant: the default configuration has no variant", () => {
+ // The derivation is RELATIVE TO THE DEFAULTS, so this tracks whatever they
+ // currently are — today chunk-local at 600 cues. Pinned explicitly (rather
+ // than via the constants) so that flipping a default without bumping
+ // PROMPT_VERSION fails here loudly instead of silently re-marking every
+ // existing digest as fresh.
assert.equal(digestPromptVariant({}), undefined);
- assert.equal(digestPromptVariant({ timestampMode: "absolute" }), undefined);
+ assert.equal(digestPromptVariant({ timestampMode: "chunk-local" }), undefined);
+ assert.equal(digestPromptVariant({ maxCues: 600 }), undefined);
+ assert.equal(
+ digestPromptVariant({ timestampMode: "chunk-local", maxCues: 600 }),
+ undefined,
+ );
assert.equal(digestPromptVariant({ promptVariant: " " }), undefined);
});
test("digestPromptVariant: a non-default timestampMode is folded in", () => {
// Folded in rather than left to the operator to remember: a knob that changes
// the output but not the identity would let a re-run under a different mode
- // skip every video as "fresh".
+ // skip every video as "fresh". Now that chunk-local is the default, it is
+ // ABSOLUTE that must carry an identity of its own.
+ assert.equal(digestPromptVariant({ timestampMode: "absolute" }), "absolute");
assert.equal(
- digestPromptVariant({ timestampMode: "chunk-local" }),
- "chunk-local",
- );
- assert.equal(
- digestPromptVariant({ promptVariant: "dense", timestampMode: "chunk-local" }),
- "dense+chunk-local",
+ digestPromptVariant({ promptVariant: "dense", timestampMode: "absolute" }),
+ "dense+absolute",
);
assert.equal(digestPromptVariant({ promptVariant: "dense" }), "dense");
});
@@ -326,22 +334,24 @@ test("isSectionFresh: a variant record is stale against the default target", ()
});
test("digestPromptVariant: the default chunk size contributes nothing", () => {
- assert.equal(digestPromptVariant({ maxCues: 1200 }), undefined);
- assert.equal(digestPromptVariant({ maxCues: 600 }), "c600");
+ assert.equal(digestPromptVariant({ maxCues: 600 }), undefined);
+ assert.equal(digestPromptVariant({ maxCues: 1200 }), "c1200");
assert.equal(
- digestPromptVariant({ timestampMode: "chunk-local", maxCues: 600 }),
- "chunk-local+c600",
+ digestPromptVariant({ timestampMode: "absolute", maxCues: 1200 }),
+ "absolute+c1200",
);
});
test("isSectionFresh: a chunk-size change invalidates the section", () => {
// Measured: halving the chunk took qwen2.5:7b from 9.34 to 13.37 chapters per
// hour and its worst coverage gap from 1:27:48 to 24:13. A change that large
- // must not be able to hide behind an unchanged identity.
+ // must not be able to hide behind an unchanged identity. 600 is now the
+ // default (and so contributes nothing), so the departure tested here is a
+ // move BACK UP to the old 1200.
assert.equal(
isSectionFresh(machineRecord(), "chapters", {
...TARGET,
- promptVariant: digestPromptVariant({ maxCues: 600 }),
+ promptVariant: digestPromptVariant({ maxCues: 1200 }),
}),
false,
);
diff --git a/common/lib/digest.ts b/common/lib/digest.ts
@@ -40,11 +40,27 @@ export type DigestLane = "local-gpu" | "remote-api";
// chunk against them without reaching a module that imports execa, and so the
// identity helper below can compare against the default without a cycle.
//
-// 16k is comfortable on an 8 GB card (qwen2.5:7b KV cache ~56 KB/token ->
-// ~0.9 GB at 16k, atop 4.7 GB of weights). digestPrompt.ts re-exports the cue
-// count as DIGEST_MAX_CUES_PER_CHUNK and derives other sizes from it.
-export const DEFAULT_DIGEST_NUM_CTX = 16384;
-export const DEFAULT_DIGEST_MAX_CUES_PER_CHUNK = 1200;
+// 8192, ON MEASUREMENT, not on comfort. Round 2 of the bake-off
+// (plans/bakeoff/round2.md) ran qwen2.5:7b at both sizes over the same 6.96
+// audio-hours, chunk-local in both cases:
+//
+// @16384 11.1% zero-yield, 9.34 chapters/h, worst gap 1:27:48, 25.1 days
+// @8192 11.8% zero-yield, 13.37 chapters/h, worst gap 0:24:13, 24.2 days
+//
+// Halving the window nearly halves the worst coverage gap and raises the
+// segmentation rate by 43% at no throughput cost — twice as many calls each
+// carry half the prompt, so the projected sweep is if anything shorter. 16k was
+// the cautious choice and it measured worse.
+//
+// A 7B model at 8k is also well within an 8 GB card (qwen2.5:7b KV cache
+// ~56 KB/token -> ~0.45 GB at 8k, atop 4.7 GB of weights). digestPrompt.ts
+// re-exports the cue count as DIGEST_MAX_CUES_PER_CHUNK and derives other sizes
+// from it; maxCuesForContext() scales the cue count with whatever numCtx an app
+// is actually configured for, so these two must stay in proportion (~10
+// tokens/cue: 600 cues ~= 6k tokens of transcript inside an 8k window, leaving
+// room for the prompt and the response).
+export const DEFAULT_DIGEST_NUM_CTX = 8192;
+export const DEFAULT_DIGEST_MAX_CUES_PER_CHUNK = 600;
// How the transcript markers inside ONE chunk are numbered, and therefore what
// the model is asked to copy.
@@ -63,11 +79,24 @@ export const DEFAULT_DIGEST_MAX_CUES_PER_CHUNK = 1200;
// caught all nine, which is exactly its job, but a caught error is still a lost
// chunk, and >4 h videos are 8.2% of the corpus by count and 46% of its tokens.
// chunk-local removes the large offset the model has to hold. Which mode is
-// actually better is a BAKE-OFF QUESTION (common/bin/digest-bakeoff.ts), which
-// is why both are shipped rather than one being pre-applied as a fix.
+// actually better was a BAKE-OFF QUESTION (common/bin/digest-bakeoff.ts), which
+// is why both were shipped rather than one being pre-applied as a fix.
+//
+// IT HAS BEEN ANSWERED. Round 2 (plans/bakeoff/round2.md), qwen2.5:7b@8192 over
+// the same 6.96 audio-hours:
+//
+// absolute 29.4% zero-yield chunks, 10.35 chapters/h, worst gap 1:05:16,
+// 47.5% of items rejected (65 of them out-of-range)
+// chunk-local 11.8% zero-yield chunks, 13.37 chapters/h, worst gap 0:24:13,
+// 19.1% rejected (13 out-of-range)
+//
+// gemma2:9b ranks the two the same way, so the effect is not model-specific.
+// chunk-local is now the DEFAULT. Changing it changes generated output, so it
+// was paired with a PROMPT_VERSION bump (digestPrompt.ts) rather than left to
+// digestPromptVariant's relative-to-default derivation — see the note there.
export const DIGEST_TIMESTAMP_MODES = ["absolute", "chunk-local"] as const;
export type DigestTimestampMode = (typeof DIGEST_TIMESTAMP_MODES)[number];
-export const DEFAULT_DIGEST_TIMESTAMP_MODE: DigestTimestampMode = "absolute";
+export const DEFAULT_DIGEST_TIMESTAMP_MODE: DigestTimestampMode = "chunk-local";
export function isDigestTimestampMode(v: unknown): v is DigestTimestampMode {
return (
diff --git a/common/lib/digestPrompt.test.ts b/common/lib/digestPrompt.test.ts
@@ -2,12 +2,17 @@ import { test } from "node:test";
import assert from "node:assert/strict";
import {
DIGEST_MAX_CUES_PER_CHUNK,
+ PROMPT_VERSION,
buildChapterPrompt,
buildTagPrompt,
maxCuesForContext,
promptOffsetSeconds,
} from "./digestPrompt";
-import { DEFAULT_DIGEST_MAX_CUES_PER_CHUNK } from "./digest";
+import {
+ DEFAULT_DIGEST_MAX_CUES_PER_CHUNK,
+ DEFAULT_DIGEST_NUM_CTX,
+ DEFAULT_DIGEST_TIMESTAMP_MODE,
+} from "./digest";
// The chunk size MUST track the configured context. Lowering num_ctx without
// lowering it feeds ollama more transcript than its window holds, and ollama
@@ -22,16 +27,40 @@ test("the two copies of the default chunk size agree", () => {
});
test("maxCuesForContext scales the slice with the window", () => {
- assert.equal(maxCuesForContext(16384), 1200);
+ // Ratio-pinned to the 8192 -> 600 default pair.
assert.equal(maxCuesForContext(8192), 600);
+ assert.equal(maxCuesForContext(16384), 1200);
assert.equal(maxCuesForContext(32768), 2400);
+ assert.equal(maxCuesForContext(4096), 300);
// Unset or nonsense falls back to the default window, never to zero cues.
- assert.equal(maxCuesForContext(undefined), 1200);
- assert.equal(maxCuesForContext(0), 1200);
+ assert.equal(maxCuesForContext(undefined), DEFAULT_DIGEST_MAX_CUES_PER_CHUNK);
+ assert.equal(maxCuesForContext(0), DEFAULT_DIGEST_MAX_CUES_PER_CHUNK);
// Floored, so a tiny window still yields a usable slice rather than one cue.
assert.equal(maxCuesForContext(256), 100);
});
+test("the shipped defaults are the configuration the bake-off measured best", () => {
+ // Round 2 (plans/bakeoff/round2.md) scored qwen2.5:7b@8192/chunk-local best on
+ // zero-yield rate, chapters/hour and worst coverage gap. The decision was made
+ // and then not shipped for a while — the defaults stayed at the losing
+ // 16384/absolute pair — so this pins the two constants that carry it.
+ assert.equal(DEFAULT_DIGEST_NUM_CTX, 8192);
+ assert.equal(DEFAULT_DIGEST_MAX_CUES_PER_CHUNK, 600);
+ assert.equal(DEFAULT_DIGEST_TIMESTAMP_MODE, "chunk-local");
+});
+
+test("PROMPT_VERSION was bumped with the defaults change", () => {
+ // The pairing is the whole point. digestPromptVariant derives its string
+ // RELATIVE to the defaults, so a record written under the old defaults carries
+ // promptVariant: undefined and would compare EQUAL under the new ones —
+ // freezing bake-off-losing output into the corpus as permanently "fresh".
+ // isSectionFresh compares promptVersion, so the bump is what invalidates them.
+ assert.ok(
+ PROMPT_VERSION >= 2,
+ "changing DEFAULT_DIGEST_* without bumping PROMPT_VERSION silently skips every existing digest as fresh",
+ );
+});
+
test("chunk-local re-bases the range the chapter prompt states", () => {
const base = {
title: "T",
diff --git a/common/lib/digestPrompt.ts b/common/lib/digestPrompt.ts
@@ -23,26 +23,42 @@ import {
type DigestTimestampMode,
} from "./digest";
-export const PROMPT_VERSION = 1;
+export const PROMPT_VERSION = 2;
-// NOT bumped by the addition of timestampMode below. In "absolute" mode — the
-// default — the rendered prompt is byte-identical to what version 1 always
-// produced, so every digest on disk stays fresh. The non-default shape is
-// distinguished by provenance.promptVariant instead (see digestPromptVariant),
-// which is the field that exists precisely so several shapes can be compared
-// without a version bump invalidating the corpus between rounds.
+// Version 1 -> 2 is a DEFAULTS change, and the bump is what makes it honest.
+//
+// The measured-best configuration (chunk-local timestamps, an 8192 context and
+// the 600-cue chunk sized to it) is now the default. digestPromptVariant()
+// derives its string RELATIVE TO THE DEFAULT CONSTANTS, so a value equal to the
+// default contributes nothing and the variant comes out `undefined`. That is
+// correct and stays as it is — but it means flipping the defaults alone would
+// have been silently destructive: the 39 pilot digests were generated under
+// absolute/1200 and recorded `promptVariant: undefined`, so under the new
+// defaults they would compare EQUAL and be skipped as fresh forever. Output from
+// the configuration the bake-off measured as worst would have been frozen into
+// the corpus, indistinguishable from output of the best one.
+//
+// isSectionFresh compares promptVersion, so bumping it invalidates every section
+// explicitly. That is the right tool for a default change; promptVariant is the
+// tool for comparing several shapes concurrently WITHOUT invalidating the
+// corpus, which is what a bake-off round needs. Cost here is 39 regenerations.
+//
+// (Version 1 also predates the addition of timestampMode. In "absolute" mode the
+// rendered prompt was byte-identical to what version 1 always produced, which is
+// why adding the knob did not itself require a bump.)
-// Chunking. 12k usable tokens of transcript per call inside a 16k context leaves
-// room for the prompt and the response. Expressed in CUES because that is what
-// the chunker slices; ~10 tokens/cue is the corpus average, so 1200 cues ≈ 12k
-// tokens. The overlap exists so a topic straddling a seam is visible whole to at
-// least one call; the parser de-dups the resulting near-identical chapters.
+// Chunking. ~6k usable tokens of transcript per call inside the default 8k
+// context leaves room for the prompt and the response. Expressed in CUES because
+// that is what the chunker slices; ~10 tokens/cue is the corpus average, so 600
+// cues ≈ 6k tokens. The overlap exists so a topic straddling a seam is visible
+// whole to at least one call; the parser de-dups the resulting near-identical
+// chapters.
export const DIGEST_MAX_CUES_PER_CHUNK = DEFAULT_DIGEST_MAX_CUES_PER_CHUNK;
// Cues per chunk, SIZED TO THE CONFIGURED CONTEXT.
//
-// The 1200-cue default is sized for a 16k window (~10 tokens/cue -> ~12k tokens
-// of transcript, leaving room for the prompt and the response). Lowering
+// The 600-cue default is sized for the default 8k window (~10 tokens/cue -> ~6k
+// tokens of transcript, leaving room for the prompt and the response). Lowering
// `numCtx` without lowering this feeds the engine more transcript than its
// window holds, and ollama TRUNCATES SILENTLY — the model then summarizes
// whatever fragment survived and the result looks like a bad model rather than a
diff --git a/common/lib/settings.ts b/common/lib/settings.ts
@@ -223,8 +223,8 @@ export type DigestSettings = {
// the call count for a smaller payoff.
sections: DigestSectionKind[];
// How each chunk's transcript markers are numbered — see DigestTimestampMode.
- // A scored variable in the bake-off, not a pre-applied fix, so its effect on
- // long videos is measured against the alternatives rather than assumed.
+ // Was a scored variable in the bake-off rather than a pre-applied fix; the
+ // measurement is in and "chunk-local" is now the shipped default.
timestampMode: DigestTimestampMode;
// A free-text label for a non-default prompt shape, folded into the recorded
// provenance by digestPromptVariant(). Setting it invalidates every digest
@@ -496,6 +496,10 @@ export function defaultDigest(): DigestSettings {
longTailSeconds: DIGEST_LONG_TAIL_DEFAULT_SECONDS,
localAppId: DEFAULT_DIGEST_APP_ID,
remoteAppId: CLAUDE_DIGEST_APP_ID,
+ // Empty on purpose: every per-app knob falls through to its own default
+ // constant (resolveNumCtx -> DEFAULT_DIGEST_NUM_CTX, now 8192, and
+ // maxCuesForContext sizes the chunk to it). Seeding a copy of those values
+ // here would give the same number two homes and let them drift.
apps: {},
digestsPaused: false,
spendCapUsd: 0,
diff --git a/editor/app/channels/[slug]/digestActions.ts b/editor/app/channels/[slug]/digestActions.ts
@@ -82,7 +82,10 @@ export async function digestChannelAction(
fn: async (onLog, signal, setProgress, ctx) => {
const stat = await readChannelStat(paths, slug);
if (stat) {
- const missing = await countMissingDigests(paths, slug);
+ // The LANE matters: the progress target must be computed against the
+ // engine that is actually about to run, not whichever one the local
+ // setting names.
+ const missing = await countMissingDigests(paths, slug, undefined, lane);
setProgress({
metric: "digests",
initial: stat.digestCount ?? 0,
@@ -157,7 +160,7 @@ export async function digestBucketAction(
fn: async (onLog, signal, setProgress, ctx) => {
const stat = await readChannelStat(paths, slug);
if (stat) {
- const missing = await countMissingDigests(paths, slug, cleaned);
+ const missing = await countMissingDigests(paths, slug, cleaned, lane);
setProgress({
metric: "digests",
initial: stat.digestCount ?? 0,
diff --git a/editor/app/channels/[slug]/lib/stageStatus.ts b/editor/app/channels/[slug]/lib/stageStatus.ts
@@ -332,9 +332,12 @@ export function computeStageStatuses(
}),
};
- // Videos with a transcript but no digest at the current prompt/model identity.
- // Counted as pending work rather than merely informational: unlike the
- // auto-captions lane, every transcribed video is eventually meant to have one.
+ // Videos with a transcript whose digest is missing OR stale against the local
+ // lane's current identity — the snapshot's noDigest bucket now computes that
+ // full freshness target (channelSnapshot.ts), so this label and that bucket
+ // finally mean the same thing. Counted as pending work rather than merely
+ // informational: unlike the auto-captions lane, every transcribed video is
+ // eventually meant to have one.
const digestRunning = runningByStage.has("digest");
const digestPending = buckets.noDigest.length;
const digest: StageStatus = {
@@ -349,10 +352,10 @@ export function computeStageStatuses(
: digestPending > 0
? pluralize(
digestPending,
- "transcript without a digest",
- "transcripts without a digest",
+ "transcript needs a digest",
+ "transcripts need a digest",
)
- : "All digested.",
+ : "All digested at the current settings.",
tone: pickTone({
running: digestRunning,
pending: digestPending,
diff --git a/plans/FACTS.md b/plans/FACTS.md
@@ -429,22 +429,73 @@ reports itself skipped corpus-wide: `6,970 shorts × 76,318 videos ≈ 531,936,4
pairs`. It still runs, unchanged, at shorts scale. Making it scale is the one
piece genuinely deferred — see STATE.md.
-### The finding that matters most: 63% of title matches are NOT duplicates
+### 63% of title matches were rejected at 0.6 — and most of those rejections were wrong
Of 8,352 pairs nominated by an exact normalized title *and* a compatible runtime,
-**5,250 were rejected once their transcripts were compared.** Only 2,717 held up.
-
-This directly corrects the planning assumption that ~8,204 title-level
-redundancies were available to prune from the digest sweep. The title-level
-*count* was right; the inference that they are duplicates was not. Banked saving
-is at most 2,717 pairs, and only 1,599 of those are timing-aligned enough for a
-shared digest to be placed correctly — roughly **a fifth of what was projected**.
-
-Caveat worth testing before treating that as final: the near threshold is a 5-gram
-Jaccard at 0.6, and the two sides of a cross-platform mirror are often different
-ASR engines (YouTube auto-captions vs. whisper). Some of the 5,250 may be real
-mirrors whose transcriptions simply disagree at the 5-gram level. Sensitivity to
-that threshold has **not** been measured. See STATE.md.
+**5,250 were rejected once their transcripts were compared** at the default 0.6
+5-gram Jaccard. Only 2,717 held up.
+
+That was read as "the title-level count was right; the inference that they are
+duplicates was not." **Measurement has since shown the opposite: the threshold was
+wrong, not the inference.** See the next section.
+
+### MEASURED: the 0.6 near threshold was rejecting real cross-platform mirrors
+
+Re-run corpus-wide, title blocking, everything else identical, only `--near`
+changed (2026-07-27):
+
+| | `--near 0.6` (default) | `--near 0.35` |
+| --- | --- | --- |
+| Nominated pairs | 8,352 | 8,352 |
+| **Confirmed by content** | 2,717 | **7,329** |
+| Rejected by content | 5,250 | **638** |
+| Untestable suspects | 385 | 385 |
+| Clusters written | 2,994 / 6,044 videos | **7,450 / 15,030 videos** |
+| `needsReview` | 335 | 334 |
+| **Timing-aligned mirrors** | 1,599 of 2,689 | **3,952 of 7,221** |
+| Wall clock | 206 s | 249 s |
+
+The confirmed set grows **+170%** and is a strict superset: **0 videos** present in
+the 0.6 clusters are absent from the 0.35 clusters.
+
+**The 4,456 entirely-new clusters were inspected, not just counted**, and they are
+real mirrors. Evidence, in descending order of strength:
+
+- **4,423 of 4,456 (99.3%) are cross-platform AND cross-channel.** 3,954 pair a
+ channel with its own platform-suffixed sibling (`the-quartering` ↔
+ `the-quartering-rumble`); of the 502 that don't, the sampled ones pair a
+ creator's alternate channels (`jeremy-hambly` / `quartering-live` /
+ `the-quartering`, `HasanAbiVODs` / `HasanAbiVODsbackup`).
+- **4,290 of 4,456 have a byte-identical title** across every copy; 4,023 agree on
+ runtime within 2 s; 3,572 share an upload date.
+- The score histogram is a tight band at **0.45–0.60** (1,995 at 0.55–0.60, 1,481
+ at 0.50–0.55, 665 at 0.45–0.50) — i.e. they pile up *just under* the old cutoff,
+ which is the signature of a systematic offset, not of noise.
+- The **138 riskiest** (score < 0.42) and the **502 non-sibling** clusters were
+ sampled across the whole band. Every one examined is the same recording: same
+ title, same runtime to the second, same upload date, mirrored platform.
+
+The mechanism is the one the caveat predicted: the two sides of a cross-platform
+mirror are transcribed by **different ASR engines**, and a 5-gram Jaccard is
+unforgiving of word-level disagreement. Two transcripts of the same audio from
+different engines land at **~0.35–0.60**, not ≥ 0.6. The old default was set for
+same-engine text and silently failed the cross-platform case, which is precisely
+the case duplicate detection exists to catch.
+
+**Consequence for the digest backfill.** The ceiling on shared-digest savings is
+not 2,717 pairs. At 0.35 it is **7,329 confirmed pairs, of which 3,952 are
+timing-aligned** enough for a shared digest to be placed correctly — **2.5× the
+1,599 previously banked**. Against a corpus of 76,318 videos that is ~5.2% of the
+sweep avoidable by sharing rather than ~2.1%.
+
+**Not yet changed:** `NEAR_THRESHOLD_DEFAULT` still ships at 0.6. Lowering it
+changes what every built site asserts to readers, so it is recorded here as a
+measured recommendation rather than applied as a side effect of a digest task.
+See STATE.md.
+
+Reproduce with:
+`pnpm exec tsx bin/duplicate-shorts.ts --all-durations --blocking title --near 0.35`
+(back up `transcripts/duplicates.json` first — the run overwrites it).
### What this unblocks
diff --git a/plans/STATE.md b/plans/STATE.md
@@ -43,13 +43,31 @@ relitigate. Earlier entries (fabric deferred, ollama-direct as the structured de
Claude Code on a second queue key, digests in their own page tree, per-section provenance,
transcripts never rewritten) still stand and are unchanged.
-**`chunk-local` timestamps are adopted, on measurement.** Each chunk's transcript is
-re-based to `00:00:00` and the parser adds the offset back before any guard runs. Measured
-on the long tail it cut zero-yield chunks from 33.3% to 11.1% (qwen2.5@16k), 29.4% to 11.8%
-(qwen2.5@8k) and 17.7% to 0% (gemma2), with the rejection rate falling in every case. The
-working hypothesis — the model cannot hold a large absolute offset across a long chunk and
-reverts to counting from zero — is confirmed. It was shipped as a *scored variable*, not a
-pre-applied fix, which is why there are numbers to quote.
+**`chunk-local` timestamps are adopted, on measurement — and (2026-07-27) SHIPPED as the
+default.** Each chunk's transcript is re-based to `00:00:00` and the parser adds the offset
+back before any guard runs. Measured on the long tail it cut zero-yield chunks from 33.3% to
+11.1% (qwen2.5@16k), 29.4% to 11.8% (qwen2.5@8k) and 17.7% to 0% (gemma2), with the
+rejection rate falling in every case. The working hypothesis — the model cannot hold a large
+absolute offset across a long chunk and reverts to counting from zero — is confirmed. It was
+shipped as a *scored variable*, not a pre-applied fix, which is why there are numbers to
+quote.
+
+For a while it was adopted *on paper only*: `DEFAULT_DIGEST_TIMESTAMP_MODE` stayed
+`"absolute"` and `DEFAULT_DIGEST_NUM_CTX` stayed 16384, so a fresh install — and the pilot —
+ran the configuration round 2 measured as **worst**. Both defaults now carry the decision
+(`chunk-local`, 8192, 600 cues).
+
+**The mechanism for that change was a `PROMPT_VERSION` bump (1 → 2), and it had to be.**
+`digestPromptVariant()` derives the variant *relative to the default constants*, so a value
+equal to the default contributes nothing and the variant is `undefined`. Flipping the
+defaults alone would therefore have made the 39 pilot digests — written under
+absolute/1200 with `promptVariant: undefined` — compare **fresh** under the new defaults:
+bake-off-losing output frozen into the corpus, indistinguishable from the best config's.
+`isSectionFresh` compares `promptVersion`, so bumping it invalidates every section
+explicitly. That is the right division of labour: **`promptVersion` for a default change,
+`promptVariant` for several shapes coexisting during a bake-off round.** Cost: 39
+regenerations. `digestPrompt.test.ts` pins both the new defaults and the bump so the pairing
+cannot be broken silently.
**Chunk size is sized to `num_ctx`, and it is a bigger lever than expected.** Halving the
context (16k → 8k, 1200 → 600 cues) took qwen2.5:7b from 9.34 to 13.37 chapters/hour and its
@@ -65,6 +83,28 @@ now derives it.
written before the field existed still compares fresh. The alternative — an operator
remembering to bump a label — fails silently by skipping the whole corpus as "fresh".
+**One freshness target, three callers (2026-07-27).** `resolveDigestTarget()`
+(`common/controller/digestTarget.ts`) is now the only place a target is derived, and the
+batch runner, `countMissingDigests` and the channel snapshot's `noDigest` bucket all use it.
+Two counters were previously lying:
+
+- `noDigest` asked only `engine && hasItems` — "is there an ai-digest.json with something in
+ it?" — while its own label and `stageStatus.ts` claimed an identity check. After any config
+ change the Digest stage read **"All digested"** while `countMissingDigests` reported the
+ whole channel stale. The two could disagree by an entire channel, which is exactly how a
+ sweep reports "nothing to do" on a corpus that needs redoing. It now computes the real
+ target; the per-video cost is zero extra I/O (the sidecar is already loaded and
+ `isSectionFresh` is pure), and only the target itself is new work, resolved once per
+ channel. Coverage-by-engine (`digestEngines`) stays deliberately identity-**blind** — a
+ section made by an older prompt version was still made by that engine.
+- `countMissingDigests` always resolved `settings.localAppId`, so a metered-lane job's
+ progress bar was sized against the *local* engine's identity (and, once `numCtx` joined the
+ identity, the local app's context size). It takes the lane as a parameter now; both call
+ sites in `digestActions.ts` already knew which lane they were.
+
+A digest shared from a duplicate cluster's canonical member counts as done in the bucket —
+the canonical member's own freshness drives regeneration and the share is re-applied from it.
+
**The sweep is ~25 days, not ~64.** Round 1 measured 59 days on short+medium; Round 2
measured ~25 on the long tail, because long videos amortise the fixed per-call overhead and
the corpus is dominated by them. Throughput is no longer the binding constraint it was
@@ -88,10 +128,15 @@ so choosing one trades recall against runtime rather than correctness. This is w
nominating aggressively safe — and it earns its keep: **5,250 of 8,352 title-nominated pairs
(63%) were rejected once their transcripts were compared.**
-**The projected digest saving was overstated by ~5x, and this is the headline finding.** The
-~8,204 "title-level redundancies" were real as a *count* of same-title pairs, but only 2,717
-survive content comparison, and only 1,599 of those are timing-aligned enough for a shared
-digest to land correctly. Plan the backfill on those numbers, not the old ones.
+**SUPERSEDED (2026-07-27) — that 63% was mostly the threshold, not the corpus.** This section
+originally read "the projected digest saving was overstated by ~5x": of ~8,204 title-level
+redundancies only 2,717 survived content comparison, and only 1,599 of those were
+timing-aligned. Measuring `--near` (the caveat that was attached to it) showed the rejections
+were overwhelmingly **real cross-platform mirrors transcribed by different ASR engines**. At
+`--near 0.35` the confirmed set is **7,329 pairs / 3,952 aligned** — 2.5× the banked saving,
+and inspection of all 4,456 new clusters found no false positives. Plan the backfill on
+**7,329 / 3,952**, and see the near-threshold entry under Open questions before treating 0.6
+as settled. Full evidence in FACTS.md.
**A previous decision is deliberately reversed: title+duration clusters DO form.** The old
rule was that only content-confirmed tiers cluster, because duration coincidence alone
@@ -132,13 +177,26 @@ playwright `webServer` with `OLLAMA_URL` pointed at it in **both** `dev:test` an
they save. gemma2 is hard-capped at an 8192 context by the model.
- **The `verylong` (>6 h) stratum is unmeasured.** Round 2 used the `long` bucket (~3.5 h).
The sample already contains an 8 h video for it.
-- **Is the 0.6 near threshold rejecting real mirrors?** 63% of title-nominated pairs failed
- content comparison. Some of that is genuinely two episodes of a daily show — which is the
- point — but the two sides of a cross-platform mirror are often *different ASR engines*
- (YouTube auto-captions vs. whisper), and a 5-gram Jaccard is unforgiving of word-level
- disagreement. Sensitivity to `--near` has **not** been measured. Cheap to answer: re-run
- `--blocking title --near 0.35` and diff the confirmed set. Do this before treating the
- 2,717 figure as the ceiling on the digest saving.
+- **ANSWERED (2026-07-27): yes — the 0.6 near threshold was rejecting real mirrors, in bulk.**
+ Re-running `--blocking title --near 0.35` corpus-wide takes confirmed pairs from **2,717 to
+ 7,329** (+170%) and timing-aligned mirrors from **1,599 to 3,952**, a strict superset (no
+ video in the 0.6 clusters is absent from the 0.35 ones). The 4,456 new clusters were
+ inspected, not merely counted: **99.3% are cross-platform *and* cross-channel**, 96% have a
+ byte-identical title, 90% agree on runtime within 2 s, and their scores pile into a tight
+ **0.45–0.60** band just under the old cutoff. Sampling the 138 riskiest (< 0.42) and all
+ 502 non-sibling cases found **no false positives**. Mechanism confirmed as hypothesised:
+ two *different ASR engines* transcribing the same audio agree at ~0.35–0.60 on 5-gram
+ Jaccard, so 0.6 was tuned for same-engine text and structurally failed the cross-platform
+ case. Full table + reproduction in FACTS.md. **The caveat is deleted, not carried.**
+
+ **Now a decision, not a question: should `NEAR_THRESHOLD_DEFAULT` drop from 0.6?** The
+ measurement says yes and the digest-saving ceiling rises 2.5× if it does. It was *not*
+ changed here, deliberately — detection output feeds every built site's Duplicates page and
+ search badges, so loosening it changes what the archive asserts to readers, and that is its
+ own reviewable change rather than a side effect of a digest task. What is still unmeasured
+ is where the floor is: 0.35 was the single probe, and the 638 pairs it still rejects have
+ not been examined. Recommended next: probe 0.25 / 0.45 to bracket the knee, eyeball the
+ remaining rejections, then change the default in its own commit.
- **Recall of the title pre-filter — now measured, and the gap is 11%.** Running both
strategies corpus-wide and diffing the clusters: duration blocking finds **672 videos title
blocking misses** (the re-titled mirrors it structurally cannot see), title finds 68 that