Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 5d1a28630eb51cf216d81fa4f104e58dc9f64349
parent 1476904975d1d186034e0f0c6a84bce9279d3adb
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon, 27 Jul 2026 00:22:37 -0400

Digest: ship the measured-best defaults, and make the counters honest

Three things that each independently made a validation run misleading.

1. The shipped defaults were the configuration the bake-off measured as
   WORST. round2.md scored qwen2.5:7b@8192/chunk-local best (11.8%
   zero-yield, 13.37 chapters/h, 24:13 worst gap) against @16384/absolute
   (33.3%, 8.19, 1:27:48) — but DEFAULT_DIGEST_TIMESTAMP_MODE was
   "absolute" and DEFAULT_DIGEST_NUM_CTX 16384, so a fresh install ran the
   loser. Now chunk-local / 8192 / 600 cues.

   PAIRED WITH A PROMPT_VERSION BUMP (1 -> 2), which is the whole point.
   digestPromptVariant derives its string RELATIVE to the defaults, so a
   value equal to the default yields `undefined`. Flipping the defaults
   alone would have made the 39 pilot digests — written under
   absolute/1200 with promptVariant: undefined — compare FRESH under the
   new defaults, freezing bake-off-losing output into the corpus forever.
   isSectionFresh compares promptVersion, so the bump invalidates them
   explicitly. Cost: 39 regenerations. Tests pin the pairing.

2. The digest pending count did not mean what it said. buckets.noDigest
   asked only `engine && hasItems`, while its label and stageStatus.ts
   claimed a full identity check, so after any config change the stage
   read "All digested" while countMissingDigests reported the whole
   channel stale — a disagreement of an entire channel. New
   resolveDigestTarget() is the single derivation; batch runner, counter
   and snapshot all use it. Per-video cost is zero extra I/O. Coverage by
   engine stays identity-blind on purpose.

3. countMissingDigests always resolved the LOCAL app, so a metered-lane
   job's progress bar was sized against the wrong engine's identity. It
   takes the lane now; both call sites already knew it.

Also §1 of the plan: measured --near sensitivity for duplicate detection.
0.6 was rejecting real cross-platform mirrors in bulk — at 0.35 confirmed
pairs go 2,717 -> 7,329 and aligned mirrors 1,599 -> 3,952, a strict
superset, and all 4,456 new clusters inspect as genuine (99.3%
cross-platform AND cross-channel, 96% byte-identical titles, scores in a
tight 0.45-0.60 band just under the old cutoff). Cause: the two sides are
different ASR engines. Recorded in FACTS.md/STATE.md; the default is NOT
changed here — it alters what every built site asserts, so it belongs in
its own commit. transcripts/duplicates.json restored byte-identical.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Diffstat:
Mcommon/controller/channelSnapshot.ts | 53++++++++++++++++++++++++++++++++++++++++++++++-------
Mcommon/controller/digestBatch.ts | 77++++++++++++++++++++++++++++++-----------------------------------------------
Acommon/controller/digestTarget.ts | 95+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/lib/digest.test.ts | 38++++++++++++++++++++++++--------------
Mcommon/lib/digest.ts | 45+++++++++++++++++++++++++++++++++++++--------
Mcommon/lib/digestPrompt.test.ts | 37+++++++++++++++++++++++++++++++++----
Mcommon/lib/digestPrompt.ts | 44++++++++++++++++++++++++++++++--------------
Mcommon/lib/settings.ts | 8++++++--
Meditor/app/channels/[slug]/digestActions.ts | 7+++++--
Meditor/app/channels/[slug]/lib/stageStatus.ts | 15+++++++++------
Mplans/FACTS.md | 79+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++--------------
Mplans/STATE.md | 94++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++---------------
12 files changed, 456 insertions(+), 136 deletions(-)

diff --git a/common/controller/channelSnapshot.ts b/common/controller/channelSnapshot.ts @@ -24,6 +24,8 @@ import { } from "../lib/availability-server"; import { isDoNotClean } from "../lib/doNotClean-server"; import { loadDigest } from "../lib/digest-server"; +import { isSectionFresh } from "../lib/digest"; +import { resolveDigestTarget } from "./digestTarget"; import { isExcludedFromTruncatedCheck } from "../lib/excludeTruncatedCheck-server"; import { loadDownloadOutcome } from "../lib/downloadOutcome-server"; import type { Paths } from "../lib/paths"; @@ -148,11 +150,17 @@ export type ChannelSnapshot = { // (retry-bucket with forceCookies). Optional: older snapshots lack it; // readers must default to []. needsCookies: string[]; - // Transcribed videos with no (or an outdated-shape) ai-digest.json — the AI - // digest layer's work list, and the denominator for corpus coverage during - // the backfill. Only TRANSCRIBED videos are listed: a video without a - // transcript is a transcription problem, not a digest one. Optional: older - // snapshots lack it; readers must default to []. + // Transcribed videos whose digest is missing or STALE against the local + // lane's current freshness target (engine, requested model, prompt version, + // prompt shape, context hash) — the AI digest layer's work list, and the + // denominator for corpus coverage during the backfill. This is the same + // target countMissingDigests and runDigestBatch compute, deliberately: a + // bucket that silently meant something weaker than its name is how a sweep + // reports "nothing to do" on a corpus that needs redoing. + // + // Only TRANSCRIBED videos are listed: a video without a transcript is a + // transcription problem, not a digest one. Optional: older snapshots lack + // it; readers must default to []. noDigest: string[]; }; undownloadedIds: string[]; @@ -458,6 +466,24 @@ export async function generateChannelSnapshot( const supersededAutoSubs: string[] = []; const noDigest: string[] = []; const digestEngines: Record<string, number> = {}; + // The SAME freshness target countMissingDigests and runDigestBatch use, so + // the Digest stage's count and the batch runner's progress target cannot + // disagree. This bucket used to ask only "is there an ai-digest.json with + // items?", which meant that after any prompt/model/context change the stage + // read "All digested." while the batch reported the whole channel as stale. + // + // Affordable in this hot path because the per-video cost is ZERO extra I/O: + // the digest sidecar is already loaded above, and isSectionFresh is a pure + // comparison. Only the target itself is new work, and it is resolved ONCE per + // channel (a settings read, a registry lookup, and one small context file). + // + // The LOCAL lane is the target on purpose: it is the lane that carries the + // corpus, and the metered lane exists only for the long tail. + const digestTarget = await resolveDigestTarget({ + paths, + channelSlug: slug, + lane: "local", + }); let transcribedWithAudioBytes = 0; let multipleAudioFormatsBytes = 0; let foreignAudioBytes = 0; @@ -610,11 +636,24 @@ export async function generateChannelSnapshot( const engine = chapters?.provenance.appId ?? tags?.provenance.appId; const hasItems = (chapters?.items.length ?? 0) > 0 || (tags?.items.length ?? 0) > 0; + // Coverage-by-engine answers "what produced what is on disk?" and stays + // identity-BLIND — a section generated by an older prompt version was + // still generated by that engine, and hiding it would make the coverage + // split lie during exactly the config change it exists to survey. if (engine && hasItems) { digestEngines[engine] = (digestEngines[engine] ?? 0) + 1; - } else { - noDigest.push(id); } + // The WORK LIST is identity-aware: a video whose digest predates the + // current identity is work, not coverage. A digest shared from a + // duplicate cluster's canonical member counts as done — the canonical + // member's own freshness is what drives regeneration, and the share is + // re-applied from it (isSharedFrom's contract). + const fresh = + digest?.derivedFrom != null || + digestTarget.sections.every((section) => + isSectionFresh(digest, section, digestTarget.target), + ); + if (!fresh) noDigest.push(id); } if (files.isUntranscribable) { untranscribable.push(id); diff --git a/common/controller/digestBatch.ts b/common/controller/digestBatch.ts @@ -24,18 +24,12 @@ import { getSettings } from "../lib/settings"; import { runPool } from "../jobs/concurrentRunner"; import type { TaskTracker } from "../jobs/taskHooks"; import type { JobProgress } from "../jobs/registry"; -import { - getDigestApp, - type DigestAppConfig, -} from "../lib/digestApps"; -import { - digestPromptVariant, - isSectionFresh, - type DigestSectionKind, -} from "../lib/digest"; +import { isSectionFresh, type DigestSectionKind } from "../lib/digest"; import { loadDigest } from "../lib/digest-server"; -import { PROMPT_VERSION, maxCuesForContext } from "../lib/digestPrompt"; -import { readDigestContext } from "../lib/digestContext-server"; +import { + resolveDigestTarget, + type DigestLaneChoice, +} from "./digestTarget"; import { CUES_JSON_FILENAME } from "../lib/videoStatus"; import { digestVideo } from "./digestVideo"; import { @@ -142,23 +136,20 @@ export async function runDigestBatch( "The metered digest lane is disabled (settings.digest.remoteEnabled). Enable it in Settings before running it.", ); } - const appId = - lane === "remote" ? digestSettings.remoteAppId : digestSettings.localAppId; - const app = getDigestApp(appId); - const config: DigestAppConfig = digestSettings.apps[app.id] ?? {}; - const sections = opts.sections ?? digestSettings.sections; - const modelRequested = config.model?.trim() || app.defaultModel(); - const timestampMode = digestSettings.timestampMode; - // Must match what digestVideo will actually chunk with, or the batch's - // freshness check and the writer would disagree on the identity and every - // video would look stale forever. - const promptVariant = digestPromptVariant({ - ...digestSettings, - maxCues: maxCuesForContext(config.numCtx), + // ONE derivation, shared with countMissingDigests and the channel snapshot's + // noDigest bucket — see digestTarget.ts. It must match what digestVideo will + // actually chunk with, or the batch's freshness check and the writer would + // disagree on the identity and every video would look stale forever. + const resolved = await resolveDigestTarget({ + paths: opts.paths, + channelSlug: opts.channelSlug, + lane, + sections: opts.sections, }); + const { app, config, sections, modelRequested, timestampMode, context } = + resolved; const dataDir = path.join(opts.paths.channelsDir, opts.channelSlug, "data"); - const context = await readDigestContext(opts.paths, opts.channelSlug); // Fail fast and loudly rather than 1,100 times in a row: an unreachable engine // is a configuration problem, and discovering it per-item wastes the log. @@ -243,13 +234,7 @@ export async function runDigestBatch( spendCapped: false, }; - const freshnessTarget = { - appId: app.id, - model: modelRequested, - promptVersion: PROMPT_VERSION, - contextHash: context.hash, - ...(promptVariant ? { promptVariant } : {}), - }; + const freshnessTarget = resolved.target; // Ids already handed out this run. Combined with the cursor below, this is what // makes the disk re-derivation O(n) overall rather than O(n²): the cursor only @@ -429,26 +414,24 @@ export async function runDigestBatch( // How many of a channel's videos still need a digest at the CURRENT prompt/model // identity. Used to scope a job's progress bar, and cheap enough to call before // starting one (a small sidecar read per video, no transcript parsing). +// +// `lane` is REQUIRED-in-spirit: it used to be absent and the local app was +// resolved unconditionally, so a metered-lane job's progress bar was sized +// against the LOCAL engine's identity — and, once numCtx joined the identity, +// against the local app's context size too. Both call sites already know their +// lane. It still defaults to "local" so the meaning of an unqualified call is +// the same as it always was, rather than silently changing under old callers. export async function countMissingDigests( paths: Paths, channelSlug: string, ids?: ReadonlyArray<string>, + lane: DigestLaneChoice = "local", ): Promise<number> { - const settings = getSettings().digest; - const app = getDigestApp(settings.localAppId); - const config = settings.apps[app.id] ?? {}; - const context = await readDigestContext(paths, channelSlug); - const promptVariant = digestPromptVariant({ - ...settings, - maxCues: maxCuesForContext(config.numCtx), + const { target, sections } = await resolveDigestTarget({ + paths, + channelSlug, + lane, }); - const target = { - appId: app.id, - model: config.model?.trim() || app.defaultModel(), - promptVersion: PROMPT_VERSION, - contextHash: context.hash, - ...(promptVariant ? { promptVariant } : {}), - }; const dataDir = path.join(paths.channelsDir, channelSlug, "data"); const dirs = ids ? [...ids] @@ -456,7 +439,7 @@ export async function countMissingDigests( let missing = 0; for (const id of dirs) { const record = await loadDigest(path.join(dataDir, id)); - const fresh = settings.sections.every((section) => + const fresh = sections.every((section) => isSectionFresh(record, section, target), ); if (!fresh) missing++; diff --git a/common/controller/digestTarget.ts b/common/controller/digestTarget.ts @@ -0,0 +1,95 @@ +// The ONE place a digest freshness target is derived. +// +// Three callers need to answer "would we regenerate this section?" — the batch +// runner (to skip), countMissingDigests (to size a progress bar), and the +// channel snapshot (to fill the noDigest bucket). They MUST agree, and before +// this module they did not: the snapshot asked only "does an ai-digest.json with +// items exist?", so after any config change the Digest stage read "All digested" +// while the batch reported the whole channel as stale. Two counters that +// disagree by an entire channel is how a sweep reports "nothing to do" on a +// corpus that needs redoing. +// +// The lane is a PARAMETER, never assumed. countMissingDigests used to resolve +// settings.localAppId unconditionally, so a metered-lane job's progress target +// was computed against the local engine's identity — and, once numCtx moved into +// the recorded identity, mis-derived promptVariant from the local app's context +// size as well. + +import type { Paths } from "../lib/paths"; +import { getSettings } from "../lib/settings"; +import { getDigestApp, type DigestApp, type DigestAppConfig } from "../lib/digestApps"; +import { + digestPromptVariant, + type DigestFreshnessTarget, + type DigestSectionKind, + type DigestTimestampMode, +} from "../lib/digest"; +import { PROMPT_VERSION, maxCuesForContext } from "../lib/digestPrompt"; +import { + readDigestContext, + type DigestContext, +} from "../lib/digestContext-server"; + +// Which lane to resolve against. The string form the editor actions already +// speak, rather than DigestLane, so call sites pass what they already hold. +export type DigestLaneChoice = "local" | "remote"; + +export type ResolvedDigestTarget = { + target: DigestFreshnessTarget; + // Which sections must ALL be fresh for a video to count as digested. + sections: DigestSectionKind[]; + app: DigestApp; + config: DigestAppConfig; + // What the config asked for. Distinct from what the engine reports having run + // — freshness compares this one (see DigestProvenance.modelRequested). + modelRequested: string; + timestampMode: DigestTimestampMode; + // Cues per chunk at this app's context size. The batch passes it on to the + // chunker, so the identity and the actual chunking cannot drift apart. + maxCues: number; + // The channel-context note + its hash. Returned rather than re-read because + // the hash is already IN the target above: reading it twice invites the two + // copies to disagree. + context: DigestContext; +}; + +export async function resolveDigestTarget(opts: { + paths: Paths; + channelSlug: string; + lane?: DigestLaneChoice; + // Overrides the lane's configured app. The bake-off harness drives this. + appId?: string; + sections?: DigestSectionKind[]; +}): Promise<ResolvedDigestTarget> { + const digestSettings = getSettings().digest; + const appId = + opts.appId ?? + (opts.lane === "remote" + ? digestSettings.remoteAppId + : digestSettings.localAppId); + const app = getDigestApp(appId); + const config: DigestAppConfig = digestSettings.apps[app.id] ?? {}; + const modelRequested = config.model?.trim() || app.defaultModel(); + const maxCues = maxCuesForContext(config.numCtx); + const promptVariant = digestPromptVariant({ + ...digestSettings, + maxCues, + }); + const context = await readDigestContext(opts.paths, opts.channelSlug); + return { + target: { + appId: app.id, + model: modelRequested, + promptVersion: PROMPT_VERSION, + contextHash: context.hash, + ...(promptVariant ? { promptVariant } : {}), + }, + sections: opts.sections ?? digestSettings.sections, + app, + config, + modelRequested, + timestampMode: digestSettings.timestampMode, + maxCues, + context, + }; +} diff --git a/common/lib/digest.test.ts b/common/lib/digest.test.ts @@ -260,22 +260,30 @@ test("isSharedFrom recognizes a digest copied from a cluster's canonical member" // --------------------------------------------------------------------------- test("digestPromptVariant: the default configuration has no variant", () => { + // The derivation is RELATIVE TO THE DEFAULTS, so this tracks whatever they + // currently are — today chunk-local at 600 cues. Pinned explicitly (rather + // than via the constants) so that flipping a default without bumping + // PROMPT_VERSION fails here loudly instead of silently re-marking every + // existing digest as fresh. assert.equal(digestPromptVariant({}), undefined); - assert.equal(digestPromptVariant({ timestampMode: "absolute" }), undefined); + assert.equal(digestPromptVariant({ timestampMode: "chunk-local" }), undefined); + assert.equal(digestPromptVariant({ maxCues: 600 }), undefined); + assert.equal( + digestPromptVariant({ timestampMode: "chunk-local", maxCues: 600 }), + undefined, + ); assert.equal(digestPromptVariant({ promptVariant: " " }), undefined); }); test("digestPromptVariant: a non-default timestampMode is folded in", () => { // Folded in rather than left to the operator to remember: a knob that changes // the output but not the identity would let a re-run under a different mode - // skip every video as "fresh". + // skip every video as "fresh". Now that chunk-local is the default, it is + // ABSOLUTE that must carry an identity of its own. + assert.equal(digestPromptVariant({ timestampMode: "absolute" }), "absolute"); assert.equal( - digestPromptVariant({ timestampMode: "chunk-local" }), - "chunk-local", - ); - assert.equal( - digestPromptVariant({ promptVariant: "dense", timestampMode: "chunk-local" }), - "dense+chunk-local", + digestPromptVariant({ promptVariant: "dense", timestampMode: "absolute" }), + "dense+absolute", ); assert.equal(digestPromptVariant({ promptVariant: "dense" }), "dense"); }); @@ -326,22 +334,24 @@ test("isSectionFresh: a variant record is stale against the default target", () }); test("digestPromptVariant: the default chunk size contributes nothing", () => { - assert.equal(digestPromptVariant({ maxCues: 1200 }), undefined); - assert.equal(digestPromptVariant({ maxCues: 600 }), "c600"); + assert.equal(digestPromptVariant({ maxCues: 600 }), undefined); + assert.equal(digestPromptVariant({ maxCues: 1200 }), "c1200"); assert.equal( - digestPromptVariant({ timestampMode: "chunk-local", maxCues: 600 }), - "chunk-local+c600", + digestPromptVariant({ timestampMode: "absolute", maxCues: 1200 }), + "absolute+c1200", ); }); test("isSectionFresh: a chunk-size change invalidates the section", () => { // Measured: halving the chunk took qwen2.5:7b from 9.34 to 13.37 chapters per // hour and its worst coverage gap from 1:27:48 to 24:13. A change that large - // must not be able to hide behind an unchanged identity. + // must not be able to hide behind an unchanged identity. 600 is now the + // default (and so contributes nothing), so the departure tested here is a + // move BACK UP to the old 1200. assert.equal( isSectionFresh(machineRecord(), "chapters", { ...TARGET, - promptVariant: digestPromptVariant({ maxCues: 600 }), + promptVariant: digestPromptVariant({ maxCues: 1200 }), }), false, ); diff --git a/common/lib/digest.ts b/common/lib/digest.ts @@ -40,11 +40,27 @@ export type DigestLane = "local-gpu" | "remote-api"; // chunk against them without reaching a module that imports execa, and so the // identity helper below can compare against the default without a cycle. // -// 16k is comfortable on an 8 GB card (qwen2.5:7b KV cache ~56 KB/token -> -// ~0.9 GB at 16k, atop 4.7 GB of weights). digestPrompt.ts re-exports the cue -// count as DIGEST_MAX_CUES_PER_CHUNK and derives other sizes from it. -export const DEFAULT_DIGEST_NUM_CTX = 16384; -export const DEFAULT_DIGEST_MAX_CUES_PER_CHUNK = 1200; +// 8192, ON MEASUREMENT, not on comfort. Round 2 of the bake-off +// (plans/bakeoff/round2.md) ran qwen2.5:7b at both sizes over the same 6.96 +// audio-hours, chunk-local in both cases: +// +// @16384 11.1% zero-yield, 9.34 chapters/h, worst gap 1:27:48, 25.1 days +// @8192 11.8% zero-yield, 13.37 chapters/h, worst gap 0:24:13, 24.2 days +// +// Halving the window nearly halves the worst coverage gap and raises the +// segmentation rate by 43% at no throughput cost — twice as many calls each +// carry half the prompt, so the projected sweep is if anything shorter. 16k was +// the cautious choice and it measured worse. +// +// A 7B model at 8k is also well within an 8 GB card (qwen2.5:7b KV cache +// ~56 KB/token -> ~0.45 GB at 8k, atop 4.7 GB of weights). digestPrompt.ts +// re-exports the cue count as DIGEST_MAX_CUES_PER_CHUNK and derives other sizes +// from it; maxCuesForContext() scales the cue count with whatever numCtx an app +// is actually configured for, so these two must stay in proportion (~10 +// tokens/cue: 600 cues ~= 6k tokens of transcript inside an 8k window, leaving +// room for the prompt and the response). +export const DEFAULT_DIGEST_NUM_CTX = 8192; +export const DEFAULT_DIGEST_MAX_CUES_PER_CHUNK = 600; // How the transcript markers inside ONE chunk are numbered, and therefore what // the model is asked to copy. @@ -63,11 +79,24 @@ export const DEFAULT_DIGEST_MAX_CUES_PER_CHUNK = 1200; // caught all nine, which is exactly its job, but a caught error is still a lost // chunk, and >4 h videos are 8.2% of the corpus by count and 46% of its tokens. // chunk-local removes the large offset the model has to hold. Which mode is -// actually better is a BAKE-OFF QUESTION (common/bin/digest-bakeoff.ts), which -// is why both are shipped rather than one being pre-applied as a fix. +// actually better was a BAKE-OFF QUESTION (common/bin/digest-bakeoff.ts), which +// is why both were shipped rather than one being pre-applied as a fix. +// +// IT HAS BEEN ANSWERED. Round 2 (plans/bakeoff/round2.md), qwen2.5:7b@8192 over +// the same 6.96 audio-hours: +// +// absolute 29.4% zero-yield chunks, 10.35 chapters/h, worst gap 1:05:16, +// 47.5% of items rejected (65 of them out-of-range) +// chunk-local 11.8% zero-yield chunks, 13.37 chapters/h, worst gap 0:24:13, +// 19.1% rejected (13 out-of-range) +// +// gemma2:9b ranks the two the same way, so the effect is not model-specific. +// chunk-local is now the DEFAULT. Changing it changes generated output, so it +// was paired with a PROMPT_VERSION bump (digestPrompt.ts) rather than left to +// digestPromptVariant's relative-to-default derivation — see the note there. export const DIGEST_TIMESTAMP_MODES = ["absolute", "chunk-local"] as const; export type DigestTimestampMode = (typeof DIGEST_TIMESTAMP_MODES)[number]; -export const DEFAULT_DIGEST_TIMESTAMP_MODE: DigestTimestampMode = "absolute"; +export const DEFAULT_DIGEST_TIMESTAMP_MODE: DigestTimestampMode = "chunk-local"; export function isDigestTimestampMode(v: unknown): v is DigestTimestampMode { return ( diff --git a/common/lib/digestPrompt.test.ts b/common/lib/digestPrompt.test.ts @@ -2,12 +2,17 @@ import { test } from "node:test"; import assert from "node:assert/strict"; import { DIGEST_MAX_CUES_PER_CHUNK, + PROMPT_VERSION, buildChapterPrompt, buildTagPrompt, maxCuesForContext, promptOffsetSeconds, } from "./digestPrompt"; -import { DEFAULT_DIGEST_MAX_CUES_PER_CHUNK } from "./digest"; +import { + DEFAULT_DIGEST_MAX_CUES_PER_CHUNK, + DEFAULT_DIGEST_NUM_CTX, + DEFAULT_DIGEST_TIMESTAMP_MODE, +} from "./digest"; // The chunk size MUST track the configured context. Lowering num_ctx without // lowering it feeds ollama more transcript than its window holds, and ollama @@ -22,16 +27,40 @@ test("the two copies of the default chunk size agree", () => { }); test("maxCuesForContext scales the slice with the window", () => { - assert.equal(maxCuesForContext(16384), 1200); + // Ratio-pinned to the 8192 -> 600 default pair. assert.equal(maxCuesForContext(8192), 600); + assert.equal(maxCuesForContext(16384), 1200); assert.equal(maxCuesForContext(32768), 2400); + assert.equal(maxCuesForContext(4096), 300); // Unset or nonsense falls back to the default window, never to zero cues. - assert.equal(maxCuesForContext(undefined), 1200); - assert.equal(maxCuesForContext(0), 1200); + assert.equal(maxCuesForContext(undefined), DEFAULT_DIGEST_MAX_CUES_PER_CHUNK); + assert.equal(maxCuesForContext(0), DEFAULT_DIGEST_MAX_CUES_PER_CHUNK); // Floored, so a tiny window still yields a usable slice rather than one cue. assert.equal(maxCuesForContext(256), 100); }); +test("the shipped defaults are the configuration the bake-off measured best", () => { + // Round 2 (plans/bakeoff/round2.md) scored qwen2.5:7b@8192/chunk-local best on + // zero-yield rate, chapters/hour and worst coverage gap. The decision was made + // and then not shipped for a while — the defaults stayed at the losing + // 16384/absolute pair — so this pins the two constants that carry it. + assert.equal(DEFAULT_DIGEST_NUM_CTX, 8192); + assert.equal(DEFAULT_DIGEST_MAX_CUES_PER_CHUNK, 600); + assert.equal(DEFAULT_DIGEST_TIMESTAMP_MODE, "chunk-local"); +}); + +test("PROMPT_VERSION was bumped with the defaults change", () => { + // The pairing is the whole point. digestPromptVariant derives its string + // RELATIVE to the defaults, so a record written under the old defaults carries + // promptVariant: undefined and would compare EQUAL under the new ones — + // freezing bake-off-losing output into the corpus as permanently "fresh". + // isSectionFresh compares promptVersion, so the bump is what invalidates them. + assert.ok( + PROMPT_VERSION >= 2, + "changing DEFAULT_DIGEST_* without bumping PROMPT_VERSION silently skips every existing digest as fresh", + ); +}); + test("chunk-local re-bases the range the chapter prompt states", () => { const base = { title: "T", diff --git a/common/lib/digestPrompt.ts b/common/lib/digestPrompt.ts @@ -23,26 +23,42 @@ import { type DigestTimestampMode, } from "./digest"; -export const PROMPT_VERSION = 1; +export const PROMPT_VERSION = 2; -// NOT bumped by the addition of timestampMode below. In "absolute" mode — the -// default — the rendered prompt is byte-identical to what version 1 always -// produced, so every digest on disk stays fresh. The non-default shape is -// distinguished by provenance.promptVariant instead (see digestPromptVariant), -// which is the field that exists precisely so several shapes can be compared -// without a version bump invalidating the corpus between rounds. +// Version 1 -> 2 is a DEFAULTS change, and the bump is what makes it honest. +// +// The measured-best configuration (chunk-local timestamps, an 8192 context and +// the 600-cue chunk sized to it) is now the default. digestPromptVariant() +// derives its string RELATIVE TO THE DEFAULT CONSTANTS, so a value equal to the +// default contributes nothing and the variant comes out `undefined`. That is +// correct and stays as it is — but it means flipping the defaults alone would +// have been silently destructive: the 39 pilot digests were generated under +// absolute/1200 and recorded `promptVariant: undefined`, so under the new +// defaults they would compare EQUAL and be skipped as fresh forever. Output from +// the configuration the bake-off measured as worst would have been frozen into +// the corpus, indistinguishable from output of the best one. +// +// isSectionFresh compares promptVersion, so bumping it invalidates every section +// explicitly. That is the right tool for a default change; promptVariant is the +// tool for comparing several shapes concurrently WITHOUT invalidating the +// corpus, which is what a bake-off round needs. Cost here is 39 regenerations. +// +// (Version 1 also predates the addition of timestampMode. In "absolute" mode the +// rendered prompt was byte-identical to what version 1 always produced, which is +// why adding the knob did not itself require a bump.) -// Chunking. 12k usable tokens of transcript per call inside a 16k context leaves -// room for the prompt and the response. Expressed in CUES because that is what -// the chunker slices; ~10 tokens/cue is the corpus average, so 1200 cues ≈ 12k -// tokens. The overlap exists so a topic straddling a seam is visible whole to at -// least one call; the parser de-dups the resulting near-identical chapters. +// Chunking. ~6k usable tokens of transcript per call inside the default 8k +// context leaves room for the prompt and the response. Expressed in CUES because +// that is what the chunker slices; ~10 tokens/cue is the corpus average, so 600 +// cues ≈ 6k tokens. The overlap exists so a topic straddling a seam is visible +// whole to at least one call; the parser de-dups the resulting near-identical +// chapters. export const DIGEST_MAX_CUES_PER_CHUNK = DEFAULT_DIGEST_MAX_CUES_PER_CHUNK; // Cues per chunk, SIZED TO THE CONFIGURED CONTEXT. // -// The 1200-cue default is sized for a 16k window (~10 tokens/cue -> ~12k tokens -// of transcript, leaving room for the prompt and the response). Lowering +// The 600-cue default is sized for the default 8k window (~10 tokens/cue -> ~6k +// tokens of transcript, leaving room for the prompt and the response). Lowering // `numCtx` without lowering this feeds the engine more transcript than its // window holds, and ollama TRUNCATES SILENTLY — the model then summarizes // whatever fragment survived and the result looks like a bad model rather than a diff --git a/common/lib/settings.ts b/common/lib/settings.ts @@ -223,8 +223,8 @@ export type DigestSettings = { // the call count for a smaller payoff. sections: DigestSectionKind[]; // How each chunk's transcript markers are numbered — see DigestTimestampMode. - // A scored variable in the bake-off, not a pre-applied fix, so its effect on - // long videos is measured against the alternatives rather than assumed. + // Was a scored variable in the bake-off rather than a pre-applied fix; the + // measurement is in and "chunk-local" is now the shipped default. timestampMode: DigestTimestampMode; // A free-text label for a non-default prompt shape, folded into the recorded // provenance by digestPromptVariant(). Setting it invalidates every digest @@ -496,6 +496,10 @@ export function defaultDigest(): DigestSettings { longTailSeconds: DIGEST_LONG_TAIL_DEFAULT_SECONDS, localAppId: DEFAULT_DIGEST_APP_ID, remoteAppId: CLAUDE_DIGEST_APP_ID, + // Empty on purpose: every per-app knob falls through to its own default + // constant (resolveNumCtx -> DEFAULT_DIGEST_NUM_CTX, now 8192, and + // maxCuesForContext sizes the chunk to it). Seeding a copy of those values + // here would give the same number two homes and let them drift. apps: {}, digestsPaused: false, spendCapUsd: 0, diff --git a/editor/app/channels/[slug]/digestActions.ts b/editor/app/channels/[slug]/digestActions.ts @@ -82,7 +82,10 @@ export async function digestChannelAction( fn: async (onLog, signal, setProgress, ctx) => { const stat = await readChannelStat(paths, slug); if (stat) { - const missing = await countMissingDigests(paths, slug); + // The LANE matters: the progress target must be computed against the + // engine that is actually about to run, not whichever one the local + // setting names. + const missing = await countMissingDigests(paths, slug, undefined, lane); setProgress({ metric: "digests", initial: stat.digestCount ?? 0, @@ -157,7 +160,7 @@ export async function digestBucketAction( fn: async (onLog, signal, setProgress, ctx) => { const stat = await readChannelStat(paths, slug); if (stat) { - const missing = await countMissingDigests(paths, slug, cleaned); + const missing = await countMissingDigests(paths, slug, cleaned, lane); setProgress({ metric: "digests", initial: stat.digestCount ?? 0, diff --git a/editor/app/channels/[slug]/lib/stageStatus.ts b/editor/app/channels/[slug]/lib/stageStatus.ts @@ -332,9 +332,12 @@ export function computeStageStatuses( }), }; - // Videos with a transcript but no digest at the current prompt/model identity. - // Counted as pending work rather than merely informational: unlike the - // auto-captions lane, every transcribed video is eventually meant to have one. + // Videos with a transcript whose digest is missing OR stale against the local + // lane's current identity — the snapshot's noDigest bucket now computes that + // full freshness target (channelSnapshot.ts), so this label and that bucket + // finally mean the same thing. Counted as pending work rather than merely + // informational: unlike the auto-captions lane, every transcribed video is + // eventually meant to have one. const digestRunning = runningByStage.has("digest"); const digestPending = buckets.noDigest.length; const digest: StageStatus = { @@ -349,10 +352,10 @@ export function computeStageStatuses( : digestPending > 0 ? pluralize( digestPending, - "transcript without a digest", - "transcripts without a digest", + "transcript needs a digest", + "transcripts need a digest", ) - : "All digested.", + : "All digested at the current settings.", tone: pickTone({ running: digestRunning, pending: digestPending, diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -429,22 +429,73 @@ reports itself skipped corpus-wide: `6,970 shorts × 76,318 videos ≈ 531,936,4 pairs`. It still runs, unchanged, at shorts scale. Making it scale is the one piece genuinely deferred — see STATE.md. -### The finding that matters most: 63% of title matches are NOT duplicates +### 63% of title matches were rejected at 0.6 — and most of those rejections were wrong Of 8,352 pairs nominated by an exact normalized title *and* a compatible runtime, -**5,250 were rejected once their transcripts were compared.** Only 2,717 held up. - -This directly corrects the planning assumption that ~8,204 title-level -redundancies were available to prune from the digest sweep. The title-level -*count* was right; the inference that they are duplicates was not. Banked saving -is at most 2,717 pairs, and only 1,599 of those are timing-aligned enough for a -shared digest to be placed correctly — roughly **a fifth of what was projected**. - -Caveat worth testing before treating that as final: the near threshold is a 5-gram -Jaccard at 0.6, and the two sides of a cross-platform mirror are often different -ASR engines (YouTube auto-captions vs. whisper). Some of the 5,250 may be real -mirrors whose transcriptions simply disagree at the 5-gram level. Sensitivity to -that threshold has **not** been measured. See STATE.md. +**5,250 were rejected once their transcripts were compared** at the default 0.6 +5-gram Jaccard. Only 2,717 held up. + +That was read as "the title-level count was right; the inference that they are +duplicates was not." **Measurement has since shown the opposite: the threshold was +wrong, not the inference.** See the next section. + +### MEASURED: the 0.6 near threshold was rejecting real cross-platform mirrors + +Re-run corpus-wide, title blocking, everything else identical, only `--near` +changed (2026-07-27): + +| | `--near 0.6` (default) | `--near 0.35` | +| --- | --- | --- | +| Nominated pairs | 8,352 | 8,352 | +| **Confirmed by content** | 2,717 | **7,329** | +| Rejected by content | 5,250 | **638** | +| Untestable suspects | 385 | 385 | +| Clusters written | 2,994 / 6,044 videos | **7,450 / 15,030 videos** | +| `needsReview` | 335 | 334 | +| **Timing-aligned mirrors** | 1,599 of 2,689 | **3,952 of 7,221** | +| Wall clock | 206 s | 249 s | + +The confirmed set grows **+170%** and is a strict superset: **0 videos** present in +the 0.6 clusters are absent from the 0.35 clusters. + +**The 4,456 entirely-new clusters were inspected, not just counted**, and they are +real mirrors. Evidence, in descending order of strength: + +- **4,423 of 4,456 (99.3%) are cross-platform AND cross-channel.** 3,954 pair a + channel with its own platform-suffixed sibling (`the-quartering` ↔ + `the-quartering-rumble`); of the 502 that don't, the sampled ones pair a + creator's alternate channels (`jeremy-hambly` / `quartering-live` / + `the-quartering`, `HasanAbiVODs` / `HasanAbiVODsbackup`). +- **4,290 of 4,456 have a byte-identical title** across every copy; 4,023 agree on + runtime within 2 s; 3,572 share an upload date. +- The score histogram is a tight band at **0.45–0.60** (1,995 at 0.55–0.60, 1,481 + at 0.50–0.55, 665 at 0.45–0.50) — i.e. they pile up *just under* the old cutoff, + which is the signature of a systematic offset, not of noise. +- The **138 riskiest** (score < 0.42) and the **502 non-sibling** clusters were + sampled across the whole band. Every one examined is the same recording: same + title, same runtime to the second, same upload date, mirrored platform. + +The mechanism is the one the caveat predicted: the two sides of a cross-platform +mirror are transcribed by **different ASR engines**, and a 5-gram Jaccard is +unforgiving of word-level disagreement. Two transcripts of the same audio from +different engines land at **~0.35–0.60**, not ≥ 0.6. The old default was set for +same-engine text and silently failed the cross-platform case, which is precisely +the case duplicate detection exists to catch. + +**Consequence for the digest backfill.** The ceiling on shared-digest savings is +not 2,717 pairs. At 0.35 it is **7,329 confirmed pairs, of which 3,952 are +timing-aligned** enough for a shared digest to be placed correctly — **2.5× the +1,599 previously banked**. Against a corpus of 76,318 videos that is ~5.2% of the +sweep avoidable by sharing rather than ~2.1%. + +**Not yet changed:** `NEAR_THRESHOLD_DEFAULT` still ships at 0.6. Lowering it +changes what every built site asserts to readers, so it is recorded here as a +measured recommendation rather than applied as a side effect of a digest task. +See STATE.md. + +Reproduce with: +`pnpm exec tsx bin/duplicate-shorts.ts --all-durations --blocking title --near 0.35` +(back up `transcripts/duplicates.json` first — the run overwrites it). ### What this unblocks diff --git a/plans/STATE.md b/plans/STATE.md @@ -43,13 +43,31 @@ relitigate. Earlier entries (fabric deferred, ollama-direct as the structured de Claude Code on a second queue key, digests in their own page tree, per-section provenance, transcripts never rewritten) still stand and are unchanged. -**`chunk-local` timestamps are adopted, on measurement.** Each chunk's transcript is -re-based to `00:00:00` and the parser adds the offset back before any guard runs. Measured -on the long tail it cut zero-yield chunks from 33.3% to 11.1% (qwen2.5@16k), 29.4% to 11.8% -(qwen2.5@8k) and 17.7% to 0% (gemma2), with the rejection rate falling in every case. The -working hypothesis — the model cannot hold a large absolute offset across a long chunk and -reverts to counting from zero — is confirmed. It was shipped as a *scored variable*, not a -pre-applied fix, which is why there are numbers to quote. +**`chunk-local` timestamps are adopted, on measurement — and (2026-07-27) SHIPPED as the +default.** Each chunk's transcript is re-based to `00:00:00` and the parser adds the offset +back before any guard runs. Measured on the long tail it cut zero-yield chunks from 33.3% to +11.1% (qwen2.5@16k), 29.4% to 11.8% (qwen2.5@8k) and 17.7% to 0% (gemma2), with the +rejection rate falling in every case. The working hypothesis — the model cannot hold a large +absolute offset across a long chunk and reverts to counting from zero — is confirmed. It was +shipped as a *scored variable*, not a pre-applied fix, which is why there are numbers to +quote. + +For a while it was adopted *on paper only*: `DEFAULT_DIGEST_TIMESTAMP_MODE` stayed +`"absolute"` and `DEFAULT_DIGEST_NUM_CTX` stayed 16384, so a fresh install — and the pilot — +ran the configuration round 2 measured as **worst**. Both defaults now carry the decision +(`chunk-local`, 8192, 600 cues). + +**The mechanism for that change was a `PROMPT_VERSION` bump (1 → 2), and it had to be.** +`digestPromptVariant()` derives the variant *relative to the default constants*, so a value +equal to the default contributes nothing and the variant is `undefined`. Flipping the +defaults alone would therefore have made the 39 pilot digests — written under +absolute/1200 with `promptVariant: undefined` — compare **fresh** under the new defaults: +bake-off-losing output frozen into the corpus, indistinguishable from the best config's. +`isSectionFresh` compares `promptVersion`, so bumping it invalidates every section +explicitly. That is the right division of labour: **`promptVersion` for a default change, +`promptVariant` for several shapes coexisting during a bake-off round.** Cost: 39 +regenerations. `digestPrompt.test.ts` pins both the new defaults and the bump so the pairing +cannot be broken silently. **Chunk size is sized to `num_ctx`, and it is a bigger lever than expected.** Halving the context (16k → 8k, 1200 → 600 cues) took qwen2.5:7b from 9.34 to 13.37 chapters/hour and its @@ -65,6 +83,28 @@ now derives it. written before the field existed still compares fresh. The alternative — an operator remembering to bump a label — fails silently by skipping the whole corpus as "fresh". +**One freshness target, three callers (2026-07-27).** `resolveDigestTarget()` +(`common/controller/digestTarget.ts`) is now the only place a target is derived, and the +batch runner, `countMissingDigests` and the channel snapshot's `noDigest` bucket all use it. +Two counters were previously lying: + +- `noDigest` asked only `engine && hasItems` — "is there an ai-digest.json with something in + it?" — while its own label and `stageStatus.ts` claimed an identity check. After any config + change the Digest stage read **"All digested"** while `countMissingDigests` reported the + whole channel stale. The two could disagree by an entire channel, which is exactly how a + sweep reports "nothing to do" on a corpus that needs redoing. It now computes the real + target; the per-video cost is zero extra I/O (the sidecar is already loaded and + `isSectionFresh` is pure), and only the target itself is new work, resolved once per + channel. Coverage-by-engine (`digestEngines`) stays deliberately identity-**blind** — a + section made by an older prompt version was still made by that engine. +- `countMissingDigests` always resolved `settings.localAppId`, so a metered-lane job's + progress bar was sized against the *local* engine's identity (and, once `numCtx` joined the + identity, the local app's context size). It takes the lane as a parameter now; both call + sites in `digestActions.ts` already knew which lane they were. + +A digest shared from a duplicate cluster's canonical member counts as done in the bucket — +the canonical member's own freshness drives regeneration and the share is re-applied from it. + **The sweep is ~25 days, not ~64.** Round 1 measured 59 days on short+medium; Round 2 measured ~25 on the long tail, because long videos amortise the fixed per-call overhead and the corpus is dominated by them. Throughput is no longer the binding constraint it was @@ -88,10 +128,15 @@ so choosing one trades recall against runtime rather than correctness. This is w nominating aggressively safe — and it earns its keep: **5,250 of 8,352 title-nominated pairs (63%) were rejected once their transcripts were compared.** -**The projected digest saving was overstated by ~5x, and this is the headline finding.** The -~8,204 "title-level redundancies" were real as a *count* of same-title pairs, but only 2,717 -survive content comparison, and only 1,599 of those are timing-aligned enough for a shared -digest to land correctly. Plan the backfill on those numbers, not the old ones. +**SUPERSEDED (2026-07-27) — that 63% was mostly the threshold, not the corpus.** This section +originally read "the projected digest saving was overstated by ~5x": of ~8,204 title-level +redundancies only 2,717 survived content comparison, and only 1,599 of those were +timing-aligned. Measuring `--near` (the caveat that was attached to it) showed the rejections +were overwhelmingly **real cross-platform mirrors transcribed by different ASR engines**. At +`--near 0.35` the confirmed set is **7,329 pairs / 3,952 aligned** — 2.5× the banked saving, +and inspection of all 4,456 new clusters found no false positives. Plan the backfill on +**7,329 / 3,952**, and see the near-threshold entry under Open questions before treating 0.6 +as settled. Full evidence in FACTS.md. **A previous decision is deliberately reversed: title+duration clusters DO form.** The old rule was that only content-confirmed tiers cluster, because duration coincidence alone @@ -132,13 +177,26 @@ playwright `webServer` with `OLLAMA_URL` pointed at it in **both** `dev:test` an they save. gemma2 is hard-capped at an 8192 context by the model. - **The `verylong` (>6 h) stratum is unmeasured.** Round 2 used the `long` bucket (~3.5 h). The sample already contains an 8 h video for it. -- **Is the 0.6 near threshold rejecting real mirrors?** 63% of title-nominated pairs failed - content comparison. Some of that is genuinely two episodes of a daily show — which is the - point — but the two sides of a cross-platform mirror are often *different ASR engines* - (YouTube auto-captions vs. whisper), and a 5-gram Jaccard is unforgiving of word-level - disagreement. Sensitivity to `--near` has **not** been measured. Cheap to answer: re-run - `--blocking title --near 0.35` and diff the confirmed set. Do this before treating the - 2,717 figure as the ceiling on the digest saving. +- **ANSWERED (2026-07-27): yes — the 0.6 near threshold was rejecting real mirrors, in bulk.** + Re-running `--blocking title --near 0.35` corpus-wide takes confirmed pairs from **2,717 to + 7,329** (+170%) and timing-aligned mirrors from **1,599 to 3,952**, a strict superset (no + video in the 0.6 clusters is absent from the 0.35 ones). The 4,456 new clusters were + inspected, not merely counted: **99.3% are cross-platform *and* cross-channel**, 96% have a + byte-identical title, 90% agree on runtime within 2 s, and their scores pile into a tight + **0.45–0.60** band just under the old cutoff. Sampling the 138 riskiest (< 0.42) and all + 502 non-sibling cases found **no false positives**. Mechanism confirmed as hypothesised: + two *different ASR engines* transcribing the same audio agree at ~0.35–0.60 on 5-gram + Jaccard, so 0.6 was tuned for same-engine text and structurally failed the cross-platform + case. Full table + reproduction in FACTS.md. **The caveat is deleted, not carried.** + + **Now a decision, not a question: should `NEAR_THRESHOLD_DEFAULT` drop from 0.6?** The + measurement says yes and the digest-saving ceiling rises 2.5× if it does. It was *not* + changed here, deliberately — detection output feeds every built site's Duplicates page and + search badges, so loosening it changes what the archive asserts to readers, and that is its + own reviewable change rather than a side effect of a digest task. What is still unmeasured + is where the floor is: 0.35 was the single probe, and the 638 pairs it still rejects have + not been examined. Recommended next: probe 0.25 / 0.45 to bracket the knee, eyeball the + remaining rejections, then change the default in its own commit. - **Recall of the title pre-filter — now measured, and the gap is 11%.** Running both strategies corpus-wide and diffing the clusters: duration blocking finds **672 videos title blocking misses** (the re-titled mirrors it structurally cannot see), title finds 68 that