Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit b0c16f2af9326934652dd4591cff6d3b4391f0ed
parent b2ad3270da6d3b219fc7466b059617dbf73e4576
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Wed, 29 Jul 2026 11:43:19 -0400

Correct the roadmap's gate list, and record what the measurements changed

Three corrections the exploration and the measurements forced:

- **Phase 1.5 is not a backfill gate.** `digestContext-server.ts:31-35` hashes
  the EMPTY note to a stable value, so adding a channel note later invalidates
  only that channel. What would invalidate the corpus is bumping the
  `digest-context-v1` prefix — so the gate is a DECISION to take before
  sweeping, not work to do first. PLAN.md said the opposite in two places.
- **"Do not backfill before 2.5 and 11a" named the wrong things.** Both were
  largely landed. The things actually missing were never on the list: no
  launcher, no boot resume, no pause writer, no GPU arbitration.
- **`updateDuplicateOverride` had zero callers**, so a fails-closed gate had no
  way to be opened and ~166 clusters could never be confirmed by anyone. Worth
  generalising, and recorded as such: "the extension point exists" is not the
  same as "the feature works".

FACTS.md gains the four-way threshold bracket with its marginal-band tables, the
audio-hour pricing of the sweep at each threshold, the heaviest-channel work
queue, and the 14.5% id-vs-directory measurement.

The correction that matters most is to a claim these docs made themselves:
duplicate sharing is worth ~1.7 sweep days of 80, not "~11% of the sweep",
because cluster members are 19% of the corpus by video count and 4.3% of its
audio-hours. The stale "consequence for the digest backfill" paragraph is marked
overstated rather than deleted — it was right about pairs and wrong about time,
and that distinction is the reusable lesson.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Diffstat:
MPLAN.md | 32+++++++++++++++++++++++++-------
Mplans/FACTS.md | 97+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++----
Mplans/STATE.md | 142+++++++++++++++++++++++++++++++++++++++++++++++++++++++++----------------------
3 files changed, 221 insertions(+), 50 deletions(-)

diff --git a/PLAN.md b/PLAN.md @@ -246,7 +246,10 @@ Markdown with optional frontmatter: `hosts`, `recurring_guests`, inference silently degrade every downstream summary for a channel, invisibly, because the output would still look plausible. - Hash the resolved context into `ai-digest.json.contextHash` so editing notes correctly - marks digests stale. **This is why backfill must not start before 1.5 exists.** + marks digests stale. **Corrected 2026-07-29:** this is NOT why backfill must wait. The + empty note already hashes to a stable value, so a note added later invalidates only its own + channel. The corpus-wide risk is the deliberate one — bumping the `digest-context-v1` hash + prefix — so what must precede a sweep is that DECISION, not this phase. ### Phase 2 — Digests as a first-class corpus — **DONE (2026-07-29)** @@ -573,16 +576,31 @@ browser only to diagnose failures. ### Ordering traps -- **Do not backfill before 1.5.** Channel context feeds `contextHash`, which correctly - invalidates every digest generated without it. At 15–60 s per video that is a costly redo. -- **Do not backfill before 2.5 and 11a.** An unobservable multi-day sweep cannot be tuned or - safely interrupted, and one with no review path is a one-shot gamble on prompt quality. +- **CORRECTED (2026-07-29): 1.5 is not a backfill gate — a hash-prefix decision is.** + `digestContext-server.ts:31-35` hashes the EMPTY note to a stable value, so adding a + channel note later invalidates only that channel, not the corpus. What would invalidate + everything is bumping the `digest-context-v1` prefix (`:42`), which 1.5-as-specced does. + So the gate is a decision to take before sweeping, not work to do first. +- **CORRECTED (2026-07-29): "before 2.5 and 11a" named the wrong things.** Both were largely + landed — the metric union, `METRIC_PREFIX`, per-job progress and the `noDigest` bucket all + existed. The things actually missing were never on this list: **no corpus-wide launcher, + no boot-time resume, no pause writer, and no GPU arbitration**. All four are now built + (`controller/digestSweep.ts`, `controller/digestYield.ts`, `instrumentation.ts`, + `DigestSweepControls`). The remaining honest gate is the review path, and its minimum — + persisted failure warnings plus a `digestWarnings` bucket — is in. - **Benchmark before committing to backfill.** The smoke test measured 15.5 s for a ~2.7k token chunk; a 176-minute podcast needs roughly a dozen chunks. Budget minutes per long video and multiply by corpus size. The remote lane changes this arithmetic — measure both, then split the backlog between them. -- **Fix the duplicated metric union before extending `JobProgressMetric`**, or Phase 2.5's - compiler-driven audit silently misses a consumer. +- **RESOLVED: the duplicated metric union is fixed.** `RunningJobsList` imports the union + and `MonitorWidget` uses a `METRIC_PREFIX` lookup table, so extending it is now + compiler-checked. +- **Price a cost lever in AUDIO-HOURS before believing it** (`common/bin/digest-plan.ts`). + Measured 2026-07-29: duplicate-cluster sharing is worth **~1.7 sweep days of 80**, not the + "~11%" `digestSharing.ts`'s header claims — cluster members are 19% of the corpus by video + count but **4.3% of its audio-hours**, because mirrors skew SHORT and the sweep is + dominated by unclustered long-form VODs. GPU contention with whisper is worth ~56 days on + the same measurement, roughly fifty times more than every duplicate lever combined. ### Hardware and prerequisites diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -490,15 +490,104 @@ timing-aligned** enough for a shared digest to be placed correctly — **2.5× t 1,599 previously banked**. Against a corpus of 76,318 videos that is ~5.2% of the sweep avoidable by sharing rather than ~2.1%. -**Not yet changed:** `NEAR_THRESHOLD_DEFAULT` still ships at 0.6. Lowering it -changes what every built site asserts to readers, so it is recorded here as a -measured recommendation rather than applied as a side effect of a digest task. -See STATE.md. +**CHANGED 2026-07-29: `DEFAULT_NEAR_THRESHOLD` now ships at 0.35**, in its own +commit, after bracketing it at four values — see the next section. Note that the +"consequence for the digest backfill" above was **overstated**: the saving is +real in pair count but nearly absent in audio-hours, which is the unit the sweep +is actually priced in. ~5.2% of pairs is ~1.7 sweep days of 80. Reproduce with: `pnpm exec tsx bin/duplicate-shorts.ts --all-durations --blocking title --near 0.35` (back up `transcripts/duplicates.json` first — the run overwrites it). +### Bracketing the near threshold — 4 runs, identical inputs (measured 2026-07-29) + +Corpus-wide, `--all-durations --blocking title`, only `--near` varied. 76,354 videos +scanned, 8,352 pairs nominated in every run. + +| `--near` | confirmed | rejected | clusters | videos | aligned | needsReview | wall | +| --- | --- | --- | --- | --- | --- | --- | --- | +| 0.6 | 2,736 | 5,404 | 2,846 | 5,746 | 4,273 | 168 | 124 s | +| 0.45 | 7,147 | 993 | 7,110 | 14,344 | 10,747 | 166 | 199 s | +| **0.35** | **7,483** | **657** | **7,434** | **14,997** | **11,217** | 166 | 225 s | +| 0.25 | 7,542 | 598 | 7,493 | 15,115 | 11,297 | 166 | ~230 s | + +Every step down is a **strict superset** — 0 videos lost at any step. Returns collapse: +0.6→0.45 admits **4,264** entirely-new clusters, 0.45→0.35 admits **324**, 0.35→0.25 admits +**59**. + +The marginal bands were read, not counted: + +| band | new clusters | byte-identical title | runtime ±2 s | cross-platform | SAME-channel | +| --- | --- | --- | --- | --- | --- | +| 0.6 → 0.45 | 4,264 | 96.2% | 91.1% | 99.5% | **3** | +| 0.45 → 0.35 | 324 | 97.2% | 83.3% | 96.0% | **2** | +| 0.35 → 0.25 | 59 | 94.9% | 86.4% | 93.2% | **0** | + +The riskiest cases in each band were printed individually. In the 0.6→0.45 band they are all +YouTube↔Rumble pairs agreeing on title (modulo whitespace), runtime to the second, and upload +date. The 0.45→0.35 band contains the only two arguable cases in the whole sweep: a +`HasanAbiVODs` pair with the same title but runtimes 9 minutes apart, and an `omnibased` pair +with identical title/duration/date. Neither is a clear false positive; the first would be +refused by the alignment gate regardless. + +**Adopted: 0.35.** The recall knee is at **0.45** — the value to take if a more conservative +assertion is ever wanted. + +### The digest sweep, priced in AUDIO-HOURS (measured 2026-07-29) + +`common/bin/digest-plan.ts --no-freshness`, at the measured 90 s/audio-hour. Total digestable +corpus: **77,298 audio-hours** over 73,367 videos (2,987 have no transcript), which reproduces +the roadmap's ~81-day headline independently. + +| `--near` | canonical | mirror-aligned (free) | mirror-unaligned | unclustered | **to generate** | **days** | +| --- | --- | --- | --- | --- | --- | --- | +| 0.6 | 853 h | 494 h | 344 h | 75,607 h | 76,804 h | 80.0 | +| 0.45 | 3,472 h | 1,448 h | 1,866 h | 70,513 h | 75,851 h | 79.0 | +| 0.35 | 4,070 h | 1,661 h | 2,206 h | 69,361 h | 75,638 h | 78.8 | +| 0.25 | 4,137 h | 1,683 h | 2,242 h | 69,236 h | 75,616 h | 78.8 | + +**Duplicate sharing is not a meaningful cost lever, and `digestSharing.ts`'s "~11% of the +sweep" header is wrong.** Cluster members are 19% of the corpus by video count but **4.3% of +its audio-hours**: mirrors skew SHORT while the sweep is dominated by unclustered long-form +VODs. Loosening the threshold as far as it goes moves the sweep from 80.0 to 78.8 days. + +Note the `mirror-unaligned` column: only ~55% of mirrors pass the alignment gate, so counting +all mirrors as free would overstate the saving by nearly half. The plan splits them. + +Heaviest channels (the sweep's work-queue order), audio-hours to generate at 0.35: + +| channel | h | days | +| --- | --- | --- | +| HasanAbiVODs3 | 8,331 | 8.7 | +| omnibased | 7,783 | 8.1 | +| HasanAbiVODs | 5,380 | 5.6 | +| rekietalaw | 4,993 | 5.2 | +| destiny | 4,020 | 4.2 | +| shondo-vods | 3,332 | 3.5 | + +`HasanAbiVODs3` alone outweighs every duplicate mirror in the corpus combined. + +### The id-vs-directory mismatch (measured 2026-07-29) + +A slug is `${channelSlug}/${id}` and an id is **not** the on-disk directory name for +**11,175 of 77,106 indexed videos (14.5%)**, concentrated almost entirely in Rumble +re-uploads: + +| channel | dir !== id | +| --- | --- | +| the-quartering-rumble | 7,870 (of 7,870 — every video) | +| leaflit-rumble | 838 | +| midwestly | 758 | +| rekietalaw-rumble | 641 | +| omnimirror | 522 | +| cornbreadman | 337 | + +Verified directly: for all 7,870 `the-quartering-rumble` entries, `stat.slug === +channel/id` and `stat.slug !== channel/dir`. `digestBatch` looked the cluster plan up by +DIRECTORY, so every one of those lookups missed silently and the mirror regenerated. See the +`videoDir` field on `DuplicateVideoRef`. + ### What this unblocks Corpus-wide detection is no longer off. `buildDigestClusterPlan` reads diff --git a/plans/STATE.md b/plans/STATE.md @@ -3,27 +3,41 @@ The working memory for the local-AI derived-corpus work. Rewritten at the end of every session, before context is cleared. See [`README.md`](README.md) for the protocol. -**Last updated:** 2026-07-29 — **Phases 2 and 3: the digest layer SHIPS.** The 102 digests +**Last updated:** 2026-07-29 — **the backfill is at the starting line.** A corpus-wide sweep +can now be started from one control, survives a server restart, can be paused without being +lost, yields the GPU to transcription, and reports coverage and an ETA in the unit the work +is actually priced in. Stages A–D of the "cheapest-first" plan are all in; see the decisions +below for what the measurements changed. + +**The headline is a negative result, and it should be read before planning any more cost +work.** The plan's own premise — shrink the work-list via duplicate sharing — does not pay. +Priced in audio-hours by the new `common/bin/digest-plan.ts`: + +| `--near` | clusters | to generate | sweep days | saved by sharing | +| --- | --- | --- | --- | --- | +| 0.6 (was) | 2,846 | 76,804 h | 80.0 | 0.5 d | +| 0.45 | 7,110 | 75,851 h | 79.0 | 1.5 d | +| **0.35 (now)** | **7,434** | **75,638 h** | **78.8** | **1.7 d** | +| 0.25 | 7,493 | 75,616 h | 78.8 | 1.8 d | + +Loosening the threshold as far as it goes buys **1.2 days of 80**. The plan assumed mirrors +skew long; **they skew short**. Cluster members are 19% of the corpus by video count and +**4.3% of its audio-hours** — the sweep is dominated by unclustered long-form VODs, and +`HasanAbiVODs3` alone (8,331 audio-hours) outweighs every mirror in the corpus combined. +Only ~55% of mirrors pass the alignment gate, so half the nominal saving is refused anyway. + +**The lever that does pay is GPU arbitration**: 90 s/audio-hour measured against 27 s +projected on an idle box is ~80 days versus ~24 — about fifty times every duplicate lever +put together. It is now implemented (`controller/digestYield.ts`) and **still unmeasured in +production**; measuring it is the single most valuable next action. + +**Previous entry —** 2026-07-29 — **Phases 2 and 3: the digest layer SHIPS.** The 102 digests already on disk now travel generation → LMDB → a shared `/digests/<slug>/` page tree → compose → the viewer, where a reader can open a Digest panel, click a chapter and seek to it. Proven end to end against the real corpus, not fixtures: `community-notes/v2cywen` renders its 20 real chapters (first at 01:09:48 — the validation run's worst coverage gap) and the playhead follows a click. The control is hidden on the 99.9% of videos with no digest, and a -borrowed digest is labelled as borrowed. Deliberately independent of the three backfill gates -(1.5 context, 2.5 observability, 11a review), none of which moved. - -Verification that matters: common 356 tests (346 + 10 new), `tsc` clean in all three -packages, export playwright **150/150** (141 + 9 new), a no-op rebuild skips digest pages, -and a hand-written `ai-digest.overrides.json` produces exactly `~1 changed` and ships the -corrected title — the `digestMs` max-of-two-sidecars requirement, which nothing else tests. - -**Previous entry —** 2026-07-27 — Stage B2: the digest layer is now INSPECTABLE and its defaults -are the ones the bake-off measured best. Settings gained a real Digest section, each video -page gained a review panel (provenance, warnings, `derivedFrom`, staleness), the two digest -counters were made to agree, `chunk-local`/8192/600 shipped as defaults behind a -`PROMPT_VERSION` bump, the pilot's "un-created job" was root-caused (a pre-hydration lost -click) and fixed, and a **102-video validation run** was completed on the real corpus. Also: -the 0.6 duplicate near-threshold was measured and found to be rejecting real mirrors in bulk. +borrowed digest is labelled as borrowed. --- @@ -33,9 +47,9 @@ the 0.6 duplicate near-threshold was measured and found to be rejecting real mir | --- | --- | --- | | 0 · Benchmark transcription engines | not started | Still unmeasured. The DIGEST bake-off is done; this is the separate transcription one. | | 1 · Generation harness | **done + validated** | Spine in `8c041fd`; correctness + bake-off + e2e in Stage B1; operator surfaces + measured-best defaults + 102-video validation run in Stage B2. | -| 1.5 · Channel context | not started | `contextHash` is plumbed and empty, so notes can be added without invalidating the corpus. | +| 1.5 · Channel context | not started | **Not a backfill gate** — the empty note hashes stably, so a note added later invalidates only its own channel. The gate is the `digest-context-v1` prefix decision. | | 2 · Digest corpus in build | **done** | Shared `/digests/<slug>/` page tree, `digests` + `digestPageHashes` + `channelDigestStats` sub-DBs at `SCHEMA_VERSION` 13, `digestMs` in the mtime record and the channel signature, compose reconcile, `CORPUS_SPEC_VERSION` 3. | -| 2.5 · Observability | **partly landed** | Its stated first step is already done — see the correction below. | +| 2.5 · Observability | **done** | Corpus coverage on the dashboard + widget, `noDigest` counters, a digest instrument, audio-hour ETAs, and the two progress bugs fixed. | | 3 · Viewer `?vm=digest` | **done** | Not `?vm=summary` — see the naming decision below. | | 4 · Search indexing | not started | | | 5 · Auto-queue | not started | | @@ -44,7 +58,7 @@ the 0.6 duplicate near-threshold was measured and found to be rejecting real mir | 8 · Visibility policy | not started | | | 9 · Attribution + quote filtering | not started | | | 10 · Lead with the derived corpus | not started | | -| 11a · Review queue | not started | **Land before backfilling.** `warnings[]` is the data it reads. | +| 11a · Review queue | **minimum landed** | Total failures now persist warnings (they used to persist none), `buckets.digestWarnings` + a `digest_warnings` filter + an /actionable section, and duplicate-cluster confirm/reject. Approve/dismiss state deferred — needs new persistence. | | 11b · Viewer feedback | not started | | **Three corrections to this document's own claims, verified against the code 2026-07-29.** @@ -63,26 +77,26 @@ Exploration for Phase 2 found the planning docs describing more unbuilt work tha line and every cross-channel surface gets a counter it is already paying for. Left unchanged here deliberately: it is Phase 2.5's to take, not a digest-corpus side effect. -**Recommended next**, in the order the validation run argues for: - -1. **Phase 2.5 observability, and treat the throughput finding as its first requirement.** - The run measured **90 s per audio-hour**, projecting **~81 sweep days** against round 2's - 24.2 — because short videos cost ~2× per audio-hour and the box was sharing a GPU with - `auto-transcribe`. A sweep of that length needs a coverage instrument and a scheduling - story (run the digest lane *against* the transcription lanes, not beside them) before it - starts, not after. Numbers in FACTS.md. -2. **Phase 11a review queue.** `warnings[]` is now proven as its data source: 93% of all - rejections are `out-of-range` from a single guard, and the per-video panel already showed - that the worst coverage gap in the run is a video where 13 proposed chapters were clamped - away rather than never proposed. That distinction is what a queue must surface. -3. **Decide the duplicate near-threshold in its own commit.** 0.6 measurably rejects real - cross-platform mirrors; 0.35 recovers 4,456 clusters with no observed false positives and - lifts the shareable-digest ceiling 2.5×. Probe 0.25/0.45 to bracket the knee first. -4. **Phase 1.5 channel context** — still a backfill gate (`contextHash` staleness), and still - not required for a validation run. - -Deferred from Stage B2 and still open: the `/actionable` duplicate-cluster confirm/reject -buttons (PLAN.md Phase 11a, extension points recorded there) and the PipelineBand instrument. +**Recommended next**, in the order the measurements argue for: + +1. **MEASURE THE GPU YIELD IN PRODUCTION.** Everything else is second by a factor of fifty. + Start a transcription job during a digest sweep, confirm the digest lane idles rather than + contending, and measure seconds-per-audio-hour with `yieldToTranscription` on and off. If + the 2.8x is real the sweep is ~24 days, not ~80, and every other estimate in these docs + changes. If it is NOT real, the contention hypothesis is wrong and the 90 s/audio-hour + needs a different explanation — which is just as valuable to know before spending 80 days. +2. **Run the sweep for a day and watch it.** The launcher, resume, pause and coverage + readouts are all built and unit-tested but have not driven a real multi-channel run. Kill + the editor mid-run and confirm it resumes with zero rework (eligibility is re-derived from + disk, so the correct result is zero). +3. **Decide the `digest-context-v1` hash-prefix question before sweeping.** This is the real + 1.5 gate, and it is a decision, not work: bumping the prefix invalidates the entire + corpus, so it must happen before 80 GPU-days go in, or not at all. +4. **Then Phase 4 search indexing / Phase 7 tags,** which are starved until coverage exists. + +Deferred deliberately: Phase 11a approve/dismiss state (needs a new sidecar field or sibling +file), Phase 1.5 as fully specced, Phase 6 Ollama `/ask` (genuinely independent — a good +parallel task), Phase 0 transcription benchmark (a separate bottleneck). --- @@ -93,6 +107,56 @@ relitigate. Earlier entries (fabric deferred, ollama-direct as the structured de Claude Code on a second queue key, digests in their own page tree, per-section provenance, transcripts never rewritten) still stand and are unchanged. +**Cluster sharing was DEAD for 14.5% of the corpus — the half it was written for +(2026-07-29).** `digestBatch` looked the plan up with `${channelSlug}/${directoryName}`; the +plan is keyed `${channelSlug}/${metadataId}`. Those differ for **11,175 of 77,106 videos**, +and not evenly: **7,870 are `the-quartering-rumble`, where dir !== id for every single +video** — precisely the mirror set `digestSharing.ts`'s header cites as the win. A miss is +silent and looks like success (the mirror is simply not recognised, so it regenerates), which +is why it survived a bake-off, a pilot and a 102-video validation run. `videoDirForSlug` had +the same bug from the other side. Fixed by recording `DuplicateVideoRef.videoDir` — which the +detector already knew and threw away — only where it differs from `id`. + +**A cost lever is only real in AUDIO-HOURS (2026-07-29).** See the table at the top. The +general lesson, and it is the same one the validation run taught about throughput: **the +corpus's video count and its audio-hours are differently distributed, so any estimate that +counts videos is measuring the wrong thing.** `bin/digest-plan.ts` exists so the next lever +gets priced before it gets built. + +**`NEAR_THRESHOLD_DEFAULT` is 0.35 (2026-07-29), and it is a PUBLISHING change.** Bracketed +at 0.6/0.45/0.35/0.25 on identical inputs; every step down is a strict superset. Marginal +bands were read, not counted: of the 4,264 clusters admitted at 0.45, 96.2% have +byte-identical titles, 99.5% are cross-platform, 3 are same-channel. The recall knee is at +**0.45** — take that if a more conservative assertion is ever wanted. 0.35 was chosen because +its band is still clean and the report's job is to surface real mirrors to readers. It was +changed in its own commit because it changes what every built site asserts, NOT because it +saves sweep time (it saves 1.2 days of 80). + +**`updateDuplicateOverride` had ZERO CALLERS before this session.** It shipped complete — +the `confirmed` flag, the read-modify-write, the atomic rename, and a careful rule that an +empty patch clears a decision without deleting a confirmation — and nothing in the repo ever +invoked it. Because `clusterMaySharePartial` fails CLOSED, that meant ~166 `needsReview` +clusters could never be confirmed by anyone, and the review queue asked a question with no +way to answer. `shareClusterFromCanonical` was in the same state. Both are now wired from +/actionable. **Worth generalising: "the extension point exists" is not the same as "the +feature works", and a fails-closed gate with no way to open it is indistinguishable from a +missing feature.** + +**Total digest failures persisted NOTHING, and that was the review queue's data source +(2026-07-29).** Both total-failure paths in `digestVideo` skipped `writeDigestSection` for a +good reason — an empty section with current provenance reads as fresh and the video is never +retried — and the consequence was that the worst outputs left evidence only in a job log that +rotates. `DigestRecord.failures` is a SIBLING of `sections`, so `isSectionFresh` cannot see +it and the retry is preserved. It records `no-output` vs `all-rejected` separately, because +"the model proposed nothing" and "the model proposed 13 chapters and a guard clamped them +all" look identical from outside and need opposite fixes. + +**Writing the test for that found a second bug, of a shape this repo has now hit twice.** +`loadDigest` rebuilds the record field by field rather than spreading, so `failures` +round-tripped to nothing — the write succeeded and only the read omitted it. Structurally +identical to Phase 2's `pageHashes: [""]`, and caught the same way: by asserting on what came +BACK, not on what went in. Anything added to `DigestRecord` must be added to that reader. + **The viewer mode is `?vm=digest`, NOT `?vm=summary` (2026-07-29).** "Summary" was already taken twice — `DisplaySummary` is a video *listing card* (`common/lib/transcripts.ts`) and `/summaries/` is the browse-index page tree the search index serves. PLAN.md's own naming