Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 73fb48fdc8fde39fec0d8f65480fd7a10d4aaf5b
parent a3170a6cd448647a36480422882e91736ccfdb5a
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon, 27 Jul 2026 02:01:33 -0400

Record the 102-video digest validation run

First digest run at corpus scale rather than on a hand-picked bake-off
sample: community-notes (39, the pilot channel, every digest invalidated by
the PROMPT_VERSION bump) + friendofrc (63, never digested) = 102 videos,
42.2 audio-hours, 153 chunks, 622 chapters, 0 failures. Shipped defaults,
production build, driven through the actual "Digest channel" button.

The bake-off was right about quality and wrong about time.

QUALITY came in at or better than round 2 on every metric: zero-yield
chunks 0.0% (0 of 153) against a projected 11.8%, chapters/hour 14.74 vs
13.37, rejection rate 18.4% vs 19.1%, generic titles 27.0% vs 31.2%,
duplicate titles 0.3% vs 3.2%. A 2-video sample turned out to be
representative of quality.

THROUGHPUT was not: 90 s per audio-hour against 27, projecting ~81 sweep
days rather than 24.2. Two compounding causes, both of which matter for
planning a multi-week sweep: short videos cost ~2x per audio-hour (143 vs
75) because a one-chunk video cannot amortize per-call overhead, and round
2 measured only the `long` bucket; and the box was not idle — auto-download
and auto-transcribe ran throughout, contending for the same GPU. The lesson
is that a sample stratified by content generalizes but one stratified by
length does not generalize to throughput.

The only metric that degraded is the worst coverage gap (1:09:48 vs
24:13), and the new per-video panel explains it instead of leaving it a
mystery: that video recorded 13 out-of-range rejections, so the model DID
propose chapters for the missing hour and the clamp threw them all away.
"Never proposed" and "proposed and rejected" need different fixes and only
a recorded-warnings artifact separates them. Median gap is 5:19 — quote the
median with the max.

Rejections are now 93% a single guard (out-of-range 130, non-monotonic 9,
seam-duplicate 1), and language-drift went to zero across 42 audio-hours
(round 2 saw 7 in 7). The range clamp is doing effectively all the work.

Two plumbing confirmations: the counters agreed at 39/39 on the real corpus
where the old bucket read "All digested", and snapshot regeneration with
the added freshness target costs 0.1 s. One new (cosmetic) defect recorded:
a progress bar cannot move during a REGENERATION, because the target is
built from digestCount — digests that exist — and a regeneration rewrites
in place, so pct sits at 0 for the whole job.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Diffstat:
Mplans/FACTS.md | 101+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mplans/STATE.md | 69+++++++++++++++++++++++++++++++++++++++++++++++++++++++++------------
2 files changed, 158 insertions(+), 12 deletions(-)

diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -505,6 +505,107 @@ stays correct either way — it just re-generates mirrors instead of sharing the --- +## Digest validation run — 102 videos, real corpus (measured 2026-07-27) + +The first digest run at corpus scale rather than on a hand-picked bake-off +sample. Shipped defaults (`qwen2.5:7b` @ 8192, chunk-local, 600 cues, +`PROMPT_VERSION` 2), production build, driven through the channel page's +**Digest channel** button. `community-notes` (39 videos, 33.5 audio-h — the +pilot channel, every digest invalidated by the version bump) plus `friendofrc` +(63 videos, 8.7 audio-h, never digested). Score with: + +``` +cd common && pnpm exec tsx bin/digest-validate.ts community-notes friendofrc +``` + +**102 digests, 42.2 audio-hours, 153 chunks, 622 chapters, 0 failures.** + +### Quality: better than projected on everything except coverage gap + +| Metric | round 2 projection | **validation (102 videos)** | | +| --- | --- | --- | --- | +| Zero-yield chunks | 11.8% (2 of 17) | **0.0% (0 of 153)** | far better | +| Chapters/hour | 13.37 | **14.74** | better | +| Rejection rate | 19.1% | **18.4%** | matches | +| Generic titles | 31.2% | **27.0%** | better | +| Duplicate titles | 3.2% | **0.3%** | far better | +| Worst coverage gap | 00:24:13 | **01:09:48** | **2.9× worse** | + +**Zero-yield went to actually zero.** Not one of 153 chunks came back empty. +Round 2's 11.8% came from 2 chunks in a 17-chunk sample; at 9× the chunks the +rate is 0. The chunk-local + 8192 pairing does not produce dead chunks. + +**Rejection rate landed within 0.7 points of a 2-video projection**, which is +the strongest evidence available that the bake-off sample was representative on +quality. + +**Rejections are now almost entirely one guard:** + +| Guard | Count | round 2 | +| --- | --- | --- | +| `out-of-range` | **130** (93%) | 13 | +| `non-monotonic` | 9 | 2 | +| `seam-duplicate` | 1 | 0 | +| `language-drift` | **0** | 7 | + +Language drift disappeared across 42 audio-hours (round 2 saw 7 in 7). The range +clamp is doing effectively all of the work, which means it is the guard whose +removal would silently corrupt the corpus. + +**The worst coverage gap is the one metric that got worse at scale, and the +per-video panel explains it.** `community-notes/v2fkbw7` (1:58:42, 20 chapters, +gap 1:09:48) recorded **13 `out-of-range` rejections** — the model proposed +chapters for that hour and every one was thrown away by the clamp. So the gap is +not "the model ignored an hour of video", it is "the model's output for that hour +was unusable". Those are different problems with different fixes, and the +distinction is only visible because the artifact records rejections instead of +dropping them. **The median gap is 00:05:19**, so this is a tail, not the norm — +report the median alongside the max. + +### Throughput: 3.3× worse than projected, and video length is why + +| | audio-h | wall | s per audio-h | +| --- | --- | --- | --- | +| round 2 (long videos, idle box) | 6.96 | — | **27** | +| `community-notes` (avg 51 min/video) | 33.5 | 42.1 min | **75** | +| `friendofrc` (avg 8 min/video) | 8.7 | 20.8 min | **143** | +| **combined** | **42.2** | **63.4 min** | **90** | + +Projected one-lane sweep over the corpus's 77,298 audio-hours: **~81 days, not +the 24.2 round 2 projected.** + +Two compounding causes, and both matter for planning: + +1. **Short videos cost ~2× per audio-hour** (143 vs 75 s). A short video is one + chunk, so per-call overhead is amortized over minutes instead of hours. Round 2 + measured only the `long` bucket, so its seconds-per-audio-hour is the corpus's + *best* case, applied to the whole corpus. The head of the corpus by count is + short videos. +2. **The box was not idle.** `auto-download` and `auto-transcribe` ran throughout + (load average 10–14, whisper contending for the same GPU). This is + representative of production but not comparable to the bake-off's conditions. + +Even the long-video channel measured 75 s/audio-h against a 27 s projection, so +contention alone accounts for roughly a 2.8× factor and video-length mix for the +rest. **A sweep plan should assume ~80–90 days on a shared box, and should +schedule against the transcription lanes rather than beside them.** + +### Two things the run confirmed about the plumbing + +- **The counters agree now.** Before the sweep, `noDigest` reported **39** and + `countMissingDigests` independently reported **39** for `community-notes` — the + version bump correctly invalidating every pilot digest. Under the old bucket the + Digest stage read "All digested" on the same corpus. Snapshot regeneration with + the added freshness target takes **0.1 s** for a 39-video channel, so the hot + path did not get meaningfully slower. +- **Progress bars cannot move during a REGENERATION.** `setProgress` uses + `initial: stat.digestCount` and `target: initial + missing`, but `digestCount` + counts digests that *exist* — and a regeneration rewrites in place, so the count + never grows and `pct` sits at 0 for the whole job (observed: + `{initial:39,current:39,target:72,pct:0}`). Cosmetic, but it makes a long + re-sweep look wedged. Fixing it means counting digests *at the current identity* + rather than digests-on-disk. + ## Editor surfaces | Surface | File | Note | diff --git a/plans/STATE.md b/plans/STATE.md @@ -3,9 +3,13 @@ The working memory for the local-AI derived-corpus work. Rewritten at the end of every session, before context is cleared. See [`README.md`](README.md) for the protocol. -**Last updated:** 2026-07-26 — Stage B1 done (chunk-local + context-sized chunking measured -and adopted, digest e2e green, pilot run); corpus-wide duplicate detection rewritten to -stream, measured on both blocking strategies, and surfaced in search results. +**Last updated:** 2026-07-27 — Stage B2: the digest layer is now INSPECTABLE and its defaults +are the ones the bake-off measured best. Settings gained a real Digest section, each video +page gained a review panel (provenance, warnings, `derivedFrom`, staleness), the two digest +counters were made to agree, `chunk-local`/8192/600 shipped as defaults behind a +`PROMPT_VERSION` bump, the pilot's "un-created job" was root-caused (a pre-hydration lost +click) and fixed, and a **102-video validation run** was completed on the real corpus. Also: +the 0.6 duplicate near-threshold was measured and found to be rejecting real mirrors in bulk. --- @@ -14,7 +18,7 @@ stream, measured on both blocking strategies, and surfaced in search results. | Phase | Status | Notes | | --- | --- | --- | | 0 · Benchmark transcription engines | not started | Still unmeasured. The DIGEST bake-off is done; this is the separate transcription one. | -| 1 · Generation harness | **done** | Spine in `8c041fd`; correctness + bake-off + e2e in Stage B1. | +| 1 · Generation harness | **done + validated** | Spine in `8c041fd`; correctness + bake-off + e2e in Stage B1; operator surfaces + measured-best defaults + 102-video validation run in Stage B2. | | 1.5 · Channel context | not started | `contextHash` is plumbed and empty, so notes can be added without invalidating the corpus. | | 2 · Digest corpus in build | not started | | | 2.5 · Observability | not started | **Blocks backfill.** | @@ -29,10 +33,26 @@ stream, measured on both blocking strategies, and surfaced in search results. | 11a · Review queue | not started | **Land before backfilling.** `warnings[]` is the data it reads. | | 11b · Viewer feedback | not started | | -**Recommended next:** the Stage B2 surfaces that were deliberately deferred — settings form -fields (including `timestampMode` / `promptVariant`), the per-video review panel, the -`/actionable` cluster section, and the PipelineBand instrument — then the ~100-video -validation run before any sweep. +**Recommended next**, in the order the validation run argues for: + +1. **Phase 2.5 observability, and treat the throughput finding as its first requirement.** + The run measured **90 s per audio-hour**, projecting **~81 sweep days** against round 2's + 24.2 — because short videos cost ~2× per audio-hour and the box was sharing a GPU with + `auto-transcribe`. A sweep of that length needs a coverage instrument and a scheduling + story (run the digest lane *against* the transcription lanes, not beside them) before it + starts, not after. Numbers in FACTS.md. +2. **Phase 11a review queue.** `warnings[]` is now proven as its data source: 93% of all + rejections are `out-of-range` from a single guard, and the per-video panel already showed + that the worst coverage gap in the run is a video where 13 proposed chapters were clamped + away rather than never proposed. That distinction is what a queue must surface. +3. **Decide the duplicate near-threshold in its own commit.** 0.6 measurably rejects real + cross-platform mirrors; 0.35 recovers 4,456 clusters with no observed false positives and + lifts the shareable-digest ceiling 2.5×. Probe 0.25/0.45 to bracket the knee first. +4. **Phase 1.5 channel context** — still a backfill gate (`contextHash` staleness), and still + not required for a validation run. + +Deferred from Stage B2 and still open: the `/actionable` duplicate-cluster confirm/reject +buttons (PLAN.md Phase 11a, extension points recorded there) and the PipelineBand instrument. --- @@ -105,6 +125,23 @@ Two counters were previously lying: A digest shared from a duplicate cluster's canonical member counts as done in the bucket — the canonical member's own freshness drives regeneration and the share is re-applied from it. +**The 102-video validation run says the bake-off was right about quality and wrong about +time (2026-07-27).** Full table in FACTS.md. Quality came in at or better than round 2's +projection on every metric — **zero-yield went to literally 0 of 153 chunks**, chapters/hour +14.74 vs 13.37, rejection rate 18.4% vs 19.1% predicted, generic titles 27.0% vs 31.2% — so a +2-video sample turned out to be representative of quality. Throughput was not: **90 s per +audio-hour vs 27**, projecting **~81 sweep days rather than 24.2**, because round 2 measured +only long videos on an idle box and the corpus is mostly short videos on a box also running +whisper. The lesson to carry: **a bake-off sample stratified by content generalizes; one +stratified by length does not generalize to throughput.** + +The one metric that degraded is the worst coverage gap (1:09:48 vs 24:13 projected), and the +new per-video panel explains it rather than leaving it a mystery: the worst video recorded 13 +`out-of-range` rejections, so the model *did* propose chapters for the missing hour and the +clamp threw all of them away. "Never proposed" and "proposed and rejected" need different +fixes, and only a recorded-warnings artifact can tell them apart. Median gap is 5:19 — quote +the median with the max. + **The sweep is ~25 days, not ~64.** Round 1 measured 59 days on short+medium; Round 2 measured ~25 on the long tail, because long videos amortise the fixed per-call overhead and the corpus is dominated by them. Throughput is no longer the binding constraint it was @@ -213,10 +250,18 @@ playwright `webServer` with `OLLAMA_URL` pointed at it in **both** `dev:test` an runtime strategies off `manifest.json` / `page-NNNN.json` paths (`:55`, `:147`, `:205`), so a root-level JSON is uncached. Affects the search-result duplicate badge in PWA mode only — it silently does not appear offline, which is the correct failure but an unstated one. -- **The pilot's un-created job.** Clicking "Digest channel" a second time produced no job - record and an empty log, with no dedupe guard in `runManagedFunction` to explain it. - Verified via `runDigestBatch` instead. **Still unexplained** — reproduce before trusting - the button to drive a long sweep. +- **ANSWERED (2026-07-27): the pilot's un-created job was a pre-hydration lost click.** + Reproduced against the real corpus on a production build. The click fires **no POST at + all**: the server-rendered button has no handler until React hydrates, so a click before + that is silently discarded — no request, no job, no log, no error. Instrumented, the first + click produced 0 POSTs and a click five seconds later started the job normally. There is no + dedupe guard in `runManagedFunction` to explain it because there was nothing to dedupe — + the action was never invoked. The window is worst exactly where the pilot hit it: a heavy + channel page under a dev server (30 s – 9.9 min to render). **Fixed** — `StreamActionLog` + disables its button until mounted, which also closes the same latent race for sync, + download and transcribe, and makes e2e clicks wait for enablement instead of losing them + (the same failure mode that made the charts specs flaky). Verified after the fix: one + click, one POST, job runs. The button is now trusted to drive a long sweep. - **Diarizer choice** — unchanged, decide at Phase 9 on real audio. - **Phase 11b collector** — **answered: no, not `r2-proxy/`.** `src/index.ts:35-41` rejects every non-GET/HEAD method and `KEY_RE` at `:27` carries a comment stating the proxy must