commit 73fb48fdc8fde39fec0d8f65480fd7a10d4aaf5b
parent a3170a6cd448647a36480422882e91736ccfdb5a
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Mon, 27 Jul 2026 02:01:33 -0400
Record the 102-video digest validation run
First digest run at corpus scale rather than on a hand-picked bake-off
sample: community-notes (39, the pilot channel, every digest invalidated by
the PROMPT_VERSION bump) + friendofrc (63, never digested) = 102 videos,
42.2 audio-hours, 153 chunks, 622 chapters, 0 failures. Shipped defaults,
production build, driven through the actual "Digest channel" button.
The bake-off was right about quality and wrong about time.
QUALITY came in at or better than round 2 on every metric: zero-yield
chunks 0.0% (0 of 153) against a projected 11.8%, chapters/hour 14.74 vs
13.37, rejection rate 18.4% vs 19.1%, generic titles 27.0% vs 31.2%,
duplicate titles 0.3% vs 3.2%. A 2-video sample turned out to be
representative of quality.
THROUGHPUT was not: 90 s per audio-hour against 27, projecting ~81 sweep
days rather than 24.2. Two compounding causes, both of which matter for
planning a multi-week sweep: short videos cost ~2x per audio-hour (143 vs
75) because a one-chunk video cannot amortize per-call overhead, and round
2 measured only the `long` bucket; and the box was not idle — auto-download
and auto-transcribe ran throughout, contending for the same GPU. The lesson
is that a sample stratified by content generalizes but one stratified by
length does not generalize to throughput.
The only metric that degraded is the worst coverage gap (1:09:48 vs
24:13), and the new per-video panel explains it instead of leaving it a
mystery: that video recorded 13 out-of-range rejections, so the model DID
propose chapters for the missing hour and the clamp threw them all away.
"Never proposed" and "proposed and rejected" need different fixes and only
a recorded-warnings artifact separates them. Median gap is 5:19 — quote the
median with the max.
Rejections are now 93% a single guard (out-of-range 130, non-monotonic 9,
seam-duplicate 1), and language-drift went to zero across 42 audio-hours
(round 2 saw 7 in 7). The range clamp is doing effectively all the work.
Two plumbing confirmations: the counters agreed at 39/39 on the real corpus
where the old bucket read "All digested", and snapshot regeneration with
the added freshness target costs 0.1 s. One new (cosmetic) defect recorded:
a progress bar cannot move during a REGENERATION, because the target is
built from digestCount — digests that exist — and a regeneration rewrites
in place, so pct sits at 0 for the whole job.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Diffstat:
| M | plans/FACTS.md | | | 101 | +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ |
| M | plans/STATE.md | | | 69 | +++++++++++++++++++++++++++++++++++++++++++++++++++++++++------------ |
2 files changed, 158 insertions(+), 12 deletions(-)
diff --git a/plans/FACTS.md b/plans/FACTS.md
@@ -505,6 +505,107 @@ stays correct either way — it just re-generates mirrors instead of sharing the
---
+## Digest validation run — 102 videos, real corpus (measured 2026-07-27)
+
+The first digest run at corpus scale rather than on a hand-picked bake-off
+sample. Shipped defaults (`qwen2.5:7b` @ 8192, chunk-local, 600 cues,
+`PROMPT_VERSION` 2), production build, driven through the channel page's
+**Digest channel** button. `community-notes` (39 videos, 33.5 audio-h — the
+pilot channel, every digest invalidated by the version bump) plus `friendofrc`
+(63 videos, 8.7 audio-h, never digested). Score with:
+
+```
+cd common && pnpm exec tsx bin/digest-validate.ts community-notes friendofrc
+```
+
+**102 digests, 42.2 audio-hours, 153 chunks, 622 chapters, 0 failures.**
+
+### Quality: better than projected on everything except coverage gap
+
+| Metric | round 2 projection | **validation (102 videos)** | |
+| --- | --- | --- | --- |
+| Zero-yield chunks | 11.8% (2 of 17) | **0.0% (0 of 153)** | far better |
+| Chapters/hour | 13.37 | **14.74** | better |
+| Rejection rate | 19.1% | **18.4%** | matches |
+| Generic titles | 31.2% | **27.0%** | better |
+| Duplicate titles | 3.2% | **0.3%** | far better |
+| Worst coverage gap | 00:24:13 | **01:09:48** | **2.9× worse** |
+
+**Zero-yield went to actually zero.** Not one of 153 chunks came back empty.
+Round 2's 11.8% came from 2 chunks in a 17-chunk sample; at 9× the chunks the
+rate is 0. The chunk-local + 8192 pairing does not produce dead chunks.
+
+**Rejection rate landed within 0.7 points of a 2-video projection**, which is
+the strongest evidence available that the bake-off sample was representative on
+quality.
+
+**Rejections are now almost entirely one guard:**
+
+| Guard | Count | round 2 |
+| --- | --- | --- |
+| `out-of-range` | **130** (93%) | 13 |
+| `non-monotonic` | 9 | 2 |
+| `seam-duplicate` | 1 | 0 |
+| `language-drift` | **0** | 7 |
+
+Language drift disappeared across 42 audio-hours (round 2 saw 7 in 7). The range
+clamp is doing effectively all of the work, which means it is the guard whose
+removal would silently corrupt the corpus.
+
+**The worst coverage gap is the one metric that got worse at scale, and the
+per-video panel explains it.** `community-notes/v2fkbw7` (1:58:42, 20 chapters,
+gap 1:09:48) recorded **13 `out-of-range` rejections** — the model proposed
+chapters for that hour and every one was thrown away by the clamp. So the gap is
+not "the model ignored an hour of video", it is "the model's output for that hour
+was unusable". Those are different problems with different fixes, and the
+distinction is only visible because the artifact records rejections instead of
+dropping them. **The median gap is 00:05:19**, so this is a tail, not the norm —
+report the median alongside the max.
+
+### Throughput: 3.3× worse than projected, and video length is why
+
+| | audio-h | wall | s per audio-h |
+| --- | --- | --- | --- |
+| round 2 (long videos, idle box) | 6.96 | — | **27** |
+| `community-notes` (avg 51 min/video) | 33.5 | 42.1 min | **75** |
+| `friendofrc` (avg 8 min/video) | 8.7 | 20.8 min | **143** |
+| **combined** | **42.2** | **63.4 min** | **90** |
+
+Projected one-lane sweep over the corpus's 77,298 audio-hours: **~81 days, not
+the 24.2 round 2 projected.**
+
+Two compounding causes, and both matter for planning:
+
+1. **Short videos cost ~2× per audio-hour** (143 vs 75 s). A short video is one
+ chunk, so per-call overhead is amortized over minutes instead of hours. Round 2
+ measured only the `long` bucket, so its seconds-per-audio-hour is the corpus's
+ *best* case, applied to the whole corpus. The head of the corpus by count is
+ short videos.
+2. **The box was not idle.** `auto-download` and `auto-transcribe` ran throughout
+ (load average 10–14, whisper contending for the same GPU). This is
+ representative of production but not comparable to the bake-off's conditions.
+
+Even the long-video channel measured 75 s/audio-h against a 27 s projection, so
+contention alone accounts for roughly a 2.8× factor and video-length mix for the
+rest. **A sweep plan should assume ~80–90 days on a shared box, and should
+schedule against the transcription lanes rather than beside them.**
+
+### Two things the run confirmed about the plumbing
+
+- **The counters agree now.** Before the sweep, `noDigest` reported **39** and
+ `countMissingDigests` independently reported **39** for `community-notes` — the
+ version bump correctly invalidating every pilot digest. Under the old bucket the
+ Digest stage read "All digested" on the same corpus. Snapshot regeneration with
+ the added freshness target takes **0.1 s** for a 39-video channel, so the hot
+ path did not get meaningfully slower.
+- **Progress bars cannot move during a REGENERATION.** `setProgress` uses
+ `initial: stat.digestCount` and `target: initial + missing`, but `digestCount`
+ counts digests that *exist* — and a regeneration rewrites in place, so the count
+ never grows and `pct` sits at 0 for the whole job (observed:
+ `{initial:39,current:39,target:72,pct:0}`). Cosmetic, but it makes a long
+ re-sweep look wedged. Fixing it means counting digests *at the current identity*
+ rather than digests-on-disk.
+
## Editor surfaces
| Surface | File | Note |
diff --git a/plans/STATE.md b/plans/STATE.md
@@ -3,9 +3,13 @@
The working memory for the local-AI derived-corpus work. Rewritten at the end of every
session, before context is cleared. See [`README.md`](README.md) for the protocol.
-**Last updated:** 2026-07-26 — Stage B1 done (chunk-local + context-sized chunking measured
-and adopted, digest e2e green, pilot run); corpus-wide duplicate detection rewritten to
-stream, measured on both blocking strategies, and surfaced in search results.
+**Last updated:** 2026-07-27 — Stage B2: the digest layer is now INSPECTABLE and its defaults
+are the ones the bake-off measured best. Settings gained a real Digest section, each video
+page gained a review panel (provenance, warnings, `derivedFrom`, staleness), the two digest
+counters were made to agree, `chunk-local`/8192/600 shipped as defaults behind a
+`PROMPT_VERSION` bump, the pilot's "un-created job" was root-caused (a pre-hydration lost
+click) and fixed, and a **102-video validation run** was completed on the real corpus. Also:
+the 0.6 duplicate near-threshold was measured and found to be rejecting real mirrors in bulk.
---
@@ -14,7 +18,7 @@ stream, measured on both blocking strategies, and surfaced in search results.
| Phase | Status | Notes |
| --- | --- | --- |
| 0 · Benchmark transcription engines | not started | Still unmeasured. The DIGEST bake-off is done; this is the separate transcription one. |
-| 1 · Generation harness | **done** | Spine in `8c041fd`; correctness + bake-off + e2e in Stage B1. |
+| 1 · Generation harness | **done + validated** | Spine in `8c041fd`; correctness + bake-off + e2e in Stage B1; operator surfaces + measured-best defaults + 102-video validation run in Stage B2. |
| 1.5 · Channel context | not started | `contextHash` is plumbed and empty, so notes can be added without invalidating the corpus. |
| 2 · Digest corpus in build | not started | |
| 2.5 · Observability | not started | **Blocks backfill.** |
@@ -29,10 +33,26 @@ stream, measured on both blocking strategies, and surfaced in search results.
| 11a · Review queue | not started | **Land before backfilling.** `warnings[]` is the data it reads. |
| 11b · Viewer feedback | not started | |
-**Recommended next:** the Stage B2 surfaces that were deliberately deferred — settings form
-fields (including `timestampMode` / `promptVariant`), the per-video review panel, the
-`/actionable` cluster section, and the PipelineBand instrument — then the ~100-video
-validation run before any sweep.
+**Recommended next**, in the order the validation run argues for:
+
+1. **Phase 2.5 observability, and treat the throughput finding as its first requirement.**
+ The run measured **90 s per audio-hour**, projecting **~81 sweep days** against round 2's
+ 24.2 — because short videos cost ~2× per audio-hour and the box was sharing a GPU with
+ `auto-transcribe`. A sweep of that length needs a coverage instrument and a scheduling
+ story (run the digest lane *against* the transcription lanes, not beside them) before it
+ starts, not after. Numbers in FACTS.md.
+2. **Phase 11a review queue.** `warnings[]` is now proven as its data source: 93% of all
+ rejections are `out-of-range` from a single guard, and the per-video panel already showed
+ that the worst coverage gap in the run is a video where 13 proposed chapters were clamped
+ away rather than never proposed. That distinction is what a queue must surface.
+3. **Decide the duplicate near-threshold in its own commit.** 0.6 measurably rejects real
+ cross-platform mirrors; 0.35 recovers 4,456 clusters with no observed false positives and
+ lifts the shareable-digest ceiling 2.5×. Probe 0.25/0.45 to bracket the knee first.
+4. **Phase 1.5 channel context** — still a backfill gate (`contextHash` staleness), and still
+ not required for a validation run.
+
+Deferred from Stage B2 and still open: the `/actionable` duplicate-cluster confirm/reject
+buttons (PLAN.md Phase 11a, extension points recorded there) and the PipelineBand instrument.
---
@@ -105,6 +125,23 @@ Two counters were previously lying:
A digest shared from a duplicate cluster's canonical member counts as done in the bucket —
the canonical member's own freshness drives regeneration and the share is re-applied from it.
+**The 102-video validation run says the bake-off was right about quality and wrong about
+time (2026-07-27).** Full table in FACTS.md. Quality came in at or better than round 2's
+projection on every metric — **zero-yield went to literally 0 of 153 chunks**, chapters/hour
+14.74 vs 13.37, rejection rate 18.4% vs 19.1% predicted, generic titles 27.0% vs 31.2% — so a
+2-video sample turned out to be representative of quality. Throughput was not: **90 s per
+audio-hour vs 27**, projecting **~81 sweep days rather than 24.2**, because round 2 measured
+only long videos on an idle box and the corpus is mostly short videos on a box also running
+whisper. The lesson to carry: **a bake-off sample stratified by content generalizes; one
+stratified by length does not generalize to throughput.**
+
+The one metric that degraded is the worst coverage gap (1:09:48 vs 24:13 projected), and the
+new per-video panel explains it rather than leaving it a mystery: the worst video recorded 13
+`out-of-range` rejections, so the model *did* propose chapters for the missing hour and the
+clamp threw all of them away. "Never proposed" and "proposed and rejected" need different
+fixes, and only a recorded-warnings artifact can tell them apart. Median gap is 5:19 — quote
+the median with the max.
+
**The sweep is ~25 days, not ~64.** Round 1 measured 59 days on short+medium; Round 2
measured ~25 on the long tail, because long videos amortise the fixed per-call overhead and
the corpus is dominated by them. Throughput is no longer the binding constraint it was
@@ -213,10 +250,18 @@ playwright `webServer` with `OLLAMA_URL` pointed at it in **both** `dev:test` an
runtime strategies off `manifest.json` / `page-NNNN.json` paths (`:55`, `:147`, `:205`), so
a root-level JSON is uncached. Affects the search-result duplicate badge in PWA mode only —
it silently does not appear offline, which is the correct failure but an unstated one.
-- **The pilot's un-created job.** Clicking "Digest channel" a second time produced no job
- record and an empty log, with no dedupe guard in `runManagedFunction` to explain it.
- Verified via `runDigestBatch` instead. **Still unexplained** — reproduce before trusting
- the button to drive a long sweep.
+- **ANSWERED (2026-07-27): the pilot's un-created job was a pre-hydration lost click.**
+ Reproduced against the real corpus on a production build. The click fires **no POST at
+ all**: the server-rendered button has no handler until React hydrates, so a click before
+ that is silently discarded — no request, no job, no log, no error. Instrumented, the first
+ click produced 0 POSTs and a click five seconds later started the job normally. There is no
+ dedupe guard in `runManagedFunction` to explain it because there was nothing to dedupe —
+ the action was never invoked. The window is worst exactly where the pilot hit it: a heavy
+ channel page under a dev server (30 s – 9.9 min to render). **Fixed** — `StreamActionLog`
+ disables its button until mounted, which also closes the same latent race for sync,
+ download and transcribe, and makes e2e clicks wait for enablement instead of losing them
+ (the same failure mode that made the charts specs flaky). Verified after the fix: one
+ click, one POST, job runs. The button is now trusted to drive a long sweep.
- **Diarizer choice** — unchanged, decide at Phase 9 on real audio.
- **Phase 11b collector** — **answered: no, not `r2-proxy/`.** `src/index.ts:35-41` rejects
every non-GET/HEAD method and `KEY_RE` at `:27` carries a comment stating the proxy must