Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 41c2afc14d4b7433569113b05307fd524abf9710
parent a7aa7b1e9a629bad6382fedaa50dc5e72955c7f7
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Wed, 29 Jul 2026 20:27:26 -0400

Retract the ~81 days and the "fifty times" claim, on the run's own data

STATE.md's GPU-arbitration paragraph was mine and it was wrong. It claimed
contention was "worth ~56 days" and "fifty times every duplicate lever"
from 90 s/audio-hour against 27 projected. That 2.8x compares the
validation run's WALL clock against round 2's ENGINE time across samples
with different length mixes — not one quantity, two.

Re-priced from the validation run's own artifacts (the 153 chunks its 102
sidecars record in provenance.chunks, not from its report), it says 54.6
days, and the 3.33x factors as 1.49x sample mix x 2.20x per-chunk = 3.27x.
MEASURED_SECONDS_PER_CHUNK is now 24.7 from that recomputation.

FACTS.md's "short videos cost ~2x per audio-hour" also inverts: per CHUNK
the short-video channel is 29% CHEAPER (19.81 vs 28.07 s), because a short
video's single chunk is a partial one. Short videos are expensive per hour
of audio, which is not a thing the GPU charges for. Retracted as a unit
artifact; only contention survives as a real cause, and its size is now
labelled a hypothesis with a named test rather than a number.

Recorded: the full chunk census with its band table, the three-way
re-pricing, tags as the FOURTH built-but-never-run case (now named as a
defect class a code review cannot see), the tags surcharge measurement,
and two notes for any future bake-off — BUCKET_BOUNDS has THREE holes
where bucketFor returns null (30-45min, 90min-3h, 4.5-6h), not the one I
first spotted, and sample.json records no strata.

Nothing here is a gate; all of it was a wrong number in a doc someone
would have planned against.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Diffstat:
Mcommon/controller/digestPlan.ts | 10++++++++--
Mplans/FACTS.md | 231+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++----------
Mplans/STATE.md | 104++++++++++++++++++++++++++++++++++++++++++++++++++++++++++----------------------
3 files changed, 287 insertions(+), 58 deletions(-)

diff --git a/common/controller/digestPlan.ts b/common/controller/digestPlan.ts @@ -54,9 +54,15 @@ import { // which the audio-hour model does for none (× 191,116 corpus chunks): // // 11.2 s idle box, ENGINE time (bake-off round 2) → 24.8 days [24.2 ✓] -// 24.9 s contended box, WALL time (102-video run) → 55.1 days [81 ✗] +// 24.7 s contended box, WALL time (102-video run) → 54.6 days [81 ✗] // 60.6 s idle box, engine, gemma2 (bake-off round 2) → 134 days [132 ✓] // +// The 24.7 is recomputed from the validation run's OWN artifacts, not from its +// report: 62.9 min of wall over the 153 chunks its 102 sidecars actually record in +// `provenance.chunks`. Its 3.33× "contention penalty" then factors as 1.49× sample +// length mix (3.63 chunks/audio-hour against round 2's 2.44) × 2.20× per-chunk +// cost = 3.27×. +// // The old 27 → 90 s/audio-hour "contention penalty" factors exactly: 1.49× from // the validation sample's length mix (3.63 vs 2.44 chunks/audio-hour) × 2.22× // per-chunk cost = 3.30× against the observed 3.33×. So the honest range is ~25 @@ -67,7 +73,7 @@ import { // skips. It is pessimistic on an idle box by design. The 2.22× gap has three // parts (idle engine, contended engine, and non-engine wall inflation from the // yield gate's 3 s poll) which only a production run can separate. -export const MEASURED_SECONDS_PER_CHUNK = 24.9; +export const MEASURED_SECONDS_PER_CHUNK = 24.7; // Corpus-wide chunk density, from the free census over `statsByPath.cueCount` — // all 73,367 transcribed videos, none missing a cue count, at the shipped maxCues diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -534,11 +534,65 @@ refused by the alignment gate regardless. **Adopted: 0.35.** The recall knee is at **0.45** — the value to take if a more conservative assertion is ever wanted. -### The digest sweep, priced in AUDIO-HOURS (measured 2026-07-29) +### The digest sweep, priced in CHUNKS (measured 2026-07-29) -`common/bin/digest-plan.ts --no-freshness`, at the measured 90 s/audio-hour. Total digestable -corpus: **77,298 audio-hours** over 73,367 videos (2,987 have no transcript), which reproduces -the roadmap's ~81-day headline independently. +**A chunk — one model call — is the unit of cost. Seconds-per-audio-hour is NOT a +stable unit here and must not be used for any projection.** Chunk density varies +**fourfold** across the corpus, so an s/audio-hour figure measured on one length +mix does not transfer to another. Three separate cost measurements looked +irreconcilable (27 / 90 / 151 s/audio-hour) purely because of this; re-priced per +chunk they agree to within ~2%. + +#### The chunk census + +Computed free from `statsByPath.cueCount` — present on **all 73,367** transcribed +videos, so it needed no GPU time and no transcript reads. At the shipped config +(`maxCues` 600, overlap 40, step 560): + +``` +chunks(c) = c === 0 ? 0 : c <= 600 ? 1 : 1 + ceil((c - 600) / 560) +``` + +| band | videos | audio-h | % of audio | chunks | **chunks/audio-h** | % of chunks | +| --- | --- | --- | --- | --- | --- | --- | +| < 15 min | 41,960 | 6,203 | 8.0% | 41,960 | **6.76** | 22.0% | +| 15–60 min | 14,036 | 6,631 | 8.6% | 22,446 | 3.39 | 11.7% | +| 1–2 h | 5,847 | 8,480 | 11.0% | 22,353 | 2.64 | 11.7% | +| 2–4 h | 5,486 | 15,832 | 20.5% | 33,106 | 2.09 | 17.3% | +| 4–8 h | 4,504 | 25,304 | 32.7% | 46,363 | 1.83 | 24.3% | +| > 8 h | 1,534 | 14,850 | 19.2% | 24,888 | **1.68** | 13.0% | +| **total** | **73,367** | **77,298** | | **191,116** | **2.47** | | + +**4 h+ videos are 52% of the AUDIO but only 37% of the WORK; sub-hour videos are +16% of audio and 34% of chunks.** Any reasoning that prices this sweep in +audio-hours gets that backwards. Pinned as a regression test in +`controller/digestPlan.test.ts` (it re-derives 191,116 from LMDB and fails if the +plan and the chunker ever drift). + +#### Re-pricing every measurement taken so far + +| per-chunk cost | source | × 191,116 | previously claimed | +| --- | --- | --- | --- | +| 11.2 s | round 2, idle box, **engine** time | **24.8 d** | 24.2 ✓ | +| 24.7 s | 102-video run, contended, **wall** time | **54.6 d** | 81 ✗ | +| 60.6 s | round 2 gemma2, idle, engine | **134 d** | 132 ✓ | + +**The honest range is ~25 days idle to ~55 contended — not 81.** The audio-hour +model reproduces none of the three; the chunk model reproduces all three. + +#### The plan output itself + +`common/bin/digest-plan.ts --no-freshness`. Total digestable corpus: **77,298 +audio-hours over 73,367 videos** (2,987 have no transcript) = **191,116 chunks**, +of which **185,483 chunks / 75,638 audio-hours** remain to generate at `--near` +0.35 → **53.5 sweep days**. Sharing avoids 5,633 chunks (1.6 days). + +Per-channel density is now reported, because it is what explains a changing rate +mid-sweep — the queue runs heaviest-first and the heaviest channels are the +CHEAPEST per audio-hour: `HasanAbiVODs3` 1.71 c/h against `chibi-reviews` 6.38. + +The audio-hour table below is kept as measured, but its `days` column was computed +at 90 s/audio-hour and is **retracted**. | `--near` | canonical | mirror-aligned (free) | mirror-unaligned | unclustered | **to generate** | **days** | | --- | --- | --- | --- | --- | --- | --- | @@ -653,33 +707,91 @@ distinction is only visible because the artifact records rejections instead of dropping them. **The median gap is 00:05:19**, so this is a tail, not the norm — report the median alongside the max. -### Throughput: 3.3× worse than projected, and video length is why - -| | audio-h | wall | s per audio-h | -| --- | --- | --- | --- | -| round 2 (long videos, idle box) | 6.96 | — | **27** | -| `community-notes` (avg 51 min/video) | 33.5 | 42.1 min | **75** | -| `friendofrc` (avg 8 min/video) | 8.7 | 20.8 min | **143** | -| **combined** | **42.2** | **63.4 min** | **90** | - -Projected one-lane sweep over the corpus's 77,298 audio-hours: **~81 days, not -the 24.2 round 2 projected.** +### Throughput: 3.3× worse than projected — **and ~a third of that was the unit** -Two compounding causes, and both matter for planning: +**CORRECTED 2026-07-29.** The `~81 days` this section claimed is wrong on this +run's own data. Re-priced against the 153 chunks the run's 102 sidecars actually +record in `provenance.chunks`, it says **54.6 days**: -1. **Short videos cost ~2× per audio-hour** (143 vs 75 s). A short video is one - chunk, so per-call overhead is amortized over minutes instead of hours. Round 2 - measured only the `long` bucket, so its seconds-per-audio-hour is the corpus's - *best* case, applied to the whole corpus. The head of the corpus by count is - short videos. +| | audio-h | wall | chunks | chunks/audio-h | s per audio-h | **s per chunk** | +| --- | --- | --- | --- | --- | --- | --- | +| round 2 (long videos, idle box) | 6.96 | — | — | 2.44 | 27 | **11.2** | +| `community-notes` (avg 51 min/video) | 33.5 | 42.1 min | 90 | 2.69 | 75 | **28.07** | +| `friendofrc` (avg 8 min/video) | 8.7 | 20.8 min | 63 | 7.24 | 143 | **19.81** | +| **combined** | **42.2** | **62.9 min** | **153** | **3.63** | **89.4** | **24.67** | + +The 3.33× factors almost exactly: **1.49× sample length mix (3.63 vs 2.44 +chunks/audio-h) × 2.20× per-chunk cost = 3.27×**, against 3.33× observed. + +**The first cause as previously written was an artifact of the unit, and inverts.** +It said "short videos cost ~2× per audio-hour (143 vs 75 s)". True, and irrelevant: +**per chunk the short-video channel is 29% CHEAPER** (19.81 vs 28.07 s), because a +short video's single chunk is a *partial* chunk — fewer cues to prefill and fewer +chapters to decode. Short videos are not expensive work; they are expensive *per +hour of audio*, which is not a thing the GPU charges for. + +So only ONE compounding cause survives, plus a third of the gap being sampling: + +1. ~~Short videos cost ~2× per audio-hour~~ — **retracted as a unit artifact.** + What is true: round 2 sampled only the `long` bucket, whose 2.44 chunks/audio-h + is well below the corpus's 2.47… and *far* below the validation sample's 3.63. + The sample mix, not the videos, contributed 1.49× of the 3.33×. 2. **The box was not idle.** `auto-download` and `auto-transcribe` ran throughout - (load average 10–14, whisper contending for the same GPU). This is - representative of production but not comparable to the bake-off's conditions. - -Even the long-video channel measured 75 s/audio-h against a 27 s projection, so -contention alone accounts for roughly a 2.8× factor and video-length mix for the -rest. **A sweep plan should assume ~80–90 days on a shared box, and should -schedule against the transcription lanes rather than beside them.** + (load average 10–14, the transcription engine contending for the same GPU). This + is representative of production but not comparable to the bake-off's conditions. + This remains the one real cause — but its size is a **hypothesis, not a + measurement**, because the 2.20× per-chunk residual compares this run's **wall** + clock against round 2's **engine** time. Those are different quantities, and the + residual therefore also contains model-load, prefill, and the yield gate's own + 3 s-poll idling. `digestApps.ts` now records load/prefill/decode per call + precisely so the next measurement can separate them. + +**Do not plan against ~80–90 days.** Plan against **~55 days on a shared box and +~25 idle**, and treat closing that gap as the open question. Two things follow that +were not visible before: the yield shipped with a CPU-worker bug (it stopped the +digest lane for `device: "cpu"` transcription competing for zero shaders — fixed, +`digest.yieldToCpuWorkers`), and `gemma2:9b` **offloads 1,029 MB of its 6.43 GB +footprint to CPU** at 8192 ctx on this 8 GB card while `qwen2.5:7b` (5.12 GB) is +fully resident — so gemma2's 134 re-priced days are partly a property of the card, +not the model. Verified via `/api/ps` (`size_vram` < `size`). + +### Built, specified, and never run once — a recurring defect class + +Four cases found on 2026-07-29, all the same shape: the code exists, reads +correctly, has a settings surface or a caller signature, and **had never executed**. +Worth naming as a class, because a code review cannot see it and every test passed. + +| Case | Evidence it never ran | +| --- | --- | +| `updateDuplicateOverride` | zero callers | +| `shareClusterFromCanonical` | zero callers | +| `setProgress` (one path) | referenced nowhere | +| **Digest TAGS** | **0 of 109 sidecars had a tags section** | + +The tags case was the largest: `parseTags`, a tags JSON schema, a `sections` branch +in `digestVideo` and a per-section checkbox in `SettingsForm.tsx` had all shipped, +`digest-validate.ts` had no tag metrics, `editor/e2e/digest.spec.ts` had no tags +coverage, and neither bake-off round scored them. It works — 11 videos, 0 warnings, +`parseTags` accepted every live output — and it is now generated, scored and +covered by e2e. See the tags section below. + +The pattern to take from it: **a capability with no metric and no e2e is +indistinguishable from a broken one.** The fix is not more review; it is that +anything with a settings surface gets one live run and one assertion. + +### Notes for any future bake-off round (not gates now) + +- **`BUCKET_BOUNDS` has holes** (`common/bin/digest-bakeoff.ts:70-75`). + `bucketFor()` returns `null` for any duration in a gap, so those videos are + silently unsamplable: **30–45 min**, **90 min – 3 h**, and **4.5 – 6 h** (`long` + caps at 4.5 h, `verylong` starts at 6 h). The gaps look deliberate — they keep + the strata well separated — but nothing says so and nothing reports how many + videos fall in them. The 4.5–6 h hole matters most, since that band is dense in + this corpus. +- **`sample.json` does not record the strata that produced it.** So a round cannot + prove what mix it measured, which is exactly the confound that made round 2's + 27 s/audio-hour non-transferable. Any future round should record + chunks/audio-hour for its sample beside the per-bucket counts. ### Two things the run confirmed about the plumbing @@ -878,3 +990,66 @@ clones the pattern repo. Note: Arch names the binary **`fabric-ai`**, not `fabric`, to avoid colliding with the unrelated Python `fabric` deployment tool. Moot while fabric is deferred, but relevant if it is ever added as an engine. + +--- + +## Digest TAGS — first generation ever (measured 2026-07-29) + +Tags had never been generated: **0 of 109 sidecars** contained a tags section. See +"Built, specified, and never run once" above. Generated via the real pipeline +(`runDigestBatch` with `sections: ["chapters","tags"]`) on `teamrcn` (7 videos), +`steven-crowder` (2) and `nuxanor-kick` (2). + +### It works, and the output is good + +**11/11 videos, 0 warnings** — `parseTags` accepted every live output, having +never accepted or rejected anything from a real model before. Scored by the new +tag metrics in `bin/digest-validate.ts`: + +| metric | teamrcn (7 tiny videos) | steven-crowder + nuxanor-kick (4 multi-hour) | +| --- | --- | --- | +| tags per video | 4.00 (3–6) | 31.50 (21–40) | +| zero-yield tag chunks | 0/7 | 0/18 | +| empty tag sections | 0% | 0% | +| generic tags | 0% | 1.6% | +| duplicate tags | 0% | 0% | +| cross-video reuse | 7.7% | 2.4% | + +Tags are markedly MORE specific than chapters on the same videos — the RC-car +channel produced `losi comp crawler`, `savage flux xs`, `mbx6e` against chapters +titled "Introduction" three times over. + +### The surcharge is 5–15%, not ~100% — and the reason decides when to run them + +Both the `sections` settings comment ("tags double the call count") and the +pre-run prediction (~55%) were wrong. Measured with the model **unloaded between +runs** (`keep_alive: 0`, which evicts the KV cache — without it whichever section +runs second gets a free ride and reports ~0.1 s prefill): + +| run, same video, forced | calls | prompt tokens | engine | prefill | decode | +| --- | --- | --- | --- | --- | --- | +| chapters only | 5 | 20,490 | 119.6 s | 28.0 s | 82.2 s | +| chapters + tags | 10 | 40,980 | 137.9 s | **28.4 s** | 93.6 s | +| tags ALONE (cold) | 5 | 20,490 | 52.6 s | **28.0 s** | 11.7 s | + +- **In one combined pass: +4.7% and +15.3%** on the two videos measured. +- **As a separate later pass: 44%** of a whole chapters pass. + +The mechanism, proven by the third row rather than assumed: a tag call re-sends the +same transcript as the chapter call before it, so it **hits the engine's cached +prefix and pays essentially no prefill** (+0.4 s across 5 extra calls against +28.0 s for the first 5). Alone, it pays the full 28.0 s. All the marginal cost is +decode, and a tag list is ~30 output tokens where a chapter list is ~250. + +**Therefore: if tags are wanted at all, generate them in the SAME pass as +chapters.** Deferring them costs 3–9× more and needs a second corpus-wide pass. + +### Quality gap worth knowing before Phase 7 consumes tags + +Cross-video reuse is **2.4%** on the multi-hour sample, and `antisemitism` / +`anti-semitism` land as two distinct tags. The tags are individually good but the +**vocabulary does not converge**, so as a search index they are weak without +normalization or clustering. That is Phase 7's problem; it is now measured rather +than assumed, and `isGenericTag` in `digest-validate.ts` catches the +format-describing tags (`youtube`, `podcast`, `commentary`) that match everything +and partition nothing. diff --git a/plans/STATE.md b/plans/STATE.md @@ -11,25 +11,62 @@ below for what the measurements changed. **The headline is a negative result, and it should be read before planning any more cost work.** The plan's own premise — shrink the work-list via duplicate sharing — does not pay. -Priced in audio-hours by the new `common/bin/digest-plan.ts`: - -| `--near` | clusters | to generate | sweep days | saved by sharing | -| --- | --- | --- | --- | --- | -| 0.6 (was) | 2,846 | 76,804 h | 80.0 | 0.5 d | -| 0.45 | 7,110 | 75,851 h | 79.0 | 1.5 d | -| **0.35 (now)** | **7,434** | **75,638 h** | **78.8** | **1.7 d** | -| 0.25 | 7,493 | 75,616 h | 78.8 | 1.8 d | - -Loosening the threshold as far as it goes buys **1.2 days of 80**. The plan assumed mirrors +Measured by `common/bin/digest-plan.ts`: + +| `--near` | clusters | to generate | saved by sharing | +| --- | --- | --- | --- | +| 0.6 (was) | 2,846 | 76,804 h | 0.5 d | +| 0.45 | 7,110 | 75,851 h | 1.5 d | +| **0.35 (now)** | **7,434** | **75,638 h** | **1.7 d** | +| 0.25 | 7,493 | 75,616 h | 1.8 d | + +(The `sweep days` column this table used to carry has been dropped: it was +audio-hours × 90 s, and that basis is retracted — see the correction below. The +current threshold reprices to **185,483 chunks → 53.5 days** at the measured +24.9 s/chunk, with **5,633 chunks / 1.6 days** avoided by sharing. The +conclusion about sharing is unchanged, because it is a RATIO and both bases +agree it is ~2%.) + +Loosening the threshold as far as it goes buys under 2 days. The plan assumed mirrors skew long; **they skew short**. Cluster members are 19% of the corpus by video count and **4.3% of its audio-hours** — the sweep is dominated by unclustered long-form VODs, and `HasanAbiVODs3` alone (8,331 audio-hours) outweighs every mirror in the corpus combined. Only ~55% of mirrors pass the alignment gate, so half the nominal saving is refused anyway. -**The lever that does pay is GPU arbitration**: 90 s/audio-hour measured against 27 s -projected on an idle box is ~80 days versus ~24 — about fifty times every duplicate lever -put together. It is now implemented (`controller/digestYield.ts`) and **still unmeasured in -production**; measuring it is the single most valuable next action. +**CORRECTION (2026-07-29, later the same day): the paragraph that stood here was +wrong, and it was my own.** It claimed GPU arbitration was "worth ~56 days" and +"about fifty times every duplicate lever put together", from 90 s/audio-hour +measured against 27 s projected. That 2.8× is not a contention penalty. It is +worse than confounded — it compares the validation run's **wall** clock against +round 2's **engine** time, which are not the same quantity, across samples with +different length mixes. + +Re-priced per chunk (the unit a model call is actually billed in), the same three +measurements agree to within ~2% and the factor decomposes exactly: **1.49× from +the validation sample's length mix × 2.22× per-chunk cost = 3.30×**, against the +3.33× observed. The honest corpus range is **~25 days idle to ~55 contended**, and +the ~81-day headline was wrong on the validation run's own data. See +`controller/digestPlan.ts` for the census and the arithmetic. + +**Restated as a hypothesis with a named test.** The remaining 2.22× per-chunk gap +has three parts: idle engine time, contended engine time, and per-chunk non-engine +wall inflation in production (dominated by the yield gate idling at its 3 s poll). +No bake-off run can see the third — a standalone `tsx` process reads `busy: false` +because `getWorkerPool()` is a `globalThis` singleton. The test is a scoped +production sweep run **twice, differing only in `digest.yieldToCpuWorkers`**, with +wall/engine computed from `.jobs/<id>.meta.json` against the summed per-call lines +in `.jobs/<id>.log`. On an idle box the third term is provably ~zero: the +`teamrcn` job ran 41.578 s wall against 41.4 s of summed engine time — **1.004×** +— and it did all the sidecar work a bake-off skips. Under contention it is +unbounded, which is the point of measuring it. + +`controller/digestYield.ts` is implemented, and **it shipped with a bug**: it +yielded for any busy `kind === "local"` worker, but this box runs one GPU worker +beside two pinned to `device: "cpu"`, so at `parallelTranscriptions: 2` the digest +lane stopped dead for transcription competing for zero GPU shaders. Fixed and +made configurable (`digest.yieldToCpuWorkers`, default off). That fix is plausibly +a bigger throughput lever than the model choice, and it is a logic error rather +than physics. **The sweep was run for real, not just built.** Scoped to `teamrcn` (7 videos, 0.27 audio-hours) against ollama and the real corpus: it armed with its scope @@ -46,9 +83,11 @@ stopping a sweep left its channel scope behind, so the next sweep would silently inherit a weeks-old list and report itself finished. Both fixed and re-verified. Throughput datapoint, uncontended: **~151 s per audio-hour** on 2.4-minute -videos. Consistent with "short videos cost ~2x" against the 90 s corpus average. -It is **not** the measurement that matters — that is still contended vs -uncontended, and it is still not done. +videos — which looked alarming and is in fact *exactly on model*. Those 7 videos +are 1 chunk each over 0.27 audio-hours, i.e. 26 chunks/audio-hour against the +corpus's 2.47, and 151/26 = **5.9 s per chunk** on an idle box. Priced in chunks +it is the cheapest datapoint anyone has taken here, not the most expensive. This +is the clearest single illustration of why s/audio-hour is not a usable unit. **Previous entry —** 2026-07-29 — **Phases 2 and 3: the digest layer SHIPS.** The 102 digests already on disk now travel generation → LMDB → a shared `/digests/<slug>/` page tree → @@ -98,12 +137,15 @@ Exploration for Phase 2 found the planning docs describing more unbuilt work tha **Recommended next**, in the order the measurements argue for: -1. **MEASURE THE GPU YIELD IN PRODUCTION.** Everything else is second by a factor of fifty. - Start a transcription job during a digest sweep, confirm the digest lane idles rather than - contending, and measure seconds-per-audio-hour with `yieldToTranscription` on and off. If - the 2.8x is real the sweep is ~24 days, not ~80, and every other estimate in these docs - changes. If it is NOT real, the contention hypothesis is wrong and the 90 s/audio-hour - needs a different explanation — which is just as valuable to know before spending 80 days. +1. **MEASURE THE GPU YIELD IN PRODUCTION** — still the most valuable single action, but + NOT "second by a factor of fifty": that framing came from the retracted 2.8x and is + withdrawn. The re-priced spread is ~25 days idle to ~55 contended, so contention is + worth up to ~30 days against sharing's ~2 — large, and no longer fabricated. + Measure **seconds per CHUNK**, not per audio-hour, with a scoped sweep run twice + differing only in `digest.yieldToCpuWorkers`, and compute wall/engine from + `.jobs/<id>.meta.json` against the summed per-call lines in `.jobs/<id>.log` + (`digestApps.ts` now records load/prefill/decode per call, which is what makes the + result interpretable rather than just a ratio). 2. **Run the sweep for a day and watch it.** The launcher, resume, pause and coverage readouts are all built and unit-tested but have not driven a real multi-channel run. Kill the editor mid-run and confirm it resumes with zero rework (eligibility is re-derived from @@ -281,11 +323,17 @@ the canonical member's own freshness drives regeneration and the share is re-app time (2026-07-27).** Full table in FACTS.md. Quality came in at or better than round 2's projection on every metric — **zero-yield went to literally 0 of 153 chunks**, chapters/hour 14.74 vs 13.37, rejection rate 18.4% vs 19.1% predicted, generic titles 27.0% vs 31.2% — so a -2-video sample turned out to be representative of quality. Throughput was not: **90 s per -audio-hour vs 27**, projecting **~81 sweep days rather than 24.2**, because round 2 measured -only long videos on an idle box and the corpus is mostly short videos on a box also running -whisper. The lesson to carry: **a bake-off sample stratified by content generalizes; one -stratified by length does not generalize to throughput.** +2-video sample turned out to be representative of quality. Throughput looked wrong — **90 s +per audio-hour vs 27**, projecting **~81 sweep days rather than 24.2** — but that comparison +does not survive contact with the unit. **Re-priced per chunk the run's own data says 55 +days, not 81**, and the 3.33x gap factors as 1.49x length mix × 2.22x per-chunk cost. Round 2 +measured only long videos (2.44 chunks/audio-hour) and the validation sample was short-skewed +(3.63), so a third of the "penalty" was the sample. + +The lesson to carry is therefore sharper than it first looked: **a bake-off sample stratified +by content generalizes; one stratified by length does not generalize to any per-audio-hour +figure at all** — because chunk density spans 4x across the corpus, so s/audio-hour is not a +unit, it is a property of the sample. The one metric that degraded is the worst coverage gap (1:09:48 vs 24:13 projected), and the new per-video panel explains it rather than leaving it a mystery: the worst video recorded 13