# Attribution bake-off — round2-closed-cast Variant **closed-cast** · 2 video(s) · model `qwen2.5:7b` @ numCtx 8192 (maxCues 600) · 2026-08-08T06:12:27.088Z **No sidecar was written.** `attribution.json` count 9 before, 9 after. ## Verdict against the thresholds fixed before the run | criterion | bar | measured | | | --- | ---: | ---: | :-: | | new labels/chunk (mean) | ≤ 0.3 | 0.29 | PASS | | transcript-text labels (worst video) | 0 | 0% | PASS | | s/chunk engine | < 25 | 24.1 | PASS | | off-cast labels (grammar held) | 0 | 0 | PASS | **ALL THRESHOLDS MET.** Round 1 (open schema, same model/context/videos) measured **14.7 new labels/chunk** and **66.2 s/chunk engine**. Digest lane baseline: 11.2 s/chunk engine. Single-label (degenerate-risk) videos: 0 of 2. A perfect headline with one label at ~100% talk time is worthless — read the two together. ## Pass A — cast discovery (1 call per video) | video | cast | rejected | call s | | --- | --- | ---: | ---: | | destiny/5nmDzKB23OU | Caller (0.50), Host (0.90) | 0 | 8.9 | | rekietalaw/EsZhaCfc8HQ | Nick Ricada (0.90), Scarlett Hampton (0.50), Jacob Frey (0.20), Tim Walls (0.20), Cecil (0.50), Trinity Beholder (0.50), Dick Masterson (0.20) | 0 | 7.0 | **Pass A is independently valuable.** If pass B churns, this alone is the cast-list product — ~1 call per video, order 10 days corpus-wide — shipping as "videos featuring X" without ever claiming who said which line. ## Cross-chunk identity, and the label space | video | chunks | labels | new/chunk | round-1 new/chunk | transcript-text | truncated | singleton | top share | | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | | destiny/5nmDzKB23OU | 4 | — | n/a | 0.33 | n/a | — | n/a | — | | rekietalaw/EsZhaCfc8HQ | 8 | 2 | 0.29 | 12.57 | 0% | 0% | 50% | 81% | ## Pass B — how often the model declined | video | turns accepted | Unknown | off-cast | | --- | ---: | ---: | ---: | | destiny/5nmDzKB23OU | 0 | 64 | 0 | | rekietalaw/EsZhaCfc8HQ | 88 | 125 | 0 | `Unknown` is a GOOD answer, not a failure — the bias must run toward dropping on uncertainty. `off-cast` must be 0: anything else means the enum did not convert to a grammar and the whole hypothesis is refuted. ## What the labels say - `destiny/5nmDzKB23OU` — (no speaker) - `rekietalaw/EsZhaCfc8HQ` — Nick Ricada 81%, Trinity Beholder 19% Warnings by code: none ## Cost 12/12 chunk(s) ok, 14 call(s) total (12 chunk call(s) with engine timing). Sample density **2.55 chunks/audio-hour** — the corpus census reads 2.46. | quantity | value | round 1 | | --- | ---: | ---: | | s/chunk, wall | 24.1 | — | | s/chunk, engine | 24.1 | 66.2 | | s/chunk, engine excl. model load | 23.3 | — | | mean model-load s/call | 0.34 | — | | decode tokens/s | 41.3 | — | | TOTAL output tokens | 8923 | — | | cast call s/VIDEO (not per chunk) | 7.9 | — | | digest chapters baseline, s/chunk engine | 11.2 | — | Output tokens are the thing to watch: engine time tracked wall time within 7% in round 1, so 66.2 s/chunk was pure decode volume (31,266 output tokens on one 13-chunk video). An enum token per turn instead of a 60-character string is the mechanism by which this variant is meant to be cheaper as well as better. ## Corpus projection — per-turn lane Census: 74,329 transcribed video(s), **194,061 chunks** at maxCues 600. **52.4 days idle to 54.0 days contended** — from n = 2 video(s) / 12 chunk(s), one box, one hour. Cast-only, for comparison: **1 call per video** over 74,329 videos at 7.9s = 6.8 days. This is a SECOND sweep on the same 8 GB card, additive to a digest sweep that has completed 0.17% of its own ~194k calls, with nothing arbitrating between them. ## Box Start: `{"loadavg":[6.23,4.5,3.66],"freeMemGb":7.83,"totalMemGb":15.53}` End: `{"loadavg":[4.92,5.16,4.2],"freeMemGb":5.42,"totalMemGb":15.53}` ## Units - The unit of per-turn work is the **chunk** (one model call). Seconds-per-audio- hour is not a unit here: chunk density varies 4× across this corpus, and pricing in it is the error behind a retracted throughput headline. This report does not print one. - The unit of cast discovery is **calls per video**, priced separately and never folded into s/chunk. - Wall and engine are different quantities and are never substituted for each other. Wall is the realistic ceiling on this shared box; engine excl. load is the floor. - Every rate carries its denominator. At n = 2 videos a bare percentage would be a lie of precision.