# Attribution bake-off — round3-closed-cast-8 Variant **closed-cast** · 8 video(s) · model `qwen2.5:7b` @ numCtx 8192 (maxCues 600) · 2026-08-08T07:07:29.963Z **No sidecar was written.** `attribution.json` count 9 before, 9 after. ## Verdict against the thresholds fixed before the run | criterion | bar | measured | | | --- | ---: | ---: | :-: | | new labels/chunk (mean) | ≤ 0.3 | 0.27 | PASS | | transcript-text labels (worst video) | 0 | 0% | PASS | | s/chunk engine | < 25 | 28.6 | FAIL | | off-cast labels (grammar held) | 0 | 0 | PASS | **NOT MET.** Round 1 (open schema, same model/context/videos) measured **14.7 new labels/chunk** and **66.2 s/chunk engine**. Digest lane baseline: 11.2 s/chunk engine. Single-label (degenerate-risk) videos: 4 of 8. A perfect headline with one label at ~100% talk time is worthless — read the two together. ## Pass A — cast discovery (1 call per video) | video | cast | rejected | call s | | --- | --- | ---: | ---: | | the-quartering-rumble/v6xl0vu | Jeremy (0.90), Stephen Colbert (0.85), Donald Trump (0.90) | 0 | 9.4 | | chibi-reviews/VQykVuHd9xQ | Reviewer (0.90) | 0 | 3.4 | | the-quartering/J0ySGwzP4Nw | Tim Pool (0.90), Twitter (0.50), Caller (0.50) | 0 | 4.2 | | destiny/5nmDzKB23OU | Caller (0.50), Host (0.90) | 0 | 3.8 | | angryjoeshow/RPJqkewZP5I | AngryJoe (0.90), Caller (0.50) | 0 | 4.2 | | rekietalaw/EsZhaCfc8HQ | Nick Ricada (0.90), Scarlett Hampton (0.50), Jacob Frey (0.20), Tim Walls (0.20), Cecil (0.50), Trinity Beholder (0.50), Dick Masterson (0.20) | 0 | 6.5 | | chrissie-mayr/cfLF2o2-0BA | Frankie McDonald (0.90), Chrissie Mayr (1.00) | 0 | 3.9 | | HasanAbiVODs/GjPX_ueTdfc | Cody (0.50), Troop AJ Hunter (0.90), Lane (0.50), Tony (0.50), Raymond (0.50) | 0 | 5.6 | **Pass A is independently valuable.** If pass B churns, this alone is the cast-list product — ~1 call per video, order 10 days corpus-wide — shipping as "videos featuring X" without ever claiming who said which line. ## Cross-chunk identity, and the label space | video | chunks | labels | new/chunk | round-1 new/chunk | transcript-text | truncated | singleton | top share | | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | | the-quartering-rumble/v6xl0vu | 1 | 2 | n/a | — | 0% | 0% | n/a | 85% | | chibi-reviews/VQykVuHd9xQ | 1 | 1 | n/a | — | 0% | 0% | n/a | 100% | | the-quartering/J0ySGwzP4Nw | 1 | 1 | n/a | — | 0% | 0% | n/a | 100% | | destiny/5nmDzKB23OU | 4 | — | n/a | 0.33 | n/a | — | n/a | — | | angryjoeshow/RPJqkewZP5I | 3 | 1 | 0.50 | — | 0% | 0% | 100% | 100% | | rekietalaw/EsZhaCfc8HQ | 8 | 2 | 0.29 | 12.57 | 0% | 0% | 50% | 81% | | chrissie-mayr/cfLF2o2-0BA | 9 | 1 | 0.13 | — | 0% | 0% | 100% | 100% | | HasanAbiVODs/GjPX_ueTdfc | 13 | 2 | 0.17 | — | 0% | 0% | 0% | 65% | ## Pass B — how often the model declined | video | turns accepted | Unknown | off-cast | out-of-range | | --- | ---: | ---: | ---: | ---: | | the-quartering-rumble/v6xl0vu | 2 | 3 | 0 | 0 | | chibi-reviews/VQykVuHd9xQ | 2 | 0 | 0 | 0 | | the-quartering/J0ySGwzP4Nw | 5 | 5 | 0 | 0 | | destiny/5nmDzKB23OU | 0 | 64 | 0 | 103 | | angryjoeshow/RPJqkewZP5I | 3 | 0 | 0 | 177 | | rekietalaw/EsZhaCfc8HQ | 88 | 125 | 0 | 0 | | chrissie-mayr/cfLF2o2-0BA | 83 | 257 | 0 | 60 | | HasanAbiVODs/GjPX_ueTdfc | 155 | 393 | 0 | 180 | `Unknown` is a GOOD answer, not a failure — the bias must run toward dropping on uncertainty. `off-cast` must be 0: anything else means the enum did not convert to a grammar and the whole hypothesis is refuted. ## What the labels say - `the-quartering-rumble/v6xl0vu` — Jeremy 15%, Donald Trump 85% - `chibi-reviews/VQykVuHd9xQ` — Reviewer 100% - `the-quartering/J0ySGwzP4Nw` — Tim Pool 100% - `destiny/5nmDzKB23OU` — (no speaker) - `angryjoeshow/RPJqkewZP5I` — AngryJoe 100% - `rekietalaw/EsZhaCfc8HQ` — Nick Ricada 81%, Trinity Beholder 19% - `chrissie-mayr/cfLF2o2-0BA` — Chrissie Mayr 100% - `HasanAbiVODs/GjPX_ueTdfc` — Tony 35%, Raymond 65% Warnings by code: none ## Cost 40/40 chunk(s) ok, 48 call(s) total (40 chunk call(s) with engine timing). Sample density **2.26 chunks/audio-hour** — the corpus census reads 2.46. | quantity | value | round 1 | | --- | ---: | ---: | | s/chunk, wall | 28.6 | — | | s/chunk, engine | 28.6 | 66.2 | | s/chunk, engine excl. model load | 28.0 | — | | mean model-load s/call | 0.22 | — | | decode tokens/s | 42.8 | — | | TOTAL output tokens | 38799 | — | | cast call s/VIDEO (not per chunk) | 5.1 | — | | digest chapters baseline, s/chunk engine | 11.2 | — | Output tokens are the thing to watch: engine time tracked wall time within 7% in round 1, so 66.2 s/chunk was pure decode volume (31,266 output tokens on one 13-chunk video). An enum token per turn instead of a 60-character string is the mechanism by which this variant is meant to be cheaper as well as better. ## Corpus projection — per-turn lane Census: 74,329 transcribed video(s), **194,061 chunks** at maxCues 600. **62.9 days idle to 64.2 days contended** — from n = 8 video(s) / 40 chunk(s), one box, one hour. Cast-only, for comparison: **1 call per video** over 74,329 videos at 5.1s = 4.4 days. This is a SECOND sweep on the same 8 GB card, additive to a digest sweep that has completed 0.17% of its own ~194k calls, with nothing arbitrating between them. ## Box Start: `{"loadavg":[0.13,0.35,1.02],"freeMemGb":11.35,"totalMemGb":15.53}` End: `{"loadavg":[0.44,0.77,1.02],"freeMemGb":3.69,"totalMemGb":15.53}` ## Units - The unit of per-turn work is the **chunk** (one model call). Seconds-per-audio- hour is not a unit here: chunk density varies 4× across this corpus, and pricing in it is the error behind a retracted throughput headline. This report does not print one. - The unit of cast discovery is **calls per video**, priced separately and never folded into s/chunk. - Wall and engine are different quantities and are never substituted for each other. Wall is the realistic ceiling on this shared box; engine excl. load is the floor. - Every rate carries its denominator. At n = 8 videos a bare percentage would be a lie of precision.