# Diarization spike — results **Date:** 2026-08-07 **Engine:** sherpa-onnx 1.13.4 (Python), pyannote segmentation-3.0 ONNX + NeMo TitaNet-small speaker embeddings, CPU. **Box:** i7-6700 (4c/8t, 2015), 16 GB RAM, RX 6600 XT (gfx1032). Phase A of the "Capture speaker diarization on new transcriptions" plan, which required evaluating on real audio before committing because *"quality here cannot be judged from code."* Both unknowns it named — throughput on this box, and quality on this content — are answered below, and both changed the design. --- ## READ THIS FIRST: the box was contended throughout The plan says to measure the box before timing anything. That warning earned its place. `/proc/loadavg` ranged from **9 to 49** during these runs, and the same 6-minute clip took anywhere from **498 to 6340 s/audio-hour** depending on when it landed. Two separate sources, and the second one is a trap worth writing down: 1. Another project's dev server and Playwright suite, running concurrently — the hazard the plan already warned about. 2. **Leaked `fake-ytdlp.mjs` fixture processes from this repo's own e2e runs.** 27 of them had accumulated, each burning ~29% of a core for over two hours, after e2e runs were interrupted. On a 4-core box that alone is most of the machine. Killing them took load from 46 to 2.7. So "check the load before timing" is not enough on its own — check *what* is producing it. A leaked fixture pool looks exactly like someone else's job, and it silently degraded a full e2e run from 45 minutes to 1.3 hours and turned 4 failures into 13. Every throughput number below is therefore a **contended** number and should be read as an upper bound on cost, not a clean measurement. The quality numbers are unaffected: clustering output does not depend on how busy the CPU is. A clean re-measurement is queued (`quiet-measure.sh` waits for load < 8). The design conclusion does not depend on it — see "Why the gate fires either way". --- ## 1. Throughput ### The comparison that matters GPU transcription was measured for free from the corpus's own `transcribe-outcome.json` records — **3,602 videos, 10,790 audio-hours**, real production runs, not a benchmark: | Lane | Rate | Realtime factor | | --- | --- | --- | | parakeet.cpp ASR, Vulkan, RX 6600 XT | **221 s/audio-hour** | 16.3× | | sherpa-onnx diarization, CPU, best observed | ~500 s/audio-hour | ~7× | | sherpa-onnx diarization, CPU, typical contended | 500–700 s/audio-hour | 5–7× | | sherpa-onnx diarization, CPU, heavily contended | up to 6340 s/audio-hour | 0.6× | **Diarization is 2–3× slower than the transcription it would follow, at best.** ### Why the gate fires either way The plan's decision gate: *"if throughput can't keep pace with GPU transcription, prefer inline-after-transcribe over a parallel CPU lane."* Even the single fastest observation (498 s/audio-hour, taken while load was lowest) is **2.3× slower** than the GPU's long-run production average. The gate fires under every reading of the data, so the design does not hinge on the pending clean measurement. But the gate's own remedy needs one correction for this corpus. Inline was preferred because a parallel lane lets audio pile up. That is true, and the disk numbers are alarming — see §3 — but inline has a cost the gate does not price: it makes the whole pipeline ~3–4× slower and idles the GPU while the CPU works. With **749 videos / 4,100 audio-hours** of retained audio pending transcription, that is the difference between ~10.5 days and ~42 days of wall clock. **So what shipped is both, switchable, with inline OFF by default:** - `diarization.enabled` arms the **cleanup guard** — the sweep stops deleting audio for a transcribed-but-undiarized video. This is what actually protects the perishable input, and it is independent of *when* diarization runs. - `diarization.inlineAfterTranscribe` (default **off**) runs it in the post-transcribe hook, for steady state. - The `diarize-channel` job backfills over retained audio at full CPU speed without holding up the GPU. For the batch that is about to run, the intended sequence is: enable capture (holds the audio), leave inline off (batch runs at full GPU speed), then backfill. ### Cost of the backfill 4,338 audio-hours of retained audio exist today (836 files, 161.5 GB). | Rate | Backfill wall time | | --- | --- | | 500 s/audio-hour | ~25 days | | 676 s/audio-hour | ~34 days | Single-threaded-lane figures. This is the same order as the digest sweep (~25–55 days) and competes with it for the same 8 threads. ### Peak RSS 482 MB on a 13-minute file. Memory is dominated by holding decoded audio as float32 at 16 kHz — ~230 MB per audio-hour — so a 5-hour VOD needs ~1.2 GB resident. On a 16 GB box with ~1 GB free and 7 GB of swap already in use, the long tail (the corpus has 8-hour VODs) is a real constraint. Not addressed here; noted for whoever runs the backfill. --- ## 2. Quality ### The default threshold is wrong for this content sherpa-onnx defaults to a clustering threshold of 0.5. On a 6-minute excerpt of **MommaOcco/v1cnNSjEmZk — a two-person interview with @ProtonJon, so ground truth is exactly 2 speakers** — that produced **22 clusters**. Sweep on the same clip: | threshold | speakers | turns | top-4 talk-time share | | --- | --- | --- | --- | | 0.4 | 23 | 62 | 38%, 31%, 13%, 3% | | 0.5 (sherpa default) | 22 | 64 | 38%, 31%, 17%, 3% | | 0.6 | 17 | 64 | 40%, 34%, 18%, 3% | | 0.7 | 12 | 65 | 39%, 38%, 19%, 1% | | 0.8 | 10 | 64 | 39%, 39%, 20%, 1% | | **0.9** | **6** | 61 | **40%, 40%, 20%, 0%** | **The shipped default is 0.9.** At 0.9 the top two clusters sit at 40%/40% — recognizably the two hosts — with a 20% third cluster and a negligible tail. The *shape* is right even where the count is not. Corroboration: on a 30-second excerpt of the same interview the wrapper returned exactly **2 speakers**, the correct answer. Over-splitting grows with duration, which is the expected failure mode for agglomerative clustering over a long recording. ### Confirmed on a full-length real file, not just an excerpt The end-to-end verification run — `ObviousRises-rumble/v6z1o2g`, the whole 12.8-minute reaction video, through the shipped `scripts/diarize.mjs` at the shipped default of 0.9: | | th=0.5 | th=0.9 | | --- | --- | --- | | speakers | 29 | **13** | | turns | 59 | 51 | | dominant cluster | 70.7% | **73%** | The tail collapses (5%, 3%, 2%, 2% after the top two) while the host's share holds. This matters because the excerpt sweep could have been an artifact of 6-minute clips; it is not — the same improvement shows on the full file. Cost of that run: 274 s of engine time for 12.8 minutes of audio, i.e. 1290 s/audio-hour, at load ~30 and only 248% of the available 400% CPU. Another contended number, and a reminder of how much contention costs here. ### Other content types | Clip | Content | th=0.9 result | | --- | --- | --- | | ObviousRises-rumble/v6z1o2g (6 min) | reaction, plays third-party clips | 8 speakers, top 60% (the host) | | kirsche/LEu6R1kn7Gs (6 min @ 1h in) | solo stream | 3 speakers @ th=0.4, top 60%/25%/15% | The reaction clip's 8 clusters are not obviously wrong — it genuinely contains several third-party voices. The solo stream returning 3 clusters at th=0.4 is over-split; higher thresholds were not reached for it before the run was cut. The full 13-minute reaction video at th=0.5 produced **29 speakers across 59 turns**, with cluster IDs running non-contiguously up to 57 — near-zero merging. One cluster held 70.7% of talk time (the host), which is the correct shape buried in noise. That single result is what motivated the whole threshold sweep. ### Verdict on quality `PLAN.md:442-443` set the expectation: *"this misfires on rapid back-and-forth, and auto-caption channels have no speaker turns at all."* That holds. What the spike adds: - **Talk-time distribution is more trustworthy than speaker count.** The dominant cluster is reliably the host across every file tested. A downstream pass that asks "which cluster is the main speaker" will do much better than one that trusts the cluster count. - **Over-splitting is the failure mode, not under-splitting.** That is the benign direction: merging clusters later is a solvable problem, and the turns and their boundaries are recorded either way. Splitting a cluster that was wrongly merged would need the audio back. - **It is worth capturing.** Not because the output is good enough to show a user — it is not — but because the boundaries and the talk-time structure are real, they are recoverable into something better, and they are unobtainable once the audio is gone. --- ## 3. Disk — the pressure the guard makes visible At the time of writing: **45 GB free, 97% full**, `transcripts/` at 387 GB, of which **161.5 GB is retained audio** across 836 files (42%). Audio averages **33 MB per audio-hour**. If diarization ran as an unguarded parallel CPU lane while the GPU transcribed, the backlog would grow at roughly (16.3 − 6) ≈ 10 audio-hours of un-diarized audio per wall hour, or **~340 MB per wall hour** — filling the remaining 45 GB in about **5.5 days**. The cleanup guard does not remove that pressure. It makes it **visible and non-destructive** instead of silent and permanent: the sweep reports `awaiting diarization` and the channel's "Est. reclaim" drops to match, rather than the audio quietly disappearing before its diarize job ran. Also worth recording: **only 82 of the 836 retained-audio videos are transcribed.** The other 754 are the pending batch. So the one-shot window this plan exists to catch is genuinely still open. --- ## 4. Why not the GPU `PLAN.md` assumed the card might be the blocker. It is not, and neither is the GPU available. - parakeet.cpp is built `GGML_VULKAN=ON` and ships `libggml-vulkan.so`, so ASR genuinely runs on the RX 6600 XT. That is why transcription hits 16.3×. - **No diarization model has been ported to ggml.** They ship as PyTorch (CUDA/ROCm) or ONNX (CUDA/ROCm/MIGraphX/OpenVINO/DirectML). Neither runtime has a Vulkan compute path on Linux. This is an ecosystem gap, not a hardware limit. The ROCm fallback (`HSA_OVERRIDE_GFX_VERSION=10.3.0` for this gfx1032 card) remains available but was **not** installed — the plan says not to do that speculatively, and the CPU path is sufficient for a backfill that is not on the critical path. Environment note: the system Python is **3.14.6**, and sherpa-onnx publishes wheels only up to cp313. The spike used a dedicated `uv`-managed 3.13 venv, which is why `diarization.python` is a configurable path rather than `python3`. --- ## 5. Reproducing ``` uv venv --python 3.13 diarize-env uv pip install --python ./diarize-env/bin/python sherpa-onnx numpy # models curl -sSLO https://github.com/k2-fsa/sherpa-onnx/releases/download/speaker-segmentation-models/sherpa-onnx-pyannote-segmentation-3-0.tar.bz2 tar xf sherpa-onnx-pyannote-segmentation-3-0.tar.bz2 curl -sSL -o titanet.onnx https://github.com/k2-fsa/sherpa-onnx/releases/download/speaker-recongition-models/nemo_en_titanet_small.onnx scripts/diarize.mjs --output diarization.json --video-id \ --python /bin/python \ --seg <...>/sherpa-onnx-pyannote-segmentation-3-0/model.onnx \ --emb <...>/titanet.onnx --threshold 0.9 ``` ## 6. End-to-end verification on real data `ObviousRises-rumble/v6z1o2g` — a real corpus video with a transcript and audio still on disk — run through the shipped wrapper, with the cleanup guard's decision evaluated read-only before and after (never by running the real sweep, which would delete the audio): ``` before after transcript true true audio files audio.mp3 audio.mp3 diarization false true sweep takes it, diarization OFF true true sweep takes it, diarization ON false --> true ``` That is the whole contract in four lines: with capture armed the sweep refuses the video until the sidecar exists, and releases it the moment it does. The `diarization OFF` column is unchanged in both states, which is what makes enabling the feature reversible rather than a one-way door. ## 7. Verification On a genuinely quiet box (load 2.7, after clearing the leaked fixtures): | Check | Result | | --- | --- | | `common` unit suite | **510 / 511** | | Editor e2e, full | **433 / 434** in 30.1 min | | Diarization e2e (5 specs) | 10/10 across two repeats, and green in-suite | | `tsc --noEmit` in common / editor / export | clean | | `pnpm build` (editor) | clean — e2e runs in dev mode and never prerenders | The two residual failures are both **pre-existing and reproduced independently of this change**: - `digestPlan.test.ts › the corpus chunk census reproduces 191,116` — the real corpus has grown to **194,053** chunks. Reproduced identically on a pristine `HEAD` worktree pointed at the same corpus, so it is not this change. The pinned number needs re-baselining by whoever owns the digest plan; it is deliberately left alone here. - `jobs-batch-tasks-drain › hard Cancel during a drain` — passes 6/6 in isolation. The rotating contention tail. One real regression was found and fixed during verification: appending `, 0 awaiting diarization` to the cleanup sweep's skip breakdown broke an exact string assertion in `pre-clean-availability.spec`. The clause is now emitted only when the count is non-zero, so the message is byte-identical for any install with the capture lane off. ## 8. What is still open - **The clean throughput number.** Queued behind a load<8 wait. Every figure here is contended. - **Thread scaling.** 2 vs 4 vs 8 threads on a 4c/8t part was not measured; queued with the above. - **Peak RSS on a genuinely long file.** Extrapolated (~230 MB per audio-hour), not measured on an 8-hour VOD. - **Whether a better embedding model closes the over-splitting gap.** Only TitaNet-small was tried. 3D-Speaker / WeSpeaker English models are the obvious next comparison, and swapping one is a settings change, not a code change.