# `.diarize/` — the diarization runtime The engines `scripts/diarize.mjs` needs. **Everything here except this file is gitignored** — `env/` and `models/` for sherpa-onnx (~148 MB) and `sortformer/` for the ggml engine (~530 MB), all machine-specific. This file is tracked so the environment is reproducible. Two engines live here, selected by `settings.diarization.engine`: | | `sherpa-onnx` (default) | `sortformer` | | --- | --- | --- | | method | pyannote segmentation + TitaNet embeddings, agglomerative clustering | end-to-end streaming Sortformer | | device | CPU only | Vulkan or CPU | | speakers on `v6z1o2g` | 13 (ground truth: 1 host + clips) | **4** | | speakers on the corpus's worst file | 35 | **4** | | dominant-speaker share | 73.1% | 73.5% | | throughput | **492–585 s/audio-hour**, ~3.9 cores | 894 s/audio-hour on Vulkan (~1 core), 1305 on tuned CPU | | memory | ~230 MB/audio-hour, O(n²) in turn count | **558 MB flat**, O(1) in duration | | knobs | threshold (0.9 here) | none | The sherpa range is two clean runs of the SAME file rather than an estimate: 585 s/audio-hour on 2026-08-08 (§ below) and 492 on 2026-08-10, both on an idle box. Treat sub-20% differences between runs as noise on this hardware. sherpa-onnx is FASTER in wall clock. sortformer is chosen for quality: over-splitting is the failure mode this lane has always had, and an end-to-end model does not have it. See `plans/diarization-spike-results.md` for the original CPU-only spike and the measurements that decided the defaults. ## Why it exists at all `scripts/diarize.mjs` and `scripts/diarize-sherpa.py` shipped and work — the 2026-08-07 spike (`plans/diarization-spike-results.md`) drove them end to end and produced the corpus's only `diarization.json`. What did *not* survive was the environment: the spike's venv was ephemeral, so the code was shipped and unrunnable. This directory is that gap closed. **sherpa-onnx publishes wheels only up to cp313 and this box's system Python is 3.14.6**, so a dedicated interpreter is not a preference — it is the reason `settings.diarization.python` is a configurable path rather than `python3`. ## Rebuilding it From the repo root, with `uv` on PATH (`~/.local/bin/uv`): ```sh uv venv --python 3.13 .diarize/env uv pip install --python ./.diarize/env/bin/python sherpa-onnx numpy mkdir -p .diarize/models && cd .diarize/models curl -sSLO https://github.com/k2-fsa/sherpa-onnx/releases/download/speaker-segmentation-models/sherpa-onnx-pyannote-segmentation-3-0.tar.bz2 tar xf sherpa-onnx-pyannote-segmentation-3-0.tar.bz2 && rm sherpa-onnx-pyannote-segmentation-3-0.tar.bz2 curl -sSL -o titanet.onnx https://github.com/k2-fsa/sherpa-onnx/releases/download/speaker-recongition-models/nemo_en_titanet_small.onnx ``` Verified versions, 2026-08-08: Python 3.13.11, **sherpa-onnx 1.13.4**, numpy 2.5.1. The sherpa version matters — it is recorded into every sidecar as `engine.version`, and it is what the reproducibility check below compares. ## Wiring it up Three fields in the (gitignored) `settings.json`; `threshold` 0.9 and `threads` 4 are already the defaults and need no change: ```json "diarization": { "python": "/.diarize/env/bin/python", "segModel": "/.diarize/models/sherpa-onnx-pyannote-segmentation-3-0/model.onnx", "embModel": "/.diarize/models/titanet.onnx" } ``` Absolute paths. `diarize.mjs` records only the BASENAME into the sidecar, because the corpus is rsynced between shards and a full path would make records non-portable. ## Checking a rebuild reproduces the record on disk Read-only, writes nothing to the corpus. `ObviousRises-rumble/v6z1o2g` is the one video with both audio and an existing `diarization.json`: ```sh node scripts/diarize.mjs --output /tmp/repro.json --video-id v6z1o2g \ --python "$PWD/.diarize/env/bin/python" \ --seg "$PWD/.diarize/models/sherpa-onnx-pyannote-segmentation-3-0/model.onnx" \ --emb "$PWD/.diarize/models/titanet.onnx" \ --threshold 0.9 --threads 4 \ transcripts/channels/ObviousRises-rumble/data/v6z1o2g/audio.mp3 ``` Then diff `.turns` against the sidecar. **On 2026-08-08 this rebuilt environment produced BYTE-IDENTICAL turns** — 51 turns, 13 speakers, `audioSeconds` 766.101 — against a sidecar generated 2026-08-07 by the spike's now-deleted venv. The result is deterministic given the same models and threshold, so a diff is a real check on the environment rather than a smoke test. ## Throughput That run took **124.5 s of wall for 766 s of audio = 585 s/audio-hour**, at 381% of the 400% available CPU, on a box at loadavg 2.7. This closes an item `plans/diarization-spike-results.md` §8 left open ("the clean throughput number" — every figure there was contended). It lands inside the spike's estimated 500–700 s/audio-hour band, so **no conclusion changes**: at 585 s/audio-hour diarization is still ~2.6× slower than the GPU's measured 221 s/audio-hour ASR, and the spike's decision gate ("if throughput can't keep pace with GPU transcription, prefer inline-after-transcribe over a parallel CPU lane") still fires. For comparison the same file took 277.7 s (1305 s/audio-hour) when the spike ran it under load ~30. Memory: ~230 MB per audio-hour of decoded float32 audio, so an 8-hour VOD needs ~1.8 GB resident. Untested on the long tail. --- ## The sortformer engine ### Rebuilding it ```sh scripts/build-sortformer.sh # Vulkan (default) SORTFORMER_BACKEND=cpu scripts/build-sortformer.sh ``` Needs `git`, `cmake`, a C++17 compiler, and for Vulkan the loader + headers and `glslc` (Arch: `vulkan-headers shaderc`). Installing `patchelf` is optional — it sets `RPATH=$ORIGIN` on the binary; without it the wrapper sets `LD_LIBRARY_PATH` itself. The script clones **openresearchtools/engine at a pinned commit** (`8bb4928c`, 2026-03-21), stages only `ggml` + `tools/realtime` + our driver, and builds those. It deliberately uses none of upstream's build system, whose scripts are PowerShell/CUDA — the Sortformer code links against `ggml` alone (no llama, no whisper, no Rust), which is what makes that possible. First build is ~10 minutes, almost all of it compiling Vulkan shaders; rebuilds after a driver edit are seconds, because `ggml` is staged once per pin. The model (471 MB, F32) is downloaded from `openresearchtools/diar_streaming_sortformer_4spk-v2.1-gguf`. **It is under the NVIDIA Open Model License, not the engine's MIT** — the GGUF is a conversion of `nvidia/diar_streaming_sortformer_4spk-v2.1`. ### Wiring it up ```json "diarization": { "engine": "sortformer", "backend": "vulkan", "sortformerBin": "/.diarize/sortformer/diarize-file", "sortformerModel": "/.diarize/sortformer/diar_streaming_sortformer_4spk-v2.1.gguf" } ``` `threshold`, `segModel`, `embModel` and `python` are ignored by this engine and do not enter its freshness identity. **Switching `engine` marks every sidecar written by the other one stale**, which is intended — see `diarizationTarget` — but on the retained audio that is weeks of rework, not a toggle. ### Checking a rebuild reproduces the record on disk Read-only, writes nothing to the corpus: ```sh node scripts/diarize.mjs --output /tmp/repro.json --video-id v6z1o2g \ --engine-kind sortformer \ --sortformer-bin "$PWD/.diarize/sortformer/diarize-file" \ --sortformer-model "$PWD/.diarize/sortformer/diar_streaming_sortformer_4spk-v2.1.gguf" \ --backend vulkan --threads 6 \ transcripts/channels/ObviousRises-rumble/data/v6z1o2g/audio.mp3 ``` Expected on 2026-08-10: **66 turns, 4 speakers, `audioSeconds` 766.101**, in 206.5 s (3.71x realtime). Two independent equivalences were checked and both held exactly: * the **CPU and Vulkan backends produce byte-identical turns**, so the backend is a throughput choice and never a quality one; and * piping mp3 through ffmpeg into the streaming stdin path gives the **same 66 turns** as running the binary directly on a decoded WAV. ### Notes for whoever touches the driver `scripts/sortformer/diarize-file.cpp` is vendored here rather than taken from upstream because upstream's `llama-realtime-smoke` is a parity tool: it retains every intermediate matrix, demands PyTorch reference fixtures, and — the fatal part — drops the event flags, so its JSON mixes 2,033 *preview* re-emissions in with the 66 real spans. Two things in it are non-obvious and both were found the hard way: * **stdin is forced back to blocking.** Node hands children non-blocking pipes, so `read()` returns `EAGAIN` long before EOF; it worked from a shell and failed under the app. * **thread count is set explicitly.** The upstream Sortformer path never sets one, so the CPU backend silently runs at ggml's default of 4. On this 4c/8t box 6 is the optimum (1305 s/audio-hour) and **8 is worse than 4** (1401) through oversubscription.