commit 112c773e1c772408b0d5974c18381b19aa9dc760
parent cf38a358bc11d38230c8d51b3f40a9a49fe380c5
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Sat, 8 Aug 2026 06:52:32 -0400
Retract the peak-RSS number, and ship a duration cap
CORRECTION: the '~505 MB per audio-hour' recorded earlier was a single ps sample
taken mid-run, not a peak, and it should not have been written down as one. The
real peak came from the kernel: it OOM-killed a diarization at 11.1 GB anon-RSS
(34.6 GB virtual) on a 6h12m Twitch VOD, after 35 minutes of engine time.
Memory is not linear in duration -- a 7h42m file finished fine and a 6h12m file
died -- because the clusterer holds a pairwise distance matrix over speech
segment embeddings, O(n^2) in SEGMENT COUNT. Dense fast-turnover streams produce
far more segments per hour, which also explains the 120 vs 373 s/audio-hour
spread already recorded. Duration is only a proxy.
That matters because 26 of the 67 remaining videos are 4h+, covering 149 of the
216 remaining audio-hours: most of the remaining work sits in the size class
that already failed once.
--max-audio-hours defers longer videos and reports loudly what it deferred, so a
cap can never read as 'the corpus is done'. NOT applied to the running job:
failures are non-destructive and scaling down approved scope is the operator's
call. Nothing is lost by deferring -- the cleanup guard holds the audio.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Diffstat:
2 files changed, 90 insertions(+), 1 deletion(-)
diff --git a/common/bin/diarize-backfill.ts b/common/bin/diarize-backfill.ts
@@ -232,6 +232,25 @@ async function main(): Promise<void> {
console.log("Scanning for diarizable media…");
const all = await collectCandidates(paths, onlyChannel, (m) => console.log(m));
+ // --max-audio-hours: skip files longer than this.
+ //
+ // MEASURED, NOT PRECAUTIONARY. A 6h12m Twitch VOD was OOM-killed by the kernel
+ // at 11.1 GB anon-RSS (34.6 GB virtual) on this 16 GB box, after burning 35
+ // minutes. Memory is NOT linear in duration — a 7h42m file completed fine at
+ // ~3.9 GB — because agglomerative clustering holds a pairwise distance matrix
+ // over speech-segment embeddings, which is O(n^2) in SEGMENT COUNT. A dense,
+ // fast-turnover stream produces far more segments per hour than a monologue,
+ // so duration is only a proxy. It is the proxy we have.
+ //
+ // An OOM costs the whole video's engine time and puts a 16 GB box under real
+ // pressure with other work on it, so capping is how a long run is made to
+ // finish rather than thrash. The skipped videos are NOT lost: they keep their
+ // audio (the cleanup guard holds it) and can be run later, on a bigger box or
+ // with a segment-bounded clusterer.
+ const maxAudioHours = flags["max-audio-hours"]
+ ? Number(flags["max-audio-hours"])
+ : undefined;
+
let selected = all;
if (onlyVideo) selected = all.filter((c) => c.slug === onlyVideo);
else if (scope === "transcribed") selected = all.filter((c) => c.transcribed);
@@ -244,7 +263,14 @@ async function main(): Promise<void> {
// freshness, and re-deriving that here would be a second definition), but they
// are excluded from the ETA so the estimate reflects real work.
const todo = force ? selected : selected.filter((c) => !c.hasDiarization);
- const ordered = limitCount ? todo.slice(0, limitCount) : todo;
+ // Deferred, not dropped — and SAID so, loudly. A cap that silently shrinks the
+ // work reads as "the corpus is done" when it is not.
+ const deferred =
+ maxAudioHours === undefined
+ ? []
+ : todo.filter((c) => c.durationSeconds > maxAudioHours * 3600);
+ const capped = maxAudioHours === undefined ? todo : todo.filter((c) => !deferred.includes(c));
+ const ordered = limitCount ? capped.slice(0, limitCount) : capped;
const audioSeconds = ordered.reduce((a, c) => a + c.durationSeconds, 0);
const unknownDuration = ordered.filter((c) => c.durationSeconds === 0).length;
@@ -258,6 +284,18 @@ async function main(): Promise<void> {
`Selected by scope: ${selected.length} · to run: ${ordered.length}` +
(limitCount ? ` (--limit ${limitCount})` : ""),
);
+ if (deferred.length) {
+ const hours = deferred.reduce((a, c) => a + c.durationSeconds, 0) / 3600;
+ console.log(
+ `DEFERRED by --max-audio-hours ${maxAudioHours}: ${deferred.length} video(s), ` +
+ `${hours.toFixed(0)} audio-hour(s). They keep their audio and can be run later — ` +
+ "the cap exists because the clusterer is O(n^2) in segment count and long dense " +
+ "streams get OOM-killed. Longest deferred:",
+ );
+ for (const c of [...deferred].sort((a, b) => b.durationSeconds - a.durationSeconds).slice(0, 5)) {
+ console.log(` ${fmtHms(c.durationSeconds)} ${c.slug}`);
+ }
+ }
console.log(
`Audio to process: ${(audioSeconds / 3600).toFixed(0)} audio-hour(s)` +
(unknownDuration ? ` (+${unknownDuration} of unknown duration)` : ""),
diff --git a/plans/FACTS.md b/plans/FACTS.md
@@ -1753,3 +1753,54 @@ Two consequences to expect, both small and both permanent until it is fixed:
The fix is to require a known media extension rather than denylisting; it is
deferred to its own branch because it changes cleanup behaviour and several
exact-string e2e assertions.
+
+### CORRECTION + escalation: the kernel OOM-killed a diarization run
+
+**The "~505 MB per audio-hour" figure recorded above is wrong, and the error was
+in the method: it was a single `ps` sample taken mid-run, not a peak.** Retract
+it as a peak-memory number.
+
+The measured peak, from `journalctl -k`, on `hasanabi/2834644359` (6h12m):
+
+```
+Out of memory: Killed process (python) total-vm:34643056kB, anon-rss:11138996kB
+```
+
+**11.1 GB resident, 34.6 GB virtual — killed by the kernel after 35 minutes of
+engine time.** No sidecar was written; nothing in the corpus was damaged; the
+backfill continued to the next video.
+
+**Memory is NOT linear in duration, so duration is a proxy and not the cause.**
+A **7h42m** file completed fine, and a **6h12m** file died. The mechanism is
+agglomerative clustering holding a pairwise distance matrix over speech-segment
+embeddings — **O(n²) in SEGMENT COUNT**. A dense, fast-turnover Twitch VOD yields
+far more segments per hour than a monologue or a reaction video, so content
+density is what decides it. This also explains the throughput spread already
+recorded (120 s/audio-hour on short files vs 373 on a long one).
+
+Consequences, measured against the remaining work:
+
+| | |
+| --- | --- |
+| remaining after video 19 | 67 videos, 216 audio-hours |
+| of those **≥ 4h** | **26 videos, 149 audio-hours** |
+| of those ≥ 6h | 10 videos, 77 audio-hours |
+
+So most of the remaining audio sits in the size class that OOM'd once. Each OOM
+costs that video's full engine time (~35 min here) and puts an 11 GB transient on
+a 16 GB box — the kernel chose python both by luck and by it being the largest
+RSS, and that is not a guarantee.
+
+**Mitigation shipped, not applied:** `--max-audio-hours` on
+`bin/diarize-backfill.ts` defers longer videos and **says loudly what it
+deferred** (a cap that silently shrinks the work reads as "the corpus is done").
+`--max-audio-hours 4` leaves 45 videos / 74 audio-hours (~6 h at the observed
+300 s/audio-hour) and defers 27 / 155.
+
+It is **deliberately not applied to the running job**: scaling down approved
+scope is the operator's call, and the failures are non-destructive. Deferred
+videos lose nothing — the cleanup guard holds their audio, so they can be run
+later on a bigger box or with a segment-bounded clusterer.
+
+`diarization.concurrency: 1` is now doubly load-bearing: two 11 GB transients
+would not merely be tight, they would be unschedulable.