Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 112c773e1c772408b0d5974c18381b19aa9dc760
parent cf38a358bc11d38230c8d51b3f40a9a49fe380c5
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sat,  8 Aug 2026 06:52:32 -0400

Retract the peak-RSS number, and ship a duration cap

CORRECTION: the '~505 MB per audio-hour' recorded earlier was a single ps sample
taken mid-run, not a peak, and it should not have been written down as one. The
real peak came from the kernel: it OOM-killed a diarization at 11.1 GB anon-RSS
(34.6 GB virtual) on a 6h12m Twitch VOD, after 35 minutes of engine time.

Memory is not linear in duration -- a 7h42m file finished fine and a 6h12m file
died -- because the clusterer holds a pairwise distance matrix over speech
segment embeddings, O(n^2) in SEGMENT COUNT. Dense fast-turnover streams produce
far more segments per hour, which also explains the 120 vs 373 s/audio-hour
spread already recorded. Duration is only a proxy.

That matters because 26 of the 67 remaining videos are 4h+, covering 149 of the
216 remaining audio-hours: most of the remaining work sits in the size class
that already failed once.

--max-audio-hours defers longer videos and reports loudly what it deferred, so a
cap can never read as 'the corpus is done'. NOT applied to the running job:
failures are non-destructive and scaling down approved scope is the operator's
call. Nothing is lost by deferring -- the cleanup guard holds the audio.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Diffstat:
Mcommon/bin/diarize-backfill.ts | 40+++++++++++++++++++++++++++++++++++++++-
Mplans/FACTS.md | 51+++++++++++++++++++++++++++++++++++++++++++++++++++
2 files changed, 90 insertions(+), 1 deletion(-)

diff --git a/common/bin/diarize-backfill.ts b/common/bin/diarize-backfill.ts @@ -232,6 +232,25 @@ async function main(): Promise<void> { console.log("Scanning for diarizable media…"); const all = await collectCandidates(paths, onlyChannel, (m) => console.log(m)); + // --max-audio-hours: skip files longer than this. + // + // MEASURED, NOT PRECAUTIONARY. A 6h12m Twitch VOD was OOM-killed by the kernel + // at 11.1 GB anon-RSS (34.6 GB virtual) on this 16 GB box, after burning 35 + // minutes. Memory is NOT linear in duration — a 7h42m file completed fine at + // ~3.9 GB — because agglomerative clustering holds a pairwise distance matrix + // over speech-segment embeddings, which is O(n^2) in SEGMENT COUNT. A dense, + // fast-turnover stream produces far more segments per hour than a monologue, + // so duration is only a proxy. It is the proxy we have. + // + // An OOM costs the whole video's engine time and puts a 16 GB box under real + // pressure with other work on it, so capping is how a long run is made to + // finish rather than thrash. The skipped videos are NOT lost: they keep their + // audio (the cleanup guard holds it) and can be run later, on a bigger box or + // with a segment-bounded clusterer. + const maxAudioHours = flags["max-audio-hours"] + ? Number(flags["max-audio-hours"]) + : undefined; + let selected = all; if (onlyVideo) selected = all.filter((c) => c.slug === onlyVideo); else if (scope === "transcribed") selected = all.filter((c) => c.transcribed); @@ -244,7 +263,14 @@ async function main(): Promise<void> { // freshness, and re-deriving that here would be a second definition), but they // are excluded from the ETA so the estimate reflects real work. const todo = force ? selected : selected.filter((c) => !c.hasDiarization); - const ordered = limitCount ? todo.slice(0, limitCount) : todo; + // Deferred, not dropped — and SAID so, loudly. A cap that silently shrinks the + // work reads as "the corpus is done" when it is not. + const deferred = + maxAudioHours === undefined + ? [] + : todo.filter((c) => c.durationSeconds > maxAudioHours * 3600); + const capped = maxAudioHours === undefined ? todo : todo.filter((c) => !deferred.includes(c)); + const ordered = limitCount ? capped.slice(0, limitCount) : capped; const audioSeconds = ordered.reduce((a, c) => a + c.durationSeconds, 0); const unknownDuration = ordered.filter((c) => c.durationSeconds === 0).length; @@ -258,6 +284,18 @@ async function main(): Promise<void> { `Selected by scope: ${selected.length} · to run: ${ordered.length}` + (limitCount ? ` (--limit ${limitCount})` : ""), ); + if (deferred.length) { + const hours = deferred.reduce((a, c) => a + c.durationSeconds, 0) / 3600; + console.log( + `DEFERRED by --max-audio-hours ${maxAudioHours}: ${deferred.length} video(s), ` + + `${hours.toFixed(0)} audio-hour(s). They keep their audio and can be run later — ` + + "the cap exists because the clusterer is O(n^2) in segment count and long dense " + + "streams get OOM-killed. Longest deferred:", + ); + for (const c of [...deferred].sort((a, b) => b.durationSeconds - a.durationSeconds).slice(0, 5)) { + console.log(` ${fmtHms(c.durationSeconds)} ${c.slug}`); + } + } console.log( `Audio to process: ${(audioSeconds / 3600).toFixed(0)} audio-hour(s)` + (unknownDuration ? ` (+${unknownDuration} of unknown duration)` : ""), diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -1753,3 +1753,54 @@ Two consequences to expect, both small and both permanent until it is fixed: The fix is to require a known media extension rather than denylisting; it is deferred to its own branch because it changes cleanup behaviour and several exact-string e2e assertions. + +### CORRECTION + escalation: the kernel OOM-killed a diarization run + +**The "~505 MB per audio-hour" figure recorded above is wrong, and the error was +in the method: it was a single `ps` sample taken mid-run, not a peak.** Retract +it as a peak-memory number. + +The measured peak, from `journalctl -k`, on `hasanabi/2834644359` (6h12m): + +``` +Out of memory: Killed process (python) total-vm:34643056kB, anon-rss:11138996kB +``` + +**11.1 GB resident, 34.6 GB virtual — killed by the kernel after 35 minutes of +engine time.** No sidecar was written; nothing in the corpus was damaged; the +backfill continued to the next video. + +**Memory is NOT linear in duration, so duration is a proxy and not the cause.** +A **7h42m** file completed fine, and a **6h12m** file died. The mechanism is +agglomerative clustering holding a pairwise distance matrix over speech-segment +embeddings — **O(n²) in SEGMENT COUNT**. A dense, fast-turnover Twitch VOD yields +far more segments per hour than a monologue or a reaction video, so content +density is what decides it. This also explains the throughput spread already +recorded (120 s/audio-hour on short files vs 373 on a long one). + +Consequences, measured against the remaining work: + +| | | +| --- | --- | +| remaining after video 19 | 67 videos, 216 audio-hours | +| of those **≥ 4h** | **26 videos, 149 audio-hours** | +| of those ≥ 6h | 10 videos, 77 audio-hours | + +So most of the remaining audio sits in the size class that OOM'd once. Each OOM +costs that video's full engine time (~35 min here) and puts an 11 GB transient on +a 16 GB box — the kernel chose python both by luck and by it being the largest +RSS, and that is not a guarantee. + +**Mitigation shipped, not applied:** `--max-audio-hours` on +`bin/diarize-backfill.ts` defers longer videos and **says loudly what it +deferred** (a cap that silently shrinks the work reads as "the corpus is done"). +`--max-audio-hours 4` leaves 45 videos / 74 audio-hours (~6 h at the observed +300 s/audio-hour) and defers 27 / 155. + +It is **deliberately not applied to the running job**: scaling down approved +scope is the operator's call, and the failures are non-destructive. Deferred +videos lose nothing — the cleanup guard holds their audio, so they can be run +later on a bigger box or with a segment-bounded clusterer. + +`diarization.concurrency: 1` is now doubly load-bearing: two 11 GB transients +would not merely be tight, they would be unschedulable.