Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 96b5989805a420125788b41dbd53d490fd90a1e5
parent 361053719d29063cd0f03f1ab523a3d6306161c8
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sat,  8 Aug 2026 04:22:59 -0400

Record what running the backfill measured

Three things the corpus told us that no amount of reading could:

PEAK RSS, which the spike listed as untested and extrapolated at ~230 MB per
audio-hour. A 7h42m file holds 3.9 GB -- ~505 MB/audio-hour, 2.2x the estimate.
That makes diarization.concurrency: 1 load-bearing on a 16 GB box rather than
merely polite: the batch contains an 8h02m VOD, and two in parallel would not
fit.

THROUGHPUT is better than the clean single-file number, not worse. 358
s/audio-hour running average against the 585 measured on one 12.8-minute file,
because short files are far cheaper per audio-hour. ~25h for the 87, not ~42h.

A PRE-EXISTING BUG: isRealAudioFile (videoStatus.ts:63) allow-lists anything
named audio.* minus a denylist, so audio.en-orig.vtt is 'audio', and
audio.live_chat.json.part-Frag114 slips past a .part rule that tests endsWith.
Exactly 3 files corpus-wide, so bounded -- but audioFilesToDelete shares the
predicate, which means the cleanup sweep can delete a subtitle believing it is
audio. Left unfixed on purpose: it changes cleanup behaviour and several
exact-string e2e assertions, so it wants its own branch, not a drive-by during
a 25-hour run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Diffstat:
Mplans/FACTS.md | 65+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
1 file changed, 65 insertions(+), 0 deletions(-)

diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -1652,3 +1652,68 @@ out-competes at 11.2 s/chunk), and the quality is thinner than the headline. The diarized lane remains the one worth having: 1 call per video, acoustic boundaries, and it already produced clean names with calibrated confidences. + +--- + +## Diarization backfill, measured in flight (2026-08-08) + +Running `common/bin/diarize-backfill.ts --scope transcribed` over the 87 +transcribed-with-audio videos (257 audio-hours), `concurrency` 1, `threads` 4. + +### Peak RSS on a genuinely long file — the spike's open item, now measured + +`plans/diarization-spike-results.md` §8 listed this as untested and extrapolated +**~230 MB per audio-hour**. Measured on `HasanAbiVODs3/bCp-EvNxBgs`, a +**7h42m** file: **3.9 GB RSS**, i.e. **~505 MB per audio-hour — 2.2× the +extrapolation.** Memory is released fully between videos (11.4 GB free again +immediately after). + +**This makes `diarization.concurrency: 1` load-bearing, not merely polite.** The +backfill list contains an 8h02m VOD, which at this rate needs ~4.1 GB resident. +Two of those in parallel would not fit alongside anything else on a 16 GB box +that already runs several GB of swap. Raising `concurrency` on this hardware +risks an OOM mid-run, and the failure would land on the longest, most valuable +videos. + +### Throughput is BETTER than the clean single-file measurement + +| | s/audio-hour | +| --- | ---: | +| spike, contended (loadavg ~30) | 1305 | +| clean single-file re-run, 12.8 min file | 585 | +| **backfill running average, 4 videos / 8.2 audio-hours** | **358** | +| the 7h42m file alone | 373 | + +Short files are dramatically cheaper per audio-hour (18 s for a 9-minute file = +~120 s/audio-hour), so the 585 figure — taken from one 12.8-minute file — was +not representative of a mixed batch. Projected wall time for the 87 fell from +~42 h to **~25 h**. Every number here is a rate over its own denominator; none +is a per-video figure. + +### A pre-existing bug found by running this: `isRealAudioFile` accepts non-audio + +`common/lib/videoStatus.ts:63` allow-lists **anything** whose name starts with +`audio.` and then subtracts a denylist. It therefore returns true for files that +are not audio at all. Corpus-wide, exactly **3 files across 3 video dirs**: + +| file | why the denylist misses it | +| --- | --- | +| `audio.en-orig.vtt` | a SUBTITLE track; no rule excludes `.vtt` | +| `audio.live_chat.json.part-Frag114` | the `.part` rule tests `endsWith`, and this ends in `-Frag114` | +| `audio.live_chat.json.part-Frag514` | same | + +Bounded (913 real audio files against 3 fakes), but it is not only cosmetic: + +- `resolveDiarizableMedia` (`controller/diarizeOne.ts`) hands the file to + ffmpeg, so those 3 videos fail the backfill. Fails fast and is visible in the + outcome tally; deliberately NOT special-cased in the backfill script, which + shares the runner's resolver on purpose so the two cannot disagree. +- **`audioFilesToDelete` uses the same predicate**, so the cleanup sweep can + delete `audio.en-orig.vtt` — a subtitle — believing it is audio. That is real, + if small, data loss. +- `readVideoFiles().audioFiles` feeds the cleanup buckets, so "Est. reclaim" + counts these too. + +**Not fixed here.** Tightening the predicate to require a known media extension +changes cleanup behaviour and touches several exact-string e2e assertions; it +wants its own branch and a full e2e run, not a drive-by during a 25-hour job.