commit 96b5989805a420125788b41dbd53d490fd90a1e5
parent 361053719d29063cd0f03f1ab523a3d6306161c8
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Sat, 8 Aug 2026 04:22:59 -0400
Record what running the backfill measured
Three things the corpus told us that no amount of reading could:
PEAK RSS, which the spike listed as untested and extrapolated at ~230 MB per
audio-hour. A 7h42m file holds 3.9 GB -- ~505 MB/audio-hour, 2.2x the estimate.
That makes diarization.concurrency: 1 load-bearing on a 16 GB box rather than
merely polite: the batch contains an 8h02m VOD, and two in parallel would not
fit.
THROUGHPUT is better than the clean single-file number, not worse. 358
s/audio-hour running average against the 585 measured on one 12.8-minute file,
because short files are far cheaper per audio-hour. ~25h for the 87, not ~42h.
A PRE-EXISTING BUG: isRealAudioFile (videoStatus.ts:63) allow-lists anything
named audio.* minus a denylist, so audio.en-orig.vtt is 'audio', and
audio.live_chat.json.part-Frag114 slips past a .part rule that tests endsWith.
Exactly 3 files corpus-wide, so bounded -- but audioFilesToDelete shares the
predicate, which means the cleanup sweep can delete a subtitle believing it is
audio. Left unfixed on purpose: it changes cleanup behaviour and several
exact-string e2e assertions, so it wants its own branch, not a drive-by during
a 25-hour run.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Diffstat:
| M | plans/FACTS.md | | | 65 | +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ |
1 file changed, 65 insertions(+), 0 deletions(-)
diff --git a/plans/FACTS.md b/plans/FACTS.md
@@ -1652,3 +1652,68 @@ out-competes at 11.2 s/chunk), and the quality is thinner than the headline.
The diarized lane remains the one worth having: 1 call per video, acoustic
boundaries, and it already produced clean names with calibrated confidences.
+
+---
+
+## Diarization backfill, measured in flight (2026-08-08)
+
+Running `common/bin/diarize-backfill.ts --scope transcribed` over the 87
+transcribed-with-audio videos (257 audio-hours), `concurrency` 1, `threads` 4.
+
+### Peak RSS on a genuinely long file — the spike's open item, now measured
+
+`plans/diarization-spike-results.md` §8 listed this as untested and extrapolated
+**~230 MB per audio-hour**. Measured on `HasanAbiVODs3/bCp-EvNxBgs`, a
+**7h42m** file: **3.9 GB RSS**, i.e. **~505 MB per audio-hour — 2.2× the
+extrapolation.** Memory is released fully between videos (11.4 GB free again
+immediately after).
+
+**This makes `diarization.concurrency: 1` load-bearing, not merely polite.** The
+backfill list contains an 8h02m VOD, which at this rate needs ~4.1 GB resident.
+Two of those in parallel would not fit alongside anything else on a 16 GB box
+that already runs several GB of swap. Raising `concurrency` on this hardware
+risks an OOM mid-run, and the failure would land on the longest, most valuable
+videos.
+
+### Throughput is BETTER than the clean single-file measurement
+
+| | s/audio-hour |
+| --- | ---: |
+| spike, contended (loadavg ~30) | 1305 |
+| clean single-file re-run, 12.8 min file | 585 |
+| **backfill running average, 4 videos / 8.2 audio-hours** | **358** |
+| the 7h42m file alone | 373 |
+
+Short files are dramatically cheaper per audio-hour (18 s for a 9-minute file =
+~120 s/audio-hour), so the 585 figure — taken from one 12.8-minute file — was
+not representative of a mixed batch. Projected wall time for the 87 fell from
+~42 h to **~25 h**. Every number here is a rate over its own denominator; none
+is a per-video figure.
+
+### A pre-existing bug found by running this: `isRealAudioFile` accepts non-audio
+
+`common/lib/videoStatus.ts:63` allow-lists **anything** whose name starts with
+`audio.` and then subtracts a denylist. It therefore returns true for files that
+are not audio at all. Corpus-wide, exactly **3 files across 3 video dirs**:
+
+| file | why the denylist misses it |
+| --- | --- |
+| `audio.en-orig.vtt` | a SUBTITLE track; no rule excludes `.vtt` |
+| `audio.live_chat.json.part-Frag114` | the `.part` rule tests `endsWith`, and this ends in `-Frag114` |
+| `audio.live_chat.json.part-Frag514` | same |
+
+Bounded (913 real audio files against 3 fakes), but it is not only cosmetic:
+
+- `resolveDiarizableMedia` (`controller/diarizeOne.ts`) hands the file to
+ ffmpeg, so those 3 videos fail the backfill. Fails fast and is visible in the
+ outcome tally; deliberately NOT special-cased in the backfill script, which
+ shares the runner's resolver on purpose so the two cannot disagree.
+- **`audioFilesToDelete` uses the same predicate**, so the cleanup sweep can
+ delete `audio.en-orig.vtt` — a subtitle — believing it is audio. That is real,
+ if small, data loss.
+- `readVideoFiles().audioFiles` feeds the cleanup buckets, so "Est. reclaim"
+ counts these too.
+
+**Not fixed here.** Tightening the predicate to require a known media extension
+changes cleanup behaviour and touches several exact-string e2e assertions; it
+wants its own branch and a full e2e run, not a drive-by during a 25-hour job.