commit 94ac368b020149163edbc855d46d15ec054f33af
parent 6af4a7ffa092a38f1f29e4bee4de91b493fa85cc
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Sat, 8 Aug 2026 22:42:39 -0400
Ask the shape of a filename, not a list of exceptions
isRealAudioFile/isSourceMediaFile/isPartAudioFile allow-listed the `audio.` /
`source-media.` PREFIX and subtracted a denylist. A denylist is open-ended and
this one had fallen behind by at least four real corpus files:
audio.en-orig.vtt (a SUBTITLE -- no rule excluded .vtt),
audio.live_chat.json.part-Frag114 (the .part rule used endsWith), and on the
container side source-media.temp.mp4, which no rule touched at all.
NOT COSMETIC IN EITHER DIRECTION. audioFilesToRemove shares the predicate, so
the cleanup sweep could delete a subtitle as audio -- an unguarded remove(),
no trash. And findSourceMedia's answer causes a WRITE: finalizeAppExtraction
transcodes from it and then MOVES it into the saved-video store as the
permanent copy, so a dir holding both source-media.mp4 and its .temp scratch
could archive the corrupt one and orphan the real container.
The fix is an ANCHORED SINGLE-SEGMENT match over AUDIO_EXTS u
VIDEO_CONTAINER_EXTS, not merely an extension allowlist. That distinction is
the whole point: this repo's own transcode temp is audio.tmp-<pid>.<fmt>
(controller/transcode.ts, observed live as audio.tmp-2760235.mp3) and carries a
perfectly valid audio extension, so an allowlist alone would re-admit it and
race the running transcoder. One anchored rule rejects it along with the other
four. The ext list is built from what the corpus ACTUALLY holds (900 mp3,
12 mp4, 1 aac) rather than from AudioFormat, which is the transcode TARGET enum
and would have dropped both mp4 and aac.
Safety is provable, not hoped for: audioFilesToRemove is a filter optionally
intersected with a second filter, both monotone in the predicate, so a strictly
narrower predicate yields a pointwise subset -- the delete set can only shrink.
No call site negates these predicates, so nothing can invert that direction.
Three more sites carried the same bug and would NOT have been fixed by
videoStatus.ts alone: cleanExtraAudioFormats and transcodeFailures each had
their own inline, staler copy of the denylist (the first an unguarded deleter,
the second feeding ffmpeg a source file), and audioCheckedDownload matched
audio.*.part, which audio.live_chat.json.part satisfies. All now share the one
predicate.
New lib/mediaFiles.ts holds the pure string rules; videoStatus.ts re-exports
them so no import site changed. The split is what lets the editor's VideoPanel
-- a client component -- share the extension lists it had been duplicating,
without dragging node:fs into the browser bundle. AUDIO_PREFERENCE was
triplicated with two different orderings; both are kept, now named for what
they mean (read vs transcode-source) so the difference reads as deliberate.
Also B2: findTurns discarded out-of-range marks with no count and no warning.
Round 3 measured 520 of 1,705 turns (30%) dropped that way. Counted per chunk
now, with a rate in the log. Observability only -- the lane stays unarmed.
Verified: common 605/605 (586 + 19 new in lib/mediaFiles.test.ts, seeded with
the real corpus filenames plus audio.tmp-2760235.mp3, audio.m4a.part.good,
audio.webm and source-media.temp.mp4); tsc clean in common and editor.
Editor e2e still to run -- B1's blast radius is wider than any single spec.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Diffstat:
13 files changed, 518 insertions(+), 181 deletions(-)
diff --git a/common/controller/attributeOne.ts b/common/controller/attributeOne.ts
@@ -438,6 +438,11 @@ async function findTurns(args: {
let reportedModel = "";
let costUsd = 0;
let chunksOk = 0;
+ // Turns thrown away for landing outside the chunk that produced them, and the
+ // total emitted, so the summary can report a RATE rather than a bare count —
+ // 520 discards means nothing without the 1,705 it is out of.
+ let outOfRangeTotal = 0;
+ let emittedTotal = 0;
for (let i = 0; i < chunks.length; i++) {
args.signal?.throwIfAborted();
@@ -472,6 +477,8 @@ async function findTurns(args: {
const raw = (result.data as { turns?: unknown })?.turns;
if (!Array.isArray(raw)) continue;
+ let outOfRange = 0;
+ emittedTotal += raw.length;
for (const item of raw) {
const e = item as { start?: unknown; speaker?: unknown };
if (typeof e?.start !== "string" || typeof e.speaker !== "string") continue;
@@ -488,11 +495,26 @@ async function findTurns(args: {
}
// Out of the chunk's own range is wrong even when the topic is real —
// clamped away rather than trusted, since a mark in another chunk's
- // territory would fight that chunk's own answer.
- if (at < startSeconds || at > endSeconds) continue;
+ // territory would fight that chunk's own answer. COUNTED, because the
+ // drop rate is a first-class quality signal: see AttributionWarning's
+ // "out-of-range".
+ if (at < startSeconds || at > endSeconds) {
+ outOfRange++;
+ continue;
+ }
if (isUselessSpeakerLabel(e.speaker)) continue;
marks.push({ at, speaker: roster.intern(e.speaker) });
}
+ // One warning per chunk carrying the count, not one per discarded turn: a
+ // 30% drop rate would otherwise bury every other warning in the record.
+ if (outOfRange > 0) {
+ outOfRangeTotal += outOfRange;
+ args.warnings.push({
+ code: "out-of-range",
+ chunk: i,
+ detail: `${outOfRange} of ${raw.length} turn(s) outside ${startSeconds}-${endSeconds}s`,
+ });
+ }
} catch (err) {
if (args.signal?.aborted) throw err;
const message = (err as Error)?.message ?? String(err);
@@ -507,6 +529,13 @@ async function findTurns(args: {
}
}
+ if (outOfRangeTotal > 0) {
+ const pct = emittedTotal > 0 ? (outOfRangeTotal / emittedTotal) * 100 : 0;
+ args.log(
+ `Attribute ${args.videoId}: discarded ${outOfRangeTotal} of ${emittedTotal} emitted turn(s) (${pct.toFixed(0)}%) for landing outside their own chunk.`,
+ );
+ }
+
if (labels.length === 0) {
return { speakers: [], segments: [], model: reportedModel, costUsd, chunks: chunks.length, chunksOk };
}
diff --git a/common/controller/cleanExtraAudioFormats.ts b/common/controller/cleanExtraAudioFormats.ts
@@ -2,6 +2,7 @@ import path from "node:path";
import fs from "fs-extra";
import type { Paths } from "../lib/paths";
import { isDoNotClean } from "../lib/doNotClean-server";
+import { audioFilesToRemove } from "../lib/videoStatus";
import { readChannelConfig } from "./channels";
const { pathExists, readdir, remove } = fs;
@@ -54,14 +55,16 @@ export async function cleanExtraAudioFormats({
const videoDir = path.join(dataDir, id);
const entries = await readdir(videoDir).catch(() => [] as string[]);
if (!entries.includes(targetAudioFile)) continue;
- const extras = entries.filter(
- (e) =>
- e.startsWith("audio.") &&
- e !== targetAudioFile &&
- !e.includes(".tmp-") &&
- !e.endsWith(".info.json") &&
- !e.endsWith(".part"),
- );
+ // Shares the app's one media-file predicate rather than re-deriving it. The
+ // inline copy that used to live here was a stale fork of the old denylist:
+ // it excluded .tmp-/.info.json/.part and nothing else, so an
+ // audio.en-orig.vtt subtitle or an audio.live_chat.json.part-Frag114
+ // fragment sitting next to the target format was deleted as an "extra audio
+ // format" — by an unguarded remove(), in a sweep with no dry run.
+ const extras = audioFilesToRemove(entries, {
+ targetAudioFile,
+ wrongFormatOnly: true,
+ });
if (extras.length === 0) continue;
if (await isDoNotClean(videoDir)) {
log(`Skipped ${id} (marked do not clean)`);
diff --git a/common/controller/diarizeOne.ts b/common/controller/diarizeOne.ts
@@ -23,12 +23,11 @@ import {
} from "../lib/diarization";
import { hasDiarization, loadDiarization } from "../lib/diarization-server";
import { findSourceMedia, isRealAudioFile } from "../lib/videoStatus";
+import { pickPreferredAudio } from "../lib/mediaFiles";
import { resolveSavedVideo } from "../lib/savedVideo-server";
const { readdir } = fs;
-const AUDIO_PREFERENCE = ["audio.mp3", "audio.m4a", "audio.opus"];
-
export type DiarizeOneOptions = {
paths: Paths;
videoDir: string;
@@ -66,13 +65,8 @@ export async function resolveDiarizableMedia(
videoDir: string,
): Promise<string | null> {
const entries = await readdir(videoDir).catch(() => [] as string[]);
- const candidates = entries.filter(isRealAudioFile);
- if (candidates.length > 0) {
- for (const preferred of AUDIO_PREFERENCE) {
- if (candidates.includes(preferred)) return preferred;
- }
- return [...candidates].sort()[0];
- }
+ const preferred = pickPreferredAudio(entries.filter(isRealAudioFile));
+ if (preferred) return preferred;
const container = findSourceMedia(entries);
if (container) return container;
const saved = await resolveSavedVideo(videoDir);
diff --git a/common/controller/transcodeFailures.ts b/common/controller/transcodeFailures.ts
@@ -3,6 +3,11 @@ import fs from "fs-extra";
import pLimit from "p-limit";
import type { Paths } from "../lib/paths";
import type { AudioFormat } from "../lib/channelConfig";
+import {
+ AUDIO_TRANSCODE_SOURCE_PREFERENCE,
+ audioFilesToRemove,
+ pickPreferredAudio,
+} from "../lib/mediaFiles";
import { transcodeAudio } from "./transcode";
import {
failedTranscodingsFile,
@@ -11,28 +16,18 @@ import {
const { pathExists, readdir, readFile, appendFile, ensureFile } = fs;
-const AUDIO_PREFERENCE = ["audio.m4a", "audio.opus", "audio.mp3"];
-
function pickSourceAudio(
entries: string[],
targetFormat: AudioFormat,
): string | null {
- const target = `audio.${targetFormat}`;
- const candidates = entries.filter(
- (e) =>
- e.startsWith("audio.") &&
- e !== target &&
- !e.includes(".tmp-") &&
- !e.endsWith(".info.json") &&
- !e.endsWith(".part"),
- );
- if (candidates.length === 0) return null;
- for (const preferred of AUDIO_PREFERENCE) {
- if (preferred !== target && candidates.includes(preferred)) {
- return preferred;
- }
- }
- return [...candidates].sort()[0];
+ // Shares the app's media-file predicate. The inline filter that used to live
+ // here was another copy of the old denylist, so it could hand ffmpeg an
+ // audio.en-orig.vtt subtitle as a "source audio" to transcode from.
+ const candidates = audioFilesToRemove(entries, {
+ targetAudioFile: `audio.${targetFormat}`,
+ wrongFormatOnly: true,
+ });
+ return pickPreferredAudio(candidates, AUDIO_TRANSCODE_SOURCE_PREFERENCE);
}
export type TranscodeBatchResult = {
diff --git a/common/controller/transcribeOne.ts b/common/controller/transcribeOne.ts
@@ -13,13 +13,12 @@ import { getSettings } from "../lib/settings";
import { pingRemoteHealth, transcribeViaRemote } from "./remoteTranscribe";
import { TranscribeError } from "./transcribeError";
import { findSourceMedia, isRealAudioFile } from "../lib/videoStatus";
+import { pickPreferredAudio } from "../lib/mediaFiles";
import { resolveSavedVideo } from "../lib/savedVideo-server";
import { writeTranscribeOutcome } from "../lib/transcribeOutcome-server";
const { pathExists, readdir, rename, writeFile } = fs;
-const AUDIO_PREFERENCE = ["audio.mp3", "audio.m4a", "audio.opus"];
-
async function resolveAudioFile(
videoDir: string,
requested: string,
@@ -28,13 +27,8 @@ async function resolveAudioFile(
if (await pathExists(path.join(videoDir, requested))) return requested;
if (strict) return null;
const entries = await readdir(videoDir);
- const candidates = entries.filter(isRealAudioFile);
- if (candidates.length > 0) {
- for (const preferred of AUDIO_PREFERENCE) {
- if (candidates.includes(preferred)) return preferred;
- }
- return [...candidates].sort()[0];
- }
+ const preferred = pickPreferredAudio(entries.filter(isRealAudioFile));
+ if (preferred) return preferred;
// No extracted audio on disk — fall back to a persisted source video
// container. Transcribers that ffmpeg-slice their input (parakeet) read a
// container directly; this is what makes a kept-but-cleaned video, or a
diff --git a/common/lib/attribution.ts b/common/lib/attribution.ts
@@ -66,7 +66,20 @@ export type AttributionSegment = {
// outcome a generator has must leave a trace somewhere other than a job log
// that rotates.
export type AttributionWarning = {
- code: "chunk-failed" | "bad-timestamp" | "unknown-cluster" | "empty";
+ code:
+ | "chunk-failed"
+ | "bad-timestamp"
+ | "unknown-cluster"
+ | "empty"
+ // The text-only lane emitted a turn whose timestamp lands outside the chunk
+ // that produced it, so the turn was discarded. Counted per chunk, not per
+ // turn, because it is not rare: the bake-off's third round measured 520 of
+ // 1,705 turns (30%) dropped this way, and the shipped findTurns was dropping
+ // them on the same condition with no count and no log line — invisible in
+ // rounds 1 and 2 and in every record on disk. A discard rate that high is
+ // the difference between "the model is doing well" and "the model is
+ // producing four usable turns in five", so it has to leave a trace.
+ | "out-of-range";
chunk?: number;
detail?: string;
};
diff --git a/common/lib/mediaFiles.test.ts b/common/lib/mediaFiles.test.ts
@@ -0,0 +1,204 @@
+import { describe, it } from "node:test";
+import assert from "node:assert/strict";
+import {
+ AUDIO_READ_PREFERENCE,
+ AUDIO_TRANSCODE_SOURCE_PREFERENCE,
+ audioFilesToRemove,
+ findSourceMedia,
+ isPartAudioFile,
+ isRealAudioFile,
+ isSourceMediaFile,
+ isVideoContainer,
+ pickPreferredAudio,
+} from "./mediaFiles";
+
+// These predicates decide what a DELETE sweep removes and what ffmpeg is handed
+// as input, and they had never had a test. The fixtures below are not invented:
+// every "junk" name is a real file surveyed out of the 78,963-video corpus, and
+// audio.tmp-2760235.mp3 was observed live in a video dir while a transcode was
+// running.
+
+// Real, finalized audio in the corpus today (900 / 12 / 1 by count).
+const REAL_AUDIO = ["audio.mp3", "audio.mp4", "audio.aac"];
+
+// The four junk names that the old prefix-plus-denylist admitted as audio, plus
+// this app's own transcode scratch, which an extension allowlist alone would
+// re-admit (its extension is perfectly valid) and which the cleanup sweep would
+// then race.
+const NOT_AUDIO = [
+ "audio.en-orig.vtt", // a SUBTITLE. No denylist rule excluded .vtt.
+ "audio.live_chat.json.part-Frag114", // the .part rule used endsWith
+ "audio.live_chat.json.part-Frag514",
+ "audio.live_chat.json.part",
+ "audio.live_chat.json",
+ "audio.info.json",
+ "audio.tmp-2760235.mp3", // controller/transcode.ts scratch, mid-transcode
+ "audio.m4a.part", // an unfinished download, not a finalized file
+ "audio.m4a.part.good", // audio-check snapshot
+ "audio.m4a.part.testing",
+ "audio.f251.webm.part", // yt-dlp per-format fragment
+ "transcript.en.vtt",
+ "metadata.info.json",
+ "audio.", // degenerate
+ "audio",
+];
+
+describe("isRealAudioFile", () => {
+ it("accepts finalized audio.<ext> for every format on disk", () => {
+ for (const name of REAL_AUDIO) {
+ assert.equal(isRealAudioFile(name), true, name);
+ }
+ });
+
+ it("accepts wrong-format containers, because those are cleanup targets", () => {
+ // bulk-actions.spec.ts renames audio.m4a -> audio.webm and requires the
+ // sweep to delete it, so webm/mp4 must stay in the list even though they
+ // are containers rather than audio-only formats.
+ for (const name of ["audio.webm", "audio.mkv", "audio.m4a", "audio.opus"]) {
+ assert.equal(isRealAudioFile(name), true, name);
+ }
+ });
+
+ it("rejects every non-media name that used to slip through", () => {
+ for (const name of NOT_AUDIO) {
+ assert.equal(isRealAudioFile(name), false, name);
+ }
+ });
+
+ it("is case-insensitive on the extension", () => {
+ assert.equal(isRealAudioFile("audio.MP3"), true);
+ });
+});
+
+describe("isSourceMediaFile / findSourceMedia", () => {
+ it("accepts a finalized container", () => {
+ assert.equal(isSourceMediaFile("source-media.mp4"), true);
+ assert.equal(isSourceMediaFile("source-media.mkv"), true);
+ });
+
+ it("rejects yt-dlp's postprocessor scratch", () => {
+ // The ONLY source-media.* file in the whole corpus is this one, and it is
+ // the corrupt 4 GiB truncated container. The old denylist had no .temp.
+ // rule, so it read as finalized media — and finalizeAppExtraction would
+ // transcode from it and then move it into the saved-video store as the
+ // permanent copy.
+ assert.equal(isSourceMediaFile("source-media.temp.mp4"), false);
+ assert.equal(isSourceMediaFile("source-media.f137.mp4"), false);
+ assert.equal(isSourceMediaFile("source-media.mp4.part"), false);
+ assert.equal(isSourceMediaFile("source-media.info.json"), false);
+ });
+
+ it("never picks the scratch file over the real container", () => {
+ assert.equal(
+ findSourceMedia(["source-media.temp.mp4", "source-media.mp4"]),
+ "source-media.mp4",
+ );
+ // Same answer whichever order readdir happened to return them in.
+ assert.equal(
+ findSourceMedia(["source-media.mp4", "source-media.temp.mp4"]),
+ "source-media.mp4",
+ );
+ });
+
+ it("returns null when only scratch is present", () => {
+ assert.equal(findSourceMedia(["source-media.temp.mp4"]), null);
+ });
+});
+
+describe("isPartAudioFile", () => {
+ it("accepts a genuine interrupted audio download", () => {
+ assert.equal(isPartAudioFile("audio.m4a.part"), true);
+ assert.equal(isPartAudioFile("audio.webm.part"), true);
+ });
+
+ it("rejects live-chat sidecars and audio-check snapshots", () => {
+ for (const name of [
+ "audio.live_chat.json.part",
+ "audio.live_chat.json.part-Frag114",
+ "audio.m4a.part.good",
+ "audio.m4a.part.testing",
+ "audio.tmp-2760235.mp3.part",
+ "audio.mp3",
+ ]) {
+ assert.equal(isPartAudioFile(name), false, name);
+ }
+ });
+});
+
+describe("audioFilesToRemove", () => {
+ // The delete set. A real video dir mid-transcode, with a subtitle leak and a
+ // live-chat fragment sitting next to the audio.
+ const DIR = [
+ "audio.mp3",
+ "audio.m4a",
+ "audio.en-orig.vtt",
+ "audio.live_chat.json.part-Frag114",
+ "audio.info.json",
+ "audio.tmp-2760235.mp3",
+ "metadata.info.json",
+ "transcript.en.vtt",
+ ];
+
+ it("removes only the two real audio files", () => {
+ assert.deepEqual(audioFilesToRemove(DIR), ["audio.mp3", "audio.m4a"]);
+ });
+
+ it("keeps the channel's target format under wrongFormatOnly", () => {
+ assert.deepEqual(
+ audioFilesToRemove(DIR, {
+ targetAudioFile: "audio.mp3",
+ wrongFormatOnly: true,
+ }),
+ ["audio.m4a"],
+ );
+ });
+
+ it("is a subset of the raw entries for any input (monotonicity)", () => {
+ const out = audioFilesToRemove(DIR);
+ for (const name of out) assert.ok(DIR.includes(name), name);
+ });
+});
+
+describe("isVideoContainer", () => {
+ it("keys off the extension, not the base name", () => {
+ assert.equal(isVideoContainer("source-media.mkv"), true);
+ assert.equal(isVideoContainer("audio.mp4"), true);
+ assert.equal(isVideoContainer("audio.mp3"), false);
+ assert.equal(isVideoContainer("noext"), false);
+ });
+});
+
+describe("pickPreferredAudio", () => {
+ it("prefers mp3 when reading", () => {
+ assert.equal(
+ pickPreferredAudio(["audio.opus", "audio.mp3", "audio.m4a"]),
+ "audio.mp3",
+ );
+ });
+
+ it("prefers m4a when picking a transcode source", () => {
+ assert.equal(
+ pickPreferredAudio(
+ ["audio.opus", "audio.mp3", "audio.m4a"],
+ AUDIO_TRANSCODE_SOURCE_PREFERENCE,
+ ),
+ "audio.m4a",
+ );
+ });
+
+ it("falls back to a stable sort, not readdir order", () => {
+ assert.equal(pickPreferredAudio(["audio.wav", "audio.aac"]), "audio.aac");
+ assert.equal(pickPreferredAudio(["audio.aac", "audio.wav"]), "audio.aac");
+ });
+
+ it("returns null for no candidates", () => {
+ assert.equal(pickPreferredAudio([]), null);
+ });
+
+ it("keeps the two orderings distinct", () => {
+ assert.notDeepEqual(
+ [...AUDIO_READ_PREFERENCE],
+ [...AUDIO_TRANSCODE_SOURCE_PREFERENCE],
+ );
+ });
+});
diff --git a/common/lib/mediaFiles.ts b/common/lib/mediaFiles.ts
@@ -0,0 +1,183 @@
+// Media file naming — the shapes this app writes into a video dir, and the
+// predicates that recognize them.
+//
+// SPLIT OUT OF lib/videoStatus.ts because these are pure string rules with no
+// filesystem in them, and videoStatus imports node:fs. The editor's VideoPanel
+// is a client component and needs the same extension lists it used to keep a
+// private copy of; importing videoStatus there would drag node:fs into the
+// browser bundle. lib/videoStatus.ts re-exports everything here, so every
+// existing import site is unchanged.
+
+// VideoPanel's "can I transcribe this here?" question is audio-only: audio.mp4
+// is a wrong-format cleanup target, not something to offer a transcribe button
+// for (editor/e2e/video-page.spec.ts pins that).
+export const AUDIO_EXTS: ReadonlyArray<string> = [
+ "mp3",
+ "m4a",
+ "aac",
+ "ogg",
+ "opus",
+ "wav",
+ "flac",
+];
+
+// Video container extensions the app may keep as a persisted source video and
+// transcribe directly from (parakeet's stitcher ffmpeg-slices any container).
+// Consolidated here so VideoPanel and the transcribe fallback share one list.
+export const VIDEO_CONTAINER_EXTS: ReadonlyArray<string> = [
+ "mp4",
+ "webm",
+ "mkv",
+ "mov",
+ "m4v",
+ "ogv",
+ "avi",
+];
+
+// Everything ffmpeg may hand us as a media file under one of the two managed
+// base names. The union is deliberate and not audio alone: a failed
+// extract-to-mp3 leaves audio.mp4 / audio.webm behind (12 audio.mp4 in this
+// corpus), and those are exactly what the wrong-format sweep exists to remove.
+export const MEDIA_EXTS: ReadonlyArray<string> = [
+ ...AUDIO_EXTS,
+ ...VIDEO_CONTAINER_EXTS,
+];
+
+// A persisted source video container is written under data/<id>/source-media.<ext>
+// (a deliberately distinct base name from audio.<ext> so it's never mistaken for
+// an extractable/cleanable audio file by isRealAudioFile). The keep-latest
+// persistence rule downloads the full video here and the app extracts audio from
+// it; the container is then kept (Phase 3 moves it to the saved-video store).
+export const SOURCE_MEDIA_BASENAME = "source-media";
+
+const MEDIA_EXT_ALT = MEDIA_EXTS.join("|");
+
+// THE SHAPE IS THE RULE, NOT A DENYLIST. The app writes exactly one finalized
+// media file per role per video dir, and it is always a SINGLE segment after the
+// base name: audio.<ext> and source-media.<ext>. So these predicates anchor on
+// that shape instead of allow-listing a prefix and subtracting known-bad
+// suffixes.
+//
+// The denylist this replaces was demonstrably losing, and not cosmetically:
+// audio.en-orig.vtt (a SUBTITLE — no rule excluded .vtt) and
+// audio.live_chat.json.part-Frag114 (the .part rule used endsWith) both passed
+// it, and audioFilesToRemove shares this predicate, so the cleanup sweep could
+// delete a subtitle believing it was audio. On the source-media side the
+// denylist also missed .temp., so yt-dlp's postprocessor scratch
+// source-media.temp.mp4 read as a finalized container — and that one causes a
+// WRITE (ytdlp/downloadOneManaged.ts finalizeAppExtraction transcodes from
+// whatever findSourceMedia returns and then moves it into the saved-video
+// store), so it was an archive-corruption path.
+//
+// An extension allowlist ALONE would not be enough: this repo's own transcode
+// scratch is audio.tmp-<pid>.<targetFormat> (controller/transcode.ts) — a
+// perfectly valid audio extension — and admitting it would race the running
+// transcoder. Anchoring to one segment rejects audio.tmp-2760235.mp3,
+// audio.m4a.part.good, audio.live_chat.json.part-Frag114 and audio.en-orig.vtt
+// with a single rule.
+const REAL_AUDIO_RE = new RegExp(`^audio\\.(?:${MEDIA_EXT_ALT})$`, "i");
+const SOURCE_MEDIA_RE = new RegExp(
+ `^${SOURCE_MEDIA_BASENAME}\\.(?:${MEDIA_EXT_ALT})$`,
+ "i",
+);
+const PART_AUDIO_RE = new RegExp(`^audio\\.(?:${MEDIA_EXT_ALT})\\.part$`, "i");
+
+// A finalized media output under data/<id>/audio.<ext>.
+export function isRealAudioFile(name: string): boolean {
+ return REAL_AUDIO_RE.test(name);
+}
+
+// The finalized audio files to delete from a video dir. Operates on the raw
+// readdir() entries: keeps only real audio (excludes .part partials, snapshots,
+// sidecars, temp files via isRealAudioFile). With wrongFormatOnly the channel's
+// target audio file is preserved and only off-target formats are returned —
+// this is how "remove wrong-format audio" strips the cornbreadman-style
+// audio.m4a/audio.mp4 leftovers that failed yt-dlp's extract-to-mp3 step.
+//
+// THIS IS THE ONE DELETE SET, and it is why narrowing isRealAudioFile is safe
+// rather than merely hoped to be: the body is a filter optionally intersected
+// with a second filter, and both are monotone in the predicate, so a strictly
+// narrower predicate yields a pointwise SUBSET for every input. The delete set
+// can only shrink. (No call site in common/, editor/ or export/ negates these
+// predicates, so there is no branch that could invert that direction.)
+export function audioFilesToRemove(
+ entries: string[],
+ opts: { targetAudioFile?: string; wrongFormatOnly?: boolean } = {},
+): string[] {
+ let files = entries.filter(isRealAudioFile);
+ if (opts.wrongFormatOnly) {
+ files = files.filter((name) => name !== opts.targetAudioFile);
+ }
+ return files;
+}
+
+export function isVideoContainer(name: string): boolean {
+ const dot = name.lastIndexOf(".");
+ if (dot < 0) return false;
+ return VIDEO_CONTAINER_EXTS.includes(name.slice(dot + 1).toLowerCase());
+}
+
+// A finalized persisted source container under data/<id>/source-media.<ext>.
+export function isSourceMediaFile(name: string): boolean {
+ return SOURCE_MEDIA_RE.test(name);
+}
+
+// The persisted source container in a video dir, if any (the keep-latest source
+// media). Null when no finalized source-media.<ext> exists.
+//
+// SORTED, not readdir order: finalizeAppExtraction transcodes from whatever this
+// returns and may then MOVE it into the saved-video store as the permanent copy,
+// so "whichever the filesystem happened to list first" is not an acceptable way
+// to choose between two containers. Sorting makes the choice reproducible; the
+// anchored predicate above is what keeps scratch files out of the running.
+export function findSourceMedia(entries: string[]): string | null {
+ const found = entries.filter(isSourceMediaFile).sort();
+ return found[0] ?? null;
+}
+
+// A genuine resumable partial: an interrupted audio download, NOT a live-chat
+// sidecar that merely ends in .part, and not an audio-check snapshot
+// (.part.good/.part.testing).
+export function isPartAudioFile(name: string): boolean {
+ return PART_AUDIO_RE.test(name);
+}
+
+// ---------------------------------------------------------------------------
+// Which audio file to use when a video dir holds more than one.
+//
+// This ordering was copy-pasted into three controllers (diarizeOne,
+// transcribeOne, transcodeFailures) with TWO DIFFERENT ORDERS, which reads like
+// drift but is not: the first two are picking something to READ and the third is
+// picking something to TRANSCODE FROM. Both orders are kept, named for what they
+// mean, so the difference is visible instead of looking like a bug.
+
+// Reading: mp3 first, because that is what this corpus overwhelmingly holds
+// (900 audio.mp3 against 12 audio.mp4 and 1 audio.aac) and what the transcribe
+// and diarize paths were tuned against.
+export const AUDIO_READ_PREFERENCE: ReadonlyArray<string> = [
+ "audio.mp3",
+ "audio.m4a",
+ "audio.opus",
+];
+
+// Transcoding: prefer the format most likely to be the ORIGINAL download rather
+// than a previous re-encode, so a repeated transcode does not stack generations
+// of lossy loss.
+export const AUDIO_TRANSCODE_SOURCE_PREFERENCE: ReadonlyArray<string> = [
+ "audio.m4a",
+ "audio.opus",
+ "audio.mp3",
+];
+
+// Pick one audio file from a set of candidates by preference, falling back to a
+// stable sort so the answer never depends on readdir order.
+export function pickPreferredAudio(
+ candidates: string[],
+ preference: ReadonlyArray<string> = AUDIO_READ_PREFERENCE,
+): string | null {
+ if (candidates.length === 0) return null;
+ for (const preferred of preference) {
+ if (candidates.includes(preferred)) return preferred;
+ }
+ return [...candidates].sort()[0] ?? null;
+}
diff --git a/common/lib/videoStatus.ts b/common/lib/videoStatus.ts
@@ -1,5 +1,6 @@
import path from "node:path";
import { readdir, readFile, stat } from "node:fs/promises";
+import { isPartAudioFile, isRealAudioFile } from "./mediaFiles";
export type VideoFiles = {
hasMeta: boolean;
@@ -54,93 +55,22 @@ export const META_FILENAME = "metadata.info.json";
// lib/diarization.ts for the record shape.
export const DIARIZATION_FILENAME = "diarization.json";
-// A finalized media output under data/<id>/audio.<ext>. Excludes yt-dlp partials
-// (.part) and audio-check snapshots (.part.good/.part.testing), the metadata
-// sidecar (*.info.json), temp work files (.tmp-...), and the live-chat sidecar
-// (audio.live_chat.json[.part]) — live_chat ignores the `subtitle:` output prefix
-// and lands under the default audio.%(ext)s template, which is the source of the
-// "no audio file found" whisper crash.
-export function isRealAudioFile(name: string): boolean {
- if (!name.startsWith("audio.")) return false;
- if (name.includes(".tmp-")) return false;
- if (name.endsWith(".info.json")) return false;
- if (name.endsWith(".live_chat.json")) return false;
- if (name.endsWith(".part")) return false; // covers .live_chat.json.part too
- if (name.endsWith(".part.good")) return false;
- if (name.endsWith(".part.testing")) return false;
- return true;
-}
-
-// The finalized audio files to delete from a video dir. Operates on the raw
-// readdir() entries: keeps only real audio (excludes .part partials, snapshots,
-// sidecars, temp files via isRealAudioFile). With wrongFormatOnly the channel's
-// target audio file is preserved and only off-target formats are returned —
-// this is how "remove wrong-format audio" strips the cornbreadman-style
-// audio.m4a/audio.mp4 leftovers that failed yt-dlp's extract-to-mp3 step.
-export function audioFilesToRemove(
- entries: string[],
- opts: { targetAudioFile?: string; wrongFormatOnly?: boolean } = {},
-): string[] {
- let files = entries.filter(isRealAudioFile);
- if (opts.wrongFormatOnly) {
- files = files.filter((name) => name !== opts.targetAudioFile);
- }
- return files;
-}
-
-// Video container extensions the app may keep as a persisted source video and
-// transcribe directly from (parakeet's stitcher ffmpeg-slices any container).
-// Consolidated here so VideoPanel and the transcribe fallback share one list.
-export const VIDEO_CONTAINER_EXTS: ReadonlyArray<string> = [
- "mp4",
- "webm",
- "mkv",
- "mov",
- "m4v",
- "ogv",
- "avi",
-];
-
-export function isVideoContainer(name: string): boolean {
- const dot = name.lastIndexOf(".");
- if (dot < 0) return false;
- return VIDEO_CONTAINER_EXTS.includes(name.slice(dot + 1).toLowerCase());
-}
-
-// A persisted source video container is written under data/<id>/source-media.<ext>
-// (a deliberately distinct base name from audio.<ext> so it's never mistaken for
-// an extractable/cleanable audio file by isRealAudioFile). The keep-latest
-// persistence rule downloads the full video here and the app extracts audio from
-// it; the container is then kept (Phase 3 moves it to the saved-video store).
-export const SOURCE_MEDIA_BASENAME = "source-media";
-
-export function isSourceMediaFile(name: string): boolean {
- if (!name.startsWith(`${SOURCE_MEDIA_BASENAME}.`)) return false;
- if (name.includes(".tmp-")) return false;
- if (name.endsWith(".info.json")) return false;
- if (name.endsWith(".part")) return false;
- if (name.endsWith(".part.good")) return false;
- if (name.endsWith(".part.testing")) return false;
- return true;
-}
-
-// The persisted source container in a video dir, if any (the keep-latest source
-// media). Null when no finalized source-media.<ext> exists.
-export function findSourceMedia(entries: string[]): string | null {
- return entries.find(isSourceMediaFile) ?? null;
-}
-
-// A genuine resumable partial: an interrupted audio download, NOT a live-chat
-// sidecar that merely ends in .part.
-export function isPartAudioFile(name: string): boolean {
- if (!name.startsWith("audio.")) return false;
- if (name.includes(".tmp-")) return false;
- if (name.endsWith(".live_chat.json.part")) return false;
- if (!name.endsWith(".part")) return false;
- if (name.endsWith(".part.good")) return false;
- if (name.endsWith(".part.testing")) return false;
- return true;
-}
+// Media-file naming and the predicates over it live in lib/mediaFiles.ts —
+// pure string rules, no filesystem — and are re-exported here so every existing
+// import of videoStatus keeps working. The split exists because the editor's
+// VideoPanel is a client component that needs the same extension lists.
+export {
+ AUDIO_EXTS,
+ VIDEO_CONTAINER_EXTS,
+ MEDIA_EXTS,
+ SOURCE_MEDIA_BASENAME,
+ isRealAudioFile,
+ audioFilesToRemove,
+ isVideoContainer,
+ isSourceMediaFile,
+ findSourceMedia,
+ isPartAudioFile,
+} from "./mediaFiles";
export type SubTrack = {
// "live_chat" | language code like "es", "en-orig", etc.
diff --git a/common/ytdlp/audioCheckedDownload.ts b/common/ytdlp/audioCheckedDownload.ts
@@ -18,6 +18,7 @@
// finally to avoid orphaning a suspended child.
import { constants as fsConstants } from "node:fs";
+import { isPartAudioFile, isRealAudioFile } from "../lib/mediaFiles";
import {
copyFile,
open,
@@ -254,11 +255,13 @@ async function findExistingPartOrFinal(
} catch {
return { partFile: null, finalFile: null };
}
+ // Shares the app's media-file predicates rather than matching on the audio.
+ // prefix: `audio.live_chat.json.part` satisfies "starts with audio., ends with
+ // .part" and is a live-chat sidecar, not a resumable download.
for (const e of entries) {
- if (e === "audio.tmp" || e.startsWith("audio.tmp-")) continue;
- if (e.endsWith(".part")) {
- if (e.startsWith("audio.")) partFile = path.join(videoDir, e);
- } else if (e.startsWith("audio.")) {
+ if (isPartAudioFile(e)) {
+ partFile = path.join(videoDir, e);
+ } else if (isRealAudioFile(e)) {
// Skip already-extracted audio.<format> outputs that aren't the source
// container; ChannelConfig.audioFormat is the target, so if the file
// matches that, it's an output (or a stale one).
@@ -339,12 +342,7 @@ async function prepareDataTree(
onLog(`Removed stale snapshot: ${p}\n`);
} else if (e.endsWith(".part.good") && e.startsWith("audio.")) {
goodFile = path.join(dir, e);
- } else if (
- e.endsWith(".part") &&
- e.startsWith("audio.") &&
- !e.endsWith(".part.good") &&
- !e.endsWith(".part.testing")
- ) {
+ } else if (isPartAudioFile(e)) {
partFile = path.join(dir, e);
}
}
@@ -561,14 +559,7 @@ async function snapshotPart(
async function findPartIn(dir: string): Promise<string | null> {
const entries = await readdir(dir).catch(() => [] as string[]);
for (const e of entries) {
- if (
- e.startsWith("audio.") &&
- e.endsWith(".part") &&
- !e.endsWith(".part.testing") &&
- !e.endsWith(".part.good")
- ) {
- return path.join(dir, e);
- }
+ if (isPartAudioFile(e)) return path.join(dir, e);
}
return null;
}
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -1,6 +1,8 @@
# Changelog
## [Unreleased]
+- **A file called `audio.en-orig.vtt` is a subtitle, and the cleanup sweep used to think it was audio.** Every place in the app that asked "is this a media file?" answered by looking at the *start* of the name — anything beginning with `audio.` counted — and then subtracting a list of known exceptions. A list of exceptions is only ever as good as the last thing someone remembered to add to it, and it had fallen behind: a leaked subtitle track (`audio.en-orig.vtt`) matched nothing on the list, and a live-chat download fragment (`audio.live_chat.json.part-Frag114`) slipped past the `.part` rule because it doesn't *end* in `.part`. Since the sweep that deletes finished audio shares that same question, it could delete a subtitle believing it was audio — a small amount of real, unrecoverable data loss, in a sweep with no dry run and no trash. The same gap on the video-container side was worse: yt-dlp's own scratch file `source-media.temp.mp4` read as a finished video, and that is the one file the app *writes* from — it would extract audio from the corrupt scratch copy and then move it into the saved-video store as the permanent archive copy, orphaning the real container. **The rule is now the shape of the name rather than a list of exceptions**: a finished media file is `audio.<ext>` or `source-media.<ext>` with exactly one extension and nothing between, and `<ext>` has to be a format we actually recognize. That one rule rejects all four junk files, and it also rejects something an extension check alone would have let back in — the app's own half-written transcode output, `audio.tmp-12345.mp3`, which has a perfectly ordinary audio extension and whose deletion would race the transcode still writing it. Three knock-on fixes ride along: the "remove extra audio formats" sweep and the transcode-retry pass each carried their own private, staler copy of the old exception list (so the former could delete that subtitle too, and the latter could hand ffmpeg a `.vtt` file as "source audio"), and the audio-integrity checker could mistake a live-chat fragment for a resumable download. All four now ask the same question in the same place. **Everything moves in the safe direction:** strictly fewer files are ever deleted, videos whose only "audio" was one of these fakes now report honestly as having no audio instead of as a failed transcription, and such a video is no longer mistaken for one that has already been downloaded — so it gets fetched. See `common/lib/mediaFiles.ts` (new, with tests seeded from the real files this corpus holds) and `common/lib/videoStatus.ts`.
+- **Speaker-attribution turns that land outside their own chunk are now counted instead of quietly dropped.** The text-only attribution lane asks the model about one slice of a transcript at a time and discards any answer whose timestamp points somewhere outside that slice — correctly, since a mark in another slice's territory would fight that slice's own answer. But it discarded them silently, with no count anywhere, and a validation run measured the drop rate at **30% — 520 of 1,705 turns**. That is the difference between "this lane is working" and "this lane is throwing away one answer in three", and nothing on disk or in the log could tell you which. The discard is now recorded on the record itself (one entry per chunk, carrying the count, so a high rate doesn't bury every other warning) and summarized in the job log as a rate. Behaviour is otherwise unchanged and the lane remains switched off by default; this is instrumentation, not a change of policy.
- **The backfill lane can now be paused from the dashboard, next to the work it pauses.** The hold itself is not new — the lane has always checked, every time it goes to pick up the next video, whether it is still switched on, and parked itself without ending if it wasn't. What was missing was anywhere sensible to flip it: the only switch lived on the Settings page, which is a strange place to look for a control over a job you are watching run. **Pause now sits beside Start Backfill Sweep, and it is a hold rather than a stop** — the job stays alive, keeps its place, and picks up within a few seconds when you resume, with nothing recomputed and nothing lost. It also survives a restart, because it is a remembered preference rather than something applied to a running pool. This matters most for speaker diarization, which is the one backfill that is genuinely heavy: it runs on the processor rather than the graphics card, pins several cores, and takes hours per long video — so on a desktop you are actually sitting at, "not right now" needs to be one click, for reasons the scheduler cannot possibly know about. **Pause and Stop are deliberately different.** Pause keeps the sweep armed and holds it; Stop ends the sweep, letting the video in flight finish first, and you re-arm it later. The button and the Settings checkbox are the same switch under the hood, so they can never disagree about whether the lane is running.
- **The archive can now put names to the speakers — two ways, both switched off until you decide they are worth the time.** Diarization already captures *that* the speaker changed; it records anonymous clusters, not people. Naming them is a separate pass, and there are two honest ways to do it that are not interchangeable. **From the audio:** the diarizer has already grouped every voice, so the model is only asked to put a name to each group given samples of what it said — **about one model call for a whole video**, and the speaker boundaries come from the audio rather than from a guess. **From the transcript alone:** for videos with no diarization, which today is nearly all of them, the model has to read the transcript to find the changes in the first place — roughly **one call per chunk**, about 194,000 across this archive, which is the same order as the AI digest sweep and would compete with it for the same card. So both are built and **neither is armed**: they are off, under a master switch that is also off, and the honest next step is a small measured pilot rather than committing 25–55 days of GPU time to the worse of the two lanes. **Quality is not oversold anywhere in the UI.** A name from the audio carries a confidence the model was actually asked for; a name from the transcript carries none, because the model was asked where the speaker changes, not how sure it is who anyone is — and inventing a number there would be exactly the wrong kind of confident. The transcript-only lane misfires on rapid back-and-forth, and on auto-caption channels there are no speaker turns to find at all. **The two lanes share one file, and the rule about that is the important part.** Naming from the audio may replace a transcript-only record — that is an upgrade, and it is queued for you automatically the moment a video gains diarization. The reverse can never happen: a transcript-only pass will not overwrite a record made from audio, even when forced, even when the audio and its diarization have since been deleted and that record is the only thing left that knows who was speaking. **Both lanes plug into the Backfill lane rather than being new machinery**, so they inherit the resource share, the corpus-wide sweep, the restart-survival and every indicator — the channel's Backfill card, `/actionable`, the dashboard and the widget all report them beside diarization, with reachable work and needs-its-media-back still counted separately. That separation matters more here than it did for diarization: only one video in this archive currently *has* the diarization the good lane needs, so a single “remaining” figure would be 73,000 videos of work no button can start. **Staleness is provenance, not age**, as everywhere else: a name is stale when the engine, the model or the prompt generation differs from what would be produced now — and, uniquely here, when the video has been **re-diarized underneath it**, because a cluster number means nothing except relative to the run that produced it. There is a Prompt generation number in Settings you can raise to force a corpus-wide redo; it cannot be set lower, because pinning it to a superseded prompt would freeze that prompt's output into the archive looking current.
- **Backfill is now a first-class thing the system knows about, with a lane of its own and a remainder you can see.** Every derived-data feature that lands runs into the same wall: the corpus that already exists does not have what it needs. Diarization hit it first — the input it wants is audio, which Clean-audio deletes once a video is transcribed — and the only catch-up was a per-channel button written for that one feature, with no way to ask how much of the corpus was missing it and no way to run it alongside new-video work without one starving the other. So a feature now **declares** what it needs and how to tell whether a video has it, and gets the lane, the resource share and the indicator for free. **The number this reports is split in two, and that is the whole design.** On this corpus 836 videos still have their media and about 76,270 do not — 91× more — so a single "remaining" figure would be dominated by work no button can start, and every channel would sit at the top of every list forever. Reachable work and needs-re-acquiring are therefore separate numbers on all four surfaces: the channel's new **Backfill** stage card, a new section on `/actionable` (with the second figure in its own column, exactly as "Est. reclaim" has), a dashboard instrument, and an optional widget strip. **The lane runs on its own queue**, so it is genuinely concurrent with transcription rather than sitting behind it, and how much of the machine it may take is one setting: **0 (the default) means idle-only** — full speed while transcription is quiet, standing aside the instant it isn't — and any value above 0 is a guaranteed share, floored at 1 so a small number is a slow lane and not a stopped one. Catch-up on a corpus that already exists must never slow down new arrivals, and the default enforces that rather than trusting it. There is a corpus-wide **Start Backfill Sweep** on the dashboard that persists its scope along with the flag and resumes itself after a restart. **Staleness now means provenance, not age.** A diarization sidecar records the models and clustering threshold that produced it, and the lane compares that against what the current settings *would* produce — so changing the threshold, which is the single knob most likely to make you want a re-run, finally shows as work instead of leaving the corpus looking finished. A field the record predates compares equal to today's default, so adding one does not invalidate everything on disk. **Re-downloading deleted media is off by default and is the part to read carefully.** With it on, each file is fetched, used, and deleted again immediately in a `finally` — whether the backfill succeeded, failed, or crashed — unless the video is marked "do not clean", and nothing starts at all when free disk is under the configured floor. On a disk at 97% a leak here fills it, which is why those four behaviours have e2e tests of their own.
diff --git a/editor/app/channels/[slug]/videos/[id]/components/VideoPanel.tsx b/editor/app/channels/[slug]/videos/[id]/components/VideoPanel.tsx
@@ -13,6 +13,10 @@ import type { SavedVideoPointer } from "yt-dlp-transcript-common/lib/savedVideo"
import type { AvailabilityHistoryEntry } from "yt-dlp-transcript-common/lib/availability";
import type { SubtitleProvenance } from "yt-dlp-transcript-common/lib/subtitleProvenance";
import { formatBytes, formatDuration } from "yt-dlp-transcript-common/lib/format";
+import {
+ AUDIO_EXTS,
+ VIDEO_CONTAINER_EXTS,
+} from "yt-dlp-transcript-common/lib/mediaFiles";
import { QueueControl } from "../../../../../components/QueueControl";
import { cancelJobAction } from "../../../../../jobs/actions";
import { PipelineStageCard } from "../../../components/PipelineStageCard";
@@ -91,25 +95,12 @@ type Props = {
position?: { index: number; total: number };
};
-const AUDIO_EXTS = new Set([
- "mp3",
- "m4a",
- "aac",
- "ogg",
- "opus",
- "wav",
- "flac",
-]);
-
-const VIDEO_EXTS = new Set([
- "mp4",
- "webm",
- "mkv",
- "mov",
- "m4v",
- "ogv",
- "avi",
-]);
+// One list, shared with the detection predicates in lib/mediaFiles.ts — these
+// were duplicated here and drifting. isAudioFile deliberately stays on
+// AUDIO_EXTS alone: audio.mp4 is a wrong-format cleanup target, not something
+// this panel offers a transcribe button for (editor/e2e/video-page.spec.ts:426).
+const AUDIO_EXT_SET = new Set(AUDIO_EXTS);
+const VIDEO_EXT_SET = new Set(VIDEO_CONTAINER_EXTS);
const TEXT_EXTS = new Set(["vtt", "json", "txt", "srt", "log"]);
@@ -119,7 +110,7 @@ function fileExt(name: string): string {
}
function isAudioFile(name: string): boolean {
- return AUDIO_EXTS.has(fileExt(name));
+ return AUDIO_EXT_SET.has(fileExt(name));
}
// Files we can transcode FROM: any media file ffmpeg can demux to audio.
@@ -130,7 +121,7 @@ function isTranscodeSource(name: string): boolean {
if (!name.startsWith("audio.")) return false;
if (name.startsWith("audio.tmp-")) return false;
const ext = fileExt(name);
- return AUDIO_EXTS.has(ext) || VIDEO_EXTS.has(ext);
+ return AUDIO_EXT_SET.has(ext) || VIDEO_EXT_SET.has(ext);
}
function audioExt(name: string): string {
@@ -140,8 +131,8 @@ function audioExt(name: string): string {
function mediaKind(name: string): "audio" | "video" | "text" | null {
const ext = fileExt(name);
- if (AUDIO_EXTS.has(ext)) return "audio";
- if (VIDEO_EXTS.has(ext)) return "video";
+ if (AUDIO_EXT_SET.has(ext)) return "audio";
+ if (VIDEO_EXT_SET.has(ext)) return "video";
if (TEXT_EXTS.has(ext)) return "text";
return null;
}
diff --git a/plans/FACTS.md b/plans/FACTS.md
@@ -1708,9 +1708,10 @@ Bounded (913 real audio files against 3 fakes), but it is not only cosmetic:
ffmpeg, so those 3 videos fail the backfill. Fails fast and is visible in the
outcome tally; deliberately NOT special-cased in the backfill script, which
shares the runner's resolver on purpose so the two cannot disagree.
-- **`audioFilesToDelete` uses the same predicate**, so the cleanup sweep can
+- **`audioFilesToRemove` uses the same predicate**, so the cleanup sweep can
delete `audio.en-orig.vtt` — a subtitle — believing it is audio. That is real,
- if small, data loss.
+ if small, data loss. (Recorded here twice as `audioFilesToDelete`, which is not
+ a function that exists; corrected 2026-08-08 along with the fix itself.)
- `readVideoFiles().audioFiles` feeds the cleanup buckets, so "Est. reclaim"
counts these too.
@@ -1747,12 +1748,19 @@ Two consequences to expect, both small and both permanent until it is fixed:
those 3 fail identically on every re-run.
2. With the guard armed they are transcribed-and-undiarized forever, so the
cleanup sweep will report **3 videos "awaiting diarization" permanently**.
- Ironically the guard is now the thing preventing `audioFilesToDelete` — which
+ Ironically the guard is now the thing preventing `audioFilesToRemove` — which
shares the broken predicate — from deleting that subtitle as if it were audio.
-The fix is to require a known media extension rather than denylisting; it is
-deferred to its own branch because it changes cleanup behaviour and several
-exact-string e2e assertions.
+**FIXED 2026-08-08.** Requiring a known media extension turned out not to be
+enough on its own — this repo's own transcode scratch is
+`audio.tmp-<pid>.<targetFormat>`, a perfectly valid audio extension, so an
+extension allowlist would have re-admitted it and raced the running transcoder.
+The shipped rule is an ANCHORED SINGLE-SEGMENT match (`^audio\.(<media exts>)$`,
+`^source-media\.(<media exts>)$`, `^audio\.(<media exts>)\.part$`) over
+`AUDIO_EXTS ∪ VIDEO_CONTAINER_EXTS`, which rejects all four junk names and the
+scratch file with one rule. Lives in the new `lib/mediaFiles.ts` (pure string
+rules, no `node:fs`, so the editor's client components can share the lists);
+`lib/videoStatus.ts` re-exports it. Tests: `lib/mediaFiles.test.ts`.
### CORRECTION + escalation: the kernel OOM-killed a diarization run