Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit c89aec1a6c95033a2e3e4f2305fe9b9a528642f4
parent d5a0674d01d84174a44ddd05b672f94740354530
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon, 21 Sep 2026 02:35:09 -0400

review: the discard was a deny-list, and four things it let through

HIGH. discardPrefetchDir is a recursive delete, so the question it has to
answer is "is EVERYTHING in here mine?" — and it was answering "is anything in
here one of the three things I thought to check for". The deny-list named
isVideoDownloaded, clips/ and saved-video.json, so it would have deleted a
chat-only corpus member (transcript.live_chat.json + live_chat.cues.json), a
transcript.es.vtt, a resumable audio.mp3.part, a diarization.json or a
digest.json. Not hypothetical: the chat-only branch calls this whenever its
chat pass came back with no file, so a chat re-fetch that failed or was
aborted would have deleted the chat that was already there.

It is an allow-list now — {metadata.info.json, download.log,
download-outcome.json}, files only, a directory is never ours — over a real
readdir whose catch is reachable (the old one sat behind readVideoFiles, which
swallows its own readdir error, so it could never fire). One unit test per
protected shape, because the difference between the two rules is invisible to
every end-to-end path that does not happen to have the right file on disk.

Three more the review found:

- A settled video has no directory now, and both sets deriveChannelSets
  produces are defined by the ABSENCE of one — so a filtered-out video that
  later left the listing read as missingNeverFetched, the loudest alarm this
  system raises, pointed at a video the operator asked us not to fetch. The
  settled set is subtracted, at both call sites, from that and from
  `undownloaded`: settled work is not work.
- A chat-only stream with no chat replay was remembered nowhere. Nothing
  fetched means the dir is discarded, which means no sidecar, which means the
  id sat in chatOnlyPending forever — markCompleted retires it for the SESSION
  only, so every runner restart re-prefetched it to run a pass that will never
  return anything. A clean pass that finds nothing now writes noLiveChat on
  the scan entry and the id settles as the ordinary filtered-out livestream it
  is. A FAILED pass does not set it; that is why the chat pass moved above the
  store write rather than staying below it.
- The chat-only verdict is decided off the FILE on disk, not the mode. Asking
  chatOnlyIds made the same directory a settled STUB the moment
  rejectedLivestreams flipped back to "skip" — and settledOnDisk is subtracted
  from totals.videos, so the site kept serving a video the report had stopped
  counting, under Diagnostics copy promising "what is already on disk stays".

And the scan pre-pick now honours a resolved focus the way strict descent
does: while the focus still has pending work only a focused channel is offered
a scan, because otherwise "only jeralyzer is moving" is false the moment
another channel has titles to read on the same network.

The e2e that 0facbdfb traded away is back, with the case that actually
survives the destination prefilter: a rejected video holding audio.mp3 on a
youtube-handling channel, whose destination is transcript.en.vtt.

Prose: a chat-only dir also holds live_chat.cues.json, the outcome sidecar and
the log; and "a settled video is never invoked" is true of the two batch paths
and not of importVideoAction, which reaches downloadOneManaged by URL.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
Mcommon/controller/autoRunner.test.ts | 35+++++++++++++++++++++++++++++++++++
Mcommon/controller/autoRunner.ts | 32+++++++++++++++++++++++++++++---
Mcommon/controller/channelSets.ts | 16+++++++++++++++-
Mcommon/controller/channelSnapshot.test.ts | 142++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-
Mcommon/controller/channelSnapshot.ts | 61++++++++++++++++++++++++++++++++++++++++++++-----------------
Mcommon/controller/metadataScanStore.ts | 17++++++++++++++++-
Mcommon/lib/downloadOutcome.ts | 9+++++----
Acommon/ytdlp/downloadOneManaged.test.ts | 140+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/ytdlp/downloadOneManaged.ts | 227+++++++++++++++++++++++++++++++++++++++++++++++--------------------------------
Mcommon/ytdlp/runYtdlp.ts | 18+++++++++++++++++-
Meditor/e2e/title-filter.spec.ts | 47+++++++++++++++++++++++++++++------------------
Mplans/FACTS.md | 61+++++++++++++++++++++++++++++++++++++++++++++++++++----------
12 files changed, 656 insertions(+), 149 deletions(-)

diff --git a/common/controller/autoRunner.test.ts b/common/controller/autoRunner.test.ts @@ -444,3 +444,38 @@ test("the first candidate in list order wins, because that order is the priority ); }); +test("a focus holds the scan the same way it holds every download pick", () => { + // The compiled tree makes a focus hold downloads by strict descent, and the + // pre-pick does not go through that tree. Without this, "only jeralyzer is + // moving" is false the moment another channel has titles to read — on the + // same network the focus is trying to have to itself. + const channels = [ + work({ slug: "other", scanUnscanned: 50, filtered: true }), + work({ slug: "focused", scanUnscanned: 4, filtered: true }), + ]; + assert.deepEqual( + pickMetadataScanChannel(channels, { + ...noSkip, + focus: { slugs: new Set(["focused"]), holding: true }, + }), + { slug: "focused", targets: 4 }, + ); + // Focus exhausted — strict descent moves on, and so does the scan. + assert.deepEqual( + pickMetadataScanChannel(channels, { + ...noSkip, + focus: { slugs: new Set(["focused"]), holding: false }, + }), + { slug: "other", targets: 50 }, + ); + // A holding focus with no scan work of its own offers nothing, rather than + // falling through to the channel it is holding back. + assert.equal( + pickMetadataScanChannel( + [work({ slug: "other", scanUnscanned: 50, filtered: true })], + { ...noSkip, focus: { slugs: new Set(["focused"]), holding: true } }, + ), + null, + ); +}); + diff --git a/common/controller/autoRunner.ts b/common/controller/autoRunner.ts @@ -1067,6 +1067,15 @@ function eligibleSlots(): number { // platform offers no scan either, because a scan is one more thing asking that // source for titles. // +// A FOCUS HOLDS THE SCAN TOO. The compiled priority tree makes a focus hold +// every download pick by strict descent (see laneDispatchRoot), and the pre-pick +// does not go through that tree — so without this, "only jeralyzer is moving" +// would be false the moment another channel had titles to read, on the same +// network the focus is trying to have to itself. The rule mirrors what strict +// descent does rather than re-implementing it: while the focus still has +// pending work, only a focused channel is offered a scan; once it is exhausted +// the lane is free and so is the scan. +// // FIRST IN `channels` ORDER, which is `metaCache` order, which is the priority // order every other pick on this lane already uses. export function pickMetadataScanChannel( @@ -1075,11 +1084,17 @@ export function pickMetadataScanChannel( scanned: ReadonlySet<string>; platformSkip: ReadonlySet<string>; platformOf: (slug: string) => string; + // The resolved focus, when one is holding the lane. Absent (or with + // `holding: false`) means every channel is eligible, which is the default + // and the byte-identical path for a corpus with no focus configured. + focus?: { slugs: ReadonlySet<string>; holding: boolean }; }, ): { slug: string; targets: number } | null { + const held = opts.focus?.holding === true; for (const c of channels) { const targets = c.scanUnscanned ?? 0; if (!c.filtered || targets <= 0) continue; + if (held && !opts.focus!.slugs.has(c.slug)) continue; if (opts.scanned.has(c.slug)) continue; if (opts.platformSkip.has(opts.platformOf(c.slug))) continue; return { slug: c.slug, targets }; @@ -1562,11 +1577,14 @@ async function runLoop( // Say — ONCE per transition — whether a focus is holding this lane. Read off // the `prio-*` leaf ids in the map just built, so it costs one pass over // keys and no new read, and it is skipped entirely while no focus resolves. - reportFocusHold( + // Computed ONCE and reused by the scan pre-pick below, which has to honour + // the same hold: `holding` is `focusPending > 0`, i.e. exactly "strict + // descent is still inside the focus group". + const focus = ctx.focusSlugs.length > 0 ? focusSummary(ctx.model, ctx.focusSlugs, pending) - : null, - ); + : null; + reportFocusHold(focus); // THE RUNNER JOB'S OWN BAR, on the metrics the per-channel jobs already // use — so a lane's runner row reads like the manual verb's row rather than // like an opaque long-lived loop. `target` moves as the corpus does (this @@ -1642,6 +1660,14 @@ async function runLoop( scanned: scannedThisRun, platformSkip, platformOf: (slug) => platformKey(slug, slugToPlatform), + ...(focus + ? { + focus: { + slugs: new Set(ctx.focusSlugs), + holding: focus.holding, + }, + } + : {}), }); if (scanCandidate) { const { slug, targets } = scanCandidate; diff --git a/common/controller/channelSets.ts b/common/controller/channelSets.ts @@ -37,15 +37,28 @@ export function deriveChannelSets(input: { roster: Roster; listedIds: ReadonlySet<string>; onDiskIds: ReadonlySet<string>; + // Ids this channel's download filter has SETTLED — the operator's own "not + // this one". Optional, and empty for every channel without a filter. + // + // WHY IT BELONGS HERE. Both sets below are defined by the ABSENCE of a + // directory, and since a title-filter rejection stopped leaving its prefetch + // dir behind, a settled video has none. Without this, a filtered-out video + // that later left the listing read as `missingNeverFetched` — "we were told + // about this, never got it, and now it is gone" — which is the highest-value + // alarm this system raises, pointed at a video the operator asked us not to + // fetch. `undownloaded` has the same shape of wrongness: it is a work list, + // and settled work is not work. + settledIds?: ReadonlySet<string>; }): ChannelSets { const { roster, listedIds, onDiskIds } = input; + const settledIds = input.settledIds ?? new Set<string>(); const rosterIds = Object.keys(roster.entries); const listed: string[] = []; const undownloaded: string[] = []; for (const id of listedIds) { listed.push(id); - if (!onDiskIds.has(id)) undownloaded.push(id); + if (!onDiskIds.has(id) && !settledIds.has(id)) undownloaded.push(id); } // missingDownloaded is derived from the DISK set, not from `roster ∩ disk`, @@ -64,6 +77,7 @@ export function deriveChannelSets(input: { const missingNeverFetched: string[] = []; for (const id of rosterIds) { + if (settledIds.has(id)) continue; if (!listedIds.has(id) && !onDiskIds.has(id)) missingNeverFetched.push(id); } diff --git a/common/controller/channelSnapshot.test.ts b/common/controller/channelSnapshot.test.ts @@ -370,7 +370,14 @@ test("generateChannelSnapshot refuses an unreachable channel rather than writing type FilterFixture = { filter?: Record<string, unknown>; listed: string[]; - scanned?: Record<string, { title?: string; liveStatus?: string }>; + scanned?: Record< + string, + { title?: string; liveStatus?: string; noLiveChat?: boolean } + >; + // Ids the channel has EVER been seen to contain. Seeded so the + // missingNeverFetched rule — "we were told about this, never got it, and it + // is gone" — can be exercised: it is derived from the roster, not the disk. + roster?: string[]; }; async function filteredChannel( @@ -405,9 +412,27 @@ async function filteredChannel( description: "", uploadDate: "20240101", ...(e.liveStatus ? { liveStatus: e.liveStatus } : {}), + ...(e.noLiveChat ? { noLiveChat: true } : {}), scannedAt: "2026-01-01T00:00:00.000Z", }; } + if (fixture.roster) { + await writeFile( + path.join(channelDir, "roster.json"), + JSON.stringify({ + version: 1, + entries: Object.fromEntries( + fixture.roster.map((id) => [ + id, + { + url: `https://www.youtube.com/watch?v=${id}`, + firstSeenAt: "2026-01-01T00:00:00.000Z", + }, + ]), + ), + }), + ); + } await writeFile( path.join(channelDir, "metadata-scan.json"), JSON.stringify({ version: 1, entries, errors: {}, lastRun: null }), @@ -541,3 +566,118 @@ test("turning the mode off re-decides the channel with nothing to migrate", asyn } }); +test("a settled video that left the listing is NOT reported as never fetched", async () => { + // THE LOUDEST ALARM THIS REPORT RAISES, pointed at the wrong video. Both + // missingNeverFetched and undownloaded are defined by the ABSENCE of a + // directory, and a settled video has none since a rejection stopped leaving + // its prefetch dir behind — so without subtracting the settled set, the sweep + // tells the operator to attempt a direct-link recovery of a video they + // configured us not to fetch. + const { dir, paths } = await filteredChannel({ + filter: { include: "keep" }, + // `bbbb0000002` is in the roster and NOT in the listing any more. + listed: ["aaaa0000001"], + roster: ["aaaa0000001", "bbbb0000002"], + scanned: { + aaaa0000001: { title: "keep this one" }, + bbbb0000002: { title: "drop this one" }, + }, + }); + try { + const snap = await generateChannelSnapshot(paths, "alpha"); + assert.deepEqual( + (snap.missingNeverFetched ?? []).map((m) => m.id), + [], + ); + // It is still settled, and still reported as such — it left the alarm, not + // the report. + assert.deepEqual(snap.buckets.skippedByTitleFilter, ["bbbb0000002"]); + } finally { + await rm(dir, { recursive: true, force: true }); + } +}); + +test("an unsettled video that left the listing still raises the alarm", async () => { + // The control for the test above: subtracting the settled set must not have + // made the category unreachable. + const { dir, paths } = await filteredChannel({ + filter: { include: "keep" }, + listed: [], + roster: ["aaaa0000001"], + scanned: { aaaa0000001: { title: "keep this one" } }, + }); + try { + const snap = await generateChannelSnapshot(paths, "alpha"); + assert.deepEqual( + (snap.missingNeverFetched ?? []).map((m) => m.id), + ["aaaa0000001"], + ); + } finally { + await rm(dir, { recursive: true, force: true }); + } +}); + +test("a stream with no chat replay leaves chatOnlyPending for good", async () => { + // A clean pass that found nothing is an ANSWER: chat replay was off, and + // asking again tomorrow gets the same nothing. Without the flag the id sits + // in chatOnlyPending forever and every runner restart re-prefetches it to run + // a pass that will never return anything. + const { dir, paths } = await filteredChannel({ + filter: { include: "keep", rejectedLivestreams: "chat-only" }, + listed: ["cccc0000003"], + scanned: { + cccc0000003: { + title: "a long stream", + liveStatus: "was_live", + noLiveChat: true, + }, + }, + }); + try { + const snap = await generateChannelSnapshot(paths, "alpha"); + assert.deepEqual(snap.buckets.chatOnlyPending, []); + assert.deepEqual(snap.buckets.chatOnly, []); + // It settles as the ordinary filtered-out livestream it is. + assert.deepEqual(snap.buckets.skippedByTitleFilter, ["cccc0000003"]); + assert.ok(!snap.undownloadedIds.includes("cccc0000003")); + } finally { + await rm(dir, { recursive: true, force: true }); + } +}); + +test("a chat already on disk stays a member after the mode is turned off", async () => { + // buildIndex has already published it as a chat track, and flipping the mode + // back to "skip" does not un-publish it. Deciding this off `chatOnlyIds` + // would make the same directory a settled STUB — and settledOnDisk is + // subtracted from totals.videos, so the site would serve a video the report + // had stopped counting. + const { dir, paths, channelDir } = await filteredChannel({ + filter: { include: "keep" }, + listed: ["cccc0000003"], + scanned: { + cccc0000003: { title: "a long stream", liveStatus: "was_live" }, + }, + }); + try { + const videoDir = path.join(channelDir, "data", "cccc0000003"); + await mkdir(videoDir, { recursive: true }); + await writeFile( + path.join(videoDir, "metadata.info.json"), + JSON.stringify({ id: "cccc0000003", title: "a long stream" }), + ); + await writeFile( + path.join(videoDir, "transcript.live_chat.json"), + '{"action":{}}\n', + ); + + const snap = await generateChannelSnapshot(paths, "alpha"); + assert.deepEqual(snap.buckets.chatOnly, ["cccc0000003"]); + assert.deepEqual(snap.buckets.skippedByTitleFilter, []); + assert.equal(snap.totals.videos, 1); + // And nothing outstanding: the mode is off, so there is no chat to fetch. + assert.deepEqual(snap.buckets.chatOnlyPending, []); + } finally { + await rm(dir, { recursive: true, force: true }); + } +}); + diff --git a/common/controller/channelSnapshot.ts b/common/controller/channelSnapshot.ts @@ -218,7 +218,14 @@ export type ChannelSnapshot = { // CHAT-ONLY videos whose chat is on disk. A rejected livestream on a channel // whose `downloadFilter.rejectedLivestreams` is "chat-only": the media was // never fetched, the live chat was, and the directory holds - // metadata.info.json + transcript.live_chat.json and nothing else. + // metadata.info.json, transcript.live_chat.json and its normalized + // live_chat.cues.json, beside the download-outcome.json and download.log + // every managed download leaves. No media, and no transcript. + // + // THE FILE ON DISK DECIDES, not the mode: an id whose chat has landed stays + // in this bucket after `rejectedLivestreams` is set back to "skip", + // because buildIndex has already published it and flipping a setting does + // not un-publish it. // // A MEMBER OF THIS CORPUS, not a stub. It is in totals.videos, it is in the // LMDB index, and the site publishes it as a chat track with no captions. @@ -1116,24 +1123,27 @@ export async function generateChannelSnapshot( // it from transcribedWithAudio and the cleanup estimates while it still // counted in totals.transcribed, i.e. a channel reporting more transcripts // than videos. Settlement only ever decides what NOT to fetch. - // CHAT ONLY, and it is a member of this corpus rather than a stub. The dir - // holds metadata.info.json and transcript.live_chat.json, which is exactly - // what buildIndex needs to publish it as a chat track with no captions. It - // is short-circuited here for the same reason the settled branch below is — - // everything under this line classifies a video by what is MISSING, and - // nothing is missing from a chat-only video — but it is NOT settledOnDisk: - // that list is subtracted from totals.videos, and this one is a video we - // deliberately have. - if (chatOnlyIds.has(id) && !videoHasAnyArtifact(files)) { + // SETTLED, and the two answers it can have. + // + // THE CHAT ON DISK DECIDES, NOT THE MODE. A dir holding + // transcript.live_chat.json is already published by buildIndex as a chat + // track with no captions — that is a fact about the corpus, and flipping + // `rejectedLivestreams` back to "skip" does not un-publish it. Asking + // `chatOnlyIds` here instead would make the same directory count as a + // settled STUB the moment the mode changed, and settledOnDisk is subtracted + // from totals.videos: the site would still serve the video while the report + // stopped counting it. So the file is the test, and the Diagnostics copy — + // "what is already on disk stays" — is true. + // + // Short-circuited for the same reason the settled branch is: everything + // under this line classifies a video by what is MISSING, and nothing is + // missing from a chat-only video. + if (settledIds.has(id) && !videoHasAnyArtifact(files)) { if (files.entries.includes(LIVE_CHAT_FILENAME)) { chatOnly.push(id); - continue; + } else { + settledOnDisk.push(id); } - // A directory with neither media nor chat is a stub like any other - // rejection's, and falls through to the settled branch below. - } - if (settledIds.has(id) && !videoHasAnyArtifact(files)) { - settledOnDisk.push(id); continue; } if (!files.hasMeta && !excludedById.has(id)) noMetadata.push(id); @@ -1328,6 +1338,12 @@ export async function generateChannelSnapshot( urls.map((u) => extractVideoId(u)).filter((id): id is string => Boolean(id)), ), onDiskIds: new Set(videoDirNames), + // See deriveChannelSets: both sets it derives are defined by the ABSENCE of + // a directory, and a settled video has none since a rejection stopped + // leaving its prefetch dir behind. Without this, `missingNeverFetched` — + // the loudest alarm this report raises — would fire for videos the operator + // asked us not to fetch. + settledIds, }); const listedIdSet = new Set(sets.listed); @@ -1366,6 +1382,10 @@ export async function generateChannelSnapshot( // artifact-derived bucket can carry it. Once the chat lands it drops out // here and appears in `chatOnly` instead. if (chatOnlyIds.has(dirId)) { + // Already a member (the chat is on disk) => nothing outstanding. The scan + // store's `noLiveChat` flag is the other way out of this list: a stream + // we asked and got nothing from is not in `chatOnlyIds` at all, so it + // settles as an ordinary rejection instead of sitting here forever. if (!f?.entries.includes(LIVE_CHAT_FILENAME)) chatOnlyPending.push(dirId); continue; } @@ -1414,6 +1434,7 @@ export async function generateChannelSnapshot( const corruptSourceSet = new Set(corruptSource); const corruptFullSourceSet = new Set(corruptFullSource); + const chatOnlySet = new Set(chatOnly); const snapshotBuckets: ChannelSnapshot["buckets"] = { noTranscript: noTranscript.sort(), downloadedNoTranscript: downloadedNoTranscript.sort(), @@ -1443,7 +1464,13 @@ export async function generateChannelSnapshot( // we decided to have its chat. skippedByTitleFilter: [...settledIds] .filter((id) => { - if (chatOnlyIds.has(id)) return false; + // Minus the two chat buckets, whichever way an id got into them — a + // chat-only video IS settled (its media is not wanted) but reporting it + // as "skipped by the title filter" would say we decided not to have it, + // when we decided to have its chat. `chatOnlySet` is what the loop + // above actually classified (the chat is on disk); `chatOnlyIds` is the + // outstanding half. + if (chatOnlySet.has(id) || chatOnlyIds.has(id)) return false; const f = filesById.get(id); return !f || !videoHasAnyArtifact(f); }) diff --git a/common/controller/metadataScanStore.ts b/common/controller/metadataScanStore.ts @@ -49,6 +49,14 @@ export type MetadataScanEntry = { uploadDate: string; liveStatus?: string; duration?: number; + // THE CHAT-ONLY DEAD END. Set when a chat-only pass ran cleanly and the + // source returned no live chat — replay was off for that stream, or it has + // since been dropped. Asking again tomorrow gets the same nothing, so this is + // an ANSWER and not a retry: `chatOnlyIdsFrom` drops the id, it leaves + // `chatOnlyPending` for good, and it settles as the ordinary filtered-out + // livestream it is. Deliberately NOT set by a pass that FAILED (a non-zero + // exit, a cooldown, an abort) — that one has to stay retryable. + noLiveChat?: boolean; scannedAt: string; }; @@ -107,6 +115,7 @@ function normalizeEntry(raw: unknown): MetadataScanEntry | null { if (typeof r.duration === "number" && Number.isFinite(r.duration)) { entry.duration = r.duration; } + if (r.noLiveChat === true) entry.noLiveChat = true; return entry; } @@ -228,7 +237,8 @@ export async function upsertMetadataScan( prev.description !== entry.description || prev.uploadDate !== entry.uploadDate || prev.liveStatus !== entry.liveStatus || - prev.duration !== entry.duration + prev.duration !== entry.duration || + prev.noLiveChat !== entry.noLiveChat ) { scan.entries[id] = entry; changed = true; @@ -304,6 +314,11 @@ export function chatOnlyIdsFrom( // that is not chat-only pays one field read for this question. if (!compiled || compiled.rejectedLivestreams !== "chat-only") return ids; for (const [id, entry] of Object.entries(scan.entries)) { + // A stream we already asked and got nothing from is not chat-only work any + // more — it is an ordinary settled rejection. Without this it sits in + // chatOnlyPending forever and every runner restart re-prefetches it to run + // a pass that will never return anything. + if (entry.noLiveChat) continue; if (titleFilterWantsChat(compiled, entry)) ids.add(id); } return ids; diff --git a/common/lib/downloadOutcome.ts b/common/lib/downloadOutcome.ts @@ -30,10 +30,11 @@ export type DownloadOutcomeStatus = | "skipped-filtered" // The download filter rejected this livestream and the channel's // `rejectedLivestreams` mode is "chat-only": the MEDIA was not fetched, the - // live chat was. The directory holds metadata.info.json and - // transcript.live_chat.json and nothing else, so the video is indexable (a - // chat track, no captions) while every "is it downloaded?" predicate — all of - // which test whisper/VTT/audio — still answers no. Terminal on the operator's + // live chat was. The directory holds metadata.info.json, + // transcript.live_chat.json and its live_chat.cues.json (plus this sidecar + // and download.log) — no media and no transcript — so the video is indexable + // as a chat track with no captions while every "is it downloaded?" predicate + // — all of which test whisper/VTT/audio — still answers no. Terminal on the operator's // terms, like skipped-filtered, and equally re-decided by editing the filter. | "chat-only"; diff --git a/common/ytdlp/downloadOneManaged.test.ts b/common/ytdlp/downloadOneManaged.test.ts @@ -0,0 +1,140 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { mkdir, mkdtemp, readdir, rm, writeFile } from "node:fs/promises"; +import { tmpdir } from "node:os"; +import path from "node:path"; +import { __discardPrefetchDirForTest as discardPrefetchDir } from "./downloadOneManaged"; + +// THE DISCARD IS A RECURSIVE DELETE, so the only test worth having is the one +// that asks what it REFUSES. +// +// It used to be a deny-list naming isVideoDownloaded, `clips/` and +// `saved-video.json`, which is a shape that cannot be right: every file nobody +// thought of is deleted by default. The cases below are the ones that were +// actually reachable — a chat-only corpus member is the worst of them, because +// the chat-only branch calls this whenever its chat pass came back with no +// file, so a failed re-fetch would have deleted the chat that was already +// there. +// +// Each case is one directory, one call, one question: did it survive? + +async function dirWith( + files: Record<string, string>, + dirs: string[] = [], +): Promise<{ root: string; videoDir: string }> { + const root = await mkdtemp(path.join(tmpdir(), "ttb-discard-")); + const videoDir = path.join(root, "data", "vid0000001"); + await mkdir(videoDir, { recursive: true }); + for (const [name, body] of Object.entries(files)) { + await writeFile(path.join(videoDir, name), body); + } + for (const name of dirs) await mkdir(path.join(videoDir, name)); + return { root, videoDir }; +} + +const noop = () => {}; + +async function discardOf( + files: Record<string, string>, + dirs: string[] = [], +): Promise<{ removed: boolean; survivors: string[] }> { + const { root, videoDir } = await dirWith(files, dirs); + try { + const removed = await discardPrefetchDir(videoDir, noop); + const survivors = await readdir(videoDir).catch(() => null); + return { removed, survivors: survivors ?? [] }; + } finally { + await rm(root, { recursive: true, force: true }); + } +} + +test("a directory holding nothing but this pass's own files is removed", async () => { + // Exactly what a rejected prefetch leaves: the info json it wrote, the + // managed download's log, and the outcome sidecar an EARLIER attempt on the + // same id left behind (including one a pre-rule rejection wrote). + const { removed, survivors } = await discardOf({ + "metadata.info.json": "{}", + "download.log": "yt-dlp\n", + "download-outcome.json": '{"status":"skipped-filtered"}', + }); + assert.equal(removed, true); + assert.deepEqual(survivors, []); +}); + +test("an empty directory is removed, and a missing one is not an error", async () => { + assert.equal((await discardOf({})).removed, true); + const root = await mkdtemp(path.join(tmpdir(), "ttb-discard-")); + try { + // Nothing to list means nothing to delete. An `rm -rf` on a path we could + // not read is the last thing to do about an unreadable path. + assert.equal( + await discardPrefetchDir(path.join(root, "data", "nope"), noop), + false, + ); + } finally { + await rm(root, { recursive: true, force: true }); + } +}); + +// Every one of these is a file the deny-list did not name and would have +// deleted. The assertion is the same each time: the directory is still there, +// with its contents intact. +const PROTECTED: Array<[string, Record<string, string>, string[]?]> = [ + [ + "a chat-only corpus member", + { + "metadata.info.json": "{}", + "transcript.live_chat.json": '{"replayChatItemAction":{}}\n', + "live_chat.cues.json": '{"cues":[]}', + "download-outcome.json": '{"status":"chat-only"}', + }, + ], + [ + "a downloaded video", + { "metadata.info.json": "{}", "audio.mp3": "bytes" }, + ], + [ + "a video whose only transcript is ours", + { "metadata.info.json": "{}", "transcript.json": "{}" }, + ], + [ + "a video whose only transcript is a foreign VTT", + { "metadata.info.json": "{}", "transcript.es.vtt": "WEBVTT\n" }, + ], + [ + "a resumable partial download", + { "metadata.info.json": "{}", "audio.mp3.part": "half" }, + ], + [ + "a diarization sidecar", + { "metadata.info.json": "{}", "diarization.json": "{}" }, + ], + [ + "a digest sidecar", + { "metadata.info.json": "{}", "digest.json": "{}" }, + ], + [ + "a saved-video pointer", + { "metadata.info.json": "{}", "saved-video.json": "{}" }, + ], + [ + "a clip window another tool asked for", + { "metadata.info.json": "{}" }, + ["clips"], + ], + [ + "anything at all that nobody has thought of yet", + { "metadata.info.json": "{}", "something-new.json": "{}" }, + ], +]; + +for (const [what, files, dirs] of PROTECTED) { + test(`${what} keeps its directory`, async () => { + const { removed, survivors } = await discardOf(files, dirs ?? []); + assert.equal(removed, false, `${what} was discarded`); + assert.deepEqual( + survivors.sort(), + [...Object.keys(files), ...(dirs ?? [])].sort(), + ); + }); +} diff --git a/common/ytdlp/downloadOneManaged.ts b/common/ytdlp/downloadOneManaged.ts @@ -1,6 +1,6 @@ import path from "node:path"; import { appendFile, mkdir, readdir, readFile, rm } from "node:fs/promises"; -import { createWriteStream, type WriteStream } from "node:fs"; +import { createWriteStream, type Dirent, type WriteStream } from "node:fs"; import { execa } from "execa"; import { AUTH_RETRY_CLASSES, @@ -28,16 +28,8 @@ import { } from "../controller/keptVideos"; import { isDoNotClean } from "../lib/doNotClean-server"; import { transcodeAudio } from "../controller/transcode"; -import { - findSourceMedia, - isVideoDownloaded, - readVideoFiles, -} from "../lib/videoStatus"; -import { - SAVED_VIDEO_POINTER_FILENAME, - savedVideoDir, - type SavedVideoOrigin, -} from "../lib/savedVideo"; +import { findSourceMedia } from "../lib/videoStatus"; +import { savedVideoDir, type SavedVideoOrigin } from "../lib/savedVideo"; import { persistSourceVideo } from "../lib/savedVideo-server"; import { type AudioCheckAttemptStats, @@ -56,7 +48,6 @@ import { import { evaluateDownloadFilters } from "../lib/downloadFilters"; import { normalizeLiveChat } from "../controller/normalizeLiveChat"; import { LIVE_CHAT_FILENAME } from "../lib/videoStatus"; -import { CLIPS_DIR_NAME } from "../lib/clipWindow"; import { upsertMetadataScan, type MetadataScanEntry, @@ -348,37 +339,63 @@ const channelConfigArgs = channelExtraArgs; // // THE BUCKET DOES NOT COME FROM THIS DIRECTORY. `skippedByTitleFilter` is // derived in the snapshot from the metadata-scan store against the channel's -// CURRENT filter (channelSnapshot.ts, settledIdsFrom) — the entry recorded a -// few lines above this call — so removing the dir costs the count nothing. The -// retryable `skippedByFilter` bucket IS read off `download-outcome.json`, which -// is why every OTHER filter (skip-live) still writes one: that skip says "we -// will try again", and a video with no directory and no outcome would silently -// leave it. +// CURRENT filter (channelSnapshot.ts, settledIdsFrom) — the entry recorded +// beside this call — so removing the dir costs the count nothing. The retryable +// `skippedByFilter` bucket IS read off `download-outcome.json`, which is why +// every OTHER filter (skip-live) still writes one: that skip says "we will try +// again", and a video with no directory and no outcome would silently leave it. +// +// ── IT IS AN ALLOW-LIST, AND IT HAS TO BE ──────────────────────────────────── // -// REFUSES RATHER THAN GUESSES. The prefetch is not the only thing that can have -// written into this dir — a video downloaded before the filter was authored is -// DOWNLOADED, which is a fact and not a preference, and a clip window or a -// saved-video pointer is somebody else's data. Any of those and the directory -// stays exactly as it is; the caller's outcome record is unaffected either way. +// This function deletes a directory recursively, so the question it must answer +// is "is EVERYTHING in here mine?", not "is anything in here one of the few +// things I thought to check for". The deny-list it replaced named +// isVideoDownloaded, `clips/` and `saved-video.json` — and would therefore have +// deleted a chat-only corpus member (`transcript.live_chat.json` + +// `live_chat.cues.json`), a foreign-language `transcript.es.vtt`, a resumable +// `audio.mp3.part`, a `diarization.json` or a `digest.json`. That is not a +// hypothetical: the chat-only branch below calls this whenever its chat pass +// did NOT come back with a file, so a chat re-fetch that failed, was aborted +// mid-write or hit a cooldown would have deleted the chat that was already +// there, and an `importVideoAction` on an existing chat-only id would have done +// the same. +// +// So: the dir goes only when every entry is something THIS pass wrote or could +// have found from a previous run of itself. Anything else — any file, any +// subdirectory, anything a future feature adds — keeps it, and the caller falls +// back to writing the outcome sidecar exactly as it always did. Being wrong in +// this direction costs one metadata stub; being wrong in the other costs bytes +// nobody can enumerate. +const PREFETCH_OWN_FILES: ReadonlySet<string> = new Set([ + // What the prefetch pass itself writes: --write-info-json under the + // `infojson:` output template (outputArgsForUrl). + "metadata.info.json", + // The managed download's own per-video log, opened before the prefetch runs. + "download.log", + // A sidecar from an EARLIER managed attempt on this same id — including the + // one a previous rejection wrote before this rule existed. It records what + // happened, never what is on disk, so it is ours to drop with the rest. + "download-outcome.json", +]); + async function discardPrefetchDir( videoDir: string, onLog: (line: string) => void, ): Promise<boolean> { - let entries: string[]; + let entries: Dirent[]; try { - const files = await readVideoFiles(videoDir); - if (isVideoDownloaded(files)) return false; - entries = files.entries; + entries = await readdir(videoDir, { withFileTypes: true }); } catch { - // No directory at all (the legacy single-call path never made one), or it - // is unreadable. Either way there is nothing of ours to remove. + // No directory at all — the legacy single-call path never made one — or it + // is unreadable. Either way there is nothing of ours to remove, and a + // `rm -rf` on a path we could not list is the last thing to do about it. return false; } - if ( - entries.includes(CLIPS_DIR_NAME) || - entries.includes(SAVED_VIDEO_POINTER_FILENAME) - ) { - return false; + for (const entry of entries) { + // A DIRECTORY IS NEVER OURS. `clips/` is the one that exists today (media + // another tool asked this editor for, invisible to every video-dir + // enumerator); the rule is shaped so the next one needs no edit here. + if (!entry.isFile() || !PREFETCH_OWN_FILES.has(entry.name)) return false; } try { await rm(videoDir, { recursive: true, force: true }); @@ -391,6 +408,13 @@ async function discardPrefetchDir( } } +// Exported for the unit test ONLY. The rule above is a delete, and the +// difference between the allow-list and the deny-list it replaced is invisible +// to every end-to-end path that does not happen to have the right file on disk +// — which is exactly the shape of bug that reaches production. See +// downloadOneManaged.test.ts. +export const __discardPrefetchDirForTest = discardPrefetchDir; + // The one-invocation runner moved to ytdlp/runOneYtdlp.ts (the clip-window // fetch needs the same log tee, stderr tail and archive scrape). This wrapper // keeps the ManagedDownloadOpts-shaped call sites below unchanged. @@ -428,7 +452,6 @@ function runOneYtdlp( async function fetchLiveChatOnly( opts: ManagedDownloadOpts, channelDir: string, - videoDir: string, cookies: string | undefined, ): Promise<AttemptOutcome> { return runOneYtdlp(opts, channelDir, [ @@ -715,57 +738,24 @@ async function runManagedDownload( opts.onLog( `Skipping ${canonicalId}: ${decision.reason} [filter=${decision.filter}]\n`, ); - // A title-filter rejection feeds the metadata-scan store, so this video is - // settled from here on WITHOUT a second metadata fetch. The store is the - // one place a settled verdict is derived from; the outcome below records - // only what happened, exactly as skip-live's does. A metadata-less skip - // (the fail-closed branch) writes nothing here — there is nothing to - // store, and it must stay retryable. - if (decision.filter === "titleFilter" && metadata) { - const entry = { - title: metadata.title ?? "", - description: metadata.description ?? "", - uploadDate: metadata.upload_date ?? "", - ...(metadata.live_status ? { liveStatus: metadata.live_status } : {}), - ...(typeof metadata.duration === "number" - ? { duration: metadata.duration } - : {}), - scannedAt: new Date().toISOString(), - }; - // A BATCH COLLECTS; A SINGLE VIDEO WRITES. The store is channel-level, - // so writing it from inside a per-video loop is a full load + stringify - // + rename per rejection — on a filtered channel that is one rewrite of - // the whole file per non-matching video. The batch runner passes a - // collector and writes once when it is done; a one-off download (no - // collector) writes for itself, because nothing else will. - if (opts.onFilterRejected) { - opts.onFilterRejected(canonicalId, entry); - } else { - try { - await upsertMetadataScan( - opts.paths, - opts.channelSlug, - { entries: { [canonicalId]: entry } }, - new Date().toISOString(), - ); - } catch (err) { - opts.onLog( - `Failed to record the metadata scan entry: ${(err as Error).message}\n`, - ); - } - } - } - // ── CHAT ONLY ──────────────────────────────────────────────────── + // ── CHAT ONLY, AND IT RUNS BEFORE THE STORE IS WRITTEN ─────────────── // // The operator asked for this livestream's chat and not its media, so // this is the one rejection that FETCHES something — and the one that - // keeps its prefetch directory, deliberately. A chat-only dir holds + // keeps its prefetch directory. A chat-only dir holds the // metadata.info.json (without it buildIndex never sees the video and the - // chat is published nowhere) plus transcript.live_chat.json, and nothing - // else. Every "is it downloaded?" predicate tests whisper, an English VTT - // or audio, so it still answers no — which is exactly right: this video - // is a chat track, not a download. + // chat is published nowhere) plus transcript.live_chat.json and its + // normalized live_chat.cues.json, beside the outcome sidecar and the log + // every managed download leaves. Every "is it downloaded?" predicate + // tests whisper, an English VTT or audio, so it still answers no — which + // is right: this is a chat track, not a download. + // + // IT RUNS FIRST because its result belongs in the scan entry below. A + // stream whose chat replay is OFF has to be remembered somewhere, or the + // id sits in chatOnlyPending forever and every runner restart re-prefetches + // it and re-runs a pass that will never return anything. let chatOnlyFetched = false; + let chatOnlyUnavailable = false; if (decision.chatOnly && !opts.signal.aborted) { opts.onLog( `Fetching the live chat for ${canonicalId} (media skipped by the download filter).\n`, @@ -773,7 +763,6 @@ async function runManagedDownload( const chatRes = await fetchLiveChatOnly( opts, channelDir, - videoDir, alwaysCookies(cookiePolicy) ?? (prefetchNeededCookies ? authRetryCookies(cookiePolicy) : undefined), ); @@ -787,13 +776,13 @@ async function runManagedDownload( ? undefined : trimError(chatRes.stderrTail), }); - // The CUES sidecar, not just the raw file: buildIndex prefers - // live_chat.cues.json when it is fresh, and normalizing here means a - // chat-only video is index-ready the moment it lands rather than - // waiting for the corpus-wide normalize pass. if (attemptSucceeded(chatRes.exitCode)) { chatOnlyFetched = await hasLiveChatOnDisk(videoDir); if (chatOnlyFetched) { + // The CUES sidecar, not just the raw file: buildIndex prefers + // live_chat.cues.json when it is fresh, so normalizing here makes + // the video index-ready the moment it lands instead of waiting for + // the corpus-wide normalize pass. try { await normalizeLiveChat({ videoDir, @@ -809,11 +798,59 @@ async function runManagedDownload( ); } } else { - // The source served no chat for this stream. Nothing was fetched, - // so the directory is not a corpus member and goes the way every - // other rejection's does. + // A CLEAN PASS THAT FOUND NOTHING IS AN ANSWER, not a retry. Chat + // replay was off for this stream, or the source has since dropped + // it; asking again tomorrow gets the same nothing. Recorded on the + // scan entry below so the video leaves chatOnlyPending for good and + // settles as the ordinary filtered-out livestream it is. + // + // A FAILED pass (non-zero exit, a cooldown, an abort) is NOT this: + // it leaves the flag alone and the id stays pending, which is the + // whole reason this is gated on the exit code. + chatOnlyUnavailable = true; + opts.onLog( + `No live chat was available for ${canonicalId}; recording that so it is not asked for again.\n`, + ); + } + } + } + // A title-filter rejection feeds the metadata-scan store, so this video is + // settled from here on WITHOUT a second metadata fetch. The store is the + // one place a settled verdict is derived from; the outcome below records + // only what happened, exactly as skip-live's does. A metadata-less skip + // (the fail-closed branch) writes nothing here — there is nothing to + // store, and it must stay retryable. + if (decision.filter === "titleFilter" && metadata) { + const entry = { + title: metadata.title ?? "", + description: metadata.description ?? "", + uploadDate: metadata.upload_date ?? "", + ...(metadata.live_status ? { liveStatus: metadata.live_status } : {}), + ...(typeof metadata.duration === "number" + ? { duration: metadata.duration } + : {}), + ...(chatOnlyUnavailable ? { noLiveChat: true } : {}), + scannedAt: new Date().toISOString(), + }; + // A BATCH COLLECTS; A SINGLE VIDEO WRITES. The store is channel-level, + // so writing it from inside a per-video loop is a full load + stringify + // + rename per rejection — on a filtered channel that is one rewrite of + // the whole file per non-matching video. The batch runner passes a + // collector and writes once when it is done; a one-off download (no + // collector) writes for itself, because nothing else will. + if (opts.onFilterRejected) { + opts.onFilterRejected(canonicalId, entry); + } else { + try { + await upsertMetadataScan( + opts.paths, + opts.channelSlug, + { entries: { [canonicalId]: entry } }, + new Date().toISOString(), + ); + } catch (err) { opts.onLog( - `No live chat was available for ${canonicalId}; nothing to keep.\n`, + `Failed to record the metadata scan entry: ${(err as Error).message}\n`, ); } } @@ -830,9 +867,13 @@ async function runManagedDownload( }; // The operator's own rejection takes its prefetch directory with it — // see discardPrefetchDir for why a metadata-only dir is not a neutral - // leftover. Every other filter's skip is RETRYABLE and its outcome - // sidecar is what `skippedByFilter` is derived from, so only this one - // discards. + // leftover, and why the rule is an allow-list. Every other filter's skip + // is RETRYABLE and its outcome sidecar is what `skippedByFilter` derives + // from, so only this one discards. + // + // A fetched chat is not "mine to delete" either way — the allow-list + // refuses the directory the moment transcript.live_chat.json is in it — + // so this condition is a shortcut past a readdir, not the safety rail. const discarded = decision.filter === "titleFilter" && metadata && !chatOnlyFetched ? await discardPrefetchDir(videoDir, opts.onLog) diff --git a/common/ytdlp/runYtdlp.ts b/common/ytdlp/runYtdlp.ts @@ -1593,7 +1593,23 @@ async function syncFullSweep(opts: RunYtdlpOpts): Promise<void> { .filter((e) => e.isDirectory()) .map((e) => e.name), ); - const sets = deriveChannelSets({ roster, listedIds, onDiskIds }); + // THE SETTLED SET IS SUBTRACTED, and it is not a nicety. A title-filter + // rejection leaves no directory now, so without this a filtered-out video + // that has since left the listing reads as `missingNeverFetched` — "we were + // told about this, never got it, and it is gone" — and the sweep tells the + // operator to attempt a direct-link recovery of the very video they + // configured us not to fetch. + const sweepSettledIds = await settledByTitleFilterIds( + opts.paths, + opts.channelSlug, + opts.channelConfig, + ); + const sets = deriveChannelSets({ + roster, + listedIds, + onDiskIds, + settledIds: sweepSettledIds, + }); maybeMissing = sets.missingDownloaded; await writeMaybeMissing(opts.paths, opts.channelSlug, { checkedAt: now, diff --git a/editor/e2e/title-filter.spec.ts b/editor/e2e/title-filter.spec.ts @@ -268,32 +268,43 @@ test("a rejection never deletes a directory that holds somebody else's data", as // every video-dir enumerator — so a discard that walked past it would delete // bytes nothing else would ever mention again. // - // (A DOWNLOADED video cannot be tested through this button: "Download videos" - // prefilters on the destination file, so one that already has a - // transcript.en.vtt is never attempted at all. That refusal is the - // isVideoDownloaded branch, covered by the unit-level guard.) + // The second video carries DOWNLOADED MEDIA. On a youtube-handling channel + // the destination is transcript.en.vtt, so an audio.mp3 does not prefilter + // the video away — the run attempts it, the filter rejects it, and the bytes + // have to survive. A video downloaded before the filter was written is + // downloaded: that is a fact, not a preference. await resetData("title-filter-channel"); - const kept = PLAIN[0]; - await mkdir(resolvePath(`${ROOT}/data/${kept}/clips`), { recursive: true }); + const [withClips, withAudio] = PLAIN; + await mkdir(resolvePath(`${ROOT}/data/${withClips}/clips`), { + recursive: true, + }); await writeFile( - resolvePath(`${ROOT}/data/${kept}/clips/0.00-30.00.mp3`), + resolvePath(`${ROOT}/data/${withClips}/clips/0.00-30.00.mp3`), "fake clip audio\n", ); + await mkdir(resolvePath(`${ROOT}/data/${withAudio}`), { recursive: true }); + await writeFile( + resolvePath(`${ROOT}/data/${withAudio}/audio.mp3`), + "fake audio bytes\n", + ); await generateReport(page, CHANNEL); await download(page); - expect(await pathExists(`${ROOT}/data/${kept}/clips/0.00-30.00.mp3`)).toBe( - true, - ); - // The directory stayed, so the retryable outcome sidecar is written for it - // exactly as it was before this rule existed. - const outcome = await readJson<{ status: string }>( - `${ROOT}/data/${kept}/download-outcome.json`, - ); - expect(outcome.status).toBe("skipped-filtered"); - // The other rejections, which had nothing of their own, are still gone. - expect(await pathExists(`${ROOT}/data/${PLAIN[1]}`)).toBe(false); + for (const [id, file] of [ + [withClips, "clips/0.00-30.00.mp3"], + [withAudio, "audio.mp3"], + ] as const) { + expect(await pathExists(`${ROOT}/data/${id}/${file}`)).toBe(true); + // The directory stayed, so the retryable outcome sidecar is written for it + // exactly as it was before this rule existed. + const outcome = await readJson<{ status: string }>( + `${ROOT}/data/${id}/download-outcome.json`, + ); + expect(outcome.status).toBe("skipped-filtered"); + } + // And the rejection that had nothing of its own is still gone. + expect(await pathExists(`${ROOT}/data/${LIVE}`)).toBe(false); }); test("changing the filter re-decides the channel with no rescan", async ({ diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -4171,9 +4171,13 @@ rejects. holding a metadata.info.json is admitted to the LMDB index and the published site by `buildIndex.ts:309-315`, transcript or not, while `deriveChannelSets` reads its NAME as "ever fetched". The scan creates none of them for exactly - those two reasons; the downloader now creates none either. It fires ONLY for a - video the filter needed a live prefetch for (one the scan has not read) — a - settled video is never invoked at all. + those two reasons; the downloader now creates none either. It fires when a + rejection reaches the DOWNLOADER at all — a video the filter needed a live + prefetch for, i.e. one the scan has not read. The two batch paths never invoke + yt-dlp for a settled video (`selectDownloadableUrls` counts it an archived + hit), but `importVideoAction` reaches `downloadOneManaged` by URL and bypasses + both, so "a settled video is never invoked" is true of a sync and a + download-missing and NOT of a hand-pasted link. - **The count does not move, and the reason is that it never came from there.** `buckets.skippedByTitleFilter` is built at `channelSnapshot.ts:1378-1383` from `settledIds` ← `settledIdsFrom(metadataScanStore, config)` at `:927`. Nothing @@ -4189,10 +4193,27 @@ rejects. predicate is how the settled short-circuit and `totals.videos` end up disagreeing about what an artifact is, and the five call sites read better asking "does this dir hold anything we fetched?". -- **It refuses rather than guesses.** Anything `isVideoDownloaded` calls - downloaded, a `clips/` window dir, or a `saved-video.json` pointer and the - directory stands. The returned `DownloadOutcomeRecord` is unchanged either way, - so `runYtdlp.ts:949` and the runner's `unitStatus` check need no edit. +- **IT IS AN ALLOW-LIST, AND THE DENY-LIST IT REPLACED WAS A BUG.** This is a + recursive delete, so the question it must answer is "is EVERYTHING in here + mine?". The first version named `isVideoDownloaded`, `clips/` and + `saved-video.json` — and would therefore have deleted a chat-only corpus + member (`transcript.live_chat.json` + `live_chat.cues.json`), a + `transcript.es.vtt`, a resumable `audio.mp3.part`, a `diarization.json` or a + `digest.json`. Reachable, not hypothetical: the chat-only branch calls this + whenever its chat pass came back with no file, so a chat re-fetch that failed + or was aborted would have deleted the chat that was already there. The rule is + now `PREFETCH_OWN_FILES` = `{metadata.info.json, download.log, + download-outcome.json}`, files only — **a directory is never ours** — and + anything else keeps the dir. `ytdlp/downloadOneManaged.test.ts` runs one case + per protected shape. The returned `DownloadOutcomeRecord` is unchanged either + way, so `runYtdlp.ts:949` and the runner's `unitStatus` check need no edit. +- **`deriveChannelSets` takes the settled set**, and both call sites pass it + (`channelSnapshot.ts`, and `syncFullSweep` in `runYtdlp.ts`). Both sets it + derives are defined by the ABSENCE of a directory, and a settled video has none + now — so without this a filtered-out video that later leaves the listing reads + as `missingNeverFetched`, the loudest alarm this system raises, pointed at a + video the operator asked us not to fetch. `undownloaded` is subtracted for the + same reason: settled work is not work. - **THE DOWNLOAD LANE AUTO-DISPATCHES THE SCAN.** `METADATA_SCAN_OPERATION` now declares `runner: "download"`, and the runner's `next()` runs a pre-pick before any video: `pickMetadataScanChannel` (pure, exported, unit-tested) takes the @@ -4207,6 +4228,14 @@ rejects. unit does, so a scan and a download never hit one source at once. `AutoRunnerInFlight` gains `note?`, and `InFlightList` renders the note (and links the CHANNEL) instead of a dead video link. +- **A FOCUS HOLDS THE SCAN TOO.** The compiled priority tree holds every + download pick by strict descent, and the pre-pick does not go through that + tree — so without this, "only jeralyzer is moving" would be false the moment + another channel had titles to read, on the same network the focus is trying to + have to itself. `pickMetadataScanChannel` takes the `focusSummary` the lane + already computes for its log line: while `holding` (i.e. `focusPending > 0`) + only a focused channel is offered a scan; once the focus is exhausted the lane + is free and so is the scan. - **ONE SCAN PER CHANNEL PER RUNNER, deliberately.** A scan that left the backlog above zero was stopped by something — a soft block, a cooldown, a batch that ended early. `scannedThisRun` means this runner does not re-offer it three @@ -4241,9 +4270,21 @@ rejects. Then `normalizeLiveChat` writes `live_chat.cues.json` on the spot. No archive line — an archive id means "downloaded" to the sync walk and to `verifyTranscripts`. -- **A chat pass that fetched nothing keeps no directory.** A stream with chat - replay off makes yt-dlp exit 0 and write nothing; `hasLiveChatOnDisk` is what - sends that dir to `discardPrefetchDir` like any other rejection's. +- **A CLEAN PASS THAT FOUND NOTHING IS AN ANSWER, not a retry.** A stream with + chat replay off makes yt-dlp exit 0 and write nothing. The directory goes like + any other rejection's, and — because nothing on disk can then remember it — + the scan entry carries **`noLiveChat: true`**, which `chatOnlyIdsFrom` drops. + Without it the id sits in `chatOnlyPending` forever: `markCompleted` retires it + for the SESSION only, so every runner restart re-prefetched it to run a pass + that will never return anything. A pass that FAILED (non-zero exit, cooldown, + abort) deliberately does not set it — that one stays retryable — which is why + the chat pass now runs BEFORE the store is written rather than after. +- **The chat-only verdict is decided off the FILE, not the mode**, in the + snapshot's settled branch: a dir holding `transcript.live_chat.json` is a + `chatOnly` member however `rejectedLivestreams` currently reads. Asking + `chatOnlyIds` there made the same directory a settled STUB the moment the mode + flipped back — and `settledOnDisk` is subtracted from `totals.videos`, so the + site kept serving a video the report had stopped counting. - **`DownloadOutcomeStatus` gains `"chat-only"`** and `DownloadAttemptKind` gains `"live-chat-only"`. The batch counts it as a SKIP (`okCount` means videos downloaded); the runner counts it as a SUCCESS (the unit did the work it was