Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 157f3c9e4562ef4a12b6f717a5a36e4491bb21bf
parent a8fd648a462d2a175ded12b8531ad223aebc0b2e
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Thu,  6 Aug 2026 14:17:29 -0400

Keep a roster, so no listed video is ever lost

The only record that a video belonged to a channel was the `playlist`
file, which a full sweep overwrites wholesale. A video that appeared in a
listing but was never downloaded has no data/<id>/ dir and no
metadata.info.json, so `playlist` was the ONLY place its URL lived — when
it dropped out of the listing it was erased with no record it had ever
existed, and no way to attempt a direct-link recovery. Downstream,
undownloadedIds walks that same file, so the auto-download runner's whole
work-list went with it.

Part A's guard only refused a COMPLETELY empty listing. A fetch that
returned 100 of 10,795 sailed through and destroyed 10,695 entries.

channels/<slug>/roster.json now records every id ever seen with the URL
it was seen at. mergeRoster has no removal branch: only a full, accepted
enumeration may compute a "missing" set, and every other writer — the
paged sync, the quick check, an import — may only add. Seeding runs from
playlist ∪ data/ BEFORE anything is overwritten, so the migration is
itself the protection; one live channel has 617 videos and no listing
file at all. Disk-seeded URLs come from availability.json (137 bytes) and
fall back to a capped scan of metadata.info.json rather than reading
550 KB per video.

acceptListing() gates both full-listing writers — the sweep and
store-playlist, which overwrote unconditionally and was the most direct
instance of the bug. A drop past max(25, previous x pct) is suspect: the
playlist and maybe-missing.json are left alone and lastFullSweepAt is not
stamped, so the next sync retries rather than waiting out the cadence. A
second enumeration reporting a similar count confirms it — a real mass
deletion repeats, a blip does not — so it needs no UI and costs at most
one cadence period. The absolute floor matters: pure percentage trips on
a 55-video channel losing six.

deriveChannelSets() is now the single definition of every set.
maybe-missing keeps its exact shape and meaning (derived from the DISK
set, so a stale roster can never un-flag a deletion), and
missingNeverFetched is the category the app could not express: knew about
it, never fetched it, now gone. It gets a Diagnostics section with the
stored URL and a "Try downloading anyway" action, and an /actionable
entry — the roster keeps the URL precisely so that is possible.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

Diffstat:
Acommon/controller/acceptListing.test.ts | 171+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/controller/acceptListing.ts | 113+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/controller/channelSets.test.ts | 113+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/controller/channelSets.ts | 80+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/controller/channelSnapshot.ts | 74++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++----------
Mcommon/controller/quickAvailabilityCheck.ts | 25+++++++++++++++++++++++++
Acommon/controller/rosterStore.test.ts | 298+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/controller/rosterStore.ts | 379+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/lib/settings.ts | 19+++++++++++++++++++
Acommon/lib/videoId.ts | 64++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/ytdlp/runYtdlp.ts | 256++++++++++++++++++++++++++++++++++++++++++++++++++-----------------------------
Meditor/CHANGELOG.md | 3+++
Meditor/app/actionable/lib/loadActionable.ts | 18++++++++++++++++++
Meditor/app/actionable/page.tsx | 23+++++++++++++++++++++++
Meditor/app/channels/[slug]/components/stages/DiagnosticsStage.tsx | 88++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-
Meditor/app/channels/[slug]/page.tsx | 1+
Meditor/app/channels/[slug]/pipelineActions.ts | 15+++++++++++++++
Meditor/app/settings/actions.ts | 4++++
Meditor/app/settings/components/SettingsForm.tsx | 9+++++++++
Meditor/e2e/sync-deep.spec.ts | 223++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-
20 files changed, 1868 insertions(+), 108 deletions(-)

diff --git a/common/controller/acceptListing.test.ts b/common/controller/acceptListing.test.ts @@ -0,0 +1,171 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { acceptListing, SHRINK_ABS_FLOOR } from "./acceptListing"; +import type { RosterSweep } from "./rosterStore"; + +// Run with: +// pnpm --filter yt-dlp-transcript-common exec tsx --test common/controller/acceptListing.test.ts + +const PCT = 10; + +function decide(input: { + listedCount: number; + previousListedCount?: number | null; + lastSweep?: RosterSweep | null; + shrinkGuardPercent?: number; +}) { + return acceptListing({ + listedCount: input.listedCount, + previousListedCount: input.previousListedCount ?? null, + lastSweep: input.lastSweep ?? null, + shrinkGuardPercent: input.shrinkGuardPercent ?? PCT, + }); +} + +function suspectAt(listedCount: number): RosterSweep { + return { at: "2026-08-05T00:00:00.000Z", listedCount, verdict: "shrink-suspect" }; +} + +test("an empty listing is rejected — the Part A behaviour, kept", () => { + const d = decide({ listedCount: 0, previousListedCount: 10795 }); + assert.equal(d.accept, false); + assert.equal(d.verdict, "empty"); +}); + +test("an empty listing is rejected even with no previous listing and the guard off", () => { + const d = decide({ + listedCount: 0, + previousListedCount: null, + shrinkGuardPercent: 0, + }); + assert.equal(d.accept, false); + assert.equal(d.verdict, "empty"); +}); + +test("the first listing for a channel is accepted — nothing to compare against", () => { + const d = decide({ listedCount: 10795, previousListedCount: 0 }); + assert.equal(d.accept, true); + assert.equal(d.verdict, "ok"); +}); + +test("a 1% drop is accepted", () => { + const d = decide({ listedCount: 10687, previousListedCount: 10795 }); + assert.equal(d.accept, true); + assert.equal(d.verdict, "ok"); +}); + +test("growth is always accepted", () => { + const d = decide({ listedCount: 20000, previousListedCount: 10795 }); + assert.equal(d.accept, true); + assert.equal(d.verdict, "ok"); +}); + +test("the reported bug: 100 of 10,795 is suspect, not acted on", () => { + const d = decide({ listedCount: 100, previousListedCount: 10795 }); + assert.equal(d.accept, false); + assert.equal(d.verdict, "shrink-suspect"); + assert.match(d.reason, /100 entries/); + assert.match(d.reason, /10795/); +}); + +test("a 40% drop is suspect, then confirmed by a second similar reading", () => { + const first = decide({ listedCount: 60, previousListedCount: 100 }); + assert.equal(first.accept, false); + assert.equal(first.verdict, "shrink-suspect"); + + // The next enumeration agrees. A real mass deletion repeats. + const second = decide({ + listedCount: 60, + previousListedCount: 100, + lastSweep: suspectAt(60), + }); + assert.equal(second.accept, true); + assert.equal(second.verdict, "shrink-confirmed"); +}); + +test("a second reading that disagrees stays suspect", () => { + const d = decide({ + listedCount: 60, + previousListedCount: 10795, + // Last time the fetch said 5,000; this time 60. Two different failures are + // not a confirmation. + lastSweep: suspectAt(5000), + }); + assert.equal(d.accept, false); + assert.equal(d.verdict, "shrink-suspect"); +}); + +test("confirmation tolerates a small wobble between the two readings", () => { + const d = decide({ + listedCount: 6000, + previousListedCount: 10795, + lastSweep: suspectAt(6100), + }); + assert.equal(d.accept, true); + assert.equal(d.verdict, "shrink-confirmed"); +}); + +test("only a previous SUSPECT confirms — an ok or empty reading does not", () => { + for (const verdict of ["ok", "empty", "shrink-confirmed"] as const) { + const d = decide({ + listedCount: 60, + previousListedCount: 100, + lastSweep: { at: "2026-08-05T00:00:00.000Z", listedCount: 60, verdict }, + }); + assert.equal(d.accept, false, verdict); + assert.equal(d.verdict, "shrink-suspect", verdict); + } +}); + +test("the absolute floor keeps ordinary churn on a small channel from tripping", () => { + // The live corpus has a 55-video channel; losing six of them is a 10.9% drop + // and entirely unremarkable. A pure percentage would have flagged it. + const d = decide({ listedCount: 49, previousListedCount: 55 }); + assert.equal(d.accept, true); + assert.equal(d.verdict, "ok"); +}); + +test("the floor is exactly SHRINK_ABS_FLOOR entries, inclusive", () => { + const prev = 100; + const atFloor = decide({ listedCount: prev - SHRINK_ABS_FLOOR, previousListedCount: prev }); + assert.equal(atFloor.accept, true, "a drop of exactly the floor is allowed"); + const overFloor = decide({ + listedCount: prev - SHRINK_ABS_FLOOR - 1, + previousListedCount: prev, + }); + assert.equal(overFloor.accept, false); +}); + +test("above the floor the percentage takes over", () => { + // 10% of 10,795 is 1,079 — far above the floor, so a 1,000-entry drop passes. + assert.equal(decide({ listedCount: 9795, previousListedCount: 10795 }).accept, true); + assert.equal(decide({ listedCount: 9000, previousListedCount: 10795 }).accept, false); +}); + +test("percent 0 disables the shrink test but not the empty rejection", () => { + const d = decide({ + listedCount: 1, + previousListedCount: 10795, + shrinkGuardPercent: 0, + }); + assert.equal(d.accept, true); + assert.equal(d.verdict, "ok"); + assert.equal( + decide({ listedCount: 0, previousListedCount: 10795, shrinkGuardPercent: 0 }) + .accept, + false, + ); +}); + +test("a stricter percent catches what the default lets through", () => { + // 500 of 10,795 is under the default 10% but over 1%. + assert.equal(decide({ listedCount: 10295, previousListedCount: 10795 }).accept, true); + assert.equal( + decide({ + listedCount: 10295, + previousListedCount: 10795, + shrinkGuardPercent: 1, + }).accept, + false, + ); +}); diff --git a/common/controller/acceptListing.ts b/common/controller/acceptListing.ts @@ -0,0 +1,113 @@ +import type { RosterSweep, SweepVerdict } from "./rosterStore"; + +// The gate every full enumeration passes through before it is allowed to +// replace what we know about a channel. +// +// With the roster in place a bad fetch is already non-destructive — nothing is +// ever removed from roster.json — so this guard's remaining job is narrow and +// worth stating exactly: don't truncate the stored `playlist`, and don't flag +// thousands of videos as missing, on the strength of one suspicious read. +// +// Part A only rejected a COMPLETELY empty listing. A fetch that returns 100 of +// 10,795 entries — a throttled connection, a partial page walk, an expiring +// cookie — sailed through and destroyed 10,695 entries. +// +// Two-observation confirmation is the whole design: a real mass deletion +// repeats, a transient blip does not. It needs no UI, no manual override, and +// costs at most one cadence period to accept a genuine deletion. + +// The floor matters as much as the percentage. A pure percentage trips on small +// channels for ordinary churn: the live corpus has a channel with 55 listed +// videos, where losing six is a 10.9% drop and entirely unremarkable. +export const SHRINK_ABS_FLOOR = 25; + +export type ListingDecision = { + // Whether the caller may act on this listing: rewrite `playlist`, rewrite + // maybe-missing.json, and stamp lastFullSweepAt. + accept: boolean; + verdict: SweepVerdict; + // One line for the job log, naming the numbers that drove the decision. + reason: string; +}; + +// How much shrink is tolerated against a reference count. Shared by the +// accept test and the "is this the same reading as last time" test so the two +// can't drift apart. +function tolerance(reference: number, percent: number): number { + return Math.max(SHRINK_ABS_FLOOR, Math.floor((reference * percent) / 100)); +} + +export function acceptListing(input: { + // Entry count of the enumeration just fetched. + listedCount: number; + // Size of the last ACCEPTED listing — in practice the stored `playlist` + // file's length. 0 or null means there is nothing to compare against. + previousListedCount: number | null; + // The previous OBSERVATION, accepted or not (roster.lastSweep). This is what + // makes the second reading confirmatory rather than just another suspect. + lastSweep: RosterSweep | null; + // syncScheduler.fullSweepShrinkGuardPercent. 0 disables the shrink test (the + // empty-listing rejection is not optional). + shrinkGuardPercent: number; +}): ListingDecision { + const { listedCount, previousListedCount, lastSweep, shrinkGuardPercent } = + input; + + // An empty listing is never trustworthy enough to act on: it is what a + // transient upstream failure, a cookie expiry and a clean 101 exit all look + // like from here. + if (listedCount <= 0) { + return { + accept: false, + verdict: "empty", + reason: + "the channel listing came back empty — leaving the stored playlist and missing-video flags untouched, and not counting this as a sweep", + }; + } + + const previous = previousListedCount ?? 0; + if (previous <= 0) { + return { + accept: true, + verdict: "ok", + reason: `${listedCount} listed, no previous listing to compare against`, + }; + } + if (shrinkGuardPercent <= 0) { + return { + accept: true, + verdict: "ok", + reason: `${listedCount} listed (shrink guard off)`, + }; + } + + const dropped = previous - listedCount; + const allowed = tolerance(previous, shrinkGuardPercent); + if (dropped <= allowed) { + return { + accept: true, + verdict: "ok", + reason: `${listedCount} listed, was ${previous}`, + }; + } + + // A big drop. Accept it only on the SECOND consecutive reading that says + // roughly the same thing. + if (lastSweep && lastSweep.verdict === "shrink-suspect") { + const reference = Math.max(listedCount, lastSweep.listedCount); + const delta = Math.abs(listedCount - lastSweep.listedCount); + if (delta <= tolerance(reference, shrinkGuardPercent)) { + return { + accept: true, + verdict: "shrink-confirmed", + reason: `${listedCount} listed, down from ${previous} — confirmed by a second enumeration (previous reading ${lastSweep.listedCount}), accepting the drop`, + }; + } + } + + return { + accept: false, + verdict: "shrink-suspect", + reason: `the listing came back with ${listedCount} entries, down ${dropped} from ${previous} (more than the ${allowed} this guard allows) — leaving the stored playlist and missing-video flags untouched, and not counting this as a sweep. A real deletion repeats: if the next enumeration agrees, it will be accepted`, + }; +} diff --git a/common/controller/channelSets.test.ts b/common/controller/channelSets.test.ts @@ -0,0 +1,113 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { deriveChannelSets } from "./channelSets"; +import { emptyRoster, mergeRoster, type Roster } from "./rosterStore"; + +// Run with: +// pnpm --filter yt-dlp-transcript-common exec tsx --test common/controller/channelSets.test.ts + +const NOW = "2026-08-06T00:00:00.000Z"; + +function rosterOf(ids: ReadonlyArray<string>): Roster { + return mergeRoster( + emptyRoster(), + ids.map((id) => ({ id, url: `https://www.youtube.com/watch?v=${id}` })), + NOW, + "listing", + ); +} + +function derive(input: { + roster: ReadonlyArray<string>; + listed: ReadonlyArray<string>; + onDisk: ReadonlyArray<string>; +}) { + return deriveChannelSets({ + roster: rosterOf(input.roster), + listedIds: new Set(input.listed), + onDiskIds: new Set(input.onDisk), + }); +} + +test("a video in the roster and on disk but not listed is missingDownloaded only", () => { + const sets = derive({ roster: ["a"], listed: [], onDisk: ["a"] }); + assert.deepEqual(sets.missingDownloaded, ["a"]); + assert.deepEqual(sets.missingNeverFetched, []); + assert.deepEqual(sets.orphaned, []); + assert.deepEqual(sets.undownloaded, []); + assert.deepEqual(sets.missing, ["a"]); +}); + +test("a roster-only video is missingNeverFetched only — the category that did not exist before", () => { + const sets = derive({ roster: ["a"], listed: [], onDisk: [] }); + assert.deepEqual(sets.missingNeverFetched, ["a"]); + assert.deepEqual(sets.missingDownloaded, []); + assert.deepEqual(sets.orphaned, []); + assert.deepEqual(sets.missing, ["a"]); +}); + +test("a disk-only video not in the roster is orphaned", () => { + const sets = derive({ roster: [], listed: [], onDisk: ["a"] }); + assert.deepEqual(sets.orphaned, ["a"]); + // Still maybe-missing: it IS on disk and the listing does not have it, which + // is exactly what maybe-missing.json has always meant. A roster that hasn't + // caught up must never be able to un-flag a deletion. + assert.deepEqual(sets.missingDownloaded, ["a"]); + assert.deepEqual(sets.missingNeverFetched, []); +}); + +test("a listed video on disk is in neither missing set", () => { + const sets = derive({ roster: ["a"], listed: ["a"], onDisk: ["a"] }); + assert.deepEqual(sets.listed, ["a"]); + assert.deepEqual(sets.undownloaded, []); + assert.deepEqual(sets.missing, []); + assert.deepEqual(sets.missingDownloaded, []); + assert.deepEqual(sets.missingNeverFetched, []); + assert.deepEqual(sets.orphaned, []); +}); + +test("a listed video with no dir is undownloaded — the work-list, and not missing", () => { + const sets = derive({ roster: ["a"], listed: ["a"], onDisk: [] }); + assert.deepEqual(sets.undownloaded, ["a"]); + assert.deepEqual(sets.missing, []); +}); + +test("every set at once, on a channel with one of each", () => { + const sets = derive({ + // gone-fetched left the listing but we have it; gone-never was listed once + // and never downloaded; orphan was imported by hand before the roster. + roster: ["listed-have", "listed-want", "gone-fetched", "gone-never"], + listed: ["listed-have", "listed-want"], + onDisk: ["listed-have", "gone-fetched", "orphan"], + }); + assert.deepEqual(sets.listed, ["listed-have", "listed-want"]); + assert.deepEqual(sets.undownloaded, ["listed-want"]); + assert.deepEqual(sets.missingDownloaded, ["gone-fetched", "orphan"]); + assert.deepEqual(sets.missingNeverFetched, ["gone-never"]); + assert.deepEqual(sets.orphaned, ["orphan"]); + assert.deepEqual(sets.missing, ["gone-fetched", "gone-never", "orphan"]); +}); + +test("missingDownloaded reproduces the old on-disk-minus-listing diff exactly", () => { + // The pre-roster rule, which maybe-missing.json's readers still depend on. + const onDisk = ["a", "b", "c", "d"]; + const listed = ["b", "d", "e"]; + const expected = onDisk.filter((id) => !listed.includes(id)).sort(); + // Whatever the roster says — full, partial, or empty — the answer is the same. + for (const roster of [onDisk, listed, [], ["a"]]) { + const sets = derive({ roster, listed, onDisk }); + assert.deepEqual(sets.missingDownloaded, expected); + } +}); + +test("an empty channel derives empty sets rather than throwing", () => { + const sets = derive({ roster: [], listed: [], onDisk: [] }); + assert.deepEqual(sets, { + listed: [], + undownloaded: [], + missing: [], + missingDownloaded: [], + missingNeverFetched: [], + orphaned: [], + }); +}); diff --git a/common/controller/channelSets.ts b/common/controller/channelSets.ts @@ -0,0 +1,80 @@ +import type { Roster } from "./rosterStore"; + +// The single place "missing" is defined. Pure, no fs, so every set can be +// unit-tested against a hand-built triple instead of a seeded temp corpus. +// +// Three inputs, and each answers a different question: +// roster — every id the channel has EVER been seen to contain (durable) +// listedIds — what the latest ACCEPTED enumeration contained (current truth) +// onDiskIds — data/<id>/ dir names (what we actually fetched) +// +// Before the roster existed only the last two were available, so "missing" +// could only ever mean "on disk but no longer listed" — a video that was listed +// and never fetched had no dir, could not be flagged, and vanished along with +// its URL when the sweep overwrote `playlist`. missingNeverFetched is the +// category that hole made inexpressible. + +export type ChannelSets = { + // The current listing, sorted. + listed: string[]; + // Listed but nothing on disk — the download work-list. Replaces walking the + // `playlist` file directly, so it can't evaporate with a bad enumeration. + undownloaded: string[]; + // Everything we know of that the current listing does not contain. + missing: string[]; + // missing ∩ on disk. EXACTLY today's maybe-missing set (see the note below). + missingDownloaded: string[]; + // missing, never fetched. We knew about this video, never downloaded it, and + // it is gone — for an archival tool, the highest-value signal in the system. + // The roster holds its URL precisely so a recovery attempt is still possible. + missingNeverFetched: string[]; + // On disk but not in the roster at all: a one-off import that predates the + // roster, or a dir dropped in by hand. Diagnostic only. + orphaned: string[]; +}; + +export function deriveChannelSets(input: { + roster: Roster; + listedIds: ReadonlySet<string>; + onDiskIds: ReadonlySet<string>; +}): ChannelSets { + const { roster, listedIds, onDiskIds } = input; + const rosterIds = Object.keys(roster.entries); + + const listed: string[] = []; + const undownloaded: string[] = []; + for (const id of listedIds) { + listed.push(id); + if (!onDiskIds.has(id)) undownloaded.push(id); + } + + // missingDownloaded is derived from the DISK set, not from `roster ∩ disk`, + // and that is deliberate: it is what maybe-missing.json has always meant + // ("we have this and the listing no longer does"), and buildIndex, + // channelSnapshot and verifyBeforeClean all read that file. Deriving it from + // the roster instead would silently drop any on-disk video the roster hadn't + // caught up with yet — a stale roster must never be able to un-flag a + // deletion. With seeding + additive writers the two sets coincide anyway. + const missingDownloaded: string[] = []; + const orphaned: string[] = []; + for (const id of onDiskIds) { + if (!listedIds.has(id)) missingDownloaded.push(id); + if (!roster.entries[id]) orphaned.push(id); + } + + const missingNeverFetched: string[] = []; + for (const id of rosterIds) { + if (!listedIds.has(id) && !onDiskIds.has(id)) missingNeverFetched.push(id); + } + + const missing = [...missingDownloaded, ...missingNeverFetched]; + + return { + listed: listed.sort(), + undownloaded: undownloaded.sort(), + missing: missing.sort(), + missingDownloaded: missingDownloaded.sort(), + missingNeverFetched: missingNeverFetched.sort(), + orphaned: orphaned.sort(), + }; +} diff --git a/common/controller/channelSnapshot.ts b/common/controller/channelSnapshot.ts @@ -40,6 +40,8 @@ import { } from "./channels"; import { computeKeptVideoIds } from "./keptVideos"; import { loadMaybeMissing } from "./quickAvailabilityCheck"; +import { loadRoster } from "./rosterStore"; +import { deriveChannelSets } from "./channelSets"; import { readTranscriptCoverage } from "./normalizeTranscript"; import { resolveVttProvenance } from "../lib/subtitleProvenance"; import { isIncompleteTranscript } from "../lib/transcriptCoverage"; @@ -195,6 +197,17 @@ export type ChannelSnapshot = { // intersected with on-disk dirs at generation time. Optional: older snapshots // lack it; readers must default via normalizeMaybeMissing. maybeMissing?: { ids: string[]; checkedAt: string }; + // Videos the channel's roster says we were told about, never downloaded, and + // that the current listing no longer contains. The app could not express this + // before: with no data/<id>/ dir there was nothing to notice their absence + // against, and `playlist` — the only place their URL lived — was overwritten + // wholesale by every sweep. + // + // The stored URL is what makes the bucket actionable rather than merely + // defensive: it is enough to attempt a direct-link download even though the + // video has left the channel. Optional: older snapshots lack it; readers must + // default to []. + missingNeverFetched?: MissingNeverFetched[]; // Estimated bytes each cleanup operation would reclaim, mirroring the bucket // counts: transcribedWithAudio sums every audio.* file in those dirs (all are // removed); multipleAudioFormats sums the non-target audio.* files (the target @@ -211,6 +224,14 @@ export type ChannelSnapshot = { }; }; +export type MissingNeverFetched = { + id: string; + // The URL the video was last seen at. "" only when it was never observed in a + // listing (a disk-seeded entry), which cannot happen for this bucket. + url: string; + firstSeenAt: string; +}; + export function emptyExcludedFromDownload(): ExcludedFromDownload { return { membersOnly: [], deleted: [], private: [] }; } @@ -296,15 +317,23 @@ export async function generateChannelSnapshot( /* ignore — the migration CLI can repair stragglers */ } - const [dirEntries, urls, archive, failedListed, config, maybeMissingRecord] = - await Promise.all([ - readdir(dataDir, { withFileTypes: true }).catch(() => [] as Dirent[]), - readPlaylistUrls(playlistPath), - readArchive(archivePath), - loadFailedTranscriptions(paths, slug), - readChannelConfig(paths, slug), - loadMaybeMissing(paths, slug), - ]); + const [ + dirEntries, + urls, + archive, + failedListed, + config, + maybeMissingRecord, + roster, + ] = await Promise.all([ + readdir(dataDir, { withFileTypes: true }).catch(() => [] as Dirent[]), + readPlaylistUrls(playlistPath), + readArchive(archivePath), + loadFailedTranscriptions(paths, slug), + readChannelConfig(paths, slug), + loadMaybeMissing(paths, slug), + loadRoster(paths, slug), + ]); const targetAudioFile = config?.audioFormat ? `audio.${config.audioFormat}` : null; @@ -713,13 +742,29 @@ export async function generateChannelSnapshot( // consumes), so they aren't re-attempted every run — they wait in the // needsCookies bucket for the manual cookie run instead. const cookiePolicy = resolveCookiePolicy(getSettings(), config ?? undefined); + + // Every set the channel is described by comes from one place, so "missing" + // can't mean two things in two files. `listed` is the stored playlist's + // canonical ids; `missingNeverFetched` is the category that used to be + // inexpressible — in the roster, never downloaded, and no longer listed. + const sets = deriveChannelSets({ + roster, + listedIds: new Set( + urls.map((u) => extractVideoId(u)).filter((id): id is string => Boolean(id)), + ), + onDiskIds: new Set(videoDirNames), + }); + const listedIdSet = new Set(sets.listed); + const undownloadedIds: string[] = []; const needsCookies: string[] = []; + // Walked in playlist order, not sorted: this is the auto-download runner's + // work queue, and the listing is newest-first. for (const url of urls) { // Post-reconcile a video's dir is its canonical id, so the URL's canonical // id is the dir name directly. const dirId = extractVideoId(url); - if (!dirId) continue; + if (!dirId || !listedIdSet.has(dirId)) continue; const f = filesById.get(dirId); if (f && videoHasAnyArtifact(f)) continue; const effective = effectiveById.get(dirId); @@ -786,6 +831,15 @@ export async function generateChannelSnapshot( keptCount: keptIds.size, availability, ...(maybeMissing ? { maybeMissing } : {}), + // Newest-first by when we first heard of them: the most recent loss is the + // one still worth chasing. + missingNeverFetched: sets.missingNeverFetched + .map((id) => ({ + id, + url: roster.entries[id]?.url ?? "", + firstSeenAt: roster.entries[id]?.firstSeenAt ?? "", + })) + .sort((a, b) => b.firstSeenAt.localeCompare(a.firstSeenAt)), cleanupBytes: { transcribedWithAudio: transcribedWithAudioBytes, multipleAudioFormats: multipleAudioFormatsBytes, diff --git a/common/controller/quickAvailabilityCheck.ts b/common/controller/quickAvailabilityCheck.ts @@ -3,6 +3,11 @@ import { readdir } from "node:fs/promises"; import { readChannelConfig } from "./channels"; import { extractVideoId, fetchFlatPlaylistUrls } from "../ytdlp/runYtdlp"; import { diffMaybeMissing, writeMaybeMissing } from "./maybeMissingStore"; +import { + mergeRosterFile, + observationsFromUrls, + seedRosterIfAbsent, +} from "./rosterStore"; import type { Paths } from "../lib/paths"; // The record itself lives in the dependency-free ./maybeMissingStore so @@ -49,6 +54,14 @@ export async function runQuickAvailabilityCheck({ ); } + // Seed the roster before the fetch, from the state that exists right now — + // the same protective migration every listing reader performs. + const now = new Date().toISOString(); + await seedRosterIfAbsent(paths, channelSlug, { + now, + onLog: (m) => log(m.trimEnd()), + }); + const freshUrls = await fetchFlatPlaylistUrls({ channelConfig: config, paths, @@ -56,6 +69,18 @@ export async function runQuickAvailabilityCheck({ onLog: log, signal, }); + + // ADD-ONLY. This check computes a missing set, but only the on-disk one — + // maybe-missing, whose meaning is unchanged below. It is not a sweep and may + // not decide that a roster entry has gone; a full, accepted enumeration is + // the only thing allowed to do that (see controller/channelSets.ts). + await mergeRosterFile( + paths, + channelSlug, + observationsFromUrls(freshUrls), + now, + "listing", + ); // Canonical, URL-derived ids — these match data-dir names like-for-like on // every platform. Drop nulls (a null would never match and would spuriously // flag every known video). diff --git a/common/controller/rosterStore.test.ts b/common/controller/rosterStore.test.ts @@ -0,0 +1,298 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { mkdir, mkdtemp, readFile, rm, writeFile } from "node:fs/promises"; +import { tmpdir } from "node:os"; +import path from "node:path"; +import type { Paths } from "../lib/paths"; +import { + emptyRoster, + loadRoster, + mergeRoster, + observationsFromUrls, + recordSweep, + rosterPath, + seedRosterIfAbsent, + writeRoster, +} from "./rosterStore"; + +// Run with: +// pnpm --filter yt-dlp-transcript-common exec tsx --test common/controller/rosterStore.test.ts + +const T1 = "2026-08-01T00:00:00.000Z"; +const T2 = "2026-08-02T00:00:00.000Z"; + +function yt(id: string): string { + return `https://www.youtube.com/watch?v=${id}`; +} + +async function withPaths(fn: (paths: Paths) => Promise<void>): Promise<void> { + const dir = await mkdtemp(path.join(tmpdir(), "ttb-roster-")); + const paths = { channelsDir: path.join(dir, "channels") } as Paths; + try { + await fn(paths); + } finally { + await rm(dir, { recursive: true, force: true }); + } +} + +// --- mergeRoster: additive, and only additive ------------------------------- + +test("mergeRoster never drops an entry, however the listing shrinks", () => { + const full = mergeRoster( + emptyRoster(), + ["a", "b", "c"].map((id) => ({ id, url: yt(id) })), + T1, + "listing", + ); + // The failure this file exists to prevent: an enumeration that returns one of + // three entries. + const after = mergeRoster(full, [{ id: "a", url: yt("a") }], T2, "listing"); + assert.deepEqual(Object.keys(after.entries).sort(), ["a", "b", "c"]); + assert.equal(after.entries.b.url, yt("b")); +}); + +test("firstSeenAt and source are stable across merges while lastListedAt advances", () => { + const first = mergeRoster( + emptyRoster(), + [{ id: "a", url: yt("a") }], + T1, + "listing", + ); + assert.equal(first.entries.a.firstSeenAt, T1); + assert.equal(first.entries.a.lastListedAt, T1); + + const second = mergeRoster(first, [{ id: "a", url: yt("a") }], T2, "listing"); + assert.equal(second.entries.a.firstSeenAt, T1, "firstSeenAt must not move"); + assert.equal(second.entries.a.lastListedAt, T2); + assert.equal(second.entries.a.source, "listing"); +}); + +test("a non-listing observation does not advance lastListedAt", () => { + const first = mergeRoster( + emptyRoster(), + [{ id: "a", url: yt("a") }], + T1, + "listing", + ); + const imported = mergeRoster(first, [{ id: "a", url: yt("a") }], T2, "import"); + assert.equal( + imported.entries.a.lastListedAt, + T1, + "an import is not an enumeration", + ); +}); + +test("a listing URL overwrites a stored one; an import only fills a gap", () => { + const seeded = mergeRoster(emptyRoster(), [{ id: "a", url: "" }], T1, "disk"); + assert.equal(seeded.entries.a.url, ""); + assert.equal(seeded.entries.a.source, "disk"); + + const filled = mergeRoster(seeded, [{ id: "a", url: yt("a") }], T1, "import"); + assert.equal(filled.entries.a.url, yt("a"), "an empty URL is a gap to fill"); + assert.equal(filled.entries.a.source, "disk", "source records first sighting"); + + const stale = mergeRoster( + filled, + [{ id: "a", url: "https://example.com/other" }], + T2, + "import", + ); + assert.equal(stale.entries.a.url, yt("a"), "an import must not overwrite"); + + const relisted = mergeRoster( + filled, + [{ id: "a", url: "https://www.youtube.com/watch?v=a&t=1" }], + T2, + "listing", + ); + assert.equal(relisted.entries.a.url, "https://www.youtube.com/watch?v=a&t=1"); +}); + +test("mergeRoster is pure and returns the same object when nothing changed", () => { + const before = mergeRoster( + emptyRoster(), + [{ id: "a", url: yt("a") }], + T1, + "listing", + ); + const snapshot = JSON.stringify(before); + const same = mergeRoster(before, [{ id: "a", url: yt("a") }], T1, "listing"); + assert.equal(same, before, "no change => same reference, so callers can skip"); + assert.equal(JSON.stringify(before), snapshot, "input must not be mutated"); +}); + +test("observationsFromUrls canonicalizes and drops what it cannot", () => { + const observed = observationsFromUrls([ + yt("abc"), + "https://rumble.com/v123abc-some-title.html", + "not a url", + ]); + assert.deepEqual( + observed.map((o) => o.id), + ["abc", "v123abc"], + ); +}); + +// --- the file --------------------------------------------------------------- + +test("a corrupt roster file reads as empty rather than throwing", async () => { + await withPaths(async (paths) => { + await mkdir(path.join(paths.channelsDir, "ch"), { recursive: true }); + await writeFile(rosterPath(paths, "ch"), "{ not json"); + assert.deepEqual(await loadRoster(paths, "ch"), emptyRoster()); + }); +}); + +test("a missing roster file reads as empty", async () => { + await withPaths(async (paths) => { + assert.deepEqual(await loadRoster(paths, "nope"), emptyRoster()); + }); +}); + +test("ill-typed entries are dropped and a half-written sweep is discarded", async () => { + await withPaths(async (paths) => { + await mkdir(path.join(paths.channelsDir, "ch"), { recursive: true }); + await writeFile( + rosterPath(paths, "ch"), + JSON.stringify({ + version: 1, + entries: { + good: { url: yt("good"), firstSeenAt: T1, lastListedAt: T1, source: "listing" }, + bad: 42, + partial: { firstSeenAt: T1 }, + }, + lastSweep: { at: T1, listedCount: "lots", verdict: "ok" }, + }), + ); + const roster = await loadRoster(paths, "ch"); + assert.deepEqual(Object.keys(roster.entries).sort(), ["good", "partial"]); + assert.equal(roster.entries.partial.url, ""); + assert.equal(roster.entries.partial.lastListedAt, T1); + assert.equal(roster.lastSweep, null); + }); +}); + +test("a written roster round-trips, sweep verdict and all", async () => { + await withPaths(async (paths) => { + const roster = recordSweep( + mergeRoster(emptyRoster(), [{ id: "a", url: yt("a") }], T1, "listing"), + { at: T1, listedCount: 1, verdict: "shrink-suspect" }, + ); + await writeRoster(paths, "ch", roster); + assert.deepEqual(await loadRoster(paths, "ch"), roster); + // No tmp file left behind by the atomic write. + const raw = await readFile(rosterPath(paths, "ch"), "utf8"); + assert.match(raw, /\n$/); + }); +}); + +// --- seeding ---------------------------------------------------------------- + +async function seedFixture( + paths: Paths, + slug: string, + opts: { + playlist?: ReadonlyArray<string>; + dirs?: ReadonlyArray<{ id: string; availabilityUrl?: string; metadataUrl?: string }>; + }, +): Promise<void> { + const channelDir = path.join(paths.channelsDir, slug); + await mkdir(channelDir, { recursive: true }); + if (opts.playlist) { + await writeFile( + path.join(channelDir, "playlist"), + opts.playlist.join("\n") + "\n", + ); + } + for (const d of opts.dirs ?? []) { + const dir = path.join(channelDir, "data", d.id); + await mkdir(dir, { recursive: true }); + if (d.availabilityUrl) { + await writeFile( + path.join(dir, "availability.json"), + JSON.stringify({ + checkedAt: T1, + availability: "public", + webpageUrl: d.availabilityUrl, + }), + ); + } + if (d.metadataUrl) { + await writeFile( + path.join(dir, "metadata.info.json"), + JSON.stringify({ id: d.id, formats: [], webpage_url: d.metadataUrl }), + ); + } + } +} + +test("seeding a channel with a playlist but no data/ preserves every URL", async () => { + await withPaths(async (paths) => { + // The exposure the roster closes: listed, never downloaded, so no dir and + // no metadata.info.json — `playlist` is the ONLY place these URLs live. + await seedFixture(paths, "ch", { playlist: [yt("a"), yt("b")] }); + const roster = await seedRosterIfAbsent(paths, "ch", { now: T1 }); + assert.deepEqual(Object.keys(roster.entries).sort(), ["a", "b"]); + assert.equal(roster.entries.a.url, yt("a")); + assert.equal(roster.entries.b.source, "listing"); + // And it is on disk, so the next sweep cannot erase it. + assert.deepEqual(await loadRoster(paths, "ch"), roster); + }); +}); + +test("seeding unions data/ dirs the playlist does not mention", async () => { + await withPaths(async (paths) => { + // The live corpus has a channel with 617 videos and no playlist file at + // all; this is that case. + await seedFixture(paths, "ch", { + dirs: [ + { id: "a", availabilityUrl: yt("a") }, + { id: "b", metadataUrl: yt("b") }, + { id: "c" }, + ], + }); + const roster = await seedRosterIfAbsent(paths, "ch", { now: T1 }); + assert.deepEqual(Object.keys(roster.entries).sort(), ["a", "b", "c"]); + assert.equal(roster.entries.a.url, yt("a"), "from availability.json"); + assert.equal(roster.entries.b.url, yt("b"), "from metadata.info.json"); + assert.equal(roster.entries.c.url, "", "no URL recoverable, but recorded"); + for (const id of ["a", "b", "c"]) { + assert.equal(roster.entries[id].source, "disk"); + } + }); +}); + +test("a listed video keeps its listing URL even when it also has a dir", async () => { + await withPaths(async (paths) => { + await seedFixture(paths, "ch", { + playlist: [yt("a")], + dirs: [{ id: "a", availabilityUrl: "https://example.com/stale" }], + }); + const roster = await seedRosterIfAbsent(paths, "ch", { now: T1 }); + assert.equal(roster.entries.a.url, yt("a")); + assert.equal(roster.entries.a.source, "listing"); + }); +}); + +test("seeding is a one-off: an existing roster is returned untouched", async () => { + await withPaths(async (paths) => { + await seedFixture(paths, "ch", { playlist: [yt("a")] }); + const first = await seedRosterIfAbsent(paths, "ch", { now: T1 }); + // Something left the listing after the seed. Re-seeding must not drop it. + await writeFile( + path.join(paths.channelsDir, "ch", "playlist"), + yt("z") + "\n", + ); + const second = await seedRosterIfAbsent(paths, "ch", { now: T2 }); + assert.deepEqual(second, first); + }); +}); + +test("seeding a channel with nothing at all writes no file", async () => { + await withPaths(async (paths) => { + await mkdir(path.join(paths.channelsDir, "ch"), { recursive: true }); + const roster = await seedRosterIfAbsent(paths, "ch", { now: T1 }); + assert.deepEqual(roster, emptyRoster()); + await assert.rejects(() => readFile(rosterPath(paths, "ch"), "utf8")); + }); +}); diff --git a/common/controller/rosterStore.ts b/common/controller/rosterStore.ts @@ -0,0 +1,379 @@ +import path from "node:path"; +import { mkdir, open, readdir, readFile, rename, writeFile } from "node:fs/promises"; +import pLimit from "p-limit"; +import type { Paths } from "../lib/paths"; +import { extractVideoId } from "../lib/videoId"; + +// The ROSTER: every video id this channel has ever been seen to contain, with +// the URL it was seen at. `channels/<slug>/roster.json`. +// +// Why it exists. Until now the only record that a video belonged to a channel +// was the `playlist` file, which a full sweep overwrites wholesale. A video +// that appeared in a listing but was never downloaded has no data/<id>/ dir and +// no metadata.info.json, so `playlist` was the ONLY place its URL lived — when +// it dropped out of the listing it was erased with no record it had ever +// existed, and no way to attempt a direct-link recovery. Downstream, +// channelSnapshot's undownloadedIds walks that same file, so the auto-download +// runner's whole work-list went with it. +// +// The roster is the durable answer: append-only, never pruned by a sweep. The +// invariant that makes every writer safe is that only a FULL, ACCEPTED +// enumeration may compute the "missing" sets (see ./channelSets). Every other +// path — the paged sync, the quick availability check, a one-off import — may +// only ADD. mergeRoster below has no removal branch at all; that is deliberate. +// +// Deliberately a leaf apart from Paths and extractVideoId, the same way +// ./maybeMissingStore is, so ytdlp/runYtdlp.ts can write it without a cycle. + +export const ROSTER_FILENAME = "roster.json"; +export const ROSTER_VERSION = 1; + +// Where an id was FIRST learned. Never revised: paired with firstSeenAt, it +// records the original sighting, so `source: "disk"` keeps meaning "this only +// ever showed up as a directory" even after a later listing confirms it. +export type RosterSource = "listing" | "disk" | "import"; + +// What an enumeration was judged to be. Only "ok" and "shrink-confirmed" are +// accepted listings — see ./acceptListing, which reads lastSweep to implement +// the two-observation confirmation. +export type SweepVerdict = + | "ok" + | "empty" + | "shrink-suspect" + | "shrink-confirmed"; + +export type RosterSweep = { + at: string; + listedCount: number; + verdict: SweepVerdict; +}; + +export type RosterEntry = { + // The URL the video was seen at. The whole point of this file: it survives + // the video leaving the listing, which is what makes recovery possible. + // May be "" for a disk-seeded entry whose URL wasn't cheaply recoverable; + // the next accepted enumeration that lists it fills it in. + url: string; + firstSeenAt: string; + // The last ACCEPTED enumeration that contained this id. Advanced only by a + // listing observation — an import or a disk scan is not an enumeration. + lastListedAt: string; + source: RosterSource; +}; + +export type Roster = { + version: number; + updatedAt: string; + lastSweep: RosterSweep | null; + entries: Record<string, RosterEntry>; +}; + +export type ObservedVideo = { id: string; url: string }; + +export function emptyRoster(): Roster { + return { + version: ROSTER_VERSION, + updatedAt: "", + lastSweep: null, + entries: {}, + }; +} + +export function rosterPath(paths: Paths, slug: string): string { + return path.join(paths.channelsDir, slug, ROSTER_FILENAME); +} + +// Canonicalize a list of listing URLs into roster observations. Nulls are +// dropped: an id we can't derive would never match a data-dir name anyway. +export function observationsFromUrls( + urls: ReadonlyArray<string>, +): ObservedVideo[] { + const out: ObservedVideo[] = []; + for (const url of urls) { + const id = extractVideoId(url); + if (id) out.push({ id, url }); + } + return out; +} + +// Fold observations into the roster. PURE and strictly ADDITIVE — there is no +// path here that drops an entry, which is what lets every non-sweep writer call +// it without needing to be trusted. +// +// Returns the SAME object when nothing changed, so callers can skip the write +// (a big channel's roster is a couple of MB and most paged syncs see nothing +// new). +export function mergeRoster( + roster: Roster, + observed: ReadonlyArray<ObservedVideo>, + now: string, + source: RosterSource, +): Roster { + const fromListing = source === "listing"; + const entries: Record<string, RosterEntry> = { ...roster.entries }; + let changed = false; + + for (const { id, url } of observed) { + if (!id) continue; + const prev = entries[id]; + if (!prev) { + entries[id] = { + url: url || "", + firstSeenAt: now, + lastListedAt: now, + source, + }; + changed = true; + continue; + } + // A listing URL is the freshest we can get, so it wins; an import or disk + // scan only fills a gap. firstSeenAt and source never move. + const nextUrl = url && (fromListing || !prev.url) ? url : prev.url; + const nextLastListed = fromListing ? now : prev.lastListedAt; + if (nextUrl !== prev.url || nextLastListed !== prev.lastListedAt) { + entries[id] = { ...prev, url: nextUrl, lastListedAt: nextLastListed }; + changed = true; + } + } + + if (!changed) return roster; + return { ...roster, version: ROSTER_VERSION, updatedAt: now, entries }; +} + +// Stamp the outcome of an enumeration. Kept separate from mergeRoster because a +// REJECTED enumeration still merges (additively, losing nothing) while recording +// that its listing was not trusted — that record is what the next enumeration +// compares against to confirm a genuine mass deletion. +export function recordSweep(roster: Roster, sweep: RosterSweep): Roster { + return { + ...roster, + version: ROSTER_VERSION, + updatedAt: sweep.at, + lastSweep: sweep, + }; +} + +function isSweepVerdict(v: unknown): v is SweepVerdict { + return ( + v === "ok" || + v === "empty" || + v === "shrink-suspect" || + v === "shrink-confirmed" + ); +} + +function isRosterSource(v: unknown): v is RosterSource { + return v === "listing" || v === "disk" || v === "import"; +} + +// Coerce a stored value into a clean Roster, dropping anything ill-typed. A +// hand-edited or half-written file must never crash a sync; the worst case is +// that it reads as empty and the next seed rebuilds it from playlist ∪ data/. +export function normalizeRoster(value: unknown): Roster { + const empty = emptyRoster(); + if (!value || typeof value !== "object") return empty; + const r = value as Record<string, unknown>; + + const entries: Record<string, RosterEntry> = {}; + if (r.entries && typeof r.entries === "object") { + for (const [id, raw] of Object.entries(r.entries as Record<string, unknown>)) { + if (!id || !raw || typeof raw !== "object") continue; + const e = raw as Record<string, unknown>; + const firstSeenAt = typeof e.firstSeenAt === "string" ? e.firstSeenAt : ""; + entries[id] = { + url: typeof e.url === "string" ? e.url : "", + firstSeenAt, + lastListedAt: + typeof e.lastListedAt === "string" ? e.lastListedAt : firstSeenAt, + source: isRosterSource(e.source) ? e.source : "listing", + }; + } + } + + const rawSweep = r.lastSweep as Record<string, unknown> | undefined | null; + const lastSweep: RosterSweep | null = + rawSweep && + typeof rawSweep === "object" && + typeof rawSweep.at === "string" && + typeof rawSweep.listedCount === "number" && + Number.isFinite(rawSweep.listedCount) && + isSweepVerdict(rawSweep.verdict) + ? { + at: rawSweep.at, + listedCount: Math.max(0, Math.floor(rawSweep.listedCount)), + verdict: rawSweep.verdict, + } + : null; + + return { + version: typeof r.version === "number" ? r.version : ROSTER_VERSION, + updatedAt: typeof r.updatedAt === "string" ? r.updatedAt : "", + lastSweep, + entries, + }; +} + +export async function loadRoster(paths: Paths, slug: string): Promise<Roster> { + try { + const raw = await readFile(rosterPath(paths, slug), "utf8"); + return normalizeRoster(JSON.parse(raw)); + } catch { + return emptyRoster(); + } +} + +// Atomic tmp+rename, like maybe-missing.json: a crashed write must never leave a +// truncated roster behind, because a truncated roster is exactly the data loss +// this file exists to prevent. +export async function writeRoster( + paths: Paths, + slug: string, + roster: Roster, +): Promise<void> { + const file = rosterPath(paths, slug); + await mkdir(path.dirname(file), { recursive: true }); + const tmp = `${file}.tmp-${process.pid}`; + await writeFile(tmp, JSON.stringify(roster, null, 2) + "\n"); + await rename(tmp, file); +} + +// Load, merge, write in one step — the shape every additive writer wants. Skips +// the write when the merge changed nothing. +export async function mergeRosterFile( + paths: Paths, + slug: string, + observed: ReadonlyArray<ObservedVideo>, + now: string, + source: RosterSource, +): Promise<Roster> { + const before = await loadRoster(paths, slug); + const after = mergeRoster(before, observed, now, source); + if (after !== before) await writeRoster(paths, slug, after); + return after; +} + +const DISK_URL_CONCURRENCY = 16; +// How far into a metadata.info.json we're willing to read looking for +// webpage_url. yt-dlp writes it AFTER the formats array, which on a real +// YouTube video puts it around 80 KB into a ~550 KB file. Seeding a channel is +// a one-off, but multiplying a full 550 KB read by 10,795 videos is not +// something to do inside a sync job; a miss just leaves the URL blank, and the +// next accepted enumeration fills it in. +const METADATA_SCAN_LIMIT_BYTES = 256 * 1024; + +// Recover a downloaded video's URL from its own dir, cheapest source first. +// availability.json is 137 bytes and is written by the availability backfill +// that runs on every sync, so it covers nearly everything; the capped scan of +// metadata.info.json is the fallback for dirs that predate it. +export async function readVideoUrlFromDisk(videoDir: string): Promise<string> { + try { + const raw = await readFile(path.join(videoDir, "availability.json"), "utf8"); + const url = (JSON.parse(raw) as { webpageUrl?: unknown }).webpageUrl; + if (typeof url === "string" && url) return url; + } catch { + /* fall through to the metadata scan */ + } + return scanMetadataForUrl(path.join(videoDir, "metadata.info.json")); +} + +async function scanMetadataForUrl(file: string): Promise<string> { + let handle: Awaited<ReturnType<typeof open>> | null = null; + try { + handle = await open(file, "r"); + const chunk = Buffer.alloc(64 * 1024); + // Carry the tail of the previous chunk so a key split across a boundary is + // still found. + let carry = ""; + let read = 0; + while (read < METADATA_SCAN_LIMIT_BYTES) { + const { bytesRead } = await handle.read(chunk, 0, chunk.length, read); + if (bytesRead === 0) break; + read += bytesRead; + const text = carry + chunk.subarray(0, bytesRead).toString("utf8"); + const match = /"webpage_url"\s*:\s*"((?:[^"\\]|\\.)*)"/.exec(text); + if (match) { + try { + return JSON.parse(`"${match[1]}"`) as string; + } catch { + return match[1]; + } + } + carry = text.slice(-64); + } + } catch { + /* no metadata, or unreadable — the URL simply isn't recoverable here */ + } finally { + await handle?.close().catch(() => {}); + } + return ""; +} + +async function readSeedPlaylist(file: string): Promise<ObservedVideo[]> { + try { + const raw = await readFile(file, "utf8"); + return observationsFromUrls( + raw.split("\n").map((s) => s.trim()).filter(Boolean), + ); + } catch { + return []; + } +} + +async function readVideoDirNames(dataDir: string): Promise<string[]> { + try { + const entries = await readdir(dataDir, { withFileTypes: true }); + return entries.filter((e) => e.isDirectory()).map((e) => e.name); + } catch { + return []; + } +} + +// Build the roster for a channel that doesn't have one yet, from the state that +// exists RIGHT NOW: the pre-existing `playlist` (ids and URLs) union the data/ +// dir names. This has to happen before anything overwrites `playlist`, and it +// is the migration's whole value — it captures today's listed-but-never-fetched +// entries before the first sweep can erase them, and gives a channel that has +// never had a listing file (there is one such channel in the live corpus, with +// 617 videos) its first record of what it contains. +// +// A corrupt or empty roster re-seeds, which is the healing behaviour we want. +export async function seedRosterIfAbsent( + paths: Paths, + slug: string, + opts: { now: string; onLog?: (msg: string) => void }, +): Promise<Roster> { + const existing = await loadRoster(paths, slug); + if (Object.keys(existing.entries).length > 0) return existing; + + const channelDir = path.join(paths.channelsDir, slug); + const [listed, dirNames] = await Promise.all([ + readSeedPlaylist(path.join(channelDir, "playlist")), + readVideoDirNames(path.join(channelDir, "data")), + ]); + + let roster = mergeRoster(existing, listed, opts.now, "listing"); + + // Only dirs the listing didn't already account for need a disk read; on a + // channel in a downloaded steady state that is a handful, not thousands. + const known = new Set(Object.keys(roster.entries)); + const diskOnly = dirNames.filter((id) => !known.has(id)); + const limit = pLimit(DISK_URL_CONCURRENCY); + const diskObserved = await Promise.all( + diskOnly.map((id) => + limit(async () => ({ + id, + url: await readVideoUrlFromDisk(path.join(channelDir, "data", id)), + })), + ), + ); + roster = mergeRoster(roster, diskObserved, opts.now, "disk"); + + const total = Object.keys(roster.entries).length; + if (total > 0) { + await writeRoster(paths, slug, roster); + opts.onLog?.( + `Roster: seeded ${total} video(s) — ${listed.length} from the stored playlist, ${diskOnly.length} from data/ dirs not in it.\n`, + ); + } + return roster; +} diff --git a/common/lib/settings.ts b/common/lib/settings.ts @@ -361,6 +361,15 @@ export type SyncSchedulerSettings = { // suspects are flagged and left for a manual check rather than firing hundreds // of probes inside a sync. 0 = never auto-confirm. fullSweepConfirmMaxSuspects: number; + // Shrink guard: how far a fresh listing may fall below the stored one before + // it is treated as suspect rather than acted on. Expressed as a percentage of + // the previous count, floored at SHRINK_ABS_FLOOR entries so ordinary churn on + // a small channel doesn't trip it. A suspect listing does not rewrite + // `playlist` or maybe-missing.json and does not count as a sweep — but a + // SECOND enumeration reporting a similar count confirms it and is accepted, so + // a genuine mass deletion costs at most one cadence period. 0 = off (the + // empty-listing rejection still applies). See controller/acceptListing.ts. + fullSweepShrinkGuardPercent: number; }; export type SocialLink = { @@ -427,6 +436,10 @@ export const KEEP_LATEST_CHECK_DEFAULT_INTERVAL_MINUTES = 1440; export const FULL_SWEEP_DEFAULT_INTERVAL_MINUTES = 1440; export const FULL_SWEEP_CONFIRM_MAX_SUSPECTS_DEFAULT = 25; export const FULL_SWEEP_CONFIRM_MAX_SUSPECTS_MAX = 10000; +// Shrink-guard default: a listing that has lost more than a tenth of its +// entries (and more than SHRINK_ABS_FLOOR of them) needs a second opinion. +export const FULL_SWEEP_SHRINK_GUARD_PERCENT_DEFAULT = 10; +export const FULL_SWEEP_SHRINK_GUARD_PERCENT_MAX = 100; export const SAVED_VIDEO_BACKUP_DEFAULT_INTERVAL_MINUTES = 1440; // Internal-heartbeat cadence bounds. 0 means "off" (use an external cron @@ -449,6 +462,7 @@ export function defaultSyncScheduler(): SyncSchedulerSettings { keepLatestCheckIntervalMinutes: KEEP_LATEST_CHECK_DEFAULT_INTERVAL_MINUTES, fullSweepIntervalMinutes: FULL_SWEEP_DEFAULT_INTERVAL_MINUTES, fullSweepConfirmMaxSuspects: FULL_SWEEP_CONFIRM_MAX_SUSPECTS_DEFAULT, + fullSweepShrinkGuardPercent: FULL_SWEEP_SHRINK_GUARD_PERCENT_DEFAULT, }; } @@ -549,6 +563,11 @@ export function sanitizeSyncScheduler(value: unknown): SyncSchedulerSettings { d.fullSweepConfirmMaxSuspects, FULL_SWEEP_CONFIRM_MAX_SUSPECTS_MAX, ), + fullSweepShrinkGuardPercent: clampIntAllowZero( + r.fullSweepShrinkGuardPercent, + d.fullSweepShrinkGuardPercent, + FULL_SWEEP_SHRINK_GUARD_PERCENT_MAX, + ), }; } diff --git a/common/lib/videoId.ts b/common/lib/videoId.ts @@ -0,0 +1,64 @@ +// The canonical video id derived from a video URL — the name every video's +// data/<id>/ dir carries, on every platform. Deliberately a leaf module with no +// imports at all: it lives here rather than in ytdlp/runYtdlp.ts (its original +// home, which still re-exports it) so that low-level stores like +// controller/rosterStore.ts can canonicalize a URL without pulling in execa, +// the settings loader and the whole download pipeline. +// +// Canonical is NOT the same as yt-dlp's native extractor id: the two coincide +// on YouTube and diverge everywhere else. See archiveIdForUrl in runYtdlp.ts +// for the native-id resolution that reads metadata.info.json. +export function extractVideoId(url: string): string | null { + try { + const u = new URL(url); + const host = u.hostname.toLowerCase(); + if (host.endsWith("youtube.com") || host === "youtu.be") { + const v = u.searchParams.get("v"); + if (v) return v; + const seg = u.pathname.split("/").filter(Boolean).pop(); + return seg ?? null; + } + if (host.endsWith("rumble.com")) { + const seg = u.pathname.split("/").filter(Boolean).pop(); + // Rumble paths often look like /v123abc-some-title.html + if (seg) { + const trimmed = seg.replace(/\.html?$/i, ""); + const dashIdx = trimmed.indexOf("-"); + return dashIdx > 0 ? trimmed.slice(0, dashIdx) : trimmed; + } + return null; + } + if (host.endsWith("odysee.com")) { + const decoded = decodeURIComponent(u.pathname); + const lastColon = decoded.lastIndexOf(":"); + if (lastColon > 0) return decoded.slice(lastColon + 1); + return null; + } + if (host.endsWith("twitch.tv")) { + // VODs: /videos/<id> or legacy /<channel>/v/<id>; clips: + // clips.twitch.tv/<slug> or /<channel>/clip/<slug>. Grab the segment + // after the marker; otherwise fall back to the last path segment. + const segs = u.pathname.split("/").filter(Boolean); + for (const marker of ["videos", "v", "clip", "clips"]) { + const idx = segs.indexOf(marker); + if (idx >= 0 && segs[idx + 1]) return segs[idx + 1]; + } + return segs.pop() ?? null; + } + if (host.endsWith("kick.com")) { + // VODs: /<channel>/videos/<uuid> or /video/<uuid>; clips: + // /<channel>/clips/<slug> or /clip/<slug>. The UUID after the marker is + // yt-dlp's native Kick id; fall back to the last segment otherwise. + const segs = u.pathname.split("/").filter(Boolean); + for (const marker of ["videos", "video", "clips", "clip"]) { + const idx = segs.indexOf(marker); + if (idx >= 0 && segs[idx + 1]) return segs[idx + 1]; + } + return segs.pop() ?? null; + } + const seg = u.pathname.split("/").filter(Boolean).pop(); + return seg ?? null; + } catch { + return null; + } +} diff --git a/common/ytdlp/runYtdlp.ts b/common/ytdlp/runYtdlp.ts @@ -15,6 +15,7 @@ import { formatBytes } from "../lib/format"; import { detectPlatform } from "../lib/platform"; import { isRealAudioFile } from "../lib/videoStatus"; import { readVttProvenance } from "../lib/subtitleProvenance"; +import { extractVideoId } from "../lib/videoId"; import type { Paths } from "../lib/paths"; import { EXCLUDED_FROM_DOWNLOAD, @@ -30,10 +31,18 @@ import { import { resolveEffectiveAvailability } from "../lib/availability-server"; import { backfillAvailabilityFromMetadata } from "../controller/backfillAvailability"; import { runAvailabilityCheck } from "../controller/checkAvailability"; +import { writeMaybeMissing } from "../controller/maybeMissingStore"; import { - diffMaybeMissing, - writeMaybeMissing, -} from "../controller/maybeMissingStore"; + mergeRoster, + mergeRosterFile, + observationsFromUrls, + recordSweep, + seedRosterIfAbsent, + writeRoster, + type Roster, +} from "../controller/rosterStore"; +import { acceptListing, type ListingDecision } from "../controller/acceptListing"; +import { deriveChannelSets } from "../controller/channelSets"; import { resolveShardItems } from "../controller/shard"; import { isFullSweepDue, @@ -450,10 +459,78 @@ async function writePlaylistFile( onLog(`Wrote ${urls.length} URLs to ${playlistPath}\n`); } +// Distinct canonical ids in the channel's currently-stored `playlist` — the +// size of the last listing we accepted, and the reference the shrink guard +// measures a fresh enumeration against. Missing file = no reference. +async function storedListingCount(root: string): Promise<number> { + try { + const raw = await readFile(path.join(root, "playlist"), "utf8"); + const urls = raw.split("\n").map((s) => s.trim()).filter(Boolean); + return new Set(observationsFromUrls(urls).map((o) => o.id)).size; + } catch { + return 0; + } +} + +// Put a full enumeration through the roster and the shrink guard. BOTH +// full-listing writers go through here — store-playlist and the sync full +// sweep — so neither can rewrite `playlist` from a listing the other would have +// refused. store-playlist overwriting unconditionally was the most direct +// instance of the reported bug. +// +// Note the order: the roster is seeded from the pre-existing state BEFORE +// anything is overwritten, and the fresh observations are merged in +// UNCONDITIONALLY — even for a rejected listing. Merging is additive, so a bad +// fetch can only ever add; an entry we decline to record is an entry we can +// lose, which is the failure this whole part exists to prevent. +async function acceptEnumeration( + opts: RunYtdlpOpts, + root: string, + urls: ReadonlyArray<string>, +): Promise<{ + decision: ListingDecision; + roster: Roster; + listedIds: Set<string>; + now: string; +}> { + const now = new Date().toISOString(); + const observed = observationsFromUrls(urls); + const listedIds = new Set(observed.map((o) => o.id)); + + const [previousListedCount, seeded] = await Promise.all([ + storedListingCount(root), + seedRosterIfAbsent(opts.paths, opts.channelSlug, { + now, + onLog: opts.onLog, + }), + ]); + + const decision = acceptListing({ + listedCount: listedIds.size, + previousListedCount, + lastSweep: seeded.lastSweep, + shrinkGuardPercent: getSettings().syncScheduler.fullSweepShrinkGuardPercent, + }); + + const roster = recordSweep(mergeRoster(seeded, observed, now, "listing"), { + at: now, + listedCount: listedIds.size, + verdict: decision.verdict, + }); + await writeRoster(opts.paths, opts.channelSlug, roster); + + return { decision, roster, listedIds, now }; +} + async function storePlaylist(opts: RunYtdlpOpts): Promise<void> { const root = channelRoot(opts); await mkdir(root, { recursive: true }); const urls = await enumeratePlaylistUrls(opts, root); + const { decision } = await acceptEnumeration(opts, root, urls); + if (!decision.accept) { + opts.onLog(`Store playlist: ${decision.reason}.\n`); + return; + } await writePlaylistFile(root, urls, opts.onLog); } @@ -1260,6 +1337,15 @@ async function syncPaged(opts: RunYtdlpOpts): Promise<void> { // before any download leaves no data/ behind. await mkdir(root, { recursive: true }); + // Seed the roster from the state that exists right now. A channel may go a + // long time between sweeps, and the seeding is itself the protection: it + // captures today's listed-but-never-fetched entries, with their URLs, before + // any later enumeration can overwrite the stored playlist. + await seedRosterIfAbsent(opts.paths, opts.channelSlug, { + now: new Date().toISOString(), + onLog: opts.onLog, + }); + // Walk the channel newest-first a page at a time. On each page, download the // entries not yet in the archive, then stop once we reach a page that // contains an already-archived entry (we've caught up to a prior sync) or a @@ -1286,6 +1372,18 @@ async function syncPaged(opts: RunYtdlpOpts): Promise<void> { const pageUrls = await enumeratePlaylistUrls(opts, root, { start, end }); if (pageUrls.length === 0) break; + // ADD-ONLY: a page is a real (if partial) sighting of the listing, so the + // ids on it belong in the roster — but a paged walk sees only the newest + // window, so it may never compute a "missing" set. That rule is what makes + // every non-sweep writer safe to call. + await mergeRosterFile( + opts.paths, + opts.channelSlug, + observationsFromUrls(pageUrls), + new Date().toISOString(), + "listing", + ); + const { newUrls, archivedHits, deferredAuthCount } = await selectDownloadableUrls(pageUrls, archive, dataDir, runCookiePolicy); opts.onLog( @@ -1354,50 +1452,64 @@ async function syncFullSweep(opts: RunYtdlpOpts): Promise<void> { `Full sweep: re-reading the whole channel listing (refreshes the video list and flags videos that have gone missing).\n`, ); - // 1. One enumeration, no range. + // 1. One enumeration, no range, then the gate: record everything it saw in + // the roster (additive, so this is safe unconditionally) and decide + // whether the listing itself may be acted on. const urls = await enumeratePlaylistUrls(opts, root); if (opts.signal.aborted) return; - // An empty listing is never trustworthy enough to act on: it is what a - // transient upstream failure, a cookie expiry, or a clean 101 exit all look - // like from here. Acting on it would truncate the stored playlist AND flag - // every video we own as missing. Leave both alone and don't stamp the sweep, - // so the next sync tries again rather than waiting out the cadence. - if (urls.length === 0) { + const { decision, roster, listedIds, now } = await acceptEnumeration( + opts, + root, + urls, + ); + opts.onLog(`Full sweep: ${decision.reason}.\n`); + + // A rejected listing must not truncate the stored playlist, must not flag + // videos missing, and must not stamp lastFullSweepAt — so the next sync + // retries instead of waiting out the cadence. It DOES still run the download + // walk below: downloading entries the listing does contain is purely + // additive, and skipping it would stall a channel's downloads for as long as + // the listing stays suspect. + let maybeMissing: string[] = []; + if (decision.accept) { + // 2. Refresh the stored playlist. Everything downstream — "download + // missing", the snapshot's undownloaded work-list — reads this file, and + // before the sweep only an explicit "store playlist" ever rewrote it. + await writePlaylistFile(root, urls, opts.onLog); + + // 3. Derive every set from one place (controller/channelSets). The + // maybe-missing record keeps its exact shape and meaning — videos we + // have on disk that the listing no longer carries — so buildIndex, + // channelSnapshot and verifyBeforeClean need no changes. What's new is + // missingNeverFetched: entries the roster knows we were told about, + // never fetched, and that are now gone. Without the roster those had no + // dir to be noticed by and no record to be noticed in. + const entries = await readdir(dataDir, { withFileTypes: true }).catch( + () => [] as Awaited<ReturnType<typeof readdir>> & { length: 0 }, + ); + const onDiskIds = new Set( + (entries as Array<{ isDirectory(): boolean; name: string }>) + .filter((e) => e.isDirectory()) + .map((e) => e.name), + ); + const sets = deriveChannelSets({ roster, listedIds, onDiskIds }); + maybeMissing = sets.missingDownloaded; + await writeMaybeMissing(opts.paths, opts.channelSlug, { + checkedAt: now, + freshPlaylistCount: listedIds.size, + ids: maybeMissing, + }); opts.onLog( - `Full sweep: the channel listing came back empty — leaving the stored playlist and missing-video flags untouched, and not counting this as a sweep.\n`, + `Full sweep: ${onDiskIds.size} known, ${listedIds.size} in fresh listing, ${maybeMissing.length} maybe-missing.\n`, ); - await touchLastSync(opts); - return; + if (sets.missingNeverFetched.length > 0) { + opts.onLog( + `Full sweep: ${sets.missingNeverFetched.length} video(s) were listed but never downloaded and have now left the listing — see "Never fetched, now gone" on the channel page to attempt a direct-link recovery.\n`, + ); + } } - // 2. Refresh the stored playlist. Everything downstream — "download missing", - // the snapshot's undownloadedIds — reads this file, and before the sweep - // only an explicit "store playlist" ever rewrote it. - await writePlaylistFile(root, urls, opts.onLog); - - // 3. Diff the fresh listing against what's on disk. Canonical, URL-derived - // ids match data-dir names on every platform; nulls are dropped because a - // null would never match and would spuriously flag every known video. - const freshIds = new Set( - urls.map((u) => extractVideoId(u)).filter((id): id is string => Boolean(id)), - ); - const entries = await readdir(dataDir, { withFileTypes: true }).catch( - () => [] as Awaited<ReturnType<typeof readdir>> & { length: 0 }, - ); - const knownIds = (entries as Array<{ isDirectory(): boolean; name: string }>) - .filter((e) => e.isDirectory()) - .map((e) => e.name); - const maybeMissing = diffMaybeMissing(knownIds, freshIds); - await writeMaybeMissing(opts.paths, opts.channelSlug, { - checkedAt: new Date().toISOString(), - freshPlaylistCount: freshIds.size, - ids: maybeMissing, - }); - opts.onLog( - `Full sweep: ${knownIds.length} known, ${freshIds.size} in fresh listing, ${maybeMissing.length} maybe-missing.\n`, - ); - // 4. Walk the same listing in newest-first windows, applying the paged walk's // filter and stopping rule to slices instead of spawns. let archive = await readArchive(archivePath); @@ -1467,7 +1579,11 @@ async function syncFullSweep(opts: RunYtdlpOpts): Promise<void> { await confirmMaybeMissing(opts, maybeMissing); await touchLastSync(opts); - await touchLastFullSweep(opts); + // Only an accepted listing counts as a sweep. Leaving the stamp alone is what + // makes the next sync retry the enumeration immediately instead of waiting + // out the cadence — and, for a genuine mass deletion, what turns the retry + // into the confirming second observation. + if (decision.accept) await touchLastFullSweep(opts); await safeBackfillAvailability(opts); } @@ -1620,57 +1736,7 @@ export async function archiveIdForUrl( return canonical; } -export function extractVideoId(url: string): string | null { - try { - const u = new URL(url); - const host = u.hostname.toLowerCase(); - if (host.endsWith("youtube.com") || host === "youtu.be") { - const v = u.searchParams.get("v"); - if (v) return v; - const seg = u.pathname.split("/").filter(Boolean).pop(); - return seg ?? null; - } - if (host.endsWith("rumble.com")) { - const seg = u.pathname.split("/").filter(Boolean).pop(); - // Rumble paths often look like /v123abc-some-title.html - if (seg) { - const trimmed = seg.replace(/\.html?$/i, ""); - const dashIdx = trimmed.indexOf("-"); - return dashIdx > 0 ? trimmed.slice(0, dashIdx) : trimmed; - } - return null; - } - if (host.endsWith("odysee.com")) { - const decoded = decodeURIComponent(u.pathname); - const lastColon = decoded.lastIndexOf(":"); - if (lastColon > 0) return decoded.slice(lastColon + 1); - return null; - } - if (host.endsWith("twitch.tv")) { - // VODs: /videos/<id> or legacy /<channel>/v/<id>; clips: - // clips.twitch.tv/<slug> or /<channel>/clip/<slug>. Grab the segment - // after the marker; otherwise fall back to the last path segment. - const segs = u.pathname.split("/").filter(Boolean); - for (const marker of ["videos", "v", "clip", "clips"]) { - const idx = segs.indexOf(marker); - if (idx >= 0 && segs[idx + 1]) return segs[idx + 1]; - } - return segs.pop() ?? null; - } - if (host.endsWith("kick.com")) { - // VODs: /<channel>/videos/<uuid> or /video/<uuid>; clips: - // /<channel>/clips/<slug> or /clip/<slug>. The UUID after the marker is - // yt-dlp's native Kick id; fall back to the last segment otherwise. - const segs = u.pathname.split("/").filter(Boolean); - for (const marker of ["videos", "video", "clips", "clip"]) { - const idx = segs.indexOf(marker); - if (idx >= 0 && segs[idx + 1]) return segs[idx + 1]; - } - return segs.pop() ?? null; - } - const seg = u.pathname.split("/").filter(Boolean).pop(); - return seg ?? null; - } catch { - return null; - } -} +// extractVideoId now lives in ../lib/videoId — a true leaf, so the roster store +// can canonicalize URLs without importing this module (which would be a cycle). +// Re-exported here because every existing caller imports it from this path. +export { extractVideoId }; diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,9 @@ # Changelog ## [Unreleased] +- **No video a channel has ever listed can be lost again, and a truncated listing can no longer destroy one.** Until now the only record that a video belonged to a channel was the stored playlist file, which the deep pass overwrites wholesale. A video that appeared in a listing but was never downloaded had no folder on disk and no metadata of its own, so the playlist was the *only* place its URL lived — when it dropped out of the listing it was erased with no trace it had ever existed, and no way to even attempt a direct-link recovery. Each channel now keeps a **roster**: every video id it has ever been seen to contain, with the URL it was seen at, added to and never pruned. It is built the first time a channel syncs, from the playlist and the folders on disk *before* anything is rewritten, so the upgrade itself is the protection. +- **A listing that comes back suspiciously small is no longer believed on the first try.** The previous guard only refused a *completely* empty listing; a fetch that returned 100 of 10,795 entries sailed straight through and destroyed the other 10,695. A fresh listing that has lost more than a tenth of its entries — and more than 25 of them, so ordinary churn on a small channel doesn't trip it — is now treated as suspect: the stored video list and the missing-video flags are left exactly as they were, and the sweep is retried on the very next sync instead of after a full day. If a second enumeration reports a similar count it is accepted and acted on, because a real mass deletion repeats and a transient blip does not. No dialog to answer and nothing to override; at worst a genuine deletion lands one cadence later. The threshold is a new **Full-sweep shrink guard** setting (0 turns it off; an empty listing is always refused). +- **New: "Never fetched, now gone".** A channel's Diagnostics stage, and the Actionable page, now list the videos that were in the listing, were never downloaded, and have since left it — a category the app previously could not express at all, and for an archive the one that matters most. Each row shows when it was first seen and the URL the roster kept, with a **Try downloading anyway** button beside it: a video that merely went unlisted still downloads from a direct link, while a deleted one will fail and tell you so. The existing maybe-missing list is unchanged and still means what it always did — videos you *have* that the listing no longer carries. - **Every schedule is set in minutes, hours, or days now — not a raw minute count.** Seven different cadences in the editor were plain number boxes measured in minutes, which meant knowing that `10080` is a week and `44640` is the documented maximum. Each one is now an amount plus a unit, with a line underneath restating both the stored number and what it works out to — "Every 1,440 minutes · about 1 full listing fetch per channel per day" — so the rate you are about to ask of a video host is visible while you set it, not after. A cadence typed by hand into a config file keeps its exact value: 137 minutes stays 137 minutes rather than being rounded to the nearest tidy unit. The **deep pass** cadence is settable at last — globally in Settings and on the Scheduler page, and per channel on the channel form, where a big archive can be told to re-read its listing weekly while everything else does it daily. **Sync all** gained a **Full sweep all** button (and each channel a **Full sweep** one) for when you don't want to wait out a day of cadence — right after upgrading, say, when nothing has been swept yet. The Scheduler page now shows both cadences for every channel at a glance and can retune many at once: tick the channels, set one or both cadences, apply. Leaving a field alone leaves that cadence alone, so forty channels' deep passes can be re-tuned without touching anyone's ordinary sync. - **Fixed: the keep-latest check interval was reset to its default every time you saved settings.** It was the one scheduler cadence with no input anywhere in the app, and saving an unrelated setting quietly overwrote whatever you had put in the config file by hand. It now has an input, and the save path preserves any field the form doesn't render. - **Sync now notices when a video disappears.** A channel's video list was being fetched three separate times for three purposes that never shared their work: "store playlist" refreshed the stored list, Sync walked the newest 50 entries at a time, and "Quick check" re-read the whole listing to find videos that had gone missing. Because Sync only ever saw the newest slice, it could never spot a deletion — and it never refreshed the stored list either, so "Download missing" and the report's not-yet-downloaded count kept working off whatever the last "store playlist" click wrote, possibly months earlier. Sync now periodically pays for **one** full read of the channel and gets all three out of it: the stored list is refreshed, videos that have left the listing are flagged, and new uploads are downloaded as before. That means **Sync all** surfaces upstream deletions across every channel on its own, where it used to take a per-channel "Quick check" click. The deep pass runs at most once a day per channel by default (it is much more expensive than a normal sync on a large channel); every sync in between stays exactly as cheap as it was. When a handful of videos are flagged — 25 or fewer by default — the same job goes on to work out which are deleted, private, or merely unlisted; past that it flags them and leaves the call to you. **What it will not do is download anything a normal sync wouldn't**: on a channel you deliberately keep only the newest few hundred of, a deep pass will not start dragging down the back catalogue. It changes what the editor *knows*, never what it *fetches*. If the channel listing comes back empty — a network blip, an expired cookie — the deep pass leaves the stored list and the missing-video flags untouched rather than concluding your whole archive vanished. diff --git a/editor/app/actionable/lib/loadActionable.ts b/editor/app/actionable/lib/loadActionable.ts @@ -27,6 +27,7 @@ export type ActionableRow = { export type ActionableSummary = { rows: ActionableRow[]; undownloaded: ActionableRow[]; + missingNeverFetched: ActionableRow[]; untranscribed: ActionableRow[]; incompleteTranscripts: ActionableRow[]; shortAudio: ActionableRow[]; @@ -71,6 +72,14 @@ export function actionableUndownloadedCount(row: ActionableRow): number { return countActionable(row.snapshot, row.snapshot?.undownloadedIds); } +// Videos the roster says we were told about, never downloaded, and that have +// since left the listing. Deliberately NOT run through countActionable: the +// availability exclusions are keyed on videos we have on disk, and these have no +// dir at all. Default 0 for snapshots written before the bucket existed. +export function actionableMissingNeverFetchedCount(row: ActionableRow): number { + return row.snapshot?.missingNeverFetched?.length ?? 0; +} + export function actionableUntranscribedCount(row: ActionableRow): number { return countActionable( row.snapshot, @@ -144,6 +153,14 @@ export async function loadActionableSummary( (a, b) => actionableUndownloadedCount(b) - actionableUndownloadedCount(a), ); + const missingNeverFetched = rows + .filter((r) => actionableMissingNeverFetchedCount(r) > 0) + .sort( + (a, b) => + actionableMissingNeverFetchedCount(b) - + actionableMissingNeverFetchedCount(a), + ); + const untranscribed = rows .filter((r) => actionableUntranscribedCount(r) > 0) .sort( @@ -191,6 +208,7 @@ export async function loadActionableSummary( return { rows, undownloaded, + missingNeverFetched, untranscribed, incompleteTranscripts, shortAudio, diff --git a/editor/app/actionable/page.tsx b/editor/app/actionable/page.tsx @@ -9,6 +9,7 @@ import { actionableCleanTranscribedCount, actionableDigestWarningsCount, actionableIncompleteTranscriptCount, + actionableMissingNeverFetchedCount, actionableShortAudioCount, actionableUndownloadedCount, actionableUntranscribedCount, @@ -53,6 +54,7 @@ export default async function ActionablePage() { const summary = await loadActionableSummary(paths); const nothingPending = summary.undownloaded.length === 0 && + summary.missingNeverFetched.length === 0 && summary.untranscribed.length === 0 && summary.incompleteTranscripts.length === 0 && summary.shortAudio.length === 0 && @@ -81,6 +83,27 @@ export default async function ActionablePage() { }, { config: { + id: "missing-never-fetched", + title: "Channels with videos lost before they were ever downloaded", + description: + "Videos that were in the channel listing, were never fetched, and have since left it. Nothing of them exists on disk — only the URL the roster kept, which is enough to attempt a direct-link download. A video that merely went unlisted still downloads; a deleted one will not. Open the channel's Diagnostics stage to try them.", + countLabel: "never fetched", + emptyLabel: "None lost.", + getCount: actionableMissingNeverFetchedCount, + primaryAction: (r) => ( + <Link + href={`/channels/${r.channel.slug}`} + aria-label={`review never-fetched videos for ${r.channel.slug}`} + className="inline-flex items-center px-2.5 py-1 rounded-md border border-border text-xs font-medium hover:bg-muted whitespace-nowrap" + > + Review + </Link> + ), + }, + rows: summary.missingNeverFetched, + }, + { + config: { id: "untranscribed", title: "Channels with downloaded videos awaiting transcription", description: diff --git a/editor/app/channels/[slug]/components/stages/DiagnosticsStage.tsx b/editor/app/channels/[slug]/components/stages/DiagnosticsStage.tsx @@ -6,8 +6,12 @@ import { AVAILABILITY_VALUES, type Availability, } from "yt-dlp-transcript-common/lib/availability"; -import type { AvailabilitySnapshot } from "yt-dlp-transcript-common/controller/channelSnapshot"; +import type { + AvailabilitySnapshot, + MissingNeverFetched, +} from "yt-dlp-transcript-common/controller/channelSnapshot"; import { verifyAction, type VerifyResult } from "../../whisperActions"; +import { importVideoAction } from "../../pipelineActions"; import { checkAvailabilityAction, checkMaybeMissingAction, @@ -52,6 +56,7 @@ type Props = { totals: { videos: number; transcribed: number; downloaded: number }; availability: AvailabilitySnapshot; maybeMissing: { ids: string[]; checkedAt: string }; + missingNeverFetched: MissingNeverFetched[]; availabilityShard: ShardConfigSummary | null; existingQueues: string[]; downloadDefaultQueueKey: string; @@ -68,6 +73,7 @@ export function DiagnosticsStage({ totals, availability, maybeMissing, + missingNeverFetched, availabilityShard, existingQueues, downloadDefaultQueueKey, @@ -168,6 +174,7 @@ export function DiagnosticsStage({ downloadDefaultQueueKey={downloadDefaultQueueKey} /> <QuickAvailabilityCheck slug={slug} maybeMissing={maybeMissing} /> + <NeverFetchedGone slug={slug} rows={missingNeverFetched} /> </div> <div className="flex flex-col gap-2" aria-label="channel health"> <div> @@ -501,6 +508,85 @@ function QuickAvailabilityCheck({ ); } +// The category the app could not express before the roster existed: videos that +// were in the channel listing, were never downloaded, and are no longer listed. +// With no data/<id>/ dir there was nothing for a missing-check to notice, and +// the stored `playlist` — the only place their URL lived — was overwritten by +// every sweep, so they disappeared without a trace. +// +// The kept URL is what makes this actionable rather than merely a record of +// loss, so every row offers a direct-link attempt. +function NeverFetchedGone({ + slug, + rows, +}: { + slug: string; + rows: MissingNeverFetched[]; +}) { + if (rows.length === 0) return null; + return ( + <div + className="mt-2 flex flex-col gap-3 rounded border border-border p-3" + aria-label="never fetched gone" + > + <div> + <h4 className="text-sm font-semibold"> + Never fetched, now gone ({rows.length}) + </h4> + <p className="text-xs text-muted-foreground"> + These videos appeared in the channel listing, were never downloaded, + and have since left it. Nothing of them survives on disk — only the URL + the roster kept, which is enough to try anyway: a video that merely + went unlisted still downloads from a direct link, while a deleted one + will fail. + </p> + </div> + <div className="rounded border border-border max-h-72 overflow-auto"> + <ul className="divide-y divide-border"> + {rows.map((r) => ( + <li + key={r.id} + aria-label={`never fetched ${r.id}`} + className="flex flex-wrap items-center gap-x-3 gap-y-1 px-3 py-2" + > + <span className="font-mono text-xs">{r.id}</span> + {r.firstSeenAt && ( + <span className="text-xs text-muted-foreground"> + first seen {new Date(r.firstSeenAt).toLocaleDateString()} + </span> + )} + {r.url ? ( + <a + href={r.url} + target="_blank" + rel="noreferrer" + aria-label={`never fetched url ${r.id}`} + className="min-w-0 truncate text-xs text-muted-foreground underline hover:text-foreground" + > + {r.url} + </a> + ) : ( + <span className="text-xs text-muted-foreground"> + no URL recorded + </span> + )} + {r.url && ( + <StreamActionLog + trigger={() => importVideoAction(slug, r.url)} + cancelAction={cancelJobAction} + buttonLabel="Try downloading anyway" + runningLabel="Trying…" + label={`try downloading ${r.id}`} + /> + )} + </li> + ))} + </ul> + </div> + </div> + ); +} + const AVAILABILITY_LABELS: Record<Availability, string> = { public: "Public", unlisted: "Unlisted", diff --git a/editor/app/channels/[slug]/page.tsx b/editor/app/channels/[slug]/page.tsx @@ -426,6 +426,7 @@ export default async function ChannelDetailPage({ totals={snapshot.totals} availability={availability} maybeMissing={maybeMissing} + missingNeverFetched={snapshot.missingNeverFetched ?? []} availabilityShard={availabilityShard} existingQueues={existingQueues} downloadDefaultQueueKey={platformDefaultQueueKey} diff --git a/editor/app/channels/[slug]/pipelineActions.ts b/editor/app/channels/[slug]/pipelineActions.ts @@ -22,6 +22,7 @@ import { readChannelStat, } from "yt-dlp-transcript-common/controller/channels"; import { extractVideoId, runYtdlp } from "yt-dlp-transcript-common/ytdlp/runYtdlp"; +import { mergeRosterFile } from "yt-dlp-transcript-common/controller/rosterStore"; import { downloadOneManaged } from "yt-dlp-transcript-common/ytdlp/downloadOneManaged"; import { getSettings } from "yt-dlp-transcript-common/lib/settings"; import { resolveCookiePolicy } from "yt-dlp-transcript-common/lib/cookiePolicy"; @@ -438,7 +439,21 @@ export async function importVideoAction( globalSkipLiveDownloads: settings.skipLiveDownloads, appendArchive: true, }); + // ADD-ONLY: record the imported video in the roster so a one-off import + // is a known member of the channel rather than an orphan dir, and so a + // later retry of the same URL is possible even if it never appears in a + // listing. An import is not an enumeration, so lastListedAt does not + // move (see common/controller/rosterStore.ts). if (videoId) { + await mergeRosterFile( + paths, + slug, + [{ id: videoId, url: videoUrl }], + new Date().toISOString(), + "import", + ).catch(() => { + /* the download succeeded; a roster write failure must not fail it */ + }); revalidatePath(`/channels/${slug}/videos/${videoId}`); } revalidatePath(`/channels/${slug}`); diff --git a/editor/app/settings/actions.ts b/editor/app/settings/actions.ts @@ -203,6 +203,10 @@ export async function saveSettingsAction( "syncSchedulerFullSweepConfirmMaxSuspects", current.syncScheduler.fullSweepConfirmMaxSuspects, ), + fullSweepShrinkGuardPercent: intOrKeep( + "syncSchedulerFullSweepShrinkGuardPercent", + current.syncScheduler.fullSweepShrinkGuardPercent, + ), }; let socialInput: unknown; diff --git a/editor/app/settings/components/SettingsForm.tsx b/editor/app/settings/components/SettingsForm.tsx @@ -372,6 +372,15 @@ export function SettingsForm({ initial, apps, digestApps }: Props) { hint="When a sweep finds at most this many videos missing from the listing, it probes each one upstream to confirm; above the cap it only flags them and leaves the probing to 'Check maybe-missing'. 0 = never auto-confirm. This is a count of videos, not a duration." /> <Field + label="Full-sweep shrink guard" + name="syncSchedulerFullSweepShrinkGuardPercent" + defaultValue={String( + initial.syncScheduler.fullSweepShrinkGuardPercent, + )} + type="number" + hint="If a fresh listing comes back smaller than the stored one by more than this percentage (and by more than 25 entries), it is treated as suspect: the stored video list and the missing-video flags are left alone and the sweep is retried on the next sync rather than after the full cadence. A second enumeration reporting a similar count confirms it and is accepted, so a genuine mass deletion still lands. 0 = off; an empty listing is always refused. This is a percentage, not a duration." + /> + <Field label="Max concurrent syncs" name="syncSchedulerMaxConcurrentSyncs" defaultValue={String(initial.syncScheduler.maxConcurrentSyncs)} diff --git a/editor/e2e/sync-deep.spec.ts b/editor/e2e/sync-deep.spec.ts @@ -1,4 +1,4 @@ -import { readFile, writeFile } from "node:fs/promises"; +import { readFile, rm, writeFile } from "node:fs/promises"; import { test, expect } from "@playwright/test"; import { generateReport, @@ -43,6 +43,73 @@ type ChannelConfigFile = { lastFullSweepAt?: string; }; +type RosterFile = { + version: number; + lastSweep: { at: string; listedCount: number; verdict: string } | null; + entries: Record< + string, + { url: string; firstSeenAt: string; lastListedAt: string; source: string } + >; +}; + +type SnapshotFile = { + missingNeverFetched?: { id: string; url: string; firstSeenAt: string }[]; +}; + +// The shrink guard has an absolute floor of 25 entries below which no drop is +// suspicious, so exercising it needs a channel bigger than the 6-video fixture. +// These ids never get directories — the archive is seeded with them so the +// sweep's download walk treats them as already fetched. +const BULK_IDS = Array.from( + { length: 60 }, + (_, i) => `bulk${String(i + 1).padStart(8, "0")}`, +); + +function ytUrl(id: string): string { + return `https://www.youtube.com/watch?v=${id}`; +} + +async function readRoster(): Promise<RosterFile> { + return readJson<RosterFile>(`${channelRoot}/roster.json`); +} + +async function writePlaylist(ids: string[]): Promise<void> { + await writeFile( + resolvePath(`${channelRoot}/playlist`), + ids.map(ytUrl).join("\n") + "\n", + ); +} + +// generateReport short-circuits when a snapshot already exists, so a spec that +// needs the report REDONE after a mutation has to clear it first. +async function regenerateReport( + page: import("@playwright/test").Page, +): Promise<void> { + await rm(resolvePath(`${channelRoot}/snapshot.json`), { force: true }); + await generateReport(page, CHANNEL); +} + +// Force a sweep regardless of the cadence — needed once lastFullSweepAt has +// been stamped by an accepted sweep. +async function runFullSweep( + page: import("@playwright/test").Page, + previousLastSyncedAt: string | undefined, +): Promise<void> { + await page.goto(`/channels/${CHANNEL}`); + for (let attempt = 0; attempt < 5; attempt++) { + await page + .getByRole("button", { name: "Full sweep", exact: true }) + .click({ timeout: 5_000 }) + .catch(() => {}); + for (let i = 0; i < 60; i++) { + const cfg = await readConfig().catch(() => null); + if (cfg?.lastSyncedAt && cfg.lastSyncedAt !== previousLastSyncedAt) return; + await new Promise((r) => setTimeout(r, 250)); + } + } + throw new Error("runFullSweep: lastSyncedAt never advanced"); +} + // Controls which ids the fake yt-dlp's flat-playlist branch emits. Its cwd // sidecar serves the sweep's single unranged call exactly as it serves the // quick check's. @@ -73,7 +140,10 @@ async function seedStalePlaylist(): Promise<void> { ); } -async function enableFullSweep(confirmMaxSuspects: number): Promise<void> { +async function enableFullSweep( + confirmMaxSuspects: number, + shrinkGuardPercent = 10, +): Promise<void> { await writeSettings({ adminTitle: "Test Admin", maxTranscriptPageBytes: 8388608, @@ -85,6 +155,7 @@ async function enableFullSweep(confirmMaxSuspects: number): Promise<void> { enabled: false, fullSweepIntervalMinutes: 1440, fullSweepConfirmMaxSuspects: confirmMaxSuspects, + fullSweepShrinkGuardPercent: shrinkGuardPercent, }, }); } @@ -219,6 +290,154 @@ test("over the confirm cap, suspects are flagged but not probed", async ({ expect(deleted.availability).not.toBe("deleted"); }); +// --- Part C: the roster, the shrink guard, and never-fetched losses --------- + +test("a suspicious shrink is refused, then accepted when a second enumeration agrees", async ({ + page, +}) => { + await resetData(FIXTURE); + await enableFullSweep(0); + const full = [...ALL_IDS, ...BULK_IDS]; + await seedArchive(full); + // The stored listing a previous sweep left behind: 66 entries, of which 60 + // were never downloaded and so have no dir and no metadata.info.json. This + // file is the ONLY place their URLs live. + await writePlaylist(full); + // ...and the fetch comes back with 6 of them. Acting on it would truncate the + // stored playlist and flag 60 videos as gone. + await setFreshPlaylist(ALL_IDS); + + await generateReport(page, CHANNEL); + await runSync(page, undefined); + + // 1. Refused. The playlist is untouched, and no missing-video flags were + // written at all. + await expect(page.getByLabel("Sync output")).toContainText("down 60 from 66"); + expect(await readPlaylist()).toHaveLength(full.length); + await expect( + readJson<MaybeMissing>(`${channelRoot}/maybe-missing.json`), + ).rejects.toThrow(); + + // 2. But the roster was still seeded and merged — additive, so a bad fetch + // can only ever add. Every URL survived the refusal. + const afterShrink = await readRoster(); + expect(Object.keys(afterShrink.entries)).toHaveLength(full.length); + expect(afterShrink.entries.bulk00000001.url).toBe(ytUrl("bulk00000001")); + expect(afterShrink.lastSweep?.verdict).toBe("shrink-suspect"); + expect(afterShrink.lastSweep?.listedCount).toBe(ALL_IDS.length); + // Not counted as a sweep, which is what makes the next ordinary sync retry + // the enumeration instead of waiting out the daily cadence. + expect((await readConfig()).lastFullSweepAt).toBeUndefined(); + + // 3. The next ORDINARY sync re-enumerates (not the cheap paged walk) and gets + // the same answer. A real mass deletion repeats; a transient blip does not. + await runSync(page, (await readConfig()).lastSyncedAt); + + await expect(page.getByLabel("Sync output")).toContainText( + "confirmed by a second enumeration", + ); + expect(typeof (await readConfig()).lastFullSweepAt).toBe("string"); + expect(await readPlaylist()).toHaveLength(ALL_IDS.length); + const confirmed = await readRoster(); + expect(confirmed.lastSweep?.verdict).toBe("shrink-confirmed"); + // Nothing was ever dropped from the roster, so the 60 departed videos are + // still recoverable by URL. + expect(Object.keys(confirmed.entries)).toHaveLength(full.length); + expect(confirmed.entries.bulk00000001.url).toBe(ytUrl("bulk00000001")); +}); + +test("a never-downloaded video that leaves the listing is surfaced with its URL", async ({ + page, +}) => { + await resetData(FIXTURE); + await enableFullSweep(0); + // ghost is listed but never fetched: no data/ dir, so nothing on disk has + // ever recorded its URL. + const listed = [...ALL_IDS, "ghost0000001"]; + await seedArchive(listed); + await setFreshPlaylist(listed); + + await generateReport(page, CHANNEL); + await runSync(page, undefined); + expect((await readRoster()).entries.ghost0000001.url).toBe( + ytUrl("ghost0000001"), + ); + + // It leaves the listing. A drop of one is far under the guard's floor, so + // this is an ordinary accepted sweep. + await setFreshPlaylist(ALL_IDS); + await runFullSweep(page, (await readConfig()).lastSyncedAt); + + await expect(page.getByLabel("Full sweep output")).toContainText( + "listed but never downloaded", + ); + // It is NOT maybe-missing: that bucket means "on disk and no longer listed", + // and its meaning is unchanged. + const record = await readJson<MaybeMissing>( + `${channelRoot}/maybe-missing.json`, + ); + expect(record.ids).toEqual([]); + + await regenerateReport(page); + const snapshot = await readJson<SnapshotFile>( + `${channelRoot}/snapshot.json`, + ); + expect(snapshot.missingNeverFetched).toEqual([ + expect.objectContaining({ + id: "ghost0000001", + url: ytUrl("ghost0000001"), + }), + ]); + + // And it is rendered, with a usable link and a recovery action. + await page.goto(`/channels/${CHANNEL}`); + const row = page.getByLabel("never fetched ghost0000001"); + await expect(row).toBeVisible(); + await expect(page.getByLabel("never fetched url ghost0000001")).toHaveAttribute( + "href", + ytUrl("ghost0000001"), + ); + await expect( + row.getByRole("button", { name: "Try downloading anyway" }), + ).toBeVisible(); +}); + +test("seeding captures an undownloaded video's URL before the first sweep can erase it", async ({ + page, +}) => { + await resetData(FIXTURE); + await enableFullSweep(0); + await seedArchive(ALL_IDS); + // The pre-existing state every channel is in today: a stored playlist, no + // roster, and an entry in it that was never downloaded. `playlist` is the + // ONLY place ghost's URL lives. + await writePlaylist([...ALL_IDS, "ghost0000001"]); + // ...and it has already gone from the channel, so the very first sweep would + // have overwritten the playlist and taken the URL with it. + await setFreshPlaylist(ALL_IDS); + + await generateReport(page, CHANNEL); + await runSync(page, undefined); + + const roster = await readRoster(); + expect(roster.entries.ghost0000001).toBeTruthy(); + expect(roster.entries.ghost0000001.url).toBe(ytUrl("ghost0000001")); + expect(roster.entries.ghost0000001.source).toBe("listing"); + // Every on-disk video is in the roster too, so nothing is left as an orphan. + for (const id of ALL_IDS) expect(roster.entries[id]).toBeTruthy(); + + // The playlist was legitimately refreshed — the migration ran first. + expect(await readPlaylist()).toHaveLength(ALL_IDS.length); + + await regenerateReport(page); + const snapshot = await readJson<SnapshotFile>( + `${channelRoot}/snapshot.json`, + ); + expect(snapshot.missingNeverFetched?.map((v) => v.id)).toEqual([ + "ghost0000001", + ]); +}); + test("the next sync stays on the cheap paged walk until the cadence elapses", async ({ page, }) => {