commit 7818ca35a03d72fc0d0f81a749c8083510d22dfb
parent 0a17900565c266621f0e141f66e6386923575717
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Sun, 20 Sep 2026 00:36:38 -0400
plans: FACTS — why the scan writes no video directory, and why settled is derived
Records the two verified readers that make `data/<id>/` the wrong place for a
metadata-only pass (buildIndex admits any dir with a metadata.info.json;
deriveChannelSets keys "ever fetched" off the dir names), the derived-not-stored
settlement rule and its two fail-closed edges, the no-artifact rule that keeps a
downloaded video from being un-counted, why sync must treat a settled id as an
archived hit, and the bot-check classification. Ends with the one thing this
branch deliberately did not do: nothing auto-dispatches the scan yet.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Diffstat:
| M | plans/FACTS.md | | | 122 | +++++++++++++++++++++++++++++++++++++++++++++++-------------------------------- |
1 file changed, 73 insertions(+), 49 deletions(-)
diff --git a/plans/FACTS.md b/plans/FACTS.md
@@ -4048,53 +4048,77 @@ never pulls `reader-fs.ts` into a client chunk.
23-min idle-machine baseline; two other Opus implementers were building in sibling
worktrees at the time. Wall time is not a signal here; the pass count is.
-## Per-channel download filter (verified 2026-09-20, branch `feat/download-title-filter`)
-
-- **A declined video is settled by an OUTCOME, never by an archive line.** Three separate
- mechanisms make an archive line the wrong tool, and each was checked:
- `common/controller/channelSnapshot.ts:1178-1196` derives `undownloadedIds` from
- ARTIFACTS on disk (`videoHasAnyArtifact`), so an archived id with no files is still
- handed to the auto-download runner; `common/controller/verifyTranscripts.ts:47` and the
- `missingFromArchive` bucket both read an archive line as "this was downloaded", so the
- channel would report ~1,800 broken videos; and an archive line carries no record of WHY,
- so changing the filter could never undo it. The outcome is
- `status: "skipped-filtered"` (no enum change — `autoRunner.ts:1980` and
- `runYtdlp.ts:916` already count that status as skipped) plus
- `filter.permanent` + `filter.signature` on `DownloadOutcomeRecord`
- (`common/lib/downloadOutcome.ts`).
-- **The signature is the expiry mechanism.** `downloadFilterSignature()` is
- `"v1\0" + include + "\0" + exclude`, null when both are blank
- (`common/lib/downloadFilters.ts`). `isSettledByFilter(outcome, config)` is the ONE
- definition of settled — a permanent filter skip whose recorded signature still equals the
- channel's current one. Editing either pattern changes the signature, so every video the
- old filter decided is re-evaluated on the next report with no migration, no sweep and no
- stored "generation" counter. `isVideoSettledByFilter(videoDir, config)`
- (`common/lib/downloadOutcome-server.ts`) is its per-video-dir form, shaped like
- `isDeferredAuthExcluded` (`common/ytdlp/runYtdlp.ts:247`) and short-circuiting on the
- signature BEFORE the file read, so a channel with no filter pays nothing.
-- **`buckets.skippedByTitleFilter` is a short-circuit, beside the other two terminal
- outcomes.** It sits immediately after the `failed-short-audio` `continue` in the per-video
- loop (`channelSnapshot.ts`), and it has to: everything below classifies a video by what is
- MISSING from its directory, and a settled video (one metadata.info.json, nothing else)
- would otherwise land in `noTranscript` AND `noMetadata` AND the retryable
- `skippedByFilter` simultaneously, once per non-match. It is also subtracted from
- `totals.videos` and excluded from `undownloadedIds`.
-- **`skippedByFilter` and `skippedByTitleFilter` are different things and must stay
- separate.** The first is RETRYABLE (skip-live: the stream ends and the video becomes
- downloadable, so every sync tries again — the decision carries no `permanent`/`signature`).
- The second is the operator's own "not this one". A configured filter that meets
- `metadata === null` deliberately produces the FIRST kind: it fails closed (does not
- download) but records a non-permanent skip, so a flaky extractor cannot settle anything.
+## Per-channel download filter + metadata scan (verified 2026-09-20, branch `feat/download-title-filter`)
+
+- **The metadata scan NEVER creates `data/<id>/`, and that is not a style
+ preference.** Two readers treat a video directory as a fact about the corpus:
+ `common/controller/buildIndex.ts:309-315` admits ANY `data/<id>/` that holds a
+ `metadata.info.json`, transcript or not — so a metadata-only pass that made
+ directories would put ~1,800 text-less videos into the LMDB index and into the
+ published site — and `deriveChannelSets` (`common/controller/channelSets.ts:58-67`,
+ fed at `channelSnapshot.ts:1167-1174`) keys "ever fetched" off those same dir
+ names, so a scanned id would stop being reported in `missingNeverFetched`, the
+ one record of a video that vanished before we ever got it. The store is
+ therefore channel-level: `channels/<slug>/metadata-scan.json`
+ (`common/controller/metadataScanStore.ts`), modelled on `rosterStore.ts`
+ (versioned, normalize-on-read, additive, write skipped when unchanged). The e2e
+ asserts the absence of the directories directly rather than inferring it.
+- **"Settled" is DERIVED, never stored.** `settledByTitleFilterIds(paths, slug,
+ config)` = the scan's entries that the channel's CURRENT `downloadFilter`
+ rejects. There is no signature, no generation counter and nothing per video:
+ editing a regex re-decides the whole channel on the next snapshot, with no
+ rescan and nothing to migrate — asserted in e2e by a byte-identical
+ `fake-ytdlp.invocations` across the flip. `DownloadOutcomeRecord.filter` is
+ `{name, reason}` and deliberately carries no verdict.
+- **It fails closed in both directions.** No filter, or an unparseable one
+ (which is inert at download time too), settles nothing. An id the scan only has
+ an ERROR for is never settled — we do not know what it is called, so we must
+ not decide it; it stays ordinary download work.
+- **Settlement applies only to a video with NO artifact on disk.** A video
+ downloaded before the filter was written is DOWNLOADED, which is a fact rather
+ than a preference. Short-circuiting it in the snapshot's per-video loop would
+ drop it from `transcribedWithAudio` and the cleanup estimates while it still
+ counted in `totals.transcribed` — a channel reporting more transcripts than
+ videos. Both the loop and the `skippedByTitleFilter` bucket apply the same
+ no-artifact rule so they cannot disagree.
+- **`buckets.skippedByTitleFilter` vs `buckets.skippedByFilter`.** The first is
+ settled (the operator's "not this one"); the second is the RETRYABLE skip that
+ skip-live has always produced, and a configured filter meeting
+ `metadata === null` deliberately lands there — a flaky extractor must not be
+ able to settle anything. The settled ids are also excluded from
+ `undownloadedIds` (artifact-derived, which is why an archive line could never
+ have settled anything: archive lines also mean "downloaded" to
+ `verifyTranscripts.ts:47` and the `missingFromArchive` bucket).
- **Sync counts a settled id as an ARCHIVED HIT** (`selectDownloadableUrls`,
- `common/ytdlp/runYtdlp.ts`). `archivedHits > 0` is the paged walk's stop signal, and a
- settled video never gets an archive line; without this a filtered channel's daily sync
- walks every page of non-matches forever and never sees a hit on the pages the filter
- emptied. The count is reported separately in the log so it doesn't claim they were
- archived. `downloadPlaylistManaged`'s exclusion pass skips them too, first and cheapest.
-- **The registry order is `[titleFilter, skipLive]`** and the order is load-bearing: a
- non-matching video that is also currently live must be settled, not classified as a
- live-stream retry. A video that PASSES the include still falls through to skip-live.
-- **An invalid regex makes the filter inert, plus one log line** — it never fails a
- download. The editor form is what refuses to save one
- (`editor/app/channels/components/parseChannelForm.ts`, `new RegExp(v, "i")`), because a bad
- pattern on disk would otherwise present as a channel that silently downloads everything.
+ `common/ytdlp/runYtdlp.ts`). `archivedHits > 0` is the paged walk's stop signal;
+ a settled video has no archive line and usually no directory at all, so without
+ this a filtered channel's daily sync walks every page of non-matches forever
+ and never sees a hit on the pages the filter emptied. The set is resolved ONCE
+ per run (one small JSON read), not per video.
+- **The registry order is `[titleFilter, skipLive]`** and it is load-bearing: a
+ non-matching video that is also live must be declined on the operator's terms,
+ not classified as a live-stream retry. A video that passes the include still
+ falls through to skip-live.
+- **The scan is a catalogued operation**, `metadata-scan`: second channel-scoped
+ entry in `operationCatalog()`, `trigger: "backlog"` (sync stays the ONE
+ `"cadence"` entry — that is how `/operations/sync` is chosen), platform download
+ queue, `needsMedia: false`, and `pauseLaneFor` answers null for it. It is
+ deliberately NOT held by the downloads pause: it fetches no media, and the
+ operator runs it precisely to decide what a paused lane should fetch when it
+ resumes. Its backlog is `snapshot.metadataScan.unscanned`, which must keep
+ meaning exactly what `metadataScanTargets()` will fetch.
+- **A scan error suppresses a re-scan of that id for 24 h**
+ (`METADATA_SCAN_ERROR_COOLDOWN_MS`). Without it a members-only video keeps the
+ backlog above zero forever and the operation page never stops offering work.
+- **YouTube's bot check is `rate_limit`-class.** "Sign in to confirm you're not
+ a bot" (right single quote in yt-dlp's output — match either) now classifies in
+ `classifyDownloadFailure`. It is not a 429 and not per-video: once it fires,
+ every remaining entry in the batch fails the same way, so the scan kills the
+ child, records the shared per-platform cooldown and keeps what it flushed
+ (every 25 records). Note it is NOT in `parseUnavailableFromStderr` — that
+ function's `/confirm your age/` branch does not match it, which is what lets it
+ fall through to the failure classifier.
+- **FOLLOW-UP, not done here: nothing auto-dispatches the scan.** The descriptor
+ has no `runner`, so the operator presses Run. The obvious next step is for the
+ download lane to run it for a channel whose filter has unscanned listed videos
+ — which is exactly the backlog the snapshot already carries.