Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 7818ca35a03d72fc0d0f81a749c8083510d22dfb
parent 0a17900565c266621f0e141f66e6386923575717
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sun, 20 Sep 2026 00:36:38 -0400

plans: FACTS — why the scan writes no video directory, and why settled is derived

Records the two verified readers that make `data/<id>/` the wrong place for a
metadata-only pass (buildIndex admits any dir with a metadata.info.json;
deriveChannelSets keys "ever fetched" off the dir names), the derived-not-stored
settlement rule and its two fail-closed edges, the no-artifact rule that keeps a
downloaded video from being un-counted, why sync must treat a settled id as an
archived hit, and the bot-check classification. Ends with the one thing this
branch deliberately did not do: nothing auto-dispatches the scan yet.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
Mplans/FACTS.md | 122+++++++++++++++++++++++++++++++++++++++++++++++--------------------------------
1 file changed, 73 insertions(+), 49 deletions(-)

diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -4048,53 +4048,77 @@ never pulls `reader-fs.ts` into a client chunk. 23-min idle-machine baseline; two other Opus implementers were building in sibling worktrees at the time. Wall time is not a signal here; the pass count is. -## Per-channel download filter (verified 2026-09-20, branch `feat/download-title-filter`) - -- **A declined video is settled by an OUTCOME, never by an archive line.** Three separate - mechanisms make an archive line the wrong tool, and each was checked: - `common/controller/channelSnapshot.ts:1178-1196` derives `undownloadedIds` from - ARTIFACTS on disk (`videoHasAnyArtifact`), so an archived id with no files is still - handed to the auto-download runner; `common/controller/verifyTranscripts.ts:47` and the - `missingFromArchive` bucket both read an archive line as "this was downloaded", so the - channel would report ~1,800 broken videos; and an archive line carries no record of WHY, - so changing the filter could never undo it. The outcome is - `status: "skipped-filtered"` (no enum change — `autoRunner.ts:1980` and - `runYtdlp.ts:916` already count that status as skipped) plus - `filter.permanent` + `filter.signature` on `DownloadOutcomeRecord` - (`common/lib/downloadOutcome.ts`). -- **The signature is the expiry mechanism.** `downloadFilterSignature()` is - `"v1\0" + include + "\0" + exclude`, null when both are blank - (`common/lib/downloadFilters.ts`). `isSettledByFilter(outcome, config)` is the ONE - definition of settled — a permanent filter skip whose recorded signature still equals the - channel's current one. Editing either pattern changes the signature, so every video the - old filter decided is re-evaluated on the next report with no migration, no sweep and no - stored "generation" counter. `isVideoSettledByFilter(videoDir, config)` - (`common/lib/downloadOutcome-server.ts`) is its per-video-dir form, shaped like - `isDeferredAuthExcluded` (`common/ytdlp/runYtdlp.ts:247`) and short-circuiting on the - signature BEFORE the file read, so a channel with no filter pays nothing. -- **`buckets.skippedByTitleFilter` is a short-circuit, beside the other two terminal - outcomes.** It sits immediately after the `failed-short-audio` `continue` in the per-video - loop (`channelSnapshot.ts`), and it has to: everything below classifies a video by what is - MISSING from its directory, and a settled video (one metadata.info.json, nothing else) - would otherwise land in `noTranscript` AND `noMetadata` AND the retryable - `skippedByFilter` simultaneously, once per non-match. It is also subtracted from - `totals.videos` and excluded from `undownloadedIds`. -- **`skippedByFilter` and `skippedByTitleFilter` are different things and must stay - separate.** The first is RETRYABLE (skip-live: the stream ends and the video becomes - downloadable, so every sync tries again — the decision carries no `permanent`/`signature`). - The second is the operator's own "not this one". A configured filter that meets - `metadata === null` deliberately produces the FIRST kind: it fails closed (does not - download) but records a non-permanent skip, so a flaky extractor cannot settle anything. +## Per-channel download filter + metadata scan (verified 2026-09-20, branch `feat/download-title-filter`) + +- **The metadata scan NEVER creates `data/<id>/`, and that is not a style + preference.** Two readers treat a video directory as a fact about the corpus: + `common/controller/buildIndex.ts:309-315` admits ANY `data/<id>/` that holds a + `metadata.info.json`, transcript or not — so a metadata-only pass that made + directories would put ~1,800 text-less videos into the LMDB index and into the + published site — and `deriveChannelSets` (`common/controller/channelSets.ts:58-67`, + fed at `channelSnapshot.ts:1167-1174`) keys "ever fetched" off those same dir + names, so a scanned id would stop being reported in `missingNeverFetched`, the + one record of a video that vanished before we ever got it. The store is + therefore channel-level: `channels/<slug>/metadata-scan.json` + (`common/controller/metadataScanStore.ts`), modelled on `rosterStore.ts` + (versioned, normalize-on-read, additive, write skipped when unchanged). The e2e + asserts the absence of the directories directly rather than inferring it. +- **"Settled" is DERIVED, never stored.** `settledByTitleFilterIds(paths, slug, + config)` = the scan's entries that the channel's CURRENT `downloadFilter` + rejects. There is no signature, no generation counter and nothing per video: + editing a regex re-decides the whole channel on the next snapshot, with no + rescan and nothing to migrate — asserted in e2e by a byte-identical + `fake-ytdlp.invocations` across the flip. `DownloadOutcomeRecord.filter` is + `{name, reason}` and deliberately carries no verdict. +- **It fails closed in both directions.** No filter, or an unparseable one + (which is inert at download time too), settles nothing. An id the scan only has + an ERROR for is never settled — we do not know what it is called, so we must + not decide it; it stays ordinary download work. +- **Settlement applies only to a video with NO artifact on disk.** A video + downloaded before the filter was written is DOWNLOADED, which is a fact rather + than a preference. Short-circuiting it in the snapshot's per-video loop would + drop it from `transcribedWithAudio` and the cleanup estimates while it still + counted in `totals.transcribed` — a channel reporting more transcripts than + videos. Both the loop and the `skippedByTitleFilter` bucket apply the same + no-artifact rule so they cannot disagree. +- **`buckets.skippedByTitleFilter` vs `buckets.skippedByFilter`.** The first is + settled (the operator's "not this one"); the second is the RETRYABLE skip that + skip-live has always produced, and a configured filter meeting + `metadata === null` deliberately lands there — a flaky extractor must not be + able to settle anything. The settled ids are also excluded from + `undownloadedIds` (artifact-derived, which is why an archive line could never + have settled anything: archive lines also mean "downloaded" to + `verifyTranscripts.ts:47` and the `missingFromArchive` bucket). - **Sync counts a settled id as an ARCHIVED HIT** (`selectDownloadableUrls`, - `common/ytdlp/runYtdlp.ts`). `archivedHits > 0` is the paged walk's stop signal, and a - settled video never gets an archive line; without this a filtered channel's daily sync - walks every page of non-matches forever and never sees a hit on the pages the filter - emptied. The count is reported separately in the log so it doesn't claim they were - archived. `downloadPlaylistManaged`'s exclusion pass skips them too, first and cheapest. -- **The registry order is `[titleFilter, skipLive]`** and the order is load-bearing: a - non-matching video that is also currently live must be settled, not classified as a - live-stream retry. A video that PASSES the include still falls through to skip-live. -- **An invalid regex makes the filter inert, plus one log line** — it never fails a - download. The editor form is what refuses to save one - (`editor/app/channels/components/parseChannelForm.ts`, `new RegExp(v, "i")`), because a bad - pattern on disk would otherwise present as a channel that silently downloads everything. + `common/ytdlp/runYtdlp.ts`). `archivedHits > 0` is the paged walk's stop signal; + a settled video has no archive line and usually no directory at all, so without + this a filtered channel's daily sync walks every page of non-matches forever + and never sees a hit on the pages the filter emptied. The set is resolved ONCE + per run (one small JSON read), not per video. +- **The registry order is `[titleFilter, skipLive]`** and it is load-bearing: a + non-matching video that is also live must be declined on the operator's terms, + not classified as a live-stream retry. A video that passes the include still + falls through to skip-live. +- **The scan is a catalogued operation**, `metadata-scan`: second channel-scoped + entry in `operationCatalog()`, `trigger: "backlog"` (sync stays the ONE + `"cadence"` entry — that is how `/operations/sync` is chosen), platform download + queue, `needsMedia: false`, and `pauseLaneFor` answers null for it. It is + deliberately NOT held by the downloads pause: it fetches no media, and the + operator runs it precisely to decide what a paused lane should fetch when it + resumes. Its backlog is `snapshot.metadataScan.unscanned`, which must keep + meaning exactly what `metadataScanTargets()` will fetch. +- **A scan error suppresses a re-scan of that id for 24 h** + (`METADATA_SCAN_ERROR_COOLDOWN_MS`). Without it a members-only video keeps the + backlog above zero forever and the operation page never stops offering work. +- **YouTube's bot check is `rate_limit`-class.** "Sign in to confirm you're not + a bot" (right single quote in yt-dlp's output — match either) now classifies in + `classifyDownloadFailure`. It is not a 429 and not per-video: once it fires, + every remaining entry in the batch fails the same way, so the scan kills the + child, records the shared per-platform cooldown and keeps what it flushed + (every 25 records). Note it is NOT in `parseUnavailableFromStderr` — that + function's `/confirm your age/` branch does not match it, which is what lets it + fall through to the failure classifier. +- **FOLLOW-UP, not done here: nothing auto-dispatches the scan.** The descriptor + has no `runner`, so the operator presses Run. The obvious next step is for the + download lane to run it for a channel whose filter has unscanned listed videos + — which is exactly the backlog the snapshot already carries.