Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 2f7002b97269125ea697294f19d3d98df68040d5
parent 1e486c64d1ffcd52a9c8a4f2e9bf7ef731640a64
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Fri, 25 Sep 2026 01:39:18 -0400

plans: release 5 rollout record — 05877ffb live on :3001 (releases 4+5), exports off on five sites, Rumble proven live; STATE + FACTS; YouTube lane pacing plan filed

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
Mplans/FACTS.md | 49+++++++++++++++++++++++++++++++++++++++++++++++++
Mplans/STATE.md | 40+++++++++++++++++++++++++++++++---------
Mplans/release-5.md | 158+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aplans/youtube-lane-pacing.md | 149+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
4 files changed, 387 insertions(+), 9 deletions(-)

diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -6194,3 +6194,52 @@ Line numbers are `plans/FACTS.md` lines at `e172749b`, before this record's in-p is historical. `normalizeTranscript.ts:141` is now `writeJsonAtomic(cuesPath, out, {indent: 0, newline: false})`. The idiom for any file is `writeFileAtomic` / `writeJsonAtomic`. + +## Release 5 — slices R, X (verified 2026-09-24/25, `main` @ `93dcb532`) + +- **One yt-dlp arg builder with a platform table.** `common/ytdlp/channelArgs.ts`: + `PLATFORM_ARGS` :30 (`rumble: --impersonate chrome --sleep-requests 1`), `platformArgs` :34, + `channelPlatform` :39 (`config.platform ?? detectPlatform(config.url)`), `channelExtraArgs` :45 + = cookies → platform args → `ytdlpExtraArgs` (a channel override is last and wins). `configArgs` + (runYtdlp), the metadata scan and `checkAvailability` all delegate to it; `probeChannelMeta` + and the no-config availability path call `platformArgs(detectPlatform(url))`. The table is code; + there is no setting. +- **A 429 mid-sweep is "incomplete".** `EnumerationIncompleteError` `runYtdlp.ts:329` + (`platform, pagesReached, count`), `lastListingPage` :353 (last `Downloading page N` — the real + wording is `[RumbleChannel] TheQuartering: Downloading page 155`). `syncFullSweep` catches it + before `acceptEnumeration`: best-effort `opts.onPlatformBackoff("rate_limit")` :1644 (= + `recordDownloadBackoff`, the download runner's helper and schedule), the `Full sweep incomplete` + line :1653, then `return syncPaged(opts)` :1658. No playlist write, no `lastFullSweepAt`. Any + other non-zero exit throws as before; exit 101 untouched. `fullSweepDue` :1408 is exported, async, + and false while `platformCooldownRemainingMs(detectPlatform(url) ?? "unknown") > 0` — the Sync + gate's key. Syncs are REFUSED during the cooldown (`pipelineActions.ts` Sync gate), and the sweep + is due again the moment it ends — no last-attempt stamp exists (follow-up if live sweeps keep + coming back incomplete). +- **A bare `HTTP Error 403` is `network`** (`common/lib/availability.ts:255`), so a Cloudflare + block backs the platform off; there is no `blocked` class because `DownloadFailureClass` is + persisted in every `download-outcome.json`. +- **Fake yt-dlp** records every flat-playlist argv as `flat-playlist:full|paged argv=…` in the + per-channel `fake-ytdlp.invocations` (`fake-ytdlp.mjs:543`); the `sweep429` URL sentinel :549 + prints 3 pages × 5 urls then the real 429 line and exits 1 on a FULL enumeration only. +- **`transcriptDownloads`** (absent = on) lives in three files and one context: `site.json` + (`siteSchema.ts:88` type, :122 docs, :245 zod, :328 write-if-false), `homepage.json` for the hub + (`homepage.ts:39,83,133` — the hub has NO site.json; `hubSite()` `export/app/lib/site.ts:40` passes + it through), and `PlayerProvider` `features` (`PlayerProvider.tsx:118`, default all on); + `TranscriptModal.tsx:79` `showDownloads` hides exactly Copy download command, the Download + dropdown and Copy MD. FOUR mounts, all in `export/` (`(workspace)/layout.tsx` → `SiteWorkspace`, + `duplicates/page.tsx`, `HubHome`, `AskHub` via their server pages); the editor mounts neither. + `SITE.md` is generated; `homepage.json` has no generated doc and no zod schema. +- **The export e2e off-test flips the committed fixture** `export/e2e/fixtures/sites/testsite/site.json` + for one test; `export/playwright.config.ts:38` strips a leftover key at config load (before any + spec), so a killed run leaves at most a whitespace diff. `workers: 1`, so no other test can see + the flip. +- **Rollout facts.** `ops channel-config` with `patch: {"<field>": null}` CLEARS a Configure-form + field (the route lays the patch over `channelConfigToFormData(existing)` and + `updateChannelAction` unsets every form field first) — that is how `fullSweepIntervalMinutes: 0` + was removed from `the-quartering-rumble`. A site's `archives` / `transcriptDownloads` are written + only by the Settings form (`saveSiteAction` → `writeSite`); there is no `ops site-config`. +- **A Settings-form save materialises missing `order` fields** on a site's groups and channels + (entries added through the CLI carry none). Group order is consumed by `sortGroups` + (`channelGroups.ts:118`, missing = last); channel `order` is parsed (`siteSchema.ts:154`) and + read by nothing. So a "no-op" site save can move bytes; `jq -S` plus the `order` fills is the + expected diff, not a bug. diff --git a/plans/STATE.md b/plans/STATE.md @@ -3,15 +3,37 @@ The working memory for the local-AI derived-corpus work. Rewritten at the end of every session, before context is cleared. See [`README.md`](README.md) for the protocol. -**Live on :3001 (2026-09-24, 18:42):** still `9ab10d77` (release 3), `BUILD_ID` -`XsKaA_drqdAbTguxUGVsn`. **Release 4 is on `main` but NOT rolled out**; the next release's -prelude rolls it out. The one-sync md5 sweep from the release-3 rollout is still owed. -YouTube held a 429 cooldown across three attempts (18:43, 19:13, 19:41). See the -[rollout record](one-core-phase-3.md#rollout-2026-09-24--9ab10d77-live-on-3001). - -**Release 5 in flight (2026-09-24, late):** slices R (Rumble) and X (exports-off) cut off -`8b7f8924` as `one-core/r5-rumble` and `one-core/r5-exports`; plan, record and rollout in -[`release-5.md`](release-5.md); rules in [`tools/implementer-rules.md`](tools/implementer-rules.md). +**Live on :3001 (2026-09-24, 23:39):** `93dcb532` (releases 4 + 5 together), `BUILD_ID` +`FKE60BTWpUiUa94WCSxTO`, next-server 888107. **Release 6 (follow-ups `4d97049f` + Phase 4 slice 1 +`cda35622`) is on `main` but NOT rolled out** — the next release's prelude rolls it out, after +`pnpm install --frozen-lockfile` in the primary (umtool→common link, aws-sdk now under common). The +one-sync md5 sweep owed since release 3 is DONE (`rekietalaw-rumble`, 4 Rumble videos archived — the +first since August). See the [release-5 rollout record](release-5.md#rollout-2026-09-24-night--93dcb532-live-on-3001-releases-4--5-together). + +**Last updated:** 2026-09-25 (early morning). **Release 5 is live; release 6 is merged.** +- Release 5 (`f4da04a9` → `93dcb532`): slice R (`one-core/r5-rumble` → `3adaea9b`): one yt-dlp arg + builder with a platform table (Rumble: `--impersonate chrome --sleep-requests 1` on every spawn), a + 429 mid-sweep is "incomplete" (paged walk + platform cooldown, never a listing, never a stamp), a + bare 403 backs the platform off. Slice X (`one-core/r5-exports` → `93dcb532`): `transcriptDownloads` + (absent = on) on `site.json` AND `homepage.json`, gating the three per-video export controls on + the export site; the site and hub forms carry the checkbox. Final suites on `93dcb532`: editor + 628/628 (no flake), export 192/192, hub 8/8. Record: [`release-5.md`](release-5.md). +- Rollout: ONE restart; smoke green; boot rewrote nothing; proof sync of `rekietalaw-rumble` and + the first complete `the-quartering-rumble` sweep in 44 days (8,045 listed, was 7,866, paced past + page 300 with no 429); exports off on the five published sites (form → build-index → build-deploy + each, verified at the public URLs; Anilyzer's owed production deploy rode along, now at spec 4). +- Release 6 (overnight, operator-approved): follow-ups (410 → deleted, umtool gets the platform args + through ONE table in `common/ytdlp/platformArgs.mjs`, atomic remote-transcript write, stale comment, + measure-nav routes) and Phase 4 slice 1 (`buildDeployCore.ts` → `common/publish/build.ts`, aws-sdk + deps to common; the move only — entry points, docker scripts and build:hub wait for the CLI slice). + Record: [`release-6.md`](release-6.md). 188 orphan `*.tmp-*` files deleted (filtered). +- Found and filed, not fixed: [`youtube-lane-pacing.md`](youtube-lane-pacing.md) — the evening's + three YouTube cooldowns were ONE Short retried 12× from the head of the queue; the fix is a 6-hour + per-video quarantine (small slice, recommended next), plus `metadataScanStore.ts:249-252` (scan + error timestamp never refreshes → 142 members-only videos re-scanned every runner start). + +**Next:** Phase 4 slice 2 (`common/bin/archilyzer.ts` + the rest of spec item 1), the pacing slice, +then the release-6 rollout (one restart) with the pacing fix riding along if it is ready. **Last updated:** 2026-09-24 (evening). **Release 4 is on `main` @ `e172749b`, and one-core Phase 3 is COMPLETE. Phase 4 is next.** The prelude was plans only: `4130aca1` (the rollout diff --git a/plans/release-5.md b/plans/release-5.md @@ -319,3 +319,161 @@ and the written file keeps both keys. and build + deploy each, then untick transcript downloads on the hub form and rebuild + deploy the hub. Verify per site as item 5 now says: no Downloads link in the header or footer, and `/downloads` shows its empty state (it still answers 200). + +## Rollout 2026-09-24 (night) — `93dcb532` live on :3001 (releases 4 + 5 together) + +ONE restart, at the END of release 5, as the plan said: R is what un-breaks Rumble syncing. +Editor only — `git diff --stat 9ab10d77..93dcb532 -- umtool` prints nothing, so umtool was +not rebuilt or restarted (still serving on :3050). The live editor had served `9ab10d77` +(release 3, `BUILD_ID` `XsKaA_drqdAbTguxUGVsn`) since the release-4 prelude. + +**Final suites on `93dcb532` first** (worktree `one-core-r5-exports` detached at the merge sha, +whose tree equals X's gated tip `d948aa0e`; ports 3211/3210/3220; logs +`final-e2e-r5-{editor,export,hub}.log`): editor **628 passed, 0 failed, 0 flaky, 39.4 min** +(release 4's `lane-runner.spec.ts:361` load flake did not recur); export **192/192, 6.2 min** +(188 + 4 new); hub **8/8, 16 s**. + +**Offline round-trips.** `phase3-settings-numbers.ts live=settings.json`: the written block +equals `jq -S . settings.json` (diff empty, 1,353 scalar paths each side — the same count as +release 4). `phase3-files-numbers.ts` over the LIVE `transcripts/`: 6 `site.json` + 71 +`config.json`, every `unknown keys: []` (77 lines), no `WRITE THREW`; the output is 3,839 +lines against release 4's 3,832 — the seven new lines are the `"transcriptDownloads": true` +each site's parse now carries (X's key, default on) plus the hub's. + +**Build and one restart.** md5 baseline over `settings.json` + 6 `site.json` + 71 +`config.json` (78 files). `pnpm --filter editor build` detached into the live `.next` while the +old server served: exit 0, compiled in 14.3 s, `ƒ /api/view/[name]` in the route table; +`BUILD_ID` `XsKaA_drqdAbTguxUGVsn` → **`FKE60BTWpUiUa94WCSxTO`**. Seven auto-lane jobs were in +flight (1 backfill, 1 digest, 2 download, 2 transcribe, 1 transcribe), as at release 4's +restart. Then TERM pnpm 525402 / next-server 525417 (cwd `editor/`), :3001 free, `setsid nohup +pnpm run start -H 0.0.0.0` → next-server **888107**, `/` 200 after 2 s, `Ready in 163ms`, +`/tags` 200. md5 sweep right after boot: identical to the baseline — boot rewrote nothing. + +**Smoke, started 4 min after boot** (`smoke10.sh`, the release-4 script). All eight old/new +pairs equal on the FIRST pass with the live fields stripped: `pulse` 109 B, `workers` 1,116 B, +`activeJobs` 2,652 B, `autoQueueStatus` 507,725 B (79 s), `schedulerStatus` 37,371 B, +`widgetSync` 1,522 B, `widgetActionable` 1,148 B, `cleanable` 1,677 B. `/api/view/bogus` and +`/api/view/invalidate-cache` 404; `/api/widget/presets` 200; `/api/view/pulse?rev=` and +`/api/pulse?rev=` both `changed:false`; `cleanableBytes: null`. Pages, all 200: `/` 0 s, +`/channels` 1 s, `/channels?sort=size` 0 s, `/jobs` 0 s, `/operations/diarization` **77 s**, +`/settings`, `/storage`, `/tags`, `/channels/FearAnd`, `/channels/FearAnd/videos/03Bgz7vkgbs` +0–1 s. No `ZodError` in the start log (0). `SMOKE_FAIL=0`. + +**The surfaces, from the served HTML.** `/channels` is the rack (`sticky top-0` thead, "sort by +Tier" and "sort by Transcribe" columns); `/channels?site=jeralyzer` renders 4 +`data-testid="group-name"` groups, each with `sync|download|transcribe|digest|speakers group +<name>` stations — the transcribe station renders its done state (`Transcribed`) for all four +groups tonight, so there was no group with untranscribed audio to show a count on. Release 5's +own surfaces: the site form (`/sites/jeralyzer`) shows "Per-video transcript downloads (Download +menu, Copy Markdown, Copy download command)" beside "Generate downloadable archive zips on +build", and the hub form on `/sites` shows the same checkbox. + +**The owed one-sync md5 sweep — and R's first live proof.** md5 sweep just before the sync: +identical to the after-boot sweep (the smoke wrote nothing). `pnpm ops sync +{"slug":"rekietalaw-rumble"} --wait` (job `01M3BARDZR528JS70BE17467WF`, 8 min 11 s, exit 0): +the full-sweep spawn was `yt-dlp --flat-playlist --skip-download --print url --impersonate +chrome --sleep-requests 1 https://rumble.com/c/RekietaLaw` — the platform table live — and it +COMPLETED: "670 listed, was 666", "669 known, 670 in fresh listing, 2 maybe-missing", the listing +accepted, 4 new videos on page 1. Every per-video spawn carried the same two flags, and all four +downloaded: `DLOM_ARCHIVE RumbleEmbed v7dn1wy / v7dh02c / v7ddwkg / v7dce86` — of the 500 +retained job logs this is the ONLY one with a Rumble archive line (four others carry the embedJS +403). Availability backfill wrote 4. md5 sweep after the sync: exactly one file moved, +`channels/rekietalaw-rumble/config.json` (`lastFullSweepAt` 2026-08-21 → 2026-09-25T03:54:54, +`lastSyncedAt` 2026-08-31 → 2026-09-25T03:54:53), nothing else. + +Live-found, not a regression: the maybe-missing confirmation of `v7e07us` reads `error` — the +sidecar says `HTTP Error 410: Gone`, and has since 2026-08-21. A 410 is a removed video, not an +error; `classifyAvailability` has no 410 pattern. Follow-up, small. `v7emb9a` was skipped ("no +webpage_url in metadata.info.json"), as before. + +**Form save.** One Configure-form save with no field changed on `legal-mindset` (`pnpm ops +channel-config` with `patch: {}`, the server action the form posts): `ok`, and the md5 sweep +after it is identical to the after-sync sweep — release 4's save had already applied the +one-time key reorder on that file, so this time nothing moved; `jq -S` equal. + +**Exports off on the five sites — the site form's one side effect.** The two switches were +unticked through the live Settings form (`/sites/<id>`, headless Playwright driving the real +form, `xr-flip.cjs`), never by hand. On `anilyzer` the save wrote the two `false` keys AND +materialised `order` on the one group and the one channel that had none (`legal-mindset`, added +by the LM-filter work through the CLI): the group became `order: 11` (it was already last, and +`sortGroups` keeps a missing order last, so nothing moved), the channel `order: 9` (placed +alphabetically after `leaflit-rumble`; the array was reordered to match). Channel `order` is +parsed into the schema (`siteSchema.ts:154`) and read by nothing under `common/lib` or +`export/app`, so there is no visible effect. Accepted as the form's normal write — a human save +does exactly this — and recorded per site below. + +**Step 7 — `the-quartering-rumble`'s first complete sweep in 44 days.** The stopgap +`fullSweepIntervalMinutes: 0` was removed through the Configure-form writer (`pnpm ops +channel-config` with `patch: {"fullSweepIntervalMinutes": null}` — null clears a field; the +before/after `jq -S` diff is exactly that one key gone). Then `pnpm ops sync` (23:55:34): with +`lastFullSweepAt` null the sync was a full sweep, and the editor's own spawn line was `yt-dlp +--flat-playlist --skip-download --print url --impersonate chrome --sleep-requests 1 +https://rumble.com/c/TheQuartering`. The paced walk went past page 155 (where the unpaced one +429'd three times on 2026-09-24) and on through page 310+ with no rate limit, ~9 min for the +listing. Accepted: **"8045 listed, was 7866"**, "7936 known, 8045 in fresh listing, 4 +maybe-missing" — the listing grew, it did not shrink, so the two-observation guard was not +exercised — and the stored playlist is 8,045 lines. Page 1 had 50 entries, all 50 new; the job +then downloads them one by one (each spawn impersonated and paced). + +**Orphan temp files, gone.** Operator-approved (2026-09-25 ~00:15) with two filters — older than +24 h, and the file's channel has no running job (`/api/view/activeJobs`): `find transcripts -name +'*.tmp-*'` listed 188, all of them older than 24 h — 175 zero-byte `.auto-queue/state.json.tmp-2514131-*` +from 2026-09-11 (the pre-priority editor's pid, before slice W's one write idiom) and 13 per-channel +sidecar temps (7 `snapshot.json`, 3 `transcript`, 2 `availability.json`, 1 `download-outcome.json`, +the newest from 2026-09-11 too). 188 deleted, 0 skipped (the one busy channel, `the-quartering-rumble`, +had none), 2,519,597 bytes freed, 0 remaining (`r5-orphans.log`). + +**Exports off, deployed.** `pnpm ops build-index` (job `01M3BBJXQ01ECGZNEZ61QE0WKG`, 12.4 min), +then one `build-deploy` job per site, serially. **anilyzer** (`01M3BC9JJQFADFX621KGPM971B`, 17 min: +"Done in 734.17s. 6 site(s): 6 built", "[archives] disabled for this build — skipping", 25 channels +composed, 227 files uploaded + 674 already there, deployed): verified at https://anilyzer.pages.dev — +`/downloads` 200 with the empty state ("No archives have been published for this site yet"), no +zip links, the home page has no link to `/downloads`, `/corpus.json` 200 at **spec 4** (26 channels, +29,751 videos), `transcripts/chibi-reviews/manifest.json` and `page-0001.json` 200, and a headless +open of a transcript modal finds none of "Copy yt-dlp download command", "Download this … as a +file", "Copy this … as Markdown" while "Copy share link" and the clip mark are still there. This is +also the Anilyzer production deploy owed since the curated-tags release, and the first of the five +sites at corpus spec 4. +**bonnellyzer** (`01M3BD94S59YG01JJF6V7PBN27`, 28.5 min), **hasanalyzer** (`01M3BEXC9QRCTN90ZJ5WDGB6RS`), +**rekietalyzer** (`01M3BFMRQJWSQKRM73NPNRQ1E5`, 11.9 min): the same six checks pass at each public URL — +`/downloads` 200 with the empty state and no zip links, no home-page link to `/downloads`, +`/corpus.json` 200 at spec 4 (bonnellyzer 4 channels / 8,079 videos; hasanalyzer 6 / 3,425; +rekietalyzer 2 / 2,925), a channel manifest + `page-0001.json` 200, and the transcript modal +without Copy download command / Download / Copy MD while "Copy share link" and the clip mark remain. +**jeralyzer** (`01M3BGADKE9S7V6X15WK2VF065`, 12.1 min): the same six checks pass — spec 4, **30 channels / +31,463 videos** (30,886 at the 2026-08-07 build). `corpus.json` on every site no longer carries +`bulkArchives`. The whole chain (index + five build-deploys, serial) took ~95 min; each build-deploy +rebuilds the index itself, so the standalone index job was redundant. Order fills per site, exactly +as predicted from the read-only check: anilyzer `legal-mindset` (channel + group), hasanalyzer +`FearAnd` + `hasanthehun-x`, jeralyzer `thequartering-X`, none on bonnellyzer and rekietalyzer (their +diffs are exactly the two keys). Group order unchanged everywhere. The checks are real: run against +anilyzer BEFORE its deploy they found the zip links, two home links to `/downloads`, `bulkArchives` +and all three modal controls. + +**Found on the way, not fixed.** (1) **jeralyzer `corpus.json` lists `thequartering-X` first — a +posts-only channel with `videoCount: 0` whose transcripts manifest is a 404** (its posts manifest is +200). A client that walks every transcripts manifest hits a 404; the composer should omit the +transcripts manifest URL (or emit an empty manifest) for a posts-only channel. Small follow-up in +`common/bin/compose-site.ts`. (2) **The hub's "Save config" form dropped `"nav": []` from +`homepage.json`** alongside adding `transcriptDownloads: false`. `nav` is not a `HomepageConfig` +field, `parseHomepageConfig` never reads it and nothing under `homepage/app`, `export/app/lib/site.ts` +or `compose-hub.ts` uses it — inert, but the writer rebuilds the file from named fields, which is +the documented behaviour (SITE.md's sibling has no doc: `homepage.json` has none). (3) **The hub +build has no deploy path.** `INSTANCE_MODE=hub` is `pnpm --filter export run build:hub` and nothing +deploys its `out` (no `pnpm ops` action, no editor entry — `plans/one-core.md:84` already says +"`build:hub` has no operator entry point"); the `homepage` package deploys to the `archilyzer` +project but mounts no transcript modal. So the hub's key is SET but no hub was rebuilt or deployed +tonight, and the federated hub could not be checked live because **https://archilyzer.pages.dev +answers HTTP 522 on every path** (`/`, `/ask`, `/corpus.json`, `/hub-sites.json`; checked twice, the +second time at 05:36 UTC) — not from this rollout, nothing deployed there. Operator: look at the +Cloudflare project. The Phase 4 CLI slice is where `build:hub` gets its entry point. + +**Release 6 rode along on `main`, not on :3001.** After this rollout the operator approved an +overnight plan (memory `overnight-2026-09-25`): the follow-ups slice merged as `4d97049f` and Phase +4 slice 1 as `cda35622` — see [`release-6.md`](release-6.md). Neither is rolled out; the next +prelude does that, after `pnpm install --frozen-lockfile` in the primary (the umtool→common link +and the aws-sdk move). The YouTube pacing investigation is filed as +[`youtube-lane-pacing.md`](youtube-lane-pacing.md): the evening's three cooldowns were one Short +retried 12× from the head of the queue, not a burst. + +**Step-7 tail, still running at this commit (01:38):** the quartering sync job keeps downloading the new videos — 252 archive lines (2 per video), 0 errors so far; its end (stamps, maybe-missing confirmation) is recorded in the morning runbook and STATE, not here. diff --git a/plans/youtube-lane-pacing.md b/plans/youtube-lane-pacing.md @@ -0,0 +1,149 @@ +# Plan: one YouTube video must not keep the whole platform in a 429 cooldown + +**Found 2026-09-25**, from STATE.md: "YouTube held a 429 cooldown across three attempts (18:43, +19:13, 19:41)", and the owed md5-sweep sync was refused all three times. Verified read-only on +`main` `93dcb532`, the live `settings.json` and the retained `transcripts/.jobs/*` (meta +`startedAt` in UTC). **The STATE times are EDT.** They are the unit attempts at 22:42, 23:10 and +23:39 UTC. + +## Findings (measured, not guessed) + +- **It was one video, not a burst.** All three cooldowns were the auto-download lane re-trying + `quarteringvlogs/ncdPaDSqt-c`, a 178 s Short. It got 12 consecutive `rate_limit` failures from + 20:28 to 23:39 UTC, and the runner logs `01M3A5034Y…` and `01M3ASBJHW…` show attempts 1→12 + (attempt 9 was the manual `download-missing` `01M3AQJ5…`, which shares the state). It succeeded + at 00:09 UTC. Attempt spacing matches `nextBackoff` exactly: +93 s, +145, +252, +477, +920, + then about 28-31 min at the cap (`platformBackoff.ts:22,24`: 60 s base, doubling, 30 min cap, + ±10 %). The cap was reached at about 21:30 UTC. +- **Why the same video every time.** On `rate_limit`/`network` the runner deliberately does NOT + `markCompleted` the video (`autoRunner.ts:1936-1962`), and `order: "listed"` puts it back at + the head of the pick. So each lapse of the cooldown (3 s idle poll, `autoRunner.ts:158`) re-picks + the same video, gets another 429, and re-arms a platform-wide cooldown with `fails+1`. The + `fails` count survives a restart: the new runner continued at attempt 10. A manual Sync reads the + same cooldown and refuses (`editor/app/channels/[slug]/pipelineActions.ts:133-140`). The one + video therefore blocked every YouTube channel for about 3.7 h, and the lane ran no other YouTube + unit in that window. +- **What 429s is YouTube's subtitle (timedtext) endpoint, per video.** All 20 retained YouTube + 429s are `Unable to download video subtitles for 'en': HTTP Error 429`. The webpage, player API + and m3u8 requests in the same spawn succeed. There are **0** `Sign in to confirm` lines in any + retained log, and the only 2 `webpage … 429`s are Rumble syncs (`01M3AVGZ…`, `01M3AH7G…`). + **It is not IP-wide.** `the-quartering/UCirTfhUP3s` got the 429 at 16:41:37 UTC, and + `HasanAbiVODs3/1FibhXko_Kw` wrote both `en` and `en-orig` subtitles 27 s later (16:42:04), as did + `nuxanor` at 16:44. `CAi9jNrHetw` 429'd 3× then succeeded. `q_UPNCELU8I` 429'd 2× then + succeeded. Every affected video is on a Quartering channel. +- **Volume before each 429 was tiny.** In the hour before 22:42, 23:10 and 23:39 UTC, the only + YouTube work was one or two attempts on the same video (2 yt-dlp spawns each). The sync at + 20:20 UTC 429'd after about 3.5 h with no YouTube spawns from this editor. Pacing between videos + could not have prevented any of the 16 unit 429s, because they were already ≥ 93 s apart. +- **Per unit, a `handling: youtube` download is 2 spawns** (`downloadOneManaged.ts`): + - the metadata prefetch at `:646-661`, with no `-t sleep`: webpage, player, m3u8; + - the primary with `-t sleep` (`:168`): 2 timedtext requests (`en`, `en-orig` from + `en.*,live_chat`) 5 s apart. + On a subtitle 429, yt-dlp re-extracts from the URL, so one attempt is about 8 requests. +- **Steady state when the lane has backlog: back-to-back, no gap.** + - The unit calls `downloadOneManaged` directly (`autoRunner.ts:2194`). `sleepBetweenDownloadsSeconds` + (live: 30) is read only by `runManagedDownloads` (`runYtdlp.ts:932-934,1061-1067`), which is + the sync path. FACTS.md:2733 records the same blindness for the backfill sweep. + - `runPool` wakes on every settle (`concurrentRunner.ts:118-126`), and `PER_PLATFORM_CAP = 1` + (`autoRunner.ts:1231`). + - Measured: units started 00:09:32, 00:09:55, 00:18:49 and 00:19:28 UTC. A success is followed + by the next start about 23-39 s later, which is ≈ 2 units/min ≈ 4 spawns ≈ 10 requests/min + at most. + - After a cooldown the lane resumes at full speed. A single success calls `clearBackoff` + (`:1957`), so the next 429 starts again at 60 s: 00:09:32 ok, then 00:09:55 429 → 55 s. +- **Lanes sharing one YouTube key.** + - Auto-download units, syncs and metadata scans all queue on `platform:youtube` (concurrency 1, + `queueKeys.ts:69-73`), so they never overlap. The 41 syncs the scheduler queued at 16:05 UTC + ran serially (and were dropped at the 16:46 restart, which left 41 stale `queued` metas). + - Backfill re-acquire calls `downloadOneManaged` in-process (`backfillReacquire.ts:226`), with + no platform queue and no cooldown check. That lane is `enabled: false` today, so it is latent. + - No availability check ran in the window. +- **No YouTube pacing args.** `PLATFORM_ARGS` has only Rumble (`channelArgs.ts:30`). YouTube + gets `-t sleep` only on the youtube-handling primary. The prefetch and transcribe-handling + downloads have no `--sleep-requests`. +- **A second, real burst: the 142 members-only legal-mindset re-scan.** + - It rescans on every runner start: 16:03, 16:46, 00:19 and 03:39 UTC, each "0 scanned, 142 + error(s)", with `--cookies-from-browser firefox` at 1 req/s. + - The 24 h suppression (`metadataScanStore.ts:334-349`) expired for good. The upsert keeps an + identical re-error untouched (`:249-252`: it rewrites only when class or message changed), so + `errors[id].at` is still 2026-09-20 and `metadataScanWanted` has said "yes" ever since. + - It was not the trigger tonight, because timing does not line up with the 429s. It is a real + cookie-authenticated burst. + +## The one change: defer a rate-limited video, not only its platform + +A `rate_limit` failure keeps the platform cooldown exactly as today, but **also takes that video +off the lane for `VIDEO_RATE_LIMIT_DEFER_MS` (6 h)**. After the platform cooldown lapses, the lane +offers the *next* video. On 09-24 ncdPaDSqt-c recovered about 3.7 h after its first 429, and the +other two recovered within 9 min. Now `fails` escalates only while *different* videos keep +getting 429s, which is a true platform signal. A single success clears it, as today. + +Why this and not the alternatives: +- **Honouring the per-video gap (`sleepBetweenDownloadsSeconds`)** is correct and wanted, but on + this evidence it changes nothing: the failing attempts were already 93 s-30 min apart. +- **A YouTube `--sleep-requests` entry** paces requests *within* a spawn. The 429 is on the first + timedtext GET after 5 s of `--sleep-subtitles`. +- **A cooldown that also slows the resume** would stretch the same loop further. It would still + re-arm on the same video every 30 min. +- **A cross-lane cap** is not needed: every YouTube path except the disabled backfill already + serializes on `platform:youtube`. + +Only the head-of-line retry explains three cooldowns in an evening at ≈ 2 spawns per 30 min. + +## Steps + +1. `common/jobs/platformBackoff.ts` (pure): export `VIDEO_RATE_LIMIT_DEFER_MS = 6 * 60 * 60_000` + and `isVideoDeferred(deferred: Map<string, number>, id, now)`, plus a prune of lapsed entries. + Do not change `nextBackoff`. +2. `common/controller/autoRunner.ts`: + - Beside `platformInFlight` (about `:1230`), add `videoDeferredUntil = new Map<string, number>()`. + - In the download branch of `next()` (`:1617-1650`), also filter ids with + `isVideoDeferred(...)`. When everything left is deferred, set `anyCooling = true` so the idle + reason reads `cooldown`, not `capped`. + - In the `backoffHit` branch (`:1945-1955`), add + `videoDeferredUntil.set(pick.videoId, Date.now() + VIDEO_RATE_LIMIT_DEFER_MS)` only when + `failureClass === "rate_limit"`. Network failures keep today's retry-same-video behaviour. + - Extend the log line: `… ${pick.videoId} deferred 6h; next video after cooldown.` + - Keep it in-memory. A runner restart re-offers the video once, which costs at most one extra + attempt per restart and needs no state-file migration. +3. `plans/FACTS.md`: record that the lane defers a rate-limited video and that the platform + cooldown now escalates across distinct videos. Note that YouTube 429s are timedtext and + per-video (evidence above). + +## Tests + +- Unit tests (`common/jobs/platformBackoff.test.ts`): `isVideoDeferred` before and after the + window, and prune. +- `common/controller/autoRunner.test.ts`: + - A fake unit returns `rate_limit` for video A and success for B. After A's 429 and the + platform cooldown, the next pick is B, not A. B's success clears the platform cooldown. A is + not offered again until the deferral lapses (injected clock). + - `network` still re-picks the same video. + - With all pending videos deferred, the idle reason is `cooldown`. +- No e2e change is needed. `auto-queue.spec.ts` should stay green, and should be run in the next + release's gate, not by this plan. + +## Rollout + +Ride the next release. Watch the auto-download runner log for one evening: +- expect at most one 429 line per video per 6 h; +- expect `attempt N` to climb only across different ids; +- expect the owed md5-sweep sync not to be refused by a single-video cooldown. + +If the lane then shows 429s on *several distinct* videos within minutes, the limit is platform-wide +after all and the cooldown is doing its job. The next lever is then the pacing gap below. + +## Found, deliberately not in this change (ranked) + +1. **Scan error `at` never refreshes** (`metadataScanStore.ts:249-252`): 142 cookie-authenticated + YouTube requests on every runner start. The fix is one line: rewrite an identical error when + `prev.at` is older than `METADATA_SCAN_ERROR_COOLDOWN_MS`. It is its own slice with its own test. +2. **The lane ignores `sleepBetweenDownloadsSeconds`**. The fix is a per-platform `nextStartAt` set + on unit settle. This is the "slow" half of the operator's rule, for when backlog is large. +3. **Backfill re-acquire bypasses the platform queue and cooldown** (`backfillReacquire.ts:226`). + It is latent while the lane is disabled. +4. **`en.*` fetches both `en` and `en-orig`**, which is 2 timedtext GETs per video. That is a + `subLangs` question for the operator, not a pacing one. + +Out of scope: changing `nextBackoff` constants; any yt-dlp client or PO-token work on the +timedtext 429 itself; adding YouTube to `PLATFORM_ARGS`; the stale `queued` metas left by restarts.