commit 2f7002b97269125ea697294f19d3d98df68040d5
parent 1e486c64d1ffcd52a9c8a4f2e9bf7ef731640a64
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Fri, 25 Sep 2026 01:39:18 -0400
plans: release 5 rollout record — 05877ffb live on :3001 (releases 4+5), exports off on five sites, Rumble proven live; STATE + FACTS; YouTube lane pacing plan filed
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Diffstat:
4 files changed, 387 insertions(+), 9 deletions(-)
diff --git a/plans/FACTS.md b/plans/FACTS.md
@@ -6194,3 +6194,52 @@ Line numbers are `plans/FACTS.md` lines at `e172749b`, before this record's in-p
is historical. `normalizeTranscript.ts:141` is now `writeJsonAtomic(cuesPath, out,
{indent: 0, newline: false})`. The idiom for any file is `writeFileAtomic` /
`writeJsonAtomic`.
+
+## Release 5 — slices R, X (verified 2026-09-24/25, `main` @ `93dcb532`)
+
+- **One yt-dlp arg builder with a platform table.** `common/ytdlp/channelArgs.ts`:
+ `PLATFORM_ARGS` :30 (`rumble: --impersonate chrome --sleep-requests 1`), `platformArgs` :34,
+ `channelPlatform` :39 (`config.platform ?? detectPlatform(config.url)`), `channelExtraArgs` :45
+ = cookies → platform args → `ytdlpExtraArgs` (a channel override is last and wins). `configArgs`
+ (runYtdlp), the metadata scan and `checkAvailability` all delegate to it; `probeChannelMeta`
+ and the no-config availability path call `platformArgs(detectPlatform(url))`. The table is code;
+ there is no setting.
+- **A 429 mid-sweep is "incomplete".** `EnumerationIncompleteError` `runYtdlp.ts:329`
+ (`platform, pagesReached, count`), `lastListingPage` :353 (last `Downloading page N` — the real
+ wording is `[RumbleChannel] TheQuartering: Downloading page 155`). `syncFullSweep` catches it
+ before `acceptEnumeration`: best-effort `opts.onPlatformBackoff("rate_limit")` :1644 (=
+ `recordDownloadBackoff`, the download runner's helper and schedule), the `Full sweep incomplete`
+ line :1653, then `return syncPaged(opts)` :1658. No playlist write, no `lastFullSweepAt`. Any
+ other non-zero exit throws as before; exit 101 untouched. `fullSweepDue` :1408 is exported, async,
+ and false while `platformCooldownRemainingMs(detectPlatform(url) ?? "unknown") > 0` — the Sync
+ gate's key. Syncs are REFUSED during the cooldown (`pipelineActions.ts` Sync gate), and the sweep
+ is due again the moment it ends — no last-attempt stamp exists (follow-up if live sweeps keep
+ coming back incomplete).
+- **A bare `HTTP Error 403` is `network`** (`common/lib/availability.ts:255`), so a Cloudflare
+ block backs the platform off; there is no `blocked` class because `DownloadFailureClass` is
+ persisted in every `download-outcome.json`.
+- **Fake yt-dlp** records every flat-playlist argv as `flat-playlist:full|paged argv=…` in the
+ per-channel `fake-ytdlp.invocations` (`fake-ytdlp.mjs:543`); the `sweep429` URL sentinel :549
+ prints 3 pages × 5 urls then the real 429 line and exits 1 on a FULL enumeration only.
+- **`transcriptDownloads`** (absent = on) lives in three files and one context: `site.json`
+ (`siteSchema.ts:88` type, :122 docs, :245 zod, :328 write-if-false), `homepage.json` for the hub
+ (`homepage.ts:39,83,133` — the hub has NO site.json; `hubSite()` `export/app/lib/site.ts:40` passes
+ it through), and `PlayerProvider` `features` (`PlayerProvider.tsx:118`, default all on);
+ `TranscriptModal.tsx:79` `showDownloads` hides exactly Copy download command, the Download
+ dropdown and Copy MD. FOUR mounts, all in `export/` (`(workspace)/layout.tsx` → `SiteWorkspace`,
+ `duplicates/page.tsx`, `HubHome`, `AskHub` via their server pages); the editor mounts neither.
+ `SITE.md` is generated; `homepage.json` has no generated doc and no zod schema.
+- **The export e2e off-test flips the committed fixture** `export/e2e/fixtures/sites/testsite/site.json`
+ for one test; `export/playwright.config.ts:38` strips a leftover key at config load (before any
+ spec), so a killed run leaves at most a whitespace diff. `workers: 1`, so no other test can see
+ the flip.
+- **Rollout facts.** `ops channel-config` with `patch: {"<field>": null}` CLEARS a Configure-form
+ field (the route lays the patch over `channelConfigToFormData(existing)` and
+ `updateChannelAction` unsets every form field first) — that is how `fullSweepIntervalMinutes: 0`
+ was removed from `the-quartering-rumble`. A site's `archives` / `transcriptDownloads` are written
+ only by the Settings form (`saveSiteAction` → `writeSite`); there is no `ops site-config`.
+- **A Settings-form save materialises missing `order` fields** on a site's groups and channels
+ (entries added through the CLI carry none). Group order is consumed by `sortGroups`
+ (`channelGroups.ts:118`, missing = last); channel `order` is parsed (`siteSchema.ts:154`) and
+ read by nothing. So a "no-op" site save can move bytes; `jq -S` plus the `order` fills is the
+ expected diff, not a bug.
diff --git a/plans/STATE.md b/plans/STATE.md
@@ -3,15 +3,37 @@
The working memory for the local-AI derived-corpus work. Rewritten at the end of every
session, before context is cleared. See [`README.md`](README.md) for the protocol.
-**Live on :3001 (2026-09-24, 18:42):** still `9ab10d77` (release 3), `BUILD_ID`
-`XsKaA_drqdAbTguxUGVsn`. **Release 4 is on `main` but NOT rolled out**; the next release's
-prelude rolls it out. The one-sync md5 sweep from the release-3 rollout is still owed.
-YouTube held a 429 cooldown across three attempts (18:43, 19:13, 19:41). See the
-[rollout record](one-core-phase-3.md#rollout-2026-09-24--9ab10d77-live-on-3001).
-
-**Release 5 in flight (2026-09-24, late):** slices R (Rumble) and X (exports-off) cut off
-`8b7f8924` as `one-core/r5-rumble` and `one-core/r5-exports`; plan, record and rollout in
-[`release-5.md`](release-5.md); rules in [`tools/implementer-rules.md`](tools/implementer-rules.md).
+**Live on :3001 (2026-09-24, 23:39):** `93dcb532` (releases 4 + 5 together), `BUILD_ID`
+`FKE60BTWpUiUa94WCSxTO`, next-server 888107. **Release 6 (follow-ups `4d97049f` + Phase 4 slice 1
+`cda35622`) is on `main` but NOT rolled out** — the next release's prelude rolls it out, after
+`pnpm install --frozen-lockfile` in the primary (umtool→common link, aws-sdk now under common). The
+one-sync md5 sweep owed since release 3 is DONE (`rekietalaw-rumble`, 4 Rumble videos archived — the
+first since August). See the [release-5 rollout record](release-5.md#rollout-2026-09-24-night--93dcb532-live-on-3001-releases-4--5-together).
+
+**Last updated:** 2026-09-25 (early morning). **Release 5 is live; release 6 is merged.**
+- Release 5 (`f4da04a9` → `93dcb532`): slice R (`one-core/r5-rumble` → `3adaea9b`): one yt-dlp arg
+ builder with a platform table (Rumble: `--impersonate chrome --sleep-requests 1` on every spawn), a
+ 429 mid-sweep is "incomplete" (paged walk + platform cooldown, never a listing, never a stamp), a
+ bare 403 backs the platform off. Slice X (`one-core/r5-exports` → `93dcb532`): `transcriptDownloads`
+ (absent = on) on `site.json` AND `homepage.json`, gating the three per-video export controls on
+ the export site; the site and hub forms carry the checkbox. Final suites on `93dcb532`: editor
+ 628/628 (no flake), export 192/192, hub 8/8. Record: [`release-5.md`](release-5.md).
+- Rollout: ONE restart; smoke green; boot rewrote nothing; proof sync of `rekietalaw-rumble` and
+ the first complete `the-quartering-rumble` sweep in 44 days (8,045 listed, was 7,866, paced past
+ page 300 with no 429); exports off on the five published sites (form → build-index → build-deploy
+ each, verified at the public URLs; Anilyzer's owed production deploy rode along, now at spec 4).
+- Release 6 (overnight, operator-approved): follow-ups (410 → deleted, umtool gets the platform args
+ through ONE table in `common/ytdlp/platformArgs.mjs`, atomic remote-transcript write, stale comment,
+ measure-nav routes) and Phase 4 slice 1 (`buildDeployCore.ts` → `common/publish/build.ts`, aws-sdk
+ deps to common; the move only — entry points, docker scripts and build:hub wait for the CLI slice).
+ Record: [`release-6.md`](release-6.md). 188 orphan `*.tmp-*` files deleted (filtered).
+- Found and filed, not fixed: [`youtube-lane-pacing.md`](youtube-lane-pacing.md) — the evening's
+ three YouTube cooldowns were ONE Short retried 12× from the head of the queue; the fix is a 6-hour
+ per-video quarantine (small slice, recommended next), plus `metadataScanStore.ts:249-252` (scan
+ error timestamp never refreshes → 142 members-only videos re-scanned every runner start).
+
+**Next:** Phase 4 slice 2 (`common/bin/archilyzer.ts` + the rest of spec item 1), the pacing slice,
+then the release-6 rollout (one restart) with the pacing fix riding along if it is ready.
**Last updated:** 2026-09-24 (evening). **Release 4 is on `main` @ `e172749b`, and one-core
Phase 3 is COMPLETE. Phase 4 is next.** The prelude was plans only: `4130aca1` (the rollout
diff --git a/plans/release-5.md b/plans/release-5.md
@@ -319,3 +319,161 @@ and the written file keeps both keys.
and build + deploy each, then untick transcript downloads on the hub form and rebuild + deploy the
hub. Verify per site as item 5 now says: no Downloads link in the header or footer, and
`/downloads` shows its empty state (it still answers 200).
+
+## Rollout 2026-09-24 (night) — `93dcb532` live on :3001 (releases 4 + 5 together)
+
+ONE restart, at the END of release 5, as the plan said: R is what un-breaks Rumble syncing.
+Editor only — `git diff --stat 9ab10d77..93dcb532 -- umtool` prints nothing, so umtool was
+not rebuilt or restarted (still serving on :3050). The live editor had served `9ab10d77`
+(release 3, `BUILD_ID` `XsKaA_drqdAbTguxUGVsn`) since the release-4 prelude.
+
+**Final suites on `93dcb532` first** (worktree `one-core-r5-exports` detached at the merge sha,
+whose tree equals X's gated tip `d948aa0e`; ports 3211/3210/3220; logs
+`final-e2e-r5-{editor,export,hub}.log`): editor **628 passed, 0 failed, 0 flaky, 39.4 min**
+(release 4's `lane-runner.spec.ts:361` load flake did not recur); export **192/192, 6.2 min**
+(188 + 4 new); hub **8/8, 16 s**.
+
+**Offline round-trips.** `phase3-settings-numbers.ts live=settings.json`: the written block
+equals `jq -S . settings.json` (diff empty, 1,353 scalar paths each side — the same count as
+release 4). `phase3-files-numbers.ts` over the LIVE `transcripts/`: 6 `site.json` + 71
+`config.json`, every `unknown keys: []` (77 lines), no `WRITE THREW`; the output is 3,839
+lines against release 4's 3,832 — the seven new lines are the `"transcriptDownloads": true`
+each site's parse now carries (X's key, default on) plus the hub's.
+
+**Build and one restart.** md5 baseline over `settings.json` + 6 `site.json` + 71
+`config.json` (78 files). `pnpm --filter editor build` detached into the live `.next` while the
+old server served: exit 0, compiled in 14.3 s, `ƒ /api/view/[name]` in the route table;
+`BUILD_ID` `XsKaA_drqdAbTguxUGVsn` → **`FKE60BTWpUiUa94WCSxTO`**. Seven auto-lane jobs were in
+flight (1 backfill, 1 digest, 2 download, 2 transcribe, 1 transcribe), as at release 4's
+restart. Then TERM pnpm 525402 / next-server 525417 (cwd `editor/`), :3001 free, `setsid nohup
+pnpm run start -H 0.0.0.0` → next-server **888107**, `/` 200 after 2 s, `Ready in 163ms`,
+`/tags` 200. md5 sweep right after boot: identical to the baseline — boot rewrote nothing.
+
+**Smoke, started 4 min after boot** (`smoke10.sh`, the release-4 script). All eight old/new
+pairs equal on the FIRST pass with the live fields stripped: `pulse` 109 B, `workers` 1,116 B,
+`activeJobs` 2,652 B, `autoQueueStatus` 507,725 B (79 s), `schedulerStatus` 37,371 B,
+`widgetSync` 1,522 B, `widgetActionable` 1,148 B, `cleanable` 1,677 B. `/api/view/bogus` and
+`/api/view/invalidate-cache` 404; `/api/widget/presets` 200; `/api/view/pulse?rev=` and
+`/api/pulse?rev=` both `changed:false`; `cleanableBytes: null`. Pages, all 200: `/` 0 s,
+`/channels` 1 s, `/channels?sort=size` 0 s, `/jobs` 0 s, `/operations/diarization` **77 s**,
+`/settings`, `/storage`, `/tags`, `/channels/FearAnd`, `/channels/FearAnd/videos/03Bgz7vkgbs`
+0–1 s. No `ZodError` in the start log (0). `SMOKE_FAIL=0`.
+
+**The surfaces, from the served HTML.** `/channels` is the rack (`sticky top-0` thead, "sort by
+Tier" and "sort by Transcribe" columns); `/channels?site=jeralyzer` renders 4
+`data-testid="group-name"` groups, each with `sync|download|transcribe|digest|speakers group
+<name>` stations — the transcribe station renders its done state (`Transcribed`) for all four
+groups tonight, so there was no group with untranscribed audio to show a count on. Release 5's
+own surfaces: the site form (`/sites/jeralyzer`) shows "Per-video transcript downloads (Download
+menu, Copy Markdown, Copy download command)" beside "Generate downloadable archive zips on
+build", and the hub form on `/sites` shows the same checkbox.
+
+**The owed one-sync md5 sweep — and R's first live proof.** md5 sweep just before the sync:
+identical to the after-boot sweep (the smoke wrote nothing). `pnpm ops sync
+{"slug":"rekietalaw-rumble"} --wait` (job `01M3BARDZR528JS70BE17467WF`, 8 min 11 s, exit 0):
+the full-sweep spawn was `yt-dlp --flat-playlist --skip-download --print url --impersonate
+chrome --sleep-requests 1 https://rumble.com/c/RekietaLaw` — the platform table live — and it
+COMPLETED: "670 listed, was 666", "669 known, 670 in fresh listing, 2 maybe-missing", the listing
+accepted, 4 new videos on page 1. Every per-video spawn carried the same two flags, and all four
+downloaded: `DLOM_ARCHIVE RumbleEmbed v7dn1wy / v7dh02c / v7ddwkg / v7dce86` — of the 500
+retained job logs this is the ONLY one with a Rumble archive line (four others carry the embedJS
+403). Availability backfill wrote 4. md5 sweep after the sync: exactly one file moved,
+`channels/rekietalaw-rumble/config.json` (`lastFullSweepAt` 2026-08-21 → 2026-09-25T03:54:54,
+`lastSyncedAt` 2026-08-31 → 2026-09-25T03:54:53), nothing else.
+
+Live-found, not a regression: the maybe-missing confirmation of `v7e07us` reads `error` — the
+sidecar says `HTTP Error 410: Gone`, and has since 2026-08-21. A 410 is a removed video, not an
+error; `classifyAvailability` has no 410 pattern. Follow-up, small. `v7emb9a` was skipped ("no
+webpage_url in metadata.info.json"), as before.
+
+**Form save.** One Configure-form save with no field changed on `legal-mindset` (`pnpm ops
+channel-config` with `patch: {}`, the server action the form posts): `ok`, and the md5 sweep
+after it is identical to the after-sync sweep — release 4's save had already applied the
+one-time key reorder on that file, so this time nothing moved; `jq -S` equal.
+
+**Exports off on the five sites — the site form's one side effect.** The two switches were
+unticked through the live Settings form (`/sites/<id>`, headless Playwright driving the real
+form, `xr-flip.cjs`), never by hand. On `anilyzer` the save wrote the two `false` keys AND
+materialised `order` on the one group and the one channel that had none (`legal-mindset`, added
+by the LM-filter work through the CLI): the group became `order: 11` (it was already last, and
+`sortGroups` keeps a missing order last, so nothing moved), the channel `order: 9` (placed
+alphabetically after `leaflit-rumble`; the array was reordered to match). Channel `order` is
+parsed into the schema (`siteSchema.ts:154`) and read by nothing under `common/lib` or
+`export/app`, so there is no visible effect. Accepted as the form's normal write — a human save
+does exactly this — and recorded per site below.
+
+**Step 7 — `the-quartering-rumble`'s first complete sweep in 44 days.** The stopgap
+`fullSweepIntervalMinutes: 0` was removed through the Configure-form writer (`pnpm ops
+channel-config` with `patch: {"fullSweepIntervalMinutes": null}` — null clears a field; the
+before/after `jq -S` diff is exactly that one key gone). Then `pnpm ops sync` (23:55:34): with
+`lastFullSweepAt` null the sync was a full sweep, and the editor's own spawn line was `yt-dlp
+--flat-playlist --skip-download --print url --impersonate chrome --sleep-requests 1
+https://rumble.com/c/TheQuartering`. The paced walk went past page 155 (where the unpaced one
+429'd three times on 2026-09-24) and on through page 310+ with no rate limit, ~9 min for the
+listing. Accepted: **"8045 listed, was 7866"**, "7936 known, 8045 in fresh listing, 4
+maybe-missing" — the listing grew, it did not shrink, so the two-observation guard was not
+exercised — and the stored playlist is 8,045 lines. Page 1 had 50 entries, all 50 new; the job
+then downloads them one by one (each spawn impersonated and paced).
+
+**Orphan temp files, gone.** Operator-approved (2026-09-25 ~00:15) with two filters — older than
+24 h, and the file's channel has no running job (`/api/view/activeJobs`): `find transcripts -name
+'*.tmp-*'` listed 188, all of them older than 24 h — 175 zero-byte `.auto-queue/state.json.tmp-2514131-*`
+from 2026-09-11 (the pre-priority editor's pid, before slice W's one write idiom) and 13 per-channel
+sidecar temps (7 `snapshot.json`, 3 `transcript`, 2 `availability.json`, 1 `download-outcome.json`,
+the newest from 2026-09-11 too). 188 deleted, 0 skipped (the one busy channel, `the-quartering-rumble`,
+had none), 2,519,597 bytes freed, 0 remaining (`r5-orphans.log`).
+
+**Exports off, deployed.** `pnpm ops build-index` (job `01M3BBJXQ01ECGZNEZ61QE0WKG`, 12.4 min),
+then one `build-deploy` job per site, serially. **anilyzer** (`01M3BC9JJQFADFX621KGPM971B`, 17 min:
+"Done in 734.17s. 6 site(s): 6 built", "[archives] disabled for this build — skipping", 25 channels
+composed, 227 files uploaded + 674 already there, deployed): verified at https://anilyzer.pages.dev —
+`/downloads` 200 with the empty state ("No archives have been published for this site yet"), no
+zip links, the home page has no link to `/downloads`, `/corpus.json` 200 at **spec 4** (26 channels,
+29,751 videos), `transcripts/chibi-reviews/manifest.json` and `page-0001.json` 200, and a headless
+open of a transcript modal finds none of "Copy yt-dlp download command", "Download this … as a
+file", "Copy this … as Markdown" while "Copy share link" and the clip mark are still there. This is
+also the Anilyzer production deploy owed since the curated-tags release, and the first of the five
+sites at corpus spec 4.
+**bonnellyzer** (`01M3BD94S59YG01JJF6V7PBN27`, 28.5 min), **hasanalyzer** (`01M3BEXC9QRCTN90ZJ5WDGB6RS`),
+**rekietalyzer** (`01M3BFMRQJWSQKRM73NPNRQ1E5`, 11.9 min): the same six checks pass at each public URL —
+`/downloads` 200 with the empty state and no zip links, no home-page link to `/downloads`,
+`/corpus.json` 200 at spec 4 (bonnellyzer 4 channels / 8,079 videos; hasanalyzer 6 / 3,425;
+rekietalyzer 2 / 2,925), a channel manifest + `page-0001.json` 200, and the transcript modal
+without Copy download command / Download / Copy MD while "Copy share link" and the clip mark remain.
+**jeralyzer** (`01M3BGADKE9S7V6X15WK2VF065`, 12.1 min): the same six checks pass — spec 4, **30 channels /
+31,463 videos** (30,886 at the 2026-08-07 build). `corpus.json` on every site no longer carries
+`bulkArchives`. The whole chain (index + five build-deploys, serial) took ~95 min; each build-deploy
+rebuilds the index itself, so the standalone index job was redundant. Order fills per site, exactly
+as predicted from the read-only check: anilyzer `legal-mindset` (channel + group), hasanalyzer
+`FearAnd` + `hasanthehun-x`, jeralyzer `thequartering-X`, none on bonnellyzer and rekietalyzer (their
+diffs are exactly the two keys). Group order unchanged everywhere. The checks are real: run against
+anilyzer BEFORE its deploy they found the zip links, two home links to `/downloads`, `bulkArchives`
+and all three modal controls.
+
+**Found on the way, not fixed.** (1) **jeralyzer `corpus.json` lists `thequartering-X` first — a
+posts-only channel with `videoCount: 0` whose transcripts manifest is a 404** (its posts manifest is
+200). A client that walks every transcripts manifest hits a 404; the composer should omit the
+transcripts manifest URL (or emit an empty manifest) for a posts-only channel. Small follow-up in
+`common/bin/compose-site.ts`. (2) **The hub's "Save config" form dropped `"nav": []` from
+`homepage.json`** alongside adding `transcriptDownloads: false`. `nav` is not a `HomepageConfig`
+field, `parseHomepageConfig` never reads it and nothing under `homepage/app`, `export/app/lib/site.ts`
+or `compose-hub.ts` uses it — inert, but the writer rebuilds the file from named fields, which is
+the documented behaviour (SITE.md's sibling has no doc: `homepage.json` has none). (3) **The hub
+build has no deploy path.** `INSTANCE_MODE=hub` is `pnpm --filter export run build:hub` and nothing
+deploys its `out` (no `pnpm ops` action, no editor entry — `plans/one-core.md:84` already says
+"`build:hub` has no operator entry point"); the `homepage` package deploys to the `archilyzer`
+project but mounts no transcript modal. So the hub's key is SET but no hub was rebuilt or deployed
+tonight, and the federated hub could not be checked live because **https://archilyzer.pages.dev
+answers HTTP 522 on every path** (`/`, `/ask`, `/corpus.json`, `/hub-sites.json`; checked twice, the
+second time at 05:36 UTC) — not from this rollout, nothing deployed there. Operator: look at the
+Cloudflare project. The Phase 4 CLI slice is where `build:hub` gets its entry point.
+
+**Release 6 rode along on `main`, not on :3001.** After this rollout the operator approved an
+overnight plan (memory `overnight-2026-09-25`): the follow-ups slice merged as `4d97049f` and Phase
+4 slice 1 as `cda35622` — see [`release-6.md`](release-6.md). Neither is rolled out; the next
+prelude does that, after `pnpm install --frozen-lockfile` in the primary (the umtool→common link
+and the aws-sdk move). The YouTube pacing investigation is filed as
+[`youtube-lane-pacing.md`](youtube-lane-pacing.md): the evening's three cooldowns were one Short
+retried 12× from the head of the queue, not a burst.
+
+**Step-7 tail, still running at this commit (01:38):** the quartering sync job keeps downloading the new videos — 252 archive lines (2 per video), 0 errors so far; its end (stamps, maybe-missing confirmation) is recorded in the morning runbook and STATE, not here.
diff --git a/plans/youtube-lane-pacing.md b/plans/youtube-lane-pacing.md
@@ -0,0 +1,149 @@
+# Plan: one YouTube video must not keep the whole platform in a 429 cooldown
+
+**Found 2026-09-25**, from STATE.md: "YouTube held a 429 cooldown across three attempts (18:43,
+19:13, 19:41)", and the owed md5-sweep sync was refused all three times. Verified read-only on
+`main` `93dcb532`, the live `settings.json` and the retained `transcripts/.jobs/*` (meta
+`startedAt` in UTC). **The STATE times are EDT.** They are the unit attempts at 22:42, 23:10 and
+23:39 UTC.
+
+## Findings (measured, not guessed)
+
+- **It was one video, not a burst.** All three cooldowns were the auto-download lane re-trying
+ `quarteringvlogs/ncdPaDSqt-c`, a 178 s Short. It got 12 consecutive `rate_limit` failures from
+ 20:28 to 23:39 UTC, and the runner logs `01M3A5034Y…` and `01M3ASBJHW…` show attempts 1→12
+ (attempt 9 was the manual `download-missing` `01M3AQJ5…`, which shares the state). It succeeded
+ at 00:09 UTC. Attempt spacing matches `nextBackoff` exactly: +93 s, +145, +252, +477, +920,
+ then about 28-31 min at the cap (`platformBackoff.ts:22,24`: 60 s base, doubling, 30 min cap,
+ ±10 %). The cap was reached at about 21:30 UTC.
+- **Why the same video every time.** On `rate_limit`/`network` the runner deliberately does NOT
+ `markCompleted` the video (`autoRunner.ts:1936-1962`), and `order: "listed"` puts it back at
+ the head of the pick. So each lapse of the cooldown (3 s idle poll, `autoRunner.ts:158`) re-picks
+ the same video, gets another 429, and re-arms a platform-wide cooldown with `fails+1`. The
+ `fails` count survives a restart: the new runner continued at attempt 10. A manual Sync reads the
+ same cooldown and refuses (`editor/app/channels/[slug]/pipelineActions.ts:133-140`). The one
+ video therefore blocked every YouTube channel for about 3.7 h, and the lane ran no other YouTube
+ unit in that window.
+- **What 429s is YouTube's subtitle (timedtext) endpoint, per video.** All 20 retained YouTube
+ 429s are `Unable to download video subtitles for 'en': HTTP Error 429`. The webpage, player API
+ and m3u8 requests in the same spawn succeed. There are **0** `Sign in to confirm` lines in any
+ retained log, and the only 2 `webpage … 429`s are Rumble syncs (`01M3AVGZ…`, `01M3AH7G…`).
+ **It is not IP-wide.** `the-quartering/UCirTfhUP3s` got the 429 at 16:41:37 UTC, and
+ `HasanAbiVODs3/1FibhXko_Kw` wrote both `en` and `en-orig` subtitles 27 s later (16:42:04), as did
+ `nuxanor` at 16:44. `CAi9jNrHetw` 429'd 3× then succeeded. `q_UPNCELU8I` 429'd 2× then
+ succeeded. Every affected video is on a Quartering channel.
+- **Volume before each 429 was tiny.** In the hour before 22:42, 23:10 and 23:39 UTC, the only
+ YouTube work was one or two attempts on the same video (2 yt-dlp spawns each). The sync at
+ 20:20 UTC 429'd after about 3.5 h with no YouTube spawns from this editor. Pacing between videos
+ could not have prevented any of the 16 unit 429s, because they were already ≥ 93 s apart.
+- **Per unit, a `handling: youtube` download is 2 spawns** (`downloadOneManaged.ts`):
+ - the metadata prefetch at `:646-661`, with no `-t sleep`: webpage, player, m3u8;
+ - the primary with `-t sleep` (`:168`): 2 timedtext requests (`en`, `en-orig` from
+ `en.*,live_chat`) 5 s apart.
+ On a subtitle 429, yt-dlp re-extracts from the URL, so one attempt is about 8 requests.
+- **Steady state when the lane has backlog: back-to-back, no gap.**
+ - The unit calls `downloadOneManaged` directly (`autoRunner.ts:2194`). `sleepBetweenDownloadsSeconds`
+ (live: 30) is read only by `runManagedDownloads` (`runYtdlp.ts:932-934,1061-1067`), which is
+ the sync path. FACTS.md:2733 records the same blindness for the backfill sweep.
+ - `runPool` wakes on every settle (`concurrentRunner.ts:118-126`), and `PER_PLATFORM_CAP = 1`
+ (`autoRunner.ts:1231`).
+ - Measured: units started 00:09:32, 00:09:55, 00:18:49 and 00:19:28 UTC. A success is followed
+ by the next start about 23-39 s later, which is ≈ 2 units/min ≈ 4 spawns ≈ 10 requests/min
+ at most.
+ - After a cooldown the lane resumes at full speed. A single success calls `clearBackoff`
+ (`:1957`), so the next 429 starts again at 60 s: 00:09:32 ok, then 00:09:55 429 → 55 s.
+- **Lanes sharing one YouTube key.**
+ - Auto-download units, syncs and metadata scans all queue on `platform:youtube` (concurrency 1,
+ `queueKeys.ts:69-73`), so they never overlap. The 41 syncs the scheduler queued at 16:05 UTC
+ ran serially (and were dropped at the 16:46 restart, which left 41 stale `queued` metas).
+ - Backfill re-acquire calls `downloadOneManaged` in-process (`backfillReacquire.ts:226`), with
+ no platform queue and no cooldown check. That lane is `enabled: false` today, so it is latent.
+ - No availability check ran in the window.
+- **No YouTube pacing args.** `PLATFORM_ARGS` has only Rumble (`channelArgs.ts:30`). YouTube
+ gets `-t sleep` only on the youtube-handling primary. The prefetch and transcribe-handling
+ downloads have no `--sleep-requests`.
+- **A second, real burst: the 142 members-only legal-mindset re-scan.**
+ - It rescans on every runner start: 16:03, 16:46, 00:19 and 03:39 UTC, each "0 scanned, 142
+ error(s)", with `--cookies-from-browser firefox` at 1 req/s.
+ - The 24 h suppression (`metadataScanStore.ts:334-349`) expired for good. The upsert keeps an
+ identical re-error untouched (`:249-252`: it rewrites only when class or message changed), so
+ `errors[id].at` is still 2026-09-20 and `metadataScanWanted` has said "yes" ever since.
+ - It was not the trigger tonight, because timing does not line up with the 429s. It is a real
+ cookie-authenticated burst.
+
+## The one change: defer a rate-limited video, not only its platform
+
+A `rate_limit` failure keeps the platform cooldown exactly as today, but **also takes that video
+off the lane for `VIDEO_RATE_LIMIT_DEFER_MS` (6 h)**. After the platform cooldown lapses, the lane
+offers the *next* video. On 09-24 ncdPaDSqt-c recovered about 3.7 h after its first 429, and the
+other two recovered within 9 min. Now `fails` escalates only while *different* videos keep
+getting 429s, which is a true platform signal. A single success clears it, as today.
+
+Why this and not the alternatives:
+- **Honouring the per-video gap (`sleepBetweenDownloadsSeconds`)** is correct and wanted, but on
+ this evidence it changes nothing: the failing attempts were already 93 s-30 min apart.
+- **A YouTube `--sleep-requests` entry** paces requests *within* a spawn. The 429 is on the first
+ timedtext GET after 5 s of `--sleep-subtitles`.
+- **A cooldown that also slows the resume** would stretch the same loop further. It would still
+ re-arm on the same video every 30 min.
+- **A cross-lane cap** is not needed: every YouTube path except the disabled backfill already
+ serializes on `platform:youtube`.
+
+Only the head-of-line retry explains three cooldowns in an evening at ≈ 2 spawns per 30 min.
+
+## Steps
+
+1. `common/jobs/platformBackoff.ts` (pure): export `VIDEO_RATE_LIMIT_DEFER_MS = 6 * 60 * 60_000`
+ and `isVideoDeferred(deferred: Map<string, number>, id, now)`, plus a prune of lapsed entries.
+ Do not change `nextBackoff`.
+2. `common/controller/autoRunner.ts`:
+ - Beside `platformInFlight` (about `:1230`), add `videoDeferredUntil = new Map<string, number>()`.
+ - In the download branch of `next()` (`:1617-1650`), also filter ids with
+ `isVideoDeferred(...)`. When everything left is deferred, set `anyCooling = true` so the idle
+ reason reads `cooldown`, not `capped`.
+ - In the `backoffHit` branch (`:1945-1955`), add
+ `videoDeferredUntil.set(pick.videoId, Date.now() + VIDEO_RATE_LIMIT_DEFER_MS)` only when
+ `failureClass === "rate_limit"`. Network failures keep today's retry-same-video behaviour.
+ - Extend the log line: `… ${pick.videoId} deferred 6h; next video after cooldown.`
+ - Keep it in-memory. A runner restart re-offers the video once, which costs at most one extra
+ attempt per restart and needs no state-file migration.
+3. `plans/FACTS.md`: record that the lane defers a rate-limited video and that the platform
+ cooldown now escalates across distinct videos. Note that YouTube 429s are timedtext and
+ per-video (evidence above).
+
+## Tests
+
+- Unit tests (`common/jobs/platformBackoff.test.ts`): `isVideoDeferred` before and after the
+ window, and prune.
+- `common/controller/autoRunner.test.ts`:
+ - A fake unit returns `rate_limit` for video A and success for B. After A's 429 and the
+ platform cooldown, the next pick is B, not A. B's success clears the platform cooldown. A is
+ not offered again until the deferral lapses (injected clock).
+ - `network` still re-picks the same video.
+ - With all pending videos deferred, the idle reason is `cooldown`.
+- No e2e change is needed. `auto-queue.spec.ts` should stay green, and should be run in the next
+ release's gate, not by this plan.
+
+## Rollout
+
+Ride the next release. Watch the auto-download runner log for one evening:
+- expect at most one 429 line per video per 6 h;
+- expect `attempt N` to climb only across different ids;
+- expect the owed md5-sweep sync not to be refused by a single-video cooldown.
+
+If the lane then shows 429s on *several distinct* videos within minutes, the limit is platform-wide
+after all and the cooldown is doing its job. The next lever is then the pacing gap below.
+
+## Found, deliberately not in this change (ranked)
+
+1. **Scan error `at` never refreshes** (`metadataScanStore.ts:249-252`): 142 cookie-authenticated
+ YouTube requests on every runner start. The fix is one line: rewrite an identical error when
+ `prev.at` is older than `METADATA_SCAN_ERROR_COOLDOWN_MS`. It is its own slice with its own test.
+2. **The lane ignores `sleepBetweenDownloadsSeconds`**. The fix is a per-platform `nextStartAt` set
+ on unit settle. This is the "slow" half of the operator's rule, for when backlog is large.
+3. **Backfill re-acquire bypasses the platform queue and cooldown** (`backfillReacquire.ts:226`).
+ It is latent while the lane is disabled.
+4. **`en.*` fetches both `en` and `en-orig`**, which is 2 timedtext GETs per video. That is a
+ `subLangs` question for the operator, not a pacing one.
+
+Out of scope: changing `nextBackoff` constants; any yt-dlp client or PO-token work on the
+timedtext 429 itself; adding YouTube to `PLATFORM_ARGS`; the stale `queued` metas left by restarts.