> **Corrected again 2026-10-01 by `plans/release-17.md` "## Slice RL — the ruling" (shipped as > "### Slice RL, as shipped" there).** A subtitle 429 — every YouTube 429 this file measured, and > all 78 of 2026-10-01's — no longer backs the platform off or defers the video: it is its own > class, `subs_rate_limit`, the download goes on to the media, and only the video's subtitles are > deferred. The "next lever" named below, honouring `sleepBetweenDownloadsSeconds` in the lane, is > built (plus an adaptive per-platform `--sleep-requests`), and a rate limit that outlasts the > cooldown cap now holds the platform to one probe at a time. The finding "pacing between videos > could not have prevented" a timedtext 429 stands. # Plan: one YouTube video must not keep the whole platform in a 429 cooldown > **Corrected by `plans/release-7.md` "## Slice Y" and shipped as "### Slice Y, as shipped" > there (2026-09-25).** The deferral is PERSISTED as `videoDeferrals` beside `platformBackoff` > in `.auto-queue/state.json`, not in memory (a rollout restart would otherwise re-hit the video > at `fails+1`). The outcome logic moved to a pure seam, `common/jobs/unitOutcome.ts`, because > `runLoop` is not exported. An all-deferred lane idles with its own reason, `deferred`, not > `cooldown`. The page lists deferred videos. The identical-scan-error `at` refresh is also in > the slice. `sleepBetweenDownloadsSeconds` stays out. The steps and tests below are the > original plan; read the release-7 record for what shipped. **Found 2026-09-25**, from STATE.md: "YouTube held a 429 cooldown across three attempts (18:43, 19:13, 19:41)", and the owed md5-sweep sync was refused all three times. Verified read-only on `main` `93dcb532`, the live `settings.json` and the retained `transcripts/.jobs/*` (meta `startedAt` in UTC). **The STATE times are EDT.** They are the unit attempts at 22:42, 23:10 and 23:39 UTC. ## Findings (measured, not guessed) - **It was one video, not a burst.** All three cooldowns were the auto-download lane re-trying `quarteringvlogs/ncdPaDSqt-c`, a 178 s Short. It got 12 consecutive `rate_limit` failures from 20:28 to 23:39 UTC, and the runner logs `01M3A5034Y…` and `01M3ASBJHW…` show attempts 1→12 (attempt 9 was the manual `download-missing` `01M3AQJ5…`, which shares the state). It succeeded at 00:09 UTC. Attempt spacing matches `nextBackoff` exactly: +93 s, +145, +252, +477, +920, then about 28-31 min at the cap (`platformBackoff.ts:22,24`: 60 s base, doubling, 30 min cap, ±10 %). The cap was reached at about 21:30 UTC. - **Why the same video every time.** On `rate_limit`/`network` the runner deliberately does NOT `markCompleted` the video (`autoRunner.ts:1936-1962`), and `order: "listed"` puts it back at the head of the pick. So each lapse of the cooldown (3 s idle poll, `autoRunner.ts:158`) re-picks the same video, gets another 429, and re-arms a platform-wide cooldown with `fails+1`. The `fails` count survives a restart: the new runner continued at attempt 10. A manual Sync reads the same cooldown and refuses (`editor/app/channels/[slug]/pipelineActions.ts:133-140`). The one video therefore blocked every YouTube channel for about 3.7 h, and the lane ran no other YouTube unit in that window. - **What 429s is YouTube's subtitle (timedtext) endpoint, per video.** All 20 retained YouTube 429s are `Unable to download video subtitles for 'en': HTTP Error 429`. The webpage, player API and m3u8 requests in the same spawn succeed. There are **0** `Sign in to confirm` lines in any retained log, and the only 2 `webpage … 429`s are Rumble syncs (`01M3AVGZ…`, `01M3AH7G…`). **It is not IP-wide.** `the-quartering/UCirTfhUP3s` got the 429 at 16:41:37 UTC, and `HasanAbiVODs3/1FibhXko_Kw` wrote both `en` and `en-orig` subtitles 27 s later (16:42:04), as did `nuxanor` at 16:44. `CAi9jNrHetw` 429'd 3× then succeeded. `q_UPNCELU8I` 429'd 2× then succeeded. Every affected video is on a Quartering channel. - **Volume before each 429 was tiny.** In the hour before 22:42, 23:10 and 23:39 UTC, the only YouTube work was one or two attempts on the same video (2 yt-dlp spawns each). The sync at 20:20 UTC 429'd after about 3.5 h with no YouTube spawns from this editor. Pacing between videos could not have prevented any of the 16 unit 429s, because they were already ≥ 93 s apart. - **Per unit, a `handling: youtube` download is 2 spawns** (`downloadOneManaged.ts`): - the metadata prefetch at `:646-661`, with no `-t sleep`: webpage, player, m3u8; - the primary with `-t sleep` (`:168`): 2 timedtext requests (`en`, `en-orig` from `en.*,live_chat`) 5 s apart. On a subtitle 429, yt-dlp re-extracts from the URL, so one attempt is about 8 requests. - **Steady state when the lane has backlog: back-to-back, no gap.** - The unit calls `downloadOneManaged` directly (`autoRunner.ts:2194`). `sleepBetweenDownloadsSeconds` (live: 30) is read only by `runManagedDownloads` (`runYtdlp.ts:932-934,1061-1067`), which is the sync path. FACTS.md:2733 records the same blindness for the backfill sweep. - `runPool` wakes on every settle (`concurrentRunner.ts:118-126`), and `PER_PLATFORM_CAP = 1` (`autoRunner.ts:1231`). - Measured: units started 00:09:32, 00:09:55, 00:18:49 and 00:19:28 UTC. A success is followed by the next start about 23-39 s later, which is ≈ 2 units/min ≈ 4 spawns ≈ 10 requests/min at most. - After a cooldown the lane resumes at full speed. A single success calls `clearBackoff` (`:1957`), so the next 429 starts again at 60 s: 00:09:32 ok, then 00:09:55 429 → 55 s. - **Lanes sharing one YouTube key.** - Auto-download units, syncs and metadata scans all queue on `platform:youtube` (concurrency 1, `queueKeys.ts:69-73`), so they never overlap. The 41 syncs the scheduler queued at 16:05 UTC ran serially (and were dropped at the 16:46 restart, which left 41 stale `queued` metas). - Backfill re-acquire calls `downloadOneManaged` in-process (`backfillReacquire.ts:226`), with no platform queue and no cooldown check. That lane is `enabled: false` today, so it is latent. - No availability check ran in the window. - **No YouTube pacing args.** `PLATFORM_ARGS` has only Rumble (`channelArgs.ts:30`). YouTube gets `-t sleep` only on the youtube-handling primary. The prefetch and transcribe-handling downloads have no `--sleep-requests`. - **A second, real burst: the 142 members-only legal-mindset re-scan.** - It rescans on every runner start: 16:03, 16:46, 00:19 and 03:39 UTC, each "0 scanned, 142 error(s)", with `--cookies-from-browser firefox` at 1 req/s. - The 24 h suppression (`metadataScanStore.ts:334-349`) expired for good. The upsert keeps an identical re-error untouched (`:249-252`: it rewrites only when class or message changed), so `errors[id].at` is still 2026-09-20 and `metadataScanWanted` has said "yes" ever since. - It was not the trigger tonight, because timing does not line up with the 429s. It is a real cookie-authenticated burst. ## The one change: defer a rate-limited video, not only its platform A `rate_limit` failure keeps the platform cooldown exactly as today, but **also takes that video off the lane for `VIDEO_RATE_LIMIT_DEFER_MS` (6 h)**. After the platform cooldown lapses, the lane offers the *next* video. On 09-24 ncdPaDSqt-c recovered about 3.7 h after its first 429, and the other two recovered within 9 min. Now `fails` escalates only while *different* videos keep getting 429s, which is a true platform signal. A single success clears it, as today. Why this and not the alternatives: - **Honouring the per-video gap (`sleepBetweenDownloadsSeconds`)** is correct and wanted, but on this evidence it changes nothing: the failing attempts were already 93 s-30 min apart. - **A YouTube `--sleep-requests` entry** paces requests *within* a spawn. The 429 is on the first timedtext GET after 5 s of `--sleep-subtitles`. - **A cooldown that also slows the resume** would stretch the same loop further. It would still re-arm on the same video every 30 min. - **A cross-lane cap** is not needed: every YouTube path except the disabled backfill already serializes on `platform:youtube`. Only the head-of-line retry explains three cooldowns in an evening at ≈ 2 spawns per 30 min. ## Steps 1. `common/jobs/platformBackoff.ts` (pure): export `VIDEO_RATE_LIMIT_DEFER_MS = 6 * 60 * 60_000` and `isVideoDeferred(deferred: Map, id, now)`, plus a prune of lapsed entries. Do not change `nextBackoff`. 2. `common/controller/autoRunner.ts`: - Beside `platformInFlight` (about `:1230`), add `videoDeferredUntil = new Map()`. - In the download branch of `next()` (`:1617-1650`), also filter ids with `isVideoDeferred(...)`. When everything left is deferred, set `anyCooling = true` so the idle reason reads `cooldown`, not `capped`. - In the `backoffHit` branch (`:1945-1955`), add `videoDeferredUntil.set(pick.videoId, Date.now() + VIDEO_RATE_LIMIT_DEFER_MS)` only when `failureClass === "rate_limit"`. Network failures keep today's retry-same-video behaviour. - Extend the log line: `… ${pick.videoId} deferred 6h; next video after cooldown.` - Keep it in-memory. A runner restart re-offers the video once, which costs at most one extra attempt per restart and needs no state-file migration. 3. `plans/FACTS.md`: record that the lane defers a rate-limited video and that the platform cooldown now escalates across distinct videos. Note that YouTube 429s are timedtext and per-video (evidence above). ## Tests - Unit tests (`common/jobs/platformBackoff.test.ts`): `isVideoDeferred` before and after the window, and prune. - `common/controller/autoRunner.test.ts`: - A fake unit returns `rate_limit` for video A and success for B. After A's 429 and the platform cooldown, the next pick is B, not A. B's success clears the platform cooldown. A is not offered again until the deferral lapses (injected clock). - `network` still re-picks the same video. - With all pending videos deferred, the idle reason is `cooldown`. - No e2e change is needed. `auto-queue.spec.ts` should stay green, and should be run in the next release's gate, not by this plan. ## Rollout Ride the next release. Watch the auto-download runner log for one evening: - expect at most one 429 line per video per 6 h; - expect `attempt N` to climb only across different ids; - expect the owed md5-sweep sync not to be refused by a single-video cooldown. If the lane then shows 429s on *several distinct* videos within minutes, the limit is platform-wide after all and the cooldown is doing its job. The next lever is then the pacing gap below. ## Found, deliberately not in this change (ranked) 1. **Scan error `at` never refreshes** (`metadataScanStore.ts:249-252`): 142 cookie-authenticated YouTube requests on every runner start. The fix is one line: rewrite an identical error when `prev.at` is older than `METADATA_SCAN_ERROR_COOLDOWN_MS`. It is its own slice with its own test. 2. **The lane ignores `sleepBetweenDownloadsSeconds`**. The fix is a per-platform `nextStartAt` set on unit settle. This is the "slow" half of the operator's rule, for when backlog is large. 3. **Backfill re-acquire bypasses the platform queue and cooldown** (`backfillReacquire.ts:226`). It is latent while the lane is disabled. 4. **`en.*` fetches both `en` and `en-orig`**, which is 2 timedtext GETs per video. That is a `subLangs` question for the operator, not a pacing one. Out of scope: changing `nextBackoff` constants; any yt-dlp client or PO-token work on the timedtext 429 itself; adding YouTube to `PLATFORM_ARGS`; the stale `queued` metas left by restarts.