commit 11f127474fc7c5edf9a18a153bd5b5a3ca85c692
parent 797c76b9a5f3236cebf194b3397bd7b942e42dfc
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Fri, 25 Sep 2026 19:12:40 -0400
plans: release 9 — slice F as shipped (four fixes + the three lows); [Unreleased] bullets
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Diffstat:
2 files changed, 183 insertions(+), 0 deletions(-)
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -1,6 +1,12 @@
# Changelog
## [Unreleased]
+- **A job that waited in a queue no longer ends `failed` after doing its work.** A sync, download or other job queued behind another on the same platform ran its final page refresh outside any request, where Next refuses it, so the job read `failed` and `pnpm ops … --wait` exited 1 even though the work was done (the teamrcn sync on 2026-09-25). The refresh is now skipped there with one warning in the server log; the pages re-read disk on their next load anyway.
+- **Downloads no longer sleep after a video that fetched nothing.** The "sleep between downloads" (30 s by default) ran after every video, including each one the channel's download filter declined and each members-only or removed video that failed before any media request. A filtered channel's download-missing slept 193 times for 14 downloads on 2026-09-25. It now sleeps only after a real fetch, success or failure, and still after every rate-limit or network failure.
+- **YouTube requests are paced at one per second.** Every yt-dlp run against YouTube now carries `--sleep-requests 1`, as Rumble's already did: the 429 investigation found YouTube had no request-level pacing at all. A channel's own `ytdlpExtraArgs` still wins, because it comes after.
+- **`/jobs` no longer calls a slow but moving job stuck.** "STUCK · POSSIBLY-STALLED" now needs the job's progress to have stood still for 10 minutes, not just the job to be 10 minutes old with nothing in flight. A metadata scan at ~10 videos a minute read stuck on 2026-09-25.
+- **Jobs left `queued` by a restart are settled at boot.** A job still waiting when the server stopped used to sit on `/jobs` as queued forever. On boot, each one that can be replayed is queued again (the same path as Retry) and the old row ends `cancelled`, naming the new job; one that cannot is ended `cancelled` with the reason. With `ARCHILYZER_IDLE_BOOT` set, all of them are cancelled and nothing is re-queued. One line per job in the server log.
+- Smaller fixes from the release 8 reviews: a manual 429 cooldown merged with the runner's keeps the higher failure count as well as the later end; a channel's video-title memo keeps its titles when a new video directory appears, instead of re-reading every one; and the homepage counts "Transcripts" from the same pass that places them on the chart, counts "gone at the source, still here" only for recordings it holds, and drops the two empty columns its stats strip had when nothing is gone. No number on today's homepage changes.
- **The yt-dlp clip command is back on sites with transcript downloads turned off.** Turning off `transcriptDownloads` (site.json, or the hub's homepage.json) hid three buttons in the transcript viewer. One of them, the yt-dlp button, only copies a `yt-dlp --download-sections` command for a marked clip to the clipboard and serves no file, so it is not a download. It now shows on every site. The switch still hides the Download menu (txt / srt / json) and Copy MD. The site and hub form labels in the editor say so. No setting changed; a site picks this up at its next build and deploy.
- **Export sites: a search restored from the last visit waits for you.** Opening a site (or `/ask`, or the hub) still loads the last query, the filters and the profile from the browser, but no longer runs the search on the first page of a visit; moving between pages after a search keeps it running, so the hub's chat still grounds in the search just done on its front page. The results show the video listing under those filters, and the bar says "Press Enter or click Search to apply", as for any unapplied edit. Search or Enter runs it; so does loading a profile. A link with a query in it (`?qt=`) still runs on arrival. Going straight to the hub's `/ask` in a new visit leaves the restored search held, and that page has no search bar: search on the hub's front page first. The restored search used to re-fetch transcript shards (up to 8 MB each) on every visit to a device that had not cached them. Needs a rebuild and deploy of every export site.
- **The project homepage is rewritten, and it can be deployed as a preview.** The front page of `archilyzer.pages.dev` now opens on one line and a chart. The chart shows the official instances' transcripts by the month each video was published, from 2009 to now, stacked by instance. It is drawn when the site is built, so no chart script loads. Under it is one strip of numbers for the official instances: hours of speech, transcripts, recordings, channels, instances, and recordings gone at the source but still here. The **Official instances** section has a card for each public site, with its channels, recordings, transcripts and hours. **What it does** is now three short paragraphs. The recent-acquisitions list, "How it works" and "What this isn't" are gone from the page. The homepage summary (`homepage/public/homepage-summary.json`) is version 5. It adds `monthly`, `official` and per-site numbers and removes nothing, so `/stats` is unchanged. `archilyzer deploy homepage --preview <branch>` deploys a preview the way `deploy hub --preview` does. Without the flag it still deploys to production (`main`).
diff --git a/plans/release-9.md b/plans/release-9.md
@@ -0,0 +1,177 @@
+# Release 9 — four fixes behind one restart
+
+`main` at `9247211e`, release 8 live on :3001 since 2026-09-25 14:43. Release 9 is the four small
+bugs the release-8 rollout found on the live editor, batched so they cost one editor restart: a
+queued job that ends `failed` after doing its work, a 30 s sleep after every video the download
+filter declined, YouTube with no request-level pacing, and a `/jobs` page that calls a slow scan
+stuck and keeps restart-orphaned jobs `queued` forever. The release-8 reviews' optional lows ride
+along. Rules: `plans/tools/implementer-rules.md`. Record file: this file.
+
+## Record
+
+### Slice F, as shipped — four fixes (2026-09-25)
+
+Branch `one-core/r9-fixes` off `main` `9247211e`, one Opus implementer, no sibling slices.
+
+**B1 — a queued job no longer ends `failed` after succeeding.** The scheduler starts the next
+queued job synchronously inside `complete()`, so that job inherits the async context of whatever
+finished the one ahead. When the chain began at a job with no request behind it (an auto-runner
+unit, a heartbeat sync), Next has no work store and `revalidatePath` throws `Invariant: static
+generation store missing`. Job bodies call it last, so the throw landed after the work was done:
+the teamrcn sync at 15:21 read `failed` and `pnpm ops … --wait` exited 1.
+`editor/app/lib/safeRevalidate.ts` — `safeRevalidate(paths, tags = [])` — swallows exactly that
+invariant (one `console.warn` per process, naming the paths) and rethrows anything else. It is
+swapped in at every job body (`fn`) and job hook (`onDone` / `afterRun` / `afterDone`), found by
+grep and by indentation, not from the prompt's list alone: 18 files, 63 call lines, including
+`repointJob.ts`'s `afterDone`, which the list did not name. Server actions that revalidate in
+their own request are untouched, including the ones that revalidate after `drainStream`
+(`review/actions.ts`, `refreshAllChannelSnapshotsAction`), because they are still inside it.
+
+The first e2e version queued the sync behind a `--test-slow` sync. That held the queue for minutes,
+and it could not fail: a job an ops request started carries that request's store. The final version
+holds `platform:youtube` with `/api/test/stuck-job?releaseAfterMs=4000`. The new parameter finishes
+the fake holder from a timer armed inside `workAsyncStorage.exit`, i.e. outside any request, and
+the route reports `detached` so the spec cannot pass for the wrong reason. The spec then queues an
+ops sync behind it. **Against the pre-fix sync body it fails with the live line** — `[error]
+Invariant: static generation store missing in revalidatePath /channels/slow-b`, status `failed`
+(`r9-e2e-prefix.log`) — and it passes with the fix.
+
+**B2 — no 30 s sleep after a video that fetched no media.** `runManagedDownloads` slept
+`sleepBetweenDownloadsSeconds` after every video. `declinedWithoutMediaFetch(outcome,
+failureClass)` is true when every attempt was the `n: 0` metadata prefetch (or its cookie retry)
+and the video ended `skipped-filtered`, or `failed` with `failureClass === "per_video"`. Any
+attempt `n >= 1` is a real fetch: the download attempts and the chat-only pass. A failed chat pass
+leaves the status `skipped-filtered` and still sleeps. A rate-limit or network failure at the
+prefetch still sleeps. `runManagedDownloads` is exported with a `deps` seam (`downloadOne`,
+`sleep`). `managedDownloadsSleep.test.ts` drives the loop with canned outcomes: fetched sleeps,
+filtered does not, a per-video failure without a fetch does not, the last video never sleeps, a
+failed real fetch and a network failure at the prefetch both sleep, and a unit table for the
+predicate. **3 of the 6 fail on the old gate.**
+
+**B3 — YouTube request pacing.** `PLATFORM_ARGS.youtube = ["--sleep-requests", "1"]`, with the
+comment citing `~/reports/release-7/data/q-429-report.md` finding 1. The prompt named
+`release-8/data`; the report is in `release-7/data`. The comment above the table now says what the
+pace is, and that the metadata scan and the clip window, which already pass `--sleep-requests 1`,
+carry it twice, harmlessly (yt-dlp keeps the last). The argv tests were updated deliberately:
+- `channelArgs.test.ts`: youtube gets the pace; a platform with no entry gets nothing; a channel's
+ own `--sleep-requests 3` comes after and wins.
+- `platform-args.test.mjs`: +1 case, youtube carries the pace and a URL on no known platform
+ carries nothing. The YouTube clip fetch and simulate now expect the pace.
+
+No e2e asserts a YouTube argv exactly. `rumble-sweep.spec.ts` only asserts Rumble lines, and
+`fake-ytdlp.mjs`'s `lastNonFlag` never sees the trailing `1`, because every spawn ends in a URL,
+`--load-info-json <path>` or `-a <file>`.
+
+**B4a — `/jobs` flags a stall by quiet time, not age.** `JobRecord.progressAt` is stamped at the
+source. `noteProgress` (the `setProgress` every `runManagedFunction` job gets) stamps it only when
+the snapshot's numbers changed, and `recordTaskDuration` stamps it when a sub-operation finishes.
+`reconcileSlots` measures quiet time from the later of the start and that stamp. This deviates from
+the prompt's suggested view-side map, on purpose. A map is only fed while someone has `/jobs` open,
+so the first render after an hour away would either flag a healthy scan (if it seeded from the
+start) or hide a real stall for 10 more minutes (if it seeded from now). The stamp has neither
+problem, and the registry is in memory, so it is still evicted with the record.
+`jobRows.test.ts` +2: an hour of 6 s steps is never stuck, and progress frozen past 10 min is
+possibly-stalled. The same count re-reported is not a move. **Both fail on the old rule.** The
+`/api/test/stuck-job` fixture backdates `startedAt` with no `progressAt`, so `queue.spec.ts`'s
+force-release case still sees its stall.
+
+**B4b — boot settles the metas a restart left `queued`.** `common/jobs/bootQueuedJobs.ts`
+(`settleQueuedJobMetas`) runs from `instrumentation.ts`: lazy-imported, voided, and best-effort,
+after the storage boot pass and before the idle gate. For each `*.meta.json` with status `queued`
+that was queued before this boot and is not held by the live registry:
+- **It has a spec:** it is re-queued through `runJobSpec` (the path Retry uses). The old meta is
+ closed `cancelled` with `cancelReason: "server restarted before it ran; re-queued as <newId>"`.
+- **It has no spec, or the re-queue is refused or throws:** it is closed `cancelled` with the
+ reason.
+
+`ARCHILYZER_IDLE_BOOT` only cancels, and so does the e2e test server (`EDITOR_TEST_ROUTES=1`),
+whose leftover metas are the previous run's fixture. There is one console line per job, and the
+reason is appended to the job's `.log`, which also makes the pair prunable, since pruning walks
+`.log` ids. `running` metas are left alone. `JobMeta` gains the optional `cancelReason`.
+`bootQueuedJobs.test.ts` has 6 cases over a temp `.jobs` dir with an injected re-queue. It was not
+run against the real corpus.
+
+**Lows, all three done.**
+- **S:** `mergeBackoffEntry` merges `until` and `fails` separately, each to its max, in both the
+ runner's merge and `recordDownloadBackoff`'s. `platformBackoff.test.ts` +1.
+- **V:** a changed `data/` mtime REBASES the channel's title memo instead of dropping it. One
+ `readdir` keeps the titles of the dirs still present, only new dirs are read, and a removed
+ dir's title goes with it. The memo case in `videoTitles.test.ts` now asserts 1 read for a new
+ dir and a fall-through for a removed one. It pins a distinct mtime, because two changes inside
+ one filesystem tick share one.
+- **H:**
+ - `official.transcripts` is a per-site accumulator from the same pass that places each
+ transcript.
+ - A site outside `keptIds` gains no monthly key.
+ - `gone` counts only with a `downloadedDate`.
+ - `FamilyStats` is `lg:grid-cols-5` without a gone cell.
+ - The stale comments in `ArchiveGrowthChart.tsx` :10-11 and `page.tsx` :125, and the test
+ title, are fixed.
+
+ `homepageSummary.test.ts` +3. No number on today's data moves. The review measured placed +
+ unplaced = `official.transcripts` (49,767) and all 480 deleted records with a `downloadedDate`.
+
+| sha | what |
+|---|---|
+| `d927dad4` | B1: `app/lib/safeRevalidate.ts` + its unit test (3), swapped in at every job body and hook (18 files); first ops-api e2e case |
+| `37d263cb` | B2: `declinedWithoutMediaFetch`, the `deps` seam on `runManagedDownloads`, `managedDownloadsSleep.test.ts` (6) |
+| `ac55a23c` | B3: `PLATFORM_ARGS.youtube`, table comment, `channelArgs.test.ts` (+2), `platform-args.test.mjs` (+1) |
+| `efd75cb0` | B4a: `JobRecord.progressAt`, `noteProgress`, quiet-time stall rule, `jobRows.test.ts` (+2) |
+| `cc884951` | B4b: `bootQueuedJobs.ts` + test (6), `JobMeta.cancelReason`, the instrumentation boot pass |
+| `db18b73d` | low S: `mergeBackoffEntry` in both merges, test (+1) |
+| `49993934` | low V: title memo rebased on a new dir, test updated |
+| `c714bc14` | low H: `homepageSummary` counts + guard, `FamilyStats` grid, comments, test (+3) |
+| `78d165f6` | B1 e2e reworked: `stuck-job?releaseAfterMs` (detached release), the spec now reproduces the live failure pre-fix |
+| _this_ | `plans:` this record, the `[Unreleased]` bullets |
+
+**Gates**, all from the worktree root. Heavy steps started at ≥ 3 GB available memory.
+- **tsc** (`pnpm -r --no-bail --workspace-concurrency=1 exec tsc --noEmit`) was clean before every
+ code commit. It was checked once to catch a planted error, so the silent `exit=0` is real.
+- **common: 1,836/1,836**, from 1,816 + 20: B2 +6, B3 +2, B4a +2, B4b +6, S +1, H +3; V's case was
+ rewritten in place.
+- **editor unit: 78/78**, from 75 + 3 (`safeRevalidate.test.ts`).
+- **`test:scripts`: 162 pass + 1 skip of 163**, from 161 + 1 skip, +1 in `platform-args.test.mjs`.
+ A first run during e2e run 1 failed `E2E_QUEUE=0 bypasses the queue entirely`, because this
+ worktree's own e2e held the lock; the rerun with no e2e running is clean.
+- **mcp: 219/219.**
+- **Builds:** `pnpm --filter editor exec next build` ok, `pnpm --filter export exec next build` ok,
+ and, because H touches `homepage/`, `pnpm --filter homepage exec next build` ok. No dangling
+ `export/public` links.
+- **EDITOR e2e.** Spec list `r9-specs.txt`: `ops-api queues auto-queue lane-runner jobs-filters
+ queue` (the grep for `possibly-stalled|STUCK|data-job-id` adds `queue.spec.ts`), plus
+ `jobs-retry jobs-active-order jobs-channel` (the retry path B4b reuses), `rumble-sweep` (argv),
+ `title-filter` (B2's filtered path), `pacing metadata-scan-botcheck sync-deep`, all `.spec.ts`.
+ - Run 1 (`r9-e2e1.log`), on `c714bc14`: **84 passed, 2 failed, 13.4 min**. One failure was the
+ first B1 spec, which timed out behind the slow sync (the rework above). The other was
+ `auto-queue.spec.ts:218`, the run's first test, where the `/api/auto-queue/status` GET outlived
+ the 30 s test timeout on a cold dev compile. It passed in run 2 with no change.
+ - Pre-fix proof (`r9-e2e-prefix.log`): the reworked B1 spec with the sync body's
+ `revalidatePath` restored: **1 failed**, with the live invariant in the job log. The fix was
+ restored before any commit.
+ - Run 2 (`r9-e2e2.log`), on `78d165f6`: **86 passed, 0 failed, 5.7 min**, with no queue wait.
+
+**Numbers: none** (per the prompt). No `settings.json`, `site.json` or `config.json` key changed.
+`JobMeta.cancelReason` and `JobRecord.progressAt` are additive and optional.
+
+**Found and left.**
+- **B2 lets filtered prefetches run back to back.** A filtered video's metadata prefetch is still
+ one yt-dlp process against YouTube. Without the 30 s sleep, a run of declined videos is a run of
+ prefetches, spaced only by process start-up and, now, B3's one-second request pace inside each.
+ That is what the prompt asked for. The metadata scan does the same work at ~10/min with no 429s.
+ If a filtered channel's download-missing ever 429s, the lever is a smaller sleep for
+ prefetch-only videos, not the full one.
+- **B4b leaves `running` metas.** A job running at a hard crash (no graceful shutdown) still reads
+ `running` → `archived` on `/jobs`. Whether a half-done job should re-run is not a boot pass's
+ call.
+- **`cancelReason` is not drawn on `/jobs`.** It is in the meta and appended to the job's log,
+ which the row's log view shows.
+- **The rollout's first boot WILL re-queue.** Every `queued` meta on the live corpus with a spec is
+ re-submitted at the restart. Count them before the restart:
+ `grep -l '"status":"queued"' transcripts/.jobs/*.meta.json | wc -l`. An operator who wants none
+ of them run should boot with `ARCHILYZER_IDLE_BOOT=1` once.
+- **The stuck-job harness reaches into a Next internal**
+ (`next/dist/server/app-render/work-async-storage.external`). It is typed, test-route-only, and
+ reported through `detached`, so a Next upgrade that moves it fails the spec loudly rather than
+ letting it pass vacuously.
+
+## Rollout