Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 21e45803f27f750d7837704cee76a7d1fbd12457
parent b636936b10f9eed83fbd9f2d00c8692ff6a63044
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Fri, 25 Sep 2026 19:27:15 -0400

plans: release 9 — the review fixes in the slice F record; CHANGELOG bullets for B2 and B4b corrected

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 4++--
Mplans/release-9.md | 121++++++++++++++++++++++++++++++++++++++++++++++++++++++++-----------------------
2 files changed, 88 insertions(+), 37 deletions(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -2,10 +2,10 @@ ## [Unreleased] - **A job that waited in a queue no longer ends `failed` after doing its work.** A sync, download or other job queued behind another on the same platform ran its final page refresh outside any request, where Next refuses it, so the job read `failed` and `pnpm ops … --wait` exited 1 even though the work was done (the teamrcn sync on 2026-09-25). The refresh is now skipped there with one warning in the server log; the pages re-read disk on their next load anyway. -- **Downloads no longer sleep after a video that fetched nothing.** The "sleep between downloads" (30 s by default) ran after every video, including each one the channel's download filter declined and each members-only or removed video that failed before any media request. A filtered channel's download-missing slept 193 times for 14 downloads on 2026-09-25. It now sleeps only after a real fetch, success or failure, and still after every rate-limit or network failure. +- **Downloads no longer sleep after a video the download filter declined.** The "sleep between downloads" (30 s by default) ran after every video, including each one the channel's download filter declined before fetching anything. A filtered channel's download-missing slept 193 times for 14 downloads on 2026-09-25. It still sleeps after every real fetch and after every failure, per-video ones included. - **YouTube requests are paced at one per second.** Every yt-dlp run against YouTube now carries `--sleep-requests 1`, as Rumble's already did: the 429 investigation found YouTube had no request-level pacing at all. A channel's own `ytdlpExtraArgs` still wins, because it comes after. - **`/jobs` no longer calls a slow but moving job stuck.** "STUCK · POSSIBLY-STALLED" now needs the job's progress to have stood still for 10 minutes, not just the job to be 10 minutes old with nothing in flight. A metadata scan at ~10 videos a minute read stuck on 2026-09-25. -- **Jobs left `queued` by a restart are settled at boot.** A job still waiting when the server stopped used to sit on `/jobs` as queued forever. On boot, each one that can be replayed is queued again (the same path as Retry) and the old row ends `cancelled`, naming the new job; one that cannot is ended `cancelled` with the reason. With `ARCHILYZER_IDLE_BOOT` set, all of them are cancelled and nothing is re-queued. One line per job in the server log. +- **Jobs left `queued` by a restart are settled at boot.** A job still waiting when the server stopped used to sit on `/jobs` as queued forever. On boot each one ends `cancelled` with its reason, and only a few are queued again, through the same path as Retry. A sync is never re-queued, because the scheduler re-derives syncs at its own pace. Nothing queued more than a day before the restart is re-queued. Of several identical jobs, only the newest is re-queued. With `ARCHILYZER_IDLE_BOOT` set, none are. The server log gets one summary line, plus one line per job queued again. - Smaller fixes from the release 8 reviews: a manual 429 cooldown merged with the runner's keeps the higher failure count as well as the later end; a channel's video-title memo keeps its titles when a new video directory appears, instead of re-reading every one; and the homepage counts "Transcripts" from the same pass that places them on the chart, counts "gone at the source, still here" only for recordings it holds, and drops the two empty columns its stats strip had when nothing is gone. No number on today's homepage changes. - **The yt-dlp clip command is back on sites with transcript downloads turned off.** Turning off `transcriptDownloads` (site.json, or the hub's homepage.json) hid three buttons in the transcript viewer. One of them, the yt-dlp button, only copies a `yt-dlp --download-sections` command for a marked clip to the clipboard and serves no file, so it is not a download. It now shows on every site. The switch still hides the Download menu (txt / srt / json) and Copy MD. The site and hub form labels in the editor say so. No setting changed; a site picks this up at its next build and deploy. - **Export sites: a search restored from the last visit waits for you.** Opening a site (or `/ask`, or the hub) still loads the last query, the filters and the profile from the browser, but no longer runs the search on the first page of a visit; moving between pages after a search keeps it running, so the hub's chat still grounds in the search just done on its front page. The results show the video listing under those filters, and the bar says "Press Enter or click Search to apply", as for any unapplied edit. Search or Enter runs it; so does loading a profile. A link with a query in it (`?qt=`) still runs on arrival. Going straight to the hub's `/ask` in a new visit leaves the restored search held, and that page has no search bar: search on the hub's front page first. The restored search used to re-fetch transcript shards (up to 8 MB each) on every visit to a device that had not cached them. Needs a rebuild and deploy of every export site. diff --git a/plans/release-9.md b/plans/release-9.md @@ -36,17 +36,21 @@ ops sync behind it. **Against the pre-fix sync body it fails with the live line* Invariant: static generation store missing in revalidatePath /channels/slow-b`, status `failed` (`r9-e2e-prefix.log`) — and it passes with the fix. -**B2 — no 30 s sleep after a video that fetched no media.** `runManagedDownloads` slept -`sleepBetweenDownloadsSeconds` after every video. `declinedWithoutMediaFetch(outcome, -failureClass)` is true when every attempt was the `n: 0` metadata prefetch (or its cookie retry) -and the video ended `skipped-filtered`, or `failed` with `failureClass === "per_video"`. Any -attempt `n >= 1` is a real fetch: the download attempts and the chat-only pass. A failed chat pass -leaves the status `skipped-filtered` and still sleeps. A rate-limit or network failure at the -prefetch still sleeps. `runManagedDownloads` is exported with a `deps` seam (`downloadOne`, -`sleep`). `managedDownloadsSleep.test.ts` drives the loop with canned outcomes: fetched sleeps, -filtered does not, a per-video failure without a fetch does not, the last video never sleeps, a -failed real fetch and a network failure at the prefetch both sleep, and a unit table for the -predicate. **3 of the 6 fail on the old gate.** +**B2 — no 30 s sleep after a video the download filter declined.** `runManagedDownloads` slept +`sleepBetweenDownloadsSeconds` after every video. `declinedWithoutMediaFetch(outcome)` is true when +every attempt was the `n: 0` metadata prefetch (or its cookie retry) and the video ended +`skipped-filtered`. Any attempt `n >= 1` is a real fetch: the download attempts and the chat-only +pass. A failed chat pass leaves the status `skipped-filtered` and still sleeps. As first shipped, a +`per_video` failure that never got past the prefetch also skipped the sleep. The review narrowed +that (see "Review fixes"): **every failure keeps the pace.** `runManagedDownloads` is exported with a +`deps` seam (`downloadOne`, `sleep`). `managedDownloadsSleep.test.ts` drives the loop with canned +outcomes: +- a fetched video sleeps; +- a filtered one does not; +- a per-video failure at the prefetch sleeps; +- the last video never sleeps; +- a failed real fetch and a network failure at the prefetch both sleep; +- a unit table for the predicate. **B3 — YouTube request pacing.** `PLATFORM_ARGS.youtube = ["--sleep-requests", "1"]`, with the comment citing `~/reports/release-7/data/q-429-report.md` finding 1. The prompt named @@ -75,21 +79,41 @@ possibly-stalled. The same count re-reported is not a move. **Both fail on the o `/api/test/stuck-job` fixture backdates `startedAt` with no `progressAt`, so `queue.spec.ts`'s force-release case still sees its stall. -**B4b — boot settles the metas a restart left `queued`.** `common/jobs/bootQueuedJobs.ts` -(`settleQueuedJobMetas`) runs from `instrumentation.ts`: lazy-imported, voided, and best-effort, -after the storage boot pass and before the idle gate. For each `*.meta.json` with status `queued` -that was queued before this boot and is not held by the live registry: -- **It has a spec:** it is re-queued through `runJobSpec` (the path Retry uses). The old meta is - closed `cancelled` with `cancelReason: "server restarted before it ran; re-queued as <newId>"`. -- **It has no spec, or the re-queue is refused or throws:** it is closed `cancelled` with the - reason. - -`ARCHILYZER_IDLE_BOOT` only cancels, and so does the e2e test server (`EDITOR_TEST_ROUTES=1`), -whose leftover metas are the previous run's fixture. There is one console line per job, and the -reason is appended to the job's `.log`, which also makes the pair prunable, since pruning walks -`.log` ids. `running` metas are left alone. `JobMeta` gains the optional `cancelReason`. -`bootQueuedJobs.test.ts` has 6 cases over a temp `.jobs` dir with an injected re-queue. It was not -run against the real corpus. +**B4b — boot settles the metas a restart left `queued`, and re-queues little.** +`common/jobs/bootQueuedJobs.ts` (`settleQueuedJobMetas`) runs from `instrumentation.ts`: +lazy-imported, voided and best-effort. It runs after the storage boot pass, and since the review +fix it WAITS for that pass. Its scope is each `*.meta.json` with status `queued` that was queued +before this boot and is not held by the live registry. A malformed meta is skipped. In order: +1. An idle boot (`ARCHILYZER_IDLE_BOOT`) and the e2e test server (`EDITOR_TEST_ROUTES=1`) cancel + everything and re-queue nothing. +2. A `sync` is never re-queued: "server restarted; the scheduler re-derives syncs". +3. A meta queued more than 24 h (`REQUEUE_MAX_AGE_MS`) before the boot is cancelled: "queued + before the last restart, stale". +4. A meta with no spec is cancelled with the reason. +5. Of what is left, only the NEWEST meta per kind + channel + bucket + params (params key-sorted) + is re-queued, through `runJobSpec` (the path Retry uses). The old meta is closed `cancelled`, + naming the new id. The others are cancelled as "superseded by a newer queued job (<id>)". A + refused or throwing re-queue is cancelled with its error. + +The log gets one line per re-queued job and one summary line: re-queued N, cancelled M, by reason. +The reason is also appended to each job's `.log`, which makes the pair prunable, since pruning walks +`.log` ids. Metas are written tmp + rename. `running` metas are left alone. `JobMeta` gains the +optional `cancelReason`. + +**Live size, from the review's read-only count.** 1,396 metas, 392 `queued`: +- 346 `sync` over 51 channels, in batches from 08-03 to 09-24; +- 29 `whisper-all` from June and July; +- about 7 others with a spec; +- 10 with none. + +The first version would have called `runJobSpec` about 382 times at boot. Under the current rules, +the syncs and everything older than 24 h are cancelled, so at most a handful of fresh non-sync jobs +come back. + +`bootQueuedJobs.test.ts` has 9 cases over a temp `.jobs` dir with an injected re-queue. One is 7 +duplicate syncs + 1 stale `whisper-all` + 1 fresh non-sync, which gives exactly 1 re-queue. Other +cases cover duplicate specs written in a different key order, and a malformed meta next to a good +one. It was not run against the real corpus. **Lows, all three done.** - **S:** `mergeBackoffEntry` merges `until` and `fails` separately, each to its max, in both the @@ -122,7 +146,11 @@ run against the real corpus. | `49993934` | low V: title memo rebased on a new dir, test updated | | `c714bc14` | low H: `homepageSummary` counts + guard, `FamilyStats` grid, comments, test (+3) | | `78d165f6` | B1 e2e reworked: `stuck-job?releaseAfterMs` (detached release), the spec now reproduces the live failure pre-fix | -| _this_ | `plans:` this record, the `[Unreleased]` bullets | +| `0c3e0b9a` | `plans:` this record, the `[Unreleased]` bullets | +| `be48f610` | (review HIGH-1, LOW-1, LOW-2, NIT-1) boot pass: no syncs, 24 h age cap, newest per spec, summary log; storage pass awaited; atomic meta writes; test 6 → 9 | +| `11c947ab` | (review MED) B2 skips the sleep for `skipped-filtered` only; the per-video test now expects a sleep | +| `23134785` | (review LOW-3) the B1 e2e releases its holder after 10 s, not 4 | +| _this_ | `plans:` the review fixes in this record, the CHANGELOG and the gates | **Gates**, all from the worktree root. Heavy steps started at ≥ 3 GB available memory. - **tsc** (`pnpm -r --no-bail --workspace-concurrency=1 exec tsc --noEmit`) was clean before every @@ -150,25 +178,48 @@ run against the real corpus. restored before any commit. - Run 2 (`r9-e2e2.log`), on `78d165f6`: **86 passed, 0 failed, 5.7 min**, with no queue wait. +**Review fixes (review verdict SHIP AFTER FIXES, `r9-review.md`).** +- **HIGH-1** (the boot-flood size), **LOW-1** (await the storage pass), **LOW-2** (malformed-meta + test) and **NIT-1** (atomic meta writes) are all in `be48f610`. +- **MED** (B2 keeps the pace after a per-video failure) is `11c947ab`. +- **LOW-3** (10 s release window) is `23134785`. +- **LOW-4** is left; see "Found and left". + +Re-gate on `23134785`: +- tsc clean. +- common **1,839/1,839** (1,836 + 3 boot-pass cases). +- editor unit **78/78**. +- EDITOR e2e `jobs-filters.spec.ts ops-api.spec.ts queues.spec.ts queue.spec.ts` + (`r9-e2e3.log`): **57 passed, 0 failed, 2.7 min**. + +No e2e covers the boot pass. The test server is cancel-only by design, so the pass is pinned by its +9 unit cases. + **Numbers: none** (per the prompt). No `settings.json`, `site.json` or `config.json` key changed. `JobMeta.cancelReason` and `JobRecord.progressAt` are additive and optional. **Found and left.** - **B2 lets filtered prefetches run back to back.** A filtered video's metadata prefetch is still one yt-dlp process against YouTube. Without the 30 s sleep, a run of declined videos is a run of - prefetches, spaced only by process start-up and, now, B3's one-second request pace inside each. - That is what the prompt asked for. The metadata scan does the same work at ~10/min with no 429s. - If a filtered channel's download-missing ever 429s, the lever is a smaller sleep for - prefetch-only videos, not the full one. + prefetches, spaced only by process start-up and B3's one-second request pace inside each. The + metadata scan does the same work at ~10/min with no 429s. If a filtered channel's + download-missing ever 429s, the lever is a smaller sleep for prefetch-only videos, not the full + one. +- **Classification quirk, left as is.** YouTube's soft block ("This content isn't available, try + again later") classifies as `deleted` → `per_video` in `classifyDownloadFailure`. So a soft block + neither triggers the platform backoff nor aborts a batch, and it reads as a per-video property. It + is why B2 keeps the sleep after every per-video failure. `classifyDownloadFailure` was not changed + (review instruction). - **B4b leaves `running` metas.** A job running at a hard crash (no graceful shutdown) still reads `running` → `archived` on `/jobs`. Whether a half-done job should re-run is not a boot pass's call. - **`cancelReason` is not drawn on `/jobs`.** It is in the meta and appended to the job's log, which the row's log view shows. -- **The rollout's first boot WILL re-queue.** Every `queued` meta on the live corpus with a spec is - re-submitted at the restart. Count them before the restart: - `grep -l '"status":"queued"' transcripts/.jobs/*.meta.json | wc -l`. An operator who wants none - of them run should boot with `ARCHILYZER_IDLE_BOOT=1` once. +- **The rollout's first boot settles the 392 live `queued` metas.** Nearly all are cancelled with + a reason, as above; only fresh (< 24 h) non-sync jobs, newest per spec, come back. Booting once with + `ARCHILYZER_IDLE_BOOT=1` re-queues none. +- **LOW-4, left:** `safeRevalidate`'s "once per process" warning is once per module instance. Next + can load a module more than once, so it may warn a second time. - **The stuck-job harness reaches into a Next internal** (`next/dist/server/app-render/work-async-storage.external`). It is typed, test-route-only, and reported through `detached`, so a Next upgrade that moves it fails the spec loudly rather than