Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 612855aad2d93e7baa7c0841813dac96b322c3a4
parent 921043fa26c46e2d2bc816c3d1081277ab695c5c
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Thu,  1 Oct 2026 23:02:20 -0400

plans: slice D0's review and its fixes — the memo keyed by settings, one queued successor, Refresh report on the queue, the yield pinned, the rollout note; e2e runs 3–5; the changelog bullets follow

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 4++--
Mplans/release-17.md | 50+++++++++++++++++++++++++++++++++++++++++++++++++-
2 files changed, 51 insertions(+), 3 deletions(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -16,8 +16,8 @@ - **A form whose save is refused keeps what you typed.** Every editor form put its plain fields back to the stored values when its save was refused — a site's ID rejected, a page size out of range, a slug already taken — so everything typed had to be typed again. A refused save now leaves every field as you left it, beside the reason: **Settings**; a site's form (new and existing); the hub's config on `/sites`; **Cut release**; a channel's form (new and **Configure**), **Rename** and **Delete**; a video's **Delete directory**; **Drive health timing** on `/storage`; the backup config on `/saved-videos`; the sync operation's controls; the **Digest**, **Diarization**, **Speaker attribution** and **Speaker work lane** settings; and the worker list on `/workers`. A save that succeeds behaves as before, with one difference you may notice: a drop-down, and a checkbox or choice that the page tracks as you change it (a cadence, a worker's **Enabled**, a social link's **Keep in header**, a site membership, a site's accent), now shows what was saved. A form's own drop-downs used to go back to what the page had loaded with until a reload, and a second save from the same page sent that old choice again; the others went back until the page next refreshed itself (every 5 seconds by default). - **A media move no longer starts over a job that is writing into the channel, holds the channel's writers while it runs, and makes its copy match the source before it verifies — so a transcription or a download during a move cannot fail it.** A move that has waited its turn behind other moves now checks again when it starts: if a job is running on the channel, or an auto-queue lane is working on one of its videos, it stops at once and says which ("a transcription of abc123 is running (Transcribe all, job …) — wait for it or cancel it"), with nothing copied — and a job you have just cancelled counts until it has actually stopped ("is stopping … — wait for it to stop"); **Preview** says the same, and the Storage panel's blocked message now names the job too. While a move's marker stands, the channel is held: every lane skips it, and every job that reads or writes its media (single-video transcriptions, downloads and transcodes and the availability checks now included) refuses to start, including one that was already queued when the move began. The rack shows a **media held** chip in the channel's Tier cell and the Storage panel says "Held: its media is moving"; both go when the move finishes or its marker is cleared. The copy is now followed by a pass that makes the destination copy match the source — files the source no longer has are removed from the copy, never from the source — so a file written or deleted during the copy (a transcriber's scratch folder, say) no longer fails the check, and **Resume move** finishes a move whose copy holds such leftovers. Every file removed from a copy is listed in the move's log, and **Preview** says so when a copy from an earlier attempt is already there. If the source keeps changing, the move stops and lists what differs: extra on the destination, missing there, or changed. A new **Reconcile and resume** button beside **Resume move** lists those differences, makes the copy match and finishes the move, so no file has to be deleted by hand. The saved-video store's move does the same matching and the same check before it starts. Needs a rebuild and restart of the editor. - **Connecting an X account opens your own browser, and the X fetchers can use your everyday browser's X login instead.** **Settings → X account session → Connect X account** used to open Playwright's bundled Chromium with its automation signals on (the "controlled by automated test software" bar, `navigator.webdriver`): Google's sign-in refused it and X's own login form stalled in it. It now opens your Chromium or Chrome when one is installed (`ARCHILYZER_X_BROWSER` names another; Playwright's bundled Chromium otherwise), without those signals. Google's sign-in may still refuse an embedded browser; X's password login is the reliable path. A new **Login source** choice (`social.x.cookieSource` in `settings.json`) says where the X fetchers' login comes from: **Browser login** hands gallery-dl `--cookies-from-browser` with your `cookiesFromBrowser` on every fetch, so the login lasts as long as you stay logged in to x.com in that browser and no window is needed; **Connected profile** is the session broker, as before. Left on **Automatic**, it is the browser login when `cookiesFromBrowser` is set and no profile is connected, and the profile otherwise. **Check** says which source is in use, whether an X login is visible in it and when it was last used (the browser's cookies are read from a private copy, never written; this reads Firefox's, and gallery-dl reads Chromium's itself). Needs a rebuild and restart of the editor. -- **The dashboard and `/jobs` keep answering while a channel's report is regenerated.** Regenerating a report walks every video of the channel inside the editor, and two regenerations of channels with a few thousand videos, running side by side, kept `/`, `/channels` and `/jobs` from loading for over an hour. Regenerations now run one at a time, on their own `refresh-report` queue on `/jobs`; a channel whose report is already waiting or regenerating is not queued a second time, whether the request came from a finished job or from **Update all reports**; and the walk pauses between batches of videos so pages are served in between. **Update all reports** therefore takes as long as all the channels' regenerations added up, not the longest one. Needs a rebuild and restart of the editor. -- **The operations pages share one count of the lanes' pending work.** Every open operations page asks for the lanes' status every 3 seconds, and each request used to count every lane's pending videos afresh from every channel's report. That count is now made once and handed to every request in the next 3 seconds, so a pending count can be up to 3 seconds old. A lane's hold, its runner and its picks are still read fresh on every request. +- **The dashboard and `/jobs` keep answering while a channel's report is regenerated.** Regenerating a report walks every video of the channel inside the editor, and two regenerations of channels with a few thousand videos, running side by side, kept `/`, `/channels` and `/jobs` from loading for over an hour. Regenerations now run one at a time, on their own `refresh-report` queue on `/jobs` — a channel's own **Refresh report** included, which now waits its turn there too; a channel whose report is already waiting is not queued a second time, whether the request came from a finished job, **Refresh report** or **Update all reports**, and a change made while a channel's report is being regenerated queues one more regeneration after it rather than being missed; and the walk pauses between batches of videos so pages are served in between. **Update all reports** therefore takes as long as all the channels' regenerations added up, not the longest one. Needs a rebuild and restart of the editor. +- **The operations pages share one count of the lanes' pending work.** Every open operations page asks for the lanes' status every 3 seconds, and each request used to count every lane's pending videos afresh from every channel's report. That count is now made once and handed to every request in the next 3 seconds. Changing a lane's rules, a focus or a channel's priority counts again at once; otherwise a pending count can be up to 3 seconds behind a report that was just rewritten or a video a lane just picked. A lane's hold, its runner and its picks are still read fresh on every request. - **Jobs a stopped editor left "running" are closed when it starts again.** A job that was still running when the editor's process ended (killed, crashed, or shut down before the job had finished unwinding) kept "running" in its record for good, and `/jobs` listed it as archived. On start the editor now marks each one **cancelled**, with "interrupted: the process running it stopped before it finished" as the reason on the job's page, and its end time is the last time its log was written. Nothing is run again; **Retry** works as for any cancelled job. A job that another live process is running, such as `archilyzer run`, is left alone, and the same check now keeps the start-up pass from closing that process's queued jobs. Such leftover jobs never blocked a media move. ## [0.11.0] - 2026-09-30 diff --git a/plans/release-17.md b/plans/release-17.md @@ -419,7 +419,7 @@ profile). | `36babcad` | `common, editor:` meta `pid`; `writerIsGone`, `processIsAlive`, `settleRunningJobMetas`; the queued pass's writer check; boot wiring; 4 boot tests + 1 `channelWriters` test | | `fef9ebc2` | `editor(e2e):` `dashboard-answers.spec.ts`; `/api/test/settle-running-metas` | | `fe1676c3` | `editor(e2e):` the spec reworked after run 1 (two 600-video channels, the serial check, the idle-floored budget) | -| this commit | `plans:` this section; the editor changelog | +| `3c90b8e9` | `plans:` this section; the editor changelog | #### Gates (logs `$T/D0-*.log`) @@ -475,4 +475,52 @@ while reports regenerate (one at a time, deduped; Update all reports takes the s share one pending count (up to 3 s old; holds, runners and picks fresh), and jobs a stopped editor left "running" are closed at start. +**Rollout note.** No `archilyzer run` may be in flight across the first editor restart after this ships: +every meta written before it has no `pid`, so the boot pass reads its writer as gone and would close a +still-running offline job's meta as interrupted (the live job rewrites it at its next meta write, but +`/jobs` offers Retry meanwhile). The same holds across pid namespaces — an editor in a container judging +a pid a host-side `archilyzer run` wrote, or the reverse, sees ESRCH. The release's rollout stops the +editor for the migration anyway. + +#### Review (SHIP AFTER FIXES) and the fixes + +| Finding | Fix | +|---|---| +| H1 — the memo served pending counts folded from settings (lane policy and tree, `channelPriority`) up to 3 s stale against a fresh tree; the payload-reading specs were not run | `5b236f8c`: `singleFlightMemo.get(key, compute)`; the shell keys it on `JSON.stringify([settings.autoQueue, settings.channelPriority])`, time the only other expiry; only the latest-started computation stores (an old key landing late never overwrites); comments say what can be one poll late (a rewritten snapshot, the runner's in-flight set) and what is fresh; +1 test. The status-reading specs ran in the full suite below | +| L2 — a change during a RUNNING regeneration was dropped until the next change (as on main) | `616a9ce2`: dedup only against a queued or starting regeneration (`isRefreshReportPending`), so a running walk gets one queued successor and no second; fire()'s trailing-edge comment is now true; +1 test | +| L3 — the per-channel Refresh report walked in the request, outside the queue | `5509ccf4`: `refreshChannelSnapshotAction` (Refresh report, ops `{slug}`, e2e `generateReport`) starts a job through `startRefreshReport` and awaits it — or polls the one already queued for the channel — and returns the walk's error sentence (`onError`) as its `{error}`; the ops route's comment follows | +| L4 — a queued refresh-report holds the Storage panel's Move (`mediaBusy.ts` counts queued jobs) | Not changed in D0, as ruled: T1's `mediaOnly` (refresh-report does not need media) clears it when T2's courtesy check passes it; to confirm at the T1/T2 merges | +| L3 follow-on — `generateReport` (85 specs) now queues a job, which `/jobs` may draw before its default filter hides it | `7ba8cd41`: `bulk-actions.spec` "queues no job" leaves `refresh-report` rows out; `dashboard-answers` compiles the ops route (a 400 on `{}`) before measuring — its first compile held `/` 15 s once under load | +| L5 — nothing pinned the yield; the e2e floor could grow without bound | `fea5b203`: `snapshotYield.test.ts` — a `setImmediate` probe queued during the first chunk runs before the second chunk's first unit (it fails with the yield removed, checked); 2 tests. The e2e budget is `min(max(5 s, 3 × slowest idle), 15 s)`, and the spec header says it pins the serial queue and gross starvation, not the yield | +| L6 — the first boot after this ships closes a live pre-slice `archilyzer run`'s meta | The rollout note above | +| N7 — `REFRESH_REPORT_QUEUE` beside the other queue keys | `616a9ce2`: defined in `lib/queueKeys.ts`, re-exported by the scheduler | +| N8 — `operations/[id]/page.tsx`'s census comment claimed the payload's listing | `5b236f8c`: says it is the same listing only on a memo miss | +| N9 — `requestCache.ts` said "no exception" | `5b236f8c`: one paragraph naming the memo and its bound | +| N10 — `channelWriters.test.ts` conflicts with T1 (both append) | Keep both, at the merge | +| the record and changelog | this commit: this subsection, the rollout note; the changelog's first two bullets say Refresh report waits its turn, a change during a regeneration queues one more, and a rules/focus/priority edit recounts at once | + +**Gates after the fixes.** tsc clean at every commit. **common 2,500/2,500** (+4: memo key 1, the +successor 1, the yield 2). **Editor unit 109/109.** **e2e** — L3 makes every `generateReport` (85 specs) a +queued job, so the whole editor suite rather than the two lists (it contains both), in three runs on a +machine that ran out of memory (15 GB used, swap 19/19 GB; a parakeet transcription of the live editor +beside several suites): + - run 3, the full suite (702 tests) at `fea5b203`: stopped at 374 passed, 8 failed, ~50 min, when + pages began to crash (`page.goto: Page crashed`). Two failures were D0's and are fixed in `7ba8cd41` + (`bulk-actions` "queues no job"; `dashboard-answers` (a): one `/` of 14.9 s at the moment the ops + route compiled, every other answer ≤ 4.6 s). The other six passed in run 4. + - run 4 at `7ba8cd41` (run 3's failures plus every spec file from `new-channel-onboarding` on, and both + lists — 69 files): **295 passed, 68 failed, 35.1 min** (54 with the queue). Every spec in both lists + passed — `jobs` 2, `channels` 8, `channel-storage` 13, `dashboard-answers` 2, `focus-banner` 3, + `auto-queue` 24, `lane-runner` 5, `channel-priority` 9, `operation-settings` 7, `backfill` 20, + `channel-work` 11, `ops-api` 23 — except `view-route`, which ran after a machine-wide OOM kill at + 22:41 took the test server (the kernel log names it, among browser tabs, the desktop session and the + live editor's transcriber): every test from the 300th on failed in under 2 s against a dead server. + Before it, four failures in `site-scope` (2), `sites-crud` and `social-channel` (the known 22–26 s + fetch-posts case). + - run 5 at `7ba8cd41`, the 25 spec files that failed in run 4 (`view-route` included): **183 passed, 0 + failed, 12.2 min** (13 with the queue). + So every editor spec has passed with the fixes in, across runs 3–5, and both lists in full. + +**test:scripts** was not re-run: the fixes touch nothing under `scripts/`. + ## Rollout