commit 45ad131f4f6b43c8316ca96cba428b3887c9c1d8
parent d66e39798d1b5a36c1f98cb90d4af88de54c3345
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Thu, 1 Oct 2026 20:44:15 -0400
plans: slice D0's record — the dashboard answers while a snapshot regenerates, as shipped; the editor changelog
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Diffstat:
2 files changed, 155 insertions(+), 0 deletions(-)
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -16,6 +16,9 @@
- **A form whose save is refused keeps what you typed.** Every editor form put its plain fields back to the stored values when its save was refused — a site's ID rejected, a page size out of range, a slug already taken — so everything typed had to be typed again. A refused save now leaves every field as you left it, beside the reason: **Settings**; a site's form (new and existing); the hub's config on `/sites`; **Cut release**; a channel's form (new and **Configure**), **Rename** and **Delete**; a video's **Delete directory**; **Drive health timing** on `/storage`; the backup config on `/saved-videos`; the sync operation's controls; the **Digest**, **Diarization**, **Speaker attribution** and **Speaker work lane** settings; and the worker list on `/workers`. A save that succeeds behaves as before, with one difference you may notice: a drop-down, and a checkbox or choice that the page tracks as you change it (a cadence, a worker's **Enabled**, a social link's **Keep in header**, a site membership, a site's accent), now shows what was saved. A form's own drop-downs used to go back to what the page had loaded with until a reload, and a second save from the same page sent that old choice again; the others went back until the page next refreshed itself (every 5 seconds by default).
- **A media move no longer starts over a job that is writing into the channel, holds the channel's writers while it runs, and makes its copy match the source before it verifies — so a transcription or a download during a move cannot fail it.** A move that has waited its turn behind other moves now checks again when it starts: if a job is running on the channel, or an auto-queue lane is working on one of its videos, it stops at once and says which ("a transcription of abc123 is running (Transcribe all, job …) — wait for it or cancel it"), with nothing copied — and a job you have just cancelled counts until it has actually stopped ("is stopping … — wait for it to stop"); **Preview** says the same, and the Storage panel's blocked message now names the job too. While a move's marker stands, the channel is held: every lane skips it, and every job that reads or writes its media (single-video transcriptions, downloads and transcodes and the availability checks now included) refuses to start, including one that was already queued when the move began. The rack shows a **media held** chip in the channel's Tier cell and the Storage panel says "Held: its media is moving"; both go when the move finishes or its marker is cleared. The copy is now followed by a pass that makes the destination copy match the source — files the source no longer has are removed from the copy, never from the source — so a file written or deleted during the copy (a transcriber's scratch folder, say) no longer fails the check, and **Resume move** finishes a move whose copy holds such leftovers. Every file removed from a copy is listed in the move's log, and **Preview** says so when a copy from an earlier attempt is already there. If the source keeps changing, the move stops and lists what differs: extra on the destination, missing there, or changed. A new **Reconcile and resume** button beside **Resume move** lists those differences, makes the copy match and finishes the move, so no file has to be deleted by hand. The saved-video store's move does the same matching and the same check before it starts. Needs a rebuild and restart of the editor.
- **Connecting an X account opens your own browser, and the X fetchers can use your everyday browser's X login instead.** **Settings → X account session → Connect X account** used to open Playwright's bundled Chromium with its automation signals on (the "controlled by automated test software" bar, `navigator.webdriver`): Google's sign-in refused it and X's own login form stalled in it. It now opens your Chromium or Chrome when one is installed (`ARCHILYZER_X_BROWSER` names another; Playwright's bundled Chromium otherwise), without those signals. Google's sign-in may still refuse an embedded browser; X's password login is the reliable path. A new **Login source** choice (`social.x.cookieSource` in `settings.json`) says where the X fetchers' login comes from: **Browser login** hands gallery-dl `--cookies-from-browser` with your `cookiesFromBrowser` on every fetch, so the login lasts as long as you stay logged in to x.com in that browser and no window is needed; **Connected profile** is the session broker, as before. Left on **Automatic**, it is the browser login when `cookiesFromBrowser` is set and no profile is connected, and the profile otherwise. **Check** says which source is in use, whether an X login is visible in it and when it was last used (the browser's cookies are read from a private copy, never written; this reads Firefox's, and gallery-dl reads Chromium's itself). Needs a rebuild and restart of the editor.
+- **The dashboard and `/jobs` keep answering while a channel's report is regenerated.** Regenerating a report walks every video of the channel inside the editor, and two regenerations of channels with a few thousand videos, running side by side, kept `/`, `/channels` and `/jobs` from loading for over an hour. Regenerations now run one at a time, on their own `refresh-report` queue on `/jobs`; a channel whose report is already waiting or regenerating is not queued a second time, whether the request came from a finished job or from **Update all reports**; and the walk pauses between batches of videos so pages are served in between. **Update all reports** therefore takes as long as all the channels' regenerations added up, not the longest one. Needs a rebuild and restart of the editor.
+- **The operations pages share one count of the lanes' pending work.** Every open operations page asks for the lanes' status every 3 seconds, and each request used to count every lane's pending videos afresh from every channel's report. That count is now made once and handed to every request in the next 3 seconds, so a pending count can be up to 3 seconds old. A lane's hold, its runner and its picks are still read fresh on every request.
+- **Jobs a stopped editor left "running" are closed when it starts again.** A job that was still running when the editor's process ended (killed, crashed, or shut down before the job had finished unwinding) kept "running" in its record for good, and `/jobs` listed it as archived. On start the editor now marks each one **cancelled**, with "interrupted: the process running it stopped before it finished" as the reason on the job's page, and its end time is the last time its log was written. Nothing is run again; **Retry** works as for any cancelled job. A job that another live process is running, such as `archilyzer run`, is left alone, and the same check now keeps the start-up pass from closing that process's queued jobs. Such leftover jobs never blocked a media move.
## [0.11.0] - 2026-09-30
- **Transcripts that arrived after a video was first seen are counted.** The stats behind the homepage, the hub and every site's charts were cached per video and refreshed only when the video's metadata changed, so a transcript that came later — a Whisper run days after the download, or a video downloaded after the last index build — never reached them, and a video with YouTube captions alone had no transcription date. Counts and charts were low; the homepage could show a site with 0 transcripts, 0 channels and 0 hours while it served its videos. A stat is now also redone whenever the index re-reads the video, every transcript has a date, and a captioned video is dated by when its captions arrived rather than by a later Normalize run, so its place on "Transcribed over time" can move. **After updating, rebuild and restart the editor before anything else:** until then, **Build stats dataset** runs the old code and would undo the new stats, while a site, hub or homepage build already runs the new code — and the first stats build of any kind re-reads every video once (about 10–30 minutes on a large archive; it can be stopped and picks up where it stopped). Then build the index, the stats, the homepage, the hub, and the sites.
diff --git a/plans/release-17.md b/plans/release-17.md
@@ -323,4 +323,156 @@ hand; a dirent `isFile()` filter over a video dir hides it."** The `.relocating.
## Record
+### Slice D0, as shipped — the dashboard answers while a snapshot regenerates (2026-10-01)
+
+Branch `r17/dashboard-answers` off `main` `7f4901f1`, worktree `~/Projects/r13-lows-editor` (editor 5401,
+test 5411, export 5410 — `pnpm wt list`'s block #24), one Opus implementer. Scratch files `D0-*` in the
+job's `tmp`. The plan is Step 0b above.
+
+**What was found before building** (the corpus only read; one video dir's text copied to scratch for the
+profile).
+- **The walk did yield — on every file read.** Each video's unit awaits its reads, so the loop turned
+ between them; what the walk does ON the loop is parse. A CPU profile of `generateChannelSnapshot` over
+ 300 copies of one omnibased video dir's text (metadata 0.62 MB, cues 0.62 MB, two 2.9 MB VTTs):
+ `readNormalizedTranscript` (the cues parse, `readTranscriptCoverage`) 1,037 ms self, `readWebpageUrl`
+ 791 ms self, everything else in the walk under 70 ms. `readWebpageUrl` is the reconcile pass at the
+ walk's start (`reconcileVideoDirs.ts`): it reads and parses EVERY `metadata.info.json` for its
+ `webpage_url`, one at a time, on every regeneration.
+- **One walk alone does not starve the pages.** Built and served by `next start` (loopback, a scratch
+ corpus of one 2,000-video channel, load average 24), with this slice: the regeneration took 26 s, and
+ `/` answered in 0.28–1.6 s and `/jobs` in 0.18–0.81 s, polled every 2 s throughout. Under `next dev`
+ at a load average of 32–35 the same walk took 118 s and the answers ran 0.3–28 s; with `main`'s five
+ files swapped back, 2.0–12.5 s. On this machine, under the dev server, one regeneration does not
+ separate the two, and an isolated tsx micro-benchmark (800 such dirs, three interleaved pairs, load
+ 25–35) put the chunked walk and `main`'s inside each other's noise (loop delay max 300–860 ms either
+ way).
+- **So the hour-long outage needed more than one walk**, and the live metas show the rest: two walks side
+ by side (the empty queue key), the omnibased one on the platter through `onDrive` (four slots per
+ location, so every page read of that drive queued behind the walk's units), two regenerations of
+ omnibased four seconds apart in the restarted process (`01M3WHY8…`, `01M3WHYC…` — two passes racing
+ between the registry check and the record's registration), and the auto-queue status poll recomputing
+ behind all of it. D0 closes the three it owns: one queue, a per-slug dedup that covers the enqueue in
+ flight, and the status memo. The platter half is T1's (text reads leave `onDrive`).
+- **The ghosts never held a move.** `channelWriters` reads this process's registry only, never a meta,
+ and the registry is in memory: a meta a dead process left `running` was never a writer. `/jobs`
+ listed such a job as `archived` (`listJobs`: a non-terminal meta not in the registry). Harmless to a
+ move, misleading on `/jobs`; the boot pass now closes them.
+- **What a SIGTERM'd process leaves behind today** (`shutdownCancel.ts`, "graceful shutdown only"): the
+ reaper cancels every live job and re-raises the signal at once, without waiting. A queued job keeps
+ `queued` on disk on purpose (the boot pass settles it). A running managed job's terminal meta is
+ written only when its function returns (`streamCommand.ts`, the `.finally`); `generateChannelSnapshot`
+ takes no signal, so a regeneration never returns before the exit, and its meta stays `running`. A
+ SIGKILL leaves the same, and orphans children.
+
+**What it does.**
+- **The walk yields between chunks.** `generateChannelSnapshot` maps its video dirs through
+ `mapInYieldingChunks` — chunks of `SNAPSHOT_YIELD_EVERY` (32: two full waves of the 16-wide limit), one
+ `setImmediate` between chunks. The unit body is untouched: the diff in that file is the helper, the
+ constant and the call's two lines, so T1's merge is trivial.
+- **One regeneration at a time, once per channel.** `snapshotScheduler.ts` gains `REFRESH_REPORT_QUEUE`
+ (`"refresh-report"`: a non-empty key is the registry's existing concurrency-1 serialization, so no new
+ mechanism) and `startRefreshReport(paths, slug)`, the one entry point: a slug with a `refresh-report`
+ queued, running or being enqueued (a `starting` set on the scheduler's global state, held across
+ `runManagedFunction`'s awaits) is answered `info` with `REFRESH_REPORT_ACTIVE`. The debounced pass and
+ **Update all reports** (`refreshAllChannelSnapshotsAction`, `editor/app/channels/actions.ts` — not in
+ the slice's file list and owned by no other slice) both go through it, so they dedup against each
+ other; the action still reports a skip as "already running". The job body is the scheduler's (it bumps
+ `generation`, which the action's copy did not).
+- **The auto-queue status poll shares one fold.** `common/views/autoQueueStatus.ts` gains
+ `singleFlightMemo` (concurrent callers share the computation in flight; a landed value is reused for
+ `AUTO_QUEUE_STATUS_MEMO_MS` = 3 s from when it LANDED; a rejection is not memoized; `clear()` detaches
+ an in-flight computation) and `autoQueueStatusMemo()` on `globalThis`. `buildAutoQueueStatusPayload`
+ is unchanged. The shell (`editor/app/operations/status.ts`) memoizes the snapshot-derived half only —
+ the channel briefs and the four lanes' `computeLeafPending` — and reads the priority view, the state
+ document, settings, the pool and the runners fresh. The e2e reset route drops the memo with the other
+ singletons.
+- **Boot closes the running metas a dead process left.** A meta now records `pid` (`jobMeta.ts`, owned by
+ no slice). `bootQueuedJobs.ts` gains `settleRunningJobMetas`, run from `instrumentation.ts` on every
+ boot (idle and test server too: it re-queues nothing, and it does not wait for the storage pass): a
+ `running` meta from before this boot, not in the registry, whose writer is gone is closed `cancelled`
+ with `cancelReason` "interrupted: the process running it stopped before it finished", `endedAt` = its
+ log's mtime (the last moment it is known to have run), else now. **The status:** `cancelled` +
+ `cancelReason` already is this file's terminal state for "the server went down under it", and every
+ reader (listJobs, `/jobs`, the job page's "Cancelled because", Retry) handles it — so there is no new
+ `interrupted` status. **The dead-process test** (`writerIsGone`): metas carried no pid or host, and
+ "predates this boot" alone is not reliable here, because `archilyzer run` (`bin/run-operation.ts`) runs
+ jobs offline into the same `.jobs/`. Gone = no pid (written before this release), or this process's
+ own pid (a container's editor restarts as the same pid; this process's own metas are already excluded
+ by `bootedAt` and `isLive`), or `kill(pid, 0)` → ESRCH (EPERM counts as alive). A live pid is left
+ alone. The queued pass now applies the same check. `channelWriters.ts` gains a comment saying why a
+ ghost cannot reach it, and a test pins it.
+- **e2e** `dashboard-answers.spec.ts` (two cases) and `/api/test/settle-running-metas` (POST, guarded by
+ `E2E_TEST_ROUTES`), which runs the running pass on demand: the e2e server boots once per suite.
+ (a) Two 600-video channels (one hardlinked ~400 KB `metadata.info.json` per dir, a Rumble archive so
+ each file is parsed twice) regenerated by Update all reports through the ops API while `/` and `/jobs`
+ are polled every 2 s: both reports land, the second job started when the first had ended (read from
+ their records), and every answer is inside the budget. (b) A `running` refresh-report meta with a dead
+ pid is closed as interrupted, one with a live pid (the test runner's) is left `running`, the job page
+ says "interrupted", and the channel's Move media completes.
+
+**Commits**
+
+| Commit | What |
+|---|---|
+| `8ce3b24d` | `common:` the chunked walk; `REFRESH_REPORT_QUEUE`, `startRefreshReport`, the `starting` set; `snapshotScheduler.test.ts` (2); Update all reports through it |
+| `56efcd74` | `common, editor:` `singleFlightMemo` + `autoQueueStatusMemo` + 5 tests; the shell memoizes the snapshot half; the reset route drops it |
+| `36babcad` | `common, editor:` meta `pid`; `writerIsGone`, `processIsAlive`, `settleRunningJobMetas`; the queued pass's writer check; boot wiring; 4 boot tests + 1 `channelWriters` test |
+| `fef9ebc2` | `editor(e2e):` `dashboard-answers.spec.ts`; `/api/test/settle-running-metas` |
+| `fe1676c3` | `editor(e2e):` the spec reworked after run 1 (two 600-video channels, the serial check, the idle-floored budget) |
+| this commit | `plans:` this section; the editor changelog |
+
+#### Gates (logs `$T/D0-*.log`)
+
+- **tsc** (all workspaces) clean at every commit.
+- **common:** **2,496/2,496** — new: `snapshotScheduler.test.ts` 2, `autoQueueStatus.test.ts` +5,
+ `bootQueuedJobs.test.ts` +4, `channelWriters.test.ts` +1. **Editor unit:** 109/109.
+ **test:scripts:** run 1: 301 passed, 1 failed, 2 skipped (304); rerun: 300 passed, 2 failed, 2
+ skipped. Every failure is in `queue-lock.test.mjs` ("prints a banner naming the holder while waiting",
+ "serves waiters in arrival order (FIFO)"), timing cases run at load averages of 16–30; the file run
+ alone passes 11/11, and the slice touches nothing under `scripts/`.
+- **Build:** the capped editor build with the corpus linked (`ln -sT`, `systemd-run --scope -p
+ MemoryMax=6G`, the link removed after): exit 0, 123 s, at `fef9ebc2`.
+- **e2e** (`jobs`, `channels`, `channel-storage`, `dashboard-answers`; `next dev`):
+ - run 1 at `fef9ebc2`: **23 passed, 2 failed, 16.7 min**. `dashboard-answers` (a) timed out at 5 min:
+ 2,000 videos under `next dev` at a load average of ~30 did not finish (the scratch-server measurement
+ above: 118 s for the walk alone, every answer slower). `channel-storage` "relocate a channel's media
+ to another root, and move it back" — the run's first test, on a cold dev server — timed out at 90 s
+ on its first request to the file route.
+ - run 2 at `fe1676c3`: **25 passed, 0 failed, 3.7 min** (34 min with the queue wait). (a) took 23.8 s:
+ nine polls of each page while the two regenerations ran, `/` 0.14–1.9 s and `/jobs` 0.34–1.6 s
+ (budget 5 s); (b) 7.8 s. The relocate case passed — run 1's failure was the cold first test.
+- **Numbers tool:** none.
+
+**Deviations from the plan, one sentence each.**
+- The chunk is 32, not 25, so every wave of the 16-wide limit is full (25 left each chunk's second wave
+ nine wide).
+- The memo covers the shell's snapshot-derived half, not the whole payload: memoizing it all would hold a
+ lane's hold, runner and picks up to 3 s behind a click, and specs read them straight after one.
+- The running-meta pass is in `bootQueuedJobs.ts` beside the queued one, not in `registry.ts`, which is
+ in memory and holds no metas.
+- "Interrupted" is `cancelled` + `cancelReason` (the existing terminal state for a job the server went
+ down under), not a new status.
+- `channelWriters.ts` gets a comment and a test, no code: a ghost cannot reach it.
+- The e2e budget is 5 s or three times the slowest idle answer measured just before, whichever is longer:
+ the dev server on this shared machine at load 30 takes 2–4 s a page with nothing regenerating, and the
+ spec pins starvation, not the machine's speed.
+- The spec regenerates two 600-video channels, not one large one: 2,000 did not finish in five minutes
+ under the dev server here, and two prove the serial queue from the jobs' records.
+
+**Found and left.**
+- The reconcile pass parses every `metadata.info.json` in full, on every regeneration, for one field — the
+ largest single cost of a walk. Reading only the head, or skipping a dir whose name is already canonical,
+ is a follow-up (`reconcileVideoDirs.ts` is not D0's).
+- `generateChannelSnapshot` takes no abort signal, so a Cancel or a graceful shutdown cannot stop a walk;
+ the chunk boundary is now the natural place to check one — a follow-up.
+- A meta whose pid has since been reused by an unrelated live process stays `running` on `/jobs` until
+ that process exits; a start-time check (`/proc/<pid>/stat`) would close it. Not worth it today.
+- `editor/app/jobs/[id]/page.tsx` and `JobRow.tsx` describe the cancel reason as the one "a restart left
+ queued"; it now also covers a dead process's running job. Comments only, left.
+
+**Changelog** (`editor/CHANGELOG.md` `[Unreleased]`): three bullets — the dashboard and `/jobs` answer
+while reports regenerate (one at a time, deduped; Update all reports takes the sum), the operations pages
+share one pending count (up to 3 s old; holds, runners and picks fresh), and jobs a stopped editor left
+"running" are closed at start.
+
## Rollout