Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 93be9ca0422fd29d023614b1d88529127fd5c1e7
parent d69f5709119eff379ff2e676fa8d0eb61c18cd44
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Wed, 30 Sep 2026 00:23:57 -0400

plans: slice DS after the review — each finding to its commit (M1–M3, L1–L10) in a Review table; the mechanism as built (the queue's deadline and overdue count, slot keys, slow-not-stalled, the health pass above the idle gate, the snapshot walk through onDrive); the stated limit said as it is; the hand-typed-root limit and the spin-down question; re-gates on the merged tree (common 2,295, editor unit 87, test:scripts 194+2, docs clean, editor build 63 s / 1,641,444 KB, e2e 73/73); FACTS "The storage health gate" rewritten; the changelog

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 2+-
Mplans/FACTS.md | 97++++++++++++++++++++++++++++++++++++++++++++++---------------------------------
Mplans/release-15.md | 217+++++++++++++++++++++++++++++++++++++++++++++++++++++--------------------------
3 files changed, 203 insertions(+), 113 deletions(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -5,7 +5,7 @@ - **A stats build keeps the stats of a channel whose drive is not mounted, and will not undo a newer version's stats.** A channel whose media is on a drive that is not mounted (or is being moved) is left as it was instead of being read as a channel with no videos; a stats rebuild that has to start over refuses until the drive is back. A stats build refuses to clear stats written by a newer version of the editor; set `ARCHILYZER_STATS_ALLOW_DOWNGRADE=1` to roll back on purpose. Its log also says apart how many videos were downloaded since the last index build (they catch up after the next one) and how many the index skipped (no upload date, or it failed on them). - **An index build keeps a channel whose drive is not mounted, instead of dropping it from the sites.** **Build index**, a site build's data phase and `archilyzer index` read a channel whose media is on a drive that is not mounted (or is being moved, or whose link and config disagree) as a channel with no videos: they removed its videos from the index, and the next site build published the channel as gone. Such a channel is now left as the last build had it — its videos stay in the index, its pages stay as they were, and the sites built next still list it — and the log names it, with its storage location: one line per channel, ` Held: N channel(s), K video(s) kept.` at the end of the `Diff:` line, and the channels again on the last line. A data folder that fails to read is held the same way, and a channel with no data folder at all is said in the log instead of passed over. An index rebuild that has to start over (after an update that changes the index's format, or with no index yet) refuses while any channel is held and says which; mount the drive first, or set `ARCHILYZER_INDEX_ALLOW_HELD=1` to rebuild without that channel until its drive is back and the index is built again — on the command for a command-line build (`ARCHILYZER_INDEX_ALLOW_HELD=1 pnpm archilyzer index`), or in the editor's own environment, with a restart, for **Build index** and the site builds started from the editor. - **umtool's build no longer lists its e2e test data, the e2e server's build folder or `.env.local` among a route's files.** The clip-audio route named its cache files in a way the bundler read as a pattern reaching into umtool's hidden folders, so its list of files took in the e2e fixture (where the tests link the song data), the e2e dev server's build folder and the env file: 1,704 of its 2,167 entries. It now lists what the other routes list (463). Those folders and env files are also excluded from every route's list, and `pnpm test:scripts` reads the last umtool build's lists back and fails on any such entry. A checkout whose umtool build predates its code (this change included) skips that check, saying so, until umtool is rebuilt (`pnpm --filter umtool exec next build`). Nothing changes when umtool runs. -- **A drive that stops answering no longer stops the editor answering.** When a storage location's drive is mounted but not answering (an SMR disk in a USB enclosure resetting under a long write), every page and poll that touched it waited on it, and a few such waits froze the whole editor until the drive came back. Every 15 seconds the editor now reads each location's disk activity counters from the kernel, which never waits on the drive: a disk with requests waiting and none finished since the last look is marked **Not answering**, and the mark comes off after two looks in a row find it working. Where no disk can be named (in a container, say) it asks the drive from a separate process with a 3-second limit instead. Any page or poll that reads the drive also gives up after 3 seconds and marks it the same way, and no more than four such reads wait on one drive at a time. While it is marked, the editor's pages and polls do not read that drive: `/storage` shows the location as **Not answering** with the time it stopped and how it is watched (**Refresh** asks again at once), the `/channels` volume chip reads "not answering since HH:MM" and its channels' badges "not answering", their videos list, video pages and Cleanup stage say so instead of reading the drive, `/saved-videos` names the channels it did not read, and the index and stats builds keep those channels as they do for an unmounted drive, and a channel that is in the middle of a move still shows as moving. Jobs for those channels are refused until the drive answers. Pages and polls also reuse each channel's media check for 5 seconds. The editor's `start` script and the container now give Node 16 threads for file access instead of 4 (`UV_THREADPOOL_SIZE`); that buys time for reads already waiting on a drive, and a read that was already waiting when the drive stalled still waits until the drive answers. +- **A drive that stops answering no longer stops the editor answering.** When a storage location's drive is mounted but not answering (an SMR disk in a USB enclosure resetting under a long write), every page and poll that touched it waited on it, and a few such waits froze the whole editor until the drive came back. Every 15 seconds the editor now reads each location's disk activity counters from the kernel, which never waits on the drive: a disk with requests waiting and none finished since the last look is marked **Not answering**, and the mark comes off after two looks in a row find it working. Where no disk can be named (in a container, say) it asks the drive from a separate process with a 3-second limit instead. Any page or poll that reads the drive also gives up after 3 seconds and marks it the same way, and no more than four such reads wait on one drive at a time. While it is marked, the editor's pages and polls do not read that drive: `/storage` shows the location as **Not answering** with the time it stopped and how it is watched (**Refresh** asks the drive again), the `/channels` volume chip reads "not answering since HH:MM" and its channels' badges "not answering", their videos list, video pages and Cleanup stage say so instead of reading the drive, `/saved-videos` names the channels it did not read, and the index and stats builds keep those channels as they do for an unmounted drive, and a channel that is in the middle of a move still shows as moving. Jobs for those channels are refused until the drive answers, and a channel paused automatically for it says the drive is not answering rather than not there. The drive check keeps running on an editor started with `ARCHILYZER_IDLE_BOOT`. Pages and polls also reuse each channel's media check for 5 seconds. The editor's `start` script and the container now give Node 16 threads for file access instead of 4 (`UV_THREADPOOL_SIZE`); that buys time for reads already waiting on a drive, and a read that was already waiting when the drive stalled still waits until the drive answers. - **Building the homepage now publishes the source: a read-only git mirror, its raw tree and a fresh tarball, behind a gate.** `archilyzer build homepage`, the `/sites` Homepage jobs and `pnpm ops build-homepage` run `archilyzer source publish` between compose and `next build`. It makes a fresh clone of the private `main` (the repository itself is never rewritten), rewrites that copy with git-filter-repo using your scrub rules (file contents and commit messages; your home directory becomes `/home/user` without a rule), and publishes it under `homepage/public` for `git clone https://archilyzer.pages.dev/source/archilyzer.git`, beside `/source/tree/` and the Downloads tarball. Before anything is written, every object of the rewritten history and every file about to be published is searched for every string you have denied; **one hit refuses the build**, and its log names the string only by where you wrote it (`denylist line 3 (len 5)`) and each hit by its object, field and byte offset — never a byte of the object. **A refusal withdraws the source**: the last publish is removed from `homepage/public` and the last build's copy from `homepage/out`, and **Deploy homepage refuses** a build whose source was not audited under today's rules and today's `main` ("run `archilyzer build homepage`, then deploy"). The rules live outside the repo, in `~/.config/archilyzer/source-scrub.txt` and `source-denylist.txt` (`ARCHILYZER_CONFIG_DIR`, `SOURCE_SCRUB_FILE`, `SOURCE_DENYLIST_FILE`); **without them the build refuses**, naming the missing file. **Put everything private in the denylist before any deploy, a preview included**: previews are public, and every deployment stays reachable at its own address until you delete it. Install git-filter-repo once (`pipx install git-filter-repo`; the editor's process needs `~/.local/bin` on its `PATH` to find it) — without it the build fetches it through `pipx run`, which needs the network — and gitleaks if you want its secret scan too. An unchanged `main` with unchanged rules is skipped, so a rebuild costs about 20 seconds only when something moved. A checkout with no git repository (the docker image, a tarball install) builds with the /source page's empty state. `archilyzer source publish --check` audits without writing, `archilyzer source audit <clone>/.git` checks any clone, `archilyzer build homepage --no-source` removes the published source instead, and `archilyzer doctor` reports the tools, the two files (rule counts and permissions, never their contents) and the last publish. `create-archives.sh` is gone. See PUBLISH.md, "The source mirror (homepage)". - **umtool reads the corpus from its checkout (or `TRANSCRIPTS_DIR`), and the song project's data defaults to `~/.local/share/archilyzer/song`.** If yours is elsewhere, link it there before restarting umtool: `mkdir -p ~/.local/share/archilyzer && ln -s <where the data is> ~/.local/share/archilyzer/song` (the data stays where it is). With no `CHANNELS_DIR`, umtool reads the corpus at `$TRANSCRIPTS_DIR/channels`, else the checkout's own `transcripts/channels`; it used to fall back to an absolute path that existed on one machine only. The song project's videos default to `~/reports/quartering-uh-song/videos`; `SONG_DIR` and `VIDEO_ROOT` still win. The song project's tracked manifests record their paths relative to the song folders, and the twenty one-off `umtool/song/*.sh` run logs, which only ever ran on the machine that wrote them, are gone. - **umtool's production build no longer reads the corpus folder.** Since umtool began finding the corpus from its checkout (the bullet above), `next build` treated the checkout's whole `transcripts/channels` as files to bundle. On a real archive it ran out of memory and was killed, so umtool could not be rebuilt. The build now ignores that folder and finishes in about 25 s at under 1 GB, the same as a checkout with no corpus. Nothing changes when umtool runs. diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -7626,63 +7626,80 @@ this section is stale, by +16 near the top and +203 at the end; they are not rew ## The storage health gate (verified 2026-09-29, branch `r15/drive-stall`) The record is [`release-15.md`](release-15.md), "Slice DS, as shipped", with the parent's rulings -Q1–Q5. Anchors are at the branch tip. +(Q1–Q5) and the review's fixes (M1–M3, L1–L10). Anchors are at the branch tip after the merge of +`main` `bab894db`. - **A drive can be mounted and not answering.** Every in-process fs call on it waits on one of libuv's threads (4 by default, 16 in the editor's `start`) until it answers (~30 s for the observed USB reset loop); only a child process isolates a call. A child `stat` of a location's ROOT does not detect it reliably: the root's inode is in the kernel's cache whenever the drive was used lately. - **The state is `common/lib/storageHealth.ts`**, one map on `globalThis.__yttStorageHealth__` (the - watch writes it from instrumentation's module copy; pages read it from theirs), with each entry's - `detector`. `recordLocationHealth` (`:143`): one `stalled` answer stalls at once; - `HEALTH_CLEAN_TO_CLEAR` (2) clean answers in a row clear it; `absent` is clean; a new root starts - over. `registerLocationHealth` (`:199`) creates entries with no answer. `stalledLocationForPath` - (`:244`) matches like `locationOfDataDir`; `stalledLocation` (`:260`) is by id AND root. + pass writes it from instrumentation's module copy; pages read it from theirs), with each entry's + `detector` and the counters' `device`. `recordLocationHealth` (`:172`): one `stalled` answer stalls + at once, and every transition to stalled refuses `onDrive`'s waiting calls; `HEALTH_CLEAN_TO_CLEAR` + (2) clean answers in a row clear it; `absent` is clean; a new root starts over. + `registerLocationHealth` (`:257`) creates entries with no answer. `stalledLocationForPath` + (`:309`) matches like `locationOfDataDir`; `stalledLocation` (`:325`) is by id AND root. - **Detector 1, every 15 s: the block device's counters** (`detectLocationHealth`, - `lib/storageVolumes.ts:573`). `findmnt -J -T <root> -o SOURCE,UUID` raced against 3 s (a timeout - reuses the last device named for that root; another volume's UUID names none), `[…]` stripped, - `/dev/mapper` resolved, basename; then `/sys/class/block/<dev>/stat`: reads completed (1) + writes - completed (5), in flight (9). Stalled ⇔ in flight at both samples AND no completion between; the - samples at least `MIN_COUNTER_INTERVAL_MS` (10 s) apart; the first gives no verdict. No device → - the child `stat -L -c %F` probe (`probeLocationHealth`, `:384`). Detector state is module state in - `storageVolumes.ts` (`resetHealthDetector`). -- **Detector 2, on every gated call: `onDrive(where, call)`** (`storageHealth.ts:410`). Refused with + `lib/storageVolumes.ts:601`). The root's device from the last pass (findmnt `-J -T <root> -o + SOURCE,UUID`, raced against 3 s, only when there is none or its `/sys` entry stops reading; + another volume's UUID names none; `[…]` stripped, `/dev/mapper` resolved, basename); then + `/sys/class/block/<dev>/stat` (`parseBlockStat`, `storageHealth.ts:682`): completed = fields 1 + 5 + + 12 + 16 (reads, writes, discards, flushes), in flight = field 9. Stalled ⇔ in flight at both + samples AND nothing completed between; samples at least `MIN_COUNTER_INTERVAL_MS` (10 s) apart; the + first gives no verdict. in_flight counts only requests dispatched to the driver: one requeued + during a host reset is not counted, so a sample in that window can read clean (the watchdog covers + it). No device → the child `stat -L -c %F` probe (`probeLocationHealth`, `:386`). The samples are + on `globalThis.__yttHealthDetector__` (the pass and /storage's Refresh share them). +- **Detector 2, on every gated call: `onDrive(where, call)`** (`storageHealth.ts:584`). Refused with no call on a stalled location; otherwise raced against `DRIVE_CALL_BUDGET_MS` (3 s; test seam - `setDriveCallBudget`); a timeout marks the location stalled (since now) and throws - `DriveNotAnsweringError` (`isDriveNotAnswering`), leaving the call to settle. At most - `DRIVE_CALLS_IN_FLIGHT` (4) calls per location in flight; the rest queue in JS and are refused on - a stall; a slot is freed when its call really returns. Do not nest it for one location. A path on - no known location is raced but marks nothing; a location object whose root is not its entry's is - raced but does not rewrite the entry. The timer is not unref'd (a fake never-settling promise would - otherwise let a test process exit). -- **The cadence** is `runStorageHealthPass` (`controller/storageWatch.ts:407`): prune, register, then - every location concurrently; every 15 s from `startStorageWatch` (`:495`), plus one at arm time; - armed only with the watch, so an idle boot and a CLI process have no pass (a CLI's inspects are - still raced). `refreshLocationHealth` (`:459`) is /storage's Refresh. -- **The gate order in `inspectChannelMedia`** (`lib/channelMedia.ts:312`): config → memo (a + `setDriveCallBudget`); the budget covers the whole unit passed in. A timeout marks the location + stalled (since now) and throws `DriveNotAnsweringError`, leaving the call to settle — unless the + location's device counters (read synchronously from `/sys` through the reader `storageVolumes.ts` + registers with `setCounterReader`, `:509`) moved since the call began: then the call is refused + as slow and nothing is marked. At most `DRIVE_CALLS_IN_FLIGHT` (4) calls per slot key in flight + (`acquireSlot`, `:488`): the rest queue in JS, raced against the budget plus a quarter of it (at + most 250 ms) and refused unmarked when that runs out; refused at once by any transition to stalled; + and when every slot is held by a call already past its budget (`overdue`), a new call is refused + at once and the location marked stalled again. A slot is freed when its call really returns. The + slot key: a configured location's id; a probe of another root under its id, that root (marks + nothing); a path on no configured location, the root it is under (`rootOfUnknownPath`, `:444`; + marks nothing). Do not nest it for one key. The timer is not unref'd. +- **The cadence** is `runStorageHealthPass` (`controller/storageWatch.ts:433`): prune, register, then + every location concurrently; every 15 s from `startStorageHealthWatch` (`:546`), plus one at arm + time, armed by `editor/instrumentation.ts` ABOVE the idle gate (it writes nothing). The five-minute + pass (`startStorageWatch`, `:521`) stays below it. A CLI process has no pass (its inspects are still + raced). `refreshLocationHealth` (`:485`) is /storage's Refresh. +- **The gate order in `inspectChannelMedia`** (`lib/channelMedia.ts:316`): config → memo (a remembered `in-transition` is returned as is, anything else is gated first) → the relocation - marker (`:358`, corpus disk) → the gate (`:376`) → the link (corpus disk) → the target's `stat` - through `onDrive`. A stall is never memoised. Other gated calls: `probeLocation` (`storageVolumes.ts:253`) - and its memo (`:723`); `volumeFreeBytes` (`controller/storageLocations.ts:263`); - `readChannelStat` (`controller/channels.ts:215`, the walk through `onDrive`, `null` on a stall); - the recency tail reads; the move-root check; the saved-video store; `listSavedVideos` with - `notAnswering`; the videos list, the video page, the Cleanup stage and the media file route. - `channelMediaStall(config)` (`channelMedia.ts:227`) is the no-I/O question for a holder of a config. + marker (`:362`, corpus disk) → the gate (`:380`) → the link (corpus disk) → the target's `stat` + through `onDrive` (`:445`). A stall is never memoised. Other gated calls: `probeLocation` + (`storageVolumes.ts:255`, its stat and statfs through `onDrive`) and its memo; `volumeFreeBytes` + (`controller/storageLocations.ts:288`, `:325`, `:334`); `readChannelStat` (`controller/channels.ts:215`, + the walk through `onDrive`, `null` on a stall); the snapshot walk (`controller/channelSnapshot.ts:751`, + its listing, keep-latest keys and per-video unit); the recency tail reads; the move-root check; the + saved-video store; `listSavedVideos` with `notAnswering`; the videos list, the video page, the + Cleanup stage, the Storage stage's statfs and the media file route. `channelMediaStall(config)` + (`channelMedia.ts:231`) is the no-I/O question for a holder of a config. - **`stalled` is a sixth `ChannelMediaStatus` and a sixth `StorageLocationStatus`** ("Not answering"). `HELD_REASON`, `MediaLocationBadge`'s two tables and `STORAGE_STATUS_LABEL` are the `Record`s that make tsc name every table a seventh would need. `isMediaHeld` holds it, so both - pool-wide builds hold a stalled channel; the storage watch counts it as down (two passes pause). -- **`inspectChannelMedia` is memoised for 5 s** (`CHANNEL_MEDIA_MEMO_MS`, `:261`), keyed by channels + pool-wide builds hold a stalled channel; the storage watch counts it as down (two passes pause), + on a location or not, and the pause record carries `cause: "not-answering"` (`ChannelAutoPause`, + `lib/channelPriority.ts`; absent = not there). +- **`inspectChannelMedia` is memoised for 5 s** (`CHANNEL_MEDIA_MEMO_MS`, `:265`), keyed by channels dir, slug and configured `dataDir`, on `globalThis.__yttChannelMediaMemo__`. `{ fresh: true }` skips it and does not store; the deciders that pass it are listed in the record (the guard and its six callers, both movers, both builds, the watch, eviction, the re-point preflight, doctor). The runners' tick shares the status poll's `buildChannelWork` and so reads the memo. - `forgetChannelMedia` (`:287`) is called by the channel mover's marker writes and clear, + `forgetChannelMedia` (`:291`) is called by the channel mover's marker writes and clear, `clearRelocationMarker`, a re-point, /storage's Refresh and the e2e `invalidate-cache` route. - **`UV_THREADPOOL_SIZE`** defaults to 16 in `editor/package.json`'s `start` and in `docker/entrypoint.sh`. `ports.test.ts` reads every `${NAME:-N}` in a script as a port and names it as the one exception (`NUMERIC_NOT_PORTS`, `common/lib/ports.test.ts:29`). -- **Not covered:** a call already in flight when the drive stalls (at most four per location through - `onDrive`); a drive already stalled at boot before the second counter sample, unless a page reaches - it; jobs already running and their own reads; per-click server actions; the file route's stream; - the corpus disk itself. +- **Not covered:** a call already in flight when the drive stalls (at most four per drive for the + calls through `onDrive` — every page and poll path and the snapshot walk; a job's own reads that do + not go through it, `measureTree`, the index build's processing phase, the snapshot's sequential + reconcile pass, are not capped); a hand-typed root is capped but never marked; a drive already + stalled at boot before the second counter sample, unless a page reaches it; per-click server + actions; the file route's stream; the corpus disk itself. diff --git a/plans/release-15.md b/plans/release-15.md @@ -470,7 +470,8 @@ Branch `r15/drive-stall` off `main` `ccf90892` (slice IG merged), worktree `~/Pr job's `tmp`. The ruling: a drive that is mounted and not answering must not stop the editor answering. Two detectors find the stall without the editor waiting on the drive, and pages and polls do not touch it in-process while it is not answering. The parent's five rulings on the first pass -(Q1–Q5, below) are applied. +(Q1–Q5) and the review's fixes (M1–M3, L1–L10) are applied; the tables at the end map each to its +commit. **What was wrong.** Node runs every filesystem call on libuv's thread pool, four threads by default. On a drive that has stalled (an SMR disk in a USB enclosure resetting under a long write) each call @@ -495,20 +496,25 @@ stalled disk is here. `common/lib/storageVolumes.ts`; ruling Q1(a)). A child `stat` of the root is answered from the kernel's inode cache whenever the drive was used lately, so it can say "ok" while the reads that reach the device wait out a reset loop. Instead, per location per pass: - - the root's device: `findmnt -J -T <root> -o SOURCE,UUID` as a child raced against 3 s (a findmnt - that does not answer reuses the device the last pass named for that root), a `[subvolume]` - suffix taken off, `/dev/mapper/*` resolved to its `dm-N` (a read of `/dev`), the basename. A UUID - other than the location's recorded one names no device (the root is then a directory on another - filesystem, not the drive). - - `/sys/class/block/<dev>/stat`, which never touches the drive: reads completed (field 1) + writes - completed (5), and requests in flight (9). Against the previous pass's sample for the location - (same device, at least `MIN_COUNTER_INTERVAL_MS` = 10 s earlier): **stalled ⇔ in flight at both - AND no completion between**; anything else is clean. The first sample gives no verdict. + - the root's device: `findmnt -J -T <root> -o SOURCE,UUID` as a child raced against 3 s, a + `[subvolume]` suffix taken off, `/dev/mapper/*` resolved to its `dm-N` (a read of `/dev`), the + basename. Asked only when the root has no device yet or its device's `/sys` entry stops reading + (a replug under another name): a findmnt per pass would leave one child stuck per pass during a + long stall (L10). A UUID other than the location's recorded one names no device (the root is + then a directory on another filesystem, not the drive). + - `/sys/class/block/<dev>/stat`, which never touches the drive: completed = reads (field 1) + + writes (5) + discards (12) + flushes (16) where the kernel counts them (in_flight counts those + too, so a long SMR media-cache flush alone moves completions; L2), and requests in flight (9). + Against the previous sample for the location (same device, at least `MIN_COUNTER_INTERVAL_MS` = + 10 s earlier): **stalled ⇔ in flight at both AND nothing completed between**; anything else is + clean. The first sample gives no verdict. The samples are on `globalThis` + (`__yttHealthDetector__`), so the pass and `/storage`'s Refresh, in different module copies, + compare against one previous sample (L5). - **No device** (a container, no findmnt, a tmpfs or network source, no `/sys` entry) falls back to the child `stat` probe (`probeLocationHealth`): `stat -L -c %F -- <root>` raced against 3 s, the child SIGKILLed and not waited for; `directory` is `ok`, anything else `absent`, no binary `ok`. - - The verdict names its detector (`"counters" | "stat"`), the health state records it, and - `/storage`'s line says which watched the drive. + - The verdict names its detector (`"counters" | "stat"`), the health state records it (and the + counters' device, which the watchdog reads), and `/storage`'s line says which watched the drive. - **Detector 2, on every gated call: a 3 s watchdog** (`onDrive(where, call)`, `lib/storageHealth.ts`; ruling Q1(b)). The detector that cannot be fooled by a cache: a page or poll that actually reaches the drive finds out. @@ -517,37 +523,55 @@ stalled disk is here. location stalled (since now, cause "a read in the editor did not answer within 3 s") and throws `DriveNotAnsweringError`; the caller answers `stalled`. The call is left to settle on its own: its thread is the stated limit. - - **At most `DRIVE_CALLS_IN_FLIGHT` (4) calls per location are in flight through it.** The rest - wait in a queue of its own (not libuv's) and are refused without a call the moment the location - stalls, so a stalled drive holds at most four of the pool's 16 threads, and a 64-wide walk that - meets a stall puts four calls on it, not 64. A slot is released when its call really returns. - Calls are not nested for one location; a unit of work (a video directory's few reads, a page's - reads of one video) goes through as one call. - - A path on no location the health state knows is raced but marks nothing (the error names no - location). A probe of another root under a location's id is raced and does not rewrite that - location's entry. The timer is not `unref`'d: it is cleared the moment the call answers. -- **The cadence** (`common/controller/storageWatch.ts`): `startStorageWatch` arms a second timer, - every 15 s, beside the five-minute pass, and runs one health pass at once. All locations are asked - concurrently, each bounded by its own timers, and an overrunning pass is not stacked. A transition - is logged (`[storage] "<id>": drive not answering — <cause>; …` / `answering again (ok)`). Stopped - with the watch; the watch is armed below the idle gate, so an idle boot has no health pass. A CLI - process has no pass either, but its gated calls still go through the watchdog. + - **Slow is not stalled** (L6): on a timeout, when the counters detector has named the location's + device, its counters are read (synchronously, from `/sys`, so the check does not wait behind the + pool it is judging) and compared with a reading taken when the call began. Requests completed + meanwhile: the drive is slow; the call is refused and nothing is marked. + - **The budget covers a whole unit of work** (L8): a video directory's reads, a page's reads of + one video, go through as one call, so a slow drive still answering can be marked by one long + unit (unless the counters show it completing, above). + - **At most `DRIVE_CALLS_IN_FLIGHT` (4) calls per slot key are in flight.** The rest wait in a + queue of its own (not libuv's), so a 64-wide walk that meets a stall puts four calls on the + drive, not 64. A slot is released when its call really returns. The queue (M1): + - every transition to `stalled` refuses the waiting calls at once, whoever decided it (the + pass, the watchdog, a Refresh); + - a wait is raced against the budget plus a grace of a quarter of it (at most 250 ms), so the + calls it waits behind, whose timers start a moment later, time out and mark first; a wait + that runs out is refused without marking; + - calls past their budget are counted per key (`overdue`); when every slot is held by one, a + new call is refused at once and the location marked stalled again, even if the pass has since + cleared it: none of those calls has returned. + - **The slot key** (M2, L7): a configured location's id; a probe of another root under a + location's id is keyed by that root and marks nothing; a path on no configured location (a root + typed by hand) is keyed by the root it is under (`<root>/<slug>/data` → `<root>`), so it holds + at most four threads too, and nothing can mark it. Calls are not nested for one key. The timer + is not `unref`'d: it is cleared the moment the call answers. +- **The cadence** (`common/controller/storageWatch.ts`): `startStorageHealthWatch` arms the 15 s + health pass and runs one at once; **`editor/instrumentation.ts` arms it above the idle gate**, + beside the storage boot probe (M2): it is in memory and writes nothing, and without it an idle + boot has no registered locations and a stall the watchdog marks is never cleared. The five-minute + pass (`startStorageWatch`), which may write an auto-pause, stays below the gate. All locations are + asked concurrently, each bounded by its own timers, and an overrunning pass is not stacked. A + transition is logged (`[storage] "<id>": drive not answering — <cause>; …` / `answering again + (ok)`). A CLI process has no pass, but its gated calls still go through the watchdog. - **The gate and the watchdog, by caller:** | Caller | On a stalled location | Through `onDrive` | |---|---|---| - | `inspectChannelMedia` (`lib/channelMedia.ts`) | **The relocation marker is read first** (ruling Q2: it is in the channel dir, on the corpus disk), so a channel mid-move on a stalled drive reads `in-transition`. Then the gate: status `stalled`, detail `drive not answering (location "<label>", since HH:MM)`, before the link and the target. With a config in hand a stalled channel costs one call, the marker read. | The target's `stat` (`:441`). | + | `inspectChannelMedia` (`lib/channelMedia.ts`) | **The relocation marker is read first** (ruling Q2: it is in the channel dir, on the corpus disk), so a channel mid-move on a stalled drive reads `in-transition`. Then the gate: status `stalled`, detail `drive not answering (location "<label>", since HH:MM)`, before the link and the target. With a config in hand a stalled channel costs one call, the marker read. | The target's `stat` (`:445`). | | `assertChannelMediaReachable` | Refuses it (`ChannelMediaUnreachableError`, status `stalled`), so `runManagedFunction`'s `needsMedia` guard, `generateChannelSnapshot` (also the channel page's **Refresh report**), the operation batch, normalise, keep-videos and the shard action refuse it. | Through inspect. | | The index and stats builds | Hold it: `HELD_REASON.stalled` is "its drive is not answering (a stalled disk)", and IG's `isMediaHeld` holds every status but `ok` and `in-place`. | Through inspect (both of the index build's looks). | - | The storage watch's five-minute pass | Counts a stalled location as down: two passes auto-pause its channels, and the pass after the drive answers restores them, as for an unmount (ruling Q3). | Through inspect and `probeLocation`. | - | `probeLocation` and `probeLocationMemo` (`lib/storageVolumes.ts`) | Probe status `stalled` (`STORAGE_STATUS_LABEL`: "Not answering"), identity unknown, no free space; no `stat`, `statfs` or `findmnt`. The memo is asked after the gate. | The root's `stat` and `statfs` (`:260`, `:277`). | + | The storage watch's five-minute pass | Counts a stalled location, or a channel whose inspect says `stalled` on no location (L9), as down: two passes auto-pause its channels, and the pass after the drive answers restores them, as for an unmount (ruling Q3). The pause record carries `cause: "not-answering"`, and `autoPauseReasonOf` says "on a drive that is not answering … when the drive answers again" (L4; a record without a cause, all written before, reads as not there). | Through inspect and `probeLocation`. | + | `probeLocation` and `probeLocationMemo` (`lib/storageVolumes.ts`) | Probe status `stalled` (`STORAGE_STATUS_LABEL`: "Not answering"), identity unknown, no free space; no `stat`, `statfs` or `findmnt`. The memo is asked after the gate. | The root's `stat` and `statfs` (`:262`, `:279`). | | `volumeFreeBytes` (`controller/storageLocations.ts`) | Unknown ("—"), with no call. | The root's and the mountpoint's `stat`, and the `statfs` (`:288`, `:325`, `:334`). | + | `generateChannelSnapshot` (`controller/channelSnapshot.ts`), after every download or sync, sixteen video directories wide (M3) | Refused by its start guard. | Its `data/` listing (a refusal is rethrown, never read as an empty channel), the keep-latest keys' metadata reads (`keyedVideosNewestFirst`, now `mapConcurrent` 16 wide instead of an unbounded `Promise.all`) and each video directory's unit (`:751`, `:844`). A throw keeps the last `snapshot.json`, as on any failed refresh. The sequential reconcile pass before it is not raced. | | `readChannelStat` (`controller/channels.ts`) | `null` (no counts, so no progress bar), with no walk. The one-second job-list poll and the home page ask it for every channel with a job listed. | The `data/` readdir, then each video directory as one call (`:107`); `null` when the drive stops answering mid-walk. The batch jobs' `listChannelStatsFromDisk` passes no drive and is unchanged. | | The recency tail reads (`controller/recencyIndex.ts`) | Skipped, and NOT remembered as misses: layers 3 and 4 until a refresh with the drive answering reads them. | Each tail read of a relocated channel (`:257`). | | `relocationRootPresenceProblem` (`controller/relocateChannelMedia.ts`) | A move onto a stalled location is refused before the root's `stat`. | The root's `stat` (`:307`). | | `inspectSavedVideosStore` (`controller/relocateSavedVideos.ts`) | `unreachable` with the stall's detail; `/storage` skips the store's size walk. | The target's `stat` (`:181`); `/storage`'s store walk too. | | `listSavedVideos` (`controller/savedVideoInventory.ts`), when a page passes `notAnswering` | The channel is skipped and named (its `data/` link is read, not followed). The backup job passes nothing and is unchanged. | Each read of a relocated channel (`:81`). | | The videos list, the video page (and its title), the channel page's Cleanup stage | A notice (`aria-label="media not answering"`, `MediaNotAnswering.tsx`) with links to the channel and `/storage`. | The list's `data/` listing and titles, then the selected video's files; the video page's whole directory read, as one unit; its title; the Cleanup stage's saved-video totals. | + | The channel page's Storage stage free space (L1) | "—". | The `statfs` of the relocated target. | | The media file route (`/api/channels/<slug>/videos/<id>/files/<name>`) | 503 with `Retry-After: 15`. | The file's `stat`; the stream after it is not raced. | - **The memo** (`inspectChannelMedia`): five seconds per channel, keyed by channels dir, slug and @@ -558,23 +582,23 @@ stalled disk is here. | Caller | Where | |---|---| - | `assertChannelMediaReachable` (so every guard below) | `lib/channelMedia.ts:510` | + | `assertChannelMediaReachable` (so every guard below) | `lib/channelMedia.ts:514` | | ↳ `runManagedFunction`'s `needsMedia` guard | `jobs/streamCommand.ts:283` | | ↳ the operation batch | `controller/operationBatch.ts:1593` | - | ↳ `generateChannelSnapshot` | `controller/channelSnapshot.ts:738` | + | ↳ `generateChannelSnapshot` | `controller/channelSnapshot.ts:739` | | ↳ `normalizeAllTranscripts` | `controller/normalizeAll.ts:63` | | ↳ keep-videos | `controller/keepVideosMatching.ts:165` | | ↳ the shard action | `editor/app/channels/[slug]/shardActions.ts:88` | | The channel mover, out and back | `controller/relocateChannelMedia.ts:638`, `:877` | | The index build, before and after the walk | `controller/buildIndex.ts:348`, `:467` | - | The stats build | `controller/buildStats.ts:286` | - | The storage watch's five-minute pass | `controller/storageWatch.ts:233` | + | The stats build | `controller/buildStats.ts:291` | + | The storage watch's five-minute pass | `controller/storageWatch.ts:235` | | Clip-window eviction | `controller/evictClipWindows.ts:99` | | The re-point preflight | `controller/storageLocations.ts:564` | | `archilyzer doctor` | `bin/doctor.ts:124` | **The memo's readers**, pages and polls: the home page (`editor/app/page.tsx:80`), `/channels` - (`editor/app/channels/page.tsx:219`), the channel page (`[slug]/page.tsx:235`), the ops channel + (`editor/app/channels/page.tsx:219`), the channel page (`[slug]/page.tsx:239`), the ops channel route (`api/ops/channel/[slug]/route.ts:70`), `channelsOnLocation`'s rollup for `/storage` (`controller/storageLocations.ts:225`), the runners' `buildChannelWork` (`controller/autoRunner.ts:622`, the tick and the three-second status poll share it), and the bulk actions' skips @@ -596,6 +620,7 @@ stalled disk is here. | `/channels` | The volume chip reads `… · not answering since HH:MM` in place of its free space, with a title saying what it means. Each row's badge reads `on <label> — not answering` (accessible name `media location: Media not answering · on <label>`). | | The channel page | The Storage stage's card is red with the inspector's sentence; its destination list names the location "Not answering"; Move back is withheld for a stalled channel, as for an unreachable one. | | `/saved-videos` | `Not read, because the drive their media is on is not answering: <slugs>.` (`aria-label="saved videos not read"`). | + | `/review`, the rack, the channel page (an auto-paused channel) | `Auto-paused — its media is on a drive that is not answering since <date>. It returns to <tier> on its own when the drive answers again.` (L4) | | The runners | `[auto] skipping <slug>: media stalled — drive not answering (…)`, once per state change. | | The logs | The health pass's transition lines, with the cause; the index and stats builds' hold lines. | @@ -613,77 +638,100 @@ stalled disk is here. | `c69ad41a` | `common:` rulings Q2 and Q1(b): the marker before the gate; `onDrive` (the 3 s watchdog and the four-call cap per location) on every common gated call; `registerLocationHealth`. Tests. | | `f6a25cf5` | `editor:` the pages' reads of a relocated drive through `onDrive` (the videos list, the video page and its title, the file route, the Cleanup totals, `/storage`'s store walk). | | `b1a30902` | `common:` ruling Q1(a): `detectLocationHealth`, the block-device counters with the child-stat fallback; `detector` in the health state and on `/storage`. Tests. | -| this commit | `plans:` this section rewritten for the rulings; FACTS; the changelog; the report. | +| `981e05a0` | `plans:` this section rewritten for the rulings; FACTS; the changelog. | +| `33094c29` | `common:` review M1, M2 (slot keys), L2, L5, L6, L7, L8, L10: the queue's deadline, the overdue count, waiters refused on every transition to stalled; the hand-typed-root and candidate-root keys; slow is not stalled; discards and flushes counted; the detector's samples on `globalThis`; findmnt only when needed. Tests. | +| `bd39579e` | `common, editor:` review M2, L4, L9: the health pass armed above the idle gate; the pause record's `cause`; a `stalled` channel on no location is down. SETTINGS.md. Tests. | +| `ff235c2f` | `common:` review M3: the snapshot walk through `onDrive`; keep-latest keys bounded. Test. | +| `19c5842d` | `editor:` review L1, L3: the Storage stage's `statfs` through `onDrive`; the stale comments. | +| `30df4193` | Merge `main` (`bab894db`: release 14 HS, S1, CF; release 15 UT). Two conflicts, both kept: the changelog's `[Unreleased]` carries UT's bullet then DS's; this file is `main`'s with DS's table row and this section after UT's. | +| this commit | `plans:` the review's findings to their commits; FACTS; the changelog; the report. | **Tests** (unit; no test stalls a real drive: a stalled call is a promise that never settles or a fake `stat` that never answers, and a stalled device is a temp `/sys` whose counters stand still) | File | What it pins | |---|---| -| `lib/storageHealth.test.ts` (19) | The rules: one miss stalls at once, one clean answer after a stall does not clear it and two in a row do, a miss in between starts over, `absent` is clean, a re-pointed root starts over, the path match is `locationOfDataDir`'s, pruning, `globalThis`, the "since" wording, registering without an answer. `onDrive`: an answer passes through (value or error) and frees its slot; a never-settling call is `stalled` on the timer, marks the location (since now) and keeps its slot until it settles; a stalled location is refused with no call; seven calls at once put four in flight and the three that waited are refused without a call when the location stalls; a freed slot runs a waiter; an unknown path marks nothing; a candidate root does not rewrite its location; the 3 s default (answered after 3–4.5 s). | -| `lib/storageHealthCounters.test.ts` (7) | The stat line parser (17 and 11 fields, garbage); the verdict over sample pairs (stuck → stalled; moving, idle or drained → ok); device names (partition, `[subvolume]` suffix, non-`/dev` sources); the detector end to end with a fake findmnt and a temp `/sys`: the first sample gives no verdict, stuck → stalled with its cause, drained → ok, busy and moving → ok; a sample sooner than 10 s gives none and keeps the first; every no-device fallback goes to the child stat and says so (tmpfs, findmnt failing, no `/sys` entry, another volume's UUID, no binary); a findmnt that does not answer is not waited for and the last device named is read. | +| `lib/storageHealth.test.ts` (22) | The rules: one miss stalls at once, one clean answer after a stall does not clear it and two in a row do, a miss in between starts over, `absent` is clean, a re-pointed root starts over, the path match is `locationOfDataDir`'s, pruning, `globalThis`, the "since" wording, registering without an answer. `onDrive`: an answer passes through (value or error) and frees its slot; a never-settling call is `stalled` on the timer, marks the location (since now) and keeps its slot until it settles; a stalled location is refused with no call; seven calls at once put four in flight and the three that waited are refused without a call when the location stalls; a freed slot runs a waiter; the 3 s default (answered after 3–4.5 s). After the review: a hand-typed root's four slots are shared by its channels and a fifth call is refused at once, with no entry made; a candidate root has its own slots and leaves the location's calls alone; **M1's case** (four hung calls, two clean answers, a fifth refused at once and the location marked again); a transition to stalled from the pass refuses the waiting call; **L6** (counters completing: four calls refused as slow, the location not marked, a waiter's wait runs out unmarked; counters standing still: marked). | +| `lib/storageHealthCounters.test.ts` (7) | The stat line parser (17 and 11 fields, garbage); the verdict over sample pairs (stuck → stalled; moving, idle or drained → ok); device names (partition, `[subvolume]` suffix, non-`/dev` sources); the detector end to end with a fake findmnt and a temp `/sys`: the first sample gives no verdict, stuck → stalled with its cause, drained → ok, busy and moving → ok; a sample sooner than 10 s gives none and keeps the first; every no-device fallback goes to the child stat and says so (tmpfs, findmnt failing, no `/sys` entry, another volume's UUID, no binary). After the review: discards and flushes are counted (a flush alone is not a stall); a known device is read without running findmnt, and findmnt runs again only when its `/sys` entry stops reading; the samples are on `globalThis`. | | `lib/storageHealthProbe.test.ts` (5) | The child `stat`: a directory is `ok`, a missing path and a file are `absent`; a fake `stat` asleep for 20 s is `stalled` on the timer without being waited for; the 3 s default; an answer inside the budget is taken; no binary is `ok`. | -| `controller/storageStall.test.ts` (21) | A spy on every `node:fs` and `node:fs/promises` call, with a hang mode that makes a matching promise-API call never settle. The gate: with the location stalled, inspect reads only the marker (with a config) or config.json and the marker (without); a channel mid-move on a stalled drive reads `in-transition` (and is remembered so); the guard, `probeLocation` and its memo (no findmnt run), `volumeFreeBytes`, `readChannelStat`, the recency layer, the move-root check, a snapshot refresh, the saved-video store and the inventory make no call on the drive. The watchdog: with the location answering and one drive call hung, inspect, `readChannelStat` (at most four video directories asked), `probeLocation`, `volumeFreeBytes`, the recency layer (not remembered as a miss) and the move-root check each answer `stalled` within the race and mark the location; after it, inspect, the guard and the walk make no call on the drive, and once cleared the drive is asked again. The memo: two inspects inside 5 s stat the target once, `fresh` and a 5 s age ask again, a fresh answer is not stored, another target is another key, `forgetChannelMedia` and `clearRelocationMarker` clear it, and a stall is seen with an `ok` remembered. | -| `controller/storageWatch.test.ts` (+7) | The pass stalls a location on one miss and a page then gets `stalled`; the five-minute pass suspects its channel; two clean passes clear it; a location no longer configured is forgotten; a probe that throws is `ok`; arming runs one pass at once and stopping stops both timers; a Refresh counts as one answer; the pass registers every location, records a verdict's detector, and a verdict with no answer changes nothing. | +| `controller/storageStall.test.ts` (22) | A spy on every `node:fs` and `node:fs/promises` call, with a hang mode that makes a matching promise-API call never settle. The gate: with the location stalled, inspect reads only the marker (with a config) or config.json and the marker (without); a channel mid-move on a stalled drive reads `in-transition` (and is remembered so); the guard, `probeLocation` and its memo (no findmnt run), `volumeFreeBytes`, `readChannelStat`, the recency layer, the move-root check, a snapshot refresh, the saved-video store and the inventory make no call on the drive. The watchdog: with the location answering and one drive call hung, inspect, `readChannelStat` (at most four video directories asked), `probeLocation`, `volumeFreeBytes`, the recency layer (not remembered as a miss) the move-root check and (M3) the snapshot walk each answer `stalled` within the race and mark the location (the snapshot writes no `snapshot.json`, and at most four video directories reach the drive); after it, inspect, the guard and the walk make no call on the drive, and once cleared the drive is asked again. The memo: two inspects inside 5 s stat the target once, `fresh` and a 5 s age ask again, a fresh answer is not stored, another target is another key, `forgetChannelMedia` and `clearRelocationMarker` clear it, and a stall is seen with an `ok` remembered. | +| `controller/storageWatch.test.ts` (+9) | The pass stalls a location on one miss and a page then gets `stalled`; the five-minute pass suspects its channel; two clean passes clear it; a location no longer configured is forgotten; a probe that throws is `ok`; the health pass arms on its own, runs once at once and stops, and the five-minute watch arms no health pass; a Refresh counts as one answer; the pass registers every location, records a verdict's detector, and a verdict with no answer changes nothing; a counters verdict records its device and a stat verdict forgets it; a stall auto-pauses after two passes with `cause: "not-answering"` and its wording, the sanitizer keeps the cause, and a record without one reads as not there. | | `views/storage.test.ts` (+1) | A stalled row reads "Not answering", carries its line, has no free space, and withholds Re-point and Mount. | #### Gates (logs `$T/ds-*.log`) -- **tsc** was clean before every commit. After the rulings: 73 s at `c69ad41a`/`f6a25cf5`, 261 s at - `b1a30902` (the machine was running other sessions' suites). -- **Unit, at `b1a30902`:** +- **tsc** was clean before every commit and on the merged tree (65 s). +- **Unit, on the merged tree (`30df4193`):** | Suite | Result | |---|---| - | common | **2,280/2,280**, 101 s (the branch point's 2,220 plus 60 new). 2,256 at the first pass. | + | common | **2,295/2,295**, 58 s: `main`'s 2,229 plus DS's 66 (60 at the rulings, 6 from the review). | | editor unit | 87/87 | - | `test:scripts` | 191 passed, 1 skipped (192), on two reruns. The first run beside another session's e2e failed `queue-lock.test.mjs:85` (the holder's details were not yet readable), as it did once on the first pass; it passes alone. | - | mcp | 271/271 (first pass; no mcp file changed since) | - -- **Docs:** `docs env --check`, `docs files --check` and `settings example --check` all exit **0**. -- **Build** (first pass, at `de4128d1`, now `4e1f7a90`): the editor's `next build`, with the - primary's `transcripts/` linked in and capped at 5 GB with no swap: 104 s, max RSS 1,643,860 KB. - The link was removed after the build, and nothing ran through it. Not rerun after the rulings (not - in the re-gate list). + | `test:scripts` | 194 passed, 2 skipped (196). The second skip is UT's post-build trace check, which skips a umtool build older than its config (this worktree's `umtool/.next` predates UT); the first is the `LIVE=1` archive check. | + | mcp | 271/271 at the first pass; no mcp file has changed since. | + +- **Docs:** `docs env --check`, `docs files --check` and `settings example --check` all exit **0** + (SETTINGS.md regenerated in `bd39579e` for the pause record's `cause`). +- **Build:** the editor's `next build` on the merged tree, with the primary's `transcripts/` linked + in and capped at 5 GB with no swap: 63 s, max RSS 1,641,444 KB, exit 0. The link was removed after the build, and + nothing ran through it. (First pass: 104 s, max RSS 1,643,860 KB.) - **e2e** (editor, detached and queued): | Run | Specs | Result | |---|---|---| - | 1, at `6a74c790` (now `480f2556`) | the `storage` and `channels` specs, `video-page`, `saved-videos`, `dashboard`, `auto-queue`, and IG's eight (`$T/ds-specs.txt`) | **124 passed, 3 failed, 12 skipped, 16.6 min**. The three were 30 s timeouts while this slice's own tsc and common suite ran beside the suite (`auto-queue.spec.ts:218` and `:352`, `tags.spec.ts:463`). | - | 2, at `de4128d1` (now `4e1f7a90`) | `auto-queue`, `tags`, `saved-videos`, four cleanup specs, `video-titles`, `video-page` (`$T/ds-specs2.txt`) | **70 passed, 1 failed, 5.9 min**: `auto-queue.spec.ts:352`, the Start-button race its own comment describes (a known flake). | - | 3, at `de4128d1` | `auto-queue` alone | **24 passed, 1.3 min**. | - | 4, at `b1a30902` (the re-gate) | `storage-locations`, `channel-storage`, `channels-storage-columns`, `channels`, `channels-actions`, `channels-counts`, `channels-sort`, `channels-rack-layers`, `channels-rack-audit` (`$T/ds-specs4.txt`) | **38 passed, 0 failed, 12 skipped, 2.5 min** (the 12 are `channels-rack-audit`, which skips without `E2E_RACK_SHOTS`). | - - No e2e fixture has a stalled drive: these confirm nothing changed for drives that answer. The - stall paths are the unit tests above. + | 1, first pass | the `storage` and `channels` specs, `video-page`, `saved-videos`, `dashboard`, `auto-queue`, and IG's eight (`$T/ds-specs.txt`) | **124 passed, 3 failed, 12 skipped, 16.6 min**: three 30 s timeouts while this slice's own tsc and common suite ran beside the suite. | + | 2, first pass | `auto-queue`, `tags`, `saved-videos`, four cleanup specs, `video-titles`, `video-page` (`$T/ds-specs2.txt`) | **70 passed, 1 failed, 5.9 min**: `auto-queue.spec.ts:352`, the Start-button race its own comment describes. | + | 3, first pass | `auto-queue` alone | **24 passed, 1.3 min**. | + | 4, after the rulings (`b1a30902`) | the 9 `storage` + `channels` specs (`$T/ds-specs4.txt`) | **38 passed, 0 failed, 12 skipped, 2.5 min**. | + | 5, after the review, on the merged tree (`30df4193`) | the 9 `storage` + `channels` specs; the snapshot scheduler's (`auto-report-refresh`, `channel-work`, `jobs-batch-tasks-drain`, `incomplete-transcript`); the keep-latest users of the bounded key fan-out (`cleanup-holds`, `saved-videos`); `reconcile`; `review` (the auto-pause wording) — `$T/ds-specs5.txt` | **73 passed, 0 failed, 12 skipped, 5.2 min** (37 min 54 s in the queue behind another session's suite). | + + The 12 skips in each are `channels-rack-audit`, which needs `E2E_RACK_SHOTS`. No e2e fixture has + a stalled drive: these confirm nothing changed for drives that answer. The stall paths are the + unit tests above. - **Numbers tool:** none. #### Found and left -- **The stated limit.** A read already in flight when the drive stalls holds its thread until the - kernel gives up (about 30 s in the observed reset loop). With the watchdog and the four-call cap, a - stall that the counters have not yet seen costs at most four threads per location for that long, - and the pages and polls asking answer within 3 s; with 16 threads the editor keeps answering. +- **The stated limit.** A call already in flight when the drive stalls holds its thread until the + kernel gives up (about 30 s in the observed reset loop). On one drive, at most four threads wait + that way for the calls that go through `onDrive`: every page and poll path, and the snapshot + walk. A job's own reads that do not go through it are not capped: `measureTree` in the movers, + the index build's processing phase (IG recorded it), the snapshot's sequential reconcile pass + (one read at a time), and the keep-latest reads of the other callers of `computeKeptVideoIds` + (cleanup, persist, prune; now 16 at a time). With 16 threads the editor keeps answering while + those four wait. +- **A root typed by hand** (a channel moved to a root no storage location names): its calls are + capped by the root they are under, four at a time, and raced, but nothing can mark it, so every + page and poll keeps asking it, each answering after 3 s or at once while its four slots are held + by calls the watchdog gave up on. Its channels' `stalled` comes from the watchdog alone, and the + five-minute pass counts it as down (L9). Naming the root as a location on `/storage` gives it the + full treatment. +- **The drive's spin-up** (L6, a question for the operator): the budget is 3 s, and a USB disk that + spins down when idle can take 3–10 s to answer its first read. The watchdog then refuses that + read, and marks the location unless the counters show requests completing meanwhile (spinning + up completes none), for 15–30 s, during which the start-of-work guard refuses the channel's jobs. + **Does the drive spin down when idle?** If it does, a longer budget for the first call after an + idle spell, or a spin-down timer on the drive, would avoid it. +- **A slow drive under load can fail a snapshot refresh.** Each video directory's unit is raced + against 3 s; a drive that is slow but completing requests is refused without being marked, and + the refresh throws, so the scheduler keeps the last `snapshot.json` and tries again on its next + trigger. - **The counters need two samples.** A drive already stalled when the editor starts is seen by the counters at the second pass (15–30 s), or at once by the watchdog when a page reaches it. -- **Units of work count as one slot.** A video directory's reads, or a page's reads of one video, go - through as one call, so a slot can hold a few sequential calls, and the video page's parallel - reads of one directory run under one slot. + in_flight counts only requests dispatched to the driver: a request requeued during a host reset + is not counted, so a sample in that window can read clean; the watchdog covers it. - **Ungated request paths**, each a click rather than a page or poll: the channel and video server actions that read a video's directory in the request (`bulkVideoActions`, `digestActions`, `videoActions`, `fixIncompleteTranscript`, `pipelineActions`); most of what they do is enqueue jobs, whose guard is fresh and refuses a stalled channel. `/api/media/fetch-window/<jobId>` (one `stat` of the fetched file, once the job is done). The media file route's stream after its `stat`. -- **Jobs are gated only at their start.** A job already running when its drive stalls (the remux that - stalls it, for one) keeps its threads and children; a job's own in-process reads (`measureTree`, - the index build's processing phase, which IG recorded) are not raced. - **The runners' tick reads the memo.** `buildChannelWork` serves both the tick and the status poll, so a drive unmounted in the last 5 s can have one unit dispatched, which fails at the dangling link. The snapshot regeneration and `runManagedFunction`'s guard are fresh. - **umtool's twin of the reachability check** (`checkChannelReachable` in `umtool/report-to-video/cues.mjs`) has no stall gate: umtool is its own process with no health state, and `umtool/**` belongs to another slice. -- **A CLI process** (`archilyzer index`) has no health pass, so nothing is gated in it; its +- **A CLI process** (`archilyzer index`) has no health pass, so nothing is marked in it; its inspects still go through the watchdog, so a target `stat` that takes over 3 s holds the channel. The corpus disk itself is not watched. @@ -699,8 +747,11 @@ fake `stat` that never answers, and a stalled device is a temp `/sys` whose coun | What I assumed | The alternative | |---|---| -| At most four gated calls per location in flight; the rest queue in JavaScript and are refused on a stall. | No cap: the watchdog alone, and a stall mid-walk fills the pool until the kernel gives up. | -| Two counter samples closer than 10 s give no verdict (a Refresh just after a pass among them). | Compare any two samples (a busy healthy drive can read "in flight, nothing completed" over a few milliseconds). | +| At most four gated calls per location in flight; the rest queue in JavaScript and are refused on a stall. **Kept at review.** | No cap: the watchdog alone, and a stall mid-walk fills the pool until the kernel gives up. | +| A wait for a slot runs out at the budget plus a quarter of it (at most 250 ms), so simultaneous timeouts of the calls it waits behind mark first. | Exactly the budget: a waiter queued in the same tick then gives up a moment before those calls and is refused unmarked, and the mark lands a few ms later. | +| The keep-latest key reads run 16 at a time for every caller (they were an unbounded `Promise.all`). | Bound them only for the snapshot. | +| An auto-pause record with no `cause` reads as not there. | Word both cases for it ("not there or not answering"). | +| Two counter samples closer than 10 s give no verdict (a Refresh just after a pass among them). **Kept at review.** | Compare any two samples (a busy healthy drive can read "in flight, nothing completed" over a few milliseconds). | | A findmnt that does not answer reuses the last device named for that root; one naming another volume's UUID names none. | Treat a findmnt that does not answer as a stall. | | An `absent` answer counts as clean toward clearing a stall. | Only `ok` clears it. | | `/storage`'s Refresh is one answer like the pass's. | Refresh clears a stall outright on one clean answer. | @@ -714,4 +765,26 @@ on an idle boot). The restart must go through `pnpm run start` in `editor/` (the script does) for `UV_THREADPOOL_SIZE` to apply; a process started another way keeps Node's 4 threads unless the variable is set. CLI builds have no health state; their inspects are raced. +#### Review + +**Verdict: SHIP AFTER FIXES** (`ds-review.md` in the job's scratch). No High. The parent's rulings +on each finding were applied as below. + +| Finding | Ruling | Where | +|---|---|---| +| M1: a queued `onDrive` call had no deadline | The review's fix: the `overdue` count, refuse at once and re-mark when every slot is overdue, the slot wait raced against the budget and refused without marking, `refuseWaiters` on every transition to stalled; the named test | `33094c29` | +| M2: nothing protected an idle boot or a hand-typed root | The 15 s health pass armed above the idle gate, the five-minute pass below it; a path on no entry capped by the root it is under; the limit stated above | `bd39579e`, `33094c29`, this commit | +| M3: the snapshot walk was uncapped and 16 wide | Its per-video unit (and its listing and keep-latest reads) through `onDrive(config.dataDir, …)`; the three "at most four threads" sentences (this record's Found and left, FACTS, the `onDrive` header) say what is true after it | `ff235c2f`, `33094c29`, this commit | +| L1: the Storage stage's `statfs` | Through `onDrive`, "—" on a refusal | `19c5842d` | +| L2: discards and flushes; requeued requests | Fields 12 and 16 added to completed; one FACTS line on requeued requests | `33094c29`, this commit | +| L3: stale comments calling the child stat the detector | Fixed (the five named, and the health pass's header) | `33094c29`, `bd39579e`, `19c5842d` | +| L4: "a drive that is not there" for a stalled drive | The pause record carries the cause; both cases worded | `bd39579e` | +| L5: the detector's samples per module copy; "Refresh asks again at once" | Samples on `globalThis`; the changelog's wording | `33094c29`, this commit | +| L6: a spin-up longer than the budget | Slow, not stalled, when the counters moved since the call began; the spin-down question above | `33094c29`, this commit | +| L7: a candidate-root probe shared the location's slots | Keyed by its root | `33094c29` | +| L8: the budget covers a whole unit | Said in the doc and above | `33094c29` | +| L9: `stalled` on no location was not down | Counted as down | `bd39579e` | +| L10: a findmnt every pass | Only on a root change or a failed `/sys` read | `33094c29` | +| The two questions | Keep the four-call cap; keep the 10 s spacing | as built | + ## Rollout