commit d9205b09248b933219751c519a04f5633f32a320
parent 9752ed15228a5708a4c7e3cd247699e9f13bf2de
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Tue, 29 Sep 2026 23:05:55 -0400
plans: slice DS after the parent's rulings — two detectors (the block device's counters every 15 s, a 3 s watchdog on every gated call with four calls per location), the marker before the gate, every memo decider listed, the ports test folded; gates (common 2,280, editor unit 87, test:scripts 191+1, docs clean, e2e storage + channels 38/38); FACTS "The storage health gate" rewritten; the changelog
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Diffstat:
3 files changed, 233 insertions(+), 156 deletions(-)
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -4,7 +4,7 @@
- **Transcripts that arrived after a video was first seen are counted.** The stats behind the homepage, the hub and every site's charts were cached per video and refreshed only when the video's metadata changed, so a transcript that came later — a Whisper run days after the download, or a video downloaded after the last index build — never reached them, and a video with YouTube captions alone had no transcription date. Counts and charts were low; the homepage could show a site with 0 transcripts, 0 channels and 0 hours while it served its videos. A stat is now also redone whenever the index re-reads the video, every transcript has a date, and a captioned video is dated by when its captions arrived rather than by a later Normalize run, so its place on "Transcribed over time" can move. **After updating, rebuild and restart the editor before anything else:** until then, **Build stats dataset** runs the old code and would undo the new stats, while a site, hub or homepage build already runs the new code — and the first stats build of any kind re-reads every video once (about 10–30 minutes on a large archive; it can be stopped and picks up where it stopped). Then build the index, the stats, the homepage, the hub, and the sites.
- **A stats build keeps the stats of a channel whose drive is not mounted, and will not undo a newer version's stats.** A channel whose media is on a drive that is not mounted (or is being moved) is left as it was instead of being read as a channel with no videos; a stats rebuild that has to start over refuses until the drive is back. A stats build refuses to clear stats written by a newer version of the editor; set `ARCHILYZER_STATS_ALLOW_DOWNGRADE=1` to roll back on purpose. Its log also says apart how many videos were downloaded since the last index build (they catch up after the next one) and how many the index skipped (no upload date, or it failed on them).
- **An index build keeps a channel whose drive is not mounted, instead of dropping it from the sites.** **Build index**, a site build's data phase and `archilyzer index` read a channel whose media is on a drive that is not mounted (or is being moved, or whose link and config disagree) as a channel with no videos: they removed its videos from the index, and the next site build published the channel as gone. Such a channel is now left as the last build had it — its videos stay in the index, its pages stay as they were, and the sites built next still list it — and the log names it, with its storage location: one line per channel, ` Held: N channel(s), K video(s) kept.` at the end of the `Diff:` line, and the channels again on the last line. A data folder that fails to read is held the same way, and a channel with no data folder at all is said in the log instead of passed over. An index rebuild that has to start over (after an update that changes the index's format, or with no index yet) refuses while any channel is held and says which; mount the drive first, or set `ARCHILYZER_INDEX_ALLOW_HELD=1` to rebuild without that channel until its drive is back and the index is built again — on the command for a command-line build (`ARCHILYZER_INDEX_ALLOW_HELD=1 pnpm archilyzer index`), or in the editor's own environment, with a restart, for **Build index** and the site builds started from the editor.
-- **A drive that stops answering no longer stops the editor answering.** When a storage location's drive is mounted but not answering (an SMR disk in a USB enclosure resetting under a long write), every page and poll that touched it waited on it, and a few such waits froze the whole editor until the drive came back. The editor now asks each location's drive every 15 seconds from a separate process, with a 3-second limit. A drive that does not answer is marked **Not answering** at once, and the mark comes off after two answers in a row. While it is marked, the editor's pages and polls do not read that drive: `/storage` shows the location as **Not answering** with the time it stopped (**Refresh** asks again at once), the `/channels` volume chip reads "not answering since HH:MM" and its channels' badges "not answering", their videos list, video pages and Cleanup stage say so instead of reading the drive, `/saved-videos` names the channels it did not read, and the index and stats builds keep those channels as they do for an unmounted drive. Jobs for those channels are refused until the drive answers. Pages and polls also reuse each channel's media check for 5 seconds. The editor's `start` script and the container now give Node 16 threads for file access instead of 4 (`UV_THREADPOOL_SIZE`); that buys time for reads already waiting on a drive, and a read that was already waiting when the drive stalled still waits until the drive answers.
+- **A drive that stops answering no longer stops the editor answering.** When a storage location's drive is mounted but not answering (an SMR disk in a USB enclosure resetting under a long write), every page and poll that touched it waited on it, and a few such waits froze the whole editor until the drive came back. Every 15 seconds the editor now reads each location's disk activity counters from the kernel, which never waits on the drive: a disk with requests waiting and none finished since the last look is marked **Not answering**, and the mark comes off after two looks in a row find it working. Where no disk can be named (in a container, say) it asks the drive from a separate process with a 3-second limit instead. Any page or poll that reads the drive also gives up after 3 seconds and marks it the same way, and no more than four such reads wait on one drive at a time. While it is marked, the editor's pages and polls do not read that drive: `/storage` shows the location as **Not answering** with the time it stopped and how it is watched (**Refresh** asks again at once), the `/channels` volume chip reads "not answering since HH:MM" and its channels' badges "not answering", their videos list, video pages and Cleanup stage say so instead of reading the drive, `/saved-videos` names the channels it did not read, and the index and stats builds keep those channels as they do for an unmounted drive, and a channel that is in the middle of a move still shows as moving. Jobs for those channels are refused until the drive answers. Pages and polls also reuse each channel's media check for 5 seconds. The editor's `start` script and the container now give Node 16 threads for file access instead of 4 (`UV_THREADPOOL_SIZE`); that buys time for reads already waiting on a drive, and a read that was already waiting when the drive stalled still waits until the drive answers.
- **Building the homepage now publishes the source: a read-only git mirror, its raw tree and a fresh tarball, behind a gate.** `archilyzer build homepage`, the `/sites` Homepage jobs and `pnpm ops build-homepage` run `archilyzer source publish` between compose and `next build`. It makes a fresh clone of the private `main` (the repository itself is never rewritten), rewrites that copy with git-filter-repo using your scrub rules (file contents and commit messages; your home directory becomes `/home/user` without a rule), and publishes it under `homepage/public` for `git clone https://archilyzer.pages.dev/source/archilyzer.git`, beside `/source/tree/` and the Downloads tarball. Before anything is written, every object of the rewritten history and every file about to be published is searched for every string you have denied; **one hit refuses the build**, and its log names the string only by where you wrote it (`denylist line 3 (len 5)`) and each hit by its object, field and byte offset — never a byte of the object. **A refusal withdraws the source**: the last publish is removed from `homepage/public` and the last build's copy from `homepage/out`, and **Deploy homepage refuses** a build whose source was not audited under today's rules and today's `main` ("run `archilyzer build homepage`, then deploy"). The rules live outside the repo, in `~/.config/archilyzer/source-scrub.txt` and `source-denylist.txt` (`ARCHILYZER_CONFIG_DIR`, `SOURCE_SCRUB_FILE`, `SOURCE_DENYLIST_FILE`); **without them the build refuses**, naming the missing file. **Put everything private in the denylist before any deploy, a preview included**: previews are public, and every deployment stays reachable at its own address until you delete it. Install git-filter-repo once (`pipx install git-filter-repo`; the editor's process needs `~/.local/bin` on its `PATH` to find it) — without it the build fetches it through `pipx run`, which needs the network — and gitleaks if you want its secret scan too. An unchanged `main` with unchanged rules is skipped, so a rebuild costs about 20 seconds only when something moved. A checkout with no git repository (the docker image, a tarball install) builds with the /source page's empty state. `archilyzer source publish --check` audits without writing, `archilyzer source audit <clone>/.git` checks any clone, `archilyzer build homepage --no-source` removes the published source instead, and `archilyzer doctor` reports the tools, the two files (rule counts and permissions, never their contents) and the last publish. `create-archives.sh` is gone. See PUBLISH.md, "The source mirror (homepage)".
- **umtool reads the corpus from its checkout (or `TRANSCRIPTS_DIR`), and the song project's data defaults to `~/.local/share/archilyzer/song`.** If yours is elsewhere, link it there before restarting umtool: `mkdir -p ~/.local/share/archilyzer && ln -s <where the data is> ~/.local/share/archilyzer/song` (the data stays where it is). With no `CHANNELS_DIR`, umtool reads the corpus at `$TRANSCRIPTS_DIR/channels`, else the checkout's own `transcripts/channels`; it used to fall back to an absolute path that existed on one machine only. The song project's videos default to `~/reports/quartering-uh-song/videos`; `SONG_DIR` and `VIDEO_ROOT` still win. The song project's tracked manifests record their paths relative to the song folders, and the twenty one-off `umtool/song/*.sh` run logs, which only ever ran on the machine that wrote them, are gone.
- **umtool's production build no longer reads the corpus folder.** Since umtool began finding the corpus from its checkout (the bullet above), `next build` treated the checkout's whole `transcripts/channels` as files to bundle. On a real archive it ran out of memory and was killed, so umtool could not be rebuilt. The build now ignores that folder and finishes in about 25 s at under 1 GB, the same as a checkout with no corpus. Nothing changes when umtool runs.
diff --git a/plans/FACTS.md b/plans/FACTS.md
@@ -7554,49 +7554,64 @@ this section is stale, by +16 near the top and +203 at the end; they are not rew
## The storage health gate (verified 2026-09-29, branch `r15/drive-stall`)
-The record is [`release-15.md`](release-15.md), "Slice DS, as shipped". Anchors are at the branch tip.
+The record is [`release-15.md`](release-15.md), "Slice DS, as shipped", with the parent's rulings
+Q1–Q5. Anchors are at the branch tip.
- **A drive can be mounted and not answering.** Every in-process fs call on it waits on one of
- libuv's threads (4 by default) until it answers (~30 s for the observed USB reset loop); only a
- child process isolates a call. `probeLocationHealth` (`common/lib/storageVolumes.ts:354`) runs
- `stat -L -c %F -- <root>` as a child raced against `HEALTH_PROBE_TIMEOUT_MS` (3 s); the timer's
- answer is `stalled` and the child is SIGKILLed, never awaited. No binary is "could not ask": `ok`.
+ libuv's threads (4 by default, 16 in the editor's `start`) until it answers (~30 s for the observed
+ USB reset loop); only a child process isolates a call. A child `stat` of a location's ROOT does not
+ detect it reliably: the root's inode is in the kernel's cache whenever the drive was used lately.
- **The state is `common/lib/storageHealth.ts`**, one map on `globalThis.__yttStorageHealth__` (the
- watch writes it from instrumentation's module copy; pages read it from theirs). `recordLocationHealth`
- (`:115`): one miss stalls at once; `HEALTH_CLEAN_TO_CLEAR` (2) clean answers in a row clear it;
- `absent` is clean; a new root starts over. `stalledLocationForPath` (`:188`) matches like
- `locationOfDataDir` (longest root, strictly under); `stalledLocation` (`:204`) is by id AND root.
-- **The cadence** is `runStorageHealthPass` (`controller/storageWatch.ts:376`), every
- `HEALTH_PROBE_INTERVAL_MS` (15 s) from `startStorageWatch` (`:458`), plus one at arm time. It is
- armed only with the watch, so an idle boot has no health state and no gate. `refreshLocationHealth`
- (`:426`) is /storage's Refresh: one answer, nothing pruned. The state is in memory: a CLI process
- (`archilyzer index`) has none, so its inspects are never gated.
-- **The gate** is asked before any fs call that would reach the drive: `inspectChannelMedia`
- (`lib/channelMedia.ts:300`; the gate at `:314`, before the marker read, so a stalled channel
- mid-move reads `stalled`); `probeLocation` (`storageVolumes.ts:241`) and `probeLocationMemo`
- (`:499`, before the memo); `volumeFreeBytes` (`controller/storageLocations.ts:257`);
- `readChannelStat` (`controller/channels.ts:197`); the recency tail reads
- (`controller/recencyIndex.ts:218`, not memoised as misses); `relocationRootPresenceProblem`
- (`controller/relocateChannelMedia.ts:293`); `inspectSavedVideosStore`
- (`controller/relocateSavedVideos.ts:167`); `listSavedVideos` when a page passes `notAnswering`;
- the videos list, the video page, the Cleanup stage (`MediaNotAnswering.tsx`) and the media file
- route (503). `channelMediaStall(config)` (`channelMedia.ts:220`) is the no-I/O question for a
- caller holding a config.
+ watch writes it from instrumentation's module copy; pages read it from theirs), with each entry's
+ `detector`. `recordLocationHealth` (`:143`): one `stalled` answer stalls at once;
+ `HEALTH_CLEAN_TO_CLEAR` (2) clean answers in a row clear it; `absent` is clean; a new root starts
+ over. `registerLocationHealth` (`:199`) creates entries with no answer. `stalledLocationForPath`
+ (`:244`) matches like `locationOfDataDir`; `stalledLocation` (`:260`) is by id AND root.
+- **Detector 1, every 15 s: the block device's counters** (`detectLocationHealth`,
+ `lib/storageVolumes.ts:573`). `findmnt -J -T <root> -o SOURCE,UUID` raced against 3 s (a timeout
+ reuses the last device named for that root; another volume's UUID names none), `[…]` stripped,
+ `/dev/mapper` resolved, basename; then `/sys/class/block/<dev>/stat`: reads completed (1) + writes
+ completed (5), in flight (9). Stalled ⇔ in flight at both samples AND no completion between; the
+ samples at least `MIN_COUNTER_INTERVAL_MS` (10 s) apart; the first gives no verdict. No device →
+ the child `stat -L -c %F` probe (`probeLocationHealth`, `:384`). Detector state is module state in
+ `storageVolumes.ts` (`resetHealthDetector`).
+- **Detector 2, on every gated call: `onDrive(where, call)`** (`storageHealth.ts:410`). Refused with
+ no call on a stalled location; otherwise raced against `DRIVE_CALL_BUDGET_MS` (3 s; test seam
+ `setDriveCallBudget`); a timeout marks the location stalled (since now) and throws
+ `DriveNotAnsweringError` (`isDriveNotAnswering`), leaving the call to settle. At most
+ `DRIVE_CALLS_IN_FLIGHT` (4) calls per location in flight; the rest queue in JS and are refused on
+ a stall; a slot is freed when its call really returns. Do not nest it for one location. A path on
+ no known location is raced but marks nothing; a location object whose root is not its entry's is
+ raced but does not rewrite the entry. The timer is not unref'd (a fake never-settling promise would
+ otherwise let a test process exit).
+- **The cadence** is `runStorageHealthPass` (`controller/storageWatch.ts:407`): prune, register, then
+ every location concurrently; every 15 s from `startStorageWatch` (`:495`), plus one at arm time;
+ armed only with the watch, so an idle boot and a CLI process have no pass (a CLI's inspects are
+ still raced). `refreshLocationHealth` (`:459`) is /storage's Refresh.
+- **The gate order in `inspectChannelMedia`** (`lib/channelMedia.ts:312`): config → memo (a
+ remembered `in-transition` is returned as is, anything else is gated first) → the relocation
+ marker (`:358`, corpus disk) → the gate (`:376`) → the link (corpus disk) → the target's `stat`
+ through `onDrive`. A stall is never memoised. Other gated calls: `probeLocation` (`storageVolumes.ts:253`)
+ and its memo (`:723`); `volumeFreeBytes` (`controller/storageLocations.ts:263`);
+ `readChannelStat` (`controller/channels.ts:215`, the walk through `onDrive`, `null` on a stall);
+ the recency tail reads; the move-root check; the saved-video store; `listSavedVideos` with
+ `notAnswering`; the videos list, the video page, the Cleanup stage and the media file route.
+ `channelMediaStall(config)` (`channelMedia.ts:227`) is the no-I/O question for a holder of a config.
- **`stalled` is a sixth `ChannelMediaStatus` and a sixth `StorageLocationStatus`** ("Not
answering"). `HELD_REASON`, `MediaLocationBadge`'s two tables and `STORAGE_STATUS_LABEL` are the
`Record`s that make tsc name every table a seventh would need. `isMediaHeld` holds it, so both
pool-wide builds hold a stalled channel; the storage watch counts it as down (two passes pause).
-- **`inspectChannelMedia` is memoised for 5 s** (`CHANNEL_MEDIA_MEMO_MS`, `:251`), keyed by channels
- dir, slug and configured `dataDir`, on `globalThis.__yttChannelMediaMemo__`, asked AFTER the gate.
- `{ fresh: true }` skips it and does not store; every decider passes it (the start-of-work guard,
- both movers, both builds, the watch, eviction, the re-point preflight, doctor). The runners' tick
- shares the status poll's `buildChannelWork` and so reads the memo. `forgetChannelMedia` (`:277`) is
- called by the channel mover's marker writes and clear, `clearRelocationMarker`, a re-point,
- /storage's Refresh and the e2e `invalidate-cache` route.
+- **`inspectChannelMedia` is memoised for 5 s** (`CHANNEL_MEDIA_MEMO_MS`, `:261`), keyed by channels
+ dir, slug and configured `dataDir`, on `globalThis.__yttChannelMediaMemo__`. `{ fresh: true }`
+ skips it and does not store; the deciders that pass it are listed in the record (the guard and its
+ six callers, both movers, both builds, the watch, eviction, the re-point preflight, doctor). The
+ runners' tick shares the status poll's `buildChannelWork` and so reads the memo.
+ `forgetChannelMedia` (`:287`) is called by the channel mover's marker writes and clear,
+ `clearRelocationMarker`, a re-point, /storage's Refresh and the e2e `invalidate-cache` route.
- **`UV_THREADPOOL_SIZE`** defaults to 16 in `editor/package.json`'s `start` and in
`docker/entrypoint.sh`. `ports.test.ts` reads every `${NAME:-N}` in a script as a port and names it
- as the one exception (`NUMERIC_NOT_PORTS`).
-- **Not covered:** a call already in flight when the drive stalls; a stall that begins between two
- probes (up to 18 s of calls); a stall the child `stat` does not see because the root's inode is in
- the kernel's cache (the probe then answers in time); jobs already running; per-click server
- actions; the corpus disk itself.
+ as the one exception (`NUMERIC_NOT_PORTS`, `common/lib/ports.test.ts:29`).
+- **Not covered:** a call already in flight when the drive stalls (at most four per location through
+ `onDrive`); a drive already stalled at boot before the second counter sample, unless a page reaches
+ it; jobs already running and their own reads; per-click server actions; the file route's stream;
+ the corpus disk itself.
diff --git a/plans/release-15.md b/plans/release-15.md
@@ -221,79 +221,136 @@ restarted.
Branch `r15/drive-stall` off `main` `ccf90892` (slice IG merged), worktree `~/Projects/r12-paths-fix`
(block #12: editor 4201, test 4211, export 4210), one Opus implementer. Scratch files `ds-*` in the
job's `tmp`. The ruling: a drive that is mounted and not answering must not stop the editor
-answering. The drive is asked out of process, and pages and polls do not touch it in-process while
-it is not answering.
+answering. Two detectors find the stall without the editor waiting on the drive, and pages and polls
+do not touch it in-process while it is not answering. The parent's five rulings on the first pass
+(Q1–Q5, below) are applied.
**What was wrong.** Node runs every filesystem call on libuv's thread pool, four threads by default.
On a drive that has stalled (an SMR disk in a USB enclosure resetting under a long write) each call
blocks its thread for about 30 s. The home page, `/channels` and the three-second auto-queue status
poll each `stat`ted every relocated channel's target in-process and uncached; the one-second job-list
-poll walked the whole `data/` of every channel with a job listed; and the recency layer read
-metadata tails 32 at a time. Four blocked calls were enough for no page, poll or job log to answer.
-The storage watch saw nothing of it: its five-minute pass asks "is the disk here", and a stalled disk
-is here.
-
-- **The health probe** (`common/lib/storageVolumes.ts`, `probeLocationHealth`):
- - `stat -L -c %F -- <root>` as a CHILD PROCESS (execa, `reject: false`), raced against a 3 s timer.
- The timer's answer is `stalled`; the child is sent SIGKILL and not waited for (a process in
- uninterruptible I/O dies only when the I/O returns, and execa's own `timeout` waits for the exit).
- - `directory` on stdout is `ok`; any other answer, or a non-zero exit, is `absent`. A `stat` that
- could not be started is "could not ask": `ok`, never `stalled` (the module's fail-open rule).
+poll walked the whole `data/` of every channel with a job listed, 64 calls wide; and the recency layer
+read metadata tails 32 at a time. Four blocked calls were enough for no page, poll or job log to
+answer. The storage watch saw nothing of it: its five-minute pass asks "is the disk here", and a
+stalled disk is here.
+
- **The health state** (new `common/lib/storageHealth.ts`): one map per process on `globalThis`
- (`__yttStorageHealth__`, the house pattern), by location id: state `ok | stalled | absent`, `since`,
- the last check, a clean streak and the cause. In memory only.
- - One missed probe marks the location `stalled` at once. Two clean answers in a row clear it (an
- `absent` answer is clean: an unmounted drive answers ENOENT at once). A miss in between starts the
- count again. A re-pointed root starts the location over. Locations no longer configured are
- pruned.
+ (`__yttStorageHealth__`, the house pattern: the watch writes it from instrumentation's module copy
+ and pages read it from theirs), by location id: state `ok | stalled | absent`, `since`, the last
+ check, a clean streak, the cause, and the `detector` that gave the last verdict. In memory only.
+ - One `stalled` answer marks the location stalled at once. Two clean answers in a row clear it (an
+ `absent` answer is clean). A miss in between starts the count again. A re-pointed root starts the
+ location over. Locations no longer configured are pruned; every configured one is registered
+ (`registerLocationHealth`) before a pass asks anything, so the watchdog can find a channel's
+ location before the first verdict.
- Pure of I/O and without execa, so `lib/channelMedia.ts` can ask it.
+- **Detector 1, every 15 s: the block device's own counters** (`detectLocationHealth`,
+ `common/lib/storageVolumes.ts`; ruling Q1(a)). A child `stat` of the root is answered from the
+ kernel's inode cache whenever the drive was used lately, so it can say "ok" while the reads that
+ reach the device wait out a reset loop. Instead, per location per pass:
+ - the root's device: `findmnt -J -T <root> -o SOURCE,UUID` as a child raced against 3 s (a findmnt
+ that does not answer reuses the device the last pass named for that root), a `[subvolume]`
+ suffix taken off, `/dev/mapper/*` resolved to its `dm-N` (a read of `/dev`), the basename. A UUID
+ other than the location's recorded one names no device (the root is then a directory on another
+ filesystem, not the drive).
+ - `/sys/class/block/<dev>/stat`, which never touches the drive: reads completed (field 1) + writes
+ completed (5), and requests in flight (9). Against the previous pass's sample for the location
+ (same device, at least `MIN_COUNTER_INTERVAL_MS` = 10 s earlier): **stalled ⇔ in flight at both
+ AND no completion between**; anything else is clean. The first sample gives no verdict.
+ - **No device** (a container, no findmnt, a tmpfs or network source, no `/sys` entry) falls back to
+ the child `stat` probe (`probeLocationHealth`): `stat -L -c %F -- <root>` raced against 3 s, the
+ child SIGKILLed and not waited for; `directory` is `ok`, anything else `absent`, no binary `ok`.
+ - The verdict names its detector (`"counters" | "stat"`), the health state records it, and
+ `/storage`'s line says which watched the drive.
+- **Detector 2, on every gated call: a 3 s watchdog** (`onDrive(where, call)`,
+ `lib/storageHealth.ts`; ruling Q1(b)). The detector that cannot be fooled by a cache: a page or poll
+ that actually reaches the drive finds out.
+ - Refused at once, with no call, when the location is stalled.
+ - Otherwise raced against `DRIVE_CALL_BUDGET_MS` (3 s). A call that has not answered marks its
+ location stalled (since now, cause "a read in the editor did not answer within 3 s") and throws
+ `DriveNotAnsweringError`; the caller answers `stalled`. The call is left to settle on its own:
+ its thread is the stated limit.
+ - **At most `DRIVE_CALLS_IN_FLIGHT` (4) calls per location are in flight through it.** The rest
+ wait in a queue of its own (not libuv's) and are refused without a call the moment the location
+ stalls, so a stalled drive holds at most four of the pool's 16 threads, and a 64-wide walk that
+ meets a stall puts four calls on it, not 64. A slot is released when its call really returns.
+ Calls are not nested for one location; a unit of work (a video directory's few reads, a page's
+ reads of one video) goes through as one call.
+ - A path on no location the health state knows is raced but marks nothing (the error names no
+ location). A probe of another root under a location's id is raced and does not rewrite that
+ location's entry. The timer is not `unref`'d: it is cleared the moment the call answers.
- **The cadence** (`common/controller/storageWatch.ts`): `startStorageWatch` arms a second timer,
- every 15 s, beside the five-minute pass, and runs one health pass at once. All locations are probed
- concurrently; each answer is bounded by the probe's timer, and an overrunning pass is not stacked.
- A transition is logged (`[storage] "<id>": drive not answering — …` / `answering again (ok)`).
- Stopped with the watch. The watch is armed below the idle gate, so an idle boot has no health probe
- and so no gate.
-- **The gate**, before any filesystem call on a stalled location:
-
- | Caller | On a stalled location |
- |---|---|
- | `inspectChannelMedia` (`lib/channelMedia.ts`) | New status `stalled`, detail `drive not answering (location "<label>", since HH:MM)`. Asked right after the configured `dataDir` is known and before the marker, link and target reads, so with a config passed in it makes no call at all (without one, only `config.json` is read). Every page and poll that inspects (home, `/channels`, the channel page, the status poll and the runners' tick, the ops channel route, the bulk actions) and every guard gets it. |
- | `assertChannelMediaReachable` | Refuses it (`ChannelMediaUnreachableError`, status `stalled`), so `runManagedFunction`'s `needsMedia` guard, `generateChannelSnapshot` (also the channel page's **Refresh report** button), the operation batch, normalise, keep-videos and the shard action refuse it. |
- | The index and stats builds | Hold it: `HELD_REASON.stalled` is "its drive is not answering (a stalled disk)", and IG's `isMediaHeld` holds every status but `ok` and `in-place`. |
- | `probeLocation` and `probeLocationMemo` (`lib/storageVolumes.ts`) | New probe status `stalled` (`STORAGE_STATUS_LABEL`: "Not answering"), identity unknown, no free space; no `stat`, `statfs` or `findmnt`. The memo is asked after the gate, so a remembered "available" does not outlive the stall. |
- | `volumeFreeBytes` (`controller/storageLocations.ts`) | Unknown ("—"), with no `stat` or `statfs`. |
- | `readChannelStat` (`controller/channels.ts`) | `null` (no counts, so no progress bar), with no walk of `data/`. The one-second job-list poll and the home page ask it for every channel with a job listed. |
- | The recency tail reads (`controller/recencyIndex.ts`) | Skipped for channels on a stalled location (from `meta[].config`), and NOT remembered as misses: they fall to layers 3 and 4 until a refresh with the drive answering reads them. |
- | `relocationRootPresenceProblem` (`controller/relocateChannelMedia.ts`) | A move onto a stalled location is refused before the root's `stat`. |
- | `inspectSavedVideosStore` (`controller/relocateSavedVideos.ts`) | The store on a stalled location reads `unreachable` with the stall's detail, before the `stat` through its link; `/storage` then skips the store's size walk. |
- | `listSavedVideos` (`controller/savedVideoInventory.ts`), when a page passes `notAnswering` | The channel is skipped and named (its `data/` link is read, not followed). `/saved-videos` names the channels it did not read. The backup job passes nothing and is unchanged. |
- | The videos list, the video page (and its title) and the channel page's Cleanup stage | A notice (`aria-label="media not answering"`, `MediaNotAnswering.tsx`) with links to the channel and `/storage`; nothing is read from the drive. |
- | The media file route (`/api/channels/<slug>/videos/<id>/files/<name>`) | 503 with `Retry-After: 15`. |
+ every 15 s, beside the five-minute pass, and runs one health pass at once. All locations are asked
+ concurrently, each bounded by its own timers, and an overrunning pass is not stacked. A transition
+ is logged (`[storage] "<id>": drive not answering — <cause>; …` / `answering again (ok)`). Stopped
+ with the watch; the watch is armed below the idle gate, so an idle boot has no health pass. A CLI
+ process has no pass either, but its gated calls still go through the watchdog.
+- **The gate and the watchdog, by caller:**
+
+ | Caller | On a stalled location | Through `onDrive` |
+ |---|---|---|
+ | `inspectChannelMedia` (`lib/channelMedia.ts`) | **The relocation marker is read first** (ruling Q2: it is in the channel dir, on the corpus disk), so a channel mid-move on a stalled drive reads `in-transition`. Then the gate: status `stalled`, detail `drive not answering (location "<label>", since HH:MM)`, before the link and the target. With a config in hand a stalled channel costs one call, the marker read. | The target's `stat` (`:441`). |
+ | `assertChannelMediaReachable` | Refuses it (`ChannelMediaUnreachableError`, status `stalled`), so `runManagedFunction`'s `needsMedia` guard, `generateChannelSnapshot` (also the channel page's **Refresh report**), the operation batch, normalise, keep-videos and the shard action refuse it. | Through inspect. |
+ | The index and stats builds | Hold it: `HELD_REASON.stalled` is "its drive is not answering (a stalled disk)", and IG's `isMediaHeld` holds every status but `ok` and `in-place`. | Through inspect (both of the index build's looks). |
+ | The storage watch's five-minute pass | Counts a stalled location as down: two passes auto-pause its channels, and the pass after the drive answers restores them, as for an unmount (ruling Q3). | Through inspect and `probeLocation`. |
+ | `probeLocation` and `probeLocationMemo` (`lib/storageVolumes.ts`) | Probe status `stalled` (`STORAGE_STATUS_LABEL`: "Not answering"), identity unknown, no free space; no `stat`, `statfs` or `findmnt`. The memo is asked after the gate. | The root's `stat` and `statfs` (`:260`, `:277`). |
+ | `volumeFreeBytes` (`controller/storageLocations.ts`) | Unknown ("—"), with no call. | The root's and the mountpoint's `stat`, and the `statfs` (`:288`, `:325`, `:334`). |
+ | `readChannelStat` (`controller/channels.ts`) | `null` (no counts, so no progress bar), with no walk. The one-second job-list poll and the home page ask it for every channel with a job listed. | The `data/` readdir, then each video directory as one call (`:107`); `null` when the drive stops answering mid-walk. The batch jobs' `listChannelStatsFromDisk` passes no drive and is unchanged. |
+ | The recency tail reads (`controller/recencyIndex.ts`) | Skipped, and NOT remembered as misses: layers 3 and 4 until a refresh with the drive answering reads them. | Each tail read of a relocated channel (`:257`). |
+ | `relocationRootPresenceProblem` (`controller/relocateChannelMedia.ts`) | A move onto a stalled location is refused before the root's `stat`. | The root's `stat` (`:307`). |
+ | `inspectSavedVideosStore` (`controller/relocateSavedVideos.ts`) | `unreachable` with the stall's detail; `/storage` skips the store's size walk. | The target's `stat` (`:181`); `/storage`'s store walk too. |
+ | `listSavedVideos` (`controller/savedVideoInventory.ts`), when a page passes `notAnswering` | The channel is skipped and named (its `data/` link is read, not followed). The backup job passes nothing and is unchanged. | Each read of a relocated channel (`:81`). |
+ | The videos list, the video page (and its title), the channel page's Cleanup stage | A notice (`aria-label="media not answering"`, `MediaNotAnswering.tsx`) with links to the channel and `/storage`. | The list's `data/` listing and titles, then the selected video's files; the video page's whole directory read, as one unit; its title; the Cleanup stage's saved-video totals. |
+ | The media file route (`/api/channels/<slug>/videos/<id>/files/<name>`) | 503 with `Retry-After: 15`. | The file's `stat`; the stream after it is not raced. |
- **The memo** (`inspectChannelMedia`): five seconds per channel, keyed by channels dir, slug and
- the configured `dataDir`, on `globalThis` (`__yttChannelMediaMemo__`). The gate is asked before
- it. `{ fresh: true }` bypasses it and does not store; every caller that decides from the answer
- passes it: `assertChannelMediaReachable`, both movers' preconditions, the index build (both
- looks), the stats build, the storage watch, eviction, the re-point preflight and `doctor`.
- `forgetChannelMedia(slug?)` clears it: the channel mover's marker writes and clear (each phase
- writes its marker after the link and config it changes), `clearRelocationMarker`, a completed
- re-point, `/storage`'s Refresh and the e2e `invalidate-cache` route.
+ the configured `dataDir`, on `globalThis` (`__yttChannelMediaMemo__`). On by default (ruling Q4).
+ A remembered `in-transition` is given as it is; any other remembered answer is gated first; a
+ stall is not remembered. `{ fresh: true }` bypasses it and does not store. **The deciders, every
+ one passing `fresh: true`:**
+
+ | Caller | Where |
+ |---|---|
+ | `assertChannelMediaReachable` (so every guard below) | `lib/channelMedia.ts:510` |
+ | ↳ `runManagedFunction`'s `needsMedia` guard | `jobs/streamCommand.ts:283` |
+ | ↳ the operation batch | `controller/operationBatch.ts:1593` |
+ | ↳ `generateChannelSnapshot` | `controller/channelSnapshot.ts:738` |
+ | ↳ `normalizeAllTranscripts` | `controller/normalizeAll.ts:63` |
+ | ↳ keep-videos | `controller/keepVideosMatching.ts:165` |
+ | ↳ the shard action | `editor/app/channels/[slug]/shardActions.ts:88` |
+ | The channel mover, out and back | `controller/relocateChannelMedia.ts:638`, `:877` |
+ | The index build, before and after the walk | `controller/buildIndex.ts:348`, `:467` |
+ | The stats build | `controller/buildStats.ts:286` |
+ | The storage watch's five-minute pass | `controller/storageWatch.ts:233` |
+ | Clip-window eviction | `controller/evictClipWindows.ts:99` |
+ | The re-point preflight | `controller/storageLocations.ts:564` |
+ | `archilyzer doctor` | `bin/doctor.ts:124` |
+
+ **The memo's readers**, pages and polls: the home page (`editor/app/page.tsx:80`), `/channels`
+ (`editor/app/channels/page.tsx:219`), the channel page (`[slug]/page.tsx:235`), the ops channel
+ route (`api/ops/channel/[slug]/route.ts:70`), `channelsOnLocation`'s rollup for `/storage`
+ (`controller/storageLocations.ts:225`), the runners' `buildChannelWork` (`controller/autoRunner.ts:622`,
+ the tick and the three-second status poll share it), and the bulk actions' skips
+ (`editor/app/channels/actions.ts:562`, `bulkStorageActions.ts:105`; the jobs they queue re-check
+ fresh). `forgetChannelMedia(slug?)` clears it: the channel mover's marker writes and clear (each
+ phase writes its marker after the link and config it changes), `clearRelocationMarker`, a
+ completed re-point, `/storage`'s Refresh and the e2e `invalidate-cache` route.
- **Headroom:** `UV_THREADPOOL_SIZE=${UV_THREADPOOL_SIZE:-16}` in `editor/package.json`'s `start`
(which the rollout restart script runs) and in `docker/entrypoint.sh` before the editor's `exec`.
Declared in `envVars.ts` (runtime) and `ENVIRONMENT.md` regenerated. It buys time for calls
- already in flight and isolates nothing; the doc line says so. `ports.test.ts` read every
- `${NAME:-N}` in a script as a port, so it now names `UV_THREADPOOL_SIZE` as the one numeric
- default that is not.
+ already in flight and isolates nothing; the doc line says so. `ports.test.ts` reads every
+ `${NAME:-N}` in a script as a port, so the same commit names `UV_THREADPOOL_SIZE` as the one
+ numeric default that is not (ruling Q5).
- **Where it shows:**
| Surface | What it says |
|---|---|
- | `/storage` | The row's status badge reads **Not answering**; a line under the details reads `not answering since HH:MM — a stat of its root did not answer within 3 s. Pages and polls skip this drive until it answers twice in a row.` (`aria-label="location not answering"`). Re-point and Mount are withheld with the status. **Refresh** asks the location's health first (one answer, counted like the timer's) and forgets the channels' remembered answers; on a stalled drive its note says so instead of probing. |
+ | `/storage` | The row's status badge reads **Not answering**; a line under the details reads `not answering since HH:MM — <cause>. Pages and polls skip this drive until it answers twice in a row. Watched through its disk's request counters.` (or `… with a stat of its root (no disk could be named here).`), `aria-label="location not answering"`. Re-point and Mount are withheld with the status. **Refresh** asks the location's detector first (one answer, counted like the pass's; the counters give none within 10 s of the last sample) and forgets the channels' remembered answers; on a stalled drive its note says so instead of probing. |
| `/channels` | The volume chip reads `… · not answering since HH:MM` in place of its free space, with a title saying what it means. Each row's badge reads `on <label> — not answering` (accessible name `media location: Media not answering · on <label>`). |
| The channel page | The Storage stage's card is red with the inspector's sentence; its destination list names the location "Not answering"; Move back is withheld for a stalled channel, as for an unreachable one. |
- | `/saved-videos` | `Not read, because the drive their media is on is not answering: <slugs>.` (`aria-label="saved videos not read"`), above the per-channel table. |
+ | `/saved-videos` | `Not read, because the drive their media is on is not answering: <slugs>.` (`aria-label="saved videos not read"`). |
| The runners | `[auto] skipping <slug>: media stalled — drive not answering (…)`, once per state change. |
- | The logs | The health pass's transition lines; the index and stats builds' hold lines. |
+ | The logs | The health pass's transition lines, with the cause; the index and stats builds' hold lines. |
Labels are contracts: no existing accessible name or test id changed.
@@ -301,47 +358,54 @@ is here.
| Commit | What |
|---|---|
-| `c0ecc55c` | `common:` `lib/storageHealth.ts`, `probeLocationHealth`, the 15 s health pass, the `stalled` statuses and the gate in `inspectChannelMedia`, `probeLocation`/its memo, `volumeFreeBytes`, `readChannelStat`, the recency tail reads, the move-root check and the saved-video store; `HELD_REASON.stalled`; the 5 s memo with `fresh` for every decider and `forgetChannelMedia` in the movers; the badge's and the stage card's words. Tests: `storageHealth`, `storageHealthProbe`, `storageStall`, `storageWatch`. |
+| `c0ecc55c` | `common:` `lib/storageHealth.ts`, `probeLocationHealth`, the 15 s health pass, the `stalled` statuses and the gate in inspect, the probe and its memo, `volumeFreeBytes`, `readChannelStat`, the recency reads, the move-root check and the saved-video store; `HELD_REASON.stalled`; the 5 s memo with `fresh` for every decider and `forgetChannelMedia` in the movers; the badge's and the stage card's words. Tests. |
| `cfe551de` | `editor:` `/storage` (the row's line, Refresh asks the health first), the `/channels` volume chip, the Storage panel's Move back, the videos list and video page notice, the media file route's 503, the e2e `invalidate-cache` route; `refreshLocationHealth`; `views/storage.ts` `notAnswering`. |
-| `6a74c790` | `editor:` `UV_THREADPOOL_SIZE=16` in the editor's `start` and `docker/entrypoint.sh`; `envVars.ts` + `ENVIRONMENT.md`. On its own this commit fails `ports.test.ts` (next row). |
-| `de4128d1` | `editor:` `listSavedVideos({ notAnswering })` for `/saved-videos` and the Cleanup stage's notice; `ports.test.ts` names `UV_THREADPOOL_SIZE` as a numeric default that is not a port. |
-| this commit | `plans:` this section; FACTS "The storage health gate"; the editor changelog. |
+| `480f2556` | `editor:` `UV_THREADPOOL_SIZE=16` in the editor's `start` and `docker/entrypoint.sh`; `envVars.ts` + `ENVIRONMENT.md`; `ports.test.ts` names it as a numeric default that is not a port (folded in from the next commit, ruling Q5; no change to the tree at the tip). |
+| `4e1f7a90` | `editor:` `listSavedVideos({ notAnswering })` for `/saved-videos` and the Cleanup stage's notice. |
+| `63f42ef5` | `plans:` the first version of this section, FACTS, the changelog. |
+| `c69ad41a` | `common:` rulings Q2 and Q1(b): the marker before the gate; `onDrive` (the 3 s watchdog and the four-call cap per location) on every common gated call; `registerLocationHealth`. Tests. |
+| `f6a25cf5` | `editor:` the pages' reads of a relocated drive through `onDrive` (the videos list, the video page and its title, the file route, the Cleanup totals, `/storage`'s store walk). |
+| `b1a30902` | `common:` ruling Q1(a): `detectLocationHealth`, the block-device counters with the child-stat fallback; `detector` in the health state and on `/storage`. Tests. |
+| this commit | `plans:` this section rewritten for the rulings; FACTS; the changelog; the report. |
-**Tests** (unit; no test stalls a real drive: a stalled `stat` is a fake binary that never answers,
-and the gated callers run against an ordinary temp directory the health state is told is stalled)
+**Tests** (unit; no test stalls a real drive: a stalled call is a promise that never settles or a
+fake `stat` that never answers, and a stalled device is a temp `/sys` whose counters stand still)
| File | What it pins |
|---|---|
-| `lib/storageHealth.test.ts` (10) | One miss stalls at once; one clean answer after a stall does not clear it and two in a row do; a miss in between starts the count again; `absent` is clean; a re-pointed root starts over; the path match is `locationOfDataDir`'s (longest root, strictly under); pruning; the map is on `globalThis`; the "since" wording. |
-| `lib/storageHealthProbe.test.ts` (5) | The real `stat`: a directory is `ok`, a missing path and a file are `absent`. A fake `stat` asleep for 20 s is `stalled` on a 400 ms timer in under 3 s, so nothing waited for the child; the default budget is 3 s (answered after 3–6 s). An answer inside the budget is taken as given; a binary that cannot start is `ok`. |
-| `controller/storageStall.test.ts` (14) | A spy on every `node:fs` and `node:fs/promises` call (the `buildIndex.test.ts` technique, reads included). With the location stalled: `inspectChannelMedia` with a config makes NO call, without one only `config.json`; the guard refuses with no call; `probeLocation` and its memo answer `stalled` with no call on the drive and without running `findmnt`; `volumeFreeBytes`, `readChannelStat`, the recency layer, the move-root check, a snapshot refresh and the saved-video store make no call on the drive; the inventory skips and names the channel (proved by what comes back: `fs-extra` bound its functions before the spy). The memo: two inspects inside 5 s stat the target once, `fresh` and a 5 s age each ask again and a fresh answer is not stored; another target is another key; `forgetChannelMedia` and `clearRelocationMarker` clear it; a stall is seen with an `ok` remembered. |
-| `controller/storageWatch.test.ts` (+6) | The health pass stalls a location on one miss and a page then gets `stalled` without asking; the five-minute pass suspects its channel; the stall clears only after two clean passes in a row; a location no longer configured is forgotten; a probe that throws is `ok`; arming the watch runs one health pass at once and stopping it stops both timers; a Refresh counts as one answer and prunes nothing. |
-| `views/storage.test.ts` (+1) | A stalled row reads "Not answering", carries its line, has no free space, and withholds Re-point and Mount; a row with no entry carries nothing. |
+| `lib/storageHealth.test.ts` (19) | The rules: one miss stalls at once, one clean answer after a stall does not clear it and two in a row do, a miss in between starts over, `absent` is clean, a re-pointed root starts over, the path match is `locationOfDataDir`'s, pruning, `globalThis`, the "since" wording, registering without an answer. `onDrive`: an answer passes through (value or error) and frees its slot; a never-settling call is `stalled` on the timer, marks the location (since now) and keeps its slot until it settles; a stalled location is refused with no call; seven calls at once put four in flight and the three that waited are refused without a call when the location stalls; a freed slot runs a waiter; an unknown path marks nothing; a candidate root does not rewrite its location; the 3 s default (answered after 3–4.5 s). |
+| `lib/storageHealthCounters.test.ts` (7) | The stat line parser (17 and 11 fields, garbage); the verdict over sample pairs (stuck → stalled; moving, idle or drained → ok); device names (partition, `[subvolume]` suffix, non-`/dev` sources); the detector end to end with a fake findmnt and a temp `/sys`: the first sample gives no verdict, stuck → stalled with its cause, drained → ok, busy and moving → ok; a sample sooner than 10 s gives none and keeps the first; every no-device fallback goes to the child stat and says so (tmpfs, findmnt failing, no `/sys` entry, another volume's UUID, no binary); a findmnt that does not answer is not waited for and the last device named is read. |
+| `lib/storageHealthProbe.test.ts` (5) | The child `stat`: a directory is `ok`, a missing path and a file are `absent`; a fake `stat` asleep for 20 s is `stalled` on the timer without being waited for; the 3 s default; an answer inside the budget is taken; no binary is `ok`. |
+| `controller/storageStall.test.ts` (21) | A spy on every `node:fs` and `node:fs/promises` call, with a hang mode that makes a matching promise-API call never settle. The gate: with the location stalled, inspect reads only the marker (with a config) or config.json and the marker (without); a channel mid-move on a stalled drive reads `in-transition` (and is remembered so); the guard, `probeLocation` and its memo (no findmnt run), `volumeFreeBytes`, `readChannelStat`, the recency layer, the move-root check, a snapshot refresh, the saved-video store and the inventory make no call on the drive. The watchdog: with the location answering and one drive call hung, inspect, `readChannelStat` (at most four video directories asked), `probeLocation`, `volumeFreeBytes`, the recency layer (not remembered as a miss) and the move-root check each answer `stalled` within the race and mark the location; after it, inspect, the guard and the walk make no call on the drive, and once cleared the drive is asked again. The memo: two inspects inside 5 s stat the target once, `fresh` and a 5 s age ask again, a fresh answer is not stored, another target is another key, `forgetChannelMedia` and `clearRelocationMarker` clear it, and a stall is seen with an `ok` remembered. |
+| `controller/storageWatch.test.ts` (+7) | The pass stalls a location on one miss and a page then gets `stalled`; the five-minute pass suspects its channel; two clean passes clear it; a location no longer configured is forgotten; a probe that throws is `ok`; arming runs one pass at once and stopping stops both timers; a Refresh counts as one answer; the pass registers every location, records a verdict's detector, and a verdict with no answer changes nothing. |
+| `views/storage.test.ts` (+1) | A stalled row reads "Not answering", carries its line, has no free space, and withholds Re-point and Mount. |
#### Gates (logs `$T/ds-*.log`)
-- **tsc** was clean before every commit: 186 s at `c0ecc55c`, 86 s before `cfe551de`, 176 s at
- `6a74c790`, 130 s at `de4128d1` (the machine was running other slices' suites).
-- **Unit:**
+- **tsc** was clean before every commit. After the rulings: 73 s at `c69ad41a`/`f6a25cf5`, 261 s at
+ `b1a30902` (the machine was running other sessions' suites).
+- **Unit, at `b1a30902`:**
| Suite | Result |
|---|---|
- | common | **2,256/2,256** at the tip, 71 s (the branch point's 2,220 plus 36 new). At `6a74c790` it was 2,254/2,255: `ports.test.ts`, fixed in `de4128d1`. |
+ | common | **2,280/2,280**, 101 s (the branch point's 2,220 plus 60 new). 2,256 at the first pass. |
| editor unit | 87/87 |
- | `test:scripts` | 191 passed, 1 skipped (192) at the tip. A first run beside the e2e suite failed `queue-lock.test.mjs:85` once (the holder's details were not yet readable); it passed on the rerun. |
- | mcp | 271/271 |
+ | `test:scripts` | 191 passed, 1 skipped (192), on two reruns. The first run beside another session's e2e failed `queue-lock.test.mjs:85` (the holder's details were not yet readable), as it did once on the first pass; it passes alone. |
+ | mcp | 271/271 (first pass; no mcp file changed since) |
- **Docs:** `docs env --check`, `docs files --check` and `settings example --check` all exit **0**.
-- **Build:** the editor's `next build`, with the primary's `transcripts/` linked in and capped at
- 5 GB with no swap: 104 s, max RSS 1,643,860 KB. The link was removed after the build, and nothing
- ran through it.
+- **Build** (first pass, at `de4128d1`, now `4e1f7a90`): the editor's `next build`, with the
+ primary's `transcripts/` linked in and capped at 5 GB with no swap: 104 s, max RSS 1,643,860 KB.
+ The link was removed after the build, and nothing ran through it. Not rerun after the rulings (not
+ in the re-gate list).
- **e2e** (editor, detached and queued):
| Run | Specs | Result |
|---|---|---|
- | 1, at `6a74c790` | `storage-locations`, `channel-storage`, `channels-storage-columns`, `channels`, `channels-actions`, `channels-counts`, `channels-sort`, `channels-rack-layers`, `channels-rack-audit` (skips without `E2E_RACK_SHOTS`), `video-page`, `saved-videos`, `dashboard`, `auto-queue`, and IG's eight (`availability`, `build`, `channel-build-toggle`, `chat-only`, `duplicate-shorts`, `jobs`, `regional-vtt-fallback`, `tags`) — `$T/ds-specs.txt` | **124 passed, 3 failed, 12 skipped, 16.6 min** (36 s in the queue). The three: `auto-queue.spec.ts:218` and `:352`, both 30 s timeouts in the first minutes while this slice's own tsc and common suite were running beside it, and `tags.spec.ts:463` (a `/sites` overlay field not yet editable at 30 s). |
- | 2, at `de4128d1` | `auto-queue`, `tags`, `saved-videos`, `cleanup-holds`, `cleanup-actionable`, `cleanup-page`, `do-not-clean`, `video-titles`, `video-page` — `$T/ds-specs2.txt` | **70 passed, 1 failed, 5.9 min** (56 s in the queue). `tags.spec.ts:463` and `auto-queue.spec.ts:218` passed. The one: `auto-queue.spec.ts:352` ("UI: build a policy…"), `locator.click` on **Start Auto-transcribe** waiting for an enabled button: the spec's own comment names the race (the three-second poll disables Start after `isEnabled()` answered true); earlier records list it as a known flake (`auto-queue.spec.ts:411` at the time). |
- | 3, at `de4128d1` | `auto-queue` alone | **24 passed, 1.3 min** (1 min 59 s in the queue), `:352` included. |
+ | 1, at `6a74c790` (now `480f2556`) | the `storage` and `channels` specs, `video-page`, `saved-videos`, `dashboard`, `auto-queue`, and IG's eight (`$T/ds-specs.txt`) | **124 passed, 3 failed, 12 skipped, 16.6 min**. The three were 30 s timeouts while this slice's own tsc and common suite ran beside the suite (`auto-queue.spec.ts:218` and `:352`, `tags.spec.ts:463`). |
+ | 2, at `de4128d1` (now `4e1f7a90`) | `auto-queue`, `tags`, `saved-videos`, four cleanup specs, `video-titles`, `video-page` (`$T/ds-specs2.txt`) | **70 passed, 1 failed, 5.9 min**: `auto-queue.spec.ts:352`, the Start-button race its own comment describes (a known flake). |
+ | 3, at `de4128d1` | `auto-queue` alone | **24 passed, 1.3 min**. |
+ | 4, at `b1a30902` (the re-gate) | `storage-locations`, `channel-storage`, `channels-storage-columns`, `channels`, `channels-actions`, `channels-counts`, `channels-sort`, `channels-rack-layers`, `channels-rack-audit` (`$T/ds-specs4.txt`) | **38 passed, 0 failed, 12 skipped, 2.5 min** (the 12 are `channels-rack-audit`, which skips without `E2E_RACK_SHOTS`). |
No e2e fixture has a stalled drive: these confirm nothing changed for drives that answer. The
stall paths are the unit tests above.
@@ -349,60 +413,58 @@ and the gated callers run against an ordinary temp directory the health state is
#### Found and left
-- **The stated limit.** A read already in flight when the drive stalls still holds its thread until
- the kernel gives up (about 30 s in the observed reset loop), and a stall that begins just after a
- probe costs every call that reaches the drive until the next one (up to 15 s, plus the 3 s
- budget). With the gate and 16 threads the editor keeps answering meanwhile.
-- **A `stat` the kernel answers from its cache does not reach the drive.** The probe stats the
- location's ROOT, whose inode is in the kernel's cache whenever anything has used the drive
- recently; a stall that only uncached reads hit (a walk of `data/`, a metadata tail) can leave the
- probe answering in time and the location `ok`. The probe then detects a stall only once the
- root's own metadata has to come off the device. Three ways to reach the device instead, none built:
- an `O_DIRECT` read of a known file on the drive (`dd iflag=direct`), the block device's in-flight
- and completed counters in `/sys/class/block/<dev>/stat` (a stall is requests in flight and none
- completing), or a watchdog on the gated in-process calls themselves (a call that has not answered
- in 3 s marks its location stalled). Left for a ruling (question 1 in the report).
-- **Request paths left ungated**, each a click rather than a page or poll: the channel and video
- server actions that read a video's directory in the request (`bulkVideoActions`,
- `digestActions`, `videoActions`, `fixIncompleteTranscript`, `pipelineActions`); most of what they
- do is enqueue jobs, whose guard is fresh and refuses a stalled channel; `/api/media/fetch-window/<jobId>` (one `stat` of the fetched file, once the job is done);
- `/api/ops/channel/<slug>` is gated through `inspectChannelMedia` and `readChannelStat` (its counts
- are `null` on a stall).
-- **Jobs are gated only at their start.** A job already running when its drive stalls (the remux
- that stalls it, for one) keeps its threads and children; a job's own in-process reads
- (`measureTree`, the index build's processing phase, which IG recorded) block their own threads.
+- **The stated limit.** A read already in flight when the drive stalls holds its thread until the
+ kernel gives up (about 30 s in the observed reset loop). With the watchdog and the four-call cap, a
+ stall that the counters have not yet seen costs at most four threads per location for that long,
+ and the pages and polls asking answer within 3 s; with 16 threads the editor keeps answering.
+- **The counters need two samples.** A drive already stalled when the editor starts is seen by the
+ counters at the second pass (15–30 s), or at once by the watchdog when a page reaches it.
+- **Units of work count as one slot.** A video directory's reads, or a page's reads of one video, go
+ through as one call, so a slot can hold a few sequential calls, and the video page's parallel
+ reads of one directory run under one slot.
+- **Ungated request paths**, each a click rather than a page or poll: the channel and video server
+ actions that read a video's directory in the request (`bulkVideoActions`, `digestActions`,
+ `videoActions`, `fixIncompleteTranscript`, `pipelineActions`); most of what they do is enqueue
+ jobs, whose guard is fresh and refuses a stalled channel. `/api/media/fetch-window/<jobId>` (one
+ `stat` of the fetched file, once the job is done). The media file route's stream after its `stat`.
+- **Jobs are gated only at their start.** A job already running when its drive stalls (the remux that
+ stalls it, for one) keeps its threads and children; a job's own in-process reads (`measureTree`,
+ the index build's processing phase, which IG recorded) are not raced.
- **The runners' tick reads the memo.** `buildChannelWork` serves both the tick and the status poll,
so a drive unmounted in the last 5 s can have one unit dispatched, which fails at the dangling
- link. The snapshot regeneration and `runManagedFunction`'s guard are fresh. The bulk actions
- (Sync all, the bulk move) decide their skips from the memo too; the jobs they queue re-check.
-- **The storage watch counts a stalled location as down:** two five-minute passes auto-pause its
- channels (the lanes already skip them from the first stalled tick), and the first watch pass
- after the health clears restores them.
+ link. The snapshot regeneration and `runManagedFunction`'s guard are fresh.
- **umtool's twin of the reachability check** (`checkChannelReachable` in
`umtool/report-to-video/cues.mjs`) has no stall gate: umtool is its own process with no health
state, and `umtool/**` belongs to another slice.
-- **An idle boot has no health probe**, so nothing is gated there; a CLI run (`archilyzer index`)
- has none either. The corpus disk itself is not probed.
-- **`6a74c790` alone fails `ports.test.ts`**; `de4128d1` fixes the test. They merge together.
+- **A CLI process** (`archilyzer index`) has no health pass, so nothing is gated in it; its
+ inspects still go through the watchdog, so a target `stat` that takes over 3 s holds the channel.
+ The corpus disk itself is not watched.
-#### Decisions the operator could overturn
+#### Rulings (parent, 2026-09-29) and decisions the operator could overturn
+
+| Question | Ruling | Where |
+|---|---|---|
+| Q1: the child stat of the root can be answered from the cache | Two detectors: the block device's counters every 15 s (child stat only with no device, and the state says which), and a 3 s watchdog on every gated call | `b1a30902`, `c69ad41a`, `f6a25cf5` |
+| Q2: the gate before or after the marker | The marker first; the gate covers everything after it | `c69ad41a` |
+| Q3: a stall auto-pauses | Yes, after two five-minute passes; it clears when the drive answers again, as for an unmount | as built |
+| Q4: the memo on by default | As built; every decider listed above | as built |
+| Q5: a commit that failed `ports.test.ts` alone | The test line folded into the thread-pool commit | `480f2556` |
| What I assumed | The alternative |
|---|---|
-| The gate is asked before the relocation marker, so a stalled channel mid-move reads `stalled`, not `in-transition` (a resumed move accepts any status; a new move refuses). With a config in hand it costs no call. | Read the marker first (one read on the corpus disk) and let `in-transition` win. |
-| A stall auto-pauses like an unmounted drive, on the watch's own two-pass cadence. | Exclude `stalled` from the watch's "down", so only the lanes' per-tick skips apply. |
+| At most four gated calls per location in flight; the rest queue in JavaScript and are refused on a stall. | No cap: the watchdog alone, and a stall mid-walk fills the pool until the kernel gives up. |
+| Two counter samples closer than 10 s give no verdict (a Refresh just after a pass among them). | Compare any two samples (a busy healthy drive can read "in flight, nothing completed" over a few milliseconds). |
+| A findmnt that does not answer reuses the last device named for that root; one naming another volume's UUID names none. | Treat a findmnt that does not answer as a stall. |
| An `absent` answer counts as clean toward clearing a stall. | Only `ok` clears it. |
-| The memo is on by default and every decider passes `fresh`. | Off by default, with pages and polls opting in. |
-| `/storage`'s Refresh is one health answer like the timer's. | Refresh clears a stall outright on one clean answer. |
+| `/storage`'s Refresh is one answer like the pass's. | Refresh clears a stall outright on one clean answer. |
| `UV_THREADPOOL_SIZE` defaults to 16 in `start` and the entrypoint; a value already set wins. | A fixed 16, or only in the rollout restart script. |
-| The videos list, the video page and the Cleanup stage show a notice instead of their content on a stalled drive. | Render what the snapshot knows and leave out only the drive's files. |
+| The videos list, the video page and the Cleanup stage show a notice instead of their content. | Render what the snapshot knows and leave out only the drive's files. |
| `readChannelStat` returns `null` on a stall (the job row draws no progress bar). | Counts marked unknown. |
-**What runs which code, for the rollout.** The health state lives in the editor's process, armed
-with the storage watch, so all of it takes effect only when the editor is rebuilt and restarted
-(and not on an idle boot). The restart must go through `pnpm run start` in `editor/` (the rollout
-restart script does) for `UV_THREADPOOL_SIZE` to apply; a process started another way keeps
-Node's 4 threads unless the variable is set. CLI builds have no health state: they hold a channel
-only for the statuses they already held.
+**What runs which code, for the rollout.** The health pass lives in the editor's process, armed with
+the storage watch, so detection takes effect only when the editor is rebuilt and restarted (and not
+on an idle boot). The restart must go through `pnpm run start` in `editor/` (the rollout restart
+script does) for `UV_THREADPOOL_SIZE` to apply; a process started another way keeps Node's 4 threads
+unless the variable is set. CLI builds have no health state; their inspects are raced.
## Rollout