commit e77ad9eb02faf162f24c95bb14f702187e308cac
parent 34eb213fc52676cccf4734cb18c0e13841b46204
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Tue, 29 Sep 2026 22:38:26 -0400
plans: slice DS as shipped — the storage health probe and its gate, the memo, the thread pool; gates (common 2,256, editor unit 87, test:scripts 191+1, mcp 271, docs clean, editor build 104 s / 1,643,860 KB, e2e 124+3 flaky / 70+1 flaky / 24); found and left (the root's inode cache, the ungated clicks, jobs in flight); FACTS "The storage health gate"; the editor changelog
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Diffstat:
3 files changed, 240 insertions(+), 1 deletion(-)
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -4,6 +4,7 @@
- **Transcripts that arrived after a video was first seen are counted.** The stats behind the homepage, the hub and every site's charts were cached per video and refreshed only when the video's metadata changed, so a transcript that came later — a Whisper run days after the download, or a video downloaded after the last index build — never reached them, and a video with YouTube captions alone had no transcription date. Counts and charts were low; the homepage could show a site with 0 transcripts, 0 channels and 0 hours while it served its videos. A stat is now also redone whenever the index re-reads the video, every transcript has a date, and a captioned video is dated by when its captions arrived rather than by a later Normalize run, so its place on "Transcribed over time" can move. **After updating, rebuild and restart the editor before anything else:** until then, **Build stats dataset** runs the old code and would undo the new stats, while a site, hub or homepage build already runs the new code — and the first stats build of any kind re-reads every video once (about 10–30 minutes on a large archive; it can be stopped and picks up where it stopped). Then build the index, the stats, the homepage, the hub, and the sites.
- **A stats build keeps the stats of a channel whose drive is not mounted, and will not undo a newer version's stats.** A channel whose media is on a drive that is not mounted (or is being moved) is left as it was instead of being read as a channel with no videos; a stats rebuild that has to start over refuses until the drive is back. A stats build refuses to clear stats written by a newer version of the editor; set `ARCHILYZER_STATS_ALLOW_DOWNGRADE=1` to roll back on purpose. Its log also says apart how many videos were downloaded since the last index build (they catch up after the next one) and how many the index skipped (no upload date, or it failed on them).
- **An index build keeps a channel whose drive is not mounted, instead of dropping it from the sites.** **Build index**, a site build's data phase and `archilyzer index` read a channel whose media is on a drive that is not mounted (or is being moved, or whose link and config disagree) as a channel with no videos: they removed its videos from the index, and the next site build published the channel as gone. Such a channel is now left as the last build had it — its videos stay in the index, its pages stay as they were, and the sites built next still list it — and the log names it, with its storage location: one line per channel, ` Held: N channel(s), K video(s) kept.` at the end of the `Diff:` line, and the channels again on the last line. A data folder that fails to read is held the same way, and a channel with no data folder at all is said in the log instead of passed over. An index rebuild that has to start over (after an update that changes the index's format, or with no index yet) refuses while any channel is held and says which; mount the drive first, or set `ARCHILYZER_INDEX_ALLOW_HELD=1` to rebuild without that channel until its drive is back and the index is built again — on the command for a command-line build (`ARCHILYZER_INDEX_ALLOW_HELD=1 pnpm archilyzer index`), or in the editor's own environment, with a restart, for **Build index** and the site builds started from the editor.
+- **A drive that stops answering no longer stops the editor answering.** When a storage location's drive is mounted but not answering (an SMR disk in a USB enclosure resetting under a long write), every page and poll that touched it waited on it, and a few such waits froze the whole editor until the drive came back. The editor now asks each location's drive every 15 seconds from a separate process, with a 3-second limit. A drive that does not answer is marked **Not answering** at once, and the mark comes off after two answers in a row. While it is marked, the editor's pages and polls do not read that drive: `/storage` shows the location as **Not answering** with the time it stopped (**Refresh** asks again at once), the `/channels` volume chip reads "not answering since HH:MM" and its channels' badges "not answering", their videos list, video pages and Cleanup stage say so instead of reading the drive, `/saved-videos` names the channels it did not read, and the index and stats builds keep those channels as they do for an unmounted drive. Jobs for those channels are refused until the drive answers. Pages and polls also reuse each channel's media check for 5 seconds. The editor's `start` script and the container now give Node 16 threads for file access instead of 4 (`UV_THREADPOOL_SIZE`); that buys time for reads already waiting on a drive, and a read that was already waiting when the drive stalled still waits until the drive answers.
- **Building the homepage now publishes the source: a read-only git mirror, its raw tree and a fresh tarball, behind a gate.** `archilyzer build homepage`, the `/sites` Homepage jobs and `pnpm ops build-homepage` run `archilyzer source publish` between compose and `next build`. It makes a fresh clone of the private `main` (the repository itself is never rewritten), rewrites that copy with git-filter-repo using your scrub rules (file contents and commit messages; your home directory becomes `/home/user` without a rule), and publishes it under `homepage/public` for `git clone https://archilyzer.pages.dev/source/archilyzer.git`, beside `/source/tree/` and the Downloads tarball. Before anything is written, every object of the rewritten history and every file about to be published is searched for every string you have denied; **one hit refuses the build**, and its log names the string only by where you wrote it (`denylist line 3 (len 5)`) and each hit by its object, field and byte offset — never a byte of the object. **A refusal withdraws the source**: the last publish is removed from `homepage/public` and the last build's copy from `homepage/out`, and **Deploy homepage refuses** a build whose source was not audited under today's rules and today's `main` ("run `archilyzer build homepage`, then deploy"). The rules live outside the repo, in `~/.config/archilyzer/source-scrub.txt` and `source-denylist.txt` (`ARCHILYZER_CONFIG_DIR`, `SOURCE_SCRUB_FILE`, `SOURCE_DENYLIST_FILE`); **without them the build refuses**, naming the missing file. **Put everything private in the denylist before any deploy, a preview included**: previews are public, and every deployment stays reachable at its own address until you delete it. Install git-filter-repo once (`pipx install git-filter-repo`; the editor's process needs `~/.local/bin` on its `PATH` to find it) — without it the build fetches it through `pipx run`, which needs the network — and gitleaks if you want its secret scan too. An unchanged `main` with unchanged rules is skipped, so a rebuild costs about 20 seconds only when something moved. A checkout with no git repository (the docker image, a tarball install) builds with the /source page's empty state. `archilyzer source publish --check` audits without writing, `archilyzer source audit <clone>/.git` checks any clone, `archilyzer build homepage --no-source` removes the published source instead, and `archilyzer doctor` reports the tools, the two files (rule counts and permissions, never their contents) and the last publish. `create-archives.sh` is gone. See PUBLISH.md, "The source mirror (homepage)".
- **umtool reads the corpus from its checkout (or `TRANSCRIPTS_DIR`), and the song project's data defaults to `~/.local/share/archilyzer/song`.** If yours is elsewhere, link it there before restarting umtool: `mkdir -p ~/.local/share/archilyzer && ln -s <where the data is> ~/.local/share/archilyzer/song` (the data stays where it is). With no `CHANNELS_DIR`, umtool reads the corpus at `$TRANSCRIPTS_DIR/channels`, else the checkout's own `transcripts/channels`; it used to fall back to an absolute path that existed on one machine only. The song project's videos default to `~/reports/quartering-uh-song/videos`; `SONG_DIR` and `VIDEO_ROOT` still win. The song project's tracked manifests record their paths relative to the song folders, and the twenty one-off `umtool/song/*.sh` run logs, which only ever ran on the machine that wrote them, are gone.
- **umtool's production build no longer reads the corpus folder.** Since umtool began finding the corpus from its checkout (the bullet above), `next build` treated the checkout's whole `transcripts/channels` as files to bundle. On a real archive it ran out of memory and was killed, so umtool could not be rebuilt. The build now ignores that folder and finishes in about 25 s at under 1 GB, the same as a checkout with no corpus. Nothing changes when umtool runs.
diff --git a/plans/FACTS.md b/plans/FACTS.md
@@ -7551,3 +7551,52 @@ this section is stale, by +16 near the top and +203 at the end; they are not rew
no digest, which is removed (`digests.remove`, `:962`/`:965`). `mtimes` is then written with the
video's current mtimes, so the loss lasts until any of its tracked mtimes (metadata, transcript,
subs, availability, digest) moves.
+
+## The storage health gate (verified 2026-09-29, branch `r15/drive-stall`)
+
+The record is [`release-15.md`](release-15.md), "Slice DS, as shipped". Anchors are at the branch tip.
+
+- **A drive can be mounted and not answering.** Every in-process fs call on it waits on one of
+ libuv's threads (4 by default) until it answers (~30 s for the observed USB reset loop); only a
+ child process isolates a call. `probeLocationHealth` (`common/lib/storageVolumes.ts:354`) runs
+ `stat -L -c %F -- <root>` as a child raced against `HEALTH_PROBE_TIMEOUT_MS` (3 s); the timer's
+ answer is `stalled` and the child is SIGKILLed, never awaited. No binary is "could not ask": `ok`.
+- **The state is `common/lib/storageHealth.ts`**, one map on `globalThis.__yttStorageHealth__` (the
+ watch writes it from instrumentation's module copy; pages read it from theirs). `recordLocationHealth`
+ (`:115`): one miss stalls at once; `HEALTH_CLEAN_TO_CLEAR` (2) clean answers in a row clear it;
+ `absent` is clean; a new root starts over. `stalledLocationForPath` (`:188`) matches like
+ `locationOfDataDir` (longest root, strictly under); `stalledLocation` (`:204`) is by id AND root.
+- **The cadence** is `runStorageHealthPass` (`controller/storageWatch.ts:376`), every
+ `HEALTH_PROBE_INTERVAL_MS` (15 s) from `startStorageWatch` (`:458`), plus one at arm time. It is
+ armed only with the watch, so an idle boot has no health state and no gate. `refreshLocationHealth`
+ (`:426`) is /storage's Refresh: one answer, nothing pruned. The state is in memory: a CLI process
+ (`archilyzer index`) has none, so its inspects are never gated.
+- **The gate** is asked before any fs call that would reach the drive: `inspectChannelMedia`
+ (`lib/channelMedia.ts:300`; the gate at `:314`, before the marker read, so a stalled channel
+ mid-move reads `stalled`); `probeLocation` (`storageVolumes.ts:241`) and `probeLocationMemo`
+ (`:499`, before the memo); `volumeFreeBytes` (`controller/storageLocations.ts:257`);
+ `readChannelStat` (`controller/channels.ts:197`); the recency tail reads
+ (`controller/recencyIndex.ts:218`, not memoised as misses); `relocationRootPresenceProblem`
+ (`controller/relocateChannelMedia.ts:293`); `inspectSavedVideosStore`
+ (`controller/relocateSavedVideos.ts:167`); `listSavedVideos` when a page passes `notAnswering`;
+ the videos list, the video page, the Cleanup stage (`MediaNotAnswering.tsx`) and the media file
+ route (503). `channelMediaStall(config)` (`channelMedia.ts:220`) is the no-I/O question for a
+ caller holding a config.
+- **`stalled` is a sixth `ChannelMediaStatus` and a sixth `StorageLocationStatus`** ("Not
+ answering"). `HELD_REASON`, `MediaLocationBadge`'s two tables and `STORAGE_STATUS_LABEL` are the
+ `Record`s that make tsc name every table a seventh would need. `isMediaHeld` holds it, so both
+ pool-wide builds hold a stalled channel; the storage watch counts it as down (two passes pause).
+- **`inspectChannelMedia` is memoised for 5 s** (`CHANNEL_MEDIA_MEMO_MS`, `:251`), keyed by channels
+ dir, slug and configured `dataDir`, on `globalThis.__yttChannelMediaMemo__`, asked AFTER the gate.
+ `{ fresh: true }` skips it and does not store; every decider passes it (the start-of-work guard,
+ both movers, both builds, the watch, eviction, the re-point preflight, doctor). The runners' tick
+ shares the status poll's `buildChannelWork` and so reads the memo. `forgetChannelMedia` (`:277`) is
+ called by the channel mover's marker writes and clear, `clearRelocationMarker`, a re-point,
+ /storage's Refresh and the e2e `invalidate-cache` route.
+- **`UV_THREADPOOL_SIZE`** defaults to 16 in `editor/package.json`'s `start` and in
+ `docker/entrypoint.sh`. `ports.test.ts` reads every `${NAME:-N}` in a script as a port and names it
+ as the one exception (`NUMERIC_NOT_PORTS`).
+- **Not covered:** a call already in flight when the drive stalls; a stall that begins between two
+ probes (up to 18 s of calls); a stall the child `stat` does not see because the root's inode is in
+ the kernel's cache (the probe then answers in time); jobs already running; per-click server
+ actions; the corpus disk itself.
diff --git a/plans/release-15.md b/plans/release-15.md
@@ -17,7 +17,7 @@ prompt carries its ruling, and this record carries what was built. Rules:
| Slice | Branch | What | Owns |
|---|---|---|---|
| IG | `r15/index-hold` | The index build holds an unreachable channel instead of emptying it | `common/controller/buildIndex.ts` + new `buildIndex.test.ts`, `common/controller/buildStats.ts` (the hold's words move to a shared module), new `common/lib/channelMediaHold.ts`, `common/lib/envVars.ts`, `ENVIRONMENT.md`; records: `plans/{STATE,FACTS,stats-cache-key}.md` |
-| DS | `r15/drive-stall` | A stalled drive does not stop the editor answering | per its prompt |
+| DS | `r15/drive-stall` | A stalled drive does not stop the editor answering | new `common/lib/storageHealth.ts`; `lib/{storageVolumes,channelMedia,channelMediaHold}.ts`, `controller/storageWatch.ts` and the gated callers; `/storage`, `/channels`, the videos pages; `UV_THREADPOOL_SIZE` (`editor/package.json`, `docker/entrypoint.sh`, `envVars.ts`) |
| UT | `r15/umtool-trace` | umtool's build stops tracing the whole `umtool/` folder | per its prompt |
**Order:** IG → DS. DS adds a health gate inside `inspectChannelMedia`, which IG's hold calls
@@ -216,4 +216,193 @@ checkout's code, so they hold from the moment `main` has this branch. The editor
**Build index** button runs its built bundle, so it holds only after the editor is rebuilt and
restarted.
+### Slice DS, as shipped — a stalled drive does not stop the editor answering (2026-09-29)
+
+Branch `r15/drive-stall` off `main` `ccf90892` (slice IG merged), worktree `~/Projects/r12-paths-fix`
+(block #12: editor 4201, test 4211, export 4210), one Opus implementer. Scratch files `ds-*` in the
+job's `tmp`. The ruling: a drive that is mounted and not answering must not stop the editor
+answering. The drive is asked out of process, and pages and polls do not touch it in-process while
+it is not answering.
+
+**What was wrong.** Node runs every filesystem call on libuv's thread pool, four threads by default.
+On a drive that has stalled (an SMR disk in a USB enclosure resetting under a long write) each call
+blocks its thread for about 30 s. The home page, `/channels` and the three-second auto-queue status
+poll each `stat`ted every relocated channel's target in-process and uncached; the one-second job-list
+poll walked the whole `data/` of every channel with a job listed; and the recency layer read
+metadata tails 32 at a time. Four blocked calls were enough for no page, poll or job log to answer.
+The storage watch saw nothing of it: its five-minute pass asks "is the disk here", and a stalled disk
+is here.
+
+- **The health probe** (`common/lib/storageVolumes.ts`, `probeLocationHealth`):
+ - `stat -L -c %F -- <root>` as a CHILD PROCESS (execa, `reject: false`), raced against a 3 s timer.
+ The timer's answer is `stalled`; the child is sent SIGKILL and not waited for (a process in
+ uninterruptible I/O dies only when the I/O returns, and execa's own `timeout` waits for the exit).
+ - `directory` on stdout is `ok`; any other answer, or a non-zero exit, is `absent`. A `stat` that
+ could not be started is "could not ask": `ok`, never `stalled` (the module's fail-open rule).
+- **The health state** (new `common/lib/storageHealth.ts`): one map per process on `globalThis`
+ (`__yttStorageHealth__`, the house pattern), by location id: state `ok | stalled | absent`, `since`,
+ the last check, a clean streak and the cause. In memory only.
+ - One missed probe marks the location `stalled` at once. Two clean answers in a row clear it (an
+ `absent` answer is clean: an unmounted drive answers ENOENT at once). A miss in between starts the
+ count again. A re-pointed root starts the location over. Locations no longer configured are
+ pruned.
+ - Pure of I/O and without execa, so `lib/channelMedia.ts` can ask it.
+- **The cadence** (`common/controller/storageWatch.ts`): `startStorageWatch` arms a second timer,
+ every 15 s, beside the five-minute pass, and runs one health pass at once. All locations are probed
+ concurrently; each answer is bounded by the probe's timer, and an overrunning pass is not stacked.
+ A transition is logged (`[storage] "<id>": drive not answering — …` / `answering again (ok)`).
+ Stopped with the watch. The watch is armed below the idle gate, so an idle boot has no health probe
+ and so no gate.
+- **The gate**, before any filesystem call on a stalled location:
+
+ | Caller | On a stalled location |
+ |---|---|
+ | `inspectChannelMedia` (`lib/channelMedia.ts`) | New status `stalled`, detail `drive not answering (location "<label>", since HH:MM)`. Asked right after the configured `dataDir` is known and before the marker, link and target reads, so with a config passed in it makes no call at all (without one, only `config.json` is read). Every page and poll that inspects (home, `/channels`, the channel page, the status poll and the runners' tick, the ops channel route, the bulk actions) and every guard gets it. |
+ | `assertChannelMediaReachable` | Refuses it (`ChannelMediaUnreachableError`, status `stalled`), so `runManagedFunction`'s `needsMedia` guard, `generateChannelSnapshot` (also the channel page's **Refresh report** button), the operation batch, normalise, keep-videos and the shard action refuse it. |
+ | The index and stats builds | Hold it: `HELD_REASON.stalled` is "its drive is not answering (a stalled disk)", and IG's `isMediaHeld` holds every status but `ok` and `in-place`. |
+ | `probeLocation` and `probeLocationMemo` (`lib/storageVolumes.ts`) | New probe status `stalled` (`STORAGE_STATUS_LABEL`: "Not answering"), identity unknown, no free space; no `stat`, `statfs` or `findmnt`. The memo is asked after the gate, so a remembered "available" does not outlive the stall. |
+ | `volumeFreeBytes` (`controller/storageLocations.ts`) | Unknown ("—"), with no `stat` or `statfs`. |
+ | `readChannelStat` (`controller/channels.ts`) | `null` (no counts, so no progress bar), with no walk of `data/`. The one-second job-list poll and the home page ask it for every channel with a job listed. |
+ | The recency tail reads (`controller/recencyIndex.ts`) | Skipped for channels on a stalled location (from `meta[].config`), and NOT remembered as misses: they fall to layers 3 and 4 until a refresh with the drive answering reads them. |
+ | `relocationRootPresenceProblem` (`controller/relocateChannelMedia.ts`) | A move onto a stalled location is refused before the root's `stat`. |
+ | `inspectSavedVideosStore` (`controller/relocateSavedVideos.ts`) | The store on a stalled location reads `unreachable` with the stall's detail, before the `stat` through its link; `/storage` then skips the store's size walk. |
+ | `listSavedVideos` (`controller/savedVideoInventory.ts`), when a page passes `notAnswering` | The channel is skipped and named (its `data/` link is read, not followed). `/saved-videos` names the channels it did not read. The backup job passes nothing and is unchanged. |
+ | The videos list, the video page (and its title) and the channel page's Cleanup stage | A notice (`aria-label="media not answering"`, `MediaNotAnswering.tsx`) with links to the channel and `/storage`; nothing is read from the drive. |
+ | The media file route (`/api/channels/<slug>/videos/<id>/files/<name>`) | 503 with `Retry-After: 15`. |
+
+- **The memo** (`inspectChannelMedia`): five seconds per channel, keyed by channels dir, slug and
+ the configured `dataDir`, on `globalThis` (`__yttChannelMediaMemo__`). The gate is asked before
+ it. `{ fresh: true }` bypasses it and does not store; every caller that decides from the answer
+ passes it: `assertChannelMediaReachable`, both movers' preconditions, the index build (both
+ looks), the stats build, the storage watch, eviction, the re-point preflight and `doctor`.
+ `forgetChannelMedia(slug?)` clears it: the channel mover's marker writes and clear (each phase
+ writes its marker after the link and config it changes), `clearRelocationMarker`, a completed
+ re-point, `/storage`'s Refresh and the e2e `invalidate-cache` route.
+- **Headroom:** `UV_THREADPOOL_SIZE=${UV_THREADPOOL_SIZE:-16}` in `editor/package.json`'s `start`
+ (which the rollout restart script runs) and in `docker/entrypoint.sh` before the editor's `exec`.
+ Declared in `envVars.ts` (runtime) and `ENVIRONMENT.md` regenerated. It buys time for calls
+ already in flight and isolates nothing; the doc line says so. `ports.test.ts` read every
+ `${NAME:-N}` in a script as a port, so it now names `UV_THREADPOOL_SIZE` as the one numeric
+ default that is not.
+- **Where it shows:**
+
+ | Surface | What it says |
+ |---|---|
+ | `/storage` | The row's status badge reads **Not answering**; a line under the details reads `not answering since HH:MM — a stat of its root did not answer within 3 s. Pages and polls skip this drive until it answers twice in a row.` (`aria-label="location not answering"`). Re-point and Mount are withheld with the status. **Refresh** asks the location's health first (one answer, counted like the timer's) and forgets the channels' remembered answers; on a stalled drive its note says so instead of probing. |
+ | `/channels` | The volume chip reads `… · not answering since HH:MM` in place of its free space, with a title saying what it means. Each row's badge reads `on <label> — not answering` (accessible name `media location: Media not answering · on <label>`). |
+ | The channel page | The Storage stage's card is red with the inspector's sentence; its destination list names the location "Not answering"; Move back is withheld for a stalled channel, as for an unreachable one. |
+ | `/saved-videos` | `Not read, because the drive their media is on is not answering: <slugs>.` (`aria-label="saved videos not read"`), above the per-channel table. |
+ | The runners | `[auto] skipping <slug>: media stalled — drive not answering (…)`, once per state change. |
+ | The logs | The health pass's transition lines; the index and stats builds' hold lines. |
+
+ Labels are contracts: no existing accessible name or test id changed.
+
+**Commits**
+
+| Commit | What |
+|---|---|
+| `c0ecc55c` | `common:` `lib/storageHealth.ts`, `probeLocationHealth`, the 15 s health pass, the `stalled` statuses and the gate in `inspectChannelMedia`, `probeLocation`/its memo, `volumeFreeBytes`, `readChannelStat`, the recency tail reads, the move-root check and the saved-video store; `HELD_REASON.stalled`; the 5 s memo with `fresh` for every decider and `forgetChannelMedia` in the movers; the badge's and the stage card's words. Tests: `storageHealth`, `storageHealthProbe`, `storageStall`, `storageWatch`. |
+| `cfe551de` | `editor:` `/storage` (the row's line, Refresh asks the health first), the `/channels` volume chip, the Storage panel's Move back, the videos list and video page notice, the media file route's 503, the e2e `invalidate-cache` route; `refreshLocationHealth`; `views/storage.ts` `notAnswering`. |
+| `6a74c790` | `editor:` `UV_THREADPOOL_SIZE=16` in the editor's `start` and `docker/entrypoint.sh`; `envVars.ts` + `ENVIRONMENT.md`. On its own this commit fails `ports.test.ts` (next row). |
+| `de4128d1` | `editor:` `listSavedVideos({ notAnswering })` for `/saved-videos` and the Cleanup stage's notice; `ports.test.ts` names `UV_THREADPOOL_SIZE` as a numeric default that is not a port. |
+| this commit | `plans:` this section; FACTS "The storage health gate"; the editor changelog. |
+
+**Tests** (unit; no test stalls a real drive: a stalled `stat` is a fake binary that never answers,
+and the gated callers run against an ordinary temp directory the health state is told is stalled)
+
+| File | What it pins |
+|---|---|
+| `lib/storageHealth.test.ts` (10) | One miss stalls at once; one clean answer after a stall does not clear it and two in a row do; a miss in between starts the count again; `absent` is clean; a re-pointed root starts over; the path match is `locationOfDataDir`'s (longest root, strictly under); pruning; the map is on `globalThis`; the "since" wording. |
+| `lib/storageHealthProbe.test.ts` (5) | The real `stat`: a directory is `ok`, a missing path and a file are `absent`. A fake `stat` asleep for 20 s is `stalled` on a 400 ms timer in under 3 s, so nothing waited for the child; the default budget is 3 s (answered after 3–6 s). An answer inside the budget is taken as given; a binary that cannot start is `ok`. |
+| `controller/storageStall.test.ts` (14) | A spy on every `node:fs` and `node:fs/promises` call (the `buildIndex.test.ts` technique, reads included). With the location stalled: `inspectChannelMedia` with a config makes NO call, without one only `config.json`; the guard refuses with no call; `probeLocation` and its memo answer `stalled` with no call on the drive and without running `findmnt`; `volumeFreeBytes`, `readChannelStat`, the recency layer, the move-root check, a snapshot refresh and the saved-video store make no call on the drive; the inventory skips and names the channel (proved by what comes back: `fs-extra` bound its functions before the spy). The memo: two inspects inside 5 s stat the target once, `fresh` and a 5 s age each ask again and a fresh answer is not stored; another target is another key; `forgetChannelMedia` and `clearRelocationMarker` clear it; a stall is seen with an `ok` remembered. |
+| `controller/storageWatch.test.ts` (+6) | The health pass stalls a location on one miss and a page then gets `stalled` without asking; the five-minute pass suspects its channel; the stall clears only after two clean passes in a row; a location no longer configured is forgotten; a probe that throws is `ok`; arming the watch runs one health pass at once and stopping it stops both timers; a Refresh counts as one answer and prunes nothing. |
+| `views/storage.test.ts` (+1) | A stalled row reads "Not answering", carries its line, has no free space, and withholds Re-point and Mount; a row with no entry carries nothing. |
+
+#### Gates (logs `$T/ds-*.log`)
+
+- **tsc** was clean before every commit: 186 s at `c0ecc55c`, 86 s before `cfe551de`, 176 s at
+ `6a74c790`, 130 s at `de4128d1` (the machine was running other slices' suites).
+- **Unit:**
+
+ | Suite | Result |
+ |---|---|
+ | common | **2,256/2,256** at the tip, 71 s (the branch point's 2,220 plus 36 new). At `6a74c790` it was 2,254/2,255: `ports.test.ts`, fixed in `de4128d1`. |
+ | editor unit | 87/87 |
+ | `test:scripts` | 191 passed, 1 skipped (192) at the tip. A first run beside the e2e suite failed `queue-lock.test.mjs:85` once (the holder's details were not yet readable); it passed on the rerun. |
+ | mcp | 271/271 |
+
+- **Docs:** `docs env --check`, `docs files --check` and `settings example --check` all exit **0**.
+- **Build:** the editor's `next build`, with the primary's `transcripts/` linked in and capped at
+ 5 GB with no swap: 104 s, max RSS 1,643,860 KB. The link was removed after the build, and nothing
+ ran through it.
+- **e2e** (editor, detached and queued):
+
+ | Run | Specs | Result |
+ |---|---|---|
+ | 1, at `6a74c790` | `storage-locations`, `channel-storage`, `channels-storage-columns`, `channels`, `channels-actions`, `channels-counts`, `channels-sort`, `channels-rack-layers`, `channels-rack-audit` (skips without `E2E_RACK_SHOTS`), `video-page`, `saved-videos`, `dashboard`, `auto-queue`, and IG's eight (`availability`, `build`, `channel-build-toggle`, `chat-only`, `duplicate-shorts`, `jobs`, `regional-vtt-fallback`, `tags`) — `$T/ds-specs.txt` | **124 passed, 3 failed, 12 skipped, 16.6 min** (36 s in the queue). The three: `auto-queue.spec.ts:218` and `:352`, both 30 s timeouts in the first minutes while this slice's own tsc and common suite were running beside it, and `tags.spec.ts:463` (a `/sites` overlay field not yet editable at 30 s). |
+ | 2, at `de4128d1` | `auto-queue`, `tags`, `saved-videos`, `cleanup-holds`, `cleanup-actionable`, `cleanup-page`, `do-not-clean`, `video-titles`, `video-page` — `$T/ds-specs2.txt` | **70 passed, 1 failed, 5.9 min** (56 s in the queue). `tags.spec.ts:463` and `auto-queue.spec.ts:218` passed. The one: `auto-queue.spec.ts:352` ("UI: build a policy…"), `locator.click` on **Start Auto-transcribe** waiting for an enabled button: the spec's own comment names the race (the three-second poll disables Start after `isEnabled()` answered true); earlier records list it as a known flake (`auto-queue.spec.ts:411` at the time). |
+ | 3, at `de4128d1` | `auto-queue` alone | **24 passed, 1.3 min** (1 min 59 s in the queue), `:352` included. |
+
+ No e2e fixture has a stalled drive: these confirm nothing changed for drives that answer. The
+ stall paths are the unit tests above.
+- **Numbers tool:** none.
+
+#### Found and left
+
+- **The stated limit.** A read already in flight when the drive stalls still holds its thread until
+ the kernel gives up (about 30 s in the observed reset loop), and a stall that begins just after a
+ probe costs every call that reaches the drive until the next one (up to 15 s, plus the 3 s
+ budget). With the gate and 16 threads the editor keeps answering meanwhile.
+- **A `stat` the kernel answers from its cache does not reach the drive.** The probe stats the
+ location's ROOT, whose inode is in the kernel's cache whenever anything has used the drive
+ recently; a stall that only uncached reads hit (a walk of `data/`, a metadata tail) can leave the
+ probe answering in time and the location `ok`. The probe then detects a stall only once the
+ root's own metadata has to come off the device. Three ways to reach the device instead, none built:
+ an `O_DIRECT` read of a known file on the drive (`dd iflag=direct`), the block device's in-flight
+ and completed counters in `/sys/class/block/<dev>/stat` (a stall is requests in flight and none
+ completing), or a watchdog on the gated in-process calls themselves (a call that has not answered
+ in 3 s marks its location stalled). Left for a ruling (question 1 in the report).
+- **Request paths left ungated**, each a click rather than a page or poll: the channel and video
+ server actions that read a video's directory in the request (`bulkVideoActions`,
+ `digestActions`, `videoActions`, `fixIncompleteTranscript`, `pipelineActions`); most of what they
+ do is enqueue jobs, whose guard is fresh and refuses a stalled channel; `/api/media/fetch-window/<jobId>` (one `stat` of the fetched file, once the job is done);
+ `/api/ops/channel/<slug>` is gated through `inspectChannelMedia` and `readChannelStat` (its counts
+ are `null` on a stall).
+- **Jobs are gated only at their start.** A job already running when its drive stalls (the remux
+ that stalls it, for one) keeps its threads and children; a job's own in-process reads
+ (`measureTree`, the index build's processing phase, which IG recorded) block their own threads.
+- **The runners' tick reads the memo.** `buildChannelWork` serves both the tick and the status poll,
+ so a drive unmounted in the last 5 s can have one unit dispatched, which fails at the dangling
+ link. The snapshot regeneration and `runManagedFunction`'s guard are fresh. The bulk actions
+ (Sync all, the bulk move) decide their skips from the memo too; the jobs they queue re-check.
+- **The storage watch counts a stalled location as down:** two five-minute passes auto-pause its
+ channels (the lanes already skip them from the first stalled tick), and the first watch pass
+ after the health clears restores them.
+- **umtool's twin of the reachability check** (`checkChannelReachable` in
+ `umtool/report-to-video/cues.mjs`) has no stall gate: umtool is its own process with no health
+ state, and `umtool/**` belongs to another slice.
+- **An idle boot has no health probe**, so nothing is gated there; a CLI run (`archilyzer index`)
+ has none either. The corpus disk itself is not probed.
+- **`6a74c790` alone fails `ports.test.ts`**; `de4128d1` fixes the test. They merge together.
+
+#### Decisions the operator could overturn
+
+| What I assumed | The alternative |
+|---|---|
+| The gate is asked before the relocation marker, so a stalled channel mid-move reads `stalled`, not `in-transition` (a resumed move accepts any status; a new move refuses). With a config in hand it costs no call. | Read the marker first (one read on the corpus disk) and let `in-transition` win. |
+| A stall auto-pauses like an unmounted drive, on the watch's own two-pass cadence. | Exclude `stalled` from the watch's "down", so only the lanes' per-tick skips apply. |
+| An `absent` answer counts as clean toward clearing a stall. | Only `ok` clears it. |
+| The memo is on by default and every decider passes `fresh`. | Off by default, with pages and polls opting in. |
+| `/storage`'s Refresh is one health answer like the timer's. | Refresh clears a stall outright on one clean answer. |
+| `UV_THREADPOOL_SIZE` defaults to 16 in `start` and the entrypoint; a value already set wins. | A fixed 16, or only in the rollout restart script. |
+| The videos list, the video page and the Cleanup stage show a notice instead of their content on a stalled drive. | Render what the snapshot knows and leave out only the drive's files. |
+| `readChannelStat` returns `null` on a stall (the job row draws no progress bar). | Counts marked unknown. |
+
+**What runs which code, for the rollout.** The health state lives in the editor's process, armed
+with the storage watch, so all of it takes effect only when the editor is rebuilt and restarted
+(and not on an idle boot). The restart must go through `pnpm run start` in `editor/` (the rollout
+restart script does) for `UV_THREADPOOL_SIZE` to apply; a process started another way keeps
+Node's 4 threads unless the variable is set. CLI builds have no health state: they hold a channel
+only for the statuses they already held.
+
## Rollout