commit 0eb8ffc8cbbdc2f3a5c7e2268c8dd81c82b2c4c1
parent 682506cca69d21e4329c3c4546ed2829ac056a5b
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Wed, 30 Sep 2026 09:08:33 -0400
plans: slice DT as shipped — the drive-health timings are settings.storage.health, edited on /storage; the slices row; FACTS' storage health gate names the setting and its accessor; the changelog
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Diffstat:
3 files changed, 236 insertions(+), 38 deletions(-)
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -15,6 +15,7 @@
- **The hub URL hints say what the setting does now.** Settings' **Family hub URL** and a site's **Hub URL** no longer promise a Hub link in the header (it was removed): the value is published as `hubUrl` in each site's `/site.json` and `/corpus.json`, so the hub can tell its member sites. `SETTINGS.md` and `SITE.md` say the same.
- **A site can be left off the homepage and the hub.** A site's settings have a new checkbox, **List on the Archilyzer homepage and hub**, on by default (`listed` in `site.json`; only `false` is written). Turned off, the site still builds and deploys at its own URL as before, but the homepage has no card, chart series, `/stats` entry or recent item for it; the hub does not list it as a member, search it, or name it in its `corpus.json` and `llms.txt`; no other site's footer links it; and `channel-sites.json` and the homepage's `stats/` leave it out. A channel only unlisted sites carry is in none of the published totals, the homepage's headline numbers included; a channel a listed site also carries is counted under the listed site. The editor's own pages still show every site. It takes effect at the next homepage, hub and site builds.
- **The sidebar's site picker shows your site from the first paint.** It used to show "All sites" on every page and then jump to the site you had picked, and Dashboard and Channels came up in your site only after a `?site=` had been added to the address. The picked site is now kept in a cookie that the editor reads before it draws a page, so the picker, Dashboard and Channels open in it at once, and the address is left alone. A link that carries `?site=<id>` still opens that page in that site, without changing the one you picked; picking a site on such a page drops the `?site=` from the address. On a site's own pages (Charts, Publish, …) the picker still follows the page, and opening one still makes that site the picked one. The first time you open the editor after updating, a site picked before is moved into the cookie; the picker may show "All sites" for a moment that once. A site picked in one tab reaches the editor's other open tabs without a reload. **New channel** starts with the picked site ticked under its sites (or the site of a `?site=` link), including when it is opened from the editor's own links. Each editor keeps its own pick, as before, when several run on one machine on different ports.
+- **The drive check's timings can be changed on `/storage`.** The numbers the editor decides a drive is "not answering" by were fixed: a read may take 3 seconds, each drive is checked every 15 seconds (a check that asks the drive from a separate process waits up to 3 seconds), two clean checks in a row put a drive back in use, and at most four reads wait on one drive at a time. They are now **Drive health timing**, a collapsed block at the foot of `/storage`, with those numbers as the defaults — for when a drive that is busy but working is marked not answering, or a stalled one is not. An empty field is its default; a number outside a field's range is refused, with the range. A save takes effect at once: the next read, the next check, and a new check interval re-times the checks. `settings.json` keeps only the values that differ from a default, under `storage.health` (see `SETTINGS.md`), so an editor that never changes them follows the defaults; `archilyzer index` and the stats build read them too. The messages that said "3 s", "every 15 s" or "twice in a row" now say the numbers in force.
## [0.10.0] - 2026-09-28
- **The homepage can be built and deployed from `/sites`.** Under a new **Homepage** section, after Hub, there is **Build homepage** (tick **Deploy after build** to ship it in the same job, only if the build succeeds) and **Deploy homepage**, which ships the build already in `homepage/out`. A **Preview branch** box beside them sends either deploy to a Cloudflare Pages preview of the `archilyzer` project instead of production, and shows the preview's address as you type; a name Cloudflare would refuse or rewrite, or `main`, greys the deploy buttons out and says why. A line under the buttons says what a deploy would ship: when `homepage/out` was built (or that it holds no build yet), and where it goes, with the live URL. Deploy homepage with nothing built is refused before any job starts. The homepage reads the search index as it stands, so run **Build index** first when its numbers should move. The jobs run the same code as `archilyzer build homepage` / `deploy homepage`, and show on `/jobs` as `build-homepage`, `deploy-homepage` and `build-deploy-homepage`. The Hub section no longer describes the homepage.
diff --git a/plans/FACTS.md b/plans/FACTS.md
@@ -7673,40 +7673,57 @@ this section is stale, by +16 near the top and +203 at the end; they are not rew
## The storage health gate (verified 2026-09-29, branch `r15/drive-stall`)
The record is [`release-15.md`](release-15.md), "Slice DS, as shipped", with the parent's rulings
-(Q1–Q5) and the review's fixes (M1–M3, L1–L10). Anchors are at the branch tip after the merge of
-`main` `bab894db`.
+(Q1–Q5) and the review's fixes (M1–M3, L1–L10); the timings became settings in "Slice DT, as
+shipped". Anchors are at slice DT's tip (`r15/drive-timings`).
- **A drive can be mounted and not answering.** Every in-process fs call on it waits on one of
libuv's threads (4 by default, 16 in the editor's `start`) until it answers (~30 s for the observed
USB reset loop); only a child process isolates a call. A child `stat` of a location's ROOT does not
detect it reliably: the root's inode is in the kernel's cache whenever the drive was used lately.
+- **The timings are settings: `settings.storage.health`** (slice DT). `budgetMs` (default 3000,
+ 500–60000), `passIntervalMs` (15000, 5000–300000), `probeTimeoutMs` (3000, 500–30000),
+ `clearAfterCleanPasses` (2, 1–10), `inFlightPerLocation` (4, 1–16); the defaults, ranges, sanitizer
+ and words are `common/lib/storageHealthTimings.ts` (pure; the /storage form imports it). A read
+ clamps into the range and keeps only a value that differs from its default (an untuned file has no
+ `health` key); the /storage form refuses out of range with a sentence. Every number is read through
+ ONE accessor, `healthTimings()` (`storageHealth.ts:182`), from the timings as last applied on
+ `globalThis` — no file read per call. `applyHealthTimings(stored)` (`:193`) sets them: the health
+ pass on every pass (from the settings it reads), /storage's save at once
+ (`saveHealthTimingsAction`), and the `index` and `build stats` bins once at start. Any other process
+ with no pass (a CLI, `archilyzer doctor`) runs on the defaults. A changed interval is told to
+ `onPassIntervalChange` (`:214`) subscribers, which re-arms the armed pass's timer; a raised cap
+ admits waiting calls, a lowered one is reached as calls return. `resetStorageHealth` keeps the
+ timings (configuration, not health); `setDriveCallBudget` is a test seam below the 500 ms floor and
+ wins over them.
- **The state is `common/lib/storageHealth.ts`**, one map on `globalThis.__yttStorageHealth__` (the
pass writes it from instrumentation's module copy; pages read it from theirs), with each entry's
- `detector` and the counters' `device`. `recordLocationHealth` (`:182`): one `stalled` answer stalls
- at once, and every transition to stalled refuses `onDrive`'s waiting calls; `HEALTH_CLEAN_TO_CLEAR`
- (2) clean answers in a row clear it; `absent` is clean; a new root starts over.
- `registerLocationHealth` (`:267`) creates entries with no answer. `stalledLocationForPath`
- (`:319`) matches like `locationOfDataDir`; `stalledLocation` (`:335`) is by id AND root.
-- **Detector 1, every 15 s: the block device's counters** (`detectLocationHealth`,
- `lib/storageVolumes.ts:601`). The root's device from the last pass (findmnt `-J -T <root> -o
- SOURCE,UUID`, raced against 3 s, only when there is none or its `/sys` entry stops reading;
+ `detector` and the counters' `device`. `recordLocationHealth` (`:251`): one `stalled` answer stalls
+ at once, and every transition to stalled refuses `onDrive`'s waiting calls; `clearAfterCleanPasses`
+ (2 by default) clean answers in a row clear it; `absent` is clean; a new root starts over.
+ `registerLocationHealth` (`:336`) creates entries with no answer. `stalledLocationForPath`
+ (`:388`) matches like `locationOfDataDir`; `stalledLocation` (`:404`) is by id AND root.
+- **Detector 1, every `passIntervalMs` (15 s): the block device's counters** (`detectLocationHealth`,
+ `lib/storageVolumes.ts:607`). The root's device from the last pass (findmnt `-J -T <root> -o
+ SOURCE,UUID`, raced against `probeTimeoutMs` (3 s), only when there is none or its `/sys` entry stops reading;
another volume's UUID names none; `[…]` stripped, `/dev/mapper` resolved, basename); then
- `/sys/class/block/<dev>/stat` (`parseBlockStat`, `storageHealth.ts:735`): completed = fields 1 + 5
+ `/sys/class/block/<dev>/stat` (`parseBlockStat`, `storageHealth.ts:826`): completed = fields 1 + 5
+ 12 + 16 (reads, writes, discards, flushes), in flight = field 9. Stalled ⇔ in flight at both
- samples AND nothing completed between; samples at least `MIN_COUNTER_INTERVAL_MS` (10 s) apart; the
- first gives no verdict. in_flight counts only requests dispatched to the driver: one requeued
+ samples AND nothing completed between; samples at least `minCounterIntervalMs()` apart
+ (`storageVolumes.ts:472`: min(10 s, interval − 5 s), floored at half the interval — 10 s at the
+ default 15 s); the first gives no verdict. in_flight counts only requests dispatched to the driver: one requeued
during a host reset is not counted, so a sample in that window can read clean (the watchdog covers
- it). No device → the child `stat -L -c %F` probe (`probeLocationHealth`, `:386`). The samples are
+ it). No device → the child `stat -L -c %F` probe (`probeLocationHealth`, `:388`, raced against
+ `probeTimeoutMs`). The samples are
on `globalThis.__yttHealthDetector__` (the pass and /storage's Refresh share them).
-- **Detector 2, on every gated call: `onDrive(where, call)`** (`storageHealth.ts:633`). Refused with
- no call on a stalled location; otherwise raced against `DRIVE_CALL_BUDGET_MS` (3 s; test seam
- `setDriveCallBudget`); the budget covers the whole unit passed in. A timeout marks the location
+- **Detector 2, on every gated call: `onDrive(where, call)`** (`storageHealth.ts:722`). Refused with
+ no call on a stalled location; otherwise raced against `budgetMs` (3 s by default, read as the call
+ starts; test seam `setDriveCallBudget`); the budget covers the whole unit passed in. A timeout marks the location
stalled (since now) and throws `DriveNotAnsweringError`, leaving the call to settle — unless the
location's device counters (read synchronously from `/sys` through the reader `storageVolumes.ts`
- registers with `setCounterReader`, `:509`) moved since the call began: then the call is refused
- as slow and nothing is marked. At most `DRIVE_CALLS_IN_FLIGHT` (4) calls per slot key in flight
- (`acquireSlot`, `storageHealth.ts:501`): the rest queue in JS. A waiting call's deadline follows progress: every call that
- returns on the key (in time or late) restarts it (`releaseSlot`, `:599`, re-arms every waiter), and a
+ registers with `setCounterReader`, `:515`) moved since the call began: then the call is refused
+ as slow and nothing is marked. At most `inFlightPerLocation` (4 by default) calls per slot key in
+ flight (`acquireSlot`, `storageHealth.ts:573`): the rest queue in JS. A waiting call's deadline follows progress: every call that
+ returns on the key (in time or late) restarts it (`releaseSlot`, `:673`, re-arms every waiter), and a
waiting call is refused unmarked only when nothing on the key has returned for the budget plus a
quarter of it (at most 250 ms) — never for the queue's depth alone. Waiting calls are refused at
once by any transition to stalled; and when every slot is held by a call already past its budget
@@ -7714,42 +7731,43 @@ The record is [`release-15.md`](release-15.md), "Slice DS, as shipped", with the
and the location is marked stalled again only if the disk has completed nothing since the oldest of
them began (otherwise "drive slow", unmarked). A slot is freed when its call really returns. The
slot key: a configured location's id; a probe of another root under its id, that root (marks
- nothing); a path on no configured location, the root it is under (`rootOfUnknownPath`, `:457`;
+ nothing); a path on no configured location, the root it is under (`rootOfUnknownPath`, `:529`;
marks nothing). Do not nest it for one key. The timer is not unref'd.
-- **The cadence** is `runStorageHealthPass` (`controller/storageWatch.ts:433`): prune, register, then
- every location concurrently; every 15 s from `startStorageHealthWatch` (`:546`), plus one at arm
- time, armed by `editor/instrumentation.ts` ABOVE the idle gate (it writes nothing). The five-minute
- pass (`startStorageWatch`, `:521`) stays below it. A CLI process has no pass (its inspects are still
- raced). `refreshLocationHealth` (`:485`) is /storage's Refresh.
-- **The gate order in `inspectChannelMedia`** (`lib/channelMedia.ts:316`): config → memo (a
+- **The cadence** is `runStorageHealthPass` (`controller/storageWatch.ts:442`): apply the timings it
+ read, prune, register, then every location concurrently; every `passIntervalMs` (15 s) from
+ `startStorageHealthWatch` (`:566`, re-armed when the interval changes), plus one at arm time,
+ armed by `editor/instrumentation.ts` ABOVE the idle gate (it writes nothing). The five-minute pass
+ (`startStorageWatch`, `:535`) stays below it. A CLI process has no pass (its inspects are still
+ raced). `refreshLocationHealth` (`:499`) is /storage's Refresh.
+- **The gate order in `inspectChannelMedia`** (`lib/channelMedia.ts:317`): config → memo (a
remembered `in-transition` is returned as is, anything else is gated first) → the relocation
- marker (`:362`, corpus disk) → the gate (`:380`) → the link (corpus disk) → the target's `stat`
- through `onDrive` (`:445`). A stall is never memoised. Other gated calls: `probeLocation`
- (`storageVolumes.ts:255`, its stat and statfs through `onDrive`) and its memo; `volumeFreeBytes`
- (`controller/storageLocations.ts:288`, `:325`, `:334`); `readChannelStat` (`controller/channels.ts:215`,
- the walk through `onDrive`, `null` on a stall); the snapshot walk (`controller/channelSnapshot.ts:751`,
+ marker (`:367`, corpus disk) → the gate (`:386`) → the link (corpus disk) → the target's `stat`
+ through `onDrive` (`:447`). A stall is never memoised. Other gated calls: `probeLocation`
+ (`storageVolumes.ts:239`, its stat and statfs through `onDrive`) and its memo; `volumeFreeBytes`
+ (`controller/storageLocations.ts:289`, `:326`, `:335`); `readChannelStat` (`controller/channels.ts:205`,
+ the walk through `onDrive`, `null` on a stall); the snapshot walk (`controller/channelSnapshot.ts:753`,
its listing, keep-latest keys and per-video unit); the recency tail reads; the move-root check; the
saved-video store; `listSavedVideos` with `notAnswering`; the videos list, the video page, the
Cleanup stage, the Storage stage's statfs and the media file route. `channelMediaStall(config)`
- (`channelMedia.ts:231`) is the no-I/O question for a holder of a config.
+ (`channelMedia.ts:232`) is the no-I/O question for a holder of a config.
- **`stalled` is a sixth `ChannelMediaStatus` and a sixth `StorageLocationStatus`** ("Not
answering"). `HELD_REASON`, `MediaLocationBadge`'s two tables and `STORAGE_STATUS_LABEL` are the
`Record`s that make tsc name every table a seventh would need. `isMediaHeld` holds it, so both
pool-wide builds hold a stalled channel; the storage watch counts it as down (two passes pause),
on a location or not, and the pause record carries `cause: "not-answering"` (`ChannelAutoPause`,
`lib/channelPriority.ts`; absent = not there).
-- **`inspectChannelMedia` is memoised for 5 s** (`CHANNEL_MEDIA_MEMO_MS`, `:265`), keyed by channels
+- **`inspectChannelMedia` is memoised for 5 s** (`CHANNEL_MEDIA_MEMO_MS`, `:266`), keyed by channels
dir, slug and configured `dataDir`, on `globalThis.__yttChannelMediaMemo__`. `{ fresh: true }`
skips it and does not store; the deciders that pass it are listed in the record (the guard and its
six callers, both movers, both builds, the watch, eviction, the re-point preflight, doctor). The
runners' tick shares the status poll's `buildChannelWork` and so reads the memo.
- `forgetChannelMedia` (`:291`) is called by the channel mover's marker writes and clear,
+ `forgetChannelMedia` (`:292`) is called by the channel mover's marker writes and clear,
`clearRelocationMarker`, a re-point, /storage's Refresh and the e2e `invalidate-cache` route.
- **`UV_THREADPOOL_SIZE`** defaults to 16 in `editor/package.json`'s `start` and in
`docker/entrypoint.sh`. `ports.test.ts` reads every `${NAME:-N}` in a script as a port and names it
as the one exception (`NUMERIC_NOT_PORTS`, `common/lib/ports.test.ts:29`).
-- **Not covered:** a call already in flight when the drive stalls (at most four per drive for the
- calls through `onDrive` — every page and poll path and the snapshot walk; a job's own reads that do
+- **Not covered:** a call already in flight when the drive stalls (at most `inFlightPerLocation` per
+ drive for the calls through `onDrive` — every page and poll path and the snapshot walk; a job's own reads that do
not go through it, `measureTree`, the index build's processing phase, the snapshot's sequential
reconcile pass, are not capped); a hand-typed root is capped but never marked; a drive already
stalled at boot before the second counter sample, unless a page reaches it; per-click server
diff --git a/plans/release-15.md b/plans/release-15.md
@@ -20,6 +20,7 @@ prompt carries its ruling, and this record carries what was built. Rules:
| DS | `r15/drive-stall` | A stalled drive does not stop the editor answering | new `common/lib/storageHealth.ts`; `lib/{storageVolumes,channelMedia,channelMediaHold}.ts`, `controller/storageWatch.ts` and the gated callers; `/storage`, `/channels`, the videos pages; `UV_THREADPOOL_SIZE` (`editor/package.json`, `docker/entrypoint.sh`, `envVars.ts`) |
| UT | `r15/umtool-trace` | umtool's build stops tracing the whole `umtool/` folder | per its prompt |
| SS | `r15/site-scope` | The editor's site picker paints the stored site at once: the selection is a cookie | `editor/app/lib/activeSite{,Server,Actions}.ts` + `activeSite.test.ts`, `editor/app/components/SiteScope{Provider,Select}.tsx`, `editor/app/layout.tsx`, the scope lines of `editor/app/page.tsx` and `editor/app/channels/page.tsx`, a comment in `editor/next.config.ts`, `editor/e2e/site-scope.spec.ts`; records: `plans/FACTS.md` |
+| DT | `r15/drive-timings` | The drive-health timings are settings (`settings.storage.health`), edited on `/storage` | new `common/lib/storageHealthTimings.ts` + test; `lib/{storageHealth,storageVolumes,channelMedia,storageLocations,settingsSchema,settingsDocs}.ts`, `controller/storageWatch.ts`, the `index` and `build stats` bins, `SETTINGS.md`; `/storage` (a form, its action and parse), the stall wording on `/channels` and the videos pages; `storage-locations.spec.ts`; records: `plans/FACTS.md` |
**Order:** IG → DS. DS adds a health gate inside `inspectChannelMedia`, which IG's hold calls
through its public signature. UT is independent. The shared files are `editor/CHANGELOG.md`'s
@@ -1088,6 +1089,184 @@ It reproduced M1 with a two-tab probe under the queue lock.
editor's built bundle, so all of it takes effect only after the editor is rebuilt and restarted.
After that, each browser's first visit migrates its localStorage selection once.
+### Slice DT, as shipped — the drive-health timings are settings (2026-09-30)
+
+Branch `r15/drive-timings` off `main` `6b8aa450` (slice DS merged), worktree
+`~/Projects/r12-paths-fix` (block #12: editor 4201, test 4211, export 4210), one Opus implementer.
+Scratch files `dt-*` in the job's `tmp`. The ruling (operator, 2026-09-30): the drive-health timings
+are configurable in `settings.json`, editable on `/storage`, with today's constants as the defaults;
+if a stall is misjudged under heavy external-disk churn, the operator tunes the numbers rather than
+the code.
+
+**What was fixed in code.** Slice DS judged "mounted and not answering" by five constants in
+`lib/storageHealth.ts`: the watchdog's budget (`DRIVE_CALL_BUDGET_MS`, 3 s), the health pass's cadence
+(`HEALTH_PROBE_INTERVAL_MS`, 15 s), the child-stat and findmnt race (`HEALTH_PROBE_TIMEOUT_MS`, 3 s),
+the clean answers that clear a stall (`HEALTH_CLEAN_TO_CLEAR`, 2) and the calls in flight per
+location (`DRIVE_CALLS_IN_FLIGHT`, 4). The constants are gone; tsc named every reader.
+
+- **The setting** is `settings.storage.health`, a nested block of five optional keys:
+
+ | Key | Default | Range | Takes effect |
+ |---|---|---|---|
+ | `budgetMs` | 3000 | 500–60000 | the next `onDrive` call (read as the call starts) |
+ | `passIntervalMs` | 15000 | 5000–300000 | at once on a save from `/storage` (the armed pass re-arms its timer); a hand edit, at the next pass |
+ | `probeTimeoutMs` | 3000 | 500–30000 | the next pass or Refresh (child `stat` and findmnt) |
+ | `clearAfterCleanPasses` | 2 | 1–10 | the next answer |
+ | `inFlightPerLocation` | 4 | 1–16 | the next slot taken; a raised cap admits waiting calls at once, a lowered one is reached as calls return |
+
+ - The type, defaults, ranges, sanitizer, words and `SETTINGS.md` docs are one pure module,
+ `common/lib/storageHealthTimings.ts` (the `/storage` form, a client file, imports it).
+ `StorageSettings.health` is optional; `sanitizeStorage` keeps a block only when something is in it.
+ - **A read clamps; only what differs from a default is kept.** `sanitizeStorageHealth` rounds each
+ number and clamps it into its range (the schema's convention: a read never throws), drops a
+ non-number, and drops a value equal to its default. An untuned `settings.json` therefore has no
+ `health` key, and a save of every default removes it.
+ - The mediaRoot migration keeps a `health` block that spells no `locations` (a hand edit).
+- **The read path: one accessor.** `healthTimings()` (`storageHealth.ts:182`) returns the timings as
+ last applied, on the health state's `globalThis` object, every absent key its default. There is no
+ settings memo in lib (`getSettings()` reads the file on every call), so the accessor reads no file:
+ the numbers are applied into memory by `applyHealthTimings(stored)` (`:193`):
+ - the health pass, at the start of every pass, from the settings it already reads for its
+ locations (so at boot, and within one pass of a hand edit);
+ - the `/storage` save, at once (`saveHealthTimingsAction`);
+ - the `index` and `build stats` bins, once, before the build (a CLI process has no pass).
+ `lib/storageHealth.ts` stays free of I/O. `resetStorageHealth` keeps the timings (configuration, not
+ health). `setDriveCallBudget` stays the test seam below the 500 ms floor and wins over them.
+- **The derived numbers stay derived.** The slot wait's grace is still a quarter of the budget, at
+ most 250 ms. The counters' sample spacing is `counterSampleMinimumMs(passIntervalMs)` =
+ min(10 s, interval − 5 s), floored at half the interval (`minCounterIntervalMs()` in
+ `storageVolumes.ts:472`): 10 s at the default 15 s, as before, and at every interval from 10 s up the
+ ruling's formula exactly (see the decisions table for the floor).
+- **The re-arm.** `startStorageHealthWatch` arms at `healthTimings().passIntervalMs` and subscribes
+ with `onPassIntervalChange` (`storageHealth.ts:214`); `applyHealthTimings` tells the subscribers when
+ the interval changed, and the watch clears and re-arms its interval (log line `[storage] health
+ pass re-armed: every N s`). The subscription is on `globalThis`, so the save (a page's module copy)
+ reaches the pass (instrumentation's). An explicit `intervalMs` (the tests') follows nothing.
+- **The cap, changed live.** `acquireSlot` reads the cap each time. `releaseSlot` hands a slot on only
+ while the key is at or under the cap, and otherwise gives it back; `admitWaiters` lets calls already
+ waiting take the slots a raised cap adds. The overdue refusal compares with the cap as it is, and
+ its words count the overdue calls ("N reads on it have not answered") instead of naming the constant.
+- **`/storage`: Drive health timing.** A collapsed `<details>` at the foot of the page
+ (`components/HealthTimingForm.tsx`, `aria-label="drive health timing"`; its summary says
+ "defaults" or "N changed from the default"). Five text inputs (`inputMode="numeric"`), each with the
+ default as its placeholder, a hint line in the operator's words, and the default and range
+ (`HEALTH_TIMING_HINTS`); the interval's hint says a save re-arms the check at once. An empty field is
+ the default. Accessible names (new): `read budget`, `health check interval`, `health check timeout`,
+ `clean checks to clear`, `reads at once per drive`, `save timing`, `timing saved`, `timing error`;
+ none contains another. No existing name changed.
+ - The server action (`saveHealthTimingsAction`, `app/storage/actions.ts`) parses with
+ `lib/healthTimingsForm.ts`: a value out of range, or not a whole number, is **refused** with a
+ sentence naming the field, the range and the value ("Read budget must be between 500 and 60000 ms
+ (got 200 ms)."), and nothing is written. Otherwise it saves a patch of the storage block through
+ `saveSettings`, applies the timings, revalidates `/storage`, and answers with the timings now in
+ force.
+- **The surfaces say the setting, not the constant.** `/storage`'s "not answering since …" line stays,
+ and its "until it answers twice in a row" follows the clear count (`clearRuleText`: once, twice in a
+ row, N times in a row). So do the Refresh note (and its "checked every N s"), the `/channels`
+ volume chip's title (a new `clears` field on `ChannelVolume`), the health pass's stall log line,
+ and the media-not-answering notice ("within N s", "checked every N s"). The `onDrive` header, the
+ module header, and the comments that said "3 s" or "every 15 s" in the gated callers now name the
+ setting and its default. FACTS' "The storage health gate" has a bullet for the setting and names the
+ keys where it named the constants.
+
+**Commits**
+
+| Commit | What |
+|---|---|
+| `13411edc` | `common:` `lib/storageHealthTimings.ts`; `settings.storage.health` in the schema, the docs table and the migration; `healthTimings()`, `applyHealthTimings()`, `onPassIntervalChange()`; the pass applies what it reads and re-arms; the live cap; `minCounterIntervalMs()`; the two bins; `SETTINGS.md`; the comments. Tests. |
+| `41ad2f65` | `editor:` the Drive health timing form, its action and parse (+ unit test), the page; the Refresh note, the stalled line, the media notice, the volume chip's title; the comments. |
+| `b92df1fe` | `editor(e2e):` `storage-locations.spec.ts`: the drive health timing case. |
+| `4372abaf` | `common:` the cap's hint says to keep it well under the editor's 16 file-access threads. |
+| this commit | `plans:` this section and the slices row; FACTS; the changelog. |
+
+**Tests** (unit; no test stalls a real drive)
+
+| File | What it pins |
+|---|---|
+| `lib/storageHealthTimings.test.ts` (7, new) | The defaults are the constants they replace, each in its range. Absent, empty, an array, a string, a number: every default, and no block (`getSettings()` with no file, `defaultSiteSettings()`). A read clamps (200 → 500, 900000 → 300000), rounds (7.6 → 8), drops a string and an unknown key, and drops a value equal to its default (also one that rounds onto it). **The settings.json round trip** through `writeSettings`/`getSettings`: `{budgetMs: 4000, clearAfterCleanPasses: 2}` is written as `{budgetMs: 4000}` and read back; a save of the default removes the key; a hand-edited 99 reads as 16. The block survives the mediaRoot migration. The spacing (15 s → 10 s, 300 s → 10 s, 12 s → 7 s, 10 s → 5 s, 8 s → 4 s, 5 s → 2.5 s). The words. |
+| `lib/storageHealth.test.ts` (+6) | **The accessor feeds the watchdog:** a stored `budgetMs: 200` applies as 500 ms (the floor), and a 700 ms unit is refused and marks the location ("within 0.5 s"); on the defaults the same unit answers. **The cap:** with `inFlightPerLocation: 2` the third call waits, and runs when a slot frees. A cap raised from 1 to 3 admits the two waiting calls at once; lowered to 1 with three in flight, two returns bring it to one and the fourth call still waits, and runs on the third return. **The clear count:** with 3, two clean answers do not clear and the third does; with 1, one does. A changed interval is told to the subscribers once, an unchanged one and other keys are not, and an unsubscribed one hears nothing. The test seam's budget wins, and a reset keeps the timings. The existing constants' assertions read `HEALTH_TIMING_DEFAULTS`. |
+| `lib/storageHealthCounters.test.ts` (+1) | The spacing follows the applied interval: at 8 s, samples 3,999 ms apart give no verdict and 4,000 ms apart compare (stalled); at 300 s it is 10 s. The existing case reads `minCounterIntervalMs()` (10 s). |
+| `lib/storageHealthProbe.test.ts` (+1) | With `probeTimeoutMs: 500` applied and no `timeoutMs` passed, a child that sleeps 20 s is `stalled` after 0.5–2.5 s. |
+| `controller/storageWatch.test.ts` (+2) | **Every pass applies what it reads:** `clearAfterCleanPasses: 3` and `budgetMs: 5000` in the settings are in force after the first pass; its stall line says "until it answers 3 times in a row"; two clean passes do not clear and the third does; a pass handed its locations reads no settings and leaves the timings. **The re-arm:** the armed watch, told `passIntervalMs: 5000` (the save's apply), logs the re-arm and runs its second pass 4.5–7 s later (15 s at the default); a stopped watch re-arms nothing; one armed with an explicit interval does not follow. |
+| `editor/app/storage/lib/healthTimingsForm.test.ts` (5, new) | One field per key, in order, with accessible names none of which contains another. Empty and blank fields write nothing. A value in range is kept and trimmed; one equal to its default is not written. Out of range is refused with the field, range and value, for a millisecond field and both counts. "3.5", "-1", "3e3", "abc", "0x10" and "two" are refused as not whole numbers. |
+
+**e2e** (`storage-locations.spec.ts`, new case "the drive health timing saves to settings.json and
+reads back"): after hydration, the block is collapsed and says "defaults"; opened, `read budget` is
+empty with placeholder 3000 (`reads at once per drive`: 4); 4000 saved → "A read may take 4 s" and
+`test-settings.json` holds `storage.health` = `{budgetMs: 4000}` and nothing else; after a reload the
+summary says "1 changed from the default", the field reads 4000 and the interval is empty; 200 is
+refused with the sentence and the file is unchanged; emptied and saved, "A read may take 3 s" and the
+`health` key is gone. The fixture's settings are its own `test-settings.json`.
+
+#### Gates (logs `$T/dt-*.log`)
+
+- **tsc** (all workspaces): clean before every commit — 58 s on the tree of the first three code
+ commits, 51 s after the hint's.
+- **Unit:**
+
+ | Suite | Result |
+ |---|---|
+ | common | **2,318/2,318**, 50 s (`main`'s 2,301 + 17) |
+ | editor unit | **100/100** (95 + 5) |
+ | `test:scripts` | 194 passed, 2 skipped (196), as at DS |
+ | mcp | not run: no mcp file and nothing it imports changed |
+
+- **Docs:** `settings example --check`, `docs env --check` and `docs files --check` all exit **0**
+ (`SETTINGS.md` regenerated in `13411edc`: the `health` row and the `storage.health` table;
+ `settings.json.example` unchanged, the default block has no `health`).
+- **Build:** the editor's `next build`, with the primary's `transcripts/` linked in (`ln -sT`) and
+ capped at 5 GB with no swap: **34 s, max RSS 1,648,128 KB**, exit 0. The link was removed after the
+ build, and nothing ran through it.
+- **e2e** (editor, detached and queued; `$T/dt-specs.txt`, the nine `storage` + `channels` specs DS
+ ran), with the new case in `storage-locations`:
+
+ | Run | At | Result |
+ |---|---|---|
+ | 1 | `b92df1fe` (the hint's edit landed on disk while it ran; nothing it asserts) | **39 passed, 0 failed, 12 skipped, 2.7 min** |
+ | 2 | `4372abaf` (the code as shipped) | **39 passed, 0 failed, 12 skipped, 2.5 min** |
+
+ DS's 38 plus the new case. The 12 skips are `channels-rack-audit`, which needs `E2E_RACK_SHOTS`.
+ Neither run waited in the queue.
+- **Numbers tool:** none.
+
+#### Found and left
+
+- **Other CLI processes run on the defaults.** The `index` and `build stats` bins apply the settings;
+ `archilyzer doctor` (`common/bin/doctor.ts`, another slice's file) and any other command whose
+ inspects go through `onDrive` race them against the default 3 s. A one-line
+ `applyHealthTimings(settingsFromFile(paths.settingsFile).storage.health)` at its start is all it
+ needs.
+- **`inFlightPerLocation` may be set to 16, the editor's whole thread pool** (`UV_THREADPOOL_SIZE`,
+ 16 in `start`). At 16, one drive that stops answering can hold every file-access thread, which is
+ what DS exists to prevent; two drives at 8 can too. The range is the ruling's; the form's hint and
+ `SETTINGS.md` say to keep it well under 16.
+- **A hand edit of `passIntervalMs`** re-arms at the next pass, so it can wait up to the old interval
+ (at most 5 minutes). A save from `/storage` re-arms at once.
+- **A refused save clears the typed value.** React resets a form after its action returns, so the
+ field shows the stored value again; the error sentence names the refused value.
+- **umtool's reachability twin** (`checkChannelReachable`) still has no health state (DS left it;
+ `umtool/**` is another slice's).
+
+#### Decisions the operator could overturn
+
+| What I assumed | The alternative |
+|---|---|
+| The counters' spacing is the ruling's min(10 s, interval − 5 s), **floored at half the interval**. Without the floor it is 0 at the 5 s minimum interval (1 s at 6 s), so a Refresh just after a pass would compare two samples milliseconds apart, the case DS's review kept the spacing for. From 10 s up the two agree. | The formula as ruled, 0 at 5 s. Or a higher minimum interval (10 s) |
+| A read clamps an out-of-range value (the schema's convention); the `/storage` form refuses it with a sentence and writes nothing. | The form clamps too, and says what it stored |
+| Only a value that differs from its default is written, so a save of 3000 for the budget writes nothing, and a later release's new default reaches it. | Write what the operator saved, pinning the default of the day |
+| The slice's test as specified (`budgetMs: 200` → a 300 ms unit refused) is below the ruled 500 ms floor, so the test stores 200, shows it clamped to 500, and refuses a 700 ms unit; the default passes the same unit. | Lower the floor so 200 applies |
+| The timings are applied into the health state (the pass on every pass, the save at once, the two bins), not read from `settings.json` on each call: lib has no settings memo, and `onDrive` is on every page's hot path. | A time-limited memo of `getSettings()` inside the accessor (a file read at most every few seconds, and lib/storageHealth.ts no longer free of I/O) |
+| A save re-arms the pass's timer at once, through a subscription on `globalThis`. | Leave the running timer; the new interval at the next restart |
+| A lowered cap is reached as calls return; calls already in flight are not refused. A raised cap admits waiting calls at once. | Refuse the calls over the new cap |
+| The form's inputs are text with a numeric keypad, so the action's sentence is the only validation. | `type="number"` with `min`/`max`: the browser's own bubble, and "3.5" blocked before the action |
+| `resetStorageHealth` (a test seam and the e2e `invalidate-cache` route) keeps the applied timings. | Reset them to the defaults until the next pass |
+
+**What runs which code, for the rollout.** The accessor, the pass's apply and re-arm, and the form
+are in the editor's built bundle and its instrumentation, so all of it takes effect after the editor
+is rebuilt and restarted. Nothing is written until the operator saves the form; with no `health` key
+the editor runs on today's numbers. A CLI `archilyzer index` or `build stats` reads the settings
+itself.
+
## Rollout
Release 15 is slices IG (`r15/index-hold`, merged `ccf90892`), UT (`r15/umtool-trace`, `07c991be`),