commit 59edf0d3b30b4ca8f23fcfa1ca09efb1124e676d
parent f2e19f2e87d17eab7a454468fc84fcd7998ad6f7
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Wed, 30 Sep 2026 00:46:17 -0400
plans: slice DS — the overdue refusal's slow-or-stalled test as a decision row (it replaces the Found-and-left note), in the mechanism, the Review table and the commit table; FACTS on it and refreshed anchors; re-gates (tsc; common 2,301; e2e storage + channels 38/38)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Diffstat:
2 files changed, 27 insertions(+), 20 deletions(-)
diff --git a/plans/FACTS.md b/plans/FACTS.md
@@ -7635,38 +7635,39 @@ The record is [`release-15.md`](release-15.md), "Slice DS, as shipped", with the
detect it reliably: the root's inode is in the kernel's cache whenever the drive was used lately.
- **The state is `common/lib/storageHealth.ts`**, one map on `globalThis.__yttStorageHealth__` (the
pass writes it from instrumentation's module copy; pages read it from theirs), with each entry's
- `detector` and the counters' `device`. `recordLocationHealth` (`:177`): one `stalled` answer stalls
+ `detector` and the counters' `device`. `recordLocationHealth` (`:182`): one `stalled` answer stalls
at once, and every transition to stalled refuses `onDrive`'s waiting calls; `HEALTH_CLEAN_TO_CLEAR`
(2) clean answers in a row clear it; `absent` is clean; a new root starts over.
- `registerLocationHealth` (`:262`) creates entries with no answer. `stalledLocationForPath`
- (`:314`) matches like `locationOfDataDir`; `stalledLocation` (`:330`) is by id AND root.
+ `registerLocationHealth` (`:267`) creates entries with no answer. `stalledLocationForPath`
+ (`:319`) matches like `locationOfDataDir`; `stalledLocation` (`:335`) is by id AND root.
- **Detector 1, every 15 s: the block device's counters** (`detectLocationHealth`,
`lib/storageVolumes.ts:601`). The root's device from the last pass (findmnt `-J -T <root> -o
SOURCE,UUID`, raced against 3 s, only when there is none or its `/sys` entry stops reading;
another volume's UUID names none; `[…]` stripped, `/dev/mapper` resolved, basename); then
- `/sys/class/block/<dev>/stat` (`parseBlockStat`, `storageHealth.ts:710`): completed = fields 1 + 5
+ `/sys/class/block/<dev>/stat` (`parseBlockStat`, `storageHealth.ts:735`): completed = fields 1 + 5
+ 12 + 16 (reads, writes, discards, flushes), in flight = field 9. Stalled ⇔ in flight at both
samples AND nothing completed between; samples at least `MIN_COUNTER_INTERVAL_MS` (10 s) apart; the
first gives no verdict. in_flight counts only requests dispatched to the driver: one requeued
during a host reset is not counted, so a sample in that window can read clean (the watchdog covers
it). No device → the child `stat -L -c %F` probe (`probeLocationHealth`, `:386`). The samples are
on `globalThis.__yttHealthDetector__` (the pass and /storage's Refresh share them).
-- **Detector 2, on every gated call: `onDrive(where, call)`** (`storageHealth.ts:612`). Refused with
+- **Detector 2, on every gated call: `onDrive(where, call)`** (`storageHealth.ts:633`). Refused with
no call on a stalled location; otherwise raced against `DRIVE_CALL_BUDGET_MS` (3 s; test seam
`setDriveCallBudget`); the budget covers the whole unit passed in. A timeout marks the location
stalled (since now) and throws `DriveNotAnsweringError`, leaving the call to settle — unless the
location's device counters (read synchronously from `/sys` through the reader `storageVolumes.ts`
registers with `setCounterReader`, `:509`) moved since the call began: then the call is refused
as slow and nothing is marked. At most `DRIVE_CALLS_IN_FLIGHT` (4) calls per slot key in flight
- (`acquireSlot`, `storageHealth.ts:494`): the rest queue in JS. A waiting call's deadline follows progress: every call that
- returns on the key (in time or late) restarts it (`releaseSlot`, `:578`, re-arms every waiter), and a
+ (`acquireSlot`, `storageHealth.ts:501`): the rest queue in JS. A waiting call's deadline follows progress: every call that
+ returns on the key (in time or late) restarts it (`releaseSlot`, `:599`, re-arms every waiter), and a
waiting call is refused unmarked only when nothing on the key has returned for the budget plus a
quarter of it (at most 250 ms) — never for the queue's depth alone. Waiting calls are refused at
once by any transition to stalled; and when every slot is held by a call already past its budget
- (`overdue`), a new call is refused at once and the location marked stalled again (even if the
- counters show the disk completing other requests). A slot is freed when its call really returns. The
+ (`overdue`, each kept with the counters reading from when it began), a new call is refused at once,
+ and the location is marked stalled again only if the disk has completed nothing since the oldest of
+ them began (otherwise "drive slow", unmarked). A slot is freed when its call really returns. The
slot key: a configured location's id; a probe of another root under its id, that root (marks
- nothing); a path on no configured location, the root it is under (`rootOfUnknownPath`, `:450`;
+ nothing); a path on no configured location, the root it is under (`rootOfUnknownPath`, `:457`;
marks nothing). Do not nest it for one key. The timer is not unref'd.
- **The cadence** is `runStorageHealthPass` (`controller/storageWatch.ts:433`): prune, register, then
every location concurrently; every 15 s from `startStorageHealthWatch` (`:546`), plus one at arm
diff --git a/plans/release-15.md b/plans/release-15.md
@@ -541,9 +541,11 @@ stalled disk is here.
grace of a quarter of it (at most 250 ms; the grace lets the calls it waits behind, whose
timers start a moment later, time out and mark first). A deep queue on a drive that is busy
but answering therefore waits as long as it takes;
- - calls past their budget are counted per key (`overdue`); when every slot is held by one, a
- new call is refused at once and the location marked stalled again, even if the pass has since
- cleared it: none of those calls has returned.
+ - calls past their budget are counted per key (`overdue`), each with the counters reading
+ taken when it began; when every slot is held by one, a new call is refused at once, and the
+ location is marked stalled again (even if the pass has since cleared it) unless the disk has
+ completed requests since the oldest of them began: then it is slow, not stalled, and nothing
+ is marked (the same test as a timeout's).
- **The slot key** (M2, L7): a configured location's id; a probe of another root under a
location's id is keyed by that root and marks nothing; a path on no configured location (a root
typed by hand) is keyed by the root it is under (`<root>/<slug>/data` → `<root>`), so it holds
@@ -650,14 +652,16 @@ stalled disk is here.
| `b34a7613` | `plans:` the review's findings to their commits; FACTS; the changelog. |
| `c367a3d7` | `editor:` re-review L11: the health pass's block above the boot probe's comment. |
| `d61bdd93` | `common:` re-review M4: a slot wait's deadline follows progress. Tests. |
-| this commit | `plans:` the re-review's findings to their commits; FACTS; the report. |
+| `ee68d225` | `plans:` the re-review's findings to their commits; FACTS. |
+| `9a12308e` | `common:` the overdue refusal tells slow from stalled by the counters, as a timeout does. Tests. |
+| this commit | `plans:` that ruling as a decision row; FACTS; the report. |
**Tests** (unit; no test stalls a real drive: a stalled call is a promise that never settles or a
fake `stat` that never answers, and a stalled device is a temp `/sys` whose counters stand still)
| File | What it pins |
|---|---|
-| `lib/storageHealth.test.ts` (26) | The rules: one miss stalls at once, one clean answer after a stall does not clear it and two in a row do, a miss in between starts over, `absent` is clean, a re-pointed root starts over, the path match is `locationOfDataDir`'s, pruning, `globalThis`, the "since" wording, registering without an answer. `onDrive`: an answer passes through (value or error) and frees its slot; a never-settling call is `stalled` on the timer, marks the location (since now) and keeps its slot until it settles; a stalled location is refused with no call; seven calls at once put four in flight and the three that waited are refused without a call when the location stalls; a freed slot runs a waiter; the 3 s default (answered after 3–4.5 s). After the review: a hand-typed root's four slots are shared by its channels and a fifth call is refused at once, with no entry made; a candidate root has its own slots and leaves the location's calls alone; **M1's case** (four hung calls, two clean answers, a fifth refused at once and the location marked again); a transition to stalled from the pass refuses the waiting call; **L6** (counters completing: four calls refused as slow, the location not marked, a waiter's wait runs out unmarked when nothing returns; counters standing still: marked). After the re-review (**M4**): a healthy 64-wide walk of units at half the budget (the last call waits about fifteen units) has no refusals and marks nothing, on a location and on a hand-typed root; a queue behind four hung calls is still refused within the budget (marked on a location, unmarked on a hand-typed root); a call that returns late hands its slot to the calls waiting behind it. |
+| `lib/storageHealth.test.ts` (28) | The rules: one miss stalls at once, one clean answer after a stall does not clear it and two in a row do, a miss in between starts over, `absent` is clean, a re-pointed root starts over, the path match is `locationOfDataDir`'s, pruning, `globalThis`, the "since" wording, registering without an answer. `onDrive`: an answer passes through (value or error) and frees its slot; a never-settling call is `stalled` on the timer, marks the location (since now) and keeps its slot until it settles; a stalled location is refused with no call; seven calls at once put four in flight and the three that waited are refused without a call when the location stalls; a freed slot runs a waiter; the 3 s default (answered after 3–4.5 s). After the review: a hand-typed root's four slots are shared by its channels and a fifth call is refused at once, with no entry made; a candidate root has its own slots and leaves the location's calls alone; **M1's case** (four hung calls, two clean answers, a fifth refused at once and the location marked again); a transition to stalled from the pass refuses the waiting call; **L6** (counters completing: four calls refused as slow, the location not marked, a waiter's wait runs out unmarked when nothing returns; counters standing still: marked). After the re-review (**M4**): a healthy 64-wide walk of units at half the budget (the last call waits about fifteen units) has no refusals and marks nothing, on a location and on a hand-typed root; a queue behind four hung calls is still refused within the budget (marked on a location, unmarked on a hand-typed root); a call that returns late hands its slot to the calls waiting behind it. Then: four units slower than the budget on a device whose completions move — the fifth call is refused at once and nothing is marked; with completions unchanged, after the pass cleared the location — refused and marked again. |
| `lib/storageHealthCounters.test.ts` (7) | The stat line parser (17 and 11 fields, garbage); the verdict over sample pairs (stuck → stalled; moving, idle or drained → ok); device names (partition, `[subvolume]` suffix, non-`/dev` sources); the detector end to end with a fake findmnt and a temp `/sys`: the first sample gives no verdict, stuck → stalled with its cause, drained → ok, busy and moving → ok; a sample sooner than 10 s gives none and keeps the first; every no-device fallback goes to the child stat and says so (tmpfs, findmnt failing, no `/sys` entry, another volume's UUID, no binary). After the review: discards and flushes are counted (a flush alone is not a stall); a known device is read without running findmnt, and findmnt runs again only when its `/sys` entry stops reading; the samples are on `globalThis`. |
| `lib/storageHealthProbe.test.ts` (5) | The child `stat`: a directory is `ok`, a missing path and a file are `absent`; a fake `stat` asleep for 20 s is `stalled` on the timer without being waited for; the 3 s default; an answer inside the budget is taken; no binary is `ok`. |
| `controller/storageStall.test.ts` (22) | A spy on every `node:fs` and `node:fs/promises` call, with a hang mode that makes a matching promise-API call never settle. The gate: with the location stalled, inspect reads only the marker (with a config) or config.json and the marker (without); a channel mid-move on a stalled drive reads `in-transition` (and is remembered so); the guard, `probeLocation` and its memo (no findmnt run), `volumeFreeBytes`, `readChannelStat`, the recency layer, the move-root check, a snapshot refresh, the saved-video store and the inventory make no call on the drive. The watchdog: with the location answering and one drive call hung, inspect, `readChannelStat` (at most four video directories asked), `probeLocation`, `volumeFreeBytes`, the recency layer (not remembered as a miss) the move-root check and (M3) the snapshot walk each answer `stalled` within the race and mark the location (the snapshot writes no `snapshot.json`, and at most four video directories reach the drive); after it, inspect, the guard and the walk make no call on the drive, and once cleared the drive is asked again. The memo: two inspects inside 5 s stat the target once, `fresh` and a 5 s age ask again, a fresh answer is not stored, another target is another key, `forgetChannelMedia` and `clearRelocationMarker` clear it, and a stall is seen with an `ok` remembered. |
@@ -673,7 +677,7 @@ fake `stat` that never answers, and a stalled device is a temp `/sys` whose coun
| Suite | Result |
|---|---|
- | common | **2,295/2,295**, 58 s: `main`'s 2,229 plus DS's 66 (60 at the rulings, 6 from the review). After the re-review: **2,299/2,299**, 53 s (4 for M4). |
+ | common | **2,295/2,295**, 58 s: `main`'s 2,229 plus DS's 66 (60 at the rulings, 6 from the review). After the re-review: **2,299/2,299**, 53 s (4 for M4); after the overdue ruling **2,301/2,301**, 75 s (2 more). |
| editor unit | 87/87, and 87/87 after the re-review |
| `test:scripts` | 194 passed, 2 skipped (196). The second skip is UT's post-build trace check, which skips a umtool build older than its config (this worktree's `umtool/.next` predates UT); the first is the `LIVE=1` archive check. |
| mcp | 271/271 at the first pass; no mcp file has changed since. |
@@ -691,6 +695,7 @@ fake `stat` that never answers, and a stalled device is a temp `/sys` whose coun
| 2, first pass | `auto-queue`, `tags`, `saved-videos`, four cleanup specs, `video-titles`, `video-page` (`$T/ds-specs2.txt`) | **70 passed, 1 failed, 5.9 min**: `auto-queue.spec.ts:352`, the Start-button race its own comment describes. |
| 3, first pass | `auto-queue` alone | **24 passed, 1.3 min**. |
| 4, after the rulings (`b1a30902`) | the 9 `storage` + `channels` specs (`$T/ds-specs4.txt`) | **38 passed, 0 failed, 12 skipped, 2.5 min**. |
+ | 7, after the overdue ruling (`9a12308e`) | the 9 `storage` + `channels` specs (`$T/ds-specs4.txt`) | **38 passed, 0 failed, 12 skipped, 2.5 min**. |
| 6, after the re-review (`d61bdd93`) | the 9 `storage` + `channels` specs (`$T/ds-specs4.txt`) | **38 passed, 0 failed, 12 skipped, 2.6 min**. |
| 5, after the review, on the merged tree (`30df4193`) | the 9 `storage` + `channels` specs; the snapshot scheduler's (`auto-report-refresh`, `channel-work`, `jobs-batch-tasks-drain`, `incomplete-transcript`); the keep-latest users of the bounded key fan-out (`cleanup-holds`, `saved-videos`); `reconcile`; `review` (the auto-pause wording) — `$T/ds-specs5.txt` | **73 passed, 0 failed, 12 skipped, 5.2 min** (37 min 54 s in the queue behind another session's suite). |
@@ -725,10 +730,9 @@ fake `stat` that never answers, and a stalled device is a temp `/sys` whose coun
longer matters (M4): a waiting call is refused only when nothing on the drive has returned for
3 s. But a single unit that takes longer than 3 s is refused: without a mark when the counters
show the disk completing other requests, with one otherwise. A snapshot refresh with such a unit
- throws, so the scheduler keeps the last `snapshot.json` and tries again on its next trigger. And
- once four such slow units hold every slot past the budget, the next call is refused at once and
- the location marked stalled (the M1 rule, kept as ruled), even when the counters show the disk
- completing: the pass then clears it after two clean samples (15–30 s).
+ throws, so the scheduler keeps the last `snapshot.json` and tries again on its next trigger. Once
+ four such units hold every slot past the budget, the next call is refused at once (a decision
+ below says when that also marks).
- **The counters need two samples.** A drive already stalled when the editor starts is seen by the
counters at the second pass (15–30 s), or at once by the watchdog when a page reaches it.
in_flight counts only requests dispatched to the driver: a request requeued during a host reset
@@ -763,6 +767,7 @@ fake `stat` that never answers, and a stalled device is a temp `/sys` whose coun
| At most four gated calls per location in flight; the rest queue in JavaScript and are refused on a stall. **Kept at review.** | No cap: the watchdog alone, and a stall mid-walk fills the pool until the kernel gives up. |
| A wait for a slot runs out at the budget plus a quarter of it (at most 250 ms), so simultaneous timeouts of the calls it waits behind mark first. | Exactly the budget: a waiter queued in the same tick then gives up a moment before those calls and is refused unmarked, and the mark lands a few ms later. |
| The keep-latest key reads run 16 at a time for every caller (they were an unbounded `Promise.all`). | Bound them only for the snapshot. |
+| **Ruled after the re-review:** when every slot is held by a call past its budget, the next call is refused at once, and marks the location stalled only when the disk has completed nothing since the oldest of those calls began; four slow units on a disk still completing requests are "drive slow", refused and unmarked. With no device named (the stat detector), it marks. | Mark whenever every slot is overdue (the M1 fix as first built): four slow units on a busy disk then marked it stalled for 15–30 s. |
| An auto-pause record with no `cause` reads as not there. | Word both cases for it ("not there or not answering"). |
| Two counter samples closer than 10 s give no verdict (a Refresh just after a pass among them). **Kept at review.** | Compare any two samples (a busy healthy drive can read "in flight, nothing completed" over a few milliseconds). |
| A findmnt that does not answer reuses the last device named for that root; one naming another volume's UUID names none. | Treat a findmnt that does not answer as a stall. |
@@ -807,5 +812,6 @@ spin-up case stays the operator's question). Two new ones:
|---|---|---|
| M4: a slot wait was timed from when the call queued, so a deep queue on a busy but answering drive was refused as "not answering" (units of about 1.1 s at 16 wide, 0.46 s at 32, 0.2 s at 64) | The deadline follows progress: every return on the key re-arms its waiters, and a wait is refused only when nothing on the key has returned for the budget; the overdue and transition refusals unchanged. The limit statement and the keep-latest comment corrected | `d61bdd93`, this commit |
| L11: the health pass's block sat inside the boot probe's comment, which still called itself the only thing an idle boot runs | Moved above it; the phrase dropped | `c367a3d7` |
+| (the implementer's note) four slow units past the budget on a disk whose counters show completions made the overdue refusal mark the location stalled | The same treatment as L6: the overdue refusal compares the counters with those taken when the oldest overdue call began; moved → refused, not marked; unchanged → marked | `9a12308e`, this commit |
## Rollout