Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 59edf0d3b30b4ca8f23fcfa1ca09efb1124e676d
parent f2e19f2e87d17eab7a454468fc84fcd7998ad6f7
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Wed, 30 Sep 2026 00:46:17 -0400

plans: slice DS — the overdue refusal's slow-or-stalled test as a decision row (it replaces the Found-and-left note), in the mechanism, the Review table and the commit table; FACTS on it and refreshed anchors; re-gates (tsc; common 2,301; e2e storage + channels 38/38)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
Mplans/FACTS.md | 21+++++++++++----------
Mplans/release-15.md | 26++++++++++++++++----------
2 files changed, 27 insertions(+), 20 deletions(-)

diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -7635,38 +7635,39 @@ The record is [`release-15.md`](release-15.md), "Slice DS, as shipped", with the detect it reliably: the root's inode is in the kernel's cache whenever the drive was used lately. - **The state is `common/lib/storageHealth.ts`**, one map on `globalThis.__yttStorageHealth__` (the pass writes it from instrumentation's module copy; pages read it from theirs), with each entry's - `detector` and the counters' `device`. `recordLocationHealth` (`:177`): one `stalled` answer stalls + `detector` and the counters' `device`. `recordLocationHealth` (`:182`): one `stalled` answer stalls at once, and every transition to stalled refuses `onDrive`'s waiting calls; `HEALTH_CLEAN_TO_CLEAR` (2) clean answers in a row clear it; `absent` is clean; a new root starts over. - `registerLocationHealth` (`:262`) creates entries with no answer. `stalledLocationForPath` - (`:314`) matches like `locationOfDataDir`; `stalledLocation` (`:330`) is by id AND root. + `registerLocationHealth` (`:267`) creates entries with no answer. `stalledLocationForPath` + (`:319`) matches like `locationOfDataDir`; `stalledLocation` (`:335`) is by id AND root. - **Detector 1, every 15 s: the block device's counters** (`detectLocationHealth`, `lib/storageVolumes.ts:601`). The root's device from the last pass (findmnt `-J -T <root> -o SOURCE,UUID`, raced against 3 s, only when there is none or its `/sys` entry stops reading; another volume's UUID names none; `[…]` stripped, `/dev/mapper` resolved, basename); then - `/sys/class/block/<dev>/stat` (`parseBlockStat`, `storageHealth.ts:710`): completed = fields 1 + 5 + `/sys/class/block/<dev>/stat` (`parseBlockStat`, `storageHealth.ts:735`): completed = fields 1 + 5 + 12 + 16 (reads, writes, discards, flushes), in flight = field 9. Stalled ⇔ in flight at both samples AND nothing completed between; samples at least `MIN_COUNTER_INTERVAL_MS` (10 s) apart; the first gives no verdict. in_flight counts only requests dispatched to the driver: one requeued during a host reset is not counted, so a sample in that window can read clean (the watchdog covers it). No device → the child `stat -L -c %F` probe (`probeLocationHealth`, `:386`). The samples are on `globalThis.__yttHealthDetector__` (the pass and /storage's Refresh share them). -- **Detector 2, on every gated call: `onDrive(where, call)`** (`storageHealth.ts:612`). Refused with +- **Detector 2, on every gated call: `onDrive(where, call)`** (`storageHealth.ts:633`). Refused with no call on a stalled location; otherwise raced against `DRIVE_CALL_BUDGET_MS` (3 s; test seam `setDriveCallBudget`); the budget covers the whole unit passed in. A timeout marks the location stalled (since now) and throws `DriveNotAnsweringError`, leaving the call to settle — unless the location's device counters (read synchronously from `/sys` through the reader `storageVolumes.ts` registers with `setCounterReader`, `:509`) moved since the call began: then the call is refused as slow and nothing is marked. At most `DRIVE_CALLS_IN_FLIGHT` (4) calls per slot key in flight - (`acquireSlot`, `storageHealth.ts:494`): the rest queue in JS. A waiting call's deadline follows progress: every call that - returns on the key (in time or late) restarts it (`releaseSlot`, `:578`, re-arms every waiter), and a + (`acquireSlot`, `storageHealth.ts:501`): the rest queue in JS. A waiting call's deadline follows progress: every call that + returns on the key (in time or late) restarts it (`releaseSlot`, `:599`, re-arms every waiter), and a waiting call is refused unmarked only when nothing on the key has returned for the budget plus a quarter of it (at most 250 ms) — never for the queue's depth alone. Waiting calls are refused at once by any transition to stalled; and when every slot is held by a call already past its budget - (`overdue`), a new call is refused at once and the location marked stalled again (even if the - counters show the disk completing other requests). A slot is freed when its call really returns. The + (`overdue`, each kept with the counters reading from when it began), a new call is refused at once, + and the location is marked stalled again only if the disk has completed nothing since the oldest of + them began (otherwise "drive slow", unmarked). A slot is freed when its call really returns. The slot key: a configured location's id; a probe of another root under its id, that root (marks - nothing); a path on no configured location, the root it is under (`rootOfUnknownPath`, `:450`; + nothing); a path on no configured location, the root it is under (`rootOfUnknownPath`, `:457`; marks nothing). Do not nest it for one key. The timer is not unref'd. - **The cadence** is `runStorageHealthPass` (`controller/storageWatch.ts:433`): prune, register, then every location concurrently; every 15 s from `startStorageHealthWatch` (`:546`), plus one at arm diff --git a/plans/release-15.md b/plans/release-15.md @@ -541,9 +541,11 @@ stalled disk is here. grace of a quarter of it (at most 250 ms; the grace lets the calls it waits behind, whose timers start a moment later, time out and mark first). A deep queue on a drive that is busy but answering therefore waits as long as it takes; - - calls past their budget are counted per key (`overdue`); when every slot is held by one, a - new call is refused at once and the location marked stalled again, even if the pass has since - cleared it: none of those calls has returned. + - calls past their budget are counted per key (`overdue`), each with the counters reading + taken when it began; when every slot is held by one, a new call is refused at once, and the + location is marked stalled again (even if the pass has since cleared it) unless the disk has + completed requests since the oldest of them began: then it is slow, not stalled, and nothing + is marked (the same test as a timeout's). - **The slot key** (M2, L7): a configured location's id; a probe of another root under a location's id is keyed by that root and marks nothing; a path on no configured location (a root typed by hand) is keyed by the root it is under (`<root>/<slug>/data` → `<root>`), so it holds @@ -650,14 +652,16 @@ stalled disk is here. | `b34a7613` | `plans:` the review's findings to their commits; FACTS; the changelog. | | `c367a3d7` | `editor:` re-review L11: the health pass's block above the boot probe's comment. | | `d61bdd93` | `common:` re-review M4: a slot wait's deadline follows progress. Tests. | -| this commit | `plans:` the re-review's findings to their commits; FACTS; the report. | +| `ee68d225` | `plans:` the re-review's findings to their commits; FACTS. | +| `9a12308e` | `common:` the overdue refusal tells slow from stalled by the counters, as a timeout does. Tests. | +| this commit | `plans:` that ruling as a decision row; FACTS; the report. | **Tests** (unit; no test stalls a real drive: a stalled call is a promise that never settles or a fake `stat` that never answers, and a stalled device is a temp `/sys` whose counters stand still) | File | What it pins | |---|---| -| `lib/storageHealth.test.ts` (26) | The rules: one miss stalls at once, one clean answer after a stall does not clear it and two in a row do, a miss in between starts over, `absent` is clean, a re-pointed root starts over, the path match is `locationOfDataDir`'s, pruning, `globalThis`, the "since" wording, registering without an answer. `onDrive`: an answer passes through (value or error) and frees its slot; a never-settling call is `stalled` on the timer, marks the location (since now) and keeps its slot until it settles; a stalled location is refused with no call; seven calls at once put four in flight and the three that waited are refused without a call when the location stalls; a freed slot runs a waiter; the 3 s default (answered after 3–4.5 s). After the review: a hand-typed root's four slots are shared by its channels and a fifth call is refused at once, with no entry made; a candidate root has its own slots and leaves the location's calls alone; **M1's case** (four hung calls, two clean answers, a fifth refused at once and the location marked again); a transition to stalled from the pass refuses the waiting call; **L6** (counters completing: four calls refused as slow, the location not marked, a waiter's wait runs out unmarked when nothing returns; counters standing still: marked). After the re-review (**M4**): a healthy 64-wide walk of units at half the budget (the last call waits about fifteen units) has no refusals and marks nothing, on a location and on a hand-typed root; a queue behind four hung calls is still refused within the budget (marked on a location, unmarked on a hand-typed root); a call that returns late hands its slot to the calls waiting behind it. | +| `lib/storageHealth.test.ts` (28) | The rules: one miss stalls at once, one clean answer after a stall does not clear it and two in a row do, a miss in between starts over, `absent` is clean, a re-pointed root starts over, the path match is `locationOfDataDir`'s, pruning, `globalThis`, the "since" wording, registering without an answer. `onDrive`: an answer passes through (value or error) and frees its slot; a never-settling call is `stalled` on the timer, marks the location (since now) and keeps its slot until it settles; a stalled location is refused with no call; seven calls at once put four in flight and the three that waited are refused without a call when the location stalls; a freed slot runs a waiter; the 3 s default (answered after 3–4.5 s). After the review: a hand-typed root's four slots are shared by its channels and a fifth call is refused at once, with no entry made; a candidate root has its own slots and leaves the location's calls alone; **M1's case** (four hung calls, two clean answers, a fifth refused at once and the location marked again); a transition to stalled from the pass refuses the waiting call; **L6** (counters completing: four calls refused as slow, the location not marked, a waiter's wait runs out unmarked when nothing returns; counters standing still: marked). After the re-review (**M4**): a healthy 64-wide walk of units at half the budget (the last call waits about fifteen units) has no refusals and marks nothing, on a location and on a hand-typed root; a queue behind four hung calls is still refused within the budget (marked on a location, unmarked on a hand-typed root); a call that returns late hands its slot to the calls waiting behind it. Then: four units slower than the budget on a device whose completions move — the fifth call is refused at once and nothing is marked; with completions unchanged, after the pass cleared the location — refused and marked again. | | `lib/storageHealthCounters.test.ts` (7) | The stat line parser (17 and 11 fields, garbage); the verdict over sample pairs (stuck → stalled; moving, idle or drained → ok); device names (partition, `[subvolume]` suffix, non-`/dev` sources); the detector end to end with a fake findmnt and a temp `/sys`: the first sample gives no verdict, stuck → stalled with its cause, drained → ok, busy and moving → ok; a sample sooner than 10 s gives none and keeps the first; every no-device fallback goes to the child stat and says so (tmpfs, findmnt failing, no `/sys` entry, another volume's UUID, no binary). After the review: discards and flushes are counted (a flush alone is not a stall); a known device is read without running findmnt, and findmnt runs again only when its `/sys` entry stops reading; the samples are on `globalThis`. | | `lib/storageHealthProbe.test.ts` (5) | The child `stat`: a directory is `ok`, a missing path and a file are `absent`; a fake `stat` asleep for 20 s is `stalled` on the timer without being waited for; the 3 s default; an answer inside the budget is taken; no binary is `ok`. | | `controller/storageStall.test.ts` (22) | A spy on every `node:fs` and `node:fs/promises` call, with a hang mode that makes a matching promise-API call never settle. The gate: with the location stalled, inspect reads only the marker (with a config) or config.json and the marker (without); a channel mid-move on a stalled drive reads `in-transition` (and is remembered so); the guard, `probeLocation` and its memo (no findmnt run), `volumeFreeBytes`, `readChannelStat`, the recency layer, the move-root check, a snapshot refresh, the saved-video store and the inventory make no call on the drive. The watchdog: with the location answering and one drive call hung, inspect, `readChannelStat` (at most four video directories asked), `probeLocation`, `volumeFreeBytes`, the recency layer (not remembered as a miss) the move-root check and (M3) the snapshot walk each answer `stalled` within the race and mark the location (the snapshot writes no `snapshot.json`, and at most four video directories reach the drive); after it, inspect, the guard and the walk make no call on the drive, and once cleared the drive is asked again. The memo: two inspects inside 5 s stat the target once, `fresh` and a 5 s age ask again, a fresh answer is not stored, another target is another key, `forgetChannelMedia` and `clearRelocationMarker` clear it, and a stall is seen with an `ok` remembered. | @@ -673,7 +677,7 @@ fake `stat` that never answers, and a stalled device is a temp `/sys` whose coun | Suite | Result | |---|---| - | common | **2,295/2,295**, 58 s: `main`'s 2,229 plus DS's 66 (60 at the rulings, 6 from the review). After the re-review: **2,299/2,299**, 53 s (4 for M4). | + | common | **2,295/2,295**, 58 s: `main`'s 2,229 plus DS's 66 (60 at the rulings, 6 from the review). After the re-review: **2,299/2,299**, 53 s (4 for M4); after the overdue ruling **2,301/2,301**, 75 s (2 more). | | editor unit | 87/87, and 87/87 after the re-review | | `test:scripts` | 194 passed, 2 skipped (196). The second skip is UT's post-build trace check, which skips a umtool build older than its config (this worktree's `umtool/.next` predates UT); the first is the `LIVE=1` archive check. | | mcp | 271/271 at the first pass; no mcp file has changed since. | @@ -691,6 +695,7 @@ fake `stat` that never answers, and a stalled device is a temp `/sys` whose coun | 2, first pass | `auto-queue`, `tags`, `saved-videos`, four cleanup specs, `video-titles`, `video-page` (`$T/ds-specs2.txt`) | **70 passed, 1 failed, 5.9 min**: `auto-queue.spec.ts:352`, the Start-button race its own comment describes. | | 3, first pass | `auto-queue` alone | **24 passed, 1.3 min**. | | 4, after the rulings (`b1a30902`) | the 9 `storage` + `channels` specs (`$T/ds-specs4.txt`) | **38 passed, 0 failed, 12 skipped, 2.5 min**. | + | 7, after the overdue ruling (`9a12308e`) | the 9 `storage` + `channels` specs (`$T/ds-specs4.txt`) | **38 passed, 0 failed, 12 skipped, 2.5 min**. | | 6, after the re-review (`d61bdd93`) | the 9 `storage` + `channels` specs (`$T/ds-specs4.txt`) | **38 passed, 0 failed, 12 skipped, 2.6 min**. | | 5, after the review, on the merged tree (`30df4193`) | the 9 `storage` + `channels` specs; the snapshot scheduler's (`auto-report-refresh`, `channel-work`, `jobs-batch-tasks-drain`, `incomplete-transcript`); the keep-latest users of the bounded key fan-out (`cleanup-holds`, `saved-videos`); `reconcile`; `review` (the auto-pause wording) — `$T/ds-specs5.txt` | **73 passed, 0 failed, 12 skipped, 5.2 min** (37 min 54 s in the queue behind another session's suite). | @@ -725,10 +730,9 @@ fake `stat` that never answers, and a stalled device is a temp `/sys` whose coun longer matters (M4): a waiting call is refused only when nothing on the drive has returned for 3 s. But a single unit that takes longer than 3 s is refused: without a mark when the counters show the disk completing other requests, with one otherwise. A snapshot refresh with such a unit - throws, so the scheduler keeps the last `snapshot.json` and tries again on its next trigger. And - once four such slow units hold every slot past the budget, the next call is refused at once and - the location marked stalled (the M1 rule, kept as ruled), even when the counters show the disk - completing: the pass then clears it after two clean samples (15–30 s). + throws, so the scheduler keeps the last `snapshot.json` and tries again on its next trigger. Once + four such units hold every slot past the budget, the next call is refused at once (a decision + below says when that also marks). - **The counters need two samples.** A drive already stalled when the editor starts is seen by the counters at the second pass (15–30 s), or at once by the watchdog when a page reaches it. in_flight counts only requests dispatched to the driver: a request requeued during a host reset @@ -763,6 +767,7 @@ fake `stat` that never answers, and a stalled device is a temp `/sys` whose coun | At most four gated calls per location in flight; the rest queue in JavaScript and are refused on a stall. **Kept at review.** | No cap: the watchdog alone, and a stall mid-walk fills the pool until the kernel gives up. | | A wait for a slot runs out at the budget plus a quarter of it (at most 250 ms), so simultaneous timeouts of the calls it waits behind mark first. | Exactly the budget: a waiter queued in the same tick then gives up a moment before those calls and is refused unmarked, and the mark lands a few ms later. | | The keep-latest key reads run 16 at a time for every caller (they were an unbounded `Promise.all`). | Bound them only for the snapshot. | +| **Ruled after the re-review:** when every slot is held by a call past its budget, the next call is refused at once, and marks the location stalled only when the disk has completed nothing since the oldest of those calls began; four slow units on a disk still completing requests are "drive slow", refused and unmarked. With no device named (the stat detector), it marks. | Mark whenever every slot is overdue (the M1 fix as first built): four slow units on a busy disk then marked it stalled for 15–30 s. | | An auto-pause record with no `cause` reads as not there. | Word both cases for it ("not there or not answering"). | | Two counter samples closer than 10 s give no verdict (a Refresh just after a pass among them). **Kept at review.** | Compare any two samples (a busy healthy drive can read "in flight, nothing completed" over a few milliseconds). | | A findmnt that does not answer reuses the last device named for that root; one naming another volume's UUID names none. | Treat a findmnt that does not answer as a stall. | @@ -807,5 +812,6 @@ spin-up case stays the operator's question). Two new ones: |---|---|---| | M4: a slot wait was timed from when the call queued, so a deep queue on a busy but answering drive was refused as "not answering" (units of about 1.1 s at 16 wide, 0.46 s at 32, 0.2 s at 64) | The deadline follows progress: every return on the key re-arms its waiters, and a wait is refused only when nothing on the key has returned for the budget; the overdue and transition refusals unchanged. The limit statement and the keep-latest comment corrected | `d61bdd93`, this commit | | L11: the health pass's block sat inside the boot probe's comment, which still called itself the only thing an idle boot runs | Moved above it; the phrase dropped | `c367a3d7` | +| (the implementer's note) four slow units past the budget on a disk whose counters show completions made the overdue refusal mark the location stalled | The same treatment as L6: the overdue refusal compares the counters with those taken when the oldest overdue call began; moved → refused, not marked; unchanged → marked | `9a12308e`, this commit | ## Rollout