Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 027d0bf362aafc668d1df68b1d13aa2b53b9087b
parent 184777e4e6548bf98cece3cd8ebe80fc3e1a83d5
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Wed, 30 Sep 2026 00:38:08 -0400

plans: slice DS after the re-review — M4 (the slot wait's deadline follows progress) and L11 to their commits; the limit statement says a drive slower than the budget per unit is treated as not answering and that queue depth no longer is; FACTS on the queue's deadline and refreshed anchors; re-gates (tsc; common 2,299; editor unit 87; e2e storage + channels 38/38)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
Mplans/FACTS.md | 23+++++++++++++----------
Mplans/release-15.md | 45+++++++++++++++++++++++++++++++++------------
2 files changed, 46 insertions(+), 22 deletions(-)

diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -7635,35 +7635,38 @@ The record is [`release-15.md`](release-15.md), "Slice DS, as shipped", with the detect it reliably: the root's inode is in the kernel's cache whenever the drive was used lately. - **The state is `common/lib/storageHealth.ts`**, one map on `globalThis.__yttStorageHealth__` (the pass writes it from instrumentation's module copy; pages read it from theirs), with each entry's - `detector` and the counters' `device`. `recordLocationHealth` (`:172`): one `stalled` answer stalls + `detector` and the counters' `device`. `recordLocationHealth` (`:177`): one `stalled` answer stalls at once, and every transition to stalled refuses `onDrive`'s waiting calls; `HEALTH_CLEAN_TO_CLEAR` (2) clean answers in a row clear it; `absent` is clean; a new root starts over. - `registerLocationHealth` (`:257`) creates entries with no answer. `stalledLocationForPath` - (`:309`) matches like `locationOfDataDir`; `stalledLocation` (`:325`) is by id AND root. + `registerLocationHealth` (`:262`) creates entries with no answer. `stalledLocationForPath` + (`:314`) matches like `locationOfDataDir`; `stalledLocation` (`:330`) is by id AND root. - **Detector 1, every 15 s: the block device's counters** (`detectLocationHealth`, `lib/storageVolumes.ts:601`). The root's device from the last pass (findmnt `-J -T <root> -o SOURCE,UUID`, raced against 3 s, only when there is none or its `/sys` entry stops reading; another volume's UUID names none; `[…]` stripped, `/dev/mapper` resolved, basename); then - `/sys/class/block/<dev>/stat` (`parseBlockStat`, `storageHealth.ts:682`): completed = fields 1 + 5 + `/sys/class/block/<dev>/stat` (`parseBlockStat`, `storageHealth.ts:710`): completed = fields 1 + 5 + 12 + 16 (reads, writes, discards, flushes), in flight = field 9. Stalled ⇔ in flight at both samples AND nothing completed between; samples at least `MIN_COUNTER_INTERVAL_MS` (10 s) apart; the first gives no verdict. in_flight counts only requests dispatched to the driver: one requeued during a host reset is not counted, so a sample in that window can read clean (the watchdog covers it). No device → the child `stat -L -c %F` probe (`probeLocationHealth`, `:386`). The samples are on `globalThis.__yttHealthDetector__` (the pass and /storage's Refresh share them). -- **Detector 2, on every gated call: `onDrive(where, call)`** (`storageHealth.ts:584`). Refused with +- **Detector 2, on every gated call: `onDrive(where, call)`** (`storageHealth.ts:612`). Refused with no call on a stalled location; otherwise raced against `DRIVE_CALL_BUDGET_MS` (3 s; test seam `setDriveCallBudget`); the budget covers the whole unit passed in. A timeout marks the location stalled (since now) and throws `DriveNotAnsweringError`, leaving the call to settle — unless the location's device counters (read synchronously from `/sys` through the reader `storageVolumes.ts` registers with `setCounterReader`, `:509`) moved since the call began: then the call is refused as slow and nothing is marked. At most `DRIVE_CALLS_IN_FLIGHT` (4) calls per slot key in flight - (`acquireSlot`, `:488`): the rest queue in JS, raced against the budget plus a quarter of it (at - most 250 ms) and refused unmarked when that runs out; refused at once by any transition to stalled; - and when every slot is held by a call already past its budget (`overdue`), a new call is refused - at once and the location marked stalled again. A slot is freed when its call really returns. The + (`acquireSlot`, `storageHealth.ts:494`): the rest queue in JS. A waiting call's deadline follows progress: every call that + returns on the key (in time or late) restarts it (`releaseSlot`, `:578`, re-arms every waiter), and a + waiting call is refused unmarked only when nothing on the key has returned for the budget plus a + quarter of it (at most 250 ms) — never for the queue's depth alone. Waiting calls are refused at + once by any transition to stalled; and when every slot is held by a call already past its budget + (`overdue`), a new call is refused at once and the location marked stalled again (even if the + counters show the disk completing other requests). A slot is freed when its call really returns. The slot key: a configured location's id; a probe of another root under its id, that root (marks - nothing); a path on no configured location, the root it is under (`rootOfUnknownPath`, `:444`; + nothing); a path on no configured location, the root it is under (`rootOfUnknownPath`, `:450`; marks nothing). Do not nest it for one key. The timer is not unref'd. - **The cadence** is `runStorageHealthPass` (`controller/storageWatch.ts:433`): prune, register, then every location concurrently; every 15 s from `startStorageHealthWatch` (`:546`), plus one at arm diff --git a/plans/release-15.md b/plans/release-15.md @@ -535,9 +535,12 @@ stalled disk is here. drive, not 64. A slot is released when its call really returns. The queue (M1): - every transition to `stalled` refuses the waiting calls at once, whoever decided it (the pass, the watchdog, a Refresh); - - a wait is raced against the budget plus a grace of a quarter of it (at most 250 ms), so the - calls it waits behind, whose timers start a moment later, time out and mark first; a wait - that runs out is refused without marking; + - a wait's deadline follows progress (re-review M4): every call that returns on the key, in + time or late, restarts the deadline of every call waiting on it, and a waiting call is + refused, without marking, only when nothing on the key has returned for the budget plus a + grace of a quarter of it (at most 250 ms; the grace lets the calls it waits behind, whose + timers start a moment later, time out and mark first). A deep queue on a drive that is busy + but answering therefore waits as long as it takes; - calls past their budget are counted per key (`overdue`); when every slot is held by one, a new call is refused at once and the location marked stalled again, even if the pass has since cleared it: none of those calls has returned. @@ -644,14 +647,17 @@ stalled disk is here. | `ff235c2f` | `common:` review M3: the snapshot walk through `onDrive`; keep-latest keys bounded. Test. | | `19c5842d` | `editor:` review L1, L3: the Storage stage's `statfs` through `onDrive`; the stale comments. | | `30df4193` | Merge `main` (`bab894db`: release 14 HS, S1, CF; release 15 UT). Two conflicts, both kept: the changelog's `[Unreleased]` carries UT's bullet then DS's; this file is `main`'s with DS's table row and this section after UT's. | -| this commit | `plans:` the review's findings to their commits; FACTS; the changelog; the report. | +| `b34a7613` | `plans:` the review's findings to their commits; FACTS; the changelog. | +| `c367a3d7` | `editor:` re-review L11: the health pass's block above the boot probe's comment. | +| `d61bdd93` | `common:` re-review M4: a slot wait's deadline follows progress. Tests. | +| this commit | `plans:` the re-review's findings to their commits; FACTS; the report. | **Tests** (unit; no test stalls a real drive: a stalled call is a promise that never settles or a fake `stat` that never answers, and a stalled device is a temp `/sys` whose counters stand still) | File | What it pins | |---|---| -| `lib/storageHealth.test.ts` (22) | The rules: one miss stalls at once, one clean answer after a stall does not clear it and two in a row do, a miss in between starts over, `absent` is clean, a re-pointed root starts over, the path match is `locationOfDataDir`'s, pruning, `globalThis`, the "since" wording, registering without an answer. `onDrive`: an answer passes through (value or error) and frees its slot; a never-settling call is `stalled` on the timer, marks the location (since now) and keeps its slot until it settles; a stalled location is refused with no call; seven calls at once put four in flight and the three that waited are refused without a call when the location stalls; a freed slot runs a waiter; the 3 s default (answered after 3–4.5 s). After the review: a hand-typed root's four slots are shared by its channels and a fifth call is refused at once, with no entry made; a candidate root has its own slots and leaves the location's calls alone; **M1's case** (four hung calls, two clean answers, a fifth refused at once and the location marked again); a transition to stalled from the pass refuses the waiting call; **L6** (counters completing: four calls refused as slow, the location not marked, a waiter's wait runs out unmarked; counters standing still: marked). | +| `lib/storageHealth.test.ts` (26) | The rules: one miss stalls at once, one clean answer after a stall does not clear it and two in a row do, a miss in between starts over, `absent` is clean, a re-pointed root starts over, the path match is `locationOfDataDir`'s, pruning, `globalThis`, the "since" wording, registering without an answer. `onDrive`: an answer passes through (value or error) and frees its slot; a never-settling call is `stalled` on the timer, marks the location (since now) and keeps its slot until it settles; a stalled location is refused with no call; seven calls at once put four in flight and the three that waited are refused without a call when the location stalls; a freed slot runs a waiter; the 3 s default (answered after 3–4.5 s). After the review: a hand-typed root's four slots are shared by its channels and a fifth call is refused at once, with no entry made; a candidate root has its own slots and leaves the location's calls alone; **M1's case** (four hung calls, two clean answers, a fifth refused at once and the location marked again); a transition to stalled from the pass refuses the waiting call; **L6** (counters completing: four calls refused as slow, the location not marked, a waiter's wait runs out unmarked when nothing returns; counters standing still: marked). After the re-review (**M4**): a healthy 64-wide walk of units at half the budget (the last call waits about fifteen units) has no refusals and marks nothing, on a location and on a hand-typed root; a queue behind four hung calls is still refused within the budget (marked on a location, unmarked on a hand-typed root); a call that returns late hands its slot to the calls waiting behind it. | | `lib/storageHealthCounters.test.ts` (7) | The stat line parser (17 and 11 fields, garbage); the verdict over sample pairs (stuck → stalled; moving, idle or drained → ok); device names (partition, `[subvolume]` suffix, non-`/dev` sources); the detector end to end with a fake findmnt and a temp `/sys`: the first sample gives no verdict, stuck → stalled with its cause, drained → ok, busy and moving → ok; a sample sooner than 10 s gives none and keeps the first; every no-device fallback goes to the child stat and says so (tmpfs, findmnt failing, no `/sys` entry, another volume's UUID, no binary). After the review: discards and flushes are counted (a flush alone is not a stall); a known device is read without running findmnt, and findmnt runs again only when its `/sys` entry stops reading; the samples are on `globalThis`. | | `lib/storageHealthProbe.test.ts` (5) | The child `stat`: a directory is `ok`, a missing path and a file are `absent`; a fake `stat` asleep for 20 s is `stalled` on the timer without being waited for; the 3 s default; an answer inside the budget is taken; no binary is `ok`. | | `controller/storageStall.test.ts` (22) | A spy on every `node:fs` and `node:fs/promises` call, with a hang mode that makes a matching promise-API call never settle. The gate: with the location stalled, inspect reads only the marker (with a config) or config.json and the marker (without); a channel mid-move on a stalled drive reads `in-transition` (and is remembered so); the guard, `probeLocation` and its memo (no findmnt run), `volumeFreeBytes`, `readChannelStat`, the recency layer, the move-root check, a snapshot refresh, the saved-video store and the inventory make no call on the drive. The watchdog: with the location answering and one drive call hung, inspect, `readChannelStat` (at most four video directories asked), `probeLocation`, `volumeFreeBytes`, the recency layer (not remembered as a miss) the move-root check and (M3) the snapshot walk each answer `stalled` within the race and mark the location (the snapshot writes no `snapshot.json`, and at most four video directories reach the drive); after it, inspect, the guard and the walk make no call on the drive, and once cleared the drive is asked again. The memo: two inspects inside 5 s stat the target once, `fresh` and a 5 s age ask again, a fresh answer is not stored, another target is another key, `forgetChannelMedia` and `clearRelocationMarker` clear it, and a stall is seen with an `ok` remembered. | @@ -660,13 +666,15 @@ fake `stat` that never answers, and a stalled device is a temp `/sys` whose coun #### Gates (logs `$T/ds-*.log`) -- **tsc** was clean before every commit and on the merged tree (65 s). +- **tsc** was clean before every commit and on the merged tree (65 s). After the re-review: 53 s, + once the worktree's gitignored `export/.next/dev` (a generated `validator.ts` an export dev server + had left truncated) was deleted; it is not a source file. - **Unit, on the merged tree (`30df4193`):** | Suite | Result | |---|---| - | common | **2,295/2,295**, 58 s: `main`'s 2,229 plus DS's 66 (60 at the rulings, 6 from the review). | - | editor unit | 87/87 | + | common | **2,295/2,295**, 58 s: `main`'s 2,229 plus DS's 66 (60 at the rulings, 6 from the review). After the re-review: **2,299/2,299**, 53 s (4 for M4). | + | editor unit | 87/87, and 87/87 after the re-review | | `test:scripts` | 194 passed, 2 skipped (196). The second skip is UT's post-build trace check, which skips a umtool build older than its config (this worktree's `umtool/.next` predates UT); the first is the `LIVE=1` archive check. | | mcp | 271/271 at the first pass; no mcp file has changed since. | @@ -683,6 +691,7 @@ fake `stat` that never answers, and a stalled device is a temp `/sys` whose coun | 2, first pass | `auto-queue`, `tags`, `saved-videos`, four cleanup specs, `video-titles`, `video-page` (`$T/ds-specs2.txt`) | **70 passed, 1 failed, 5.9 min**: `auto-queue.spec.ts:352`, the Start-button race its own comment describes. | | 3, first pass | `auto-queue` alone | **24 passed, 1.3 min**. | | 4, after the rulings (`b1a30902`) | the 9 `storage` + `channels` specs (`$T/ds-specs4.txt`) | **38 passed, 0 failed, 12 skipped, 2.5 min**. | + | 6, after the re-review (`d61bdd93`) | the 9 `storage` + `channels` specs (`$T/ds-specs4.txt`) | **38 passed, 0 failed, 12 skipped, 2.6 min**. | | 5, after the review, on the merged tree (`30df4193`) | the 9 `storage` + `channels` specs; the snapshot scheduler's (`auto-report-refresh`, `channel-work`, `jobs-batch-tasks-drain`, `incomplete-transcript`); the keep-latest users of the bounded key fan-out (`cleanup-holds`, `saved-videos`); `reconcile`; `review` (the auto-pause wording) — `$T/ds-specs5.txt` | **73 passed, 0 failed, 12 skipped, 5.2 min** (37 min 54 s in the queue behind another session's suite). | The 12 skips in each are `channels-rack-audit`, which needs `E2E_RACK_SHOTS`. No e2e fixture has @@ -712,10 +721,14 @@ fake `stat` that never answers, and a stalled device is a temp `/sys` whose coun up completes none), for 15–30 s, during which the start-of-work guard refuses the channel's jobs. **Does the drive spin down when idle?** If it does, a longer budget for the first call after an idle spell, or a spin-down timer on the drive, would avoid it. -- **A slow drive under load can fail a snapshot refresh.** Each video directory's unit is raced - against 3 s; a drive that is slow but completing requests is refused without being marked, and - the refresh throws, so the scheduler keeps the last `snapshot.json` and tries again on its next - trigger. +- **A drive slower than the budget per unit is treated as not answering.** A queue's depth no + longer matters (M4): a waiting call is refused only when nothing on the drive has returned for + 3 s. But a single unit that takes longer than 3 s is refused: without a mark when the counters + show the disk completing other requests, with one otherwise. A snapshot refresh with such a unit + throws, so the scheduler keeps the last `snapshot.json` and tries again on its next trigger. And + once four such slow units hold every slot past the budget, the next call is refused at once and + the location marked stalled (the M1 rule, kept as ruled), even when the counters show the disk + completing: the pass then clears it after two clean samples (15–30 s). - **The counters need two samples.** A drive already stalled when the editor starts is seen by the counters at the second pass (15–30 s), or at once by the watchdog when a page reaches it. in_flight counts only requests dispatched to the driver: a request requeued during a host reset @@ -787,4 +800,12 @@ on each finding were applied as below. | L10: a findmnt every pass | Only on a root change or a failed `/sys` read | `33094c29` | | The two questions | Keep the four-call cap; keep the 10 s spacing | as built | +**Re-review: SHIP AFTER FIXES.** Every finding above was confirmed closed (L6 for a busy drive; the +spin-up case stays the operator's question). Two new ones: + +| Finding | Ruling | Where | +|---|---|---| +| M4: a slot wait was timed from when the call queued, so a deep queue on a busy but answering drive was refused as "not answering" (units of about 1.1 s at 16 wide, 0.46 s at 32, 0.2 s at 64) | The deadline follows progress: every return on the key re-arms its waiters, and a wait is refused only when nothing on the key has returned for the budget; the overdue and transition refusals unchanged. The limit statement and the keep-latest comment corrected | `d61bdd93`, this commit | +| L11: the health pass's block sat inside the boot probe's comment, which still called itself the only thing an idle boot runs | Moved above it; the phrase dropped | `c367a3d7` | + ## Rollout