Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit b845ae30867285af38079e67f93e015b1d32b901
parent 82b5f9f1786bea9eb648e5d4526a480731fef03b
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Wed, 30 Sep 2026 09:37:33 -0400

Merge main (release 15 slice DT) into r15/stagit — release-15.md keeps IG, UT, DS, SS, DT, then SG before the Rollout, and both slices-table rows; the editor changelog keeps DT's bullet, then SG's

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
MSETTINGS.md | 11+++++++++++
Mcommon/bin/build-index.ts | 8+++++++-
Mcommon/bin/build-stats.ts | 8+++++++-
Mcommon/controller/channelSnapshot.ts | 3++-
Mcommon/controller/channels.ts | 5+++--
Mcommon/controller/recencyIndex.ts | 5+++--
Mcommon/controller/relocateChannelMedia.ts | 5+++--
Mcommon/controller/storageLocations.ts | 3++-
Mcommon/controller/storageWatch.test.ts | 70++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/controller/storageWatch.ts | 87++++++++++++++++++++++++++++++++++++++++++++++++++++++-------------------------
Mcommon/lib/channelMedia.ts | 8+++++---
Mcommon/lib/settingsDocs.ts | 11+++++++++++
Mcommon/lib/settingsSchema.ts | 6++++++
Mcommon/lib/storageHealth.test.ts | 151+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++----
Mcommon/lib/storageHealth.ts | 218+++++++++++++++++++++++++++++++++++++++++++++++++++++++++----------------------
Mcommon/lib/storageHealthCounters.test.ts | 37+++++++++++++++++++++++++++++++++----
Mcommon/lib/storageHealthProbe.test.ts | 20++++++++++++++++++--
Acommon/lib/storageHealthTimings.test.ts | 138+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/storageHealthTimings.ts | 166+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mcommon/lib/storageLocations.ts | 16+++++++++++++++-
Mcommon/lib/storageVolumes.ts | 28+++++++++++++++++-----------
Meditor/CHANGELOG.md | 1+
Meditor/app/channels/[slug]/components/MediaNotAnswering.tsx | 13+++++++++----
Meditor/app/channels/[slug]/page.tsx | 2+-
Meditor/app/channels/[slug]/videos/[id]/page.tsx | 6++++--
Meditor/app/channels/[slug]/videos/page.tsx | 5+++--
Meditor/app/channels/components/ChannelVolumeBar.tsx | 5++++-
Meditor/app/channels/page.tsx | 12+++++++++---
Meditor/app/settings/saveSettings.test.ts | 22++++++++++++++++++++++
Meditor/app/storage/actions.ts | 67++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++-------
Meditor/app/storage/buildStorage.ts | 9++++++---
Aeditor/app/storage/components/HealthTimingForm.tsx | 119+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aeditor/app/storage/lib/healthTimingsForm.test.ts | 95+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Aeditor/app/storage/lib/healthTimingsForm.ts | 91+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/app/storage/page.tsx | 6++++++
Meditor/e2e/storage-locations.spec.ts | 82+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/instrumentation.ts | 3++-
Mplans/FACTS.md | 95+++++++++++++++++++++++++++++++++++++++++++++++--------------------------------
Mplans/release-15.md | 210+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
39 files changed, 1660 insertions(+), 187 deletions(-)

diff --git a/SETTINGS.md b/SETTINGS.md @@ -522,6 +522,7 @@ Where a channel's downloaded media goes when it is relocated off the corpus disk | `locations` | `[]` | The named storage locations a channel's media may be relocated to — one entry per root, each with an id, label, root, `autoRepoint` and the learned volume identity. Order is display order. Managed on /storage. | | `defaultLocationId` | `""` | The location prefilled as the destination of a move. "" = no default. | | `savedVideosLocationId` | absent | WHERE THE SAVED-VIDEO STORE IS, by location id. "" = in place, under the corpus at `paths.savedVideosDir`.<br><br>A RECORD OF WHAT IS ON DISK, never an intention — the same contract as a channel's `config.dataDir`. It is written by the move, on success, after the copy has verified and the symlink is in place; nothing else writes it, and a reader that disagrees with the disk trusts the disk. Optional so an older settings.json parses (and an older binary that drops it leaves a store that still works, because the symlink is what every reader follows). | +| `health` | absent | THE DRIVE-HEALTH TIMINGS: how long a read may take before a drive counts as not answering, how often the health pass looks, how long its look may take, how many clean looks clear a stall, and how many reads may be on one drive at once. Edited on /storage (Drive health timing). Absent = every default, and only a value that differs from its default is written, so an untuned install follows a default changed later. See `storage.health` below. | #### `storage.locations[]` @@ -547,6 +548,16 @@ Per entry — each entry spells its own values. | `mountpoint` | Where the volume was mounted at the last successful probe, and the path of the location's root RELATIVE to that mountpoint. Invariant: `root === join(mountpoint, relPath)`. Keeping the two halves is what lets a probe compute a candidate root when the volume reappears elsewhere. | | `relPath` | The location root's path RELATIVE to `mountpoint` (see there). Invariant: `root === join(mountpoint, relPath)`. | +#### `storage.health` + +| Key | Default | Description | +|---|---|---| +| `budgetMs` | `3000` | How long one read may take before the drive counts as not answering, in ms (default 3000, 500–60000). The watchdog's budget (`onDrive`, lib/storageHealth.ts) for one unit of work — a video directory's reads, a page's reads of one video: a unit that has not answered by then is refused, and marks its location not answering unless the disk's request counters show it still completing others (slow, not stalled). A read waiting for a slot is refused when nothing on the drive has returned for this long plus a quarter of it (at most 250 ms). Takes effect on the next read after a save. | +| `passIntervalMs` | `15000` | How often the health pass reads each location's disk counters, in ms (default 15000, 5000–300000). A stall that starts between two passes is seen by the next, or at once by a page's read. A save on /storage re-arms the pass's timer at once; a hand edit, at the next pass. Two counter samples are compared only when at least min(10 s, this − 5 s) apart, a spacing never less than half of this. | +| `probeTimeoutMs` | `3000` | How long the health pass waits, in ms (default 3000, 500–30000), for the child `stat` of a root where no disk can be named (a timeout counts as not answering) and for the `findmnt` that names a root's disk (a timeout names none that pass). Takes effect on the next pass. | +| `clearAfterCleanPasses` | `2` | How many clean answers in a row clear a location marked not answering (default 2, 1–10). Each health pass is one answer, and so is a Refresh on /storage; a miss in between starts the count again. Takes effect on the next answer. | +| `inFlightPerLocation` | `4` | How many reads through the watchdog may be on one location's drive at once (default 4, 1–8); the rest wait in the editor's own queue, so a stall mid-walk holds this many of Node's threads, not all of them. At most 8, half of `UV_THREADPOOL_SIZE` (16 in the editor's start script and the container), so one drive that stops answering cannot hold every thread. Takes effect on the next read. | + Default: ```json diff --git a/common/bin/build-index.ts b/common/bin/build-index.ts @@ -3,10 +3,16 @@ // as `tsx bin/build-index.ts` (export's build:index script). import { getPaths, type Paths } from "../lib/paths"; import { buildIndex } from "../controller/buildIndex"; +import { settingsFromFile } from "../lib/settings"; +import { applyHealthTimings } from "../lib/storageHealth"; import { runIfEntryPoint } from "./_cli"; export async function main(opts: { paths?: Paths } = {}): Promise<void> { - await buildIndex({ paths: opts.paths ?? getPaths() }); + const paths = opts.paths ?? getPaths(); + // A CLI process has no health pass: the drive-health timings the build's + // watchdog runs on (settings.storage.health) are applied here, once. + applyHealthTimings(settingsFromFile(paths.settingsFile).storage.health); + await buildIndex({ paths }); } runIfEntryPoint(import.meta.url, () => main()); diff --git a/common/bin/build-stats.ts b/common/bin/build-stats.ts @@ -3,10 +3,16 @@ // build:stats script); `main` is exported for the archilyzer CLI. import { getPaths, type Paths } from "../lib/paths"; import { buildStats } from "../controller/buildStats"; +import { settingsFromFile } from "../lib/settings"; +import { applyHealthTimings } from "../lib/storageHealth"; import { runIfEntryPoint } from "./_cli"; export async function main(opts: { paths?: Paths } = {}): Promise<void> { - await buildStats({ paths: opts.paths ?? getPaths() }); + const paths = opts.paths ?? getPaths(); + // A CLI process has no health pass: the drive-health timings the build's + // watchdog runs on (settings.storage.health) are applied here, once. + applyHealthTimings(settingsFromFile(paths.settingsFile).storage.health); + await buildStats({ paths }); } runIfEntryPoint(import.meta.url, () => main()); diff --git a/common/controller/channelSnapshot.ts b/common/controller/channelSnapshot.ts @@ -743,7 +743,8 @@ export async function generateChannelSnapshot( // each video directory's unit below. At most four of them wait on that drive // at once — this walk runs in the editor's own process after every download // or sync of the channel, sixteen wide, which is exactly while a long write is - // stressing the drive — and one that does not answer in 3 s throws, so the + // stressing the drive — and one that does not answer within the budget + // (`storage.health.budgetMs`, 3 s by default) throws, so the // scheduler keeps the last good snapshot.json, as on any failed refresh. // The reconcile pass just below is sequential (one read at a time) and is // not raced. diff --git a/common/controller/channels.ts b/common/controller/channels.ts @@ -90,8 +90,9 @@ async function hasDigestWithItems(videoDir: string): Promise<boolean> { // `drive` is the channel's configured target (`config.dataDir`) when its media // is on another drive. Then every read goes through `onDrive`: none while that -// location is stalled, at most four in flight on it, and one that has not -// answered in 3 s marks it stalled — and the walk answers null ("the drive did +// location is stalled, at most `inFlightPerLocation` (4) in flight on it, and +// one that has not answered within the budget (`storage.health.budgetMs`, 3 s +// by default) marks it stalled — and the walk answers null ("the drive did // not answer") instead of counts. The rest of the walk is refused without a // call. An in-place channel's walk is on the corpus disk and is not wrapped. async function countDataFiles( diff --git a/common/controller/recencyIndex.ts b/common/controller/recencyIndex.ts @@ -199,8 +199,9 @@ async function readTailUploadDate(file: string): Promise<string | null> { // `drives` maps a channel whose media is on another drive to its configured // target. Its reads go through `onDrive`: none while that drive's location is // stalled (lib/storageHealth.ts) — each would hold an I/O thread until the drive -// came back, 32 at a time — at most four in flight on it, and one that has not -// answered in 3 s marks it stalled. An id not read for that reason is NOT +// came back, 32 at a time — at most `inFlightPerLocation` (4) in flight on it, +// and one that has not answered within the budget (3 s by default) marks it +// stalled. An id not read for that reason is NOT // memoized as a miss: it falls through to layers 3 and 4 for now, and a later // refresh with the drive answering reads it. const NOT_READ = Symbol("not read: the drive is not answering"); diff --git a/common/controller/relocateChannelMedia.ts b/common/controller/relocateChannelMedia.ts @@ -300,8 +300,9 @@ export async function relocationRootPresenceProblem( ); } - // Through the watchdog: a stat that has not answered in 3 s marks the - // location stalled and refuses the same way. + // Through the watchdog: a stat that has not answered within the budget + // (`storage.health.budgetMs`, 3 s by default) marks the location stalled + // and refuses the same way. let isDir = false; try { isDir = (await onDrive(named ?? r, () => stat(r))).isDirectory(); diff --git a/common/controller/storageLocations.ts b/common/controller/storageLocations.ts @@ -259,7 +259,8 @@ export async function volumeFreeBytes(opts: { // NOT ASKED WHILE ITS DRIVE IS NOT ANSWERING: each stat and the statfs // below would hold an I/O thread until it did. "—", like unmounted. The // calls that reach the drive go through `onDrive`'s watchdog; one that - // has not answered in 3 s marks the location stalled, and reads "—" too. + // has not answered within the budget (3 s by default) marks the location + // stalled, and reads "—" too. if (stalledLocation(loc)) { out[loc.id] = undefined; return; diff --git a/common/controller/storageWatch.test.ts b/common/controller/storageWatch.test.ts @@ -23,6 +23,8 @@ import { } from "./storageWatch"; import { inspectChannelMedia } from "../lib/channelMedia"; import { + applyHealthTimings, + healthTimings, locationHealth, resetStorageHealth, type LocationHealthState, @@ -35,6 +37,7 @@ import { beforeEach(() => { resetStorageWatchSuspicion(); resetStorageHealth(); + applyHealthTimings(); }); // Run with: @@ -589,3 +592,70 @@ test("a stall auto-pauses after two passes, and says the drive is not answering assert.match(String(autoPauseReasonOf(old, "a")), /drive that is not there/); }); }); + +// ── the timings are settings (release 15 slice DT) ───────────────────────── + +test("DT: every pass applies the timings it reads — the clear count and the log line follow storage.health", async () => { + await withTmp(async (h) => { + await seedRelocated(h, "slow", { targetExists: true }); + h.io.read().storage.health = { clearAfterCleanPasses: 3, budgetMs: 5_000 }; + const probe = scripted(["stalled", "ok", "ok", "ok"]); + const lines: string[] = []; + const pass = () => runStorageHealthPass({ io: h.io, probe, log: (l) => lines.push(l) }); + await pass(); + assert.equal(healthTimings().budgetMs, 5_000, "applied before anything was asked"); + assert.match(lines.join("\n"), /until it answers 3 times in a row/); + await pass(); + await pass(); + assert.equal(locationHealth("cold")?.state, "stalled", "two clean passes are not three"); + await pass(); + assert.equal(locationHealth("cold")?.state, "ok"); + // A pass handed its locations reads no settings: the timings stay. + await runStorageHealthPass({ locations: h.io.read().storage.locations, probe: scripted(["ok"]) }); + assert.equal(healthTimings().clearAfterCleanPasses, 3); + }); +}); + +test("DT: a changed pass interval re-arms the armed pass; an explicit interval follows nothing", async () => { + await withTmp(async (h) => { + let asked = 0; + const lines: string[] = []; + startStorageHealthWatch({ + io: h.io, + probe: async () => { + asked += 1; + return "ok"; + }, + log: (l) => lines.push(l), + }); + try { + assert.equal(asked, 1, "one pass at once"); + // The save on /storage: written to settings, then applied at once. + h.io.read().storage.health = { passIntervalMs: 5_000 }; + applyHealthTimings(h.io.read().storage.health); + assert.match(lines.join("\n"), /health pass re-armed: every 5 s/); + // At the default 15 s nothing would run for another 15 s; re-armed at + // 5 s, the next pass comes within about 5 s. + const started = Date.now(); + while (asked < 2 && Date.now() - started < 7_000) { + await new Promise((r) => setTimeout(r, 50)); + } + assert.equal(asked, 2, `a second pass after ${Date.now() - started} ms`); + assert.ok(Date.now() - started >= 4_500); + } finally { + stopStorageHealthWatch(); + } + // Stopped: a later change re-arms nothing. + const before = lines.length; + applyHealthTimings({ passIntervalMs: 20_000 }); + assert.equal(lines.length, before); + // Armed with an explicit interval, a change is not followed. + startStorageHealthWatch({ io: h.io, probe: async () => "ok", intervalMs: 60_000, log: (l) => lines.push(l) }); + try { + applyHealthTimings({ passIntervalMs: 6_000 }); + assert.equal(lines.some((l) => /re-armed/.test(l) && /6 s/.test(l)), false); + } finally { + stopStorageHealthWatch(); + } + }); +}); diff --git a/common/controller/storageWatch.ts b/common/controller/storageWatch.ts @@ -31,16 +31,18 @@ import { type VolumeBins, } from "../lib/storageVolumes"; import { - HEALTH_PROBE_INTERVAL_MS, - HEALTH_PROBE_TIMEOUT_MS, NOT_ANSWERING, + applyHealthTimings, + healthTimings, noteLocationDetector, + onPassIntervalChange, pruneLocationHealth, recordLocationHealth, registerLocationHealth, type HealthTransition, type LocationHealthState, } from "../lib/storageHealth"; +import { clearRuleText, secondsText } from "../lib/storageHealthTimings"; import { listChannelConfigs } from "./channels"; import { maybeAutoRepoint, probeAllLocations } from "./storageLocations"; @@ -361,12 +363,17 @@ export const STORAGE_WATCH_INTERVAL_MS = 5 * 60_000; // cable pulled out. It is no help for a drive that is here and stalled — an // SMR disk in a USB enclosure resetting under a long write — because every // in-process call on that drive waits for it, and four waits stop the editor -// answering at all. So every 15 s this reads each location's block device -// counters in /sys (`detectLocationHealth`; a child `stat` of the root only -// where no device can be named) and records the answer in `lib/storageHealth.ts`, -// which every page and poll consults before it touches a drive. One `stalled` -// answer marks a location at once; two clean answers in a row clear it (the -// rules are that module's). +// answering at all. So every `storage.health.passIntervalMs` (15 s by default) +// this reads each location's block device counters in /sys +// (`detectLocationHealth`; a child `stat` of the root only where no device can +// be named) and records the answer in `lib/storageHealth.ts`, which every page +// and poll consults before it touches a drive. One `stalled` answer marks a +// location at once; `clearAfterCleanPasses` clean answers in a row (two by +// default) clear it (the rules are that module's). +// +// EVERY PASS APPLIES THE TIMINGS from the settings it reads (`applyHealthTimings`), +// so a value changed by hand takes effect within one pass, and a changed +// interval re-arms the pass's own timer (below). // // READ-ONLY AND IN MEMORY. It writes no settings and pauses nothing: the pass // above sees a stalled location as down (its probe answers `stalled` without @@ -382,7 +389,7 @@ export type StorageHealthPassOpts = { // findmnt, for naming each root's block device. Default: getPaths(). bins?: Pick<VolumeBins, "findmntBin">; // Test seam. Default: `detectLocationHealth` — the block device's counters, - // or a child `stat` against a 3 s timer when no device can be named. + // or a child `stat` against `probeTimeoutMs` when no device can be named. probe?: LocationHealthProbe; now?: () => number; log?: (line: string) => void; @@ -401,7 +408,9 @@ function asVerdict(answer: LocationHealthState | HealthVerdict): HealthVerdict { return typeof answer === "string" ? { answer } : answer; } -const STAT_CAUSE = `a stat of its root did not answer within ${HEALTH_PROBE_TIMEOUT_MS / 1000} s`; +function statCause(): string { + return `a stat of its root did not answer within ${secondsText(healthTimings().probeTimeoutMs)}`; +} // Record one verdict. A verdict with no answer records nothing but the // detector that gave it. @@ -424,7 +433,7 @@ function recordVerdict( } return recordLocationHealth(loc, verdict.answer, { now, - cause: verdict.cause ?? STAT_CAUSE, + cause: verdict.cause ?? statCause(), ...(verdict.detector ? { detector: verdict.detector } : {}), ...(device !== undefined ? { device } : {}), }); @@ -434,8 +443,12 @@ export async function runStorageHealthPass( opts: StorageHealthPassOpts = {}, ): Promise<StorageHealthPassResult> { const log = opts.log ?? (() => {}); - const locations = - opts.locations ?? (opts.io ?? DEFAULT_IO).read().storage.locations; + // The settings this pass runs on: its locations, and the drive-health + // timings, applied before anything is asked. A caller that hands in the + // locations reads no settings, and the timings stay as they were. + const storage = opts.locations ? null : (opts.io ?? DEFAULT_IO).read().storage; + if (storage) applyHealthTimings(storage.health); + const locations = opts.locations ?? storage?.locations ?? []; pruneLocationHealth(locations.map((l) => l.id)); // Every configured location has an entry before anything is asked, so the // watchdog (lib/storageHealth.ts `onDrive`) can find a channel's location @@ -467,8 +480,9 @@ export async function runStorageHealthPass( out.transitions.push(t); if (t.to === "stalled") { log( - `[storage] "${loc.id}": ${NOT_ANSWERING} — ${verdict.cause ?? STAT_CAUSE}; ` + - `pages and polls skip it until two passes in a row find it answering`, + `[storage] "${loc.id}": ${NOT_ANSWERING} — ${verdict.cause ?? statCause()}; ` + + `pages and polls skip it until it answers ` + + `${clearRuleText(healthTimings().clearAfterCleanPasses)}`, ); } else if (t.from === "stalled") { log(`[storage] "${loc.id}": answering again (${t.to})`); @@ -481,7 +495,7 @@ export async function runStorageHealthPass( // operator pressing it after doing something about the drive gets an answer // taken afterwards. It counts as one answer like any other — a stalled location // still needs two clean ones in a row — and the counters give none when their -// last sample is under MIN_COUNTER_INTERVAL_MS old. Nothing is pruned. +// last sample is under `minCounterIntervalMs()` old. Nothing is pruned. export async function refreshLocationHealth( loc: StorageLocation, probe?: LocationHealthProbe, @@ -500,11 +514,10 @@ export async function refreshLocationHealth( // The cadence // --------------------------------------------------------------------------- -// Fifteen seconds (lib/storageHealth.ts says why). Each pass reads each -// location's device counters in /sys (a findmnt only when the device is not -// known yet or its /sys entry stopped reading; with no device, one short-lived -// `stat`). -export const STORAGE_HEALTH_INTERVAL_MS = HEALTH_PROBE_INTERVAL_MS; +// `storage.health.passIntervalMs`, fifteen seconds by default +// (lib/storageHealth.ts says why). Each pass reads each location's device +// counters in /sys (a findmnt only when the device is not known yet or its /sys +// entry stopped reading; with no device, one short-lived `stat`). // A per-module-copy singleton, deliberately left so: it is not a temp-file // name (slice W folded every tmp + rename onto lib/jsonFile-server.ts, whose @@ -514,6 +527,7 @@ export const STORAGE_HEALTH_INTERVAL_MS = HEALTH_PROBE_INTERVAL_MS; let timer: ReturnType<typeof setInterval> | null = null; let healthTimer: ReturnType<typeof setInterval> | null = null; let healthInFlight = false; +let stopFollowingInterval: (() => void) | null = null; // THE FIVE-MINUTE PASS, ARMED ONCE PER PROCESS, below the idle gate: it writes // settings (an auto-pause). `unref()` so it never holds the event loop open — a @@ -539,10 +553,16 @@ export function stopStorageWatch(): void { timer = null; } -// THE FIFTEEN-SECOND HEALTH PASS, ARMED ONCE PER PROCESS, above the idle gate: -// it is in memory and writes nothing (see the section header). It also runs -// once at once, so the locations are registered and a drive that is already -// stalled is on its way to being known before the first page. +// THE HEALTH PASS, ARMED ONCE PER PROCESS, above the idle gate: it is in +// memory and writes nothing (see the section header). It also runs once at +// once, so the locations are registered and a drive that is already stalled is +// on its way to being known before the first page. +// +// ITS INTERVAL FOLLOWS THE SETTING. Armed at `healthTimings().passIntervalMs`, +// and RE-ARMED whenever an applied change moves it — a save on /storage (at +// once, from the page's module copy: the subscription is on globalThis), or a +// hand edit (at the next pass, which applies what it read). An explicit +// `intervalMs` (the tests') is fixed and follows nothing. export function startStorageHealthWatch( opts: { io?: { read: () => SiteSettings }; @@ -573,8 +593,19 @@ export function startStorageHealthWatch( healthInFlight = false; }); }; - healthTimer = setInterval(health, opts.intervalMs ?? STORAGE_HEALTH_INTERVAL_MS); - healthTimer.unref?.(); + const arm = (every: number) => { + if (healthTimer) clearInterval(healthTimer); + healthTimer = setInterval(health, every); + healthTimer.unref?.(); + }; + arm(opts.intervalMs ?? healthTimings().passIntervalMs); + if (opts.intervalMs === undefined) { + stopFollowingInterval = onPassIntervalChange((every) => { + if (!healthTimer) return; + (opts.log ?? console.log)(`[storage] health pass re-armed: every ${secondsText(every)}`); + arm(every); + }); + } health(); return true; } @@ -582,4 +613,6 @@ export function startStorageHealthWatch( export function stopStorageHealthWatch(): void { if (healthTimer) clearInterval(healthTimer); healthTimer = null; + stopFollowingInterval?.(); + stopFollowingInterval = null; } diff --git a/common/lib/channelMedia.ts b/common/lib/channelMedia.ts @@ -3,14 +3,15 @@ import { lstat, readFile, readlink, rm, stat } from "node:fs/promises"; import type { Paths } from "./paths"; import type { ChannelConfig } from "./channelConfig"; import { - DRIVE_CALL_BUDGET_MS, NOT_ANSWERING, + healthTimings, isDriveNotAnswering, onDrive, sinceText, stalledLocationForPath, type LocationHealth, } from "./storageHealth"; +import { secondsText } from "./storageHealthTimings"; // WHERE A CHANNEL'S MEDIA ACTUALLY IS, and whether it can be reached. // @@ -221,7 +222,7 @@ export function stalledMediaLocation( detail: health ? `${NOT_ANSWERING} (location "${health.label}", ${sinceText(health.since)})` : (detail ?? - `${NOT_ANSWERING} (a read did not answer within ${DRIVE_CALL_BUDGET_MS / 1000} s)`), + `${NOT_ANSWERING} (a read did not answer within ${secondsText(healthTimings().budgetMs)})`), }; } @@ -440,7 +441,8 @@ async function inspectOnDisk( // // THE ONE CALL HERE THAT REACHES THE DRIVE, so it goes through the // watchdog: not made while the location is stalled, and a stat that has - // not answered in 3 s marks it stalled and answers `stalled` now. + // not answered within the budget (`storage.health.budgetMs`, 3 s by + // default) marks it stalled and answers `stalled` now. try { const st = await onDrive(configured, () => stat(configured)); if (!st.isDirectory()) { diff --git a/common/lib/settingsDocs.ts b/common/lib/settingsDocs.ts @@ -48,6 +48,10 @@ import { STORAGE_SETTINGS_FIELD_DOCS, STORAGE_VOLUME_FIELD_DOCS, } from "./storageLocations"; +import { + HEALTH_TIMING_DEFAULTS, + STORAGE_HEALTH_SETTINGS_FIELD_DOCS, +} from "./storageHealthTimings"; // `workers` IS LEFT OUT OF THE EXAMPLE, and that is the one place the example // is not the literal default object. Its default is `[]`, and a settings.json @@ -172,6 +176,13 @@ export function blockTables(d: SiteSettings): Partial<Record<keyof SiteSettings, }, { path: "storage.locations[]", docs: STORAGE_LOCATION_FIELD_DOCS }, { path: "storage.locations[].volume", docs: STORAGE_VOLUME_FIELD_DOCS }, + // Absent from the default block (only a tuned value is written), so the + // Default column is each timing's default, not the block's. + { + path: "storage.health", + docs: STORAGE_HEALTH_SETTINGS_FIELD_DOCS, + defaults: fromObject(HEALTH_TIMING_DEFAULTS), + }, ], buildPipeline: [ { diff --git a/common/lib/settingsSchema.ts b/common/lib/settingsSchema.ts @@ -59,6 +59,7 @@ import { type StorageSettings, type StorageVolume, } from "./storageLocations"; +import { sanitizeStorageHealth } from "./storageHealthTimings"; import { DEFAULT_DIARIZATION_ENGINE, DEFAULT_DIARIZATION_THRESHOLD, @@ -978,10 +979,15 @@ export function sanitizeStorage(value: unknown): StorageSettings { const savedVideosLocationId = locations.some((l) => l.id === savedWanted) ? savedWanted : ""; + // THE DRIVE-HEALTH TIMINGS: each clamped into its range, and kept only where + // it differs from its default; a block with nothing left is not written + // (lib/storageHealthTimings.ts). + const health = sanitizeStorageHealth(r.health); return { locations, defaultLocationId, ...(savedVideosLocationId ? { savedVideosLocationId } : {}), + ...(Object.keys(health).length > 0 ? { health } : {}), }; } diff --git a/common/lib/storageHealth.test.ts b/common/lib/storageHealth.test.ts @@ -1,15 +1,15 @@ import { beforeEach, test } from "node:test"; import assert from "node:assert/strict"; import { - DRIVE_CALLS_IN_FLIGHT, - DRIVE_CALL_BUDGET_MS, DriveNotAnsweringError, - HEALTH_CLEAN_TO_CLEAR, allLocationHealth, + applyHealthTimings, driveCallsInFlight, + healthTimings, isDriveNotAnswering, noteLocationDetector, onDrive, + onPassIntervalChange, registerLocationHealth, setCounterReader, setDriveCallBudget, @@ -22,6 +22,7 @@ import { stalledLocation, stalledLocationForPath, } from "./storageHealth"; +import { HEALTH_TIMING_DEFAULTS } from "./storageHealthTimings"; // Run with: // pnpm --filter yt-dlp-transcript-common exec tsx --test lib/storageHealth.test.ts @@ -32,6 +33,7 @@ import { beforeEach(() => { resetStorageHealth(); + applyHealthTimings(); setDriveCallBudget(); setCounterReader(undefined); }); @@ -55,7 +57,8 @@ test("a location first seen stalled is stalled", () => { }); test("one clean probe after a stall does not clear it; two in a row do", () => { - assert.equal(HEALTH_CLEAN_TO_CLEAR, 2); + assert.equal(HEALTH_TIMING_DEFAULTS.clearAfterCleanPasses, 2); + assert.equal(healthTimings().clearAfterCleanPasses, 2); recordLocationHealth(USB, "ok", { now: 0 }); recordLocationHealth(USB, "stalled", { now: 10 }); assert.equal(recordLocationHealth(USB, "ok", { now: 20 }), null); @@ -209,7 +212,7 @@ test("a stalled location is refused without the call being made", async () => { }); test("at most four calls in flight on a location; the rest wait, and are refused without a call when it stalls", async () => { - assert.equal(DRIVE_CALLS_IN_FLIGHT, 4); + assert.equal(HEALTH_TIMING_DEFAULTS.inFlightPerLocation, 4); registerLocationHealth([USB]); setDriveCallBudget(80); let made = 0; @@ -380,7 +383,8 @@ test("L6: a call that times out while its disk is still completing requests is s }); test("the default budget is 3 s", async () => { - assert.equal(DRIVE_CALL_BUDGET_MS, 3_000); + assert.equal(HEALTH_TIMING_DEFAULTS.budgetMs, 3_000); + assert.equal(healthTimings().budgetMs, 3_000); registerLocationHealth([USB]); const started = Date.now(); await assert.rejects(() => onDrive(USB, never), DriveNotAnsweringError); @@ -542,3 +546,138 @@ test("every slot held by an overdue call on a disk completing nothing: the next assert.equal(locationHealth("usb")?.state, "stalled"); assert.match(String(locationHealth("usb")?.cause), /4 reads on it have not answered/); }); + +// ── the timings are settings (release 15 slice DT) ───────────────────────── +// `applyHealthTimings` takes the stored `settings.storage.health` block, as the +// health pass and /storage's save hand it over; `healthTimings()` is what every +// number above is read through. + +test("DT: the applied budget feeds the watchdog — a unit slower than it is refused, and the same unit passes on the default", async () => { + registerLocationHealth([USB]); + // The settings' floor, 500 ms (a stored 200 clamps to it; storageHealthTimings.test.ts). + applyHealthTimings({ budgetMs: 200 }); + assert.equal(healthTimings().budgetMs, 500); + const unit = () => sleep(700).then(() => "done"); + await assert.rejects(() => onDrive(USB, unit), DriveNotAnsweringError); + assert.equal(locationHealth("usb")?.state, "stalled"); + assert.match(String(locationHealth("usb")?.cause), /did not answer within 0.5 s/); + // Back on the defaults (3 s), with the location clear, the same unit answers. + await sleep(250); + applyHealthTimings(); + resetStorageHealth(); + registerLocationHealth([USB]); + assert.equal(await onDrive(USB, unit), "done"); + assert.equal(locationHealth("usb")?.state, "ok"); +}); + +test("DT: the applied cap — with inFlightPerLocation 2, the third call waits for a slot", async () => { + registerLocationHealth([USB]); + applyHealthTimings({ inFlightPerLocation: 2 }); + const gates = Array.from({ length: 2 }, () => deferred<number>()); + const first = gates.map((g) => onDrive(USB, () => g.promise)); + let third = false; + const waiting = onDrive(USB, async () => { + third = true; + return 3; + }); + await new Promise((r) => setImmediate(r)); + assert.equal(driveCallsInFlight("usb"), 2); + assert.equal(third, false, "the third call is queued, not on the drive"); + gates[0].resolve(1); + assert.equal(await waiting, 3); + gates[1].resolve(2); + await Promise.all(first); + assert.equal(driveCallsInFlight("usb"), 0); +}); + +test("DT: a raised cap admits the calls already waiting; a lowered one is reached as calls return", async () => { + registerLocationHealth([USB]); + applyHealthTimings({ inFlightPerLocation: 1 }); + const gates = Array.from({ length: 3 }, () => deferred<number>()); + let made = 0; + const calls = gates.map((g) => + onDrive(USB, () => { + made += 1; + return g.promise; + }), + ); + await new Promise((r) => setImmediate(r)); + assert.equal(made, 1); + // Raised to 3 on a save: the two waiting take the new slots now. + applyHealthTimings({ inFlightPerLocation: 3 }); + await new Promise((r) => setImmediate(r)); + assert.equal(made, 3); + assert.equal(driveCallsInFlight("usb"), 3); + // Lowered to 1 with three in flight: a fourth call waits, and a return gives + // its slot back rather than handing it on while more than one is in flight. + applyHealthTimings({ inFlightPerLocation: 1 }); + let fourth = false; + const late = onDrive(USB, async () => { + fourth = true; + return 4; + }); + gates[0].resolve(0); + gates[1].resolve(0); + await new Promise((r) => setImmediate(r)); + assert.equal(fourth, false, "two returns bring three down to one: no slot for the fourth yet"); + assert.equal(driveCallsInFlight("usb"), 1); + gates[2].resolve(0); + assert.equal(await late, 4); + await Promise.all(calls); + assert.equal(driveCallsInFlight("usb"), 0); +}); + +test("DT: the applied clear count — three clean answers with clearAfterCleanPasses 3, one with 1", () => { + applyHealthTimings({ clearAfterCleanPasses: 3 }); + recordLocationHealth(USB, "stalled", { now: 0 }); + recordLocationHealth(USB, "ok", { now: 1 }); + recordLocationHealth(USB, "ok", { now: 2 }); + assert.equal(locationHealth("usb")?.state, "stalled", "two are not three"); + assert.deepEqual(recordLocationHealth(USB, "ok", { now: 3 }), { + id: "usb", + from: "stalled", + to: "ok", + }); + applyHealthTimings({ clearAfterCleanPasses: 1 }); + recordLocationHealth(USB, "stalled", { now: 4 }); + recordLocationHealth(USB, "ok", { now: 5 }); + assert.equal(locationHealth("usb")?.state, "ok"); +}); + +test("DT: a changed pass interval is told to the subscribers; the same one, or another timing, is not", () => { + const heard: number[] = []; + const stop = onPassIntervalChange((ms) => heard.push(ms)); + try { + applyHealthTimings({ passIntervalMs: 30_000 }); + applyHealthTimings({ passIntervalMs: 30_000, budgetMs: 5_000 }); + applyHealthTimings(); + assert.deepEqual(heard, [30_000, 15_000]); + } finally { + stop(); + } + applyHealthTimings({ passIntervalMs: 60_000 }); + assert.deepEqual(heard, [30_000, 15_000], "unsubscribed"); +}); + +test("DT: the test seam's budget wins over the applied one, and the timings survive a reset", () => { + applyHealthTimings({ budgetMs: 8_000, inFlightPerLocation: 6 }); + setDriveCallBudget(80); + assert.equal(healthTimings().budgetMs, 80); + setDriveCallBudget(); + assert.equal(healthTimings().budgetMs, 8_000); + resetStorageHealth(); + assert.equal(healthTimings().inFlightPerLocation, 6, "configuration, not health"); +}); + +test("DT review L4: a timeout on no known location names the budget the call ran against", async () => { + setDriveCallBudget(80); + const refused = onDrive("/hand/typed/ch/data", never).then( + () => "answered", + (err: Error) => err.message, + ); + // Changed once the call is out (it reads its budget after taking a slot): + // its words keep the budget it was given. + await new Promise((r) => setImmediate(r)); + setDriveCallBudget(5_000); + assert.equal(await refused, "drive not answering (a read did not answer within 0.08 s)"); +}); diff --git a/common/lib/storageHealth.ts b/common/lib/storageHealth.ts @@ -1,4 +1,11 @@ import { locationOfDataDir, type StorageLocation } from "./storageLocations"; +import { + HEALTH_TIMING_DEFAULTS, + resolveHealthTimings, + secondsText, + type HealthTimings, + type StorageHealthSettings, +} from "./storageHealthTimings"; // IS A STORAGE LOCATION'S DRIVE ANSWERING RIGHT NOW — the in-memory answer every // page and poll asks before it touches the drive. @@ -13,11 +20,14 @@ import { locationOfDataDir, type StorageLocation } from "./storageLocations"; // // SO SOMETHING THAT CANNOT HANG ASKS, AND THIS MODULE REMEMBERS WHAT IT SAID. // Two detectors write here. The health pass (`controller/storageWatch.ts`, -// every 15 s) reads each location's block device counters in /sys, which never -// touch the drive (`detectLocationHealth` in `storageVolumes.ts`; a child `stat` -// of the root raced against 3 s only where no device can be named). And -// `onDrive` below races every in-process call the gate covers against 3 s and -// marks the location the moment one does not answer. Everything that would +// every `passIntervalMs`) reads each location's block device counters in /sys, +// which never touch the drive (`detectLocationHealth` in `storageVolumes.ts`; a +// child `stat` of the root raced against `probeTimeoutMs` only where no device +// can be named). And `onDrive` below races every in-process call the gate +// covers against `budgetMs` and marks the location the moment one does not +// answer. The numbers are `settings.storage.health`, read through +// `healthTimings()` below (by default a pass every 15 s, a 3 s probe and a 3 s +// budget). Everything that would // touch the drive in-process asks this state first and, on a stalled location, // answers without the call: `inspectChannelMedia` reports `stalled`, the // free-space column reads "—", the probe reads "Not answering", the recency @@ -27,12 +37,12 @@ import { locationOfDataDir, type StorageLocation } from "./storageLocations"; // - ONE `stalled` answer marks the location `stalled` at once. A drive that // did not answer will not answer the next page either, and every page that // asks costs a thread. -// - TWO consecutive clean answers clear it. A clean answer is anything else: -// the counters moving or idle, or a child `stat` answering in time whether -// the root was there (`ok`) or not (`absent`: an unmounted drive answers -// ENOENT at once, and that is a different problem, which -// `inspectChannelMedia` already reports). One clean answer in the middle of -// a reset loop is not recovery. +// - `clearAfterCleanPasses` (TWO by default) consecutive clean answers clear +// it. A clean answer is anything else: the counters moving or idle, or a +// child `stat` answering in time whether the root was there (`ok`) or not +// (`absent`: an unmounted drive answers ENOENT at once, and that is a +// different problem, which `inspectChannelMedia` already reports). One +// clean answer in the middle of a reset loop is not recovery. // - A ROOT CHANGE (a re-point) starts the location over: the old root's stall // says nothing about the new one. // @@ -42,6 +52,13 @@ import { locationOfDataDir, type StorageLocation } from "./storageLocations"; // boot too) registers the locations; the watchdog re-learns a stall the moment // a page reaches the drive. // +// THE TIMINGS ARE SETTINGS (`settings.storage.health`, release 15 slice DT), +// held here as the process last applied them: the health pass applies them +// from the settings it reads on every pass, and /storage's save applies them at +// once. A process with no pass (a CLI) runs on the defaults unless its entry +// applies them (the index and stats bins do). `healthTimings()` is the one +// accessor. This module still reads no file. +// // ONE MAP PER PROCESS, NOT PER MODULE COPY. The watch that probes is armed from // `editor/instrumentation.ts`, and the pages that read are another bundle // layer; Next can load this module once for each. The map lives on @@ -78,7 +95,7 @@ export type LocationHealth = { // When the last probe (or observation) was recorded. checkedAt: number; // Consecutive clean answers since the location was marked stalled. It clears - // at HEALTH_CLEAN_TO_CLEAR. + // at `clearAfterCleanPasses`. cleanStreak: number; // What did not answer, for a stalled location: the probe's own words. cause?: string; @@ -87,16 +104,16 @@ export type LocationHealth = { device?: string; }; -// The probe's budget. A `stat` of a directory on a healthy disk answers in -// microseconds; three seconds is a thousand times that, and short enough that -// a page asking during a stall has not waited long. -export const HEALTH_PROBE_TIMEOUT_MS = 3_000; -// How often the watch asks. Short, because the gate is only as current as the -// last answer: a stall that began just after a probe costs every page that -// touches the drive until the next one. -export const HEALTH_PROBE_INTERVAL_MS = 15_000; -// Clean answers in a row that clear a stall. -export const HEALTH_CLEAN_TO_CLEAR = 2; +// THE DEFAULTS, and why they are what they are (lib/storageHealthTimings.ts +// holds them, with their ranges): +// - `probeTimeoutMs` 3 s: a `stat` of a directory on a healthy disk answers in +// microseconds; three seconds is a thousand times that, and short enough +// that a page asking during a stall has not waited long. +// - `passIntervalMs` 15 s: short, because the gate is only as current as the +// last answer — a stall that began just after a pass costs every page that +// touches the drive until the next one. +// - `clearAfterCleanPasses` 2: one clean answer in a reset loop is not +// recovery. // The one wording of the state, for every surface that shows it. export const NOT_ANSWERING = "drive not answering"; @@ -123,7 +140,13 @@ type HealthState = { inFlight?: Map<string, number>; overdue?: Map<string, OverdueCall[]>; waiters?: Map<string, Waiter[]>; - // Test seam: the watchdog's budget. + // The timings as last applied (`applyHealthTimings`); absent = the defaults. + timings?: HealthTimings; + // Told when the pass interval changes, so the armed pass re-arms its timer + // (controller/storageWatch.ts). On globalThis like the rest: the save that + // changes it runs in a page's module copy, the pass in instrumentation's. + intervalListeners?: Set<(passIntervalMs: number) => void>; + // Test seam: the watchdog's budget, below the settings' 500 ms floor. budgetMs?: number; // Reads a block device's counters, synchronously and without touching the // drive (set by `storageVolumes.ts`, which owns /sys). Absent: no check. @@ -136,7 +159,7 @@ declare global { } type FilledState = HealthState & - Required<Pick<HealthState, "inFlight" | "overdue" | "waiters">>; + Required<Pick<HealthState, "inFlight" | "overdue" | "waiters" | "intervalListeners">>; function healthState(): FilledState { if (!globalThis.__yttStorageHealth__) { @@ -148,11 +171,57 @@ function healthState(): FilledState { s.inFlight ??= new Map(); s.overdue ??= new Map(); s.waiters ??= new Map(); + s.intervalListeners ??= new Set(); return s as FilledState; } +// THE ONE ACCESSOR of the drive-health timings: `settings.storage.health` as +// this process last applied it, every absent key its default. Read at the +// moment a number is needed (a call's timer, a slot, an answer's count, a +// pass's probe), so an applied change takes effect on the next of each. +export function healthTimings(): HealthTimings { + const s = healthState(); + const t = s.timings ?? HEALTH_TIMING_DEFAULTS; + return s.budgetMs === undefined ? t : { ...t, budgetMs: s.budgetMs }; +} + +// Apply `settings.storage.health` (the stored block; absent = every default). +// The health pass calls this with the settings it reads on every pass, and +// /storage's save calls it at once. A changed pass interval is told to the +// armed pass (`onPassIntervalChange`); a raised cap lets calls already waiting +// for a slot take the new ones. +export function applyHealthTimings(stored?: StorageHealthSettings): HealthTimings { + const s = healthState(); + const before = healthTimings(); + s.timings = resolveHealthTimings(stored); + const after = healthTimings(); + if (after.inFlightPerLocation > before.inFlightPerLocation) { + for (const key of [...s.waiters.keys()]) admitWaiters(key); + } + if (after.passIntervalMs !== before.passIntervalMs) { + for (const fn of [...s.intervalListeners]) { + try { + fn(after.passIntervalMs); + } catch { + /* a listener that throws must not stop a save */ + } + } + } + return after; +} + +// Subscribe to pass-interval changes; returns the unsubscribe. +export function onPassIntervalChange(fn: (passIntervalMs: number) => void): () => void { + const set = healthState().intervalListeners; + set.add(fn); + return () => { + set.delete(fn); + }; +} + // Test seam, and the escape hatch for a process that wants to forget. Calls -// still waiting for a slot are released to run. +// still waiting for a slot are released to run. The applied timings are +// configuration, not health, and are kept. export function resetStorageHealth(): void { const s = healthState(); s.byId.clear(); @@ -246,7 +315,7 @@ function recordAnswer( } if (prev.state === "stalled") { prev.cleanStreak += 1; - if (prev.cleanStreak < HEALTH_CLEAN_TO_CLEAR) return null; + if (prev.cleanStreak < healthTimings().clearAfterCleanPasses) return null; prev.state = answer; prev.since = now; prev.cleanStreak = 0; @@ -362,17 +431,18 @@ export function notAnsweringText(h: Pick<LocationHealth, "since">, now?: number) } // --------------------------------------------------------------------------- -// The watchdog: every call the gate covers, raced against 3 s +// The watchdog: every call the gate covers, raced against the budget // --------------------------------------------------------------------------- // -// THE DETECTOR THAT CANNOT BE FOOLED BY A CACHE. The 15 s pass reads the -// block device's counters (or, with no device, a child `stat`), and a stall -// that starts between two passes is not seen by it. A page or a poll that -// actually reaches the drive is: `onDrive` runs the call against a 3 s timer, -// and a call that has not answered by then marks its location `stalled` at -// once (since now) and throws `DriveNotAnsweringError`. The call itself is left -// to settle on its own: its thread is held until the drive answers, which is -// the stated limit. +// THE DETECTOR THAT CANNOT BE FOOLED BY A CACHE. The pass reads the block +// device's counters (or, with no device, a child `stat`), and a stall that +// starts between two passes is not seen by it. A page or a poll that actually +// reaches the drive is: `onDrive` runs the call against a timer of +// `settings.storage.health.budgetMs` (3 s by default; `healthTimings()`), and a +// call that has not answered by then marks its location `stalled` at once +// (since now) and throws `DriveNotAnsweringError`. The call itself is left to +// settle on its own: its thread is held until the drive answers, which is the +// stated limit. // // THE BUDGET COVERS A WHOLE UNIT OF WORK. A caller sends a unit through as one // call (a video directory's few reads, a page's reads of one video), so a slow @@ -383,7 +453,8 @@ export function notAnsweringText(h: Pick<LocationHealth, "since">, now?: number) // requests completed meanwhile the drive is slow, not stalled, and only this // call is refused. // -// AT MOST DRIVE_CALLS_IN_FLIGHT CALLS PER SLOT KEY ARE IN FLIGHT. A walk of +// AT MOST `inFlightPerLocation` (4 by default) CALLS PER SLOT KEY ARE IN +// FLIGHT. A walk of // `data/` fans out 64 wide, and a stall mid-walk would otherwise put all 64 in // libuv's queue before the watchdog fired. The rest wait in a queue of our own: // - a transition to `stalled` (from any detector) refuses them at once; @@ -397,9 +468,12 @@ export function notAnsweringText(h: Pick<LocationHealth, "since">, now?: number) // began (slow, not stalled — the same test as a timeout's): none of those // calls has returned, whatever the last pass said. // A slot is released when its call really returns, not when the watchdog gave -// up on it. So on one location at most four threads wait on its drive for the -// calls that come through here — every page and poll path, and the snapshot -// walk. A job's own reads that do not come through here are not capped. +// up on it. So on one location at most `inFlightPerLocation` threads wait on +// its drive for the calls that come through here — every page and poll path, +// and the snapshot walk. A job's own reads that do not come through here are +// not capped. A cap lowered while calls are in flight is reached as they +// return (a returning call gives its slot back instead of handing it on); a +// cap raised lets waiting calls take the new slots at once. // // THE SLOT KEY: a configured location's id; for a probe of another root under a // location's id, that root; for a path on no configured location (a root typed @@ -409,16 +483,14 @@ export function notAnsweringText(h: Pick<LocationHealth, "since">, now?: number) // Do not nest `onDrive` for one key: the inner call would wait for a slot the // outer one holds. -export const DRIVE_CALL_BUDGET_MS = 3_000; -export const DRIVE_CALLS_IN_FLIGHT = 4; - -// Test seam: shorten (or restore, with no argument) the watchdog's budget. +// Test seam: shorten (or restore, with no argument) the watchdog's budget, +// past the settings' 500 ms floor. It wins over the applied timings. export function setDriveCallBudget(ms?: number): void { healthState().budgetMs = ms; } function driveCallBudget(): number { - return healthState().budgetMs ?? DRIVE_CALL_BUDGET_MS; + return healthTimings().budgetMs; } export class DriveNotAnsweringError extends Error { @@ -429,7 +501,7 @@ export class DriveNotAnsweringError extends Error { detail ?? (health ? `${NOT_ANSWERING} (location "${health.label}", ${sinceText(health.since)})` - : `${NOT_ANSWERING} (a read did not answer within ${driveCallBudget() / 1000} s)`), + : `${NOT_ANSWERING} (a read did not answer within ${secondsText(driveCallBudget())})`), ); this.name = "DriveNotAnsweringError"; this.health = health; @@ -500,13 +572,14 @@ function refuseWaiters(key: string, health: LocationHealth | null): void { // refused at once when every slot is held by an overdue call. async function acquireSlot(r: Resolved): Promise<void> { const s = healthState(); + const cap = healthTimings().inFlightPerLocation; const n = s.inFlight.get(r.key) ?? 0; - if (n < DRIVE_CALLS_IN_FLIGHT) { + if (n < cap) { s.inFlight.set(r.key, n + 1); return; } const overdue = s.overdue.get(r.key) ?? []; - if (overdue.length >= DRIVE_CALLS_IN_FLIGHT) { + if (overdue.length >= cap) { // EVERY SLOT IS HELD BY A CALL THE WATCHDOG GAVE UP ON: refused at once. // Marked stalled only when the disk has not been completing requests // since the oldest of them began — the watchdog's own slow-or-stalled test @@ -516,14 +589,14 @@ async function acquireSlot(r: Resolved): Promise<void> { if (oldest.before && now && now.completed > oldest.before.completed) { throw new DriveNotAnsweringError( null, - `drive slow (${DRIVE_CALLS_IN_FLIGHT} reads on it are past ` + - `${driveCallBudget() / 1000} s, while its disk is still completing others)`, + `drive slow (${overdue.length} reads on it are past ` + + `${secondsText(driveCallBudget())}, while its disk is still completing others)`, ); } let health: LocationHealth | null = null; if (r.loc) { recordLocationHealth(r.loc, "stalled", { - cause: `${DRIVE_CALLS_IN_FLIGHT} reads on it have not answered`, + cause: `${overdue.length} reads on it have not answered`, }); health = stalledLocation(r.loc); } @@ -559,7 +632,7 @@ async function acquireSlot(r: Resolved): Promise<void> { reject( new DriveNotAnsweringError( r.loc ? stalledLocation(r.loc) : null, - `${NOT_ANSWERING} (nothing on it answered for ${budget / 1000} s while a read waited)`, + `${NOT_ANSWERING} (nothing on it answered for ${secondsText(budget)} while a read waited)`, ), ); }; @@ -595,16 +668,32 @@ async function acquireSlot(r: Resolved): Promise<void> { // A slot is released when its call returns (or, on a refusal after a wait, // without a call). A return is progress: every waiter on the key has its -// deadline restarted, then the slot goes to the first waiter still waiting. +// deadline restarted, then the slot goes to the first waiter still waiting — +// unless the cap was lowered below what is in flight, when it is given back. function releaseSlot(key: string): void { const s = healthState(); const q = s.waiters.get(key); if (q) for (const w of q) w.rearm(); - while (q && q.length > 0) { + const n = s.inFlight.get(key) ?? 1; + if (n <= healthTimings().inFlightPerLocation) { + while (q && q.length > 0) { + const next = q.shift() as Waiter; + if (next.resolve()) return; + } + } + s.inFlight.set(key, Math.max(0, n - 1)); +} + +// A raised cap: the calls already waiting on `key` take the new slots now, +// rather than one at a time as calls return. +function admitWaiters(key: string): void { + const s = healthState(); + const q = s.waiters.get(key); + const cap = healthTimings().inFlightPerLocation; + while (q && q.length > 0 && (s.inFlight.get(key) ?? 0) < cap) { const next = q.shift() as Waiter; - if (next.resolve()) return; + if (next.resolve()) s.inFlight.set(key, (s.inFlight.get(key) ?? 0) + 1); } - s.inFlight.set(key, Math.max(0, (s.inFlight.get(key) ?? 1) - 1)); } // How many calls are in flight on a location through `onDrive` (for tests and @@ -626,8 +715,8 @@ function readDeviceCounters(loc: Resolved["loc"]): BlockStatSample | null { } // Run `call` against the drive `where` is on: refused at once when that -// location is stalled, queued behind DRIVE_CALLS_IN_FLIGHT calls already in -// flight on its key, and raced against the budget. Throws +// location is stalled, queued behind `inFlightPerLocation` calls already in +// flight on its key, and raced against `budgetMs` (`healthTimings()`). Throws // DriveNotAnsweringError for a refusal or a timeout; any other error is the // call's own. export async function onDrive<T>(where: Where, call: () => Promise<T>): Promise<T> { @@ -658,11 +747,13 @@ export async function onDrive<T>(where: Where, call: () => Promise<T>): Promise< throw err; } const TIMED_OUT = Symbol("timed out"); + // The budget this call runs against, read as it starts. + const budget = driveCallBudget(); let timer: ReturnType<typeof setTimeout> | undefined; // Not unref'd: it is cleared the moment the call answers, and while the call // is outstanding the timer is what must fire. const timeout = new Promise<typeof TIMED_OUT>((resolve) => { - timer = setTimeout(() => resolve(TIMED_OUT), driveCallBudget()); + timer = setTimeout(() => resolve(TIMED_OUT), budget); }); let answer: T | typeof TIMED_OUT; try { @@ -693,18 +784,23 @@ export async function onDrive<T>(where: Where, call: () => Promise<T>): Promise< // SLOW, NOT STALLED: the disk completed requests while this call waited. throw new DriveNotAnsweringError( null, - `drive slow (a read did not answer within ${driveCallBudget() / 1000} s, ` + + `drive slow (a read did not answer within ${secondsText(budget)}, ` + `while its disk was still completing others)`, ); } if (r.loc) { recordLocationHealth(r.loc, "stalled", { - cause: `a read in the editor did not answer within ${driveCallBudget() / 1000} s`, + cause: `a read in the editor did not answer within ${secondsText(budget)}`, }); throw new DriveNotAnsweringError(stalledLocation(r.loc)); } refuseWaiters(r.key, null); - throw new DriveNotAnsweringError(null); + // The budget this call ran against, not the one in force now: a save during + // the call must not rewrite what it was given. + throw new DriveNotAnsweringError( + null, + `${NOT_ANSWERING} (a read did not answer within ${secondsText(budget)})`, + ); } // --------------------------------------------------------------------------- diff --git a/common/lib/storageHealthCounters.test.ts b/common/lib/storageHealthCounters.test.ts @@ -6,10 +6,10 @@ import { chmod, mkdir, mkdtemp, rm, writeFile } from "node:fs/promises"; import { blockDeviceName, detectLocationHealth, - MIN_COUNTER_INTERVAL_MS, + minCounterIntervalMs, resetHealthDetector, } from "./storageVolumes"; -import { countersVerdict, parseBlockStat } from "./storageHealth"; +import { applyHealthTimings, countersVerdict, parseBlockStat } from "./storageHealth"; // Run with: // pnpm --filter yt-dlp-transcript-common exec tsx --test lib/storageHealthCounters.test.ts @@ -19,7 +19,10 @@ import { countersVerdict, parseBlockStat } from "./storageHealth"; // the test writes — a stalled device is one whose in-flight count stays up // while its completions stand still. -beforeEach(() => resetHealthDetector()); +beforeEach(() => { + resetHealthDetector(); + applyHealthTimings(); +}); // A real line from this machine's /sys/class/block/<dev>/stat (17 fields). const LINE = @@ -147,8 +150,9 @@ test("counters: a second sample sooner than the minimum interval gives no verdic await detectLocationHealth(loc(h.root), h.bins, { sysBlockDir: h.sys, now: 0 }); const soon = await detectLocationHealth(loc(h.root), h.bins, { sysBlockDir: h.sys, - now: MIN_COUNTER_INTERVAL_MS - 1, + now: minCounterIntervalMs() - 1, }); + assert.equal(minCounterIntervalMs(), 10_000); assert.equal(soon.answer, null); // Compared with the FIRST sample, not the refused one. const later = await detectLocationHealth(loc(h.root), h.bins, { @@ -159,6 +163,31 @@ test("counters: a second sample sooner than the minimum interval gives no verdic }); }); +test("counters: the minimum spacing follows the pass interval (storage.health.passIntervalMs)", async () => { + await withHarness(async (h) => { + await h.control({ source: "/dev/fakedisk1" }); + await h.counters("fakedisk1", 100, 2); + // A pass every 8 s: samples 4 s apart are compared (8 − 5 = 3 s, floored + // at half the interval), which the default 15 s pass would refuse. + applyHealthTimings({ passIntervalMs: 8_000 }); + assert.equal(minCounterIntervalMs(), 4_000); + await detectLocationHealth(loc(h.root), h.bins, { sysBlockDir: h.sys, now: 0 }); + const early = await detectLocationHealth(loc(h.root), h.bins, { + sysBlockDir: h.sys, + now: 3_999, + }); + assert.equal(early.answer, null); + const due = await detectLocationHealth(loc(h.root), h.bins, { + sysBlockDir: h.sys, + now: 4_000, + }); + assert.equal(due.answer, "stalled"); + // A pass every 5 minutes: 10 s, as at the default. + applyHealthTimings({ passIntervalMs: 300_000 }); + assert.equal(minCounterIntervalMs(), 10_000); + }); +}); + test("no device → the child stat, and the verdict says so", async () => { await withHarness(async (h) => { const opts = { sysBlockDir: h.sys, now: 0 }; diff --git a/common/lib/storageHealthProbe.test.ts b/common/lib/storageHealthProbe.test.ts @@ -4,7 +4,8 @@ import path from "node:path"; import { tmpdir } from "node:os"; import { chmod, mkdir, mkdtemp, rm, writeFile } from "node:fs/promises"; import { probeLocationHealth } from "./storageVolumes"; -import { HEALTH_PROBE_TIMEOUT_MS } from "./storageHealth"; +import { HEALTH_TIMING_DEFAULTS } from "./storageHealthTimings"; +import { applyHealthTimings } from "./storageHealth"; // Run with: // pnpm --filter yt-dlp-transcript-common exec tsx --test lib/storageHealthProbe.test.ts @@ -75,7 +76,7 @@ test("a child that never answers is 'stalled' on the timer, without waiting for }); test("the default budget is 3 s", async () => { - assert.equal(HEALTH_PROBE_TIMEOUT_MS, 3_000); + assert.equal(HEALTH_TIMING_DEFAULTS.probeTimeoutMs, 3_000); await withDir(async (dir) => { const statBin = await fakeStat(dir, { sleepMs: 20_000, out: "directory" }); const started = Date.now(); @@ -85,6 +86,21 @@ test("the default budget is 3 s", async () => { }); }); +test("the budget is storage.health.probeTimeoutMs when the caller names none", async () => { + applyHealthTimings({ probeTimeoutMs: 500 }); + try { + await withDir(async (dir) => { + const statBin = await fakeStat(dir, { sleepMs: 20_000, out: "directory" }); + const started = Date.now(); + assert.equal(await probeLocationHealth({ root: dir }, { statBin }), "stalled"); + const took = Date.now() - started; + assert.ok(took >= 500 && took < 2_500, `answered after ${took} ms`); + }); + } finally { + applyHealthTimings(); + } +}); + test("an answer inside the budget is taken as given", async () => { await withDir(async (dir) => { const slowDir = await fakeStat(dir, { sleepMs: 100, out: "directory" }); diff --git a/common/lib/storageHealthTimings.test.ts b/common/lib/storageHealthTimings.test.ts @@ -0,0 +1,138 @@ +// THE DRIVE-HEALTH TIMINGS AS SETTINGS: `settings.storage.health`. +// +// Run with: +// pnpm --filter yt-dlp-transcript-common exec tsx --test lib/storageHealthTimings.test.ts +// +// The claims: absent is every default; a read clamps an out-of-range value into +// its range (the schema's rule: a read never throws); only a value that differs +// from its default is kept, so an untuned file has no `health` key and a save +// writes none; the block survives a save and a read, and the mediaRoot +// migration. Its own file for the SETTINGS SEAM (settingsWrite.test.ts says +// why): SETTINGS_FILE is set before the settings module is imported. + +import { mkdtempSync, readFileSync, writeFileSync } from "node:fs"; +import { rm } from "node:fs/promises"; +import os from "node:os"; +import path from "node:path"; +import { after, test } from "node:test"; +import assert from "node:assert/strict"; +import { + HEALTH_TIMING_BOUNDS, + HEALTH_TIMING_DEFAULTS, + HEALTH_TIMING_KEYS, + clearRuleText, + counterSampleMinimumMs, + resolveHealthTimings, + sanitizeStorageHealth, + secondsText, +} from "./storageHealthTimings"; + +const ROOT = mkdtempSync(path.join(os.tmpdir(), "health-timings-")); +process.env.TRANSCRIPTS_DIR = ROOT; +process.env.SETTINGS_FILE = path.join(ROOT, "settings.json"); + +const { getSettings, writeSettings, siteSettingsSchema, defaultSiteSettings } = + await import("./settings"); + +after(() => rm(ROOT, { recursive: true, force: true })); + +const onDisk = () => + JSON.parse(readFileSync(process.env.SETTINGS_FILE!, "utf8")) as { + storage: Record<string, unknown>; + }; + +test("the defaults are the constants they replace", () => { + assert.deepEqual(HEALTH_TIMING_DEFAULTS, { + budgetMs: 3_000, + passIntervalMs: 15_000, + probeTimeoutMs: 3_000, + clearAfterCleanPasses: 2, + inFlightPerLocation: 4, + }); + assert.deepEqual([...HEALTH_TIMING_KEYS].sort(), Object.keys(HEALTH_TIMING_DEFAULTS).sort()); + for (const key of HEALTH_TIMING_KEYS) { + const { min, max } = HEALTH_TIMING_BOUNDS[key]; + const d = HEALTH_TIMING_DEFAULTS[key]; + assert.ok(min <= d && d <= max, `${key}'s default is in its range`); + } +}); + +test("absent is every default: no block, an empty block, junk", () => { + for (const raw of [undefined, null, {}, [], "3000", 7]) { + assert.deepEqual(sanitizeStorageHealth(raw), {}); + assert.deepEqual(resolveHealthTimings(raw as never), HEALTH_TIMING_DEFAULTS); + } + assert.equal(getSettings().storage.health, undefined, "no file: no block"); + assert.equal(defaultSiteSettings().storage.health, undefined); +}); + +test("a read clamps into the range, rounds, drops non-numbers and keeps only what differs from the default", () => { + assert.deepEqual( + sanitizeStorageHealth({ + budgetMs: 200, // below 500 + passIntervalMs: 900_000, // above 300000 + probeTimeoutMs: "4000", // a string is not a number + clearAfterCleanPasses: 2, // the default + inFlightPerLocation: 7.6, + unknown: 1, + }), + { budgetMs: 500, passIntervalMs: 300_000, inFlightPerLocation: 8 }, + ); + // A value that clamps ONTO its default is the default, and is not kept. + assert.deepEqual(sanitizeStorageHealth({ clearAfterCleanPasses: 2.2 }), {}); + assert.deepEqual(resolveHealthTimings({ budgetMs: 200 }), { + ...HEALTH_TIMING_DEFAULTS, + budgetMs: 500, + }); +}); + +test("a settings.json round trip: a tuned value is written, read back, and a default is not written", async () => { + const base = defaultSiteSettings(); + await writeSettings({ + ...base, + storage: { ...base.storage, health: { budgetMs: 4_000, clearAfterCleanPasses: 2 } }, + }); + assert.deepEqual(onDisk().storage.health, { budgetMs: 4_000 }); + assert.deepEqual(getSettings().storage.health, { budgetMs: 4_000 }); + // Every value back to its default: the key is not written at all. + await writeSettings({ ...base, storage: { ...base.storage, health: { budgetMs: 3_000 } } }); + assert.equal("health" in onDisk().storage, false); + // A hand-edited file out of range reads clamped. + writeFileSync( + process.env.SETTINGS_FILE!, + JSON.stringify({ storage: { locations: [], health: { inFlightPerLocation: 99 } } }), + ); + assert.deepEqual(getSettings().storage.health, { inFlightPerLocation: 8 }); +}); + +test("the timings survive the mediaRoot migration (a block with no locations)", () => { + const parsed = siteSettingsSchema.parse({ + storage: { health: { passIntervalMs: 30_000 } }, + }); + assert.deepEqual(parsed.storage.health, { passIntervalMs: 30_000 }); + writeFileSync( + process.env.SETTINGS_FILE!, + JSON.stringify({ storage: { mediaRoot: "/mnt/cold", health: { budgetMs: 6_000 } } }), + ); + const s = getSettings(); + assert.deepEqual(s.storage.locations.map((l) => l.root), ["/mnt/cold"]); + assert.deepEqual(s.storage.health, { budgetMs: 6_000 }); +}); + +test("the counters' sample spacing: min(10 s, interval − 5 s), floored at half the interval", () => { + assert.equal(counterSampleMinimumMs(15_000), 10_000, "the default, as it was"); + assert.equal(counterSampleMinimumMs(300_000), 10_000); + assert.equal(counterSampleMinimumMs(12_000), 7_000); + assert.equal(counterSampleMinimumMs(10_000), 5_000); + assert.equal(counterSampleMinimumMs(8_000), 4_000, "the floor: 3 s would be under half"); + assert.equal(counterSampleMinimumMs(5_000), 2_500, "never 0 at the minimum interval"); +}); + +test("the words", () => { + assert.equal(secondsText(3_000), "3 s"); + assert.equal(secondsText(2_500), "2.5 s"); + assert.equal(secondsText(80), "0.08 s"); + assert.equal(clearRuleText(1), "once"); + assert.equal(clearRuleText(2), "twice in a row"); + assert.equal(clearRuleText(5), "5 times in a row"); +}); diff --git a/common/lib/storageHealthTimings.ts b/common/lib/storageHealthTimings.ts @@ -0,0 +1,166 @@ +import type { FieldDocs } from "./fieldDocs"; + +// THE DRIVE-HEALTH TIMINGS — `settings.storage.health` — and their defaults. +// +// The health gate (lib/storageHealth.ts) decides that a drive is not answering +// from a handful of numbers: how long one read may take, how often the health +// pass looks at the drive, how long that look may take, how many clean looks in +// a row clear a stall, and how many reads may wait on one drive at once. They +// were constants. Under heavy external-disk churn a stall can be misjudged, and +// the ruling (release 15, slice DT) is that the operator then tunes the numbers +// on /storage rather than the code. Today's constants are the defaults. +// +// THE STORED BLOCK HOLDS ONLY WHAT DIFFERS FROM A DEFAULT. Every key is +// optional, an absent key is its default, and `sanitizeStorageHealth` drops a +// value equal to its default — so a settings.json that never tuned anything has +// no `health` key at all, and a default changed in a later release reaches it. +// +// READS CLAMP, THE FORM REFUSES. A hand-edited value outside its range is +// clamped into it on read (the schema's rule: a read never throws); the /storage +// form refuses one with a sentence instead, because a save that quietly stored +// another number than the one typed reads as a form that did not listen. +// +// PURE, NO IMPORTS BUT A TYPE: the /storage form (a "use client" file) imports +// the defaults, the ranges and the words from here. + +export type StorageHealthSettings = { + budgetMs?: number; + passIntervalMs?: number; + probeTimeoutMs?: number; + clearAfterCleanPasses?: number; + inFlightPerLocation?: number; +}; + +// Every timing, resolved: what the health module runs on. +export type HealthTimings = Required<StorageHealthSettings>; + +export type HealthTimingKey = keyof HealthTimings; + +// Today's constants (release 15 slice DS), now the defaults. +export const HEALTH_TIMING_DEFAULTS: Readonly<HealthTimings> = Object.freeze({ + budgetMs: 3_000, + passIntervalMs: 15_000, + probeTimeoutMs: 3_000, + clearAfterCleanPasses: 2, + inFlightPerLocation: 4, +}); + +export const HEALTH_TIMING_BOUNDS: Readonly< + Record<HealthTimingKey, { min: number; max: number }> +> = Object.freeze({ + budgetMs: { min: 500, max: 60_000 }, + passIntervalMs: { min: 5_000, max: 300_000 }, + probeTimeoutMs: { min: 500, max: 30_000 }, + clearAfterCleanPasses: { min: 1, max: 10 }, + // At most HALF the editor's 16 file-access threads (`UV_THREADPOOL_SIZE` in + // `start` and the container): one drive that stops answering holds this + // many of them, and at 16 it would hold every one — the outage the cap + // exists to bound. + inFlightPerLocation: { min: 1, max: 8 }, +}); + +// In the order the form shows them. +export const HEALTH_TIMING_KEYS: readonly HealthTimingKey[] = [ + "budgetMs", + "passIntervalMs", + "probeTimeoutMs", + "clearAfterCleanPasses", + "inFlightPerLocation", +]; + +// A whole number in range, or undefined (not a finite number). +export function clampHealthTiming(key: HealthTimingKey, value: unknown): number | undefined { + if (typeof value !== "number" || !Number.isFinite(value)) return undefined; + const { min, max } = HEALTH_TIMING_BOUNDS[key]; + return Math.min(max, Math.max(min, Math.round(value))); +} + +// Coerce a raw `settings.storage.health` into the stored block: each value a +// number clamped into its range, and only when it differs from its default. +// Anything else — a string, a key the block does not name — is dropped. +export function sanitizeStorageHealth(value: unknown): StorageHealthSettings { + if (!value || typeof value !== "object" || Array.isArray(value)) return {}; + const r = value as Record<string, unknown>; + const out: StorageHealthSettings = {}; + for (const key of HEALTH_TIMING_KEYS) { + const n = clampHealthTiming(key, r[key]); + if (n !== undefined && n !== HEALTH_TIMING_DEFAULTS[key]) out[key] = n; + } + return out; +} + +// The stored block with every absent key filled from the defaults. +export function resolveHealthTimings(stored?: StorageHealthSettings): HealthTimings { + const clean = sanitizeStorageHealth(stored); + return { ...HEALTH_TIMING_DEFAULTS, ...clean }; +} + +// THE COUNTERS' SAMPLE SPACING, derived from the pass interval. Two samples of a +// disk's request counters closer than this are not compared: a healthy drive can +// have a request in flight at two instants a moment apart without completing +// one (release 15 DS, "kept at review"). The ruling: min(10 s, interval − 5 s), +// so consecutive passes always compare — which is 10 s at the default 15 s, as +// it was. Below a 10 s interval that formula falls toward nothing (0 at the 5 s +// minimum, where a /storage Refresh just after a pass would compare two samples +// taken milliseconds apart), so it is floored at half the interval. +export function counterSampleMinimumMs(passIntervalMs: number): number { + return Math.min(10_000, Math.max(passIntervalMs - 5_000, Math.round(passIntervalMs / 2))); +} + +// "3 s", "2.5 s", "0.08 s": a millisecond figure in the seconds every surface +// uses. +export function secondsText(ms: number): string { + return `${ms / 1000} s`; +} + +// How many clean answers clear a stall, in words: "once", "twice in a row", +// "3 times in a row". +export function clearRuleText(n: number): string { + return n === 1 ? "once" : n === 2 ? "twice in a row" : `${n} times in a row`; +} + +// What each field means, in the operator's words: the /storage form's hint line. +export const HEALTH_TIMING_HINTS: Readonly<Record<HealthTimingKey, string>> = Object.freeze({ + budgetMs: + "How long one read may take before the drive counts as not answering. Takes effect on the next read.", + passIntervalMs: + "How often each drive is checked, from its disk's own request counters. A save re-arms the check at once.", + probeTimeoutMs: + "How long one check may wait for the drive's root where no disk can be named (and for naming the disk).", + clearAfterCleanPasses: + "How many clean checks in a row it takes before a drive marked not answering is used again.", + inFlightPerLocation: + "How many reads may be on one drive at once; the rest wait their turn, and are refused if it stops answering. At most 8, half the editor's 16 file-access threads: a drive that stops answering holds this many of them.", +}); + +// SETTINGS.md's `storage.health` table. +export const STORAGE_HEALTH_SETTINGS_FIELD_DOCS: FieldDocs<StorageHealthSettings> = { + budgetMs: + "How long one read may take before the drive counts as not answering, in ms (default 3000, " + + "500–60000). The watchdog's budget (`onDrive`, lib/storageHealth.ts) for one unit of work — a " + + "video directory's reads, a page's reads of one video: a unit that has not answered by then is " + + "refused, and marks its location not answering unless the disk's request counters show it still " + + "completing others (slow, not stalled). A read waiting for a slot is refused when nothing on the " + + "drive has returned for this long plus a quarter of it (at most 250 ms). Takes effect on the next " + + "read after a save.", + passIntervalMs: + "How often the health pass reads each location's disk counters, in ms (default 15000, " + + "5000–300000). A stall that starts between two passes is seen by the next, or at once by a page's " + + "read. A save on /storage re-arms the pass's timer at once; a hand edit, at the next pass. Two " + + "counter samples are compared only when at least min(10 s, this − 5 s) apart, a spacing never " + + "less than half of this.", + probeTimeoutMs: + "How long the health pass waits, in ms (default 3000, 500–30000), for the child `stat` of a root " + + "where no disk can be named (a timeout counts as not answering) and for the `findmnt` that names a " + + "root's disk (a timeout names none that pass). Takes effect on the next pass.", + clearAfterCleanPasses: + "How many clean answers in a row clear a location marked not answering (default 2, 1–10). Each " + + "health pass is one answer, and so is a Refresh on /storage; a miss in between starts the count " + + "again. Takes effect on the next answer.", + inFlightPerLocation: + "How many reads through the watchdog may be on one location's drive at once (default 4, 1–8); the " + + "rest wait in the editor's own queue, so a stall mid-walk holds this many of Node's threads, not " + + "all of them. At most 8, half of `UV_THREADPOOL_SIZE` (16 in the editor's start script and the " + + "container), so one drive that stops answering cannot hold every thread. Takes effect on the next " + + "read.", +}; diff --git a/common/lib/storageLocations.ts b/common/lib/storageLocations.ts @@ -1,5 +1,6 @@ import path from "node:path"; import type { FieldDocs } from "./fieldDocs"; +import type { StorageHealthSettings } from "./storageHealthTimings"; // STORAGE LOCATIONS — the named places a channel's media may live. // @@ -95,6 +96,7 @@ export type StorageSettings = { locations: StorageLocation[]; defaultLocationId: string; savedVideosLocationId?: string; + health?: StorageHealthSettings; }; export const STORAGE_SETTINGS_FIELD_DOCS: FieldDocs<StorageSettings> = { @@ -112,6 +114,14 @@ export const STORAGE_SETTINGS_FIELD_DOCS: FieldDocs<StorageSettings> = { "Optional so an older settings.json parses (and an older binary that " + "drops it leaves a store that still works, because the symlink is what " + "every reader follows).", + health: + "THE DRIVE-HEALTH TIMINGS: how long a read may take before a drive " + + "counts as not answering, how often the health pass looks, how long " + + "its look may take, how many clean looks clear a stall, and how many " + + "reads may be on one drive at once. Edited on /storage (Drive health " + + "timing). Absent = every default, and only a value that differs from " + + "its default is written, so an untuned install follows a default " + + "changed later. See `storage.health` below.", }; // Strip trailing slashes so "/mnt/platter/" and "/mnt/platter" are one root. @@ -214,9 +224,12 @@ export function migrateMediaRootToLocations(raw: unknown): unknown { if (!raw || typeof raw !== "object" || Array.isArray(raw)) return raw; const r = raw as Record<string, unknown>; if (r.locations !== undefined) return raw; + // The drive-health timings are not about the locations: a block that spells + // them and no `locations` (a hand edit) keeps them through the migration. + const health = r.health !== undefined ? { health: r.health } : {}; const mediaRoot = typeof r.mediaRoot === "string" ? r.mediaRoot.trim() : ""; if (mediaRoot === "" || !path.isAbsolute(mediaRoot)) { - return { locations: [], defaultLocationId: "" }; + return { locations: [], defaultLocationId: "", ...health }; } return { locations: [ @@ -228,5 +241,6 @@ export function migrateMediaRootToLocations(raw: unknown): unknown { }, ], defaultLocationId: "default", + ...health, }; } diff --git a/common/lib/storageVolumes.ts b/common/lib/storageVolumes.ts @@ -6,8 +6,8 @@ import type { Paths } from "./paths"; import { getFreeBytes } from "./diskSpace"; import type { StorageLocation, StorageVolume } from "./storageLocations"; import { - HEALTH_PROBE_TIMEOUT_MS, countersVerdict, + healthTimings, isDriveNotAnswering, onDrive, parseBlockStat, @@ -17,6 +17,7 @@ import { type HealthDetector, type LocationHealthState, } from "./storageHealth"; +import { counterSampleMinimumMs, secondsText } from "./storageHealthTimings"; // STORAGE VOLUME PROBES — is this location's disk here, and if not, where? // @@ -247,7 +248,8 @@ export async function probeLocation( // run in-process, and on a stalled disk each holds an I/O thread until the // drive comes back; the health pass (below) already asked without touching // it. Both go through `onDrive`'s watchdog, too: one that has not answered - // in 3 s marks the location stalled and this probe answers so. + // within the budget (`storage.health.budgetMs`, 3 s by default) marks the + // location stalled and this probe answers so. const stalledProbe: StorageLocationProbe = { status: "stalled", identity: { known: false }, @@ -389,7 +391,7 @@ export async function probeLocationHealth( ): Promise<LocationHealthState> { const root = loc.root.trim(); if (root === "") return "absent"; - const timeoutMs = opts.timeoutMs ?? HEALTH_PROBE_TIMEOUT_MS; + const timeoutMs = opts.timeoutMs ?? healthTimings().probeTimeoutMs; let child: ReturnType<typeof execa>; try { child = execa(opts.statBin ?? "stat", ["-L", "-c", "%F", "--", root], { @@ -440,8 +442,9 @@ export async function probeLocationHealth( // reading them touches only /sys. So the health pass asks this first: // // 1. the root's device: `findmnt -J -T <root> -o SOURCE,UUID` as a child -// raced against 3 s (a findmnt stuck resolving the root holds nothing of -// ours), the `[subvolume]` suffix a bind or btrfs mount adds taken off, +// raced against `probeTimeoutMs` (3 s by default; a findmnt stuck +// resolving the root holds nothing of ours), the `[subvolume]` suffix a +// bind or btrfs mount adds taken off, // `/dev/mapper/<x>` resolved to its `dm-N`, then the basename. A partition // and a mapper device both have `/sys/class/block/<name>/stat`. Asked only // when the root has no device yet or its device's /sys entry cannot be read @@ -450,7 +453,7 @@ export async function probeLocationHealth( // none this pass; one whose UUID is not the location's recorded one names // none (the root is then a directory on some other filesystem). // 2. that device's `stat` line, compared with the previous pass's sample for -// the location (same device, at least MIN_COUNTER_INTERVAL_MS earlier). +// the location (same device, at least `minCounterIntervalMs()` earlier). // // NO DEVICE (a container, no findmnt, a network or tmpfs mount, no /sys entry) // FALLS BACK TO THE CHILD `stat`, and the verdict says which detector answered. @@ -464,8 +467,11 @@ export type DetectorOptions = HealthProbeOptions & { export const SYS_BLOCK_DIR = "/sys/class/block"; // Two samples closer than this are not compared: a healthy drive can have a // request in flight at two instants a moment apart without completing one. -// The pass is 15 s apart; a /storage Refresh just after a pass gives no verdict. -export const MIN_COUNTER_INTERVAL_MS = 10_000; +// Derived from the pass interval (`counterSampleMinimumMs`): 10 s at the +// default 15 s, so a /storage Refresh just after a pass gives no verdict. +export function minCounterIntervalMs(): number { + return counterSampleMinimumMs(healthTimings().passIntervalMs); +} type CounterSample = BlockStatSample & { device: string; at: number }; @@ -604,7 +610,7 @@ export async function detectLocationHealth( opts: DetectorOptions = {}, ): Promise<HealthVerdict> { const now = opts.now ?? Date.now(); - const timeoutMs = opts.timeoutMs ?? HEALTH_PROBE_TIMEOUT_MS; + const timeoutMs = opts.timeoutMs ?? healthTimings().probeTimeoutMs; const sysBlockDir = opts.sysBlockDir ?? SYS_BLOCK_DIR; const detector = detectorState(); // The device this root was last known on, if its /sys entry still reads; @@ -623,7 +629,7 @@ export async function detectLocationHealth( if (sample) { detector.deviceByRoot.set(loc.root, device); const prev = detector.samples.get(loc.id); - if (prev && prev.device === device && now - prev.at < MIN_COUNTER_INTERVAL_MS) { + if (prev && prev.device === device && now - prev.at < minCounterIntervalMs()) { return { answer: null, detector: "counters", device }; } detector.samples.set(loc.id, { ...sample, device, at: now }); @@ -651,7 +657,7 @@ export async function detectLocationHealth( answer, detector: "stat", ...(answer === "stalled" - ? { cause: `a stat of its root did not answer within ${timeoutMs / 1000} s` } + ? { cause: `a stat of its root did not answer within ${secondsText(timeoutMs)}` } : {}), }; } diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -15,6 +15,7 @@ - **The hub URL hints say what the setting does now.** Settings' **Family hub URL** and a site's **Hub URL** no longer promise a Hub link in the header (it was removed): the value is published as `hubUrl` in each site's `/site.json` and `/corpus.json`, so the hub can tell its member sites. `SETTINGS.md` and `SITE.md` say the same. - **A site can be left off the homepage and the hub.** A site's settings have a new checkbox, **List on the Archilyzer homepage and hub**, on by default (`listed` in `site.json`; only `false` is written). Turned off, the site still builds and deploys at its own URL as before, but the homepage has no card, chart series, `/stats` entry or recent item for it; the hub does not list it as a member, search it, or name it in its `corpus.json` and `llms.txt`; no other site's footer links it; and `channel-sites.json` and the homepage's `stats/` leave it out. A channel only unlisted sites carry is in none of the published totals, the homepage's headline numbers included; a channel a listed site also carries is counted under the listed site. The editor's own pages still show every site. It takes effect at the next homepage, hub and site builds. - **The sidebar's site picker shows your site from the first paint.** It used to show "All sites" on every page and then jump to the site you had picked, and Dashboard and Channels came up in your site only after a `?site=` had been added to the address. The picked site is now kept in a cookie that the editor reads before it draws a page, so the picker, Dashboard and Channels open in it at once, and the address is left alone. A link that carries `?site=<id>` still opens that page in that site, without changing the one you picked; picking a site on such a page drops the `?site=` from the address. On a site's own pages (Charts, Publish, …) the picker still follows the page, and opening one still makes that site the picked one. The first time you open the editor after updating, a site picked before is moved into the cookie; the picker may show "All sites" for a moment that once. A site picked in one tab reaches the editor's other open tabs without a reload. **New channel** starts with the picked site ticked under its sites (or the site of a `?site=` link), including when it is opened from the editor's own links. Each editor keeps its own pick, as before, when several run on one machine on different ports. +- **The drive check's timings can be changed on `/storage`.** The numbers the editor decides a drive is "not answering" by were fixed: a read may take 3 seconds, each drive is checked every 15 seconds (a check that asks the drive from a separate process waits up to 3 seconds), two clean checks in a row put a drive back in use, and at most four reads wait on one drive at a time. They are now **Drive health timing**, a collapsed block at the foot of `/storage`, with those numbers as the defaults — for when a drive that is busy but working is marked not answering, or a stalled one is not. An empty field is its default; a number outside a field's range is refused, with the range. A save takes effect at once: the next read, the next check, and a new check interval re-times the checks. `settings.json` keeps only the values that differ from a default, under `storage.health` (see `SETTINGS.md`), so an editor that never changes them follows the defaults; `archilyzer index` and the stats build read them too. The messages that said "3 s", "every 15 s" or "twice in a row" now say the numbers in force. - **Building the homepage publishes the source's history too: every commit of main with its diff, at `/source/git/`, when stagit is installed.** `archilyzer build homepage`, the `/sites` Homepage jobs and `pnpm ops build-homepage` render the scrubbed mirror with stagit into a log, a page per commit, the refs and two Atom feeds; each page carries the homepage's grounds and one line back to `/source/`, its Files page links into the raw tree, and every page goes through the same gate as the mirror (a denied literal in one refuses the publish). stagit is installed once, outside the repository: `git clone git://git.codemadness.org/stagit && make -C stagit && cp stagit/stagit ~/.local/bin/` (or point `STAGIT_BIN` at it). Without it the build goes on without the history and says so in one line; a render that fails, or pages that would pass the host's limits, are a warning, never a failed build. `archilyzer doctor` shows where stagit is, beside git-filter-repo. The newest 10,000 commits have pages; past that the log says how many more there are. The first build after this change publishes the source again (the step's version is 4). A render cache of about 140 MB is kept in `~/.cache/archilyzer/source-history` (or under `XDG_CACHE_HOME`), so later builds render only the new commits; `archilyzer doctor` shows where it is and its size. ## [0.10.0] - 2026-09-28 diff --git a/editor/app/channels/[slug]/components/MediaNotAnswering.tsx b/editor/app/channels/[slug]/components/MediaNotAnswering.tsx @@ -1,8 +1,10 @@ import Link from "next/link"; import { + healthTimings, notAnsweringText, type LocationHealth, } from "yt-dlp-transcript-common/lib/storageHealth"; +import { secondsText } from "yt-dlp-transcript-common/lib/storageHealthTimings"; // WHAT A PAGE THAT READS A CHANNEL'S `data/` SHOWS WHILE ITS DRIVE IS NOT // ANSWERING, instead of reading it. @@ -26,11 +28,13 @@ export function MediaNotAnswering({ }: { slug: string; // The stalled location, or null when the drive is on no location the health - // state knows and a read of it did not answer within 3 s. + // state knows and a read of it did not answer within the budget + // (`storage.health.budgetMs`). stall: LocationHealth | null; // What would have been shown: "The video list", "This video". what: string; }) { + const timings = healthTimings(); return ( <div className="flex flex-col gap-3"> <p @@ -42,13 +46,14 @@ export function MediaNotAnswering({ <> {what} reads this channel&apos;s media, which is on &ldquo; {stall.label}&rdquo; — a drive that is {notAnsweringText(stall)}. - Nothing is read from it until it answers again (it is checked every - 15 s). + Nothing is read from it until it answers again (it is checked every{" "} + {secondsText(timings.passIntervalMs)}). </> ) : ( <> {what} reads this channel&apos;s media, and a read of its drive did - not answer within 3 s. Nothing more is read from it on this page. + not answer within {secondsText(timings.budgetMs)}. Nothing more is + read from it on this page. </> )} </p> diff --git a/editor/app/channels/[slug]/page.tsx b/editor/app/channels/[slug]/page.tsx @@ -618,7 +618,7 @@ export default async function ChannelDetailPage({ const marker = media.marker ?? null; // A statfs of the channel's drive goes through the watchdog // (lib/storageHealth.ts): refused on a stalled location, given up on - // after 3 s, and read "—" either way. + // after the budget (3 s by default), and read "—" either way. let freeBytes: number | null; try { freeBytes = diff --git a/editor/app/channels/[slug]/videos/[id]/page.tsx b/editor/app/channels/[slug]/videos/[id]/page.tsx @@ -87,7 +87,8 @@ export async function generateMetadata({ }): Promise<Metadata> { const { slug, id } = await params; // The title is read off the drive; a drive that is not answering is not - // asked, and one that does not answer in 3 s is given up on. + // asked, and one that does not answer within the budget + // (`storage.health.budgetMs`, 3 s by default) is given up on. const drive = (await readChannelConfig(getPaths(), slug))?.dataDir?.trim(); let subject = id; try { @@ -117,7 +118,8 @@ export default async function VideoDetailPage({ } // EVERY READ BELOW IS OF THIS VIDEO'S DIRECTORY, on the channel's drive when // it is relocated, so they go through the watchdog as one unit: a drive that - // has not answered them in 3 s is marked stalled and the page says so. + // has not answered them within the budget (3 s by default) is marked stalled + // and the page says so. const loadAll = async () => { const dirData = await loadVideoDir(slug, id); const meta = await loadMeta(slug, id); diff --git a/editor/app/channels/[slug]/videos/page.tsx b/editor/app/channels/[slug]/videos/page.tsx @@ -145,8 +145,9 @@ export default async function ChannelVideosPage({ const channelDataDir = path.join(paths.channelsDir, slug, "data"); // THE READS OF THE DRIVE go through the watchdog when the channel is - // relocated: a drive that has not answered them in 3 s is marked stalled and - // the page says so instead (see MediaNotAnswering). + // relocated: a drive that has not answered them within the budget + // (`storage.health.budgetMs`, 3 s by default) is marked stalled and the page + // says so instead (see MediaNotAnswering). const drive = config.dataDir?.trim(); const onMedia = <T,>(call: () => Promise<T>): Promise<T> => drive ? onDrive(drive, call) : call(); diff --git a/editor/app/channels/components/ChannelVolumeBar.tsx b/editor/app/channels/components/ChannelVolumeBar.tsx @@ -44,6 +44,9 @@ export type ChannelVolume = { // "not answering since 11:35", while the health probe finds the drive not // answering (its free space is then not asked either). notAnswering?: string; + // With `notAnswering`: how many clean checks clear it, in words ("twice in a + // row"; settings.storage.health.clearAfterCleanPasses). + clears?: string; }; export function ChannelVolumeBar({ @@ -105,7 +108,7 @@ export function ChannelVolumeBar({ } title={ v.notAnswering - ? `${v.label}: the drive is ${v.notAnswering}. Pages and polls do not touch it until it answers twice in a row; its channels read "not answering".` + ? `${v.label}: the drive is ${v.notAnswering}. Pages and polls do not touch it until it answers ${v.clears ?? "again"}; its channels read "not answering".` : v.freeBytes === undefined ? `${v.label}: the root is not there — an unmounted drive reports no free space rather than its parent's.` : `${v.label}: ${v.channels} channel(s) hold ${formatBytes(v.bytes)}${ diff --git a/editor/app/channels/page.tsx b/editor/app/channels/page.tsx @@ -8,9 +8,11 @@ import { import { getPaths } from "yt-dlp-transcript-common/lib/paths"; import { inspectChannelMedia } from "yt-dlp-transcript-common/lib/channelMedia"; import { + healthTimings, notAnsweringText, stalledLocation, } from "yt-dlp-transcript-common/lib/storageHealth"; +import { clearRuleText } from "yt-dlp-transcript-common/lib/storageHealthTimings"; import { getSite, listSiteIds, @@ -226,8 +228,9 @@ export default async function ChannelsPage({ // 71 rows on every auto-refresh and the rule is that tables never shell out. const freeByVolume = await volumeFreeBytes({ paths, locations }); // A DRIVE THAT IS NOT ANSWERING, said on its chip. From memory — the health - // pass (the block device's counters every 15 s) or the watchdog on a read is - // what found it (lib/storageHealth.ts); the table asks nothing. + // pass (the block device's counters, every `storage.health.passIntervalMs`) + // or the watchdog on a read is what found it (lib/storageHealth.ts); the + // table asks nothing. const notAnsweringByVolume: Record<string, string | undefined> = {}; for (const loc of locations) { const stall = stalledLocation(loc); @@ -306,7 +309,10 @@ export default async function ChannelsPage({ unmeasured: rows.length - measured.length, freeBytes: freeByVolume[id], ...(notAnsweringByVolume[id] - ? { notAnswering: notAnsweringByVolume[id] } + ? { + notAnswering: notAnsweringByVolume[id], + clears: clearRuleText(healthTimings().clearAfterCleanPasses), + } : {}), }; }) diff --git a/editor/app/settings/saveSettings.test.ts b/editor/app/settings/saveSettings.test.ts @@ -62,6 +62,28 @@ test("an object nested inside a block replaces, it is not merged", () => { assert.deepEqual(out.channelPriority.channels, { other: { tier: "paused" } }); }); +// A LOCATION WRITE KEEPS THE DRIVE-HEALTH TIMINGS. The /storage location +// actions patch `storage` with `{ locations, defaultLocationId }` only; the +// one-level merge is what keeps `storage.health` (and `savedVideosLocationId`, +// which an add once erased before slice 4a). +test("a storage patch of the locations keeps storage.health", () => { + const base = defaultSiteSettings(); + base.storage = { + ...base.storage, + savedVideosLocationId: "cold", + health: { budgetMs: 4_000, inFlightPerLocation: 2 }, + }; + const out = mergeSettingsPatch(base, { + storage: { + locations: [{ id: "cold", label: "Cold", root: "/mnt/cold", autoRepoint: false }], + defaultLocationId: "cold", + }, + }); + assert.deepEqual(out.storage.health, { budgetMs: 4_000, inFlightPerLocation: 2 }); + assert.equal(out.storage.savedVideosLocationId, "cold"); + assert.equal(out.storage.locations.length, 1); +}); + test("saveSettings writes the merged result and touches nothing else", async () => { writeFileSync( process.env.SETTINGS_FILE!, diff --git a/editor/app/storage/actions.ts b/editor/app/storage/actions.ts @@ -36,11 +36,19 @@ import { savedVideosStoreBusyReason } from "./lib/storeBusy"; import { refreshLocationHealth } from "yt-dlp-transcript-common/controller/storageWatch"; import { forgetChannelMedia } from "yt-dlp-transcript-common/lib/channelMedia"; import { + applyHealthTimings, + healthTimings, locationHealth, notAnsweringText, } from "yt-dlp-transcript-common/lib/storageHealth"; +import { + clearRuleText, + secondsText, +} from "yt-dlp-transcript-common/lib/storageHealthTimings"; +import { parseHealthTimingsForm } from "./lib/healthTimingsForm"; -// THE SIX THINGS AN OPERATOR MAY DO TO A STORAGE LOCATION. +// THE SIX THINGS AN OPERATOR MAY DO TO A STORAGE LOCATION (and, at the end, +// the drive-health timings every location is judged by). // // Five of them are small settings writes or one subprocess and run INLINE: // their whole output is a sentence, and a queued job with a log would be a @@ -251,21 +259,24 @@ export async function refreshStorageLocationAction( const location = settings.storage.locations.find((l) => l.id === id); if (!location) return { ok: false, error: `There is no storage location "${id}".` }; // IS IT ANSWERING, asked first and without touching the drive (the block - // device's counters in /sys, or a child `stat` raced against 3 s where no - // device can be named): the probe below runs in-process, and on a stalled - // drive it is not asked at all. The counters give no answer within 10 s of - // the pass's last sample. The channels' remembered answers go too — the - // operator has just done something about the drive. + // device's counters in /sys, or a child `stat` raced against + // `storage.health.probeTimeoutMs` where no device can be named): the probe + // below runs in-process, and on a stalled drive it is not asked at all. The + // counters give no answer within 10 s (by default) of the pass's last + // sample. The channels' remembered answers go too — the operator has just + // done something about the drive. await refreshLocationHealth(location); forgetChannelMedia(); const health = locationHealth(location.id); if (health?.state === "stalled") { revalidateStorage(); + const timings = healthTimings(); return { ok: true, note: `${location.label}: ${notAnsweringText(health)} — ${health.cause ?? "its root did not answer"}. ` + - `Pages skip this drive until it answers twice in a row (checked every 15 s).`, + `Pages skip this drive until it answers ${clearRuleText(timings.clearAfterCleanPasses)} ` + + `(checked every ${secondsText(timings.passIntervalMs)}).`, }; } const probe = await probeLocationMemo(location, paths, { refresh: true }); @@ -475,3 +486,45 @@ export async function evictClipWindowsAction(opts: { ...(opts.dryRun ? { dryRun: true } : {}), }); } + +// --------------------------------------------------------------------------- +// The drive-health timings +// --------------------------------------------------------------------------- + +export type HealthTimingsResult = { ok: true; note: string } | { ok: false; error: string }; + +// SAVE `settings.storage.health` FROM THE /storage FORM (HealthTimingForm). +// +// Parsed by `parseHealthTimingsForm`: an empty field is the default and is not +// written, and a value out of range is REFUSED with a sentence (the schema +// would clamp it; a save that stored another number than the one typed would +// read as a form that did not listen). Then written through the one settings +// writer, as a patch of the storage block, and APPLIED AT ONCE to this +// process's health state (`applyHealthTimings`, on globalThis): the next read +// runs on the new budget and cap, the next answer on the new clear count, the +// next pass on the new probe timeout, and a new pass interval re-arms the +// pass's timer now. The pass would apply them anyway, at its next run. +export async function saveHealthTimingsAction( + _prev: HealthTimingsResult | undefined, + formData: FormData, +): Promise<HealthTimingsResult> { + const parsed = parseHealthTimingsForm(formData); + if (!parsed.ok) return parsed; + const settings = getSettings(); + try { + await saveSettings({ storage: { ...settings.storage, health: parsed.health } }); + } catch (e) { + return { ok: false, error: (e as Error).message }; + } + const t = applyHealthTimings(getSettings().storage.health); + revalidatePath("/storage"); + return { + ok: true, + note: + `Saved. A read may take ${secondsText(t.budgetMs)}; drives are checked every ` + + `${secondsText(t.passIntervalMs)} (a check waits up to ${secondsText(t.probeTimeoutMs)}); ` + + `a drive marked not answering is used again after ` + + `${t.clearAfterCleanPasses === 1 ? "one clean check" : `${t.clearAfterCleanPasses} clean checks in a row`}; ` + + `${t.inFlightPerLocation} read(s) at once per drive.`, + }; +} diff --git a/editor/app/storage/buildStorage.ts b/editor/app/storage/buildStorage.ts @@ -4,11 +4,13 @@ import { getRegistry } from "yt-dlp-transcript-common/jobs/registry"; import { getFreeBytes } from "yt-dlp-transcript-common/lib/diskSpace"; import { udisksctlAvailable } from "yt-dlp-transcript-common/lib/storageVolumes"; import { + healthTimings, isDriveNotAnswering, notAnsweringText, onDrive, stalledLocation, } from "yt-dlp-transcript-common/lib/storageHealth"; +import { clearRuleText } from "yt-dlp-transcript-common/lib/storageHealthTimings"; import { listChannelBriefs } from "yt-dlp-transcript-common/controller/channels"; import { channelsOnLocation, @@ -88,9 +90,10 @@ export async function buildStorage(): Promise<StorageRowsPayload> { // the only thing cached; its location, status and marker are read fresh every // render, because those are the safety facts. // THE DRIVES THAT ARE NOT ANSWERING, in words. From memory: the health pass - // (the block device's counters every 15 s) or the watchdog on a page's read - // is what found it (lib/storageHealth.ts). + // (the block device's counters, every `storage.health.passIntervalMs`) or the + // watchdog on a page's read is what found it (lib/storageHealth.ts). const now = Date.now(); + const clears = clearRuleText(healthTimings().clearAfterCleanPasses); const notAnswering: Record<string, string> = {}; for (const loc of locations) { const stall = stalledLocation(loc); @@ -103,7 +106,7 @@ export async function buildStorage(): Promise<StorageRowsPayload> { : ""; notAnswering[loc.id] = `${notAnsweringText(stall, now)} — ${stall.cause ?? "its root did not answer"}. ` + - `Pages and polls skip this drive until it answers twice in a row.${watched}`; + `Pages and polls skip this drive until it answers ${clears}.${watched}`; } } const store = await inspectSavedVideosStore(paths, settings); diff --git a/editor/app/storage/components/HealthTimingForm.tsx b/editor/app/storage/components/HealthTimingForm.tsx @@ -0,0 +1,119 @@ +"use client"; + +import { useActionState } from "react"; +import { + HEALTH_TIMING_BOUNDS, + HEALTH_TIMING_DEFAULTS, + HEALTH_TIMING_HINTS, + type StorageHealthSettings, +} from "yt-dlp-transcript-common/lib/storageHealthTimings"; +import { HEALTH_TIMING_FIELDS } from "../lib/healthTimingsForm"; +import { saveHealthTimingsAction, type HealthTimingsResult } from "../actions"; + +// THE DRIVE HEALTH TIMING — `settings.storage.health`, the five numbers the +// editor decides "this drive is mounted and not answering" by +// (lib/storageHealth.ts). The ruling (release 15, slice DT): when a stall is +// misjudged under heavy external-disk churn, the operator tunes these rather +// than the code. +// +// COLLAPSED, AND LAST. The defaults suit a healthy disk and most operators will +// never open it; the locations are what the page is for. +// +// A FIELD LEFT EMPTY IS THE DEFAULT, which is its placeholder. A field holds a +// value only where one was saved, so clearing it goes back to the default (and +// nothing is written for it). The server refuses a value out of range with a +// sentence, rather than storing another number than the one typed. +// +// TEXT INPUTS, NOT `type="number"`: a number input's own validation blocks the +// submit on "3.5" with a browser bubble instead of the sentence the action +// returns, and the action is the one validator. +// +// VALUES FROM THE TIMINGS MODULE ONLY (pure): this is a "use client" file. +// +// Accessible names: "drive health timing" (the block), the five fields' +// (lib/healthTimingsForm.ts), "save timing", "timing saved", "timing error" — +// none contains another (Playwright's getByLabel matches substrings). + +export function HealthTimingForm({ stored }: { stored: StorageHealthSettings }) { + const [state, formAction, pending] = useActionState< + HealthTimingsResult | undefined, + FormData + >(saveHealthTimingsAction, undefined); + const tuned = HEALTH_TIMING_FIELDS.filter((f) => stored[f.key] !== undefined).length; + + return ( + <details + aria-label="drive health timing" + className="rounded-xl border border-border bg-card px-4 py-3" + > + <summary className="cursor-pointer text-base font-semibold"> + Drive health timing + <span className="ml-2 text-sm font-normal text-muted-foreground"> + {tuned === 0 ? "defaults" : `${tuned} changed from the default`} + </span> + </summary> + <div className="mt-3 flex flex-col gap-3"> + <p className="text-sm text-muted-foreground max-w-3xl"> + How the editor decides that a drive is mounted but not answering, and + stops reading it until it answers again. The defaults suit a disk in + good health; change them only when a busy drive is being called not + answering (or a stalled one is not). An empty field is its default. + </p> + <form action={formAction} className="flex flex-col gap-3"> + <div className="grid gap-3 sm:grid-cols-2 lg:grid-cols-3"> + {HEALTH_TIMING_FIELDS.map((f) => { + const { min, max } = HEALTH_TIMING_BOUNDS[f.key]; + const unit = f.unit ? ` ${f.unit}` : ""; + return ( + <label key={f.key} className="flex flex-col gap-1 text-sm"> + <span className="font-medium"> + {f.label} + {f.unit && <span className="font-normal text-muted-foreground"> ({f.unit})</span>} + </span> + <input + type="text" + inputMode="numeric" + name={f.key} + aria-label={f.ariaLabel} + defaultValue={stored[f.key] === undefined ? "" : String(stored[f.key])} + placeholder={String(HEALTH_TIMING_DEFAULTS[f.key])} + className="rounded border border-border bg-background px-2 py-1 text-sm font-mono tabular-nums" + /> + <span className="text-xs text-muted-foreground"> + {HEALTH_TIMING_HINTS[f.key]} Default {HEALTH_TIMING_DEFAULTS[f.key]} + {unit}; {min}–{max} + {unit}. + </span> + </label> + ); + })} + </div> + <div className="flex flex-wrap items-center gap-3"> + <button + type="submit" + disabled={pending} + aria-label="save timing" + className="px-3 py-1.5 rounded-md border border-border text-sm font-medium disabled:opacity-50" + > + {pending ? "Saving…" : "Save timing"} + </button> + {state?.ok === true && ( + <span role="status" aria-label="timing saved" className="text-sm"> + {state.note} + </span> + )} + {state?.ok === false && ( + <span + role="alert" + aria-label="timing error" + className="text-sm text-destructive" + > + {state.error} + </span> + )} + </div> + </form> + </div> + </details> + ); +} diff --git a/editor/app/storage/lib/healthTimingsForm.test.ts b/editor/app/storage/lib/healthTimingsForm.test.ts @@ -0,0 +1,95 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { + HEALTH_TIMING_DEFAULTS, + HEALTH_TIMING_KEYS, +} from "yt-dlp-transcript-common/lib/storageHealthTimings"; +import { HEALTH_TIMING_FIELDS, parseHealthTimingsForm } from "./healthTimingsForm"; + +// Run with: pnpm -C editor exec tsx --test "app/**/*.test.ts" +// +// THE /storage DRIVE HEALTH TIMING FORM'S PARSE: an empty field is the default +// and is not written; a value equal to its default is not written either; a +// whole number in range is; anything else is refused with a sentence naming +// the field, and nothing is saved. + +function form(fields: Record<string, string>): FormData { + const f = new FormData(); + for (const [k, v] of Object.entries(fields)) f.set(k, v); + return f; +} + +test("every timing has one field, in the settings' order", () => { + assert.deepEqual( + HEALTH_TIMING_FIELDS.map((f) => f.key), + [...HEALTH_TIMING_KEYS], + ); + const names = HEALTH_TIMING_FIELDS.map((f) => f.ariaLabel); + assert.equal(new Set(names).size, names.length); + // No accessible name contains another (Playwright's getByLabel is a + // substring match). + for (const a of names) { + for (const b of names) if (a !== b) assert.ok(!a.includes(b), `${a} contains ${b}`); + } +}); + +test("empty fields are the defaults: nothing is written", () => { + assert.deepEqual(parseHealthTimingsForm(form({})), { ok: true, health: {} }); + assert.deepEqual( + parseHealthTimingsForm(form({ budgetMs: "", passIntervalMs: " " })), + { ok: true, health: {} }, + ); +}); + +test("a value in range is kept, trimmed; one equal to its default is not written", () => { + assert.deepEqual( + parseHealthTimingsForm( + form({ + budgetMs: " 4000 ", + passIntervalMs: String(HEALTH_TIMING_DEFAULTS.passIntervalMs), + probeTimeoutMs: "500", + clearAfterCleanPasses: "3", + inFlightPerLocation: "8", + }), + ), + { + ok: true, + health: { + budgetMs: 4_000, + probeTimeoutMs: 500, + clearAfterCleanPasses: 3, + inFlightPerLocation: 8, + }, + }, + ); +}); + +test("out of range is refused with the field, the range and the value — not clamped", () => { + assert.deepEqual(parseHealthTimingsForm(form({ budgetMs: "200" })), { + ok: false, + error: "Read budget must be between 500 and 60000 ms (got 200 ms).", + }); + assert.deepEqual(parseHealthTimingsForm(form({ inFlightPerLocation: "9" })), { + ok: false, + error: "Reads at once per drive must be between 1 and 8 (got 9).", + }); + assert.deepEqual(parseHealthTimingsForm(form({ clearAfterCleanPasses: "0" })), { + ok: false, + error: "Clean checks to clear must be between 1 and 10 (got 0).", + }); +}); + +test("not a whole number is refused", () => { + for (const bad of ["3.5", "-1", "3e3", "abc", "0x10"]) { + const r = parseHealthTimingsForm(form({ passIntervalMs: bad })); + assert.equal(r.ok, false, bad); + if (!r.ok) { + assert.equal(r.error, `Health check interval must be a whole number of milliseconds (got "${bad}").`); + } + } + const count = parseHealthTimingsForm(form({ clearAfterCleanPasses: "two" })); + assert.deepEqual(count, { + ok: false, + error: 'Clean checks to clear must be a whole number (got "two").', + }); +}); diff --git a/editor/app/storage/lib/healthTimingsForm.ts b/editor/app/storage/lib/healthTimingsForm.ts @@ -0,0 +1,91 @@ +import { + HEALTH_TIMING_BOUNDS, + HEALTH_TIMING_DEFAULTS, + type HealthTimingKey, + type StorageHealthSettings, +} from "yt-dlp-transcript-common/lib/storageHealthTimings"; + +// THE DRIVE HEALTH TIMING FORM ON /storage: its fields' names and units, and +// the one parse of what was submitted. Pure, and client-safe (it imports only +// the pure timings module), so the form draws its labels from here and the +// server action parses with it. +// +// AN EMPTY FIELD IS THE DEFAULT. The inputs show each default as a placeholder +// and hold a value only where the operator set one, so clearing a field puts +// that timing back on its default — and nothing is written for it +// (`sanitizeStorageHealth` keeps only what differs from a default). +// +// OUT OF RANGE IS REFUSED, NOT CLAMPED. The schema clamps a hand-edited value on +// read (a read never throws); a save from this form says which field is out of +// range and writes nothing, because storing another number than the one typed +// reads as a form that did not listen. + +export type HealthTimingField = { + key: HealthTimingKey; + // The visible label, and the input's accessible name (a contract once an + // e2e spec names it). + label: string; + ariaLabel: string; + // "ms", or "" for a count. + unit: string; +}; + +export const HEALTH_TIMING_FIELDS: readonly HealthTimingField[] = [ + { key: "budgetMs", label: "Read budget", ariaLabel: "read budget", unit: "ms" }, + { + key: "passIntervalMs", + label: "Health check interval", + ariaLabel: "health check interval", + unit: "ms", + }, + { + key: "probeTimeoutMs", + label: "Health check timeout", + ariaLabel: "health check timeout", + unit: "ms", + }, + { + key: "clearAfterCleanPasses", + label: "Clean checks to clear", + ariaLabel: "clean checks to clear", + unit: "", + }, + { + key: "inFlightPerLocation", + label: "Reads at once per drive", + ariaLabel: "reads at once per drive", + unit: "", + }, +]; + +export type HealthTimingsParse = + | { ok: true; health: StorageHealthSettings } + | { ok: false; error: string }; + +// What the form posted, as the stored block: each non-empty field a whole +// number in its range, and only the ones that differ from their default. +export function parseHealthTimingsForm(form: { + get(name: string): FormDataEntryValue | null; +}): HealthTimingsParse { + const health: StorageHealthSettings = {}; + for (const field of HEALTH_TIMING_FIELDS) { + const raw = form.get(field.key); + const text = typeof raw === "string" ? raw.trim() : ""; + if (text === "") continue; + const unit = field.unit ? ` ${field.unit}` : ""; + const n = Number(text); + if (!/^\d+$/.test(text) || !Number.isSafeInteger(n)) { + const what = field.unit === "ms" ? "a whole number of milliseconds" : "a whole number"; + return { ok: false, error: `${field.label} must be ${what} (got "${text}").` }; + } + const { min, max } = HEALTH_TIMING_BOUNDS[field.key]; + if (n < min || n > max) { + return { + ok: false, + error: `${field.label} must be between ${min} and ${max}${unit} (got ${n}${unit}).`, + }; + } + if (n !== HEALTH_TIMING_DEFAULTS[field.key]) health[field.key] = n; + } + return { ok: true, health }; +} diff --git a/editor/app/storage/page.tsx b/editor/app/storage/page.tsx @@ -1,7 +1,9 @@ import type { Metadata } from "next"; import Link from "next/link"; +import { getSettings } from "yt-dlp-transcript-common/lib/settings"; import { buildStorage } from "./buildStorage"; import { StorageLocationsTable } from "./components/StorageLocationsTable"; +import { HealthTimingForm } from "./components/HealthTimingForm"; // FORCE-DYNAMIC, and not as a formality. Every number on this page comes from a // probe of the machine taken when the page was asked for — whether a disk is @@ -33,6 +35,10 @@ export default async function StoragePage() { </p> <StorageLocationsTable payload={payload} /> + + {/* THE TIMINGS EVERY LOCATION IS JUDGED BY (settings.storage.health), + collapsed and last: a page-wide setting, not a fact about a row. */} + <HealthTimingForm stored={getSettings().storage.health ?? {}} /> </div> ); } diff --git a/editor/e2e/storage-locations.spec.ts b/editor/e2e/storage-locations.spec.ts @@ -678,3 +678,85 @@ test("evicting at any age needs the tick as well as the preview", async ({ await page.getByLabel("clip eviction age").selectOption("0"); await expect(page.getByLabel("confirm evicting every window")).not.toBeChecked(); }); + +// --- THE DRIVE HEALTH TIMING (release 15 slice DT) --------------------------- +// +// The five numbers the editor judges "mounted but not answering" by are +// `settings.storage.health`, edited in a collapsed block at the foot of the +// page. The claim: a value saved there is in settings.json and is what the page +// shows on the next load; an empty field is the default and writes nothing; a +// value out of range is refused with a sentence and writes nothing. (That the +// saved numbers feed the watchdog is unit-tested in lib/storageHealth.test.ts: +// no fixture here has a drive that stalls.) + +async function storedHealth(): Promise<Record<string, unknown> | undefined> { + const s = await readJson<{ storage?: { health?: Record<string, unknown> } }>( + "test-settings.json", + ); + return s.storage?.health; +} + +// React has attached its props to the form's button: a click now runs the +// action through React rather than as a pre-hydration form post. +async function timingFormHydrated(page: Page): Promise<void> { + await page.waitForFunction(() => { + const el = document.querySelector('[aria-label="save timing"]'); + return !!el && Object.keys(el).some((k) => k.startsWith("__reactProps")); + }); +} + +test("the drive health timing saves to settings.json and reads back", async ({ + page, +}) => { + test.setTimeout(90_000); + await resetData("one-youtube-channel-with-data"); + await writeSettings({ adminTitle: "Test Admin", minFreeDiskGB: 0 }); + expect(await storedHealth()).toBeUndefined(); + + await page.goto("/storage"); + const block = page.getByLabel("drive health timing"); + const summary = block.locator("summary"); + const budget = block.getByLabel("read budget"); + // Opened after hydration, so React never meets a `<details open>` it did + // not render. + await timingFormHydrated(page); + // COLLAPSED: the defaults suit a healthy disk. + await expect(summary).toContainText("defaults"); + await expect(budget).toBeHidden(); + await summary.click(); + await expect(budget).toBeVisible(); + // An empty field is its default, which is its placeholder. + await expect(budget).toHaveValue(""); + await expect(budget).toHaveAttribute("placeholder", "3000"); + await expect(block.getByLabel("reads at once per drive")).toHaveAttribute( + "placeholder", + "4", + ); + + await budget.fill("4000"); + await block.getByLabel("save timing").click(); + await expect(block.getByLabel("timing saved")).toContainText("A read may take 4 s"); + // ONLY WHAT DIFFERS FROM A DEFAULT IS WRITTEN: the four empty fields are not. + expect(await storedHealth()).toEqual({ budgetMs: 4000 }); + + await page.reload(); + await timingFormHydrated(page); + await expect(summary).toContainText("1 changed from the default"); + await summary.click(); + await expect(budget).toHaveValue("4000"); + await expect(block.getByLabel("health check interval")).toHaveValue(""); + + // OUT OF RANGE IS REFUSED, not clamped, and nothing is written. + await budget.fill("200"); + await block.getByLabel("save timing").click(); + await expect(block.getByLabel("timing error")).toHaveText( + "Read budget must be between 500 and 60000 ms (got 200 ms).", + ); + expect(await storedHealth()).toEqual({ budgetMs: 4000 }); + + // Emptied, it is the default again, and the block is gone from the file. + await budget.fill(""); + await block.getByLabel("save timing").click(); + await expect(block.getByLabel("timing saved")).toContainText("A read may take 3 s"); + expect(await storedHealth()).toBeUndefined(); +}); diff --git a/editor/instrumentation.ts b/editor/instrumentation.ts @@ -85,7 +85,8 @@ export async function register() { /* a failed re-pause must not block server readiness */ } - // THE DRIVE HEALTH PASS, every 15 s — and ON AN IDLE BOOT TOO, like the + // THE DRIVE HEALTH PASS, every `storage.health.passIntervalMs` (15 s by + // default; a save on /storage re-arms it) — and ON AN IDLE BOOT TOO, like the // probe below. It reads each location's block device counters (never the // drive) and keeps, in memory only, which drives are not answering; every // page and poll asks it before touching a drive, and its watchdog marks a diff --git a/plans/FACTS.md b/plans/FACTS.md @@ -7673,40 +7673,58 @@ this section is stale, by +16 near the top and +203 at the end; they are not rew ## The storage health gate (verified 2026-09-29, branch `r15/drive-stall`) The record is [`release-15.md`](release-15.md), "Slice DS, as shipped", with the parent's rulings -(Q1–Q5) and the review's fixes (M1–M3, L1–L10). Anchors are at the branch tip after the merge of -`main` `bab894db`. +(Q1–Q5) and the review's fixes (M1–M3, L1–L10); the timings became settings in "Slice DT, as +shipped". Anchors are at slice DT's tip (`r15/drive-timings`). - **A drive can be mounted and not answering.** Every in-process fs call on it waits on one of libuv's threads (4 by default, 16 in the editor's `start`) until it answers (~30 s for the observed USB reset loop); only a child process isolates a call. A child `stat` of a location's ROOT does not detect it reliably: the root's inode is in the kernel's cache whenever the drive was used lately. +- **The timings are settings: `settings.storage.health`** (slice DT). `budgetMs` (default 3000, + 500–60000), `passIntervalMs` (15000, 5000–300000), `probeTimeoutMs` (3000, 500–30000), + `clearAfterCleanPasses` (2, 1–10), `inFlightPerLocation` (4, 1–8: at most half the editor's 16 + file-access threads, review M1); the defaults, ranges, sanitizer + and words are `common/lib/storageHealthTimings.ts` (pure; the /storage form imports it). A read + clamps into the range and keeps only a value that differs from its default (an untuned file has no + `health` key); the /storage form refuses out of range with a sentence. Every number is read through + ONE accessor, `healthTimings()` (`storageHealth.ts:182`), from the timings as last applied on + `globalThis` — no file read per call. `applyHealthTimings(stored)` (`:193`) sets them: the health + pass on every pass (from the settings it reads), /storage's save at once + (`saveHealthTimingsAction`), and the `index` and `build stats` bins once at start. Any other process + with no pass (a CLI, `archilyzer doctor` until slice SG adds its line) runs on the defaults. A changed interval is told to + `onPassIntervalChange` (`:214`) subscribers, which re-arms the armed pass's timer; a raised cap + admits waiting calls, a lowered one is reached as calls return. `resetStorageHealth` keeps the + timings (configuration, not health); `setDriveCallBudget` is a test seam below the 500 ms floor and + wins over them. - **The state is `common/lib/storageHealth.ts`**, one map on `globalThis.__yttStorageHealth__` (the pass writes it from instrumentation's module copy; pages read it from theirs), with each entry's - `detector` and the counters' `device`. `recordLocationHealth` (`:182`): one `stalled` answer stalls - at once, and every transition to stalled refuses `onDrive`'s waiting calls; `HEALTH_CLEAN_TO_CLEAR` - (2) clean answers in a row clear it; `absent` is clean; a new root starts over. - `registerLocationHealth` (`:267`) creates entries with no answer. `stalledLocationForPath` - (`:319`) matches like `locationOfDataDir`; `stalledLocation` (`:335`) is by id AND root. -- **Detector 1, every 15 s: the block device's counters** (`detectLocationHealth`, - `lib/storageVolumes.ts:601`). The root's device from the last pass (findmnt `-J -T <root> -o - SOURCE,UUID`, raced against 3 s, only when there is none or its `/sys` entry stops reading; + `detector` and the counters' `device`. `recordLocationHealth` (`:251`): one `stalled` answer stalls + at once, and every transition to stalled refuses `onDrive`'s waiting calls; `clearAfterCleanPasses` + (2 by default) clean answers in a row clear it; `absent` is clean; a new root starts over. + `registerLocationHealth` (`:336`) creates entries with no answer. `stalledLocationForPath` + (`:388`) matches like `locationOfDataDir`; `stalledLocation` (`:404`) is by id AND root. +- **Detector 1, every `passIntervalMs` (15 s): the block device's counters** (`detectLocationHealth`, + `lib/storageVolumes.ts:607`). The root's device from the last pass (findmnt `-J -T <root> -o + SOURCE,UUID`, raced against `probeTimeoutMs` (3 s), only when there is none or its `/sys` entry stops reading; another volume's UUID names none; `[…]` stripped, `/dev/mapper` resolved, basename); then - `/sys/class/block/<dev>/stat` (`parseBlockStat`, `storageHealth.ts:735`): completed = fields 1 + 5 + `/sys/class/block/<dev>/stat` (`parseBlockStat`, `storageHealth.ts:831`): completed = fields 1 + 5 + 12 + 16 (reads, writes, discards, flushes), in flight = field 9. Stalled ⇔ in flight at both - samples AND nothing completed between; samples at least `MIN_COUNTER_INTERVAL_MS` (10 s) apart; the - first gives no verdict. in_flight counts only requests dispatched to the driver: one requeued + samples AND nothing completed between; samples at least `minCounterIntervalMs()` apart + (`storageVolumes.ts:472`: min(10 s, interval − 5 s), floored at half the interval — 10 s at the + default 15 s); the first gives no verdict. in_flight counts only requests dispatched to the driver: one requeued during a host reset is not counted, so a sample in that window can read clean (the watchdog covers - it). No device → the child `stat -L -c %F` probe (`probeLocationHealth`, `:386`). The samples are + it). No device → the child `stat -L -c %F` probe (`probeLocationHealth`, `:388`, raced against + `probeTimeoutMs`). The samples are on `globalThis.__yttHealthDetector__` (the pass and /storage's Refresh share them). -- **Detector 2, on every gated call: `onDrive(where, call)`** (`storageHealth.ts:633`). Refused with - no call on a stalled location; otherwise raced against `DRIVE_CALL_BUDGET_MS` (3 s; test seam - `setDriveCallBudget`); the budget covers the whole unit passed in. A timeout marks the location +- **Detector 2, on every gated call: `onDrive(where, call)`** (`storageHealth.ts:722`). Refused with + no call on a stalled location; otherwise raced against `budgetMs` (3 s by default, read as the call + starts; test seam `setDriveCallBudget`); the budget covers the whole unit passed in. A timeout marks the location stalled (since now) and throws `DriveNotAnsweringError`, leaving the call to settle — unless the location's device counters (read synchronously from `/sys` through the reader `storageVolumes.ts` - registers with `setCounterReader`, `:509`) moved since the call began: then the call is refused - as slow and nothing is marked. At most `DRIVE_CALLS_IN_FLIGHT` (4) calls per slot key in flight - (`acquireSlot`, `storageHealth.ts:501`): the rest queue in JS. A waiting call's deadline follows progress: every call that - returns on the key (in time or late) restarts it (`releaseSlot`, `:599`, re-arms every waiter), and a + registers with `setCounterReader`, `:515`) moved since the call began: then the call is refused + as slow and nothing is marked. At most `inFlightPerLocation` (4 by default) calls per slot key in + flight (`acquireSlot`, `storageHealth.ts:573`): the rest queue in JS. A waiting call's deadline follows progress: every call that + returns on the key (in time or late) restarts it (`releaseSlot`, `:673`, re-arms every waiter), and a waiting call is refused unmarked only when nothing on the key has returned for the budget plus a quarter of it (at most 250 ms) — never for the queue's depth alone. Waiting calls are refused at once by any transition to stalled; and when every slot is held by a call already past its budget @@ -7714,42 +7732,43 @@ The record is [`release-15.md`](release-15.md), "Slice DS, as shipped", with the and the location is marked stalled again only if the disk has completed nothing since the oldest of them began (otherwise "drive slow", unmarked). A slot is freed when its call really returns. The slot key: a configured location's id; a probe of another root under its id, that root (marks - nothing); a path on no configured location, the root it is under (`rootOfUnknownPath`, `:457`; + nothing); a path on no configured location, the root it is under (`rootOfUnknownPath`, `:529`; marks nothing). Do not nest it for one key. The timer is not unref'd. -- **The cadence** is `runStorageHealthPass` (`controller/storageWatch.ts:433`): prune, register, then - every location concurrently; every 15 s from `startStorageHealthWatch` (`:546`), plus one at arm - time, armed by `editor/instrumentation.ts` ABOVE the idle gate (it writes nothing). The five-minute - pass (`startStorageWatch`, `:521`) stays below it. A CLI process has no pass (its inspects are still - raced). `refreshLocationHealth` (`:485`) is /storage's Refresh. -- **The gate order in `inspectChannelMedia`** (`lib/channelMedia.ts:316`): config → memo (a +- **The cadence** is `runStorageHealthPass` (`controller/storageWatch.ts:442`): apply the timings it + read, prune, register, then every location concurrently; every `passIntervalMs` (15 s) from + `startStorageHealthWatch` (`:566`, re-armed when the interval changes), plus one at arm time, + armed by `editor/instrumentation.ts` ABOVE the idle gate (it writes nothing). The five-minute pass + (`startStorageWatch`, `:535`) stays below it. A CLI process has no pass (its inspects are still + raced). `refreshLocationHealth` (`:499`) is /storage's Refresh. +- **The gate order in `inspectChannelMedia`** (`lib/channelMedia.ts:317`): config → memo (a remembered `in-transition` is returned as is, anything else is gated first) → the relocation - marker (`:362`, corpus disk) → the gate (`:380`) → the link (corpus disk) → the target's `stat` - through `onDrive` (`:445`). A stall is never memoised. Other gated calls: `probeLocation` - (`storageVolumes.ts:255`, its stat and statfs through `onDrive`) and its memo; `volumeFreeBytes` - (`controller/storageLocations.ts:288`, `:325`, `:334`); `readChannelStat` (`controller/channels.ts:215`, - the walk through `onDrive`, `null` on a stall); the snapshot walk (`controller/channelSnapshot.ts:751`, + marker (`:367`, corpus disk) → the gate (`:386`) → the link (corpus disk) → the target's `stat` + through `onDrive` (`:447`). A stall is never memoised. Other gated calls: `probeLocation` + (`storageVolumes.ts:239`, its stat and statfs through `onDrive`) and its memo; `volumeFreeBytes` + (`controller/storageLocations.ts:289`, `:326`, `:335`); `readChannelStat` (`controller/channels.ts:205`, + the walk through `onDrive`, `null` on a stall); the snapshot walk (`controller/channelSnapshot.ts:753`, its listing, keep-latest keys and per-video unit); the recency tail reads; the move-root check; the saved-video store; `listSavedVideos` with `notAnswering`; the videos list, the video page, the Cleanup stage, the Storage stage's statfs and the media file route. `channelMediaStall(config)` - (`channelMedia.ts:231`) is the no-I/O question for a holder of a config. + (`channelMedia.ts:232`) is the no-I/O question for a holder of a config. - **`stalled` is a sixth `ChannelMediaStatus` and a sixth `StorageLocationStatus`** ("Not answering"). `HELD_REASON`, `MediaLocationBadge`'s two tables and `STORAGE_STATUS_LABEL` are the `Record`s that make tsc name every table a seventh would need. `isMediaHeld` holds it, so both pool-wide builds hold a stalled channel; the storage watch counts it as down (two passes pause), on a location or not, and the pause record carries `cause: "not-answering"` (`ChannelAutoPause`, `lib/channelPriority.ts`; absent = not there). -- **`inspectChannelMedia` is memoised for 5 s** (`CHANNEL_MEDIA_MEMO_MS`, `:265`), keyed by channels +- **`inspectChannelMedia` is memoised for 5 s** (`CHANNEL_MEDIA_MEMO_MS`, `:266`), keyed by channels dir, slug and configured `dataDir`, on `globalThis.__yttChannelMediaMemo__`. `{ fresh: true }` skips it and does not store; the deciders that pass it are listed in the record (the guard and its six callers, both movers, both builds, the watch, eviction, the re-point preflight, doctor). The runners' tick shares the status poll's `buildChannelWork` and so reads the memo. - `forgetChannelMedia` (`:291`) is called by the channel mover's marker writes and clear, + `forgetChannelMedia` (`:292`) is called by the channel mover's marker writes and clear, `clearRelocationMarker`, a re-point, /storage's Refresh and the e2e `invalidate-cache` route. - **`UV_THREADPOOL_SIZE`** defaults to 16 in `editor/package.json`'s `start` and in `docker/entrypoint.sh`. `ports.test.ts` reads every `${NAME:-N}` in a script as a port and names it as the one exception (`NUMERIC_NOT_PORTS`, `common/lib/ports.test.ts:29`). -- **Not covered:** a call already in flight when the drive stalls (at most four per drive for the - calls through `onDrive` — every page and poll path and the snapshot walk; a job's own reads that do +- **Not covered:** a call already in flight when the drive stalls (at most `inFlightPerLocation` per + drive for the calls through `onDrive` — every page and poll path and the snapshot walk; a job's own reads that do not go through it, `measureTree`, the index build's processing phase, the snapshot's sequential reconcile pass, are not capped); a hand-typed root is capped but never marked; a drive already stalled at boot before the second counter sample, unless a page reaches it; per-click server diff --git a/plans/release-15.md b/plans/release-15.md @@ -20,6 +20,7 @@ prompt carries its ruling, and this record carries what was built. Rules: | DS | `r15/drive-stall` | A stalled drive does not stop the editor answering | new `common/lib/storageHealth.ts`; `lib/{storageVolumes,channelMedia,channelMediaHold}.ts`, `controller/storageWatch.ts` and the gated callers; `/storage`, `/channels`, the videos pages; `UV_THREADPOOL_SIZE` (`editor/package.json`, `docker/entrypoint.sh`, `envVars.ts`) | | UT | `r15/umtool-trace` | umtool's build stops tracing the whole `umtool/` folder | per its prompt | | SS | `r15/site-scope` | The editor's site picker paints the stored site at once: the selection is a cookie | `editor/app/lib/activeSite{,Server,Actions}.ts` + `activeSite.test.ts`, `editor/app/components/SiteScope{Provider,Select}.tsx`, `editor/app/layout.tsx`, the scope lines of `editor/app/page.tsx` and `editor/app/channels/page.tsx`, a comment in `editor/next.config.ts`, `editor/e2e/site-scope.spec.ts`; records: `plans/FACTS.md` | +| DT | `r15/drive-timings` | The drive-health timings are settings (`settings.storage.health`), edited on `/storage` | new `common/lib/storageHealthTimings.ts` + test; `lib/{storageHealth,storageVolumes,channelMedia,storageLocations,settingsSchema,settingsDocs}.ts`, `controller/storageWatch.ts`, the `index` and `build stats` bins, `SETTINGS.md`; `/storage` (a form, its action and parse), the stall wording on `/channels` and the videos pages; `storage-locations.spec.ts`; records: `plans/FACTS.md` | | SG | `r15/stagit` | The source's history and diffs on the homepage, rendered by stagit at `/source/git/` | new `common/publish/sourceHistory.ts` + test; `common/publish/source.ts` (step 12b, the key, the digest, the deploy check) + test; `common/lib/{sourceManifest,paths,envVars}.ts`, `common/lib/themeConfig.ts` (moved from `components/`, which re-exports it); `common/bin/doctor.ts`; `homepage/app/{source/page.tsx,lib/source.ts,layout.tsx}`, `homepage/e2e/{source,source-history}.spec.ts`, `homepage/e2e/fixture-source.ts`, `homepage/playwright.config.ts`; `ENVIRONMENT.md`, `PUBLISH.md`; records: `plans/FACTS.md` | **Order:** IG → DS. DS adds a health gate inside `inspectChannelMedia`, which IG's hold calls @@ -1089,6 +1090,215 @@ It reproduced M1 with a two-tab probe under the queue lock. editor's built bundle, so all of it takes effect only after the editor is rebuilt and restarted. After that, each browser's first visit migrates its localStorage selection once. +### Slice DT, as shipped — the drive-health timings are settings (2026-09-30) + +Branch `r15/drive-timings` off `main` `6b8aa450` (slice DS merged), worktree +`~/Projects/r12-paths-fix` (block #12: editor 4201, test 4211, export 4210), one Opus implementer. +Scratch files `dt-*` in the job's `tmp`. The ruling (operator, 2026-09-30): the drive-health timings +are configurable in `settings.json`, editable on `/storage`, with today's constants as the defaults; +if a stall is misjudged under heavy external-disk churn, the operator tunes the numbers rather than +the code. + +**What was fixed in code.** Slice DS judged "mounted and not answering" by five constants in +`lib/storageHealth.ts`: the watchdog's budget (`DRIVE_CALL_BUDGET_MS`, 3 s), the health pass's cadence +(`HEALTH_PROBE_INTERVAL_MS`, 15 s), the child-stat and findmnt race (`HEALTH_PROBE_TIMEOUT_MS`, 3 s), +the clean answers that clear a stall (`HEALTH_CLEAN_TO_CLEAR`, 2) and the calls in flight per +location (`DRIVE_CALLS_IN_FLIGHT`, 4). The constants are gone; tsc named every reader. + +- **The setting** is `settings.storage.health`, a nested block of five optional keys: + + | Key | Default | Range | Takes effect | + |---|---|---|---| + | `budgetMs` | 3000 | 500–60000 | the next `onDrive` call (read as the call starts) | + | `passIntervalMs` | 15000 | 5000–300000 | at once on a save from `/storage` (the armed pass re-arms its timer); a hand edit, at the next pass | + | `probeTimeoutMs` | 3000 | 500–30000 | the next pass or Refresh (child `stat` and findmnt) | + | `clearAfterCleanPasses` | 2 | 1–10 | the next answer | + | `inFlightPerLocation` | 4 | 1–8 (review M1: at most half the editor's 16 file-access threads) | the next slot taken; a raised cap admits waiting calls at once, a lowered one is reached as calls return | + + - The type, defaults, ranges, sanitizer, words and `SETTINGS.md` docs are one pure module, + `common/lib/storageHealthTimings.ts` (the `/storage` form, a client file, imports it). + `StorageSettings.health` is optional; `sanitizeStorage` keeps a block only when something is in it. + - **A read clamps; only what differs from a default is kept.** `sanitizeStorageHealth` rounds each + number and clamps it into its range (the schema's convention: a read never throws), drops a + non-number, and drops a value equal to its default. An untuned `settings.json` therefore has no + `health` key, and a save of every default removes it. + - The mediaRoot migration keeps a `health` block that spells no `locations` (a hand edit). +- **The read path: one accessor.** `healthTimings()` (`storageHealth.ts:182`) returns the timings as + last applied, on the health state's `globalThis` object, every absent key its default. There is no + settings memo in lib (`getSettings()` reads the file on every call), so the accessor reads no file: + the numbers are applied into memory by `applyHealthTimings(stored)` (`:193`): + - the health pass, at the start of every pass, from the settings it already reads for its + locations (so at boot, and within one pass of a hand edit); + - the `/storage` save, at once (`saveHealthTimingsAction`); + - the `index` and `build stats` bins, once, before the build (a CLI process has no pass). + `lib/storageHealth.ts` stays free of I/O. `resetStorageHealth` keeps the timings (configuration, not + health). `setDriveCallBudget` stays the test seam below the 500 ms floor and wins over them. +- **The derived numbers stay derived.** The slot wait's grace is still a quarter of the budget, at + most 250 ms. The counters' sample spacing is `counterSampleMinimumMs(passIntervalMs)` = + min(10 s, interval − 5 s), floored at half the interval (`minCounterIntervalMs()` in + `storageVolumes.ts:472`): 10 s at the default 15 s, as before, and at every interval from 10 s up the + ruling's formula exactly (see the decisions table for the floor). +- **The re-arm.** `startStorageHealthWatch` arms at `healthTimings().passIntervalMs` and subscribes + with `onPassIntervalChange` (`storageHealth.ts:214`); `applyHealthTimings` tells the subscribers when + the interval changed, and the watch clears and re-arms its interval (log line `[storage] health + pass re-armed: every N s`). The subscription is on `globalThis`, so the save (a page's module copy) + reaches the pass (instrumentation's). An explicit `intervalMs` (the tests') follows nothing. +- **The cap, changed live.** `acquireSlot` reads the cap each time. `releaseSlot` hands a slot on only + while the key is at or under the cap, and otherwise gives it back; `admitWaiters` lets calls already + waiting take the slots a raised cap adds. The overdue refusal compares with the cap as it is, and + its words count the overdue calls ("N reads on it have not answered") instead of naming the constant. +- **`/storage`: Drive health timing.** A collapsed `<details>` at the foot of the page + (`components/HealthTimingForm.tsx`, `aria-label="drive health timing"`; its summary says + "defaults" or "N changed from the default"). Five text inputs (`inputMode="numeric"`), each with the + default as its placeholder, a hint line in the operator's words, and the default and range + (`HEALTH_TIMING_HINTS`); the interval's hint says a save re-arms the check at once. An empty field is + the default. Accessible names (new): `read budget`, `health check interval`, `health check timeout`, + `clean checks to clear`, `reads at once per drive`, `save timing`, `timing saved`, `timing error`; + none contains another. No existing name changed. + - The server action (`saveHealthTimingsAction`, `app/storage/actions.ts`) parses with + `lib/healthTimingsForm.ts`: a value out of range, or not a whole number, is **refused** with a + sentence naming the field, the range and the value ("Read budget must be between 500 and 60000 ms + (got 200 ms)."), and nothing is written. Otherwise it saves a patch of the storage block through + `saveSettings`, applies the timings, revalidates `/storage`, and answers with the timings now in + force. +- **The surfaces say the setting, not the constant.** `/storage`'s "not answering since …" line stays, + and its "until it answers twice in a row" follows the clear count (`clearRuleText`: once, twice in a + row, N times in a row). So do the Refresh note (and its "checked every N s"), the `/channels` + volume chip's title (a new `clears` field on `ChannelVolume`), the health pass's stall log line, + and the media-not-answering notice ("within N s", "checked every N s"). The `onDrive` header, the + module header, and the comments that said "3 s" or "every 15 s" in the gated callers now name the + setting and its default. FACTS' "The storage health gate" has a bullet for the setting and names the + keys where it named the constants. + +**Commits** + +| Commit | What | +|---|---| +| `13411edc` | `common:` `lib/storageHealthTimings.ts`; `settings.storage.health` in the schema, the docs table and the migration; `healthTimings()`, `applyHealthTimings()`, `onPassIntervalChange()`; the pass applies what it reads and re-arms; the live cap; `minCounterIntervalMs()`; the two bins; `SETTINGS.md`; the comments. Tests. | +| `41ad2f65` | `editor:` the Drive health timing form, its action and parse (+ unit test), the page; the Refresh note, the stalled line, the media notice, the volume chip's title; the comments. | +| `b92df1fe` | `editor(e2e):` `storage-locations.spec.ts`: the drive health timing case. | +| `4372abaf` | `common:` the cap's hint says to keep it well under the editor's 16 file-access threads. | +| `fd2b8d63` | `plans:` the first version of this section and the slices row; FACTS; the changelog. | +| `85c2407b` | `common:` review M1: `inFlightPerLocation` at most 8; the hint, the docs string, `SETTINGS.md`, the test values (the form's too). | +| `6a24ac4e` | `common:` review L4: a timeout on no known location names the budget the call ran against. Test. | +| `e7a525f3` | `editor:` review L2: a storage patch of the locations keeps `storage.health` (a `saveSettings` merge case). | +| this commit | `plans:` the rulings and the review in this section; FACTS; the report. | + +**Tests** (unit; no test stalls a real drive) + +| File | What it pins | +|---|---| +| `lib/storageHealthTimings.test.ts` (7, new) | The defaults are the constants they replace, each in its range. Absent, empty, an array, a string, a number: every default, and no block (`getSettings()` with no file, `defaultSiteSettings()`). A read clamps (200 → 500, 900000 → 300000), rounds (7.6 → 8), drops a string and an unknown key, and drops a value equal to its default (also one that rounds onto it). **The settings.json round trip** through `writeSettings`/`getSettings`: `{budgetMs: 4000, clearAfterCleanPasses: 2}` is written as `{budgetMs: 4000}` and read back; a save of the default removes the key; a hand-edited 99 reads as 8 (16 before review M1). The block survives the mediaRoot migration. The spacing (15 s → 10 s, 300 s → 10 s, 12 s → 7 s, 10 s → 5 s, 8 s → 4 s, 5 s → 2.5 s). The words. | +| `lib/storageHealth.test.ts` (+7) | **The accessor feeds the watchdog:** a stored `budgetMs: 200` applies as 500 ms (the floor), and a 700 ms unit is refused and marks the location ("within 0.5 s"); on the defaults the same unit answers. **The cap:** with `inFlightPerLocation: 2` the third call waits, and runs when a slot frees. A cap raised from 1 to 3 admits the two waiting calls at once; lowered to 1 with three in flight, two returns bring it to one and the fourth call still waits, and runs on the third return. **The clear count:** with 3, two clean answers do not clear and the third does; with 1, one does. A changed interval is told to the subscribers once, an unchanged one and other keys are not, and an unsubscribed one hears nothing. The test seam's budget wins, and a reset keeps the timings. The existing constants' assertions read `HEALTH_TIMING_DEFAULTS`. After review L4: a call on a hand-typed root whose budget changes while it is out (80 ms, then 5 s) is refused naming 0.08 s (the pre-fix code named 5 s). | +| `lib/storageHealthCounters.test.ts` (+1) | The spacing follows the applied interval: at 8 s, samples 3,999 ms apart give no verdict and 4,000 ms apart compare (stalled); at 300 s it is 10 s. The existing case reads `minCounterIntervalMs()` (10 s). | +| `lib/storageHealthProbe.test.ts` (+1) | With `probeTimeoutMs: 500` applied and no `timeoutMs` passed, a child that sleeps 20 s is `stalled` after 0.5–2.5 s. | +| `controller/storageWatch.test.ts` (+2) | **Every pass applies what it reads:** `clearAfterCleanPasses: 3` and `budgetMs: 5000` in the settings are in force after the first pass; its stall line says "until it answers 3 times in a row"; two clean passes do not clear and the third does; a pass handed its locations reads no settings and leaves the timings. **The re-arm:** the armed watch, told `passIntervalMs: 5000` (the save's apply), logs the re-arm and runs its second pass 4.5–7 s later (15 s at the default); a stopped watch re-arms nothing; one armed with an explicit interval does not follow. | +| `editor/app/storage/lib/healthTimingsForm.test.ts` (5, new) | One field per key, in order, with accessible names none of which contains another. Empty and blank fields write nothing. A value in range is kept and trimmed; one equal to its default is not written. Out of range is refused with the field, range and value, for a millisecond field and both counts (the cap: 8 kept, 9 refused, after review M1). "3.5", "-1", "3e3", "abc", "0x10" and "two" are refused as not whole numbers. | +| `editor/app/settings/saveSettings.test.ts` (+1, review L2) | A storage patch of `{ locations, defaultLocationId }` (what the /storage location actions write) keeps `storage.health` and `savedVideosLocationId`. | + +**e2e** (`storage-locations.spec.ts`, new case "the drive health timing saves to settings.json and +reads back"): after hydration, the block is collapsed and says "defaults"; opened, `read budget` is +empty with placeholder 3000 (`reads at once per drive`: 4); 4000 saved → "A read may take 4 s" and +`test-settings.json` holds `storage.health` = `{budgetMs: 4000}` and nothing else; after a reload the +summary says "1 changed from the default", the field reads 4000 and the interval is empty; 200 is +refused with the sentence and the file is unchanged; emptied and saved, "A read may take 3 s" and the +`health` key is gone. The fixture's settings are its own `test-settings.json`. + +#### Gates (logs `$T/dt-*.log`) + +- **tsc** (all workspaces): clean before every commit — 58 s on the tree of the first three code + commits, 51 s after the hint's; after the review, 44 s (M1) and 58 s (L4, L2). +- **Unit:** + + | Suite | Result | + |---|---| + | common | **2,318/2,318**, 50 s (`main`'s 2,301 + 17); after the review **2,319/2,319**, 50 s (+1, L4) | + | editor unit | **100/100** (95 + 5); after the review **101/101** (+1, L2) | + | `test:scripts` | 194 passed, 2 skipped (196), as at DS | + | mcp | not run: no mcp file and nothing it imports changed | + +- **Docs:** `settings example --check`, `docs env --check` and `docs files --check` all exit **0**, + before and after the review (`SETTINGS.md` regenerated in `13411edc`: the `health` row and the + `storage.health` table; and in `85c2407b` for the cap's range; `settings.json.example` unchanged, + the default block has no `health`). +- **Build:** the editor's `next build`, with the primary's `transcripts/` linked in (`ln -sT`) and + capped at 5 GB with no swap: **34 s, max RSS 1,648,128 KB**, exit 0. The link was removed after the + build, and nothing ran through it. +- **e2e** (editor, detached and queued; `$T/dt-specs.txt`, the nine `storage` + `channels` specs DS + ran), with the new case in `storage-locations`: + + | Run | At | Result | + |---|---|---| + | 1 | `b92df1fe` (the hint's edit landed on disk while it ran; nothing it asserts) | **39 passed, 0 failed, 12 skipped, 2.7 min** | + | 2 | `4372abaf` (the code as shipped) | **39 passed, 0 failed, 12 skipped, 2.5 min** | + + DS's 38 plus the new case. The 12 skips are `channels-rack-audit`, which needs `E2E_RACK_SHOTS`. + Neither run waited in the queue. Not re-run after the review, as the parent directed: the fixes + change a range, a message and a unit test, and the review's own run of `storage-locations` at + `fd2b8d63` passed 9/9. +- **Numbers tool:** none. + +#### Found and left + +- **Other CLI processes run on the defaults.** The `index` and `build stats` bins apply the settings; + `archilyzer doctor` and any other command whose inspects go through `onDrive` race them against the + default 3 s. **Doctor's line belongs to slice SG**, which owns `common/bin/doctor.ts` (ruled below): + `applyHealthTimings(settingsFromFile(paths.settingsFile).storage.health)` at its start. +- **Two drives at the cap's maximum can still hold every thread.** Review M1 lowered the maximum to + 8, half of `UV_THREADPOOL_SIZE` (16), so one drive that stops answering cannot hold them all; two + such drives at 8 each can, as two at 4 hold half. The cap is per location, not per process. +- **A hand edit of `passIntervalMs`** re-arms at the next pass, so it can wait up to the old interval + (at most 5 minutes). A save from `/storage` re-arms at once. +- **A refused save clears the typed value.** React resets a form after its action returns, so the + field shows the stored value again; the error sentence names the refused value. +- **umtool's reachability twin** (`checkChannelReachable`) still has no health state (DS left it; + `umtool/**` is another slice's). + +#### Decisions the operator could overturn + +| What I assumed | The alternative | +|---|---| +| The counters' spacing is the ruling's min(10 s, interval − 5 s), **floored at half the interval**. Without the floor it is 0 at the 5 s minimum interval (1 s at 6 s), so a Refresh just after a pass would compare two samples milliseconds apart, the case DS's review kept the spacing for. From 10 s up the two agree. **Ruled at review: the floor stays.** | The formula as ruled, 0 at 5 s. Or a higher minimum interval (10 s) | +| A read clamps an out-of-range value (the schema's convention); the `/storage` form refuses it with a sentence and writes nothing. | The form clamps too, and says what it stored | +| Only a value that differs from its default is written, so a save of 3000 for the budget writes nothing, and a later release's new default reaches it. | Write what the operator saved, pinning the default of the day | +| The slice's test as specified (`budgetMs: 200` → a 300 ms unit refused) is below the ruled 500 ms floor, so the test stores 200, shows it clamped to 500, and refuses a 700 ms unit; the default passes the same unit. | Lower the floor so 200 applies | +| The timings are applied into the health state (the pass on every pass, the save at once, the two bins), not read from `settings.json` on each call: lib has no settings memo, and `onDrive` is on every page's hot path. | A time-limited memo of `getSettings()` inside the accessor (a file read at most every few seconds, and lib/storageHealth.ts no longer free of I/O) | +| A save re-arms the pass's timer at once, through a subscription on `globalThis`. | Leave the running timer; the new interval at the next restart | +| A lowered cap is reached as calls return; calls already in flight are not refused. A raised cap admits waiting calls at once. **Ruled at review: it stays.** | Refuse the calls over the new cap | +| **After review M1:** the cap's maximum is 8, half the editor's 16 file-access threads (the ruling said 16). | 16, with a warning on the form and in `SETTINGS.md` (as first shipped) | +| The form's inputs are text with a numeric keypad, so the action's sentence is the only validation. | `type="number"` with `min`/`max`: the browser's own bubble, and "3.5" blocked before the action | +| `resetStorageHealth` (a test seam and the e2e `invalidate-cache` route) keeps the applied timings. | Reset them to the defaults until the next pass | + +**What runs which code, for the rollout.** The accessor, the pass's apply and re-arm, and the form +are in the editor's built bundle and its instrumentation, so all of it takes effect after the editor +is rebuilt and restarted. Nothing is written until the operator saves the form; with no `health` key +the editor runs on today's numbers. A CLI `archilyzer index` or `build stats` reads the settings +itself. + +#### Rulings (parent, 2026-09-30) + +| Question | Ruling | Where | +|---|---|---| +| The counters' spacing: the ruled min(10 s, interval − 5 s) is 0 at a 5 s interval | The half-interval floor stays | as built (`13411edc`) | +| A lowered cap: refuse the calls in flight over it, or reach it as they return | Reached as calls return; nothing in flight is refused | as built (`13411edc`) | +| `archilyzer doctor` runs on the default timings | Its one `applyHealthTimings` line goes to slice SG, which owns `common/bin/doctor.ts` | "Found and left" | + +#### Review + +**Verdict: SHIP AFTER FIXES** (`dt-review.md` in the job's scratch). No High. The review re-ran every +gate (tsc, common 2,318, editor unit 100, the three docs checks, `storage-locations` 9/9, a clean +`merge-tree` against `main`), checked the writer paths, the migration, the accessor's `globalThis` +home, the re-arm, the live cap's arithmetic and 20 FACTS anchors, and found no accessible-name +collision in the three specs that open `/storage`. + +| Finding | Ruling | Where | +|---|---|---| +| M1: `inFlightPerLocation` could be 16, the editor's whole thread pool, so one drive that stops answering could hold every thread; the hint only warned | The maximum is 8; the hint, the docs string, `SETTINGS.md`, the test values and these records | `85c2407b`, this commit | +| L1: the parent's rulings were not in the record | The block above; the doctor bullet names SG | this commit | +| L2: no test pinned that a location write keeps `storage.health` | One `mergeSettingsPatch` case | `e7a525f3` | +| L3: a refused save clears the typed value (React's form reset) | As recorded in "Found and left" | as built | +| L4: the no-location timeout's words read the budget at throw time, not the one the call ran against | The detail carries `secondsText(budget)`; a test that changes the budget mid-call | `6a24ac4e` | + ### Slice SG, as shipped — the source's history and diffs, rendered by stagit (2026-09-30) Branch `r15/stagit` off `main` `6b8aa450`, worktree `~/Projects/r12-source-mirror` (block #13: