# Release 15 — storage and build hardening `main` at `99d4d76a` (release 14's T1, H1 and H2 merged). No plan file of its own: each slice's prompt carries its ruling, and this record carries what was built. Rules: `plans/tools/implementer-rules.md`, with the commit trailer this release's prompts give. **The standing choices** (not re-opened): - **An unmounted drive is not an empty channel,** for any build that walks the pool. The stats build already holds such a channel ([`stats-cache-key.md`](stats-cache-key.md)); the index build gets the same hold here. - **A slice that needs another slice's file stops and says so**; it does not edit it. - **Nothing is edited in the primary checkout**; each slice has its own worktree, and the parent merges with `git merge --no-ff` only on a clean tree. ## The slices | Slice | Branch | What | Owns | |---|---|---|---| | IG | `r15/index-hold` | The index build holds an unreachable channel instead of emptying it | `common/controller/buildIndex.ts` + new `buildIndex.test.ts`, `common/controller/buildStats.ts` (the hold's words move to a shared module), new `common/lib/channelMediaHold.ts`, `common/lib/envVars.ts`, `ENVIRONMENT.md`; records: `plans/{STATE,FACTS,stats-cache-key}.md` | | DS | `r15/drive-stall` | A stalled drive does not stop the editor answering | new `common/lib/storageHealth.ts`; `lib/{storageVolumes,channelMedia,channelMediaHold}.ts`, `controller/storageWatch.ts` and the gated callers; `/storage`, `/channels`, the videos pages; `UV_THREADPOOL_SIZE` (`editor/package.json`, `docker/entrypoint.sh`, `envVars.ts`) | | UT | `r15/umtool-trace` | umtool's build stops tracing the whole `umtool/` folder | per its prompt | | SS | `r15/site-scope` | The editor's site picker paints the stored site at once: the selection is a cookie | `editor/app/lib/activeSite{,Server,Actions}.ts` + `activeSite.test.ts`, `editor/app/components/SiteScope{Provider,Select}.tsx`, `editor/app/layout.tsx`, the scope lines of `editor/app/page.tsx` and `editor/app/channels/page.tsx`, a comment in `editor/next.config.ts`, `editor/e2e/site-scope.spec.ts`; records: `plans/FACTS.md` | | DT | `r15/drive-timings` | The drive-health timings are settings (`settings.storage.health`), edited on `/storage` | new `common/lib/storageHealthTimings.ts` + test; `lib/{storageHealth,storageVolumes,channelMedia,storageLocations,settingsSchema,settingsDocs}.ts`, `controller/storageWatch.ts`, the `index` and `build stats` bins, `SETTINGS.md`; `/storage` (a form, its action and parse), the stall wording on `/channels` and the videos pages; `storage-locations.spec.ts`; records: `plans/FACTS.md` | | SG | `r15/stagit` | The source's history and diffs on the homepage, rendered by stagit at `/source/git/` | new `common/publish/sourceHistory.ts` + test, `common/publish/__fixtures__/fakeStagit.ts`; `common/publish/source.ts` (step 12b, the key, the digest, the deploy check) + test; `common/lib/{sourceManifest,paths,envVars}.ts`, `common/lib/themeConfig.ts` (moved from `components/`, which re-exports it); `common/bin/doctor.ts`; `homepage/app/{source/page.tsx,lib/source.ts,layout.tsx}`, `homepage/e2e/{source,source-history}.spec.ts`, `homepage/e2e/fixture-source.ts`, `homepage/playwright.config.ts`; `ENVIRONMENT.md`, `PUBLISH.md`; records: `plans/FACTS.md` | **Order:** IG → DS. DS adds a health gate inside `inspectChannelMedia`, which IG's hold calls through its public signature. UT is independent. The shared files are `editor/CHANGELOG.md`'s `[Unreleased]` and this record. ## Record ### Slice IG, as shipped — the index build holds an unreachable channel (2026-09-29) Branch `r15/index-hold` off `main` `99d4d76a`, worktree `~/Projects/r12-paths-fix` (block #12: editor 4201, test 4211, export 4210), one Opus implementer. Scratch files `ig-*` in the job's `tmp`. The ruling: an index build meets a channel it cannot read the way the stats build already does. It holds the channel instead of reading it as empty, and a full rebuild with one held refuses unless `ARCHILYZER_INDEX_ALLOW_HELD=1`. **What was wrong.** `scanSource` read `data/` with a bare `catch { continue }`. A relocated channel whose drive was unmounted (a dangling `data/` link) therefore contributed no videos. The removal pass then dropped every record the channel had, the page writers rewrote its shared transcript tree empty and removed its subs tree, and the next site build published the channel as gone. Only the `Diff: … -R removed` line showed it. - **The hold** (`common/controller/buildIndex.ts`): - `scanSource` asks `inspectChannelMedia({ channelsDir }, slug, cfg)` for every channel that is not excluded and not social, before it reads `data/`. - Any status but `ok` or `in-place` holds the channel, with a reason and its storage location's label, and no path. - It asks again after the walk, so a drive that goes away mid-walk holds the channel instead of dropping the videos after that point. - A `readdir(data/)` that fails holds the channel too: `its data directory could not be read ()`. - The exception is ENOENT on an `in-place` channel: a channel with nothing downloaded (or its media deleted), which is emptied as before. The log now says it: `Channel : no data/ directory; indexed as a channel with no videos.` - A per-video metadata `stat` failing with anything but ENOENT or ENOTDIR holds the channel, logged as `a video in its data directory could not be read ()`. - **A held channel keeps everything:** - its `mtimes` records, since the removal pass skips its keys, and so its `sums`, `cues`, `subs`, `digests` and `byChannel` entries; - its shared transcript, subs and digest trees: not rewritten, not pruned, and not removed by the top-level cleanups; - its subs and digest stats, carried from their sub-DBs, so the per-site manifests still list it; - its availability states. The maybe-missing overlay skips it, since its `availability.json` files are on the missing drive and a missing one reads as "maybe missing". The last build's `videoState` entries for it are carried over. - The per-site summaries are built from LMDB, so the sites built next still list its videos. - **A curated-tag change while held:** the re-apply pass re-derives the held channel's records in LMDB as usual, but its pages are not written. The pages-pending flag is therefore kept (and logged) while any channel is held, and the first build with the drive back writes them. - **A full rebuild refuses** (a schema change, or a first build with no index). - The scan now runs BEFORE the clear, so a refusal leaves the index untouched. `scanStartedAt` is still taken at the scan's start. - The message names each channel with its location's label and says why a clear would publish it as gone. It gives the ways out, mounting first (the same words as the stats build's refusal), then the override by name and where it is set: the command's own environment for a CLI run, and the editor's own environment (which takes a restart) for its Build index job or a site build started from it. Their children inherit the editor's `process.env`. - With `ARCHILYZER_INDEX_ALLOW_HELD=1` (1/true/yes/on, declared in `envVars.ts`, `ENVIRONMENT.md` regenerated), the build proceeds. The held channel's records go with the clear; they cannot be carried across a format change. Its shared trees are left on disk, and it is out of the index until its media is back and an index build runs. - **The words are shared:** new `common/lib/channelMediaHold.ts` (`isMediaHeld`, `HELD_REASON`, `heldReason`, `describeHeld`, `HELD_WAYS_OUT`). `buildStats.ts` uses it, and its messages are byte-identical (its case (i) passes unchanged). - **The result and the log:** `BuildIndexResult.heldChannels`. The log carries one line per held channel (`Channel : ; its N indexed video(s) are kept as they are, not rescanned, and its transcript, subtitle and digest pages are left as they are.`). The `Diff:` line ends in ` Held: N channel(s), K video(s) kept.` and the `Done in` line in ` Held, their media not readable: .` Both prefixes are unchanged: the e2e helper waits for `Done`. **Who reports it** (every caller of `buildIndex`): | Caller | What it shows | |---|---| | `pnpm archilyzer index` (also the root, export and homepage `build:index` scripts) | The log on stdout, with the three lines above. It exits 0 on a hold. On the refusal it exits 1 with the message on stderr (`runIfEntryPoint`). `common/bin/build-index.ts` is unchanged: it prints the log and discards the result. | | A site build's data phase (`build:data`, from `buildSite`'s steps and `buildAll`'s Phase A in `common/publish/build.ts`) | The same lines, in the site build's job log. A refusal stops the steps at the data phase (`Build failed (exit 1)` / `Data phase failed (exit 1)`), so nothing is composed or deployed from an emptied index. The hub build does not run the index. | | The editor's **Build index** job (`buildIndexAction`, `editor/app/sites/lib/buildAction.ts`) | The lines in the job's log on `/sites` and `/jobs`. A refusal fails the job with the message. The action discards the result, and it was left unchanged: the log already carries every held channel, and `editor/app/sites/**` belongs to another slice this release. | | `pnpm ops build-index --wait` | Follows the job's log, so it prints the same lines. | | The tests | `heldChannels`. | **Commits** | Commit | What | |---|---| | `c6ae51b0` | `plans:` this record: the header, the slices, and empty Record and Rollout sections. | | `51328098` | `common:` the hold's words move to `lib/channelMediaHold.ts`; the stats build uses them, with its messages unchanged. | | `ffb01d8e` | `common:` the index build's hold, the refusal and its override, `heldChannels`, the log lines; `envVars.ts` + `ENVIRONMENT.md`; new `buildIndex.test.ts` (9 cases). | | `702cd0da` | `plans:` this section; FACTS "The index build's hold" (and the two stale statements in "The stats cache key" corrected); the STATE follow-up closed; `stats-cache-key.md` "Left" marked closed; the editor changelog. | | `9ec48351` | `common:` review L3 + L5: one unreadable video directory is logged apart from an unreadable `data/`; the refusal says where the override is set. | | `5af516b7` | `common(test):` review M1: case (i), the drive lost mid-walk. | | this commit | `plans:` the review's findings to their commits; FACTS L1, L2 and the anchors; the changelog's override sentence (L5). | **Tests** (`common/controller/buildIndex.test.ts`, the real `buildIndex` over a temp corpus). The drive channel is seeded with `buildStats.test.ts`'s `seedDriveChannel` shape and unmounted by renaming its media root away. Every path is pinned under a temp root, and case (z) spies on node:fs writes. | Case | What it pins | On the pre-change `buildIndex.ts` | |---|---|---| | (a) | Unmounted, incremental build with another change: records kept, the drive's transcript and subs trees byte-identical (manifest included), the site still lists both videos and its subs count; the log lines, with the label and no path | removed 2, not 0 | | (b) | Availability carried: `maybe_missing` stays and a post-scan confirmation stays `available`; `videoState` unchanged | d1 no longer published | | (c) | A full rebuild refuses (message, ways out, the variable, no path); the index and pages untouched; the CLI exits non-zero; with the override it holds (records cleared, pages kept); the drive back re-adds both | no refusal | | (d) | A first build (no index) with a channel held refuses too | no refusal | | (e) | The drive back: held set empty, a video added to the drive meanwhile indexed, 0 changed | removed 2 while away | | (f) | A really empty in-place channel (an empty `data/`, and no `data/`) is emptied, not held; the missing `data/` is logged | the new log line only (the emptying matched, as it should) | | (g) | An unreadable `data/`, and an unreadable video dir mid-walk (mode 000), hold their channels with `EACCES` | removed 3 | | (h) | A tag rule added while held: the held pages untouched, the flag kept; the drive back writes the tag to them and clears the flag | the drive's pages emptied | | (i) | The drive lost MID-WALK: `node:fs/promises` `stat` unmounts it right after the walk's first drive metadata stat, so every later stat in the channel is ENOENT. The second look holds the channel: 0 removed, records and pages unchanged, the log line | removed 3 of 4; the same with only the second look deleted from the new code | | (z) | No write outside the temp root | passes on both | The pre-change column was run with the old `buildIndex.ts` swapped in once, with the `heldChannels` assertions removed so each case reached its first substantive assertion. #### Gates (at `ffb01d8e`, and after the review at `5af516b7`; logs `$T/ig-*.log`) - **tsc** was clean before every commit: 78 s at the branch point, 48 s at `ffb01d8e`, 50 s at `5af516b7`. - **Unit:** | Suite | Result | |---|---| | common | 2,219/2,219 at `ffb01d8e` (the branch point's 2,210 plus the 9 new cases), 77 s; **2,220/2,220** at `5af516b7` (case (i) added), 78 s | | editor unit | 87/87 | | `test:scripts` | 191 passed, 1 skipped (192) | | mcp | 271/271 | - **Docs:** `docs env --check`, `docs files --check` and `settings example --check` all exit **0**, at both points. - **Build:** the editor's `next build`, with the primary's `transcripts/` linked in and capped at 5 GB with no swap: 64 s, max RSS 1,642 MB. The link was removed after the build, and nothing ran through it. - **e2e** (editor, detached and queued; the spec list is every spec that runs Build index: `availability`, `build`, `channel-build-toggle`, `chat-only`, `duplicate-shorts`, `jobs`, `regional-vtt-fallback`, `tags`, passed as `e2e/.spec.ts` so `availability` does not also match `pre-clean-availability`): **35 passed, 0 failed, 3.6 min**, after 1 min 45 s in the queue. No e2e fixture has an unreachable channel, so these confirm the hold changes nothing for a readable corpus. Not rerun after the review: its two log-wording changes touch no spec (`git grep "could not be read" editor/e2e` finds only the curated-tag preview and the title filter). - **Numbers tool:** none. #### Found and left - **A drive that drops during the processing phase** (after the scan) is not covered, for a changed video whose metadata was read before the drop. Its transcript read is caught as "no cues", so its cues are removed. A failed sub-track read is skipped, and with none left its subs are removed. A failed digest load leaves no digest, which is removed. `mtimes` is then written with the video's current mtimes, so the loss lasts until any of its tracked mtimes (metadata, transcript, subs, availability, digest) moves. These are per-video reads inside the worker, left to the slice that handles drive stalls. - **A drive that drops and comes back inside one walk** is not covered either: the second look finds it, and the videos skipped in between are removed. - **`inspectChannelMedia` is asked twice per channel** (before and after the walk), three syscalls each today. Slice DS puts a health gate inside it, and that gate runs twice per channel per index build. - **Under the override,** a held channel is listed on its sites with 0 videos, and its subs and digest counts are left out of the site manifests. Its old shared page trees stay on disk, and a compose copies them. - **While a channel is held after a tag change,** the pages-pending flag stays set. Every build until the drive is back is then a full page walk; it skips unchanged pages by hash, so it rewrites none. - **The digest tree's retention has no test.** The code path mirrors the subs tree's, and no fixture carries a digest sidecar. - **The editor shows a hold only in the job log.** A `/sites` badge would be in `editor/app/sites/**`. #### Decisions the operator could overturn | What I assumed | The alternative | |---|---| | A held channel's shared pages are not written at all. **Ruled at review: keep the skip.** | Rewrite them from the kept records: byte-identical on an incremental build, and a tag change would reach them at once. After an override rebuild the records are gone, and a rewrite would publish the channel empty. | | Under the override, the held channel's records go with the clear. | Carry them across the clear. A schema change means the stored format moved, so the old records cannot be trusted. | | An unreadable `data/`, or a per-video `stat` failing with anything but ENOENT, holds the channel. | Hold only for the statuses `inspectChannelMedia` reports, and keep treating other read errors as "no videos", as the old code did. | | A channel with no `data/` is logged, one line per build. | Stay silent, as before; the ruling asked for a log. | | The CLI exits 0 on a hold, since the hold is the safe outcome; the refusal exits 1. **Ruled at review: both stay.** | Exit non-zero so scripts notice; a site build's data phase would then fail whenever a drive is out. | | The words live in a new `lib/channelMediaHold.ts` shared with the stats build. | Duplicate them in `buildIndex.ts` and leave `buildStats.ts` untouched. | #### Review **Verdict: SHIP AFTER FIXES** (`ig-review.md` in the job's scratch). No High. Every write path a held channel could reach was traced, and the refusal fires before anything is deleted. | Finding | Where | |---|---| | M1: the second look after the walk had no test | `5af516b7`: case (i). It fails with 3 of 4 removed when that look is deleted. | | L1: FACTS said an unreachable drive reaches `notIndexable` only through the override | this commit: a video that reached the drive after the last index build that could read it is counted there by a stats-only run between the drive's return and the next index build, which heals it. | | L2: the processing-phase gap also loses subs and digests, until any tracked mtime moves | this commit, in "Found and left" and FACTS. | | L3: one unreadable video dir was logged as the whole data directory | `9ec48351`: `a video in its data directory could not be read ()`; case (g) expects it. | | L4 (optional): a narrower `HELD_REASON` type | **Left**, as the review allowed. The current contract already fails tsc on a new status. | | L5: the refusal did not say where the override is set | `9ec48351` (message, case (c)) and this commit (changelog). | | L6: check the held set before routine builds resume | A rollout note; the parent records it. | | L7: stale trees under the override | Already in "Found and left". | | Q2: the commit trailer | Ruled correct. | **What runs which code, for the rollout.** Every CLI command and every spawned data phase runs the checkout's code, so they hold from the moment `main` has this branch. The editor's in-process **Build index** button runs its built bundle, so it holds only after the editor is rebuilt and restarted. ### Slice UT, as shipped — umtool's build stops tracing its dot-directories (2026-09-29) Branch `r15/umtool-trace` off `main` `ccf90892`, worktree `~/Projects/r12-source-mirror` (block #13: editor 4301, test 4311, export 4310), one Opus implementer. Scratch files `ut-*` in the job's `tmp`. The ruling: find the one expression that widens the clip-audio route's trace and fix it there; exclude the fixture, the e2e build and env files as a second line; narrow or drop the `ignoreIssue`; make the trace guard read a build back and close the release 14 review's L1 and L3. **What was wrong.** With the primary's e2e fixture in place, umtool's `app/api/clip/[key]/audio/route.js.nft.json` listed 2,167 files: the 463 its sibling routes list, 178 under `.e2e-song/` (the fixture, where `make-fixture.mjs` links the song data), 1,525 under `.next-e2e/` (the e2e dev server's build directory, 1.1 GB) and `.env.local`. Turbopack's warning for it ("Encountered unexpected file in NFT list", the "whole project was traced" text) was silenced by the config's `ignoreIssue`. The traces are not consumed while `output: "standalone"` stays off, so nothing broke; a fixture with more in it, or a standalone build, would have carried it. **The bisect.** One change per build, in this worktree with four probe files planted in `.e2e-song/probe/` and `.next-e2e/probe/`; the audio route's trace, total / under dot-directories. The route as on `main`: **467 / 4**. | Change (line on `main`) | Entries | |---|---| | `existsSync(file)` :51 stubbed | 467 / 4 | | **the join `path.join(CACHE_DIR, `${stamp}.${asMp3 ? "mp3" : "wav"}`)` :58 written as a string concatenation** | **463 / 0** | | `existsSync(cached)` :61, `readFile(cached)` :62, `mkdir(CACHE_DIR)` :78, `readFile(tmpMp3)` :89, `rename(tmpMp3, cached)` :90, `writeFile(tmp)` :93 or `rename(tmp, cached)` :94 stubbed, each alone | 467 / 4 each | | the `tmpWav` / `tmpMp3` joins :82-83 as concatenations; `writeFile(tmpWav)` :84 stubbed; both `writeFile`s stubbed; either `writeFile` opted out | 467 / 4 each | | the opt-out on the :58 join | 467 / 4 | | the :58 ternary hoisted into a `const ext` | 467 / 4 | | **:58 without the ternary** (`${stamp}.wav`) | **463 / 0** | | :58 as `path.join(CACHE_DIR, asMp3 ? `${stamp}.mp3` : `${stamp}.wav`)` | 463 / 0 | | :58 as a ternary of two joins | 463 / 0 | | opt-outs on `existsSync(cached)` and `readFile(cached)`, or either alone; both stubbed | 467 / 4 each | | all five readers and writers of `cached` stubbed | 467 / 4 | | **all five stubbed, and the opt-out on the :58 join** | **463 / 0** | | all five stubbed, and the join as a concatenation | 463 / 0 | With the primary's fixture (the table's 4 are 1,704 there), on top of the last-but-one row: | Change | Entries | |---|---| | one `existsSync(path.join(/* opt-out */ CACHE_DIR, …))` | 2,167 / 1,704 | | the same `existsSync` opted out as well | 463 / 0 | | `existsSync(/* opt-out */ cached)` as the only reader of the opted-out join | 463 / 0 | **The expression** is the :58 join, and in it the ternary inside the template literal. The join and the fs calls on its value each trace the pattern (the join alone with every reader stubbed; the readers alone with the join opted out; `existsSync` is one such reader), so no single opt-out or stub cleared it. Without the ternary, the pattern stays out of the dot-directories. **The fix** (`a305b956`). `umtool/lib/paths.mjs` gains `cacheFile(name)`, `path.join(/* opt-out */ CACHE_DIR, name)`, re-exported by `lib/paths.ts`. The audio route names all four of its cache files through it (`cached`, `tmpWav`, `tmpMp3`, and `tmp`, now `cacheFile(`${stamp}.wav.tmp`)`, the same path as `${cached}.tmp` in that branch). A value returned by a function from another module is opaque to the tracer, so the call site traces nothing: the route lists 463, as its siblings do. The video and face-frame routes join `CACHE_DIR` with a fixed extension; they were measured clean (the video route's `.mp4` join did not reach the `.mp4` in `.e2e-song`) and name their cache files the same way, so no route joins `CACHE_DIR` itself. The run-time paths are unchanged. **The second line** (`f18fa034`). `outputFileTracingExcludes: { "/*": ["./.e2e-song/**/*", "./.next-e2e/**/*", "./.env*"] }` (Next 16.2.3's `05-config/01-next-config-js/output.md`: route globs to globs from the project root; Turbopack reads the key natively, `collect-build-traces.js` is the webpack path). Measured alone, with `main`'s route: 2,167 → 463. **The warning stays silenced, narrow as it was (path + title), with its measured reason.** Dropping it shows one warning on every build, and after the fix it is still true: - 66 of the 68 routes trace umtool's whole tree outside dot-directories: its 361 files, `next.config.ts` (the file the warning names) among them. Only `_global-error` and `_not-found` do not. - Path and fs calls on env, home-directory and parameter values do it. Opting out every path op in `song/paths.mjs`, `lib/paths.mjs` and `lib/paths.ts` left 31 of the 68 routes clean (the audio route 463 → 102). Opting out all 319 path ops in the 53 modules that have one left 49 clean. The two clean before are among them. The rest come through fs calls; the next import trace the warning names is `lib/report/snapshots.mjs`. - That walk skips dot-directories and does not enter symlinks. The warning names the same file for it as for the audio route's, so it cannot tell the two apart; the second line and the guard below cover the dot-directories instead. The config's comment says all of this. **A symlinked directory is not entered by these patterns.** The planted link `umtool/.e2e-song/data/planted` → `/transcripts/channels` gave 0 entries before and after the fix, and an in-root link to `common/` (675 files) gave 0 on `main`'s route. What the pattern reached was the fixture's real files. **The guard** (`scripts/next-build-trace.test.mjs`, `9a375ecd`, and after the review `a4d100b4`, `c3e8a2c7`), 6 → 10 tests: - **(a) It reads umtool's last build back.** Every `.nft.json` under `umtool/.next` but the build's own `cache/` and `dev/` fails on an entry outside the repo, under `transcripts/`, or through any name starting with a dot but the build's own directory and `node_modules/.pnpm`. Unit tests pin `forbiddenTrace` and which directories are read. - **It skips, saying so,** with no build, and (review M1) when `umtool/.next/BUILD_ID` is older than `umtool/next.config.ts` or any umtool module the guard scans. The message gives the build's time, the first newer file and how to rebuild. A merge or checkout gives the changed files new mtimes, and umtool runs under `next dev`, which does not refresh `.next`, so a stale build skips instead of failing on a call already fixed. - **What it can see** (review L2). Without the excludes, `main`'s route failed it with 1,704 entries (the first 20 listed); the primary's pre-fix build fails it with 1,705 (the review: `test-results/.last-run.json` too). With the excludes in place, a pattern like that one shows only through a name they miss: `test-results/.last-run.json` after an e2e run, `.next-shots`, the corpus, a path outside the repo. A checkout with no e2e run behind it is blind to it; the fix at the call is what keeps the route clean. - **(b) The scan set follows relative imports** out of the listed folders, to any depth. It adds the review's L1 modules and no others: `common/bin/_publicFile.ts`, `homepage/content/docs.ts` (895 → 897) and seven `umtool/song` modules, `reasons`, `archive-url`, `pitch`, `flatness`, `clipwindow`, `deplosive`, `orderfeat` (211 → 218; `song/paths.mjs` was listed by hand before and is now reached). A test pins them, and that a song CLI and a common CLI stay out. The comments that called them CLI-only are gone. - **(c) The checked calls** add `open`, `writeFile`, `appendFile`, `createWriteStream` and their Sync forms, and the `fs.promises.` / `fsPromises.` prefixes. No new finding. - **(d)** `common/lib/paths.ts` `under(first, ...rest)` (`8b3409c2`): the opt-out sits before a named first argument. `getPaths()` hashed identical, old module against new, from the repo root, `editor/` and a directory outside the repo (56 keys). The guard passes; every caller type-checks. - The header no longer says the env and home directory are unfollowed or that the opt-out is documented. The nested-call exemption's comment says it is a simplification (below). **FACTS**, "A path joined from `process.cwd()` …", corrected: - "documented": the Next docs list `turbopackIgnore` only for `import()`, `require()`, `require.resolve()` and `new Worker()`. The path form is Turbopack's own advice, in the warning's text in the 16.2.3 binary, and Next's own server uses it on the join and on the fs call around it (`next/dist/server/next-server.js:620`; review L4). - "an outer fs call on an opted-out `path.join` is covered": it is not (the second bisect table). - "Not followed by the tracer: `os.homedir()` and `process.env.*`": such values are dynamic parts, and make patterns over the app directory. The synthetic-HOME build showed that a pattern walk does not enter symlinks. - The guard's entry, the worktree caveat (no fixture either) and `under()`'s shape. **Commits** | Commit | What | |---|---| | `a305b956` | `umtool:` `cacheFile`; the audio, video and face-frame routes name their cache files through it | | `f18fa034` | `umtool:` `outputFileTracingExcludes`; the `ignoreIssue` kept, its comment the measured reason | | `8b3409c2` | `common:` `under(first, ...rest)` | | `9a375ecd` | `scripts:` the post-build check, the relative-import scan set, the opens and writes, the header | | `750fc850` | `plans:` this section; FACTS; the editor changelog | | `a4d100b4` | `scripts:` review M1: the post-build check skips a build older than the code it judges | | `c3e8a2c7` | `scripts:` review L1 (only the build's own `cache/` and `dev/` unread, pinned), L2 (the check's comment says what it can see), L3 | | `b4d6607d` | `umtool:` review L3 in the `ignoreIssue` comment | | `31bb0f8c` | `plans:` the review's findings to their commits; FACTS (M1, L2, L3, L4); the changelog (M1); the rollout note | | `7c6d4969` | merge of `main` (`721ed0eb`, release 14 HS and S1); clean, the changelog bullet still under `[Unreleased]` | | this commit | `plans:` the gates after the review and the merge | #### Gates (logs `$T/ut-*.log`) - **tsc** clean over the combined tree before the first commit, 241 s (load average ~26). The commits are independent pieces of that tree; since then only a comment changed in a type-checked file (`cacheFile`'s, in `umtool/lib/paths.mjs`). - **`test:scripts`:** 194 passed, 1 skipped (195): `main`'s 191 + 1 and the guard's three new tests. Before the final build it failed exactly the post-build test, on a build of `main`'s route. - **common:** 2,220/2,220, 107 s. - **umtool's build, capped at 5 GB with no swap, with the primary's `transcripts/` linked in, the primary's `.e2e-song` and `.next-e2e` hard-linked in, an empty `.env.local`, and the planted link** (all removed afterwards; none committed): | Tree | Wall | User | Max RSS | Audio route | `.e2e-song` | `.next-e2e` | `.env*` | `transcripts` / planted | |---|---|---|---|---|---|---|---|---| | `main`'s route and config | 23 s | 57 s | 809 MB | 2,167 | 178 | 1,525 | 1 | 0 / 0 | | the branch | 34 s | 64 s | 765 MB | 463 | 0 | 0 | 0 | 0 / 0 | The wall times swing with the machine's load (another slice's e2e and builds ran alongside); compile was 7.5 s and 10.7 s, TypeScript 12 s and 19 s. - **The editor's build**, capped, with the corpus linked in (`paths.ts` changed): 97 s wall, 159 s user, max RSS 1,576 MB (IG's: 64 s / 1,642 MB, under less load); 0 of its 81 traces' entries under `transcripts/`, and none that `forbiddenTrace` refuses. - **e2e** (umtool's own filter, `SONG_DIR=~/reports/quartering-uh-song/data`, queued): `find.spec.ts` and `triage.spec.ts` fetch the audio route, `faces.spec.ts` the face-frame route, and `find.spec.ts` names the video route: **6 passed, 29 skipped, 0 failed, 27 s** (6.3 min with the queue). The skips are the fixture's: this machine has no `wav48/`, `asr/` or `media/`, so every spec that fetches one of the three routes skipped, and they are not exercised at run time here. What stands for that: calling `cacheFile` gives the same path as the old expression for all six names (the four audio names, and the video and frame temporaries). - **After the e2e run** (which built this worktree its own fixture and `.next-e2e`), a last capped build with the corpus linked: the audio route 463, none under a dot-directory; `test:scripts` 194 passed, 1 skipped. The first run of that `test:scripts` failed `queue-lock.test.mjs`'s FIFO case once (`S1E1S3E3S2E2`) under a load average of about 26; it passed on the rerun, and this slice does not touch the queue lock. - **Numbers tool:** none. - **After the review and the merge of `main`** (at `7c6d4969`): - tsc clean, 98 s; - common **2,229/2,229** (`main`'s 2,229), 108 s; - `test:scripts` **195 passed, 1 skipped (196)**: `main`'s 191 + 1 and the guard's four new tests. The two runs before it each failed `queue-lock.test.mjs`'s "prints a banner naming the holder while waiting" under a load average of about 26, the known flake (alone, 11/11 twice); this slice does not touch the queue lock; - umtool's build, capped, with the corpus linked and this worktree's own e2e fixture present: 48 s wall, max RSS 796 MB, the audio route 463, none under a dot-directory; the post-build check passes on it; - the same build with its `BUILD_ID` set back to 2026-09-01 (in place, then put back; nothing committed): the check skips, `umtool/.next was built 2026-09-01T04:00:00.000Z, before umtool/next.config.ts (219 changed since); rebuild umtool (…) to check its traces`. With the mtime put back it passes again. #### Found and left - **The whole-folder trace in 66 routes** (above). Bounded to umtool's own files; cleaning it means opt-outs on hundreds of path and fs calls on unknown values, which no static check can find. - **The guard's nested-call exemption.** Without it, four calls would need an outer opt-out: `common/lib/paths.ts:311` (`existsSync` of one file), `export/app/changelog/page.tsx:9` and `homepage/app/changelog/page.tsx:24` (one file each; other slices own them), and `umtool/report-to-video/brand.mjs:82` (the brand kits, which a standalone build needs). Each traces the file or files it reads. Left, and the comment says so. - **Which fs calls Turbopack traces is not established per call.** `existsSync` does (the second table); the writes were added to the guard without a measurement, since an extra opt-out costs nothing. - **The post-build check reads umtool only.** In this worktree the editor's 81 traces pass the same rule. The review's Info saw `editor/.env` in the primary's; a worktree has none, so it was not re-measured. - **The gate command in `implementer-rules.md`, in a worktree that has a `transcripts/` directory,** makes `transcripts/transcripts` and builds without the corpus where the paths point. This worktree had one (an `index.mdb` from 2026-09-28), and my first two corpus-linked builds ran like that. The numbers above are from builds that set it aside and put it back. `ln -sT` would refuse instead. - **A checkout whose umtool build predates its code skips the post-build check** until umtool is rebuilt (review M1). The primary's `umtool/.next` is from before this slice, so the check skips there until the rollout rebuilds it. #### Decisions the operator could overturn | What I did | The alternative | |---|---| | A `cacheFile` helper in `lib/paths.mjs`, used by all three routes that join `CACHE_DIR`. **Ruled at review: keep; one way to name a cache file.** | Opt-outs on the audio route's join and on every fs call on its value, in that route only | | The `ignoreIssue` stays, narrow, with the measured reason in its comment | Drop it: one warning on every build, naming one route's import trace | | The post-build check refuses any dot-named path but the build's own and `node_modules/.pnpm` | Refuse only `.e2e-song`, `.next-e2e`, `.env*` and `.git` | | The post-build check covers umtool only. **Ruled at review: umtool only.** In the primary the editor's traces would fail it (75 of 81, `editor/.env` among them) and the export's `.export-index` entries would be misjudged. | Also read the editor's, the export's and the homepage's builds | | The static check keeps its nested-call exemption, documented as a simplification. **Ruled at review: it stays; the four sites are safe.** | Require the outer opt-out: four new findings, two in files other slices own | | A stale build skips the post-build check (review M1) | Fail on it, as first shipped | #### Review **Verdict: SHIP AFTER FIXES** (`ut-review.md` in the job's scratch). No High. The reviewer found the fix's nine call-site paths byte-identical, the bisect logs in agreement with the tables, the excludes' key and shape right, and 0 dot entries across all 70 traces of a fresh build. | Finding | Where | |---|---| | M1: a stale umtool build turns `test:scripts` red, pointing at a call already fixed | `a4d100b4`: the check skips a build older than `next.config.ts` or any module it scans; the changelog, FACTS and "Found and left" say so; the rollout note below | | L1: `cache`/`dev` skipped at any depth | `c3e8a2c7`: only directly under `.next`; a test with a route directory named each | | L2: with the excludes in place the check cannot see the original defect | `c3e8a2c7` (the test's comment), guard (a) above and FACTS: it sees only a name the excludes miss | | L3: "cleaned 31 / 49" | `c3e8a2c7`, `b4d6607d`, this commit: "left 31 / 49 of the 68 clean" | | L4: FACTS cited the native binary's string as Next's runtime | this commit: `next-server.js:620` | | L5: the build-gate command's `ln -s` | The parent's (`implementer-rules.md`) | | Info: the editor's and export's primary builds carry the same class of widening | Recorded in the decisions table; a later slice's | **For the rollout.** Rebuild umtool in the primary first, under the cap (`timeout -s KILL 240 systemd-run --user --scope -q -p MemoryMax=5G -p MemorySwapMax=0 pnpm --filter umtool exec next build`), then run `pnpm run test:scripts` there. That is the only proof on the real fixture, which carries `.env.local`, `.next-shots` and `test-results/.last-run.json`, two of them names the excludes do not cover. Until that rebuild the post-build check skips in the primary, saying why. umtool's code changes nothing at run time. ### Slice DS, as shipped — a stalled drive does not stop the editor answering (2026-09-29) Branch `r15/drive-stall` off `main` `ccf90892` (slice IG merged), worktree `~/Projects/r12-paths-fix` (block #12: editor 4201, test 4211, export 4210), one Opus implementer. Scratch files `ds-*` in the job's `tmp`. The ruling: a drive that is mounted and not answering must not stop the editor answering. Two detectors find the stall without the editor waiting on the drive, and pages and polls do not touch it in-process while it is not answering. The parent's five rulings on the first pass (Q1–Q5) and the review's fixes (M1–M3, L1–L10) are applied; the tables at the end map each to its commit. **What was wrong.** Node runs every filesystem call on libuv's thread pool, four threads by default. On a drive that has stalled (an SMR disk in a USB enclosure resetting under a long write) each call blocks its thread for about 30 s. The home page, `/channels` and the three-second auto-queue status poll each `stat`ted every relocated channel's target in-process and uncached; the one-second job-list poll walked the whole `data/` of every channel with a job listed, 64 calls wide; and the recency layer read metadata tails 32 at a time. Four blocked calls were enough for no page, poll or job log to answer. The storage watch saw nothing of it: its five-minute pass asks "is the disk here", and a stalled disk is here. - **The health state** (new `common/lib/storageHealth.ts`): one map per process on `globalThis` (`__yttStorageHealth__`, the house pattern: the watch writes it from instrumentation's module copy and pages read it from theirs), by location id: state `ok | stalled | absent`, `since`, the last check, a clean streak, the cause, and the `detector` that gave the last verdict. In memory only. - One `stalled` answer marks the location stalled at once. Two clean answers in a row clear it (an `absent` answer is clean). A miss in between starts the count again. A re-pointed root starts the location over. Locations no longer configured are pruned; every configured one is registered (`registerLocationHealth`) before a pass asks anything, so the watchdog can find a channel's location before the first verdict. - Pure of I/O and without execa, so `lib/channelMedia.ts` can ask it. - **Detector 1, every 15 s: the block device's own counters** (`detectLocationHealth`, `common/lib/storageVolumes.ts`; ruling Q1(a)). A child `stat` of the root is answered from the kernel's inode cache whenever the drive was used lately, so it can say "ok" while the reads that reach the device wait out a reset loop. Instead, per location per pass: - the root's device: `findmnt -J -T -o SOURCE,UUID` as a child raced against 3 s, a `[subvolume]` suffix taken off, `/dev/mapper/*` resolved to its `dm-N` (a read of `/dev`), the basename. Asked only when the root has no device yet or its device's `/sys` entry stops reading (a replug under another name): a findmnt per pass would leave one child stuck per pass during a long stall (L10). A UUID other than the location's recorded one names no device (the root is then a directory on another filesystem, not the drive). - `/sys/class/block//stat`, which never touches the drive: completed = reads (field 1) + writes (5) + discards (12) + flushes (16) where the kernel counts them (in_flight counts those too, so a long SMR media-cache flush alone moves completions; L2), and requests in flight (9). Against the previous sample for the location (same device, at least `MIN_COUNTER_INTERVAL_MS` = 10 s earlier): **stalled ⇔ in flight at both AND nothing completed between**; anything else is clean. The first sample gives no verdict. The samples are on `globalThis` (`__yttHealthDetector__`), so the pass and `/storage`'s Refresh, in different module copies, compare against one previous sample (L5). - **No device** (a container, no findmnt, a tmpfs or network source, no `/sys` entry) falls back to the child `stat` probe (`probeLocationHealth`): `stat -L -c %F -- ` raced against 3 s, the child SIGKILLed and not waited for; `directory` is `ok`, anything else `absent`, no binary `ok`. - The verdict names its detector (`"counters" | "stat"`), the health state records it (and the counters' device, which the watchdog reads), and `/storage`'s line says which watched the drive. - **Detector 2, on every gated call: a 3 s watchdog** (`onDrive(where, call)`, `lib/storageHealth.ts`; ruling Q1(b)). The detector that cannot be fooled by a cache: a page or poll that actually reaches the drive finds out. - Refused at once, with no call, when the location is stalled. - Otherwise raced against `DRIVE_CALL_BUDGET_MS` (3 s). A call that has not answered marks its location stalled (since now, cause "a read in the editor did not answer within 3 s") and throws `DriveNotAnsweringError`; the caller answers `stalled`. The call is left to settle on its own: its thread is the stated limit. - **Slow is not stalled** (L6): on a timeout, when the counters detector has named the location's device, its counters are read (synchronously, from `/sys`, so the check does not wait behind the pool it is judging) and compared with a reading taken when the call began. Requests completed meanwhile: the drive is slow; the call is refused and nothing is marked. - **The budget covers a whole unit of work** (L8): a video directory's reads, a page's reads of one video, go through as one call, so a slow drive still answering can be marked by one long unit (unless the counters show it completing, above). - **At most `DRIVE_CALLS_IN_FLIGHT` (4) calls per slot key are in flight.** The rest wait in a queue of its own (not libuv's), so a 64-wide walk that meets a stall puts four calls on the drive, not 64. A slot is released when its call really returns. The queue (M1): - every transition to `stalled` refuses the waiting calls at once, whoever decided it (the pass, the watchdog, a Refresh); - a wait's deadline follows progress (re-review M4): every call that returns on the key, in time or late, restarts the deadline of every call waiting on it, and a waiting call is refused, without marking, only when nothing on the key has returned for the budget plus a grace of a quarter of it (at most 250 ms; the grace lets the calls it waits behind, whose timers start a moment later, time out and mark first). A deep queue on a drive that is busy but answering therefore waits as long as it takes; - calls past their budget are counted per key (`overdue`), each with the counters reading taken when it began; when every slot is held by one, a new call is refused at once, and the location is marked stalled again (even if the pass has since cleared it) unless the disk has completed requests since the oldest of them began: then it is slow, not stalled, and nothing is marked (the same test as a timeout's). - **The slot key** (M2, L7): a configured location's id; a probe of another root under a location's id is keyed by that root and marks nothing; a path on no configured location (a root typed by hand) is keyed by the root it is under (`//data` → ``), so it holds at most four threads too, and nothing can mark it. Calls are not nested for one key. The timer is not `unref`'d: it is cleared the moment the call answers. - **The cadence** (`common/controller/storageWatch.ts`): `startStorageHealthWatch` arms the 15 s health pass and runs one at once; **`editor/instrumentation.ts` arms it above the idle gate**, beside the storage boot probe (M2): it is in memory and writes nothing, and without it an idle boot has no registered locations and a stall the watchdog marks is never cleared. The five-minute pass (`startStorageWatch`), which may write an auto-pause, stays below the gate. All locations are asked concurrently, each bounded by its own timers, and an overrunning pass is not stacked. A transition is logged (`[storage] "": drive not answering — ; …` / `answering again (ok)`). A CLI process has no pass, but its gated calls still go through the watchdog. - **The gate and the watchdog, by caller:** | Caller | On a stalled location | Through `onDrive` | |---|---|---| | `inspectChannelMedia` (`lib/channelMedia.ts`) | **The relocation marker is read first** (ruling Q2: it is in the channel dir, on the corpus disk), so a channel mid-move on a stalled drive reads `in-transition`. Then the gate: status `stalled`, detail `drive not answering (location "