commit def2ab2ec238f59fc89f052182fd8c87aeafb939
parent d1c283701e325f94d51e467900f610f5cda0108b
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Tue, 29 Sep 2026 21:01:41 -0400
plans: slice IG as shipped — the index build's hold, who reports it, the tests against the old code, gates (common 2,219, editor unit 87, test:scripts 191+1, mcp 271, docs clean, editor build 64 s / 1,642 MB, e2e 35/35 in 3.6 min); FACTS "The index build's hold"; the STATE follow-up closed; the editor changelog
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Diffstat:
5 files changed, 231 insertions(+), 14 deletions(-)
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -3,6 +3,7 @@
## [Unreleased]
- **Transcripts that arrived after a video was first seen are counted.** The stats behind the homepage, the hub and every site's charts were cached per video and refreshed only when the video's metadata changed, so a transcript that came later — a Whisper run days after the download, or a video downloaded after the last index build — never reached them, and a video with YouTube captions alone had no transcription date. Counts and charts were low; the homepage could show a site with 0 transcripts, 0 channels and 0 hours while it served its videos. A stat is now also redone whenever the index re-reads the video, every transcript has a date, and a captioned video is dated by when its captions arrived rather than by a later Normalize run, so its place on "Transcribed over time" can move. **After updating, rebuild and restart the editor before anything else:** until then, **Build stats dataset** runs the old code and would undo the new stats, while a site, hub or homepage build already runs the new code — and the first stats build of any kind re-reads every video once (about 10–30 minutes on a large archive; it can be stopped and picks up where it stopped). Then build the index, the stats, the homepage, the hub, and the sites.
- **A stats build keeps the stats of a channel whose drive is not mounted, and will not undo a newer version's stats.** A channel whose media is on a drive that is not mounted (or is being moved) is left as it was instead of being read as a channel with no videos; a stats rebuild that has to start over refuses until the drive is back. A stats build refuses to clear stats written by a newer version of the editor; set `ARCHILYZER_STATS_ALLOW_DOWNGRADE=1` to roll back on purpose. Its log also says apart how many videos were downloaded since the last index build (they catch up after the next one) and how many the index skipped (no upload date, or it failed on them).
+- **An index build keeps a channel whose drive is not mounted, instead of dropping it from the sites.** **Build index**, a site build's data phase and `archilyzer index` read a channel whose media is on a drive that is not mounted (or is being moved, or whose link and config disagree) as a channel with no videos: they removed its videos from the index, and the next site build published the channel as gone. Such a channel is now left as the last build had it — its videos stay in the index, its pages stay as they were, and the sites built next still list it — and the log names it, with its storage location: one line per channel, ` Held: N channel(s), K video(s) kept.` at the end of the `Diff:` line, and the channels again on the last line. A data folder that fails to read is held the same way, and a channel with no data folder at all is said in the log instead of passed over. An index rebuild that has to start over (after an update that changes the index's format, or with no index yet) refuses while any channel is held and says which; mount the drive first, or set `ARCHILYZER_INDEX_ALLOW_HELD=1` to rebuild without that channel until its drive is back and the index is built again.
- **Building the homepage now publishes the source: a read-only git mirror, its raw tree and a fresh tarball, behind a gate.** `archilyzer build homepage`, the `/sites` Homepage jobs and `pnpm ops build-homepage` run `archilyzer source publish` between compose and `next build`. It makes a fresh clone of the private `main` (the repository itself is never rewritten), rewrites that copy with git-filter-repo using your scrub rules (file contents and commit messages; your home directory becomes `/home/user` without a rule), and publishes it under `homepage/public` for `git clone https://archilyzer.pages.dev/source/archilyzer.git`, beside `/source/tree/` and the Downloads tarball. Before anything is written, every object of the rewritten history and every file about to be published is searched for every string you have denied; **one hit refuses the build**, and its log names the string only by where you wrote it (`denylist line 3 (len 5)`) and each hit by its object, field and byte offset — never a byte of the object. **A refusal withdraws the source**: the last publish is removed from `homepage/public` and the last build's copy from `homepage/out`, and **Deploy homepage refuses** a build whose source was not audited under today's rules and today's `main` ("run `archilyzer build homepage`, then deploy"). The rules live outside the repo, in `~/.config/archilyzer/source-scrub.txt` and `source-denylist.txt` (`ARCHILYZER_CONFIG_DIR`, `SOURCE_SCRUB_FILE`, `SOURCE_DENYLIST_FILE`); **without them the build refuses**, naming the missing file. **Put everything private in the denylist before any deploy, a preview included**: previews are public, and every deployment stays reachable at its own address until you delete it. Install git-filter-repo once (`pipx install git-filter-repo`; the editor's process needs `~/.local/bin` on its `PATH` to find it) — without it the build fetches it through `pipx run`, which needs the network — and gitleaks if you want its secret scan too. An unchanged `main` with unchanged rules is skipped, so a rebuild costs about 20 seconds only when something moved. A checkout with no git repository (the docker image, a tarball install) builds with the /source page's empty state. `archilyzer source publish --check` audits without writing, `archilyzer source audit <clone>/.git` checks any clone, `archilyzer build homepage --no-source` removes the published source instead, and `archilyzer doctor` reports the tools, the two files (rule counts and permissions, never their contents) and the last publish. `create-archives.sh` is gone. See PUBLISH.md, "The source mirror (homepage)".
- **umtool reads the corpus from its checkout (or `TRANSCRIPTS_DIR`), and the song project's data defaults to `~/.local/share/archilyzer/song`.** If yours is elsewhere, link it there before restarting umtool: `mkdir -p ~/.local/share/archilyzer && ln -s <where the data is> ~/.local/share/archilyzer/song` (the data stays where it is). With no `CHANNELS_DIR`, umtool reads the corpus at `$TRANSCRIPTS_DIR/channels`, else the checkout's own `transcripts/channels`; it used to fall back to an absolute path that existed on one machine only. The song project's videos default to `~/reports/quartering-uh-song/videos`; `SONG_DIR` and `VIDEO_ROOT` still win. The song project's tracked manifests record their paths relative to the song folders, and the twenty one-off `umtool/song/*.sh` run logs, which only ever ran on the machine that wrote them, are gone.
- **umtool's production build no longer reads the corpus folder.** Since umtool began finding the corpus from its checkout (the bullet above), `next build` treated the checkout's whole `transcripts/channels` as files to bundle. On a real archive it ran out of memory and was killed, so umtool could not be rebuilt. The build now ignores that folder and finishes in about 25 s at under 1 GB, the same as a checkout with no corpus. Nothing changes when umtool runs.
diff --git a/plans/FACTS.md b/plans/FACTS.md
@@ -7413,8 +7413,9 @@ on anchors elsewhere in this file:
index it:
- no `upload_date` (`buildIndex.ts:701`);
- a processing failure;
- - or the channel's media was unreachable during that build, since buildIndex has no drive
- guard.
+ - or the channel's media was unreachable during that build. Since release 15 (slice IG) that
+ happens only through a full rebuild run with `ARCHILYZER_INDEX_ALLOW_HELD=1`: an incremental
+ build keeps a held channel's records (see "The index build's hold").
It stays until fixed, and is logged as such rather than as pending (`:500`).
- **A transcript always has a date, and a caption video takes its captions' arrival.**
@@ -7446,10 +7447,9 @@ on anchors elsewhere in this file:
- This is the build's own guard. The job registry's `needsMedia` check
(`streamCommand.ts refuseForUnreachableMedia`) is per channel and needs a `channelSlug`, so it
never covered this pool-wide build, from the editor or from the CLI.
- - **buildIndex has no such guard.** An index build with a drive unmounted drops those channels'
- index records, and the site pages built from it lose them.
- - Its `Diff: … -R removed` line (`buildIndex.ts:641`) shows it.
- - A proper hold there is a follow-up slice (`stats-cache-key.md`, "Left").
+ - **buildIndex had no such guard until release 15.** An index build with a drive unmounted
+ dropped those channels' index records, and the site pages built from it lost them. It holds
+ them now: see "The index build's hold" below.
- **One stats build at a time: an operator rule, not a lock.**
- Two concurrent runs are harmless unless one clears the cache (a schema change) after the other
has scanned. The other then collects a partly refilled `statsByPath` and publishes truncated
@@ -7484,3 +7484,59 @@ on anchors elsewhere in this file:
finishes the rest.
- Run in the editor, it stalls the editor's event loop for the length of the pass. Prefer the CLI
with the editor idle.
+
+## The index build's hold (verified 2026-09-29, branch `r15/index-hold`)
+
+The record is [`release-15.md`](release-15.md), "Slice IG, as shipped". Anchors are at the branch
+tip. The branch added lines to `buildIndex.ts` from `:10` on, so every `buildIndex.ts` anchor above
+this section is stale, by +16 near the top and +197 at the end; they are not rewritten in place.
+
+- **An unmounted drive is not an empty channel, for the index either.**
+ - `scanSource` (`common/controller/buildIndex.ts:309`) calls `inspectChannelMedia({ channelsDir },
+ slug, cfg)` for every channel that is not excluded and not social (`:348`), before it reads
+ `data/`. A status other than `ok` or `in-place` (`isMediaHeld`,
+ `common/lib/channelMediaHold.ts:26`) puts the channel in `held` with a reason and no path, and
+ it is not scanned.
+ - It calls it again after the walk (`:462`), so a drive that goes away mid-walk holds the
+ channel instead of dropping the videos after that point.
+ - A failed `readdir(data/)` holds too (`its data directory could not be read (<code>)`), except
+ ENOENT on an `in-place` channel, which is a channel with no downloads and is logged as such
+ (`:362`). A per-video metadata `stat` failing with anything but ENOENT or ENOTDIR holds the
+ channel (`:387`).
+- **What a held channel keeps, on an incremental build:**
+ - its `mtimes` records: the removal pass skips its keys and counts them (`:709`), so `sums`,
+ `cues`, `subs`, `digests` and `byChannel` keep them too;
+ - its shared transcript, subs and digest trees: not rewritten, not pruned, not removed
+ (`:1198`, `:1387`, `:1766`, and the top-level cleanups `:1507`, `:1873`);
+ - its subs and digest stats, carried from `channelStatsDb` / `channelDigestStatsDb`, so the
+ per-site subs and digest manifests still list it;
+ - its availability states: the maybe-missing overlay skips it (`:1318`), and the last build's
+ `videoState` entries for it are carried over (`:1350`). Its `availability.json` files are on
+ the missing drive, and a missing one reads as `maybe_missing`.
+ - The per-site summaries come from LMDB, so the site built next still lists its videos.
+- **A curated-tag change while a channel is held:** the re-apply pass re-derives its records in
+ LMDB (no disk read), but its pages are not written, so `curatedPagesPending` is NOT cleared while
+ a channel is held and the pass had pages pending (`:1545`). The first build with the drive back
+ rewrites them.
+- **A full rebuild with a channel held refuses** (`:646`). A full rebuild is a schema change or a
+ first build (no `meta.schema`); it clears every sub-DB, and a held channel cannot be re-read.
+ - The scan now runs BEFORE the clear (`:636`), so the refusal leaves the index untouched.
+ `scanStartedAt` is still taken at the scan's start.
+ - The message names each channel with its location's label, the ways out (`HELD_WAYS_OUT`,
+ shared with the stats build's refusal), and `ARCHILYZER_INDEX_ALLOW_HELD` (`:509`, declared in
+ `envVars.ts`). The CLI exits 1 on it, so a site build's data phase fails with it.
+ - With the variable set, the build proceeds: the held channel's records go with the clear, its
+ shared trees are left on disk, and it is out of the index (the site lists it with 0 videos)
+ until its media is back and an index build runs.
+- **Who sees it:** `BuildIndexResult.heldChannels` (`:499`); the log's per-channel line, the
+ `Diff:` line's ` Held: N channel(s), K video(s) kept.` suffix (`:774`), and the `Done in` line's
+ ` Held, their media not readable: <slugs>.` suffix. The CLI (`archilyzer index`, the export's
+ and homepage's `build:index`, so every site build's data phase) prints the log; the editor's
+ **Build index** job (`buildIndexAction`) streams it into the job log, which `pnpm ops build-index
+ --wait` follows. Nothing reads the result's field outside the tests.
+- **The words are shared with the stats build** (`lib/channelMediaHold.ts`): `HELD_REASON` is a
+ `Record<ChannelMediaStatus, string>`, so a new status added to `inspectChannelMedia` fails tsc
+ until it has a reason; `isMediaHeld` treats any status but `ok` and `in-place` as held.
+- **Not covered:** a drive that drops during the PROCESSING phase (after the scan). A transcript
+ read that fails there is caught as "no cues" (`cueList = undefined`), so that video's cues are
+ removed until it next changes.
diff --git a/plans/STATE.md b/plans/STATE.md
@@ -20,13 +20,13 @@ holds the record, the review and the rollout. FACTS has "The stats cache key".
step 0: it runs the source publish.
- **Merge note:** `homepage/social-visible` merged `main` (`10cefd15`) at `4d11542c`; the one
conflict, `homepage/CHANGELOG.md`'s `[Unreleased]`, kept both sides.
-- **FOLLOW-UP, its own slice: the index build still treats an unmounted drive as an empty
- channel.** It drops that channel's index records, and the next site build publishes the channel
- as gone.
- - The fix: give `buildIndex` the stats build's hold, or at least a refusal with an override.
- - Schedule it before routine builds resume after this rollout.
- - Until then, the rollout's step 3 `Diff:` check is the safeguard: thousands removed means a
- drive was missing.
+- **CLOSED by release 15 slice IG (branch `r15/index-hold`, [`release-15.md`](release-15.md)):
+ the index build no longer treats an unmounted drive as an empty channel.** It holds the channel:
+ not rescanned, its index records and shared pages kept, and the `Diff:` line says
+ ` Held: N channel(s), K video(s) kept.` A full rebuild with a channel held refuses unless
+ `ARCHILYZER_INDEX_ALLOW_HELD=1`. FACTS has "The index build's hold". Until that branch is merged
+ and rolled out, the rollout's step 3 `Diff:` check stays the safeguard: thousands removed means
+ a drive was missing.
**Now (2026-09-28, evening): release 12 — the source mirror — is merged to `main` and NOT rolled
out.** [`release-12.md`](release-12.md) holds Q's and R's records, their reviews, "Merged" and
diff --git a/plans/release-15.md b/plans/release-15.md
@@ -16,7 +16,7 @@ prompt carries its ruling, and this record carries what was built. Rules:
| Slice | Branch | What | Owns |
|---|---|---|---|
-| IG | `r15/index-hold` | The index build holds an unreachable channel instead of emptying it | `common/controller/buildIndex.ts` + new `buildIndex.test.ts`, `common/controller/buildStats.ts` (the hold's words move to a shared module), new `common/lib/channelMediaHold.ts`, `common/bin/build-index.ts`, `common/lib/envVars.ts`, `ENVIRONMENT.md` |
+| IG | `r15/index-hold` | The index build holds an unreachable channel instead of emptying it | `common/controller/buildIndex.ts` + new `buildIndex.test.ts`, `common/controller/buildStats.ts` (the hold's words move to a shared module), new `common/lib/channelMediaHold.ts`, `common/lib/envVars.ts`, `ENVIRONMENT.md`; records: `plans/{STATE,FACTS,stats-cache-key}.md` |
| DS | `r15/drive-stall` | A stalled drive does not stop the editor answering | per its prompt |
| UT | `r15/umtool-trace` | umtool's build stops tracing the whole `umtool/` folder | per its prompt |
@@ -26,4 +26,162 @@ through its public signature. UT is independent. The shared files are `editor/CH
## Record
+### Slice IG, as shipped — the index build holds an unreachable channel (2026-09-29)
+
+Branch `r15/index-hold` off `main` `99d4d76a`, worktree `~/Projects/r12-paths-fix` (block #12:
+editor 4201, test 4211, export 4210), one Opus implementer. Scratch files `ig-*` in the job's
+`tmp`. The ruling: an index build meets a channel it cannot read the way the stats build already
+does. It holds the channel instead of reading it as empty, and a full rebuild with one held refuses
+unless `ARCHILYZER_INDEX_ALLOW_HELD=1`.
+
+**What was wrong.** `scanSource` read `data/` with a bare `catch { continue }`. A relocated
+channel whose drive was unmounted (a dangling `data/` link) therefore contributed no videos. The
+removal pass then dropped every record the channel had, the page writers rewrote its shared
+transcript tree empty and removed its subs tree, and the next site build published the channel as
+gone. Only the `Diff: … -R removed` line showed it.
+
+- **The hold** (`common/controller/buildIndex.ts`):
+ - `scanSource` asks `inspectChannelMedia({ channelsDir }, slug, cfg)` for every channel that is
+ not excluded and not social, before it reads `data/`.
+ - Any status but `ok` or `in-place` holds the channel, with a reason and its storage
+ location's label, and no path.
+ - It asks again after the walk, so a drive that goes away mid-walk holds the channel instead of
+ dropping the videos after that point.
+ - A `readdir(data/)` that fails holds the channel too: `its data directory could not be read
+ (<code>)`.
+ - The exception is ENOENT on an `in-place` channel: a channel with nothing downloaded (or its
+ media deleted), which is emptied as before. The log now says it: `Channel <slug>: no data/
+ directory; indexed as a channel with no videos.`
+ - A per-video metadata `stat` failing with anything but ENOENT or ENOTDIR holds the channel.
+ - **A held channel keeps everything:**
+ - its `mtimes` records, since the removal pass skips its keys, and so its `sums`, `cues`,
+ `subs`, `digests` and `byChannel` entries;
+ - its shared transcript, subs and digest trees: not rewritten, not pruned, and not removed by
+ the top-level cleanups;
+ - its subs and digest stats, carried from their sub-DBs, so the per-site manifests still list
+ it;
+ - its availability states. The maybe-missing overlay skips it, since its `availability.json`
+ files are on the missing drive and a missing one reads as "maybe missing". The last build's
+ `videoState` entries for it are carried over.
+ - The per-site summaries are built from LMDB, so the sites built next still list its videos.
+ - **A curated-tag change while held:** the re-apply pass re-derives the held channel's records
+ in LMDB as usual, but its pages are not written. The pages-pending flag is therefore kept (and
+ logged) while any channel is held, and the first build with the drive back writes them.
+- **A full rebuild refuses** (a schema change, or a first build with no index).
+ - The scan now runs BEFORE the clear, so a refusal leaves the index untouched. `scanStartedAt`
+ is still taken at the scan's start.
+ - The message names each channel with its location's label and says why a clear would publish
+ it as gone. It gives the ways out, mounting first (the same words as the stats build's
+ refusal), then the override by name.
+ - With `ARCHILYZER_INDEX_ALLOW_HELD=1` (1/true/yes/on, declared in `envVars.ts`, `ENVIRONMENT.md`
+ regenerated), the build proceeds. The held channel's records go with the clear; they cannot be
+ carried across a format change. Its shared trees are left on disk, and it is out of the index
+ until its media is back and an index build runs.
+- **The words are shared:** new `common/lib/channelMediaHold.ts` (`isMediaHeld`, `HELD_REASON`,
+ `heldReason`, `describeHeld`, `HELD_WAYS_OUT`). `buildStats.ts` uses it, and its messages are
+ byte-identical (its case (i) passes unchanged).
+- **The result and the log:** `BuildIndexResult.heldChannels`. The log carries one line per held
+ channel (`Channel <slug>: <why>; its N indexed video(s) are kept as they are, not rescanned, and
+ its transcript, subtitle and digest pages are left as they are.`). The `Diff:` line ends in
+ ` Held: N channel(s), K video(s) kept.` and the `Done in` line in ` Held, their media not
+ readable: <slugs>.` Both prefixes are unchanged: the e2e helper waits for `Done`.
+
+**Who reports it** (every caller of `buildIndex`):
+
+| Caller | What it shows |
+|---|---|
+| `pnpm archilyzer index` (also the root, export and homepage `build:index` scripts) | The log on stdout, with the three lines above. It exits 0 on a hold. On the refusal it exits 1 with the message on stderr (`runIfEntryPoint`). `common/bin/build-index.ts` is unchanged: it prints the log and discards the result. |
+| A site build's data phase (`build:data`, from `buildSite`'s steps and `buildAll`'s Phase A in `common/publish/build.ts`) | The same lines, in the site build's job log. A refusal stops the steps at the data phase (`Build failed (exit 1)` / `Data phase failed (exit 1)`), so nothing is composed or deployed from an emptied index. The hub build does not run the index. |
+| The editor's **Build index** job (`buildIndexAction`, `editor/app/sites/lib/buildAction.ts`) | The lines in the job's log on `/sites` and `/jobs`. A refusal fails the job with the message. The action discards the result, and it was left unchanged: the log already carries every held channel, and `editor/app/sites/**` belongs to another slice this release. |
+| `pnpm ops build-index --wait` | Follows the job's log, so it prints the same lines. |
+| The tests | `heldChannels`. |
+
+**Commits**
+
+| Commit | What |
+|---|---|
+| `c6ae51b0` | `plans:` this record: the header, the slices, and empty Record and Rollout sections. |
+| `51328098` | `common:` the hold's words move to `lib/channelMediaHold.ts`; the stats build uses them, with its messages unchanged. |
+| `ffb01d8e` | `common:` the index build's hold, the refusal and its override, `heldChannels`, the log lines; `envVars.ts` + `ENVIRONMENT.md`; new `buildIndex.test.ts` (9 cases). |
+| this commit | `plans:` this section; FACTS "The index build's hold" (and the two stale statements in "The stats cache key" corrected); the STATE follow-up closed; `stats-cache-key.md` "Left" marked closed; the editor changelog. |
+
+**Tests** (`common/controller/buildIndex.test.ts`, the real `buildIndex` over a temp corpus). The
+drive channel is seeded with `buildStats.test.ts`'s `seedDriveChannel` shape and unmounted by
+renaming its media root away. Every path is pinned under a temp root, and case (z) spies on
+node:fs writes.
+
+| Case | What it pins | On the pre-change `buildIndex.ts` |
+|---|---|---|
+| (a) | Unmounted, incremental build with another change: records kept, the drive's transcript and subs trees byte-identical (manifest included), the site still lists both videos and its subs count; the log lines, with the label and no path | removed 2, not 0 |
+| (b) | Availability carried: `maybe_missing` stays and a post-scan confirmation stays `available`; `videoState` unchanged | d1 no longer published |
+| (c) | A full rebuild refuses (message, ways out, the variable, no path); the index and pages untouched; the CLI exits non-zero; with the override it holds (records cleared, pages kept); the drive back re-adds both | no refusal |
+| (d) | A first build (no index) with a channel held refuses too | no refusal |
+| (e) | The drive back: held set empty, a video added to the drive meanwhile indexed, 0 changed | removed 2 while away |
+| (f) | A really empty in-place channel (an empty `data/`, and no `data/`) is emptied, not held; the missing `data/` is logged | the new log line only (the emptying matched, as it should) |
+| (g) | An unreadable `data/`, and an unreadable video dir mid-walk (mode 000), hold their channels with `EACCES` | removed 3 |
+| (h) | A tag rule added while held: the held pages untouched, the flag kept; the drive back writes the tag to them and clears the flag | the drive's pages emptied |
+| (z) | No write outside the temp root | passes on both |
+
+The pre-change column was run with the old `buildIndex.ts` swapped in once, with the
+`heldChannels` assertions removed so each case reached its first substantive assertion.
+
+#### Gates (at `ffb01d8e`; logs `$T/ig-*.log`)
+
+- **tsc** was clean before every commit: 78 s at the branch point, 48 s at `ffb01d8e`.
+- **Unit:**
+
+ | Suite | Result |
+ |---|---|
+ | common | 2,219/2,219 (the branch point's 2,210 plus the 9 new cases), 77 s |
+ | editor unit | 87/87 |
+ | `test:scripts` | 191 passed, 1 skipped (192) |
+ | mcp | 271/271 |
+
+- **Docs:** `docs env --check`, `docs files --check` and `settings example --check` all exit **0**.
+- **Build:** the editor's `next build`, with the primary's `transcripts/` linked in and capped at
+ 5 GB with no swap: 64 s, max RSS 1,642 MB. The link was removed after the build, and nothing
+ ran through it.
+- **e2e** (editor, detached and queued; the spec list is every spec that runs Build index:
+ `availability`, `build`, `channel-build-toggle`, `chat-only`, `duplicate-shorts`, `jobs`,
+ `regional-vtt-fallback`, `tags`, passed as `e2e/<name>.spec.ts` so `availability` does not also
+ match `pre-clean-availability`): **35 passed, 0 failed, 3.6 min**, after 1 min 45 s in the
+ queue. No e2e fixture has an unreachable channel, so these confirm the hold changes nothing for a
+ readable corpus.
+- **Numbers tool:** none.
+
+#### Found and left
+
+- **A drive that drops during the processing phase** (after the scan) is not covered. A transcript
+ read that fails there is caught as "no cues", so that video's cues are removed until its
+ transcript next changes. That is a per-video read inside the worker, and it is left to the slice
+ that handles drive stalls.
+- **`inspectChannelMedia` is asked twice per channel** (before and after the walk), three syscalls
+ each today. Slice DS puts a health gate inside it, and that gate runs twice per channel per
+ index build.
+- **Under the override,** a held channel is listed on its sites with 0 videos, and its subs and
+ digest counts are left out of the site manifests. Its old shared page trees stay on disk, and a
+ compose copies them.
+- **While a channel is held after a tag change,** the pages-pending flag stays set. Every build
+ until the drive is back is then a full page walk; it skips unchanged pages by hash, so it rewrites
+ none.
+- **The digest tree's retention has no test.** The code path mirrors the subs tree's, and no
+ fixture carries a digest sidecar.
+- **The editor shows a hold only in the job log.** A `/sites` badge would be in `editor/app/sites/**`.
+
+#### Decisions the operator could overturn
+
+| What I assumed | The alternative |
+|---|---|
+| A held channel's shared pages are not written at all. | Rewrite them from the kept records: byte-identical on an incremental build, and a tag change would reach them at once. It would still need the skip after an override, when the records are gone. |
+| Under the override, the held channel's records go with the clear. | Carry them across the clear. A schema change means the stored format moved, so the old records cannot be trusted. |
+| An unreadable `data/`, or a per-video `stat` failing with anything but ENOENT, holds the channel. | Hold only for the statuses `inspectChannelMedia` reports, and keep treating other read errors as "no videos", as the old code did. |
+| A channel with no `data/` is logged, one line per build. | Stay silent, as before; the ruling asked for a log. |
+| The CLI exits 0 on a hold, since the hold is the safe outcome. | Exit non-zero so scripts notice; a site build's data phase would then fail whenever a drive is out. |
+| The words live in a new `lib/channelMediaHold.ts` shared with the stats build. | Duplicate them in `buildIndex.ts` and leave `buildStats.ts` untouched. |
+
+**What runs which code, for the rollout.** Every CLI command and every spawned data phase runs the
+checkout's code, so they hold from the moment `main` has this branch. The editor's in-process
+**Build index** button runs its built bundle, so it holds only after the editor is rebuilt and
+restarted.
+
## Rollout
diff --git a/plans/stats-cache-key.md b/plans/stats-cache-key.md
@@ -273,6 +273,8 @@ PATH, so it is `pnpm archilyzer …`.
`mtimes`, cues and pages), or at least a refusal with an override.
- When: schedule it before routine builds resume after this rollout.
- Until then, the rollout's step 3 `Diff:` check is the safeguard.
+ - **Closed by release 15 slice IG** ([`release-15.md`](release-15.md), "Slice IG, as shipped"):
+ the hold, and a refusal with an override for a full rebuild.
- **A cross-process lock for builds** (see "Concurrent stats builds" above). The rule stands in
for it.
- **O5:** a manual English caption (no inline timing tags) indexes as 0 cues (FACTS).