commit 82f159af8b73f615ed2294339010fda9de9b6995
parent 1eb37bde03bab3949985c10cf14e6cf7047e91ea
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Fri, 2 Oct 2026 02:30:01 -0400
Merge r17/media-tier-migrate (release 17 slice T3) — archilyzer storage migrate-tier <slug>|--all: the one-off migration of a whole-data/ channel to the media tier (editor stopped; copy of text, clips and scratch to the SSD from a classifier list, verified twice; relative links carrying the file's mtime; the platter's data renamed media; resumable from each step; --all smallest first with a stop before the three big-text channels; --reclaim later, guarded by a note and a twin check); AGENTS.md's seventh non-optional thing; FACTS, CHANNEL.md, ENVIRONMENT.md; reviewed SHIP
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Diffstat:
13 files changed, 2210 insertions(+), 25 deletions(-)
diff --git a/AGENTS.md b/AGENTS.md
@@ -228,7 +228,7 @@ apps*, not [PUBLISH.md](PUBLISH.md), which is about *building and publishing sit
| Path | What it is |
|---|---|
-| `transcripts/channels/<slug>/` | One channel: `config.json` (every key in [CHANNEL.md](CHANNEL.md)), `playlist`, `archive`, `snapshot.json`, and `data/<videoId>/` holding media, transcripts and sidecars. **`data/` may be an absolute SYMLINK** — see below. |
+| `transcripts/channels/<slug>/` | One channel: `config.json` (every key in [CHANNEL.md](CHANNEL.md)), `playlist`, `archive`, `snapshot.json`, and `data/<videoId>/` holding transcripts, sidecars and media. **A big file in `data/<id>/` may be a relative link into `media/`, which may be an absolute SYMLINK to another drive** (and a legacy channel's whole `data/` is one, until migrated) — see below. |
| `transcripts/sites/<id>/site.json` | **Per-site config, including the public URL** (every key in [SITE.md](SITE.md)). This is where deployed-site facts live — *not* under `channels/`. |
| `transcripts/index.mdb` | The LMDB transcript index. Key-only range scans over its `byChannel` sub-DB are cheap; see `common/controller/recencyIndex.ts`. |
| `transcripts/saved-videos/` | Persisted source-video store. |
@@ -263,32 +263,46 @@ changed by `patchChannelConfig` (`common/controller/channels.ts`), never by spre
config read earlier; a per-video sidecar is declared once with `sidecar()`
(`common/lib/sidecar-server.ts`), which refuses a `transcript.<x>.<y>` name.
-## A channel's `data/` may live on another drive
-
-`channels/<slug>/data` can be an **absolute symlink** to `<root>/<slug>/data` on
-another disk, with `config.dataDir` recording the target. The editor's Storage panel
-(channel page → Storage) moves it — to one of the locations configured on `/storage`,
-or to a root typed by hand; nothing else writes that field. **A channel is not tagged
-with its location**: it is on location L iff its `config.dataDir` is under `L.root`,
-which is why re-pointing a location rewrites only links and `dataDir`. The on-disk contract
-`channelDir/data/<id>/…` is unchanged, so **no reader needs to know** — yt-dlp's
-cwd-relative writes, the LMDB index (it stores mtimes, and `rsync -a` preserves them)
-and the export build all keep working with no call-site changes.
-
-Six things that are not optional:
+## A channel's media may live on another drive
+
+A channel's text — transcripts, cues, `metadata.info.json`, every sidecar, `clips/` — is
+always in a REAL `channels/<slug>/data/` on the corpus disk. Only its big files move
+(release 17, the media tier; `common/lib/mediaTier.ts` says which, by name): each becomes a
+RELATIVE link `data/<id>/<name> -> ../../media/<id>/<name>`, and `channels/<slug>/media` is a
+real directory on the corpus disk, or ONE absolute symlink to `<root>/<slug>/media` on
+another disk with `config.mediaDir` recording the target, or absent (a classic channel, its
+big files still real in `data/<id>/`). The editor's Storage panel (channel page → Storage)
+moves `media/` — to one of the locations configured on `/storage`, or to a root typed by
+hand; nothing else writes `mediaDir` but the re-point, the rename and the one-off
+migration. **A channel is not tagged with its location**: it is on location L iff its
+`config.mediaDir` is under `L.root`, which is why re-pointing a location rewrites only the
+`media` links and `mediaDir`. The on-disk contract `channelDir/data/<id>/…` is unchanged,
+so **no reader needs to know** — yt-dlp's cwd-relative writes, the LMDB index (it stores
+mtimes; it `lstat`s a tierable name, and a tier link carries its file's times) and the
+export build all keep working with no call-site changes. **`config.dataDir` is RETIRED**:
+a channel that still carries it, or whose `data/` is itself a link (the pre-release-17
+whole-directory move), is `legacy` — held by every guard, text included — until
+`archilyzer storage migrate-tier <slug>` (editor stopped) brings its text home.
+
+Seven things that are not optional:
- **Never symlink a whole channel dir.** Channel listing filters `isDirectory()` on
`channelsDir` entries (`channels.ts:246,363`), so a symlinked `<slug>/` vanishes from
- the corpus. Only `data/` may be a link.
+ the corpus. Only `media` may be a link (and a legacy channel's `data/`, until it is
+ migrated).
- **An unmounted drive is not an empty channel.** Every enumerator swallows ENOENT on
`data/` as "no videos", which to a runner means *everything is undownloaded*.
`common/lib/channelMedia.ts` is the one module that can tell the two apart;
- `inspectChannelMedia` / `assertChannelMediaReachable` are what the guards call, and a
- job kind declares `needsMedia` in `common/jobs/jobKinds.ts` to be covered by the one
- in `runManagedFunction`. If you add a path that reads `data/`, guard it there.
+ `inspectChannelMedia` / `assertChannelMediaReachable` (a job that opens a big file) /
+ `assertChannelTextReadable` (a reader of the text) are what the guards call, and a job
+ kind declares `needsMedia` or `needsText` in `common/jobs/jobKinds.ts` to be covered by
+ the one in `runManagedFunction`. If you add a path that reads `data/`, guard it there.
- **`channels/<slug>/.relocating.json`** is the in-flight marker. Its presence means
"media is in transition" to every guard and lets an interrupted move resume from its
- `phase`. `deleteChannel` and `renameChannel` refuse while it exists. The saved-video
+ `phase`. Its `scope` says what is moving: `"media"` (the Storage panel's move — the
+ text stays readable, only media writers are held) or `"tier-migration"` (the one-off
+ migration, which rebuilds `data/` — the text is held too, and only `migrate-tier`
+ resumes or clears it). `deleteChannel` and `renameChannel` refuse while it exists. The saved-video
store has its own, one level up: **`transcripts/.relocating-saved-videos.json`**, the
same `{target, direction, startedAt, phase}` shape. Its reader and
`assertSavedVideosStoreWritable` live in `common/lib/savedVideoStore.ts` — in *lib*
@@ -307,10 +321,14 @@ Six things that are not optional:
(`controller/relocateChannelMedia.ts`) stats the root and, when it belongs to a location
carrying a `volume.uuid`, requires the probe's identity to match. It runs from the
preview AND immediately before each copy phase's mkdir.
-- **`data/<id>/clips/` is a cache, and `evictClipWindows` is the only thing that prunes
- it — BY AGE.** Nothing in the editor can know whether a umtool report still cites a
+- **`data/<id>/clips/` is a cache, on the SSD, never tiered, and `evictClipWindows` is
+ the only thing that prunes it — BY AGE.** Nothing in the editor can know whether a umtool report still cites a
window (the manifests are in a umtool project), so there is no reference count and every
surface says so. An evicted window is re-fetchable: the cost is a fetch, not data.
+- **A media file in `data/<id>/` may be a relative symlink into `channels/<slug>/media/`.
+ Remove one with `removeMediaFile`, never `rm`/`remove`; tier one with the hook after
+ every media finalisation, never by hand; a dirent `isFile()` filter over a video dir
+ hides it.**
## The live instances
diff --git a/CHANNEL.md b/CHANNEL.md
@@ -27,7 +27,7 @@ Regenerate this file with `pnpm --filter yt-dlp-transcript-common exec tsx bin/f
| `keepLatest` | config | Keep-latest window: the newest N videos (by upload date) are protected from the Clean-audio sweep AND have their source video persisted to the saved-video store. 0 or absent = disabled; positives clamp to [1, 100000]. A kept video later found deleted at the source is pinned permanently via the do-not-clean marker. |
| `extractionMode` | config | `"ytdlp"` (default — yt-dlp's own `-x --audio-format` postprocessor, no source container kept) or `"app"` (yt-dlp downloads the source container and the app runs ffmpeg). The keep-latest persistence rule forces `"app"` for the videos it persists. |
| `savedVideosDir` | config | Per-channel override for the saved-video store root: this channel's persisted source videos live under `<savedVideosDir>/<slug>/<videoId>/`. Trimmed; blank = the global store. |
-| `dataDir` | config | RETIRED (release 17). The whole-directory layout's record: the absolute path `channels/<slug>/data` was a symlink to. Still parsed for one release so a write never erases it: a channel that carries it — or whose `data/` is a link — is `legacy`, and every media job, lane and build holds it until `archilyzer storage migrate-tier <slug>` moves its text back and its media into `mediaDir`. Never written by anything but that migration, which removes it. |
+| `dataDir` | config | RETIRED (release 17). The whole-directory layout's record: the absolute path `channels/<slug>/data` was a symlink to. Still parsed for one release so a write never erases it: a channel that carries it — or whose `data/` is a link — is `legacy`, and every media job, lane and build holds it until `archilyzer storage migrate-tier <slug>` moves its text back and its media into `mediaDir`. Written only by the re-point of a storage location (which keeps it `<root>/<slug>/data` on the new root); removed by that migration. |
| `mediaDir` | config | Where this channel's big files live when relocated: `channels/<slug>/media` is a symlink to it, `<root>/<slug>/media`. Absent = in place. Written only by relocate / re-point / the tier migration — a record of what is on disk, never free text, because a value that disagrees with the link is an "inconsistent" channel every media guard refuses. The text (`data/`) never moves. |
| `ytdlpExtraArgs` | config | Extra yt-dlp arguments, appended verbatim. Must be an array of strings or it is dropped. |
| `subLangs` | config | yt-dlp `--sub-langs` value for caption downloads. |
diff --git a/ENVIRONMENT.md b/ENVIRONMENT.md
@@ -82,7 +82,7 @@ Tokens, credentials and knobs a running process reads. Most configuration is not
| `ARCHILYZER_INDEX_ALLOW_HELD` | off | `1` lets a FULL index rebuild (a schema change, or no index yet) proceed while a channel's media cannot be read; that channel stays out of the index until its media is back and the index is built again. Unset, such a build refuses and names each channel. | common/controller/buildIndex.ts |
| `UV_THREADPOOL_SIZE` | `16` for the editor (`4` is Node's own) | Threads in Node's pool for filesystem calls. A call on a stalled drive holds one until the drive answers, so the editor starts with 16. It buys time for calls already in flight and isolates nothing: the storage health probe and its gate keep new calls off a stalled drive. | Node's libuv (set by editor/package.json `start` and docker/entrypoint.sh) |
| `MCP_IO_STATS` | off | `1` turns on per-call I/O accounting, for `mcp/bench`. | common/lib/archive/io-stats.ts |
-| `ARCHILYZER_EDITOR_URL` | `http://localhost:3001` | Which editor `pnpm ops` and the MCP's `fetch_clip` talk to. | scripts/archilyzer-ops.mjs, mcp/src/fetchClip.ts, umtool |
+| `ARCHILYZER_EDITOR_URL` | `http://localhost:3001` | Which editor `pnpm ops` and the MCP's `fetch_clip` talk to, and whose `/api/pulse` `archilyzer storage migrate-tier` asks before it refuses to run beside it. | scripts/archilyzer-ops.mjs, mcp/src/fetchClip.ts, umtool, common/bin/migrate-media-tier.ts |
| `ARCHILYZER_AGENT` | `cli` | Who is asking, recorded as the provenance of a curated-tag write through `pnpm ops`. | scripts/archilyzer-ops.mjs |
| `DIARIZE_ENGINE_KIND` | `sherpa-onnx` | The diarization engine: `sherpa-onnx` or `sortformer`. | scripts/diarize.mjs |
| `DIARIZE_ENGINE_CMD` | the bundled sherpa script | The engine command the wrapper runs. | scripts/diarize.mjs |
diff --git a/common/bin/archilyzer.ts b/common/bin/archilyzer.ts
@@ -267,6 +267,8 @@ export const COMMANDS: Command[] = [
"--channel <slug> list duplicate and missing transcripts"),
script(["migrate", "channel-priority"], "migrate-channel-priority.ts",
"[--dry-run] the one-shot channel-priority migration (plans/channel-priority.md, S5)"),
+ script(["storage", "migrate-tier"], "migrate-media-tier.ts",
+ "<slug>…|--all [--order smallest] [--include-large] [--dry-run] [--reclaim] bring a channel off the retired whole-directory layout onto the media tier, its text home to the corpus disk (editor stopped; --all stops before the three big-text channels)"),
{
path: ["brand", "media"],
usage: "[--out <dir>] [--video-kit] render the Archilyzer Media channel's assets",
diff --git a/common/bin/migrate-media-tier.test.ts b/common/bin/migrate-media-tier.test.ts
@@ -0,0 +1,652 @@
+import { after, test } from "node:test";
+import assert from "node:assert/strict";
+import {
+ lstatSync,
+ mkdirSync,
+ mkdtempSync,
+ readdirSync,
+ readFileSync,
+ readlinkSync,
+ renameSync,
+ rmSync,
+ statSync,
+ symlinkSync,
+ utimesSync,
+ writeFileSync,
+} from "node:fs";
+import { tmpdir } from "node:os";
+import path from "node:path";
+
+// Run with:
+// pnpm --filter yt-dlp-transcript-common exec tsx --test bin/migrate-media-tier.test.ts
+//
+// The one-off media-tier migration over a fixture: a "platter" dir beside a
+// corpus, a legacy channel (its `data/` an absolute link to
+// `<platter>/<slug>/data`, `config.dataDir`) holding text, audio, a raw live
+// chat, a clip window, a source container and scratch. The REAL rsync runs.
+// The environment is pinned before any module reads it (getPaths caches), as
+// buildIndex.test.ts does, so the index case can build over the same corpus.
+
+const ROOT = mkdtempSync(path.join(tmpdir(), "migrate-tier-"));
+const PINNED: Record<string, string> = {
+ TRANSCRIPTS_DIR: path.join(ROOT, "transcripts"),
+ SAVED_VIDEOS_DIR: path.join(ROOT, "saved-videos"),
+ SITES_DIR: path.join(ROOT, "transcripts", "sites"),
+ SETTINGS_FILE: path.join(ROOT, "settings.json"),
+ EXPORT_PUBLIC_DIR: path.join(ROOT, "public"),
+ EXPORT_INDEX_DIR: path.join(ROOT, ".export-index"),
+ EXPORT_BUILDS_DIR: path.join(ROOT, ".export-builds"),
+ EDITOR_CHANGELOG_FILE: path.join(ROOT, "editor-CHANGELOG.md"),
+ EXPORT_CHANGELOG_FILE: path.join(ROOT, "export-CHANGELOG.md"),
+ CHARTS_CONFIG_FILE: path.join(ROOT, "chart-templates.json"),
+ SEARCH_ALIASES_FILE: path.join(ROOT, "transcripts", "search-aliases.json"),
+ CURATED_TAGS_FILE: path.join(ROOT, "transcripts", "tags.json"),
+ ARCHILYZER_CONFIG_DIR: path.join(ROOT, "config"),
+ ARCHILYZER_SOURCE_SCRATCH: path.join(ROOT, "source-scratch"),
+};
+Object.assign(process.env, PINNED);
+delete process.env.ARCHILYZER_INDEX_ALLOW_HELD;
+after(() => rmSync(ROOT, { recursive: true, force: true }));
+
+const { getPaths } = await import("../lib/paths");
+const {
+ editorRunningReason,
+ legacyChannels,
+ migrateMediaTier,
+ INCOMING_NAME,
+ LIST_NAME,
+ RECLAIM_NOTE,
+} = await import("./migrate-media-tier");
+const { inspectChannelMedia, readRelocationMarker } = await import("../lib/channelMedia");
+const { classifyEntry } = await import("../lib/mediaTier");
+const { readChannelConfig } = await import("../controller/channels");
+const { buildIndex } = await import("../controller/buildIndex");
+const { isLiveChatCuesFresh, normalizeLiveChat } = await import(
+ "../controller/normalizeLiveChat"
+);
+
+type MigrateOpts = Parameters<typeof migrateMediaTier>[0];
+type MigrationStep = NonNullable<
+ NonNullable<MigrateOpts["deps"]>["checkpoint"]
+> extends (slug: string, step: infer S) => unknown
+ ? S
+ : never;
+
+const paths = getPaths();
+const PLATTER = path.join(ROOT, "platter");
+const SLUG = "legacy-ch";
+const SITE = "testsite";
+
+const OLD = new Date("2021-03-04T05:06:07.000Z");
+const CUES_AT = new Date("2021-03-05T05:06:07.000Z");
+
+// Settings with the disk gate off, so the tmpfs's free space is not the
+// subject (relocateChannelMedia.test.ts says why at length); the free-space
+// cases inject `freeBytes` and turn the floor on.
+const SETTINGS = {
+ minFreeDiskGB: 0,
+ resumeMarginGB: 2,
+ storage: { locations: [], defaultLocationId: "" },
+} as unknown as MigrateOpts["settings"];
+
+const CHAT_LINE =
+ JSON.stringify({
+ replayChatItemAction: {
+ actions: [
+ {
+ addChatItemAction: {
+ item: {
+ liveChatTextMessageRenderer: {
+ message: { runs: [{ text: "hello" }] },
+ authorName: { simpleText: "a" },
+ timestampUsec: "1000000",
+ },
+ },
+ },
+ },
+ ],
+ videoOffsetTimeMsec: "1000",
+ },
+ }) + "\n";
+
+const VTT =
+ "WEBVTT\nKind: captions\nLanguage: en\n\n" +
+ "00:00:00.000 --> 00:00:05.000\nFirst caption line.\n\n" +
+ "00:01:00.000 --> 00:01:50.000\nSecond caption line.\n";
+
+// One video dir's files: name -> content. `clips/` is a directory.
+const VIDEOS: Record<string, Record<string, string>> = {
+ v1: {
+ "transcript.en.vtt": VTT,
+ "audio.mp3": "A".repeat(4096),
+ "transcript.live_chat.json": CHAT_LINE.repeat(20),
+ "source-media.mp4": "S".repeat(2048),
+ "audio.m4a.part": "P".repeat(512),
+ "audio.tmp-1234.mp3": "T".repeat(256),
+ "clips/w-0001.mp4": "C".repeat(1024),
+ // A postprocessor's dead temp: left on the platter, never copied.
+ "source-media.temp.mp4": "D".repeat(1536),
+ },
+ v2: {
+ "transcript.json": JSON.stringify({ segments: [{ start: 0, end: 1, text: "hi there" }] }),
+ "audio.opus": "O".repeat(3000),
+ // Syncthing's unfinished transfer: scratch, carried.
+ ".syncthing.audio.mp3.tmp": "Y".repeat(128),
+ },
+ v3: {},
+};
+
+function writeJson(file: string, value: unknown): void {
+ mkdirSync(path.dirname(file), { recursive: true });
+ writeFileSync(file, JSON.stringify(value, null, 2));
+}
+
+function channelDir(slug = SLUG): string {
+ return path.join(paths.channelsDir, slug);
+}
+function oldTree(slug = SLUG): string {
+ return path.join(PLATTER, slug, "data");
+}
+function mediaTarget(slug = SLUG): string {
+ return path.join(PLATTER, slug, "media");
+}
+
+function writeConfig(slug: string, extra: Record<string, unknown> = {}): void {
+ writeJson(path.join(channelDir(slug), "config.json"), {
+ handling: "youtube",
+ name: slug,
+ url: `https://www.youtube.com/@${slug}/videos`,
+ ...extra,
+ });
+}
+
+function resetCorpus(channels: string[] = [SLUG]): void {
+ for (const p of [paths.transcriptsDir, PINNED.EXPORT_INDEX_DIR, PLATTER]) {
+ rmSync(p, { recursive: true, force: true });
+ }
+ mkdirSync(paths.channelsDir, { recursive: true });
+ mkdirSync(PLATTER, { recursive: true });
+ writeFileSync(paths.settingsFile, JSON.stringify({}));
+ writeJson(path.join(paths.sitesDir, SITE, "site.json"), {
+ siteId: SITE,
+ siteTitle: "Test Site",
+ siteDescription: "fixture",
+ headerTitle: "Test Site",
+ homeTagline: "",
+ socialLinks: [],
+ groups: [{ id: "default", name: "All channels", selectedByDefault: true }],
+ defaultGroupId: "default",
+ channels: channels.map((slug) => ({ slug, groupId: "default" })),
+ });
+}
+
+// Write the videos into `dataDir` (every file at OLD, the cues a day later).
+function writeVideos(dataDir: string, slug: string, scale = 1): void {
+ for (const [id, files] of Object.entries(VIDEOS)) {
+ const dir = path.join(dataDir, id);
+ mkdirSync(dir, { recursive: true });
+ writeJson(path.join(dir, "metadata.info.json"), {
+ id,
+ title: `Video ${id}`,
+ channel: slug,
+ upload_date: "20260601",
+ duration: 120,
+ description: "fixture",
+ webpage_url: `https://www.youtube.com/watch?v=${id}`,
+ extractor_key: "Youtube",
+ });
+ utimesSync(path.join(dir, "metadata.info.json"), OLD, OLD);
+ for (const [name, content] of Object.entries(files)) {
+ const file = path.join(dir, name);
+ mkdirSync(path.dirname(file), { recursive: true });
+ writeFileSync(file, content.repeat(scale));
+ utimesSync(file, OLD, OLD);
+ }
+ }
+}
+
+// A channel on the RETIRED layout, as the pre-release-17 mover left it.
+function seedLegacy(slug = SLUG, scale = 1): void {
+ writeVideos(oldTree(slug), slug, scale);
+ writeConfig(slug, { dataDir: oldTree(slug) });
+ symlinkSync(oldTree(slug), path.join(channelDir(slug), "data"));
+}
+
+// What is on disk under ROOT's corpus and platter: type, size, mtime, link
+// target, content of small files. Equal snapshots = nothing was written.
+function snapshot(): Record<string, string> {
+ const out: Record<string, string> = {};
+ const walk = (p: string) => {
+ const st = lstatSync(p);
+ const rel = path.relative(ROOT, p);
+ if (st.isSymbolicLink()) {
+ out[rel] = `link ${readlinkSync(p)} ${st.mtimeMs}`;
+ return;
+ }
+ if (st.isDirectory()) {
+ out[rel] = `dir`;
+ for (const n of readdirSync(p)) walk(path.join(p, n));
+ return;
+ }
+ out[rel] = `file ${st.size} ${st.mtimeMs} ${readFileSync(p, "utf8").slice(0, 64)}`;
+ };
+ walk(paths.transcriptsDir);
+ walk(PLATTER);
+ return out;
+}
+
+function run(over: Partial<MigrateOpts> = {}, lines: string[] = []) {
+ return migrateMediaTier({
+ paths,
+ settings: SETTINGS,
+ slugs: [SLUG],
+ log: (l) => lines.push(l),
+ ...over,
+ deps: {
+ editorProbe: async () => null,
+ ...over.deps,
+ },
+ });
+}
+
+// THE MIGRATED STATE, every rule the slice names.
+const DEAD_TEMP = "source-media.temp.mp4";
+
+async function assertMigrated(
+ slug = SLUG,
+ scale = 1,
+ opts: { reclaimed?: boolean } = {},
+): Promise<void> {
+ const ch = channelDir(slug);
+ const config = await readChannelConfig(paths, slug);
+ assert.equal(config?.mediaDir, mediaTarget(slug));
+ assert.equal(config?.dataDir, undefined, "dataDir unset");
+ assert.ok(lstatSync(path.join(ch, "data")).isDirectory(), "data/ is a real dir");
+ assert.equal(readlinkSync(path.join(ch, "media")), mediaTarget(slug));
+ assert.throws(() => lstatSync(oldTree(slug)), "the old tree is renamed");
+ assert.throws(() => lstatSync(path.join(ch, INCOMING_NAME)));
+ assert.throws(() => lstatSync(path.join(ch, LIST_NAME)));
+ assert.equal(await readRelocationMarker(paths, slug), null, "no marker");
+
+ for (const [id, files] of Object.entries(VIDEOS)) {
+ const dir = path.join(ch, "data", id);
+ const meta = lstatSync(path.join(dir, "metadata.info.json"));
+ assert.ok(meta.isFile());
+ assert.equal(meta.mtimeMs, OLD.getTime(), `${id} metadata keeps its mtime`);
+ for (const name of Object.keys(files)) {
+ const top = name.split("/")[0];
+ const p = path.join(dir, name);
+ if (name === DEAD_TEMP) {
+ assert.throws(() => lstatSync(p), "a dead temp is not carried to the corpus disk");
+ const onPlatter = path.join(mediaTarget(slug), id, name);
+ if (opts.reclaimed) assert.throws(() => lstatSync(onPlatter), "reclaimed");
+ else assert.ok(lstatSync(onPlatter).isFile(), "left on the platter, unlinked");
+ continue;
+ }
+ const st = lstatSync(p);
+ if (name === "audio.mp3" || name === "audio.opus" || name === "transcript.live_chat.json") {
+ assert.ok(st.isSymbolicLink(), `${id}/${name} is a link`);
+ assert.equal(readlinkSync(p), `../../media/${id}/${name}`, "relative");
+ assert.equal(st.mtimeMs, OLD.getTime(), `${id}/${name}: the link carries the file's mtime`);
+ assert.equal(readFileSync(p, "utf8"), files[name].repeat(scale), "readable through the link");
+ assert.ok(lstatSync(path.join(mediaTarget(slug), id, name)).isFile(), "bytes on the platter");
+ } else {
+ assert.ok(st.isFile(), `${id}/${name} stays real (${top})`);
+ assert.equal(st.mtimeMs, OLD.getTime(), `${id}/${name} keeps its mtime`);
+ assert.equal(readFileSync(p, "utf8"), files[name].repeat(scale));
+ }
+ }
+ }
+ const loc = await inspectChannelMedia(paths, slug, undefined, { fresh: true });
+ assert.equal(loc.status, "ok", loc.detail);
+ assert.equal(loc.text.readable, true);
+}
+
+test("clips/ is text by the classifier, and so is carried to the corpus disk", () => {
+ assert.equal(classifyEntry("clips"), "text");
+});
+
+test("a dry run writes nothing, and says what it would do", async () => {
+ resetCorpus();
+ seedLegacy();
+ const before = snapshot();
+ const lines: string[] = [];
+ const res = await run({ dryRun: true }, lines);
+ assert.equal(res.exitCode, 0);
+ const [r] = res.channels;
+ assert.equal(r.outcome, "would-migrate", r.detail);
+ assert.equal(r.tierableFiles, 3);
+ assert.equal(r.copy?.clips.files, 1);
+ assert.equal(r.copy?.source.files, 1, "source-media is carried, not tiered");
+ assert.equal(r.copy?.scratch.files, 3, "the part, the tmp and the Syncthing temp");
+ assert.equal(r.leftOnPlatter?.files, 1, "the dead *.temp.* stays on the platter");
+ assert.deepEqual(snapshot(), before);
+ assert.ok(lines.some((l) => l.includes("DRY RUN")));
+});
+
+test("a run migrates the channel; a rerun is a no-op", async () => {
+ resetCorpus();
+ seedLegacy();
+ const lines: string[] = [];
+ const res = await run({}, lines);
+ assert.equal(res.exitCode, 0, lines.join("\n"));
+ const [r] = res.channels;
+ assert.equal(r.outcome, "migrated", r.detail);
+ assert.equal(r.linksMade, 3);
+ assert.equal(r.copy?.clips.files, 1);
+ await assertMigrated();
+ // The live chat's cues (none here) aside, the raw replay's link answers the
+ // freshness check from the corpus disk with the file's time.
+ assert.equal(lstatSync(path.join(channelDir(), "data", "v1", "transcript.live_chat.json")).mtimeMs, OLD.getTime());
+ // Without --reclaim the platter keeps its text copies, and the reclaim note
+ // says so.
+ assert.ok(lstatSync(path.join(mediaTarget(), "v1", "transcript.en.vtt")).isFile());
+ assert.ok(lstatSync(path.join(channelDir(), RECLAIM_NOTE)).isFile());
+
+ const before = snapshot();
+ const again = await run();
+ assert.equal(again.exitCode, 0);
+ assert.equal(again.channels[0].outcome, "already");
+ assert.deepEqual(snapshot(), before, "the rerun writes nothing");
+});
+
+test("--reclaim deletes the platter's text copies and keeps the media", async () => {
+ resetCorpus();
+ seedLegacy();
+ const res = await run({ reclaim: true });
+ assert.equal(res.exitCode, 0, res.channels[0].detail);
+ const r = res.channels[0];
+ assert.equal(r.outcome, "migrated");
+ assert.ok((r.reclaimedFiles ?? 0) > 0);
+ assert.deepEqual(r.reclaimKept, []);
+ await assertMigrated(SLUG, 1, { reclaimed: true });
+ assert.throws(() => lstatSync(path.join(channelDir(), RECLAIM_NOTE)), "the note goes");
+ assert.deepEqual(readdirSync(path.join(mediaTarget(), "v1")).sort(), [
+ "audio.mp3",
+ "transcript.live_chat.json",
+ ]);
+ assert.deepEqual(readdirSync(path.join(mediaTarget(), "v2")), ["audio.opus"]);
+ assert.throws(() => lstatSync(path.join(mediaTarget(), "v3")), "an emptied dir goes");
+
+ // Reclaim again, later: the note is gone, so it refuses and touches nothing.
+ const before = snapshot();
+ const again = await run({ reclaim: true });
+ assert.equal(again.exitCode, 1);
+ assert.equal(again.channels[0].outcome, "refused");
+ assert.match(again.channels[0].detail ?? "", /no \.tier-migration\.reclaim beside its config/);
+ assert.deepEqual(snapshot(), before);
+});
+
+test("--reclaim after a plain run takes the copies a first run left", async () => {
+ resetCorpus();
+ seedLegacy();
+ await run();
+ const res = await run({ reclaim: true, dryRun: true });
+ assert.equal(res.channels[0].outcome, "already");
+ const wouldTake = res.channels[0].reclaimedFiles ?? 0;
+ assert.ok(wouldTake > 0);
+ assert.ok(lstatSync(path.join(mediaTarget(), "v1", "transcript.en.vtt")).isFile(), "dry: kept");
+ const real = await run({ reclaim: true });
+ assert.equal(real.channels[0].reclaimedFiles, wouldTake);
+ assert.throws(() => lstatSync(path.join(mediaTarget(), "v1", "transcript.en.vtt")));
+ await assertMigrated(SLUG, 1, { reclaimed: true });
+});
+
+test("--reclaim keeps a platter copy whose twin on the corpus disk is gone or differs", async () => {
+ resetCorpus();
+ seedLegacy();
+ await run();
+ // Since the migration: one text file removed, one rewritten to another size.
+ rmSync(path.join(channelDir(), "data", "v1", "transcript.en.vtt"));
+ writeFileSync(path.join(channelDir(), "data", "v2", "transcript.json"), "{}");
+ const lines: string[] = [];
+ const res = await run({ reclaim: true }, lines);
+ assert.equal(res.exitCode, 0, res.channels[0].detail);
+ assert.deepEqual(res.channels[0].reclaimKept?.sort(), ["v1/transcript.en.vtt", "v2/transcript.json"]);
+ assert.ok(lstatSync(path.join(mediaTarget(), "v1", "transcript.en.vtt")).isFile(), "kept");
+ assert.ok(lstatSync(path.join(mediaTarget(), "v2", "transcript.json")).isFile(), "kept");
+ assert.throws(() => lstatSync(path.join(mediaTarget(), "v1", "metadata.info.json")), "its twin is there: taken");
+ assert.throws(() => lstatSync(path.join(mediaTarget(), "v1", DEAD_TEMP)), "the dead temp is taken");
+ assert.ok(lines.some((l) => /kept 2 entries on the drive with no same-size copy/.test(l)));
+});
+
+test("a reclaim killed after its deletes resumes from its marker and ends clean", async () => {
+ resetCorpus();
+ seedLegacy();
+ const killed = await run({
+ reclaim: true,
+ deps: {
+ checkpoint: (_slug, s) => {
+ if (s === "reclaimed") throw new Error("killed");
+ },
+ },
+ });
+ assert.equal(killed.channels[0].outcome, "failed");
+ assert.equal((await readRelocationMarker(paths, SLUG))?.phase, "reclaim");
+ const resumed = await run();
+ assert.equal(resumed.channels[0].outcome, "migrated", resumed.channels[0].detail);
+ assert.equal(resumed.channels[0].resumedFrom, "reclaim");
+ await assertMigrated(SLUG, 1, { reclaimed: true });
+ assert.throws(() => lstatSync(path.join(channelDir(), RECLAIM_NOTE)));
+});
+
+test("--all --reclaim takes only channels this tool migrated", async () => {
+ resetCorpus(["moved", SLUG]);
+ seedLegacy();
+ // A channel the Storage panel moved: on the media tier, no reclaim note, a
+ // non-tierable file in its media tree that a reclaim would take.
+ writeConfig("moved", { mediaDir: mediaTarget("moved") });
+ mkdirSync(path.join(mediaTarget("moved"), "x1"), { recursive: true });
+ writeFileSync(path.join(mediaTarget("moved"), "x1", "stray.txt"), "s");
+ mkdirSync(path.join(channelDir("moved"), "data", "x1"), { recursive: true });
+ writeFileSync(path.join(channelDir("moved"), "data", "x1", "stray.txt"), "s");
+ symlinkSync(mediaTarget("moved"), path.join(channelDir("moved"), "media"));
+
+ const res = await run({ slugs: "all", reclaim: true });
+ assert.equal(res.exitCode, 0, res.channels.map((c) => c.detail).join("; "));
+ assert.deepEqual(res.channels.map((c) => c.slug), [SLUG]);
+ await assertMigrated(SLUG, 1, { reclaimed: true });
+ assert.ok(lstatSync(path.join(mediaTarget("moved"), "x1", "stray.txt")).isFile(), "not walked");
+
+ const named = await run({ slugs: ["moved"], reclaim: true });
+ assert.equal(named.channels[0].outcome, "refused");
+ assert.match(named.channels[0].detail ?? "", /migrate-tier did not migrate it/);
+});
+
+test("a media move's marker is refused, with or without a scope; nothing is touched", async () => {
+ for (const marker of [
+ { target: mediaTarget(), direction: "out", startedAt: "", phase: "copy", scope: "media" },
+ { target: path.join(PLATTER, "elsewhere", SLUG, "data"), direction: "out", startedAt: "", phase: "swap" },
+ ]) {
+ resetCorpus();
+ seedLegacy();
+ writeJson(path.join(channelDir(), ".relocating.json"), marker);
+ const before = snapshot();
+ const res = await run();
+ assert.equal(res.exitCode, 1);
+ assert.equal(res.channels[0].outcome, "refused");
+ assert.match(res.channels[0].detail ?? "", /media move .* Storage panel/);
+ assert.deepEqual(snapshot(), before);
+ }
+});
+
+test("a running editor refuses a real run; a dry run only notes it", async () => {
+ resetCorpus();
+ seedLegacy();
+ const before = snapshot();
+ const lines: string[] = [];
+ const editorProbe = async () => "the editor answers at http://localhost:3001/api/pulse (HTTP 200)";
+ const res = await run({ deps: { editorProbe } }, lines);
+ assert.equal(res.exitCode, 2);
+ assert.equal(res.channels.length, 0);
+ assert.ok(lines.some((l) => /Refusing: the editor must be stopped/.test(l)));
+ assert.deepEqual(snapshot(), before);
+ const dry = await run({ dryRun: true, deps: { editorProbe } });
+ assert.equal(dry.exitCode, 0);
+ assert.equal(dry.channels[0].outcome, "would-migrate");
+ assert.deepEqual(snapshot(), before);
+});
+
+test("editorRunningReason: the port, then a live job meta", async () => {
+ const jobsDir = path.join(ROOT, "jobs-probe");
+ rmSync(jobsDir, { recursive: true, force: true });
+ mkdirSync(jobsDir, { recursive: true });
+ const refused = async () => {
+ throw Object.assign(new TypeError("fetch failed"), { cause: { code: "ECONNREFUSED" } });
+ };
+ const timedOut = async () => {
+ throw Object.assign(new Error("The operation was aborted due to timeout"), { name: "TimeoutError" });
+ };
+ const answers = async () => new Response("ok", { status: 200 });
+ const opts = { url: "http://localhost:3999", machineBootMs: 0 };
+
+ assert.equal(await editorRunningReason({ jobsDir }, { ...opts, fetchImpl: refused as typeof fetch }), null);
+ assert.match(
+ (await editorRunningReason({ jobsDir }, { ...opts, fetchImpl: answers as typeof fetch })) ?? "",
+ /answers at http:\/\/localhost:3999\/api\/pulse \(HTTP 200\)/,
+ );
+ assert.match(
+ (await editorRunningReason({ jobsDir }, { ...opts, fetchImpl: timedOut as typeof fetch })) ?? "",
+ /did not answer/,
+ );
+
+ const meta = { id: "j1", kind: "whisper-video", queueKey: "q", channelSlug: "c", status: "running", queuedAt: Date.now(), pid: 424242 };
+ writeFileSync(path.join(jobsDir, "j1.meta.json"), JSON.stringify(meta));
+ assert.match(
+ (await editorRunningReason({ jobsDir }, { ...opts, fetchImpl: refused as typeof fetch, isAlive: () => true })) ?? "",
+ /job j1 \(whisper-video, c\) is running in a live process \(pid 424242\)/,
+ );
+ assert.equal(
+ await editorRunningReason({ jobsDir }, { ...opts, fetchImpl: refused as typeof fetch, isAlive: () => false }),
+ null,
+ "a dead writer's meta is a ghost",
+ );
+ writeFileSync(path.join(jobsDir, "j1.meta.json"), JSON.stringify({ ...meta, pid: undefined }));
+ assert.match(
+ (await editorRunningReason({ jobsDir }, { ...opts, fetchImpl: refused as typeof fetch })) ?? "",
+ /no recorded process/,
+ );
+ assert.equal(
+ await editorRunningReason({ jobsDir }, { ...opts, machineBootMs: Date.now() + 60_000, fetchImpl: refused as typeof fetch }),
+ null,
+ "a meta from before the machine booted is not read",
+ );
+});
+
+test("killed at any step between copy and swap, a rerun resumes and finishes", async () => {
+ const steps: MigrationStep[] = [
+ "copied",
+ "linked",
+ "marker-swap",
+ "platter-renamed",
+ "media-linked",
+ "data-unlinked",
+ "data-renamed",
+ "config-written",
+ ];
+ for (const step of steps) {
+ resetCorpus();
+ seedLegacy();
+ const killed = await run({
+ deps: {
+ checkpoint: (_slug, s) => {
+ if (s === step) throw new Error(`killed at ${s}`);
+ },
+ },
+ });
+ assert.equal(killed.channels[0].outcome, "failed", step);
+ assert.equal(killed.exitCode, 1);
+ const marker = await readRelocationMarker(paths, SLUG);
+ assert.equal(marker?.scope, "tier-migration", step);
+ const loc = await inspectChannelMedia(paths, SLUG, undefined, { fresh: true });
+ assert.equal(loc.status, "in-transition", step);
+ assert.equal(loc.text.readable, false, `${step}: the text is held mid-migration`);
+ assert.deepEqual(await legacyChannels(paths), [SLUG], `${step}: --all would resume it`);
+
+ const resumed = await run();
+ assert.equal(resumed.exitCode, 0, `${step}: ${resumed.channels[0].detail}`);
+ assert.equal(resumed.channels[0].outcome, "migrated", step);
+ assert.equal(
+ resumed.channels[0].resumedFrom,
+ ["copied", "linked"].includes(step) ? "copy" : "swap",
+ step,
+ );
+ await assertMigrated();
+ }
+});
+
+test("the free-space stop: refused before anything is written", async () => {
+ resetCorpus();
+ seedLegacy();
+ const before = snapshot();
+ const settings = { ...SETTINGS, minFreeDiskGB: 5 } as MigrateOpts["settings"];
+ // 6 GB free, a 5 GB floor and a 2 GB margin: the copy cannot fit.
+ const res = await run({ settings, deps: { freeBytes: async () => 6 * 1024 ** 3 } });
+ assert.equal(res.exitCode, 1);
+ assert.equal(res.channels[0].outcome, "no-space");
+ assert.match(res.channels[0].detail ?? "", /not enough space on the corpus disk: 6\.00 GB free .*5 GB floor.*2 GB resume margin/);
+ assert.deepEqual(snapshot(), before);
+ // 8 GB free clears it.
+ const ok = await run({ settings, deps: { freeBytes: async () => 8 * 1024 ** 3 } });
+ assert.equal(ok.channels[0].outcome, "migrated", ok.channels[0].detail);
+});
+
+test("--all: smallest text first, and a stop before the big three", async () => {
+ const big = "omnibased";
+ resetCorpus(["bigger", "small", big]);
+ seedLegacy("bigger", 4);
+ seedLegacy("small", 1);
+ seedLegacy(big, 1);
+ // A classic channel and a social one are not candidates.
+ writeConfig("classic");
+ mkdirSync(path.join(channelDir("classic"), "data", "x1"), { recursive: true });
+
+ const lines: string[] = [];
+ const res = await run({ slugs: "all" }, lines);
+ assert.equal(res.exitCode, 0, lines.join("\n"));
+ assert.deepEqual(res.channels.map((c) => c.slug), ["small", "bigger"]);
+ assert.deepEqual(res.stoppedBefore, [big]);
+ assert.ok(lines.some((l) => l.includes("STOPPED before the big-text channels: omnibased")));
+ assert.ok(lines.some((l) => l.includes("--all --include-large")));
+ await assertMigrated("small");
+ await assertMigrated("bigger", 4);
+ assert.equal(lstatSync(path.join(channelDir(big), "data")).isSymbolicLink(), true, "untouched");
+
+ const rest = await run({ slugs: "all", includeLarge: true });
+ assert.deepEqual(rest.channels.map((c) => c.slug), [big]);
+ await assertMigrated(big);
+ assert.deepEqual(await legacyChannels(paths), []);
+});
+
+test("the index build after the migration rewrites no record", async () => {
+ resetCorpus();
+ // Built while the channel's text was readable (before release 17 the index
+ // read a legacy channel through its link) — here, a classic layout with the
+ // same files and times.
+ writeConfig(SLUG);
+ writeVideos(path.join(channelDir(), "data"), SLUG);
+ const v1 = path.join(channelDir(), "data", "v1");
+ const outcome = await normalizeLiveChat({ videoDir: v1, channelSlug: SLUG, tier: false });
+ assert.equal(outcome.status, "wrote");
+ utimesSync(path.join(v1, "live_chat.cues.json"), CUES_AT, CUES_AT);
+ const first = await buildIndex({ paths, onLog: () => {} });
+ assert.equal(first.added, 3);
+ const subsBefore = statSync(path.join(v1, "transcript.live_chat.json")).mtimeMs;
+
+ // Onto the retired layout, as the old mover left it (rename keeps mtimes).
+ mkdirSync(path.dirname(oldTree()), { recursive: true });
+ renameSync(path.join(channelDir(), "data"), oldTree());
+ symlinkSync(oldTree(), path.join(channelDir(), "data"));
+ writeConfig(SLUG, { dataDir: oldTree() });
+
+ const res = await run();
+ assert.equal(res.channels[0].outcome, "migrated", res.channels[0].detail);
+ assert.equal(lstatSync(path.join(v1, "transcript.live_chat.json")).mtimeMs, subsBefore);
+ assert.equal((await isLiveChatCuesFresh(v1)).fresh, true, "the cues are not stale");
+
+ const second = await buildIndex({ paths, onLog: () => {} });
+ assert.deepEqual(second.heldChannels, []);
+ assert.equal(second.added, 0);
+ assert.equal(second.changed, 0, "no record rewritten: every mtime the index keys on is unchanged");
+ assert.equal(second.removed, 0);
+});
diff --git a/common/bin/migrate-media-tier.ts b/common/bin/migrate-media-tier.ts
@@ -0,0 +1,1277 @@
+#!/usr/bin/env tsx
+// THE ONE-OFF MEDIA-TIER MIGRATION (release 17, slice T3).
+//
+// archilyzer storage migrate-tier <slug>… [--dry-run] [--reclaim]
+// archilyzer storage migrate-tier --all [--order smallest] [--include-large] [--dry-run] [--reclaim]
+//
+// Before release 17 a relocation moved a channel's WHOLE `data/` to another
+// drive: `channels/<slug>/data` an absolute symlink to `<root>/<slug>/data`,
+// `config.dataDir` recording it. Such a channel is `legacy`
+// (lib/channelMedia.ts): its text is on the far drive, so it is held by every
+// guard until this runs. This brings the TEXT home and leaves the big files
+// where they are, in the release-17 layout (plans/release-17.md, "The model"):
+//
+// channels/<slug>/data/ a REAL directory on the corpus disk again
+// <id>/audio.mp3 -> ../../media/<id>/audio.mp3 (relative, the file's mtime)
+// channels/<slug>/media -> <root>/<slug>/media (ONE absolute link)
+// <root>/<slug>/media/<id>/… (the old `<root>/<slug>/data`, renamed)
+// config.json: mediaDir = <root>/<slug>/media, dataDir removed
+//
+// THE PHASES, per channel (plans/release-17.md, "The one-off migration"):
+//
+// preflight the channel is on the retired layout and its drive answers;
+// `assertRelocationRootPresent(root)`; the old tree is walked and
+// classified BY NAME (lib/mediaTier.ts); the space rule
+// `copy bytes + resume margin <= free(channelsDir) - floor`.
+// copy marker `{target: <root>/<slug>/media, direction: "out",
+// phase: "copy", scope: "tier-migration"}`; every entry that is
+// NOT tierable (the text, `clips/`, scratch, `source-media.*`) —
+// except a postprocessor's dead `*.temp.*`, left on the platter —
+// is copied to `channels/<slug>/data.incoming/` by rsync from a
+// NUL-separated list the classifier wrote (never a glob); the
+// copy is verified by an itemized dry run AND per-kind counts and
+// bytes; then every tierable file gets its relative link, carrying
+// the file's times (`lutimes` — without them every migrated live
+// chat's cues would read as stale).
+// swap platter `rename(<root>/<slug>/data -> <root>/<slug>/media)`;
+// `symlink(<root>/<slug>/media, channels/<slug>/media)`;
+// `unlink(channels/<slug>/data)`; `rename(data.incoming -> data)`;
+// `patchChannelConfig(slug, { mediaDir }, { unset: ["dataDir"] })`.
+// Each step looks at the disk first, so a rerun after a crash
+// between any two of them finishes the rest. The reclaim note
+// (`.tier-migration.reclaim`) is left beside the config.
+// reclaim only with --reclaim, and only on a channel carrying the note:
+// the platter's copies of the text (an entry of
+// `<root>/<slug>/media/<id>/` that is not tierable, and whose twin
+// on the corpus disk `lstat`s the same) and the dead temps are
+// deleted, empty video dirs dropped; anything else is kept and
+// reported.
+// done the marker is cleared; bytes by kind and links made are printed.
+//
+// THE EDITOR MUST BE STOPPED. It holds its job registry in memory, so nothing
+// on disk can say a job of its is about to write into the channel; a real run
+// refuses while the editor's port answers or `.jobs/` shows a job in a live
+// process (`editorRunningReason`). A DRY RUN WRITES NOTHING — no marker, no
+// list, no directory — and only warns.
+//
+// IDEMPOTENT AND RESUMABLE. A channel already on the new layout is a no-op; a
+// `tier-migration` marker is resumed from its phase. Any OTHER marker (a media
+// move's) is refused: that is the Storage panel's to finish.
+//
+// No index rebuild follows: rsync -a keeps every text file's mtime, the links
+// carry the media files' mtimes, and the index stats only sidecars.
+
+import os from "node:os";
+import path from "node:path";
+import {
+ lstat,
+ lutimes,
+ mkdir,
+ readdir,
+ readlink,
+ rename,
+ rm,
+ rmdir,
+ stat,
+ symlink,
+ unlink,
+ writeFile,
+} from "node:fs/promises";
+import type { Stats } from "node:fs";
+import type { Paths } from "../lib/paths";
+import type { SiteSettings } from "../lib/settings";
+import { formatBytes } from "../lib/format";
+import { getFreeBytes } from "../lib/diskSpace";
+import { classifyEntry, CLIPS_DIR, isTierable } from "../lib/mediaTier";
+import {
+ channelMediaLink,
+ relocatedMediaDir,
+ tierLinkTarget,
+} from "../lib/mediaTier-server";
+import {
+ forgetChannelMedia,
+ inspectChannelMedia,
+ readRelocationMarker,
+ relocatedDataDir,
+ relocationMarkerPath,
+ type RelocationMarker,
+} from "../lib/channelMedia";
+import {
+ clearDirMarker,
+ linkOrDirState,
+ makeProgressSink,
+ rsyncTree,
+ writeDirMarker,
+} from "../controller/relocateDir";
+import {
+ assertRelocationRootPresent,
+ rootOfRelocatedMediaDir,
+} from "../controller/relocateChannelMedia";
+import { patchChannelConfig, readChannelConfig } from "../controller/channels";
+import { readJobMeta } from "../jobs/jobMeta";
+import { processIsAlive } from "../jobs/bootQueuedJobs";
+import { parseArgv } from "./_parseFlags";
+import { runIfEntryPoint } from "./_cli";
+
+const GB = 1024 ** 3;
+
+// THE THREE BIG-TEXT CHANNELS (plans/release-17.md, Rulings: "a stop before
+// the three big-text channels … with the free space reported; the operator
+// decides"). `--all` migrates everything else, smallest text first, and stops
+// before these; `--include-large` (or naming one) goes on.
+export const LARGE_TEXT_CHANNELS: readonly string[] = [
+ "omnibased",
+ "rekietalaw",
+ "the-quartering-rumble",
+];
+
+// The rebuilt text tree, beside `data` in the channel dir until the swap.
+export const INCOMING_NAME = "data.incoming";
+// The NUL-separated rsync list, beside the marker; removed when the channel is done.
+export const LIST_NAME = ".tier-migration.files";
+// THE RECLAIM NOTE, beside the config from the swap until a reclaim: "this
+// tool migrated this channel, and the drive still holds its text copies".
+// `--all --reclaim` takes only channels carrying it, and a named `--reclaim`
+// refuses one without it — a channel the Storage panel moved has nothing of
+// the retired layout to take, and walking its whole tree to find that out
+// costs minutes on a platter.
+export const RECLAIM_NOTE = ".tier-migration.reclaim";
+
+// ---------------------------------------------------------------------------
+// The inventory: the old tree, by kind
+// ---------------------------------------------------------------------------
+
+// What the copy carries, by kind. `source` is a media file that is NOT
+// tierable — `source-media.*` stays a real file in `data/<id>/` (the
+// saved-video store renames it), so it is copied with the text rather than
+// left on the platter with no link (and then deleted by --reclaim).
+export type CopyKind = "text" | "clips" | "scratch" | "source";
+export const COPY_KINDS: readonly CopyKind[] = ["text", "clips", "scratch", "source"];
+
+export type Tally = { files: number; bytes: number };
+export type KindTallies = Record<CopyKind, Tally>;
+
+function emptyTallies(): KindTallies {
+ return {
+ text: { files: 0, bytes: 0 },
+ clips: { files: 0, bytes: 0 },
+ scratch: { files: 0, bytes: 0 },
+ source: { files: 0, bytes: 0 },
+ };
+}
+
+export type TierableFile = {
+ id: string;
+ name: string;
+ bytes: number;
+ atime: Date;
+ mtime: Date;
+};
+
+export type Inventory = {
+ // Every top-level directory of the old tree (a video dir), in order.
+ videoIds: string[];
+ // Paths relative to the old tree, for rsync's --files-from.
+ copyList: string[];
+ copy: KindTallies;
+ copyBytes: number;
+ tierable: TierableFile[];
+ tierableBytes: number;
+ // Dead postprocessor temps (`*.temp.*`) left on the platter, neither copied
+ // nor linked; `--reclaim` deletes them.
+ left: Tally;
+};
+
+// Which copy kind a video-dir entry is. `stats` decides only one thing: a
+// tierable NAME that is not a regular file (a directory, a link) is not
+// something the link step can stand in for, so it is copied as it is.
+function copyKindOf(name: string): CopyKind {
+ if (name === CLIPS_DIR) return "clips";
+ const kind = classifyEntry(name);
+ if (kind === "scratch") return "scratch";
+ if (kind === "media") return "source";
+ return "text";
+}
+
+// Files and bytes under one entry: a file or link counts itself (`lstat`
+// size — rsync -a copies a link as a link), a directory everything below it.
+async function measureEntry(p: string, st?: Stats): Promise<Tally> {
+ const s = st ?? (await lstat(p));
+ if (!s.isDirectory()) return { files: 1, bytes: s.size };
+ const out = { files: 0, bytes: 0 };
+ for (const name of await readdir(p)) {
+ const t = await measureEntry(path.join(p, name));
+ out.files += t.files;
+ out.bytes += t.bytes;
+ }
+ return out;
+}
+
+function add(t: KindTallies, kind: CopyKind, by: Tally): void {
+ t[kind].files += by.files;
+ t[kind].bytes += by.bytes;
+}
+
+// A POSTPROCESSOR'S DEAD TEMP — `source-media.temp.mp4`, `audio.temp.mp3`:
+// yt-dlp's postprocessor writes it and renames it over the final name, so one
+// still there after the run is a leftover (15 GB of them on one channel,
+// measured 2026-10-02). It is scratch by the classifier; the migration leaves
+// it on the platter, unlinked, rather than carry it to the corpus disk, and
+// `--reclaim` deletes it with the text copies.
+export function isLeftOnPlatter(name: string): boolean {
+ return classifyEntry(name) === "scratch" && /\.temp\./.test(name);
+}
+
+// ONE WALK OF THE OLD TREE: each video dir's entries classified by name. A
+// tierable regular file is left where it is (a link stands in for it); every
+// other entry is listed for the copy and measured.
+export async function inventoryTree(dir: string): Promise<Inventory> {
+ const inv: Inventory = {
+ videoIds: [],
+ copyList: [],
+ copy: emptyTallies(),
+ copyBytes: 0,
+ tierable: [],
+ tierableBytes: 0,
+ left: { files: 0, bytes: 0 },
+ };
+ const top = (await readdir(dir, { withFileTypes: true })).sort((a, b) =>
+ a.name.localeCompare(b.name),
+ );
+ for (const e of top) {
+ const p = path.join(dir, e.name);
+ if (!e.isDirectory()) {
+ // A stray file at the top of `data/`: text, carried.
+ inv.copyList.push(e.name);
+ add(inv.copy, copyKindOf(e.name), await measureEntry(p));
+ continue;
+ }
+ inv.videoIds.push(e.name);
+ const names = (await readdir(p)).sort();
+ for (const name of names) {
+ const fp = path.join(p, name);
+ const st = await lstat(fp);
+ if (isTierable(name) && st.isFile()) {
+ inv.tierable.push({
+ id: e.name,
+ name,
+ bytes: st.size,
+ atime: st.atime,
+ mtime: st.mtime,
+ });
+ inv.tierableBytes += st.size;
+ continue;
+ }
+ if (isLeftOnPlatter(name)) {
+ const t = await measureEntry(fp, st);
+ inv.left.files += t.files;
+ inv.left.bytes += t.bytes;
+ continue;
+ }
+ inv.copyList.push(`${e.name}/${name}`);
+ add(inv.copy, copyKindOf(name), await measureEntry(fp, st));
+ }
+ }
+ inv.copyBytes = COPY_KINDS.reduce((n, k) => n + inv.copy[k].bytes, 0);
+ return inv;
+}
+
+// THE REBUILT TREE, measured the same way, minus the links the migration made
+// (a link whose name is tierable and whose target is the tier link). A resumed
+// copy's tree already holds them.
+export async function measureIncoming(dir: string): Promise<KindTallies> {
+ const out = emptyTallies();
+ let top;
+ try {
+ top = await readdir(dir, { withFileTypes: true });
+ } catch {
+ return out;
+ }
+ for (const e of top) {
+ const p = path.join(dir, e.name);
+ if (!e.isDirectory()) {
+ add(out, copyKindOf(e.name), await measureEntry(p));
+ continue;
+ }
+ for (const name of await readdir(p)) {
+ const fp = path.join(p, name);
+ const st = await lstat(fp);
+ if (st.isSymbolicLink() && isTierable(name)) {
+ const target = await readlink(fp).catch(() => "");
+ if (target === tierLinkTarget(e.name, name)) continue;
+ }
+ add(out, copyKindOf(name), await measureEntry(fp, st));
+ }
+ }
+ return out;
+}
+
+function tallyText(t: Tally): string {
+ return `${t.files} file(s), ${formatBytes(t.bytes)}`;
+}
+
+function kindsLine(t: KindTallies): string {
+ return COPY_KINDS.map((k) => `${k} ${tallyText(t[k])}`).join("; ");
+}
+
+// ---------------------------------------------------------------------------
+// The editor must be stopped
+// ---------------------------------------------------------------------------
+
+export const DEFAULT_EDITOR_URL = "http://localhost:3001";
+
+export type EditorCheckOptions = {
+ url?: string;
+ fetchImpl?: typeof fetch;
+ isAlive?: (pid: number) => boolean;
+ // Ms since epoch the machine booted (a pid-less `running` meta from before
+ // it is a ghost).
+ machineBootMs?: number;
+ timeoutMs?: number;
+};
+
+// WHY THE EDITOR IS (OR MAY BE) RUNNING, or null. Two questions, either
+// enough:
+// 1. Does its port answer? `GET <ARCHILYZER_EDITOR_URL>/api/pulse`: any HTTP
+// answer is an editor; a refused connection is none; a connection that
+// does not answer within the timeout is treated as one (the editor's
+// main thread can be busy for minutes — release 17 step 0).
+// 2. Does `.jobs/` show a job in a live process? A `running` or `queued`
+// meta whose writer (`pid`) is alive and is not this process — the editor,
+// or an `archilyzer run`. A `running` meta with no pid (written before
+// release 17) counts when it was written after the machine booted: the
+// editor's boot pass closes those, so one left is a process that has not
+// been through it.
+export async function editorRunningReason(
+ paths: Pick<Paths, "jobsDir">,
+ opts: EditorCheckOptions = {},
+): Promise<string | null> {
+ const base = (opts.url ?? process.env.ARCHILYZER_EDITOR_URL ?? DEFAULT_EDITOR_URL)
+ .trim()
+ .replace(/\/+$/, "");
+ const url = `${base}/api/pulse`;
+ const doFetch = opts.fetchImpl ?? fetch;
+ const timeoutMs = opts.timeoutMs ?? 3000;
+ try {
+ const res = await doFetch(url, { signal: AbortSignal.timeout(timeoutMs) });
+ return `the editor answers at ${url} (HTTP ${res.status})`;
+ } catch (err) {
+ const name = (err as Error | null)?.name;
+ if (name === "TimeoutError" || name === "AbortError") {
+ return (
+ `something is listening at ${base} but did not answer ${url} within ` +
+ `${Math.round(timeoutMs / 1000)} s — if it is the editor, stop it`
+ );
+ }
+ /* refused, unresolvable: nothing is listening there */
+ }
+
+ const isAlive = opts.isAlive ?? processIsAlive;
+ const bootMs = opts.machineBootMs ?? Date.now() - os.uptime() * 1000;
+ let names: string[];
+ try {
+ names = await readdir(paths.jobsDir);
+ } catch {
+ return null;
+ }
+ const live: string[] = [];
+ for (const file of names) {
+ if (!file.endsWith(".meta.json")) continue;
+ // A meta a live process wrote was written after the machine booted; an
+ // older file cannot be one, and is not read.
+ try {
+ if ((await stat(path.join(paths.jobsDir, file))).mtimeMs < bootMs) continue;
+ } catch {
+ continue;
+ }
+ const meta = await readJobMeta(
+ paths as Paths,
+ file.slice(0, -".meta.json".length),
+ );
+ if (!meta || (meta.status !== "running" && meta.status !== "queued")) continue;
+ const what = `job ${meta.id} (${meta.kind}${meta.channelSlug ? `, ${meta.channelSlug}` : ""}) is ${meta.status}`;
+ if (typeof meta.pid === "number") {
+ if (meta.pid !== process.pid && isAlive(meta.pid)) {
+ live.push(`${what} in a live process (pid ${meta.pid})`);
+ }
+ } else if (meta.status === "running") {
+ const at = meta.startedAt ?? meta.queuedAt;
+ if (typeof at === "number" && at >= bootMs) {
+ live.push(`${what} with no recorded process, since after this machine booted`);
+ }
+ }
+ }
+ if (live.length === 0) return null;
+ const more = live.length > 3 ? ` (and ${live.length - 3} more)` : "";
+ return `${live.slice(0, 3).join("; ")}${more}`;
+}
+
+// ---------------------------------------------------------------------------
+// One channel
+// ---------------------------------------------------------------------------
+
+export type MigrationStep =
+ | "copied"
+ | "linked"
+ | "marker-swap"
+ | "platter-renamed"
+ | "media-linked"
+ | "data-unlinked"
+ | "data-renamed"
+ | "config-written"
+ | "reclaimed";
+
+export type MigrateDeps = {
+ // Why the editor is running, or null. Default: `editorRunningReason`.
+ editorProbe?: (paths: Paths) => Promise<string | null>;
+ freeBytes?: (dir: string) => Promise<number>;
+ // Test seam: called after each named step; a throw is a kill there.
+ checkpoint?: (slug: string, step: MigrationStep) => void | Promise<void>;
+};
+
+export type MigrateSettings = Pick<
+ SiteSettings,
+ "minFreeDiskGB" | "resumeMarginGB" | "storage"
+>;
+
+export type MigrateOptions = {
+ paths: Paths;
+ settings: MigrateSettings;
+ // Slugs named on the command line, or every legacy channel.
+ slugs: string[] | "all";
+ dryRun?: boolean;
+ reclaim?: boolean;
+ // `--all` goes on past the three big-text channels.
+ includeLarge?: boolean;
+ log?: (line: string) => void;
+ deps?: MigrateDeps;
+};
+
+export type ChannelOutcome =
+ | "migrated"
+ | "already"
+ | "not-legacy"
+ | "would-migrate"
+ | "no-space"
+ | "refused"
+ | "failed";
+
+export type ChannelReport = {
+ slug: string;
+ outcome: ChannelOutcome;
+ detail?: string;
+ copy?: KindTallies;
+ copyBytes?: number;
+ tierableFiles?: number;
+ tierableBytes?: number;
+ // Dead `*.temp.*` left on the platter (not copied, not linked).
+ leftOnPlatter?: Tally;
+ linksMade?: number;
+ linksAlready?: number;
+ reclaimedFiles?: number;
+ reclaimedBytes?: number;
+ // Platter entries `--reclaim` kept: no same-size copy on the corpus disk.
+ reclaimKept?: string[];
+ // The phase a marker was resumed from.
+ resumedFrom?: string;
+};
+
+export type MigrateResult = {
+ exitCode: number;
+ channels: ChannelReport[];
+ // `--all` stopped before these (the big three), with this much free.
+ stoppedBefore?: string[];
+ freeBytes?: number;
+ editorRunning?: string | null;
+};
+
+class Refusal extends Error {
+ constructor(
+ message: string,
+ readonly outcome: "refused" | "no-space" = "refused",
+ ) {
+ super(message);
+ this.name = "Refusal";
+ }
+}
+
+type Ctx = {
+ paths: Paths;
+ settings: MigrateSettings;
+ dryRun: boolean;
+ reclaim: boolean;
+ log: (line: string) => void;
+ freeBytes: (dir: string) => Promise<number>;
+ checkpoint: (slug: string, step: MigrationStep) => Promise<void>;
+ // A dry run of several channels: the space the earlier ones would take.
+ projectedUse: number;
+};
+
+// Where a channel stands, read from the corpus disk alone.
+type Plan =
+ | { kind: "migrate"; D: string; root: string; target: string; marker: RelocationMarker | null }
+ | { kind: "migrated"; mediaDir: string }
+ | { kind: "not-legacy" };
+
+function sameDir(a: string, b: string): boolean {
+ return path.resolve(a) === path.resolve(b);
+}
+
+async function planChannel(ctx: Ctx, slug: string): Promise<Plan> {
+ const { paths } = ctx;
+ const channelDir = path.join(paths.channelsDir, slug);
+ const config = await readChannelConfig(paths, slug);
+ if (!config) throw new Refusal(`no channel "${slug}" (no readable config.json)`);
+ const dataLink = path.join(channelDir, "data");
+ const dataState = await linkOrDirState(dataLink);
+ const marker = await readRelocationMarker(paths, slug);
+
+ if (marker && marker.scope !== "tier-migration") {
+ throw new Refusal(
+ `a media move (${marker.direction}, phase "${marker.phase}", target ` +
+ `${marker.target}) is in flight or was interrupted — finish it from the ` +
+ `channel's Storage panel first; this tool resumes only its own marker`,
+ );
+ }
+
+ if (marker) {
+ // RESUMING: the marker is the source of truth for the target — the swap
+ // unsets `dataDir` before the marker is cleared.
+ const target = marker.target;
+ if (path.basename(target) !== "media" || path.basename(path.dirname(target)) !== slug) {
+ throw new Refusal(`its tier-migration marker names ${target}, not <root>/${slug}/media`);
+ }
+ const root = rootOfRelocatedMediaDir(target, slug);
+ const D = relocatedDataDir(root, slug);
+ const recorded = config.dataDir?.trim();
+ if (recorded && !sameDir(recorded, D)) {
+ throw new Refusal(
+ `its tier-migration marker names ${target} but config.json records dataDir ${recorded}`,
+ );
+ }
+ return { kind: "migrate", D, root, target, marker };
+ }
+
+ const recorded = config.dataDir?.trim();
+ const linked = dataState.kind === "link" ? path.resolve(channelDir, dataState.linkTarget) : undefined;
+ if (!recorded && !linked) {
+ const mediaDir = config.mediaDir?.trim();
+ if (mediaDir && dataState.kind !== "other") return { kind: "migrated", mediaDir };
+ return { kind: "not-legacy" };
+ }
+ const D = recorded ?? (linked as string);
+ if (recorded && linked && !sameDir(recorded, linked)) {
+ throw new Refusal(
+ `config.json records dataDir ${recorded} but ${dataLink} points at ${linked} — ` +
+ `inconsistent; fix one by hand first`,
+ );
+ }
+ if (dataState.kind !== "link") {
+ throw new Refusal(
+ `config.json records dataDir ${recorded} but ${dataLink} is ` +
+ `${dataState.kind === "real-dir" ? "a real directory" : dataState.kind} — ` +
+ `inconsistent; fix one by hand first`,
+ );
+ }
+ if (path.basename(D) !== "data" || path.basename(path.dirname(D)) !== slug) {
+ throw new Refusal(`its dataDir ${D} is not <root>/${slug}/data`);
+ }
+ if (config.mediaDir?.trim()) {
+ throw new Refusal(
+ `config.json records both dataDir ${D} and mediaDir ${config.mediaDir.trim()} — ` +
+ `inconsistent; fix one by hand first`,
+ );
+ }
+ const root = rootOfRelocatedMediaDir(D, slug);
+ return { kind: "migrate", D, root, target: relocatedMediaDir(root, slug), marker: null };
+}
+
+async function isRealDir(p: string): Promise<boolean> {
+ try {
+ return (await stat(p)).isDirectory();
+ } catch {
+ return false;
+ }
+}
+
+async function writeMarker(ctx: Ctx, slug: string, marker: RelocationMarker): Promise<void> {
+ await writeDirMarker(relocationMarkerPath(ctx.paths, slug), marker);
+ forgetChannelMedia(slug);
+}
+
+// THE VERIFY'S DRY RUN: what rsync would still send, by line. A directory
+// whose only difference is its time (`.d..t`) is not content.
+function contentDrift(output: string): string[] {
+ return output
+ .split(/[\r\n]+/)
+ .map((l) => l.trim())
+ .filter((l) => l.length > 0)
+ .filter((l) => !l.startsWith("sending incremental"))
+ .filter((l) => !/^(sent |total size |building file list|done$)/.test(l))
+ .filter((l) => !/^\.d\.\.t/.test(l));
+}
+
+function rsyncListArgs(listFile: string): string[] {
+ // -r EXPLICITLY: --files-from turns off -a's implied recursion, and
+ // `clips/` and a scratch dir are directories to carry whole.
+ return ["-a", "-r", "--from0", `--files-from=${listFile}`];
+}
+
+async function migrateOne(ctx: Ctx, slug: string): Promise<ChannelReport> {
+ const { paths, log } = ctx;
+ const plan = await planChannel(ctx, slug);
+ if (plan.kind === "not-legacy") {
+ return { slug, outcome: "not-legacy", detail: "not on the retired layout — nothing to migrate" };
+ }
+ if (plan.kind === "migrated") {
+ const report: ChannelReport = { slug, outcome: "already", detail: "already on the media tier" };
+ if (ctx.reclaim) Object.assign(report, await reclaimChannel(ctx, slug, plan.mediaDir));
+ return report;
+ }
+
+ const { D, root, target, marker } = plan;
+ const channelDir = path.join(paths.channelsDir, slug);
+ const dataLink = path.join(channelDir, "data");
+ const incoming = path.join(channelDir, INCOMING_NAME);
+ const listFile = path.join(channelDir, LIST_NAME);
+ const mediaLink = channelMediaLink(paths, slug);
+ let phase = marker?.phase ?? "copy";
+ const report: ChannelReport = { slug, outcome: "migrated" };
+ if (marker) {
+ report.resumedFrom = phase;
+ log(`${slug}: resuming a tier migration from phase "${phase}"`);
+ }
+
+ if (phase === "copy") {
+ // PREFLIGHT. Everything that can refuse does so here, before the marker.
+ if (!(await isRealDir(D))) {
+ throw new Refusal(`${D} is not reachable (is the drive mounted?)`);
+ }
+ await assertRelocationRootPresent(root, ctx.settings.storage, paths);
+ const targetState = await linkOrDirState(target);
+ if (targetState.kind !== "missing") {
+ throw new Refusal(
+ `${target} already exists — a media move or an earlier attempt left it; ` +
+ `look at it before migrating`,
+ );
+ }
+ const mediaState = await linkOrDirState(mediaLink);
+ if (mediaState.kind !== "missing") {
+ throw new Refusal(`${mediaLink} already exists — look at it before migrating`);
+ }
+ if (!marker && (await linkOrDirState(incoming)).kind !== "missing") {
+ throw new Refusal(
+ `${incoming} exists with no tier-migration marker — a leftover; remove it ` +
+ `by hand after checking it`,
+ );
+ }
+
+ log(`${slug}: walking ${D}…`);
+ const inv = await inventoryTree(D);
+ report.copy = inv.copy;
+ report.copyBytes = inv.copyBytes;
+ report.tierableFiles = inv.tierable.length;
+ report.tierableBytes = inv.tierableBytes;
+ report.leftOnPlatter = inv.left;
+ log(
+ `${slug}: ${inv.videoIds.length} video dir(s); to the corpus disk: ` +
+ `${formatBytes(inv.copyBytes)} (${kindsLine(inv.copy)}); staying on the ` +
+ `media drive: ${inv.tierable.length} file(s), ${formatBytes(inv.tierableBytes)}` +
+ (inv.left.files > 0
+ ? `; dead postprocessor temps left there unlinked (--reclaim deletes them): ${tallyText(inv.left)}`
+ : ""),
+ );
+
+ // THE SPACE RULE: the copy plus the resume margin must fit above the disk
+ // gate's floor on the corpus disk — landing the text with nothing to spare
+ // would stop the next download at the gate. The margin is zero when the
+ // gate is off (the mover's rule). A resumed copy needs only what it lacks.
+ const floorGB = ctx.settings.minFreeDiskGB;
+ const marginGB = floorGB > 0 ? ctx.settings.resumeMarginGB : 0;
+ const already = marker ? sumBytes(await measureIncoming(incoming)) : 0;
+ const needed = Math.max(0, inv.copyBytes - already) + marginGB * GB;
+ const free = (await ctx.freeBytes(paths.channelsDir)) - ctx.projectedUse;
+ const usable = free - floorGB * GB;
+ if (needed > usable) {
+ throw new Refusal(
+ `not enough space on the corpus disk: ${formatBytes(Math.max(0, free))} free` +
+ (floorGB > 0 ? ` (${formatBytes(Math.max(0, usable))} above the ${floorGB} GB floor)` : "") +
+ `; ${formatBytes(Math.max(0, inv.copyBytes - already))} to copy` +
+ (marginGB > 0 ? ` plus a ${marginGB} GB resume margin` : ""),
+ "no-space",
+ );
+ }
+
+ if (ctx.dryRun) {
+ ctx.projectedUse += Math.max(0, inv.copyBytes - already);
+ report.outcome = "would-migrate";
+ report.detail =
+ `would copy ${formatBytes(inv.copyBytes - already)} to ${incoming}, link ` +
+ `${inv.tierable.length} file(s), rename ${D} to ${target}` +
+ (ctx.reclaim ? `, then reclaim the platter's text copies` : "");
+ return report;
+ }
+
+ // COPY.
+ await writeMarker(ctx, slug, {
+ target,
+ direction: "out",
+ startedAt: marker?.startedAt || new Date().toISOString(),
+ phase: "copy",
+ scope: "tier-migration",
+ });
+ await writeFile(listFile, inv.copyList.map((p) => `${p}\0`).join(""));
+ await mkdir(incoming).catch((err: NodeJS.ErrnoException) => {
+ if (err.code !== "EEXIST") throw err;
+ });
+ const sink = makeProgressSink({ totalBytes: inv.copyBytes, alreadyBytes: already, log });
+ const copied = await rsyncTree({
+ rsyncBin: paths.rsyncBin,
+ src: D,
+ dest: incoming,
+ args: [...rsyncListArgs(listFile), "--partial", "--info=progress2"],
+ log,
+ progress: sink,
+ });
+ if (copied.exitCode !== 0) {
+ throw new Error(
+ `rsync failed (exit ${copied.exitCode}); nothing on ${D} was touched — run again to resume`,
+ );
+ }
+ // VERIFY: rsync has nothing left to send…
+ const dry = await rsyncTree({
+ rsyncBin: paths.rsyncBin,
+ src: D,
+ dest: incoming,
+ args: [...rsyncListArgs(listFile), "--dry-run", "--itemize-changes"],
+ log: (m) => {
+ if (m.startsWith("$ ")) log(m);
+ },
+ });
+ if (dry.exitCode !== 0) throw new Error(`the verification rsync failed (exit ${dry.exitCode})`);
+ const drift = contentDrift(dry.output);
+ if (drift.length > 0) {
+ throw new Error(
+ `verification failed: ${drift.length} item(s) still differ (first: ${drift[0]}); ` +
+ `nothing on ${D} was touched`,
+ );
+ }
+ // …and the rebuilt tree holds what the old one does, kind by kind.
+ const got = await measureIncoming(incoming);
+ const off = COPY_KINDS.filter(
+ (k) => got[k].files !== inv.copy[k].files || got[k].bytes !== inv.copy[k].bytes,
+ );
+ if (off.length > 0) {
+ throw new Error(
+ `verification failed: ${off
+ .map((k) => `${k} ${tallyText(got[k])} copied, ${tallyText(inv.copy[k])} on ${D}`)
+ .join("; ")}; nothing on ${D} was touched`,
+ );
+ }
+ log(`${slug}: copy verified (${kindsLine(got)})`);
+ await ctx.checkpoint(slug, "copied");
+
+ // THE LINKS, each carrying its file's times.
+ let made = 0;
+ let had = 0;
+ for (const id of inv.videoIds) {
+ await mkdir(path.join(incoming, id)).catch((err: NodeJS.ErrnoException) => {
+ if (err.code !== "EEXIST") throw err;
+ });
+ }
+ for (const f of inv.tierable) {
+ const p = path.join(incoming, f.id, f.name);
+ const want = tierLinkTarget(f.id, f.name);
+ try {
+ await symlink(want, p);
+ made++;
+ } catch (err) {
+ if ((err as NodeJS.ErrnoException).code !== "EEXIST") throw err;
+ const st = await linkOrDirState(p);
+ if (st.kind !== "link" || st.linkTarget !== want) {
+ throw new Error(`${p} exists and is not the link ${want}`);
+ }
+ had++;
+ }
+ await lutimes(p, f.atime, f.mtime);
+ }
+ report.linksMade = made;
+ report.linksAlready = had;
+ log(`${slug}: ${made} link(s) made${had ? `, ${had} already there` : ""}`);
+ await ctx.checkpoint(slug, "linked");
+
+ await writeMarker(ctx, slug, {
+ target,
+ direction: "out",
+ startedAt: marker?.startedAt || new Date().toISOString(),
+ phase: "swap",
+ scope: "tier-migration",
+ });
+ phase = "swap";
+ await ctx.checkpoint(slug, "marker-swap");
+ } else if (ctx.dryRun) {
+ report.outcome = "would-migrate";
+ report.detail = `would resume its tier migration from phase "${phase}"`;
+ return report;
+ }
+
+ if (phase === "swap") {
+ await swap(ctx, slug, { D, target, dataLink, incoming, mediaLink });
+ await writeFile(
+ path.join(channelDir, RECLAIM_NOTE),
+ JSON.stringify({ target, migratedAt: new Date().toISOString() }) + "\n",
+ );
+ await rm(listFile, { force: true });
+ phase = "reclaim";
+ }
+
+ if (ctx.reclaim || marker?.phase === "reclaim") {
+ await writeMarker(ctx, slug, {
+ target,
+ direction: "out",
+ startedAt: marker?.startedAt || new Date().toISOString(),
+ phase: "reclaim",
+ scope: "tier-migration",
+ });
+ Object.assign(report, await reclaimTree(ctx, slug, target));
+ await ctx.checkpoint(slug, "reclaimed");
+ await rm(path.join(channelDir, RECLAIM_NOTE), { force: true });
+ }
+
+ await clearDirMarker(relocationMarkerPath(paths, slug));
+ forgetChannelMedia(slug);
+ const after = await inspectChannelMedia(paths, slug, undefined, { fresh: true });
+ if (after.status !== "ok") {
+ report.outcome = "failed";
+ report.detail = `migrated, but its media reads ${after.status}: ${after.detail ?? ""}`;
+ }
+ return report;
+}
+
+function sumBytes(t: KindTallies): number {
+ return COPY_KINDS.reduce((n, k) => n + t[k].bytes, 0);
+}
+
+// THE SWAP, every step asking the disk what is there first, so a rerun after a
+// crash between any two of them does the rest and nothing twice.
+async function swap(
+ ctx: Ctx,
+ slug: string,
+ p: { D: string; target: string; dataLink: string; incoming: string; mediaLink: string },
+): Promise<void> {
+ const { log } = ctx;
+ // 1. The platter: <root>/<slug>/data -> <root>/<slug>/media.
+ const d = await linkOrDirState(p.D);
+ const t = await linkOrDirState(p.target);
+ if (d.kind === "real-dir" && t.kind === "missing") {
+ await rename(p.D, p.target);
+ log(`${slug}: renamed ${p.D} -> ${p.target}`);
+ } else if (!(d.kind === "missing" && t.kind === "real-dir")) {
+ throw new Error(
+ `cannot swap: ${p.D} is ${d.kind} and ${p.target} is ${t.kind} — the drive may ` +
+ `not be mounted; nothing was changed by this step`,
+ );
+ }
+ await ctx.checkpoint(slug, "platter-renamed");
+
+ // 2. The corpus disk: channels/<slug>/media -> <root>/<slug>/media.
+ const m = await linkOrDirState(p.mediaLink);
+ if (m.kind === "missing") {
+ await symlink(p.target, p.mediaLink);
+ } else if (!(m.kind === "link" && sameDir(path.resolve(path.dirname(p.mediaLink), m.linkTarget), p.target))) {
+ throw new Error(`cannot swap: ${p.mediaLink} is ${m.kind}, not the link to ${p.target}`);
+ }
+ await ctx.checkpoint(slug, "media-linked");
+
+ // 3. The old whole-directory link goes.
+ const dl = await linkOrDirState(p.dataLink);
+ if (dl.kind === "link") {
+ if (!sameDir(path.resolve(path.dirname(p.dataLink), dl.linkTarget), p.D)) {
+ throw new Error(`cannot swap: ${p.dataLink} points at ${dl.linkTarget}, not ${p.D}`);
+ }
+ await unlink(p.dataLink);
+ } else if (dl.kind === "other") {
+ throw new Error(`cannot swap: ${p.dataLink} is neither a link nor a directory`);
+ }
+ await ctx.checkpoint(slug, "data-unlinked");
+
+ // 4. The rebuilt text tree takes its name.
+ const inc = await linkOrDirState(p.incoming);
+ const now = await linkOrDirState(p.dataLink);
+ if (inc.kind === "real-dir" && now.kind === "missing") {
+ await rename(p.incoming, p.dataLink);
+ } else if (!(inc.kind === "missing" && now.kind === "real-dir")) {
+ throw new Error(`cannot swap: ${p.incoming} is ${inc.kind} and ${p.dataLink} is ${now.kind}`);
+ }
+ await ctx.checkpoint(slug, "data-renamed");
+
+ // 5. The record.
+ await patchChannelConfig(ctx.paths, slug, { mediaDir: p.target }, { unset: ["dataDir"] });
+ forgetChannelMedia(slug);
+ log(`${slug}: config.json records mediaDir ${p.target}; dataDir removed`);
+ await ctx.checkpoint(slug, "config-written");
+}
+
+// RECLAIM: the platter's copies of the text. In `<root>/<slug>/media/<id>/`
+// a tierable regular file is the media and stays; a dead `*.temp.*` (left
+// there by the migration, never copied) goes; every other entry goes ONLY when
+// one `lstat` of its twin on the corpus disk — `channels/<slug>/data/<id>/<name>`
+// — finds it there, a directory for a directory or a file of the same size.
+// Anything else is KEPT and reported: the platter copy may be the only one (a
+// reclaim long after the migration, a file changed or removed since). A video
+// dir left empty goes (non-recursive `rmdir`). A top-level non-directory entry
+// is a stray file copied from the top of the old tree, checked the same way.
+async function reclaimTree(
+ ctx: Ctx,
+ slug: string,
+ mediaDir: string,
+): Promise<Pick<ChannelReport, "reclaimedFiles" | "reclaimedBytes" | "reclaimKept">> {
+ const ssd = path.join(ctx.paths.channelsDir, slug, "data");
+ let files = 0;
+ let bytes = 0;
+ const kept: string[] = [];
+ const hasTwin = async (rel: string, st: Stats): Promise<boolean> => {
+ let twin: Stats;
+ try {
+ twin = await lstat(path.join(ssd, rel));
+ } catch {
+ return false;
+ }
+ if (st.isDirectory()) return twin.isDirectory();
+ return !twin.isDirectory() && twin.size === st.size;
+ };
+ const take = async (p: string, st: Stats): Promise<void> => {
+ const t = await measureEntry(p, st);
+ files += t.files;
+ bytes += t.bytes;
+ if (!ctx.dryRun) await rm(p, { recursive: true, force: true });
+ };
+ const top = await readdir(mediaDir, { withFileTypes: true });
+ for (const e of top) {
+ const p = path.join(mediaDir, e.name);
+ if (!e.isDirectory()) {
+ const st = await lstat(p);
+ if (await hasTwin(e.name, st)) await take(p, st);
+ else kept.push(e.name);
+ continue;
+ }
+ for (const name of await readdir(p)) {
+ const fp = path.join(p, name);
+ const st = await lstat(fp);
+ if (isTierable(name) && st.isFile()) continue;
+ const rel = `${e.name}/${name}`;
+ if (isLeftOnPlatter(name) || (await hasTwin(rel, st))) await take(fp, st);
+ else kept.push(rel);
+ }
+ if (!ctx.dryRun) await rmdir(p).catch(() => {});
+ }
+ ctx.log(
+ `${slug}: ${ctx.dryRun ? "would reclaim" : "reclaimed"} ${files} file(s), ` +
+ `${formatBytes(bytes)} of text copies and dead temps from ${mediaDir}`,
+ );
+ if (kept.length > 0) {
+ ctx.log(
+ `${slug}: kept ${kept.length} entr${kept.length === 1 ? "y" : "ies"} on the drive with no ` +
+ `same-size copy on the corpus disk (${kept.slice(0, 5).join(", ")}` +
+ `${kept.length > 5 ? `, … ${kept.length - 5} more` : ""})`,
+ );
+ }
+ return { reclaimedFiles: files, reclaimedBytes: bytes, reclaimKept: kept };
+}
+
+// --reclaim on a channel already on the media tier: only one this tool
+// migrated (its reclaim note stands), whose media drive answers and whose
+// `data/` is a real directory.
+async function reclaimChannel(
+ ctx: Ctx,
+ slug: string,
+ mediaDir: string,
+): Promise<Pick<ChannelReport, "reclaimedFiles" | "reclaimedBytes" | "reclaimKept">> {
+ const note = path.join(ctx.paths.channelsDir, slug, RECLAIM_NOTE);
+ if ((await linkOrDirState(note)).kind === "missing") {
+ throw new Refusal(
+ `cannot reclaim: no ${RECLAIM_NOTE} beside its config — migrate-tier did not ` +
+ `migrate it, or its reclaim is done; its drive holds no text copies of the ` +
+ `retired layout to take`,
+ );
+ }
+ const loc = await inspectChannelMedia(ctx.paths, slug, undefined, { fresh: true });
+ if (loc.status !== "ok") {
+ throw new Refusal(
+ `cannot reclaim: its media reads ${loc.status}${loc.detail ? ` (${loc.detail})` : ""}`,
+ );
+ }
+ if (!(await isRealDir(path.join(ctx.paths.channelsDir, slug, "data")))) {
+ throw new Refusal(`cannot reclaim: its data/ is not a real directory`);
+ }
+ if (ctx.dryRun) return reclaimTree(ctx, slug, mediaDir);
+ await writeMarker(ctx, slug, {
+ target: mediaDir,
+ direction: "out",
+ startedAt: new Date().toISOString(),
+ phase: "reclaim",
+ scope: "tier-migration",
+ });
+ const out = await reclaimTree(ctx, slug, mediaDir);
+ await ctx.checkpoint(slug, "reclaimed");
+ await rm(note, { force: true });
+ await clearDirMarker(relocationMarkerPath(ctx.paths, slug));
+ forgetChannelMedia(slug);
+ return out;
+}
+
+// ---------------------------------------------------------------------------
+// The run
+// ---------------------------------------------------------------------------
+
+// Every channel this tool has work on: on the retired layout (a `data` link or
+// a recorded `dataDir`), or carrying its own marker.
+export async function legacyChannels(paths: Paths): Promise<string[]> {
+ const entries = await readdir(paths.channelsDir, { withFileTypes: true }).catch(() => []);
+ const out: string[] = [];
+ for (const e of entries) {
+ if (!e.isDirectory()) continue;
+ const slug = e.name;
+ const marker = await readRelocationMarker(paths, slug);
+ if (marker?.scope === "tier-migration") {
+ out.push(slug);
+ continue;
+ }
+ const config = await readChannelConfig(paths, slug).catch(() => null);
+ if (!config) continue;
+ const data = await linkOrDirState(path.join(paths.channelsDir, slug, "data"));
+ if (config.dataDir?.trim() || data.kind === "link") out.push(slug);
+ }
+ return out.sort();
+}
+
+// Channels this tool migrated whose drive still holds the text copies (their
+// reclaim note stands), for `--all --reclaim`.
+async function reclaimableChannels(paths: Paths): Promise<string[]> {
+ const entries = await readdir(paths.channelsDir, { withFileTypes: true }).catch(() => []);
+ const out: string[] = [];
+ for (const e of entries) {
+ if (!e.isDirectory()) continue;
+ const note = await linkOrDirState(path.join(paths.channelsDir, e.name, RECLAIM_NOTE));
+ if (note.kind !== "missing") out.push(e.name);
+ }
+ return out.sort();
+}
+
+// The text bytes a legacy channel would copy, for `--order smallest`; a
+// channel whose tree cannot be walked sorts last (and is refused in its turn).
+async function copyBytesOf(paths: Paths, slug: string): Promise<number> {
+ const marker = await readRelocationMarker(paths, slug);
+ if (marker?.scope === "tier-migration") return -1; // resume these first
+ const config = await readChannelConfig(paths, slug).catch(() => null);
+ const data = path.join(paths.channelsDir, slug, "data");
+ const D =
+ config?.dataDir?.trim() ||
+ (await readlink(data).then((t) => path.resolve(path.dirname(data), t)).catch(() => ""));
+ if (!D) return Number.MAX_SAFE_INTEGER;
+ try {
+ return (await inventoryTree(D)).copyBytes;
+ } catch {
+ return Number.MAX_SAFE_INTEGER;
+ }
+}
+
+function describe(r: ChannelReport): string {
+ const parts: string[] = [`${r.slug}: ${r.outcome}`];
+ if (r.detail) parts.push(`— ${r.detail}`);
+ return parts.join(" ");
+}
+
+export async function migrateMediaTier(opts: MigrateOptions): Promise<MigrateResult> {
+ const log = opts.log ?? ((l: string) => console.log(l));
+ const deps = opts.deps ?? {};
+ const ctx: Ctx = {
+ paths: opts.paths,
+ settings: opts.settings,
+ dryRun: opts.dryRun === true,
+ reclaim: opts.reclaim === true,
+ log,
+ freeBytes: deps.freeBytes ?? getFreeBytes,
+ checkpoint: async (slug, step) => {
+ await deps.checkpoint?.(slug, step);
+ },
+ projectedUse: 0,
+ };
+ const result: MigrateResult = { exitCode: 0, channels: [] };
+
+ const probe = deps.editorProbe ?? ((p: Paths) => editorRunningReason(p));
+ const running = await probe(opts.paths);
+ result.editorRunning = running;
+ if (running) {
+ if (!ctx.dryRun) {
+ log(
+ `Refusing: the editor must be stopped for the migration — ${running}. ` +
+ `It holds its job registry in memory, so nothing on disk can say a job of ` +
+ `its is about to write into a channel. Stop it and run this again.`,
+ );
+ result.exitCode = 2;
+ return result;
+ }
+ log(`Note: ${running}. A dry run reads only; a real run would refuse until it is stopped.`);
+ }
+ if (ctx.dryRun) log("DRY RUN — nothing will be written.");
+
+ let order: string[];
+ let deferred: string[] = [];
+ // Every legacy channel's copy bytes, from ONE walk each: the order, and the
+ // stop's report on the big three.
+ const sizes = new Map<string, number>();
+ if (opts.slugs === "all") {
+ const legacy = await legacyChannels(opts.paths);
+ const large = legacy.filter((s) => LARGE_TEXT_CHANNELS.includes(s));
+ const rest = legacy.filter((s) => !LARGE_TEXT_CHANNELS.includes(s));
+ const measure = async (slugs: string[]) => {
+ const sized: { slug: string; bytes: number }[] = [];
+ for (const slug of slugs) {
+ const bytes = await copyBytesOf(opts.paths, slug);
+ sizes.set(slug, bytes);
+ sized.push({ slug, bytes });
+ }
+ return sized
+ .sort((a, b) => a.bytes - b.bytes || a.slug.localeCompare(b.slug))
+ .map((s) => s.slug);
+ };
+ log(`${legacy.length} channel(s) on the retired layout: ${legacy.join(", ") || "none"}`);
+ order = await measure(rest);
+ const largeOrder = await measure(large);
+ if (opts.includeLarge) order.push(...largeOrder);
+ else deferred = largeOrder;
+ if (ctx.reclaim) {
+ for (const slug of await reclaimableChannels(opts.paths)) {
+ if (!order.includes(slug)) order.push(slug);
+ }
+ }
+ } else {
+ order = opts.slugs;
+ }
+
+ for (const slug of order) {
+ let report: ChannelReport;
+ try {
+ report = await migrateOne(ctx, slug);
+ } catch (err) {
+ const refusal = err instanceof Refusal;
+ report = {
+ slug,
+ outcome: refusal ? (err as Refusal).outcome : "failed",
+ detail: (err as Error).message,
+ };
+ }
+ result.channels.push(report);
+ log(describe(report));
+ if (report.outcome === "migrated" || report.outcome === "already") {
+ if (report.copy) {
+ log(
+ ` ${slug}: copied ${formatBytes(report.copyBytes ?? 0)} (${kindsLine(report.copy)}); ` +
+ `${report.linksMade ?? 0} link(s) made; ${report.tierableFiles ?? 0} media file(s), ` +
+ `${formatBytes(report.tierableBytes ?? 0)} stay on the media drive` +
+ (report.leftOnPlatter?.files
+ ? `; ${tallyText(report.leftOnPlatter)} of dead temps left there for --reclaim`
+ : ""),
+ );
+ }
+ }
+ const bad = report.outcome === "refused" || report.outcome === "failed" || report.outcome === "no-space";
+ if (bad) {
+ result.exitCode = 1;
+ // A real run stops at the first channel it could not finish: the
+ // operator decides. A dry run reports every channel.
+ if (!ctx.dryRun) {
+ if (opts.slugs === "all") {
+ log(
+ `Stopped at ${slug}. Channels done so far are migrated; run ` +
+ `\`archilyzer storage migrate-tier --all\` again after fixing it (finished ` +
+ `channels are a no-op).`,
+ );
+ }
+ return result;
+ }
+ }
+ }
+
+ if (deferred.length > 0) {
+ result.stoppedBefore = deferred;
+ const free = await ctx.freeBytes(opts.paths.channelsDir);
+ result.freeBytes = free;
+ log("");
+ log(
+ `STOPPED before the big-text channels: ${deferred.join(", ")}. The corpus disk has ` +
+ `${formatBytes(free - (ctx.dryRun ? ctx.projectedUse : 0))} free${ctx.dryRun ? " (after the channels above)" : ""}; ` +
+ `the gate floor is ${opts.settings.minFreeDiskGB} GB.`,
+ );
+ for (const slug of deferred) {
+ const bytes = sizes.get(slug) ?? (await copyBytesOf(opts.paths, slug));
+ log(
+ ` ${slug}: ${
+ bytes < 0
+ ? "a tier migration is in flight — resume it by name"
+ : bytes === Number.MAX_SAFE_INTEGER
+ ? "its old tree could not be walked"
+ : `${formatBytes(bytes)} to copy`
+ }`,
+ );
+ }
+ log(
+ `Continue with \`archilyzer storage migrate-tier <slug>\` one at a time, or ` +
+ `\`archilyzer storage migrate-tier --all --include-large\`.`,
+ );
+ }
+ return result;
+}
+
+// ---------------------------------------------------------------------------
+// The command line
+// ---------------------------------------------------------------------------
+
+const USAGE =
+ "usage: archilyzer storage migrate-tier <slug>… | --all [--order smallest] " +
+ "[--include-large] [--dry-run] [--reclaim]";
+
+export async function main(argv: string[]): Promise<number> {
+ const BOOLS = ["all", "dry-run", "reclaim", "include-large", "help"];
+ const { flags, positionals } = parseArgv(argv, BOOLS);
+ if (flags.help) {
+ console.log(USAGE);
+ return 0;
+ }
+ const known = new Set([...BOOLS, "order"]);
+ const unknown = Object.keys(flags).filter((k) => !known.has(k));
+ if (unknown.length > 0) {
+ console.error(`migrate-tier: unknown flag(s) ${unknown.map((k) => `--${k}`).join(", ")}\n${USAGE}`);
+ return 2;
+ }
+ const all = flags.all === true;
+ if (all === positionals.length > 0) {
+ console.error(`migrate-tier: name channel slugs or --all, not both and not neither\n${USAGE}`);
+ return 2;
+ }
+ if (flags.order !== undefined && (flags.order !== "smallest" || !all)) {
+ console.error(`migrate-tier: --order takes "smallest", with --all\n${USAGE}`);
+ return 2;
+ }
+ if (flags["include-large"] && !all) {
+ console.error(`migrate-tier: --include-large goes with --all (a named channel always runs)\n${USAGE}`);
+ return 2;
+ }
+ const { getPaths } = await import("../lib/paths");
+ const { getSettings } = await import("../lib/settings");
+ const paths = getPaths();
+ const settings = getSettings();
+ console.log(`channels: ${paths.channelsDir}`);
+ const res = await migrateMediaTier({
+ paths,
+ settings,
+ slugs: all ? "all" : positionals,
+ dryRun: flags["dry-run"] === true,
+ reclaim: flags.reclaim === true,
+ includeLarge: flags["include-large"] === true,
+ });
+ return res.exitCode;
+}
+
+runIfEntryPoint(import.meta.url, () => main(process.argv.slice(2)));
diff --git a/common/lib/channelConfig.ts b/common/lib/channelConfig.ts
@@ -168,7 +168,7 @@ export const CHANNEL_CONFIG_FIELD_DOCS: FieldDocs<ChannelConfig> = {
savedVideosDir:
"Per-channel override for the saved-video store root: this channel's persisted source videos live under `<savedVideosDir>/<slug>/<videoId>/`. Trimmed; blank = the global store.",
dataDir:
- "RETIRED (release 17). The whole-directory layout's record: the absolute path `channels/<slug>/data` was a symlink to. Still parsed for one release so a write never erases it: a channel that carries it — or whose `data/` is a link — is `legacy`, and every media job, lane and build holds it until `archilyzer storage migrate-tier <slug>` moves its text back and its media into `mediaDir`. Never written by anything but that migration, which removes it.",
+ "RETIRED (release 17). The whole-directory layout's record: the absolute path `channels/<slug>/data` was a symlink to. Still parsed for one release so a write never erases it: a channel that carries it — or whose `data/` is a link — is `legacy`, and every media job, lane and build holds it until `archilyzer storage migrate-tier <slug>` moves its text back and its media into `mediaDir`. Written only by the re-point of a storage location (which keeps it `<root>/<slug>/data` on the new root); removed by that migration.",
mediaDir:
"Where this channel's big files live when relocated: `channels/<slug>/media` is a symlink to it, `<root>/<slug>/media`. Absent = in place. Written only by relocate / re-point / the tier migration — a record of what is on disk, never free text, because a value that disagrees with the link is an \"inconsistent\" channel every media guard refuses. The text (`data/`) never moves.",
ytdlpExtraArgs: "Extra yt-dlp arguments, appended verbatim. Must be an array of strings or it is dropped.",
diff --git a/common/lib/envVars.ts b/common/lib/envVars.ts
@@ -118,7 +118,7 @@ const DECLARED: EnvVarDecl[] = [
{ name: "ARCHILYZER_INDEX_ALLOW_HELD", audience: "runtime", default: "off", readBy: "common/controller/buildIndex.ts", doc: "`1` lets a FULL index rebuild (a schema change, or no index yet) proceed while a channel's media cannot be read; that channel stays out of the index until its media is back and the index is built again. Unset, such a build refuses and names each channel." },
{ name: "UV_THREADPOOL_SIZE", audience: "runtime", default: "`16` for the editor (`4` is Node's own)", readBy: "Node's libuv (set by editor/package.json `start` and docker/entrypoint.sh)", doc: "Threads in Node's pool for filesystem calls. A call on a stalled drive holds one until the drive answers, so the editor starts with 16. It buys time for calls already in flight and isolates nothing: the storage health probe and its gate keep new calls off a stalled drive." },
{ name: "MCP_IO_STATS", audience: "runtime", default: "off", readBy: "common/lib/archive/io-stats.ts", doc: "`1` turns on per-call I/O accounting, for `mcp/bench`." },
- { name: "ARCHILYZER_EDITOR_URL", audience: "runtime", default: "`http://localhost:3001`", readBy: "scripts/archilyzer-ops.mjs, mcp/src/fetchClip.ts, umtool", doc: "Which editor `pnpm ops` and the MCP's `fetch_clip` talk to." },
+ { name: "ARCHILYZER_EDITOR_URL", audience: "runtime", default: "`http://localhost:3001`", readBy: "scripts/archilyzer-ops.mjs, mcp/src/fetchClip.ts, umtool, common/bin/migrate-media-tier.ts", doc: "Which editor `pnpm ops` and the MCP's `fetch_clip` talk to, and whose `/api/pulse` `archilyzer storage migrate-tier` asks before it refuses to run beside it." },
{ name: "ARCHILYZER_AGENT", audience: "runtime", default: "`cli`", readBy: "scripts/archilyzer-ops.mjs", doc: "Who is asking, recorded as the provenance of a curated-tag write through `pnpm ops`." },
{ name: "DIARIZE_ENGINE_KIND", audience: "runtime", default: "`sherpa-onnx`", readBy: "scripts/diarize.mjs", doc: "The diarization engine: `sherpa-onnx` or `sortformer`." },
{ name: "DIARIZE_ENGINE_CMD", audience: "runtime", default: "the bundled sherpa script", readBy: "scripts/diarize.mjs", doc: "The engine command the wrapper runs." },
diff --git a/common/lib/mediaTier.test.ts b/common/lib/mediaTier.test.ts
@@ -45,6 +45,7 @@ const TABLE: ReadonlyArray<[string, "media" | "text" | "scratch", boolean]> = [
[".audio.mp3.parakeet", "scratch", false],
[".audio.mp3.tierlink-4242", "scratch", false],
[".audio.mp3.tiering-4242", "scratch", false],
+ [".syncthing.audio.mp3.tmp", "scratch", false],
// The hot text.
["transcript.json", "text", false],
["transcript.en.vtt", "text", false],
diff --git a/common/lib/mediaTier.ts b/common/lib/mediaTier.ts
@@ -63,6 +63,9 @@ const SCRATCH_PATTERNS: ReadonlyArray<RegExp> = [
// beside the name, and the bytes being placed in `media/<id>/`.
/\.tierlink-\d+$/,
/\.tiering-\d+$/,
+ // Syncthing's in-flight temp (`.syncthing.audio.mp3.tmp`): a transfer the
+ // sync tool never finished.
+ /^\.syncthing\..*\.tmp$/,
];
export function isScratchEntry(name: string): boolean {
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -35,6 +35,7 @@
- **A channel's text stays on the fast disk when its media moves, so a slow or unplugged media drive no longer holds its transcripts.** A channel's big files — the audio and the raw live-chat replay — can now live in the channel's own `media` folder, on this disk or another, while its transcripts, cues, metadata and every other small file stay in `data/` where they always were; each big file that moves leaves a small link behind, so everything that opens it by name still finds it. New downloads, transcodes and live-chat normalizes put their big files there as they finish, and every cleanup that deletes audio removes the file the link points to, not just the link. What that changes when a media drive is stalled, unplugged or mid-move: **the index and stats builds never wait on it or are held by it** (a live chat whose transcript cues are out of date keeps the cues the last build read until the drive answers), the channel's report still refreshes (its media size reads as unknown until the drive answers), **digests keep running — even during a move of that channel's media** — and so do normalize, the availability checks, the metadata scan and clip eviction. Transcription, downloads, the backfill lane and anything else that opens the audio are held as before. **A channel moved the old way — its whole `data/` on the other drive — is now shown as "Media layout retired" and held by everything, the builds and digests included, until `archilyzer storage migrate-tier <channel>` brings its text home;** every refusal says so. Deleting a video from its page is refused while its channel's media drive is not reachable, so its audio is never left behind on the drive. Needs a rebuild and restart of the editor.
- **Moving a channel's media now moves only its big files.** The Storage panel's **Move media** copies the channel's audio and raw live-chat replays to `<root>/<channel>/media` on the destination and leaves its transcripts, metadata and every other small file in `data/` on this disk; a channel whose audio is still in `data/` has it put into the channel's `media` folder on this disk first, and the preview says how many files ("2 file(s) tiered first"). **Move back in place** brings the `media` folder home; the links in `data/` are not touched either way. The panel shows the media path and the text path, each with its size. While a channel's media is held — moving, or on a drive that is unplugged or not answering — its text stays readable: the video page and the videos list open as usual and list the audio with its size unknown, the transcript opens, the audio answers "not reachable, try again" (503) instead of "not found", and a digest may run, during a move too; only the jobs and lanes that open the audio wait. A move, its preview and **Resume move** refuse a channel still on the retired whole-directory layout and name `archilyzer storage migrate-tier`; `/storage` counts such channels as unreachable "(n to migrate)". On `/storage` a location's figure is now the media on it, and the corpus volume's row adds every channel's text and clip windows ("text … + clips … on the corpus volume, plus … media of in-place channels"). Re-pointing a location, renaming a channel and deleting one follow the `media` folder; a re-point also carries a channel not yet migrated (its whole `data/` link), so a drive holding both kinds still follows its disk, while renaming such a channel is refused until it is migrated, and **Clear marker** never removes an interrupted migration's marker. Needs a rebuild and restart of the editor.
- **An export build no longer reads a raw live-chat replay from a media drive that is unplugged, not answering or mid-move.** It publishes the live-chat cues already on disk for that channel, and its log says how many are older than their replay and how many videos with a replay and no cues yet were left out of that build.
+- **`archilyzer storage migrate-tier` brings a channel moved the old way onto the media tier: its text comes home to this disk, its big files stay on the drive.** Run it with the editor **stopped** — a real run refuses while the editor answers on its port (`ARCHILYZER_EDITOR_URL`, default `http://localhost:3001`) or while a job on disk belongs to a live process; a dry run only notes it. `archilyzer storage migrate-tier <channel>` does one channel, `--all` every channel still shown as "Media layout retired". Start with `--dry-run`: it writes nothing and prints, per channel, how much text, clip windows, scratch and source video would be copied to this disk, how many audio files and live-chat replays stay on the drive, and whether it fits. A real run copies everything but the audio, the raw live-chat replays and any dead `*.temp.*` left by a download's postprocessor into the channel's `data/` on this disk, checks the copy file by file and kind by kind, links each audio file and replay back under its old name with its own time, renames the drive's `<root>/<channel>/data` to `<root>/<channel>/media`, and records `mediaDir` in the channel's `config.json` in place of `dataDir`; then it prints the bytes copied by kind and the links made. Nothing needs re-indexing: every file keeps its time. It needs the copy plus the resume margin free above the disk gate's floor on this disk, and refuses before writing anything when that is not there. `--all` takes the smallest channels first and **stops before omnibased, rekietalaw and the-quartering-rumble**, printing the free space and what each of them would copy; go on with the channel's name, or `--all --include-large`. A run that is stopped or fails partway carries on where it left off when run again, and a channel already done is left alone. The drive keeps its copy of the text, and those dead temps, until you run it again with `--reclaim`, which deletes them and keeps the media. `--reclaim` takes only channels this command migrated, and deletes a copy only where the same file, at the same size, is in the channel's `data/` on this disk; anything else stays on the drive and is listed. Leave it until the editor has run on the migrated channels for a while: until then the drive's copy is a second one. **Superseded auto-subtitles are purged after the migration, not before:** a channel on the retired layout refuses `purge-superseded-auto-subs` like every text job, so a channel's superseded `en-orig` subtitles (7.6 GB on omnibased) are copied to this disk first and purged from there.
- **A video whose YouTube subtitles answer "Too Many Requests" (HTTP 429) is downloaded anyway, and YouTube is not put in a cooldown for it.** YouTube refuses a subtitle file per video while the video itself downloads fine; every one of the day's 429s on 2026-10-01 was a subtitle fetch, and each failed its download, put all of YouTube in a cooldown that reached 30 minutes, and deferred the video for 6 hours to fail the same way again. Now that refusal is noted and the download goes on to the audio, as for a video with no captions, so the transcription lane transcribes it; nothing platform-wide is backed off. The video's subtitles are deferred: **Download missing subs** skips them for 6 hours, and from the third time they are refused, for 7 days; a download from the video's own page still fetches them. The download lane's page lists them under **Deferred subtitles**, and the video page says how many times and when. **Download missing subs** also goes on to the next video when one video's subtitles are refused, instead of stopping. Any other subtitle failure (a 403, a missing file, a chat replay that fails) still fails the download, as before. Needs a rebuild and restart of the editor.
- **The download pace adapts to rate limits, a rate limit that outlasts the cooldown holds the platform, and the auto-download lane waits between downloads.** Every yt-dlp run against a platform now waits its platform's current pace between requests: 1 second for YouTube and Rumble, doubled by each real rate limit (up to 16 seconds) and eased back one step after every 5 clean downloads, and one step for every hour with no rate limit; a subtitle-only refusal never raises it. When a platform has failed three times in a row at the 30-minute cooldown, it is **held**: auto-download tries it once an hour instead of every 30 minutes, and until that try is due a manual **Sync**, download or metadata scan on it is refused with a sentence giving its time. A clean try lifts the hold — the lane's, or a manual Sync, download or scan once the try is due, which is how a hold ends while auto-download is off or has nothing to fetch on that platform — and so does a try whose video came down although its subtitles were refused. The lane page keeps a held platform listed until then (saying when the lane is off), with a **Clear hold** button that drops the hold, the cooldown and the raised pace at once and says so in a job log. The auto-download lane now waits **Sleep between downloads** between two downloads on one platform, as a channel's batch downloads always did, plus whatever the pace was raised by; batch downloads add that too. The lane page's **Rate-limit cooldown** box shows held platforms, the raised paces and the deferred subtitles, the lane says when it is idle because a platform is held or it is pausing between downloads, and `archilyzer doctor` warns about a platform in a cooldown or held, and about a raised pace. **Download missing subs** also waits that gap between videos. The four numbers are the new `pacing` block in `settings.json` (SETTINGS.md). **After the restart, YouTube may be held at its first real failure:** its cooldown count from before the update (the subtitle refusals) still stands, so one failure puts it straight past the cap — **Clear hold** on the download lane's page resets it, and a clean download does too. Needs a rebuild and restart of the editor.
diff --git a/plans/FACTS.md b/plans/FACTS.md
@@ -3983,6 +3983,65 @@ Supersedes, where they differ, the mover and storage-location facts above (the w
- **`noCorpusWalkInRenderPaths.test.ts` greps render-path files for the walkers' NAMES**, comments
included: naming `measureTree` in a comment of a file a page imports fails it.
+### A channel's media is tiered — the model, the hook, the guards, the migration (release 17, verified 2026-10-02)
+
+The whole of release 17's storage model in one place; the T1 and T2 sections above carry the detail
+and the measurements.
+- **The model.** `channels/<slug>/data/` is ALWAYS a real directory on the corpus disk and holds the
+ text: transcripts, cues, `metadata.info.json`, every sidecar, `clips/` (never tiered), scratch, and
+ `source-media.*` (the saved-video store renames it, so it is never a link). A tierable file
+ (`isTierable`: `audio.<ext>`, `transcript.live_chat.json`) may be a RELATIVE link
+ `data/<id>/<name> -> ../../media/<id>/<name>` carrying the file's times. `channels/<slug>/media` is a
+ real dir (tiered in place), ONE absolute link to `<root>/<slug>/media` with `config.mediaDir`
+ (relocated), or absent (classic). Classification is BY NAME (`lib/mediaTier.ts`, pure).
+- **The hook** (`lib/mediaTier-server.ts`): `tierMediaFile`/`tierVideoDir`/`tierChannelMedia` after every
+ media finalisation, never throwing; `removeMediaFile`/`removeVideoDirMedia` for every delete. It writes
+ nothing while a marker stands.
+- **The guards** (`lib/channelMedia.ts`): `assertChannelMediaReachable` for a kind with `needsMedia`,
+ `assertChannelTextReadable` for one with `needsText`; statuses `in-place | ok | unreachable |
+ in-transition | inconsistent | stalled | legacy`. A `media`-scoped marker holds media writers only; a
+ `tier-migration` marker holds the text too.
+- **The mover** (`controller/relocateChannelMedia.ts`) moves `media/` only, tiering a classic channel
+ first; it refuses a `legacy` channel and a `tier-migration` marker with `tierMigrationRefusal`.
+- **`legacy`** = `config.dataDir` recorded or `data/` a link: the pre-release-17 whole-directory move.
+ T2's re-point and rename keep `dataDir` = `<root>/<slug>/data` (rename refuses a legacy channel), so the
+ migration may derive everything from that shape.
+- **The migration** (`common/bin/migrate-media-tier.ts`, `archilyzer storage migrate-tier
+ <slug>…|--all`, release 17 slice T3). Editor STOPPED: a real run refuses while
+ `<ARCHILYZER_EDITOR_URL>/api/pulse` answers (a timeout counts as running) or a `running`/`queued` meta in
+ `.jobs/` belongs to a live pid that is not its own (a pid-less `running` meta written after the
+ machine booted counts too; files older than the boot are not read); a dry run only notes it and writes
+ nothing. Per channel: the plan reads the corpus disk (config, `data` link, marker) — a marker whose
+ `scope` is not `tier-migration` is refused, `dataDir` must be `<root>/<slug>/data` and agree with the
+ link, `mediaDir` must be absent; preflight: the old tree is a directory, `assertRelocationRootPresent`,
+ `<root>/<slug>/media` and `channels/<slug>/media` absent, one walk classifying every video dir's entries
+ (a tierable REGULAR file stays; a postprocessor's dead `*.temp.*` — `isLeftOnPlatter` — stays too,
+ unlinked; everything else is listed), the space rule `copy bytes − bytes already
+ in data.incoming + margin ≤ free(channelsDir) − minFreeDiskGB` (margin 0 when the gate is off, the
+ mover's rule); copy: marker `{target: <root>/<slug>/media, direction: "out", phase: "copy", scope:
+ "tier-migration"}`, the NUL list at `channels/<slug>/.tier-migration.files`, `rsync -a -r --from0
+ --files-from=… --partial --info=progress2` into `channels/<slug>/data.incoming/` (`-r` explicitly:
+ `--files-from` turns off `-a`'s recursion and `clips/` is a directory), verify = the same list's
+ `--dry-run --itemize-changes` empty (a `.d..t` directory-time line is not content) AND per-kind
+ (`text | clips | scratch | source`) files and bytes equal, then a link per tierable file
+ (EEXIST = already, if it is the same link) with `lutimes` to the file's times; marker → `swap`;
+ swap: platter rename `data → media`, the `media` link, unlink the old `data` link, rename
+ `data.incoming → data`, `patchChannelConfig(slug, { mediaDir }, { unset: ["dataDir"] })` — each step
+ `linkOrDirState`s first, so a rerun after a kill at any point finishes; the swap leaves
+ `channels/<slug>/.tier-migration.reclaim` (the reclaim note); `--reclaim` (marker phase `reclaim`, only
+ on a channel carrying the note — a named slug without it is refused, `--all --reclaim` takes only those):
+ in `<root>/<slug>/media/<id>/` a tierable regular file stays, a dead `*.temp.*` goes, and any other entry
+ goes only when one `lstat` finds its twin at `channels/<slug>/data/<id>/<name>` (a dir for a dir, else the
+ same size) — the rest is kept and listed; an emptied dir goes; the note is removed. Done: marker cleared,
+ the channel must then inspect `ok`. A channel on the new layout is a no-op. `--all` = every legacy channel plus any carrying a
+ `tier-migration` marker (those first), smallest copy first, the three big-text channels
+ (`LARGE_TEXT_CHANNELS`: omnibased, rekietalaw, the-quartering-rumble) deferred with the free space and
+ each one's copy bytes (from the ordering walk) printed unless `--include-large`; a real run stops at the first channel it cannot
+ finish, a dry run reports every channel and projects the space the earlier ones would take. No index
+ rebuild: `buildIndex` over a migrated fixture reports 0 added, 0 changed, 0 removed
+ (`bin/migrate-media-tier.test.ts`). `purge-superseded-auto-subs` is `needsText`, so it is refused on a
+ legacy channel and runs AFTER the migration. `.syncthing.<name>.tmp` is scratch (`lib/mediaTier.ts`).
+
## Channel priority (verified 2026-09-11) — one tier per channel, four compiled trees
Branch `channel-priority/s5`, off S0's `28bfee3`, merging `s1`–`s4` and closing the twelve
diff --git a/plans/release-17.md b/plans/release-17.md
@@ -1913,4 +1913,176 @@ stay placed by their retired `dataDir`; the digest job stands for the digest lan
exit 0, 44 s; e2e (`channels-storage-columns`, `bulk-actions`, `maybe-missing`, `fetch-window`)
**22 passed, 0 failed, 1.7 min**.
+### Slice T3, as shipped — the one-off migration, and the records (2026-10-02)
+
+Branch `r17/media-tier-migrate` off `main` `b6d9c1e2` (U1, U2, XP, D0, T1, RL and T2 merged), worktree
+`~/Projects/r13-build-image` (editor 5201, test 5211, export 5210), one Opus implementer. Scratch files
+`T3-*` in the job's `tmp`. The plan is "The one-off migration" above and the T3 row; T1 left the
+`lutimes` rule, T2 left `relocatedDataDir`, the `<root>/<slug>/data` invariant and `tierMigrationRefusal`.
+
+**What it does.**
+- **`archilyzer storage migrate-tier <slug>…|--all [--order smallest] [--include-large] [--dry-run]
+ [--reclaim]`** (`common/bin/migrate-media-tier.ts`, one `script([...])` row in `archilyzer.ts` beside
+ `migrate channel-priority`) brings a `legacy` channel — `data/` an absolute link to `<root>/<slug>/data`,
+ `config.dataDir` — onto the media tier: its text is copied home and its big files stay on the drive.
+- **Editor stopped.** A real run refuses (exit 2, nothing read further) while `<ARCHILYZER_EDITOR_URL,
+ default http://localhost:3001>/api/pulse` answers — any HTTP answer, or a connection that does not
+ answer within 3 s — or while `.jobs/` holds a `running`/`queued` meta whose `pid` is alive and is not
+ its own (a pid-less `running` meta written after the machine booted counts too; meta files older than
+ the boot are not read). `migrate channel-priority` checked nothing; this is new. A dry run only notes it.
+- **Per channel**, exactly the plan's phases, each refusing before it writes: the plan reads the corpus
+ disk (a marker whose `scope` is not `tier-migration` is refused naming the Storage panel; `dataDir` must
+ agree with the `data` link and be `<root>/<slug>/data`; `mediaDir` must be absent); preflight — the old
+ tree is a directory, `assertRelocationRootPresent(root)`, `<root>/<slug>/media` and
+ `channels/<slug>/media` absent, one walk classifying every video dir (`inventoryTree`: a tierable
+ REGULAR file stays, everything else is listed and measured as `text | clips | scratch | source`), the
+ space rule `copy − already in data.incoming + margin ≤ free(channelsDir) − minFreeDiskGB` (margin 0 with
+ the gate off, the mover's rule); `copy` — the marker, the NUL list (`channels/<slug>/.tier-migration.files`),
+ `rsync -a -r --from0 --files-from=… --partial --info=progress2` into `data.incoming/` with the mover's
+ decile progress lines, the verify (the same list's `--dry-run --itemize-changes` empty, a `.d..t`
+ directory-time line not counted, AND per-kind files and bytes equal), then one relative link per
+ tierable file with `lutimes` to its times (EEXIST = already when it is the same link); `swap` — the
+ platter rename, the `media` link, the old `data` link unlinked, `data.incoming → data`,
+ `patchChannelConfig(slug, { mediaDir }, { unset: ["dataDir"] })`, every step `linkOrDirState` first;
+ `reclaim` (only `--reclaim`, or a reclaim marker being resumed) — every entry of
+ `<root>/<slug>/media/<id>/` that is not a tierable regular file goes, an emptied dir goes; done — the
+ marker cleared, the channel must then inspect `ok`, and the bytes by kind, links made and media left on
+ the drive are printed.
+- **Resumable and idempotent.** A `tier-migration` marker is resumed from its phase (`copy` redoes the
+ copy and verify — rsync sends only what is missing — and finds its links; `swap` and `reclaim` finish);
+ a channel already on the new layout is `already` and writes nothing (with `--reclaim` it reclaims). A
+ dry run writes nothing at all — no marker, list, directory or config — and says what it would do.
+- **`--all`** = every legacy channel plus any carrying a `tier-migration` marker (resumed first), smallest
+ copy first, with the three big-text channels (`LARGE_TEXT_CHANNELS`) deferred: the run ends with "STOPPED
+ before the big-text channels: …", the corpus disk's free space, each one's copy bytes, and "Continue with
+ `archilyzer storage migrate-tier <slug>` one at a time, or `archilyzer storage migrate-tier --all
+ --include-large`" (exit 0). A real run stops at the first channel it cannot finish (exit 1, "run --all
+ again after fixing it"); a dry run reports every channel and subtracts the space the earlier ones would
+ take.
+- **`clips/`** is `text` by the classifier (`classifyEntry("clips")`, pinned) and is carried to the corpus
+ disk with the text, its own kind in the counts.
+
+**Deviations from the plan** (one sentence each):
+1. The copy carries `source-media.*` (media but never tierable) with the text, as kind `source`: the plan's
+ "text + scratch" would have left it on the platter with no link, and `--reclaim` would then delete it.
+2. rsync takes `-r` explicitly: `--files-from` turns off `-a`'s recursion, and `clips/` and a scratch dir
+ are directories to carry whole.
+3. "Status must be `ok`" is read for the retired layout (which `inspectChannelMedia` answers `legacy`,
+ never `ok`, without touching the drive): `dataDir` and the link agree, have the fixed shape, and the old
+ tree is a directory.
+4. A dry run does not refuse a running editor, it notes it: it reads only, and the live dry run may come
+ before the restart.
+5. "A running meta newer than the last boot": nothing on disk records the editor's boot, so the rule is a
+ live writer pid (or a pid-less `running` meta after the machine's boot).
+6. The continuation after the stop is `--all --include-large` or per slug; `--order smallest` is the only
+ order and `--all`'s default (any other value is a usage error).
+7. A reclaim writes its own marker phase (`reclaim`, `scope: "tier-migration"`) so a killed reclaim resumes,
+ and `--all --reclaim` also reclaims every channel already on the new layout (`mediaDir`, no `dataDir`) —
+ one the mover moved has only tierable files there, so nothing is taken.
+8. The NUL list lives beside the marker in the channel dir (the tool runs on the operator's machine, with
+ no job scratch dir), and is removed after the swap.
+
+No helper was added to `mediaTier-server.ts`; the migration uses `relocatedDataDir`, `relocatedMediaDir`,
+`tierLinkTarget`, `channelMediaLink`, `linkOrDirState`, `rsyncTree`, `makeProgressSink`, the marker
+writers, `assertRelocationRootPresent`, `rootOfRelocatedMediaDir` and `patchChannelConfig` as they are.
+
+**Records.** `AGENTS.md`: the section is now "A channel's media may live on another drive" (the `media`
+link, `mediaDir`, `dataDir` retired and `legacy` until `migrate-tier`); "Six things" → seven — the seventh
+exactly as "Order" above states it, the `.relocating.json` bullet gains `scope`, the `clips/` bullet "on
+the SSD, never tiered", and the first two bullets name the `media` link and the text guard. `plans/FACTS.md`:
+"A channel's media is tiered — the model, the hook, the guards, the migration". SETTINGS.md, CHANNEL.md and
+ENVIRONMENT.md: no schema text changed (`docs files --check`, `settings example --check`, `docs env
+--check` all clean), so nothing regenerated. The `[Unreleased]` bullet in `editor/CHANGELOG.md`.
+
+**Commits**
+
+| Commit | What |
+|---|---|
+| `38a12ea3` | `common:` `archilyzer storage migrate-tier` — the migration, the editor check, `--all` with the stop, the CLI row; `bin/migrate-media-tier.test.ts` (12) |
+| `587b0946` | `docs:` AGENTS.md's media section and the seven things; the changelog bullet |
+| this commit | `plans:` this section, FACTS "A channel's media is tiered" |
+
+#### Gates (logs `$T/T3-*.log`)
+
+- **tsc** (all workspaces) clean at `38a12ea3` (the later commits change no code).
+- **common:** **2,667 passed, 0 failed, 0 skipped** (2,655 at T2's tip + this slice's 12).
+- **The migration test** (`bin/migrate-media-tier.test.ts`, real rsync, a tmp "platter" beside a corpus, a
+ legacy channel with text, `audio.mp3`, `audio.opus`, a raw live chat, `clips/`, `source-media.mp4`,
+ `audio.m4a.part`, `audio.tmp-1234.mp3` and a media-less video): **12/12** — `clips/` is text; a dry run
+ writes nothing (a byte/mtime/link snapshot of the corpus and the platter is unchanged) and counts 3
+ tierable, 1 clip, 1 source, 2 scratch; a run migrates (every tierable file a relative link with the file's
+ mtime by `lstat`, readable through it, its bytes on the platter; `source-media.*`, `.part` and the temp
+ stay real with their mtimes; `clips/` carried; no marker, list or `data.incoming` left; `inspectChannelMedia`
+ `ok`, text readable) and the rerun writes nothing; `--reclaim` in the same run and later (`already`, 0 taken,
+ nothing written), and after a plain run (dry count = real count); a media move's marker refused with and
+ without a `scope`, nothing touched; a running editor refuses a real run (exit 2, nothing touched) and a dry
+ run notes it; `editorRunningReason` (refused port, an answer, a timeout, a live pid, a dead pid, a pid-less
+ meta, a meta from before the boot); a kill at each of the eight steps from `copied` to `config-written` —
+ the channel reads `in-transition` with its text held and `--all` would pick it — and the rerun resumes
+ from `copy` or `swap` and ends migrated; the free-space stop (6 GB free, a 5 GB floor and a 2 GB margin:
+ refused, nothing written; 8 GB: migrated); `--all` takes the smaller channel first and stops before
+ `omnibased` with the free space and `--all --include-large`, which then migrates it; and an index built
+ over the classic layout, the channel put on the retired layout and migrated: the live chat's cues read
+ fresh and the next `buildIndex` reports **0 added, 0 changed, 0 removed**, none held.
+- **Editor unit:** 109/109 (no editor code touched). **test:scripts:** not run — no file it covers
+ (`scripts/`, umtool) was touched. **Capped editor build:** not needed — no editor code touched.
+- **CLI smoke** against a scratch corpus in `$T`: usage via `archilyzer storage migrate-tier --help`;
+ `--all --dry-run` notes the running editor and finds no channel; `--all` refuses (exit 2) naming the
+ editor's `/api/pulse` answer; no slug and no `--all`, and `--order biggest`, are usage errors (exit 2). The
+ probe's one GET reached the live editor's `/api/pulse` (read-only); nothing else was pointed at the
+ primary checkout.
+- **e2e:** none named for this slice.
+- **Privacy gate:** 0 added lines carry the user or host name (`git diff main`, counts only; the one file the
+ whole-file grep names is `plans/FACTS.md`, with the same count as on `main`). No identifier ends in the
+ refused parent suffix.
+- **Numbers tool:** none. The live dry run is the parent's.
+
+#### Found and left
+
+- `lib/envVars.ts`'s `ARCHILYZER_EDITOR_URL.readBy` and CHANNEL.md's `dataDir` row were stale; both fixed in
+ the review round (below).
+- The dry run (and `--all`'s ordering) walks every legacy channel's old tree on its drive, one `lstat` per
+ entry, and the stop walks the big three to print what each would copy; a real run walks a channel again in
+ its own preflight. Read-only, but minutes on a platter. (Since the review the stop no longer walks the big
+ three a second time.)
+- The space rule counts `clips/` (nuxanor-kick's 15 GB) as text, as the plan measured it.
+
+#### Operator notes
+
+- **`purge-superseded-auto-subs` runs AFTER the migration, not before** (the plan said before): since T1
+ the kind is `needsText` and a `legacy` channel's text is refused, so the editor refuses the purge on
+ omnibased and HasanAbiVODs3 until they are migrated. omnibased's ≈ 7.6 GB of `en-orig.vtt` lands on the
+ corpus disk first and is purged there.
+- **`transcripts/channels` is a Syncthing folder (paused).** Syncthing syncs a symlink as a link; after the
+ migration the 13 channels' text (≈ 64 GB with clips) is a real tree inside that folder, where today it is
+ behind a `data` link. Nothing moves while the folder is paused; unpausing it would send that text to the
+ peers.
+- `--reclaim` after the editor has run on the migrated channels for a while: until then the drive's text
+ copy is a second copy.
+
+#### Review (SHIP AFTER FIXES) and the fixes
+
+The review (`$T/T3-review.md`) found four LOWs and five NITs. Rulings: `--all --reclaim` is limited to
+channels this tool migrated; a reclaim deletes only what has a same-size twin on the corpus disk; dead
+`*.temp.*` stay out of the copy; a pulse that does not answer within 3 s counts as a running editor (as
+built).
+
+| Finding | Fix |
+|---|---|
+| L1 `--all --reclaim` walked every channel on the media tier | `beaa6a44`: the swap leaves `channels/<slug>/.tier-migration.reclaim` (`RECLAIM_NOTE`); `--all --reclaim` takes only channels carrying it, a named `--reclaim` without it is refused ("migrate-tier did not migrate it, or its reclaim is done"), a reclaim removes it. Tests: a Storage-panel-moved channel is not walked by `--all --reclaim` and is refused by name; a second `--reclaim` refuses and touches nothing; a reclaim killed after its deletes resumes from its marker and removes the note. |
+| L2 a later `--reclaim` deleted without asking whether the corpus disk still has the file | `beaa6a44`: an entry goes only when one `lstat` finds its twin at `channels/<slug>/data/<id>/<name>` (a directory for a directory, else the same size); everything else is kept and listed (`reclaimKept`, and a "kept N entries" line). Test: a removed and a rewritten text file are kept on the platter, the rest taken. |
+| L3 the purge cannot run first | Operator note above and in the changelog bullet. |
+| L4 dead `*.temp.*` carried to the corpus disk (15 GB on nuxanor-kick) | `beaa6a44`: `isLeftOnPlatter` (scratch by the classifier AND `*.temp.*`) is neither copied nor linked; it stays on the platter and `--reclaim` deletes it; the walk reports it ("dead postprocessor temps left there unlinked"). `96191600`: `/^\.syncthing\..*\.tmp$/` joins the classifier's scratch patterns (+ the table row). Tests: the fixture carries `source-media.temp.mp4` (left, then reclaimed) and `.syncthing.audio.mp3.tmp` (scratch, carried). |
+| NIT AGENTS.md's corpus table | `e01db98a`: "a big file in `data/<id>/` may be a relative link into `media/`, which may be an absolute SYMLINK to another drive (and a legacy channel's whole `data/` is one, until migrated)". |
+| NIT `ARCHILYZER_EDITOR_URL.readBy` | `e01db98a`: names `common/bin/migrate-media-tier.ts`; ENVIRONMENT.md regenerated. |
+| NIT the `dataDir` field doc | `e01db98a`: "Written only by the re-point of a storage location …; removed by that migration" (`lib/channelConfig.ts`); CHANNEL.md regenerated. |
+| NIT the stop re-walked the big three | `beaa6a44`: the stop reports the sizes the ordering walk measured (the big three are measured once, in the ordering pass, whether or not `--include-large`). |
+| NIT apparent bytes vs block slack | Left: the 2 GB margin covers it. |
+
+**Gates after the fixes** (logs `$T/T3-test4.log`, `$T/T3-tsc2.log`, `$T/T3-common2.log`): tsc (all
+workspaces) clean; the migration test **15/15** (12 + 3 new); the classifier test 37/37; `docs files
+--check`, `docs env --check`, `settings example --check` clean after the regeneration; **common 2,671 passed, 0 failed, 0 skipped** (2,667 + the three new
+migration cases + the classifier's Syncthing row). Not re-run: editor unit, test:scripts, the editor build
+(no editor, script or umtool file touched). Privacy: 0 added lines carry the user or host name.
+
## Rollout