commit 136c974fe3337f03ddac1718f0f24ea790ab2ad3
parent f196bfe246958aa7125b27f3ea87af3f8d23e3d6
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Sat, 12 Sep 2026 01:12:12 -0400
merge: storage/relocate-media into integrate/2026-09-storage-priority
The relocate branch is the first of the two interlude shipments to land on the
integration branch. It branched at 1939ef9 = f196bfe^, so the only conflict is
plans/STATE.md: relocate's own "Next" (straight to phase 2) against this tip's
sequencing of channel-priority after it. Kept relocate's shipped entry and its
three-step rollout block whole, and kept the channel-priority sequencing in the
Next line — priority is merged second on this branch, so the order the tip
recorded is still the order.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Diffstat:
58 files changed, 5769 insertions(+), 78 deletions(-)
diff --git a/AGENTS.md b/AGENTS.md
@@ -171,7 +171,7 @@ apps*, not [DEPLOY_DOCKER.md](DEPLOY_DOCKER.md), which is about *building sites*
| Path | What it is |
|---|---|
-| `transcripts/channels/<slug>/` | One channel: `config.json`, `playlist`, `archive`, `snapshot.json`, and `data/<videoId>/` holding media, transcripts and sidecars. |
+| `transcripts/channels/<slug>/` | One channel: `config.json`, `playlist`, `archive`, `snapshot.json`, and `data/<videoId>/` holding media, transcripts and sidecars. **`data/` may be an absolute SYMLINK** — see below. |
| `transcripts/sites/<id>/site.json` | **Per-site config, including the public URL.** This is where deployed-site facts live — *not* under `channels/`. |
| `transcripts/index.mdb` | The LMDB transcript index. Key-only range scans over its `byChannel` sub-DB are cheap; see `common/controller/recencyIndex.ts`. |
| `transcripts/saved-videos/` | Persisted source-video store. |
@@ -180,6 +180,30 @@ apps*, not [DEPLOY_DOCKER.md](DEPLOY_DOCKER.md), which is about *building sites*
**The public-URL key in `site.json` is `siteUrl`.** The editor form labels the field
"Public URL", so grepping for `publicUrl` finds the UI hint and misses the data.
+## A channel's `data/` may live on another drive
+
+`channels/<slug>/data` can be an **absolute symlink** to `<root>/<slug>/data` on
+another disk, with `config.dataDir` recording the target. The editor's Storage panel
+(channel page → Storage) moves it; nothing else writes that field. The on-disk contract
+`channelDir/data/<id>/…` is unchanged, so **no reader needs to know** — yt-dlp's
+cwd-relative writes, the LMDB index (it stores mtimes, and `rsync -a` preserves them)
+and the export build all keep working with no call-site changes.
+
+Three things that are not optional:
+
+- **Never symlink a whole channel dir.** Channel listing filters `isDirectory()` on
+ `channelsDir` entries (`channels.ts:246,363`), so a symlinked `<slug>/` vanishes from
+ the corpus. Only `data/` may be a link.
+- **An unmounted drive is not an empty channel.** Every enumerator swallows ENOENT on
+ `data/` as "no videos", which to a runner means *everything is undownloaded*.
+ `common/lib/channelMedia.ts` is the one module that can tell the two apart;
+ `inspectChannelMedia` / `assertChannelMediaReachable` are what the guards call, and a
+ job kind declares `needsMedia` in `common/jobs/jobKinds.ts` to be covered by the one
+ in `runManagedFunction`. If you add a path that reads `data/`, guard it there.
+- **`channels/<slug>/.relocating.json`** is the in-flight marker. Its presence means
+ "media is in transition" to every guard and lets an interrupted move resume from its
+ `phase`. `deleteChannel` and `renameChannel` refuse while it exists.
+
## The live instances
Read these out of `transcripts/sites/*/site.json` rather than hardcoding them — this
diff --git a/RUNNING_IN_DOCKER.md b/RUNNING_IN_DOCKER.md
@@ -377,6 +377,28 @@ docker run --rm -v archilyzer_corpus:/corpus -v "$PWD:/backup" \
debian:bookworm-slim tar czf /backup/corpus.tar.gz -C /corpus .
```
+### A channel whose media is on another drive
+
+The Storage panel can move a channel's `data/` to another root (see AGENTS.md). In a
+container that root is a path **inside the container**, and `channels/<slug>/data`
+becomes an absolute symlink to it — so the drive must be bind-mounted **at the same
+absolute path the editor recorded**:
+
+```yaml
+services:
+ editor:
+ volumes:
+ - /mnt/platter/archilyzer-media:/mnt/platter/archilyzer-media
+```
+
+Mount it somewhere else and the link dangles. That is **reported as unreachable, by
+design** — the channel is skipped by the lane runners, its media jobs are refused and
+its report is not regenerated, rather than the alternative, which is every count on that
+channel reading zero and the download runner treating the whole archive as missing. The
+badge on `/channels` and the channel's Storage panel name the path they cannot reach.
+
+The `site` profile does not need the mount: an export build never reads `data/`.
+
### Useful commands
```sh
diff --git a/WORKTREES.md b/WORKTREES.md
@@ -143,6 +143,17 @@ writes a `.worktree-env` file pointing `TRANSCRIPTS_DIR` at the **main** worktre
`transcripts/`, and `wt run` loads it. This lets a worktree reuse the already-downloaded
corpus for read-mostly work and builds.
+### Copying a corpus that has a relocated channel
+
+`channels/<slug>/data` may be an absolute symlink to another drive (see AGENTS.md).
+`rsync -a` copies a symlink **as a symlink**, which is usually right — both checkouts
+then read the same media through the same absolute path, and nothing is duplicated. Use
+`rsync -a --copy-links` only when the copy has to carry the media itself (a shard bound
+for a machine that will not have that drive). Note what `--copy-links` gives you: a real
+directory where the source had a link, so the copy's `config.dataDir` still names a
+target it is no longer using — `inspectChannelMedia` reports that as `inconsistent`, and
+clearing `dataDir` in the copy's `config.json` is what makes it in-place again.
+
> ⚠️ **Caveat:** the shared LMDB index (`transcripts/index.mdb`) is not safe for concurrent
> **writes**. Use shared mode for reading/building, not for running ingestion (downloads /
> indexing) in two worktrees at the same time — concurrent writers can corrupt the index.
diff --git a/common/controller/autoRunner.ts b/common/controller/autoRunner.ts
@@ -1,5 +1,6 @@
import path from "node:path";
import type { Paths } from "../lib/paths";
+import type { ChannelConfig } from "../lib/channelConfig";
import { mapConcurrent } from "../lib/concurrency";
import { getPaths } from "../lib/paths";
import { getSettings } from "../lib/settings";
@@ -78,6 +79,10 @@ import { type DownloadOutcomeStatus } from "../lib/downloadOutcome";
import { downloadQueueKey } from "../lib/queueKeys";
import { isGateHeld } from "../lib/pauseGates";
import {
+ inspectChannelMedia,
+ type ChannelMediaStatus,
+} from "../lib/channelMedia";
+import {
listChannelConfigs,
readChannelConfig,
readChannelSnapshotShared,
@@ -279,7 +284,14 @@ export function getAutoRunnerStatus(kind: AutoQueueKind): AutoRunnerStatus {
// --- Pending-work construction from snapshots ------------------------------
-type ChannelMeta = { slug: string; platform: ReturnType<typeof detectPlatform> };
+type ChannelMeta = {
+ slug: string;
+ platform: ReturnType<typeof detectPlatform>;
+ // The parsed config, carried so the media guard below does not re-read all 68
+ // config.json files on every tick of four lanes and on every three-second
+ // status poll. listChannelConfigs has already parsed them.
+ config: ChannelConfig;
+};
const SNAPSHOT_READ_CONCURRENCY = 64;
@@ -290,9 +302,41 @@ async function listChannelMeta(paths: Paths): Promise<ChannelMeta[]> {
return configs.map(({ slug, config }) => ({
slug,
platform: detectPlatform(config.url),
+ config,
}));
}
+// ONCE PER STATE CHANGE, not once per tick. buildChannelWork runs on every
+// scheduling tick of all four lanes and on a three-second status poll; logging a
+// skip each time would write the same line thousands of times an hour and bury
+// the one that matters. Keyed by slug, so mounting the drive logs the recovery
+// too — an operator watching the log sees the channel leave and come back.
+const mediaSkipLogged = new Map<string, ChannelMediaStatus>();
+
+function noteSkippedForMedia(
+ slug: string,
+ status: ChannelMediaStatus,
+ detail?: string,
+): void {
+ if (mediaSkipLogged.get(slug) === status) return;
+ mediaSkipLogged.set(slug, status);
+ console.log(
+ `[auto] skipping ${slug}: media ${status}${detail ? ` — ${detail}` : ""}`,
+ );
+}
+
+function noteMediaReachable(slug: string): void {
+ if (!mediaSkipLogged.has(slug)) return;
+ mediaSkipLogged.delete(slug);
+ console.log(`[auto] ${slug}: media reachable again`);
+}
+
+// For tests and for /api/test/invalidate-cache: the log-once memory is
+// process-local state, not a decision, and nothing downstream reads it.
+export function resetChannelMediaSkipLog(): void {
+ mediaSkipLogged.clear();
+}
+
// Read each channel's snapshot and project the buckets this runner kind cares
// about into ChannelWork, plus a videoId -> owning channel map (a platform/all
// leaf spans channels, so the pick needs the owner to locate the video dir).
@@ -323,9 +367,29 @@ async function buildChannelWork(
const snaps = await mapConcurrent(meta, SNAPSHOT_READ_CONCURRENCY, (m) =>
readChannelSnapshotShared(paths, m.slug),
);
+ // GUARD 2 OF FOUR (see plans/relocate-channel-media.md). This is the runners'
+ // ONLY channel-level chokepoint, and it is the one that matters most: the
+ // download lane reading an unmounted channel's snapshot as "everything
+ // undownloaded" is an instruction to re-fetch the whole channel onto the disk
+ // that was too full to hold it. Three syscalls per channel per tick, run
+ // concurrently alongside the snapshot reads.
+ const media = await mapConcurrent(meta, SNAPSHOT_READ_CONCURRENCY, (m) =>
+ // The config is passed, not re-read: without it this is a 69th, 70th …
+ // config.json read per tick for a field listChannelConfigs already parsed.
+ inspectChannelMedia(paths, m.slug, m.config),
+ );
for (const [i, { slug, platform }] of meta.entries()) {
const snap = snaps[i];
if (!snap) continue;
+ // A SKIP IS A SKIP. The lane keeps running every other channel: this is
+ // never a lane stop and it is not a hold — "a zero limit is a hold, never a
+ // stop" (pauseGates.ts) is a different mechanism and is untouched by it.
+ const location = media[i];
+ if (location && location.status !== "ok" && location.status !== "in-place") {
+ noteSkippedForMedia(slug, location.status, location.detail);
+ continue;
+ }
+ noteMediaReachable(slug);
const buckets: Record<string, string[]> = {};
// Project every bucket a leaf could be pointed at — including the opt-in
// auto-caption ones. A leaf that names a bucket the runner never projected
@@ -1056,6 +1120,12 @@ async function runLoop(
// only — transcription writes a transcript.json next to audio it already
// has, so stopping it frees nothing.
if (kind === "download") {
+ // The CORPUS volume, deliberately, and not a channel's. This check runs
+ // before the pick, so there is no channel yet — and the lane spans 68 of
+ // them, on however many volumes. The per-channel answer is taken further
+ // down, in runYtdlp, where the slug is known and the bytes are about to
+ // be written; a full SSD idling the lane here is the conservative half of
+ // the pair, not the whole of it.
const gate = await diskGate(paths, settings);
if (!gate.ok) {
if (!diskIdle) {
diff --git a/common/controller/backfillReacquire.ts b/common/controller/backfillReacquire.ts
@@ -163,7 +163,10 @@ export async function reacquireMediaFor(opts: {
}
const settings = getSettings();
- const gate = await diskGate(opts.paths, settings);
+ // The channel's own volume: a relocated channel re-acquires onto the platter.
+ const gate = await diskGate(opts.paths, settings, {
+ dir: path.join(opts.paths.channelsDir, opts.channelSlug, "data"),
+ });
if (!gate.ok) {
log(`Not re-acquiring ${opts.videoId}: ${gate.message}.`);
return { status: "disk-floor", cleanup: NOTHING_TO_CLEAN };
diff --git a/common/controller/channelSnapshot.test.ts b/common/controller/channelSnapshot.test.ts
@@ -1,5 +1,8 @@
import { test } from "node:test";
import assert from "node:assert/strict";
+import { mkdir, mkdtemp, rm, symlink, writeFile } from "node:fs/promises";
+import { tmpdir } from "node:os";
+import path from "node:path";
import {
attributeAudioHold,
diarizationWillNeverClear,
@@ -7,7 +10,10 @@ import {
emptyHeldAudio,
foldBackfillEntry,
foldBucketLaneEntry,
+ generateChannelSnapshot,
} from "./channelSnapshot";
+import { ChannelMediaUnreachableError } from "../lib/channelMedia";
+import type { Paths } from "../lib/paths";
import {
emptyOperationCounts,
presentOperationWork,
@@ -325,3 +331,30 @@ test("only the reachable diarization states will ever clear a hold", () => {
assert.equal(diarizationWillNeverClear(state), true);
}
});
+
+test("generateChannelSnapshot refuses an unreachable channel rather than writing an empty snapshot", async () => {
+ // The readdir inside it swallows ENOENT as "no videos", so without the guard a
+ // relocated channel whose drive is unmounted would publish a snapshot saying
+ // every video is undownloaded — and all four lanes read that as work to do.
+ // The throw is what makes the scheduler keep the last good snapshot.json.
+ const dir = await mkdtemp(path.join(tmpdir(), "ttb-snap-media-"));
+ try {
+ const paths = { channelsDir: path.join(dir, "channels") } as Paths;
+ const channelDir = path.join(paths.channelsDir, "alpha");
+ await mkdir(channelDir, { recursive: true });
+ const target = path.join(dir, "platter", "alpha", "data");
+ await writeFile(
+ path.join(channelDir, "config.json"),
+ JSON.stringify({ url: "https://example.com/c", dataDir: target }),
+ );
+ // A link with no target: an unmounted drive, exactly.
+ await symlink(target, path.join(channelDir, "data"));
+
+ await assert.rejects(
+ () => generateChannelSnapshot(paths, "alpha"),
+ ChannelMediaUnreachableError,
+ );
+ } finally {
+ await rm(dir, { recursive: true, force: true });
+ }
+});
diff --git a/common/controller/channelSnapshot.ts b/common/controller/channelSnapshot.ts
@@ -18,6 +18,7 @@ import {
type Availability,
} from "../lib/availability";
import { resolveCookiePolicy } from "../lib/cookiePolicy";
+import { assertChannelMediaReachable } from "../lib/channelMedia";
import { getSettings } from "../lib/settings";
import {
loadAvailability,
@@ -603,6 +604,20 @@ export async function generateChannelSnapshot(
const archivePath = path.join(channelDir, "archive");
const playlistPath = path.join(channelDir, "playlist");
+ // GUARD 3 OF FOUR (see plans/relocate-channel-media.md). The readdir below
+ // swallows ENOENT as "this channel has no videos", so a relocated channel
+ // whose drive is not mounted would generate a snapshot saying every video is
+ // undownloaded and every transcript is missing — and everything downstream
+ // (/cleanup, the channel bands, all four lanes' work lists) reads that
+ // snapshot as the truth. Throw instead: the scheduler keeps the last good
+ // snapshot.json on a failed refresh, which is exactly the right outcome.
+ // The config is read HERE rather than in the fan-out below and passed in, so
+ // the guard does not open config.json a second time for the one field it
+ // needs. One extra await in front of a function that then walks the whole
+ // channel.
+ const config = await readChannelConfig(paths, slug);
+ await assertChannelMediaReachable(paths, slug, config);
+
// Heal any video dir that drifted from the canonical id layout before we read
// data/* (best-effort; never fail snapshot generation on a reconcile error).
try {
@@ -616,7 +631,6 @@ export async function generateChannelSnapshot(
urls,
archive,
failedListed,
- config,
maybeMissingRecord,
roster,
] = await Promise.all([
@@ -624,7 +638,6 @@ export async function generateChannelSnapshot(
readPlaylistUrls(playlistPath),
readArchive(archivePath),
loadFailedTranscriptions(paths, slug),
- readChannelConfig(paths, slug),
loadMaybeMissing(paths, slug),
loadRoster(paths, slug),
]);
diff --git a/common/controller/channels.test.ts b/common/controller/channels.test.ts
@@ -0,0 +1,132 @@
+import { test } from "node:test";
+import assert from "node:assert/strict";
+import {
+ mkdir,
+ mkdtemp,
+ rm,
+ stat,
+ symlink,
+ writeFile,
+} from "node:fs/promises";
+import { tmpdir } from "node:os";
+import path from "node:path";
+import type { Paths } from "../lib/paths";
+import type { ChannelConfig } from "../lib/channelConfig";
+import {
+ relocatedDataDir,
+ RELOCATION_MARKER_FILENAME,
+} from "../lib/channelMedia";
+import { channelExists, deleteChannel, writeChannelConfig } from "./channels";
+
+// Run with:
+// pnpm --filter yt-dlp-transcript-common exec tsx --test controller/channels.test.ts
+//
+// Deleting a RELOCATED channel is the one place `rm -r` on the channel dir is
+// not the whole job: node's rm does not follow symlinks, so it takes the link
+// and leaves the media behind as an orphan nothing in the app can see.
+
+const config: ChannelConfig = { handling: "youtube", name: "A channel" };
+
+async function withPaths(
+ fn: (paths: Paths, mediaRoot: string) => Promise<void>,
+): Promise<void> {
+ const dir = await mkdtemp(path.join(tmpdir(), "ttb-channels-"));
+ const paths = { channelsDir: path.join(dir, "channels") } as Paths;
+ const mediaRoot = path.join(dir, "platter");
+ try {
+ await fn(paths, mediaRoot);
+ } finally {
+ await rm(dir, { recursive: true, force: true });
+ }
+}
+
+async function there(p: string): Promise<boolean> {
+ try {
+ await stat(p);
+ return true;
+ } catch {
+ return false;
+ }
+}
+
+test("deleting an in-place channel removes its directory", async () => {
+ await withPaths(async (paths) => {
+ await writeChannelConfig(paths, "alpha", config);
+ await mkdir(path.join(paths.channelsDir, "alpha", "data", "v1"), {
+ recursive: true,
+ });
+ await deleteChannel(paths, "alpha");
+ assert.equal(await channelExists(paths, "alpha"), false);
+ });
+});
+
+test("deleting a relocated channel reclaims the media on the other drive", async () => {
+ await withPaths(async (paths, mediaRoot) => {
+ const target = relocatedDataDir(mediaRoot, "alpha");
+ await writeChannelConfig(paths, "alpha", { ...config, dataDir: target });
+ await mkdir(path.join(target, "v1"), { recursive: true });
+ await writeFile(path.join(target, "v1", "audio.m4a"), "BYTES");
+ await symlink(target, path.join(paths.channelsDir, "alpha", "data"));
+
+ await deleteChannel(paths, "alpha");
+
+ assert.equal(await channelExists(paths, "alpha"), false);
+ assert.equal(await there(target), false);
+ // <root>/<slug> goes too, once it holds nothing.
+ assert.equal(await there(path.join(mediaRoot, "alpha")), false);
+ // ...but the root itself is the operator's and is never touched.
+ assert.equal(await there(mediaRoot), true);
+ });
+});
+
+test("a non-empty <root>/<slug> survives the delete", async () => {
+ await withPaths(async (paths, mediaRoot) => {
+ const target = relocatedDataDir(mediaRoot, "alpha");
+ await writeChannelConfig(paths, "alpha", { ...config, dataDir: target });
+ await mkdir(target, { recursive: true });
+ await symlink(target, path.join(paths.channelsDir, "alpha", "data"));
+ // Something of the operator's, alongside the data dir we own.
+ await writeFile(path.join(mediaRoot, "alpha", "NOTES.txt"), "mine");
+
+ await deleteChannel(paths, "alpha");
+ assert.equal(await there(target), false);
+ assert.equal(await there(path.join(mediaRoot, "alpha", "NOTES.txt")), true);
+ });
+});
+
+test("deleting is refused while a relocation is in flight", async () => {
+ await withPaths(async (paths, mediaRoot) => {
+ const target = relocatedDataDir(mediaRoot, "alpha");
+ await writeChannelConfig(paths, "alpha", config);
+ await mkdir(path.join(paths.channelsDir, "alpha", "data"), {
+ recursive: true,
+ });
+ await writeFile(
+ path.join(paths.channelsDir, "alpha", RELOCATION_MARKER_FILENAME),
+ JSON.stringify({
+ target,
+ direction: "out",
+ startedAt: new Date().toISOString(),
+ phase: "copy",
+ }),
+ );
+ await assert.rejects(
+ () => deleteChannel(paths, "alpha"),
+ /relocation in progress/,
+ );
+ assert.equal(await channelExists(paths, "alpha"), true);
+ });
+});
+
+test("an unmounted target does not make the channel undeletable", async () => {
+ await withPaths(async (paths, mediaRoot) => {
+ // `force: true` makes a missing target a no-op, so the channel goes and its
+ // media stays on the unmounted platter. Deliberate: refusing here would
+ // make a channel whose drive is gone for good impossible to remove.
+ const target = relocatedDataDir(mediaRoot, "alpha");
+ await writeChannelConfig(paths, "alpha", { ...config, dataDir: target });
+ await symlink(target, path.join(paths.channelsDir, "alpha", "data"));
+ await deleteChannel(paths, "alpha");
+ assert.equal(await channelExists(paths, "alpha"), false);
+ });
+});
diff --git a/common/controller/channels.ts b/common/controller/channels.ts
@@ -13,6 +13,7 @@ import {
readVideoFiles,
} from "../lib/videoStatus";
import { loadDigest } from "../lib/digest-server";
+import { readRelocationMarker } from "../lib/channelMedia";
// TYPE-ONLY, and it must stay that way: ./channelSnapshot imports
// readChannelConfig from this module, and it drags in the snapshot generator's
// whole dependency graph (lmdb, the archive reader, the digest layer). A value
@@ -473,10 +474,39 @@ export async function createChannel(
await writeChannelConfig(paths, slug, config);
}
+// DELETING A RELOCATED CHANNEL HAS TO REACH THE OTHER DRIVE. `rm -r` on the
+// channel dir removes the symlink, not what it points at (node's rm does not
+// follow links), so without this the media would survive the channel as an
+// orphan nothing in the app can see or reclaim. The target is removed FIRST, so
+// a target that cannot be removed for a reason OTHER than being absent takes
+// the delete down with it and leaves the channel intact and retryable. An
+// unmounted drive is NOT that case — `force: true` makes a missing path a
+// no-op, and the channel is deleted with its media left on the platter. That is
+// deliberate: refusing would make a channel whose drive is gone for good
+// undeletable, and the marker check below is the refusal that matters.
export async function deleteChannel(
paths: Paths,
slug: string,
): Promise<void> {
+ const marker = await readRelocationMarker(paths, slug);
+ if (marker) {
+ throw new Error(
+ `Channel "${slug}" has a media relocation in progress (phase ` +
+ `"${marker.phase}", target ${marker.target}). Finish or cancel it ` +
+ `before deleting the channel.`,
+ );
+ }
const dir = path.join(paths.channelsDir, slug);
+ const target = (await readChannelConfig(paths, slug))?.dataDir?.trim();
+ if (target) {
+ await rm(target, { recursive: true, force: true });
+ // <root>/<slug> is ours by construction (relocatedDataDir fixes the
+ // suffix), so take it too — but only when nothing else landed in it.
+ const slugRoot = path.dirname(target);
+ if (path.basename(slugRoot) === slug) {
+ const left = await readdir(slugRoot).catch(() => ["keep"]);
+ if (left.length === 0) await rm(slugRoot, { recursive: true, force: true });
+ }
+ }
await rm(dir, { recursive: true, force: true });
}
diff --git a/common/controller/normalizeAll.ts b/common/controller/normalizeAll.ts
@@ -9,6 +9,7 @@ import pLimit from "p-limit";
import { listChannelStatsFromDisk } from "./channels";
import { normalizeTranscript } from "./normalizeTranscript";
import type { Paths } from "../lib/paths";
+import { assertChannelMediaReachable } from "../lib/channelMedia";
export type NormalizeAllOptions = {
paths: Paths;
@@ -53,6 +54,18 @@ export async function normalizeAllTranscripts(
for (const ch of channels) {
if (opts.signal?.aborted) break;
const dataDir = path.join(opts.paths.channelsDir, ch.slug, "data");
+ // GUARD, same reason as the snapshot's and the batch's: `readdir(dataDir)
+ // .catch(() => [])` cannot tell "this channel has downloaded nothing" from
+ // "this channel's drive is not mounted", and here the second reads as a
+ // clean run over zero videos that reports 0/0/0/0 and moves on. A sweep
+ // that silently skips a channel is worse than one that stops on it.
+ try {
+ await assertChannelMediaReachable(opts.paths, ch.slug, ch.config);
+ } catch (err) {
+ log(`Normalize ${ch.slug}: SKIPPED — ${(err as Error).message}`);
+ result.failed++;
+ continue;
+ }
const videoIds = await readdir(dataDir).catch(() => [] as string[]);
log(`Normalize ${ch.slug}: ${videoIds.length} videos`);
let wrote = 0;
diff --git a/common/controller/operationBatch.ts b/common/controller/operationBatch.ts
@@ -38,6 +38,7 @@ import { open, readFile, readdir } from "node:fs/promises";
import type { Paths } from "../lib/paths";
import { getSettings, type SiteSettings } from "../lib/settings";
import { isGateHeld } from "../lib/pauseGates";
+import { assertChannelMediaReachable } from "../lib/channelMedia";
import { runPool } from "../jobs/concurrentRunner";
import type { TaskTracker } from "../jobs/taskHooks";
import type { JobProgress } from "../jobs/registry";
@@ -1576,6 +1577,12 @@ export async function runOperationBatch(
? await resolveDigestFor(run, opts.channelSlug)
: null;
+ // GUARD 4 OF FOUR (see plans/relocate-channel-media.md). Same swallow as the
+ // snapshot's, deciding the candidate list for both operation lanes. All three
+ // of this function's callers are already behind guard 1 today; this is here so
+ // a fourth in-process caller that is not cannot quietly find "no candidates"
+ // on a channel whose drive is unmounted.
+ await assertChannelMediaReachable(opts.paths, opts.channelSlug);
const dataDir = path.join(opts.paths.channelsDir, opts.channelSlug, "data");
const allDirs = await readdir(dataDir).catch(() => [] as string[]);
const onDisk = new Set(allDirs);
diff --git a/common/controller/relocateChannelMedia.test.ts b/common/controller/relocateChannelMedia.test.ts
@@ -0,0 +1,900 @@
+import { test } from "node:test";
+import assert from "node:assert/strict";
+import {
+ lstat,
+ mkdir,
+ mkdtemp,
+ readdir,
+ readFile,
+ readlink,
+ rename,
+ rm,
+ stat,
+ symlink,
+ utimes,
+ writeFile,
+} from "node:fs/promises";
+import { tmpdir } from "node:os";
+import path from "node:path";
+import type { Paths } from "../lib/paths";
+import {
+ inspectChannelMedia,
+ readRelocationMarker,
+ relocatedDataDir,
+} from "../lib/channelMedia";
+import { previewRelocation, relocateChannelMedia } from "./relocateChannelMedia";
+import { readChannelConfig } from "./channels";
+
+// Run with:
+// pnpm --filter yt-dlp-transcript-common exec tsx --test controller/relocateChannelMedia.test.ts
+//
+// These exercise the REAL rsync binary (paths.rsyncBin -> "rsync"), like
+// controller/backupSavedVideos.test.ts. Everything happens inside one mkdtemp.
+
+const MTIME = new Date("2021-03-04T05:06:07.000Z");
+
+async function withTmp(
+ fn: (paths: Paths, root: string, dir: string) => Promise<void>,
+): Promise<void> {
+ const dir = await mkdtemp(path.join(tmpdir(), "ttb-relocate-"));
+ // THE CORPUS AND THE PLATTER ARE SIBLINGS under the tmpdir, never nested. A
+ // root inside the corpus is refused by design (relocationRootProblem: it
+ // would copy the channel onto itself and then reclaim the only copy), so a
+ // harness that nested them would be exercising a shape the controller
+ // rejects — and would have hidden exactly the bug that check exists for.
+ const transcriptsDir = path.join(dir, "corpus");
+ const paths = {
+ transcriptsDir,
+ channelsDir: path.join(transcriptsDir, "channels"),
+ rsyncBin: "rsync",
+ } as Paths;
+ const root = path.join(dir, "platter");
+ await mkdir(paths.channelsDir, { recursive: true });
+ await mkdir(root, { recursive: true });
+ try {
+ await fn(paths, root, dir);
+ } finally {
+ await rm(dir, { recursive: true, force: true });
+ }
+}
+
+async function seed(
+ paths: Paths,
+ slug: string,
+ videos: Record<string, Record<string, string>>,
+ config: Record<string, unknown> = {},
+): Promise<string> {
+ const channelDir = path.join(paths.channelsDir, slug);
+ await mkdir(channelDir, { recursive: true });
+ await writeFile(
+ path.join(channelDir, "config.json"),
+ JSON.stringify(
+ { handling: "transcribe", url: "https://example.com/c", ...config },
+ null,
+ 2,
+ ) + "\n",
+ );
+ // These live in the CHANNEL dir, not data/ — they must not move.
+ await writeFile(path.join(channelDir, "archive"), "youtube v1\n");
+ await writeFile(path.join(channelDir, "playlist"), "https://x/v1\n");
+ for (const [id, files] of Object.entries(videos)) {
+ const videoDir = path.join(channelDir, "data", id);
+ await mkdir(videoDir, { recursive: true });
+ for (const [name, content] of Object.entries(files)) {
+ const file = path.join(videoDir, name);
+ await writeFile(file, content);
+ // A known mtime, so "rsync -a preserved it" is an assertion and not a
+ // coincidence. The LMDB index keys on mtimes, so this is what makes a
+ // relocation free of a reindex.
+ await utimes(file, MTIME, MTIME);
+ }
+ }
+ return channelDir;
+}
+
+test("out: copies, links, records the target, keeps mtimes and reclaims the source", async () => {
+ await withTmp(async (paths, root) => {
+ const channelDir = await seed(paths, "alpha", {
+ v1: { "audio.m4a": "one".repeat(500), "transcript.json": "{}" },
+ v2: { "audio.m4a": "two".repeat(500) },
+ });
+ const target = relocatedDataDir(root, "alpha");
+
+ const res = await relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "out",
+ root,
+ onLog: () => {},
+ });
+ assert.equal(res.direction, "out");
+ assert.equal(res.target, target);
+ assert.equal(res.files, 3);
+ assert.equal(res.resumed, false);
+
+ // data/ is an absolute symlink at the target...
+ const link = await lstat(path.join(channelDir, "data"));
+ assert.ok(link.isSymbolicLink());
+ assert.equal(await readlink(path.join(channelDir, "data")), target);
+
+ // ...so the on-disk contract channelDir/data/<id>/<file> still resolves.
+ assert.equal(
+ await readFile(
+ path.join(channelDir, "data", "v1", "transcript.json"),
+ "utf8",
+ ),
+ "{}",
+ );
+ assert.equal(
+ (await stat(path.join(target, "v1", "audio.m4a"))).mtime.getTime(),
+ MTIME.getTime(),
+ );
+
+ // The record is in config.json, written only now, on success.
+ assert.equal((await readChannelConfig(paths, "alpha"))?.dataDir, target);
+ assert.equal((await inspectChannelMedia(paths, "alpha")).status, "ok");
+
+ // The parked copy is gone and the marker with it — this is the step that
+ // actually frees the source volume.
+ const left = await readdir(channelDir);
+ assert.deepEqual(
+ left.filter((n) => n.startsWith("data.")),
+ [],
+ );
+ assert.equal(await readRelocationMarker(paths, "alpha"), null);
+ // Channel-level files never moved.
+ assert.ok(left.includes("archive"));
+ assert.ok(left.includes("playlist"));
+ });
+});
+
+test("abort from an onLog hook leaves the source intact, and the rerun completes", async () => {
+ await withTmp(async (paths, root) => {
+ const channelDir = await seed(paths, "alpha", {
+ v1: { "audio.m4a": "x".repeat(20000) },
+ });
+ const target = relocatedDataDir(root, "alpha");
+
+ // Aborting from the log hook is how a Cancel button reaches this code: the
+ // job's onLog and its AbortSignal belong to the same run. Firing on the
+ // first line is deterministic — what it pins is the contract (the source is
+ // untouched, the marker is kept, a rerun resumes), not rsync's own
+ // --partial mechanics, which are rsync's to test.
+ const controller = new AbortController();
+ await assert.rejects(
+ () =>
+ relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "out",
+ root,
+ signal: controller.signal,
+ onLog: () => controller.abort(),
+ }),
+ /Cancelled/,
+ );
+
+ // The source is a REAL directory still, with its file.
+ assert.ok((await lstat(path.join(channelDir, "data"))).isDirectory());
+ assert.equal(
+ (await readFile(path.join(channelDir, "data", "v1", "audio.m4a"), "utf8"))
+ .length,
+ 20000,
+ );
+ // Nothing was recorded, and the marker says where it was going.
+ assert.equal((await readChannelConfig(paths, "alpha"))?.dataDir, undefined);
+ const marker = await readRelocationMarker(paths, "alpha");
+ assert.equal(marker?.target, target);
+ assert.equal(marker?.direction, "out");
+ assert.equal(marker?.phase, "copy");
+ // And every guard now reads the channel as in transition.
+ assert.equal(
+ (await inspectChannelMedia(paths, "alpha")).status,
+ "in-transition",
+ );
+
+ // The rerun finishes it. It is the ONE caller allowed to look past its own
+ // marker.
+ const res = await relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "out",
+ root,
+ onLog: () => {},
+ });
+ assert.equal(res.resumed, true);
+ assert.equal((await inspectChannelMedia(paths, "alpha")).status, "ok");
+ assert.equal((await readChannelConfig(paths, "alpha"))?.dataDir, target);
+ });
+});
+
+test("a verify failure keeps the source and does not write the config", async () => {
+ await withTmp(async (paths, root) => {
+ const channelDir = await seed(paths, "alpha", {
+ v1: { "audio.m4a": "one" },
+ });
+ const target = relocatedDataDir(root, "alpha");
+ // A stray file at the target from some earlier, unrelated write. rsync
+ // without --delete leaves it, so the itemize dry-run is clean and only the
+ // two-sided measurement catches the drift — which is exactly why the verify
+ // is both checks and not just the dry run.
+ await mkdir(path.join(target, "v9"), { recursive: true });
+ await writeFile(path.join(target, "v9", "stray.m4a"), "not ours");
+
+ await assert.rejects(
+ () =>
+ relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "out",
+ root,
+ onLog: () => {},
+ }),
+ /Verification failed/,
+ );
+
+ assert.ok((await lstat(path.join(channelDir, "data"))).isDirectory());
+ assert.equal(
+ await readFile(path.join(channelDir, "data", "v1", "audio.m4a"), "utf8"),
+ "one",
+ );
+ assert.equal((await readChannelConfig(paths, "alpha"))?.dataDir, undefined);
+ });
+});
+
+test("back: restores a real directory, clears the config and reclaims the target", async () => {
+ await withTmp(async (paths, root) => {
+ const channelDir = await seed(paths, "alpha", {
+ v1: { "audio.m4a": "one", "transcript.json": "{}" },
+ });
+ const target = relocatedDataDir(root, "alpha");
+ await relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "out",
+ root,
+ onLog: () => {},
+ });
+
+ const res = await relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "back",
+ onLog: () => {},
+ });
+ assert.equal(res.direction, "back");
+ assert.equal(res.files, 2);
+
+ const back = await lstat(path.join(channelDir, "data"));
+ assert.ok(back.isDirectory());
+ assert.ok(!back.isSymbolicLink());
+ assert.equal(
+ await readFile(path.join(channelDir, "data", "v1", "audio.m4a"), "utf8"),
+ "one",
+ );
+ assert.equal(
+ (await stat(path.join(channelDir, "data", "v1", "audio.m4a"))).mtime.getTime(),
+ MTIME.getTime(),
+ );
+ assert.equal((await readChannelConfig(paths, "alpha"))?.dataDir, undefined);
+ assert.equal((await inspectChannelMedia(paths, "alpha")).status, "in-place");
+ // The target is reclaimed, and so is <root>/<slug> once it is empty.
+ assert.equal(await pathThere(target), false);
+ assert.equal(await pathThere(path.dirname(target)), false);
+ assert.equal(await readRelocationMarker(paths, "alpha"), null);
+ assert.equal(
+ (await readdir(channelDir)).includes("data.incoming"),
+ false,
+ );
+ });
+});
+
+test("moving back a channel that was never moved is refused", async () => {
+ await withTmp(async (paths) => {
+ await seed(paths, "alpha", { v1: { "audio.m4a": "one" } });
+ await assert.rejects(
+ () =>
+ relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "back",
+ onLog: () => {},
+ }),
+ /already in place/,
+ );
+ });
+});
+
+test("a relocated channel is refused a second move without a move back", async () => {
+ await withTmp(async (paths, root, dir) => {
+ await seed(paths, "alpha", { v1: { "audio.m4a": "one" } });
+ await relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "out",
+ root,
+ onLog: () => {},
+ });
+ const other = path.join(dir, "platter2");
+ await mkdir(other, { recursive: true });
+ await assert.rejects(
+ () =>
+ relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "out",
+ root: other,
+ onLog: () => {},
+ }),
+ /already relocated/,
+ );
+ });
+});
+
+test("an unmounted root is refused before anything is written", async () => {
+ await withTmp(async (paths, _root, dir) => {
+ const channelDir = await seed(paths, "alpha", { v1: { "audio.m4a": "one" } });
+ await assert.rejects(
+ () =>
+ relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "out",
+ root: path.join(dir, "not-mounted"),
+ onLog: () => {},
+ }),
+ /does not exist or is not a directory/,
+ );
+ assert.equal(await readRelocationMarker(paths, "alpha"), null);
+ assert.ok((await lstat(path.join(channelDir, "data"))).isDirectory());
+ });
+});
+
+test("a social channel has no media to relocate", async () => {
+ await withTmp(async (paths, root) => {
+ await seed(paths, "poster", {}, { sourceKind: "social" });
+ await assert.rejects(
+ () =>
+ relocateChannelMedia({
+ paths,
+ slug: "poster",
+ direction: "out",
+ root,
+ onLog: () => {},
+ }),
+ /social channel/,
+ );
+ });
+});
+
+test("preview measures the tree and both volumes without moving anything", async () => {
+ await withTmp(async (paths, root) => {
+ const channelDir = await seed(paths, "alpha", {
+ v1: { "audio.m4a": "abcde", "transcript.json": "{}" },
+ v2: { "audio.m4a": "fgh" },
+ });
+ const preview = await previewRelocation({ paths, slug: "alpha", root });
+ assert.equal(preview.target, relocatedDataDir(root, "alpha"));
+ assert.equal(preview.files, 3);
+ assert.equal(preview.bytesToMove, 5 + 2 + 3);
+ assert.equal(preview.existingPartial, false);
+ assert.ok(preview.freeOnRoot > 0);
+ assert.ok(preview.freeOnSource > 0);
+ // Both dirs are under one tmpdir, so this really is the same volume.
+ assert.equal(preview.sameDevice, true);
+ assert.ok((await lstat(path.join(channelDir, "data"))).isDirectory());
+ });
+});
+
+async function pathThere(p: string): Promise<boolean> {
+ try {
+ await stat(p);
+ return true;
+ } catch {
+ return false;
+ }
+}
+
+// ---------------------------------------------------------------------------
+// CRASH RECOVERY: one case per phase per direction.
+//
+// The plan's contract is "a rerun with a matching marker re-verifies and
+// continues from swap/reclaim", and every one of these states is reachable from
+// a kill -9 between two syscalls. What they pin is that a rerun OBSERVES the
+// disk rather than assuming the previous run finished the step the marker names:
+// the marker says how far the last attempt got, it cannot say how far it got
+// THROUGH a phase, and the disk is the only evidence.
+//
+// Each seeds the exact on-disk state of a crash at that point and asserts the
+// rerun reaches a clean `ok` / `in-place` — with the config right, the marker
+// gone and NO `data.relocated-*` or `data.incoming` left holding a second copy
+// on the volume the move exists to free.
+
+async function seedMarker(
+ paths: Paths,
+ slug: string,
+ marker: {
+ target: string;
+ direction: "out" | "back";
+ phase: "copy" | "swap" | "reclaim";
+ },
+): Promise<void> {
+ await writeFile(
+ path.join(paths.channelsDir, slug, ".relocating.json"),
+ JSON.stringify({ ...marker, startedAt: new Date().toISOString() }) + "\n",
+ );
+}
+
+async function setConfigDataDir(
+ paths: Paths,
+ slug: string,
+ value: string | null,
+): Promise<void> {
+ const file = path.join(paths.channelsDir, slug, "config.json");
+ const config = JSON.parse(await readFile(file, "utf8")) as Record<
+ string,
+ unknown
+ >;
+ if (value === null) delete config.dataDir;
+ else config.dataDir = value;
+ await writeFile(file, JSON.stringify(config, null, 2) + "\n");
+}
+
+async function leftoverCopies(channelDir: string): Promise<string[]> {
+ return (await readdir(channelDir))
+ .filter((n) => n.startsWith("data.relocated-") || n === "data.incoming")
+ .sort();
+}
+
+// rsync -a of the tree, so the seeded "already copied" target is byte-for-byte
+// what a completed copy phase would have left — including mtimes, which the
+// re-verify compares.
+async function copyTree(src: string, dest: string): Promise<void> {
+ await mkdir(dest, { recursive: true });
+ for (const entry of await readdir(src, { withFileTypes: true })) {
+ const from = path.join(src, entry.name);
+ const to = path.join(dest, entry.name);
+ if (entry.isDirectory()) {
+ await copyTree(from, to);
+ } else {
+ await writeFile(to, await readFile(from));
+ await utimes(to, MTIME, MTIME);
+ }
+ }
+}
+
+test("out @ swap: crash before the rename — the rerun re-verifies and completes", async () => {
+ await withTmp(async (paths, root) => {
+ const channelDir = await seed(paths, "alpha", {
+ v1: { "audio.m4a": "one".repeat(500) },
+ });
+ const target = relocatedDataDir(root, "alpha");
+ // The copy finished; the process died before `data/` was parked.
+ await copyTree(path.join(channelDir, "data"), target);
+ await seedMarker(paths, "alpha", { target, direction: "out", phase: "swap" });
+
+ const res = await relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "out",
+ root,
+ onLog: () => {},
+ });
+ assert.equal(res.resumed, true);
+ assert.equal((await inspectChannelMedia(paths, "alpha")).status, "ok");
+ assert.equal((await readChannelConfig(paths, "alpha"))?.dataDir, target);
+ assert.deepEqual(await leftoverCopies(channelDir), []);
+ assert.equal(await readRelocationMarker(paths, "alpha"), null);
+ });
+});
+
+test("out @ swap: crash after the config write — the first run's parked copy is still reclaimed", async () => {
+ await withTmp(async (paths, root) => {
+ const channelDir = await seed(paths, "alpha", {
+ v1: { "audio.m4a": "one".repeat(500) },
+ });
+ const dataDir = path.join(channelDir, "data");
+ const target = relocatedDataDir(root, "alpha");
+ await copyTree(dataDir, target);
+ // Everything through writeChannelConfig ran; the reclaim marker never
+ // landed. A full second copy of the channel is parked on the source volume
+ // — the disk the whole move exists to free — and the marker still says
+ // "swap", so the old code minted a SECOND parked name and reclaimed only
+ // that one, orphaning this forever.
+ const parked = path.join(channelDir, "data.relocated-1700000000000");
+ await rename(dataDir, parked);
+ await symlink(target, dataDir);
+ await setConfigDataDir(paths, "alpha", target);
+ await seedMarker(paths, "alpha", { target, direction: "out", phase: "swap" });
+
+ await relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "out",
+ root,
+ onLog: () => {},
+ });
+ assert.equal((await inspectChannelMedia(paths, "alpha")).status, "ok");
+ assert.deepEqual(await leftoverCopies(channelDir), []);
+ assert.equal(await readRelocationMarker(paths, "alpha"), null);
+ // The media itself is untouched by the sweep.
+ assert.equal(
+ await readFile(path.join(dataDir, "v1", "audio.m4a"), "utf8"),
+ "one".repeat(500),
+ );
+ });
+});
+
+test("out @ reclaim: the rerun sweeps every parked copy and clears the marker", async () => {
+ await withTmp(async (paths, root) => {
+ const channelDir = await seed(paths, "alpha", {
+ v1: { "audio.m4a": "one".repeat(500) },
+ });
+ const dataDir = path.join(channelDir, "data");
+ const target = relocatedDataDir(root, "alpha");
+ await copyTree(dataDir, target);
+ await rename(dataDir, path.join(channelDir, "data.relocated-1700000000000"));
+ await symlink(target, dataDir);
+ await setConfigDataDir(paths, "alpha", target);
+ await seedMarker(paths, "alpha", {
+ target,
+ direction: "out",
+ phase: "reclaim",
+ });
+
+ await relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "out",
+ root,
+ onLog: () => {},
+ });
+ assert.equal((await inspectChannelMedia(paths, "alpha")).status, "ok");
+ assert.deepEqual(await leftoverCopies(channelDir), []);
+ assert.equal(await readRelocationMarker(paths, "alpha"), null);
+ });
+});
+
+// THE MISSING CELL of the resume matrix: out @ {swap, reclaim} and back @
+// {swap(before), swap(after), reclaim} all had cases, and back @ copy had none —
+// the only one where the rerun has to finish an rsync rather than a rename, and
+// the only one where the space check is asked to credit a PARTIAL copy.
+test("back @ copy: a half-copied data.incoming is resumed, not restarted", async () => {
+ await withTmp(async (paths, root) => {
+ const channelDir = await seed(paths, "alpha", {
+ v1: { "audio.m4a": "one".repeat(500), "transcript.json": "{}" },
+ v2: { "audio.m4a": "two".repeat(500) },
+ });
+ const dataDir = path.join(channelDir, "data");
+ const target = relocatedDataDir(root, "alpha");
+ await copyTree(dataDir, target);
+ await rm(dataDir, { recursive: true, force: true });
+ await symlink(target, dataDir);
+ await setConfigDataDir(paths, "alpha", target);
+ // The copy got one video in and died. The marker says `copy`, so the rerun
+ // resumes the rsync — and the surviving file keeps its mtime, which is what
+ // says rsync skipped it rather than re-sending it.
+ const incoming = path.join(channelDir, "data.incoming");
+ await copyTree(path.join(target, "v1"), path.join(incoming, "v1"));
+ await seedMarker(paths, "alpha", { target, direction: "back", phase: "copy" });
+
+ const res = await relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "back",
+ onLog: () => {},
+ });
+ assert.equal(res.resumed, true);
+ assert.equal(res.files, 3);
+ assert.equal((await inspectChannelMedia(paths, "alpha")).status, "in-place");
+ assert.ok((await lstat(dataDir)).isDirectory());
+ assert.equal(
+ await readFile(path.join(dataDir, "v2", "audio.m4a"), "utf8"),
+ "two".repeat(500),
+ );
+ assert.equal(
+ (await stat(path.join(dataDir, "v1", "audio.m4a"))).mtime.getTime(),
+ MTIME.getTime(),
+ );
+ assert.equal((await readChannelConfig(paths, "alpha"))?.dataDir, undefined);
+ assert.equal(await pathThere(target), false);
+ assert.deepEqual(await leftoverCopies(channelDir), []);
+ assert.equal(await readRelocationMarker(paths, "alpha"), null);
+ });
+});
+
+test("back @ swap: crash before the rename — the rerun finishes the swap", async () => {
+ await withTmp(async (paths, root) => {
+ const channelDir = await seed(paths, "alpha", {
+ v1: { "audio.m4a": "one".repeat(500) },
+ });
+ const dataDir = path.join(channelDir, "data");
+ const target = relocatedDataDir(root, "alpha");
+ // A completed relocation, then a move-back whose copy finished.
+ await copyTree(dataDir, target);
+ await rm(dataDir, { recursive: true, force: true });
+ await symlink(target, dataDir);
+ await setConfigDataDir(paths, "alpha", target);
+ await copyTree(target, path.join(channelDir, "data.incoming"));
+ await seedMarker(paths, "alpha", { target, direction: "back", phase: "swap" });
+
+ const res = await relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "back",
+ onLog: () => {},
+ });
+ assert.equal(res.resumed, true);
+ assert.equal((await inspectChannelMedia(paths, "alpha")).status, "in-place");
+ assert.ok((await lstat(dataDir)).isDirectory());
+ assert.equal((await readChannelConfig(paths, "alpha"))?.dataDir, undefined);
+ assert.deepEqual(await leftoverCopies(channelDir), []);
+ assert.equal(await readRelocationMarker(paths, "alpha"), null);
+ });
+});
+
+test("back @ swap: crash AFTER the rename — the rerun does not ENOENT forever", async () => {
+ await withTmp(async (paths, root) => {
+ const channelDir = await seed(paths, "alpha", {
+ v1: { "audio.m4a": "one".repeat(500) },
+ });
+ const dataDir = path.join(channelDir, "data");
+ const target = relocatedDataDir(root, "alpha");
+ await copyTree(dataDir, target);
+ // rename(incoming, data) committed; the process died before the config was
+ // cleared. `data/` is a real directory and `incoming` is gone — so the old
+ // sequence unlinked nothing (swallowed), then renamed a path that no longer
+ // exists, and failed identically on every rerun.
+ await setConfigDataDir(paths, "alpha", target);
+ await seedMarker(paths, "alpha", { target, direction: "back", phase: "swap" });
+
+ await relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "back",
+ onLog: () => {},
+ });
+ assert.equal((await inspectChannelMedia(paths, "alpha")).status, "in-place");
+ assert.equal((await readChannelConfig(paths, "alpha"))?.dataDir, undefined);
+ assert.equal(await pathIsThere(target), false);
+ assert.deepEqual(await leftoverCopies(channelDir), []);
+ assert.equal(await readRelocationMarker(paths, "alpha"), null);
+ });
+});
+
+test("back @ reclaim: the config is already clear, and the rerun still finishes", async () => {
+ await withTmp(async (paths, root) => {
+ const channelDir = await seed(paths, "alpha", {
+ v1: { "audio.m4a": "one".repeat(500) },
+ });
+ const target = relocatedDataDir(root, "alpha");
+ // The swap completed and cleared config.dataDir — which is why this case
+ // was UNREACHABLE: with no dataDir the entry point threw "is not relocated"
+ // and the marker named the only place that knew where the media had been.
+ await copyTree(path.join(channelDir, "data"), target);
+ await seedMarker(paths, "alpha", {
+ target,
+ direction: "back",
+ phase: "reclaim",
+ });
+
+ await relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "back",
+ onLog: () => {},
+ });
+ assert.equal((await inspectChannelMedia(paths, "alpha")).status, "in-place");
+ assert.equal(await pathIsThere(target), false);
+ assert.deepEqual(await leftoverCopies(channelDir), []);
+ assert.equal(await readRelocationMarker(paths, "alpha"), null);
+ });
+});
+
+test("a marker for the other direction is never resumed into", async () => {
+ await withTmp(async (paths, root) => {
+ await seed(paths, "alpha", { v1: { "audio.m4a": "one" } });
+ const target = relocatedDataDir(root, "alpha");
+ await setConfigDataDir(paths, "alpha", target);
+ await seedMarker(paths, "alpha", { target, direction: "out", phase: "swap" });
+ await assert.rejects(
+ relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "back",
+ onLog: () => {},
+ }),
+ /move-out .* in flight|interrupted/,
+ );
+ });
+});
+
+async function pathIsThere(p: string): Promise<boolean> {
+ try {
+ await lstat(p);
+ return true;
+ } catch {
+ return false;
+ }
+}
+
+// ---------------------------------------------------------------------------
+// TWO WAYS THE MOVE USED TO EAT THE MEDIA. Both were reachable from the shipped
+// UI and both ended with rm -r on the only copy, so both get their own cases.
+// ---------------------------------------------------------------------------
+
+// `inconsistent` — config.json records a target while `data/` is a REAL
+// directory — is what `rsync --copy-links` of a channel produces, which
+// WORKTREES.md documents as the way to carry media into a shard. inspect()
+// reports relocated:true for it, so the panel offered "Move back in place", and
+// the run copied the target to data.incoming, verified it, found data/ already a
+// real dir, logged "the swap had completed", cleared the config, rm -r'd the
+// target and then swept data.incoming. Three copies in, zero out.
+test("back: an inconsistent channel is refused, and the target keeps its bytes", async () => {
+ await withTmp(async (paths, root) => {
+ const channelDir = await seed(paths, "alpha", {
+ v1: { "audio.m4a": "one" },
+ });
+ const target = relocatedDataDir(root, "alpha");
+ await relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "out",
+ root,
+ onLog: () => {},
+ });
+ // Turn the link back into a real directory WITHOUT clearing config.dataDir —
+ // the --copy-links shape, reproduced.
+ await rm(path.join(channelDir, "data"));
+ await copyTree(target, path.join(channelDir, "data"));
+ assert.equal(
+ (await inspectChannelMedia(paths, "alpha")).status,
+ "inconsistent",
+ );
+
+ await assert.rejects(
+ () =>
+ relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "back",
+ onLog: () => {},
+ }),
+ /not in a movable state/,
+ );
+
+ // The relocated copy is still there, the local one is still there, and
+ // nothing was parked or marked.
+ assert.equal(
+ await readFile(path.join(target, "v1", "audio.m4a"), "utf8"),
+ "one",
+ );
+ assert.equal(
+ await readFile(path.join(channelDir, "data", "v1", "audio.m4a"), "utf8"),
+ "one",
+ );
+ assert.equal((await readChannelConfig(paths, "alpha"))?.dataDir, target);
+ assert.equal(await readRelocationMarker(paths, "alpha"), null);
+ assert.equal((await readdir(channelDir)).includes("data.incoming"), false);
+ });
+});
+
+// An `unreachable` channel (the drive is not mounted) is refused for the same
+// reason: nothing can vouch for what the target holds, and the run ends by
+// deleting it.
+test("back: an unreachable channel is refused", async () => {
+ await withTmp(async (paths, root) => {
+ await seed(paths, "alpha", { v1: { "audio.m4a": "one" } });
+ const target = relocatedDataDir(root, "alpha");
+ await relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "out",
+ root,
+ onLog: () => {},
+ });
+ // The platter goes away. The link dangles; config still names the target.
+ await rm(path.dirname(target), { recursive: true, force: true });
+ assert.equal(
+ (await inspectChannelMedia(paths, "alpha")).status,
+ "unreachable",
+ );
+ await assert.rejects(
+ () =>
+ relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "back",
+ onLog: () => {},
+ }),
+ /not in a movable state|not reachable/,
+ );
+ });
+});
+
+// THE SELF-RELOCATION. root = <transcriptsDir>/channels makes the target the
+// source: rsync src/ src/ succeeds, verifyCopy compares the tree with itself,
+// the swap parks `data/` (which IS the verified "target") and links to a path
+// that no longer exists, and the reclaim sweep deletes the only copy. The
+// panel's own help text describes transcripts/channels almost word for word.
+test("a root inside the corpus is refused by the job and by the preview", async () => {
+ await withTmp(async (paths) => {
+ const channelDir = await seed(paths, "alpha", {
+ v1: { "audio.m4a": "one" },
+ });
+ const roots = [
+ paths.channelsDir,
+ paths.transcriptsDir,
+ channelDir,
+ path.join(channelDir, "data"),
+ path.join(paths.channelsDir, "beta"),
+ ];
+ for (const bad of roots) {
+ await assert.rejects(
+ () =>
+ relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "out",
+ root: bad,
+ onLog: () => {},
+ }),
+ /inside the corpus|inside the channel directory/,
+ `job accepted ${bad}`,
+ );
+ await assert.rejects(
+ () => previewRelocation({ paths, slug: "alpha", root: bad }),
+ /inside the corpus|inside the channel directory/,
+ `preview accepted ${bad}`,
+ );
+ }
+ // Untouched: still a real directory, no config, no marker, no parked copy.
+ assert.ok((await lstat(path.join(channelDir, "data"))).isDirectory());
+ assert.equal(
+ await readFile(path.join(channelDir, "data", "v1", "audio.m4a"), "utf8"),
+ "one",
+ );
+ assert.equal((await readChannelConfig(paths, "alpha"))?.dataDir, undefined);
+ assert.equal(await readRelocationMarker(paths, "alpha"), null);
+ assert.equal(
+ (await readdir(channelDir)).filter((n) => n.startsWith("data.")).length,
+ 0,
+ );
+ });
+});
+
+// A root OUTSIDE the corpus whose `<slug>` level is a symlink back into it. The
+// root's own realpath cannot see this — the link is one level down — which is
+// why the target is resolved separately.
+test("a root that links back into the channel dir is refused", async () => {
+ await withTmp(async (paths, _root, dir) => {
+ const channelDir = await seed(paths, "alpha", {
+ v1: { "audio.m4a": "one" },
+ });
+ const sneaky = path.join(dir, "sneaky");
+ await mkdir(sneaky, { recursive: true });
+ await symlink(channelDir, path.join(sneaky, "alpha"));
+ await assert.rejects(
+ () =>
+ relocateChannelMedia({
+ paths,
+ slug: "alpha",
+ direction: "out",
+ root: sneaky,
+ onLog: () => {},
+ }),
+ /inside the channel directory|inside the corpus/,
+ );
+ assert.ok((await lstat(path.join(channelDir, "data"))).isDirectory());
+ });
+});
+
+test("a relative root is refused by the preview, not only by the job", async () => {
+ await withTmp(async (paths) => {
+ await seed(paths, "alpha", { v1: { "audio.m4a": "one" } });
+ await assert.rejects(
+ () => previewRelocation({ paths, slug: "alpha", root: "platter/media" }),
+ /must be an absolute path/,
+ );
+ });
+});
diff --git a/common/controller/relocateChannelMedia.ts b/common/controller/relocateChannelMedia.ts
@@ -0,0 +1,926 @@
+import path from "node:path";
+import {
+ access,
+ constants as fsConstants,
+ lstat,
+ mkdir,
+ readdir,
+ readFile,
+ readlink,
+ realpath,
+ rename,
+ rm,
+ stat,
+ symlink,
+ unlink,
+ writeFile,
+} from "node:fs/promises";
+import { execa } from "execa";
+import type { Paths } from "../lib/paths";
+import { isSocialChannel } from "../lib/channelConfig";
+import { getFreeBytes } from "../lib/diskSpace";
+import { getSettings } from "../lib/settings";
+import { formatBytes } from "../lib/format";
+import {
+ inspectChannelMedia,
+ relocatedDataDir,
+ relocationMarkerPath,
+ type RelocationDirection,
+ type RelocationMarker,
+ type RelocationPhase,
+} from "../lib/channelMedia";
+import { readChannelConfig, writeChannelConfig } from "./channels";
+
+// MOVE A CHANNEL'S MEDIA TO ANOTHER DRIVE, AND BACK.
+//
+// The mechanism is a symlink (see common/lib/channelMedia.ts for why): the media
+// is copied to `<root>/<slug>/data`, verified, and `channels/<slug>/data` becomes
+// an absolute link to it while `config.json` records the target in `dataDir`.
+// Nothing that reads a channel changes, because the on-disk contract
+// `channelDir/data/<id>/…` is preserved exactly.
+//
+// THE SOURCE IS NEVER TOUCHED UNTIL THE COPY IS VERIFIED. An abort, a full
+// target, a crash, an rsync failure — all of them leave `data/` exactly where it
+// was, and leave the partial copy resumable. The only destructive step is the
+// reclaim at the very end, after the link is in place and the config is written.
+//
+// `rsync -a` preserves mtimes, which is what makes this free for the LMDB index:
+// it stores ids and mtimes, no paths, so a relocated channel needs no reindex.
+
+const BYTES_PER_GB = 1024 ** 3;
+
+export type RelocateChannelMediaResult = {
+ slug: string;
+ direction: RelocationDirection;
+ // Absolute path of the relocated data dir. For "back" this is what was
+ // reclaimed, not where the media now lives.
+ target: string;
+ bytes: number;
+ files: number;
+ // True when the run picked up an interrupted one from its marker rather than
+ // starting from scratch.
+ resumed: boolean;
+};
+
+type RelocateOpts = {
+ paths: Paths;
+ slug: string;
+ direction: RelocationDirection;
+ // Required for "out"; ignored for "back", which reads the target from config.
+ root?: string;
+ onLog?: (line: string) => void;
+ signal?: AbortSignal;
+};
+
+export type RelocationPreview = {
+ slug: string;
+ target: string;
+ bytesToMove: number;
+ files: number;
+ freeOnRoot: number;
+ freeOnSource: number;
+ // The move is pointless (and rename-unsafe assumptions change) when the root
+ // is the volume the corpus is already on.
+ sameDevice: boolean;
+ // A resumable partial copy from an earlier attempt is already at the target.
+ existingPartial: boolean;
+};
+
+// One walk, used by the preview, the space check and the verify. Follows no
+// symlinks (a relocated channel is never the SOURCE of another relocation).
+async function measureTree(
+ dir: string,
+): Promise<{ bytes: number; files: number }> {
+ let bytes = 0;
+ let files = 0;
+ const stack = [dir];
+ while (stack.length > 0) {
+ const cur = stack.pop() as string;
+ let entries;
+ try {
+ entries = await readdir(cur, { withFileTypes: true });
+ } catch {
+ continue;
+ }
+ for (const e of entries) {
+ const p = path.join(cur, e.name);
+ if (e.isDirectory()) {
+ stack.push(p);
+ } else if (e.isFile()) {
+ try {
+ const st = await stat(p);
+ bytes += st.size;
+ files++;
+ } catch {
+ /* vanished mid-walk; the verify is what catches real drift */
+ }
+ }
+ }
+ }
+ return { bytes, files };
+}
+
+async function pathExists(p: string): Promise<boolean> {
+ try {
+ await stat(p);
+ return true;
+ } catch {
+ return false;
+ }
+}
+
+async function isDirectory(p: string): Promise<boolean> {
+ try {
+ return (await stat(p)).isDirectory();
+ } catch {
+ return false;
+ }
+}
+
+// Resolved through every symlink when the path exists, lexically when it does
+// not. A destination root that does not exist yet is refused elsewhere; a root
+// that exists and is a symlink back into the corpus is exactly what this is for.
+async function realOrResolved(p: string): Promise<string> {
+ return await realpath(p).catch(() => path.resolve(p));
+}
+
+// `child` IS `parent`, or lives under it.
+function isWithin(parent: string, child: string): boolean {
+ const rel = path.relative(parent, child);
+ return rel === "" || (!rel.startsWith("..") && !path.isAbsolute(rel));
+}
+
+// WHY A ROOT INSIDE THE CORPUS IS NOT MERELY POINTLESS BUT DESTRUCTIVE.
+//
+// Take `root = <transcriptsDir>/channels`. Then `relocatedDataDir(root, slug)`
+// is `<channels>/<slug>/data` — the SOURCE. `rsync -a src/ src/` succeeds,
+// verifyCopy compares the tree with itself and passes, the swap renames `data/`
+// to `data.relocated-<ts>` (which moves the "target" it just verified), creates
+// a symlink pointing at a path that no longer exists, and the reclaim sweep then
+// `rm -r`s the parked directory — the only copy of the media. Nothing in the
+// happy path can notice, because every check it runs is comparing the tree with
+// itself. The panel's own help text ("an absolute directory that already
+// exists") describes `transcripts/channels` almost word for word.
+//
+// So containment is checked, in one place, and called from all three: the
+// preview the operator reads, the action that enqueues, and the job that moves.
+// A root inside the corpus is refused whether it frees anything or not — the
+// same-device hint stays a hint, and this is a refusal.
+export async function relocationRootProblem(opts: {
+ paths: Paths;
+ slug: string;
+ root: string;
+}): Promise<string | null> {
+ const root = opts.root.trim();
+ if (!root) return "No destination root given";
+ if (!path.isAbsolute(root)) {
+ return `The destination root must be an absolute path (got "${root}")`;
+ }
+ const channelDir = path.join(opts.paths.channelsDir, opts.slug);
+ const [realRoot, realCorpus, realChannel] = await Promise.all([
+ realOrResolved(root),
+ realOrResolved(opts.paths.transcriptsDir),
+ realOrResolved(channelDir),
+ ]);
+ if (isWithin(realCorpus, realRoot)) {
+ return (
+ `The destination root ${root} is inside the corpus at ` +
+ `${opts.paths.transcriptsDir}. A relocation there would copy the channel ` +
+ `onto itself and then reclaim the only copy — pick a directory on the ` +
+ `other drive.`
+ );
+ }
+ // Belt and braces for a root that is outside the corpus but whose `<slug>`
+ // level is a link back into it: the target is resolved separately, because
+ // realpath of the root cannot see through a link one level down.
+ const realTarget = await realOrResolved(relocatedDataDir(realRoot, opts.slug));
+ if (isWithin(realChannel, realTarget) || isWithin(realTarget, realChannel)) {
+ return (
+ `The destination ${relocatedDataDir(root, opts.slug)} resolves inside ` +
+ `the channel directory ${channelDir} — the media would be copied onto ` +
+ `itself and then reclaimed.`
+ );
+ }
+ if (isWithin(realCorpus, realTarget)) {
+ return (
+ `The destination ${relocatedDataDir(root, opts.slug)} resolves inside ` +
+ `the corpus at ${opts.paths.transcriptsDir}. Pick a directory on the ` +
+ `other drive.`
+ );
+ }
+ return null;
+}
+
+// WHAT `data/` ACTUALLY IS RIGHT NOW. lstat, never stat: a dangling symlink —
+// the exact state a half-finished move or an unmounted drive leaves behind —
+// reads as ABSENT through stat, and the code that then tries to create the link
+// fails with EEXIST on a path it was just told was not there.
+//
+// Every phase past the copy dispatches on this rather than on what the previous
+// phase is supposed to have done, which is what makes a rerun idempotent: the
+// marker says how far the last run GOT, the disk says what is actually there,
+// and only the disk is evidence.
+type DataDirState =
+ | { kind: "missing" }
+ | { kind: "real-dir" }
+ | { kind: "link"; linkTarget: string }
+ | { kind: "other" };
+
+async function dataDirState(p: string): Promise<DataDirState> {
+ let st;
+ try {
+ st = await lstat(p);
+ } catch {
+ return { kind: "missing" };
+ }
+ if (st.isSymbolicLink()) {
+ return { kind: "link", linkTarget: await readlink(p).catch(() => "") };
+ }
+ if (st.isDirectory()) return { kind: "real-dir" };
+ return { kind: "other" };
+}
+
+// Every leftover a crashed run can have parked next to `data/`, in one list.
+// The reclaim phase sweeps ALL of them rather than the one name the run that is
+// finishing happens to hold: a crash between the config write and the reclaim
+// marker leaves a full second copy of the channel on the source volume, and
+// reclaiming only the name this process minted would orphan it forever — on the
+// disk the move exists to free.
+async function parkedSiblings(channelDir: string): Promise<string[]> {
+ const names = await readdir(channelDir).catch(() => [] as string[]);
+ return names
+ .filter((n) => n.startsWith("data.relocated-") || n === "data.incoming")
+ .map((n) => path.join(channelDir, n));
+}
+
+async function sweepParked(
+ channelDir: string,
+ log: (m: string) => void,
+): Promise<void> {
+ for (const p of await parkedSiblings(channelDir)) {
+ log(`Reclaiming ${p}`);
+ await rm(p, { recursive: true, force: true });
+ }
+}
+
+async function writeMarker(
+ paths: Paths,
+ slug: string,
+ marker: RelocationMarker,
+): Promise<void> {
+ const file = relocationMarkerPath(paths, slug);
+ const tmp = `${file}.tmp-${process.pid}`;
+ await writeFile(tmp, JSON.stringify(marker, null, 2) + "\n");
+ await rename(tmp, file);
+}
+
+async function clearMarker(paths: Paths, slug: string): Promise<void> {
+ await rm(relocationMarkerPath(paths, slug), { force: true });
+}
+
+async function readMarkerRaw(
+ paths: Paths,
+ slug: string,
+): Promise<RelocationMarker | null> {
+ try {
+ const raw = await readFile(relocationMarkerPath(paths, slug), "utf8");
+ const r = JSON.parse(raw) as Partial<RelocationMarker>;
+ if (typeof r.target !== "string") return null;
+ return {
+ target: r.target,
+ direction: r.direction === "back" ? "back" : "out",
+ startedAt: typeof r.startedAt === "string" ? r.startedAt : "",
+ phase:
+ r.phase === "swap" || r.phase === "reclaim"
+ ? r.phase
+ : ("copy" as RelocationPhase),
+ };
+ } catch {
+ return null;
+ }
+}
+
+// rsync, exactly as backupSavedVideos does it: the real binary from
+// paths.rsyncBin, cancellable, streaming its output into the job log.
+async function rsyncTree(opts: {
+ paths: Paths;
+ src: string;
+ dest: string;
+ args: string[];
+ log: (m: string) => void;
+ signal?: AbortSignal;
+}): Promise<{ exitCode: number; output: string }> {
+ // Trailing slash on src: copy the CONTENTS, so <src>/ -> <dest>/ and not
+ // <dest>/data/. Getting this wrong is a silently nested corpus.
+ const args = [...opts.args, `${opts.src}/`, `${opts.dest}/`];
+ opts.log(`$ ${opts.paths.rsyncBin} ${args.join(" ")}`);
+ const child = execa(opts.paths.rsyncBin, args, {
+ cancelSignal: opts.signal,
+ all: true,
+ buffer: false,
+ reject: false,
+ });
+ let output = "";
+ child.all?.on("data", (c: Buffer) => {
+ const text = c.toString("utf8");
+ output += text;
+ opts.log(text);
+ });
+ const result = await child;
+ return { exitCode: result.exitCode ?? 1, output };
+}
+
+// What a copy has to clear before the swap: rsync itself agrees there is
+// nothing left to send, AND the two trees measure the same. The dry run alone
+// would accept a target that is byte-identical for the wrong reason; the counts
+// alone would accept two trees of equal size with different contents.
+async function verifyCopy(opts: {
+ paths: Paths;
+ src: string;
+ dest: string;
+ log: (m: string) => void;
+ signal?: AbortSignal;
+}): Promise<{ bytes: number; files: number }> {
+ const { exitCode, output } = await rsyncTree({
+ ...opts,
+ args: ["-a", "--dry-run", "--itemize-changes"],
+ });
+ if (exitCode !== 0) {
+ throw new Error(`Verification rsync failed (exit ${exitCode})`);
+ }
+ const drift = output
+ .split("\n")
+ .map((l) => l.trim())
+ .filter((l) => l.length > 0 && !l.startsWith("sending incremental"))
+ .filter((l) => !/^(sent|total size|$)/.test(l));
+ if (drift.length > 0) {
+ throw new Error(
+ `Verification failed: ${drift.length} file(s) still differ ` +
+ `(first: ${drift[0]}). The source has NOT been touched.`,
+ );
+ }
+ const [a, b] = await Promise.all([measureTree(opts.src), measureTree(opts.dest)]);
+ if (a.files !== b.files || a.bytes !== b.bytes) {
+ throw new Error(
+ `Verification failed: source has ${a.files} file(s)/${formatBytes(a.bytes)}, ` +
+ `target has ${b.files}/${formatBytes(b.bytes)}. The source has NOT been touched.`,
+ );
+ }
+ return a;
+}
+
+// What the operator sees before committing to a move. Cheap enough to run on a
+// form keystroke debounce: one tree walk of the channel plus two statfs calls.
+export async function previewRelocation({
+ paths,
+ slug,
+ root,
+}: {
+ paths: Paths;
+ slug: string;
+ root: string;
+}): Promise<RelocationPreview> {
+ // The preview REFUSES rather than reporting numbers for a root the job will
+ // reject: its whole job is to answer "is this root usable" before the
+ // operator commits, and a plausible pair of figures for a destructive root is
+ // the worst possible answer.
+ const problem = await relocationRootProblem({ paths, slug, root });
+ if (problem) throw new Error(problem);
+ // EXISTENCE, HERE AS WELL AS IN THE JOB. getFreeBytes walks up to the nearest
+ // existing ancestor, so a typo'd root statfs's its parent and previews with
+ // perfectly plausible numbers — and the move is then refused by the job, after
+ // the operator has already read a confirmation.
+ if (!(await isDirectory(root))) {
+ throw new Error(
+ `The destination root ${root} does not exist or is not a directory ` +
+ `(is the drive mounted?)`,
+ );
+ }
+ const source = path.join(paths.channelsDir, slug, "data");
+ const target = relocatedDataDir(root, slug);
+ const [measured, freeOnRoot, freeOnSource, existingPartial] = await Promise.all([
+ measureTree(source),
+ getFreeBytes(root),
+ getFreeBytes(paths.channelsDir),
+ pathExists(target),
+ ]);
+ let sameDevice = false;
+ try {
+ const [a, b] = await Promise.all([stat(source), stat(root)]);
+ sameDevice = a.dev === b.dev;
+ } catch {
+ /* an unmounted or absent root is not "same device" */
+ }
+ return {
+ slug,
+ target,
+ bytesToMove: measured.bytes,
+ files: measured.files,
+ freeOnRoot,
+ freeOnSource,
+ sameDevice,
+ existingPartial,
+ };
+}
+
+export async function relocateChannelMedia(
+ opts: RelocateOpts,
+): Promise<RelocateChannelMediaResult> {
+ const { paths, slug, direction, signal } = opts;
+ const log = opts.onLog ?? ((m: string) => console.log(m));
+ const config = await readChannelConfig(paths, slug);
+ if (!config) throw new Error(`Channel "${slug}" not found`);
+ if (isSocialChannel(config)) {
+ throw new Error(
+ `Channel "${slug}" is a social channel — it has no downloaded media to relocate`,
+ );
+ }
+
+ const channelDir = path.join(paths.channelsDir, slug);
+ const dataDir = path.join(channelDir, "data");
+ const existingMarker = await readMarkerRaw(paths, slug);
+
+ if (direction === "back") {
+ // THE MARKER IS THE SECOND SOURCE OF TRUTH FOR THE TARGET, and without it
+ // the resume path was unreachable. moveBack clears config.dataDir as part
+ // of its swap, so a crash after that point left a rerun reading "this
+ // channel is not relocated" from the config and throwing — while a marker
+ // sat next to it naming the very target still holding the media, and
+ // inspect() reported in-transition forever.
+ const resume = existingMarker?.direction === "back" ? existingMarker : null;
+ const target = config.dataDir?.trim() || resume?.target;
+ if (!target) {
+ throw new Error(
+ `Channel "${slug}" is not relocated — its media is already in place`,
+ );
+ }
+ if (existingMarker && existingMarker.target !== target) {
+ throw new Error(
+ `A relocation to ${existingMarker.target} is already in progress for "${slug}"`,
+ );
+ }
+ // A marker for the OTHER direction is never resumed into this one. Same
+ // target, opposite intent: an interrupted move-out has a full copy on the
+ // platter and a real dir here, and treating that as a move-back to resume
+ // would swap the wrong way round.
+ if (existingMarker && !resume) {
+ throw new Error(
+ `A move-out to ${existingMarker.target} is in flight or was interrupted ` +
+ `for "${slug}" — finish that before moving back`,
+ );
+ }
+ return moveBack({
+ paths,
+ slug,
+ config,
+ channelDir,
+ dataDir,
+ target,
+ log,
+ signal,
+ resumed: Boolean(resume),
+ phase: resume?.phase ?? "copy",
+ });
+ }
+
+ const root = opts.root?.trim() ?? "";
+ // Blank, relative, and inside-the-corpus, in one list — see
+ // relocationRootProblem. The action and the preview ask the same question
+ // earlier so the operator does not find out from a job log, but this is the
+ // one that is load-bearing.
+ const rootProblem = await relocationRootProblem({ paths, slug, root });
+ if (rootProblem) throw new Error(rootProblem);
+ const target = relocatedDataDir(root, slug);
+ if (existingMarker && existingMarker.target !== target) {
+ throw new Error(
+ `A relocation to ${existingMarker.target} is already in progress for "${slug}" — ` +
+ `finish or clear it before moving to ${target}`,
+ );
+ }
+ if (config.dataDir?.trim() && config.dataDir.trim() !== target) {
+ throw new Error(
+ `Channel "${slug}" is already relocated to ${config.dataDir.trim()}. ` +
+ `Move it back in place first.`,
+ );
+ }
+
+ const resume = existingMarker?.direction === "out" ? existingMarker : null;
+ if (existingMarker && !resume) {
+ throw new Error(
+ `A move-back from ${existingMarker.target} is in flight or was interrupted ` +
+ `for "${slug}" — finish that before moving out again`,
+ );
+ }
+
+ return moveOut({
+ paths,
+ slug,
+ config,
+ channelDir,
+ dataDir,
+ root,
+ target,
+ log,
+ signal,
+ // `resumed` relaxes the "must be in-place" precondition, so it must mean
+ // "this run continues an interrupted move OUT" and nothing looser.
+ resumed: Boolean(resume),
+ phase: resume?.phase ?? "copy",
+ });
+}
+
+async function moveOut(args: {
+ paths: Paths;
+ slug: string;
+ config: NonNullable<Awaited<ReturnType<typeof readChannelConfig>>>;
+ channelDir: string;
+ dataDir: string;
+ root: string;
+ target: string;
+ log: (m: string) => void;
+ signal?: AbortSignal;
+ resumed: boolean;
+ phase: RelocationPhase;
+}): Promise<RelocateChannelMediaResult> {
+ const { paths, slug, channelDir, dataDir, root, target, log, signal } = args;
+
+ // PREFLIGHT. Everything that can refuse does so here, before a single byte is
+ // written and before the marker exists.
+ if (!(await isDirectory(root))) {
+ throw new Error(
+ `The destination root ${root} does not exist or is not a directory ` +
+ `(is the drive mounted?)`,
+ );
+ }
+ // Writability, checked explicitly rather than discovered by rsync's exit code
+ // three minutes in. A read-only mount is the ordinary way a platter comes
+ // back after a bad shutdown.
+ try {
+ await access(root, fsConstants.W_OK);
+ } catch {
+ throw new Error(`The destination root ${root} is not writable`);
+ }
+ // A RESUMING RUN MUST TOLERATE ITS OWN MARKER. `inspectChannelMedia` reports
+ // "in-transition" for any channel carrying one, which is exactly right for
+ // every guard and exactly wrong here: the rerun that finishes an interrupted
+ // move is the one caller allowed to see it. A marker for a DIFFERENT target
+ // was already refused above.
+ const location = await inspectChannelMedia(paths, slug, args.config);
+ if (!args.resumed && location.status !== "in-place") {
+ throw new Error(
+ `Channel "${slug}" is not in a movable state: ${
+ location.detail ?? location.status
+ }`,
+ );
+ }
+
+ // A channel that has downloaded nothing has no data/ at all, and rsync exits
+ // 23 on a missing source — an error message about a partial transfer for
+ // something that is not a transfer. Say what is actually the matter.
+ if (args.phase === "copy" && !(await isDirectory(dataDir))) {
+ throw new Error(
+ `Channel "${slug}" has no ${dataDir} to move — nothing has been ` +
+ `downloaded for it yet`,
+ );
+ }
+
+ const measured = await measureTree(dataDir);
+ log(
+ `Relocating ${slug}: ${measured.files} file(s), ${formatBytes(measured.bytes)} ` +
+ `-> ${target}`,
+ );
+
+ let phase = args.phase;
+ if (phase === "copy") {
+ // THE BAR IS THE BYTES PLUS THE RESUME MARGIN, not the bytes. Landing the
+ // media with nothing to spare puts the destination volume under the disk
+ // gate's own floor the moment it arrives, so the channel's next download is
+ // refused by the gate on the drive it was just moved to. The margin is the
+ // operator's configured one, so the two numbers cannot drift apart — and it
+ // is ZERO when the gate is switched off, because the margin exists to clear
+ // a bar that then does not exist. Demanding headroom for a gate nobody
+ // armed would refuse a move on a disk with room for it.
+ const settings = getSettings();
+ const marginGB =
+ settings.minFreeDiskGB > 0 ? settings.resumeMarginGB : 0;
+ const needed = measured.bytes + marginGB * BYTES_PER_GB;
+ const free = await getFreeBytes(root);
+ if (free < needed) {
+ throw new Error(
+ `Not enough space on ${root}: ${formatBytes(free)} free, ` +
+ `${formatBytes(measured.bytes)} to move` +
+ (marginGB > 0
+ ? ` plus a ${marginGB} GB resume margin = ${formatBytes(needed)} required`
+ : " required"),
+ );
+ }
+ await mkdir(target, { recursive: true });
+ await writeMarker(paths, slug, {
+ target,
+ direction: "out",
+ startedAt: new Date().toISOString(),
+ phase: "copy",
+ });
+ // --partial keeps an aborted transfer resumable; -a preserves mtimes, which
+ // is what makes the LMDB index a no-op afterwards.
+ const { exitCode } = await rsyncTree({
+ paths,
+ src: dataDir,
+ dest: target,
+ args: ["-a", "--partial", "--info=progress2"],
+ log,
+ signal,
+ });
+ if (signal?.aborted) {
+ throw new Error(
+ `Cancelled. ${dataDir} is untouched and the partial copy at ${target} ` +
+ `is resumable — rerun to continue.`,
+ );
+ }
+ if (exitCode !== 0) throw new Error(`rsync failed (exit ${exitCode})`);
+
+ log("Verifying the copy…");
+ await verifyCopy({ paths, src: dataDir, dest: target, log, signal });
+ phase = "swap";
+ }
+
+ if (phase === "swap") {
+ await writeMarker(paths, slug, {
+ target,
+ direction: "out",
+ startedAt: new Date().toISOString(),
+ phase: "swap",
+ });
+
+ // EVERY STEP BELOW OBSERVES THE DISK INSTEAD OF ASSUMING THE LAST ONE RAN.
+ // The marker says how far the previous attempt got; it cannot say how far
+ // it got THROUGH a phase, and a crash lands between any two syscalls.
+ const state = await dataDirState(dataDir);
+ const parked = (await parkedSiblings(channelDir)).filter((p) =>
+ path.basename(p).startsWith("data.relocated-"),
+ );
+
+ // Re-verify, because a resumed run did not do the copy in this process and
+ // must not take the interrupted one's word for it. The source to verify
+ // against is whichever copy of it still exists; once the swap has committed
+ // there is none, and there is nothing left to check.
+ const verifySrc =
+ state.kind === "real-dir" ? dataDir : (parked[0] ?? null);
+ if (verifySrc) {
+ log("Verifying the copy…");
+ await verifyCopy({ paths, src: verifySrc, dest: target, log, signal });
+ }
+
+ if (state.kind === "real-dir") {
+ // Same device, so the rename is atomic: `data/` is a real dir one instant
+ // and the parked copy the next, never half of each. A fresh name even
+ // when a parked dir already exists — renaming onto a non-empty directory
+ // is ENOTEMPTY, and the reclaim sweep takes all of them anyway.
+ await rename(dataDir, path.join(channelDir, `data.relocated-${Date.now()}`));
+ } else if (state.kind === "other") {
+ throw new Error(
+ `${dataDir} is neither a directory nor a symlink — refusing to replace it`,
+ );
+ }
+
+ const after = await dataDirState(dataDir);
+ if (
+ after.kind === "link" &&
+ path.resolve(after.linkTarget) !== path.resolve(target)
+ ) {
+ // A link to somewhere else is not this move's work to reinterpret, and
+ // silently repointing it would strand whatever it does point at.
+ throw new Error(
+ `${dataDir} already points at ${after.linkTarget}, not ${target}`,
+ );
+ }
+ if (after.kind === "missing") {
+ await symlink(target, dataDir);
+ }
+
+ // Written only now, on success: config.dataDir is a record of what is on
+ // disk, never an intention. Skipped when it already says so, so a rerun
+ // does not rewrite a file it agrees with.
+ const fresh = (await readChannelConfig(paths, slug)) ?? args.config;
+ if (fresh.dataDir?.trim() !== target) {
+ await writeChannelConfig(paths, slug, { ...fresh, dataDir: target });
+ }
+ log(`Swapped: ${dataDir} -> ${target}`);
+ await writeMarker(paths, slug, {
+ target,
+ direction: "out",
+ startedAt: new Date().toISOString(),
+ phase: "reclaim",
+ });
+ }
+
+ // RECLAIM RUNS ON EVERY PATH, not only on a resume. A crash between the
+ // config write and the reclaim marker used to orphan a full second copy of
+ // the channel on the source volume — the disk the move exists to free — and
+ // the sweep that would have caught it was in the branch a fresh run never
+ // takes. It sweeps every sibling, not the one name this process minted.
+ await sweepParked(channelDir, log);
+
+ await clearMarker(paths, slug);
+ const freeNow = await getFreeBytes(paths.channelsDir);
+ log(
+ `Done. ${formatBytes(measured.bytes)} now on ${root}; ` +
+ `${formatBytes(freeNow)} free on the source volume.`,
+ );
+ return {
+ slug,
+ direction: "out",
+ target,
+ bytes: measured.bytes,
+ files: measured.files,
+ resumed: args.resumed,
+ };
+}
+
+async function moveBack(args: {
+ paths: Paths;
+ slug: string;
+ config: NonNullable<Awaited<ReturnType<typeof readChannelConfig>>>;
+ channelDir: string;
+ dataDir: string;
+ target: string;
+ log: (m: string) => void;
+ signal?: AbortSignal;
+ resumed: boolean;
+ phase: RelocationPhase;
+}): Promise<RelocateChannelMediaResult> {
+ const { paths, slug, channelDir, dataDir, target, log, signal } = args;
+ const incoming = path.join(channelDir, "data.incoming");
+
+ // THE MIRROR OF moveOut's "must be in-place", and it is not symmetry for its
+ // own sake: move back is the only direction that ENDS by deleting the target.
+ //
+ // `relocated` is true for `inconsistent` and `unreachable` as well as `ok` —
+ // config.dataDir is set in all three — so without this the UI offers Move back
+ // for a channel whose config records a target while `data/` is a REAL
+ // directory. That state is not hypothetical: it is what `rsync --copy-links`
+ // of a channel produces, which WORKTREES.md documents as the way to carry
+ // media into a shard. The run then copies the target to `data.incoming`,
+ // verifies it, finds `data/` already a real dir, logs "the swap had
+ // completed", clears the config, `rm -r`s the target and finally sweeps
+ // `data.incoming` — three copies in, zero out.
+ //
+ // `in-transition` is allowed because a marker is what a resume carries, and a
+ // rerun is the caller this precondition must not refuse.
+ const location = await inspectChannelMedia(paths, slug, args.config);
+ if (
+ !args.resumed &&
+ location.status !== "ok" &&
+ location.status !== "in-transition"
+ ) {
+ throw new Error(
+ `Channel "${slug}" is not in a movable state: ${
+ location.detail ?? location.status
+ }. Nothing has been touched.`,
+ );
+ }
+ // THE PRECONDITION BELONGS TO THE COPY, NOT TO THE RERUN. Past the swap the
+ // media is already back in the channel dir and the target may well be gone —
+ // demanding it be reachable there would refuse the very run that finishes
+ // cleaning up after an interrupted move.
+ if (args.phase === "copy" && !(await isDirectory(target))) {
+ throw new Error(
+ `The relocated media at ${target} is not reachable (is the drive mounted?)`,
+ );
+ }
+ // Measured off whichever copy still exists, in the order they stop existing.
+ const measured = (await isDirectory(target))
+ ? await measureTree(target)
+ : (await isDirectory(incoming))
+ ? await measureTree(incoming)
+ : await measureTree(dataDir);
+ log(
+ `Moving ${slug} back in place: ${measured.files} file(s), ` +
+ `${formatBytes(measured.bytes)} <- ${target}`,
+ );
+
+ let phase = args.phase;
+ if (phase === "copy") {
+ // THE SAME BAR MOVE-OUT USES, and for the same reason: landing the media
+ // with nothing to spare puts the corpus volume under the disk gate's floor
+ // the moment it arrives. And a resumed move-back must only be charged for
+ // what is still MISSING — the bytes already sitting in `data.incoming` are
+ // not about to be written twice, and counting them refused reruns on a disk
+ // that had room for the remainder.
+ const settings = getSettings();
+ const marginGB = settings.minFreeDiskGB > 0 ? settings.resumeMarginGB : 0;
+ const already = (await isDirectory(incoming))
+ ? (await measureTree(incoming)).bytes
+ : 0;
+ const needed =
+ Math.max(0, measured.bytes - already) + marginGB * BYTES_PER_GB;
+ const free = await getFreeBytes(paths.channelsDir);
+ if (free < needed) {
+ throw new Error(
+ `Not enough space on the corpus volume: ${formatBytes(free)} free, ` +
+ `${formatBytes(Math.max(0, measured.bytes - already))} still to move back` +
+ (marginGB > 0
+ ? ` plus a ${marginGB} GB resume margin = ${formatBytes(needed)} required`
+ : " required"),
+ );
+ }
+ await mkdir(incoming, { recursive: true });
+ await writeMarker(paths, slug, {
+ target,
+ direction: "back",
+ startedAt: new Date().toISOString(),
+ phase: "copy",
+ });
+ const { exitCode } = await rsyncTree({
+ paths,
+ src: target,
+ dest: incoming,
+ args: ["-a", "--partial", "--info=progress2"],
+ log,
+ signal,
+ });
+ if (signal?.aborted) {
+ throw new Error(
+ `Cancelled. ${target} is untouched and ${incoming} is resumable — ` +
+ `rerun to continue.`,
+ );
+ }
+ if (exitCode !== 0) throw new Error(`rsync failed (exit ${exitCode})`);
+ log("Verifying the copy…");
+ await verifyCopy({ paths, src: target, dest: incoming, log, signal });
+ phase = "swap";
+ }
+
+ if (phase === "swap") {
+ await writeMarker(paths, slug, {
+ target,
+ direction: "back",
+ startedAt: new Date().toISOString(),
+ phase: "swap",
+ });
+
+ // OBSERVE, DO NOT ASSUME. The old sequence was `unlink(data)` (swallowing
+ // its error) then `rename(incoming, data)`, which is idempotent in exactly
+ // the case that never happens: a crash AFTER the rename left `data/` a real
+ // directory, the swallowed unlink then failed on it, and the rename ENOENTed
+ // on an `incoming` that no longer existed — forever, on every rerun.
+ const state = await dataDirState(dataDir);
+ if (state.kind === "real-dir") {
+ // The rename already committed. Nothing to swap; the leftovers are the
+ // reclaim's business.
+ log(`${dataDir} is already a real directory — the swap had completed`);
+ } else if (state.kind === "other") {
+ // A regular file (or a socket, or a fifo) where `data/` should be is not
+ // a link to replace and not a directory to keep. moveOut refuses the same
+ // shape at its own swap; refusing here too is what keeps `unlink` below
+ // meaning "remove the symlink" and nothing else.
+ throw new Error(
+ `${dataDir} is neither a directory nor a symlink — refusing to replace it`,
+ );
+ } else {
+ if (!(await isDirectory(incoming))) {
+ throw new Error(
+ `Cannot finish moving "${slug}" back: ${dataDir} is not a directory ` +
+ `and there is no verified copy at ${incoming}`,
+ );
+ }
+ // unlink, not rm -r: `data` is the LINK here, and removing it recursively
+ // would be the one way this whole design eats the media.
+ if (state.kind !== "missing") await unlink(dataDir);
+ await rename(incoming, dataDir);
+ log(`Swapped: ${dataDir} is a real directory again`);
+ }
+
+ const fresh = await readChannelConfig(paths, slug);
+ if (fresh?.dataDir !== undefined) {
+ const { dataDir: _dropped, ...rest } = fresh;
+ await writeChannelConfig(paths, slug, rest);
+ }
+ await writeMarker(paths, slug, {
+ target,
+ direction: "back",
+ startedAt: new Date().toISOString(),
+ phase: "reclaim",
+ });
+ }
+
+ await rm(target, { recursive: true, force: true });
+ // Leave <root>/<slug> behind only if something else is in it.
+ const slugRoot = path.dirname(target);
+ if ((await readdir(slugRoot).catch(() => ["keep"])).length === 0) {
+ await rm(slugRoot, { recursive: true, force: true });
+ }
+ // Any half-copied `data.incoming` (or a parked dir from an earlier move out)
+ // goes with it — the same sweep, for the same reason.
+ await sweepParked(channelDir, log);
+ await clearMarker(paths, slug);
+ log(`Done. ${formatBytes(measured.bytes)} back in place.`);
+ return {
+ slug,
+ direction: "back",
+ target,
+ bytes: measured.bytes,
+ files: measured.files,
+ resumed: args.resumed,
+ };
+}
diff --git a/common/controller/renameChannel.test.ts b/common/controller/renameChannel.test.ts
@@ -1,6 +1,15 @@
import { test } from "node:test";
import assert from "node:assert/strict";
-import { mkdir, mkdtemp, rm, stat, writeFile } from "node:fs/promises";
+import {
+ mkdir,
+ mkdtemp,
+ readFile,
+ readlink,
+ rm,
+ stat,
+ symlink,
+ writeFile,
+} from "node:fs/promises";
import { tmpdir } from "node:os";
import path from "node:path";
import type { Paths } from "../lib/paths";
@@ -22,6 +31,11 @@ import {
writeSchedulerState,
} from "../jobs/syncSchedulerState";
import { renameChannel } from "./renameChannel";
+import {
+ inspectChannelMedia,
+ relocatedDataDir,
+ RELOCATION_MARKER_FILENAME,
+} from "../lib/channelMedia";
// Run with:
// pnpm --filter yt-dlp-transcript-common exec tsx --test controller/renameChannel.test.ts
@@ -133,3 +147,113 @@ test("renameChannel rejects invalid, same, and existing targets", async () => {
assert.equal(await channelExists(paths, "old"), true);
});
});
+
+// --- relocated media ------------------------------------------------------
+
+test("rename re-points a convention-shaped relocated media dir", async () => {
+ await withPaths(async (paths) => {
+ const dir = path.dirname(paths.channelsDir);
+ const mediaRoot = path.join(dir, "platter");
+ const target = relocatedDataDir(mediaRoot, "old");
+ await writeChannelConfig(paths, "old", { ...config, dataDir: target });
+ await mkdir(path.join(target, "vid1"), { recursive: true });
+ await writeFile(path.join(target, "vid1", "audio.m4a"), "BYTES");
+ await symlink(target, path.join(paths.channelsDir, "old", "data"));
+
+ const result = await renameChannel(paths, "old", "new", {
+ ...config,
+ dataDir: target,
+ });
+ assert.deepEqual(result.warnings, []);
+
+ const newTarget = relocatedDataDir(mediaRoot, "new");
+ assert.equal((await readChannelConfig(paths, "new"))?.dataDir, newTarget);
+ assert.equal(
+ await readlink(path.join(paths.channelsDir, "new", "data")),
+ newTarget,
+ );
+ // The media reads through the new link at the old on-disk contract path.
+ assert.equal(
+ await readFile(
+ path.join(paths.channelsDir, "new", "data", "vid1", "audio.m4a"),
+ "utf8",
+ ),
+ "BYTES",
+ );
+ assert.equal(
+ (await inspectChannelMedia(paths, "new")).status,
+ "ok",
+ );
+ // The old <root>/<slug> is gone, not left as a duplicate.
+ await assert.rejects(() => stat(path.join(mediaRoot, "old")));
+ });
+});
+
+test("a media dir that does not follow the convention is left alone", async () => {
+ await withPaths(async (paths) => {
+ const dir = path.dirname(paths.channelsDir);
+ // <root>/<something-else>/data — the link is absolute and still works after
+ // the rename, so moving a directory whose name is not ours would be worse
+ // than leaving it.
+ const target = path.join(dir, "platter", "handpicked", "data");
+ await writeChannelConfig(paths, "old", { ...config, dataDir: target });
+ await mkdir(path.join(target, "vid1"), { recursive: true });
+ await symlink(target, path.join(paths.channelsDir, "old", "data"));
+
+ await renameChannel(paths, "old", "new", { ...config, dataDir: target });
+
+ assert.equal((await readChannelConfig(paths, "new"))?.dataDir, target);
+ assert.equal(
+ await readlink(path.join(paths.channelsDir, "new", "data")),
+ target,
+ );
+ await stat(path.join(target, "vid1"));
+ });
+});
+
+test("rename rolls the channel dir back when the media move fails", async () => {
+ await withPaths(async (paths) => {
+ const dir = path.dirname(paths.channelsDir);
+ const mediaRoot = path.join(dir, "platter");
+ const target = relocatedDataDir(mediaRoot, "old");
+ await writeChannelConfig(paths, "old", { ...config, dataDir: target });
+ await mkdir(path.join(target, "vid1"), { recursive: true });
+ await symlink(target, path.join(paths.channelsDir, "old", "data"));
+ // Something is already sitting at <root>/new.
+ await mkdir(path.join(mediaRoot, "new"), { recursive: true });
+
+ await assert.rejects(
+ () => renameChannel(paths, "old", "new", { ...config, dataDir: target }),
+ /already exists/,
+ );
+ // Nothing half-renamed: the channel is still "old", still linked, still
+ // pointing at its media.
+ assert.equal(await channelExists(paths, "old"), true);
+ assert.equal(await channelExists(paths, "new"), false);
+ assert.equal(
+ await readlink(path.join(paths.channelsDir, "old", "data")),
+ target,
+ );
+ await stat(path.join(target, "vid1"));
+ });
+});
+
+test("rename is refused while a relocation is in flight", async () => {
+ await withPaths(async (paths) => {
+ await writeChannelConfig(paths, "old", config);
+ await writeFile(
+ path.join(paths.channelsDir, "old", RELOCATION_MARKER_FILENAME),
+ JSON.stringify({
+ target: "/mnt/platter/old/data",
+ direction: "out",
+ startedAt: new Date().toISOString(),
+ phase: "copy",
+ }),
+ );
+ await assert.rejects(
+ () => renameChannel(paths, "old", "new", config),
+ /relocation in progress/,
+ );
+ assert.equal(await channelExists(paths, "old"), true);
+ });
+});
diff --git a/common/controller/renameChannel.ts b/common/controller/renameChannel.ts
@@ -1,9 +1,18 @@
import path from "node:path";
-import { readdir, rename, stat } from "node:fs/promises";
+import { readdir, rename, stat, symlink, unlink } from "node:fs/promises";
import type { Paths } from "../lib/paths";
import type { ChannelConfig } from "../lib/channelConfig";
-import { channelExists, isValidChannelSlug } from "./channels";
+import {
+ channelExists,
+ isValidChannelSlug,
+ readChannelConfig,
+ writeChannelConfig,
+} from "./channels";
import { savedVideoRoot } from "../lib/savedVideo";
+import {
+ readRelocationMarker,
+ relocatedDataDir,
+} from "../lib/channelMedia";
import { rewriteSavedVideoDir } from "../lib/savedVideo-server";
import { getSite, listSiteIds, writeSite } from "../lib/site";
import {
@@ -15,6 +24,8 @@ import {
// (transcripts/channels/<slug>/), this moves the channel directory AND migrates
// every other store that keys by slug and would otherwise be orphaned:
// - the saved-video store dir + each saved-video.json pointer's absolute `dir`
+// - a relocated media dir on another drive, when it follows the
+// <root>/<slug>/data convention, plus the symlink and config.dataDir
// - site.json memberships across all sites
// - the sync scheduler's per-channel backoff state
//
@@ -70,6 +81,16 @@ export async function renameChannel(
throw new Error(`A directory already exists at channels/${newSlug}`);
}
+ // A rename mid-relocation would move the channel dir out from under a running
+ // copy and leave the marker pointing at a target named for the old slug.
+ const marker = await readRelocationMarker(paths, oldSlug);
+ if (marker) {
+ throw new Error(
+ `Channel "${oldSlug}" has a media relocation in progress (phase ` +
+ `"${marker.phase}"). Finish or cancel it before renaming.`,
+ );
+ }
+
const storeRoot = savedVideoRoot(paths, config);
const oldStoreDir = path.join(storeRoot, oldSlug);
const newStoreDir = path.join(storeRoot, newSlug);
@@ -101,9 +122,88 @@ export async function renameChannel(
}
}
+ // 3. Move the relocated media dir when it follows the <root>/<slug>/data
+ // convention, and re-point the symlink at it. The link is ABSOLUTE, so a
+ // target that does NOT follow the convention is deliberately left alone —
+ // it still works, and moving someone else's directory because its name
+ // happened to match would be worse than leaving it. Rolls the channel-dir
+ // (and store) move back on failure, same shape as step 2.
+ const relocated = config.dataDir?.trim();
+ const conventional =
+ relocated && relocated === relocatedDataDir(path.dirname(path.dirname(relocated)), oldSlug)
+ ? relocated
+ : null;
+ if (conventional) {
+ const mediaRoot = path.dirname(path.dirname(conventional));
+ const newTarget = relocatedDataDir(mediaRoot, newSlug);
+ // A ROLLBACK MUST ONLY UNDO WHAT ACTUALLY RAN. The pre-check below fails
+ // BECAUSE something unrelated already occupies <root>/<newSlug> — and the
+ // old catch then renamed that stranger to <root>/<oldSlug>, destroying a
+ // directory this function had never touched, in the name of undoing a move
+ // it had not made. Each step records that it happened; the catch replays
+ // only those, in reverse.
+ let movedMedia = false;
+ // THE UNLINK IS A STEP TOO. It was untracked, so a throw from the symlink
+ // below (EACCES on a read-only channel dir, ENOSPC) rolled the media
+ // directory back while `data/` stayed DELETED — config still naming the old
+ // target, nothing on disk pointing at it: `inconsistent`, which is now a
+ // state move-back refuses. The catch replays only what ran, and removing the
+ // link ran.
+ let unlinked = false;
+ let relinked = false;
+ try {
+ if (await pathExists(path.dirname(newTarget))) {
+ throw new Error(`A media directory already exists at ${path.dirname(newTarget)}`);
+ }
+ await rename(path.dirname(conventional), path.dirname(newTarget));
+ movedMedia = true;
+ await unlink(path.join(newChannelDir, "data"))
+ .then(() => {
+ unlinked = true;
+ })
+ .catch(() => {});
+ await symlink(newTarget, path.join(newChannelDir, "data"));
+ relinked = true;
+ // Re-read: the channel dir has already moved, so this is the file that
+ // will actually be on disk afterwards.
+ const fresh = (await readChannelConfig(paths, newSlug)) ?? config;
+ await writeChannelConfig(paths, newSlug, {
+ ...fresh,
+ dataDir: newTarget,
+ });
+ } catch (err) {
+ // The link goes back too, and before the directory under it moves: a
+ // failure between the symlink and the config write left `data/` pointing
+ // at <root>/<newSlug>/data while everything else was rolled back to the
+ // old slug — a dangling link, which reads as an unmounted drive.
+ if (relinked || unlinked) {
+ await unlink(path.join(newChannelDir, "data")).catch(() => {});
+ await symlink(conventional, path.join(newChannelDir, "data")).catch(
+ () => {},
+ );
+ }
+ if (movedMedia) {
+ await rename(path.dirname(newTarget), path.dirname(conventional)).catch(
+ () => {},
+ );
+ }
+ if (hadStore) {
+ await rename(newStoreDir, oldStoreDir).catch(() => {});
+ }
+ await rename(newChannelDir, oldChannelDir).catch(() => {});
+ if ((err as NodeJS.ErrnoException).code === "EXDEV") {
+ throw new Error(
+ `Cannot rename across filesystems: the relocated media at ${conventional} ` +
+ `is on a different device from its own root. Move it manually, then retry.`,
+ );
+ }
+ throw err;
+ }
+ }
+
const warnings: string[] = [];
- // 3. Repoint each saved-video.json at the moved store dir. The pointer stores
+ // 4. Repoint each saved-video.json at the moved store dir. The pointer stores
// an absolute `dir` that includes the slug, and resolveSavedVideo trusts it
// verbatim, so a stale `dir` makes persisted source videos unresolvable.
if (hadStore) {
@@ -127,7 +227,7 @@ export async function renameChannel(
}
}
- // 4. Rewrite site memberships that reference the old slug.
+ // 5. Rewrite site memberships that reference the old slug.
try {
for (const siteId of listSiteIds(paths)) {
const site = getSite(siteId, paths);
@@ -141,7 +241,7 @@ export async function renameChannel(
warnings.push(`Site membership update failed: ${(err as Error).message}`);
}
- // 5. Move the scheduler's per-channel backoff entry so auto-sync state carries
+ // 6. Move the scheduler's per-channel backoff entry so auto-sync state carries
// over (the historical run log is left as-is — it's observability only).
try {
const state = await readSchedulerState(paths);
diff --git a/common/jobs/jobKinds.ts b/common/jobs/jobKinds.ts
@@ -37,6 +37,27 @@ export type JobKindMeta = {
// Fallback scheduler tier when a record is not explicitly background. Left
// undefined today so Phase 4 maps it to "foreground" (unchanged behavior).
defaultTier?: SchedulerTier;
+ // Whether this kind reads or writes files under `channels/<slug>/data/`.
+ // Declarative, and consumed by exactly one thing: runManagedFunction refuses
+ // to enqueue a media kind for a channel whose media is not reachable (a
+ // relocated channel whose drive is unmounted, or one mid-relocation) rather
+ // than letting it read an empty dir as the truth. See
+ // common/lib/channelMedia.ts.
+ //
+ // ABSENT MEANS FALSE, and that is deliberate rather than lazy. The guard is
+ // opt-in so a kind that is not listed keeps exactly its current behavior, and
+ // so that `refresh-report` — which is not in this table at all — still runs
+ // and lets the snapshot generator itself report the reason it refused. A kind
+ // that genuinely never opens a video dir (store playlist, clear markers) must
+ // not be refused for a drive it does not read.
+ //
+ // THE TEST IS "DOES IT OPEN data/", NOT "IS IT BOOKKEEPING". Two kinds that
+ // read as bookkeeping declare it anyway, and both were misses:
+ // `normalize-transcripts` walks every video dir and writes a sidecar into
+ // each, and `sync` reads the data dir to decide what is already downloaded
+ // before downloading into it. Against an unmounted drive the first reports a
+ // clean run over zero videos and the second concludes nothing is downloaded.
+ needsMedia?: boolean;
};
// One entry per kind known to the system. `label` is included only where the
@@ -48,6 +69,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: true,
replayable: false,
queueKeyStrategy: "parallel",
+ needsMedia: true,
},
"auto-download": {
kind: "auto-download",
@@ -55,6 +77,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: true,
replayable: false,
queueKeyStrategy: "parallel",
+ needsMedia: true,
},
"auto-download-unit": {
kind: "auto-download-unit",
@@ -62,6 +85,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: false,
replayable: false,
queueKeyStrategy: "platform",
+ needsMedia: true,
},
// The two OPERATION lanes' runners. Same shape as the two above — one
// long-lived job per lane on queueKey "", drainable, never replayable — with
@@ -77,6 +101,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: true,
replayable: false,
queueKeyStrategy: "parallel",
+ needsMedia: true,
},
"auto-backfill": {
kind: "auto-backfill",
@@ -84,6 +109,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: true,
replayable: false,
queueKeyStrategy: "parallel",
+ needsMedia: true,
},
"whisper-all": {
kind: "whisper-all",
@@ -91,6 +117,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: true,
replayable: true,
queueKeyStrategy: "custom",
+ needsMedia: true,
},
"whisper-bucket-downloaded-no-transcript": {
kind: "whisper-bucket-downloaded-no-transcript",
@@ -98,6 +125,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: true,
replayable: true,
queueKeyStrategy: "custom",
+ needsMedia: true,
},
// Replace-auto-captions lane, transcribe half: whisper over videos whose only
// transcript is a YouTube ASR VTT (the downloadedAutoSubsOnly bucket). Same
@@ -109,6 +137,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: true,
replayable: true,
queueKeyStrategy: "custom",
+ needsMedia: true,
},
// Delete the superseded English ASR VTTs kept as backups next to a finished
// whisper transcript. Manual only — never auto-queued — and the single
@@ -120,6 +149,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: false,
replayable: true,
queueKeyStrategy: "custom",
+ needsMedia: true,
},
// AI digest sweep, local (ollama) lane — the one that carries the corpus. Both
// digest kinds are drainable (the batch honors the drain signal: it stops
@@ -131,6 +161,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: true,
replayable: true,
queueKeyStrategy: "custom",
+ needsMedia: true,
},
// Same batch, metered lane. Off unless settings.digest.remoteEnabled is true,
// and it lands on its own queue key so it runs CONCURRENTLY with the local lane
@@ -141,6 +172,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: true,
replayable: true,
queueKeyStrategy: "custom",
+ needsMedia: true,
},
// Copy a duplicate cluster's canonical digest onto its aligned mirrors. A fast
// file operation gated by the timestamp-alignment check, so it is not drainable
@@ -151,6 +183,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: false,
replayable: false,
queueKeyStrategy: "parallel",
+ needsMedia: true,
},
// Write the compact transcript.cues.json sidecar next to every raw transcript
// that lacks a current one — corpus-wide from the Pool on /sites, or one
@@ -163,12 +196,16 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
// write, so "let the in-flight one finish" is already how it behaves. Not
// replayable either — that would need a JobSpec and a jobReplayRegistry
// handler, and the button is one click from the card that reports the count.
+ // needsMedia: it walks `data/<id>/` for every video and writes a sidecar into
+ // each. Against an unmounted drive it finds nothing, reports a clean
+ // 0/0/0/0 run and moves on — a sweep that silently skips a channel.
"normalize-transcripts": {
kind: "normalize-transcripts",
label: "Normalize transcripts",
drainable: false,
replayable: false,
queueKeyStrategy: "custom",
+ needsMedia: true,
},
"redownload-incomplete-bucket": {
kind: "redownload-incomplete-bucket",
@@ -176,6 +213,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: true,
replayable: true,
queueKeyStrategy: "custom",
+ needsMedia: true,
},
"download-from-playlist": {
kind: "download-from-playlist",
@@ -183,6 +221,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: true,
replayable: true,
queueKeyStrategy: "platform",
+ needsMedia: true,
},
"download-missing": {
kind: "download-missing",
@@ -190,6 +229,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: true,
replayable: true,
queueKeyStrategy: "platform",
+ needsMedia: true,
},
"download-missing-subs": {
kind: "download-missing-subs",
@@ -197,6 +237,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: true,
replayable: true,
queueKeyStrategy: "platform",
+ needsMedia: true,
},
"import-one": {
kind: "import-one",
@@ -204,6 +245,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: false,
replayable: false,
queueKeyStrategy: "custom",
+ needsMedia: true,
},
"redownload-archive": {
kind: "redownload-archive",
@@ -211,6 +253,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: false,
replayable: false,
queueKeyStrategy: "custom",
+ needsMedia: true,
},
"retry-bucket": {
kind: "retry-bucket",
@@ -218,6 +261,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: true,
replayable: true,
queueKeyStrategy: "platform",
+ needsMedia: true,
},
"clean-audio-transcribed": {
kind: "clean-audio-transcribed",
@@ -225,6 +269,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: false,
replayable: true,
queueKeyStrategy: "custom",
+ needsMedia: true,
},
// Speaker-diarization backfill over a channel's retained audio. The capture
// lane's catch-all: it picks up everything the post-transcribe hook missed
@@ -239,6 +284,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: true,
replayable: true,
queueKeyStrategy: "custom",
+ needsMedia: true,
},
// One channel through the backfill lane. Drainable (the batch stops taking new
// videos and lets the in-flight one finish) and replayable, because it
@@ -250,6 +296,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: true,
replayable: true,
queueKeyStrategy: "custom",
+ needsMedia: true,
},
// The corpus-wide corrupt-media scan and its per-channel twin. Registered
// properly, unlike check-availability / refresh-report / detect-duplicates,
@@ -265,6 +312,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
// escape hatch refresh-report and detect-duplicates use.
queueKeyStrategy: "parallel",
defaultTier: "background",
+ needsMedia: true,
},
"scan-media-channel": {
kind: "scan-media-channel",
@@ -273,6 +321,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
replayable: false,
queueKeyStrategy: "parallel",
defaultTier: "background",
+ needsMedia: true,
},
"check-kept-deleted": {
kind: "check-kept-deleted",
@@ -280,6 +329,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: false,
replayable: true,
queueKeyStrategy: "custom",
+ needsMedia: true,
},
"persist-kept": {
kind: "persist-kept",
@@ -287,6 +337,7 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: false,
replayable: true,
queueKeyStrategy: "custom",
+ needsMedia: true,
},
"backup-saved-videos": {
kind: "backup-saved-videos",
@@ -302,6 +353,39 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
replayable: false,
queueKeyStrategy: "custom",
},
+ // MOVE A CHANNEL'S MEDIA TO ANOTHER DRIVE, AND BACK
+ // (plans/relocate-channel-media.md).
+ //
+ // `needsMedia: false` is written out rather than omitted, and this is the one
+ // entry where the explicit `false` earns its line: this kind is what FIXES an
+ // unreachable channel. Guard 1 in runManagedFunction refuses a media kind for
+ // a channel whose media it cannot reach — so a kind that declared `true` here
+ // would be refused precisely when the operator needs it (a `back` run after
+ // remounting, or a retry of an interrupted move), and the only action that can
+ // clear the condition would be the one the condition blocks.
+ //
+ // Not drainable: the work is one rsync child, and stopping it is a cancel —
+ // which the AbortSignal already does, leaving the source untouched and the
+ // partial copy resumable. There is no "stop starting new sub-operations" to
+ // honor. Not replayable: a replay carries no direction and no root, and
+ // re-running a move against a channel that has since moved is not a retry.
+ //
+ // Queue key is relocationQueueKey() (set by the one enqueue both actions
+ // share), so EVERY relocation in the process serializes against every other
+ // one: the registry caps a key at concurrency 1 and caps nothing across keys,
+ // and a bulk move's jobs all write to the same destination volume and all run
+ // their space check when they START. Serializing against the channel's own
+ // bookkeeping jobs is not what the key buys — both actions refuse a channel
+ // that has running or queued jobs before they enqueue, which refuses rather
+ // than waits.
+ "relocate-channel-media": {
+ kind: "relocate-channel-media",
+ label: "Relocate channel media",
+ drainable: false,
+ replayable: false,
+ queueKeyStrategy: "custom",
+ needsMedia: false,
+ },
// Social-post ingest for a `sourceKind: "social"` channel. Drainable (the
// fetcher stops paging on the drain signal and keeps what it already has) and
// replayable. queueKeyForUrl() routes x.com / bsky.app to
@@ -323,12 +407,17 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
replayable: true,
queueKeyStrategy: "platform",
},
+ // needsMedia: syncPaged does not merely refresh a playlist — it reads the
+ // channel's data dir to decide what is already there and downloads into
+ // `data/<id>/`. Against an unmounted drive its listing is empty, which means
+ // "nothing is downloaded", which means download everything.
sync: {
kind: "sync",
label: "Sync",
drainable: true,
replayable: true,
queueKeyStrategy: "platform",
+ needsMedia: true,
},
// Replayable kinds that never had a JOB_KIND_LABELS entry: label omitted so
// jobKindLabel() keeps falling back to the raw kind (unchanged behavior).
@@ -349,12 +438,14 @@ const JOB_KINDS: Record<string, JobKindMeta> = {
drainable: false,
replayable: true,
queueKeyStrategy: "custom",
+ needsMedia: true,
},
"remove-wrong-format-audio": {
kind: "remove-wrong-format-audio",
drainable: false,
replayable: true,
queueKeyStrategy: "custom",
+ needsMedia: true,
},
};
@@ -371,3 +462,10 @@ export function jobKindLabel(kind: string): string {
export function isDrainableKind(kind: string): boolean {
return JOB_KINDS[kind]?.drainable ?? false;
}
+
+// Whether a kind's work reaches `channels/<slug>/data/`. Absent = false: the
+// media guard is opt-in, so an unlisted (or unknown) kind behaves exactly as it
+// did before the guard existed.
+export function kindNeedsMedia(kind: string): boolean {
+ return JOB_KINDS[kind]?.needsMedia ?? false;
+}
diff --git a/common/jobs/snapshotScheduler.ts b/common/jobs/snapshotScheduler.ts
@@ -47,6 +47,19 @@ const NO_REGEN_KINDS = new Set<string>([
// than doing the backfill. The per-channel backfill job regenerates ONCE, at
// job end.
"backfill-channel",
+ // A RELOCATION MOVES BYTES, IT DOES NOT CHANGE THEM. Every video dir, every
+ // transcript and every mtime is identical afterwards — `rsync -a` preserves
+ // them, which is the same property that makes the LMDB index a no-op — so a
+ // regen would walk the whole channel to write a byte-identical snapshot with
+ // a newer generatedAt, claiming a measurement it did not take. On the 130 GB
+ // channel this exists for that is a 16-way walk of 11,000 video dirs for
+ // nothing, immediately after a job that just moved 130 GB.
+ //
+ // It also removes a race the operator would feel: the regen is a queued job
+ // for this channel, and the Storage panel refuses a move while the channel
+ // has one — so "move out, then move back" would be blocked by a report
+ // nobody needed.
+ "relocate-channel-media",
]);
export function shouldRequestSnapshot(kind: string): boolean {
diff --git a/common/jobs/streamCommand.ts b/common/jobs/streamCommand.ts
@@ -19,6 +19,8 @@ import {
import { writeJobMeta } from "./jobMeta";
import { maybePruneJobLogs } from "./listJobs";
import type { JobSpec } from "./jobSpec";
+import { kindNeedsMedia } from "./jobKinds";
+import { assertChannelMediaReachable } from "../lib/channelMedia";
// Mark a job's channel report dirty so the debounced scheduler regenerates the
// snapshot — called both on each completed sub-operation and on the job's
@@ -261,9 +263,34 @@ export async function runManagedCommand(
return { ok: true, jobId: id, stream, done };
}
+// GUARD 1 OF FOUR (see plans/relocate-channel-media.md). runManagedFunction is
+// the funnel every job-shaped action goes through, and 35 of its 53 call sites
+// already pass a channelSlug — so one check here covers every per-channel media
+// action without touching any of them. A refusal happens BEFORE the job record
+// is made: no queued job, no log, no sidecar, just `{ ok: false }` carrying the
+// reason, which every caller already renders.
+//
+// The kind decides. `needsMedia` is declarative on JobKindMeta and absent means
+// false, so a bookkeeping kind is never refused for a drive it does not read,
+// and the relocate job itself — the thing that FIXES an unreachable channel —
+// must never declare it.
+async function refuseForUnreachableMedia(
+ opts: CommonOpts,
+): Promise<string | null> {
+ if (!opts.channelSlug || !kindNeedsMedia(opts.kind)) return null;
+ try {
+ await assertChannelMediaReachable(opts.paths, opts.channelSlug);
+ return null;
+ } catch (err) {
+ return (err as Error).message;
+ }
+}
+
export async function runManagedFunction(
opts: RunManagedFunctionOpts,
): Promise<StreamActionResult> {
+ const refusal = await refuseForUnreachableMedia(opts);
+ if (refusal) return { ok: false, error: refusal };
const registry = getRegistry();
await ensureJobsDir(opts.paths);
const { id, logPath, record } = makeJob(
diff --git a/common/lib/channelConfig.ts b/common/lib/channelConfig.ts
@@ -99,6 +99,14 @@ export type ChannelConfig = {
// channel's large videos land on a different disk than the rest. Resolved by
// savedVideoDir() in common/lib/savedVideo.ts. Empty/whitespace = use global.
savedVideosDir?: string;
+ // Where this channel's downloaded media ACTUALLY lives, when it has been
+ // relocated to another drive: the absolute path `channels/<slug>/data` is a
+ // symlink to. Blank/absent = in place. Written ONLY by the relocate job on
+ // success (common/controller/relocateChannelMedia.ts) — it is a record of
+ // what is on disk, never a free-text field, because a value that disagrees
+ // with the link is an "inconsistent" channel that every guard refuses. See
+ // common/lib/channelMedia.ts.
+ dataDir?: string;
ytdlpExtraArgs?: string[];
subLangs?: string;
lastSyncedAt?: string;
@@ -258,6 +266,9 @@ export function parseChannelConfig(raw: unknown): ChannelConfig | null {
if (typeof r.savedVideosDir === "string" && r.savedVideosDir.trim() !== "") {
config.savedVideosDir = r.savedVideosDir.trim();
}
+ if (typeof r.dataDir === "string" && r.dataDir.trim() !== "") {
+ config.dataDir = r.dataDir.trim();
+ }
if (
Array.isArray(r.ytdlpExtraArgs) &&
r.ytdlpExtraArgs.every((x) => typeof x === "string")
diff --git a/common/lib/channelMedia.test.ts b/common/lib/channelMedia.test.ts
@@ -0,0 +1,281 @@
+import { test } from "node:test";
+import assert from "node:assert/strict";
+import { mkdir, mkdtemp, rm, symlink, writeFile } from "node:fs/promises";
+import { tmpdir } from "node:os";
+import path from "node:path";
+import {
+ ChannelMediaUnreachableError,
+ assertChannelMediaReachable,
+ clearRelocationMarker,
+ inspectChannelMedia,
+ readRelocationMarker,
+ relocatedDataDir,
+ RELOCATION_MARKER_FILENAME,
+ type ChannelMediaPaths,
+} from "./channelMedia";
+
+// Run with:
+// pnpm --filter yt-dlp-transcript-common exec tsx --test lib/channelMedia.test.ts
+//
+// Everything here happens inside one mkdtemp; nothing reads a real corpus.
+
+async function withTmp(
+ fn: (paths: ChannelMediaPaths, root: string, dir: string) => Promise<void>,
+): Promise<void> {
+ const dir = await mkdtemp(path.join(tmpdir(), "ttb-media-"));
+ const paths: ChannelMediaPaths = { channelsDir: path.join(dir, "channels") };
+ const root = path.join(dir, "platter");
+ await mkdir(paths.channelsDir, { recursive: true });
+ try {
+ await fn(paths, root, dir);
+ } finally {
+ await rm(dir, { recursive: true, force: true });
+ }
+}
+
+async function seedChannel(
+ paths: ChannelMediaPaths,
+ slug: string,
+ config: Record<string, unknown> = {},
+): Promise<string> {
+ const channelDir = path.join(paths.channelsDir, slug);
+ await mkdir(channelDir, { recursive: true });
+ await writeFile(
+ path.join(channelDir, "config.json"),
+ JSON.stringify({ url: "https://example.com/c", ...config }, null, 2),
+ );
+ return channelDir;
+}
+
+test("relocatedDataDir fixes the <root>/<slug>/data suffix", () => {
+ assert.equal(relocatedDataDir("/mnt/p", "alpha"), "/mnt/p/alpha/data");
+ assert.equal(relocatedDataDir(" /mnt/p ", "alpha"), "/mnt/p/alpha/data");
+});
+
+test("a real data dir with no config.dataDir is in-place", async () => {
+ await withTmp(async (paths) => {
+ const channelDir = await seedChannel(paths, "alpha");
+ await mkdir(path.join(channelDir, "data", "v1"), { recursive: true });
+ const loc = await inspectChannelMedia(paths, "alpha");
+ assert.equal(loc.status, "in-place");
+ assert.equal(loc.relocated, false);
+ assert.equal(loc.target, undefined);
+ assert.equal(loc.dataDir, path.join(channelDir, "data"));
+ await assertChannelMediaReachable(paths, "alpha");
+ });
+});
+
+test("a channel that has downloaded nothing is in-place, not an error", async () => {
+ await withTmp(async (paths) => {
+ await seedChannel(paths, "alpha");
+ const loc = await inspectChannelMedia(paths, "alpha");
+ assert.equal(loc.status, "in-place");
+ await assertChannelMediaReachable(paths, "alpha");
+ });
+});
+
+test("a link agreeing with config and pointing at a live dir is ok", async () => {
+ await withTmp(async (paths, root) => {
+ const target = relocatedDataDir(root, "alpha");
+ await mkdir(path.join(target, "v1"), { recursive: true });
+ const channelDir = await seedChannel(paths, "alpha", { dataDir: target });
+ await symlink(target, path.join(channelDir, "data"));
+ const loc = await inspectChannelMedia(paths, "alpha");
+ assert.equal(loc.status, "ok");
+ assert.equal(loc.relocated, true);
+ assert.equal(loc.target, target);
+ await assertChannelMediaReachable(paths, "alpha");
+ });
+});
+
+test("a dangling link (drive not mounted) is unreachable and throws", async () => {
+ await withTmp(async (paths, root) => {
+ const target = relocatedDataDir(root, "alpha");
+ const channelDir = await seedChannel(paths, "alpha", { dataDir: target });
+ // The link is made WITHOUT creating the target: exactly an unmounted drive.
+ await symlink(target, path.join(channelDir, "data"));
+ const loc = await inspectChannelMedia(paths, "alpha");
+ assert.equal(loc.status, "unreachable");
+ assert.equal(loc.relocated, true);
+ assert.match(loc.detail ?? "", /does not exist/);
+ await assert.rejects(
+ () => assertChannelMediaReachable(paths, "alpha"),
+ (err: unknown) => {
+ assert.ok(err instanceof ChannelMediaUnreachableError);
+ assert.equal(err.slug, "alpha");
+ assert.equal(err.status, "unreachable");
+ assert.match(err.message, /not reachable/);
+ return true;
+ },
+ );
+ });
+});
+
+test("an EMPTY mountpoint is still unreachable — the link points deep", async () => {
+ await withTmp(async (paths, root) => {
+ const target = relocatedDataDir(root, "alpha");
+ // The root exists (mountpoint present) but holds nothing.
+ await mkdir(root, { recursive: true });
+ const channelDir = await seedChannel(paths, "alpha", { dataDir: target });
+ await symlink(target, path.join(channelDir, "data"));
+ const loc = await inspectChannelMedia(paths, "alpha");
+ assert.equal(loc.status, "unreachable");
+ });
+});
+
+test("a marker makes the channel in-transition whatever the disk says", async () => {
+ await withTmp(async (paths, root) => {
+ const target = relocatedDataDir(root, "alpha");
+ const channelDir = await seedChannel(paths, "alpha");
+ await mkdir(path.join(channelDir, "data"), { recursive: true });
+ await writeFile(
+ path.join(channelDir, RELOCATION_MARKER_FILENAME),
+ JSON.stringify({
+ target,
+ direction: "out",
+ startedAt: new Date().toISOString(),
+ phase: "copy",
+ }),
+ );
+ const loc = await inspectChannelMedia(paths, "alpha");
+ assert.equal(loc.status, "in-transition");
+ assert.equal(loc.marker?.phase, "copy");
+ assert.equal(loc.marker?.direction, "out");
+ assert.equal(loc.target, target);
+ await assert.rejects(
+ () => assertChannelMediaReachable(paths, "alpha"),
+ ChannelMediaUnreachableError,
+ );
+
+ const marker = await readRelocationMarker(paths, "alpha");
+ assert.equal(marker?.target, target);
+ });
+});
+
+test("a link that disagrees with config is inconsistent, never guessed past", async () => {
+ await withTmp(async (paths, root) => {
+ const real = relocatedDataDir(root, "alpha");
+ const recorded = relocatedDataDir(path.join(root, "other"), "alpha");
+ await mkdir(real, { recursive: true });
+ const channelDir = await seedChannel(paths, "alpha", { dataDir: recorded });
+ await symlink(real, path.join(channelDir, "data"));
+ const loc = await inspectChannelMedia(paths, "alpha");
+ assert.equal(loc.status, "inconsistent");
+ assert.match(loc.detail ?? "", /points at/);
+ await assert.rejects(
+ () => assertChannelMediaReachable(paths, "alpha"),
+ ChannelMediaUnreachableError,
+ );
+ });
+});
+
+test("a link with no config.dataDir is inconsistent", async () => {
+ await withTmp(async (paths, root) => {
+ const target = relocatedDataDir(root, "alpha");
+ await mkdir(target, { recursive: true });
+ const channelDir = await seedChannel(paths, "alpha");
+ await symlink(target, path.join(channelDir, "data"));
+ const loc = await inspectChannelMedia(paths, "alpha");
+ assert.equal(loc.status, "inconsistent");
+ assert.equal(loc.relocated, false);
+ assert.match(loc.detail ?? "", /records no dataDir/);
+ });
+});
+
+test("config.dataDir with a real directory on disk is inconsistent", async () => {
+ await withTmp(async (paths, root) => {
+ const target = relocatedDataDir(root, "alpha");
+ const channelDir = await seedChannel(paths, "alpha", { dataDir: target });
+ await mkdir(path.join(channelDir, "data"), { recursive: true });
+ const loc = await inspectChannelMedia(paths, "alpha");
+ assert.equal(loc.status, "inconsistent");
+ assert.match(loc.detail ?? "", /never moved/);
+ });
+});
+
+test("config.dataDir with no data/ at all is inconsistent (link gone)", async () => {
+ await withTmp(async (paths, root) => {
+ const target = relocatedDataDir(root, "alpha");
+ await seedChannel(paths, "alpha", { dataDir: target });
+ const loc = await inspectChannelMedia(paths, "alpha");
+ assert.equal(loc.status, "inconsistent");
+ assert.match(loc.detail ?? "", /symlink is missing/);
+ });
+});
+
+test("a passed config is used verbatim; config.json is only read when it is absent", async () => {
+ await withTmp(async (paths, root) => {
+ const target = relocatedDataDir(root, "alpha");
+ await mkdir(target, { recursive: true });
+ // config.json on disk says NOTHING about a relocation...
+ const channelDir = await seedChannel(paths, "alpha");
+ await symlink(target, path.join(channelDir, "data"));
+
+ // ...so reading it itself gives "inconsistent"...
+ assert.equal(
+ (await inspectChannelMedia(paths, "alpha")).status,
+ "inconsistent",
+ );
+ // ...while a caller that hands over the config it already holds gets the
+ // answer for THAT config, with no second read.
+ const passed = await inspectChannelMedia(paths, "alpha", {
+ dataDir: target,
+ });
+ assert.equal(passed.status, "ok");
+ assert.equal(passed.target, target);
+
+ // An explicit null means "I have no config" and must not silently fall back
+ // to reading the file.
+ const nulled = await inspectChannelMedia(paths, "alpha", null);
+ assert.equal(nulled.status, "inconsistent");
+ });
+});
+
+test("a blank config.dataDir means in place", async () => {
+ await withTmp(async (paths) => {
+ const channelDir = await seedChannel(paths, "alpha", { dataDir: " " });
+ await mkdir(path.join(channelDir, "data"), { recursive: true });
+ assert.equal((await inspectChannelMedia(paths, "alpha")).status, "in-place");
+ assert.equal(
+ (await inspectChannelMedia(paths, "alpha", { dataDir: " " })).status,
+ "in-place",
+ );
+ });
+});
+
+test("an unreadable or missing config.json is not a relocation", async () => {
+ await withTmp(async (paths) => {
+ const channelDir = path.join(paths.channelsDir, "alpha");
+ await mkdir(path.join(channelDir, "data"), { recursive: true });
+ await writeFile(path.join(channelDir, "config.json"), "{ not json");
+ assert.equal((await inspectChannelMedia(paths, "alpha")).status, "in-place");
+ });
+});
+
+test("clearRelocationMarker removes the marker and touches nothing else", async () => {
+ await withTmp(async (paths) => {
+ const target = path.join(paths.channelsDir, "..", "platter", "alpha", "data");
+ await mkdir(target, { recursive: true });
+ const channelDir = await seedChannel(paths, "alpha", { dataDir: target });
+ await symlink(target, path.join(channelDir, "data"));
+ await writeFile(
+ path.join(channelDir, RELOCATION_MARKER_FILENAME),
+ JSON.stringify({ target, direction: "out", phase: "swap", startedAt: "" }),
+ );
+ // A marker outranks everything: the channel is in transition, which is the
+ // dead end this escape hatch exists for when the run that wrote it is gone.
+ assert.equal(
+ (await inspectChannelMedia(paths, "alpha")).status,
+ "in-transition",
+ );
+
+ await clearRelocationMarker(paths, "alpha");
+
+ assert.equal(await readRelocationMarker(paths, "alpha"), null);
+ // The link, the config and the media are exactly as they were — what is
+ // left is the truth the disk was already telling underneath the marker.
+ assert.equal((await inspectChannelMedia(paths, "alpha")).status, "ok");
+ // Idempotent: clearing a marker that is not there is not an error.
+ await clearRelocationMarker(paths, "alpha");
+ });
+});
diff --git a/common/lib/channelMedia.ts b/common/lib/channelMedia.ts
@@ -0,0 +1,329 @@
+import path from "node:path";
+import { lstat, readFile, readlink, rm, stat } from "node:fs/promises";
+import type { Paths } from "./paths";
+import type { ChannelConfig } from "./channelConfig";
+
+// WHERE A CHANNEL'S MEDIA ACTUALLY IS, and whether it can be reached.
+//
+// A channel's downloaded media lives at `channels/<slug>/data/`. That path is
+// joined inline at ~74 call sites and is the on-disk contract every reader,
+// yt-dlp's cwd-relative output template and the LMDB index depend on, so
+// relocating a channel to another drive does NOT change it: `data/` becomes an
+// absolute SYMLINK to `<root>/<slug>/data` and `config.json` records the target
+// in `dataDir`. Every existing reader follows the link transparently — there is
+// no symlink-aware code anywhere in common/, editor/ or export/, and there does
+// not need to be.
+//
+// What that buys in call-site churn it owes in one new failure mode: an
+// unmounted drive. A dangling link reads as ENOENT, and the three places that
+// enumerate `data/` swallow ENOENT as "this channel has no videos" — which to an
+// unattended runner means *everything is undownloaded* and is an instruction to
+// re-download hundreds of gigabytes onto the volume that was too full to hold
+// them. This module is the one place that can tell those two apart, and the
+// guards that call it are what make the symlink safe.
+//
+// IT LIVES IN lib/ AND MAY NOT IMPORT controller/ (architecture.test.ts), which
+// is where readChannelConfig is. Hence the optional `config` argument: a caller
+// holding a parsed config passes it and pays nothing, and when it is absent this
+// module reads `<channelDir>/config.json` itself and pulls out the one field it
+// needs. That is a four-line JSON read, not a second config parser — nothing
+// here validates or defaults anything else in the file.
+
+// Only the paths field this module needs, so a caller (and a test) does not have
+// to build a whole Paths to ask where a channel's media is.
+export type ChannelMediaPaths = Pick<Paths, "channelsDir">;
+
+// Marker written in the CHANNEL dir (never in data/, which is the thing being
+// moved) for the duration of a relocation. Its presence means "media is in
+// transition" to every guard, and its `phase` is what lets an interrupted job
+// resume rather than restart.
+export const RELOCATION_MARKER_FILENAME = ".relocating.json";
+
+export type RelocationDirection = "out" | "back";
+export type RelocationPhase = "copy" | "swap" | "reclaim";
+
+export type RelocationMarker = {
+ // Absolute path of the relocated data dir: <root>/<slug>/data.
+ target: string;
+ direction: RelocationDirection;
+ startedAt: string;
+ phase: RelocationPhase;
+};
+
+export type ChannelMediaStatus =
+ // No relocation: `data/` is a real directory (or does not exist yet).
+ | "in-place"
+ // Relocated, link and config agree, and the target is a reachable directory.
+ | "ok"
+ // Relocated, but the target is not there — almost always an unmounted drive.
+ | "unreachable"
+ // A relocation is in flight (or was interrupted): the marker is present.
+ | "in-transition"
+ // Disk and config disagree, in either direction. Never guessed past.
+ | "inconsistent";
+
+export type ChannelMediaLocation = {
+ // Always channelDir/data — the path every reader uses, relocated or not.
+ dataDir: string;
+ // Whether config.json records a relocation target.
+ relocated: boolean;
+ // config.dataDir (or, mid-transition with no config yet, the marker's target).
+ target?: string;
+ status: ChannelMediaStatus;
+ // Operator-readable reason, set for every status except "in-place" and "ok".
+ detail?: string;
+ // The in-flight marker, when one is present.
+ marker?: RelocationMarker;
+};
+
+// Thrown by assertChannelMediaReachable. A distinct class so a caller can tell
+// "this channel's drive is not mounted" from any other I/O failure and skip
+// rather than fail the whole lane.
+export class ChannelMediaUnreachableError extends Error {
+ readonly slug: string;
+ readonly status: ChannelMediaStatus;
+ readonly location: ChannelMediaLocation;
+ constructor(slug: string, location: ChannelMediaLocation) {
+ super(
+ `Channel "${slug}": media is not reachable — ${
+ location.detail ?? location.status
+ }`,
+ );
+ this.name = "ChannelMediaUnreachableError";
+ this.slug = slug;
+ this.status = location.status;
+ this.location = location;
+ }
+}
+
+// The relocated layout, fixed so one root can hold many channels and the shape
+// mirrors the saved-video store (<root>/<slug>/<...>). The `<slug>/data` suffix
+// is not configurable: deleteChannel and renameChannel recognise a target by it.
+export function relocatedDataDir(root: string, slug: string): string {
+ return path.join(root.trim(), slug, "data");
+}
+
+export function channelMediaDir(
+ paths: ChannelMediaPaths,
+ slug: string,
+): string {
+ return path.join(paths.channelsDir, slug, "data");
+}
+
+export function relocationMarkerPath(
+ paths: ChannelMediaPaths,
+ slug: string,
+): string {
+ return path.join(paths.channelsDir, slug, RELOCATION_MARKER_FILENAME);
+}
+
+function parseMarker(raw: unknown): RelocationMarker | null {
+ if (!raw || typeof raw !== "object") return null;
+ const r = raw as Record<string, unknown>;
+ if (typeof r.target !== "string" || r.target.trim() === "") return null;
+ const direction = r.direction === "back" ? "back" : "out";
+ const phase =
+ r.phase === "swap" || r.phase === "reclaim" ? r.phase : "copy";
+ return {
+ target: r.target,
+ direction,
+ startedAt: typeof r.startedAt === "string" ? r.startedAt : "",
+ phase,
+ };
+}
+
+export async function readRelocationMarker(
+ paths: ChannelMediaPaths,
+ slug: string,
+): Promise<RelocationMarker | null> {
+ try {
+ const raw = await readFile(relocationMarkerPath(paths, slug), "utf8");
+ return parseMarker(JSON.parse(raw));
+ } catch {
+ return null;
+ }
+}
+
+// THE OPERATOR'S LAST RESORT, and the only writer in this module.
+//
+// A channel carrying a marker is "in-transition" to every guard, which is
+// correct while a move is running and a dead end once one is not: the runners
+// skip the channel, runManagedFunction refuses its media jobs, and the snapshot
+// will not regenerate. The relocate job itself resumes from a marker and clears
+// it on success, so this is not the normal way out — it is for a marker whose
+// run is gone (a killed process, a container replaced mid-copy) and whose state
+// on disk the operator has looked at.
+//
+// It removes the marker and NOTHING else: no link is touched, no config is
+// rewritten, nothing is deleted. Whatever inspect() says afterwards is the truth
+// the disk was already telling, with the transition claim taken off the top.
+export async function clearRelocationMarker(
+ paths: ChannelMediaPaths,
+ slug: string,
+): Promise<void> {
+ await rm(relocationMarkerPath(paths, slug), { force: true });
+}
+
+// The `dataDir` field alone, read straight off config.json. Deliberately NOT
+// parseChannelConfig: this runs in guards on hot paths and must not depend on
+// the controller that owns the rest of the schema.
+async function readConfiguredDataDir(
+ paths: ChannelMediaPaths,
+ slug: string,
+): Promise<string | undefined> {
+ try {
+ const raw = await readFile(
+ path.join(paths.channelsDir, slug, "config.json"),
+ "utf8",
+ );
+ const parsed = JSON.parse(raw) as { dataDir?: unknown };
+ if (typeof parsed.dataDir !== "string") return undefined;
+ const trimmed = parsed.dataDir.trim();
+ return trimmed === "" ? undefined : trimmed;
+ } catch {
+ return undefined;
+ }
+}
+
+// Two stats and (at most) one small JSON read. Render-safe: nothing here walks a
+// directory, so calling it per channel on a listing page costs three syscalls a
+// row.
+export async function inspectChannelMedia(
+ paths: ChannelMediaPaths,
+ slug: string,
+ config?: Pick<ChannelConfig, "dataDir"> | null,
+): Promise<ChannelMediaLocation> {
+ const dataDir = channelMediaDir(paths, slug);
+ const configured =
+ config === undefined
+ ? await readConfiguredDataDir(paths, slug)
+ : config?.dataDir && config.dataDir.trim() !== ""
+ ? config.dataDir.trim()
+ : undefined;
+
+ const marker = await readRelocationMarker(paths, slug);
+ if (marker) {
+ return {
+ dataDir,
+ relocated: Boolean(configured),
+ target: configured ?? marker.target,
+ status: "in-transition",
+ detail:
+ `a media relocation (${marker.direction}) is in progress or was ` +
+ `interrupted at phase "${marker.phase}" — target ${marker.target}`,
+ marker,
+ };
+ }
+
+ let link: Awaited<ReturnType<typeof lstat>> | null = null;
+ try {
+ link = await lstat(dataDir);
+ } catch {
+ // No data/ at all. With no configured target that is just a channel that
+ // has downloaded nothing yet — the overwhelmingly common case, and not an
+ // error. With one, the link this channel is supposed to have is gone.
+ if (!configured) return { dataDir, relocated: false, status: "in-place" };
+ return {
+ dataDir,
+ relocated: true,
+ target: configured,
+ status: "inconsistent",
+ detail:
+ `config.json records dataDir ${configured} but ${dataDir} does not ` +
+ `exist — the symlink is missing`,
+ };
+ }
+
+ if (link.isSymbolicLink()) {
+ let linkTarget = "";
+ try {
+ linkTarget = await readlink(dataDir);
+ } catch {
+ /* readlink of a link we just lstat'd: treat as unreadable below */
+ }
+ if (!configured) {
+ return {
+ dataDir,
+ relocated: false,
+ target: linkTarget || undefined,
+ status: "inconsistent",
+ detail:
+ `${dataDir} is a symlink to ${linkTarget || "(unreadable)"} but ` +
+ `config.json records no dataDir`,
+ };
+ }
+ if (path.resolve(linkTarget) !== path.resolve(configured)) {
+ return {
+ dataDir,
+ relocated: true,
+ target: configured,
+ status: "inconsistent",
+ detail:
+ `${dataDir} points at ${linkTarget || "(unreadable)"} but ` +
+ `config.json records ${configured}`,
+ };
+ }
+ // The link points at a DEEP path (<root>/<slug>/data), so an unmounted root
+ // gives ENOENT here. An empty mountpoint can never be mistaken for the
+ // media, which is the whole reason the suffix is fixed.
+ try {
+ const st = await stat(configured);
+ if (!st.isDirectory()) {
+ return {
+ dataDir,
+ relocated: true,
+ target: configured,
+ status: "unreachable",
+ detail: `${configured} exists but is not a directory`,
+ };
+ }
+ } catch {
+ return {
+ dataDir,
+ relocated: true,
+ target: configured,
+ status: "unreachable",
+ detail: `${configured} does not exist (drive not mounted?)`,
+ };
+ }
+ return { dataDir, relocated: true, target: configured, status: "ok" };
+ }
+
+ if (!link.isDirectory()) {
+ return {
+ dataDir,
+ relocated: Boolean(configured),
+ target: configured,
+ status: "inconsistent",
+ detail: `${dataDir} is neither a directory nor a symlink`,
+ };
+ }
+
+ if (configured) {
+ return {
+ dataDir,
+ relocated: true,
+ target: configured,
+ status: "inconsistent",
+ detail:
+ `config.json records dataDir ${configured} but ${dataDir} is a real ` +
+ `directory — the media was never moved, or was moved back by hand`,
+ };
+ }
+ return { dataDir, relocated: false, status: "in-place" };
+}
+
+// "ok" and "in-place" pass; everything else throws. An in-transition or
+// inconsistent channel is refused for the same reason an unreachable one is:
+// the caller would otherwise read a half-populated or empty dir as the truth.
+export async function assertChannelMediaReachable(
+ paths: ChannelMediaPaths,
+ slug: string,
+ config?: Pick<ChannelConfig, "dataDir"> | null,
+): Promise<ChannelMediaLocation> {
+ const location = await inspectChannelMedia(paths, slug, config);
+ if (location.status === "ok" || location.status === "in-place") {
+ return location;
+ }
+ throw new ChannelMediaUnreachableError(slug, location);
+}
diff --git a/common/lib/diskSpace.test.ts b/common/lib/diskSpace.test.ts
@@ -1,6 +1,17 @@
import { test } from "node:test";
import assert from "node:assert/strict";
-import { evaluateDiskGate } from "./diskSpace";
+import os from "node:os";
+import path from "node:path";
+import { mkdtemp, rm, statfs, writeFile } from "node:fs/promises";
+import {
+ diskGate,
+ evaluateDiskGate,
+ getFreeBytes,
+ isDiskGateLatched,
+ resetDiskGate,
+} from "./diskSpace";
+import type { Paths } from "./paths";
+import type { SiteSettings } from "./settings";
// The gate's whole job is the pair (floor, hysteresis). evaluateDiskGate is the
// pure core precisely so both can be pinned without a filesystem: the flapping
@@ -106,3 +117,142 @@ test("a statfs failure fails open", () => {
assert.equal(g.ok, true);
assert.equal(g.latched, false);
});
+
+// --- THE LATCH IS PER VOLUME ----------------------------------------------
+//
+// There is no longer one disk. A channel whose media has been relocated writes
+// its downloads to another drive through the data/ symlink
+// (common/lib/channelMedia.ts), so one shared boolean would have let a full SSD
+// pause work landing on the platter and a healthy platter reopen the gate the
+// SSD had closed — in both directions, and reading as the manual pause while it
+// did it.
+//
+// These take a real measurement, which is what makes them worth having next to
+// the pure cases above. BOTH DIRS ARE REAL AND EXIST: a path that does not exist
+// is no longer a stand-in for "a volume with room", because getFreeBytes now
+// measures a missing dir's nearest existing ANCESTOR (see below — that fix is
+// what makes the per-channel gate work at all). So "full" is a real path under
+// an absurd floor and "roomy" is a real path under a floor of about a kilobyte.
+// No filesystem is written to.
+
+const FULL = "/"; // a real path, measured, and always under the absurd floor
+const ROOMY = os.tmpdir(); // a real path, measured, and always over a 1 KB floor
+
+const settingsWithFloor = (minFreeDiskGB: number) =>
+ ({ minFreeDiskGB, resumeMarginGB: 1 }) as SiteSettings;
+const somePaths = { transcriptsDir: FULL } as Paths;
+
+// A floor of ~1 KB: enabled (a floor of 0 disables the gate entirely) and
+// cleared by any filesystem with room on it.
+const TINY_FLOOR = 1e-6;
+
+test("a latch on one volume does not hold another", async () => {
+ resetDiskGate();
+ // An absurd floor no real filesystem clears.
+ const full = await diskGate(somePaths, settingsWithFloor(1e9), { dir: FULL });
+ assert.equal(full.ok, false);
+ assert.equal(isDiskGateLatched(FULL), true);
+ assert.equal(isDiskGateLatched(ROOMY), false);
+
+ // The other volume is unaffected — and, crucially, asking about it does not
+ // clear the first one's latch. With one shared boolean it would have.
+ const roomy = await diskGate(somePaths, settingsWithFloor(TINY_FLOOR), {
+ dir: ROOMY,
+ });
+ assert.equal(roomy.ok, true);
+ assert.equal(isDiskGateLatched(FULL), true);
+ assert.equal(isDiskGateLatched(ROOMY), false);
+ resetDiskGate();
+});
+
+// --- A DIR THAT DOES NOT EXIST YET IS MEASURED ON ITS PARENT'S VOLUME -------
+//
+// The six per-channel gate callers pass `channels/<slug>/data`, which does not
+// exist until the channel's first download creates it. Fail-open on ENOENT
+// answered Infinity there, so the gate waved through exactly the download it
+// was installed to stop. It measures the nearest existing ancestor instead.
+
+test("getFreeBytes measures a not-yet-existing dir on its nearest existing ancestor", async () => {
+ const root = os.tmpdir();
+ const missing = path.join(root, "ttb-no-such-dir", "slug", "data");
+ const measured = await getFreeBytes(missing);
+ const onRoot = await statfs(root);
+ assert.equal(Number.isFinite(measured), true, "must not fail open");
+ // Same volume, so the same figure — modulo whatever the machine wrote
+ // between the two calls, which is why this is a band and not an equality.
+ const expected = onRoot.bsize * onRoot.bavail;
+ assert.ok(
+ Math.abs(measured - expected) < expected * 0.05 + 1024 ** 3,
+ `expected ~${expected}, got ${measured}`,
+ );
+});
+
+test("a non-ENOENT failure still fails open", async () => {
+ // A path UNDER A FILE is ENOTDIR, not ENOENT — nonsense rather than
+ // not-there-yet — so the walk does not run and the original contract holds:
+ // a measurement glitch never blocks a download. This is the half of the old
+ // fail-open that survives, and it is deliberately still Infinity.
+ const dir = await mkdtemp(path.join(os.tmpdir(), "ttb-disk-"));
+ try {
+ const file = path.join(dir, "a-file");
+ await writeFile(file, "x");
+ assert.equal(await getFreeBytes(path.join(file, "nope")), Infinity);
+ } finally {
+ await rm(dir, { recursive: true, force: true });
+ }
+});
+
+test("the gate latches for a data dir that does not exist yet", async () => {
+ resetDiskGate();
+ const notYet = path.join(os.tmpdir(), "ttb-unborn-channel", "data");
+ const status = await diskGate(somePaths, settingsWithFloor(1e9), {
+ dir: notYet,
+ });
+ // Measured on tmpdir's volume, which no absurd floor clears. Before the
+ // ancestor walk this was Infinity and `ok: true`.
+ assert.equal(status.ok, false);
+ assert.equal(isDiskGateLatched(notYet), true);
+ resetDiskGate();
+});
+
+test("the hysteresis is still per volume: a latched dir is held to the higher bar", async () => {
+ resetDiskGate();
+ await diskGate(somePaths, settingsWithFloor(1e9), { dir: FULL });
+ assert.equal(isDiskGateLatched(FULL), true);
+ // A floor of 0 disables the gate entirely, which clears every volume: the
+ // operator switching the floor off means nothing is held anywhere.
+ const off = await diskGate(somePaths, settingsWithFloor(0), { dir: FULL });
+ assert.equal(off.ok, true);
+ assert.equal(off.enabled, false);
+ assert.equal(isDiskGateLatched(), false);
+});
+
+test("observe reads a volume's latch without writing it", async () => {
+ resetDiskGate();
+ const observed = await diskGate(somePaths, settingsWithFloor(1e9), {
+ dir: FULL,
+ mode: "observe",
+ });
+ assert.equal(observed.ok, false);
+ // Nothing was recorded: a dashboard poll is not the thing that latches.
+ assert.equal(isDiskGateLatched(FULL), false);
+ resetDiskGate();
+});
+
+test("resetDiskGate drops one volume or all of them", async () => {
+ resetDiskGate();
+ await diskGate(somePaths, settingsWithFloor(1e9), { dir: FULL });
+ assert.equal(isDiskGateLatched(), true);
+ resetDiskGate(ROOMY);
+ assert.equal(isDiskGateLatched(FULL), true);
+ resetDiskGate(FULL);
+ assert.equal(isDiskGateLatched(FULL), false);
+ assert.equal(isDiskGateLatched(), false);
+});
+
+test("with no dir the gate measures paths.transcriptsDir, as it always has", async () => {
+ resetDiskGate();
+ await diskGate(somePaths, settingsWithFloor(1e9));
+ assert.equal(isDiskGateLatched(somePaths.transcriptsDir), true);
+ resetDiskGate();
+});
diff --git a/common/lib/diskSpace.ts b/common/lib/diskSpace.ts
@@ -1,3 +1,4 @@
+import path from "node:path";
import { statfs } from "node:fs/promises";
import type { Paths } from "./paths";
import type { SiteSettings } from "./settings";
@@ -6,16 +7,39 @@ import { formatBytes } from "./format";
const BYTES_PER_GB = 1024 ** 3;
// Free space (bytes) available to an unprivileged process on the filesystem
-// holding `dir`. Fail-open: any error (statfs unsupported, path missing,
-// permissions) reports Infinity so a measurement glitch never blocks a
-// download. statfs reports blocks in `bsize`-sized units; `bavail` excludes
+// holding `dir`. statfs reports blocks in `bsize`-sized units; `bavail` excludes
// blocks reserved for root, which is what a normal write can actually use.
+//
+// A DIRECTORY THAT DOES NOT EXIST YET IS MEASURED ON ITS NEAREST EXISTING
+// ANCESTOR, and that is the difference between this gate working and this gate
+// being decorative. Since the gate became per-volume, six callers pass
+// `channels/<slug>/data` — the volume the bytes are ABOUT to land on — and that
+// directory does not exist until the channel's first download creates it
+// (neither createChannel nor syncPaged makes it, deliberately). A bare statfs
+// there is ENOENT, the old fail-open answered Infinity, and the gate then waved
+// through exactly the first download onto a disk it had been asked to protect.
+// The ancestor is the right answer and not an approximation: a directory about
+// to be created lands on the filesystem its parent is on.
+//
+// Everything else still FAILS OPEN (Infinity): statfs unsupported, permissions,
+// a nonsense path (ENOTDIR). A measurement glitch must never block a download —
+// that is the original contract and it is unchanged. Only ENOENT walks, because
+// only ENOENT means "not there YET"; the walk terminates at the filesystem root,
+// where dirname is a fixed point.
export async function getFreeBytes(dir: string): Promise<number> {
- try {
- const stats = await statfs(dir);
- return stats.bsize * stats.bavail;
- } catch {
- return Number.POSITIVE_INFINITY;
+ let current = path.resolve(dir);
+ for (;;) {
+ try {
+ const stats = await statfs(current);
+ return stats.bsize * stats.bavail;
+ } catch (err) {
+ if ((err as NodeJS.ErrnoException).code !== "ENOENT") {
+ return Number.POSITIVE_INFINITY;
+ }
+ const parent = path.dirname(current);
+ if (parent === current) return Number.POSITIVE_INFINITY;
+ current = parent;
+ }
}
}
@@ -83,11 +107,18 @@ export async function checkDiskSpace(
// unnoticed for a week. Every result here carries a reason and a
// preformatted message so no caller has to invent its own wording.
//
-// The latch is module-level on purpose: there is one disk, so there is one
-// rule, and every caller in the process shares it. It is self-healing — any
-// check that sees enough headroom (or a disabled gate) clears it — so no caller
-// has to remember to reset it, and a crashed or cancelled job cannot leave the
-// pipeline wedged.
+// The latch is module-level on purpose: every caller in the process shares one
+// rule. It is self-healing — any check that sees enough headroom (or a disabled
+// gate) clears it — so no caller has to remember to reset it, and a crashed or
+// cancelled job cannot leave the pipeline wedged.
+//
+// IT IS KEYED BY DIRECTORY, because there is no longer one disk. A channel whose
+// media has been relocated to another drive writes its downloads THERE
+// (common/lib/channelMedia.ts), so a full SSD must not pause work landing on the
+// platter and a full platter must not pause everything else. One shared boolean
+// would have done exactly that, in both directions, and would have read as the
+// manual pause while it did it. Callers that write media for a KNOWN channel
+// pass that channel's data dir; the rest keep the default, paths.transcriptsDir.
//
// DELIBERATELY NOT GATED: derived sidecars (digest.json, diarization.json,
// attribution.json). Those are kilobytes. Stopping them frees nothing and costs
@@ -178,7 +209,10 @@ export function evaluateDiskGate(opts: {
};
}
-let gateLatched = false;
+// dir -> latched. One entry per volume anything has actually asked about, which
+// is the corpus plus however many media roots the operator is using — single
+// digits, and it never grows on its own.
+const gateLatched = new Map<string, boolean>();
// Which of the three kinds of caller is asking. The distinctions are the whole
// reason this is one shared function rather than three private checks:
@@ -201,17 +235,23 @@ let gateLatched = false;
// silently reopen the gate for the unattended runner.
export type DiskGateMode = "enforce" | "observe" | "manual";
-// Measure the transcripts filesystem and apply the gate. Skips the statfs
-// syscall entirely when the gate is disabled, like checkDiskSpaceFor — this
-// runs before every item of every byte-writing batch.
+// Measure a filesystem and apply the gate. `dir` defaults to
+// paths.transcriptsDir — pass the channel's data dir when the caller knows which
+// volume the bytes are about to land on. Skips the statfs syscall entirely when
+// the gate is disabled, like checkDiskSpaceFor — this runs before every item of
+// every byte-writing batch.
export async function diskGate(
paths: Paths,
settings: SiteSettings,
- opts?: { mode?: DiskGateMode },
+ opts?: { mode?: DiskGateMode; dir?: string },
): Promise<DiskGateStatus> {
const mode = opts?.mode ?? "enforce";
+ const dir = opts?.dir ?? paths.transcriptsDir;
if (settings.minFreeDiskGB <= 0) {
- if (mode === "enforce") gateLatched = false;
+ // A disabled gate clears EVERY volume's latch, not just this one: the
+ // operator turning the floor off means nothing is held anywhere, and a
+ // surviving entry would keep isDiskGateLatched() answering yes forever.
+ if (mode === "enforce") gateLatched.clear();
return {
ok: true,
enabled: false,
@@ -222,25 +262,36 @@ export async function diskGate(
message: "",
};
}
- const freeBytes = await getFreeBytes(paths.transcriptsDir);
+ const freeBytes = await getFreeBytes(dir);
const { latched, ...status } = evaluateDiskGate({
freeBytes,
minFreeDiskGB: settings.minFreeDiskGB,
resumeMarginGB: settings.resumeMarginGB,
- latched: mode === "manual" ? false : gateLatched,
+ latched: mode === "manual" ? false : (gateLatched.get(dir) ?? false),
});
- if (mode === "enforce") gateLatched = latched;
+ if (mode === "enforce") {
+ // Delete rather than store false: the map is "what is currently held", so
+ // an unlatched volume leaves no trace and the map stays the size of the
+ // problem.
+ if (latched) gateLatched.set(dir, true);
+ else gateLatched.delete(dir);
+ }
return status;
}
// Whether the gate is currently holding. Read-only; for surfaces that want to
-// say "stopped by disk" without taking another measurement.
-export function isDiskGateLatched(): boolean {
- return gateLatched;
+// say "stopped by disk" without taking another measurement. With no `dir` this
+// answers for ANY volume — which is what a single global "downloads are stopped
+// by disk" indicator wants.
+export function isDiskGateLatched(dir?: string): boolean {
+ if (dir === undefined) return gateLatched.size > 0;
+ return gateLatched.get(dir) ?? false;
}
// Drop the latch. For tests and for an operator action that means "try again
-// now" — nothing in normal operation needs it, since the gate self-heals.
-export function resetDiskGate(): void {
- gateLatched = false;
+// now" — nothing in normal operation needs it, since the gate self-heals. With
+// no `dir`, drops every volume's.
+export function resetDiskGate(dir?: string): void {
+ if (dir === undefined) gateLatched.clear();
+ else gateLatched.delete(dir);
}
diff --git a/common/lib/queueKeys.test.ts b/common/lib/queueKeys.test.ts
@@ -0,0 +1,52 @@
+import test from "node:test";
+import assert from "node:assert/strict";
+
+import {
+ BACKFILL_QUEUE,
+ DIGEST_LOCAL_QUEUE,
+ DIGEST_REMOTE_QUEUE,
+ channelQueueKey,
+ relocationQueueKey,
+ resolveQueueKey,
+} from "./queueKeys";
+
+// THE ONE PROPERTY THAT IS NOT OBVIOUS FROM READING THE CALLER.
+//
+// registry.ts submits every non-empty queueKey at concurrency 1 and caps
+// nothing across keys, so "do these two jobs serialize?" is answered entirely by
+// whether their keys are EQUAL. For a media relocation the answer has to be yes
+// even across channels — N moves selected on /channels write to one destination
+// volume, and each job's space check runs when it starts, so concurrent starts
+// each measure a root the others have not written to yet and jointly overrun it.
+test("two channels' relocations land on one queue, and their bookkeeping does not", () => {
+ // THE CLAIM: two channels, one key. It is stated as a function of the slug on
+ // both sides so the test still reads as the claim if someone gives
+ // relocationQueueKey a slug parameter — at which point it fails, which is the
+ // point. channelQueueKey is the control: that one SHOULD differ per channel.
+ const relocationKeyFor = (_slug: string) => relocationQueueKey();
+ assert.equal(relocationKeyFor("omnimirror"), relocationKeyFor("rekietalaw"));
+ assert.notEqual(channelQueueKey("omnimirror"), channelQueueKey("rekietalaw"));
+ // A relocation does not sit on the channel's bookkeeping queue either — that
+ // key is per-channel, so using it is the bug above by another route.
+ assert.notEqual(relocationQueueKey(), channelQueueKey("omnimirror"));
+ // Non-empty, or the registry would run it immediately and untracked.
+ assert.notEqual(relocationQueueKey(), "");
+});
+
+// The lane keys are distinct from each other and from the relocation queue, so
+// a move never blocks (or is blocked by) a sweep.
+test("the named lane queues stay distinct", () => {
+ const keys = [
+ BACKFILL_QUEUE,
+ DIGEST_LOCAL_QUEUE,
+ DIGEST_REMOTE_QUEUE,
+ relocationQueueKey(),
+ ];
+ assert.equal(new Set(keys).size, keys.length);
+});
+
+test("resolveQueueKey: undefined keeps the default, a string overrides, blank is immediate", () => {
+ assert.equal(resolveQueueKey("channel:a", undefined), "channel:a");
+ assert.equal(resolveQueueKey("channel:a", " relocate "), "relocate");
+ assert.equal(resolveQueueKey("channel:a", ""), "");
+});
diff --git a/common/lib/queueKeys.ts b/common/lib/queueKeys.ts
@@ -35,6 +35,35 @@ export function channelQueueKey(slug: string): string {
return `channel:${slug}`;
}
+// ONE QUEUE FOR EVERY MEDIA RELOCATION IN THE PROCESS, and it takes no slug on
+// purpose.
+//
+// The obvious key here is channelQueueKey(slug) — a move is a channel-local
+// operation and it does have to serialize against that channel's own downloads
+// and transcriptions. But registry.ts submits every non-empty key at
+// concurrency 1 and has no cap ACROSS keys, so a per-channel key serializes a
+// channel only against itself: ticking ten rows on /channels and pressing Move
+// starts ten rsyncs at once, all writing to the SAME destination volume. That
+// is wrong twice over. It is slower — ten interleaved sequential writes to one
+// spinning platter is the worst access pattern the drive has — and it breaks
+// the space check, which each job runs when it STARTS: ten jobs that all start
+// together each measure a root none of the others has written to yet, all nine
+// of the later ones are credited room that is already spoken for, and they
+// jointly overrun it. The failure is recoverable (ENOSPC aborts the copy and the
+// source is untouched until a verify passes) but it costs hours of copying to
+// learn something one queue slot would have known.
+//
+// So a relocation runs behind every other relocation, and the run-time space
+// check then measures a root that no other move is writing to. Serializing
+// against the channel's OWN jobs is not lost: both callers refuse a channel that
+// has running or queued jobs before they enqueue, which is a stronger rule than
+// the queue's (it refuses rather than waits, for a move the operator can see is
+// already in flight).
+export const RELOCATION_QUEUE = "relocate";
+export function relocationQueueKey(): string {
+ return RELOCATION_QUEUE;
+}
+
// Queue a network-bound download/pipeline job lands on: the channel's platform
// queue, falling back to a per-domain queue for unrecognized hosts.
export function downloadQueueKey(config: ChannelConfig): string {
diff --git a/common/lib/settings.ts b/common/lib/settings.ts
@@ -214,6 +214,12 @@ export type SiteSettings = {
// sync scheduler runs the backup on the configured cadence. See
// common/controller/backupSavedVideos.ts.
savedVideoBackup: SavedVideoBackupSettings;
+ // Where a channel's downloaded media goes when it is relocated off the corpus
+ // disk. A DEFAULT ONLY: the relocate controller never reads it and always
+ // takes an explicit root, so this is the value the per-channel Storage panel
+ // prefills and the /channels bulk move falls back to. Blank = no default.
+ // See StorageSettings.
+ storage: StorageSettings;
// How the static export is built: "basic" reuses the single export/ tree and
// serializes builds on one queue (the long-standing behavior); "docker" runs
// each site's build in an isolated container for safe parallelism. The Docker
@@ -820,6 +826,47 @@ export function sanitizeSavedVideoBackup(
};
}
+// Where relocated channel media goes by default.
+//
+// ONE FIELD, and deliberately no more. A channel's media is moved by
+// common/controller/relocateChannelMedia.ts, which ALWAYS takes an explicit
+// root: this setting is never read there. It is read by the two places an
+// operator picks a root — the channel's Storage panel (which prefills its input
+// with it) and the /channels bulk move (which falls back to it when the bulk
+// bar's input is blank) — so "the cold drive" is typed once instead of once per
+// channel. It is not a policy: a relocated channel is not thereby
+// deprioritized, and nothing auto-relocates anything because this is set.
+export type StorageSettings = {
+ // Absolute directory holding relocated channels, one `<slug>/data` under it.
+ // Blank = no default; every move then names its own root.
+ mediaRoot: string;
+};
+
+export function defaultStorage(): StorageSettings {
+ return { mediaRoot: "" };
+}
+
+// Coerce a raw settings.storage value into a clean StorageSettings.
+//
+// EXISTENCE IS NOT CHECKED, on purpose: the whole point of a cold root is that
+// it is a drive that may not be mounted when settings are read, and a sanitizer
+// that dropped the field on an unmounted platter would silently erase the
+// operator's choice on the next save.
+//
+// ABSOLUTENESS *IS* checked, and a relative value is dropped to blank rather
+// than resolved. Resolving it would anchor the default to whatever cwd the
+// editor happened to boot in — a different directory under docker, under a
+// worktree, and under `pnpm dev` — so the same settings.json would name three
+// different drives. The settings form rejects a relative path with a message
+// before it ever gets here; this is the last line, not the only one.
+export function sanitizeStorage(value: unknown): StorageSettings {
+ const d = defaultStorage();
+ if (!value || typeof value !== "object") return d;
+ const r = value as Record<string, unknown>;
+ const mediaRoot = typeof r.mediaRoot === "string" ? r.mediaRoot.trim() : "";
+ return { mediaRoot: path.isAbsolute(mediaRoot) ? mediaRoot : "" };
+}
+
// 4 hours. Measured: videos over this are 8.2% of the corpus by count but hold
// 46% of all transcript tokens, so they are where a sweep's wall-clock actually
// goes and where chunk-seam bugs live.
@@ -1019,6 +1066,7 @@ function defaults(): SiteSettings {
socialLinks: [],
homepageUrl: "",
savedVideoBackup: defaultSavedVideoBackup(),
+ storage: defaultStorage(),
buildPipeline: defaultBuildPipeline(),
// Must be listed here or the allowlist loop in getSettings() drops the key
// entirely and the whole section is never read from disk.
@@ -1404,6 +1452,7 @@ export function getSettings(): SiteSettings {
merged.socialLinks = parseSocialLinks(merged.socialLinks);
merged.homepageUrl = normalizeHomepageUrl(merged.homepageUrl);
merged.savedVideoBackup = sanitizeSavedVideoBackup(merged.savedVideoBackup);
+ merged.storage = sanitizeStorage(merged.storage);
merged.buildPipeline = sanitizeBuildPipeline(merged.buildPipeline);
merged.digest = sanitizeDigest(merged.digest);
merged.diarization = sanitizeDiarization(merged.diarization);
@@ -1614,6 +1663,7 @@ export async function writeSettings(next: SiteSettings): Promise<void> {
socialLinks,
homepageUrl: normalizeHomepageUrl(next.homepageUrl),
savedVideoBackup: sanitizeSavedVideoBackup(next.savedVideoBackup),
+ storage: sanitizeStorage(next.storage),
buildPipeline: sanitizeBuildPipeline(next.buildPipeline),
digest: sanitizeDigest(next.digest),
diarization: sanitizeDiarization(next.diarization),
diff --git a/common/lib/storageSettings.test.ts b/common/lib/storageSettings.test.ts
@@ -0,0 +1,66 @@
+// THE COLD-STORAGE DEFAULT, and the two things its sanitizer refuses to do.
+//
+// Run with: node_modules/.bin/tsx --test common/lib/storageSettings.test.ts
+//
+// Its own file rather than a `settings.test.ts`: getSettings() memoizes paths at
+// module scope (see controller/laneForOperation.test.ts), and sanitizeStorage is
+// a pure function that needs none of that seam. Naming the file after the block
+// keeps it that way.
+
+import { test } from "node:test";
+import assert from "node:assert/strict";
+import { defaultStorage, sanitizeStorage } from "./settings";
+
+test("absent, blank and ill-typed all mean no default root", () => {
+ assert.deepEqual(sanitizeStorage(undefined), { mediaRoot: "" });
+ assert.deepEqual(sanitizeStorage(null), { mediaRoot: "" });
+ assert.deepEqual(sanitizeStorage("/mnt/platter"), { mediaRoot: "" });
+ assert.deepEqual(sanitizeStorage({}), { mediaRoot: "" });
+ assert.deepEqual(sanitizeStorage({ mediaRoot: "" }), { mediaRoot: "" });
+ assert.deepEqual(sanitizeStorage({ mediaRoot: " " }), { mediaRoot: "" });
+ assert.deepEqual(sanitizeStorage({ mediaRoot: 7 }), { mediaRoot: "" });
+ assert.deepEqual(defaultStorage(), { mediaRoot: "" });
+});
+
+test("an absolute root is kept, trimmed", () => {
+ assert.deepEqual(sanitizeStorage({ mediaRoot: "/mnt/platter/media" }), {
+ mediaRoot: "/mnt/platter/media",
+ });
+ assert.deepEqual(sanitizeStorage({ mediaRoot: " /mnt/platter/media \n" }), {
+ mediaRoot: "/mnt/platter/media",
+ });
+});
+
+// DROPPED, NOT RESOLVED. Resolving would anchor the default to whatever cwd the
+// editor booted in — different under docker, under a worktree and under `pnpm
+// dev` — so one settings.json would name three different drives.
+test("a relative root is dropped rather than resolved", () => {
+ assert.deepEqual(sanitizeStorage({ mediaRoot: "platter/media" }), {
+ mediaRoot: "",
+ });
+ assert.deepEqual(sanitizeStorage({ mediaRoot: "../platter" }), {
+ mediaRoot: "",
+ });
+ assert.deepEqual(sanitizeStorage({ mediaRoot: "./platter" }), {
+ mediaRoot: "",
+ });
+});
+
+// EXISTENCE IS NEVER CHECKED: the whole point of a cold root is a drive that may
+// be unmounted when settings are read, and a sanitizer that dropped it then
+// would erase the operator's choice on the next save.
+test("a root that does not exist is kept", () => {
+ assert.deepEqual(
+ sanitizeStorage({ mediaRoot: "/definitely/not/mounted/anywhere" }),
+ { mediaRoot: "/definitely/not/mounted/anywhere" },
+ );
+});
+
+// Unknown keys are not carried through: the block is a closed shape, so a
+// settings.json written by a newer build cannot smuggle a field back onto disk.
+test("only mediaRoot survives", () => {
+ assert.deepEqual(
+ sanitizeStorage({ mediaRoot: "/mnt/a", tier: "cold", extra: 1 }),
+ { mediaRoot: "/mnt/a" },
+ );
+});
diff --git a/common/ytdlp/runYtdlp.ts b/common/ytdlp/runYtdlp.ts
@@ -873,7 +873,12 @@ async function runManagedDownloads(
if (opts.drainSignal?.aborted) return;
if (firstFailure && abortOnError) return;
if (lowDiskStopped) return;
- const gate = await diskGate(opts.paths, settings);
+ // THIS channel's volume, not the corpus's: a relocated channel's
+ // downloads land on the platter through the data/ symlink, so the SSD
+ // being full is not a reason to stop them (and vice versa).
+ const gate = await diskGate(opts.paths, settings, {
+ dir: path.join(channelRoot(opts), "data"),
+ });
if (!gate.ok) {
lowDiskStopped = true;
opts.onLog(
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -1,6 +1,7 @@
# Changelog
## [Unreleased]
+- **A channel's media can live on another drive.** A channel page has a **Storage** panel: where its media actually is, how much audio is on disk, how much room is free on the volume holding it, and **Move media to…** — give it a directory on another disk, press *Preview* to see the bytes and the free space there, and the move copies, **verifies**, and only then swaps `data/` for a link to the new location and records it. **Move back in place** reverses it. The source is never touched until the copy has verified, so a cancelled or crashed move leaves everything where it was and the partial copy resumable; re-running finishes it. Nothing else changes: every page, every job, yt-dlp and the search index read the channel exactly as before, because the path they use is unchanged. **The point is what happens when the drive is not mounted.** `data/` reads as empty then, and an empty `data/` means "nothing has been downloaded" to the download runner — an instruction to re-fetch the entire channel onto the disk that was too full to hold it. So an unreachable channel is **refused rather than guessed at**: its media jobs will not start, the four lane runners skip it (and keep running every other channel — this is not a lane stop), its report will not regenerate over an empty directory, and a red **Media unreachable** badge names the path on `/channels`, on the dashboard and on the channel itself. A relocated-and-reachable channel gets a neutral badge saying where; a channel in place gets none. The low-disk floor now measures **the volume the bytes are actually going to** rather than always the corpus disk, and holds each volume separately — a full SSD no longer pauses downloads landing on the platter. The **Media location** line on a channel's Configure form is read-only on purpose: it is a record of what is on disk, written only by a move that succeeded. The cold drive is typed **once**: **Settings → Default media root** seeds the root box in every channel's Storage panel, and `/channels` rows can now be ticked — select several and **Move media to…** queues one job per channel on that channel's own queue, so they serialize instead of fanning out, each one running its own space check at run time rather than at enqueue time (a root that fills partway through refuses the remainder cleanly, and a channel already on that root is skipped rather than failed). The default is a default and nothing more: it is never read by the move itself, which always takes an explicit root, and a relocated channel is not thereby deprioritized. **Nothing moves on its own, and nothing on disk changes until you move a channel.**
- **Every pipeline is dispatched by one thing now: its lane’s runner. The two corpus sweeps and the arbiter are gone.** Digest and Speaker work were driven by a *sweep* — a corpus walk armed by its own switch, with its own scope, its own order and its own console — while Download and Transcription were driven by the auto-queue runner, with rules, a claim ladder, a next-up and a pick log. Two mechanisms, two vocabularies, two sets of bugs. There is one: **each of the four lanes has a runner, a rule list, and Start / Drain / Stop beside its pause**, on the operation’s own page. Arming a corpus pass is switching the lane on; scoping it to particular channels or operations is a *rule*, written the same way auto-transcribe’s have been written since it shipped. The dashboard and the widget keep a one-click switch per lane — **Run every channel** / **Stop the lane** where they said *Sweep every channel* / *Stop sweeping* — and the scope lives on the lane’s page, where you can see what it would do next. **Your armed scope is carried over, and no lane is switched on that was not.** The ten settings fields the sweeps used (`digest.sweepEnabled`, `sweepChannels`, `recencyOrder`, `recencyReach`; `backfill.sweepEnabled`, `sweepKinds`, `sweepChannels`, `order`, `reach`, `weight`) are read once and written into the lane’s rules the first time the editor starts: a sweep armed on three channels becomes three rules, an unscoped one becomes a single *every channel* rule, and a disarmed sweep becomes a switched-off lane. What is retired rather than migrated: **Reach**, because a rule already orders every video it claims across every channel — which rule goes first is the rule list’s job; the digest **order**, whose real meaning was always *newest day first, shortest video within a day* and which the lane spells as **Shortest first** (pick *Newest first* there if you want the date order alone); and the backfill lane’s **Resource share**, which was one number answering two different questions. A lane now stands aside for transcription when it would actually compete for the graphics card, and keeps its slots when it would not — so speaker-naming over an LLM endpoint no longer parks itself behind a transcription it was not competing with. **The arbiter, which never ran a single unit in production, is deleted**; the runner is what dispatches an operation-named rule. **Nothing on disk changes**, and the retired keys are left in `settings.json` — harmless, ignored, and yours to delete.
- **The transcode operation is gone — it never fired.** A channel page had a *Transcode* stage, `/operations/transcode` had a "no console here" panel, `/cleanup` offered "Clear failed transcodings", and the video list drew a third status dot — all for a re-encode step built against two failures that never happened in production: in 68 channels, no snapshot has ever listed a video as missing its target format, no `failed-transcodings` file has ever held an id, and only four channels even met the stage's gate. Transcription never needed it — a video whose audio is in another format transcribes from that file. What stayed is everything that was never the operation's: the download path still re-encodes what it extracts itself, the video page still offers **Transcode audio.\<ext\> → \<fmt\>** per file, and both audio-format sweeps on the Cleanup stage and `/cleanup` are unchanged (gated on the channel having an `audioFormat`, which is what they compare against). The snapshot bucket behind the sweep is `wrongFormatAudio` now — its operator-facing name — and old reports keep their stray key until their next refresh. A `?stage=transcode` bookmark opens the channel overview. **Nothing on disk changes.** Also: the Pool's running-jobs list names the eight kinds its buttons enqueue, and the site's Search aliases tab no longer carries a "no site selected" branch that could not run.
- **A site has tabs, and the family has one page.** Charts, Search aliases, Deploy, Build and Homepage were five sidebar entries beside *Sites*, three of them reading the site from a `?site=` parameter the sidebar picker had to seed, one of them (Build) about no site at all, and one (Homepage) about the family's own hub. A site is one thing now: **`/sites/<id>` is Settings · Charts · Search aliases · Publish**, the site named in the path, the picker following it (and Dashboard and Channels following the picker). **`/sites` is the family page**: the list, then *Release notes*, *Build all sites* with the Basic/Docker mode, the *Hub*, and the *Pool* — the corpus-wide index, stats, sidecar and archive jobs — folded under a disclosure. Search aliases keep both sections on the site's tab: the global dictionary and the site's overrides. Every button, label and log is unchanged; "Select a specific site from the sidebar" is gone because a site's page always has one. The five routes redirect — a `?site=<id>` bookmark lands on that site's tab (the query rides along), `?site=__all__` and the bare routes on `/sites`; a bookmark to a deleted site 404s there exactly as `/sites/<id>` does. The Sites group is one entry; the nav is **eleven**, the IA doc's end state. **Nothing on disk changes.**
diff --git a/editor/app/channels/[slug]/components/stages/StorageStage.tsx b/editor/app/channels/[slug]/components/stages/StorageStage.tsx
@@ -0,0 +1,413 @@
+"use client";
+
+import { useState } from "react";
+import { StreamActionLog } from "yt-dlp-transcript-common/components/StreamActionLog";
+import { formatBytes } from "yt-dlp-transcript-common/lib/format";
+import type { ChannelMediaLocation } from "yt-dlp-transcript-common/lib/channelMedia";
+import type { RelocationPreview } from "yt-dlp-transcript-common/controller/relocateChannelMedia";
+import { MediaLocationBadge } from "../../../../components/MediaLocationBadge";
+import { cancelJobAction } from "../../../../jobs/actions";
+import {
+ clearRelocationMarkerAction,
+ moveChannelMediaBackAction,
+ previewRelocationAction,
+ relocateChannelMediaAction,
+} from "../../storageActions";
+
+// WHERE THIS CHANNEL'S MEDIA LIVES, and the two buttons that change it.
+//
+// The move itself is common/controller/relocateChannelMedia.ts; what this panel
+// owns is the operator's decision. Three numbers are enough to make it: how much
+// there is to move, how much room is free where it is now, and — once a root is
+// named — how much room is free there. The third is what the preview is for, and
+// it is the reason the Move button is gated behind one: "is there space" is not
+// a question this panel should let anyone skip.
+//
+// ⚠️ NEITHER RUN PANEL IS EVER UNMOUNTED BY ITS OWN RESULT.
+// StreamActionLog holds its streamed log in React state and calls
+// router.refresh() the instant a run ends (plans/FACTS.md, "a run log lives in
+// the panel's React state"). That refresh re-renders this panel from the server
+// with `location.relocated` FLIPPED — so a naive `{!relocated && <MoveOut/>}`
+// would delete the log of the move that just succeeded, at the exact moment the
+// operator wants to read it. Both halves therefore follow the documented shape:
+// the parent renders them unconditionally and passes the CONDITION down; each
+// holds a `ranHere` flag set inside its own trigger; each returns null only
+// while `!condition && !ranHere`; and each puts its log LAST, keyed, in a fixed
+// slot, with the now-cleared condition fed to StreamActionLog's `disabled` so
+// the panel that stays for its log is not a second Run button.
+
+type Props = {
+ slug: string;
+ location: ChannelMediaLocation;
+ // Audio bytes this channel holds, from the loaded snapshot — NOT a walk. Null
+ // when the snapshot predates the field (or there is no snapshot), and rendered
+ // as "—" rather than "0": a zero here would claim a measurement nobody took.
+ mediaBytes: number | null;
+ // Free space on the volume the media is on RIGHT NOW — the platter for a
+ // relocated channel, the corpus disk otherwise.
+ freeBytes: number;
+ volumeDir: string;
+ // Why both buttons are off, or null when they are live. Running/queued jobs
+ // for this channel, or a relocation marker left by an interrupted move.
+ blockedReason: string | null;
+ // Whether to offer the marker escape hatch. Computed on the server as "a
+ // marker is present AND no job is running" — a separate answer from
+ // blockedReason, which the marker itself sets: the two cannot be derived from
+ // each other, and reading the hatch off `blockedReason === null` would hide
+ // it in exactly the state it exists for.
+ canClearMarker: boolean;
+ // settings.storage.mediaRoot — the cold root the operator configured once, or
+ // "" when there is none. READ ONLY: this panel prefills its destination box
+ // with it and never writes it back. The settings page is the one writer.
+ defaultRoot: string;
+};
+
+export function StorageStage({
+ slug,
+ location,
+ mediaBytes,
+ freeBytes,
+ volumeDir,
+ blockedReason,
+ canClearMarker,
+ defaultRoot,
+}: Props) {
+ return (
+ <div className="flex flex-col gap-6">
+ <section className="flex flex-col gap-2">
+ <div className="flex items-center gap-3">
+ <h3 className="text-base font-semibold">Location</h3>
+ <MediaLocationBadge media={location} />
+ </div>
+ <dl className="grid grid-cols-[max-content_1fr] gap-x-4 gap-y-1 text-sm">
+ <dt className="text-muted-foreground">Media path</dt>
+ <dd className="font-mono text-xs break-all" aria-label="media path">
+ {location.relocated && location.target
+ ? location.target
+ : location.dataDir}
+ </dd>
+ <dt className="text-muted-foreground">Read as</dt>
+ <dd className="font-mono text-xs break-all">
+ {location.dataDir}
+ {location.relocated ? " (symlink)" : ""}
+ </dd>
+ <dt className="text-muted-foreground">Audio on disk</dt>
+ <dd aria-label="media bytes">
+ {mediaBytes === null ? "—" : formatBytes(mediaBytes)}
+ </dd>
+ <dt className="text-muted-foreground">Free on that volume</dt>
+ <dd aria-label="free on media volume">
+ {formatBytes(freeBytes)}{" "}
+ <span className="text-xs text-muted-foreground font-mono">
+ ({volumeDir})
+ </span>
+ </dd>
+ </dl>
+ {location.detail && (
+ <p className="text-sm text-muted-foreground">{location.detail}</p>
+ )}
+ {mediaBytes === null && (
+ <p className="text-xs text-muted-foreground">
+ No audio total in this channel’s report yet — refresh the
+ report for a figure. The preview below measures the real tree
+ regardless, and it is the number the move acts on.
+ </p>
+ )}
+ </section>
+
+ {blockedReason && (
+ <p
+ role="status"
+ className="text-sm rounded border border-border bg-muted px-3 py-2"
+ >
+ {blockedReason}
+ </p>
+ )}
+
+ {canClearMarker && <StaleMarker key="stale-marker" slug={slug} />}
+
+ <MoveOut
+ key="move-out"
+ slug={slug}
+ canMoveOut={!location.relocated}
+ blockedReason={blockedReason}
+ defaultRoot={defaultRoot}
+ />
+ <MoveBack
+ key="move-back"
+ slug={slug}
+ relocated={location.relocated}
+ target={location.target}
+ blockedReason={blockedReason}
+ // MOVE BACK IS THE ONE DIRECTION THAT ENDS BY DELETING THE TARGET, so
+ // it is not offered for a location nobody can vouch for. `relocated` is
+ // true for `inconsistent` and `unreachable` as well as `ok` — config
+ // records a target in all three — and an inconsistent channel (config
+ // set, `data/` a real directory, which is what `rsync --copy-links`
+ // produces) would have its relocated copy reclaimed while the local one
+ // is swept too. The controller refuses it; this is the same refusal with
+ // a reason, one click earlier.
+ unvouched={
+ location.status === "inconsistent" ||
+ location.status === "unreachable"
+ ? `This channel's media location is ${location.status}: ${
+ location.detail ?? "disk and config do not agree"
+ } Moving back would delete the relocated copy, so it is refused until the location reads "relocated · reachable".`
+ : null
+ }
+ />
+ </div>
+ );
+}
+
+function MoveOut({
+ slug,
+ canMoveOut,
+ blockedReason,
+ defaultRoot,
+}: {
+ slug: string;
+ canMoveOut: boolean;
+ blockedReason: string | null;
+ defaultRoot: string;
+}) {
+ // Seeded from settings.storage.mediaRoot, then owned by the operator. It is
+ // an initial value and NOT a controlled default: a router.refresh() (which
+ // every finished run triggers) must not throw away a root being typed. It
+ // also does not enable the move — the preview gate is unchanged, so a
+ // prefilled root still has to be previewed before Move lights up.
+ const [root, setRoot] = useState(defaultRoot);
+ const [preview, setPreview] = useState<RelocationPreview | null>(null);
+ // The root the preview above describes. Edit the input and the confirmation
+ // goes stale — the numbers were measured against a different volume, and
+ // letting them authorise a move to this one is exactly the mistake the gate
+ // exists to prevent.
+ const [previewedRoot, setPreviewedRoot] = useState("");
+ const [previewError, setPreviewError] = useState<string | null>(null);
+ const [previewing, setPreviewing] = useState(false);
+ const [ranHere, setRanHere] = useState(false);
+
+ if (!canMoveOut && !ranHere) return null;
+
+ const trimmed = root.trim();
+ const confirmed = trimmed !== "" && trimmed === previewedRoot.trim();
+ const disabled = !canMoveOut || !confirmed || blockedReason !== null;
+
+ async function runPreview() {
+ setPreviewing(true);
+ setPreviewError(null);
+ const result = await previewRelocationAction(slug, root);
+ setPreviewing(false);
+ if (result.ok) {
+ setPreview(result.preview);
+ setPreviewedRoot(root);
+ } else {
+ setPreview(null);
+ setPreviewedRoot("");
+ setPreviewError(result.error);
+ }
+ }
+
+ return (
+ <section className="flex flex-col gap-2">
+ <div>
+ <h3 className="text-base font-semibold">Move media to…</h3>
+ <p className="text-sm text-muted-foreground">
+ Copies <code>data/</code> to <code><root>/{slug}/data</code>,
+ verifies it, and leaves a symlink behind so every reader, yt-dlp and
+ the index keep working unchanged. The source is not touched until the
+ copy verifies.
+ </p>
+ </div>
+ <label className="flex flex-col gap-1 text-sm">
+ <span className="font-medium">Destination root</span>
+ <input
+ type="text"
+ name="mediaRoot"
+ aria-label="destination root"
+ value={root}
+ onChange={(e) => setRoot(e.target.value)}
+ placeholder="/mnt/platter/archilyzer-media"
+ disabled={!canMoveOut}
+ className="rounded border border-border bg-card px-2 py-1 text-sm font-mono"
+ />
+ <span className="text-xs text-muted-foreground">
+ An absolute directory that already exists. One root holds many
+ channels; each gets its own <code><slug>/data</code> under it.
+ </span>
+ </label>
+ <div>
+ <button
+ type="button"
+ onClick={runPreview}
+ disabled={!canMoveOut || trimmed === "" || previewing}
+ className="px-3 py-1.5 rounded-md border border-border text-sm font-medium disabled:opacity-50"
+ >
+ {previewing ? "Checking…" : "Preview"}
+ </button>
+ </div>
+ {previewError && (
+ <p role="alert" className="text-sm text-destructive">
+ {previewError}
+ </p>
+ )}
+ {preview && (
+ <dl
+ aria-label="relocation preview"
+ className="grid grid-cols-[max-content_1fr] gap-x-4 gap-y-1 text-sm rounded border border-border bg-muted/40 px-3 py-2"
+ >
+ <dt className="text-muted-foreground">To move</dt>
+ <dd aria-label="bytes to move">
+ {formatBytes(preview.bytesToMove)} in{" "}
+ {preview.files.toLocaleString()} file(s)
+ </dd>
+ <dt className="text-muted-foreground">Target</dt>
+ <dd className="font-mono text-xs break-all">{preview.target}</dd>
+ <dt className="text-muted-foreground">Free there</dt>
+ <dd aria-label="free on destination">
+ {formatBytes(preview.freeOnRoot)}
+ </dd>
+ <dt className="text-muted-foreground">Free here</dt>
+ <dd>{formatBytes(preview.freeOnSource)}</dd>
+ </dl>
+ )}
+ {preview && preview.freeOnRoot < preview.bytesToMove && (
+ <p role="alert" className="text-sm text-destructive">
+ The destination has less free space than the media needs. The job will
+ refuse before copying anything.
+ </p>
+ )}
+ {preview?.sameDevice && (
+ <p className="text-sm text-muted-foreground">
+ That root is on the same volume the corpus is already on, so the move
+ frees nothing.
+ </p>
+ )}
+ {preview?.existingPartial && (
+ <p className="text-sm text-muted-foreground">
+ A partial copy from an earlier attempt is already at the target; this
+ run resumes it rather than starting over.
+ </p>
+ )}
+ {!confirmed && trimmed !== "" && (
+ <p className="text-xs text-muted-foreground">
+ Preview this root to enable the move.
+ </p>
+ )}
+ <StreamActionLog
+ key="move-out-log"
+ trigger={() => {
+ setRanHere(true);
+ return relocateChannelMediaAction(slug, root);
+ }}
+ cancelAction={cancelJobAction}
+ buttonLabel="Move media"
+ runningLabel="Moving media…"
+ label="Move media"
+ disabled={disabled}
+ />
+ </section>
+ );
+}
+
+function MoveBack({
+ slug,
+ relocated,
+ target,
+ blockedReason,
+ unvouched,
+}: {
+ slug: string;
+ relocated: boolean;
+ target: string | undefined;
+ blockedReason: string | null;
+ // Why moving back is refused for THIS location, or null. Separate from
+ // blockedReason, which is about jobs and markers and disables both halves.
+ unvouched: string | null;
+}) {
+ const [ranHere, setRanHere] = useState(false);
+ if (!relocated && !ranHere) return null;
+ return (
+ <section className="flex flex-col gap-2">
+ <div>
+ <h3 className="text-base font-semibold">Move back in place</h3>
+ <p className="text-sm text-muted-foreground">
+ {relocated
+ ? `Copies ${target ?? "the target"} back into the channel dir, verifies it, replaces the symlink with a real directory and clears the recorded location.`
+ : "This channel's media is in place."}
+ </p>
+ </div>
+ {unvouched && (
+ <p
+ role="status"
+ aria-label="move back refused"
+ className="text-sm rounded border border-destructive/50 bg-destructive/5 px-3 py-2"
+ >
+ {unvouched}
+ </p>
+ )}
+ <StreamActionLog
+ key="move-back-log"
+ trigger={() => {
+ setRanHere(true);
+ return moveChannelMediaBackAction(slug);
+ }}
+ cancelAction={cancelJobAction}
+ buttonLabel="Move back in place"
+ runningLabel="Moving back…"
+ label="Move back in place"
+ disabled={!relocated || blockedReason !== null || unvouched !== null}
+ />
+ </section>
+ );
+}
+
+// The way out of a marker whose run is gone.
+//
+// It is offered ONLY when the channel is in-transition AND no job is running —
+// with a live job the marker is not stale, it belongs to that run. Nothing else
+// on this panel is available in that state (both moves are disabled while a
+// marker stands), so without this a killed copy leaves the channel skipped by
+// every lane, refused by every media job and unable to regenerate its snapshot,
+// with the only fix being to delete a dotfile over SSH.
+//
+// It clears the marker and nothing else, which is why the copy says to look
+// first: whatever is on disk stays on disk, and the status underneath may well
+// be `inconsistent`. That is the honest answer.
+function StaleMarker({ slug }: { slug: string }) {
+ const [busy, setBusy] = useState(false);
+ const [error, setError] = useState<string | null>(null);
+ return (
+ <section className="flex flex-col gap-2 rounded border border-destructive/50 bg-destructive/5 px-3 py-2">
+ <h3 className="text-base font-semibold">Clear the relocation marker</h3>
+ <p className="text-sm text-muted-foreground">
+ A move is in flight, or one was interrupted. While the marker stands this
+ channel is skipped by every lane and its media jobs are refused. If no
+ move is actually running, clear the marker — it removes the marker file
+ and nothing else: no files are moved, copied or deleted. Check what is on
+ the drive first.
+ </p>
+ {error && (
+ <p role="alert" className="text-sm text-destructive">
+ {error}
+ </p>
+ )}
+ <div>
+ <button
+ type="button"
+ disabled={busy}
+ onClick={async () => {
+ setBusy(true);
+ setError(null);
+ const result = await clearRelocationMarkerAction(slug);
+ setBusy(false);
+ if (!result.ok) setError(result.error);
+ }}
+ className="px-3 py-1.5 rounded-md border border-destructive text-destructive text-sm font-medium disabled:opacity-50"
+ >
+ {busy ? "Clearing…" : "Clear marker"}
+ </button>
+ </div>
+ </section>
+ );
+}
diff --git a/editor/app/channels/[slug]/lib/fixIncompleteTranscript.ts b/editor/app/channels/[slug]/lib/fixIncompleteTranscript.ts
@@ -75,7 +75,9 @@ export async function fixIncompleteTranscriptOne(opts: {
// full disk once that audio is gone would leave the video with neither the
// stub nor a replacement. Latching, because the channel-level batch calls this
// in a loop unattended.
- const gate = await diskGate(paths, getSettings());
+ const gate = await diskGate(paths, getSettings(), {
+ dir: path.join(paths.channelsDir, slug, "data"),
+ });
if (!gate.ok) {
throw new Error(
`Cannot re-download audio for ${videoId}: ${gate.message}. ` +
diff --git a/editor/app/channels/[slug]/lib/stageStatus.ts b/editor/app/channels/[slug]/lib/stageStatus.ts
@@ -10,6 +10,10 @@ import {
reachableOperationWork,
type OperationGroup,
} from "yt-dlp-transcript-common/lib/operations";
+// TYPE ONLY. channelMedia.ts imports node:fs, and this module is imported by
+// three client components (AttentionStrip, NextAction, OverviewPanel) for its
+// types. A type import is erased, a value import would not be.
+import type { ChannelMediaLocation } from "yt-dlp-transcript-common/lib/channelMedia";
export type SnapshotBuckets = ChannelSnapshot["buckets"];
@@ -59,6 +63,11 @@ export type StageId =
| "speakers"
| "cleanup"
| "diagnostics"
+ // WHERE THE MEDIA PHYSICALLY IS (plans/relocate-channel-media.md). A channel
+ // CHORE like cleanup and diagnostics, not an operation — nothing registers it
+ // and no group owns it — so it is hand-listed at the call site alongside them
+ // and is deliberately absent from GROUP_STAGES below.
+ | "storage"
| "danger";
// THE STAGES EACH OPERATION GROUP OWNS, in the group's own order.
@@ -168,6 +177,12 @@ export type ComputeStageStatusesInput = {
// generic name, because a caller that only needs tone and counts should not
// have to resolve the registry.
backfillKindIds?: ReadonlyArray<string>;
+ // Where this channel's media actually is, from inspectChannelMedia. Optional
+ // because the two stats it costs belong to the caller that already has the
+ // config in hand, and a caller that only wants tone and counts should not
+ // have to do I/O to get them — an omitted location reads as "in place", which
+ // is what every channel was before relocation existed.
+ media?: ChannelMediaLocation | null;
};
export function computeStageStatuses(
@@ -180,6 +195,7 @@ export function computeStageStatuses(
runningJobs,
backfillEnabled = true,
backfillKindIds,
+ media,
} = input;
const buckets = normalizeBuckets(snapshot.buckets);
@@ -557,6 +573,35 @@ export function computeStageStatuses(
}),
};
+ // WHERE THE MEDIA IS. The only stage whose tone comes from a filesystem fact
+ // rather than from a count: an unreachable channel is a channel whose numbers
+ // everywhere else on this page are about to be wrong (an unmounted drive reads
+ // as "nothing downloaded"), so this card is red the moment inspect() says so
+ // and neutral the rest of the time. "in-place" is not an achievement, so it is
+ // never "ok" — the fallback tone for a healthy relocation is neutral too.
+ const mediaStatus = media?.status ?? "in-place";
+ const storage: StageStatus = {
+ id: "storage",
+ title: "Storage",
+ pending: 0,
+ failed: 0,
+ running: mediaStatus === "in-transition",
+ defaultOpen: true,
+ summary:
+ mediaStatus === "in-place"
+ ? "Media is in the channel directory."
+ : mediaStatus === "ok"
+ ? `Media relocated to ${media?.target ?? "another drive"}.`
+ : (media?.detail ?? mediaStatus),
+ tone:
+ mediaStatus === "in-transition"
+ ? "running"
+ : mediaStatus === "unreachable" ||
+ mediaStatus === "inconsistent"
+ ? "danger"
+ : "neutral",
+ };
+
const danger: StageStatus = {
id: "danger",
title: "Danger zone",
@@ -577,6 +622,7 @@ export function computeStageStatuses(
speakers,
cleanup,
diagnostics,
+ storage,
danger,
};
}
diff --git a/editor/app/channels/[slug]/page.tsx b/editor/app/channels/[slug]/page.tsx
@@ -37,6 +37,8 @@ import {
type ShardOp,
} from "yt-dlp-transcript-common/controller/shard";
import { getPaths } from "yt-dlp-transcript-common/lib/paths";
+import { inspectChannelMedia } from "yt-dlp-transcript-common/lib/channelMedia";
+import { getFreeBytes } from "yt-dlp-transcript-common/lib/diskSpace";
import {
platformQueueKey,
queueKeyForUrl,
@@ -66,6 +68,7 @@ import { PlaylistStage } from "./components/stages/PlaylistStage";
import { TranscribeStage } from "./components/stages/TranscribeStage";
import { DigestStage } from "./components/stages/DigestStage";
import { SpeakersStage } from "./components/stages/SpeakersStage";
+import { StorageStage } from "./components/stages/StorageStage";
import {
getOperation,
backfillLaneOperations,
@@ -202,6 +205,14 @@ export default async function ChannelDetailPage({
path.join(paths.channelsDir, slug, "playlist"),
);
+ // WHERE THE MEDIA IS. Two stats and one small JSON read, with the config
+ // already in hand so nothing re-reads it — explicitly not the corpus walk
+ // noCorpusWalkInRenderPaths.test.ts bans. It is read on every render of this
+ // page rather than cached because an unmounted drive is exactly the kind of
+ // fact that must never be served stale: the whole point of the badge is that
+ // it is true NOW.
+ const media = await inspectChannelMedia(paths, slug, config);
+
const buckets = normalizeBuckets(snapshot.buckets);
const undownloadedIds = snapshot.undownloadedIds ?? [];
const excludedFromDownload = normalizeExcludedFromDownload(
@@ -233,6 +244,7 @@ export default async function ChannelDetailPage({
runningJobs,
backfillEnabled: laneOperations.length > 0,
backfillKindIds: laneOperations.map((k) => k.id),
+ media,
});
// THE MIDDLE COMES FROM THE REGISTRY, in OPERATION_GROUP_ORDER. Same array as
@@ -253,6 +265,7 @@ export default async function ChannelDetailPage({
...OPERATION_GROUP_ORDER.flatMap((g) => GROUP_STAGES[g]),
"cleanup",
"diagnostics",
+ "storage",
"danger",
];
@@ -493,6 +506,54 @@ export default async function ChannelDetailPage({
downloadDefaultQueueKey={platformDefaultQueueKey}
/>
);
+ case "storage": {
+ // The free-space figure names the volume the media is on RIGHT NOW —
+ // the platter for a relocated channel, the corpus disk otherwise — so
+ // it answers "can this channel keep downloading", not "how full is
+ // /home". getFreeBytes returns Infinity for a path it cannot statfs,
+ // which formatBytes renders "∞"; that is the honest answer for an
+ // unmounted target and is why the status line above it is the thing to
+ // read first.
+ const volumeDir =
+ media.status === "ok" && media.target ? media.target : paths.channelsDir;
+ // `media` is inspectChannelMedia's answer and it already CARRIES the
+ // marker — reading the file again here was a second read of the same
+ // bytes that could disagree with the status rendered beside it.
+ const marker = media.marker ?? null;
+ const freeBytes = await getFreeBytes(volumeDir);
+ const activeJobs = runningJobs.filter(
+ (j) => j.status === "running" || j.status === "queued",
+ ).length;
+ // The SAME two conditions storageActions.ts refuses on, stated here as
+ // prose so the button is off with a reason rather than off and silent —
+ // and stated in the action too, because a disabled button is a courtesy
+ // and the server is the guard.
+ const blockedReason =
+ activeJobs > 0
+ ? `Finish or cancel ${activeJobs} running/queued job(s) for this channel before moving its media.`
+ : marker
+ ? `A relocation (${marker.direction}) to ${marker.target} is in flight, or was interrupted at phase "${marker.phase}". A channel in transition is not moved again from here — the running job finishes it, and an interrupted one is resumed by rerunning the move.`
+ : null;
+ return (
+ <StorageStage
+ slug={slug}
+ location={media}
+ // From the loaded snapshot, not a walk. Null (rendered "—") when the
+ // snapshot predates the field or does not exist: a 0 would claim a
+ // measurement nobody took.
+ mediaBytes={snapshot.totalAudioBytes ?? null}
+ freeBytes={freeBytes}
+ volumeDir={volumeDir}
+ blockedReason={blockedReason}
+ // A marker with no job behind it is a stale marker: the run that
+ // wrote it is gone, and nothing else will ever clear it.
+ canClearMarker={marker !== null && activeJobs === 0}
+ // The configured cold root, prefilled into the destination box so
+ // it is typed once in Settings instead of once per channel.
+ defaultRoot={settings.storage.mediaRoot}
+ />
+ );
+ }
case "danger":
return (
<div className="flex flex-col gap-4">
diff --git a/editor/app/channels/[slug]/pipelineActions.ts b/editor/app/channels/[slug]/pipelineActions.ts
@@ -1,5 +1,6 @@
"use server";
+import path from "node:path";
import { revalidatePath } from "next/cache";
import {
HANDLING_VALUES,
@@ -43,10 +44,16 @@ import type { Paths } from "yt-dlp-transcript-common/lib/paths";
// between videos via the gate in runManagedDownloads (common/ytdlp/runYtdlp.ts).
async function lowDiskError(
paths: Paths,
+ slug: string,
): Promise<{ ok: false; error: string } | null> {
// The operator clicked this: only the floor applies and the shared hysteresis
- // latch is left alone (see diskGate). runYtdlp re-checks per video mid-batch.
- const disk = await diskGate(paths, getSettings(), { mode: "manual" });
+ // latch is left alone (see diskGate). runYtdlp re-checks per video mid-batch,
+ // against this same volume — the channel's own, which for a relocated channel
+ // is the platter and not the corpus disk.
+ const disk = await diskGate(paths, getSettings(), {
+ mode: "manual",
+ dir: path.join(paths.channelsDir, slug, "data"),
+ });
if (disk.ok) return null;
return {
ok: false,
@@ -139,7 +146,7 @@ async function runPipelineAction(
// store-playlist only fetches the video list (no media written), so it is not
// gated; every other mode downloads audio/subtitles into transcriptsDir.
if (mode !== "store-playlist") {
- const err = await lowDiskError(paths);
+ const err = await lowDiskError(paths, slug);
if (err) return err;
}
return runManagedFunction({
@@ -415,7 +422,7 @@ export async function importVideoAction(
// downloadOneManaged falls back to %(id)s and the reconcile pass repairs the
// dir, so we don't hard-fail here.
const videoId = extractVideoId(videoUrl) ?? undefined;
- const err = await lowDiskError(paths);
+ const err = await lowDiskError(paths, slug);
if (err) return err;
const settings = getSettings();
return runManagedFunction({
diff --git a/editor/app/channels/[slug]/shardActions.ts b/editor/app/channels/[slug]/shardActions.ts
@@ -9,6 +9,7 @@ import {
type ShardOp,
} from "yt-dlp-transcript-common/controller/shard";
import { readChannelConfig } from "yt-dlp-transcript-common/controller/channels";
+import { assertChannelMediaReachable } from "yt-dlp-transcript-common/lib/channelMedia";
import { runYtdlp } from "yt-dlp-transcript-common/ytdlp/runYtdlp";
import { runWhisperBatch } from "yt-dlp-transcript-common/controller/whisperBatch";
import { runAvailabilityCheck } from "yt-dlp-transcript-common/controller/checkAvailability";
@@ -68,6 +69,27 @@ export async function saveShardConfigAction(
const noopLog = () => {};
const signal = new AbortController().signal;
+ // THE ONE TRUE BYPASS (plans/relocate-channel-media.md). Every other path to a
+ // channel's media goes through runManagedFunction and is covered by the guard
+ // there; all three branches below call their runner DIRECTLY, with no job
+ // record, so nothing upstream has checked anything.
+ //
+ // What that costs on an unmounted drive is not a failed save, it is a
+ // confident wrong one: each runner computes its slice over the population it
+ // finds on disk (all video dirs / the post-prefilter missing set / the data
+ // dirs), and an unreachable data/ makes that "every video is missing". The
+ // saved slice then says the whole channel needs re-downloading, is written to
+ // disk, and outlives the mount long enough for another machine to act on it.
+ // A shard is precisely the artifact you do not want silently wrong.
+ //
+ // It is at the top rather than in the download branch alone because all three
+ // read the same dirs for the same reason.
+ try {
+ await assertChannelMediaReachable(paths, slug);
+ } catch (e) {
+ return { ok: false, error: (e as Error).message };
+ }
+
try {
if (op === "transcribe-missing") {
await runWhisperBatch({
diff --git a/editor/app/channels/[slug]/storageActions.ts b/editor/app/channels/[slug]/storageActions.ts
@@ -0,0 +1,137 @@
+"use server";
+
+// WHERE A CHANNEL'S MEDIA LIVES — the three actions the Storage panel drives.
+//
+// All of the mechanism is in common/controller/relocateChannelMedia.ts. What is
+// here is the editor's half: the preview (a plain async call, no job — it is
+// read-only and the operator is waiting on its numbers before committing), and
+// the two directions of the move, each wrapped in runManagedFunction so it gets
+// a job record, a streamed log, a queue slot and a cancel button like every
+// other long action in the editor.
+//
+// THE ACTIVE-JOBS GUARD IS THE RENAME'S, deliberately copied rather than shared:
+// renameChannelAction refuses while the channel has running or queued jobs
+// because the in-memory registry keys by slug and those jobs would be orphaned
+// by the move. A relocation has the same hazard with a sharper edge — a download
+// or a transcribe running against `data/` WHILE its bytes are being copied out
+// would write into the directory the swap is about to replace, and the verify
+// would then fail (which is the safe outcome) or the write would be lost (which
+// is not). The check is cheap and refuses early, before any bytes move.
+//
+// The relocate job's own queue key is channelQueueKey(slug), so a second one is
+// also serialized by the queue — but the queue would make it WAIT, and waiting
+// is the wrong answer for a move the operator can see is already in flight.
+
+import { revalidatePath } from "next/cache";
+import { getPaths } from "yt-dlp-transcript-common/lib/paths";
+import { getRegistry } from "yt-dlp-transcript-common/jobs/registry";
+import type { StreamActionResult } from "yt-dlp-transcript-common/jobs/streamCommand";
+import { clearRelocationMarker } from "yt-dlp-transcript-common/lib/channelMedia";
+import {
+ previewRelocation,
+ relocationRootProblem,
+ type RelocationPreview,
+} from "yt-dlp-transcript-common/controller/relocateChannelMedia";
+import { enqueueRelocation } from "../lib/relocationJob";
+
+export type PreviewRelocationResult =
+ | { ok: true; preview: RelocationPreview }
+ | { ok: false; error: string };
+
+// Read-only: one tree walk of the channel's data dir plus two statfs calls. It
+// runs INLINE rather than as a job because its whole purpose is to answer a
+// question the operator is holding a form open for; a queued job with a log
+// would be a worse way to show two numbers.
+export async function previewRelocationAction(
+ slug: string,
+ root: string,
+): Promise<PreviewRelocationResult> {
+ const trimmed = root.trim();
+ if (!trimmed) return { ok: false, error: "Enter a destination root." };
+ try {
+ const preview = await previewRelocation({
+ paths: getPaths(),
+ slug,
+ root: trimmed,
+ });
+ return { ok: true, preview };
+ } catch (e) {
+ return { ok: false, error: (e as Error).message };
+ }
+}
+
+// The rename's guard (editor/app/channels/actions.ts), returning the shape
+// StreamActionLog already renders rather than an ActionResult.
+function activeJobsRefusal(slug: string, what: string): string | null {
+ const active = getRegistry()
+ .list()
+ .filter(
+ (j) =>
+ j.channelSlug === slug &&
+ (j.status === "running" || j.status === "queued"),
+ );
+ if (active.length === 0) return null;
+ return (
+ `Finish or cancel ${active.length} running/queued job(s) for this channel ` +
+ `before ${what}.`
+ );
+}
+
+export async function relocateChannelMediaAction(
+ slug: string,
+ root: string,
+): Promise<StreamActionResult> {
+ const trimmed = root.trim();
+ if (!trimmed) return { ok: false, error: "Enter a destination root." };
+ // Relative, or inside the corpus. The job refuses both too — it is the guard —
+ // but a root that would copy the channel onto itself should not become a job
+ // record and a log the operator has to open to read the reason.
+ const rootProblem = await relocationRootProblem({
+ paths: getPaths(),
+ slug,
+ root: trimmed,
+ });
+ if (rootProblem) return { ok: false, error: rootProblem };
+ const refusal = activeJobsRefusal(slug, "moving its media");
+ if (refusal) return { ok: false, error: refusal };
+ return enqueueRelocation({ slug, direction: "out", root: trimmed });
+}
+
+export async function moveChannelMediaBackAction(
+ slug: string,
+): Promise<StreamActionResult> {
+ const refusal = activeJobsRefusal(slug, "moving its media back");
+ if (refusal) return { ok: false, error: refusal };
+ return enqueueRelocation({ slug, direction: "back" });
+}
+
+// THE LAST RESORT, and the only one of the three that is not a move.
+//
+// A channel carrying a relocation marker is "in-transition" to every guard: the
+// lane runners skip it, runManagedFunction refuses its media jobs, and its
+// snapshot will not regenerate. That is correct while a move is running and a
+// dead end once one is not — a killed process or a replaced container leaves a
+// marker nothing will ever clear, and the relocate job itself is refused from
+// the panel while the marker stands.
+//
+// So this removes the marker and NOTHING else: no link, no config, no bytes.
+// Whatever inspect() reports afterwards is the truth the disk was already
+// telling underneath it — which may well be `inconsistent`, and that is the
+// honest answer rather than a repair nobody asked for. It is refused while the
+// channel has a live job, because then the marker is not stale: it belongs to
+// the run that is holding it.
+export async function clearRelocationMarkerAction(
+ slug: string,
+): Promise<{ ok: true } | { ok: false; error: string }> {
+ const refusal = activeJobsRefusal(slug, "clearing its relocation marker");
+ if (refusal) return { ok: false, error: refusal };
+ try {
+ await clearRelocationMarker(getPaths(), slug);
+ } catch (e) {
+ return { ok: false, error: (e as Error).message };
+ }
+ revalidatePath(`/channels/${slug}`);
+ revalidatePath("/channels");
+ revalidatePath("/");
+ return { ok: true };
+}
diff --git a/editor/app/channels/[slug]/videos/[id]/videoActions.ts b/editor/app/channels/[slug]/videos/[id]/videoActions.ts
@@ -142,7 +142,7 @@ export async function downloadVideoPipelineAction(
const paths = getPaths();
// This action fetches a full container. Its sibling redownloadToArchiveAction
// has always preflighted the disk; this one never did.
- const lowDisk = await lowDiskError(paths);
+ const lowDisk = await lowDiskError(paths, slug);
if (lowDisk) return lowDisk;
const url = await findVideoSourceUrl(paths, slug, videoId, r.config);
if (!url) {
@@ -218,7 +218,7 @@ export async function redownloadToArchiveAction(
"Could not determine the video URL: no metadata.info.json and the playlist does not contain a matching entry.",
};
}
- const err = await lowDiskError(paths);
+ const err = await lowDiskError(paths, slug);
if (err) return err;
const settings = getSettings();
return runManagedFunction({
@@ -262,8 +262,14 @@ export async function redownloadToArchiveAction(
// applies and the shared hysteresis latch is left alone (see diskGate).
async function lowDiskError(
paths: Paths,
+ slug: string,
): Promise<{ ok: false; error: string } | null> {
- const gate = await diskGate(paths, getSettings(), { mode: "manual" });
+ // The channel's own volume. A relocated channel's bytes land on the platter
+ // through the data/ symlink, so a full SSD is not a reason to refuse them.
+ const gate = await diskGate(paths, getSettings(), {
+ mode: "manual",
+ dir: path.join(paths.channelsDir, slug, "data"),
+ });
if (gate.ok) return null;
return {
ok: false,
@@ -298,7 +304,10 @@ export async function whisperVideoAction(
// disk produces a transcript.json measured in kilobytes, so a low disk
// is no reason to refuse it — the gate belongs on the fetch, not on the
// whole action.
- const gate = await diskGate(paths, getSettings(), { mode: "manual" });
+ const gate = await diskGate(paths, getSettings(), {
+ mode: "manual",
+ dir: path.join(paths.channelsDir, slug, "data"),
+ });
if (!gate.ok) {
throw new Error(
`Cannot download audio for ${videoId}: ${gate.message}. ` +
diff --git a/editor/app/channels/actions.ts b/editor/app/channels/actions.ts
@@ -303,7 +303,18 @@ export async function refreshChannelSnapshotAction(
if (!(await channelExists(paths, slug))) {
return { error: `Channel "${slug}" not found` };
}
- await generateChannelSnapshot(paths, slug);
+ // generateChannelSnapshot now THROWS on a channel whose media is not
+ // reachable (guard 3) rather than writing a snapshot that says every video is
+ // undownloaded. That is the right behaviour and the wrong exception to let
+ // out of a server action: an uncaught throw here reaches the client as a
+ // digest-only "an error occurred", and the one thing the operator needs is
+ // the sentence naming the unmounted drive. Every sibling in this file returns
+ // { error }; so does this.
+ try {
+ await generateChannelSnapshot(paths, slug);
+ } catch (e) {
+ return { error: (e as Error).message };
+ }
revalidatePath(`/channels/${slug}`);
// The Report column on /channels is read off this snapshot, and the row
// action sits next to the marker it flips — so revalidate the list too, not
diff --git a/editor/app/channels/bulkStorageActions.ts b/editor/app/channels/bulkStorageActions.ts
@@ -0,0 +1,138 @@
+"use server";
+
+// MOVE SEVERAL CHANNELS' MEDIA TO THE COLD ROOT, FROM /channels.
+//
+// ONE JOB PER CHANNEL, ALL ON ONE QUEUE — `relocationQueueKey()`, the shared key
+// the per-channel Storage panel uses too. There is no batch controller and there
+// is deliberately not going to be one: the job queue is the batch, and a
+// cross-channel controller would have to reinvent cancellation, resume and the
+// per-channel log that already exist.
+//
+// THE SHARED KEY IS LOAD-BEARING, not tidiness. registry.ts caps a queue key at
+// concurrency 1 and caps NOTHING across keys, so a per-channel key would start
+// every selected channel's rsync at once, onto one destination volume. See
+// relocationQueueKey() for why that breaks the space check as well as the disk.
+//
+// EVERY CHECK RUNS TWICE, AND THAT IS THE DESIGN. What this action refuses at
+// ENQUEUE time is what it can see now — a social channel, a channel with no
+// media to move, one already relocated, one mid-transition, one with live jobs.
+// What it cannot see is the state of the destination twenty minutes from now, so
+// each job runs its own `previewRelocation`-equivalent preflight (space,
+// writability, containment, movable state) when it STARTS. Because the jobs run
+// one at a time, that check measures a root no other move is writing to — so a
+// root that fills up partway through a selection refuses the remainder cleanly,
+// one job at a time, instead of the whole bulk failing at the point the first
+// one is enqueued.
+//
+// A skip is never a failure: the result names every slug that did not queue and
+// why, and the bar renders both numbers.
+
+import path from "node:path";
+import { stat } from "node:fs/promises";
+import { getPaths } from "yt-dlp-transcript-common/lib/paths";
+import { getSettings } from "yt-dlp-transcript-common/lib/settings";
+import { getRegistry } from "yt-dlp-transcript-common/jobs/registry";
+import { isSocialChannel } from "yt-dlp-transcript-common/lib/channelConfig";
+import { inspectChannelMedia } from "yt-dlp-transcript-common/lib/channelMedia";
+import { readChannelConfig } from "yt-dlp-transcript-common/controller/channels";
+import { relocationRootProblem } from "yt-dlp-transcript-common/controller/relocateChannelMedia";
+import { enqueueRelocation } from "./lib/relocationJob";
+import { queueForSlugs, type QueueOutcome } from "./lib/queueForSlugs";
+
+// The same shape syncAllChannelsAction and the per-group stage buttons return,
+// so the bar renders it the same way they do.
+export type BulkRelocateResult = QueueOutcome;
+
+// Follows the link, deliberately: for a relocated channel `data/` is a symlink
+// and what matters is whether the thing it points at is there. (This path only
+// reaches it for a channel inspect() already called in-place, so in practice it
+// is a real directory or nothing.)
+async function isDirectory(p: string): Promise<boolean> {
+ try {
+ return (await stat(p)).isDirectory();
+ } catch {
+ return false;
+ }
+}
+
+// The rename's guard, per channel — the copy in [slug]/storageActions.ts, for
+// the reason stated there: a download writing into `data/` while its bytes are
+// being copied out either fails the verify (safe) or is lost (not).
+function activeJobsRefusal(slug: string): string | null {
+ const active = getRegistry()
+ .list()
+ .filter(
+ (j) =>
+ j.channelSlug === slug &&
+ (j.status === "running" || j.status === "queued"),
+ );
+ if (active.length === 0) return null;
+ return `${active.length} running/queued job(s) for this channel`;
+}
+
+export async function bulkRelocateChannelMediaAction(
+ slugs: string[],
+ root?: string,
+): Promise<BulkRelocateResult> {
+ const paths = getPaths();
+ // The bar's own box wins; blank falls back to the configured cold root. The
+ // settings page is the only writer of that value — this only reads it.
+ const chosen = (root ?? "").trim() || getSettings().storage.mediaRoot.trim();
+ if (!chosen) {
+ return {
+ queued: [],
+ skipped: slugs.map((slug) => ({
+ slug,
+ reason:
+ "no destination root — set a default media root in Settings, or type one here",
+ })),
+ };
+ }
+ if (!path.isAbsolute(chosen)) {
+ return {
+ queued: [],
+ skipped: slugs.map((slug) => ({
+ slug,
+ reason: `the destination root must be an absolute path (got "${chosen}")`,
+ })),
+ };
+ }
+
+ return queueForSlugs(slugs, {
+ skip: async (slug) => {
+ const config = await readChannelConfig(paths, slug);
+ if (!config) return "channel not found";
+ if (isSocialChannel(config)) {
+ return "social channel — it has no downloaded media";
+ }
+ const jobs = activeJobsRefusal(slug);
+ if (jobs) return jobs;
+ const media = await inspectChannelMedia(paths, slug, config);
+ // NOTHING TO MOVE IS A SKIP, NOT A JOB. A channel that has downloaded
+ // nothing has no `data/` at all, and inspect() calls that `in-place` —
+ // correctly, since it is not relocated. Without this the bulk path queues
+ // a job whose only act is to throw the same sentence from
+ // relocateChannelMedia.ts's copy phase, after a job record, a log and a
+ // queue slot. Selecting a whole page of channels is the ordinary way this
+ // is used, and on a fresh corpus most of them are this case.
+ if (!media.relocated && !(await isDirectory(media.dataDir))) {
+ return "nothing to move — no media has been downloaded for it yet";
+ }
+ if (media.marker) {
+ return `a relocation (${media.marker.direction}) to ${media.marker.target} is already in flight`;
+ }
+ // ALREADY RELOCATED IS A SKIP, NOT A FAILURE — including to a DIFFERENT
+ // root. Selecting the whole page and pressing Move is the ordinary way
+ // this gets used, and the channels already on the platter are exactly the
+ // ones that should quietly drop out of the batch.
+ if (media.relocated) {
+ return `already relocated to ${media.target ?? "another root"}`;
+ }
+ // Containment is per-channel because the target is: `<root>/<slug>/data`
+ // can resolve into one channel's directory and not another's.
+ return relocationRootProblem({ paths, slug, root: chosen });
+ },
+ run: (slug) =>
+ enqueueRelocation({ slug, direction: "out", root: chosen }),
+ });
+}
diff --git a/editor/app/channels/components/ChannelForm.tsx b/editor/app/channels/components/ChannelForm.tsx
@@ -645,6 +645,33 @@ export function ChannelForm({
placeholder="(global default)"
hint="Per-channel override for where this channel's persisted source videos live (e.g. a larger disk). Blank uses the global SAVED_VIDEOS_DIR default."
/>
+ {/* READ-ONLY, AND NOT AN INPUT — the one field on this page that is a
+ record of the disk rather than an instruction to it.
+ `config.dataDir` is written ONLY by the relocate job, on success,
+ after the bytes are copied, verified and the symlink is in place.
+ An editable text box here would let the two disagree with a
+ keystroke: type a path nothing was moved to and every reader
+ follows a `data/` link that still points somewhere else, which
+ inspectChannelMedia reports as `inconsistent` and refuses to guess
+ past. So the move is the only writer, and this line just says what
+ it wrote. (It is also not in CHANNEL_FORM_FIELDS, so saving this
+ form preserves it rather than clearing it.) */}
+ <div className="flex flex-col gap-1 text-sm">
+ <span className="font-medium">Media location</span>
+ <span
+ aria-label="media location"
+ className="font-mono text-xs break-all rounded border border-border bg-muted px-2 py-1"
+ >
+ {c?.dataDir?.trim()
+ ? c.dataDir.trim()
+ : "In the channel directory (data/)"}
+ </span>
+ <span className="text-xs text-muted-foreground">
+ Where this channel’s downloaded media actually lives. Change it
+ from the Storage panel, which copies and verifies the bytes before
+ recording anything here.
+ </span>
+ </div>
</Section>
</>
)}
diff --git a/editor/app/channels/components/ChannelStorageBulkBar.tsx b/editor/app/channels/components/ChannelStorageBulkBar.tsx
@@ -0,0 +1,123 @@
+"use client";
+
+// THE SELECTION BAR for moving several channels' media to the cold root.
+//
+// Appears only with a selection, the idiom SyncConsole's bulk bar established.
+// The root box is seeded from settings.storage.mediaRoot and is the only thing
+// on this bar that is not a button: one root for the whole batch, because a
+// per-row destination is a per-row decision and that is what the channel's own
+// Storage panel is for.
+//
+// NO PREVIEW GATE HERE, unlike that panel, and the asymmetry is deliberate: a
+// preview measures ONE channel's tree against one volume, and the useful answer
+// for a batch ("will all of these fit") is not the sum of the previews — the
+// root fills up as the jobs run. Each job therefore re-checks space when it
+// STARTS, and because every relocation in the process shares one queue key the
+// jobs run ONE AT A TIME — so that check measures a root no other move is
+// writing to, and a root that fills partway refuses the remainder one job at a
+// time with the reason in that job's log. What the bar owes the operator is the
+// count and the skips, not a number that would be stale before the second job.
+
+import { useState, useTransition } from "react";
+import Link from "next/link";
+import {
+ bulkRelocateChannelMediaAction,
+ type BulkRelocateResult,
+} from "../bulkStorageActions";
+
+export function ChannelStorageBulkBar({
+ slugs,
+ onClear,
+ defaultRoot,
+}: {
+ slugs: string[];
+ onClear: () => void;
+ // settings.storage.mediaRoot, read on the server. "" when none is configured,
+ // which leaves the box empty and the action refusing with that as the reason.
+ defaultRoot: string;
+}) {
+ const [pending, startTransition] = useTransition();
+ const [root, setRoot] = useState(defaultRoot);
+ const [result, setResult] = useState<BulkRelocateResult | null>(null);
+ const [error, setError] = useState<string | null>(null);
+
+ if (slugs.length === 0) return null;
+ const trimmed = root.trim();
+
+ return (
+ <div
+ aria-label="channel media bulk"
+ className="flex flex-wrap items-center gap-2 rounded border border-border bg-card px-3 py-2 text-sm"
+ >
+ <span className="text-muted-foreground">
+ {slugs.length} selected
+ </span>
+ <input
+ type="text"
+ aria-label="bulk media root"
+ value={root}
+ disabled={pending}
+ onChange={(e) => setRoot(e.target.value)}
+ placeholder="/mnt/platter/archilyzer-media"
+ className="rounded-md border border-border bg-card px-2 py-1 text-xs font-mono disabled:opacity-50"
+ />
+ <button
+ type="button"
+ disabled={pending || trimmed === ""}
+ aria-label="move media for selected channels"
+ title="One relocate job per channel, on that channel's own queue. Channels that are already relocated, mid-move, social, or busy are skipped with a reason."
+ onClick={() =>
+ startTransition(async () => {
+ setError(null);
+ setResult(null);
+ try {
+ setResult(await bulkRelocateChannelMediaAction(slugs, trimmed));
+ } catch (e) {
+ setError((e as Error).message);
+ }
+ })
+ }
+ className="rounded-md border border-border px-2 py-1 text-xs hover:bg-muted disabled:opacity-50"
+ >
+ {pending
+ ? "Queueing…"
+ : `Move media to ${trimmed === "" ? "…" : trimmed}`}
+ </button>
+ <button
+ type="button"
+ disabled={pending}
+ onClick={onClear}
+ className="rounded-md border border-border px-2 py-1 text-xs hover:bg-muted disabled:opacity-50"
+ >
+ Clear
+ </button>
+ {result && (
+ <span
+ aria-label="bulk media move result"
+ className="text-xs text-muted-foreground"
+ title={
+ result.skipped.length === 0
+ ? undefined
+ : result.skipped
+ .map((s) => `${s.slug}: ${s.reason}`)
+ .join("\n")
+ }
+ >
+ Queued {result.queued.length} · skipped {result.skipped.length} ·{" "}
+ <Link href="/jobs" className="underline hover:text-foreground">
+ view jobs
+ </Link>
+ </span>
+ )}
+ {error && (
+ <span
+ role="alert"
+ aria-label="bulk media move error"
+ className="text-xs text-destructive"
+ >
+ {error}
+ </span>
+ )}
+ </div>
+ );
+}
diff --git a/editor/app/channels/components/ChannelsTable.tsx b/editor/app/channels/components/ChannelsTable.tsx
@@ -18,6 +18,11 @@ import { ChannelBuildToggle } from "./ChannelBuildToggle";
import { ChannelSyncButton } from "./ChannelSyncButton";
import { ChannelSyncToggle } from "./ChannelSyncToggle";
import { InlineActionButton } from "../../components/actions/InlineActionButton";
+import {
+ MediaLocationBadge,
+ type MediaBadgeInput,
+} from "../../components/MediaLocationBadge";
+import { ChannelStorageBulkBar } from "./ChannelStorageBulkBar";
// A row is a stat plus its pipeline bands, in column order. The bands are
// projected on the server from the same snapshot the counts come from, so a
@@ -28,6 +33,11 @@ export type ChannelRow = ChannelStat & {
// projected from that report, so its age is the caveat on all of them — which
// is why it belongs beside them rather than on a page of its own.
report: { generatedAt: string | null; state: "current" | "stale" | "missing" };
+ // Where this channel's media physically is, from inspectChannelMedia on the
+ // server. Null for an in-place channel — the overwhelming majority — so the
+ // badge column is empty for them and the two that matter stand out. See
+ // components/MediaLocationBadge.tsx.
+ media: MediaBadgeInput | null;
};
// A column heading for one pipeline. Comes off the operation registry on the
@@ -178,6 +188,7 @@ export function ChannelsTable({
columns,
sections = null,
siteId,
+ defaultMediaRoot = "",
}: {
channels: ChannelRow[];
// Which pipelines to draw, in group order, resolved on the server from the
@@ -190,8 +201,15 @@ export function ChannelsTable({
// pool. That path is today's flat table, unchanged.
sections?: ChannelGroupSection[] | null;
siteId?: string;
+ // settings.storage.mediaRoot, resolved on the server. Seeds the bulk bar's
+ // root box; "" when no cold root is configured.
+ defaultMediaRoot?: string;
}) {
const [sort, setSort] = useState<SortState>(null);
+ // Slugs ticked for a bulk edit. A Set of SLUGS, not indices, so a
+ // re-render that reorders or drops a row cannot retarget the selection —
+ // the same reason SyncConsole keys its selection this way.
+ const [selected, setSelected] = useState<ReadonlySet<string>>(new Set());
// Plain component state, deliberately NOT the URL: router.replace races the
// global AutoRefresh's router.refresh() and gets dropped.
const [grouped, setGrouped] = useState(true);
@@ -206,8 +224,24 @@ export function ChannelsTable({
[channels],
);
const showSections = grouped && !!sections && sections.length > 0 && !!siteId;
- // Eight fixed columns, one per pipeline, then Actions.
- const colSpan = 9 + columns.length;
+ // The select column, eight fixed columns, one per pipeline, then Actions.
+ const colSpan = 10 + columns.length;
+ // Always the intersection with what is on screen: a slug can leave the table
+ // between renders (a scope change, a deletion), and a bulk edit must not act
+ // on a row nobody can see.
+ const selectedSlugs = channels
+ .map((c) => c.slug)
+ .filter((s) => selected.has(s));
+ const allSelected =
+ channels.length > 0 && selectedSlugs.length === channels.length;
+
+ function toggleOne(slug: string) {
+ setSelected((prev) => {
+ const next = new Set(prev);
+ if (!next.delete(slug)) next.add(slug);
+ return next;
+ });
+ }
function onHeaderClick(key: SortKey) {
setSort((prev) => {
@@ -234,6 +268,21 @@ export function ChannelsTable({
<table className="text-sm border-y md:border border-border md:rounded-md md:overflow-hidden w-full">
<thead className="bg-muted">
<tr>
+ <th className="px-2 py-2">
+ <input
+ type="checkbox"
+ aria-label="select all channels"
+ checked={allSelected}
+ onChange={(e) =>
+ setSelected(
+ e.target.checked
+ ? new Set(channels.map((c) => c.slug))
+ : new Set(),
+ )
+ }
+ className="accent-primary"
+ />
+ </th>
<SortableTh
label="Slug"
sortKey="slug"
@@ -319,18 +368,35 @@ export function ChannelsTable({
section.channels.flatMap((c) => rowBySlug.get(c.slug) ?? []),
sort,
).map((c) => (
- <ChannelTableRow key={c.slug} channel={c} columns={columns} />
+ <ChannelTableRow
+ key={c.slug}
+ channel={c}
+ columns={columns}
+ selected={selected.has(c.slug)}
+ onToggle={toggleOne}
+ />
))}
</tbody>
))
) : (
<tbody>
{rows.map((c) => (
- <ChannelTableRow key={c.slug} channel={c} columns={columns} />
+ <ChannelTableRow
+ key={c.slug}
+ channel={c}
+ columns={columns}
+ selected={selected.has(c.slug)}
+ onToggle={toggleOne}
+ />
))}
</tbody>
)}
</table>
+ <ChannelStorageBulkBar
+ slugs={selectedSlugs}
+ onClear={() => setSelected(new Set())}
+ defaultRoot={defaultMediaRoot}
+ />
<div className="flex flex-col gap-1 px-3 py-2 md:px-0">
<BandLegend />
{columns.some((c) => c.id.startsWith("attribution-")) && (
@@ -397,9 +463,13 @@ function PipelineCell({
function ChannelTableRow({
channel: c,
columns,
+ selected,
+ onToggle,
}: {
channel: ChannelRow;
columns: PipelineColumn[];
+ selected: boolean;
+ onToggle: (slug: string) => void;
}) {
return (
<tr
@@ -410,13 +480,25 @@ function ChannelTableRow({
: "")
}
>
+ <td className="px-2 py-2">
+ <input
+ type="checkbox"
+ aria-label={`select ${c.slug}`}
+ checked={selected}
+ onChange={() => onToggle(c.slug)}
+ className="accent-primary"
+ />
+ </td>
<Td className="font-mono">
- <Link
- href={`/channels/${c.slug}`}
- className="underline hover:text-foreground"
- >
- {c.slug}
- </Link>
+ <span className="inline-flex items-center gap-1.5">
+ <Link
+ href={`/channels/${c.slug}`}
+ className="underline hover:text-foreground"
+ >
+ {c.slug}
+ </Link>
+ <MediaLocationBadge media={c.media} compact />
+ </span>
</Td>
<Td>{c.config.name ?? ""}</Td>
<Td>{c.config.handling}</Td>
diff --git a/editor/app/channels/lib/relocationJob.ts b/editor/app/channels/lib/relocationJob.ts
@@ -0,0 +1,74 @@
+import { revalidatePath } from "next/cache";
+import { getPaths } from "yt-dlp-transcript-common/lib/paths";
+import { relocationQueueKey } from "yt-dlp-transcript-common/lib/queueKeys";
+import {
+ runManagedFunction,
+ type StreamActionResult,
+} from "yt-dlp-transcript-common/jobs/streamCommand";
+import { formatBytes } from "yt-dlp-transcript-common/lib/format";
+import { relocateChannelMedia } from "yt-dlp-transcript-common/controller/relocateChannelMedia";
+import type { RelocationDirection } from "yt-dlp-transcript-common/lib/channelMedia";
+
+// ONE ENQUEUE OF THE RELOCATION JOB, for the two callers that have one: the
+// per-channel Storage panel ([slug]/storageActions.ts) and the /channels bulk
+// move (bulkStorageActions.ts).
+//
+// Deliberately NOT in either of those files: both carry "use server", which
+// means every non-type export in them is a server action — a shared helper
+// cannot live there, and exporting this one from a "use server" file would put
+// an unguarded enqueue on the wire under its own endpoint. Same reason
+// lib/queueForSlugs.ts exists.
+//
+// THE GUARDS ARE THE CALLERS'. This function refuses nothing: each caller runs
+// its own active-jobs check and its own root check before reaching here, and the
+// controller re-checks everything at run time regardless. What is shared is only
+// the job record's shape — kind, queue key, channel slug, the summary line and
+// the three revalidations — because those are what must not drift between a
+// single move and a bulk one.
+//
+// THE QUEUE KEY IS THE SHARED ONE AND TAKES NO SLUG, which is the whole reason
+// this helper exists rather than each caller building its own record. See
+// relocationQueueKey() in common/lib/queueKeys.ts: a per-channel key would
+// serialize a channel only against itself, so a bulk move of ten channels would
+// start ten rsyncs onto one destination volume whose run-time space checks would
+// then each be credited room the others had already claimed.
+export async function enqueueRelocation(opts: {
+ slug: string;
+ direction: RelocationDirection;
+ root?: string;
+}): Promise<StreamActionResult> {
+ const { slug, direction, root } = opts;
+ const paths = getPaths();
+ return runManagedFunction({
+ kind: "relocate-channel-media",
+ queueKey: relocationQueueKey(),
+ paths,
+ channelSlug: slug,
+ fn: async (onLog, signal) => {
+ const result = await relocateChannelMedia({
+ paths,
+ slug,
+ direction,
+ root,
+ onLog,
+ signal,
+ });
+ onLog(
+ `${direction === "out" ? "Moved" : "Moved back"} ${result.files} file(s) / ` +
+ `${formatBytes(result.bytes)} — ${result.target}` +
+ (result.resumed ? " (resumed an interrupted move)" : ""),
+ );
+ // No snapshot regen — deliberately, and `relocate-channel-media` is in
+ // NO_REGEN_KINDS so the central hook does not arm one either. The move
+ // changes where the bytes are, not what they are: every count in the
+ // report is identical afterwards, and a regen would be a full walk of the
+ // channel to rewrite the same numbers under a newer timestamp.
+ //
+ // What DOES have to change is what the pages read per render — the badge,
+ // the location line, the free-space figure — so those are revalidated.
+ revalidatePath(`/channels/${slug}`);
+ revalidatePath("/channels");
+ revalidatePath("/");
+ },
+ });
+}
diff --git a/editor/app/channels/page.tsx b/editor/app/channels/page.tsx
@@ -6,6 +6,7 @@ import {
type ChannelBrief,
} from "yt-dlp-transcript-common/controller/channels";
import { getPaths } from "yt-dlp-transcript-common/lib/paths";
+import { inspectChannelMedia } from "yt-dlp-transcript-common/lib/channelMedia";
import {
getSite,
listSiteIds,
@@ -139,8 +140,24 @@ export default async function ChannelsPage({
// change exists to remove.
const snapshots = new Map(briefs.map((b) => [b.slug, b.snapshot]));
const briefBySlug = new Map(briefs.map((b) => [b.slug, b]));
+ // WHERE EACH CHANNEL'S MEDIA IS. Two stats and a small JSON read per channel,
+ // with the config already in hand from the brief — dozens of channels, so a
+ // few hundred syscalls, and explicitly not a corpus walk. It is done here
+ // rather than skipped because an unmounted drive makes every OTHER number on
+ // this row wrong (an unreachable data/ reads as "nothing downloaded"), and a
+ // page of confidently wrong counts with no marking is the failure this badge
+ // exists to prevent. Only a non-in-place location is carried into the row.
+ const mediaBySlug = new Map(
+ await Promise.all(
+ briefs.map(
+ async (b) =>
+ [b.slug, await inspectChannelMedia(paths, b.slug, b.config)] as const,
+ ),
+ ),
+ );
const all: ChannelRow[] = stats.map((stat) => {
const brief = briefBySlug.get(stat.slug);
+ const media = mediaBySlug.get(stat.slug);
return {
...stat,
pipelines: buildChannelBands(snapshots.get(stat.slug) ?? null, ids),
@@ -148,6 +165,14 @@ export default async function ChannelsPage({
generatedAt: brief?.snapshot?.generatedAt ?? null,
state: brief ? reportStateOf(brief) : ("missing" as const),
},
+ media:
+ media && media.status !== "in-place"
+ ? {
+ status: media.status,
+ target: media.target,
+ detail: media.detail,
+ }
+ : null,
};
});
// Scope to the active site's membership; "all sites" shows the full pool.
@@ -194,6 +219,9 @@ export default async function ChannelsPage({
columns={columns}
sections={sections}
siteId={activeSite?.siteId}
+ // The configured cold root, for the bulk bar's box. Read here, not
+ // in the client component — the settings page is its one writer.
+ defaultMediaRoot={getSettings().storage.mediaRoot}
/>
<p
className="text-xs text-muted-foreground"
diff --git a/editor/app/components/MediaLocationBadge.tsx b/editor/app/components/MediaLocationBadge.tsx
@@ -0,0 +1,104 @@
+import type {
+ ChannelMediaLocation,
+ ChannelMediaStatus,
+} from "yt-dlp-transcript-common/lib/channelMedia";
+
+// THE ONE RENDERING OF "where is this channel's media, and can we reach it".
+//
+// Three surfaces draw it — the channel page's Storage panel, the /channels
+// table and the dashboard's channels table — and two of those are client
+// components, so this file must stay import-clean: the only thing it takes from
+// common/lib/channelMedia.ts is a TYPE, which is erased. Nothing here touches
+// the filesystem; the inspect() call that produces the location happens on the
+// server, once per row, and only its result travels.
+//
+// WHAT THE FIVE STATUSES LOOK LIKE, and why there are only three appearances:
+//
+// in-place → NOTHING. The overwhelming majority of channels are in place,
+// and a badge on every row saying "normal" is noise that makes
+// the two that matter harder to see, not easier.
+// ok → neutral. Relocated and reachable is a fact worth stating (the
+// bytes are not on the corpus disk) but it is not a problem.
+// everything → red. unreachable, in-transition and inconsistent are all
+// else "do not trust what this channel's dirs say right now": the
+// first because the drive is not mounted, the second because a
+// move is half-done, the third because disk and config disagree
+// and nothing here is willing to guess which one is right.
+//
+// The `detail` string is the operator's prose from inspect() — the drive path,
+// the phase, the disagreement — and it goes on `title` so a row badge carries
+// the reason without spending a column on it.
+
+// The prop shape, deliberately narrower than ChannelMediaLocation: a row only
+// needs what it draws, so a server page can project three fields onto a client
+// component instead of serializing a whole location per channel.
+export type MediaBadgeInput = Pick<
+ ChannelMediaLocation,
+ "status" | "target" | "detail"
+>;
+
+export type MediaBadgeTone = "neutral" | "danger";
+
+export type MediaBadge = {
+ label: string;
+ title: string;
+ tone: MediaBadgeTone;
+};
+
+const LABELS: Record<ChannelMediaStatus, string | null> = {
+ "in-place": null,
+ ok: "Media relocated",
+ unreachable: "Media unreachable",
+ "in-transition": "Media moving",
+ inconsistent: "Media inconsistent",
+};
+
+// Null means "draw nothing" — an in-place channel, or no location at all (a
+// caller that could not inspect). Both are the same instruction to a renderer.
+export function mediaBadgeOf(
+ media: MediaBadgeInput | null | undefined,
+): MediaBadge | null {
+ if (!media) return null;
+ const label = LABELS[media.status] ?? null;
+ if (!label) return null;
+ const where = media.target ? ` to ${media.target}` : "";
+ return {
+ label: media.status === "ok" ? `${label}${where}` : label,
+ // The detail carries the reason; the target alone is the fallback so a
+ // location written by an older inspect() still says where it points.
+ title: media.detail ?? `${label}${where}`,
+ tone: media.status === "ok" ? "neutral" : "danger",
+ };
+}
+
+const TONE_CLASS: Record<MediaBadgeTone, string> = {
+ neutral: "border-border bg-muted text-muted-foreground",
+ danger: "border-destructive bg-destructive/10 text-destructive",
+};
+
+// `compact` drops the target from the label: a table cell wants the four-word
+// state, and the full path is one hover away on the title.
+export function MediaLocationBadge({
+ media,
+ compact = false,
+}: {
+ media: MediaBadgeInput | null | undefined;
+ compact?: boolean;
+}) {
+ const badge = mediaBadgeOf(media);
+ if (!badge) return null;
+ const label =
+ compact && media?.status === "ok" ? (LABELS.ok as string) : badge.label;
+ return (
+ <span
+ title={badge.title}
+ aria-label={`media location: ${badge.label}`}
+ className={
+ "inline-flex items-center rounded border px-1.5 py-0.5 text-xs font-medium whitespace-nowrap " +
+ TONE_CLASS[badge.tone]
+ }
+ >
+ {label}
+ </span>
+ );
+}
diff --git a/editor/app/components/dashboard/ChannelsTable.tsx b/editor/app/components/dashboard/ChannelsTable.tsx
@@ -7,6 +7,7 @@ import { prioritizeChannelDownloadAction } from "../../operations/actions";
import { InlineActionButton } from "../actions/InlineActionButton";
import { fmtTime } from "../../widget/lib/relativeTime";
import type { DashboardChannel } from "./types";
+import { MediaLocationBadge } from "../MediaLocationBadge";
// The enriched channels table: the plain slug/handling/videos list plus a
// relative "last sync" that ticks live, and per-row inline actions (Sync,
@@ -52,12 +53,15 @@ export function ChannelsTable({
className="border-t border-border align-top"
>
<td className="px-3 py-2 font-mono">
- <Link
- href={`/channels/${c.slug}`}
- className="underline underline-offset-2 hover:text-brand transition-colors"
- >
- {c.slug}
- </Link>
+ <span className="inline-flex items-center gap-1.5">
+ <Link
+ href={`/channels/${c.slug}`}
+ className="underline underline-offset-2 hover:text-brand transition-colors"
+ >
+ {c.slug}
+ </Link>
+ <MediaLocationBadge media={c.media} compact />
+ </span>
</td>
<td className="px-3 py-2">{c.handling}</td>
<td className="px-3 py-2 text-right tabular-nums">
diff --git a/editor/app/components/dashboard/types.ts b/editor/app/components/dashboard/types.ts
@@ -1,3 +1,5 @@
+import type { MediaBadgeInput } from "../MediaLocationBadge";
+
// Enriched per-channel row for the dashboard cockpit's Channels table. Built
// server-side in app/page.tsx from the actionable summary and passed straight
// through (not into client state), so a global AutoRefresh re-seeds it live.
@@ -19,4 +21,10 @@ export type DashboardChannel = {
// Videos with no transcript are NOT in it: they are classified as waiting on
// transcription and counted separately.
digestReachable: number;
+ // Where this channel's media physically is, when that is not "in the channel
+ // directory". Null for an in-place channel — the overwhelming majority — so
+ // the badge marks only the rows whose other numbers may not be trustworthy:
+ // an unmounted drive reads as "nothing downloaded" to every count on this
+ // row. See app/components/MediaLocationBadge.tsx.
+ media: MediaBadgeInput | null;
};
diff --git a/editor/app/page.tsx b/editor/app/page.tsx
@@ -1,6 +1,7 @@
import { readFileSync } from "node:fs";
import type { Metadata } from "next";
import { getPaths } from "yt-dlp-transcript-common/lib/paths";
+import { inspectChannelMedia } from "yt-dlp-transcript-common/lib/channelMedia";
import { getSettings } from "yt-dlp-transcript-common/lib/settings";
import {
getSite,
@@ -67,6 +68,21 @@ export default async function Dashboard({
// Enriched channels table rows (pass-through, not client state, so a global
// AutoRefresh re-seeds them live).
+ // WHERE EACH CHANNEL'S MEDIA IS — two stats and a small JSON read per row,
+ // with the config already in hand. Not a corpus walk. A row whose drive is
+ // not mounted reports every other number as zero, which is exactly why the
+ // dashboard has to be able to say so rather than drawing a confident zero.
+ const mediaBySlug = new Map(
+ await Promise.all(
+ rows.map(
+ async (r) =>
+ [
+ r.channel.slug,
+ await inspectChannelMedia(paths, r.channel.slug, r.channel.config),
+ ] as const,
+ ),
+ ),
+ );
const channels: DashboardChannel[] = rows.map((r) => ({
slug: r.channel.slug,
handling: r.channel.config.handling,
@@ -78,6 +94,12 @@ export default async function Dashboard({
undownloaded: actionableUndownloadedCount(r),
untranscribed: actionableUntranscribedCount(r),
digestReachable: actionableDigestReachableCount(r),
+ media: (() => {
+ const m = mediaBySlug.get(r.channel.slug);
+ return m && m.status !== "in-place"
+ ? { status: m.status, target: m.target, detail: m.detail }
+ : null;
+ })(),
}));
// "Needs work" payload, same shape the widget endpoint the cockpit polls
diff --git a/editor/app/settings/actions.ts b/editor/app/settings/actions.ts
@@ -1,5 +1,6 @@
"use server";
+import path from "node:path";
import { revalidatePath } from "next/cache";
import {
AUTO_REFRESH_INTERVAL_MAX_SECONDS,
@@ -54,6 +55,10 @@ export async function saveSettingsAction(
formData.get("autoRefreshIntervalSeconds") ?? "",
).trim();
const minFreeDiskRaw = String(formData.get("minFreeDiskGB") ?? "").trim();
+ // The cold-storage default. Blank is a legitimate value ("no default root"),
+ // so the only thing rejected is a non-absolute one — see sanitizeStorage for
+ // why a relative root is never resolved.
+ const mediaRoot = String(formData.get("mediaRoot") ?? "").trim();
const resumeMarginRaw = String(formData.get("resumeMarginGB") ?? "").trim();
const inlineTranscribeOnFallback =
formData.get("inlineTranscribeOnFallback") === "on";
@@ -152,6 +157,13 @@ export async function saveSettingsAction(
};
}
+ if (mediaRoot !== "" && !path.isAbsolute(mediaRoot)) {
+ return {
+ ok: false,
+ error: "Default media root must be an absolute path",
+ };
+ }
+
let socialInput: unknown;
try {
socialInput = JSON.parse(String(formData.get("socialLinksJson") ?? "[]"));
@@ -242,6 +254,9 @@ export async function saveSettingsAction(
// Preserve the saved-video backup config on an unrelated settings save (the
// Saved Videos page edits it). writeSettings re-sanitizes it regardless.
savedVideoBackup: getSettings().savedVideoBackup,
+ // THIS FORM IS THE ONE WRITER of the cold-storage default. The Storage
+ // panel and the /channels bulk move only read it.
+ storage: { mediaRoot },
buildPipeline,
// Each edited on its own operation page; preserved here. Slice 3 moved
// these four fieldsets to /operations/<id>, and with them the hidden
diff --git a/editor/app/settings/components/SettingsForm.tsx b/editor/app/settings/components/SettingsForm.tsx
@@ -132,6 +132,12 @@ export function SettingsForm({ initial }: Props) {
hint="Extra headroom above the floor that a disk-stopped pipeline must see before it starts writing again. Without it the first resumed download drops free space back under the floor and the pipeline flaps. Default 2 GB. Set to 0 to resume at the floor."
/>
<Field
+ label="Default media root (cold storage)"
+ name="mediaRoot"
+ defaultValue={initial.storage.mediaRoot}
+ hint="Absolute directory that holds channels moved off the corpus disk, one <slug>/data under it (e.g. /mnt/platter/archilyzer-media). Prefills the root on each channel's Storage panel and is what the bulk move on /channels uses when its own box is blank. Never checked for existence — the drive may be unmounted — and never resolved, so it must be absolute. Leave blank for no default."
+ />
+ <Field
label="Auto-refresh interval (seconds)"
name="autoRefreshIntervalSeconds"
defaultValue={String(initial.autoRefreshIntervalSeconds)}
diff --git a/editor/e2e/channel-storage.spec.ts b/editor/e2e/channel-storage.spec.ts
@@ -0,0 +1,393 @@
+import { lstat, mkdir, readdir, symlink, writeFile } from "node:fs/promises";
+import { join } from "node:path";
+import { test, expect, type Page } from "@playwright/test";
+import { baseUrl } from "./baseUrl";
+import {
+ channelStage,
+ channelVideos,
+ generateReport,
+ pathExists,
+ readJson,
+ resetData,
+ resolvePath,
+ writeChannelConfig,
+ writeSettings,
+} from "./helpers";
+
+// MOVING A CHANNEL'S MEDIA TO ANOTHER DIRECTORY, AND BACK.
+//
+// The thing this spec is really pinning is the claim the whole design rests on:
+// that a relocated channel is indistinguishable from an in-place one to every
+// reader. So the assertions after the move are deliberately NOT about the move —
+// they are the videos list still listing the video and the file route still
+// serving the transcript, through the same URLs, with the bytes now somewhere
+// else entirely.
+//
+// THE REPORT IS GENERATED FIRST AND THEN WAITED OUT. The Storage panel refuses
+// a move while the channel has a queued or running job — the rename's guard, for
+// the sharper reason that a download writing into data/ mid-copy either fails
+// the verify or loses the write — and generateReport leaves exactly such a job
+// behind for a moment after the snapshot lands. quiet() is that wait, and it is
+// the honest one: the guard is real and the spec has to satisfy it rather than
+// race it.
+//
+// A relocation itself queues NO report (`relocate-channel-media` is in
+// NO_REGEN_KINDS: the move changes where the bytes are, not what they are), so
+// there is nothing to wait for between the move out and the move back.
+//
+// The destination is the TEST'S OWN tmp dir (testInfo.outputPath), never a real
+// drive and never a shared path — with --repeat-each each repeat gets its own.
+
+const SLUG = "test-youtube";
+const VIDEO = "20240101_test1234567";
+
+const dataDir = () => resolvePath(`test-transcripts/channels/${SLUG}/data`);
+
+// Wait until the channel has no running or queued job. Polled from the same
+// endpoint /jobs draws, because the guard reads the same registry.
+async function quiet(page: Page): Promise<void> {
+ await expect
+ .poll(
+ async () => {
+ const res = await page.request.get(`${baseUrl}/api/jobs/active`);
+ const body = await res.json();
+ const jobs: { channelSlug?: string; status: string }[] = Array.isArray(
+ body,
+ )
+ ? body
+ : (body.jobs ?? []);
+ return jobs.filter(
+ (j) =>
+ j.channelSlug === SLUG &&
+ (j.status === "running" || j.status === "queued"),
+ ).length;
+ },
+ { timeout: 30_000 },
+ )
+ .toBe(0);
+}
+
+test("relocate a channel's media to another root, and move it back", async ({
+ page,
+}, testInfo) => {
+ test.setTimeout(90_000);
+ await resetData("one-youtube-channel-with-data");
+ await generateReport(page, SLUG);
+ await quiet(page);
+ const root = testInfo.outputPath("media-root");
+ await mkdir(root, { recursive: true });
+ const target = join(root, SLUG, "data");
+
+ // --- before -------------------------------------------------------------
+ await page.goto(channelStage(SLUG, "storage"));
+ await expect(page.getByLabel("media path")).toHaveText(dataDir());
+ // In place draws no badge at all — a badge on every channel saying "normal"
+ // is what makes the one that matters hard to find.
+ await expect(page.getByLabel(/^media location:/)).toHaveCount(0);
+
+ // --- preview gates the move --------------------------------------------
+ const moveButton = page.getByRole("button", { name: "Move media" });
+ // Nothing previewed yet, so the move is not offered even with a root typed.
+ // Asserted through the HINT rather than through the disabled button alone:
+ // the button is also disabled before hydration, so `toBeDisabled()` on its
+ // own would pass for the wrong reason.
+ await page.getByLabel("destination root").fill(root);
+ await expect(
+ page.getByText("Preview this root to enable the move."),
+ ).toBeVisible();
+ await expect(moveButton).toBeDisabled();
+
+ await page.getByRole("button", { name: "Preview" }).click();
+ const preview = page.getByLabel("relocation preview");
+ await expect(preview).toBeVisible({ timeout: 15_000 });
+ // The fixture is two files in one video dir.
+ await expect(page.getByLabel("bytes to move")).toContainText("2 file(s)");
+
+ // --- the move -----------------------------------------------------------
+ await expect(moveButton).toBeEnabled();
+ await moveButton.click();
+ await expect(page.getByLabel("Move media output")).toContainText("Moved", {
+ timeout: 60_000,
+ });
+
+ // data/ is a symlink now, and config.json records the target — written only
+ // by the job, on success, after the copy verified.
+ expect((await lstat(dataDir())).isSymbolicLink()).toBe(true);
+ expect(
+ (await readJson<{ dataDir?: string }>(
+ `test-transcripts/channels/${SLUG}/config.json`,
+ )).dataDir,
+ ).toBe(target);
+ // The source was reclaimed: no parked copy left holding a second copy of the
+ // channel on the volume the move exists to free, and no marker.
+ const siblings = await readdir(
+ resolvePath(`test-transcripts/channels/${SLUG}`),
+ );
+ expect(siblings.filter((n) => n.startsWith("data."))).toEqual([]);
+ expect(
+ await pathExists(`test-transcripts/channels/${SLUG}/.relocating.json`),
+ ).toBe(false);
+
+ // --- NOTHING ELSE NOTICED ----------------------------------------------
+ // The videos list is read off data/ through the same joined path it always
+ // was, and resolves through the link.
+ await page.goto(channelVideos(SLUG));
+ await expect(page.getByText(VIDEO).first()).toBeVisible();
+
+ // So does the per-file route, which resolves channelsDir/<slug>/data/<id>/
+ // and has no idea any of this happened.
+ const file = await page.request.get(
+ `${baseUrl}/api/channels/${SLUG}/videos/${VIDEO}/files/transcript.en.vtt`,
+ );
+ expect(file.status()).toBe(200);
+ expect(await file.text()).toContain("WEBVTT");
+
+ // --- the badge ----------------------------------------------------------
+ await page.goto("/channels");
+ await expect(
+ page.getByLabel(/^media location: Media relocated/),
+ ).toBeVisible();
+
+ // --- back ---------------------------------------------------------------
+ await page.goto(channelStage(SLUG, "storage"));
+ await expect(page.getByLabel("media path")).toHaveText(target);
+ const backButton = page.getByRole("button", {
+ name: "Move back in place",
+ });
+ await expect(backButton).toBeEnabled();
+ await backButton.click();
+ await expect(page.getByLabel("Move back in place output")).toContainText(
+ "Moved back",
+ { timeout: 60_000 },
+ );
+
+ // A real directory again, the config field gone, and the target reclaimed.
+ expect((await lstat(dataDir())).isDirectory()).toBe(true);
+ expect(
+ (await readJson<{ dataDir?: string }>(
+ `test-transcripts/channels/${SLUG}/config.json`,
+ )).dataDir,
+ ).toBe(undefined);
+ const back = await page.request.get(
+ `${baseUrl}/api/channels/${SLUG}/videos/${VIDEO}/files/transcript.en.vtt`,
+ );
+ expect(back.status()).toBe(200);
+});
+
+test("the Configure form shows the media location read-only", async ({
+ page,
+}) => {
+ await resetData("one-youtube-channel-with-data");
+ await page.goto(channelStage(SLUG, "configure"));
+ // exact: the Storage panel's badge is "media location: …", and a substring
+ // match would find either.
+ const line = page.getByLabel("media location", { exact: true });
+ await expect(line).toHaveText("In the channel directory (data/)");
+ // Not an input: it is a record of what is on disk, and the only writer is a
+ // move that succeeded. A text box here would let config and disk disagree
+ // with a keystroke.
+ expect(await line.evaluate((el) => el.tagName)).not.toBe("INPUT");
+});
+
+// THE BULK MOVE, from /channels, with the cold root coming out of Settings.
+//
+// Two channels are selected and exactly one moves. The other is a channel that
+// is ALREADY relocated, built directly on disk — an absolute `data` symlink plus
+// `config.dataDir`, which is precisely what a finished move leaves behind —
+// rather than by running a second relocation first. A real move here would cost
+// a second rsync, a second job wait and a second 60s timeout to assert a skip
+// that is decided before any byte is read; what is under test is the skip, and
+// the state it keys off is three filesystem calls to produce.
+//
+// THE HAPPENS-BEFORE EDGE IS THE BADGE. The bulk action returns when the jobs
+// are ENQUEUED, not when they finish, so "Queued 1" is not permission to read
+// config.json. The row's badge is rendered from inspectChannelMedia on the
+// server, so it cannot appear before the swap has committed — waiting for it is
+// waiting for the job, through the same surface an operator watches. The videos
+// list afterwards is the claim the whole design rests on, checked once more from
+// a channel that got there by a different route.
+test("the /channels bulk move queues one job per channel and skips the rest", async ({
+ page,
+}, testInfo) => {
+ test.setTimeout(120_000);
+ await resetData("one-youtube-channel-with-data");
+ const root = testInfo.outputPath("bulk-root");
+ await mkdir(root, { recursive: true });
+ // THE DEFAULT ROOT, set the way the operator does. The bulk bar's box is
+ // seeded from it, so this is also what asserts the settings value reaches the
+ // client. The other four keys mirror fixtures/test-settings.default.json,
+ // which writeSettings replaces wholesale — notably minFreeDiskGB 0, so the
+ // move is not charged the resume margin on a nearly-full disk.
+ await writeSettings({
+ adminTitle: "Test Admin",
+ minFreeDiskGB: 0,
+ verifyAvailabilityBeforeClean: false,
+ syncScheduler: { fullSweepIntervalMinutes: 0 },
+ storage: { mediaRoot: root },
+ });
+
+ const PRE = "pre-moved";
+ const preTarget = join(root, PRE, "data");
+ await mkdir(join(preTarget, "20240102_pre1234567"), { recursive: true });
+ await writeFile(
+ join(preTarget, "20240102_pre1234567", "transcript.en.vtt"),
+ "WEBVTT\n\n00:00.000 --> 00:01.000\nhello\n",
+ );
+ await writeChannelConfig(PRE, { dataDir: preTarget });
+ await symlink(preTarget, resolvePath(`test-transcripts/channels/${PRE}/data`));
+
+ await page.goto("/channels");
+ // One badge before the move: the channel that is already on the root.
+ const badges = page.getByLabel(/^media location: Media relocated/);
+ await expect(badges).toHaveCount(1);
+
+ // The bar only exists with a selection — it is a selection bar, and a root box
+ // with nothing to apply it to is a control that cannot do anything.
+ await expect(page.getByLabel("bulk media root")).toHaveCount(0);
+ await page.getByLabel("select all channels").check();
+ await expect(page.getByLabel(`select ${SLUG}`)).toBeChecked();
+ await expect(page.getByLabel(`select ${PRE}`)).toBeChecked();
+ // And it carries the configured root with no typing.
+ await expect(page.getByLabel("bulk media root")).toHaveValue(root);
+
+ await page.getByLabel("move media for selected channels").click();
+ const result = page.getByLabel("bulk media move result");
+ await expect(result).toContainText("Queued 1 · skipped 1", {
+ timeout: 30_000,
+ });
+ // The skip names the channel and why — the same title-attribute shape the
+ // "Sync every channel" result uses for its own skips.
+ await expect(result).toHaveAttribute(
+ "title",
+ `${PRE}: already relocated to ${preTarget}`,
+ );
+
+ // Wait for the job through the badge: two relocated channels, not one.
+ await expect
+ .poll(
+ async () => {
+ await page.reload();
+ return badges.count();
+ },
+ { timeout: 90_000 },
+ )
+ .toBe(2);
+
+ // The moved channel: config records the target under the tmp root, `data/` is
+ // a link, and the videos list still lists the video through it.
+ expect(
+ (await readJson<{ dataDir?: string }>(
+ `test-transcripts/channels/${SLUG}/config.json`,
+ )).dataDir,
+ ).toBe(join(root, SLUG, "data"));
+ expect((await lstat(dataDir())).isSymbolicLink()).toBe(true);
+ await page.goto(channelVideos(SLUG));
+ await expect(page.getByText(VIDEO).first()).toBeVisible();
+
+ // The skipped channel was not touched: same target, still a link, no marker.
+ expect(
+ (await readJson<{ dataDir?: string }>(
+ `test-transcripts/channels/${PRE}/config.json`,
+ )).dataDir,
+ ).toBe(preTarget);
+ expect(
+ (await lstat(resolvePath(`test-transcripts/channels/${PRE}/data`)))
+ .isSymbolicLink(),
+ ).toBe(true);
+ expect(
+ await pathExists(`test-transcripts/channels/${PRE}/.relocating.json`),
+ ).toBe(false);
+});
+
+// ONE QUEUE FOR EVERY MOVE, AND A CHANNEL WITH NOTHING TO MOVE NEVER BECOMES A
+// JOB.
+//
+// The queue key is the ONLY thing that decides whether two relocations run at
+// once: registry.ts caps a key at concurrency 1 and caps nothing across keys.
+// A per-channel key therefore starts every selected channel's rsync together,
+// onto one destination volume, and each job's space check runs when it STARTS —
+// so concurrent starts are each credited room the others have already claimed.
+// The shared key is what the bulk bar's "refuses the remainder one job at a
+// time" sentence actually rests on, so it is asserted rather than described:
+// both jobs land on `relocate`, read off the /jobs table's own queue column.
+//
+// The empty channel is the second half: `inspect()` calls a channel that has
+// downloaded nothing `in-place` — correctly, it is not relocated — so without an
+// up-front skip it becomes a job record, a log and a queue slot whose only act
+// is to throw "has no data/ to move". On a fresh corpus most of a page is that.
+test("a bulk move puts every job on one queue and skips a channel with nothing to move", async ({
+ page,
+}, testInfo) => {
+ test.setTimeout(120_000);
+ await resetData("one-youtube-channel-with-data");
+ const root = testInfo.outputPath("queue-root");
+ await mkdir(root, { recursive: true });
+ await writeSettings({
+ adminTitle: "Test Admin",
+ storage: { mediaRoot: root },
+ });
+
+ // A second channel with real media, so the selection queues TWO moves — one
+ // queue key is only a claim about two of them.
+ const SECOND = "second-mover";
+ await writeChannelConfig(SECOND, {
+ handling: "youtube",
+ name: "Second Mover",
+ url: "https://www.youtube.com/@second/videos",
+ });
+ await mkdir(
+ resolvePath(`test-transcripts/channels/${SECOND}/data/20240103_second12345`),
+ { recursive: true },
+ );
+ await writeFile(
+ resolvePath(
+ `test-transcripts/channels/${SECOND}/data/20240103_second12345/transcript.en.vtt`,
+ ),
+ "WEBVTT\n\n00:00.000 --> 00:01.000\nhello\n",
+ );
+
+ // ...and a third with a config and no data/ at all: the skip.
+ const EMPTY = "no-media-yet";
+ await writeChannelConfig(EMPTY, {
+ handling: "youtube",
+ name: "No Media Yet",
+ url: "https://www.youtube.com/@empty/videos",
+ });
+
+ await page.goto("/channels");
+ await page.getByLabel("select all channels").check();
+ await page.getByLabel("move media for selected channels").click();
+
+ const result = page.getByLabel("bulk media move result");
+ await expect(result).toContainText("Queued 2 · skipped 1", {
+ timeout: 30_000,
+ });
+ await expect(result).toHaveAttribute(
+ "title",
+ `${EMPTY}: nothing to move — no media has been downloaded for it yet`,
+ );
+
+ // Both moves land, through the badge, and then both jobs are on ONE queue.
+ await expect
+ .poll(
+ async () => {
+ await page.reload();
+ return page.getByLabel(/^media location: Media relocated/).count();
+ },
+ { timeout: 90_000 },
+ )
+ .toBe(2);
+
+ await page.goto("/jobs");
+ const rows = page
+ .getByRole("row")
+ .filter({ hasText: "Relocate channel media" });
+ await expect(rows).toHaveCount(2);
+ // The queue column: `relocate` for both. The discriminating half is the
+ // negative — a per-channel key prints `channel:<slug>` there, and that is
+ // exactly the shape that would let the two run at once.
+ for (const i of [0, 1]) {
+ await expect(rows.nth(i)).toContainText("relocate");
+ await expect(rows.nth(i)).not.toContainText("channel:");
+ }
+});
diff --git a/editor/e2e/disk-space.spec.ts b/editor/e2e/disk-space.spec.ts
@@ -272,3 +272,59 @@ test("active-jobs API and monitor widget report low disk", async ({ page }) => {
expect(offPayload.disk.enabled).toBe(false);
expect(offPayload.disk.low).toBe(false);
});
+
+// THE HARNESS'S OWN INVARIANT, pinned here because breaking it is invisible.
+//
+// writeSettings() used to REPLACE test-settings.json rather than merge onto
+// fixtures/test-settings.default.json, so a spec that did not name a key
+// inherited the PRODUCT default instead of the fixture's — a different set of
+// numbers, chosen for operators and not for a test host. Two of them cost:
+// minFreeDiskGB 5 GB instead of 0 arms the low-disk gate, so every media action
+// in that spec depends on how much room the host has left; and
+// sleepBetweenDownloadsSeconds 10 instead of 0 puts ten seconds between every
+// download in a fixture batch. Eighteen spec files write settings without
+// naming the floor and fourteen without naming the sleep; none of them is about
+// either.
+//
+// The failure that costs is not a red assertion about disk or about time, it is
+// a red assertion about something else: the preflight in pipelineActions returns
+// `{ ok: false }`, the action never runs, and the spec reports an empty list —
+// or the batch simply does not finish inside the timeout. This asserts the
+// fixture's intent survives a spec's own write, and that a spec's own value
+// still wins over it.
+test("a spec's settings are the fixture's plus what it names", async ({
+ page,
+}) => {
+ await resetData("test-pipeline");
+ await writeSettings({ adminTitle: "Test Admin" });
+
+ await page.goto("/settings");
+ // Unnamed: both come from the fixture, not from the product defaults (5, 10).
+ await expect(page.locator('input[name="minFreeDiskGB"]')).toHaveValue("0");
+ await expect(
+ page.locator('input[name="sleepBetweenDownloadsSeconds"]'),
+ ).toHaveValue("0");
+ // Named: the spec's own value survives the merge.
+ await expect(page.locator('input[name="adminTitle"]')).toHaveValue(
+ "Test Admin",
+ );
+
+ // And the gate itself agrees — the same payload the indicator reads.
+ const res = await page.request.get("/api/jobs/active");
+ const payload = (await res.json()) as {
+ disk: { enabled: boolean; low: boolean };
+ };
+ expect(payload.disk.enabled).toBe(false);
+ expect(payload.disk.low).toBe(false);
+
+ // A named floor still wins, which is what every case above this one needs.
+ await writeSettings({ minFreeDiskGB: HUGE_FLOOR_GB });
+ await page.goto("/settings");
+ await expect(page.locator('input[name="minFreeDiskGB"]')).toHaveValue(
+ String(HUGE_FLOOR_GB),
+ );
+ // ...and the fixture's other keys are still underneath it.
+ await expect(
+ page.locator('input[name="sleepBetweenDownloadsSeconds"]'),
+ ).toHaveValue("0");
+});
diff --git a/editor/e2e/helpers.ts b/editor/e2e/helpers.ts
@@ -73,8 +73,55 @@ export async function resetData(fixtureName: string | null = null) {
await fetch(`${baseUrl}/api/test/invalidate-cache`).catch(() => {});
}
-export async function writeSettings(settings: Record<string, unknown>) {
- await writeFile(testSettingsFile, JSON.stringify(settings, null, 2));
+// A SPEC'S SETTINGS ARE THE FIXTURE'S PLUS WHAT THE SPEC NAMES.
+//
+// This used to write the file WHOLESALE, which meant every key in
+// fixtures/test-settings.default.json — copied in by resetData() moments earlier
+// — was gone the instant a spec called this. The keys did not fall back to
+// something neutral: they fell back to the PRODUCT defaults, which is a
+// different fixture, chosen for operators and not for a test host. Four of them
+// diverge, and each one makes a spec depend on something it never mentions:
+//
+// minFreeDiskGB 0 vs 5 GB — arms the low-disk gate,
+// so every media action in the spec depends on how much room the HOST has
+// left, and the refusal is a returned `{ ok: false }` from
+// pipelineActions.lowDiskError() that no assertion reads.
+// sleepBetweenDownloadsSeconds 0 vs 10 s — 10 s between every
+// download in a fixture batch, which is the difference between a spec that
+// finishes and one that times out on a slow host.
+// verifyAvailabilityBeforeClean false vs true — an extra source probe on
+// the cleanup path.
+// syncScheduler.fullSweepIntervalMinutes 0 vs 1440 — whether a never-swept
+// channel is due for a full sweep or the cheap paged walk.
+//
+// Fourteen spec files call this without naming the sleep; eighteen without
+// naming the floor. Merging is what makes "a spec writes the settings it cares
+// about" true. EXPLICIT WINS at every level, which is what disk-space.spec.ts,
+// widget.spec.ts and backfill.spec.ts's "the disk floor refuses to re-acquire
+// anything" rely on.
+//
+// ONE LEVEL DEEP, and no deeper. `syncScheduler` is the only nested block the
+// fixture sets, and a spec that names it (scheduler.spec.ts, cadence-ui.spec.ts)
+// names the whole scheduler except that one key. Arrays and every other object
+// are REPLACED, not merged — `workers: []` has to mean no workers, and a
+// half-merged policy tree would be a worse surprise than a replaced one.
+type SettingsPatch = Record<string, unknown>;
+
+function isPlainObject(v: unknown): v is SettingsPatch {
+ return typeof v === "object" && v !== null && !Array.isArray(v);
+}
+
+export async function writeSettings(settings: SettingsPatch) {
+ const base = JSON.parse(
+ await readFile(defaultTestSettingsFile, "utf8"),
+ ) as SettingsPatch;
+ const merged: SettingsPatch = { ...base, ...settings };
+ for (const [key, value] of Object.entries(settings)) {
+ if (isPlainObject(value) && isPlainObject(base[key])) {
+ merged[key] = { ...base[key], ...value };
+ }
+ }
+ await writeFile(testSettingsFile, JSON.stringify(merged, null, 2));
await fetch(`${baseUrl}/api/test/invalidate-cache`).catch(() => {});
}
@@ -352,6 +399,7 @@ export type ChannelStage =
| "speakers"
| "cleanup"
| "diagnostics"
+ | "storage"
| "danger";
// The channel overview with one stage panel open. Landing here is what the old
diff --git a/editor/e2e/scheduler.spec.ts b/editor/e2e/scheduler.spec.ts
@@ -67,10 +67,27 @@ test("scheduler queues due channels, skips not-due, and dedups running ones", as
});
// First tick: slow-a is due and gets queued; slow-b is not due.
+ //
+ // ASSERTED WITH THE TICK'S OWN EXPLANATION ALONGSIDE IT, deliberately. This
+ // read `expect(r1.queued).toEqual(["slow-a"])` on its own, and a tick has
+ // FOUR ways to queue nothing: it never ran (`reason: "tick already
+ // running"`), the scheduler is off (`reason: "scheduler disabled"`), the
+ // channel was held back by the selector, or the launch was refused and its
+ // error recorded — `skipped.push({ slug, reason: result.error })` in
+ // runTick.ts, which is where a low-disk preflight, a rate-limit cooldown and
+ // the unreachable-media guard all land. All four print the same
+ // `Received: []`, so a red run names none of them; this one names whichever
+ // it was.
const tick1 = await page.request.post("/api/scheduler/tick");
expect(tick1.ok()).toBeTruthy();
const r1 = (await tick1.json()) as TickResult;
- expect(r1.queued).toEqual(["slow-a"]);
+ expect({
+ queued: r1.queued,
+ reason: r1.reason,
+ // slow-b's "not due" line belongs in skipped and is not a fault, so only
+ // slow-a's own skip is part of the claim.
+ refusedSlowA: r1.skipped.filter((s) => s.slug === "slow-a"),
+ }).toEqual({ queued: ["slow-a"], reason: undefined, refusedSlowA: [] });
expect(r1.queued).not.toContain("slow-b");
// Second tick (no time passed): slow-a's sync is still running, so it's
diff --git a/editor/e2e/widget.spec.ts b/editor/e2e/widget.spec.ts
@@ -81,6 +81,16 @@ const BACKFILL_SETTINGS = {
};
const TWO_WORKERS = {
+ // NAMED, not inherited. Several cases below reach for the widget's "Disk
+ // space" cell, which the layout only draws while the gate is CONFIGURED, and
+ // the e2e fixture switches the gate OFF (minFreeDiskGB 0) as every other spec
+ // wants. This spec used to get it armed by accident — writeSettings replaced
+ // the fixture wholesale and the product default filled the gap. It merges
+ // now, so the one spec that wants the gate on says so. A positive floor
+ // because the cell has to exist, not because the number matters: every
+ // assertion here is about position, and the cell renders whether or not the
+ // floor is met.
+ minFreeDiskGB: 5,
workers: [
{ id: "gpu", name: "GPU", kind: "local", enabled: true, priority: 0, appId: "whisper-cpp", config: {} },
{ id: "cpu", name: "CPU", kind: "local", enabled: true, priority: 1, appId: "whisper-cpp", config: {} },
diff --git a/plans/FACTS.md b/plans/FACTS.md
@@ -3587,3 +3587,62 @@ Unchanged from the Phase 0 entry, plus one thing Phase 1 learned twice: **`git a
it** (it is untracked, not ignored), which carries the Turbopack panic into every worktree of
that commit. Add by path. A worktree also needs a composed fixture site copied into its
`export/public` or the export webServer 500s and Playwright times out at 120 s.
+
+### A channel's media may be on another drive (verified 2026-09-11, branch `storage/relocate-media`)
+
+`channels/<slug>/data` may be an **absolute symlink** to `<root>/<slug>/data` with
+`config.dataDir` recording the target (`common/lib/channelConfig.ts`, written only by the
+relocate job), and the `<slug>/data` suffix is fixed rather than configurable
+(`relocatedDataDir`, `common/lib/channelMedia.ts:102`) so an empty mountpoint can never be
+mistaken for the media and `deleteChannel`/`renameChannel` can recognise a target by shape.
+**No reader changed**: the on-disk contract `channelDir/data/<id>/…` is what yt-dlp's
+cwd-relative writes, the LMDB index (mtimes, which `rsync -a` preserves) and the export build
+already use, and the only `lstat`/`readlink`/`realpath`/`symlink` calls anywhere in `common/`,
+`editor/` or `export/` are in `channelMedia.ts`, `relocateChannelMedia.ts` and
+`renameChannel.ts` — every other hit is a comment or a test. What that buys
+in call-site churn it owes in one failure mode — a dangling link reads as ENOENT and every
+enumerator swallows ENOENT as "this channel has no videos", which to a runner means
+*everything is undownloaded* — so `inspectChannelMedia` (`channelMedia.ts:191`, two stats and
+at most one small JSON read: `in-place` / `ok` / `unreachable` / `in-transition` /
+`inconsistent`) and `assertChannelMediaReachable` (`:319`, only the first two pass) are the one
+place that can tell the two apart, and **four guards plus one bypass** call them:
+`runManagedFunction` refuses before any job record exists (`common/jobs/streamCommand.ts:277`
+and `:292`) for a kind that declares `needsMedia` — **opt-in, absent means false**
+(`common/jobs/jobKinds.ts:60`, `kindNeedsMedia` at `:464`), so an unlisted kind behaves exactly
+as it did before the guard and `relocate-channel-media` writes `false` out because it is the
+thing that fixes an unreachable channel; the four lane runners skip the channel and keep
+running every other one (`common/controller/autoRunner.ts:370-386`, three syscalls per channel
+per tick, the parsed config passed in so it is not a 69th config read); snapshot generation
+throws rather than write a snapshot saying every video is undownloaded
+(`common/controller/channelSnapshot.ts:607-619` — the scheduler keeps the last good
+`snapshot.json` on a failed refresh); `runOperationBatch` asks before deciding the candidate
+list for both operation lanes (`common/controller/operationBatch.ts:1580-1585`), as does
+`normalizeAllTranscripts` (`common/controller/normalizeAll.ts:63`, which skips and counts a
+failure rather than reporting a clean 0/0/0/0); and the ONE path with no job record,
+`saveShardConfigAction` (`editor/app/channels/[slug]/shardActions.ts:88`), asks at the top for
+all three of its branches. The low-disk gate is now **per volume**: `diskGate(paths, settings,
+{ mode, dir })` defaults `dir` to `paths.transcriptsDir` and six media callers pass the
+channel's own `data` path, the hysteresis latch is a `Map<dir, boolean>`
+(`common/lib/diskSpace.ts:215`, `diskGate` at `:243`, `isDiskGateLatched(dir?)` at `:286` —
+no-arg answers for ANY volume), and `getFreeBytes` (`:29`) **walks up to the nearest existing
+ancestor on ENOENT only**, because `channels/<slug>/data` does not exist until a channel's
+first download and the old fail-open Infinity waved that download through; every other errno
+still fails open. `autoRunner`'s lane-level gate deliberately keeps the corpus volume — it runs
+before the pick, so there is no channel yet. **A relocation root is contained by rule, not by
+hint**: `relocationRootProblem` (`common/controller/relocateChannelMedia.ts:169`, asked by the
+preview, the action and the job — `:387`, `:491`) refuses blank, relative, inside the corpus,
+and a root whose `<slug>` level resolves back into the channel dir, comparing REAL paths and
+resolving the target separately from the root; without it `root = <transcriptsDir>/channels`
+makes the source its own target, `rsync -a src/ src/` succeeds, the verify compares the tree
+with itself, and the reclaim deletes the only copy. **Every relocation in the process runs on ONE queue**, `relocationQueueKey()` =
+`"relocate"` (`common/lib/queueKeys.ts`), set by the single enqueue both the Storage panel and
+the `/channels` bulk move share (`editor/app/channels/lib/relocationJob.ts`) — because
+`registry.ts` caps a queue key at concurrency 1 and caps nothing across keys, so a per-channel
+key would start every selected channel's rsync at once onto one destination volume and each
+job's start-time space check would be credited room the others had already claimed.
+Serializing a move against the channel's OWN jobs is done by refusal, not by the queue: both
+actions reject a channel with running or queued jobs before they enqueue.
+`channels/<slug>/.relocating.json`
+(`channelMedia.ts:40`) is the in-flight marker: present means "in transition" to every guard,
+its `phase` is what lets an interrupted move resume, `deleteChannel` and `renameChannel` refuse
+while it exists, and `clearRelocationMarker` (`:160`) removes it and nothing else.
diff --git a/plans/STATE.md b/plans/STATE.md
@@ -3,7 +3,11 @@
The working memory for the local-AI derived-corpus work. Rewritten at the end of every
session, before context is cleared. See [`README.md`](README.md) for the protocol.
-**Last updated:** 2026-09-08 — **one-core Phase 1 shipped** on branch `one-core/phase-1`
+**Last updated:** 2026-09-11 — **`relocate-channel-media` shipped**, all three slices, on
+branch `storage/relocate-media` off `61eae05` and unmerged; the entry with the rollout order is
+below, under the Phase 1 record it builds on.
+
+**2026-09-08 — one-core Phase 1 shipped** on branch `one-core/phase-1`
(`7f294df` → `81a663f` plus a docs commit, 36 commits, not merged): **dispatch is one scheduler, and the lane is the
noun.** The slice-level record — every sha range, every divergence, both operator gates — is
[`one-core-phase-1.md`](one-core-phase-1.md); the umbrella is
@@ -155,15 +159,40 @@ assertion, and it was right: writing the two new snapshot entries doubled `/chan
and Transcribe coverage, because that page passes the external ids into `buildOperationBands`
and `addRegistryEntry` folded them on top of `addExternalBands`. Fixed in `d8754d2`.
-**Next:** [`relocate-channel-media.md`](relocate-channel-media.md) FIRST, 3 slices —
-`/home` is at **100 %, 6.9 G free**, and the mechanism (a symlinked `data/` plus a
-`config.dataDir` record and four guards) is orthogonal to phase 2, so it neither waits on
-nor complicates the contract work. Then [`channel-priority.md`](channel-priority.md) (one channel priority model that
+**2026-09-11 — [`relocate-channel-media.md`](relocate-channel-media.md) SHIPPED**, all three
+slices, on branch `storage/relocate-media` (`4059dad` → `affe525`, 24 commits off `61eae05`,
+**unmerged**; the suite is **523/523 in 22.3 min** at `affe525`, and
+`common` is 971/971). A channel's `data/` can be an absolute symlink to
+another drive with `config.dataDir` recording the target, moved by a Storage panel on the
+channel page or in bulk from `/channels`; four guards plus a `needsMedia` flag on the job kind
+stand between an unmounted drive and a re-download, and the low-disk gate now measures the
+volume the bytes are going to and latches per volume. **Nothing on disk moved** — the mechanism
+shipped, the bytes did not. The slice table, the two review rounds and eight divergences are in
+that plan's "As shipped (2026-09-11)"; the anchors are in
+[`FACTS.md`](FACTS.md#one-core-phase-1-verified-2026-09-08).
+
+**The rollout is the operator's, in this order:**
+
+1. **Mount the platter.** `sdb1` (1.8 T, ext4) at `/mnt/platter` with `nofail` in fstab;
+ create `/mnt/platter/archilyzer-media` and `/mnt/platter/archilyzer-saved-videos`, owned by
+ `user`. Nothing below works before this, and no code needed it.
+2. **Step 0, the saved-video store, by hand** — the plan's runbook: rsync
+ `transcripts/saved-videos/` out, verify with `--dry-run --itemize-changes`, move the
+ original aside, symlink the store root, restart, then delete the original. **That is the
+ step that frees the 130 GB**, it needs no code at all (every pointer carries an absolute
+ `dir` and nothing walks the store root), and it goes first because a 100 %-full `/home` is
+ a hazard to every unrelated writer while the channel moves run.
+3. **Channels through the UI, largest deprioritized first.** Storage panel per channel, or tick
+ rows on `/channels` and use the bulk bar; `omnimirror` (130.3 GB) is the obvious first move.
+ Sizes are in the plan's table — they are sizes, not priorities.
+
+**Next:** [`channel-priority.md`](channel-priority.md) (one channel priority model that
replaces the two `excludeFrom*` flags and compiles to the four lane trees; S0 contract, then
-S1–S4 in parallel, S5 migration last; built on a branch off this tip, merged after relocate).
-Then phase 2 (the contract: one `ArchiveReader`), 3 slices. Read
+S1–S4 in parallel, S5 migration last), merged after relocate. Then phase 2 (the contract:
+one `ArchiveReader`), 3 slices. Read
`common/architecture.test.ts`'s allow-list first — it is the shortest accurate statement of
-what is still tangled, it shrank by one across phase 1, and no slice added an entry.
+what is still tangled, it shrank by one across phase 1, and no slice added an entry, relocate
+included.
**Previously:** 2026-09-07 — **one-core Phase 0 shipped** on branch `one-core/phase-0`
(`df5eb48` → `1691c4f`, six commits, not merged): **guardrails and dead weight**, every item a
diff --git a/plans/relocate-channel-media.md b/plans/relocate-channel-media.md
@@ -441,3 +441,172 @@ platter through the link, and the disk gate watches **that** volume for it.
- Any coupling between relocation and auto-queue weights or lane priorities. Noted in §8 as a
possible later step; not designed, not built.
- Moving `index.mdb` (13 GB) — it is hot and small relative to the media.
+
+---
+
+## As shipped (2026-09-11)
+
+Branch `storage/relocate-media`, **`4059dad` → `affe525`** — 24 commits off `61eae05` to that
+code tip (the plans commit that opened the branch, 19 of code and tests, the suite fix, the
+records, and the two merge-review fixes), plus this record. **Unmerged.** All three slices landed on the day they were planned. **No data
+moved**: every byte in `transcripts/` is where it was, and the rollout below is still the
+operator's.
+
+| Slice | Commits | What landed |
+|---|---|---|
+| 1 — core | `1cf4e64` `5b0b18f` `a371832` `cb616e7` `c26b1d5` | `common/lib/channelMedia.ts`, the four guards + `needsMedia`, `common/controller/relocateChannelMedia.ts`, delete/rename reaching the other drive, the per-dir disk gate |
+| 2 — job + UI | `2cf7e37` `edbe6bd` `520a683` `a5dae31` `4c1b849` `1c5ef7b` `0d185a8` `0111a86` | the `relocate-channel-media` kind + `storageActions.ts`, the Storage panel and the badges, the ancestor walk, resume-observes-the-disk, two more `data/` readers, the shardActions bypass, the three docs, `editor/e2e/channel-storage.spec.ts` |
+| 3 — cold root + bulk | `bb92c00` `4039355` `00c1116` `45172e4` `96d2e25` `22a3b35` | `storage.mediaRoot` + `sanitizeStorage`, the two review rounds, per-row selection on `/channels`, `bulkStorageActions.ts`, the bulk e2e case |
+
+### The three review rounds, and what closed them
+
+**Round one (`4039355`) — the two ways a move ended by deleting the only copy.** Both were
+reachable from the shipped panel and both ended in `rm -r` on media nothing else held.
+*Move back had no state precondition*: `moveOut` refuses unless the channel is in-place,
+`moveBack` never asked, and `inspect()` reports `relocated: true` for `inconsistent` and
+`unreachable` as well as `ok` — so the panel offered "Move back in place" for a channel whose
+config records a target while `data/` is a real directory, which is exactly what
+`rsync --copy-links` of a channel produces. Closed by the mirror refusal in `moveBack` plus a
+disabled button with the reason. *Nothing checked where the root was*: `root =
+<transcriptsDir>/channels` makes `relocatedDataDir(root, slug)` the SOURCE, `rsync -a src/
+src/` succeeds, `verifyCopy` compares the tree with itself, and the reclaim sweep then deletes
+the only copy — every check in the happy path being the tree against itself. Closed by
+`relocationRootProblem`, asked by all three callers (the preview the operator reads, the
+action that enqueues, the job that moves), comparing REAL paths and resolving the target
+separately from the root. The unit harness had nested its "platter" inside the corpus, which
+is the shape now refused, and is why 16 green tests never saw it.
+
+**Round two (`00c1116`) — a rollback that undid a step it had not recorded, and four smaller
+edges.** `renameChannel` recorded `movedMedia` and `relinked` but not the `unlink` of `data/`
+between them, so a throw from the `symlink` rolled the media directory back while `data/`
+stayed deleted — `inconsistent`, which round one had just made a state move-back refuses, so
+the rollback left a channel with no way back through the UI. Also closed: the missing resume
+cell (back @ copy — the only resume that finishes an rsync rather than a rename, and the only
+one asked to credit a partial copy); a preview that never checked the root existed (the
+ancestor walk statfs'd the parent of a typo and previewed plausible numbers the job then
+refused); move back's space check using neither the resume margin nor credit for the bytes
+already in `data.incoming`; and three nits (`moveBack`'s swap refusing a `data/` of kind
+"other", `sweepParked`'s unused `keep`, and the channel page reading the marker twice while
+telling the operator to delete it by hand above the button that does it).
+
+### Round three (`7722acf`, `affe525`) — the merge review
+
+Both round-one blockers re-verified closed; two should-fixes landed before this became the
+merge candidate.
+
+**Ten ticked rows started ten rsyncs onto one drive.** The queue key is the only thing that
+decides whether two moves run at once — `registry.ts` submits every non-empty key at
+concurrency 1 and caps NOTHING across keys — so `channelQueueKey(slug)` serialized a channel
+against its own downloads and against no other channel. A bulk move therefore started one
+rsync per selected channel, simultaneously, all writing to one destination volume: the worst
+access pattern a platter has, and a broken space check, because each job runs its preflight
+when it STARTS and jobs that start together are each credited room the others have already
+claimed. Recoverable (ENOSPC aborts the copy, the source is untouched until a verify passes)
+but hours of copying to learn what one queue slot knew. Both comments asserted the opposite,
+and the bulk bar's no-preview-gate argument rests entirely on it. `relocationQueueKey()` takes
+no slug and lives in `common/lib/queueKeys.ts` beside `BACKFILL_QUEUE`; the one enqueue both
+actions share sets it, so there is no per-caller key to drift. Serializing against the
+channel's own jobs is not lost — both actions refuse a channel with running or queued jobs
+before they enqueue, which refuses rather than waits. `queueKeys.test.ts` states the claim as a
+function of the slug on both sides so it still reads as the claim (and fails) if someone gives
+the function a slug parameter; the e2e case asserts it where an operator sees it, two rows in
+the `/jobs` table both on `relocate` and neither on `channel:`. **And a channel with nothing to
+move is a skip, not a job**: no `data/` is `in-place` to `inspect()` — correctly — so the bulk
+path used to queue a job whose only act was to throw from the copy phase, which on a fresh
+corpus is most of a page.
+
+**A spec's settings are the fixture's plus what it names.** The suite fix below defaulted one
+key; that was the symptom. `writeSettings` REPLACED `test-settings.json`, so every key in
+`fixtures/test-settings.default.json` fell back not to something neutral but to the PRODUCT
+defaults — a different fixture, chosen for operators. Four diverge: `minFreeDiskGB` 0 vs 5 GB,
+`sleepBetweenDownloadsSeconds` 0 vs 10 s (ten seconds between every download in a fixture
+batch), `verifyAvailabilityBeforeClean` false vs true, `syncScheduler.fullSweepIntervalMinutes`
+0 vs 1440. Eighteen spec files write settings without naming the floor, fourteen without naming
+the sleep. It merges now, one level deep for nested blocks, arrays and every other object
+replaced rather than merged (`workers: []` has to mean no workers), explicit winning at every
+level. The two other keys the fixture sets are inert: no spec asserts the admin title, and
+8388608 IS `TRANSCRIPT_PAGE_DEFAULT_BYTES`.
+
+### Merging with `channel-priority/s5`
+
+One semantic conflict, and it is in the file both branches changed for the same reason.
+`autoRunner.ts`'s `listChannelMeta` returns `{ meta, slugs }` on s5, with the paused channels
+filtered out of `meta` and every slug kept in `slugs` — and it DROPS `config` from
+`ChannelMeta`. This branch adds `config` precisely so the media guard can pass it:
+`inspectChannelMedia(paths, m.slug, m.config)`, which is what keeps the guard from re-reading
+68 `config.json` files on every tick of four lanes and every three-second status poll.
+**Resolution: keep s5's return shape and its paused filter, AND keep `config` on each meta
+entry plus the media skip in `buildChannelWork`.** The two are orthogonal — one decides which
+channels are eligible, the other which of those can be reached.
+
+Everything else is keep-both: `ChannelsTable.tsx` (four hunks), `channels/page.tsx` (two),
+`CHANGELOG.md`, `STATE.md`, `FACTS.md`. `jobKinds.ts` does not conflict. The other files the
+two branches share — `backfillReacquire.ts`, `channelConfig.ts`, `settings.ts`,
+`channels/actions.ts`, `settings/actions.ts` — touch disjoint hunks.
+
+### Divergences from the plan
+
+- **`needsMedia` is opt-in and absent means false** (`common/jobs/jobKinds.ts:457-462`). The
+ plan described the guard, not the default. An unlisted or unknown kind behaves exactly as it
+ did before the guard existed, which is what keeps `refresh-report` — not in the table at all
+ — running and reporting its own refusal. `relocate-channel-media` writes `needsMedia: false`
+ out rather than omitting it: it is the thing that FIXES an unreachable channel, so a `true`
+ there would refuse it precisely when the operator needs it.
+- **`relocate-channel-media` joins `NO_REGEN_KINDS`** (`common/jobs/snapshotScheduler.ts:50`),
+ which the plan did not call for. A move preserves every byte and mtime (`rsync -a`, the same
+ property that makes the LMDB index a no-op), so a regen would walk 11,000 video dirs to write
+ a byte-identical snapshot with a newer `generatedAt` immediately after moving 130 GB. It also
+ removes a race the operator would feel: the Storage panel refuses a move while the channel
+ has a queued job, so "move out, then move back" would be blocked by a report nobody needed.
+- **Preview is the confirm.** The plan said "live preview, confirm" as two things. The panel
+ makes them one: Move is disabled until the root in the box is the root a preview described
+ (`StorageStage.tsx:193-194`), and editing the input un-confirms it. There is no second
+ dialog, and there is no way to move to a root whose numbers the operator has not seen.
+- **A Clear-marker hatch.** Not in the plan, and needed once the marker became a state every
+ guard refuses: a marker whose run is gone (a killed process, a container replaced mid-copy)
+ is otherwise a dead end. `clearRelocationMarker` removes the marker and nothing else — no
+ link touched, no config rewritten, nothing deleted — so what `inspect()` says afterwards is
+ the truth the disk was already telling.
+- **Root containment is a rule, not a hint.** `relocationRootProblem` refuses blank, relative,
+ inside-the-corpus and resolves-into-the-channel-dir. The plan's UI copy ("an absolute
+ directory that already exists") described `transcripts/channels` almost word for word.
+- **`sanitizeStorage` DROPS a relative root to blank rather than resolving it**
+ (`common/lib/settings.ts`). Resolving would anchor the default to whatever cwd the editor
+ booted in — a different directory under docker, under a worktree and under `pnpm dev` — so
+ one settings.json would name three drives. Existence is deliberately not checked: the whole
+ point of a cold root is a drive that may be unmounted when settings are read.
+- **`autoRunner.ts`'s lane gate kept the corpus volume**, though the plan listed it among the
+ callers to thread. That check runs before the pick, so there is no channel yet and the lane
+ spans 68 of them on however many volumes; the per-channel answer is taken further down in
+ `runYtdlp`, where the slug is known and the bytes are about to be written.
+- **The worktree port block drifted under the work.** `pnpm wt list` assigns by position in
+ `git worktree list`, so `storage/relocate-media` moved from index 7 to index 8 (3711/3710 →
+ 3811/3810) while the branch was in flight, and `queue-lock --ports` only VERIFIES a block is
+ free — it does not set it. A hardcoded block in a runbook goes stale; read `pnpm wt ports`
+ each session, and pass the values as env.
+- **The e2e settings fixture is replaced, not merged, and it bit once here.**
+ `helpers.writeSettings` overwrites `test-settings.json` wholesale, so a spec that does not
+ name `minFreeDiskGB` inherits the PRODUCT default of 5 GB rather than the fixture's 0 — and
+ `/home` was at 6.9 GB free. `channel-storage.spec.ts:209` failed 4/4 on exactly that (the
+ move's space check is charged floor + resume margin = 7 GB) and was fixed by writing the key
+ out by hand at `:219-221`. The general fix landed with the suite below, and took `widget.spec.ts` — which wanted the gate ARMED and was getting it from the same gap — with it.
+
+### The suite
+
+**523 passed, 0 failed of 523, 22.3 min** at the merge-candidate tip `affe525`, one worker
+behind the machine-global queue lock, from a worktree with a composed fixture site in
+`export/public`. `common` is **971/971** and `tsc --noEmit` is clean in both packages. It was
+522/522 at `549dd2e` before the two merge-review fixes, which added the queue-key case. The
+run before the fix, at `22a3b35`, was **519 passed / 2 failed** — `backfill.spec.ts:457` and
+`scheduler.spec.ts:29` — and the first run WITH it was 521/1, the one failure being
+`widget.spec.ts:591`, which wanted the disk gate armed and had been getting it by accident
+from the same fixture gap; it names its floor now.
+
+Neither of the original two is this branch's: `backfill.spec.ts:457` is an old
+`uncheck()`-did-not-take flake, reproduced twice in six repeats here; `scheduler.spec.ts:29` did not
+reproduce at all — 8 green runs at the same sha with a clean tree, including an exact
+replication of the failing command — and the tick path this branch touched cannot empty
+`queued` on that fixture (the media guard answers `in-place` for `slow-a`, and the disk gate
+measures the same volume before and after the change, proved offline). Its assertion now
+carries the tick's own `reason` and slow-a's own skip line, because a tick has four ways to
+queue nothing and all four printed the same `Received: []`.