Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit eb5ba0f4866406ce042919f8388e7c33b7d9d1df
parent 12c231722b8042a2f0807a1153f1c87f7134287e
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Fri, 11 Sep 2026 11:43:46 -0400

common+editor: two more readers of data/, and a rollback that undid what never ran

The needsMedia census had two misses, and the test is "does it open data/", not
"does it read like bookkeeping". `sync` reads the data dir to decide what is
already downloaded before downloading into `data/<id>/` — against an unmounted
drive it concludes nothing is downloaded, which is the 130 GB re-fetch this
whole guard exists to prevent. `normalize-transcripts` walks every video dir and
writes a sidecar into each; unmounted, it reports a clean 0/0/0/0 and moves on.
normalizeAll.ts is also the FIFTH `readdir(dataDir).catch(() => [])` swallow, so
it takes the same guard as the snapshot and the batch: a sweep that silently
skips a channel is worse than one that stops on it.

renameChannel's media rollback ran unconditionally, including when the failure
was the pre-check that fires BECAUSE something unrelated already occupies
<root>/<newSlug>. It then renamed that stranger to <root>/<oldSlug> — destroying
a directory the function had never touched, in the name of undoing a move it had
not made. Each step records that it ran and the catch replays only those, link
first: a failure between the symlink and the config write used to leave `data/`
pointing at the new slug while everything else rolled back to the old one, which
is a dangling link, which reads as an unmounted drive.

The media guard no longer re-reads all 68 config.json on every tick of four
lanes and every three-second status poll: listChannelMeta already holds each
parsed config and now carries it. The snapshot generator reads its config once,
ahead of the guard, instead of twice.

And refreshChannelSnapshotAction catches. generateChannelSnapshot throws on an
unreachable channel now, which is right, and an uncaught throw out of a server
action reaches the client as a digest-only "an error occurred" — while the one
thing the operator needs is the sentence naming the unmounted drive.

The panel gets the marker escape hatch: offered only when a marker is present
AND no job is running, because with a live job the marker is not stale. That is
a separate answer from "why are the buttons disabled" — the marker itself
disables them — so it is computed on the server and passed as its own flag
rather than read off the absence of a block reason, which would have hidden the
hatch in exactly the state it exists for.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
Mcommon/controller/autoRunner.ts | 15+++++++++++++--
Mcommon/controller/channelSnapshot.ts | 9++++++---
Mcommon/controller/normalizeAll.ts | 13+++++++++++++
Mcommon/controller/renameChannel.ts | 28+++++++++++++++++++++++++---
Mcommon/jobs/jobKinds.ts | 23+++++++++++++++++++----
Meditor/app/channels/[slug]/components/stages/StorageStage.tsx | 60++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/app/channels/[slug]/page.tsx | 3+++
Meditor/app/channels/[slug]/storageActions.ts | 38+++++++++++++++++++++++++++++++++++---
Meditor/app/channels/actions.ts | 13++++++++++++-
9 files changed, 186 insertions(+), 16 deletions(-)

diff --git a/common/controller/autoRunner.ts b/common/controller/autoRunner.ts @@ -1,5 +1,6 @@ import path from "node:path"; import type { Paths } from "../lib/paths"; +import type { ChannelConfig } from "../lib/channelConfig"; import { mapConcurrent } from "../lib/concurrency"; import { getPaths } from "../lib/paths"; import { getSettings } from "../lib/settings"; @@ -283,7 +284,14 @@ export function getAutoRunnerStatus(kind: AutoQueueKind): AutoRunnerStatus { // --- Pending-work construction from snapshots ------------------------------ -type ChannelMeta = { slug: string; platform: ReturnType<typeof detectPlatform> }; +type ChannelMeta = { + slug: string; + platform: ReturnType<typeof detectPlatform>; + // The parsed config, carried so the media guard below does not re-read all 68 + // config.json files on every tick of four lanes and on every three-second + // status poll. listChannelConfigs has already parsed them. + config: ChannelConfig; +}; const SNAPSHOT_READ_CONCURRENCY = 64; @@ -294,6 +302,7 @@ async function listChannelMeta(paths: Paths): Promise<ChannelMeta[]> { return configs.map(({ slug, config }) => ({ slug, platform: detectPlatform(config.url), + config, })); } @@ -365,7 +374,9 @@ async function buildChannelWork( // that was too full to hold it. Three syscalls per channel per tick, run // concurrently alongside the snapshot reads. const media = await mapConcurrent(meta, SNAPSHOT_READ_CONCURRENCY, (m) => - inspectChannelMedia(paths, m.slug), + // The config is passed, not re-read: without it this is a 69th, 70th … + // config.json read per tick for a field listChannelConfigs already parsed. + inspectChannelMedia(paths, m.slug, m.config), ); for (const [i, { slug, platform }] of meta.entries()) { const snap = snaps[i]; diff --git a/common/controller/channelSnapshot.ts b/common/controller/channelSnapshot.ts @@ -611,7 +611,12 @@ export async function generateChannelSnapshot( // (/cleanup, the channel bands, all four lanes' work lists) reads that // snapshot as the truth. Throw instead: the scheduler keeps the last good // snapshot.json on a failed refresh, which is exactly the right outcome. - await assertChannelMediaReachable(paths, slug); + // The config is read HERE rather than in the fan-out below and passed in, so + // the guard does not open config.json a second time for the one field it + // needs. One extra await in front of a function that then walks the whole + // channel. + const config = await readChannelConfig(paths, slug); + await assertChannelMediaReachable(paths, slug, config); // Heal any video dir that drifted from the canonical id layout before we read // data/* (best-effort; never fail snapshot generation on a reconcile error). @@ -626,7 +631,6 @@ export async function generateChannelSnapshot( urls, archive, failedListed, - config, maybeMissingRecord, roster, ] = await Promise.all([ @@ -634,7 +638,6 @@ export async function generateChannelSnapshot( readPlaylistUrls(playlistPath), readArchive(archivePath), loadFailedTranscriptions(paths, slug), - readChannelConfig(paths, slug), loadMaybeMissing(paths, slug), loadRoster(paths, slug), ]); diff --git a/common/controller/normalizeAll.ts b/common/controller/normalizeAll.ts @@ -9,6 +9,7 @@ import pLimit from "p-limit"; import { listChannelStatsFromDisk } from "./channels"; import { normalizeTranscript } from "./normalizeTranscript"; import type { Paths } from "../lib/paths"; +import { assertChannelMediaReachable } from "../lib/channelMedia"; export type NormalizeAllOptions = { paths: Paths; @@ -53,6 +54,18 @@ export async function normalizeAllTranscripts( for (const ch of channels) { if (opts.signal?.aborted) break; const dataDir = path.join(opts.paths.channelsDir, ch.slug, "data"); + // GUARD, same reason as the snapshot's and the batch's: `readdir(dataDir) + // .catch(() => [])` cannot tell "this channel has downloaded nothing" from + // "this channel's drive is not mounted", and here the second reads as a + // clean run over zero videos that reports 0/0/0/0 and moves on. A sweep + // that silently skips a channel is worse than one that stops on it. + try { + await assertChannelMediaReachable(opts.paths, ch.slug, ch.config); + } catch (err) { + log(`Normalize ${ch.slug}: SKIPPED — ${(err as Error).message}`); + result.failed++; + continue; + } const videoIds = await readdir(dataDir).catch(() => [] as string[]); log(`Normalize ${ch.slug}: ${videoIds.length} videos`); let wrote = 0; diff --git a/common/controller/renameChannel.ts b/common/controller/renameChannel.ts @@ -136,13 +136,23 @@ export async function renameChannel( if (conventional) { const mediaRoot = path.dirname(path.dirname(conventional)); const newTarget = relocatedDataDir(mediaRoot, newSlug); + // A ROLLBACK MUST ONLY UNDO WHAT ACTUALLY RAN. The pre-check below fails + // BECAUSE something unrelated already occupies <root>/<newSlug> — and the + // old catch then renamed that stranger to <root>/<oldSlug>, destroying a + // directory this function had never touched, in the name of undoing a move + // it had not made. Each step records that it happened; the catch replays + // only those, in reverse. + let movedMedia = false; + let relinked = false; try { if (await pathExists(path.dirname(newTarget))) { throw new Error(`A media directory already exists at ${path.dirname(newTarget)}`); } await rename(path.dirname(conventional), path.dirname(newTarget)); + movedMedia = true; await unlink(path.join(newChannelDir, "data")).catch(() => {}); await symlink(newTarget, path.join(newChannelDir, "data")); + relinked = true; // Re-read: the channel dir has already moved, so this is the file that // will actually be on disk afterwards. const fresh = (await readChannelConfig(paths, newSlug)) ?? config; @@ -151,9 +161,21 @@ export async function renameChannel( dataDir: newTarget, }); } catch (err) { - await rename(path.dirname(newTarget), path.dirname(conventional)).catch( - () => {}, - ); + // The link goes back too, and before the directory under it moves: a + // failure between the symlink and the config write left `data/` pointing + // at <root>/<newSlug>/data while everything else was rolled back to the + // old slug — a dangling link, which reads as an unmounted drive. + if (relinked) { + await unlink(path.join(newChannelDir, "data")).catch(() => {}); + await symlink(conventional, path.join(newChannelDir, "data")).catch( + () => {}, + ); + } + if (movedMedia) { + await rename(path.dirname(newTarget), path.dirname(conventional)).catch( + () => {}, + ); + } if (hadStore) { await rename(newStoreDir, oldStoreDir).catch(() => {}); } diff --git a/common/jobs/jobKinds.ts b/common/jobs/jobKinds.ts @@ -47,10 +47,16 @@ export type JobKindMeta = { // ABSENT MEANS FALSE, and that is deliberate rather than lazy. The guard is // opt-in so a kind that is not listed keeps exactly its current behavior, and // so that `refresh-report` — which is not in this table at all — still runs - // and lets the snapshot generator itself report the reason it refused. A - // bookkeeping kind (normalize, availability check, store playlist, clear - // markers) never declares it: it must not be refused for a drive it never - // reads. + // and lets the snapshot generator itself report the reason it refused. A kind + // that genuinely never opens a video dir (store playlist, clear markers) must + // not be refused for a drive it does not read. + // + // THE TEST IS "DOES IT OPEN data/", NOT "IS IT BOOKKEEPING". Two kinds that + // read as bookkeeping declare it anyway, and both were misses: + // `normalize-transcripts` walks every video dir and writes a sidecar into + // each, and `sync` reads the data dir to decide what is already downloaded + // before downloading into it. Against an unmounted drive the first reports a + // clean run over zero videos and the second concludes nothing is downloaded. needsMedia?: boolean; }; @@ -190,12 +196,16 @@ const JOB_KINDS: Record<string, JobKindMeta> = { // write, so "let the in-flight one finish" is already how it behaves. Not // replayable either — that would need a JobSpec and a jobReplayRegistry // handler, and the button is one click from the card that reports the count. + // needsMedia: it walks `data/<id>/` for every video and writes a sidecar into + // each. Against an unmounted drive it finds nothing, reports a clean + // 0/0/0/0 run and moves on — a sweep that silently skips a channel. "normalize-transcripts": { kind: "normalize-transcripts", label: "Normalize transcripts", drainable: false, replayable: false, queueKeyStrategy: "custom", + needsMedia: true, }, "redownload-incomplete-bucket": { kind: "redownload-incomplete-bucket", @@ -392,12 +402,17 @@ const JOB_KINDS: Record<string, JobKindMeta> = { replayable: true, queueKeyStrategy: "platform", }, + // needsMedia: syncPaged does not merely refresh a playlist — it reads the + // channel's data dir to decide what is already there and downloads into + // `data/<id>/`. Against an unmounted drive its listing is empty, which means + // "nothing is downloaded", which means download everything. sync: { kind: "sync", label: "Sync", drainable: true, replayable: true, queueKeyStrategy: "platform", + needsMedia: true, }, // Replayable kinds that never had a JOB_KIND_LABELS entry: label omitted so // jobKindLabel() keeps falling back to the raw kind (unchanged behavior). diff --git a/editor/app/channels/[slug]/components/stages/StorageStage.tsx b/editor/app/channels/[slug]/components/stages/StorageStage.tsx @@ -8,6 +8,7 @@ import type { RelocationPreview } from "yt-dlp-transcript-common/controller/relo import { MediaLocationBadge } from "../../../../components/MediaLocationBadge"; import { cancelJobAction } from "../../../../jobs/actions"; import { + clearRelocationMarkerAction, moveChannelMediaBackAction, previewRelocationAction, relocateChannelMediaAction, @@ -49,6 +50,12 @@ type Props = { // Why both buttons are off, or null when they are live. Running/queued jobs // for this channel, or a relocation marker left by an interrupted move. blockedReason: string | null; + // Whether to offer the marker escape hatch. Computed on the server as "a + // marker is present AND no job is running" — a separate answer from + // blockedReason, which the marker itself sets: the two cannot be derived from + // each other, and reading the hatch off `blockedReason === null` would hide + // it in exactly the state it exists for. + canClearMarker: boolean; }; export function StorageStage({ @@ -58,6 +65,7 @@ export function StorageStage({ freeBytes, volumeDir, blockedReason, + canClearMarker, }: Props) { return ( <div className="flex flex-col gap-6"> @@ -111,6 +119,8 @@ export function StorageStage({ </p> )} + {canClearMarker && <StaleMarker key="stale-marker" slug={slug} />} + <MoveOut key="move-out" slug={slug} @@ -309,3 +319,53 @@ function MoveBack({ </section> ); } + +// The way out of a marker whose run is gone. +// +// It is offered ONLY when the channel is in-transition AND no job is running — +// with a live job the marker is not stale, it belongs to that run. Nothing else +// on this panel is available in that state (both moves are disabled while a +// marker stands), so without this a killed copy leaves the channel skipped by +// every lane, refused by every media job and unable to regenerate its snapshot, +// with the only fix being to delete a dotfile over SSH. +// +// It clears the marker and nothing else, which is why the copy says to look +// first: whatever is on disk stays on disk, and the status underneath may well +// be `inconsistent`. That is the honest answer. +function StaleMarker({ slug }: { slug: string }) { + const [busy, setBusy] = useState(false); + const [error, setError] = useState<string | null>(null); + return ( + <section className="flex flex-col gap-2 rounded border border-destructive/50 bg-destructive/5 px-3 py-2"> + <h3 className="text-base font-semibold">Clear the relocation marker</h3> + <p className="text-sm text-muted-foreground"> + A move is in flight, or one was interrupted. While the marker stands this + channel is skipped by every lane and its media jobs are refused. If no + move is actually running, clear the marker — it removes the marker file + and nothing else: no files are moved, copied or deleted. Check what is on + the drive first. + </p> + {error && ( + <p role="alert" className="text-sm text-destructive"> + {error} + </p> + )} + <div> + <button + type="button" + disabled={busy} + onClick={async () => { + setBusy(true); + setError(null); + const result = await clearRelocationMarkerAction(slug); + setBusy(false); + if (!result.ok) setError(result.error); + }} + className="px-3 py-1.5 rounded-md border border-destructive text-destructive text-sm font-medium disabled:opacity-50" + > + {busy ? "Clearing…" : "Clear marker"} + </button> + </div> + </section> + ); +} diff --git a/editor/app/channels/[slug]/page.tsx b/editor/app/channels/[slug]/page.tsx @@ -547,6 +547,9 @@ export default async function ChannelDetailPage({ freeBytes={freeBytes} volumeDir={volumeDir} blockedReason={blockedReason} + // A marker with no job behind it is a stale marker: the run that + // wrote it is gone, and nothing else will ever clear it. + canClearMarker={marker !== null && activeJobs === 0} /> ); } diff --git a/editor/app/channels/[slug]/storageActions.ts b/editor/app/channels/[slug]/storageActions.ts @@ -32,6 +32,7 @@ import { } from "yt-dlp-transcript-common/jobs/streamCommand"; import { requestChannelSnapshot } from "yt-dlp-transcript-common/jobs/snapshotScheduler"; import { formatBytes } from "yt-dlp-transcript-common/lib/format"; +import { clearRelocationMarker } from "yt-dlp-transcript-common/lib/channelMedia"; import { previewRelocation, relocateChannelMedia, @@ -77,7 +78,7 @@ function activeJobsRefusal(slug: string, what: string): string | null { if (active.length === 0) return null; return ( `Finish or cancel ${active.length} running/queued job(s) for this channel ` + - `before ${what} its media.` + `before ${what}.` ); } @@ -87,7 +88,7 @@ export async function relocateChannelMediaAction( ): Promise<StreamActionResult> { const trimmed = root.trim(); if (!trimmed) return { ok: false, error: "Enter a destination root." }; - const refusal = activeJobsRefusal(slug, "moving"); + const refusal = activeJobsRefusal(slug, "moving its media"); if (refusal) return { ok: false, error: refusal }; return runMove(slug, "out", trimmed); } @@ -95,7 +96,7 @@ export async function relocateChannelMediaAction( export async function moveChannelMediaBackAction( slug: string, ): Promise<StreamActionResult> { - const refusal = activeJobsRefusal(slug, "moving back"); + const refusal = activeJobsRefusal(slug, "moving its media back"); if (refusal) return { ok: false, error: refusal }; return runMove(slug, "back"); } @@ -137,3 +138,34 @@ async function runMove( }, }); } + +// THE LAST RESORT, and the only one of the three that is not a move. +// +// A channel carrying a relocation marker is "in-transition" to every guard: the +// lane runners skip it, runManagedFunction refuses its media jobs, and its +// snapshot will not regenerate. That is correct while a move is running and a +// dead end once one is not — a killed process or a replaced container leaves a +// marker nothing will ever clear, and the relocate job itself is refused from +// the panel while the marker stands. +// +// So this removes the marker and NOTHING else: no link, no config, no bytes. +// Whatever inspect() reports afterwards is the truth the disk was already +// telling underneath it — which may well be `inconsistent`, and that is the +// honest answer rather than a repair nobody asked for. It is refused while the +// channel has a live job, because then the marker is not stale: it belongs to +// the run that is holding it. +export async function clearRelocationMarkerAction( + slug: string, +): Promise<{ ok: true } | { ok: false; error: string }> { + const refusal = activeJobsRefusal(slug, "clearing its relocation marker"); + if (refusal) return { ok: false, error: refusal }; + try { + await clearRelocationMarker(getPaths(), slug); + } catch (e) { + return { ok: false, error: (e as Error).message }; + } + revalidatePath(`/channels/${slug}`); + revalidatePath("/channels"); + revalidatePath("/"); + return { ok: true }; +} diff --git a/editor/app/channels/actions.ts b/editor/app/channels/actions.ts @@ -303,7 +303,18 @@ export async function refreshChannelSnapshotAction( if (!(await channelExists(paths, slug))) { return { error: `Channel "${slug}" not found` }; } - await generateChannelSnapshot(paths, slug); + // generateChannelSnapshot now THROWS on a channel whose media is not + // reachable (guard 3) rather than writing a snapshot that says every video is + // undownloaded. That is the right behaviour and the wrong exception to let + // out of a server action: an uncaught throw here reaches the client as a + // digest-only "an error occurred", and the one thing the operator needs is + // the sentence naming the unmounted drive. Every sibling in this file returns + // { error }; so does this. + try { + await generateChannelSnapshot(paths, slug); + } catch (e) { + return { error: (e as Error).message }; + } revalidatePath(`/channels/${slug}`); // The Report column on /channels is read off this snapshot, and the row // action sits next to the marker it flips — so revalidate the list too, not