Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 964979cc3e701d2b7ccf6c28739667bcc412858f
parent 909ba617994953364553651c1d40873566c39c43
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Fri, 11 Sep 2026 11:45:39 -0400

docs: data/ may be a symlink, and three ways that bites

AGENTS.md's corpus table now says `data/` may be an absolute symlink, with the
three rules that are not optional beside it: never symlink a whole channel dir
(the listing filters isDirectory(), so it would vanish), an unmounted drive is
not an empty channel (which is the whole reason channelMedia.ts exists, and
where a new reader of data/ goes to be guarded), and the marker file that makes
a move resumable and that delete and rename refuse under.

RUNNING_IN_DOCKER.md: the corpus is a named volume, so a relocated target must
be bind-mounted at the SAME absolute path the editor recorded. Mounted elsewhere
the link dangles — and the doc says why that is reported as unreachable rather
than fallen back from, because the fallback is every count on the channel
reading zero and the download runner treating the archive as missing.

WORKTREES.md: `rsync -a` copies the link AS a link, which is usually what you
want — both checkouts read the same media through the same absolute path.
`--copy-links` is for a shard going to a machine that will not have the drive,
and it leaves the copy's config.dataDir naming a target it no longer uses, which
reads as `inconsistent` until the field is cleared.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
MAGENTS.md | 26+++++++++++++++++++++++++-
MRUNNING_IN_DOCKER.md | 22++++++++++++++++++++++
MWORKTREES.md | 11+++++++++++
Meditor/CHANGELOG.md | 1+
4 files changed, 59 insertions(+), 1 deletion(-)

diff --git a/AGENTS.md b/AGENTS.md @@ -171,7 +171,7 @@ apps*, not [DEPLOY_DOCKER.md](DEPLOY_DOCKER.md), which is about *building sites* | Path | What it is | |---|---| -| `transcripts/channels/<slug>/` | One channel: `config.json`, `playlist`, `archive`, `snapshot.json`, and `data/<videoId>/` holding media, transcripts and sidecars. | +| `transcripts/channels/<slug>/` | One channel: `config.json`, `playlist`, `archive`, `snapshot.json`, and `data/<videoId>/` holding media, transcripts and sidecars. **`data/` may be an absolute SYMLINK** — see below. | | `transcripts/sites/<id>/site.json` | **Per-site config, including the public URL.** This is where deployed-site facts live — *not* under `channels/`. | | `transcripts/index.mdb` | The LMDB transcript index. Key-only range scans over its `byChannel` sub-DB are cheap; see `common/controller/recencyIndex.ts`. | | `transcripts/saved-videos/` | Persisted source-video store. | @@ -180,6 +180,30 @@ apps*, not [DEPLOY_DOCKER.md](DEPLOY_DOCKER.md), which is about *building sites* **The public-URL key in `site.json` is `siteUrl`.** The editor form labels the field "Public URL", so grepping for `publicUrl` finds the UI hint and misses the data. +## A channel's `data/` may live on another drive + +`channels/<slug>/data` can be an **absolute symlink** to `<root>/<slug>/data` on +another disk, with `config.dataDir` recording the target. The editor's Storage panel +(channel page → Storage) moves it; nothing else writes that field. The on-disk contract +`channelDir/data/<id>/…` is unchanged, so **no reader needs to know** — yt-dlp's +cwd-relative writes, the LMDB index (it stores mtimes, and `rsync -a` preserves them) +and the export build all keep working with no call-site changes. + +Three things that are not optional: + +- **Never symlink a whole channel dir.** Channel listing filters `isDirectory()` on + `channelsDir` entries (`channels.ts:246,363`), so a symlinked `<slug>/` vanishes from + the corpus. Only `data/` may be a link. +- **An unmounted drive is not an empty channel.** Every enumerator swallows ENOENT on + `data/` as "no videos", which to a runner means *everything is undownloaded*. + `common/lib/channelMedia.ts` is the one module that can tell the two apart; + `inspectChannelMedia` / `assertChannelMediaReachable` are what the guards call, and a + job kind declares `needsMedia` in `common/jobs/jobKinds.ts` to be covered by the one + in `runManagedFunction`. If you add a path that reads `data/`, guard it there. +- **`channels/<slug>/.relocating.json`** is the in-flight marker. Its presence means + "media is in transition" to every guard and lets an interrupted move resume from its + `phase`. `deleteChannel` and `renameChannel` refuse while it exists. + ## The live instances Read these out of `transcripts/sites/*/site.json` rather than hardcoding them — this diff --git a/RUNNING_IN_DOCKER.md b/RUNNING_IN_DOCKER.md @@ -377,6 +377,28 @@ docker run --rm -v archilyzer_corpus:/corpus -v "$PWD:/backup" \ debian:bookworm-slim tar czf /backup/corpus.tar.gz -C /corpus . ``` +### A channel whose media is on another drive + +The Storage panel can move a channel's `data/` to another root (see AGENTS.md). In a +container that root is a path **inside the container**, and `channels/<slug>/data` +becomes an absolute symlink to it — so the drive must be bind-mounted **at the same +absolute path the editor recorded**: + +```yaml +services: + editor: + volumes: + - /mnt/platter/archilyzer-media:/mnt/platter/archilyzer-media +``` + +Mount it somewhere else and the link dangles. That is **reported as unreachable, by +design** — the channel is skipped by the lane runners, its media jobs are refused and +its report is not regenerated, rather than the alternative, which is every count on that +channel reading zero and the download runner treating the whole archive as missing. The +badge on `/channels` and the channel's Storage panel name the path they cannot reach. + +The `site` profile does not need the mount: an export build never reads `data/`. + ### Useful commands ```sh diff --git a/WORKTREES.md b/WORKTREES.md @@ -143,6 +143,17 @@ writes a `.worktree-env` file pointing `TRANSCRIPTS_DIR` at the **main** worktre `transcripts/`, and `wt run` loads it. This lets a worktree reuse the already-downloaded corpus for read-mostly work and builds. +### Copying a corpus that has a relocated channel + +`channels/<slug>/data` may be an absolute symlink to another drive (see AGENTS.md). +`rsync -a` copies a symlink **as a symlink**, which is usually right — both checkouts +then read the same media through the same absolute path, and nothing is duplicated. Use +`rsync -a --copy-links` only when the copy has to carry the media itself (a shard bound +for a machine that will not have that drive). Note what `--copy-links` gives you: a real +directory where the source had a link, so the copy's `config.dataDir` still names a +target it is no longer using — `inspectChannelMedia` reports that as `inconsistent`, and +clearing `dataDir` in the copy's `config.json` is what makes it in-place again. + > ⚠️ **Caveat:** the shared LMDB index (`transcripts/index.mdb`) is not safe for concurrent > **writes**. Use shared mode for reading/building, not for running ingestion (downloads / > indexing) in two worktrees at the same time — concurrent writers can corrupt the index. diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,7 @@ # Changelog ## [Unreleased] +- **A channel's media can live on another drive.** A channel page has a **Storage** panel: where its media actually is, how much audio is on disk, how much room is free on the volume holding it, and **Move media to…** — give it a directory on another disk, press *Preview* to see the bytes and the free space there, and the move copies, **verifies**, and only then swaps `data/` for a link to the new location and records it. **Move back in place** reverses it. The source is never touched until the copy has verified, so a cancelled or crashed move leaves everything where it was and the partial copy resumable; re-running finishes it. Nothing else changes: every page, every job, yt-dlp and the search index read the channel exactly as before, because the path they use is unchanged. **The point is what happens when the drive is not mounted.** `data/` reads as empty then, and an empty `data/` means "nothing has been downloaded" to the download runner — an instruction to re-fetch the entire channel onto the disk that was too full to hold it. So an unreachable channel is **refused rather than guessed at**: its media jobs will not start, the four lane runners skip it (and keep running every other channel — this is not a lane stop), its report will not regenerate over an empty directory, and a red **Media unreachable** badge names the path on `/channels`, on the dashboard and on the channel itself. A relocated-and-reachable channel gets a neutral badge saying where; a channel in place gets none. The low-disk floor now measures **the volume the bytes are actually going to** rather than always the corpus disk, and holds each volume separately — a full SSD no longer pauses downloads landing on the platter. The **Media location** line on a channel's Configure form is read-only on purpose: it is a record of what is on disk, written only by a move that succeeded. **Nothing moves on its own, and nothing on disk changes until you move a channel.** - **Every pipeline is dispatched by one thing now: its lane’s runner. The two corpus sweeps and the arbiter are gone.** Digest and Speaker work were driven by a *sweep* — a corpus walk armed by its own switch, with its own scope, its own order and its own console — while Download and Transcription were driven by the auto-queue runner, with rules, a claim ladder, a next-up and a pick log. Two mechanisms, two vocabularies, two sets of bugs. There is one: **each of the four lanes has a runner, a rule list, and Start / Drain / Stop beside its pause**, on the operation’s own page. Arming a corpus pass is switching the lane on; scoping it to particular channels or operations is a *rule*, written the same way auto-transcribe’s have been written since it shipped. The dashboard and the widget keep a one-click switch per lane — **Run every channel** / **Stop the lane** where they said *Sweep every channel* / *Stop sweeping* — and the scope lives on the lane’s page, where you can see what it would do next. **Your armed scope is carried over, and no lane is switched on that was not.** The ten settings fields the sweeps used (`digest.sweepEnabled`, `sweepChannels`, `recencyOrder`, `recencyReach`; `backfill.sweepEnabled`, `sweepKinds`, `sweepChannels`, `order`, `reach`, `weight`) are read once and written into the lane’s rules the first time the editor starts: a sweep armed on three channels becomes three rules, an unscoped one becomes a single *every channel* rule, and a disarmed sweep becomes a switched-off lane. What is retired rather than migrated: **Reach**, because a rule already orders every video it claims across every channel — which rule goes first is the rule list’s job; the digest **order**, whose real meaning was always *newest day first, shortest video within a day* and which the lane spells as **Shortest first** (pick *Newest first* there if you want the date order alone); and the backfill lane’s **Resource share**, which was one number answering two different questions. A lane now stands aside for transcription when it would actually compete for the graphics card, and keeps its slots when it would not — so speaker-naming over an LLM endpoint no longer parks itself behind a transcription it was not competing with. **The arbiter, which never ran a single unit in production, is deleted**; the runner is what dispatches an operation-named rule. **Nothing on disk changes**, and the retired keys are left in `settings.json` — harmless, ignored, and yours to delete. - **The transcode operation is gone — it never fired.** A channel page had a *Transcode* stage, `/operations/transcode` had a "no console here" panel, `/cleanup` offered "Clear failed transcodings", and the video list drew a third status dot — all for a re-encode step built against two failures that never happened in production: in 68 channels, no snapshot has ever listed a video as missing its target format, no `failed-transcodings` file has ever held an id, and only four channels even met the stage's gate. Transcription never needed it — a video whose audio is in another format transcribes from that file. What stayed is everything that was never the operation's: the download path still re-encodes what it extracts itself, the video page still offers **Transcode audio.\<ext\> → \<fmt\>** per file, and both audio-format sweeps on the Cleanup stage and `/cleanup` are unchanged (gated on the channel having an `audioFormat`, which is what they compare against). The snapshot bucket behind the sweep is `wrongFormatAudio` now — its operator-facing name — and old reports keep their stray key until their next refresh. A `?stage=transcode` bookmark opens the channel overview. **Nothing on disk changes.** Also: the Pool's running-jobs list names the eight kinds its buttons enqueue, and the site's Search aliases tab no longer carries a "no site selected" branch that could not run. - **A site has tabs, and the family has one page.** Charts, Search aliases, Deploy, Build and Homepage were five sidebar entries beside *Sites*, three of them reading the site from a `?site=` parameter the sidebar picker had to seed, one of them (Build) about no site at all, and one (Homepage) about the family's own hub. A site is one thing now: **`/sites/<id>` is Settings · Charts · Search aliases · Publish**, the site named in the path, the picker following it (and Dashboard and Channels following the picker). **`/sites` is the family page**: the list, then *Release notes*, *Build all sites* with the Basic/Docker mode, the *Hub*, and the *Pool* — the corpus-wide index, stats, sidecar and archive jobs — folded under a disclosure. Search aliases keep both sections on the site's tab: the global dictionary and the site's overrides. Every button, label and log is unchanged; "Select a specific site from the sidebar" is gone because a site's page always has one. The five routes redirect — a `?site=<id>` bookmark lands on that site's tab (the query rides along), `?site=__all__` and the bare routes on `/sites`; a bookmark to a deleted site 404s there exactly as `/sites/<id>` does. The Sites group is one entry; the nav is **eleven**, the IA doc's end state. **Nothing on disk changes.**