Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit c3963bd69359a69891f185619ff5fe594b0d67dc
parent 9d6c8e32974ba0aaff93221347fa31b71b0a3236
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon, 28 Sep 2026 12:05:05 -0400

changelog: umtool's path defaults in [Unreleased]; a released entry's path example says <user>

One [Unreleased] bullet for slice Q. The storage-locations entry's
`/run/media/user/<uuid>` becomes `/run/media/<user>/<uuid>`: a released
entry is normally left as written, but this is a path example, not a
name a reader searches for.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 3++-
1 file changed, 2 insertions(+), 1 deletion(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -9,6 +9,7 @@ - **`archilyzer` checks the machine, runs one operation offline, starts the MCP server, and is one command from the repo root.** `pnpm archilyzer <command>` is the short form (`pnpm archilyzer --help` lists them all). `pnpm archilyzer doctor` is a read-only report: Node, the checkout, the corpus and whether each channel's media is reachable, `settings.json`, every tool the paths name plus each enabled worker's engine and model, umtool's report-pipeline tools, and this checkout's ports; it exits 1 only for something the machine is set up to do and cannot. `pnpm archilyzer run <operation> <channel> [ids…]` runs diarization, either attribution pass or digest over one channel as the editor's job does (a job record and log under `.jobs/`, the same summary line, the same refusal for an unmounted drive); sync, the metadata scan, downloads and transcription are refused with the reason, because they run on the editor's paced download queue and worker pool. `pnpm archilyzer mcp` starts the MCP server, so it can be registered as `-- pnpm -C "$PWD" archilyzer mcp`. Every other script in `common/bin/` is a subcommand too (`duplicates`, `posts fetch`, `digest plan`, `verify transcripts`, …), and export's `detect:duplicates` script is now `archilyzer duplicates`. Every environment variable is listed, by audience, in the new `ENVIRONMENT.md`, and `DEPLOY_CLOUDFLARE.md` and `DEPLOY_DOCKER.md` are now one `PUBLISH.md`. - **The build-mode toggle and setting are gone, and the test-only environment variables start with `E2E_`.** The Basic / Docker choice on `/sites` and in **Settings → Build pipeline** was a label nothing read, so it is removed; a `settings.json` that still has `buildPipeline.mode` loads as before and drops the key on the next save. **Build all sites** builds every site in parallel in containers whenever a container engine answers, and serially on the host when none does, as it always did; a single site's build runs one at a time. Max parallel builds, the image tag and the Dockerfile path stay. For whoever runs the test suites: every variable only a test harness reads is renamed with an `E2E_` prefix and listed in its package's `playwright.config.ts` (and in `ENVIRONMENT.md`); `SHARDS=N pnpm e2e:sharded` is now `E2E_SHARDS=N`. - **In high-contrast mode the sidebar's Archilyzer mark keeps its edge.** In Windows' high-contrast mode (forced colours) the reader's own background replaces the page on every ground and can be as dark as the mark's slate tile, whose thin ring is only drawn on Dark. In that mode the tile now gets a 1-pixel outline in the reader's text colour, on every ground, following its rounded corners. Nothing changes outside that mode. +- **umtool finds its data without one machine's home directory in the code.** With no `CHANNELS_DIR` it reads the corpus at `$TRANSCRIPTS_DIR/channels`, else the checkout's own `transcripts/channels`; it used to fall back to an absolute path that existed on one machine only. The song project's bulk data now defaults to `~/.local/share/archilyzer/song` and its videos to `~/reports/quartering-uh-song/videos`; `SONG_DIR` and `VIDEO_ROOT` still win. If the song data lives elsewhere, link it there **before restarting umtool** (`ln -s <where it is> ~/.local/share/archilyzer/song`), which keeps it where it is. The song project's tracked manifests record their paths relative to the song folders, and the twenty one-off `umtool/song/*.sh` run logs, which only ever ran on the machine that wrote them, are gone. ## [0.9.4] - 2026-09-28 - **On the Dark ground the sidebar's Archilyzer mark has a thin outline.** Its slate tile now has a 1-pixel ring just outside it, following its rounded corners, in the colour of the mark's unlit lines, so the tile's edge shows against the dark page. Light and Sepia are unchanged, and so is the favicon. @@ -80,7 +81,7 @@ - **A channel's Storage panel has one verb.** *Move media to…* and *Move back in place* were two sections with two buttons whose availability was the inverse of each other — one decision split across two controls. The corpus volume is now a destination in the same select; picking it moves the media back. While the media is on a location it is the only destination offered, because a move straight from one location to another is refused by the mover itself. - **The saved-video store can be moved to another drive.** It is the one large thing in the corpus that belongs to no channel, so no channel move could ever reach it. `/storage` now shows it with its size and where it is, and moves it onto a location — and back — by exactly the mechanism a channel uses: the copy is verified before the source is touched, a symlink is left behind, and every reader keeps working unchanged. An interrupted move can be **resumed** rather than restarted, and its marker cleared if it cannot. - **A drive that disappears now disables and flags the channels on it, and un-does that when it comes back.** Every guard in the app refuses work on unreachable media at the moment the work starts; none of them was a *detector*, so a channel whose drive fell off a cable sat there being refused with nothing anywhere saying why. A five-minute pass now notices, **pauses those channels and marks them** — the tier column shows a *storage* badge with the reason — and restores each one to the tier it had when the drive returns. It never claims a channel you paused yourself, and changing a tier by hand permanently takes it out of the machine's hands. It writes only when something has actually changed, and never arms at all under `ARCHILYZER_IDLE_BOOT`. -- **The drives a channel's media lives on are named places now, and one click re-points them.** The cold root used to be a single string typed into Settings, and a relocated channel's `config.dataDir` an absolute path — so when the platter was automounted at `/run/media/user/<uuid>` and came back somewhere else, every channel on it read *unreachable* and the only remedy was SSH and hand edits. **`/storage`** (twelfth entry, under Machine) lists each media root as a **location** with a name, a status and the channels on it: `Available`, `Not mounted`, `Not attached`, or **`Mounted elsewhere`** — which is the one that matters, because it means the disk is here under a different mountpoint, and the row then offers **Re-point**, which rewrites every channel's `data/` symlink and `config.dataDir` and **moves no bytes at all**. There is a **Refresh** per row (a probe is `findmnt`, memoised for ten seconds, and it never writes availability to `settings.json`), a **Mount** for an attached-but-unmounted volume, a per-location **auto re-point** opt-in for operators who would rather it just happened, and a boot pass that checks every location as the editor starts. Identity is the volume's filesystem **UUID**, learned at the last successful probe — in a container there are no block devices to learn it from, every probe **fails open to "unknown", and re-point by path is the whole story** (see RUNNING_IN_DOCKER.md). The old `settings.storage.mediaRoot` **migrates on read** into a one-entry list called *Default*; the Settings field is now a link to the page. Everywhere a move starts, the destination is a **name picked from a list** rather than a path retyped per channel: the channel's Storage panel has a **destination select** showing each drive's current state (with *Another root…* keeping the free-text box), and `/channels`' selection deck has the same select for a whole batch — and what reaches the server is the **id**, never the root, so a page rendered before a re-point cannot aim a batch at a root that has since moved. The badge on every channel row says **`on Platter`** instead of sixty columns of absolute path, or **`on Platter — unreachable`** when the drive is not there. **And an interrupted move can be finished.** The controller has always resumed a half-done copy; nothing in the editor could reach it, so the only offered way out of a killed rsync was *Clear marker* and a full re-copy — which for the incident behind this work meant re-copying 131 GB that was already correctly on the far side. The panel now offers **Resume move** beside it, and the same release closes the three ways that incident happened: the relocate copy raced an auto-queue digest unit that made **no job record**, so "is this channel busy" now asks the lanes as well as the registry, and a lane that finds a relocation marker stops instead of writing into a directory being copied; a sidecar written mid-copy left one directory timestamp differing and `verifyCopy` refused the whole 131 GB, so a drift that is *only* directory mtimes now gets one more `rsync -a` pass and a re-verify (content drift still refuses, and still says the source has not been touched). Also fixed: on `/channels` a dimmed row's **Advanced menu drew underneath the rows below it** — `opacity` on a `<tr>` makes a stacking context, so the row is dimmed cell by cell now, and the cell hosting the popover is left alone. +- **The drives a channel's media lives on are named places now, and one click re-points them.** The cold root used to be a single string typed into Settings, and a relocated channel's `config.dataDir` an absolute path — so when the platter was automounted at `/run/media/<user>/<uuid>` and came back somewhere else, every channel on it read *unreachable* and the only remedy was SSH and hand edits. **`/storage`** (twelfth entry, under Machine) lists each media root as a **location** with a name, a status and the channels on it: `Available`, `Not mounted`, `Not attached`, or **`Mounted elsewhere`** — which is the one that matters, because it means the disk is here under a different mountpoint, and the row then offers **Re-point**, which rewrites every channel's `data/` symlink and `config.dataDir` and **moves no bytes at all**. There is a **Refresh** per row (a probe is `findmnt`, memoised for ten seconds, and it never writes availability to `settings.json`), a **Mount** for an attached-but-unmounted volume, a per-location **auto re-point** opt-in for operators who would rather it just happened, and a boot pass that checks every location as the editor starts. Identity is the volume's filesystem **UUID**, learned at the last successful probe — in a container there are no block devices to learn it from, every probe **fails open to "unknown", and re-point by path is the whole story** (see RUNNING_IN_DOCKER.md). The old `settings.storage.mediaRoot` **migrates on read** into a one-entry list called *Default*; the Settings field is now a link to the page. Everywhere a move starts, the destination is a **name picked from a list** rather than a path retyped per channel: the channel's Storage panel has a **destination select** showing each drive's current state (with *Another root…* keeping the free-text box), and `/channels`' selection deck has the same select for a whole batch — and what reaches the server is the **id**, never the root, so a page rendered before a re-point cannot aim a batch at a root that has since moved. The badge on every channel row says **`on Platter`** instead of sixty columns of absolute path, or **`on Platter — unreachable`** when the drive is not there. **And an interrupted move can be finished.** The controller has always resumed a half-done copy; nothing in the editor could reach it, so the only offered way out of a killed rsync was *Clear marker* and a full re-copy — which for the incident behind this work meant re-copying 131 GB that was already correctly on the far side. The panel now offers **Resume move** beside it, and the same release closes the three ways that incident happened: the relocate copy raced an auto-queue digest unit that made **no job record**, so "is this channel busy" now asks the lanes as well as the registry, and a lane that finds a relocation marker stops instead of writing into a directory being copied; a sidecar written mid-copy left one directory timestamp differing and `verifyCopy` refused the whole 131 GB, so a drift that is *only* directory mtimes now gets one more `rsync -a` pass and a re-verify (content drift still refuses, and still says the source has not been touched). Also fixed: on `/channels` a dimmed row's **Advanced menu drew underneath the rows below it** — `opacity` on a `<tr>` makes a stacking context, so the row is dimmed cell by cell now, and the cell hosting the popover is left alone. - **The operations poll reads the auto-queue's state file once instead of four times.** Every payload the editor draws — the jobs head, the workers grid, the operations board, the sync schedule, the widget's tiles, the pulse token — used to be computed by a function that did its own reading, so each of the four lanes on `/operations` opened `.auto-queue/state.json` for itself: four parses of the same document every three seconds, on a page whose four lanes were always reading one document. Those builders are pure functions in the shared core now — they are handed the settings, the registry, the scheduler, the pool, the clock and their readings, and they cannot reach disk or construct a singleton, which a layer test enforces rather than a comment asking nicely. The reading happens once, at the edge, and is shared. **Nothing moved that you can see**: same pages, same URLs, same JSON on every endpoint, same numbers — the difference is that each payload now has unit tests of its own (the console's cooldown filter, the pulse token's sensitivity, the workers grid's task grouping), where previously the only way to test one was to render the page that showed it. - **`/channels` is a rack now, with one selection deck and a meter bridge.** The page had two selection bars for one selection — a floating one for Tier and Focus, and a second block below sixty-seven rows for Move media, both saying "N selected" and both offering Clear. There is **one deck**: it docks under the table when you tick a row, carries **Tier**, **Focus** and **Media** side by side, and unmounts when you untick. The destination root lives in its own box beside the button (the button used to carry it in its label, where it truncated to *Move media to…* and you could not read where the files were going). **The table stops spilling off the screen.** It lives in one scroll region: the column headers pin to its top, the checkbox and slug cells pin to its left, and a section's name pins under the headers — so the identity column and the meter bridge header stay on screen while sixteen columns scroll sideways. The six pipeline columns read as **one block** rather than six loose dashes: a shared *Pipeline* eyebrow, a surface behind them, a rule at each end. **Rows are 41 px instead of ~90.** The tier cell is one line, and being held by a focus is a small **held** chip rather than the same orange sentence repeated on sixty-one rows — the sentence is stated once, with a count, on the focus line above the table, and each chip still carries the full reason for a screen reader and on hover. Opening a row's *Advanced* overlays the rows below instead of pushing them down. **The page's caveat is at the top.** The note saying every number here is read from each channel's last report, and how old the oldest one is, used to be the last thing on the page in 11 px type under a floating bar; it is the subtitle beside the title now, with the channel count. The band legend and the Names·A / Names·T explainer moved above the table too, beside *Group by section*. In the header, *Sync every channel*, *Full sweep every channel* and *Update all reports* are outlines under an **Every channel** eyebrow that says what they sweep, and **New channel** is the only filled button. Nothing on disk moved and no control changed its name. - **A channel's media can live on another drive.** A channel page has a **Storage** panel: where its media actually is, how much audio is on disk, how much room is free on the volume holding it, and **Move media to…** — give it a directory on another disk, press *Preview* to see the bytes and the free space there, and the move copies, **verifies**, and only then swaps `data/` for a link to the new location and records it. **Move back in place** reverses it. The source is never touched until the copy has verified, so a cancelled or crashed move leaves everything where it was and the partial copy resumable; re-running finishes it. Nothing else changes: every page, every job, yt-dlp and the search index read the channel exactly as before, because the path they use is unchanged. **The point is what happens when the drive is not mounted.** `data/` reads as empty then, and an empty `data/` means "nothing has been downloaded" to the download runner — an instruction to re-fetch the entire channel onto the disk that was too full to hold it. So an unreachable channel is **refused rather than guessed at**: its media jobs will not start, the four lane runners skip it (and keep running every other channel — this is not a lane stop), its report will not regenerate over an empty directory, and a red **Media unreachable** badge names the path on `/channels`, on the dashboard and on the channel itself. A relocated-and-reachable channel gets a neutral badge saying where; a channel in place gets none. The low-disk floor now measures **the volume the bytes are actually going to** rather than always the corpus disk, and holds each volume separately — a full SSD no longer pauses downloads landing on the platter. The **Media location** line on a channel's Configure form is read-only on purpose: it is a record of what is on disk, written only by a move that succeeded. The cold drive is typed **once**: **Settings → Default media root** seeds the root box in every channel's Storage panel, and `/channels` rows can now be ticked — select several and **Move media to…** queues one job per channel on that channel's own queue, so they serialize instead of fanning out, each one running its own space check at run time rather than at enqueue time (a root that fills partway through refuses the remainder cleanly, and a channel already on that root is skipped rather than failed). The default is a default and nothing more: it is never read by the move itself, which always takes an explicit root, and a relocated channel is not thereby deprioritized. **Nothing moves on its own, and nothing on disk changes until you move a channel.** *(Superseded above: **Settings → Default media root** is gone — the roots are named locations on `/storage` now, and every destination is picked from that list by name rather than typed.)*