Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 7544c6c10d4fa873579b4484f7d21e4c4060fee0
parent 5215b918d701fd00c025feee84149853dad83771
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Sat, 22 Aug 2026 19:24:37 -0400

changelog: restore the armed-cost entry

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Diffstat:
Meditor/CHANGELOG.md | 2+-
1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -3,7 +3,7 @@ ## [Unreleased] - **One word was standing in for three different things, and hiding a fourth.** "Backfill" named a *lane* (three operations sharing a CPU queue), a *per-channel action*, and — on the download stage — a completely unrelated **download job** for subtitle tracks. Nobody arms, pauses or runs "a backfill"; it is a queue key. Every screen where an operator reads a **figure** now reads an **operation name**: the channel transit line's station is *Speakers*, its stage panel is *Speaker work · 3 operations* with a *Run speaker work* button, the download stage says *Fetch missing subtitle tracks*, and the lane card and `/auto-queue` name their members. The lane name survives in exactly one place — the settings fieldset and the lane card, where the one shared pause and the one shared sweep live — and there it now **lists what it holds**, because "these operations share a queue and a pause" is a true and load-bearing fact rather than a leaked implementation detail. - **The transit line's Speakers station stopped summing three operations into one number.** It set its coverage by adding every operation on the lane together — the one thing this codebase forbids everywhere else — which on the live corpus meant adding diarization (one audio pass per video, 4 done of 11,338) to speaker-names-from-the-transcript (**~1 model call per transcript chunk**, 1 done of 11,338) and printing the total under a label that named the queue. The numeral now belongs to exactly one operation, and each member states itself, with its own band and its own denominator, in the station foot. - +- **An armed operation now says what one unit of it costs.** The fact that hid: speaker-names-from-the-transcript is switched **on**, can reach 11,337 videos of one channel — on the order of **194,000 model calls corpus-wide** — has completed one video, and read as a quiet row on every screen, because "11,337 reachable" is the same shape of number whether the unit is an audio pass or a per-chunk model call. Each operation declares its cost basis, printed beside its backlog wherever it is armed. Deliberately **no threshold and no editorialising**: what is affordable is the operator's call, and a "this is a lot" cutoff would be a magic number the next operation gets wrong. Nothing here changes a setting. - **Every pipeline is now visible on /channels, not just two of them.** The table printed two bare integers — `Downloads` and `Transcripts` — and said nothing whatsoever about the four derived pipelines, which were collapsed everywhere else behind the word *backfill*. All six now draw the same **state band** the comparison rail on /auto-queue uses, at table scale: one fill per population, `can run now` the only saturated colour on the page, and pattern (solid / hatched / dotted / hollow) carrying the meaning ahead of hue, because no four-colour palette clears all-pairs colour-blindness. **Not a percent bar, deliberately** — digest sits at 0 done on every large channel and diarization is 99.96% media-gone, so "% complete" renders `0%` on all 68 rows and says nothing; what varies, and what an operator needs, is the *shape* of the remainder. Every pipeline column **sorts by what can run now**, which answers a question the page has never been able to answer: *which channel has the most diarizable audio left right now* previously meant opening 68 channel pages one at a time. The costs nothing extra to draw — `/channels` was already reading every channel's full snapshot for its counts and throwing the rest away. - **`docker compose up -d` now stands up a working archive.** The repo had two Dockerfiles and neither ran the app: one fans per-site export builds out across containers, the other runs sharded e2e. So the only way to host this was to install the whole Unix toolchain by hand, which is why the Windows instructions said "use WSL2 and follow the Linux steps". There is now a runtime image and a compose stack — the editor plus Caddy by default, with the published site, the project homepage and umtool behind compose **profiles**, so somebody who only wants an archive runs two containers rather than five. First boot creates the volumes, downloads a speech model, and seeds a `settings.json` **carrying one enabled worker**: the defaults ship `workers: []`, zero workers means zero transcription slots, and a fresh container that looks healthy and silently transcribes nothing is the worst possible first run. Two things are deliberately not baked into the image and cannot be: the corpus, and the export site — that site is a static render *of* a corpus, and there is no corpus at image-build time, so `docker/publish-site.sh` builds it at run time into the volume the `site` service serves.