# Running an archive in Docker One command stands up a working archive server: the editor, the tools it drives (yt-dlp, ffmpeg, and a transcription engine — whisper.cpp, or parakeet.cpp in the Vulkan image), and a reverse proxy that is the only thing on the box with an open port. > This document is about **running the apps** — publishing included: every publish > stage (index, build, deploy) runs inside the editor's container. Fanning per-site > *export builds* out across containers is a different thing — the opt-in docker > runner of a HOST install, from `Dockerfile.build`, in > [PUBLISH.md](PUBLISH.md#building-every-site-in-containers). Neither affects the > other. --- ## Quickstart ```sh git clone archilyzer cd archilyzer cp .env.example .env docker compose up -d --build ``` Then open **http://localhost:8081**. The first build takes a while — it compiles whisper.cpp and builds two Next.js apps. Afterwards, `docker compose up -d` is seconds. On first boot the container also: - creates the corpus, config, models and builds volumes; - writes a `settings.json` **with one enabled worker for the image's engine** (whisper.cpp here, parakeet.cpp in the Vulkan image). With no `workers` key, or no file, `getSettings()` would synthesize `parallelTranscriptions` (default 2) enabled whisper.cpp workers — two CPU slots, and never parakeet; a file that says `"workers": []` has zero slots, and auto-transcribe would report `no-workers` and quietly do nothing; - downloads the `base.en` whisper model (~142 MB) into the models volume. Every key that `settings.json` can carry, with its default, is in [SETTINGS.md](SETTINGS.md). Watch it happen with `docker compose logs -f editor`. ### What is running | | | | |---|---|---| | **editor** | http://localhost:8081 | the admin app: channels, downloads, transcription, builds | | site | http://localhost:8080 | the published archive (profile `site`) | | homepage | http://localhost:8082 | the project's own site (profile `homepage`) | | umtool | http://localhost:8083 | the clip/report bench (profile `umtool`) | Only the editor and Caddy run by default. Somebody who wants an archive runs two containers, not five: ```sh docker compose --profile site up -d # + the archive server docker compose --profile umtool up -d # + umtool ``` --- ## The exposure model **Read this before you change a bind address.** The editor has **no authentication of any kind**. It shells out to yt-dlp, deletes media, streams job logs, and rewrites your corpus. It was built to run on a machine you are sitting at. So the stack is arranged to make exposing it a decision rather than an accident: 1. **No application container publishes a port.** The editor, the site server, the homepage and umtool sit on an internal Docker network. Nothing on your host can reach them directly. This removes the whole class of "I published 3001 once to try something and forgot". 2. **Caddy is the single front door**, and every port it publishes binds `127.0.0.1` **by default — the public sites included**. Nothing is reachable off the box until you say so. 3. **A startup rail.** If a private app is bound to anything but loopback while nothing is checking credentials, the containers refuse to start and tell you the four ways to fix it. | App | Internal port | Caddy port | Default bind | Auth | |---|---|---|---|---| | editor | 3001 | 8081 | `127.0.0.1` | `basic_auth` | | export site | 3000 | 8080 | `127.0.0.1` | none — public by design | | homepage | 3031 | 8082 | `127.0.0.1` | none | | umtool | 3050 | 8083 | `127.0.0.1` | `basic_auth` | ### Serving your archive to the world The archive is static files with no write path. Publishing it is the point: ```sh # .env SITE_BIND=0.0.0.0 ``` ```sh docker compose --profile site up -d ``` Nothing else becomes reachable. The editor stays on loopback. ### Reaching the editor from another machine Pick one. In rough order of how much you should like it: **Tailscale — exposes nothing.** Install Tailscale on the host, leave every bind at `127.0.0.1`, and run `tailscale serve 8081`. The editor is reachable from your own devices and from nowhere else, with no port open on any router and no password to leak. This is the recommended answer for a machine at home. **An SSH tunnel** — the same idea with no new software: ```sh ssh -N -L 8081:127.0.0.1:8081 you@archive-box ``` **A password.** Caddy's `basic_auth` is built in, so this works with nothing installed and cannot break on your first evening: ```sh docker run --rm caddy:2.11-alpine caddy hash-password --plaintext 'your-password' ``` ```sh # .env — the $ characters need no escaping here EDITOR_BIND=0.0.0.0 ARCHILYZER_AUTH_USER=you ARCHILYZER_AUTH_HASH=$2a$14$... ``` Basic auth is a password in a browser dialog sent on every request. It is real protection over a trusted network and nothing more; put HTTPS in front before it crosses one you don't trust. **A real identity provider** — sessions, TOTP, SSO. Two drop-in overlays ship here, and neither is required: ```sh docker compose -f docker-compose.yml -f docker-compose.tinyauth.yml up -d docker compose -f docker-compose.yml -f docker-compose.authelia.yml up -d ``` *Tinyauth* is configured entirely by environment variables, is under 10 MB, and has been OpenID Certified since v5.1.0. It is AGPL-3.0, and its config keys churn between major releases — the overlay pins the tag for exactly that reason, so read the release notes before moving it. *Authelia* is Apache-2.0 and the established forward-auth standard, at the cost of a YAML config file and a mandatory database. Templates are in `docker/authelia/`; copy the two `.example` files and fill them in. Both attach through Caddy's `forward_auth`. The overlays set `ARCHILYZER_AUTH_MODE=forward` for you. ### The rail, and its escape hatch ``` EDITOR_BIND=0.0.0.0 exposes the editor beyond this machine, and nothing is checking credentials in front of it. ``` If you genuinely have your own auth in front — a reverse proxy you trust, a tailnet-only address — say so explicitly: ```sh ARCHILYZER_AUTH_MODE=none ``` That is the only thing that disables the check. It is deliberately a line you have to write. --- ## Everyday operation ### Add a channel and get transcripts Everything happens in the editor at http://localhost:8081 — this is the same app a host install runs, so [README.md](README.md) and [SETUP.md](SETUP.md) describe it accurately. 1. **Channels → Add**, paste a channel URL. 2. **Store playlist**, then **Sync**. Media and any published captions land in the `corpus` volume. 3. **Transcribe** for videos with no captions. This uses the seeded whisper.cpp worker and the model in the `models` volume. ### Publish the archive The site server serves a *static build of your corpus*, which does not exist until you make one. Publishing is a set of **stages** — update the index, build a site, deploy it — and every one of them runs **inside the editor's container**: the queued jobs the editor starts (Sites → Publish, the publish lane) and the commands below alike. The quick way, for the local `site` service: ```sh docker compose exec editor /repo/docker/publish-site.sh docker compose --profile site up -d ``` Create the site itself in the editor first (**Sites → New**); run the script with no argument to list the ones you have. The script is three stages in a row, and you can run them yourself: ```sh docker compose exec editor pnpm archilyzer publish index docker compose exec editor pnpm archilyzer publish build docker compose exec editor pnpm archilyzer publish deploy --to local docker compose exec editor pnpm archilyzer publish status ``` A stage that is already fresh does nothing, so running one twice is cheap. A local deploy copies the built bundle into the volume the `site` service serves; no restart needed afterwards. A **private** site (`audience: "private"`) is built but never deployed, locally or anywhere else. **`exec`, never `run --rm`.** `docker compose run --rm editor …` starts a SECOND container with its own copy of the image's `export/public` and a second writer on the index, and the publish lock (which keeps a stage you start from colliding with one the editor is running) cannot see across containers — worse, the second container carries the editor's host identity below with its own pids, so it would judge the editor's live lock dead and take it. `exec` runs in the editor's own container, beside its jobs, under the same lock. `pnpm ops publish` from the host goes through the editor too. The lock names its holder by host and pid. A container's hostname changes every time it is recreated, so compose gives the editor a fixed identity (`ARCHILYZER_HOST_ID=archilyzer-editor`): a lock left by a stage that died with the container is then recognised as this editor's own, and taken over once its pid is gone. A lock naming any OTHER host is waited on, never taken. If one is left behind — say, from before this setting, or by a host install sharing the volume that is gone for good — and you are sure nothing is publishing, delete it: `docker compose exec editor rm /data/builds/.export-builds/.publish.lock`. #### Deploying to Cloudflare from the container The same stages deploy to Cloudflare Pages. The container has no browser for `wrangler login`, so it authenticates with a token from `.env`, which the editor reads and every stage it runs inherits: ```sh # .env CLOUDFLARE_API_TOKEN=... # an API token with "Cloudflare Pages: Edit" CLOUDFLARE_ACCOUNT_ID=... # only when Settings → archive storage names an R2 bucket for oversize archives: R2_ACCESS_KEY_ID=... R2_SECRET_ACCESS_KEY=... ``` ```sh docker compose up -d # picks up .env docker compose exec editor pnpm archilyzer publish deploy --preview docker compose exec editor pnpm archilyzer publish deploy # production ``` Nothing is fetched at deploy time: wrangler is pinned in the workspace the image ships. With no token a deploy is refused before wrangler runs, with a sentence saying so; `archilyzer doctor` reports whether each credential is **set** (never its value). Every deploy is checked live afterwards. PUBLISH.md has the whole publish flow; this is only what is different in a container. #### The homepage, and its /source mirror The image carries a build of the project's own homepage, served by the `homepage` service. A real one — the corpus's numbers, the `/source` mirror of this repo — is a publish like any other: ```sh docker compose exec editor pnpm archilyzer publish homepage --deploy --to local docker compose --profile homepage restart homepage # once, after the first ``` The `homepage` service serves the local deploy (`/data/builds/homepage`) once there is one, else the baked build — chosen at boot, hence the one restart. The `/source` mirror needs two things the image deliberately lacks: - **The repository.** The image has no `.git`. `docker-compose.source.yml` mounts the host's git directory read-only at `/data/source.git` and points `ARCHILYZER_SOURCE_REPO` at it (default `./.git`; from a git worktree set `ARCHILYZER_SOURCE_HOST_DIR` to the primary checkout's `.git`). Without it the homepage builds with an empty `/source` page. - **The operator's scrub rules and denylist.** They live in the config volume — `ARCHILYZER_CONFIG_DIR` is `/data/config/archilyzer` in the container, made (mode 700, empty) on the first boot — and are never printed by anything: ```sh docker compose -f docker-compose.yml -f docker-compose.source.yml up -d docker compose cp ~/.config/archilyzer/source-scrub.txt editor:/data/config/archilyzer/ docker compose cp ~/.config/archilyzer/source-denylist.txt editor:/data/config/archilyzer/ docker compose exec editor chmod 600 /data/config/archilyzer/source-scrub.txt /data/config/archilyzer/source-denylist.txt ``` `git-filter-repo` (pinned) is in the image. **gitleaks and stagit are not**, so a source publish from the container is weaker than a host's in two ways it tells you about: it skips the secret scan, with a `WARNING` in its log (the literal audit against your denylist still runs, and still refuses), and it publishes the source without its history pages (`/source/git/`). `archilyzer doctor` lists both as absent. Publish the mirror from a host checkout that has them if you want either; a pinned gitleaks in the image is a planned follow-up. ### Model choice `ARCHILYZER_FETCH_MODEL` picks what gets downloaded on boot. For the default and CUDA images that is a whisper model: `tiny.en` (~75 MB), `base.en` (~142 MB), `small.en` (~466 MB), `medium.en` (~1.5 GB), `large-v3` (~3 GB). Bigger is better and slower, and on CPU the difference is large. Changing it later fetches the new model but does **not** switch the worker over — model choice lives on the editor's Workers page, which is the right place for it. For the Vulkan image it names a parakeet GGUF instead — see [GPU transcription](#gpu-transcription). `ARCHILYZER_FETCH_MODEL=none` skips the download entirely. ### Keeping yt-dlp current An image pins yt-dlp at build time, and a stale yt-dlp is the most common reason downloads suddenly start failing. Set `YTDLP_AUTO_UPDATE=1` in `.env` and each boot self-updates it. Exactly `1`, `true`, `yes` or `on` turns it on, as written — `TRUE` or `Yes` does not (the entrypoint and `archilyzer doctor` read it the same way). Or, once: ```sh docker compose exec editor yt-dlp -U ``` Every editor boot logs which yt-dlp it will run, and whether it runs: ``` [entrypoint] yt-dlp: /usr/local/bin/yt-dlp 2026.09.30 (image) ``` ### Substituting yt-dlp To run your own yt-dlp — a patched build, say — set `YTDLP_BIN`. It is a swap at run time: nothing is rebuilt, and the image's release binary stays where it is (`ARCHILYZER_IMAGE_YTDLP`, `/usr/local/bin/yt-dlp`). Two ways: **A source checkout**, run with the image's python (yt-dlp's optional modules are installed beside it): ```sh # .env YTDLP_SOURCE_HOST_DIR=/home/you/yt-dlp-patched # the directory holding yt_dlp/ ``` ```sh docker compose -f docker-compose.yml -f docker-compose.ytdlp.yml up -d ``` The overlay mounts the checkout read-only at `/opt/yt-dlp-src` and sets `YTDLP_BIN=/usr/local/bin/yt-dlp-from-source`, a wrapper baked into the image. **A zipapp** built on the host (`make yt-dlp` in the checkout), copied into the config volume: ```sh docker compose exec editor mkdir -p /data/config/bin docker compose cp ./yt-dlp editor:/data/config/bin/yt-dlp docker compose exec editor chmod 755 /data/config/bin/yt-dlp # .env YTDLP_BIN=/data/config/bin/yt-dlp ``` Either way the boot line says `(override)`, and `YTDLP_AUTO_UPDATE` leaves an override alone with a warning — update it where you build it. A substitute that does not run is reported as `yt-dlp: MISSING — … does not run: …` (the editor still starts), and by `archilyzer doctor`. ### Driving the editor without a browser Every editor gesture is a server action, which is fine for a person and hostile to a script: there is no URL to POST to. `/api/ops/*` is a thin layer over the **same actions** — one route per gesture, no rule of its own — so a shell, a cron job or an agent can run the archive without Playwright. It is gated by the **same `WORKER_TOKEN`** as `/api/worker/*`, deliberately: that variable already means "this instance takes instructions from something that is not the browser in front of it". Unset on the server and every route answers **503** (the surface is off until you opt in); wrong or missing on the caller and it answers **401**. ```sh export ARCHILYZER_EDITOR_URL=http://localhost:3001 export WORKER_TOKEN= pnpm ops sync --json '{"slug":"the-quartering"}' --wait pnpm ops metadata-scan --json '{"slug":"the-quartering"}' pnpm ops refresh-metadata --json '{"slug":"the-quartering","id":""}' --wait pnpm ops channel-config --json '{"slug":"the-quartering","patch":{"downloadFilterExclude":"rerun"}}' pnpm ops channel-config --json '{"slug":"example-x","sites":[],"excludeFromBuild":true}' pnpm ops create-channel --json '{"fields":{"name":"Example (X)","handling":"transcribe","url":"https://x.com/example"}}' pnpm ops rename-channel --json '{"slug":"exmaple-x","newSlug":"example-x"}' pnpm ops delete-channel --json '{"slug":"example-x","confirm":"example-x"}' pnpm ops channel-priority --json '{"slugs":["the-quartering"],"operation":"download","tier":"paused"}' pnpm ops lane --json '{"lane":"download","held":true}' pnpm ops refresh-report --json '{"all":true}' pnpm ops keep-videos --json '{"slug":"paramount-tactical","match":"TheQuartering","dryRun":true}' pnpm ops fetch-posts --json '{"slug":"example-x","older":true}' --wait pnpm ops capture-posts --json '{"slug":"example-x","ids":["1234567890"]}' --wait pnpm ops persist-videos --json '{"items":[{"slug":"example-channel","id":"abc123"}],"dryRun":true}' pnpm ops get channel the-quartering pnpm ops get channels # every channel, its kind and its sites pnpm ops list # every action name ``` These are examples. Every action and getter, with its body, is in [COMMANDS.md](COMMANDS.md) (generated from `pnpm ops --help` and `pnpm archilyzer --help`); recipes that chain them are in [OPERATING.md](OPERATING.md). Four things to know before you script against it: - **A job-starting action returns a `jobId` and does not stream.** The job may sit in a platform queue behind other work for hours, so "started" is the honest answer; `--wait` follows `/api/jobs//log` to the end and exits with the job's status. - **Unknown body keys are a 400.** A misspelled `downloadFilterExclude` would otherwise save cleanly and leave a channel downloading everything. - **`channel-config` patch keys are the CONFIGURE FORM's field names**, not `config.json`'s — `downloadFilterInclude` / `downloadFilterExclude` rather than a `downloadFilter` object. That is what routes them through the form's own validators, so a bad regex is refused here with the sentence the form shows. `""` clears a field, exactly as clearing the input does. `"sites"` is the form's Sites section — the WHOLE membership set, `[]` for on no site — and `create-channel`'s `"fields"` take the same names. - **`keep-videos` sets the do-not-clean marker** — the video page's "Do not clean" toggle, over every video of one channel whose title or description matches `match`. `match` is matched exactly as a `downloadFilterInclude` is (a case-insensitive regex over title + description, the same compile), and `fields: ["title"]` or `["description"]` narrows it to one half; `dryRun: true` reports without writing. The marker lives in the video's `data//`, so it is two steps on a fresh channel: run `metadata-scan` first (an id with no text cannot match — the reply counts those as `unscanned`), then `keep-videos`. Matches that were never downloaded come back in `notDownloaded` rather than getting a directory; pass them to `download-missing` and run `keep-videos` again to mark them. The read side needs no new routes for jobs: `/api/jobs/active`, `/api/jobs//log`, `/api/scheduler/status` and `/api/auto-queue/status` already exist. `GET /api/ops/channel/` is the one addition — config, report totals, bucket sizes, priority and, the part no directory listing can tell you, whether the channel's media is actually **reachable**. Add `--counts` (`?counts=1`) for the live on-disk counts; it is opt-in because it walks every video directory, eleven thousand of them on the largest channel here. ### Booting without resuming work `editor/instrumentation.ts` arms the sync heartbeat and every enabled auto-queue lane runner — transcription, download, digest and backfill — on every boot. That is right for a host install, where a restart interrupts work you own. It is wrong the first time you point a container at somebody else's corpus: its stored policies may have the digest lane switched on, and a corpus-wide digest pass is GPU-*weeks*. ```sh ARCHILYZER_IDLE_BOOT=1 ``` Boots the server with all of that stopped. You can start any of it from the UI afterwards. The shutdown reaper and the persisted-pause restore stay armed either way — both only ever *stop* work. ### Running as a unit executor for another machine A second machine can take whole work *units* (diarization, speaker attribution, CPU transcription) from a primary instance without holding any corpus at all. Two shapes, by cost: - **LLM calls only** (digest / attribution — the bulk of any backlog): run nothing but `ollama serve` with the primary's exact model tag pulled, and register the box on the primary as an **LLM endpoint** worker (Settings → Transcription workers). No repo, no container, no token. The primary verifies the model tag before use and refuses an endpoint that lacks it — freshness pins the model identity, so "almost the right model" would write permanently-stale records. - **A unit executor** (everything else): boot this app with ```sh ARCHILYZER_IDLE_BOOT=1 # never arm runners or resume sweeps here WORKER_TOKEN= # enables /api/worker/*; off without it TRANSCRIPTS_DIR=/some/empty/dir SETTINGS_FILE=/some/where/settings.json # pin it — the cwd fallback is a trap off-repo ``` and register it on the primary as a **Remote** worker with matching token. Units arrive with their inputs, run against a scratch corpus under `TRANSCRIPTS_DIR/.worker-scratch/`, and the primary pulls the produced sidecars back and applies them through its own guarded writers. The executor's own settings are never consulted for the work's identity — the primary injects its model/prompt configuration into every unit, and a unit that would fall back to local defaults fails loudly instead. **Tag the worker** (e.g. `cpu, diarization`) — units are shipped only to *tagged* remote workers, so an untagged remote from before this protocol keeps doing transcription only. The executor never opens the primary's LMDB, never downloads media, and never arms a sweep — the primary stays the sole scheduler. ### GPU transcription Two overlays, because two different engines get the GPU. Neither is required — the default image transcribes on the CPU and works everywhere. **Vulkan + parakeet.cpp — the one that runs on any GPU.** ```sh docker compose -f docker-compose.yml -f docker-compose.vulkan.yml up -d --build ``` Vulkan rather than a vendor SDK, so the same image drives AMD (RADV), Intel (ANV) and NVIDIA. On AMD it is the only GPU path here that works at all: ROCm is not packaged and CUDA is NVIDIA-only. The engine is [parakeet.cpp](https://github.com/mudler/parakeet.cpp), driven by the overlapping-window wrapper the app already ships (`scripts/parakeet-stitch.mjs`), and the first boot seeds a **parakeet** worker and fetches a parakeet GGUF instead of a whisper model. All the host has to provide is a render node: ```sh ls /dev/dri # renderD128, card0, … vulkaninfo --summary | head # names your GPU ``` The overlay passes `/dev/dri` straight through — no container toolkit, no runtime shim, no privileged mode. **Check that the container actually sees it.** This is the one thing worth verifying, because the failure is silent: with no render node inside, parakeet falls back to the CPU and produces perfectly correct transcripts an order of magnitude slower, and nothing raises an error. The entrypoint prints a `vulkan:` line on every boot saying which it got, and the image ships `vulkaninfo` so you can ask directly: ```sh docker compose -f docker-compose.yml -f docker-compose.vulkan.yml \ exec editor vulkaninfo --summary | head -20 ``` Pin a device with `PARAKEET_DEVICE` in `.env` — `Vulkan0`/`Vulkan1` for a specific GPU, or `cpu` to take the GPU out of the picture with everything else identical. That last one is also the **surest test that the GPU is really being used**, because it needs no log parsing: run the same audio both ways and compare. On an RX 6600 XT, one 33-second clip through the app's own wrapper: ``` default (Vulkan) 3.4 s --device cpu 36.3 s ``` A ~10× gap means the GPU is doing the work. No gap means you are on the CPU whatever anything else claims. Models live at [mudler/parakeet-cpp-gguf](https://huggingface.co/mudler/parakeet-cpp-gguf); `ARCHILYZER_FETCH_MODEL` names one without the `.gguf` (default `tdt_ctc-110m-q8_0` — small and quick; `tdt-0.6b-v3-q5_k` and `tdt_ctc-1.1b-q5_k` are better and slower). The image also carries CPU `whisper-cli`, so you can add a whisper worker on the Workers page to compare without rebuilding. **CUDA + whisper.cpp — NVIDIA only.** ```sh docker compose -f docker-compose.yml -f docker-compose.gpu.yml up -d --build ``` Needs the NVIDIA Container Toolkit on the host. Verify it first: ```sh docker run --rm --gpus all nvidia/cuda:12.6.3-base-ubuntu22.04 nvidia-smi ``` This builds a second whisper.cpp with CUDA offload and pulls a CUDA base image — several GB, and a long first build. Two build args take the edge off it: ```sh # one nvcc per job, and each is hungry — lower it on a shared or small machine --build-arg WHISPER_BUILD_JOBS=4 # every CUDA translation unit is compiled once per architecture; the default list # spans Pascal..Hopper, a single value is much faster if you know your card --build-arg CUDA_ARCHITECTURES=86 ``` Apple Metal is not covered: whisper.cpp supports it, but a Mac GPU is not reachable from a Linux container at all, so it would have to be a host install. ### Backup `corpus` is the volume that matters — real media, real transcripts, hundreds of GB when it grows up. `config` holds `settings.json`. `models` and `builds` are reproducible; losing them costs a download and a rebuild. ```sh docker run --rm -v archilyzer_corpus:/corpus -v "$PWD:/backup" \ debian:bookworm-slim tar czf /backup/corpus.tar.gz -C /corpus . ``` ### A channel whose media is on another drive The Storage panel can move a channel's `data/` to another root (see AGENTS.md). In a container that root is a path **inside the container**, and `channels//data` becomes an absolute symlink to it — so the drive must be bind-mounted **at the same absolute path the editor recorded**: ```yaml services: editor: volumes: - /mnt/platter/archilyzer-media:/mnt/platter/archilyzer-media ``` Mount it somewhere else and the link dangles. That is **reported as unreachable, by design** — the channel is skipped by the lane runners, its media jobs are refused and its report is not regenerated, rather than the alternative, which is every count on that channel reading zero and the download runner treating the whole archive as missing. The badge on `/channels` and the channel's Storage panel name the path they cannot reach. #### Storage locations in a container: re-point by path, and that is the whole story `/storage` names each media root as a **location** and reports whether it is there. On a host it can do more than that: it learns the volume's filesystem UUID from `findmnt`, so when a drive comes back at a different mountpoint the page offers **Re-point** and the operator takes it in one click. **Inside a container none of that identity exists.** Block devices are not passed through, so `findmnt -J -T ` describes the bind mount and not the disk behind it: no UUID, no `/dev/disk/by-uuid` entry, nothing to mount with `udisksctl`. Every identity probe **fails open to "unknown"** — deliberately, because a probe that turned "I could not ask" into "your disk is gone" would declare every containerised corpus broken. What a location reports here is what `stat` says and nothing more: | Status | What it means in a container | |---|---| | Available | The root is a directory. The bind mount is up. | | Missing | The root is not there, and there is no identity to look for. | | Not attached | The root is not there and a **recorded** UUID was found nowhere. | | Mounted elsewhere / Not mounted | **Never reported.** Both need a UUID the probe can find *now*. | `Not attached` does appear here, and only for one reason: a `settings.json` authored on a host carries the `volume.uuid` learned there, and it survives being bind-mounted into the container. Inside, `findmnt -S UUID=…` and `/dev/disk/by-uuid/` both miss, so the probe reports `absent` — which reads as "Not attached" but **means exactly what Missing means here**: the root is not at that path. There is no disk to go looking for and nothing to mount; fix the path. So the container's remedy is the manual one, and it is not a downgrade: **change the location's root to the path the media is actually at, and re-point.** Edit the location on `/storage`, or press Re-point after correcting the root; the job rewrites each channel's `data/` symlink and its `config.dataDir` and moves no bytes. Equivalently, fix the compose file so the bind mount lands where the editor recorded — the same `-v /host/path:/container/path` line above — and nothing needs re-pointing at all. `autoRepoint` has nothing to act on here and can stay off. The `site` profile does not need the mount: an export build never reads `data/`. ### Useful commands ```sh docker compose logs -f editor docker compose exec editor bash # a shell in the image docker compose exec editor yt-dlp --version docker compose exec editor pnpm archilyzer doctor # what this container can do docker compose exec editor pnpm archilyzer publish status docker compose restart editor docker compose down # stop; volumes survive docker compose down -v # stop AND DELETE the corpus ``` --- ## What is in the image, and what is not **Baked in:** Node 22 (the pinned wrangler needs it) + the installed workspace, the built editor / umtool / homepage, `yt-dlp`, a JavaScript runtime for it (`deno` — without one yt-dlp warns and silently loses formats), `ffmpeg`/`ffprobe`, `whisper-cli` (statically linked), `zip`/`tar`/`xz`/`gzip`, `rsync`, `git`, `curl`, `ps`; `python3` with yt-dlp's optional modules (for a substituted yt-dlp), `pipx` and a pinned `git-filter-repo` (the homepage's source mirror), and — in the workspace's `node_modules` — a pinned `wrangler`, so a deploy fetches nothing. The Vulkan image adds `parakeet-cli`, the Mesa Vulkan drivers and `vulkaninfo`. **Not baked, on purpose:** - **The corpus.** It is a separate git repo holding production data, and it lives in the `corpus` volume. - **whisper models.** 142 MB to 3 GB, and the choice is yours. Fetched on boot. - **The export site build.** It is a static render *of a corpus*, and there is no corpus at image-build time. The publish stages make it at run time. - **Credentials.** `CLOUDFLARE_API_TOKEN` and friends come from `.env` at run time, never from the image. - **The repository.** No `.git`; `docker-compose.source.yml` mounts the host's for the `/source` mirror. The build's commit and branch are baked as `ARCHILYZER_COMMIT`/`ARCHILYZER_BRANCH` when the build passes them (`ARCHILYZER_COMMIT=$(git rev-parse HEAD) ARCHILYZER_BRANCH=$(git branch --show-current) docker compose build`) — the publish stamps record them. - **ImageMagick with Pango, and `qrencode`.** Only `umtool/report-to-video/` needs them, and rendering a report to video is a workstation task, not something a server does. Run that part on a host checkout. - **gallery-dl.** The X/Twitter post fetcher. Install it and point `GALLERY_DL_BIN` at it if you want the social corpus. - **ollama / claude.** The digest lanes reach a local ollama or the `claude` CLI. Point `OLLAMA_URL` at a host ollama (`http://host.docker.internal:11434`) if you want digests. ### Two build runners, and the container has one Every site build goes through one stage contract with two runners: - **local** — the default everywhere, in a container or not: each stage is a child process of the editor (or of the CLI you ran), one at a time. - **docker** — a HOST install's opt-in fan-out (`settings.publish.runner: "docker"`, or `publish build all --runner docker`): one `Dockerfile.build` container per site (see [PUBLISH.md](PUBLISH.md#building-every-site-in-containers)). Both write the same bundles and stamps, so a deploy does not care which built it. Inside the container there is no engine, and the docker runner **refuses** ("the docker runner needs an engine on this host") rather than pretend. The editor never gets the docker socket — **do not** mount it, and do not try to make docker-in-docker work; run the fan-out from a host checkout if you want it. ### It runs as root Volume ownership across Docker Desktop, WSL2 and Linux is not worth the class of bug that comes with getting it subtly wrong, and everything the container touches is either a volume or the image itself. The apps are not reachable from the network by default, which is the control that matters here. --- ## Windows This is the reason the stack exists. Install **Docker Desktop** (which uses WSL2 underneath, without you having to run anything inside it), then: ```powershell git clone archilyzer cd archilyzer copy .env.example .env docker compose up -d --build ``` Open http://localhost:8081. Keep the clone on the Windows filesystem — the build context is small, and the corpus lives in a Docker volume inside the WSL2 VM, which is where the I/O actually happens. Publishing needs nothing else on Windows — Docker Desktop is the whole requirement. A checklist, in PowerShell from the clone: 1. `copy .env.example .env`, then set in `.env`: `WORKER_TOKEN` (for `pnpm ops` and agents), `CLOUDFLARE_API_TOKEN` and `CLOUDFLARE_ACCOUNT_ID` (to deploy). 2. `docker compose up -d --build` 3. Open http://localhost:8081 → **Sites** → **Publish now**, or: `docker compose exec editor pnpm archilyzer publish now` 4. A local preview: `docker compose --profile site up -d site`, then `docker compose exec editor pnpm archilyzer publish deploy --to local`, and open http://localhost:8080. 5. `docker compose exec editor pnpm archilyzer doctor` — every check, including whether the Cloudflare token is set (never its value). 6. Optional, your own yt-dlp: set `YTDLP_SOURCE_HOST_DIR` (a Windows path such as `C:\src\yt-dlp` works) and run `docker compose -f docker-compose.yml -f docker-compose.ytdlp.yml up -d`. You still want WSL2 directly if you intend to *develop* the project or run Claude Code against it — see [README.md](README.md#claude-code-on-windows). For running an archive, this is enough. --- ## Troubleshooting **`REFUSING TO START`** — a private app is bound off-loopback with no auth. The message lists the four fixes; see [the rail](#the-rail-and-its-escape-hatch). Both the app and Caddy refuse, and `restart: unless-stopped` means they keep retrying: `docker compose ps` shows them `Restarting (1)` until you fix `.env`. Nothing is listening on the exposed port while that is true. **A port answers 502.** Its service is not running. `site`, `homepage` and `umtool` are behind profiles; start the one you want. Caddy publishes all four ports regardless, and this is what an unstarted one looks like. **Transcription fails with a missing model.** The first-boot download failed (check `docker compose logs editor`). `docker compose restart editor` retries it. **Transcription is slow on the Vulkan image.** Almost certainly the silent CPU fallback: check `docker compose logs editor | grep vulkan:` and the `--device cpu` comparison above. The usual cause is a missing `/dev/dri` — a container without the render node still transcribes, just slowly. **Downloads started failing on every channel.** Almost always a stale yt-dlp: `docker compose exec editor yt-dlp -U`. Check `docker compose logs editor | grep yt-dlp:` first — `(override)` means `YTDLP_BIN` names your own build, which `-U` does not touch, and `MISSING` means it does not run at all (a from-source override with no checkout mounted, say). **A deploy is refused before wrangler runs.** `CLOUDFLARE_API_TOKEN` is not set in `.env`, or the editor was started before it was: `docker compose up -d` again (a restart does not re-read `.env`; `up` recreates the container). `archilyzer doctor`'s `cloudflare-auth` line says which. **Port 8080 is already taken.** `SITE_HTTP_PORT=9080` in `.env`. Those variables move the *outside* port; nothing inside the containers changes. **A build resumed by itself after `docker compose up`.** That is the boot behaviour described under [booting without resuming work](#booting-without-resuming-work) — set `ARCHILYZER_IDLE_BOOT=1`.