Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 114b64e7cdc9e5d72e7f3801f9dee8fdc8865094
parent 0d33a251c076ed4338fa98e4c539a3fc855e95ad
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Fri, 21 Aug 2026 08:47:41 -0400

Merge branch 'feat/runtime-container'

Diffstat:
M.dockerignore | 2++
A.env.example | 109+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
M.gitignore | 3+++
MAGENTS.md | 77+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
MDEPLOY_DOCKER.md | 6++++++
ADockerfile | 477+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
MREADME.md | 40+++++++++++++++++++++++++++++++---------
ARUNNING_IN_DOCKER.md | 454+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Acommon/lib/idleBoot.test.ts | 27+++++++++++++++++++++++++++
Acommon/lib/idleBoot.ts | 34++++++++++++++++++++++++++++++++++
Adocker-compose.authelia.yml | 42++++++++++++++++++++++++++++++++++++++++++
Adocker-compose.gpu.yml | 44++++++++++++++++++++++++++++++++++++++++++++
Adocker-compose.tinyauth.yml | 47+++++++++++++++++++++++++++++++++++++++++++++++
Adocker-compose.vulkan.yml | 75+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Adocker-compose.yml | 178+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Adocker/Caddyfile | 84+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Adocker/authelia/configuration.yml.example | 47+++++++++++++++++++++++++++++++++++++++++++++++
Adocker/authelia/users_database.yml.example | 13+++++++++++++
Adocker/caddy-start.sh | 47+++++++++++++++++++++++++++++++++++++++++++++++
Adocker/entrypoint.sh | 286+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Adocker/guard-exposure.sh | 96+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Adocker/publish-site.sh | 50++++++++++++++++++++++++++++++++++++++++++++++++++
Meditor/CHANGELOG.md | 4++++
Meditor/instrumentation.ts | 34+++++++++++++++++++++++++++++-----
24 files changed, 2262 insertions(+), 14 deletions(-)

diff --git a/.dockerignore b/.dockerignore @@ -35,6 +35,8 @@ export/.r2-staging/ *.pem .DS_Store +Dockerfile Dockerfile.test Dockerfile.build +docker-compose*.yml .dockerignore diff --git a/.env.example b/.env.example @@ -0,0 +1,109 @@ +# Copy to .env before the first `docker compose up`. Every value here is +# optional — the defaults are the safe ones. +# +# cp .env.example .env +# +# See RUNNING_IN_DOCKER.md. + +# --------------------------------------------------------------------------- +# EXPOSURE. This is the part worth reading. +# --------------------------------------------------------------------------- +# Caddy publishes one port per app. Each of these is the HOST INTERFACE that +# port binds to — not a hostname, not a URL. 127.0.0.1 means "this machine +# only": not the LAN, not the internet, not another container. +# +# All four default to loopback, INCLUDING the two public sites. Opening one is a +# deliberate edit of a single line. +# +# 127.0.0.1 this machine only (the default) +# 0.0.0.0 every interface — LAN, and the internet if your router forwards +# 192.168.x.y one specific interface +# +# The archive itself. Serving this to the world is what it is for; it is static +# files with no write path. Still off by default. +#SITE_BIND=0.0.0.0 + +# The EDITOR. Read the auth section below before changing this. The editor has +# no login of its own: anyone who reaches it can download, delete and +# re-transcribe anything in your corpus. The container REFUSES TO START if you +# put it on a network with no auth in front. +#EDITOR_BIND=0.0.0.0 + +# The project's own site (marketing + docs). Only useful if you're mirroring it. +#HOMEPAGE_BIND=0.0.0.0 + +# umtool. Private, same rules as the editor. +#UMTOOL_BIND=0.0.0.0 + +# The published port numbers, if 8080-8083 collide with something you run. +# These are the OUTSIDE ports; the ports inside the containers never change. +#SITE_HTTP_PORT=8080 +#EDITOR_HTTP_PORT=8081 +#HOMEPAGE_HTTP_PORT=8082 +#UMTOOL_HTTP_PORT=8083 + +# --------------------------------------------------------------------------- +# AUTH for the two private apps (editor, umtool). +# --------------------------------------------------------------------------- +# basic (default) Caddy's built-in basic_auth, when a hash is set below. +# Nothing to install, nothing extra to keep running. +# forward hand the decision to Tinyauth or Authelia — see the +# docker-compose.tinyauth.yml / docker-compose.authelia.yml +# overlays, which set this for you. +# none you have your own auth in front, or you are on a tailnet and +# mean it. This is the explicit escape hatch from the startup +# refusal — nothing else disables it. +#ARCHILYZER_AUTH_MODE=basic + +# Generate the hash (the `$` characters need no escaping in this file): +# +# docker run --rm caddy:2.11-alpine caddy hash-password --plaintext 'your-password' +# +#ARCHILYZER_AUTH_USER=you +#ARCHILYZER_AUTH_HASH=$2a$14$replace.this.with.the.hash.the.command.printed + +# --------------------------------------------------------------------------- +# First run +# --------------------------------------------------------------------------- +# The speech model fetched into the models volume on first boot, and used by the +# seeded transcription worker. Which FAMILY depends on the image you run: +# +# default / CUDA image a whisper.cpp model. base.en (~142 MB) is a sane +# starting point; small.en (~466 MB) and medium.en +# (~1.5 GB) are better and slower; tiny.en (~75 MB) is +# faster and noticeably worse. +# Vulkan image a parakeet GGUF from mudler/parakeet-cpp-gguf, named +# without the extension. Default tdt_ctc-110m-q8_0; +# tdt-0.6b-v3-q5_k and tdt_ctc-1.1b-q5_k are better and +# slower. +# +# `none` skips the download entirely. +# +# Changing this AFTER the first boot fetches the new model but does not switch +# the worker over — do that on the editor's Workers page, which is where model +# choice actually lives. +#ARCHILYZER_FETCH_MODEL=base.en + +# Run `yt-dlp -U` on every boot. An image pins yt-dlp at build time and a stale +# yt-dlp is the most common reason downloads start failing. Off by default +# because it is a network call at startup. +#YTDLP_AUTO_UPDATE=1 + +# Boot the editor WITHOUT resuming the schedulers, the auto-queue runners or the +# digest/backfill sweeps. Set this the first time you point a container at a +# corpus somebody else configured: its stored policies may say "sweep", and you +# probably want to look around before GPU-weeks of work starts on its own. +# Unset (the default) matches a host install: everything armed on boot. +#ARCHILYZER_IDLE_BOOT=1 + +# Which compute device parakeet uses — Vulkan image only (see +# docker-compose.vulkan.yml). Unset lets parakeet.cpp choose, which is usually +# right. "Vulkan0"/"Vulkan1" pin a specific GPU (`vulkaninfo --summary` lists +# them in order). "cpu" takes the GPU out of the picture with everything else +# identical — the honest way to measure what it is buying you. +#PARAKEET_DEVICE=Vulkan0 + +# --------------------------------------------------------------------------- +# Image +# --------------------------------------------------------------------------- +#ARCHILYZER_TAG=local diff --git a/.gitignore b/.gitignore @@ -32,6 +32,9 @@ yarn-error.log* # env files (can opt-in for committing if needed) .env* +# ...except the template the docker stack tells you to copy. It carries the +# exposure model in comments and is worth more in git than out of it. +!.env.example # vercel .vercel diff --git a/AGENTS.md b/AGENTS.md @@ -87,6 +87,83 @@ better), whose clock is not the same. This tooling is on `main` as of 2026-08-20. +# The runtime container + +`docker compose up -d` stands up a working archive: the editor plus Caddy, with +`site`, `homepage` and `umtool` behind compose **profiles**. The root `Dockerfile` +is new and is a THIRD Dockerfile — `Dockerfile.build` (per-site export build +fan-out) and `Dockerfile.test` (sharded e2e) are untouched and unrelated. + +Three runtime targets share one build: `runtime` (CPU whisper.cpp, the default), +`runtime-vulkan` (parakeet.cpp on Vulkan — AMD/Intel/NVIDIA, `/dev/dri` passed +through, `docker-compose.vulkan.yml`) and `runtime-cuda` (NVIDIA-only whisper, +`docker-compose.gpu.yml`). `ARCHILYZER_TRANSCRIBER` is baked per target and is +what makes `docker/entrypoint.sh` seed a parakeet worker and fetch a GGUF rather +than a whisper worker and a `.bin`. + +**A Vulkan container with no `/dev/dri` does not fail — it transcribes on the +CPU**, correctly and ~10× slower (measured on an RX 6600 XT: 3.4 s vs 36.3 s for +the same 33-second clip). The entrypoint prints a `vulkan:` line on every boot +for exactly that reason. If you touch this, keep that line honest. + +Two ways the Vulkan build silently degrades or breaks, both already paid for: + +- **`-DGGML_VULKAN=ON` is the wrong flag** for parakeet.cpp. Its CMakeLists does + `set(GGML_VULKAN ${PARAKEET_GGML_VULKAN} CACHE BOOL "" FORCE)`, so passing + `GGML_VULKAN` is not ignored — it is OVERWRITTEN with OFF. The build succeeds, + ships no `libggml-vulkan.so`, and every transcription runs on the CPU. Use + `PARAKEET_GGML_VULKAN=ON`; the stage now asserts the library exists. +- **The Vulkan stage builds on trixie, not bookworm.** ggml-vulkan needs Vulkan + headers ≥ ~1.3.272 for `vk::LayerSettingEXT`; bookworm ships 1.3.239 and the + compile fails inside ggml-vulkan.cpp. That is also why the Vulkan RUNTIME is + trixie, and why `RUNTIME_IMAGE` is a build arg. + +**glibc only goes forward.** The workspace (with native modules) is compiled once +in the `build` stage and copied into every runtime, so `NODE_IMAGE` must have the +OLDEST glibc of any runtime it lands in — bookworm 2.36 < ubuntu 24.04 2.39 < +trixie 2.41. This is why the CUDA bases are ubuntu24.04 and not 22.04 (2.35, +older than the build stage — the native modules would not load). + +**`CMAKE_CUDA_ARCHITECTURES=all-major` does not work here** and the error names +neither cmake nor the flag: the CUDA base ships CMake 3.22, `all-major` landed in +3.23, and nvcc receives the literal string +(`nvcc fatal : Unsupported gpu architecture 'compute_'`). The Dockerfile pins an +explicit list. + +It ships the repo plus `node_modules`, deliberately **not** `output: "standalone"`. +`editor/next.config.ts` explains why: the server does runtime-dynamic `fs` reads the +tracer cannot bound, so the `.nft.json` traces are never consumed and a bundle built +from them would be missing files nobody can enumerate. + +**Everything is guarded on the local docker host by default.** No application +container publishes a port; Caddy is the single front door and binds `127.0.0.1` on +all four ports, public sites included. `docker/guard-exposure.sh` then **refuses to +start** if a private app (editor, umtool) is bound off-loopback with no auth in +front — run by the app entrypoint *and* by the caddy container, which is the process +that actually opens the ports. `ARCHILYZER_AUTH_MODE=none` is the only escape hatch. + +Two things the image cannot bake, and the reasons matter: + +- **The export site.** It is a static render OF a corpus, and there is no corpus at + image-build time. `docker/publish-site.sh` builds it at run time into the volume + the `site` service serves. +- **whisper models.** 142 MB to 3 GB, and the choice is the operator's. + `docker/entrypoint.sh` fetches one on first boot — and seeds a `settings.json` + carrying one enabled worker, because `defaults()` returns `workers: []` and zero + workers means auto-transcribe silently does nothing. + +`ARCHILYZER_IDLE_BOOT=1` boots the editor without arming the heartbeat, the +auto-queue runners or the digest/backfill sweeps (`common/lib/idleBoot.ts`) — for +pointing a fresh container at a corpus whose stored policies would otherwise resume +GPU-weeks of work. The shutdown reaper stays armed regardless. + +Inside a container the multi-site build pipeline has no `docker` binary and falls +back to the serial host build it already handles. Do not try to make +docker-in-docker work. + +See [RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md) — which is about *running the +apps*, not [DEPLOY_DOCKER.md](DEPLOY_DOCKER.md), which is about *building sites*. + # The corpus, and where the live sites are configured `transcripts/` is **its own git repo**, gitignored by the workspace (`/transcripts` in diff --git a/DEPLOY_DOCKER.md b/DEPLOY_DOCKER.md @@ -1,5 +1,11 @@ # Docker export build pipeline +> **Not the document you want if you are trying to *run* the apps in containers.** +> That is [RUNNING_IN_DOCKER.md](./RUNNING_IN_DOCKER.md) — `docker compose up` and a +> working archive server, built from the root `Dockerfile`. This page is about +> *building sites*: fanning per-site export builds out across containers, using +> `Dockerfile.build`. The two share nothing but the word "docker". + Archilyzer's Docker build mode builds **every site in parallel** in isolated containers, then deploys them serially — a large speedup when you host several sites, and stronger isolation than the basic single-process build. This is opt-in: set **Build diff --git a/Dockerfile b/Dockerfile @@ -0,0 +1,477 @@ +# Runtime image for the whole app: the editor (admin), the static site server, +# the project homepage and umtool, plus the tools they shell out to (yt-dlp, +# ffmpeg, and a transcription backend). +# +# This is NOT one of the other two Dockerfiles and does not replace either: +# Dockerfile.build — one hermetic container per SITE for the export build fan-out. +# Dockerfile.test — the sharded playwright e2e route. +# Both keep working exactly as they did; nothing here touches them. +# +# Three published targets, all built from the same app code: +# runtime CPU whisper.cpp. The default. (docker-compose.yml) +# runtime-vulkan parakeet.cpp on Vulkan. Any GPU with a Vulkan driver — +# AMD/RADV, Intel, NVIDIA. (docker-compose.vulkan.yml) +# runtime-cuda whisper.cpp on CUDA. NVIDIA only. (docker-compose.gpu.yml) +# +# Used by docker-compose.yml. See RUNNING_IN_DOCKER.md. +# +# --------------------------------------------------------------------------- +# Why the whole repo ships instead of `output: "standalone"` +# --------------------------------------------------------------------------- +# editor/next.config.ts (and umtool/next.config.ts) explain it: the server does +# runtime-dynamic fs reads the file tracer cannot statically bound — readdir over +# a video dir, per-job logs, settings.json — so the .nft.json traces Next emits +# are deliberately never consumed. A standalone bundle built from them would be +# missing files nobody can enumerate ahead of time. So we ship the repo plus +# node_modules, exactly as Dockerfile.test already does. Read those comments +# before "optimising" this. + +# --------------------------------------------------------------------------- +# Base images, and the one rule that ties them together +# --------------------------------------------------------------------------- +# GLIBC ONLY GOES FORWARD. A binary built against an older glibc runs on a newer +# one; the reverse fails at load time with "GLIBC_2.xx not found". The workspace +# (including native modules — lmdb, msgpackr-extract) is compiled ONCE in the +# `build` stage and copied into every runtime, so: +# +# NODE_IMAGE must have the OLDEST glibc of anything it is copied into. +# +# bookworm is 2.36, trixie 2.41, ubuntu 24.04 2.39 — so building on bookworm and +# running on any of them is safe, and moving NODE_IMAGE to trixie would silently +# break the CUDA target. If you bump one of these, check that direction first. +ARG NODE_IMAGE=node:20-bookworm-slim +# The runtime base, overridable per target: docker-compose.vulkan.yml builds with +# trixie because the Vulkan stack needs it (see the parakeet stage below). +ARG RUNTIME_IMAGE=node:20-bookworm-slim +# ubuntu24.04, not 22.04: 22.04 is glibc 2.35, OLDER than the bookworm the +# workspace is built on, and the native modules would not load. +ARG CUDA_DEVEL_IMAGE=nvidia/cuda:12.6.3-devel-ubuntu24.04 +ARG CUDA_RUNTIME_IMAGE=nvidia/cuda:12.6.3-runtime-ubuntu24.04 +# whisper.cpp release to build. Pinned: this is a compiled dependency, and +# "whatever master was that day" is not a thing you can reproduce later. +ARG WHISPER_REF=v1.7.6 +# parakeet.cpp, same reasoning. https://github.com/mudler/parakeet.cpp +ARG PARAKEET_REF=v0.5.0 +# Compile parallelism for the transcription backends. Defaults to every core; lower it when the +# machine is shared or short on RAM (the CUDA build in particular runs one nvcc +# per job and each is hungry). `--build-arg WHISPER_BUILD_JOBS=4`. +ARG WHISPER_BUILD_JOBS= +# Which CUDA architectures the GPU build targets. Turing through Hopper, plus +# Pascal, so the image runs on more than the card it was built next to — at a +# real cost in build time, since every CUDA translation unit is compiled once per +# architecture. Narrow it (e.g. "86") when you are building for one known GPU. +# +# NOT "all-major", which is the obvious thing to write here and does not work: +# CMake only understands that value from 3.23, the CUDA base image ships CMake +# 3.22, and the failure is not a cmake error but +# nvcc fatal : Unsupported gpu architecture 'compute_' +# i.e. the literal string reaching nvcc as if it were an architecture. +ARG CUDA_ARCHITECTURES="61;70;75;80;86;89;90" +# ggml compiles the CPU backend with -march=native by default, which bakes THIS +# machine's instruction set into the image. That is right for the normal case — +# you build on the box you run on, and it is measurably faster — and wrong the +# moment the image moves: an older CPU dies with SIGILL on the first tensor op. +# Set GGML_NATIVE=OFF when building an image you intend to publish or run +# elsewhere. +ARG GGML_NATIVE=ON + +# --------------------------------------------------------------------------- +# whisper-cpu — build whisper-cli, the default transcription backend. +# +# Statically linked (BUILD_SHARED_LIBS=OFF) so the runtime stage copies ONE file +# and inherits no libwhisper/libggml search-path problems. Models are NOT baked: +# they are 142 MB (base.en) to 3 GB (large-v3), they change independently of the +# code, and they belong in the `models` volume. docker/entrypoint.sh fetches one +# on first boot. +# --------------------------------------------------------------------------- +FROM debian:bookworm-slim AS whisper-cpu +ARG WHISPER_REF +ARG WHISPER_BUILD_JOBS +ARG GGML_NATIVE +RUN apt-get update \ + && apt-get install -y --no-install-recommends \ + build-essential cmake git ca-certificates \ + && rm -rf /var/lib/apt/lists/* +RUN git clone --depth 1 --branch "${WHISPER_REF}" \ + https://github.com/ggml-org/whisper.cpp /src/whisper.cpp +RUN cmake -S /src/whisper.cpp -B /src/whisper.cpp/build \ + -DCMAKE_BUILD_TYPE=Release \ + -DBUILD_SHARED_LIBS=OFF \ + -DGGML_NATIVE="${GGML_NATIVE}" \ + -DWHISPER_BUILD_TESTS=OFF \ + -DWHISPER_BUILD_SERVER=OFF \ + && cmake --build /src/whisper.cpp/build --config Release \ + -j "${WHISPER_BUILD_JOBS:-$(nproc)}" --target whisper-cli \ + && cp /src/whisper.cpp/build/bin/whisper-cli /usr/local/bin/whisper-cli \ + && strip /usr/local/bin/whisper-cli + +# --------------------------------------------------------------------------- +# parakeet-vulkan — build parakeet-cli with the ggml Vulkan backend. +# +# On TRIXIE, and not by preference: ggml-vulkan uses vk::LayerSettingEXT, which +# arrived in the Vulkan headers around 1.3.272. Bookworm ships 1.3.239, so the +# build dies in ggml-vulkan.cpp with "'LayerSettingEXT' is not a member of 'vk'". +# Trixie has 1.4.309 and the matching glslc/glslangValidator. That is also why +# the Vulkan RUNTIME is trixie — this binary links trixie's libstdc++/glibc. +# +# Vulkan rather than a vendor SDK on purpose: one build runs on AMD (RADV), +# Intel and NVIDIA alike, and needs no proprietary runtime in the image — just +# the driver's ICD, which comes from mesa-vulkan-drivers in the runtime stage. +# +# The app already knows how to drive this and none of it is new app code: +# common/lib/transcriptionApps.ts registers a `parakeet` app whose binary is +# scripts/parakeet-stitch.mjs (the overlapping-window wrapper), which shells out +# to PARAKEET_CLI and passes the worker's device through as PARAKEET_DEVICE. +# +# Submodules are NOT optional — ggml lives in third_party/ggml, and a plain +# --depth 1 clone produces a tree that fails to configure. +# --------------------------------------------------------------------------- +FROM debian:trixie-slim AS parakeet-vulkan +ARG PARAKEET_REF +ARG WHISPER_BUILD_JOBS +ARG GGML_NATIVE +# The Vulkan backend needs more than the loader: ggml-vulkan compiles its +# compute shaders at BUILD time, so it wants glslc AND glslangValidator +# (glslang-tools), and it find_package()s SPIRV-Headers — without which cmake +# stops at "Could not find a package configuration file provided by +# SPIRV-Headers", several lines below a cheerful "Found Vulkan". +RUN apt-get update \ + && apt-get install -y --no-install-recommends \ + build-essential cmake git ca-certificates \ + libvulkan-dev glslc glslang-tools spirv-headers spirv-tools \ + && rm -rf /var/lib/apt/lists/* +RUN git clone --depth 1 --branch "${PARAKEET_REF}" --recurse-submodules \ + https://github.com/mudler/parakeet.cpp /src/parakeet.cpp +# PARAKEET_GGML_VULKAN, *not* GGML_VULKAN — and the difference is silent. +# parakeet.cpp's CMakeLists does +# set(GGML_VULKAN ${PARAKEET_GGML_VULKAN} CACHE BOOL "" FORCE) +# so passing GGML_VULKAN=ON is not ignored, it is OVERWRITTEN with OFF. The build +# then succeeds, prints only "Including CPU backend", ships no +# libggml-vulkan.so, and every transcription runs on the CPU with nothing +# anywhere saying why. Measured: that is exactly what the first build of this +# image did. +# Three steps, not one `&&` chain: chaining lets a cmake CONFIGURE error fall +# through to the guard's message, which then reports the wrong problem. +RUN set -eux; \ + cmake -S /src/parakeet.cpp -B /src/parakeet.cpp/build \ + -DCMAKE_BUILD_TYPE=Release \ + -DPARAKEET_GGML_VULKAN=ON \ + -DGGML_NATIVE="${GGML_NATIVE}" \ + -DPARAKEET_BUILD_TESTS=OFF \ + -DPARAKEET_BUILD_SERVER=OFF; \ + cmake --build /src/parakeet.cpp/build --config Release \ + -j "${WHISPER_BUILD_JOBS:-$(nproc)}"; \ + if [ -z "$(find /src/parakeet.cpp/build -name 'libggml-vulkan.so*' -print -quit)" ]; then \ + echo "FATAL: this build produced no Vulkan backend — see the comment above" >&2; \ + exit 1; \ + fi +# ggml builds each backend as its own shared object, nested a level deeper than +# the rest (third_party/ggml/src/ggml-vulkan/). A COPY glob does not recurse, so +# collect everything into one directory here — a missing libggml-vulkan.so is +# what silently turns a "GPU image" into a CPU one. +RUN set -eux; \ + mkdir -p /out/lib; \ + cp "$(find /src/parakeet.cpp/build -name parakeet-cli -type f -perm -u+x | head -1)" \ + /out/parakeet-cli; \ + find /src/parakeet.cpp/build -name '*.so*' -type f -exec cp -a {} /out/lib/ \;; \ + ls -1 /out/lib + +# --------------------------------------------------------------------------- +# whisper-cuda — the same binary with CUDA offload, for the GPU overlay. +# +# Only ever built when something asks for the runtime-cuda target (BuildKit +# builds a target's graph, not the file), so an unreachable or wrong CUDA base +# never blocks the default `docker compose build`. +# --------------------------------------------------------------------------- +FROM ${CUDA_DEVEL_IMAGE} AS whisper-cuda +ARG WHISPER_REF +ARG WHISPER_BUILD_JOBS +ARG CUDA_ARCHITECTURES +ARG GGML_NATIVE +RUN apt-get update \ + && apt-get install -y --no-install-recommends \ + build-essential cmake git ca-certificates \ + && rm -rf /var/lib/apt/lists/* +RUN git clone --depth 1 --branch "${WHISPER_REF}" \ + https://github.com/ggml-org/whisper.cpp /src/whisper.cpp +# Shared libs here: the CUDA backend links against the driver stack anyway, so a +# static build buys nothing, and the CUDA_ARCHITECTURES list keeps the image +# usable on more than the card it was built next to. +# The stubs directory is not optional. ggml's CUDA backend calls the DRIVER API +# (cuMemCreate, cuMemMap, … for its virtual-memory allocator), which lives in +# libcuda.so — shipped by the NVIDIA DRIVER, not by the toolkit, and therefore +# absent from a build container. The toolkit provides a link-time stub for +# exactly this case; without it the compile succeeds and the final link dies with +# a wall of "undefined reference to `cuMemCreate'". The real libcuda.so.1 is +# injected at run time by the NVIDIA container runtime. +# +# BOTH halves are needed: -L points at the stub directory, -lcuda actually links +# it. Adding only the -L changes nothing and fails identically — measured. +RUN cmake -S /src/whisper.cpp -B /src/whisper.cpp/build \ + -DCMAKE_BUILD_TYPE=Release \ + -DGGML_CUDA=ON \ + -DGGML_NATIVE="${GGML_NATIVE}" \ + -DCMAKE_CUDA_ARCHITECTURES="${CUDA_ARCHITECTURES}" \ + -DCMAKE_EXE_LINKER_FLAGS="-L/usr/local/cuda/lib64/stubs -lcuda" \ + -DCMAKE_SHARED_LINKER_FLAGS="-L/usr/local/cuda/lib64/stubs -lcuda" \ + -DWHISPER_BUILD_TESTS=OFF \ + -DWHISPER_BUILD_SERVER=OFF \ + && cmake --build /src/whisper.cpp/build --config Release \ + -j "${WHISPER_BUILD_JOBS:-$(nproc)}" --target whisper-cli +# Collect the binary and EVERY shared object the build produced into one dir. +# A COPY glob does not recurse, and the CUDA backend's library is nested a level +# deeper than the rest (ggml/src/ggml-cuda/) — copying "the .so files in +# ggml/src" silently misses the one that makes this variant a GPU build at all. +RUN mkdir -p /out/lib \ + && cp /src/whisper.cpp/build/bin/whisper-cli /out/whisper-cli \ + && find /src/whisper.cpp/build -name '*.so*' -type f -exec cp -a {} /out/lib/ \; + +# --------------------------------------------------------------------------- +# deps — install the workspace. +# --------------------------------------------------------------------------- +FROM ${NODE_IMAGE} AS deps + +# pnpm as a plain global binary, NOT corepack — same reason Dockerfile.build +# gives: there is no `packageManager` pin in package.json, so corepack would try +# to fetch pnpm from the registry at run time. +RUN npm install -g pnpm@9.15.4 + +WORKDIR /repo + +# Manifests first, for layer caching: this layer rebuilds only when a package.json +# or the lockfile moves. ALL SEVEN workspace packages listed in pnpm-workspace.yaml +# must be copied — a --frozen-lockfile install of a workspace with a missing member +# fails outright. (Dockerfile.build copies two because it only ever builds the +# export; that shortcut does not transfer here.) +COPY pnpm-lock.yaml pnpm-workspace.yaml package.json ./ +COPY common/package.json common/package.json +COPY editor/package.json editor/package.json +COPY export/package.json export/package.json +COPY homepage/package.json homepage/package.json +COPY mcp/package.json mcp/package.json +COPY scripts/report-to-video/package.json scripts/report-to-video/package.json +COPY umtool/package.json umtool/package.json + +# `allowBuilds` / `onlyBuiltDependencies` in pnpm-workspace.yaml rebuild the +# native modules (lmdb, msgpackr-extract, esbuild) for THIS image's arch. +RUN pnpm install --frozen-lockfile + +# --------------------------------------------------------------------------- +# build — compile the two server apps and the homepage. +# --------------------------------------------------------------------------- +FROM deps AS build + +COPY . . + +# There is no corpus at image-build time and there must not be: it is hundreds of +# GB, it is a separate git repo, and .dockerignore excludes it. Point the build +# at an empty one so anything that stats a corpus path finds an empty archive +# rather than a path that cannot exist. +ENV TRANSCRIPTS_DIR=/tmp/empty-corpus \ + NEXT_TELEMETRY_DISABLED=1 +RUN mkdir -p /tmp/empty-corpus/channels /tmp/empty-corpus/sites + +RUN pnpm --filter editor exec next build +RUN pnpm --filter umtool exec next build + +# The homepage is the project's own marketing/docs site and is corpus-independent +# — build:nodata skips the compose+index data phase that needs one. The EXPORT +# site is deliberately NOT built here: it is a static render OF a corpus, so +# there is nothing to render until the operator has one. docker/publish-site.sh +# builds it at run time and publishes it into the shared volume the `site` +# service serves. +RUN pnpm --filter homepage run build:nodata + +# --------------------------------------------------------------------------- +# runtime-base — everything the app needs EXCEPT a transcription backend. +# +# Shared by `runtime` and `runtime-vulkan` — the two differ only in RUNTIME_IMAGE +# (bookworm vs trixie) and which transcription backend they copy in. +# runtime-cuda cannot derive from it (it starts from an NVIDIA base image, not +# the node one) and so repeats it — that duplication is the price of a different +# base, not an oversight. +# --------------------------------------------------------------------------- +FROM ${RUNTIME_IMAGE} AS runtime-base + +# ffmpeg/ffprobe: audio extraction and the download-time duration guard. +# zip/tar/xz/gzip: the export build's archive formats. +# rsync: the saved-video backup mirror. +# curl: model fetch + the compose healthcheck. +# git: the "cut release" flow commits, and yt-dlp resolves some extractors better +# with it present. +# procps: `ps`, so "is the transcriber actually running in there?" is answerable +# from a `docker compose exec` shell. debian:slim ships without it. +RUN apt-get update \ + && apt-get install -y --no-install-recommends \ + ffmpeg zip unzip tar xz-utils gzip rsync curl ca-certificates git procps \ + && rm -rf /var/lib/apt/lists/* + +# yt-dlp as the standalone release binary (it bundles its own python), installed +# writable so `yt-dlp -U` works — see YTDLP_AUTO_UPDATE in docker/entrypoint.sh. +# A pinned yt-dlp goes stale fast, and a stale yt-dlp is the single most common +# reason downloads start failing. +ARG TARGETARCH=amd64 +RUN set -eux; \ + case "${TARGETARCH}" in \ + amd64) asset=yt-dlp_linux ;; \ + arm64) asset=yt-dlp_linux_aarch64 ;; \ + *) echo "unsupported TARGETARCH=${TARGETARCH}" >&2; exit 1 ;; \ + esac; \ + curl -fsSL "https://github.com/yt-dlp/yt-dlp/releases/latest/download/${asset}" \ + -o /usr/local/bin/yt-dlp; \ + chmod 0755 /usr/local/bin/yt-dlp; \ + /usr/local/bin/yt-dlp --version + +# A JavaScript runtime for yt-dlp. +# +# Not optional any more, and easy to miss because it degrades rather than fails: +# without one, yt-dlp prints "No supported JavaScript runtime could be found ... +# some formats may be missing" and carries on with a reduced format list. deno is +# the one it enables by default. Measured inside this image before it was added — +# every YouTube extraction warned. +ARG DENO_VERSION=v2.9.5 +RUN set -eux; \ + case "${TARGETARCH}" in \ + amd64) deno_asset=deno-x86_64-unknown-linux-gnu.zip ;; \ + arm64) deno_asset=deno-aarch64-unknown-linux-gnu.zip ;; \ + *) echo "unsupported TARGETARCH=${TARGETARCH}" >&2; exit 1 ;; \ + esac; \ + curl -fsSL "https://github.com/denoland/deno/releases/download/${DENO_VERSION}/${deno_asset}" \ + -o /tmp/deno.zip; \ + unzip -q /tmp/deno.zip -d /usr/local/bin; \ + rm /tmp/deno.zip; \ + chmod 0755 /usr/local/bin/deno; \ + /usr/local/bin/deno --version + +RUN npm install -g pnpm@9.15.4 + +# WORKDIR matters: findMonorepoRoot() (common/lib/paths.ts) walks UP from cwd +# looking for pnpm-workspace.yaml, and everything under it — the export staging +# dirs, the changelog paths, chart-templates.json — hangs off what it finds. +WORKDIR /repo +COPY --from=build /repo /repo + +# Data lives in volumes, never in the image. See docker-compose.yml. +ENV NODE_ENV=production \ + NEXT_TELEMETRY_DISABLED=1 \ + TRANSCRIPTS_DIR=/data/transcripts \ + SETTINGS_FILE=/data/config/settings.json \ + YTDLP_BIN=/usr/local/bin/yt-dlp \ + EXPORT_INDEX_DIR=/data/builds/.export-index \ + EXPORT_BUILDS_DIR=/data/builds/.export-builds \ + ARCHILYZER_SITE_OUT=/data/builds/site +RUN mkdir -p /data/transcripts /data/config /data/models /data/builds + +ENTRYPOINT ["/repo/docker/entrypoint.sh"] +CMD ["editor"] + +# --------------------------------------------------------------------------- +# runtime — the published default target. CPU whisper.cpp. +# --------------------------------------------------------------------------- +FROM runtime-base AS runtime +COPY --from=whisper-cpu /usr/local/bin/whisper-cli /usr/local/bin/whisper-cli +ENV ARCHILYZER_TRANSCRIBER=whisper-cpp \ + WHISPER_BIN=/usr/local/bin/whisper-cli \ + WHISPER_MODEL=/data/models/ggml-base.en.bin + +# --------------------------------------------------------------------------- +# runtime-vulkan — parakeet.cpp on any Vulkan GPU. docker-compose.vulkan.yml. +# +# Ships whisper-cli too: the GPU is for parakeet, but a CPU whisper worker is a +# useful thing to have in the same image — to compare against, or for when the +# GPU is busy. +# --------------------------------------------------------------------------- +FROM runtime-base AS runtime-vulkan + +# libvulkan1 is the loader; mesa-vulkan-drivers supplies the ICDs that actually +# talk to the hardware (RADV for AMD, ANV for Intel). vulkan-tools is here so +# `vulkaninfo` can answer "does the container see the GPU at all?", which is the +# first question every time this does not work. +RUN apt-get update \ + && apt-get install -y --no-install-recommends \ + libvulkan1 mesa-vulkan-drivers vulkan-tools \ + && rm -rf /var/lib/apt/lists/* + +COPY --from=parakeet-vulkan /out/parakeet-cli /usr/local/bin/parakeet-cli +COPY --from=parakeet-vulkan /out/lib/ /usr/local/lib/ +COPY --from=whisper-cpu /usr/local/bin/whisper-cli /usr/local/bin/whisper-cli +RUN ldconfig + +ENV ARCHILYZER_TRANSCRIBER=parakeet \ + PARAKEET_CLI=/usr/local/bin/parakeet-cli \ + PARAKEET_MODEL=/data/models/tdt_ctc-110m-q8_0.gguf \ + WHISPER_BIN=/usr/local/bin/whisper-cli \ + WHISPER_MODEL=/data/models/ggml-base.en.bin + +# --------------------------------------------------------------------------- +# runtime-cuda — whisper.cpp on CUDA. NVIDIA only. docker-compose.gpu.yml. +# --------------------------------------------------------------------------- +FROM ${CUDA_RUNTIME_IMAGE} AS runtime-cuda + +ARG NODE_MAJOR=20 +RUN apt-get update \ + && apt-get install -y --no-install-recommends \ + ffmpeg zip unzip tar xz-utils gzip rsync curl ca-certificates git gnupg procps \ + && curl -fsSL "https://deb.nodesource.com/setup_${NODE_MAJOR}.x" | bash - \ + && apt-get install -y --no-install-recommends nodejs \ + && rm -rf /var/lib/apt/lists/* + +ARG TARGETARCH=amd64 +RUN set -eux; \ + case "${TARGETARCH}" in \ + amd64) asset=yt-dlp_linux ;; \ + arm64) asset=yt-dlp_linux_aarch64 ;; \ + *) echo "unsupported TARGETARCH=${TARGETARCH}" >&2; exit 1 ;; \ + esac; \ + curl -fsSL "https://github.com/yt-dlp/yt-dlp/releases/latest/download/${asset}" \ + -o /usr/local/bin/yt-dlp; \ + chmod 0755 /usr/local/bin/yt-dlp; \ + /usr/local/bin/yt-dlp --version + +# A JavaScript runtime for yt-dlp. +# +# Not optional any more, and easy to miss because it degrades rather than fails: +# without one, yt-dlp prints "No supported JavaScript runtime could be found ... +# some formats may be missing" and carries on with a reduced format list. deno is +# the one it enables by default. Measured inside this image before it was added — +# every YouTube extraction warned. +ARG DENO_VERSION=v2.9.5 +RUN set -eux; \ + case "${TARGETARCH}" in \ + amd64) deno_asset=deno-x86_64-unknown-linux-gnu.zip ;; \ + arm64) deno_asset=deno-aarch64-unknown-linux-gnu.zip ;; \ + *) echo "unsupported TARGETARCH=${TARGETARCH}" >&2; exit 1 ;; \ + esac; \ + curl -fsSL "https://github.com/denoland/deno/releases/download/${DENO_VERSION}/${deno_asset}" \ + -o /tmp/deno.zip; \ + unzip -q /tmp/deno.zip -d /usr/local/bin; \ + rm /tmp/deno.zip; \ + chmod 0755 /usr/local/bin/deno; \ + /usr/local/bin/deno --version + +RUN npm install -g pnpm@9.15.4 + +# The CUDA whisper build is dynamically linked, so its ggml/whisper libraries +# come along with it. +COPY --from=whisper-cuda /out/whisper-cli /usr/local/bin/whisper-cli +COPY --from=whisper-cuda /out/lib/ /usr/local/lib/ +RUN ldconfig + +WORKDIR /repo +COPY --from=build /repo /repo + +ENV NODE_ENV=production \ + NEXT_TELEMETRY_DISABLED=1 \ + TRANSCRIPTS_DIR=/data/transcripts \ + SETTINGS_FILE=/data/config/settings.json \ + ARCHILYZER_TRANSCRIBER=whisper-cpp \ + WHISPER_BIN=/usr/local/bin/whisper-cli \ + WHISPER_MODEL=/data/models/ggml-base.en.bin \ + YTDLP_BIN=/usr/local/bin/yt-dlp \ + EXPORT_INDEX_DIR=/data/builds/.export-index \ + EXPORT_BUILDS_DIR=/data/builds/.export-builds \ + ARCHILYZER_SITE_OUT=/data/builds/site +RUN mkdir -p /data/transcripts /data/config /data/models /data/builds + +ENTRYPOINT ["/repo/docker/entrypoint.sh"] +CMD ["editor"] diff --git a/README.md b/README.md @@ -143,7 +143,7 @@ Needed only for the pipeline — install what you will actually use: | ffmpeg + ffprobe | Audio transcode and duration checks. | | A transcription backend | Channels you transcribe yourself (whisper.cpp by default; `chough` and `parakeet.cpp` also supported). | | rsync | Backing up the saved-video store, if you enable it. | -| Docker or podman | Only for the parallel multi-site export build. | +| Docker or podman | Running the whole stack in containers ([RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md)), and the parallel multi-site export build. | **The apps run with none of the second table installed.** You just cannot fetch anything yet, which makes it easy to try the UI first and commit to the toolchain @@ -173,8 +173,29 @@ Homebrew's `whisper-cpp` saves you building the default transcription backend. ### Windows -**Use WSL2.** The transcription backends and several helper scripts are Unix-oriented, -so the reliable path is a Linux userland on your Windows box: +**To run an archive, use Docker Desktop.** One command, no Linux userland to set up +by hand, and nothing published beyond your own machine: + +```powershell +git clone <your-remote> archilyzer +cd archilyzer +copy .env.example .env +docker compose up -d --build +``` + +Then open **http://localhost:8081**. The first build compiles whisper.cpp and both +Next.js apps, so it takes a while; after that, starting is seconds. On first boot it +seeds a settings file with a working transcription worker and downloads a whisper +model, so the editor is usable the moment it comes up. + +**[RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md)** covers the rest: the exposure model +(every port binds `127.0.0.1` by default, including the public ones), how to serve +your archive to the world without exposing the admin app, GPU transcription, and +backup. The same stack runs on Linux and macOS. + +**To develop the project — or to run Claude Code against it — use WSL2 instead.** The +transcription backends and several helper scripts are Unix-oriented, so the reliable +path for working *on* the code is a Linux userland: ```powershell wsl --install -d Ubuntu @@ -227,12 +248,12 @@ Two Windows-specific things to keep in mind: here: `pnpm install`, file watching and the corpus I/O all get dramatically slower across the Windows filesystem boundary. -> **Status: a first-class container image is not shipped yet.** The two Dockerfiles in -> this repo are for the parallel multi-site export build (`Dockerfile.build`) and for -> sharded e2e runs (`Dockerfile.test`) — neither runs the editor as a long-lived -> server. A single `docker compose up` that stands up an archive on Docker Desktop is -> the intended next step for exactly the case above; until it lands, WSL2 is the -> Windows path. +> **Which of the two do you want?** The container stack *runs* an archive: it is the +> shortest path from a clone to a working editor, and it is what to reach for if you +> want the thing rather than the toolchain. WSL2 gives you the repo itself — a dev +> server with hot reload, the e2e suite, the MCP server, and Claude Code sitting in +> the same filesystem as the corpus. Running both is fine; they share nothing but the +> source. --- @@ -463,6 +484,7 @@ See [CONTRIBUTING.md](CONTRIBUTING.md) to work on the code. |---|---| | [SETUP.md](SETUP.md) | Full per-OS install, every environment variable, transcription backends. | | [CONTRIBUTING.md](CONTRIBUTING.md) | Workspace layout, tests, CLI shims, internals. | +| [RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md) | `docker compose up` for the whole stack: exposure model, auth, GPU. | | [DEPLOY_CLOUDFLARE.md](DEPLOY_CLOUDFLARE.md) | Pages + R2, and cost-abuse protection. | | [DEPLOY_DOCKER.md](DEPLOY_DOCKER.md) | Parallel multi-site export builds in containers. | | [SCHEDULED_SYNC.md](SCHEDULED_SYNC.md) | Unattended per-channel syncing. | diff --git a/RUNNING_IN_DOCKER.md b/RUNNING_IN_DOCKER.md @@ -0,0 +1,454 @@ +# Running an archive in Docker + +One command stands up a working archive server: the editor, the tools it drives +(yt-dlp, ffmpeg, whisper.cpp), and a reverse proxy that is the only thing on the +box with an open port. + +> This document is about **running the apps**. [DEPLOY_DOCKER.md](DEPLOY_DOCKER.md) +> is a different thing entirely — it is about fanning per-site *export builds* out +> across containers, and it uses `Dockerfile.build`. Neither file affects the +> other. + +--- + +## Quickstart + +```sh +git clone <your-remote> archilyzer +cd archilyzer +cp .env.example .env +docker compose up -d --build +``` + +Then open **http://localhost:8081**. + +The first build takes a while — it compiles whisper.cpp and builds two Next.js +apps. Afterwards, `docker compose up -d` is seconds. + +On first boot the container also: + +- creates the corpus, config, models and builds volumes; +- writes a `settings.json` **with one enabled whisper.cpp worker** (a settings + file with zero workers means zero transcription slots, and auto-transcribe + would report `no-workers` and quietly do nothing); +- downloads the `base.en` whisper model (~142 MB) into the models volume. + +Watch it happen with `docker compose logs -f editor`. + +### What is running + +| | | | +|---|---|---| +| **editor** | http://localhost:8081 | the admin app: channels, downloads, transcription, builds | +| site | http://localhost:8080 | the published archive (profile `site`) | +| homepage | http://localhost:8082 | the project's own site (profile `homepage`) | +| umtool | http://localhost:8083 | the clip/report bench (profile `umtool`) | + +Only the editor and Caddy run by default. Somebody who wants an archive runs two +containers, not five: + +```sh +docker compose --profile site up -d # + the archive server +docker compose --profile umtool up -d # + umtool +``` + +--- + +## The exposure model + +**Read this before you change a bind address.** + +The editor has **no authentication of any kind**. It shells out to yt-dlp, +deletes media, streams job logs, and rewrites your corpus. It was built to run on +a machine you are sitting at. + +So the stack is arranged to make exposing it a decision rather than an accident: + +1. **No application container publishes a port.** The editor, the site server, + the homepage and umtool sit on an internal Docker network. Nothing on your + host can reach them directly. This removes the whole class of "I published + 3001 once to try something and forgot". + +2. **Caddy is the single front door**, and every port it publishes binds + `127.0.0.1` **by default — the public sites included**. Nothing is reachable + off the box until you say so. + +3. **A startup rail.** If a private app is bound to anything but loopback while + nothing is checking credentials, the containers refuse to start and tell you + the four ways to fix it. + +| App | Internal port | Caddy port | Default bind | Auth | +|---|---|---|---|---| +| editor | 3001 | 8081 | `127.0.0.1` | `basic_auth` | +| export site | 3000 | 8080 | `127.0.0.1` | none — public by design | +| homepage | 3031 | 8082 | `127.0.0.1` | none | +| umtool | 3050 | 8083 | `127.0.0.1` | `basic_auth` | + +### Serving your archive to the world + +The archive is static files with no write path. Publishing it is the point: + +```sh +# .env +SITE_BIND=0.0.0.0 +``` + +```sh +docker compose --profile site up -d +``` + +Nothing else becomes reachable. The editor stays on loopback. + +### Reaching the editor from another machine + +Pick one. In rough order of how much you should like it: + +**Tailscale — exposes nothing.** Install Tailscale on the host, leave every bind +at `127.0.0.1`, and run `tailscale serve 8081`. The editor is reachable from your +own devices and from nowhere else, with no port open on any router and no +password to leak. This is the recommended answer for a machine at home. + +**An SSH tunnel** — the same idea with no new software: + +```sh +ssh -N -L 8081:127.0.0.1:8081 you@archive-box +``` + +**A password.** Caddy's `basic_auth` is built in, so this works with nothing +installed and cannot break on your first evening: + +```sh +docker run --rm caddy:2.11-alpine caddy hash-password --plaintext 'your-password' +``` + +```sh +# .env — the $ characters need no escaping here +EDITOR_BIND=0.0.0.0 +ARCHILYZER_AUTH_USER=you +ARCHILYZER_AUTH_HASH=$2a$14$... +``` + +Basic auth is a password in a browser dialog sent on every request. It is real +protection over a trusted network and nothing more; put HTTPS in front before +it crosses one you don't trust. + +**A real identity provider** — sessions, TOTP, SSO. Two drop-in overlays ship +here, and neither is required: + +```sh +docker compose -f docker-compose.yml -f docker-compose.tinyauth.yml up -d +docker compose -f docker-compose.yml -f docker-compose.authelia.yml up -d +``` + +*Tinyauth* is configured entirely by environment variables, is under 10 MB, and +has been OpenID Certified since v5.1.0. It is AGPL-3.0, and its config keys churn +between major releases — the overlay pins the tag for exactly that reason, so +read the release notes before moving it. + +*Authelia* is Apache-2.0 and the established forward-auth standard, at the cost +of a YAML config file and a mandatory database. Templates are in +`docker/authelia/`; copy the two `.example` files and fill them in. + +Both attach through Caddy's `forward_auth`. The overlays set +`ARCHILYZER_AUTH_MODE=forward` for you. + +### The rail, and its escape hatch + +``` +EDITOR_BIND=0.0.0.0 exposes the editor beyond this machine, and nothing is +checking credentials in front of it. +``` + +If you genuinely have your own auth in front — a reverse proxy you trust, a +tailnet-only address — say so explicitly: + +```sh +ARCHILYZER_AUTH_MODE=none +``` + +That is the only thing that disables the check. It is deliberately a line you +have to write. + +--- + +## Everyday operation + +### Add a channel and get transcripts + +Everything happens in the editor at http://localhost:8081 — this is the same app +a host install runs, so [README.md](README.md) and [SETUP.md](SETUP.md) describe +it accurately. + +1. **Channels → Add**, paste a channel URL. +2. **Store playlist**, then **Sync**. Media and any published captions land in + the `corpus` volume. +3. **Transcribe** for videos with no captions. This uses the seeded whisper.cpp + worker and the model in the `models` volume. + +### Publish the archive + +The site server serves a *static build of your corpus*, which does not exist +until you make one: + +```sh +docker compose exec editor /repo/docker/publish-site.sh <site-id> +docker compose --profile site up -d +``` + +Create the site itself in the editor first (**Sites → New**); run the script with +no argument to list the ones you have. It runs the same `pnpm --filter export run +build` a host install runs, then copies the output into the volume the `site` +service serves. No restart needed afterwards. + +### Model choice + +`ARCHILYZER_FETCH_MODEL` picks what gets downloaded on boot. For the default and +CUDA images that is a whisper model: `tiny.en` (~75 MB), +`base.en` (~142 MB), `small.en` (~466 MB), `medium.en` (~1.5 GB), `large-v3` +(~3 GB). Bigger is better and slower, and on CPU the difference is large. +Changing it later fetches the new model but does **not** switch the worker over — +model choice lives on the editor's Workers page, which is the right place for it. + +For the Vulkan image it names a parakeet GGUF instead — see +[GPU transcription](#gpu-transcription). + +`ARCHILYZER_FETCH_MODEL=none` skips the download entirely. + +### Keeping yt-dlp current + +An image pins yt-dlp at build time, and a stale yt-dlp is the most common reason +downloads suddenly start failing. Set `YTDLP_AUTO_UPDATE=1` in `.env` and each +boot self-updates it. Or, once: + +```sh +docker compose exec editor yt-dlp -U +``` + +### Booting without resuming work + +`editor/instrumentation.ts` arms the sync heartbeat, both auto-queue runners and +the digest and backfill sweeps on every boot. That is right for a host install, +where a restart interrupts work you own. It is wrong the first time you point a +container at somebody else's corpus: its stored policies may say "sweep", and a +corpus-wide digest sweep is GPU-*weeks*. + +```sh +ARCHILYZER_IDLE_BOOT=1 +``` + +Boots the server with all of that stopped. You can start any of it from the UI +afterwards. The shutdown reaper and the persisted-pause restore stay armed +either way — both only ever *stop* work. + +### GPU transcription + +Two overlays, because two different engines get the GPU. Neither is required — +the default image transcribes on the CPU and works everywhere. + +**Vulkan + parakeet.cpp — the one that runs on any GPU.** + +```sh +docker compose -f docker-compose.yml -f docker-compose.vulkan.yml up -d --build +``` + +Vulkan rather than a vendor SDK, so the same image drives AMD (RADV), Intel +(ANV) and NVIDIA. On AMD it is the only GPU path here that works at all: ROCm is +not packaged and CUDA is NVIDIA-only. The engine is +[parakeet.cpp](https://github.com/mudler/parakeet.cpp), driven by the +overlapping-window wrapper the app already ships +(`scripts/parakeet-stitch.mjs`), and the first boot seeds a **parakeet** worker +and fetches a parakeet GGUF instead of a whisper model. + +All the host has to provide is a render node: + +```sh +ls /dev/dri # renderD128, card0, … +vulkaninfo --summary | head # names your GPU +``` + +The overlay passes `/dev/dri` straight through — no container toolkit, no +runtime shim, no privileged mode. + +**Check that the container actually sees it.** This is the one thing worth +verifying, because the failure is silent: with no render node inside, parakeet +falls back to the CPU and produces perfectly correct transcripts an order of +magnitude slower, and nothing raises an error. The entrypoint prints a +`vulkan:` line on every boot saying which it got, and the image ships +`vulkaninfo` so you can ask directly: + +```sh +docker compose -f docker-compose.yml -f docker-compose.vulkan.yml \ + exec editor vulkaninfo --summary | head -20 +``` + +Pin a device with `PARAKEET_DEVICE` in `.env` — `Vulkan0`/`Vulkan1` for a +specific GPU, or `cpu` to take the GPU out of the picture with everything else +identical. + +That last one is also the **surest test that the GPU is really being used**, +because it needs no log parsing: run the same audio both ways and compare. On an +RX 6600 XT, one 33-second clip through the app's own wrapper: + +``` +default (Vulkan) 3.4 s +--device cpu 36.3 s +``` + +A ~10× gap means the GPU is doing the work. No gap means you are on the CPU +whatever anything else claims. + +Models live at [mudler/parakeet-cpp-gguf](https://huggingface.co/mudler/parakeet-cpp-gguf); +`ARCHILYZER_FETCH_MODEL` names one without the `.gguf` (default +`tdt_ctc-110m-q8_0` — small and quick; `tdt-0.6b-v3-q5_k` and `tdt_ctc-1.1b-q5_k` +are better and slower). The image also carries CPU `whisper-cli`, so you can add +a whisper worker on the Workers page to compare without rebuilding. + +**CUDA + whisper.cpp — NVIDIA only.** + +```sh +docker compose -f docker-compose.yml -f docker-compose.gpu.yml up -d --build +``` + +Needs the NVIDIA Container Toolkit on the host. Verify it first: + +```sh +docker run --rm --gpus all nvidia/cuda:12.6.3-base-ubuntu22.04 nvidia-smi +``` + +This builds a second whisper.cpp with CUDA offload and pulls a CUDA base image — +several GB, and a long first build. Two build args take the edge off it: + +```sh +# one nvcc per job, and each is hungry — lower it on a shared or small machine +--build-arg WHISPER_BUILD_JOBS=4 +# every CUDA translation unit is compiled once per architecture; the default list +# spans Pascal..Hopper, a single value is much faster if you know your card +--build-arg CUDA_ARCHITECTURES=86 +``` + +Apple Metal is not covered: whisper.cpp supports it, but a Mac GPU is not +reachable from a Linux container at all, so it would have to be a host install. + +### Backup + +`corpus` is the volume that matters — real media, real transcripts, hundreds of +GB when it grows up. `config` holds `settings.json`. `models` and `builds` are +reproducible; losing them costs a download and a rebuild. + +```sh +docker run --rm -v archilyzer_corpus:/corpus -v "$PWD:/backup" \ + debian:bookworm-slim tar czf /backup/corpus.tar.gz -C /corpus . +``` + +### Useful commands + +```sh +docker compose logs -f editor +docker compose exec editor bash # a shell in the image +docker compose exec editor yt-dlp --version +docker compose restart editor +docker compose down # stop; volumes survive +docker compose down -v # stop AND DELETE the corpus +``` + +--- + +## What is in the image, and what is not + +**Baked in:** node + the installed workspace, the built editor / umtool / homepage, +`yt-dlp`, a JavaScript runtime for it (`deno` — without one yt-dlp warns and +silently loses formats), `ffmpeg`/`ffprobe`, `whisper-cli` (statically linked), +`zip`/`tar`/`xz`/`gzip`, `rsync`, `git`, `curl`, `ps`. The Vulkan image adds +`parakeet-cli`, the Mesa Vulkan drivers and `vulkaninfo`. + +**Not baked, on purpose:** + +- **The corpus.** It is a separate git repo holding production data, and it lives + in the `corpus` volume. +- **whisper models.** 142 MB to 3 GB, and the choice is yours. Fetched on boot. +- **The export site build.** It is a static render *of a corpus*, and there is no + corpus at image-build time. `docker/publish-site.sh` makes it at run time. +- **ImageMagick with Pango, and `qrencode`.** Only `scripts/report-to-video/` + needs them, and rendering a report to video is a workstation task, not + something a server does. Run that part on a host checkout. +- **gallery-dl.** The X/Twitter post fetcher. Install it and point + `GALLERY_DL_BIN` at it if you want the social corpus. +- **ollama / claude.** The digest lanes reach a local ollama or the `claude` CLI. + Point `OLLAMA_URL` at a host ollama (`http://host.docker.internal:11434`) if you + want digests. + +### The multi-site build pipeline falls back inside a container + +The editor can fan per-site export builds out across containers +(`buildPipeline.mode = "docker"`, see DEPLOY_DOCKER.md). Inside a container there +is no `docker` binary, so that path is unavailable. It already handles this — the +build logs + +``` +[notice] No container engine available — building sites serially on the host. +``` + +and does the builds one after another in-process. Correct, just not parallel. **Do +not** try to make docker-in-docker work for this; run the parallel pipeline from a +host checkout if you need it. + +### It runs as root + +Volume ownership across Docker Desktop, WSL2 and Linux is not worth the class of +bug that comes with getting it subtly wrong, and everything the container touches +is either a volume or the image itself. The apps are not reachable from the +network by default, which is the control that matters here. + +--- + +## Windows + +This is the reason the stack exists. Install **Docker Desktop** (which uses WSL2 +underneath, without you having to run anything inside it), then: + +```powershell +git clone <your-remote> archilyzer +cd archilyzer +copy .env.example .env +docker compose up -d --build +``` + +Open http://localhost:8081. Keep the clone on the Windows filesystem — the build +context is small, and the corpus lives in a Docker volume inside the WSL2 VM, +which is where the I/O actually happens. + +You still want WSL2 directly if you intend to *develop* the project or run Claude +Code against it — see [README.md](README.md#claude-code-on-windows). For running +an archive, this is enough. + +--- + +## Troubleshooting + +**`REFUSING TO START`** — a private app is bound off-loopback with no auth. The +message lists the four fixes; see [the rail](#the-rail-and-its-escape-hatch). Both +the app and Caddy refuse, and `restart: unless-stopped` means they keep retrying: +`docker compose ps` shows them `Restarting (1)` until you fix `.env`. Nothing is +listening on the exposed port while that is true. + +**A port answers 502.** Its service is not running. `site`, `homepage` and +`umtool` are behind profiles; start the one you want. Caddy publishes all four +ports regardless, and this is what an unstarted one looks like. + +**Transcription fails with a missing model.** The first-boot download failed +(check `docker compose logs editor`). `docker compose restart editor` retries it. + +**Transcription is slow on the Vulkan image.** Almost certainly the silent CPU +fallback: check `docker compose logs editor | grep vulkan:` and the `--device cpu` +comparison above. The usual cause is a missing `/dev/dri` — a container without +the render node still transcribes, just slowly. + +**Downloads started failing on every channel.** Almost always a stale yt-dlp: +`docker compose exec editor yt-dlp -U`. + +**Port 8080 is already taken.** `SITE_HTTP_PORT=9080` in `.env`. Those variables +move the *outside* port; nothing inside the containers changes. + +**A build resumed by itself after `docker compose up`.** That is the boot +behaviour described under [booting without resuming +work](#booting-without-resuming-work) — set `ARCHILYZER_IDLE_BOOT=1`. diff --git a/common/lib/idleBoot.test.ts b/common/lib/idleBoot.test.ts @@ -0,0 +1,27 @@ +import { test } from "node:test"; +import assert from "node:assert/strict"; +import { IDLE_BOOT_ENV, isIdleBoot } from "./idleBoot"; + +// Run with: pnpm --filter yt-dlp-transcript-common exec tsx --test common/lib/idleBoot.test.ts + +test("an unset variable keeps the host default: arm everything", () => { + assert.equal(isIdleBoot({}), false); +}); + +test("the affirmative spellings all mean idle", () => { + for (const value of ["1", "true", "yes", "on", "TRUE", " On "]) { + assert.equal(isIdleBoot({ [IDLE_BOOT_ENV]: value }), true, value); + } +}); + +test("empty and negative values arm the runners", () => { + // A compose file writing an unset variable produces "", and that must not + // read as "idle" — it is the difference between auto-queue running and not. + for (const value of ["", "0", "false", "no", "off", " ", "maybe"]) { + assert.equal(isIdleBoot({ [IDLE_BOOT_ENV]: value }), false, JSON.stringify(value)); + } +}); + +test("only this one variable is consulted", () => { + assert.equal(isIdleBoot({ IDLE_BOOT: "1", ARCHILYZER_IDLE: "1" }), false); +}); diff --git a/common/lib/idleBoot.ts b/common/lib/idleBoot.ts @@ -0,0 +1,34 @@ +// "Boot idle": start the server WITHOUT resuming any of the long-running work +// a previous process was doing. +// +// editor/instrumentation.ts arms four things on every boot — the sync heartbeat, +// the auto-queue runners, the digest sweep and the backfill sweep. That is right +// for a host install, where a restart is an interruption of work you own. It is +// wrong the first time a CONTAINER is pointed at a corpus somebody else built: +// the policies stored in that corpus's settings.json say "yes, sweep", and the +// container obliges by starting GPU-weeks of work seconds after `docker compose +// up`, before the operator has seen a single page. +// +// So: an env switch, read once at boot. Unset (the host default) keeps the old +// behavior exactly. See RUNNING_IN_DOCKER.md. +// +// Deliberately dependency-free — instrumentation.ts is bundled for the Edge +// runtime as well as Node, and every import it makes is lazy for that reason. +// This module imports nothing and touches no Node API, so it is safe to import +// statically from there. + +export const IDLE_BOOT_ENV = "ARCHILYZER_IDLE_BOOT"; + +// Values that mean "yes". Anything else — including the empty string a compose +// file writes for an unset variable — means "no", so a half-configured .env +// cannot silently disable the runners. +const TRUTHY = new Set(["1", "true", "yes", "on"]); + +// True when this process should skip arming the boot-time runners. +export function isIdleBoot( + env: Record<string, string | undefined> = process.env, +): boolean { + const raw = env[IDLE_BOOT_ENV]; + if (typeof raw !== "string") return false; + return TRUTHY.has(raw.trim().toLowerCase()); +} diff --git a/docker-compose.authelia.yml b/docker-compose.authelia.yml @@ -0,0 +1,42 @@ +# Drop-in: Authelia in front of the private apps. +# +# cp docker/authelia/configuration.yml.example docker/authelia/configuration.yml +# cp docker/authelia/users_database.yml.example docker/authelia/users_database.yml +# # edit both, then: +# docker compose -f docker-compose.yml -f docker-compose.authelia.yml up -d +# +# Authelia is the established forward-auth standard and Apache-2.0 licensed. It +# is also the heavier of the two options on purpose: it wants a YAML config file +# and a database (SQLite is the smallest supported), so it is not a one-variable +# setup the way Tinyauth is. Pick it if you already run it, or if you want 2FA +# and per-app access rules that outlive this stack. +# +# Set in .env: +# +# ARCHILYZER_AUTH_MODE=forward +# +# The config file in docker/authelia/ is a STARTING POINT, not a hardened +# deployment: read Authelia's docs on session/domain settings before exposing +# any of this beyond your own machine. + +services: + authelia: + image: authelia/authelia:4.39 + restart: unless-stopped + networks: [archilyzer] + volumes: + - ./docker/authelia:/config + - authelia-data:/data + environment: + X_AUTHELIA_CONFIG: /config/configuration.yml + ports: + - "${AUTHELIA_BIND:-127.0.0.1}:${AUTHELIA_HTTP_PORT:-8085}:9091" + + caddy: + environment: + ARCHILYZER_AUTH_MODE: forward + ARCHILYZER_FORWARD_AUTH_UPSTREAM: authelia:9091 + ARCHILYZER_FORWARD_AUTH_URI: /api/authz/forward-auth + +volumes: + authelia-data: diff --git a/docker-compose.gpu.yml b/docker-compose.gpu.yml @@ -0,0 +1,44 @@ +# NVIDIA GPU overlay: whisper.cpp with CUDA offload instead of the CPU build. +# +# docker compose -f docker-compose.yml -f docker-compose.gpu.yml up -d --build +# +# Needs the NVIDIA Container Toolkit on the host (`nvidia-ctk`), which is what +# makes `deploy.resources.reservations.devices` work. Check it first: +# +# docker run --rm --gpus all nvidia/cuda:12.6.3-base-ubuntu22.04 nvidia-smi +# +# This builds a SECOND whisper.cpp (the runtime-cuda target) and pulls a CUDA +# base image — expect several GB and a long first build. The CPU image is +# untouched and remains the default; nothing here is required to run an archive. +# +# AMD/ROCm and Apple Metal are not covered: whisper.cpp supports both, but each +# needs its own base image and device plumbing. Build a variant of the +# `whisper-cpu` stage with the right -DGGML_* flag if you need one. + +services: + editor: + build: + target: runtime-cuda + image: archilyzer:${ARCHILYZER_TAG:-local}-cuda + deploy: + resources: + reservations: + devices: + - driver: nvidia + count: all + capabilities: [gpu] + + # The other three run the same image so `docker compose ... up` never rebuilds + # a second variant behind your back. None of them touches the GPU. + site: + build: + target: runtime-cuda + image: archilyzer:${ARCHILYZER_TAG:-local}-cuda + homepage: + build: + target: runtime-cuda + image: archilyzer:${ARCHILYZER_TAG:-local}-cuda + umtool: + build: + target: runtime-cuda + image: archilyzer:${ARCHILYZER_TAG:-local}-cuda diff --git a/docker-compose.tinyauth.yml b/docker-compose.tinyauth.yml @@ -0,0 +1,47 @@ +# Drop-in: real sessions (and TOTP, and OAuth) in front of the private apps, +# instead of a browser basic-auth box. +# +# docker compose -f docker-compose.yml -f docker-compose.tinyauth.yml up -d +# +# Tinyauth is the small end of this: configured entirely by environment +# variables, under 10 MB, and OpenID Certified as of v5.1.0. Two caveats, both +# real: +# +# * AGPL-3.0. Fine for running it; read it before you build on it. +# * Its configuration keys churn between major releases. The tag below is +# PINNED for that reason — check the release notes before moving it, and +# expect to touch the environment block when you do. +# +# Set in .env first: +# +# ARCHILYZER_AUTH_MODE=forward +# TINYAUTH_SECRET=<32+ random chars> +# TINYAUTH_USERS=you:$$2a$$14$$... # htpasswd-style bcrypt, see below +# TINYAUTH_APP_URL=http://localhost:8084 # where YOU reach the login page +# +# Generate the user entry: +# docker run --rm ghcr.io/steveiliop56/tinyauth:v5.1.0 user create --interactive +# +# The login page needs to be reachable from the browser, so it is the one extra +# published port here. It binds loopback like everything else. + +services: + tinyauth: + image: ghcr.io/steveiliop56/tinyauth:v5.1.0 + restart: unless-stopped + networks: [archilyzer] + env_file: + - path: .env + required: false + environment: + APP_URL: ${TINYAUTH_APP_URL:-http://localhost:8084} + USERS: ${TINYAUTH_USERS:-} + SECRET: ${TINYAUTH_SECRET:-} + ports: + - "${TINYAUTH_BIND:-127.0.0.1}:${TINYAUTH_HTTP_PORT:-8084}:3000" + + caddy: + environment: + ARCHILYZER_AUTH_MODE: forward + ARCHILYZER_FORWARD_AUTH_UPSTREAM: tinyauth:3000 + ARCHILYZER_FORWARD_AUTH_URI: /api/auth/caddy diff --git a/docker-compose.vulkan.yml b/docker-compose.vulkan.yml @@ -0,0 +1,75 @@ +# GPU overlay: parakeet.cpp on Vulkan instead of CPU whisper.cpp. +# +# docker compose -f docker-compose.yml -f docker-compose.vulkan.yml up -d --build +# +# Vulkan, not a vendor SDK, so one image covers AMD (RADV), Intel (ANV) and +# NVIDIA. On AMD this is the only GPU path that works at all — ROCm is not +# packaged here and CUDA is NVIDIA-only. +# +# What the host needs: a working Vulkan driver and a render node at /dev/dri. +# Check before you build — it takes a second and saves a long compile: +# +# ls /dev/dri # renderD128, card0, … +# vulkaninfo --summary | head # names your GPU +# +# Then confirm the CONTAINER sees it (the image ships vulkan-tools for exactly +# this, and the entrypoint prints a line about it on every boot): +# +# docker compose -f docker-compose.yml -f docker-compose.vulkan.yml \ +# exec editor vulkaninfo --summary | head -20 +# +# If /dev/dri is missing inside, parakeet quietly runs on the CPU — correct +# output, an order of magnitude slower, and nothing says so except that boot +# line. That is the failure mode to watch for. +# +# Model: the first boot fetches a parakeet GGUF into the models volume +# (ARCHILYZER_FETCH_MODEL, default tdt_ctc-110m-q8_0 — small and quick). The +# whole family is at https://huggingface.co/mudler/parakeet-cpp-gguf; the bigger +# tdt-0.6b-v3 and tdt_ctc-1.1b models are better and slower. +# +# The image also carries CPU whisper-cli, so you can add a whisper worker on the +# Workers page to compare without rebuilding anything. + +services: + editor: + build: + target: runtime-vulkan + args: + # The Vulkan stack needs a newer Debian than the default image runs on; + # the Dockerfile explains why, and why the workspace is still BUILT on + # the older one. + RUNTIME_IMAGE: node:20-trixie-slim + image: archilyzer:${ARCHILYZER_TAG:-local}-vulkan + devices: + # The render node. This is the whole GPU passthrough — no toolkit, no + # runtime shim, no privileged mode. The container runs as root, so the + # host's ownership of renderD128 (usually root:render) needs no group + # mapping on top. + - /dev/dri:/dev/dri + # To pin which device parakeet uses, put PARAKEET_DEVICE in .env — every + # service already reads that file. "Vulkan0"/"Vulkan1" select a specific GPU + # (`vulkaninfo --summary` lists them in order); "cpu" takes the GPU out of + # the picture with everything else identical, which is the honest way to + # measure what the GPU is buying you. Unset — the default — lets + # parakeet.cpp choose. + + # The other three run the same image so `up` never builds a second variant + # behind your back. None of them touches the GPU. + site: + build: + target: runtime-vulkan + args: + RUNTIME_IMAGE: node:20-trixie-slim + image: archilyzer:${ARCHILYZER_TAG:-local}-vulkan + homepage: + build: + target: runtime-vulkan + args: + RUNTIME_IMAGE: node:20-trixie-slim + image: archilyzer:${ARCHILYZER_TAG:-local}-vulkan + umtool: + build: + target: runtime-vulkan + args: + RUNTIME_IMAGE: node:20-trixie-slim + image: archilyzer:${ARCHILYZER_TAG:-local}-vulkan diff --git a/docker-compose.yml b/docker-compose.yml @@ -0,0 +1,178 @@ +# `docker compose up -d` stands up an archive. +# +# Two containers by default: caddy (the only thing that publishes a port) and +# editor (the admin app). The three other apps are behind profiles, so somebody +# who only wants an archive runs two containers, not five: +# +# docker compose up -d editor + caddy +# docker compose --profile site up -d + the published archive +# docker compose --profile site --profile homepage up -d +# docker compose --profile umtool up -d +# +# THE EXPOSURE MODEL, which is the load-bearing part of this file: +# +# * No application container publishes a port. They talk to caddy over an +# internal network and are unreachable from the host except through it. +# * Caddy publishes four ports and EVERY ONE binds 127.0.0.1 by default — +# the public sites included. Nothing is reachable off this box until you +# change a bind in .env. +# * The editor has no authentication of its own. Binding it to anything but +# loopback without an auth layer is refused at startup, loudly. See +# docker/guard-exposure.sh. +# +# See RUNNING_IN_DOCKER.md. Copy .env.example to .env before your first run. + +name: archilyzer + +# --------------------------------------------------------------------------- +# Every app runs the same image; the command selects which one. +# --------------------------------------------------------------------------- +x-app: &app + build: + context: . + target: runtime + image: archilyzer:${ARCHILYZER_TAG:-local} + restart: unless-stopped + networks: [archilyzer] + # env_file rather than `environment:` for the auth variables, and this is not + # cosmetic: a bcrypt hash is full of `$`, and compose interpolates `${...}` in + # the YAML but passes env_file values through verbatim. Put the hash in .env + # and it arrives intact, with nothing to escape. + env_file: + - path: .env + required: false + +x-app-env: &app-env + TRANSCRIPTS_DIR: /data/transcripts + SETTINGS_FILE: /data/config/settings.json + # NOTHING ABOUT THE TRANSCRIPTION ENGINE BELONGS HERE. + # + # Which engine, which binary, which model file and which model to fetch are + # properties of the IMAGE (each runtime target sets them), and the entrypoint + # defaults the download per engine. Naming a whisper model here looks harmless + # and quietly breaks the Vulkan image, which wanted a parakeet GGUF and got + # sent to fetch "base.en.gguf" — measured, on the first boot of that image. + # Override any of them from .env, which every service already reads. + # + # Build staging and the published site live in a volume, not in the container's + # writable layer — a full build's staging is tens of GB. + EXPORT_INDEX_DIR: /data/builds/.export-index + EXPORT_BUILDS_DIR: /data/builds/.export-builds + ARCHILYZER_SITE_OUT: /data/builds/site + # Set explicitly so a stray EDITOR_PORT/UMTOOL_PORT in .env cannot move an app + # off the port Caddy proxies to. These are INTERNAL ports; the published ones + # are on the caddy service below. + EDITOR_PORT: "3001" + UMTOOL_PORT: "3050" + +services: + # ------------------------------------------------------------------------- + # The front door. The only service with `ports:`. + # ------------------------------------------------------------------------- + caddy: + image: caddy:2.11-alpine + restart: unless-stopped + networks: [archilyzer] + env_file: + - path: .env + required: false + # Runs the same exposure rail the app entrypoint does — this is the process + # that actually opens the ports, so it refuses first. + entrypoint: ["/bin/sh", "/etc/archilyzer/caddy-start.sh"] + ports: + # published-bind : published-port : caddy-port + - "${SITE_BIND:-127.0.0.1}:${SITE_HTTP_PORT:-8080}:8080" + - "${EDITOR_BIND:-127.0.0.1}:${EDITOR_HTTP_PORT:-8081}:8081" + - "${HOMEPAGE_BIND:-127.0.0.1}:${HOMEPAGE_HTTP_PORT:-8082}:8082" + - "${UMTOOL_BIND:-127.0.0.1}:${UMTOOL_HTTP_PORT:-8083}:8083" + volumes: + - ./docker/Caddyfile:/etc/caddy/Caddyfile:ro + - ./docker:/etc/archilyzer:ro + - caddy-data:/data + - caddy-config:/config + + # ------------------------------------------------------------------------- + # The admin app. Downloads, transcribes, builds, deletes. No login of its own. + # ------------------------------------------------------------------------- + editor: + <<: *app + command: ["editor"] + environment: + <<: *app-env + volumes: + - corpus:/data/transcripts + - config:/data/config + - models:/data/models + - builds:/data/builds + healthcheck: + # /api/pulse answers from in-memory state plus two stat() calls — it is + # built to be polled forever, which is exactly what a healthcheck does. + test: ["CMD", "curl", "-fsS", "-o", "/dev/null", "http://127.0.0.1:3001/api/pulse"] + interval: 15s + timeout: 10s + retries: 10 + start_period: 40s + # Long enough for armShutdownCancel() to reap a running yt-dlp or whisper + # child rather than have it killed out from under a half-written file. + stop_grace_period: 30s + + # ------------------------------------------------------------------------- + # The published archive: static files, built by docker/publish-site.sh. + # ------------------------------------------------------------------------- + site: + <<: *app + profiles: [site] + command: ["site"] + environment: + <<: *app-env + volumes: + # Writable, and only for one reason: before anything has been published, + # the entrypoint drops a placeholder page into the site directory that says + # how to publish. Serving that beats serving a 404 nobody can interpret. + # Publishing later needs no restart — `serve` reads from disk per request. + - builds:/data/builds + + # ------------------------------------------------------------------------- + # The project's own site (marketing + docs). Baked into the image. + # ------------------------------------------------------------------------- + homepage: + <<: *app + profiles: [homepage] + command: ["homepage"] + environment: + <<: *app-env + + # ------------------------------------------------------------------------- + # umtool: the clip/report bench. Private, like the editor. + # ------------------------------------------------------------------------- + umtool: + <<: *app + profiles: [umtool] + command: ["umtool"] + environment: + <<: *app-env + volumes: + - corpus:/data/transcripts + - config:/data/config + - builds:/data/builds + +networks: + archilyzer: + # A normal bridge, deliberately NOT `internal: true` — yt-dlp needs the + # internet. What keeps these containers private is that none of them + # publishes a port, not that they are cut off from the network. + driver: bridge + +volumes: + # The corpus. This is the one that matters: real media, real transcripts, + # hundreds of GB when it grows up. Back it up. + corpus: + # settings.json, and anything else operational that must outlive the container. + config: + # whisper .bin models — 142 MB to 3 GB, fetched once. + models: + # Export build staging + the published static site. Reproducible; losing it + # costs a rebuild, not data. + builds: + caddy-data: + caddy-config: diff --git a/docker/Caddyfile b/docker/Caddyfile @@ -0,0 +1,84 @@ +# The single front door for the whole stack. +# +# NO application container publishes a port. The editor has no authentication of +# any kind, shells out to yt-dlp, and deletes files — so it sits on an internal +# docker network and is reachable only through here. That removes the entire +# class of "I published 3001 once to try something and forgot". +# +# Caddy publishes one port per app, and docker-compose.yml binds every one of +# them to 127.0.0.1 by default — the public sites included. Nothing is reachable +# off the box until you change a bind in .env. +# +# 8080 export site -> site:3000 public by design, no auth +# 8081 editor -> editor:3001 PRIVATE: admin, no auth of its own +# 8082 homepage -> homepage:3031 public by design, no auth +# 8083 umtool -> umtool:3050 PRIVATE +# +# A port whose service isn't running (site/homepage/umtool are behind compose +# profiles) answers 502. That is expected, not a misconfiguration. +# +# Auth is pluggable and applies to the two PRIVATE apps. docker/caddy-start.sh +# picks the mode and expands {$ARCHILYZER_AUTH_IMPORT} to one of the snippets +# below — or to nothing at all. See RUNNING_IN_DOCKER.md. + +{ + # No admin API: it is an unauthenticated control plane by default, and + # nothing here needs it. + admin off + # These are plain HTTP ports behind whatever you put in front of them (or + # behind nothing, on loopback). Certificate management is not this file's job. + auto_https off + log { + output stdout + format console + } +} + +# --- auth modes ------------------------------------------------------------ +# Built into Caddy: nothing to install, nothing to keep running, and it cannot +# break on somebody's first evening. Generate the hash with: +# docker run --rm caddy:2.11-alpine caddy hash-password --plaintext 'secret' +(basicauth) { + basic_auth { + {$ARCHILYZER_AUTH_USER:archilyzer} {$ARCHILYZER_AUTH_HASH} + } +} + +# Hand the decision to an external identity provider (Tinyauth, Authelia, …). +# See docker-compose.tinyauth.yml and docker-compose.authelia.yml. +(forwardauth) { + forward_auth {$ARCHILYZER_FORWARD_AUTH_UPSTREAM} { + uri {$ARCHILYZER_FORWARD_AUTH_URI:/api/auth/caddy} + copy_headers Remote-User Remote-Groups Remote-Name Remote-Email + } +} + +# --- the four apps --------------------------------------------------------- + +# The export site: a static archive. Public is the whole point of it. +:8080 { + reverse_proxy site:3000 +} + +# The editor. PRIVATE — this is the admin app. +:8081 { + {$ARCHILYZER_AUTH_IMPORT} + # Server actions and log streams; the editor streams job output for as long + # as a job runs, which is longer than any sensible proxy default. + reverse_proxy editor:3001 { + flush_interval -1 + } +} + +# The project's own site (marketing + docs). Public. +:8082 { + reverse_proxy homepage:3031 +} + +# umtool. PRIVATE. +:8083 { + {$ARCHILYZER_AUTH_IMPORT} + reverse_proxy umtool:3050 { + flush_interval -1 + } +} diff --git a/docker/authelia/configuration.yml.example b/docker/authelia/configuration.yml.example @@ -0,0 +1,47 @@ +# A STARTING POINT for Authelia in front of the private apps. Copy to +# configuration.yml, replace every <...>, and read Authelia's docs on session +# cookies and domains before exposing any of this beyond this machine. +# +# Generate each secret with: openssl rand -hex 32 + +theme: dark + +server: + address: tcp://0.0.0.0:9091 + +log: + level: info + +identity_validation: + reset_password: + jwt_secret: <random-hex-32> + +authentication_backend: + file: + path: /config/users_database.yml + +# Everything behind this stack is admin surface, so the default is the strict +# one. Loosen deliberately, per rule, not by changing this line. +access_control: + default_policy: two_factor + +session: + secret: <random-hex-32> + cookies: + # `domain` must match the hostname you actually type in the browser, and + # `authelia_url` must be reachable from that browser — the login page is + # a redirect target, not an internal call. + - domain: localhost + authelia_url: http://localhost:8085 + default_redirection_url: http://localhost:8081 + +storage: + encryption_key: <random-hex-32> + local: + path: /data/db.sqlite3 + +# Writes "emails" to a file. Enough to complete a first-time 2FA registration +# without configuring SMTP; swap for a real notifier when you have one. +notifier: + filesystem: + filename: /data/notification.txt diff --git a/docker/authelia/users_database.yml.example b/docker/authelia/users_database.yml.example @@ -0,0 +1,13 @@ +# Copy to users_database.yml. Generate the hash with: +# +# docker run --rm authelia/authelia:4.39 authelia crypto hash generate argon2 \ +# --password 'your-password' + +users: + you: + disabled: false + displayname: "You" + password: "<the $argon2id$... string it printed>" + email: you@example.com + groups: + - admins diff --git a/docker/caddy-start.sh b/docker/caddy-start.sh @@ -0,0 +1,47 @@ +#!/bin/sh +# Entrypoint for the caddy service. Runs the exposure rail, decides which auth +# snippet the Caddyfile imports, then execs caddy. +# +# The Caddyfile cannot decide this for itself: there is no "apply this directive +# only if an env var is non-empty" in Caddyfile syntax. What there IS is env +# substitution over the raw file before it is parsed, and that substitution can +# expand to more than one token — so {$ARCHILYZER_AUTH_IMPORT} becomes either +# `import basicauth`, `import forwardauth`, or nothing at all. +# +# POSIX sh: caddy:alpine ships busybox. +set -eu + +/etc/archilyzer/guard-exposure.sh + +case "${ARCHILYZER_AUTH_MODE:-basic}" in +none) + ARCHILYZER_AUTH_IMPORT="" + echo "[caddy] auth: none (ARCHILYZER_AUTH_MODE=none)" + ;; +forward) + if [ -z "${ARCHILYZER_FORWARD_AUTH_UPSTREAM:-}" ]; then + echo "[caddy] ARCHILYZER_AUTH_MODE=forward needs ARCHILYZER_FORWARD_AUTH_UPSTREAM" >&2 + exit 1 + fi + ARCHILYZER_AUTH_IMPORT="import forwardauth" + echo "[caddy] auth: forward_auth -> ${ARCHILYZER_FORWARD_AUTH_UPSTREAM}" + ;; +basic) + if [ -n "${ARCHILYZER_AUTH_HASH:-}" ]; then + ARCHILYZER_AUTH_IMPORT="import basicauth" + echo "[caddy] auth: basic_auth as '${ARCHILYZER_AUTH_USER:-archilyzer}'" + else + # Allowed only because the rail above already proved every private app + # is on loopback. + ARCHILYZER_AUTH_IMPORT="" + echo "[caddy] auth: none — no ARCHILYZER_AUTH_HASH set, admin apps are loopback-only" + fi + ;; +*) + echo "[caddy] unknown ARCHILYZER_AUTH_MODE='${ARCHILYZER_AUTH_MODE}' (basic|forward|none)" >&2 + exit 1 + ;; +esac +export ARCHILYZER_AUTH_IMPORT + +exec caddy run --config /etc/caddy/Caddyfile --adapter caddyfile diff --git a/docker/entrypoint.sh b/docker/entrypoint.sh @@ -0,0 +1,286 @@ +#!/usr/bin/env bash +# First run, every run: prepare the data volumes, then hand the container over to +# one of the four apps. +# +# editor the admin app (next start, :3001) +# site the published archive (serve, :3000) +# homepage the project's own site (serve, :3031) +# umtool the clip/report bench (next start, :3050) +# shell drop into bash — for `docker compose run --rm editor shell` +# +# Nothing here publishes a port. Caddy does that; see docker/Caddyfile. +# +# The app is exec'd, deliberately: armShutdownCancel() in editor/instrumentation.ts +# reaps in-flight yt-dlp/whisper children on SIGTERM, and it can only see the +# signal if the app is PID 1's own process rather than a grandchild of a wrapper. +set -euo pipefail + +APP="${1:-editor}" + +TRANSCRIPTS_DIR="${TRANSCRIPTS_DIR:-/data/transcripts}" +SETTINGS_FILE="${SETTINGS_FILE:-/data/config/settings.json}" +MODELS_DIR="${ARCHILYZER_MODELS_DIR:-/data/models}" +BUILDS_DIR="${ARCHILYZER_BUILDS_DIR:-/data/builds}" +SITE_OUT="${ARCHILYZER_SITE_OUT:-${BUILDS_DIR}/site}" +WHISPER_MODEL="${WHISPER_MODEL:-${MODELS_DIR}/ggml-base.en.bin}" +# Which transcription backend THIS IMAGE was built with. Set by the Dockerfile +# per target, not by the operator: `runtime`/`runtime-cuda` carry whisper.cpp, +# `runtime-vulkan` carries parakeet.cpp on the GPU. It decides which worker the +# seed below writes and which model family gets fetched. +TRANSCRIBER="${ARCHILYZER_TRANSCRIBER:-whisper-cpp}" +export TRANSCRIPTS_DIR SETTINGS_FILE WHISPER_MODEL + +log() { printf '[entrypoint] %s\n' "$*"; } +die() { printf '[entrypoint] %s\n' "$*" >&2; exit 1; } + +# --------------------------------------------------------------------------- +# 1. The data volumes. +# --------------------------------------------------------------------------- +mkdir -p \ + "${TRANSCRIPTS_DIR}/channels" \ + "${TRANSCRIPTS_DIR}/sites" \ + "$(dirname "${SETTINGS_FILE}")" \ + "${MODELS_DIR}" \ + "${BUILDS_DIR}" \ + "${SITE_OUT}" + +# --------------------------------------------------------------------------- +# 2. Seed settings.json — with a worker. +# +# defaults() in common/lib/settings.ts returns `workers: []`, and zero workers +# means zero transcription slots: auto-transcribe reports `no-workers` and does +# nothing at all, silently. A fresh container that looks healthy and transcribes +# nothing is the worst possible first run, so the seed carries exactly one local +# whisper.cpp worker. +# +# Everything else is left out on purpose. getSettings() merges a partial file +# over defaults(), so a short seed is a FEATURE: keys we don't write here keep +# tracking the app's own defaults as those move, instead of being frozen at +# whatever they were the day this image was built. +# +# The worker's config block is empty for the same reason — an empty `bin`/`model` +# falls back to WHISPER_BIN / WHISPER_MODEL, which the image already sets. One +# source of truth for where the binary and the model are. +# --------------------------------------------------------------------------- +if [ ! -f "${SETTINGS_FILE}" ]; then + case "${TRANSCRIBER}" in + parakeet) + # The GPU image. `device` is left unset so parakeet.cpp picks the best + # backend it can see — Vulkan when /dev/dri is passed through, CPU when + # it is not, rather than failing outright. Name a device explicitly + # (Vulkan0, cpu, …) on the Workers page when you want to pin it. + cat >"${SETTINGS_FILE}" <<'JSON' +{ + "adminTitle": "Archilyzer", + "transcriptionApp": "parakeet", + "workers": [ + { + "id": "parakeet", + "name": "parakeet (GPU)", + "kind": "local", + "enabled": true, + "priority": 0, + "appId": "parakeet", + "config": {} + } + ], + "parallelTranscriptions": 1 +} +JSON + log "seeded ${SETTINGS_FILE} with one local parakeet worker" + ;; + *) + cat >"${SETTINGS_FILE}" <<'JSON' +{ + "adminTitle": "Archilyzer", + "transcriptionApp": "whisper-cpp", + "workers": [ + { + "id": "whisper-cpp", + "name": "whisper.cpp", + "kind": "local", + "enabled": true, + "priority": 0, + "appId": "whisper-cpp", + "config": {} + } + ], + "parallelTranscriptions": 1 +} +JSON + log "seeded ${SETTINGS_FILE} with one local whisper.cpp worker" + ;; + esac +fi + +# --------------------------------------------------------------------------- +# 3. The speech model. +# +# Not baked into the image: 142 MB for whisper base.en, 3 GB for large-v3, and +# the choice is the operator's. Fetched once into the models volume. Set +# ARCHILYZER_FETCH_MODEL= (empty) or =none to skip entirely — e.g. when you mount +# your own models directory. +# +# Two families, because the two images have different engines: +# whisper-cpp ggml-<name>.bin from ggerganov/whisper.cpp (base.en, small.en, …) +# parakeet <name>.gguf from mudler/parakeet-cpp-gguf (tdt_ctc-110m-q8_0, …) +# --------------------------------------------------------------------------- +fetch_model() { + local name target url + if [ "${TRANSCRIBER}" = "parakeet" ]; then + name="${ARCHILYZER_FETCH_MODEL-tdt_ctc-110m-q8_0}" + target="${MODELS_DIR}/${name}.gguf" + url="https://huggingface.co/mudler/parakeet-cpp-gguf/resolve/main/${name}.gguf" + else + name="${ARCHILYZER_FETCH_MODEL-base.en}" + target="${MODELS_DIR}/ggml-${name}.bin" + url="https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-${name}.bin" + fi + [ -z "${name}" ] && return 0 + [ "${name}" = "none" ] && return 0 + + if [ -s "${target}" ]; then + log "speech model present: ${target}" + return 0 + fi + + log "fetching ${TRANSCRIBER} model '${name}' -> ${target}" + log " (this is a one-time download; set ARCHILYZER_FETCH_MODEL=none to skip)" + # Download to a temp name and move: a container killed mid-download must not + # leave a truncated file that looks like a model forever after. + if curl -fL --retry 3 --retry-delay 2 -o "${target}.partial" "${url}"; then + mv "${target}.partial" "${target}" + log "model ready: ${target} ($(du -h "${target}" | cut -f1))" + else + rm -f "${target}.partial" + # Not fatal. The editor starts fine without a model; transcription is + # what fails, and it fails with a message naming the missing file. + log "WARNING: model download failed — transcription will not work until" + log " ${target} exists. Retry with: docker compose restart editor" + fi +} + +# --------------------------------------------------------------------------- +# 4. yt-dlp self-update. +# +# An image pins yt-dlp at build time, and a stale yt-dlp is the most common +# reason downloads suddenly start failing — YouTube changes, yt-dlp ships a fix +# within days, and an image from last month has none of them. Opt-in because it +# is a network call on every boot. +# --------------------------------------------------------------------------- +update_ytdlp() { + case "${YTDLP_AUTO_UPDATE:-0}" in + 1 | true | yes | on) ;; + *) return 0 ;; + esac + log "yt-dlp -U (YTDLP_AUTO_UPDATE is set)" + "${YTDLP_BIN:-yt-dlp}" -U || log "WARNING: yt-dlp self-update failed — continuing with $("${YTDLP_BIN:-yt-dlp}" --version 2>/dev/null || echo unknown)" +} + +# --------------------------------------------------------------------------- +# 5. The exposure rail. Shared with the caddy container — one implementation. +# --------------------------------------------------------------------------- +/repo/docker/guard-exposure.sh + +# --------------------------------------------------------------------------- +# 6. Hand over. +# --------------------------------------------------------------------------- +next_bin() { + local dir="$1" + [ -x "${dir}/node_modules/.bin/next" ] || die "next not found in ${dir} — was the image built?" + printf '%s' "${dir}/node_modules/.bin/next" +} + +serve_bin() { + local dir="$1" + [ -x "${dir}/node_modules/.bin/serve" ] || die "serve not found in ${dir} — was the image built?" + printf '%s' "${dir}/node_modules/.bin/serve" +} + +case "${APP}" in +editor) + fetch_model + update_ytdlp + log "corpus: ${TRANSCRIPTS_DIR}" + log "settings: ${SETTINGS_FILE}" + if [ "${TRANSCRIBER}" = "parakeet" ]; then + log "engine: parakeet ${PARAKEET_CLI:-parakeet-cli} (model ${PARAKEET_MODEL:-unset})" + # One line that answers "is the GPU actually visible in here?" before a + # transcription has to answer it the slow way. + if [ -e /dev/dri ] && command -v vulkaninfo >/dev/null 2>&1; then + log "vulkan: $(vulkaninfo --summary 2>/dev/null | grep -m2 -E 'deviceName' | sed 's/^[[:space:]]*//' | tr '\n' ' ' || echo 'no device reported')" + else + log "vulkan: NO /dev/dri IN THIS CONTAINER — parakeet will run on CPU." + log " Pass the GPU through: see docker-compose.vulkan.yml." + fi + else + log "engine: whisper ${WHISPER_BIN:-whisper-cli} (model ${WHISPER_MODEL})" + fi + if [ "${ARCHILYZER_IDLE_BOOT:-}" = "1" ]; then + log "idle boot: schedulers, auto-queue runners and sweeps stay STOPPED" + fi + cd /repo/editor + exec "$(next_bin /repo/editor)" start --port "${EDITOR_PORT:-3001}" + ;; + +site) + # The published archive. Built at RUN time, not baked: it is a static render + # of a corpus, and there is no corpus in the image. docker/publish-site.sh + # fills this directory; until it has, serve a page that says so rather than + # a 404 nobody can interpret. + if [ -z "$(ls -A "${SITE_OUT}" 2>/dev/null)" ]; then + log "no site published yet at ${SITE_OUT} — serving a placeholder" + cat >"${SITE_OUT}/index.html" <<'HTML' +<!doctype html> +<meta charset="utf-8"> +<title>No site published yet</title> +<style> + body { font: 16px/1.6 system-ui, sans-serif; max-width: 42rem; margin: 12vh auto; padding: 0 1.5rem; } + code { background: #1112; padding: .15em .4em; border-radius: .3em; } + pre { background: #1112; padding: 1rem; border-radius: .5em; overflow-x: auto; } +</style> +<h1>No site published yet</h1> +<p>This is the static archive server. It has nothing to serve because no site has +been built into its volume yet.</p> +<p>Add a channel in the editor, sync it, then publish:</p> +<pre>docker compose exec editor /repo/docker/publish-site.sh &lt;site-id&gt;</pre> +<p>See <code>RUNNING_IN_DOCKER.md</code>.</p> +HTML + fi + log "serving ${SITE_OUT} on :3000" + # -c: the project's OWN serve config — `access-control-allow-origin: *` and a + # cache lifetime on the JSON, which is what makes /corpus.json readable by the + # MCP server and federatable by a hub. A published archive gets that from the + # _headers file compose-site.ts emits, and compose-site.ts says in as many + # words that `serve` ignores _headers and reads serve.json instead. It looks + # for one in the SERVED directory, which is a volume holding build output, so + # the path has to be given. (`serve --cors` looks like the answer and does not + # set the header — measured.) + # --no-port-switching: if :3000 were taken, serve would silently move to + # another port and Caddy would 502 at a healthy-looking container. + exec "$(serve_bin /repo/export)" "${SITE_OUT}" -l 3000 \ + -c /repo/export/serve.json --no-clipboard --no-port-switching + ;; + +homepage) + # Corpus-independent, so this one IS baked into the image at build time. + log "serving /repo/homepage/out on :3031" + exec "$(serve_bin /repo/homepage)" /repo/homepage/out -l 3031 \ + -c /repo/homepage/serve.json --no-clipboard --no-port-switching + ;; + +umtool) + cd /repo/umtool + exec "$(next_bin /repo/umtool)" start --port "${UMTOOL_PORT:-3050}" + ;; + +shell) + exec bash + ;; + +*) + # Anything else is run verbatim, so `docker compose run --rm editor yt-dlp --version` + # works without a special case per tool. + exec "$@" + ;; +esac diff --git a/docker/guard-exposure.sh b/docker/guard-exposure.sh @@ -0,0 +1,96 @@ +#!/bin/sh +# The safety rail: refuse to run with an unauthenticated ADMIN app bound to +# anything but loopback. +# +# The editor has no authentication of its own. It shells out to yt-dlp, deletes +# media, and streams job logs. umtool is the same shape. Exposing either of them +# to a network without an auth layer in front should take deliberate effort, not +# a typo in a bind address — so this runs both in the app entrypoint and in the +# Caddy container (the one that actually publishes the ports). +# +# POSIX sh: this has to run under the busybox shell in caddy:alpine as well as +# under bash in the app image. No arrays, no [[ ]], no local. +# +# Exit 0 = safe to start. Exit 1 = refused, with the reason on stderr. +set -eu + +# Loopback, or unset (compose's default is 127.0.0.1 — an unset value here means +# the variable was never passed, not that the app is on 0.0.0.0). +is_loopback() { + case "$1" in + "" | localhost | ::1 | "[::1]") return 0 ;; + 127.*) return 0 ;; + *) return 1 ;; + esac +} + +auth_mode="${ARCHILYZER_AUTH_MODE:-basic}" + +# Is the admin apps' front door accounted for? +# +# "none" counts. It is not "no protection is needed" — it is the operator saying +# in writing that they have their own (a proxy they trust, a tailnet-only +# address). That assertion is the entire point of the escape hatch, so honour it +# here rather than making it the one setting that cannot be used. +auth_accounted_for() { + case "${auth_mode}" in + none) return 0 ;; + forward) [ -n "${ARCHILYZER_FORWARD_AUTH_UPSTREAM:-}" ] && return 0 || return 1 ;; + *) [ -n "${ARCHILYZER_AUTH_HASH:-}" ] && return 0 || return 1 ;; + esac +} + +refuse() { + app="$1" + bind="$2" + var="$3" + cat >&2 <<EOF + + ┌─ REFUSING TO START ────────────────────────────────────────────────────┐ + + ${var}=${bind} exposes the ${app} beyond this machine, and nothing is + checking credentials in front of it. + + The ${app} is an ADMIN app with no login of its own: anyone who can reach + it can download, delete and re-transcribe anything in your corpus. + + Pick one: + + 1. Put a password on it (the built-in option, nothing to install): + + docker run --rm caddy:2.11-alpine caddy hash-password --plaintext 'your-password' + + then in .env: + + ARCHILYZER_AUTH_USER=you + ARCHILYZER_AUTH_HASH=<the hash it printed> + + 2. Leave it on loopback and reach it over Tailscale or an SSH tunnel: + + ${var}=127.0.0.1 + ssh -N -L 8081:127.0.0.1:8081 you@this-machine + + 3. Put a real identity provider in front of it — see the forward-auth + overlays in RUNNING_IN_DOCKER.md (ARCHILYZER_AUTH_MODE=forward). + + 4. You genuinely have your own auth in front (a reverse proxy you trust, + a tailnet-only address). Say so explicitly: + + ARCHILYZER_AUTH_MODE=none + + └────────────────────────────────────────────────────────────────────────┘ + +EOF + exit 1 +} + +# Only the two PRIVATE apps are guarded. The export site and the homepage are +# public archives; serving them to the world is what they are for. +check() { + is_loopback "$2" && return 0 + auth_accounted_for && return 0 + refuse "$1" "$2" "$3" +} + +check "editor" "${EDITOR_BIND:-}" "EDITOR_BIND" +check "umtool" "${UMTOOL_BIND:-}" "UMTOOL_BIND" diff --git a/docker/publish-site.sh b/docker/publish-site.sh @@ -0,0 +1,50 @@ +#!/usr/bin/env bash +# Build one site's static archive and publish it into the volume the `site` +# service serves. +# +# docker compose exec editor /repo/docker/publish-site.sh <site-id> +# +# Why this exists rather than a baked build: the export site is a static render +# OF A CORPUS, and there is no corpus at image-build time — so the image ships +# the code and this publishes the output. It is the same `pnpm --filter export +# run build` a host install runs; the only container-specific part is the copy +# at the end. +# +# The copy is a copy and not a symlink on purpose. `next build` REMOVES and +# recreates export/out (docker/build-site.sh has the same note), so a symlink +# there survives exactly one build, and a bind mount there breaks the build +# outright — the rm fails on a busy mount point. +set -euo pipefail + +SITE_ID="${1:-${SITE_ID:-}}" +SITE_OUT="${ARCHILYZER_SITE_OUT:-/data/builds/site}" + +if [ -z "${SITE_ID}" ]; then + echo "usage: publish-site.sh <site-id>" >&2 + echo >&2 + echo "configured sites:" >&2 + ls -1 "${TRANSCRIPTS_DIR:-/data/transcripts}/sites" 2>/dev/null | + grep -v '^_' | sed 's/^/ /' >&2 || + echo " (none yet — create one in the editor under Sites)" >&2 + exit 2 +fi + +export SITE_ID +cd /repo + +echo "[publish-site] building '${SITE_ID}'" +# The full build: prebuild (shared LMDB index + stats + chart templates), then +# compose:site for THIS site, then next build. +pnpm --filter export run build + +[ -d /repo/export/out ] || { echo "[publish-site] no export/out after build" >&2; exit 1; } + +echo "[publish-site] publishing -> ${SITE_OUT}" +mkdir -p "${SITE_OUT}" +# Empty the directory's CONTENTS, never the directory: it is a volume mount +# point, and removing it fails. +find "${SITE_OUT}" -mindepth 1 -maxdepth 1 -exec rm -rf {} + +cp -a /repo/export/out/. "${SITE_OUT}/" + +echo "[publish-site] done — $(find "${SITE_OUT}" -type f | wc -l) files" +echo "[publish-site] the site service serves it immediately; no restart needed." diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md @@ -1,6 +1,10 @@ # Changelog ## [Unreleased] +- **`docker compose up -d` now stands up a working archive.** The repo had two Dockerfiles and neither ran the app: one fans per-site export builds out across containers, the other runs sharded e2e. So the only way to host this was to install the whole Unix toolchain by hand, which is why the Windows instructions said "use WSL2 and follow the Linux steps". There is now a runtime image and a compose stack — the editor plus Caddy by default, with the published site, the project homepage and umtool behind compose **profiles**, so somebody who only wants an archive runs two containers rather than five. First boot creates the volumes, downloads a speech model, and seeds a `settings.json` **carrying one enabled worker**: the defaults ship `workers: []`, zero workers means zero transcription slots, and a fresh container that looks healthy and silently transcribes nothing is the worst possible first run. Two things are deliberately not baked into the image and cannot be: the corpus, and the export site — that site is a static render *of* a corpus, and there is no corpus at image-build time, so `docker/publish-site.sh` builds it at run time into the volume the `site` service serves. +- **Nothing the container runs is reachable from outside the machine until you say so.** The editor has no authentication of any kind, shells out to yt-dlp, and deletes media — so **no application container publishes a port at all**. Caddy is the single front door, and every one of the four ports it publishes binds `127.0.0.1` by default, the public sites included; opening one is a deliberate edit of a single line in `.env`. Because a bind address is exactly the sort of thing that gets changed in a hurry, there is also a rail: if a private app is bound off-loopback with nothing checking credentials, **the containers refuse to start** — both the app and Caddy, which is the process that actually opens the ports — and print the four ways to fix it. `basic_auth` is built into Caddy so a password needs nothing installed; Tinyauth and Authelia attach through `forward_auth` as documented drop-in overlays; `ARCHILYZER_AUTH_MODE=none` is the one explicit escape hatch for people who already have their own front door. +- **The GPU is usable from the container, including on AMD.** Alongside the default CPU whisper.cpp image there is a **Vulkan** target running parakeet.cpp — one build that covers AMD (RADV), Intel and NVIDIA, needing nothing on the host but a render node at `/dev/dri` and no vendor container toolkit — and a CUDA target for NVIDIA whisper. Measured on an RX 6600 XT, through the app's own overlapping-window wrapper: **3.4 s against 36.3 s** for the same 33-second clip pinned to the CPU. That gap is also the thing to watch for, because the failure here is silent — a Vulkan container with no `/dev/dri` does not error, it transcribes correctly on the CPU about ten times slower. The entrypoint prints which one it got on every boot, and the image ships `vulkaninfo` so you can ask directly. +- **The editor can be started without resuming whatever it was in the middle of.** Booting arms the sync heartbeat, both auto-queue runners and the digest and backfill sweeps — right for a host install, where a restart interrupts work you own, and wrong the first time a container is pointed at a corpus somebody else configured: its stored policies may say *sweep*, and a corpus-wide digest sweep is GPU-**weeks** that would start seconds after `docker compose up`. `ARCHILYZER_IDLE_BOOT=1` starts the server with all of that stopped, and you can start any of it from the UI afterwards. The shutdown reaper and the persisted-pause restore stay armed either way, because both only ever *stop* work. - **The auto-queue can be told to do the newest uploads first, and it now genuinely does.** Both runners always took the first video off a rule's pile, and that pile's order came straight from the channel snapshot, which sorts most buckets **alphabetically by video id** — arbitrary for YouTube ids, and oldest-first for the date-prefixed folder names some sites use. So when a channel uploaded today, nothing made that video jump the nine-thousand-video backlog in front of it; the only reason auto-download roughly worked was that its one bucket happens to be left in playlist order. There is now an **Order** setting per runner — *Listed order* (what you have today, and still the default), *Newest first*, *Oldest first*. It sorts the videos **inside** each rule, across every channel and bucket that rule claims; the rule list still decides which rule goes first, because that is what the rule list is for. For a straight newest-first archive, use one catch-all rule. Worth knowing before you switch it on: under *Newest first* the retry and partial-download buckets lose their head start, so a half-finished download can end up waiting behind fresh work. The page says so next to the setting. - **Working out how recent 79,000 videos are turned out to be nearly free, once we stopped guessing where the dates were.** The obvious source — reading each video's metadata file — is about six and a half minutes and 41 GB of reading, on every scheduling decision, which is a non-starter. The transcript index already holds a date per video in a form that can be scanned without decoding anything: **78,583 videos in well under a second**. That covers the corpus, but it turned out **not** to cover the videos auto-transcribe actually queues, because the index only holds videos that already *have* a transcript and auto-transcribe's whole job is the ones that don't — of the 870 videos genuinely pending here, it knew the date of **122**. The gap is closed by reading the last 8 KB of each remaining video's metadata file, where the upload date happens to sit: **868 of the 870, at a fifth of a millisecond each**, and remembered afterwards so it is paid once rather than every few seconds. Videos not downloaded yet have no date anywhere on disk at all, so auto-download estimates one from the video's position in the channel's newest-first listing; those show with a `≈`, and a brand-new upload with nothing dated above it goes to the front, which is the entire point. Anything still undatable sorts to the back rather than disappearing, and a missing or busy index degrades to the old ordering instead of stopping the runner. - **The auto-queue page now answers "what is it doing, and why not?".** It was two identical panels of pending counts and a pick log — and the payload it was already receiving contained the answers to both questions, thrown away on arrival. A **dispatch deck** across the top gives both runners at a glance, so you never scroll to find out about the other one. Each lane then reads top to bottom as the questions you actually arrive with. **Next up** names the exact video the policy would hand out next, with its upload date, the rule that claimed it, and *why that rule* — "rule 1 has no pending work" — which is also the fastest way to see that an ordering change did what you asked. **In flight** lists what is running right now with how long each has been going, which the page received and rendered as a bare count. The policy tree became a **claim ladder**: one rung per rule, numbered by its real priority, saying in plain language what it matches, with a hairline down its left edge whose fill shows how much of the runner that rung is currently holding. The pending count on each rung opens to show the actual videos, in the exact order they will be handed out. This folds three previously separate sections — the tree, "Pending per rule" and "Recent picks", which you had to cross-reference by eye — into one object. The lane header gains how long the runner has been up and a link to its job log, both of which were on the wire and discarded. diff --git a/editor/instrumentation.ts b/editor/instrumentation.ts @@ -3,6 +3,15 @@ // heartbeat so auto-sync can run without an external cron job. // // See editor/app/scheduler/heartbeat.ts and SCHEDULED_SYNC.md. +// +// Everything armed here except the shutdown reaper is skipped when the process +// boots idle (ARCHILYZER_IDLE_BOOT) — see isIdleBoot below. +// +// The one STATIC import in this file, and safe as one because idleBoot.ts +// imports nothing and touches no Node API: the Edge bundle's static Node-API +// scan has nothing to object to. Every other import stays lazy. +import { isIdleBoot } from "yt-dlp-transcript-common/lib/idleBoot"; + export async function register() { // register() is called in every runtime (Node.js and Edge). The heartbeat and // its transitive imports (runTick -> server actions, lmdb, fs) are Node-only, @@ -26,11 +35,23 @@ export async function register() { /* failing to arm the reaper must not block server readiness */ } - // Lazy import inside the guard keeps server-only code out of the Edge bundle. - // startSyncHeartbeat only arms a timer (no synchronous tick), so it never - // blocks the server from becoming ready. - const { startSyncHeartbeat } = await import("./app/scheduler/heartbeat"); - startSyncHeartbeat(); + // Boot idle: the operator asked for a server, not for whatever the corpus's + // stored policies were in the middle of. Pointing a fresh container at an + // existing corpus would otherwise resume a GPU-weeks digest sweep seconds + // after `docker compose up`. See common/lib/idleBoot.ts. + // + // Two things stay armed regardless, because both only ever STOP work: the + // shutdown reaper above (a server that cannot reap its children is never + // correct) and the persisted-pause restore below. + const idle = isIdleBoot(); + + if (!idle) { + // Lazy import inside the guard keeps server-only code out of the Edge bundle. + // startSyncHeartbeat only arms a timer (no synchronous tick), so it never + // blocks the server from becoming ready. + const { startSyncHeartbeat } = await import("./app/scheduler/heartbeat"); + startSyncHeartbeat(); + } // Restore a persisted transcription pause: if the operator paused // transcriptions and the server later restarted, re-pause the worker pool so @@ -51,6 +72,9 @@ export async function register() { /* a failed re-pause must not block server readiness */ } + // Everything past here STARTS work. On an idle boot, nothing does. + if (idle) return; + // Start the automatic priority-queue runners (auto-transcribe / auto-download) // if their policies are enabled. Each is a self-managed registry job; this only // kicks them off and returns. Best-effort — a failure here must not stop the