commit 685a84c8e0e13f334d705a82f21d446e8933e204
parent 114b64e7cdc9e5d72e7f3801f9dee8fdc8865094
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Fri, 21 Aug 2026 08:53:43 -0400
Merge branch 'docs/windows-container-path'
Diffstat:
3 files changed, 67 insertions(+), 8 deletions(-)
diff --git a/SETUP.md b/SETUP.md
@@ -46,7 +46,7 @@ homepage):
| **ffmpeg** + **ffprobe** | Audio transcode + duration checks for `transcribe` channels. | `ffmpeg` / `ffprobe` on `PATH` |
| A **transcription backend** | `handling: "transcribe"` channels only. Default is **whisper.cpp** (`whisper-cli`); `chough` and `parakeet.cpp` are alternatives. | `whisper-cli` on `PATH` |
| **rsync** | Backing up the saved-video store. | `rsync` on `PATH` |
-| **Docker** | The sharded parallel e2e run (`pnpm e2e:sharded`) only. | — |
+| **Docker** | Running the whole stack in containers ([RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md)), the parallel multi-site build ([DEPLOY_DOCKER.md](DEPLOY_DOCKER.md)), and the sharded e2e run (`pnpm e2e:sharded`). | — |
Every binary above is overridable by an environment variable (e.g. `YTDLP_BIN`) —
see [Configuration & environment variables](#configuration--environment-variables).
@@ -107,8 +107,32 @@ corepack enable
### Windows
-**Recommended: use WSL2.** The transcription toolchain (whisper.cpp / parakeet.cpp)
-and helper shell scripts assume a Unix shell, so the smoothest path on Windows is
+Two paths, and they answer different questions.
+
+**To run an archive: Docker Desktop.** The container image already contains
+yt-dlp, ffmpeg and a transcription backend, so none of the toolchain in the tables
+above has to be installed on the host at all. Get the source first
+([Get the code running](#get-the-code-running) below), then, from the repo root:
+
+```powershell
+copy .env.example .env
+docker compose up -d --build
+```
+
+Then open **http://localhost:8081**. The first build compiles whisper.cpp and two
+Next.js apps and is slow; afterwards, starting is seconds. First boot fetches a
+speech model and seeds a `settings.json` with one enabled worker, so transcription
+works immediately instead of silently doing nothing.
+
+Every port binds `127.0.0.1` by default — including the published site — and the
+containers **refuse to start** if the editor is exposed to a network with no auth
+in front of it. See [RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md) for the exposure
+model, the auth options, GPU transcription (Vulkan/parakeet or CUDA/whisper) and
+backups. The same stack runs on Linux and macOS.
+
+**To develop, or to run the e2e suite: WSL2.** The transcription toolchain
+(whisper.cpp / parakeet.cpp) and helper shell scripts assume a Unix shell, so the
+smoothest path for working *on* the code is
[WSL2](https://learn.microsoft.com/windows/wsl/install):
```powershell
@@ -133,7 +157,8 @@ corepack enable
Caveats on native Windows:
- Building the **whisper.cpp / parakeet.cpp** transcription backends is Unix-oriented;
- do transcription work under WSL2.
+ do transcription work under WSL2, or let the container do it (it ships a backend
+ already built).
- `create-archives.sh` and the external-cron sync path (`pnpm sync:tick` from cron)
assume a Unix shell — use WSL2 or Task Scheduler equivalents.
diff --git a/homepage/content/README.md b/homepage/content/README.md
@@ -29,7 +29,7 @@ its public counterpart needs the same change.
| Public page | Derived from | Watch for drift in |
|---|---|---|
| `docs/what-is-archilyzer.md` | `README.md` | package list, pipeline modes |
-| `docs/install.md` | `SETUP.md` | tool versions, env-var table, backend list |
+| `docs/install.md` | `SETUP.md` | tool versions, env-var table, backend list, the Windows path |
| `docs/operate.md` | `README.md`, `SCHEDULED_SYNC.md` | editor routes, scheduler settings |
| `docs/deploy-cloudflare.md` | `DEPLOY_CLOUDFLARE.md` | the 25 MB Pages limit, R2 options |
| `docs/deploy-docker.md` | `DEPLOY_DOCKER.md` | phase structure, settings names |
diff --git a/homepage/content/docs/install.md b/homepage/content/docs/install.md
@@ -22,7 +22,7 @@ Needed only for the pipeline — install what you will actually use:
| ffmpeg and ffprobe | Audio transcode and duration checks. |
| A transcription backend | Channels you transcribe yourself. whisper.cpp by default. |
| rsync | Backing up the saved-video store, if you enable it. |
-| Docker or podman | Only for the parallel multi-site build. |
+| Docker or podman | Running the whole stack in containers — the simplest path on Windows, and it needs none of the tools above — and the parallel multi-site build. |
The apps start and run with **none** of the second table installed. You just
can't fetch anything yet.
@@ -95,8 +95,42 @@ whisper-cpp`), which saves you building it.
### Windows
-**Use WSL2.** The transcription toolchain and the helper shell scripts assume a
-Unix shell, so the smooth path is:
+There are two Windows paths, and which one you want depends on whether you are
+**running** an archive or **working on** the software.
+
+**To run an archive: Docker Desktop.** One command, and none of the toolchain
+above has to be installed by hand — the image already contains yt-dlp, ffmpeg
+and a transcription backend. Install
+[Docker Desktop](https://www.docker.com/products/docker-desktop/), then in
+PowerShell:
+
+```powershell
+Invoke-WebRequest https://archilyzer.pages.dev/downloads/archilyzer-source.tar.gz -OutFile archilyzer-source.tar.gz
+tar xzf archilyzer-source.tar.gz
+cd archilyzer
+copy .env.example .env
+docker compose up -d --build
+```
+
+Then open **http://localhost:8081**. The first build compiles a transcription
+backend and two web apps, so it takes a while; after that, starting is seconds.
+The first boot also downloads a speech model and writes a settings file with one
+transcription worker already enabled, so the editor is usable the moment it comes
+up rather than looking healthy and quietly transcribing nothing.
+
+**Nothing it runs is reachable from anywhere else by default.** Every port binds
+`127.0.0.1` — the published archive included — so opening one up is a deliberate
+one-line edit in `.env`. The editor has no login of its own, so the containers
+refuse to start if you expose *it* without putting a password or an identity
+provider in front; the error says how. `RUNNING_IN_DOCKER.md`, in the snapshot
+you just unpacked, covers that, GPU transcription, and backups.
+
+The same stack runs on Linux and macOS. It is described here because it is the
+one place where it is clearly the *better* option, not because it is
+Windows-only.
+
+**To work on the software: WSL2.** The transcription toolchain and the helper
+shell scripts assume a Unix shell:
```powershell
wsl --install -d Ubuntu