# Setup & requirements This guide takes you from a fresh copy of **Archilyzer** to a running editor and static export on **Linux, macOS, or Windows**. It covers what to install, in what order, and which pieces you only need for specific features. > An operator-facing version of this guide is published at > [archilyzer.pages.dev/docs/install/](https://archilyzer.pages.dev/docs/install/). This > one is the contributor's copy and additionally covers worktrees, the e2e queue and the > sharded test run. - [Requirements at a glance](#requirements-at-a-glance) - [Install the toolchain](#install-the-toolchain) - [Linux](#linux) - [macOS](#macos) - [Windows](#windows) - [Get the code running](#get-the-code-running) - [The `transcripts/` data directory](#the-transcripts-data-directory) - [Transcription backends (optional)](#transcription-backends-optional) - [Configuration & environment variables](#configuration--environment-variables) - [Running the tests](#running-the-tests) - [Where to go next](#where-to-go-next) --- ## Requirements at a glance **Always required** to install dependencies and run the web apps (editor, export, homepage): | Tool | Version | Notes | | --- | --- | --- | | **Node.js** | **22** (≥ 20.9 runs the apps) | Next.js 16 needs ≥ 20.9, but **deploying** runs the wrangler pinned in `common/package.json`, which refuses anything below Node 22 — so use 22. The Docker image ships 22. | | **pnpm** | **9+** | Lockfile is v9. Easiest via Corepack (bundled with Node) — see below. | | **git** | any recent | To clone the repo. | | **C/C++ toolchain** | platform default | Only if pnpm can't find a prebuilt binary for a native module (`lmdb`, `sharp`, …). Usually not needed on mainstream platforms. | > The web apps run with **none** of the media binaries below. You only need those to > use the download/transcribe **pipeline** (fetching videos, generating transcripts). **Feature-specific — install only what you'll use:** | Tool | Needed for | Default lookup | | --- | --- | --- | | **yt-dlp** | Downloading / syncing channels (the whole pipeline). | `yt-dlp` on `PATH` | | **ffmpeg** + **ffprobe** | Audio transcode + duration checks for `transcribe` channels. | `ffmpeg` / `ffprobe` on `PATH` | | A **transcription backend** | `handling: "transcribe"` channels only. Default is **whisper.cpp** (`whisper-cli`); `chough` and `parakeet.cpp` are alternatives. | `whisper-cli` on `PATH` | | **rsync** | Backing up the saved-video store. | `rsync` on `PATH` | | **Docker** | Running the whole stack in containers ([RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md)), the parallel multi-site build ([PUBLISH.md](PUBLISH.md#building-every-site-in-containers)), and the sharded e2e run (`pnpm e2e:sharded`). | — | Every binary above is overridable by an environment variable (e.g. `YTDLP_BIN`) — see [Configuration & environment variables](#configuration--environment-variables). --- ## Install the toolchain Enable **Corepack** once (it ships with Node) so the repo's pnpm version is used automatically: ```sh corepack enable ``` ### Linux Node (use your distro, or a version manager like [`fnm`](https://github.com/Schniz/fnm) / `nvm` to get ≥ 20.9), then the media tools: ```sh # Debian / Ubuntu sudo apt install nodejs yt-dlp ffmpeg rsync build-essential corepack enable # Fedora sudo dnf install nodejs yt-dlp ffmpeg rsync gcc-c++ make corepack enable # Arch sudo pacman -S nodejs yt-dlp ffmpeg rsync base-devel corepack enable ``` If your distro's `nodejs` is older than 20.9, install Node via `fnm`/`nvm` instead: ```sh fnm install 22 && fnm use 22 # or: nvm install 22 && nvm use 22 corepack enable ``` `build-essential` / `base-devel` / `gcc-c++ make` provides the C/C++ compiler pnpm falls back to when a native module (`lmdb`, `sharp`, …) has no prebuilt binary for your platform. ### macOS Install the Xcode Command Line Tools (compiler for native modules), then use [Homebrew](https://brew.sh): ```sh xcode-select --install # C/C++ toolchain brew install node pnpm yt-dlp ffmpeg rsync # rsync ships with macOS but brew's is newer corepack enable ``` `brew install pnpm` is optional if you use Corepack — either works. ### Windows Two paths, and they answer different questions. **To run an archive: Docker Desktop.** The container image already contains yt-dlp, ffmpeg and a transcription backend, so none of the toolchain in the tables above has to be installed on the host at all. Get the source first ([Get the code running](#get-the-code-running) below), then, from the repo root: ```powershell copy .env.example .env docker compose up -d --build ``` Then open **http://localhost:8081**. The first build compiles whisper.cpp and two Next.js apps and is slow; afterwards, starting is seconds. First boot fetches a speech model and seeds a `settings.json` with one enabled worker, so transcription works immediately instead of silently doing nothing. Every port binds `127.0.0.1` by default — including the published site — and the containers **refuse to start** if the editor is exposed to a network with no auth in front of it. See [RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md) for the exposure model, the auth options, GPU transcription (Vulkan/parakeet or CUDA/whisper) and backups. The same stack runs on Linux and macOS. **To develop, or to run the e2e suite: WSL2.** The transcription toolchain (whisper.cpp / parakeet.cpp) and helper shell scripts assume a Unix shell, so the smoothest path for working *on* the code is [WSL2](https://learn.microsoft.com/windows/wsl/install): ```powershell wsl --install -d Ubuntu ``` Then open the Ubuntu shell and follow the [**Linux**](#linux) steps above. Keep the repo **inside** the WSL filesystem (e.g. `~/yt-dlp-transcript-browser`), not under `/mnt/c/…`, for good file-watch and I/O performance. **Native-Windows notes.** The web apps and the yt-dlp pipeline work natively too. Install with [winget](https://learn.microsoft.com/windows/package-manager/) (or [Scoop](https://scoop.sh)): ```powershell winget install OpenJS.NodeJS.LTS winget install yt-dlp.yt-dlp winget install Gyan.FFmpeg corepack enable ``` Caveats on native Windows: - Building the **whisper.cpp / parakeet.cpp** transcription backends is Unix-oriented; do transcription work under WSL2, or let the container do it (it ships a backend already built). - The external-cron sync path (`pnpm sync:tick` from cron) assumes a Unix shell — use WSL2 or Task Scheduler equivalents. So does publishing the source mirror (`archilyzer source publish`: git-filter-repo, `tar`). --- ## Get the code running Archilyzer is MIT licensed and **forge-neutral** — there is no provider-specific CI config or `.github/` directory, and nothing in the build assumes a particular host. Any complete copy of the tree works. The canonical public copy is the read-only git mirror on the project site — `main` with its whole history: ```sh git clone https://archilyzer.pages.dev/source/archilyzer.git archilyzer cd archilyzer ``` `git pull` updates it; nothing takes a push. Its commit ids differ from the private repository's (machine paths are scrubbed on the way out), and every homepage build regenerates it (`archilyzer source publish`; [PUBLISH.md](PUBLISH.md#the-source-mirror-homepage)). No git? The same tree without history is a tarball, with a `snapshot.json` sidecar (commit, size, SHA-256) beside it: ```sh curl -LO https://archilyzer.pages.dev/downloads/archilyzer-source.tar.gz tar xzf archilyzer-source.tar.gz cd archilyzer ``` The snapshot is a working tree at one commit — no history, no branches, no remote — so updating means fetching a newer one. Once you have a tree, from a clone or a snapshot alike: ```sh pnpm install # installs all workspace packages; compiles native modules # (lmdb, msgpackr-extract, sharp, …) per the build allowlist # in pnpm-workspace.yaml pnpm dev:editor # editor (admin UI) at http://localhost:3001 ``` To build and serve the read-only public site: ```sh pnpm build # publish index + publish build $SITE_ID (export/out links to the bundle) pnpm start:export # serve export/out/ at http://localhost:3000 ``` The editor **starts fine with no data** — a fresh clone has no `transcripts/` directory (see below) and that's expected. Create your first channel through the editor's **/channels** page; the pipeline populates `transcripts/` from there. --- ## The `transcripts/` data directory All downloaded data — channels, archives, the LMDB index, job logs — lives under `transcripts/` at the repo root. It is a **separate git repo**, deliberately **not** part of this workspace (`/transcripts` is in the root `.gitignore`), so a fresh clone of this repo does **not** include it. There is nothing to bootstrap: the app **creates the directory and its subdirectories lazily** as you use it. Start the editor, add channels through **/channels**, and run the pipeline — `transcripts/channels//…`, `transcripts/index.mdb`, etc. appear on demand. Override the location with `TRANSCRIPTS_DIR` (e.g. to put data on another disk). `common/lib/paths.ts` is the source of truth for every path and its env override. --- ## Transcription backends (optional) You only need a transcription backend for channels with `handling: "transcribe"` (audio-only downloads that get transcribed locally). YouTube channels (`handling: "youtube"`) fetch existing captions and need no backend. The active backend is chosen globally in **Settings → Transcription**; three are registered in `common/lib/transcriptionApps.ts`: - **whisper.cpp** (default, `whisper-cli`) — build [whisper.cpp](https://github.com/ggml-org/whisper.cpp), download a `ggml` model, and either put `whisper-cli` on your `PATH` or set `WHISPER_BIN`. Point `WHISPER_MODEL` at the model (default `~/whispercpp/whisper.cpp/models/ggml-base.en.bin`), or set the binary/model in the Settings UI. On macOS you can skip the build: `brew install whisper-cpp`. - **chough** — set the binary via `CHOUGH_BIN`; auto-downloads a model when none is configured. Supports a remote server (`CHOUGH_URL`). - **parakeet.cpp** — driven by the bundled overlapping-segment wrapper (`scripts/parakeet-stitch.mjs`); needs `parakeet-cli` (`PARAKEET_CLI`) and a `.gguf` model (`PARAKEET_MODEL`). Interruptible (stops and stitches a partial transcript). All three read `ffmpeg`/`ffprobe`, so make sure those are installed for transcribe channels. --- ## Configuration & environment variables Most configuration now lives in the editor's **/settings** page, persisted to `settings.json` at the repo root (gitignored). Every key, its default and what it does is in [SETTINGS.md](SETTINGS.md); `settings.json.example` is the defaults as a starting template. Both are generated from the settings schema (`common/lib/settingsSchema.ts`). The per-site `site.json` and the per-channel `config.json` have generated key tables of their own: [SITE.md](SITE.md) and [CHANNEL.md](CHANNEL.md). Settings are optional — a missing/partial `settings.json` falls back to built-in defaults, so the app runs out of the box. Paths and binaries resolve through `getPaths()` in `common/lib/paths.ts`, and each is overridden by an environment variable before launching — `TRANSCRIPTS_DIR` (the corpus, default `/transcripts`), `SETTINGS_FILE`, `YTDLP_BIN`, `FFMPEG_BIN` / `FFPROBE_BIN`, `WHISPER_BIN` / `WHISPER_MODEL`, `PARAKEET_CLI` / `PARAKEET_MODEL` and the rest. **Every variable the code reads is in [ENVIRONMENT.md](ENVIRONMENT.md)**, by audience: the path overrides, the runtime tokens and knobs (`WORKER_TOKEN`, `SYNC_TICK_URL`, the R2 credentials, …), the ports, the docker `ARCHILYZER_*` set and the test-only ones. It is generated from `common/lib/envVars.ts`, and a test fails when the code reads a variable that list does not declare. To see what this machine has, run: ```sh pnpm archilyzer doctor ``` It is read-only: the checkout, the corpus and each channel's media, `settings.json`, every binary the paths name plus each enabled worker's engine and model, umtool's report-pipeline tools, and this checkout's port block. It exits 1 only for something the machine is configured to do and cannot (an enabled worker's engine missing beside a corpus, a settings file that does not parse, an override naming a missing binary). --- ## Running the tests The editor's Playwright suite drives the admin UI. It uses fake binaries for yt-dlp/whisper/ffmpeg, so you do **not** need the media tools installed to run it — but you do need the Playwright browser: ```sh npx playwright install chromium # one-time: download the test browser pnpm e2e # sequential run on the host (next dev, port 3011) ``` For a faster parallel run, `pnpm e2e:sharded` splits the suite across N Docker containers (needs **Docker**; the container image ships the browser, so no `playwright install` there). Shard count defaults to `min(max(2, cpus/2), 8)`; override with `E2E_SHARDS=N pnpm e2e:sharded`. To run multiple checkouts / dev servers / e2e suites at once without port collisions, see [WORKTREES.md](WORKTREES.md). --- ## Where to go next - [README.md](README.md) — project overview, pipeline modes. - [SCHEDULED_SYNC.md](SCHEDULED_SYNC.md) — automatic per-channel sync (internal heartbeat or external cron). - [PUBLISH.md](PUBLISH.md) — building and deploying sites: Cloudflare Pages, R2 archive overflow, parallel builds in containers. - [ENVIRONMENT.md](ENVIRONMENT.md) — every environment variable, by audience. - [WORKTREES.md](WORKTREES.md) — parallel development with per-worktree ports. - [mcp/README.md](mcp/README.md) — MCP server exposing the archive to Claude Code / Desktop / Cursor. - [r2-proxy/README.md](r2-proxy/README.md) — the Cloudflare Worker that serves oversize archives from R2.