Install
What to install, in what order, and which parts you can skip until you need them.
From a clone (or an unpacked snapshot) to a running editor. The web apps need very little; the download-and-transcribe pipeline needs the media tools, and you can add those later.
What you need
Always required, to install dependencies and run the apps:
| Tool | Version | Notes |
|---|---|---|
| Node.js | 20.9 or newer | LTS 20 or 22. Required by Next.js 16. |
| pnpm | 9 or newer | Easiest via Corepack, which ships with Node. |
| A C/C++ toolchain | platform default | Only if no prebuilt binary exists for a native module on your platform. Usually not needed. |
Needed only for the pipeline — install what you will actually use:
| Tool | Needed for |
|---|---|
| yt-dlp | Downloading and syncing channels. The whole pipeline. |
| ffmpeg and ffprobe | Audio transcode and duration checks. |
| A transcription backend | Channels you transcribe yourself. whisper.cpp by default. |
| rsync | Backing up the saved-video store, if you enable it. |
| Docker or podman | Running the whole stack in containers — the simplest path on Windows, and it needs none of the tools above — and the parallel multi-site build. |
The apps start and run with none of the second table installed. You just can't fetch anything yet.
Get the code
The source lives on this site as a read-only git mirror of the main branch — there is no GitHub. Clone it:
git clone https://archilyzer.pages.dev/source/archilyzer.git archilyzer
cd archilyzer
git pull brings you up to date; nothing takes a push. The commit ids differ
from the private repository's, because machine paths are scrubbed on the way
out. Source has the details, a browsable copy of every file, and the
history with every commit's diff.
No git? The same tree, without history, is a tarball:
curl -LO https://archilyzer.pages.dev/downloads/archilyzer-source.tar.gz
tar xzf archilyzer-source.tar.gz
cd archilyzer
That is a working tree, not a clone. Updating means downloading a newer snapshot. See Downloads for what is and isn't inside.
Install and start
corepack enable # once, so the pinned pnpm version is used
pnpm install # installs every workspace package
pnpm dev:editor # the editor, at http://localhost:3001
The editor starts fine with no data. A fresh unpack has no corpus directory, and that is expected — it is created as you use it. Add your first channel from the editor's Channels page.
To build and serve the public site:
pnpm build # static site under export/out/
pnpm start:export # serve it at http://localhost:3000
Platform notes
Linux
Install Node from your distribution or a version manager, then the media tools:
# Debian / Ubuntu
sudo apt install nodejs yt-dlp ffmpeg rsync build-essential
# Fedora
sudo dnf install nodejs yt-dlp ffmpeg rsync gcc-c++ make
# Arch
sudo pacman -S nodejs yt-dlp ffmpeg rsync base-devel
If your distribution's Node is older than 20.9, install it with fnm or nvm
instead. The build-tools package is the compiler pnpm falls back to when a
native module has no prebuilt binary for your platform.
macOS
xcode-select --install # C/C++ toolchain
brew install node pnpm yt-dlp ffmpeg rsync
corepack enable
whisper.cpp can be installed straight from Homebrew (brew install
whisper-cpp), which saves you building it.
Windows
There are two Windows paths, and which one you want depends on whether you are running an archive or working on the software.
To run an archive: Docker Desktop. One command, and none of the toolchain above has to be installed by hand — the image already contains yt-dlp, ffmpeg and a transcription backend. Install Docker Desktop, then in PowerShell:
Invoke-WebRequest https://archilyzer.pages.dev/downloads/archilyzer-source.tar.gz -OutFile archilyzer-source.tar.gz
tar xzf archilyzer-source.tar.gz
cd archilyzer
copy .env.example .env
docker compose up -d --build
Then open http://localhost:8081. The first build compiles a transcription backend and two web apps, so it takes a while; after that, starting is seconds. The first boot also downloads a speech model and writes a settings file with one transcription worker already enabled, so the editor is usable the moment it comes up rather than looking healthy and quietly transcribing nothing.
Nothing it runs is reachable from anywhere else by default. Every port binds
127.0.0.1 — the published archive included — so opening one up is a deliberate
one-line edit in .env. The editor has no login of its own, so the containers
refuse to start if you expose it without putting a password or an identity
provider in front; the error says how. RUNNING_IN_DOCKER.md, in the tree you
just unpacked, covers that, GPU transcription, and backups.
The same stack runs on Linux and macOS. It is described here because it is the one place where it is clearly the better option, not because it is Windows-only.
To work on the software: WSL2. The transcription toolchain and the helper shell scripts assume a Unix shell:
wsl --install -d Ubuntu
then follow the Linux steps inside it. Keep the files inside the WSL
filesystem rather than under /mnt/c/, or file watching and I/O will be slow.
The apps and the yt-dlp pipeline do work natively on Windows — install Node, yt-dlp and ffmpeg with winget or Scoop — but building the transcription backends and running the shell scripts is Unix-oriented.
Transcription backends
You only need one of these for channels whose audio you transcribe yourself. Channels where the platform publishes captions need no backend at all.
- whisper.cpp — the default. Build it, download a
ggmlmodel, and either putwhisper-clion yourPATHor pointWHISPER_BINat it. SetWHISPER_MODELto the model file. - chough — set
CHOUGH_BIN; it downloads a model when none is configured, and can talk to a remote server viaCHOUGH_URL. - parakeet.cpp — driven through a bundled overlapping-segment wrapper. Needs
parakeet-cliand a.ggufmodel. It is interruptible: stopping it stitches together the partial transcript rather than throwing the work away.
The active backend is chosen once, globally, in the editor's settings. All three call ffmpeg, so that must be installed either way.
Configuration
Most configuration lives in the editor's settings page and is written to a settings file at the root. It is entirely optional — a missing or partial file falls back to built-in defaults, so the app runs out of the box.
Paths and binaries can also be overridden by environment variables set before launch. The ones worth knowing:
| Variable | Default | Purpose |
|---|---|---|
TRANSCRIPTS_DIR | <root>/transcripts | The corpus: channels, media, index, job logs. |
SAVED_VIDEOS_DIR | inside TRANSCRIPTS_DIR | Persisted source-video store; can live on another disk. |
SITES_DIR | inside TRANSCRIPTS_DIR | Per-site configuration. |
EXPORT_PUBLIC_DIR | <root>/export/public | Where the index writes its paginated JSON. |
YTDLP_BIN | yt-dlp on PATH | The downloader. |
WHISPER_BIN | whisper-cli on PATH | whisper.cpp binary. |
WHISPER_MODEL | a path under your home directory | whisper.cpp model file. |
FFMPEG_BIN / FFPROBE_BIN | on PATH | Transcode and duration checks. |
Pointing TRANSCRIPTS_DIR at a large disk before you start is the one decision
worth making early — the corpus grows to whatever your channels amount to, and
moving it later means moving everything.
Next: Running an archive.