Archilyzer
← All docs

Install

What to install, in what order, and which parts you can skip until you need them.

From a clone (or an unpacked snapshot) to a running editor. The web apps need very little; the download-and-transcribe pipeline needs the media tools, and you can add those later.

What you need

Always required, to install dependencies and run the apps:

Tool Version Notes
Node.js 20.9 or newer LTS 20 or 22. Required by Next.js 16.
pnpm 9 or newer Easiest via Corepack, which ships with Node.
A C/C++ toolchain platform default Only if no prebuilt binary exists for a native module on your platform. Usually not needed.

Needed only for the pipeline — install what you will actually use:

Tool Needed for
yt-dlp Downloading and syncing channels. The whole pipeline.
ffmpeg and ffprobe Audio transcode and duration checks.
A transcription backend Channels you transcribe yourself. whisper.cpp by default.
rsync Backing up the saved-video store, if you enable it.
Docker or podman Running the whole stack in containers — the simplest path on Windows, and it needs none of the tools above — and the parallel multi-site build.

The apps start and run with none of the second table installed. You just can't fetch anything yet.

Get the code

The source lives on this site as a read-only git mirror of the main branch — there is no GitHub. Clone it:

git clone https://archilyzer.pages.dev/source/archilyzer.git archilyzer
cd archilyzer

git pull brings you up to date; nothing takes a push. The commit ids differ from the private repository's, because machine paths are scrubbed on the way out. Source has the details, a browsable copy of every file, and the history with every commit's diff.

No git? The same tree, without history, is a tarball:

curl -LO https://archilyzer.pages.dev/downloads/archilyzer-source.tar.gz
tar xzf archilyzer-source.tar.gz
cd archilyzer

That is a working tree, not a clone. Updating means downloading a newer snapshot. See Downloads for what is and isn't inside.

Install and start

corepack enable        # once, so the pinned pnpm version is used
pnpm install           # installs every workspace package
pnpm dev:editor        # the editor, at http://localhost:3001

The editor starts fine with no data. A fresh unpack has no corpus directory, and that is expected — it is created as you use it. Add your first channel from the editor's Channels page.

To build and serve the public site:

pnpm build             # static site under export/out/
pnpm start:export      # serve it at http://localhost:3000

Platform notes

Linux

Install Node from your distribution or a version manager, then the media tools:

# Debian / Ubuntu
sudo apt install nodejs yt-dlp ffmpeg rsync build-essential

# Fedora
sudo dnf install nodejs yt-dlp ffmpeg rsync gcc-c++ make

# Arch
sudo pacman -S nodejs yt-dlp ffmpeg rsync base-devel

If your distribution's Node is older than 20.9, install it with fnm or nvm instead. The build-tools package is the compiler pnpm falls back to when a native module has no prebuilt binary for your platform.

macOS

xcode-select --install                       # C/C++ toolchain
brew install node pnpm yt-dlp ffmpeg rsync
corepack enable

whisper.cpp can be installed straight from Homebrew (brew install whisper-cpp), which saves you building it.

Windows

There are two Windows paths, and which one you want depends on whether you are running an archive or working on the software.

To run an archive: Docker Desktop. One command, and none of the toolchain above has to be installed by hand — the image already contains yt-dlp, ffmpeg and a transcription backend. Install Docker Desktop, then in PowerShell:

Invoke-WebRequest https://archilyzer.pages.dev/downloads/archilyzer-source.tar.gz -OutFile archilyzer-source.tar.gz
tar xzf archilyzer-source.tar.gz
cd archilyzer
copy .env.example .env
docker compose up -d --build

Then open http://localhost:8081. The first build compiles a transcription backend and two web apps, so it takes a while; after that, starting is seconds. The first boot also downloads a speech model and writes a settings file with one transcription worker already enabled, so the editor is usable the moment it comes up rather than looking healthy and quietly transcribing nothing.

Nothing it runs is reachable from anywhere else by default. Every port binds 127.0.0.1 — the published archive included — so opening one up is a deliberate one-line edit in .env. The editor has no login of its own, so the containers refuse to start if you expose it without putting a password or an identity provider in front; the error says how. RUNNING_IN_DOCKER.md, in the tree you just unpacked, covers that, GPU transcription, and backups.

The same stack runs on Linux and macOS. It is described here because it is the one place where it is clearly the better option, not because it is Windows-only.

To work on the software: WSL2. The transcription toolchain and the helper shell scripts assume a Unix shell:

wsl --install -d Ubuntu

then follow the Linux steps inside it. Keep the files inside the WSL filesystem rather than under /mnt/c/, or file watching and I/O will be slow.

The apps and the yt-dlp pipeline do work natively on Windows — install Node, yt-dlp and ffmpeg with winget or Scoop — but building the transcription backends and running the shell scripts is Unix-oriented.

Transcription backends

You only need one of these for channels whose audio you transcribe yourself. Channels where the platform publishes captions need no backend at all.

  • whisper.cpp — the default. Build it, download a ggml model, and either put whisper-cli on your PATH or point WHISPER_BIN at it. Set WHISPER_MODEL to the model file.
  • chough — set CHOUGH_BIN; it downloads a model when none is configured, and can talk to a remote server via CHOUGH_URL.
  • parakeet.cpp — driven through a bundled overlapping-segment wrapper. Needs parakeet-cli and a .gguf model. It is interruptible: stopping it stitches together the partial transcript rather than throwing the work away.

The active backend is chosen once, globally, in the editor's settings. All three call ffmpeg, so that must be installed either way.

Configuration

Most configuration lives in the editor's settings page and is written to a settings file at the root. It is entirely optional — a missing or partial file falls back to built-in defaults, so the app runs out of the box.

Paths and binaries can also be overridden by environment variables set before launch. The ones worth knowing:

Variable Default Purpose
TRANSCRIPTS_DIR <root>/transcripts The corpus: channels, media, index, job logs.
SAVED_VIDEOS_DIR inside TRANSCRIPTS_DIR Persisted source-video store; can live on another disk.
SITES_DIR inside TRANSCRIPTS_DIR Per-site configuration.
EXPORT_PUBLIC_DIR <root>/export/public Where the index writes its paginated JSON.
YTDLP_BIN yt-dlp on PATH The downloader.
WHISPER_BIN whisper-cli on PATH whisper.cpp binary.
WHISPER_MODEL a path under your home directory whisper.cpp model file.
FFMPEG_BIN / FFPROBE_BIN on PATH Transcode and duration checks.

Pointing TRANSCRIPTS_DIR at a large disk before you start is the one decision worth making early — the corpus grows to whatever your channels amount to, and moving it later means moving everything.

Next: Running an archive.