# Install From a clone (or an unpacked snapshot) to a running editor. The web apps need very little; the download-and-transcribe pipeline needs the media tools, and you can add those later. ## What you need Always required, to install dependencies and run the apps: | Tool | Version | Notes | |---|---|---| | Node.js | 20.9 or newer | LTS 20 or 22. Required by Next.js 16. | | pnpm | 9 or newer | Easiest via Corepack, which ships with Node. | | A C/C++ toolchain | platform default | Only if no prebuilt binary exists for a native module on your platform. Usually not needed. | Needed only for the pipeline — install what you will actually use: | Tool | Needed for | |---|---| | yt-dlp | Downloading and syncing channels. The whole pipeline. | | ffmpeg and ffprobe | Audio transcode and duration checks. | | A transcription backend | Channels you transcribe yourself. whisper.cpp by default. | | rsync | Backing up the saved-video store, if you enable it. | | Docker or podman | Running the whole stack in containers — the simplest path on Windows, and it needs none of the tools above — and the parallel multi-site build. | The apps start and run with **none** of the second table installed. You just can't fetch anything yet. ## Get the code The source lives on this site as a read-only git mirror of the main branch — there is no GitHub. Clone it: ```sh git clone https://archilyzer.pages.dev/source/archilyzer.git archilyzer cd archilyzer ``` `git pull` brings you up to date; nothing takes a push. The commit ids differ from the private repository's, because machine paths are scrubbed on the way out. [Source](/source/) has the details, a browsable copy of every file, and the history with every commit's diff. No git? The same tree, without history, is a tarball: ```sh curl -LO https://archilyzer.pages.dev/downloads/archilyzer-source.tar.gz tar xzf archilyzer-source.tar.gz cd archilyzer ``` That is a working tree, not a clone. Updating means downloading a newer snapshot. See [Downloads](/downloads/) for what is and isn't inside. ## Install and start ```sh corepack enable # once, so the pinned pnpm version is used pnpm install # installs every workspace package pnpm dev:editor # the editor, at http://localhost:3001 ``` The editor **starts fine with no data**. A fresh unpack has no corpus directory, and that is expected — it is created as you use it. Add your first channel from the editor's Channels page. To build and serve the public site: ```sh pnpm build # static site under export/out/ pnpm start:export # serve it at http://localhost:3000 ``` ## Platform notes ### Linux Install Node from your distribution or a version manager, then the media tools: ```sh # Debian / Ubuntu sudo apt install nodejs yt-dlp ffmpeg rsync build-essential # Fedora sudo dnf install nodejs yt-dlp ffmpeg rsync gcc-c++ make # Arch sudo pacman -S nodejs yt-dlp ffmpeg rsync base-devel ``` If your distribution's Node is older than 20.9, install it with `fnm` or `nvm` instead. The build-tools package is the compiler pnpm falls back to when a native module has no prebuilt binary for your platform. ### macOS ```sh xcode-select --install # C/C++ toolchain brew install node pnpm yt-dlp ffmpeg rsync corepack enable ``` whisper.cpp can be installed straight from Homebrew (`brew install whisper-cpp`), which saves you building it. ### Windows There are two Windows paths, and which one you want depends on whether you are **running** an archive or **working on** the software. **To run an archive: Docker Desktop.** One command, and none of the toolchain above has to be installed by hand — the image already contains yt-dlp, ffmpeg and a transcription backend. Install [Docker Desktop](https://www.docker.com/products/docker-desktop/), then in PowerShell: ```powershell Invoke-WebRequest https://archilyzer.pages.dev/downloads/archilyzer-source.tar.gz -OutFile archilyzer-source.tar.gz tar xzf archilyzer-source.tar.gz cd archilyzer copy .env.example .env docker compose up -d --build ``` Then open **http://localhost:8081**. The first build compiles a transcription backend and two web apps, so it takes a while; after that, starting is seconds. The first boot also downloads a speech model and writes a settings file with one transcription worker already enabled, so the editor is usable the moment it comes up rather than looking healthy and quietly transcribing nothing. **Nothing it runs is reachable from anywhere else by default.** Every port binds `127.0.0.1` — the published archive included — so opening one up is a deliberate one-line edit in `.env`. The editor has no login of its own, so the containers refuse to start if you expose *it* without putting a password or an identity provider in front; the error says how. `RUNNING_IN_DOCKER.md`, in the tree you just unpacked, covers that, GPU transcription, and backups. The same stack runs on Linux and macOS. It is described here because it is the one place where it is clearly the *better* option, not because it is Windows-only. **To work on the software: WSL2.** The transcription toolchain and the helper shell scripts assume a Unix shell: ```powershell wsl --install -d Ubuntu ``` then follow the Linux steps inside it. Keep the files **inside** the WSL filesystem rather than under `/mnt/c/`, or file watching and I/O will be slow. The apps and the yt-dlp pipeline do work natively on Windows — install Node, yt-dlp and ffmpeg with winget or Scoop — but building the transcription backends and running the shell scripts is Unix-oriented. ## Transcription backends You only need one of these for channels whose audio you transcribe yourself. Channels where the platform publishes captions need no backend at all. - **whisper.cpp** — the default. Build it, download a `ggml` model, and either put `whisper-cli` on your `PATH` or point `WHISPER_BIN` at it. Set `WHISPER_MODEL` to the model file. - **chough** — set `CHOUGH_BIN`; it downloads a model when none is configured, and can talk to a remote server via `CHOUGH_URL`. - **parakeet.cpp** — driven through a bundled overlapping-segment wrapper. Needs `parakeet-cli` and a `.gguf` model. It is interruptible: stopping it stitches together the partial transcript rather than throwing the work away. The active backend is chosen once, globally, in the editor's settings. All three call ffmpeg, so that must be installed either way. ## Configuration Most configuration lives in the editor's settings page and is written to a settings file at the root. It is entirely optional — a missing or partial file falls back to built-in defaults, so the app runs out of the box. Paths and binaries can also be overridden by environment variables set before launch. The ones worth knowing: | Variable | Default | Purpose | |---|---|---| | `TRANSCRIPTS_DIR` | `/transcripts` | The corpus: channels, media, index, job logs. | | `SAVED_VIDEOS_DIR` | inside `TRANSCRIPTS_DIR` | Persisted source-video store; can live on another disk. | | `SITES_DIR` | inside `TRANSCRIPTS_DIR` | Per-site configuration. | | `EXPORT_PUBLIC_DIR` | `/export/public` | Where the index writes its paginated JSON. | | `YTDLP_BIN` | `yt-dlp` on `PATH` | The downloader. | | `WHISPER_BIN` | `whisper-cli` on `PATH` | whisper.cpp binary. | | `WHISPER_MODEL` | a path under your home directory | whisper.cpp model file. | | `FFMPEG_BIN` / `FFPROBE_BIN` | on `PATH` | Transcode and duration checks. | Pointing `TRANSCRIPTS_DIR` at a large disk before you start is the one decision worth making early — the corpus grows to whatever your channels amount to, and moving it later means moving everything. Next: [Running an archive](/docs/operate/).