commit 5708cc3ee7b394cccf515aea581d8c9dc6427004
parent 64f0a6b2ea8d9b79580df8ed14d8b69b84e4870d
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Thu, 16 Jul 2026 01:34:14 -0400
Add setup docs
Diffstat:
| M | README.md | | | 20 | +++++++++++++++----- |
| A | SETUP.md | | | 272 | +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ |
2 files changed, 287 insertions(+), 5 deletions(-)
diff --git a/README.md b/README.md
@@ -8,6 +8,13 @@ A pnpm-workspace monorepo for archiving and browsing video transcripts. The proj
Transcripts and per-channel state live at `<repo>/transcripts/` (its own git repo, untouched by the workspace).
+## Requirements
+
+- **Node.js ≥ 20.9** (LTS 20 or 22) and **pnpm 9+** (easiest via `corepack enable`) and **git** — needed to install and run the web apps.
+- The download/transcribe **pipeline** additionally needs **yt-dlp**, **ffmpeg**/**ffprobe**, and a transcription backend (whisper.cpp by default). The web apps run without them.
+
+**See [SETUP.md](SETUP.md) for full per-OS install instructions (Linux, macOS, Windows).**
+
## Getting started
```bash
@@ -15,6 +22,7 @@ pnpm install
pnpm dev:editor # editor at http://localhost:3001
pnpm build # build the static site under export/out/
pnpm start:export # serve export/out/ at http://localhost:3000
+npx playwright install chromium # one-time, before the first e2e run
pnpm e2e # run the editor's Playwright suite (sequential)
pnpm e2e:sharded # same suite, sharded across N Docker containers in parallel
```
@@ -31,9 +39,9 @@ pnpm e2e:sharded # same suite, sharded across N Docker containers in paralle
Channels can also sync automatically on a per-channel cadence via a cron heartbeat — see [SCHEDULED_SYNC.md](SCHEDULED_SYNC.md).
-## Environment variables
+## Configuration & environment variables
-All paths and binaries used by the editor and CLI shims resolve through `getPaths()`. Override any of them before launching:
+Most configuration lives in the editor's **/settings** page (persisted to `settings.json`); it's optional and falls back to defaults. All paths and binaries resolve through `getPaths()` (`common/lib/paths.ts`) and can be overridden by environment variables. Common ones:
| Variable | Default | Purpose |
| --- | --- | --- |
@@ -41,9 +49,11 @@ All paths and binaries used by the editor and CLI shims resolve through `getPath
| `EXPORT_PUBLIC_DIR` | `<repo>/export/public` | Where the index writes paginated JSON. |
| `SETTINGS_FILE` | `<repo>/settings.json` | Site-settings file. |
| `YTDLP_BIN` | `yt-dlp` (PATH lookup) | Binary used by the pipeline. |
-| `WHISPER_BIN` | `whisper-cli` (PATH lookup) | Binary used by the whisper batch. |
+| `WHISPER_BIN` | `whisper-cli` (PATH lookup) | Default transcription binary. |
| `WHISPER_MODEL` | `~/whispercpp/whisper.cpp/models/ggml-base.en.bin` | Model file passed to whisper-cli. |
-| `PARALLEL_TRANSCRIBE_LIMIT` | `4` | Max parallel whisper jobs. |
+| `PARALLEL_TRANSCRIBE_LIMIT` | `4` | Max parallel transcription jobs. |
+
+See **[SETUP.md](SETUP.md#configuration--environment-variables)** for the full list (ffmpeg/ffprobe, rsync, the `chough`/`parakeet` backends, saved-video and sites dirs, and the remote-worker token).
## Pipeline modes
@@ -53,7 +63,7 @@ The editor's pipeline panel and `common/ytdlp/runYtdlp.ts` support three modes p
- **Download from playlist** — reads `<channel>/playlist`, parses `<channel>/archive`, writes `<channel>/playlist.tofetch` containing only URLs whose ID is *not* already archived (this app-side prefilter avoids forcing yt-dlp to re-fetch metadata for known entries on platforms like Odysee), then runs yt-dlp with `--download-archive ./archive --break-on-existing -a playlist.tofetch` plus per-handling args.
- **Sync** — `yt-dlp --lazy-playlist --download-archive ./archive --break-on-existing <url>`. Quick incremental fetch.
-YouTube channels (`handling: "youtube"`) use `--write-auto-subs --skip-download`. Transcribe channels (`handling: "transcribe"`) download audio only (`-f bestaudio -x --audio-format <m4a|mp3|opus>`); whisper-cpp transcribes them later via the whisper panel.
+YouTube channels (`handling: "youtube"`) use `--write-auto-subs --skip-download`. Transcribe channels (`handling: "transcribe"`) download audio only (`-f bestaudio -x --audio-format <m4a|mp3|opus>`); a transcription backend transcribes them later via the whisper panel. whisper.cpp is the default; `chough` and `parakeet.cpp` are selectable in **Settings → Transcription** (see [SETUP.md](SETUP.md#transcription-backends-optional)).
## End-to-end tests
diff --git a/SETUP.md b/SETUP.md
@@ -0,0 +1,272 @@
+# Setup & requirements
+
+This guide takes you from a fresh clone to a running editor and static export on
+**Linux, macOS, or Windows**. It covers what to install, in what order, and which
+pieces you only need for specific features.
+
+- [Requirements at a glance](#requirements-at-a-glance)
+- [Install the toolchain](#install-the-toolchain)
+ - [Linux](#linux)
+ - [macOS](#macos)
+ - [Windows](#windows)
+- [Get the code running](#get-the-code-running)
+- [The `transcripts/` data directory](#the-transcripts-data-directory)
+- [Transcription backends (optional)](#transcription-backends-optional)
+- [Configuration & environment variables](#configuration--environment-variables)
+- [Running the tests](#running-the-tests)
+- [Where to go next](#where-to-go-next)
+
+---
+
+## Requirements at a glance
+
+**Always required** to install dependencies and run the web apps (editor, export,
+homepage):
+
+| Tool | Version | Notes |
+| --- | --- | --- |
+| **Node.js** | **≥ 20.9** (LTS 20 or 22) | Required by Next.js 16. The code is typed against Node 20. |
+| **pnpm** | **9+** | Lockfile is v9. Easiest via Corepack (bundled with Node) — see below. |
+| **git** | any recent | To clone the repo. |
+| **C/C++ toolchain** | platform default | Only if pnpm can't find a prebuilt binary for a native module (`lmdb`, `sharp`, …). Usually not needed on mainstream platforms. |
+
+> The web apps run with **none** of the media binaries below. You only need those to
+> use the download/transcribe **pipeline** (fetching videos, generating transcripts).
+
+**Feature-specific — install only what you'll use:**
+
+| Tool | Needed for | Default lookup |
+| --- | --- | --- |
+| **yt-dlp** | Downloading / syncing channels (the whole pipeline). | `yt-dlp` on `PATH` |
+| **ffmpeg** + **ffprobe** | Audio transcode + duration checks for `transcribe` channels. | `ffmpeg` / `ffprobe` on `PATH` |
+| A **transcription backend** | `handling: "transcribe"` channels only. Default is **whisper.cpp** (`whisper-cli`); `chough` and `parakeet.cpp` are alternatives. | `whisper-cli` on `PATH` |
+| **rsync** | Backing up the saved-video store. | `rsync` on `PATH` |
+| **Docker** | The sharded parallel e2e run (`pnpm e2e:sharded`) only. | — |
+
+Every binary above is overridable by an environment variable (e.g. `YTDLP_BIN`) —
+see [Configuration & environment variables](#configuration--environment-variables).
+
+---
+
+## Install the toolchain
+
+Enable **Corepack** once (it ships with Node) so the repo's pnpm version is used
+automatically:
+
+```sh
+corepack enable
+```
+
+### Linux
+
+Node (use your distro, or a version manager like [`fnm`](https://github.com/Schniz/fnm)
+/ `nvm` to get ≥ 20.9), then the media tools:
+
+```sh
+# Debian / Ubuntu
+sudo apt install nodejs yt-dlp ffmpeg rsync build-essential
+corepack enable
+
+# Fedora
+sudo dnf install nodejs yt-dlp ffmpeg rsync gcc-c++ make
+corepack enable
+
+# Arch
+sudo pacman -S nodejs yt-dlp ffmpeg rsync base-devel
+corepack enable
+```
+
+If your distro's `nodejs` is older than 20.9, install Node via `fnm`/`nvm` instead:
+
+```sh
+fnm install 22 && fnm use 22 # or: nvm install 22 && nvm use 22
+corepack enable
+```
+
+`build-essential` / `base-devel` / `gcc-c++ make` provides the C/C++ compiler pnpm
+falls back to when a native module (`lmdb`, `sharp`, …) has no prebuilt binary for
+your platform.
+
+### macOS
+
+Install the Xcode Command Line Tools (compiler for native modules), then use
+[Homebrew](https://brew.sh):
+
+```sh
+xcode-select --install # C/C++ toolchain
+brew install node pnpm yt-dlp ffmpeg rsync # rsync ships with macOS but brew's is newer
+corepack enable
+```
+
+`brew install pnpm` is optional if you use Corepack — either works.
+
+### Windows
+
+**Recommended: use WSL2.** The transcription toolchain (whisper.cpp / parakeet.cpp)
+and helper shell scripts assume a Unix shell, so the smoothest path on Windows is
+[WSL2](https://learn.microsoft.com/windows/wsl/install):
+
+```powershell
+wsl --install -d Ubuntu
+```
+
+Then open the Ubuntu shell and follow the [**Linux**](#linux) steps above. Keep the
+repo **inside** the WSL filesystem (e.g. `~/yt-dlp-transcript-browser`), not under
+`/mnt/c/…`, for good file-watch and I/O performance.
+
+**Native-Windows notes.** The web apps and the yt-dlp pipeline work natively too.
+Install with [winget](https://learn.microsoft.com/windows/package-manager/) (or
+[Scoop](https://scoop.sh)):
+
+```powershell
+winget install OpenJS.NodeJS.LTS
+winget install yt-dlp.yt-dlp
+winget install Gyan.FFmpeg
+corepack enable
+```
+
+Caveats on native Windows:
+
+- Building the **whisper.cpp / parakeet.cpp** transcription backends is Unix-oriented;
+ do transcription work under WSL2.
+- `create-archives.sh` and the external-cron sync path (`pnpm sync:tick` from cron)
+ assume a Unix shell — use WSL2 or Task Scheduler equivalents.
+
+---
+
+## Get the code running
+
+```sh
+git clone <this-repo-url> yt-dlp-transcript-browser
+cd yt-dlp-transcript-browser
+
+pnpm install # installs all workspace packages; compiles native modules
+ # (lmdb, msgpackr-extract, sharp, …) per the build allowlist
+ # in pnpm-workspace.yaml
+
+pnpm dev:editor # editor (admin UI) at http://localhost:3001
+```
+
+To build and serve the read-only public site:
+
+```sh
+pnpm build # static site under export/out/
+pnpm start:export # serve export/out/ at http://localhost:3000
+```
+
+The editor **starts fine with no data** — a fresh clone has no `transcripts/`
+directory (see below) and that's expected. Create your first channel through the
+editor's **/channels** page; the pipeline populates `transcripts/` from there.
+
+---
+
+## The `transcripts/` data directory
+
+All downloaded data — channels, archives, the LMDB index, job logs — lives under
+`transcripts/` at the repo root. It is a **separate git repo**, deliberately
+**not** part of this workspace (`/transcripts` is in the root `.gitignore`), so a
+fresh clone of this repo does **not** include it.
+
+There is nothing to bootstrap: the app **creates the directory and its
+subdirectories lazily** as you use it. Start the editor, add channels through
+**/channels**, and run the pipeline — `transcripts/channels/<slug>/…`,
+`transcripts/index.mdb`, etc. appear on demand.
+
+Override the location with `TRANSCRIPTS_DIR` (e.g. to put data on another disk).
+`common/lib/paths.ts` is the source of truth for every path and its env override.
+
+---
+
+## Transcription backends (optional)
+
+You only need a transcription backend for channels with `handling: "transcribe"`
+(audio-only downloads that get transcribed locally). YouTube channels
+(`handling: "youtube"`) fetch existing captions and need no backend.
+
+The active backend is chosen globally in **Settings → Transcription**; three are
+registered in `common/lib/transcriptionApps.ts`:
+
+- **whisper.cpp** (default, `whisper-cli`) — build
+ [whisper.cpp](https://github.com/ggml-org/whisper.cpp), download a `ggml` model,
+ and either put `whisper-cli` on your `PATH` or set `WHISPER_BIN`. Point
+ `WHISPER_MODEL` at the model (default
+ `~/whispercpp/whisper.cpp/models/ggml-base.en.bin`), or set the binary/model in the
+ Settings UI.
+ On macOS you can skip the build: `brew install whisper-cpp`.
+- **chough** — set the binary via `CHOUGH_BIN`; auto-downloads a model when none is
+ configured. Supports a remote server (`CHOUGH_URL`).
+- **parakeet.cpp** — driven by the bundled overlapping-segment wrapper
+ (`scripts/parakeet-stitch.mjs`); needs `parakeet-cli` (`PARAKEET_CLI`) and a `.gguf`
+ model (`PARAKEET_MODEL`). Interruptible (stops and stitches a partial transcript).
+
+All three read `ffmpeg`/`ffprobe`, so make sure those are installed for transcribe
+channels.
+
+---
+
+## Configuration & environment variables
+
+Most configuration now lives in the editor's **/settings** page, persisted to
+`settings.json` at the repo root (gitignored). `settings.json.example` is a minimal
+starting template. Settings are optional — a missing/partial `settings.json` falls
+back to built-in defaults, so the app runs out of the box.
+
+Paths and binaries resolve through `getPaths()` in `common/lib/paths.ts`. Override
+any of them via environment variables before launching:
+
+| Variable | Default | Purpose |
+| --- | --- | --- |
+| `TRANSCRIPTS_DIR` | `<repo>/transcripts` | Channels, archives, LMDB index, job logs. |
+| `SAVED_VIDEOS_DIR` | `<TRANSCRIPTS_DIR>/saved-videos` | Persisted source-video store (can live on a separate disk). |
+| `SITES_DIR` | `<TRANSCRIPTS_DIR>/sites` | Per-site config (`sites/<id>/site.json`). |
+| `EXPORT_PUBLIC_DIR` | `<repo>/export/public` | Where the index writes paginated JSON. |
+| `SETTINGS_FILE` | `<repo>/settings.json` | Site-settings file. |
+| `YTDLP_BIN` | `yt-dlp` (PATH) | Pipeline downloader. |
+| `WHISPER_BIN` | `whisper-cli` (PATH) | whisper.cpp binary. |
+| `WHISPER_MODEL` | `~/whispercpp/whisper.cpp/models/ggml-base.en.bin` | whisper.cpp model file. |
+| `CHOUGH_BIN` / `CHOUGH_URL` / `CHOUGH_MODEL` | `chough` / — / — | chough backend binary, remote server, model. |
+| `PARAKEET_CLI` / `PARAKEET_MODEL` / `PARAKEET_STITCH_BIN` | `parakeet-cli` / — / `scripts/parakeet-stitch.mjs` | parakeet.cpp CLI, model, and wrapper. |
+| `FFMPEG_BIN` / `FFPROBE_BIN` | `ffmpeg` / `ffprobe` (PATH) | Audio transcode + duration checks. |
+| `RSYNC_BIN` | `rsync` (PATH) | Saved-video backup. |
+| `PARALLEL_TRANSCRIBE_LIMIT` | `4` | Max parallel transcription jobs. |
+| `WORKER_TOKEN` | — | Bearer token for the remote-worker transcription API (set on both ends when used). |
+
+Feature-area docs cover their own env vars: [SCHEDULED_SYNC.md](SCHEDULED_SYNC.md)
+(`SYNC_HEARTBEAT_SECONDS`, `SYNC_TICK_URL`, `SYNC_TICK_TOKEN`) and
+[DEPLOY_CLOUDFLARE.md](DEPLOY_CLOUDFLARE.md) (R2 credentials).
+
+---
+
+## Running the tests
+
+The editor's Playwright suite drives the admin UI. It uses fake binaries for
+yt-dlp/whisper/ffmpeg, so you do **not** need the media tools installed to run it —
+but you do need the Playwright browser:
+
+```sh
+npx playwright install chromium # one-time: download the test browser
+pnpm e2e # sequential run on the host (next dev, port 3001)
+```
+
+For a faster parallel run, `pnpm e2e:sharded` splits the suite across N Docker
+containers (needs **Docker**; the container image ships the browser, so no
+`playwright install` there). Shard count defaults to `min(max(2, cpus/2), 8)`;
+override with `SHARDS=N pnpm e2e:sharded`.
+
+To run multiple checkouts / dev servers / e2e suites at once without port
+collisions, see [WORKTREES.md](WORKTREES.md).
+
+---
+
+## Where to go next
+
+- [README.md](README.md) — project overview, pipeline modes, CLI shims.
+- [SCHEDULED_SYNC.md](SCHEDULED_SYNC.md) — automatic per-channel sync (internal
+ heartbeat or external cron).
+- [DEPLOY_CLOUDFLARE.md](DEPLOY_CLOUDFLARE.md) — deploying to Cloudflare Pages + R2
+ archive overflow.
+- [WORKTREES.md](WORKTREES.md) — parallel development with per-worktree ports.
+- [mcp/README.md](mcp/README.md) — MCP server exposing the archive to Claude Code /
+ Desktop / Cursor.
+- [r2-proxy/README.md](r2-proxy/README.md) — the Cloudflare Worker that serves
+ oversize archives from R2.