Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 5708cc3ee7b394cccf515aea581d8c9dc6427004
parent 64f0a6b2ea8d9b79580df8ed14d8b69b84e4870d
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Thu, 16 Jul 2026 01:34:14 -0400

Add setup docs

Diffstat:
MREADME.md | 20+++++++++++++++-----
ASETUP.md | 272+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
2 files changed, 287 insertions(+), 5 deletions(-)

diff --git a/README.md b/README.md @@ -8,6 +8,13 @@ A pnpm-workspace monorepo for archiving and browsing video transcripts. The proj Transcripts and per-channel state live at `<repo>/transcripts/` (its own git repo, untouched by the workspace). +## Requirements + +- **Node.js ≥ 20.9** (LTS 20 or 22) and **pnpm 9+** (easiest via `corepack enable`) and **git** — needed to install and run the web apps. +- The download/transcribe **pipeline** additionally needs **yt-dlp**, **ffmpeg**/**ffprobe**, and a transcription backend (whisper.cpp by default). The web apps run without them. + +**See [SETUP.md](SETUP.md) for full per-OS install instructions (Linux, macOS, Windows).** + ## Getting started ```bash @@ -15,6 +22,7 @@ pnpm install pnpm dev:editor # editor at http://localhost:3001 pnpm build # build the static site under export/out/ pnpm start:export # serve export/out/ at http://localhost:3000 +npx playwright install chromium # one-time, before the first e2e run pnpm e2e # run the editor's Playwright suite (sequential) pnpm e2e:sharded # same suite, sharded across N Docker containers in parallel ``` @@ -31,9 +39,9 @@ pnpm e2e:sharded # same suite, sharded across N Docker containers in paralle Channels can also sync automatically on a per-channel cadence via a cron heartbeat — see [SCHEDULED_SYNC.md](SCHEDULED_SYNC.md). -## Environment variables +## Configuration & environment variables -All paths and binaries used by the editor and CLI shims resolve through `getPaths()`. Override any of them before launching: +Most configuration lives in the editor's **/settings** page (persisted to `settings.json`); it's optional and falls back to defaults. All paths and binaries resolve through `getPaths()` (`common/lib/paths.ts`) and can be overridden by environment variables. Common ones: | Variable | Default | Purpose | | --- | --- | --- | @@ -41,9 +49,11 @@ All paths and binaries used by the editor and CLI shims resolve through `getPath | `EXPORT_PUBLIC_DIR` | `<repo>/export/public` | Where the index writes paginated JSON. | | `SETTINGS_FILE` | `<repo>/settings.json` | Site-settings file. | | `YTDLP_BIN` | `yt-dlp` (PATH lookup) | Binary used by the pipeline. | -| `WHISPER_BIN` | `whisper-cli` (PATH lookup) | Binary used by the whisper batch. | +| `WHISPER_BIN` | `whisper-cli` (PATH lookup) | Default transcription binary. | | `WHISPER_MODEL` | `~/whispercpp/whisper.cpp/models/ggml-base.en.bin` | Model file passed to whisper-cli. | -| `PARALLEL_TRANSCRIBE_LIMIT` | `4` | Max parallel whisper jobs. | +| `PARALLEL_TRANSCRIBE_LIMIT` | `4` | Max parallel transcription jobs. | + +See **[SETUP.md](SETUP.md#configuration--environment-variables)** for the full list (ffmpeg/ffprobe, rsync, the `chough`/`parakeet` backends, saved-video and sites dirs, and the remote-worker token). ## Pipeline modes @@ -53,7 +63,7 @@ The editor's pipeline panel and `common/ytdlp/runYtdlp.ts` support three modes p - **Download from playlist** — reads `<channel>/playlist`, parses `<channel>/archive`, writes `<channel>/playlist.tofetch` containing only URLs whose ID is *not* already archived (this app-side prefilter avoids forcing yt-dlp to re-fetch metadata for known entries on platforms like Odysee), then runs yt-dlp with `--download-archive ./archive --break-on-existing -a playlist.tofetch` plus per-handling args. - **Sync** — `yt-dlp --lazy-playlist --download-archive ./archive --break-on-existing <url>`. Quick incremental fetch. -YouTube channels (`handling: "youtube"`) use `--write-auto-subs --skip-download`. Transcribe channels (`handling: "transcribe"`) download audio only (`-f bestaudio -x --audio-format <m4a|mp3|opus>`); whisper-cpp transcribes them later via the whisper panel. +YouTube channels (`handling: "youtube"`) use `--write-auto-subs --skip-download`. Transcribe channels (`handling: "transcribe"`) download audio only (`-f bestaudio -x --audio-format <m4a|mp3|opus>`); a transcription backend transcribes them later via the whisper panel. whisper.cpp is the default; `chough` and `parakeet.cpp` are selectable in **Settings → Transcription** (see [SETUP.md](SETUP.md#transcription-backends-optional)). ## End-to-end tests diff --git a/SETUP.md b/SETUP.md @@ -0,0 +1,272 @@ +# Setup & requirements + +This guide takes you from a fresh clone to a running editor and static export on +**Linux, macOS, or Windows**. It covers what to install, in what order, and which +pieces you only need for specific features. + +- [Requirements at a glance](#requirements-at-a-glance) +- [Install the toolchain](#install-the-toolchain) + - [Linux](#linux) + - [macOS](#macos) + - [Windows](#windows) +- [Get the code running](#get-the-code-running) +- [The `transcripts/` data directory](#the-transcripts-data-directory) +- [Transcription backends (optional)](#transcription-backends-optional) +- [Configuration & environment variables](#configuration--environment-variables) +- [Running the tests](#running-the-tests) +- [Where to go next](#where-to-go-next) + +--- + +## Requirements at a glance + +**Always required** to install dependencies and run the web apps (editor, export, +homepage): + +| Tool | Version | Notes | +| --- | --- | --- | +| **Node.js** | **≥ 20.9** (LTS 20 or 22) | Required by Next.js 16. The code is typed against Node 20. | +| **pnpm** | **9+** | Lockfile is v9. Easiest via Corepack (bundled with Node) — see below. | +| **git** | any recent | To clone the repo. | +| **C/C++ toolchain** | platform default | Only if pnpm can't find a prebuilt binary for a native module (`lmdb`, `sharp`, …). Usually not needed on mainstream platforms. | + +> The web apps run with **none** of the media binaries below. You only need those to +> use the download/transcribe **pipeline** (fetching videos, generating transcripts). + +**Feature-specific — install only what you'll use:** + +| Tool | Needed for | Default lookup | +| --- | --- | --- | +| **yt-dlp** | Downloading / syncing channels (the whole pipeline). | `yt-dlp` on `PATH` | +| **ffmpeg** + **ffprobe** | Audio transcode + duration checks for `transcribe` channels. | `ffmpeg` / `ffprobe` on `PATH` | +| A **transcription backend** | `handling: "transcribe"` channels only. Default is **whisper.cpp** (`whisper-cli`); `chough` and `parakeet.cpp` are alternatives. | `whisper-cli` on `PATH` | +| **rsync** | Backing up the saved-video store. | `rsync` on `PATH` | +| **Docker** | The sharded parallel e2e run (`pnpm e2e:sharded`) only. | — | + +Every binary above is overridable by an environment variable (e.g. `YTDLP_BIN`) — +see [Configuration & environment variables](#configuration--environment-variables). + +--- + +## Install the toolchain + +Enable **Corepack** once (it ships with Node) so the repo's pnpm version is used +automatically: + +```sh +corepack enable +``` + +### Linux + +Node (use your distro, or a version manager like [`fnm`](https://github.com/Schniz/fnm) +/ `nvm` to get ≥ 20.9), then the media tools: + +```sh +# Debian / Ubuntu +sudo apt install nodejs yt-dlp ffmpeg rsync build-essential +corepack enable + +# Fedora +sudo dnf install nodejs yt-dlp ffmpeg rsync gcc-c++ make +corepack enable + +# Arch +sudo pacman -S nodejs yt-dlp ffmpeg rsync base-devel +corepack enable +``` + +If your distro's `nodejs` is older than 20.9, install Node via `fnm`/`nvm` instead: + +```sh +fnm install 22 && fnm use 22 # or: nvm install 22 && nvm use 22 +corepack enable +``` + +`build-essential` / `base-devel` / `gcc-c++ make` provides the C/C++ compiler pnpm +falls back to when a native module (`lmdb`, `sharp`, …) has no prebuilt binary for +your platform. + +### macOS + +Install the Xcode Command Line Tools (compiler for native modules), then use +[Homebrew](https://brew.sh): + +```sh +xcode-select --install # C/C++ toolchain +brew install node pnpm yt-dlp ffmpeg rsync # rsync ships with macOS but brew's is newer +corepack enable +``` + +`brew install pnpm` is optional if you use Corepack — either works. + +### Windows + +**Recommended: use WSL2.** The transcription toolchain (whisper.cpp / parakeet.cpp) +and helper shell scripts assume a Unix shell, so the smoothest path on Windows is +[WSL2](https://learn.microsoft.com/windows/wsl/install): + +```powershell +wsl --install -d Ubuntu +``` + +Then open the Ubuntu shell and follow the [**Linux**](#linux) steps above. Keep the +repo **inside** the WSL filesystem (e.g. `~/yt-dlp-transcript-browser`), not under +`/mnt/c/…`, for good file-watch and I/O performance. + +**Native-Windows notes.** The web apps and the yt-dlp pipeline work natively too. +Install with [winget](https://learn.microsoft.com/windows/package-manager/) (or +[Scoop](https://scoop.sh)): + +```powershell +winget install OpenJS.NodeJS.LTS +winget install yt-dlp.yt-dlp +winget install Gyan.FFmpeg +corepack enable +``` + +Caveats on native Windows: + +- Building the **whisper.cpp / parakeet.cpp** transcription backends is Unix-oriented; + do transcription work under WSL2. +- `create-archives.sh` and the external-cron sync path (`pnpm sync:tick` from cron) + assume a Unix shell — use WSL2 or Task Scheduler equivalents. + +--- + +## Get the code running + +```sh +git clone <this-repo-url> yt-dlp-transcript-browser +cd yt-dlp-transcript-browser + +pnpm install # installs all workspace packages; compiles native modules + # (lmdb, msgpackr-extract, sharp, …) per the build allowlist + # in pnpm-workspace.yaml + +pnpm dev:editor # editor (admin UI) at http://localhost:3001 +``` + +To build and serve the read-only public site: + +```sh +pnpm build # static site under export/out/ +pnpm start:export # serve export/out/ at http://localhost:3000 +``` + +The editor **starts fine with no data** — a fresh clone has no `transcripts/` +directory (see below) and that's expected. Create your first channel through the +editor's **/channels** page; the pipeline populates `transcripts/` from there. + +--- + +## The `transcripts/` data directory + +All downloaded data — channels, archives, the LMDB index, job logs — lives under +`transcripts/` at the repo root. It is a **separate git repo**, deliberately +**not** part of this workspace (`/transcripts` is in the root `.gitignore`), so a +fresh clone of this repo does **not** include it. + +There is nothing to bootstrap: the app **creates the directory and its +subdirectories lazily** as you use it. Start the editor, add channels through +**/channels**, and run the pipeline — `transcripts/channels/<slug>/…`, +`transcripts/index.mdb`, etc. appear on demand. + +Override the location with `TRANSCRIPTS_DIR` (e.g. to put data on another disk). +`common/lib/paths.ts` is the source of truth for every path and its env override. + +--- + +## Transcription backends (optional) + +You only need a transcription backend for channels with `handling: "transcribe"` +(audio-only downloads that get transcribed locally). YouTube channels +(`handling: "youtube"`) fetch existing captions and need no backend. + +The active backend is chosen globally in **Settings → Transcription**; three are +registered in `common/lib/transcriptionApps.ts`: + +- **whisper.cpp** (default, `whisper-cli`) — build + [whisper.cpp](https://github.com/ggml-org/whisper.cpp), download a `ggml` model, + and either put `whisper-cli` on your `PATH` or set `WHISPER_BIN`. Point + `WHISPER_MODEL` at the model (default + `~/whispercpp/whisper.cpp/models/ggml-base.en.bin`), or set the binary/model in the + Settings UI. + On macOS you can skip the build: `brew install whisper-cpp`. +- **chough** — set the binary via `CHOUGH_BIN`; auto-downloads a model when none is + configured. Supports a remote server (`CHOUGH_URL`). +- **parakeet.cpp** — driven by the bundled overlapping-segment wrapper + (`scripts/parakeet-stitch.mjs`); needs `parakeet-cli` (`PARAKEET_CLI`) and a `.gguf` + model (`PARAKEET_MODEL`). Interruptible (stops and stitches a partial transcript). + +All three read `ffmpeg`/`ffprobe`, so make sure those are installed for transcribe +channels. + +--- + +## Configuration & environment variables + +Most configuration now lives in the editor's **/settings** page, persisted to +`settings.json` at the repo root (gitignored). `settings.json.example` is a minimal +starting template. Settings are optional — a missing/partial `settings.json` falls +back to built-in defaults, so the app runs out of the box. + +Paths and binaries resolve through `getPaths()` in `common/lib/paths.ts`. Override +any of them via environment variables before launching: + +| Variable | Default | Purpose | +| --- | --- | --- | +| `TRANSCRIPTS_DIR` | `<repo>/transcripts` | Channels, archives, LMDB index, job logs. | +| `SAVED_VIDEOS_DIR` | `<TRANSCRIPTS_DIR>/saved-videos` | Persisted source-video store (can live on a separate disk). | +| `SITES_DIR` | `<TRANSCRIPTS_DIR>/sites` | Per-site config (`sites/<id>/site.json`). | +| `EXPORT_PUBLIC_DIR` | `<repo>/export/public` | Where the index writes paginated JSON. | +| `SETTINGS_FILE` | `<repo>/settings.json` | Site-settings file. | +| `YTDLP_BIN` | `yt-dlp` (PATH) | Pipeline downloader. | +| `WHISPER_BIN` | `whisper-cli` (PATH) | whisper.cpp binary. | +| `WHISPER_MODEL` | `~/whispercpp/whisper.cpp/models/ggml-base.en.bin` | whisper.cpp model file. | +| `CHOUGH_BIN` / `CHOUGH_URL` / `CHOUGH_MODEL` | `chough` / — / — | chough backend binary, remote server, model. | +| `PARAKEET_CLI` / `PARAKEET_MODEL` / `PARAKEET_STITCH_BIN` | `parakeet-cli` / — / `scripts/parakeet-stitch.mjs` | parakeet.cpp CLI, model, and wrapper. | +| `FFMPEG_BIN` / `FFPROBE_BIN` | `ffmpeg` / `ffprobe` (PATH) | Audio transcode + duration checks. | +| `RSYNC_BIN` | `rsync` (PATH) | Saved-video backup. | +| `PARALLEL_TRANSCRIBE_LIMIT` | `4` | Max parallel transcription jobs. | +| `WORKER_TOKEN` | — | Bearer token for the remote-worker transcription API (set on both ends when used). | + +Feature-area docs cover their own env vars: [SCHEDULED_SYNC.md](SCHEDULED_SYNC.md) +(`SYNC_HEARTBEAT_SECONDS`, `SYNC_TICK_URL`, `SYNC_TICK_TOKEN`) and +[DEPLOY_CLOUDFLARE.md](DEPLOY_CLOUDFLARE.md) (R2 credentials). + +--- + +## Running the tests + +The editor's Playwright suite drives the admin UI. It uses fake binaries for +yt-dlp/whisper/ffmpeg, so you do **not** need the media tools installed to run it — +but you do need the Playwright browser: + +```sh +npx playwright install chromium # one-time: download the test browser +pnpm e2e # sequential run on the host (next dev, port 3001) +``` + +For a faster parallel run, `pnpm e2e:sharded` splits the suite across N Docker +containers (needs **Docker**; the container image ships the browser, so no +`playwright install` there). Shard count defaults to `min(max(2, cpus/2), 8)`; +override with `SHARDS=N pnpm e2e:sharded`. + +To run multiple checkouts / dev servers / e2e suites at once without port +collisions, see [WORKTREES.md](WORKTREES.md). + +--- + +## Where to go next + +- [README.md](README.md) — project overview, pipeline modes, CLI shims. +- [SCHEDULED_SYNC.md](SCHEDULED_SYNC.md) — automatic per-channel sync (internal + heartbeat or external cron). +- [DEPLOY_CLOUDFLARE.md](DEPLOY_CLOUDFLARE.md) — deploying to Cloudflare Pages + R2 + archive overflow. +- [WORKTREES.md](WORKTREES.md) — parallel development with per-worktree ports. +- [mcp/README.md](mcp/README.md) — MCP server exposing the archive to Claude Code / + Desktop / Cursor. +- [r2-proxy/README.md](r2-proxy/README.md) — the Cloudflare Worker that serves + oversize archives from R2.