Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 73018db4c2daf5dca526c63f4741a58326bd55c6
parent d9204a6e301d8ac1774fac6c51a53edfefa1d63d
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Mon, 28 Sep 2026 02:04:13 -0400

docs: PUBLISH.md absorbs DEPLOY_CLOUDFLARE.md and DEPLOY_DOCKER.md; env vars point at ENVIRONMENT.md

PUBLISH.md is the one document for publishing: what gets published (a site,
the hub, the homepage — two Pages projects), the three ways to drive it (the
editor's /sites, `pnpm ops`, `pnpm archilyzer`) in one table, Cloudflare
Pages (a project must exist before its first deploy; a production deploy
inherits the checkout's branch), previews, download archives and R2, the
cost-abuse defenses, the container build pipeline, and registering the MCP
server against a published archive with `-- pnpm -C "$PWD" archilyzer mcp`.
The two old files are removed and every link to them outside plans/ and the
released changelog bullets now names PUBLISH.md (README, SETUP, AGENTS,
RUNNING_IN_DOCKER, r2-proxy, homepage/content/README.md).

README's and SETUP's hand-kept env-var tables become a paragraph and a
pointer to the generated ENVIRONMENT.md plus `pnpm archilyzer doctor`.
SETUP: `pnpm e2e` is port 3011, not 3001. CONTRIBUTING's "CLI shims"
(which named a `retry-failures.ts` that does not exist) is "The archilyzer
CLI"; a new path override is declared in envVars.ts, a new port in
ports.mjs. RUNNING_IN_DOCKER names parakeet.cpp beside whisper.cpp.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
MAGENTS.md | 2+-
MCONTRIBUTING.md | 34+++++++++++++++++++++++++---------
DDEPLOY_CLOUDFLARE.md | 332-------------------------------------------------------------------------------
DDEPLOY_DOCKER.md | 82-------------------------------------------------------------------------------
APUBLISH.md | 443+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
MREADME.md | 42+++++++++++++++++-------------------------
MRUNNING_IN_DOCKER.md | 15++++++++-------
MSETUP.md | 52++++++++++++++++++++++++++--------------------------
Mhomepage/content/README.md | 6+++---
Mr2-proxy/README.md | 2+-
Mr2-proxy/src/index.ts | 2+-
Mr2-proxy/wrangler.toml | 2+-
12 files changed, 526 insertions(+), 488 deletions(-)

diff --git a/AGENTS.md b/AGENTS.md @@ -208,7 +208,7 @@ back to the serial host build it already handles. Do not try to make docker-in-docker work. See [RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md) — which is about *running the -apps*, not [DEPLOY_DOCKER.md](DEPLOY_DOCKER.md), which is about *building sites*. +apps*, not [PUBLISH.md](PUBLISH.md), which is about *building and publishing sites*. # The corpus, and where the live sites are configured diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md @@ -34,7 +34,8 @@ pnpm install pnpm dev:editor # editor at http://localhost:3001 pnpm build # build the static site under export/out/ pnpm start:export # serve export/out/ at http://localhost:3000 -pnpm build:index # rebuild the search index in-process +pnpm build:index # rebuild the search index in-process (= pnpm archilyzer index) +pnpm archilyzer doctor # what this machine has: tools, corpus, settings, ports ``` `pnpm lint` at the root only lints `export/`. **The editor has no eslint config**, so @@ -128,18 +129,29 @@ buttons" is red under `next dev` (which injects a Dev Tools button) and green un **The editor e2e suite runs in dev mode and therefore never prerenders.** A change to a layout or a client component is not verified until `pnpm build` passes too. -## CLI shims +## The `archilyzer` CLI -The same controllers the editor uses are exposed as terminal shims under -`common/bin/`, which is how you drive the pipeline headlessly or from cron: +The same controllers the editor uses are one command line, `common/bin/archilyzer.ts`, +which is how you drive the pipeline headlessly or from cron. `pnpm archilyzer +<command>` from the repo root is the short form of `pnpm --filter +yt-dlp-transcript-common exec tsx bin/archilyzer.ts <command>`; `pnpm archilyzer +--help` lists every command. ```bash -pnpm --filter yt-dlp-transcript-common exec tsx bin/build-index.ts -pnpm --filter yt-dlp-transcript-common exec tsx bin/transform.ts --channel <slug> -pnpm --filter yt-dlp-transcript-common exec tsx bin/retry-failures.ts --channel <slug> -pnpm --filter yt-dlp-transcript-common exec tsx bin/verify-transcripts.ts --channel <slug> +pnpm archilyzer doctor # read-only: can this machine do what it is configured to? +pnpm archilyzer index # the LMDB index +pnpm archilyzer run diarization <channel> [ids…] # one catalogued operation, offline, as the editor's job +pnpm archilyzer build site <id> # publish: see PUBLISH.md +pnpm archilyzer verify transcripts --channel <slug> +pnpm archilyzer mcp # the MCP server on stdio ``` +The table is `archilyzer.ts`; the machinery (parser, lookup, usage) is `_cli.ts`. Every +file in `common/bin/` is reachable from a row — a test fails otherwise. A bin that +parses its own flags is a *passthrough* row, run as a child with its argv untouched. +`run` refuses sync, the metadata scan, downloads and transcription: they run on the +editor's paced download queue and worker pool, which a second process must not race. + ## Internals worth knowing **This is not the Next.js you may know.** The workspace tracks a recent major and its @@ -155,7 +167,11 @@ deprecation notices. `export/public/{summaries,transcripts}/` and renders a search/filter UI into a static `out/`. - `getPaths()` (`common/lib/paths.ts`) is the single resolver for every path and - binary. Nothing should hardcode a location; add an env override there instead. + binary. Nothing should hardcode a location; add an env override there instead, and + declare it in `common/lib/envVars.ts` (a test fails until you do, and + `pnpm archilyzer docs env` regenerates [ENVIRONMENT.md](ENVIRONMENT.md)). A new + local server's port goes in `common/lib/ports.mjs`, which `pnpm wt` offsets per + worktree. - `common/lib/project.ts` holds **product** identity (the name, the project URL) and is deliberately import-free. **Operator** identity — what a given deployment calls itself — lives in settings and per-site config. A string that should change when diff --git a/DEPLOY_CLOUDFLARE.md b/DEPLOY_CLOUDFLARE.md @@ -1,332 +0,0 @@ -# Deploying to Cloudflare (Pages + R2 archive overflow) - -An **Archilyzer** site is deployed to **Cloudflare Pages**; large download archives that -exceed Pages' per-file limit overflow to **Cloudflare R2**. This guide covers the R2 -setup and — importantly — how to configure Cloudflare so that **public archive -downloads can't be abused to drive up your bill**. - -Everything here fits inside Cloudflare's **free tier**. - -- [How archives are served](#how-archives-are-served) -- [Preview deployments](#preview-deployments) -- [One-time R2 setup](#one-time-r2-setup) -- [Securing downloads against cost-abuse](#securing-downloads-against-cost-abuse) -- [Cost expectations](#cost-expectations) - ---- - -## How archives are served - -At build time, `compose-site.ts` generates one transcript zip and one live-chat zip -**per channel** into `export/public/archives/`, and records them in -`public/archives/manifest.json`. The `/downloads` page and the header **Downloads** -link read that manifest. - -Cloudflare Pages rejects any single asset larger than **25 MB**. Real channels blow -past that easily (a channel's live-chat zip can be hundreds of MB). So the pipeline -splits archives by size: - -| Archive size | Where it's served from | Manifest entry | -|---|---|---| -| ≤ 25 MB | Cloudflare **Pages** (shipped in `out/`, free) | `filename`, no `url` | -| > 25 MB, R2 configured | Cloudflare **R2**, uploaded on deploy | `url` → R2 | -| > 25 MB, R2 **not** configured | not served | `oversize: true`, shown as "Too large to host" | - -Oversize archives are staged during compose into `export/.r2-staging/<siteId>/archives/` -(gitignored, kept out of `public/`), then uploaded by the editor's **Deploy** / -**Build & deploy** actions to `<bucket>/<siteId>/archives/<file>.zip` **before** the -Pages deploy runs, so the manifest URLs resolve immediately. - -> **Note:** every deploy path uploads oversize archives to R2 before the Pages -> deploy: the editor's **Deploy** / **Build & deploy**, `pnpm ops deploy-site`, and -> `archilyzer deploy site <id>` (which `pnpm run deploy` in `export/` runs, with -> `SITE_ID`). A build alone (`archilyzer build site`, `pnpm run build`) only -> *stages* them in `export/.r2-staging/`. The bucket is read from `settings.json`; -> the R2 credentials come from the environment (see "3. Authenticate" below). - -Uploads go through R2's **S3 API** (via the AWS SDK's multipart uploader), not -`wrangler r2 object put` — wrangler caps a single upload at **300 MiB**, and real -live-chat archives are larger (multipart has no such limit). That's why archive -uploads need S3 credentials (step 3) in addition to the wrangler auth your Pages -deploy already uses. - ---- - -## Preview deployments - -A **preview** is the same built bundle deployed to a branch that is not the Pages -project's production branch. Cloudflare publishes it at a **branch alias** — - -``` -https://<branch>.<project>.pages.dev -``` - -— and leaves the live site alone. Each deploy also gets an immutable -per-deployment URL (`https://<hash>.<project>.pages.dev`), which wrangler prints as -"Deployment complete! Take a peek over at …"; the editor repeats both on a -`[preview]` line at the end of the job log, because the streamed log scrolls. - -The alias is a function of the project and the branch and nothing else, so it is -known *before* the deploy runs — which is why the editor can link it while you are -still typing the name. - -**Branch names** must be 1–28 lowercase letters, digits and dashes, starting and -ending with a letter or digit. That is exactly what Cloudflare's alias sanitizer -preserves verbatim, so the alias shown is the alias that resolves. `main`, `master` -and `production` are refused: a deploy to the production branch is not a preview, it -is the live site. - -### The three surfaces - -| Surface | How | -|---|---| -| Editor | A site's **Publish** tab → *Individual steps* → **Deploy a preview**: type a branch, press **Deploy preview**. The production button beside it says **Deploy to production**. | -| Ops API | `POST /api/ops/deploy-site` `{ "siteId": "...", "preview": "<branch>" }` — deploy-only, of the already-built `export/out`. `POST /api/ops/build-deploy` takes `preview` too (build *then* preview-deploy). Both answer with `previewUrl`. | -| CLI | `pnpm ops deploy-site --json '{"siteId":"anilyzer","preview":"tags-exclude"}' --wait` — the alias is printed on its own line after the log. | - -The high-value loop is **build once, preview, then promote**: `build-site` (or the -Publish tab's *Build static export*), then `deploy-site` with a `preview`, look at -it, then `deploy-site` again with no `preview` — the same `export/out`, unrebuilt. - -Deploy-only ships whatever is in `export/out`, which the basic build composes one -site at a time into a single shared directory — so it **refuses, before starting a -job, if `export/out` holds a build of another site** (or no build at all), naming -the site to build first. `build-deploy` cannot hit this: it builds. - -### Two things to know - -**A preview shares the production R2 archive bucket.** R2 has no per-branch -namespace, and the keys are `<siteId>/archives/<file>.zip` either way. In practice -this is cheap and harmless — the upload skips any object R2 already holds at the same -size, and an unchanged channel re-zips byte-stable — but a *changed* archive -replaces the one production's manifest links to. The editor says so once at the top -of every preview deploy. - -**Who can open a preview is a Cloudflare setting, not ours.** Pages projects have a -*preview deployment access* setting (Settings → General): **public** by default, or -restricted to Cloudflare Access. A default-configured project's preview URL is -world-readable by anyone who has the link. - -> **A production deploy inherits the checkout's git branch.** `wrangler pages deploy` -> with no `--branch` infers one from the git repository it runs in, so running a -> *production* deploy from a feature-branch checkout silently produces a preview -> instead. This predates the preview feature and is unchanged: only the preview path -> passes `--branch`. If a "production" deploy did not go live, check what branch the -> editor's checkout is on. - ---- - -## One-time R2 setup - -### 1. Create the bucket - -One bucket serves **all** your sites — objects are namespaced by `siteId`, so you do -**not** need one bucket per site. In the dashboard (**R2 → Create bucket**) or via the -same `wrangler` CLI the deploy already uses: - -```sh -wrangler r2 bucket create my-archives-bucket -``` - -### 2. Expose it publicly - -R2 buckets are private by default, and the Downloads page links to plain -`https://…/<file>.zip` URLs, so the bucket needs public read access. There are three -ways to expose it; the first needs **no domain purchase** and is the recommended one: - -- **A Worker on `*.workers.dev` (recommended — no domain, no WHOIS):** deploy the - bundled `r2-proxy/` Worker to a free `<name>.workers.dev` subdomain (just like Pages' - `*.pages.dev`). It streams objects from the bucket and gives you edge caching **and** - tunable rate limiting in code — the same cost defenses a custom domain would, without - owning a domain. One Worker serves **every** site. See - [Serving via a Worker on workers.dev](#option-a--serving-via-a-worker-on-workersdev-no-domain) - below. Set the editor's public URL to `https://<name>.workers.dev`. -- **Custom domain:** bucket **Settings → Public access → Custom Domains → Connect - Domain**, e.g. `archives.example.com` (the domain must be on your Cloudflare account). - Routes downloads through Cloudflare's CDN, WAF, caching, and dashboard rate-limiting - rules. See [Securing via a custom domain](#option-b--securing-via-a-custom-domain). -- **`r2.dev` subdomain (quick test only):** bucket **Settings → Public access → Allow - Access**. You get `https://pub-abc123.r2.dev` — free and domain-free, but Cloudflare - throttles `r2.dev` and gives you **no** cache / rate-limit control of your own. Fine - for a smoke test; use the Worker for a real instance. - -### 3. Authenticate — two credentials - -**a) `wrangler` (for the Pages deploy).** The `wrangler pages deploy` step reuses -whatever auth you already use. If it works today, nothing to do. Otherwise run -`wrangler login`, or set `CLOUDFLARE_API_TOKEN`. - -**b) R2 S3 API keys (for the archive uploads).** Archive objects are uploaded over -R2's S3-compatible API, which needs an Access Key ID + Secret. In the dashboard: -**R2 → Manage R2 API Tokens → Create API token**, permission **Object Read & Write**, -scoped to your bucket. Then set three environment variables where the editor runs: - -```sh -export R2_ACCESS_KEY_ID=<access key id> -export R2_SECRET_ACCESS_KEY=<secret access key> -export CLOUDFLARE_ACCOUNT_ID=<your account id> # used to build the S3 endpoint -``` - -The account id is on the R2 overview page; the S3 endpoint is derived as -`https://<CLOUDFLARE_ACCOUNT_ID>.r2.cloudflarestorage.com`. Keep these in the -environment (a shell profile, a systemd unit, a `.env` the editor loads) — **not** in -the settings JSON, which isn't a place for secrets. If a deploy has oversize archives -to upload but these are unset, it fails *before* the Pages deploy (so the site never -links to a missing file) with a message pointing back here. - -### 4. Point the editor at the bucket - -In the editor, open **Settings** and set: - -| Field | Value | -|---|---| -| **Archive overflow storage (R2 bucket)** | the bucket name, e.g. `my-archives-bucket` | -| **Archive overflow public URL** | the public base from step 2, e.g. `https://archives.example.com` (no trailing slash needed) | - -Leave both blank to keep the old behavior (oversize archives dropped, shown as "Too -large to host"). - -On the next **Build & deploy**, oversize archives upload to R2 and the Downloads page -(and the header link, which reappears once anything is hostable) point at them. - ---- - -## Securing downloads against cost-abuse - -The threat: someone scripts repeated downloads of large archives to run up your bill. - -**The reassuring part — R2 egress is free.** Unlike S3, Cloudflare R2 charges **$0** -for bandwidth/egress. An attacker looping downloads of a 480 MB zip cannot run up a -bandwidth bill. The *only* metered cost from reads is **Class B operations** (10M free -per month, then $0.36/M) — and the defenses below make even that hard to reach. - -You get these controls one of two ways — **a Worker on `workers.dev`** (no domain) or -**a custom domain**. Pick one; both are covered below. The Worker path is recommended -if you don't want to own a domain. - -## Option A — Serving via a Worker on workers.dev (no domain) - -The bundled **`r2-proxy/`** Worker is a small, self-contained project that binds the R2 -bucket and serves archive objects on a free `<name>.workers.dev` subdomain — the same -domain-free model as Pages' `*.pages.dev`. It's a **pure passthrough**: the request path -`<siteId>/archives/<file>.zip` maps straight to the bucket key, so **one Worker serves -every site** (deploy it once, not per site — that's the whole point of keying objects by -site id). What it gives you, all on the free tier: - -- **Edge caching** — full downloads are cached with the Cache API, so repeat pulls skip - R2 (no billable Class B op). It honors the `Cache-Control` we set on each object. -- **Rate limiting** — Cloudflare's native, free rate-limit binding caps requests per - client IP + file (dashboard rate-limit rules need a paid zone; this doesn't). -- **Path allow-listing** — it only serves `*/archives/*.zip`, never arbitrary keys. -- **Range / resumable downloads** — honors `Range` requests so big zips can resume. - -**Deploy it (once for the whole instance):** - -1. Edit `r2-proxy/wrangler.toml` and set `bucket_name` to the **same bucket** you use in - the editor's Settings. Optionally rename the Worker (`name`) and tune the rate limit - (`limit` / `period`). -2. From the repo root: - ```sh - cd r2-proxy - pnpm install # first time only - pnpm run deploy # = wrangler deploy, reusing your host wrangler auth - ``` - wrangler prints the deployed URL, e.g. `https://ytdlp-archive-proxy.<you>.workers.dev`. - (An "unsafe fields are experimental" warning for the rate-limit binding is expected.) -3. In the editor's **Settings**, set **Archive overflow public URL** to that - `workers.dev` URL. Re-deploy a site and its Downloads links resolve through the Worker. - -The Worker code lives in `r2-proxy/src/index.ts` — the caching and rate-limit logic are -small and commented if you want to adjust them. - -> **Free-tier limit:** Workers Free allows **100,000 requests/day** (resets daily). Far -> more than a downloads endpoint needs; if you ever exceed it, requests get a `429` -> (fail closed — no surprise bill) until the next day, or upgrade to Workers Paid ($5/mo). - -The archive `Cache-Control` is also set at upload time (`Cache-Control: public, -max-age=3600`, constant `ARCHIVE_CACHE_CONTROL` in -`common/publish/build.ts`); the Worker reads it back when caching. Archive -filenames are stable and overwritten in place on re-deploy, so this 1-hour bound is what -keeps a re-uploaded archive from being served stale for long — raise it if your archives -rarely change. - -## Option B — Securing via a custom domain - -If you'd rather use a custom domain (its own upsides: dashboard WAF, managed bot rules, -and rate-limiting rules without touching code), connect it per -[step 2 above](#2-expose-it-publicly) and add these, in order of impact: - -### 1. Edge caching - -Served through a custom domain, Cloudflare's CDN caches each archive at the edge (it -honors the `Cache-Control: public, max-age=3600` we set at upload), so repeated -downloads of the same file are served from cache and **never hit R2**. To make caching -aggressive, add a **Cache Rule** (dashboard: **Caching → Cache Rules → Create**): - -- **When:** `URI Path` contains `/archives/` -- **Then:** *Eligible for cache*, **Edge TTL → Override → 1 day** (or longer). - -If you raise the TTL a lot, **purge the cache on deploy** (dashboard **Caching → Purge**, -or `wrangler`/API) so a re-uploaded archive isn't served stale. - -### 2. Rate limiting — the hard backstop - -A Rate Limiting rule caps how fast any single client can pull archives, stopping a -flood that misses cache. Dashboard: **Security → WAF → Rate limiting rules → Create** -(the free plan includes one rule): - -- **When incoming requests match:** `URI Path` contains `/archives/` -- **Rate:** e.g. **20 requests per 1 minute** per client IP -- **Then:** *Block* for 10 minutes (or *Managed Challenge*). - -Tune the threshold to real usage — legitimate users download a handful of files, not -dozens per minute. - -### 3. Bot Fight Mode + WAF managed rules - -Dashboard: **Security → Bots → Bot Fight Mode** (free). Blocks the low-effort scripted -abuse that makes up most of this traffic. The free **WAF managed ruleset** adds a -baseline of protection at no cost. - -### 4. Hotlink protection (optional) - -Stops other sites embedding your archives and spending your ops budget serving their -audience. A WAF custom rule (**Security → WAF → Custom rules**): - -- **When:** `URI Path` contains `/archives/` **and** `Referer` does not contain your - domain **and** `Referer` is not empty -- **Then:** *Block*. - -(Allow an empty `Referer` so direct clicks and privacy-conscious browsers still work.) - -## Billing / usage alerts (either option) - -R2 has no hard spend cap, but Cloudflare **Notifications** (dashboard: **Notifications -→ Add**) can email you when R2 storage or Class A/B operations cross a threshold — -cheap insurance so nothing surprises you. - -## What to skip - -**Signed URLs / token-gated downloads** are the heavyweight option — they add key -management and friction for legitimate users. Given egress is free and caching -neutralizes the ops cost, they're overkill for *cost* defense (the `r2-proxy` Worker is -a plain passthrough, not an access gate). Only reach for signed URLs if you want -*access control* (private archives), not cost control. - ---- - -## Cost expectations - -Measured across all sites in this project, total compressed archives are **≈ 2–4.5 GB**. -Against R2's free tier: - -| Resource | Free tier / month | This project's usage | -|---|---|---| -| **Egress / bandwidth** | unlimited, **$0** | irrelevant — no egress charge exists | -| **Storage** | 10 GB-month | ~2–4.5 GB — comfortable | -| **Class A ops** (writes) | 1,000,000 | ~one PUT per archive per deploy — negligible | -| **Class B ops** (reads) | 10,000,000 | mostly absorbed by CDN cache | - -The realistic bill for hosting these archives is **$0**. The configuration above exists -to keep it that way under adversarial traffic, not because normal usage is close to any -limit. diff --git a/DEPLOY_DOCKER.md b/DEPLOY_DOCKER.md @@ -1,82 +0,0 @@ -# Docker export build pipeline - -> **Not the document you want if you are trying to *run* the apps in containers.** -> That is [RUNNING_IN_DOCKER.md](./RUNNING_IN_DOCKER.md) — `docker compose up` and a -> working archive server, built from the root `Dockerfile`. This page is about -> *building sites*: fanning per-site export builds out across containers, using -> `Dockerfile.build`. The two share nothing but the word "docker". - -Archilyzer's Docker build mode builds **every site in parallel** in isolated containers, then -deploys them serially — a large speedup when you host several sites, and stronger -isolation than the basic single-process build. This is opt-in: set **Build -pipeline → Docker** in Settings (or the toggle on the Deploy page). Basic mode is -unchanged and remains the default. - -## Prerequisites - -- A container engine: **Docker**, or **podman** (set `DOCKER_BIN=podman`). Rootless - podman is a good fit — it maps container files to your host user automatically. -- The editor host still needs Node + pnpm (Phase A and deploy run on the host) and - your Cloudflare/R2 credentials in the environment (see - [DEPLOY_CLOUDFLARE.md](./DEPLOY_CLOUDFLARE.md)). Credentials are **never** passed - into a container — deploy runs on the host. - -The build image is built (and cached) automatically from `Dockerfile.build` the -first time you run; edit **Build image** / **Dockerfile** in Settings to override -the tag/path. - -## How it works - -Trigger it with **Build all sites** on the Deploy page. One managed job runs three -ordered phases: - -1. **Phase A — shared, on the host, once.** `build:data` (search index + - `.export-index` staging) then `build:archives` (warm the shared archive-zip - cache for the union of all sites' channels). Only the host writes this shared - state, so containers never race it. This phase is serial and is the long pole on - a cold build; on a warm rebuild it's near-instant (unchanged channels are - skipped). -2. **Phase B — per-site, in parallel containers.** Each site's `compose:site + - next build` runs in its own container, capped by **Max parallel builds**. Each - writes an isolated `out/` under `export/.export-builds/<siteId>/`. Containers - mount the corpus/index/staging/archive cache **read-only**. (Network is left on: - `next build` fetches the site's fonts from Google via `next/font/google`; - isolation comes from the read-only mounts, per-site output dir, and non-root - user.) -3. **Phase C — deploy, on the host, serially.** After every build finishes, each - built site is deployed in turn (oversize-archive R2 upload, then `wrangler pages - deploy`). A single site failing to build or deploy is reported and skipped; the - rest still ship. - -If no container engine is available, the action logs a notice and falls back to a -serial host build+deploy (one site at a time). - -## Mounts (per Phase-B container) - -| Host | Container | Mode | -|---|---|---| -| `transcripts/` (corpus + `index.mdb` + archive cache) | `/data/transcripts` | ro | -| `export/.export-index` (shared + per-site staging) | `/data/export/.export-index` | ro | -| `export/.export-builds/<siteId>` (public/out/.next/caches) | `/site` | rw | -| `settings.json` (build config, mounted fresh — not baked) | `/data/settings.json` | ro | - -The per-site `/site` mount is persistent, so incremental `next build` (`.next`) and -incremental compose (`.compose-cache`) stay warm across builds. - -## Tuning & environment - -- **Max parallel builds** (setting) — how many site containers run at once. Each - `next build` can use up to ~8 GB; a safe starting point is `floor(RAM_GB / 9)`. -- `DOCKER_BIN` — container binary (default `docker`; e.g. `podman`). -- `DOCKER_BUILD_MEMORY`, `DOCKER_BUILD_CPUS` — optional per-container `--memory` / - `--cpus` caps so a fan-out can't OOM/peg the host. -- `BUILD_ARCHIVES=0` (or the **Skip archive zips** checkbox) — skip the archive - warm + per-site archive materialize for a faster build with no download bundles. - -## Notes - -- Containers run as your host uid/gid (`-u`), so files under `.export-builds/` are - host-owned, not root-owned. -- The image bakes the repo source + deps; a code change rebuilds it, but Docker - layer caching keeps that cheap (deps re-install only when the lockfile moves). -- `.export-builds/` is gitignored and excluded from the image build context. diff --git a/PUBLISH.md b/PUBLISH.md @@ -0,0 +1,443 @@ +# Publishing + +What gets published, how to build and deploy it, and how to keep a public archive +cheap and safe. For *running* the apps in containers see +[RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md); every environment variable named here +is in [ENVIRONMENT.md](ENVIRONMENT.md). + +- [What gets published](#what-gets-published) +- [The three ways to drive it](#the-three-ways-to-drive-it) +- [Cloudflare Pages](#cloudflare-pages) +- [Preview deployments](#preview-deployments) +- [Download archives and R2](#download-archives-and-r2) +- [Securing downloads against cost-abuse](#securing-downloads-against-cost-abuse) +- [Building every site in containers](#building-every-site-in-containers) +- [Reading a published archive from Claude Code](#reading-a-published-archive-from-claude-code) + +--- + +## What gets published + +Three static artefacts, each pre-rendered to plain HTML and JSON — no database, no +server-side code, servable by anything: + +| Artefact | What it is | Built into | Config | +|---|---|---|---| +| A **site** | The export app over the channels one site selects: search, transcripts, charts, downloads. One corpus can publish several. | `export/out` (docker mode: `export/.export-builds/<siteId>/out`) | `transcripts/sites/<id>/site.json` ([SITE.md](SITE.md)) | +| The **hub** | The export app in hub mode: federated search across the family of sites. | `export/out` | `transcripts/sites/_homepage/homepage.json` | +| The **homepage** | The project's own site (`homepage/`). | `homepage/out` | — | + +The hub and the homepage are two different Pages projects: the hub deploys to the +project `homepage.json` names (e.g. `archilyzer-hub`) and is refused `archilyzer`, +which is the homepage's. + +A site build has three steps, run in `export/`: the **data phase** (the LMDB index, +the stats datasets and the chart templates — `archilyzer index`, `build stats`, +`build templates`), **compose** (the site's slice of the shared index into +`export/public`, plus its download archives — `archilyzer compose site <id>`) and +`next build`. `--nodata` skips the data phase and reuses the last one's staging. + +## The three ways to drive it + +The editor's **/sites** page, `pnpm ops` (HTTP to a running editor, with its +`WORKER_TOKEN`) and the `archilyzer` CLI (local, no editor needed) call the same +entry points in `common/publish/build.ts`. + +| To | Editor | `pnpm ops` | `pnpm archilyzer …` | +|---|---|---|---| +| Build one site | a site's **Publish** tab → *Build static export* | `build-site` | `build site <id> [--nodata] [--skip-archives]` | +| Deploy the built site | *Deploy to production* / *Deploy preview* | `deploy-site` | `deploy site <id> [--preview <branch>]` | +| Build, then deploy | *Build & deploy* | `build-deploy` | `build site <id>` then `deploy site <id>` | +| Build every site | /sites → **Build all sites** | — | `build all [--skip-archives]` | +| The hub | /sites → Hub → **Build hub** / **Deploy hub** | `build-hub`, `deploy-hub` | `build hub`, `deploy hub [--preview <branch>]` | +| The homepage | — | — | `build homepage`, `deploy homepage` | + +`pnpm archilyzer <command>` is the short form of +`pnpm --filter yt-dlp-transcript-common exec tsx bin/archilyzer.ts <command>`; +`pnpm archilyzer --help` lists every command, and `pnpm archilyzer doctor` checks that +this machine has what a build needs. From `export/`, `pnpm run build` is +`archilyzer build site` (with `SITE_ID`) and `pnpm run deploy` is `archilyzer deploy +site`. + +Deploy-only ships whatever is in `export/out`, which the basic build composes one +site at a time into a single shared directory — so it **refuses, before starting a +job, if `export/out` holds a build of another site** (or no build at all), naming the +site to build first. A deploy queued behind another site's build re-checks when it +starts. `build-deploy` cannot hit this: it builds. + +--- + +## Cloudflare Pages + +Sites deploy to **Cloudflare Pages** with `wrangler pages deploy`; download archives +too large for Pages' per-file limit overflow to **Cloudflare R2**. Everything here +fits inside Cloudflare's **free tier**. + +- **A Pages project must exist before its first deploy.** wrangler offers to create a + missing project only on an interactive terminal, and the deploy's stdin is a pipe, + so a missing project fails at once with wrangler's own "does not exist" sentence. + Create it first: `pnpm dlx wrangler pages project create <name> --production-branch + main`. A site's project is `site.json`'s `cloudflareProject`. +- **wrangler's own auth.** The deploy reuses whatever auth wrangler already has — + `wrangler login`, or `CLOUDFLARE_API_TOKEN` in the environment. +- **A production deploy inherits the checkout's git branch.** `wrangler pages deploy` + with no `--branch` infers one from the repository it runs in, so a *production* + deploy from a feature-branch checkout silently produces a preview instead. Only the + preview path passes `--branch`; `deploy homepage` passes `--branch main`. If a + "production" deploy did not go live, check what branch the checkout is on. + +## Preview deployments + +A **preview** is the same built bundle deployed to a branch that is not the Pages +project's production branch. Cloudflare publishes it at a **branch alias** — + +``` +https://<branch>.<project>.pages.dev +``` + +— and leaves the live site alone. Each deploy also gets an immutable +per-deployment URL (`https://<hash>.<project>.pages.dev`), which wrangler prints as +"Deployment complete! Take a peek over at …"; the editor repeats both on a +`[preview]` line at the end of the job log, because the streamed log scrolls. + +The alias is a function of the project and the branch and nothing else, so it is +known *before* the deploy runs — which is why the editor can link it while you are +still typing the name. + +**Branch names** must be 1–28 lowercase letters, digits and dashes, starting and +ending with a letter or digit. That is exactly what Cloudflare's alias sanitizer +preserves verbatim, so the alias shown is the alias that resolves. `main`, `master` +and `production` are refused: a deploy to the production branch is not a preview, it +is the live site. + +| Surface | How | +|---|---| +| Editor | A site's **Publish** tab → *Individual steps* → **Deploy a preview**: type a branch, press **Deploy preview**. The production button beside it says **Deploy to production**. | +| Ops API | `POST /api/ops/deploy-site` `{ "siteId": "...", "preview": "<branch>" }` — deploy-only, of the already-built `export/out`. `POST /api/ops/build-deploy` takes `preview` too (build *then* preview-deploy). Both answer with `previewUrl`. `pnpm ops deploy-site --json '{"siteId":"anilyzer","preview":"tags-exclude"}' --wait` prints the alias on its own line after the log. | +| CLI | `pnpm archilyzer deploy site anilyzer --preview tags-exclude` | + +The high-value loop is **build once, preview, then promote**: build the site, deploy +it with a preview branch, look at it, then deploy again with no preview — the same +`export/out`, unrebuilt. + +**A preview shares the production R2 archive bucket.** R2 has no per-branch +namespace, and the keys are `<siteId>/archives/<file>.zip` either way. In practice +this is cheap and harmless — the upload skips any object R2 already holds at the same +size, and an unchanged channel re-zips byte-stable — but a *changed* archive replaces +the one production's manifest links to. The editor says so once at the top of every +preview deploy. + +**Who can open a preview is a Cloudflare setting, not ours.** Pages projects have a +*preview deployment access* setting (Settings → General): **public** by default, or +restricted to Cloudflare Access. A default-configured project's preview URL is +world-readable by anyone who has the link. + +--- + +## Download archives and R2 + +At compose time, `archilyzer compose site` generates one transcript zip and one +live-chat zip **per channel** into `export/public/archives/`, and records them in +`public/archives/manifest.json`. The `/downloads` page and the header **Downloads** +link read that manifest. A site can turn its archives off (`site.json` `archives: false`), and a build +can skip them (`--skip-archives`, `BUILD_ARCHIVES=0`, or the **Skip archive zips** +checkbox). + +Cloudflare Pages rejects any single asset larger than **25 MB**, and real channels +blow past that easily (a channel's live-chat zip can be hundreds of MB). So the +pipeline splits archives by size: + +| Archive size | Where it's served from | Manifest entry | +|---|---|---| +| ≤ 25 MB | Cloudflare **Pages** (shipped in `out/`, free) | `filename`, no `url` | +| > 25 MB, R2 configured | Cloudflare **R2**, uploaded on deploy | `url` → R2 | +| > 25 MB, R2 **not** configured | not served | `oversize: true`, shown as "Too large to host" | + +(The cap is `MAX_ARCHIVE_BYTES`, or a site's own `archiveMaxBytes`; `0` = no cap.) + +Oversize archives are staged during compose into `export/.r2-staging/<siteId>/archives/` +(gitignored, kept out of `public/`), then uploaded to +`<bucket>/<siteId>/archives/<file>.zip` **before** the Pages deploy runs, so the +manifest URLs resolve immediately. **Every deploy path uploads them** — the editor's +**Deploy** / **Build & deploy**, `pnpm ops deploy-site`, and `archilyzer deploy site` +(which `pnpm run deploy` in `export/` runs). A build alone only *stages* them. The +bucket is read from `settings.json`; the R2 credentials come from the environment +(step 3 below). + +Uploads go through R2's **S3 API** (the AWS SDK's multipart uploader), not +`wrangler r2 object put` — wrangler caps a single upload at **300 MiB**, and real +live-chat archives are larger (multipart has no such limit). That is why archive +uploads need S3 credentials in addition to the wrangler auth the Pages deploy uses. + +### One-time R2 setup + +**1. Create the bucket.** One bucket serves **all** your sites — objects are +namespaced by `siteId`, so you do **not** need one bucket per site. In the dashboard +(**R2 → Create bucket**) or with the same wrangler the deploy already uses: + +```sh +wrangler r2 bucket create my-archives-bucket +``` + +**2. Expose it publicly.** R2 buckets are private by default, and the Downloads page +links to plain `https://…/<file>.zip` URLs, so the bucket needs public read access. +There are three ways; the first needs **no domain purchase** and is the recommended +one: + +- **A Worker on `*.workers.dev` (recommended — no domain, no WHOIS):** deploy the + bundled `r2-proxy/` Worker to a free `<name>.workers.dev` subdomain (just like Pages' + `*.pages.dev`). It streams objects from the bucket and gives you edge caching **and** + tunable rate limiting in code — the same cost defenses a custom domain would, without + owning a domain. One Worker serves **every** site. See + [Option A](#option-a--a-worker-on-workersdev-no-domain) below. Set the editor's + public URL to `https://<name>.workers.dev`. +- **Custom domain:** bucket **Settings → Public access → Custom Domains → Connect + Domain**, e.g. `archives.example.com` (the domain must be on your Cloudflare account). + Routes downloads through Cloudflare's CDN, WAF, caching, and dashboard rate-limiting + rules. See [Option B](#option-b--a-custom-domain). +- **`r2.dev` subdomain (quick test only):** bucket **Settings → Public access → Allow + Access**. You get `https://pub-abc123.r2.dev` — free and domain-free, but Cloudflare + throttles `r2.dev` and gives you **no** cache / rate-limit control of your own. Fine + for a smoke test; use the Worker for a real instance. + +**3. The R2 S3 credentials.** Archive objects are uploaded over R2's S3-compatible +API, which needs an Access Key ID + Secret. In the dashboard: **R2 → Manage R2 API +Tokens → Create API token**, permission **Object Read & Write**, scoped to your +bucket. Then set three environment variables where the editor (or the CLI) runs: + +```sh +export R2_ACCESS_KEY_ID=<access key id> +export R2_SECRET_ACCESS_KEY=<secret access key> +export CLOUDFLARE_ACCOUNT_ID=<your account id> # used to build the S3 endpoint +``` + +The account id is on the R2 overview page; the S3 endpoint is derived as +`https://<CLOUDFLARE_ACCOUNT_ID>.r2.cloudflarestorage.com`. Keep these in the +environment (a shell profile, a systemd unit, a `.env` the editor loads) — **not** in +the settings JSON, which is not a place for secrets. If a deploy has oversize +archives to upload but these are unset, it fails *before* the Pages deploy (so the +site never links to a missing file) with a message pointing back here. + +**4. Point the editor at the bucket.** In **Settings**: + +| Field | Value | +|---|---| +| **Archive overflow storage (R2 bucket)** | the bucket name, e.g. `my-archives-bucket` | +| **Archive overflow public URL** | the public base from step 2, e.g. `https://archives.example.com` (no trailing slash needed) | + +Leave both blank to keep oversize archives dropped and shown as "Too large to host". +On the next build and deploy, oversize archives upload to R2 and the Downloads page +(and the header link, which reappears once anything is hostable) point at them. + +--- + +## Securing downloads against cost-abuse + +The threat: someone scripts repeated downloads of large archives to run up your bill. + +**The reassuring part — R2 egress is free.** Unlike S3, Cloudflare R2 charges **$0** +for bandwidth/egress. An attacker looping downloads of a 480 MB zip cannot run up a +bandwidth bill. The *only* metered cost from reads is **Class B operations** (10M free +per month, then $0.36/M) — and the defenses below make even that hard to reach. + +You get these controls one of two ways — **a Worker on `workers.dev`** (no domain) or +**a custom domain**. Pick one. The Worker is recommended if you do not want to own a +domain. + +### Option A — a Worker on workers.dev (no domain) + +The bundled **`r2-proxy/`** Worker is a small, self-contained project that binds the R2 +bucket and serves archive objects on a free `<name>.workers.dev` subdomain. It is a +**pure passthrough**: the request path `<siteId>/archives/<file>.zip` maps straight to +the bucket key, so **one Worker serves every site** (deploy it once, not per site). +What it gives you, all on the free tier: + +- **Edge caching** — full downloads are cached with the Cache API, so repeat pulls skip + R2 (no billable Class B op). It honors the `Cache-Control` set on each object. +- **Rate limiting** — Cloudflare's native, free rate-limit binding caps requests per + client IP + file (dashboard rate-limit rules need a paid zone; this does not). +- **Path allow-listing** — it only serves `*/archives/*.zip`, never arbitrary keys. +- **Range / resumable downloads** — honors `Range` requests so big zips can resume. + +**Deploy it (once for the whole instance):** + +1. Edit `r2-proxy/wrangler.toml` and set `bucket_name` to the **same bucket** you use in + the editor's Settings. Optionally rename the Worker (`name`) and tune the rate limit + (`limit` / `period`). +2. From the repo root: + ```sh + cd r2-proxy + pnpm install # first time only + pnpm run deploy # = wrangler deploy, reusing your host wrangler auth + ``` + wrangler prints the deployed URL, e.g. `https://ytdlp-archive-proxy.<you>.workers.dev`. + (An "unsafe fields are experimental" warning for the rate-limit binding is expected.) +3. In the editor's **Settings**, set **Archive overflow public URL** to that + `workers.dev` URL. Re-deploy a site and its Downloads links resolve through the Worker. + +The Worker code is `r2-proxy/src/index.ts`; the caching and rate-limit logic are small +and commented. + +> **Free-tier limit:** Workers Free allows **100,000 requests/day** (resets daily). Far +> more than a downloads endpoint needs; if you ever exceed it, requests get a `429` +> (fail closed — no surprise bill) until the next day, or upgrade to Workers Paid ($5/mo). + +The archive `Cache-Control` is set at upload time (`Cache-Control: public, +max-age=3600`, constant `ARCHIVE_CACHE_CONTROL` in `common/publish/build.ts`); the +Worker reads it back when caching. Archive filenames are stable and overwritten in +place on re-deploy, so this 1-hour bound is what keeps a re-uploaded archive from being +served stale for long — raise it if your archives rarely change. + +### Option B — a custom domain + +A custom domain has its own upsides — dashboard WAF, managed bot rules, and +rate-limiting rules without touching code. Connect it per step 2 above and add these, +in order of impact: + +1. **Edge caching.** Through a custom domain, Cloudflare's CDN caches each archive at + the edge (it honors the `Cache-Control: public, max-age=3600` set at upload), so + repeated downloads are served from cache and **never hit R2**. To make caching + aggressive, add a **Cache Rule** (**Caching → Cache Rules → Create**): *when* `URI + Path` contains `/archives/`, *then* Eligible for cache, **Edge TTL → Override → 1 + day** (or longer). If you raise the TTL a lot, **purge the cache on deploy** so a + re-uploaded archive is not served stale. +2. **Rate limiting — the hard backstop.** A Rate Limiting rule caps how fast any + single client can pull archives, stopping a flood that misses cache (**Security → + WAF → Rate limiting rules → Create**; the free plan includes one rule): *when* + `URI Path` contains `/archives/`, *rate* e.g. **20 requests per 1 minute** per + client IP, *then* Block for 10 minutes (or Managed Challenge). Tune the threshold + to real usage — legitimate users download a handful of files, not dozens per + minute. +3. **Bot Fight Mode + WAF managed rules.** **Security → Bots → Bot Fight Mode** + (free) blocks the low-effort scripted abuse that makes up most of this traffic; the + free **WAF managed ruleset** adds a baseline at no cost. +4. **Hotlink protection (optional).** Stops other sites embedding your archives and + spending your ops budget serving their audience. A WAF custom rule: *when* `URI + Path` contains `/archives/` **and** `Referer` does not contain your domain **and** + `Referer` is not empty, *then* Block. (Allow an empty `Referer` so direct clicks and + privacy-conscious browsers still work.) + +### Billing alerts, and what to skip + +R2 has no hard spend cap, but Cloudflare **Notifications** (**Notifications → Add**) +can email you when R2 storage or Class A/B operations cross a threshold — cheap +insurance so nothing surprises you. + +**Signed URLs / token-gated downloads** are the heavyweight option — they add key +management and friction for legitimate users. Given egress is free and caching +neutralizes the ops cost, they are overkill for *cost* defense (the `r2-proxy` Worker +is a plain passthrough, not an access gate). Reach for them only if you want *access +control* (private archives), not cost control. + +### Cost expectations + +Measured across all sites in this project, total compressed archives are **≈ 2–4.5 GB**. +Against R2's free tier: + +| Resource | Free tier / month | This project's usage | +|---|---|---| +| **Egress / bandwidth** | unlimited, **$0** | irrelevant — no egress charge exists | +| **Storage** | 10 GB-month | ~2–4.5 GB — comfortable | +| **Class A ops** (writes) | 1,000,000 | ~one PUT per archive per deploy — negligible | +| **Class B ops** (reads) | 10,000,000 | mostly absorbed by CDN cache | + +The realistic bill for hosting these archives is **$0**. The configuration above exists +to keep it that way under adversarial traffic, not because normal usage is close to +any limit. + +--- + +## Building every site in containers + +> **Not about running the apps in containers** — that is +> [RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md), built from the root `Dockerfile`. This +> is about *building sites*: fanning per-site export builds out across containers, +> using `Dockerfile.build`. The two share nothing but the word "docker". + +The docker build mode builds **every site in parallel** in isolated containers, then +deploys them serially — a large speedup when you host several sites, and stronger +isolation than the basic single-process build. It is opt-in: **Settings → Build +pipeline → Docker**. Basic mode, the default, builds one site at a time in `export/`. + +**Prerequisites.** + +- A container engine: **Docker**, or **podman** (set `DOCKER_BIN=podman`). Rootless + podman is a good fit — it maps container files to your host user automatically. +- The editor host still needs Node + pnpm (Phase A and the deploys run on the host) + and the Cloudflare/R2 credentials in the environment. Credentials are **never** + passed into a container — deploy runs on the host. +- Inside the runtime container (`docker compose up`) there is no `docker` binary, and + the pipeline falls back to the serial host build. + +The build image is built (and cached) automatically from `Dockerfile.build` the first +time you run; **Build image** / **Dockerfile** in Settings override the tag and path. + +**How it works.** Trigger it with **Build all sites** on /sites (or `pnpm archilyzer +build all`, which uses docker when `docker version` answers). One job runs three +ordered phases: + +1. **Phase A — shared, on the host, once.** The data phase (the search index and the + `.export-index` staging), then `archilyzer build archives` (the shared archive-zip + cache for the union of all sites' channels). Only the host writes this shared + state, so containers never race it. This phase is serial and is the long pole on a + cold build; on a warm rebuild it is near-instant (unchanged channels are skipped). +2. **Phase B — per-site, in parallel containers.** Each site's compose + `next build` + runs in its own container (`docker/build-site.sh`: `archilyzer build site <id> + --nodata`), capped by **Max parallel builds**. Each writes an isolated `out/` under + `export/.export-builds/<siteId>/`. Containers mount the corpus, index, staging and + archive cache **read-only** (`ARCHIVES_READONLY=1`). Network is left on: `next + build` fetches the site's fonts through `next/font/google`; isolation comes from the + read-only mounts, the per-site output dir and the non-root user. +3. **Phase C — deploy, on the host, serially.** After every build finishes, each built + site is deployed in turn (the R2 upload, then `wrangler pages deploy`). A single + site failing to build or deploy is reported and skipped; the rest still ship. + +If no container engine is available, the action logs a notice and falls back to a +serial host build+deploy (one site at a time). + +**Mounts, per Phase-B container.** + +| Host | Container | Mode | +|---|---|---| +| `transcripts/` (corpus + `index.mdb` + archive cache) | `/data/transcripts` | ro | +| `export/.export-index` (shared + per-site staging) | `/data/export/.export-index` | ro | +| `export/.export-builds/<siteId>` (public/out/.next/caches) | `/site` | rw | +| `settings.json` (build config, mounted fresh — not baked) | `/data/settings.json` | ro | + +The per-site `/site` mount is persistent, so incremental `next build` (`.next`) and +incremental compose (`.compose-cache`) stay warm across builds. + +**Tuning.** + +- **Max parallel builds** (setting) — how many site containers run at once. Each `next + build` can use up to ~8 GB; a safe starting point is `floor(RAM_GB / 9)`. +- `DOCKER_BIN` — the container binary (default `docker`; e.g. `podman`). +- `DOCKER_BUILD_MEMORY`, `DOCKER_BUILD_CPUS` — optional per-container `--memory` / + `--cpus` caps so a fan-out cannot OOM or peg the host. +- `BUILD_ARCHIVES=0` (or **Skip archive zips**) — skip the archive warm and the + per-site archive materialize for a faster build with no download bundles. + +**Notes.** Containers run as your host uid/gid (`-u`), so files under +`.export-builds/` are host-owned, not root-owned. The image bakes the repo source and +deps; a code change rebuilds it, but layer caching keeps that cheap (deps re-install +only when the lockfile moves). `.export-builds/` is gitignored and excluded from the +image build context. The editor mounts the host's `docker/build-site.sh` over the +baked one, so an image older than the checkout still runs today's script. + +--- + +## Reading a published archive from Claude Code + +Every published archive serves a machine contract (`/corpus.json`, `/llms.txt`), and +the MCP server reads it over HTTP. Register it against your public URL — as +`archilyzer`, which the tracked `/ask` and `/sweep` commands expect: + +```sh +claude mcp add archilyzer \ + --env TRANSCRIPT_SITE_URL=https://<your-site>.pages.dev \ + -- pnpm -C "$PWD" archilyzer mcp +``` + +`archilyzer mcp` starts the same server as `pnpm --filter yt-dlp-transcript-mcp exec +tsx src/index.ts` (the form [mcp/README.md](mcp/README.md) uses), with the +environment passed through and nothing on stdout but the protocol. diff --git a/README.md b/README.md @@ -329,6 +329,7 @@ claude # then try: /ask what has he said about magic tournaments? The two editor lines are optional: they let `fetch_clip` ask a local editor for clip media (`WORKER_TOKEN` is the editor's own). Leave them out for research alone. +`-- pnpm -C "$PWD" archilyzer mcp` starts the same server through the repo's CLI. > **Register the server as `archilyzer`.** The shipped commands call > `mcp__archilyzer__ask_plan` / `mcp__archilyzer__sweep_plan`, and that tool name @@ -449,34 +450,25 @@ Two more pieces round out the workspace: the **project site** (`homepage/`) and Transcripts and per-channel state live at `<repo>/transcripts/` — **its own git repo**, untouched by the workspace. -Everything resolves through `getPaths()` (`common/lib/paths.ts`) and can be overridden -by environment variables: - -| Variable | Default | Purpose | -| --- | --- | --- | -| `TRANSCRIPTS_DIR` | `<repo>/transcripts` | The corpus: channels, media, index, job logs. | -| `SAVED_VIDEOS_DIR` | inside `TRANSCRIPTS_DIR` | Persisted source-video store; can live on another disk. | -| `SITES_DIR` | inside `TRANSCRIPTS_DIR` | Per-site configuration. | -| `EXPORT_PUBLIC_DIR` | `<repo>/export/public` | Where the index writes paginated JSON. | -| `SETTINGS_FILE` | `<repo>/settings.json` | Operational settings — every key, default and meaning is in [SETTINGS.md](SETTINGS.md). | -| `YTDLP_BIN` | `yt-dlp` on PATH | The downloader. | -| `WHISPER_BIN` / `WHISPER_MODEL` | `whisper-cli` on PATH | Default transcription backend and its model. | -| `FFMPEG_BIN` / `FFPROBE_BIN` | on PATH | Transcode and duration checks. | +Every path and binary resolves through `getPaths()` (`common/lib/paths.ts`), and each +can be overridden by an environment variable — `TRANSCRIPTS_DIR` moves the whole corpus, +`YTDLP_BIN` / `FFMPEG_BIN` / `WHISPER_BIN` name the tools. The full list, with every +other variable the code reads, is **[ENVIRONMENT.md](ENVIRONMENT.md)** (generated from +a list the tests hold to the code). `pnpm archilyzer doctor` prints which overrides are set +and whether every tool this machine is configured to use is there. Everything else lives in the editor's **Settings** page and is optional — a missing or -partial settings file falls back to defaults. Full list in -[SETUP.md](SETUP.md#configuration--environment-variables). +partial settings file falls back to defaults. Every key is in [SETTINGS.md](SETTINGS.md). ## Publishing The static export can be served by anything. The path with the most support is Cloudflare Pages, with download archives too large for Pages' 25 MB per-file limit -overflowing to R2 — see **[DEPLOY_CLOUDFLARE.md](DEPLOY_CLOUDFLARE.md)**, which also -covers the configuration that keeps public archive downloads from being abused to run -up costs. - -If you host several sites from one corpus, the opt-in Docker build pipeline builds them -all in parallel in isolated containers — see **[DEPLOY_DOCKER.md](DEPLOY_DOCKER.md)**. +overflowing to R2. **[PUBLISH.md](PUBLISH.md)** covers building and deploying a site, +the hub and the homepage (from the editor, `pnpm ops` or `pnpm archilyzer`), previews, +the R2 setup, the configuration that keeps public archive downloads from being abused +to run up costs, and the opt-in pipeline that builds several sites in parallel in +isolated containers. Channels can sync automatically on a per-channel cadence via a cron heartbeat — see **[SCHEDULED_SYNC.md](SCHEDULED_SYNC.md)**. @@ -499,11 +491,11 @@ See [CONTRIBUTING.md](CONTRIBUTING.md) to work on the code. | Document | Covers | |---|---| -| [SETUP.md](SETUP.md) | Full per-OS install, every environment variable, transcription backends. | -| [CONTRIBUTING.md](CONTRIBUTING.md) | Workspace layout, tests, CLI shims, internals. | +| [SETUP.md](SETUP.md) | Full per-OS install, transcription backends. | +| [ENVIRONMENT.md](ENVIRONMENT.md) | Every environment variable, by audience (generated). | +| [CONTRIBUTING.md](CONTRIBUTING.md) | Workspace layout, tests, the `archilyzer` CLI, internals. | | [RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md) | `docker compose up` for the whole stack: exposure model, auth, GPU. | -| [DEPLOY_CLOUDFLARE.md](DEPLOY_CLOUDFLARE.md) | Pages + R2, and cost-abuse protection. | -| [DEPLOY_DOCKER.md](DEPLOY_DOCKER.md) | Parallel multi-site export builds in containers. | +| [PUBLISH.md](PUBLISH.md) | Building and deploying sites: Pages + R2, cost-abuse protection, parallel builds in containers. | | [SCHEDULED_SYNC.md](SCHEDULED_SYNC.md) | Unattended per-channel syncing. | | [WORKTREES.md](WORKTREES.md) | Parallel checkouts and the port scheme. | | [mcp/README.md](mcp/README.md) | The MCP server: tools, links, client setup. | diff --git a/RUNNING_IN_DOCKER.md b/RUNNING_IN_DOCKER.md @@ -1,13 +1,14 @@ # Running an archive in Docker One command stands up a working archive server: the editor, the tools it drives -(yt-dlp, ffmpeg, whisper.cpp), and a reverse proxy that is the only thing on the -box with an open port. +(yt-dlp, ffmpeg, and a transcription engine — whisper.cpp, or parakeet.cpp in the +Vulkan image), and a reverse proxy that is the only thing on the box with an open +port. -> This document is about **running the apps**. [DEPLOY_DOCKER.md](DEPLOY_DOCKER.md) -> is a different thing entirely — it is about fanning per-site *export builds* out -> across containers, and it uses `Dockerfile.build`. Neither file affects the -> other. +> This document is about **running the apps**. Fanning per-site *export builds* out +> across containers is a different thing entirely — it uses `Dockerfile.build`, and +> it is in [PUBLISH.md](PUBLISH.md#building-every-site-in-containers). Neither +> affects the other. --- @@ -538,7 +539,7 @@ silently loses formats), `ffmpeg`/`ffprobe`, `whisper-cli` (statically linked), ### The multi-site build pipeline falls back inside a container The editor can fan per-site export builds out across containers -(`buildPipeline.mode = "docker"`, see DEPLOY_DOCKER.md). Inside a container there +(`buildPipeline.mode = "docker"`, see [PUBLISH.md](PUBLISH.md#building-every-site-in-containers)). Inside a container there is no `docker` binary, so that path is unavailable. It already handles this — the build logs diff --git a/SETUP.md b/SETUP.md @@ -46,7 +46,7 @@ homepage): | **ffmpeg** + **ffprobe** | Audio transcode + duration checks for `transcribe` channels. | `ffmpeg` / `ffprobe` on `PATH` | | A **transcription backend** | `handling: "transcribe"` channels only. Default is **whisper.cpp** (`whisper-cli`); `chough` and `parakeet.cpp` are alternatives. | `whisper-cli` on `PATH` | | **rsync** | Backing up the saved-video store. | `rsync` on `PATH` | -| **Docker** | Running the whole stack in containers ([RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md)), the parallel multi-site build ([DEPLOY_DOCKER.md](DEPLOY_DOCKER.md)), and the sharded e2e run (`pnpm e2e:sharded`). | — | +| **Docker** | Running the whole stack in containers ([RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md)), the parallel multi-site build ([PUBLISH.md](PUBLISH.md#building-every-site-in-containers)), and the sharded e2e run (`pnpm e2e:sharded`). | — | Every binary above is overridable by an environment variable (e.g. `YTDLP_BIN`) — see [Configuration & environment variables](#configuration--environment-variables). @@ -260,28 +260,27 @@ starting template. Both are generated from the settings schema [CHANNEL.md](CHANNEL.md). Settings are optional — a missing/partial `settings.json` falls back to built-in defaults, so the app runs out of the box. -Paths and binaries resolve through `getPaths()` in `common/lib/paths.ts`. Override -any of them via environment variables before launching: +Paths and binaries resolve through `getPaths()` in `common/lib/paths.ts`, and each is +overridden by an environment variable before launching — `TRANSCRIPTS_DIR` (the corpus, +default `<repo>/transcripts`), `SETTINGS_FILE`, `YTDLP_BIN`, `FFMPEG_BIN` / +`FFPROBE_BIN`, `WHISPER_BIN` / `WHISPER_MODEL`, `PARAKEET_CLI` / `PARAKEET_MODEL` and +the rest. **Every variable the code reads is in [ENVIRONMENT.md](ENVIRONMENT.md)**, by +audience: the path overrides, the runtime tokens and knobs (`WORKER_TOKEN`, +`SYNC_TICK_URL`, the R2 credentials, …), the ports, the docker `ARCHILYZER_*` set and +the test-only ones. It is generated from `common/lib/envVars.ts`, and a test fails when +the code reads a variable that list does not declare. -| Variable | Default | Purpose | -| --- | --- | --- | -| `TRANSCRIPTS_DIR` | `<repo>/transcripts` | Channels, archives, LMDB index, job logs. | -| `SAVED_VIDEOS_DIR` | `<TRANSCRIPTS_DIR>/saved-videos` | Persisted source-video store (can live on a separate disk). | -| `SITES_DIR` | `<TRANSCRIPTS_DIR>/sites` | Per-site config (`sites/<id>/site.json` — every key in [SITE.md](SITE.md)). | -| `EXPORT_PUBLIC_DIR` | `<repo>/export/public` | Where the index writes paginated JSON. | -| `SETTINGS_FILE` | `<repo>/settings.json` | Site-settings file. | -| `YTDLP_BIN` | `yt-dlp` (PATH) | Pipeline downloader. | -| `WHISPER_BIN` | `whisper-cli` (PATH) | whisper.cpp binary. | -| `WHISPER_MODEL` | `~/whispercpp/whisper.cpp/models/ggml-base.en.bin` | whisper.cpp model file. | -| `CHOUGH_BIN` / `CHOUGH_URL` / `CHOUGH_MODEL` | `chough` / — / — | chough backend binary, remote server, model. | -| `PARAKEET_CLI` / `PARAKEET_MODEL` / `PARAKEET_STITCH_BIN` | `parakeet-cli` / — / `scripts/parakeet-stitch.mjs` | parakeet.cpp CLI, model, and wrapper. | -| `FFMPEG_BIN` / `FFPROBE_BIN` | `ffmpeg` / `ffprobe` (PATH) | Audio transcode + duration checks. | -| `RSYNC_BIN` | `rsync` (PATH) | Saved-video backup. | -| `WORKER_TOKEN` | — | Bearer token for the remote-worker transcription API (set on both ends when used), and for the `/api/ops/*` HTTP layer over the editor's actions — see [RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md#driving-the-editor-without-a-browser) and `pnpm ops`. Unset means both surfaces are off. | - -Feature-area docs cover their own env vars: [SCHEDULED_SYNC.md](SCHEDULED_SYNC.md) -(`SYNC_HEARTBEAT_SECONDS`, `SYNC_TICK_URL`, `SYNC_TICK_TOKEN`) and -[DEPLOY_CLOUDFLARE.md](DEPLOY_CLOUDFLARE.md) (R2 credentials). +To see what this machine has, run: + +```sh +pnpm archilyzer doctor +``` + +It is read-only: the checkout, the corpus and each channel's media, `settings.json`, +every binary the paths name plus each enabled worker's engine and model, umtool's +report-pipeline tools, and this checkout's port block. It exits 1 only for something +the machine is configured to do and cannot (an enabled worker's engine missing beside +a corpus, a settings file that does not parse, an override naming a missing binary). --- @@ -293,7 +292,7 @@ but you do need the Playwright browser: ```sh npx playwright install chromium # one-time: download the test browser -pnpm e2e # sequential run on the host (next dev, port 3001) +pnpm e2e # sequential run on the host (next dev, port 3011) ``` For a faster parallel run, `pnpm e2e:sharded` splits the suite across N Docker @@ -308,11 +307,12 @@ collisions, see [WORKTREES.md](WORKTREES.md). ## Where to go next -- [README.md](README.md) — project overview, pipeline modes, CLI shims. +- [README.md](README.md) — project overview, pipeline modes. - [SCHEDULED_SYNC.md](SCHEDULED_SYNC.md) — automatic per-channel sync (internal heartbeat or external cron). -- [DEPLOY_CLOUDFLARE.md](DEPLOY_CLOUDFLARE.md) — deploying to Cloudflare Pages + R2 - archive overflow. +- [PUBLISH.md](PUBLISH.md) — building and deploying sites: Cloudflare Pages, R2 + archive overflow, parallel builds in containers. +- [ENVIRONMENT.md](ENVIRONMENT.md) — every environment variable, by audience. - [WORKTREES.md](WORKTREES.md) — parallel development with per-worktree ports. - [mcp/README.md](mcp/README.md) — MCP server exposing the archive to Claude Code / Desktop / Cursor. diff --git a/homepage/content/README.md b/homepage/content/README.md @@ -6,7 +6,7 @@ exists for whoever edits the docs next. ## Why these are hand-written rather than rendered from the root docs -The obvious move is to render `SETUP.md`, `DEPLOY_CLOUDFLARE.md` and friends +The obvious move is to render `SETUP.md`, `PUBLISH.md` and friends directly, and keep one copy. That was rejected for a decisive reason: > `SETUP.md` says `git clone <this-repo-url>`. **There is no public repository.** @@ -31,8 +31,8 @@ its public counterpart needs the same change. | `docs/what-is-archilyzer.md` | `README.md` | package list, pipeline modes | | `docs/install.md` | `SETUP.md` | tool versions, env-var table, backend list, the Windows path | | `docs/operate.md` | `README.md`, `SCHEDULED_SYNC.md` | editor routes, scheduler settings | -| `docs/deploy-cloudflare.md` | `DEPLOY_CLOUDFLARE.md` | the 25 MB Pages limit, R2 options | -| `docs/deploy-docker.md` | `DEPLOY_DOCKER.md` | phase structure, settings names | +| `docs/deploy-cloudflare.md` | `PUBLISH.md` (Cloudflare, R2, cost-abuse) | the 25 MB Pages limit, R2 options | +| `docs/deploy-docker.md` | `PUBLISH.md` (building every site in containers) | phase structure, settings names | | `docs/ai-and-mcp.md` | `mcp/README.md` | tool names, `corpus.json` shape | | `docs/faq.md` | — (written for this site) | claims about cost and hardware | diff --git a/r2-proxy/README.md b/r2-proxy/README.md @@ -22,7 +22,7 @@ Then set the editor's **Settings → Archive overflow public URL** to the deploy `https://<name>.<you>.workers.dev` URL. Full setup and the alternative custom-domain path are documented in -[../DEPLOY_CLOUDFLARE.md](../DEPLOY_CLOUDFLARE.md). +[../PUBLISH.md](../PUBLISH.md#securing-downloads-against-cost-abuse). ## Scripts diff --git a/r2-proxy/src/index.ts b/r2-proxy/src/index.ts @@ -7,7 +7,7 @@ // request path straight to the bucket key — there is nothing per-site about it. // Deploy it ONCE to a free `<name>.workers.dev` subdomain (no custom domain, no // domain purchase, no WHOIS), then point every site's "Archive overflow public -// URL" at that single subdomain. See ../DEPLOY_CLOUDFLARE.md. +// URL" at that single subdomain. See ../PUBLISH.md. // // Why a Worker instead of the raw r2.dev URL: it gives us tunable, in-code rate // limiting (the cost/abuse backstop — Cloudflare's dashboard rate-limit rules diff --git a/r2-proxy/wrangler.toml b/r2-proxy/wrangler.toml @@ -4,7 +4,7 @@ # cd r2-proxy && pnpm dlx wrangler deploy # # One bucket + one Worker serves every export site (keys are namespaced by site -# id). See ../DEPLOY_CLOUDFLARE.md for the full setup. +# id). See ../PUBLISH.md for the full setup. name = "archilyzer-exports" main = "src/index.ts"