commit 73018db4c2daf5dca526c63f4741a58326bd55c6
parent d9204a6e301d8ac1774fac6c51a53edfefa1d63d
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Mon, 28 Sep 2026 02:04:13 -0400
docs: PUBLISH.md absorbs DEPLOY_CLOUDFLARE.md and DEPLOY_DOCKER.md; env vars point at ENVIRONMENT.md
PUBLISH.md is the one document for publishing: what gets published (a site,
the hub, the homepage — two Pages projects), the three ways to drive it (the
editor's /sites, `pnpm ops`, `pnpm archilyzer`) in one table, Cloudflare
Pages (a project must exist before its first deploy; a production deploy
inherits the checkout's branch), previews, download archives and R2, the
cost-abuse defenses, the container build pipeline, and registering the MCP
server against a published archive with `-- pnpm -C "$PWD" archilyzer mcp`.
The two old files are removed and every link to them outside plans/ and the
released changelog bullets now names PUBLISH.md (README, SETUP, AGENTS,
RUNNING_IN_DOCKER, r2-proxy, homepage/content/README.md).
README's and SETUP's hand-kept env-var tables become a paragraph and a
pointer to the generated ENVIRONMENT.md plus `pnpm archilyzer doctor`.
SETUP: `pnpm e2e` is port 3011, not 3001. CONTRIBUTING's "CLI shims"
(which named a `retry-failures.ts` that does not exist) is "The archilyzer
CLI"; a new path override is declared in envVars.ts, a new port in
ports.mjs. RUNNING_IN_DOCKER names parakeet.cpp beside whisper.cpp.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Diffstat:
12 files changed, 526 insertions(+), 488 deletions(-)
diff --git a/AGENTS.md b/AGENTS.md
@@ -208,7 +208,7 @@ back to the serial host build it already handles. Do not try to make
docker-in-docker work.
See [RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md) — which is about *running the
-apps*, not [DEPLOY_DOCKER.md](DEPLOY_DOCKER.md), which is about *building sites*.
+apps*, not [PUBLISH.md](PUBLISH.md), which is about *building and publishing sites*.
# The corpus, and where the live sites are configured
diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md
@@ -34,7 +34,8 @@ pnpm install
pnpm dev:editor # editor at http://localhost:3001
pnpm build # build the static site under export/out/
pnpm start:export # serve export/out/ at http://localhost:3000
-pnpm build:index # rebuild the search index in-process
+pnpm build:index # rebuild the search index in-process (= pnpm archilyzer index)
+pnpm archilyzer doctor # what this machine has: tools, corpus, settings, ports
```
`pnpm lint` at the root only lints `export/`. **The editor has no eslint config**, so
@@ -128,18 +129,29 @@ buttons" is red under `next dev` (which injects a Dev Tools button) and green un
**The editor e2e suite runs in dev mode and therefore never prerenders.** A change to
a layout or a client component is not verified until `pnpm build` passes too.
-## CLI shims
+## The `archilyzer` CLI
-The same controllers the editor uses are exposed as terminal shims under
-`common/bin/`, which is how you drive the pipeline headlessly or from cron:
+The same controllers the editor uses are one command line, `common/bin/archilyzer.ts`,
+which is how you drive the pipeline headlessly or from cron. `pnpm archilyzer
+<command>` from the repo root is the short form of `pnpm --filter
+yt-dlp-transcript-common exec tsx bin/archilyzer.ts <command>`; `pnpm archilyzer
+--help` lists every command.
```bash
-pnpm --filter yt-dlp-transcript-common exec tsx bin/build-index.ts
-pnpm --filter yt-dlp-transcript-common exec tsx bin/transform.ts --channel <slug>
-pnpm --filter yt-dlp-transcript-common exec tsx bin/retry-failures.ts --channel <slug>
-pnpm --filter yt-dlp-transcript-common exec tsx bin/verify-transcripts.ts --channel <slug>
+pnpm archilyzer doctor # read-only: can this machine do what it is configured to?
+pnpm archilyzer index # the LMDB index
+pnpm archilyzer run diarization <channel> [ids…] # one catalogued operation, offline, as the editor's job
+pnpm archilyzer build site <id> # publish: see PUBLISH.md
+pnpm archilyzer verify transcripts --channel <slug>
+pnpm archilyzer mcp # the MCP server on stdio
```
+The table is `archilyzer.ts`; the machinery (parser, lookup, usage) is `_cli.ts`. Every
+file in `common/bin/` is reachable from a row — a test fails otherwise. A bin that
+parses its own flags is a *passthrough* row, run as a child with its argv untouched.
+`run` refuses sync, the metadata scan, downloads and transcription: they run on the
+editor's paced download queue and worker pool, which a second process must not race.
+
## Internals worth knowing
**This is not the Next.js you may know.** The workspace tracks a recent major and its
@@ -155,7 +167,11 @@ deprecation notices.
`export/public/{summaries,transcripts}/` and renders a search/filter UI into a
static `out/`.
- `getPaths()` (`common/lib/paths.ts`) is the single resolver for every path and
- binary. Nothing should hardcode a location; add an env override there instead.
+ binary. Nothing should hardcode a location; add an env override there instead, and
+ declare it in `common/lib/envVars.ts` (a test fails until you do, and
+ `pnpm archilyzer docs env` regenerates [ENVIRONMENT.md](ENVIRONMENT.md)). A new
+ local server's port goes in `common/lib/ports.mjs`, which `pnpm wt` offsets per
+ worktree.
- `common/lib/project.ts` holds **product** identity (the name, the project URL) and
is deliberately import-free. **Operator** identity — what a given deployment calls
itself — lives in settings and per-site config. A string that should change when
diff --git a/DEPLOY_CLOUDFLARE.md b/DEPLOY_CLOUDFLARE.md
@@ -1,332 +0,0 @@
-# Deploying to Cloudflare (Pages + R2 archive overflow)
-
-An **Archilyzer** site is deployed to **Cloudflare Pages**; large download archives that
-exceed Pages' per-file limit overflow to **Cloudflare R2**. This guide covers the R2
-setup and — importantly — how to configure Cloudflare so that **public archive
-downloads can't be abused to drive up your bill**.
-
-Everything here fits inside Cloudflare's **free tier**.
-
-- [How archives are served](#how-archives-are-served)
-- [Preview deployments](#preview-deployments)
-- [One-time R2 setup](#one-time-r2-setup)
-- [Securing downloads against cost-abuse](#securing-downloads-against-cost-abuse)
-- [Cost expectations](#cost-expectations)
-
----
-
-## How archives are served
-
-At build time, `compose-site.ts` generates one transcript zip and one live-chat zip
-**per channel** into `export/public/archives/`, and records them in
-`public/archives/manifest.json`. The `/downloads` page and the header **Downloads**
-link read that manifest.
-
-Cloudflare Pages rejects any single asset larger than **25 MB**. Real channels blow
-past that easily (a channel's live-chat zip can be hundreds of MB). So the pipeline
-splits archives by size:
-
-| Archive size | Where it's served from | Manifest entry |
-|---|---|---|
-| ≤ 25 MB | Cloudflare **Pages** (shipped in `out/`, free) | `filename`, no `url` |
-| > 25 MB, R2 configured | Cloudflare **R2**, uploaded on deploy | `url` → R2 |
-| > 25 MB, R2 **not** configured | not served | `oversize: true`, shown as "Too large to host" |
-
-Oversize archives are staged during compose into `export/.r2-staging/<siteId>/archives/`
-(gitignored, kept out of `public/`), then uploaded by the editor's **Deploy** /
-**Build & deploy** actions to `<bucket>/<siteId>/archives/<file>.zip` **before** the
-Pages deploy runs, so the manifest URLs resolve immediately.
-
-> **Note:** every deploy path uploads oversize archives to R2 before the Pages
-> deploy: the editor's **Deploy** / **Build & deploy**, `pnpm ops deploy-site`, and
-> `archilyzer deploy site <id>` (which `pnpm run deploy` in `export/` runs, with
-> `SITE_ID`). A build alone (`archilyzer build site`, `pnpm run build`) only
-> *stages* them in `export/.r2-staging/`. The bucket is read from `settings.json`;
-> the R2 credentials come from the environment (see "3. Authenticate" below).
-
-Uploads go through R2's **S3 API** (via the AWS SDK's multipart uploader), not
-`wrangler r2 object put` — wrangler caps a single upload at **300 MiB**, and real
-live-chat archives are larger (multipart has no such limit). That's why archive
-uploads need S3 credentials (step 3) in addition to the wrangler auth your Pages
-deploy already uses.
-
----
-
-## Preview deployments
-
-A **preview** is the same built bundle deployed to a branch that is not the Pages
-project's production branch. Cloudflare publishes it at a **branch alias** —
-
-```
-https://<branch>.<project>.pages.dev
-```
-
-— and leaves the live site alone. Each deploy also gets an immutable
-per-deployment URL (`https://<hash>.<project>.pages.dev`), which wrangler prints as
-"Deployment complete! Take a peek over at …"; the editor repeats both on a
-`[preview]` line at the end of the job log, because the streamed log scrolls.
-
-The alias is a function of the project and the branch and nothing else, so it is
-known *before* the deploy runs — which is why the editor can link it while you are
-still typing the name.
-
-**Branch names** must be 1–28 lowercase letters, digits and dashes, starting and
-ending with a letter or digit. That is exactly what Cloudflare's alias sanitizer
-preserves verbatim, so the alias shown is the alias that resolves. `main`, `master`
-and `production` are refused: a deploy to the production branch is not a preview, it
-is the live site.
-
-### The three surfaces
-
-| Surface | How |
-|---|---|
-| Editor | A site's **Publish** tab → *Individual steps* → **Deploy a preview**: type a branch, press **Deploy preview**. The production button beside it says **Deploy to production**. |
-| Ops API | `POST /api/ops/deploy-site` `{ "siteId": "...", "preview": "<branch>" }` — deploy-only, of the already-built `export/out`. `POST /api/ops/build-deploy` takes `preview` too (build *then* preview-deploy). Both answer with `previewUrl`. |
-| CLI | `pnpm ops deploy-site --json '{"siteId":"anilyzer","preview":"tags-exclude"}' --wait` — the alias is printed on its own line after the log. |
-
-The high-value loop is **build once, preview, then promote**: `build-site` (or the
-Publish tab's *Build static export*), then `deploy-site` with a `preview`, look at
-it, then `deploy-site` again with no `preview` — the same `export/out`, unrebuilt.
-
-Deploy-only ships whatever is in `export/out`, which the basic build composes one
-site at a time into a single shared directory — so it **refuses, before starting a
-job, if `export/out` holds a build of another site** (or no build at all), naming
-the site to build first. `build-deploy` cannot hit this: it builds.
-
-### Two things to know
-
-**A preview shares the production R2 archive bucket.** R2 has no per-branch
-namespace, and the keys are `<siteId>/archives/<file>.zip` either way. In practice
-this is cheap and harmless — the upload skips any object R2 already holds at the same
-size, and an unchanged channel re-zips byte-stable — but a *changed* archive
-replaces the one production's manifest links to. The editor says so once at the top
-of every preview deploy.
-
-**Who can open a preview is a Cloudflare setting, not ours.** Pages projects have a
-*preview deployment access* setting (Settings → General): **public** by default, or
-restricted to Cloudflare Access. A default-configured project's preview URL is
-world-readable by anyone who has the link.
-
-> **A production deploy inherits the checkout's git branch.** `wrangler pages deploy`
-> with no `--branch` infers one from the git repository it runs in, so running a
-> *production* deploy from a feature-branch checkout silently produces a preview
-> instead. This predates the preview feature and is unchanged: only the preview path
-> passes `--branch`. If a "production" deploy did not go live, check what branch the
-> editor's checkout is on.
-
----
-
-## One-time R2 setup
-
-### 1. Create the bucket
-
-One bucket serves **all** your sites — objects are namespaced by `siteId`, so you do
-**not** need one bucket per site. In the dashboard (**R2 → Create bucket**) or via the
-same `wrangler` CLI the deploy already uses:
-
-```sh
-wrangler r2 bucket create my-archives-bucket
-```
-
-### 2. Expose it publicly
-
-R2 buckets are private by default, and the Downloads page links to plain
-`https://…/<file>.zip` URLs, so the bucket needs public read access. There are three
-ways to expose it; the first needs **no domain purchase** and is the recommended one:
-
-- **A Worker on `*.workers.dev` (recommended — no domain, no WHOIS):** deploy the
- bundled `r2-proxy/` Worker to a free `<name>.workers.dev` subdomain (just like Pages'
- `*.pages.dev`). It streams objects from the bucket and gives you edge caching **and**
- tunable rate limiting in code — the same cost defenses a custom domain would, without
- owning a domain. One Worker serves **every** site. See
- [Serving via a Worker on workers.dev](#option-a--serving-via-a-worker-on-workersdev-no-domain)
- below. Set the editor's public URL to `https://<name>.workers.dev`.
-- **Custom domain:** bucket **Settings → Public access → Custom Domains → Connect
- Domain**, e.g. `archives.example.com` (the domain must be on your Cloudflare account).
- Routes downloads through Cloudflare's CDN, WAF, caching, and dashboard rate-limiting
- rules. See [Securing via a custom domain](#option-b--securing-via-a-custom-domain).
-- **`r2.dev` subdomain (quick test only):** bucket **Settings → Public access → Allow
- Access**. You get `https://pub-abc123.r2.dev` — free and domain-free, but Cloudflare
- throttles `r2.dev` and gives you **no** cache / rate-limit control of your own. Fine
- for a smoke test; use the Worker for a real instance.
-
-### 3. Authenticate — two credentials
-
-**a) `wrangler` (for the Pages deploy).** The `wrangler pages deploy` step reuses
-whatever auth you already use. If it works today, nothing to do. Otherwise run
-`wrangler login`, or set `CLOUDFLARE_API_TOKEN`.
-
-**b) R2 S3 API keys (for the archive uploads).** Archive objects are uploaded over
-R2's S3-compatible API, which needs an Access Key ID + Secret. In the dashboard:
-**R2 → Manage R2 API Tokens → Create API token**, permission **Object Read & Write**,
-scoped to your bucket. Then set three environment variables where the editor runs:
-
-```sh
-export R2_ACCESS_KEY_ID=<access key id>
-export R2_SECRET_ACCESS_KEY=<secret access key>
-export CLOUDFLARE_ACCOUNT_ID=<your account id> # used to build the S3 endpoint
-```
-
-The account id is on the R2 overview page; the S3 endpoint is derived as
-`https://<CLOUDFLARE_ACCOUNT_ID>.r2.cloudflarestorage.com`. Keep these in the
-environment (a shell profile, a systemd unit, a `.env` the editor loads) — **not** in
-the settings JSON, which isn't a place for secrets. If a deploy has oversize archives
-to upload but these are unset, it fails *before* the Pages deploy (so the site never
-links to a missing file) with a message pointing back here.
-
-### 4. Point the editor at the bucket
-
-In the editor, open **Settings** and set:
-
-| Field | Value |
-|---|---|
-| **Archive overflow storage (R2 bucket)** | the bucket name, e.g. `my-archives-bucket` |
-| **Archive overflow public URL** | the public base from step 2, e.g. `https://archives.example.com` (no trailing slash needed) |
-
-Leave both blank to keep the old behavior (oversize archives dropped, shown as "Too
-large to host").
-
-On the next **Build & deploy**, oversize archives upload to R2 and the Downloads page
-(and the header link, which reappears once anything is hostable) point at them.
-
----
-
-## Securing downloads against cost-abuse
-
-The threat: someone scripts repeated downloads of large archives to run up your bill.
-
-**The reassuring part — R2 egress is free.** Unlike S3, Cloudflare R2 charges **$0**
-for bandwidth/egress. An attacker looping downloads of a 480 MB zip cannot run up a
-bandwidth bill. The *only* metered cost from reads is **Class B operations** (10M free
-per month, then $0.36/M) — and the defenses below make even that hard to reach.
-
-You get these controls one of two ways — **a Worker on `workers.dev`** (no domain) or
-**a custom domain**. Pick one; both are covered below. The Worker path is recommended
-if you don't want to own a domain.
-
-## Option A — Serving via a Worker on workers.dev (no domain)
-
-The bundled **`r2-proxy/`** Worker is a small, self-contained project that binds the R2
-bucket and serves archive objects on a free `<name>.workers.dev` subdomain — the same
-domain-free model as Pages' `*.pages.dev`. It's a **pure passthrough**: the request path
-`<siteId>/archives/<file>.zip` maps straight to the bucket key, so **one Worker serves
-every site** (deploy it once, not per site — that's the whole point of keying objects by
-site id). What it gives you, all on the free tier:
-
-- **Edge caching** — full downloads are cached with the Cache API, so repeat pulls skip
- R2 (no billable Class B op). It honors the `Cache-Control` we set on each object.
-- **Rate limiting** — Cloudflare's native, free rate-limit binding caps requests per
- client IP + file (dashboard rate-limit rules need a paid zone; this doesn't).
-- **Path allow-listing** — it only serves `*/archives/*.zip`, never arbitrary keys.
-- **Range / resumable downloads** — honors `Range` requests so big zips can resume.
-
-**Deploy it (once for the whole instance):**
-
-1. Edit `r2-proxy/wrangler.toml` and set `bucket_name` to the **same bucket** you use in
- the editor's Settings. Optionally rename the Worker (`name`) and tune the rate limit
- (`limit` / `period`).
-2. From the repo root:
- ```sh
- cd r2-proxy
- pnpm install # first time only
- pnpm run deploy # = wrangler deploy, reusing your host wrangler auth
- ```
- wrangler prints the deployed URL, e.g. `https://ytdlp-archive-proxy.<you>.workers.dev`.
- (An "unsafe fields are experimental" warning for the rate-limit binding is expected.)
-3. In the editor's **Settings**, set **Archive overflow public URL** to that
- `workers.dev` URL. Re-deploy a site and its Downloads links resolve through the Worker.
-
-The Worker code lives in `r2-proxy/src/index.ts` — the caching and rate-limit logic are
-small and commented if you want to adjust them.
-
-> **Free-tier limit:** Workers Free allows **100,000 requests/day** (resets daily). Far
-> more than a downloads endpoint needs; if you ever exceed it, requests get a `429`
-> (fail closed — no surprise bill) until the next day, or upgrade to Workers Paid ($5/mo).
-
-The archive `Cache-Control` is also set at upload time (`Cache-Control: public,
-max-age=3600`, constant `ARCHIVE_CACHE_CONTROL` in
-`common/publish/build.ts`); the Worker reads it back when caching. Archive
-filenames are stable and overwritten in place on re-deploy, so this 1-hour bound is what
-keeps a re-uploaded archive from being served stale for long — raise it if your archives
-rarely change.
-
-## Option B — Securing via a custom domain
-
-If you'd rather use a custom domain (its own upsides: dashboard WAF, managed bot rules,
-and rate-limiting rules without touching code), connect it per
-[step 2 above](#2-expose-it-publicly) and add these, in order of impact:
-
-### 1. Edge caching
-
-Served through a custom domain, Cloudflare's CDN caches each archive at the edge (it
-honors the `Cache-Control: public, max-age=3600` we set at upload), so repeated
-downloads of the same file are served from cache and **never hit R2**. To make caching
-aggressive, add a **Cache Rule** (dashboard: **Caching → Cache Rules → Create**):
-
-- **When:** `URI Path` contains `/archives/`
-- **Then:** *Eligible for cache*, **Edge TTL → Override → 1 day** (or longer).
-
-If you raise the TTL a lot, **purge the cache on deploy** (dashboard **Caching → Purge**,
-or `wrangler`/API) so a re-uploaded archive isn't served stale.
-
-### 2. Rate limiting — the hard backstop
-
-A Rate Limiting rule caps how fast any single client can pull archives, stopping a
-flood that misses cache. Dashboard: **Security → WAF → Rate limiting rules → Create**
-(the free plan includes one rule):
-
-- **When incoming requests match:** `URI Path` contains `/archives/`
-- **Rate:** e.g. **20 requests per 1 minute** per client IP
-- **Then:** *Block* for 10 minutes (or *Managed Challenge*).
-
-Tune the threshold to real usage — legitimate users download a handful of files, not
-dozens per minute.
-
-### 3. Bot Fight Mode + WAF managed rules
-
-Dashboard: **Security → Bots → Bot Fight Mode** (free). Blocks the low-effort scripted
-abuse that makes up most of this traffic. The free **WAF managed ruleset** adds a
-baseline of protection at no cost.
-
-### 4. Hotlink protection (optional)
-
-Stops other sites embedding your archives and spending your ops budget serving their
-audience. A WAF custom rule (**Security → WAF → Custom rules**):
-
-- **When:** `URI Path` contains `/archives/` **and** `Referer` does not contain your
- domain **and** `Referer` is not empty
-- **Then:** *Block*.
-
-(Allow an empty `Referer` so direct clicks and privacy-conscious browsers still work.)
-
-## Billing / usage alerts (either option)
-
-R2 has no hard spend cap, but Cloudflare **Notifications** (dashboard: **Notifications
-→ Add**) can email you when R2 storage or Class A/B operations cross a threshold —
-cheap insurance so nothing surprises you.
-
-## What to skip
-
-**Signed URLs / token-gated downloads** are the heavyweight option — they add key
-management and friction for legitimate users. Given egress is free and caching
-neutralizes the ops cost, they're overkill for *cost* defense (the `r2-proxy` Worker is
-a plain passthrough, not an access gate). Only reach for signed URLs if you want
-*access control* (private archives), not cost control.
-
----
-
-## Cost expectations
-
-Measured across all sites in this project, total compressed archives are **≈ 2–4.5 GB**.
-Against R2's free tier:
-
-| Resource | Free tier / month | This project's usage |
-|---|---|---|
-| **Egress / bandwidth** | unlimited, **$0** | irrelevant — no egress charge exists |
-| **Storage** | 10 GB-month | ~2–4.5 GB — comfortable |
-| **Class A ops** (writes) | 1,000,000 | ~one PUT per archive per deploy — negligible |
-| **Class B ops** (reads) | 10,000,000 | mostly absorbed by CDN cache |
-
-The realistic bill for hosting these archives is **$0**. The configuration above exists
-to keep it that way under adversarial traffic, not because normal usage is close to any
-limit.
diff --git a/DEPLOY_DOCKER.md b/DEPLOY_DOCKER.md
@@ -1,82 +0,0 @@
-# Docker export build pipeline
-
-> **Not the document you want if you are trying to *run* the apps in containers.**
-> That is [RUNNING_IN_DOCKER.md](./RUNNING_IN_DOCKER.md) — `docker compose up` and a
-> working archive server, built from the root `Dockerfile`. This page is about
-> *building sites*: fanning per-site export builds out across containers, using
-> `Dockerfile.build`. The two share nothing but the word "docker".
-
-Archilyzer's Docker build mode builds **every site in parallel** in isolated containers, then
-deploys them serially — a large speedup when you host several sites, and stronger
-isolation than the basic single-process build. This is opt-in: set **Build
-pipeline → Docker** in Settings (or the toggle on the Deploy page). Basic mode is
-unchanged and remains the default.
-
-## Prerequisites
-
-- A container engine: **Docker**, or **podman** (set `DOCKER_BIN=podman`). Rootless
- podman is a good fit — it maps container files to your host user automatically.
-- The editor host still needs Node + pnpm (Phase A and deploy run on the host) and
- your Cloudflare/R2 credentials in the environment (see
- [DEPLOY_CLOUDFLARE.md](./DEPLOY_CLOUDFLARE.md)). Credentials are **never** passed
- into a container — deploy runs on the host.
-
-The build image is built (and cached) automatically from `Dockerfile.build` the
-first time you run; edit **Build image** / **Dockerfile** in Settings to override
-the tag/path.
-
-## How it works
-
-Trigger it with **Build all sites** on the Deploy page. One managed job runs three
-ordered phases:
-
-1. **Phase A — shared, on the host, once.** `build:data` (search index +
- `.export-index` staging) then `build:archives` (warm the shared archive-zip
- cache for the union of all sites' channels). Only the host writes this shared
- state, so containers never race it. This phase is serial and is the long pole on
- a cold build; on a warm rebuild it's near-instant (unchanged channels are
- skipped).
-2. **Phase B — per-site, in parallel containers.** Each site's `compose:site +
- next build` runs in its own container, capped by **Max parallel builds**. Each
- writes an isolated `out/` under `export/.export-builds/<siteId>/`. Containers
- mount the corpus/index/staging/archive cache **read-only**. (Network is left on:
- `next build` fetches the site's fonts from Google via `next/font/google`;
- isolation comes from the read-only mounts, per-site output dir, and non-root
- user.)
-3. **Phase C — deploy, on the host, serially.** After every build finishes, each
- built site is deployed in turn (oversize-archive R2 upload, then `wrangler pages
- deploy`). A single site failing to build or deploy is reported and skipped; the
- rest still ship.
-
-If no container engine is available, the action logs a notice and falls back to a
-serial host build+deploy (one site at a time).
-
-## Mounts (per Phase-B container)
-
-| Host | Container | Mode |
-|---|---|---|
-| `transcripts/` (corpus + `index.mdb` + archive cache) | `/data/transcripts` | ro |
-| `export/.export-index` (shared + per-site staging) | `/data/export/.export-index` | ro |
-| `export/.export-builds/<siteId>` (public/out/.next/caches) | `/site` | rw |
-| `settings.json` (build config, mounted fresh — not baked) | `/data/settings.json` | ro |
-
-The per-site `/site` mount is persistent, so incremental `next build` (`.next`) and
-incremental compose (`.compose-cache`) stay warm across builds.
-
-## Tuning & environment
-
-- **Max parallel builds** (setting) — how many site containers run at once. Each
- `next build` can use up to ~8 GB; a safe starting point is `floor(RAM_GB / 9)`.
-- `DOCKER_BIN` — container binary (default `docker`; e.g. `podman`).
-- `DOCKER_BUILD_MEMORY`, `DOCKER_BUILD_CPUS` — optional per-container `--memory` /
- `--cpus` caps so a fan-out can't OOM/peg the host.
-- `BUILD_ARCHIVES=0` (or the **Skip archive zips** checkbox) — skip the archive
- warm + per-site archive materialize for a faster build with no download bundles.
-
-## Notes
-
-- Containers run as your host uid/gid (`-u`), so files under `.export-builds/` are
- host-owned, not root-owned.
-- The image bakes the repo source + deps; a code change rebuilds it, but Docker
- layer caching keeps that cheap (deps re-install only when the lockfile moves).
-- `.export-builds/` is gitignored and excluded from the image build context.
diff --git a/PUBLISH.md b/PUBLISH.md
@@ -0,0 +1,443 @@
+# Publishing
+
+What gets published, how to build and deploy it, and how to keep a public archive
+cheap and safe. For *running* the apps in containers see
+[RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md); every environment variable named here
+is in [ENVIRONMENT.md](ENVIRONMENT.md).
+
+- [What gets published](#what-gets-published)
+- [The three ways to drive it](#the-three-ways-to-drive-it)
+- [Cloudflare Pages](#cloudflare-pages)
+- [Preview deployments](#preview-deployments)
+- [Download archives and R2](#download-archives-and-r2)
+- [Securing downloads against cost-abuse](#securing-downloads-against-cost-abuse)
+- [Building every site in containers](#building-every-site-in-containers)
+- [Reading a published archive from Claude Code](#reading-a-published-archive-from-claude-code)
+
+---
+
+## What gets published
+
+Three static artefacts, each pre-rendered to plain HTML and JSON — no database, no
+server-side code, servable by anything:
+
+| Artefact | What it is | Built into | Config |
+|---|---|---|---|
+| A **site** | The export app over the channels one site selects: search, transcripts, charts, downloads. One corpus can publish several. | `export/out` (docker mode: `export/.export-builds/<siteId>/out`) | `transcripts/sites/<id>/site.json` ([SITE.md](SITE.md)) |
+| The **hub** | The export app in hub mode: federated search across the family of sites. | `export/out` | `transcripts/sites/_homepage/homepage.json` |
+| The **homepage** | The project's own site (`homepage/`). | `homepage/out` | — |
+
+The hub and the homepage are two different Pages projects: the hub deploys to the
+project `homepage.json` names (e.g. `archilyzer-hub`) and is refused `archilyzer`,
+which is the homepage's.
+
+A site build has three steps, run in `export/`: the **data phase** (the LMDB index,
+the stats datasets and the chart templates — `archilyzer index`, `build stats`,
+`build templates`), **compose** (the site's slice of the shared index into
+`export/public`, plus its download archives — `archilyzer compose site <id>`) and
+`next build`. `--nodata` skips the data phase and reuses the last one's staging.
+
+## The three ways to drive it
+
+The editor's **/sites** page, `pnpm ops` (HTTP to a running editor, with its
+`WORKER_TOKEN`) and the `archilyzer` CLI (local, no editor needed) call the same
+entry points in `common/publish/build.ts`.
+
+| To | Editor | `pnpm ops` | `pnpm archilyzer …` |
+|---|---|---|---|
+| Build one site | a site's **Publish** tab → *Build static export* | `build-site` | `build site <id> [--nodata] [--skip-archives]` |
+| Deploy the built site | *Deploy to production* / *Deploy preview* | `deploy-site` | `deploy site <id> [--preview <branch>]` |
+| Build, then deploy | *Build & deploy* | `build-deploy` | `build site <id>` then `deploy site <id>` |
+| Build every site | /sites → **Build all sites** | — | `build all [--skip-archives]` |
+| The hub | /sites → Hub → **Build hub** / **Deploy hub** | `build-hub`, `deploy-hub` | `build hub`, `deploy hub [--preview <branch>]` |
+| The homepage | — | — | `build homepage`, `deploy homepage` |
+
+`pnpm archilyzer <command>` is the short form of
+`pnpm --filter yt-dlp-transcript-common exec tsx bin/archilyzer.ts <command>`;
+`pnpm archilyzer --help` lists every command, and `pnpm archilyzer doctor` checks that
+this machine has what a build needs. From `export/`, `pnpm run build` is
+`archilyzer build site` (with `SITE_ID`) and `pnpm run deploy` is `archilyzer deploy
+site`.
+
+Deploy-only ships whatever is in `export/out`, which the basic build composes one
+site at a time into a single shared directory — so it **refuses, before starting a
+job, if `export/out` holds a build of another site** (or no build at all), naming the
+site to build first. A deploy queued behind another site's build re-checks when it
+starts. `build-deploy` cannot hit this: it builds.
+
+---
+
+## Cloudflare Pages
+
+Sites deploy to **Cloudflare Pages** with `wrangler pages deploy`; download archives
+too large for Pages' per-file limit overflow to **Cloudflare R2**. Everything here
+fits inside Cloudflare's **free tier**.
+
+- **A Pages project must exist before its first deploy.** wrangler offers to create a
+ missing project only on an interactive terminal, and the deploy's stdin is a pipe,
+ so a missing project fails at once with wrangler's own "does not exist" sentence.
+ Create it first: `pnpm dlx wrangler pages project create <name> --production-branch
+ main`. A site's project is `site.json`'s `cloudflareProject`.
+- **wrangler's own auth.** The deploy reuses whatever auth wrangler already has —
+ `wrangler login`, or `CLOUDFLARE_API_TOKEN` in the environment.
+- **A production deploy inherits the checkout's git branch.** `wrangler pages deploy`
+ with no `--branch` infers one from the repository it runs in, so a *production*
+ deploy from a feature-branch checkout silently produces a preview instead. Only the
+ preview path passes `--branch`; `deploy homepage` passes `--branch main`. If a
+ "production" deploy did not go live, check what branch the checkout is on.
+
+## Preview deployments
+
+A **preview** is the same built bundle deployed to a branch that is not the Pages
+project's production branch. Cloudflare publishes it at a **branch alias** —
+
+```
+https://<branch>.<project>.pages.dev
+```
+
+— and leaves the live site alone. Each deploy also gets an immutable
+per-deployment URL (`https://<hash>.<project>.pages.dev`), which wrangler prints as
+"Deployment complete! Take a peek over at …"; the editor repeats both on a
+`[preview]` line at the end of the job log, because the streamed log scrolls.
+
+The alias is a function of the project and the branch and nothing else, so it is
+known *before* the deploy runs — which is why the editor can link it while you are
+still typing the name.
+
+**Branch names** must be 1–28 lowercase letters, digits and dashes, starting and
+ending with a letter or digit. That is exactly what Cloudflare's alias sanitizer
+preserves verbatim, so the alias shown is the alias that resolves. `main`, `master`
+and `production` are refused: a deploy to the production branch is not a preview, it
+is the live site.
+
+| Surface | How |
+|---|---|
+| Editor | A site's **Publish** tab → *Individual steps* → **Deploy a preview**: type a branch, press **Deploy preview**. The production button beside it says **Deploy to production**. |
+| Ops API | `POST /api/ops/deploy-site` `{ "siteId": "...", "preview": "<branch>" }` — deploy-only, of the already-built `export/out`. `POST /api/ops/build-deploy` takes `preview` too (build *then* preview-deploy). Both answer with `previewUrl`. `pnpm ops deploy-site --json '{"siteId":"anilyzer","preview":"tags-exclude"}' --wait` prints the alias on its own line after the log. |
+| CLI | `pnpm archilyzer deploy site anilyzer --preview tags-exclude` |
+
+The high-value loop is **build once, preview, then promote**: build the site, deploy
+it with a preview branch, look at it, then deploy again with no preview — the same
+`export/out`, unrebuilt.
+
+**A preview shares the production R2 archive bucket.** R2 has no per-branch
+namespace, and the keys are `<siteId>/archives/<file>.zip` either way. In practice
+this is cheap and harmless — the upload skips any object R2 already holds at the same
+size, and an unchanged channel re-zips byte-stable — but a *changed* archive replaces
+the one production's manifest links to. The editor says so once at the top of every
+preview deploy.
+
+**Who can open a preview is a Cloudflare setting, not ours.** Pages projects have a
+*preview deployment access* setting (Settings → General): **public** by default, or
+restricted to Cloudflare Access. A default-configured project's preview URL is
+world-readable by anyone who has the link.
+
+---
+
+## Download archives and R2
+
+At compose time, `archilyzer compose site` generates one transcript zip and one
+live-chat zip **per channel** into `export/public/archives/`, and records them in
+`public/archives/manifest.json`. The `/downloads` page and the header **Downloads**
+link read that manifest. A site can turn its archives off (`site.json` `archives: false`), and a build
+can skip them (`--skip-archives`, `BUILD_ARCHIVES=0`, or the **Skip archive zips**
+checkbox).
+
+Cloudflare Pages rejects any single asset larger than **25 MB**, and real channels
+blow past that easily (a channel's live-chat zip can be hundreds of MB). So the
+pipeline splits archives by size:
+
+| Archive size | Where it's served from | Manifest entry |
+|---|---|---|
+| ≤ 25 MB | Cloudflare **Pages** (shipped in `out/`, free) | `filename`, no `url` |
+| > 25 MB, R2 configured | Cloudflare **R2**, uploaded on deploy | `url` → R2 |
+| > 25 MB, R2 **not** configured | not served | `oversize: true`, shown as "Too large to host" |
+
+(The cap is `MAX_ARCHIVE_BYTES`, or a site's own `archiveMaxBytes`; `0` = no cap.)
+
+Oversize archives are staged during compose into `export/.r2-staging/<siteId>/archives/`
+(gitignored, kept out of `public/`), then uploaded to
+`<bucket>/<siteId>/archives/<file>.zip` **before** the Pages deploy runs, so the
+manifest URLs resolve immediately. **Every deploy path uploads them** — the editor's
+**Deploy** / **Build & deploy**, `pnpm ops deploy-site`, and `archilyzer deploy site`
+(which `pnpm run deploy` in `export/` runs). A build alone only *stages* them. The
+bucket is read from `settings.json`; the R2 credentials come from the environment
+(step 3 below).
+
+Uploads go through R2's **S3 API** (the AWS SDK's multipart uploader), not
+`wrangler r2 object put` — wrangler caps a single upload at **300 MiB**, and real
+live-chat archives are larger (multipart has no such limit). That is why archive
+uploads need S3 credentials in addition to the wrangler auth the Pages deploy uses.
+
+### One-time R2 setup
+
+**1. Create the bucket.** One bucket serves **all** your sites — objects are
+namespaced by `siteId`, so you do **not** need one bucket per site. In the dashboard
+(**R2 → Create bucket**) or with the same wrangler the deploy already uses:
+
+```sh
+wrangler r2 bucket create my-archives-bucket
+```
+
+**2. Expose it publicly.** R2 buckets are private by default, and the Downloads page
+links to plain `https://…/<file>.zip` URLs, so the bucket needs public read access.
+There are three ways; the first needs **no domain purchase** and is the recommended
+one:
+
+- **A Worker on `*.workers.dev` (recommended — no domain, no WHOIS):** deploy the
+ bundled `r2-proxy/` Worker to a free `<name>.workers.dev` subdomain (just like Pages'
+ `*.pages.dev`). It streams objects from the bucket and gives you edge caching **and**
+ tunable rate limiting in code — the same cost defenses a custom domain would, without
+ owning a domain. One Worker serves **every** site. See
+ [Option A](#option-a--a-worker-on-workersdev-no-domain) below. Set the editor's
+ public URL to `https://<name>.workers.dev`.
+- **Custom domain:** bucket **Settings → Public access → Custom Domains → Connect
+ Domain**, e.g. `archives.example.com` (the domain must be on your Cloudflare account).
+ Routes downloads through Cloudflare's CDN, WAF, caching, and dashboard rate-limiting
+ rules. See [Option B](#option-b--a-custom-domain).
+- **`r2.dev` subdomain (quick test only):** bucket **Settings → Public access → Allow
+ Access**. You get `https://pub-abc123.r2.dev` — free and domain-free, but Cloudflare
+ throttles `r2.dev` and gives you **no** cache / rate-limit control of your own. Fine
+ for a smoke test; use the Worker for a real instance.
+
+**3. The R2 S3 credentials.** Archive objects are uploaded over R2's S3-compatible
+API, which needs an Access Key ID + Secret. In the dashboard: **R2 → Manage R2 API
+Tokens → Create API token**, permission **Object Read & Write**, scoped to your
+bucket. Then set three environment variables where the editor (or the CLI) runs:
+
+```sh
+export R2_ACCESS_KEY_ID=<access key id>
+export R2_SECRET_ACCESS_KEY=<secret access key>
+export CLOUDFLARE_ACCOUNT_ID=<your account id> # used to build the S3 endpoint
+```
+
+The account id is on the R2 overview page; the S3 endpoint is derived as
+`https://<CLOUDFLARE_ACCOUNT_ID>.r2.cloudflarestorage.com`. Keep these in the
+environment (a shell profile, a systemd unit, a `.env` the editor loads) — **not** in
+the settings JSON, which is not a place for secrets. If a deploy has oversize
+archives to upload but these are unset, it fails *before* the Pages deploy (so the
+site never links to a missing file) with a message pointing back here.
+
+**4. Point the editor at the bucket.** In **Settings**:
+
+| Field | Value |
+|---|---|
+| **Archive overflow storage (R2 bucket)** | the bucket name, e.g. `my-archives-bucket` |
+| **Archive overflow public URL** | the public base from step 2, e.g. `https://archives.example.com` (no trailing slash needed) |
+
+Leave both blank to keep oversize archives dropped and shown as "Too large to host".
+On the next build and deploy, oversize archives upload to R2 and the Downloads page
+(and the header link, which reappears once anything is hostable) point at them.
+
+---
+
+## Securing downloads against cost-abuse
+
+The threat: someone scripts repeated downloads of large archives to run up your bill.
+
+**The reassuring part — R2 egress is free.** Unlike S3, Cloudflare R2 charges **$0**
+for bandwidth/egress. An attacker looping downloads of a 480 MB zip cannot run up a
+bandwidth bill. The *only* metered cost from reads is **Class B operations** (10M free
+per month, then $0.36/M) — and the defenses below make even that hard to reach.
+
+You get these controls one of two ways — **a Worker on `workers.dev`** (no domain) or
+**a custom domain**. Pick one. The Worker is recommended if you do not want to own a
+domain.
+
+### Option A — a Worker on workers.dev (no domain)
+
+The bundled **`r2-proxy/`** Worker is a small, self-contained project that binds the R2
+bucket and serves archive objects on a free `<name>.workers.dev` subdomain. It is a
+**pure passthrough**: the request path `<siteId>/archives/<file>.zip` maps straight to
+the bucket key, so **one Worker serves every site** (deploy it once, not per site).
+What it gives you, all on the free tier:
+
+- **Edge caching** — full downloads are cached with the Cache API, so repeat pulls skip
+ R2 (no billable Class B op). It honors the `Cache-Control` set on each object.
+- **Rate limiting** — Cloudflare's native, free rate-limit binding caps requests per
+ client IP + file (dashboard rate-limit rules need a paid zone; this does not).
+- **Path allow-listing** — it only serves `*/archives/*.zip`, never arbitrary keys.
+- **Range / resumable downloads** — honors `Range` requests so big zips can resume.
+
+**Deploy it (once for the whole instance):**
+
+1. Edit `r2-proxy/wrangler.toml` and set `bucket_name` to the **same bucket** you use in
+ the editor's Settings. Optionally rename the Worker (`name`) and tune the rate limit
+ (`limit` / `period`).
+2. From the repo root:
+ ```sh
+ cd r2-proxy
+ pnpm install # first time only
+ pnpm run deploy # = wrangler deploy, reusing your host wrangler auth
+ ```
+ wrangler prints the deployed URL, e.g. `https://ytdlp-archive-proxy.<you>.workers.dev`.
+ (An "unsafe fields are experimental" warning for the rate-limit binding is expected.)
+3. In the editor's **Settings**, set **Archive overflow public URL** to that
+ `workers.dev` URL. Re-deploy a site and its Downloads links resolve through the Worker.
+
+The Worker code is `r2-proxy/src/index.ts`; the caching and rate-limit logic are small
+and commented.
+
+> **Free-tier limit:** Workers Free allows **100,000 requests/day** (resets daily). Far
+> more than a downloads endpoint needs; if you ever exceed it, requests get a `429`
+> (fail closed — no surprise bill) until the next day, or upgrade to Workers Paid ($5/mo).
+
+The archive `Cache-Control` is set at upload time (`Cache-Control: public,
+max-age=3600`, constant `ARCHIVE_CACHE_CONTROL` in `common/publish/build.ts`); the
+Worker reads it back when caching. Archive filenames are stable and overwritten in
+place on re-deploy, so this 1-hour bound is what keeps a re-uploaded archive from being
+served stale for long — raise it if your archives rarely change.
+
+### Option B — a custom domain
+
+A custom domain has its own upsides — dashboard WAF, managed bot rules, and
+rate-limiting rules without touching code. Connect it per step 2 above and add these,
+in order of impact:
+
+1. **Edge caching.** Through a custom domain, Cloudflare's CDN caches each archive at
+ the edge (it honors the `Cache-Control: public, max-age=3600` set at upload), so
+ repeated downloads are served from cache and **never hit R2**. To make caching
+ aggressive, add a **Cache Rule** (**Caching → Cache Rules → Create**): *when* `URI
+ Path` contains `/archives/`, *then* Eligible for cache, **Edge TTL → Override → 1
+ day** (or longer). If you raise the TTL a lot, **purge the cache on deploy** so a
+ re-uploaded archive is not served stale.
+2. **Rate limiting — the hard backstop.** A Rate Limiting rule caps how fast any
+ single client can pull archives, stopping a flood that misses cache (**Security →
+ WAF → Rate limiting rules → Create**; the free plan includes one rule): *when*
+ `URI Path` contains `/archives/`, *rate* e.g. **20 requests per 1 minute** per
+ client IP, *then* Block for 10 minutes (or Managed Challenge). Tune the threshold
+ to real usage — legitimate users download a handful of files, not dozens per
+ minute.
+3. **Bot Fight Mode + WAF managed rules.** **Security → Bots → Bot Fight Mode**
+ (free) blocks the low-effort scripted abuse that makes up most of this traffic; the
+ free **WAF managed ruleset** adds a baseline at no cost.
+4. **Hotlink protection (optional).** Stops other sites embedding your archives and
+ spending your ops budget serving their audience. A WAF custom rule: *when* `URI
+ Path` contains `/archives/` **and** `Referer` does not contain your domain **and**
+ `Referer` is not empty, *then* Block. (Allow an empty `Referer` so direct clicks and
+ privacy-conscious browsers still work.)
+
+### Billing alerts, and what to skip
+
+R2 has no hard spend cap, but Cloudflare **Notifications** (**Notifications → Add**)
+can email you when R2 storage or Class A/B operations cross a threshold — cheap
+insurance so nothing surprises you.
+
+**Signed URLs / token-gated downloads** are the heavyweight option — they add key
+management and friction for legitimate users. Given egress is free and caching
+neutralizes the ops cost, they are overkill for *cost* defense (the `r2-proxy` Worker
+is a plain passthrough, not an access gate). Reach for them only if you want *access
+control* (private archives), not cost control.
+
+### Cost expectations
+
+Measured across all sites in this project, total compressed archives are **≈ 2–4.5 GB**.
+Against R2's free tier:
+
+| Resource | Free tier / month | This project's usage |
+|---|---|---|
+| **Egress / bandwidth** | unlimited, **$0** | irrelevant — no egress charge exists |
+| **Storage** | 10 GB-month | ~2–4.5 GB — comfortable |
+| **Class A ops** (writes) | 1,000,000 | ~one PUT per archive per deploy — negligible |
+| **Class B ops** (reads) | 10,000,000 | mostly absorbed by CDN cache |
+
+The realistic bill for hosting these archives is **$0**. The configuration above exists
+to keep it that way under adversarial traffic, not because normal usage is close to
+any limit.
+
+---
+
+## Building every site in containers
+
+> **Not about running the apps in containers** — that is
+> [RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md), built from the root `Dockerfile`. This
+> is about *building sites*: fanning per-site export builds out across containers,
+> using `Dockerfile.build`. The two share nothing but the word "docker".
+
+The docker build mode builds **every site in parallel** in isolated containers, then
+deploys them serially — a large speedup when you host several sites, and stronger
+isolation than the basic single-process build. It is opt-in: **Settings → Build
+pipeline → Docker**. Basic mode, the default, builds one site at a time in `export/`.
+
+**Prerequisites.**
+
+- A container engine: **Docker**, or **podman** (set `DOCKER_BIN=podman`). Rootless
+ podman is a good fit — it maps container files to your host user automatically.
+- The editor host still needs Node + pnpm (Phase A and the deploys run on the host)
+ and the Cloudflare/R2 credentials in the environment. Credentials are **never**
+ passed into a container — deploy runs on the host.
+- Inside the runtime container (`docker compose up`) there is no `docker` binary, and
+ the pipeline falls back to the serial host build.
+
+The build image is built (and cached) automatically from `Dockerfile.build` the first
+time you run; **Build image** / **Dockerfile** in Settings override the tag and path.
+
+**How it works.** Trigger it with **Build all sites** on /sites (or `pnpm archilyzer
+build all`, which uses docker when `docker version` answers). One job runs three
+ordered phases:
+
+1. **Phase A — shared, on the host, once.** The data phase (the search index and the
+ `.export-index` staging), then `archilyzer build archives` (the shared archive-zip
+ cache for the union of all sites' channels). Only the host writes this shared
+ state, so containers never race it. This phase is serial and is the long pole on a
+ cold build; on a warm rebuild it is near-instant (unchanged channels are skipped).
+2. **Phase B — per-site, in parallel containers.** Each site's compose + `next build`
+ runs in its own container (`docker/build-site.sh`: `archilyzer build site <id>
+ --nodata`), capped by **Max parallel builds**. Each writes an isolated `out/` under
+ `export/.export-builds/<siteId>/`. Containers mount the corpus, index, staging and
+ archive cache **read-only** (`ARCHIVES_READONLY=1`). Network is left on: `next
+ build` fetches the site's fonts through `next/font/google`; isolation comes from the
+ read-only mounts, the per-site output dir and the non-root user.
+3. **Phase C — deploy, on the host, serially.** After every build finishes, each built
+ site is deployed in turn (the R2 upload, then `wrangler pages deploy`). A single
+ site failing to build or deploy is reported and skipped; the rest still ship.
+
+If no container engine is available, the action logs a notice and falls back to a
+serial host build+deploy (one site at a time).
+
+**Mounts, per Phase-B container.**
+
+| Host | Container | Mode |
+|---|---|---|
+| `transcripts/` (corpus + `index.mdb` + archive cache) | `/data/transcripts` | ro |
+| `export/.export-index` (shared + per-site staging) | `/data/export/.export-index` | ro |
+| `export/.export-builds/<siteId>` (public/out/.next/caches) | `/site` | rw |
+| `settings.json` (build config, mounted fresh — not baked) | `/data/settings.json` | ro |
+
+The per-site `/site` mount is persistent, so incremental `next build` (`.next`) and
+incremental compose (`.compose-cache`) stay warm across builds.
+
+**Tuning.**
+
+- **Max parallel builds** (setting) — how many site containers run at once. Each `next
+ build` can use up to ~8 GB; a safe starting point is `floor(RAM_GB / 9)`.
+- `DOCKER_BIN` — the container binary (default `docker`; e.g. `podman`).
+- `DOCKER_BUILD_MEMORY`, `DOCKER_BUILD_CPUS` — optional per-container `--memory` /
+ `--cpus` caps so a fan-out cannot OOM or peg the host.
+- `BUILD_ARCHIVES=0` (or **Skip archive zips**) — skip the archive warm and the
+ per-site archive materialize for a faster build with no download bundles.
+
+**Notes.** Containers run as your host uid/gid (`-u`), so files under
+`.export-builds/` are host-owned, not root-owned. The image bakes the repo source and
+deps; a code change rebuilds it, but layer caching keeps that cheap (deps re-install
+only when the lockfile moves). `.export-builds/` is gitignored and excluded from the
+image build context. The editor mounts the host's `docker/build-site.sh` over the
+baked one, so an image older than the checkout still runs today's script.
+
+---
+
+## Reading a published archive from Claude Code
+
+Every published archive serves a machine contract (`/corpus.json`, `/llms.txt`), and
+the MCP server reads it over HTTP. Register it against your public URL — as
+`archilyzer`, which the tracked `/ask` and `/sweep` commands expect:
+
+```sh
+claude mcp add archilyzer \
+ --env TRANSCRIPT_SITE_URL=https://<your-site>.pages.dev \
+ -- pnpm -C "$PWD" archilyzer mcp
+```
+
+`archilyzer mcp` starts the same server as `pnpm --filter yt-dlp-transcript-mcp exec
+tsx src/index.ts` (the form [mcp/README.md](mcp/README.md) uses), with the
+environment passed through and nothing on stdout but the protocol.
diff --git a/README.md b/README.md
@@ -329,6 +329,7 @@ claude # then try: /ask what has he said about magic tournaments?
The two editor lines are optional: they let `fetch_clip` ask a local editor for clip
media (`WORKER_TOKEN` is the editor's own). Leave them out for research alone.
+`-- pnpm -C "$PWD" archilyzer mcp` starts the same server through the repo's CLI.
> **Register the server as `archilyzer`.** The shipped commands call
> `mcp__archilyzer__ask_plan` / `mcp__archilyzer__sweep_plan`, and that tool name
@@ -449,34 +450,25 @@ Two more pieces round out the workspace: the **project site** (`homepage/`) and
Transcripts and per-channel state live at `<repo>/transcripts/` — **its own git repo**,
untouched by the workspace.
-Everything resolves through `getPaths()` (`common/lib/paths.ts`) and can be overridden
-by environment variables:
-
-| Variable | Default | Purpose |
-| --- | --- | --- |
-| `TRANSCRIPTS_DIR` | `<repo>/transcripts` | The corpus: channels, media, index, job logs. |
-| `SAVED_VIDEOS_DIR` | inside `TRANSCRIPTS_DIR` | Persisted source-video store; can live on another disk. |
-| `SITES_DIR` | inside `TRANSCRIPTS_DIR` | Per-site configuration. |
-| `EXPORT_PUBLIC_DIR` | `<repo>/export/public` | Where the index writes paginated JSON. |
-| `SETTINGS_FILE` | `<repo>/settings.json` | Operational settings — every key, default and meaning is in [SETTINGS.md](SETTINGS.md). |
-| `YTDLP_BIN` | `yt-dlp` on PATH | The downloader. |
-| `WHISPER_BIN` / `WHISPER_MODEL` | `whisper-cli` on PATH | Default transcription backend and its model. |
-| `FFMPEG_BIN` / `FFPROBE_BIN` | on PATH | Transcode and duration checks. |
+Every path and binary resolves through `getPaths()` (`common/lib/paths.ts`), and each
+can be overridden by an environment variable — `TRANSCRIPTS_DIR` moves the whole corpus,
+`YTDLP_BIN` / `FFMPEG_BIN` / `WHISPER_BIN` name the tools. The full list, with every
+other variable the code reads, is **[ENVIRONMENT.md](ENVIRONMENT.md)** (generated from
+a list the tests hold to the code). `pnpm archilyzer doctor` prints which overrides are set
+and whether every tool this machine is configured to use is there.
Everything else lives in the editor's **Settings** page and is optional — a missing or
-partial settings file falls back to defaults. Full list in
-[SETUP.md](SETUP.md#configuration--environment-variables).
+partial settings file falls back to defaults. Every key is in [SETTINGS.md](SETTINGS.md).
## Publishing
The static export can be served by anything. The path with the most support is
Cloudflare Pages, with download archives too large for Pages' 25 MB per-file limit
-overflowing to R2 — see **[DEPLOY_CLOUDFLARE.md](DEPLOY_CLOUDFLARE.md)**, which also
-covers the configuration that keeps public archive downloads from being abused to run
-up costs.
-
-If you host several sites from one corpus, the opt-in Docker build pipeline builds them
-all in parallel in isolated containers — see **[DEPLOY_DOCKER.md](DEPLOY_DOCKER.md)**.
+overflowing to R2. **[PUBLISH.md](PUBLISH.md)** covers building and deploying a site,
+the hub and the homepage (from the editor, `pnpm ops` or `pnpm archilyzer`), previews,
+the R2 setup, the configuration that keeps public archive downloads from being abused
+to run up costs, and the opt-in pipeline that builds several sites in parallel in
+isolated containers.
Channels can sync automatically on a per-channel cadence via a cron heartbeat — see
**[SCHEDULED_SYNC.md](SCHEDULED_SYNC.md)**.
@@ -499,11 +491,11 @@ See [CONTRIBUTING.md](CONTRIBUTING.md) to work on the code.
| Document | Covers |
|---|---|
-| [SETUP.md](SETUP.md) | Full per-OS install, every environment variable, transcription backends. |
-| [CONTRIBUTING.md](CONTRIBUTING.md) | Workspace layout, tests, CLI shims, internals. |
+| [SETUP.md](SETUP.md) | Full per-OS install, transcription backends. |
+| [ENVIRONMENT.md](ENVIRONMENT.md) | Every environment variable, by audience (generated). |
+| [CONTRIBUTING.md](CONTRIBUTING.md) | Workspace layout, tests, the `archilyzer` CLI, internals. |
| [RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md) | `docker compose up` for the whole stack: exposure model, auth, GPU. |
-| [DEPLOY_CLOUDFLARE.md](DEPLOY_CLOUDFLARE.md) | Pages + R2, and cost-abuse protection. |
-| [DEPLOY_DOCKER.md](DEPLOY_DOCKER.md) | Parallel multi-site export builds in containers. |
+| [PUBLISH.md](PUBLISH.md) | Building and deploying sites: Pages + R2, cost-abuse protection, parallel builds in containers. |
| [SCHEDULED_SYNC.md](SCHEDULED_SYNC.md) | Unattended per-channel syncing. |
| [WORKTREES.md](WORKTREES.md) | Parallel checkouts and the port scheme. |
| [mcp/README.md](mcp/README.md) | The MCP server: tools, links, client setup. |
diff --git a/RUNNING_IN_DOCKER.md b/RUNNING_IN_DOCKER.md
@@ -1,13 +1,14 @@
# Running an archive in Docker
One command stands up a working archive server: the editor, the tools it drives
-(yt-dlp, ffmpeg, whisper.cpp), and a reverse proxy that is the only thing on the
-box with an open port.
+(yt-dlp, ffmpeg, and a transcription engine — whisper.cpp, or parakeet.cpp in the
+Vulkan image), and a reverse proxy that is the only thing on the box with an open
+port.
-> This document is about **running the apps**. [DEPLOY_DOCKER.md](DEPLOY_DOCKER.md)
-> is a different thing entirely — it is about fanning per-site *export builds* out
-> across containers, and it uses `Dockerfile.build`. Neither file affects the
-> other.
+> This document is about **running the apps**. Fanning per-site *export builds* out
+> across containers is a different thing entirely — it uses `Dockerfile.build`, and
+> it is in [PUBLISH.md](PUBLISH.md#building-every-site-in-containers). Neither
+> affects the other.
---
@@ -538,7 +539,7 @@ silently loses formats), `ffmpeg`/`ffprobe`, `whisper-cli` (statically linked),
### The multi-site build pipeline falls back inside a container
The editor can fan per-site export builds out across containers
-(`buildPipeline.mode = "docker"`, see DEPLOY_DOCKER.md). Inside a container there
+(`buildPipeline.mode = "docker"`, see [PUBLISH.md](PUBLISH.md#building-every-site-in-containers)). Inside a container there
is no `docker` binary, so that path is unavailable. It already handles this — the
build logs
diff --git a/SETUP.md b/SETUP.md
@@ -46,7 +46,7 @@ homepage):
| **ffmpeg** + **ffprobe** | Audio transcode + duration checks for `transcribe` channels. | `ffmpeg` / `ffprobe` on `PATH` |
| A **transcription backend** | `handling: "transcribe"` channels only. Default is **whisper.cpp** (`whisper-cli`); `chough` and `parakeet.cpp` are alternatives. | `whisper-cli` on `PATH` |
| **rsync** | Backing up the saved-video store. | `rsync` on `PATH` |
-| **Docker** | Running the whole stack in containers ([RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md)), the parallel multi-site build ([DEPLOY_DOCKER.md](DEPLOY_DOCKER.md)), and the sharded e2e run (`pnpm e2e:sharded`). | — |
+| **Docker** | Running the whole stack in containers ([RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md)), the parallel multi-site build ([PUBLISH.md](PUBLISH.md#building-every-site-in-containers)), and the sharded e2e run (`pnpm e2e:sharded`). | — |
Every binary above is overridable by an environment variable (e.g. `YTDLP_BIN`) —
see [Configuration & environment variables](#configuration--environment-variables).
@@ -260,28 +260,27 @@ starting template. Both are generated from the settings schema
[CHANNEL.md](CHANNEL.md). Settings are optional — a missing/partial `settings.json` falls
back to built-in defaults, so the app runs out of the box.
-Paths and binaries resolve through `getPaths()` in `common/lib/paths.ts`. Override
-any of them via environment variables before launching:
+Paths and binaries resolve through `getPaths()` in `common/lib/paths.ts`, and each is
+overridden by an environment variable before launching — `TRANSCRIPTS_DIR` (the corpus,
+default `<repo>/transcripts`), `SETTINGS_FILE`, `YTDLP_BIN`, `FFMPEG_BIN` /
+`FFPROBE_BIN`, `WHISPER_BIN` / `WHISPER_MODEL`, `PARAKEET_CLI` / `PARAKEET_MODEL` and
+the rest. **Every variable the code reads is in [ENVIRONMENT.md](ENVIRONMENT.md)**, by
+audience: the path overrides, the runtime tokens and knobs (`WORKER_TOKEN`,
+`SYNC_TICK_URL`, the R2 credentials, …), the ports, the docker `ARCHILYZER_*` set and
+the test-only ones. It is generated from `common/lib/envVars.ts`, and a test fails when
+the code reads a variable that list does not declare.
-| Variable | Default | Purpose |
-| --- | --- | --- |
-| `TRANSCRIPTS_DIR` | `<repo>/transcripts` | Channels, archives, LMDB index, job logs. |
-| `SAVED_VIDEOS_DIR` | `<TRANSCRIPTS_DIR>/saved-videos` | Persisted source-video store (can live on a separate disk). |
-| `SITES_DIR` | `<TRANSCRIPTS_DIR>/sites` | Per-site config (`sites/<id>/site.json` — every key in [SITE.md](SITE.md)). |
-| `EXPORT_PUBLIC_DIR` | `<repo>/export/public` | Where the index writes paginated JSON. |
-| `SETTINGS_FILE` | `<repo>/settings.json` | Site-settings file. |
-| `YTDLP_BIN` | `yt-dlp` (PATH) | Pipeline downloader. |
-| `WHISPER_BIN` | `whisper-cli` (PATH) | whisper.cpp binary. |
-| `WHISPER_MODEL` | `~/whispercpp/whisper.cpp/models/ggml-base.en.bin` | whisper.cpp model file. |
-| `CHOUGH_BIN` / `CHOUGH_URL` / `CHOUGH_MODEL` | `chough` / — / — | chough backend binary, remote server, model. |
-| `PARAKEET_CLI` / `PARAKEET_MODEL` / `PARAKEET_STITCH_BIN` | `parakeet-cli` / — / `scripts/parakeet-stitch.mjs` | parakeet.cpp CLI, model, and wrapper. |
-| `FFMPEG_BIN` / `FFPROBE_BIN` | `ffmpeg` / `ffprobe` (PATH) | Audio transcode + duration checks. |
-| `RSYNC_BIN` | `rsync` (PATH) | Saved-video backup. |
-| `WORKER_TOKEN` | — | Bearer token for the remote-worker transcription API (set on both ends when used), and for the `/api/ops/*` HTTP layer over the editor's actions — see [RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md#driving-the-editor-without-a-browser) and `pnpm ops`. Unset means both surfaces are off. |
-
-Feature-area docs cover their own env vars: [SCHEDULED_SYNC.md](SCHEDULED_SYNC.md)
-(`SYNC_HEARTBEAT_SECONDS`, `SYNC_TICK_URL`, `SYNC_TICK_TOKEN`) and
-[DEPLOY_CLOUDFLARE.md](DEPLOY_CLOUDFLARE.md) (R2 credentials).
+To see what this machine has, run:
+
+```sh
+pnpm archilyzer doctor
+```
+
+It is read-only: the checkout, the corpus and each channel's media, `settings.json`,
+every binary the paths name plus each enabled worker's engine and model, umtool's
+report-pipeline tools, and this checkout's port block. It exits 1 only for something
+the machine is configured to do and cannot (an enabled worker's engine missing beside
+a corpus, a settings file that does not parse, an override naming a missing binary).
---
@@ -293,7 +292,7 @@ but you do need the Playwright browser:
```sh
npx playwright install chromium # one-time: download the test browser
-pnpm e2e # sequential run on the host (next dev, port 3001)
+pnpm e2e # sequential run on the host (next dev, port 3011)
```
For a faster parallel run, `pnpm e2e:sharded` splits the suite across N Docker
@@ -308,11 +307,12 @@ collisions, see [WORKTREES.md](WORKTREES.md).
## Where to go next
-- [README.md](README.md) — project overview, pipeline modes, CLI shims.
+- [README.md](README.md) — project overview, pipeline modes.
- [SCHEDULED_SYNC.md](SCHEDULED_SYNC.md) — automatic per-channel sync (internal
heartbeat or external cron).
-- [DEPLOY_CLOUDFLARE.md](DEPLOY_CLOUDFLARE.md) — deploying to Cloudflare Pages + R2
- archive overflow.
+- [PUBLISH.md](PUBLISH.md) — building and deploying sites: Cloudflare Pages, R2
+ archive overflow, parallel builds in containers.
+- [ENVIRONMENT.md](ENVIRONMENT.md) — every environment variable, by audience.
- [WORKTREES.md](WORKTREES.md) — parallel development with per-worktree ports.
- [mcp/README.md](mcp/README.md) — MCP server exposing the archive to Claude Code /
Desktop / Cursor.
diff --git a/homepage/content/README.md b/homepage/content/README.md
@@ -6,7 +6,7 @@ exists for whoever edits the docs next.
## Why these are hand-written rather than rendered from the root docs
-The obvious move is to render `SETUP.md`, `DEPLOY_CLOUDFLARE.md` and friends
+The obvious move is to render `SETUP.md`, `PUBLISH.md` and friends
directly, and keep one copy. That was rejected for a decisive reason:
> `SETUP.md` says `git clone <this-repo-url>`. **There is no public repository.**
@@ -31,8 +31,8 @@ its public counterpart needs the same change.
| `docs/what-is-archilyzer.md` | `README.md` | package list, pipeline modes |
| `docs/install.md` | `SETUP.md` | tool versions, env-var table, backend list, the Windows path |
| `docs/operate.md` | `README.md`, `SCHEDULED_SYNC.md` | editor routes, scheduler settings |
-| `docs/deploy-cloudflare.md` | `DEPLOY_CLOUDFLARE.md` | the 25 MB Pages limit, R2 options |
-| `docs/deploy-docker.md` | `DEPLOY_DOCKER.md` | phase structure, settings names |
+| `docs/deploy-cloudflare.md` | `PUBLISH.md` (Cloudflare, R2, cost-abuse) | the 25 MB Pages limit, R2 options |
+| `docs/deploy-docker.md` | `PUBLISH.md` (building every site in containers) | phase structure, settings names |
| `docs/ai-and-mcp.md` | `mcp/README.md` | tool names, `corpus.json` shape |
| `docs/faq.md` | — (written for this site) | claims about cost and hardware |
diff --git a/r2-proxy/README.md b/r2-proxy/README.md
@@ -22,7 +22,7 @@ Then set the editor's **Settings → Archive overflow public URL** to the deploy
`https://<name>.<you>.workers.dev` URL.
Full setup and the alternative custom-domain path are documented in
-[../DEPLOY_CLOUDFLARE.md](../DEPLOY_CLOUDFLARE.md).
+[../PUBLISH.md](../PUBLISH.md#securing-downloads-against-cost-abuse).
## Scripts
diff --git a/r2-proxy/src/index.ts b/r2-proxy/src/index.ts
@@ -7,7 +7,7 @@
// request path straight to the bucket key — there is nothing per-site about it.
// Deploy it ONCE to a free `<name>.workers.dev` subdomain (no custom domain, no
// domain purchase, no WHOIS), then point every site's "Archive overflow public
-// URL" at that single subdomain. See ../DEPLOY_CLOUDFLARE.md.
+// URL" at that single subdomain. See ../PUBLISH.md.
//
// Why a Worker instead of the raw r2.dev URL: it gives us tunable, in-code rate
// limiting (the cost/abuse backstop — Cloudflare's dashboard rate-limit rules
diff --git a/r2-proxy/wrangler.toml b/r2-proxy/wrangler.toml
@@ -4,7 +4,7 @@
# cd r2-proxy && pnpm dlx wrangler deploy
#
# One bucket + one Worker serves every export site (keys are namespaced by site
-# id). See ../DEPLOY_CLOUDFLARE.md for the full setup.
+# id). See ../PUBLISH.md for the full setup.
name = "archilyzer-exports"
main = "src/index.ts"