# Publishing What gets published, how the publish stages build and deploy it, and how to keep a public archive cheap and safe. For *running* the apps in containers see [RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md); every environment variable named here is in [ENVIRONMENT.md](ENVIRONMENT.md), every key of `settings.json` and `site.json` in [SETTINGS.md](SETTINGS.md) and [SITE.md](SITE.md). - [What gets published](#what-gets-published) - [Publishing is stages](#publishing-is-stages) - [The three ways to drive it](#the-three-ways-to-drive-it) - [The publish lane](#the-publish-lane) - [Cloudflare Pages](#cloudflare-pages) - [The source mirror (homepage)](#the-source-mirror-homepage) - [Preview deployments](#preview-deployments) - [Download archives and R2](#download-archives-and-r2) - [Securing downloads against cost-abuse](#securing-downloads-against-cost-abuse) - [Building every site in containers](#building-every-site-in-containers) - [Reading a published archive from Claude Code](#reading-a-published-archive-from-claude-code) --- ## What gets published Three static artefacts, each pre-rendered to plain HTML and JSON — no database, no server-side code, servable by anything: | Artefact | What it is | Its bundle | Config | |---|---|---|---| | A **site** | The export app over the channels one site selects: search, transcripts, charts, downloads. One corpus can publish several. | `export/.export-builds//out` | `transcripts/sites//site.json` ([SITE.md](SITE.md)) | | The **hub** | The export app in hub mode: federated search across the family of sites. | `export/.export-builds/_hub/out` | `transcripts/sites/_homepage/homepage.json` | | The **homepage** | The project's own site (`homepage/`), with the source mirror, its raw tree, its history pages and the source tarball. | `homepage/out` (its stamps in `export/.export-builds/_homepage/`) | `~/.config/archilyzer/` (the source mirror's two operator files) | **Each target keeps its own bundle, and a deploy ships that bundle and nothing else.** `export/out` is a link to the bundle built last, so `serve out` and anything else that read it keep working. The hub and the homepage are two different Pages projects: the hub deploys to the project `homepage.json` names (e.g. `archilyzer-hub`) and is refused `archilyzer`, which is the homepage's. **Every target is built from ONE index.** The update-index stage (`archilyzer publish index`) runs the data phase once — the LMDB index, the stats datasets and the chart templates — and writes the index stamp. A site's build is then **compose** (the site's slice of the shared index into `export/public`, plus its download archives — `archilyzer compose site `) and `next build`; it has no data phase of its own. A transcript record in the shared pages carries its other English caption tracks (`altTracks`, with the transcript's own `track`) only where one's words differ from the transcript's — the served `en` beside `en-orig`, a regional or auto-translated track, the captions a local transcription replaced (`common/lib/captionTracks.ts`). Identical tracks add nothing, so most records' bytes are what they were. Search reads those tracks with the transcript and names the track of a hit only one holds; the reader switches to them. English VTTs are not published as subtitle tracks. The first index build after this reads them once (`Alternate tracks v1: N record(s) re-read.`), for exactly the records that can hold one. ### Reports and cited sites A site with `reports` (site.json) has two more steps, before its build and on the host. First, **fetch the missing evidence**: *Fetch missing evidence* on the Reports tab, or `pnpm ops fetch-windows --json '{"siteId":""}' --wait`, fetches the window of every cited span whose media is not on disk, through the editor's managed path, as one paced job per platform (YouTube and Rumble side by side; at least 30 s between two windows on one platform, a 429 or two 403s in a row backing the platform off and stopping its job). `"dryRun": true` (*Preview missing evidence*) lists the windows, and the spans no window can fill — a video the source says is deleted or private, a channel off the site, a drive that is not mounted — without fetching. Running it again resumes: fetched windows are on disk. Then **reports prepare** cuts the evidence clip of every span its published reports cite — from the editor's clip windows, the saved-video store or a record's audio, fitted inside 1280×720 (H.264 crf 23, AAC; an audio span is an `.m4a`) — and copies the screenshot and media of every cited post, only those, into `.export-index/sites//report-media/`, with a manifest (`index.json`) of each moment's file, size, hash and duration. Prepare fetches nothing: a citation whose media is not on disk, a clip over 24 MiB or an invalid report is listed and fails the run (exit 1, or a failed job) — fetch the missing evidence or persist the video, capture the post, and run it again; what is already cut is reused. When nothing is missing, prepare ends by **exporting** the reports (`reports export` runs the same step alone): each published report is written, as compose would resolve it, to `.export-index/sites//report-exports//` as `report.html` (one self-contained file — inline style, no script, its stills and post screenshots inlined; clips linked on the site), `report.pdf` (that page printed by headless Chromium; skipped with a note on a host without Playwright's browser), `report.md` (plain Markdown with numbered references), `slides.html` (the report as slides — the same slides the page's Slides view shows, `common/lib/report/slides.ts` — one self-contained file: the site's slide stylesheet in its accent, every picture inlined, a few lines of script for the keys), `slides.pdf` (those slides printed, one 1280 × 720 page per slide; skipped like `report.pdf`; neither slides file for a report whose `slides.hide` says so) and `evidence-pack.zip` (the page and the slides with its clips, stills and screenshots as files, so the clips play offline, plus the Markdown and the citations; packed by the system `zip`, which a host must have), with an `export.json` naming each file's size and checksum and the hash of the report.json it was made from. Every export ends with a footer naming the report's revision, its date and the first 12 characters of that hash. Each export also keeps the report's **revision history** (`common/publish/reportHistory.ts`): a bare git repository per report at `sites//reports//history-git/` (add `history-git/` to the corpus repository's `.gitignore`). When the report.json's sha256 differs from the newest revision's, the export commits a new revision holding `report.json` (byte for byte), `report.md` and `exports.json` (the sha256 of every export file and of the citations JSON and CSV), with the message `Revision N` and a summary of what changed. A re-export of an unchanged report commits nothing. Author and committer are the site (its title, `noreply@.invalid`), dated in UTC; git runs with no system or global config and none of the operator's environment. A footer names the revision and the report's sha256, never a commit (the commit holds the footer); the history page maps the sha256 to its commit, and `export.json` records it. The build's compose then writes the reports from what prepare left (`common/publish/composeReports.ts`): each report's page, its citations as `citations.json` and `citations.csv`, its exports (`report.html`, `report.pdf`, `report.md`, `slides.html`, `slides.pdf`, `evidence-pack.zip` — only those made from the report.json as it is now, and only a file of at most 24 MiB: a larger pack stays on the host; the page's download line lists what was published), its cited stills (never a saved source copy), one page per cited moment with the transcript lines around it, the prepared clips and captures — only those the published reports cite — and each report's revision history: `reports//history/history.json` (every revision's number, date, commit, report sha256, change summary and claim-level word diff, which `/reports//history/` renders) and a dumb-HTTP clone at `reports//history/repo/`, staged as the source mirror is (a fresh bare clone, `repack -a -d`, `update-server-info`, then only `HEAD`, `packed-refs`, `info/refs`, `objects/info/packs`, the packs and `refs/heads/main`), so `git clone /reports//history/repo` works. The deploy audit refuses any other file in a clone or one over 24 MiB, and counts the clone's files toward Pages' 20,000. Every quote is checked as it is composed: a span's against its cues within 5 s either side (read under the caption-track rule — `en-orig` first, then the first English track with cues — and against every other English track, the best match shown), a post's against its text, scored as the share of the quote's words found there; the score, the time and the method are written into the citation, replacing any typed by hand. Compose fails, before it writes any of it, with the list of everything wrong: an invalid report, a citation of a channel outside the site or a post the site may not carry, a missing record, still or post, a quote below 60 %, and a citation whose media was not prepared, or was cut for a span the report no longer cites. `--allow-missing-media` (on `compose site`, `publish build` and `build site`) lets the last two through; their pages render without a clip. A **cited** (report-only) site is one with search off (`site.json` `search: false`; a legacy `publish: "cited"` reads the same): it publishes its reports and nothing else. Its compose removes every corpus-shaped file the shared `export/public` holds — summaries, stats, transcripts, subs, posts, digests, archives, the duplicates, tags, aliases and chart files, the service worker — writes `site.json` with no channels and `corpus.json` with `site.scope: "cited"` (spec 5), and an `llms.txt` and sitemap listing the reports and moment pages. After `next build`, the built `out/` is audited: a cited build that holds anything but `_next/`, the reports, the moment pages, the cited media, the shell's own pages and files, or a file over 25 MiB, or more than 20,000 files, fails the build, and every deploy refuses it. A site switched to search off is refused at deploy until it is built again. With no published reports it is an empty report index. A cited site is not a searchable archive, so the family never lists it, whatever `listed` says (`isListedSite`): it is no hub member, has no homepage card, counts in no family total and no other site's footer links to it. A full site with reports is listed as `listed` says. Readers of the contract see a cited site as an empty corpus with reports: the MCP server's `list_reports` / `get_report` read them (and its `list_channels` says "cited-only site: N report(s)"), and umtool's cue walk refuses one by name — it has no transcripts to cut a report video from. `pnpm --filter export run e2e:report` builds a cited fixture site for real (one fact-check citing every kind, its moments and their media) in a stage project at the repo root, `.e2e-report-site/` — never in `export/`, whose `public/` and `out/` are the checkout's — audits its `out/` as above and reads it in a browser. `pnpm --filter export run build:report-fixture` builds and audits it without the browser. The suite, and export's `start` (`pnpm start:export`), serve the build with `export/scripts/serve-out.mjs`, not `serve`: `serve` (still the container's `site` service) takes a video or audio moment's directory name, `-` in seconds with decimals, for a file with an extension, and lists the directory instead of serving its page. --- ## Publishing is stages Publishing is seven queueable **stages** driven by what is on disk (`common/publish/stages.ts`). Each asks a pure `needs()` — is its target **fresh**, **stale** (and why) or **blocked** — runs only when stale or forced, and writes the stamp the next stage reads. One index update serves every build; builds and deploys run one at a time. Paths below are under `export/.export-builds/` unless named. | Stage (job `publish-`) | Target | Fresh when | Writes | |---|---|---|---| | `update-index` | `_index` | a stamp exists and nothing it reads is newer (below) | `export/.export-index/stamp.json` | | `build-site` | ``, or `_all` (each stale site) | the bundle's `inputSig` is the stamp's, no member channel and no config of the site changed since it was built or last checked, and the bundle passes `builtBundleProblem` | `/built.json` | | `deploy-site` | `` | the slot it deploys to — `production`, `local` or `previews.` — already holds this build | `/deployed.json` | | `build-hub` | `_hub` | the bundle's `inputSig` is the stamp's `hubSig` and nothing of a listed site changed | `_hub/built.json` | | `deploy-hub` | `_hub` | as `deploy-site` (the hub has no local target) | `_hub/deployed.json` | | `build-homepage` | `_homepage` | built from the current stamp, and its source is of `main`'s head (when a repository answers) | `_homepage/built.json` | | `deploy-homepage` | `_homepage` | as `deploy-site` | `_homepage/deployed.json` | - **update-index** runs `buildIndex`, `buildStats` and the chart templates in ONE child (8 GB heap), then signs every site and the hub. It runs whenever asked; a rerun that finds nothing short-circuits and **keeps its stamp id** ("nothing changed — the stamp stands"), so the bundles built from it stay current. - **A build** composes into the one shared `export/public`, runs `next build` into `export/out`, audits it and installs it as `/out` (a swap through `out.next`; across filesystems, as in the container, a copy), moves the archive staging beside it and points `export/out` at it. The homepage's runs the source mirror into `homepage/out`. - **A deploy** ships the target's bundle ([Cloudflare Pages](#cloudflare-pages)), or with `--to local` copies it into `ARCHILYZER_SITE_OUT` / `ARCHILYZER_HOMEPAGE_OUT`. A stage whose input is not there is **blocked** and exits 3 with a sentence that says so: "update the index first"; "no build of jeralyzer in export/.export-builds/jeralyzer — archilyzer publish build jeralyzer"; a private site, no Pages project, a bundle that fails its guards, a production deploy of a build not made on `main`. `--force` runs a fresh stage; Publish now and the lane never force. A build by older code shows a "code newer" chip and is not stale. ### The stamps Written atomically, each only by its stage. A missing or malformed file is no stamp, and no stamp is stale. - **`stamp.json`** (the index): its id, the LMDB `generation`, `scannedAt`, `builtAt` and the commit; per site an **`inputSig`** over everything the site's compose and export build read (its `.export-index` tree, each published member's shared trees, its `site.json`, the curated tags, aliases and duplicates files, the archive and X visibility settings, the social links, the hub URL, its sibling sites) — fresh means the build would produce the same bundle; the **`hubSig`**; the settings signature. A record added anywhere bumps `generation` and moves every site's signature: more rebuilds, never a wrong skip. - **`/built.json`**: its id, the index stamp it came from, `inputSig`, `builtAt`, `checkedAt` (a later no-op build found it current), the **commit and branch** (`ARCHILYZER_COMMIT` / `ARCHILYZER_BRANCH` when set — the image bakes them — else git; a detached HEAD records none), the runner, the bundle's `generatedAt`, its size and the archives staged for R2. - **`/deployed.json`**: one record per slot — the build's stamp id, when, the deployment URL, the preview alias, the wrangler version and the live check. **What makes the index stale.** The signals decide *when* to run; the index child decides *what* changed. (1) An ingest job — a download, transcription, digest, normalize, import, post fetch, tag write — ended `done` after the stamp's `scannedAt`, or a channel's report regenerated since; (2) a config file newer than the stamp: `tags.json`, `search-aliases.json`, `duplicates*.json`, any `site.json`, `homepage.json`, the charts config; (3) **the settings the index reads, by a signature of those keys, not `settings.json`'s mtime** (which every pause click moves): `socialLinks`, `homepageUrl`, `buildArchives`, `archiveStorage`, `social.x.visibility`, `maxTranscriptPageBytes`, the storage locations' roots. The same channel read marks a **site** stale before any index runs — "stale: 3 channels changed (a, b, c)" — and the signature confirms it after ("stale: data changed"). ### The publish lock **One stage at a time on a machine, whoever started it.** Every stage — a child the editor spawned, or `archilyzer publish …` in a terminal — takes `export/.export-builds/.publish.lock` (`{pid, host, kind, target, since}`, created exclusively). A second one **waits**, saying once whom for ("[publish] waiting for the publish lock — held by build-site jeralyzer (pid 4242 on , since …)"); Cancel or Ctrl-C gives up the wait. The host is `ARCHILYZER_HOST_ID`, else the hostname. A lock left by a process that is gone is **taken over** by the next stage: on this host, a pid that does not answer, answers with another start time, or started after the lock was taken (where there is no `/proc`, pid-alive alone); a file that does not parse for 60 s. **A lock naming another host is never taken over**; its wait line names both hosts. `archilyzer doctor`'s `publish-lock` row says free, held, stale, torn or another host's, and prints the `rm` that clears it — it never removes one. Remove it by hand only when nothing is publishing. ### Exit codes, Cancel `0` ran, or nothing to do · `1` failed (a build, R2, wrangler, a credential refusal) · `2` usage (a bad flag, an unknown site, a bad preview name) · `3` precondition not met (blocked, above; the docker runner with no engine) · `130` cancelled. **`pnpm archilyzer` reports 2, 3 and 130 as 1**: it is `pnpm --filter … exec`. The editor spawns `tsx bin/archilyzer.ts` from `common/` directly and sees the real code. `publish build all` and `publish deploy all` try every site and exit 1 when any failed; `publish now` exits with the worst code (1 over 3). **Cancel takes the stage's whole tree** — `next build`'s workers, wrangler, docker; in a terminal a second Ctrl-C kills it at once. --- ## The three ways to drive it The editor's **/sites** page, `pnpm ops` (HTTP to a running editor, with its `WORKER_TOKEN`) and the `archilyzer` CLI (no editor needed) run the same stages: the first two as jobs on the editor's `publish` queue, the CLI in its own process, under the same lock. | To | Editor | `pnpm ops publish`, `{"verb": …}` | `pnpm archilyzer …` | |---|---|---|---| | Update the index | /sites → Pool → **Build index** | `"index"` | `publish index` | | Build a site | its /sites row → **Build**; its **Publish** tab → *Build static export* | `"build"` + `"siteId"`/`"siteIds"` | `publish build [--force] [--skip-archives]` | | Build every stale site | /sites → **Build all stale** | `"stale"` | `publish build all [--runner docker]` | | Deploy a built site | its row → **Deploy preview** / **Deploy production** / **Deploy local**; its Publish tab | `"deploy"` + `"preview"`, `"to": "local"`, `"force"` | `publish deploy [--preview ] [--to local] [--force]` | | Build, then deploy | its Publish tab → **Build & deploy** | the `build-deploy` alias | `publish build `, then `publish deploy ` | | The hub | the Hub row → **Build hub** (tick *Deploy after build*), **Deploy hub** | `"hub"` + `"deploy": true`, `"preview"` | `publish hub [--deploy \| --deploy-only] [--preview ]` | | The homepage | the Homepage row → **Build homepage**, **Deploy homepage**, **Deploy local** | `"homepage"` + `"deploy"`, `"preview"`, `"to"` | `publish homepage [--deploy \| --deploy-only] [--preview ] [--to local]` | | What the lane would run | /sites → **Publish now** | `"now"` | `publish now` | | The status | the /sites Publish panel; /operations/publish | `pnpm ops get publish` | `publish status [--json]` | | The source mirror alone | — (every homepage build runs it) | — | `source publish [--force] [--check] [--keep-scratch]`, `source audit []` | | Fetch the clip windows a site's reports cite and the disk lacks | its **Reports** tab → *Fetch missing evidence* (*Preview missing evidence* lists them) | `fetch-windows` (`{"siteId", "dryRun"?}`) | — | | A site's report evidence media (then its exports) | its **Reports** tab → *Prepare evidence media* | `reports-prepare` (`{"siteId"}`) | `reports prepare ` | | A site's report exports (HTML, PDF, Markdown, slides, evidence pack) | its **Reports** tab → *Export reports* | `reports-export` (`{"siteId", "reportId"?, "formats"?}`) | `reports export [--report ] [--formats html,pdf,md,slides,slides-pdf,zip]` | | A report from a /sweep report, an /ask answer or a report-to-video manifest, and a starter manifest from a report | — | — | `reports convert --out [--channels-dir ]`, `reports to-manifest --out ` | **The editor.** The /sites **Publish** panel has a row per site, then the hub and the homepage, each with four chips — `index`, `built`, `deployed`, `live` — its policy and what is next; Deploy local shows where `ARCHILYZER_SITE_OUT` / `ARCHILYZER_HOMEPAGE_OUT` is set. Above the rows, beside **Publish now** and **Build all stale**, is the plan Publish now would enqueue and what it leaves out. A site's Publish tab shows when its bundle was built and from which branch, what it last deployed and how the live check read. **A button always runs**: a Build is forced, and the index update goes first when the index is not fresh; a Deploy is forced too. The CLI runs only the stage you name — `publish build` builds from the index as it stands. **A run is ordered on disk.** A run's stages are enqueued together under one run id; a build carries `indexAfter` and a deploy `builtAfter`, and each is blocked ("waiting for the index update this run started") until that step's stamp is new enough, so a failed or cancelled step leaves the next one refused rather than shipping old data. Cancel on a run's console cancels every stage of that run still queued or running. /jobs reads each stage as `run · `. A stage still queued when the editor restarts is cancelled, never re-queued: the lane works out again what is stale. **`pnpm ops publish`** answers `{runId, jobs: [{target, kind, jobId, previewUrl?}], skipped, refused}`; `--wait` follows every job, and a request refused before any job exists is named in `refused` (all refused: a 400). The routes it replaced are **aliases** with their old bodies and answers (`"skipData"` is ignored): `build-index`, `build-site`, `build-deploy`, `deploy-site`, `build-hub`, `deploy-hub`, `build-homepage`, `deploy-homepage`. `build-deploy {"all": true}` builds and deploys every deployable site to production, whatever its policy. **The CLI.** `pnpm archilyzer ` is the short form of `pnpm --filter yt-dlp-transcript-common exec tsx bin/archilyzer.ts `; `--help` lists every command, and `doctor` checks what a publish needs (`wrangler`, `cloudflare-auth`, `r2-keys`, `export-builds`, `publish-lock`, `index-stamp`). The old rows are **aliases** that print what they run: `build site [--nodata]` = `publish index` (not with `--nodata`) + `publish build --force`; `build all` = `publish index` + `publish build all --runner auto`; `deploy site ` = `publish deploy `; `deploy hub` and `deploy homepage` = `publish hub --deploy-only` and `publish homepage --deploy-only`. From `export/`, `pnpm run build` and `pnpm run deploy` are `build site` and `deploy site` (the site from `SITE_ID`). `build hub` and `build homepage [--no-source]` are still bare builds that stamp nothing, so no deploy stage ships them: build with `publish hub` / `publish homepage`. ### In the runtime container Every stage runs inside the editor's container too ([RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md#publish-the-archive)). Run commands with **`docker compose exec editor pnpm archilyzer publish …`, never `docker compose run --rm`**: a second container has its own `export/public`, is a second writer on the index, and carries the editor's fixed host id (`ARCHILYZER_HOST_ID=archilyzer-editor`) with pids of its own, so it would judge the editor's live lock dead and take it. Cloudflare and R2 credentials come from `.env`. `--to local` fills the volumes the `site` and `homepage` services serve. The homepage's source mirror builds there too, with `docker-compose.source.yml` (the host's git directory, read-only) and the operator's files in the config volume; the image has the pinned git-filter-repo but not gitleaks or stagit, so the secret scan is skipped with its WARNING and no history pages ship. The docker runner is refused there. --- ## The publish lane The lane keeps the published targets current by itself: when the index is stale it updates it, then builds what changed and deploys where each target's **policy** says, one stage at a time. It is **off by default** (`settings.publish.enabled`); the buttons and the CLI work either way. **Policies**: `off` (default; the lane leaves it alone), `build`, `preview` (built, then deployed to `settings.publish.previewBranch`, default `preview`) or `production`. A site's is `site.json` `publish.auto` (the **Publish policy** on its settings form). **A private site is only ever built** — its policy is clamped to `build` — and `preview`/`production` need a `cloudflareProject` (a save without one is refused; a file that says so reads as `build`). The hub's is `settings.publish.hub` (production needs its Pages project in `homepage.json`), the homepage's `settings.publish.homepage` (which also publishes the source mirror). **A pass.** The runner (`auto-publish` on /jobs) wakes every `checkEveryMinutes` (10) and starts no pass while the lane is **held**, in its `quietHours`, or while any publish stage is queued or running. A pass is due with no index stamp; when the index is stale and its last update is at least `refreshEveryMinutes` (360; 0 = whenever stale) old; or when a policy target is left stale and the last pass is that old. It dispatches ONE stage, waits for it, then re-plans; it never forces. A failed index update ends the pass and a failed build drops its deploy. A hold, quiet hours or the lane switched off stop the dispatching between stages and never kill one. **Drain** finishes the stage in flight and ends the runner; **Stop** ends the runner and leaves a running stage to finish (Cancel it on /jobs). An idle boot leaves the runner off. **Publish now** is the same plan enqueued at once: the index update when stale, each site's build and its policy's deploy, then the hub, then the homepage — with every policy off, the index update alone. **Build all stale** builds every stale site whatever its policy and deploys nothing. **/operations/publish** is the lane's page: Start, Drain, Stop, **Hold the lane** (`settings.publish.held`), when a pass is due and why, the last pass, what a pass would run now, and the lane's settings — including `runner` ([docker](#building-every-site-in-containers), host only). --- ## Cloudflare Pages Sites deploy to **Cloudflare Pages** with `wrangler pages deploy`; download archives too large for Pages' per-file limit overflow to **Cloudflare R2**. Everything here fits inside Cloudflare's **free tier**. - **A Pages project must exist before its first deploy.** wrangler offers to create a missing project only on an interactive terminal, and the deploy's stdin is a pipe, so a missing project fails at once with wrangler's own "does not exist" sentence. Create it first: `pnpm --filter yt-dlp-transcript-common exec wrangler pages project create --production-branch main`. A site's project is `site.json`'s `cloudflareProject`. - **wrangler is pinned**: an exact devDependency of `common` (4.147.0), run as `common/node_modules/.bin/wrangler` — nothing is fetched at deploy time. `WRANGLER_BIN` overrides it (the e2e suite's fake). It needs Node 22 or later. - **The credential preflight**: `CLOUDFLARE_API_TOKEN`, or a `wrangler login` on this machine. With neither, the deploy is refused before wrangler runs: "[deploy] REFUSED — no Cloudflare credentials: set CLOUDFLARE_API_TOKEN in .env (or run `wrangler login` on this machine). Nothing was sent to Cloudflare." A token Cloudflare rejects — wrong, expired or malformed — ends the log on **"[deploy] REFUSED by Cloudflare — the API token was not accepted"**. Both exit 1. - **Every deploy names its branch**: production `--branch main`, a preview `--branch `, never inferred from the checkout. **Production ships only a build of `main`**: a bundle whose `built.json` records another branch, or none (a detached HEAD, an image built without `ARCHILYZER_BRANCH`), is refused — build it from `main`, or deploy it as a preview. A project whose production branch is not `main` would take `--branch main` as a preview. **A deploy, in order**: the target's refusals; its `built.json` (already in this slot: a no-op unless forced); the bundle guards (its `site.json` and `corpus.json` must name the site); the preflight; a site's oversize archives to R2; wrangler; the live check; then `deployed.json`, written only when everything before it succeeded. **The live check.** A wrangler exit 0 says a bundle was uploaded, not that a visitor gets it. The stage reads `/corpus.json` **plain** and **cache-busted** (`?cb=`), up to three tries 10 s apart, and compares each `generatedAt` with the build's. The URL is the preview alias, or for production the target's public URL (a site's `siteUrl`, the hub's, the project's), else the deployment URL wrangler printed; the homepage's check reads `/` for a 2xx. | Verdict | Means | |---|---| | `ok` | both reads serve this build (or the plain one does and the busted one failed) | | `stale-edge` | the busted read serves this build, the plain one an older copy: Cloudflare's edge still holds the old object | | `mismatch` | the busted read serves another build | | `unreachable` | neither read answered 2xx, or the plain read failed | | `skipped` | `E2E_LIVE_CHECK=skip` | **A verdict short of `ok` is a warning, never a failure**: `[live] WARNING — …` with each read's status, `generatedAt`, `cf-cache-status`, `age` and `cache-control`; the job ends `done` and the verdict is recorded. A rebuild with no data change carries the same `generatedAt`, so the check tells a stale or wrong deployment from the build's data, not two builds of the same data apart. **Tombstones for withdrawn X posts.** Leaving a path out of a deploy does not take it off Cloudflare's edge (the hub served a withdrawn X shard from a week-long cache), so while `social.x.visibility` is `private` withdrawn posts are **replaced, never deleted** (`common/publish/tombstones.ts`). A public site writes, per withheld X channel, `posts//manifest.json` with no pages and `[]` for every page the shared tree has (and an empty `posts/manifest.json` when it lists no posts at all); the hub writes the same for every X channel a non-private site carries, plus an empty posts manifest — a channel only private sites carry is never named on the hub. `_headers` serves them `Cache-Control: no-store`. **The hub's deploy reads each tombstone back**, plain and busted, into its live check. ## The source mirror (homepage) The homepage carries the project's own source, read-only, as static files: > **Before ANY deploy of the homepage — a preview included — the denylist must hold > everything private.** A Pages preview is a public publication, and every Pages > deployment stays reachable at its own `..pages.dev` URL until that > deployment is deleted; a preview's branch alias is guessable (the plans that name > it ship in the mirror). If a deploy ever carries something private, DELETE that > deployment in the Pages dashboard: a newer deploy does not remove the old one. | Path on the site | What | |---|---| | `/source/` | The page: the clone command, the mirror head, the private `main` it reflects | | `/source/archilyzer.git/` | A git repository for git's **dumb HTTP** protocol: `git clone https://archilyzer.pages.dev/source/archilyzer.git` | | `/source/tree/` | Every tracked file of `main`, raw, with a generated `index.html` per directory | | `/source/git/` | The history: `log.html`, a page per commit with its diff (`commit/.html`), `refs.html`, `files.html` (an index into the raw tree), `atom.xml` and `tags.xml` — rendered by **stagit** when it is installed (below) | | `/downloads/archilyzer-source.tar.gz` + `snapshot.json` | The same tree without history | | `/source/manifest.json` | What was published, from which private commit, audited how | **`archilyzer source publish`** makes them, and **the build-homepage stage runs it** between compose and `next build` — `archilyzer publish homepage`, the Homepage row's Build homepage on /sites, `pnpm ops publish {"verb": "homepage"}` and the lane. A stage is a child of the editor (or of the CLI) and runs in its environment and `PATH`: the editor's process needs `~/.local/bin` on its `PATH` to find a pipx-installed `git filter-repo`; without it the step falls back to `pipx run`, which needs the network. The bare `archilyzer build homepage` runs it too. **A refusal withdraws the source, everywhere it could ship from.** It fails the build before `next build`, and: - the step removes the LAST publish from `homepage/public` (manifest first, then the mirror, the tree, the history pages, the tarball, `snapshot.json` and the skip key, and the history's render cache) — it was audited under rules that may not be today's, and the refusal is the best evidence that they are not. Any outcome but success does this once the rules are loaded (an audit hit, a limit, a missing tool, a cancel, a crash); `--check` writes nothing, this included; - the build removes the last BUILD's copy from `homepage/out` (`out/source`, the tarball, `snapshot.json`); - **the deploy-homepage stage refuses** an `out/` holding a source unless the skip key says that publish was made under today's rules and step version, of today's `main`, by today's gitleaks, and is exactly the one in `out/`: its mirror head, and a digest over every published file (the mirror, the tree, the history pages, the manifest, the tarball, `snapshot.json` — sorted path, size and sha256, recomputed over `out/`), so a mixed or edited `out/` refuses too: "run `archilyzer build homepage` (it re-audits), then deploy" — rebuild with `archilyzer publish homepage`, the stage, which stamps the bundle the deploy reads. History pages that are not the audited ones (edited, missing, or present when none were published) are named on their own: "homepage/out's history pages (/source/git/) are not the ones that were audited". An `out/` whose `/source` page shows the empty state (`--no-source`) deploys as before; one with no `/source` page (a refused build) does not. **A stage runs the checkout's source, not the editor's bundle.** Each is `tsx bin/archilyzer.ts stage …` spawned from `common/`, so a merged change to this step applies to the next build-homepage stage without an editor rebuild. What the editor decides itself — the panel, the plan, the refusals before a job exists — is its built bundle, rebuilt and restarted as usual. `build homepage --no-source` (CLI only, the bare build) removes the previously published source instead, because it was audited against the rules of its own day. A checkout with no git repository and no `ARCHILYZER_SOURCE_REPO` — a tarball install, the runtime container without `docker-compose.source.yml` — has nothing to mirror: the build says `no git repository here; nothing to mirror — the /source page will show its empty state`, removes any old publish and goes on; `source publish` run directly there exits 1 with the same sentence. What one publish does: 1. A **fresh** `git clone --no-local --bare --single-branch --no-tags --branch main` of `ARCHILYZER_SOURCE_REPO` when it is set (a git DIR — in a container, the read-only mount of the host's), else the checkout's git *common* dir — so a worktree build mirrors the primary's `main`. A variable naming a path that is not there is a refusal naming it. The private repository's history is never rewritten. 2. **git-filter-repo** rewrites that copy: file contents (`--replace-text`) AND commit messages (`--replace-message`) with the operator's scrub rules, then two built-in ones (`SESSION_LINK_RULES`): the `Claude-Session` trailer line goes from every message and file, and any other Claude session link becomes `[session link removed]`; the gate then denies `claude.ai` session links in any case (`built-in session-link rule`). Author and committer identities are not rewritten — the gate still reads them. 3. `git repack -a -d --max-pack-size=20m`, `pack-refs`, `update-server-info`: the dumb-HTTP file set is `HEAD`, `refs/heads/main`, `packed-refs`, `info/refs`, `objects/info/packs` and `objects/pack/*.{pack,idx}`, copied from an **allowlist** — never `config` (it names the clone's origin path), `hooks/`, `filter-repo/` (the PRIVATE commit ids) or `logs/`. 4. **The gate** (`common/publish/sourceAudit.ts`): every object of the rewritten mirror — blobs, commits and tags whole (messages and identity lines), trees by entry name — then every staged file and path (the tarball decompressed), searched for every denied literal; then gitleaks over the history, when it is installed. **One hit and nothing is written** (and the last publish is withdrawn, above). The report names a literal only by where it was written — `denylist line 3 (len 5)`, `scrub line 2 lhs (len 11)`, `built-in home rule (len 11)` — never by any of its characters, and a hit only by its object: kind, id, the blob's path in history, the byte offset, and for a commit or tag the field (`author`, `committer`, `tagger`, `message`). It prints no byte from the object, because what sits beside a denied name (a surname, the rest of an address) is as private as the name, and every refusal message and path it prints is masked (`[REDACTED]`), so a literal that spans path components (`a/b`) is not printed by a refusal that names a tree path. It ends with `add a rule to ~/.config/archilyzer/source-scrub.txt or drop the file from history, then re-run.` 5. The tree (`git archive` → `tar -x`, a page per directory; a tracked `index.html`, a tracked `404.html` — Pages would serve it as HTML for every missing path below it — or a symlink is a refusal) and the tarball (`git archive --format=tar.gz -9`). Then the history pages (stagit, below), into the same stage — so the file sweep of step 4 reads every one of them too. 6. Limits, then the install: link-safe (`common/bin/_publicFile.ts`), the manifest removed first and written **last**, so a crash leaves the page's empty state, never a manifest over a half-copied tree. An unchanged `main` with unchanged rules, step version, filter-repo and gitleaks, and published files that still match their digest, skips (`[source] up to date at …`; `--force` rebuilds). `--check` does everything but the install and writes nothing. `--keep-scratch` leaves the scratch clone (`ARCHILYZER_SOURCE_SCRATCH`, default the OS temp dir) for a look, minus `replace.txt` (the scrub rules), which is always deleted; a scratch root inside the checkout or the public dir is refused, since a kept scratch there could be committed or published. `archilyzer source audit []` runs the gate alone — over the published mirror by default, or over a clone: `archilyzer source audit /.git`. **The two operator files** live outside the repo, in `${ARCHILYZER_CONFIG_DIR:-~/.config/archilyzer}/` (`SOURCE_SCRUB_FILE` and `SOURCE_DENYLIST_FILE` override each path; [ENVIRONMENT.md](ENVIRONMENT.md)). They are never committed; keep them mode 600. A missing file is a refusal naming it. - `source-scrub.txt` — git-filter-repo `--replace-text` lines: `lhs==>rhs` (split at the last `==>`), `literal:…`, `regex:…==>…`, `glob:…==>…`; a line with no `==>` is replaced by `***REMOVED***`. Lines whose first non-blank character is `#` are comments (the step drops them — filter-repo itself would treat one as a literal). `==>/home/user` is built in and runs first (a trailing `/` on the home dir is dropped). Rules apply in order. An empty left side is a refusal. - `source-denylist.txt` — one literal per line, `i:` in front for any ASCII case; `#` comments. - In both, a UTF-8 byte-order mark and CRLF line endings are dropped first: either would silently turn the first rule into one that matches nothing. - **Every literal scrub rule's left side is denied too** — its exact bytes, so a rule that stopped matching because its text was mistyped refuses instead of leaking. A NEW spelling in history (another case, another path) is caught only by the denylist: deny the name itself, not just the paths it appears in. - The rules' hash (with the step's version) is the skip key and the deploy check, kept in `homepage/.source-publish.json` — beside `public/`, never in it: a published hash of the denylist would confirm a guess at it. For the same reason the manifest does not say how many literals there are. **A known limit: compressed content is opaque.** The gate is a byte search, so it cannot see inside a ZIP member, a PNG's compressed text chunks, a PDF stream or a woff2 (the tarball is the exception: it is decompressed). And git-filter-repo never scrubs a blob with a NUL in its first 8 KiB, so a binary is audited, never scrubbed. At release 12 the review decompressed every such blob in the published history (a zip, PNGs and icons) and found no denied literal. A private string inside a binary must be removed from history by hand. **Tools.** `git filter-repo` (`pipx install git-filter-repo`; without it the step runs `pipx run --spec git-filter-repo==2.47.0`, which needs the network on first use, and without pipx it refuses with the install line; the runtime image has that version installed). gitleaks is optional: without it the secret scan is skipped with a WARNING and the literal audit still runs. stagit is optional too (next section). Neither is in the runtime image. `archilyzer doctor` has a "source publish" block: which filter-repo would run, stagit (its path, or not found, and the render cache's path and size), gitleaks, the repository it would mirror (`source-repo`), the config dir, the two files (rule counts and modes, never contents) and the last publish. The mirror's ids are deterministic for a given filter-repo version; an upgrade that changes its rewriting changes every id (readers re-clone). The manifest records both tools. **The history pages (`/source/git/`) are stagit's.** [stagit](https://codemadness.org/stagit.html) (C over libgit2) renders the scrubbed mirror as static pages: the log, a page per commit with its diffstat and diff, the refs, and two Atom feeds. It is an operator-installed tool like git-filter-repo — never vendored, never committed. Install it once (libgit2 and its headers are the one dependency): ```sh git clone git://git.codemadness.org/stagit && make -C stagit && cp stagit/stagit ~/.local/bin/ ``` The step finds it as `STAGIT_BIN`, else `stagit` on `PATH`, else `~/.local/bin/stagit` (the editor's process needs `~/.local/bin` on its `PATH` for a pipx `git filter-repo`; for stagit the fallback covers it). **Without it the source is published without the history**: one line says so and how to install it, the manifest has no `history` block, and `/source/` shows no History links. A render that fails, or history pages that would break the host's limits, are the same, with a WARNING — never a failed build, and the next build tries again (a stagit present and no history never skips). What the step does with it: - It runs after the mirror is built and its objects audited: `stagit -c -u https://archilyzer.pages.dev/source/git/ ` (past the cap, `-l 10000` in place of `-c`, below). The clone is named `archilyzer.git` (stagit names the repository after its directory); its `description` is set to "Archilyzer" and its `url` to the clone URL, for stagit's header. Neither file is published. - **Only an allowlist is published:** `log.html`, `files.html`, `refs.html`, `atom.xml`, `tags.xml`, the page of each commit the log lists (`commit/.html`), and a `style.css` the step writes from `common/styles/tokens.css` (the homepage's two grounds, light and dark; diff insertions in `--success`, deletions in `--destructive`). **stagit's per-file pages (`file/…`) are not**: browsing is the raw tree, so every link into `file/` — the Files index, the README and LICENSE links in the header, a diff's file names — is rewritten to `../tree/`. - Every `.html` page (never the feeds) gets exactly two additions: the homepage's pre-paint theme script (the one `ThemeScript` emits, `lib/themeConfig.ts`), so a page opens on the visitor's stored base or the homepage's dark default — without JavaScript, `prefers-color-scheme` decides — and one line at the top, "Archilyzer · Source", linking to `/source/`. stagit's logo and favicon point at the site's `/icons/icon-32.png`. A link to what is not published keeps its text and loses its `href`: a diff of a file main no longer has (deleted or renamed since — stagit links every diff side), and a commit with no page (past the cap, the oldest page's parent); the raw tree's file list decides. Nothing else is changed; the pages are rewritten byte for byte. - **The gate reads every page**: they are in the stage the file sweep reads, so a denied literal in one refuses the publish, named by its path (`file source/git/commit/.html (contents, byte N)`) and never by its bytes. The pages are rendered from objects the object sweep already read, so a hit there means something stagit or the step added. - The manifest's `history` block has the log's href, the commits with a page and the commits in all (`commits`, `total`), the head, the file count and bytes, a sha256 over the pages, and stagit's identity (the sha256 of its binary; it has no version flag). The skip key has stagit's identity and the pages' digest. - **The cap: the newest 10,000 commits** (`SOURCE_HISTORY_MAX_COMMITS`). Past it, stagit runs with `-l 10000`: its log lists the newest 10,000 and ends "N more commits remaining, fetch the repository", and `/source/` says "the latest 10,000 of M commits". `-l` does not cap the pages — stagit still writes one for every commit — so the step publishes the page of each commit the log lists and no other; and stagit refuses `-c` with `-l`, so past the cap every publish computes 10,000 diffstats (about 5 ms each here) where `-c` computes only the new ones. The 15,000-file drop below stays, as the last resort. - **The render cache is `${XDG_CACHE_HOME:-~/.cache}/archilyzer/source-history/`** (about 140 MB at 1,872 commits; made on first use; never inside the checkout or the public dir — the step renders without it there, and says so). stagit keeps a commit page once rendered, and `-c` keeps its log lines, so a publish renders only the new commits. At 1,834 commits, stagit alone took 0.6 s with nothing new against 8 s for every page; the whole history step (with the post-pass and the copy) 2 s against 8 s. **But `-c`'s walk is in commit-date order and stops at the head it rendered last**, so a `--no-ff` merge of commits older than that head leaves them out: the step's count check sees the log come up short and renders every page again (about 8 s), with a line that says so. With this repo's merges of parallel slices that is common, and correct. The cache holds stagit's `-c` file, its output and a key. Its pages are kept only under the same rules, step, filter-repo, stagit and header text, and after a run that finished; its log lines only when they end at an ancestor of today's head. One publish holds it at a time; a lock whose pid is not running, or over an hour old, is stale and replaced (one line). A cache that cannot be made or written (EACCES, EROFS, ENOSPC) is one line and a render without it, never a failed build. `--force` renders every page again, `--check` never touches it (it renders in its own scratch), and a refusal removes it. `archilyzer doctor` ends its stagit line with `cache: , `. - **Its size, and the limit ahead.** At 1,872 commits (2026-09-30): 1,878 files, 141.3 MB, the largest page 5.0 MB (stagit prints "Diff is too large, output suppressed" past 1,000 files or 100,000 lines in one commit); the whole publish is 4,396 files, and `homepage/out` 4,576. One file per commit counts against the step's 15,000 (of Pages' 20,000); the cap keeps it at 10,006 at most. Should the rest of the publish ever grow past 5,000 files, the history is dropped with a WARNING, never the publish. **Cloudflare Pages, and the traps it sets.** - **Never a path segment named `.git`.** wrangler's upload ignore list drops `**/.git` (and `**/node_modules`) *silently*, and Cloudflare's managed rules block `/.git/` requests. The mirror is `archilyzer.git`, which passes both. Extensionless files (`HEAD`, `info/refs`) upload and serve fine. - **The caps:** 20,000 files per deployment and 25 MiB per file. The step refuses at **15,000 files** (the rest of the site needs room) and **24 MiB**; packs are split at 20 MB. - **`_headers` makes the raw tree plain text.** wrangler's mime map would serve `.ts` as `video/mp2t`, `.mjs` as JavaScript and `.html` as a page on this origin, so `homepage/public/_headers` sets `/source/tree/*` to `text/plain; charset=utf-8` with `nosniff` and `noindex`, then puts back `text/html` for the directory pages and the real types of the few binaries (`.ttf`, `.m4a`, `.zip`, `.onnx`, `.svg` — the SVG under a `default-src 'none'` CSP). `/source/git/*`, the history pages, is `noindex` too (ruled), and keeps its own types. **Every matching rule applies, and a header a later rule sets again is APPENDED** (`text/plain; charset=utf-8, text/html; charset=utf-8`), so each of those overrides first detaches the tree's type with `! Content-Type`. Only the ROOT `_headers` is read; the tree's own copies are served as text. Locally, only `wrangler pages dev` honours `_headers` — `next dev` and `serve` do not, and neither serves a directory's `index.html` at `//` the way Pages does. - **Tailwind and tsc must not see the mirror.** Tailwind v4 scans every file `.gitignore` does not exclude (a binary pack yields "class names" that break the stylesheet), and `homepage/tsconfig.json` includes `**/*.ts`: `/homepage/public/source` is gitignored, and both `public` and `out` (where `next build` copies it) are excluded from the tsconfig. Keep all three. `.dockerignore` leaves the mirror and the skip key out of every image. ## Preview deployments A **preview** is the same bundle — the target's own, as built — deployed to a branch that is not the Pages project's production branch. Cloudflare publishes it at a **branch alias** — ``` https://..pages.dev ``` — and leaves the live site alone. Each deploy also gets an immutable per-deployment URL (`https://..pages.dev`), which wrangler prints as "Deployment complete! Take a peek over at …"; the deploy repeats both on a `[preview]` line at the end of its log, because the streamed log scrolls, and records them in the preview's slot of `deployed.json`. Its live check reads the alias. The alias is a function of the project and the branch and nothing else, so it is known *before* the deploy runs — which is why the editor can link it while you are still typing the name. **Branch names** must be 1–28 lowercase letters, digits and dashes, starting and ending with a letter or digit. That is exactly what Cloudflare's alias sanitizer preserves verbatim, so the alias shown is the alias that resolves. `main`, `master` and `production` are refused: a deploy to the production branch is not a preview, it is the live site. | Surface | How | |---|---| | Editor | A site's row on /sites: type a branch (it starts on `settings.publish.previewBranch`), press **Deploy preview**; **Deploy production** beside it. A site's **Publish** tab → *Individual steps* → **Deploy a preview**. The hub's and the homepage's rows take a branch in their preview box (empty = production). | | Ops API | `pnpm ops publish --json '{"verb":"deploy","siteId":"anilyzer","preview":"tags-exclude"}' --wait` deploys the built bundle and prints the alias on its own line after the log; the answer carries `previewUrl`. `"hub"` / `"homepage"` take `"deploy": true, "preview"`. The `deploy-site` and `build-deploy` aliases take `"preview"` as before. | | CLI | `pnpm archilyzer publish deploy anilyzer --preview tags-exclude` | | The lane | a target whose policy is `preview` deploys to `settings.publish.previewBranch` | The high-value loop is **build once, preview, then promote**: build the site, deploy it with a preview branch, look at it, then deploy again with no preview — the same bundle, unrebuilt. Promotion is a production deploy, so the bundle must have been built on `main`; a build from a feature branch previews but never promotes. Each branch is its own slot: deploying the same build to the same branch again does nothing unless forced. **A preview shares the production R2 archive bucket.** R2 has no per-branch namespace, and the keys are `/archives/.zip` either way. In practice this is cheap and harmless — the upload skips any object R2 already holds at the same size, and an unchanged channel re-zips byte-stable — but a *changed* archive replaces the one production's manifest links to. Every preview deploy of a site says so once, at the top of its log. **Who can open a preview is a Cloudflare setting, not ours.** Pages projects have a *preview deployment access* setting (Settings → General): **public** by default, or restricted to Cloudflare Access. A default-configured project's preview URL is world-readable by anyone who has the link. --- ## Download archives and R2 At compose time, `archilyzer compose site` generates one transcript zip and one live-chat zip **per channel** into `export/public/archives/`, and records them in `public/archives/manifest.json`. The `/downloads` page and the header **Downloads** link read that manifest. A site can turn its archives off (`site.json` `archives: false`), and a build can skip them (`--skip-archives`, `BUILD_ARCHIVES=0`, or the **Skip archive zips** checkbox on a site's Publish tab). Cloudflare Pages rejects any single asset larger than **25 MB**, and real channels blow past that easily (a channel's live-chat zip can be hundreds of MB). So the pipeline splits archives by size: | Archive size | Where it's served from | Manifest entry | |---|---|---| | ≤ 25 MB | Cloudflare **Pages** (shipped in the bundle, free) | `filename`, no `url` | | > 25 MB, R2 configured | Cloudflare **R2**, uploaded on deploy | `url` → R2 | | > 25 MB, R2 **not** configured | not served | `oversize: true`, shown as "Too large to host" | (The cap is `MAX_ARCHIVE_BYTES`, or a site's own `archiveMaxBytes`; `0` = no cap.) Oversize archives are staged by compose into `export/.r2-staging//archives/` (gitignored, kept out of `public/`), and **the build moves them beside the site's bundle**, `export/.export-builds//.r2-staging//archives/`, replacing the last build's (the docker runner's containers write them there directly); `built.json` counts them. They are uploaded to `//archives/.zip` **before** the Pages deploy runs, so the manifest URLs resolve immediately: **every Pages deploy of a site uploads them**, however the deploy stage was started. A build alone only stages them, and a `--to local` deploy uploads nothing. The bucket is read from `settings.json`; the R2 credentials come from the environment (step 3 below). Uploads go through R2's **S3 API** (the AWS SDK's multipart uploader), not `wrangler r2 object put` — wrangler caps a single upload at **300 MiB**, and real live-chat archives are larger (multipart has no such limit). That is why archive uploads need S3 credentials in addition to the wrangler auth the Pages deploy uses. ### One-time R2 setup **1. Create the bucket.** One bucket serves **all** your sites — objects are namespaced by `siteId`, so you do **not** need one bucket per site. In the dashboard (**R2 → Create bucket**) or with the same wrangler the deploy already uses: ```sh wrangler r2 bucket create my-archives-bucket ``` **2. Expose it publicly.** R2 buckets are private by default, and the Downloads page links to plain `https://…/.zip` URLs, so the bucket needs public read access. There are three ways; the first needs **no domain purchase** and is the recommended one: - **A Worker on `*.workers.dev` (recommended — no domain, no WHOIS):** deploy the bundled `r2-proxy/` Worker to a free `.workers.dev` subdomain (just like Pages' `*.pages.dev`). It streams objects from the bucket and gives you edge caching **and** tunable rate limiting in code — the same cost defenses a custom domain would, without owning a domain. One Worker serves **every** site. See [Option A](#option-a--a-worker-on-workersdev-no-domain) below. Set the editor's public URL to `https://.workers.dev`. - **Custom domain:** bucket **Settings → Public access → Custom Domains → Connect Domain**, e.g. `archives.example.com` (the domain must be on your Cloudflare account). Routes downloads through Cloudflare's CDN, WAF, caching, and dashboard rate-limiting rules. See [Option B](#option-b--a-custom-domain). - **`r2.dev` subdomain (quick test only):** bucket **Settings → Public access → Allow Access**. You get `https://pub-abc123.r2.dev` — free and domain-free, but Cloudflare throttles `r2.dev` and gives you **no** cache / rate-limit control of your own. Fine for a smoke test; use the Worker for a real instance. **3. The R2 S3 credentials.** Archive objects are uploaded over R2's S3-compatible API, which needs an Access Key ID + Secret. In the dashboard: **R2 → Manage R2 API Tokens → Create API token**, permission **Object Read & Write**, scoped to your bucket. Then set three environment variables where the editor (or the CLI) runs — in Docker, in `.env`: ```sh export R2_ACCESS_KEY_ID= export R2_SECRET_ACCESS_KEY= export CLOUDFLARE_ACCOUNT_ID= # used to build the S3 endpoint ``` The account id is on the R2 overview page; the S3 endpoint is derived as `https://.r2.cloudflarestorage.com`. Keep these in the environment (a shell profile, a systemd unit, a `.env` the editor loads) — **not** in the settings JSON, which is not a place for secrets. If a deploy has oversize archives to upload but these are unset, it fails *before* the Pages deploy (so the site never links to a missing file) with a message pointing back here. `archilyzer doctor`'s `r2-keys` names any that are unset, once a bucket is configured. **4. Point the editor at the bucket.** In **Settings**: | Field | Value | |---|---| | **Archive overflow storage (R2 bucket)** | the bucket name, e.g. `my-archives-bucket` | | **Archive overflow public URL** | the public base from step 2, e.g. `https://archives.example.com` (no trailing slash needed) | Leave both blank to keep oversize archives dropped and shown as "Too large to host". On the next build and deploy, oversize archives upload to R2 and the Downloads page (and the header link, which reappears once anything is hostable) point at them. --- ## Securing downloads against cost-abuse The threat: someone scripts repeated downloads of large archives to run up your bill. **The reassuring part — R2 egress is free.** Unlike S3, Cloudflare R2 charges **$0** for bandwidth/egress. An attacker looping downloads of a 480 MB zip cannot run up a bandwidth bill. The *only* metered cost from reads is **Class B operations** (10M free per month, then $0.36/M) — and the defenses below make even that hard to reach. You get these controls one of two ways — **a Worker on `workers.dev`** (no domain) or **a custom domain**. Pick one. The Worker is recommended if you do not want to own a domain. ### Option A — a Worker on workers.dev (no domain) The bundled **`r2-proxy/`** Worker is a small, self-contained project that binds the R2 bucket and serves archive objects on a free `.workers.dev` subdomain. It is a **pure passthrough**: the request path `/archives/.zip` maps straight to the bucket key, so **one Worker serves every site** (deploy it once, not per site). What it gives you, all on the free tier: - **Edge caching** — full downloads are cached with the Cache API, so repeat pulls skip R2 (no billable Class B op). It honors the `Cache-Control` set on each object. - **Rate limiting** — Cloudflare's native, free rate-limit binding caps requests per client IP + file (dashboard rate-limit rules need a paid zone; this does not). - **Path allow-listing** — it only serves `*/archives/*.zip`, never arbitrary keys. - **Range / resumable downloads** — honors `Range` requests so big zips can resume. **Deploy it (once for the whole instance):** 1. Edit `r2-proxy/wrangler.toml` and set `bucket_name` to the **same bucket** you use in the editor's Settings. Optionally rename the Worker (`name`) and tune the rate limit (`limit` / `period`). 2. From the repo root: ```sh cd r2-proxy pnpm install # first time only pnpm run deploy # = wrangler deploy, reusing your host wrangler auth ``` wrangler prints the deployed URL, e.g. `https://ytdlp-archive-proxy..workers.dev`. (An "unsafe fields are experimental" warning for the rate-limit binding is expected.) 3. In the editor's **Settings**, set **Archive overflow public URL** to that `workers.dev` URL. Re-deploy a site and its Downloads links resolve through the Worker. The Worker code is `r2-proxy/src/index.ts`; the caching and rate-limit logic are small and commented. > **Free-tier limit:** Workers Free allows **100,000 requests/day** (resets daily). Far > more than a downloads endpoint needs; if you ever exceed it, requests get a `429` > (fail closed — no surprise bill) until the next day, or upgrade to Workers Paid ($5/mo). The archive `Cache-Control` is set at upload time (`Cache-Control: public, max-age=3600`, constant `ARCHIVE_CACHE_CONTROL` in `common/publish/build.ts`); the Worker reads it back when caching. Archive filenames are stable and overwritten in place on re-deploy, so this 1-hour bound is what keeps a re-uploaded archive from being served stale for long — raise it if your archives rarely change. ### Option B — a custom domain A custom domain has its own upsides — dashboard WAF, managed bot rules, and rate-limiting rules without touching code. Connect it per step 2 above and add these, in order of impact: 1. **Edge caching.** Through a custom domain, Cloudflare's CDN caches each archive at the edge (it honors the `Cache-Control: public, max-age=3600` set at upload), so repeated downloads are served from cache and **never hit R2**. To make caching aggressive, add a **Cache Rule** (**Caching → Cache Rules → Create**): *when* `URI Path` contains `/archives/`, *then* Eligible for cache, **Edge TTL → Override → 1 day** (or longer). If you raise the TTL a lot, **purge the cache on deploy** (dashboard **Caching → Purge**, or `wrangler`/the API) so a re-uploaded archive is not served stale. 2. **Rate limiting — the hard backstop.** A Rate Limiting rule caps how fast any single client can pull archives, stopping a flood that misses cache (**Security → WAF → Rate limiting rules → Create**; the free plan includes one rule): *when* `URI Path` contains `/archives/`, *rate* e.g. **20 requests per 1 minute** per client IP, *then* Block for 10 minutes (or Managed Challenge). Tune the threshold to real usage — legitimate users download a handful of files, not dozens per minute. 3. **Bot Fight Mode + WAF managed rules.** **Security → Bots → Bot Fight Mode** (free) blocks the low-effort scripted abuse that makes up most of this traffic; the free **WAF managed ruleset** adds a baseline at no cost. 4. **Hotlink protection (optional).** Stops other sites embedding your archives and spending your ops budget serving their audience. A WAF custom rule (**Security → WAF → Custom rules**): *when* `URI Path` contains `/archives/` **and** `Referer` does not contain your domain **and** `Referer` is not empty, *then* Block. (Allow an empty `Referer` so direct clicks and privacy-conscious browsers still work.) ### Billing alerts, and what to skip R2 has no hard spend cap, but Cloudflare **Notifications** (**Notifications → Add**) can email you when R2 storage or Class A/B operations cross a threshold — cheap insurance so nothing surprises you. **Signed URLs / token-gated downloads** are the heavyweight option — they add key management and friction for legitimate users. Given egress is free and caching neutralizes the ops cost, they are overkill for *cost* defense (the `r2-proxy` Worker is a plain passthrough, not an access gate). Reach for them only if you want *access control* (private archives), not cost control. ### Cost expectations Measured across all sites in this project, total compressed archives are **≈ 2–4.5 GB**. Against R2's free tier: | Resource | Free tier / month | This project's usage | |---|---|---| | **Egress / bandwidth** | unlimited, **$0** | irrelevant — no egress charge exists | | **Storage** | 10 GB-month | ~2–4.5 GB — comfortable | | **Class A ops** (writes) | 1,000,000 | ~one PUT per archive per deploy — negligible | | **Class B ops** (reads) | 10,000,000 | mostly absorbed by CDN cache | The realistic bill for hosting these archives is **$0**. The configuration above exists to keep it that way under adversarial traffic, not because normal usage is close to any limit. --- ## Building every site in containers > **Not about running the apps in containers** — that is > [RUNNING_IN_DOCKER.md](RUNNING_IN_DOCKER.md), built from the root `Dockerfile`. This > is about *building sites*: fanning per-site export builds out across containers, > using `Dockerfile.build`. The two share nothing but the word "docker". Every build goes through one stage contract with two **runners**. **local** is the default everywhere, in a container or not: each site in turn, as a child of the editor or the CLI. **docker** is an opt-in for a **Linux host with a container engine**: every stale site at once, each in its own isolated container. Both write the same bundles and `built.json` (`runner: "docker"`), so a deploy does not care which built a site; the docker runner only builds. **Choosing it**: `settings.publish.runner: "docker"` (/operations/publish → Build runner) for the lane and Publish now, or `archilyzer publish build all --runner docker` (`pnpm ops publish {"verb": "build", "runner": "docker"}`); `--runner auto` (the `build all` alias) takes docker when an engine answers. Where none answers `docker version` — **the runtime container included**, which has no engine and never gets the socket — the stage is **refused**, exit 3: "the docker runner needs an engine on this host". No docker-in-docker. **Prerequisites.** Docker, or **podman** (`DOCKER_BIN=podman`; rootless podman maps container files to your host user). A current index stamp — the runner builds from the update-index stage's index and is blocked without one. Node + pnpm on the host (the archive warm and the deploys run there) and the Cloudflare/R2 credentials in its environment; credentials **never** enter a container. The image is built (and cached) from `Dockerfile.build` at the start of every fan-out; Settings → Build pipeline → **Docker image tag** and **Dockerfile path** override the tag and path. **How it works.** One stage, `build-site _all --runner docker`, holding the publish lock throughout: 1. **Which sites**: each site's freshness, as the build stage judges it; fresh sites are skipped (`--force` builds every one). 2. **The archive cache, on the host, once** (`build:archives`, skipped by `--skip-archives`): the shared archive-zip cache for the union of the sites' channels. Only the host writes it, so containers never race it. 3. **Per site, in parallel containers**, up to **Max parallel builds** (Settings → Build pipeline): `docker/build-site.sh` runs `archilyzer build site --nodata`, which in the container builds in place and stamps nothing (`ARCHIVES_READONLY=1`; its lock stays inside the container), and copies `out/` to `export/.export-builds//out`. Network is left on: `next build` fetches fonts through `next/font/google`; isolation is the read-only mounts, the per-site output dir and the non-root user. 4. **On the host, per site**: `builtBundleProblem`, then `built.json`. A failed site is named and the rest are still stamped; the stage then exits 1. | Host | Container | Mode | |---|---|---| | `transcripts/` (corpus + `index.mdb` + archive cache) | `/data/transcripts` | ro | | `export/.export-index` (shared + per-site staging) | `/data/export/.export-index` | ro | | `export/.export-builds/` (public/, out/, .compose-cache) | `/site` | rw | | `settings.json` (build config, mounted fresh — not baked) | `/data/settings.json` | ro | The per-site `/site` mount is persistent, so incremental compose (`.compose-cache`) stays warm; `next build` starts fresh each time in the container's own `.next`. The container's `export/public` is a link to `/site/public`, so `out/` carries this site's data and nothing baked into the image. The container checks that `out/`'s `site.json` and `corpus.json` both name the site before handing it back, the host before it stamps, and the deploy stage again. **Tuning.** **Max parallel builds** — each `next build` can use ~8 GB; start at `floor(RAM_GB / 9)`. `DOCKER_BUILD_MEMORY`, `DOCKER_BUILD_CPUS` — per-container `--memory` / `--cpus` caps. `--skip-archives` (or `BUILD_ARCHIVES=0`) — no archive warm and no download bundles. **Notes.** Containers run as your host uid/gid (`-u`), so `.export-builds/` stays host-owned. The image bakes common/ and export/ with their deps; layer caching keeps a code change cheap. Its build context is `Dockerfile.build.dockerignore`'s allow-list (about 7 MB, read by BuildKit; podman is unverified). The host's `docker/build-site.sh` is mounted over the baked one. `archilyzer doctor` says whether the image is there and older than its Dockerfile. --- ## Reading a published archive from Claude Code Every published archive serves a machine contract (`/corpus.json`, `/llms.txt`), and the MCP server reads it over HTTP. Register it against your public URL — as `archilyzer`, which the tracked `/ask` and `/sweep` commands expect: ```sh claude mcp add archilyzer \ --env TRANSCRIPT_SITE_URL=https://.pages.dev \ -- pnpm --silent -C "$PWD" archilyzer mcp ``` `archilyzer mcp` starts the same server as `pnpm --filter yt-dlp-transcript-mcp exec tsx src/index.ts`, with the environment passed through. `--silent` is what keeps stdout the protocol's alone: without it pnpm 9 and 10 print their `> …` script banner there first (pnpm 11 prints `$ …` to stderr). [mcp/README.md](mcp/README.md) has the other ways to run and register it.