commit 0278e0af2c2dfefde5332afce1cd9ec53772875b
parent 9b649912101756b6a746db374b3b40af5966564d
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Tue, 6 Oct 2026 08:01:21 -0400
plans: release 18 — publishing as queueable stages (the plan, record file)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Diffstat:
| A | plans/release-18.md | | | 358 | +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ |
1 file changed, 358 insertions(+), 0 deletions(-)
diff --git a/plans/release-18.md b/plans/release-18.md
@@ -0,0 +1,358 @@
+# Release 18 — publishing as queueable stages (one index, serial builds, deploys checked live, stages run in the container)
+
+Written 2026-10-04 in plan mode against `main` `0a7a02e9` (releases 13–17 live; today's manual rollout of all six
+sites + hub done 13:34–14:09); implementation began 2026-10-06 against `main` `b8466d24`. Self-contained: the
+operator clears context after this. Rules: `plans/tools/implementer-rules.md` — one Opus implementer per slice in its
+own worktree (`pnpm wt add r18/<slice>`), one read-only Opus review per slice, the parent merges `--no-ff` on a clean
+tree (into `r18/integration`, fast-forwarded to `main` at the end — other sessions share the primary checkout);
+records state rulings never reasons; counts-only privacy greps before every merge;
+changelog bullets checked by eye; never run a whole e2e suite while agents work; the homepage build gate is
+`build:nodata`; the export build gate is `pnpm --filter export exec next build`; never boot a second editor against
+the real corpus; implementers' commits carry their OWN model's `Co-Authored-By` + the session line. Record file:
+this file (shape of `plans/release-17.md`).
+
+Path corrections against the tree as of 2026-10-06: `pagesDeploy.ts` is `common/lib/pagesDeploy.ts`; the e2e fake
+binaries live in `editor/e2e/fixtures/bin/` (so the fake wrangler is `editor/e2e/fixtures/bin/fake-wrangler.mjs`);
+`e2e/helpers.ts` is `editor/e2e/helpers.ts`.
+
+## Context
+
+Publishing today is a handful of one-shot jobs: `build site <id>` reruns the whole data phase (index + stats +
+templates, 14 min today) for every site unless `--nodata`; from `/sites`, "Build index"/"Build stats" call
+`buildIndex()`/`buildStats()` IN-PROCESS on the editor's event loop (`editor/app/sites/lib/buildAction.ts:47-76`), which
+starves the dashboard (`scripts/archilyzer-ops.mjs:379-383`); every site and the hub build into ONE `export/out` and ONE
+`export/public` (`common/publish/build.ts:39-41`; `:347` "another job can rewrite export/out"); jobs are flat (no
+dependency, group or run id; `common/jobs/registry.ts:88-138`), the only chain is `build-deploy` = one function doing
+both (`buildAction.ts:120-195`); wrangler is fetched UNPINNED by `pnpm dlx` at deploy time (`build.ts:382`); nothing
+checks a deploy live; a withdrawn path stays on Cloudflare's edge until its TTL (the hub still serves 313 private X posts
+from a 7-day cache after today's deploy); the runtime image cannot run the homepage's source mirror (no python/pipx/
+git-filter-repo) and "deploy runs on the host" (PUBLISH.md:639-641). Today's rollout proved the shape that works: ONE
+data phase, then each site `--nodata` (~1 min) + deploy (~1 min), serially.
+
+**Measured/verified 2026-10-04:** `buildIndex`/`buildStats` are incremental with internal LMDB fingerprints
+(`generation`, `siteFp:<id>`, `statsFp:<id>`, page sha1; result `shortCircuited`, `buildIndex.ts:2312`, `buildStats.ts:778`;
+`INDEX_SCANNED_AT_KEY` `:2289`) but expose NO external stamp; `build templates` is unconditional; `buildAll` basic mode
+already passes `skipData: i > 0` (`build.ts:838-847`); the Docker fan-out already builds into `exportBuildsDir/<id>/out`
+(`dockerSiteOutDir`, `:45-47`) over a read-only `.export-index`; `runManagedCommand` sets `record.child` so Cancel kills
+the process, `runManagedFunction` never does (`common/jobs/shutdownCancel.ts:15-18`); the scheduler is one FIFO per
+`queueKey`, concurrency 1, `queueKey ""` = unserialized (lane runners, `autoRunner.ts:2366-2404`); boot settles ghosts
+(`bootQueuedJobs.ts` `settleRunningJobMetas`); `docker/entrypoint.sh:289-293` `exec "$@"` runs any command; the image
+locates the downloader only via `YTDLP_BIN` (`common/lib/paths.ts:269`) and curls the latest release at build time
+(`Dockerfile:~310-323`); `~/Projects/yt-dlp-patched` is a yt-dlp SOURCE tree (no built binary).
+
+## Rulings (operator, 2026-10-04; do not re-open)
+
+- Stages are independent queueable jobs driven by ON-DISK STATE (stamps), like the ingest lanes; ONE index build
+ shared by every site build; builds and deploys queue one at a time.
+- A `publish` LANE (policy, pause gate, the ingest runner pattern) AND manual per-stage buttons / "Publish now".
+- The lane may deploy PRODUCTION only for a site that opts in via `site.json` (`publish.auto: "production"`); previews
+ per site likewise; default off; private sites never deploy.
+- Two runners behind ONE stage contract: **single-container/local** (default everywhere: each stage is a child
+ process of the editor's own process, in or out of a container — the Windows target is "Docker Desktop and nothing
+ else") and **docker fan-out** (Linux hosts, opt-in, host-driven: the existing `Dockerfile.build` per-site containers).
+ **No docker-in-docker** — the editor never gets the socket. Rejected alternative, recorded: a socket mount for memory
+ isolation; answered by a heap cap on the child and serial stages.
+- The runtime image bakes python3 + pipx + pinned git-filter-repo and a PINNED wrangler, so the homepage (source mirror
+ included) and deploys run from the container; Cloudflare credentials enter via `.env`.
+- The image allows substituting yt-dlp (the operator's patched build) WITHOUT bundling it: a runtime swap by default.
+- Withdrawn private content ships TOMBSTONES, never deletions; every deploy is checked live (plain + cache-busted).
+
+## Model
+
+**Stage contract — `common/publish/stages.ts`** (pure table + runners; the only module that knows all stages):
+```ts
+type StageKind = "update-index"|"build-site"|"deploy-site"|"build-hub"|"deploy-hub"|"build-homepage"|"deploy-homepage";
+type StageRequest = { kind: StageKind; target: string /* "_index" | siteId | "_all" | "_hub" | "_homepage" */;
+ runId: string; preview?: string; to?: "pages"|"local"; runner?: "local"|"docker";
+ force?: boolean; skipArchives?: boolean; indexAfter?: number; builtAfter?: number };
+type Freshness = {state:"fresh"} | {state:"stale"; reason:string} | {state:"blocked"; reason:string};
+type Stage = { kind; label; jobKind: `publish-${StageKind}`; queueKey: "publish";
+ needs(s: PublishStatus, r: StageRequest): Freshness; // pure, over the view
+ argv(r: StageRequest): string[]; // ["stage", kind, target, ...flags]
+ run(ctx: StageContext, r: StageRequest): Promise<StageOutcome> }; // inside the child, or the CLI
+type StageOutcome = { status: "ran"|"noop"; stamp: string; summary: string };
+```
+- **Every stage job is a `runManagedCommand`** (`streamCommand.ts:198`): `<common>/node_modules/.bin/tsx
+ bin/archilyzer.ts stage <kind> <target> [flags]`, cwd `common/`, the editor's env (+ `NODE_OPTIONS=--max-old-space-
+ size=8192` for update-index). Cancel kills the child; the child traps SIGTERM and `runChildIntoLog` kills grandchildren
+ (`next build`, wrangler, docker). Exit codes: 0 ran/no-op, 1 failed, 2 usage, 3 precondition not met, 130 cancelled.
+- **The publish lock — `common/publish/stageLock.ts`**: `<exportBuildsDir>/.publish.lock` `{pid, host, kind, target,
+ since}`, taken `O_EXCL` by the stage child AND the CLI; stale when same host and the pid is dead (`processIsAlive`,
+ `bootQueuedJobs.ts:111`); a live holder makes the taker wait (poll 5 s, one log line, cancellable). This is what keeps the
+ editor and a CLI started with `docker compose exec` from building at once.
+- **Stamps — `common/publish/stamps.ts`** (atomic temp+rename; malformed = null = stale):
+```ts
+// <exportIndexDir>/stamp.json
+IndexStamp = { v:1; stampId; generation; scannedAt /*INDEX_SCANNED_AT_KEY*/; builtAt; templatesAt; commit|null;
+ index:{shortCircuited; added; changed; removed; heldChannels[]}; stats:{shortCircuited; notIndexedYet; notIndexable};
+ sites: Record<siteId,{siteFp; statsFp; inputSig}>; hubSig }
+// <exportBuildsDir>/<target>/built.json (target = siteId | "_hub" | "_homepage")
+BuiltStamp = { v:1; stampId; target; kind:"site"|"hub"|"homepage"; indexStampId; inputSig; builtAt; commit|null;
+ branch|null; runner:"local"|"docker"; audience; corpusGeneratedAt|null; files; bytes; archivesStaged; sourceCommit? }
+// <exportBuildsDir>/<target>/deployed.json
+DeployRecord = { builtStampId; builtAt; kind:"production"|"preview"|"local"; branch?; url|null; alias?; at; wrangler?;
+ liveCheck: LiveCheck|null }
+DeployedFile = { v:1; target; production?; local?; previews: Record<branch, DeployRecord> }
+LiveCheck = { at; url; plain: Probe; busted: Probe; expected: string|null;
+ verdict:"ok"|"stale-edge"|"mismatch"|"unreachable"|"skipped"; tombstones?: {path; plain; busted; ok}[] }
+Probe = { status|null; generatedAt?; cfCacheStatus?; age?; cacheControl?; error? }
+```
+- **`inputSig`** (per site, computed by update-index): sha1 over the site's `.export-index/sites/<id>/` tree, its member
+ channels' shared transcripts/subs/posts/digests trees (ignoring `manifest.json`), its stats dir and chart templates, the
+ bytes of `site.json`, and the `archiveStorage` + `social.x.visibility` settings — using compose's own `dirSignature`
+ (`compose-site.ts:593`) moved unchanged to `common/lib/dirSignature.ts`. Rule: a site is fresh exactly when compose would
+ skip everything. `hubSig` = sha1(stampId, `homepage.json`, each listed site's `siteUrl` + title).
+- **Bundle layout**: `<exportBuildsDir>/<id>/out` is THE bundle every deploy ships, whichever runner built it. A build
+ writes `<id>/out.next`, then swaps (`out → out.prev`, `out.next → out`, rm `out.prev`); a leftover `out.next` is deleted
+ first. Archive staging = `dockerSiteStagingDir` for both runners. `export/out` becomes a SYMLINK to the bundle built last
+ (older readers and `serve out` keep working; `next build` removes the link itself, never its target). The hub's bundle is
+ `_hub/out`; the homepage bundle stays `homepage/out` (the source gate withdraws from there, `build.ts:1106`), its stamps
+ in `_homepage/`.
+- **Decided: compose into the shared `export/public`, then move `out/` per site.** `next build` copies `public/` into
+ `out/` and `distDir` may not leave the project (Next docs `distDir.md`); a per-site public dir would mean replacing
+ `export/public` by a symlink (what `build-site.sh` does in its throwaway container) — breaks the dev server and the
+ worktree links on a host. The serial queue + the lock make the shared dir safe; compose's per-site cache still works
+ (`public/site.json` names the site composed last, `compose-site.ts:743-749`). On the host the move is a same-fs `rename`;
+ in the container `export/out` is an image layer and the builds dir a volume → EXDEV → `fs.cp` + swap (~1 min).
+
+## Stages
+
+| job kind | target | fresh when | the child does | writes |
+|---|---|---|---|---|
+| `publish-update-index` | `_index` | a stamp exists and nothing below is newer than `stamp.scannedAt` | `buildIndex` → `buildStats` → `build templates` in ONE child; writes the stamp (a rerun is cheap: both report `shortCircuited`) | `stamp.json` |
+| `publish-build-site` | `<id>` | no `changedChannels` since `built.builtAt`, `built.inputSig === stamp.sites[id].inputSig`, and `builtBundleProblem(<id>/out)` null | `buildSiteSteps({skipData:true})` (`build.ts:79-109`) → `builtBundleProblem(export/out)` → move out + staging → symlink. Blocked "update the index first" with no stamp | `<id>/built.json` |
+| `publish-build-site` | `_all --runner docker` | per site, as above | `build archives` on the host → `ensureBuildImage` → `runDockerBuildOne` per site (`build.ts:486-600`) → check + stamp per site. No engine → refused "the docker runner needs an engine on this host" | each `<id>/built.json` |
+| `publish-deploy-site` | `<id> [--preview b] [--to local]` | `deployed[kind/branch].builtStampId === built.stampId` | today's guards over `<id>/out` (`siteDeployProblem`, `builtSiteProblem`, `builtAudienceProblem`, `builtBundleProblem`); production refused when `built.branch` set and ≠ `main`; credential preflight; R2 upload from `<id>/.r2-staging`; pinned wrangler `--branch main` (production) / `--branch b`; live check. `--to local` copies to `ARCHILYZER_SITE_OUT` (private still refused) | `<id>/deployed.json` |
+| `publish-build-hub` | `_hub` | `built.inputSig === stamp.hubSig` | rm `public/site.json`, compose:hub (writes tombstones), `next build` `INSTANCE_MODE=hub`, `builtHubProblem`, move | `_hub/built.json` |
+| `publish-deploy-hub` | `_hub` | as deploy-site | `hubProjectProblem`, wrangler, live check + tombstone probes | `_hub/deployed.json` |
+| `publish-build-homepage` | `_homepage` | `indexStampId` current and `sourceCommit` = `main` HEAD when a repo is reachable | `buildHomepage()` (`build.ts:1071-1102`) incl. the source mirror | `_homepage/built.json` |
+| `publish-deploy-homepage` | `[--to local]` | as deploy-site | `publishedSourceProblem`, `homepageDeployArgs` (`--branch main`) | `_homepage/deployed.json` |
+
+- **What makes the index stale** (`needs()` of update-index, cheap, no LMDB): (1) any ingest job meta in `.jobs`
+ (download/transcribe/digest/normalize/import/fetch-posts/tag kinds — the drainable kinds of `jobKinds.ts`) that ENDED
+ `done` after `stamp.scannedAt` — read through the registry's archive reader, no new writer anywhere; (2) a config file
+ newer than the stamp: `tags.json`, `search-aliases.json`, `duplicates*.json`, `sites/*/site.json`, `homepage.json`, the
+ settings file, the charts config. Decided: mtimes/job completions decide WHEN to run, the index child decides WHAT
+ changed (its fingerprints stat every sidecar and are the only correct detector; LMDB `generation` cannot see new files).
+- **A site is marked stale by its channels, before any index runs** (operator, 2026-10-05): per site,
+ `changedChannels` = the member channels (by `site.json` membership, groups expanded) that have an ingest job meta
+ (same drainable kinds as above, which carry `channelSlug`) ended `done` after `built.builtAt`, plus a config change
+ (its `site.json`, `tags.json`, `search-aliases.json`, `duplicates*.json`) newer than the build. Two signals, two
+ chips: `built` shows "stale: N channels changed (slug, slug, …)" from this cheap read, and after update-index runs
+ the stamp's `inputSig` mismatch is the authoritative confirmation ("stale: data changed"). "Build all stale", Publish
+ now and the lane use `changedChannels ∪ inputSig-mismatch` to choose sites; a site with neither is skipped. The same
+ per-channel read gives the hub `changedChannels` across listed sites and the homepage its pool-level flag.
+- `--force` (manual buttons only) builds/deploys a fresh stage; Publish now and the lane never force. A code change
+ (`built.commit` differs) shows a "code newer" chip and does not make a build stale.
+
+## Lane
+
+- **`publish` is NOT added to `LANES`** (nine loops compile a channel tree per lane: `channelPriority.ts:831`,
+ `channelWriters`, `storageWatch`, `operationBatch`, `channelSnapshot`, `autoQueueSchema:224,281`). Add
+ `PIPELINE_LANES = ["publish"]` in `autoQueueTypes.ts`; widen `PauseLane` to `AutoQueueKind | "publish"`;
+ `isGateHeld`/`withGateHeld` (`pauseGates.ts:97-120`) read/write `settings.publish.held`.
+- **Settings `settings.publish`** (`settingsSchema.ts`, SETTINGS.md regenerated): `{ enabled:false, held:false,
+ checkEveryMinutes:10, refreshEveryMinutes:360, quietHours:{start,end}|null, runner:"local"|"docker" ("local"),
+ previewBranch:"preview", hub:Policy ("off"), homepage:Policy ("off") }`.
+- **Per-site policy `site.json` `publish: { auto: "off"|"build"|"preview"|"production" }`** (`siteSchema.ts`, SITE.md, the
+ site form): default off; a private site is clamped to "build"; "preview"/"production" require `cloudflareProject`.
+- **Runner `common/controller/publishRunner.ts`**: `startPublishRunner` copies `startAutoRunner`'s skeleton
+ (`autoRunner.ts:2366-2404`): `runManagedFunction` kind `auto-publish`, queueKey `""`. Wakes every `checkEveryMinutes`;
+ skips a pass when held, in quiet hours, or while any publish job is queued/running. A pass runs when the index is stale
+ and `now − stamp.builtAt ≥ refreshEveryMinutes`, or there is no stamp: `planPublishRun(status, {deploys:"policy"})`,
+ enqueued ONE stage at a time via `enqueueStage({…, background:true})`, awaiting each `done`; the gate is re-checked
+ between stages (a hold stops dispatching, never kills a stage); drain finishes the current stage. `startAutoRunnersIfEnabled`
+ also starts it; an idle boot leaves it off (`instrumentation.ts:218`).
+- **Enqueueing `common/controller/publishStages.ts`**: `enqueueStage(paths, req)` refuses a duplicate (same kind + target
+ queued/running on `publish`) with `info:true` + the existing id; spec `{kind: jobKind, slug: target, params:{runId,…req}}`
+ (`parseJobSpec` requires `slug`, `jobSpec.ts:66-82`; `channelSlug` unset). `enqueuePublishRun(paths, plan)` enqueues the
+ whole plan at once; ordering is enforced by ON-DISK preconditions, not memory: builds carry `indexAfter=runStart` when
+ update-index is in the run, deploys `builtAfter=runStart` when their build is.
+- **/jobs + boot**: `jobDetail.ts:12` shows `run <runId…6> · <target>` for publish kinds (one run reads as a group);
+ `jobKinds.ts` gets the seven `publish-*` entries, replayable (`jobReplayRegistry.ts` re-enqueues from the spec); at boot
+ (`bootQueuedJobs.ts:280`) publish kinds are treated like `sync`: cancelled "the publish lane re-derives stages from
+ on-disk state", never requeued.
+
+## Surfaces
+
+- **One state builder**: `common/controller/publishState.ts` `readPublishInputs(paths)` (stamps, sites, settings, input
+ mtimes, recent job metas, live publish jobs) → `common/views/publishStatus.ts` `buildPublishStatus(inputs, now)` (pure):
+ per target `{indexFresh, built, builtFromCurrentIndex, buildStale?: {reason:"channels"|"data"|"config";
+ changedChannels: string[]}, deployed, deployedIsBuilt, previewUrl?, liveCheck?, policy, deployProblem?, next, running?,
+ queued[]}` + `{index, lane, plan}`; `planPublishRun` pure in the same view. The
+ /sites panel, the lane, `archilyzer publish status` and `GET /api/ops/publish` all read it (one-core doctrine).
+- **/sites "Publish" panel** (`PublishPanel.tsx`, replaces "Build all sites"): one row per site + hub + homepage, chips
+ `index | built | deployed | live`, buttons Build / Deploy preview / Deploy production / Deploy local (only with
+ `ARCHILYZER_SITE_OUT`); **Publish now** = `enqueuePublishRun(plan(deploys:"policy"))` (every policy off ⇒ update-index
+ alone); **Build all stale** builds every stale site regardless of policy; consoles are `JobLane`
+ (`editor/app/sites/components/JobLane.tsx`). Pool: "Build index" keeps its name + `Build index output` label
+ (`editor/e2e/helpers.ts:457`) and now enqueues update-index; "Build stats dataset" removed; normalize/archive unchanged.
+ `/sites/<id>/publish`: Build & deploy = `enqueuePublishRun([build --force, deploy builtAfter])`; preview/production
+ split stays. New `/operations/publish` (static route beside `[id]`): enable/hold/start/stop/drain, the settings form,
+ last/next pass, the plan. `editor/app/sites/lib/publishActions.ts` is new; `buildIndexAction`/`buildStatsAction` deleted;
+ `buildAction.ts`, `deployAction.ts`, `hubActions.ts`, `homepageDeployActions.ts` become thin `enqueueStage` calls.
+- **CLI (`archilyzer.ts`)**: `publish status [--json]`, `publish index`, `publish build <id|all> [--runner docker]
+ [--force]`, `publish deploy <id|all> [--preview b] [--to local] [--force]`, `publish hub [--deploy] [--preview b]`,
+ `publish homepage [--deploy]`, `publish now` — run the stage bodies in the CLI's process under the lock; `stage …` is the
+ internal row the editor spawns. Old rows become printed aliases: `build site` = `publish index` (skipped by `--nodata`) +
+ `publish build --force`; `deploy site` = `publish deploy`; `build all` = `publish build all --runner auto`.
+- **ops: ONE route `editor/app/api/ops/publish/route.ts`** (the `lane` route's precedent; adapters only): body
+ `{verb:"index"|"build"|"deploy"|"hub"|"homepage"|"now", siteId|siteIds?, preview?, to?, force?, runner?, deploy?}` →
+ `{jobs:[{target,kind,jobId,previewUrl?}], skipped}` (`--wait` follows every id); `GET` returns `PublishStatus`;
+ `archilyzer-ops.mjs` gains `publish` in ACTIONS and `get publish`. Old routes stay as aliases (same bodies/responses;
+ `skipData` accepted and ignored).
+
+## Docker
+
+- **(a) Runtime image** (`Dockerfile` runtime-base `:304-307`): apt `python3 pipx python3-certifi python3-brotli
+ python3-websockets python3-mutagen python3-pycryptodome python3-requests`; `PIPX_HOME=/opt/pipx PIPX_BIN_DIR=/usr/local/bin
+ pipx install git-filter-repo==2.47.0` (a drift test in `buildImage.test.ts` pins it to `FILTER_REPO_PIPX_SPEC`).
+ **Wrangler pinned as an exact devDependency of `common`** (one pin for host and image; the image ships workspace
+ `node_modules`; nothing is fetched at deploy time); `wranglerBin(paths) = WRANGLER_BIN ?? <common>/node_modules/.bin/
+ wrangler` replaces `pnpm dlx` (`build.ts:382,961`); `pagesDeploy.ts` gains `WRANGLER_MAJOR` for doctor. Build args
+ `ARCHILYZER_COMMIT`/`ARCHILYZER_BRANCH` → ENV, filling the stamps' `commit`/`branch` where there is no `.git`.
+- **(b) Credentials + config dir**: `.env` carries `CLOUDFLARE_API_TOKEN`, `CLOUDFLARE_ACCOUNT_ID`, the `R2_*` keys
+ (documented in `.env.example` + `envVars.ts` → ENVIRONMENT.md); the editor inherits them, so its deploy children do.
+ Compose `x-app-env` gains `ARCHILYZER_CONFIG_DIR=/data/config/archilyzer` (`paths.ts:216` honours it): the scrub and
+ denylist files live in the config volume (`docker compose cp` them in; never printed).
+- **(c) Host one-shots: `docker compose exec editor pnpm archilyzer publish <verb>`** (or `pnpm ops publish`). NOT
+ `docker compose run --rm`: a second container means a second `export/public` and a second LMDB writer, and the lock
+ cannot see pids across namespaces. The entrypoint catch-all stays for tools.
+- **(d) Linux fan-out runner**: `settings.publish.runner:"docker"` or `--runner docker` on a HOST editor runs the
+ `build-site _all` stage above (same `built.json` stamps ⇒ deploy stages are runner-agnostic). In a container there is no
+ engine and the stage refuses — no docker-in-docker.
+- **(e) Local targets**: `docker/publish-site.sh <id>` becomes `publish index && publish build "$1" && publish deploy "$1"
+ --to local`; the `homepage` service mounts `builds` and serves `/data/builds/homepage` when non-empty, else the image's
+ baked out; `deploy-homepage --to local` fills it.
+- **(f) Source mirror from the container**: `source.ts:766` defaults `sourceRepo` to `ARCHILYZER_SOURCE_REPO`;
+ `docker-compose.source.yml` bind-mounts the host's git common dir read-only at `/data/source.git` and sets it; the
+ entrypoint adds `git config --global --add safe.directory /data/source.git`.
+- **(g) Substituting yt-dlp** — default: the image's release binary; the hook is a RUNTIME swap, no rebuild:
+ `YTDLP_BIN` in `.env` (passed by `env_file`) points at either a zipapp built on the host (`make yt-dlp` in the patched
+ checkout) bind-mounted or copied into `/data/config/bin/`, or `/usr/local/bin/yt-dlp-from-source`, a wrapper baked into
+ the image (`PYTHONPATH=${YTDLP_SOURCE_DIR:-/opt/yt-dlp-src} exec python3 -m yt_dlp "$@"`) used with
+ `docker-compose.ytdlp.yml` mounting `${YTDLP_SOURCE_HOST_DIR}:/opt/yt-dlp-src:ro`. The image already sets `YTDLP_BIN`
+ (`Dockerfile:359`), so "overridden" = differs from a new `ARCHILYZER_IMAGE_YTDLP=/usr/local/bin/yt-dlp`. Every boot
+ prints `yt-dlp: <path> <version> (image|override)` (like the `vulkan:` line; `MISSING` warning, editor still starts);
+ `update_ytdlp` (`entrypoint.sh:173`) skips with a warning when overridden. A build-time `--build-context ytdlp=…` zipapp
+ stage is a follow-up.
+- **`archilyzer doctor` adds**: `yt-dlp` (path, version, image|override, auto-update conflict), `wrangler` (binary, major
+ = `WRANGLER_MAJOR`), `cloudflare-auth` (token set or OAuth config present; never a value), `r2-keys` (only with a bucket),
+ `config-dir` (exists, writable; counts only), `export-builds` (writable, free ≥ 1.5× bundles), `publish-lock` (dead pid),
+ `index-stamp` (age; sites built from an older stamp), `source-repo` (homepage policy on, no `.git`); `filter-repo` stays.
+
+## Deploy hardening
+
+- **Credential preflight**: no token and no OAuth config → refused before wrangler ("set CLOUDFLARE_API_TOKEN in .env");
+ wrangler's `Authentication error [code: 10000]` → `[deploy] REFUSED by Cloudflare — the API token was not accepted`.
+- **Live check `common/publish/liveCheck.ts`** (injectable `fetch`): `<siteUrl | alias | deployment URL>/corpus.json`
+ plain and `?cb=<builtStampId>`, 3 tries 10 s apart; records `cf-cache-status`, `age`, `cache-control`; compares
+ `generatedAt` with `built.corpusGeneratedAt`; mismatch/stale-edge = WARNING, job still `done`; `E2E_LIVE_CHECK=skip`.
+- **Tombstones `common/publish/tombstones.ts`** while `social.x.visibility` is private: compose-site writes, per
+ withheld member (`publishedMemberSlugs`), `posts/<slug>/manifest.json` (manifest-only shape, `pageCount: 0`) and
+ `posts/<slug>/page-NNNN.json = []` for every N below the shared tree's `pageCount`; a site whose only posts were X gets
+ an empty `posts/manifest.json` (`corpus.json` still advertises no posts). compose-hub writes the same for every X channel
+ with a shared posts tree + `posts/manifest.json`, after removing `SITE_ONLY_PUBLIC_ENTRIES`. `renderHeadersFile`
+ (`headers.ts:80`) gains `noStore: string[]` → `Cache-Control: no-store` (+ CORS) for `/posts/manifest.json` and
+ `/posts/<slug>/*` on a site, `/posts/*` on the hub; the committed fixture updated. **The hub's next deploy must contain,
+ path for path:** `posts/manifest.json` (empty), `posts/thequartering-X/manifest.json` (`pageCount` 0),
+ `posts/thequartering-X/page-0000.json` (`[]`), and the `_headers` `no-store` block for `/posts/*`; deploy-hub probes each
+ plain + cache-busted (`LiveCheck.tombstones`). If the edge still serves the old objects, they expire ~2026-10-08 21:00.
+
+## Migration
+
+- Kinds no longer created keep label-only `jobKinds.ts` entries (archived metas still read): `build-index`,
+ `build-stats`, `build-export`, `build-deploy`, `build-all`, `build-deploy-all`, `deploy-export`, `build-hub`,
+ `deploy-hub`, `build-deploy-hub`, `build-homepage`, `deploy-homepage`, `build-deploy-homepage`.
+- Deleted: `runDocker*AllPhase`, `basicBuildAndDeployAll`, `BuildAllSitesButton`, `BuildSitesPanel`, the index/stats
+ parts of `BuildButtons`. ops aliases: `build-deploy {all}` = build every site + deploy each deployable site to production
+ (explicit request ⇒ policy ignored); deploying a never-built site refuses "no build of X in <dir> — archilyzer publish
+ build X"; `docker/publish-site.sh` is a wrapper.
+- Docs: PUBLISH.md rewritten around stages (replaces "deploy runs on the host" `:639-641` and the branch hazard `:85-88`);
+ RUNNING_IN_DOCKER.md (Publish, Windows `:568`, yt-dlp `:223`, fan-out `:544`); AGENTS.md runtime-container section
+ (stages run in the container; `exec`, not `run`); FACTS "The publish stages"; SITE.md, SETTINGS.md, ENVIRONMENT.md
+ regenerated. Changelog bullets: Publishing is stages · One index for every site · Publish lane · Deploys are pinned
+ and checked live · Withdrawn X posts ship tombstones · Substitute your own yt-dlp.
+
+## Slices
+
+| slice | branch | owns | after | tests |
+|---|---|---|---|---|
+| **S1** stage core | `r18/stage-core` | `common/publish/{stages,stamps,stageLock,stageRun}.ts`; `common/lib/dirSignature.ts` (+ the import line in compose-site); `build.ts` (move-out, symlink, staging); `archilyzer.ts` (`stage` row, `publish index/build/deploy/hub/homepage`, aliases) | — | stamp round-trip + tolerance; lock (stale, wait, other host); the `needs()` table; move-out with injected EXDEV; argv pins in `build.test.ts`; aliases |
+| **S2** deploy hardening | `r18/deploy-hardening` | `common/lib/pagesDeploy.ts`; `common/publish/{deployStage,liveCheck,tombstones}.ts`; `headers.ts` + fixture; `compose-site.ts`; `compose-hub.ts`; `common/package.json` + lockfile (wrangler pin); `editor/e2e/fixtures/bin/fake-wrangler.mjs`; playwright env `WRANGLER_BIN`/`E2E_LIVE_CHECK` | S1's types (interfaces above; may start in parallel) | `--branch main`; auth classifier; live-check verdicts; tombstone paths; postsVisibility integration; fake-wrangler argv sidecar + `E2E_FAKE_WRANGLER_AUTH_FAIL` |
+| **S3** state + lane | `r18/publish-lane` | `common/controller/{publishState,publishStages,publishRunner}.ts`; `common/views/publishStatus.ts`; `settingsSchema`, `siteSchema` + generated docs; `autoQueueTypes`, `pauseGates`; `jobKinds`, `jobDetail`, `bootQueuedJobs`, `jobReplayRegistry`; `instrumentation.ts`; `archilyzer.ts` `publish status/now` | S1 | view/plan purity; dedupe; preconditions; runner (hold between stages, drain, quiet hours); boot category; sanitizers |
+| **S4** surfaces | `r18/publish-surfaces` | editor `sites/**`, `operations/publish/`, `api/ops/publish` + aliases, `archilyzer-ops.mjs`, specs | S3, S2 | new `publish.spec` (chips; Build index writes the stamp; build; preview deploy through the fake wrangler; Publish now with policies off = index only; one runId on /jobs; the auth-fail sentence); `publish-lane.spec` (hold + drain via stuck-job); `ops-api` (verbs, aliases, refusals); updated `build`, `deploy-page`, `site-scope`, `site-publish-preview`, `sites-homepage`, `duplicate-shorts`, `sites-crud` |
+| **S5** docker + doctor | `r18/docker-publish` | `Dockerfile`, `docker/*`, `docker-compose*.yml` (+ `source`, `ytdlp` overlays), `.env.example`, `doctor.ts` + test, `source.ts` (env default), `envVars.ts`, RUNNING_IN_DOCKER.md | S2 (`wranglerBin`); the image half may start at once | doctor cases; `buildImage.test` drift; the compose smoke |
+| **S6** records | `r18/records` | PUBLISH.md, AGENTS.md, FACTS, `plans/release-18.md` "as shipped", CHANGELOGs | all | docs gates |
+
+**Graph:** S1 → {S2, S3} → S4; S2 → S5; all → S6. Start together: S1, S2 (against the interfaces above), S5's image half.
+
+**Rollout (operator-owned restarts; the live editor must never be booted twice):**
+1. Merge all slices (`--no-ff`, clean tree, counts-only privacy greps before each), run the gates.
+2. ONE editor rebuild + restart (`~/reports/release-17/scripts/r17-{build,restart,smoke}.sh` pattern; update the guard sha).
+3. Every site's policy still "off": `pnpm ops publish --json '{"verb":"now"}' --wait` = update-index only; poll `/`
+ every 2 s meanwhile, every answer under 5 s (the starvation is gone).
+4. `publish build jeralyzer` → `publish deploy jeralyzer --preview r18` → read the live-check verdict.
+5. `publish hub --deploy` → read the tombstone probes (this closes today's loose end: the cached X shard).
+6. Then, per site, `publish deploy <id>` (production) in today's order; then set policies (recommended: `build` for
+ all six, `preview` where previews are wanted, `production` for none until a week of lane runs reads clean).
+7. umtool/homepage untouched by the restart; `archilyzer doctor` after.
+
+## Verification
+
+- Unit + build gates per slice (common, editor, export, homepage, scripts, mcp suites; tsc; editor/export/homepage
+ builds; umtool's capped build). e2e: the editor suite with S4's spec list — wrangler never runs
+ (`WRANGLER_BIN=fake-wrangler.mjs`). Capped export build: `pnpm --filter export exec next build` in a worktree.
+- Compose smoke on Linux (project `-p r18smoke`, a FIXTURE corpus copied into `$T`, never the real one): `exec editor pnpm
+ archilyzer publish index && publish build <fixture> && publish deploy <fixture> --preview smoke`; with no token the
+ deploy is refused before wrangler; with `CLOUDFLARE_API_TOKEN=bogus` the real pinned wrangler is refused by Cloudflare;
+ both: `deployed.json` untouched, exit 1; `doctor` lists every new check; the boot log shows the `yt-dlp:` line.
+- Windows checklist (Docker Desktop, PowerShell): `copy .env.example .env` (set `WORKER_TOKEN`, `CLOUDFLARE_API_TOKEN`,
+ `CLOUDFLARE_ACCOUNT_ID`); `docker compose up -d --build`; open `http://localhost:8081` → /sites → Publish now (or
+ `docker compose exec editor pnpm archilyzer publish now`); local preview: `docker compose --profile site up -d site`,
+ `docker compose exec editor pnpm archilyzer publish deploy <id> --to local`, open `http://localhost:8080`;
+ `docker compose exec editor pnpm archilyzer doctor`; optional patched yt-dlp: set `YTDLP_SOURCE_HOST_DIR` +
+ `YTDLP_BIN=/usr/local/bin/yt-dlp-from-source`, `docker compose -f docker-compose.yml -f docker-compose.ytdlp.yml up -d`.
+- Live: rollout steps 3–5 above, each with its verdict recorded in this file's "As it went".
+
+## Open questions (settle during S1/S2, record the answer)
+
+1. Where does the hub's `s-maxage=604800` come from (no repo rule sets it)? If a Cloudflare rule overrides edge TTLs,
+ `no-store` may be ignored — the live check will show it; a purge-by-URL is not designed (no zone for `pages.dev`).
+2. Is `main` the production branch of every Pages project (jeralyzer … jasolyzer, hub, archilyzer)?
+3. Does a `generation` bump rewrite summaries of sites whose data did not change? If so `inputSig` is conservative
+ (more rebuilds, never a wrong skip).
+4. Should the lane also rebuild on a code change (`built.commit` differs)? Default no.
+5. Six per-site bundles cost disk (`export/out` is ~491 MB for the largest site); hardlink or accept.
+
+## Assumptions the operator can overturn
+
+| assumed | alternative |
+|---|---|
+| `publish` is a pipeline lane outside `LANES`, configured in `settings.publish` | add it to `LANES`/`autoQueue` and special-case the nine channel-tree loops |
+| compose into the shared `export/public`, move `out/` per site | per-site `EXPORT_PUBLIC_DIR` behind an `export/public` symlink |
+| the index child runs and is trusted; job completions + config mtimes decide when | compare LMDB `generation` to the stamp (cannot see new files) |
+| wrangler as an exact devDependency of `common` | `npm i -g wrangler@ARG` in the image only |
+| the lane enqueues one stage at a time; Publish now enqueues all at once with on-disk preconditions | an in-memory chain / a parent job |
+| Publish now obeys policies (all off = index only) | Publish now builds everything stale |
+| `exec`, not `run --rm`, for host one-shots | dedicated compose `publish` services |
+| yt-dlp swapped at runtime (`YTDLP_BIN` + wrapper or zipapp); build-time context is a follow-up | a `--build-context ytdlp=…` zipapp stage now |
+| `export/out` stays as a symlink to the last bundle built | remove `export/out` |
+| the fan-out runner is host-only, opt-in | drop the fan-out entirely (single runner) |
+
+## Follow-ups carried over (not scheduled; keep)
+- normalize/archive pool jobs still run in-process (move them to child processes).
+- `migrate-tier --reclaim` on the platter; the en/en-orig duplicate-track ruling + the `en` → 0 cues bug (212 degraded);
+ "text on platter" per-channel option (low priority, /home has 225 GB free); Syncthing folder paused.
+- Release 16/17 leftovers unchanged: "Load more results" leaf resume, `/changelog/` phone overflow, the three flaky-spec
+ fixes, the JSX entity sweep, build traces listing dot-dirs, Docker image for others (needs a GitHub home), X section
+ by-hand check, `channel-rename.spec` race, RM L5 force-release, XL Firefox store discovery, the channel export/import
+ bundle, realcandaceo's Rumble fact-check (Candace data stays local-only).
+
+## Record
+
+(Each slice adds a "### Slice <X>, as shipped" section here, before "## Rollout".)
+
+## Rollout
+
+(Steps 1–7 above; "### As it went" is written as the rollout runs.)