commit e89a360e943897413544f213f266e72e5bce0c9d
parent 6c175b0d312cb6aaf6accbc065064f7952707797
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Tue, 6 Oct 2026 12:54:54 -0400
e2e: the alerts are filtered past Next's route announcer; the job page's first compile gets 30 s; changelog: the surfaces fold into Publishing is stages and The publish lane
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Diffstat:
4 files changed, 12 insertions(+), 8 deletions(-)
diff --git a/editor/CHANGELOG.md b/editor/CHANGELOG.md
@@ -1,14 +1,14 @@
# Changelog
## [Unreleased]
-- **The publish lane.** Publishing can run itself: turn on `publish.enabled` in settings and the lane checks every `checkEveryMinutes` (10) whether the index is stale; when it is — and its last update is at least `refreshEveryMinutes` (360) old — it updates it, then builds every site whose channels changed or whose data the new index moved, one stage at a time on the `publish` queue. What it may do with a site is the site's own: `site.json` `publish.auto` is `off` (the default: left alone), `build`, `preview` (built and deployed to the preview branch `publish.previewBranch`) or `production`; the hub and the homepage have `publish.hub` and `publish.homepage`. A private site is only ever built, and a site needs its Cloudflare Pages project before it may deploy. Hold the lane and the stage running finishes and no next one starts; quiet hours (`publish.quietHours`) do the same; Drain finishes the stage and ends the runner. The lane never forces a stage: a stage that finds its target current does nothing. On /jobs every stage of one run reads `run <id> · <target>`, and a stage still queued when the editor restarts is cancelled, never re-queued — the lane works out again what is stale from what is on disk. `archilyzer publish now` runs the same plan from the command line, one stage after another in its own process.
+- **The publish lane.** Publishing can run itself: turn it on at **/operations/publish** (the runner's Start, Drain and Stop, the hold, and the lane's settings; or `publish.enabled` in settings) and the lane checks every `checkEveryMinutes` (10) whether the index is stale; when it is — and its last update is at least `refreshEveryMinutes` (360) old — it updates it, then builds every site whose channels changed or whose data the new index moved, one stage at a time on the `publish` queue. What it may do with a site is the site's own — the **Publish policy** on the site's settings form, `site.json` `publish.auto` —: `off` (the default: left alone), `build`, `preview` (built and deployed to the preview branch `publish.previewBranch`) or `production`; the hub and the homepage have `publish.hub` and `publish.homepage`. A private site is only ever built, and a site needs its Cloudflare Pages project before it may deploy. Hold the lane and the stage running finishes and no next one starts; quiet hours (`publish.quietHours`) do the same; Drain finishes the stage and ends the runner. The lane never forces a stage: a stage that finds its target current does nothing. On /jobs every stage of one run reads `run <id> · <target>`, and a stage still queued when the editor restarts is cancelled, never re-queued — the lane works out again what is stale from what is on disk. `archilyzer publish now` runs the same plan from the command line, one stage after another in its own process.
- **One index for every site.** The index is updated once and every site, the hub and the homepage are built from it; `archilyzer publish status` says, per site, whether its build is current — "stale: 3 channels changed (a, b, c)" as soon as a download, transcription or digest on one of its channels finishes, before any index runs; "stale: data changed" once the index has run and the site's data moved; "stale: config changed" after its site.json, tags or aliases changed — and whether what is deployed is that build, with a build made by older code marked "code newer" but not stale.
- **Deploys are pinned and checked live.** wrangler is an exact dependency of the workspace (4.147.0), so a deploy runs the version installed with the code instead of whatever `pnpm dlx` fetched that day, and every deploy names its branch: production is `--branch main`, never taken from the checkout it ran in (where a "production" deploy from a feature branch used to land as a preview). The publish stages' deploy (release 18) refuses before wrangler runs when there is no Cloudflare credential at all — "set CLOUDFLARE_API_TOKEN in .env" — and says "REFUSED by Cloudflare — the API token was not accepted" when Cloudflare rejects one; it refuses a production deploy of a build made from a branch other than `main`. After each deploy it reads `corpus.json` at the site's address twice, as a visitor would and cache-busted, and records the verdict: ok, stale-edge (the deployment is right, Cloudflare's edge still serves an older copy), mismatch, or unreachable. A verdict short of ok is a warning in the log; the deploy itself succeeded. What each target last shipped, where, and how it read is kept in `deployed.json` beside its build.
- **Withdrawn X posts ship tombstones.** While X posts are private, a public site's build no longer just leaves an X channel's posts out: at every path they were served from it ships an empty stand-in — the channel's posts manifest with no pages, and an empty page for each page the channel has — served uncached. The hub, which carries no posts, ships the same for every X channel a public site carries, with an empty posts manifest; a channel only on a private site, or on no site, is never named on the hub. Leaving a path out of a deploy does not take it off Cloudflare's edge, which kept serving a withdrawn copy for up to a week; a changed object at the same path replaces it. The hub's deploy reads each of those paths back.
-- **Publishing is stages, from the command line: `archilyzer publish`.** `publish index` updates the index — the LMDB index, the stats datasets and the chart templates, in one child process with an 8 GB heap — and writes an index stamp (`export/.export-index/stamp.json`) naming, for each site, a signature of everything that site's build reads. `publish build <id|all>` builds a site from that index (no data phase of its own) into its own bundle, `export/.export-builds/<id>/out`, and stamps it (`built.json`); a site whose bundle already matches the index is a no-op unless `--force`. `publish deploy <id|all> [--preview <branch>] [--to local]` ships that bundle — to Cloudflare Pages, or with `--to local` into the directory the docker `site` service serves — and records the deploy (`deployed.json`); deploying the same build again is a no-op unless `--force`. `all` passes over private sites and, to Pages, sites with no Pages project; any other site it cannot deploy is a failure, said after the rest are tried. `publish hub [--deploy]` and `publish homepage [--deploy]` do the same for the hub (`_hub/out`) and the homepage. A stage whose input is not there says so and exits 3: "update the index first", "no build of jeralyzer — archilyzer publish build jeralyzer". Production refuses a bundle built on a branch other than `main`, or with no branch recorded (a detached checkout; an image sets `ARCHILYZER_BRANCH`) — a preview of it is fine. Exit codes: 0 done or nothing to do, 1 failed, 2 usage, 3 precondition not met, 130 cancelled.
+- **Publishing is stages, from the command line: `archilyzer publish`.** `publish index` updates the index — the LMDB index, the stats datasets and the chart templates, in one child process with an 8 GB heap — and writes an index stamp (`export/.export-index/stamp.json`) naming, for each site, a signature of everything that site's build reads. `publish build <id|all>` builds a site from that index (no data phase of its own) into its own bundle, `export/.export-builds/<id>/out`, and stamps it (`built.json`); a site whose bundle already matches the index is a no-op unless `--force`. `publish deploy <id|all> [--preview <branch>] [--to local]` ships that bundle — to Cloudflare Pages, or with `--to local` into the directory the docker `site` service serves — and records the deploy (`deployed.json`); deploying the same build again is a no-op unless `--force`. `all` passes over private sites and, to Pages, sites with no Pages project; any other site it cannot deploy is a failure, said after the rest are tried. `publish hub [--deploy]` and `publish homepage [--deploy]` do the same for the hub (`_hub/out`) and the homepage. A stage whose input is not there says so and exits 3: "update the index first", "no build of jeralyzer — archilyzer publish build jeralyzer". Production refuses a bundle built on a branch other than `main`, or with no branch recorded (a detached checkout; an image sets `ARCHILYZER_BRANCH`) — a preview of it is fine. Exit codes: 0 done or nothing to do, 1 failed, 2 usage, 3 precondition not met, 130 cancelled. In the editor, **/sites has a Publish panel** in place of "Build all sites" and the hub's and the homepage's build sections: a row per site, the hub and the homepage, each with four chips — index, built, deployed, live — and Build, Deploy preview, Deploy production (and Deploy local where the container serves one); **Publish now** runs what the lane would, **Build all stale** builds every stale site whatever its policy, and the plan Publish now would run is listed above the rows. Every button is stages on the `publish` queue, followed in one log, and a manual Build or Deploy always runs (the index is updated first when it is stale). The Pool's **Build index** is the index stage, and **Build stats dataset is gone**: the stats are part of it. A site's Publish tab is stages too, and its "Last deployed" is what that site last shipped and how its live check read. Over HTTP, `pnpm ops publish` takes `{"verb": "index" | "build" | "deploy" | "hub" | "homepage" | "now" | "stale"}` and `pnpm ops get publish` is the status; `build-index`, `build-site`, `build-deploy`, `deploy-site`, `build-hub`, `deploy-hub`, `build-homepage` and `deploy-homepage` still answer as before, as stages.
- **One publish at a time on a machine.** Every stage takes `export/.export-builds/.publish.lock`; a second one — an `archilyzer publish` beside the editor, say — waits for it, saying once whom it waits for, and Ctrl-C ends the wait. A lock left by a process that is gone is taken over. A cancelled stage takes the whole process tree it started with it (`next build`'s workers, wrangler, docker).
- **`export/out` is now a link to the bundle built last.** Each site, and the hub, keeps its own bundle, so building one site no longer replaces another's; `export/out` points at whichever was built most recently, so `serve out` and anything else that read it keeps working.
-- **`build site`, `build all` and `deploy site` are aliases of the publish commands** and print what they run: `build site <id>` is `publish index` (skipped with `--nodata`) then `publish build <id> --force`; `build all` is `publish index` then `publish build all --runner auto` (containers when an engine answers, else one site at a time on the host); `deploy site <id>` is `publish deploy <id>`, which now ships the site's own bundle and refuses a site never built that way. `publish build all --runner docker` builds every stale site in containers on a Linux host and refuses with "the docker runner needs an engine on this host" where there is none.
+- **`build site`, `build all` and `deploy site` are aliases of the publish commands** and print what they run: `build site <id>` is `publish index` (skipped with `--nodata`) then `publish build <id> --force`; `build all` is `publish index` then `publish build all --runner auto` (containers when an engine answers, else one site at a time on the host); `deploy site <id>` is `publish deploy <id>`, which now ships the site's own bundle and refuses a site never built that way; `deploy hub` and `deploy homepage` are `publish hub --deploy-only` and `publish homepage --deploy-only`. `publish build all --runner docker` builds every stale site in containers on a Linux host and refuses with "the docker runner needs an engine on this host" where there is none.
- **Substitute your own yt-dlp in Docker.** Point `YTDLP_BIN` at a zipapp you built, or set `YTDLP_SOURCE_HOST_DIR` to a yt-dlp checkout and start with `docker-compose.ytdlp.yml`: the image runs it with its own python, and nothing is rebuilt. Every editor boot logs `yt-dlp: <path> <version> (image|override)` (`MISSING` when it does not run; the editor still starts), and `YTDLP_AUTO_UPDATE` updates the image's yt-dlp only, warning instead of touching yours.
- **The Docker image can publish.** It carries python, `pipx` and a pinned `git-filter-repo`, so the homepage's `/source` mirror builds in the container; `docker-compose.source.yml` mounts your repository read-only for it, and the scrub rules and denylist live in the config volume (`/data/config/archilyzer`). Cloudflare and R2 credentials come from `.env`. Run publish commands with `docker compose exec editor pnpm archilyzer …`, not `run --rm`. The `homepage` service serves a local deploy from the builds volume once there is one. RUNNING_IN_DOCKER.md has a Windows checklist.
- **`archilyzer doctor` checks what a publish needs.** Which yt-dlp runs (the image's, the host's or an override, and whether it runs), whether the Cloudflare token and the R2 keys are set (never their values; R2 only when a bucket is configured) — judged exactly as a deploy judges them —, the wrangler a deploy runs (the pinned one or your `WRANGLER_BIN`, and that it starts and is the expected major), free space for the site bundles, the publish lock (free, held by a running stage, or left by one that is gone — with the command to clear it; never cleared for you), the index stamp's age and which sites were built from an older one, the repository the source mirror reads, the private config dir, and whether this Node is new enough for the pinned wrangler (deploys need 22).
diff --git a/editor/e2e/jobs.spec.ts b/editor/e2e/jobs.spec.ts
@@ -31,8 +31,10 @@ test("kicks off the index update, lists it, and tails its log", async ({
// Click into the most recent job. The job-id link is the first link in the row.
await buildRow.getByRole("link").first().click();
+ // 30 s: under next dev the job page compiles on its first visit, and in a
+ // long run the index stage's child is still winding down beside it.
await expect(page.getByLabel("Job log")).toContainText("Done", {
- timeout: 10_000,
+ timeout: 30_000,
});
const status = await page.getByLabel("Job status").textContent();
expect(["done", "running"]).toContain((status ?? "").trim());
diff --git a/editor/e2e/publish-lane.spec.ts b/editor/e2e/publish-lane.spec.ts
@@ -55,7 +55,8 @@ test("the settings form saves settings.publish, and Start says why when the lane
// Off: Start is refused, in the runner's words.
await page.getByRole("button", { name: "Start publish lane" }).click();
- await expect(page.getByRole("alert")).toContainText("The publish lane is off");
+ // Filtered: Next's route announcer is an alert too.
+ await expect(page.getByRole("alert").filter({ hasText: "The publish lane is off" })).toBeVisible();
const form = page.locator('form[data-settings-block="publish"]');
await form.getByLabel("Check every (minutes)").fill("5");
diff --git a/editor/e2e/publish.spec.ts b/editor/e2e/publish.spec.ts
@@ -250,9 +250,10 @@ test("the site form's publish policy: preview without a Pages project is refused
await expect(policy).toHaveValue("off");
await policy.selectOption("preview");
await page.getByRole("button", { name: /save site/i }).click();
- await expect(page.getByRole("alert")).toContainText(
- 'publish.auto "preview" deploys the site, and it has no cloudflareProject',
- );
+ // Filtered: Next's route announcer is an alert too.
+ await expect(
+ page.getByRole("alert").filter({ hasText: 'publish.auto "preview" deploys the site, and it has no cloudflareProject' }),
+ ).toBeVisible();
const site = await readJson<Record<string, unknown>>("test-transcripts/sites/noproj/site.json");
expect("publish" in site).toBe(false);