# e2e speed — plan (2026-10-09) The editor suite took **71 min** on Wave 0's tip (`next dev`, 733 tests, one worker, a loaded machine). A gate round runs it once and the report, umtool, export, hub and homepage suites one after another behind the same global lock, so a release's e2e gate costs most of an evening. This plan makes it a fraction of that without giving up what the serial design protects (no fixture races, no OOM). ## What was measured Sources: Wave 0's full run (`next dev`, 2026-10-09 11:27–12:38) and release 19's gate run (`E2E_MODE=start` on an editor built at the tip, 2026-10-09 evening). Durations from the `list` reporter, per test. | | | |---|---| | dev, wall | 71 min 22 s for 733 tests (711 passed); 4 min of it outside any test (boot, compile, teardown) | | start vs dev, the same 333 tests | 672 s vs 2,087 s — **0.32×** | | start, full suite (696 tests after B6) | **23 min 56 s** (681 passed, 3 failed, 12 skipped; 2026-10-09 20:40–21:04, a peer's light jobs running) — **~3× faster** | | the biggest movers dev → start | `audio-check-scenarios` 350 s → 46 s; `auto-queue` 260 s → 25 s (on the tests covered); `attribution` 81 s → 7 s | | fixed sleeps (`waitForTimeout`, literal `setTimeout`) | ~33 s in the whole suite — not a lever | | real product timers | `pacing.spec:193` 60 s (a real 60 s cooldown), `publish-lane.spec:78` 52 s, `fetch-window.spec:337` 39 s, `channel-storage.spec:872` 30 s — the same in both modes | | per-test setup | 519 `resetData` + 188 `writeSettings`; the fixture tree is 528 KB / 53 files — cheap; ~1,200 `invalidate-cache` round-trips per run | | memory, one set of test servers (start) | editor ~0.3 GB, export (`next dev`) ~0.45 GB, Playwright + Chromium ~0.7 GB — **~1.5 GB** | | the sharded docker route | last measured 2026-07-30 at 392 tests: 7.1 min on 4 shards vs 24–30 serial; rebuilds the image (and a `next build`) on every source change; never measured at today's size | The dev → start gap is Next's per-route compile under polling: a polled route recompiles in dev, and the specs poll (170+ `expect.poll` sites). ## Why it is serial (and stays safe) `editor/playwright.config.ts` names it: one data root that 519 `resetData` calls destroy and recreate, one `test-settings.json` the export server also reads, and the job registry, worker pool, scheduler, auto-runner and auto-queue state as `globalThis` singletons inside the one Next server. A directory per worker does not fix the last; **a server per worker does.** Separately, every package's suite takes the same `e2e` lock (`scripts/queue-lock.mjs`, default name), though their ports (`common/lib/ports.mjs`) and fixture files are disjoint — the lock exists so two runs never drive each other's `reuseExistingServer`. ## Slices, in order ### S1 — start mode by default, with a build that cannot be stale `[unit]` + one full run - `pnpm e2e` (editor) runs against `next start`. Before it starts, a **build stamp** decides whether `editor/.next` is the tree under test: `next build` writes `editor/.next/e2e-stamp.json` (HEAD + a hash of the dirty files under `editor/`, `common/`); the runner rebuilds through the heavy slot when the stamp differs, and says so. "`E2E_MODE=start` serves a stale build" (FACTS) is then impossible rather than a rule to remember. - `E2E_MODE=dev` stays for iterating on one spec (no rebuild per edit). - The editor suite's own export server likewise serves a build when the stamp matches (it is `next dev` today). - Fix the start-only reds tonight's run names (each re-run alone first); `widget.spec`'s Dev Tools case already passes only under start. - umtool, export (default, hub) and homepage get the same `E2E_MODE` switch and stamp. - **Measured:** editor 71 → 24 min (release 19's gate run, by hand). The other suites: similar ratios expected, not measured. ### S2 — a timing record per run `[unit]` - Each suite adds Playwright's `json` reporter beside `list` (`test-results/timings.json`), and `scripts/e2e-timings.mjs` prints per-spec totals and the delta against the last run on the same branch. - Gives every later slice its before/after, and catches a spec that quietly grows a 60 s wait. ### S3 — product timers behind test knobs `[unit]` + the named specs - The four timers above (≈3 min per run) take a test-only duration through `E2E_SERVER_ENV`, as the audio check already does (`E2E_AUDIO_CHECK_INTERVAL_MS`). Each knob is declared in `common/lib/envVars.ts` as test-only. - The spec still asserts the behaviour (backs off, defers, resumes); only the clock shrinks. ### S4 — the suites run side by side `[unit]` + a two-suite run - Each package's suite queues under its own name (`--name e2e-`), so editor, export, umtool and homepage can run at once; the same suite never runs twice. - The heavy slot becomes **memory-weighted**: a job declares its estimate (an e2e suite ~1.5 GB, a `next build` 5 GB — its cap), and is admitted while `MemAvailable − Σ(running estimates) ≥ HEAVY_MIN_FREE_MB`. One build at a time stays a hard rule. The two OOMs this slot exists for were builds and renders, not suites. - `pnpm e2e:all` runs every suite through this queue and prints one summary (the release gate in one command). - **Expected:** the gate's e2e wall time becomes the longest suite, not the sum. ### S5 — a server per worker in the editor suite (the real fix) `[unit]` + full runs at 1, 2 and 3 workers - **Paths through one helper first:** 104 spec files name `test-transcripts` / `test-settings.json` directly. A mechanical slice routes every one through `e2e/helpers.ts` (`testPaths()`), and `dev:test` / `start:test` take `TRANSCRIPTS_DIR`, `SETTINGS_FILE` and the fake binaries' log paths as `${VAR:-default}`. One worker, same results — a no-behaviour-change commit. - **Then the servers:** a worker-scoped Playwright fixture starts that worker's editor (`next start`, port = base + worker index), its export server and its Ollama stub, each over the worker's own data root and settings file, and stops them at the end. The config's single `webServer` goes; `globalSetup`'s stray reaping covers every worker's ports. - `workers` defaults to what memory allows at start (`floor((MemAvailable − floor) / 1.5 GB)`, capped at 3 on this 8-core machine; `E2E_WORKERS` overrides); one worker remains a supported mode. - Specs that time wall-clock behaviour under contention are tagged `@serial` and run in a final one-worker project (the timing specs that fail under load today are the first candidates). - **Expected:** ~24 → ~10–12 min at 2–3 workers (estimate; the 2026-07 sharded run gave 2.1–3.4× on 4 shards). - This makes the docker-sharded route a remote-machine option only (a VM over `docker context`), not the local one. ### S6 — test economy, continued `[unit]` - B6 moved 37 request/response specs to route unit tests. The same audit over the rest: a spec that drives no page (only `request.*` and file reads) becomes a `route.test.ts`. ## Order and gates S1 → S2 first (small, and tonight's run is S1's measurement). S3 any time after S2. S4 and S5 are independent of each other; S5's path refactor can start at once. Every slice: the full editor suite once, its result and wall time recorded here under "As it went", load flakes re-run alone before they count. The heavy slot and the e2e queue apply throughout; no slice weakens the memory floor. ## As it went ### S1–S3, as shipped (2026-10-10, Track E of the overnight batch, branch `r20/e2e-speed`) **S1 — start mode by default, the build stamped.** - `scripts/e2e-stamp.mjs` (`ensure | check | build `). The stamp is a FINGERPRINT, not HEAD: the index's blob id of every file the build reads, and git's blob id of the working-tree bytes of every dirty or untracked one, over the package, `common/` and `pnpm-lock.yaml`, minus what no build reads (`/e2e/`, the package's `playwright.config.ts`, `*.test.*`, `CHANGELOG.md`). An edit and its commit fingerprint the same; a plans-only commit costs no rebuild. A stale stamp names where the tree moved (`common, editor changed since the build (4 uncommitted files there, e.g. common/controller/fetchWindows.ts)`). The build runs through `queue-lock.mjs --heavy` and `systemd-run --user --scope -p MemoryMax=5G -p MemorySwapMax=0` (inside an e2e run the slot is already held and passes through), and the stamp written is the tree the build STARTED from. - **The build has its own directory** — ruled: the primary checkout's `.next` is what the live editor and umtool serve. Editor: `.next/e2e` through `E2E_NEXT_DIST_DIR` (a one-line `distDir` in `editor/next.config.ts`; inside the ignored `.next/`, so Tailwind never scans it). umtool: `.next-e2e-start` (its config already reads `NEXT_DIST_DIR`; ignored by `umtool/.next-*/`). The stamp is `editor/.next/e2e/e2e-stamp.json`, not `editor/.next/e2e-stamp.json` as briefed. A `next build` with the custom dist dir leaves `tsconfig.json` alone and rewrites only the ignored `next-env.d.ts`. - `editor/playwright.config.ts` and `umtool/playwright.config.ts` run `ensure` before any server starts (`E2E_BUILD_CHECKED` keeps a worker's second load from checking again). `E2E_MODE=dev` builds nothing. `CI` (the sharded image, which builds its own `.next`) keeps today's path. - **The export (default, hub) and homepage suites take the switch and stay `next dev`**, printing so for `E2E_MODE=start` — ruled: they are static exports, several export specs (header, transcript-downloads, brand, site-branding, …) rewrite the fixture site mid-run and assert the next page, and the homepage's fixture summary and publish are read only outside a production build. **The editor suite's export server stays `next dev` too** — ruled: its two spec files took 42 s in all in release 19's start run, under one export build's cost, and a static build would bake in the `test-settings.json` export-search.spec rewrites. - **Flakes from release 19's start run:** `channel-work.spec:208` was a real race — `fs.cp` mkdirs each directory after finding it absent, and a write still landing from the previous spec's work created `test-transcripts/channels` in between. `resetData` now retries the clear-and-copy on `EEXIST` (four attempts). `sites-crud.spec:392` (the first Save's status not seen in 5 s) and `whisper.spec:184` (the clear job's output still "Waiting for output…" at 15 s) are recorded as load flakes: both pass alone, and nothing in either names a race. - **Stale-safety, by hand** (the S1 tree, before its commit): first run `e2e build: rebuilding editor — no build stamp here yet` → `done in 54s, stamped f534007e1d7a`, `channel-work.spec` 11 passed. Then a comment appended to `common/lib/project.ts`: `check editor` → `stale — common changed since the build …`; the runner → `e2e build: rebuilding editor — common changed since the build …` (stopped before the suite). Killing that build left NO stamp (`next build` empties its dist dir first), so the next run rebuilt rather than trusting a half-written build: `done in 84s, stamped f534007e1d7a` — the same key for the same tree. **S2 — a timing record per run.** Every suite's config writes Playwright's `json` report beside `list` (`test-results*/timings.json`; editor, umtool, export default/hub/report/2origin, homepage). `scripts/e2e-timings.mjs []` (default: the editor's) totals it per spec file (every result, retries included), prints the files slowest first, each against the last run on the same branch that ran it, and the wall time against the last run of the same spec files; it records the run in the git common dir (`.git/e2e-timings//.json`, the last 30), untracked and shared by every worktree. **S3 — the four timers.** Two were product timers and take a test-only duration from `E2E_SERVER_ENV`, declared `test` in `common/lib/envVars.ts`: - `E2E_BACKOFF_BASE_MS=20000` — the first rate-limit cooldown, read in `nextBackoff` (`common/jobs/platformBackoff.ts`); the cap and the hold arithmetic (`FAILS_TO_REACH_CAP`) keep the real constants. 20 s, not less — ruled: rumble-sweep, metadata-scan-softblock and fetch-window assert a cooldown still in force a page load after it was recorded. - `E2E_CLIP_WINDOW_GAP_MS=2000` — the pause between two clip-window fetches (`common/controller/fetchWindows.ts`: `testGapMs()` beside `CLIP_WINDOW_PLATFORM_MIN_GAP_SECONDS`, and one line, `testGapMs() ??`, in the gap expression). fetch-window.spec now asserts the one pause the batch owes, from the job log. The other two were never product timers: publish-lane.spec's 45 s and channel-storage.spec's 25 s were the stuck-job holder's own `releaseAfterMs`, sized for a cold dev server. `GET /api/test/stuck-job?release=` finishes a holder now, detached from the request exactly as `releaseAfterMs` does (one helper, `armDetached`); the two specs release it once they have seen what it holds (publish-lane also asserts the stage is still queued at that point). Before/after, start mode, same worktree, per test (and per spec file, from `e2e-timings.mjs`): | test | before | after | file before → after | |---|---|---|---| | `pacing.spec:195` (a live 429 backs youtube off …) | 1.1 m | 22.4 s | 66.5 s → 24.0 s | | `publish-lane.spec:78` (a hold mid-stage …) | 49.7 s | 5.6 s | 51.2 s → 7.0 s | | `fetch-window.spec:337` (a batch skips what is cached …) | 34.1 s | 3.1 s | 45.2 s → 6.4 s | | `channel-storage.spec:872` (a job cancelled but still stopping …) | 28.1 s | 2.8 s | 85.4 s → 35.3 s | Four files: 248 s → 73 s of test time. Commits: `d1e80e42` S1 · `5ef76aa0` S2 · `977c8e72` S3 · `0f509eac` a comment. Gates: tsc clean before each commit (3 runs, 3.6–4.7 min); `scripts/e2e-stamp.test.mjs` 9/9, `scripts/e2e-timings.test.mjs` 8/8; `test:scripts` 725 passed + 3 skipped (728); editor unit 218/218; common 3635/3636 — the one, `fetchPosts.test.ts`'s drain-mid-page case, passes 3/3 alone and touches nothing here; `platformBackoff.test.ts` + `fetchWindows.test.ts` 33/33 with a case for each knob; `docs env --check` clean. e2e, single files in the foreground: `channel-work.spec` 11/11 (twice, each after a build); umtool `article-notes` + `usage` 15/15 in start mode (first umtool build 43 s); export `header.spec` 33/33 under `E2E_MODE=start` (prints the dev line); homepage `source.spec` 6/6; the four S3 files 4/4 + 24/24 before and after. With the knobs on: `rumble-sweep` + `metadata-scan-softblock` 2/2 and `rate-limit` + `ops-api` 18/18 (start; ops-api's `detached` holder still detached); `pacing` + `publish-lane` 4/4 under `E2E_MODE=dev` (1.0 m; pacing:195 21.7 s), and the stamp still fresh after it — a `next dev` in `.next` leaves `.next/e2e` alone. A commit that changed only `editor/playwright.config.ts` (`0f509eac`) cost no rebuild. Not done here, by the brief: the full suites (the orchestrator runs them, start then dev). S4–S6.