commit 2c850a8dc7dfd95f8e9ace36b916250c974680c2
parent afc918a01e571703f50a8f12ee3335a9ea39f5d6
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Fri, 9 Oct 2026 21:07:01 -0400
plans: release 19's landing as it went (gates, the two fixes, start-mode e2e 24 min); e2e-speed planned
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Diffstat:
2 files changed, 133 insertions(+), 0 deletions(-)
diff --git a/plans/e2e-speed.md b/plans/e2e-speed.md
@@ -0,0 +1,99 @@
+# e2e speed — plan (2026-10-09)
+
+The editor suite took **71 min** on Wave 0's tip (`next dev`, 733 tests, one worker, a loaded machine). A gate
+round runs it once and the report, umtool, export, hub and homepage suites one after another behind the same
+global lock, so a release's e2e gate costs most of an evening. This plan makes it a fraction of that without
+giving up what the serial design protects (no fixture races, no OOM).
+
+## What was measured
+
+Sources: Wave 0's full run (`next dev`, 2026-10-09 11:27–12:38) and release 19's gate run (`E2E_MODE=start` on an
+editor built at the tip, 2026-10-09 evening). Durations from the `list` reporter, per test.
+
+| | |
+|---|---|
+| dev, wall | 71 min 22 s for 733 tests (711 passed); 4 min of it outside any test (boot, compile, teardown) |
+| start vs dev, the same 333 tests | 672 s vs 2,087 s — **0.32×** |
+| start, full suite (696 tests after B6) | **23 min 56 s** (681 passed, 3 failed, 12 skipped; 2026-10-09 20:40–21:04, a peer's light jobs running) — **~3× faster** |
+| the biggest movers dev → start | `audio-check-scenarios` 350 s → 46 s; `auto-queue` 260 s → 25 s (on the tests covered); `attribution` 81 s → 7 s |
+| fixed sleeps (`waitForTimeout`, literal `setTimeout`) | ~33 s in the whole suite — not a lever |
+| real product timers | `pacing.spec:193` 60 s (a real 60 s cooldown), `publish-lane.spec:78` 52 s, `fetch-window.spec:337` 39 s, `channel-storage.spec:872` 30 s — the same in both modes |
+| per-test setup | 519 `resetData` + 188 `writeSettings`; the fixture tree is 528 KB / 53 files — cheap; ~1,200 `invalidate-cache` round-trips per run |
+| memory, one set of test servers (start) | editor ~0.3 GB, export (`next dev`) ~0.45 GB, Playwright + Chromium ~0.7 GB — **~1.5 GB** |
+| the sharded docker route | last measured 2026-07-30 at 392 tests: 7.1 min on 4 shards vs 24–30 serial; rebuilds the image (and a `next build`) on every source change; never measured at today's size |
+
+The dev → start gap is Next's per-route compile under polling: a polled route recompiles in dev, and the specs
+poll (170+ `expect.poll` sites).
+
+## Why it is serial (and stays safe)
+
+`editor/playwright.config.ts` names it: one data root that 519 `resetData` calls destroy and recreate, one
+`test-settings.json` the export server also reads, and the job registry, worker pool, scheduler, auto-runner and
+auto-queue state as `globalThis` singletons inside the one Next server. A directory per worker does not fix the
+last; **a server per worker does.** Separately, every package's suite takes the same `e2e` lock
+(`scripts/queue-lock.mjs`, default name), though their ports (`common/lib/ports.mjs`) and fixture files are
+disjoint — the lock exists so two runs never drive each other's `reuseExistingServer`.
+
+## Slices, in order
+
+### S1 — start mode by default, with a build that cannot be stale `[unit]` + one full run
+- `pnpm e2e` (editor) runs against `next start`. Before it starts, a **build stamp** decides whether
+ `editor/.next` is the tree under test: `next build` writes `editor/.next/e2e-stamp.json` (HEAD + a hash of the
+ dirty files under `editor/`, `common/`); the runner rebuilds through the heavy slot when the stamp differs, and
+ says so. "`E2E_MODE=start` serves a stale build" (FACTS) is then impossible rather than a rule to remember.
+- `E2E_MODE=dev` stays for iterating on one spec (no rebuild per edit).
+- The editor suite's own export server likewise serves a build when the stamp matches (it is `next dev` today).
+- Fix the start-only reds tonight's run names (each re-run alone first); `widget.spec`'s Dev Tools case already
+ passes only under start.
+- umtool, export (default, hub) and homepage get the same `E2E_MODE` switch and stamp.
+- **Measured:** editor 71 → 24 min (release 19's gate run, by hand). The other suites: similar ratios expected, not measured.
+
+### S2 — a timing record per run `[unit]`
+- Each suite adds Playwright's `json` reporter beside `list` (`test-results/timings.json`), and
+ `scripts/e2e-timings.mjs` prints per-spec totals and the delta against the last run on the same branch.
+- Gives every later slice its before/after, and catches a spec that quietly grows a 60 s wait.
+
+### S3 — product timers behind test knobs `[unit]` + the named specs
+- The four timers above (≈3 min per run) take a test-only duration through `E2E_SERVER_ENV`, as the audio check
+ already does (`E2E_AUDIO_CHECK_INTERVAL_MS`). Each knob is declared in `common/lib/envVars.ts` as test-only.
+- The spec still asserts the behaviour (backs off, defers, resumes); only the clock shrinks.
+
+### S4 — the suites run side by side `[unit]` + a two-suite run
+- Each package's suite queues under its own name (`--name e2e-<pkg>`), so editor, export, umtool and homepage
+ can run at once; the same suite never runs twice.
+- The heavy slot becomes **memory-weighted**: a job declares its estimate (an e2e suite ~1.5 GB, a `next build`
+ 5 GB — its cap), and is admitted while `MemAvailable − Σ(running estimates) ≥ HEAVY_MIN_FREE_MB`. One build at a
+ time stays a hard rule. The two OOMs this slot exists for were builds and renders, not suites.
+- `pnpm e2e:all` runs every suite through this queue and prints one summary (the release gate in one command).
+- **Expected:** the gate's e2e wall time becomes the longest suite, not the sum.
+
+### S5 — a server per worker in the editor suite (the real fix) `[unit]` + full runs at 1, 2 and 3 workers
+- **Paths through one helper first:** 104 spec files name `test-transcripts` / `test-settings.json` directly. A
+ mechanical slice routes every one through `e2e/helpers.ts` (`testPaths()`), and `dev:test` / `start:test` take
+ `TRANSCRIPTS_DIR`, `SETTINGS_FILE` and the fake binaries' log paths as `${VAR:-default}`. One worker, same
+ results — a no-behaviour-change commit.
+- **Then the servers:** a worker-scoped Playwright fixture starts that worker's editor (`next start`, port = base +
+ worker index), its export server and its Ollama stub, each over the worker's own data root and settings file,
+ and stops them at the end. The config's single `webServer` goes; `globalSetup`'s stray reaping covers every
+ worker's ports.
+- `workers` defaults to what memory allows at start (`floor((MemAvailable − floor) / 1.5 GB)`, capped at 3 on this
+ 8-core machine; `E2E_WORKERS` overrides); one worker remains a supported mode.
+- Specs that time wall-clock behaviour under contention are tagged `@serial` and run in a final one-worker
+ project (the timing specs that fail under load today are the first candidates).
+- **Expected:** ~24 → ~10–12 min at 2–3 workers (estimate; the 2026-07 sharded run gave 2.1–3.4× on 4 shards).
+- This makes the docker-sharded route a remote-machine option only (a VM over `docker context`), not the local one.
+
+### S6 — test economy, continued `[unit]`
+- B6 moved 37 request/response specs to route unit tests. The same audit over the rest: a spec that drives no
+ page (only `request.*` and file reads) becomes a `route.test.ts`.
+
+## Order and gates
+
+S1 → S2 first (small, and tonight's run is S1's measurement). S3 any time after S2. S4 and S5 are independent of
+each other; S5's path refactor can start at once. Every slice: the full editor suite once, its result and wall time
+recorded here under "As it went", load flakes re-run alone before they count. The heavy slot and the e2e queue
+apply throughout; no slice weakens the memory floor.
+
+## As it went
+
+(Each slice adds its record here.)
diff --git a/plans/release-19.md b/plans/release-19.md
@@ -434,4 +434,38 @@ running); only the changed files' tests ran. No e2e (class `[none]`/`[unit]`).
- fetch-windows: the cross-run platform gap is in memory, and dedupe is within one request (FACTS; A5).
- C2 is regenerated after each Track A merge that adds an action or a CLI row (`pnpm archilyzer docs cli`).
+## Landing, as it went (2026-10-09 evening)
+
+`r19/integration` from `8eabed56` (releases 19 A1–A5/B1/B3–B6/C1–C5, 21 D1/D3/D4b, 22 E1–E6, the timeline feeds,
+the session-link hotfix, the X `--post-range` fix), gated in one quiet window (peers asked to hold renders, builds,
+e2e and lanes; the live editor's transcription lane held through `pnpm ops lane`).
+
+**Light gates** (`$T/r19-*.log`): `pnpm typecheck` clean (108 s). `pnpm test` 2 red of 3,634 in common, both real:
+- `fetchOlderPosts` "a limit stops the walk part-way" — its inline fake gallery-dl, and the editor e2e's
+ `fake-gallery-dl.mjs`, still capped on `--range` after `8eabed56` moved the cap to `--post-range`;
+- `architecture.test.ts` — `lib/report/exportSlides.test.ts` imported `components/report/slides/Slide`, a new
+ lib → components back-edge. The test moved to `components/report/exportSlides.test.ts` (it is the parity test
+ between the string slides and `Slide.tsx`); the allow-list stays as it was.
+
+Both fixed in `904b3d91`; common then 3,634/3,634. Editor 218, scripts 708 (+3 skipped), mcp 293, export 118,
+homepage 23. `docs env|files|cli --check` and `settings example --check` 0.
+
+**Builds** (each through `pnpm heavy`, 5 GB scope): editor 121 s, export 32 s (over the committed fixture
+compose), homepage `build:nodata` 24 s, umtool 40 s (corpus linked).
+
+**e2e:**
+- Editor suite, **`E2E_MODE=start`** on the editor built at the tip: 681 passed, 3 failed, 12 skipped, **23 min
+ 56 s** (Wave 0's dev-mode run: 71 min). The three — `channel-work.spec:208` (`resetData`'s EEXIST race),
+ `sites-crud.spec:392`, `whisper.spec:184` — pass alone in start mode (13 s) and in dev (54 s): load flakes.
+ Measurements and the plan that follows from them: [e2e-speed.md](e2e-speed.md).
+- Export report suite: 42 passed, 3 failed — release 22's slides/overview specs, first run (release-22.md, "Fix
+ round").
+- umtool article specs: 18 passed, 3 failed — the same (release-22.md).
+
+**Privacy, counts only, `main..r19/integration`** (diff added lines / commit messages): the operator's name,
+email local part and guarded suffix 0/0; home paths 0/0; `2bbb46d0` 0/0; corpus data paths 0/0; author and
+committer one identity. Session trailers: 0 in any file; 40 commit messages written before the rule (11:19–14:07)
+carry one, as 2,060 of main's 2,720 do — the mirror strips them (`SESSION_LINK_RULES`) and its audit refuses any
+left, so the source check after the fast-forward is the gate.
+
## Rollout