Archilyzer · Source

archilyzer

Archilyzer
git clone https://archilyzer.pages.dev/source/archilyzer.git
Log | Files | Refs | README | LICENSE

commit 2c850a8dc7dfd95f8e9ace36b916250c974680c2
parent afc918a01e571703f50a8f12ee3335a9ea39f5d6
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date:   Fri,  9 Oct 2026 21:07:01 -0400

plans: release 19's landing as it went (gates, the two fixes, start-mode e2e 24 min); e2e-speed planned

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Diffstat:
Aplans/e2e-speed.md | 99+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Mplans/release-19.md | 34++++++++++++++++++++++++++++++++++
2 files changed, 133 insertions(+), 0 deletions(-)

diff --git a/plans/e2e-speed.md b/plans/e2e-speed.md @@ -0,0 +1,99 @@ +# e2e speed — plan (2026-10-09) + +The editor suite took **71 min** on Wave 0's tip (`next dev`, 733 tests, one worker, a loaded machine). A gate +round runs it once and the report, umtool, export, hub and homepage suites one after another behind the same +global lock, so a release's e2e gate costs most of an evening. This plan makes it a fraction of that without +giving up what the serial design protects (no fixture races, no OOM). + +## What was measured + +Sources: Wave 0's full run (`next dev`, 2026-10-09 11:27–12:38) and release 19's gate run (`E2E_MODE=start` on an +editor built at the tip, 2026-10-09 evening). Durations from the `list` reporter, per test. + +| | | +|---|---| +| dev, wall | 71 min 22 s for 733 tests (711 passed); 4 min of it outside any test (boot, compile, teardown) | +| start vs dev, the same 333 tests | 672 s vs 2,087 s — **0.32×** | +| start, full suite (696 tests after B6) | **23 min 56 s** (681 passed, 3 failed, 12 skipped; 2026-10-09 20:40–21:04, a peer's light jobs running) — **~3× faster** | +| the biggest movers dev → start | `audio-check-scenarios` 350 s → 46 s; `auto-queue` 260 s → 25 s (on the tests covered); `attribution` 81 s → 7 s | +| fixed sleeps (`waitForTimeout`, literal `setTimeout`) | ~33 s in the whole suite — not a lever | +| real product timers | `pacing.spec:193` 60 s (a real 60 s cooldown), `publish-lane.spec:78` 52 s, `fetch-window.spec:337` 39 s, `channel-storage.spec:872` 30 s — the same in both modes | +| per-test setup | 519 `resetData` + 188 `writeSettings`; the fixture tree is 528 KB / 53 files — cheap; ~1,200 `invalidate-cache` round-trips per run | +| memory, one set of test servers (start) | editor ~0.3 GB, export (`next dev`) ~0.45 GB, Playwright + Chromium ~0.7 GB — **~1.5 GB** | +| the sharded docker route | last measured 2026-07-30 at 392 tests: 7.1 min on 4 shards vs 24–30 serial; rebuilds the image (and a `next build`) on every source change; never measured at today's size | + +The dev → start gap is Next's per-route compile under polling: a polled route recompiles in dev, and the specs +poll (170+ `expect.poll` sites). + +## Why it is serial (and stays safe) + +`editor/playwright.config.ts` names it: one data root that 519 `resetData` calls destroy and recreate, one +`test-settings.json` the export server also reads, and the job registry, worker pool, scheduler, auto-runner and +auto-queue state as `globalThis` singletons inside the one Next server. A directory per worker does not fix the +last; **a server per worker does.** Separately, every package's suite takes the same `e2e` lock +(`scripts/queue-lock.mjs`, default name), though their ports (`common/lib/ports.mjs`) and fixture files are +disjoint — the lock exists so two runs never drive each other's `reuseExistingServer`. + +## Slices, in order + +### S1 — start mode by default, with a build that cannot be stale `[unit]` + one full run +- `pnpm e2e` (editor) runs against `next start`. Before it starts, a **build stamp** decides whether + `editor/.next` is the tree under test: `next build` writes `editor/.next/e2e-stamp.json` (HEAD + a hash of the + dirty files under `editor/`, `common/`); the runner rebuilds through the heavy slot when the stamp differs, and + says so. "`E2E_MODE=start` serves a stale build" (FACTS) is then impossible rather than a rule to remember. +- `E2E_MODE=dev` stays for iterating on one spec (no rebuild per edit). +- The editor suite's own export server likewise serves a build when the stamp matches (it is `next dev` today). +- Fix the start-only reds tonight's run names (each re-run alone first); `widget.spec`'s Dev Tools case already + passes only under start. +- umtool, export (default, hub) and homepage get the same `E2E_MODE` switch and stamp. +- **Measured:** editor 71 → 24 min (release 19's gate run, by hand). The other suites: similar ratios expected, not measured. + +### S2 — a timing record per run `[unit]` +- Each suite adds Playwright's `json` reporter beside `list` (`test-results/timings.json`), and + `scripts/e2e-timings.mjs` prints per-spec totals and the delta against the last run on the same branch. +- Gives every later slice its before/after, and catches a spec that quietly grows a 60 s wait. + +### S3 — product timers behind test knobs `[unit]` + the named specs +- The four timers above (≈3 min per run) take a test-only duration through `E2E_SERVER_ENV`, as the audio check + already does (`E2E_AUDIO_CHECK_INTERVAL_MS`). Each knob is declared in `common/lib/envVars.ts` as test-only. +- The spec still asserts the behaviour (backs off, defers, resumes); only the clock shrinks. + +### S4 — the suites run side by side `[unit]` + a two-suite run +- Each package's suite queues under its own name (`--name e2e-<pkg>`), so editor, export, umtool and homepage + can run at once; the same suite never runs twice. +- The heavy slot becomes **memory-weighted**: a job declares its estimate (an e2e suite ~1.5 GB, a `next build` + 5 GB — its cap), and is admitted while `MemAvailable − Σ(running estimates) ≥ HEAVY_MIN_FREE_MB`. One build at a + time stays a hard rule. The two OOMs this slot exists for were builds and renders, not suites. +- `pnpm e2e:all` runs every suite through this queue and prints one summary (the release gate in one command). +- **Expected:** the gate's e2e wall time becomes the longest suite, not the sum. + +### S5 — a server per worker in the editor suite (the real fix) `[unit]` + full runs at 1, 2 and 3 workers +- **Paths through one helper first:** 104 spec files name `test-transcripts` / `test-settings.json` directly. A + mechanical slice routes every one through `e2e/helpers.ts` (`testPaths()`), and `dev:test` / `start:test` take + `TRANSCRIPTS_DIR`, `SETTINGS_FILE` and the fake binaries' log paths as `${VAR:-default}`. One worker, same + results — a no-behaviour-change commit. +- **Then the servers:** a worker-scoped Playwright fixture starts that worker's editor (`next start`, port = base + + worker index), its export server and its Ollama stub, each over the worker's own data root and settings file, + and stops them at the end. The config's single `webServer` goes; `globalSetup`'s stray reaping covers every + worker's ports. +- `workers` defaults to what memory allows at start (`floor((MemAvailable − floor) / 1.5 GB)`, capped at 3 on this + 8-core machine; `E2E_WORKERS` overrides); one worker remains a supported mode. +- Specs that time wall-clock behaviour under contention are tagged `@serial` and run in a final one-worker + project (the timing specs that fail under load today are the first candidates). +- **Expected:** ~24 → ~10–12 min at 2–3 workers (estimate; the 2026-07 sharded run gave 2.1–3.4× on 4 shards). +- This makes the docker-sharded route a remote-machine option only (a VM over `docker context`), not the local one. + +### S6 — test economy, continued `[unit]` +- B6 moved 37 request/response specs to route unit tests. The same audit over the rest: a spec that drives no + page (only `request.*` and file reads) becomes a `route.test.ts`. + +## Order and gates + +S1 → S2 first (small, and tonight's run is S1's measurement). S3 any time after S2. S4 and S5 are independent of +each other; S5's path refactor can start at once. Every slice: the full editor suite once, its result and wall time +recorded here under "As it went", load flakes re-run alone before they count. The heavy slot and the e2e queue +apply throughout; no slice weakens the memory floor. + +## As it went + +(Each slice adds its record here.) diff --git a/plans/release-19.md b/plans/release-19.md @@ -434,4 +434,38 @@ running); only the changed files' tests ran. No e2e (class `[none]`/`[unit]`). - fetch-windows: the cross-run platform gap is in memory, and dedupe is within one request (FACTS; A5). - C2 is regenerated after each Track A merge that adds an action or a CLI row (`pnpm archilyzer docs cli`). +## Landing, as it went (2026-10-09 evening) + +`r19/integration` from `8eabed56` (releases 19 A1–A5/B1/B3–B6/C1–C5, 21 D1/D3/D4b, 22 E1–E6, the timeline feeds, +the session-link hotfix, the X `--post-range` fix), gated in one quiet window (peers asked to hold renders, builds, +e2e and lanes; the live editor's transcription lane held through `pnpm ops lane`). + +**Light gates** (`$T/r19-*.log`): `pnpm typecheck` clean (108 s). `pnpm test` 2 red of 3,634 in common, both real: +- `fetchOlderPosts` "a limit stops the walk part-way" — its inline fake gallery-dl, and the editor e2e's + `fake-gallery-dl.mjs`, still capped on `--range` after `8eabed56` moved the cap to `--post-range`; +- `architecture.test.ts` — `lib/report/exportSlides.test.ts` imported `components/report/slides/Slide`, a new + lib → components back-edge. The test moved to `components/report/exportSlides.test.ts` (it is the parity test + between the string slides and `Slide.tsx`); the allow-list stays as it was. + +Both fixed in `904b3d91`; common then 3,634/3,634. Editor 218, scripts 708 (+3 skipped), mcp 293, export 118, +homepage 23. `docs env|files|cli --check` and `settings example --check` 0. + +**Builds** (each through `pnpm heavy`, 5 GB scope): editor 121 s, export 32 s (over the committed fixture +compose), homepage `build:nodata` 24 s, umtool 40 s (corpus linked). + +**e2e:** +- Editor suite, **`E2E_MODE=start`** on the editor built at the tip: 681 passed, 3 failed, 12 skipped, **23 min + 56 s** (Wave 0's dev-mode run: 71 min). The three — `channel-work.spec:208` (`resetData`'s EEXIST race), + `sites-crud.spec:392`, `whisper.spec:184` — pass alone in start mode (13 s) and in dev (54 s): load flakes. + Measurements and the plan that follows from them: [e2e-speed.md](e2e-speed.md). +- Export report suite: 42 passed, 3 failed — release 22's slides/overview specs, first run (release-22.md, "Fix + round"). +- umtool article specs: 18 passed, 3 failed — the same (release-22.md). + +**Privacy, counts only, `main..r19/integration`** (diff added lines / commit messages): the operator's name, +email local part and guarded suffix 0/0; home paths 0/0; `2bbb46d0` 0/0; corpus data paths 0/0; author and +committer one identity. Session trailers: 0 in any file; 40 commit messages written before the rule (11:19–14:07) +carry one, as 2,060 of main's 2,720 do — the mirror strips them (`SESSION_LINK_RULES`) and its audit refuses any +left, so the source check after the fast-forward is the gate. + ## Rollout