# De-flake the e2e suites Pinned to `d1acca7` (`one-core/phase-1`). Written 2026-09-08. ## Context Phase 1 of one-core shipped on `one-core/phase-1` (`7f294df` → `d1acca7`, closing e2e 515/515). Across that job's sixteen e2e runs one editor spec reddened once without a code cause (`video-page.spec.ts:216`), one attribution spec timed out once in a run that was later killed, and memory records three export-site specs that flake on the baseline (the two "no FOUC" theme tests and ask-chat "Stop aborts mid-sweep"). The operator asked to fix the flake and any others we can. A read-only investigation (2026-09-08) root-caused three of the four; the fourth is unattributed and gets a bounded experiment. Branch: commit on top of `one-core/phase-1`. Nothing here touches the lane model. The same commits cherry-pick cleanly onto `main`, since every file involved exists there. ## Findings, one per flake ### 1. `editor/e2e/video-page.spec.ts:216` "Delete directory wrong-id confirmation" — root-caused, fixture + helper Failure: `pathExists(".../channels/test-youtube/data/20240101_test1234567")` is `false` at `:237` right after test `:196` ("Delete file") in the same file. The directory is not deleted, it is **renamed**: every snapshot generation calls `reconcileVideoDirs` (`common/controller/channelSnapshot.ts:606-611` → `common/controller/reconcileVideoDirs.ts:133-145`), which renames `data/` to `data/` when they differ. The fixture's `metadata.info.json` for that video has `webpage_url` ending `watch?v=test1234567`, the only directory in the whole fixture tree whose canonical id ≠ directory name. Test `:196`'s delete-file action ends with `requestChannelSnapshot` (`videoActions.ts:413`), arming the debounced regen; `resetData` (`editor/e2e/helpers.ts:34-63`) does `rm` + `cp` of the fixture FIRST and only then calls `/api/test/invalidate-cache`, so a regen that fires in the gap renames the freshly copied directory. The page still renders because `app/channels/[slug]/videos/[id]/page.tsx:44-54` swallows the readdir failure. Solo runs pass because nothing arms the scheduler beforehand. Product behaviour is correct; the fixture is inconsistent with it and the reset ordering is backwards. Fix: - `editor/e2e/fixtures/test-transcripts/one-youtube-channel-with-data/channels/test-youtube/data/20240101_test1234567/metadata.info.json`: set `id` and `webpage_url` so the canonical id equals the directory name (`watch?v=20240101_test1234567`). Do NOT rename the directory to `test1234567` — `video-page.spec.ts:148` creates a different video by that id in the same channel. Nothing asserts this fixture's `webpage_url` (the one `watch?v=test1234567` literal, at `:156`, is the unrelated created video). Grep the editor e2e tree for any other reader of this fixture's `id` before committing. - `editor/e2e/helpers.ts:45`: call `GET /api/test/invalidate-cache` BEFORE the `rm`, keeping the existing call at `:63`. Quiesce first, then mutate the tree. Both calls are idempotent; this also subsumes the ENOTEMPTY retry note at `:35-44` (update the comment). - `editor/e2e/helpers.ts:25-32` `fileExists`: re-throw when `err.code !== "ENOENT"` so a transient errno never reads as "gone". - `plans/STATE.md` (the "One known flake, reported and not fixed" paragraph): the note says nothing deletes the directory; correct it to "nothing deletes it, `reconcileVideoDirs` renamed it" and mark it fixed. ### 2. `export/e2e/theme.spec.ts:40,56` and `theme-family.spec.ts:34` "no FOUC" — test-side timing `page.reload({waitUntil: "commit"})` resolves before the inline `` script in `common/components/ThemeScript.tsx:26-47` is guaranteed to have run, so the following `page.evaluate` can read `data-theme: null`. Product is fine. Fix: the script sets `d.setAttribute('data-theme-ready','1')` in the same synchronous `try` block (just before the `catch`); each of the three tests gates its read on `page.waitForFunction(() => document.documentElement.dataset.themeReady === "1")` after the reload. The marker and the theme attributes are set by the same statement sequence and no React has run before a head script, so the test still proves "set before hydration". ### 3. `export/e2e/ask-chat.spec.ts:961` "corpus sweep: Stop aborts mid-sweep" — test-side timing The route mock delays batch 2 with a wall-clock `setTimeout(r, 500)` at `:985`; the test must click Stop (`:1027`) inside that window, and on a loaded box batch 2 lands first. Fix: replace the timer with a latch the test controls — a deferred promise declared above the route; the batch-2 branch awaits it; the test resolves it immediately after the Stop click at `:1027`. Assertions at `:1032-1035` unchanged. ### 4. `editor/e2e/attribution.spec.ts:349` "Run from the video page runs that video and no other" — unattributed One timeout (≈96 s = setup + the 90 s `toContainText("1 done")` wait at `:361-364`) in a run later killed with exit 137; the ✘ is 54 tests before the kill, so not a kill victim, but the kill suppressed the reporter detail. Passed in every later full run. Ruled out: lost pre-hydration click (`StreamActionLog.tsx:184` disables until hydrated) and a stale ollama stub (port preflight aborts on a bound 11435). Bounded experiment, not a fix: run `attribution.spec.ts` alone with `--repeat-each 10` from the e2e worktree, editor server stdout captured to a file. If it reproduces, the stream log's last line says whether the unit was never dispatched (slot starvation from the previous test) or dispatched and never answered (stub); fix accordingly and add the finding here. If 10/10 pass, record it in `plans/STATE.md` as unreproduced with the evidence and stop. **Result (2026-09-08): it reproduced — 2 failures in 10 — and it is NEITHER of those two. Not fixed; the evidence is below so the next attempt starts from it.** The unit was dispatched and the stub answered. Every failure froze on exactly the same three lines: the batch header, `Attribute attrvid0002: 120 cues → 1 chunk(s), text-only.`, and `ollama qwen2.5:7b: 100 in / 50 out tokens …`. Four separate 10× runs, six failures, identical content each time. What the failure snapshot adds: `attribution-text freshness` already reads **`current`**, so `attribution.json` HAD been written — and `attributeOneVideo` logs its closing `Attribute (text-only): N speaker(s)…` line *immediately* after `await writeAttribution(...)`, with nothing between the two statements. So the job got past its write and then produced nothing further for 90 s. The editor server keeps answering `/api/pulse` throughout with an UNCHANGING `rev`, so the event loop is alive and the job's registry record never moves: **the job never reaches a terminal state.** `runOperationBatch` never returns, so `operationJobs.ts`'s single summary line is never emitted, so the test's wait is never satisfied. It is a hang in the batch, somewhere after the unit's write. **Two fixes were tried against the panel and BOTH failed, which is itself the evidence that the panel is not the problem.** (a) Reconciling `StreamActionLog`'s log against the job's log FILE once the stream ends: 2/10, unchanged. (b) Reconciling from the 1 s status poll on a terminal job status, and on either finish route: 1/10 then 2/10 — noise. Neither could find anything the panel did not already have, because the file does not have it either. Both were reverted; the component is untouched. If someone picks this up: the surviving suspects are inside `runOperationBatch` between `runOperationUnit` returning and `runPool` draining — `concurrentRunner.ts`'s documented "a transiently-zero limit ... does NOT terminate" wait at a zero `laneLimit`, and `taskHooks.ts`'s `onLog` progress parsers, which run on every line the unit emits. Instrument the job's registry record, not the log panel. ### 4, as shipped — the panel was remounted, and the batch was innocent (`fae99f1`) Root-caused under the plan in [`attribution-batch-hang.md`](attribution-batch-hang.md). Instrumented at `fea2996` (`console.error`, ISO timestamps, the job id, never through the job's own `onLog`, editor webServer `stdout`/`stderr` piped): **3 failures in 25**, all identical. The decisive trace, one job: ``` [TRACE 2026-09-08T20:25:21.598Z] onLog job=01M21B4PDQZC6YWGZH38WYV37N n=1 closed=false :: Backfill attribution-channel: attribution-text over 1 video [TRACE 2026-09-08T20:25:21.611Z] onLog job=01M21B4PDQZC6YWGZH38WYV37N n=3 closed=false :: ollama qwen2.5:7b: 100 in / 50 out tokens in 0s wall [TRACE 2026-09-08T20:25:21.661Z] attributeOne.post-write attrvid0002 [TRACE 2026-09-08T20:25:21.661Z] onLog job=01M21B4PDQZC6YWGZH38WYV37N n=4 closed=false :: Attribute attrvid0002 (text-only): 2 speaker(s), 2 segment(s) [TRACE 2026-09-08T20:25:21.662Z] runPool returned attribution-channel [TRACE 2026-09-08T20:25:21.662Z] onLog job=01M21B4PDQZC6YWGZH38WYV37N n=5 closed=false :: Backfill attribution-channel: 1 done, 0 already current, 0 failed [TRACE 2026-09-08T20:25:21.663Z] fn.finally job=01M21B4PDQZC6YWGZH38WYV37N status=done Error: locator.getAttribute: Test timeout of 120000ms exceeded. - waiting for getByLabel('Run Speaker names (from the transcript) output') ``` The whole job took **65 ms**. Every line reached `onLog` with `closed=false`, the summary included, and the record finalized `done`. **The trace killed all five hypotheses the plan ranked, H1's mechanism included.** H2: `limit()` never read 0 (`pool.wait-capacity target=1` throughout). H3: `runOne`'s `finally` timestamps are all present. H4: finalize ran. H5: no signal aborted. And H1 as WRITTEN — an early SSE teardown, to be fixed by polling the job's log file until terminal — is dead too: `stream.cancel` fired **zero** times in 240 instrumented tests (40 of this spec) and every line enqueued `closed=false`. All that held of H1 was its DISPOSITION: the server is innocent and the failure is client-side. The real cause is a sixth mechanism, outside the plan's table, and **the fix the plan prescribed for H1 was not shipped and would not have worked** — see below. **The sixth mechanism: React unmounted the panel.** The log was not stale — the log ELEMENT was gone. Both bodies returned two fragments of different SHAPES, each with `RunOne` at a different index in the two branches: | body | no record on disk | with a record | RunOne moves | |---|---|---|---| | `DiarizationBody` | `[p, RunOne]` | `[dl, p, RunOne]` | index 1 → 2 | | `AttributionBody` | `[p, RunOne]` | `[dl, ul, p, RunOne]` | index 1 → 3 | Static JSX fragment children compile to one array, so `reconcileChildrenArray` runs: the index fast-loop breaks at 0 (`p` vs `dl`) and the map pass then finds `p`/`ul` at index 1 where `RunOne` used to be. **A positional TYPE MISMATCH on an UNKEYED child is what destroyed it** — not the move as such: React preserves state across a KEYED move, so a `key` alone would have matched the old fiber at its new index, and a fixed position alone would never have reached the map pass. Either was sufficient. With `log` empty and `running` false the remounted component renders no `
` at all, which is exactly what the failure snapshot
shows: the record body, the Run button, and no log. The refresh that lands the record is the
one `StreamActionLog` fires itself at the end of a run (`:163`), so the panel is wiped at the
instant the summary line reaches it, and the spec could only pass in the ~200 ms window
before the refresh landed.

**This retires two claims made here on 2026-09-08 and both were measurement artifacts**: "the
job never reaches a terminal state" (it reached it in 65 ms) and "the log FILE has no more
than the panel does" (the file had everything). It also explains why the two panel-side
reconciles measured no change, and why the plan's prescribed H1 fix would have measured no
change either — no amount of reconciling helps a component React has just destroyed and
rebuilt with empty state.

Fix: both bodies are ONE two-child list, the record body a sibling of the button, `RunOne`
last AND keyed in both branches — belt and brace, since each alone would have done. No DOM
change, no server change, `StreamActionLog` untouched. The durable rule this generalises to
(and three pre-existing, unfixed instances of it in `VideoPanel.tsx`) is in `plans/FACTS.md`.
Test: the spec asserts the panel STILL contains "1 done" AFTER the freshness pill reads
"current" — the pill is the proof the refresh landed, which makes it deterministic rather than
a second race. At `fea2996` that assertion is red **5/5** (`element(s) not found`); at
`fae99f1` `attribution.spec.ts --repeat-each 20` is **20/20**.

**Follow-up:** the three `VideoPanel.tsx` instances this section left unfixed are closed in
`c19098b` (specs `b207e84`), with a fourth found in `PerFileTranscribeRow` and the review's
button/alert/timeout fixes in `12d1778` — same rule, the shape to copy is in `FACTS.md`. Full
suite at that tip: **518/518**.

## Not changed, deliberately

- `reconcileVideoDirs` keeps mutating during snapshot generation; that is the product's
  design. The e2e hazard is closed by quiescing before the fixture copy, and the STATE.md note
  names the class ("a debounced write survives a test boundary until the next reset").
- `workers: 1`, `fullyParallel: false`, `retries: 0` stay as documented in
  `editor/playwright.config.ts:42-72`. No `retries`, `test.slow`, or serial-mode markers are
  added anywhere — the point is determinism, not tolerance.

## Verification

- `tsc --noEmit` in common and export (ThemeScript change); `pnpm --filter yt-dlp-transcript-common test`.
- Stress the fixed specs from a worktree (the primary checkout cannot run editor e2e because
  of the untracked `editor/content` symlink; worktree recipe in `plans/FACTS.md`'s one-core
  sections — copy a composed fixture site into the worktree's `export/public`, e.g. from
  `.claude/worktrees/duplicates-page/export/public`):
  - editor: `pnpm e2e -- video-page.spec.ts --repeat-each 10` (the full file, so `:196`
    precedes `:216` every time) — 10/10. If `--repeat-each` is swallowed after `--`, use
    Playwright's config/CLI another way (e.g. `pnpm exec playwright test ... --repeat-each 10`
    with the queue lock taken the way `pnpm e2e` takes it) and say which you used.
  - editor: `pnpm e2e -- attribution.spec.ts --repeat-each 10` with server stdout captured
    (the item-4 experiment).
  - export: `theme.spec.ts theme-family.spec.ts ask-chat.spec.ts --repeat-each 10` — all green
    (use the export package's own e2e entry; `e2e:2origin` is known-broken on main and is not
    used).
- Full editor e2e once at the end from the worktree, detached behind the queue lock, watched
  with Monitor (not polled): expect 515/515.
- One commit per flake (fixture+helper; theme; ask-chat; STATE.md + the item-4 record), each
  with the stress result in its message.