commit c1c563dc262bfcebcdf8b5a269381b63c5491195
parent 49e71c0fb4a0cd9b3e698651c523218c1d414948
Author: I Mean I'm Just Saying <imeanimjustsaying@kiwifarms.st>
Date: Tue, 8 Sep 2026 09:11:26 -0400
plans: what the four flakes really were, including the one that is still open
Corrects the "one known flake, reported and not fixed" paragraph, which said
nothing deletes that video directory: nothing did, `reconcileVideoDirs` RENAMED
it. The three fixed flakes are recorded with the class they share — a debounced
write survives a test boundary until the next reset — and the fixture/reset,
theme-marker and ask-chat-latch fixes are named against it.
The fourth was the unattributed attribution timeout, and it gets the honest
entry: it REPRODUCED (~2 runs in 10 under `--repeat-each 10`, four runs and six
failures), it is NOT fixed, and it is neither of the two things the plan
guessed. The unit is dispatched and the stub answers; every failure freezes on
the same three lines with the freshness pill already reading "current", so the
sidecar was written and `attributeOneVideo`'s closing line — the very next
statement — never arrived; and the server keeps answering `/api/pulse` with an
unchanging rev, so the job never reaches a terminal state at all. The hang is
in the batch, after the unit's write.
Two panel-side fixes were tried and reverted, and that is recorded too, because
it is the evidence that rules the panel out: reconciling `StreamActionLog`
against the job's log FILE when the stream ends measured 2/10, and doing it from
the status poll measured 1/10 then 2/10. Neither found anything the panel did
not have, because the file does not have it either. `StreamActionLog` is
untouched. The note names where to look next and says to instrument the job's
registry record rather than the log.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Diffstat:
2 files changed, 60 insertions(+), 4 deletions(-)
diff --git a/plans/STATE.md b/plans/STATE.md
@@ -71,10 +71,37 @@ instead — which needs one extra step, a composed fixture site copied into its
so the export webServer answers 200 instead of 500. **`git add -A` stages that symlink**, and
a commit made that way carries the panic into every worktree of it: add by path.
-**One known flake, reported and not fixed:** `video-page.spec.ts:216` ("Delete directory
-wrong-id confirmation shows an error") asserts a video directory still exists after a refused
-delete. It red once in 1.3's full run and passed in 1.4's and 1.5's, and passes 20/20 re-run
-alone on the same commit; nothing in this phase deletes a video directory.
+**Three of the four e2e flakes are fixed** (`plans/deflake-e2e.md`, `a32612a` onward), and the
+class is worth remembering: **a debounced write survives a test boundary
+until the next reset.** `video-page.spec.ts:216` ("Delete directory wrong-id confirmation")
+asserted a video directory still existed after a refused delete and found it gone — nothing
+deleted it, `reconcileVideoDirs` RENAMED it. Snapshot generation moves `data/<name>` to
+`data/<extractVideoId(webpage_url)>` when the two differ, and the
+`one-youtube-channel-with-data` fixture was the only directory in the tree where they did;
+the previous test armed the debounced regen, and `resetData` copied the fixture back in
+BEFORE it invalidated, so the regen renamed the fresh copy. The fixture's id now matches its
+directory and `resetData` quiesces first (that route's `resetSnapshotScheduler()` cancels the
+pending regen), so the two halves of the hazard are both closed. The export "no FOUC" theme
+tests read `data-theme` after a `waitUntil: "commit"` reload that does not promise the head
+script has run — the script now marks itself `data-theme-ready` and they wait for it. And
+ask-chat's "Stop aborts mid-sweep" raced a 500 ms route timer; the test holds batch 2 open on
+a latch it releases itself. No retries, `test.slow`, or serial markers were added anywhere.
+
+**The fourth REPRODUCED, is NOT fixed, and is not what the plan guessed.**
+`attribution.spec.ts` "Run from the video page runs that video and no other" fails ~2 runs in 10
+under `--repeat-each 10` (four runs, six failures), always waiting the full 90 s for the batch's
+closing "1 done" line. The unit IS dispatched and the stub DOES answer: every failure freezes on
+the same three lines, and the failure snapshot shows the freshness pill already reading
+"current" — so `attribution.json` was written, and `attributeOneVideo` logs its closing line
+immediately after that write with nothing in between. Meanwhile the server keeps answering
+`/api/pulse` with an unchanging `rev`, so **the job never reaches a terminal state**:
+`runOperationBatch` does not return and `operationJobs.ts`'s summary line is never emitted. The
+hang is in the BATCH, after the unit's write — not in the log panel. Two panel-side fixes were
+tried (reconcile against the job's log file when the stream ends; reconcile from the status poll
+on a terminal status) and both measured 1-2/10, i.e. no change: the log FILE has no more than
+the panel does. Both reverted; `StreamActionLog` is untouched. Next attempt should instrument
+the job's registry record and look at `concurrentRunner.ts`'s "a transiently-zero limit does NOT
+terminate" wait and `taskHooks.ts`'s per-line progress parsers.
**The last full suite: 514 passed, 1 failed of 515, 25.1 min** from a worktree of `2133d94`.
The red was slice 1.5's own new band assertion and it was right — writing the two new
diff --git a/plans/deflake-e2e.md b/plans/deflake-e2e.md
@@ -86,6 +86,35 @@ line says whether the unit was never dispatched (slot starvation from the previo
dispatched and never answered (stub); fix accordingly and add the finding here. If 10/10 pass,
record it in `plans/STATE.md` as unreproduced with the evidence and stop.
+**Result (2026-09-08): it reproduced — 2 failures in 10 — and it is NEITHER of those two.
+Not fixed; the evidence is below so the next attempt starts from it.**
+
+The unit was dispatched and the stub answered. Every failure froze on exactly the same three
+lines: the batch header, `Attribute attrvid0002: 120 cues → 1 chunk(s), text-only.`, and
+`ollama qwen2.5:7b: 100 in / 50 out tokens …`. Four separate 10× runs, six failures, identical
+content each time.
+
+What the failure snapshot adds: `attribution-text freshness` already reads **`current`**, so
+`attribution.json` HAD been written — and `attributeOneVideo` logs its closing
+`Attribute <id> (text-only): N speaker(s)…` line *immediately* after `await
+writeAttribution(...)`, with nothing between the two statements. So the job got past its write
+and then produced nothing further for 90 s. The editor server keeps answering `/api/pulse`
+throughout with an UNCHANGING `rev`, so the event loop is alive and the job's registry record
+never moves: **the job never reaches a terminal state.** `runOperationBatch` never returns, so
+`operationJobs.ts`'s single summary line is never emitted, so the test's wait is never
+satisfied. It is a hang in the batch, somewhere after the unit's write.
+
+**Two fixes were tried against the panel and BOTH failed, which is itself the evidence that the
+panel is not the problem.** (a) Reconciling `StreamActionLog`'s log against the job's log FILE
+once the stream ends: 2/10, unchanged. (b) Reconciling from the 1 s status poll on a terminal
+job status, and on either finish route: 1/10 then 2/10 — noise. Neither could find anything the
+panel did not already have, because the file does not have it either. Both were reverted; the
+component is untouched. If someone picks this up: the surviving suspects are inside
+`runOperationBatch` between `runOperationUnit` returning and `runPool` draining —
+`concurrentRunner.ts`'s documented "a transiently-zero limit ... does NOT terminate" wait at a
+zero `laneLimit`, and `taskHooks.ts`'s `onLog` progress parsers, which run on every line the
+unit emits. Instrument the job's registry record, not the log panel.
+
## Not changed, deliberately
- `reconcileVideoDirs` keeps mutating during snapshot generation; that is the product's